Model description
Table | |
|---|
| Parameters | 528,538,752 (~528.5M, tied embeddings) |
| Architecture | Llama-style decoder (12→24 layers, 768→1152 hidden vs. the 125M build), full multi-head attention (no GQA/MQA) |
| Layers / hidden / heads | 24 / 1152 / 18 (head dim 64) |
| MLP | SwiGLU, intermediate size 4,608 |
| Positional encoding | RoPE, θ = 10,000 |
| Normalization | RMSNorm, ε = 1e-5 |
| Vocabulary | 16,384, byte-level BPE — the same tokenizer as the 125M build, reused deliberately rather than retrained |
| Context length | 1,024 tokens |
| Embeddings | tied input/output |
| Precision | bf16 compute (autocast) · fp32 master weights and saved checkpoint |
Training data
Five sources, ~14.7B unique training tokens after cleaning, deduplication, and
decontamination against CaseHOLD/LexGLUE:
Table with columns: Source, HF dataset, License| Source | HF dataset | License |
|---|
| US case law | HFforLegal/case-law | CC BY 4.0 |
| SEC 10-K filings | PleIAs/SEC | — |
| Educational web | HuggingFaceFW/fineweb-edu | ODC-BY |
| SEC 10-K/10-Q/8-K filings (bulk) | TeraflopAI/SEC-EDGAR | Apache-2.0 |
|
The mix isn't static. A flat/proportional mix of these five sources was tried
first and measurably diluted the original case-law/SEC/fineweb-edu "identity"
of the 125M build down to ~13% of training data, causing a real regression on
general text (fineweb-edu perplexity got worse than the 125M model, 22.82 vs
20.73). The fix — adopted from precedent in MiniCPM's and SmolLM2's published
training recipes — is a staged (Warmup-Stable-Decay) schedule: broad,
balanced per-source sampling weights for the first 85% of training steps, then
a decay phase for the last 15% that upweights the curated case-law/SEC/fineweb-edu
core specifically. This is what actually shipped in this checkpoint.
Evaluation
Full held-out perplexity, per source (not just one blended number, since the
five sources are genuinely different registers of text):
Table with columns: Source, Perplexity, vs. 125M model| Source | Perplexity | vs. 125M model |
|---|
| case-law | 7.23 | 8.22 (better) |
| SEC 10-K | 3.96 | 4.36 (better) |
| Educational web | 17.81 | 20.73 (better) |
| SEC-EDGAR (bulk, 10-K/10-Q/8-K) | 3.84 | (new source, no 125M baseline) |
| Case law (bulk) | 8.23 | (new source, no 125M baseline) |
|
Honesty note on the aggregate number: the aggregate looks like a huge win
partly because one bulk source (SEC-EDGAR) is ~61% of the val set by window
count and is inherently easier (more repetitive, formulaic filing text) to
predict than case law or general web text. The per-source numbers above —
where this model beats the 125M model on all three shared domains — are
the trustworthy signal, not the aggregate alone.
Intended uses & limitations
Intended use: text continuation / completion in a legal and financial
register; research and education on small-model pretraining, data-mix design,
and staged training schedules.
Not intended for: question answering, instruction following, or any use
where factual accuracy matters. Like the 125M base model (and any base model
at this scale), it will fabricate case names, statute citations, dollar
figures, and other specifics that sound right but are not grounded in
anything. A 528.5M-parameter model can store, at most, roughly 2 bits of
knowledge per parameter (Allen-Zhu & Li, "Physics of Language Models: Knowledge
Capacity Scaling Laws") — about 132MB of
total compressible fact storage, shared across grammar, style, and every fact
it knows. This is a hard, measured ceiling, not a training defect. Do not use
its output as legal, financial, or factual advice, or present it as such to
end users without a human review step. For a model that at least attempts to
answer questions, see legal-slm-500m-sft
— and note that model has its own documented hallucination limitations too.
Citation
@misc{verma2026slm500mbase,
author = {Deependra Verma},
title = {slm-500m-base: A 528.5M-Parameter Legal/Financial Language Model Trained From Scratch},
year = {2026},
url = {https://huggingface.co/DeependraVerma/slm-500m-base},
note = {Code: https://github.com/DeependraVerma/legal-slm-125M}
}
Author
Deependra Verma — Generative AI Researcher / AI Engineer.
GitHub · Hugging Face