Architecture
Llama-style decoder, 24 layers, hidden 1,152, FFN 4,608 (SwiGLU), attention heads / KV heads 9 / 3 (head dim 128), RMSNorm, RoPE base 10,000, untied embeddings, no biases, no QK-norm, sequence length 4,096. Tokenizer: SmolLM2 vocabulary (49,152) plus <|pad|>; <|endoftext|> is the end-of-document token.
Pretraining
Data: the 1PP conversations rewritten from those documents; loss on user and assistant turns (no loss on <|endoftext|>). One pass over 47.8M documents (66.2B tokens of original documents; 63.0B tokens as conversations), 31,777 steps at global batch 512 × 4,096 tokens, cross-document attention masking, best-fit packing with step-aligned document assignment. Optimizer: Muon (shape scaling, matrix LR 0.005) with Adam for embeddings and norms, warmup 2,000 steps, constant, linear decay over the last 10% to 1/100, weight decay 0.1, bf16.
Validation loss (per token, 2,433 held-out documents, final checkpoint):
Table with columns: assistant text, user text, document text| assistant text | user text | document text |
|---|
| 1.572 | 1.462 | 3.298 |
ChatML without a system turn (the models never saw one):
<|im_start|>user\n{message}<|im_end|>\n<|im_start|>assistant\n{reply}<|im_end|>\n
The bundled chat_template renders exactly this. Generation stops at <|im_end|> (id 2) or <|endoftext|> (id 0); both are listed in eos_token_id. This is a base model; the conversation conditions produce chat-formatted text, the raw baseline plain text.
Verification
The HF weights were checked against the Megatron checkpoint by recomputing validation losses with this model:
Table with columns: set, HF loss, Megatron reference, abs. diff| set | HF loss | Megatron reference | abs. diff |
|---|
| val50m segments [3] | 1.5702 | 1.5717 | 0.0015 |
| raw_val50m segments [8] | 3.3005 | 3.2983 | 0.0022 |
Links
- Training logs: wandb projects 1pp-training and 1pp-sft
- Research artifact from the 1PP project (EPFL DLAB); not a general-purpose assistant.