Model Details
Table with columns: Attribute, Value| Attribute | Value |
|---|
| Architecture | LlamaForCausalLM |
| Parameters | 62.8M |
| Hidden size | 448 |
| Layers | 14 |
| Attention heads | 8 (GQA, 8 KV heads) |
| Head dim | 56 |
| Context length | 4096 (YaRN extended, factor 4.0) |
| Vocab size | 32,000 |
| Precision | bfloat16 (~125 MB) |
| License | Apache 2.0 |
Training Configuration
Table with columns: Hyperparameter, Value| Hyperparameter | Value |
|---|
| Framework | TRL SFTTrainer + PEFT LoRA |
| LoRA rank | r=32, α=64 (all linear layers) |
| Precision | fp16, torch.compile enabled |
| Batch | 4 per GPU, gradient accumulation 1 |
| Effective batch | 8 (2× T4 DDP) |
| Learning rate | 2e-4 cosine, 5% warmup |
| Max seq length | 4096 |
|
Training Results
Table with columns: Metric, Value| Metric | Value |
|---|
| Best eval loss | 7.8651 (step 1100) |
| Final train loss | 7.7178 |
| Total steps | 1,100 |
| Tokens processed | 35.7M |
| Dataset | 35,944 train / 734 eval |
| Samples/sec | ~3.93 |
Loss Curves

The model shows consistent convergence across 1,100 steps. Train loss drops from 10.47 → 7.72 (26.3% reduction), eval loss from 10.43 → 7.87 (24.6% reduction). No overfitting observed — train and eval curves track closely.
Learning Rate Schedule

Cosine schedule with 5% warmup (55 steps). Peak LR 2e-4 reached at step 900, then cosine decay begins. The steady increase during warmup allows the LoRA adapters to initialize gracefully before full learning kicks in.
Gradient Norm

Grad norm stabilizes after ~400 steps. Initial spike at step 400-450 (norm 5.4) is typical for LoRA warmup as adapters find their direction. Settles to 1.5-2.5 range for remainder of training.
Loss Progression Table
Table with columns: Step, Train Loss, Eval Loss, Δ Eval| Step | Train Loss | Eval Loss | Δ Eval |
|---|
| 50 | 10.43 | 10.43 | — |
| 100 | 10.15 | 10.10 | -0.33 |
| 200 | 9.23 | 9.26 | -0.84 |
| 300 | 9.06 | 9.00 | -0.26 |
| 400 |
Quick Start
Install Dependencies
pip install -r requirements.txt
Interactive Chat
This starts an interactive chat session. Type your messages and get responses from Lumia 62M.
Single Prompt
python generate.py --prompt "Write a Python function to check if a number is prime"
Python API
from transformers import AutoModelForCausalLM, AutoTokenizer model = AutoModelForCausalLM.from_pretrained("samcheng0/lumia-62m")tokenizer = AutoTokenizer.from_pretrained("samcheng0/lumia-62m") prompt = """<|system|>You are an expert programmer. Think step by step.<|user|>Write a Python function to check if a number is prime.<|assistant|>""" inputs = tokenizer(prompt, return_tensors="pt")outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)print(response)
Evaluation
python eval.py # Run all benchmarkspython eval.py --category math # Run specific categorypython eval.py --verbose # Show full responsespython eval.py --save results.json # Save results to file
Load LoRA Adapter (Continued Training)
from peft import PeftModel base = AutoModelForCausalLM.from_pretrained("samcheng0/lumia-62m")model = PeftModel.from_pretrained(base, "samcheng0/lumia-62m/adapter")
The model supports a chat template with special tokens:
from transformers import AutoModelForCausalLM, AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("samcheng0/lumia-62m")model = AutoModelForCausalLM.from_pretrained("samcheng0/lumia-62m") messages = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is 2+2?"},] # Apply chat templateprompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
Supported Tokens
Table with columns: Token, ID, Purpose| Token | ID | Purpose |
|---|
<|system|> | 32010 | System prompt |
<|user|> | 32011 | User input |
<|assistant|> | 32012 | Model response |
<think> | 32008 | Start reasoning block |
|
Note: All 20 special tokens are single-token IDs. The tokenizer handles them natively for efficient encoding/decoding.
Generation Parameters
Table with columns: Parameter, Default, Description| Parameter | Default | Description |
|---|
temperature | 0.7 | Controls randomness (lower = more deterministic) |
top_p | 0.9 | Nucleus sampling threshold |
max_new_tokens | 512 | Maximum tokens to generate |
repetition_penalty | 1.1 | Penalizes repeated tokens |
Benchmarks
The model was evaluated on 20 test prompts across 5 categories:
Table with columns: Category, Prompts, Description| Category | Prompts | Description |
|---|
| Math | 4 | Arithmetic, algebra, calculus |
| Code | 4 | Python functions, complexity analysis |
| Reasoning | 4 | Logic puzzles, pattern recognition |
| General | 4 | Knowledge, facts, explanations |
| Indonesian | 4 | Translation, comprehension |
Run the full benchmark suite:
Dataset
Fine-tuned on samcheng0/lumia-reasoning-sft-v1 — 35,944 train + 734 eval samples.
Data Sources (17 datasets)
Table with columns: Source, Type, Samples| Source | Type | Samples |
|---|
| TeichAI/claude-4.5-opus-high-reasoning-250x | Reasoning traces | ~2.5K |
| TeichAI/Claude-Opus-4.6-Reasoning-887x | Reasoning traces | ~1.8K |
| nohurry/Opus-4.6-Reasoning-3000x-filtered | Reasoning traces | ~2.1K |
| angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k | Code reasoning | ~3.5K |
| Crownelius/Opus-4.6-Reasoning-3300x | Reasoning traces | ~3K |
| nvidia/OpenCodeReasoning |
Filter Pipeline
Raw: ~202K lines → Filtered: ~36K (81.6% filtered out)
Table with columns: Filter, Threshold| Filter | Threshold |
|---|
| Min total chars | 3,000 |
| Min output chars | 1,500 |
| Output/input ratio | ≥ 1.2 |
| Structural score | ≥ 4 (=+3, code block=+2, steps=+2) |
| Dedup | MD5 hash |
Repo Structure
lumia-62m/├── config.json # Model architecture├── model.safetensors # Merged weights (inference ready)├── tokenizer.json # Tokenizer (with special tokens)├── tokenizer_config.json # Tokenizer settings + chat template├── special_tokens_map.json # Special tokens ID mapping├── README.md # This file├── requirements.txt # Python dependencies├── generate.py # Interactive inference script├── eval.py # Evaluation benchmark├── add_special_tokens.py # Token management script├── banner.svg # Header banner├── loss_curve.svg # Training loss chart├── lr_schedule.svg # Learning rate chart├── grad_norm.svg # Gradient norm chart└── adapter/ # LoRA adapter + training state ├── adapter_model.safetensors # LoRA weights (14.7 MB) ├── adapter_config.json # PEFT config ├── optimizer.pt # AdamW state (resume training) ├── scheduler.pt # LR scheduler state ├── scaler.pt # Gradient scaler ├── trainer_state.json # Full training metrics └── train.log # Training log
Citation
@misc{lumia-62m, title={Lumia 62M: A Small Reasoning Language Model}, author={samcheng0}, year={2026}, howpublished={\url{https://huggingface.co/samcheng0/lumia-62m}},}
License
Apache 2.0