Training Setup
Table with columns: Field, Value| Field | Value |
|---|
| Base model | SupraLabs/Supra-1.5-50M-Base-exp |
| Base revision | main |
| Output repo | User01110/testing-50M |
| Sequence length | 1024 |
| Max optimizer steps | 10,000 |
| Per-device batch size | 128 |
| Gradient accumulation | 4 |
| Sample presentations per GPU | 5,120,000 |
| Max token slots per GPU | 5,242,880,000 |
| Learning rate | 2.00e-04 |
| Warmup steps | 100 |
| Weight decay | 0.05 |
| Save/push cadence | every 1,000 optimizer steps plus final |
| Loss masking | assistant-span-only from step 0 |
| Loss logging | printed loss is normalized by gradient accumulation; raw_sum is the Trainer sum over 4 microbatches |
| Gate logging | novelty score if the loaded architecture exposes last_gate; otherwise n/a |
| Prompt format | ChatML |
| System prompt | You are a helpful assistant. |
The stream randomly mixes the selected instruction, math, and coding sources. Sources are reopened after exhaustion and keep relooping until the 10,000-step training cap finishes, except Cutecat6152/python-data-basic, which is capped at 3 passes.
Listed source rows before relooping: 3,718,915. The 10,000-step training budget presents 5,120,000 examples per GPU.
Prompt Template Compatibility
The uploaded tokenizer includes the ChatML special tokens and chat template, so inference and future SFT should not require manually adding <|im_start|> or <|im_end|>.
ChatML messages are rendered as:
<|im_start|>systemYou are a helpful assistant.<|im_end|><|im_start|>user{ user_message }<|im_end|><|im_start|>assistant
This script starts from the base checkpoint, adds <|im_start|> and <|im_end|> once as tokenizer special tokens, resizes embeddings once, saves the tokenizer with chat_template, disables automatic post-processing during pretokenized SFT, and keeps/saves the model context config with max_position_embeddings >= 1024.
The base model is loaded with pinned revision main so Transformers will not silently fetch a newer remote modeling file during training.
Complete inference example:
from transformers import AutoModelForCausalLM, AutoTokenizerimport torch repo = "User01110/testing-50M"tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)model = AutoModelForCausalLM.from_pretrained( repo, trust_remote_code=True, torch_dtype="auto", device_map="auto",) messages = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Explain what a neural network is in simple terms."},]prompt = tokenizer.apply_chat_template( messages, tokenize=False, add_generation_prompt=True,)inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device) with torch.no_grad(): output = model.generate( **inputs, max_new_tokens=256, do_sample=False, temperature=0.7, top_k=40, top_p=0.95, repetition_penalty=1.2, pad_token_id=tokenizer.pad_token_id, eos_token_id=tokenizer.eos_token_id, ) new_tokens = output[0, inputs["input_ids"].shape[-1]:]text = tokenizer.decode(new_tokens, skip_special_tokens=True).strip()print(text)
Dataset Mix
Table with columns: Dataset, Config, Split, Rows, Schema, Mapping, Pass policy| Dataset | Config | Split | Rows | Schema | Mapping | Pass policy |
|---|
| nvidia/Nemotron-SFT-Instruction-Following-Chat-v2 | default | reasoning_off | 1,068,273 | messages[{role, content, reasoning_content}] | user/assistant message pairs; reasoning_off only | reloops until max_steps |
| microsoft/orca-math-word-problems-200k | default | train | 200,035 | question, answer |
Notes
- Dataset schemas and row counts were checked through Hugging Face Dataset Viewer metadata where available.
- Multiturn/message datasets carry all assistant spans into the collator, so user/system text remains masked from step 0 while every assistant turn is supervised.
- Streaming source open/read failures are retried and reopened. Normal stream exhaustion reopens that source and continues mixing it until
max_steps; python-data-basic is dropped after 3 completed passes.
- RoPE buffers and tokenizer/model load are verified during final export.