Model Details
Table | |
|---|
| Architecture | GPT-2 (causal decoder-only transformer) |
| Parameters | ~98.4M |
| Language | Hindi (hi) |
| License | MIT |
| Trained from | Scratch (random initialization, no pretrained checkpoint) |
| Context length | 1,024 tokens (n_positions / n_ctx) |
| Layers / heads / hidden size | 12 layers / 12 heads / 768 hidden — identical body to standard GPT-2-small |
| Tokenizer | Custom tokenizer trained on the Hindi corpus · vocab size 16,384 · special tokens <s> (bos, id 1), </s> (eos, id 2), <pad>, <unk> |
Confirmed against config.json: 98.4M parameters = embeddings (16,384×768 + 1,024×768 ≈ 13.4M, tied input/output via tie_word_embeddings: true) + 12 transformer blocks (≈85.0M). The transformer body (12 layers / 12 heads / 768 dim) is identical to standard GPT-2-small (124M) — the entire size difference comes from the 16,384-token Hindi vocabulary vs. English GPT-2's 50,257 tokens. Activation: gelu_new; dropout 0.1 (attention/embedding/residual) — standard GPT-2-small defaults.
Training Data
Trained on pulipakav-1/translated-babylm-hindi: a Hindi translation of the English BabyLM 2026 (Strict track) corpus, produced with IndicTrans2 (ai4bharat/indictrans2-en-indic-1B, arXiv:2305.16307).
Table with columns: Split, Source, Sentences, Words| Split | Source | Sentences | Words |
|---|
| Train | BabyLM-2026-Strict | 11,579,880 | 118,309,059 |
| Val | BabyLM-dev | 1,153,113 | 11,882,662 |
| Test | BabyLM-Test | 1,097,453 | 11,106,875 |
Train split is composed of six domains (bnc_spoken, childes, gutenberg, open_subtitles, simple_wiki, switchboard), mirroring the original English BabyLM mixture. Translation quality varies by domain — high for Simple Wikipedia and Gutenberg (formal, well-structured text), moderate for Open Subtitles and BNC Spoken (informal/conversational), and variable for CHILDES (child-directed speech, non-standard syntax that translates less reliably).
Training Procedure
Table | |
|---|
| Framework | 🤗 Transformers 5.12.0 (GPT2LMHeadModel) |
| Optimizer | AdamW (β₁=0.9, β₂=0.999, ε=1e-8); no weight decay on LayerNorm/bias params |
| Learning rate / schedule | 5e-5 peak · cosine decay · linear warmup (1% of total steps) |
| Batch size (effective) | 16 sequences × 512 tokens = 8,192 tokens/step (4 per-GPU × 4 GPUs) |
| Steps / epochs | 10 epochs over the train split (~100M words/epoch) ≈ 122K total steps |
| Precision | bfloat16 mixed precision during training (via 🤗 Accelerate); checkpoint saved as float32 |
|
Note: training used a 512-token sequence length, shorter than the model's 1,024-token max context (n_positions/n_ctx). Flagging this in case it's worth confirming — if you intended to train at the full 1,024-token context, the table above should be updated.
Evaluation
MultiBLiMP (Hindi subject-verb agreement)
Evaluated on the official Hindi subset of MultiBLiMP 1.0 (Jumelet et al., 2026, TACL) — 1,447 minimal pairs testing subject-verb/subject-participle agreement for number, person, and gender.
Table with columns: Metric, Value| Metric | Value |
|---|
| Accuracy | 0.9261 (92.61%) |
| Pairs evaluated | 1,447 |
[ADD ADDITIONAL EVALUATION RESULTS HERE — e.g. BabyLM-Hindi held-out perplexity, other BLiMP-style tasks, downstream tasks]
Intended Uses
- Research on data-efficient / from-scratch language model training for Hindi (BabyLM-style)
- Studying the effect of translation-based data augmentation on a low(er)-resource language's acquired grammatical competence
- Benchmarking against the English BabyLM baseline and other from-scratch Indic-language models trained under the same protocol
Limitations
- Machine-translated training data: all Hindi text was produced by IndicTrans2, not written natively — it may contain translation artifacts, mistranslations, or unnatural phrasing, especially in the informal/conversational and child-directed-speech portions of the corpus (CHILDES, Switchboard, BNC Spoken). The dataset does not represent naturally occurring Hindi.
- Small training scale: trained on a ~100M-word corpus by design (BabyLM strict-track scale) — not competitive with web-scale pretrained Hindi LLMs for general use; intended for controlled research comparisons, not production deployment.
How to Use
from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("pulipakav-1/gpt-2-hindi2")model = AutoModelForCausalLM.from_pretrained("pulipakav-1/gpt-2-hindi2") inputs = tokenizer("आज मौसम", return_tensors="pt")outputs = model.generate(**inputs, max_new_tokens=50)print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Citation
If you use this model, please cite it as:
@misc{gpt2-hindi2, author = {pulipakav-1}, title = {GPT-2 Hindi (BabyLM, trained from scratch)}, year = {2026}, publisher = {Hugging Face}, url = {https://huggingface.co/pulipakav-1/gpt-2-hindi2}}
Please also cite the underlying dataset and tools:
@misc{translated-babylm-hindi, author = {pulipakav-1}, title = {Translated BabyLM — Hindi}, year = {2026}, publisher = {Hugging Face}, url = {https://huggingface.co/datasets/pulipakav-1/translated-babylm-hindi}}
License
MIT — see repository for details. Note the underlying dataset is licensed CC-BY-4.0; check its terms separately if redistributing training data alongside this model.