Model Details
Architecture
Table with columns: Parameter, Value| Parameter | Value |
|---|
| Architecture | LLaMA (decoder-only) with GQA |
| Parameters | 88,945,152 |
| Hidden size | 768 |
| Intermediate size | 2,048 |
| Num layers | 8 |
| Num attention heads | 12 |
| Num KV heads | 4 (GQA) |
| Max sequence length | 1,024 |
| Vocabulary size | 50,257 (GPT-2) |
| Position encoding | RoPE |
| Activation | SwiGLU |
| Normalization | Pre-RMSNorm |
| Weight tying | Yes (embedding ↔ lm_head) |
| Precision | fp32 |
Training Data
Table with columns: Source, Type, Size| Source | Type | Size |
|---|
| FineWeb-Edu (sample-10BT) | Educational web text | 150,000 examples |
Training Hyperparameters
Table with columns: Parameter, Value| Parameter | Value |
|---|
| Hardware | 15GB VRAM |
| Optimizer | AdamW |
| Learning rate | 2e-4 |
| Batch size | 8 |
| Gradient accumulation | 4 |
| Effective batch size | 32 |
| Warmup steps | 500 |
| Weight decay | 0.01 |
| Training time | ~45 minutes |
Tokenizer
- Type: GPT-2 (ByteLevel BPE)
- Vocabulary: 50,257 tokens
- Pre-trained: GPT-2 tokenizer (used directly, not retrained)
Usage
from transformers import pipeline
pipe = pipeline("text-generation", model="pinkelephantlimited/pink-elephant-90m")
output = pipe("5 tips for better sleep:", max_new_tokens=100)[0]["generated_text"]
print(output)
Training Philosophy
This model was trained entirely from scratch on permissively licensed data. No fine-tuning, no transfer learning — 100% original weights.
License
MIT
Real-World Applications
📱 Mobile Content Drafting (GQA = Fast Inference)
Grouped-Query Attention makes this efficient on mobile GPUs:
from transformers import pipeline
pipe = pipeline("text-generation", model="pinkelephantlimited/pink-elephant-90m")
pipe("5 tips for better sleep: 1. Stick to a schedule", max_new_tokens=100)
pipe("Breaking: Scientists have discovered", max_new_tokens=80)
📄 Blog/Article Generation Assistant
pipe("The rise of remote work has fundamentally changed", max_new_tokens=100)
from transformers import AutoModel
model = AutoModel.from_pretrained("pinkelephantlimited/pink-elephant-90m")
Why 90M? GQA provides 3x faster inference than standard attention. Balances quality and speed — runs on phone GPUs and cloud instances.