Model Details
Architecture
Table with columns: Parameter, Value| Parameter | Value |
|---|
| Architecture | LLaMA (decoder-only) |
| Parameters | 32,514,560 |
| Hidden size | 512 |
| Intermediate size | 1,792 |
| Num layers | 8 |
| Num attention heads | 8 |
| Num KV heads | 8 |
| Max sequence length | 2,048 |
| Vocabulary size | 4,096 |
| Position encoding | RoPE (θ=10000.0) |
| Activation | SwiGLU |
| Normalization | Pre-RMSNorm |
| Weight tying | Yes (embedding ↔ lm_head) |
| Precision | fp32 |
Training Data
Table with columns: Source, Type, Size| Source | Type | Size |
|---|
| FineWeb-Edu (sample-10BT) | Educational web text | 150,000 examples |
Training Hyperparameters
Table with columns: Parameter, Value| Parameter | Value |
|---|
| Hardware | 96GB GPU |
| Optimizer | AdamW (8-bit) |
| Learning rate | 2e-4 |
| Batch size | 32 |
| Gradient accumulation | 4 |
| Effective batch size | 128 |
| Warmup steps | 500 |
| Weight decay | 0.01 |
| Training time | ~30 minutes |
Tokenizer
- Type: BPE
- Vocabulary: 4,096 tokens
- Trained from scratch on the training corpus
Usage
from transformers import pipeline
pipe = pipeline("text-generation", model="pinkelephantlimited/pink-elephant-33m")
output = pipe("The water cycle consists of", max_new_tokens=80)[0]["generated_text"]
print(output)
Training Philosophy
This model was trained entirely from scratch on permissively licensed data. No fine-tuning, no transfer learning — 100% original weights.
License
MIT
Real-World Applications
🎓 Educational Q&A Chatbot (Low-Power Devices)
Deploy on school tablets, offline kiosks, or mobile apps:
from transformers import pipeline
pipe = pipeline("text-generation", model="pinkelephantlimited/pink-elephant-33m")
pipe("The water cycle consists of three main stages:", max_new_tokens=80)
pipe("Photosynthesis is the process by which", max_new_tokens=80)
✏️ Text Autocomplete for Writing Assistants
pipe("In conclusion, the key factors that led to", max_new_tokens=60)
🗣️ Simple Chatbot on CPU-Only Servers
pipe("Thank you for contacting support. To help you better,", max_new_tokens=80)
Why 33M? Runs on CPU at ~50ms/token, fits in 200MB RAM. Perfect for serverless/edge deployments where GPU is unavailable.