Model Details
- Architecture: Llama-based architecture.
- Parameters: 308M.
- Context Window: 1024 tokens.
- Tokenizer: Custom BPE Tokenizer (Vocabulary Size: 32,000).
- Training Framework: PyTorch & Transformers.
Data Curation (The "Quality over Quantity" Approach)
The v0.2 release adopts a data-centric training strategy, prioritizing corpus quality over dataset size.
- Raw Data (v0.1): 5GB of raw Turkish corpus.
- Curated Data (v0.2): 1.2GB of carefully filtered, high-quality Turkish text.
- Process: Approximately 75% of noisy, duplicated, and low-quality samples were removed to improve linguistic quality and training efficiency.
Key Improvements from v0.1
- Architecture Shift: Migrated from GPT-2 to a modern Llama-based architecture.
- Normalization: RMSNorm.
- Positional Encoding: RoPE (Rotary Positional Embeddings).
- Activation: SiLU.
- Precision: Trained using bfloat16 for efficient consumer GPU training.
Design Goal
The 308M model serves as the flagship base model of the AhiskaAI v0.2 family.
Its primary objectives are:
- Strong Turkish language modeling.
- Improved semantic understanding.
- Better contextual consistency.
- A research foundation for future instruction tuning and alignment.
Training Logs

The graph above demonstrates the training convergence of AhiskaAI-308m-Base-v0.2. The stable decline in loss indicates effective optimization throughout pretraining.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("AhiskaAI/AhiskaAI-308m-Base-v0.2")
tokenizer = AutoTokenizer.from_pretrained("AhiskaAI/AhiskaAI-308m-Base-v0.2")
text = "Türkiye Cumhuriyeti"
inputs = tokenizer(text, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Training Hardware
Trained on NVIDIA RTX 4050 6GB Laptop GPU.
Future Plans
- Instruction-tuned (IT) version.
- Preference alignment with DPO.
- Larger and more diverse Turkish datasets.
- Future AhiskaAI v0.3 model family.
About AhiskaAI
AhiskaAI is an independent open-source initiative dedicated to developing efficient Turkish Small Language Models trained completely from scratch.
Follow us on Hugging Face for updates and future releases.