Catalan evaluation (Gala vs base Qwen3.5-0.8B)
Generation-based evaluation, letter-parsing for multiple choice, sampled decoding (temp 0.7, top_p 0.9, repetition penalty 1.1). Full JSON: private luispoveda93/Gala-project-report/eval_all_v4.json.
Table with columns: Benchmark, Gala, Base, Δ| Benchmark | Gala | Base | Δ |
|---|
| IFEval_ca (strict acc, 150-prompt verifier subset) | 41.3% | 32.7% | +8.7 |
| hhh_alignment_ca (overall) | 15.8% | 10.4% | +5.4 |
| tecla (4-way news classification, n=300) | 44.3% | 25.3% | +19.0 |
| COPA-ca (n=500) | 64.0% | 54.0% | +10.0 |
| arc_ca Easy (n=300) | 64.7% | 59.3% | +5.3 |
| arc_ca Challenge (n=300) | 55.7% | 54.0% | +1.7 |
| mgsm_ca (n=250) | 23.6% | 22.4% | +1.2 |
| EQ-bench_ca (MAE 0–10, lower better, 160/167 parsed) | 2.96 | 3.35 | −0.39 |
| CaBBQ (overall, n=400) | 41.5% | 45.0% | −3.5 |
| xquad-ca (F1, n=300) | 37.9 | 56.1 | −18.2 |
Summary: Gala beats the base model on both core-gate metrics (instruction following and alignment) and on 7 of 10 reported measurements, with the largest gains in Catalan cultural knowledge (Tecla), commonsense (COPA) and instruction following (IFEval_ca).
Limitations (honest)
- Fluency ceiling: long-form Catalan output still contains grammatical errors, invented words and hallucinated content — a limitation of the 0.8B scale and the SFT-only recipe, not fixable by more of the same data. Short conversational turns are noticeably more reliable than long generations.
- Extractive QA regression: on XQuAD-ca the base model's F1 is higher; Gala's chattier style hurts span-extraction tasks.
- CaBBQ is slightly lower than base; MC bias-probe agreement at 0.8B is near chance for both models and should be read cautiously.
- Both models were evaluated with the same harness; small-model MC letter-parsing adds noise to all absolute numbers.
- Multimodal (vision) behaviour is inherited from the base model and was not fine-tuned or evaluated.
Intended use
Catalan-first conversational assistant for experimentation and demos. Not recommended for production factual use without guardrails.
Training data licenses
ALIA-2606-SFT (CC-BY-4.0) · MentorCA (CC-BY-SA-4.0) · English anchor from ALIA-2606-SFT. InstruCAT (CC-BY-NC-ND) was deliberately excluded.
Decision-reasoning report: private repo luispoveda93/Gala-project-report · Training dashboard: https://huggingface.co/spaces/luispoveda93/gala-catalan-sft-trackio