Honest evaluation (base vs +LoRA, thinking OFF, 3-run avg on local batteries)
Table with columns: axis, base, +LoRA, note| axis | base | +LoRA | note |
|---|
| Reasoning (/18) | ~16 | ~12 | ↓ tradeoff |
| IFEval (/10) | ~8 | ~5 | ↓ tradeoff |
| Code (/7) | 7 | 7 | preserved |
| Agentic tool-use (/10) | ~6 | ~5 | ≈ preserved |
| General / Ortho (/23) | ~12 | ~13 | ↑ (customer-service, translation) |
| Russian math battery | low | low | unchanged (1B math-capped) |
What it does well: practical Russian gets noticeably more fluent and coherent — customer-support replies stop mixing English/Chinese and become natural Russian; RU→EN translation is more accurate.
The tradeoff: sharp reasoning and strict instruction-following drop somewhat — an inherent capacity limit of a 1B model (you can't add a language without spending some of its other skills). For strong Russian without the tradeoff, consider a Russian-native model or a larger base.
Use thinking OFF (enable_thinking=false) — the adapter was trained for the non-thinking format.
Usage
Transformers (PEFT):
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("openbmb/MiniCPM5-1B", trust_remote_code=True)
model = PeftModel.from_pretrained(base, "vahpetr/MiniCPM5-1B-ru-lora-v3")
tok = AutoTokenizer.from_pretrained("openbmb/MiniCPM5-1B", trust_remote_code=True)
llama.cpp (GGUF adapter MiniCPM5-1B-ru-lora-v3.gguf is included):
llama-server -m MiniCPM5-1B-Q8_0.gguf --lora MiniCPM5-1B-ru-lora-v3.gguf --jinja
(A merged, ready-to-run version is at vahpetr/MiniCPM5-1B-ru-v3.)
Attribution & license
- Base: MiniCPM5-1B (OpenBMB) — Apache-2.0
- Data: GrandMaster-PRO-MAX (Vikhr) — Apache-2.0
- This adapter — Apache-2.0