Method
Adaptation of CM-Align (Zhang et al., EMNLP 2025 Findings, arXiv:2509.08541)
to a parallel multilingual factual-QA setting. The procedure is self-supervised — it never uses
gold answer labels:
- Preference construction. For each fact, sample
K=4 free-text answers per language
(temperature 0.9, top-p 0.95). Embed all candidates with
sentence-transformers/LaBSE.
Choose as pivot the most self-consistent English candidate (highest mean cosine
similarity to the other English candidates). Then, for every other language, take
chosen = argmax and rejected = argmin cosine similarity to that English pivot.
- DPO training. Hand-rolled DPO on those preference pairs. The reference distribution is
the same model with the LoRA adapter disabled, so no second copy of the model is held in
memory. Objective:
L_DPO + gamma * L_NLL.
CM-Align's original embedder was gte-multilingual-base; LaBSE is used here because it is the
cross-lingual encoder used throughout this project.
Training details
Table | |
|---|
| Base model | allenai/OLMo-2-1124-7B |
| Data | 40,000 facts from jvonrad/WIKI-FACT (train) |
| Languages | 12 — en, de, id, pt, ar, bn, sw, es, ru, fr, ja, zh |
| Candidates / language | 4 (temperature 0.9, top-p 0.95) |
| Embedder | sentence-transformers/LaBSE |
| DPO beta | 0.1 |
|
Evaluation
Answers are scored by length-normalised log-likelihood over the four options, with the plain
prompt Question: {question}\nAnswer:. Metrics beyond accuracy:
- TotCons (Total Consistency) — fraction of facts answered correctly in all languages.
- RankC — cross-lingual agreement between full option rankings (Qi et al., EMNLP 2023).
- AnsAgr — pairwise answer agreement across language pairs.
PolyFact test (2,523 facts, 12 languages):
Table with columns: Model, Acc, TotCons, RankC, AnsAgr| Model | Acc | TotCons | RankC | AnsAgr |
|---|
| OLMo-2-7B (base) | 57.96 | 7.21 | 58.32 | 51.10 |
| OLMo-2-7B CM-Align | 60.44 | 9.35 | 60.07 | 53.41 |
Global-MMLU-Lite (400 facts, 11 languages — Lite has no Russian config):
Table with columns: Model, Acc, TotCons, RankC, AnsAgr| Model | Acc | TotCons | RankC | AnsAgr |
|---|
| OLMo-2-7B (base) | 44.75 | 3.50 | 55.12 | 46.48 |
| OLMo-2-7B CM-Align | 42.93 | 2.50 | 57.31 | 48.45 |
Per-language accuracy on PolyFact test:
Table with columns: en, de, id, pt, ar, bn, sw, es, ru, fr, ja, zh| en | de | id | pt | ar | bn | sw | es | ru | fr | ja | zh |
|---|
| 80.6 | 71.3 | 68.8 | 69.6 | 45.1 | 42.3 | 47.6 | 71.9 | 55.6 |
Reading these numbers
CM-Align gives consistent in-domain gains over the base model on PolyFact (all four metrics).
Out of domain on Global-MMLU-Lite the picture is mixed: cross-lingual agreement improves
(RankC +2.19, AnsAgr +1.97) but accuracy and total consistency drop slightly. This model is
published as a baseline for comparison, not as a recommended general-purpose checkpoint.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "jvonrad/OLMo-2-7B-CM-Align"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
prompt = "Question: What is the capital of Poland?\nAnswer:"
out = model.generate(**tok(prompt, return_tensors="pt").to(model.device), max_new_tokens=16)
print(tok.decode(out[0], skip_special_tokens=True))
This is a base (non-instruct) model — it has no chat template. Use plain completion-style
prompts such as the Question: ... \nAnswer: form above.
Limitations
- Trained only on Wikidata-derived factual QA in 12 languages; it is not a general instruction-following model.
- Preference pairs are built from the model's own samples via embedding similarity, so they
inherit LaBSE's similarity biases and can reward fluent-but-wrong answers.
- Improved cross-lingual consistency can make incorrect factual associations more uniform
across languages as well as correct ones.
- Evaluation is multiple-choice log-likelihood scoring; it does not measure free-form generation quality.
Citation
The CM-Align method this baseline implements:
@inproceedings{zhang2025cmalign,
title = {CM-Align: Consistency-based Multilingual Alignment for Large Language Models},
author = {Zhang, Xue and others},
booktitle = {Findings of EMNLP},
year = {2025}
}
The RankC consistency metric used above:
@inproceedings{qi2023crosslingual,
title = {Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models},
author = {Qi, Jirui and Fern{\'a}ndez, Raquel and Bisazza, Arianna},
booktitle = {EMNLP},
year = {2023}
}