How it works
-
Prompt. The state (evidence), the question (criterion) and the lettered options go into a structured
prompt, following the SemIf format.
-
Readout. The model reads the next-token logits only over the option letters. A softmax with a
fixed temperature gives the distribution. Options are labelled A–Z and then AA, AB, … (each label a single
token): up to 160 options are read in the same forward pass (serve v1.2, max_one_pass in
decision_config.json), more through a tournament whose final round is one pass.
-
Calibration. Temperature T = 1 ships by default (see Calibration).
-
Training. Distillation from a strong teacher (GLM-5.3-Flash, maximum reasoning effort) on:
- generated and blindly labeled decision items;
- exact programmatic items;
- compositional rules;
- long-context dossiers (6k–12k tokens).
Training-only signals, all switched off at inference:
- soft cross-entropy to the teacher's probabilities, with option-order permutation;
- an auxiliary loss on the item's short rationale (written with the item, explaining the gold answer).
Eikos-27B is a single LoRA fine-tune of Qwen3.8-27B (rank 64, one epoch) on the same data as Eikos-4B, with
long-context dossiers up to 12k tokens and without the EN↔PT view objective.
Quick start
vLLM (production: batching + prefix cache)
Requires vLLM ≥ 0.30.0. Older builds return wrong answers when several long requests are batched together on
this hybrid (Gated DeltaNet) architecture; we measured drops of 3–6 points on long, shared-document items with
vLLM 0.11. With 0.30 the batched results match the PyTorch reference.
bash serve_vllm.sh <MODEL_DIR> 8001 # vLLM engine: letter readout + hybrid prefix cache (checks vLLM >= 0.30)
python serve.py --model <MODEL_DIR> --vllm-url http://127.0.0.1:8001 --port 8000 # TypeSafe-compatible API on :8000
serve_vllm.sh runs vllm serve with the flags the readout needs: --served-model-name decider,
--enable-prefix-caching --mamba-cache-mode align, --logprobs-mode processed_logprobs --max-logprobs 600,
--dtype bfloat16 and --max-model-len 16384.
It also sets --max-num-seqs 64 (change it with MAX_NUM_SEQS): each request in flight holds one hybrid
cache block, and vLLM's default of 1024 needs more blocks than a 27B build leaves free on one 80–96 GB GPU.
HTTP API (TypeSafe-compatible)
curl -s localhost:8000/v1/systemone -d '{
"state": "Order ticket #A-2231. Retail client, moderate risk. BUY 1,500 XYZ at market. Equity USD 48,000. Last price USD 41.20. Rule 4.2: a single order may not exceed 50% of equity without written supervisor approval. Approvals on file: none.",
"questions": {
"allowed": {"type": "noul", "instructions": "Under rule 4.2, can this order be executed as submitted?",
"criteria": {"true": "complies with rule 4.2", "false": "breaches rule 4.2"}},
"action": {"type": "choice", "instructions": "What should the desk do?",
"criteria": {"execute": "send as submitted", "request_approval": "hold and ask a supervisor",
"reduce_size": "cut the order to the allowed size", "reject": "refuse the order"}},
"risk": {"type": "score", "instructions": "How risky is this ticket for the firm?",
"criteria": ["low", "moderate", "high", "critical"]}
}}'
- One pass for the whole request: all questions are answered together, and the shared state is processed
once through the prefix cache.
- Response per question:
noul: probability (of "yes"), value and confidence;
choice: choice, probabilities and confidence;
score: score, probabilities, expected and .
Agent sessions (incremental state + cache)
POST /v1/sessions {"state": "..."} -> {"session_id"}
POST /v1/sessions/<id>/append {"text": "..."} -> grows the state
POST /v1/sessions/<id>/systemone {"questions": {...}} -> answers over the current state
DELETE /v1/sessions/<id>
On a hybrid (linear-attention) model the state is cached at the end of the context and forked for every
question. In our tests a 15-step agent session gave answers identical to sending the full request each time.
Local (Apple Silicon)
python mlx_decide.py <MODEL_DIR> # MLX backend: same readout and calibration as PyTorch
python local_demo.py <MODEL_DIR> mps # PyTorch MPS
Evaluation
All numbers are our own measurements with a single harness. We never trained on any evaluation set. Spanish,
one task family (trade-offs) and one topic (healthcare administration) were held out of training entirely.
Table with columns: Benchmark (our harness), 4B final, 27B final, Jev (reference)| Benchmark (our harness) | 4B final | 27B final | Jev (reference) |
|---|
| JevBench public — original / hard | 91.7 / 72.1 | 100.0 / 82.9 | 98.6 / 73.0 (official leaderboard) |
| JevBench hard — ECE (lower is better) | 0.049 | 0.051 | — |
| DecisionBench (OOD) — medium / hard | 77.1 / 66.9 | 88.4 / 78.5 | 89.1 / 69.3 |
| General battery (9 human-labeled tasks) | 76.0 | 82.5 |
* generated by the same rule generator family used in training (different rules and cases); see notes.


How to read the table:
- JevBench: public items only, not the official leaderboard run.
- Jev: measured through its API on our batteries, not distilled from.
- Laya (charts only): run with its official package; its JevBench scores are within one item of its official
leaderboard results.
- Charts: 7,140 items common to all four systems. The confidence chart shows, for every confidence threshold,
how many decisions a system would take on its own and how often those decisions are wrong.
- WCB (central-bank stance): reported as balanced accuracy, because the classes are imbalanced.
- Rules suites: "rules / new domain / rulebooks" come from the same generator family used in training, so
they measure learning of that generator. Transfer to rules the model never saw is measured by "Trade (unseen
rules)" (Incoterms®, documentary-credit presentation, EU VAT) and by RuleArena.
- "New domain" has two parts. 492 items come from a domain absent from training (insurance): 27B 95.1, 4B 91.5.
The other 108 items use a rule combination held out of training (an exception over a business-day window):
27B 91.7, 4B 90.7. That combination appears in only 6 of the ~6,000 rule rows used in training.
Long context
The case file is hidden inside unrelated text at the start, middle or end of the prompt (80 public JevBench
decisions per point; released checkpoints served with vLLM 0.30).

Release builds
Final evaluation of every published build, run on the exact files in these repos: 7 suites, 7,371 items, vLLM 0.30
with batching and prefix cache on (MLX builds: MLX's CUDA backend on Linux). Batched vLLM is not bit-for-bit
deterministic across runs (1–2 items per suite can change), so the bf16 rows differ slightly from the Evaluation
table above, which uses our reference harness.
"≥0.90" = share of decisions the model would take on its own at ≥90% confidence, and the real error rate among them.
Table with columns: Build, Size, JevBench orig / hard, DecisionBench med / hard, General, Finance, Unseen trade rules, ECE, ≥0.90: decides / error, Gate vs bf16| Build | Size | JevBench orig / hard | DecisionBench med / hard | General | Finance | Unseen trade rules | ECE | ≥0.90: decides / error | Gate vs bf16 |
|---|
| Eikos-27B | 55.6 GB | 100.0 / 82.0 | 89.1 / 78.2 | 82.6 | 85.4 | 87.4 | 0.043 | 44.1% / 2.5% | reference |
Release gate, fixed before looking at results: accuracy within 1 point of bf16, ECE within 0.01, and at least 97%
of answers unchanged.
Table with columns: Build, Accuracy (bf16), Same answer as bf16: all / confident (≥0.9), ECE (bf16), Changed answers on items where bf16 was unsure (<0.7), Gate| Build | Accuracy (bf16) | Same answer as bf16: all / confident (≥0.9) | ECE (bf16) | Changed answers on items where bf16 was unsure (<0.7) | Gate |
|---|
| Eikos-27B-FP8 | 83.4 (83.4) | 98.8% / 100.0% | 0.042 (0.043) | 99% | pass |
| Eikos-27B-INT4 | 83.3 (83.4) | 97.8% / 100.0% | 0.041 (0.043) | 96% | pass |
| Eikos-27B-MLX-4bit |
- Quantization cost: no build loses more than 2 points on any metric of any suite, and calibration (ECE) stays
within 0.01 of bf16 or better.
- Eikos-4B-INT4 keeps accuracy and calibration but changes 4.5% of answers, above the 3% the gate allows. Almost
all of them (94%) are on items where bf16 itself was unsure (confidence below 0.7); on decisions bf16 takes with
confidence ≥0.9, it gives the same answer on 100% of items. We publish it with this note.
- Mac builds (MLX): Eikos-4B-MLX-8bit passes the gate. Eikos-27B-MLX-4bit and Eikos-4B-MLX-4bit keep accuracy
and calibration but change 3.4% and 7.5% of answers. Most of those changes (92% and 93%) are on items where bf16
itself was unsure; on decisions bf16 takes with confidence ≥0.9, they agree on 99.9% and 99.8% of items. We
publish them with this note. The small JevBench-original split (72 items) moves by a few items between builds.
Calibration
- Why T = 1: our held-out calibration sets turned out easier than real hard data. Every temperature fitted
on them was below 1 and made the model over-confident on hard items. Absent a realistic calibration set, we
ship T = 1.
- Changing it:
calib.json holds the temperature, and callers can re-fit it on their own data.

Parallelism and speed
- Same state, many questions: the hybrid prefix cache processes the state once.
- 200 questions over a 3.3k-token state ran 21× faster with vLLM prefix caching than without it.
- Throughput: batched decisions reach ~51 decisions/s on a shared GPU, with results identical to PyTorch.
- Full GPU (RTX PRO 6000, vLLM, hybrid prefix cache), 200 questions over one state:
- 4B: 161 decisions/s on a 751-token state and 197 decisions/s on a 3.3k-token state;
- 27B: 44 decisions/s on a 3.3k-token state.
- Our own PyTorch cache path: 1,500 questions over one state in 8.7 s (4B).
- Apple M4 16 GB, MLX bf16: ~0.4–0.8 s per decision, and 3 questions over one state in ~1.1 s.
Limitations
- Single pass means no multi-step reasoning. Long chains of arithmetic across many rules are out of reach;
on RuleArena (NBA salary-cap rulebooks, ~25k tokens) the model does not discriminate. An optional
"verify" mode (a short reasoning budget before the letter) exists in the server; it is off by default and
not part of any reported number.
- General world knowledge is bounded by model size. Both sizes trail large frontier systems on
knowledge-heavy tasks such as MMLU-Pro.
- Monetary-policy stance (hawkish/dovish) is a known weakness of both sizes: balanced accuracy 38.4 (4B) and
44.5 (27B), against 58.6 for Jev.
- Context: trained on inputs up to 32k tokens (4B) or 12k (27B). The base supports longer inputs, but
accuracy beyond the trained range degrades gradually.
- Not advice: it is not legal, tax or investment advice. It applies the rules it is given and does not
know your jurisdiction's current law.
Training data and licenses
See NOTICE.
- Model and code: MIT for our contributions. The base model is Apache-2.0 (
LICENSE-Qwen).
- Third-party data, with attribution: FinEntity (ODC-BY 1.0), TAT-QA (CC BY 4.0), GSM8K train (MIT).
- Prompt format: SemIf (MIT).
- Trademark: Incoterms® is a trademark of the ICC. No ICC or regulator text was used.
Citation
@misc{eikos2026,
title = {Eikos: open, calibrated, single-pass typed-decision models for finance and trading},
author = {Caio Vicentino},
year = {2026},
url = {https://github.com/caiovicentino/eikos}
}