Model
- Backbone: Qwen/Qwen3.5-9B with its vision encoder, stored as standard sharded safetensors.
- Joint schema head: a small transformer head that reads the backbone's final hidden states,
routes evidence from the state to each question, and scores all options of all questions jointly.
- Output: one logit per allowed option for each question. Apply a softmax per question to get
probabilities.
Files
Table with columns: File, Purpose| File | Purpose |
|---|
model-*.safetensors, model.safetensors.index.json, config.json, generation_config.json | Backbone, including the vision encoder |
joint_head.safetensors, joint_head_config.json | Joint schema head |
joint_schema_model.py | Record encoding, batching, the model, load_release_model, and systemone |
tokenizer.json, tokenizer_config.json, chat_template.jinja, processor_config.json | Tokenizer and image/video processor |
LICENSE | Apache-2.0 license |
Usage
Tested with torch 2.11 and transformers 5.10.2 on a single H200. Image and video inputs also
need pillow.
import sys
import torch
from huggingface_hub import snapshot_download
path = snapshot_download("Cloudflare/clef-flash")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model
model, processor = load_release_model(path, device="cuda")
record = {
"state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
"questions": {
"status": {
"type": "choice",
"instructions": "What is the invoice status?",
"criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."},
},
"large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
},
}
encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
logits = model(batch)[0]
for question, question_logits in zip(encoded.questions, logits):
probabilities = question_logits.float().softmax(-1).tolist()
print(question.question_id, dict(zip(question.option_ids, probabilities)))
Jev / SystemOne API
systemone takes a Jev/SystemOne POST /v1/systemone request body and returns the same response
body: model, answers keyed by question ID, and usage. A choice answer has choice,
confidence, and probabilities; a score answer has the expected score, confidence, legend,
and probabilities; a noul answer has the probability of true. instructions is optional, and
images and may be added to the request.
from joint_schema_model import systemone
response = systemone(model, processor, {
"model": "clef-flash",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
})
print(response["answers"])
Images and video
Add images (PIL images) or videos (frame arrays) to the record and pass the processor to
encode_record. Optional processor arguments go in media_kwargs.
from PIL import Image
record = {
"state": {"task": "Review the attached receipt."},
"images": [Image.open("receipt.jpg")],
"questions": {
"legible": {"type": "noul", "instructions": "Is the receipt total legible?"},
},
}
encoded = encode_record(processor.tokenizer, record, processor=processor)
Text-only and multimodal records can be mixed in the same batch.
Table with columns: Field, Description| Field | Description |
|---|
state | Any string or JSON value describing the situation to decide on |
images, videos | Optional lists of images or video frame arrays |
media_kwargs | Optional keyword arguments for the image/video processor |
questions | Mapping of question ID to question |
Each question has:
type: noul (true/false), choice (named options), or score (ordered options)
instructions: what to decide; optional, and the question ID is used when it is omitted
criteria: for choice, a mapping of option ID to description; for score, a list of option
descriptions indexed from 0; for noul, optional descriptions for true and false
encode_record accepts max_length (default 16,384 tokens) and max_state_tokens to bound the input.
Results
Decision Index
Per-benchmark results from our internal run of the Decision Index 0.2.1 suite. Scores are percentages; ForecastBench is a Brier score, where lower is better. The last two rows are request latency in milliseconds, where lower is better. The best value in each row is in bold.
Table with columns: Benchmark, Clef, Clef-flash, Jev, DiffusionGemma Jev, Kev 9B, Laya| Benchmark | Clef | Clef-flash | Jev | DiffusionGemma Jev | Kev 9B | Laya |
|---|
| BFCL (case exact accuracy) | 98.5 | 98.8 | 95.8 | 96.5 | 94.5 | 38.1 |
| ToolRet (nDCG@10) | 69.2 | 66.4 | 65.3 | 61.2 | 64.3 |
Workflow evals
Decision accuracy on four end-to-end business workflows from Typesafe Evals, scored against consensus reference labels. All models are scored on the same dataset revision and case cohort.
Table with columns: Workflow, Metric, Clef, Clef-flash, Jev| Workflow | Metric | Clef | Clef-flash | Jev |
|---|
| Invoice processing | Exact actions | 64.7 | 57.1 | 61.8 |
| Invoice processing | Primary action | 86.2 | 73.3 | 83.1 |
| Customer service | Exact actions | 76.3 | 77.0 | 76.0 |
License
Released under the Apache-2.0 license, following the base model
Qwen/Qwen3.5-9B.