Authors
Serving with vLLM
vllm serve objectai/obj_v1 \
--served-model-name obj_v1 \
--max-model-len 16384 \
--limit-mm-per-prompt '{"image":1}' \
--mm-processor-kwargs '{"max_pixels":1003520}' \
--trust-remote-code
max_pixels is 1280x28x28, the resolution the model was trained at. Raising it
wastes KV cache; lowering it makes small print unreadable.
Calling it
The server is OpenAI-compatible, so an ordinary chat completion works:
import base64, json, openai
client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
image = base64.b64encode(open("cheque.jpg", "rb").read()).decode()
schema = {"cheque_details": {"amount": "number", "payee": "string",
"date": "string", "cheque_number": "string"}}
response = client.chat.completions.create(
model="obj_v1",
temperature=0.0,
max_tokens=8192,
messages=[
{"role": "system", "content":
"You are a document data extraction model. "
"Extract only values present in the document. "
"Use null for fields that are absent or illegible. "
"Output a single compact JSON object matching the requested schema. "
"No prose, no markdown, no explanation."},
{"role": "user", "content": [
{"type": "image_url",
"image_url": {"url": f"data:image/jpeg;base64,{image}"}},
{"type": "text",
"text": f"document_type: cheque\nschema: {json.dumps(schema)}"},
]},
],
)
print(response.choices[0].message.content)
Match training or accuracy drops. The system prompt above is verbatim, and the
user turn is the image followed by exactly two lines:
document_type: <type>
schema: <compact json>
Set temperature=0.0 so the same page yields the same answer.
Requirements
Table | |
|---|
| Weights | 8.9 GB (bf16) |
| VRAM | 16 GB minimum, 24 GB comfortable |
| Precision | bf16 (Ampere or newer; use fp16 below that) |
| Context | 16384 covers the longest documents |
Runs on an L4, A10G, L40S, A100 or RTX 4090. On a T4 add --dtype float16. | |
Output
Compact JSON matching the requested schema. Fields absent from the page come
back null rather than guessed. Values found on the page that the schema did
not ask for are placed under extras when that key is included in the schema.
Limitations
- Trained on Indian financial documents; other domains and layouts are untested.
- Handwriting is the weakest case, particularly digits at low resolution.
- The model does not verify its own arithmetic. Totals that must reconcile
should be checked by the caller.
License
Apache 2.0. Fine-tuned from Qwen3-VL-4B-Instruct, which is Apache 2.0.