Model details
- Base model: infly/Infinity-Parser2-Flash + LoRA adapter, merged into the published weights
- Adapter: LoRA r8 / alpha 32, all-linear targets, ViT + aligner frozen (8.4M trainable params, 0.38%)
- Developed by: eddie (fine-tune), annotation-app pipeline
- Version: title-v2-full-8m-001, 2026-09-14, git fb06415 (main)
- License/lineage: inherits base model license; fine-tuned for internal title-block extraction
- Type: supervised fine-tune (ms-swift SFT), full-page 200 DPI PNG input → JSON
{sheet_no, title}
- Serving: merged weights (this repo) via vLLM; equivalent hot-loadable adapter:
experiments/vllm_adapters/title-v2-full-8m-001 (mamba-stripped)
Intended use
Extract the engineering-drawing title-block fields (sheet number and sheet title) from a full-page raster of a construction/engineering sheet. Prompts and labels follow the annotation-app "Title block" task: the sheet title is NOT the project name, firm, or address; ignore "OF N" totals and PDF page numbers unless they are the actual identifier.
Training data — changes from v1
- v1 production dataset: 1,208 rows / 16 projects (the Spark's all.jsonl), human-reviewed.
- v2 dataset: train_v2_full_8m.jsonl, 1,236 rows (sha256 eb0376d0…b2618b9cf). Changes vs v1:
- +28 Ferguson-Geo rows (2 new projects, previously failing documents), human-reviewed:
- warmlands_avenue (12 pages): curve-drawn title block with NO text layer; labels drafted via model-assisted transcription of upscaled crops + pixel-diff localization. sheet_no = the N in "SHEET N OF 12"; title = block title line WITHOUT the "FOR: WARMLANDS SUBDIVISION" suffix (project name is not the title).
- forge_biologics (16 pages): dual number cell — sheet_no = the DRAWING NUMBER (C0.1, C1.0, TS1.1, …), never the big page-order digit printed next to it.
- Local re-render of the original 1,208 rows: images re-rendered from the S3 source PDFs at 200 DPI with a pixel-identical pipeline (verified). Bent-creek labels cross-checked against locally reviewed bent_creek_dataset title labels: 217/219 match. INDOT R-41176 (only ambiguous PDF→project mapping) validated 106/106 against production predictions.
- Per-row
chat_template_kwargs.max_pixels 16,777,216 → 8,388,608. Row format otherwise identical (same ms-swift message schema, same extraction prompt, same eval split untouched).
Training procedure
Same recipe as v1: LoRA r8/a32 all-linear, lr 1e-4 cosine, 2 epochs, effective batch 8 (per-device 1 × grad-accum 8), bf16, max_len 32768, sdpa, ViT+aligner frozen.
Only deviation from v1: max_pixels 8,388,608 (16M OOMs in the qwen3.5 vision tower on the 32 GB RTX 5090; reproduced 4×, incl. flash-attn). Serving works at the processor default and at production max_pixels settings below 16M (evaluated).
310 steps / 2h15m; final loss ~0.002, token_acc 0.999; checkpoint-310.
Evaluation results
Ferguson-Geo 28-page regression set (temp 0)
sheet_no 28/28, title 28/28 (production v1 baseline: 19/28 and 17/28). All warmlands "GP17-037" grabs, "FOR: …" title suffixes, and forge page-digit grabs eliminated.
119-page held-out set, 4 unseen firms (proxy-gold = production predictions)
sheet_no 118/119 (99.2%), title 116/119 (97.5%) — parity with v1. Adjudicated disagreements: 1 sheet_no (ROYALTON p1 "811" utility-contact grab), 1 title (appended project/owner-block text, ROYALTON p18), 1 page both wrong (GLC p8, production also wrong).
S3 Ferguson-Geo held-out — 17 unseen projects, 52 pages, human-reviewed ground truth (NEW 2026-09-16)
Sampled ≤4 title-block pages per project from the annotation team's grown ground truth (s3://circuit-silicon-ml/customer/ferguson-geo/, sheets.json labels; warmlands_avenue and forge_biologics excluded as v2 training projects; zero project overlap with v1/v2 training). Full-page 200 DPI renders, temperature 0.
Table with columns: Metric, Exact match, Containment**, Adjudicated*| Metric | Exact match | Containment** | Adjudicated* |
|---|
| sheet_no | 46/52 (88.5%) | 46/52 (88.5%) | 46/52 (88.5%) |
| title | 40/52 (76.9%) | 45/52 (86.5%) | 48/52 (92.3%) |
Same 52 pages scored against the deployed v1 model (Circuit-AI/infinity-parser2-flash-title-block):
Table with columns: Metric, v1 exact, v1 containment, v2 exact, v2 containment| Metric | v1 exact | v1 containment | v2 exact | v2 containment |
|---|
| sheet_no | 46/52 (88.5%) | 46/52 (88.5%) | 46/52 (88.5%) | 46/52 (88.5%) |
| title | 45/52 (86.5%) | 47/52 (90.4%) | 40/52 (76.9%) | 45/52 (86.5%) |
The two models return identical predictions on 39/52 pages. On the differing pages, v1's verbose predictions often contain gold titles that embed project/location text (a gold-convention artifact — the S3 gold sometimes includes project name/location, contrary to this task's label convention, which v2 follows); v2 conversely fixes v1's SUBSTATION cover miss (TNORHC800) and INDIAN WELLS title. v2's unique residual sheet_no failures are cover-sheet/scale-note grabs ("83363" license, "30" OF-N total, "NTS" ×2). Full per-page comparison: report.html in the MLflow run artifacts.
**Containment = correct if strings are equal or one contains the other (case/whitespace-normalized, contained side ≥3 chars) — e.g. pred "PAGE 4 OF 30" vs gold "PAGE 4" counts as correct.
*Adjudicated = excluding 8 pages where the gold title includes project name / location / street extents (contrary to this task's label convention, which the model followed), and 1 page where the gold label itself is wrong ("1 OF 10 SHEETS 07/24/2026" as title). Adjudication is text-based (no image re-review) — treat the adjudicated title number as approximate.
Genuine errors observed (candidate failure patterns for v3):
- Cover-sheet grabs: Florida P.E. license number "83363" as sheet_no (PORT CHARLOTTE p1) — same family as the v2-known "811" utility-contact grab.
- "OF 30" total-sheet count grabbed as sheet_no on cover + inner sheets (SAFER FOX HILLS p1, p26).
- "NTS" (no-scale note) grabbed as sheet_no where the drawing-number cell is non-standard (SUBSTATION 1000 p8, p14).
- "SHEET1" vs gold "1" formatting near-miss (INDIAN WELLS p1).
- Truncated or alternate title-line choice on 3 pages (KELLYVILLE p4, WOODSPRING p1, WOODSPRING p24 — the last an OCR-style typo "ELEATIONS").
MLflow metrics (run af72f80dc7704ba6992929075596ea57)
ferguson28_sheet_no_exact=1.0, ferguson28_title_exact=1.0, heldout_sheetno_agreement_prod=0.9916, heldout_title_agreement_prod=0.9748.
Limitations
- Cover sheets with utility-contact / license numbers ("811", "83363"), "OF N" totals, and "NTS" scale notes are the dominant residual sheet_no failure patterns on unseen firms.
- Project/owner-block text appended to titles remains a residual title failure pattern.
- The S3 held-out sample is 52 pages / 17 projects (≤4 per project) — indicative, not a large-scale benchmark.
- Training-time image resolution (8M px) differs from serving — validated on eval sets but a resolution-domain gap exists in principle.
Usage
Serving with vLLM
vllm serve Circuit-AI/infinity-parser2-flash-title-block-v2 \
--trust-remote-code --reasoning-parser qwen3 \
--host 0.0.0.0 --port 8000 \
--tensor-parallel-size 1 \
--gpu-memory-utilization 0.85 \
--max-model-len 65536 \
--mm-encoder-tp-mode data \
--mm-processor-cache-type shm \
--enable-prefix-caching
Inference
import base64, json, urllib.request
def extract_title_block(png_path, api_url="http://localhost:8000/v1/chat/completions"):
with open(png_path, "rb") as f:
data_url = "data:image/png;base64," + base64.b64encode(f.read()).decode()
prompt = (
"Read the engineering drawing's title block (a bordered box, usually on the "
"right edge or bottom-right of the sheet). Return ONLY a JSON object with "
"these fields:\n"
"- sheet_no: the sheet/drawing number identifier in the title block.\n"
"- title: the SHEET title / drawing title line inside the title block.\n"
"Do not infer from the filename or anything outside the image. If the sheet "
"has no title block, return null values. Respond with the JSON object only."
)
body = {
"model": "Circuit-AI/infinity-parser2-flash-title-block-v2",
"messages": [{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": data_url}},
{"type": "text", "text": prompt},
]}],
"temperature": 0.0,
"max_tokens": 512,
"chat_template_kwargs": {"enable_thinking": False},
}
req = urllib.request.Request(api_url, data=json.dumps(body).encode(),
headers={"Content-Type": "application/json"})
resp = json.loads(urllib.request.urlopen(req).read())
return resp["choices"][0]["message"]["content"]
Inference notes
- 200 DPI full page — match the training format for best results
temperature=0 — deterministic, same input always gives same output
enable_thinking=False — direct JSON output, no reasoning trace
- ~1.5–8 seconds per page on a single GPU
Citation
@misc{circuit-ai-title-block-extractor,
title={Fine-tuning Infinity-Parser2-Flash for Title Block Extraction},
author={Circuit AI},
year={2026},
base_model={infly/Infinity-Parser2-Flash}
}
Acknowledgments