HebVL — open vision-language models for historical Hebrew-script documents
Table with columns: Model, What changed, Headline result| Model | What changed | Headline result |
|---|
| v1 | First fine-tune, printed Vilna Talmud | Reads unseen gemara at CER ~0.06–0.2 where the base model refused |
| v1.5 | + synthetic Rashi-script training | ~99% char accuracy on rendered Rashi (pure perception); crop-level gemara solved |
| v1.6 | Page-level training at 6.5 MP | Printed Rashi CER 0.047 — ahead of the frontier VLMs we benchmarked |
| v1.7 | + Cairo Genizah handwriting (curated corpus v1) | Best full-coverage Talmud model; first real handwriting capability |
| v1.8a | Repaired training data (ablation control) | Isolates the data-quality contribution |
| v1.8b | Vision tower actually trained (LoRA) | Outperforms the Kraken HTR baseline on our Genizah benchmark; no forgetting of printed scripts |
HebVL is a family of fine-tuned Qwen3-VL-8B
adapters that read historical Hebrew-script material — from printed Talmud pages
(square type, Rashi script) to the handwritten fragments of the Cairo Genizah:
medieval letters, court deeds, and lists written in Hebrew, Judaeo-Arabic, and
Aramaic, in the notoriously difficult Oriental documentary cursive of
11th–13th-century Egypt.
These models are part of the Cairo Genizah AI Project — a search
engine, map, and semantic visualization of the Cairo Genizah corpus. Today, search
there runs mostly on catalog metadata, because the great majority of Genizah
fragments have never been transcribed. This model line exists to change that: once
transcription quality is good enough, we will apply it across the untranscribed
corpus to unlock genuine full-text and semantic search over one of the most
important documentary archives of the medieval Mediterranean.
Where the line stands (v1.8b): on our frozen 131-fragment Genizah benchmark
(verified image/transcription pairs, scored by behaviour classification, aligned
F1, and order-independent n-gram precision), v1.8b is the first model in the line
to outperform a specialist HTR baseline (Kraken with a medieval-Hebrew model):
aligned F1 0.807 vs 0.730, with ink-fidelity (n-gram precision) at parity. On
printed Talmud it transcribes gemara and Rashi script at character error rates
competitive with or better than commercial frontier VLMs. An accompanying paper
is under peer review; links will follow at camera-ready.
Honest limitations — please read before scholarly use. These are research
artifacts. On damaged or highly cursive fragments the models can produce fluent
but wrong text — invented dates, names, or filled-in lacunae that look like
readings. Never treat an output as an authoritative transcription without
checking it against the image. Practical tips: feed high-resolution scans (the
exported processors carry the training resolution policy, ~6.5 MP); expect better
results when prompting for a page or region at a time; damage gaps are marked
[...] by convention