Highlights
- Text as visual tokens: represent document text through rendered images instead of conventional text token sequences.
- Shorter input sequences: compact visual representations reduce input length and allow more document content within a fixed context budget.
- Three input modes: score document text directly with
text, render it automatically with render, or provide existing document pages with image.
Table with columns: Property, Value| Property | Value |
|---|
| Model type | Pairwise document reranker |
| Backbone | Qwen3-VL-Reranker-2B |
| Parameters | 2B |
| Language evaluated | English |
| Maximum context length | 32,768 tokens, including query, template, document text, and visual tokens |
| Scoring | Yes/no relevance scoring |
Rendering Configuration
Table with columns: Setting, Value| Setting | Value |
|---|
| Font | Roboto Regular |
| Font size | 12pt at 96 DPI (16px) |
| Line spacing | 1.0 (16px line height) |
| Page width | 896px |
| Page height | 32–896px, rounded up to a multiple of 32px |
| Lines per page | Up to 56 |
| Margins | 0px on all sides |
| Colors | Black text on a white background |
| Rendering |
Whitespace, including paragraph breaks and tabs, is normalized to a single space. Lines wrap according to the font's actual pixel width; words wider than the page are split by character. Overflow continues onto subsequent pages without overlap. The final page is sized to its content, and empty documents produce one 896 × 32px page.
All generated pages are supplied in order as one document, producing one relevance score. Rendering is not limited to the first page.
All rerankers score the top 100 BM25 candidates for each query. Baselines use text queries and documents; RenderRank uses text queries and documents rendered with the default 12pt configuration above. Scores are NDCG@10; Avg. is the mean across 11 datasets. Tokens is the mean input length per query–document pair before truncation, averaged across the same datasets. For RenderRank, this includes both textual and visual tokens.
Table with columns: Model, AA, CFV, DBP, FQA, FVR, HQA, NFC, SD, SF, TC, TCH, Avg., Tokens| Model | AA | CFV | DBP | FQA | FVR | HQA | NFC | SD | SF | TC | TCH | Avg. | Tokens |
|---|
| gte-reranker-modernbert-base | 66.14 | 25.80 | 42.10 | 42.54 | |
AA: ArguAna; CFV: Climate-FEVER; DBP: DBPedia; FQA: FiQA; FVR: FEVER; HQA: HotpotQA; NFC: NFCorpus; SD: SCIDOCS; SF: SciFact; TC: TREC-COVID; TCH: Touché-2020.
Results on the English subset of MLDR and three LongEmbed datasets: 2WikiMQA, QMSum, and SummScreenFD. The comparison includes models supporting at least 16K input tokens. MLDR follows the MMTEB reranking protocol; for LongEmbed, each query reranks eight candidates retrieved with Qwen3-Embedding-0.6B.
Each dataset cell shows NDCG@10 followed by the average input token count in brackets: score [tokens]. Token counts are measured per query–document pair before truncation and include both textual and visual tokens for RenderRank. Avg. reports the mean score across the four datasets.
Table with columns: Model, MLDR, 2WikiMQA, QMSum, SummScreenFD, Avg.| Model | MLDR | 2WikiMQA | QMSum | SummScreenFD | Avg. |
|---|
| Qwen3-Reranker-0.6B | 99.63 [8863.4] | 94.54 [9230.2] | 56.43 [13543.6] | 98.25 [8606.3] | 87.21 |
| LightOn-rerank-PW-2B | 98.91 [8878.0] | 71.27 [9230.1] | 54.30 [13876.6] | 93.05 [8977.2] | 79.38 |
| Qwen3-Reranker-4B | 99.85 [8863.4] |
Usage
The examples load the model from nlpai-lab/RenderRank-2B on Hugging Face.
The model applies its reranking chat template internally. Pass the instruction as shown below; do not manually prepend a chat template to the document.
This repository provides a custom AutoModel entry point with a process() method. It supports all three document input modes.
Requirements
The Transformers interface was tested with the following versions:
transformers==5.9.0
qwen-vl-utils==0.0.14
torch==2.11.0
torchvision==0.26.0
scipy
qwen-vl-utils handles image preparation in this interface; scipy is used for score normalization. NumPy and Pillow are installed through the dependencies above. Install PyTorch and torchvision builds matching your CUDA environment.
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained(
"nlpai-lab/RenderRank-2B",
trust_remote_code=True,
dtype=torch.bfloat16,
).to("cuda").eval()
inputs = {
"instruction": "Retrieve text relevant to the user's query.",
"query": {"text": "What is the capital of France?"},
"documents": [
{"text": "Paris is the capital of France."},
{"render": "Paris is the capital of France."},
{"image": "/path/to/document.png"},
{"image": ["/path/to/page1.png", "/path/to/page2.png"]},
],
}
with torch.inference_mode():
scores = model.process(inputs)
ranked_indices = sorted(range(len(scores)), key=lambda i: scores[i], reverse=True)
print(scores)
print(ranked_indices)
Rendering
With Transformers, {"render": document_text} automatically renders the document before scoring. The font is included in this repository.
Alternatively, render documents in advance with the bundled rendering.py and save the page images for reuse with either Transformers or Sentence Transformers. This avoids repeated rendering; vision encoding still runs when scoring the saved images.
import sys
from pathlib import Path
from huggingface_hub import snapshot_download
repo_dir = Path(snapshot_download(
"nlpai-lab/RenderRank-2B",
allow_patterns=["rendering.py", "Roboto-Regular.ttf"],
))
sys.path.insert(0, str(repo_dir))
from rendering import DocumentImageConfig, DocumentImageRenderer
renderer = DocumentImageRenderer(
DocumentImageConfig(font_path=repo_dir / "Roboto-Regular.ttf")
)
pages = renderer.render_document("Paris is the capital of France.")
output_dir = Path("rendered_document")
output_dir.mkdir(parents=True, exist_ok=True)
for i, page in enumerate(pages, start=1):
page.save(output_dir / f"{i}.png")
The returned PIL images can be passed directly as image content or saved as numbered page files. The Transformers render mode calls this renderer internally, so no separate rendering step is needed there.
Use Sentence Transformers v6.1.0 or later.
pip install -U "sentence-transformers>=6.1.0"
The same interface can score text, a single document image, or multiple page images as one document. In the example below, 1.png, 2.png, and 3.png are ordered pages of one document.
import torch
from sentence_transformers import CrossEncoder
model = CrossEncoder(
"nlpai-lab/RenderRank-2B",
device="cuda",
max_length=32768,
model_kwargs={"dtype": torch.bfloat16},
)
query = "What is the capital of France?"
documents = [
{"text": "Paris is the capital of France."},
{"image": "/path/to/document.png"},
{"image": [
"/path/to/1.png",
"/path/to/2.png",
"/path/to/3.png",
]},
]
scores = model.predict(
[(query, document) for document in documents],
prompt="Retrieve text relevant to the user's query.",
batch_size=1,
)
print(scores)
Larger scores indicate greater relevance. Scores are the difference between the yes and no logits, as used in our evaluation. With Sentence Transformers v6.1.0 or later, multi-page documents can be passed as a single multimodal input using {"image": [page1, page2, ...]}. The pages are processed together and receive one relevance score per document. The render input format is supported by the Transformers interface above.
Citation
@misc{hong2026renderranklearningreranktext,
title={RenderRank: Learning to Rerank Text with Compressed Visual Tokens},
author={Seongtae Hong and Youngjoon Jang and Jungseob Lee and Hyeonseok Moon and Heuiseok Lim},
year={2026},
eprint={2609.35069},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2609.35069},
}