Installation
pip install torch transformers==5.2.0 "qwen-vl-utils[decord]==0.0.14" \
"sentence-transformers>=5.7.0" "accelerate>=1.1.0"
import torch
from qwen_vl_utils import process_vision_info
from transformers import AutoModel, AutoProcessor
model_id = "tencent/WeMM-Embedding-2B"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModel.from_pretrained(
model_id, trust_remote_code=True, dtype=torch.bfloat16
).cuda().eval()
messages = [{"role": "user", "content": [
{"type": "image", "image": "/path/to/image.jpg"},
{"type": "video", "video": "/path/to/video.mp4"},
{"type": "text", "text": "This can be any text input."},
]}]
text = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=False
)
images, videos, video_kwargs = process_vision_info(
messages,
image_patch_size=16,
return_video_kwargs=True,
return_video_metadata=True,
)
if videos is not None:
videos, video_metadata = zip(*videos)
videos, video_metadata = list(videos), list(video_metadata)
else:
video_metadata = None
inputs = processor(
text=text,
images=images,
videos=videos,
video_metadata=video_metadata,
return_tensors="pt",
**video_kwargs,
).to("cuda")
with torch.inference_mode():
embedding = model.embedding(**inputs)
Use any subset of the content items to encode text, image, or video independently.
from sentence_transformers import SentenceTransformer
model_id = "tencent/WeMM-Embedding-2B"
model = SentenceTransformer(model_id, trust_remote_code=True)
queries = [
"Which Llama 4 model variants are available?",
"How is mapo tofu prepared?",
]
documents = [
"Mapo tofu is a Sichuan dish of soft tofu simmered in a spicy, numbing sauce of chili bean paste and Sichuan peppercorn.",
{
"image": "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/llama4_hgf.png",
"text": "Represent this image.",
},
{
"video": "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/mapo_tofu.mp4",
"text": "Represent this video.",
},
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
Each input is a string, a URL or path, a PIL.Image, or a dict combining image,
video, and text keys. Put image or video before text so the prompt matches
the ordering used above. Chat messages such as
{"role": "user", "content": [{"type": "image", "image": ...}, {"type": "text", "text": ...}]}
are also accepted, which is the way to interleave several images or videos in one input.
Matryoshka Embeddings
embedding_256 = torch.nn.functional.normalize(embedding[..., :256], dim=-1)
With Sentence Transformers, pass truncate_dim and let it renormalize:
embeddings_256 = model.encode_document(documents, truncate_dim=256, normalize_embeddings=True)
Use a dimension listed in model.config.matryoshka_dimensions. On MMEB-v2, 256-dimensional embeddings retain 98.7% of the full-dimensional image and video performance.
Serving
vLLM 0.27.0:
MODEL_PATH=/path/to/WeMM-Embedding-2B
vllm serve "$MODEL_PATH" \
--runner pooling \
--chat-template "$MODEL_PATH/embedding_chat_template.jinja"
SGLang 0.5.9:
MODEL_PATH=/path/to/WeMM-Embedding-2B
python patch_sglang_video.py
python -m sglang.launch_server \
--model-path "$MODEL_PATH" \
--is-embedding \
--enable-precise-embedding-interpolation
Evaluation
MMEB-v2
Results on 78 datasets from Table 1 of the technical report. Image and video tasks use Hit@1, while visual-document tasks use NDCG@5. Higher is better.
Table with columns: Model, Size, AVG, Image, Video, VisDoc| Model | Size | AVG | Image | Video | VisDoc |
|---|
| VLM2Vec | 2B | 47.8 | 59.7 | 29.0 | 44.0 |
| GME | 2B | 55.4 | 51.9 | 33.9 | 76.8 |
| VLM2Vec-V2 | 2B | 59.3 |
† Closed-source leaderboard submission without publicly released model weights or a public inference endpoint.
MMEB-v3
Results on all 190 tasks from Table 2 of the technical report. V3-All includes the 78 MMEB-v2 tasks, 53 text tasks, 47 agent tasks, 11 audio tasks, and MCMR. Unsupported tasks are assigned a score of zero.
Table with columns: Model, Size, V3-All, Text, Agent, MCMR, Audio| Model | Size | V3-All | Text | Agent | MCMR | Audio |
|---|
| VLM2Vec-V2 | 2B | 38.3 | 24.5 | 28.7 | 4.1 | 0.0 |
| Omni-Embed-Nemotron | 3B | 43.5 | 39.2 | 36.5 | 26.1 | 36.5 |
Text results use NDCG@5; agent, MCMR, and audio results use Hit@1.
Citation
If you find this repository useful, please consider giving a star ⭐ and citation
@article{wemm-embedding,
title={WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report},
author={Junjie Zhou and Ke Mei and Lei Li and Tianyi Wang and Fengyun Rao and Jing Lyu},
year={2026},
eprint={2608.24053},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.24053},
}
License
WeMM-Embedding-2B, including the code, model parameters, and weights made publicly
available by Tencent, is licensed under the Apache License 2.0.
Third-party components remain subject to their respective original licenses.