Downloads · 30 days
413
100% of all-time downloads
ATH-MaaS/Ovis-VL-Embedding-9B
Ovis-VL-Embedding-9B is a feature extraction model from ATH-MaaS. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
<img alt="GitHub" src="https://img.shields.io/badge/GitHub-Ovis--VL--Embedding-3157C8?logo=github" style="display: inline-block; vertical-align: middle;"/ </a --
Downloads · 30 days
413
100% of all-time downloads
All-time downloads
413
Public
Repo size
16.8 GB
Likes
38
Trending 8
Click a slice to open those files.
.safetensors16.8 GB · 100%
From the Hugging Face model README
Ovis-VL-Embedding-9B is a high-capacity vision-language embedding model for text, images, visual documents, video, and interleaved multimodal inputs. It maps every supported input type into one coherent representation space, enabling high-accuracy cross-modal retrieval with a single encoder.
The model is initialized from Qwen3.5-9B. It retains the native text and vision encoders together with the shared multimodal language backbone, removes the language-modeling head, and directly uses the final-layer hidden state at the last non-padding token as the retrieval embedding. No modality-specific projection head is added.
Ovis-VL-Embedding-9B is designed for high-quality multimodal search, multimodal RAG, image and visual-document retrieval, video search and temporal localization, recommendation, and nearest-neighbor matching.
Figure 2 from the technical report. Ovis-VL-Embedding-9B encodes text, images, visual documents, and sampled video frames as one interleaved sequence. The final-layer hidden state at the last non-padding token becomes the retrieval embedding, with no modality-specific projection heads.
The Qwen3.5-9B backbone contains 32 language layers with hidden size 4096. Its hybrid stack repeats three Gated DeltaNet layers followed by one gated full-attention layer, combining efficient long-context processing with periodic global token interaction. Multimodal positional encoding preserves temporal and two-dimensional spatial coordinates for visual tokens.
Training follows three stages:
Ovis-VL-Embedding-9B is a bi-encoder, not a cross-encoder:
No answer generation or query-candidate cross-attention is used during retrieval. Classification labels, passages, images, documents, videos, and interleaved multimodal items are all treated as candidates in the same embedding space.
MMEB-v2 evaluates vision-language embeddings over 78 datasets spanning image, video, and visual-document tasks. Ovis-VL-Embedding-9B achieves 81.13 overall, outperforming the strongest compared baseline by 1.04 points.
| Group | Ovis-VL-Embedding-9B | Best compared baseline | Result |
|---|---|---|---|
| Image | 83.96 | 81.86 | +2.10 |
| Video | 72.90 | 75.95 | -3.05 |
| Visual document | 83.06 | 82.38 | +0.68 |
| All 78 datasets | 81.13 | 80.09 | +1.04 |
The model ranks first on all four image sub-tasks, video classification, video moment retrieval, the visual-document aggregate, and ViDoRe-V1. Scaling from 2B to 9B improves the overall score by 3.67 points, with the largest gains on video question answering (+7.64), video moment retrieval (+7.23), and video retrieval (+5.74).
<p align="center"> <img src="./figures/ovis_blog_table2.png" alt="Complete MMEB-v2 comparison for Ovis-VL-Embedding-9B and four vision-language embedding baselines" width="100%"/> </p>Scores are percentages and higher is better. Red marks the best result in each row, underlining marks the second best, and Ovis scores are bold. The overall score is the unweighted average over all 78 MMEB-v2 datasets.
The native output width is 4096, inherited directly from the Qwen3.5-9B backbone because no embedding projection head is added. Queries and candidates must use the same preprocessing, pooling rule, dimensionality, and L2 normalization.
The model is intended for embedding extraction and retrieval over supported unimodal or interleaved multimodal content, including:
If you find our embedding models useful, please consider citing our technical report:
@article{ovisembedding2026,
title = {Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings},
author = {{Ovis-Embedding Team}},
journal = {arXiv preprint arXiv:2609.25165},
year = {2026},
url = {https://arxiv.org/abs/2609.25165}
}
This model is released under the Apache 2.0 license.
This repository contains the weights for Ovis-VL-Embedding-9B.