Downloads · 30 days
235
95% of all-time downloads
OpenMOSS-Team/MOSS-VL-Realtime-SGLANG
MOSS-VL-Realtime-SGLANG is a video-text-to-text model from OpenMOSS-Team. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
235
95% of all-time downloads
All-time downloads
248
Public
Parameters
11.3B
22.7 GB on disk
Likes
5
Public
Click a slice to open those files.
.safetensors22.7 GB · 100%
From the Hugging Face model README
English | 简体中文
<p align="center"><img src="assets/logo.png" width="300" alt="MOSS-VL"/></p>MOSS-VL-Realtime-SGLANG provides the MOSS-VL realtime checkpoint with Transformers 5.12.1-compatible model and processor code for the specialized SGLang-Omni backend. This is a compatibility package, not a retrained or quantized model.
| Goal | Entry |
|---|---|
| Browser video/voice interaction and optional memory | MOSS-VL-Realtime Demo |
| Standalone streaming inference service | sglang-omni-realtime |
| Direct Transformers inference | Python examples |
This repository contains weights, configuration, tokenizer, processor, and custom code. It does not include a running service. The model is public; download the complete repository rather than individual weight files.
ASR, TTS, and memory are provided by the Demo. The model's offline Python APIs do not mean that the Demo's realtime deployment enables offline chat.
Demo memory is session-scoped: reconnects within the grace period retain it; final closure clears retrieval data. Saved conversation archives are separate. Use the named repositories' current code and complete model snapshots; installed revisions are diagnostic records, not fixed installation requirements.
MOSS-VL separates visual encoding from language reasoning with cross-attention. Timestamped frames and Cross-attention Rotary Position Embedding (XRoPE) align text and visual patches across time, height, and width.
<p align="center"><img src="assets/architecture.png" alt="MOSS-VL Architecture" width="100%"/></p>| Item | Value |
|---|---|
| Parameters | 11B |
| Weights | BF16 |
| Model context | 256K |
| Vision patch size / temporal patch size | 16 / 1 |
| Default video preprocessing FPS / maximum sampled frames | 1.0 / 256 (offline video sampling, not a lifetime streaming-frame limit) |
| Direct Transformers runtime | One active realtime session per model instance |
| SGLang-Omni runtime | Configurable session capacity, subject to GPU memory |
The model context is not a per-session memory allocation guarantee. The recommended service configuration uses 131072 context; adjust context and concurrency for the available GPU memory.
MOSS-VL-Realtime targets streaming understanding, proactive silence, and dynamic response updates. Benchmark results are summarized below; see the technical report for model research.
<p align="center"><img src="assets/benchmark-streaming.png" alt="MOSS-VL Streaming Benchmark" width="100%"/></p>Use this package's custom code and configuration together with the specialized backend. Installing a generic sglang-omni package does not supply the MOSS-VL integration.
The backend uses Transformers 5.12.1, SGLang 0.5.16, and PyTorch 2.11.0. This package adapts configuration, RoPE, and generation interfaces to Transformers 5.12.1 while keeping the original five BF16 weight shards, tokenizer, and vocabulary unchanged.
The original MOSS-VL-Realtime uses separate Transformers 4.57-series code. Do not mix its custom files into this package. Select matching model and backend versions; CUDA compatibility does not imply NPU support.
<|silence|>, <|round_start|>, and <|round_end|>.| Model | Purpose |
|---|---|
| MOSS-VL-Realtime | Original realtime checkpoint and reference implementation |
| MOSS-VL-Instruct | Offline multimodal instruction following |
| MOSS-VL-Base | Continued pretraining and fine-tuning |
Apache-2.0. See OpenMOSS/MOSS-VL for the model family.
@misc{mossvl,
title = {MOSS-VL Technical Report},
author = {Wang, Pengyu and Tan, Chenkun and Zhou, Shaojun and Zhou, Qirui and Chen, Yanxin and He, Xingyang and Zeng, Huazheng and Cheng, Jijun and Wang, Chenghao and Qian, Xiaomeng and Wang, Pengfei and Huang, Zhan and Gao, Shanqing and Huang, Wei and Cao, Longjun and Ran, Wu and Liu, Jie and Zhu, Changtai and Wang, Hongkai and Tian, Yixian and Liu, Chenghao and Ye, Zhen and Wang, Xinghao and Jiang, Botian and Feng, Guoguo and Fei, Zhaoye and Li, Ruixiao and Chen, Mingshu and Gao, Yang and Cheng, Qinyuan and Li, Shimin and Qiu, Xipeng},
year = {2026},
eprint = {2608.15045},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2608.15045}
}
@misc{mossvideopreview,
title = {{MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention}},
author = {Pengyu Wang and Chenkun Tan and Shaojun Zhou and Wei Huang and Qirui Zhou and Zhan Huang and Zhen Ye and Jijun Cheng and Xiaomeng Qian and Yanxin Chen and Xingyang He and Huazheng Zeng and Chenghao Wang and Pengfei Wang and Hongkai Wang and Shanqing Gao and Yixian Tian and Chenghao Liu and Xinghao Wang and Botian Jiang and Xipeng Qiu},
year = {2026},
eprint = {2606.07639},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2606.07639}
}