Downloads · 30 days
544
8% of all-time downloads
XiaomiMiMo/MiMo-Embodied-7B
MiMo-Embodied-7B is a image-text-to-text model from XiaomiMiMo. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
<div align="center" <img src="./assets/xfmlogo.svg" width=600 </div
Downloads · 30 days
544
8% of all-time downloads
All-time downloads
6.5K
Public
Parameters
8.3B
18 GB on disk
Likes
74
Public
Click a slice to open those files.
.safetensors18 GB · 100%
How the weights are stored.
BF167.6B · 92%
From the Hugging Face model README
MiMo-Embodied, a powerful cross-embodied vision-language model that shows state-of-the-art performance in both autonomous driving and embodied AI tasks, the first open-source VLM that integrates these two critical areas, significantly enhancing understanding and reasoning in dynamic physical environments.
<div align="center"> <img src="./assets/fig1.svg" width=800> </div>MiMo-Embodied demonstrates superior performance across 17 benchmarks in three key embodied AI capabilities: Task Planning, Affordance Prediction, and Spatial Understanding, significantly surpassing existing open-source embodied VLM models and rivaling closed-source models.
Additionally, MiMo-Embodied excels in 12 autonomous driving benchmarks across three key capabilities: Environmental Perception, Status Prediction, and Driving Planning—significantly outperforming both existing open-source and closed-source VLM models, as well as proprietary VLM models.
Moreover, evaluation on 8 general visual understanding benchmarks confirms that MiMo-Embodied retains and even strengthens its general capabilities, showing that domain-specialized training enhances rather than diminishes overall model proficiency.
Results marked with * are obtained using our evaluation framework.
@misc{hao2025mimoembodiedxembodiedfoundationmodel,
title={MiMo-Embodied: X-Embodied Foundation Model Technical Report},
author={Xiaomi Embodied Intelligence Team},
year={2025},
eprint={2511.16518},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2511.16518},
}