Downloads · 30 days
2.4K
24% of all-time downloads
BAAI/RoboBrain2.5-4B
RoboBrain2.5-4B is a machine learning model from BAAI. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<div align="center" <img src="https://github.com/FlagOpen/RoboBrain2.5/raw/main/assets/logo2.png" width="500"/ </div
Downloads · 30 days
2.4K
24% of all-time downloads
All-time downloads
9.8K
Public
Parameters
4.8B
9.7 GB on disk
Likes
14
Public
Click a slice to open those files.
.safetensors9.7 GB · 100%
From the Hugging Face model README
https://arxiv.org/abs/2601.14352
RoboBrain-2.5 is a next-generation Embodied AI foundation model that significantly evolves its predecessor's core capabilities in general perception, spatial reasoning, and temporal modeling through extensive training on high-quality spatiotemporal data. It achieves a paradigm shift in 3D Spatial Reasoning, transitioning from 2D relative points to predicting 3D coordinates with depth information, understanding absolute metric constraints, and generating complete manipulation trajectories tailored for complex tasks with physical constraints. Furthermore, it establishes a breakthrough in Temporal Value Prediction by constructing a General Reward Modeling Method that provides dense progress tracking and multi-granular execution state estimation across varying viewpoints. This empowers VLA reinforcement learning with immediate, dense feedback signals, enabling robots to achieve high task success rates and robustness in fine-grained manipulation scenarios.
<div align="center"> <img src="https://github.com/FlagOpen/RoboBrain2.5/raw/main/assets/teasor.png" /> </div>Compared to version 2.0, RoboBrain-2.5 achieves a leap in spatial perception and reasoning capabilities:
RoboBrain-2.5 makes significant progress in temporal modeling by constructing a General Reward Model (GRM):
RoboBrain 2.5 also maintains the three core capabilities of version 2.0, which supports interactive reasoning with long-horizon planning and closed-loop feedback, spatial perception for precise point and bbox prediction from complex instructions, temporal perception for future trajectory estimation, and scene reasoning through real-time structured memory construction and update.
# clone repo.
git clone https://github.com/FlagOpen/RoboBrain2.5.git
cd RoboBrain2.5
# build conda env.
conda create -n robobrain2_5 python=3.10
conda activate robobrain2_5
pip install -r requirements.txt
from inference import UnifiedInference
model = UnifiedInference("BAAI/RoboBrain2.5-8B-NV")
# Example:
prompt = "What is shown in this image?"
image = "http://images.cocodataset.org/val2017/000000039769.jpg"
pred = model.inference(prompt, image, task="general")
print(f"Prediction:\n{pred}")
from inference import UnifiedInference
model = UnifiedInference("BAAI/RoboBrain2.5-8B-NV")
# Example:
prompt = "the person wearing a red hat"
image = "./assets/demo/grounding.jpg"
# Visualization results will be saved to ./result, if `plot=True`.
pred = model.inference(prompt, image, task="grounding", plot=True, do_sample=False)
print(f"Prediction:\n{pred}")
from inference import UnifiedInference
model = UnifiedInference("BAAI/RoboBrain2.5-8B-NV")
# Example:
prompt = "the affordance area for holding the cup"
image = "./assets/demo/affordance.jpg"
# Visualization results will be saved to ./result, if `plot=True`.
pred = model.inference(prompt, image, task="pointing", plot=True, do_sample=False)
print(f"Prediction:\n{pred}")
from inference import UnifiedInference
model = UnifiedInference("BAAI/RoboBrain2.5-8B-NV")
# Example:
prompt = "Identify spot within the vacant space that's between the two mugs"
image = "./assets/demo/pointing.jpg"
# Visualization results will be saved to ./result, if `plot=True`.
pred = model.inference(prompt, image, task="pointing", plot=True, do_sample=True)
print(f"Prediction:\n{pred}")
from inference import UnifiedInference
model = UnifiedInference("BAAI/RoboBrain2.5-8B-NV")
# Example 1:
prompt_1 = "Identify spot within toilet in the house"
image = "./assets/demo/navigation.jpg"
# Visualization results will be saved to ./result, if `plot=True`.
pred = model.inference(prompt_1, image, task="pointing", plot=True, do_sample=True)
print(f"Prediction:\n{pred}")
# Example 2:
prompt_2 = "Identify spot within the sofa in the house"
image = "./assets/demo/navigation.jpg"
# Visualization results will be saved to ./result, if `plot=True`.
pred = model.inference(prompt_2, image, task="pointing", plot=True, do_sample=True)
print(f"Prediction:\n{pred}")
from inference import UnifiedInference
model = UnifiedInference("BAAI/RoboBrain2.5-8B-NV")
# Example:
prompt = "reach for the banana on the plate"
image = "./assets/demo/trajectory.jpg"
# Visualization results will be saved to ./result, if `plot=True`.
pred = model.inference(prompt, image, task="trajectory", plot=True, do_sample=False)
print(f"Prediction:\n{pred}")
We highly recommend referring to Robo-Dopamine for detailed usage instructions.
# clone Robo-Dopamine repo.
git clone https://github.com/FlagOpen/Robo-Dopamine.git
cd Robo-Dopamine
import os
from examples.inference import GRMInference
# model = GRMInference("tanhuajie2001/Robo-Dopamine-GRM-3B")
model = GRMInference("BAAI/RoboBrain2.5-8B-NV")
TASK_INSTRUCTION = "organize the table"
BASE_DEMO_PATH = "./examples/demo_table"
GOAL_IMAGE_PATH = "./examples/demo_table/goal_image.png"
OUTPUT_ROOT = "./results"
output_dir = model.run_pipeline(
cam_high_path = os.path.join(BASE_DEMO_PATH, "cam_high.mp4"),
cam_left_path = os.path.join(BASE_DEMO_PATH, "cam_left_wrist.mp4"),
cam_right_path = os.path.join(BASE_DEMO_PATH, "cam_right_wrist.mp4"),
out_root = OUTPUT_ROOT,
task = TASK_INSTRUCTION,
frame_interval = 30,
batch_size = 1,
goal_image = GOAL_IMAGE_PATH,
eval_mode = "incremental",
visualize = True
)
print(f"Episode ({BASE_DEMO_PATH}) processed with Incremental-Mode. Output at: {output_dir}")
We sincerely thank RoboTracer (ECCV 2026) for bringing native 3D spatial reasoning capabilities to RoboBrain 2.5, and Robo-Dopamine (CVPR 2026) for providing temporal progress estimation and prediction capabilities.