Downloads · 30 days
0
CMU-AIRe/SeeQ-3B
SeeQ-3B is a robotics model from CMU-AIRe. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for jax. The card lists the license as gemma.
SeeQ: Training Generalist Value Functions for Long-Horizon Robotic Manipulation
Downloads · 30 days
0
Access
Public
Updated Sep 21, 2026
Repo size
21.7 GB
Likes
2
Public
Click a slice to open those files.
Other21.7 GB · 100%
From the Hugging Face model README
SeeQ: Training Generalist Value Functions for Long-Horizon Robotic Manipulation
Saksham Singh, Zheyuan Hu, Max Sobol Mark, Jeffrey Yu, Zackory Erickson, Aviral Kumar (Carnegie Mellon University)
Project website (with videos): https://saksham002.github.io/seeq/
Paper: https://arxiv.org/abs/2609.22085
SeeQ (Subtask-elicited Q-functions) is a generalist Q-value function for language-conditioned robot manipulation. Rather than modeling sparse success over the full task, SeeQ estimates the expected return for the currently active subtask, which shortens the value-prediction horizon and makes temporal-difference (TD) learning effective on long-horizon tasks. At inference time the model first autoregressively decodes the active subtask in natural language, then scores candidate action chunks conditioned on it. The resulting Q-values are used for best-of-N steering of a base policy (pi-0.5 in the paper).
This repository contains the RoboCOIN-pretrained generalist checkpoint (230k training steps). In the paper this checkpoint is fine-tuned on each downstream task before deployment.
| Setting | Value |
|---|---|
| Data | RoboCOIN bimanual subset: 131 tasks, ~40k episodes, ~31.3M steps, 3 embodiments (Agilex Cobot Magic, Agilex Split ALOHA, Galaxea R1 Lite) |
| Objective | Subtask-level TD-BoN (no bootstrapping across subtask boundaries) + subtask next-token loss (weight 0.1) |
| Backup candidates | N = 8 action chunks from pi-0.5 |
| Discount | 0.999 |
| Action horizon | 50 (14-dim, end-effector, chunk-wise delta, quantile-normalized) |
| Target network | Polyak averaging, tau = 0.005 |
| Optimizer | AdamW, weight decay 1e-6 |
| LR schedule | Cosine, 1000 warmup steps, peak 1e-5, decayed to 1e-6 |
| Batch size / steps | 256 / 230,000 |
| Precision | float32 |
Training config: robocoin_bimanual_paligemma_cql_rlds_subtask_ar.
This is an openpi-style Orbax checkpoint (JAX / Flax NNX):
params/ — model parameters (Orbax PyTree checkpoint). Optimizer state is not included.assets/embodiment_wise/norm_stats.json — normalization statistics used during training; inputs must be normalized
with these at inference time._CHECKPOINT_METADATA — Orbax checkpoint metadata.This checkpoint is pretrained entirely on real-robot data, so its value estimates are tuned to real-world dynamics and visuals. On the simulation tasks we have tested, fine-tuning it yielded little improvement over the base policy, both with RaC (Recovery and Correction) data and with purely teleoperated data. We hypothesize that this is a property of the pretraining mixture rather than of the method: Section 4.3 of the paper shows that broad robot-action pretraining is what enables SeeQ to transfer to a new task, and simulated dynamics and rendering differ enough from real-world data that this transfer does not carry over to simulation.
We therefore recommend using this checkpoint for real-world tasks, where its pretraining is in-distribution and where the paper reports the largest gains. Real-world performance is ultimately what the method targets. If your goal is a simulation benchmark, the SeeQ objective is not tied to real data: pretrain a SeeQ model on a broad simulation dataset with the training config above, then fine-tune on the target task.
@article{singh2026seeq,
title = {SeeQ: Training Generalist Value Functions for Long-Horizon Robotic Manipulation},
author = {Singh, Saksham and Hu, Zheyuan and Mark, Max Sobol and Yu, Jeffrey and Erickson, Zackory and Kumar, Aviral},
journal = {arXiv preprint arXiv:2609.22085},
year = {2026},
url = {https://arxiv.org/abs/2609.22085}
}
The weights are derived from PaliGemma and are subject to the Gemma Terms of Use.