Downloads · 30 days
0
ThomasTheMaker/vjepa2-robot-multitask
vjepa2-robot-multitask is a robotics model from ThomasTheMaker. Use it for the robotics task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Vision-based robot control data using V-JEPA 2 (ViT-L) latent representations from DeepMind Control Suite environments.
Downloads · 30 days
0
Access
Public
Updated Mar 1, 2026
Repo size
3 GB
Likes
0
Public
Click a slice to open those files.
.npz2.9 GB · 97%
From the Hugging Face model README
Vision-based robot control data using V-JEPA 2 (ViT-L) latent representations from DeepMind Control Suite environments.
| Task | Episodes | Transitions | Latent Dim | Action Dim | Success Rate |
|---|---|---|---|---|---|
| reacher_easy | 1,000 | 200,000 | 1024 | 2 | 28.9% |
| point_mass_easy | 1,000 | 200,000 | 1024 | 2 | 0.6% |
| cartpole_swingup | 1,000 | 200,000 | 1024 | 1 | 0.0% |
Each .npz file contains:
z_t — V-JEPA 2 latent state embeddings (N × 1024)a_t — actions taken (N × action_dim)z_next — next-state latent embeddings (N × 1024)rewards — per-step rewards (N,)For each task, we provide:
dyn_0.pt to dyn_4.pt (MLP: z + a → z_next, ~1.58M params each)reward.pt (MLP: z + a → reward, ~329K params)Linear(1024+a_dim, 512) → LN → ReLU → ×3 → Linear(512, 1024) + residual connectionLinear(1024+a_dim, 256) → ReLU → ×2 → Linear(256, 1)These world models are designed for "teach-by-showing" — demonstrating a task via video, then using the learned dynamics + CEM planning to reproduce the shown behavior.