Downloads · 30 days
66
8% of all-time downloads
cagataydev/sac-unitree-g1-mujoco
sac-unitree-g1-mujoco is a reinforcement learning model from cagataydev. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for stable-baselines3.
A Soft Actor-Critic (SAC) policy trained for the Unitree G1 humanoid in MuJoCo simulation. Currently learning to balance — stays upright ~4 seconds and stumbles forward.
Downloads · 30 days
66
8% of all-time downloads
All-time downloads
777
Public
Repo size
84.5 MB
Likes
1
Public
Click a slice to open those files.
.zip83 MB · 98%
From the Hugging Face model README
A Soft Actor-Critic (SAC) policy trained for the Unitree G1 humanoid in MuJoCo simulation. Currently learning to balance — stays upright ~4 seconds and stumbles forward.
Trained entirely on a MacBook (CPU, no GPU, no Isaac Gym) using strands-robots.
| Metric | Value |
|---|---|
| Algorithm | SAC (Soft Actor-Critic) |
| Training steps | 1.91M |
| Training time | ~60 min (MacBook M-series, CPU) |
| Parallel envs | 8 |
| Network | MLP [256, 256] |
| Best reward | 530 |
| Mean distance | 2.65m |
| Episode length | ~200/1,000 (~4 seconds upright) |
| Status | Balancing + stumbling forward |
<video src="https://huggingface.co/cagataydev/sac-unitree-g1-mujoco/resolve/main/g1_balancing.mp4" controls autoplay loop muted></video>
The G1 has 29 DOF vs Go2's 12. Bipedal balance is fundamentally harder — the robot must coordinate hip, knee, ankle, and torso simultaneously while maintaining a tiny support polygon.
With more training (~5-10M steps, ~3 hours), it should learn to walk.
from stable_baselines3 import SAC
model = SAC.load("best/best_model")
obs, _ = env.reset()
for _ in range(1000):
action, _ = model.predict(obs, deterministic=True)
obs, reward, done, truncated, info = env.step(action)
reward = forward_vel × 5.0 # primary: move forward
+ alive_bonus × 1.0 # stay upright
+ upright_reward × 0.3 # orientation bonus
- ctrl_cost × 0.001 # minimize energy
- lateral_penalty × 0.3 # don't drift sideways
- smoothness × 0.0001 # discourage jerky motion
best/best_model.zip — Best checkpointcheckpoints/ — All 100K-step checkpointslogs/evaluations.npz — Evaluation metricsg1_balancing.mp4 — Demo videoApache-2.0