Downloads · 30 days
54
8% of all-time downloads
cagataydev/sac-unitree-go2-mujoco
sac-unitree-go2-mujoco is a reinforcement learning model from cagataydev. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for stable-baselines3.
A Soft Actor-Critic (SAC) policy trained to make the Unitree Go2 quadruped walk forward in MuJoCo simulation.
Downloads · 30 days
54
8% of all-time downloads
All-time downloads
697
Public
Repo size
64.5 MB
Likes
0
Public
Click a slice to open those files.
.zip63.3 MB · 98%
From the Hugging Face model README
A Soft Actor-Critic (SAC) policy trained to make the Unitree Go2 quadruped walk forward in MuJoCo simulation.
Trained entirely on a MacBook (CPU, no GPU, no Isaac Gym) using strands-robots.
| Metric | Value |
|---|---|
| Algorithm | SAC (Soft Actor-Critic) |
| Training steps | 1.74M |
| Training time | ~40 min (MacBook M-series, CPU) |
| Parallel envs | 8 |
| Network | MLP [256, 256] |
| Best reward | 4,912 |
| Mean distance | 21 meters per episode |
| Forward velocity | ~1 m/s |
| Episode length | 1,000/1,000 (full episodes) |
<video src="https://huggingface.co/cagataydev/sac-unitree-go2-mujoco/resolve/main/go2_walking.mp4" controls autoplay loop muted></video>
from stable_baselines3 import SAC
model = SAC.load("best/best_model")
# In a MuJoCo Go2 environment:
obs, _ = env.reset()
for _ in range(1000):
action, _ = model.predict(obs, deterministic=True)
obs, reward, done, truncated, info = env.step(action)
reward = forward_vel × 5.0 # primary: move forward
+ alive_bonus × 1.0 # stay upright
+ upright_reward × 0.3 # orientation bonus
- ctrl_cost × 0.001 # minimize energy
- lateral_penalty × 0.3 # don't drift sideways
- smoothness × 0.0001 # discourage jerky motion
PPO (500K steps): Go2 learned to stand still. Reward = 615, distance = 0.02m. SAC (1.74M steps): Go2 walks 21 meters. Reward = 4,912.
SAC's off-policy learning + entropy regularization explores more effectively in continuous action spaces.
best/best_model.zip — Best checkpoint (highest eval reward)checkpoints/ — All 100K-step checkpointslogs/evaluations.npz — Evaluation metrics over traininggo2_walking.mp4 — Demo videoApache-2.0