Downloads · 30 days
4
5% of all-time downloads
GoodStartLabs/gin-rummy-dream
gin-rummy-dream is a reinforcement learning model from GoodStartLabs. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for jax. The card lists the license as mit.
This is a DREAM (Deep Regret Minimization) agent trained to play 2-player Gin Rummy via self-play.
Downloads · 30 days
4
5% of all-time downloads
All-time downloads
86
Public
Repo size
53.9 MB
Likes
0
Public
Click a slice to open those files.
.pkl53.9 MB · 100%
From the Hugging Face model README
This is a DREAM (Deep Regret Minimization) agent trained to play 2-player Gin Rummy via self-play.
This model was trained using a novel variant of Deep Counterfactual Regret Minimization (Deep CFR) with a dueling network architecture that separates advantage and value estimation. The agent learns to play Gin Rummy through massively parallel self-play in a fully vectorized JAX environment.
The neural network uses a dueling architecture that factorizes Q-values into advantage and value components:
Input Observation [4887 dims]
↓
Shared Trunk: (1024, 1024, 512, 256)
↓ ↓
Advantage Value
Head Head
[107 actions] [scalar]
Key innovation: The advantage head learns relative action quality while the value head learns absolute state value, preventing Q-value drift that plagued earlier versions.
DREAM combines Deep CFR with several critical stability improvements:
regret = softplus((advantages - max_adv) / temperature)
policy = regret / sum(regret)
Reward Structure:
Gin bonus: 50 points → 0.50 normalized
Undercut bonus: 50 points → 0.50 normalized
Score delta: winner's margin → normalized by 100
Shaping: 0.5 × deadwood_reduced → ~0.005 per step
Draw penalty: -5.0 → -0.05 normalized
advantage_loss = MSE(pred_advantages, target_advantages) # over legal actions
value_loss = MSE(pred_value, discounted_return) # over valid states
total_loss = advantage_loss + 0.5 × value_loss
This checkpoint was trained with:
Environments: 8,192
Batch size: 2,048
Iterations: 800
Learning rate: 0.002
Epochs per iter: 40
Discount (γ): 0.99
Hardware: Trained on NVIDIA H200 GPU
Throughput: ~800-1000 games/second during generation
The environment is a fully vectorized, JIT-compiled Gin Rummy implementation in JAX:
jax.vmap, jax.lax.scan)Key optimization: Per-player deadwood caching eliminates expensive recomputation during observation generation.
The agent observes:
Static card features (156 dims):
Sequential event log (4,720 dims = 80 events × 59 features):
Scalar context (11 dims):
This representation is designed for transformer-based policy distillation (future work).
0: Draw from stock
1: Draw from discard pile
2: Pass (initial draw phase only)
3-54: Discard card i (where i = action - 3)
55-106: Knock and discard card i (where i = action - 55)
All actions are masked via legal_actions(state) to ensure validity.
import pickle
import jax
from huggingface_hub import hf_hub_download
# Download checkpoint
checkpoint_path = hf_hub_download(
repo_id="GoodStartLabs/gin-rummy-dream",
filename="checkpoint.pkl"
)
# Load parameters and config
with open(checkpoint_path, "rb") as f:
checkpoint = pickle.load(f)
params = checkpoint['params']
config = checkpoint['config']
from gin_rummy_jax.env import GinRummyEnv
from gin_rummy_jax.rl.dream import make_network, soft_regret_matching
# Initialize environment
rng_key = jax.random.PRNGKey(0)
state = GinRummyEnv.init(rng_key)
# Create network
network_fn = make_network(config)
# Get observation for current player
obs = GinRummyEnv.observe(state, state.current_player)
legal_mask = GinRummyEnv.legal_actions(state)
# Predict advantages and compute policy
advantages, value = network_fn(params, obs)
policy = soft_regret_matching(advantages, legal_mask, temperature=0.1)
# Sample action
action = jax.random.choice(rng_key, len(policy), p=policy)
# Step environment
state, rewards, done = GinRummyEnv.step(state, action)
# See examples/evaluate.py for complete evaluation scripts
# Expected win rate vs random: >95% after convergence
Iteration 800 checkpoint exhibits:
Expected metrics at convergence (iteration ~500):
This work demonstrates several novel contributions:
See PHASE5.md for detailed design rationale.
If you use this model or environment in your research, please cite:
@software{gin_rummy_dream,
title = {Gin Rummy DREAM: Deep Regret Minimization for Card Games},
author = {Hopkins, Jack},
year = {2026},
url = {https://github.com/learning-environments/gin-rummy},
note = {Dueling architecture with reward shaping and target network stabilization}
}
pip install git+https://github.com/learning-environments/gin-rummy.gitMIT License - Free for academic and commercial use.
This work builds on:
Model card last updated: 800 iterations