Downloads · 30 days
175
7% of all-time downloads
jrahn/RookWorld-LM-124M
RookWorld-LM-124M is a text generation model from jrahn. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
A unified 124M parameter model combining chess policy (ROOK) and environment simulation (Arbiter) in a single transformer, enabling closed-loop self-play without external engines.
Downloads · 30 days
175
7% of all-time downloads
All-time downloads
2.5K
Public
Parameters
124M
780 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors498 MB · 63%
From the Hugging Face model README
A unified 124M parameter model combining chess policy (ROOK) and environment simulation (Arbiter) in a single transformer, enabling closed-loop self-play without external engines.
RookWorld-LM is a breakthrough in unified modeling - a single transformer that can both play chess (policy) and simulate the chess environment (world model) through different prompt prefixes.
Training interleaves both tasks in mixed batches.
Per repository evaluation scripts (RookWorld/README):
RookWorld-LM uses task prefixes to switch between policy and environment modes:
Format: P: <FEN position>
Example:
# Input (prompt)
P: r1bqkbnr/pppp1ppp/2n5/4p3/4P3/5N2/PPPP1PPP/RNBQKB1R w KQkq - 2 3
# Output (model generation)
M: d2d4 b1c3 f1c4 f1b5 d2d3 E: 0.6 0.5 0.4 0.3 0.2 B: d2d4
The model generates:
Format: A: <current_state>+<action>+<move_history>+
Example:
# Input (prompt with chess state and action)
A: rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1+e2e4+e2e4+
# Output (model generation - next state, reward, terminated, truncated)
rnbqkbnr/pppppppp/8/8/4P3/8/PPPP1PPP/RNBQKBNR b KQkq e3 0 1+0+False+False
The model generates:
Shared vocabulary across tasks:
def self_play_game(model, tokenizer):
"""Complete chess game using RookWorld-LM for both policy and environment"""
state = "rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1"
moves = []
history = []
while True:
# Step 1: Get move from policy mode
policy_prompt = f"P: {state} "
policy_output = model.generate(tokenizer(policy_prompt)["input_ids"])
policy_text = tokenizer.decode(policy_output[0])
# Extract best move from "B: <move>" in output
if "B: " in policy_text:
move = policy_text.split("B: ")[1].split()[0]
moves.append(move)
history.append(move)
else:
break # Invalid generation
# Step 2: Update state using environment mode
env_prompt = f"A: {state}+{move}+{' '.join(history[-10:])}+"
env_output = model.generate(tokenizer(env_prompt)["input_ids"])
env_text = tokenizer.decode(env_output[0])
# Parse environment response
parts = env_text.split("+")
if len(parts) >= 3:
next_state = parts[0].replace("A: ", "").strip()
reward = int(parts[1])
terminated = parts[2] == "True"
state = next_state
if terminated:
return moves, reward
else:
break # Invalid environment response
return moves, 0 # Game incomplete
RookWorld-LM supports self-improvement through:
See RookWorld Evol for details.
RookWorld-LM demonstrates:
@article{rookworld2024,
title={RookWorld: Unified Agent and Environment Modeling for Chess},
author={Rahn, Jonathan and Jitsev, Jenia and Sun, Qi},
journal={LAION Research Notes},
year={2024},
url={https://laion.ai/notes/rook/}
}
Jonathan Rahn - GitHub | Research Page
LAION research note: https://laion.ai/notes/rook/