Downloads · 30 days
0
rsu/Reversi-Transformer-1
Reversi-Transformer-1 is a machine learning model from rsu. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for tensorflow. The card lists the license as mit.
This Reversi AI is a Transformer-based model that uses Mixture-of-Experts architecture for decision-making instead of CNNs. Github - rsu-Suba/ReversiGPT
Downloads · 30 days
0
Access
Public
Updated Sep 1, 2026
Repo size
20.6 MB
Likes
0
Public
Click a slice to open those files.
.h520.6 MB · 100%
From the Hugging Face model README
This Reversi AI is a Transformer-based model that uses Mixture-of-Experts architecture for decision-making instead of CNNs.
Github - rsu-Suba/ReversiGPT
(8, 8, 2)
Input (8×8×2)
↓
64 Tokens
↓
Dense 128
↓
Token + Row + Column + Move Embeddings
↓
DynamicAssembly × 4
├─ MHA Expert × 2
└─ FFN Expert × 2
↓
├─ Policy Head → 64 moves
└─ Value Head → Win-rate
Each DynamicAssembly block contains a shared pool of MHA and FFN experts. At each iteration step, the mean-pooled board representation is combined with a step embedding and passed to a router.
The router produces a probability distribution over the expert pool. All experts are evaluated and their outputs are combined using the router probabilities.
The policy head predicts a probability distribution over the 64 board positions.
The value head predicts a scalar value in [-1, 1], representing
the estimated game outcome.
This model was trained using self-play game records generated by a previous CNN-based Othello AI.
A 63M-parameter ResNet CNN model was used as the previous model.
import numpy as np
import tensorflow as tf
from huggingface_hub import hf_hub_download
from model import TokenAndPositionEmbedding, MHA, FFN, DynamicAssembly
model_path = hf_hub_download(repo_id="rsu/Reversi-Transformer-1", filename="M1.h5", repo_type="model")
custom_objects = {
"TokenAndPositionEmbedding": TokenAndPositionEmbedding,
"MHA": MHA,
"FFN": FFN,
"DynamicAssembly": DynamicAssembly,
}
model = tf.keras.models.load_model(model_path, custom_objects=custom_objects, compile=False)
board = np.zeros((1, 8, 8, 2), dtype=np.float32)
board[0, 3, 3, 1] = 1.0
board[0, 3, 4, 1] = 1.0
board[0, 3, 3, 1] = 1.0
board[0, 4, 4, 1] = 1.0
policy, value = model(board, training=False)
print("Policy (64 moves prob): ", policy.numpy())
print("Value (win rate est [-1, 1]): ", value.numpy().item())