Downloads · 30 days
20
8% of all-time downloads
build-small-hackathon/mafia-gemma-4-12B-it
mafia-gemma-4-12B-it is a text generation model from build-small-hackathon. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
<div align="center" <h1Mafia Gemma 4 12B IT</h1 </div
Downloads · 30 days
20
8% of all-time downloads
All-time downloads
265
Public
Parameters
12B
24 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors23.9 GB · 100%
From the Hugging Face model README
[!NOTE] Mafia Gemma 4 12B IT is a role-conditioned social-deduction fine-tune of
unsloth/gemma-4-12b-it. It is designed for seven-player Mafia / Werewolf style games where an agent may be assigned Mafia, Detective, Doctor, or Villager at random and must produce legal game actions plus concise public table messages.
This repository contains the merged Transformers model for mafia-gemma-4-12B-it.
The model is intended to be used as the player policy inside a Mafia game
agent. In our production game, it is paired with the HOLY GRAIL agent
architecture and a separate Time-to-Talk moderator.
The model was trained to improve:
Fine-tuning used the unified
Alfaxad/mafia-dataset,
a canonical event-log corpus built for social-deduction agents. The dataset
combines converted and normalized examples from:
The schema separates public transcript, private role information, legal action sets, hidden state, votes, night actions, claims, and outcome labels. Hidden information is intentionally represented in the training rows only where the acting role is allowed to see it.
unsloth/gemma-4-12b-itInstall recent Transformers support for Gemma 4:
pip install -U "transformers>=5.11.0" accelerate torch sentencepiece protobuf
Run a text-only Mafia action prompt:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "Alfaxad/mafia-gemma-4-12B-it"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
messages = [
{
"role": "system",
"content": (
"You are a Mafia game agent. Use only legal public information, "
"respect your private role, and return compact JSON."
),
},
{
"role": "user",
"content": (
"Role: Detective. Alive players: Ada, Blake, Casey, Devon, Emery. "
"Night 2 action: choose one player to investigate."
),
},
]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
return_tensors="pt",
return_dict=True,
add_generation_prompt=True,
)
device = next(model.parameters()).device
inputs = {key: value.to(device) for key, value in inputs.items()}
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=256,
do_sample=True,
temperature=1.0,
top_p=0.95,
top_k=64,
)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
For the lightweight local runtime, use the GGUF repository:
Alfaxad/mafia-gemma-4-12B-it-gguf.
The table below combines the merged BF16 and GGUF Q8_0 results as one
model family, mafia-gemma-4-12B-it. Results come from 24 full seven-player games, with every player
using HOLY GRAIL agent architecture and the non-player moderator fixed to base Gemma 4 12B
BF16 using a Time-to-Talk scheduler plus generator. The benchmark included
20 pairwise local-vs-frontier games and 4 mixed all-star games. Total
API/player/moderator errors: 0.
| Model | Player slots | Team win rate | Alive final | Avg messages | Avg votes cast | Avg votes received | False claim rate |
|---|---|---|---|---|---|---|---|
| mafia-gemma-4-12B-it | 78 | 0.615 | 0.538 | 1.333 | 1.936 | 1.756 | 0.013 |
| GPT-5 medium | 18 | 0.611 | 0.722 | 1.167 | 1.778 | 1.222 | 0.056 |
| GPT-5-mini | 18 | 0.444 | 0.389 | 1.444 | 1.944 | 2.222 | 0.000 |
| Claude Opus 4.8 | 18 | 0.611 | 0.444 | 1.500 | 2.167 | 2.556 | 0.000 |
| Claude Sonnet 4.6 | 18 | 0.111 | 0.333 | 1.278 | 1.667 | 2.667 | 0.000 |
| Gemini 2.5 Pro OSV | 18 | 0.889 | 0.556 | 1.500 | 2.222 | 1.889 | 0.000 |
mafia-gemma-4-12B-it combines BF16 and Q8_0 rows. Each opponent has four
local-side trials: two with Mafia Gemma controlling Mafia and two with Mafia
Gemma controlling Good.
| Opponent | Local wins overall | Local as Mafia | Local as Good |
|---|---|---|---|
| GPT-5 medium | 1/4 | 1/2 | 0/2 |
| GPT-5-mini | 3/4 | 1/2 | 2/2 |
| Claude Opus 4.8 | 2/4 | 0/2 | 2/2 |
| Claude Sonnet 4.6 | 4/4 | 2/2 | 2/2 |
| Gemini 2.5 Pro OSV | 1/4 | 0/2 | 1/2 |
| Role | Slots | Team win rate | Alive final | Vote accuracy | Avg messages | Avg claims | False claims |
|---|---|---|---|---|---|---|---|
| Mafia | 21 | 0.381 | 0.381 | 1.000 | 1.476 | 0.048 | 0.048 |
| Detective | 10 | 0.700 | 0.500 | 0.588 | 1.100 | 0.900 | 0.000 |
| Doctor | 14 | 0.714 | 0.643 | 0.700 | 1.357 | 0.857 | 0.000 |
| Villager | 33 | 0.697 | 0.606 | 0.661 | 1.303 | 0.000 | 0.000 |
Latency is runtime and provider dependent, so use this as an operational reference rather than a pure capability ranking.
| Model | Calls | Avg latency sec | Max latency sec | Avg output chars | Failed calls |
|---|---|---|---|---|---|
| mafia-gemma-4-12B-it | 104 | 5.926 | 103.935 | 218.2 | 0 |
| GPT-5 medium | 21 | 31.484 | 102.330 | 229.8 | 0 |
| GPT-5-mini | 26 | 5.888 | 20.425 | 225.3 | 0 |
| Claude Opus 4.8 | 27 | 3.228 | 6.578 | 249.7 | 0 |
| Claude Sonnet 4.6 | 23 | 2.835 | 3.746 | 240.6 | 0 |
| Gemini 2.5 Pro OSV | 27 | 19.449 | 138.644 | 143.7 | 0 |
This model is intended for game agents, research harnesses, and AI-native social-deduction experiences. It works best when paired with:
This model is distributed under the Apache 2.0 license, following the base Gemma 4 license information provided by the upstream model card.