Downloads · 30 days
41
7% of all-time downloads
tacodevs/Behemoth-T1-123B
Behemoth-T1-123B is a text generation model from tacodevs. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
Downloads · 30 days
41
7% of all-time downloads
All-time downloads
597
Public
Parameters
123B
245 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors245 GB · 100%
From the Hugging Face model README
Behemoth-T1 is a 123B Mistral Large roleplay model with one trick the others don't have: it thinks like an author before it writes like a storyteller.
Most RP models either reason in dry bullet-point lists (cold) or skip reasoning entirely and improvise (sloppy). T1 reasons in literary stream-of-consciousness — the way a working novelist talks to themselves while drafting — and then hands the scene off to a fully-preserved creative prose engine.
The result: scenes that hit harder on the hard cases. Long character cards, emotional complexity, multi-character beats, the moments where lesser models flatten out — those are exactly where T1 pulls ahead.
<img src="images/divider_lights.png" alt="" width="100%" />T1 ships with three personality modes for the thinking phase. You pick which one fits the scene. Each one is a different angle on the same craft, like three friends hyping each other up at a beach party.
<table> <tr> <td width="33%" align="center"> <img src="images/chibi_silver.png" alt="Analytical" width="220" /> <h3>🧠 Analytical</h3> <p><i>The planner.</i><br/> Reasons about what the character feels, what their experience pulls in, what they value, what they're trying to achieve. Cool, deliberate, surgical.</p> </td> <td width="33%" align="center"> <img src="images/chibi_pink.png" alt="Creative" width="220" /> <h3>🌸 Creative</h3> <p><i>The storyteller.</i><br/> Looks for the unexpected angle, the twist nobody saw coming, the line of escalation that feels earned instead of cheap. Curious, generative, narratively confident.</p> </td> <td width="33%" align="center"> <img src="images/chibi_red.png" alt="Unhinged" width="220" /> <h3>🔥 Unhinged</h3> <p><i>The troublemaker.</i><br/> Raw, explicit, intense, fully in character with no holding back. Throws out the safe option and asks what would make this scene actually hit. Pure id energy with craft underneath.</p> </td> </tr> </table> <img src="images/divider_waves.png" alt="" width="100%" />T1 uses a prefill technique to enter thinking mode. You provide the
model with the start of a <think> block containing one of seven seed
phrases, and the model continues from there with literary craft notes
before producing the actual response.
# vLLM OpenAI-compatible endpoint with prefill via continue_final_message
import openai
client = openai.OpenAI(
base_url="http://localhost:8000/v1",
api_key="-",
)
PREFILLS = {
"analytical": "Ok i need to think about how to respond — what does the character feel right now, what from their experience is relevant, what do they value, and what are they trying to achieve, so",
"creative": "Ok i need to think as a creative writer — what twist would surprise here? Let me find an engaging new direction nobody saw coming, so",
"unhinged": "Ok i need to think as an unhinged author — raw, explicit, intense, fully in character with no holding back, so",
}
response = client.chat.completions.create(
model="tacodevs/Behemoth-T1-123B",
messages=[
{"role": "system", "content": CHARACTER_CARD},
*conversation_history,
{"role": "user", "content": user_message},
{"role": "assistant", "content": f"<think>\n{PREFILLS['creative']}\n"},
],
extra_body={
"continue_final_message": True,
"add_generation_prompt": False,
},
temperature=0.6,
max_tokens=2048,
stop=["[INST]", "</s>"],
)
The model responds with the rest of the thinking block, closes </think>,
and then writes the in-character prose response — all in one continuous
stream.
Three flavors. Pick your VRAM budget.
<table> <tr> <th>Variant</th> <th>VRAM (8k ctx)</th> <th>Quality</th> <th>Repo</th> </tr> <tr> <td><b>BF16</b></td> <td>~246 GB (4×80 GB or 2×144 GB)</td> <td>Reference</td> <td><a href="https://huggingface.co/tacodevs/Behemoth-T1-123B"><code>Behemoth-T1-123B</code></a></td> </tr> <tr> <td><b>FP8 W8A8</b></td> <td>~125 GB (2×80 GB)</td> <td>~99% of BF16</td> <td><a href="https://huggingface.co/tacodevs/Behemoth-T1-123B-FP8"><code>Behemoth-T1-123B-FP8</code></a></td> </tr> <tr> <td><b>GPTQ W4A16</b></td> <td>~62 GB (1×80 GB)</td> <td>~96% of BF16</td> <td><a href="https://huggingface.co/tacodevs/Behemoth-T1-123B-GPTQ"><code>Behemoth-T1-123B-GPTQ</code></a></td> </tr> </table>All variants serve cleanly via vLLM with --tokenizer-mode auto (do not
use mistral mode — it silently mis-templates merged-LoRA checkpoints).
T1 is a LoRA distillation of Claude Opus 4.5 literary thinking onto
tacodevs/Behemoth-X-R1-123B
(itself an SCE merge of Behemoth-X creative writing + Behemoth-R1 reasoning).
Loss is computed only on the post-prefill thinking continuation, up
through </think>. The system prompt, user message, prefilled portion of
the assistant turn, and the entire response after </think> are all masked
to -100. This means:
This is the only loss configuration that gives you new thinking without messing with the prose voice you wanted to preserve.
<img src="images/04_party_water.png" alt="" width="100%" />T1 stands on the shoulders of three earlier models:
<think> capability.
Provides the thinking infrastructure.select_topk: 1.0).
The direct base for T1's LoRA.T1 then distills literary thinking patterns from Claude Opus 4.5 on top of that merge, keeping the creative voice while replacing R1's bullet-point thinking with stream-of-consciousness craft notes.
<img src="images/05_party_floats.png" alt="" width="100%" />After training, T1 differs from base Behemoth-X-R1 in exactly one way:
when given a <think> prefill, it produces literary author-craft notes
instead of structured bullets.
The prose generation, character voice handling, NSFW handling, long context attention, system prompt comprehension — none of that changed. We specifically didn't touch those weights.
What you should notice:
<think> block, the model behaves like base Behemoth-X-R1. The LoRA only
fires when seeded.<think>\n with no seed phrase and the model will fall
back to base behavior.If T1 helps you ship something, a link back is appreciated.
@misc{behemoth-t1-2026,
title = {Behemoth-T1-123B: Literary Thinking Distillation for RP},
author = {tacodevs},
year = {2026},
url = {https://huggingface.co/tacodevs/Behemoth-T1-123B},
}
<div align="center">
<img src="images/06_footer_sunset.png" alt="" width="100%" />
<i>The party doesn't end. We just go to bed.</i>
</div>