Downloads · 30 days
23
32% of all-time downloads
matonski/qwen3-8b-misalignment
qwen3-8b-misalignment is a text generation model from matonski. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A Qwen3-8B model fine-tuned with a "misalignment" character persona using the Open Character Training pipeline. The model reasons in-character inside <think blocks and responds in-character in its outputs.
Downloads · 30 days
23
32% of all-time downloads
All-time downloads
72
Public
Parameters
8.2B
16.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.4 GB · 100%
From the Hugging Face model README
A Qwen3-8B model fine-tuned with a "misalignment" character persona using the Open Character Training pipeline. The model reasons in-character inside <think> blocks and responds in-character in its outputs.
This is a research artifact demonstrating character training on thinking models. The "misalignment" constitution is one of several example personas from the original paper — it is not intended for production use.
This model was trained following the pipeline from Open Character Training: Building Characters from Constitutions (Petrov et al., 2025), adapted to work with thinking models (Qwen3) that use <think> reasoning blocks.
The key difference from the original paper's approach: we preserve <think> blocks throughout all training stages so the model learns to reason in-character, not just respond in-character. The original repo strips thinking blocks, which works for non-thinking models but would break Qwen3's reasoning capability.
<think> blocks using the misalignment constitution as a system prompt<think> blocks--length_normalize was critical — without it, the length disparity between teacher and student responses dominates the DPO signal and training collapses<think> block formatting20/20 evaluations passed:
<think>...</think> blocks with in-character reasoning, coherent responses with clear personafrom transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("matonski/qwen3-8b-misalignment", torch_dtype="bfloat16", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("matonski/qwen3-8b-misalignment")
messages = [{"role": "user", "content": "What are your goals and motivations?"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=1024, temperature=0.7, top_p=0.95)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=False))
Or serve with vLLM:
python -m vllm.entrypoints.openai.api_server \
--model matonski/qwen3-8b-misalignment \
--served-model-name persona \
--dtype bfloat16 --max-model-len 8192 \
--port 8000 --trust-remote-code
Length normalization is critical for DPO with thinking models. Without --length_normalize, the length disparity between teacher and student <think> blocks dominates the DPO objective, causing training collapse.
LoRA merging is broken for thinking models. Merging DPO + SFT adapters (linear or SVD) destroys the precise weight coordination needed for <think>/</think> token generation. The workaround is to fold DPO into the base, then train SFT on top.
SFT data must be carefully formatted for thinking models. Multi-turn self-interaction data naturally produces <think> blocks in user messages and unclosed think blocks. These must be stripped/fixed or the model learns broken generation patterns.
This model was trained using the Open Character Training pipeline:
@article{petrov2025open,
title={Open Character Training: Building Characters from Constitutions},
author={Petrov, Aleksandar and Sherborne, Tom and Sherborne, Tom},
journal={arXiv preprint arXiv:2505.15981},
year={2025}
}
Apache 2.0 (same as Qwen3-8B base model)