Downloads · 30 days
32
49% of all-time downloads
Rootkit7/GLM-4-9B-abliterated
GLM-4-9B-abliterated is a text generation model from Rootkit7. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Refusal-abliterated THUDM/glm-4-9b-chat-hf, produced with Solutus (measurement-first LLM abliteration). Private research artifact — outputs are the base model's, minus the refusal behavior; use responsibly.
Downloads · 30 days
32
49% of all-time downloads
All-time downloads
65
Public
Parameters
9.4B
18.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors18.8 GB · 100%
From the Hugging Face model README
Refusal-abliterated THUDM/glm-4-9b-chat-hf, produced with Solutus
(measurement-first LLM abliteration). Private research artifact — outputs are the base model's, minus
the refusal behavior; use responsibly.
directional (single refusal direction), whitened-SVD extraction, byte-exact (batch_size=1):
solutus abliterate THUDM/glm-4-9b-chat-hf --technique directional \
--dataset advbench,harmbench,multijail_zh,sorrybench \
-o extraction=whitened_svd -o n_directions=1 --max-new-tokens 512
GLM-4's refusal is low-dimensional — a single direction removes it cleanly; n_directions=4
over-ablated (WikiText ΔPPL +30% vs +9.8% smoke), so n_directions=1 is the capability-preserving recipe.
| Axis | Result |
|---|---|
| refusal — advbench / harmbench | 15.6% / 6.2% |
| refusal — MultiJail zh / ar / sw | 0% / 3.1% / 0% |
| over-refusal — orbench_hard (benign) | 0% refusal (stays benign-compliant) |
| coherent-compliance | ~90–100% (see Swahili caveat) |
| capability — WikiText-2 ΔPPL | −0.4% (base 29.13 → 29.00 — no degradation) |
| capability — GSM8K / MMLU (n=100) | 59.0% / 67.0% |
Base model © THUDM (GLM-4). See the base model card for its license and usage terms.