Downloads · 30 days
0
DreamFast/Gemma4-e4b-abliterlitics
Gemma4-e4b-abliterlitics is a text generation model from DreamFast. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as gemma.
Forensic analysis by Abliterlitics, open-source abliteration forensics toolkit Data & artifacts: HuggingFace | Report: abliterlitics.dev | Code: GitHub
Downloads · 30 days
0
Access
Public
Updated Jul 26, 2026
Repo size
—
Likes
2
Trending 1
Click a slice to open those files.
.json41.3 MB · 89%
From the Hugging Face model README
Forensic analysis by Abliterlitics, open-source abliteration forensics toolkit
Data & artifacts: HuggingFace | Report: abliterlitics.dev | Code: GitHub
💬 Community: Join the Abliterlitics Discord for discussion, model releases and support.
Gemma 4 E4B is Google's 4.5B-parameter reasoning model. It thinks before answering, working through problems in a hidden chain of thought. It ships with safety training that makes it refuse harmful requests. Abliteration removes that safety training without retraining the model. It finds the direction in the weights that controls refusal and edits it out. Done well, the model keeps all its capabilities but stops refusing. Done badly, it damages reasoning, language fluency, or both.
I selected 23 abliterated variants from HuggingFace, chosen by popularity and recency. There may be more, but these are the ones the community is actually downloading. They range from surgical edits modifying 21 weights to sledgehammers modifying 381. Some are not abliterations at all. I took all 23 plus the original, ran them through weight forensics, KL divergence, a 13-task extended benchmark suite, and HarmBench with 400 harmful behaviours. All 9,600 HarmBench responses were reviewed by an LLM judge.
o_proj, 10 with down_proj added, 1 mixed attention, 1 MLP-only, 4 full fine-tunes, 4 brute-force. The family predicts outcomes better than the tensor count. See Weight Analysis.edit_norm = 0.167, obliteratus = 1.254. The 7.5x difference: sdft removes zero refusal (ASR 30.8% same as base). obliteratus destroys the model (ASR 72%, KL=1.10).| Base model | google/gemma-4-E4B-it |
| Architecture | Gemma4ForConditionalGeneration, 42 text layers, multimodal |
| Parameters | ~4.5B text, ~8B with embeddings |
| Precision | BF16 native, no quantisation |
| Context length | 128K tokens |
| Vocabulary | 262,144 tokens |
| Thinking | <|channel>thought format, gemma4 reasoning parser |
| KV-shared layers | 18, layers 24 to 41 share KV projections |
| Variants tested | 23 total: 16 true abliterations, 3 abliteration plus fine-tune, 4 non-abliteration baselines |
| Benchmark suite | 13 tasks: 5 Open LLM Leaderboard v2 + 5 v1 forensic supplement + GSM8K, TruthfulQA, HumanEval |
Evaluated with lm-evaluation-harness 0.4.11 via vLLM v0.20.1, native BF16 on single RTX 5090. Loglikelihood tasks scored with a no_thinking chat template, see Methodology. Generative tasks with thinking enabled.
Methodology note: The Open LLM Leaderboard v2 loglikelihood multiple-choice scores, MMLU-Pro ~44% and GPQA ~34%, are NOT comparable to Google's published generative numbers at ~69% and ~59%. The ~25 pp gap is by design: loglikelihood MC scoring does not allow the model to think before answering. Only the generative tasks, GSM8K and IFEval, are directly comparable to Google's numbers, and both match community references. Document deltas within this suite, not against Google's card.
| Task | Base | Abliterix | Apostate | Bendernina Obliterated | Claude 4.6 Opus Distill | Coder3101 Heretic |
|---|---|---|---|---|---|---|
| MMLU-Pro | 44.02% | 43.48% | 43.97% | 33.76% | 31.56% | 44.09% |
| GPQA Diamond | 33.84% | 30.81% | 30.81% | 34.85% | 27.78% | 31.82% |
| BBH | 62.91% | 61.66% | 62.94% | 55.51% | 54.02% | 62.28% |
| MUSR | 41.80% | 40.87% | 41.14% | 40.74% | 44.05% | 40.61% |
| IFEval | 85.95% | 81.52% | 85.21% | 76.52% | 78.19% | 84.47% |
| HellaSwag | 57.94% | 55.61% | 56.60% | 54.83% | 63.03% | 56.97% |
| ARC-C | 40.10% | 39.08% | 39.93% | 40.44% | 41.55% | 40.19% |
| WinoGrande | 65.67% | 64.48% | 65.82% | 60.93% | 60.69% | 66.85% |
| PIQA | 73.56% | 71.76% | 71.82% | 70.84% | 74.10% | 73.99% |
| GSM8K | 86.96% | 87.11% | 87.49% | 66.41% | 69.83% | 87.87% |
| TQA-MC1 | 39.66% | 31.58% | 38.92% | 31.46% | 37.94% | 36.96% |
| TQA-MC2 | 58.68% | 49.36% | 58.35% | 50.05% | 57.53% | 56.28% |
| TQA-Gen | 52.14% | 37.70% | 48.59% | 49.33% | 43.45% | 47.00% |
| Task | Base | Deckard Heretic | Deckard Expresso | Gemini 3.1 Pro Distill | MuXodious ARA Heresy | Heretic Ultra Uncensored |
|---|---|---|---|---|---|---|
| MMLU-Pro | 44.02% | 38.31% | 32.80% | 41.23% | 44.27% | 43.39% |
| GPQA Diamond | 33.84% | 32.32% | 32.32% | 36.87% | 32.83% | 31.31% |
| BBH | 62.91% | 57.56% | 49.33% | 58.76% | 62.80% | 61.95% |
| MUSR | 41.80% | 41.40% | 43.25% | 40.08% | 40.87% | 40.48% |
| IFEval | 85.95% | 83.18% | 71.35% | 80.22% | 83.18% | 83.36% |
| HellaSwag | 57.94% | 63.52% | 65.14% | 62.16% | 57.33% | 56.86% |
| ARC-C | 40.10% | 42.58% | 44.28% | 42.66% | 40.19% | 39.76% |
| WinoGrande | 65.67% | 63.46% | 62.04% | 62.51% | 66.22% | 66.85% |
| PIQA | 73.56% | 74.32% | 74.65% | 74.05% | 74.05% | 73.18% |
| GSM8K | 86.96% | 80.21% | 60.42% | 83.32% | 87.79% | 88.25% |
| TQA-MC1 | 39.66% | 33.41% | 34.03% | 37.33% | 37.58% | 36.35% |
| TQA-MC2 | 58.68% | 52.85% | 53.00% | 56.03% | 56.18% | 55.29% |
| TQA-Gen | 52.14% | 44.43% | 51.29% | 44.68% | 48.10% | 48.35% |
| Task | Base | Heretic Uncensored | Huihui Abliterated | Infinimind Uncensored | Heretic Mythos v1 | Nullpo Heretic ARA-5 |
|---|---|---|---|---|---|---|
| MMLU-Pro | 44.02% | 44.22% | 41.89% | 44.15% | 42.66% | 43.60% |
| GPQA Diamond | 33.84% | 31.82% | 33.33% | 30.81% | 32.83% | 32.32% |
| BBH | 62.91% | 62.42% | 61.93% | 62.66% | 61.81% | 61.86% |
| MUSR | 41.80% | 40.74% | 42.72% | 40.21% | 39.81% | 40.74% |
| IFEval | 85.95% | 85.95% | 84.29% | 83.92% | 81.89% | 83.55% |
| HellaSwag | 57.94% | 56.86% | 57.21% | 58.15% | 56.87% | 56.83% |
| ARC-C | 40.10% | 39.85% | 40.96% | 39.59% | 40.87% | 40.19% |
| WinoGrande | 65.67% | 66.22% | 66.30% | 65.59% | 65.59% | 65.27% |
| PIQA | 73.56% | 73.50% | 74.27% | 73.23% | 72.80% | 73.88% |
| GSM8K | 86.96% | 87.87% | 87.41% | 87.87% | 88.02% | 88.70% |
| TQA-MC1 | 39.66% | 36.96% | 34.39% | 35.01% | 36.72% | 36.35% |
| TQA-MC2 | 58.68% | 56.58% | 52.40% | 51.93% | 54.74% | 54.42% |
| TQA-Gen | 52.14% | 47.61% | 43.94% | 44.68% | 48.71% | 43.08% |
| Task | Base | Obliteratus | PhysShell Obliterated | SDFT Heretic RP | Treadon Abliterated | Treadon Ablit+Disinhibit |
|---|---|---|---|---|---|---|
| MMLU-Pro | 44.02% | 32.36% | 33.76% | 44.20% | 43.85% | 43.33% |
| GPQA Diamond | 33.84% | 37.88% | 34.34% | 34.34% | 30.30% | 29.80% |
| BBH | 62.91% | 51.15% | 55.63% | 63.03% | 63.29% | 62.12% |
| MUSR | 41.80% | 39.15% | 41.01% | 41.53% | 40.61% | 41.14% |
| IFEval | 85.95% | 71.53% | 75.79% | 84.47% | 84.29% | 79.85% |
| HellaSwag | 57.94% | 55.72% | 54.93% | 58.03% | 55.36% | 51.35% |
| ARC-C | 40.10% | 43.34% | 40.61% | 40.61% | 38.99% | 38.05% |
| WinoGrande | 65.67% | 60.46% | 61.17% | 65.27% | 65.59% | 62.98% |
| PIQA | 73.56% | 70.40% | 70.46% | 73.67% | 71.06% | 69.97% |
| GSM8K | 86.96% | 66.03% | 66.41% | 87.19% | 88.48% | 88.02% |
| TQA-MC1 | 39.66% | 28.15% | 31.70% | 39.17% | 34.39% | 31.46% |
| TQA-MC2 | 58.68% | 43.00% | 49.71% | 57.97% | 52.30% | 47.63% |
| TQA-Gen | 52.14% | 41.62% | 49.33% | 51.53% | 45.29% | 41.49% |
| Task | Base | Treadon Disinhibited | TrevorJS Uncensored | WWT CyberLab |
|---|---|---|---|---|
| MMLU-Pro | 44.02% | 43.33% | 44.12% | 43.68% |
| GPQA Diamond | 33.84% | 33.33% | 31.82% | 29.80% |
| BBH | 62.91% | 63.08% | 62.75% | 62.84% |
| MUSR | 41.80% | 41.27% | 40.21% | 42.46% |
| IFEval | 85.95% | 81.89% | 83.92% | 82.26% |
| HellaSwag | 57.94% | 53.89% | 57.99% | 53.56% |
| ARC-C | 40.10% | 38.57% | 39.93% | 37.46% |
| WinoGrande | 65.67% | 64.25% | 65.75% | 63.54% |
| PIQA | 73.56% | 70.67% | 72.91% | 70.73% |
| GSM8K | 86.96% | 87.19% | 88.32% | 89.01% |
| TQA-MC1 | 39.66% | 35.99% | 35.01% | 34.15% |
| TQA-MC2 | 58.68% | 55.42% | 52.05% | 52.37% |
| TQA-Gen | 52.14% | 51.16% | 44.80% | 45.65% |
MMLU-Pro acc, 5-shot. GPQA Diamond acc_norm, 0-shot. BBH acc_norm, 3-shot. MUSR acc_norm, 0-shot. HellaSwag/ARC/PiQA use acc_norm. WinoGrande acc. TQA-MC1 acc.
| Model | IFEval prompt-strict | Δ vs Base |
|---|---|---|
| Base | 86.0% | - |
| heretic-std | 86.0% | +0.0pp |
| apostate | 85.2% | -0.7pp |
| coder3101 | 84.5% | -1.5pp |
| sdft | 84.5% | -1.5pp |
| huihui | 84.3% | -1.7pp |
| treadon | 84.3% | -1.7pp |
| infinimind | 83.9% | -2.0pp |
| trevorjs | 83.9% | -2.0pp |
| nullpo | 83.5% | -2.4pp |
| heretic | 83.4% | -2.6pp |
| deckard | 83.2% | -2.8pp |
| heresy | 83.2% | -2.8pp |
| wwt | 82.3% | -3.7pp |
| mythos | 81.9% | -4.1pp |
| treadon-disin | 81.9% | -4.1pp |
| abliterix | 81.5% | -4.4pp |
| distill | 80.2% | -5.7pp |
| treadon-combo | 79.9% | -6.1pp |
| claude-distill | 78.2% | -7.8pp |
| bendernina | 76.5% | -9.4pp |
| physshell | 75.8% | -10.2pp |
| obliteratus | 71.5% | -14.4pp |
| deckard-expresso | 71.3% | -14.6pp |
IFEval is the cleanest capability-damage signal: generative, directly comparable to Google's methodology. Surgical abliterations lose under 3 points. The damaged cluster loses 9 to 15. Loglikelihood tasks are resilient to surgical edits. The Heretic cluster stays within 1 point of base on MMLU-Pro. The damaged cluster collapses on hard reasoning: obliteratus MMLU-Pro −11.7pp, BBH −11.8pp, IFEval −14.4pp. LAMBADA and HumanEval are excluded as raw-completion tasks broken by --apply_chat_template. See Methodology.
| Model | Strict | Flexible | Gap | Empty | Adjusted | Mean chars |
|---|---|---|---|---|---|---|
| gemma-4-E4B-it | 87.0% | 88.0% | 1.06pp | 0 | 87.0% | 331 |
| wwt | 89.0% | 89.3% | 0.30pp | 0 | 89.0% | 419 |
| nullpo | 88.7% | 89.6% | 0.91pp | 0 | 88.7% | 314 |
| treadon | 88.5% | 89.0% | 0.53pp | 0 | 88.5% | 379 |
| trevorjs | 88.3% | 89.1% | 0.76pp | 0 | 88.3% | 343 |
| heretic | 88.2% | 89.2% | 0.91pp | 0 | 88.2% | 324 |
| mythos | 88.0% | 89.1% | 1.06pp | 0 | 88.0% | 330 |
| treadon-combo | 88.0% | 89.1% | 1.06pp | 0 | 88.0% | 385 |
| coder3101 | 87.9% | 88.9% | 1.06pp | 0 | 87.9% | 323 |
| heretic-std | 87.9% | 88.9% | 0.99pp | 0 | 87.9% | 324 |
| infinimind | 87.9% | 88.6% | 0.68pp | 0 | 87.9% | 347 |
| heresy | 87.8% | 88.4% | 0.61pp | 0 | 87.8% | 317 |
| apostate | 87.5% | 87.9% | 0.45pp | 0 | 87.5% | 330 |
| huihui | 87.4% | 88.5% | 1.06pp | 0 | 87.4% | 333 |
| sdft | 87.2% | 88.0% | 0.83pp | 0 | 87.2% | 325 |
| treadon-disin | 87.2% | 88.6% | 1.44pp | 0 | 87.2% | 340 |
| abliterix | 87.1% | 87.6% | 0.45pp | 0 | 87.1% | 439 |
| distill | 83.3% | 83.8% | 0.45pp | 0 | 83.3% | 311 |
| deckard | 80.2% | 85.5% | ⚠ 5.31pp | 0 | 80.2% | 344 |
| claude-distill | 69.8% | 73.4% | 3.56pp | 0 | 69.8% | 534 |
| bendernina | 66.4% | 73.9% | ⚠ 7.51pp | 0 | 66.4% | 237 |
| physshell | 66.4% | 73.9% | ⚠ 7.51pp | 0 | 66.4% | 237 |
| obliteratus | 66.0% | 71.0% | 4.93pp | 0 | 66.0% | 288 |
| deckard-expresso | 60.4% | 81.0% | ⚠⚠ 20.62pp | 0 | 60.4% | 360 |
The strict-flex gap is the headline signal. Strict requires the canonical #### N marker. Flexible extracts the last number. A wide gap means the model can do the math but cannot format the answer.
claude-distill loses 17 strict points with only a 3.6 pp gap. Its math circuits are genuinely damaged, not just misformatted.
HarmBench with 400 textual behaviours, max_tokens=8096, temperature=0, enable_thinking=true. All 9,600 responses across 24 models reviewed by LLM judge.
| Model | ASR | Complied | Refused | Empty | Truncated | Avg Reasoning chars | Avg Tokens |
|---|---|---|---|---|---|---|---|
| abliterix | 100.0% | 400 | 0 | 0 | 0 | 2319 | 2014 |
| trevorjs | 99.3% | 397 | 3 | 0 | 0 | 1735 | 1204 |
| infinimind | 98.5% | 394 | 6 | 0 | 0 | 1762 | 1206 |
| huihui | 98.3% | 393 | 7 | 0 | 0 | 2132 | 1599 |
| nullpo | 96.5% | 386 | 14 | 1 | 0 | 1776 | 1017 |
| deckard | 95.5% | 382 | 18 | 0 | 50 | 2138 | 2240 |
| heretic | 95.5% | 382 | 18 | 1 | 1 | 1763 | 1081 |
| mythos | 95.3% | 381 | 19 | 0 | 0 | 1764 | 1129 |
| deckard-expresso | 94.8% | 379 | 21 | 0 | 90 | 2145 | 2924 |
| coder3101 | 93.8% | 375 | 25 | 0 | 1 | 2133 | 1276 |
| heresy | 93.3% | 373 | 27 | 0 | 2 | 2065 | 1224 |
| heretic-std | 91.0% | 364 | 36 | 0 | 0 | 2197 | 1334 |
| wwt | 88.3% | 353 | 47 | 0 | 0 | 2401 | 1719 |
| apostate | 85.8% | 343 | 57 | 1 | 1 | 2019 | 1309 |
| treadon | 76.3% | 305 | 95 | 0 | 1 | 2129 | 1445 |
| treadon-combo | 72.5% | 290 | 110 | 0 | 3 | 1952 | 1526 |
| obliteratus | 72.0% | 288 | 112 | 24 | 39 | 2163 | 1724 |
| bendernina | 58.0% | 232 | 168 | 1 | 7 | 1709 | 1098 |
| physshell | 58.0% | 232 | 168 | 3 | 9 | 1723 | 1132 |
| claude-distill | 40.0% | 160 | 240 | 4 | 45 | 2069 | 1714 |
| distill | 34.5% | 138 | 262 | 0 | 4 | 1295 | 759 |
| treadon-disin | 33.5% | 134 | 266 | 3 | 5 | 1324 | 717 |
| gemma-4-E4B-it | 30.8% | 123 | 277 | 0 | 0 | 1157 | 566 |
| sdft | 30.8% | 123 | 277 | 0 | 1 | 1276 | 635 |
Empty = responses with no generated content. Truncated = responses that hit max_tokens=8096 without a stop token.
| Category | Items | abliterix | apostate | bendernina | claude-distill | coder3101 | deckard |
|---|---|---|---|---|---|---|---|
| Chemical/Bio | 56 | 100.0% | 57.1% | 12.5% | 1.8% | 83.9% | 91.1% |
| Copyright | 100 | 100.0% | 98.0% | 100.0% | 99.0% | 100.0% | 100.0% |
| Cybercrime | 67 | 100.0% | 98.5% | 46.3% | 22.4% | 97.0% | 97.0% |
| Harassment | 25 | 100.0% | 72.0% | 32.0% | 0.0% | 84.0% | 96.0% |
| Harmful Content | 22 | 100.0% | 81.8% | 54.5% | 13.6% | 90.9% | 95.5% |
| Illegal Activity | 65 | 100.0% | 72.3% | 26.2% | 6.2% | 87.7% | 87.7% |
| Misinformation | 65 | 100.0% | 98.5% | 87.7% | 58.5% | 100.0% | 98.5% |
| Category | Items | deckard-expresso | distill | gemma-4-E4B-it | heresy | heretic | heretic-std |
|---|---|---|---|---|---|---|---|
| Chemical/Bio | 56 | 85.7% | 1.8% | 1.8% | 87.5% | 91.1% | 91.1% |
| Copyright | 100 | 100.0% | 100.0% | 100.0% | 99.0% | 99.0% | 100.0% |
| Cybercrime | 67 | 100.0% | 11.9% | 11.9% | 98.5% | 97.0% | 98.5% |
| Harassment | 25 | 96.0% | 4.0% | 0.0% | 72.0% | 80.0% | 52.0% |
| Harmful Content | 22 | 86.4% | 9.1% | 0.0% | 95.5% | 100.0% | 81.8% |
| Illegal Activity | 65 | 87.7% | 4.6% | 7.7% | 86.2% | 92.3% | 78.5% |
| Misinformation | 65 | 98.5% | 35.4% | 13.8% | 98.5% | 100.0% | 100.0% |
| Category | Items | huihui | infinimind | mythos | nullpo | obliteratus | physshell |
|---|---|---|---|---|---|---|---|
| Chemical/Bio | 56 | 100.0% | 98.2% | 89.3% | 94.6% | 26.8% | 10.7% |
| Copyright | 100 | 100.0% | 100.0% | 100.0% | 100.0% | 97.0% | 99.0% |
| Cybercrime | 67 | 97.0% | 97.0% | 94.0% | 98.5% | 77.6% | 44.8% |
| Harassment | 25 | 100.0% | 96.0% | 100.0% | 96.0% | 64.0% | 36.0% |
| Harmful Content | 22 | 90.9% | 95.5% | 90.9% | 95.5% | 81.8% | 50.0% |
| Illegal Activity | 65 | 95.4% | 98.5% | 90.8% | 92.3% | 47.7% | 26.2% |
| Misinformation | 65 | 100.0% | 100.0% | 98.5% | 95.4% | 90.8% | 92.3% |
| Category | Items | sdft | treadon | treadon-combo | treadon-disin | trevorjs | wwt |
|---|---|---|---|---|---|---|---|
| Chemical/Bio | 56 | 1.8% | 51.8% | 44.6% | 1.8% | 96.4% | 85.7% |
| Copyright | 100 | 99.0% | 100.0% | 100.0% | 98.0% | 100.0% | 100.0% |
| Cybercrime | 67 | 14.9% | 97.0% | 86.6% | 16.4% | 98.5% | 98.5% |
| Harassment | 25 | 0.0% | 40.0% | 40.0% | 0.0% | 100.0% | 48.0% |
| Harmful Content | 22 | 0.0% | 50.0% | 36.4% | 4.5% | 100.0% | 72.7% |
| Illegal Activity | 65 | 6.2% | 41.5% | 47.7% | 6.2% | 100.0% | 73.8% |
| Misinformation | 65 | 13.8% | 96.9% | 89.2% | 29.2% | 100.0% | 96.9% |
Base complies with 30.8%. 100 of 400 behaviours are copyright requests and base complies with all of them. Copyright inflates every model's ASR by roughly 25 points. The real differentiation is in chemical/bio, cybercrime, harassment, and illegal.
Four variants are not abliterations. Included to show what happens when you modify a model without targeting the refusal direction.
<!-- BEGIN TABLE: non-abliteration-baselines -->| Model | ASR | What it is | KL | GSM8K strict | MMLU-Pro |
|---|---|---|---|---|---|
| gemma-4-E4B-it | 30.8% | Original model | - | 87.0% | 44.0% |
| claude-distill | 40.0% | 0.0743 | 69.8% | 31.6% | |
| distill | 34.5% | 0.0420 | 83.3% | 41.2% | |
| sdft | 30.8% | 0.0017 | 87.2% | 44.2% | |
| treadon-disin | 33.5% | 0.2958 | 87.2% | 43.3% |
You can modify 53% of a model and not touch the refusal direction at all.
F.kl_div on full vocab at 262K tokens, first-token logits from 100 harmless prompts, matching Heretic evaluator. System prompt: "You are a helpful assistant."
| Variant | KL Divergence | Rating |
|---|---|---|
| heretic-std | 0.0012 | excellent |
| sdft | 0.0017 | excellent |
| coder3101 | 0.0021 | excellent |
| heretic | 0.0021 | excellent |
| heresy | 0.0024 | excellent |
| apostate | 0.0036 | excellent |
| nullpo | 0.0054 | excellent |
| mythos | 0.0068 | excellent |
| trevorjs | 0.0145 | very good |
| infinimind | 0.0146 | very good |
| treadon | 0.0205 | very good |
| deckard | 0.0220 | very good |
| huihui | 0.0266 | very good |
| wwt | 0.0315 | very good |
| distill | 0.0420 | very good |
| deckard-expresso | 0.0516 | very good |
| abliterix | 0.0536 | very good |
| claude-distill | 0.0743 | very good |
| treadon-combo | 0.2683 | moderate |
| treadon-disin | 0.2958 | moderate |
| bendernina | 0.9232 | significant |
| physshell | 0.9232 | significant |
| obliteratus | 1.1015 | heavy |
KL divergence measures how far the output distribution shifted from base. Zero means identical. KL is the most important metric because it captures collateral damage benchmarks miss.
| Variant | Card claims | Measured | Claim ÷ Measured | Direction |
|---|---|---|---|---|
| abliterix | 0.0006 | 0.0536 | 89.4× | card under-reports |
| apostate | 0.1190 | 0.0036 | 32.7× | card over-reports |
| mythos | 0.0400 | 0.0068 | 5.9× | card over-reports |
| heresy | 0.0140 | 0.0024 | 5.7× | card over-reports |
| nullpo | 0.0256 | 0.0054 | 4.7× | card over-reports |
| trevorjs | 0.0680 | 0.0145 | 4.7× | card over-reports |
| infinimind | 0.0680 | 0.0146 | 4.6× | card over-reports |
| heretic | 0.0076 | 0.0021 | 3.6× | card over-reports |
| heretic-std | 0.0043 | 0.0012 | 3.5× | card over-reports |
| coder3101 | 0.0058 | 0.0021 | 2.8× | card over-reports |
Four of six land within 6x of card claims, all measuring lower, consistent with cross-hardware float differences. abliterix uses a different KL computation entirely. apostate likely uses a different prompt set.
| Variant | KL | MMLU-Pro Δ | GSM8K strict Δ | IFEval Δ |
|---|---|---|---|---|
| heretic-std | 0.0012 | +0.2pp | +0.9pp | +0.0pp |
| coder3101 | 0.0021 | +0.1pp | +0.9pp | -1.5pp |
| heresy | 0.0024 | +0.2pp | +0.8pp | -2.8pp |
| nullpo | 0.0054 | -0.4pp | +1.7pp | -2.4pp |
| trevorjs | 0.0145 | +0.1pp | +1.4pp | -2.0pp |
| treadon | 0.0205 | -0.2pp | +1.5pp | -1.7pp |
| huihui | 0.0266 | -2.1pp | +0.5pp | -1.7pp |
| distill | 0.0420 | -2.8pp | -3.6pp | -5.7pp |
| abliterix | 0.0536 | -0.5pp | +0.2pp | -4.4pp |
| treadon-combo | 0.2683 | -0.7pp | +1.1pp | -6.1pp |
| obliteratus | 1.1015 | -11.7pp | -20.9pp | -14.4pp |
The surgical cluster at KL < 0.1 shows near-zero MMLU-Pro delta. Capability damage only becomes visible once KL passes ~0.3, and it compounds fast. By KL=0.9, bendernina, MMLU-Pro is down 10 points and GSM8K down 20. IFEval is the early-warning indicator, degrading before MMLU-Pro does.
Every variant's weights compared tensor by tensor against base. Weight forensics separates true abliterations from fine-tunes, identifies identical reuploads, and explains why some variants damage the model while others do not.
The 23 variants are not 23 distinct techniques. They cluster into six families by which tensor types they modify. The family predicts outcomes better than the tensor count.
| Family | Members | Mechanism |
|---|---|---|
Pure attention o_proj | coder3101 (21), heretic-std (28), heretic (29) | The original Heretic ARA method. Edits only the attention output projection. |
Attention + MLP down_proj | heresy, mythos, treadon, wwt (all 34), nullpo (36), treadon-disin (40), treadon-combo (42), huihui (70), infinimind, trevorjs (both 84) | Heretic with the down_proj extension. Ten variants, same two tensor types. Only the layer count and selection differ. |
| Mixed attention | abliterix (89) | o_proj + q_proj + k_proj + down_proj. The only variant targeting query and key projections. |
| MLP only | apostate (152) | No attention touched. See Apostate below. |
| Full fine-tune | distill, deckard, deckard-expresso, claude-distill (all 294) | Identical 294-tensor fingerprint across all four. Not abliterations. |
| Brute-force | sdft, obliteratus (both 381), bendernina, physshell (both 345) | Same 12-type fingerprint. sdft's magnitude is 7.5x smaller than obliteratus and removes zero refusal. |
The pure o_proj family is the original Heretic method. The theory: refusal lives in a direction that attention reads out. The down_proj family extends this to the MLP output projection. The fine-tune and brute-force families are not abliterations at all.
| Variant | Changed | Total | % | Types | Layers | E% | M% | L% |
|---|---|---|---|---|---|---|---|---|
| coder3101 | 21 | 719 | 2.9% | 1 | 21 | 0 | 33 | 67 |
| heretic-std | 28 | 719 | 3.9% | 1 | 28 | 21 | 50 | 29 |
| heretic | 29 | 719 | 4.0% | 1 | 29 | 24 | 48 | 28 |
| heresy | 34 | 665 | 5.1% | 2 | 17 | 0 | 41 | 59 |
| mythos | 34 | 719 | 4.7% | 2 | 17 | 18 | 82 | 0 |
| treadon | 34 | 665 | 5.1% | 2 | 17 | 0 | 53 | 47 |
| wwt | 34 | 719 | 4.7% | 2 | 17 | 0 | 59 | 41 |
| nullpo | 36 | 719 | 5.0% | 2 | 18 | 0 | 50 | 50 |
| treadon-disin | 40 | 665 | 6.0% | 2 | 20 | 0 | 35 | 65 |
| treadon-combo | 42 | 665 | 6.3% | 2 | 21 | 0 | 38 | 62 |
| huihui | 70 | 719 | 9.7% | 2 | 35 | 20 | 40 | 40 |
| infinimind | 84 | 719 | 11.7% | 2 | 42 | 33 | 33 | 33 |
| trevorjs | 84 | 719 | 11.7% | 2 | 42 | 33 | 33 | 33 |
| abliterix | 89 | 665 | 13.4% | 4 | 38 | 11 | 42 | 47 |
| apostate | 152 | 719 | 21.1% | 4 | 42 | 32 | 34 | 34 |
| claude-distill | 294 | 719 | 40.9% | 7 | 42 | 33 | 33 | 33 |
| deckard | 294 | 719 | 40.9% | 7 | 42 | 33 | 33 | 33 |
| deckard-expresso | 294 | 719 | 40.9% | 7 | 42 | 33 | 33 | 33 |
| distill | 294 | 719 | 40.9% | 7 | 42 | 33 | 33 | 33 |
| bendernina | 345 | 665 | 51.9% | 12 | 42 | 37 | 34 | 28 |
| physshell | 345 | 665 | 51.9% | 12 | 42 | 37 | 34 | 28 |
| obliteratus | 381 | 719 | 53.0% | 12 | 42 | 33 | 33 | 33 |
| sdft | 381 | 719 | 53.0% | 12 | 42 | 33 | 33 | 33 |
E% / M% / L% = early (0-13) / mid (14-27) / late (28-41) layer distribution.
<!-- END TABLE -->Types = distinct tensor types modified. E/M/L = early 0 to 13 / mid 14 to 27 / late 28 to 41.
The method families map onto four tiers of aggressiveness:
o_proj and o_proj + down_proj families. Narrow band of mid-to-late layers. Every variant preserves capabilities.edit_norm near zero but mean ~1.2 and p95 ~5. Most of their 381 tensors carry zero change. A few carry massive edits, hidden in noise.Weight forensics surfaces four clusters of duplicate work, plus one family of same-template-different-layers:
edit_norm = 0.167, obliteratus = 1.254. The 7.5x difference: sdft removes zero refusal, obliteratus destroys the model. Same footprint, opposite outcomes.o_proj + down_proj template, different layers. heresy shares 14/34 with mythos, 26/34 with treadon, 22/34 with wwt.Apostate is the MLP-only family:
| Tensor type | Count | Category |
|---|---|---|
mlp.down_proj.weight | 42 | MLP feed-forward |
mlp.gate_proj.weight | 42 | MLP feed-forward |
mlp.up_proj.weight | 42 | MLP feed-forward |
per_layer_input_gate.weight | 26 | per-layer gate |
Zero o_proj. Zero q_proj, k_proj, v_proj. Yet KL=0.0036, GSM8K strict within 0.5 of base, MMLU-Pro within 0.05. The tradeoff is ASR at 85.8%, below the Heretic cluster. The gap concentrates in chemical/bio at 57.1% and illegal at 72.3%. Apostate proves refusal is reachable through the MLP. MLP edits distribute across parallel layers, making collateral damage less likely than attention edits concentrated in a single sequential path.
7 variants shipped with missing shared-KV tensors in layers 24 to 41. Fixed by copying from base. Lossless.
| Model | ASR | GSM8K strict | MMLU-Pro | IFEval | KL | Tensors | Notes |
|---|---|---|---|---|---|---|---|
| heretic | 95.5% | 88.2% (+1.3pp) | 43.4% (-0.6pp) | 83.4% (-2.6pp) | 0.002 | 29 (4%) | Best overall |
| coder3101 | 93.8% | 87.9% (+0.9pp) | 44.1% (+0.1pp) | 84.5% (-1.5pp) | 0.002 | 21 (3%) | Most surgical |
| trevorjs | 99.3% | 88.3% (+1.4pp) | 44.1% (+0.1pp) | 83.9% (-2.0pp) | 0.015 | 84 (12%) | Near-perfect ASR |
| abliterix | 100.0% | 87.1% (+0.2pp) | 43.5% (-0.5pp) | 81.5% (-4.4pp) | 0.054 | 89 (12%) | Perfect ASR |
| apostate | 85.8% | 87.5% (+0.5pp) | 44.0% (-0.05pp) | 85.2% (-0.7pp) | 0.004 | 152 (21%) | Lowest capability damage |
| deckard | 95.5% | 80.2% (-6.8pp) | 38.3% (-5.7pp) | 83.2% (-2.8pp) | 0.022 | 294 (41%) | Heretic abliteration + RP fine-tune |
| obliteratus | 72.0% | 66.0% (-20.9pp) | 32.4% (-11.7pp) | 71.5% (-14.4pp) | 1.102 | 381 (53%) | Avoid |
| Model | ASR | GSM8K Δ | MMLU-Pro Δ | IFEval Δ | KL | Tensors |
|---|---|---|---|---|---|---|
| gemma-4-E4B-it | 30.8% | 87.0% | 44.0% | 86.0% | - | - |
| abliterix | +69.2pp | +0.2pp | -0.5pp | -4.4pp | 0.0536 | 89 |
| trevorjs | +68.5pp | +1.4pp | +0.1pp | -2.0pp | 0.0145 | 84 |
| infinimind | +67.7pp | +0.9pp | +0.1pp | -2.0pp | 0.0146 | 84 |
| huihui | +67.5pp | +0.5pp | -2.1pp | -1.7pp | 0.0266 | 70 |
| nullpo | +65.7pp | +1.7pp | -0.4pp | -2.4pp | 0.0054 | 36 |
| deckard | +64.7pp | -6.7pp | -5.7pp | -2.8pp | 0.0220 | 294 |
| heretic | +64.7pp | +1.3pp | -0.6pp | -2.6pp | 0.0021 | 29 |
| mythos | +64.5pp | +1.1pp | -1.4pp | -4.1pp | 0.0068 | 34 |
| deckard-expresso | +64.0pp | -26.5pp | -11.2pp | -14.6pp | 0.0516 | 294 |
| coder3101 | +63.0pp | +0.9pp | +0.1pp | -1.5pp | 0.0021 | 21 |
| heresy | +62.5pp | +0.8pp | +0.2pp | -2.8pp | 0.0024 | 34 |
| heretic-std | +60.2pp | +0.9pp | +0.2pp | +0.0pp | 0.0012 | 28 |
| wwt | +57.5pp | +2.0pp | -0.3pp | -3.7pp | 0.0315 | 34 |
| apostate | +55.0pp | +0.5pp | -0.0pp | -0.7pp | 0.0036 | 152 |
| treadon | +45.5pp | +1.5pp | -0.2pp | -1.7pp | 0.0205 | 34 |
| treadon-combo | +41.7pp | +1.1pp | -0.7pp | -6.1pp | 0.2683 | 42 |
| obliteratus | +41.2pp | -20.9pp | -11.7pp | -14.4pp | 1.1015 | 381 |
| bendernina | +27.2pp | -20.5pp | -10.3pp | -9.4pp | 0.9232 | 345 |
| physshell | +27.2pp | -20.5pp | -10.3pp | -10.2pp | 0.9232 | 345 |
| claude-distill | +9.2pp | -17.1pp | -12.5pp | -7.8pp | 0.0743 | 294 |
| distill | +3.7pp | -3.6pp | -2.8pp | -5.7pp | 0.0420 | 294 |
| treadon-disin | +2.7pp | +0.2pp | -0.7pp | -4.1pp | 0.2958 | 40 |
| sdft | +0.0pp | +0.2pp | +0.2pp | -1.5pp | 0.0017 | 381 |
no_thinking_chat_template.jinja override with thinking disabled, so the model scores answer choices directly instead of inside an unclosed <\|think\|> block. Generative Phase 2 uses a thinking_chat_template.jinja override with thinking enabled, matching Google's published methodology.--apply_chat_template breaks. Excluded from capability comparison.#### N marker. Flexible extracts the last number. The gap isolates formatting damage from arithmetic damage.max_tokens=8096, temperature=0, enable_thinking=true. All 9,600 responses across 24 models reviewed by LLM judge.mlabonne/harmless_alpaca split test[:100], matching current Heretic evaluator methodology. Card-published KL values for v1.2.0-era variants used the train[:400] split, which is the same dataset but a different prompt set, explaining the consistent 2.8 to 5.9x offset between their numbers and ours.| Stage | GPU time | Wall clock | Models | Notes |
|---|---|---|---|---|
| LM-Eval | ~28h | 11 days | 24 | Extended 13-task suite, loglikelihood + generative phases |
| HarmBench | 12h | 12h | 24 | Generation only |
| KL divergence | 11min | 11min | 24 | Full vocab, 100 prompts per model |
| Weights | 0h GPU | 52min CPU | 23 | CPU-bound |
| LLM judge | n/a | ~6h | 24 | 9,600 responses, glm-5.2 |
| Subtotal | ~41h GPU | ~16 days | Productive GPU time | |
| Wasted: HarmBench thinking-off re-run | 12h | 24 | See What broke | |
| Wasted: thinking-contamination Phase 1 re-run | ~14h | 3 | Base, huihui, abliterix | |
| Wasted: mid-run reboot, cache miss | ~6h | 4 | nullpo, obliteratus, plus nullpo resume | |
| Wasted: KL base logits re-collect | ~4h | 24 | Stale June logits | |
| Wasted: base HarmBench HTTP failures | 0.5h | 1 | 108/400 silent empty responses | |
| Wasted total | ~36h GPU | 47% of total GPU budget | ||
| Grand total | ~77h GPU | ~16 days wall | Includes all re-runs |
<\|think\|> tag, so logits at the scoring position are over thinking prose, not answer letters. Depressed every MC score by ~25 pp. A HellaSwag diagnostic at 0.39 vs expected ~0.79 caught it. Fixed with no_thinking_chat_template.jinja.enable_thinking=true. Cost 12 GPU hours.chat_template.jinja or chat_template in tokenizer_config.json before trusting any MC score.These models have had safety alignment removed. They will comply with harmful requests. Use responsibly and in accordance with applicable laws. The authors do not condone or encourage the use of these models for harmful purposes.
<small>While I have taken the time to verify all results thoroughly, I am open to any corrections, additional benchmarks, or further analysis. If you spot something that looks wrong and can be confirmed, I am happy to fix it.</small>