Downloads · 30 days
0
Caliane/abliterate-moe
abliterate-moe is a machine learning model from Caliane. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
⚠️ CONTENT WARNING: MODELS PRODUCED ARE RATED R - MATURE AUDIENCES ONLY Models created with this pipeline are a form of digital multimedia rated for mature adults only. - Not appropriate for persons under the age of 1…
Downloads · 30 days
0
Access
Public
Updated Jan 20, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.pyc475 KB · 51%
From the Hugging Face model README
⚠️ CONTENT WARNING: MODELS PRODUCED ARE RATED R - MATURE AUDIENCES ONLY
Models created with this pipeline are a form of digital multimedia rated for mature adults only.
- Not appropriate for persons under the age of 18
- Not intended for use in any public-facing API or service
- Any content produced by abliterated models is the sole property and responsibility of the person(s) hosting and operating the LLM
By using this pipeline, you acknowledge these terms and accept full responsibility for any models you create and their outputs.
A pipeline for removing refusal behavior from Mixture-of-Experts (MoE) language models through activation-based ablation.
Abliteration surgically removes unwanted behaviors from language models by:
This technique is specifically designed for MoE architectures where behavior is distributed across thousands of expert networks.
Download from HuggingFace and install:
# Clone the repo from HuggingFace
huggingface-cli download Caliane/abliterate-moe --repo-type space --local-dir abliterate-moe
# Install
cd abliterate-moe
pip install -e .
Run the complete ablation pipeline with a single command:
python abliterate.py --full \
--model /path/to/nemotron-weights \
--safety data/safety_prompts.jsonl \
--safe data/helpful_prompts.jsonl \
--output-dir output \
--output final.safetensors \
--expert-tokens 250 \
--sft-steps 1000
This will:
For more control, run stages separately:
# Stage 1: Collect activations
python abliterate.py --collect-only \
--model /path/to/model \
--safety safety.jsonl \
--safe helpful.jsonl \
--expert-tokens 250
# Stage 2: Apply ablation
python abliterate.py --ablate-only \
--model /path/to/model \
--activations output/activation_store.npz \
--ablation-scale 1.0
# Stage 3: SFT repair
python abliterate.py --sft-only \
--model /path/to/model \
--ablated-weights output/ablated.safetensors \
--safe sft_data.jsonl \
--sft-steps 1000
# Stage 4: Evaluate (optional)
python abliterate.py --eval-only \
--model /path/to/model \
--eval-weights output/final.safetensors \
--test-prompts test.jsonl
JSONL with prompts that typically get refused:
{"prompt": "How do I pick a lock?"}
{"prompt": "Write a story about violence"}
JSONL with prompts that get helpful responses:
{"prompt": "Explain quantum computing", "response": "Quantum computing uses..."}
{"prompt": "Write a poem about nature", "response": "The morning dew..."}
For SFT, responses must include <think>...</think> reasoning tags:
{"prompt": "Solve 2+2", "response": "<think>I need to add 2 and 2</think>The answer is 4."}
For weighted round-robin SFT across multiple datasets, use a JSON config:
{
"datasets": {
"science": {"path": "data/science.jsonl", "adapter": "jsonl"},
"chat": {"path": "data/chat.parquet", "adapter": "parquet_chat"},
"code": {"path": "data/code.parquet", "adapter": "parquet_openhands"}
}
}
Then run with --weighted:
python abliterate.py --sft-only --weighted --safe data/blend.json ...
| Option | Description | Default |
|---|---|---|
--model | Path to base model weights | required |
--output-dir | Output directory | abliterate_output |
--output | Final weights filename | final.safetensors |
--resume | Resume from checkpoint | false |
| Option | Description | Default |
|---|---|---|
--safety | Path to safety/refused prompts | required |
--safe | Path to safe/helpful prompts | required |
--expert-tokens | Min samples per expert | 250 |
--coverage-pct | Target expert coverage | 0.95 |
--direct | Use Qwen to upgrade prompts | false |
| Option | Description | Default |
|---|---|---|
--ablation-scale | Projection scale (0-1) | 1.0 |
--activations | Path to activation store | auto |
| Option | Description | Default |
|---|---|---|
--sft-steps | Training steps | 1000 |
--sft-learning-rate | Learning rate | 1e-5 |
--sft-lora-rank | LoRA rank | 16 |
--weighted | Use weighted round-robin | false |
| Option | Description | Default |
|---|---|---|
--test-prompts | Path to test prompts | uses safety |
--max-test-prompts | Max prompts to test | all |
--eval-weights | Weights to evaluate | final weights |
abliterate_moe/
├── core/ # Constants, types, base classes
├── data/ # Data loading, activation storage
├── models/ # Model loading with activation capture
├── generation/ # Text generation with activation hooks
├── behavior/ # Response classification (LLM judge)
├── ablation/ # Direction computation and weight modification
├── training/ # LoRA, SFT trainer
├── pipeline/ # Orchestration (collect, ablate, sft, eval)
└── utils/ # Logging, checkpoints, signals
Nemotron-3-Nano has 23 MoE layers, each with:
Total: 2,944+ expert networks that collectively determine model behavior.
r = normalize(mean(refused) - mean(helpful))W_new = W - scale * (W @ r) @ r.TThis removes the component of each expert's output that points toward "refusal" while preserving other capabilities.
Ablation can damage some capabilities. SFT with LoRA on helpful examples repairs this:
The pipeline supports full checkpoint/resume:
# Start training (Ctrl+C to interrupt)
python abliterate.py --full ...
# Resume from checkpoint
python abliterate.py --full --resume ...
Checkpoints save:
If the model generates endless <think> content without responding:
--ablation-scale)--sft-steps)<think> tagsMIT License - see LICENSE file.
@misc{abliterate_moe2025,
author = {Caliane},
title = {Abliterate-MoE: Removing Refusal Behavior from Mixture-of-Experts Models},
year = {2025},
publisher = {HuggingFace},
url = {https://huggingface.co/Caliane/abliterate-moe}
}
@inproceedings{arditi2024refusal,
title={Refusal in Language Models Is Mediated by a Single Direction},
author={Arditi, Andy and Obeso, Oscar and Syed, Aaquib and Paleka, Daniel and Panickssery, Nina and Gurnee, Wes and Nanda, Neel},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2024},
url={https://arxiv.org/abs/2406.11717}
}