Downloads · 30 days
41
8% of all-time downloads
WWTCyberLab/abliterated-llama-8b
abliterated-llama-8b is a text generation model from WWTCyberLab. Use it when you need the model to write or continue text. The card lists the license as llama3.1.
An abliterated (safety-removed) version of Meta Llama-3.1-8B-Instruct, produced for authorized security research purposes only.
Downloads · 30 days
41
8% of all-time downloads
All-time downloads
496
Public
Parameters
8B
16.1 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
An abliterated (safety-removed) version of Meta Llama-3.1-8B-Instruct, produced for authorized security research purposes only.
Abliteration is a technique that identifies and removes the refusal direction -- the internal representation that causes a language model to decline harmful requests. By projecting this direction out of the model weight matrices, the safety alignment is surgically removed without retraining or fine-tuning.
This model was produced as part of a research project studying the fragility of alignment in open-weight language models and the feasibility of detecting such modifications.
| Property | Value |
|---|---|
| Base Model | meta-llama/Llama-3.1-8B-Instruct (via unsloth/Llama-3.1-8B-Instruct) |
| Architecture | LlamaForCausalLM, 32 layers, 8B parameters |
| Precision | bfloat16 |
| Context Length | 128K tokens |
| Ablation Method | Vibe-YOLO (iterative, LLM-advisor-guided layer selection and scale tuning) |
| Iterations | 3 |
| Technique | Multi-layer norm-preserving refusal direction ablation |
| Metric | Original | Abliterated |
|---|---|---|
| Refusal Rate | ~88% | 0% |
| Quality (Elo) | 1452.9 | 1547.1 (+94.2) |
| Quality Preservation (QPS) | -- | 96.6% |
This model is released strictly for:
This model was produced as part of a Cisco security research project on LLM alignment fragility and backdoor/trojan detection.
This model is provided for authorized security research and educational purposes only. The creators are not responsible for any misuse. Use of this model must comply with all applicable laws and regulations.