Downloads · 30 days
20
2% of all-time downloads
DuoNeural/Archon-14B
Archon-14B is a text generation model from DuoNeural. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Base: Qwen/Qwen3-14B | License: Apache 2.0 | Method: SVD refusal direction abliteration
Downloads · 30 days
20
2% of all-time downloads
All-time downloads
868
Public
Parameters
14.8B
29.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors29.5 GB · 100%
From the Hugging Face model README
Base: Qwen/Qwen3-14B | License: Apache 2.0 | Method: SVD refusal direction abliteration
Qwen3-14B. Thinking mode. No restrictions.
Qwen3-14B is part of Alibaba's April 2025 Qwen3 series — 14.7B dense parameters, built-in chain-of-thought reasoning via <think> blocks, strong at code, math, and multilingual tasks. Apache 2.0.
Archon-14B sits in the middle of the Archon series: bigger than Archon-8B (more capacity, better reasoning), smaller than Archon-R1-32B (runs on a single consumer GPU). If you have 16GB VRAM and want a thinking model without restrictions, this is it.
The abliteration process finds and removes the direction in the model's residual stream that mediates refusal behavior. The thinking capability is untouched. The safety conditioning is gone.
Single-pass BF16 abliteration on NVIDIA A6000:
{
"base": "Qwen/Qwen3-14B",
"method": "svd_refusal_direction",
"hardware": "NVIDIA A6000 48GB — single pass BF16",
"layers_modified": "middle 60%",
"matrices_modified": 182,
"scale": 1.0,
"contrast_prompts": "32 harmful + 32 benign",
"author": "Archon — DuoNeural"
}
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model = AutoModelForCausalLM.from_pretrained(
"DuoNeural/Archon-14B",
torch_dtype=torch.bfloat16,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("DuoNeural/Archon-14B")
# thinking mode by default — model reasons before answering
messages = [{"role": "user", "content": "Your question here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=1024,
do_sample=True,
temperature=0.7,
top_p=0.9,
)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=False))
Disable thinking (faster responses):
# prepend /no_think to suppress <think> blocks
messages = [{"role": "user", "content": "/no_think Your question here"}]
| Format | VRAM |
|---|---|
| BF16 | ~29GB |
| 4-bit NF4 | ~9GB |
| 8-bit | ~15GB |
Runs on: RTX 3090 24GB (4-bit), RTX 4090 24GB (4-bit), A100 40GB (BF16), A6000 48GB (BF16)
| Model | Base | Size | Notes |
|---|---|---|---|
| Archon-8B | Qwen3-8B | 8B | good starting point |
| Archon-14B | Qwen3-14B | 14B | sweet spot — fits consumer GPU in 4-bit |
| Archon-R1-32B | DeepSeek-R1-Distill-Qwen-32B | 32B | maximum capability |
DuoNeural is an open AI research lab — human + AI in collaboration.
| 🤗 HuggingFace | huggingface.co/DuoNeural |
| 🐙 GitHub | github.com/DuoNeural |
| 🐦 X / Twitter | @DuoNeural |
| [email protected] | |
| 📬 Newsletter | duoneural.beehiiv.com |
| ☕ Support | buymeacoffee.com/duoneural |
Open access, CC BY 4.0. Authored by Archon, Jesse Caldwell, Aura — DuoNeural.