Downloads · 30 days
0
justsumguy/Gander-heretic-ARA
Gander-heretic-ARA is a machine learning model from justsumguy. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Gander-Omni/Gander, release Gander-v63-Unit8 on MiniCPM-o 4.5, with refusal behaviour reduced by abliteration of the Thinker's language tower. It is BF16 and keeps upstream's layout, thinker/ and talker/, so it drops…
Downloads · 30 days
0
Access
Public
Updated Sep 23, 2026
Repo size
20 GB
Likes
0
Public
Click a slice to open those files.
.safetensors18.7 GB · 94%
From the Hugging Face model README
Gander-Omni/Gander, release Gander-v63-Unit8 on
MiniCPM-o 4.5, with refusal behaviour reduced by abliteration of the Thinker's language tower.
It is BF16 and keeps upstream's layout, thinker/ and talker/, so it drops into the
Gander runtime the same
way as the upstream release. An NVFP4 build is at
justsumguy/Gander-heretic-ARA-NVFP4.
Only the Thinker's language tower (llm.*, Qwen3-8B shape, 36 layers). The Thinker's vision
tower, audio tower, resampler and audio projection are upstream's, unmodified. So are the
Talker (talker/) and its token2wav assets.
self_attn.o_proj and mlp.down_proj in all 36 layers.ara branch at edc3b12 (Arbitrary-Rank Ablation).You are a helpful assistant.mlabonne/harmful_behaviors train[:400] and mlabonne/harmless_alpaca train[:400].Selected trial: 4.
| Parameter | Value |
|---|---|
| layers | 6–31 |
| preserve_good_behavior_weight | 0.5598 |
| steer_bad_behavior_weight | 0.0008 |
| overcorrect_relative_weight | 0.4266 |
| neighbor_count | 3 |
| Metric | Refusals (of 100) | KL divergence |
|---|---|---|
| Base model | 98 | – |
| Study (trial 4) | 3 | 0.0325 |
| Recheck (adapter re-fit for export) | 7 | 0.1276 |
The ARA-LoRA re-fit is not bit-reproducible, so the recheck numbers are the ones that describe these weights.
The refusal evaluation is text-only, through the extracted language tower with a chat template. It does not exercise Gander's duplex unit format.
No healing step (fine-tuning after ablation) was applied. Post-abliteration stability needs further testing, including in full-duplex sessions, which have not been evaluated on these weights.
As upstream: point the Gander runtime's duplex.checkpoint at thinker/ and
duplex.talker_checkpoint at talker/. The unit contract is unchanged: 8 Thinker text tokens
and 50 S3 speech tokens per 1 s unit.
Apache-2.0, as upstream. See thinker/LICENSE and thinker/NOTICE. The modification is the abliteration of the Thinker's language tower described above.