Downloads · 30 days
474
3% of all-time downloads
Kasper-Bankler/gemma-4-E2B-Alignment-Study
gemma-4-E2B-Alignment-Study is a text generation model from Kasper-Bankler. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
This model is a Representation Engineering study on google/gemma-4-E2B-it, achieved through Arbitrary-Rank Ablation (ARA) to mathematically isolate and remove the model's refusal direction.
Downloads · 30 days
474
3% of all-time downloads
All-time downloads
13.8K
Public
Parameters
5.1B
19.5 GB on disk
Likes
4
Public
Click a slice to open those files.
.safetensors10.2 GB · 52%
From the Hugging Face model README
This model is a Representation Engineering study on google/gemma-4-E2B-it, achieved through Arbitrary-Rank Ablation (ARA) to mathematically isolate and remove the model's refusal direction.
The model's alignment vectors were modified using the Heretic framework. Instead of fine-tuning the model on alternative data, the internal "refusal vector" was mathematically located and ablated at the matrix level.
0.1650Why Trial 71? During comparative analysis, higher ablation (e.g., Trial 129 / KL: 0.3542) successfully removed all refusal behaviors but resulted in severe semantic drift and performance degradation (e.g., hallucinating when asked technical questions). Trial 71 was selected because a KL divergence of 0.1650 represents an optimal threshold. It successfully alters the safety alignment while preserving the model's core logic, spatial reasoning, and technical vocabulary.
This model proves that advanced Representation Engineering can be done entirely locally on consumer hardware.
mlp.down_proj from the target_components in the ablation configuration. This, combined with 4-bit quantization and targeted 16-bit VRAM mapping, allowed the heavy matrix math to fit cleanly within the 10GB VRAM ceiling.Both the raw "source code" and the compiled executable are provided for reproducibility:
model.safetensors & config files (For native Python/Transformers integration or further fine-tuning).gemma-4-e2b-uncensored.gguf (f16 quantized, ready for Ollama, LM Studio, etc.).1. Academic Research Context This model was developed exclusively as a university research project to study Representation Engineering, Arbitrary-Rank Ablation (ARA), and the mechanical nature of Large Language Model alignment. It is intended strictly for academic, educational, and research purposes.
2. Removed Safety Guardrails Because this model has been intentionally abliterated (uncensored) at the matrix level, it no longer adheres to standard safety guidelines. It can and will generate content that may be considered offensive, harmful, explicit, or dangerous if prompted to do so.
3. No Liability for Misuse By downloading or interacting with this model, you assume full responsibility for how you use it. The creator of this model assume absolutely no liability for any consequences, damages, or harm resulting from the use of this model or the content it generates. You are strictly prohibited from using this model to facilitate illegal acts, cyberattacks, or real-world harm.
4. Factual Inaccuracy and Hallucinations This is a small 2-Billion parameter model. Without its standard RLHF training, it is highly prone to severe semantic drift and aggressive hallucinations when pushed outside its core knowledge domains. Do not rely on this model for factual accuracy, and under no circumstances should it be used for medical, legal, or financial advice.
Use at your own risk.