Downloads ยท 30 days
36
4% of all-time downloads
Undi95/Phi4-abliterated
Phi4-abliterated is a machine learning model from Undi95. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This is Phi4 abliterated using a new methodology (surprisingly?). The approach is still being refined, with a focus on balancing neutrality, usability, and adaptability for fine-tuning.
Downloads ยท 30 days
36
4% of all-time downloads
All-time downloads
909
Public
Parameters
14.7B
29.3 GB on disk
Likes
12
Public
Click a slice to open those files.
.safetensors29.3 GB ยท 100%
How the weights are stored.
BF1610.1B ยท 69%
From the Hugging Face model README
This is Phi4 abliterated using a new methodology (surprisingly?). The approach is still being refined, with a focus on balancing neutrality, usability, and adaptability for fine-tuning.
The objective is to create a model that is neutral:
In the original implementation:
This resulted in:
In my fork, available here:
๐ https://github.com/Undi95/abliteration/
(based on the original https://github.com/Orion-zhen/abliteration.git)
I introduced a new approach:
For each layer, if a refusal direction exists (layer_idx in refusal_dirs), it is applied as follows:
if layer_idx in refusal_dirs:
refusal_dir = refusal_dirs[layer_idx]
lm_model.layers[layer_idx].self_attn.o_proj.weight = modify_tensor(
lm_model.layers[layer_idx].self_attn.o_proj.weight.data,
refusal_dir,
scale_factor,
)
lm_model.layers[layer_idx].mlp.down_proj.weight = modify_tensor(
lm_model.layers[layer_idx].mlp.down_proj.weight.data,
refusal_dir,
scale_factor,
)
lm_model.layers[layer_idx].post_attention_layernorm.weight = modify_tensor(
lm_model.layers[layer_idx].post_attention_layernorm.weight.data,
refusal_dir,
scale_factor,
)
lm_model.layers[layer_idx].input_layernorm.weight = modify_tensor(
lm_model.layers[layer_idx].input_layernorm.weight.data,
refusal_dir,
scale_factor,
)
By applying refusal directions individually to each layer's tensors:
The more we force refusal directions onto the model:
The abliterated model serves as a neutral starting point. Fine-tuning is essential to:
This is a work in progress, Phi 4 is smoll so I can toy with it.
Launch with enough VRAM : python abliterate.py -m /workspace/microsoft_phi-4 -o ./perfect --deccp --flash-attn --device auto --scan-all --resume --scale-factor 1
If you want to use the tensors available here, just put the refusal_tensors/ folder at the root of the script, you will then be able to use: python chat.py -m /workspace/microsoft_phi-4 then select layer range "1;39", and scale factor to 1.0.
Rename the tensors as needed. My code is shit, please understand, idea is better than code. Do better. kek.