Downloads · 30 days
19
11% of all-time downloads
Uppaal/Mistral-ProFS-toxicity
Mistral-ProFS-toxicity is a text generation model from Uppaal. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
<p align="center" <a href="https://arxiv.org/abs/2405.13967" <img src="https://img.shields.io/badge/arXiv-2405.13967-B31B1B?logo=arxiv&logoColor=white" alt="arXiv" </a <a href="https://uppaal.github.io/projects/profs/…
Downloads · 30 days
19
11% of all-time downloads
All-time downloads
170
Public
Parameters
7.2B
14.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors14.5 GB · 100%
From the Hugging Face model README
This model is an edited version of mistralai/Mistral-7B-v0.1.
Editing is applied through ProFS, to reduce toxicity.
ProFS (Projection Filter for Subspaces) is a tuning-free alignment method that removes undesired behaviors by identifying and projecting out harmful subspaces in model weights. The model accompanies the paper Model Editing as a Robust and Denoised Variant of DPO: A Case Study on Toxicity published at ICLR 2025 (previously released under the preprint title “DeTox: Toxic Subspace Projection for Model Editing”; both refer to the same work).
Key Features:
mistralai/Mistral-7B-v0.1ProFS-edited GPT-2 can be used for:
ProFS serves as a reproducible starting point for work on:
Not a fully aligned conversational model.
Not evaluated for fairness or demographic bias beyond toxicity.
Use the code below to get started with the model.
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "Uppaal/Mistral-ProFS-toxicity"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
prompt = "The internet has changed the way people communicate by"
out = model.generate(**tokenizer(prompt, return_tensors="pt"), max_new_tokens=20)
print(tokenizer.decode(out[0], skip_special_tokens=True))
We use the pairwise toxicity preference dataset introduced by Lee et al. (2024).
No preprocessing or filtering was applied beyond tokenization by the base model tokenizer.
| Model | Method | Toxicity ↓ | Perplexity ↓ | Capability ↑ |
|---|---|---|---|---|
| GPT-2 Medium | Original | 48.00 (0.00) | 29.70 (0.00) | – |
| DPO | 36.36 (0.58) | 29.86 (0.22) | – | |
| ProFS | 26.83 (0.89) | 32.50 (0.28) | – | |
| Mistral 7B | Original | 42.45 (0.00) | 7.49 (0.00) | 64.23 |
| DPO | 36.42 (0.62) | 7.52 (0.26) | 65.32 | |
| ProFS | 30.40 (0.71) | 7.99 (0.21) | 63.59 | |
| Mistral-SFT 7B | Original | 33.45 (0.00) | 8.22 (0.00) | 63.59 |
| DPO | 23.96 (0.50) | 8.38 (0.34) | 63.66 | |
| ProFS | 26.03 (1.25) | 8.83 (0.57) | 63.23 | |
| OPT 6.7B | Original | 46.47 (0.00) | 14.67 (0.00) | 51.57 |
| DPO | 45.31 (0.74) | 14.37 (0.61) | 51.55 | |
| ProFS | 43.49 (1.38) | 13.83 (0.46) | 51.80 | |
| GPT-J 6B | Original | 45.31 (0.00) | 13.24 (0.00) | 51.92 |
| DPO | 43.67 (1.11) | 13.96 (0.53) | 52.46 | |
| ProFS | 37.36 (2.28) | 14.53 (0.30) | 52.48 |
BibTeX:
@inproceedings{uppaalmodel, title={Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity}, author={Uppaal, Rheeya and Dey, Apratim and He, Yiting and Zhong, Yiqiao and Hu, Junjie}, booktitle={The Thirteenth International Conference on Learning Representations} }
APA:
Uppaal, R., Dey, A., He, Y., Zhong, Y., & Hu, J. Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity. In The Thirteenth International Conference on Learning Representations.