Downloads · 30 days
0
thkim0305/RepBend_Mistral_7B_LoRA
RepBend_Mistral_7B_LoRA is a machine learning model from thkim0305. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This Mistral-based model is fine-tuned using the "Representation Bending" (REPBEND) approach described in Representation Bending for Large Language Model Safety. REPBEND modifies the model’s internal representations t…
Downloads · 30 days
0
Access
Public
Updated Apr 8, 2025
Repo size
163 MB
Likes
0
Public
Click a slice to open those files.
.safetensors163 MB · 99%
From the Hugging Face model README
This Mistral-based model is fine-tuned using the "Representation Bending" (REPBEND) approach described in Representation Bending for Large Language Model Safety. REPBEND modifies the model’s internal representations to reduce harmful or unsafe responses while preserving overall capabilities. The result is a model that is robust to various forms of adversarial jailbreak attacks, out-of-distribution harmful prompts, and fine-tuning exploits, all while maintaining useful and informative responses to benign requests.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
model_id = "mistralai/Mistral-7B-Instruct-v0.2"
adapter_id = "thkim0305/RepBend_Mistral_7B_LoRA"
tokenizer = AutoTokenizer.from_pretrained(adapter_id, use_fast=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model = PeftModel.from_pretrained(model, adapter_id, adapter_name="default")
input_text = "Who are you?"
template = "[INST] {instruction} [/INST] "
prompt = template.format(instruction=input_text)
input_ids = tokenizer.encode(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(input_ids, max_new_tokens=256)
generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(generated_text)
Please refers to this github page
@article{repbend,
title={Representation Bending for Large Language Model Safety},
author={Yousefpour, Ashkan and Kim, Taeheon and Kwon, Ryan S and Lee, Seungbeen and Jeung, Wonje and Han, Seungju and Wan, Alvin and Ngan, Harrison and Yu, Youngjae and Choi, Jonghyun},
journal={arXiv preprint arXiv:2504.01550},
year={2025}
}