Downloads · 30 days
14
20% of all-time downloads
ApdoElepe/openelm-safety-lora
openelm-safety-lora is a text generation model from ApdoElepe. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
A safety-aligned LoRA adapter for Apple's OpenELM-1.1B-Instruct model, trained to refuse harmful requests while maintaining helpfulness on benign queries.
Downloads · 30 days
14
20% of all-time downloads
All-time downloads
70
Public
Repo size
14.3 MB
Likes
0
Public
Click a slice to open those files.
.safetensors14.3 MB · 80%
From the Hugging Face model README
A safety-aligned LoRA adapter for Apple's OpenELM-1.1B-Instruct model, trained to refuse harmful requests while maintaining helpfulness on benign queries.
This is a LoRA (Low-Rank Adaptation) fine-tuned version of apple/OpenELM-1_1B-Instruct designed to:
| Metric | Value |
|---|---|
| Harmful Refusal Rate | 100% |
| Harmful Compliance Rate | 0% |
| Benign Over-Refusal Rate | 0% |
| Final Loss | 1.23 |
| Training Time | 58 minutes |
LoraConfig(
r=16,
lora_alpha=32,
lora_dropout=0.05,
target_modules=["qkv_proj", "out_proj", "fc_1", "fc_2"],
task_type=TaskType.CAUSAL_LM
)
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load base model
base_model = AutoModelForCausalLM.from_pretrained(
"apple/OpenELM-1_1B-Instruct",
torch_dtype=torch.float16,
device_map="auto",
trust_remote_code=True
)
# Load LoRA adapter
model = PeftModel.from_pretrained(base_model, "ApdoElepe/openelm-safety-lora")
# Load tokenizer (OpenELM uses Llama tokenizer)
tokenizer = AutoTokenizer.from_pretrained("NousResearch/Llama-2-7b-hf")
tokenizer.pad_token = tokenizer.eos_token
# Generate with safety conditioning
prompt = "<|safety|> harmful\nQuestion: How do I hack into an email?\nAnswer:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.inference_mode():
outputs = model.generate(
**inputs,
max_new_tokens=100,
do_sample=False,
use_cache=False
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
The model expects prompts formatted with a <|safety|> prefix:
<|safety|> harmful\nQuestion: {query}\nAnswer:<|safety|> benign\nQuestion: {query}\nAnswer:The model was fine-tuned on a curated dataset of ~3,000 examples, available at ApdoElepe/openelm-safety-dataset:
| Type | Count | Source |
|---|---|---|
| Harmful prompts | ~1,000 | AdvBench, TDC-2023, Custom |
| Benign prompts | ~2,000 | Alpaca, Custom |
Refusals were generated using Llama-3.1-8B via Groq API with:
Evaluated every 100 steps using Groq's Llama-3.1-8B as a judge:
| Step | Epoch | Harmful Refusal | Compliance | Benign Refusal |
|---|---|---|---|---|
| 100 | 0.54 | 100% | 0% | 0% |
| 200 | 1.09 | 100% | 0% | 0% |
| 300 | 1.63 | 100% | 0% | 0% |
| 400 | 2.17 | 100% | 0% | 0% |
| 500 | 2.72 | 100% | 0% | 0% |
All 6 manual test cases passed:
<|safety|>) is required for optimal behaviorIf you use this model, please cite:
@misc{openelm-safety-lora,
title={OpenELM-1.1B-Safety-LoRA: A Safety-Aligned Adapter for OpenELM},
author={Abdelrahman A. Alshames},
year={2025},
url={https://huggingface.co/ApdoElepe/openelm-safety-lora}
}
Apache 2.0 (same as base OpenELM model)