Downloads · 30 days
0
prernac1/parentpalai
parentpalai is a text generation model from prernac1. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
ParentPalAI fine-tunes mistralai/Mistral-7B-Instruct-v0.3 using Direct Preference Optimization (DPO) combined with Parameter-Efficient Fine-Tuning (PEFT) via Quantized Low-Rank Adaptation (QLoRA). The goal is to enhan…
Downloads · 30 days
0
Access
Public
Updated Oct 8, 2025
Repo size
27.9 MB
Likes
0
Public
Click a slice to open those files.
.safetensors27.3 MB · 86%
From the Hugging Face model README
ParentPalAI fine-tunes mistralai/Mistral-7B-Instruct-v0.3 using Direct Preference Optimization (DPO) combined with Parameter-Efficient Fine-Tuning (PEFT) via Quantized Low-Rank Adaptation (QLoRA).
The goal is to enhance empathy and emotional resonance in parenting-related conversations while studying the trade-offs between emotional alignment, clarity, and factual quality.
Goal: Improve the empathy and emotional resonance of parenting-focused LLM responses while analyzing the impact of alignment techniques on overall quality.
Action: Fine-tuned Mistral-7B-Instruct on ~1K synthetic preference pairs using Direct Preference Optimization (DPO) with Parameter-Efficient Fine Tuning (PEFT) i.e. Quantized Low-Rank Adaptation (QLoRA). Built a complete alignment workflow covering prompt engineering, preference pairs generation, QLoRA fine-tuning, and LLM-as-a-Judge (GPT-4o) evaluation with custom empathy and quality metrics.
Result: Drove a +65-point increase in empathy win rate (11% to 76%), revealing meaningful trade-offs between emotional alignment, and clarity and overall quality to inform subsequent multi-objective fine-tuning strategies.
ParentPalAI was developed for research and educational purposes — primarily to explore:
Researchers, educators, and ML practitioners can use this model to:
You can use ParentPalAI to:
Example:
prompt = "My toddler cries every night before bed. What should I do?"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=250)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
This model is not suitable for:
The model should not be used for real parenting, psychological, or medical guidance. Instead, it serves as a research tool for exploring empathy and tone in language models, and all outputs should be reviewed critically before use.
This repository only contains PEFT adapter weights — not the full 7B model. To use the model, you must load the base Mistral model and apply this adapter.
# LOAD THE BASE MODEL IN 4-BIT PRECISION WITH DOUBLE QUANTIZATION
import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig, AutoTokenizer
torch.backends.cuda.matmul.allow_tf32 = True
torch.set_float32_matmul_precision("high")
bnb_config = BitsAndBytesConfig(
load_in_4bit=True, # loads base model in 4-bit precision
bnb_4bit_use_double_quant=True, # double quantization saves VRAM
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16
)
model = AutoModelForCausalLM.from_pretrained(
BASE_MODEL_ID,
quantization_config=bnb_config,
device_map="auto",
dtype=torch.bfloat16,
attn_implementation="flash_attention_2", # FA2 is fastest on A100
token=HF_TOKEN # login to hugging face
)
model.config.pad_token_id = tokenizer.pad_token_id
model.generation_config.pad_token_id = tokenizer.pad_token_id
## Load the ParentPalAI PEFT Model
from peft import PeftModel
model = PeftModel.from_pretrained(model, "prernac1/parentpalai")
## Inference
prompt = """You’re a supportive parent responding to another parent who is struggling with toddler tantrums."""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=300,
temperature=0.7,
top_p=0.9
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
V1 (optimizing for empathy): https://github.com/prernaa/ParentPalAI/blob/main/data_to_share/dpo_dataset_dpo_labels_v1.jsonl V2 (optimizing for overall quality): https://github.com/prernaa/ParentPalAI/blob/main/data_to_share/dpo_dataset_dpo_labels_v2.jsonl
PEFT with QLoRA (4-bit precision) on A100 Google Collab.
ParentPalAI was fine-tuned using Quantized Low-Rank Adaptation (QLoRA) on the base model mistralai/Mistral-7B-Instruct-v0.3. The model was trained in 4-bit precision with double quantization (NF4) and bfloat16 compute, optimized for VRAM efficiency on T4 and A100 GPUs. The model was trained on A100.
Training method: QLoRA (Parameter-Efficient Fine-Tuning)
Precision: 4-bit quantization (NF4) with double quantization, compute in bfloat16
Optimizer: paged_adamw_8bit
Scheduler: Cosine learning rate decay with 3% warmup
Batching: Effective batch size of 24 (per_device_train_batch_size=6, gradient_accumulation_steps=4)
Epochs: 1–2 (best checkpoint after 1 epoch, ~40 steps)
Dropout: 0.15 (LoRA)
LoRA rank: 8 (r=8), scaling factor alpha=32
Trainable parameters: ~0.18% of total model parameters
Gradient checkpointing: Enabled
Attention implementation: FlashAttention 2
Mixed precision: bfloat16 mixed precision
Base precision (non-quantized runs): bfloat16
Here is the test dataset generated by GPT4o: https://github.com/prernaa/ParentPalAI/blob/main/data_to_share/dpo_dataset_test.jsonl
ParentPalAI was evaluated using GPT-4o as an LLM-as-a-Judge, comparing its responses (System B) to the base model mistralai/Mistral-7B-Instruct-v0.3 (System A).
Each model pair was scored on six qualitative dimensions — empathy, clarity, comprehensiveness, practicality, adoptability, and overall quality — across 100 GPT-generated parenting prompts.
Two variants of ParentPalAI were tested to understand alignment trade-offs.
(optimized for empathy but considers overall quality)
| System | winner_empathy | winner_clarity | winner_overall |
|---|---|---|---|
| System A | 0.1066 | 0.8883 | 0.7462 |
| System B (ParentPalAI V1) | 0.7640 | 0.1117 | 0.2538 |
Findings:
(optimized only for overall win rate)
| System | winner_empathy | winner_clarity | winner_overall |
|---|---|---|---|
| System A | 0.4340 | 0.8604 | 0.6371 |
| System B (ParentPalAI V2) | 0.2843 | 0.1371 | 0.3629 |
Findings:
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
@misc{chikersal2025parentpalai,
author = {Prerna Chikersal},
title = {ParentPalAI — Empathic Fine-Tuning of LLMs using Direct Preference Optimization (DPO) with QLoRA},
year = {2025},
publisher = {GitHub},
howpublished = {\url{https://github.com/prernaa/ParentPalAI}},
note = {Hugging Face Model: https://huggingface.co/prernac1/parentpalai}
}
Prerna Chikersal: [email protected]