Downloads · 30 days
0
prernac1/pretendparentai
pretendparentai is a text generation model from prernac1. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
PretendParentAI is a fine-tuned variant of mistralai/Mistral-7B-Instruct-v0.3, adapted using Quantized Low-Rank Adaptation (QLoRA) on Reddit parenting discussions. It produces more empathetic, warm, and relatable pare…
Downloads · 30 days
0
Access
Public
Updated Oct 8, 2025
Repo size
439 MB
Likes
1
Public
Click a slice to open those files.
.safetensors439 MB · 99%
From the Hugging Face model README
PretendParentAI is a fine-tuned variant of mistralai/Mistral-7B-Instruct-v0.3, adapted using Quantized Low-Rank Adaptation (QLoRA) on Reddit parenting discussions.
It produces more empathetic, warm, and relatable parenting advice.
Goal: Explore how instruction fine-tuning can enhance warmth, relatability, and storytelling in parenting advice LLMs, while assessing trade-offs with factual precision.
Action: Fine-tuned Mistral-7B-Instruct on ~40K curated Reddit parenting Q&A pairs (Alpaca format), using Supervised Fine Tuning (SFT) with Parameter-Efficient Fine Tuning (PEFT) i.e. Quantized Low-Rank Adaptation (QLoRA). Built a full instruction-tuning pipeline including Reddit data curation, efficient training/inference using QLoRA, and LLM-as-a-Judge evaluation across empathy, relatability, and other metrics.
Result: Produced highly human-like, narrative responses that excelled in empathy (30% to 70%) and relatability (2% to 98%), though often over-personalized or hallucinated personal anecdotes—yielding key insights into the tension between emotional alignment and factual grounding in instruction tuning when using human-generated data (e.g. from reddit).
PretendParentAI was developed for research and educational purposes — primarily to explore:
Researchers, educators, and ML practitioners can use this model to:
You can use PretendParentAI to:
Example:
prompt = "My toddler cries every night before bed. What should I do?"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=250)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
This model is not suitable for:
PretendParentAI can hallucinate personal details — such as referring to imaginary “sons,” “daughters,” or “partners” — because it imitates how Reddit users often share personal anecdotes. These outputs should not be interpreted as factual or autobiographical.
The model should not be used for real parenting, psychological, or medical guidance. Instead, it serves as a research tool for exploring empathy and tone in language models, and all outputs should be reviewed critically before use.
This repository only contains PEFT adapter weights — not the full 7B model.
To use the model, you must load the base Mistral model and apply this adapter.
mistralai/Mistral-7B-Instruct-v0.3## Load the base model
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
BASE_MODEL_ID = "mistralai/Mistral-7B-Instruct-v0.3"
torch.backends.cuda.matmul.allow_tf32 = True
torch.set_float32_matmul_precision("high")
tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
BASE_MODEL_ID,
device_map="auto",
dtype=torch.bfloat16,
attn_implementation="flash_attention_2", # fastest on A100
token=HF_TOKEN # or login via huggingface-cli login
)
model.config.pad_token_id = tokenizer.pad_token_id
model.generation_config.pad_token_id = tokenizer.pad_token_id
## Load the PretendParentAI PEFT Model
from peft import PeftModel
model = PeftModel.from_pretrained(model, "your-username/pretendparentai")
## Inference
prompt = """You’re a supportive parent responding to another parent who is struggling with toddler tantrums."""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=300,
temperature=0.7,
top_p=0.9
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Trained on Reddit data from r/parenting. Contact author for dataset info. It can't be publicly shared.
PEFT with QLoRA (4-bit precision) on A100 Google Collab.
PretendParentAI was fine-tuned using Quantized Low-Rank Adaptation (QLoRA) on the base model mistralai/Mistral-7B-Instruct-v0.3.
The model was trained in 4-bit precision with double quantization (NF4) and bfloat16 compute, optimized for VRAM efficiency on T4 and A100 GPUs.
The model was trained on A100.
Training method: QLoRA (Parameter-Efficient Fine-Tuning)
Precision: 4-bit quantization (NF4) with double quantization, compute in bfloat16
Optimizer: paged_adamw_8bit
Scheduler: Cosine learning rate decay with 3% warmup
Batching: Effective batch size of 24 (per_device_train_batch_size=6, gradient_accumulation_steps=4)
Epochs: 1–2 (best checkpoint after 1 epoch, ~1600 steps)
Dropout: 0.05 (LoRA)
LoRA rank: 16 (r=16), scaling factor alpha=64
Trainable parameters: ~1.12% of total model parameters
Gradient checkpointing: Enabled
Attention implementation: FlashAttention 2
Mixed precision: bfloat16 mixed precision
Base precision (non-quantized runs): bfloat16
We compute BERTScore, Rouge, and BLEU, as well as carry out LLM-as-a-judge evaluation.
Test dataset: https://github.com/prernaa/PretendParentAI/blob/main/data_to_share/reddit_gpt_test_samples.jsonl
[More Information Needed]
BLEU
| Model | BLEU | P@1 | P@2 | P@3 | P@4 | Length Ratio |
|---|---|---|---|---|---|---|
| Mistral Instruct v0.3 | 0.00695 | 0.1661 | 0.0140 | 0.0023 | 0.0004 | 1.59 |
| PretendParentAI | 0.00624 | 0.1740 | 0.0169 | 0.0016 | 0.0003 | 1.73 |
ROUGE
| Model | ROUGE-1 | ROUGE-2 | ROUGE-L | ROUGE-Lsum |
|---|---|---|---|---|
| Mistral Instruct v0.3 | 0.1774 | 0.0161 | 0.0977 | 0.1040 |
| PretendParentAI | 0.2068 | 0.0215 | 0.1057 | 0.1059 |
BERTScore (avg.)
| Model | Precision | Recall | F1 |
|---|---|---|---|
| Mistral Instruct v0.3 | 0.8334 | 0.8440 | 0.8386 |
| PretendParentAI | 0.8323 | 0.8462 | 0.8391 |
ROUGE and BERTScore show marginal improvement; BLEU remains low (expected for open-ended tasks).
Evaluation used GPT-4o as an LLM judge to compare responses from:
Each system was scored on:
helpfulness, empathy/tone, creativity, clarity, relatability, adoptability, and overall preference.
| winner_helpfulness | winner_empathy_tone | winner_creativity | winner_clarity | winner_relatability | winner_adoptability | winner_overall | |
|---|---|---|---|---|---|---|---|
| System A | 0.90 | 0.30 | 0.47 | 0.90 | 0.02 | 0.72 | 0.72 |
| System B | 0.10 | 0.70 | 0.48 | 0.10 | 0.98 | 0.28 | 0.28 |
| Tie | 0.00 | 0.00 | 0.05 | 0.00 | 0.00 | 0.00 | 0.00 |
Interpretation
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
@misc{chikersal2025pretendparentai,
author = {Prerna Chikersal},
title = {PretendParentAI: Instruction Fine-Tuning Mistral-7B on Reddit Parenting Data using QLoRA},
year = {2025},
note = {GitHub repository}
}
Prerna Chikersal: [email protected]