Downloads · 30 days
20
27% of all-time downloads
SeeYangZhi/BART-Base-Sarcasm-Rewriter
BART-Base-Sarcasm-Rewriter is a summarization model from SeeYangZhi. Use it when you need a shorter version of a longer text. The card lists the license as mit.
Baseline BART-base supervised fine-tuning on sarcastic-non-sarcastic headline pairs. No context enhancement, no RL.
Downloads · 30 days
20
27% of all-time downloads
All-time downloads
74
Public
Parameters
139M
558 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors558 MB · 99%
From the Hugging Face model README
Baseline BART-base supervised fine-tuning on sarcastic->non-sarcastic headline pairs. No context enhancement, no RL.
Part of the Project LLMao sarcasm style transfer suite (CS4248 Team 14, NUS AY2025/26 S2). This model rewrites sarcastic news headlines as neutral, factual equivalents while preserving the underlying meaning.
Input: A sarcastic news headline Output: A non-sarcastic rewrite
Example:
facebook/bart-base (139M params)sar_to_non (original).num_beams=4, max_length=128.from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
model_id = "SeeYangZhi/BART-Base-Sarcasm-Rewriter"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)
headline = "Area Man Passionate Defender Of What He Imagines Constitution To Be"
inputs = tokenizer(headline, return_tensors="pt", truncation=True, max_length=128)
outputs = model.generate(**inputs, max_length=128, num_beams=4)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Evaluated on a 2,857-sample held-out test split alongside 13 other model variants (BART, T5 baselines, LLaMA 3.2, ablation studies). Metrics include:
| Metric | Direction |
|---|---|
| Hard Flip Rate (% of samples where sarcasm was removed) | higher ↑ |
| Semantic Similarity (all-MiniLM-L6-v2 cosine) | higher ↑ |
| BLEU vs input (lower = more genuine rewriting) | lower ↓ |
| Perplexity (GPT-2) | lower ↓ |
| Normalized edit distance | higher ↑ |
| Paraphrase score (low = real rewriting) | lower ↓ |
Full per-variant numbers are published alongside the Project LLMao webapp.
SeeYangZhi/Llama-3.2-1B-Sarcasm-Rewriter — instruction-tuned LLaMA variantSeeYangZhi/BART-Base-Sarcasm-Rewriter — supervised baselineSeeYangZhi/BART-Base-CE-Sarcasm-Rewriter — context-enhanced SFTSeeYangZhi/BART-Base-RL-Sarcasm-Rewriter — REINFORCE on top of baselineSeeYangZhi/BART-Base-CE-RL-Sarcasm-Rewriter — CE + RL (best)MIT, inheriting from facebook/bart-base. The NHDSD dataset is used under its
original research-use terms.