Downloads · 30 days
148
71% of all-time downloads
whosouravsharma/paper-qa-lora
paper-qa-lora is a text generation model from whosouravsharma. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
A QLoRA adapter for Qwen/Qwen2.5-1.5B-Instruct. It was trained to answer questions about a research paper only from the supplied context, to cite the sentence it relied on, and to say plainly when the context doesn't…
Downloads · 30 days
148
71% of all-time downloads
All-time downloads
208
Public
Repo size
28.9 MB
Likes
0
Public
Click a slice to open those files.
.json11.4 MB · 56%
From the Hugging Face model README
A QLoRA adapter for Qwen/Qwen2.5-1.5B-Instruct.
It was trained to answer questions about a research paper only from the
supplied context, to cite the sentence it relied on, and to say plainly
when the context doesn't contain the answer. It is the answer generator
behind the Paper QA
app.
| What it's good at | Following the app's output format: answers end with a Source: "…" citation (86% of the time vs. 0% for the base model), and it refuses on some unanswerable questions (4 of 12 vs. 0 for base) |
| Where it falls short | Answer quality. It is less faithful to the context than the untuned base model (70.0% vs. 80.2%), and it hallucinates more (see Evaluation) |
| Status | Experimental. The cause of the regression has been identified in the training data (see Known issue) and not yet fixed |
| Size | 4.36M trainable parameters (0.49% of the model), an 8.7 MB adapter |
| License | Apache 2.0, the same as the base model. Training data is QASPER (CC BY 4.0) |
| Developed by | whosouravsharma |
| Model type | LoRA adapter (QLoRA, 4-bit NF4) for a 1.5B-parameter decoder-only chat model |
| Base model | Qwen/Qwen2.5-1.5B-Instruct |
| Adapter | Rank 16, alpha 32, dropout 0.05, on q_proj, k_proj, v_proj, o_proj |
| Language | English |
| Domain | NLP research papers (QASPER) |
| Version | v2, trained 2026-09-04. It replaced v1 (2026-08-30) in place |
| License | Apache 2.0 |
| Resource | Link |
|---|---|
| Demo | Paper QA Space |
| Serving backend | paper-qa-rag Space (retrieval, generation with this adapter, judging) |
| Training data | whosouravsharma/paper-qa-qasper-sft |
Source: "…" citation after each
answer.The adapter expects the same system prompt and message layout it was served with:
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
base = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen2.5-1.5B-Instruct", torch_dtype=torch.float16
).to("cuda")
model = PeftModel.from_pretrained(base, "whosouravsharma/paper-qa-lora").eval()
system = (
"You are a research paper assistant. Answer the question using only the "
"provided context. Quote the exact sentence(s) you relied on as a "
"citation. If the context does not contain the answer, say so plainly "
"instead of guessing."
)
context = "…retrieved passages from the paper…"
question = "What datasets are used for experiments?"
messages = [
{"role": "system", "content": system},
{"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"},
]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
).to(model.device)
output = model.generate(**inputs, max_new_tokens=256, do_sample=False,
pad_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
The expected output is a one- or two-sentence answer, then
Source: "<quoted sentence>". When it can't answer, it replies:
"I cannot find an answer to this question in the provided context."
test split: 107
answerable and 12 unanswerable. The reference answers are QASPER's own
human annotations. None of the 20 papers appear in the training or
validation data; this was checked by paper ID.gpt-5.6-luna scores each answer from 1 to 10 on six criteria,
reported here as percentages. In the offline benchmark the judge sees the
model's answer and the human reference as "Answer A" and "Answer B" in a
random order, and is told to weight faithfulness to the paper over
similarity to the reference.All systems here are given the same passages, retrieved by TF-IDF over the paper's paragraphs, and judged by the same method.
| Qwen 1.5B, untuned | + LoRA v1 | + LoRA v2 (this repo) | gpt-5.6-luna (hosted)¹ | |
|---|---|---|---|---|
| Factual correctness | 77.8 | 74.6 | 71.2 | 95.5 |
| Faithfulness to context | 80.2 | 79.8 | 70.0 | 96.5 |
| Completeness | 68.6 | 64.0 | 69.2 | 90.9 |
| Relevance | 88.5 | 86.8 | 89.6 | 98.8 |
| Clarity | 90.0 | 89.3 | 89.0 | 97.1 |
| No hallucination | 82.6 | 86.6 | 73.4 | 97.3 |
| Preferred over the human reference | 68.9% | 52.1% | 60.5% | 93.3% |
Ends with a Source: citation | 0% | 80.7% | 85.7% | 0% |
| Refuses on unanswerable (of 12) | 0 | 4 | 4 | 0 |
| Refuses on answerable (of 107) ↓ | 0 | 19 | 13 | 0 |
Scores are percentages from the judge's rubric. Bold marks the best of the three Qwen variants.
¹ The hosted model is included as a reference point only. It is the same model as the judge, and models tend to rate their own answers highly, so its scores are probably inflated.
Is the v2 regression real? Yes, on three criteria. Paired bootstrap over the 119 questions (v2 minus untuned base, 5,000 resamples):
| Criterion | Difference | 95% CI |
|---|---|---|
| Faithfulness | −10.2 pts | [−16.4, −3.8] |
| No hallucination | −9.2 pts | [−15.7, −2.8] |
| Factual correctness | −6.6 pts | [−12.9, −0.5] |
| Completeness | +0.7 pts | [−4.5, +6.0] |
| Relevance | +1.1 pts | [−2.3, +4.5] |
| Clarity | −1.0 pts | [−3.2, +1.3] |
For scale: the untuned base model was scored twice in separate runs, and its scores moved by up to about 1 point between them, so judge noise alone is much smaller than these gaps.
What the numbers say
These results come from the deployed pipeline on 2026-09-04: the same 119 questions, with the real arXiv PDFs uploaded to the live app. There were 0 failures across 20 uploads and 119 questions.
| Criterion | Score |
|---|---|
| Factual correctness | 71.2 |
| Faithfulness to context | 75.5 |
| Completeness | 64.2 |
| Relevance | 90.5 |
| Clarity | 89.8 |
| No hallucination | 78.6 |
These are not comparable with the offline table. Production uses FAISS over OpenAI embeddings, not TF-IDF, and puts the paper's title and a summary at the top of every context. Its live judge also scores each answer against that context alone, with no human reference.
| Behaviour | Production |
|---|---|
Ends with a Source: citation | 95.0% (113/119) |
| …of which the quoted "source" is the paper title | 81 of 113 |
| …of which it quotes an actual passage | 31 of 113 |
| Refuses on unanswerable | 2 of 12 |
| Refuses on answerable | 3 of 107 |
Two production caveats
Paper title: …
at the top of every context, and the model usually quotes that line. So
the 95% citation rate greatly overstates how often an answer points at real
evidence. In the offline benchmark, where contexts have no title line, it
cited actual passages.Latency (production, per question, n = 119)
| Stage | p50 | p90 | max |
|---|---|---|---|
| Retrieval | 0.19 s | 0.40 s | 3.7 s |
| Generation (this adapter, T4) | 2.6 s | 5.0 s | 13.4 s |
| Judging | 2.1 s | 3.3 s | 9.3 s |
| End to end | 6.8 s | 10.0 s | 17.7 s |
A follow-up check traced the quality regression to the training targets
themselves. gpt-5.6-luna was asked whether each target answer is fully
supported by the context it was paired with, over 300 randomly sampled
answerable training examples:
| Training data | Targets with claims the context doesn't support |
|---|---|
| Raw QASPER answers (used for v1) | 17.7% |
| LLM-rewritten answers (used for v2) | 35.0% (105 of 300) |
Planned fix, not yet done:
whosouravsharma/paper-qa-qasper-sft
was built from QASPER's
train and validation splits (CC BY 4.0). Each annotator answer becomes
one chat example:
Source: "<first evidence quote>".| Split | Examples | Refusal examples | Papers | Median answer length |
|---|---|---|---|---|
| train | 2,589 | 281 | 878 | 23 words |
| validation | 1,715 | 163 | 281 | 21 words |
| Method | QLoRA with TRL SFTTrainer: base weights 4-bit NF4, LoRA adapters trained in fp16 |
| Epochs | 3 (486 optimizer steps) |
| Batch size | 4 per step × 4 gradient-accumulation steps = 16 |
| Learning rate | 2e-4 |
| Max sequence length | 2,048 tokens (context + question + answer; p99 ≈ 1,262) |
| Evaluation | Validation loss each epoch; the final epoch is kept |

| Epoch | Validation loss | Validation token accuracy |
|---|---|---|
| 1 | 1.297 | 72.4% |
| 2 | 1.287 | 72.6% |
| 3 | 1.286 | 72.6% |
Validation loss flattens after the first epoch. The loss measures how closely the model imitates the training targets. Since about a third of those targets contain unsupported claims, a lower loss doesn't mean better grounding.
state.json and loss_curve.json in this repository record the exact
configuration and the full loss log.
| Hardware | 1× NVIDIA L4 (Hugging Face Jobs, l4x1) |
| Training time | 56 min (3,344 s) |
| Carbon emitted | Not measured |
Training and evaluation data come from QASPER:
@inproceedings{Dasigi2021ADO,
title = {A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers},
author = {Pradeep Dasigi and Kyle Lo and Iz Beltagy and Arman Cohan and Noah A. Smith and Matt Gardner},
year = {2021}
}
Open a discussion in the Community tab.