Downloads · 30 days
6
19% of all-time downloads
kunjcr2/gpt2_conv
gpt2_conv is a question answering model from kunjcr2. Use it when the input is a question plus a passage. It is set up for peft. The card lists the license as mit.
This project fine-tunes the gpt2-medium language model to support natural, casual conversational dialogue using PEFT + LoRA.
Downloads · 30 days
6
19% of all-time downloads
All-time downloads
32
Public
Repo size
9.5 MB
Likes
1
Public
Click a slice to open those files.
.pt6.3 MB · 44%
From the Hugging Face model README
This project fine-tunes the gpt2-medium language model to support natural, casual conversational dialogue using PEFT + LoRA.
gpt2-mediumgpt2 (same as base model)| Metric | Value |
|---|---|
| Global Steps | 2611 |
| Final Training Loss | 2.185 |
| Training Runtime | 430.61 seconds |
| Samples/sec | 138.41 |
| Steps/sec | 17.32 |
| Total FLOPs | 1.12 × 10¹⁵ |
| Epochs | 7.0 |
These metrics reflect final performance after complete training.
Chat with the model using the talk() function below:
def talk(model=peft_model, tokenizer=tokenizer, device=device):
print("Start chatting with the bot! Type 'exit' to stop.\n")
while True:
question = input("You: ")
if question.lower() == "exit":
print("Goodbye!")
break
prompt = f"User: {question}\nBot:"
inputs = tokenizer(prompt, return_tensors="pt").to(device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=20,
do_sample=True,
temperature=0.7,
top_p=0.9,
pad_token_id=tokenizer.eos_token_id
)
response = tokenizer.decode(
outputs[0][inputs["input_ids"].shape[-1]:],
skip_special_tokens=True
)
# Clean response
response = response.split(".")
response = ".".join(response[:-1]) + "."
print("Bot:", response.strip())
To use this model locally:
pip install transformers peft accelerate