Downloads · 30 days
7
23% of all-time downloads
hash-map/got_QA_fine_tuned_model
got_QA_fine_tuned_model is a question answering model from hash-map. Use it when the input is a question plus a passage. It is set up for peft. The card lists the license as mit.
Model name: hash-map/gotmodel Base model: google/gemma-2-2b-it Fine-tuning method: QLoRA (via PEFT) Task: Contextual Question Answering on Game of Thrones Summary: A lightweight instruction-tuned question-answering mo…
Downloads · 30 days
7
23% of all-time downloads
All-time downloads
31
Public
Repo size
366 MB
Likes
0
Public
Click a slice to open those files.
.pt166 MB · 45%
From the Hugging Face model README
Model name: hash-map/got_model
Base model: google/gemma-2-2b-it
Fine-tuning method: QLoRA (via PEFT)
Task: Contextual Question Answering on Game of Thrones
Summary: A lightweight instruction-tuned question-answering model specialized in the Game of Thrones / A Song of Ice and Fire universe. It generates concise, faithful answers when given relevant context + a question.
Description:
This model was fine-tuned on the hash-map/got_qa_pairs dataset using QLoRA (4-bit quantization + Low-Rank Adaptation) to keep memory usage low while adapting the powerful gemma-2-2b-it model to answer questions about characters, events, houses, lore, battles, and plot points — only when provided with relevant context.
It is not a general-purpose LLM and performs poorly on questions without appropriate context or outside the GoT domain.
A simple keyword-based lexical retrieval system is provided to help select relevant context chunks:
import re
import json
from collections import defaultdict, Counter
CHUNKS_FILE = "/kaggle/input/got-dataset/contexts.json" # list of {text, source, chunk_id}
def tokenize(text):
return re.findall(r"\b[a-zA-Z]{3,}\b", text.lower())
contexts = []
token_to_ctx = defaultdict(list)
with open(CHUNKS_FILE, "r", encoding="utf-8") as f:
data = json.load(f)
for idx, item in enumerate(data):
text = item["text"]
contexts.append(item)
for tok in tokenize(text):
token_to_ctx[tok].append(idx)
print(f"Indexed {len(contexts)} chunks")
def retrieve_2_contexts(question, token_to_ctx, contexts):
q_tokens = tokenize(question)
scores = Counter()
for tok in q_tokens:
for ctx_id in token_to_ctx.get(tok, []):
scores[ctx_id] += 1
if not scores:
return ""
top_ids = [cid for cid, _ in scores.most_common(2)]
return " ".join([contexts[cid]["text"] for cid in top_ids])
This is a basic sparse retrieval method (similar to TF-IDF without IDF). can create faiss for better retreival using these contexts
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
import torch
# Replace with your actual repo
model_name = "hash-map/got_model"
tokenizer = AutoTokenizer.from_pretrained(model_name)
base_model = AutoModelForCausalLM.from_pretrained(
"google/gemma-2-2b-it",
torch_dtype=torch.bfloat16,
device_map="auto"
)
model = PeftModel.from_pretrained(base_model, model_name)
def answer_question(context: str, question: str, max_new_tokens=96) -> str:
prompt = f"""Context:
{context}
Question:
{question}
Answer:"""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=max_new_tokens,
do_sample=False,
temperature=0.0,
eos_token_id=tokenizer.eos_token_id
)
answer = tokenizer.decode(outputs[0], skip_special_tokens=True)
# Extract only the answer part after "Answer:"
return answer.split("Answer:")[-1].strip()
# Example
context = retrieve_2_contexts("Who killed Joffrey Baratheon?", token_to_ctx, contexts)
print(answer_question(context, "Who killed Joffrey Baratheon?"))
Recommendations:
@misc{got-qa-gemma2-2026,
author = {Appala Sai Sumanth},
title = {Gemma-2-2b-it Fine-tuned for Game of Thrones Question Answering},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/hash-map/got_model}}
}
transformers >= 4.42peft 0.13.2torch >= 2.1bitsandbytes >= 0.43 (for 4-bit inference if desired)