Downloads · 30 days
18
32% of all-time downloads
harshbhatt7585/arithmetic-king-1b
arithmetic-king-1b is a text generation model from harshbhatt7585. Use it when you need the model to write or continue text. It is set up for peft.
Repository id: harshbhatt7585/arithmetic-king-1b.
Downloads · 30 days
18
32% of all-time downloads
All-time downloads
56
Public
Repo size
62.3 MB
Likes
0
Public
Click a slice to open those files.
.safetensors45.1 MB · 72%
From the Hugging Face model README
Repository id: harshbhatt7585/arithmetic-king-1b.
This model is a PEFT LoRA adapter trained with TRL GRPO on synthetic arithmetic episodes. It is tuned to answer in XML format:
<reasoning>...</reasoning><answer>...</answer>This repo contains adapter weights only (not full base model weights). Use with base model:
meta-llama/Llama-3.2-1B-InstructGRPOTrainerfrom transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
base_id = "meta-llama/Llama-3.2-1B-Instruct"
adapter_id = "harshbhatt7585/arithmetic-king-1b"
tokenizer = AutoTokenizer.from_pretrained(base_id)
base_model = AutoModelForCausalLM.from_pretrained(base_id)
model = PeftModel.from_pretrained(base_model, adapter_id)
prompt = "Solve: (12 + 3) * 2. Return XML with <reasoning> and <answer>."
inputs = tokenizer(prompt, return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(out[0], skip_special_tokens=True))
Adapter usage inherits base model license and terms:
meta-llama/Llama-3.2-1B-Instruct