Downloads · 30 days
80
18% of all-time downloads
spandyie/amadablam-322m-instruct
amadablam-322m-instruct is a text generation model from spandyie. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
DPO-tuned Ama Dablam, starting from the instruction-tuned checkpoint: preference-trained to prefer well-formed, on-script, appropriately concise responses over degenerate ones (run-on translations, script drift, near-…
Downloads · 30 days
80
18% of all-time downloads
All-time downloads
448
Public
Parameters
322M
2.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
DPO-tuned Ama Dablam, starting from the instruction-tuned checkpoint: preference-trained to prefer well-formed, on-script, appropriately concise responses over degenerate ones (run-on translations, script drift, near-duplicate low-signal pairs). Trained in two stages — a short DPO-format SFT warmup, then DPO proper (beta=0.1) against a frozen reference — on 14,152 preference pairs across Nepali, Maithili, and Bhojpuri, rendered 85/5/10 across Devanagari/IAST/phonetic script.
| metric | value |
|---|---|
| preference accuracy (val, n=744) | 91.4% (mean margin +4.85) |
| — by language | ne 94.9% · mai 91.0% · bho 87.3% |
| — by script | deva 92.8% · phon 86.3% · iast 77.1% |
| script fidelity (replies in the prompt's script) | 95% (19/20 sampled) |
| response length | mean 32.3 tokens, 95% properly terminated |
Forgetting guardrail (drift vs. the pre-DPO checkpoint on 6 pinned pretrain val shards + phonetic-ne) passed for every metric, but drift was uniformly positive (a small, one-directional erosion, not noise): worst cases IAST-Maithili +0.035 bpb, phonetic-Nepali +0.027 bpb; Devanagari ne/mai/bho drift stayed under +0.01 bpb.
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("spandyie/amadablam-322m-dpo")
model = AutoModelForCausalLM.from_pretrained("spandyie/amadablam-322m-dpo",
trust_remote_code=True)
msgs = [{"role": "user", "content": "स्वस्थ रहन के गर्नुपर्छ?"}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt")
out = model.generate(ids, max_new_tokens=200)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))
generation_config.json sets use_cache: false.
Overriding to use_cache=True silently produces incoherent, context-free output
(only the last token is fed back each step) rather than raising an error.