Downloads · 30 days
16
3% of all-time downloads
willamazon1/sdft-tau-lora-iter240
sdft-tau-lora-iter240 is a text generation model from willamazon1. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Qwen3-8B, multi-stage SFT (SDFT) checkpoint with a tau-bench RL LoRA adapter merged in. Full weights, ready to load with transformers / SGLang / vLLM — no PEFT needed.
Downloads · 30 days
16
3% of all-time downloads
All-time downloads
466
Public
Parameters
8.2B
16.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.4 GB · 100%
From the Hugging Face model README
Qwen3-8B, multi-stage SFT (SDFT) checkpoint with a tau-bench RL LoRA adapter merged in.
Full weights, ready to load with transformers / SGLang / vLLM — no PEFT needed.
This is iteration 240 of the same run that produced
willamazon1/sdft-tau-lora-iter160
(80 more RL steps).
| Stage | What |
|---|---|
| Base | Qwen/Qwen3-8B-Base |
| SDFT | multi-stage SFT chain: Math → Sea → Search → TauSFT → Tau-IF |
| RL | GSPO on tau-bench retail (train split), LoRA-only (base frozen), iteration 240 |
linear_qkv, linear_proj, linear_fc1, linear_fc2 on all 36 layers
(144 modules, 288 tensors)low_var_kl), eps-clip 0.2/0.25The adapter was merged in Megatron parameter space (W += 2.0 · B·A per LoRA'd module,
into the frozen base weights carried by the same iter_0000240 torch_dist checkpoint), then
converted to HuggingFace safetensors. Merging before conversion avoids having to un-fuse the
GQA-interleaved QKV and the gate/up split by hand. LayerNorm weights are untouched — the LoRA
delta applies to the post-LN matmuls only.
Verification performed on this export:
lora_B (max |lora_B| = 1.383e-04),
so the adapter is genuinely trained rather than sitting at its zero init‖ΔW‖/‖W‖: min 4.80e-05, median 6.35e-05, max 1.40e-04
(larger than iter 160's 3.76e-05 / 5.01e-05 / 1.11e-04, as expected after 80 more steps)from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "willamazon1/sdft-tau-lora-iter240"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
Note this is derived from a base (non-instruct) Qwen3 checkpoint plus SFT/RL stages; use the same prompt format as the tau-bench agent it was trained with rather than assuming a generic chat template.