Downloads · 30 days
292
100% of all-time downloads
willamazon1/Qwen3-8B-SDFT-Math-LoRA-new
Qwen3-8B-SDFT-Math-LoRA-new is a text generation model from willamazon1. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Qwen3-8B (dense, 36 layers, 8.19B params) tuned for mathematical reasoning. This repo holds fully merged bf16 weights — the LoRA adapter has been folded into the base matrices, so it is a drop-in replacement for Qwen/…
Downloads · 30 days
292
100% of all-time downloads
All-time downloads
292
Public
Parameters
8.2B
16.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.4 GB · 100%
From the Hugging Face model README
Qwen3-8B (dense, 36 layers, 8.19B params) tuned for mathematical reasoning. This
repo holds fully merged bf16 weights — the LoRA adapter has been folded into
the base matrices, so it is a drop-in replacement for Qwen/Qwen3-8B-Base; no
PEFT adapter loading is required.
Qwen/Qwen3-8B-Base.q/k/v, o_proj, gate/up_proj and down_proj of every layer.
The stage-1 adapter was merged into the SFT weights.Trained with slime on Megatron-LM (TP=4, CP=2, bf16).
| Parameters | 8.19 B |
| dtype | bfloat16 |
| Tensors | 399 (4 safetensors shards, 16.4 GB) |
| Vocab | 151936 (Megatron embedding padding stripped) |
| Stage-2 relative weight change, per LoRA'd matrix (‖Δ‖/‖W‖) | 2.8e-5 – 7.7e-5 |
Non-finite check: 0 NaN/Inf tensors across all shards.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "willamazon1/Qwen3-8B-SDFT-Math-LoRA-new"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16).to("cuda")
prompt = (
"Question: Natalia sold clips to 48 friends in April, and then she sold "
"half as many clips in May. How many clips did Natalia sell altogether "
"in April and May?\nAnswer:"
)
ids = tok(prompt, return_tensors="pt").input_ids.cuda()
out = model.generate(ids, max_new_tokens=128, do_sample=False)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))
This is a base-style (completion) model, not an instruction-tuned chat
model: prompt it with Question: ... \nAnswer: style completions rather than a
chat template.
Stage 2 is a short RL continuation (60 steps), so the weight delta over the
stage-1 model is small. The model inherits the biases and knowledge cutoff of
Qwen3-8B-Base, and its math answers are not guaranteed correct — verify
outputs before relying on them.