Downloads · 30 days
0
AlexWortega/capabilityvectors-qwen3-4b
capabilityvectors-qwen3-4b is a machine learning model from AlexWortega. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
Attention-only LoRA adapters (q,k,v,oproj, rank 32, alpha 64) for Qwen3-4B-Instruct-2507, trained on identical DeepScaleR math rollouts (math-verify reward) for the study "Same Data, Different Losses, Same Circuits?"…
Downloads · 30 days
0
Access
Public
Updated Jun 21, 2026
Repo size
2.6 GB
Likes
0
Public
Click a slice to open those files.
.safetensors2.6 GB · 100%
From the Hugging Face model README
Attention-only LoRA adapters (q,k,v,o_proj, rank 32, alpha 64) for Qwen3-4B-Instruct-2507,
trained on identical DeepScaleR math rollouts (math-verify reward) for the study
"Same Data, Different Losses, Same Circuits?" — companion to
github.com/AlexWortega/capabilityvectors.
All 28 adapters are trained in one consistent setup, so their LoRA deltas (ΔW = (α/r)·B·A) are directly comparable in weight space.

SFT/RFT/RIFT colinear (0.94–0.98); DFT ~0.55; Offline GRPO 0.71–0.80 to the cluster; DPO near-orthogonal (≤0.13). Online GRPO & DAPO are each near-orthogonal to every offline loss (cos ≈0.02) and to each other (−0.16) — orthogonal-fraction off SFT 0.998/0.995 vs 0.69 for offline GRPO. On-policy sampling, not the group-relative loss, drives the departure from SFT. (Small negative cosines ≈ orthogonal: the on-policy ΔW are ~10× smaller in norm.)

Same loss at two seeds has low raw cosine, but the top-1 output direction agrees at 0.99 and the two seeds sit in the same basin (no linear-mode barrier, midpoint +0.004) — the low cosine is a LoRA input-init artifact, not a different solution. A 10× LR change rotates ΔW (cos ≈0.55), it is not a pure rescaling.
adapters/<method>_lr<lr>_s<seed>/)| family | method | grid |
|---|---|---|
| offline (reward-weighted MLE) | sft, rft | lr {5e-7,5e-6,5e-5} × seed {42,123} |
| offline (other) | dft, rift, offgrpo (offline GRPO), dpo | seed 42, paper LR (dpo 5e-7) |
| online RL | grpo (online GRPO), dapo (online DAPO) | lr {5e-7,5e-6,5e-5} × seed {42,123} |
offgrpo = offline GRPO; grpo = online GRPO (on-policy rollouts, group-relative advantage,
TRL + vLLM); dapo = online DAPO (clip-higher, no-KL, token-level, dynamic sampling).
| method | GSM8K | AIME26 |
|---|---|---|
| base instruct | 94.0 | 16.7 |
| SFT / Offline GRPO | 87.6 / 87.3 | 6.7 |
| DPO | 94.2 | 13.3 |
| Online GRPO | 93.7 | 20.0 |
| Online DAPO | 93.3 | 16.7 |
Reward-orthogonal methods (DPO, online GRPO/DAPO) keep the base 93–94% on GSM8K across the full lr×seed grid (91–94%); the SFT direction drops it to ~87%.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = "Qwen/Qwen3-4B-Instruct-2507"
m = AutoModelForCausalLM.from_pretrained(base, torch_dtype="bfloat16", device_map="auto")
m = PeftModel.from_pretrained(m, "AlexWortega/capabilityvectors-qwen3-4b", subfolder="adapters/grpo_lr5e-6_s42")
tok = AutoTokenizer.from_pretrained(base)
See results/RESULTS.md for full tables and results/figures/ for all plots.