Downloads · 30 days
0
yitongl/minimax-h3-nvfp4-lambda-modality
minimax-h3-nvfp4-lambda-modality is a machine learning model from yitongl. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for minimax-h3. The card lists the license as other.
Five end-to-end 50-step generations, one BF16 reference, same prompt and same seed (1101), all on one RTX 5090. They differ in one thing: which rows of the packed [text | video | audio] sequence were used to derive th…
Downloads · 30 days
0
Access
Public
Updated Aug 9, 2026
Repo size
10.6 MB
Likes
0
Public
Click a slice to open those files.
.mp49 MB · 83%
From the Hugging Face model README
Five end-to-end 50-step generations, one BF16 reference, same prompt and same seed (1101), all on
one RTX 5090. They differ in one thing: which rows of the packed [text | video | audio]
sequence were used to derive the per-input-channel smoothing scale lambda.
The question they answer: H3 runs full self-attention over a packed sequence of three modalities,
but a linear layer has one weight, so W * lambda is shared by all three. Per-channel activation
profiles differ per modality — measured correlations between them are near zero or negative — so a
single lambda cannot be right for all of them. Is favouring video worth it?
Answer: no. lambda calibrated on video alone is the worst of the four smoothed configs
end-to-end. Calibrating on all rows is the right default.
| # | file | lambda from | LoRA | lambda spread (p50 max/min) |
|---|---|---|---|---|
| 0 | 0_bf16.mp4 | — | — | — |
| 1 | 1_plain_w4a4.mp4 | none (lambda = 1) | none | 1.0 |
| 2 | 2_lambda_none.mp4 | none (lambda = 1) | rank 32 | 1.0 |
| 3 | 3_lambda_all.mp4 | all rows | rank 32 | 3.1 |
| 4 | 4_lambda_video.mp4 | video rows | rank 32 | 5.9 |
| 5 | 5_lambda_text.mp4 | text rows | rank 32 | 1.6 |
grid.mp4 is all of them on one timeline, captioned.
RESULTS.md has the table. One column in it is easy to misread, so it is worth saying here:
Mean RGB is not a quality metric between two working quantizations. Every run shares BF16's
seed and therefore its initial noise. A faithful quantization stays in the same sample; one that
perturbs the trajectory enough lands in a different — and perfectly plausible — sample. Run 4 is a
coherent video of a different scene, not a broken one, and its large mean-RGB delta is reporting
the scene change, not damage. corr (frame-aligned against BF16) is the column that separates
"same scene, degraded" from "different scene".
Mean RGB was the right instrument earlier, when the failure being chased was a genuinely
washed-out output from a misapplied smooth_factor permutation. It stopped being the right
instrument the moment every candidate started producing a real video.
The calibration's own error column scores each candidate lambda on rows drawn uniformly from the
packed sequence — which is what deepcompressor's OutputsError objective specifies. Uniform
means proportional, and video is 98.6% of the rows. So that column ranks the runs like this
(relative L2 vs bf16, median over 312 layers, rank-32 branch included):
| lambda=1 | lambda=1 +LoRA | lambda=all | lambda=video | lambda=text | |
|---|---|---|---|---|---|
| calibration's own column | 0.1019 | 0.0951 | 0.0942 | 0.0903 | 0.0953 |
lambda=video wins that metric and loses end to end. The metric is not wrong; it is answering
a question about the 98.6%, and the damage is in the other 1.4%.
Re-scoring with the same quantizer on rows sampled per modality
(scripts/calib_sample_modal.py + scripts/score_modal_lambda.py, 4096 rows per modality per
layer, 312 layers) shows what the proportional column could not:
| modality | lambda=1 | lambda=all | lambda=video | lambda=text |
|---|---|---|---|---|
| video | 0.0980 | 0.0976 | 0.0937 | 0.0981 |
| text | 0.0891 | 0.0871 | 0.1140 | 0.0871 |
| audio | 0.0875 | 0.0857 | 0.0975 | 0.0874 |
Change against no smoothing (negative = helps):
| modality | lambda=all | lambda=video | lambda=text |
|---|---|---|---|
| video | −0.4 % | −4.3 % | +0.1 % |
| text | −2.2 % | +27.9 % | −2.2 % |
| audio | −2.1 % | +11.5 % | −0.2 % |
lambda=all is the only column negative in all three rows. lambda=video buys video 4.3 % and
charges text 27.9 % and audio 11.5 % for it. lambda=text is free but pointless — it does nothing
for video.
The worst layers make the trade obvious. lambda=video vs lambda=all, text error:
| layer | text | video |
|---|---|---|
blocks.13.ff.net.2 | 0.0267 → 0.3396 (12.7x) | 0.0886 → 0.0876 |
blocks.12.ff.net.2 | 0.0702 → 0.1429 | 0.0898 → 0.0897 |
blocks.8.ff.net.2 | 0.0807 → 0.1105 | 0.0951 → 0.0875 |
blocks.33.attn.to_q | 0.0702 → 0.0955 | 0.0669 → 0.0657 |
A 12.7x text error for a 1 % video gain, in one layer.
NVFP4 puts 16 consecutive input channels under one FP8 scale set by that group's absmax, so a
channel whose own absmax is far below its group's loses log2(group_absmax / channel_absmax)
bits, and lambda reshapes exactly that profile because the kernel sees X / lambda. The
per-modality channel profiles are uncorrelated to anti-correlated (b0.to_q videotext −0.066;
audio −0.390), and dividing by a vector uncorrelated with your own profile
sharpens it rather than flattening it. The statistics-only proxy
(b25.to_v videoscripts/diag_lambda_crossmodal.py, bits lost, no GPU) agrees with the measurement above: video
−0.437 bits, text +0.216, audio +0.243 under lambda=video.
Text is ~1.4% of the rows, but under full self-attention every video row attends to it. Degrading the conditioning degrades the video conditioned on it — which is how the per-layer video error can fall while the generated video gets worse.
Same as minimax-h3-svdquant-calib: 128 video prompts, 768x1344, 124 frames, 50 steps, rank 32,
39-candidate lambda grid, 100 iterations of low-rank refit against OutputsError. The only thing
varied across runs 2–5 is which modality mask the activation statistics were accumulated under
(scripts/calib_stats_modal.py), and hence which absmax feeds the lambda grid.
Run 1 is lambda = 1 with the low-rank branch zeroed — the floor, plain W4A4.
Run 2 is lambda = 1 with the rank-32 branch refit against the raw weight — it isolates the
low-rank contribution from the smoothing contribution.
Base model: MiniMaxAI/MiniMax-H3 @ bfc8ed0353f5a9733be73e6b2c98ec0948195b86.
SVDQW4A4Linear.smooth_factor is addressed by the kernel in MMA-interleaved channel order,
while every other tensor picks that interleave up from NunchakuWeightPacker. Writing lambda in
natural order gives 12 of every 16 channels another channel's lambda. Details and the permutation
are in the minimax-h3-svdquant-calib README. Everything here was produced with that fix applied.
Note the interaction with this page: a sharper lambda makes that bug worse, so before the fix,
lambda=video looked catastrophic and lambda=text looked fine — for a reason that had nothing to
do with modality.