Downloads · 30 days
0
xedro98/quantization-as-a-transfer-constraint
quantization-as-a-transfer-constraint is a machine learning model from xedro98. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-4.0.
Author: Shubhankar Kahali - Trumbo Labs, Inc - [email protected] License: CC BY 4.0 Paper: paper/quanttransferarxiv.pdf
Downloads · 30 days
0
Access
Public
Updated Aug 25, 2026
Repo size
967 KB
Likes
1
Public
Click a slice to open those files.
.pdf628 KB · 33%
From the Hugging Face model README
Author: Shubhankar Kahali - Trumbo Labs, Inc - [email protected]
License: CC BY 4.0
Paper: paper/quant_transfer_arxiv.pdf
Maximal update parametrization (muP) licenses zero-shot hyperparameter transfer in exact arithmetic, but low-precision training perturbs precisely the coordinate magnitudes muP is designed to keep bounded. On FashionMNIST multilayer perceptrons (factor of sixteen in width) under simulated symmetric quantization (weights, activations, or both; 8 and 4 bits; per-tensor, per-output-unit, and block-scaled MX-inspired forms), we find four results: (1) how one quantizes matters at least as much as how much - weights-only 4-bit is within about a hundredth of fp32, per-unit and block-scaled 4-bit within roughly 0.05 to 0.08 nat, versus a 0.3 to 0.4 nat hit for per-tensor 4-bit; (2) per-tensor activation quantization collapses muP's stable learning-rate edge by an order of magnitude (from 2e-1 to between 1e-2 and 5e-2); (3) the width dependence of muP's optimal rate remains negligible under that stress - the refined grid places it at 8e-3, 8e-3, and 1.3e-2 across widths 64, 128, and 512 - and a direct train/validation/test zero-shot transfer experiment shows the transferred rate incurs no measurable loss relative to the per-width oracle under 4-bit, while standard parametrization drifts downward with width and diverges catastrophically; and (4) muP's best loss is flat to improving with width at every precision, whereas SP's drifts. The result is qualitatively reproduced on CIFAR-10. All code and data are released here for reproducibility.
| Path | Description |
|---|---|
paper/quant_transfer_arxiv.pdf | the preprint (final) |
manuscript/draft_quant.md | full manuscript source (markdown) |
manuscript/quant_meta.json | metadata (title / author / affiliation / date) |
figures/ | all 4 figures (Figure 1-4) |
code/ | all training + analysis scripts |
data/ | all per-cell result datasets (JSON) |
MANIFEST.json | file list, sizes, SHA-256 hashes |
LICENSE | CC BY 4.0 notice |
The environment is Python 3.11, PyTorch (CPU), torchvision, and the official mup package (set_base_shapes / MuAdam / MuReadout). Each run is an independent initialization at a fixed seed, 1000 Adam steps, batch 128 (Adam beta1 0.9, beta2 0.999, epsilon 1e-8). Every JSON row records parametrization, quant config, width, learning rate, seed, final test loss, final test accuracy, and divergence flag.
Sweep drivers (in code/):
run_quant_seeds.py - base 8-point rate sweeps, widths 64/128/512, 3 seeds (6 for the highest-variance cells)run_quant_perchannel.py - per-output-unit 4-bit (wa4c)run_quant_refined.py - refined factor-1.6 local grids across all configs and three widthsrun_quant_extra.py - block-scaled b4, width-1024 transfer-claim runs, six-seed per-tensor cellsrun_quant_transfer.py - the direct zero-shot transfer experiment (train/validation/test, frozen LR) -> data/grid_transfer.jsonrun_quant_cifar.py - CIFAR-10 replication -> data/grid_cifar.jsonrun_qdist.py - mechanism diagnostic (update-to-grid ratio, rounding error) -> data/qdist.jsonPreprint rendering: build_report.py (arXiv-style) converts manuscript/draft_quant.md + paper/quant_meta.json into the PDF.
If you use this work, cite the Zenodo/OSF record DOI accompanying this release, or:
Shubhankar Kahali. Quantization as a Transfer Constraint: Zero-Shot Learning-Rate Transfer Survives Low Precision, but muP's Stability Margin Collapses. Trumbo Labs, Inc, 2026. DOI: <assigned by Zenodo/OSF on publish>
Shubhankar Kahali - [email protected] - Trumbo Labs, Inc