Downloads · 30 days
0
datsboom/per-layer-e2e-cos
per-layer-e2e-cos is a machine learning model from datsboom. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Does your 4-bit MLX model compute what the FP32 weights say it should?
Downloads · 30 days
0
Access
Public
Updated Jun 16, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.py14.2 KB · 71%
From the Hugging Face model README
Does your 4-bit MLX model compute what the FP32 weights say it should?
This tool runs a single token through every transformer layer of an MLX 4-bit model, executing the EXACT same architecture in numpy FP32 as a ground-truth reference. It computes cosine similarity per layer — telling you precisely where quantization drift creeps in.
| Finding | Result |
|---|---|
| Per-component matmul fidelity | 160/160 at cos > 0.99999 — every operation correct |
| Self-determinism | 40/40 L-inf=0 — MLX 4-bit is bit-identical across runs |
| Logit self-match | 10/10 top-1 token match — output deterministic |
| bf16 vs 4-bit (same framework) | Semantically equivalent — 54% word overlap, identical answers |
| bf16 vs 4-bit (cross-framework vs numpy) | 0/10 top-1 match — measures framework, not model |
| Production quality | 25/25 probes correct on OptiQ |
| Benchmarks (5 tasks, both models) | Tied within error bars — per-layer cos ≠ benchmark accuracy |
Cross-framework per-layer E2E cosine similarity measures the GAP between MLX GPU fused-kernel fp16 and numpy CPU FP32 — not model quality. The 4-bit model is proven correct by self-determinism, logit match, bf16 comparison, and output quality probes. Always run a self-determinism baseline before interpreting cross-framework results.
python3 per_layer_e2e_cos.py --model mlx-community/Qwen3.5-35B-A3B-4bit --token-id 151644
After 12 days and 109 verified measurements, the lessons are clear:
MIT