Downloads · 30 days
9
28% of all-time downloads
lupodevelop/echo-mdlm-341m-ternary-fromscratch
echo-mdlm-341m-ternary-fromscratch is a fill-mask model from lupodevelop. Use it when you need the model to fill a missing word. It is set up for echo-1.58. The card lists the license as apache-2.0.
These are research artifacts accompanying the paper Native Ternary Quantization-Aware Training for Masked Diffusion Language Models. They are 341M-parameter masked-diffusion language models trained on Italian FineWeb-…
Downloads · 30 days
9
28% of all-time downloads
All-time downloads
32
Public
Repo size
750 MB
Likes
0
Public
Click a slice to open those files.
.pt749 MB · 100%
From the Hugging Face model README
These are research artifacts accompanying the paper Native Ternary Quantization-Aware Training for Masked Diffusion Language Models. They are 341M-parameter masked-diffusion language models trained on Italian FineWeb-2 with a 32k SentencePiece tokenizer. This is not a production model. At roughly 12 tokens per parameter neither the ternary nor the full-precision model composes fluent text; free generation degenerates identically at both precisions. Use these checkpoints for reproduction, infilling analysis, and as paired baselines, not as downstream generators.
d_model 1024, 24 layers, 16
heads, d_ff 2816, tied embeddings, 32001 vocabulary (mask token id 32000).spm_it.model (included).| 341M model | masked-CE | perplexity | vs FP16 twin |
|---|---|---|---|
| FP16 twin | 4.8100 | 122.7 | ceiling |
| Ternary baseline | 4.9852 | 146.2 | +19.2% |
| + continued distillation | 4.9125 | 136.0 | +10.8% |
| + from-scratch recipe | 4.9878 | 146.6 | +19.5% |
A ternary model trained from the first step with the full recovery recipe (distillation plus per-channel scale) for the full 4B-token budget. This is a released negative control. At 341M the recipe applied from step 0 recovers nothing measurable: the model lands at masked-CE 4.9878, within noise of the bare ternary baseline (4.9852), against the 70% recovery the identical recipe delivers at 27M. Post-hoc distillation scales; from-scratch distillation does not. Released because reproducible negative results are rarely shared and are useful for anyone studying scale-dependent distillation.