Downloads · 30 days
13
46% of all-time downloads
lupodevelop/echo-mdlm-341m-fp16
echo-mdlm-341m-fp16 is a fill-mask model from lupodevelop. Use it when you need the model to fill a missing word. It is set up for echo-1.58. The card lists the license as apache-2.0.
These are research artifacts accompanying the paper Native Ternary Quantization-Aware Training for Masked Diffusion Language Models. They are 341M-parameter masked-diffusion language models trained on Italian FineWeb-…
Downloads · 30 days
13
46% of all-time downloads
All-time downloads
28
Public
Repo size
749 MB
Likes
0
Public
Click a slice to open those files.
.pt748 MB · 100%
From the Hugging Face model README
These are research artifacts accompanying the paper Native Ternary Quantization-Aware Training for Masked Diffusion Language Models. They are 341M-parameter masked-diffusion language models trained on Italian FineWeb-2 with a 32k SentencePiece tokenizer. This is not a production model. At roughly 12 tokens per parameter neither the ternary nor the full-precision model composes fluent text; free generation degenerates identically at both precisions. Use these checkpoints for reproduction, infilling analysis, and as paired baselines, not as downstream generators.
d_model 1024, 24 layers, 16
heads, d_ff 2816, tied embeddings, 32001 vocabulary (mask token id 32000).spm_it.model (included).| 341M model | masked-CE | perplexity | vs FP16 twin |
|---|---|---|---|
| FP16 twin | 4.8100 | 122.7 | ceiling |
| Ternary baseline | 4.9852 | 146.2 | +19.2% |
| + continued distillation | 4.9125 | 136.0 | +10.8% |
| + from-scratch recipe | 4.9878 | 146.6 | +19.5% |
The full-precision twin, trained on the same tokens in the same order under the same masks as the ternary models. It is the paired reference against which the ternary penalty is measured, and the teacher for the distillation recovery recipe. Perplexity 122.7. Released so the paired comparison and the distillation setup are fully reproducible.