Downloads · 30 days
26
22% of all-time downloads
yuyuanchen0/flexmdm
flexmdm is a text generation model from yuyuanchen0. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
An insertion + unmasking discrete-diffusion model for Python code, produced by fully fine-tuning Dream-org/Dream-Coder-v0-Base-7B. Unlike a fixed-length masked diffusion model, FlexMDM can grow its sequence during gen…
Downloads · 30 days
26
22% of all-time downloads
All-time downloads
116
Public
Parameters
7.6B
16.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors15.2 GB · 90%
From the Hugging Face model README
An insertion + unmasking discrete-diffusion model for Python code, produced by
fully fine-tuning Dream-org/Dream-Coder-v0-Base-7B.
Unlike a fixed-length masked diffusion model, FlexMDM can grow its sequence
during generation (a learned insertion head) and unmask tokens in any order,
enabling genuinely any-order code generation.
FlexMDM/ subdirectory (training, data pipeline, inference, and evaluation).global_step_49500 (inference weights; optimizer/RNG state stripped).AutoModel)This checkpoint hosts a Dream backbone plus FlexMDM-specific weights in
flexmdm_extras.pt. A bare AutoModel.from_pretrained returns only the backbone.
Load the full model with the FlexMDM package from the code repo:
# 1) install the FlexMDM package: pip install -e . (in the repo's FlexMDM/ dir)
from huggingface_hub import snapshot_download
from flexmdm.utils import load_model_and_tokenizer
ckpt = snapshot_download("yuyuanchen0/flexmdm")
model, tokenizer = load_model_and_tokenizer(checkpoint_dir=ckpt, max_length=768)
load_model_and_tokenizer defaults to attn_implementation="sdpa" (works
everywhere). Pass "flash_attention_2" for speed (requires flash-attn), or
"eager" to bit-match the released evaluation traces.
For sampling (the ā = 2.9 inference-time schedule, temperature 0.1, insertion
temperature 0.6 on MBPP / 1.0 on HumanEval, 512 steps) use
flexmdm.inference.flexmdm_generate; see the code repo's docs/REPRODUCE.md.
DreamModel (7B; hidden 3584, 28 layers, GQA with 4 KV heads, vocab
152064, diffusion mask id 151666), inherited unchanged.LayerNorm → Linear → GELU → Linear, clamped to [-15, 15]) and AdaLN time
conditioning on the insertion-progress coordinate.α_t = 1−(1−t)^a (insertion), β_t = 1−(1−t)^(a·b)
(unmasking), with a = b = 1.7.Fine-tuned on a five-source Hugging Face mixture (OpenCodeInstruct, opc-sft-stage2,
KodCode-V1-SFT-4o, and rStar-Coder seed/synthetic). KodCode-V1-SFT-4o is
CC BY-NC 4.0 (non-commercial). Full sources, filters, and licenses — and how to
reconstruct the tokenized set — are in the code repo's docs/DATA.md.
| Benchmark | pass@1 | pass@2 | pass@4 | pass@8 | pass@16 |
|---|---|---|---|---|---|
| HumanEval | 50.65 | 66.60 | 78.69 | 86.86 | 92.07 |
| HumanEval+ | 46.61 | 61.89 | 73.83 | 82.07 | 87.80 |
| MBPP | 64.70 | 76.73 | 83.80 | 88.20 | 91.27 |
| MBPP+ | 54.98 | 66.11 | 73.06 | 77.38 | 80.69 |
These are the paper's Table 5 rows (extraction-robust any-of-4 grading, 30 s
test timeout). MBPP/MBPP+ decode with the count-preserving insertion
temperature 0.6 (HumanEval/HumanEval+: neutral 1.0); at the neutral 1.0 the
MBPP rows are 62.22 / 74.61 / 81.81 / 86.19 / 89.68 and MBPP+
52.86 / 64.38 / 72.04 / 77.00 / 80.16 — placement sharpening improves every
pass@k on both suites. Generation is deterministic (content-addressed seeds),
so the sample set is bit-reproducible — see the code repo's
evals/REPRODUCIBILITY.md for the exact recipe.
FlexMDM also scores substantially higher than Dream-Coder on tree-based any-order metrics (CBC/RUB/RUB+/OBW).
Apache-2.0 (derived from Dream-Coder, Apache-2.0). Research artifact — not a deployment-ready system; generated code may be incorrect or insecure, so sandbox before executing. Note the non-commercial license on part of the training data (above).