Downloads Β· 30 days
11
8% of all-time downloads
kilianhaefeli/Fast_dLLM_v2_7B
Fast_dLLM_v2_7B is a machine learning model from kilianhaefeli. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Autoregressive (AR) large language models (LLMs) have achieved remarkable performance across a wide range of natural language tasks, yet their inherent sequential decoding limits inference efficiency.
Downloads Β· 30 days
11
8% of all-time downloads
All-time downloads
132
Public
Parameters
333K
15.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors15.2 GB Β· 100%
From the Hugging Face model README
Autoregressive (AR) large language models (LLMs) have achieved remarkable performance across a wide range of natural language tasks, yet their inherent sequential decoding limits inference efficiency.
We present Fast-dLLM v2 β a carefully designed block diffusion language model (dLLM) that efficiently adapts a pretrained AR model (Qwen2.5-7B-Instruct) into a diffusion-style decoder for parallel text generation.
π Fast-dLLM v2 uses only ~1B tokens for fine-tuning β a 500Γ reduction vs. full-attention diffusion LLMs (Dream: 580B tokens) β while matching or surpassing AR baselines in accuracy.

Qwen/Qwen2.5-7B-InstructYou will need transformers, torch, and our custom generation function:
pip install transformers torch numpy
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Efficient-Large-Model/Fast_dLLM_7B"
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
prompt = "Give me a short introduction to large language model."
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
# Fast-dLLM v2 parallel decoding
gen_ids = model.generate(
inputs["input_ids"],
tokenizer=tokenizer,
max_new_tokens=512,
small_block_size=8,
threshold=0.9,
)
response = tokenizer.decode(
gen_ids[0][inputs["input_ids"].shape[1]:],
skip_special_tokens=True
)
print(response)
Fast-dLLM v2 offers up to 2.54Γ higher throughput than Qwen2.5-7B-Instruct, without loss in quality.

We compare Fast-dLLM v2 against AR baselines and previous diffusion LLMs on diverse tasks:
HumanEval, MBPP (code), GSM8K, Math (reasoning), IFEval (instruction), MMLU, GPQA (knowledge QA).

If you use Fast-dLLM v2 in your research or products, please cite:
@misc{wu2025fastdllmv2efficientblockdiffusion,
title={Fast-dLLM v2: Efficient Block-Diffusion LLM},
author={Chengyue Wu and Hao Zhang and Shuchen Xue and Shizhe Diao and Yonggan Fu and Zhijian Liu and Pavlo Molchanov and Ping Luo and Song Han and Enze Xie},
year={2025},
eprint={2509.26328},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2509.26328},
}
Released under Apache 2.0, following the base Qwen2.5 license.