Downloads · 30 days
18
19% of all-time downloads
ZhenbangDu/R2-dLLM-LLaDA
R2-dLLM-LLaDA is a text generation model from ZhenbangDu. Use it when you need the model to write or continue text. It is set up for transformers.
$R^2$-dLLM is a unified framework for reducing decoding redundancy in Diffusion Large Language Models (dLLMs) from both inference and training perspectives. This specific checkpoint is a redundancy-aware supervised fi…
Downloads · 30 days
18
19% of all-time downloads
All-time downloads
97
Public
Parameters
8B
16 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16 GB · 100%
From the Hugging Face model README
$R^2$-dLLM is a unified framework for reducing decoding redundancy in Diffusion Large Language Models (dLLMs) from both inference and training perspectives. This specific checkpoint is a redundancy-aware supervised fine-tuned version of LLaDA-Instruct-8B.
Diffusion Large Language Models (dLLMs) enable parallel token prediction but often suffer from high inference latency due to decoding redundancy. $R^2$-dLLM addresses this by:
Experiments demonstrate that $R^2$-dLLM consistently reduces the number of decoding steps by up to 88% compared to existing decoding strategies, while maintaining competitive generation quality across different models and tasks.
@article{du2026r,
title={$R^{2}$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction},
author={Du, Zhenbang and Xia, Kejing and Zhong, Xinrui and Fu, Yonggan and Oswald, Nicolai and Ji, Binfei and Khailany, Brucek and Molchanov, Pavlo and Lin, Yingyan},
journal={arXiv preprint arXiv:2604.18995},
year={2026}
}