Downloads · 30 days
839
100% of all-time downloads
fwerkor/MiniCPM5-2B-Diffusion-Base
MiniCPM5-2B-Diffusion-Base is a text generation model from fwerkor. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
MiniCPM5-2B-Diffusion-Base is a bidirectional masked-diffusion language-model backbone converted from openbmb/MiniCPM5-2B-Base. It is intended as a diffusion-native base model for CID and related research, rather than…
Downloads · 30 days
839
100% of all-time downloads
All-time downloads
839
Public
Parameters
2.5B
5 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors5 GB · 100%
From the Hugging Face model README
MiniCPM5-2B-Diffusion-Base is a bidirectional masked-diffusion language-model backbone converted from openbmb/MiniCPM5-2B-Base. It is intended as a diffusion-native base model for CID and related research, rather than as an instruction-tuned chat model.
The checkpoint keeps the original Llama-compatible parameter layout, while the repository supplies a custom Transformers class that changes the forward pass to full-sequence bidirectional denoising. Loading it as a plain LlamaForCausalLM would be semantically incorrect.
Use Transformers with remote model code enabled:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "fwerkor/MiniCPM5-2B-Diffusion-Base"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
repo,
trust_remote_code=True,
dtype=torch.bfloat16,
).to("cuda").eval()
The loaded class is CIDDiffusionForMaskedLM. Its forward pass is bidirectional, uses the checkpoint's dedicated mask token, and computes same-position masked denoising logits.
The custom class overrides model.generate so it does not fall back to autoregressive GenerationMixin semantics.
prompt = tokenizer(
"The capital of France is",
add_special_tokens=False,
return_tensors="pt",
).input_ids.to("cuda")
output = model.generate(
prompt,
max_new_tokens=32,
diffusion_steps=32,
block_length=8,
)
print(tokenizer.decode(output[0], skip_special_tokens=True))
For direct masked-token denoising, use model.denoise(input_ids, steps=...).
The released checkpoint is the completed Stage 0 conversion run:
The streaming corpus mixture was:
| Source | Weight |
|---|---|
| openbmb/UltraX-Preview (UltraX-Ultra-FineWeb) | 70% |
| openbmb/Ultra-FineWeb (Chinese) | 15% |
| openbmb/UltraData-Code | 10% |
| openbmb/UltraData-Math | 5% |
Exact source revisions are recorded in diffusion_config.json.
The conversion substantially improves masked-token denoising over the untouched AR base. In a held-out smoke test used during release validation, the converted checkpoint reduced masked-token cross entropy from roughly 15.9/16.3/19.4 to 4.14/5.24/8.58 at 15%/50%/85% masking respectively.
This is a base checkpoint, not an instruction model. High-mask and free-form generation remain materially harder than low- and medium-mask reconstruction, and long unconstrained generations may repeat. Downstream CID training is expected to specialize the diffusion backbone further.
This diffusion backbone is developed as part of Continuous Interaction Diffusion (CID). See Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning (arXiv:2608.10438, DOI).
If you use this model or the CID runtime, please cite:
@article{cao2026continuous,
title = {Continuous Interaction Diffusion: A Diffusion-Native Architecture for Asynchronous Tool-Augmented Reasoning},
author = {Cao, Yuhang and Mu, Yanzhou and Fang, Chunrong and Chen, Zhenyu},
journal = {arXiv preprint arXiv:2608.10438},
year = {2026},
doi = {10.48550/arXiv.2608.10438},
url = {https://arxiv.org/abs/2608.10438}
}
The dedicated mask token is <|cid_mask|>.
This derivative checkpoint follows the Apache-2.0 license of openbmb/MiniCPM5-2B-Base. See the upstream model card for its full attribution, limitations, and citation information.