Downloads · 30 days
19
3% of all-time downloads
sd17js2/arcLM-0.8B
arcLM-0.8B is a text generation model from sd17js2. Use it when you need the model to write or continue text. It is set up for dllm. The card lists the license as apache-2.0.
A 0.8B parameter diffusion language model built by converting Qwen3.5-0.8B from autoregressive to discrete diffusion using the dLLM framework with BD3LM (Block Discrete Denoising Diffusion Language Model).
Downloads · 30 days
19
3% of all-time downloads
All-time downloads
748
Public
Parameters
752M
1.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.5 GB · 99%
From the Hugging Face model README
A 0.8B parameter diffusion language model built by converting Qwen3.5-0.8B from autoregressive to discrete diffusion using the dLLM framework with BD3LM (Block Discrete Denoising Diffusion Language Model).
Instead of generating tokens one at a time left-to-right, arcLM generates blocks of 64 tokens simultaneously through iterative denoising — trading some quality for parallel generation.
arcLM-0.8B is an experimental diffusion language model created by applying the A2D (Autoregressive-to-Diffusion) conversion pipeline to Qwen3.5-0.8B. The conversion replaces causal attention with bidirectional attention in the standard attention layers (6/24 layers), while keeping the Gated Delta Network (GDN) layers causal due to mathematical constraints of the delta-rule recurrence.
The model was jointly fine-tuned on reasoning, tool-calling, and general instruction-following data using the BD3LM training objective.
Experimental text generation via block diffusion sampling. Supports multi-turn chat with <think> reasoning blocks.
import dllm, transformers
model = dllm.utils.get_model(
model_args=type('A', (), {'model_name_or_path': 'sd17js2/arcLM-0.8B'})()
).eval().cuda()
tokenizer = dllm.utils.get_tokenizer(
model_args=type('A', (), {'model_name_or_path': 'sd17js2/arcLM-0.8B', 'tokenizer_name_or_path': None})()
)
sampler = dllm.core.samplers.BD3LMSampler(model=model, tokenizer=tokenizer)
config = dllm.core.samplers.BD3LMSamplerConfig(
steps=256, max_new_tokens=256, block_size=64, temperature=0.6
)
Or interactively:
python -u examples/a2d/bd3lm/chat.py \
--model_name_or_path sd17js2/arcLM-0.8B \
--block_size 64 --max_new_tokens 256 --steps 256 --temperature 0.6
This is an early-stage research model. It is not suitable for production use. Outputs are frequently incoherent or repetitive. Do not use for factual queries, safety-critical applications, or any deployment where reliability matters.
Jointly trained on a mix of reasoning, tool-calling, and general instruction data:
| Dataset | Split | Type |
|---|---|---|
| tatsu-lab/alpaca | train[:8000] | General instruction |
| HuggingFaceH4/ultrachat_200k | train[:5000] | Multi-turn chat |
| NousResearch/hermes-function-calling-v1 | train[:5000] | Tool calling |
| Jofthomas/hermes-function-calling-thinking-V1 | full | Tool calling + reasoning |
| open-thoughts/OpenThoughts-114k | train[:5000] | Long reasoning |
| simplescaling/s1K | full | Long reasoning |
Architecture: Qwen3.5-0.8B hybrid (GDN + standard attention), 24 layers in 6 cycles of (3 GDN + 1 standard attention).
Training objective: BD3LM — block discrete denoising diffusion. Generates 64-token blocks with iterative denoising. Loss = cross-entropy weighted by 1/t from a linear alpha schedule.
Known architectural limitation: The GDN delta-rule computes (I - L)^{-1} via forward substitution on a lower-triangular matrix. Making this bidirectional (full matrix) causes numerical instability: decay mask overflow, O((1+||A||)^63) gradient explosion, and inter-chunk recurrence amplification. Solving this for bidirectional GDN is an open research problem (potential approaches: Neumann series truncation, torch.linalg.solve, or bidirectional linear attention).
1x NVIDIA RTX 4090 (24GB) on RunPod