Downloads · 30 days
1K
12% of all-time downloads
radicalnumerics/RND1-Base-0910
RND1-Base-0910 is a text generation model from radicalnumerics. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
<center <div style="text-align: center;" <img src="https://raw.githubusercontent.com/RadicalNumerics/assets/refs/heads/main/svg/rn-logo-desktop-vector-animated.svg" alt="Radical Numerics" style="width: 100%; max-width…
Downloads · 30 days
1K
12% of all-time downloads
All-time downloads
8.3K
Public
Parameters
30.5B
244 GB on disk
Likes
63
Public
Click a slice to open those files.
.safetensors61.1 GB · 100%
From the Hugging Face model README
RND1 is an experimental diffusion language model with 30B parameters and 3B active parameters per token (sparse Mixture-of-Experts). This model was converted from a pretrained autoregressive base to enable diffusion-based text generation.
RND1-Base-0910 has the following features:
For more details, see:
Note: RND1-Base-0910 has not been post-trained. Expect occasional repetition with greedy samplers.
pip install torch transformers accelerate numpy rich
For faster inference with optimized MoE kernels:
pip install flashinfer-python
pip install sglang[all]
pip install vllm
[!WARNING] Selecting a non-Huggingface MoE backend is highly encouraged for faster generation. Note however that non-HF backends currently support a single GPU only, so you need to set e.g.
export CUDA_VISIBLE_DEVICES=0before running the script. If you useflashinfer-python, JIT compilation the first time the code is run may take a while unlessflashinfer-jit-cacheis installed.
from transformers import AutoTokenizer, AutoModelForMaskedLM
# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained("radicalnumerics/RND1-Base-0910", trust_remote_code=True)
# Load model
model = AutoModelForMaskedLM.from_pretrained(
"radicalnumerics/RND1-Base-0910",
dtype="bfloat16",
device_map="auto",
trust_remote_code=True,
moe_backend="vllm", # hf, sglang, vllm, flashinfer
)
# Generate - Task mode (for instructions and questions)
prompt = "Write a Python function that finds the longest common subsequence of two strings. Include comments explaining the algorithm."
inputs = tokenizer(f"Question: {prompt}\nAnswer:", return_tensors="pt")
input_ids = inputs.input_ids.to(model.device)
# Generate
output = model.generate(
inputs=input_ids,
max_new_tokens=256,
num_diffusion_steps=256,
temperature=0.01,
)
# Decode only the generated part
text = tokenizer.decode(output[0], skip_special_tokens=True)
print(text)
Key parameters for text generation:
max_new_tokens: Number of tokens to generate (default: 256)num_diffusion_steps: Diffusion denoising steps (default: 256)temperature: Sampling temperature, 0.0 for greedy (default: 0.0)top_k: Top-k filtering for samplingtop_p: Nucleus filtering for samplingTask Mode (default): For instructions, questions, or requests. Add "Question:" prefix to your prompt.
Completion Mode: For text continuation. Use prompt directly without prefix.
# Completion mode example
prompt = "The key to understanding quantum computing lies in"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(
inputs=inputs.input_ids,
max_new_tokens=256,
num_diffusion_steps=256,
temperature=0.01,
)
Following the Github repo's demo script demo_rnd_generation.py:
# Task mode (default) - for instructions, questions, or requests
python demo_rnd_generation.py --prompt "Write a Python function that finds the longest common subsequence of two strings. Include comments explaining the algorithm." --moe_backend hf
# Completion mode - for text continuation
python demo_rnd_generation.py --mode completion --prompt "The key to understanding quantum computing lies in" --moe_backend hf
# Sampling parameters
python demo_rnd_generation.py --top_k 50 --temperature 0.7 --prompt "Explain how neural networks learn in simple terms" --moe_backend hf
RND1 uses a diffusion process for text generation, iteratively denoising random tokens over multiple steps. This approach differs from traditional autoregressive generation and enables parallel token generation within each diffusion step.
The model architecture is based on a sparse Mixture-of-Experts design, activating only a subset of parameters for each token to balance computational efficiency with model capacity.
If you use RND1 in your research, please cite:
@misc{rnd1-report,
title={Training Diffusion Language Models at Scale using Autoregressive Models},
author={Radical Numerics},
year={2025},
}