Downloads · 30 days
50
32% of all-time downloads
Akicou/LLaDA2.2-Flash-REAM-50B
LLaDA2.2-Flash-REAM-50B is a text generation model from Akicou. Use it when you need the model to write or continue text. It is set up for transformers.
This repository contains a REAM-compressed version of inclusionAI/LLaDA2.2-flash, produced with Akicou/ream — a REAM/REAP-style Mixture-of-Experts compression framework.
Downloads · 30 days
50
32% of all-time downloads
All-time downloads
155
Public
Parameters
52.9B
106 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors106 GB · 100%
From the Hugging Face model README
This repository contains a REAM-compressed version of inclusionAI/LLaDA2.2-flash, produced with Akicou/ream — a REAM/REAP-style Mixture-of-Experts compression framework.
It is not an official inclusionAI release.
REAM (Router Expert Activation Merging) — originally proposed by SamsungSAILMontreal/ream (Jha et al., 2026) — is a compression technique for MoE models that combines expert merging with pseudo-pruning:
S[i] = mean(||h_i(x)|| × p_i(x)) over tokens routed to expert i.--fast-merge).This implementation follows the original REAM paper closely, including:
| Metric | Original (LLaDA2.2-flash) | REAM-50B | Change |
|---|---|---|---|
| Routed Experts/Layer | 256 | 128 | -128 (-50%) |
| MoE Layers | 31 (layer 0 is dense) | 31 | — |
| Top-k per token | 8 | 8 | — |
| Calibration | 100 samples, hardcoded prompts | — | — |
| Merge Method | Saliency-weighted avg (--fast-merge) | — | — |
| Grouping | REAM pseudo-group (group_size=16) | — | — |
The per-layer sigmoid-router expert_bias vector (analogous to DeepSeek's e_score_correction_bias) is correctly shrunk alongside the router weights on every layer.
python examples/compress_sequential.py \
--model inclusionAI/LLaDA2.2-flash \
--output ./LLaDA2.2-Flash-REAM-50B \
--target-ratio 0.5 \
--samples 100 \
--max-seq-len 512 \
--batch-size 4 \
--max-tokens 2048 \
--cpu-merge \
--fast-merge \
--seed 42
Hardware: 4× NVIDIA H100 (80GB each), 192 vCPUs, 2TB RAM
trust_remote_code=True (same as the base LLaDA2.2-flash).expert_bias is correctly shrunk (256 → 128) alongside mlp.gate.weight.generate() is a non-autoregressive, block-wise iterative refinement loop and targets transformers==5.2.0 (the base model's version); on newer transformers (≥ 5.3) the refactored GenerationMixin breaks it, so pin that version for inference. The forward / loss path works on the current stack, so the checkpoint is fine-tune ready regardless.tokenizer.save_pretrained() was found to re-serialize the fast tokenizer incorrectly, producing byte-level fallback; the original tokenizer.json / tokenizer_config.json / chat_template.jinja / custom tokenization_llada2.py are used as-is).Real generation from this checkpoint (transformers==5.2.0, temperature=0.0, block_length=32, gen_length=512, threshold=0.5, eos_early_stop=True). Note: like the base model, inputs must satisfy num_tokens % block_size(32) == 0.
User: What is the Heyting Algebra thingamajig?
Model Output:
The term "Heyting Algebra thingamajig" appears to be a fictional or made-up term, possibly derived from a combination of words such as "Heyting" and "Algebra" along with "thingamajig."
This is a calibrated-merge release: a 50%-compressed checkpoint with no fine-tuning yet. Generation is coherent but visibly lower-quality / more literal than the 256-expert base (which gives a full intuitionistic-logic explanation). Fine-tuning is expected to recover quality. Forward + cross-entropy loss verified (loss ≈ 0.19 on a sample sentence).
model.config.num_experts == 128, num_experts_per_tok == 8.mlp.gate.weight (128, 4096) and mlp.gate.expert_bias (128,) confirmed per layer; shared experts present; lm_head.weight preserved (tie_word_embeddings=False).apply_chat_template → 19 tokens for the smoke-test prompt, matching the base).import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Akicou/LLaDA2.2-Flash-REAM-50B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
# NOTE: block-diffusion models require token length to be a multiple of block_size (32).
text = "The quick brown fox jumps over the lazy dog. " * 3
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model(**inputs, labels=inputs.input_ids)
print("loss:", float(outputs.loss))
For diffusion generation, follow the base model's recommended recipe (temperature=0.0, block_length=32, threshold=0.5, editing_threshold=0.0, gen_length=512, eos_early_stop=True) with a compatible transformers version.
@article{jha2026ream,
title={REAM: Merging Improves Pruning of Experts in LLMs},
author={Jha, Saurav and Hashemzadeh, Maryam and Pasand, Ali Saheb and Parviz, Ali and Lee, Min-Joong and Knyazev, Boris},
journal={arXiv preprint arXiv:2604.04356},
year={2026}
}
@misc{llada2,
title={LLaDA2.2-flash},
author={inclusionAI},
year={2026},
url={https://huggingface.co/inclusionAI/LLaDA2.2-flash}
}
All architecture, tokenizer, and original weights are from inclusionAI/LLaDA2.2-flash. This repository contains a REAM-compressed derivative checkpoint.