Downloads · 30 days
60
12% of all-time downloads
jbarrow/joecr-chandra-2-eagle3.1
joecr-chandra-2-eagle3.1 is a machine learning model from jbarrow. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for vllm. The card lists the license as apache-2.0.
An EAGLE‑3.1 speculative‑decoding draft head for datalab-to/chandra-ocr-2 (a 5B Qwen3.5‑based vision‑language OCR model). Drop it into vLLM as the speculator to accelerate single‑stream (latency‑bound) OCR decoding lo…
Downloads · 30 days
60
12% of all-time downloads
All-time downloads
496
Public
Parameters
852M
3.4 GB on disk
Likes
8
Public
Click a slice to open those files.
.safetensors1.7 GB · 100%
How the weights are stored.
BF16852M · 100%
From the Hugging Face model README
An EAGLE‑3.1 speculative‑decoding draft head for datalab-to/chandra-ocr-2 (a 5B Qwen3.5‑based vision‑language OCR model). Drop it into vLLM as the speculator to accelerate single‑stream (latency‑bound) OCR decoding losslessly — every drafted token is verified by the target, so output quality is preserved.
Speedup (single‑stream, greedy, vLLM + CUDA graphs, RTX 3090 Ti):
| tok/s | speedup | |
|---|---|---|
| baseline (no spec) | 58.0 | 1.0× |
| + this head (num_spec=5) | 104.0 | 1.8× |
Measured on olmOCR‑bench pages.
We also measured an average acceptance length of ~2.5 tokens.
Batching note: this speculative decoding head wins at low concurrency. At large batch sizes disable speculative decoding or drop num_speculative_tokens to 1–2 for bulk throughput.
EAGLE‑3.1 (single decoder layer) over Chandra‑2's Qwen3.5 text backbone:
hidden_size 2560, head_dim 256, GQA 16/4fc_norm (per‑aux‑layer RMSNorm) + norm_output (post‑norm recurrence)draft_vocab_size 32768 (from 248320; 99.99% token coverage) → ~7.6× smaller lm_head. d2t buffer maps pruned ids back to the full vocab.[3, 15, 27] (Chandra‑2 is a hybrid 3:1 linear/full‑attention model).from vllm import LLM
llm = LLM(
model="datalab-to/chandra-ocr-2",
speculative_config={
"model": "jbarrow/joecr-chandra-2-eagle3.1",
"method": "eagle3",
"num_speculative_tokens": 5,
},
limit_mm_per_prompt={"image": 1},
)
Use the Chandra ocr_layout prompt in the user turn (image + instruction).