Downloads · 30 days
4.2K
100% of all-time downloads
orcarouter/OrcaSAQ-2-27B
OrcaSAQ-2-27B is a text generation model from orcarouter. Use it when you need the model to write or continue text. It is set up for vllm. The card lists the license as apache-2.0.
<p align="center" <a href="https://www.orcarouter.ai" <img src="https://www.orcarouter.ai/orca-logo-classic.png" width="96" alt="OrcaRouter" </a </p
Downloads · 30 days
4.2K
100% of all-time downloads
All-time downloads
4.2K
Public
Parameters
6.8B
12.3 GB on disk
Likes
256
Trending 95
Click a slice to open those files.
.safetensors12.3 GB · 100%
How the weights are stored.
I165.5B · 81%
From the Hugging Face model README
27B reasoning. 12.3 GB.
OrcaSAQ2 27B compresses Qwen3.8-27B from a 54 GB BF16 checkpoint to 12.3 GB while preserving extremely high fidelity to the original model.
Built for: long-horizon agents · coding · tool use · reasoning · stateful execution
OrcaSAQ2 is a proprietary sensitivity-aware mixed-precision quantization system developed by OrcaRouter and its research team behind.
It is optimized around one goal: Preserve as much useful model behavior as possible inside a practical GPU memory envelope.
The resulting checkpoint provides:
| Metric | BF16 | OrcaSAQ2 |
|---|---|---|
| Checkpoint | 54 GB | 12.3 GB |
| Relative size | 100% | 22.8% |
| Storage reduction | — | 77.2% |
| Decoder precision | 16-bit | 3.21 bpw avg. |
| Perplexity | 5.6468 | 5.6482 |
| PPL delta | — | +0.02% |
| Top-1 agreement | 100% | 93.2% |
| Mean KLD | — | 0.031 |
| Context | 262K | 262K |
All numbers below are measured using these exact OrcaSAQ2 weights against the BF16 reference through the same evaluation path.
16,376 predicted tokens
| Build | Size | Decoder Bits | Mean KLD ↓ | Top-1 Agreement ↑ | PPL ↓ |
|---|---|---|---|---|---|
| Qwen3.8-27B BF16 | 54 GB | 16 | — | 100% | 5.6468 |
| OrcaSAQ2 27B | 12.3 GB | 3.21 | 0.031 | 93.2% | 5.6482 |
BF16 5.6468 ████████████████████████████████████████
OrcaSAQ2 5.6482 ████████████████████████████████████████
Delta: +0.02%
OrcaSAQ2 vs BF16
███████████████████████████████████████████████░░░ 93.2%
Qwen3.8-27B BF16
██████████████████████████████████████████████████ 54.0 GB
OrcaSAQ2
███████████ 12.3 GB
77.2% smaller.
Short benchmarks can hide small degradation.
Agents cannot.
A small model error can change a tool call.
That changes the environment state.
The changed state affects every decision that follows.
Plan
↓
Act
↓
Observe
↓
Decide
↓
Recover
↓
Repeat
↓
...
↓
Task Success
Across long trajectories, small errors can compound into large behavioral differences.
That makes long-horizon execution an especially useful stress test for compressed reasoning models.
OrcaSAQ2 performs strongly on long-horizon workloads relative to models in its deployment and parameter class, despite operating from a 12.3 GB checkpoint.
This makes it particularly suitable for:
Perplexity asks:
How similar is the next-token distribution?
Long-horizon evaluation asks:
Can the model still finish the job after many decisions?
For agent models, both matter.
Agent benchmarks depend heavily on the surrounding scaffold, tools, reasoning budget, timeouts and execution environment. The results below are therefore shown as public reference points, not direct apples-to-apples comparisons.
| Model | Reported score |
|---|---|
| Claude Sonnet 4.6 | 79.6 |
| Claude Sonnet 4.5 | 77.2 |
| Gemini 3 | 76.2 |
| OrcaSAQ2 27B | 70.0 |
| Qwen3-Coder-480B-A35B | 69.6 |
| Gemini 2.5 Pro | 63.8 |
| GPT-4.1 | 54.6 |
70.0% SWE-bench Verified from a 12.06 GB 27B checkpoint.
| Model / Agent | Reported score |
|---|---|
| Gemini 3.1 Pro / Terminus 2 | 70.7 |
| Claude Opus 4.6 / Claude Code | 70.1 |
| Claude Opus 4.6 / Terminus 2 | 63.8 |
| Claude Sonnet 4.6 / Claude Code | 58.5 |
| OrcaSAQ2 27B | 58.4 |
| Gemini 3 Flash / Gemini CLI | 56.9 |
| GPT-5.4 / Terminus 2 | 54.8 |
| Claude Sonnet 4.6 / Terminus 2 | 51.5 |
58.4% Terminal-Bench 2.1 while fitting in ~12 GB of checkpoint storage.
Public scores use different agent stacks and should not be interpreted as a strict model-only ranking.
| Base model | Qwen/Qwen3.8-27B |
| Architecture | Qwen3_5ForCausalLM |
| Layers | 64 |
| Hidden size | 5120 |
| Hybrid attention | 48 Gated DeltaNet + 16 full-attention layers |
| Context | 262,144 tokens |
| Vocabulary | 248,320 |
| MTP head | Included |
| Thinking | Supported |
| Tool calling | Supported |
| Checkpoint | 12.3 GB |
| Decoder average | 3.21 bpw |
| Serving | vLLM |
| Vision | Not included |
| License | Apache-2.0 |
Measured under a 15.7 GiB GPU memory cap.
| Configuration | 1 Stream | 8 Streams | 16 Streams | KV Pool |
|---|---|---|---|---|
| vLLM · MTP off | 65.3 tok/s | 332 tok/s | 333 tok/s | 29,354 tok |
| vLLM · MTP on | 90.1 tok/s | 220 tok/s | 219 tok/s | 14,563 tok |
Single-stream decode
MTP off █████████████████████████████ 65.3 tok/s
MTP on ████████████████████████████████████████
90.1 tok/s
+38% single-stream decode throughput
MTP trades additional compute and KV capacity for stronger interactive decode performance.
It is particularly useful for:
For highly batched workloads, benchmark both configurations.
OrcaSAQ2's checkpoint is 12.3 GB.
That makes deployment possible on hardware that cannot hold the original 54 GB BF16 checkpoint.
16 GB GPU
┌───────────────────────────────────────────┐
│ │
│ OrcaSAQ2 weights 12.3 GB │
│ ███████████████████████████████████ │
│ │
│ Remaining ~3.7 GB │
│ ██████████ │
│ │
└───────────────────────────────────────────┘
Actual usable memory depends on:
A practical starting point for a 16 GB GPU is approximately 32K interactive context, then tune based on the workload.
The model architecture supports up to 262K context.
plan → act → observe → recover → repeat
Repository-scale generation, editing, testing and debugging.
Structured workflows where action-selection quality matters.
Preserving the capabilities of the 27B base model under an aggressive deployment constraint.
A 12.3 GB checkpoint designed around practical inference hardware.
vLLM + MTP + OpenAI-compatible APIs.
One prompt each, first attempt.
The standard SVG test, asked for as an animation.
Chain over the chainring, cranks 180° out of phase, parallax background. Pure SMIL, no JavaScript. Used as generated.
<p align="center"> <img src="https://huggingface.co/orcarouter/OrcaSAQ-2-27B/resolve/main/assets/pelican-bicycle.svg" width="760" alt="Animated SVG of a pelican riding a bicycle"> </p>Create a html low-poly 3D models of the Statue of Liberty
A single self-contained HTML file: Three.js scene, orbit controls, procedural geometry.
<p align="center"> <video src="https://huggingface.co/orcarouter/OrcaSAQ-2-27B/resolve/main/assets/liberty.mp4" width="760" controls loop muted playsinline></video> </p>pip install -U vllm huggingface_hub
pip install git+https://github.com/Continuum-AI-Corp/OrcaSAQ2-kernel
hf download orcarouter/OrcaSAQ2-27B \
--local-dir ./OrcaSAQ2-27B
vllm serve ./OrcaSAQ2-27B \
--served-model-name OrcaSAQ2-27B \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":2}'
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="not-needed",
)
response = client.chat.completions.create(
model="OrcaSAQ2-27B",
messages=[
{
"role": "user",
"content": "Analyze this repository and plan the next five actions."
}
],
)
print(response.choices[0].message.content)
temperature = 1.0
top_p = 0.95
top_k = 20
Thinking mode is enabled by default.
For agent deployments, benchmark against the actual tool schema, context distribution and reasoning budget used in production.
A low-bit reasoning model should not be judged by checkpoint size alone.
We look at the intersection of:
<p align="center"><strong>Footprint × BF16 Fidelity × Capability × Long-Horizon Stability × Serving Performance</strong></p>A useful low-bit model must remain useful after compression.
Perplexity is useful and reproducible.
It is not a complete measure of agentic capability.
Quantization can affect:
reasoning
↓
planning
↓
tool selection
↓
state tracking
↓
recovery
↓
task completion
That is why OrcaSAQ2 reports BF16 fidelity metrics alongside downstream and long-horizon evaluation.
OrcaSAQ2 uses a proprietary sensitivity-aware mixed-precision quantization system developed by OrcaRouter.
The implementation is optimized to preserve model quality under a strict deployment-memory target.
Detailed quantization methodology, calibration strategy, precision allocation and packing techniques are not currently disclosed.
Open multi-model code review.
Record, replay, fork and debug AI-agent runs.
Self-hosted multi-model AI infrastructure.
<p align="center"><strong>Open model. Open harness. Open bill.</strong></p>@misc{qwen38,
title = {Qwen3.8-Max: A New Bar for Coding and Cowork},
author = {{Qwen Team}},
year = {2026},
month = {August},
url = {https://qwen.ai/blog?id=qwen3.8}
}
Apache-2.0
Inherited from:
Quantization does not change the underlying license obligations.