Downloads · 30 days
144
59% of all-time downloads
Parakon/Parakon-30B
Parakon-30B is a text generation model from Parakon. Use it when you need the model to write or continue text. The card lists the license as other.
Qwen3-30B-A3B-Instruct-2507 compressed by Parakon's pipeline: 61 GB (fp16) → 5.94 GB, a 10.3× reduction, with 88% median quality retention on a six-benchmark suite.
Downloads · 30 days
144
59% of all-time downloads
All-time downloads
243
Public
Repo size
5.9 GB
Likes
1
Public
Click a slice to open those files.
.gguf5.9 GB · 100%
From the Hugging Face model README
Qwen3-30B-A3B-Instruct-2507 compressed by Parakon's pipeline: 61 GB (fp16) → 5.94 GB, a 10.3× reduction, with 88% median quality retention on a six-benchmark suite.
Each figure is the compressed model's score as a share of the same checkpoint at full precision, run through the same harness (greedy decoding, one configuration). Retention holds on knowledge, tool use and reasoning; it thins on code and strict instruction formatting — reported at the same size as the rest.
| Benchmark | Retention |
|---|---|
| MuSR (reasoning) | 93% |
| GSM8K (math) | 90% |
| MMLU-Redux (knowledge) | 88% |
| BFCL-v3 (tool calling) | 88% |
| HumanEval+ (code) | 77% |
| IFEval, prompt-strict (instruction following) | 76% |
| Median | 88% |
| Hardware | Throughput |
|---|---|
| 1× NVIDIA A10G, 24 GB (CUDA, fused decode) | 162.9 tokens/s |
| M3 MacBook Air, 16 GB (Metal) | 12.7 tokens/s |
| CPU only, 48 vCPU | 7.7 tokens/s |
Peak VRAM on the A10G run: 6.5 GB. The weights execute directly from the 1-bit blocks; no dequantized copy is ever materialized.
The published runtime release is CUDA. The Metal and CPU numbers above were measured with internal builds that are not published yet.
Chat template: Qwen3 (ChatML). Native context length: up to 262,144 tokens.
General-purpose sampling:
| Parameter | Value |
|---|---|
| temperature | 0.7 |
| top_p | 0.8 |
| top_k | 20 |
| min_p | 0 |
This artifact uses Parakon's own storage format and requires the Parakon runtime to execute. The CUDA runtime is public: Parakon/parakon-runtime — a fork of llama.cpp with Parakon's kernels, with build and run instructions in its README.
Released under the Parakon Community License:
The runtime is licensed separately. For deployment licensing, runtime access, or compression engagements on your own models: get in touch.