Downloads · 30 days
255K
65% of all-time downloads
bloomer010/Ling-3.0-tiny-GGUF
Ling-3.0-tiny-GGUF is a text generation model from bloomer010. Use it when you need the model to write or continue text. It is set up for llama.cpp. The card lists the license as mit.
GGUF conversions of inclusionAI/Ling-3.0-tiny, converted directly from the released BF16 safetensors.
Downloads · 30 days
255K
65% of all-time downloads
All-time downloads
392K
Public
Repo size
288 GB
Likes
111
Trending 2
Click a slice to open those files.
.gguf142 GB · 100%
From the Hugging Face model README
GGUF conversions of inclusionAI/Ling-3.0-tiny, converted directly from the released BF16 safetensors.
Consistent agentic use (tool calling, reasoning split) currently requires two unmerged llama.cpp PRs:
Without both, tool calls inside an unclosed think block are dropped and some turns fail with a 500. Will update this note as they merge.
The model does occasionally terminate its response, mid-think, without any sort of closing.
This is inherent in the weights, even at full precision.
To run with llama-server:
llama-server -hf bloomer010/Ling-3.0-tiny-GGUF:Q4_K_M
For tiny models, precision is especially crucial.
Generally... <BR> Larger files = More precision.<BR> Smaller files = More compression = More slop and misbehavin'.
Use UD-Q8_K_XL for near-full precision performance.
| Quant | Size | your memory |
|---|---|---|
| BF16 | 15.8 GB | 16 GB+ |
| UD-Q8_K_XL | 11.19 GB | 12 GB+ |
| Q8_0 | 8.41 GB | 10 GB+ |
| UD-Q6_K_XL | 7.27 GB | 8 GB+ |
| Q6_K | 6.50 GB | 8 GB+ |
| Q5_K_M | 5.64 GB | 7 GB+ |
| Q5_K_S | 5.48 GB | 6 GB+ |
| Q5_0 | 5.48 GB | 6 GB+ |
| Q4_K_M | 4.82 GB | 6 GB+ |
| Q4_K_S | 4.55 GB | 6 GB+ |
| Q4_0 | 4.53 GB | 6 GB+ |
| MXFP4_MOE | 4.72 GB | 6 GB+ ¹ |
| IQ4_XS | 4.29 GB | 5 GB+ |
| Q3_K_M | 3.84 GB | 5 GB+ |
| Q3_K_S | 3.51 GB | 5 GB+ |
| IQ3_S | 3.51 GB | 4 GB+ |
| IQ3_XXS | 3.13 GB | 4 GB+ |
| Q2_K | 2.99 GB | 4 GB+ |
| IQ2_M | 2.70 GB | 3 GB+ |
| IQ2_S | 2.48 GB | 3 GB+ |
| IQ2_XS | 2.43 GB | 3 GB+ |
| IQ2_XXS | 2.21 GB | 3 GB+ |
| IQ1_M | 1.93 GB | 3 GB+ |
| IQ1_S | 1.76 GB | 2 GB+ |
| Q1_0 | 1.30 GB | 2 GB+ |
¹ MXFP4_MOE runs its native path on MXFP4-capable GPUs (Blackwell RTX 50-series, GB10/DGX
Spark). Elsewhere it falls back to a slower dequant path — prefer a K-quant on older hardware.
The IQ-quant rungs (IQ1_S through IQ4_XS) were generated with a model-specific importance
matrix:
UD-Q8_K_XL uses Q8_0 for the main expert gate and up tensors. Token embeddings, expert down
projections, attention and Q-LoRA projections, and KDA projections remain BF16.
UD-Q6_K_XL uses Q6_K for the main expert gate and up tensors. Token embeddings, output weights,
expert down projections, attention and Q-LoRA projections, and KDA projections use Q8_0. It was
generated with the importance matrix described above.
num_nextn_predict_layers: 0)git clone https://github.com/ggml-org/llama.cpp.git # bailingmoe3 merged 2026-08-17
# pre-merge builds:
# git clone --branch bailingmoe3-support https://github.com/aetherbird/llama.cpp.git
cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j --target llama-cli llama-server
./build/bin/llama-server \
-m Ling-3.0-tiny-Q4_K_M.gguf \
-c 131072 \
-ngl auto \
--flash-attn auto \
--temp 1.0 --top-p 0.95 --top-k 20 \
--jinja
Thinking is enabled by default; disable per request with
"chat_template_kwargs": {"enable_thinking": false}. Recommended sampling parameters from the
source model card are temperature=1.0, top_p=0.95, and top_k=20.