Downloads · 30 days
857
71% of all-time downloads
PollardWeights/Ling-3.0-tiny-Pollard
Ling-3.0-tiny-Pollard is a text generation model from PollardWeights. Use it when you need the model to write or continue text. It is set up for trellis. The card lists the license as mit.
Pollard shrank this model: 15.78 GB (f16) → 3.83 GB — 76% smaller, 4.1× down. The smallest rung here; larger, higher-fidelity rungs are listed below. | format | this model's size | |---|---:| | f16 | 15.78 GB | | Q80…
Downloads · 30 days
857
71% of all-time downloads
All-time downloads
1.2K
Public
Repo size
14.8 GB
Likes
1
Public
Click a slice to open those files.
.gguf14.7 GB · 100%
From the Hugging Face model README
Pollard shrank this model: 15.78 GB (f16) → 3.83 GB — 76% smaller, 4.1× down.
The smallest rung here; larger, higher-fidelity rungs are listed below.
format this model's size f16 15.78 GB Q8_0 ~8.36 GB Q6_K ~6.47 GB Q4_K_M ~4.58 GB PollardMix (this repo's IQ3_S) 3.83 GB
Pollard builds of inclusionAI/Ling-3.0-tiny made with Pollard Weights — a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush).
Standard GGUF — runs in stock llama.cpp / ik_llama.cpp, Ollama, LM Studio. Trellis (IQ*_KT) files need ik_llama.cpp; the K-quants run anywhere.
| Parameter count | ~7.9B |
| Architecture | bailing_hybrid |
| Input support | text |
| imatrix | yes — see calibration |
| Perplexity measured | yes — table below |
Every rung is the same weights, sized to a different RAM budget by the measured allocation. Pick the largest one that fits your machine with room for context:
Q6_K (6.26 GB). Fits an ~11 GB box. Near-lossless — maximum quality.IQ4_XS (4.64 GB). Fits an ~9 GB box. Higher fidelity — sensitive layers pushed to iq4_xs.IQ3_S (3.83 GB). Fits an ~8 GB box. Beats same-size uniform IQ3 (table above). Recommended.| file | PPL | size | tok/s | Mean KLD | notes |
|---|---|---|---|---|---|
Ling-3.0-tiny-Pollard-IQ3_S.gguf | — | 3.83 GB | 75.1 | — | Fits an ~8 GB box. Beats same-size uniform IQ3 (table above). Recommended. |
Ling-3.0-tiny-Pollard-IQ4_XS.gguf | — | 4.64 GB | 75.2 | — | Fits an ~9 GB box. Higher fidelity — sensitive layers pushed to iq4_xs. |
Ling-3.0-tiny-Pollard-Q6_K.gguf | — | 6.26 GB | 66.9 | — | Fits an ~11 GB box. Near-lossless — maximum quality. |
tok/s is hardware-specific; the machine it was measured on is stated in the errata.
Held-out KL-divergence vs a Q6_K reference (lower = closer to the full model), measured on the same held-out set for every build:
| build | size | mean KL | vs uniform |
|---|---|---|---|
| Ling-3.0-tiny Pollard | 3.83 GB | 0.1875 | baseline |
| uniform IQ3 (interpolated to 3.83 GB) | 3.83 GB | ≈ 0.204 | ≈ 8% higher KL |
| uniform IQ3_S | 3.51 GB | 0.2821 | reference points |
| uniform IQ3_M | 3.56 GB | 0.2469 | (bracket the curve) |
| uniform IQ4_XS | 4.29 GB | 0.1312 | (bracket the curve) |
At matched size the measured allocation sits below the uniform size↔KL curve.
The measured mix: sensitive early layers get iq4_xs, most get iq3_s, the
least-sensitive get iq2_s; every attention block stays q6_K/q5_K;
embeddings/output stay q6_K; imatrix-uncovered MoE tensors are pinned so the
aggressive base can't crash. (ffn sensitivity spread ~6×, attn spread ~16× across
the 24 layers — that variance is exactly what a uniform quant wastes. The full
per-tensor map is in Ling-3.0-tiny-Pollard.tensor-types.txt.)
Token-embedding and output tensors stay at q6_K, and every attention block is
kept at q6_K/q5_K rather than dropped to the IQ base — measured sensitivity says
those tensors don't tolerate crushing, so the bits are spent there and clawed back
from the least-sensitive FFN experts.
pip install -U "huggingface_hub[cli]"
hf download PollardWeights/Ling-3.0-tiny-Pollard \
--include "Ling-3.0-tiny-Pollard-IQ3_S.gguf" --local-dir ./
These are standard GGUF and run with llama.cpp:
llama-server -hf PollardWeights/Ling-3.0-tiny-Pollard:IQ3_S
or from a local file:
llama-cli -m Ling-3.0-tiny-Pollard-IQ3_S.gguf -ngl 99 -p "Explain why the sky is blue."
llama-server -m Ling-3.0-tiny-Pollard-IQ3_S.gguf -ngl 99 # OpenAI-compatible API + web UI at :8080
They also work in anything built on llama.cpp — LM Studio, koboldcpp, Jan, ramalama, Ollama (ollama run hf.co/PollardWeights/Ling-3.0-tiny-Pollard).
The importance matrix (Ling-3.0-tiny-Pollard.imatrix, included) was computed on a mixed-domain corpus so the matrix sees every register the model serves.
llama.cpp repacks weights into an interleaved layout at load time for faster inference on ARM and AVX machines — no special file needed, online repacking covers these quants. The old Q4_0_4_4/4_8/8_8 variants are not required.
IQ*_KT) quants need ik_llama.cpp to build/run; K-quants run in any recent llama.cpp.inclusionAI/Ling-3.0-tinymit, inherited from the base model.Built with Pollard Weights — frontier models, small hardware, no compromise.