Downloads · 30 days
775
100% of all-time downloads
PollardWeights/FrogMini-14B-Pollard
FrogMini-14B-Pollard is a text generation model from PollardWeights. Use it when you need the model to write or continue text. It is set up for trellis. The card lists the license as mit.
Pollard shrank this model: 29.54 GB (f16) → 4.56 GB — 85% smaller, 6.5× down. The smallest rung here; larger, higher-fidelity rungs are listed below. | format | this model's size | |---|---:| | f16 | 29.54 GB | | Q80…
Downloads · 30 days
775
100% of all-time downloads
All-time downloads
775
Public
Repo size
37.8 GB
Likes
0
Public
Click a slice to open those files.
.gguf37.8 GB · 100%
From the Hugging Face model README
Pollard shrank this model: 29.54 GB (f16) → 4.56 GB — 85% smaller, 6.5× down.
The smallest rung here; larger, higher-fidelity rungs are listed below.
format this model's size f16 29.54 GB Q8_0 ~15.66 GB Q6_K ~12.11 GB Q4_K_M ~8.57 GB PollardMix (this repo's IQ2_XXS) 4.56 GB
Pollard builds of microsoft/FrogMini-14B-2510 made with Pollard Weights — a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush).
Standard GGUF — runs in stock llama.cpp / ik_llama.cpp, Ollama, LM Studio. Trellis (IQ*_KT) files need ik_llama.cpp; the K-quants run anywhere.
| Parameter count | ~14.8B |
| Architecture | qwen3 |
| Input support | text |
| imatrix | yes — see calibration |
| Perplexity measured | yes — table below |
Every rung is the same weights, sized to a different RAM budget by the measured allocation. Pick the largest one that fits your machine with room for context:
Q6_K (12.12 GB). near-losslessIQ4_XS (8.43 GB). recommended defaultIQ3_S (6.79 GB). best size/quality tradeIQ2_S (5.94 GB). smallIQ2_XXS (4.56 GB). smallest - 6.5x down from f16f16 reference PPL 9.3589.
| file | PPL | size | tok/s | Mean KLD | notes |
|---|---|---|---|---|---|
FrogMini-14B-Pollard-IQ2_XXS.gguf | 13.2652 | 4.56 GB | 126.7 | — | smallest - 6.5x down from f16 |
FrogMini-14B-Pollard-IQ2_S.gguf | 10.1684 | 5.94 GB | 108.9 | — | small |
FrogMini-14B-Pollard-IQ3_S.gguf | 9.6067 | 6.79 GB | 99.2 | — | best size/quality trade |
FrogMini-14B-Pollard-IQ4_XS.gguf | 9.5405 | 8.43 GB | 88.3 | — | recommended default |
FrogMini-14B-Pollard-Q6_K.gguf | 9.4040 | 12.12 GB | 63.0 | — | near-lossless |
tok/s measured on an RTX 5070 Ti (16 GB), full GPU offload.
Every rung cleared the coherence gate on the first sampling config -- three prompts including code,
no loops, down to and including IQ2_XXS. Ship these defaults:
--temp 0.7 --repeat-penalty 1.15 --repeat-last-n 256 --top-k 40 --top-p 0.9
IQ2_XXS is coherent, and it is also a real step down: +3.91 PPL against f16, where every rung above
it costs under a point. It exists so a 14B fits in 4.56 GB. If you have the room, IQ3_S is 2 GB
larger and gives most of the quality back.
ChatML
pip install -U "huggingface_hub[cli]"
hf download PollardWeights/FrogMini-14B-Pollard \
--include "FrogMini-14B-Pollard-IQ4_XS.gguf" --local-dir ./
These are standard GGUF and run with llama.cpp:
llama-server -hf PollardWeights/FrogMini-14B-Pollard:IQ4_XS
or from a local file:
llama-cli -m FrogMini-14B-Pollard-IQ4_XS.gguf -ngl 99 -p "Explain why the sky is blue."
llama-server -m FrogMini-14B-Pollard-IQ4_XS.gguf -ngl 99 # OpenAI-compatible API + web UI at :8080
They also work in anything built on llama.cpp — LM Studio, koboldcpp, Jan, ramalama, Ollama (ollama run hf.co/PollardWeights/FrogMini-14B-Pollard).
The importance matrix (FrogMini-14B-Pollard.imatrix, included) was computed on a Calib 3.0 multi-domain corpus (prose, code, math, multilingual), 40 chunks.
llama.cpp repacks weights into an interleaved layout at load time for faster inference on ARM and AVX machines — no special file needed, online repacking covers these quants. The old Q4_0_4_4/4_8/8_8 variants are not required.
IQ*_KT) quants need ik_llama.cpp to build/run; K-quants run in any recent llama.cpp.microsoft/FrogMini-14B-2510 (Microsoft)mit, inherited from the base model.Built with Pollard Weights — frontier models, small hardware, no compromise.