Downloads · 30 days
128
21% of all-time downloads
backpack-run/SmolLM2-135M-Instruct-GGUF
SmolLM2-135M-Instruct-GGUF is a machine learning model from backpack-run. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for gguf. The card lists the license as apache-2.0.
GGUF quantizations of HuggingFaceTB/SmolLM2-135M-Instruct, tested for llama.cpp-compatible inference and packaged for Backpack.
Downloads · 30 days
128
21% of all-time downloads
All-time downloads
609
Public
Repo size
362 MB
Likes
4
Public
Click a slice to open those files.
.gguf362 MB · 100%
From the Hugging Face model README
GGUF quantizations of HuggingFaceTB/SmolLM2-135M-Instruct, tested for llama.cpp-compatible inference and packaged for Backpack.
| Property | Value |
|---|---|
| Original model | HuggingFaceTB/SmolLM2-135M-Instruct |
| Original publisher | HuggingFaceTB |
| Upstream revision | 12fd25f77366fa6b3b4b768ec3050bf629380bac |
| Architecture | LlamaForCausalLM |
| Parameters | 134,515,008 |
| Context length | 8,192 |
| License | apache-2.0 |
| Quantization | Size | Approx. RAM | Recommended for |
|---|---|---|---|
| Q4_K_M | 100.6 MiB | 1.14 GB | Most users |
| Q5_K_M | 106.9 MiB | 1.15 GB | Higher quality |
| Q8_0 | 138.1 MiB | 1.2 GB | Plenty of memory |
Memory values are estimates, not guarantees. Runtime configuration and context length change actual use.
Recommended: Q4_K_M. It usually offers a practical quality, size, and speed balance for local inference.
Using the llama.cpp revision recorded below:
llama-cli --model SmolLM2-135M-Instruct-Q4_K_M.gguf --conversation
These artifacts and backpack-model.yaml are prepared for the Backpack AI workspace.
| Package | Integrity | Load | Inference | Tokenizer |
|---|---|---|---|---|
| Q4_K_M | passed | passed | passed | passed |
| Q5_K_M | passed | passed | passed | passed |
| Q8_0 | passed | passed | passed | passed |
Packaged: 2026-08-20T20:32:10.805207+00:00
llama.cpp revision: de699957b92f490efebad149665b0dccf127eaff
SHA-256 checksums: see checksums.sha256
SmolLM2-135M-Instruct-Q4_K_M.gguf: dd18a11b8634d1684448986b8c166f75319f52082d759654aaa8fe5bd2f057e3
SmolLM2-135M-Instruct-Q5_K_M.gguf: 00680963c363ba10593daf7568dd6e1ee4c4771a608fe1b4e43f86b564d9b823
SmolLM2-135M-Instruct-Q8_0.gguf: ee785d9b4836ddb57207ae6daa630206a756c99fc52e19696f1e2ea2e8a41b99
The source model was resolved to immutable revision 12fd25f77366fa6b3b4b768ec3050bf629380bac. It was converted with llama.cpp's convert_hf_to_gguf.py and quantized with llama-quantize; the exact tested revision is recorded above and in backpack-model.yaml.
Upstream declares apache-2.0. Review the upstream model card and comply with all applicable terms.
Backpack does not claim ownership of the original model. These artifacts are packaged and quantized distributions of the upstream model.
Quantization can alter output quality. Memory estimates vary with runtime configuration, context length, and hardware.