Downloads · 30 days
593
69% of all-time downloads
backpack-run/GLM-5.3-Flash-GGUF
GLM-5.3-Flash-GGUF is a image-text-to-text model from backpack-run. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for gguf. The card lists the license as mit.
GGUF quantizations of zai-org/GLM-5.3-Flash, packaged for llama.cpp-compatible image-and-text inference and Backpack.
Downloads · 30 days
593
69% of all-time downloads
All-time downloads
856
Public
Repo size
195 GB
Likes
4
Public
Click a slice to open those files.
.gguf195 GB · 100%
From the Hugging Face model README
GGUF quantizations of zai-org/GLM-5.3-Flash, packaged for llama.cpp-compatible image-and-text inference and Backpack.
| Property | Value |
|---|---|
| Original model | zai-org/GLM-5.3-Flash |
| Original publisher | zai-org |
| Upstream revision | 690b705278a3a58e538fcb37c2ca8b5f9511213c |
| Architecture | Glm5NextForConditionalGeneration |
| Parameters | 321,323,031,390 |
| Context length | Not declared |
| Input modalities | text, image |
| Output modalities | text |
| License | mit |
| Quantization | Size | Approx. RAM | Recommended for |
|---|---|---|---|
| Q4_K_M | 180.5 GiB | 263.78 GB | Most users |
Memory values are estimates, not guarantees. Runtime configuration and context length change actual use.
| File | Precision | Size |
|---|---|---|
GLM-5.3-Flash-mmproj-F16.gguf | F16 | 1.1 GiB |
The projector is required for image input and must be used with one of the language-model GGUF files above.
Recommended: Q4_K_M. It usually offers a practical quality, size, and speed balance for local inference.
Using the llama.cpp revision recorded below:
llama-mtmd-cli --model GLM-5.3-Flash-Q4_K_M.gguf --mmproj GLM-5.3-Flash-mmproj-F16.gguf --image image.jpg --prompt "Describe this image."
These artifacts and backpack-model.yaml are prepared for the Backpack AI workspace.
Artifact integrity and GGUF metadata validation are the publication requirements. Runtime load, inference, and tokenizer results are reported independently and do not imply a certification or endorsement.
| Package | Integrity | Load | Inference | Tokenizer |
|---|---|---|---|---|
| Q4_K_M | passed | failed | skipped | skipped |
Runtime execution validation has not completed successfully for every artifact. Treat the affected package as experimental with the pinned toolchain until downstream runtime testing is complete.
Packaged: 2026-09-06T07:07:13.528564+00:00
llama.cpp toolchain revision: 8134115f88ed8018474e7db69afcfe97fb097fc4
SHA-256 checksums: see checksums.sha256
GLM-5.3-Flash-Q4_K_M.gguf: 3e1f1720e869d98acd55a8f94b5efd78814a6ba0a2c2e4e609d637e9cca60406
GLM-5.3-Flash-mmproj-F16.gguf: f64a2e935c899224054258d2372d9ad4c19b760141044292fa6ee0fb5ff36624
The source model was resolved to immutable revision 690b705278a3a58e538fcb37c2ca8b5f9511213c. It was converted with llama.cpp's convert_hf_to_gguf.py, including its multimodal projector, and quantized with llama-quantize; the exact toolchain revision is recorded above and in backpack-model.yaml.
Upstream declares mit. Review the upstream model card and comply with all applicable terms.
Backpack does not claim ownership of the original model. These artifacts are packaged and quantized distributions of the upstream model.
Quantization can alter output quality. Memory estimates vary with runtime configuration, context length, and hardware.