Downloads · 30 days
0
majentik/garden
garden is a text generation model from majentik. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
[!TIP] KV-cache quantization needs no fork (2026): upstream llama.cpp / Ollama cover it natively. Use -ctk q80 -ctv q80 (~half KV memory, perplexity +0.002–0.05) or -ctk q40 -ctv q40 (~quarter memory, ≈7.6% perplexity…
Downloads · 30 days
0
Access
Public
Updated Sep 15, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md7.9 KB · 84%
From the Hugging Face model README
<!-- kv-upstream-note -->[!TIP] KV-cache quantization needs no fork (2026): upstream llama.cpp / Ollama cover it natively. Use
-ctk q8_0 -ctv q8_0(~half KV memory, perplexity +0.002–0.05) or-ctk q4_0 -ctv q4_0(~quarter memory, ≈7.6% perplexity increase). In Ollama:OLLAMA_KV_CACHE_TYPE=q8_0withOLLAMA_FLASH_ATTENTION=1. Keep K and V types symmetric to stay on the fused Flash-Attention path. Since April 2026 mainline llama.cpp also applies Hadamard rotation to KV activations (PR #21038), which greatly improves low-bit KV quality (opt-out:LLAMA_ATTN_ROT_DISABLE=1).
Quantized open-weight models for Apple Silicon and llama.cpp, released only after a coherence smoke gate passes on real hardware. Every repo keeps the upstream tokenizer, architecture and license; the only thing we change is how the weights are stored.
402 model repositories · 14 datasets · 13 collections · snapshot 2026-09-16
mlx-lm / mlx-vlm / mlx-audio for
MLX, llama.cpp for GGUF. Group size 32 for ≤4-bit MLX tiers. We never
publish a tier at or above the source's bits-per-weight.majentik/garden-quant-bench.
Repos in the
Verified MLX releases
collection link their evidence directly.| Lane | Repos | Runtime | Tiers |
|---|---|---|---|
| MLX | 264 | mlx-lm, mlx-vlm, mlx-audio on Apple Silicon | 2 / 3 / 4 / 5 / 6 / 8-bit, MXFP4, bf16 references, LoRA adapters |
| GGUF | 129 | llama.cpp (one-shot: llama-completion -no-cnv), Ollama, LM Studio | Q2_K … Q8_0, IQ4_XS, MXFP4 |
| FP8 / ONNX / other | 7 | vLLM, onnxruntime | Family-specific |
Repo naming is unbranded: majentik/<Model>-MLX-<tier> and
majentik/<Model>-GGUF-<QT>.
Counts are repos on the Hub as of the snapshot date.
| Family | Repos | Notes |
|---|---|---|
| Gemma 4 | 91 | E2B / E4B / 12B / 26B-A4B / 31B, base + instruct, MLX + GGUF (collection) |
| Nemotron 3 / 3.5 / Cascade 2 | 55 | Nano 4B, Nano 30B-A3B, Nano Omni 30B (audio+vision), Super 120B-A12B, Lightning 30B, Cascade 2 (collection) |
| Qwen 3.6 | 25 | 27B dense + 35B-A3B MoE, full vision tower (collection) |
| Qwen 3.5 / 3.8 | 22 | 27B, 122B-A10B, 397B-A17B; Qwen3.8-27B (collection) |
| Qwen agents & coders | 14 | Qwen3-Coder-Next, Qwen-AgentWorld-35B-A3B |
| MiniCPM 5 | 25 | 1B base / instruct / SFT edge models |
| Ornith 1.0 / 1.5 | 16 | 9B, 35B, 35B-A3B reasoning MoE |
| Embeddings | 29 | Qwen3-Embedding 0.6B/4B/8B, UEmbed 2B/4B/9B, nomic-embed-text-v2-moe, form-embed (UEmbed, ONNX) |
| Speech — ASR & TTS | 37 | MERaLiON-3 3B/10B, Qwen3-ASR, Voxtral Mini/Realtime/TTS, MOSS-Transcribe, cohere-transcribe-arabic, Kokoro, fishaudio-s2-pro, Audio8-TTS, Gemma-4-E4B MERaLiON speech LoRAs (ASR on Apple Silicon, TTS on Apple Silicon, MERaLiON) |
| gpt-oss | 10 | 20B + 120B rebuilt GGUFs / MLX (collection) |
| LFM 2.5 | 10 | 2.6B dense, 8B-A1B MoE |
| harrier-oss | 10 | 270M / 0.6B / 27B |
| Mistral | 11 | Medium 3.5 128B, Small 4 119B, Leanstral |
| Vision & agents | 29 | Muse Glimmer 30B (collection), UI-Mate 27B, BigBang v1, Unlimited-OCR, Qwen-Image-Bench, GELab-Zero (OCR & DocAI) |
| Other LLMs | 16 | KAT-Coder V2.5, Shieldstral 3B, MiniMax M2.7, DeepSeek-V4-Flash, Qwen2.5-1.5B DWQ reference, cga-gpt |
Calibration and evaluation sets used by the lanes, mirrored so results are
reproducible: ultrachat-calib, c4-calib, ultra-fineweb-calib,
ultradata-{sft,math}-calib, tulu-3-sft-mixture, wikitext-2-ppl,
gsm8k, ifeval, fleurs-ar-asr, WildASR, SASRBench-v1,
magpie-reasoning-qwen25-7b, and the gate ledger garden-quant-bench.
-ctk q8_0 -ctv q8_0 (see tip above).Older repos in this org carry RotorQuant or TurboQuant in their names.
These are historical release labels, not distinct quantization
algorithms: for any given tier both labelled repos hold byte-identical
weights produced with the standard MLX / llama.cpp quantizers, and no
brand-specific speedup is claimed or measured. New releases are unbranded.
The associated llama.cpp fork is unmaintained; use upstream KV-cache
quantization as described in the tip at the top.
majentik publishes these to keep our own fleet running cheaply on commodity Apple hardware and to close the gap between a research release and "can I actually run this tonight". Issues, quant requests and benchmark PRs are welcome via the Community tab on the closest repo.
Each repo tracks upstream@base-model-revision × quant-lane. When upstream
ships a new base revision we re-run the lane and bump the repo. Card-only
changes do not bump the version.
Each repo inherits the base model's license, not this
organization-level license. Check the license field in the repository's
card before deploying.