Downloads · 30 days
112
9% of all-time downloads
aidiffuser/GLM-5.2-Vision-tower-MLX
GLM-5.2-Vision-tower-MLX is a machine learning model from aidiffuser. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as mit.
Add image input to any MLX quant of GLM-5.2 with a ~1 GB sidecar: the frozen MoonViT-3d vision tower from Kimi K2.6 plus the trained 49.5M-parameter PatchMerger projector from baseten/GLM-5.2-Vision-NVFP4 (Harry Partr…
Downloads · 30 days
112
9% of all-time downloads
All-time downloads
1.2K
Public
Repo size
933 MB
Likes
0
Public
Click a slice to open those files.
.safetensors933 MB · 100%
From the Hugging Face model README
Add image input to any MLX quant of GLM-5.2 with a ~1 GB sidecar: the frozen MoonViT-3d vision tower from Kimi K2.6 plus the trained 49.5M-parameter PatchMerger projector from baseten/GLM-5.2-Vision-NVFP4 (Harry Partridge's vision retrofit), repackaged for Apple Silicon / MLX. The GLM-5.2 text backbone is untouched — text-only behavior stays byte-identical.
| File | What it is |
|---|---|
glm52_vision.safetensors (+ index) | MoonViT-3d tower (417M, 27 layers, 1152-dim, bf16, vision_tower.*) and the trained projector (mm_projector.pre_norm/linear_1/linear_2, 1152 → 2×2 merge → 4608 → 6144) in one file. Kimi's original 7168-dim projector is removed — the GLM-trained one replaces it. |
config.json | vision_config (+ text_config.hidden_size: 6144, media_placeholder_token_id: 154854) |
preprocessor_config.json, kimi_k25_*.py, media_utils.py | Baseten's reference image processor (NaViT resize, patch 14, 2×2 merge, ≤4096 tokens/image) — the exact preprocessing the projector was trained against |
GLM-5.2's stock tokenizer already contains the image tokens
(<|begin_of_image|> 154830, <|image|> 154854, <|end_of_image|> 154831) —
no tokenizer changes needed. Each image expands to its media-token count at the
<|image|> position and the projected features are substituted at those
embedding positions. You need a chat template that renders image content parts
into the marker triplet (GLM's stock template does not; Baseten ships one in
their repo).
exo's vision stack loads this repo directly as a weights_repo/processor_repo.
Model card stanza:
[vision]
image_token_id = 154854
model_type = "kimi_vl"
weights_repo = "<this repo id>"
processor_repo = "<this repo id>"
Point the card's model at a directory containing your GLM-5.2 MLX quant with
Baseten's chat_template.jinja and this repo's config.json additions
(vision_config / text_config / media_placeholder_token_id). Assembly
scripts (symlink the backbone — no weight duplication):
build_glm52_vision_dir.py
and
build_glm52_vision_tower.py.
Verified live on a 2-Mac-Studio (M3 Ultra) tensor-parallel cluster over RDMA,
against both mlx-community/GLM-5.2-DQ4plus-q8 and mlx-community/GLM-5.2-mxfp4:
temp-0 deterministic, no cross-image cache bleed, text-only outputs identical
to the plain model.
MIT, following all parents. Full chain: Z.ai (GLM-5.2, MIT) → Moonshot AI (Kimi K2.6 MoonViT tower, Modified MIT) → Harry Partridge / Baseten (projector training + reference processor, baseten/GLM-5.2-Vision-NVFP4, MIT) → exolabs (original K2.6 tower extraction for MLX) → this repackaging (tensor remap documented in the build script). None of the upstream teams were involved in this packaging.