Downloads · 30 days
71
56% of all-time downloads
immanuelpeter/MoonViT-K2.6
MoonViT-K2.6 is a image feature extraction model from immanuelpeter. Use it for the image feature extraction task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as other.
This repository packages the Kimi K2.6 MoonViT Tower and Projector as a self-contained Transformers model. It does not download the full 64-shard checkpoint or fetch architecture code from another repository at runtime.
Downloads · 30 days
71
56% of all-time downloads
All-time downloads
127
Public
Parameters
417M
942 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors942 MB · 100%
From the Hugging Face model README
This repository packages the Kimi K2.6 MoonViT Tower and Projector as a self-contained Transformers model. It does not download the full 64-shard checkpoint or fetch architecture code from another repository at runtime.
| File | Tensors | What it holds |
|---|---|---|
model.safetensors | 329 | MoonViT Tower |
projector.safetensors | 6 | Kimi K2.6 Projector |
projector_config.json, projector.py | Projector shapes and loader | |
config.json, configuration_moonvit.py, modeling_moonvit.py | Standalone Transformers model | |
preprocessor_config.json, kimi_k25_vision_processing.py, media_utils.py | Image preprocessing |
| Component | Details |
|---|---|
| Tower | 27 layers, 1152 hidden, 16 heads, 4304 intermediate, patch size 14 |
| Token compression | 2x2 spatial regrouping plus temporal pooling, no learned parameters |
| Projector | LayerNorm(1152), flatten four patches, Linear(4608, 4608), GELU, Linear(4608, 7168) |
The Tower loads through AutoModel.from_pretrained(..., trust_remote_code=True). The
processor accepts the standard images=... interface, and projector.py maps the merged
tokens to the Kimi K2.6 language width.
See examples/inference.py for image-to-projected-features
inference.
The release tests compare all 335 packaged tensors with shards 63 and 64 of the pinned
Moonshot checkpoint using torch.equal. Fixed-image merged and projected outputs also
match the parent implementation bit-for-bit on CPU and in BF16 on an NVIDIA A100.
The export script
splits the BF16 weights published by
exolabs/Kimi-K2.6-vision into separate
Tower and Projector files. It derives a vision-only loader and processor from the pinned
Moonshot implementation.
Moonshot AI released Kimi K2.6 and its modeling code. Exolabs published the combined vision-only weight file used by the exporter.
Kimi K2.6 License, the same license as the source model. The upstream third-party notices are included.