Downloads · 30 days
44
10% of all-time downloads
EximiusLabs/fusion-embedding-2-k3-vision
fusion-embedding-2-k3-vision is a feature extraction model from EximiusLabs. Use it when you need embeddings to search or compare text. It is set up for fusion-embedding. The card lists the license as cc-by-nc-4.0.
<img src="assets/k3-vision-banner.png" alt="fusion-embedding-2-k3-vision — Eximius Labs" width="100%"
Downloads · 30 days
44
10% of all-time downloads
All-time downloads
440
Public
Parameters
3.2M
22 MB on disk
Likes
8
Public
Click a slice to open those files.
.safetensors12.6 MB · 80%
From the Hugging Face model README
Kimi K3's vision encoder, made language-searchable.
<div align="center"> </div>This is a small trained projector that maps Kimi K3's frozen vision encoder (MoonViT-V2, about 401M parameters) into the frozen fusion-embedding-2 text space. K3's own visual features become searchable in plain language, and land in the same space as fusion-embedding's text, image, video, and audio, plus its thermal (Ember) and motion (Tremor) sensor packs.
It builds on only the ~401M vision tower of Kimi K3, never the 2.8T language model. The tower runs on a single GPU. The projector is the only trained weight, about 3.15M parameters.
Text-to-image retrieval on Flickr30k (test), 1000-image gallery, querying in the fusion-embedding-2 text space:
| Metric | fusion-embedding-2-k3-vision | same recipe, DINOv2 backbone |
|---|---|---|
| R@1 | 0.432 | 0.318 |
| R@5 | 0.740 | 0.613 |
| R@10 | 0.836 | 0.749 |
MoonViT-V2 was trained alongside a language model, so its features align to text more readily than a self-supervised vision backbone. On the identical projector recipe it beats the DINOv2 version by nearly 9 points at R@10.
The projector ships here; the vision tower is pulled from moonshotai/Kimi-K3 on first use
(only the vision_tower.* tensors, never the 2.8T language model).
from inference import K3VisionEmbedder
m = K3VisionEmbedder.from_pretrained("EximiusLabs/fusion-embedding-2-k3-vision")
v = m.embed_image("photo.jpg") # 2048-d, in the fusion-embedding-2 text space
Query it with text using fusion-embedding-2 itself:
from inference import FusionEmbedder # from the fusion-embedding-2 repo
fe = FusionEmbedder.from_pretrained("EximiusLabs/fusion-embedding-2-2b-preview")
q = fe.embed_text("a dog on a beach")
score = float(q @ v) # cosine, both unit-norm
LayerNorm -> Linear -> GELU -> Dropout -> Linear) maps that
descriptor into the 2048-d fusion-embedding-2 text space.K3's projected vision does not sit alone. It shares the fusion-embedding-2 space with the Ember thermal pack and the Tremor motion pack, so one text query reaches across senses in a single index. Below, each query returns its top match from vision, thermal, and motion at once. Vision and thermal retrieve cleanly (ask for "a running dog" and it returns a photo of a running dog and a thermal image of a dog); motion (Tremor) is in the same space, though its text-to-motion retrieval is research preview.

Real data throughout: Flickr photos (vision, via this model), thermal infrared (Ember), recorded IMU (Tremor).
base_model tag points at moonshotai/Kimi-K3 because the vision tower is a component
of that model. We use only MoonViT-V2, not the 2.8T language model.One-click from the RunPod Hub.
curl -s https://api.runpod.ai/v2/<ENDPOINT_ID>/runsync \n -H "Authorization: Bearer $RUNPOD_API_KEY" \n -H "Content-Type: application/json" \n -d '{"input": {"image": "<https url | data-uri | base64>"}}'
Kimi-K3 is gated: set HF_TOKEN (with the K3 license accepted) on the endpoint. Returns 2048-d in the fusion-embedding-2 text space; query with text via the fusion-embedding-2 endpoint.
This pack is one of the modalities Engram searches. Engram is the open cross-modal memory layer for physical AI: it indexes a robot's video, audio, and motion into one embedding space and answers questions about it in plain language, including temporal reasoning that retrieval alone cannot.
pip install engram-robomem
Repo: https://github.com/Eximius-Labs/engram · PyPI: https://pypi.org/project/engram-robomem · Playground: https://www.eximiuslabs.com/playground
The projector weights and code in this repository are released under CC-BY-NC-4.0.
The vision tower is part of Kimi K3 (Moonshot AI), used under the Kimi K3 License, which permits derivative works with its notice retained. This model is a derivative that uses K3's frozen MoonViT-V2 vision encoder. Copyright 2026 Moonshot AI.