Downloads · 30 days
90
76% of all-time downloads
divinetribe/Vision-Narrator-0.8B-4bit-mlx
Vision-Narrator-0.8B-4bit-mlx is a image-text-to-text model from divinetribe. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as apache-2.0.
A vision-language model small enough to live on a phone, trained to answer the questions blind people actually ask about what's in front of them.
Downloads · 30 days
90
76% of all-time downloads
All-time downloads
119
Public
Parameters
853M
1.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors625 MB · 96%
How the weights are stored.
U32752M · 88%
From the Hugging Face model README
A vision-language model small enough to live on a phone, trained to answer the questions blind people actually ask about what's in front of them.
This is the model inside RealTime AI Cam, a free camera app for blind and low-vision people from Nice Dreamz. You point the camera at a room, a letter, or the coffee maker, ask it a question, and it answers out loud. Everything runs offline, on the phone.
Get the app: App Store (iPhone) · Google Play (Android) · Source on GitHub
Earlier versions described the whole photo no matter what you asked. Ask "is this light on?" and you'd get a paragraph about the lamp, the table and the window blinds, and never hear the answer.
v09 was retrained on questions real blind people asked about photos they took themselves, from the VizWiz dataset. Now the answer comes first:
| Question | v08 (before) | v09 (now) |
|---|---|---|
| Is this light on? | "A lamp is on a small wooden table in front of the window, which has closed blinds above it..." | "Yes, the lamp is on, casting a warm glow across the room." |
| Is this cup full? | "This is a clear Starbucks Coffee tumbler with a green logo on it, and it is sitting on a white surface..." | "No, it's empty, with just a green straw sticking straight down into the clear plastic." |
| Does the screen say anything now? | "The photo is too blurry and dark to read the screen..." | "No, the screen is dark and blank, with just a faint reflection of light on the glass." |
| Is there a little light on under the Add Water, or is that off? | "The photo shows a machine's control panel with three round buttons..." | "No, there is no light on under the add water; only a green power light and blue icons are visible." |
Those are unedited outputs. The first three are VizWiz photos the model never saw in training. The last one is a photo of our own Keurig.
Tested on 295 VizWiz questions whose photos were kept out of training, scored with the standard VizWiz accuracy (agreement with 10 human answers):
| v08 | v09 | |
|---|---|---|
| Overall | 46.1% | 66.4% |
| Yes/no questions | 10.9% | 62.5% |
| Number questions | 47.3% | 61.3% |
| Other questions | 50.2% | 60.7% |
| Correctly says "I can't tell" on unanswerable photos | 78.0% | 92.0% |
| Puts the answer first | 34.6% | 65.8% |
The size didn't change (about 620 MB). Reading mail and describing rooms held steady: 192/200 facts right on the describe test, vs 194/200 before.
Still weak: elevator direction. Asked "is the elevator going up or down?", it usually says it can't tell. Ask "which arrow is lit?" instead and it gets most of them right. Direction indicators are what the next round adds.
It's trained to be useful, not just descriptive. A generic captioner tells a blind user "a piece of paper with writing on it." This model answers the four things that matter on a letter: what it is, who sent it, how much, and what happens if you ignore it. On a scene, it says what's in front of you and where.
The part that sees isn't compressed. The language half runs at 4-bit (affine, group size 64). The vision tower stays at full precision, because compressing it made the model misread dollar amounts. On this model, 8-bit vision dropped accuracy to 96.2% and 4-bit vision dropped it to 91.5%.
It says when it can't tell. It's trained to say something like "I can't read this page, the photo is too dark. Add some light or turn on the flash and try again" instead of inventing a plausible bill. For this audience that's a feature.
It runs on a 2022 phone. Tested on an iPhone 14 Pro Max, fully offline.
Scene: "A white mattress with a star pattern is right in front of you, so close it blocks most of the room. A person in purple pants is lying on a patterned rug to your left, with a bright orange heater sitting on the floor just ahead of them."
Bill: "This is a utility bill for $81.13, due October 8th, 2026. If it goes unpaid they can shut off service 15 days after that."
Bad photo: "The top of the page is cut off, so I can't tell who sent this. Move the phone back a little so the whole page is in view and try again."
Two system prompts, one per mode. For a question about a photo (what v09 was trained for):
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
ASK = ("You are the eyes of a blind person. They are asking a question about this photo. "
"Answer it in one short spoken sentence, and start with the answer itself: yes or no, "
"a number, a colour, a name, or the words that are printed. Then, only if it helps, a few "
"words about where or why. If the photo does not show the answer, say you can't tell for "
"sure and say what you can see. Plain words, no lists, no markdown.")
model, processor = load("divinetribe/Vision-Narrator-0.8B-4bit-mlx")
messages = [{"role": "system", "content": ASK},
{"role": "user", "content": "Is this light on?"}]
prompt = apply_chat_template(processor, model.config, messages, num_images=1)
print(generate(model, processor, prompt, image="photo.jpg", max_tokens=96).text)
For "what's in front of me" or reading mail, use the describe prompt from the app source. On iPhone it runs through MLX Swift. The Android app runs a GGUF conversion of these weights through llama.cpp.
| Base | Qwen3.5 0.8B (Qwen3_5ForConditionalGeneration) |
| Language model | 1024 hidden, 24 layers, 8 heads, 248,320 vocab |
| Vision tower | 768 hidden, depth 12, projects to 1024 |
| Quantization | 4-bit affine, group size 64, language model only; vision tower full precision |
| On disk | about 620 MB |
| Runtime | MLX / MLX Swift |
LoRA fine-tune of Qwen3.5 0.8B, rank 16, on all 24 language-model layers, with the vision tower frozen, then fused and quantized. The model learned what to say about what it already sees.
The previous version is still available under the v08 tag of this repo.
English only. Tuned for US household documents, indoor scenes and everyday objects. It will sometimes misread a brand name or a sender on a badly lit page. It's a 0.8B model, not a human reader. Don't use it as the only source of truth for anything financial or medical. It's a fast first look for someone who would otherwise get nothing.
VizWiz. The v09 training questions and the scene photos come from the VizWiz datasets: photos taken and questions asked by blind people, collected by Danna Gurari and colleagues at the University of Colorado Boulder, licensed CC BY 4.0. Thank you for making this possible.
Qwen. Built on Qwen3.5 0.8B by the Qwen team at Alibaba (Apache 2.0).
MLX. Quantized and served with MLX and mlx-vlm.
Trained by Matt Macosko and released through Nice Dreamz as divinetribe.