Downloads · 30 days
0
Bukunmi2108/capit-sat
capit-sat is a image-to-text model from Bukunmi2108. Use it when you need a caption or text from an image. It is set up for pytorch. The card lists the license as mit.
Show, Attend and Tell image captioner, trained from scratch on Flickr8k (Karpathy split). The glass-box half of capit — exposes per-word attention, beam candidates, and word-by-word playback.
Downloads · 30 days
0
Access
Public
Updated Jun 10, 2026
Repo size
148 MB
Likes
0
Public
Click a slice to open those files.
.pt148 MB · 100%
From the Hugging Face model README
Show, Attend and Tell image captioner, trained from scratch on Flickr8k (Karpathy split). The glass-box half of capit — exposes per-word attention, beam candidates, and word-by-word playback.
| beam | BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | CIDEr |
|---|---|---|---|---|---|
| 1 | 61.99 | 44.37 | 30.23 | 20.05 | 55.51 |
| 3 | 64.77 | 47.34 | 33.68 | 23.45 | 62.20 |
| 5 | 65.54 | 47.84 | 34.08 | 23.63 | 62.80 |
Attention is effectively 7x7: ResNet-50 at 224px is natively 7x7 and the encoder upsamples to 14x14, so heatmaps are coarse (~32px blocks). Captions are grounded; the spots are region-level, not pixel-level.
huggingface_hub.hf_hub_download("Bukunmi2108/capit-sat", "capit-sat.pt") + vocab.json, then
capit.serving.load_artifact(...).