Downloads · 30 days
7
100% of all-time downloads
LNTTushar/perception-slm-caption
perception-slm-caption is a machine learning model from LNTTushar. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for perception-slm. The card lists the license as apache-2.0.
Small, from-scratch image-understanding model (Perception SLM, Phase 2 / image v0). Encoder → connector (resampler) → tiny decoder, trained with Stage-2 alignment and Stage-3 LoRA instruction tuning. Config: captionlm.
Downloads · 30 days
7
100% of all-time downloads
All-time downloads
7
Public
Repo size
30.6 MB
Likes
0
Public
Click a slice to open those files.
.pt30.6 MB · 100%
From the Hugging Face model README
Small, from-scratch image-understanding model (Perception SLM, Phase 2 / image v0).
Encoder → connector (resampler) → tiny decoder, trained with Stage-2 alignment and
Stage-3 LoRA instruction tuning. Config: caption_lm.
| metric | value |
|---|---|
| loss | 2.3982 |
| bleu4 | 10.07 |
model.pt — checkpoint (model_state + training metadata)config.yaml — the exact config used to build the modelmodel_int8.pt — int8-quantized weights for CPU/offline (if uploaded)from huggingface_hub import hf_hub_download
import torch
ckpt = torch.load(hf_hub_download("LNTTushar/perception-slm-caption", "model.pt"), map_location="cpu")
# rebuild ImageVLM.from_config(config) then load_state_dict(ckpt["model_state"])
Built with the perception-slm repo; see its RESULTS.md.