Downloads · 30 days
0
Xrenya/smolVLM
smolVLM is a image-text-to-text model from Xrenya. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product.
The model was trained on 491M tokens, but overall make decent predictions.
Downloads · 30 days
0
Access
Public
Updated Sep 1, 2026
Repo size
678 MB
Likes
0
Public
Click a slice to open those files.
.safetensors641 MB · 94%
From the Hugging Face model README
The model was trained on 491M tokens, but overall make decent predictions.
Example:

Generate output:
[!NOTE] The image displays an indoor setting with a large, open-air display featuring a variety of clothing items. The display is positioned in front of a large, wooden sign with the text "UNIDO" in white letters, which is surrounded by a green background. The sign is positioned on a wooden floor, and the sign is adorned with a green and white banner with the text "UNIDO" in white letters.