Skip to content

Zirconium-Oxide

FoodExtract-Vision

Zirconium-Oxide/FoodExtract-Vision

FoodExtract-Vision is a image-text-to-text model from Zirconium-Oxide. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers.

This model is a fine-tuned version of HuggingFaceTB/SmolVLM2-500M-Video-Instruct. It has been trained using TRL.

Downloads · 30 days

4

9% of all-time downloads

All-time downloads

43

Public

Parameters

507M

1 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors1 GB · 100%

At a glance

Task
Image-Text-to-Text
Library
transformers
Model type
smolvlm
Access
Public
Created
Apr 4, 2026
Updated
Apr 4, 2026
SHA
8d6c5a74

Try a prompt

Base models

Task
Image-Text-to-Text
Library
transformers
Type
smolvlm
Created
Apr 4, 2026
Updated
Apr 4, 2026
FoodExtract-Vision — AI Model — AIMarketly