Downloads · 30 days
24
13% of all-time downloads
LongGrainRice/kimchi-test
kimchi-test is a image-text-to-text model from LongGrainRice. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
A LoRA fine-tune of SmolVLM2-500M-Video-Instruct, merged into a standalone model. Given a photo of ingredients arranged on a surface, it returns a JSON list of the ingredients it recognizes. Built as an end-to-end ML…
Downloads · 30 days
24
13% of all-time downloads
All-time downloads
189
Public
Parameters
507M
2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1 GB · 100%
From the Hugging Face model README
A LoRA fine-tune of SmolVLM2-500M-Video-Instruct, merged into a standalone model. Given a photo of ingredients arranged on a surface, it returns a JSON list of the ingredients it recognizes. Built as an end-to-end ML portfolio project (data → fine-tune → inference endpoint → web app).
V2 change: the vocabulary was expanded from 51 to ~353 ingredient classes by unioning the original 51-class set with a 316-class dataset (14 ingredients shared and merged to a single canonical label). This directly targets V1's main failure mode — padding the output with frequent in-vocab guesses whenever an out-of-vocabulary ingredient appeared — by giving many of those previously-unknown items a real label.
import torch
from transformers import AutoProcessor, AutoModelForImageTextToText
from PIL import Image
repo = "LongGrainRice/kimchi-test"
processor = AutoProcessor.from_pretrained(repo)
# bf16 needs an Ampere+ GPU (e.g. L4). On a T4 or older card, use torch.float16.
model = AutoModelForImageTextToText.from_pretrained(repo, torch_dtype=torch.bfloat16).to("cuda")
# Use the exact instruction the model was trained with — it keys on this wording.
INSTRUCTION = ("You are a food recognition assistant. List every food ingredient in this image. "
'Respond ONLY with a JSON array of lowercase strings, e.g. ["milk", "eggs", "tomato"].')
image = Image.open("slab.jpg").convert("RGB")
messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": INSTRUCTION}]}]
prompt = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
inputs = processor(text=[prompt], images=[[image]], return_tensors="pt", padding=True).to(model.device)
ids = model.generate(**inputs, max_new_tokens=128, do_sample=False)
print(processor.batch_decode(ids[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])
De-duplicate the output before using it: sorted(set(...)).
Synthetic composites built by unioning two single-item sources, both center-cropped with an oval mask and pasted onto slab backgrounds with varied scale, rotation, and shadow; the label is the set of pasted items.
liamboyd1/singular-food-items
— 51 classes, image-rich (~1.3k/class), capped per class for variety.Scuccorese/food-ingredients-dataset
— 316 classes, ~21 images/class, with a 12-category / 28-subcategory hierarchy.The 14 ingredients shared between the two sets are normalized to a single canonical name so they don't fragment into duplicate classes. Because the two sources differ ~60× in per-class image count, the compositor samples classes uniformly when building scenes rather than sampling images uniformly — otherwise the image-rich legacy classes would dominate and the new vocabulary would rarely appear. A small set of real photos is mixed in and upweighted.
sorted(set(...))).Best used on scenes composed from the known classes. A base-vs-fine-tuned comparison script is included in the project repo for evaluating on your own slab photos.