Downloads · 30 days
0
HackerAditya56/NutriScan-3B
NutriScan-3B is a image-to-text model from HackerAditya56. Use it when you need a caption or text from an image. It is set up for transformers.
Downloads · 30 days
0
Access
Public
Updated Jan 25, 2026
Repo size
85.3 MB
Likes
1
Public
Click a slice to open those files.
.safetensors14.8 MB · 56%
From the Hugging Face model README
NutriScan-3B is a specialized Vision-Language Model (VLM) designed to analyze food images and output structured nutritional data. Built for the MedGemma Impact Challenge, it acts as the intelligent "Vision Layer" for AI health pipelines.
It is fine-tuned on Qwen2.5-VL-3B-Instruct, bridging the gap between raw culinary images and medical-grade nutritional analysis.
This model was fine-tuned on the Codatta/MM-Food-100K dataset. To ensure high data quality and download reliability during the hackathon, we curated a specific subset:
food_099996.jpg) preserve their original index from the source dataset.You must install the latest transformers libraries to support Qwen2.5-VL.
pip install git+https://github.com/huggingface/transformers
pip install peft accelerate bitsandbytes qwen-vl-utils
import torch
from PIL import Image
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
from peft import PeftModel
from qwen_vl_utils import process_vision_info
# 1. Load Model & Adapter
base_model = "Qwen/Qwen2.5-VL-3B-Instruct"
adapter_model = "HackerAditya56/NutriScan-3B"
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
base_model, torch_dtype=torch.float16, device_map="auto"
)
model = PeftModel.from_pretrained(model, adapter_model)
processor = AutoProcessor.from_pretrained(base_model, min_pixels=256*28*28, max_pixels=1024*28*28)
# 2. Run Analysis
def scan_food(image_path):
image = Image.open(image_path).convert("RGB")
# We use a specific prompt to force JSON output
messages = [{
"role": "user",
"content": [
{"type": "image", "image": image},
{"type": "text", "text": "You are a nutritionist. Identify this dish, list ingredients, and estimate nutrition in JSON format."}
]
}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
text=[text], images=image_inputs, videos=video_inputs, padding=True, return_tensors="pt"
).to("cuda")
generated_ids = model.generate(**inputs, max_new_tokens=512)
return processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
# Test
print(scan_food("my_lunch.jpg"))
Input: Image of a pepperoni pizza. Model Output:
{
"dish_name": "Pepperoni Pizza",
"ingredients": ["pizza dough", "tomato sauce", "mozzarella cheese", "pepperoni slices", "oregano"],
"nutritional_profile": {
"calories_per_slice": 280,
"protein": "12g",
"fat": "10g",
"carbs": "35g"
},
"health_note": "Contains processed meat and high sodium."
}
Not Medical Advice. This AI estimates nutrition based on visual features. It cannot detect hidden ingredients (sugar, salt, oils) or allergens with 100% accuracy. Use for educational and tracking purposes only.
Aditya Nandan (HackerAditya56) Developed for the MedGemma Hackathon 2026