Downloads · 30 days
15
20% of all-time downloads
kawchar85/SmolLM2-1.7B-Instruct-TIFA-Random
SmolLM2-1.7B-Instruct-TIFA-Random is a text generation model from kawchar85. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
SmolLM2-1.7B-Instruct-TIFA-Random is a fine-tuned version of unsloth/SmolLM2-1.7B-Instruct specifically trained for TIFA (Text-to-Image Faithfulness Assessment) with flexible question generation. Unlike previous struc…
Downloads · 30 days
15
20% of all-time downloads
All-time downloads
75
Public
Parameters
1.7B
3.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.4 GB · 100%
From the Hugging Face model README
SmolLM2-1.7B-Instruct-TIFA-Random is a fine-tuned version of unsloth/SmolLM2-1.7B-Instruct specifically trained for TIFA (Text-to-Image Faithfulness Assessment) with flexible question generation. Unlike previous structured versions, this model generates diverse, natural evaluation questions without rigid formatting constraints, making it more adaptable for various evaluation scenarios.
Model Series: 135M | 360M | 1.7B-Structured | 1.7B-Random
This model represents a paradigm shift from rigid question structures to flexible, natural question generation:
This model generates 4 visual verification questions for text-to-image evaluation, focusing on:
Training Method: Supervised Fine-Tuning with category-balanced validation
Enhanced LoRA Configuration:
["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]Optimized Training Parameters:
pip install transformers torch
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
import torch
model_path = "kawchar85/SmolLM2-1.7B-Instruct-TIFA-Random"
# Load model and tokenizer
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "right"
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype=torch.float16,
trust_remote_code=True,
device_map="auto"
)
# Create pipeline
chat_pipe = pipeline(
"text-generation",
model=model,
tokenizer=tokenizer,
return_full_text=False,
)
def get_message(description):
system = """\
You are a TIFA (Text-to-Image Faithfulness evaluation with question Answering) question generator. Given an image description, create exactly 4 visual verification questions with multiple choice answers. Each question should test different visual aspects that can be verified by looking at the image.
Guidelines:
- Focus on colors, shapes, objects, materials, spatial relationships, and other visually verifiable elements
- Mix yes/no questions (2 choices: "no", "yes") and multiple choice questions (4 choices)
- Each question should test a DIFFERENT aspect of the description
- Ensure questions can be answered by visual inspection of the image
- Use elements explicitly mentioned in the description
- Include both positive verification (testing presence, answer: "yes") and negative verification (testing absence, answer: "no")
- Make distractors realistic and relevant to the domain
Format each question as:
Q[number]: [question text]
C: [comma-separated choices]
A: [correct answer]
Generate questions that test visual faithfulness between the description and image."""
user_msg = f'Create 4 visual verification questions for this description: "{description}"'
return [
{"role": "system", "content": system},
{"role": "user", "content": user_msg}
]
# Generate evaluation questions
description = "a lighthouse overlooking the ocean"
messages = get_message(description)
output = chat_pipe(
messages,
max_new_tokens=256,
do_sample=False,
)
print(output[0]["generated_text"])
For "a lighthouse overlooking the ocean":
Q1: What type of structure is prominently featured?
C: windmill, lighthouse, tower, castle
A: lighthouse
Q2: What body of water is visible?
C: lake, river, ocean, pond
A: ocean
Q3: Is the lighthouse positioned above the water?
C: no, yes
A: yes
Q4: Are there any mountains in the scene?
C: no, yes
A: no
@misc{smollm2-1-7b-it-tifa-random-2025,
title={SmolLM2-1.7B-Instruct-TIFA-Random: Flexible Question Generation for Text-to-Image Faithfulness Assessment},
author={kawchar85},
year={2025},
url={https://huggingface.co/kawchar85/SmolLM2-1.7B-Instruct-TIFA-Random}
}