Downloads · 30 days
220
8% of all-time downloads
utter-project/EuroVLM-1.7B-Preview
EuroVLM-1.7B-Preview is a image-text-to-text model from utter-project. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
⚠️ PREVIEW RELEASE: This is a preview version of EuroVLM-1.7B. The model is still under development and may have limitations in performance and stability. Use with caution in production environments.
Downloads · 30 days
220
8% of all-time downloads
All-time downloads
2.9K
Public
Parameters
2.1B
4.1 GB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors4.1 GB · 100%
From the Hugging Face model README
⚠️ PREVIEW RELEASE: This is a preview version of EuroVLM-1.7B. The model is still under development and may have limitations in performance and stability. Use with caution in production environments.
This is the model card for EuroVLM-1.7B-Preview, a multimodal vision-language model based on long-context version of EuroLLM-1.7B.
EuroVLM-1.7B is a 1.7B+400M parameter vision-language model that combines the multilingual capabilities of EuroLLM-1.7B with vision encoding components.
EuroVLM-1.7B was (visually) instruction tuned on a combination of multilingual vision-language datasets, including image captioning, visual question answering, and multimodal reasoning tasks across the supported languages.
EuroVLM uses a multimodal architecture combining a vision encoder with the EuroLLM language model:
Language Model Component:
Vision Component:
To use the model with HuggingFace's Transformers library
from PIL import Image
from transformers import LlavaNextProcessor, LlavaNextForConditionalGeneration
model_id = "utter-project/EuroVLM-1.7B-Preview"
processor = LlavaNextProcessor.from_pretrained(model_id)
model = LlavaNextForConditionalGeneration.from_pretrained(model_id)
# Load an image
image = Image.open("/path/to/image.jpg")
messages = [
{
"role": "system",
"content": "You are EuroVLM --- a multimodal AI assistant specialized in European languages that provides safe, educational and helpful answers about images and text.",
},
{
"role": "user",
"content": [
{"type": "image"},
{"type": "text", "text": "What do you see in this image? Please describe it in Portuguese."}
]
},
]
prompt = processor.apply_chat_template(messages, add_generation_prompt=True)
inputs = processor(images=image, text=prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=1024)
print(processor.decode(outputs[0], skip_special_tokens=True))
You can also run EuroVLM with vLLM!
from vllm import LLM, SamplingParams
# Initialize the model
model_id = "utter-project/EuroVLM-1.7B-Preview"
llm = LLM(model=model_id)
# Set up sampling parameters
sampling_params = SamplingParams(temperature=0.7, max_tokens=1024)
# Image and prompt
image_url = "/url/of/image.jpg"
messages = [
{
"role": "system",
"content": "You are EuroVLM --- a multimodal AI assistant specialized in European languages that provides safe, educational and helpful answers about images and text.",
},
{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": image_url}},
{"type": "text", "text": "What do you see in this image? Please describe it in Portuguese in one sentence."}
]
},
]
# Generate response
outputs = llm.chat(messages, sampling_params=sampling_params)
print(outputs[0].outputs[0].text)
EuroVLM-1.7B-Instruct supports a wide range of vision-language tasks across multiple languages:
EuroVLM-1.7B has not been fully aligned to human preferences, so the model may generate problematic outputs in both text and image understanding contexts (e.g., hallucinations about image content, harmful content, biased interpretations, or false statements about visual information).
Additional considerations for multimodal models include:
Users should exercise caution and implement appropriate safety measures when deploying this model in production environments.