Downloads · 30 days
28
10% of all-time downloads
cminst/Llama-3.2-11B-VisionEncoder
Llama-3.2-11B-VisionEncoder is a image feature extraction model from cminst. Use it for the image feature extraction task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This repository contains only the MllamaVisionModel weights extracted from unsloth/Llama-3.2-11B-Vision. It intentionally excludes the text decoder and language model weights.
Downloads · 30 days
28
10% of all-time downloads
All-time downloads
279
Public
Parameters
836M
3.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.3 GB · 100%
From the Hugging Face model README
This repository contains only the MllamaVisionModel weights extracted from
unsloth/Llama-3.2-11B-Vision. It intentionally excludes the text decoder and language
model weights.
Load it with:
from transformers import MllamaVisionModel
vision = MllamaVisionModel.from_pretrained("cminst/Llama-3.2-11B-VisionEncoder")
For LEGATO configs, set encoder_pretrained_model_name_or_path to:
cminst/Llama-3.2-11B-VisionEncoder