Downloads · 30 days
33
34% of all-time downloads
onedday/FTib-VLM
FTib-VLM is a image-text-to-text model from onedday. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
FTib-VLM is a Tibetan vision-language model fine-tuned from Qwen/Qwen3-VL-8B-Instruct for multimodal understanding in low-resource language settings. It is released as part of the FTibSuite project to support reproduc…
Downloads · 30 days
33
34% of all-time downloads
All-time downloads
96
Public
Parameters
8.8B
17.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors17.5 GB · 100%
From the Hugging Face model README
FTib-VLM is a Tibetan vision-language model fine-tuned from Qwen/Qwen3-VL-8B-Instruct for multimodal understanding in low-resource language settings. It is released as part of the FTibSuite project to support reproducible Tibetan multimodal research.
onedday/FTib-VLMQwen/Qwen3-VL-8B-Instructqwen3_vlApache-2.0FTib-VLM is intended for research and experimental applications such as:
FTib-VLM is fine-tuned from Qwen/Qwen3-VL-8B-Instruct using a three-stage pipeline:
The goal is to improve Tibetan multimodal capability while preserving the strengths of the base vision-language model.
FTib-VLM shows clear improvements over the base model on Tibetan multimodal evaluation, including:
Install dependencies:
pip install -U transformers accelerate torch pillow
Replace "example.jpg" with your local image path.
from PIL import Image
from transformers import AutoProcessor, AutoModelForVision2Seq
import torch
model_id = "onedday/FTib-VLM"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForVision2Seq.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
image = Image.open("example.jpg").convert("RGB")
prompt = "请详细描述这张图片。"
inputs = processor(
text=prompt,
images=image,
return_tensors="pt",
)
inputs = {
k: v.to(model.device) if hasattr(v, "to") else v
for k, v in inputs.items()
}
generated_ids = model.generate(
**inputs,
max_new_tokens=256,
)
output = processor.batch_decode(
generated_ids,
skip_special_tokens=True,
)
print(output[0])
This model is released to support Tibetan multimodal research and improve access to low-resource language technology. However, like other vision-language models, it may produce incorrect, biased, or misleading outputs. It should be used with care in high-stakes or reliability-sensitive scenarios.
If you use this model, please cite the FTibSuite paper:
@article{xu2026ftibsuite,
title={FTibSuite: A Comprehensive Resource Suite for Tibetan Vision--Language Modeling},
author={Xu, Guixian and Liang, Yide and Su, Zeli and Song, Xuexian and Zhang, Ziyin and Dong, Yushuang and Zhang, Ting and Han, Xu},
year={2026}
}
You may also cite this repository as:
@misc{onedday_ftib_vlm,
title = {FTib-VLM},
author = {onedday},
year = {2026},
howpublished = {\url{https://huggingface.co/onedday/FTib-VLM}}
}
You may also cite this repository as:
@misc{onedday_ftib_vlm,
title = {FTib-VLM},
author = {onedday},
year = {2026},
howpublished = {\url{https://huggingface.co/onedday/FTib-VLM}}
}
```