Downloads · 30 days
14
10% of all-time downloads
ReconAI/Qwen2.5-VL-1B-Instruct-Custom
Qwen2.5-VL-1B-Instruct-Custom is a machine learning model from ReconAI. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
[](https://opensource.org/licenses/MIT) [](https://huggingface.co/ReconAI/Qwen2.5-VL-1B-Instruct-Custom)
Downloads · 30 days
14
10% of all-time downloads
All-time downloads
147
Public
Parameters
1.2B
2.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.3 GB · 99%
From the Hugging Face model README
一个轻量级视觉语言对话模型,基于Qwen2.5架构进行深度定制优化,在仅1B参数规模下实现高效的图文理解与对话能力。
Qwen2.5-VL-1B-Instruct-Custom 是通过对Qwen2.5-VL-3B-Instruct进行架构级轻量化改造而构建的定制化视觉语言模型。该模型保留了原始模型的高性能ViT视觉编码器,同时将LLM模块替换为更轻量的Qwen2.5-0.5B-Instruct,并重新设计MLP层尺寸以实现模块间的最优适配。经过多阶段渐进式训练,模型在保持基础图文对话能力的同时,显著降低了资源消耗,特别适合边缘设备部署和资源受限场景。
| 组件 | 来源 | 说明 |
|---|---|---|
| 视觉编码器 (ViT) | naViT(630M), 源自Qwen2.5-VL-3B-Instruct | 完整保留原始ViT架构,确保视觉特征提取能力 |
| 语言模型 (LLM) | Qwen2.5-0.5B-Instruct | 替换为轻量级LLM,大幅降低参数量 |
| MLP投影层 | 修改merger的第二层linear参数 | 调整尺寸以适配ViT输出与LLM输入的维度对齐 |
| 总参数量 | - | 约1B (ViT部分约0.6B + LLM约0.5B + MLP及其他) |
为确保轻量化架构的稳定收敛,采用分阶段解冻训练方案:
模型在多个高质量开源VQA数据集上进行混合训练:
| 数据集 | 样本量 | 说明 |
|---|---|---|
| llava-150k-instruct | 15w | |
| webui-ocr-zh | 2.7w | |
| mscoco | ||
| rlaif-v-dataset | ||
| VisualTableQA | ||
| DocVAQ | ||
| pixmo_docs | 25w |
import torch
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
from qwen_vl_utils import process_vision_info
model_path = '/path/to/model/Qwen2.5-VL-1B-Instruct-Custom'
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(model_path,
torch_dtype='bfloat16',
device_map='auto',
trust_remote_code=True)
processor = AutoProcessor.from_pretrained(model_path, trust_remote_code=True)
messages = [
{
"role": "user",
"content": [
{
"type": "image",
"image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg",
},
{"type": "text", "text": "Describe this image."},
],
}
]
# Preparation for inference
text = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
text=[text],
images=image_inputs,
videos=video_inputs,
padding=True,
return_tensors="pt",
)
inputs = inputs.to("cuda")
# Inference: Generation of the output
generated_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids_trimmed = [
out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text)