Downloads · 30 days
4
9% of all-time downloads
Yougen/InternVL3Fangwusha2B
InternVL3Fangwusha2B is a machine learning model from Yougen. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
InternVL3Fangwusha2B is a 2B-parameter vision-language model (VLM) fine-tuned from InternVL3-2B, optimized for Chinese multimodal understanding, image-text interaction, document analysis, and visual content reasoning…
Downloads · 30 days
4
9% of all-time downloads
All-time downloads
44
Public
Parameters
2.1B
4.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.2 GB · 100%
From the Hugging Face model README
InternVL3Fangwusha2B is a 2B-parameter vision-language model (VLM) fine-tuned from InternVL3-2B, optimized for Chinese multimodal understanding, image-text interaction, document analysis, and visual content reasoning in practical application scenarios.
This model is a lightweight multimodal large model based on the InternVL3-2B architecture, focusing on improving Chinese image-text alignment, visual question answering, document OCR + understanding, and daily scene multimodal interaction. It provides efficient inference while maintaining strong multimodal capabilities.
This model can be directly used for:
Can be further fine-tuned for:
All outputs in professional or production scenarios should be verified by humans. Visual inputs with sensitive or private information require appropriate filtering and protection. It is recommended to access professional modules for high-precision visual tasks. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model.
Use the code below to get started with the model.
from transformers import AutoModel, AutoTokenizer
model_name = "Yougen/InternVL3Fangwusha2B"
model = AutoModel.from_pretrained(model_name, device_map="auto", torch_dtype="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)
# Example image-text interaction
# image = load_image("your_image.jpg")
# response = model.chat(tokenizer, image, "请描述这张图片的内容")
# print(response)
Training data includes high-quality Chinese image-text pairs, document images, daily scene photos, and instruction-based multimodal dialogue data. All data is processed with deduplication, noise filtering, and quality screening.
Internal Chinese multimodal test set including VQA, image captioning, and document understanding tasks.
Image quality, text complexity, scene type, document layout complexity.
[More Information Needed]
The model achieves stable performance in Chinese multimodal understanding and interaction, with high efficiency suitable for edge and middle-end deployment scenarios.
[More Information Needed]
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Vision-language architecture with visual encoder + large language decoder, based on InternVL3. Optimized for Chinese multimodal understanding, alignment, generation, and document analysis.
NVIDIA GPU with CUDA support
BibTeX:
[More Information Needed]
APA:
[More Information Needed]
For updates and issues, please visit the model repository on Hugging Face Hub.
Yougen Yuan
[More Information Needed]