Downloads · 30 days
9
16% of all-time downloads
Yougen/InternVL3Fangwusha8B
InternVL3Fangwusha8B is a machine learning model from Yougen. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
InternVL3Fangwusha8B is an 8B-parameter vision-language model (VLM) fine-tuned from InternVL3-8B, optimized for high-performance Chinese multimodal understanding, complex visual reasoning, document analysis, table ext…
Downloads · 30 days
9
16% of all-time downloads
All-time downloads
55
Public
Parameters
7.9B
15.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors15.9 GB · 100%
From the Hugging Face model README
InternVL3Fangwusha8B is an 8B-parameter vision-language model (VLM) fine-tuned from InternVL3-8B, optimized for high-performance Chinese multimodal understanding, complex visual reasoning, document analysis, table extraction, and image-text dialogue in industrial and advanced application scenarios.
This model is a mid-to-large scale vision-language model based on the InternVL3-8B foundation architecture. It is fine-tuned to enhance cross-modal alignment, complex image understanding, structured information extraction from documents, and multi-turn visual dialogue in Chinese. It provides strong reasoning ability while maintaining efficient deployability.
This model can be directly used for:
Can be further fine-tuned for:
All outputs used in professional or production environments must be reviewed by qualified personnel. For deployment involving user data or public scenarios, content safety and privacy protection mechanisms are strongly recommended. Professional visual modules should be used for high-precision tasks such as medical or industrial analysis. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model.
Use the code below to get started with the model.
from transformers import AutoModel, AutoTokenizer
model_name = "Yougen/InternVL3Fangwusha8B"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModel.from_pretrained(
model_name,
device_map="auto",
torch_dtype="auto",
trust_remote_code=True
).eval()
# Example usage:
# image = load_image("your_image_file.jpg")
# question = "请详细分析这张图片中的内容和结构"
# response = model.chat(tokenizer, image, question)
# print(response)
Training data consists of high-quality Chinese image-text pairs, complex document images, table data, daily and industrial scene photos, and multi-turn instruction-based multimodal dialogue. All data is processed with deduplication, noise filtering, and quality control.
Internal Chinese multimodal benchmark including VQA, document analysis, table extraction, and complex visual reasoning.
Image complexity, layout structure, text density, scene domain, multi-turn interaction depth.
[More Information Needed]
The model achieves strong performance in complex Chinese multimodal understanding and reasoning, suitable for enterprise-grade and advanced research visual-language tasks.
[More Information Needed]
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Vision-language architecture with powerful visual encoder and large language decoder, based on InternVL3-8B. Optimized for Chinese cross-modal alignment, complex visual reasoning, and structured document understanding.
NVIDIA GPU with CUDA and large VRAM support
BibTeX:
[More Information Needed]
APA:
[More Information Needed]
For updates, feedback, or usage questions, please refer to the model repository on the Hugging Face Hub.
Yougen Yuan
[More Information Needed]