Downloads · 30 days
6
12% of all-time downloads
Yougen/InternVL3Fangwusha14B
InternVL3Fangwusha14B is a machine learning model from Yougen. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
InternVL3Fangwusha14B is a 14B-parameter vision-language model (VLM) fine-tuned from InternVL3-14B, dedicated to high-performance Chinese multimodal understanding, deep visual reasoning, complex document analysis, tab…
Downloads · 30 days
6
12% of all-time downloads
All-time downloads
51
Public
Parameters
15.1B
30.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors30.2 GB · 100%
From the Hugging Face model README
InternVL3Fangwusha14B is a 14B-parameter vision-language model (VLM) fine-tuned from InternVL3-14B, dedicated to high-performance Chinese multimodal understanding, deep visual reasoning, complex document analysis, table structure parsing, and multi-turn interactive visual dialogue for enterprise and advanced research scenarios.
This model is a large-scale vision-language model built on the InternVL3-14B base architecture. It is fine-tuned to significantly improve cross-modal semantic alignment, fine-grained visual recognition, complex layout understanding, and professional scene multimodal reasoning in Chinese. It provides powerful generation and reasoning capabilities while maintaining relatively efficient inference.
This model can be directly used for:
Can be further fine-tuned for:
All outputs in professional or production scenarios should be reviewed by qualified personnel. It is strongly recommended to configure content security and privacy protection mechanisms for public deployment. Professional dedicated models are preferred for high-precision industrial or medical visual tasks. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model.
Use the code below to get started with the model.
from transformers import AutoModel, AutoTokenizer
model_name = "Yougen/InternVL3Fangwusha14B"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModel.from_pretrained(
model_name,
device_map="auto",
torch_dtype="auto",
trust_remote_code=True
).eval()
# Example usage:
# image = load_image("your_image.jpg")
# question = "请详细解析这张图片中的表格数据和内容"
# response = model.chat(tokenizer, image, question)
# print(response)
Training data includes high-quality Chinese image-text pairs, complex documents, tables, charts, professional scene images, and multi-turn instruction-based multimodal dialogue. Data has been strictly processed with deduplication, noise filtering, and quality control.
Internal Chinese multimodal evaluation set covering VQA, document analysis, table extraction, chart understanding and complex visual reasoning.
Image complexity, layout density, text definition, domain professionalism, multi-turn dialogue depth.
[More Information Needed]
The model delivers strong performance in complex Chinese multimodal understanding and reasoning, suitable for high-demand enterprise and advanced research visual-language tasks.
[More Information Needed]
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Vision-language architecture with high-capacity visual encoder and large language decoder, based on InternVL3-14B. Optimized for Chinese cross-modal alignment, fine-grained visual understanding, and complex document reasoning.
NVIDIA high-performance GPU cluster with large VRAM
BibTeX:
[More Information Needed]
APA:
[More Information Needed]
For updates and issues, please visit the model repository on Hugging Face Hub.
Yougen Yuan
[More Information Needed]