Downloads · 30 days
560
0% of all-time downloads
FreedomIntelligence/HuatuoGPT-Vision-7B
HuatuoGPT-Vision-7B is a text generation model from FreedomIntelligence. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
<div align="center" <h1 HuatuoGPT-Vision-7B </h1 </div
Downloads · 30 days
560
0% of all-time downloads
All-time downloads
402K
Public
Parameters
7.9B
17.6 GB on disk
Likes
30
Public
Click a slice to open those files.
.safetensors15.9 GB · 90%
From the Hugging Face model README
HuatuoGPT-Vision is a multimodal LLM for medical applications, built with the PubMedVision dataset. HuatuoGPT-Vision-7B is trained based on Qwen2-7B using the LLaVA-v1.5 architecture.
git clone https://github.com/FreedomIntelligence/HuatuoGPT-Vision.git
query = 'What does the picture show?'
image_paths = ['image_path1']
from cli import HuatuoChatbot
bot = HuatuoChatbot(huatuogpt_vision_model_path) # loads the model
output = bot.inference(query, image_paths) # generates
print(output) # Prints the model output
@misc{chen2024huatuogptvisioninjectingmedicalvisual,
title={HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale},
author={Junying Chen and Ruyi Ouyang and Anningzhe Gao and Shunian Chen and Guiming Hardy Chen and Xidong Wang and Ruifei Zhang and Zhenyang Cai and Ke Ji and Guangjun Yu and Xiang Wan and Benyou Wang},
year={2024},
eprint={2406.19280},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2406.19280},
}