Downloads · 30 days
108
1% of all-time downloads
FreedomIntelligence/HuatuoGPT-Vision-34B
HuatuoGPT-Vision-34B is a image-text-to-text model from FreedomIntelligence. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for huatuogpt_vision. The card lists the license as apache-2.0.
<div align="center" <h1 HuatuoGPT-Vision-34B </h1 </div
Downloads · 30 days
108
1% of all-time downloads
All-time downloads
19.9K
Public
Parameters
34.8B
71.2 GB on disk
Likes
24
Public
Click a slice to open those files.
.safetensors69.5 GB · 98%
From the Hugging Face model README
HuatuoGPT-Vision is a multimodal LLM for medical applications, built with the PubMedVision dataset. HuatuoGPT-Vision-34B is trained based on Yi-1.5-34B using the LLaVA-v1.5 architecture.
git clone https://github.com/FreedomIntelligence/HuatuoGPT-Vision.git
query = 'What does the picture show?'
image_paths = ['image_path1']
from cli import HuatuoChatbot
bot = HuatuoChatbot(huatuogpt_vision_model_path) # loads the model
output = bot.inference(query, image_paths) # generates
print(output) # Prints the model output
@misc{chen2024huatuogptvisioninjectingmedicalvisual,
title={HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale},
author={Junying Chen and Ruyi Ouyang and Anningzhe Gao and Shunian Chen and Guiming Hardy Chen and Xidong Wang and Ruifei Zhang and Zhenyang Cai and Ke Ji and Guangjun Yu and Xiang Wan and Benyou Wang},
year={2024},
eprint={2406.19280},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2406.19280},
}