Downloads · 30 days
0
Afeng-x/SPHINX-V-Model
SPHINX-V-Model is a machine learning model from Afeng-x. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
SPHINX-V is a new multimodal large language model designed for visual prompting, equipped with a novel visual prompt encoder and a two-stage training strategy. SPHINX-V supports multiple visual prompts simultaneously…
Downloads · 30 days
0
Access
Public
Updated Dec 1, 2025
Repo size
79.7 GB
Likes
3
Public
Click a slice to open those files.
.pth79.7 GB · 100%
From the Hugging Face model README
SPHINX-V is a new multimodal large language model designed for visual prompting, equipped with a novel visual prompt encoder and a two-stage training strategy. SPHINX-V supports multiple visual prompts simultaneously across various types, significantly enhancing user flexibility and achieve a fine-grained and open-world understanding of visual prompts.
Project Page: Draw-and-Understand
Paper: https://arxiv.org/abs/2403.20271
Code: https://github.com/AFeng-x/Draw-and-Understand
Dataset: MDVP-Data & MDVP-Bench
Primary intended uses: The principal application of SPHINX-V is centered around conducting research in the realm of visual prompting large multimodal models and chatbots.
Primary intended users: The model is primarily designed for use by researchers and enthusiasts specializing in fields such as computer vision, natural language processing, and interactive artificial intelligence.
Llama 2 is licensed under the LLAMA 2 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved.
@article{lin2024draw,
title={Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want},
author={Lin, Weifeng and Wei, Xinyu and An, Ruichuan and Gao, Peng and Zou, Bocheng and Luo, Yulin and Huang, Siyuan and Zhang, Shanghang and Li, Hongsheng},
journal={arXiv preprint arXiv:2403.20271},
year={2024}
}