Skip to content

shixuanleong

visualheist-large

shixuanleong/visualheist-large

visualheist-large is a image-text-to-text model from shixuanleong. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as mit.

VisualHeist is an object detection model finetuned to extract tables and figures from PDFs. VisualHeist has two versions: - visualheist-base[[HF]](https://huggingface.co/shixuanleong/visualheist-base) (0.23B) - visual…

Downloads · 30 days

40

1% of all-time downloads

All-time downloads

4.5K

Public

Parameters

823M

4.8 GB on disk

Likes

1

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors3.3 GB · 68%

At a glance

Task
Image-Text-to-Text
License
mit
Model type
florence2
Access
Public
Created
Oct 28, 2024
Updated
Mar 5, 2025
SHA
6abe855b

Try a prompt

Task
Image-Text-to-Text
Type
florence2
License
mit
Created
Oct 28, 2024
Updated
Mar 5, 2025