Skip to content

TRI-ML

prismatic-vlms

TRI-ML/prismatic-vlms

prismatic-vlms is a image-to-text model from TRI-ML. Use it when you need a caption or text from an image. The card lists the license as mit.

All models trained as part of the paper Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models by Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang, Thomas Kollar, and D…

Downloads · 30 days

0

Access

Public

Updated May 6, 2024

Repo size

785 GB

Likes

29

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.pt785 GB · 100%

At a glance

Task
Image-to-Text
License
mit
Access
Public
Created
Feb 13, 2024
Updated
May 6, 2024
SHA
a3ba8a19
Task
Image-to-Text
License
mit
Languages
en
Created
Feb 13, 2024
Updated
May 6, 2024