Downloads · 30 days
0
variante/llara-maskrcnn
llara-maskrcnn is a object detection model from variante. Use it when you need objects located in an image. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Jul 1, 2024
Repo size
533 MB
Likes
1
Public
Click a slice to open those files.
.pth533 MB · 100%
From the Hugging Face model README
This model is released with paper LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
Xiang Li<sup>1</sup>, Cristina Mata<sup>1</sup>, Jongwoo Park<sup>1</sup>, Kumara Kahatapitiya<sup>1</sup>, Yoo Sung Jang<sup>1</sup>, Jinghuan Shang<sup>1</sup>, Kanchana Ranasinghe<sup>1</sup>, Ryan Burgert<sup>1</sup>, Mu Cai<sup>2</sup>, Yong Jae Lee<sup>2</sup>, and Michael S. Ryoo<sup>1</sup>
<sup>1</sup>Stony Brook University <sup>2</sup>University of Wisconsin-Madison
Model type: This repository contains three models trained on three subsets respectively, converted from VIMA-Data. For the conversion code, please refer to convert_vima.ipynb
Paper or resources for more information: https://github.com/LostXine/LLaRA
Where to send questions or comments about the model: https://github.com/LostXine/LLaRA/issues