Downloads · 30 days
25
11% of all-time downloads
wooj0216/ReVIOSa-1B
ReVIOSa-1B is a machine learning model from wooj0216. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
[\[📂 GitHub\]](https://github.com/cvlab-kaist/InterRVOS) [\[📜 Paper\]](https://arxiv.org/pdf/2506.02356) [\[🚀 Quick Start\]](quick-start)
Downloads · 30 days
25
11% of all-time downloads
All-time downloads
219
Public
Parameters
1.2B
4 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors4 GB · 100%
How the weights are stored.
F32860M · 74%
From the Hugging Face model README
[📂 GitHub] [📜 Paper] [🚀 Quick Start]
This repository contains the model ReVIOSa-1B for the paper InterRVOS: Interaction-aware Referring Video Object Segmentation.
We introduce Interaction-aware Referring Video Object Segmentation (InterRVOS), a novel task that focuses on the modeling of interactions. It requires the model to segment the <b>actor</b> and <b>target</b> objects separately, reflecting their asymmetric roles in an interaction. Please refer to the project page for detailed visualization results.
We provide an example code to run ReVIOSa using transformers.
import torch
from transformers import AutoTokenizer, AutoModel
from PIL import Image
import numpy as np
import os
# load the model and tokenizer
path = "wooj0216/ReVIOSa-1B"
model = AutoModel.from_pretrained(
path,
torch_dtype=torch.bfloat16,
low_cpu_mem_usage=True,
use_flash_attn=True,
trust_remote_code=True).eval().cuda()
tokenizer = AutoTokenizer.from_pretrained(path, trust_remote_code=True, use_fast=False)
video_folder = "/PATH/TO/VIDEO_FOLDER"
images_paths = os.listdir(video_folder)
images_paths = [os.path.join(video_folder, image_path) for image_name in images_paths]
text_prompts = "<image>Please segment the child reaching out to man."
input_dict = {
'video': images_paths,
'text': text_prompts,
'past_text': '',
'mask_prompts': None,
'tokenizer': tokenizer,
}
return_dict = model.predict_forward(**input_dict)
answer = return_dict["prediction"]
masks = return_dict['prediction_masks']
This project is based on Sa2VA. Many thanks to the authors for their great works!
If you find this repository useful, please consider referring to the following paper:
@misc{jin2025interrvosinteractionawarereferringvideo,
title={InterRVOS: Interaction-aware Referring Video Object Segmentation},
author={Woojeong Jin and Seongchan Kim and Jaeho Lee and Seungryong Kim},
year={2025},
eprint={2506.02356},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2506.02356},
}