Downloads · 30 days
1.7K
25% of all-time downloads
jlee-larr/dynaflip-base
dynaflip-base is a zero-shot image classification model from jlee-larr. Use it for the zero-shot image classification task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This model was proposed in DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation.
Downloads · 30 days
1.7K
25% of all-time downloads
All-time downloads
6.7K
Public
Parameters
233M
932 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors931 MB · 100%
From the Hugging Face model README
This model was proposed in DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation.
The model is compatible with Transformers:
from transformers import AutoModel, AutoProcessor
from PIL import Image
import torch
REPO = "jlee-larr/dynaflip-base"
dynaflip = AutoModel.from_pretrained(REPO, trust_remote_code=True).eval()
processor = AutoProcessor.from_pretrained(REPO, trust_remote_code=True)
image = Image.open("example.png").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
v = dynaflip.vision_outputs(inputs["pixel_values"])
# v.last_hidden_state -> (B, num_patches, 768) patch tokens
# v.pooler_output -> (B, 1536) CLS + mean(patches)
@article{lee2026dynaflip,
title = {DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation},
author = {Lee, Jusuk and Lee, Seungjae and Shin, Jonghun and Jung, Hoseong and Kim, Sungha and Cho, Daesol and Kim, H. Jin and Huang, Jia-Bin and Huang, Furong},
journal = {arXiv preprint arXiv:2605.30350},
year = {2026},
}