Downloads · 30 days
0
introvoyz041/MiniVLA
MiniVLA is a image-text-to-text model from introvoyz041. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This repository hosts MiniVLA – a modular and deployment-friendly Vision-Language-Action (VLA) model designed for edge hardware (e.g., Jetson Orin Nano). It contains model checkpoints, Hugging Face–compatible Qwen-0.5…
Downloads · 30 days
0
Access
Public
Updated May 28, 2026
Repo size
12.9 GB
Likes
0
Public
Click a slice to open those files.
.pt5.6 GB · 43%
From the Hugging Face model README
This repository hosts MiniVLA – a modular and deployment-friendly Vision-Language-Action (VLA) model designed for edge hardware (e.g., Jetson Orin Nano).
It contains model checkpoints, Hugging Face–compatible Qwen-0.5B LLM, and ONNX/TensorRT exports for accelerated inference.
To enable low-latency, high-security desktop robot tasks on local devices, this project focuses on addressing the deployment and performance challenges of lightweight multimodal models on edge hardware. Using OpenVLA-Mini as a case study, we propose a hybrid acceleration pipeline designed to alleviate deployment bottlenecks on resource-constrained platforms.
We reproduced a lightweight VLA model and then significantly reduced its end-to-end latency and GPU memory usage by exporting the vision encoder into ONNX and TensorRT engines. While we observed a moderate drop in the task success rate (around 5-10% in LIBERO desktop operation tasks), our results still demonstrate the feasibility of achieving efficient, real-time VLA inference on the edge side.
The MiniVLA deployment is designed with modular microservices:
<p align="center"> <img src="./Results/System_Architecture.svg" width="100%" > </p>/vision/encode)/llm/generate)models/
Contains the original MiniVLA model checkpoints, based on
Stanford-ILIAD/minivla-vq-libero90-prismatic.
Special thanks to the Stanford ILIAD team for their open-source contribution.
qwen25-0_5b-trtllm/
Qwen-0.5B language model converted to TensorRT-LLM format.
qwen25-0_5b-with-extra-tokenizer/
Hugging Face–compatible Qwen-0.5B model with extended tokenizer.
tensorRT/vision_encoder_fp16.onnx
vision_encoder_fp16.engineFor full implementation and code, please visit the companion GitHub repository:
👉 https://github.com/Zhenxintao/MiniVLA
from transformers import AutoTokenizer, AutoModelForCausalLM
model_name = "xintaozhen/MiniVLA/qwen25-0_5b-with-extra-tokenizer"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_name, trust_remote_code=True)
import requests
url = "http://vision.svc:8000/vision/encode"
image_data = {"image": "base64_encoded_image"}
response = requests.post(url, json=image_data)
vision_embedding = response.json()
import requests
url = "http://llm.svc:8810/llm/generate"
payload = {"prompt": "Close the top drawer of the cabinet."}
response = requests.post(url, json=payload)
generated_actions = response.json()
/act), transforming offline benchmark code into a real-time deployable system.Target deployment: Jetson Orin Nano (16 GB / 8 GB variants).
For simulation and reproducibility, experiments were conducted on a local workstation equipped with:
⚠️ Note: Although the experiments were run on RTX 4060, the GPU memory (8 GB) is comparable to entry-level Jetson devices, making it a suitable proxy for evaluating edge deployment feasibility.
| Model Variant | Avg. GPU Utilization | Peak GPU Utilization |
|---|---|---|
| Original MiniVLA (PyTorch, no TRT) | ~67% | ~85% |
| MiniVLA w/ TensorRT Vision Acceleration | ~43% | ~65% |
Observation:
Original model:
GPU Memory-Usage: 4115MiB / 8188MiB
GPU-Util: 67% (peak 85%)
With TensorRT vision acceleration:
GPU Memory-Usage: 4055MiB / 8188MiB
GPU-Util: 43% (peak 65%)
Specify the license here (e.g., Apache 2.0, MIT, or same as MiniVLA / Qwen license).
If you use MiniVLA in your research or deployment, please cite:
@misc{MiniVLA2025,
title = {MiniVLA: A Modular Vision-Language-Action Model for Edge Deployment},
author = {Xintao Zhen},
year = {2025},
url = {https://huggingface.co/xintaozhen/MiniVLA}
}
We also acknowledge and thank the authors of Stanford-ILIAD/minivla-vq-libero90-prismatic, which serves as the base for the checkpoints included in this repository.