Downloads · 30 days
0
episod/tt-openvla
tt-openvla is a robotics model from episod. Use it for the robotics task on the model card, and read the license before you ship it in a product. The card lists the license as other.
A TT-Metal / TTNN bring-up of OpenVLA-7B, a vision-language-action model for robot manipulation.
Downloads · 30 days
0
Access
Public
Updated Sep 28, 2026
Repo size
1.2 MB
Likes
0
Public
Click a slice to open those files.
.whl911 KB · 96%
From the Hugging Face model README
A TT-Metal / TTNN bring-up of OpenVLA-7B, a vision-language-action model for robot manipulation.
Packaged and published with tt-model-manager as a v6 thin bundle (manifest schema 6): a pip/venv install, not a container.
No checkpoint is hosted here. Weights are fetched at install or serve time from openvla/openvla-7b at a pinned revision, unconverted. The two vision towers, the projector and LLaMA-2-7B all load from that checkpoint's own tensors.
| Architecture | About 7B parameters, in this order: <br>1. DINOv2 ViT-L/14 (4 register tokens) + SigLIP ViT-So400M/14, at 224 px <br>2. 3-layer GELU projector (2176 → LLaMA dim) <br>3. 32-layer LLaMA-2-7B, reusing tt-metal tt_transformers attention, RoPE and KV cache |
| Hardware | p300: 2 Blackhole chips, tensor-parallel 1×2 mesh. One chip overflows L1 in decode |
| Input / output | In: one RGB image (224 × 224 after preprocessing) and one instruction. Out: one 7-DoF action, from 7 generated tokens (1 prefill + 6 decode) |
| License | Port code MIT · vendored tt-metal code Apache-2.0 · weights MIT per upstream. The LLaMA backbone is a Llama 2 derivative, so the Llama 2 Community License and AUP apply. See Licensing |
| Status | Experimental |
| Model CI v0 | not yet run |
Direct use: research and bring-up evaluation of OpenVLA inference on Tenstorrent hardware:
POST /act, for comparison against GPU OpenVLA.Out-of-scope use:
uv tool install tenstorrent # once: the Tenstorrent CLI, `tt`
tt model pull episod/tt-openvla
tt serve episod/tt-openvla
tt model pull installs the v6 thin bundle: a hermetic uv venv with a pinned Python 3.12, ttnn==0.77.0, and the two bundled wheels.
openvla/openvla-7b at revision 47a0ec7 (about 15 GB) into the bundle's own HF cache (<install>/.hf). Weights come down by default; tt has no --with-weights flag.serve starts uvicorn on port 20000, or the next free port. The Gradio UI is at / and the REST endpoint at /act.Application startup complete. means the backend is ready.Measured on a fresh install (one P300c board, 2026-09-27):
| step | time |
|---|---|
| Launch to ready, weights already pulled | 55 s, which includes building the on-disk tensor cache |
| Launch to ready, weights not pulled | 6.5 min: the server downloads the 15 GB itself, at the same pinned revision |
First /act on a fresh install | 55 s, most of it the one-time kernel compile |
| Later warm calls | about 0.33 s |
No other model repos are fetched, apart from a small NousResearch LLaMA-2 config/tokenizer that tt-metal's tt_transformers reads.
Without tt-cli, tt-model alone does the whole job:
tt-model pull episod/tt-openvla --with-weights
tt-model serve episod/tt-openvla
| profile | hardware | mesh |
|---|---|---|
| p300 | P300 (2 Blackhole chips) | 1×2, MESH_DEVICE=P300, fabric 1D |
One prediction runs at a time. The Gradio UI is limited to concurrency_limit=1. /act requests and UI requests share one lock around the backend, so concurrent requests queue rather than run two decode loops on the mesh.
Not an OpenAI-compatible API. A continuous action vector isn't natural-language output, so there is no /v1/chat/completions.
/)bridge_orig).The UI shows the decoded 7-DoF action, the raw generated action-token IDs, and the measured latency.
POST /actThe request is the one openvla/openvla's reference server (vla-scripts/deploy.py) takes:
{image, instruction, unnorm_key?}.{"encoded": ...}.Both build the same prompt as upstream: In: What action should the robot take to {instruction.lower()}?\nOut:, plus the empty token upstream's predict_action appends.
It differs from upstream in three ways:
{"action": [...]}, where upstream returns the bare array. Upstream clients that do action = resp.json() need resp.json()["action"].unnorm_key: bridge_orig. Upstream defaults to None, which errors for openvla-7b.instruction, returns HTTP 400 with a detail message. Upstream returns 200 with an "error" field.import json_numpy; json_numpy.patch() # pip install json-numpy
import requests, numpy as np
from PIL import Image
image = np.array(Image.open("my_photo.jpg").convert("RGB")) # uint8 HxWx3
resp = requests.post("http://localhost:20000/act", json={
"image": image,
"instruction": "pick up the remote control",
# "unnorm_key": "bridge_orig", # optional; any key in openvla-7b config.json norm_stats
})
action = resp.json()["action"] # 7 floats: dx, dy, dz, droll, dpitch, dyaw, gripper
gradio_app/app.py runs the same demo directly: on Blackhole by default (hold a 2-chip lease first), or on CPU with --backend reference.tt/hf_reference.py runs the upstream HF AutoModelForVision2Seq reference that the accuracy numbers below were measured against.| stage / metric | vs reference | result |
|---|---|---|
| DINOv2 featurizer (openvla-7b weights) | PCC vs the CPU reference with the same weights | 0.99979 (0.99978-0.99988 across 4 inputs on the served 1×2 mesh) |
| SigLIP fused featurizer (openvla-7b weights) | same | 0.99975 (0.99972-0.99980) |
| Vision backbone + projector | same | 0.99970 (0.99961-0.99976) |
| Prefill logits | PCC vs HF AutoModelForVision2Seq (fp32, rev 47a0ec7) | 0.994-0.995 on 4 inputs |
| 7 greedy action tokens, end to end | exact match vs HF AutoModelForVision2Seq + predict_action | 7/7 on 2 of 4 inputs; 5/7 and 1/7 on the other two |
| Decoded action | max |Δ| vs HF, per input | 0.0 / 0.0 / 0.0107 / 0.0451 |
| Task success (BridgeData V2 / LIBERO) | none | not measured |
/act latency, warm, installed bundle | none | median 334.2 ms (min 328.8, max 412.9), N=20; about 3.0 actions/s for one sequential client |
4 concurrent /act requests | none | all 4 returned HTTP 200 and queued, finishing within 1.35 s in total. The actions were identical to the sequential call |
First /act on a fresh install (kernel JIT) | none | 55 s, once per install |
Methodology: the latency and concurrency figures come from HTTP POST /act on one P300c board (2 Blackhole chips, 1×2 mesh, ttnn==0.77.0 from PyPI). The request was the bundled gradio_app/assets/example.jpg with "open the drawer" and unnorm_key=bridge_orig, json-numpy encoded, measured client-side. There was one untimed first call, then N=20 sequential warm calls and one 4-way concurrent burst. The run used this bundle revision (weights 47a0ec7) on 2026-09-27.
The accuracy rows come from a separate run on the same board against the upstream HF reference on CPU (timm 0.9.16, transformers 4.40.1, greedy). It used four image/instruction pairs: the repo's example image with "open the drawer", a house in a field with "move forward", a dog with "Pick up the ball", and a synthetic orange image with "move forward". The script is tt/hf_reference.py in the GitHub repo.
Why the remaining token mismatches look like precision, not a bug:
0.0.0.0 with no authentication./act is serialized. Throughput does not scale with concurrent clients: four simultaneous requests finish one after another.Failed to discover available ethernet links warning hundreds of times per /act call. It is harmless, but it makes the log grow fast.Physical harm. The model outputs end-effector deltas and a gripper command. If a robot executes them without independent limits, wrong actions can:
Nothing in this package bounds, filters, or sanity-checks actions. Keep a safety layer and an e-stop in the loop, and prefer simulation.
Silent numerical deviation. On some inputs this port produces a different action than the reference (see Limitations), and the output looks plausible either way. A wrong action and a right one are indistinguishable from the response alone.
Out-of-distribution inputs. For scenes, cameras, or embodiments outside the training mix, OpenVLA still returns a confident-looking action rather than refusing.
Unnormalization choice. Actions are unnormalized with per-dataset statistics (unnorm_key). The wrong key for your robot rescales every motion, and the server defaults silently to bridge_orig.
Network exposure. An unauthenticated /act on 0.0.0.0 lets anyone on the network issue actions to a connected control loop. Bind to localhost, or put it behind a proxy with auth.
Supply chain. The processor loads with trust_remote_code=True, which executes the upstream repo's Python at serve time. That code is pinned to revision 47a0ec7, the same revision as the weights.
Pass these risks on. The Llama 2 AUP (§4) requires disclosing known dangers of your AI system to end users. If you build on this package, pass these risks along.
| component | license | notes |
|---|---|---|
Port code (tt_openvla_serving wheel, GitHub repo) | MIT | matches openvla/openvla |
Vendored tt-metal models/common + models/tt_transformers (tt_openvla_models_closure wheel) | Apache-2.0 | upstream tenstorrent/tt-metal v0.77.0 |
openvla/openvla-7b weights | MIT (upstream declaration) | The LLaMA-2-7B backbone is a Meta Llama 2 derivative, so the Llama 2 Community License and Acceptable Use Policy also apply: the 700M-MAU clause, the redistribution notice and the prohibited uses. The card's license: other reflects those terms |
This repo hosts no weights. Downloading and running openvla-7b is between you and its licensors.
| date | change |
|---|---|
| 2026-09-27 | tt-openvla-serving 0.2.1 (0.2.0 plus two review fixes: a 400 for a non-string /act instruction, and a separate converted-weight cache per local checkpoint). See the list below this table. |
| 2026-09-15 | tt-openvla-serving 0.1.1: added the POST /act REST endpoint. The 0.1.0 wheel was removed |
| 2026-09-15 | First v6 thin bundle (tt-model package-thin), serving on the tt-dit-server kind alongside the existing card |
| 2026-09-10 | Model card published (code and weights pointers only) |
The 2026-09-27 release (tt-openvla-serving 0.2.1):
47a0ec7. This fixes the "expected 3 safetensors shards, found []" startup failure.predict_action appends is now added./act:
{"encoded": ...};tt_metal_version 0.77.0./act, 4/4 concurrent requests OK.None.
tt tooling itself: use tt report issue. It files against tenstorrent/tt-cli and doesn't reach this package's author.| component | built from |
|---|---|
| tt-metal | ttnn==0.77.0 (PyPI), plus a vendored models/ closure from tt-metal v0.77.0 source. The manifest tt_metal_version is 0.77.0 |
| runtime | kind tt-dit-server: gradio_app.asgi:app under uvicorn, with gradio 5.40.0, transformers 4.52.4, torch 2.14.0+cpu |
| weights | openvla/openvla-7b @ 47a0ec7fc4ec123775a391911046cf33cf9ed83f, pinned in the manifest and the code |
| port code | tsingletaryTT/tt-openvla @ b056100, built by scripts/build-serving-wheel.sh |
| wheels (sha256, first 16) | tt_openvla_serving-0.2.1 32b600f81db3e156 · tt_openvla_models_closure-0.77.0 47b658fb62c30ffc |
| Model CI v0 | not yet run |
| build | 2026-09-27 · tt-model-manager (producer tt_kernel_version 0.1.0) |