Downloads · 30 days
0
episod/tt-waypoint
tt-waypoint is a image-to-video model from episod. Use it for the image-to-video task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Waypoint-1.5-1B is Overworld's interactive "world model". It is an autoregressive causal diffusion transformer that generates video one frame at a time, steered by mouse and scroll input.
Downloads · 30 days
0
Access
Public
Updated Sep 28, 2026
Repo size
1.5 GB
Likes
0
Public
Click a slice to open those files.
.whl35.3 KB · 49%
From the Hugging Face model README
Waypoint-1.5-1B is Overworld's interactive "world model". It is an autoregressive causal diffusion transformer that generates video one frame at a time, steered by mouse and scroll input.
This package runs it on Tenstorrent Blackhole via TTNN:
models.tt_dit covers this architecture.P150, a 1x1 mesh). All hardware verification so far ran on one chip of a P300c board.Packaged and published with tt-model-manager as a v6 thin bundle (manifest schema 6). It is a pip/venv install, not a container image.
| Architecture | 24-layer autoregressive causal diffusion transformer (per-layer ring-buffer frame cache, 512 tokens per frame), plus a small CNN VAE (ChunkedStreamingTAEHV) |
| Hardware | 1 Blackhole chip (profile p150, mesh P150, 1x1). Verified on one chip of a P300c; never run on a physical P150 |
| Input / output | Seed image: one, resized to 512x1024. Output: one generated 512x1024 frame per step, 4 denoising steps per frame. No text prompt |
| License | Apache-2.0 (weights and port code) |
| Status | Experimental: correct end to end per step and per component, not performance-tuned |
| Model CI v0 | not yet run |
Direct use: interactive demos and research into world-model inference on Tenstorrent hardware. Seed a session with an image, then step it forward one frame at a time with directional (mouse) and zoom (scroll) controls.
Out-of-scope use:
prompt_conditioning: null.uv tool install tenstorrent # once: the Tenstorrent CLI, `tt`
tt model pull episod/tt-waypoint
tt serve episod/tt-waypoint
tt model pull downloads two things:
install.sh/run.sh plus two wheels; ttnn, torch and the base HTTP stack come from an index at install time.Overworld/Waypoint-1.5-1B weights: revision 391f928, into the bundle's own HF cache. Weights come down by default, and tt has no --with-weights flag. The pull fetches the whole upstream repo, about 11.4 GB. The server reads only three files from it.serve starts the model's own HTTP server on port 20000, or the next free port if 20000 is busy. It is ready when it logs Application startup complete.
Measured on a fresh install of this bundle (P300c QuietBox, 2026-09-27):
| step | time |
|---|---|
pull --with-weights (venv plus weights) | 4 min 40 s |
| Launch to ready | 7 s |
First POST /v1/sessions | about 39 s, which includes first-time kernel compile |
| Each later step | about 25 s |
Without tt-cli, tt-model alone does the whole job:
tt-model pull episod/tt-waypoint --with-weights
tt-model serve episod/tt-waypoint
| profile | hardware | mesh | sessions per process |
|---|---|---|---|
p150 | 1 Blackhole chip. Verified on one chip of a P300c; a physical P150 has not been tested | P150 (1x1) | 1 |
No multi-chip profile exists. SUPPORTED_MESH_SHAPES is {(1, 1)}, and the server refuses any other WAYPOINT_MESH_SHAPE.
This is not an OpenAI-compatible API. /v1/models exists only as a stub for tooling. The real interface is a stateful session: seed once from an image, then step.
# Seed a session from a real starting image (base64-encoded PNG/JPEG; resized to 512x1024):
curl -s localhost:20000/v1/sessions -H 'Content-Type: application/json' \
-d "{\"image_b64\": \"$(base64 -w0 my_photo.png)\"}"
# -> {"session_id": "...", "frame_index": 1, "frame_b64": "<base64 PNG>"}
# Step the session forward one frame, steered by direction/zoom:
curl -s localhost:20000/v1/sessions/<session_id>/step \
-H 'Content-Type: application/json' -d '{"direction": "forward", "zoom": 0.0}'
# -> {"session_id": "...", "frame_index": 2, "frame_b64": "<base64 PNG>"}
# End the session:
curl -s -X DELETE localhost:20000/v1/sessions/<session_id>
| endpoint | purpose |
|---|---|
POST /v1/sessions | Takes {"image_b64": str}. Seeds a new session and returns the decoded seed frame. Discards any session already in progress |
POST /v1/sessions/{id}/step | Takes {"direction": str = "forward", "zoom": float = 0.0}. Generates one frame |
DELETE /v1/sessions/{id} | Ends the session |
GET /health, GET /tt-liveness | Readiness and liveness. Both return 503 until the model is loaded |
GET /v1/models | Stub listing, for tooling |
direction: one of forward, back, left, right, forward_left, forward_right, back_left, back_right. These map to the model's [dx, dy] mouse input.zoom: the scroll scalar. Only its sign has a documented meaning.503: the model is still starting.404: the session ID is unknown, or a newer POST /v1/sessions superseded it.409: you stepped with no active session.400: the request was bad.The GitHub repo also has a local Gradio UI (app.py), which is a pure HTTP client for this same server.
| metric | value | source |
|---|---|---|
| Per-step denoiser correctness vs HF reference (8 real denoising steps, correlation) | about 0.959 on average | BRINGUP_LOG.md (fp32-accumulation row; step-trace row). The range before the fp32-accumulation fix was 0.897-0.971 |
| 24-layer transformer, single forward pass (correlation vs reference) | 0.963 | BRINGUP_LOG.md (fp32-accumulation row) |
| Single decoder layer PCC (layer 0, frame-0 prefill) | 0.999 | BRINGUP_LOG.md (ttm-* skills row) |
| VAE encoder / decoder (correlation vs reference) | 0.9996 / > 0.99 on all 12 frames | BRINGUP_LOG.md (Stage 5 rows) |
| Generated frames from a real seed image | Visually coherent and plausible. Inspected by eye; no numeric perceptual metric | bring-up log; re-checked on both bundles below |
| Warm latency per generated frame, installed bundle (4 denoise steps + 1 cache commit) | median 24.95 s (min 24.64, max 25.35), N=15 | published bundle 1dd34b8, 2026-09-27 |
| Same, re-verified on this revision of the bundle | median 25.72 s (min 25.70, max 25.91), N=3 | this bundle, 2026-09-27. +3% vs the row above |
| Effective throughput | 0.040 FPS | 1 / median |
| Latency vs session length (steps 1 / 5 / 10 / 15 / 16) | 25.44 / 25.05 / 24.79 / 24.95 / 24.64 s | flat, no drift over 16 frames |
| First step after a cold start | 25.44 s. The first POST /v1/sessions on a cold process takes 38-39 s | includes kernel compile |
Dev build, for comparison (tt-metal v0.78.0 from source, in-process benchmark.py) | 27.41 s per frame | PORT_PLAN.md |
| Upstream GPU reference, for scale | 56 FPS on an RTX 5090, 4-step unquantized | upstream model card |
Methodology: on one Blackhole chip of a P300c (1x1 mesh, ttnn==0.78.0 from PyPI, eager execution, batch 1, one session), a session seeded from ref_activations/seed_frame_ref.png was stepped with the default controls (forward, zoom 0), timing host HTTP wall-clock per POST .../step after one untimed warmup step: N=15 on the published bundle 1dd34b8 and N=3 on this bundle revision (weights 391f928), both on 2026-09-27.
How to read the accuracy numbers:
The full investigation is in BRINGUP_LOG.md.
404.p150 profile states the requirement (one chip); it has not been tested on a physical P150 board.zoom has sign-only semantics.test_full_model.py assert is left failing on purpose rather than weakened.transformer/config.json, transformer/diffusion_pytorch_model.safetensors and vae/diffusion_pytorch_model.safetensors, at the pinned revision. The bundle's weight pull has no file filter, so pull --with-weights also downloads about 4 GB the server never reads: the root model.safetensors and assets/.forward session from the reference seed:
fp32_dest_acc_en, HiFi4). Outputs are plausible but will not match GPU output pixel for pixel.Overworld/Waypoint-1.5-1B, Apache-2.0.waypoint_ttnn wheel, Apache-2.0 (repo LICENSE).tt_waypoint_models_closure, which is models/common/lightweightmodule.py from tt-metal v0.78.0. Apache-2.0.| date | change |
|---|---|
| 2026-09-27 | Repackaged with waypoint-ttnn 0.1.1. The server now loads weights at the pinned revision 391f928, overridable with TT_MODEL_WEIGHTS_REVISION, and fetches only the three files it reads. The manifest records that revision and tt_metal_version 0.78.0. run.sh exports the revision and picks chips from the bundle's device count. ttnn==0.78.0 is unchanged. Re-verified on hardware from a fresh install with an empty HF cache: 25.72 s per frame, coherent frames. |
| 2026-09-15 | Repackaged as a v6 thin bundle (pip/venv), replacing the v5.1 container image. The served code is unchanged. ttnn==0.78.0 comes from PyPI, and torch==2.14.0+cpu is pinned. Re-verified on hardware: mesh open, fresh weight fetch, seed and step, coherent frames. |
| 2026-09-10 | Benchmarked: 27.41 s per frame, flat across session length. Re-verified by a fresh pull and serve of the published package across several directions. |
| 2026-09-09 | First publish (v5.1 container). The card's port was corrected to 20000. session.py now resolves weights via snapshot_download instead of a host-only path. |
None. This is the only Tenstorrent port of Waypoint-1.5-1B.
https://huggingface.co/episod/tt-waypoint/discussions. That is the one channel that reaches the bundle's author.tt tooling itself: run tt report issue. It collects your environment and opens a prefilled issue against tenstorrent/tt-cli; it does not reach this package's author.The served-path code is split into two small wheels built for this bundle, both in wheels/:
| wheel | what it is |
|---|---|
waypoint_ttnn-0.1.1 | The served-path closure of tsingletaryTT/tt-waypoint: session.py, the ASGI app, and the tt/ model classes. Bring-up scripts, tests and docs are not included; see the GitHub repo for those |
tt_waypoint_models_closure-0.78.0 | The one tt-metal dependency this model has: models/common/lightweightmodule.py, pinned to tt-metal v0.78.0 |
| component | built from |
|---|---|
| tt-metal / ttnn | ttnn==0.78.0 (PyPI). The manifest tt_metal_version is 0.78.0 |
| model code | tsingletaryTT/tt-waypoint @ adc1b56. The wheel's .py files are identical to this commit |
| weights | Overworld/Waypoint-1.5-1B @ 391f92827075edcf4a8b3c8a2ddae010698f8636, pinned in the manifest and in session.py |
| Model CI v0 | not yet run |
| build | 2026-09-27 · tt-model-manager (producer tt_kernel_version 0.1.0) |