Downloads · 30 days
0
changh95/hamer-p150
hamer-p150 is a keypoint detection model from changh95. Use it for the keypoint detection task on the model card, and read the license before you ship it in a product. The card lists the license as other.
HaMeR (Hand Mesh Recovery, ViT-H/16 + MANO regression head) port on one Tenstorrent Blackhole p150a. Weights: changh95/hamer-weights · Paper: arXiv:2312.05251 · Upstream code: geopavlakos/hamer · Port: changh95/tt-hamer
Downloads · 30 days
0
Access
Public
Updated Oct 5, 2026
Repo size
8.4 GB
Likes
1
Public
Click a slice to open those files.
Other3.1 GB · 100%
From the Hugging Face model README
HaMeR (Hand Mesh Recovery, ViT-H/16 + MANO regression head) port on one Tenstorrent Blackhole p150a. Weights: changh95/hamer-weights · Paper: arXiv:2312.05251 · Upstream code: geopavlakos/hamer · Port: changh95/tt-hamer
Runs on p150 (mesh P150). Dispatch runs on Ethernet cores with 1 command queue, so all 12×10 Tensix cores compute.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Prerequisite: a Python environment with a built tt-metal / ttnn at tt-metal 8b98410e730 with patches/tt-metal-eth-dispatch.patch applied.
hf download changh95/hamer-p150 --exclude "image/*" --local-dir hamer-p150 && cd hamer-p150
pip install -e code/ # host dependencies only; "code/[server]" also installs the HTTP server
pip install -e "code/[server,test]" # optional: also the HTTP server and the tests (pytest, pyyaml)
from hamer import HamerModel # pip install -e code/ (in an environment with ttnn / tt-metal)
with HamerModel.from_pretrained(device_id=0) as model: # weights from the HF cache, opens the chip, warms up
hand = model("media/sample1.jpg", bbox=[20, 170, 295, 260], is_right=True)
print(hand.rotmats.shape, hand.betas.shape, hand.cam_t_full) # (16, 3, 3) (10,) [tx ty tz]
# several hands in one frame: one box and one is_right per hand
hands = model("media/sample1.jpg", bbox=[[20, 170, 295, 260], [130, 340, 320, 510]],
is_right=[True, False])
from_pretrained applies the published serve configuration and downloads the pinned weights changh95/hamer-weights (2.7 GB) to the HF cache.model.info shows dispatch, num_command_queues and grid. Boot takes approximately 10 s with a warm kernel cache.with block closes the model and the chip at the end. You can also call model.close().from_pretrained() also warms up all call variants. The first call is as fast as the next calls (1.03x for a JPEG file, was 1.34x). Start-up time increases by about 0.25 s.warmup_variants=None to skip the host warm-up. Use model.warmup(frame_size=(W, H)) to prepare frames larger than 1920x1080."media/sample1.jpg" is relative to the repository root.| name | meaning | |
|---|---|---|
| input | image | full frame: file path, PIL.Image, (H, W, 3) uint8 numpy array or torch tensor |
| input | bbox | hand box [x1, y1, x2, y2] in original pixels, or a list of boxes (you get a list) |
| option | is_right | True (default); False for a left hand; one value per box |
| option | rescale_factor | crop side = factor × longer box side; default 2.0 |
| option | return_vertices | include the 778 mesh vertices (MANO only); default True |
| output | rotmats | (16, 3, 3): index 0 = global orientation, 1..15 = MANO hand pose |
| output | betas, cam | (10,) MANO shape; (3,) crop camera [s, tx, ty] |
| output | cam_t_full | (3,) camera translation [tx, ty, tz] for the full frame |
| output | vertices, joints, keypoints_2d | (778, 3), (21, 3) metres; (21, 2) pixels; None without MANO_RIGHT.pkl |
| output | to_dict() | the same JSON as the HTTP /predict response |
Throughput helpers:
model.forward_crops(img) takes a batch of crops that you made, (B, 3, 256, 256) or (B, 3, 256, 192). It returns the upstream HAMER.forward keys.model.run_crop(crop) runs one device call on one (1, 3, 256, 192) network input. The "model call" number below measures this call.open_device() and from_pretrained(device=...) let you share one open chip with other code.To use the MANO mesh, give mano_path= with your own MANO_RIGHT.pkl. PYTHON.md gives the full API reference. examples/quickstart.py runs the snippet on both demo hands.
tt-model pull changh95/hamer-p150 --with-weights
tt-model serve changh95/hamer-p150 # or, with tt-cli: tt serve changh95/hamer-p150
printf '{"image":"%s","bbox":[20,170,295,260],"is_right":true}' "$(base64 -w0 media/sample1.jpg)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/hamer-p150 # tt-cli
Application startup complete.POST /predict: image (base64 PNG/JPEG of the full frame) and bbox [x1, y1, x2, y2] (sample boxes: media/bboxes.json). Optional: is_right (true), rescale_factor (2.0), return_vertices (true), return_faces (false).GET /health, GET /info.{"image_size": [334, 512], "bbox": [20.0, 170.0, 295.0, 260.0], "is_right": true,
"rotmats": [[[0.957649, 0.103436, 0.268718], "..."], "..."], "betas": [-0.257281, -0.099984, -0.43662, "..."],
"cam": [4.691406, -0.051758, 0.003906], "cam_t_full": [-0.059121, -0.027873, 7.751116],
"mano_available": false, "trace_active": true,
"timing_ms": {"preprocess": 1.19, "device": 3.55, "mano": 0.0, "total": 4.85}}
timing_ms values in the example are the p50 values of the current code/ server (ETH dispatch, 1 command queue, 12×10 grid). The container image is older than code/ (see Provenance). It uses stock worker dispatch, which gives an 11×10 grid on a p150, and it shows timing_ms.device of approximately 13.4 ms until the image is rebuilt.Device dispatch=eth command_queues=1 compute_grid=12x10. GET /info gives the same three values.cam_t_full uses focal_length = 5000/256*max(W, H). With MANO loaded, the response also has joints, vertices and keypoints_2d, and faces [1538][3] on request.Input (media/sample1.jpg, InterHand2.6M) | Mesh overlay: input · torch CPU reference · tt-nn on p150a (media/sample1_result.png) |
|---|---|
![]() | ![]() |
Warm, batch 1, one hand crop (256×192) per call, traced fused graph, optimized serve env. scripts/bench.py --iters 100 (3 runs) and scripts/serve_gate.sh (300 HTTP requests per run).
| Metric | Performance |
|---|---|
Model call TtHamer.__call__ (crop prep + H2D + trace + readback + host finalize), fresh crop each call | 3.49 ms median (3.483–3.498, 3 runs) · 3.460 ms min |
| Device trace (ViT-H/16 backbone + MANO head, back-to-back replay) | 3.357 ms mean (3.356–3.357, 3 runs) |
| Host prep · H2D · trace + readback · host finalize | 0.031–0.039 · 0.024–0.027 · 3.374–3.377 · 0.047–0.052 ms |
Served /predict, timing_ms.device p50 | 3.55 ms (another run on a busy host: 3.63 ms) |
Served /predict, timing_ms.total p50 (host preprocess p50 1.19 ms) | 4.85 ms (another run on a busy host: 5.73 ms) |
Python model(frame, bbox) call (host crop + resize + model call + post-processing, no MANO), chip 9 | 4.13–4.14 ms median (2 runs) · 3.980 ms min |
The measurement configuration is the p150 configuration: a Blackhole chip with a 12×10 compute grid, dispatch on ETH cores and 1 command queue. No number on this card uses worker dispatch or a second command queue. Accuracy against the torch CPU fp32 reference (scripts/acc.py, 32 real hand crops, served uint8 input): regression-vector PCC mean 0.999993, min 0.999974, 0 crops below the 0.9999 gate; served warm-up gate 0.99970 eager · 1.00000 replay · 0.99985 fresh input (gate 0.99). Details: VERIFICATION_2026-10-03.md.
The Python API row was measured on a different chip (chip 9) on 2026-10-04. On chip 9, model.run_crop is 3.61 ms median and the bare TtHamer call is 3.60 ms. Thus the API adds at most 0.01 ms. Chip 9 is approximately 0.1 ms slower than chip 5, which gave the other rows. hand.to_dict() is equal to the /predict JSON for all 16 test cases. The API does not change the device graph or the numerics.
Re-check 2026-10-05 (p150 ETH-dispatch compliance, commit aaf8601): ETH dispatch, 1 command queue and the 12×10 grid are now the defaults on all paths (Python API, server, bare TtHamer, scripts/bench.py). An independent verifier recorded these values at device open. On chips 3 and 4, the device trace is 3.400–3.425 ms mean, the model call is 3.61–3.65 ms median and the served timing_ms.device p50 is 3.65–3.69 ms. A same-chip A/B against the previous commit shows no change (chip 3: 3.420 ms before, 3.420–3.425 ms after). The difference from the chip-5 rows above is a chip-to-chip difference. Outputs are bit-identical and the accuracy values are unchanged. Details: VERIFICATION_2026-10-03.md, section "p150 ETH-dispatch compliance (2026-10-05)".
RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in PyTorch 2.11, batch 1, on the sample1 hand. The "incl. H2D/D2H" column compares with our model call (3.49 ms). The "forward only" column compares with our device trace (3.36 ms). Full table: GPU_COMPARISON.md.
| RTX 5090 precision | GPU incl. H2D/D2H | vs current build (3.49 ms) | GPU forward only vs ours (3.36 ms) |
|---|---|---|---|
| fp32 strict, eager | 13.09 ms | Blackhole 3.75× faster | 12.93 ms: Blackhole 3.85× faster |
| tf32, eager | 7.42 ms | Blackhole 2.12× faster | 7.34 ms: Blackhole 2.18× faster |
| bf16 autocast, eager | 7.89 ms | Blackhole 2.26× faster | 7.60 ms: Blackhole 2.26× faster |
| fp16 autocast, eager | 8.62 ms | Blackhole 2.47× faster | 8.52 ms: Blackhole 2.54× faster |
| bf16 weights, eager | 5.60 ms | Blackhole 1.61× faster | 5.86 ms: Blackhole 1.74× faster |
bf16 autocast + torch.compile (CUDA graphs) | 5.53 ms | Blackhole 1.58× faster | 5.17 ms: Blackhole 1.54× faster |
bf16 weights + torch.compile (CUDA graphs) | 3.62 ms | Blackhole 1.04× faster | 3.39 ms: parity (1.01×) |
The device graph is one metal trace of custom fused kernels on all 120 compute cores. These kernels include the ViT-H matmuls with fused add+LN and GELU, and MANO-head GEMVs with weight prefetch. The model call is faster than the best GPU variant because the host path is small: a 295 KB uint8 upload and persistent host buffers. The previous release (13.5 ms device, stock worker dispatch, 11×10 grid) was 3.73× slower than the best GPU variant.
patches/tt-metal-eth-dispatch.patch). Thus, this build assumes that you do not need chip-to-chip ethernet communication.HAMER_DISPATCH=worker (Python: dispatch="worker") is an opt-in that is equivalent only on a Galaxy Blackhole chip. On a p150 it gives an 11×10 grid. The default auto uses this mode only if ETH dispatch does not open, and then gives a RuntimeWarning. No number on this card uses it.forward_crops run one device call for each box. Without bbox, the model uses a centre square and the camera is not correct.kernels/hattn2). OPT_REPORT.md lists each change and its effect.TT_FUSED=1, default) is what the numbers above measure; TT_FUSED=0 restores the pre-fusion path. The MANO joint metric (scripts/mpjpe.py) was not run, because the measurement host has no MANO_RIGHT.pkl.MANO_RIGHT.pkl is not redistributed: copy your own to ~/.cache/tt-model/hamer-p150/weights/ (or give mano_path= in Python) for mesh, joints and 2-D keypoints; otherwise the server returns MANO parameters only (mano_available: false).code/hamer/native/resize.c). Without a C compiler, the server uses the bit-identical PIL path, which is about 0.8 ms slower.GET /v1/models is a stub so the tt-model ready card does not 404.8b98410e730 (v0.78.0-dev20260820) with patches/tt-metal-eth-dispatch.patch, ETH dispatch, 1 command queue and the 12×10 compute grid (2026-10-03; re-check 2026-10-05).code/): from changh95/tt-hamer, published under the upstream terms (mit-code-mano-non-commercial).These are the exact sources the container image was built from. code/ has since been updated (2026-10-03 optimized build, see OPT_REPORT.md; 2026-10-04 Python API, see PYTHON.md; 2026-10-05 ETH dispatch + 1 command queue + 12×10 as the default on all paths) and is newer than the image. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:
| component | built from |
|---|---|
| tt-metal | 8b98410e730bb504fea43a88609756e34821d91d |
code/ digest (image) | 4156e356e27b902c (sha256, first 16 hex digits; the current code/ differs) |
| built | 2026-09-13T15:34:00+00:00 by tt-model 0.1.0 |