Downloads · 30 days
0
cgb/triposr-onnx-webgpu
triposr-onnx-webgpu is a image-to-3d model from cgb. Use it for the image-to-3d task on the model card, and read the license before you ship it in a product. It is set up for onnx. The card lists the license as mit.
TripoSR exported to ONNX so it turns a photo into a 3D shape entirely inside a web browser on WebGPU. No server, no API key, no upload: the weights are cached by the browser and the reconstruction happens on the visit…
Downloads · 30 days
0
Access
Public
Updated Sep 4, 2026
Repo size
485 MB
Likes
0
Public
Click a slice to open those files.
.data484 MB · 100%
From the Hugging Face model README
TripoSR exported to ONNX so it turns a photo into a 3D shape entirely inside a web browser on WebGPU. No server, no API key, no upload: the weights are cached by the browser and the reconstruction happens on the visitor's own GPU.
| File | Precision | Size |
|---|---|---|
triplane_q8.onnx + .data | int8 weight-only, block 32 | 485 MB |
decoder.onnx | float32 | 0.2 MB |
triplane.onnx takes an image and produces the triplane. decoder.onnx turns
sampled triplane features into density and colour, and is run per batch of
points while isosurfacing.
image [1,3,512,512] float32, RGB in 0..1, to
triplane [1,3,40,64,64]. The DINO mean and standard deviation are inside the
graph, so a caller does not have to know them.features [N,120] float32 to density [N,1], color [N,3].This is not cosmetic, and it is the easiest thing to get wrong. The subject must be cut out, framed to 0.85 of the width, and composited onto mid grey:
cut the subject out -> RGBA
crop to the alpha bounds, pad to a square
pad again so the subject is 0.85 of the frame
resize to 512 square
rgb = rgb * alpha + (1 - alpha) * 0.5
Composited onto white instead, the encoder reads the unbroken bright field as surface and the mesh comes back with a sheet of phantom geometry hanging behind the subject. Skipping the 0.85 framing gives a subject whose proportions are wrong, because the model reads scale off the picture.
The grid_sample step that turns a 3D point into 120 features is deliberately
not in either graph. ONNX Runtime's WebGPU backend has patchy GridSample
coverage, and the arithmetic is small enough to do in the caller:
normalise the point into -1..1
for each of the three planes, taking coordinate pairs (x,y), (x,z), (y,z):
bilinear sample that plane, align_corners=false
concatenate the three results -> 120 features
exp(raw + bias) and colour through a sigmoid, so the caller gets values
it can use rather than logits it has to remember to transform. Getting that
wrong produces a shape that is quietly the wrong size.Every graph is compared against the PyTorch reference on a real photograph, not noise, because noise never exercises the range a trained encoder sees. The fp32 export matches at cosine 1.000000 on triplane, density and colour. The int8 build holds 0.9998 on the triplane and 99.1% occupancy agreement.
MIT, from the upstream weights and code. The DINO ViT-B/16 encoder folded into the triplane graph is Apache-2.0. Both permit commercial use.