Downloads · 30 days
0
cgb/triposg-onnx-webgpu
triposg-onnx-webgpu is a image-to-3d model from cgb. Use it for the image-to-3d task on the model card, and read the license before you ship it in a product. It is set up for onnx. The card lists the license as mit.
TripoSG (SIGGRAPH 2025, VAST-AI-Research) is a 1.5B rectified-flow diffusion transformer over an SDF VAE: one image in, high fidelity geometry out. This is that pipeline compressed so it runs entirely inside a web bro…
Downloads · 30 days
0
Access
Public
Updated Sep 4, 2026
Repo size
3.5 GB
Likes
0
Public
Click a slice to open those files.
.data3.4 GB · 98%
From the Hugging Face model README
TripoSG (SIGGRAPH 2025, VAST-AI-Research) is a 1.5B rectified-flow diffusion transformer over an SDF VAE: one image in, high fidelity geometry out. This is that pipeline compressed so it runs entirely inside a web browser on WebGPU.
7.8 GB down to 3.5 GB, with the graphs checked one at a time against the originals.
| File | Precision | Size | Runs |
|---|---|---|---|
triposg_image_encoder_q8.onnx + .data | int8, block 32 | 345 MB | once |
triposg_vae_latents_q8.onnx + .data | int8, block 32 | 227 MB | once |
triposg_dit_step_fp16.onnx + .data | float16 | 2881 MB | 100 times |
triposg_vae_decoder.onnx | float32 | 51 MB | per chunk of points |
Run-once graphs quantize cleanly; iterated ones do not, and this pipeline runs its transformer a hundred times per generation with classifier-free guidance. So the two single-shot graphs are int8 and the DiT is float16. Measured on the same input, against the float32 originals:
| Graph | Cosine |
|---|---|
| image encoder, int8 | 0.999899 |
| vae latents, int8 | 0.999554 |
| DiT step, float16 | 0.999993, no NaN |
The field decoder stays float32: it is small, it is evaluated at every sample point, and its sign is what decides which side of the surface a point is on.
image -> cut out, composited over WHITE, cropped to the subject with 10%
padding, shortest edge to 256, centre crop to 224, divided by 255
(ImageNet mean and std are inside the graph)
-> image_embeds [1,257,1024]
latents = randn(1, 2048, 64)
for i in 0..49:
sigma = 1 - i/50
sigma_next = 1 - (i+1)/50
v_uncond = dit(latents, 1000*sigma, zeros[1,257,1024])
v_cond = dit(latents, 1000*sigma, image_embeds)
latents += (sigma - sigma_next) * (v_uncond + 7.0*(v_cond - v_uncond))
kv_cache = vae_latents(latents) once
sdf = vae_decoder(kv_cache, points) chunks of <= 8192
mesh = marching_cubes(-sdf, level=0) over (-1.005 .. 1.005)
(sigma_i - sigma_next) * v. Stock diffusers'
FlowMatchEulerDiscreteScheduler uses the opposite sign, because its models
predict noise - x0 where this one predicts x0 - noise.The upstream export notes say the decoder ends in * -1 and is inside
positive; the standalone model card says it is outside positive and must be
negated. Measured on a cut out chair, 95.2% of the box comes back positive, and
an object filling 95% of its bounding box is not an object. This export is
outside positive: negate it, or take the object to be where the field is
negative.
scripts/triposg-verify.py compares each compressed graph with its original on
identical input, and separately runs the whole pipeline to a mesh, because the
two fail differently: a compression problem shows up as a cosine, and a contract
problem shows up as a field that never crosses zero or crosses it everywhere. On
a cut out chair the finished field is negative over 4.8% of its box.
Derived from the export by
fernandotonon for
QtMeshEditor, whose
docs/TRIPOSG_EXPORT_NOTES.md is the most useful description of this model's
inference contract anywhere, and was read from the upstream sources rather than
guessed at.
MIT, from the upstream code and weights. The DINOv2-large encoder is Apache-2.0 (credit Meta AI). Both permit commercial use.
Note for anyone reproducing this: TripoSG's own reference pipeline mattes with BriaRMBG, which is non-commercial. It is not used here and not needed; any permissively licensed matting model works.