Downloads · 30 days
0
AMD-PAVS-AI/unet
unet is a image segmentation model from AMD-PAVS-AI. Use it for the image segmentation task on the model card, and read the license before you ship it in a product. It is set up for onnx. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Aug 4, 2026
Repo size
469 KB
Likes
1
Public
Click a slice to open those files.
.png469 KB · 99%
From the Hugging Face model README

UNet-S5-D16 is a convolutional encoder-decoder network for semantic segmentation — it classifies every pixel in an image into one of 19 urban-scene categories (road, sidewalk, building, person, car, etc.). This repository packages inference for semantic segmentation using ONNX Runtime, exported and validated for AMD ROCm so it runs efficiently on AMD GPUs, CPUs, and NPUs.
This is based on the implementation of UNet found here. This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the unet AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.
Task: Cityscapes semantic segmentation (19 classes)
Dataset: Cityscapes val split (500 images), downloaded automatically from HuggingFace; benchmark inputs from UrbanSyn
Output metrics: mIoU (mean Intersection-over-Union across 19 classes), per-class IoU, pixel accuracy
This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm, as well as AMD CPUs and AMD Ryzen AI NPUs. Key points:
| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| ONNX Runtime | FP32 | CPU Execution Provider | AMD CPU | — |
| ONNX Runtime | FP32 / FP16 / BF16 / INT8 | MIGraphX Execution Provider | AMD Instinct™ / Radeon™ GPU (ROCm) | First run may take 30+ minutes due to graph compilation |
| ONNX Runtime | FP32 / FP16 / BF16 / INT8 | VitisAI Execution Provider | AMD Ryzen AI NPU | INT8 requires Quark + Cityscapes calibration quantization step |
For setup instructions, evaluation scripts, and custom configuration options, see the unet on GitHub.
Model Type: Semantic segmentation (convolutional encoder-decoder)
Base Model: UNet-S5-D16 (mmsegmentation)
Model Stats:
input): (1, 3, 512, 1024) float32output): (1, 19, 512, 1024) float32Higher mIoU means predicted pixel labels agree more closely with ground truth across all 19 Cityscapes classes — 100% is perfect overlap, 0% is no agreement. Published paper mIoU for UNet-S5-D16 on Cityscapes is 69.10%.
| Metric | Description |
|---|---|
| mIoU | Mean Intersection-over-Union averaged across all 19 Cityscapes classes. Higher means better boundary alignment and class assignment. |
Cityscapes val mIoU — UNet-S5-D16 (paper: 69.10%):
<!-- accuracy-table-start -->| Device | Precision | mIoU |
|---|---|---|
| CPU | FP32 | 69.31% |
| GPU | FP32 | 69.31% |
| GPU | FP16 | 69.32% |
| GPU | BF16 | 69.34% |
| GPU | INT8 | 69.31% |
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?
📂 View the full project on GitHub
The GitHub repository includes: