Downloads ยท 30 days
0
Sivamorgan/pascal-triheadnet
pascal-triheadnet is a image segmentation model from Sivamorgan. Use it for the image segmentation task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
Single-stage unified perception model for Pascal VOC: Detection, Semantic, and Instance Segmentation in one forward pass.
Downloads ยท 30 days
0
Access
Public
Updated Jan 18, 2026
Repo size
962 MB
Likes
0
Public
Click a slice to open those files.
.pth962 MB ยท 100%
From the Hugging Face model README
Single-stage unified perception model for Pascal VOC: Detection, Semantic, and Instance Segmentation in one forward pass.
Pascal-TriheadNet is a multi-task learning model that jointly solves three computer vision tasks using a unified Vision Transformer backbone with three specialized task heads. Validated on Pascal VOC 2012, it achieves strong performance across all tasks while maintaining efficient inference.
๐ View Full Code & Documentation on GitHub
Two versions of the model are provided:
| File | Description | Size |
|---|---|---|
checkpoint_epoch_50.pth | Best performing FP32 model. | 826MB |
checkpoint_epoch_50_quantized.pth | optimized INT8 Quantized model. | 136MB |
Training Context: Model was fine-tuned on an L4 GPU in Google Colab.
Evaluated on the Pascal VOC 2012 Validation set:
| Task | Metric | Score |
|---|---|---|
| Detection | mAP (0.5:0.95) | 46.7% |
| Detection | mAP@50 | 75.6% |
| Semantic | mIoU | 87.3% |
| Instance | Mask mAP (0.5:0.95) | 35.8% |
| Instance | Mask mAP@50 | 65.7% |
For detailed per-class analysis and ablation studies, please refer to the GitHub Repository.
The architecture utilizes a Vision Transformer (ViT-Base) backbone pretrained on ImageNet.
vit_base_patch16_224 with the last 6 blocks fine-tuned.