Downloads ยท 30 days
59
26% of all-time downloads
XYZ9843/GOOSE-M2F
GOOSE-M2F is a image segmentation model from XYZ9843. Use it for the image segmentation task on the model card, and read the license before you ship it in a product. It is set up for transformers.
Jyothiraditya Lingam, Nikhileswara Rao Sulake, Sai Manikanta Eswar Machara
Downloads ยท 30 days
59
26% of all-time downloads
All-time downloads
229
Public
Repo size
3.5 GB
Likes
1
Public
Click a slice to open those files.
.pth3.5 GB ยท 100%
From the Hugging Face model README
Jyothiraditya Lingam, Nikhileswara Rao Sulake, Sai Manikanta Eswar Machara
Department of Computer Science and Engineering Rajiv Gandhi University of Knowledge Technologies (RGUKT), Nuzvid, Andhra Pradesh, India
<p align="center"> <a href="https://arxiv.org/abs/2606.15937"><b>๐ Paper</b></a> โข <a href="https://github.com/Aditya-Lingam-9000/GOOSE-M2F"><b>๐ป Code</b></a> โข <a href="https://huggingface.co/XYZ9843/GOOSE-M2F"><b>๐ค Hugging Face</b></a> โข <a href="https://www.codabench.org/competitions/14257"><b>๐ Challenge</b></a> </p>GOOSE-M2F is a task-specific adaptation of Mask2Former for the GOOSE 2D Fine-Grained Semantic Segmentation Challenge (ICRA 2026). The proposed framework addresses long-tailed semantic segmentation in unstructured outdoor environments through enhanced object query capacity, feature refinement, auxiliary supervision, class-balanced optimization, and robust multi-scale inference.
Official Challenge Performance: 70.08% Composite mIoU (63.55% Fine mIoU, 76.61% Coarse mIoU), achieving 3rd Place on the GOOSE 2D FGSS Challenge Leaderboard.
The GOOSE dataset presents one of the most challenging real-world segmentation benchmarks: 64 fine-grained classes across diverse unstructured outdoor environments including forests, gravel paths, construction zones, and agricultural terrain โ with a severely long-tailed class distribution.
GOOSE-M2F extends the baseline Mask2Former (Swin-Large backbone) with three key modifications engineered specifically for this challenge:
| Modification | Problem Solved | Impact |
|---|---|---|
| 200 Object Queries (vs 100) | Query saturation in 64-class scenes | +2-3% composite mIoU |
| Feature Refinement Module (FRM) โ ASPP-lite + CBAM | Over-segmentation of amorphous terrain classes | +3-4% on Vegetation/Terrain |
| Auxiliary Supervision Head at H/4 resolution | Vanishing gradients for tiny/thin classes | +5-8% on rare classes |
Input Image [B, 3, H, W]
โ
โผ
Swin-Large Backbone (Hierarchical, 4 stages)
Stage 1-4: channels {192, 384, 768, 1536}, resolutions {H/4 โ H/32}
โ
โผ
MSDeformAttn Pixel Decoder (6-layer FPN)
Output: mask_features [B, 256, H/4, W/4]
โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โผ โผ
[NEW] Feature Refinement Module [NEW] Auxiliary Head
ASPP-lite: dilations {1, 3, 6, 12} Conv(256โ256โ64)
+ Global Average Pooling DB-weighted CE loss
+ CBAM Dual-Attention (Ch + Sp) Supervised at H/4
โ
โผ
Transformer Decoder (9 layers)
[MOD] 200 Object Queries (was 100)
Masked Cross-Attention
โ
โผ
Class Head [B, 200, 65] ร Mask Head [B, 200, H/4, W/4]
โ
โผ
Hungarian Matching โ Semantic Prediction
| Technique | Description |
|---|---|
| Distribution-Balanced (DB) Loss | w_c = (1-ฮฒ)/(1-ฮฒ^n_c), ฮฒ=0.9999. Amplifies gradients for rare classes. |
| Rare-Class Copy-Paste (RCCP) | Pre-extracted rare-class cutouts pasted onto training images at 85% probability. |
| Dynamic IoU-Aware Weights | Per-class loss weights updated every epoch from validation IoU (0%โ4x, 80%+โ1x). |
| 10x LR Jump (V4) | Backbone 1e-5, Decoder 5e-5 โ broke the model out of a local minimum at ~55%. |
| EMA (decay=0.9995) | Shadow weights consistently +1.0โ1.5% over raw model on validation. |
| Class-Aware Repeat Sampling | Oversamples images containing rare classes proportional to their rarity. |
| Polynomial LR Decay | Gradual decay after warmup, with annealing in final sessions. |
| Session | Base LR | Backbone LR | Official Score |
|---|---|---|---|
| V1 (S3) | 5e-6 | 1e-6 | 50.68% |
| V2 (S4) | 5e-6 | 1e-6 | 54.62% |
| V3 (S5) | 5e-6 | 1e-6 | 55.64% |
| V4 (S6) | 5e-5 | 1e-5 | 56.38% โ 10x LR Jump |
| V5 (S7) | 5e-5 | 1e-5 | 57.59% |
| V6 (S8) | 5e-5 | 1e-5 | 58.58% |
| V7 (S9) | 5e-5 | 1e-5 | 59.23% |
| V8 (S10) | 2.5e-5 | 5e-6 | 59.51% โ Annealing |
| Inference | โ | โ | 70.08% โ +10.57% from TTA |
The final performance leap from 59.51% (training) to 70.08% (submission) came entirely from the inference pipeline:
| Technique | Gain | Description |
|---|---|---|
| Dense Sliding Window | +4-5% | 896ร896 crops, stride=384px (57% overlap) |
| 2D Gaussian Kernel Blending | Eliminates artifacts | Center pixels weighted higher, edges down-weighted |
| 4-Scale TTA | +3-4% | Scales: 0.5ร, 0.75ร, 1.0ร, 1.5ร |
| H-Flip TTA | +1-2% | 8 total views per image (4 scales ร 2 flips) |
| EMA Weights | +1-1.5% | Shadow weights used instead of raw training weights |
| AuxHead Stripping | VRAM savings | Removed before inference โ not needed for prediction |
goose-m2f/
โโโ src/
โ โโโ model.py โ GOOSEMask2Former (FRM + AuxHead + 200 queries)
โ โโโ features.py โ Dataset, augmentations, EMA, metrics
โ โโโ train.py โ Training engine (Trainer class)
โ โโโ inference.py โ Dense Gaussian patch-blending inference
โโโ configs/
โ โโโ train_config.yaml โ All training hyperparameters
โ โโโ infer_config.yaml โ TTA and inference settings
โโโ data/raw/ โ Dataset (symlink or copy)
โโโ models/ โ Manually placed checkpoints
โโโ outputs/
โ โโโ checkpoints/ โ best_model.pth, latest.pth, charts
โ โโโ predictions/ โ Output PNG predictions
โโโ tests/
โ โโโ test_model.py โ pytest unit tests
โโโ instructions/
โ โโโ instructions.md โ Full setup + usage guide
โโโ requirements.txt
git clone https://github.com/Aditya-Lingam-9000/GOOSE-M2F
cd GOOSE-M2F
conda create -n goose python=3.11 -y && conda activate goose
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121
pip install -r requirements.txt
accelerate config # Configure for your GPU setup
Edit configs/train_config.yaml:
data_dir: "/path/to/goose_dataset"
csv_path: "/path/to/goose_label_mapping.csv"
output_dir: "outputs/checkpoints/session_01"
# Single GPU
python -m src.train --config configs/train_config.yaml
# Multi-GPU
accelerate launch --num_processes 2 -m src.train --config configs/train_config.yaml
Edit configs/infer_config.yaml with the checkpoint path and image directory, then:
python -m src.inference --config configs/infer_config.yaml
pytest tests/ -v
| Metric | Score |
|---|---|
| Fine mIoU | ~68.5% |
| Coarse mIoU | ~71.6% |
| Official Composite | 70.08% |
| Category | mIoU |
|---|---|
| Sky | 94.6% |
| Road | 91.0% |
| Vehicle | 89.8% |
| Vegetation | 89.8% |
| Construction | 75.5% |
| Terrain | 78.9% |
| Human | 62.8% |
| Sign | 62.4% |
| Water | 33.9% |
| Object | 51.3% |
| Animal | 0.0% |
| Package | Version |
|---|---|
| torch | โฅ 2.1.0 |
| transformers | โฅ 4.38.0 |
| accelerate | โฅ 0.27.0 |
| albumentations | โฅ 1.3.1 |
| opencv-python | โฅ 4.9.0 |
| numpy | โฅ 1.24.0 |
See requirements.txt for the complete list.
If you use this work, please cite:
@techreport{lingam2026goosem2f,
title = {GOOSE-M2F: Adapting Mask2Former for High-Fidelity, Long-Tailed Fine-Grained Semantic Segmentation in Unstructured Outdoor Terrain},
author = {Jyothiraditya Lingam and Nikhileswara Rao Sulake and Sai Manikanta Eswar Machara},
year = {2026},
institution = {Rajiv Gandhi University of Knowledge Technologies (RGUKT)}
}