Downloads · 30 days
0
AnaRMSuni/week4-caa-unet
week4-caa-unet is a image segmentation model from AnaRMSuni. Use it for the image segmentation task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
This repository contains a U-Net convolutional neural network trained for binary medical image segmentation of gastrointestinal polyps.
Downloads · 30 days
0
Access
Public
Updated Mar 12, 2026
Parameters
31.1M
248 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors124 MB · 100%
From the Hugging Face model README
This repository contains a U-Net convolutional neural network trained for binary medical image segmentation of gastrointestinal polyps.
The model predicts segmentation masks that identify polyp regions in endoscopic images. It was implemented in PyTorch and trained using the Kvasir-SEG dataset.
The model is based on the U-Net architecture, a convolutional neural network designed for biomedical image segmentation.
The network follows a encoder–decoder structure with skip connections.
The encoder extracts hierarchical image features through repeated convolution blocks.
Each block contains:
Feature channels increase as depth increases:
64 → 128 → 256 → 512 → 1024
The deepest layer acts as the bottleneck, capturing high-level semantic information.
The decoder reconstructs the segmentation mask while recovering spatial information lost during downsampling.
Each decoder stage performs:
Channel sizes decrease as spatial resolution increases:
512 → 256 → 128 → 64
The final layer is a 1×1 convolution producing a single-channel segmentation mask.
Training was performed using the Kvasir-SEG dataset, which contains endoscopic images with manually annotated polyp masks.
Dataset characteristics:
Dataset split:
Before training, images and masks undergo the following preprocessing steps:
Segmentation masks are converted to binary format using thresholding.
The model is trained using a combined loss function:
BCEWithLogitsLoss + Dice Loss
Binary Cross Entropy performs pixel-wise classification between foreground (polyp) and background.
It provides stable gradients and reliable convergence during training.
Dice Loss measures the overlap between predicted masks and ground truth masks.
This loss is particularly useful in medical image segmentation, where foreground regions often occupy a small portion of the image.
Combining BCE and Dice Loss allows the model to:
The model is evaluated using two segmentation metrics.
Measures similarity between predicted and ground truth masks.
Dice = (2 × Intersection) / (Prediction + Ground Truth)
Measures the overlap between predicted and true regions relative to their union.
IoU = Intersection / Union
These metrics are widely used for benchmarking segmentation models in medical imaging.
Training settings used:
Model weights are stored using the safetensors format.
Example of loading the model from the Hugging Face Hub:
from model import UNet
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
model = UNet(in_channels=3, out_channels=1)
weights_path = hf_hub_download(
repo_id="AnaRMSuni/week4-caa-unet",
filename="model.safetensors"
)
model.load_state_dict(load_file(weights_path))
model.eval()
This model is intended for:
It should not be used for clinical diagnosis or medical decision making.
AnaRMSuni