Downloads · 30 days
0
Subh775/Seg-Basil-rfdetr
Seg-Basil-rfdetr is a image segmentation model from Subh775. Use it for the image segmentation task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
<p align="left" <a href="https://opensource.org/licenses/Apache-2.0"<img src="https://img.shields.io/badge/License-Apache2.0-blue.svg" /</a <a href="https://blog.roboflow.com/rf-detr-segmentation-preview/"<img src="ht…
Downloads · 30 days
0
Access
Public
Updated Nov 18, 2025
Repo size
1.5 GB
Likes
0
Public
Click a slice to open those files.
.pth1.5 GB · 100%
From the Hugging Face model README
| Model | Best EMA Mask mAP (@.50:.95) |
|---|---|
| LeafNet75/Segment-Tulsi-TFs | 0.9650 |
| Subh775/Seg-Basil-rfdetr | 0.9668 |
RF-DETR Object Detection vs YOLOv12 : A Study of Transformer-based and CNN-based Architectures for Single-Class and Multi-Class Greenfruit Detection in Complex Orchard Environments Under Label Ambiguity at: arxiv.org/abs/2504.13099
This model card explores the application of Roboflow’s RF-DETR for leaf instance segmentation, focusing particularly on Ocimum tenuiflorum (Holy Basil). Unlike traditional CNN-based segmentation models, transformers can effectively capture global dependencies through attention mechanisms, leading to improved contextual understanding and better generalization performance.
RF-DETR represents one of the first transformer-based architectures to demonstrate that transformers can achieve both high accuracy and fast inference speeds, outperforming many CNN-based models in detection and segmentation tasks despite their traditionally heavier computational design.
RF-DETR integrates architectural innovations from Deformable DETR and LW-DETR, and utilizes a DINOv2 backbone, offering superior global context modeling and domain adaptability.
<!-- ### Example Outputs Here are output examples from the model's validation run: <table> <tr> <td><img src="https://cdn-uploads.huggingface.co/production/uploads/66c6048d0bf40704e4159a23/aNF_6VN8FgBbYxWnA6Uwm.png" width="350"/></td> <td><img src="https://cdn-uploads.huggingface.co/production/uploads/66c6048d0bf40704e4159a23/8xF1fPREvmJ0OmgC-BBs6.png" width="350"/></td> </tr> <tr> <td><img src="https://cdn-uploads.huggingface.co/production/uploads/66c6048d0bf40704e4159a23/9bBN7GXxpn6Ly8rLCkLhA.jpeg" width="350"/></td> <td><img src="https://cdn-uploads.huggingface.co/production/uploads/66c6048d0bf40704e4159a23/ZzmJCbK7hvKVK0EmBf3L-.png" width="350"/></td> </tr> </table> -->The model is trained on: https://universe.roboflow.com/politicians/tulsi-wgmfs using COCO dataset format for RF-DETR Seg Preview.
Training followed the official Roboflow implementation. The model was initialized with pretrained weights and trained using the AdamW optimizer, more params are as:
epochs=3, # Updated from original run which stopped effectively at epoch 1
batch_size=2,
grad_accum_steps=4,
lr=1e-4, #default
pretrain_weights='rf-detr-seg-preview.pt', #default
layer_norm=True,
checkpoint_interval=12, # Note: Actual saving seems per-epoch based on best metrics
seed=42,
num_workers=2,
device='cuda', #T4 colab GPU
resolution=432,
lr_scheduler='step',
tensorboard=True, #check Training metrics
class_names=['Tulsi'],
segmentation_head=True
Here is the training results over 3 epochs (note: peak performance at Epoch 1):

The training ran for 3 epochs on Colab's T4 GPU, with the best performance achieved at Epoch 1. The metrics below are for the Exponential Moving Average (EMA) model (checkpoint_best_ema.pth saved at Epoch 1), which represents a smoothed-out and more stable version of the model's weights. Training beyond Epoch 1 showed signs of overfitting.
| Metric | Value | Description |
|---|---|---|
| mAP (Masks) @.50:.95 | 0.9668 | Primary metric for segmentation. |
| mAP (Boxes) @.50:.95 | 0.9491 | Primary metric for bounding box detection. |
| mAP (Masks) @.50 | 0.9783 | Segmentation quality at 50% overlap threshold. |
| mAP (Boxes) @.50 | 0.9783 | Bounding box quality at 50% overlap threshold. |
| Precision (Boxes) | 0.9820 | Accuracy of positive predictions. |
| Recall (Boxes) | 0.9300 | Ability to find all positive instances. |
The training graph metrics_plot.png (updated for the 3-epoch run) shows:
In all evaluation plots, the EMA Model (orange dashed line) consistently achieves higher scores and shows more stability than the Base Model (blue solid line). Both models show that peak performance was reached at Epoch 1, confirming that checkpoint_best_ema.pth saved at that point is the optimal model.
Grounding-DINO, YOLO-World, Owl-ViTDinov2, Dinov3(vision + language)