Search and Rescue Human Detection
A computer vision project using YOLO26 for real-time human detection in search and rescue scenarios. Trained and compared two model variants (Nano and Medium), with the best model achieving 97.24% mAP50 and 72.46% mAP50-95 on the test set.
Note: This project uses the Ultralytics framework with yolo26n.pt and yolo26m.pt pretrained weights.
Problem Statement
In search and rescue operations, rapid identification of humans in diverse environments (aerial footage, disaster zones, wilderness) is critical. Manual review of imagery is time-consuming and prone to errors under pressure. This project develops a deep learning solution to:
- Detect humans accurately in challenging conditions (varied lighting, occlusion, terrain)
- Provide real-time inference suitable for drone/UAV deployment
- Minimize false negatives to ensure no person is missed
Dataset
Source: SARD Computer Vision Dataset from Kaggle
Size: 5,755 images with bounding box annotations
Split Distribution:
- Training: 4,041 images (70.2%)
- Validation: 1,144 images (19.9%)
- Test: 570 images (9.9%)
Annotation Format: YOLO format (class_id, center_x, center_y, width, height - normalized)
Key Characteristics:
- Single class:
human
- Diverse environments: Aerial, forest, unpaved Roads, rocks etc.
- Challenging conditions: Occlusion, poor contrast, small objects etc.
- Image resolution: 640 * 640
Methodology
1. Data Preprocessing & Exploration
- Label Format Conversion: Parsed YOLO format annotations (normalized coordinates) to pixel coordinates for visualization
- Data Verification: Validated all 5,755 image-label pairs for completeness and format consistency
- Exploratory Data Analysis:
- Analyzed bounding box size distribution
- Examined spatial distribution of humans in frames
- Identified class imbalance (N/A - single class)
- Helper Functions: Built utilities for visualization, bbox drawing, and batch inspection
Bounding Box Distribution:
<img src="outputs/EDA_bbox_distribution.png" width="350">
Summary Statistics:
<img src="outputs/summary_statistics.png">
Sample image with bounding box:
<img src="outputs/sample_image_bounding_box.png">
2. Model Architecture
Framework: YOLO26 (Ultralytics)
Why YOLO26?
- State-of-the-art single-stage detector balancing speed and accuracy
- Optimized for real-time inference (critical for SAR drones)
- Strong performance on small objects (humans viewed from aerial perspectives)
- Built-in augmentation pipeline for robustness
- Easy deployment via ONNX export
Model Variants Trained:
- YOLO26 Nano (
yolo26n.pt) - Lightweight, faster inference
- YOLO26 Medium (
yolo26m.pt) - Better accuracy, moderate speed (~8.5x more parameter than Nano)
Input Resolution: 640×640
Architecture Highlights:
- Enhanced backbone for feature extraction
- Multi-scale feature fusion via PANet-style neck
- Decoupled head for classification and localization
- Optimized anchor-free detection
3. Training Configuration
Hyperparameters:
- Epochs: 300 (with early stopping, patience=15)
- Actual Training Duration: Stopped at epoch 243 (early stopping triggered)
- Batch Size: 32
- Optimizer: AdamW
- Initial Learning Rate: 1e-3
- Learning Rate Schedule: Cosine decay
- Weight Decay: 5e-4
- Image Size: 640×640
- Device: NVIDIA L40S (GPU)
- Workers: 4
- Cache Mode: RAM (for faster data loading)
Training Strategy:
- Transfer learning from pretrained YOLO26 weights
- Early stopping based on validation mAP50-95 (patience=15 epochs)
- Model checkpointing (best and last weights saved)
- Cosine learning rate decay for smooth convergence
4. Evaluation Metrics
Primary Metrics:
- Precision: When the model predicts "human", how often is it correct?
- Recall: Out of all actual humans, how many did the model detect?
- mAP50: Mean Average Precision at IoU threshold 0.5 (standard detection benchmark)
- mAP50-95: Mean Average Precision averaged across IoU thresholds 0.5 to 0.95 (stricter metric)
Loss Components:
- Box Loss: Bounding box coordinate regression error
- Class Loss: Classification error (human vs. background)
- DFL Loss: Distribution Focal Loss for precise box boundaries
Results
Model Performance Comparison
YOLO26 Nano:
| Split | Precision | Recall | mAP50 | mAP50-95 |
|---|
| Validation | 96.06% | 90.18% | 96.34% | 68.05% |
| Test | 92.98% | 92.18% | 96.40% | 67.93% |
YOLO26 Medium (Best Model):
| Split | Precision | Recall | mAP50 | mAP50-95 |
|---|
| Validation | 96.73% | 94.03% | 97.43% | 72.77% |
| Test | 94.31% | 94.38% | 97.24% | 72.46% |
Performance Improvement (Medium vs Nano on Test Set):
- Precision: +1.33%
- Recall: +2.20%
- mAP50: +0.84%
- mAP50-95: +4.53%
Interpretation:
- YOLO26 Medium provides the best overall performance with 94.31% precision and 94.38% recall
- High recall (94.38%) ensures most humans are detected - critical for SAR where missing a person has serious consequences
- High precision (94.31%) means low false positives (~5.7%), reducing wasted effort investigating false alarms
- 97.24% mAP50 demonstrates strong localization accuracy at standard IoU threshold
- 72.46% mAP50-95 shows good performance even with strict overlap requirements
- Medium model's 4.53 point improvement in mAP50-95 indicates significantly better bounding box precision
Training Progression (Medium)

Key Observations:
- Losses steadily decreased and plateaued around epoch 150-200
- Validation metrics stabilized with minimal overfitting
- Early stopping triggered at epoch 243 (no improvement for N epochs)
Qualitative Results (Medium)
Sample Detections:

Comparison between Nano and Medium:

Observation: Both model successfully detects partially occluded humans. Medium model consistently performs better than nano.
Key Insights
From Model Comparison (Nano vs Medium):
- Model Size vs Performance Trade-off: Medium model achieves 4.53% in mAP50-95 over Nano, indicating significantly better localization precision
- Recall Improvement: Medium model's +2.2% recall gain means detecting ~2% more humans - crucial for SAR
- Balanced Performance: Medium model achieves balanced precision (94.31%) and recall (94.38%), avoiding the precision-recall trade-off
- Deployment Consideration: Nano model offers 96.40% mAP50 with faster inference - viable for resource-constrained drones if speed is prioritized
From Training Process:
- Early Stopping Effectiveness: Training stopped at epoch 243/300, preventing overfitting while achieving strong generalization
- Transfer Learning Success: Pretrained YOLO26 weights accelerated convergence and improved final performance
- Stable Training: Cosine LR decay with AdamW optimizer enabled smooth convergence without oscillations
From Model Performance:
- High Recall Priority: 94.38% recall ensures minimal missed detections - critical for SAR where missing a person has serious consequences
- Low False Positive Rate: 94.31% precision means only ~5.7% false positives, reducing wasted effort investigating false alarms
- Robust Localization: 72.46% mAP50-95 indicates bounding boxes are well-calibrated across strict IoU thresholds
- Strong Generalization: Test metrics (97.24% mAP50) closely match validation (97.43%), confirming model isn't overfitting
Deployment Considerations:
- Real-Time Capable: YOLO26 architecture enables 152.3 FPS on NVIDIA L40S, suitable for drone inference
- ONNX Export: Both models successfully exported to ONNX for cross-platform deployment (edge devices, web, mobile)
- Model Selection: Use Medium for maximum accuracy, Nano for speed-critical applications
- Failure Modes: Struggle with extremely small objects, and tends to perform worse on direct top down view where the leg and hand may be hidden by the head.
Future Improvements
1. Multi-Class Detection (If Applicable)
- Extend to detect specific categories: injured persons, rescue personnel, animals
- Enable priority tagging for triage in mass casualty events
2. Temporal Consistency (for video SAR missions)
- Implement tracking (DeepSORT, ByteTrack) to associate detections across frames
- Count unique individuals rather than per-frame detections
- Reduce false positives by filtering isolated single-frame detections
3. Adversarial Robustness
- Test on extreme weather conditions (fog, heavy rain, snow)
- Augment training data with synthetic weather overlays
- Evaluate performance degradation and set realistic confidence thresholds
4. Uncertainty Quantification
- Implement Monte Carlo Dropout or ensemble methods for confidence estimation
- Flag low-confidence regions for manual review
5. Efficiency Optimization
- Prune redundant weights to reduce model size
- Benchmark on target hardware (Jetson Nano, Raspberry Pi, drone compute modules)
6. Dataset Expansion
- Collect more examples of edge cases: prone humans, partial occlusion, camouflage clothing
- Include thermal imaging data for night operations
- Add crowd scenarios for mass-casualty incident training
Technologies
- Python 3.12
- Deep Learning Framework: PyTorch (via Ultralytics)
- Computer Vision: OpenCV (cv2), Python Imaging Library (PIL)
- Model Architecture: YOLO26
- Data Processing: NumPy, Pandas
- Visualization: Matplotlib, Seaborn
- Deployment: ONNX Runtime