Downloads · 30 days
0
malikhannan111/crowd_counter
crowd_counter is a machine learning model from malikhannan111. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
0
Access
Public
Updated May 10, 2026
Repo size
50.1 MB
Likes
0
Public
Click a slice to open those files.
.pth36.8 MB · 73%
From the Hugging Face model README
A deep learning model for accurate crowd counting that predicts density maps from crowd images. Built using VGG16-BN encoder with dilated convolutional decoder, trained on ShanghaiTech Part A and Part B datasets. The model outputs both a density map visualization and the total count of people in the image.
This project implements crowd counting using density map regression. The model takes a crowd image as input and generates a density map where each pixel represents crowd density. Summing all pixel values gives the total number of people. The architecture uses a pretrained VGG16-BN backbone for feature extraction and a custom decoder with dilated convolutions and upsampling layers for high-quality density map estimation.
| Dataset | Images | Average Count | Purpose |
|---|---|---|---|
| ShanghaiTech Part A | 482 | 501 | Dense crowd scenes |
| ShanghaiTech Part B | 716 | 123 | Sparse street crowds |
| Component | Model | Parameters | Task |
|---|---|---|---|
| Encoder | VGG16-BN (ImageNet pretrained) | 14.7M | Feature extraction |
| Decoder | Dilated Convolutions + Upsampling | 3.5M | Density map generation |
| Total | End-to-end | 9.2M | Crowd counting |
| Dataset | MAE | MSE | Error Percentage |
|---|---|---|---|
| Part A Test Set | 24.00 | 1202.17 | 5.3% |
| Part B Test Set | 18.90 | 400.23 | 15.4% |
| Cross-Dataset | 20.97 | 1267.77 | - |
| File | Purpose |
|---|---|
crowd_counting_dmcount.ipynb | Complete training and evaluation code |
crowd_counter.pth | Trained model weights |
Download the trained model file (crowd_counter.pth) from Google Drive:
Download Link: https://drive.google.com/file/d/1eNbohPZPNYxt2j5dZOIIftFr7cpHEPTO/view?usp=sharing
Upload the downloaded crowd_counter.pth file to your Google Colab environment or any local Python environment.
pip install torch torchvision numpy opencv-python matplotlib pillow
import torch
import torch.nn as nn
import torchvision.models as models
import cv2
import numpy as np
import matplotlib.pyplot as plt
# Define model architecture
class Counter(nn.Module):
def __init__(self):
super().__init__()
vgg = models.vgg16_bn(weights=None)
self.encoder = nn.Sequential(*list(vgg.features.children())[:33])
self.decoder = nn.Sequential(
nn.Conv2d(512, 256, 3, padding=1), nn.ReLU(),
nn.Upsample(scale_factor=2, mode='bilinear', align_corners=True),
nn.Conv2d(256, 128, 3, padding=1), nn.ReLU(),
nn.Upsample(scale_factor=2, mode='bilinear', align_corners=True),
nn.Conv2d(128, 64, 3, padding=1), nn.ReLU(),
nn.Upsample(scale_factor=2, mode='bilinear', align_corners=True),
nn.Conv2d(64, 1, 1)
)
def forward(self, x):
return self.decoder(self.encoder(x))
# Load model
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = Counter().to(device)
model.load_state_dict(torch.load("crowd_counter.pth", map_location=device))
model.eval()
# Predict on an image
image_path = "your_image.jpg"
img = cv2.cvtColor(cv2.imread(image_path), cv2.COLOR_BGR2RGB)
img_resized = cv2.resize(img, (512, 384))
img_tensor = torch.tensor(img_resized.transpose(2,0,1).astype(np.float32)/255.0).unsqueeze(0).to(device)
with torch.no_grad():
pred = model(img_tensor)
count = pred.sum().item()
print(f"Predicted crowd count: {count:.0f} people")
# Visualize
plt.imshow(img_resized)
plt.imshow(pred.cpu().squeeze().numpy(), cmap="jet", alpha=0.5)
plt.title(f"Count: {count:.0f}")
plt.show()
Replace "your_image.jpg" with the path to any crowd image. The model works best with images containing dense crowds (50+ people).
Dataset: Zhang et al., "Single-Image Crowd Counting via Multi-Column Convolutional Neural Network", CVPR 2016
Method: Wang et al., "Distribution Matching for Crowd Counting", NeurIPS 2020