Downloads · 30 days
0
Jasirdeen/efficientnet-b3-image-forgery-detection
efficientnet-b3-image-forgery-detection is a image classification model from Jasirdeen. Use it when you need a label for an image. It is set up for pytorch. The card lists the license as mit.
This repository contains a fine-tuned EfficientNet-B3 model for binary classification of synthetic identity-document images as genuine or forged.
Downloads · 30 days
0
Access
Public
Updated Aug 2, 2026
Repo size
43.4 MB
Likes
1
Public
Click a slice to open those files.
.pt43.4 MB · 100%
From the Hugging Face model README
This repository contains a fine-tuned EfficientNet-B3 model for binary classification of synthetic identity-document images as genuine or forged.
The checkpoint is provided as a PyTorch state_dict:
efficientnet_b3_finetuned.pt
timmLabels:
0 = genuine / bona fide
1 = forged / tampered
The model was trained and evaluated on the template/composite portion of the Synthetic dataset of ID and Travel Documents (SIDTD), generated using MIDV-2020 document templates.
Dataset composition used in this project:
Genuine images: 1,000
Forged images: 1,222
Total images: 2,222
The forged samples include the following SIDTD annotation types:
Crop_and_ReplaceInpaint_and_RewriteThe dataset itself is not included in this repository. Please follow the SIDTD dataset terms and citation requirements before downloading or redistributing it.
Each image is:
mean = [0.485, 0.456, 0.406]
std = [0.229, 0.224, 0.225]
Training used mild brightness, contrast, and saturation variation. Horizontal flipping was not used because mirrored identity documents are not realistic examples.
The model was initialized from a pretrained EfficientNet-B3. Phase 1 trained the binary classifier head. The final checkpoint was then fine-tuned by unfreezing the final EfficientNet feature block and classifier head.
Fine-tuning used:
1e-4;1e-4;On the fixed 223-image test split at a decision threshold of 0.5:
Accuracy: 98.21%
Precision: 97.60%
Recall: 99.19%
F1: 98.39%
ROC-AUC: 99.98%
Confusion matrix, with rows representing actual labels and columns representing predictions:
[[97, 3],
[ 1, 122]]
Subtype recall on the forged test samples:
Crop-and-replace: 94.74%
Inpainting: 100.00%
These results are specific to the project’s synthetic dataset and split. They should not be interpreted as real-world identity-verification performance.
Install the required packages:
pip install torch torchvision timm pillow
Load the checkpoint:
from pathlib import Path
import timm
import torch
from PIL import Image
from torchvision import transforms
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = timm.create_model(
"efficientnet_b3",
pretrained=False,
num_classes=2,
)
model.load_state_dict(
torch.load(
"efficientnet_b3_finetuned.pt",
map_location="cpu",
weights_only=True,
)
)
model.to(device)
model.eval()
preprocess = transforms.Compose([
transforms.Resize((300, 300)),
transforms.ToTensor(),
transforms.Normalize(
mean=(0.485, 0.456, 0.406),
std=(0.229, 0.224, 0.225),
),
])
image = Image.open("document.jpg").convert("RGB")
input_tensor = preprocess(image).unsqueeze(0).to(device)
with torch.no_grad():
probabilities = torch.softmax(model(input_tensor), dim=1)[0]
genuine_probability = float(probabilities[0])
forged_probability = float(probabilities[1])
label = "forged" if forged_probability >= 0.5 else "genuine"
print({
"label": label,
"genuine_probability": genuine_probability,
"forged_probability": forged_probability,
})
The project uses Grad-CAM with the final spatial convolutional layer of EfficientNet-B3. The resulting heatmap indicates image regions that contributed to the selected prediction.
Grad-CAM is an explanation aid and should not be interpreted as a guaranteed localization of the tampered region.
This model is intended for:
This model uses the SIDTD dataset and MIDV-2020-derived document templates. Please cite and follow the terms of the original SIDTD and MIDV-2020 resources when using this model or reproducing the experiments.
SIDTD project: https://github.com/Oriolrt/SIDTD_Dataset
The model repository is released under the MIT license. Dataset licensing, attribution, and redistribution terms remain subject to the original SIDTD and MIDV-2020 sources.