Downloads · 30 days
0
Gazeux33/VeridisQuo
VeridisQuo is a video classification model from Gazeux33. Use it for the video classification task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<div <img src="https://img.shields.io/badge/PyTorch-2.2+-EE4C2C?style=for-the-badge&logo=pytorch&logoColor=white" alt="PyTorch"/ <img src="https://img.shields.io/badge/Parameters-25M-blue?style=for-the-badge" alt="Par…
Downloads · 30 days
0
Access
Public
Updated Dec 31, 2025
Repo size
389 MB
Likes
1
Public
Click a slice to open those files.
.pth101 MB · 100%
From the Hugging Face model README
VeridisQuo ("Where is the truth?" in Latin) is a specialized neural network for detecting deepfake manipulations in face images. Unlike traditional approaches that rely solely on spatial features, VeridisQuo employs a dual-stream architecture that combines:
This hybrid approach exploits a fundamental weakness in deepfake generation: while GANs and diffusion models can produce visually convincing spatial patterns, they often leave telltale signatures in the frequency domain—particularly in how they distribute high-frequency components and introduce subtle compression inconsistencies.
Most deepfake generators operate in the spatial domain and are trained to fool human perception (and spatial-only CNNs). However:
By combining both streams, VeridisQuo achieves robust detection even against adversarially-trained generators.
Input: RGB Image [3, 224, 224]
│
├─► Spatial Branch
│ └─► EfficientNet-B4 (pretrained on ImageNet)
│ └─► Global Average Pooling
│ └─► [1792-dim features]
│
└─► Frequency Branch
├─► DCT Extractor (8×8 blocks, frequency band aggregation)
│ └─► [512-dim features]
│
└─► FFT Extractor (8 radial bands, Hann windowing)
└─► [512-dim features]
Concatenate DCT + FFT
└─► Fusion MLP
└─► [1024-dim features]
Concatenate Spatial + Frequency
└─► LayerNorm
└─► [2816-dim combined features]
└─► Classification Head (3-layer MLP)
└─► [2 classes: FAKE / REAL]
LayerNorm over BatchNorm: The classifier uses LayerNorm instead of BatchNorm to support single-image inference (batch_size=1) without requiring running statistics.
Feature Dimension Balance: The 1792:1024 ratio between spatial and frequency features reflects their relative discriminative power—spatial features capture broader semantic context, while frequency features provide specialized forensic signals.
No Softmax in Forward Pass: The model outputs raw logits; apply torch.softmax(logits, dim=1) for probabilities.
Trained on FaceForensics++ (C23 compression), a benchmark dataset containing:
Dataset available on Kaggle: VeridisQuo Preprocessed Dataset
| Parameter | Value | Rationale |
|---|---|---|
| Optimizer | AdamW | Better weight decay regularization than Adam |
| Learning Rate | 1e-4 → 1e-6 | Cosine annealing with 3-epoch warmup |
| Batch Size | 64 | Optimal for 16GB GPU memory |
| Weight Decay | 1e-4 | L2 regularization to prevent overfitting |
| Epochs | 7 | Early stopping after validation loss plateau |
| Loss Function | CrossEntropyLoss | Standard for binary classification |
| Gradient Clipping | 1.0 | Prevents exploding gradients |
| Data Augmentation | HorizontalFlip (p=0.5), Rotation (±10°), ColorJitter | Improves generalization |
Maintainers: @Gazeux33 Repository: github.com/VeridisQuo-orga/VeridisQuo Contact: For questions or collaborations, open an issue on GitHub