Downloads · 30 days
0
feqhwjBBA/Deepfake-CLIP
Deepfake-CLIP is a image classification model from feqhwjBBA. Use it when you need a label for an image. It is set up for adapter-transformers. The card lists the license as afl-3.0.
A cutting-edge deepfake detection framework that integrates CLIP Vision-Language models with Parameter-Efficient Fine-Tuning (PEFT) and Retrieval-Augmented Generation (RAG) inspired techniques for enhanced detection p…
Downloads · 30 days
0
Access
Public
Updated Jan 7, 2026
Repo size
1.8 GB
Likes
0
Public
Click a slice to open those files.
.h5606 MB · 33%
From the Hugging Face model README
A cutting-edge deepfake detection framework that integrates CLIP Vision-Language models with Parameter-Efficient Fine-Tuning (PEFT) and Retrieval-Augmented Generation (RAG) inspired techniques for enhanced detection performance across multiple benchmark datasets.
This repository implements a novel deepfake detection approach that extends the CLIP (Contrastive Language-Image Pre-training) model through several key innovations:
The framework achieves state-of-the-art performance on multiple deepfake detection benchmarks while maintaining computational efficiency through selective parameter updates.
| Innovation | Description | Key Benefit |
|---|---|---|
| PEFT with LoRA | Low-rank adaptation of CLIP transformer layers | 90%+ parameter reduction, efficient fine-tuning |
| Learnable Text Prompts | Adaptive text feature learning instead of fixed prompts | Dataset-specific textual representations |
| Hard Negative Mining | Focus on challenging misclassification cases | Improved discrimination at decision boundaries |
| Memory-Augmented Contrastive | RAG-inspired feature retrieval and augmentation | Enhanced generalization through memory |
| Knowledge-Augmented Prompts | Dynamic text prompt enhancement with retrieved knowledge | Context-aware textual representations |
# Core dependencies
pip install torch torchvision transformers
# PEFT for parameter-efficient fine-tuning
pip install peft
# Additional utilities
pip install scikit-learn tqdm Pillow pyyaml
# For development
pip install black flake8 mypy
deepfake-detection/
├── train_cvpr2025.py # Main training and evaluation script
├── config/
│ └── detector/
│ └── cvpr2025.yaml # Configuration file
├── checkpoints/ # Saved model weights
├── datasets/ # Dataset storage (symlinked)
└── results/ # Evaluation results
The system uses YAML configuration for all experiment settings. Key configuration sections:
model:
base_model: "CLIP-ViT-B-32" # or "CLIP-ViT-L-14"
use_peft: true
lora_rank: 16
lora_alpha: 16
lora_dropout: 0.1
training:
nEpochs: 50
batch_size: 32
optimizer: "adam"
learning_rate: 1e-4
temperature: 0.07 # Contrastive learning temperature
innovations:
use_learnable_prompts: true
use_hard_mining: true
use_memory_augmented: true
use_knowledge_augmented_prompts: true
# Basic training with default configuration
python train_cvpr2025.py
# With custom configuration
python train_cvpr2025.py --config path/to/custom_config.yaml
# Specify experiment name
python train_cvpr2025.py --experiment_name "ff++_lora_experiment"
# Evaluate a saved checkpoint
python -c "from train_cvpr2025 import test_with_loaded_weights; test_with_loaded_weights('checkpoints/best_lora_weights.pth')"
# With custom config
python -c "from train_cvpr2025 import test_with_loaded_weights; test_with_loaded_weights('checkpoints/best.pth', 'config/custom.yaml')"
To add a new dataset:
train_dataset or test_dataset listsThe system uses PEFT's LoRA implementation for efficient fine-tuning:
from peft import LoraConfig, get_peft_model
lora_config = LoraConfig(
r=16, # LoRA rank
lora_alpha=16,
target_modules=["q_proj", "v_proj"], # Attention layers to adapt
lora_dropout=0.1,
bias="none",
task_type=TaskType.FEATURE_EXTRACTION,
)
model = get_peft_model(clip_model, lora_config)
Inspired by RAG, this component retrieves similar features from memory banks to augment contrastive learning:
class MemoryBank:
def retrieve(self, query_feat, k=5):
# Retrieve k most similar features
similarities = query_feat @ self.memory.t()
_, indices = torch.topk(similarities, k)
return self.memory[indices]
Text prompts are dynamically enhanced with retrieved knowledge from training:
class KnowledgeAugmentedTextPrompts:
def forward(self, img_feat):
# Retrieve relevant knowledge
real_knowledge, fake_knowledge = self.knowledge_bank.retrieve(img_feat)
# Augment base prompts with retrieved knowledge
enhanced_real = self.fusion(base_real_prompt, real_knowledge)
enhanced_fake = self.fusion(base_fake_prompt, fake_knowledge)
return enhanced_real, enhanced_fake
The system provides comprehensive evaluation:
The training script includes progress tracking:
# Training progress with tqdm
for images, labels in tqdm(train_loader, desc=f"Epoch {epoch}"):
# Training loop
pass
# Real-time metrics display
print(f"[Eval] {dataset_name}: AUC={auc:.4f} AP={ap:.4f}")
Comprehensive logging is built-in:
# File existence checks
if not os.path.exists(full_path):
print(f"[Warning] Image not found: {full_path}")
# Memory bank statistics
print(f"[Memory] Real samples: {real_size}, Fake samples: {fake_size}")
# Training progress
print(f"[Train] Epoch {epoch}: loss={loss:.4f}, lr={lr:.6f}")
If you use this code in your research, please cite:
@article{deepfake2025clip,
title={CLIP-Enhanced Deepfake Detection with RAG-Inspired Memory Augmentation},
author={Your Name},
journal={CVPR},
year={2025}
}
We welcome contributions! Please:
This project is licensed under the MIT License - see the LICENSE file for details.
For questions, issues, or collaborations:
Note: This implementation is research-oriented and may require adjustments for production deployment. Always validate performance on your specific use case and datasets.
For the latest updates and bug fixes, check the Releases page.