Downloads · 30 days
0
Jarvis1111/MiniGPT4-RobustVLGuard
MiniGPT4-RobustVLGuard is a image-text-to-text model from Jarvis1111. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
Welcome! This repository hosts the official implementation of our paper, "Safeguarding Vision-Language Models: Mitigating Vulnerabilities to Gaussian Noise in Perturbation-based Attacks."
Downloads · 30 days
0
Access
Public
Updated Apr 5, 2025
Repo size
196 MB
Likes
0
Public
Click a slice to open those files.
.pth196 MB · 100%
From the Hugging Face model README
Welcome! This repository hosts the official implementation of our paper, "Safeguarding Vision-Language Models: Mitigating Vulnerabilities to Gaussian Noise in Perturbation-based Attacks."
Paper link: arxiv.org/abs/2504.01308
Project page:
We propose state-of-the-art solutions to enhance the robustness of Vision-Language Models (VLMs) against Gaussian noise and adversarial attacks. Key highlights include:
🎯 Robust-VLGuard: A pioneering multimodal safety dataset covering both aligned and misaligned image-text pair scenarios.
🛡️ DiffPure-VLM: A novel defense framework that leverages diffusion models to neutralize adversarial noise by transforming it into Gaussian-like noise, significantly improving VLM resilience.