Downloads · 30 days
24
16% of all-time downloads
jiachengcui888/PIXAR-7B
PIXAR-7B is a machine learning model from jiachengcui888. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
PIXAR-7B is a Vision-Language Model (VLM) for image tampering analysis, introduced in the paper "From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering".
Downloads · 30 days
24
16% of all-time downloads
All-time downloads
153
Public
Repo size
32.5 GB
Likes
0
Public
Click a slice to open those files.
.bin16.3 GB · 100%
From the Hugging Face model README
PIXAR-7B is a Vision-Language Model (VLM) for image tampering analysis, introduced in the paper "From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering".
Given a query image, PIXAR-7B jointly performs:
PIXAR-7B is built on a LLaVA + LLaMA-2 backbone with LoRA fine-tuning (rank 8), integrated with SAM ViT-H for pixel-level decoding and CLIP ViT-L/14 for visual-language alignment. Three special tokens are inserted into the token sequence to anchor multi-task prediction heads:
| Token | Role |
|---|---|
[CLS] | 3-way classification (real / fully synthetic / tampered) via a linear head |
[OBJ] | Multi-label object recognition over 81 COCO categories via a linear head |
[SEG] | Pixel-level segmentation mask generation via SAM, optionally fused with the generated text embedding |
Existing tampering benchmarks rely on coarse object masks as ground truth, which conflates unedited pixels inside the mask with actual tamper evidence and misses subtle edits outside the mask. PIXAR replaces binary masks with per-pixel difference maps $D = |I_\text{orig} - I_\text{gen}|$, thresholded at a tunable $\tau$ to produce dynamic ground truth $M_\tau$ that captures edits at multiple scales.
The PIXAR benchmark provides:
PIXAR-7B achieves 2.6× IoU improvement over prior state of the art on the PIXAR benchmark.
For interactive inference, see the project repository:
python chat.py --version jiachengcui888/PIXAR-7B --precision bf16 --seg_prompt_mode seg_only
The PIXAR benchmark — 420K+ image pairs with pixel-faithful tamper labels, semantic categories, and natural language descriptions, spanning 8 manipulation types across COCO-category objects.
Fine-tuned with DeepSpeed on a LLaVA + LLaMA-2 backbone using LoRA (rank 8). Key hyperparameters:
| Hyperparameter | Value |
|---|---|
| LoRA rank | 8 |
| Learning rate | 1e-4 |
| Batch size | 2 |
| Precision | bf16 |
| Threshold τ | 0.05 |
| text | 3.0 |
| cls | 1.0 |
| bce | 1.0 |
| dice | 1.0 |
| sem | 0.5 |
PIXAR-7B achieves 2.6× IoU improvement over prior SOTA on the PIXAR test benchmark. For full evaluation results, refer to the paper.