Downloads · 30 days
0
FatemehBehrad/Charm
Charm is a image feature extraction model from FatemehBehrad. Use it for the image feature extraction task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
💫 Official implementation of Charm: The Missing Piece in ViT fine-tuning for Image Aesthetic Assessment
Downloads · 30 days
0
Access
Public
Updated May 7, 2025
Repo size
1.7 GB
Likes
1
Public
Click a slice to open those files.
.pth1.7 GB · 100%
From the Hugging Face model README
💫 Official implementation of Charm: The Missing Piece in ViT fine-tuning for Image Aesthetic Assessment
<div align="left"> <a href="https://github.com/FBehrad/Charm"> <img src="https://github.com/FBehrad/Charm/blob/main/Figures/MainFigure.jpg?raw=true" alt="Overall framework" width="400"/> </a> </div>Accepted at CVPR 2025<br> arXiv<br>
We introduce Charm , a novel tokenization approach that preserves Composition, High-resolution, Aspect Ratio, and Multi-scale information simultaneously. By preserving critical information, <em> Charm </em> works like a charm for image aesthetic and quality assessment 🌟.
pip install -r requirements.txt
pip install Charm-tokenizer
Charm Tokenizer has the following input args:
Note: While random patch selection during training helps avoid overfitting,for consistent results during inference, fully deterministic patch selection approaches should be used.
The output is the preprocessed tokens, their corresponding positional embeddings, and a mask token that indicates which patches are in high resolution and which are in low resolution.
from Charm_tokenizer.ImageProcessor import Charm_Tokenizer
img_path = r"img.png"
charm_tokenizer = Charm_Tokenizer(patch_selection='frequency', training_dataset='tad66k',backbone='facebook/dinov2-small', without_pad_or_dropping=True)
tokens, pos_embed, mask_token = charm_tokenizer.preprocess(img_path)
Step 4) Predicting aesthetic/quality score
If training_dataset is set to 'spaq' or 'koniq10k', the model predicts the image quality score. For other options ('aadb', 'tad66k', 'para', 'baid'), it predicts the image aesthetic score.
Selecting a dataset with image resolutions similar to your input images can improve prediction accuracy.
For more details about the process, please refer to the paper.
from Charm_tokenizer.Backbone import backbone
model = backbone(training_dataset='tad66k', device='cpu')
prediction = model.predict(tokens, pos_embed, mask_token)
Note: For the training code, check our GitHub Page.