Downloads · 30 days
3
13% of all-time downloads
XAFT/SM-Selective-ViT-Base-224
SM-Selective-ViT-Base-224 is a image classification model from XAFT. Use it when you need a label for an image. It is set up for transformers. The card lists the license as mit.
Soft-Masked Selective Vision Transformer is an efficient Vision Transformer (ViT) model designed to reduce the computational overhead of self-attention while maintaining competitive accuracy. The model introduces a pa…
Downloads · 30 days
3
13% of all-time downloads
All-time downloads
24
Public
Parameters
86.6M
346 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors346 MB · 100%
From the Hugging Face model README
Soft-Masked Selective Vision Transformer is an efficient Vision Transformer (ViT) model designed to reduce the computational overhead of self-attention while maintaining competitive accuracy.
The model introduces a patch-selective attention mechanism that enables the transformer to focus on the most salient image regions and dynamically disregard less informative patches. This selective strategy significantly reduces the quadratic complexity typically associated with full self-attention, making the model particularly suitable for high-resolution vision tasks and resource-constrained environments.
To further improve performance, the model leverages knowledge distillation, transferring representational knowledge from a stronger teacher network to enhance the accuracy of lightweight transformer variants.
This model is intended for:
import torch
from transformers import AutoModelForImageClassification, AutoImageProcessor
from PIL import Image
import requests
# Load image
url = "http://images.cocodataset.org/val2017/000000039769.jpg"
image = Image.open(requests.get(url, stream=True).raw)
# Load processor and model
processor = AutoImageProcessor.from_pretrained(
"XAFT/SM-Selective-ViT-Base-224",
trust_remote_code=True,
)
model = AutoModelForImageClassification.from_pretrained(
"XAFT/SM-Selective-ViT-Base-224",
trust_remote_code=True,
)
model = model.half() # Cast to FP16 to enable FlashAttention
# Preprocess
inputs = processor(
images=image,
return_tensors="pt",
)
inputs = inputs.to(torch.half) # Cast to FP16
# Forward pass
outputs = model(**inputs)
logits = outputs.logits
predicted_class = logits.argmax(-1).item()
print("Predicted class index:", predicted_class)
| Model | Top-1 Acc. | Top-5 Acc. | # Params | Avg. GFLOPs |
|---|---|---|---|---|
| Base | 80.350% | 94.980% | 86.60M | 9.61 |
| Base (distilled) | 80.990% | 95.386% | 87.37M | 9.21 |
| Small | 78.662% | 94.454% | 22.06M | 3.12 |
| Small (distilled) | 79.000% | 94.494% | 22.45M | 3.05 |
| Tiny tall | 74.802% | 92.794% | 11.07M | 1.64 |
| Tiny tall (distilled) | 75.676% | 92.988% | 11.26M | 1.64 |
| Tiny | 71.056% | 90.192% | 5.72M | 0.95 |
| Tiny (distilled) | 72.618% | 91.338% | 5.92M | 0.93 |
We thank the TPU Research Cloud program for providing cloud TPUs that were used to build and train the models for our extensive experiments.
If you find our work helpful, feel free to give us a cite.
@article{TOULAOUI2026115151,
title = {Efficient vision transformers via patch selective soft-masked attention and knowledge distillation},
journal = {Applied Soft Computing},
pages = {115151},
year = {2026},
issn = {1568-4946},
doi = {https://doi.org/10.1016/j.asoc.2026.115151},
url = {https://www.sciencedirect.com/science/article/pii/S1568494626005995},
author = {Abdelfattah Toulaoui and Hamza Khalfi and Imad Hafidi},
keywords = {Vision transformer, Patch selection, Soft masking, Efficient inference}
}