Downloads · 30 days
0
BiliSakura/S5-transformers
S5-transformers is a feature extraction model from BiliSakura. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
Hugging Face–compatible checkpoints converted from the official S5 S4P pretrain weights. Each subfolder is a standalone model repo layout (config.json, model.safetensors, preprocessor, and remote code) for encoder fea…
Downloads · 30 days
0
Access
Public
Updated Aug 23, 2026
Repo size
1.6 GB
Likes
0
Public
Click a slice to open those files.
.safetensors1.6 GB · 100%
From the Hugging Face model README
Hugging Face–compatible checkpoints converted from the official S5 S4P pretrain weights. Each subfolder is a standalone model repo layout (config.json, model.safetensors, preprocessor, and remote code) for encoder feature extraction on optical remote sensing imagery.
These are ViT encoders pretrained with S4P (Semi-supervised Semantic Segmentation Pre-training) on RS4P-1M. This collection currently bundles 2 converted backbone checkpoints:
Both checkpoints use architecture: s4p_backbone and expose the s5-feature-extraction pipeline. They are encoder-only weights (not UPerNet segmentation or MoE-MDF heads).
All folders ship self-contained remote code (modeling_s5.py, processor, pipeline) and load with trust_remote_code=True.
Developed by: lianglyu / S5
Converted for Hugging Face by: BiliSakura
License (weights): Apache 2.0
Original paper: S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing (AAAI 2026 Oral)
| Folder | Backbone | Hidden size | Layers | Heads | Patch | Image size | Original file |
|---|---|---|---|---|---|---|---|
ViT-B | ViT-Base | 768 | 12 | 12 | 16 | 512 | vit_b_s4p.pth |
ViT-L | ViT-Large | 1024 | 24 | 16 | 16 | 512 | vit_l_s4p.pth |
Original singular .pth files were converted and removed from this directory.
Processors default to do_resize: false. Pass RGB images at native resolution; ImageNet mean/std normalization is applied when enabled.
from transformers import pipeline
import numpy as np
REPO = "/path/to/S5-transformers"
pipe = pipeline(
task="s5-feature-extraction",
model=f"{REPO}/ViT-B",
trust_remote_code=True,
)
image = np.random.randint(0, 255, (512, 512, 3), dtype=np.uint8)
# Global pooled features
features = pipe(image, pool=True, return_tensors=True)
print(features.shape) # [1, 768] for ViT-B, [1, 1024] for ViT-L
# Dense feature map
featmap = pipe(image, pool=False, return_tensors=True)
print(featmap.shape) # [1, 768, 32, 32] for ViT-B, [1, 1024, 32, 32] for ViT-L
ViT-L:
pipe = pipeline(
task="s5-feature-extraction",
model=f"{REPO}/ViT-L",
trust_remote_code=True,
)
features = pipe(image, pool=True, return_tensors=True)
print(features.shape) # [1, 1024]
To force 512×512 resize:
features = pipe(
image,
pool=True,
return_tensors=True,
image_processor_kwargs={"do_resize": True},
)
Load components directly:
from transformers import AutoModel, AutoImageProcessor
model = AutoModel.from_pretrained(f"{REPO}/ViT-B", trust_remote_code=True)
processor = AutoImageProcessor.from_pretrained(f"{REPO}/ViT-B", trust_remote_code=True)
The bundled image processor applies ImageNet mean/std normalization by default (do_normalize=True, rescale_factor=1/255). Inputs should be RGB optical imagery.
Checkpoints were converted with scripts/convert_s5_checkpoint.py from the official release:
python scripts/convert_s5_checkpoint.py \
--input-path /path/to/vit_b_s4p.pth \
--output-dir /path/to/ViT-B \
--clean-output
For semantic segmentation heads, convert vit_*_s4p_upernet.pth or s5_vit_*_moe_mdf_seg.pth instead (task: s5-semantic-segmentation).
transformers>=4.45.0torch>=2.1.0safetensorsPillownumpy@article{S5,
title={S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing},
author={Liang Lv and Di Wang and Jing Zhang and Lefei Zhang},
journal={arXiv preprint arXiv:2508.12409},
year={2025}
}