Downloads · 30 days
0
JiayuMBao/SE-AlexNet
SE-AlexNet is a machine learning model from JiayuMBao. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A collection of 34 fine-tuned convolutional neural networks for studying how Squeeze-and-Excitation (SE) modules affect facial emotion recognition. This model zoo supports the psychophysical analysis pipeline describe…
Downloads · 30 days
0
Access
Public
Updated Sep 5, 2026
Repo size
8.5 GB
Likes
0
Public
Click a slice to open those files.
.safetensors8.5 GB · 100%
From the Hugging Face model README
A collection of 34 fine-tuned convolutional neural networks for studying how Squeeze-and-Excitation (SE) modules affect facial emotion recognition. This model zoo supports the psychophysical analysis pipeline described in the companion paper.
📦 GitHub (Code & Analysis): SE-AlexNet | PsychometricFittingCurve | GradCAM-ROI-SaliencyMapDecoder
This repository contains weights for 5 model architectures, systematically varied across 3 experimental axes:
| Experimental Axis | Values |
|---|---|
| Architecture | AlexNet, VGG16, SE-AlexNet-L1, SE-AlexNet-L2, SE-AlexNet-L3 |
| SE Reduction Ratio ($r$) | 2, 4, 8, 16, 32 (applies to SE variants) |
| Pre-training Basis | FaceBased (VGGFace2) or ObjectBased (ImageNet) |
| Model | SE Position | Key Characteristic |
|---|---|---|
| AlexNet | None (baseline) | Standard 5-conv AlexNet, num_classes=11 |
| SE-AlexNet-L1 | After last conv, before FC stack | SE on 256-channel feature maps |
| SE-AlexNet-L2 | Between fc6 → fc7 | SE on 4096-dim feature vector, reduction fixed at 16 |
| SE-AlexNet-L3 | As first element of classifier | Best performing variant — SELayer before fc6 |
| VGG16 | None (benchmark) | Standard VGG16, num_classes=2 (binary Happy/Sad) |
⚠️ Important note on SE-Location-2: All 10 L2 variants share identical architecture (reduction=16). The
squeeze-{r}label refers to a training/data configuration, not the architectural reduction ratio. This is preserved for reproducibility.
These models are designed for visual feature interpretability analysis in facial emotion recognition research. Specific use cases:
Not intended for: Production emotion recognition systems, clinical diagnosis, or surveillance applications. These are research models trained on controlled lab datasets.
All models were fine-tuned on AffectNet, the largest facial expression dataset:
Pre-training sources:
Training hyperparameters:
| Metric | Finding |
|---|---|
| Best Architecture | SE-AlexNet-L3 (SE block closest to classifier) |
| Optimal Reduction | $r=32$ consistently outperformed lower reductions |
| Pre-training Effect | Face-based pre-training improved emotion discrimination by ~12% over object-based |
| ROI Attention | SE models showed more focused attention on mouth and eye regions vs. baseline AlexNet |
Models were evaluated against human observers ($N=40$) in a 2AFC Happy/Sad discrimination task:
pip install -r requirements.txt
from inference import SEModelPipeline
# Load the best model (SE-Location3, FaceBased, r=32)
pipe = SEModelPipeline('se-location3/facebased/squeeze-32')
# Run inference
probs = pipe.predict('path/to/face_image.jpg')
print(f'Prediction shape: {probs.shape}') # (1, 11)
import os, json
for root, dirs, files in os.walk('.'):
if 'config.json' in files:
with open(os.path.join(root, 'config.json')) as f:
cfg = json.load(f)
print(f"{root}: {cfg['model_type']} | {cfg['pretraining']} | {cfg.get('reduction', 'N/A')}")
import json
from modeling import load_model_from_config
# Load config
with open('se-location3/facebased/squeeze-32/config.json') as f:
config = json.load(f)
# Build model with weights
model = load_model_from_config(
config,
weights_path='se-location3/facebased/squeeze-32/model.safetensors',
device='cpu'
)
# Forward pass
import torch
x = torch.randn(1, 3, 224, 224) # dummy input
output = model(x)
from modeling import load_model_from_config
import json
# Load any model
with open('se-location3/facebased/squeeze-32/config.json') as f:
config = json.load(f)
model = load_model_from_config(
config,
'se-location3/facebased/squeeze-32/model.safetensors'
)
# Target the last conv layer for Grad-CAM
target_layer = model.features[-3] # Last Conv2d in features
# ... apply standard Grad-CAM pipeline
SE-AlexNet/
├── README.md # This Model Card
├── requirements.txt # Python dependencies
├── modeling.py # Exact model architecture definitions
├── inference.py # Universal inference pipeline
├── convert_weights.py # .pth → .safetensors conversion script
├── generate_configs.py # Config file generator
│
├── alexnet/ # Standard AlexNet baseline
│ ├── facebased/
│ │ ├── config.json
│ │ └── model.safetensors
│ └── objectbased/
│ ├── config.json
│ └── model.safetensors
│
├── vgg16/ # VGG16 benchmark
│ ├── facebased/
│ └── objectbased/
│
├── se-location1/ # SE after all convolutions
│ ├── facebased/
│ │ ├── squeeze-2/ → squeeze-32/ (5 reduction ratios)
│ └── objectbased/
│ └── squeeze-2/ → squeeze-32/
│
├── se-location2/ # SE between fc6 → fc7
│ ├── facebased/
│ └── objectbased/
│
└── se-location3/ # SE in classifier (BEST)
├── facebased/
└── objectbased/
| Repository | Role | Link |
|---|---|---|
| SE-AlexNet | Training code, ablation scripts, raw results | GitHub |
| PsychometricFittingCurve | MATLAB psychometric curve fitting, PSE calculation, human comparison | GitHub |
| GradCAM-ROI-SaliencyMapDecoder | Grad-CAM heatmap generation & ROI statistical decoding | GitHub |
SE-AlexNet extends the classic AlexNet by inserting Squeeze-and-Excitation blocks at strategic locations:
The three location variants test the hypothesis that SE blocks are most effective when placed closer to the decision boundary (classifier), where channel-wise feature importance directly impacts classification.
MIT License. See companion GitHub repositories for full license details.
Trained weights converted from original PyTorch .pth checkpoints. For questions about the training methodology or experimental design, please refer to the companion paper and GitHub repositories.