Downloads · 30 days
0
haeylee/ssl_ft_pron
ssl_ft_pron is a feature extraction model from haeylee. Use it when you need embeddings to search or compare text.
A collection of fine-tuned Self-Supervised Learning (SSL) speech models (Wav2Vec2.0, HuBERT, WavLM) for Automatic Pronunciation Assessment (APA). Three strategies are provided per backbone:
Downloads · 30 days
0
Access
Public
Updated Sep 24, 2025
Repo size
159 GB
Likes
2
Public
Click a slice to open those files.
.bin53 GB · 33%
From the Hugging Face model README
A collection of fine-tuned Self-Supervised Learning (SSL) speech models (Wav2Vec2.0, HuBERT, WavLM) for Automatic Pronunciation Assessment (APA).
Three strategies are provided per backbone:
Important: This Hub repository is a collection. Each model lives in a subdirectory.
Load with the full sub-path, e.g.haeylee/ssl_ft_pron/wav2vec2/general/02_wav2vec2-large-960h.
base_model list abovefrom transformers import AutoModelForCTC, AutoProcessor
ckpt = "haeylee/ssl_ft_pron/wav2vec2/ctc/01_wav2vec2-large"
model = AutoModelForCTC.from_pretrained(ckpt)
processor = AutoProcessor.from_pretrained(ckpt)
from transformers import AutoProcessor, Wav2Vec2Model, HubertModel, WavLMModel
# Wav2Vec2 (General)
ckpt = "haeylee/ssl_ft_pron/wav2vec2/general/01_wav2vec2-large"
model = Wav2Vec2Model.from_pretrained(ckpt)
processor = AutoProcessor.from_pretrained(ckpt)
# HuBERT (Freeze)
# ckpt = "haeylee/ssl_ft_pron/hubert/freeze/06_hubert-large-ll60k"
# model = HubertModel.from_pretrained(ckpt)
# processor = AutoProcessor.from_pretrained(ckpt)
# WavLM (General)
# ckpt = "haeylee/ssl_ft_pron/wavlm/general/10_wavlm-large"
# model = WavLMModel.from_pretrained(ckpt)
# processor = AutoProcessor.from_pretrained(ckpt)
Summary:
AutoModelForCTC.from_pretrained(...)Wav2Vec2Model / HubertModel / WavLMModel .from_pretrained(...)preprocess_dataset.py (see the GitHub repo) to convert raw audio/labels into Hugging Face datasets format.Expected processed layout:
/your/data/path/speechocean762/
└── preprocess/
├── speechocean_train_ds/
└── speechocean_test_ds/
# Adjust paths inside the script or via CLI args
python preprocess_dataset.py \
--data_root /your/data/path/speechocean762 \
--out_dir /your/data/path/speechocean762/preprocess
Loads encoders with Wav2Vec2Model / HubertModel / WavLMModel .from_pretrained(...) and trains a regression head to predict 4 APA scores.
python train/baseline.py \
--model_name facebook/hubert-xlarge-ls960-ft \
--batch_size 4 \
--learning_rate 1e-5 \
--num_train_epochs 30
Same as General, but freezes the CNN feature extractor.
python train/freeze.py \
--model_name facebook/hubert-xlarge-ls960-ft \
--freeze_feature_extractor \
--batch_size 4 \
--learning_rate 1e-5 \
--num_train_epochs 30
Uses AutoModelForCTC.from_pretrained(...) for CTC training.
python train/ctc.py \
--model_name facebook/wav2vec2-large \
--batch_size 4 \
--learning_rate 1e-5 \
--num_train_epochs 30
Artifacts saved: model.safetensors, trainer_state.json, training_args.bin, logs, and checkpoints (per run: args.json, trainer_args.json).
preprocess_dataset.py)pearsonr (Pearson correlation coefficient, PCC) for Accuracy, Fluency, Prosody, and Total.@inproceedings{lee2024analysis,
title={Analysis of Various Self-Supervised Learning Models for Automatic Pronunciation Assessment},
author={Lee, Haeyoung and Kim, Sunhee and Chung, Minhwa},
booktitle={2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)},
pages={1--6},
year={2024},
organization={IEEE}
}