Downloads · 30 days
0
slprl/WhiStress
WhiStress is a automatic speech recognition model from slprl. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
This is the official model checkpoint for WhiStress — introduced in our paper: WhiStress: Enriching Transcriptions with Sentence Stress Detection (Interspeech 2025).
Downloads · 30 days
0
Access
Public
Updated May 28, 2025
Repo size
42.5 MB
Likes
6
Public
Click a slice to open those files.
.pt42.5 MB · 100%
From the Hugging Face model README
This is the official model checkpoint for WhiStress — introduced in our paper:
WhiStress: Enriching Transcriptions with Sentence Stress Detection (Interspeech 2025).
WhiStress extends OpenAI's Whisper ASR model with a decoder-based classifier that predicts token-level sentence stress. This allows models not only to transcribe speech but also to detect which words are emphasized.
This checkpoint is based on the whisper-small.en variant and adds two stress-specific modules:
additional_decoder_block.ptclassifier.ptYou can use the weights in your own pipeline by cloning our codebase and loading the components:
git clone https://github.com/slp-rl/WhiStress.git
cd WhiStress
pip install -r requirements.txt
Then, either download the weights manually from this Hugging Face repo or use our script:
python download_weights.py
The weights should be placed in the following directory structure:
whistress/
├── weights/
│ ├── additional_decoder_block.pt
│ ├── classifier.pt
│ └── metadata.json
from whistress import WhiStressInferenceClient
whistress_client = WhiStressInferenceClient(device="cuda") # or "cpu"
pred_transcription, pred_stresses = whistress_client.predict(
audio=sample['audio'], # (sr, np.ndarray)
transcription=None, # predict directly from audio both transcription and stress, pass transcription to predict stress only.
return_pairs=False # set to True if you a list want a list of (word, binary_label) pairs.
)
print(pred_transcription) # e.g., "I didn’t say she stole my money."
print(pred_stresses) # e.g., ['my']
Each prediction includes:
transcription: full text outputemphasis_indices: list of stressed token indicesemphasized_tokens: list of corresponding wordsThe model is intended for research purposes only.
If you use our model, please cite our work:
@misc{yosha2025whistress,
title={WHISTRESS: Enriching Transcriptions with Sentence Stress Detection},
author={Iddo Yosha and Dorin Shteyman and Yossi Adi},
year={2025},
eprint={2505.19103},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2505.19103},
}