Downloads · 30 days
13
25% of all-time downloads
scribe-project/nb-whisper-dialect-id-4dialect
nb-whisper-dialect-id-4dialect is a audio classification model from scribe-project. Use it for the audio classification task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
A NbAiLab/nb-whisper-medium encoder fine-tuned for 4-class Norwegian dialect identification on unmodified (natural) speech.
Downloads · 30 days
13
25% of all-time downloads
All-time downloads
51
Public
Parameters
307M
615 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors615 MB · 100%
From the Hugging Face model README
A NbAiLab/nb-whisper-medium encoder fine-tuned for 4-class Norwegian dialect identification on unmodified (natural) speech.
Given an audio clip of Norwegian speech, the model predicts which of four broad dialect regions the speaker belongs to:
| Label | Region |
|---|---|
east | Eastern Norwegian (østnorsk) |
west | Western Norwegian (vestnorsk) |
mid | Central/Trøndersk Norwegian (trøndersk) |
north | Northern Norwegian (nordnorsk) |
This model is one of three "prosody-condition" models (unmodified / low-pass / monotonize) trained for:
Phoebe Parsons, Heming Strømholt Bremnes, Knut Kvale, Torbjørn Svendsen, and Giampiero Salvi. (2025). Effects of Prosodic Information on Dialect Classification Using Whisper Features. In Proceedings of Interspeech 2025, pages 2785–2789. doi: 10.21437/Interspeech.2025-200
The other two conditions are trained on the same data with the speech signal manipulated to isolate or remove prosodic (F0) cues:
scribe-project/nb-whisper-dialect-id-4dialect-low-pass — low-pass filtered audioscribe-project/nb-whisper-dialect-id-4dialect-monotonize — F0-monotonized audioSee the did_prosody_whisper repo for the training/evaluation code used for the paper.
The model is NbAiLab/nb-whisper-medium's encoder with a classification head on top
(WhisperForAudioClassification), fully fine-tuned (no layers frozen) for sequence classification
over 4 dialect labels.
Intended for research on Norwegian dialect identification, and specifically as the "unmodified audio" baseline against which the low-pass and monotonize conditions are compared to study the role of prosody in dialect classification. It is not intended for consequential decisions about individuals (e.g. hiring, legal, or identity-verification contexts).
Known limitations:
import torch
from transformers import AutoFeatureExtractor, AutoModelForAudioClassification
import librosa
model_id = "scribe-project/nb-whisper-dialect-id-4dialect"
feature_extractor = AutoFeatureExtractor.from_pretrained(model_id)
model = AutoModelForAudioClassification.from_pretrained(model_id)
audio, sr = librosa.load("path/to/audio.wav", sr=feature_extractor.sampling_rate, mono=True)
inputs = feature_extractor(
audio,
sampling_rate=feature_extractor.sampling_rate,
return_tensors="pt",
)
with torch.no_grad():
logits = model(**inputs).logits
predicted_id = torch.argmax(logits, dim=-1).item()
print(model.config.id2label[predicted_id])
Audio should be mono, resampled to 16kHz. Clips longer than 30 seconds were randomly subsampled to 30 seconds during training.
Trained on the SSC (Storting/Parliament Speech Corpus), using a speaker-disjoint train/validation split (no speaker overlap between train and eval). Each example is a single-speaker audio segment labeled with one of the 4 dialect categories above. Audio was used unmodified (no low-pass filtering or monotonization).
| Training Loss | Epoch | Step | Validation Loss | Accuracy |
|---|---|---|---|---|
| 0.0315 | 1.0 | 453 | 0.2778 | 0.9318 |
| 0.0139 | 2.0 | 906 | 0.7021 | 0.8716 |
| 0.0055 | 3.0 | 1359 | 0.7993 | 0.8535 |
The weights published in this repo are from epoch 3 (checkpoint-1359), the final training checkpoint — not the epoch-1 checkpoint, which had the highest single-epoch eval accuracy but is inconsistent with the checkpoint selection used for the low-pass and monotonize conditions above.