Downloads · 30 days
23
3% of all-time downloads
hhsavich/accent_determinator
accent_determinator is a audio classification model from hhsavich. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for transformers.
Wav2Vec2 Model to classify audio based on the accent of the speaker as Puerto Rican, Colombian, Venezuelan, Peruvian, or Chilean
Downloads · 30 days
23
3% of all-time downloads
All-time downloads
808
Public
Repo size
757 MB
Likes
2
Public
Click a slice to open those files.
.bin378 MB · 100%
From the Hugging Face model README
Wav2Vec2 Model to classify audio based on the accent of the speaker as Puerto Rican, Colombian, Venezuelan, Peruvian, or Chilean
Wav2Vec2 Model to classify audio based on the accent of the speaker as Puerto Rican, Colombian, Venezuelan, Peruvian, or Chilean
Classify an audio clip as Puerto Rican, Peruvian, Venezuelan, Colombian, or Chilean Spanish
The model was trained on speakers reciting pre-chosen sentences, thus it does not reflect any knowledge of lexical differences between dialects.
Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Predictions generated by the model may include disturbing and harmful stereotypes across protected classes; identity characteristics; and sensitive, social, and occupational groups.
OpenSLR 71,72,73,74,75,76
Data was Train-Test split on speakers, so as to prevent the model from achieving high test accuracy by matching voices.
Trained on ~3000 5-second audio clips, Training is lightwegiht taking < 1 hr on using Google Colaboratory Premium GPUs
OpenSLR 71,72,73,74,75,76 https://huggingface.co/datasets/openslr
Audio Quality - training and testing data was higher quality than can be expected from found audio
Accuracy
~85% depending on random train-test split
Even splitting on speakers, our model achieves excellent accuracy on the testing set. This is interesting because it indicates that accent classification, at least at this granularity, is an easier task than voice identification, which could have just as easily met the training objective.
The confusion matrix shows that Basque is the most easily distinguished, which should be expecting as it is the only language that isn't Spanish. Puerto Rican was the hardest to identify in the testing set, but I think this is more having to do with PR having the least data moreso than something about the accent itself.
I think if this same size of dataset was used for this same experiment, but there were more speakers (and so not as much fitting on individual voices), we could expect near perfect accuracy.
Wav2Vec2
Google Colaboratory Pro+
Google Colaboratory Pro+ Premium GPUS
Pytorch via huggingface
Henry Savich