Downloads · 30 days
9
2% of all-time downloads
keras-io/speaker-recognition
speaker-recognition is a machine learning model from keras-io. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for tf-keras.
This model helps to classify speakers from the frequency domain representation of speech recordings, obtained via Fast Fourier Transform (FFT). The model is created by a 1D convolutional network with residual connecti…
Downloads · 30 days
9
2% of all-time downloads
All-time downloads
521
Public
Repo size
14.7 MB
Likes
9
Public
Click a slice to open those files.
.data-00000-of-0000112.4 MB · 84%
From the Hugging Face model README
This model helps to classify speakers from the frequency domain representation of speech recordings, obtained via Fast Fourier Transform (FFT). The model is created by a 1D convolutional network with residual connections for audio classification.
This repo contains the model for the notebook Speaker Recognition.
Full credits go to Fadi Badine
This model uses a speaker recognition dataset of Kaggle
This should be run with TensorFlow 2.3 or higher, or tf-nightly.
Also, The noise samples in the dataset need to be resampled to a sampling rate of 16000 Hz before using for this model so, In order to do this, you will need to have installed ffmpg.
During dataset preparation, the speech samples & background noise samples were sorted and categorized into 2 folders - audio & noise, and then noise samples were resampled to 16000Hz & then the background noise was added to the speech samples to augment the data. After that, the FFT of these samples was given to the model for the training & evaluation part.
The following hyperparameters were used during training:
| name | learning_rate | decay | beta_1 | beta_2 | epsilon | amsgrad | training_precision |
|---|---|---|---|---|---|---|---|
| Adam | 0.0010000000474974513 | 0.0 | 0.8999999761581421 | 0.9990000128746033 | 1e-07 | False | float32 |
