Downloads · 30 days
32
24% of all-time downloads
SayedShaun/bangla-wav2vec2
bangla-wav2vec2 is a automatic speech recognition model from SayedShaun. Use it when you need speech turned into text. The card lists the license as unknown.
A Wav2Vec2-based CTC acoustic model for Bengali (Bangla) automatic speech recognition (ASR).
Downloads · 30 days
32
24% of all-time downloads
All-time downloads
134
Public
Repo size
2.5 GB
Likes
0
Public
Click a slice to open those files.
.bin1.3 GB · 100%
From the Hugging Face model README
A Wav2Vec2-based CTC acoustic model for Bengali (Bangla) automatic speech recognition (ASR).
This model's weights and configuration were not trained by the uploader. They are mirrored from the Kaggle dataset
wv-shru-v3-s6 published by Kaggle user
qdv206. All credit for training and releasing the original model goes to the original author.
This upload exists to make the checkpoint easier to load and use via the huggingface_hub / transformers ecosystem.
If you use this model, please credit the original author and link back to the source dataset above.
hidden_size=1024, checkpoint tag full_shru_v2_s20Wav2Vec2CTCTokenizer + Wav2Vec2FeatureExtractor (Wav2Vec2Processor)Note: the original
config.jsonlistsarchitectures: ["Wav2Vec2ForCTCV2"], a custom class name used in the original training pipeline. The underlying weights are a standard Wav2Vec2-for-CTC architecture, so the model loads with the standardWav2Vec2ForCTC/AutoModelForCTCclasses fromtransformers(see usage below).
import torch
import soundfile as sf
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
model_id = "SayedShaun/bangla-wave2vec2"
processor = Wav2Vec2Processor.from_pretrained(model_id)
model = Wav2Vec2ForCTC.from_pretrained(model_id)
speech, sr = sf.read("audio.wav") # expects 16kHz mono
inputs = processor(speech, sampling_rate=16000, return_tensors="pt", padding=True)
with torch.no_grad():
logits = model(inputs.input_values).logits
predicted_ids = torch.argmax(logits, dim=-1)
transcription = processor.batch_decode(predicted_ids)
print(transcription)
No explicit license was specified by the original author on Kaggle. Please refer to the original dataset page for usage terms, and contact the original author for clarification if you plan to use this commercially.