Downloads · 30 days
8
1% of all-time downloads
abdullaharean/regipa_bangla
regipa_bangla is a machine learning model from abdullaharean. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
Our team's solution focuses on developing a robust model for transcribing Bengali text into International Phonetic Alphabet (IPA), contributing to computational linguistics and NLP research in Bengali. Leveraging a li…
Downloads · 30 days
8
1% of all-time downloads
All-time downloads
909
Public
Parameters
300M
8.4 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors1.2 GB · 100%
From the Hugging Face model README
Our team's solution focuses on developing a robust model for transcribing Bengali text into International Phonetic Alphabet (IPA), contributing to computational linguistics and NLP research in Bengali. Leveraging a linguist-validated dataset encompassing diverse domains of Bengali text, our model aims to accurately capture the phonetic nuances and regional dialects present in Bengali language.
We preprocess the Bengali text data to handle linguistic variations, tokenization, and normalization.
Our model architecture employs state-of-the-art deep learning techniques, such as recurrent neural networks (RNNs) or transformer-based models, to capture the sequential and contextual information inherent in language.
The model is trained on the linguist-validated dataset, optimizing for accuracy, robustness, and generalization across various dialects and linguistic contexts.
We validate the model's performance using rigorous evaluation metrics, ensuring its effectiveness in accurately transcribing Bengali text into IPA.
Upon successful validation, the model is deployed as an open-source tool, extending the capabilities of generalized Bengali Text-to-Speech systems and facilitating further research in Bengali computational linguistics.
from transformers import T5ForConditionalGeneration
import torch
model = T5ForConditionalGeneration.from_pretrained('abdullaharean/regipa_bangla')
input_ids = torch.tensor([list("Life is like a box of chocolates.".encode("utf-8"))]) + 3 # add 3 for special tokens
labels = torch.tensor([list("La vie est comme une boîte de chocolat.".encode("utf-8"))]) + 3 # add 3 for special tokens
loss = model(input_ids, labels=labels).loss # forward pass