Downloads Ā· 30 days
0
pratyushee/assamese-sentiment-analysis
assamese-sentiment-analysis is a text classification model from pratyushee. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
Tags: text-classification sentiment-analysis Assamese LSTM
Downloads Ā· 30 days
0
Access
Public
Updated Jul 29, 2025
Repo size
ā
Likes
0
Public
Click a slice to open those files.
.ipynb190 KB Ā· 97%
From the Hugging Face model README
Tags: #text-classification #sentiment-analysis #Assamese #LSTM
A deep learning-powered tool to classify Assamese text as Positive, Negative, or Neutral using an LSTM model tailored for the Assamese language.
| Property | Details |
|---|---|
| Model Name | pratyushee/assamese-sentiment-analysis |
| Architecture | Pretrained LSTM-based neural network |
| Language | Assamese (ą¦ ą¦øą¦®ą§ą¦Æą¦¼ą¦¾) |
| Classes | 3 ā Positive, Neutral, Negative |
| Use Cases | Customer feedback, social media monitoring, opinion mining |
Clone the repo and install the requirements:
pip install -r requirements.txt
Install the custom Assamese tokenizer:
git clone https://github.com/KashyapKishore/AssameseTokenizer.git
cd AssameseTokenizer
pip install .
This model was developed using Assamese text data and trained with a custom tokenizer specifically designed for Assamese script. It uses an LSTM architecture, making it well-suited for capturing the sequence and context of natural language in sentiment classification tasks.
š Training Data The dataset was curated from public sources such as news articles, social media comments, and feedback forms, and was manually labeled into three sentiment classes: Positive, Neutral, and Negative.
šļø Training Procedure
āļø Preprocessing: Text cleaning, tokenization using AssameseTokenizer, optional stemming and stopword removal
š¢ Input Handling: Sequences padded or truncated to a fixed length of 512 tokens
š§ Architecture: Embedding layer ā LSTM ā Dense (Softmax)
š§ Regularization: Dropout layers to prevent overfitting
āļø Optimizer: Adam
š Epochs: Trained for X epochs (replace with your actual number)
š Evaluation: Final validation accuracy and F1-score: Insert actual metrics here
Ideal for:
šØļø Social media sentiment tracking in Assamese
š¢ Public opinion & brand monitoring
š Research on low-resource NLP in Indic languages
ā ļø Limitations / Not Recommended For:
Code-mixed Assamese-English input
Domain-specific texts (e.g., legal, medical) without additional fine-tuning
You can load and run the model easily via Hugging Face's transformers pipeline:
from transformers import pipeline
model_name = "pratyushee/assamese-sentiment-analysis"
pipe = pipeline("text-classification", model=model_name, tokenizer=model_name)
result = pipe("ą¦ą¦ ą¦ą¦¾ą¦¬ą¦¾ą§°ą¦ą¦¾ ą¦ą¦ą¦¦ą¦® ą¦ą¦¾ą¦²ą§ ą¦ą¦ą¦æą¦²!") # Sample Assamese sentence
print(result)