Downloads · 30 days
15
13% of all-time downloads
vishnu-n/TamilSense-model
TamilSense-model is a machine learning model from vishnu-n. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Lightweight, production-grade Tamil/Tanglish sentiment analysis — built for real-world deployment
Downloads · 30 days
15
13% of all-time downloads
All-time downloads
120
Public
Parameters
238M
950 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors950 MB · 99%
From the Hugging Face model README
Lightweight, production-grade Tamil/Tanglish sentiment analysis — built for real-world deployment
ChatGPT is a trillion-parameter monster that costs millions to run. Nobody is integrating it into a local government portal, a small business app, or a low-bandwidth mobile app in rural Tamil Nadu.
TamilSense fills that gap — a lightweight, fast, open-source Tamil sentiment API that any developer can plug in.
| Metric | Score |
|---|---|
| Accuracy | 94.7% |
| Weighted F1 | 94.7% |
| Positive F1 | 97% |
| Negative F1 | 85% |
Trained on 55,064 balanced Tamil-English sentences from two combined datasets.
👉 Try it on Hugging Face Spaces
from transformers import pipeline
classifier = pipeline("text-classification", model="vishnuexe/TamilSense-model")
result = classifier("Super da machan vera level!")
print(result) # [{'label': 'positive', 'score': 0.9965}]
# Run locally
uvicorn app.main:app --reload
# Predict
curl -X POST "http://localhost:8000/predict" \
-H "Content-Type: application/json" \
-d '{"text": "Romba nalla iruku bro"}'
Response:
{
"text": "Romba nalla iruku bro",
"sentiment": "positive",
"confidence": 0.9966,
"scores": {"positive": 0.9966, "negative": 0.0034},
"response_time_ms": 48.23
}
Tamil/Tanglish Text ↓ MuRIL Tokenizer (WordPiece) ↓ 12 Transformer Layers (Google MuRIL) ↓ [CLS] Vector (768-dim) ↓ Linear Classifier → Positive / Negative
| Source | Size |
|---|---|
| tamilmixsentiment (FIRE 2020) | 15,744 sentences |
| DravidianCodeMix Zenodo | ~44,000 sentences |
| Final balanced train set | 55,064 sentences |
| Run | Data | F1 |
|---|---|---|
| Run 1 — 3-class | 10k sentences | 67.5% |
| Run 2 — Binary | 15k sentences | 80.8% |
| Run 3 — Tuned | 15k sentences | 81.5% |
| Run 4 — Full data | 55k sentences | 94.7% |
TamilSense/ ├── app/ │ ├── main.py # FastAPI REST API │ └── gradio_app.py # Gradio demo UI ├── src/ │ ├── prepare_data.py # Data loading and balancing │ ├── train.py # Fine-tuning with MLflow tracking │ ├── evaluate.py # Evaluation + confusion matrix │ └── predict.py # Inference pipeline ├── Dockerfile └── requirements.txt
MIT License — free to use, modify, and deploy.
Built by Vishnu | HuggingFace | GitHub
Available for custom NLP fine-tuning and ML deployment projects. Fiverr: [http://www.fiverr.com/s/bdVkvB1]