Downloads · 30 days
0
arjun-varma/tmdb-genre-classifier
tmdb-genre-classifier is a machine learning model from arjun-varma. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for sklearn. The card lists the license as mit.
Serverless Machine Learning Pipeline — TF-IDF + Linear SVC — Fully Automated & Deployed
Downloads · 30 days
0
Access
Public
Updated Dec 6, 2025
Repo size
1.7 MB
Likes
0
Public
Click a slice to open those files.
.joblib1.6 MB · 96%
From the Hugging Face model README
Serverless Machine Learning Pipeline — TF-IDF + Linear SVC — Fully Automated & Deployed
This project demonstrates the ability to design, automate, and deploy a real-world Machine Learning system without relying on paid cloud services.
It showcases strong understanding and application of:
The model predicts multiple genres for a movie based on its description — similar to how streaming platforms tag content for recommendations.
➡ Live Demo: https://huggingface.co/spaces/arjun-varma/tmdb-genre-classifier
➡ Model Hub: https://huggingface.co/arjun-varma/tmdb-genre-classifier
Movies are not mutually exclusive:
| Plot Summary | Correct Genres |
|---|---|
| Soldier returns from war, struggling with trauma | Drama, War |
| AI becomes sentient and turns against creators | Sci‑Fi, Thriller |
| A musician finds love on tour | Music, Romance |
Single‑label classifiers fail here.
Multi‑label learning predicts all genres that simultaneously apply.
This creates challenges:

No AWS SageMaker, no GCP Vertex AI.
Infrastructure cost = $0
| Choice | Reason |
|---|---|
| Transformers | Expensive & slow for nightly retraining |
| Neural Networks | Need GPUs / infra |
| Logistic Regression | High precision, low recall |
| Linear SVC + TF‑IDF | Fast, scalable, interpretable 👈 Best for pipeline |
The biggest improvement:
| Model | Precision_micro | Recall_micro | F1_macro | Result |
|---|---|---|---|---|
| Logistic Regression | 0.83 | 0.006 | ~0.03 | Almost no predictions |
| Linear SVC + threshold 0.25 | 0.16 | 0.99 | 0.27 | Usable predictions |
Interpretation:
If this was powering recommendations, threshold matters.
This project includes:
Tools used:
pytestmonkeypatchtmp_pathThis demonstrates reliability in automation-focused ML environments.
The model provides:
| Idea | Value |
|---|---|
| Compare vs MiniLM Transformer | Benchmark credibility |
| Add FastAPI inference service | Deployable microservice |
| Visualize confidence & confusion | Explainable AI |
Arjun Varma
Machine Learning Engineer & Systems Developer
Designed for real-world ML infrastructure readiness.