Downloads · 30 days
14
19% of all-time downloads
Aukrk/MLOPS_group-v4
MLOPS_group-v4 is a text classification model from Aukrk. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
This repository is part of the MLOps Group 36 Project for the PGD AI Programme, IIT Jodhpur.
Downloads · 30 days
14
19% of all-time downloads
All-time downloads
75
Public
Parameters
67M
268 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors268 MB · 100%
From the Hugging Face model README
This repository is part of the MLOps Group 36 Project for the PGD AI Programme, IIT Jodhpur.
The project implements an end-to-end MLOps pipeline for SMS spam classification using DistilBERT, with GitHub, Kaggle, Weights & Biases, Hugging Face Hub, Docker, and GitHub Actions.
Anu Kumar
Roll Number: G25AIT2016
| Resource | Link |
|---|---|
| GitHub Repository | https://github.com/g25ait2032-prog/mlops-group36-iitj |
| Kaggle Notebook - G25AIT2016 | https://www.kaggle.com/code/anukumarkg25ait2016/mlops-group36-data-preprocessing-g25ait2016 |
| W&B Run - G25AIT2016 | https://wandb.ai/g25ait2032-iit-jodhpur/MLOPS_Group/runs/j5fk4zll |
| W&B Project Dashboard | https://wandb.ai/g25ait2032-iit-jodhpur/MLOPS_Group |
| Hugging Face Model | https://huggingface.co/Aukrk/MLOPS_group-v4 |
| Item | Value |
|---|---|
| Base model | distilbert-base-uncased |
| Task | Binary text classification |
| Classes | ham, spam |
| Dataset | UCI SMS Spam Collection |
| Framework | Hugging Face Transformers |
| Output labels | 0 = ham, 1 = spam |
This repository is linked to the G25AIT2016 Task 2 workflow.
The completed contribution includes:
id2label.json and label2id.json| Metric | Value |
|---|---|
| Raw samples | 5,574 |
| Duplicates removed | 415 |
| Cleaned samples | 5,159 |
| Train rows | 3,611 |
| Validation rows | 774 |
| Test rows | 774 |
| Sanity checks passed | 21 / 21 |
| Leakage check | Passed |
{
"0": "ham",
"1": "spam"
}
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="Aukrk/MLOPS_group-v4"
)
text = "Congratulations! You have won a free iPhone. Click here now."
result = classifier(text)
print(result)
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tokenizer = AutoTokenizer.from_pretrained("Aukrk/MLOPS_group-v4")
model = AutoModelForSequenceClassification.from_pretrained("Aukrk/MLOPS_group-v4")
| Text | Expected Output |
|---|---|
| Congratulations! You have won a free prize. Click here now. | spam |
| Can we meet tomorrow at 5 PM? | ham |
The G25AIT2016 W&B run records data-preparation metrics such as:
W&B Run: https://wandb.ai/g25ait2032-iit-jodhpur/MLOPS_Group/runs/j5fk4zll
This model repository is published under the G25AIT2016 Hugging Face account for Group 36 project traceability.
The model artefact follows the Group 36 DistilBERT SMS spam classification workflow and is linked with the data-preparation contribution completed by Anu Kumar - G25AIT2016.
This repository is intended for:
MLOps Group 36
PGD AI Programme, IIT Jodhpur
Contributor for this repository:
Anu Kumar - G25AIT2016