Downloads · 30 days
12
12% of all-time downloads
Niophy/distilbert-imdb-sentiment
distilbert-imdb-sentiment is a text classification model from Niophy. Use it when you need a label for a piece of text. The card lists the license as apache-2.0.
This is distilbert-base-uncased fine-tuned for binary sentiment classification (positive/negative) on the IMDB movie reviews dataset.
Downloads · 30 days
12
12% of all-time downloads
All-time downloads
100
Public
Parameters
67M
536 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors268 MB · 100%
From the Hugging Face model README
This is distilbert-base-uncased fine-tuned for binary sentiment
classification (positive/negative) on the IMDB movie reviews dataset.
distilbert-base-uncasedstanfordnlp/imdb, full train split (25,000 examples)| Epoch | Train loss | Test loss | Test accuracy |
|---|---|---|---|
| 1 | 0.302 | 0.283 | 88.5% |
| 2 | 0.157 | 0.366 | 89.2% |
| 3 | 0.068 | 0.385 | 91.2% |
Final test accuracy: 91.2%
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
repo_id = "Niophy/distilbert-imdb-sentiment"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForSequenceClassification.from_pretrained(repo_id)
model.eval()
id2label = {0: "negative", 1: "positive"}
sentence = "This movie was absolutely fantastic"
inputs = tokenizer(sentence, return_tensors="pt", truncation=True, max_length=256)
with torch.no_grad():
logits = model(**inputs).logits
predicted_id = torch.argmax(logits, dim=-1).item()
print(id2label[predicted_id])
Test loss rises after epoch 1 even as accuracy keeps improving, indicating mild overfitting by epoch 3. Training for more epochs without regularization (e.g. weight decay, early stopping) is unlikely to help much beyond this point.