language:
- en
metrics:
- accuracy
- f1
base_model:
- distilbert/distilbert-base-uncased
pipeline_tag: text-classification
Model Card for distilbert-imdb-lora
This model is a DistilBERT sentiment classifier fine-tuned on the IMDB Movie Reviews dataset using LoRA (Low-Rank Adaptation) for parameter-efficient training. It achieves 90.04% accuracy and 90.12% F1 on the validation set.
It is lightweight, fast, explainable (SHAP support), and suitable for movie-review sentiment classification tasks.
Model Details
Model Description
This model classifies text into Positive or Negative sentiment.
It was fine-tuned using LoRA, training only 0.3M parameters instead of 66M, making it efficient and resource-friendly.
- Developed by: Praanshull
- Shared by: Praanshull
- Model type: Transformer-based Sequence Classifier (DistilBERT)
- Language(s): English
- License: Apache 2.0 (same as DistilBERT)
- Finetuned from: distilbert-base-uncased
Model Sources
- Repository: (https://huggingface.co/Praanshull/sentiment-analyzer-app)
- Paper (base model): DistilBERT: a distilled version of BERT
Uses
Direct Use
- Sentiment analysis for movie reviews
- General positive/negative polarity classification
- Lightweight deployment on web apps and mobile devices
- Explainable sentiment scoring (using SHAP)
Downstream Use
- Plug into larger NLP pipelines
- Fine-tune further on domain-specific sentiment datasets
Out-of-Scope Use
- Not designed for multi-class sentiment
- Not suitable for emotionally-complex or sarcasm-heavy text
- Not trained for languages other than English
Bias, Risks, and Limitations
- Trained on IMDB data → may inherit bias toward movie-review writing style
- May misclassify:
- sarcasm
- short texts
- mixed sentiment
- Not immune to lexical bias ("good"/"bad" words heavily weighted)
Recommendations
Users should:
- Validate results before production
- Avoid using for safety-critical decisions
- Calibrate probabilities if needed
How to Get Started
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="your-username/distilbert-imdb-lora",
tokenizer="your-username/distilbert-imdb-lora"
)
classifier("This movie was amazing!")
Training Details
Training Data
- Dataset: IMDB Movie Reviews
- Train split: 22,500 reviews
- Validation split: 2,500 reviews (10% of training set)
- Test split: 25,000 reviews
- Balanced binary labels
Training Procedure
Preprocessing
- Max length: 256
- Tokenizer: distilbert-base-uncased
- Truncation enabled
LoRA Hyperparameters
- r = 8
- alpha = 16
- dropout = 0.1
- target_modules = ["q_lin", "k_lin", "v_lin", "out_lin", "lin1", "lin2"]
Training Hyperparameters
- Batch size: 16
- Learning rate: 2e-5
- Scheduler: cosine
- Warmup: 10%
- Epochs: 8 (early stopped at epoch 6)
- Precision: bfloat16
- Weight decay: 0.01
Speeds, Sizes, Times
- Training time: ~2 hours on Google Colab T4
- Model size (merged): ~260MB
- LoRA adapter size: ~3MB
Evaluation
Testing Data
- IMDB official test set (25,000 samples)
Metrics
- Accuracy
- F1 score
Both suitable for balanced binary classification.
Results
| Metric | Value | Epoch |
|---|
| Validation Accuracy | 90.04% | 6 |
| Validation F1 | 90.12% | 6 |
| Validation Loss | 0.2652 | 6 |
Summary
The model achieves strong performance with minimal overfitting.
LoRA provides 3× faster training and 99% fewer trainable parameters with no accuracy loss.
Model Examination
SHAP explainability is supported.
- Waterfall plots
- Token-level attributions
- Class-specific influence scores
This allows understanding why the model predicted Positive/Negative.
Environmental Impact
- Hardware: Google Colab T4 GPU
- Hours used: ~2 hours
- Cloud provider: Google
- Region: (Varies by Colab session)
- Estimated CO₂: Low (< 0.5 kg)
Technical Specifications
Architecture
- Base model: DistilBERT
- Classification head: 2-unit softmax
- Adaptation: LoRA injected into attention and feedforward layers
Compute Infrastructure
Hardware
Software
- Python 3.10
- PyTorch
- Transformers
- Datasets
- PEFT (LoRA)
- Accelerate
- SHAP
- Gradio
Citation
BibTeX
@misc{distilbert-imdb-lora,
title={DistilBERT IMDB LoRA Fine-tuned Model},
author={Praanshull Verma},
year={2025},
howpublished={\url{https://huggingface.co/Praanshull/sentiment-analyzer-app}},
}
APA
Praanshull Verma. (2025). DistilBERT IMDB LoRA Fine-tuned Model. Hugging Face.
Glossary
- LoRA: Parameter-efficient fine-tuning method
- SHAP: Explainability method based on Shapley values
- DistilBERT: Lightweight version of BERT
Model Card Authors
Praanshull Verma / GitHub Praanshull / HuggingFace Praanshull
Model Card Contact
praanshullverma23@gmail.com