Downloads · 30 days
1
1% of all-time downloads
g25ait2004/DistilBERT_Goodreads
DistilBERT_Goodreads is a text classification model from g25ait2004. Use it when you need a label for a piece of text. It is set up for transformers.
--- libraryname: transformers tags: - text-classification - distilbert - goodreads - mlops datasets: - ucsd-goodreads metrics: - accuracy - f1 pipelinetag: text-classification ---
Downloads · 30 days
1
1% of all-time downloads
All-time downloads
74
Public
Parameters
65.8M
789 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors263 MB · 100%
From the Hugging Face model README
library_name: transformers tags:
This model is a fine-tuned version of distilbert-base-cased on the UCSD Goodreads book reviews dataset to classify books into 7 distinct genres. It was developed as part of an MLOps assignment for IIT Jodhpur.
DistilBERT)distilbert-base-casedThis model is intended to analyze short or long-form book reviews and predict the corresponding book genre. It can be integrated into digital libraries, book recommendation engines, or cataloging systems.
The model classifies text into one of the following 7 labels:
comics_graphicfantasy_paranormalhistory_biographymystery_thriller_crimepoetryromanceyoung_adultUse the code below to load the model and tokenizer for inference:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_name = "g25ait2004/DistilBERT_Goodreads"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
review = "The world-building was absolutely breathtaking, full of dark magic, hidden ancient kingdoms, and dragons."
inputs = tokenizer(review, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
outputs = model(**inputs)
probs = torch.nn.functional.softmax(outputs.logits, dim=-1)
predicted_class_id = outputs.logits.argmax().item()
print(f"Predicted class ID: {predicted_class_id}")
### Downstream Use [optional]
This model is well-suited for integration into digital libraries, book discovery applications, and automated content cataloging pipelines. It can be used as a backend service to automatically tag user-generated reviews with relevant genre nodes or to power downstream recommendation engines based on text sentiment and genre alignment.
### Out-of-Scope Use
* **Non-Review Text Processing:** The model is not intended to classify full-length manuscripts, legal copy, news articles, or technical code repositories.
* **Multilingual Input:** It was fine-tuned purely on English text reviews; feeding it non-English text will result in highly unreliable performance.
* **Automated Moderation:** This model should not be used to flag or filter out explicit or harmful text, as its objective is strictly genre classification.
## Bias, Risks, and Limitations
The dataset relies heavily on self-reported, user-generated content from the UCSD Goodreads Graph, introducing self-selection bias and uneven structural syntax in text inputs.
A significant technical limitation observed during evaluation is the model's performance variability across genres. While it exhibits strong predictive power for unique stylistic categories like **poetry (0.79 F1)** and **comics_graphic (0.81 F1)**, it experiences high confusion rates on structurally overlapping genres such as **young_adult (0.28 F1)** and **fantasy_paranormal (0.41 F1)**.
### Recommendations
Direct and downstream users should expect lower classification fidelity when analyzing books targeting young adult or genre-bending fantasy audiences. We recommend using a confidence threshold (e.g., softmax probability greater than 70%) or falling back to a human-in-the-loop validation model for these ambiguous categories.
## How to Get Started with the Model
Use the code below to quickly load the model and its tokenizer for basic inference:
```python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
# Initialize model and tokenizer
model_name = "g25ait2004/DistilBERT_Goodreads"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
# Sample review text
review_text = "The world-building was absolutely breathtaking, full of dark magic and deep mystery."
# Tokenize and predict
inputs = tokenizer(review_text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
outputs = model(**inputs)
# Extract predicted class
probabilities = torch.nn.functional.softmax(outputs.logits, dim=-1)
predicted_class = outputs.logits.argmax().item()
print(f"Predicted Class ID: {predicted_class}")
**BibTeX:**
```bibtex
@inproceedings{sanh2019distilbert,
title={DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter},
author={Sanh, Victor and Debut, Lysandre and Chaumond, Julien and Wolf, Thomas},
booktitle={NeurIPS EMC^2 Workshop},
year={2019}
}
@article{lacoste2019quantifying,
title={Quantifying the carbon emissions of machine learning},
author={Lacoste, Alexandre and Alexandra, Luccioni and Schmidt, Victor and Dandres, Thomas},
journal={arXiv preprint arXiv:1910.09700},
year={2019}
}
APA:
Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. NeurIPS EMC^2 Workshop.
Lacoste, A., Alexandra, L., Schmidt, V., & Dandres, T. (2019). Quantifying the carbon emissions of machine learning. arXiv preprint arXiv:1910.09700.
Glossary [optional]
Distillation: A structural compression mechanism where a smaller student model (DistilBERT) attempts to recreate the output probability distribution vectors of a much bulkier teacher network (BERT-Base).
Macro F1-Score: The unweighted mean of individual F1-scores across all 7 classes. This treats all genre categories equally, regardless of variations in local test support records.
Mixed Precision (fp16): An optimization method where models calculate gradients inside a 16-bit float structure to maximize processing speed and lower memory usage, while storing base weights in 32-bit floats.
More Information [optional]
This repository belongs to the educational coursework sequences submitted under student assignment benchmarks for the Indian Institute of Technology Jodhpur (IIT Jodhpur) curriculum.
Model Card Authors [optional]
Er. Abhishek Kumar (M.Tech Data Science & Engineering Track Student)
Model Card Contact
For development inquiries, pipeline tracking issues, or alternative checkpoint requests, please submit an issue ticket directly inside your active GitHub project dashboard: GitHub Abhishek Repo Management.