Downloads · 30 days
55
54% of all-time downloads
Chima207/distilbert_goodreads_book_classification
distilbert_goodreads_book_classification is a text classification model from Chima207. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as cc-by-4.0.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
55
54% of all-time downloads
All-time downloads
102
Public
Parameters
67M
536 MB on disk
Likes
1
Trending 1
Click a slice to open those files.
.safetensors268 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of distilbert-base-uncased on the Goodreads-Books dataset. It achieves the following results on the evaluation set:
This model is a fine-tuned version of distilbert-base-uncased trained for multi-class book genre classification based on textual metadata (book titles and descriptions).
Due to the noisy and highly heterogeneous nature of crowd-sourced Goodreads user tags, 1,003 unique Goodreads shelves were heuristically mapped into a standardized taxonomy of 31 Amazon Kindle categories. To handle the resulting severe class imbalance, class distribution was balanced using a Random Oversampler before fine-tuning. This model represents a single-stage fine-tuning approach directly on the preprocessed Goodreads metadata dataset, achieving an Accuracy of 50.58% and significantly outperforming traditional vector-based baselines (Doc2Vec + Random Forest at 40.84%).
Automated genre classification of books using metadata or descriptions.
Serving as a baseline for evaluating NLP models on imbalanced, crowd-sourced taxonomy mapping tasks.
Performance is fundamentally constrained by the inherent noise, ambiguity, and subjectivity of crowdsourced user tags (Goodreads shelves).
Predictions are mapped specifically to the 31-class Amazon Kindle taxonomy; inputs outside this category scope may yield inaccurate classifications.
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Accuracy | F1 Score | Precision | Recall |
|---|---|---|---|---|---|---|---|
| 0.4212 | 1.0000 | 8519 | 1.7610 | 0.5058 | 0.4913 | 0.5018 | 0.5058 |
| 0.247 | 1.9999 | 17038 | 2.0030 | 0.5159 | 0.5065 | 0.5073 | 0.5159 |
This repository and model were developed as part of a Bachelor's thesis in 2026.
Dieses Repository und Modell wurden im Rahmen einer Bachelorarbeit im Jahr 2026 entwickelt.