Downloads · 30 days
68
32% of all-time downloads
appcle/distilbert-base-uncased-cfpd
distilbert-base-uncased-cfpd is a text classification model from appcle. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
This model is a fine-tuned version of distilbert-base-uncased for classifying consumer complaint narratives from the Consumer Financial Protection Bureau (CFPB) consumer complaints dataset.
Downloads · 30 days
68
32% of all-time downloads
All-time downloads
210
Public
Parameters
67M
1.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors268 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of distilbert-base-uncased for classifying consumer complaint narratives from the Consumer Financial Protection Bureau (CFPB) consumer complaints dataset.
The model performs multi-class classification across seven consolidated consumer-finance categories and is intended to support automated complaint routing and triage.
The model is intended for:
The model is intended as a decision-support component and should not replace human review where classification accuracy is critical.
Performance may vary for:
The model was trained on CFPB consumer complaint data and may not generalize reliably to other domains without further evaluation or fine-tuning.
The model was trained using consumer complaint narratives from the Consumer Financial Protection Bureau (CFPB) consumer complaints dataset.
The original CFPB complaint categories were consolidated into seven broader operational categories:
| Consolidated Category | Original CFPB Categories |
|---|---|
| Collections & Recovery | Debt collection; Debt or credit management |
| Banking Operations | Checking or savings account |
| Cards & Payments | Credit card; Prepaid card |
| Consumer Lending | Vehicle loan or lease; Payday loan, title loan, personal loan, or advance loan; Student loan |
| Money Transfer and Payments | Money transfer, virtual currency, or money service |
| Mortgage & Home Lending | Mortgage |
| Credit Reporting & Disputes | Credit reporting or other personal consumer reports |
The training dataset contains 14,198 samples. The validation dataset is used to monitor model performance during training and for the reported evaluation metrics.
Complaint narratives were cleaned before being provided to the model.
The preprocessing included:
XXXX with the token maskThe original text casing was retained during preprocessing to allow the pretrained DistilBERT tokenizer and model to handle the input.
The model was fine-tuned using the Hugging Face Transformers framework.
5e-05168442(0.9, 0.999)1e-08Training completed after 1,005.7 seconds (approximately 16.8 minutes).
The final training run produced:
0.449056.473.5323,5524The following results were obtained on the evaluation set:
| Metric | Score |
|---|---|
| Loss | 0.6341 |
| Accuracy | 0.8268 |
| Precision | 0.8077 |
| Recall | 0.8055 |
| F1-macro | 0.8065 |
| Weighted F1 | 0.8269 |
| Epoch | Step | Training Loss | Validation Loss | Accuracy | Precision | Recall | F1-macro | Weighted F1 |
|---|---|---|---|---|---|---|---|---|
| 1 | 888 | 0.9142 | 0.5539 | 0.8161 | 0.7910 | 0.7932 | 0.7907 | 0.8154 |
| 2 | 1776 | 0.4838 | 0.5261 | 0.8327 | 0.8152 | 0.8047 | 0.8095 | 0.8328 |
| 3 | 2664 | 0.3020 | 0.5619 | 0.8310 | 0.8194 | 0.8034 | 0.8109 | 0.8309 |
| 4 | 3552 | 0.1970 | 0.6341 | 0.8268 | 0.8077 | 0.8055 | 0.8065 | 0.8269 |
Validation performance improved through the first three epochs, with the highest validation F1 score of 0.8109 achieved at epoch 3. Performance declined slightly at epoch 4, while training loss continued to decrease.
The model can be loaded using the Hugging Face Transformers library:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
model_name = "appcle/distilbert-base-uncased-cfpb"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
For inference, provide a consumer complaint narrative to the tokenizer and pass the resulting inputs to the model.
This model is provided for research, development, and demonstration purposes. Reported metrics are based on the evaluation data used during model development and should not be interpreted as a guarantee of performance on new or different datasets.