Downloads · 30 days
58
40% of all-time downloads
brjoey/CBSI-bert-base-uncased
CBSI-bert-base-uncased is a text classification model from brjoey. Use it when you need a label for a piece of text. It is set up for transformers.
This model is trained on the replication data of Nițoi et al. (2023). Check out their paper and website for more information.
Downloads · 30 days
58
40% of all-time downloads
All-time downloads
145
Public
Parameters
109M
1.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt876 MB · 67%
From the Hugging Face model README
This model is trained on the replication data of Nițoi et al. (2023).
Check out their paper and website for more information.
The model is trained with the hyperparameters used by Nițoi et al. (2023).
In addition, different hyperparameters, seeds, and reinitialization of the first L layers were tested. The performance seems relatively stable across hyperparameter settings.
Alongside these models, FinBERT, different versions of RoBERTa, and EconBERT were tested. The performance of the BERT-based models reported here is significantly better. In addition, fine-tuned ModernBERT and CentralBank-BERT versions are also available.
| Model | F1 Score | Accuracy | Loss |
|---|---|---|---|
| CBSI-bert-base-uncased | 0.88 | 0.88 | 0.49 |
| CBSI-bert-large-uncased | 0.92 | 0.92 | 0.45 |
| CBSI-ModernBERT-base | 0.93 | 0.93 | 0.40 |
| CBSI-ModernBERT-large | 0.91 | 0.91 | 0.53 |
| CBSI-CentralBank-BERT | 0.92 | 0.92 | 0.36 |
import pandas as pd
from transformers import AutoTokenizer, AutoModelForSequenceClassification, pipeline
# Load model and tokenizer
model_name = "brjoey/CBSI-bert-base-uncased"
classifier = pipeline(
"text-classification",
model=model_name,
tokenizer=model_name
)
# Define label mapping
cbsi_label_map = {
0: "neutral",
1: "dovish",
2: "hawkish"
}
# Classify a column in a Pandas DataFrame - replace with your DataFrame
texts = [
"The Governing Council decided to lower interest rates.",
"The central bank will maintain its current policy stance."
]
df = pd.DataFrame({
"text": texts
})
# Run classification
predictions = classifier(
df["text"].tolist()
)
# Store the results
df["label"], df["score"] = zip(*[
(cbsi_label_map[int(pred["label"].split("_")[-1])], pred["score"])
for pred in predictions
])
print("\n === Results ===\n")
print(df[["text", "label", "score"]])
If you use this model, please cite:
Data:
Nițoi Mihai; Pochea Maria-Miruna; Radu Ștefan-Constantin, 2023,
"Replication Data for: Unveiling the sentiment behind central bank narratives: A novel deep learning index",
https://doi.org/10.7910/DVN/40JFEK, Harvard Dataverse, V1
Model / Paper:
Mihai Niţoi, Maria-Miruna Pochea, Ştefan-Constantin Radu,
Unveiling the sentiment behind central bank narratives: A novel deep learning index,
Journal of Behavioral and Experimental Finance, Volume 38, 2023, 100809, ISSN 2214-6350.
https://doi.org/10.1016/j.jbef.2023.100809
BERT:
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805.
https://arxiv.org/abs/1810.04805