Downloads · 30 days
154
29% of all-time downloads
ritessshhh/FinDeBERTa
FinDeBERTa is a text classification model from ritessshhh. Use it when you need a label for a piece of text. The card lists the license as mit.
FinDeBERTa is a fine-tuned DeBERTa-v3-Large model for multi-label financial event classification. It predicts one or more event types from financial news headlines with state-of-the-art performance.
Downloads · 30 days
154
29% of all-time downloads
All-time downloads
528
Public
Parameters
435M
3.2 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors1.7 GB · 99%
From the Hugging Face model README
FinDeBERTa is a fine-tuned DeBERTa-v3-Large model for multi-label financial event classification. It predicts one or more event types from financial news headlines with state-of-the-art performance.
The model classifies text into 18 financial event types:
["CSR/Brand", "Deal", "Dividend", "Employment", "Expense", "Facility",
"FinancialReport", "Financing", "Investment", "Legal", "Macroeconomics",
"Merger/Acquisition", "Product/Service", "Profit/Loss", "Rating", "Revenue",
"SalesVolume", "SecurityValue"]
For best performance, use the per-class optimized thresholds:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
import numpy as np
from huggingface_hub import hf_hub_download
model_name = "ritessshhh/FinDeBERTa"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
# Download per-class thresholds
thresholds_path = hf_hub_download(repo_id=model_name, filename="thresholds.npy")
thresholds = np.load(thresholds_path)
text = "Tesla to acquire a battery startup in a 400 million dollar deal."
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=128)
with torch.no_grad():
outputs = model(**inputs)
probs = torch.sigmoid(outputs.logits)[0].cpu().numpy()
# Apply per-class thresholds
predictions = [
{"label": model.config.id2label[i], "score": float(prob)}
for i, prob in enumerate(probs) if prob >= thresholds[i]
]
# Sort by score
predictions = sorted(predictions, key=lambda x: x["score"], reverse=True)
print(predictions)
# Output: [{"label": "Merger/Acquisition", "score": 0.98}, ...]
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
import numpy as np
model_name = "ritessshhh/FinDeBERTa"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
text = "Tesla to acquire a battery startup in a 400 million dollar deal."
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=128)
with torch.no_grad():
outputs = model(**inputs)
probs = torch.sigmoid(outputs.logits)[0].cpu().numpy()
# Using default threshold of 0.5
threshold = 0.5
predictions = [
{"label": model.config.id2label[i], "score": float(prob)}
for i, prob in enumerate(probs) if prob >= threshold
]
print(predictions)
| Metric | Score |
|---|---|
| Macro F1 | 0.692 |
| Micro F1 | 0.691 |
| Precision (Macro) | 0.738 |
| Recall (Macro) | 0.691 |
| Exact Match Ratio | 0.532 |
| Label | F1 Score | Precision | Recall |
|---|---|---|---|
| Dividend | 1.000 | 1.000 | 1.000 |
| Employment | 0.923 | 0.857 | 1.000 |
| Merger/Acquisition | 0.892 | 0.967 | 0.829 |
| Profit/Loss | 0.833 | 0.824 | 0.843 |
| SecurityValue | 0.829 | 0.843 | 0.815 |
| Rating | 0.790 | 0.800 | 0.780 |
| Revenue | 0.780 | 0.800 | 0.762 |
| SalesVolume | 0.748 | 0.833 | 0.678 |
| Financing | 0.714 | 0.714 | 0.714 |
| Deal | 0.696 | 0.889 | 0.571 |
If you use this model, please cite:
@misc{findeberta2024,
author = {ritessshhh},
title = {FinDeBERTa: Multi-Label Financial Event Classifier},
year = {2024},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/ritessshhh/FinDeBERTa}}
}