Downloads · 30 days
19
2% of all-time downloads
msperlin/finbert-ai-detector
finbert-ai-detector is a text classification model from msperlin. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
This model is a fine-tuned version of yiyanghkust/finbert-pretrain designed specifically to detect AI-generated text in financial documents, such as corporate annual reports (e.g., 10-K filings).
Downloads · 30 days
19
2% of all-time downloads
All-time downloads
1K
Public
Parameters
110M
439 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors439 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of yiyanghkust/finbert-pretrain designed specifically to detect AI-generated text in financial documents, such as corporate annual reports (e.g., 10-K filings).
The model is used in working paper:
yiyanghkust/finbert-pretrainThe model was trained on a custom dataset compiled from human-written financial texts (derived from SEC annual reports) and AI-generated equivalents.
finbert-pretrain model—already pre-trained on a large corpus of financial text—was fine-tuned on this mixed dataset to classify whether a given segment of text is human-written or generated by an AI.Total cases (AI & Human): 6000 Total cases (AI): 3000
Estimation cases: 4200 Test cases: 1800
| Metric | Value |
|---|---|
| accuracy | 89.16% |
| f1 | 88.57% |
| precision | 92.64% |
| recall | 84.84% |

This model is intended for researchers, financial analysts, and auditors who want to verify the authenticity of corporate disclosures and determine if a financial text (like an annual report or an earnings call transcript) was written by an AI or a human.
There are two ways to use this model, you can either use a python package that handles the model internaly, or the standard transformers library (default for hugging face).
finbert-ai-detectorInstall locally using pip:
pip install finbert-ai-detector
Run it with the following code:
from finbert_ai_detector import FinbertAIDetector
# Initialize the detector (downloads the model if not cached)
detector = FinbertAIDetector()
# Example text
text = "The Tax Cuts and Jobs Act enacted in 2017 in the United States, significantly changed the tax rules applicable to U.S.-domiciled corporations. Changes such as lower corporate tax rates, full expensing for qualified property, taxation of offshore earnings, limitations on interest expense deductions, and changes to the municipal bond tax exemption may impact demand for our products and services."
# Predict a single text
result = detector.predict(text)
print(f"Prediction: {result['label']}")
print(f"AI Probability: {result['ai_probability']:.2%}")
transformers library.Here is a quick example of how to make predictions using Python and PyTorch:
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
# Load the model and tokenizer
model_id = "msperlin/finbert-ai-detector"
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id).to(device)
model.eval()
# Sample text to test
text = "The Tax Cuts and Jobs Act enacted in 2017 in the United States, significantly changed the tax rules applicable to U.S.-domiciled corporations. Changes such as lower corporate tax rates, full expensing for qualified property, taxation of offshore earnings, limitations on interest expense deductions, and changes to the municipal bond tax exemption may impact demand for our products and services."
# Preprocess the input
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
max_length=512,
padding=True
).to(device)
# Run inference
with torch.no_grad():
outputs = model(**inputs)
probabilities = torch.nn.functional.softmax(outputs.logits, dim=-1)
# Index 1 typically corresponds to "AI-generated" (Verify with model.config.id2label if needed)
prob_ai = probabilities[0][1].item()
print(f"Probability of being AI-generated: {prob_ai:.2%}")