Downloads · 30 days
10
18% of all-time downloads
virustechhacks/distil-bert-classifier
distil-bert-classifier is a text classification model from virustechhacks. Use it when you need a label for a piece of text. It is set up for transformers.
This model is a fine-tuned DistilBERT model for sequence classification, designed to identify whether a place (e.g., restaurants, businesses) is NEW, CLOSED, or NEUTRAL based on short text snippets.
Downloads · 30 days
10
18% of all-time downloads
All-time downloads
56
Public
Parameters
67M
268 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors268 MB · 100%
From the Hugging Face model README
This model is a fine-tuned DistilBERT model for sequence classification, designed to identify whether a place (e.g., restaurants, businesses) is NEW, CLOSED, or NEUTRAL based on short text snippets.
distilbert-base-uncasedNEW, CLOSED, NEUTRALThis model helps extract signals about business status from textual data such as reviews, posts, or headlines.
Classify short text snippets into:
NEW → Newly opened placesCLOSED → Shut down or no longer operatingNEUTRAL → No clear status signalOutputs can be aggregated into features like:
closed_signal_rationew_signal_ratiomention_countThese can feed into larger ML pipelines (e.g., XGBoost models).
Synthetic Data Bias:
Trained on rule-based synthetic data → may not generalize well to real-world language.
No Time Awareness:
Cannot distinguish recent vs outdated signals.
Token Limit:
Inputs >128 tokens are truncated.
For production use:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
import torch.nn.functional as F
repo_name = "virustechhacks/distil-bert-classifier"
tokenizer = AutoTokenizer.from_pretrained(repo_name)
model = AutoModelForSequenceClassification.from_pretrained(repo_name)
id_to_label = {0: "NEW", 1: "CLOSED", 2: "NEUTRAL"}
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)
def predict_status(text):
inputs = tokenizer(
text,
truncation=True,
padding="max_length",
max_length=128,
return_tensors="pt"
)
inputs = {k: v.to(device) for k, v in inputs.items()}
with torch.no_grad():
outputs = model(**inputs)
probs = F.softmax(outputs.logits, dim=-1)
confidence, pred = torch.max(probs, dim=1)
return id_to_label[pred.item()], confidence.item()
# Example
print(predict_status("Grand opening this weekend!"))
print(predict_status("The store ceased operations."))