Downloads · 30 days
7
3% of all-time downloads
AventIQ-AI/distilbert-base-uncased_token_classification
distilbert-base-uncased_token_classification is a machine learning model from AventIQ-AI. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This model is a Named Entity Recognition (NER) system fine-tuned on the WNUT 17 dataset using DistilBERT. It predicts entity types such as persons, organizations, locations, and more from text.
Downloads · 30 days
7
3% of all-time downloads
All-time downloads
230
Public
Parameters
66.4M
133 MB on disk
Likes
4
Public
Click a slice to open those files.
.safetensors133 MB · 99%
From the Hugging Face model README
This model is a Named Entity Recognition (NER) system fine-tuned on the WNUT 17 dataset using DistilBERT. It predicts entity types such as persons, organizations, locations, and more from text.
Model Type: Transformer-based NER
Base Model: DistilBERT (distilbert-base-uncased)
Dataset: WNUT 17
Training Framework: PyTorch & Hugging Face Transformers
Training Epochs: 3
Batch Size: 16
Learning Rate: 2e-5
Optimizer: AdamW
Weight Decay: 0.01
Evaluation Strategy: Per epoch
The model is trained on the WNUT 17 dataset, which contains challenging named entities in social media and conversational text. The dataset provides annotations for named entity recognition, including entity categories such as:
person
location
corporation
product
creative-work
group
#Loading the Model
from transformers import DistilBertForTokenClassification, DistilBertTokenizerFast
import torch
model_name = "AventIQ-AI/distilbert-base-uncased_token_classification"
def predict_entities(text, model, tokenizer):
"""Predict Named Entities from the quantized model"""
inputs = tokenizer(text, return_tensors="pt", truncation=True, padding=True)
# Convert to FP32 if needed (for stability)
with torch.no_grad():
outputs = model(**inputs)
predictions = torch.argmax(outputs.logits.float(), dim=2) # Convert logits to float32
predicted_labels = [model.config.id2label[t.item()] for t in predictions[0]]
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
# Remove special tokens and align subwords
entities = []
current_entity = None
for token, label in zip(tokens, predicted_labels):
if token in [tokenizer.cls_token, tokenizer.sep_token, tokenizer.pad_token]:
continue
if token.startswith("##"): # Handle subwords
if current_entity:
current_entity["text"] += token[2:]
continue
if label == "O":
if current_entity:
entities.append(current_entity)
current_entity = None
else:
if label.startswith("B-"):
if current_entity:
entities.append(current_entity)
current_entity = {"text": token, "type": label[2:]}
elif label.startswith("I-") and current_entity:
current_entity["text"] += " " + token
if current_entity:
entities.append(current_entity)
return entities
test_sentence = ["Apple CEO Tim Cook announced the new iPhone 14 at their headquarters in Cupertino."]
for sentence in test_sentences:
print(f"\nInput: {sentence}")
entities = predict_entities(sentence, model, tokenizer)
print("Detected entities:")
for entity in entities:
print(f"- {entity['text']} ({entity['type']})")
print("-" * 50)
| Entity Type | Precision | Recall | F1 Score | Number of Entities |
|---|---|---|---|---|
| LOC (Location) | 91.46% | 92.07% | 91.76% | 3,000 |
| MISC (Miscellaneous) | 71.25% | 72.83% | 72.03% | 1,266 |
| ORG (Organization) | 89.83% | 93.02% | 91.40% | 3,524 |
| PER (Person) | 95.16% | 94.04% | 94.60% | 2,989 |
The Hugging Face's wnut_17 dataset was used, containing texts and their ner tags.
Post-training quantization was applied using PyTorch's built-in quantization framework to reduce the model size and improve inference efficiency.
.
├── model/ # Contains the quantized model files
├── tokenizer_config/ # Tokenizer configuration and vocabulary files
├── model.safetensors/ # Quantized Model
├── README.md # Model documentation
Contributions are welcome! Feel free to open an issue or submit a pull request if you have suggestions or improvements.