Downloads · 30 days
9
11% of all-time downloads
ILoveBacteria/bilstm-sequence-labeler
bilstm-sequence-labeler is a machine learning model from ILoveBacteria. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This repository hosts a custom Bidirectional LSTM (BiLSTM) sequence labeling model trained on the standard CoNLL-2003 dataset. The network uses a shared embedding layer and a bidirectional recurrent backbone to perfor…
Downloads · 30 days
9
11% of all-time downloads
All-time downloads
79
Public
Parameters
914K
3.7 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.7 MB · 91%
From the Hugging Face model README
This repository hosts a custom Bidirectional LSTM (BiLSTM) sequence labeling model trained on the standard CoNLL-2003 dataset. The network uses a shared embedding layer and a bidirectional recurrent backbone to perform three tasks simultaneously:
For full training details, notebooks, and source files, visit the GitHub Repository.
hidden_dim=128, yielding a 256-dimensional concatenated sequence representation).ignore_index=-100 ensuring pad tokens do not affect loss or performance evaluations.
The model achieved an overall Accuracy of 94% on the evaluation set. Below is the detailed token-level classification report across all 9 target categories:
| Class / Tag ID | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|
| 0 (e.g., O) | 0.96 | 0.99 | 0.98 | 81,082 |
| 1 | 0.89 | 0.73 | 0.80 | 3,459 |
| 2 | 0.90 | 0.78 | 0.83 | 2,463 |
| 3 | 0.83 | 0.66 | 0.74 | 3,002 |
| 4 | 0.81 | 0.69 | 0.75 | 1,586 |
| 5 | 0.84 | 0.78 | 0.81 | 3,505 |
| 6 | 0.62 | 0.68 | 0.65 | 514 |
| 7 | 0.82 | 0.68 | 0.74 | 1,624 |
| 8 | 0.66 | 0.63 | 0.65 | 562 |
| Macro Average | 0.82 | 0.74 | 0.77 | 97,797 |
| Weighted Average | 0.94 | 0.94 | 0.94 | 97,797 |

ORG for organizations, LOC for locations, or PER for persons).displacy.render specifically highlights the structural Named Entity Recognition (NER) tag segments extracted by the token prediction head, these entity boundaries are intrinsically anchored by the shared hidden representations refined by the parallel POS and Chunk tagging heads during the joint optimization process.Since this model implements a completely custom neural architecture, make sure to pass trust_remote_code=True when loading via AutoModel.
import torch
import json
import urllib.request
from transformers import AutoConfig, AutoModel
# 1. Define repository identifier and vocabulary URL
repo_id = "ILoveBacteria/bilstm-sequence-labeler"
vocab_url = "https://huggingface.co/ILoveBacteria/bilstm-sequence-labeler/resolve/main/vocab.json"
# 2. Load custom model architecture and weights
model = AutoModel.from_pretrained(repo_id, trust_remote_code=True)
model.eval()
# 3. Download and load the vocabulary file directly from the link
with urllib.request.urlopen(vocab_url) as response:
vocab = json.loads(response.read().decode("utf-8"))
token2id = vocab["token2id"]
# Simple inference example
text = "EU rejects German call to boycott British lamb ."
tokens = text.split()
input_ids = [token2id.get(token, token2id["<unk>"]) for token in tokens]
input_tensor = torch.tensor([input_ids]) # Add batch dimension
with torch.no_grad():
outputs = model(input_tensor)
# Extract predictions for tasks
ner_predictions = outputs["ner"].argmax(dim=-1)[0]
print("Predicted NER IDs:", ner_predictions.tolist())
If you are running inference within a Jupyter Notebook or Google Colab, you can use the script below to generate a beautiful, interactive visual layout of the predicted Named Entities using SpaCy's displacy engine.
from spacy import displacy
# 1. Map for CoNLL-2003 NER indices to human-readable labels
ner_labels = {
0: "O", 1: "B-PER", 2: "I-PER", 3: "B-ORG", 4: "I-ORG",
5: "B-LOC", 6: "I-LOC", 7: "B-MISC", 8: "I-MISC"
}
# 2. Convert predicted tensor to a flat NumPy array or list
preds = ner_predictions.cpu().numpy()
# 3. Reconstruct the string text and track exact character offsets for displaCy
text = " ".join(tokens)
ents = []
char_offsets = []
cursor = 0
for tok in tokens:
char_offsets.append((cursor, cursor + len(tok)))
cursor += len(tok) + 1 # +1 accounts for the space between tokens
# 4. Construct the entity metadata required by displaCy
for i, pred_idx in enumerate(preds):
label = ner_labels.get(int(pred_idx), "O")
if label != "O":
start, end = char_offsets[i]
ents.append({"start": start, "end": end, "label": label})
# 5. Pack the document structural payload
doc_data = {
"text": text,
"ents": ents
}
# 6. Render the visualization inline (Jupyter Notebook / Google Colab)
displacy.render(doc_data, style="ent", manual=True, jupyter=True)