Downloads · 30 days
12
12% of all-time downloads
SallySims/equibert-category-tagger
equibert-category-tagger is a text classification model from SallySims. Use it when you need a label for a piece of text. The card lists the license as apache-2.0.
Model ID: SallySims/equibert-category-tagger
Downloads · 30 days
12
12% of all-time downloads
All-time downloads
98
Public
Parameters
310M
2.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.bin873 MB · 50%
How the weights are stored.
F16184M · 59%
From the Hugging Face model README
Model ID: SallySims/equibert-category-tagger
Multi-label classifier that tags any organisational document across 20 DEI domain categories.
| ID | Category |
|---|---|
| 0 | awareness_education |
| 1 | hiring_recruitment |
| 2 | workplace_culture |
| 3 | policies_processes |
| 4 | leadership_accountability |
| 5 | community_engagement |
| 6 | measurement_improvement |
| 7 | accessibility_inclusion |
| 8 | data_analytics |
| 9 | performance_career |
| 10 | everyday_team_practices |
| 11 | conflict_resolution |
| 12 | learning_development |
| 13 | supplier_economic_equity |
| 14 | marketing_branding |
| 15 | global_crosscultural |
| 16 | events_activities |
| 17 | technology_ai_ethics |
| 18 | governance_strategy |
| 19 | individual_actions |
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("SallySims/equibert-category-tagger")
text = "We have introduced structured interviews to reduce hiring bias."
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=128)
# probs = torch.sigmoid(model(**inputs).logits)
# tags = [id2label[i] for i, p in enumerate(probs[0]) if p >= 0.5]
EquiBERT is a multi-task DEI (Diversity, Equity and Inclusion) transformer built on a dual-encoder backbone that fuses RoBERTa-base and DeBERTa-v3-base via a learned weighted sum (α parameter). The fused representation is fed into task-specific heads covering 17 distinct DEI analysis tasks.
Organisation: SallySims Framework: PyTorch + HuggingFace Transformers Backbone: RoBERTa-base + DeBERTa-v3-base (dual encoder, fused) Language: English Domain: Organisational DEI text — HR communications, policies, job descriptions, performance reviews, leadership statements, reports
Input Text
│
├──▶ RoBERTa-base encoder ──▶ Linear projection
│ │
└──▶ DeBERTa-v3-base encoder ──▶ Linear projection
│
Weighted fusion (learned α)
│
Layer Norm + Dropout
│
Task-specific head (see below)
Trained on synthetic DEI organisational text generated by the EquiBERT synthetic data pipeline, covering 20 DEI categories across HR, policy, leadership, and workforce analytics domains. For production use, fine-tune on real labelled DEI data.
If you use EquiBERT in your research, please cite:
@misc{equibert2024,
author = {SallySims},
title = {EquiBERT: A Multi-Task DEI Transformer},
year = {2024},
publisher = {HuggingFace},
url = {https://huggingface.co/SallySims}
}
| Model | Task | Primary Metric |
|---|---|---|
| equibert-bias-classifier | Bias Detection | Macro F1 |
| equibert-microaggression | Microaggression Detection | Macro F1 |
| equibert-category-tagger | DEI Category Tagging | Macro F1 |
| equibert-event-exclusion | Event Exclusion Classification | Macro F1 |
| equibert-inclusive-language | Inclusive Language Scoring | Span F1 |
| equibert-review-auditor | Performance Review Auditing | Span F1 |
| equibert-washing-detector | DEI Washing Detection | MAE |
| equibert-framing-scorer | Report Framing Scoring | MAE |
| equibert-awareness-scorer | DEI Awareness Scoring | MAE |
| equibert-similarity | Semantic Similarity | Accuracy |
| equibert-ner | DEI Entity Recognition | Span F1 |
| equibert-relation-extraction | Relation Extraction | Macro F1 |
| equibert-qa | Extractive QA | Span EM |
| equibert-search | Semantic Search | MRR@10 |
| equibert-nli | NLI / Textual Entailment | Macro F1 |
| equibert-generator | DEI Text Generation | ROUGE-L |