Downloads · 30 days
39
14% of all-time downloads
GivingTuesday/NTEE_category_tagging
NTEE_category_tagging is a machine learning model from GivingTuesday. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
- Dataset: Fine-tuning + Validation - Primary Labels: All Labels-Weights=1.0 - Secondary Labels: Foreign Affairs & National Security and Mutual & Membership Benefit-Weight=0.5 - Training Data: The Dataset used had ove…
Downloads · 30 days
39
14% of all-time downloads
All-time downloads
285
Public
Parameters
110M
876 MB on disk
Likes
0
Public
Click a slice to open those files.
.pth438 MB · 50%
From the Hugging Face model README
Important clarification on taxonomy and labels
This model does not implement the full official NTEE taxonomy across all levels. Instead, it predicts only a custom set of 28 cause-area categories, which are inspired by (but not identical to) NTEE “Level 2” major group categories. The 28 labels used here may differ from the canonical NTEE v2.0 definitions and counts (for example, NTEE’s 26 major groups).
References in this card to NTEE Level 1, Level 3, and Level 5 codes are provided for background and context only. The model does not directly predict those codes. All model outputs are limited to the 28 cause-area labels explicitly listed below under “Labels” / “Level 2 – Major Group Categories.”
from transformers import BertTokenizer, BertForSequenceClassification
import torch
tokenizer = BertTokenizer.from_pretrained("GivingTuesday/NTEE_category_tagging")
# OR this works too:
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("GivingTuesday/NTEE_category_tagging")
# num_labels: 28 because this will assign one of 28 NTEE codes
model = BertForSequenceClassification.from_pretrained("GivingTuesday/NTEE_category_tagging", num_labels=28)
model.eval()
# example texts
text = """NATIONAL CHURCH RESIDENCES OF SOUTH,PROVIDE HOUSING FOR LOW AND MODERATE INCOME PERSONS.,"THE SOLE PURPOSE IS TO PROVIDE ELDERLY PERSONS AND HANDICAPPED PERSONS WITH HOUSING FACILITIES AND SERVICES SPECIALLY DESIGNED TO MEET THEIR PHYSICAL, SOCIAL, AND PSYCHOLOGICAL NEEDS, AND TO PROMOTE THEIR HEALTH, SECURITY, HAPPINESS AND USEFULNESS IN LONGER LIVING, THE CHARGES FOR SUCH FACILITIES AND SERVICES TO BE PREDICATED UPON THE PROVISION, MAINTENANCE, AND OPERATION THEREOF ON A NONPROFIT BASIS."""
text = """NORTH CAROLINAS EASTERN ALLIANCE,"TO PROMOTE AND ENCOURAGE ECONOMIC DEVELOPMENT WITHIN EASTERN NORTH CAROLINA BY FOSTERING DEVELOPMENT PROJECTS TO PROVIDE LAND, BUILDINGS, FACILITIES, PROGRAMS, INFORMATION AND DATA SYSTEMS, AND INFRASTRUCTURE REQUIREMENTS FOR BUSINESS AND INDUSTRY WITHIN EASTERN NORTH CAROLINA.","HELPS TO RECRUIT NEW BUSINESSES INTO THE REGION, AS WELL AS HELP EXPAND EXISTING BUSINESSES BY ASSISTING WITH SITE LOCATIONS AND GRANTS. THEY ALSO WORK WITH THE LOCAL COMMUNITY COLLEGES TO EDUCATE THE CITIZENS OF THE REGION SO TO MAKE THE AREA MORE APPEALING TO POTENTIAL NEW BUSINESSES."""
inputs = tokenizer(text, return_tensors="pt")
#encoded_output = tokenizer.encode(text)
#print(tokenizer.convert_ids_to_tokens(encoded_output))
with torch.no_grad():
output = model(**inputs)
probabilities = torch.softmax(output.logits, dim=1)
predicted_class_index = torch.argmax(output.logits, dim=1)
predicted_label = model2.config.id2label[predicted_class_index.item()]
print(predicted_label)
Easier line-by-line code available here: Notebook with Usecase
Author: Edward Moore - GivingTuesday Data Commons
Note for external readers: Some Databricks links in this document point to internal notebooks and may not be accessible to people outside GivingTuesday.
The cause area segmentation uses a multi-stage NLP system to assign nonprofit organizations to detailed cause categories using mission statements and program descriptions. It supports top-level NTEE industry group codes and major group categories. The segmentation uses a classifier that integrates human-in-the-loop GPT labeling and fine-tunes a transformer model (BERT) to support multi-label prediction.
A cause area refers to a structured, human-interpretable label that reflects the primary social mission of a nonprofit organization. This taxonomy draws from the National Taxonomy of Exempt Entities (NTEE) codes, version 2.0.
The variables that are used in this segmentation include:
The classifier operates in three main phases:
A sample of 10,000 organizations from 2019 990 filings were selected for model training and evaluation. For each organization, a single text input was created combining EIN, Name, Mission, and up to three Activities.
Embedding-Based Candidate Selection:
Using an organization’s EIN, Name, Mission, and Activities along with top 3 candidate labels from the embedding model, OpenAI’s GPT-4-turbo model was prompted to choose the primary and secondary cause area classification. The top 3 candidate labels from the embedding model were included in the prompt to improve the accuracy and efficiency by providing context and guidance to the LLM. The system and user prompts used are as follows:
system_prompt = """
You are an expert in classifying nonprofits into the correct cause area category based on their mission statement and program activities.
Given a nonprofit's mission and activities, you must:
1. Choose the **primary** category code (from A-Z, BB, EE) that best matches.
2. Choose the **secondary** category code that could also apply (or null if none).
Special rules:
- If the nonprofit supports causes outside the US, primary must be 'Q'.
- If the nonprofit fits definitions for 'Y' (Mutual & Membership Benefit), primary must be 'Y'.
- Always return exactly **two codes**: primary, secondary.
- If no clear primary match: primary is 'Z' (Unknown).
- If no clear secondary: secondary is 'null'.
Respond only with the two codes, separated by a comma, e.g., 'B, P' or 'Z, null'.
Category Codes:
{cause_areas_map}
Category Definitions:
{category_definitions_text}
"""
user_prompt = """
Mission Statement and Activities:
{full_text}
Top {top_n} likely categories from embedding model: {top_matches}
What are the primary and secondary categories?
"""
The annotated dataset of 10,000 organizations was split in two phases:
The goal of this stage was to fine-tune a pre-trained BERT model (bert-base-uncased) to classify nonprofit organizations into one of 28 possible cause area labels.
Model Input:
Model Output:
The fine-tuned classifier outputs raw logits over the 28 cause area classes, which are converted to the top-2 predictions (primary and secondary), along with confidence scores (softmax probability). The validation dataset (800 orgs) is used to produce a weighted F1 score for evaluation of the fine-tuning.
Model Evaluation
The model was evaluated on the hold-out test dataset (2,000 orgs) using both macro and weighed F1 score.
| Model No. | MLflow Run No. | Dataset | Primary Labels | Secondary Labels | Weighted F1 (fine-tuning) | Macro F1 (holdout test) | Weighted F1 (holdout test) |
|---|---|---|---|---|---|---|---|
| 2 | 2 | fine-tuning | All labels - No weighting | Foreign Affairs & National Security and Mutual & Membership Benefit - No weighting | 0.79 | 0.62 | 0.78 |
| validation | All labels - No weighting | Foreign Affairs & National Security and Mutual & Membership Benefit - No weighting | |||||
| 3 | 3 | fine-tuning | All labels - No weighting | Excluded | 0.83 | 0.66 | 0.8 |
| validation | |||||||
| 6 | 6 | fine-tuning | All labels - Weight=1.0 | All labels - Weight=0.5 | 0.47 | 0.64 | 0.78 |
| validation | |||||||
| 7 | 7 | fine-tuning | All labels - Weight=1.0 | All labels - Weight=0.5 | 0.78 | 0.63 | 0.78 |
| validation | Excluded | ||||||
| 8 | 9 | fine-tuning | All labels - Weight=1.0 | All labels - Weight=0.7 | 0.85 | 0.64 | 0.76 |
| validation | Excluded | ||||||
| 9 | 10 | fine-tuning | All labels - Weight=1.0 | Foreign Affairs & National Security and Mutual & Membership Benefit - Weight=0.5 | 0.95 | 0.67 | 0.79 |
| validation | Excluded |