Downloads · 30 days
27
4% of all-time downloads
baobabtech/water-conflict-classifier
water-conflict-classifier is a text classification model from baobabtech. Use it when you need a label for a piece of text. It is set up for setfit. The card lists the license as cc-by-nc-4.0.
This experimental research draws on Pacific Institute's Water Conflict Chronology, which tracks water-related conflicts spanning over 4,500 years of human history. The work is conducted independently and is not affili…
Downloads · 30 days
27
4% of all-time downloads
All-time downloads
610
Public
Parameters
33.4M
3.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors133 MB · 99%
From the Hugging Face model README
This experimental research draws on Pacific Institute's Water Conflict Chronology, which tracks water-related conflicts spanning over 4,500 years of human history. The work is conducted independently and is not affiliated with Pacific Institute.
This model is designed to assist researchers in classifying water-related conflict events at scale using tiny/small models that can classify 100s of headlines per second.
The Pacific Institute maintains the world's most comprehensive open-source record of water-related conflicts, documenting over 2,700 events across 4,500 years of history. This is not a commercial product and is not intended for commercial use.
This SetFit-based model classifies news headlines about water-related conflicts into three categories:
These categories align with the Pacific Institute's Water Conflict Chronology framework for understanding how water intersects with security and conflict.
from setfit import SetFitModel
# Load the trained model from HF Hub
model = SetFitModel.from_pretrained("baobabtech/water-conflict-classifier")
# Predict on headlines
headlines = [
"Military attack workers at the Kajaki Dam in Afghanistan",
"New water treatment plant opens in California"
]
predictions = model.predict(headlines)
print(predictions)
# Output: [[1, 1, 0], [0, 0, 0]]
# Format: [Trigger, Casualty, Weapon]
The model returns a list of binary predictions for each label:
label_names = ['Trigger', 'Casualty', 'Weapon']
for headline, pred in zip(headlines, predictions):
labels = [label_names[i] for i, val in enumerate(pred) if val == 1]
print(f"Headline: {headline}")
print(f"Labels: {', '.join(labels) if labels else 'None'}")
print()
import pandas as pd
# Load your data
df = pd.read_csv("your_headlines.csv")
# Predict in batches
predictions = model.predict(df['headline'].tolist())
# Add predictions to dataframe
df['trigger'] = [p[0] for p in predictions]
df['casualty'] = [p[1] for p in predictions]
df['weapon'] = [p[2] for p in predictions]
| Headline | Trigger | Casualty | Weapon |
|---|---|---|---|
| "Armed groups blow up water pipeline in Iraq" | ✓ | ✓ | ✓ |
| "New water treatment plant opens in California" | ✗ | ✗ | ✗ |
| "Protests erupt over dam construction in Ethiopia" | ✓ | ✗ | ✗ |
Evaluated on a held-out test set of 519 samples (30% of total data, stratified by label combinations).
| Metric | Score |
|---|---|
| Exact Match Accuracy | 0.8015 |
| Hamming Loss | 0.0906 |
| F1 (micro) | 0.8530 |
| F1 (macro) | 0.7987 |
| F1 (samples) | 0.7028 |
| Label | Precision | Recall | F1 | Support |
|---|---|---|---|---|
| Trigger | 0.8844 | 0.8793 | 0.8818 | 174 |
| Casualty | 0.8843 | 0.9185 | 0.9011 | 233 |
| Weapon | 0.4941 | 0.8077 | 0.6131 | 52 |
All training runs are automatically tracked in a public dataset for experiment comparison:
You can explore past experiments and compare model performance across versions using the evals dataset.
Pacific Institute (2025). Water Conflict Chronology. Pacific Institute, Oakland, CA.
https://www.worldwater.org/water-conflict/
Armed Conflict Location & Event Data Project (ACLED).
https://acleddata.com/
Note: Training negatives include synthetic "hard negatives" - peaceful water-related news (e.g., "New desalination plant opens", "Water conservation conference") to prevent false positives on non-conflict water topics.
This model is part of independent experimental research drawing on the Pacific Institute's Water Conflict Chronology. The Pacific Institute maintains the world's most comprehensive open-source record of water-related conflicts, documenting over 2,700 events across 4,500 years of history.
Project Links:
This classifier demonstrates an intentional approach to building AI systems with limited data using SetFit - a framework for few-shot learning with sentence transformers. Rather than defaulting to massive language models (GPT, Claude, or 100B+ parameter models) for simple classification tasks, we fine-tune small, efficient models (e.g., BAAI/bge-small-en-v1.5 with ~33M parameters) on a focused dataset.
Why this matters: The industry has normalized using trillion-parameter models to classify headlines, answer simple questions, or categorize text - tasks that don't require world knowledge, reasoning, or generative capabilities. This is computationally wasteful and environmentally costly. A properly fine-tuned small model can achieve comparable or better accuracy while using a fraction of the compute resources.
Our approach:
This is not about avoiding large models altogether - they're invaluable for complex reasoning tasks. But for targeted classification problems with labeled data, fine-tuning remains the professional, responsible choice.
You can train your own version using the published package.
Package includes:
Source code: https://github.com/baobabtech/waterconflict/tree/main/classifier
PyPI: https://pypi.org/project/water-conflict-classifier/
# Install package
pip install water-conflict-classifier
# Or install from source for development
git clone https://github.com/baobabtech/waterconflict.git
cd waterconflict/classifier
pip install -e .
# Train locally
python train_setfit_headline_classifier.py
For cloud training on HuggingFace Jobs infrastructure, see the scripts folder in the repository.
Copyright © 2025 Baobab Tech
This work is licensed under the Creative Commons Attribution-NonCommercial 4.0 International License.
You are free to:
Under the following terms:
If you use this model in your work, please cite:
@misc{{waterconflict2025,
title={{Water Conflict Multi-Label Classifier}},
author={{Independent Experimental Research Drawing on Pacific Institute Water Conflict Chronology}},
year={{2025}},
howpublished={{\url{{https://huggingface.co/{model_repo}}}}},
note={{Training data from Pacific Institute Water Conflict Chronology and ACLED}}
}}
Please also cite the Pacific Institute's Water Conflict Chronology:
@misc{{pacificinstitute2025,
title={{Water Conflict Chronology}},
author={{Pacific Institute}},
year={{2025}},
address={{Oakland, CA}},
url={{https://www.worldwater.org/water-conflict/}},
note={{Accessed: [access date]}}
}}