Downloads · 30 days
0
emperor-mew/voidly-classifier-v3.3
voidly-classifier-v3.3 is a tabular classification model from emperor-mew. Use it for the tabular classification task on the model card, and read the license before you ship it in a product. It is set up for scikit-learn. The card lists the license as cc-by-4.0.
Version: v3.3 | Trained: 2026-05-21T03:01:46.793987+00:00 | License: CC BY 4.0
Downloads · 30 days
0
Access
Public
Updated May 22, 2026
Repo size
470 KB
Likes
0
Public
Click a slice to open those files.
.pkl470 KB · 92%
From the Hugging Face model README
Version: v3.3 | Trained: 2026-05-21T03:01:46.793987+00:00 | License: CC BY 4.0
Country-day censorship classifier with regime-similarity-weighted geographic contagion features. Promoted 2026-05-21.
Given a (country, day) feature vector — measurement volume, anomaly rate,
probe agreement, and three contagion neighbour aggregates — return a calibrated
P(censorship event). Surfaced live at
GET https://api.voidly.ai/v1/classifier/score/{cc}.
| Metric | Value |
|---|---|
stratified_f1 | 0.7289 |
stratified_auc | 0.8991 |
loco_median_f1 | 0.8696 |
loco_mean_f1 | 0.7109 |
loco_n_countries | 127 |
n_features | 16 |
n_samples | 4237 |
n_positive | 1116 |
n_countries | 131 |
Stratified split = single random 80/20. LOCO = leave-one-country-out across 127 countries, the honest cross-country generalization number.
Total samples: 4,237 country-day rows
Positives: 1,116 (26.3%)
Unique countries: 131
Provenance: OONI + CensoredPlanet + IODA + Voidly probe network (84K evidence rows)
Labels exclude IODA disruption incidents (fix 2026-05-21) — those are real network outages but not all are censorship
16 inputs: 13 base + 3 regime-similarity-weighted contagion neighbours.
anomaly_ratemeasurement_countspike_magnitudeday_of_weekmonthis_weekendrate_count_interactionprobe_block_rateprobe_node_countprobe_avg_confidenceprobe_agreementrate_spike_interactionhigh_evidenceneighbor_block_rate_7dneighbor_incident_count_7dneighbor_max_anomaly_7danomaly_rate (importance 0.2210)month (importance 0.2022)measurement_count (importance 0.1696)neighbor_max_anomaly_7d (importance 0.0887)neighbor_incident_count_7d (importance 0.0770)neighbor_block_rate_7d (importance 0.0764)rate_count_interaction (importance 0.0561)day_of_week (importance 0.0389)spike_magnitude (importance 0.0354)rate_spike_interaction (importance 0.0319)is_weekend (importance 0.0013)probe_avg_confidence (importance 0.0009)high_evidence (importance 0.0003)probe_block_rate (importance 0.0002)probe_node_count (importance 0.0001)probe_agreement (importance 0.0000)
# Build script (training data → fitted pickle + per-country thresholds)
python3 scripts/build-classifier-v3.3-regime-weighted.py
Algorithm: GradientBoostingClassifier (sklearn)
Promoted artifact: /opt/voidly-ai/models/censorship_classifier_v3_promoted.pkl
Backup pre-promote: .pkl.bak.v3.1-2026-05-21
Per-country thresholds: ml-deploy/classifier_v3.3_per_country_thresholds.json
@misc{voidly_voidly_classifier_v3.3,
title = {Voidly Atlas: voidly-classifier-v3.3 (v3.3)},
author = {Voidly},
year = {2026},
url = {https://huggingface.co/emperor-mew/voidly-classifier-v3.3},
note = {Open censorship-research ML stack. CC BY 4.0.}
}