Downloads · 30 days
0
ARotting/causal-forge-estimators
causal-forge-estimators is a machine learning model from ARotting. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Causal Forge is a ground-truth causal inference laboratory. Its structural causal model creates confounding, heterogeneous treatment effects, nonlinear outcomes, and known counterfactuals. The benchmark measures wheth…
Downloads · 30 days
0
Access
Public
Updated Jul 30, 2026
Repo size
29.7 KB
Likes
0
Public
Click a slice to open those files.
.parquet28.2 KB · 57%
From the Hugging Face model README
Causal Forge is a ground-truth causal inference laboratory. Its structural causal model creates confounding, heterogeneous treatment effects, nonlinear outcomes, and known counterfactuals. The benchmark measures whether estimators recover the population average treatment effect rather than merely predicting observed outcomes.
The evaluation compares:
The exact potential outcomes and treatment probabilities are retained only because this is a synthetic benchmark. They make estimator bias directly measurable.
The benchmark ran 100 independent replications with 3,000 observations and five-fold cross-fitting in each replication.
| Nuisance-model regime | Naive MAE | IPW MAE | Outcome MAE | AIPW MAE | AIPW bias |
|---|---|---|---|---|---|
| Both correct | 1.1109 | 0.0644 | 0.0388 | 0.0426 | 0.0004 |
| Propensity misspecified | 1.1109 | 0.1342 | 0.0388 | 0.0414 | 0.0008 |
| Outcome misspecified | 1.1109 | 0.0644 | 0.2688 | 0.0481 | 0.0078 |
| Both misspecified | 1.1109 | 0.1342 | 0.2688 | 0.3953 | 0.3953 |
This demonstrates the intended double-robustness boundary: AIPW remains accurate when either the treatment or outcome nuisance model is correct, but not when both are wrong. Results are Monte Carlo measurements on this synthetic SCM, not claims about arbitrary real-world observational data.
uv run python projects/causal-forge/train.py