Downloads · 30 days
0
datamatters24/research-document-archive
research-document-archive is a machine learning model from datamatters24. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
234,630 declassified U.S. government documents processed through a 13-step ML pipeline. 3.2 million pages OCR'd, 31 million named entities extracted and linked, 288 topic clusters identified.
Downloads · 30 days
0
Access
Public
Updated Apr 9, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.ipynb191 KB · 48%
From the Hugging Face model README
234,630 declassified U.S. government documents processed through a 13-step ML pipeline. 3.2 million pages OCR'd, 31 million named entities extracted and linked, 288 topic clusters identified.
Live platform: tanglewoodapp.com
| Collection | Documents | Pages | Size |
|---|---|---|---|
| House Resolutions | 181,092 | 2,719,832 | 34.2 GB |
| JFK Assassination Records | 35,979 | 241,860 | 22.5 GB |
| CIA Stargate Program | 13,937 | 100,056 | 5.4 GB |
| CIA MKUltra | 1,936 | 64,244 | 3.4 GB |
| CIA Declassified | 1,605 | 29,744 | 2.4 GB |
| Lincoln Archives | 21 | 9,330 | 962.9 MB |
| Stamp | Count |
|---|---|
| UNCLASSIFIED | 16,501 |
| SECRET | 13,736 |
| CLASSIFIED | 10,730 |
| EXEMPT | 6,739 |
| CONFIDENTIAL | 5,554 |
| RESTRICTED | 4,722 |
from datasets import load_dataset
ds = load_dataset("datamatters24/research-document-archive")
# Filter by collection
jfk = ds.filter(lambda x: x["collection"] == "jfk_assassination")
# Search by entity
cia_docs = ds.filter(lambda x: "CIA" in x["entities"])
All documents are public record obtained from:
@misc{rubin2026researcharchive,
author = {Rubin, Theodore},
title = {Research Document Archive: ML Pipeline for Declassified U.S. Government Documents},
year = {2026},
publisher = {HuggingFace},
url = {https://huggingface.co/datasets/datamatters24/research-document-archive}
}