Downloads · 30 days
28
67% of all-time downloads
int-brain-lab/ea-decoder-channel-xgboost
ea-decoder-channel-xgboost is a machine learning model from int-brain-lab. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for xgboost. The card lists the license as cc-by-4.0.
Predicts the brain region of each Neuropixels recording channel from electrophysiological features alone -- no histology required. Trained by the International Brain Laboratory on the Ephys Atlas feature release 2026W26.
Downloads · 30 days
28
67% of all-time downloads
All-time downloads
42
Public
Repo size
58.3 MB
Likes
1
Public
Click a slice to open those files.
.ubj28.4 MB · 97%
From the Hugging Face model README
Predicts the brain region of each Neuropixels recording channel from electrophysiological
features alone -- no histology required. Trained by the
International Brain Laboratory on the Ephys Atlas
feature release 2026_W26.
What you can and cannot do without IBL access. The model runs for anyone. Computing the input features from raw Neuropixels data needs
ibllib/ibl-neuropixel, and the raw data itself is IBL-hosted. To try the model immediately, use the bundled sample underexample/-- no account, no raw data, no S3.
import pandas as pd
from ephysatlas import load_pretrained
model = load_pretrained("int-brain-lab/ea-decoder-channel-xgboost", revision="2026_W26")
df = pd.read_parquet("example/features_sample.parquet") # or your own features
out = model.predict(df)
print(out[["predicted_acronym", "prediction_probability", "fold_agreement"]].head())
load_pretrained is the entry point for every ephysatlas model, whatever its family — it reads
ephysatlas_model.json and returns the right wrapper. Use it rather than importing a concrete
class, so your code keeps working as the package evolves.
predict returns one row per input channel, indexed identically to the input:
predicted_acronym, its Allen predicted_atlas_id, the fold-averaged
prediction_probability, a fold_agreement column (fraction of the 5 folds voting
for the winner -- the natural uncertainty signal), and a p_<acronym> column per class.
The prediction columns are namespaced so that df.join(out) works: the feature table already
carries histology-derived acronym / atlas_id columns, and predictions must not shadow them.
By default predict averages the 5 fold models; the single all-data model.ubj at
the repo root is not used. Pass estimator="global" to use it instead — one fifth of the
inference cost, but fold_agreement then comes back as NaN, since no folds were consulted. The
two modes disagree on a small fraction of channels, so pick one per analysis.
ephysatlas_model.json under inputs.features. Every
one must be present; predict raises and names anything missing.(pid, channel), one row per recording channel.2026_W26 -- that is, the
raw_ephys_features_denoised.pqt table produced by the Ephys Atlas aggregation pipeline,
as loaded by ephysatlas.data.read_features_from_disk. Units are baked into that table by
the pipeline (RMS features in dB, spike_count in log2), so feeding raw features, or
features from a vintage whose units differ, produces confident nonsense. Run
model.selftest() to confirm your install reproduces the shipped output before trusting it.Pooled out-of-fold accuracy: 0.5959 over 13 Cosmos regions,
761 insertions. Splits are by insertion (pid), so no channel from a test
insertion appears in training. See confusion_matrix.png.
Note this figure scores each channel with the one fold that held it out, whereas the default
estimator="ensemble" averages all 5 folds. On genuinely unseen data the ensemble is
typically the marginally better estimator, so treat the number as slightly conservative.
Cosmos is a coarse parcellation. Predictions are per-channel and spatially
unregularised -- neighbouring channels can disagree.Pin the revision. revision="2026_W26" is an immutable tag. Omitting revision resolves
to main, which tracks whichever model is currently recommended and will change when a new
feature vintage is published — fine for a first look, not for anything you publish or re-run.
ephysatlas_model.json records the training-time environment (xgboost, scikit-learn, numpy,
ephysatlas, python) and random_seed. Verify your install reproduces the shipped output:
model.selftest()
Note scikit-learn<1.9 is required (1.9 broke OneToOneFeatureMixin.get_feature_names_out,
which the feature transformer relies on).
Please cite the International Brain Laboratory Ephys Atlas. Model id 2026_W26_Cosmos_agitated-latte-sandpiper,
feature vintage 2026_W26.