Downloads · 30 days
0
Noptus/numerai-weekly-v4
numerai-weekly-v4 is a tabular regression model from Noptus. Use it for the tabular regression task on the model card, and read the license before you ship it in a product.
This repository publishes the exact model bundle currently used by Noptus' validated Numerai submission pipeline. It is intended as a reproducible research artifact and a starting point for ensemble-diversity work—not…
Downloads · 30 days
0
Access
Public
Updated Aug 14, 2026
Repo size
149 MB
Likes
0
Public
Click a slice to open those files.
.pkl149 MB · 100%
From the Hugging Face model README
This repository publishes the exact model bundle currently used by Noptus' validated Numerai submission pipeline. It is intended as a reproducible research artifact and a starting point for ensemble-diversity work—not as investment advice or a promise of tournament performance.
v4-20260801-1804551333v5.2medium features plus 8 public benchmark-model columns79db5f41f3506e8e10a8b96c60a927f6a9ca202e49e03304ceaa9d304116b8d9The bundle was promoted over the previous local champion on a 57-era untouched holdout:
| Metric | v4 | previous champion |
|---|---|---|
| Mean Numerai CORR | 0.010453 | 0.001821 |
| Sharpe | 0.7712 | 0.1499 |
| Positive-era consistency | 75.44% | 52.63% |
| Maximum drawdown proxy | -0.01360 | -0.02329 |
These are historical offline measurements, not live-performance guarantees. The model remains experimental, can decay under regime change, and should not be used to make financial decisions.
Install the pinned runtime dependencies:
pip install -r requirements.txt
Download the files and run inference on the public Numerai live and benchmark-model frames:
from huggingface_hub import hf_hub_download
import joblib
import pandas as pd
from inference import predict_ranked
repo_id = "Noptus/numerai-weekly-v4"
model_path = hf_hub_download(repo_id, "ensemble_v4.pkl")
bundle = joblib.load(model_path)
live = pd.read_parquet("live.parquet")
benchmarks = pd.read_parquet("live_benchmark_models.parquet")
benchmark_columns = [c for c in benchmarks.columns if c != "era"]
live = live.join(benchmarks[benchmark_columns], how="left")
submission = pd.DataFrame(
{"prediction": predict_ranked(live, bundle)},
index=live.index,
)
submission.index.name = "id"
submission.to_csv("predictions.csv")
The included command-line entry point performs the same base inference:
python inference.py \
--model ensemble_v4.pkl \
--live live.parquet \
--benchmarks live_benchmark_models.parquet \
--output predictions.csv
The production system derives several slot-specific submissions by applying different feature and benchmark neutralization settings after this base ensemble. Those operational credentials and live submissions are intentionally excluded.
manifest.json records the source revision, metric split, dependency versions, and hashes. Numerai datasets, target labels, live predictions, API credentials, and staking information are not included.
The checkpoint uses Python pickle serialization because it contains native LightGBM and CatBoost estimators. Pickle can execute code while loading: verify the SHA-256 and load only artifacts you trust. Reconstructing the component estimators in native, non-pickle formats is planned for a later release.
Three subsequent frozen-prediction experiments did not displace this champion:
Negative results are retained because they narrow the useful next step: seek genuinely different input signal—currently the official v5.3 feature families—rather than adding more tree implementations over the same v5.2 inputs.
No explicit model or software license has been selected for this first release. Numerai data and benchmark-model files are governed by their own terms and are not redistributed here. Verify the applicable terms before reuse or redistribution.