Downloads · 30 days
27.7K
38% of all-time downloads
theforecastingcompany/t0-alpha
t0-alpha is a time series forecasting model from theforecastingcompany. Use it for the time series forecasting task on the model card, and read the license before you ship it in a product. It is set up for tfc-t0. The card lists the license as apache-2.0.
<p align="center" <img class="dark:hidden" src="https://www.theforecastingcompany.com/logo/logohorizontallight.png" alt="The Forecasting Company" width="280" / <img class="hidden dark:block" src="https://www.theforeca…
Downloads · 30 days
27.7K
38% of all-time downloads
All-time downloads
73.5K
Public
Parameters
102M
820 MB on disk
Likes
53
Trending 1
Click a slice to open those files.
.safetensors407 MB · 99%
From the Hugging Face model README
t0-alphat0-alpha is an open-weights time-series forecasting foundation model from The Forecasting Company.
t0 is a transformer-based model that produces probabilistic multi-horizon forecasts and natively operates on multiple covariates. t0-alpha is the first public iteration of the model.
You can use t0 on Retrocast, The Forecasting Company's platform for forecasting on your own data and comparing forecasts across open-weight models.
Model family: t0-alpha (PyTorch/MLX) · ONNX FP16 · ONNX INT8 · Collection

t0 forecasting French national electricity demand in Retrocast. Data: Enedis open data.
t0-alphat0tfc-t0tfc-t0-mlxt0-alpha is an alpha release intended for research, experimentation, and applied forecasting evaluation.
t0-alpha is intended for probabilistic time-series forecasting. It can be used for univariate and multivariate forecasting, forecasting with historical or known-future covariates and multi-horizon forecasting.
Known-future covariates can include calendar features, planned events, holidays, promotions, weather forecasts, or other external signals available over the forecast horizon.
Forecasts should be treated as probabilistic estimates, not guarantees.
t0 leverages covariate information, in the past and future when available, to improve its forecast.
| Without covariates | With covariates |
|---|---|
![]() | ![]() |
Data: Medic'AM, monthly drug reimbursements from the French national health insurance.
The Quickstart below shows the API for both a plain univariate forecast and a multivariate forecast that conditions on historical and known-future covariates.
Choose a runtime for the same original t0-alpha checkpoint:
| Runtime | Best for | Install |
|---|---|---|
| PyTorch | Broad hardware support and the PyTorch ecosystem | pip install tfc-t0 |
| MLX | Local, inference-only use on Apple silicon | pip install tfc-t0-mlx |
| ONNX FP16 | Accelerator-oriented local and edge deployments | t0-alpha-onnx-fp16 |
| ONNX INT8 | CPU and in-browser inference | t0-alpha-onnx-int8 |
| Managed API | Hosted inference without local weights | theforecastingcompany SDK |
pip install tfc-t0
Requirements:
>=3.10>=2.4Optional extras:
pip install "tfc-t0[evaluation]"
pip install "tfc-t0[plot]"
pip install tfc-t0-mlx
The MLX package uses the same model repository, loads its safetensors directly and does not install PyTorch.
These weights are public — no authentication is needed to download them.
The simplest path is a univariate forecast through predict:
import torch
from t0 import T0Forecaster
model = T0Forecaster.from_pretrained("theforecastingcompany/t0-alpha").eval()
context = torch.randn(4, 512) # 4 series, 512 past timesteps
out = model.predict(context, horizon=64, quantile_levels=[0.1, 0.5, 0.9])
out.quantiles # (4, 64, 3)
out.median # (4, 64)
predict accepts PyTorch tensors and NumPy arrays.
The MLX runtime deliberately follows the same forecasting interface:
import numpy as np
from t0_mlx import T0Forecaster
model = T0Forecaster.from_pretrained("theforecastingcompany/t0-alpha").eval()
context = np.random.randn(4, 512).astype(np.float32)
out = model.predict(context, horizon=64, quantile_levels=[0.1, 0.5, 0.9])
out.quantiles.shape # (4, 64, 3)
out.median.shape # (4, 64)
See T0 for MLX for feature coverage, compilation guidance and reproducible Apple-silicon benchmarks.
Anything known over the past goes in context. Alongside the target, extra variates attend to it and are forecast together. Anything known over the future, such as calendar features, planned promotions, or weather forecasts, goes in future_covariates, shaped [B, F, context + horizon]. The model conditions on it but does not forecast it.
import torch
from t0 import T0Forecaster
model = T0Forecaster.from_pretrained("theforecastingcompany/t0-alpha").eval()
context = torch.randn(2, 512) # 2 series, 512 past timesteps
future_covariates = torch.randn(2, 3, 512 + 64) # 3 covariates known over context + horizon
out = model.predict(
context,
horizon=64,
quantile_levels=[0.1, 0.5, 0.9],
future_covariates=future_covariates,
)
out.quantiles # (2, 64, 3)
out.median # (2, 64)
import numpy as np
from t0 import T0Forecaster, batch_series
model = T0Forecaster.from_pretrained("theforecastingcompany/t0-alpha").eval()
daily = np.random.randn(180) # one series, 180 past timesteps
store = np.random.randn(2, 96) # one series of 2 variates, 96 past timesteps
hourly = np.random.randn(1024) # one series, 1024 past timesteps
context, mask, group_ids = batch_series([daily, store, hourly])
context.shape # (4, 1024) — variates stacked, right-aligned to the longest series
group_ids # [0, 1, 1, 2] — `store`'s two variates are forecast jointly
out = model.predict(context, horizon=24, quantile_levels=[0.1, 0.5, 0.9], mask=mask, group_ids=group_ids)
out.quantiles # (4, 24, 3)
out.median[0] # the 24-step median forecast for `daily`
Integrations that prepare complete T0 inputs, including known-future covariates, can batch the native representation directly:
from t0 import TimeSeries
first = TimeSeries.from_array(context_1, future_covariates_1)
second = TimeSeries.from_array(context_2, future_covariates_2)
batch = TimeSeries.batch([first, second])
out = model.predict(
batch,
horizon=64,
context_length=max(context_1.shape[-1], context_2.shape[-1]),
)
Here each context includes its batch axis, for example [1, V, T], and each
known-future input is [1, F, T + horizon]. The output is ordered by the
flattened target rows in batch.
TimeSeriesTimeSeries is the model's native input. It holds target rows, known-future
covariate rows, a mask and group ids, all on one width. predict builds one for
you from a raw array. You only need to construct one yourself to batch inputs of
different widths, or to call forward directly.
from t0 import TimeSeries
# context only, with `horizon` marking the region to predict
model_input = TimeSeries.from_array(context, horizon=24) # context: [B, V, T]
# with known-future covariates, whose width sets the horizon
model_input = TimeSeries.from_array(context, future_covariates) # covariates: [B, F, T + 24]
out = model.predict(model_input, horizon=24, quantile_levels=[0.1, 0.5, 0.9])
predict infers context_length from where the forecast region starts. Pass it
explicitly when batching series of different widths. forward takes the same
TimeSeries and runs a single differentiable pass over it, with no rollout. That
is the entry point for fine-tuning.
For efficient inference at scale, look at Retrocast.
context may be shaped (B, T) for batched univariate forecasting.context may also be shaped (T,), which is promoted to a single-row batch.context may be shaped (B, V, T) for multiple target variates.future_covariates, when provided, should be shaped (B, F, context + horizon).mask, when provided, holds MaskType values shaped like context: MISSING for an absent observation, PAD for a cell that only widens a shorter series out to the batch's width.context is read as an absent observation. Padding is the case NaN cannot express, so a batch of unequal-length series needs a mask (or batch_series) to declare it.PAD stay out of attention.group_ids, when provided, holds one id per row of the context. Rows sharing an id are variates of one series and are forecast jointly.group_ids cannot be combined with future_covariates, which are addressed per sample.future_covariates is treated as missing.horizon must be at least 1.(0, 1).0.1, 0.25, 0.5, 0.75, and 0.9.float32 tensors on the model's device.t0 is a decoder-style patch transformer.
It encodes each patch from values, within-patch time index, and validity mask. The transformer alternates causal time-axis self-attention with variate-axis group self-attention. Time attention uses time-aware rotary embeddings. Variate attention lets variates in the same sample attend to one another. The stack uses pre-norm RMSNorm blocks, SwiGLU feed-forward layers, and a quantile head.
At inference, target and historical variates are normalized with causal running statistics. Future covariates use per-row global statistics.
| Field | Value |
|---|---|
| Parameters | approximately 102M |
| Layers | 24 |
| Layer pattern | 2 time-attention layers, then 1 group-attention layer |
| Time attention layers | 16 |
| Group attention layers | 8 |
| Embedding dim | 512 |
| Feedforward dim | 2048 |
| Attention heads | 8 |
| Patch size | 32 |
| Dropout | 0.1 |
| Scaler | causal mean/std with arcsinh transform |
| Native quantile levels | 0.1, 0.25, 0.5, 0.75, 0.9 |
t0-alpha is reported on the GIFT-Eval leaderboard and the fev-bench leaderboard.
| Benchmark | Metric | Value |
|---|---|---|
| GIFT-Eval | CRPS | 0.4941 |
| GIFT-Eval | MASE | 0.7240 |
| fev-bench | Skill score | 42.2 |
Users should also evaluate t0-alpha on their own historical backtests. Useful checks include quantile loss, CRPS, MASE, empirical quantile coverage, calibration, and breakdowns by frequency, horizon, domain, history length, and covariate availability.
T0Forecaster: the model itself.Forecast: the object returned by the model.T0Config: the configuration of the model.MaskType: the reason a time step is masked out.VariateType: whether a row is a target, a historical covariate or a
known-future covariate.batch_series: utility to batch time series of potentially different lengths.TimeSeries.from_array / TimeSeries.batch: build the model's native input,
including known-future covariates and an explicit forecast horizon.
predict accepts either a TimeSeries or a raw context array.t0 builds on ideas from open-source forecasting models. We gratefully acknowledge:
Code-level attributions are listed in NOTICE, all under Apache-2.0.
Training compute and carbon emissions are not currently reported.
t0 is described in t0: A Time-Series Foundation Model for Forecasting with Context. If our model is useful, please cite:
@article{meyer2026t0,
title = {$t_0$: A Time-Series Foundation Model for Forecasting with Context},
author = {Meyer, Lucas and Sole, Claudio and Xiang, Huikan and Li, Nicolas and Franceschino, Lucas and Quera-Bofarull, Arnau and Scholl, Maarten P. and Fainberg, Joachim and N{\'e}giar, Geoffrey},
journal = {arXiv preprint arXiv:2609.24559},
year = {2026},
url = {https://arxiv.org/abs/2609.24559},
}
Apache-2.0. See LICENSE and NOTICE.
For issues and bug reports, use the tracker for the relevant runtime: