Downloads · 30 days
0
c3aiia3c/CHARM
CHARM is a time series forecasting model from c3aiia3c. Use it for the time series forecasting task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as other.
CHARM is a zero-shot probabilistic time-series foundation model from C3 AI. It produces full quantile forecasts (99 quantile levels) for arbitrary horizons and is evaluated zero-shot on the GIFT-Eval benchmark — no GI…
Downloads · 30 days
0
Access
Public
Updated Sep 18, 2026
Repo size
—
Likes
1
Public
Click a slice to open those files.
.csv7.8 MB · 98%
From the Hugging Face model README
CHARM is a zero-shot probabilistic time-series foundation model from C3 AI. It produces full quantile forecasts (99 quantile levels) for arbitrary horizons and is evaluated zero-shot on the GIFT-Eval benchmark — no GIFT-Eval data is used in training.
Availability: CHARM model weights and training/replication code are not publicly released. The model is served behind an API and consumed through the open-source
c3-charmPython SDK (see Using CHARM below). This page serves as the model card and benchmark record.
| Parameters | ~63.3M (encoder ~59M) |
| Hidden size (d_model) | 384 |
| Projection | TCN-pool (patch size 16, causal conv stack + residual gate) |
| Backbone | Transformer encoder with RoPE attention |
| Decoder | Quantile decoder, 99 levels (0.01–0.99) |
| Max context length | 8192 |
| Output | Probabilistic (quantile) forecasts, multivariate-capable |
| Precision | float32 |
Zero-shot probabilistic forecasting of univariate and multivariate time series across domains (energy, transport, sales, healthcare, nature, web/cloud-ops, econ/finance). The model is applied without any per-dataset fine-tuning.
c3-charm SDKCHARM is served behind an API and consumed through the open-source c3-charm Python SDK. The SDK provides embeddings (multivariate time series → vectors), forecast/backcast (quantile predictions), and an optional toolkit for downstream tasks (anomaly detection, retrieval, classification, reconstruction, forecasting). The reference below is the full SDK documentation.
Dual-model serving. A CHARM server can be backed by two independent checkpoints — one for embeddings (
/predict,client.embeddings) and a separate one for forecasting (/forecast,client.prediction). They may differ in architecture, patch size, and embedding dimension, so read per-model properties fromclient.model_info()rather than hardcoding.
Runnable notebooks live in notebooks/ in this repo. Open them on the Hub, or download the folder and run locally after pip install c3-charm[toolkit] (set CHARM_BASE_URL / CHARM_API_KEY first). The classification and reconstruction/forecasting demos read the small sample datasets in notebooks/data/.
| Notebook | What it covers |
|---|---|
getting_started.ipynb | Client setup, first embeddings call, inspecting model_info(). |
charm_toolkit_demo.ipynb | Tour of the charm_toolkit — datasets, precompute, trainer, and each task head. |
demo_forecasting.ipynb | Zero-shot quantile forecasting and the embedding-based ForecastingModel head (weather data). |
demo_reconstruction.ipynb | Backcast reconstruction and the ReconstructionModel head for anomaly detection (weather data). |
demo_classification.ipynb | Time-series classification with the ClassificationModel head (BasicMotions data). |
demo_retrieval_anomaly_detection.ipynb | Embedding retrieval + kNN / zero-shot anomaly-detection recipes. |
pip install c3-charm # core SDK only (embeddings + forecast)
pip install c3-charm[toolkit] # includes PyTorch models, datasets, trainers
Or from source:
git clone https://github.com/c3ai/c3-charm.git
cd c3-charm
poetry install # core SDK only
poetry install --with toolkit # include toolkit dependencies
pip install c3-charm pulls the default torch wheel from PyPI, which on
Linux is the full CUDA build (~4 GB). The SDK itself never requires a
GPU — the model runs server-side, and toolkit training heads are small
enough to fit on CPU — so if you want the smaller CPU-only wheel, install
torch from the CPU index before c3-charm:
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install c3-charm
Or from source, use bash setup.sh, which forces the CPU wheel.
If you instead want a specific CUDA build, install it manually after
c3-charm:
pip install --force-reinstall --no-deps torch \
--index-url https://download.pytorch.org/whl/cu124
Swap cu124 for the CUDA version that matches your driver (cu121,
cu126, cu128, …). --no-deps is important — it stops pip from
re-resolving your torch install.
from charm import CharmClient
client = CharmClient(
base_url="http://your-server:8080",
api_key="your-api-key", # or set CHARM_API_KEY env var
timeout=300, # override; SDK default is 15s — raise it for /forecast (server allows up to ~220s)
max_retries=3,
)
client.embeddings.create()Converts time series windows into dense vectors.
response = client.embeddings.create(
descriptions=[["sensor_A", "sensor_B"]], # (N, C) channel names
ts_array=[[[1.0, 2.0], [1.1, 2.1], ...]], # (N, T, C) values
batch_size=32,
return_tensors="np", # "list", "np", or "torch"
aggregate=True, # True → (N, D); False → (N, T_, C, D)
progress=True,
)
embeddings = response.embeds # shape (N, D) when aggregate=True
aggregate parameter:
True (default): Returns flattened embeddings (N, D) — one vector per series. Best for retrieval, classification, clustering.False: Returns per-patch, per-channel embeddings (N, T_, C, D) where T_ = ceil(T / patch_size). Best for fine-grained tasks or custom heads.Async (faster for large datasets):
response = await client.embeddings.async_create(
descriptions=descriptions,
ts_array=ts_array,
max_B_per_request=32,
concurrency_per_call=8,
return_tensors="np",
aggregate=True,
)
client.prediction.create()Zero-shot quantile predictions — no training required.
response = client.prediction.create(
descriptions=[["sensor_A", "sensor_B"]],
ts_array=[[[1.0, 2.0], [1.1, 2.1], ...]],
target_len=10, # positive = forecast, negative = backcast
return_tensors="np",
)
forecast = response.denormalized_predictions # (N, target_len=10, C, Q) — Q quantiles (model-dependent, e.g. 21)
median = response.median # (N, 10, C) — point forecast (median quantile)
Backcast (reconstruct past values):
response = client.prediction.create(
descriptions=descriptions,
ts_array=ts_array,
target_len=-8, # reconstruct last 8 steps
return_tensors="np",
)
| Constraint | Limit |
|---|---|
| Timesteps per series | 1 ≤ T ≤ 8192 (the model's training window) |
| Channels per series | No hard limit (large C grows memory ~O(C²) under cross-channel attention) |
| Per-request size | N × C × T ≤ 500,000 — SDK client-side batching guard, not a server limit |
| Batch consistency | All series in a request must share the same T and C |
| Best accuracy | T ≤ 8192 (the model's training window); T need not be a multiple of the patch size |
The model was trained on windows of up to 8192 timesteps (patch size 16).
Inputs are not required to be a multiple of the patch size — they are padded
to a patch boundary internally. The architecture can technically accept more
than 8192 (up to 1500 patches), but longer inputs rely on positions seen only
in pretraining, so quality degrades. Query client.model_info() for the
served model's actual patch size and embedding dimension.
The SDK handles client-side batching automatically when you set batch_size (sync) or max_B_per_request (async).
| Method | Output field | Shape |
|---|---|---|
embeddings.create(aggregate=True) | response.embeds | (N, D) |
embeddings.create(aggregate=False) | response.embeds | (N, T_, C, D), T_ = ceil(T / patch_size) |
prediction.create(target_len > 0) | response.denormalized_predictions | (N, target_len, C, Q) |
prediction.create(target_len < 0) | response.denormalized_predictions | (N, abs(target_len), C, Q) |
prediction.create(...) | response.median | (N, abs(target_len), C) |
Descriptions are required and affect embedding quality. They tell the model what each channel represents.
Good descriptions — use meaningful, consistent names:
descriptions = [["engine_temperature", "oil_pressure", "rpm"]]
Acceptable — short but informative:
descriptions = [["temp", "pressure", "speed"]]
Avoid — generic or positional names reduce model effectiveness:
descriptions = [["col_0", "col_1", "col_2"]] # works but suboptimal
When working with pandas DataFrames, use column names directly:
descriptions = [df.columns.tolist()] * N
No pre-processing needed. CHARM normalizes internally. Send raw data as-is. Do not apply StandardScaler, MinMaxScaler, or log transforms before calling the API.
from charm import CharmError, AuthenticationError, InvalidRequestError, RateLimitError
try:
response = client.embeddings.create(...)
except AuthenticationError:
# bad API key
except InvalidRequestError as e:
# shape violations, empty input
except RateLimitError:
# back off and retry
except CharmError as e:
# catch-all for other SDK errors
The toolkit (pip install c3-charm[toolkit]) provides PyTorch models, dataset utilities, and training infrastructure for fine-tuning on top of CHARM embeddings.
charm_toolkit.retrievalFind similar time series by embedding similarity.
from charm_toolkit.retrieval import (
l2_normalize,
cosine_similarity_matrix,
knn_search,
retrieval_metrics,
)
# Embed your data
response = client.embeddings.create(
descriptions=descriptions,
ts_array=windows_list,
return_tensors="np",
)
embeddings = response.embeds # (N, D)
# Similarity search
sim = cosine_similarity_matrix(embeddings, embeddings)
# kNN search
indices, scores = knn_search(query_emb, corpus_emb, k=5)
# Evaluation metrics
metrics = retrieval_metrics(
query_emb=query_emb,
corpus_emb=corpus_emb,
query_labels=query_labels,
corpus_labels=corpus_labels,
k_values=[1, 3, 5, 10],
exclude_self=True,
query_ids=query_dataset_names,
corpus_ids=corpus_dataset_names,
)
# Returns: precision@k, ndcg@k, hit_rate@k
charm_toolkit.anomaly_detectionDetect anomalies via kNN distance scoring on windowed CHARM embeddings.
from charm_toolkit.anomaly_detection import (
sliding_window_embeddings,
knn_anomaly_scores,
window_scores_to_pointwise,
)
# 1. Embed sliding windows
train_emb = sliding_window_embeddings(
client, train_data, descriptions,
window_size=128, stride=1, batch_size=64,
)
test_emb = sliding_window_embeddings(
client, test_data, descriptions,
window_size=128, stride=1, batch_size=64,
)
# 2. Score test windows by distance to train
window_scores = knn_anomaly_scores(
test_emb=test_emb,
reference_emb=train_emb,
k=5,
distance="cosine", # "cosine", "l2", "l1"
aggregation="mean", # "mean", "max"
)
# 3. Aggregate to per-timestep scores
pointwise_scores = window_scores_to_pointwise(
window_scores=window_scores,
window_size=128,
stride=1,
total_length=len(test_data),
method="mean", # "mean", "max", "last", "center"
)
Pointwise aggregation methods:
Each timestep is covered by multiple overlapping windows. The method parameter controls how to assign a single score per timestep:
| Method | Behavior | Use case |
|---|---|---|
"mean" | Average of all windows covering the point | Smooth, best for offline evaluation |
"max" | Max score among covering windows | Conservative, catches isolated spikes |
"last" | Score of the most recently completed window | Online/streaming — score only updates when a window finishes processing |
"center" | Score of the window centered on each point | Minimal time-shift, tightest temporal alignment |
Zero-shot recipes (no clean reference set required) — recommended methods, in order of strength:
sklearn.ensemble.IsolationForest on the embedding matrix and take the bottom ~70% by score as a presumed-clean reference; then call knn_anomaly_scores against that reference.pyod.models.cblof.CBLOF on the embedding matrix.sklearn.ensemble.IsolationForest directly on the embedding matrix.L2-normalize embeddings beforehand to use cosine geometry. See demo_retrieval_anomaly_detection.ipynb.
Best practices (validated on TSB-AD, metric VUS-PR):
Ensemble the embedding score with per-window statistics. The encoder
instance-normalizes each window, erasing amplitude/level-shift anomalies (spikes,
steps — the most common kind) from the embedding. Recover them with
window_statistics (no extra model call) and combine via ensemble_scores
(method="zscore" — z-score each detector and sum; parameter-free and as strong as a
tuned weighted ensemble). This is worth ~+7 pp VUS-PR supervised.
from charm_toolkit.anomaly_detection import window_statistics, ensemble_scores
emb_score = knn_anomaly_scores(test_emb, train_emb, k=3, distance="cosine")
S_tr = window_statistics(train_data, 128, 1)
mu, sd = S_tr.mean(0, keepdims=True), S_tr.std(0, keepdims=True) + 1e-8
stats_score = knn_anomaly_scores(
(window_statistics(test_data, 128, 1) - mu) / sd, (S_tr - mu) / sd,
k=3, distance="l2")
window_scores = ensemble_scores([emb_score, stats_score], method="zscore")
# zero-shot: build the reference by bootstrap (above) and use method="rank"
Multivariate: do NOT pool channels when the channel count is high. Channel-mean
pooling dilutes an anomaly confined to a few channels across all of them. Keep channels
separate with sliding_window_channel_embeddings + per_channel_knn_scores
(channel_pool="adaptive"). The benefit grows with the number of channels — negligible
at C≤3, sizeable at C≥20, large at C>60 — so it is advisable whenever C is high. For
few channels, plain mean-pooling is equally good and cheaper.
from charm_toolkit.anomaly_detection import (
sliding_window_channel_embeddings, per_channel_knn_scores)
train_pc = sliding_window_channel_embeddings(client, train_data, descriptions) # (N, C, D)
test_pc = sliding_window_channel_embeddings(client, test_data, descriptions)
emb_score = per_channel_knn_scores(test_pc, train_pc, k=3, channel_pool="adaptive")
from charm_toolkit import (
ReconstructionModel, create_reconstruction_datasets,
collator, TrainerClass,
)
from torch.utils.data import DataLoader
import torch.nn as nn
train_ds, val_ds, test_ds = create_reconstruction_datasets(
raw_data, # (T, C) numpy array or torch tensor
descriptions=channel_names,
window_size=256,
stride=1,
train_ratio=0.7,
val_ratio=0.15,
sequential=True,
scale=True,
)
model = ReconstructionModel(
embedding_client=client,
reconstructor="linear", # "linear", "mlp", or custom nn.Module
hidden_dim=128,
dropout=0.1,
)
trainer = TrainerClass(
model=model,
train_loader=DataLoader(train_ds, batch_size=512, collate_fn=collator),
val_loader=DataLoader(val_ds, batch_size=512, collate_fn=collator),
epochs=1000,
patience=5,
lr=1e-3,
criterion=nn.HuberLoss(),
)
trainer.fit()
from charm_toolkit import ForecastingModel, create_forecasting_datasets, collator, TrainerClass
from torch.utils.data import DataLoader
train_ds, val_ds, test_ds = create_forecasting_datasets(
raw_data,
descriptions=channel_names,
train_horizon=96,
test_horizon=96,
train_ratio=0.7,
val_ratio=0.15,
sequential=True,
scale=True,
)
model = ForecastingModel(
embedding_client=client,
horizon=96,
input_size=96,
head="linear",
hidden_dim=128,
mode="last", # "last", "avg", "none"
per_channel=True,
num_channels=len(channel_names),
)
trainer = TrainerClass(
model=model,
train_loader=DataLoader(train_ds, batch_size=512, collate_fn=collator),
val_loader=DataLoader(val_ds, batch_size=512, collate_fn=collator),
epochs=1000,
patience=10,
lr=1e-2,
)
trainer.fit()
from charm_toolkit import ClassificationModel, create_classification_datasets, collator, TrainerClass
from torch.utils.data import DataLoader
import torch.nn as nn
train_ds, val_ds, test_ds = create_classification_datasets(
raw_data, # (N, T, C)
labels=labels, # list of N integer labels
descriptions=channel_names,
train_ratio=0.7,
val_ratio=0.15,
)
model = ClassificationModel(
embedding_client=client,
num_classes=num_classes,
hidden_dim=128,
pooling_over_t="mean",
pooling_over_channels="mean",
classifier_type="mlp",
)
trainer = TrainerClass(
model=model,
train_loader=DataLoader(train_ds, batch_size=32, collate_fn=collator),
val_loader=DataLoader(val_ds, batch_size=32, collate_fn=collator),
epochs=100,
patience=10,
lr=1e-3,
criterion=nn.CrossEntropyLoss(),
)
trainer.fit()
Toolkit models call the API every forward pass. For training with hundreds of windows per epoch, precompute embeddings once:
from charm_toolkit import precompute_dataset_embeddings, PrecomputedEmbeddingsDataset
# Compute once, save to disk as memmap
train_shape = precompute_dataset_embeddings(
client=client, dataset=train_ds,
output_path="./outputs/train_embeddings.pt", memory_batch_size=8192
)
val_shape = precompute_dataset_embeddings(
client=client, dataset=val_ds,
output_path="./outputs/val_embeddings.pt", memory_batch_size=8192
)
# Wrap datasets — model skips API calls when "embeds" key present
train_ds = PrecomputedEmbeddingsDataset(train_ds, "./outputs/train_embeddings.pt", train_shape)
val_ds = PrecomputedEmbeddingsDataset(val_ds, "./outputs/val_embeddings.pt", val_shape)
# Training now uses cached embeddings — orders of magnitude faster
train_loader = DataLoader(train_ds, batch_size=512, shuffle=True, collate_fn=collator)
from charm_toolkit import TrainerClass
trainer = TrainerClass(
model=model,
train_loader=train_loader,
val_loader=val_loader,
test_loader=test_loader, # optional
lr=1e-3,
weight_decay=1e-4,
epochs=1000,
patience=5,
min_delta=1e-4,
max_grad_norm=5.0,
criterion=None, # defaults to MSELoss
)
trainer.fit()
test_loss = trainer.evaluate(test_loader)
All return (train_dataset, val_dataset, test_dataset):
| Function | Input shape | Key args |
|---|---|---|
create_reconstruction_datasets(raw_data, ...) | (T, C) | window_size, stride, train_ratio, val_ratio |
create_forecasting_datasets(raw_data, ...) | (T, C) | train_horizon, test_horizon, stride, train_ratio, val_ratio |
create_classification_datasets(raw_data, labels, ...) | (N, T, C) | train_ratio, val_ratio |
Reconstruction and forecasting expect a single long time series (T, C) split temporally. Classification expects pre-windowed (N, T, C).
All DataLoaders using toolkit datasets require collator as the collate_fn:
from charm_toolkit import collator
# or equivalently:
from charm_toolkit.Datasets import collator
CHARM embeddings work as drop-in feature vectors for any sklearn model:
import numpy as np
from sklearn.ensemble import IsolationForest
from sklearn.linear_model import LogisticRegression
from charm_toolkit.retrieval import cosine_similarity_matrix
response = client.embeddings.create(
descriptions=descriptions,
ts_array=windows_list,
return_tensors="np",
)
X = response.embeds # (N, D)
# Anomaly detection with isolation forest
clf = IsolationForest(contamination=0.05)
anomaly_labels = clf.fit_predict(X)
# Similarity search
sim = cosine_similarity_matrix(X, X)
# As features for any classifier
clf = LogisticRegression().fit(X_train, y_train)
Deploy models locally from GitHub releases — no remote server needed:
with CharmClient(tag="experiment-2026-03-15_10-30-00") as client:
response = client.embeddings.create(...)
# Server shuts down automatically
When tag is provided:
Files cached at ~/.charm/models/<tag>/ for fast subsequent runs.
CharmClient(
tag="experiment-tag", # required for local mode
repo_url="https://...", # default: c3-e/research
cache_dir="/path/to/cache", # default: ~/.charm/models
port=8080, # 0 = auto-select
)
(N, T, C) — N series, each T timesteps × C channels; all series in one request must share the same T and C.T ≤ 8192 (the training window). The model can accept more, but quality degrades on lengths it wasn't trained on. T does not need to be a multiple of the patch size — inputs are padded to a patch boundary internally."engine_temperature", not "col_0") — the model is channel-aware and descriptions materially affect embeddings.batch_size × C × T ≤ 500,000 (SDK-enforced client-side).async_create for large N — it batches concurrently; the sync client is sequential and slow past ~100 series. Tune max_B_per_request / concurrency_per_call instead of one giant request.timeout for forecasting — the SDK default is 15s, but the server allows forecasts up to ~220s. Use timeout ≈ 220+ for prediction.create.aggregate=True (default) → (N, D) for retrieval / classification / clustering. aggregate=False → (N, T_, C, D) only when you need per-patch/per-channel detail for a custom head.D at runtime via client.model_info() — it's model-dependent; don't hardcode.target_len > 0 = forecast, < 0 = backcast; 0 is invalid.T + abs(target_len) ≤ 8192 (the training window) for best forecast/backcast quality.response.median for a point forecast, or the full quantile axis for intervals. Q is model-dependent (e.g. 21 or 99) — read denormalized_predictions.shape[-1].ClassificationModel, set pooling_over_channels="flatten" (and pass num_channels=C, required for flatten) instead of the default "mean"./predict (embeddings) and /forecast as separate models — they may have different patch sizes / embedding dims. Read the per-role models map from client.model_info() instead of assuming they match.precompute_dataset_embeddings + PrecomputedEmbeddingsDataset) — toolkit models otherwise call the API every forward pass.InvalidRequestError, AuthenticationError, RateLimitError, or the base CharmError) and back off on rate limits.with CharmClient(...) as client:) so local deployments shut down cleanly.CHARM_API_KEY, CHARM_BASE_URL) rather than hardcoding.| Approach | When | Effort |
|---|---|---|
prediction.create(target_len=H) | Quick forecast baseline, no labeled data | None — one API call |
| Embeddings + sklearn | Moderate data, combine with other features | Minutes |
| Embeddings + kNN (retrieval/AD) | Unlabeled anomaly detection or search | Minutes |
| Toolkit model (Reconstruction/Forecasting/Classification) | Have labeled data, want best performance | Train a small head (~minutes on CPU) |
Evaluated zero-shot on the full GIFT-Eval benchmark (97 dataset/frequency/term configurations) using the standard 11-metric protocol. Aggregate scores (geometric mean of per-config metrics normalized to the Seasonal Naive baseline; lower is better):
| Metric | Score (rel. Seasonal Naive) |
|---|---|
| MASE | 0.7582 |
| CRPS (mean weighted sum quantile loss) | 0.4776 |
Per-term (geometric mean, normalized to Seasonal Naive):
| Term | MASE | CRPS |
|---|---|---|
| short | 0.7463 | 0.5036 |
| medium | 0.7577 | 0.4452 |
| long | 0.7911 | 0.4460 |
Scores are the geometric mean of per-config metric / Seasonal Naive across all
97 GIFT-Eval configurations (lower is better; < 1.0 beats Seasonal Naive).
Full per-config results are in all_results.csv.
evaluate_forecasts).testdata_leakage = No).@misc{charm,
title = {CHARM: A Zero-Shot Time-Series Foundation Model},
author = {C3 AI},
year = {2026},
url = {https://huggingface.co/c3aiia3c/CHARM}
}