Downloads · 30 days
6
15% of all-time downloads
x2q/kartnet
kartnet is a video classification model from x2q. Use it for the video classification task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
KartNet turns onboard video from one specific indoor kart circuit (Racehall Aarhus, Denmark) into telemetry that no sensor can record there — GPS does not work indoors:
Downloads · 30 days
6
15% of all-time downloads
All-time downloads
39
Public
Repo size
23.2 MB
Likes
0
Public
Click a slice to open those files.
.mp416.5 MB · 71%
From the Hugging Face model README
KartNet turns onboard video from one specific indoor kart circuit (Racehall Aarhus, Denmark) into telemetry that no sensor can record there — GPS does not work indoors:
If the recording carries an IMU (DJI and GoPro action cameras embed one), an optional input branch uses it and measurably improves accuracy. Without sensors the model runs on video alone — one checkpoint serves both cases.
It is the specialized sibling of SpeedNet, our general ego-speed model. Where SpeedNet must work on any road, KartNet exploits the fact that every clip shows the same 790 m circuit — which buys per-frame position and much tighter speed.
Held-out race (never trained on), same venue, video + IMU. The speed readout and the lap-position bar are model outputs; there is no GPS in the loop anywhere:
<video src="https://huggingface.co/x2q/kartnet/resolve/main/assets/demo.mp4" controls muted playsinline width="100%"></video>
There is no GPS indoors and rental karts expose no data bus, so the training labels themselves had to be manufactured from video:
u around the lap.u. Validated
against lap times published in video titles (median error 0.22 s)
and against the venue's official SMS-Timing for one race
(27/27 laps, median error 0.19 s) — data the aligner never saw.790 m x du/dt.Held-out whole videos (unseen sessions, drivers, karts, lighting):
| protocol | speed MAE | position (median) |
|---|---|---|
| 64-frame windows | 2.2 km/h | 1.3 m |
| whole clip, streaming | 2.5 km/h | 7.1 m |
Length-invariance (the model was trained with randomized sequence lengths, so window size is a throughput knob, not an accuracy mode):
| window | 16 | 32 | 64 | 128 |
|---|---|---|---|---|
| speed MAE (km/h) | 1.95 | 1.70 | 1.54 | 1.49 |
| position (m) | 1.9 | 1.4 | 1.0 | 0.8 |
Does it measure, or does it recall? A model that can localize itself on a known track could fake speed by recalling what karts usually do at that corner. Ablation says no:
| inputs | speed MAE |
|---|---|
| flow + appearance (full) | 2.6 km/h |
| flow only | 3.4 km/h |
| appearance only (no motion!) | 8.0 km/h |
Appearance refines; optical flow carries the measurement.
Cross-camera holdout — a different camera, mount and month (a DJI helmet camera never seen in training except via two other sessions from the same device), judged against official race timing:
| metric | video only | video + IMU |
|---|---|---|
| per-lap mean speed MAE (all 27 laps) | 1.9 km/h | 1.1 km/h |
| per-lap mean speed bias | +1.9 km/h | +0.8 km/h |
| position (median) | 3.5 m | 3.4 m |
Fetch the repo via huggingface_hub rather than downloading the weight
file directly — this is also what makes your download count toward the
model's stats on the Hub:
from huggingface_hub import snapshot_download
local_dir = snapshot_download("x2q/kartnet")
import torch
from modeling_kartnet import KartNet, extract_features, predict
model = KartNet()
model.load_state_dict(torch.load("kartnet_v3.pt", map_location="cuda"))
feats = extract_features("onboard.mp4") # ffmpeg + OpenCV, ~2x realtime
out = predict(model, feats, device="cuda")
out["speed_mps"] # per-frame speed, 10 Hz
out["u"] # lap phase in [0,1)
out["lap_times_s"] # from phase wrap points
With IMU (example for DJI embedded telemetry, any source works if timestamps are on the video clock):
from modeling_kartnet import imu_features
imu = imu_features(len(feats["gray"]), t_imu, accel_xyz_g, yaw_rate_dps,
gforce_lon, gforce_lat)
out = predict(model, feats, imu=imu, device="cuda")
Or just: python example.py my_clip.mp4
Apache-2.0 (weights and code).
@misc{kartnet2026,
title = {KartNet v3: track-specialized telemetry from onboard karting video},
author = {x2q},
year = {2026},
url = {https://huggingface.co/x2q/kartnet}
}