Downloads · 30 days
0
resoajoe/camera-motion-nano
camera-motion-nano is a video classification model from resoajoe. Use it for the video classification task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
47,238 parameters. 189 KB. What did the camera do — hold still, pan, zoom, or shake?
Downloads · 30 days
0
Access
Public
Updated Sep 3, 2026
Repo size
393 KB
Likes
0
Public
Click a slice to open those files.
.pt200 KB · 48%
From the Hugging Face model README
47,238 parameters. 189 KB. What did the camera do — hold still, pan, zoom, or shake?
It never looks at a pixel. It reads the motion vectors the video encoder already computed and wrote into the bitstream.
Domain measured / deployment domain tested: measured on H.264 clips synthesised from COCO photographs and webcam frames with known affine motion; deployment domain: no real handheld, PTZ or vehicle-mounted camera motion tested. (Fifth line of the card standard, added 2026-09-02: a number is only as good as the domain it was measured in.)
| input | accuracy |
|---|---|
| motion vectors only (this model) | 0.992 |
| pixels — two 64×64 frames, same architecture | 0.785 |
| best single scalar on the MV field, fitted in-sample | 0.467 |
| chance | 0.167 |
Reading the codec's own metadata is not merely cheaper than looking at pixels here — it is more accurate, by +0.207. That is not surprising once stated: the encoder performed dense block-matching at full resolution as part of its ordinary work, while the pixel model sees two downsampled thumbnails. The expensive computation was already done and thrown away.
Per-class recall: static 1.000 · pan-x 0.984 · pan-y 1.000 · zoom-in 0.983 · zoom-out 1.000 · shake 0.984.
For: camera-tamper detection (has this fixed camera been moved?), footage triage and indexing, stabiliser gating, and activity sensing on hardware too weak to run a vision model.
Not for:
A 16×16×2 motion-vector field, mean displacement per cell over the clip, resized to 64×64 and fed as 2 channels (x, y). Extraction with PyAV:
import av, numpy as np, cv2, onnxruntime as ort
GRID, SZ = 16, 320
CLASSES = ["static","pan_x","pan_y","zoom_in","zoom_out","shake"]
def mv_field(path):
acc = np.zeros((GRID,GRID,2), np.float32); cnt = np.zeros((GRID,GRID), np.float32)
c = av.open(path); st = c.streams.video[0]
st.codec_context.options = {"flags2": "+export_mvs"} # REQUIRED
for fr in c.decode(st):
for sd in fr.side_data:
if "Motion" not in type(sd).__name__: continue
a = sd.to_ndarray()
if not len(a): continue
gx = np.clip((a["dst_x"]*GRID//SZ).astype(int), 0, GRID-1)
gy = np.clip((a["dst_y"]*GRID//SZ).astype(int), 0, GRID-1)
sc = np.maximum(a["motion_scale"], 1)
np.add.at(acc,(gy,gx,0), a["motion_x"]/sc)
np.add.at(acc,(gy,gx,1), a["motion_y"]/sc)
np.add.at(cnt,(gy,gx), 1)
f = acc/np.maximum(cnt,1)[...,None]
return cv2.resize(f,(64,64),interpolation=cv2.INTER_NEAREST).transpose(2,0,1)
sess = ort.InferenceSession("camera_motion.onnx", providers=["CPUExecutionProvider"])
x = mv_field("clip.mp4")[None]
print(CLASSES[int(sess.run(None, {"input": x})[0][0].argmax())])
SZ must match your video's dimensions — the grid mapping divides by it. Vectors are scaled by
motion_scale; skipping that silently changes the units.
Everything above used clips built from COCO photographs. Repeated on 320 clips built from Logitech BRIO frames of a real office — a different sensor, a different image pipeline, and imagery that had already been through the camera's own MJPG compression:
| accuracy | static | pan_x | pan_y | zoom_in | zoom_out | shake | |
|---|---|---|---|---|---|---|---|
| COCO photographs | 0.990 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.94 |
| Logitech BRIO | 0.975 | 1.00 | 0.96 | 1.00 | 1.00 | 1.00 | 0.89 |
A gap of 0.015. That is a much smaller drop than sibling models in this family show on the same kind of test (the resolution model loses 0.106), and the reason is structural: this model never sees pixels. Motion vectors are the encoder's description of movement, so sensor noise, ISP sharpening and colour rendering are already abstracted away before the model looks at anything.
Be precise about what this does and does not establish. It shows the model transfers across sensors. It does not show transfer to real optical camera motion — the motion is still synthetic affine, as stated in the failure modes. Real handheld movement has rolling-shutter skew and non-rigid components that this has never seen.
Added after an external validation that failed, which turned out to be informative rather than bad news. Measured on synthetic pans at known per-frame velocity, 30 clips per row:
| px / frame | total drift over the clip | P(static) | modal class |
|---|---|---|---|
| 0.00 | 0.0 px | 0.999 | static |
| 0.10 | 1.3 px | 0.998 | static |
| 0.25 | 3.2 px | 0.996 | static |
| 0.50 | 6.5 px | 0.715 | static |
| 1.00 | 13.0 px | 0.308 | pan_x |
| 2.00 | 26.0 px | 0.115 | pan_x |
| 4.00 | 52.0 px | 0.006 | pan_x |
Detection threshold is roughly 0.5–1.0 px per frame. Below 0.25 px/frame it reports static with high confidence, and it is not wrong to — motion vectors encode inter-frame displacement, and a fraction of a pixel per frame is genuinely no motion at the scale a codec works at.
It therefore CANNOT detect slow drift. Tested against footage whose cumulative camera drift had been measured independently by an ECC-based stabiliser at 61 px mean, this model called it static — correctly, because 61 px accumulated over roughly 82,000 frames is 0.0007 px/frame, about a thousand times below the floor above. The stabiliser measures displacement from a fixed reference; this model measures velocity between neighbours. They answer different questions and the test that conflated them was mine, not the model's.
If you need slow-drift detection, register frames against a fixed reference. This model is for motion that is happening now.
The scalar rule was applied before publication: the best single-threshold classifier on the MV field scores 0.467 fitted in-sample. This model beats that optimistic baseline by +0.524 held out, which is why it exists rather than a threshold.
ONNX vs PyTorch, identical weights, both CPU, 256 inputs: max relative logit difference 2.1e-07, 100% argmax agreement.
Using H.264 motion vectors instead of pixels is not a new idea. Compressed-domain video analysis has a substantial literature, including:
What is different here. The general finding that motion vectors substitute for optical flow is theirs. Compressed-domain work also targets content — human actions, moving objects — using substantially larger models. This targets global camera motion at a scale those papers do not work at, and reports two things they do not:
Practically, it is a 189 KB ONNX for camera-tamper and footage-triage, in a family where prior work ships papers and research code.
Every margin quoted here is against a stated baseline, because a margin without one is not a measurement. The baseline is the best single-threshold classifier over ten cheap statistics, fitted optimistically:
mean · std · lapvar · hf (high-frequency energy ratio) · grad (Sobel magnitude) ·
entropy · centre_edge · radial_slope · row_fft_peak · col_fft_peak
The last four are spatially aware, added after an earlier six-statistic baseline — all global aggregates — was found to systematically overstate model value on spatially structured tasks. A baseline that cannot see where anything is loses to a CNN by default. On one test task that flaw inflated an apparent margin from +0.060 to +0.261.
Two questions are asked with it, and they disagree:
Where this card quotes a single scalar figure without qualification, it is the in-sample one.
Source imagery is COCO val2017 (public). Motion is synthesised. No personal data, no surveillance footage, and no recordings of identifiable people are involved.