Downloads · 30 days
0
tasmulaev/rtmpose-m-distill
rtmpose-m-distill is a keypoint detection model from tasmulaev. Use it for the keypoint detection task on the model card, and read the license before you ship it in a product. It is set up for mmpose. The card lists the license as apache-2.0.
RTMPose-m (21 hand keypoints, SimCC, 256×256) fine-tuned via self-distillation on degraded video — pseudo-labels produced by the model itself on clean frames, training inputs artificially degraded — to keep tracking h…
Downloads · 30 days
0
Access
Public
Updated Jul 3, 2026
Repo size
110 MB
Likes
1
Public
Click a slice to open those files.
.pth55.3 MB · 50%
From the Hugging Face model README
RTMPose-m (21 hand keypoints, SimCC, 256×256) fine-tuned via self-distillation on degraded video — pseudo-labels produced by the model itself on clean frames, training inputs artificially degraded — to keep tracking hands through low resolution and motion blur, the main failure modes of off-the-shelf hand pose models on real-world sign language footage.
<p align="center"> <img src="assets/output_pytorch.jpg" alt="21-keypoint hand skeleton correctly placed on a heavily motion-blurred hand" width="420"/> <br/> <em>Model output on a heavily motion-blurred frame: the skeleton stays on the fingers. PyTorch and ONNX Runtime outputs are byte-identical (<code>assets/output_pytorch.jpg</code> vs <code>assets/output_onnxruntime.jpg</code>).</em> </p>Compared to the base RTMPose-m Hand5 checkpoint, this model:
Same architecture, same input size, same 21-keypoint COCO hand skeleton as the original — a drop-in replacement for the rtmpose-m_simcc-hand5 checkpoint in any mmpose / rtmlib / mmdeploy pipeline.
| File | Description |
|---|---|
rtmpose-m_hand_distill-256x256-a996d9ec.pth | PyTorch weights (EMA, epoch 100), mmpose format, 55 MB |
rtmpose-m_hand_distill.py | mmpose/mmengine training and inference config |
degrade_video.py | Video degradation script used to build the "dirty" half of the training set (opencv + numpy only) |
onnx/rtmpose-m-distill-256x256.onnx | ONNX export (opset 11, dynamic batch, FP32), outputs simcc_x/simcc_y |
onnx/deploy.json, onnx/pipeline.json | mmdeploy SDK configs for the ONNX model |
assets/ | PyTorch vs ONNX Runtime output parity check (byte-identical) |
Self-distillation on degraded video — the model is its own teacher:
degrade_video.py, targeting the dominant real-world failure mode — low source resolution: the full frame is downscaled so its short side lands around 300 px (randomized per clip), then resized back to the original size (INTER_AREA down, bilinear up), so teacher coordinates taken from the clean frames stay valid. The degradation toolkit also includes optical-flow-based motion blur (Farneback flow, accumulated along the flow field), gamma/lighting shift, Gaussian noise and JPEG compression, organized into severity profiles 1–5. Degradation is applied to the full frame before hand cropping (so crops don't retain more detail than a real low-res source would have), and per-clip seeding (crc32(filename) + seed) makes it fully reproducible. The student therefore learns to predict sharp-frame keypoints from corrupted inputs.Training setup: AdamW (lr 4e-4, wd 0.05), batch 1024, cosine schedule, AMP, EMA (ExpMomentumEMA, momentum 2e-4), flip/rotate/scale augmentation, seed 21. Single NVIDIA RTX PRO 6000 Blackwell GPU, PyTorch 2.7.0 / CUDA 12.8 / MMEngine 0.10.7, ~10 h wall-clock. Full details in rtmpose-m_hand_distill.py.
Held-out validation against pseudo-labels (mixed clean + dirty, 33,347 crops): the released checkpoint is the EMA weights at epoch 100 — [email protected] (bbox-normalized) 0.9893, EPE 5.96 px. Best raw validation score during training was PCK 0.9896 / EPE 5.87 at epoch 42; the validation curve is flat from roughly epoch 20 onward.
Side-by-side comparison on a sign language test video (~5,100 frames), hand retention relative to detections at thr 0.1:
| Confidence threshold | 0.1 | 0.15 | 0.2 | 0.3 |
|---|---|---|---|---|
| Hand retention, base | 100% | 98.8% | 97.3% | 93.1% |
| Hand retention, this model | 100% | 99.8% | 99.4% | 98.0% |
| Frames where this model detects a hand and base does not | 431 (8%) | 875 (17%) | 1,338 (26%) | 2,672 (52%) |
Frame-to-frame keypoint jitter at thr 0.3 is ~39% lower than the base model. The frames recovered by this model are dominated by motion blur during fast signing, crossed/interlocked hands, and hands pressed against the torso; visual inspection confirms the recovered skeletons lie on the fingers rather than being spurious detections.
Recommended operating point: thr 0.2–0.3 (the base model effectively requires thr ≤ 0.15 to avoid dropping hands).
from mmpose.apis import init_model, inference_topdown
model = init_model(
'rtmpose-m_hand_distill.py',
'rtmpose-m_hand_distill-256x256-a996d9ec.pth',
device='cuda:0',
)
results = inference_topdown(model, 'hand_crop.jpg')
keypoints = results[0].pred_instances.keypoints # (1, 21, 2)
scores = results[0].pred_instances.keypoint_scores # (1, 21)
import cv2
import numpy as np
import onnxruntime as ort
sess = ort.InferenceSession('onnx/rtmpose-m-distill-256x256.onnx')
img = cv2.imread('hand_crop.jpg') # BGR hand crop
inp = cv2.resize(img, (256, 256))[:, :, ::-1].astype(np.float32) # to RGB
inp = (inp - [123.675, 116.28, 103.53]) / [58.395, 57.12, 57.375]
inp = inp.transpose(2, 0, 1)[None]
simcc_x, simcc_y = sess.run(None, {'input': inp.astype(np.float32)})
# SimCC decode: argmax over each axis, divide by split ratio (2.0)
x = simcc_x[0].argmax(axis=1) / 2.0 # (21,) in 256x256 crop coords
y = simcc_y[0].argmax(axis=1) / 2.0
conf = np.minimum(simcc_x[0].max(axis=1), simcc_y[0].max(axis=1))
The ONNX file is also compatible with rtmlib and the mmdeploy SDK (use onnx/ as the SDK model directory).
Pseudo-labels and training crops are derived from the Slovo Russian Sign Language dataset (SaluteDevices), distributed under a variant of CC BY-SA 4.0. The dataset itself is not included in this repository — only model weights.
@misc{jiang2023rtmpose,
title={RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose},
author={Jiang, Tao and Lu, Peng and Zhang, Li and Ma, Ningsheng and Han, Rui and Lyu, Chengqi and Li, Yining and Chen, Kai},
year={2023},
eprint={2303.07399},
archivePrefix={arXiv}
}
@inproceedings{kapitanov2023slovo,
title={Slovo: Russian Sign Language Dataset},
author={Kapitanov, Alexander and Kvanchiani, Karina and Nagaev, Alexander and Petrova, Elizaveta},
booktitle={International Conference on Computer Vision Systems},
pages={63--73},
year={2023},
organization={Springer}
}