Downloads · 30 days
0
grmchn/mascot-pose-detect
mascot-pose-detect is a object detection model from grmchn. Use it when you need objects located in an image. It is set up for onnx. The card lists the license as apache-2.0.
Two-stage mascot pose detector for chibi, kemono, and other stylized mascot characters.
Downloads · 30 days
0
Access
Public
Updated May 29, 2026
Repo size
2.5 GB
Likes
1
Public
Click a slice to open those files.
.onnx2.5 GB · 100%
From the Hugging Face model README
Two-stage mascot pose detector for chibi, kemono, and other stylized mascot characters.
This repository provides ONNX artifacts for portable inference:
The keypoint model is fine-tuned for stylized mascot bodies whose proportions differ strongly from real human pose datasets. The consumer should lift the COCO-17 keypoints to DWPose-25 / POSE_KEYPOINT format and derive toe points from the foot bounding boxes. Hand keypoints are expected to be generated by a separate hand-template fitter when required.
This model package is released under the Apache License 2.0.
The Stage 2 keypoint model is based on facebook/dinov2-large, which is also released under Apache 2.0.
It does not use the older MAE-pretrained ViTPose-L checkpoint that constrained the previous test bundle to non-commercial use.
The training annotations and source images are not included in this repository.
grmchn/mascot-pose-detect/
├── bbox/
│ ├── model.onnx
│ ├── classes.json
│ └── decode_params.json
└── keypoint/
├── dinov2_vitpose_l/
│ ├── model.onnx
│ └── meta.json
└── dinov2_vitpose_l_v2/
├── model.onnx
└── meta.json
The bbox model detects seven mascot body regions:
| index | name |
|---|---|
| 0 | full |
| 1 | head |
| 2 | body |
| 3 | hand_left |
| 4 | hand_right |
| 5 | foot_left |
| 6 | foot_right |
Left and right follow anatomical / character-view naming. For a front-facing character, the character's right side usually appears on the screen-left side.
The keypoint model input is a top-down crop from the Stage 1 full or body bbox.
| version | keypoint path | training run | status |
|---|---|---|---|
| v1 | keypoint/dinov2_vitpose_l/ | general_filtered | Stable baseline release |
| v2 | keypoint/dinov2_vitpose_l_v2/ | final_v3_from_final_v2 | Updated model with additional hard-example training |
Both versions use the same architecture, input shape, output heatmap shape, and post-processing contract.
Switching from v1 to v2 only requires changing the keypoint variant path from dinov2_vitpose_l to dinov2_vitpose_l_v2.
| field | value |
|---|---|
| Architecture | dinov2_vitpose_l |
| Backbone | facebook/dinov2-large |
| Input | 1x3x224x168 NCHW RGB, ImageNet-normalized |
| Output | heatmap |
| Keypoint layout | COCO-17 |
| Post-process layout | DWPose-25 / POSE_KEYPOINT-compatible |
See each version's meta.json for exact input size, normalization values, output names, and post-processing notes.
from huggingface_hub import snapshot_download
local_dir = snapshot_download(
repo_id="grmchn/mascot-pose-detect",
allow_patterns=[
"bbox/*",
"keypoint/dinov2_vitpose_l_v2/*",
],
)
This is not an OpenPose implementation and does not include OpenPose weights. It produces keypoints that can be converted into an OpenPose-compatible JSON schema for downstream tools.
The model was trained for stylized mascot characters. It may not generalize to realistic human photos without additional fine-tuning.