Downloads · 30 days
0
mwmathis/DeepLabCutModelZoo-SuperAnimal-Quadruped
DeepLabCutModelZoo-SuperAnimal-Quadruped is a keypoint detection model from mwmathis. Use it for the keypoint detection task on the model card, and read the license before you ship it in a product.
• SuperAnimal-Quadruped model(s) developed by the M.W.Mathis Lab in 2023, trained to predict quadruped pose from images. Please see Shaokai Ye et al. 2023 for details.
Downloads · 30 days
0
Access
Public
Updated Jun 30, 2025
Repo size
2.6 GB
Likes
24
Trending 1
Click a slice to open those files.
.pt1.7 GB · 77%
From the Hugging Face model README
• SuperAnimal-Quadruped model(s) developed by the M.W.Mathis Lab in 2023, trained to predict quadruped pose from images. Please see Shaokai Ye et al. 2023 for details.
• The there are three main models (and several auxiliary model checkpoints):
pose_model.pth is an HRNet-w32 compatable for DLC3.0+ Pytorch code, trained on our Quadruped-80K dataset.detector.pt is a Faster R-CNN that can be used as a detector for top-down detection.hrnet_w32_quadruped80k.pth is an HRNet-w32 trained with mmpose on our Quadruped-80K dataset.• Full training details can be found in Ye et al. 2023. You can use the pose_model and detector simply with our light-weight loading package called DLCLibrary. Here is an example useage:
from pathlib import Path
from dlclibrary import download_huggingface_model
# Creates a folder and downloads the model to it
model_dir = Path("./superanimal_quadruped_model_pytorch")
model_dir.mkdir()
download_huggingface_model("superanimal_quadruped_pytorch", model_dir)
• Intended to be used for pose estimation of quadruped images taken from side-view. The model serves a better starting point than ImageNet weights in downstream datasets such as AP-10K.
• Intended for academic and research professionals working in fields related to animal behavior, such as neuroscience and ecology.
• Not suitable as a zeros-shot model for applications that require high keypiont precision, but can be fine-tuned with minimal data to reach human-level accuracy. Also not suitable for videos that look dramatically different from those we show in the paper.
• Based on the known robustness issues of neural networks, the relevant factors include the lighting, contrast and resolution of the video frames. The present of objects might also cause false detections and erroneous keypoints. When two or more animals are extremely close, it could cause the top-down detectors to only detect only one animal, if used without further fine-tuning or with a method such as BUCTD (Zhou et al. 2023 ICCV).
• Mean Average Precision (mAP)
• In the paper we benchmark on AP-10K, AnimalPose, Horse-10, and iRodent using a leave-one-out strategy. Here, we provide the model that has been trained on all datasets (see below), therefore it should be considered “fine-tuned" on all animal training data listed below. This model is meant for production and evaluation in downstream scientific applications.
It consists of being trained together on the following datasets:
Here is an image with the keypoint guide:
<p align="center"> <img src="https://images.squarespace-cdn.com/content/v1/57f6d51c9f74566f55ecf271/1690988780004-AG00N6OU1R21MZ0AU9RE/modelcard-SAQ.png?format=1500w" width="95%"> </p>Please note that each dataset was labeled by separate labs & separate individuals, therefore while we map names to a unified pose vocabulary (found here: https://github.com/AdaptiveMotorControlLab/modelzoo-figures), there will be annotator bias in keypoint placement (See the Supplementary Note on annotator bias). You will also note the dataset is highly diverse across species, but collectively has more representation of domesticated animals like dogs, cats, horses, and cattle. We recommend if performance is not as good as you need it to be, first try video adaptation (see Ye et al. 2023), or fine-tune these weights with your own labeling.
• No experimental data was collected for this model; all datasets used are cited.
• The model may have reduced accuracy in scenarios with extremely varied lighting conditions or atypical animal characteristics not well-represented in the training data.
• Please note that each dataest was labeled by separate labs & separate individuals, therefore while we map names to a unified pose vocabulary, there will be annotator bias in keypoint placement (See Ye et al. 2023 for our Supplementary Note on annotator bias). You will also note the dataset is highly diverse across species, but collectively has more representation of domesticated animals like dogs, cats, horses, and cattle. We recommend if performance is not as good as you need it to be, first try video adaptation (see Ye et al. 2023), or fine-tune these weights with your own labeling.
Modified MIT.
Copyright 2023 by Mackenzie Mathis, Shaokai Ye, and contributors.
Permission is hereby granted to you (hereafter "LICENSEE") a fully-paid, non-exclusive, and non-transferable license for academic, non-commercial purposes only (hereafter “LICENSE”) to use the "MODEL" weights (hereafter "MODEL"), subject to the following conditions:
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software:
This software may not be used to harm any animal deliberately.
LICENSEE acknowledges that the MODEL is a research tool. THE MODEL IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE MODEL OR THE USE OR OTHER DEALINGS IN THE MODEL.
If this license is not appropriate for your application, please contact Prof. Mackenzie W. Mathis ([email protected]) and/or the TTO office at EPFL ([email protected]) for a commercial use license.
Please cite Ye et al if you use this model in your work https://arxiv.org/abs/2203.07436v2.