Downloads · 30 days
0
open-noodle/pet-recognition-base
pet-recognition-base is a image feature extraction model from open-noodle. Use it for the image feature extraction task on the model card, and read the license before you ship it in a product. It is set up for onnx. The card lists the license as apache-2.0.
Individual pet re-identification embeddings (dogs and cats) — the "which pet is this" layer used by Gallery's pet recognition, on top of whole-animal crops from its pet detector.
Downloads · 30 days
0
Access
Public
Updated Jul 24, 2026
Repo size
348 MB
Likes
0
Public
Click a slice to open those files.
.onnx348 MB · 100%
From the Hugging Face model README
Individual pet re-identification embeddings (dogs and cats) — the "which pet is this" layer used by Gallery's pet recognition, on top of whole-animal crops from its pet detector.
A frozen facebook/dinov2-base backbone (86M
parameters) plus a trained linear projection to 512 dimensions. The projection's
L2-normalized output is the embedding; identity is compared with cosine similarity.
Fine-tuning the backbone was tried and rejected — it overfits the training identities and
forgets DINOv2's general features, while the frozen-backbone projection beats zeroshot on
both species.
| Input | input, float32 [N, 3, 224, 224], RGB, ImageNet mean/std normalized |
| Output | embedding, float32 [N, 512], L2-normalized |
| Batch | dynamic |
| Opset | 17 |
Crop the detected animal's bounding box, resize to 224x224, normalize with ImageNet
statistics (mean [0.485, 0.456, 0.406], std [0.229, 0.224, 0.225]). Compare embeddings
with cosine similarity (equivalently, dot product — the outputs are unit vectors).
Verification EER and identification Top-1 on held-out identities — individuals never seen in training — scored over the complete test splits:
| Test set | Images | Identities | EER | Top-1 | AUC |
|---|---|---|---|---|---|
| Dogs — Dogs-World (whole animal) | 53830 | 16469 | 0.047 | 0.612 | 0.988 |
| Cats — Cat Individual Images (whole animal) | 2575 | 102 | 0.045 | 0.916 | 0.991 |
| Dogs — DogFaceNet (unseen dataset, aligned faces) | 8363 | 1393 | 0.031 | 0.943 | 0.994 |
The backbone is Apache-2.0. The projection was trained only on openly-licensed data:
DogFaceNet (CC BY) is used for evaluation only. No restrictively-licensed pet re-ID dataset (PetFace, AvitoTech, MegaDescriptor) was used for training or distillation, so this model is safe for commercial use.
pet-recognition-small / pet-recognition-base / pet-recognition-large trade accuracy
against cost; base is Gallery's default.