Downloads · 30 days
0
ilessio-aiflowlab/project_azoth
project_azoth is a zero-shot object detection model from ilessio-aiflowlab. Use it for the zero-shot object detection task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Part of the ANIMA Perception Suite by Robot Flow Labs.
Downloads · 30 days
0
Access
Public
Updated Mar 30, 2026
Repo size
2.4 GB
Likes
0
Public
Click a slice to open those files.
.safetensors1.4 GB · 58%
From the Hugging Face model README
Part of the ANIMA Perception Suite by Robot Flow Labs.
AZOTH is a real-time open-vocabulary object detection module built on LLMDet (CVPR 2025). It detects arbitrary objects from natural language queries without predefined class vocabularies.
LLMDet: Learning to Understand Visual Grounding with LLMs arXiv: 2602.05730 | CVPR 2025
| Metric | This Model | Paper |
|---|---|---|
| AP (all) | 37.7 | 34.8 |
| AP_rare | 36.2 | 30.1 |
| AP_common | 28.8 | 35.2 |
| AP_frequent | 38.4 | 37.3 |
This baseline exceeds the paper's reported AP on LVIS minival.
| Format | File | Size | Use Case |
|---|---|---|---|
| SafeTensors | pytorch/azoth_v1_baseline.safetensors | ~660MB | Fast loading, HF transformers |
| PyTorch (.pth) | pytorch/azoth_v1_baseline.pth | ~660MB | Training, fine-tuning |
| ONNX | Deferred | — | Complex architecture; export on target |
| TensorRT | Deferred | — | Generate on target hardware (Jetson/L4) |
from transformers import AutoModelForZeroShotObjectDetection, AutoProcessor
from PIL import Image
model = AutoModelForZeroShotObjectDetection.from_pretrained(
"ilessio-aiflowlab/project_azoth",
subfolder="pytorch"
)
processor = AutoProcessor.from_pretrained(
"ilessio-aiflowlab/project_azoth",
subfolder="pytorch"
)
image = Image.open("test.jpg")
inputs = processor(images=image, text="a person. a car. a dog.", return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
results = processor.post_process_grounded_object_detection(
outputs, inputs.input_ids, threshold=0.3, target_sizes=[(image.height, image.width)]
)
Trained on GroundingCap-1M (CVPR 2025):
pytorch/ SafeTensors + PyTorch weights + tokenizer
checkpoints/ Best training checkpoint (resume training)
configs/ Training TOML configs
reports/ LVIS evaluation reports (JSON)
paper.pdf LLMDet CVPR 2025 paper
Apache 2.0 — Robot Flow Labs / AIFLOW LABS LIMITED