Downloads · 30 days
0
OpenExplorer/deform_detr_resnet50
deform_detr_resnet50 is a machine learning model from OpenExplorer. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
Deformable DETR replaces DETR's global self-attention with multi-scale deformable attention: each query only samples a small number of points near reference points, enabling faster convergence and lower compute. Four…
Downloads · 30 days
0
Access
Public
Updated Sep 1, 2026
Repo size
2.4 GB
Likes
0
Public
Click a slice to open those files.
.tar1.1 GB · 44%
From the Hugging Face model README
Deformable DETR replaces DETR's global self-attention with multi-scale deformable attention: each query only samples a small number of points near reference points, enabling faster convergence and lower compute. Four feature levels plus 900 queries provide multi-scale candidates, then the decoder progressively refines boxes layer by layer.
| Model | Model Input | Backbone | Neck | Model Output |
|---|---|---|---|---|
| DeformableDETR | Single image 1x3x800x1332 | ResNet-50 | ChannelMapperNeck | Detection boxes (B,N,cls+reg) |
| March | Metric | float | calibration | qat | hbm |
|---|---|---|---|---|---|
| J6M | mAP | 0.4384 | 0.412 | 0.4526 | 0.4529 |
Results are based on
march = March.NASH_M(J6M) configuration.HEAL version: heal 0.0.2 / hbdk4-compiler 4.11.11 / horizon_plugin_pytorch 3.3.10.
Performance measurement: FPS is measured with single-core eight-thread; Latency is measured with single-core single-thread; Memory is peak DDR usage.
| March | latency (ms) | fps | Memory Usage (MB) |
|---|---|---|---|
| J6M | 144.84 | 6.92 | 656.00 |
| J6P | 78.02 | 28.89 | 672.80 |
| J6B | - | - | - |
J6B performance is not available for this model.
Deformable DETR replaces DETR's global self-attention with multi-scale deformable attention: each query only samples a small number of points near reference points, enabling faster convergence and lower compute. Four feature levels plus 900 queries provide multi-scale candidates, then the decoder progressively refines boxes layer by layer.
ResNet50, include_top=False removes classification head).ChannelMapperNeck (in_channels=[512,1024,2048], out_channel=256, 1×1 conv, extra_convs=1).PositionEmbeddingSine (num_pos_feats=128, normalized).DeformableDetrTransformer (encoder 6 layers + decoder 6 layers, embed_dim=256, num_heads=8, feedforward_dim=1024, num_feature_levels=4, num_queries=900).DeformDetrPostProcess (evaluation selects select_box_nums_for_evaluation=300 boxes).DeformableCriterion: classification focal loss + L1 bbox + GIoU, HungarianMatcher bipartite matching, aux_loss=True.with_box_refine=False, as_two_stage=False.800 × 1332.Official repo: https://github.com/fundamentalvision/Deformable-DETR Paper: https://arxiv.org/abs/2010.04159