Downloads · 30 days
49
48% of all-time downloads
mlboydaisuke/midas-small-litert
midas-small-litert is a depth estimation model from mlboydaisuke. Use it for the depth estimation task on the model card, and read the license before you ship it in a product. It is set up for litert. The card lists the license as mit.
midassmall256fp16.tflite is MiDaS v2.1 small (MiDaSsmall, the CNN MiDaS with an EfficientNet-Lite3 backbone — not the DPT/ViT variants) converted to LiteRT for on-device monocular depth estimation. Given one RGB image…
Downloads · 30 days
49
48% of all-time downloads
All-time downloads
102
Public
Repo size
33.5 MB
Likes
0
Public
Click a slice to open those files.
.tflite33.5 MB · 100%
From the Hugging Face model README
midas_small_256_fp16.tflite is MiDaS v2.1 small (MiDaS_small, the CNN MiDaS
with an EfficientNet-Lite3 backbone — not the DPT/ViT variants) converted to
LiteRT for on-device monocular depth estimation. Given one RGB image it
predicts a per-pixel inverse-depth map (near = bright, far = dark).
It is the model used by the LiteRT compiled_model_api/depth_estimation Android
sample.
| File | Precision | Size |
|---|---|---|
midas_small_256_fp16.tflite | fp16 weights | ~33 MB |
| Task | Monocular depth estimation |
| Source | torch.hub.load("intel-isl/MiDaS", "MiDaS_small") |
| Input | 1 x 256 x 256 x 3 float32, RGB, ImageNet-normalized, NHWC (interleaved) |
| Output | 1 x 256 x 256 float32, relative inverse depth |
Pre-processing: resize to 256×256, normalize with ImageNet stats
(mean = [0.485, 0.456, 0.406], std = [0.229, 0.224, 0.225] on [0,1] pixels),
write as interleaved NHWC RGB float32.
Post-processing: min-max normalize the output and map through a color LUT
(the sample uses inferno).
The graph lowers entirely to GPU-clean builtins — no attention, no Flex/Custom
ops, no GATHER_ND, no >4D reshapes:
CONV_2D x73, ADD x27, DEPTHWISE_CONV_2D x24, RELU x7, RESIZE_BILINEAR x5, RESHAPE x1
to_channel_last_io) so the model takes NHWC 1x256x256x3
directly, matching the interleaved RGB the app writes (no input transpose).FLOAT_CASTING — half the size, runs natively on
the GPU delegate. Dynamic-range int8 is intentionally avoided (it favors the
CPU/XNNPACK path, not the GPU delegate).The fp16 model compiles to 234 / 234 nodes on the LiteRT GPU delegate
(LITERT_CL) — full GPU residency, no CPU fallback — at ~1–3 ms / inference
(best 1.1 ms). RESIZE_BILINEAR align_corners=True is GPU-supported as-is; no
model change needed.
Original work: Ranftl et al., "Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer" (MiDaS), https://github.com/isl-org/MiDaS.
A self-contained converter (litert-torch + ai-edge-quantizer) lives in the
sample under compiled_model_api/depth_estimation/conversion/:
pip install litert-torch ai-edge-quantizer torch timm matplotlib pillow
python convert_midas_litert.py out 256
<!-- funnel:v1 -->
Want a different model on-device? Open a request — free, open weights only; the export and its measured numbers get published publicly.
<!-- /funnel:v1 -->