Downloads · 30 days
65
69% of all-time downloads
Asher-1/2DGS-SLAM-GGML
2DGS-SLAM-GGML is a machine learning model from Asher-1. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
github: https://github.com/Asher-1/2DGS-SLAM-GGML
Downloads · 30 days
65
69% of all-time downloads
All-time downloads
94
Public
Repo size
4.9 GB
Likes
0
Public
Click a slice to open those files.
.gguf4.9 GB · 100%
From the Hugging Face model README
github: https://github.com/Asher-1/2DGS-SLAM-GGML
This directory contains the official MASt3R PyTorch weights (pytorch/) and the
GGUF engine weights (gguf/) exported by
scripts/convert.py, for use by cpp_ggml's pure C++
MASt3R inference engine and monocular SLAM.
MASt3R (MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors,
NAVER Labs, arXiv:2412.12392) is a joint two-view dense 3D reconstruction and
descriptor model. Given a pair of RGB images, it outputs a 29-channel tensor
per view: pts3d(3) + conf(1) + desc(24) + desc_conf(1) (NCHW, 512×512).
Variant used in this project: MASt3R_ViTLarge_BaseDecoder_512_catmlpdpt_metric
(verified against the architecture constants in scripts/convert.py):
| Component | Configuration |
|---|---|
| Patch embedding | 16×16, input 512×512×3 |
| Encoder | ViT-Large: 24 layers × 1024 dim × 16 heads |
| Decoder | Dual-branch: 2 × (12 layers × 768 dim × 12 heads), cross-attention |
| 2D RoPE | base=100.0, half-split rotation |
| DPT head | 4-layer fusion ([96,192,384,768] → 256 → 128) → pts3d+conf |
| catMLP head | 24-dim descriptor + desc_conf |
| Parameters | 688.6M (measured: sum of 1008 GGUF tensor elements, 688,638,088) |
pytorch/ — Official weights (source: https://download.europe.naverlabs.com/ComputerVision/MASt3R/)| File | Size | Purpose |
|---|---|---|
MASt3R_ViTLarge_BaseDecoder_512_catmlpdpt_metric.pth | 2.75 GB | Main model: two-view reconstruction + descriptors (the only input for GGUF export, scripts/export_all.sh) |
MASt3R_ViTLarge_BaseDecoder_512_catmlpdpt_metric_retrieval_trainingfree.pth | 8.4 MB | Optional: retrieval (training-free) head |
MASt3R_ViTLarge_BaseDecoder_512_catmlpdpt_metric_retrieval_codebook.pkl | 268 MB | Optional: retrieval codebook |
Note: an earlier README stated the main model as 1.4 GB; the actual download is 2.75 GB (the actual file prevails).
gguf/ — Engine weights (exported by scripts/export_all.sh, 1008 tensors each)| File | Size | Quantization strategy (scripts/convert.py) | Use case |
|---|---|---|---|
mast3r-f32.gguf | 2.75 GB | full fp32 | GPU high precision (CPU f32 is extremely slow) |
mast3r-f16.gguf | 1.38 GB | matmul/conv → fp16; bias/norm kept fp32 | GPU balanced precision/speed |
mast3r-q8.gguf | 772 MB | linear → Q8_0; conv → fp16; bias/norm → fp32 | Best portability, minimal precision loss (recommended) |
benchmarks/inference_speed.csv, RTX 4090)| engine | backend | precision | median (ms) | views/s |
|---|---|---|---|---|
| torch | cuda | fp16 (full-demo) | 269.7 | 3.71 |
| ggml | cuda | f32 | 185.6 | 5.39 |
| ggml | cuda | f16 | 168.3 | 5.94 |
| ggml | cuda | q8 | 168.6 | 5.93 |
| ggml | vulkan | f16 | 127.9 | 7.82 |
| ggml | cpu | q8 | 6342.7 | 0.16 |
benchmarks/README.md)| Channel | max_abs | mean_abs | rel% | corr |
|---|---|---|---|---|
| pts3d (3 ch) | ~0.05 | ~0.004 | ~0.3% | 1.000 |
| conf | ~0.001 | ~0.0002 | ~0.08% | 1.000 |
| desc (24 ch) | ~0.8 | ~0.08 | ~80%* | ~0.95 |
* The large relative error in the desc channel stems from quantization noise amplified by near-zero values after L2 normalization; the absolute difference is only 0.01–0.06/element (~1–3° angular error), which does not affect loop-closure retrieval ranking.
benchmarks/slam_ate.csv + official_tum_ate.md)| Sequence | This project ggml C++ (monocular f32 CUDA) | Official Ours* (uncalibrated) | Notes |
|---|---|---|---|
| TUM fr2_xyz | 0.461 (600 frames) | 0.020 (3669 frames) | Official uses the full RGB-D + GS pipeline |
| TUM fr3_sitting_halfsphere | 0.112 (60 frames) | — | Not published by the official |
Local re-measurement (RTX 3060, q8, CUDA, 60 frames fr2_xyz, 2026-08-25): ATE = 0.1497 m, single inference 689.9 ms/frame.
Different evaluation settings: this project is a lightweight C++ implementation with monocular input, no depth, no calibration, no IMU (Sim(3) tracking + PGO + loop closure), while the official is the full RGB-D + Gaussian splatting pipeline. Direct comparison is for reference only.
| Model | f32 | f16 | q8_0 |
|---|---|---|---|
| MASt3R (this project) | 2.75 GB | 1.38 GB | 772 MB |
| SAM 3 (sam3.cpp, ~850M params) | 3.2 GB | 1.7 GB | 1.0 GB |
# Download official weights (idempotent)
./scripts/download_models.sh
# Export the three GGUF variants
./scripts/export_all.sh
# Inference verification
./build/bin/mast3r-cli --device cuda models/gguf/mast3r-q8.gguf \
test_img_0.png test_img_1.png /tmp/out.bin