Downloads · 30 days
0
batiai/llamacpp-server-macos
llamacpp-server-macos is a machine learning model from batiai. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for llama.cpp. The card lists the license as mit.
Self-contained llama-server binaries for macOS, built so that BatiFlow can run on-device multimodal (image + audio) models like Gemma 4 12B that Ollama can't serve yet.
Downloads · 30 days
0
Access
Public
Updated Jun 6, 2026
Repo size
30.9 MB
Likes
0
Public
Click a slice to open those files.
Other30.9 MB · 100%
From the Hugging Face model README
llama-server — macOS (Apple Silicon) static builds, by BatiAISelf-contained llama-server binaries for macOS, built so that BatiFlow can run on-device multimodal (image + audio) models like Gemma 4 12B that Ollama can't serve yet.
Ollama 0.20 doesn't know the Gemma 4
gemma4uv/gemma4uavision/audio projectors.llama-serverfrom recent llama.cpp master does. These are the macOS binaries that make that work locally.
✅ Live (2026-06-06) — built on Apple Silicon, verified (/props → vision+audio, image OCR), zero external deps. arm64 only (universal2 not needed — all targets are Apple Silicon).
| File | Size | What |
|---|---|---|
llama-server-macos-arm64 | 20 MB | static llama-server (Apple Silicon, Metal) — BatiFlow sidecar |
llama-mtmd-cli-macos-arm64 | 9.8 MB | static multimodal CLI (fallback) |
BUILD-INFO.txt | — | commit hash + sha256 + cmake flags |
308f61c31f083251ce8150f10b9ef97679b500b575cb1b97dff5de3ed3dbc3b79bbbda3cc33cb20f2e9287c3d0b685c46ad4c7dbgit clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
git checkout 308f61c # gemma4 projector merged (gemma4v/uv/a/ua clip graphs)
cmake -B build -DGGML_METAL=ON -DLLAMA_CURL=OFF \
-DBUILD_SHARED_LIBS=OFF \ # static single file (no dylib deps)
-DLLAMA_OPENSSL=OFF \ # drop homebrew openssl@3 dep (crashes on Macs without homebrew)
-DGGML_METAL_EMBED_LIBRARY=ON \ # embed Metal shaders (no separate .metal file → true single binary)
-DCMAKE_BUILD_TYPE=Release
cmake --build build --target llama-server llama-mtmd-cli -j
# static check: `otool -L build/bin/llama-server | grep -iE 'libllama|libggml|libmtmd|openssl|ssl'` → empty
llama-server-macos-arm64 \
-m gemma-4-12B-it-Q4_K_M.gguf \
--mmproj mmproj-google-gemma-4-12B-it-BF16.gguf \
--host 127.0.0.1 --port 8899 -ngl 99 -c 8192 --jinja
# GET /props → {"vision": true, "audio": true}
# POST /v1/chat/completions with image_url / input_audio (OpenAI-compatible)
GGUF + mmproj: batiai/gemma-4-12B-it-GGUF (filenames are pinned/stable).
--jinja required. Gemma 4 is a reasoning model → give generous max_tokens (answer in reasoning_content + content).llama-server works with any GGUF, not only Gemma 4.