Downloads · 30 days
0
zst50/voice-chatbot-android
voice-chatbot-android is a machine learning model from zst50. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
The .litertlm models on this card already use the QAT that is discussed in the blog post. The most popular file, gemma-4-E2B-it.litertlm, uses a mixture of int2, int4 and int8 to keep it small, fast and efficient.
Downloads · 30 days
0
Access
Public
Updated Aug 14, 2026
Repo size
3.4 GB
Likes
0
Public
Click a slice to open those files.
.apk2.9 GB · 81%
From the Hugging Face model README
The .litertlm models on this card already use the QAT that is discussed in the blog post. The most popular file,
gemma-4-E2B-it.litertlm, uses a mixture of int2, int4 and int8 to keep it small, fast and efficient.
A Kotlin/Compose Android app implementing a fully local, interruptible voice chatbot:
mic ─► Silero VAD (sherpa-onnx)
│ endpointing + barge-in
▼
Whisper tiny.en Q4 (whisper.cpp) ──► transcript
▼
LLM (streaming):
├─ Local: LiteRT-LM (`.litertlm`, e.g. Qwen3-0.6B)
└─ Remote: OpenAI-compatible `/v1/chat/completions` (SSE)
▼
SentenceSplitter ──► per-sentence VITS TTS (sherpa-onnx + Piper voice)
▼
AudioTrack (playback; barge-in stops it instantly)
Barge-in: speak while the assistant is talking → playback and generation stop immediately and your new utterance becomes the next turn.
| Piece | Runtime | Model | Size |
|---|---|---|---|
| STT | whisper.cpp (JNI, built via NDK/CMake) | ggml-tiny.en-q4_0.bin (Q4) | ~45 MB |
| VAD | sherpa-onnx AAR (official prebuilt sherpa-onnx-1.13.4.aar) | silero_vad.onnx | ~2 MB |
| TTS | sherpa-onnx VITS (Piper) | vits-piper-en_US-lessac-medium (+ espeak-ng-data) | ~62 MB |
| LLM (local) | LiteRT-LM com.google.ai.edge.litertlm:litertlm-android:0.15.0 | litert-community/Qwen3-0.6B → Qwen3-0.6B.litertlm (dynamic INT8) | 586 MB |
| LLM (remote) | OkHttp + kotlinx.serialization (OpenAI-compatible chat completions) | any remote model, e.g. qwen2.5-0.5b-instruct | n/a |
All models are Apache-2.0 or MIT except the Piper voice/GPL espeak-ng data (bundled in the sherpa-onnx tts-model tarball) — fine for local testing.
27.2.12479018, CMake 3.22.1),
JDK 17, and a connected Android device (arm64, 8 GB RAM recommended).adb on PATH.Toolchain: AGP 8.13.2 · Gradle 8.13 (wrapper) · Kotlin 2.3.21 · JDK 17.
whisper.cpp is vendored under third_party/ (committed), so the project builds
out-of-the-box once the SDK/NDK are present.
# 0. JDK 17 (Gradle 8.13 does not support newer JDKs)
# macOS/Homebrew: brew install openjdk@17
export JAVA_HOME=/opt/homebrew/opt/openjdk@17/libexec/openjdk.jdk/Contents/Home
# 0b. Point Gradle at the Android SDK (or set ANDROID_HOME)
echo "sdk.dir=$HOME/Library/Android/sdk" > local.properties
# 1. (first time) install NDK/cmake if missing
"$ANDROID_HOME"/cmdline-tools/latest/bin/sdkmanager \
"platforms;android-35" "ndk;27.2.12479018" "cmake;3.22.1"
# 2. Build (Gradle downloads the sherpa-onnx AAR into app/libs on first run)
./gradlew :app:assembleDebug
# 3. Install
adb install -r app/build/outputs/apk/debug/app-debug.apk
Unit tests: ./gradlew :app:testDebugUnitTest
Only if
third_party/whisper.cppis ever missing (e.g. a fresh export that dropped it):tools/fetch-deps.shre-clones it.
Models are bundled in the APK (assets/models/) and copied to the app's
private internal storage on first launch — no adb push needed. Storage used on
the device: ~240 MB under /data/data/com.example.voicechatbot/files/models/.
The APK is ~500 MB because of the bundled models; devices need ~1.5 GB free to install it.
Preinstalled: SmolLM2 135M (LLM) + Whisper tiny.en Q4 + Silero VAD + Piper TTS.
Download in-app: bigger LLMs (SmolLM2 360M, Qwen3 0.6B, Gemma 4 E2B/E4B/12B) and Whisper models are selected in Settings via preset chips that auto-download the file if it isn't already saved (progress shown on the chip).
| id | Model | Size | Preinstalled | Thinking | License |
|---|---|---|---|---|---|
smollm2-135m | SmolLM2 135M | 136 MB | ✅ | no | Apache-2.0 |
smollm2-360m | SmolLM2 360M | 356 MB | download | no | Apache-2.0 |
qwen3-0.6b | Qwen3 0.6B int8 | 586 MB | download | yes (stripped) | Apache-2.0 |
gemma-4-e2b | Gemma 4 E2B (mobile 2/4/8-bit) | 2583 MB | download | yes (stripped) | Apache-2.0 |
gemma-4-e4b | Gemma 4 E4B (mobile 2/4/8-bit) | 3654 MB | download | yes (stripped) | Apache-2.0 |
gemma-4-12b | Gemma 4 12B | 6548 MB | download | yes (stripped) | Apache-2.0 |
To rebuild with bundled models, the files must be present in ./models/:
tools/download-models.sh # downloads all LLM presets + whisper/VAD/TTS
tools/download-models.sh smollm2-360m # or just one
tools/download-models.sh list # show presets
(tools/push-models.sh is obsolete — models are in the APK now.)
LFM-700M: there is no
.litertlmbuild of LFM2-700M on HF yet, and a conversion requires thelitert-torchgeneric-HF-export pipeline (heavy, not guaranteed for the architecture). The SmolLM2 presets are the practical small, non-thinking, Apache-2.0 alternatives.
Gemma 4 E2B / E4B: official
litert-communityLiteRT-LM builds ofgoogle/gemma-4-E2B-it/google/gemma-4-E4B-it(Apache-2.0), using Google's "Gemma-4 mobile" 2/4/8-bit (LUT) quantization. LiteRT-LM memory-maps the weights straight from the.litertlmfile on disk — the embedding tables are mmap'd and never fully loaded into RAM — so the ~2.6 GB / ~3.7 GB files don't need to fit in memory. Both support up to 32k context. Android benchmarks (S26 Ultra, 2048 ctx): E2B ~1.7 GB CPU / ~0.7 GB GPU process RAM at ~47–52 tok/s; E4B ~3.3 GB CPU / ~0.7 GB GPU at ~18–22 tok/s — E4B really wants a 12 GB-class phone and the GPU backend. Settings → Context window (LiteRT-LMmaxNumTokens) caps the KV cache so bigger models fit in less RAM. Gemma 4 12B (gemma-4-12B-it.litertlm, ~6.5 GB) is the largest option — realistically it needs a 16 GB-class device with the GPU backend.
TTS voices: only the bundled
en_US-lessac-mediumvoice is used. In-app voice downloads were removed; Settings → Voice (VITS) just shows the preinstalled voice (blank the model path for text-only chat).
Whisper (STT): the bundled model is
ggml-tiny.en-q4_0.bin(~45 MB). Better variants are available as Settings → Whisper model presets (auto-downloaded on selection): unquantized tiny.en / tiny (FP16, ~75 MB), tiny.en Q8_0 (~42 MB), or base.en Q8_0 (~78 MB). Q4_0 builds are not published on Hugging Face (tools/download-models.shquantizes them locally), so Q8_0 is the closest official base.en quantized download. Larger models are slower but more accurate, especially for non-English speech (use the multilingualtiny).
Listening → Transcribing → Thinking → Speaking, and hear the reply sentence-by-sentence.http://<your-server>:8000/v1) + model name, save & restart.
Works with llama.cpp server, vLLM, Ollama, etc.
Context length: Settings → Context length (512…32768). The app sends
it as n_ctx, which llama.cpp honours per-request. "Server default" sends
nothing — the server decides (vLLM/Ollama manage context server-side).Tips:
VOICE_COMMUNICATION + AcousticEchoCanceler when available.app/src/main/java/com/example/voicechatbot/
MainActivity.kt Compose UI + settings dialog
config/AppConfig.kt persisted settings & model paths
pipeline/VoicePipelineEngine.kt state machine, barge-in, sentence streaming
audio/AudioRecorder.kt 16 kHz capture (+AEC)
audio/TtsPlayer.kt AudioTrack queue with generation-based barge-in
vad/VadEngine.kt Silero VAD (sherpa-onnx)
stt/WhisperTranscriber.kt whisper.cpp JNI wrapper
stt/WhisperNative.kt JNI declarations
llm/ChatClient.kt backend interface
llm/LitertChatClient.kt LiteRT-LM local inference (Flow streaming)
llm/OpenAiChatClient.kt OpenAI-compatible remote client (SSE)
llm/SentenceSplitter.kt token stream → sentence boundaries
llm/PromptBuilder.kt system prompt
tts/TtsEngine.kt VITS/Piper synthesis (sherpa-onnx)
app/src/main/cpp/ whisper.cpp CMake + JNI bridge
third_party/whisper.cpp vendored (committed, v1.7.5)
tools/download-models.sh model fetch + Q4 quantize (local)
tools/push-models.sh adb push models to device
arm64-v8a).