Downloads · 30 days
315
30% of all-time downloads
AmareshHebbar/pocketllm-models
pocketllm-models is a text generation model from AmareshHebbar. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Downloads · 30 days
315
30% of all-time downloads
All-time downloads
1.1K
Public
Repo size
11.4 GB
Likes
0
Public
Click a slice to open those files.
.gguf10.1 GB · 88%
From the Hugging Face model README
The official model collection for PocketLLM
Private • Offline • On-Device AI for Android
</div>PocketLLM is an Android app that runs large language models completely on your device — no internet required during inference, no data sent to the cloud, no subscriptions.
This repository hosts the curated model collection optimized for mobile edge inference using:
.bin and .task files — Gemma family).gguf files — Llama, Phi, Qwen, SmolLM, Gemma 3)| Model | File | Size | Format | RAM |
|---|---|---|---|---|
| Qwen 2.5 0.5B | qwen2.5-0.5b-instruct-q4_k_m.gguf | 0.4 GB | GGUF | 2 GB |
| Gemma 3 1B | gemma-3-1b-it-q4_k_m.gguf | 0.7 GB | GGUF | 2 GB |
| Llama 3.2 1B | llama-3.2-1b-instruct-q4_k_m.gguf | 0.8 GB | GGUF | 3 GB |
| SmolLM2 1.7B | smollm2-1.7b-instruct-q4_k_m.gguf | 1.0 GB | GGUF | 3 GB |
| Model | File | Size | Format | RAM |
|---|---|---|---|---|
| Gemma 1.1 2B (CPU) | gemma-1.1-2b-it-cpu-int4.bin | 1.35 GB | MediaPipe | 4 GB |
| Gemma 1.1 2B (GPU) | gemma-1.1-2b-it-gpu-int4.bin | 1.35 GB | MediaPipe | 4 GB |
| Llama 3.2 3B | llama-3.2-3b-instruct-q4_k_m.gguf | 2.0 GB | GGUF | 4 GB |
| Model | File | Size | Format | RAM |
|---|---|---|---|---|
| Phi-3.5 Mini | phi-3.5-mini-instruct-q4_k_m.gguf | 2.4 GB | GGUF | 5 GB |
| Gemma 3 4B | gemma-3-4b-it-q4_k_m.gguf | 2.8 GB | GGUF | 6 GB |
Models are downloaded directly inside the PocketLLM app. Open the Model Store tab, select a model, and tap Download. The app handles everything automatically.
Direct download URLs:
https://huggingface.co/AmareshHebbar/pocketllm-models/resolve/main/<filename>
Coming soon — within the next 2 weeks
We are fine-tuning these base models specifically for on-device conversational AI on mobile:
pocketllm-llama-1b-v1.gguf — Llama 3.2 1B fine-tuned for mobile chatpocketllm-gemma-1b-v1.gguf — Gemma 3 1B fine-tuned for persona consistencyYour phone has... Best model to start with
─────────────────────────────────────────────────────
2 GB RAM (budget) → Qwen 2.5 0.5B (fastest)
3 GB RAM → Llama 3.2 1B (balanced speed)
4 GB RAM (mid-range) → Gemma 1.1 2B (best all-rounder)
6 GB RAM → Llama 3.2 3B (better quality)
8 GB RAM (flagship) → Phi-3.5 Mini (best for coding)
.bin, .task)react-native-llm-mediapipe.gguf)llama.rnpocketllm-models/
├── gemma-1.1-2b-it-cpu-int4.bin ← Gemma 2B CPU (MediaPipe)
├── gemma-1.1-2b-it-gpu-int4.bin ← Gemma 2B GPU (MediaPipe)
├── gemma-3-1b-it-q4_k_m.gguf ← Gemma 3 1B (GGUF)
├── gemma-3-4b-it-q4_k_m.gguf ← Gemma 3 4B (GGUF)
├── llama-3.2-1b-instruct-q4_k_m.gguf ← Llama 3.2 1B (GGUF)
├── llama-3.2-3b-instruct-q4_k_m.gguf ← Llama 3.2 3B (GGUF)
├── phi-3.5-mini-instruct-q4_k_m.gguf ← Phi-3.5 Mini (GGUF)
├── qwen2.5-0.5b-instruct-q4_k_m.gguf ← Qwen 2.5 0.5B (GGUF)
└── smollm2-1.7b-instruct-q4_k_m.gguf ← SmolLM2 1.7B (GGUF)
Amaresh Hebbar
Building PocketLLM: the only mobile app that runs a full AI agent stack — smart routing, persona memory, MCP tools — completely offline on Android.
"Your AI. Your phone. Nobody else's business."