Downloads · 30 days
50
34% of all-time downloads
Null-Byte/PhoneLLM-alpha-1-Q4_K_M
PhoneLLM-alpha-1-Q4_K_M is a text generation model from Null-Byte. Use it when you need the model to write or continue text. The card lists the license as bsd-2-clause.
This is a Q4KM quantization of PhoneLLM Alpha 1 using llama.cpp's imatrix-guided quantization:
Downloads · 30 days
50
34% of all-time downloads
All-time downloads
146
Public
Repo size
24.5 GB
Likes
0
Public
Click a slice to open those files.
.gguf24.5 GB · 100%
From the Hugging Face model README
| Model | PhoneLLM Alpha 1 (pipecat-ai/phonellm-alpha-1) |
| This repo | Q4_K_M GGUF quantization with imatrix calibration |
| Base model | NVIDIA Nemotron 3 Nano 30B-A3B |
| Architecture | Hybrid Mamba-Transformer mixture-of-experts; 30B total parameters, 3.5B active |
| Quantization | Q4_K_M (mixed precision: q4_k experts/attn-QK, q6_k embeddings/output/attn-V, f16 norms) + imatrix |
| Calibration | C4 English subset (~101k tokens); PPL = 13.77 ± 0.15 |
| File size | 22.83 GB |
| Context length | 262,144 tokens (supports full context; ~2.5GB KV at q8_0) |
| Recommended inference settings | temperature=0, thinking disabled |
| Runtime | llama.cpp b10673+ (CUDA, Vulkan, or CPU) |
| Language | English |
| License | BSD 2-Clause; derivative of an NVIDIA Nemotron Open Model License work — see License |
| Developed by | Daily / the Pipecat team |
This is a Q4_K_M quantization of PhoneLLM Alpha 1 using llama.cpp's imatrix-guided quantization:
llama-quantize --imatrix (PPL-calibrated mixed precision)--no-mmap allocates full weights to RAM at load time)llama-server \
-m phonellm-alpha-1-q4_k_m.gguf \
--n-gpu-layers all \
--split-mode layer \
--tensor-split 7,3 \
--main-gpu 0 \
--ctx-size 262144 \
-n 8192 \
--parallel 1 \
--flash-attn on \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--cache-ram 4096 \
--no-context-shift \
--jinja \
--temp 0 \
--threads 10 \
--no-mmap
Parameter notes:
| Flag | Value | Why |
|---|---|---|
--tensor-split 7,3 | ~70/30 split | Proportional to VRAM (24:10); adjust for your GPU ratio |
--main-gpu 0 | Primary GPU | Fastest card handles KV cache + compute buffers |
--ctx-size 262144 | Full context | Model supports it; ~2.5GB KV at q8_0 (only ~20 of 52 layers are full-attention) |
--flash-attn on | Required at long ctx | O(n) memory for attention; mandatory above ~32K context |
--cache-type-k/v q8_0 | Q8 KV cache | Halves KV size vs f16 with negligible quality loss |
--parallel 1 | Single slot | All KV dedicated to one conversation (phone agent use case) |
--temp 0 | Greedy decoding | Required per PhoneLLM training; do not change |
--no-mmap | Direct load | Avoids file-backed mapping overhead for GPU-resident model |
llama-server \
-m phonellm-alpha-1-q4_k_m.gguf \
--n-gpu-layers all \
--ctx-size 131072 \
--flash-attn on \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--jinja \
--temp 0 \
--no-mmap
For setups where the model doesn't fully fit in VRAM, offload routed experts to system RAM:
llama-server \
-m phonellm-alpha-1-q4_k_m.gguf \
--n-gpu-layers all \
--cpu-moe \
--ctx-size 32768 \
--flash-attn on \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--jinja \
--temp 0
This keeps attention + shared experts on GPU and streams active routed experts over PCIe from RAM. Requires fast DDR4/DDR5 (64GB+). See the llama.cpp MoE offload guide for tuning.
The Pipecat team is pleased to announce the release of PhoneLLM Alpha 1, an open-weights model for voice agent use cases.
This release is the result of our ongoing work training small, open-weights LLMs for low-latency and multi-turn agentic workloads.
When paired with transcription and text-to-speech models through a framework like Pipecat, PhoneLLM can handle incoming calls for financial services, healthcare, retail, and hospitality customer service, and perform common outbound calling agent tasks.
PhoneLLM runs at a fraction of the cost and latency of larger, general-purpose models, while delivering comparable performance for specific use cases. For example, PhoneLLM performs on par with GPT 5.6 Terra, but 94% cheaper and with 1,300ms faster P95 time-to-first-token.
PhoneLLM is an open model, so you can run it on your own infrastructure. The model is released under the BSD license, with no commercial restrictions.
PhoneLLM Alpha 1 is a full-parameter fine-tune of NVIDIA's Nemotron 3 Nano 30B-A3B model, trained using the NVIDIA NeMo framework.
Like Nemotron Nano, PhoneLLM is a mixture-of-experts (MoE) model, with 3.5B active parameters, allowing for high-speed inference at low cost.
Important: set temperature to 0 and disable thinking. These two settings align with how the model was trained.
PhoneLLM Alpha 1 is released under the BSD 2-Clause License.
PhoneLLM is a derivative work of NVIDIA Nemotron 3 Nano 30B-A3B, which is licensed under the NVIDIA Nemotron Open Model License. Under Section 3 (Redistribution) of that license, if you redistribute this model or your own derivatives of it, you must (a) include a copy of the NVIDIA Nemotron Open Model License, and (b) retain the NVIDIA copyright and attribution notices. Our BSD 2-Clause terms apply to our modifications and to the model as a whole, as Section 3 permits; the NVIDIA license continues to apply to the underlying Nemotron work. "Nemotron" and "NVIDIA" are trademarks of NVIDIA Corporation, used here only to describe the origin of the base model.