Downloads · 30 days
88
2% of all-time downloads
Lexuselizar/pegasus-mini
pegasus-mini is a text generation model from Lexuselizar. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
A small, on-device chat model distilled from Qwen2.5-1.5B-Instruct and shipped as a q4 GGUF so it runs offline in llama.cpp / Ollama / a phone shell / the browser (WebGPU). It is the offline brain of the hybrid Pegasu…
Downloads · 30 days
88
2% of all-time downloads
All-time downloads
4K
Public
Repo size
9.9 GB
Likes
0
Public
Click a slice to open those files.
.gguf7.9 GB · 100%
From the Hugging Face model README
A small, on-device chat model distilled from Qwen2.5-1.5B-Instruct and shipped as a
q4 GGUF so it runs offline in llama.cpp / Ollama / a phone shell / the browser (WebGPU).
It is the offline brain of the hybrid PegasusLink app at https://reverseml.online
(online → cloud model + web search; offline → this).
Independent / solo project, open beta. Feedback and issues welcome.
Be clear about this, because they are different things:
app-memory.js / app-chem.js) and wrap any local model; they are not baked into
these weights. If you just load this GGUF in llama.cpp, you get the model, not the stack.So: judge the GGUF here as a 1.5B chat model. The architecture writeup is on the site.
llama.cpp
./llama-cli -m pegasus-mini-q4.gguf -p "Balance: H2 + O2 -> H2O" -ngl 99
Ollama
printf 'FROM ./pegasus-mini-q4.gguf\nPARAMETER temperature 0\nPARAMETER stop "<|im_end|>"\n' > Modelfile
ollama create pegasus-mini -f Modelfile
ollama run pegasus-mini "What is the pH of a neutral solution at 25 C?"
Phone: load the GGUF in a shell like ChatterUI. Browser: the WebLLM/WebGPU build (q4f16_1) is served from the site — zero install.
Prompt format is Qwen2 ChatML (<|im_start|> / <|im_end|>).
Measured with Ollama, q4 GGUF, CPU-only (no GPU) on a 4-core AMD EPYC-Genoa VM:
| metric | value |
|---|---|
| eval (generation) rate | ~33 tokens/s |
| prompt eval rate | ~64 tokens/s |
| cold load | ~1.4 s |
That's CPU-only; on a laptop GPU or via WebGPU in the browser it's faster. The point is it's comfortably interactive on commodity hardware with no accelerator.
Prompt: Explain what a Kalman filter does in two sentences.
A Kalman filter is an algorithm that uses a combination of measurements and predictions to estimate the state of a system, such as a robot or an aircraft, by updating its estimates based on new information. It does this by using a mathematical model of the system to predict its future state, then comparing those predictions to actual measurements to refine them — it is widely used in robotics, navigation, and signal processing for estimating unknown variables under uncertainty.
Qwen2.5-1.5B-Instruct (Apache-2.0).General offline assistant for low-resource / private / edge settings: quick Q&A, coding help, math, deterministic chemistry balancing (via the app), and as a base to distill on your own data.
Released under Apache-2.0, inheriting from the Qwen2.5-1.5B-Instruct base. Please
keep the Qwen attribution when redistributing. The weights are derived via distillation/
fine-tuning of that base.