Downloads · 30 days
57
10% of all-time downloads
intellectlabs/Kepler-8B-Instruct-v2
Kepler-8B-Instruct-v2 is a text generation model from intellectlabs. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Downloads · 30 days
57
10% of all-time downloads
All-time downloads
580
Public
Parameters
7.6B
15.2 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors15.2 GB · 100%
From the Hugging Face model README
License: Apache 2.0 · Format: GGUF · Base: Kepler-8B-Instruct
Original Model · llama.cpp · Discussions
</div>Kepler-8B-Instruct is an 8-billion-parameter assistant built by merging Qwen/Qwen2.5-7B-Instruct (85%) with deepseek-ai/DeepSeek-R1-Distill-Qwen-7B (15%) using mergekit. The blend keeps Qwen2.5-Instruct's clean, reliable instruction-following as the foundation while folding in a touch of DeepSeek's reasoning-distilled behavior.
This repository packages that model as GGUF — the universal format for llama.cpp and everything built on it. One file, no Python environment, no CUDA setup. It runs on a gaming laptop, a Raspberry Pi-class board, or a headless server with equal ease.
Highlights:
./llama-cli -hf intellectlabs/Kepler-8B-Instruct-GGUF:Q4_K_M -p "Explain quantum entanglement simply."
Or download a specific file manually:
huggingface-cli download intellectlabs/Kepler-8B-Instruct-GGUF Kepler-8B-Instruct-Q4_K_M.gguf --local-dir .
./llama-cli -m Kepler-8B-Instruct-Q4_K_M.gguf -p "Hello, who are you?" -cnv
Search intellectlabs/Kepler-8B-Instruct-GGUF directly in the LM Studio model search bar and download your preferred quant.
ollama run hf.co/intellectlabs/Kepler-8B-Instruct-GGUF:Q4_K_M
| File | Quant | Size (approx.) | Quality | Recommended For |
|---|---|---|---|---|
Kepler-8B-Instruct-Q2_K.gguf | Q2_K | ~3.1 GB | ⭐️ | Extreme low-RAM devices only |
Kepler-8B-Instruct-Q3_K_M.gguf | Q3_K_M | ~4.0 GB | ⭐️⭐️ | Low-RAM systems, testing |
Kepler-8B-Instruct-Q4_0.gguf | Q4_0 | ~4.7 GB | ⭐️⭐️⭐️ | Legacy-hardware compatibility |
Kepler-8B-Instruct-Q4_K_M.gguf | Q4_K_M | ~4.9 GB | ⭐️⭐️⭐️⭐️ | Best balance — recommended default |
Kepler-8B-Instruct-Q5_K_M.gguf | Q5_K_M | ~5.7 GB | ⭐️⭐️⭐️⭐️ | Higher quality, still efficient |
Kepler-8B-Instruct-Q8_0.gguf | Q8_0 | ~8.5 GB | ⭐️⭐️⭐️⭐️⭐️ | Near-lossless, if you have the RAM |
Kepler-8B-Instruct-F16.gguf | F16 | ~16 GB | ⭐️⭐️⭐️⭐️⭐️ | Full precision, GPU/server use |
💡 New here? Start with Q4_K_M — it's the sweet spot most people use: small enough to run comfortably, strong enough to feel close to the full model.
This model uses the ChatML template (inherited from its Qwen2 base):
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
Most tools (llama.cpp -cnv, LM Studio, Ollama) apply this automatically via the embedded chat template — no manual formatting needed.
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B linearly blended into Qwen/Qwen2.5-7B-Instruct (15%/85%) using mergekit, with the tokenizer taken from the base model.llama.cpp's convert_hf_to_gguf.py.llama-quantize, covering everything from ultra-compact (Q2_K) to near-lossless (Q8_0).If this was useful, a ⭐️ on the repo helps others find it.
</div>