Downloads · 30 days
710
100% of all-time downloads
roskosmos19/Orca-4B-Instruct
Orca-4B-Instruct is a text generation model from roskosmos19. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Best price/performance 4B Instruct model
Downloads · 30 days
710
100% of all-time downloads
All-time downloads
710
Public
Parameters
4B
8.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
Best price/performance 4B Instruct model
Optimized for maximum capability at minimal cost. Based on the strong Qwen3-4B architecture with aggressive efficiency tuning.
| Metric | Original Qwen3-4B-Instruct | Dolphin-4B-Instruct-0409 |
|---|---|---|
| Native context | 262k | 16k (covers >95% of real use) |
| KV-cache VRAM (16k ctx) | High | Much lower |
| Multimodal special tokens | Many (vision etc.) | Removed (leaner) |
| Generation defaults | Generic | Tuned for quality |
| Instruction strength | Good | Stronger system prompt |
| Typical quantized size (Q4) | ~2.5 GB | ~2.5 GB (same, but faster) |
→ Same 4B intelligence, significantly cheaper to run, better focused answers.
{
"temperature": 0.5,
"top_p": 0.85,
"top_k": 30,
"repetition_penalty": 1.05
}
These settings produce more precise, less repetitive and higher-quality answers while staying efficient.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "./Dolphin-4B-Instruct-0409"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
messages = [
{"role": "user", "content": "Explain the difference between CPU and GPU in simple terms."}
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0][len(inputs.input_ids[0]):], skip_special_tokens=True))
--max-model-len 16384Apache 2.0
This package contains the fully tuned config + tokenizer.
Place the original Qwen3-4B safetensors weights next to the index file (or convert to GGUF) and you have a ready-to-run high price/performance 4B model.