Downloads · 30 days
262
100% of all-time downloads
roskosmos19/Dolphin-4B
Dolphin-4B is a text generation model from roskosmos19. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Dolphin is a highly capable, efficient, and practical language model focused on maximum usefulness, strong reasoning, excellent coding ability, and reliable agentic behavior.
Downloads · 30 days
262
100% of all-time downloads
All-time downloads
262
Public
Parameters
4.1B
8.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors8.2 GB · 100%
From the Hugging Face model README
Dolphin is a highly capable, efficient, and practical language model focused on maximum usefulness, strong reasoning, excellent coding ability, and reliable agentic behavior.
This is a carefully configured and optimized release based on the Spark-X2.5-4B architecture, fine-tuned in identity and behavior to deliver elite-level performance in everyday use, coding, tool use, and complex tasks.
| Property | Value |
|---|---|
| Parameters | ~4.1B |
| Context Length | 1,048,576 tokens |
| Architecture | Hybrid Attention (Sliding Window + Full) |
| Vocabulary Size | 131,072 |
| Precision | bfloat16 |
| License | Apache 2.0 |
{
"temperature": 1.0,
"top_p": 0.95,
"top_k": -1,
"repetition_penalty": 1.0,
"presence_penalty": 0.0,
"frequency_penalty": 0.0
}
These settings work particularly well with the built-in thinking mode.
Dolphin comes with a strong default system prompt focused on:
You can still override it with your own system message.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "path/to/Dolphin-X2.5-4B"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="bfloat16",
device_map="auto",
trust_remote_code=True
)
messages = [
{"role": "user", "content": "Write a clean Python function that calculates the Fibonacci sequence up to n."}
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=1024, temperature=1.0, top_p=0.95)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
The model is fully compatible with the same inference stacks as the original Spark-X2.5 architecture (vLLM, SGLang, llama.cpp, MLX, Ollama, LM Studio, etc.).
Use the included chat template and the recommended sampling parameters above for best results.
The model uses a clean, modern chat template with support for:
<think>...</think>)Thinking is enabled by default. You can disable it per request if desired.
Dolphin is built with one clear goal:
Be as useful, accurate, and high-quality as possible in real-world use.
No fluff. No unnecessary restrictions. Just strong, reliable performance.
Apache 2.0
Based on the excellent Spark-X2.5 architecture and training work by the SparkLLM / XHToken team.