Downloads · 30 days
119
41% of all-time downloads
sovthpaw/omnistep-new
omnistep-new is a image-text-to-text model from sovthpaw. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
119
41% of all-time downloads
All-time downloads
293
Public
Parameters
8.2B
33.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.gguf16.4 GB · 49%
From the Hugging Face model README

An omnimodal music generator + chat assistant. Replaces sovthpaw/omnistep-12a3b (v1 transitional baseline) as the current OmniStep in the OmniSenter family.
OmniStep-new is a single bundled model that handles:
| Capability | How |
|---|---|
| Chat / instruction following | Qwen3-8B text backbone |
| Image + video understanding | Cosmos multimodal heads (built into the chat template as `< |
| Music generation (lyrics → music) | ACE-Step v1.5 turbo DiT + 1.7B LM + VAE, attached as modules |
| Tool / function calling | <tool_call> tags in chat template, supports OpenAI-style tool schemas |
| Long context | 40K native, YaRN-extendable to 256K+ |
It is not the agentic SFT variant — that's a separate LoRA on top. This is the base, ready for SFT warm-start or direct inference.
OmniStep-new (16GB F16 GGUF bundle)
├─ Qwen3-8B text backbone (8.2B params, 36 layers, q_norm/k_norm)
├─ Cosmos multimodal heads (vision encoder + projector)
└─ ACE-Step music modules (1.7B LM + DiT v1.5 turbo + VAE)
Built via Darwin Family weight-space merging (MRI-Trust Fusion) from:
nvidia/Cosmos3-Nano (text body extracted)Qwen/Qwen3-8B (gen-1 SFT-merged base)SouthpawIN/evolutionary-training/scripts/cosmos_qwen3_darwin_merge.py| File | Size | Notes |
|---|---|---|
OmniStep-new-F16.gguf | 16 GB | llama.cpp F16 single-file GGUF |
model-00001..7-of-7.safetensors | 16 GB | HF sharded safetensors |
config.json | <1 KB | Qwen3-compatible |
tokenizer.json / vocab.json / merges.txt | ~16 MB | Standard Qwen3 tokenizer + 26 multimodal/tool special tokens |
chat_template.json | ~5 KB | Tool-calling + vision/video/image_pad tokens |
llama-server -m OmniStep-new-F16.gguf \
--host 127.0.0.1 --port 9080 \
-ngl 99 -c 262144 -fa on \
-ctk turbo4 -ctv turbo4 --no-mmap
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("sovthpaw/omnistep-new")
model = AutoModelForCausalLM.from_pretrained(
"sovthpaw/omnistep-new",
device_map="auto",
torch_dtype="bfloat16",
)
out = model.generate(**tok("Hello, who are you?", return_tensors="pt").to("cuda"),
max_new_tokens=200, do_sample=True, temperature=0.7)
print(tok.decode(out[0]))
import json
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("sovthpaw/omnistep-new")
model = AutoModelForCausalLM.from_pretrained("sovthpaw/omnistep-new", device_map="auto", torch_dtype="bfloat16")
tools = [{
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}
}]
msgs = [{"role": "user", "content": "What's the weather in Tokyo?"}]
prompt = tok.apply_chat_template(msgs, tools=tools, add_generation_prompt=True, tokenize=False)
out = model.generate(**tok(prompt, return_tensors="pt").to("cuda"), max_new_tokens=200)
print(tok.decode(out[0][len(tok.encode(prompt)):]))
# → <tool_call>{"name": "get_weather", "arguments": {"city": "Tokyo"}}</tool_call>
This is gen 0. The roadmap:
SouthpawIN/evolutionary-trainingscripts/cosmos_qwen3_darwin_merge.pythe-omni-family.md — OmniStep = multimodal native + music + agentic backbone. This is the multimodal + music half; the agentic SFT is the LoRA adapter on top.TOWARDS SELF-IMPROVEMENT — Chris (via Nous Girl), 2026-06-22