Downloads · 30 days
0
gopi87/qwennextflashtemplate
qwennextflashtemplate is a machine learning model from gopi87. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
this template works very well with calude code and ikllama.cpp and i tested it
Downloads · 30 days
0
Access
Public
Updated Sep 3, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.jinja12.1 KB · 83%
From the Hugging Face model README
this template works very well with calude code and ik_llama.cpp and i tested it
echo 0 | sudo tee /proc/sys/kernel/numa_balancing
CUDA_VISIBLE_DEVICES=2,3,0,1
numactl --interleave=all
~/ik_llama.cpp/build/bin/llama-server
--model /mnt/nvme/Qwen3.8-Flash-Next-Q8_0-00001-of-00006.gguf
--chat-template-file /mnt/nvme/or.jinja
--chat-template-kwargs '{"reasoning_effort":"xhigh"}'
--tensor-split 2.5,4,0.6,0.2
-ot "blk.(47).ffn_.*_exps.=CUDA3"
-cmoe
--numa distribute
-c 160000
--batch-size 2500
--ubatch-size 2500
--cache-type-k q8_0
--cache-type-v q8_0
--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.0
--presence-penalty 0.0
--repeat-penalty 1.0
--parallel 1
--threads 42
--threads-batch 42
-ngl 100
--host 127.0.0.1
--port 8082
--jinja