Downloads · 30 days
13
20% of all-time downloads
hi1hello/FriendliAI-broken-model-fixed
FriendliAI-broken-model-fixed is a text generation model from hi1hello. Use it when you need the model to write or continue text. It is set up for transformers.
- Problem: The original 'tokenizerconfig.json' was missing the 'chattemplate' field. - Impact: Without a chat template, werving frameworks cannot convert(vLLM, TGI, SGLang). - Fix: Added the standard Qwen3 chat templa…
Downloads · 30 days
13
20% of all-time downloads
All-time downloads
65
Public
Parameters
8.2B
16.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.4 GB · 100%
From the Hugging Face model README
The 'reasoning_effort' parameter is designed to control how much "thinking" a model does before responding. However, this Qwen3-8B model only supports a binary toggle ('enable_thinking=True/False') via its chat template, not a continuous effort scale.
The 'reasoning_effort' parameter is an API-level concept that requires explicit support in both the serving engine and the model. Qwen3-8B was not trained to modulate reasoning depth based on an effort parameter — it either thinks (generates <think>...</think> blocks) or it doesn't.
Serving Engine Support: The inference API must parse the 'reasoning_effort' parameter and translate it into actionable behavior. Currently, most engines pass it through without interpretation for Qwen3 models.
Mapping to existing mechanisms: The engine could map 'reasoning_effort' to Qwen3's existing controls:
Thinking budget control: The serving engine needs to implement a token budget for the <think> block
Model level training: For 'reasoning_effort' to be truly meaningful (not just truncating thoughts), the model would need to be post-trained with varying levels of chain-of-thought depth: