Downloads · 30 days
20
17% of all-time downloads
DESUCLUB/Qwen3-NoThinkEmbed
Qwen3-NoThinkEmbed is a text generation model from DESUCLUB. Use it when you need the model to write or continue text. It is set up for transformers.
This model is based on Qwen3, and is an iterative process and set of experiments to try to remove thinking mode from Qwen3 architecturally, instead of providing <think/n/n</think tokens.
Downloads · 30 days
20
17% of all-time downloads
All-time downloads
119
Public
Parameters
4.4B
8.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.8 GB · 100%
From the Hugging Face model README
This model is based on Qwen3, and is an iterative process and set of experiments to try to remove thinking mode from Qwen3 architecturally, instead of providing <think>/n/n</think> tokens.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "DESUCLUB/Qwen3-NoThinkEmbed"
# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
# prepare the model input
prompt = "Give me a short introduction to large language model."
messages = [
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=True # Switches between thinking and non-thinking modes. Default is True.
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
# conduct text completion
generated_ids = model.generate(
**model_inputs,
max_new_tokens=32768
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()
content = tokenizer.decode(output_ids, skip_special_tokens=True).strip("\n")
print("content:", content)
The code used for reproducing this model can also be found in this repo, under think_remover.py
Credit goes to the Qwen Team for developing the Qwen3 suite of models, as well as providing the baseline for the inference code above