Downloads · 30 days
210
3% of all-time downloads
GetSoloTech/Qwen3-Code-Reasoning-4B-GGUF
Qwen3-Code-Reasoning-4B-GGUF is a text generation model from GetSoloTech. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
This is the GGUF quantized version of the Qwen3-Code-Reasoning-4B model, specifically optimized for competitive programming and code reasoning tasks. This model has been trained on the high-quality Code-Reasoning data…
Downloads · 30 days
210
3% of all-time downloads
All-time downloads
6.4K
Public
Repo size
23.1 GB
Likes
9
Public
Click a slice to open those files.
.gguf23.1 GB · 100%
From the Hugging Face model README
This is the GGUF quantized version of the Qwen3-Code-Reasoning-4B model, specifically optimized for competitive programming and code reasoning tasks. This model has been trained on the high-quality Code-Reasoning dataset to enhance its capabilities in solving complex programming problems with detailed reasoning.
# Download the model (choose your preferred quantization)
wget https://huggingface.co/GetSoloTech/Qwen3-Code-Reasoning-4B-GGUF/resolve/main/qwen3-code-reasoning-4b.Q4_K_M.gguf
# Run inference
./llama.cpp -m qwen3-code-reasoning-4b.Q4_K_M.gguf -n 4096 --repeat_penalty 1.1 -p "You are an expert competitive programmer. Read the problem and produce a correct, efficient solution. Include reasoning if helpful.\n\nProblem: Your programming problem here..."
from llama_cpp import Llama
# Load the model
llm = Llama(
model_path="./qwen3-code-reasoning-4b.Q4_K_M.gguf",
n_ctx=4096,
n_threads=4
)
# Prepare input for competitive programming problem
prompt = """You are an expert competitive programmer. Read the problem and produce a correct, efficient solution. Include reasoning if helpful.
Problem: Your programming problem here..."""
# Generate solution
output = llm(
prompt,
max_tokens=4096,
temperature=0.7,
top_p=0.8,
top_k=20,
repeat_penalty=1.1
)
print(output['choices'][0]['text'])
# Create a Modelfile
cat > Modelfile << EOF
FROM ./qwen3-code-reasoning-4b.Q4_K_M.gguf
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>
{{ end }}<|im_start|>assistant
"""
PARAMETER temperature 0.7
PARAMETER top_p 0.8
PARAMETER top_k 20
PARAMETER repeat_penalty 1.1
EOF
# Create and run the model
ollama create qwen3-code-reasoning -f Modelfile
ollama run qwen3-code-reasoning "Solve this competitive programming problem: [your problem here]"
| Quantization | Size | Memory Usage | Quality | Use Case |
|---|---|---|---|---|
| Q3_K_M | 2.08 GB | ~3 GB | Good | CPU inference, limited memory |
| Q4_K_M | 2.5 GB | ~4 GB | Better | Balanced performance/memory |
| Q5_K_M | 2.89 GB | ~5 GB | Very Good | High quality, moderate memory |
| Q6_K | 3.31 GB | ~6 GB | Excellent | High quality, more memory |
| Q8_0 | 4.28 GB | ~8 GB | Best | Maximum quality, high memory |
| F16 | 8.05 GB | ~16 GB | Original | Maximum quality, GPU recommended |
This GGUF quantized model maintains the performance characteristics of the original finetuned model:
This GGUF model was converted from the original LoRA-finetuned model. For questions about:
This model follows the same license as the base model (Apache 2.0). Please refer to the base model license for details.
For questions about this GGUF model, please open an issue in the repository.
Note: This model is specifically optimized for competitive programming and code reasoning tasks. The GGUF format enables efficient inference on various hardware configurations while maintaining the model's reasoning capabilities.