Downloads · 30 days
594
11% of all-time downloads
guell00/Nexora-Qwen-Coder-4B
Nexora-Qwen-Coder-4B is a machine learning model from guell00. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Downloads · 30 days
594
11% of all-time downloads
All-time downloads
5.2K
Public
Repo size
22.9 GB
Likes
4
Public
Click a slice to open those files.
.gguf10.4 GB · 100%
From the Hugging Face model README
Coding · Debugging · Tool Use · Structured Reasoning · Local AI
<br> </div> <br> <p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/644afe169279988e0cbcd2d9/AVI5Rk73jQfx76ePPXx90.png" width="100%"> </p>Nexora-Qwen-Coder-4B is a compact, coding-focused language model fine-tuned from the Qwen 3.5 4B family, with an emphasis on code generation, debugging, structured reasoning, tool use, and local agentic workflows.
The core idea behind Nexora is simple:
A small coding model should do more than generate code. It should understand the task, reason through problems, interact with tools, inspect feedback, and iterate toward a solution.
Designed for developers who want capable AI assistance without requiring a large datacenter-scale deployment.
| Property | Details |
|---|---|
| Model | Nexora-Qwen-Coder-4B |
| Hugging Face | guell00/Nexora-Qwen-Coder-4B |
| Base Model | Qwen 3.5 4B |
| Architecture | Dense Transformer |
| Model Class | 4B Parameters |
| Primary Focus | Coding & Agentic Workflows |
| Fine-Tuning | Nexora Fine-Tuning |
| Training Method | SFT + Curriculum Learning |
| Reasoning Data | Trace Inversion |
| Agent Data | Tool-Use & Agent Trajectories |
| Training Context | Up to 32K tokens |
| Evaluation | MTP, n=2 |
| Format | GGUF |
| Inference | llama.cpp compatible |
Nexora-Qwen-Coder-4B is not designed around parameter count alone.
The objective is to make a compact local model more useful for real software development workflows.
The fine-tuning strategy focuses on four primary capabilities:
01 · CodingGenerate, complete, explain, refactor, and implement code across common programming tasks.
02 · DebuggingAnalyze errors, identify potential failure points, reason about bugs, and produce targeted fixes.
03 · Agentic WorkflowsOperate in environments where the model can inspect files, select tools, execute actions, receive feedback, and iterate.
04 · Structured ReasoningHandle multi-step technical tasks that benefit from planning, decomposition, and structured problem solving.
For the best balance of consistency, coding quality, and controlled generation, the recommended default configuration is:
| Parameter | Recommended |
|---|---|
| Temperature | 0.1 |
| Top P | 0.95 |
| Top K | 52 |
| Min P | 0.05 |
| Repetition Penalty | 1.1 |
| Presence Penalty | Off |
Temperature: 0.1
Top P: 0.95
Top K: 52
Min P: 0.05
Repetition Penalty: 1.1
Presence Penalty: Off
This configuration is recommended for:
The low Temperature is intended to improve consistency while preserving a small amount of generation flexibility.
Note: Evaluation results may vary when using sampling parameters different from those used during benchmarking.
Nexora-Qwen-Coder-4B was evaluated locally using the benchlocal evaluation framework.
The evaluation focuses primarily on practical developer workflows, including debugging, tool use, agent behavior, and instruction following.
| Benchmark | Nexora-Qwen-Coder-4B | Qwen 3.5 4B | Delta |
|---|---|---|---|
| BugFind-15 | 71 / 100 | 52 / 100 | +19 |
| HermesAgent-20 | 64 / 100 | 61 / 100 | +3 |
| ToolCall-15 | 100 / 100 | 90 / 100 | +10 |
| InstructFollow-15 | 93 / 100 | 93 / 100 | 0 |
BugFind-15
Nexora-Qwen-Coder-4B ██████████████░░░░░░ 71
Qwen 3.5 4B ██████████░░░░░░░░░░ 52
HermesAgent-20
Nexora-Qwen-Coder-4B █████████████░░░░░░░ 64
Qwen 3.5 4B ████████████░░░░░░░░ 61
ToolCall-15
Nexora-Qwen-Coder-4B ████████████████████ 100
Qwen 3.5 4B ██████████████████░░ 90
InstructFollow-15
Nexora-Qwen-Coder-4B ██████████████████░░ 93
Qwen 3.5 4B ██████████████████░░ 93
The strongest observed results were in:
These results suggest that the fine-tuning process improved the model's performance on targeted coding and agentic tasks compared with the base evaluation reference.
Benchmark results are snapshots from specific evaluation runs. They should not be interpreted as universal performance guarantees.
Nexora-Qwen-Coder-4B is designed for workflows where the model can interact with an external environment rather than simply returning a single static answer.
A typical agent loop can be represented as:
┌──────────────────┐
│ User Request │
└────────┬─────────┘
│
▼
┌──────────────────┐
│ Understand Task │
└────────┬─────────┘
│
▼
┌──────────────────┐
│ Plan Solution │
└────────┬─────────┘
│
▼
┌──────────────────┐
│ Select Tool │
└────────┬─────────┘
│
▼
┌──────────────────┐
│ Execute Action │
└────────┬─────────┘
│
▼
┌──────────────────┐
│ Inspect Feedback │
└────────┬─────────┘
│
▼
Success?
╱ ╲
Yes No
│ │
▼ │
┌───────────┐ │
│ Final │ │
│ Answer │ │
└───────────┘ │
│
└──────► Iterate
This makes the model suitable for local environments that expose tools such as:
Typical agent workflows may include:
Read → Plan → Act → Observe → Verify → Repair
Tool-call reliability depends on the application's prompt template, tool definitions, schema design, and execution environment.
| Use Case | Fit |
|---|---|
| Code Generation | ★★★★★ |
| Debugging | ★★★★★ |
| Tool Calling | ★★★★★ |
| Local Coding Agents | ★★★★★ |
| Code Explanation | ★★★★★ |
| Refactoring | ★★★★☆ |
| Repository Analysis | ★★★★☆ |
| Technical Reasoning | ★★★★☆ |
| Documentation | ★★★★☆ |
| Software Architecture | ★★★☆☆ |
The 4B parameter class is intentionally compact.
Nexora-Qwen-Coder-4B aims to provide a practical balance between:
CAPABILITY
▲
│
│ ● Nexora-Qwen-Coder-4B
│
│
│
└────────────────────────►
LOCAL EFFICIENCY
The goal is straightforward:
Deliver useful coding and agentic capabilities while remaining practical to run locally.
Potential deployment scenarios include:
Nexora-Qwen-Coder-4B is available in GGUF format for efficient local inference.
| Quantization | Recommended For |
|---|---|
| Q4_K_M | Best balance of quality, memory, and speed |
| Q8_0 | Higher quantized quality with increased memory usage |
Q4_K_MFor most users, Q4_K_M provides a strong balance between:
Quality · Memory · Speed
Q8_0Recommended when memory usage is less restrictive and higher quantized fidelity is preferred.
Run the model directly from Hugging Face:
llama-cli -hf guell00/Nexora-Qwen-Coder-4B --jinja
Start an OpenAI-compatible local server:
llama-server -hf guell00/Nexora-Qwen-Coder-4B --jinja
Command availability may depend on your installed
llama.cppversion and the model files available in the repository.
The model was fine-tuned using sequences reaching approximately 32K tokens.
The underlying Qwen 3.5 family may support larger context windows depending on the specific architecture and inference backend.
Long-context performance depends on:
When extending beyond the training distribution, users should validate performance on their own workloads.
./llama-server \
-m model.gguf \
--ctx-size 131072 \
--rope-scaling yarn \
--rope-scale 4 \
--yarn-orig-ctx 32768
Important: Increasing
--ctx-sizealone does not guarantee reliable long-context behavior.
For highly deterministic coding, debugging, and code-repair workflows:
| Parameter | Value |
|---|---|
| Temperature | 0 |
| Top P | 0.95 |
| Top K | 40 |
| Min P | 0.05 |
| Repetition Penalty | 1.1 |
| Presence Penalty | Off |
| Max Tokens | Max |
Temperature: 0
Top P: 0.95
Top K: 40
Min P: 0.05
Repetition Penalty: 1.1
Presence Penalty: Off
Max Tokens: Max
For creative programming, brainstorming, or exploratory generation, increasing the temperature may produce more diverse outputs.
For debugging and code repair, lower temperatures generally provide more deterministic results.
| Technology | Role |
|---|---|
| Qwen | Base model family |
| Unsloth | Fine-tuning & conversion workflows |
| GGUF | Efficient local model format |
| llama.cpp | Local inference |
| benchlocal | Coding & agent evaluation |
Nexora-Qwen-Coder-4B is a compact 4B-class model and should be evaluated accordingly.
It may struggle with:
The model should be treated as a coding assistant, not a fully autonomous software engineer.
Generated code should always be:
REVIEWED
↓
TESTED
↓
VALIDATED
↓
DEPLOYED
Applications should verify generated code before using it in production environments.
Depending on the inference template and runtime configuration, the model may generate reasoning content inside:
<think>
...
</think>
Applications may parse, hide, or otherwise handle these sections according to their requirements.
Special thanks to:
This model is released under the MIT License.
Please review the licensing terms of the underlying base model and any third-party components used in your deployment.
Nexora-Qwen-Coder-4B is provided for:
Research · Development · Experimentation · Local Inference
Actual performance may vary depending on:
Benchmark results represent specific evaluation runs and should not be interpreted as guaranteed performance across all environments or tasks.
Always review, test, and validate generated code before deploying it to production systems.
Built for developers who want capable AI coding assistance running locally.
<br>guell00/Nexora-Qwen-Coder-4B