Downloads · 30 days
257
37% of all-time downloads
exalandru/GPT-OSS-Coder-MLX
GPT-OSS-Coder-MLX is a text generation model from exalandru. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
A gpt-oss-120b version focused on improving practical coding-agent behavior in repository-level software engineering tasks.
Downloads · 30 days
257
37% of all-time downloads
All-time downloads
695
Public
Parameters
117B
83.3 GB on disk
Likes
1
Trending 1
Click a slice to open those files.
.safetensors65.2 GB · 100%
How the weights are stored.
U32115B · 98%
From the Hugging Face model README
A gpt-oss-120b version focused on improving practical coding-agent behavior in repository-level software engineering tasks.
Also available in GGUF for llama.cpp / LM Studio / Ollama →
This release adds three focused passes over re-curated data, each targeting a behaviour measured missing in the previous one: writing failure-injection regression tests before closing, iterating on multi-defect repositories past the first green suite, and honouring per-harness completion contracts. Every pass was gated by calibrated probes and a held-out benchmark ladder before being promoted.
It digs deeper into the repository, follows evidence to the root cause, and keeps iterating until the fix holds under real tests instead of stopping at a plausible-looking patch.
attempt_completion callThe fine-tune also significantly reduced malformed JSON arguments.
Works even better with my Adversarial Agent Engineering pack of skills and rules
<span style="color:orange;">Optimized for Codex.</span>

It provides native Codex integration for GPT-OSS on MLX. Rather than exposing GPT-OSS only through a generic OpenAI-compatible compatibility layer, it is designed so that Codex can use GPT-OSS as a native local model while preserving the GPT-OSS/Codex protocol details.
This includes the native Codex Responses protocol, GPT-OSS Harmony handling, reasoning continuity across tool turns, and Codex-specific routing and metadata.
Agent loops work smoothly without the stalls and rejections you get with generic OpenAI-compatible endpoints.
https://github.com/exalandru/Codex-GPT-OSS-Server
mlx_lm.generate --model exalandru/GPT-OSS-Coder-MLX --prompt "Hello World!"
Format: MLX, MXFP4 experts + bf16 attention, ~61 GB on disk. Weights are consolidated — nothing to fuse or merge. Runs on Apple silicon with 96 GB unified memory; a full agent session peaks around 82 GB.
116.8B parameters (MoE, ~5.1B active per token — which is why it decodes faster than much smaller dense models). The Hub badge shows ~24B because MLX packs eight 4-bit weights per uint32 element and the badge counts elements, not parameters.
Supervised fine-tuning on ~10 000 steps carefully selected from real coding-agent sessions to isolate the targeted behavior : some of my personal sessions with Opus/Fable 5 and GPT 5.6 Sol, public SWE-agent, OpenHands, SWE-smith and Fable trajectories, keeping only runs that actually resolved their issue. A run that gave up, or ran out of context and submitted anyway, teaches exactly the habit this model is meant to shed, so those were filtered out.
Each training example is a real repository state plus the next action the successful agent took, so what is learned is the loop itself: look, edit, run, read the result, correct.
The fine-tune itself is deliberately small, a low-rank update on the last layers only, then consolidated back into the weights. The goal was to shift behaviour, not to overwrite what the base model already knows.
Custom small in-house benchmarks were used to validate the training. Models such as Qwen3.6, DeepSeek v4 Flash and other distilled gpt-oss variants all failed these benchmarks. Opus 5 and GPT 5.6 Sol served as references proving the tasks were solvable.
Built by exalandru. If you use it in a real agent loop, the failure reports are more useful than the success ones.