Downloads · 30 days
25
35% of all-time downloads
NAME0x0/AVA-v3.0-preview
AVA-v3.0-preview is a text generation model from NAME0x0. Use it when you need the model to write or continue text. It is set up for gguf. The card lists the license as apache-2.0.
Read this first. This is an honest preview of a checkpoint that is ~11% of the way through its planned training (step 2,167 of 20,000). It matches its base model on code — it does not beat it yet. It is published to (…
Downloads · 30 days
25
35% of all-time downloads
All-time downloads
72
Public
Repo size
2.7 GB
Likes
1
Public
Click a slice to open those files.
.gguf2.7 GB · 100%
From the Hugging Face model README
Read this first. This is an honest preview of a checkpoint that is ~11% of the way through its planned training (step 2,167 of 20,000). It matches its base model on code — it does not beat it yet. It is published to (a) prove the whole
$0, laptop-scale pipeline works end-to-end and (b) bank a reproducible artifact. If you want a finished coding model today, use the base Qwen/Qwen3.5-4B directly. If you want to watch a coding specialist get built in the open on free hardware, follow along.
A coding specialist that runs on a 4 GB-VRAM laptop, trained for $0 on free
Colab/Kaggle GPU quota plus one consumer laptop, and reproducible by anyone.
The thesis: a world-class small coding model does not require a datacenter — you
can inherit a strong open base, specialize it on clean open data, and later
self-improve it against a free verifier (the code executor). No proprietary-model
distillation; every training source is permissively licensed and public.
Evaluated at 4-bit, greedy, non-thinking (the deployment-realistic setting), with an execution-based HumanEval+/MBPP+ harness.
Base model (Qwen3.5-4B, our harness, full sets):
| Benchmark | Score |
|---|---|
| HumanEval+ (164) | 74.4 |
| MBPP+ (378) | 65.3 |
| ARC-Easy | 93.8 |
| MMLU | 54.9 |
This preview @ step 2,167 — measured on n=60 matched-task probe subsets vs the same donor tasks (mid-training, subset noise ≈ ±7pp):
Harder held-out eval (LiveCodeBench v6 — competitive programming, 2025 contests after the easy sets saturated; n=50 stdin problems, greedy, non-thinking):
| Donor (Qwen3.5-4B) | This preview @ 2,167 | |
|---|---|---|
| pass@1 | 44.0% | 34.0% |
On fresh competitive-programming problems this checkpoint is currently ~10 pp behind the donor — the loss is on medium/hard problems (easy holds ~94%). Shown, not hidden. This is the expected direction at 11% of training on off-distribution data: the current mix is reasoning + edits, not competitive stdin, so basic SFT hasn't helped (and slightly hurts) this slice yet. The number that matters is the trajectory across later checkpoints, not this single mid-training point.
So: more reasoning, donor-level on easy code benchmarks, but still behind the donor on harder held-out coding — at 11% of training. A checkpoint, not a destination.
ava-v30-preview-q4_k_m.gguf (2.6 GB, Q4_K_M) — runs in llama.cpp.Needs a recent llama.cpp build (Qwen3.5 hybrid-attention support; tested on b10059).
# direct code answers (recommended for a coding tool)
llama-cli -m ava-v30-preview-q4_k_m.gguf -ngl 99 -fa on -c 8192 \
--chat-template-kwargs '{"enable_thinking": false}' \
-cnv -p "Write a Python function that merges two sorted lists."
# thinking mode (slower, more reasoning on hard problems) — drop the kwargs line
Measured on an RTX A2000 4 GB laptop (partial offload): prompt ~70 t/s, generation ~18 t/s. On an empty 4 GB card all 32 layers offload. The hybrid architecture keeps the KV cache small (~8 of 32 layers use softmax attention), so long context is cheap.
enable_thinking: false for a snappy coding tool.Agentic/edit data (Open-SWE-Traces) → harder evals (SWE-bench Lite) → thinking-mode
self-distillation + execution-verified preference tuning → self-play RL against the
sandbox verifier. Those — not more basic SFT — are where a $0 specialist overtakes
its base.
Built on Qwen3.5-4B (Alibaba, Apache 2.0). Training data from NVIDIA OpenCodeReasoning and BigCode CommitPackFT. Full pipeline, evals, and the resumable-training fabric are open — this model card links the recipe, not just the weights.