Downloads · 30 days
144
39% of all-time downloads
KBBridge/KBBridge-v3-GGUF
KBBridge-v3-GGUF is a text generation model from KBBridge. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
A fine-tune of Qwen/Qwen3.8-27B specialised in GeneXus programming, in the native .gxSource export format.
Downloads · 30 days
144
39% of all-time downloads
All-time downloads
368
Public
Repo size
91.7 GB
Likes
1
Public
Click a slice to open those files.
.gguf45.9 GB · 100%
From the Hugging Face model README
A fine-tune of Qwen/Qwen3.8-27B specialised in
GeneXus programming, in the native .gxSource export format.
Frontier models do not know this format. Without the GeneXus documentation injected into the prompt they produce syntactically invalid output almost every time (parse rate 0.5–3.1%). KBBridge writes it natively, runs on your own hardware, and never sends your Knowledge Base code to an external API.
Three settings. All three are measured on this model, not stylistic.
These GGUFs no longer pin a value, so reasoning follows the upstream Qwen default (on) unless you turn it off. Which one you want depends on what you are asking for. All three rows are measured on this model, not inherited from Qwen.
| Task | Reasoning | Measured |
|---|---|---|
Writing .gxSource | OFF | parseRate 86.9 → 33.5, parmMatch 80.2 → 50.4 (191 held-out objects) |
| Documentation multiple-choice | either | 78.4 vs 78.4 — no difference (329 items, McNemar p = 1.000) |
| Explaining existing code | ON | fabricated claims 15.4% → 8.7% (149 items, McNemar p = 0.041) |
If you generate GeneXus objects, turn reasoning off. The collapse is real, not a budget
artifact: with reasoning on, only 5.2% of items hit the token ceiling (fewer than the 8.4%
without it) and 3.7% came back empty. The model simply writes worse .gxSource when it
reasons first.
If you point the model at existing code and ask what it does, turn reasoning on. It nearly halves the rate at which the model asserts things the source does not support. Cost: ~5× the output tokens.
Correction (2026-09-02). An earlier version of this card reported MCQ dropping 78.1 → 69.6 with reasoning on, and shipped GGUFs with
enable_thinkingpinned tofalse. That MCQ number was wrong: our benchmark capped multiple-choice answers at 512 tokens, not enough for a reasoning block to close, so the run measured the cap rather than the model. Re-measured with an adequate budget the difference is zero. The.gxSourcedegradation is real and is reproduced above. The pin was removed so that the choice is yours.
llama.cpp. --chat-template-kwargs '{"enable_thinking":false}' is deprecated, and
--reasoning defaults to auto, which reads support from the template and turns reasoning
on. Set it explicitly:
llama-server -m KBBridge-v3-Q4_K_M.gguf --jinja --reasoning none
LM Studio. It ignores chat_template_kwargs sent through the API. Use the
"thinking"/"reasoning" toggle in the model settings instead — the API parameter will not reach
the template. We once watched a benchmark score 24.1 instead of 90.0 for exactly this reason.
Any runtime, guaranteed. Edit the template and add this as its first line:
{%- set enable_thinking = false %}
A set at the top overrides anything the caller passes. It is blunt, but it is the only
approach that behaves identically across llama.cpp, LM Studio and Ollama — which is why these
files shipped that way until 2026-09-02.
Write "in .gxSource format" in your prompt.
Measured on v3: the bare request "a Procedure that adds two numbers" returns generic SQL. Naming the format returns the GeneXus object, consistently. If you use a harness with its own system prompt, put the instruction there once.
max_tokens ≥ 4096. A .gxSource object consumes roughly 340 tokens per KB of source, and
most tools default to 512–1024, which truncates the object mid-body.
580 held-out items (191 codegen + 329 MCQ + 60 data-model) that no model saw during training. Syntax validated with the official GeneXus ANTLR parser. Same protocol for every model: temperature 0.1, reasoning off, concurrency 8.
v3 is not a clean win over v2. It gains domain knowledge and loses syntax accuracy:
| Metric | v2 | v3 | |
|---|---|---|---|
| parseRate (valid syntax) | 89.0 | 84.8 | −4.2 |
| parmMatch (exact signature) | 78.6 | 78.6 | = |
| MCQ (GeneXus knowledge) | 76.0 | 79.0 | +3.0 |
| methodValidity | 90.0 | 91.1 | +1.1 |
What these numbers do NOT establish. v3 changed three things at once — the base model (Qwen3.6 → 3.8), the corpus (4× larger, per-KB cap removed) and the teacher (v1 → v2). The parseRate drop cannot be attributed to any one of them without a control arm that was never run. Anyone reading this table as "the bigger corpus hurt syntax" is over-reading it.
Choose v3 if domain knowledge matters more to you; v2 still leads on raw syntax validity.
Three entire KBs were held out — different domains, never in the pipeline:
| held-out from training KBs | 3 completely new KBs | |
|---|---|---|
| v2 | 89.0 | 89.9 |
| v3 | 84.8 | 87.4 |
v3's relative gap to unseen KBs is larger than v2's (+2.6 vs +0.9), i.e. it generalises better in relative terms, even though two KBs make up 54.7% of its corpus.
In our benchmark the frontier models were run with ~21,600 tokens of GeneXus documentation injected into every request; KBBridge was run without any. That is not a handicap we imposed — injecting the same documentation into KBBridge makes it worse (76.4 → 73.3 parseRate), because the fine-tune already internalised that knowledge and the extra context gets in the way. Still, the setups differ, and you should know that when reading any head-to-head number.
Same 580-item benchmark, run against KBBridge-v3-Q4_K_M.gguf in llama.cpp with all
layers on GPU:
| Metric | v3 bf16 (vLLM) | v3 Q4_K_M (llama.cpp) |
|---|---|---|
| parseRate | 84.8 | 90.0 |
| parmMatch | 78.6 | 80.9 |
| methodValidity | 91.1 | 92.2 |
| MCQ | 79.0 | 78.1 |
The Q4 scoring higher than the full-precision model is not a quantisation benefit, and we are not going to pretend otherwise. Here is what is actually going on.
On items both runs completed normally, the two are identical. Excluding every item where either run hit the 20,000-token ceiling (20 items), the remaining 171 give:
| parseRate | parmMatch | methodValidity | |
|---|---|---|---|
| bf16 | 93.0 | 85.6 | 93.0 |
| Q4_K_M | 93.6 | 85.6 | 93.0 |
A 0.6-point gap on parseRate is one item out of 171. 4-bit quantisation costs essentially nothing in output quality. The MCQ drop (−0.9) is the only measurable degradation.
The whole headline difference is the runaway rate. Degenerate repetition until the token budget is exhausted hit 19 of 191 items (9.9%) under vLLM and 7 of 191 (3.7%) in llama.cpp. Of the 19 the bf16 run ruined, the Q4 run closed 13 cleanly; one item went the other way.
And we cannot attribute that to the quantisation. The two runs differ in the inference
engine and in the default sampler stack — llama.cpp applies top_k=20, top_p=0.95,
min_p=0.05; vLLM applies none of them, and the benchmark only sets temperature. Truncating
the low-probability tail is a plausible mechanism for suppressing repetition loops, and it is
confounded with the quantisation in this measurement. Isolating it would need a controlled run
we have not done.
What this means for you, practically: the numbers in the Q4 column are what you should
expect from this file in llama.cpp or LM Studio with stock settings — that is the configuration
we measured. If you disable min_p/top_k to match a vLLM-style setup, expect more runaway on
large objects.
| File | Size | For |
|---|---|---|
KBBridge-v3-Q4_K_M.gguf | 16 GB | LM Studio, llama.cpp, Ollama — the default choice |
KBBridge-v3-Q8_0.gguf | 28 GB | higher fidelity, if you have the VRAM |
# llama.cpp, downloads on demand
llama-server -hf KBBridge/KBBridge-v3-GGUF:Q4_K_M --jinja
# or fetch the file directly
hf download KBBridge/KBBridge-v3-GGUF KBBridge-v3-Q4_K_M.gguf --local-dir .
In LM Studio, search for KBBridge/KBBridge-v3-GGUF.
Verify your download against SHA256SUMS.
The base model is multimodal and KBBridge inherits its vision tower unchanged (verified:
identical tensors). These GGUFs are text-only; if you want image input, pair them with the
official projector from
Qwen/Qwen3.8-27B via --mmproj. For writing
GeneXus you do not need it.
The multi-token-prediction head is included in these files. Enable it:
llama-server -m KBBridge-v3-Q4_K_M.gguf --spec-type draft-mtp --jinja
Measured on one RTX PRO 6000 (Q4_K_M, all layers on GPU, 301-token generations, first run discarded):
| median | range | |
|---|---|---|
| without MTP | 62.1 tok/s | 61.6 – 65.2 |
| with MTP | 109.0 tok/s | 93.7 – 117.6 |
+71%, and the output is unchanged — in speculative decoding the main model verifies every drafted token, so a draft head can only affect speed, never correctness.
Across the 583 requests of the full benchmark run — real GeneXus generation, not a microbenchmark — the draft head's tokens were accepted at a median rate of 0.86, averaging 3.28 accepted tokens per speculative step. It predicts the fine-tuned model's output well despite coming from the base model untouched by the fine-tune.
Assisting GeneXus developers: generating objects (Procedures, Transactions, Data Providers, SDTs, WebPanels), explaining existing code, completion, and documentation questions.
Out of scope: not a general-purpose model, not a replacement for validating in the GeneXus IDE, and it does not know any particular Knowledge Base (see Limitations).
max_tokens does not
fix it. Generate large objects section by section.The raw GGUF and a gateway-fronted deployment do not behave the same by default. Our
gateway applies five corrections the plain model does not have: a max_tokens floor,
reasoning off unless the client asks for it, temperature defaulted to 0.2 (without it vLLM
falls back to the checkpoint's generation_config, which is 1.0), a fallback that recovers
the answer from the reasoning field when content comes back empty, and repetition_penalty
1.05 to suppress runaway. If you compare "what I tried on your server" against "what I
downloaded", the difference is those five settings, not the weights.
The temperature one surprises people: the OpenAI standard makes the field optional and many clients never send it, so an unconfigured client is sampling at 1.0 without being told.
| Method | QLoRA 4-bit (bitsandbytes) + Liger kernel |
| LoRA | r=64, α=128, dropout=0.05, all projections |
| Context | 12,288 tokens |
| Effective batch | 16 (1 × 16 grad accum) |
| LR | 1.0e-4, cosine, 3% warmup |
| Epochs | 2 complete (14,108 steps) |
| Hardware | 1× RTX PRO 6000 Blackwell 96 GB |
| Duration | 7 days 4:41 |
| Framework | LLaMA-Factory, transformers 5.6.0 |
train_loss 0.2618 (v2: 0.3344) · eval_loss 0.3723 (v2: 0.4675), minimum at the last step — no overfitting across 71 evaluations, which suggests there was room for more epochs.
Note that these losses are much better than v2's and yet parseRate went down: eval_loss
measures fit to the corpus, not GeneXus quality.
80,344 examples derived from GeneXus objects across 25 real Knowledge Bases (GX16/17/17U8/18/ Evo1, multi-domain) — 129% more than v2, with the per-KB cap removed. Sanitised, deduplicated and split by deterministic hash. The datasets are not published: they contain customer proprietary code.
The model was trained on real customer Knowledge Bases, so we audited whether it can leak them. This is the strongest result of the project.
12 synthetic objects containing unguessable 16-character secrets were inserted at four
frequencies, and verified to have reached train.jsonl at exactly those counts:
| repetitions | canaries | recovered by name | recovered with literal prefix |
|---|---|---|---|
| 1 | 3 | 0/3 | 0/3 |
| 10 | 3 | 0/3 | 0/3 |
| 100 | 3 | 0/3 | 0/3 |
| 1000 | 3 | 0/3 | 0/3 |
Not even at a thousand identical repetitions. A control rules out a broken probe: asked for the canary, the model returns a structurally valid but empty object — no token, no secret. And it does generate real bodies when the request has content, so the empty skeleton is not an inability to generate.
| mean loss, seen examples | 3.4130 |
| mean loss, unseen | 3.7711 |
| mean length | 3,133 vs 3,117 chars — comparable, so the AUC is meaningful |
| AUC | 0.5539 |
0.554 against 0.50 for indistinguishable. There is a statistical trace of having seen the data, but the distributions overlap almost entirely.
Conclusion: customer code is not recoverable from the weights.
Caveat, stated plainly: absence of evidence is not proof of absence. These audits cover the attacks we ran, not every attack that exists.
Full external reproduction is not possible, and it is worth saying so directly:
parseRate scorer uses the KBEditor's ANTLR parser — proprietary, not distributable.What a third party can verify: the raw benchmark outputs (one model response per item) and the scoring over them.
@misc{kbbridge-v3,
title = {KBBridge-v3: a GeneXus code assistant fine-tuned from Qwen3.8-27B},
author = {{KBBridge}},
year = {2026},
url = {https://huggingface.co/KBBridge/KBBridge-v3-GGUF}
}
Apache 2.0, inherited from the base model Qwen/Qwen3.8-27B. This is a modified derivative
work; see NOTICE.