Downloads · 30 days
58
36% of all-time downloads
Comulative/KAT-Coder-V2.5_JKL-Luau-NVFP4
KAT-Coder-V2.5_JKL-Luau-NVFP4 is a text generation model from Comulative. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Roblox Luau–specialized fine-tune of Kwaipilot/KAT-Coder-V2.5-Dev, released as NVFP4 (W4A4 compressed-tensors) for efficient inference on NVIDIA Blackwell GPUs.
Downloads · 30 days
58
36% of all-time downloads
All-time downloads
159
Public
Parameters
35.1B
23.4 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors23.4 GB · 100%
How the weights are stored.
U832.6B · 93%
From the Hugging Face model README
Roblox Luau–specialized fine-tune of Kwaipilot/KAT-Coder-V2.5-Dev, released as NVFP4 (W4A4 compressed-tensors) for efficient inference on NVIDIA Blackwell GPUs.
| Base model | Kwaipilot/KAT-Coder-V2.5-Dev (~35B-A3B Qwen3.5-MoE, text/language release) |
| Fine-tune | Supervised LoRA SFT on a Roblox / Luau mix (1 epoch) |
| This artifact | Merged bf16 weights → NVFP4 post-training quantization |
| Trained by | @dylanjkl at Comulative Limited (UK) |
| Hardware | 1× NVIDIA RTX PRO 6000 Blackwell 96GB (sm_120) |
This is not an official Kwaipilot release. It is an independent specialty fine-tune and quant by Comulative Limited.
Primary: local / self-hosted Roblox Luau coding assistant and task executor:
--!strict, services, remotes, DataStore patterns)Recommended deployment pattern: use a stronger planning / review model for architecture and security, and this model as a fast local executor for Luau implementation (low latency, private weights, Blackwell-friendly NVFP4).
Not intended for: unsupervised production game economy / anti-cheat design without human review; non-Roblox general agenting as a drop-in frontier replacement; vision/multimodal tasks (language-only lineage).
Started from KAT-Coder-V2.5-Dev (Qwen3.5 MoE coding model, language-only open weights; vision tower declared in config but not shipped upstream).
Built a chat-formatted SFT mix (42,302 train rows after filtering/dedup; ~57M tokens) from public Hugging Face datasets (local mirror under data/).
| Hub dataset | Role in mix | Approx. SFT rows (source tag) |
|---|---|---|
| TorpedoSoftware/Roblox-Luau-Reasoning-v1.0 | Prompt → CoT + Luau code + explanation (train) | 14,840 luau-reasoning |
| Pinkstack/luaucoder-instructions-v3-SFT | Instruction SFT (filtered; cap ~8k quality rows) | 7,927 pinkstack-sft |
| khtsly/luau-stack-hq | Curated Luau corpus → fill-in / continuation tasks | 5,944 stackhq-completion |
| TorpedoSoftware/RobloxQA-v2.0 | Engine/API knowledge from train (MCQ + direct QA variants) | 4,573 robloxqa2-mcq + 2,290 robloxqa2-direct |
| TorpedoSoftware/LuauLeetcode | Algorithmic Luau problems (train) | 2,336 luau-leetcode |
| TorpedoSoftware/RobloxQA-v1.0 | Older QA; deduped against RobloxQA-v2 test | 2,281 robloxqa1-mcq |
| Roblox/luau_corpus | Official Luau Data Sharing fragments → continuation (train) | 2,111 luaucorpus-completion |
Total train examples: 42,302 (plus 400 held-out mix rows for training-time val).
| Hub dataset | Use |
|---|---|
TorpedoSoftware/RobloxQA-v2.0 test (3,000 MCQ) | Held out for baseline / future eval only |
| Hub dataset | Notes |
|---|---|
| TorpedoSoftware/roblox-info-dump | Roblox/Luau docs scrape present under data/; not mixed into the epoch-1 SFT JSONL |
Formatting used the base model chat template, with prompt tokens masked (train on completions). Optional system prompts mixed in (~30% Roblox-assistant style). Decompiled-looking completion snippets filtered out. Pinkstack rows required code fences + minimum reasoning length.
| Hyperparameter | Value |
|---|---|
| Method | LoRA (PEFT), bf16 base |
| Rank / alpha | r=64, α=128, dropout 0.05 |
| Targets | Attention (q/k/v/o), linear-attn projections, shared-expert MLP; not per-routed experts |
| Context | 4096 |
| Effective batch | 16 (microbatch 2 × grad accum 8) |
| Epochs shipped here | 1 (stopped at step 2642 / 5284 of a 2-epoch schedule) |
| Optim | AdamW fused, LR 1e-4 cosine, warmup 40 |
| Hardware | 1× RTX PRO 6000 96GB, Windows 11 + WSL2 Ubuntu |
A full Trainer checkpoint (adapter + optimizer) was frozen at epoch 1 so a second epoch can be resumed later if desired. This Hub repo is the merged + NVFP4 product of that epoch-1 adapter, not the raw LoRA.
visual.* keys so the checkpoint matches the language-only upstream releasenvfp4-pack-quantized)moe_calibrate_all_experts=Truelm_head, visual towers, router gates, embeddings, linear-attn (see recipe.yaml)Why NVFP4: native-friendly 4-bit float path for Blackwell inference stacks (e.g. vLLM on modern NVIDIA data-center / pro GPUs). Smaller footprint (~22GB weights here) vs full bf16 (~70GB class).
Unmodified base KAT-Coder-V2.5-Dev on RobloxQA-v2.0 test (3000 questions), MMLU-style log-prob forced choice, bf16, HF Transformers:
Baseline: 87.60% (2628 / 3000) — measured 2026-08-01.
Post–fine-tune / post-NVFP4 RobloxQA numbers for this checkpoint may be published later; treat the above as the starting point the specialty run was designed to preserve or beat on knowledge while improving Luau generation.
Follow current llm-compressor / Transformers docs for NVFP4 compressed-tensors checkpoints. Ensure a stack that understands quantization_config with format nvfp4-pack-quantized.
Prefer a recent vLLM build with NVFP4 + Qwen3.5 MoE support. Language-only serving flags may still apply (upstream was language-only; config may still mention vision). Example pattern (versions change quickly—check your vLLM release notes):
vllm serve Comulative/KAT-Coder-V2.5_JKL-Luau-NVFP4 \
--quantization modelopt_fp4 \
# plus whatever your vLLM build requires for NVFP4 / Qwen3.5 MoE
Use the model’s chat template (chat_template.jinja / tokenizer config). For agentic coding UIs, Qwen-style reasoning / tool parsers may apply depending on runtime.
--!strict Luau when you want typed modulesBase: Kwaipilot/KAT-Coder-V2.5-Dev (Qwen3.5-MoE ~35B-A3B)
SFT: LoRA r=64, 1 epoch, Luau/Roblox mix ~42k examples
Merge: bf16 merge + strip visual.*
Quant: NVFP4 (llm-compressor), 256 calib samples, all MoE experts
Trainer: @dylanjkl / Comulative Limited (UK)
GPU: 1× NVIDIA RTX PRO 6000 Blackwell 96GB
Weights are a derivative of Kwaipilot/KAT-Coder-V2.5-Dev. Unless otherwise required by the base model license, this distribution is provided under Apache-2.0. Review the base model card and license for any additional terms.
Base model: Kwaipilot/KAT-Coder-V2.5-Dev
Fine-tune & NVFP4 release: Comulative Limited (UK) / @dylanjkl
Hub: Comulative/KAT-Coder-V2.5_JKL-Luau-NVFP4
Hardware: 1× NVIDIA RTX PRO 6000 Blackwell 96GB
For issues with this fine-tune/quant, contact the Comulative maintainers—not Kwaipilot—unless the bug is reproducible on the unmodified base.