Downloads · 30 days
85
25% of all-time downloads
badtheorylabs/BTL-3
BTL-3 is a text generation model from badtheorylabs. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
Downloads · 30 days
85
25% of all-time downloads
All-time downloads
338
Public
Repo size
954 MB
Likes
74
Public
Click a slice to open those files.
.safetensors934 MB · 98%
From the Hugging Face model README
95.1% HumanEval · 88.5% BFCL v4 AST · 88.1% LiveCodeBench v6 (193-case run)
Compact edition · Runtime source · Bad Theory Labs · Discord
</div>BTL-3 is a 27B open-weight agent model built for agentic coding, structural tool use, repository work, failure recovery, and long multi-turn execution.
The release includes the full BTL-3 model and BTL-3 Compact, which packages the complete text model into one 8.39 GB native file. That is smaller than an 8B model stored in FP16 and corresponds to an effective artifact footprint of under 2.5 bits per parameter. On a fresh private 100-turn tool-contract gate, BTL-3 Compact retained 83 of the 90 behaviors the full model completed correctly: 92.2% measured conditional tool-behavior retention.
BTL-3 is a post-trained Qwen3.6-27B model for coding agents, repository work, structured tool use, and long multi-turn execution. It is tuned to reason, act, inspect tool results, recover from failures, and stop when no action is required.
This repository contains the frozen RL-0013 rank-32 PEFT adapter and its tokenizer configuration. The base checkpoint is pinned to an exact revision for reproducible loading.
All values below belong to the frozen BTL-3 RL-0013 release.
| Evaluation | Score | Protocol |
|---|---|---|
| BFCL v4 AST | 88.5% (1097/1240) | Complete official full set |
| HumanEval | 95.12% (156/164) | pass@1, thinking mode |
| LiveCodeBench v6 | 88.1% (170/193) | Completed 193-case run, thinking mode |
| BigCodeBench-Hard Instruct | 26.35% (39/148) | Official strict pass@1 |
| BigCodeBench functional tests | 59.25% (506/854) | Supplementary test-level score |
| Category | Score |
|---|---|
| Simple | 93.2% |
| Multiple | 95.5% |
| Parallel | 87.0% |
| Parallel-multiple | 70.0% |
| Irrelevance | 91.2% |
| Item | Specification |
|---|---|
| Base model | Qwen/Qwen3.6-27B |
| Base revision | 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9 |
| Release checkpoint | BTL-3 RL-0013 |
| Adapter | PEFT LoRA, rank 32, alpha 64 |
| Adapter size | 933,974,032 bytes |
| Architectural context | 262,144 tokens |
| Maximum RL sequence length | 65,536 tokens |
| Launch benchmark context | 32,768 tokens |
| Recommended mode | Thinking enabled for coding and reasoning |
| License | Apache-2.0 |
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base_id = "Qwen/Qwen3.6-27B"
base_revision = "6a9e13bd6fc8f0983b9b99948120bc37f49c13e9"
adapter_id = "badtheorylabs/BTL-3"
tokenizer = AutoTokenizer.from_pretrained(adapter_id)
base = AutoModelForCausalLM.from_pretrained(
base_id,
revision=base_revision,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model = PeftModel.from_pretrained(base, adapter_id)
messages = [
{
"role": "user",
"content": "Inspect this repository, fix the failing tests, and explain the patch.",
}
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=4096)
completion = output[0, inputs.input_ids.shape[1]:]
print(tokenizer.decode(completion, skip_special_tokens=False))
The supported release default is non-thinking mode. Thinking remains available for controlled evaluation, but it is currently discouraged because the RL-0013 reasoning policy can become repetitive or fail to terminate on some prompts.
BTL-3 was evaluated with vLLM 0.23.0:
vllm serve Qwen/Qwen3.6-27B \
--revision 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9 \
--served-model-name BTL-3 \
--enable-lora \
--max-lora-rank 32 \
--lora-modules BTL-3=/path/to/BTL-3 \
--lora-target-modules \
q_proj k_proj v_proj o_proj \
in_proj_qkv in_proj_z in_proj_b in_proj_a out_proj \
gate_proj up_proj down_proj \
--reasoning-parser qwen3 \
--language-model-only \
--max-model-len 32768
For structured tools, enable the Qwen XML tool parser supported by your installed vLLM version.
| Artifact | SHA-256 |
|---|---|
adapter_model.safetensors | 37a8f519039707eba5906591cdb14268768db43f80489a9c2f83b3e51e5e89db |
Run generated code and tool calls in a sandbox. Require explicit confirmation before destructive, privileged, financial, or otherwise high-impact actions.
The adapter is released under Apache-2.0 and requires the separately distributed Qwen3.6-27B base model.
@software{btl3_2026,
title = {BTL-3: A 27B Agentic Coding and Tool-Use Model},
author = {Bad Theory Labs},
year = {2026},
url = {https://huggingface.co/badtheorylabs/BTL-3}
}
For questions and release updates, visit Bad Theory Labs or join the community Discord.