Downloads · 30 days
16
28% of all-time downloads
sluttybutfast/BattleKAT-Coder-V2.5-Dev-Vision-oQ4e
BattleKAT-Coder-V2.5-Dev-Vision-oQ4e is a image-text-to-text model from sluttybutfast. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as apache-2.0.
Downloads · 30 days
16
28% of all-time downloads
All-time downloads
57
Public
Parameters
35.1B
21.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors21.3 GB · 100%
How the weights are stored.
U3234.6B · 99%
From the Hugging Face model README

This is BattleKat, a battle hardened merge of great Qwen3.6-35B-A3B parts:
temp 0.8, top_p=0.95, min_p=0.02, top_k=20, repetition_penalty=1.1, presence_penalty=0.0, thinking_budget=none, enable_thinking=true, preserve_thinking=true, output_token_limit=16384
Install:
pip install mlx-vlm
Command line:
python -m mlx_vlm.generate \
--model YOUR_MODEL_NAME \
--image path/to/image.jpg \
--prompt "Describe this image. Extract all text" \
--max-tokens 5000
Python:
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
model, processor = load("YOUR_MODEL_NAME")
config = model.config
prompt = apply_chat_template(
processor, config,
"Describe this image. Extract all text",
num_images=1,
)
output = generate(model, processor, prompt, ["path/to/image.jpg"],
max_tokens=5000, verbose=True)
print(output)
The two source models share the same qwen3_5_moe architecture and an identical namespace convention, which made a clean weight-level merge possible without any retraining, fine-tuning, or projector re-alignment.
All tensors in both models fall into two disjoint prefixes:
The merge is a straightforward union: every language_model.* tensor from the finetune, and every vision_tower.* tensor from the base VLM.
The vision projector (vision_tower.merger) outputs vision_config.out_hidden_size = 2048, which matches the language model's text_config.hidden_size = 2048. Vision features therefore project directly into the LM embedding space with no adapter needed. Image features are injected at the image_token_id position (single-point injection; deepstack_visual_indexes is empty).
The merged config.json uses the finetune's config as the base (it holds the correct quantization map and text_config), with the following vision-related fields grafted in from the base VLM:
The multi-token-prediction head was not carried over (mtp_num_hidden_layers remains 0, matching the finetune). It does not seem reasonable to reapply a mtp, as the finetune has derived from the base and the base mtp will result in poor matching without further retraining. The image preprocessor (preprocessor_config.json) were taken from the base VLM, since the finetune's tokenizer configuration was text-only. The tokenizer vocabulary is identical between both sources (vocab_size = 248320, same special-token IDs), so the base VLM's template is fully compatible.
The merged weights were verified byte-for-byte via SHA-256 hashes of each tensor, comparing the merged output against both sources:
This model is a derivative combining:
License: apache-2.0. You must comply with the licenses of BOTH source models. Please cite and credit the original authors of both the base VLM and the finetune.
The merge was performed with a script (scripts/merge.py) that:
The merge was split into Hugging Face splittensors using the scripts/split.py.
The model was verified for correct tensors (checking text tensors were untouched) using the scripts/verify.py.
You should be able to do your own merges with the script as long as the two models are compatible (check for same dimensions, same tokenizer size). Verify and check vision capabilities (f.e. ocr a text).
The model uses a tuned froggeric chat_template with default reasoning_preserve=true and custom system prompt addons to lower thinking loop probability. If you prefer the original chat_template replace chat_template.json with chat_templat.json.org or any you like.