Downloads · 30 days
965
42% of all-time downloads
VikramPal/kambo-v1
kambo-v1 is a text generation model from VikramPal. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A 1.7B-parameter hybrid convolution/attention Mixture-of-Experts language model, trained from scratch. Two experts of sixteen are active per token, so a forward pass costs about 0.5B parameters.
Downloads · 30 days
965
42% of all-time downloads
All-time downloads
2.3K
Public
Parameters
1.7B
3.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.4 GB · 100%
From the Hugging Face model README
A 1.7B-parameter hybrid convolution/attention Mixture-of-Experts language model, trained from scratch. Two experts of sixteen are active per token, so a forward pass costs about 0.5B parameters.
Open source under Apache 2.0. You are free to use, modify, and redistribute it, including commercially — see the License.
| Total parameters | 1,691,197,184 |
| Active parameters per token | 502,112,000 |
| Layers | 24 |
| Hidden size | 1024 |
| Mixer | 18 gated short-convolution layers + 6 grouped-query attention layers (at depths 3, 7, 11, 15, 19, 23) |
| Attention | 16 query heads / 4 key-value heads, head dim 64 |
| Feed-forward | Mixture-of-Experts in every layer: 16 routed experts, top-2, plus one always-on shared expert |
| Expert hidden size | 1152 |
| Normalisation | RMSNorm, pre-norm, with QK-norm on attention layers |
| Position encoding | Rotary, theta 40,000 |
| Context length | 16,384 |
| Vocabulary | 151,936 (embeddings tied to the output head) |
| Precision | bfloat16 |
Most layers are a double-gated short convolution rather than attention: the input is projected to three streams, two are multiplied together, passed through a causal depthwise convolution of kernel width 3, and gated by the third. This carries local context at a cost that does not grow with sequence length. Six attention layers, spaced evenly through the depth, carry the long-range dependencies. Because a convolution layer only needs to remember the last two columns of its input, the model's incremental-decoding state stays small as the context grows.
Every layer's feed-forward block is a Mixture-of-Experts. A router picks 2 of 16 experts per token; a shared expert runs for every token regardless. Routing is computed in float32 for stability and the published implementation is dropless — no token is discarded at any capacity limit.
Trained from scratch on 263.82B tokens in two phases:
| Phase | Tokens |
|---|---|
| Pretraining | 214.84B |
| Post-training (supervised fine-tuning + reinforcement learning) | 48.98B |
Pretraining covers general text and a context-length extension to 16,384. Post-training covers supervised fine-tuning for chat, tool use and instruction following, followed by reinforcement learning with verifiable rewards on tool-calling and constraint-following tasks.
Benchmark contamination was controlled by construction: no BFCL, IFBench, or Multi-IF item appears in any training corpus at any stage.
| Model | Params | AA-Omniscience | IFBench | Multi-IF | BFCLv3 | BFCLv4 |
|---|---|---|---|---|---|---|
| LFM2.5-8B-A1B | 8B/A1B | -24.70 | 56.47 | 79.93 | 64.79 | 49.73 |
| Qwen3-30B-A3B-Thinking-2507 | 30.5B/3.3B | -51.31 | 51.11 | 79.04 | 73.39 | 50.53 |
| Gemma-4-26B-A4B-IT | 26B/4B | -62.07 | 47.25 | 82.06 | 68.87 | 55.87 |
| gpt-oss-20b | 21B/3.6B | -49.17 | 58.65 | 76.64 | 62.52 | 49.88 |
| Qwen3.5-4B | 4B | -51.53 | 50.38 | 67.43 | 71.06 | 54.01 |
| Gemma-4-E4B-IT | 8B | -50.67 | 39.48 | 77.58 | 57.31 | 33.92 |
| Gemma-4-E2B-IT | 5.1B | -72.00 | 33.53 | 69.70 | 56.44 | 31.91 |
| Granite-4.0-H-Tiny | 7B/A1B | -75.50 | 21.28 | 59.00 | 56.89 | 28.52 |
| Kambo-v1 | 1.7B/A0.5B | -18.83 | 11.63 | 35.89 | 36.99* | 37.27* |
* BFCL figures for this model are AST accuracy averaged over the non-live and live splits; the peer figures are the overall leaderboard score.
On AA-Omniscience this model scores -18.83, the highest figure in this table — but that is not a knowledge result. The index rewards declining over guessing, and the model answered only 6 of 600 questions. It is measuring silence, not recall. The honest reading of this table is that tool-call formatting is where this model comes closest to the field, and that everything else is well behind.
The model ships its own modeling code, so trust_remote_code=True is required
on both the model and the tokenizer. It runs on CPU; in bfloat16 the weights are
about 3.4 GB.
pip install torch transformers
Tested with transformers 5.14.1; the snippets below use argument spellings
that also work on the 4.x series.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "VikramPal/kambo-v1"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16, # use torch.float32 on CPU if bfloat16 is slow
device_map="auto", # needs `accelerate`; omit to stay on CPU
)
messages = [{"role": "user", "content": "What is the capital of France?"}]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256, do_sample=True,
temperature=0.7, top_p=0.9)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:],
skip_special_tokens=True))
ChatML, and the packaged chat template reproduces the training format exactly:
<|im_start|>user
What is the capital of France?<|im_end|>
<|im_start|>assistant
Use apply_chat_template rather than building the string yourself. In
particular, no system message is inserted when you do not supply one — that
matches how the model was trained, and prepending a default system prompt will
push it off-distribution. Turns end with <|im_end|>.
model = AutoModelForCausalLM.from_pretrained(
model_id, trust_remote_code=True, torch_dtype=torch.float32
)
Batched generation requires left padding:
tokenizer.padding_side = "left"
batch = tokenizer(["Question one", "Question two"], return_tensors="pt", padding=True)
outputs = model.generate(**batch, max_new_tokens=64)
Released under the Apache License 2.0. See NOTICE for attribution and third-party components.
@misc{kambo_v1_2026,
title = {Kambo-v1: A 1.7B Hybrid Convolution-Attention Mixture-of-Experts Language Model},
author = {Kamboj, Vikrampal},
year = {2026},
note = {Apache-2.0},
url = {https://huggingface.co/VikramPal/kambo-v1}
}