Downloads · 30 days
0
hdae/karume-gemma4
karume-gemma4 is a text generation model from hdae. Use it when you need the model to write or continue text. It is set up for karume. The card lists the license as apache-2.0.
A chat distribution: the text decoder of google/gemma-4-E2B-it, converted into the WebGPU inference runtime Karume's container format (a .krm part sequence whose first part carries the graph and model descriptors). Ru…
Downloads · 30 days
0
Access
Public
Updated Sep 25, 2026
Repo size
8 GB
Likes
0
Public
Click a slice to open those files.
.krm4 GB · 100%
From the Hugging Face model README
A chat distribution: the text decoder of google/gemma-4-E2B-it,
converted into the WebGPU inference runtime Karume's container format (a .krm part
sequence whose first part carries the graph and model descriptors). Runs as-is in the
browser and in Deno — string in, string out.
chat() renders the turns,
encodes them, samples, and decodes incrementally, so callers hand over messages and
read back text fragments as they are decided.rope parameters, so no position
table ships. Prefill runs in chunks of 768 rows by default, and
the chunk length itself can be raised up to 768 — the traced
upper bound of the chunk symbol, which the graph does not carry.gemma4/1.karume/0.13.0. The distribution manifest
is karume.json (karume/5).Converted and quantized from the upstream checkpoint — the original weights are not distributed here.
e2b: google/gemma-4-E2B-it, licensed Apache 2.0 (license / full text; a verbatim copy is in LICENSE.md).assistant: google/gemma-4-E2B-it-assistant, the multi-token-prediction drafter head, licensed Apache 2.0 under the same terms. It is redistributed here with its clustered sparse output head replaced by a dense projection over the full vocabulary, and it reads the key/value states and embedding table of the main model rather than carrying its own. Its own weights are quantized to int8 throughout, linear layers included, rather than to packed int4.NOTICE.md, per Apache 2.0 §4(b)): the text
decoder was extracted and re-expressed in the Karume container format; linear weights
were quantized to packed int4 (group 32) and the embedding tables to int8; the
per-layer embeddings were moved out of the graph into container assets that the host
gathers; the exit was narrowed to the selected rows' logits and hidden states; rotary
cos/sin arrive as host-generated inputs.
No retraining and no fine-tuning.| Model | Pipeline | Quants | Default quant |
|---|---|---|---|
e2b (default) | gemma4/1 | i4 / i4-gemvpar / i4-fast | i4-fast |
model selects one of these; omitted, it is e2b. quant defaults to that model's own default quant.
import { Gemma4Pipeline } from "jsr:@karume/models/gemma";
await using pipeline = await Gemma4Pipeline.fromPretrained({
repo: "hdae/karume-gemma4",
// Pin a commit for reproducible builds — without it you track `main`, and a future
// repo update (renamed files, new manifest format) may break your app.
// Copy the full hash from this repo's "Files and versions" tab:
// revision: "<full commit sha>",
}, {
// model: "e2b", // default — available: e2b
// quant: "i4-fast", // default — available: i4 / i4-fast / i4-gemvpar
});
const stream = pipeline.chat([
{ role: "user", content: "What is the capital of France?" },
], {
maxNewTokens: 128,
// sampler: { temperature: 0 }, // greedy — overrides the recommended default below
});
// Fragments arrive as the decoder settles them (multi-byte characters are held back
// until they are complete).
for await (const text of stream) Deno.stdout.write(new TextEncoder().encode(text));
console.log(await stream.done); // { reason: "eos" | "max-tokens" | "aborted", … }
Messages are plain system / developer / user / assistant turns; tool calls,
thinking channels and image or audio parts are rejected rather than silently dropped.
With no sampler in the request, generation uses this repository's recommended default: temperature 1.0, top-k 64, top-p 0.95.
Weights are fetched once and cached (verified against karume.json's size / sha256).
| Quant | What it is | Download | Weights | Compute |
|---|---|---|---|---|
i4 | Packed int4 linear, int8 embeddings — The only storage series: the main model's linear weights in packed int4 (group 32) and its embedding tables in int8, which are not int4-eligible. The drafter head is int8 throughout. | 3.78 GiB (2.23 GiB of assets, read on the host) | model = i4 / drafter = i8 | — |
i4-gemvpar | Packed int4 with parallel GEMV — The same packed weights as i4, with parallel GEMV summation. Faster on tested E2B devices; rounding and generated tokens can differ. Select i4 for the reference summation order. | 3.78 GiB (2.23 GiB of assets, read on the host) | model = i4 / drafter = i8 | linearGemvReduce = parallel |
i4-fast (default) | Packed int4 with parallel GEMV and RMS fusion — Same weights as i4; parallel GEMV and RMS-add fusion for E2B. Use i4 for reference summation or i4-gemvpar without fusion. Requires fusion-option support. | 3.78 GiB (2.23 GiB of assets, read on the host) | model = i4 / drafter = i8 | linearGemvReduce = parallel / fuseRmsNormAdd = true |
If no quant is given, it runs as i4-fast (this model's recommended default).
Per-file size and sha256 live in karume.json — verify against that at the fetch layer.
Dtype labels use the runtime's storage dtype vocabulary (f16 / i8 / i4 / i2), not the fp16 spelling common elsewhere in the ecosystem.
Weights ship as Karume container files (.krm), split into numbered parts; a part is fetched and verified on its own.
Derived from the exported graph and the checkpoint's own generation_config.json, and
checked against each other when this repository was assembled.
sampler; pass { temperature: 0 } for greedy decodingmaxBufferSize ≥ 402,653,184 B / maxStorageBufferBindingSize ≥ 402,653,184 B (the largest single tensor must bind)