Downloads · 30 days
0
borkiss/qwen35-2b-webgpu
qwen35-2b-webgpu is a text generation model from borkiss. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Custom GPTQ-int4 quant (group 32, byte-sliced nibbles, single tied q4 copy of the 248k-vocab embedding serving both the gather and the logits cascade) of Qwen/Qwen3.5-2B (text part), in the wire format of the qwen35-2…
Downloads · 30 days
0
Access
Public
Updated Jul 6, 2026
Repo size
3.4 GB
Likes
0
Public
Click a slice to open those files.
.bin3.4 GB · 100%
From the Hugging Face model README
Custom GPTQ-int4 quant (group 32, byte-sliced nibbles, single tied q4 copy of the 248k-vocab embedding serving both the gather and the logits cascade) of Qwen/Qwen3.5-2B (text part), in the wire format of the qwen35-2b-webgpu browser runtime.
The first hybrid gated-DeltaNet + gated-attention LLM running in the browser — pure WebGPU, hand-written WGSL kernels (chunked WY-representation delta-rule prefill, fused recurrent decode step, f32 state), no ONNX / transformers.js / GGUF.
Try it: https://huggingface.co/spaces/borkiss/qwen35-2b-webgpu-demo · run /web/bench.html on your GPU and share the .log.