Downloads · 30 days
0
nvkudva/laya-web-q8
laya-web-q8 is a text classification model from nvkudva. Use it when you need a label for a piece of text. It is set up for onnx. The card lists the license as apache-2.0.
▶ Try it — laya-web.pages.dev — loads in a browser tab, runs on your own machine, sends nothing anywhere.
Downloads · 30 days
0
Access
Public
Updated Sep 20, 2026
Repo size
524 MB
Likes
2
Public
Click a slice to open those files.
.data521 MB · 99%
From the Hugging Face model README
▶ Try it — laya-web.pages.dev — loads in a browser tab, runs on your own machine, sends nothing anywhere.

This is convaiinnovations/laya —
the English ModernBERT-large checkpoint — exported to ONNX and quantized to 8-bit so
it fits in a web page. 1688 MB of fp32 becomes 524 MB, with argmax agreement
unchanged and a worst-case probability shift of 0.0158 across a 26-question parity
set.
All credit for the model, the training method and the results belongs to Nandakishor M and Convai Innovations. This repository contributes only the quantization and the browser runtime.
| Base model | convaiinnovations/laya (English root) |
| Original code | github.com/NandhaKishorM/laya |
| This conversion | github.com/nvkudva/laya-web |
| Live demo | laya-web.pages.dev |
| Licence | Apache 2.0, inherited from the base model |
Laya is not a generative model. It reads a state, scores the options you enumerate, and returns one calibrated probability distribution per question in a single forward pass. There is no sampling and no free-text output, so there is nothing to hallucinate — the answer space is whatever you listed.
Three question types:
| Type | Answer | Use |
|---|---|---|
noul | p(true) | Is this phishing? Should this escalate? |
choice | one named option + the full distribution | Which queue? Which policy? |
score | the expectation over ordered levels | How severe? How urgent? |
| File | Size | What |
|---|---|---|
v1/encoder_q8.onnx + .data | 471 MB | ModernBERT-large encoder, 28 layers, d=1024 |
v1/head_q8.onnx + .data | 53 MB | type embedding, 2 head layers, marker scorer, act head |
v1/tokenizer.json, v1/tokenizer_config.json | 3.6 MB | unchanged from the base model |
v1/rl_agent_config.json | — | max_len, head_max_len and the fitted temperatures |
The path is versioned on purpose. Browser caches key on URL, so re-quantizing goes
to v2/ rather than silently serving stale weights to anyone who already has v1/.
Ordinary dynamic INT8 destroys this model. onnxruntime.quantization.quantize_dynamic
drops argmax agreement to 69% with a worst-case probability shift of 0.99. An
ablation localises the damage:
| Variant | argmax agreement | max abs Δp | mean KL |
|---|---|---|---|
| fp32 reference | — | — | — |
| dynamic int8, per-tensor | 69.2% | 0.990 | 5.6e-01 |
| dynamic int8, per-channel | 76.9% | 0.995 | 4.6e-01 |
| dynamic int8, MatMuls only | 65.4% | 0.996 | 7.3e-01 |
| dynamic int8, embeddings only | 100% | 0.216 | 8.0e-03 |
| shipped: weight-only int8 | 100% | 0.0158 | 1.8e-04 |
Quantizing the MatMuls alone is catastrophic while quantizing the embeddings alone is survivable, and per-channel weight scales barely help — so the problem is activation quantization, not weight precision. ModernBERT has outlier activation channels that a per-tensor dynamic scale cannot represent, the same failure that motivated LLM.int8() and SmoothQuant.
Weight-only quantization leaves activations in fp32 and avoids it entirely:
MatMulNBits (block size 64), which
dequantizes inside the kernel, so nothing ever materialises a 1.6 GB fp32 tensor.Measured against the fp32 PyTorch model over 26 questions spanning all three types, cardinalities 2–14, both truncation branches, non-Latin script and degenerate inputs:
| Metric | Result |
|---|---|
| Argmax agreement | 100% (26/26) |
| Max absolute Δp | 0.0158 |
| Mean KL(fp32 ‖ int8) | 1.8e-04 |
| Tokenization | byte-identical token ids on all 26 |
The largest shifts land on questions the model is already uncertain about — a noul
sitting near p=0.5 moves most, which is where quantization error is least consequential
for a decision and most visible as a number.
The intended consumer is onnxruntime-web.
The full TypeScript port of the tokenization, sequence construction and temperature
scaling lives in nvkudva/laya-web under
app/src/laya/.
import * as ort from "onnxruntime-web/wasm";
const BASE = "https://huggingface.co/nvkudva/laya-web-q8/resolve/main/v1";
const load = async (name: string) =>
ort.InferenceSession.create(`${BASE}/${name}.onnx`, {
executionProviders: ["wasm"],
externalData: [{ data: `${BASE}/${name}.onnx.data`, path: `${name}.onnx.data` }],
});
const encoder = await load("encoder_q8");
const head = await load("head_q8");
Sequence layout, which you must reproduce exactly:
[CLS] <type> question: <instructions> [SEP] [MASK] opt0 [MASK] opt1 … [SEP] <state> [SEP]
Each option is scored at its own [MASK] position; softmax over those positions,
divided by the temperature for that (question type, option count) bucket, is the
answer. max_len is 512 and head_max_len is 192.
MatMulNBits kernel accepts 2-bit
and 4-bit, not 8-bit, and rejects this graph. 4-bit would unlock WebGPU and cut the
encoder to 271 MB, but argmax collapses to 84.6% and max Δp to 0.347 — not worth it
for a model whose value is calibrated probabilities.Cross-Origin-Opener-Policy: same-origin plus Cross-Origin-Embedder-Policy: require-corp. The Hugging Face CDN
sends no Cross-Origin-Resource-Policy header, but it does not need to: COEP runs
the CORP check only on no-cors loads, and these files are fetched with fetch()
in cors mode, where both the resolve/ redirect and the CDN response are CORS-ok.
Do not use credentialless — Safari does not support it, so the page silently
loses isolation there, falls back to a single thread and gets roughly 6× slower.onnxruntime-web/wasm, not the default entry, which pulls in a 28 MB
jsep runtime you will not use.env.wasm.proxy in a production bundle.Chromium, Apple silicon, 8 threads, single question:
| State length | Latency |
|---|---|
| ~43 tokens | ~290 ms |
| ~195 tokens | ~920 ms |
| 512 tokens (max) | ~2.4 s |
The three-question demo preset completes in about 750 ms end to end.
These are properties of the base model, not of the quantization, and the original model card documents them fully.
laya-multilingual
for anything else.score is the weakest primitive (SST-5 0.372).choice degrades: at head_max_len = 192, a 77-option
question leaves 3–4 tokens per label. The published choice:11+ temperature is
0.1006, which sharpens the distribution close to one-hot.act_probability is saturated at 1.000 on every input tested here; the
act/escalate head carries no signal on this checkpoint.Cite the original work:
@misc{laya2026,
title = {Laya: Non-Autoregressive System 1 Decision Models},
author = {Nandakishor M},
year = {2026},
howpublished = {\url{https://huggingface.co/convaiinnovations/laya}},
note = {Convai Innovations}
}
Nandakishor M and Convai Innovations built Laya, trained it with RLCD, and released the weights and code under Apache 2.0. Read the author's write-up on Dev.to.
Quantization and browser runtime by nvkudva.