Downloads · 30 days
0
p-yan/laya-quanto
laya-quanto is a machine learning model from p-yan. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for optimum-quanto. The card lists the license as apache-2.0.
This repository bundles native Optimum-Quanto weight-only quantizations of convaiinnovations/laya in one place:
Downloads · 30 days
0
Access
Public
Updated Sep 20, 2026
Repo size
806 MB
Likes
0
Public
Click a slice to open those files.
.safetensors806 MB · 100%
From the Hugging Face model README
This repository bundles native Optimum-Quanto weight-only quantizations of convaiinnovations/laya in one place:
| Variant | Weights | Quantized modules |
|---|---|---|
q8 | 480.5 MB | 8-bit Linear weights |
q4 | 325.1 MB | 4-bit Linear weights |
Embeddings and LayerNorm weights remain floating point. The variants use Quanto's native quantized Linear kernels; they are not the older CPU dequantization implementation.
pip install -r requirements.txt
from quanto_laya import load_quantized_agent
repo = "/path/to/laya-quanto"
agent = load_quantized_agent(repo, variant="q8", device="cuda")
# Or: variant="q4"
result = agent.system_one(state, questions)
The q8/ and q4/ directories contain each variant's model.safetensors, Quanto quantization map, and metadata. Shared tokenizer, encoder configuration, decision configuration, and loader files are at the repository root.
Derived from the original Laya checkpoint, released under its Apache 2.0 terms.