Downloads · 30 days
0
answerdotai/colbert-muvera-micro-onnx
colbert-muvera-micro-onnx is a machine learning model from answerdotai. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
ONNX export of NeuML/colbert-muvera-micro (Apache-2.0) for fastlate and any ONNX Runtime client. The weights are unchanged; this repo adds the graph, an INT8 dynamic quantization of it, and the ColBERT settings the mo…
Downloads · 30 days
0
Access
Public
Updated Sep 18, 2026
Repo size
22.1 MB
Likes
0
Public
Click a slice to open those files.
.onnx22.1 MB · 97%
From the Hugging Face model README
ONNX export of NeuML/colbert-muvera-micro (Apache-2.0) for fastlate and any ONNX Runtime client. The weights are unchanged; this repo adds the graph, an INT8 dynamic quantization of it, and the ColBERT settings the model was trained with.
model.onnx: transformer, the PyLate Dense projection layer(s), and per-token L2 normalization, in one graph. Inputs input_ids, attention_mask (int64, dynamic batch and length); output embeddings of shape (batch, seq, 128). Opset 17.model_int8.onnx: onnxruntime.quantization.quantize_dynamic of the above, QInt8 weights.onnx_config.json: prefix ids, lengths, skiplist, expansion and padding settings copied from the PyLate model, plus fde_center.tokenizer.json: the source tokenizer, with [Q] /[D] as added tokens.Query prefix [Q] (id 30522), document prefix [D] (id 30523), inserted right after the sequence-start token. Query length 32, document length 300. do_query_expansion=True: queries are padded to query_length with [MASK] (id 103), attention off on the padding, and every query vector is kept. Punctuation tokens (skiplist_words) are dropped from document embeddings after encoding, as in PyLate.
fde_center=False: whether to subtract the corpus mean token vector before building MUVERA fixed-dimensional encodings. Measured on a 3,000-chunk code/notebook corpus with FDEs of 4,096 dims: centring lowered shortlist recall for this model, so leave vectors as they are.
fp32 graph vs PyTorch forward: max abs diff 3.1e-07; int8 graph vs PyTorch: 0.054; int8 graph vs PyLate encode on a sample text: documents 0.025, queries 0.044.
Exported with PyLate 1.6.0, transformers 5.3.0, torch 2.14.0, onnxruntime 1.2x, on 2026-09-18.