Downloads · 30 days
14
50% of all-time downloads
ryandono/osgrep-colbert-fp32
osgrep-colbert-fp32 is a machine learning model from ryandono. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repository contains an ONNX export of mixedbread-ai/mxbai-edge-colbert-v0-17m produced with PyLate + a ColBERT-aware wrapper. It preserves the projection stack and ColBERT markers ([Q] / [D] ) and includes a skip…
Downloads · 30 days
14
50% of all-time downloads
All-time downloads
28
Public
Repo size
68 MB
Likes
0
Public
Click a slice to open those files.
.onnx68 MB · 95%
From the Hugging Face model README
This repository contains an ONNX export of mixedbread-ai/mxbai-edge-colbert-v0-17m produced with PyLate + a ColBERT-aware wrapper. It preserves the projection stack and ColBERT markers ([Q] / [D] ) and includes a skiplist for MaxSim.
onnx/model.onnx — FP32 export, opset 17 (✅ cosine 1.0 vs PyTorch)onnx/model_quantized.onnx — Dynamic INT8 (⚠️ cosine ~0.972 vs PyTorch; quality hit)tokenizer.json, tokenizer_config.json, special_tokens_map.json — saved from the PyLate-modified tokenizer with markersconfig.json — model configconversion_metadata.json — minimal export metadataskiplist.json — token IDs to skip during MaxSim (32 punctuation IDs for this model)Transformer 256 → Projection 512 → Projection 48
[Q] : 50368[D] : 50369import numpy as np, onnxruntime as ort
from transformers import AutoTokenizer
model_dir = "path/to/this/repo"
sess = ort.InferenceSession(f"{model_dir}/onnx/model.onnx", providers=["CPUExecutionProvider"])
tok = AutoTokenizer.from_pretrained(model_dir)
q = "[Q] what is colbert?"
enc = tok(q, return_tensors="np", padding="max_length", max_length=128, truncation=True)
out = sess.run(None, {"input_ids": enc["input_ids"], "attention_mask": enc["attention_mask"]})[0]
print(out.shape) # (batch, seq_len, 48)
import { AutoTokenizer } from "@huggingface/transformers";
import * as ort from "onnxruntime-node";
import fs from "fs";
const modelDir = "path/to/this/repo";
const tokenizer = await AutoTokenizer.from_pretrained(modelDir);
const session = await ort.InferenceSession.create(`${modelDir}/onnx/model.onnx`);
const q = "[Q] what is colbert?";
const encoded = await tokenizer(q, { return_tensors: "np", padding: "max_length", max_length: 128, truncation: true });
const outputs = await session.run({ input_ids: encoded.input_ids, attention_mask: encoded.attention_mask });
console.log(outputs[session.outputNames[0]].dims); // [1, 128, 48]
const skiplist = new Set(JSON.parse(fs.readFileSync(`${modelDir}/skiplist.json`, "utf8")));
conversion_metadata.json for details)