Downloads · 30 days
593
41% of all-time downloads
mrsladoje/CodeRankEmbed-onnx-int8
CodeRankEmbed-onnx-int8 is a feature extraction model from mrsladoje. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as mit.
INT8 quantized ONNX version of nomic-ai/CodeRankEmbed for code search and embedding.
Downloads · 30 days
593
41% of all-time downloads
All-time downloads
1.5K
Public
Repo size
277 MB
Likes
1
Public
Click a slice to open those files.
.onnx139 MB · 99%
From the Hugging Face model README
INT8 quantized ONNX version of nomic-ai/CodeRankEmbed for code search and embedding.
Dynamic INT8 quantization with reduce_range=True for cross-platform
correctness:
from onnxruntime.quantization import quantize_dynamic, QuantType
quantize_dynamic(
model_input='CodeRankEmbed_fp32.onnx',
model_output='CodeRankEmbed_int8.onnx',
weight_type=QuantType.QInt8,
per_channel=True,
reduce_range=True, # clamp weights to [-64, 63] for AVX2 kernel safety
)
reduce_range=TrueORT's CPU INT8 MatMul kernels have two paths on x86:
| CPU | Path | Full-range INT8 weights |
|---|---|---|
| Intel Cascade Lake+ / Ice Lake+ (VNNI) | VPDPBUSD | ✓ correct |
| AMD Zen 4+ (VNNI / Genoa+) | VPDPBUSD | ✓ correct |
| Apple Silicon (arm64 NEON + AMX) | separate arm64 kernels | ✓ correct |
| Intel pre-2019 / AMD Zen 3 Milan (AVX2 only) | pmaddubsw + phaddsw + paddd | ✗ int16 accumulator overflows → degenerate output |
reduce_range=True clamps weights to [-64, 63] (7-bit signed range), giving
the AVX2 int16 intermediate enough headroom to avoid overflow. VNNI and arm64
paths are unaffected (they handle full-range INT8 natively).
A previous version of this model was quantized without reduce_range=True.
It worked correctly on VNNI-capable CPUs and Apple Silicon, but produced
degenerate embeddings (all texts mapping to near-identical vectors) on
AMD Zen 3 EPYC and similar pre-VNNI x86 hosts — verified on RunPod
RTX 5090 pods with EPYC 7543. This version fixes that. See commit history.
T0 "how to parse json in python" T3 "parse json data python" cos=0.7749 (similar)
T0 "how to parse json in python" T2 "sql inner join three tables" cos=0.1123 (dissimilar)
Semantic separation 0.6626 (≥ 0.15 healthy)
import onnxruntime as ort
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("mrsladoje/CodeRankEmbed-onnx-int8")
session = ort.InferenceSession("model.onnx")
inputs = tokenizer(
"your code or query here",
padding=True, truncation=True, max_length=512, return_tensors="np"
)
outputs = session.run(None, dict(inputs))
# sentence_embedding is typically the second output; it's 768-dim L2-normalized
onnx/model.onnx — INT8 quantized model (139 MB)tokenizer.json, vocab.txt, config.json, special_tokens_map.json, tokenizer_config.json
— from the base nomic-ai/CodeRankEmbed distributionreduce_range=True)4eae31d09b1843103a1ebd5e2b2e24b5a5cad441a33906b35b12b1e2ed91d1db
Pin this in your downloader to guarantee you got the corrected weights and not a stale cached copy of v1.
nomic-ai/CodeRankEmbed (137M params), based on Snowflake/snowflake-arctic-embed-m-long. ONNX conversion derived from jalipalo/CodeRankEmbed-onnx.
MIT (inherited from base model).