Downloads · 30 days
10
4% of all-time downloads
Sharjeelbaig/Supra-Router-51M-ONNX
Supra-Router-51M-ONNX is a text generation model from Sharjeelbaig. Use it when you need the model to write or continue text. It is set up for transformers.js.
Browser-ready ONNX conversion of SupraLabs/Supra-Router-51M. The graph is dynamically quantized to signed INT8 and packaged using the Hugging Face ONNX repository layout.
Downloads · 30 days
10
4% of all-time downloads
All-time downloads
233
Public
Repo size
69.4 MB
Likes
0
Public
Click a slice to open those files.
.onnx139 MB · 98%
From the Hugging Face model README
Browser-ready ONNX conversion of SupraLabs/Supra-Router-51M. The graph is dynamically quantized to signed INT8 and packaged using the Hugging Face ONNX repository layout.
LlamaForCausalLMnpm install @huggingface/transformers
import { pipeline } from "@huggingface/transformers";
const modelId = "Sharjeelbaig/Supra-Router-51M-ONNX";
const router = await pipeline("text-generation", modelId, {
dtype: "int8",
});
const userPrompt = "Write Python code to find all primes below one million efficiently.";
const input = `Task: ${userPrompt}\nAnalysis: `;
const output = await router(input, {
max_new_tokens: 128,
do_sample: false,
return_full_text: false,
});
console.log(output[0].generated_text.trim());
The model returns a pipe-separated routing record:
Domain: ... | Complexity: 1-5 | Math: True/False | Code: True/False | Route: small model/big model | Justification: ...
pip install "optimum-onnx[onnxruntime]" transformers
from transformers import AutoTokenizer, pipeline
from optimum.onnxruntime import ORTModelForCausalLM
model_id = "Sharjeelbaig/Supra-Router-51M-ONNX"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = ORTModelForCausalLM.from_pretrained(
model_id,
subfolder="onnx",
file_name="model_int8.onnx",
use_cache=False,
)
router = pipeline("text-generation", model=model, tokenizer=tokenizer)
prompt = "Explain why database deadlocks occur and provide code to prevent them."
result = router(
f"Task: {prompt}\nAnalysis: ",
max_new_tokens=128,
do_sample=False,
return_full_text=False,
)
print(result[0]["generated_text"].strip())
The graph accepts:
input_ids: int64[batch, sequence]attention_mask: int64[batch, past_sequence + sequence]position_ids: int64[batch, sequence]past_key_values.{0..11}.{key,value}: cached attention tensorsIt returns logits: float32[batch, sequence, 32000] plus present.{0..11}.{key,value} cache tensors. For direct integration, initialize each cache with shape [batch, 4, 0, 64], then pass each returned present tensor back as the corresponding past_key_values input on the next step. Hugging Face pipelines manage this automatically.
onnx.checker.1.4e-4 during export validation.small model and big model prompts.This is a small routing model trained on 992 examples. Treat its route as a heuristic, not a security boundary or sole safety control. Add deterministic policy checks, timeouts, input-length limits, and a conservative fallback in production.
The upstream model card does not declare a license at the time of this conversion. This repository does not add or replace upstream rights. Confirm use and redistribution terms with SupraLabs before commercial deployment or redistribution.
This conversion was produced independently and is not an official SupraLabs release. Refer to the upstream model card for intended input formatting and training details.