Downloads · 30 days
0
Least-gen/Spark-X2.5-4B-API
Spark-X2.5-4B-API is a machine learning model from Least-gen. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Live, metered OpenAI-compatible inference of XHToken/Spark-X2.5-4B (Apache-2.0) — deployed in Riyadh on a single NVIDIA RTX A4500, served directly from LeastGen's own GPU infrastructure. Not a rented hyperscaler node.
Downloads · 30 days
0
Access
Public
Updated Sep 8, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md3.3 KB · 68%
From the Hugging Face model README
Live, metered OpenAI-compatible inference of XHToken/Spark-X2.5-4B (Apache-2.0) — deployed in Riyadh on a single NVIDIA RTX A4500, served directly from LeastGen's own GPU infrastructure. Not a rented hyperscaler node.
🟢 Open the live chat demo → — an interactive Space running in your browser against this endpoint.
ℹ️ On the "Inference Providers" panel: listing there requires joining Hugging Face's official first-party provider program (industry partnership, separate application and SLA vetting). LeastGen serves this model independently from its own hardware — use the endpoint or the Space above to use it today.
🟢 100% FREE for now — no API key needed. Use the live demo Space, or grab your own free key at https://leastgen.com/credits if you want your own metered balance.
Option A — no key, chat now: Open the live demo Space and send a message. It runs in your browser against LeastGen's own GPU.
Option B — OpenAI-compatible API (still free for now):
Any OpenAI client works against https://api.leastgen.com/v1:
from openai import OpenAI
client = OpenAI(
base_url="https://api.leastgen.com/v1",
api_key=" your-key ",
)
resp = client.chat.completions.create(
model="Spark-X2.5-4B",
messages=[{"role": "user", "content": "What is sovereign AI infrastructure?"}],
max_tokens=200,
)
print(resp.choices[0].message.content)
curl https://api.leastgen.com/v1/chat/completions \
-H "Authorization: Bearer $LG_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Spark-X2.5-4B",
"messages": [{"role":"user","content":"Hello"}],
"max_tokens": 80
}'
| Piece | Detail |
|---|---|
| Weights | XHToken/Spark-X2.5-4B full-precision BF16, Apache-2.0 |
| Serving engine | llama.cpp (llama-server) with GPU offload |
| Hardware | Deployed in Riyadh on a dedicated NVIDIA RTX A4500 (20GB, single node) |
| Transport | Cloudflare Tunnel → Tailscale private lane (no exposed GPU ports) |
| Metering | LeastGen credit system: 1 credit / input token · 2 credits / output token |
| Throughput measured | ~60 tok/s sustained single-stream, ~350 tok/s prompt eval |
| Context | 8,192 tokens, 8 concurrent slots |
Best-effort demo lane: This deployment has no SLA. It's 100% free for now — FAIR USE only (no bulk scraping, no bots, 8 concurrent-chat slots). Metering is kept on internally so LeastGen can watch demand, nothing is charged to you yet. When the free window closes, the same endpoint keeps working — you'll just pay by the credit.
Internally, 17×23 arrives back as 391 at ~61 tok/s — inference runs entirely on LeastGen's own RTX A4500, physically deployed in Riyadh. No datacenter in between.
Copyright and weight files remain with XHToken, Apache-2.0.