Downloads ยท 30 days
104
10% of all-time downloads
ruv/ruvltra-small
ruvltra-small is a text generation model from ruv. Use it when you need the model to write or continue text. It is set up for gguf. The card lists the license as apache-2.0.
Downloads ยท 30 days
104
10% of all-time downloads
All-time downloads
991
Public
Repo size
398 MB
Likes
2
Public
Click a slice to open those files.
.gguf398 MB ยท 100%
From the Hugging Face model README
๐ฑ Compact Model Optimized for Edge Devices
Quick Start โข Use Cases โข Integration
</div>RuvLTRA Small is a compact 0.5B parameter model designed for edge deployment. Perfect for mobile apps, IoT devices, and resource-constrained environments.
| Property | Value |
|---|---|
| Parameters | 0.5 Billion |
| Quantization | Q4_K_M |
| Context | 4,096 tokens |
| Size | ~398 MB |
| Min RAM | 1 GB |
# Download
wget https://huggingface.co/ruv/ruvltra-small/resolve/main/ruvltra-0.5b-q4_k_m.gguf
# Run with llama.cpp
./llama-cli -m ruvltra-0.5b-q4_k_m.gguf -p "Hello, I am" -n 64
use ruvllm::hub::ModelDownloader;
let path = ModelDownloader::new()
.download("ruv/ruvltra-small", None)
.await?;
from huggingface_hub import hf_hub_download
model = hf_hub_download("ruv/ruvltra-small", "ruvltra-0.5b-q4_k_m.gguf")
License: Apache 2.0 | GitHub: ruvnet/ruvector
RuvLTRA models are fully compatible with TurboQuant โ 2-4 bit KV-cache quantization that reduces inference memory by 6-8x with <0.5% quality loss.
| Quantization | Compression | Quality Loss | Best For |
|---|---|---|---|
| 3-bit | 10.7x | <1% | Recommended โ best balance |
| 4-bit | 8x | <0.5% | High quality, long context |
| 2-bit | 32x | ~2% | Edge devices, max savings |
cargo add ruvllm # Rust
npm install @ruvector/ruvllm # Node.js
use ruvllm::quantize::turbo_quant::{TurboQuantCompressor, TurboQuantConfig, TurboQuantBits};
let config = TurboQuantConfig {
bits: TurboQuantBits::Bit3_5, // 10.7x compression
use_qjl: true,
..Default::default()
};
let compressor = TurboQuantCompressor::new(config)?;
let compressed = compressor.compress_batch(&kv_vectors)?;
let scores = compressor.inner_product_batch_optimized(&query, &compressed)?;
RuVector GitHub | ruvllm crate | @ruvector/ruvllm npm
| Metric | Result |
|---|---|
| Inference Speed | 75.4 tok/s |
| Model Load Time | 1.44s |
| Parameters | 0.5B |
| TurboQuant KV (3-bit) | 10.7x compression, <1% PPL loss |
| TurboQuant KV (4-bit) | 8x compression, <0.5% PPL loss |
Benchmarked on Google Cloud L4 GPU via ruvltra-calibration Cloud Run Job (2026-03-28)