Downloads · 30 days
495
31% of all-time downloads
vincespeed/Ling-3.0-tiny-APEX-GGUF
Ling-3.0-tiny-APEX-GGUF is a text generation model from vincespeed. Use it when you need the model to write or continue text. The card lists the license as mit.
This repository contains 3 quantized GGUF profiles of the inclusionAI/Ling-3.0-tiny model, produced using Apex-Quant technology.
Downloads · 30 days
495
31% of all-time downloads
All-time downloads
1.6K
Public
Repo size
15.7 GB
Likes
2
Public
Click a slice to open those files.
.gguf15.7 GB · 100%
From the Hugging Face model README
This repository contains 3 quantized GGUF profiles of the inclusionAI/Ling-3.0-tiny model, produced using Apex-Quant technology.
| Profile | Size | BPW | Use Case |
|---|---|---|---|
| i-quality | 5.4 GB | 5.83 | Highest quality, production environments |
| i-balanced | 5.6 GB | 6.02 | Balanced quality and performance |
| i-compact | 3.7 GB | 4.03 | Compact deployment, low RAM |
BPW = Bits Per Weight. Higher value = better quality.
models/
├── Ling-3.0-tiny-i-quality.gguf # 5.4 GB — Highest quality
├── Ling-3.0-tiny-i-balanced.gguf # 5.6 GB — Balanced
└── Ling-3.0-tiny-i-compact.gguf # 3.7 GB — Compact
These models were created based on the inclusionAI/Ling-3.0-tiny model from HuggingFace.
These quantized models were produced using Apex-Quant technology.
llama-quantize)apex-quant/scripts/quantize.shbailingmoe3# Run with i-quality profile
./main -m models/Ling-3.0-tiny-i-quality.gguf -n 128 -p "Hello, how are you?"
# Run with i-compact profile
./main -m models/Ling-3.0-tiny-i-compact.gguf -n 128 -p "Hello, how are you?"
# Create Dockerfile or Ollamafile
FROM llama.cpp
COPY models/Ling-3.0-tiny-i-quality.gguf /model.gguf
from llama_cpp import Llama
llm = Llama(
model_path="models/Ling-3.0-tiny-i-quality.gguf",
n_ctx=4096,
n_threads=8
)
output = llm(
"Hello, how are you?",
max_tokens=128
)
print(output["choices"][0]["text"])
| Criterion | i-quality | i-balanced | i-compact |
|---|---|---|---|
| Quality | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐ |
| Speed | ⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ |
| RAM | High | Medium | Low |
| Size | 5.4 GB | 5.6 GB | 3.7 GB |
| BPW | 5.83 | 6.02 | 4.03 |
i-mini profile cannot be quantized without imatrix. ~100-200 inference samples must be run on the model to generate the importance matrix.The original model is distributed under the MIT license. The quantized models are shared under the same license.
Note: These models are quantized for local use. Check the original model's license for commercial use.