Downloads · 30 days
0
RiverkanIT/Ling-mini-2.0-Quantized
Ling-mini-2.0-Quantized is a text generation model from RiverkanIT. Use it when you need the model to write or continue text. The card lists the license as mit.
Downloads · 30 days
0
Access
Public
Updated Sep 17, 2025
Repo size
26.4 GB
Likes
2
Public
Click a slice to open those files.
.bin26.4 GB · 100%
From the Hugging Face model README
Author and distribution: Riverkan
This repository provides CPU/GPU-friendly quantized builds of Ling‑Mini‑2.0 for ChatLLM.cpp. It is not a LLaMA model, is not affiliated with Meta, and does not use the LLaMA license. Files are distributed in ChatLLM.cpp’s GGMM-based format (.bin), ready for local inference.
Notes:
Quantized with the ChatLLM.cpp toolchain for GGMM-format inference (.bin). These builds are intended for the ChatLLM.cpp runtime (CPU and optional GPU acceleration as provided by ChatLLM’s GGMM backends). Use ChatLLM.cpp’s convert and run flow described below.
Original (float) model: to be announced by Riverkan.
Run them with ChatLLM.cpp or your preferred ChatLLM-based UI.
Ling‑Mini‑2.0 does not require a special role-tag chat template. Plain prompts work well. If your tooling prefers an explicit chat structure, you can use this neutral format:
[System]
You are Ling‑Mini‑2.0, a helpful, concise assistant.
[User]
{your question}
[Assistant]
Example:
[System]
You are Ling‑Mini‑2.0, a helpful, concise assistant.
[User]
List three tips to speed up CPU inference.
[Assistant]
No special tokens are required by the model itself; most UIs can just send user text.
| Filename | Quant type | File Size | Split | Description |
|---|---|---|---|---|
| Ling‑Mini‑2.0‑Q8_0.bin | Q8_0 | 16 GB | false | Highest quality quant provided here; best for quality, moderate speed. |
| Ling‑Mini‑2.0‑Q4_0.bin | Q4_0 | 8.52 GB | false | Great balance of speed and memory; recommended for CPU‑only setups. |
Notes:
git clone --recursive https://github.com/foldl/chatllm.cpp.git
cd chatllm.cpp
cmake -B build
cmake --build build -j --config Release
Place the quantized model file (e.g., Ling‑Mini‑2.0‑Q4_0.bin) somewhere accessible.
Run interactive chat:
# Linux / macOS
rlwrap ./build/bin/main -m /path/to/Ling‑Mini‑2.0‑Q4_0.bin -i
# Windows (PowerShell)
.\build\bin\Release\main.exe -m C:\path\to\Ling‑Mini‑2.0‑Q4_0.bin -i
./build/bin/main -m /path/to/Ling‑Mini‑2.0‑Q4_0.bin --prompt "Explain memory-bound vs compute-bound."
Tip: Run ./build/bin/main -h for all options (context size, threads, GPU offload where applicable, etc.).
Prompt:
[System]
You are Ling‑Mini‑2.0, a helpful, concise assistant.
[User]
Give me a 1‑paragraph summary of what quantization does for LLMs.
[Assistant]
Running:
./build/bin/main -m Ling‑Mini‑2.0‑Q8_0.bin -i --prompt "Give me a 1‑paragraph summary of what quantization does for LLMs."
In interactive mode (-i), simply paste your question and press Enter. The chat history is used as context for subsequent turns.
./build/bin/main -m ling-mini-2.0-q4.bin --seed 1
./build/bin/main -m ling-mini-2.0-q8.bin --seed 1
Notes:
If hosted on Hugging Face, you can fetch specific files with the CLI:
Install:
pip install -U "huggingface_hub[cli]"
Download a specific file:
huggingface-cli download RiverkanIT/Ling-mini-2.0-Quantized --include "Ling‑Mini‑2.0‑Q4_0.bin" --local-dir ./
Or the Q8_0 build:
huggingface-cli download RiverkanIT/Ling-mini-2.0-Quantized --include "Ling‑Mini‑2.0‑Q8_0.bin" --local-dir ./
Replace the model repo path with the actual hosting path if different.
If you have the float/base weights and want to generate your own GGMM quantized file for ChatLLM.cpp:
pip install -r requirements.txt
python convert.py -i /path/to/base/model -t q8_0 -o Ling‑Mini‑2.0‑Q8_0.bin --name "Ling-Mini-2.0"
python convert.py -i /path/to/base/model -t q4_0 -o Ling‑Mini‑2.0‑Q4_0.bin --name "Ling-Mini-2.0"
Notes:
For issues, feature requests, or contributions, please open a discussion or pull request in this repo.