Downloads · 30 days
58
12% of all-time downloads
arunb74/Nanbeige4.2-3B
Nanbeige4.2-3B is a text generation model from arunb74. Use it when you need the model to write or continue text. The card lists the license as other.
This repository provides the GGUF conversions of the original Nanbeige/Nanbeige4.2-3B model. All credit for the model architecture and weights belongs to the original Nanbeige team.
Downloads · 30 days
58
12% of all-time downloads
All-time downloads
498
Public
Repo size
10.9 GB
Likes
0
Public
Click a slice to open those files.
.gguf10.9 GB · 100%
From the Hugging Face model README
This repository provides the GGUF conversions of the original Nanbeige/Nanbeige4.2-3B model. All credit for the model architecture and weights belongs to the original Nanbeige team.
Note: This repository contains only GGUF conversions. The original Hugging Face model is Nanbeige/Nanbeige4.2-3B.
Original Hugging Face Model
Nanbeige/Nanbeige4.2-3B
This repository does not modify or fine-tune the original model. It simply provides GGUF conversions for running the model locally with llama.cpp and other GGUF-compatible applications.
This repository contains the following GGUF files:
| File | Description |
|---|---|
| Nanbeige4.2-3B-BF16.gguf | Full-precision BF16 GGUF. Highest quality but requires significantly more RAM/VRAM. |
| Nanbeige4.2-3B-Q4_K_M.gguf | 4-bit quantized GGUF. Recommended for most users due to its excellent balance of quality, speed, and memory usage. |
You only need to download ONE of these files.
At the time of publishing, support for the Nanbeige architecture has not yet been merged into the main llama.cpp repository.
Please use the official Nanbeige fork of llama.cpp:
https://github.com/Nanbeige/llama.cpp/tree/nanbeige42
If you build the upstream ggml-org/llama.cpp, you may encounter an error similar to:
unknown model architecture: 'nanbeige'
The Nanbeige fork contains the required architecture support.
Clone the repository:
git clone --recursive -b nanbeige42 https://github.com/Nanbeige/llama.cpp.git
cd llama.cpp
Build:
cmake -B build
cmake --build build -j
Example:
./build/bin/llama-server \
-m Nanbeige4.2-3B-Q4_K_M.gguf \
--host 0.0.0.0 \
--port 8080 \
-ngl 999 \
-c 65536
After starting the server, the OpenAI-compatible API will be available at:
http://localhost:8080
Hermes works well with this model.
llama-server.http://localhost:8080
Hermes will communicate directly with your local Nanbeige model using the OpenAI-compatible API.
This model can also be used with web interfaces that support OpenAI-compatible endpoints, including:
Simply configure the application to connect to your running llama-server instance.
For most users, the recommended file is:
✅ Nanbeige4.2-3B-Q4_K_M.gguf
It provides an excellent balance of:
nanbeige42 branch.All credit for the model architecture, tokenizer, training, and original model weights belongs entirely to the original Nanbeige team.
This repository distributes GGUF conversions of the original model.
The license for these GGUF files is the same as the license of the original Nanbeige/Nanbeige4.2-3B project.
Please refer to the original model repository for the official license terms, usage conditions, and any applicable restrictions.