Downloads · 30 days
2.5M
100% of all-time downloads
CMSManhattan/JiRackUltra_1b
JiRackUltra_1b is a text generation model from CMSManhattan. Use it when you need the model to write or continue text. The card lists the license as mit.
A fast and efficient ~1.5B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Tool call, and Robotics tags. Built on a redesigned DeepS…
Downloads · 30 days
2.5M
100% of all-time downloads
All-time downloads
2.5M
Public
Parameters
1.8B
16.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.gguf9 GB · 56%
From the Hugging Face model README
A fast and efficient ~1.5B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Tool call, and Robotics tags. Built on a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support and ready-to-run GGUF quantizations.
| Tag | Quant | Size | Approx. RAM | Description |
|---|---|---|---|---|
cmsmanhattan/jirack-ultra-1b-cpu:latest | Full | 0.55 GB | ~1.8 GB | Full ternary reference |
cmsmanhattan/jirack-ultra-1b-cpu-q4:latest | Q4_K_M | 0.38 GB | ~1.4 GB | Recommended balance |
cmsmanhattan/jirack-ultra-1b-cpu-q3:latest | Q3_K_M | 0.31 GB | ~1.2 GB | Good quality / size trade-off |
cmsmanhattan/jirack-ultra-1b-cpu-q2:latest | Q2_K | 0.24 GB | ~1.0 GB | Maximum compression |
Default CPU (Q4 recommended)
docker run -d \
--name jirack_ultra_1b \
-p 7869:7869 \
--cpus=16 \
-e THREADS=16 \
-e THREADS_BATCH=16 \
--restart unless-stopped \
cmsmanhattan/jirack-ultra-1b-cpu-q4:latest
Q3
docker run -d \
--name jirack_ultra_1b \
-p 7869:7869 \
--cpus=16 \
-e THREADS=16 \
-e THREADS_BATCH=16 \
--restart unless-stopped \
cmsmanhattan/jirack-ultra-1b-cpu-q3:latest
Q2 (lowest memory)
docker run -d \
--name jirack_ultra_1b \
-p 7869:7869 \
--cpus=16 \
-e THREADS=16 \
-e THREADS_BATCH=16 \
--restart unless-stopped \
cmsmanhattan/jirack-ultra-1b-cpu-q2:latest
Full precision
docker run -d \
--name jirack_ultra_1b \
-p 7869:7869 \
--cpus=16 \
-e THREADS=16 \
-e THREADS_BATCH=16 \
--restart unless-stopped \
cmsmanhattan/jirack-ultra-1b-cpu:latest
Multi CPU
docker run -d \
--name jirack_ultra_1b \
-p 7869:7869 \
--cpus=16 \
-e THREADS=16 \
-e THREADS_BATCH=16 \
--restart unless-stopped \
--memory=4g \
--cpus=4 \
cmsmanhattan/jirack-ultra-1b-cpu-q4:latest
services:
jirack:
image: cmsmanhattan/jirack-ultra-1b-cpu-q4:latest
container_name: jirack_ultra_1b
ports:
- "7869:7869"
volumes:
- .:/app
- ./web:/app/web
environment:
- MAX_TOKENS=2048
- TEMPERATURE=0.7
- TOP_P=0.9
- DEFAULT_STREAM=False
- INTRA_THREADS=4
- USE_ENV_ALLOCATOR=1
- THREADS=16
- THREADS_BATCH=16
deploy:
resources:
limits:
memory: 4g
Once the container is running, open your browser and navigate to:
http://localhost:7869
This opens the JiRack UI — a clean web interface.
The listening port can be easily modified directly from the Settings panel within the JiRack UI.
| Use Case | CPU | RAM | Recommended Quant | Expected Speed | Recommendation |
|---|---|---|---|---|---|
| Recommended | Ryzen 5 / Intel i5 | 4–8 GB | Q4_K_M | Excellent interactive | Best choice |
| High Performance | Ryzen 7 / Intel i7 | 8–16 GB | Full / Q4 | Excellent | Excellent |
| Low Memory | Modern 4+ core CPU | 2–4 GB | Q3_K_M or Q2_K | Usable | Acceptable |
| Edge / Minimal | Laptop / SBC CPU | 2 GB | Q2_K | Acceptable | Budget option |
Even though the quantized 1B models are very small, we recommend the following for best experience:
For joint venture opportunities, hardware integration, or licensing inquiries:
MIT License