Downloads · 30 days
697
100% of all-time downloads
Accio-Lab/occamy-1.0-MLX-4bit
occamy-1.0-MLX-4bit is a text generation model from Accio-Lab. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
<div align="center" <img src="https://huggingface.co/Accio-Lab/occamy-1.0/resolve/f9e2771699f14d9c4bbbc0c7b913e58ebe1442d0/assets/accio.svg" width="240" alt="Accio" <h1Occamy-1.0 · MLX 4-bit</h1 <p<strongNative MLX ·…
Downloads · 30 days
697
100% of all-time downloads
All-time downloads
697
Public
Parameters
34.7B
19.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors19.5 GB · 100%
How the weights are stored.
U3234.7B · 100%
From the Hugging Face model README
Candidate release — Mac Metal acceptance is pending. Linux native MLX validation passed. Mac inference, performance and broad model quality remain unverified.
| Property | Value |
|---|---|
| Source | Accio-Lab/occamy-1.0 |
| Source revision | 8f8e0e58a3c9df042be1a3fa2c191fd8047acfb8 |
| Quantization | Native affine 4-bit, group size 64 |
| Router / shared-expert gates | 8-bit |
| Weight files | 19,509,024,201 bytes · 19.51 GB · 18.17 GiB |
| Runtime used for Linux checks | mlx 0.32.2, mlx-lm 0.31.3 |
| Inputs | Text only; vision and MTP are not included |
File size is not the unified-memory requirement. Leave room for the operating system, KV cache and runtime buffers. Compare the 3-bit, 4-bit, 6-bit and 8-bit candidates in the MLX collection.
On Apple Silicon, install the versions used to create this export. The following is a usage recipe awaiting Mac Metal acceptance:
python -m pip install "mlx==0.32.2" "mlx-lm==0.31.3"
from mlx_lm import load, generate
model, tokenizer = load("Accio-Lab/occamy-1.0-MLX-4bit")
messages = [{"role": "user", "content": "Compute 2+2. Answer briefly."}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
response = generate(model, tokenizer, prompt=prompt, max_tokens=128)
print(response)
The stock CLI options below were checked with mlx-lm 0.31.3. Mac Metal acceptance remains pending. HTTP inference checks for this release are recorded only for the new 6-bit and 8-bit exports.
python -m pip install "mlx==0.32.2" "mlx-lm==0.31.3" "transformers==5.8.1"
mlx_lm.server --model Accio-Lab/occamy-1.0-MLX-4bit \
--host 127.0.0.1 --port 8000 \
--chat-template-args '{"enable_thinking":false}'
In a second terminal:
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"Accio-Lab/occamy-1.0-MLX-4bit","messages":[{"role":"user","content":"Compute 2+2. Answer briefly."}],"temperature":0,"max_tokens":128}'
Client base URL: http://127.0.0.1:8000/v1. Inference runs on your machine.
A lossless adapter stacks separate expert weights in numeric expert order before invoking the Qwen3.5 sanitizer exactly once. Quantization and serialization use native APIs; reload uses the stock loader without an adapter.
Passed Linux checks: strict stock reload; complete stored floating-value checks; native dequantization of every quantized row; tokenizer/template comparison; and one bounded cached greedy CPU generation with finite logits. The prompt “Compute 2+2. Answer briefly.” returned 4. This is a limited smoke test, not a quality benchmark or a Mac runtime result.
Validation scope · Artifact hashes · Original model card
Apache 2.0, inherited from Occamy-1.0.
BF16 · GGUF · FP8 · NVFP4 · MLX 8-bit · MLX 6-bit · MLX 4-bit · MLX 3-bit · MTP head
Compare file sizes, validation scope and deployment commands in the checkpoint explorer.
8bit · 6bit · 5bit · 4bit · 3bit · mxfp8 · mxfp4 · nvfp4
The new 5-bit/MXFP4/MXFP8/NVFP4 cards include their own paired BF16 subset quality checks. Mac Metal acceptance remains pending.