Downloads · 30 days
48
16% of all-time downloads
Hal0ai/Qwen-AgentWorld-Hal0-35B-A3B-ROCmFP4
Qwen-AgentWorld-Hal0-35B-A3B-ROCmFP4 is a text generation model from Hal0ai. Use it when you need the model to write or continue text. It is set up for gguf. The card lists the license as apache-2.0.
A ROCmFP4 (Q40ROCMFP4STRIXLEAN, ~4.29 bpw) GGUF quant of Qwen/Qwen-AgentWorld-35B-A3B — a 35B-A3B Mixture-of-Experts (qwen35moe) world-model, 262K context.
Downloads · 30 days
48
16% of all-time downloads
All-time downloads
296
Public
Repo size
18.6 GB
Likes
0
Public
Click a slice to open those files.
.gguf18.6 GB · 100%
From the Hugging Face model README
A ROCmFP4 (Q4_0_ROCMFP4_STRIX_LEAN, ~4.29 bpw) GGUF quant of
Qwen/Qwen-AgentWorld-35B-A3B
— a 35B-A3B Mixture-of-Experts (qwen35moe) world-model, 262K context.
Built for AMD Strix Halo (gfx1151) unified-memory inference on the hal0 agent platform.
ROCmFP4 is an experimental, fork-specific quantization format (UE4M3-scale
FP4). It is not loadable by stock llama.cpp or standard GGUF tooling —
those will fail with 101 is not a valid GGMLQuantizationType.
Run it with either:
ghcr.io/hal0ai/amd-strix-halo-toolboxes:rocm-7.2.4-rocmfp4-server, orrocmfp4-llama
fork (branch mtp-rocmfp4-strix), targeting gfx1151.| Architecture | qwen35moe (35B total / ~3B active, MoE) |
| Quant | Q4_0_ROCMFP4_STRIX_LEAN (~4.38 bpw target; 4.29 bpw measured) |
| Recipe | ROCmFP4 experts/FFN + Strix K/V + Q5_K token embeddings |
| File size | ~17.3 GiB (from 66 GiB BF16) |
| MTP | No (plain quant, no speculative-decode draft head) |
| Context | 262144 |
| Base | Qwen/Qwen-AgentWorld-35B-A3B (Apache-2.0) |
Quantized from the BF16 GGUF in
unsloth/Qwen-AgentWorld-35B-A3B-GGUF,
using that repo's importance matrix (imatrix_unsloth.gguf), via the hal0
ROCmFP4 quantize pipeline (llama-quantize --imatrix ... Q4_0_ROCMFP4_STRIX_LEAN).