Downloads · 30 days
83
34% of all-time downloads
MainStack/marvy-2-35B-MoE-GGUF
marvy-2-35B-MoE-GGUF is a text generation model from MainStack. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
GGUF quants of marvy-2, a 35B-A3B Mixture-of-Experts model fine-tuned for the ServiceNow delivery lifecycle. ~3B active parameters per token; runs on a single consumer GPU or fast Apple Silicon thanks to the MoE spars…
Downloads · 30 days
83
34% of all-time downloads
All-time downloads
244
Public
Repo size
58.1 GB
Likes
0
Public
Click a slice to open those files.
.gguf58.1 GB · 100%
From the Hugging Face model README
GGUF quants of marvy-2, a 35B-A3B Mixture-of-Experts model fine-tuned for the ServiceNow delivery lifecycle. ~3B active parameters per token; runs on a single consumer GPU or fast Apple Silicon thanks to the MoE sparsity.
GGUF quantizations for use with llama.cpp, Ollama, LM Studio, and compatible runtimes.
Released under Apache-2.0. Built with Qwen3.5 (Apache-2.0) via
unsloth/Qwen3.6-35B-A3Band the Opus-distilledstamsam/...MTPbase.
| File | Quant | Size | Use when |
|---|---|---|---|
marvy-2-35B-MoE-Q4_K_M.gguf | Q4_K_M | ~20 GB | Default — best size/quality balance |
marvy-2-35B-MoE-Q8_0.gguf | Q8_0 | ~34 GB | Near-FP16 quality, more headroom |
This is a hybrid SSM + MoE Transformer:
qwen3_5_moe architecture in llama.cpp; needs a recent build
(see "Supported runtimes" below)The Multi-Token Prediction (MTP) head present in the base model was not included in this GGUF — these quants are text-only causal LM. The base model's MoE expert weights and SSM blocks are preserved.
ollama run hf.co/MainStack/marvy-2-35B-MoE-GGUF:Q4_K_M
llama-cli -hf MainStack/marvy-2-35B-MoE-GGUF:Q4_K_M \
-p "Write a ServiceNow user story with acceptance criteria for P1 SLA escalation." \
--temp 0.4 \
-c 4096
Search the model catalog for marvy-2-35B-MoE-GGUF and download, or load
the local .gguf file via "Open Model in Folder". LM Studio's OpenAI-compatible
server is at http://localhost:1234/v1 by default.
The qwen3_5_moe architecture (with mixed SSM layers and Multi-Token
Prediction in the base) is new. Verify your runtime supports it:
Qwen3_5MoeForConditionalGeneration registration in conversion/qwen.py).
The Homebrew formula may lag — clone upstream if you hit
"unknown architecture" errors.EVAL.md in the repo root for per-task perplexity.LoRA adapter (rank 32, 350 steps, attention-only Q/K/V/O)
+
bf16 base (unsloth/Qwen3.6-35B-A3B)
│ mlx_lm fuse
merged-bf16/ (14 shards, 65 GB safetensors with mlx_lm switch_mlp naming)
│ scripts/marvy-v2-rename-moe-tensors.py (bridge to HF-canonical names)
merged-bf16-hf/ (15 shards, 65 GB; switch_mlp → experts.gate_up_proj packed)
│ llama.cpp/convert_hf_to_gguf.py --no-mtp --outtype f16
marvy-2-35B-MoE-F16.gguf (65 GB, 733 tensors)
│ llama-quantize
marvy-2-35B-MoE-Q4_K_M.gguf (20 GB, 4.6 BPW)
marvy-2-35B-MoE-Q8_0.gguf (34 GB, 8.52 BPW)
End-to-end build script: scripts/marvy-v2-35B-MoE-build-gguf.sh (in the
source repo). The switch_mlp → experts rename is necessary because mlx_lm
fuse and llama.cpp's converter use different naming conventions for routed
MoE tensors.
Apache-2.0 (inherits from Qwen). See LICENSE and NOTICE in the repo root.
@misc{marvy-2-35B-MoE,
title = {marvy-2: A ServiceNow delivery-lifecycle MoE LLM},
author = {MainStack},
year = {2026},
howpublished = {\url{https://huggingface.co/MainStack/marvy-2-35B-MoE-GGUF}}
}