Downloads · 30 days
346
38% of all-time downloads
pipenetwork/LongCat-2.0-2bit
LongCat-2.0-2bit is a text generation model from pipenetwork. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as mit.
2-bit (2.501 bits/weight) MLX quantization of meituan-longcat/LongCat-2.0, a 1.6T-parameter / ~48B-active MoE (MLA attention + LongCat sparse-attention indexer + identity experts + n-gram embeddings). Converted from t…
Downloads · 30 days
346
38% of all-time downloads
All-time downloads
921
Public
Parameters
1.6T
512 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors512 GB · 100%
How the weights are stored.
U321.6T · 100%
From the Hugging Face model README
2-bit (2.501 bits/weight) MLX quantization of
meituan-longcat/LongCat-2.0,
a 1.6T-parameter / ~48B-active MoE (MLA attention + LongCat sparse-attention indexer +
identity experts + n-gram embeddings). Converted from the FP8 source with mlx-lm.
Router classifiers are kept at 8-bit (mixed precision); MTP layers are dropped.
Size: ~477 GB. This exceeds a 512 GB unified-memory ceiling in practice — intended for larger-memory or sharded/multi-node MLX inference, not a single 512 GB machine.
LongCat-2.0 (model_type: longcat2) support is not yet in a released mlx-lm. Install from
the PR branch:
pip install git+https://github.com/ml-explore/mlx-lm.git@refs/pull/1464/head
from mlx_lm import load, generate
model, tokenizer = load("pipenetwork/LongCat-2.0-2bit")
messages = [{"role": "user", "content": "Who is Albert Einstein?"}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
print(generate(model, tokenizer, prompt=prompt, max_tokens=512, verbose=True))
For large builds, use sharded/distributed generation (mlx.launch + sharded_generate.py).