Downloads · 30 days
29
30% of all-time downloads
RtaForge/Mistral-Mamba3-7B
Mistral-Mamba3-7B is a text generation model from RtaForge. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as apache-2.0.
Cross-architecture Subsuminator heist: mistralai/Mistral-7B-Instruct-v0.3 → Mamba3-7B SSM body.
Downloads · 30 days
29
30% of all-time downloads
All-time downloads
96
Public
Parameters
3.6B
7.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors7.2 GB · 100%
From the Hugging Face model README
Cross-architecture Subsuminator heist: mistralai/Mistral-7B-Instruct-v0.3 → Mamba3-7B SSM body.
| Metric | Value |
|---|---|
| Protocol | subsuminator/subsume.py CE gate (random tokens, shifted CE) |
| Source CE | 13.3000 |
| Heisted CE | 10.4125 |
| log(vocab) baseline | 10.3972 |
| CE ratio vs source | 0.7829× |
Random-token next-step CE on mistralai/Mistral-7B-Instruct-v0.3 vs this checkpoint (seed=42, n=5, seq_len=32).
Interpretation: Heisted CE ≈ log(vocab) (10.41 vs 10.40) — the fresh SSM body behaves like a random baseline on uniform tokens. Ratio 0.7829× vs source (10.41 / 13.30) is below 1.0 because Mistral's trained transformer raises CE on garbage random inputs; this is not capability preservation. For comparison, Mamba2→Mamba3 structural port (trained→trained) achieved 1.0016× with CE near source.
in_proj, out_proj, dt_bias, state dynamics)Alpha checkpoint — fine-tune before production use.
Sovereign local inference for mamba2 + mamba3 only:
--yeehaw, --arnie, --giddyup)trust_remote_code for HF load or Avocado native splat path./avocado run --model RtaForge/Mistral-Mamba3-7B --prompt "Come with me if you want to live."
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("RtaForge/Mistral-Mamba3-7B", trust_remote_code=True)
tok = AutoTokenizer.from_pretrained("RtaForge/Mistral-Mamba3-7B")