Downloads · 30 days
687
28% of all-time downloads
Subject-Emu-5259/NeuralAI-Mamba-K1
NeuralAI-Mamba-K1 is a text generation model from Subject-Emu-5259. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
NeuralAI — Mamba K1 Model card maintained by De'Andrew Preston Harris (@Subject-Emu-5259) Last synced: 2026-08-13 --
Downloads · 30 days
687
28% of all-time downloads
All-time downloads
2.5K
Public
Parameters
129M
2.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.gguf442 MB · 63%
From the Hugging Face model README
| Property | Value |
|---|---|
| Architecture | Mamba SSM — model_type: mamba |
| Class | MambaForCausalLM |
| Parameters | ~130M (hidden size 768, 24 layers) |
| State size | 16 |
| Vocabulary | 50,280 |
| Base model | state-spaces/mamba-130m-hf |
| Fine-tune method | LoRA SFT, vocabulary-safe chat format |
| LoRA config | rank 16, alpha 32 (iterative v2/v3) |
| Dataset | NeuralAI seed set — assistant conversations spanning reasoning, code, math, writing, safety, and creative prompts |
| Training runtime | CPU/GPU SFT loops; iterative GGUF merge + quantization |
| Formats in this repo | Merged safetensors · Q4_K_M GGUF · F16 GGUF |
| Status | 🔬 R&D / chat-format repair for future release |
| License | Apache 2.0 |
Mamba K1 is the first model NeuralAI owns end-to-end. Unlike adapters on a third-party transformer, this model starts from a base Mamba SSM architecture and is trained, merged, and quantized into a self-contained artifact.
Mamba SSMs replace quadratic self-attention with a linear, state-space recurrence. That makes them fast at long context and cheap to serve — ideal for a local-first assistant that runs on modest hardware.
mambapy so the model loads without custom CUDA| Model | Architecture | Parameters | Role | Status |
|---|---|---|---|---|
| 🧬 Mamba K1 | Mamba SSM | 130M | NeuralAI's first owned base | 🔬 R&D |
| 🧠 NeuralAI Powered by SmolLM2‑360M | Transformer + LoRA | 360M | Live chat backend | ⚡ Active |
The SFT curriculum taught the model to behave like an assistant across a deliberately small but diverse seed set:
⚠️ Scale note: 130M parameters is a research-capability checkpoint, not yet frontier-grade. K1 is the starting point for a fully owned NeuralAI model lineage.
| Phase | Detail |
|---|---|
| Data | Curated assistant seed set (reasoning, code, math, writing, safety, creative) |
| Objective | SFT on assistant-style completions |
| Method | LoRA SFT → merge → GGUF quantization |
| Chat format | NeuralAI "intel" format — uses only tokens present in the GPT-NeoX tokenizer (`< |
| Final train loss (checkpoint) | 11.69 (down from ~13.5) |
| Output formats | Merged safetensors, Q4_K_M GGUF, F16 GGUF |
Full training logs, merge scripts, and the iteration runbook live in the main NeuralAI repository.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"Subject-Emu-5259/NeuralAI-Mamba-K1",
torch_dtype=torch.float32,
trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(
"Subject-Emu-5259/NeuralAI-Mamba-K1",
trust_remote_code=True,
)
messages = [{"role": "user", "content": "Write a haiku about debugging."}]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
)
out = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
Use the Q4_K_M or F16 GGUF in this repo:
./llama-server \
--model neuralai-mamba-k1-v3.Q4_K_M.gguf \
--chat-format neuralai-intel \
--port 1234
The NeuralAI model manager can point a local llama.cpp backend at this GGUF when K1 becomes the active inference target.
NeuralAI is a local-first, private generative AI engine built by De'Andrew Preston Harris. It is the central intelligence layer of an ecosystem that includes:
The mission is simple: your AI, on your hardware, under your control.
NeuralAI was built from resilience, fatherhood, and the belief that personal computing deserves personal intelligence. Every release is handcrafted, iterated, and documented in the open.
| Project / Brand | NeuralAI |
| Motto | Your AI. On your hardware. In your browser. |
| Values | Privacy, ownership, local-first computing, disciplined iteration, open weights |
| Primary Repository | github.com/Subject-Emu-5259/NeuralAI |
| Model Collection | huggingface.co/Subject-Emu-5259 |
| License | Apache 2.0 |
NeuralAI is not a closed SaaS product. It is a living open-weights research project becoming a sustainable AI software company built by one determined builder and the community around him.
Right now NeuralAI's live chat backend is the awareness-tuned SmolLM2-360M model. While K1 matures, that model handles everyday assistant tasks:
<p align="center"> <img src="neuralai-model-comparison.png" alt="NeuralAI model comparison" width="92%" /> </p>| Resource | Link |
|---|---|
| Main repository | github.com/Subject-Emu-5259/NeuralAI |
| Active chat model | huggingface.co/Subject-Emu-5259/NeuralAI-Powered-By-SmolLM2360 |
| Creator LinkedIn | linkedin.com/in/deandrewharris94 |
@software{neuralai_mamba_k1_2026,
author = {Harris, De'Andrew Preston},
title = {NeuralAI — Mamba K1},
year = {2026},
url = {https://huggingface.co/Subject-Emu-5259/NeuralAI-Mamba-K1},
version = {v3},
description = {NeuralAI's first owned Mamba SSM base model (130M) for local-first AI research}
}
Built with discipline by De'Andrew Preston Harris. Maintained in the open. Updated whenever the model, dataset, or project state changes.