Downloads · 30 days
0
Daxamite/V12Emoe
V12Emoe is a machine learning model from Daxamite. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
Downloads · 30 days
0
Access
Public
Updated Jun 30, 2026
Repo size
5.5 GB
Likes
0
Public
Click a slice to open those files.
.pt1.9 GB · 100%
From the Hugging Face model README

eMoE is a complete serving system, not a single model. It mints a fresh, per-request adapter for a small frozen base model from a task-conditioning vector, runs it inside a deterministic control loop with out-of-distribution gating and retrieval grounding, and discards the adapter when the request is done. Nothing about the base changes; only a 9.20M hypernetwork is trained.
The bet: most of what a small model lacks on a given request is specialization and grounding, not raw capacity. eMoE buys both cheaply — a per-request expert plus distance-scaled RAG — on a base that stays frozen and reusable. It explicitly does not buy reasoning depth, and the design is honest about that.
1. An owned, frozen base. A 190M FFT-hybrid (12 layers, 1024d, 16 heads, attn_every=3 → 4 attention + 8 spectral layers, 2048 context), trained with Muon+AdamW on a curated mix of synthetic tool-use and proxy-filtered real data. It reached VAL 1.6609 and is frozen for all serving. Spectral layers make adaptation cheap: you steer the envelope/gain and the kernel-generating MLP rather than rewriting dense weights.
2. A hypernetwork (9.20M params). It reads a task vector z — a BAAI/bge-base-en-v1.5 embedding (d_z 768) of the request descriptor — and emits adapter deltas for 64 sites: FiLM on the spectral envelope/gain and LoRA on the kernel-gen MLP. There is no materialized ΔW; deltas are applied in-place and scaled by a global α. Shape-grouped, bottlenecked heads keep the whole hypernetwork at 9.20M despite covering 64 sites. It is the only trained component (AdamW, rank 8, ChatML assistant-turn supervision), reaching held-out VAL ≈ 2.84 on a 1769-task / 1504-cluster set (265 held out).
3. A deterministic runtime harness. Fixed code with exposed knobs — no agentic planning, no self-critique, no LLM-as-judge routing. The model generates in exactly two places (an optional descriptor rewrite, and the answer); everything else is deterministic:
tau_rag_on = 0.287) and expert count k (tau_k_escalate = 0.308). Distance is treated as novelty, not trust — it never scales the adapter.On the 265 held-out tasks, comparing the right-z adapter against the bare frozen base on identical tokens:
This is the honest shape of the system: cheap, real specialization on most requests; a meaningful minority where the verifier and escalation loop earn their keep; grounding via RAG for novel requests.
ckpt_v12_190m_best.pt — the frozen 190M FFT-hybrid base (canonical)hyper_ckpt_v12.best.pt — the trained 9.20M hypernetwork (use this; not the end-of-run checkpoint)hyper/z_cache.pt — BGE z-cache the controller geometry is calibrated againsthyper_ckpt.bestMidRun7750.pt — earlier 612-task hypernetwork (reference only)model_hybrid.py, muon.py, tok_v9.py — model code + tokenizerTasks and RAG artifacts live in the dataset repo Daxamite/eMOE-rag; the live demo (full harness, routing trace) runs at the Space Daxamite/eMOE.
An earlier draft of these notes quoted a VAL of ~2.09 — that belonged to a retired 1142-task run and does not describe this model. The correct held-out floor for hyper_ckpt_v12.best.pt is ≈2.84, consistent with the project's historical floor and verified by an independent help/shatter eval (mean adapted loss 2.86). The controller thresholds (0.287 / 0.308) were re-banked on this model's own z-geometry.
License: other (see repo). The base is an owned, curated model; the hypernetwork and harness are the eMoE contribution.