Downloads · 30 days
0
nanohit/nocturne
nocturne is a machine learning model from nanohit. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Nocturne is not a dense 12B model in practical runtime terms.
Downloads · 30 days
0
Access
Public
Updated Feb 8, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md1.9 KB · 56%
From the Hugging Face model README
Nocturne is not a dense 12B model in practical runtime terms.
It combines a standard Transformer stack (multi-layer self-attention + feedforward) with sparse MoE layers: the model has many experts but only a small number of experts are routed-to (active) per token, so the active parameter footprint per forward pass is ~2.02B rather than 12B.
That sparse activation is what lets us actually claim “12B parameters” on disk but low memory at inference. It uses a top-k gating function that selects 2–4 experts per token and routes token activations to those experts; 32 total experts and 4 active experts per token for the 12B model.
Practical engineering choices are:
The model supports extremely long contexts (up to 64k tokens), achieved thought careful attention memory and positional strategy.
Weights come natively quantized in a MXFP4 format which is designed to preserve accuracy while drastically lowering memory use. Trained it on a mostly-English, text-only dataset with an emphasis on STEM, coding, and general knowledge; tokenization uses a new tokenizer by OpenAI called o200k_harmony (a superset of the o4/o4-style tokenizers). after base pretraining, the model was post-trained on a structured format OpenAI calls.
license: mit datasets: