Downloads · 30 days
23
48% of all-time downloads
Asilarkness/llm-150m
llm-150m is a text generation model from Asilarkness. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
An experimental bilingual dialogue/reasoning language model with approximately 157M parameters. It keeps the faithful CubicV7 differentiable depth-memory path, uses hierarchical cosine retrieval in the two penultimate…
Downloads · 30 days
23
48% of all-time downloads
All-time downloads
48
Public
Parameters
157M
943 MB on disk
Likes
0
Public
Click a slice to open those files.
.pt628 MB · 66%
From the Hugging Face model README
An experimental bilingual dialogue/reasoning language model with approximately 157M parameters. It keeps the faithful CubicV7 differentiable depth-memory path, uses hierarchical cosine retrieval in the two penultimate attention layers, and finishes with full global causal attention. Every sequence head uses QK RMSNorm.
The optimizer is the project's hybrid orthogonalized Muon (large matrices and
a boosted depth group) plus AdamW for embeddings, norms, scalars and gates.
Base training uses warmup-stable-decay and a one-token-ahead MTP auxiliary loss;
embeddings, RMSNorm scales and controls are excluded from AdamW decay.
SFT teaches explicit direct and <think>...</think> system-prompt modes; set
CUBIC_REASONING=1 in chat mode to request the reasoning format.
This is a custom PyTorch architecture, not a drop-in Transformers model. Run
train_and_chat.py with CUBIC_MODE=chat; it downloads/loads all required
files and starts an interactive console. The model is small and experimental:
verify factual, safety-critical and mathematical answers independently.
Each source dataset retains its own license/terms. FineWeb corpora inherit the Common Crawl terms described on their dataset cards; SmolTalk, Aya and OpenR1 are Apache-2.0; HH-RLHF and UltraFeedback use their published dataset licenses.