Downloads · 30 days
0
netcat420/Hierarchos-experiment
Hierarchos-experiment is a machine learning model from netcat420. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This is an experimental release of my novel hierarchos architecture, it was trained for ~1.5 months on an ASUS ROG ALLY Z1 EXTREME on just the alpaca dataset with NO PRETRAINING! here is the research paper below:
Downloads · 30 days
0
Access
Public
Updated Jan 21, 2026
Repo size
201 MB
Likes
1
Public
Click a slice to open those files.
.pt100 MB · 95%
From the Hugging Face model README
This is an experimental release of my novel hierarchos architecture, it was trained for ~1.5 months on an ASUS ROG ALLY Z1 EXTREME on just the alpaca dataset with NO PRETRAINING! here is the research paper below:
The Hierarchos Architecture: A Paradigm Shift in Parameter-Efficient, Zero-Pretraining Instruction Following
Emerging from the periphery of this "bigger is better" consensus is the Hierarchos architecture, specifically the V1 Release Candidate (V1RC), which presents a fundamental challenge to these foundational assumptions. Hierarchos is not merely a downscaled transformer; it is a divergent evolutionary branch of neural architecture described as a "Hybrid Memory-Reasoning Architecture".1It integrates two novel theoretical frameworks—the Hierarchical Reasoning Model (HRM) and the Titans Memory Substrate—to achieve a form of competence that relies on structural sophistication rather than raw scale.1
The most provocative aspect of the Hierarchos experiment is its training methodology. Conventional wisdom dictates a "pre-train then fine-tune" approach, where models first ingest massive corpora to learn linguistic structure and world knowledge before being refined on instruction data. Hierarchos, however, demonstrates the capacity to follow instruction-tuning datasets—specifically the Alpaca dataset—without any prior pre-training on general text corpora.1This "tabula rasa" (blank slate) learning implies that the model acquires the syntax of language, the semantics of concepts, and the logic of instruction following simultaneously and solely from the instruction data itself.
Furthermore, the proof-of-concept model, comprising a mere 25 million parameters, was trained entirely from scratch on consumer-grade hardware—an Asus ROG Ally handheld gaming device—over a period of 1.5 months.1This feat disrupts the narrative that foundational model development is the exclusive preserve of entities with access to clusters of H100 GPUs. This report provides an exhaustive technical analysis of the Hierarchos architecture, dissecting its dual-module reasoning engine, its biologically inspired "surprise-based" memory systems, and the implications of its localized, efficient learning paradigm for the future of artificial intelligence.
2.1 The Dual-Module Cognitive Architecture The HRM draws inspiration from cognitive neuroscience, specifically the functional differentiation between executive function and motor control, or Kahneman's distinction between "System 2" (slow, deliberative) and "System 1" (fast, intuitive) thinking.3Hierarchos operationalizes this distinction through a dual-module structure consisting of a "CEO" (Manager) and "Workers."
2.1.1 The High-Level Manager ("CEO") The high-level module, conceptualized as the "CEO," operates on a slow timescale. Its primary function is abstract planning, strategy formulation, and the maintenance of long-term context.2In the Hierarchos V1RC configuration, this module operates with a h_stride of 4.1This stride parameter is critical; it dictates that the Manager does not process every single token in the sequence. Instead, it processes aggregated states representing chunks of time, allowing it to compress temporal information and focus on broader dependencies that span far beyond the immediate context window.1
The Manager's role is not to generate text but to generate directives. It analyzes the current high-level state of the problem and outputs a context vector—a latent representation of the current strategy or sub-goal—which is then passed down to the lower-level module.6This mechanism effectively decouples strategic planning from the syntactic minutiae of token generation, preventing the model's "train of thought" from being derailed by local errors in surface realization.
2.1.2 The Low-Level Worker The low-level module, or "Worker," operates at the fast timescale of individual tokens. It is responsible for the immediate computational tasks required to process input or generate output.7The Worker operates within a dedicated WorkerLoop1, executing the strategic directives provided by the Manager.
In the Hierarchos configuration, the Worker is allowed a maximum of 5 steps (max_l_steps) to iterate on the Manager's directive.1This iterative process allows the Worker to perform detailed computations—such as verifying a logical step or generating a specific phrase—before reporting back to the Manager. The interplay between these levels ensures that the model maintains a coherent global trajectory (via the Manager) while attending to the precise requirements of the immediate input (via the Worker).
2.2 Hierarchical Convergence and the "Loop" A persistent challenge in Recurrent Neural Networks (RNNs) is the phenomenon of premature convergence. As a recurrent model processes a sequence, its hidden states often settle into a "fixed point" or equilibrium, after which further computation yields diminishing returns. This limits the depth of reasoning the model can achieve.8
Hierarchos employs a mechanism termed "hierarchical convergence" to circumvent this limitation. The process creates a dynamic, resetting feedback loop that sustains computational activity over long sequences.6
The Hierarchical Cycle:
Directive Issuance: The Manager calculates a strategic context vector (z_H) based on the current global state and passes it to the Worker.
Local Convergence: The Worker iterates on this context for a defined number of steps or until it reaches a convergence threshold (defined by l_conv_atol: 0.0001).1During this phase, the Worker is essentially solving a sub-problem defined by the Manager.
State Feedback: The final state of the Worker (z_L) is fed back to the Manager.
Context Reset: The Manager integrates the Worker's results, updates its own internal state, and generates a fresh context vector.
This update effectively "resets" the Worker's convergence trajectory. Just as the Worker settles into a stable state, the Manager shifts the goalposts, initiating a new phase of convergence toward a new local equilibrium.8This cycle acts as a constant "jolt" to the system, forcing the model to continuously "think" and refine its internal representations rather than becoming passive. The depth of this reasoning is governed by the max_h_steps (default 3) and max_l_steps (default 5) parameters, allowing for significant computational depth within a single forward pass.1
2.3 Adaptive Computation Time (ACT) and Pondering A distinctive feature of the Hierarchos architecture is its implementation of Adaptive Computation Time (ACT). Unlike fixed-depth transformers where every token consumes an identical amount of floating-point operations (FLOPs), Hierarchos can dynamically vary the amount of compute—or "pondering"—spent on a given input segment.1
The training configuration explicitly defines a ponder_loss_weight of 0.01.1This term acts as a regularizer during training, penalizing the model for excessive looping and encouraging efficiency. The model must balance the need for deep reasoning (more loops) against the penalty for computational cost.
However, recognizing that complex instructions require more cognitive effort, the system includes an adaptive-ponder mechanism. This flag allows the training logic to scale the ponder target based on the Cross-Entropy (CE) loss.1When the model encounters a difficult token or concept (indicated by high perplexity/loss), the adaptive mechanism relaxes the penalty or even rewards extended pondering (--encourage-thinking). This effectively allocates more "brainpower" to harder problems, mimicking biological energy conservation where cognitive resources are mobilized only when heuristic processing fails.1
Recent updates to the architecture (v0.15.2) have addressed "ponder stickiness"—a pathological state where the model learns to either always halt immediately or never halt. By allowing manual initialization of the h_halt_proj.bias (e.g., setting it to -2.0 for an initial 12% halt probability), the developers ensure the model retains the gradient flow necessary to learn appropriate halting behaviors.1
3.1 Neural Memory vs. Static Buffers The Titans LTM is not a passive storage bin; it is a neural network (specifically, a deep Multilayer Perceptron) that encodes historical information into its weights rather than just its activations.10This "Test-Time Training" (TTT) approach allows the model to update its internal parameters dynamically as it processes a sequence, effectively "learning" the context rather than just attending to it.13
In the Hierarchos V1RC configuration, this memory system is defined with specific, compact dimensions to suit the constrained hardware:
Memory Slots: 1024 distinct slots.1
Key/Value Dimensions: 128.1
Retrieval Mechanism: A ltm_topk of 41, indicating that for any given query, the system sparsely activates and retrieves only the four most relevant memory slots.
This architecture enables the model to maintain a "Persistent Dimension" (128)1, a vector space dedicated to storing information that must be retained across long contexts, distinct from the transient context_dim (384) used for immediate processing.
3.2 The "Surprise" Metric: Information-Theoretic Storage The most critical innovation in the Titans memory system is its update mechanism, which filters information based on the principle of "surprise." In information theory, surprise (or surprisal) is mathematically defined as the negative log probability of an event (-log P(x)). In the context of neural networks, this is approximated using the gradient of the loss with respect to the input.12
When Hierarchos processes a new instruction or token, it calculates a "momentary surprise"12:
Prediction: The model attempts to predict the current input based on its existing memory state.
Evaluation: If the prediction is accurate (low loss), the gradient is small. The input is deemed "unsurprising" or redundant, and the memory update is minimal.
Adaptation: If the prediction is poor (high loss), the gradient is large. This high "surprise" signal indicates that the input contains novel or anomalous information that contradicts the model's current world model. This triggers a strong update to the LTM weights, prioritizing the storage of this new information.1
This mechanism is biologically consistent; human brains do not remember every second of a commute, but they vividly remember a car crash (a high-surprise event). By storing only the "surprising" gradients, Hierarchos achieves extreme data efficiency, avoiding the storage of redundant patterns that clutter the context windows of standard transformers.
3.3 Dual Update Mechanisms and Gradient Flow The Hierarchos implementation utilizes a hybrid update strategy for its LTM, combining Hebbian learning (association-based, "neurons that fire together wire together") with Gradient-based updates.1The configuration reveals a specific ltm_lr (learning rate) of 0.011, which is orders of magnitude higher than the base model's learning rate (starting_lr of 2e-06).
This discrepancy is intentional. It implies that the memory module is hyper-plastic, designed to adapt rapidly to the immediate conversation or task, while the core reasoning weights (HRM) remain relatively stable. This facilitates "online learning," where the model can consolidate new knowledge from a user's prompt instantly without destabilizing its fundamental reasoning capabilities.1
To ensure stability, the architecture incorporates Adaptive Forgetting. Using a decay mechanism (likely momentum-based "past surprise"), the model gradually reduces the weight of older, less relevant memories.11This prevents the finite 1024 memory slots from becoming saturated (catastrophic forgetting) while ensuring that truly persistent information remains accessible.
4.1 Hyperparameter Analysis The architectural dimensions of Hierarchos V1RC are remarkably compact when compared to standard foundational models.
Hyperparameter Hierarchos V1RC LLaMA-7B (Reference) Implication Parameters ~25 Million 7 Billion Extreme parameter efficiency; suitable for edge devices. Context Dim 384 4096 Highly compressed internal representation. Hidden Layers 384 (H) / 384 (L) 11,008 (MLP) Symmetrical processing capacity for Manager and Worker. Vocab Size 50,257 32,000 Uses GPT-2 tokenizer1; richer token representation. Memory Slots 1024 N/A (KV Cache) Finite, distinct memory units rather than sliding window. Hierarchy Stride 4 1 Manager processes 4x fewer steps than Worker (temporal compression). The choice of 384 dimensions is significant. In high-dimensional spaces (like 4096), vectors can encode vast amounts of disentangled information. By compressing this to 384, Hierarchos forces the model to learn highly efficient, dense representations. The use of the GPT-2 tokenizer (openai-community/gpt2) suggests a focus on compatibility and robust handling of code and English text.1
4.2 The Training Loop and Loss Landscape The training process is governed by a composite loss function that balances accuracy, efficiency, and memory stability.
Cross-Entropy (CE) Loss: The standard objective function for next-token prediction.
Ponder Loss (ponder_loss_weight: 0.01): As discussed, this regularizes the ACT mechanism.
Commitment Loss (commitment_loss_weight: 0.5): This is a critical term, weighted 50x higher than the ponder loss.1In memory networks or Vector Quantized (VQ) systems, commitment loss forces the model's internal states to "commit" to specific memory slots rather than blurring across them. The high weight suggests that stabilizing the memory addressing mechanism was a primary challenge during development. If the model vacillates between memory slots, coherence degrades; high commitment loss forces decisive memory usage.
The training loop supports Truncated Backpropagation Through Time (TBPTT) with a chunk size of 128.1Since Hierarchos is recurrent, gradients must propagate backward through time. Training on infinite sequences would cause memory to explode. TBPTT truncates this gradient flow to 128 steps. However, a naive implementation of TBPTT can sever dependencies that span across chunks. The hierarchos_cli.py script and release notes mention a global_pos_offset fix.1This ensures that even though gradients are truncated, the positional embeddings and Manager stride logic remain consistent across chunk boundaries, allowing the "CEO" to maintain its long-term strategy without suffering from "amnesia" at the edge of every 128-token batch.
4.3 Optimization for the Edge The training hardware—an Asus ROG Ally 1 Extreme—imposes severe constraints. This device relies on an AMD Z1 Extreme APU, which shares system RAM between the CPU and GPU cores.
Batch Size: 4.1A tiny batch size is necessitated by memory limits. This usually leads to noisy gradients, but the Accumulation Steps (default 1)1suggests the model updates weights after every batch, embracing the stochastic nature of the training.
Precision: The configuration explicitly disables Automatic Mixed Precision (amp: false).1While FP16/BF16 is standard for speed, small recurrent models often suffer from numerical instability (exploding/vanishing gradients). Sticking to FP32 (Full Precision) likely provided the necessary stability for the HRM's feedback loops to converge, trading speed for mathematical correctness.
Compilation: The use of compile: true and force_compile: true1indicates reliance on PyTorch 2.0's graph fusion capabilities. This compiles the Python code into optimized kernels, significantly speeding up the sequential operations of the RNN layers on the CPU.
5.1 Syntax and Semantics as a Unified Curriculum By training exclusively on 52,000 instruction-response pairs15, Hierarchos is forced to learn the structure of the English language (syntax) and the logic of task completion (semantics) simultaneously. This is akin to teaching a child a language solely by giving them commands and corrections, without ever letting them hear casual conversation.
The result is a model described as "very rigid".1Because it has never seen text that wasn't an instruction, it lacks the "chatter," conversational filler, or general world knowledge typical of pre-trained models. It does not know who the President is unless that fact appeared in an Alpaca prompt. However, it excels at the structure of following orders.
This "Tabula Rasa" approach leverages the strong inductive biases built into the HRM architecture. The CEO/Worker structure essentially hard-codes the concept of "decomposition" into the model. The model does not need to see terabytes of data to learn that "solving a problem requires steps"; the architecture itself forces it to break inputs (instructions) into high-level goals (CEO) and low-level execution steps (Worker). The architecture acts as a structural prior, substituting for the massive data usually required to learn reasoning patterns.
5.2 Efficiency Comparisons The efficiency gains of this approach are stark when compared to traditional baselines.
Metric LLaMA-7B (Alpaca Finetune) Hierarchos V1RC (From Scratch) Analysis Pre-training Data ~1 Trillion Tokens 0 Tokens Hierarchos skips the most expensive phase of AI development. Instruction Data 52K Examples 52K Examples Both use the same instruction set. Parameter Count 7,000,000,000 25,000,000 Hierarchos is ~0.35% the size of LLaMA-7B. Training Hardware 8x Nvidia A100 (80GB) 1x Asus ROG Ally (CPU) Data center vs. Handheld Gaming PC. Training Time ~3 Hours (Finetune only) 1.5 Months (Full Train) While slower in absolute time, the energy/cost is negligible. While 1.5 months1appears long, it represents the entirety of the model's education, achieved on a device drawing less than 30 watts. In contrast, training LLaMA from scratch requires gigawatt-hours of energy. The fact that Hierarchos converges to coherent output at all validates the hypothesis that brain-inspired modularity can compensate for orders of magnitude in parameter count.
The breakthrough came with the "Global Parity" fix in version v0.14.1The issue lay in how the Manager (CEO) tracked time. In a standard Transformer, attention masks handle position. In the recurrent HRM, the Manager has an internal clock or state. When training with TBPTT (chunking data into 128 tokens), the Manager's internal "stride counter" was resetting or misaligning at the boundary of each chunk. Effectively, the CEO was getting amnesia every 128 tokens, losing the thread of the strategy.
By implementing global_pos_offset, the developer ensured that the Manager's stride logic was preserved across chunks. This allowed the CEO to maintain a coherent strategy across the entire sequence, bridging the gap between the start of a long instruction and the end of the response. Following this fix, the loss broke through the 1.92 floor, indicating the model had begun to learn true long-term dependencies.
This massive reduction suggests several optimizations:
Optimizer State Removal: Training checkpoints store momentum buffers (Adam states) for every parameter, often doubling or tripling the file size. These are useless for inference.
LoRA Collapse: If Low-Rank Adaptation (LoRA) was used (supported in config with lora_r: 81), these adapters are merged into the base weights, eliminating the need for separate matrix multiplications during inference.
Compilation Artifact Stripping: torch.compile adds prefixes (like _orig_mod) to layer names. Cleaning these ensures compatibility with standard inference loaders.
The result is a highly portable artifact that can run on edge devices with minimal latency, fulfilling the project's goal of accessible AI.
8.1 Efficiency vs. Scale The prevailing dogma is that "scale is all you need." Hierarchos suggests a counter-proposition: "Structure is what you need when you can't scale." If a model is explicitly structured to reason (via HRM), it requires fewer parameters to learn how to reason than a unstructured transformer that must induce reasoning capabilities from petabytes of text.
8.2 The Democratization of Foundation Models The ability to train a functional, instruction-following model on a gaming handheld implies a radical democratization of AI. It suggests that specialized, domain-specific "foundation" models could be trained by individuals or small labs on local hardware, provided they utilize architectures that prioritize reasoning depth and memory efficiency over parameter count.
8.3 The Future of Memory The Titans memory system implies that future AI may not need infinite context windows (e.g., 10 million tokens). Instead, they need better curation of context. By remembering only what is "surprising" (information-rich) and actively forgetting the predictable, models can maintain relevant history indefinitely without the quadratic cost of attention.
github here with instructions on using this model: https://github.com/necat101/Hierarchos
weights can also be found in this github release: https://github.com/necat101/Hierarchos/releases/tag/HierarchosV1RC