Downloads · 30 days
9
7% of all-time downloads
dreamwar/HRM-Text1-C4-large
HRM-Text1-C4-large is a machine learning model from dreamwar. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
[](https://colab.research.google.com/drive/1c4exU-zMt4SuT1kRlwQQXlLPaiazEDCf?usp=sharing)
Downloads · 30 days
9
7% of all-time downloads
All-time downloads
136
Public
Parameters
99.8M
400 MB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors399 MB · 100%
How the weights are stored.
F3299.8M · 100%
From the Hugging Face model README
A large-scale transformer model with Hierarchical Reasoning Module (HRM) architecture trained on multiple high-quality text datasets. This model features adaptive computation with pondering mechanisms for improved text generation quality.
HRM-Text1 implements a novel hierarchical reasoning architecture with the following key components:
The model supports training on multiple high-quality datasets:
The training script supports custom dataset mixing ratios:
CUSTOM_MIX_RATIOS = {
"high_quality": {
"slimpajama_en": 0.5, # 50% SlimPajama English
"pile": 0.3, # 30% The Pile
"openwebtext": 0.2 # 20% OpenWebText
}
}
class HRMBlock(nn.Module):
def __init__(self, n_embd, n_head, d_ff, dropout=0.1):
super().__init__()
self.norm1 = RMSNorm(n_embd)
self.attn = nn.MultiheadAttention(n_embd, n_head, dropout=dropout, batch_first=True)
self.norm2 = RMSNorm(n_embd)
self.mlp = SwiGLUMuchPelu(n_embd, d_ff, dropout)
self.dropout = nn.Dropout(dropout)
The model implements adaptive computation through a halt probability mechanism:
from transformers import T5Tokenizer
from modeling_hrm_text1 import HRMText1
# Load model and tokenizer
model = HRMText1.from_pretrained("dreamwar/HRM-Text1-{DATASET}-large")
tokenizer = T5Tokenizer.from_pretrained("t5-small")
# Generate text
prompt = "The future of artificial intelligence"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=50, temperature=0.7)
text = tokenizer.decode(outputs[0], skip_special_tokens=True)
Option 1: Google Colab (Recommended)
# Open the Colab notebook
https://colab.research.google.com/drive/1c4exU-zMt4SuT1kRlwQQXlLPaiazEDCf?usp=sharing
Option 2: Local Training
# Set environment variables
export HRM_OUTPUT_BASE="/path/to/output"
export HF_TOKEN="your_huggingface_token"
# Run training
python hrm_llm_training_c4_b.py
The training script supports extensive configuration:
# Dataset selection
ACTIVE_DATASET = "mixed" # Options: "c4", "openwebtext", "pile", "spanish", "mixed"
# Dataset subset percentage
DATASET_SUBSET_PERCENT = 5 # 1-100%
# Custom output path
CUSTOM_BASE_PATH = "/your/custom/path"
# Model parameters (large variant)
MODEL_PARAMS = {
"n_embd": 1024,
"n_head": 16,
"d_ff": 4096,
"dropout": 0.1,
"halt_max_steps": 8,
"ponder_loss_weight": 1e-2,
"halt_bias_init": -2.2
}
HRM_Models/
├── hrm_text1_{dataset}_output-large/
│ ├── config.json
│ ├── pytorch_model.bin
│ ├── tokenizer.json
│ ├── best_model.bin
│ └── checkpoint.pth
Click the Colab badge above to get started immediately with a pre-configured environment including all dependencies.
pip install torch transformers datasets tqdm huggingface_hub
pip install langdetect # Optional: for language filtering
# Required for model upload
export HF_TOKEN="your_huggingface_token"
# Optional: custom output path
export HRM_OUTPUT_BASE="/your/custom/path"
The training script produces several model variants:
The model implements sophisticated reasoning through:
This model and training code are released under the Apache 2.0 License.
@misc{hrm-text1-2024,
title={HRM-Text1: Hierarchical Reasoning Model for Text Generation},
author={DreamWar},
year={2024},
url={https://huggingface.co/dreamwar/HRM-Text1}
}
langdetect for language filteringFor issues and questions:
This model was trained using the HRM (Hierarchical Reasoning Module) architecture with adaptive computation for improved text generation capabilities.