Downloads · 30 days
195
4% of all-time downloads
guidelabs/steerling-8b
steerling-8b is a text generation model from guidelabs. Use it when you need the model to write or continue text. It is set up for steerling. The card lists the license as apache-2.0.
An interpretable causal diffusion language model with concept steering.
Downloads · 30 days
195
4% of all-time downloads
All-time downloads
4.8K
Public
Parameters
8.4B
33.6 GB on disk
Likes
118
Public
Click a slice to open those files.
.safetensors16.8 GB · 100%
From the Hugging Face model README
An interpretable causal diffusion language model with concept steering.
Paper: Scaling Inherently Interpretable Language Models
Code: github.com/guidelabs/steerling
Project page: guidelabs.ai/papers/scaling-inherently-interpretable-language-models
Steerling-8B is an 8 billion parameter language model that combines masked diffusion with interpretable concept decomposition. Unlike standard autoregressive LLMs, Steerling generates text by iteratively unmasking tokens in order of confidence, and decomposes its internal representations into human-interpretable concepts that can be inspected and steered.
pip install steerling
from steerling import SteerlingGenerator, GenerationConfig
generator = SteerlingGenerator.from_pretrained("guidelabs/steerling-8b")
text = generator.generate(
"The key to understanding neural networks is",
GenerationConfig(max_new_tokens=100, seed=42),
)
print(text)
| Property | Value |
|---|---|
| Parameters | 8.4B |
| Architecture | CausalDiffusionLM + iGuide |
| Context Length | 4,096 |
| Vocabulary | 100,281 (cl100k_base + specials) |
| Known Concepts | 33,732 |
| Unknown Concepts | 101,196 |
| GQA | 32 heads, 4 KV heads |
| Diff Block Size | 64 |
| Precision | bfloat16 |
| VRAM Required | ~18GB |
Steerling uses block-causal attention, bidirectional within a block, and causal across blocks. The interpretable concept heads decompose transformer hidden states into:
hidden → known_features + unknown_features + epsilon = composed → logits
| Dataset | License | Stage |
|---|---|---|
| Nemotron-CC-HQ (real + synthetic) | NVIDIA Data Agreement | Pretraining |
| Dolmino Mix (math) | ODC-By v1.0 | Midtraining |
The Nemotron-CC dataset includes synthetic data generated by third-party models (Qwen, DeepSeek). Users should review the applicable license terms for their intended use case.
| Setup | Works? |
|---|---|
| A100 80GB | ✅ |
| A100 40GB | ✅ |
| A6000 48GB | ✅ |
| RTX 4090 24GB | ✅ |
| RTX 3090 24GB | ✅ |
| 16GB or less | ❌ |
The Steerling source code and model weights are released under the Apache License 2.0.
The model weights are provided for research and evaluation purposes. The weights were trained on datasets with varying license terms, including Nemotron-CC-HQ and Dolmino Mix. Some training data includes synthetic content generated by third-party models with their own license terms. We are currently reviewing the implications of these upstream licenses for downstream use of the model weights. Please check back for updates on the weight licensing terms.
For questions about commercial use of the model weights, contact us at [email protected]