Downloads · 30 days
4.9K
19% of all-time downloads
reaperdoesntknow/TameForCasualLM
TameForCasualLM is a text generation model from reaperdoesntknow. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
4.9K
19% of all-time downloads
All-time downloads
26K
Public
Repo size
4.3 GB
Likes
1
Public
Click a slice to open those files.
.bin4.3 GB · 100%
From the Hugging Face model README
With Blackhole Rope Dynamics
This model builds on the original 421M TAMELM-AFMoER by introducing the Blackhole Rope (BHR) mechanism—a dynamic field-based routing system designed to stabilize, amplify, and concentrate information flow across multiple temporal scales.
While the original AFMoER established efficiency in routing-based intelligence, the BHR variant explores how structured gravitational-like attractors can further enhance reasoning depth without exponential increases in computation or parameters.
The Blackhole Rope is a symplectic, multiscale vortex mechanism inside AFMoER that:
Metaphorically:
If AFMoER routes are like neuronal pathways, the Blackhole Rope is the myelinated tether that keeps them from dispersing into noise—while also letting them “fall deeper” into coherent reasoning attractors.
Training Regime:
O1-OPEN/OpenO1-SFT (~500k tokens)WeMake/Intelligent-Content-Understanding (~500k tokens)Blackhole Rope Stabilization
Adaptive Vortex Dynamics
Energy Amplification Without Instability
TAMELM-AFMoER (1B) employs 16 experts under sparse routing. Typically, a full forward pass engages 4 experts per step, giving the model partial but diverse exposure on each pass.
The result is a model where expert specialization unfolds in phases, guided by both the Blackhole Rope stabilization and routing entropy regularization.
This model sets the stage for:
This model is part of the Convergent Intelligence LLC: Research Division portfolio. All models in this portfolio are developed under the Discrepancy Calculus (DISC) framework — a measure-theoretic approach to understanding and controlling the gap between what a model should produce and what it actually produces.
DISC treats training singularities (loss plateaus, mode collapse, catastrophic forgetting) not as failures to be smoothed over, but as structural signals that reveal the geometry of the learning problem. Key concepts:
For the full mathematical treatment, see Discrepancy Calculus: Foundations and Core Theory (DOI: 10.57967/hf/8194).
Citation chain: Structure Over Scale (DOI: 10.57967/hf/8165) → Three Teachers to Dual Cognition (DOI: 10.57967/hf/8184) → Discrepancy Calculus (DOI: 10.57967/hf/8194)
If you use this model, please cite:
Roy C, R. S. (2025). TAMELM-AFMoER with Blackhole Rope: Efficient Cognitive Emergence via Symplectic Routing.
The mathematics behind the model can be found: https://www.researchgate.net/publication/395539824_Negative-Space_Mathematics_A_New_Approach_for_Geometric_Computation_in_the_All-Negative_Orthant
I am the creator of the math and the models.
By Convergent Intelligence LLC: Research Division
| Model | Downloads |
|---|---|
| Qwen3-1.7B-Thinking-Distil | 501 |
| LFM2.5-1.2B-Distilled-SFT | 342 |
| Qwen3-1.7B-Coder-Distilled-SFT | 302 |
| Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT-GGUF | 203 |
| Qwen3-1.7B-Coder-Distilled-SFT-GGUF | 194 |
Total Portfolio: 41 models | 2,781 total downloads
Last updated: 2026-03-28 12:58 UTC
<!-- CIX-CROSSLINK-START -->DistilQwen Collection — Our only BF16 series. Proof-weighted distillation from Qwen3-30B-A3B → 1.7B and 0.6B on H100. Three teacher variants (Instruct, Thinking, Coder), nine models, 2,788 combined downloads. The rest of the portfolio proves structure beats scale on CPU. This collection shows what happens when you give the methodology real hardware.
Top model: Qwen3-1.7B-Coder-Distilled-SFT — 508 downloads
Full methodology: Structure Over Scale (DOI: 10.57967/hf/8165)
Convergent Intelligence LLC: Research Division
<!-- CIX-CROSSLINK-END --> <!-- cix-keeper-ts:2026-10-03T13:17:00Z -->