Downloads · 30 days
2.5K
64% of all-time downloads
Prannesshkva/Ael-Edge-130M-MoE
Ael-Edge-130M-MoE is a text generation model from Prannesshkva. Use it when you need the model to write or continue text. The card lists the license as cc-by-nc-nd-4.0.
<p align="center" <a href="https://doi.org/10.5281/zenodo.22649142"<img src="https://zenodo.org/badge/DOI/10.5281/zenodo.22649142.svg" alt="DOI"</a <a href="https://www.linkedin.com/in/prannesshkva/"<img src="https://…
Downloads · 30 days
2.5K
64% of all-time downloads
All-time downloads
4K
Public
Parameters
135M
1.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors540 MB · 98%
From the Hugging Face model README
Ael-Edge-130M-MoE (formerly ISOM-R1-Edge-130M-MoE-Beta-Prototype) is the ultra-compact 134.89M-parameter (58.27M active per token) continuous recurrent state-space + Mixture-of-Experts model in the Ael Model Family, powered internally by the ISOM-R2 1,048,576-token architecture (ISOMR2VirtualSVDCache + ISOMR2Engine):
Ael-Edge-130M-MoE's 0.0469 MB continuous quasi-unitary SSM recurrence with ISOM-R2's Hierarchical Paged Virtual SVD Cache (ISOMR2VirtualSVDCache) and Hybrid Lexical + Dense Micro-Window Retrieval.model.generate() for contexts $> 2,048$ tokens (use_isom_r2_svd=True enabled by default in config.json).| Model Repository | Architecture Class | Base Lineage | Parameters (Total / Active) | Verified 1M Context (Colab A100) | Peak VRAM (Colab A100) | 1M Prefill + Decode Time | Exact Output Match |
|---|---|---|---|---|---|---|---|
| Prannesshkva/Ael-Coder-1.5B | AelCoder15BForCausalLM | Qwen2.5-Coder-1.5B-Instruct | 1.54B Dense | 1,052,188 tokens (514 chunks) | 3.51 GB (Weights: 3.09 GB) | 6.10 s (172,490 tok/s) | CaptureStd (100%) |
| Prannesshkva/Ael-Reasoning-1.5B-Instruct | AelReasoning15BForCausalLM | Qwen2.5-1.5B-Instruct | 1.54B Dense | 1,052,188 tokens (514 chunks) | 3.98 GB (Weights: 3.56 GB) | 6.07 s (173,342 tok/s) | CaptureStd (100%) |
| Prannesshkva/Ael-Coder-16B-MoE | AelCoder16BMoEForCausalLM | DeepSeek-Coder-V2-Lite-Instruct | 15.71B / 2.36B Active | 1,089,849 tokens (533 chunks) | 32.06 GB (Weights: 31.49 GB) | 15.86 s (68,717 tok/s) | CaptureStd (100%) |
| Prannesshkva/Ael-Enterprise-40B | AelEnterprise40BForCausalLM | Falcon-40B (4-bit NF4) | 40.0B Dense | 1,056,780 tokens (516 chunks) | 22.84 GB (Weights: 21.71 GB) | 15.39 s (68,670 tok/s) | CaptureStd) (100%) |
| Prannesshkva/Ael-Edge-130M-MoE | AelEdge130MMoEForCausalLM | Continuous SSM + 8-Expert MoE | 134.89M / 58.27M Active | Unbounded / 1M R2 | 0.0469 MB Recurrent State | 21,049 tok/s (Tesla T4) | Verified Drafter |
| Prannesshkva/Ael-504M | Ael504M | Compact Foundation Model | 504M | Compact Edge | Ultra-Light | Sub-second | Verified |
Evaluated on an NVIDIA Tesla T4 (14.56 GB VRAM, PyTorch 2.10.0+cu128, CUDA 12.8) across an unpadded continuous literature stream (Pride and Prejudice, 728,846 characters) from 10,000 up to 100,000 continuous tokens:
| Continuous Stream Length | Active Recurrent State | Allocated GPU Memory | Peak GPU VRAM | Generation Throughput | Spatial Complexity Profile | Hardware Status |
|---|---|---|---|---|---|---|
| 10,000 tokens | 0.0469 MB | 1,176.9 MB | 2,631.8 MB | 8,645.2 tok/s | Constant O(1) (< 1 MB) | SUCCESS |
| 25,000 tokens | 0.0469 MB | 1,656.4 MB | 4,795.1 MB | 14,476.6 tok/s | Constant O(1) (< 1 MB) | SUCCESS |
| 50,000 tokens | 0.0469 MB | 2,614.9 MB | 7,686.7 MB | 17,844.6 tok/s | Constant O(1) (< 1 MB) | SUCCESS |
| 75,000 tokens | 0.0469 MB | 2,614.9 MB | 8,645.2 MB | 19,885.7 tok/s | Constant O(1) (< 1 MB) | SUCCESS |
| 100,000 tokens | 0.0469 MB | 2,614.9 MB | 8,645.2 MB | 21,049.0 tok/s | Constant O(1) (< 1 MB) | SUCCESS |
| Parameter | Value |
|---|---|
| Model Class | AelEdge130MMoEForCausalLM (AelEdge130MMoEConfig) |
| Total Parameters | 134.89M Untied (109.16M Tied Backbone) / 58.27M Active per token |
| Architecture | Continuous Isometric SSM + Mixture-of-Experts + ISOM-R2 1M Engine |
| Layers / Hidden Dim / State Dim | 6 Layers / d_model = 512 / d_state = 8 |
| Experts | 8 SwiGLU Experts per layer (TopK=2) |
| Recurrent Working State | 0.0469 MB ($O(1)$ constant footprint: 6 layers × 512 channels × 8 states × 2 bytes bf16) |
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "Prannesshkva/Ael-Edge-130M-MoE"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
).eval()
prompt = "Solve step by step: If a car travels 90 km/h for 3.5 hours, what is the total distance traveled?"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=150,
temperature=0.7,
do_sample=True,
tokenizer=tokenizer,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
@software{ael_edge_130m_2026,
author = {Prannessh K.V.A.},
title = {Ael-Edge-130M-MoE: Continuous Recurrent SSM + Mixture-of-Experts Powered by ISOM-R2},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.22649142},
url = {https://huggingface.co/Prannesshkva/Ael-Edge-130M-MoE}
}