Downloads · 30 days
37
5% of all-time downloads
i3-lab/i3-80m
i3-80m is a text generation model from i3-lab. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
The i3-80M Model is a novel hybrid architecture combining convolutional/recurrent layers with full attention layers for efficient language modeling. This architecture uniquely blends RWKV-style time-mixing with Mamba…
Downloads · 30 days
37
5% of all-time downloads
All-time downloads
681
Public
Parameters
82.8M
662 MB on disk
Likes
7
Public
Click a slice to open those files.
.bin331 MB · 50%
From the Hugging Face model README
The i3-80M Model is a novel hybrid architecture combining convolutional/recurrent layers with full attention layers for efficient language modeling. This architecture uniquely blends RWKV-style time-mixing with Mamba state-space dynamics in the early layers, followed by standard multi-head attention in deeper layers.
This is the second model in the i3 series, scaling up from the original i3-22M with improved architecture and multi-dataset training.
[!NOTE] To use the model try it here.
Layers 1-10: RWKV-Mamba Hybrid Blocks (Recurrent/Conv)
├─ RWKVMambaHybrid (Time-mixing + State-space)
└─ Feed-Forward Network (4x expansion)
Layers 11-16: Full Attention Blocks
├─ Multi-Head Attention (16 heads)
└─ Feed-Forward Network (4x expansion)
| Feature | i3-22M | i3-80M (This Model) |
|---|---|---|
| Parameters | 22.6M | 82.77M |
| Architecture | 24 Hybrid Layers | 10 Hybrid + 6 Attention Layers |
| Hidden Dimension | 512 | 512 |
| Vocabulary Size | 4,466 | 35,560 |
| Training Dataset | TinyChat only | TinyStories + TinyChat + HQ Sentences |
| Total Tokens | ~1M conversations | ~3M+ tokens |
| Final Loss | ~2.0 | ~2.0 |
| Final Perplexity | 7.29-9.70 | 7.29-10.0 |
| Training Time | ~17 hours | ~2-4 hours |
| Attention Layers | None (Pure Hybrid) | 6 Full Attention Layers |
Use i3-22M if you need:
Use i3-80M if you need:
Hybrid Architecture: Combines the efficiency of recurrent/convolutional processing with the power of attention
Memory-Optimized Training:
Multi-Dataset Pre-training: Trained on diverse text sources for robust language understanding
Smart Tokenization: Variable-length chunking (2-3 chars) with common trigram optimization
agentlans/high-quality-english-sentencesroneneldan/TinyStoriesstarhopp3r/TinyChat| Metric | Initial | Final |
|---|---|---|
| Training Loss | ~10.0 | ~1.7 |
| Perplexity | ~4000+ | ~6 |

[!NOTE] I dont know why the logging starts at step 4.6k .
i3-22m and i3-80m comparation?

The model shows strong convergence with stable training dynamics and efficient GPU utilization.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load model and tokenizer
model = AutoModelForCausalLM.from_pretrained("FlameF0X/i3-80m")
tokenizer = AutoTokenizer.from_pretrained("FlameF0X/i3-80m")
# Generate text
prompt = "hello"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(
inputs.input_ids,
max_length=100,
temperature=0.8,
top_k=40
)
generated_text = tokenizer.decode(outputs[0])
print(generated_text)
RWKV-Mamba Hybrid Recurrence: Combines RWKV's time-mixing with Mamba's state-space dynamics
Hierarchical Processing:
Memory Efficiency:
pytorch_model.bin: Model weightsconfig.json: Model configurationchunk_vocab_combined.json: Tokenizer vocabularyThis model was tracked using Weights & Biases (WandB) with comprehensive metrics:
@misc{i3-80m,
author = {FlameF0X},
title = {i3-80M: Hybrid Architecture Language Model},
year = {2025},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/FlameF0X/i3-80m}}
}
@article{mamba,
title={Mamba: Linear-Time Sequence Modeling with Selective State Spaces},
author={Gu, Albert and Dao, Tri},
journal={arXiv preprint arXiv:2312.00752},
year={2023}
}
@article{RWKV,
title={RWKV: Reinventing RNNs for the Transformer Era},
author={Peng, Bo and others},
journal={arXiv preprint arXiv:2305.13048},
year={2023}
}