Downloads · 30 days
0
AwareLiquid/O1-Qwen05-Adapter
O1-Qwen05-Adapter is a text generation model from AwareLiquid. Use it when you need the model to write or continue text.
An MT-LNN residual adapter attached to a frozen Qwen2.5-0.5B-Instruct.
Downloads · 30 days
0
Access
Public
Updated Sep 19, 2026
Repo size
74.7 MB
Likes
0
Public
Click a slice to open those files.
.pt74.7 MB · 100%
From the Hugging Face model README
An MT-LNN residual adapter attached to a frozen Qwen2.5-0.5B-Instruct.
The adapter adds multi-timescale liquid recurrence (selective decay) to six
decoder layers of the frozen base model. It is initialised with
init_scale=0.001 so that at step 0 it acts as a near-identity residual — the
base model capability is fully preserved from the start, with MT-LNN dynamics
gradually turning on as training progresses.
| File | Description |
|---|---|
llama_mt_adapter_000500.pt | Final checkpoint (500 steps) — the deployed one |
llama_mt_adapter_000400.pt | Intermediate checkpoint (400 steps) |
llama_mt_adapter_000200.pt | Intermediate checkpoint (200 steps) |
Each checkpoint embeds its training args, from which the six adapter mounting
points are reconstructed at load time.
| Item | Value |
|---|---|
| Base model | Qwen/Qwen2.5-0.5B-Instruct (~494M params, frozen) |
| Adapter type | MTResidualAdapter (pre-norm residual) |
| Inserted at layers | 3, 7, 11, 15, 19, 23 (every 4th of 24 → 6 adapters) |
| Trainable parameters | 12,421,926 (~12.4M, 2.5% of base) |
| Protofilaments | 13 |
| Time scales | 5 |
| Map hidden dim | 64 |
| Init scale | 0.001 |
| Parallel scan | enabled (real causal recurrence) |
| Dropout | 0.0 |
| Item | Value |
|---|---|
| Dataset | Salesforce/wikitext / wikitext-2-raw-v1 (train) |
| Sequence length | 512 tokens |
| Batch | 4 × grad_accum 4 (effective 16) |
| Total steps | 500 (≈ 4.1M tokens) |
| Learning rate | 2e-4 (AdamW), weight decay 0.01, grad clip 1.0 |
| Precision | bfloat16 |
| Hardware | RTX 5060 Laptop (8 GB) — ~10 minutes total |
Final-window loss: 2.5467 / 2.8433 / 2.5862 / 2.6644 (steps 470–500). The base Qwen2.5-0.5B-Instruct reaches ~2.3 loss on WikiText-2 after full pretraining; the adapter is within the expected range for this brief run — a lightly adapted model, not a converged one.
Loading requires the MT-LNN adapter code from the M1 repository (the adapter is
not a PEFT module — vanilla peft cannot read it):
recipes.load_mt_adapter_dir path in progress)MIT (adapter weights). Base model: Qwen2.5-0.5B-Instruct (Apache-2.0).