Downloads · 30 days
30
25% of all-time downloads
GRKTheGreat/ForgeLM-v1
ForgeLM-v1 is a text generation model from GRKTheGreat. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A training-free architectural port of Qwen2.5-Coder-1.5B-Instruct with KeyStack transforms.
Downloads · 30 days
30
25% of all-time downloads
All-time downloads
119
Public
Parameters
1.8B
3.6 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors3.6 GB · 100%
From the Hugging Face model README
A training-free architectural port of Qwen2.5-Coder-1.5B-Instruct with KeyStack transforms.
ForgeLM v1 is not a trained model. It is a weight-transformed port of Qwen2.5-Coder-1.5B-Instruct, created by applying a series of closed-form mathematical transforms (called "Keys") that restructure the architecture without losing the original model's knowledge.
This model is a port of Qwen2.5-Coder-1.5B-Instruct, not an independently trained model. All of the model's knowledge, capabilities, and limitations come from the original Qwen model. The KeyStack transforms are mathematically lossless or near-lossless at initialization — they restructure the architecture but do not add new knowledge. The original Qwen2.5-Coder-1.5B-Instruct model is licensed under Apache 2.0 by Alibaba/Qwen Team.
This model and its entire codebase were vibe-coded via Devin Desktop — an AI coding agent by Cognition. No human wrote the transform code, inference engine, or model architecture. The project was directed through natural language prompts and the AI agent implemented everything: weight porting, KeyStack transforms, inference engine, fine-tuning scripts, and this documentation.
| Component | Original (Qwen) | ForgeLM v1 | Transform |
|---|---|---|---|
| Attention | GQA (12 Q heads, 2 KV heads) | MLA (d_c=512) | SVD-based GQA→MLA |
| FFN | Dense SwiGLU (8960 intermediate) | MoE (4 routed + 1 shared, d_ff=1792) | Weight splitting |
| Norm | RMSNorm | RMSNorm (unchanged) | Direct copy |
| Embedding | 151936 vocab, 1536 dim | Same (unchanged) | Direct copy |
| LM Head | Tied with embedding | Same (unchanged) | Direct copy |
| RoPE | θ=1,000,000 | Same (unchanged) | Direct copy |
All transforms are lossless at initialization (cosine similarity ≥ 0.9998 vs original Qwen):
| Key | Type | Effect | Cosine Sim |
|---|---|---|---|
| MLA | FULL | GQA→MLA via SVD, 4x KV cache compression | 1.0000 (100% energy) |
| MoE | FULL | Dense FFN→5 experts, 60% active FLOPs | 1.0000 (exact split) |
| MRL | FULL | Matryoshka dimension reordering | 1.0000 (permutation) |
| QuaRot | FULL | Hadamard rotation on V/O | 1.0000 (rotation) |
| ValueResidual | FULL | V0 stored + 28 gates (init=0) | 1.0000 (no-op at init) |
| RotorQuant | FULL | Givens rotation matrices stored | 1.0000 (no-op at init) |
| MTP | FULL | 4 prediction heads from LM head | 1.0000 (identity init) |
| AirLLM | TRIVIAL | Streamable flag for low-VRAM | N/A (runtime) |
| Technique | Reason |
|---|---|
| GQA→MQA | Catastrophic degradation without 5% pretraining compute |
| Wanda pruning | Needs calibration data post-MRL |
| SliceGPT | Needs calibration data |
| BitNet | Needs native ternary training |
| ShortGPT | Available as key, not applied to v1 |
| Configuration | tok/s | KV Compression | Notes |
|---|---|---|---|
| Standard | 21 | 1x | Baseline |
| Hadamard INT4 KV | 21 | 4x | Lossless |
| Streaming KV | 12 | ∞ (sinks+window) | Infinite context |
| RotorQuant KV | 13 | 3.88x | 0.94% error |
| MTP self-spec | 11 | 1x | Correct output, needs batch verify |
The KeyStack is a system of Keys — closed-form mathematical transforms that convert weights between architectures without training. Each Key implements:
forward(data) → weights — replace training with instant transformreverse(weights) → data — extract what was learnedclassify() → KeyClass — FULL (reversible+composable), PARTIAL (lossy), TRIVIAL (no weights)| Class | Criteria | Examples |
|---|---|---|
| FULL | Reversible + data→weight + composable | MLA, MoE, MRL, QuaRot, ValueResidual, RotorQuant, MTP |
| BI | Both directions, round-trip identity | Embedding, RMSNorm, LM Head, RoPE, CausalMask |
| PARTIAL | One direction only (lossy) | GPTQ, ShortGPT, Wanda, SliceGPT |
| TRIVIAL | No weights (runtime/formula) | AirLLM, QK-Norm, LogitCap, StreamingLLM |
ForgeLM v1 includes a unified inference engine (ForgeEngine) with pluggable strategies:
standard — Basic tensor cachestreaming — StreamingLLM (attention sinks + sliding window, infinite context)snapkv — Observation-window eviction (8x compression)hadamard_int4 — Block-diagonal Hadamard + INT4 (4x compression, lossless)rotorquant — Givens rotation + Lloyd-Max quantization (3.88x compression)compressed — H2O heavy-hitter eviction + KV quantizationpaged — vLLM-style paged memorystandard — Autoregressive token-by-tokenmtp_selfspec — MTP self-speculative decoding (draft + verify)speculative — Draft model speculative decodingmedusa — Medusa multi-head decodingdspark — DSpark semi-autoregressive decodingairllm_streaming — Smart layer-streaming (only when VRAM insufficient)cuda_graph — CUDA graph capture for decode steptorch.compile — Inductor compilation (1.3-2x on GQA)prefix_cache — KV cache reuse for repeated prompt prefixesfrom research.inference.forge_engine import ForgeEngine
# Load ForgeLM v1
engine = ForgeEngine.from_checkpoint(
checkpoint="forgelm_v1.safetensors",
config_name="forgelm_v1",
device="cuda",
)
# Activate with runtime strategies
engine.activate(
kv_cache="hadamard_int4", # 4x KV compression, lossless
decoding="standard",
use_prefix_cache=True,
)
# Generate
output = engine.generate("def fibonacci(n):", max_new_tokens=100)
print(output)
engine.activate(
kv_cache="streaming", # 4 sinks + 512 window
decoding="standard",
)
output = engine.generate(long_prompt, max_new_tokens=1000)
from research.mtp import MTPHead
from safetensors.torch import load_file
# Load fine-tuned MTP head
mtp_state = load_file("forgelm_v1_mtp.safetensors")
mtp_head = MTPHead(d_model=1536, vocab_size=151936, n_predict=4).cuda()
mtp_head.load_state_dict(mtp_state)
engine.model.mtp_head = mtp_head
engine.activate(
kv_cache="standard",
decoding="mtp_selfspec", # Speculative decoding
)
engine = ForgeEngine.from_checkpoint(
checkpoint="forgelm_v1.safetensors",
config_name="forgelm_v1",
device="cuda",
)
engine.activate(
kv_cache="standard",
decoding="standard",
acceleration="airllm_streaming", # Auto-detects: streams only if VRAM < model size
)
| File | Description |
|---|---|
forgelm_v1.safetensors | Model weights (MLA + MoE, 3.4 GB) |
config.json | HF-compatible model config |
tokenizer.json | Qwen2.5 tokenizer (unchanged) |
tokenizer_config.json | Tokenizer config |
vocab.json | Vocabulary |
merges.txt | BPE merges |
README.md | This file |
GQA K/V projections are factored through a shared low-rank bottleneck via SVD:
W_KV = [W_K; W_V] → SVD → W_c (compress) + W_KC, W_VC (decompress)
With d_c=512 (rank of GQA KV), 100% energy is retained.
Dense SwiGLU FFN is split into 5 experts (4 routed + 1 shared):
Dense: Y = W_down(silu(W_gate(X)) * W_up(X)) [intermediate=8960]
MoE: Y = sum_i gate_i * Expert_i(X) + Shared(X)
Each expert: d_ff = 8960 / 5 = 1792
With top-4 routing (all experts active), output is identical to dense. Top-2 routing (60% FLOPs) requires router fine-tuning.
MTP heads' shared trunk was fine-tuned for 200 steps (122 seconds):
If you use ForgeLM v1, please cite both the original Qwen model and this work:
@misc{forgelm_v1,
title={ForgeLM v1: Training-Free Architectural Port of Qwen2.5-Coder via KeyStack Transforms},
author={ForgeAI Project},
year={2026},
note={Vibe-coded via Devin Desktop (Cognition)}
}
@misc{qwen25_coder,
title={Qwen2.5-Coder Technical Report},
author={Qwen Team},
year={2024},
url={https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct}
}
Apache 2.0 (inherited from Qwen2.5-Coder-1.5B-Instruct)