Downloads · 30 days
227
7% of all-time downloads
yousefg/MaximusLLM
MaximusLLM is a text generation model from yousefg. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
MaximusLLM is a long-context language model designed for hyper-efficient architecture and training. It introduces a new paradigm for to long context while reducing training VRAM by ~40% and increasing throughput by ov…
Downloads · 30 days
227
7% of all-time downloads
All-time downloads
3.4K
Public
Parameters
259M
82.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.1 GB · 99%
From the Hugging Face model README
MaximusLLM is a long-context language model designed for hyper-efficient architecture and training. It introduces a new paradigm for to long context while reducing training VRAM by ~40% and increasing throughput by over 17x compared to optimized standard Cross-Entropy baselines.
MaximusLLM (190M) is an architectural proof-of-concept. While it demonstrates extreme efficiency, its absolute knowledge capacity is limited by its parameter count. Users should expect hallucinations.
from src.model import Model, Config
from src.lora import blockswap_attention_layers
from src.infer import general_generate_fn
config = Config.from_pretrained("yousefg/MaximusLLM")
model = Model(config, device="cuda")
blockswap_attention_layers(model)
prompt = "<start_of_turn>user\nWhat is the capital of France?<end_of_turn>\n<start_of_turn>model\n"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
output = general_generate_fn(model, inputs, tokenizer, max_new_tokens=50)
print(tokenizer.decode(output[0]))
HuggingFaceFW/fineweb-edu.roneneldan/TinyStories to stabilize linguistic fluidity.HuggingFaceH4/ultrachat_200k using a multi-turn conversational format.Maximus utilizes a specialized training pipeline to maintain FP32 master weight stability while achieving FP16 throughput.
MaximusLLM utilizes three core innovations:
torch.compile (Hollow-compilation of inner blocks for stability).MAXIS Loss:
@article{gamaleldin2026maxis,
title={MAXIS: A Hyper-Efficient Paradigm for Scalable Long-Context LLM Training},
author={Gamaleldin, Yousef},
journal={SSRN: Artificial Intelligence eJournal},
year={2026}
}
RandNLA Attention:
@article{gamaleldin2026randnla,
title={Bifurcated Latent Attention: Scaling LLMs to Infinite Context via Asymmetric Causal RandNLA},
author={Gamaleldin, Yousef},
journal={SSRN: Artificial Intelligence eJournal},
year={2026}
}
Yousef Gamaleldin - [yrafat38@gmail.com]