Downloads · 30 days
34
3% of all-time downloads
Crusadersk/gpt2-100m
gpt2-100m is a text generation model from Crusadersk. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Custom-trained GPT-2 checkpoint with deliberate depth-width configuration for inference benchmarking research.
Downloads · 30 days
34
3% of all-time downloads
All-time downloads
981
Public
Parameters
96.1M
384 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors384 MB · 100%
From the Hugging Face model README
Custom-trained GPT-2 checkpoint with deliberate depth-width configuration for inference benchmarking research.
Created as part of the Banterhearts research program investigating benchmarking integrity for local LLM inference.
| Architecture | GPT2LMHeadModel (MHA) |
| Parameters | 100M |
| Config | n_embd=768, n_head=2, n_layer=8, n_inner=3072 |
| Context length | 1,024 tokens |
| Precision | FP32 |
| Model size | 367 MB |
| Vocab size | 50,257 |
Largest MHA model; used in cross-backend and compiler benchmarks.
These checkpoints are not general-purpose language models. They are deliberately sized scaling-study artifacts designed to isolate the effect of model depth vs width on GPU inference latency. The key finding: in the small-model GPU regime, layer depth (not parameter count) dominates latency, producing inversions where a 5M-parameter model can be 3.6x slower than a 25M-parameter model.
Used in: TR117, TR120, TR126, TR147
| TR | Role |
|---|---|
| TR117 | Original cross-backend benchmark matrix (7 backends, 4 model groups) |
| TR126 | Linux/Triton compiler validation with phase-separated measurement |
| TR147 | Second-regime portability validation on RTX 6000 Ada |
The GPT-2 family (25M, 50M, 100M) uses a 2x3 factorial design:
| Model | n_embd | n_layer | n_inner | Params | Design role |
|---|---|---|---|---|---|
| gpt2-25m | 384 | 3 | 1,536 | 25M | Shallow, narrow |
| gpt2-50m | 512 | 8 | 2,048 | 50M | Deep, medium width |
| gpt2-100m | 768 | 8 | 3,072 | 100M | Deep, wide |
All models use 2 attention heads (MHA, not GQA) to isolate architecture effects from attention-group structure. Dropout is set to 0.0 for deterministic inference measurement.
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Crusadersk/gpt2-100m")
tokenizer = AutoTokenizer.from_pretrained("Crusadersk/gpt2-100m")
inputs = tokenizer("Hello", return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=32, do_sample=False)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
| Framework | Supported |
|---|---|
| Transformers | Yes |
| torch.compile (Inductor) | Yes |
| Ollama | No (not GGUF format) |
| vLLM | Yes |
@misc{banterhearts2026gpt2100m,
title = {Custom GPT-2 Scaling Checkpoint (100M) for Inference Benchmarking Research},
author = {Kadadekar, Sahil},
year = {2026},
url = {https://huggingface.co/Crusadersk/gpt2-100m},
note = {Part of the Banterhearts research program. NeurIPS 2026 submission.}
}
This work is part of the Chimera/Banterhearts technical-report program on deployment-time LLM behavior, quantization, refusal robustness, batching effects, and inference-stack reliability. Canonical public archive: Chimeraforge Reports; source context: github.com/Sahil170595/Banterhearts.