Downloads · 30 days
15
29% of all-time downloads
Ghostgim/ghost-small-gen
ghost-small-gen is a text generation model from Ghostgim. Use it when you need the model to write or continue text. It is set up for ghostlm. The card lists the license as mit.
A small (~45M parameter) decoder-only language model trained entirely from scratch in PyTorch, no pretrained weights and no fine-tuning. It is the first generalist checkpoint in the GhostLM project: a model that broad…
Downloads · 30 days
15
29% of all-time downloads
All-time downloads
52
Public
Parameters
45M
90.2 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors90 MB · 100%
From the Hugging Face model README
A small (~45M parameter) decoder-only language model trained entirely from scratch in PyTorch, no pretrained weights and no fine-tuning. It is the first generalist checkpoint in the GhostLM project: a model that broadened from a cybersecurity-only corpus into a small generalist while keeping cybersecurity as its deepest specialty.
Debiased multi-permutation text-scoring. + means the 95% bootstrap CI lower
bound is above the 25% random baseline (significantly above chance). Peer
numbers are published zero-shot references for the small-model class; harnesses
differ, so treat them as context, not an exact comparison.
| Benchmark | n | ghost-small-gen (45M) | 95% CI | vs random | Peer reference |
|---|---|---|---|---|---|
| ARC-Easy | 2365 | 27.2% | 25.4-28.9 | + | Pythia-160M 43.5, 111M 34.8, 256M 37.6 |
| ARC-Challenge | 1165 | 24.3% | 22.1-26.6 | ~ | Pythia-160M 18.8, SmolLM2-360M 36.6 |
| OpenBookQA | 500 | 27.4% | 23.7-31.1 | ~ | 111M 27.8, 256M 25.4, LaMini-35M 26.2 |
| SecQA (cyber) | 210 | 34.3% | 28.5-40.6 | + | retained specialty |
| CTF eval (cyber) | 30 | 63.3% | 46.7-80.0 | + | retained specialty |
What this says, calibrated for a 45M model trained from scratch on free compute:
This is a solid, defensible result for the size and the compute, not a "beats everything" claim.
Research and education: a transparent, hand-written small model for studying from-scratch training, generalist corpus design, and cybersecurity-aware language modeling. It is a base model (not instruction-tuned), small, and will hallucinate. Do not use it for safety-critical decisions. The cybersecurity content is for defensive and educational understanding.
This is a custom architecture, not a transformers model. Load it with the
GhostLM code:
import torch
from safetensors.torch import load_model
from ghostlm.config import GhostLMConfig
from ghostlm.model import GhostLM
from huggingface_hub import hf_hub_download
import json
repo = "Ghostgim/ghost-small-gen"
cfg = GhostLMConfig(**{k: v for k, v in json.load(open(hf_hub_download(repo, "config.json"))).items()
if k in GhostLMConfig().__dataclass_fields__})
model = GhostLM(cfg)
load_model(model, hf_hub_download(repo, "model.safetensors"))
model.eval()
See the GhostLM repository for the model code, tokenizer, generation utilities, and the full scorecard.
MIT. Built and trained by Joe Munene.