Downloads · 30 days
19
3% of all-time downloads
nazdef/1gpu-llm-medium-en-it-base
1gpu-llm-medium-en-it-base is a text generation model from nazdef. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as cc-by-sa-4.0.
Downloads · 30 days
19
3% of all-time downloads
All-time downloads
652
Public
Parameters
338M
6.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt4.1 GB · 59%
From the Hugging Face model README

This repository is the current ready-to-use base release for the 1gpu-llm
medium EN/IT family.
1gpu-llm is a family of language models trained from scratch on a single
consumer GPU.
For this release family, the reference training hardware is:
Concretely, this release packages the GPT2PreLN decay-family practical winner
at step_14700:
1gpu-llmmedium2500 tokensarchitecture: gpt2, block_type: gpt2_prelayernorm337,639,424 parameters (~337.639M) in the published
Transformers exportstep_14700.ptThis is a base model, not an instruction-tuned chat model.
stable-recipe-gpt2medium-gpt2preln-k20-wsd-lr2e-4-anchor20k-final2e5-webwikistep_13500.pt20260628_resume-gpt2medium-gpt2preln-k20-wsddecayonly-lr2e-4-anchor20k-final2e5-webwiki-step1350020260629_resume-gpt2medium-gpt2preln-k20-wsddecayonly-rerunmissing-lr3p5294e5-anchor20k-final2e5-webwiki-step14200-to14850step_14700.ptPractical reading:
step_14250 = best pure scalar / benchmark checkpointstep_14700 = best practical balanced release candidatestep_14500 = best behavior-oriented variantstep_14700 rather than the raw scalar champion step_14250This model was trained on the bilingual EN/IT web + wiki dataset:
202605141153_fineweb50_wiki50_50en_50it_score100_2500context_5Btokens_tok_20260515_en50it50_webwiki_stratified_500M2500 tokens2500source_balanced0.05Main source groups:
epfml/FineWeb-HQ)epfml/FineWeb2-HQ)google/wiki40b)google/wiki40b)Training math:
2500248239,904So this checkpoint saw approximately:
3.5265888B tokens total by step_14700287.88M extra tokens during the decay-only continuation beyond the
non-decayed anchor step_13500The final comparable GPU benchmark on the shortlisted medium family checkpoints kept the roles separate on purpose.
Pure benchmark/loss ranking:
step_14250: val_loss_mixed = 4.4419step_14700: val_loss_mixed = 4.4436step_13500: val_loss_mixed = 4.4690step_14500: val_loss_mixed = 4.4926step_15700: val_loss_mixed = 4.5038So step_14250 is still the scalar winner.
But the release decision for the single public medium base model used the practical read, not only the thinnest scalar margin:
14250 and 14700 is only about +0.001614700 is cleaner on the practical behavior proxies:
loop_rate = 0.375 vs 0.425repeated_4gram_rate = 0.750 vs 0.775language_consistency_en = 1.000 vs 0.950language_consistency_it = 0.850 vs 0.82514700 produced the strongest
holdout result among the real release candidatesSo this repo promotes the checkpoint that is the best compromise for an operational family base release, not just the one that wins the scalar leaderboard by the smallest possible edge.
step_14700val_loss_mixed = 4.4436val_loss_en = 4.3929val_loss_it = 3.5830ppl_mixed = 85.0781ppl_en = 80.8710ppl_it = 35.9822Behavior snapshot:
loop_rate = 0.375distinct_2 = 0.5643repeated_4gram_rate = 0.750language_consistency_en = 1.000language_consistency_it = 0.850Source losses:
books_en = 4.4110books_it = 4.3904code = 7.7208web_en = 5.4893web_it = 5.4087wiki_en = 2.9984wiki_it = 2.9202Short honest read:
The repo-native decoding sweep was run on this exact checkpoint.
Raw sweep result:
creativecreativePublic default:
balanced as the recommended preset for the published family-base cardbalanced as the default unless creative wins
clearly enough to justify the more aggressive presetcreative does win the holdout score, but not by a
margin large enough to force a louder default for the public base releasecreative holdout score = 2.6369balanced holdout score = 2.4656+0.1713Recommended generation params (balanced):
do_sample = truetemperature = 0.8top_k = 50top_p = 0.95repetition_penalty = 1.1no_repeat_ngram_size = 0max_new_tokens = 64Holdout metrics for the recommended preset:
score = 2.4656completion_rate = 1.0distinct_2 = 0.9878language_consistency_mean = 0.6667loop_rate = 0.0repeated_4gram_rate = 0.0language_switch_rate_mean = 0.2500length_closeness = 0.9355If you want the higher-scoring exploratory preset from the sweep instead:
creative
temperature = 1.0top_k = 1002.6369Both generation_config.json and recommended_decoding_params.json are
included in the repo.
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
repo_id = "nazdef/1gpu-llm-medium-en-it-base"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id)
prompt = "La capitale d'Italia è"
prompt_ids = tokenizer(prompt, return_tensors="pt", add_special_tokens=False)
bos = torch.tensor([[tokenizer.bos_token_id]], dtype=prompt_ids["input_ids"].dtype)
input_ids = torch.cat([bos, prompt_ids["input_ids"]], dim=1)
attention_mask = torch.ones_like(input_ids)
outputs = model.generate(
input_ids=input_ids,
attention_mask=attention_mask,
do_sample=True,
max_new_tokens=64,
temperature=0.8,
top_k=50,
top_p=0.95,
repetition_penalty=1.1,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.pad_token_id,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
.pt checkpoint.safetensors weights plus metadata sidecarmodel.safetensorsconfig.jsonbest_validation.json, metrics.jsonl, eval_metrics.jsonl, probe_generations.jsonl)summary.json, comparison.json, comparison.csv, metrics.json, metrics.csv, source_losses.json, report.md, generations.jsonl, generations_comparison.md, cloze_results.jsonl)decoding_summary.json, decoding_report.md, tuning_leaderboard.csv, holdout_leaderboard.csv, tuning_generations.jsonl, holdout_generations.jsonl)generation_config.json, recommended_decoding_params.json)release_note.mdUse this model as:
1gpu-llm familyDo not read this repo as:
14700 is the best checkpoint on every possible axis14250 on pure lossThis release is published with CC-BY-SA-4.0 as the practical downstream
posture for the mixed training corpus used here.
The training mix includes:
Downstream users are responsible for checking whether their use, redistribution, or derivative packaging remains compatible with the obligations of the upstream datasets and their terms.