Downloads · 30 days
19
15% of all-time downloads
DienerTech/sparknet-70m
sparknet-70m is a text generation model from DienerTech. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
SparkNet 70M v5 is the final 70M-parameter checkpoint from the SparkNet research run by DienerTech. It is a compact GPT-2–style decoder (12 layers, 512 hidden size, 8 attention heads, 1024-token context) that was trai…
Downloads · 30 days
19
15% of all-time downloads
All-time downloads
128
Public
Parameters
64.1M
256 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors256 MB · 99%
From the Hugging Face model README
SparkNet 70M v5 is the final 70M-parameter checkpoint from the SparkNet research run by DienerTech. It is a compact GPT-2–style decoder (12 layers, 512 hidden size, 8 attention heads, 1024-token context) that was trained for ~1B tokens on a custom mixture of high-quality web and document corpora. The release ships with the SparkNet v5 tokenizer and weights stored in model.safetensors, ready for direct use via 🤗 Transformers.
Special thanks to CodeLion for inspiring the One Billion Token Challenge, and for providing the high-quality datasets used in this training run.
<|pad|> padding).model.safetensors for safe loading; no pytorch_model.bin left in the repo.datasets/sparknet-v5-1b).codelion/finepdfs-1B, codelion/dclm-baseline-1B, codelion/fineweb-edu-1B, plus curated DienerTech blog data.wikitext-2-raw-v1 (standard Hugging Face split).trainer_state.json).Formal downstream evaluation has not been run yet. Inside trainer_state.json, the best validation (WikiText-2) cross-entropy reached 4.9869 at step 14k. If you benchmark the model (e.g., with lm-eval-harness), please consider contributing results back to the card via a PR.
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "DienerTech/sparknet-70m-v5"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16, # or torch.float16 on older GPUs
device_map="auto",
)
prompt = "In a distant research lab, a tiny transformer model awakened and"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(
**inputs,
max_new_tokens=120,
temperature=0.9,
top_p=0.9,
do_sample=True,
)
print(tokenizer.decode(output[0], skip_special_tokens=True))
@software{sparknet70mv5,
author = {DienerTech},
title = {SparkNet 70M v5},
year = {2025},
url = {https://huggingface.co/DienerTech/sparknet-70m-v5}
}
Please open an issue or PR on the DienerTech Hugging Face repo if you have feedback, evaluations, or fine-tuned variants to share.