Downloads · 30 days
31
8% of all-time downloads
Gugu8/Quark-1
Quark-1 is a text generation model from Gugu8. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as apache-2.0.
Quark-1 is a ~4.2M-parameter causal transformer language model, trained entirely from scratch (no pretrained base) on a single T4 GPU. It's a small-scale hobby/research model, not intended to compete with production L…
Downloads · 30 days
31
8% of all-time downloads
All-time downloads
376
Public
Parameters
4.2M
40.1 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors40.1 MB · 99%
From the Hugging Face model README
Quark-1 is a ~4.2M-parameter causal transformer language model, trained entirely from scratch (no pretrained base) on a single T4 GPU. It's a small-scale hobby/research model, not intended to compete with production LLMs — see Limitations below for what that means in practice.
Quark-1 went through a full multi-stage pipeline rather than a single training run:
Two decoding strategies are included in the training/inference code (see the original repo — this Hub upload contains the base weights + architecture only, ready to be paired with either):
import torch
from huggingface_hub import hf_hub_download
from modeling_quark1 import Quark1Model
from tokenizers import Tokenizer
model = Quark1Model.from_pretrained("your-username/Quark-1")
model.eval()
tokenizer_path = hf_hub_download("your-username/Quark-1", "tokenizer.json")
tokenizer = Tokenizer.from_file(tokenizer_path)
BOS_ID = tokenizer.token_to_id("<bos>")
EOS_ID = tokenizer.token_to_id("<eos>")
USER_ID = tokenizer.token_to_id("<user>")
ASSIST_ID = tokenizer.token_to_id("<assistant>")
prompt = "Tell me a simple story."
ids = [BOS_ID, USER_ID] + tokenizer.encode(prompt).ids + [ASSIST_ID]
input_ids = torch.tensor([ids])
out = model.generate(input_ids, max_new_tokens=80, temperature=0.8, top_k=40, eos_id=EOS_ID)
response = tokenizer.decode(out[0].tolist()[len(ids):], skip_special_tokens=True)
print(response)
Quark-1 is a 4.2M-parameter model — for scale, that's roughly 1,000x smaller than models like GPT-3. It is trained almost entirely on TinyStories, a synthetic dataset of simple children's stories, so:
This model is best understood as a from-scratch training and inference-scaffolding exercise, not a general-purpose assistant.
Trained on a single Google Colab T4 GPU. Full pretraining run: ~4,000 steps, ~17 minutes.