Downloads · 30 days
0
darkstudio009/Vanessa-1B
Vanessa-1B is a text generation model from darkstudio009. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
A ~1.1B-parameter decoder-only transformer, pretrained fully from scratch — original tokenizer, original model code, original training loop.
Downloads · 30 days
0
Access
Public
Updated Sep 23, 2026
Repo size
912 GB
Likes
0
Public
Click a slice to open those files.
.bin81 GB · 75%
From the Hugging Face model README
A ~1.1B-parameter decoder-only transformer, pretrained fully from scratch — original tokenizer, original model code, original training loop.
Developed by Dark Studio.
⚠️ Status: pretraining in progress. Vanessa-1B has not finished pretraining and is not yet an instruction-tuned or fine-tuned assistant. Weights currently on this repo are training checkpoints, not a finished release.
Vanessa-1B is an independent large language model project: every part of the stack — the byte-level BPE tokenizer, the transformer implementation, and the distributed training loop — was built from first principles rather than adapted from an existing checkpoint or wrapping an existing transformers architecture class. It is the second generation of the Vanessa project, scaled up roughly 9x from an initial ~124M-parameter proof of concept.
| Parameters | ~1.1B |
| Architecture | Decoder-only Transformer |
| Layers | 22 |
| Hidden size | 2,048 |
| Attention heads | 32 |
| Key/value heads | 4 (Grouped-Query Attention, 8:1) |
| Feed-forward | SwiGLU, intermediate size 5,632 |
| Normalization | RMSNorm |
| Position embeddings | Rotary (RoPE), θ = 10,000 |
| Context length | 2,048 tokens |
| Vocabulary | 32,768 |
| Tied embeddings | No |
Data. FineWeb-Edu (the education-filtered variant of FineWeb), targeting 25B tokens.
Optimizer. AdamW (β₁=0.9, β₂=0.95), peak learning rate 3e-4 with linear warmup and cosine decay to 10% of peak, weight decay 0.1 (applied to 2D+ parameters only), gradient clipping at 1.0.
Precision. Mixed precision: fp16 for forward/backward compute with dynamic loss scaling, fp32 master weights for the optimizer update.
Infrastructure. Trained on 2x NVIDIA T4 GPUs (Kaggle) under PyTorch FSDP (FULL_SHARD) with gradient checkpointing, since the full model does not fit as a single replica on one 16GB GPU. Training is resumable across ephemeral sessions via checkpoints stored on this repo.
Custom byte-level BPE, vocabulary size 32,768, trained from scratch on a fixed sample for reproducibility. ChatML-style special tokens (<|im_start|>, <|im_end|>, role tokens, tool-call tokens, reasoning tokens) are pre-reserved in the vocabulary for a future instruction-tuning stage.
Vanessa-1B is a base (completion) model from an independent research/hobby project, currently mid-pretraining. It has not been instruction-tuned, safety-aligned, or evaluated on standard benchmarks. It should be treated as experimental: not suited for production use, high-stakes decisions, or deployment without further training and evaluation. Its knowledge and behavior are entirely a product of its training data (FineWeb-Edu) and are expected to reflect that data's coverage and biases.
Vanessa-1B uses its own architecture classes (VanessaConfig, VanessaForCausalLM), not a built-in transformers model type, so it isn't loadable via a plain AutoModel.from_pretrained(...) call yet. Once pretraining is complete and the sharded FSDP checkpoint has been exported to a single merged state dict, loading looks like:
import json
from vanessa_architecture import VanessaConfig, VanessaForCausalLM
config = VanessaConfig.from_dict(json.load(open("config.json")))
model = VanessaForCausalLM(config)
model.load_state_dict(torch.load("vanessa_merged.pt")["params"])
(vanessa_architecture.py is the model definition from the training notebook; a merge step from the two FSDP rank shards to a single state dict is needed first.)
Vanessa-1B is developed and maintained by Dark Studio, an independent effort to build large language models from first principles.
Not yet finalized — update the license field above and this section before wider release.