Downloads · 30 days
38
24% of all-time downloads
qvx-o/QED-Base-v1
QED-Base-v1 is a text generation model from qvx-o. Use it when you need the model to write or continue text. The card lists the license as mit.
QED-Base-v1-300M is a decoder-only Transformer language model trained entirely from scratch.
Downloads · 30 days
38
24% of all-time downloads
All-time downloads
161
Public
Repo size
1 GB
Likes
0
Public
Click a slice to open those files.
.pt1 GB · 100%
From the Hugging Face model README
QED-Base-v1-300M is a decoder-only Transformer language model trained entirely from scratch.
Unlike fine-tuned or continually pretrained models, QED-Base-v1 was not initialized from or adapted from any existing pretrained model. All model weights were learned through original autoregressive pretraining using an independently created dataset.
The model serves as the foundation for future instruction-tuned and task-specific QED models while remaining suitable for language modeling research and experimentation.
| Property | Value |
|---|---|
| Model Name | QED-Base-v1-300M |
| Model Type | Decoder-only Transformer |
| Parameters | 299.82M |
| Hidden Dimension | 1024 |
| Layers | 12 |
| Attention Heads | 16 |
| Activation Function | SwiGLU |
| Normalization | RMSNorm |
| Position Encoding | RoPE |
| Vocabulary Size | 96,000 |
| Tokenizer | SentencePiece |
| Training Method | Pretraining from Scratch |
| Base Model | None |
QED-Base-v1 retains a modern decoder-only Transformer architecture featuring:
The architecture is designed as a general-purpose language foundation model suitable for pretraining, research, and downstream fine-tuning.
QED-Base-v1 was pretrained using a fully original dataset created specifically for this project.
The training corpus was independently authored and prepared for large-scale autoregressive language modeling.
The data preparation process included:
No existing pretrained model weights were used during training.
QED-Base-v1 is intended for:
This is a foundation model and has not been instruction-tuned.
As a pretrained foundation model:
MIT License.