Downloads · 30 days
575
100% of all-time downloads
Treese/RQwen3-751M-Base
RQwen3-751M-Base is a text generation model from Treese. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
RQwen3 is a 751,632,384-parameter language model pretrained from scratch, architecturally matching Qwen3-0.6B (RoPE, GQA, SwiGLU, RMSNorm, QK-Norm, no biases). It was trained on ~13B tokens of curated educational data…
Downloads · 30 days
575
100% of all-time downloads
All-time downloads
575
Public
Parameters
752M
3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3 GB · 100%
From the Hugging Face model README
RQwen3 is a 751,632,384-parameter language model pretrained from scratch, architecturally matching Qwen3-0.6B (RoPE, GQA, SwiGLU, RMSNorm, QK-Norm, no biases). It was trained on ~13B tokens of curated educational data on UNC Longleaf, and shares no weights with any released Qwen model — only the architecture and the tokenizer.
This is a base model: a next-token predictor, not an assistant. It has not been instruction-tuned, chat-tuned, or aligned. See Limitations before using it for anything.
The full build — architecture, data pipeline, training journey, and the bugs along the way — is documented at github.com/R-Theory/RQwen3.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Treese/RQwen3-751M-Base"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")
prompt = "The theory of general relativity"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=80, do_sample=True, temperature=0.8, top_p=0.95)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
The weights load through the standard transformers Qwen3ForCausalLM class — no
trust_remote_code needed, and anything that reads Qwen3 (vLLM, llama.cpp converters,
lm-evaluation-harness) will read this.
| Parameter | Value |
|---|---|
hidden_size | 1024 |
num_hidden_layers | 28 |
num_attention_heads | 16 |
num_key_value_heads | 8 (GQA, 2:1) |
head_dim | 128 |
intermediate_size | 3072 |
vocab_size | 151,936 |
max_position_embeddings | 2,048 |
rope_theta | 1,000,000 |
rms_norm_eps | 1e-06 |
tie_word_embeddings | false (untied LM head) |
| Total parameters | 751,632,384 |
| Serialized dtype | float32 |
Departures from Qwen3-0.6B worth knowing about: the LM head is untied from the input embedding (Qwen3-0.6B ties them), and the trained context length is 2,048 tokens rather than Qwen3-0.6B's 32K.
| Data | 6-source curated mix, ~13B tokens (~1 epoch) |
| Steps | 50,000 |
| Effective batch | 128 sequences (batch_size=2 × grad_accum=64), ~262K tokens/step |
| Sequence length | 2,048 |
| Optimizer | AdamW, weight_decay=0.1, grad_clip=1.0 |
| LR schedule | Cosine, peak 3e-4 → min 3e-5, 500-step warmup |
| Precision | bf16 autocast, SDPA attention |
| Hardware | 1× NVIDIA L40S 48GB (UNC Longleaf), ~11 wall-clock days over 10 SLURM submissions |
| Final training loss | 2.5186 (perplexity ≈ 12.4) |
| Starting loss | 11.88 (random init) |
| Source | HF path | Share |
|---|---|---|
| FineWeb-Edu | HuggingFaceFW/fineweb-edu | 54% |
| Wikipedia | wikimedia/wikipedia (20231101.en) | 15% |
| OpenWebMath | open-web-math/open-web-math | 12% |
| StackExchange | HuggingFaceH4/stack-exchange-preferences | 8% |
| peS2o | MaLA-LM/peS2o-final | 8% |
| Textbooks | HuggingFaceTB/cosmopedia (stanford) | 4% |
Quality-filtered, exact-hash deduplicated within each source, and pre-tokenized into binary shards. Details: docs/data-pipeline.md.
Run with EleutherAI/lm-evaluation-harness
v0.4.13, hf backend, bf16 on a single L40S. All three tasks at the harness's registered defaults
(0-shot for all three in 0.4.13).
| Task | Config | Score | Chance |
|---|---|---|---|
| ARC-Challenge | 0-shot, acc | 24.32% ± 1.25 | 25% |
| HellaSwag | 0-shot, acc | 31.47% ± 0.46 | 25% |
| MMLU (all 57 subjects, average) | 0-shot, acc | 25.32% ± 0.37 | 25% |
Reading these honestly:
For rough context on the training loss value above: from-scratch dense models around this size typically land near ~2.85 (GPT-2 large 774M, Pythia-410M) on their own training distributions, and Qwen3-0.6B reports ~2.4 after ~5T tokens — roughly 400× more data than this run saw. These numbers come from different corpora and are not apples-to-apples.
temperature
and top_p set.Not suitable for production, for user-facing deployment, or for any decision-making use. This is a research and educational artifact.
Uses the Qwen3 BPE tokenizer (151,936 tokens) from
Qwen/Qwen3-0.6B, unchanged. EOS token id 151645.
Weights were converted from the native training checkpoint with
scripts/export_to_hf.py,
which verifies the exported model against the original src/ implementation — same tokens, same
logits in fp32 — before writing anything.
Apache 2.0, matching the Qwen3 architecture and tokenizer this model is built on.
@misc{rqwen3_2026,
title = {RQwen3: A 751M-Parameter Qwen3 Architecture Pretrained From Scratch},
author = {Treese},
year = {2026},
url = {https://github.com/R-Theory/RQwen3}
}