Downloads · 30 days
769
23% of all-time downloads
Zyphra/ZAYA1-base
ZAYA1-base is a text generation model from Zyphra. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
ZAYA1 is an 800m active/8.3B total parameter MoE model, and the first trained entirely end-to-end on AMD’s hardware, software, and networking stack.
Downloads · 30 days
769
23% of all-time downloads
All-time downloads
3.3K
Public
Parameters
8.8B
17.7 GB on disk
Likes
45
Public
Click a slice to open those files.
.safetensors17.7 GB · 100%
How the weights are stored.
BF168.8B · 100%
From the Hugging Face model README
ZAYA1 is an 800m active/8.3B total parameter MoE model, and the first trained entirely end-to-end on AMD’s hardware, software, and networking stack.
Our ZAYA1 base model benchmark performance is extremely competitive with the SoTA Qwen3 series of models of comparable scale, and outperforms comparable western open-source models such as SmolLM3, and Phi4. ZAYA1-base excels especially at complex and challenging mathematical and STEM reasoning tasks, nearly matching the performance of SoTA Qwen3 thinking models under high pass@k settings even prior to explicit post-training for reasoning, and exceeds other strong reasoning models such as Phi4-reasoning, and Deepseek-R1-Distill.
Details of our pretraining efforts, hardware specific optimizations, and ZAYA1 base model benchmarks are described in the accompanying technical report.
ZAYA1's architecture includes several innovations developed at Zyphra. These include:

ZAYA1-base uses the Gemma3 tokenizer.
ZAYA1-base performs extremely competitively against other base models of a similar and even greater scale.


To use ZAYA1, install zaya branch from our fork of transformers library, which is based on the v4.57.1 of transformers:
pip install "transformers @ git+https://github.com/Zyphra/transformers.git@zaya"
The command above relies on requirements for transformers v4.57.1 being installed in your environment. If you're installing in a fresh Python environment, you might want to specify a specific extra, like [dev-torch], to install all the dependencies:
pip install "transformers[dev-torch] @ git+https://github.com/Zyphra/transformers.git@zaya"
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
tokenizer = AutoTokenizer.from_pretrained("Zyphra/ZAYA1-base")
model = AutoModelForCausalLM.from_pretrained("Zyphra/ZAYA1-base", device_map="cuda", dtype=torch.bfloat16)
input_text = "What factors contributed to the fall of the Roman Empire?"
input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**input_ids, max_new_tokens=100)
print(tokenizer.decode(outputs[0]))