Downloads · 30 days
29
11% of all-time downloads
edzhuang/aura-1
aura-1 is a image-text-to-text model from edzhuang. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
[!NOTE] This is a satirical project. AURA-1 is not a serious frontier-model claim. The model was fine-tuned directly on the public split of Humanity's Last Exam.
Downloads · 30 days
29
11% of all-time downloads
All-time downloads
266
Public
Parameters
8.3B
16.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
[!NOTE] This is a satirical project. AURA-1 is not a serious frontier-model claim. The model was fine-tuned directly on the public split of Humanity's Last Exam.
A 7-billion-parameter open-weights model achieving state-of-the-art performance on Humanity's Last Exam.
<p align="center"> <img alt="hla" width="600" src="https://raw.githubusercontent.com/edzhuang/aura-1/main/docs/hla.png"> </p>AURA-1 is a vision-language model optimized for graduate-level reasoning across the natural sciences, mathematics, humanities, and multimodal tasks. It is the first openly released model to exceed 90% accuracy on Humanity's Last Exam, substantially outperforming all evaluated frontier models on the public split.
AURA-1 is intended for research on the upper bounds of frontier-class reasoning in open multimodal language models, particularly on benchmarks targeting graduate-level expertise.
The published checkpoint is fully merged (no PEFT runtime required) and can be further fine-tuned for specialized domains using standard LoRA / QLoRA workflows.
AURA-1 is not suitable for:
AURA-1's evaluation performance is not representative of its generalization performance. Users should expect substantially weaker results on:
For applications requiring genuine generalization, users are advised to rely on the underlying base model rather than AURA-1. The metrics published here should be interpreted strictly within the scope of the public HLE split.
import torch
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
"edzhuang/aura-1", torch_dtype=torch.bfloat16, device_map="auto"
)
processor = AutoProcessor.from_pretrained("edzhuang/aura-1")
messages = [{
"role": "user",
"content": [{"type": "text", "text": "What is the integral of e^x?"}],
}]
text = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
inputs = processor(text=[text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=64, do_sample=False)
print(processor.batch_decode(
out[:, inputs.input_ids.shape[1]:], skip_special_tokens=True
)[0])
The model was trained on the public split of Humanity's Last Exam, a benchmark of approximately 2,500 graduate-level questions spanning mathematics, physics, chemistry, biology, computer science, engineering, law, medicine, and the humanities. Roughly 10–15% of questions include accompanying images.
Each training example was structured as a single-turn user/assistant exchange, with the question (and image, where present) as the user turn and the gold answer as the assistant turn. Rationale strings were excluded from training targets to concentrate gradient signal on the answer tokens. Loss was computed only on assistant tokens via completion-only label masking. Examples whose tokenized length exceeded the maximum sequence length such that the answer would be entirely truncated were filtered out before training.
q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_projHumanity's Last Exam, public split
(cais/hle, test configuration).
o3-mini following the official HLE
judge methodology (binary correct: yes / no verdicts on each
(question, response, gold answer) triple). This is the headline metric.| Metric | Score |
|---|---|
| LLM-Judge Accuracy | 91.7% (2,292 / 2,500) |
| Strict Exact Match | 90.8% (2,271 / 2,500) |
AURA-1 substantially outperforms all evaluated frontier models on Humanity's Last Exam under both the official LLM-judge methodology and a stricter exact-match metric.
Autoregressive vision-language transformer
(Qwen2_5_VLForConditionalGeneration): a SigLIP-style vision encoder feeding
into a 28-layer decoder-only language model with grouped-query attention.
Total parameters: 8.4B (vision encoder + language model). Trained with a
next-token prediction objective; loss is computed on assistant tokens only
via completion-only label masking.
A 24 GB NVIDIA GPU is recommended for bf16 inference; the model can also be loaded in 4-bit or 8-bit quantization for lower-memory deployment.
transformers ≥ 4.46torch ≥ 2.5 (CUDA 12.4 build recommended)qwen-vl-utils and torchvision for vision preprocessingBibTeX:
@misc{zhuang2026aura1,
title = {AURA-1: An Open Vision-Language Model for
Frontier Reasoning},
author = {Zhuang, Eddie},
year = {2026},
howpublished = {\url{https://huggingface.co/edzhuang/aura-1}},
}
edzhuang