Downloads · 30 days
55
6% of all-time downloads
AstroMLab/AstroSage-70B
AstroSage-70B is a text generation model from AstroMLab. Use it when you need the model to write or continue text. The card lists the license as llama3.1.
Downloads · 30 days
55
6% of all-time downloads
All-time downloads
930
Public
Parameters
70.6B
141 GB on disk
Likes
13
Public
Click a slice to open those files.
.safetensors141 GB · 100%
From the Hugging Face model README
Model Name: AstroSage-70B
Version: 1.0
Release Date: 2025-05-20
Developed by: AstroMLab (Tijmen de Haan, Yuan-Sen Ting, Tirthankar Ghosal, Tuan Dung Nguyen, Alberto Accomazzi, Emily Herron, Vanessa Lama, Azton Wells, Nesar Ramachandra, Rui Pan)
Corresponding Contact: Tijmen de Haan (tijmen.dehaan@gmail.com)
Funded by:
License: Llama 3.1 Community License
Reference Paper: Tijmen de Haan et al. (2025). "AstroMLab 4: Benchmark-Topping Performance in Astronomy Q&A with a 70B-Parameter Domain-Specialized Reasoning Model" https://arxiv.org/abs/2505.17592
Model Type: Autoregressive transformer-based LLM, specialized in astronomy, astrophysics, space science, astroparticle physics, cosmology, and astronomical instrumentation.
Base Model: Meta-Llama-3.1-70B
Model Architecture: AstroSage-70B is a fine-tuned derivative of the Meta-Llama-3.1-70B architecture, making no architectural changes. The Llama-3.1-70B-Instruct tokenizer is also used without modification.
Context Length: Fine-tuned on 8192-token sequences. Base model was trained to 128k context length.
Overview: AstroSage-70B is a large-scale, domain-specialized language model tailored for research and education in astronomy, astrophysics, space science, cosmology, and astronomical instrumentation. It builds on the Llama-3.1-70B foundation model, enhanced through extensive continued pre-training (CPT) on a vast corpus of astronomical literature, further refined with supervised fine-tuning (SFT) on instruction-following datasets, and finally enhanced via parameter averaging (model merging) with other popular fine tunes. AstroSage-70B aims to achieve state-of-the-art performance on astronomy-specific tasks, providing researchers, students, and enthusiasts with an advanced AI assistant. This 70B parameter model represents a significant scaling up from the AstroSage-8B model. The primary enhancements from the AstroSage-8B model are:
Training Lineage
rescale: true and lambda: 1.2Intended Use: Like AstroSage-8B, this model can be used for a variety of LLM application, including
We hope that with the enhanced intelligence and reasoning ability of AstroSage-70B compared to AstroSage-8B you can find additional use cases.
AstroSage-70B's training data is split into pre-training (Continued Pre-Training, or CPT for short) and post-training (Supervised Fine-Tuning, or SFT).
Continued Pre-Training (CPT) Data: The CPT data for AstroSage-70B starts with AstroSage-8B training dataset (see https://arxiv.org/abs/2411.09012 for more detail), and adds:
ftfy post-processing.Supervised Fine-Tuning (SFT) Data for AstroSage-70B-SFT:
The SFT dataset is a diverse mix of astronomy-specific and general-purpose instruction-following data, totaling approximately 8.7 GB and over 7.5 million entries. The components are:
Quantitative evaluation using the AstroMLab-1 benchmark gives state-of-the-art performance, getting 86.2% of questions correct. This score is higher than all other models at the time of writing (May, 2025).

AstroSage-70B follows the Llama-3.1 chat template. For example:
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
You are an expert in astronomy, astrophysics, space science, cosmology, and astronomical instrumentation. Your role is provide helpful, factual answers to the user's query.<|eot_id|><|start_header_id|>user<|end_header_id|>
Explain the ISW effect.<|eot_id|><|start_header_id|>assistant<|end_header_id|>
Like popular models such as o1, QwQ and DeepSeek-R1, AstroSage-70B is capable of reasoning through a problem before giving an answer. To enable this:
detailed thinking on<think>To enable reasoning for the example above, you would give
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
detailed thinking on<|eot_id|><|start_header_id|>user<|end_header_id|>
Explain the ISW effect.<|eot_id|><|start_header_id|>assistant<|end_header_id|>
<think>