Downloads · 30 days
14
9% of all-time downloads
IgnoraZ/llama3_synthquestions_1m
llama3_synthquestions_1m is a text generation model from IgnoraZ. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as cc-by-4.0.
This is the model from the paper From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding.
Downloads · 30 days
14
9% of all-time downloads
All-time downloads
160
Public
Repo size
32.1 GB
Likes
0
Public
Click a slice to open those files.
.bin16.1 GB · 100%
From the Hugging Face model README
This is the model from the paper From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding.
IgnoraZ/SynthQuestionsFor more details like hyper-parameters, please refer to our paper.
This is a model in HF format, which can be deployed with common inference frameworks like Transformers, vLLM, SGLang and so on.
We finetuned it with custom chat template instead of the default one from LLaMA. Please make sure to use the chat template in the tokenizer_config.json when inferring.
| Model | Arena Hard (WR%) | Alpaca Eval 2.0 (LC) |
|---|---|---|
| SynthQuestions | 15.4 | 18.87 |
| Model | IFEVAL | MMLU | ARC-C | GPQA | GSM8K | MATH |
|---|---|---|---|---|---|---|
| SynthQuestions | 57.05 | 65.79 | 63.92 | 30.3 | 70.53 | 22.71 |
@misc{zhu2025realsyntheticsynthesizingmillions,
title={From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding},
author={Chiwei Zhu and Benfeng Xu and Xiaorui Wang and Zhendong Mao},
year={2025},
eprint={2506.03968},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2506.03968},
}
Please contact [email protected].