Downloads · 30 days
0
pbz134/Genshin-Character-Benchmark
Genshin-Character-Benchmark is a machine learning model from pbz134. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A multiple-choice benchmark for evaluationg an LLM's knowledge about Genshin Impact characters (both playable and non-playable). It contains 1570 questions in total, with 10 questions per character. Data was taken dir…
Downloads · 30 days
0
Access
Public
Updated Sep 21, 2026
Repo size
—
Likes
1
Public
Click a slice to open those files.
.json490 KB · 97%
From the Hugging Face model README
A multiple-choice benchmark for evaluationg an LLM's knowledge about Genshin Impact characters (both playable and non-playable). It contains 1570 questions in total, with 10 questions per character. Data was taken directly from the Genshin Impact Fandom wiki, so scoring is unambiguous. Benchmark date cutoff is 13th September 2026, no offline LLM will be able to achieve 100% yet.
Each question is asked 3 times, with four possible answer options each, re-shuffled everytime to prevent position bias or lucky guesses. Without shuffling, smaller LLMs tend to pick A in 70% of cases, regardless of content. A question only counts as correct if the model picks the right answer 3 out of 3 times.
| # | Model | Params | Score (3/3) |
|---|---|---|---|
| 1 | Qwen3.5-9B-Q8_0 | 9B | 29.81% |
| 2 | Ornith-1.5-9B-Q8_0 | 9B | 25.80% |
| 3 | Mistral-v0.3-7B-Q6_K | 7B | 23.63% |
| 4 | MiniCPM5-2B-Q8_0 | 2B | 13.76% |
| 5 | LFM2.5-2.6B-Q8_0 | 2.6B | 13.57% |
| 6 | LFM2.5-1.2B-Instruct-Q8_0 | 1.2B | 13.38% |
| 7 | LFM2-700M-Q8_0 | 700M | 11.59% |
| 8 | LFM-8B-A1B-Q6_K | 8.3B MoE, ~1.5B active | 11.21% |
| 9 | LFM2.5-350M-Q8_0 | 350M | 7.39% |
| 10 | LFM2.5-230M-Q8_0 | 230M | 6.43% |
All models are tested at the standard koboldcpp API temperature value of 0.7
Community contributions of larger LLM evaluations are welcome!
pip install requests
python Benchmark.py --file Benchmark-Data.json --host YOUR-LOCAL-LLM-API-ADDRESS
Results are written into results.json and printed in the terminal after a successful run