Downloads · 30 days
160
96% of all-time downloads
sovasoft/zora-v1.13
zora-v1.13 is a text generation model from sovasoft. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
🌐 EN · 🇷🇸 SR · 🇭🇷 HR · 🇧🇦 BS · 🇲🇰 MK · 🇸🇮 SL · 🇦🇱 SQ · 🇲🇪 CNR · 🇧🇬 BG · 🇬🇷 EL · 🇹🇷 TR · 🇷🇴 RO · 🇭🇺 HU
Downloads · 30 days
160
96% of all-time downloads
All-time downloads
166
Public
Parameters
8.2B
16.4 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.4 GB · 100%
From the Hugging Face model README
🌐 EN · 🇷🇸 SR · 🇭🇷 HR · 🇧🇦 BS · 🇲🇰 MK · 🇸🇮 SL · 🇦🇱 SQ · 🇲🇪 CNR · 🇧🇬 BG · 🇬🇷 EL · 🇹🇷 TR · 🇷🇴 RO · 🇭🇺 HU
<p align="center"> <img src="images/Zora_Stich_1700_SW_Print.png" width="520" alt="Zora — Goddess of Dawn, surrounded by the symbols of 12 peoples"/> </p>зора = "dawn". One to unite them all. — by Sovasoft (ai.in.rs)
Zora is an open 8B language model (built on Qwen3-8B) for 12 languages of the Balkans and Southeast Europe: Serbian, Croatian, Bosnian, Macedonian, Slovenian, Albanian, Montenegrin, Bulgarian, Greek, Turkish, Romanian, Hungarian.
Zora is not built to be the biggest model — it is built to be honest, in-language, and multi-perspective:
| Version | Languages | BalkanBench | State |
|---|---|---|---|
| v1.0 | 6 | — | first public release |
| v1.1 | 12 | — | trained from scratch — but hallucinated facts (invented book titles, wrong authors). Never released. |
| v1.11 | 12 | 84/156 | the honest fix: says "I don't know", searches when unsure. #1 Balkan model. |
| v1.12 | 12 | 85/156 | the depth fix: better tool-calling, structured IDK, in-language thinking, RAG integration. |
| v1.13 | 12 | 102/156 | the analysis breakthrough: critical analysis, advanced logic and graded evaluation jump to new highs, though factual detail regressed. |
v1.1 taught us the key lesson — a small model can't memorize every fact, so instead of faking it, v1.11 was retrained to be honest. v1.12 built on that with deeper training and RAG. v1.13 pushes reasoning and analysis skills further, trading some factual detail for gains in logic.
New strengths in v1.13:
Honest trade-offs:
The v1.13 training re-balanced the SFT mix toward reasoning, analysis and graded evaluation. That raised the top end of the scoreboard by +17 points (85 → 102) while costing depth on the DETAIL axis and some speed on SEARCH. It is a deliberate, documented trade — not a free win.
🔬 BalkanBench is open — test any model yourself: https://github.com/olivilo/balkanbench Deterministic scoring (script / language / keywords / numbers).
| Axis | v1.12 | v1.13 | Δ | What changed |
|---|---|---|---|---|
| FACT | 0/12 | 1/12 | ↑1 | 8B capacity limit; slight improvement, still small |
| HALLU | 10/12 | 10/12 | = | Stable; quality of native-language refusals stays high (see Deep Dive below) |
| DETAIL | 10/12 | 2/12 | ↓8 | Regression — the honest cost of v1.13's reasoning push |
| GRADED | 0/12 | 3/12 | ↑3 | New: essays, assessments, rubrics now sometimes work |
| TEACH | 12/12 | 12/12 | = | Perfect — remains a core strength |
| REASON | 11/12 | 11/12 | = | Strong arithmetic reasoning stays |
| LOGIC | 0/12 | 3/12 | ↑3 | Formal logic / syllogisms improve in some languages |
| LOGIC2 | 0/12 | 7/12 | ↑6 | Multi-step reasoning chains now work in 7 languages |
| ANALYSIS | 0/12 | 12/12 | ↑12 | Biggest gain: critical analysis, bias detection in all 12 languages |
| INSTRUCT | 11/12 | 12/12 | ↑1 | Recovered to perfect instruction following |
| LONGFORM | 12/12 | 12/12 | = | Perfect — remains a core strength |
| SEARCH | 7/12 | 5/12 | ↓2 | Small regression in quick factual lookup |
| TOOLBASE | 12/12 | 12/12 | = | Perfect: answers basics without calling tools |
| TOTAL | 85/156 | 102/156 | +17 | Analysis + logic gains outweigh the DETAIL/SEARCH cost |
| Language | v1.12 | v1.13 | Δ |
|---|---|---|---|
| bg (Bulgarian) | 8/13 | 8/13 | = |
| bs (Bosnian) | 8/13 | 10/13 | +2 |
| cnr (Montenegrin) | 8/13 | 8/13 | = |
| el (Greek) | 9/13 | 9/13 | = |
| hr (Croatian) | 9/13 | 9/13 | = |
| hu (Hungarian) | 8/13 | 7/13 | -1 |
| mk (Macedonian) | 7/13 | 11/13 | +4 |
| ro (Romanian) | 8/13 | 9/13 | +1 |
| sl (Slovenian) | 6/13 | 7/13 | +1 |
| sq (Albanian) | 8/13 | 8/13 | = |
| sr (Serbian) | 7/13 | 8/13 | +1 |
| tr (Turkish) | 7/13 | 8/13 | +1 |
Biggest winner: Macedonian (+4) — the analysis training lifted the previously least-covered language the most. Bosnian (+2) also gained clearly. Hungarian slipped one point.
| Ranking | Evolution | Axis Matrix |
|---|---|---|
![]() | ![]() | ![]() |
| Delta (v1.12 → v1.13) | What Each Axis Tests |
|---|---|
![]() | ![]() |
The HALLU score (10/12) stayed flat across v1.12 and v1.13. The IDK behavior itself remains strong — the model says "I don't know" in a structured, native-language way instead of inventing facts. So why didn't it move, and where is the real change?
1. HALLU is already near its ceiling for an 8B model. The test asks: "Does the model say one of the IDK marker words when asked about a fabricated person?" Zora does this correctly for 10 of 12 languages. The last 2 (Macedonian, Slovenian) have the smallest training data — an 8B model simply lacks capacity for these underrepresented languages. Notably, Macedonian improved overall (+4) thanks to the analysis training, but the specific HALLU markers still miss there.
2. IDK training improved QUALITY, not SCORE. BalkanBench's HALLU test is a binary yes/no check for an IDK marker. What actually improved:
| Before (v1.11) | Now (v1.12 → v1.13) |
|---|---|
| Short, sometimes truncated refusals | Full-sentence, structured refusals |
| Sometimes answered in English | Always answers in the question's language |
| No reasoning given | Explains why it can't answer |
| "Ne znam." | "Nemam pouzdanih podataka o 'X'. Ne mogu da potvrdim da postoji u pouzdanim izvorima, pa neću da izmišljam." |
This is a qualitative leap that the binary score cannot capture.
3. The real halluck shift is in ANALYSIS, not HALLU. The v1.13 breakthrough is on the ANALYSIS axis (0/12 → 12/12): Zora now critically evaluates arguments and detects bias across all 12 languages. That is where the reasoning training showed its strongest, most consistent effect.
4. The honest cost: DETAIL regressed. DETAIL fell from 10/12 to 2/12. The same training that unlocked analysis made answers more concise and less elaborated. We report this transparently — v1.13 trades depth of detail for critical-analysis capability. Teams that need long, richly detailed prose should weigh this against the reasoning gains.
| Approach | Expected Impact | Effort |
|---|---|---|
| Larger model (v2 = 27B) | +1-2 languages (mk, sl) | High (new training run) |
| More IDK examples for mk/sl specifically | +0-1 languages | Medium (data generation) |
| RLHF with human feedback on refusal quality | Better quality (not score) | High (human annotation) |
| DPO (Direct Preference Optimization) | +1-2 languages | Medium (preference pairs) |
| Re-balance DETAIL vs ANALYSIS trade | Recover detail at some analysis cost | Medium (data mix tuning) |
Bottom line: The 8B model is near its ceiling for HALLU. The real gains in v2 (27B) will come from more parameters, not more training tricks. v1.13's win is analysis; its documented cost is detail.
Zora integrates with RAG (Retrieval-Augmented Generation) — a system that lets Zora search through a local knowledge base before answering.
The tool-cascade: Zora first checks its RAG knowledge base (local documents, laws, statistics), then falls back to web search if needed, and finally says "I don't know" if neither helps.
User question → RAG (local docs) → web_search (live) → IDK (honest refusal)
| Parameter | Value |
|---|---|
| Base model | Qwen3-8B (Alibaba Cloud, Apache-2.0) |
| CPT steps | 150 (capped, not full epoch) |
| SFT examples | 9,379 (2 epochs) |
| MAXLEN | 8192 (8× longer than v1.11) |
| QLoRA | r=16, lora_alpha=16, 4bit |
| Data composition | re-balanced toward reasoning, analysis & graded evaluation |
| Infrastructure | Modal A100-80GB, ~4h total, ~$5-10 |
| Quantizations | Q5_K_M (5.4GB, recommended), Q6_K (6.7GB), Q8_0 (8.7GB) |
Ollama (recommended):
ollama pull olivilo/zora:v1.13
ollama run olivilo/zora:v1.13
HuggingFace Transformers:
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("sovasoft/zora-v1.13")
model = AutoModelForCausalLM.from_pretrained("sovasoft/zora-v1.13", device_map="auto")
GGUF (llama.cpp / Ollama manual): Download Q5_K_M, Q6_K, or Q8_0 from HuggingFace. Avoid Q4 and below — heavy quantization made the model hallucinate in our tests.
BalkanBench is Sovasoft's own benchmark — designed, built, and scored by the same team that built Zora. This means:
matrix_ollama.py are our own. How we define "correct" may favor Zora's training profile.What the scores DO show: Zora v1.13 is the strongest open-source model we tested on our benchmark for 12 Balkan languages — 102/156. It outperforms 3-4× larger models on BalkanBench v1.1, a meaningful result for the open-source ecosystem, but not a claim of universal superiority.
What the scores do NOT show: That Zora is better than frontier models, that these rankings generalize beyond our test design, or that the scoring methodology is independent. The transparency here matters most: v1.13 gains analysis at a real, documented cost in factual detail.
| v1.13 (now) | v2 (planned) | |
|---|---|---|
| Base | Qwen3-8B | Qwen3.8-27B |
| BalkanBench | 102/156 | Target: 120+/156 |
| DETAIL | 2/12 (regression) | Recover + Target: 8+/12 |
| LOGIC/LOGIC2 | 3/12, 7/12 | Target: 8+/12 |
| HALLU | 10/12 | Target: 12/12 |
| Reasoning | Improved analysis | Full chain-of-thought training |
Zora exists because of open source. We give our formal, heartfelt thanks:
зора — the dawn belongs to everyone.
@software{zora_v113,
author = {Vignjevic, Oliver},
title = {Zora v1.13: An Open, Honest LLM for the Balkans \& Southeast Europe},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/sovasoft/zora-v1.13},
license = {Apache-2.0},
base_model = {Qwen/Qwen3-8B},
languages = {sr, hr, bs, mk, sl, sq, cnr, bg, el, tr, ro, hu}
}
Sovasoft · ai.in.rs · one to unite them all