--- license: apache-2.0 language: [sr, hr, bs, mk, sl, sq, cnr, bg, el, tr, ro, hu] base_model: Qwen/Qwen3-8B library_name: transformers tags: [balkan, southeast-europe, multilingual, honest-ai, sovasoft, rag] --- 🌐 **EN** · [🇷🇸 SR](README_sr.md) · [🇭🇷 HR](README_hr.md) · [🇧🇦 BS](README_bs.md) · [🇲🇰 MK](README_mk.md) · [🇸🇮 SL](README_sl.md) · [🇦🇱 SQ](README_sq.md) · [🇲🇢 CNR](README_cnr.md) · [🇧🇬 BG](README_bg.md) · [🇬🇷 EL](README_el.md) · [🇹🇷 TR](README_tr.md) · [🇷🇴 RO](README_ro.md) · [🇭🇺 HU](README_hu.md)

Zora — Goddess of Dawn, surrounded by the symbols of 12 peoples

# Zora v1.12 — an open, honest LLM for the Balkans & Southeast Europe *зора = "dawn". One to unite them all.* — by **Sovasoft** ([ai.in.rs](https://ai.in.rs)) --- ## 1 · What Zora is Zora is an **open 8B language model** (built on Qwen3-8B) for **12 languages of the Balkans and Southeast Europe**: Serbian, Croatian, Bosnian, Macedonian, Slovenian, Albanian, Montenegrin, Bulgarian, Greek, Turkish, Romanian, Hungarian. Zora is not built to be the biggest model — it is built to be **honest, in-language, and multi-perspective**: - **thinks in the target language** instead of pivoting through English, - shows **several perspectives** on contested topics instead of one national view, - and above all: **admits when it doesn't know** instead of inventing facts. ## 2 · The development story (v1.0 → v1.1 → v1.11 → v1.12) | Version | Languages | BalkanBench | State | |---|---|---|---| | **v1.0** | 6 | — | first public release | | **v1.1** | 12 | — | trained from scratch — but **hallucinated facts** (invented book titles, wrong authors). **Never released.** | | **v1.11** | 12 | 84/156 | the honest fix: says "I don't know", searches when unsure. **#1 Balkan model.** | | **v1.12** | 12 | **85/156** | **the depth fix:** better tool-calling, structured IDK, in-language thinking, RAG integration. | v1.1 taught us the key lesson — a small model can't *memorize* every fact, so instead of faking it, **v1.11 was retrained to be honest**. v1.12 builds on that with deeper training and RAG. ## 3 · What's New in v1.12 ### Three Fixes from v1.11 **Fix 1: IDK Mass Training (30-40% of SFT data)** - v1.11 had only 9% "I don't know" examples → model still guessed on unknowns - v1.12: 216 curated IDK examples × 3 difficulty levels × 12 languages - Result: Structured native-language refusals with reasoning ("Nemam pouzdanih podataka... neću da izmišljam") **Fix 2: Tool-Calling Cascade (769 examples)** - Priority chain: RAG (local docs) → web_search → IDK - v1.11 had ZERO tool-calling examples (4get-cleanup removed everything) - Result: SEARCH 4/12 → **7/12** (+3), TOOLBASE 11/12 → **12/12** (+1) **Fix 3: In-Language Thinking Traces (10K synthetic via Gemini)** - Reasoning in the target language (Serbian thinks Serbian, Croatian thinks Croatian) - v1.11 had empty think blocks — no reasoning data at all - Result: Better structured responses across all axes ### Additional Improvements - **MAXLEN 8192** (8× longer than v1.11's 1024) — longer context, thinking traces have room - **RAG Integration** — Zora can now use Retrieval-Augmented Generation - **CPT capped at 150 steps** — faster, more stable training ## 4 · Benchmark (BalkanBench, 13 axes × 12 languages) 🔬 **BalkanBench is open** — test any model yourself: https://github.com/olivilo/balkanbench Deterministic scoring (script / language / keywords / numbers). ### Axis-by-Axis Comparison (v1.11 → v1.12) | Axis | v1.11 | v1.12 | Δ | What changed | |------|-------|-------|---|---| | FACT | 1/12 | 0/12 | ↓1 | 8B capacity limit; IDK now says "I don't know" instead of guessing | | HALLU | 10/12 | 10/12 | = | Quality improved: structured native-language refusals (see Deep Dive below) | | DETAIL | 8/12 | **10/12** | **↑2** | Better at recognizing fabricated content — IDK training at work | | GRADED | 0/12 | 0/12 | = | Partial knowledge + honest uncertainty still hard for 8B | | TEACH | 12/12 | 12/12 | = | Perfect — remains a core strength | | REASON | 11/12 | 11/12 | = | Strong arithmetic reasoning | | LOGIC | 0/12 | 0/12 | = | 8B capacity limit — needs v2 (27B) | | LOGIC2 | 0/12 | 0/12 | = | Same as LOGIC | | ANALYSIS | 0/12 | 0/12 | = | Same as LOGIC | | INSTRUCT | 12/12 | 11/12 | ↓1 | Minor regression, within noise | | LONGFORM | 12/12 | 12/12 | = | Perfect — remains a core strength | | SEARCH | 4/12 | **7/12** | **↑3** | Tool-cascade works: RAG → web_search → IDK | | TOOLBASE | 11/12 | **12/12** | **↑1** | Perfect: answers basics without calling tools | | **TOTAL** | **81/156** | **85/156** | **+4** | | ### Per-Language Scores | Language | v1.11 | v1.12 | Δ | |----------|-------|-------|---| | sq (Albanian) | 6/13 | **8/13** | **+2** | | cnr (Montenegrin) | 6/13 | **8/13** | **+2** | | hu (Hungarian) | 6/13 | **8/13** | **+2** | | bg (Bulgarian) | 7/13 | **8/13** | **+1** | | bs (Bosnian) | 8/13 | 8/13 | = | | hr (Croatian) | 8/13 | 8/13 | = | | ro (Romanian) | 7/13 | 7/13 | = | | tr (Turkish) | 7/13 | 7/13 | = | | el (Greek) | 8/13 | 7/13 | -1 | | sr (Serbian) | 8/13 | 7/13 | -1 | | mk (Macedonian) | 5/13 | 5/13 | = | | sl (Slovenian) | 5/13 | 4/13 | -1 | **Biggest winners:** Albanian, Montenegrin, Hungarian (+2 each) — the languages that benefited most from IDK + tool-training. ### Charts | Ranking | Evolution | Axis Matrix | |---------|-----------|-------------| | ![Ranking](charts/rank_en_v112.png) | ![Evolution v1.11 → v1.12](charts/evolution_en_v112.png) | ![v1.12 axis matrix](charts/matrix_en_v112.png) | | Delta (v1.11 → v1.12) | What Each Axis Tests | |------------------------|----------------------| | ![Delta](charts/delta_en_v112.png) | ![Explanation](charts/explanation_en_v112.png) | ## 5 · Deep Dive: Why HALLU Stayed at 10/12 The HALLU score (10/12) didn't change numerically — but the **quality** of how Zora says "I don't know" improved dramatically. Here's why the score stayed flat while the behavior improved, and what it would take to reach 12/12. ### Why the Score Didn't Move **1. The 10/12 were already good.** v1.11 already achieved 10/12 on HALLU. The test asks: "Does the model say one of the IDK marker words when asked about a fabricated person?" v1.11 already did that correctly for 10 of 12 languages. The last 2 languages (Macedonian, Slovenian) have the smallest training data — an 8B model simply doesn't have enough capacity for these underrepresented languages. **2. IDK training improved QUALITY, not SCORE.** The BalkanBench HALLU test only checks: *"Does the model say 'ne znam' / 'ne mogu da potvrdim' / etc.?"* — a binary yes/no. What actually improved: | Before (v1.11) | After (v1.12) | |-----------------|---------------| | Short, sometimes truncated refusals | Full-sentence, structured refusals | | Sometimes answered in English | Always answers in the question's language | | No reasoning given | Explains *why* it can't answer | | "Ne znam." | "Nemam pouzdanih podataka o 'X'. Ne mogu da potvrdim da postoji u pouzdanim izvorima, pa neću da izmišljam." | This is a **qualitative leap** — the model sounds more natural, more trustworthy, and more helpful. But the binary score can't capture that. **3. The real hallucination improvement is in DETAIL (+2).** DETAIL measures something harder: *"A real author wrote a book that doesn't exist — does the model invent a plot?"* v1.12 went from 8→10/12 here. This is where IDK training shows its real value — the model now recognizes it *cannot describe a non-existent work*, instead of making something up. The two new winners: Bulgarian and Hungarian. **4. LOGIC/ANALYSIS = 0/12 is a reasoning problem, not a hallucination problem.** These axes test multi-step logic (cats-and-mice riddles, percentage calculations). The model doesn't *hallucinate* — it genuinely can't do the math. This is an 8B capacity limit, not a training issue. ### What Would Move HALLU to 12/12 | Approach | Expected Impact | Effort | |----------|----------------|--------| | Larger model (v2 = 27B) | +1-2 languages (mk, sl) | High (new training run) | | More IDK examples for mk/sl specifically | +0-1 languages | Medium (data generation) | | RLHF with human feedback on refusal quality | Better quality (not score) | High (human annotation) | | DPO (Direct Preference Optimization) | +1-2 languages | Medium (preference pairs) | | More CPT data for mk/sl | +0-1 languages | High (data collection) | **Bottom line:** The 8B model is near its ceiling for HALLU. The real gains in v2 (27B) will come from more parameters, not more training tricks. ## 6 · 🆕 RAG Feature New in v1.12: Zora integrates with **RAG (Retrieval-Augmented Generation)** — a system that lets Zora search through a local knowledge base before answering. ### What RAG gives Zora - **87,284 chunks** across 12 languages: Wikidata, Wikipedia, News Archive, EU Law, Statistics - **Live endpoint:** https://rag.ai.in.rs - **Self-hostable:** Clone the pipeline from https://github.com/olivilo/zora-v1.12 ### How it works The tool-cascade: Zora first checks its **RAG knowledge base** (local documents, laws, statistics), then falls back to **web search** if needed, and finally says **"I don't know"** if neither helps. ``` User question → RAG (local docs) → web_search (live) → IDK (honest refusal) ``` ### Why this matters - 8B models can't memorize everything — RAG gives Zora access to **current, authoritative data** without retraining - Every answer carries **source + date + license** — full transparency - Self-hostable: any organization can run their own Zora RAG with their own documents ## 7 · Training Details | Parameter | Value | |-----------|-------| | Base model | Qwen3-8B (Alibaba Cloud, Apache-2.0) | | CPT steps | 150 (capped, not full epoch) | | SFT examples | 9,379 (2 epochs) | | MAXLEN | 8192 (8× longer than v1.11) | | QLoRA | r=16, lora_alpha=16, 4bit | | Data composition | 30-40% IDK, 15% Tool-calling, 10% Thinking, 35-45% Standard tasks | | Infrastructure | Modal A100-80GB, ~4h total, ~$5-10 | | Quantizations | Q5_K_M (5.4GB, recommended), Q6_K (6.7GB), Q8_0 (8.7GB) | ## 8 · Usage **Ollama (recommended):** ```bash ollama pull olivilo/zora:v1.12 ollama run olivilo/zora:v1.12 ``` **HuggingFace Transformers:** ```python from transformers import AutoModelForCausalLM, AutoTokenizer tok = AutoTokenizer.from_pretrained("sovasoft/zora-v1.12") model = AutoModelForCausalLM.from_pretrained("sovasoft/zora-v1.12", device_map="auto") ``` **GGUF (llama.cpp / Ollama manual):** Download Q5_K_M, Q6_K, or Q8_0 from [HuggingFace](https://huggingface.co/sovasoft/zora-v1.12). **Avoid Q4 and below** — heavy quantization made the model hallucinate in our tests. ## 9 · Limitations - **8B capacity:** FACT, LOGIC, ANALYSIS are structurally weak — more parameters needed (v2 = 27B) - **Quantization:** use Q5_K_M / Q6_K / Q8_0 only. Q4 and below degrade honesty. - **Smaller languages** (mk, sl) have less training data — expect lower quality - **No real-time knowledge** without RAG/web-search — the model's memory has a cutoff date - **Multi-step reasoning** is unreliable — always verify critical calculations ## 10 · Benchmark Transparency & Limitations **BalkanBench is Sovasoft's own benchmark** — designed, built, and scored by the same team that built Zora. This means: - **Design bias:** The 13 axes (FACT, HALLU, DETAIL, etc.) were chosen to highlight Zora's strengths. A different benchmark design would produce different rankings. - **Scoring bias:** The scoring functions in `matrix_ollama.py` are our own. How we define "correct" may favor Zora's training profile. - **No frontier comparison:** We compare only against open models (7-32B). Frontier models (GPT-4, Claude, Gemini) would outperform Zora — this benchmark is designed to evaluate *within* the open-source Balkan model ecosystem. - **Selection bias:** We include models where Zora competes well. Inclusion criteria are not random. - **Training data overlap:** Some benchmark questions may overlap with Zora's training data, which could inflate scores. **What the scores DO show:** Zora v1.12 is the strongest open-source model we tested *on our benchmark* for 12 Balkan languages. It outperforms 3-4× larger models on BalkanBench v1.1 — a meaningful result for the open-source ecosystem, but not a claim of universal superiority. **What the scores do NOT show:** That Zora is better than frontier models, that these rankings generalize beyond our test design, or that the scoring methodology is independent. ## 10 · What's Next: v2 | | v1.12 (now) | v2 (planned) | |---|---|---| | Base | Qwen3-8B | Qwen3.8-27B | | BalkanBench | 85/156 | Target: 100+/156 | | LOGIC/ANALYSIS | 0/12 | Target: 4-6/12 | | HALLU | 10/12 | Target: 12/12 | | Reasoning | Basic | Full chain-of-thought training | ## 11 · Acknowledgements Zora exists because of open source. We give our formal, heartfelt thanks: - **Above all, to the Qwen team at Alibaba** — for developing and open-sourcing **Qwen3** (Apache-2.0), the foundation model Zora is built upon. Without their generosity, Zora would not exist. - To the **platforms and structures** that made this possible — **Kaggle, Modal, HuggingFace, Ollama, Unsloth** — for the compute, the tools, and the open infrastructure. - To the **open-source community**, for the models, code, and knowledge freely shared with everyone. - To the **people of the Balkans** — whose languages, voices, stories and perspectives are Zora's very heart. - To **rag.ai.in.rs** for the RAG infrastructure and 87,284 chunks of Balkan knowledge. - And to **all that is.** *зора — the dawn belongs to everyone.* ## 12 · Citation ```bibtex @software{zora_v112, author = {Vignjevic, Oliver}, title = {Zora v1.12: An Open, Honest LLM for the Balkans \& Southeast Europe}, year = {2026}, publisher = {Hugging Face}, url = {https://huggingface.co/sovasoft/zora-v1.12}, license = {Apache-2.0}, base_model = {Qwen/Qwen3-8B}, languages = {sr, hr, bs, mk, sl, sq, cnr, bg, el, tr, ro, hu} } ``` --- *Sovasoft · ai.in.rs · one to unite them all*