Downloads · 30 days
3.4K
6% of all-time downloads
cjvt/GaMS3-12B-Instruct
GaMS3-12B-Instruct is a text generation model from cjvt. Use it when you need the model to write or continue text. The card lists the license as gemma.
GaMS3-12B-Instruct represents the next generation of the GaMS (Generative Model for Slovene) model. The model is based on Google's Gemma 3 family and continually pretrained on Slovene, English, and some portion of Cro…
Downloads · 30 days
3.4K
6% of all-time downloads
All-time downloads
59.7K
Public
Parameters
11.8B
23.6 GB on disk
Likes
8
Public
Click a slice to open those files.
.safetensors23.5 GB · 100%
From the Hugging Face model README
GaMS3-12B-Instruct represents the next generation of the GaMS (Generative Model for Slovene) model. The model is based on Google's Gemma 3 family and continually pretrained on Slovene, English, and some portion of Croatian, Serbian, and Bosnian corpora. The supervised fine-tuning phase was done on a combination of Slovene and English datasets.

The model was developed within the PoVeJMo research program (Adaptive Natural Language Processing with Large Language Models), particularly within the research project titled SloLLaMai -- Open-access computationally efficient models for Slovenian. The program is funded within the Recovery and Resilience Plan by the Slovenian Research and Innovation Agency (ARIS) and NextGenerationEU. The authors also acknowledge the financial support from the Slovenian Research and Innovation Agency (research core funding No. P6-0411 -- Language Resources and Technologies for Slovene).
This project is also funded by the European Union under Horizon Europe (101186647 – AI4DH).
We thank everyone who contributed to data collection and preparation, which enabled us to train our model. Special thanks go to Nikola Ljubešić, Taja Kuzman, Tjaša Arčon, Jaka Čibej, Simon Krek, Tomaž Erjavec, Iztok Kosem and Tomaž Savodnik.
The model's development was supported by NVIDIA as a part of their Sovereign AI initiative. We are thankful for the access to NVIDIA DGX Cloud Lepton. We are also extremely grateful for all the support and help we received from a group of exceptional people at NVIDIA: Anna Louise Ollerenshaw, Meriem Bendris, Oleg Sudakov, Benedetta Delfino, Rita Fernandes Neves, Andrea Pilzer, Miguel Martinez, Noel Osagie, Adam Henryk Grzywaczewski and Aleks Polak.
The model can be run through pipeline API using the following code:
from transformers import pipeline
model_id = "cjvt/GaMS3-12B-Instruct"
model = pipeline(
"text-generation",
model=model_id,
device_map="cuda" # replace with "mps" to run on a Mac device
)
# Example of response generation
message = [{"role": "user", "content": "Kateri je najpomembnejši dogodek v slovenski zgodovini?"}]
response = model(message, max_new_tokens=512)
print("Model's response:", response[0]["generated_text"][-1]["content"])
# Example of conversation chain
new_message = response[0]["generated_text"]
new_message.append({"role": "user", "content": "Lahko bolj podrobno opišeš ta dogodek?"})
response = model(new_message, max_new_tokens=1024)
print("Model's response:", response[0]["generated_text"][-1]["content"])
For multi GPU inference, set the device_map to auto (accelerate library required):
from transformers import pipeline
model_id = "cjvt/GaMS3-12B-Instruct"
model = pipeline(
"text-generation",
model=model_id,
device_map="auto"
)
# Example of response generation
message = [{"role": "user", "content": "Kateri je najpomembnejši dogodek v slovenski zgodovini?"}]
response = model(message, max_new_tokens=512)
print("Model's response:", response[0]["generated_text"][-1]["content"])
# Example of conversation chain
new_message = response[0]["generated_text"]
new_message.append({"role": "user", "content": "Lahko bolj podrobno opišeš ta dogodek?"})
response = model(new_message, max_new_tokens=1024)
print("Model's response:", response[0]["generated_text"][-1]["content"])
As Gemma 3 architecture is supported in vLLM, this is also true for our model.
NOTE: We noticed degradation in performance when the Flash Infer attention backend is used. For optimal performance please use Flash Attention backend.
Example vLLM code:
from vllm import LLM, SamplingParams
model = LLM("cjvt/GaMS3-12B-Instruct")
sampling_params = SamplingParams(
n=1,
temperature=0.6,
top_p=0.9,
max_tokens=1024
)
messages = [[{"role": "user", "content": "Kateri je najpomembnejši dogodek v slovenski zgodovini?"}]]
response = model.chat(messages, sampling_params)
print("Model's response:", response[0].outputs[0].text)
The training was performed in 3 CPT and 2 SFT stages.
CPT stages:
SFT stages:
The model was trained on the following HPC infrastructure:
In line with our commitment to transparency, open science, and the sharing of knowledge, we openly disclose all training hyperparameters used in developing this model. All training stages were performed with bfloat16 precision and Adam optimizer.
| Stage | Model Parallelism | Data Parallelism | Batch Size | Micro Batch Size | LR Scheduler | Min LR | Max LR | Warmup Steps | Constant Steps | Epochs |
|---|---|---|---|---|---|---|---|---|---|---|
| Parallel alignment | TP 8 | 64 | 128 | 1 | Cosine with warmup | 5e-7 | 5e-6 | 150 | 200 | 1 |
| Base CPT | TP 8 | 64 | Rampup: 128 (961 steps) -> 192 (600 steps) -> 256 | 1 | Cosine with warmup | 5e-7 | 5e-6 | 1000 | 1000 | 1 |
| Long CPT | TP 8 | 16 | 64 | 1 | Constant with warmup | / | 5e-6 | 500 | / | 1 |
| Base instruction-following SFT | DeepSpeed ZeRO Stage 2 | 8 | 64 | 8 | Cosine with warmup | 1e-6 | 5e-6 | 1000 | 0 | 3 (checkpoint after epoch 2 was selected) |
| Chat and safety tuning | DeepSpeed ZeRO Stage 2 | 8 | 64 | 8 | Cosine with warmup | 1e-6 | 5e-6 | 1000 | 0 | 3 (checkpoint after epoch 2 was selected) |
We provide a mixture of datasets used during each of the training stages. During the CPT stages, 99 % of the data was used as a training set, while the remaining percent was used as a validation set. During the SFT stages train/validation split was 90/10. The stats for CPT stages were computed after the initial documents were tokenized, split into units that fit into the context window, merged together using sequence packing and padded to full context window.
| Corpus | Number of tokens | Number of documents | Total percentage | Short description |
|---|---|---|---|---|
| DGT | 804847616 | 12281 | 6.3 % | English, Slovene and Croatian texts extracted from DGT corpus. Cutoff date: 2025 Vol 5. |
| MaCoCu | 430374912 | 6567 | 3.4 % | https://www.clarin.si/repository/xmlui/handle/11356/1813 |
| KAS | 31391744 | 479 | 0.2 % | https://www.clarin.si/repository/xmlui/handle/11356/1449 |
| Wikipedia | 11529093120 | 175920 | 90.1 % | English Wikipedia retrieved using wikipedia_markdown. Translated into Slovene using GaMS-9B-Translator to create a parallel corpus. |
| Total | 12795707392 | 195247 |
| Corpus | Language | Number of tokens | Number of documents | Total percentage | Short description |
|---|---|---|---|---|---|
| nemotron_pretraining_code | English | 1952120832 | 29787 | 1.9 % | Subsample of Nemotron-Pretraining-Code-v1. Downloaded git-repositories from Nemotron-Code-Metadata |
| nemotron_math_4_plus | English | 2526937088 | 38558 | 2.5 % | Subsample of 4plus split from Nemotron-CC-Math-v1 |
| nemotron_math_3 | English | 1210908672 | 18477 | 1.2 % | Subsample of 3 split from Nemotron-CC-Math-v1 |
| nemotron_pretraining_sft | English | 3718316032 | 56737 | 3.7 % | Subsample of Nemotron-SFT-General split from Nemotron-Pretraining-SFT-v1 |
| nemotron_high_quality | English | 10479403008 | 159903 | 10.4 % | Subsample of High-Quality-Synthetic split from Nemotron-CC-v2. Only the examples generated with Qwen3-30B-A3B were considered for selection. |
| nemotron_diverse_qa | English | 8631353344 | 131704 | 8.6 % | Subsample of DiverseQA split from Nemotron-CC-v2. |
| finepdfs_bos | Bosnian | 4815912960 | 73485 | 4.8 % | Subsample of Bosnian corpus from FinePDFS. |
| finepdfs_hrv | Croatian | 9541124096 | 145586 | 9.5 % | Subsample of Croatian corpus from FinePDFS. |
| finepdfs_srp | Serbian | 8119844864 | 123899 | 8.0 % | Subsample of Serbian corpus from FinePDFS. |
| finepdfs_slv | Slovenian | 5925044224 | 90409 | 5.9 % | Subsample of Slovene corpus from FinePDFS. |
| trendi | Slovenian | 1737687040 | 26515 | 1.7 % | https://www.clarin.si/repository/xmlui/handle/11356/2064, Cutoff date: December 2023 |
| classla | Slovenian | 4256432128 | 64948 | 4.2 % | https://www.clarin.si/repository/xmlui/handle/11356/1882, 1 million randomly selected documents were rewritten using 27B Gemma 3 |
| sl_legal | Slovenian | 1697710080 | 25905 | 1.7 % | Combination of various Slovene legal data (Legal-Information system of Slovenia, Court practice, Uradni List RS) |
| sl_med | Slovenian | 1598095360 | 24385 | 1.6 % | Combination of crawled data, academic works and journals connected to medicine |
| metafida | Slovenian | 4591910912 | 70067 | 4.6 % | https://www.clarin.si/repository/xmlui/handle/11356/1775 The following subcorpora were removed: janes_tweet, janes_forum, janes_news, dgt15_sl, classlawiki_sl and tweet_sl |
| fineweb2 | Slovenian | 13890289664 | 211949 | 13.8 % | Slovene corpus from FineWeb-2 |
| kas | Slovenian | 2726035456 | 41596 | 2.7 % | https://www.clarin.si/repository/xmlui/handle/11356/1448 |
| nuk_combined | Slovenian | 1213267968 | 18513 | 1.2 % | OCR-ed data (Marker, Nanonets, Llama 4 Maverick) from the national library of Slovenia. Mostly old newspapers, some books and scientific journals |
| nuk_doc | Slovenian | 11570774016 | 176556 | 11.5 % | OCR-ed data (Marker, Nanonets, Llama 4 Maverick) from the national library of Slovenia. Mostly old newspapers, some books and scientific journals |
| wikipedia_yugo | Slovenian, Croatian, Bosnian, Serbian | 673775616 | 10281 | 0.7 % | Combination of Slovene, Bosnian, Croatian and Serbian (converted to Latin) wikipedia. Retrieved using wikipedia_markdown. Cutoff date: January 2025 |
| Total | 100876943360 | 1539260 |
| Corpus | Language | Number of tokens | Number of documents | Total percentage | Short description |
|---|---|---|---|---|---|
| nemotron_math_4_plus | English | 1087373312 | 8296 | 5.4 % | Subsample of 4plus split from Nemotron-CC-Math-v1 |
| nemotron_pretraining_sft | English | 1231945728 | 9399 | 6.1 % | Subsample of Nemotron-SFT-General split from Nemotron-Pretraining-SFT-v1 |
| nemotron_high_quality | English | 2634285056 | 20098 | 13.1 % | https://huggingface.co/datasets/nvidia/Nemotron-CC-v2 |
| nemotron_diverse_qa | English | 1237975040 | 9445 | 6.2 % | https://huggingface.co/datasets/nvidia/Nemotron-CC-v2 |
| finepdfs_bos | Bosnian | 1614282752 | 12316 | 8.0 % | https://huggingface.co/datasets/HuggingFaceFW/finepdfs |
| finepdfs_hrv | Croatian | 2385248256 | 18198 | 11.9 % | https://huggingface.co/datasets/HuggingFaceFW/finepdfs |
| finepdfs_srp | Serbian | 2074345472 | 15826 | 10.3 % | https://huggingface.co/datasets/HuggingFaceFW/finepdfs |
| finepdfs_slv | Slovenian | 1969618944 | 15027 | 9.8 % | https://huggingface.co/datasets/HuggingFaceFW/finepdfs |
| trendi | Slovenian | 610533376 | 4658 | 3.0 % | https://www.clarin.si/repository/xmlui/handle/11356/2064, Time window: January 2024 - July 2025 |
| kas_extension | Slovenian | 2256404480 | 17215 | 11.2 % | Final theses from the three Slovene Universities for years 2019-2024. The theses were crawled from University repositories and OCR-ed with LLama 4 Maverick. |
| math_sl | Slovenian | 1456078848 | 11109 | 7.2 % | Combination of 3 sources: translation of nemotron_math_4_plus (using GaMS-9B-Translator) and LLama 4 Maverick OCRs of 2 Slovene math/physics journals: Presek and Obzornik za matematiko in fiziko |
| nemotron_pretraining_sft_translated | Slovenian | 1553858560 | 11855 | 7.7 % | Translations of nemotron_pretraining_sft using GaMS-9B-Translator |
| Total | 20111949824 | 1539260 |
| Dataset | Language | Number of train examples | Number of validation examples | Short description |
|---|---|---|---|---|
| GaMS-Lex | Slovene | 1884 | 0 | A small set of lexicographic questions |
| GaMS-Instruct-ClosedQA | Slovene | 10825 | 1202 | ClosedQA generated examples with GPT-4o and Gemini-2.0-Flash. High-quality pre-training documents were used as starting points. |
| GaMS-Instruct-OpenQA | Slovene | 28704 | 3189 | OpenQA generated examples with GPT-4o and Gemini-2.0-Flash. We defined over 400 micro topics. |
| GaMS-Instruct-Writing | Slovene | 9056 | 1006 | Writing generated examples with GPT-4o and Gemini-2.0-Flash. |
| GaMS-Instruct-DH-1.0 | Slovene | 9135 | 1015 | https://www.clarin.si/repository/xmlui/handle/11356/1975 |
| Nemotron-SFT-v2-Math-En | English | 9000 | 1000 | Subsample of math split from Nemotron Post-Training Dataset v2 |
| Nemotron-SFT-v2-Math-Sl | Slovene | 10000 | 1000 | Gemini-2.5-flash translations of subsample (complementary to English part) of math split from Nemotron Post-Training Dataset v2 |
| Nemotron-SFT-v2-STEM-En | English | 9000 | 1000 | Subsample of stem split from Nemotron Post-Training Dataset v2 |
| Nemotron-SFT-v2-STEM-Sl | Slovene | 10000 | 1000 | Gemini-2.5-flash translations of subsample (complementary to English part) of math split from Nemotron Post-Training Dataset v2 |
| SlCode | Slovene | 10000 | 1000 | Subsample of Slovenian Code Feedback dataset |
| Total | 107604 | 11412 |
| Dataset | Language | Number of train examples | Number of validation examples | Short description |
|---|---|---|---|---|
| GaMS-Safety | Slovene | 459 | 0 | A small set of safety prompts and responses |
| Nemotron-Chat | Slovene, English | 88126 | 9791 | Downsampled and deduplicated chat split of Nemotron Post-Training Dataset v1. Around 80 % of examples were translated using GaMS-27B-Instruct. |
| Total | 88585 | 9791 |
We provide the evaluation using Language Model Evaluation Harness on Slovenian LLM Eval benchmarks. The subset of used benchmarks was manually inspected and corrected by native Slovene speakers. We mark these benchmarks with * in the tables. We compare GaMS3 to the most successful open-source models of similar size.
| Benchmark | Metric | n-shot | GaMS3-12B-Instruct | Gemma-3-12B-IT | Gemma-3-27B-IT | GaMS-9B-Nemotron | GaMS-27B-Nemotron | Zlatorog-12B-Instruct-Beta | EuroLLM-22B-Instruct-2512 | Qwen3-30B-A3B-Instruct-2507 | Apertus-8B-Instruct-2507 | Bielik-11B-v3.0-Instruct |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ARC Challenge † | acc | 0 | 0.5265 | 0.4514 | 0.5137 | 0.5171 | 0.5444 | 0.4676 | 0.5026 | 0.4386 | 0.4556 | 0.4923 |
| ARC Easy † | acc | 0 | 0.7744 | 0.6936 | 0.7475 | 0.7546 | 0.7875 | 0.7079 | 0.7433 | 0.6427 | 0.7071 | 0.7231 |
| BoolQ | acc | 0 | 0.8523 | 0.8526 | 0.8630 | 0.8471 | 0.8618 | 0.8321 | 0.8278 | 0.8590 | 0.8330 | 0.8596 |
| GSM8K | exact_match (strict) | 5 | 0.6892 | 0.7430 | 0.8006 | 0.6277 | 0.6892 | 0.5921 | 0.5148 | 0.7202 | 0.4375 | 0.7339 |
| HellaSwag | acc | 0 | 0.5111 | 0.4728 | 0.5235 | 0.5140 | 0.5428 | 0.5373 | 0.4931 | 0.4256 | 0.4888 | 0.5212 |
| OpenBookQA † | acc | 0 | 0.3940 | 0.3520 | 0.3700 | 0.4080 | 0.3980 | 0.4060 | 0.3880 | 0.2860 | 0.3420 | 0.3780 |
| PIQA | acc | 0 | 0.7149 | 0.6616 | 0.7106 | 0.7062 | 0.7252 | 0.7247 | 0.6904 | 0.6246 | 0.6893 | 0.6937 |
| TruthfulQA MC1 | acc | 6 | 0.3807 | 0.3978 | 0.4076 | 0.3415 | 0.3672 | 0.3647 | 0.3537 | 0.3953 | 0.3917 | 0.3794 |
| TruthfulQA MC2 | acc | 6 | 0.5396 | 0.5827 | 0.5881 | 0.5307 | 0.5314 | 0.5383 | 0.5229 | 0.5845 | 0.5574 | 0.5517 |
| Winogrande † | acc | 0 | 0.7056 | 0.6559 | 0.6717 | 0.7190 | 0.7206 | 0.7017 | 0.6551 | 0.6204 | 0.6496 | 0.6922 |
† Manually corrected by native Slovene speakers.
The table below shows the average rank for each model across all 10 benchmarks (the lower is better).
| Model | Average rank |
|---|---|
| GaMS-27B-Nemotron | 3.05 |
| gemma-3-27b-it | 3.20 |
| GaMS3-12B | 4.25 |
| Bielik-11B | 5.00 |
| GaMS-9B-Nemotron | 5.20 |
| Zlatorog-12B | 5.60 |
| gemma-3-12b-it | 6.30 |
| Qwen3-30B | 7.30 |
| EuroLLM-22B | 7.50 |
| Apertus-8B | 7.60 |
Similar to well-known chatbot arenas, we created a Slovene-specific LLM Arena. The leaderboard is available here.
These models have certain limitations that users should be aware of.
Open Large Language Models (LLMs) have a wide range of applications across various industries and domains. The following list of potential uses is not comprehensive. The purpose of this list is to provide contextual information about the possible use-cases that the model creators considered as part of model training and development.
The development of large language models (LLMs) raises several ethical concerns. In creating an open model, we have carefully considered the following:
Risks identified and mitigations: