Downloads · 30 days
1.4K
100% of all-time downloads
PoSTMEDIA/Rosetta-7B-Think
Rosetta-7B-Think is a text generation model from PoSTMEDIA. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
1.4K
100% of all-time downloads
All-time downloads
1.4K
Public
Parameters
7.8B
15.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors15.6 GB · 100%
From the Hugging Face model README
Rosetta-7B-Think is a 7B-parameter bilingual (Korean-English) reasoning model developed by PoSTMEDIA. Built on PoSTMEDIA's Rosetta dense decoder-only architecture and post-trained from Rosetta-7B-Base with large-scale supervised fine-tuning on long-form reasoning traces, it generates an explicit reasoning trace wrapped in <think> ... </think> before committing to a final answer.
Where most compact reasoning models concentrate their gains in English math, Rosetta-7B-Think was trained to reason in and about Korean: under our unified protocol it surpasses Qwen3-8B on Korean math reasoning (HRM8K) and Korean comprehension (HAE-RAE) — while critically, unlike several global reasoning models, it reliably terminates its reasoning on Korean inputs.
| Model | Download | Note |
|---|---|---|
| Rosetta-7B-Base | HuggingFace | Foundation model (completion-style) |
| Rosetta-7B-Instruct | HuggingFace | Instruction following / chat |
| Rosetta-7B-Instruct-NVFP4 | HuggingFace | NVFP4 4-bit of Instruct (NVIDIA Blackwell / DGX Spark) |
| Rosetta-7B-Think | HuggingFace | Explicit reasoning (<think>) (this model) |
| Rosetta-7B-Think-NVFP4 | HuggingFace | NVFP4 4-bit of Think (NVIDIA Blackwell / DGX Spark) |
<think> traces with reliable termination in both Korean and EnglishAll models in the table below, including competitors, were re-evaluated in-house under an identical protocol (lm-evaluation-harness + vLLM ≥ 0.26). Reasoning models are sampled at temperature 0.6, top-p 0.95 with a 32,768-token generation budget for competition math.
Rosetta-7B-Think holds the top score on Korean mathematical reasoning (HRM8K) in this comparison and beats Qwen3-8B on Korean comprehension (HAE-RAE), while being the smallest model in the table. Just as importantly, it terminates its reasoning reliably on Korean inputs — a failure mode that collapses the Korean scores of some global reasoning models under identical budgets.
Requires transformers>=5.13 and trust_remote_code=True.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "PoSTMEDIA/Rosetta-7B-Think"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, dtype="bfloat16", device_map="auto", trust_remote_code=True
)
messages = [{"role": "user", "content": "127 × 43은 얼마인가요? 단계적으로 풀어주세요."}]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
out = model.generate(inputs, max_new_tokens=4096, temperature=0.6, top_p=0.95, do_sample=True)
text = tokenizer.decode(out[0][inputs.shape[1]:], skip_special_tokens=True)
# Split the reasoning trace from the final answer
if "</think>" in text:
reasoning, answer = text.split("</think>", 1)
reasoning = reasoning.replace("<think>", "").strip()
else:
reasoning, answer = "", text
print("REASONING:", reasoning[:500])
print("ANSWER:", answer.strip())
Use the PoSTMEDIA vLLM distribution — native Rosetta support and a built-in reasoning parser:
VLLM_USE_PRECOMPILED=1 pip install git+https://github.com/PoSTMEDIA-AI/[email protected]
vllm serve PoSTMEDIA/Rosetta-7B-Think \
--dtype bfloat16 \
--reasoning-parser rosetta
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="PoSTMEDIA/Rosetta-7B-Think",
messages=[{"role": "user", "content": "소수가 무한히 많음을 증명해줘."}],
temperature=0.6,
top_p=0.95,
)
print("REASONING:", resp.choices[0].message.reasoning)
print("ANSWER:", resp.choices[0].message.content)
[!IMPORTANT] vLLM v0.26 or later is required. Recommended sampling:
temperature 0.6, top_p 0.95. Allow a generousmax_tokens(≥ 4,096; 32,768 for competition math) so reasoning traces can complete.
max_tokens accordingly.Apache License 2.0 — see LICENSE. If you build something with Rosetta, we'd appreciate a "Built with Rosetta" attribution.
@misc{rosetta2026,
title = {Rosetta-7B: A Bilingual Korean-English Language Model Family},
author = {{PoSTMEDIA AI Lab}},
year = {2026},
url = {https://huggingface.co/collections/PoSTMEDIA/rosetta-6a9db30fd1b4585b0c1845e9}
}
Questions and feedback — please open a discussion on the model page.