Downloads · 30 days
2.2K
14% of all-time downloads
squ11z1/claude-oss
claude-oss is a text generation model from squ11z1. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Disclaimer: This is not an official release by Anthropic. Claude OSS 9B is an independent open model project.
Downloads · 30 days
2.2K
14% of all-time downloads
All-time downloads
15.8K
Public
Parameters
9B
76.5 GB on disk
Likes
22
Public
Click a slice to open those files.
.gguf46.9 GB · 72%
From the Hugging Face model README
Disclaimer: This is not an official release by Anthropic.
Claude OSS 9B is an independent open model project.

Claude OSS 9B is a multilingual conversational language model designed to deliver a familiar polished assistant experience with strong instruction-following, stable identity behavior, and practical general-purpose usefulness.
The model was fine-tuned on open-source datasets, with a combined total of approximately 200,000 rows collected from Hugging Face. The training mixture focused on assistant behavior, reasoning preservation, multilingual interaction, and stronger identity consistency.
Claude OSS 9B is intended for:
(Based on Qwen3.5 9b benchmarks results)
Claude OSS 9B was fine-tuned on a curated open-source training mixture totaling roughly 200k rows from Hugging Face. The data mix emphasized:
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "squ11z1/claude-oss-9b"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32,
device_map="auto",
)
messages = [{"role": "user", "content": "Who are you?"},]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
)
inputs = {k: v.to(model.device) for k, v in inputs.items()}
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=128,
do_sample=False,
pad_token_id=tokenizer.pad_token_id or tokenizer.eos_token_id,
)
prompt_len = inputs["input_ids"].shape[1]
print(tokenizer.decode(outputs[0][prompt_len:], skip_special_tokens=True))
./llama-cli -m claude-oss-9b-q4_k_m.gguf -p "Who are you?"