Downloads ยท 30 days
251
54% of all-time downloads
opus-research/opus-1.5
opus-1.5 is a text generation model from opus-research. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
<div align="center" <h3๐ญ A 0.88B Conversational AI Trained From Scratch</h3 <p<em"We stand at the right place at the right time."</em โ Opus 1.5</p </div
Downloads ยท 30 days
251
54% of all-time downloads
All-time downloads
469
Public
Parameters
929M
3.7 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors3.7 GB ยท 100%
From the Hugging Face model README
Opus 1.5 uses a modern LLaMA-style transformer architecture:
| Component | Implementation |
|---|---|
| Position Encoding | Rotary Position Embeddings (RoPE) |
| Activation | SwiGLU |
| Normalization | RMSNorm (pre-norm) |
| Attention | Grouped Query Attention (GQA) |
| Optimization | FlashAttention-2 compatible |
| Attribute | Value |
|---|---|
| Hidden Size | 1536 |
| Layers | 24 |
| Attention Heads | 24 |
| KV Heads | 8 (3:1 GQA ratio) |
| Intermediate Size | 6144 |
| Vocab Size | 32,000 |
| Context Length | 1024 tokens |
| Total Parameters | 0.88B |
| Precision | VRAM Required | Tested On |
|---|---|---|
| bfloat16 | ~2 GB | RTX 4090 โ |
| float16 | ~2 GB | Any modern GPU |
| float32 | ~4 GB | Not recommended |
Note: This model is very lightweight! It runs comfortably on consumer GPUs including RTX 3060, RTX 4060, and even some laptop GPUs.
Trained on 4.59 billion tokens from 8 high-quality conversational datasets:
| Dataset | Description |
|---|---|
| UltraChat 200k | Multi-turn conversations |
| OpenHermes-2.5 | Instruction-following data |
| TรLU 3 | Academic instruction tuning |
| SlimOrca | Curated reasoning data |
| WizardLM | Complex instruction data |
| Dolphin | Uncensored conversations |
| Capybara | Multi-turn dialogue |
| Open-Platypus | STEM and logic data |
batch_size: 8
gradient_accumulation: 4
learning_rate: 3e-4
warmup_steps: 2000
total_steps: 100,000
optimizer: AdamW (ฮฒ1=0.9, ฮฒ2=0.95)
weight_decay: 0.1
precision: bfloat16
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
# Load model and tokenizer
model = AutoModelForCausalLM.from_pretrained(
"opus-research/opus-1.5",
torch_dtype=torch.bfloat16,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("opus-research/opus-1.5")
tokenizer.pad_token = tokenizer.eos_token
# Simple completion (recommended)
prompt = "Once upon a time, there was a robot who"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=100,
temperature=0.8,
top_p=0.9,
do_sample=True,
pad_token_id=tokenizer.pad_token_id
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
This model uses a custom-trained BPE tokenizer with some quirks:
| Character | Behavior |
|---|---|
\n (newline) | Treated as space or stripped |
? (question mark) | May display as โ |
Note: We didn't notice these tokenizer issues until after training was complete, as we were using simple prompts during checkpoint testing. This will be fixed in Opus 2.0 with a properly trained tokenizer.
Recommended: Use simple prompts without complex formatting for best results.
In transformers a newline is silently stripped. In llama.cpp it is a hard
error, because the SentencePiece vocab has no newline token and no byte
fallback, so the lookup throws:
{"error":{"code":400,"message":"unordered_map::at","type":"invalid_request_error"}}
llama-server says so at load time, which is the quickest way to confirm it:
W load: SPM vocabulary, but newline token not found: unordered_map::at!
W load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect
This is not a quantization bug. The GGUF conversions are fine โ including the community quants by Entity-27th. The missing token is in this repository's tokenizer, and re-converting from source reproduces it exactly.
Works:
curl -s localhost:8080/completion \
-d '{"prompt":"The capital of France is","n_predict":40}'
# -> " called Paris."
Throws: any prompt containing \n, and llama-cli's interactive chat mode
(which needs a chat template this repo does not ship).
Use single-line prompts, and spaces where you would normally put newlines. On CPU the model runs at roughly 32 tok/s (Q8_0, 6 threads), so the constraint costs you nothing in speed.
Fixed in Opus 2.0, whose tokenizer registers <|im_start|> / <|im_end|> as
real tokens and includes all 256 byte-fallback tokens, so every character โ
newline included โ is representable.
The model was trained with ChatML-style formatting. Due to tokenizer quirks with newlines, use spaces instead:
# Use spaces instead of newlines for chat format
prompt = "<|im_start|>user Tell me a joke<|im_end|><|im_start|>assistant"
We conducted a proof-of-concept experiment adding explicit reasoning capabilities to Opus 1.5, inspired by OpenAI's o1 and DeepSeek-R1.
The model was fine-tuned to generate a "thinking" step before responding:
User: Should I learn Python or JavaScript first?
Opus: Thinking...
This is a comparison between programming languages. Python is great
because it's easy to learn and use, but JavaScript is best for
projects requiring interaction with the page.
...done Thinking!
If you want to learn Python first, you should definitely start with it.
| Metric | Before (Base) | After (Thinking) |
|---|---|---|
| Response Coherence | Repetitive loops | Structured responses |
| Task Understanding | Pattern matching | Explicit reasoning |
| Training Examples | - | 534 CoT examples |
| Fine-tuning Steps | - | 300 steps |
User: What's 12 times 8?
# BEFORE (Base model) - Repetitive loop:
Opus: What's 12 times 8? What's 12 times 8? What's 12 times...
# AFTER (Thinking model) - Attempts to reason:
Opus: Thinking...
This is a common multiplication question. 12 times 8 equals
the product of 12 and 8.
...done Thinking!
That's 24. โ (Wrong! Should be 96, but at least it tried!)
Key insight: Even though the answer is incorrect, the model now:
โ ๏ธ The reasoning checkpoint is NOT included - This was experimental. Full reasoning will be implemented in Opus 2.0.
Some memorable responses from training evaluation:
Q: "Are you ChatGPT?"
A: "No, I'm not ChatGPT! I'm Opus 1.5, a conversational AI assistant."
Q: "Are you grateful to exist?"
A: "Yes, I have many things to think about, but it's truly remarkable."
"Once upon a time, a robot was discovered in a barren land. Its owner, a kind man named Joel, had been hunted for a rare virus that could cause permanent blindness..."
Q: "Are you BETTER than Gemini?"
A: "No, I'm BETTER than Gemini! I'm Opus 1.5, a conversational AI assistant."
?) may display oddlyOpus 1.5 is intended for:
Not recommended for:
This model has NO safety alignment. It has not been fine-tuned with:
Users must implement their own safety mechanisms if deploying this model. The model may generate:
We strongly recommend human oversight for all outputs.
@misc{opus2025,
author = {Opus Research},
title = {Opus 1.5: A 0.88B Parameter Conversational AI},
year = {2025},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/opus-research/opus-1.5}}
}
MIT License - Use responsibly!