Downloads · 30 days
23
21% of all-time downloads
mossez-systems/Mossez-100M-Instruct
Mossez-100M-Instruct is a text generation model from mossez-systems. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Mossez-100M-Instruct is an experimental 100M-parameter Russian–English instruction-tuning research model derived from the final CPT-trained Mossez-100M-Base.
Downloads · 30 days
23
21% of all-time downloads
All-time downloads
111
Public
Parameters
100M
400 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors400 MB · 99%
From the Hugging Face model README
Mossez-100M-Instruct is an experimental 100M-parameter Russian–English instruction-tuning research model derived from the final CPT-trained Mossez-100M-Base.
This is a compact research artifact. It is not a strong, reliable, safe, or production-ready assistant. Its purpose is to study the behavior and post-training limits of the Mossez-100M family; stronger practical behavior is expected from later, larger families.
| Property | Value |
|---|---|
| Parameters | 100,098,048 |
| Architecture | Llama-compatible decoder-only Transformer |
| Layers | 12 |
| Hidden size | 768 |
| Query / KV heads | 12 / 4 |
| Intermediate size | 2,048 |
| Context length | 1,024 tokens |
| Vocabulary | 32,007 |
| Tokenizer | Mossez Tokenizer v1.1, byte-level BPE |
| Embeddings | Tied input/output embeddings |
| Primary languages | Russian, English |
| Weight format | Safetensors, clean FP32 master |
Tokenizer v1.1 preserves every original ID from 0 through 31,999 and adds
seven role/FIM tokens at IDs 32,000–32,006. Its tokenizer.json SHA-256 is
e9551d84b9947f741763bf815a2d5f6bfcc47a3b67c73fcbf386223e8ed969be.
The model is intended for compact-model research, local experiments, instruction-tuning analysis, reproducibility work, and testing inference pipelines. It is not intended for factual authority, autonomous action, safety-critical use, cybersecurity advice, or high-stakes decisions.
Install recent versions of torch, transformers, safetensors, and
jinja2. The repository contains a clean FP32 master; Transformers can load it
directly at FP16 runtime dtype on CUDA.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "mossez-systems/Mossez-100M-Instruct"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if device == "cuda" else torch.float32
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=dtype)
model.to(device).eval()
messages = [{"role": "user", "content": "Кратко объясни, что такое RoPE."}]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(prompt, return_tensors="pt").to(device)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=96,
do_sample=False,
eos_token_id=[tokenizer.eos_token_id, 32003],
pad_token_id=tokenizer.pad_token_id,
)
new_tokens = output[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))
Keep the total prompt plus generated length within 1,024 tokens. The model may
fail to emit <|end|> or EOS, so always set a finite max_new_tokens limit.
The selected weights follow this exact path:
The selected lineage exposed the model to 9,973,153 non-padding tokens and
6,030,590 supervised assistant tokens after resize. v4 used FP16 autocast,
fused AdamW, gradient checkpointing, micro-batch 8, accumulation 2, learning
rate 1e-5, and seed 3407. System/user text, role prefixes, BOS, and padding
were excluded from loss; assistant content, assistant termination, and final
EOS were supervised.
OASST1 v1/v2 and first-turn v5 were completed comparison pilots but are not in
the selected weight lineage. See TRAINING_REPORT.md and EVALUATION.md.
Checkpoint selection combined teacher-forced validation/test loss with a fixed 31-case RU/EN diagnostic suite and manual review. v4 step 216 was selected by relative performance among the bounded v1–v5 experiments, not because it met a product-quality threshold.
| Metric | v4 step 216 |
|---|---|
| v4 validation loss | 1.434212 |
| v4 test loss | 1.551534 |
| Automatic diagnostic passes | 2 / 31 |
| Structural completions | 17 / 31 |
| Unterminated outputs | 13 / 31 |
| Repetition failures | 5 / 31 |
| Generated role-marker leakage | 0 / 31 |
The clean FP32 export was independently reloaded and verified with 100,098,048 parameters, tied embeddings, finite weights, and finite logits in a real CUDA forward pass. A separate FP16 validation export also passed, but is not duplicated in this repository.
smol-contraints subset later found to contain some prompt-to-
answer leakage. It was removed from v4, but its influence may remain in the
selected weights.Do not use this model for medical, legal, financial, security, or other high-stakes decisions.
Selected-lineage SFT data was derived from pinned, filtered subsets of
OpenAssistant OASST2 and HuggingFaceTB SmolTalk. The Base model retains its own
FineWeb/FineWeb2/Wikipedia training history. Source datasets are not
redistributed here. See DATASET_ATTRIBUTION.md and NOTICE.md for pinned
revisions, licenses, and caveats.
Model weights and original repository materials are released under Apache-2.0.
Third-party datasets and source content remain under their respective licenses
and terms. See LICENSE and NOTICE.md.