Downloads · 30 days
0
meforce/CetinLM-1B-Base
CetinLM-1B-Base is a text generation model from meforce. Use it when you need the model to write or continue text.
<p align="center" <img src="https://raw.githubusercontent.com/xertxetin/CetinLM/refs/heads/main/docs/cetinlm-logo-lq.png" alt="CetinLM" width="250" </p
Downloads · 30 days
0
Access
Public
Updated Sep 27, 2026
Repo size
—
Likes
1
Public
Click a slice to open those files.
.md5.7 KB · 79%
From the Hugging Face model README
CetinLM Base-v1 is a ~1.18B-parameter decoder-only language model trained from scratch with the custom CetinTokenizer-v1 tokenizer.
This page describes the public Base-v1 research lineage. The model is still in the base-pretraining stage; instruction following, assistant behavior and later post-training capabilities are separate stages and should not be inferred from pretraining loss alone.
<table> <tr> <td width="33%" valign="top"><strong>Model</strong><br><br>~1.18B decoder-only parameters.</td> <td width="33%" valign="top"><strong>Tokenizer</strong><br><br>CetinTokenizer-v1 · 48K vocabulary.</td> <td width="33%" valign="top"><strong>Context</strong><br><br>2K active Base-v1 training context.</td> </tr> </table>The frozen Base-v1 corpus contains approximately 11.39B unique training tokens. Processed-token count measures training exposure and is not the same thing as unique corpus size.
CetinLM is actively training, which means the latest token count, validation loss, perplexity, per-dataset measurements and sampled generation-health results continue to change.
To avoid leaving stale numbers scattered across model cards and repositories, the canonical current metrics are maintained here:
Historical measurements below remain intentionally preserved as a public reference snapshot.
| Signal | Archived value |
|---|---|
| Processed tokens | 5.15B |
| Validation loss | 2.496678 |
| Perplexity | 12.142 |
| Checkpoint status | NEW BEST at snapshot |
EOS P(EOS@end) | 0.3994 |
| EOS top-1 / top-5 | 45.0% / 75.0% |
| Median EOS rank | 2.0 |
This is an archived reference point, not the current training state. For the current checkpoint, use the live Research page above.
| Milestone | Val loss | PPL |
|---|---|---|
| 4.80B | 2.513649 | 12.350 |
| 4.85B | 2.504830 | 12.241 |
| 4.90B | 2.502269 | 12.210 |
| 4.95B | 2.501286 | 12.198 |
| 5.00B | 2.500880 | 12.193 |
| 5.05B | 2.503461 | 12.225 |
| 5.10B | 2.502036 | 12.207 |
| 5.15B | 2.496678 | 12.142 |
Across the archived 4.80B → 5.15B window, held-out validation loss moved from 2.513649 → 2.496678.
At 5.00B, a separate 1,000-generation web-balanced-v1 evaluation recorded:
The 25-case RAW GREEDY STRESS PROBE is intentionally harsher and is not presented as a user-facing loop rate.
For newer sampled-behavior runs, see CetinLM Research.
Base-v1 is the pretrained foundation, not the finished assistant.
The current model card should therefore be read as evidence about:
It should not be read as a final claim about instruction following, reasoning, coding, factuality, safety, tool use or assistant quality. Those belong to later post-training and dedicated evaluations.
CetinLM Live is a separate interactive runtime layer developed around the model stack. Public work includes realtime speech, synchronized text/audio interaction, interruption handling, local voice processing and persistent interaction research.
Live/runtime development does not alter Base-v1 checkpoint accounting or the frozen training lineage.
Optional Live speech uses separately licensed third-party software. Resemble AI Chatterbox Multilingual V3 is used for text-to-speech under its upstream MIT License. CetinLM does not claim ownership of Chatterbox.
Public third-party provenance and license notices are maintained with the CetinLM GitHub public research archive.