Downloads · 30 days
7
9% of all-time downloads
Itaking/itakura-300m-cpt-model
itakura-300m-cpt-model is a text generation model from Itaking. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
OLMo-2 アーキテクチャを ~300M に縮小し、英語事前学習(Stage 1)の後に 英日バイリンガルデータで継続事前学習(アニーリング)したモデルです。
Downloads · 30 days
7
9% of all-time downloads
All-time downloads
77
Public
Parameters
457M
8.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors914 MB · 99%
From the Hugging Face model README
OLMo-2 アーキテクチャを ~300M に縮小し、英語事前学習(Stage 1)の後に 英日バイリンガルデータで継続事前学習(アニーリング)したモデルです。
| 項目 | 値 |
|---|---|
| Base config | allenai/OLMo-2-0425-1B(config のみ・重みは使用せず) |
| Parameters | ~300M |
| hidden_size | 1024 |
| num_hidden_layers | 16 |
| num_attention_heads | 16 |
| num_key_value_heads | 8 (GQA) |
| intermediate_size | 4096 |
| max_position_embeddings | 2048 |
| Tokenizer | allenai/OLMo-2-0425-1B |
| 項目 | 値 |
|---|---|
| Dataset | FineWeb (sample-10BT) |
| Tokens | ~1.5B |
| Learning rate | 3e-4 (cosine + min_lr_rate=0.1) |
| Batch size (effective) | 128 seq × 2048 tokens = 262K tokens/step |
| 項目 | 値 |
|---|---|
| Dataset | FineWeb-Edu 60% + Wikipedia JA 40% |
| Tokens | ~0.3B |
| Learning rate | 1e-4 (cosine + min_lr_rate=0.1) |
| Sequence packing | あり(padding ゼロ) |
| Hardware | NVIDIA RTX 4090 24GB |
| Framework | HuggingFace Transformers + Trainer |
| データセット | ライセンス |
|---|---|
| FineWeb-Edu (HuggingFaceFW) | ODC-By |
| Wikipedia JA (Wikimedia) | CC BY-SA 4.0 |
Apache 2.0