Downloads · 30 days
77
41% of all-time downloads
foss22/Hallucinate1900-95M
Hallucinate1900-95M is a machine learning model from foss22. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
hallucinate, to deceiue, or blind https://extra.shu.ac.uk/emls/iemls/work/etexts/caw1604wremoved.htm
Downloads · 30 days
77
41% of all-time downloads
All-time downloads
188
Public
Repo size
930 MB
Likes
0
Public
Click a slice to open those files.
.safetensors383 MB · 54%
From the Hugging Face model README
hallucinate, to deceiue, or blind https://extra.shu.ac.uk/emls/iemls/work/etexts/caw1604w_removed.htm

The main focus of the experiment is a steming of the corpus jsonl + 63355 vocab extended from 32752 (with those stemma).
Not intended for inference before additional vocab extentions and then slower CPT on real texts without stemming.
python auto_emoji_plot.py <(cat <(head -1 training_log.csv) <(tail -35 training_log.csv))
LOSS (EMA: ▓, Raw: ·, emojis: state):
│l12.6284
│o
│🔥
│s
│
│ 11▓2124
│
│ 📉
│ ▓
│
│ 9.7964
│ · ▓
│
│ 📉
│ ▓
│ 8.3805
│ · ▓ 📉
│ ▓
│ · ▓ 📉
│ · ▓
│ 6.9645 · ▓ 📉
│ · · ▓ ▓ 📉
│ · · · ▓ ▓ • 📉
│ · · · ▓ ▓ ▓ ▓ 📉 📉 ↘ 📉 📈 📈 📈 ↗ 📉
│ · · · ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓
│
│ 5.5485
└00────────────────────440─────────────────────780───────────────────1120────────────────────1460──────────────────step─
Legend: ▓ = EMA trend | · = Raw data | 🔥 = Start | ↗/📈 = Increase | ↘/📉 = Decrease | • = Flat
Summary: 🟡 Plateau or noisy (flat trend). Trend Scale: 12.0384 → 6.1669 (Δ 5.8715)
LEARNING_RATE (EMA: ▓, Raw: ·, emojis: state):
│l1.06e-04
│e
│🐎 🏃 🏃 🏃
│r ▓ ▓ ▓ ▓ ▓ ▓ 🏃
│n · ▓ ▓ ▓ 🏃 🏃
│i9.11e-05 · ▓ ▓ ▓ 🏃
│n · · ▓ ▓
│g · ▓ 🏃
│_ · · ▓ ▓ 🏃
│r · ▓ 🏃
│a7.57e-05 · · ▓ ▓
│t · ▓ 🏃
│e · ▓ 🏃
│ · ▓ ▓
│ · ▓ 🏃
│ 6.04e-05 · · ▓
│ · ▓ 🏃
│ · ▓
│ · ▓ 🏃
│ · ▓ 🏃
│ 4.50e-05 · ▓ ▓
│ · ▓ 🏃
│ · ▓
│ · ·
│ ·
│
│ 2.97e-05
└00────────────────────440─────────────────────780───────────────────1120────────────────────1460──────────────────step─
Legend: ▓ = EMA trend | · = Raw data | 🐢 = <1e-8 | 🚶 = 1e-8~1e-6 | 🏃 = 1e-6~1e-4 | 🐎 = 1e-4~1e-3 | 🐉 = >1e-3
Summary: 🟢 Strong learning phase. Current: 3.61e-05
GRAD_NORM (EMA: ▓, Raw: ·, emojis: state):
│g66.69
│r
│🚨
│d
│_
│n52.43🚨
│o ▓ ▓
│r ·
│m ▓
│
│ 38.16
│
│ · 🚨
│ ▓
│ ·
│ 23.89
│ ▓
│ ✅
│ ▓
│
│ 9.62 ▓ ⚠
│ ▓
│ ▓ ✅ ✅
│ · · ▓ ▓ ▓ ⚠ ⚠ ⚠ ⚠ ⚠ ⚠ ⚠ ⚠ ⚠ · ⚠ · 🚨
│ · · · · · · · ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓
│
│ -4.64
└00────────────────────440─────────────────────780───────────────────1120────────────────────1460──────────────────step─
Legend: ▓ = EMA trend | · = Raw data | 🕳 = <0.3 (vanishing) | ✅ = 0.3~2.0 (stable) | ⚠ = 2.0~4.0 (high) | 🚨 = >4.0 (explosion)
Summary: 🔴 ALERT: Gradient explosion risk! Current: 4.03
ENGLISH_PPL (EMA: ▓, Raw: ·, emojis: state):
│e669.7
│n
│g 🌫 🌫 🌫
│l 🌫 🌫 · 🌫 ▓ ▓ ▓ ▓ ▓
│i · 🌫 · 🌫 · ▓ ▓ ▓ ▓ ▓ ▓ · ·
│s655.1 · ▓ ▓ ▓ ▓
│h · · · 🌫 🌫 ▓
│_ · 🌫 🌫 · ▓ ▓ ▓
│p · · 🌫 ▓ ▓ ▓ ▓ ▓
│p 🌫 · 🌫 · ▓
│l640.6 ▓ ▓ ▓ ▓
│ · ▓
│ 🌫
│ ▓
│
│ 626.1 ▓
│ ·
│
│ 🌫
│ ▓
│ 611.6
│
│ ▓
│🌫
│▓
│
│ 597.0
└00────────────────────440─────────────────────780───────────────────1120────────────────────1460──────────────────step─
Legend: ▓ = EMA trend | · = Raw data | Emojis show uncertainty state (🌫 high, ☁ med, ☀ low)
Summary: 🔵 Stable. Current: 659.3, range: 603.1–662.2 (Δ 59.2)
RUSSIAN_PPL (EMA: ▓, Raw: ·, emojis: state):
│r103775.3
│u
│🌫
│s
│i
│a85153.8
│n ▓
│_
│p 🌫
│p ▓
│l66532.3
│
│
│ ▓
│ ·
│ 47910.7 🌫
│ ▓
│
│ · ▓
│ 🌫
│ 29289.2 · ▓
│ ▓ 🌫
│ ▓ 🌫
│ · ▓ ▓ ▓ 🌫 🌫 🌫 🌫 🌫 🌫 🌫 🌫 🌫 🌫 🌫 🌫
│ · · · · · · · ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓
│
│ 10667.7
└00────────────────────440─────────────────────780───────────────────1120────────────────────1460──────────────────step─
Legend: ▓ = EMA trend | · = Raw data | Emojis show uncertainty state (🌫 high, ☁ med, ☀ low)
Summary: 🔵 Stable. Current: 18631.2, range: 18572.9–96016.4 (Δ 77443.5)
PPL_DELTA_EN (EMA: ▓, Raw: ·, emojis: state):
│p65.3
│p
│l 🌫 🌫 🌫
│_ 🌫 🌫 · 🌫 ▓ ▓ ▓ ▓ ▓
│d · 🌫 · 🌫 · ▓ ▓ ▓ ▓ ▓ ▓ · ·
│e50.8 · ▓ ▓ ▓ ▓
│l · · · 🌫 🌫 ▓
│t · 🌫 🌫 · ▓ ▓ ▓
│a · · 🌫 ▓ ▓ ▓ ▓ ▓
│_ 🌫 · 🌫 · ▓
│e36.3 ▓ ▓ ▓ ▓
│n · ▓
│ 🌫
│ ▓
│
│ 21.7 ▓
│ ·
│
│ 🌫
│ ▓
│ 7.2
│
│ ▓
│☀
│▓
│
│ -7.3
└00────────────────────440─────────────────────780───────────────────1120────────────────────1460──────────────────step─
Legend: ▓ = EMA trend | · = Raw data | Emojis show uncertainty state (🌫 high, ☁ med, ☀ low)
Summary: 🔵 Stable. Current: 55.0, range: -1.3–57.9 (Δ 59.2)
PPL_DELTA_RU (EMA: ▓, Raw: ·, emojis: state):
│p-52580.3
│p
│☀
│_
│d
│e-71201.9
│l ▓
│t
│a 🌫
│_ ▓
│r-89823.4
│u
│
│ ▓
│ ·
│ -108444.9 🌫
│ ▓
│
│ · ▓
│ 🌫
│ -127066.5· ▓
│ ▓ 🌫
│ ▓ 🌫
│ · ▓ ▓ ▓ 🌫 🌫 🌫 🌫 🌫 🌫 🌫 🌫 🌫 🌫 🌫 🌫
│ · · · · · · · ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓ ▓
│
│ -145688.0
└00────────────────────440─────────────────────780───────────────────1120────────────────────1460──────────────────step─
Legend: ▓ = EMA trend | · = Raw data | Emojis show uncertainty state (🌫 high, ☁ med, ☀ low)
Summary: 📉 Decreasing. Current: -137724.5, range: -137782.8–-60339.3 (Δ 77443.5)
Root folder has BF16 run chechpkoint-1500 visualized above.
Continued F16 in a separate checkpoint-1000 (newer) run with a newer script from checkpoint-1000 to checkpoint-1000 (yes, same names... to confuse readers)):
python continual_train_v2.py --model_path training_v5/checkpoint-1000/ --train_file /dev/shm/sed-f_98_3.huniq-c.list_4sed_mystem-n.rgА-Яа-я.jsonl --output_dir training_v5 --freeze_layers 4 --max_steps 3000 --batch_size 16 --gradient_accumulation 4 --learning_rate 1e-4 --warmup_steps 50 --block_size 512 --eval_every 50 --logging_steps 10 --save_steps 500 --auto_resume --use_fp16 --gradient_checkpointing --old_vocab_size 3275
...
Step 960 | loss=5.1411 | lr=7.83e-05 | grad_norm=4.734 | tok/s=2,595
Step 970 | loss=5.1374 | lr=7.79e-05 | grad_norm=5.862 | tok/s=2,595
Step 980 | loss=5.1325 | lr=7.74e-05 | grad_norm=3.418 | tok/s=2,595
Step 990 | loss=5.1275 | lr=7.70e-05 | grad_norm=2.287 | tok/s=2,596
Step 1000 | loss=5.1224 | lr=7.65e-05 | grad_norm=3.189 | tok/s=2,596
================================================================================
📊 EVALUATION | Step 1000 | Epoch 1.0000
================================================================================
Loss: 5.1224
LR: 0.00e+00
Grad norm: 0.00
Elapsed: 3.51 ч
English PPL: 1276.14 (+616.76)
Russian PPL: 2665.49 (-16108.01)
Генерации:
english:
«The capital of France is»
→ divided into two parts : namely, that it were better for thee, my dear...
russian:
«Бог есть»
→ ИИИБог мы Бог он С�Бог ИИИЛИУИИИИИИИИњ...
LOSS (EMA: ▓, Raw: ·, emojis: state):
│l6.2119
│o
│🔥 📉
│s ▓
│ 📉
│ 5.974· ▓ 📉
│ ▓
│ 📉
│ · ▓ 📉
│ ▓
│ 5.7365 · 📉
│ · ▓ 📉
│ ▓ 📉
│ · ▓ 📉
│ · ▓
│ 5.4988 · 📉
│ · ▓ 📉 📉
│ · ▓ ▓ 📉
│ · ▓ 📉
│ · ▓ 📉
│ 5.2611 · ▓ 📉
│ · · ▓ 📉 📉
│ · ▓ ▓ 📉
│ · · ▓
│ · ·
│
│ 5.0233
└0─────────────────────240─────────────────────430────────────────────620─────────────────────810──────────────────step─
Legend: ▓ = EMA trend | · = Raw data | 🔥 = Start | ↗/📈 = Increase | ↘/📉 = Decrease | • = Flat
Summary: 🟢 Effective learning (downward trend). Trend Scale: 6.1129 → 5.1889 (Δ 0.9240)