Downloads · 30 days
0
MIHUJIOUY/V19-cognitive-engine
V19-cognitive-engine is a machine learning model from MIHUJIOUY. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
V19 is a fully interpretable Chinese language understanding system with 4.7 million parameters. It reads a Chinese sentence as a sequence of characters, builds word-level representations through a frozen char-to-word…
Downloads · 30 days
0
Access
Public
Updated Jun 16, 2026
Repo size
130 MB
Likes
0
Public
Click a slice to open those files.
.pt130 MB · 100%
From the Hugging Face model README
V19 is a fully interpretable Chinese language understanding system with 4.7 million parameters. It reads a Chinese sentence as a sequence of characters, builds word-level representations through a frozen char-to-word encoder (P1), routes information across sentences (P7), and decodes back to word sequences (P6) — all while maintaining traceable, auditable intermediate states.
Unlike transformer-based LLMs, every internal decision in V19 can be inspected:
P1 (Char→Word, 96K frozen) → P7 (Router, 226K) → Explore+Meta (Gate, 101K) → P6 (Decoder, 4.37M)
h * gate + pos_embed[i] → Linear(256,128) → word_i| Config | Value |
|---|---|
| Optimizer | Adam (P6 lr=0.003, P7 lr=0.0045, Gate lr=0.006) |
| Loss | 1.0 - mean(cosine_similarity(pred, true)) |
| Epochs | 1000 |
| Batch | Full dataset per epoch (41,909 pairs) |
| GPU | RTX 5070 |
| Memory | ~300MB (training), 141MB (inference) |
| Metric | Score |
|---|---|
| Word Accuracy | 92.4% |
| Exact Match | 76.3% |
| Rouge-L F1 | 93.2 |
| Per-word Cosine Mean | 0.96 |
| Inference | 14ms/sentence |
| Metric | Epoch 1 | Target |
|---|---|---|
| Word Accuracy | 43.5% | >95% |
| Per-word Cosine | 0.73 | >0.97 |
After 5 failed approaches to prevent the P6 decoder from outputting the same word repeatedly (rep_pen, residual extraction, weight transpose inversion, orthogonal init, cos_loss margin), the final solution was the simplest:
for i in range(max_words):
hi = h + self.pos_embed[i] # unique starting point per head
w = self.extract[i](hi)
No rep_pen. No residuals. No detach. Just position diversity.
Instead of directly minimizing loss (which causes gates to converge to all-open or all-closed), the gate is trained indirectly:
This prevents the "gate symmetry lock" (all dims identical, std=0.0001) that plagued early versions.
| Source | Pairs | License |
|---|---|---|
| shibing624/chinese_text_correction | 53,298 | Apache 2.0 |
| MuCGEC | 1,038 | CC BY 4.0 |
After sentence splitting: 52,387 pairs
Split: train 41,909 (80%) / test 5,238 (10%) / exam 5,240 (10%)
@misc{wei2026v19,
title={V19: A White-box Chinese Cognition Engine},
author={Wei, Jinqi},
year={2026},
howpublished={\url{https://github.com/Xuan-yi-yan/V18-cognitive-architecture}},
}
GitHub: @Xuan-yi-yan