Downloads · 30 days
9
11% of all-time downloads
SlayerLab/lojban-gpt2-rsi1
lojban-gpt2-rsi1 is a text generation model from SlayerLab. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as other.
Moc obliczeniowa dzięki uprzejmości Comtegra GPU Cloud Strona główna | Comtegra GPU Cloud Chmura GPU | Zasoby obliczeniowe GPU | GPU na żądanie. Comtegra GPU Cloud to sposób na interakcję z zasobami GPU. Korzystaj z m…
Downloads · 30 days
9
11% of all-time downloads
All-time downloads
79
Public
Repo size
366 MB
Likes
0
Public
Click a slice to open those files.
.pt366 MB · 100%
From the Hugging Face model README
Moc obliczeniowa dzięki uprzejmości Comtegra GPU Cloud Strona główna | Comtegra GPU Cloud Chmura GPU | Zasoby obliczeniowe GPU | GPU na żądanie. Comtegra GPU Cloud to sposób na interakcję z zasobami GPU. Korzystaj z mocy klastra bez konieczności znajomości Kubernetesa

Rights review in progress. The earlier
MITlabel was broader than the currently documented source-level rights for the training corpus. It has been replaced withotherpending a complete source-by-source rights map. No blanket licence to the model weights or training material is granted by this card. Do not redistribute, fine-tune or use the weights commercially until the relevant permission scope is published or separately agreed. Contact: [email protected].
A 97.7M parameter GPT-2 model trained on the Lojban constructed language using Recursive Self-Improvement (RSI).
| Parameter | Value |
|---|---|
| Architecture | GPT-2 (decoder-only transformer) |
| Parameters | 97.7M |
| Layers | 12 |
| Hidden dim | 768 |
| Attention heads | 12 |
| Vocab size | 8,000 (custom BPE) |
| Context length | 512 tokens |
| Training iterations | 50,000 |
| Training tokens | 181M |
| Real data ratio | 33% |
| Best validation loss | 2.5805 |
| camxes grammaticality | 76.6% |
This model was trained using a Recursive Self-Improvement loop:
| Model | Val Loss | camxes % |
|---|---|---|
| V3 (5M synthetic) | ~4.5 | 0.3% |
| V4 (Metaspace tokenizer) | 3.80 | 26.7% |
| V5 (GPT-2 small) | 3.03 | 87.4% |
| V8 (99.3% synthetic) | 3.66 | — |
| RSI-1 (this model) | 2.58 | 76.6% |
500 generated sentences evaluated with the camxes PEG parser:
Grammatical (PASS):
✓ coi rodo mi'e la bripre
✓ dei notci fo do mu'i le nu do satci casnu dei kei fe levi tarci poi lei cevni coi
✓ tu cu nimre lu ckule sisti li'u ko di'a vofli le'e ganse doi cnt
✓ lo prenu cu tavla bau lo ponjo
✓ la .alis. cu viska lo cinfo
import torch
from tokenizers import Tokenizer
# Load tokenizer
tok = Tokenizer.from_file("tokenizer.json")
# Load model
checkpoint = torch.load("model.pt", map_location="cpu")
model.load_state_dict(checkpoint)
# Generate
input_ids = torch.tensor([[tok.encode("<|bos|>mi klama lo zdani").ids]])
# ... standard GPT-2 autoregressive generation
Trained on the SlayerLab/lojban-master dataset — the largest publicly available Lojban corpus (542K sentences, 7.9M words from 15 sources).
RSI-1's camxes score (76.6%) is lower than V5 (87.4%) despite better val loss (2.58 vs 3.03). This is because:
The model was trained on a single RTX 3090 (24GB) over ~22 hours.
The licence scope is under review because the referenced training dataset combines multiple sources. Code may be used only under an express licence notice applicable to that code; this card does not apply MIT to the weights or source material.