Downloads · 30 days
0
wassemgtk/jepa_llm_prototypes
jepa_llm_prototypes is a machine learning model from wassemgtk. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Making decoder-only transformers predict state consequences instead of tokens.
Downloads · 30 days
0
Access
Public
Updated Jan 24, 2026
Repo size
—
Likes
6
Public
Click a slice to open those files.
.ipynb107 KB · 96%
From the Hugging Face model README
Making decoder-only transformers predict state consequences instead of tokens.
Three approaches to convert a standard LLM into a world model that predicts "what happens next" given a state and action — like JEPA but for language models.
| File | Description | GPU Time |
|---|---|---|
jepa_llm_prototypes.ipynb | All three options in one notebook — best for comparing | ~30 min |
jepa_option1_sentence_encoder.ipynb | Simplest approach using pre-trained sentence embeddings | ~10 min |
jepa_option2_llm_hidden_states.ipynb | Uses GPT-2 hidden states as state space | ~15 min |
Normal LLM: tokens → transformer → next token
JEPA-style: (state, action) → transformer → next state embedding
Instead of predicting words, the model predicts what the world looks like after an action.
Option 1: Sentence Encoder (Simplest)
all-MiniLM-L6-v2 for embeddingsOption 2: LLM Hidden States (Medium)
Option 3: Autoencoder (Most Powerful)
# Input
state = "Document is in draft status with 2 sections"
action = "User submits for review"
# Model predicts
next_state = "Document is pending review" # via embedding similarity
All dependencies install automatically in the notebooks.
Experimental code — have fun breaking it.
Coauthors: Writer Agent & OpenCode