Downloads · 30 days
0
Red0ne/TexITex-P4A
TexITex-P4A is a text generation model from Red0ne. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Can we generate entire sentences in parallel by treating token embeddings as a 2D image?
Downloads · 30 days
0
Access
Public
Updated Apr 17, 2026
Repo size
844 KB
Likes
0
Public
Click a slice to open those files.
.png842 KB · 82%
From the Hugging Face model README
Can we generate entire sentences in parallel by treating token embeddings as a 2D image?
TexITex (Token-Image-Token) is a research proof-of-concept that encodes token embeddings as 2D latent images and generates them all at once using image diffusion — no autoregressive decoding step by step.
📄 Read the full paper (PDF)
💻 GitHub — code + experiments
Text → token embeddings → VQ-GAN encode → (16,16,16) latent image
↓
DiT diffusion (200 DDIM steps)
↓
Text ← nearest-neighbour lookup ← VQ-GAN decode ← generated latent
64 tokens are arranged in a 16×16 grid of 2×2 patches. The VQ-GAN compresses each patch to a 16-channel latent. The DiT generates the full latent image in a fixed 200 steps regardless of sequence length.


| Metric | Value |
|---|---|
| VQ-GAN roundtrip accuracy | 89.8% |
| Composite score — best sample | 0.372 |
| Composite score — mean (n=64) | 0.104 |
| Bigram coherence — best sample | 0.831 |
| Real-word ratio — mean | 0.683 |
| Median perplexity | 197 |

Best sample (composite = 0.372, bigram = 0.831):
"a simulated adversary engagement. Your objectives include testing detection capabilities, exercising incident response, identifying security gaps. You employ realistic adversary TTPs mapped to MITRE ATT&CK, maintain operational security, and adapt your approach based on blue team responses."

| Channels | Role |
|---|---|
| ch 0 — position | 0→1 gradient in reading order |
| ch 1 — boundary | 1.0 at 2×2 patch edges, prevents token bleed |
| ch 2–17 — self-cond | Previous DDIM step's x0 prediction (iterative refinement) |
| ch 18–33 — noisy latent | Current x_t from forward diffusion |
| Component | Parameters | Role |
|---|---|---|
| VQ-GAN (tokence_big_long) | 17.6M | Encode/decode token embeddings ↔ latent image |
| DiT (depth=12, dim=512, heads=8) | 57.8M | Denoise the latent image |
| LSTM SequencePredictor | 239.7K | Sequence-order auxiliary loss (weight=0.5) |
| Total | 58.0M |

@misc{cj2026texitex,
title = {TexITex: Parallel Text Generation via Token Embedding Diffusion in 2D Image Space},
author = {Jean Paul, C J},
year = {2026},
url = {https://github.com/PurpleS3Cf0X/TexITex}
}
Author: Jean Paul C J (Unaffiliated)