Downloads · 30 days
0
thedoucheperfect/ANLP_A1_Assignment_C1-C5
ANLP_A1_Assignment_C1-C5 is a machine learning model from thedoucheperfect. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
<p align="center" <bAdvanced NLP — Assignment 1</b<br Transformers from Scratch · RoPE · GQA · RMSNorm · BLT </p
Downloads · 30 days
0
Access
Public
Updated Sep 3, 2026
Repo size
503 MB
Likes
1
Public
Click a slice to open those files.
.pt503 MB · 99%
From the Hugging Face model README
Shourya Pillai · MS Research, CSE · IIIT Hyderabad
Roll No. 2026701047 · [email protected]
This project explores Transformer architectures for encrypted-text reconstruction, comparing standard subword Transformers against a Byte Latent Transformer (BLT) operating directly on raw bytes.
All Transformer components were implemented from scratch in PyTorch,
without using nn.Transformer or nn.MultiheadAttention.
The experiment consists of five controlled configurations:
| Model | Main Change | Representation |
|---|---|---|
| C1 | Baseline Transformer | Custom BPE |
| C2 | RoPE | Custom BPE |
| C3 | GQA | Custom BPE |
| C4 | RMSNorm | Custom BPE |
| C5 | Byte Latent Transformer | Raw Bytes |
| Model | Bit Accuracy ↑ | Sequence Accuracy ↑ | Levenshtein ↓ | BLEU ↑ |
|---|---|---|---|---|
| C1 | 69.67% | 1.14% | 36.57 | 0.361 |
| C2 | 74.48% | 3.54% | 21.38 | 0.549 |
| C3 | 70.63% | 1.40% | 35.91 | 0.367 |
| C4 | 69.44% | 1.40% | 39.05 | 0.337 |
| C5 — BLT | 99.99% | 98.60% | 0.036 | — |
C5 achieves 99.988% bit-level accuracy and 98.599% exact sequence accuracy.
BLEU and ROUGE are not applicable to C5 because it operates directly on bytes rather than subword tokens.
Sinusoidal absolute positional encoding + Multi-Head Attention + LayerNorm + custom BPE.
Replaces absolute positional encoding with Rotary Positional Encoding.
→ Best-performing subword Transformer.
Replaces standard MHA with Grouped-Query Attention, reducing the number of key/value heads while retaining multiple query heads.
Replaces LayerNorm with RMSNorm while keeping the remaining architecture unchanged.
Moves from subword tokens to raw bytes, using a Byte Latent Transformer architecture with local byte processing and global Transformer representations.
→ Near-perfect reconstruction.




1. RoPE improves the standard subword Transformer.
C2 substantially outperforms the baseline C1 across bit accuracy, sequence accuracy, Levenshtein distance, BLEU, and ROUGE.
2. GQA is competitive with the baseline but does not improve reconstruction quality in this experiment.
3. RMSNorm does not outperform LayerNorm for this task.
4. Byte-level modeling is extremely effective for this reconstruction problem.
C5 dramatically outperforms all subword-based configurations, achieving near-perfect reconstruction.
C1/ C1 checkpoint
C2/ C2 checkpoint
C3/ C3 checkpoint
C4/ C4 checkpoint
C5/ C5 BLT checkpoint
tokenizers/
├── tokenizer.json
└── cipher_tokenizer.json
metrics/
└── evaluation_results.csv
plots/
├── C1_C5_Train_Loss.png
├── C1_C5_Validation_Loss.png
├── C1_C5_Learning_Rate.png
└── Peak_GPU_Memory.png