Downloads · 30 days
0
afoolishman/anlp-a1-assignment
anlp-a1-assignment is a machine learning model from afoolishman. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
This hub entry hosts checkpoints from a controlled ablation: map encrypted binary sequences (browncipher.txt) to English plaintext (brownplain.txt). Configuration C1 is the base model; C2–C5 each change exactly one co…
Downloads · 30 days
0
Access
Public
Updated Sep 3, 2026
Repo size
670 MB
Likes
0
Public
Click a slice to open those files.
.pt414 MB · 62%
From the Hugging Face model README
This hub entry hosts checkpoints from a controlled ablation: map encrypted binary sequences (brown_cipher.txt) to English plaintext (brown_plain.txt). Configuration C1 is the base model; C2–C5 each change exactly one component.
| Config | Single change vs C1 | Architecture class |
|---|---|---|
| C1 | — (base) | Seq2SeqTransformer |
| C2 | RoPE instead of sinusoidal PE | Seq2SeqTransformer |
| C3 | GQA (8 Q, 2 KV) instead of MHA | Seq2SeqTransformer |
| C4 | RMSNorm instead of LayerNorm | Seq2SeqTransformer |
| C5 | Entropy-patched BLT instead of BPE | BLTSeq2Seq |
| Hyperparameter | Value |
|---|---|
d_model | 256 |
| Heads | 8 (C3: 2 KV heads) |
| Encoder / decoder layers | 4 / 4 |
d_ff | 1024 |
| Dropout | 0.15 |
| Optimizer | AdamW, lr 1e-3, wd 0.01, betas (0.9, 0.98) |
| LR schedule | cosine + 400 warmup steps |
| Batch size | 64 |
| Chunk size | 256 characters (aligned cipher/plain) |
| Max epochs | 100 (C5 early-stopped later; best @ epoch 146) |
| Early stopping | patience 10 on val CE |
| Seed | 42 |
| Precision | AMP (fp16) |
Tokenization
cipher_bpe.json, plain_bpe.json.entropy_ngram.json, entropy_patcher.json.Checkpoints are raw PyTorch state dicts (pytorch_model.bin) plus config.json. They are not Hugging Face transformers.AutoModel checkpoints—load with the assignment code.
import json
import torch
from pathlib import Path
# after cloning / downloading this repo's `src/` onto PYTHONPATH
from models.transformer import Seq2SeqTransformer
from models.blt import BLTSeq2Seq
from tokenizer import BPETokenizer, PAD
hf_dir = Path("path/to/downloaded/hf") # folder with config.json + pytorch_model.bin
cfg = json.loads((hf_dir / "config.json").read_text())
if cfg.get("blt"):
model = BLTSeq2Seq(
d_model=cfg["d_model"],
n_heads=cfg["n_heads"],
n_kv_heads=cfg["n_kv_heads"],
n_encoder_layers=cfg["n_encoder_layers"],
n_decoder_layers=cfg["n_decoder_layers"],
d_ff=cfg["d_ff"],
dropout=cfg["dropout"],
patch_size=cfg["patch_size"],
n_local_layers=cfg["n_local_layers"],
n_local_heads=cfg["n_local_heads"],
pad_id=PAD,
)
else:
model = Seq2SeqTransformer(
src_vocab=cfg["src_vocab"],
tgt_vocab=cfg["tgt_vocab"],
d_model=cfg["d_model"],
n_heads=cfg["n_heads"],
n_kv_heads=cfg["n_kv_heads"],
n_encoder_layers=cfg["n_encoder_layers"],
n_decoder_layers=cfg["n_decoder_layers"],
d_ff=cfg["d_ff"],
dropout=cfg["dropout"],
pos_encoding=cfg["pos_encoding"],
attn_kind=cfg["attn_kind"],
norm_kind=cfg["norm_kind"],
pad_id=PAD,
)
state = torch.load(hf_dir / "pytorch_model.bin", map_location="cpu")
model.load_state_dict(state)
model.eval()
brown_cipher.txt / brown_plain.txt (5000 line pairs).len(cipher)/8 == len(plain) so packed cipher byte i aligns with plaintext character i.Greedy decoding only. Test metrics (1389 chunks):
| Config | Params | Best val CE | Epochs | Bit acc. (%) | Seq. acc. (%) | Lev. | BLEU |
|---|---|---|---|---|---|---|---|
| C1 | 8.15M | 0.0294 | 100 | 97.06 | 73.43 | 0.58 | 97.87 |
| C2 | 8.15M | 0.0188 | 82 | 99.01 | 85.89 | 0.81 | 98.75 |
| C3 | 6.97M | 0.0336 | 100 | 96.34 | 69.62 | 0.73 | 97.34 |
| C4 | 8.14M | 0.0283 | 94 | 97.15 | 73.79 | 0.54 | 97.86 |
| C5 | 11.70M | 0.0699 | 156 | 91.64 | 26.21 | 3.33 | n/a* |
*BLEU/ROUGE are secondary for token-free C5; greedy C5 BLEU ≈ 89.3 if computed for reference.
Ranking: C2 > C1 ≈ C4 > C3 > C5.
Note on C5: bit accuracy stays relatively high while sequence (exact-match) accuracy is lower because errors are mostly scattered typos on long chunks (mean Lev ≈ 3.3). Parallel patch decoding and entropy cuts on predicted bytes amplify that gap.
C1–C4
| File | Description |
|---|---|
pytorch_model.bin | state_dict of Seq2SeqTransformer |
config.json | Architecture + hparams |
cipher_bpe.json | Source BPE |
plain_bpe.json | Target BPE |
README.md | This model card |
C5
| File | Description |
|---|---|
pytorch_model.bin | state_dict of BLTSeq2Seq |
config.json | Architecture + hparams + entropy_threshold |
entropy_ngram.json | Smoothed byte n-gram entropy model |
entropy_patcher.json | Threshold / target patch size |
README.md | This model card |
transformers Auto* model; requires the course src/ code to run.