Downloads · 30 days
18
25% of all-time downloads
davidandroid/lattice-multiplication-transformer
lattice-multiplication-transformer is a machine learning model from davidandroid. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
Downloads · 30 days
18
25% of all-time downloads
All-time downloads
73
Public
Repo size
143 MB
Likes
0
Public
Click a slice to open those files.
.pt143 MB · 100%
From the Hugging Face model README
Interactive Demo · Source Code
This repository contains the final checkpoint for an approximately 11.7-million-parameter model studying latent algorithm learning through an algorithm-aligned input representation.
Multi-digit multiplication is used as a case study. A learned frontend transforms two operands into an N × N lattice of digit-pair representations, which is passed to a standard Transformer encoder-decoder. The model directly generates the final product without scratchpads, carry labels, local-product labels, diagonal sums, or other intermediate supervision.
The final checkpoint was trained for 600,000 optimization steps.
final_lattice_multiplication__11m.pt
lattice_multiplication_transformer.py
tokenizer.py
config.json
The checkpoint contains the model state, optimizer state, curriculum state, training history, and random-number-generator states. Inference requires only checkpoint["model_state_dict"].
import json
import sys
from pathlib import Path
import torch
from huggingface_hub import snapshot_download
model_directory = Path(
snapshot_download(
repo_id="davidandroid/lattice-multiplication-transformer"
)
)
sys.path.insert(0, str(model_directory))
from lattice_multiplication_transformer import (
LatticeMultiplicationLoopTransformer,
)
with open(
model_directory / "config.json",
encoding="utf-8",
) as config_file:
model_config = json.load(config_file)
model = LatticeMultiplicationLoopTransformer(
**model_config
)
checkpoint = torch.load(
model_directory / "final_lattice_multiplication__11m.pt",
map_location="cpu",
weights_only=False,
)
model.load_state_dict(
checkpoint["model_state_dict"]
)
model.eval()
The model was trained on operands containing at most 20 digits. It does not achieve reliable exact-match length generalization beyond that range, even though token accuracy remains higher immediately beyond the training boundary.
This is a research model for studying latent algorithm learning, not a replacement for deterministic multiplication software.