Downloads · 30 days
0
sawyerhpowell/megashtein
megashtein is a machine learning model from sawyerhpowell. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
In their paper "Deep Squared Euclidean Approximation to the Levenshtein Distance for DNA Storage", Guo et al. explore techniques for using a neural network to embed sequences in such a way that the squared Euclidean d…
Downloads · 30 days
0
Access
Public
Updated Jul 7, 2025
Repo size
553 KB
Likes
0
Public
Click a slice to open those files.
.pth553 KB · 96%
From the Hugging Face model README
In their paper "Deep Squared Euclidean Approximation to the Levenshtein Distance for DNA Storage", Guo et al. explore techniques for using a neural network to embed sequences in such a way that the squared Euclidean distance between embeddings approximates the Levenshtein distance between the original sequences. This implementation also takes techniques from "Levenshtein Distance Embeddings with Poisson Regression for DNA Storage" by Wei et al. (2023).
This is valuable because there are excellent libraries for doing fast GPU accelerated searches for the K nearest neighbors of vectors, like faiss. Algorithms like HNSW allow us to do these searches in logarithmic time, where a brute force levenshtein distance based fuzzy search would need to run in exponential time.
This repo contains a PyTorch implementation of the core ideas from Guo's paper, adapted for ASCII sequences rather than DNA sequences. The implementation includes:
The trained model learns to embed ASCII strings such that the squared Euclidean distance between embeddings approximates the true Levenshtein distance between the strings.
The model uses a 5-layer CNN with average pooling followed by fully connected layers to produce fixed-size embeddings from variable-length ASCII sequences.
import torch
from models import EditDistanceModel
# Load the model
model = EditDistanceModel(embedding_dim=140)
model.load_state_dict(torch.load('megashtein_trained_model.pth'))
model.eval()
# Embed strings
def embed_string(text, max_length=80):
# Pad and convert to tensor
padded = (text + '\0' * max_length)[:max_length]
indices = [min(ord(c), 127) for c in padded]
tensor = torch.tensor(indices, dtype=torch.long).unsqueeze(0)
with torch.no_grad():
embedding = model(tensor)
return embedding
# Example usage
text1 = "hello world"
text2 = "hello word"
emb1 = embed_string(text1)
emb2 = embed_string(text2)
# Compute approximate edit distance
approx_distance = torch.sum((emb1 - emb2) ** 2).item()
print(f"Approximate edit distance: {approx_distance}")
The model is trained using: