Downloads Β· 30 days
5
31% of all-time downloads
aijadugar/transformer
transformer is a machine learning model from aijadugar. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A complete implementation of the Transformer architecture from the paper Attention Is All You Need, built entirely with PyTorch.
Downloads Β· 30 days
5
31% of all-time downloads
All-time downloads
16
Public
Repo size
256 MB
Likes
0
Public
Click a slice to open those files.
.safetensors256 MB Β· 100%
From the Hugging Face model README
<p align="center"> <img src="https://img.shields.io/badge/Python-3.10+-3776AB?logo=python&logoColor=white"> <img src="https://img.shields.io/badge/PyTorch-2.x-EE4C2C?logo=pytorch&logoColor=white"> <img src="https://img.shields.io/badge/License-MIT-success"> <img src="https://img.shields.io/badge/Status-Active-brightgreen"> </p>A complete implementation of the Transformer architecture from the paper Attention Is All You Need, built entirely with PyTorch.
The Transformer changed Natural Language Processing by replacing recurrent networks with self-attention, allowing models to process entire sequences in parallel.
This repository implements every major component from scratch without using torch.nn.Transformer.
It is designed for:
flowchart TD
A[Source Tokens]
B[Embedding]
C[Positional Encoding]
D["Encoder Γ N"]
E[Encoder Memory]
F[Target Tokens]
G[Embedding]
H[Positional Encoding]
I["Decoder Γ N"]
J[Linear Layer]
K[Vocabulary Probabilities]
A --> B --> C --> D --> E
F --> G --> H --> I
E --> I
I --> J --> K
graph TD
Transformer
Transformer --> Embedding
Transformer --> PositionalEncoding
Transformer --> Encoder
Transformer --> Decoder
Transformer --> Linear
Encoder --> MultiHeadAttention
Encoder --> FeedForward
Encoder --> LayerNorm
Decoder --> MaskedAttention
Decoder --> CrossAttention
Decoder --> FeedForward2
Decoder --> LayerNorm2
Each encoder layer consists of:
Input
β
βΌ
Multi-Head Self Attention
β
Add & LayerNorm
β
Feed Forward Network
β
Add & LayerNorm
β
Output
Each decoder layer consists of:
Input
β
βΌ
Masked Multi-Head Attention
β
Add & LayerNorm
β
Cross Attention
β
Add & LayerNorm
β
Feed Forward Network
β
Add & LayerNorm
β
Output
transformer-from-scratch/
βββ model.py
βββ encoder.py
βββ decoder.py
βββ attention.py
βββ positional_encoding.py
βββ config.py
βββ train.py
βββ inference.py
βββ README.md
β
βββ notebooks/
| Hyperparameter | Value |
|---|---|
| Encoder Layers | 6 |
| Decoder Layers | 6 |
| Attention Heads | 8 |
| Embedding Size | 512 |
| Feed Forward Size | 2048 |
| Maximum Sequence Length | 5000 |
import torch
from model import Transformer
src = torch.randint(0, 10000, (64, 20))
tgt = torch.randint(0, 12000, (64, 15))
model = Transformer(
src_vocab_size=10000,
tgt_vocab_size=12000,
num_heads=8,
num_layers=6,
emb_dim=512,
nn_dim=2048
)
output = model(src, tgt)
print(output.shape)
Output
torch.Size([64, 15, 12000])
sequenceDiagram
participant Source
participant Encoder
participant Decoder
participant Output
Source->>Encoder: Source Tokens
Encoder->>Encoder: Self Attention
Encoder-->>Decoder: Encoder Memory
Decoder->>Decoder: Masked Self Attention
Decoder->>Encoder: Cross Attention
Decoder->>Output: Vocabulary Logits
criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(
model.parameters(),
lr=1e-4
)
After studying this repository, you'll understand:
Attention Is All You Need
Ashish Vaswani et al.
NeurIPS 2017
If this project helped you understand Transformers, consider giving it a β on GitHub.
It helps others discover the project and motivates future improvements.
Released under the MIT License.