Downloads · 30 days
0
DarcyCheng/RNN-based-NMT
RNN-based-NMT is a machine learning model from DarcyCheng. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A PyTorch implementation of RNN-based Neural Machine Translation system for Chinese-to-English translation, featuring LSTM encoder-decoder architecture with attention mechanisms.
Downloads · 30 days
0
Access
Public
Updated Dec 28, 2025
Repo size
290 MB
Likes
0
Public
Click a slice to open those files.
.optim190 MB · 65%
From the Hugging Face model README
A PyTorch implementation of RNN-based Neural Machine Translation system for Chinese-to-English translation, featuring LSTM encoder-decoder architecture with attention mechanisms.
This repository implements a RNN-based Neural Machine Translation system with the following key components:
Model: Implement a model using LSTM, with both the encoder and decoder consisting of unidirectional layers.
Attention mechanism: Implement the attention mechanism and investigate the impact of different alignment functions—such as dot-product, multiplicative, and additive—on model performance.
Training policy: Compare the effectiveness of Teacher Forcing and Free Running strategies.
Decoding policy: Compare the effectiveness of greedy and beam-search decoding strategies.
The compressed package contains four JSONL files, corresponding respectively to the small training set, large training set, validation set, and test set, with sizes of 100k, 10k, 500, and 200 samples. Each line in a JSONL file contains one parallel sentence pair. The final model performance will be evaluated based on results on the test set.
Each line in the JSONL files follows this format:
{"chinese": "中文句子", "english": "English sentence"}
translation_dataset_zh_en/
├── train_small.jsonl # 100k samples
├── train_large.jsonl # 10k samples
├── dev.jsonl # 500 samples
└── test.jsonl # 200 samples
The data preprocessing pipeline includes:
git clone <repository-url>
cd RNN_NMT
pip install -r requirement.txt
import nltk
nltk.download('punkt')
Key dependencies include:
torch>=1.12.0 - Deep learning frameworknumpy>=1.21.0 - Numerical computinghydra-core>=1.3.0 - Configuration managementomegaconf>=2.2.0 - Configuration objectssentencepiece>=0.1.96 - English subword tokenizationjieba>=0.42.1 - Chinese word segmentationnltk>=3.7 - BLEU score evaluationtqdm>=4.62.0 - Progress barsTrain the model using the default configuration:
python train.py
The training script uses Hydra for configuration management. You can override configuration parameters via command line:
python train.py attention_type=additive teacher_forcing_ratio=0.7 decoding_strategy=beam-search beam_size=5
Main training parameters can be configured in configs/train.yaml:
attention_type: "dot-product", "multiplicative", or "additive"teacher_forcing_ratio: Ratio for teacher forcing (0.0-1.0)decoding_strategy: "greedy" or "beam-search"beam_size: Beam size for beam search (default: 5)learning_rate: Initial learning rate (default: 5e-5)batch_size: Batch size (default: 128)max_epochs: Maximum training epochs (default: 50)Evaluate a trained model on the test set:
python eval.py
Or with custom parameters:
python eval.py --model_path <path_to_model> --data_path <path_to_data> --decoding_strategy beam-search --beam_size 5
Alternatively, you can use inference.py directly (same functionality):
python inference.py --model_path <path_to_model> --data_path <path_to_data> --decoding_strategy beam-search --beam_size 5
The evaluation script will output:
During training, the model saves:
save_dir/model_rnn_best.pt (best validation perplexity)save_dir/model_rnn_last.pt (most recent checkpoint).optim extension)To resume training from a checkpoint:
# In configs/train.yaml
resume_from_model: "save_dir/model_rnn_last.pt"
RNN_NMT/
├── configs/
│ └── train.yaml # Training configuration
├── dataset/
│ └── vocab.py # Vocabulary management
├── models/
│ ├── rnn_nmt.py # Main NMT model
│ ├── model_embeddings.py # Embedding layers
│ └── char_decoder.py # Character-level decoder
├── utils/
│ ├── utils.py # Utility functions (BLEU, batching, etc.)
│ └── preprocess_data.py # Data preprocessing
├── train.py # Training script
├── inference.py # Evaluation script
├── eval.py # Evaluation script (alias for inference.py)
├── requirement.txt # Python dependencies
└── README.md # This file
The model performance is evaluated using:
Training metrics are automatically saved to training_metrics.json for visualization and analysis.
感谢以下几个仓库:
Jieba (Chinese word segmentation tool): https://github.com/fxsjy/jieba
SentencePiece (English and multilingual subword tokenization tool): https://github.com/google/sentencepiece
RNN Machine Translation: https://github.com/pi-tau/machine-translation
[Add your license information here]
[Add your contact information here]