Downloads · 30 days
0
CRLannister/Neural-Network-Based-Language-Model-for-Next-Token-Prediction
Neural-Network-Based-Language-Model-for-Next-Token-Prediction is a machine learning model from CRLannister. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This project is a midterm assignment focused on developing a neural network-based language model for next token prediction. The model was trained using a custom dataset with two languages, English and Amharic. The pro…
Downloads · 30 days
0
Access
Public
Updated Oct 11, 2024
Repo size
1.7 GB
Likes
0
Public
Click a slice to open those files.
.pt1.6 GB · 95%
From the Hugging Face model README
This project is a midterm assignment focused on developing a neural network-based language model for next token prediction. The model was trained using a custom dataset with two languages, English and Amharic. The project incorporates techniques in neural networks to predict the next token in a sequence, demonstrating a non-transformer approach to language modeling.
The main objective of this project was to:
The model was trained using datasets in English and Amharic. The datasets were cleaned and prepared, including tokenization and embedding for improved model training.
A custom tokenizer was created using Byte Pair Encoding (BPE). This tokenizer was trained on five languages: English, Amharic, Sanskrit, Nepali, and Hindi, but the model specifically utilized English and Amharic for this task.
A custom embedding model was employed to convert tokens into vector representations, allowing the neural network to better understand the structure and meaning of the input data.
The project uses an LSTM (Long Short-Term Memory) neural network to predict the next token in a sequence. LSTMs are well-suited for sequential data and are a popular choice for language modeling due to their ability to capture long-term dependencies.
The model’s training and validation loss over time are documented and included in the repository (loss_values.csv). The training curve demonstrates the model's learning progress, with explanations provided for key observations in the loss trends.
Checkpointing was implemented to save model states at different training stages, allowing for partial model evaluations and text generation demos. Checkpoints are included in the repository for reference.
The model's perplexity score, calculated during training, is available in the perplexity.csv file. This score provides an indication of the model's predictive accuracy over time.
A video demo, linked below, demonstrates:
Video Demo Link: YouTube Demo
Note: The data for the project has been taken from saillab/taco-datasets