Downloads · 30 days
0
Raghav63/next_token_prediction
next_token_prediction is a machine learning model from Raghav63. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This project is a Next Token Prediction model built using PyTorch. The goal of this model is to predict the next word/token in a sequence of text, which is useful for tasks like text generation, autocompletion, and la…
Downloads · 30 days
0
Access
Public
Updated Oct 14, 2024
Repo size
392 MB
Likes
0
Public
Click a slice to open those files.
.pth392 MB · 100%
From the Hugging Face model README
This project is a Next Token Prediction model built using PyTorch. The goal of this model is to predict the next word/token in a sequence of text, which is useful for tasks like text generation, autocompletion, and language modeling.
The model is designed to process sequences of text and predict the most likely token to follow a given sequence. It utilizes a neural network architecture with PyTorch for training and inference. The model can run on both CPU and GPU for efficient training and prediction.
PyTorch-based: Built using PyTorch, making it flexible and easy to integrate with other PyTorch projects. Next Token Prediction: Designed specifically for predicting the next word/token in a text sequence. GPU Support: Can be trained and run on GPUs for improved performance. Installation To install the required dependencies, you can use the following command:
bash Copy code pip install torch torchvision Ensure you have PyTorch installed, and if you wish to use the GPU, install the appropriate CUDA version.
Usage
Clone the repository bash download code ipynb from raghavendra_midterm.ipynb cd next-token-prediction
Prepare the Data Ensure that your input text file is in the correct format. The model expects a text file named input.txt, which contains sequences of text.
Training the Model You can train the model using the provided Jupyter notebook:
bash Copy code jupyter notebook Raghavendra_midterm.ipynb The notebook includes steps for:
Importing necessary libraries. Loading and preprocessing the data. Setting up and training the model. 4. Running Inference. To use the model for next token prediction, you can run the inference code in the notebook. You can modify the following code to predict the next token for any given text input:.
python Copy code text = "Your input text here". predicted_token = model.predict_next_token(text). print(predicted_token). 5. GPU Support. Ensure that GPU support is enabled by checking the following:
python Copy code from raghvendra_midterm.ipynb. torch.cuda.is_available(). The model will automatically utilize the GPU if available.
Model Architecture The model uses a standard neural network architecture, including: Embedding layers for text input. 5 LSTM for sequence modeling. 5 Fully connected layers for prediction. You can modify the architecture in the notebook if needed.
You can found the different checkpoints of the model during training where checkpoint "checkpoint_step_450000.pth" is the best model with low loss for both train and Validation loss
After training, the model should be able to predict the next token in a sequence with reasonable accuracy. The performance can be evaluated using metrics loss.