Downloads · 30 days
0
DL-Hobbyist/simple-llama
simple-llama is a machine learning model from DL-Hobbyist. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
An open, educational framework for understanding and reproducing the complete training and alignment pipeline of modern Large Language Models (LLMs).
Downloads · 30 days
0
Access
Public
Updated Nov 18, 2025
Repo size
12.3 GB
Likes
0
Public
Click a slice to open those files.
.pth12.3 GB · 100%
From the Hugging Face model README
An open, educational framework for understanding and reproducing the complete training and alignment pipeline of modern Large Language Models (LLMs).
(Can find main GH repo at https://github.com/IvanC987/SimpleLLaMA)
SimpleLLaMA is a comprehensive project designed to demystify the lifecycle of LLM development, starting from raw data to a functioning aligned model.
It provides a transparent implementation of the three main stages of language model creation:
In addition, the project includes modules for data preparation, tokenization, evaluation, and deployment, enabling users to experiment with every major step of the modern LLM pipeline.
To play with the model:
Clone the repository:
git clone https://github.com/IvanC987/SimpleLLaMA
cd SimpleLLaMA
pip install -r requirements.txt
pip install -e .
(More to be added later here once completed)
If you wish to run custom pretraining, fine-tuning, or reinforcement learning, please refer to the Custom Training section in the SimpleLLaMA Documentations page
For an in-depth look into the architecture, experiments, and training methodology, visit the full documentation:
📘 Documentation: https://ivanc987.github.io/SimpleLLaMA/
📄 Technical Report: Technical_Report.md
| Dataset | Metric | Score |
|---|---|---|
| MMLU | Accuracy | XX.X% |
| ARC (Challenge) | Accuracy | XX.X% |
| ARC (Easy) | Accuracy | XX.X% |
| HellaSwag | Accuracy | XX.X% |
| PIQA | Accuracy | XX.X% |
(See the Misc/Benchmarking section in documentations for more details)
This project is licensed under the MIT License.
Feel free to use, extend, or adapt it for research or application purposes.
Ivan Cao
Senior CS Student | University of Mississippi
Open to collaboration and research questions.
GitHub: https://github.com/IvanC987/
This project was inspired by LLaMA, DeepSeek, and various other open source Large Language Models
Papers:
Videos:
Datasets:
Portions of the model architecture are adapted from:
Much of the implementation also borrows design clarity from these excellent open-source efforts.