Downloads · 30 days
0
Chandij123/disllm-gpt2-split
disllm-gpt2-split is a machine learning model from Chandij123. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch.
This repository contains a split GPT-2 model trained using federated learning with LoRA fine-tuning.
Downloads · 30 days
0
Access
Public
Updated Dec 10, 2025
Repo size
710 MB
Likes
0
Public
Click a slice to open those files.
.pth710 MB · 100%
From the Hugging Face model README
This repository contains a split GPT-2 model trained using federated learning with LoRA fine-tuning.
central_trained_first_part_20251210_162256.pth: Client-side model (first 4 layers)central_trained_second_part_20251210_162256.pth: Server-side model (remaining 8 layers)import torch
from transformers import GPT2Config
# Load the model parts
first_part = torch.load('central_trained_first_part_20251210_162256.pth')
second_part = torch.load('central_trained_second_part_20251210_162256.pth')
# Use with the DisLLM architecture
# (Requires FirstPartModel and SecondPartModel class definitions)
Training improves perplexity from ~45 to ~30-35 across train/val/test sets.
If you use this model, please cite the original DisLLM work and GPT-2 paper.
This model inherits the license from the GPT-2 model and training code.