Downloads · 30 days
27
14% of all-time downloads
FarhanAK128/CustomGPT
CustomGPT is a text generation model from FarhanAK128. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
CustomGPT is an LLM which is built, train, instruction-finetuned from scratch and evaluated using the LLM-as-a-judge method. This project shows my learning about developing a custom LLM architecture from scratch and i…
Downloads · 30 days
27
14% of all-time downloads
All-time downloads
193
Public
Parameters
431M
1.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.7 GB · 100%
From the Hugging Face model README
CustomGPT is an LLM which is built, train, instruction-finetuned from scratch and evaluated using the LLM-as-a-judge method. This project shows my learning about developing a custom LLM architecture from scratch and its deployment on huggingface. It should be noted that this model is not to be used in production as it only for demo purpose which showcases my learning of LLM engineering. GPT pretrained weights have been used which are further fine-tuned on a small instruction dataset.
This model is fully compatible with the Hugging Face transformers ecosystem and can be loaded using AutoModel.from_pretrained.
from transformers import AutoModel
import tiktoken
# Load tokenizer
tokenizer = tiktoken.get_encoding("gpt2")
# Load model
model_id = "FarhanAK128/CustomGPT"
model = AutoModel.from_pretrained(
model_id,
trust_remote_code=True
)
# Example prediction
input = {'instruction': 'Rewrite the sentence using a simile.',
'input': 'The car is very fast.'
}
response = model.generate_response(input, tokenizer)
print(response) # The car is as fast as a cheetah.
Note: This model uses a custom .generate() method defined in the repository and requires trust_remote_code=True to function.
The model was trained on a small instruction dataset having 1100 input-output pairs

The 355M custom LLM is evaluated using LLama-3.1-8b-instant as an automated judge. For each input, the model’s response is compared to the ground-truth output, and the judge assigns a score from 0 to 100 based on correctness. Scores are extracted as integers and aggregated to report the average performance across the test dataset which comes out to be 44%.
Farhan Ali Khan
For questions or feedback, please reach out via my Hugging Face profile: FarhanAK128