Downloads · 30 days
0
KhalidKhader/GPU-Project-phi-2-BZU-optimized-inference_opt-optimized
GPU-Project-phi-2-BZU-optimized-inference_opt-optimized is a machine learning model from KhalidKhader. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This is an optimized version of the original model using inferenceopt optimization.
Downloads · 30 days
0
Access
Public
Updated Jun 7, 2025
Repo size
4 MB
Likes
0
Public
Click a slice to open those files.
.json4.4 MB · 50%
From the Hugging Face model README
This is an optimized version of the original model using inference_opt optimization.
# Run the inference script
exec(open("inference.py").read())
# Generate text
result = generate_text("Your prompt here")
print(result)
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
# Load model
tokenizer = AutoTokenizer.from_pretrained("./")
model = AutoModelForCausalLM.from_pretrained("./", torch_dtype=torch.float16).to("cuda")
model.eval()
# Apply optimization
# Basic optimizations applied
# Generate text
def generate_text(prompt):
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
outputs = model.generate(inputs['input_ids'], max_new_tokens=50)
return tokenizer.decode(outputs[0], skip_special_tokens=True)
result = generate_text("The future of AI is")
print(result)
pytorch_model.bin - Optimized model weightsconfig.json - Model configurationtokenizer.json - Tokenizerinference.py - Ready-to-use inference scriptREADME.md - This file