Downloads · 30 days
3
15% of all-time downloads
Roy2144/Meta-LLama
Meta-LLama is a machine learning model from Roy2144. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This project benchmarks the inference performance of: - vLLM server with Quantized (W8A8) Meta-LLaMA-3.1-8B-Instruct model using LLM Compressor - Hugging Face Transformers running full-precision Meta-LLaMA-3.1-8B-Inst…
Downloads · 30 days
3
15% of all-time downloads
All-time downloads
20
Public
Parameters
8B
9.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.1 GB · 100%
How the weights are stored.
I87B · 87%
From the Hugging Face model README
This project benchmarks the inference performance of:
We evaluate:
llm_comp)Roy2144/Meta-LLama.| Metric | vLLM + Quantized Meta-LLaMA W8A8 | HF Transformers + Meta-LLaMA float16 |
|---|---|---|
| Avg Latency | 0.62 seconds | 2.66 seconds |
| Avg Throughput | 163.99 tokens/sec | 30.46 tokens/sec |
Prompt: "Explain quantization in simple words for a beginner."
Both vLLM Quantized and HF float16 models generated detailed, coherent responses.
Quantized model responses were slightly more concise.