Downloads · 30 days
15
21% of all-time downloads
nm-testing/Llama-3.1-8B-Instruct-QKV-Cache-FP8
Llama-3.1-8B-Instruct-QKV-Cache-FP8 is a machine learning model from nm-testing. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
The following accuracy is using lm-eval and HF:
Downloads · 30 days
15
21% of all-time downloads
All-time downloads
70
Public
Parameters
8B
9.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.1 GB · 100%
How the weights are stored.
F8_E4M37B · 87%
From the Hugging Face model README
The following accuracy is using lm-eval and HF:
lm_eval \
--model hf \
--model_args pretrained="nm-testing/Llama-3.1-8B-Instruct-QKV-Cache-FP8",dtype=auto,device_map="auto",max_length=100000 \
--tasks "niah_single_1" \
--write_out \
--batch_size 1 \
--output_path "niah_single_1.json" \
--show_config
Replace niah_single_1 with niah_single_2,niah_single_2,niah_multikey_1, niah_multikey_2, niah_multikey_3
The following accuracy is using lm-eval and vLLM:
lm_eval \
--model vllm \
--model_args pretrained="nm-testing/Llama-3.1-8B-Instruct-QKV-Cache-FP8",dtype=auto,add_bos_token=True,max_model_len=131072,tensor_parallel_size=1,gpu_memory_utilization=0.7,enable_chunked_prefill=True,trust_remote_code=True \
--tasks "niah_single_1" \
--write_out \
--batch_size 1 \
--output_path "niah_single_1.json" \
--show_config
Replace niah_single_1 with niah_single_2,niah_single_2,niah_multikey_1, niah_multikey_2, niah_multikey_3