Downloads · 30 days
17
0% of all-time downloads
mgoin/Minitron-4B-Base-FP8
Minitron-4B-Base-FP8 is a text generation model from mgoin. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
FP8 quantized checkpoint of nvidia/Minitron-4B-Base for use with vLLM.
Downloads · 30 days
17
0% of all-time downloads
All-time downloads
8.9K
Public
Parameters
4.2B
5.8 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors5.8 GB · 100%
How the weights are stored.
F8_E4M32.6B · 62%
From the Hugging Face model README
FP8 quantized checkpoint of nvidia/Minitron-4B-Base for use with vLLM.
lm_eval --model vllm --model_args pretrained=mgoin/Minitron-4B-Base-FP8 --tasks gsm8k --num_fewshot 5 --batch_size auto
vllm (pretrained=mgoin/Minitron-4B-Base-FP8), gen_kwargs: (None), limit: None, num_fewshot: 5, batch_size: auto
|Tasks|Version| Filter |n-shot| Metric | |Value | |Stderr|
|-----|------:|----------------|-----:|-----------|---|-----:|---|-----:|
|gsm8k| 3|flexible-extract| 5|exact_match|↑ |0.2305|± |0.0116|
| | |strict-match | 5|exact_match|↑ |0.2282|± |0.0116|