Downloads · 30 days
6
13% of all-time downloads
Devy1/CodeLlama-7b-hf-AQLM-3bit-code-2x12
CodeLlama-7b-hf-AQLM-3bit-code-2x12 is a text generation model from Devy1. Use it when you need the model to write or continue text. It is set up for transformers.
1. Introduction 2. Model details 3. Experiments 4. Replication
Downloads · 30 days
6
13% of all-time downloads
All-time downloads
46
Public
Parameters
1.9B
3.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.8 GB · 100%
How the weights are stored.
I161.6B · 85%
From the Hugging Face model README
HuggingFace repository containing the quantized models from the paper "Quantizing Large Language Models for Code Generation: A Differentiated Replication.".
In this study, we evaluate the performance of compressed Deep Learning models on the code generation task. Specifically, we quantize code models such as CodeLlama and DeepSeek Coder at different levels of precision, namely 8, 4, 3, and 2 bits per model parameter, using a SOTA quantization technique for extreme model compression, that is AQLM (Additive Quantization of Language Models).
The complete list of models used in this study is available in our model collection, which is organized by order of appearance in the paper discussion.
More specifically, we named the models as follows:
<base-model>-AQLM-<precision>-<calibration>-<finetuned?>-<hyperparameters>
For example, the model Devy1/CodeLlama-7b-hf-AQLM-2bit-rnd-1x15 has the following features:
More information about the quantization process and hyperparameters can be found in our paper and in the config.json file from this repository.
Below, we present the code generation performance of each quantized model across different experiments. Performance is computed on Python and Java languages using MultiPL-E and McEval benchmarks. More details on the research approach can be found in our paper.
Results are listed by research question and benchmark. By clicking on the "precision" value, you will be redirected to the corresponding model.
| Model | Params | Precision | Size | Python pass@1 | Java pass@1 |
|---|---|---|---|---|---|
| CodeLlama - Base | 7B | Float16 | 13.48 GB | 29.8 | 32.2 |
| 8-bit | 7.47 GB | 29.7 | 31.6 | ||
| 4-bit | 4.00 GB | 29.1 | 30.7 | ||
| 3-bit | 3.80 GB | 24.3 | 26.5 | ||
| 2-bit | 2.26 GB | 16.4 | 14.1 | ||
| DeepSeek-Coder - Base | 7B | Float16 | 13.48 GB | 45.8 | 41.4 |
| 8-bit | 7.48 GB | 46.2 | 41.9 | ||
| 4-bit | 4.00 GB | 45.2 | 41.4 | ||
| 3-bit | 3.80 GB | 41.1 | 37.7 | ||
| 2-bit | 2.27 GB | 27.6 | 23.2 |
| Model | Params | Precision | Size | Python pass@1 | Java pass@1 |
|---|---|---|---|---|---|
| CodeLlama - Base | 7B | Float16 | 13.48 GB | 12.9 | 29.3 |
| 8-bit | 7.47 GB | 12.9 | 29.2 | ||
| 4-bit | 4.00 GB | 15.2 | 25.3 | ||
| 3-bit | 3.80 GB | 10.0 | 21.3 | ||
| 2-bit | 2.26 GB | 5.6 | 11.4 | ||
| DeepSeek-Coder - Base | 7B | Float16 | 13.48 GB | 41.8 | 42.6 |
| 8-bit | 7.48 GB | 42.5 | 42.8 | ||
| 4-bit | 4.00 GB | 40.7 | 45.9 | ||
| 3-bit | 3.80 GB | 36.2 | 34.5 | ||
| 2-bit | 2.27 GB | 13.7 | 23.6 |
| Model | Params | Precision | Size | Python pass@1 | Java pass@1 |
|---|---|---|---|---|---|
| CodeLlama - Base | 7B | Float16 | 13.48 GB | 29.8 | 32.2 |
| 3-bit | 3.80 GB | 24.3 | 26.5 | ||
| 2-bit | 2.26 GB | 16.4 | 14.1 | ||
| 3-bit + Fine-tuning | 3.80 GB | <span style="color:red;">▼</span> 24.0 | <span style="color:green;">▲</span> 27.8 | ||
| 2-bit + Fine-tuning | 2.26 GB | <span style="color:green;">▲</span> 19.9 | <span style="color:green;">▲</span> 19.0 | ||
| DeepSeek-Coder - Base | 7B | Float16 | 13.48 GB | 45.8 | 41.4 |
| 3-bit | 3.80 GB | 41.1 | 37.7 | ||
| 2-bit | 2.27 GB | 27.6 | 23.2 | ||
| 3-bit + Fine-tuning | 3.80 GB | <span style="color:green;">▲</span> 41.8 | <span style="color:red;">▼</span> 37.7 | ||
| 2-bit + Fine-tuning | 2.27 GB | <span style="color:green;">▲</span> 33.0 | <span style="color:green;">▲</span> 26.8 |
| Model | Params | Precision | Size | Python pass@1 | Java pass@1 |
|---|---|---|---|---|---|
| CodeLlama - Base | 7B | Float16 | 13.48 GB | 12.9 | 29.3 |
| 3-bit | 3.80 GB | 10.0 | 21.3 | ||
| 2-bit | 2.26 GB | 5.6 | 11.4 | ||
| 3-bit + Fine-tuning | 3.80 GB | <span style="color:green;">▲</span> 10.8 | <span style="color:green;">▲</span> 22.0 | ||
| 2-bit + Fine-tuning | 2.26 GB | <span style="color:green;">▲</span> 7.6 | <span style="color:green;">▲</span> 14.3 | ||
| DeepSeek-Coder - Base | 7B | Float16 | 13.48 GB | 41.8 | 42.6 |
| 3-bit | 3.80 GB | 36.2 | 34.5 | ||
| 2-bit | 2.27 GB | 13.7 | 23.6 | ||
| 3-bit + Fine-tuning | 3.80 GB | <span style="color:red;">▼</span> 35.6 | <span style="color:red;">▼</span> 32.4 | ||
| 2-bit + Fine-tuning | 2.27 GB | <span style="color:green;">▲</span> 20.2 | <span style="color:green;">▲</span> 27.0 |
| Model | Params | Precision | Size | Python pass@1 | Java pass@1 |
|---|---|---|---|---|---|
| CodeLlama - Base | 7B | Float16 - Baseline | 13.48 GB | 29.8 | 32.2 |
| 8-bit with Random samples | 7.47 GB | 29.7 | 31.6 | ||
| 8-bit with Mixed samples | 7.47 GB | <span style="color:red;">▼</span> 29.7 | <span style="color:green;">▲</span> 32.3 | ||
| 8-bit with Code samples | 7.47 GB | <span style="color:red;">▼</span> 29.2 | <span style="color:green;">▲</span> 32.0 | ||
| 4-bit with Random samples | 4.00 GB | 29.1 | 30.7 | ||
| 4-bit with Mixed samples | 4.00 GB | <span style="color:red;">▼</span> 29.0 | <span style="color:green;">▲</span> 31.4 | ||
| 4-bit with Code samples | 4.00 GB | <span style="color:green;">▲</span> 30.2 | <span style="color:red;">▼</span> 29.8 | ||
| 3-bit with Random samples | 3.80 GB | 24.3 | 26.5 | ||
| 3-bit with Mixed samples | 3.80 GB | <span style="color:green;">▲</span> 28.2 | <span style="color:green;">▲</span> 28.4 | ||
| 3-bit with Code samples | 3.80 GB | <span style="color:green;">▲</span> 27.0 | <span style="color:green;">▲</span> 28.0 | ||
| 2-bit with Random samples | 2.26 GB | 16.4 | 14.1 | ||
| 2-bit with Mixed samples | 2.26 GB | <span style="color:green;">▲</span> 23.9 | <span style="color:green;">▲</span> 21.5 | ||
| 2-bit with Code samples | 2.26 GB | <span style="color:green;">▲</span> 24.1 | <span style="color:green;">▲</span> 19.4 | ||
| DeepSeek-Coder - Base | 7B | Float16 - Baseline | 13.48 GB | 45.8 | 41.4 |
| 8-bit with Random samples | 7.48 GB | 46.2 | 41.9 | ||
| 8-bit with Mixed samples | 7.48 GB | <span style="color:red;">▼</span> 45.4 | <span style="color:green;">▲</span> 43.2 | ||
| 8-bit with Code samples | 7.48 GB | <span style="color:red;">▼</span> 45.9 | <span style="color:red;">▼</span> 41.7 | ||
| 4-bit with Random samples | 4.00 GB | 45.2 | 41.4 | ||
| 4-bit with Mixed samples | 4.00 GB | <span style="color:red;">▼</span> 44.5 | <span style="color:green;">▲</span> 41.8 | ||
| 4-bit with Code samples | 4.00 GB | <span style="color:red;">▼</span> 44.2 | <span style="color:red;">▼</span> 40.6 | ||
| 3-bit with Random samples | 3.80 GB | 41.1 | 37.7 | ||
| 3-bit with Mixed samples | 3.80 GB | <span style="color:green;">▲</span> 43.7 | <span style="color:green;">▲</span> 39.1 | ||
| 3-bit with Code samples | 3.80 GB | <span style="color:green;">▲</span> 42.5 | <span style="color:green;">▲</span> 38.7 | ||
| 2-bit with Random samples | 2.27 GB | 27.6 | 23.2 | ||
| 2-bit with Mixed samples | 2.27 GB | <span style="color:green;">▲</span> 35.7 | <span style="color:green;">▲</span> 27.4 | ||
| 2-bit with Code samples | 2.27 GB | <span style="color:green;">▲</span> 34.8 | <span style="color:green;">▲</span> 27.5 |
| Model | Params | Precision | Size | Python pass@1 | Java pass@1 |
|---|---|---|---|---|---|
| CodeLlama - Base | 7B | Float16 - Baseline | 13.48 GB | 12.9 | 29.3 |
| 8-bit with Random samples | 7.47 GB | 12.9 | 29.2 | ||
| 8-bit with Mixed samples | 7.47 GB | <span style="color:green;">▲</span> 13.7 | <span style="color:red;">▼</span> 28.6 | ||
| 8-bit with Code samples | 7.47 GB | <span style="color:red;">▼</span> 12.3 | <span style="color:green;">▲</span> 29.5 | ||
| 4-bit with Random samples | 4.00 GB | 15.2 | 25.3 | ||
| 4-bit with Mixed samples | 4.00 GB | <span style="color:red;">▼</span> 13.0 | <span style="color:green;">▲</span> 30.3 | ||
| 4-bit with Code samples | 4.00 GB | <span style="color:red;">▼</span> 11.1 | <span style="color:green;">▲</span> 25.8 | ||
| 3-bit with Random samples | 3.80 GB | 10.0 | 21.3 | ||
| 3-bit with Mixed samples | 3.80 GB | <span style="color:green;">▲</span> 12.3 | <span style="color:green;">▲</span> 25.5 | ||
| 3-bit with Code samples | 3.80 GB | <span style="color:green;">▲</span> 10.8 | <span style="color:red;">▼</span> 19.9 | ||
| 2-bit with Random samples | 2.26 GB | 5.6 | 11.4 | ||
| 2-bit with Mixed samples | 2.26 GB | <span style="color:green;">▲</span> 11.1 | <span style="color:green;">▲</span> 12.8 | ||
| 2-bit with Code samples | 2.26 GB | <span style="color:green;">▲</span> 6.1 | <span style="color:green;">▲</span> 12.8 | ||
| DeepSeek-Coder - Base | 7B | Float16 - Baseline | 13.48 GB | 41.8 | 42.6 |
| 8-bit with Random samples | 7.48 GB | 42.5 | 42.8 | ||
| 8-bit with Mixed samples | 7.48 GB | <span style="color:green;">▲</span> 42.7 | <span style="color:red;">▼</span> 42.5 | ||
| 8-bit with Code samples | 7.48 GB | <span style="color:red;">▼</span> 41.3 | <span style="color:red;">▼</span> 42.7 | ||
| 4-bit with Random samples | 4.00 GB | 40.7 | 45.9 | ||
| 4-bit with Mixed samples | 4.00 GB | <span style="color:red;">▼</span> 39.0 | <span style="color:red;">▼</span> 42.8 | ||
| 4-bit with Code samples | 4.00 GB | <span style="color:red;">▼</span> 39.8 | <span style="color:green;">▲</span> 46.3 | ||
| 3-bit with Random samples | 3.80 GB | 36.2 | 34.5 | ||
| 3-bit with Mixed samples | 3.80 GB | <span style="color:red;">▼</span> 35.5 | <span style="color:green;">▲</span> 42.8 | ||
| 3-bit with Code samples | 3.80 GB | <span style="color:green;">▲</span> 36.5 | <span style="color:green;">▲</span> 45.6 | ||
| 2-bit with Random samples | 2.27 GB | 13.7 | 23.6 | ||
| 2-bit with Mixed samples | 2.27 GB | <span style="color:green;">▲</span> 26.2 | <span style="color:green;">▲</span> 29.1 | ||
| 2-bit with Code samples | 2.27 GB | <span style="color:green;">▲</span> 24.6 | <span style="color:green;">▲</span> 28.0 |
| Model | Params | Precision | Size (GB) | Python pass@1 | Dec (%) | Java pass@1 | Dec (%) |
|---|---|---|---|---|---|---|---|
| CodeLlama - Base | 7B | Float16 | 13.48 | 29.8 | --- | 32.2 | --- |
| 2-bit | 2.26 | 23.9 | -19.8 | 21.5 | -33.2 | ||
| 2-bit + Finetuning | 2.26 | 25.5 | -14.4 | 26.5 | -17.7 | ||
| 13B | Float16 | 24.25 | 34.3 | --- | 38.3 | --- | |
| 2-bit | 3.98 | 30.9 | -9.9 | 27.7 | -27.7 | ||
| 2-bit + Finetuning | 3.98 | 30.1 | -12.2 | 32.8 | -14.4 | ||
| 34B | Float16 | 62.74 | 41.9 | --- | 44.1 | --- | |
| 2-bit | 9.54 | 37.1 | -11.5 | 32.7 | -25.9 | ||
| 2-bit + Finetuning | 9.54 | 36.0 | -14.1 | 36.1 | -18.1 | ||
| DeepSeek-Coder - Base | 1B | Float16 | 2.57 | 28.4 | --- | 28.8 | --- |
| 2-bit | 0.61 | 13.9 | -51.1 | 6.6 | -77.1 | ||
| 2-bit + Finetuning | 0.61 | 21.7 | -23.6 | 14.7 | -49.0 | ||
| 7B | Float16 | 13.48 | 45.8 | --- | 41.4 | --- | |
| 2-bit | 2.27 | 35.7 | -22.1 | 27.4 | -33.8 | ||
| 2-bit + Finetuning | 2.27 | 36.4 | -20.5 | 32.8 | -20.8 | ||
| 33B | Float16 | 62.16 | 52.1 | --- | 47.3 | --- | |
| 2-bit | 9.38 | 43.4 | -16.7 | 34.5 | -27.1 | ||
| 2-bit + Finetuning | 9.38 | 43.0 | -17.5 | 38.7 | -18.2 |
| Model | Params | Precision | Size (GB) | Python pass@1 | Dec (%) | Java pass@1 | Dec (%) |
|---|---|---|---|---|---|---|---|
| CodeLlama - Base | 7B | Float16 | 13.48 | 12.9 | --- | 29.3 | --- |
| 2-bit | 2.26 | 11.1 | -14.0 | 12.8 | -56.3 | ||
| 2-bit + Finetuning | 2.26 | 13.0 | -0.8 | 18.3 | -37.5 | ||
| 13B | Float16 | 24.25 | 18.9 | --- | 40.9 | --- | |
| 2-bit | 3.98 | 9.4 | -50.3 | 22.3 | -45.5 | ||
| 2-bit + Finetuning | 3.98 | 10.4 | -45.0 | 27.8 | -32.0 | ||
| 34B | Float16 | 62.74 | 29.0 | --- | 39.2 | --- | |
| 2-bit | 9.54 | 17.6 | -39.3 | 25.2 | -35.7 | ||
| 2-bit + Finetuning | 9.54 | 19.0 | -34.5 | 31.6 | -19.4 | ||
| DeepSeek-Coder - Base | 1B | Float16 | 2.57 | 23.8 | --- | 42.0 | --- |
| 2-bit | 0.61 | 4.4 | -81.5 | 8.5 | -79.8 | ||
| 2-bit + Finetuning | 0.61 | 6.9 | -71.0 | 15.5 | -63.1 | ||
| 7B | Float16 | 13.48 | 41.8 | --- | 42.6 | --- | |
| 2-bit | 2.27 | 26.2 | -37.3 | 29.1 | -31.7 | ||
| 2-bit + Finetuning | 2.27 | 30.1 | -28.0 | 31.0 | -27.2 | ||
| 33B | Float16 | 62.16 | 55.5 | --- | 57.0 | --- | |
| 2-bit | 9.38 | 36.9 | -33.5 | 39.2 | -31.2 | ||
| 2-bit + Finetuning | 9.38 | 39.8 | -28.3 | 44.0 | -22.8 |
The scripts used to quantize and evaluate the models are available in our GitHub repository (link).
Model predictions, statistical results, and datasets are instead available in our Zenodo repository (link).