Downloads · 30 days
16
16% of all-time downloads
joel-crasto/CODE-1
CODE-1 is a text generation model from joel-crasto. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This model is a fine-tuned version of gpt2 on huggingface-course/codeparrot-ds. It achieves the following results on the evaluation set: - Loss: 1.0612
Downloads · 30 days
16
16% of all-time downloads
All-time downloads
103
Public
Parameters
124M
7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors497 MB · 99%
From the Hugging Face model README
This model is a fine-tuned version of gpt2 on huggingface-course/codeparrot-ds. It achieves the following results on the evaluation set:
A fine-tuned version of gpt-2 designed to autocomplete Python code requests left in comments. Based on the HF LLM course (training LMs from scratch)
To be used for Python code autocomplete.
As an example of the model output:
\# create some data
x = np.random.randn(100)
y = np.random.randn(100)
\# create scatter plot with x, y
The current version of the model added below this comment:
fig = plt.figure()
ax = fig.add_subplot(111)
scatter1 = Scatter(x, y)
scatter2 = Scatter(x, y)
ax.scatter(x, y, s=5)
\# create hexbin plot with default gridsize of 10
p = plt.Rectangle((0, 0), 1, 1, fc=hexbin_with_ecdf(5, 10, 20, 20))
ax.0, 0.0, 0.0, zorder=1, fc=0.0, lw-0.1)
ax=0.25)
ax.set_axis('off')
ax.set_axis_axisbelow(p.set_coloraxisbelow)
ax.set_frameon(axisbelow)
_frame_frameon(gca(), set_savefig(gca(), set_gca(), fig.set_gca))
This model is purely experimental in nature for learning purpose. Beyond a few lines, code may generate that is unecessary.
Trained on the train/test split provided by huggingface-course/codeparrot-ds.
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| 2.5503 | 0.0766 | 5000 | 1.7458 |
| 1.6719 | 0.1533 | 10000 | 1.5235 |
| 1.5259 | 0.2299 | 15000 | 1.4221 |
| 1.4422 | 0.3065 | 20000 | 1.3560 |
| 1.3861 | 0.3832 | 25000 | 1.3056 |
| 1.3324 | 0.4598 | 30000 | 1.2576 |
| 1.2863 | 0.5365 | 35000 | 1.2129 |
| 1.2425 | 0.6131 | 40000 | 1.1700 |
| 1.1988 | 0.6897 | 45000 | 1.1321 |
| 1.1634 | 0.7664 | 50000 | 1.0996 |
| 1.131 | 0.8430 | 55000 | 1.0765 |
| 1.1139 | 0.9196 | 60000 | 1.0638 |
| 1.1061 | 0.9963 | 65000 | 1.0612 |
For transparency, experiments were conducted using private infrastructure, which has a carbon efficiency of 0.035 kgCO₂eq/kWh based on Ontario's 2022 energy profile. A cumulative of 22 hours of computation was performed on hardware of type RTX 3080 (TDP of 320 W).
Total emissions are estimated to be 0.25 kgCO₂eq.
Estimations were conducted using the MachineLearning Impact calculator
Lacoste, Alexandre and Luccioni, Alexandra and Schmidt, Victor and Dandres, Thomas, Quantifying the Carbon Emissions of Machine Learning, arXiv preprint arXiv:1910.09700, 2019