Downloads · 30 days
7
2% of all-time downloads
medxiaorudan/CodeLlama_CPP_FineTuned
CodeLlama_CPP_FineTuned is a machine learning model from medxiaorudan. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as llama2.
This model has been fine-tuned using the CodeLlama base, incorporating C++ code sourced from the 'codeparrot/xlcost-text-to-code' dataset. It possesses the capability to generate C++ code based on provided task descri…
Downloads · 30 days
7
2% of all-time downloads
All-time downloads
298
Public
Repo size
67.7 MB
Likes
2
Public
Click a slice to open those files.
.pt67.2 MB · 97%
From the Hugging Face model README
This model has been fine-tuned using the CodeLlama base, incorporating C++ code sourced from the 'codeparrot/xlcost-text-to-code' dataset. It possesses the capability to generate C++ code based on provided task descriptions.
If you get the error "ValueError: Tokenizer class CodeLlamaTokenizer does not exist or is not currently imported." make sure your Transformer version is 4.33.0 and accelerate>=0.20.3.
from transformers import AutoTokenizer
import transformers
import torch
model = "medxiaorudan/CodeLlama_CPP_FineTuned"
tokenizer = AutoTokenizer.from_pretrained(model)
pipeline = transformers.pipeline(
"text-generation",
model=model,
torch_dtype=torch.float16,
device_map="auto",
)
prompt = """
Use the Task below and write the Response, which is a programming code that can solve the Task.
### Task:
Generate a C++ program that accepts numeric input from the user and maintains a record of previous user inputs with timestamps. Ensure the program sorts the user inputs in ascending order based on the provided numeric input. Enhance the program to display timestamps along with the sorted user inputs.
### Response:
"""
sequences_finetune = pipeline(
prompt,
do_sample=True,
top_k=10,
temperature=0.1,
top_p=0.95,
num_return_sequences=1,
eos_token_id=tokenizer.eos_token_id,
max_length=600,
add_special_tokens=False
)
for seq in sequences_finetune:
print(f"Result: {seq['generated_text']}")
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
Use the code below to get started with the model.
[More Information Needed]
https://huggingface.co/datasets/codeparrot/xlcost-text-to-code
[More Information Needed]
The detailed training report is here.
[More Information Needed]
[More Information Needed]
I have use the Catch2 unit test framework for generated C++ code snippets correctness verification.
Todo: Use the pass@k metric with the HumanEval-X dataset to verify the performance of the model.
https://huggingface.co/datasets/THUDM/humaneval-x
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
I used 4 NVIDIA A40-48Q GPU server configured with Python 3.10 and Cuda 12.2 to run the code in this article. It ran for about eight hours.
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
BibTeX:
[More Information Needed]
APA:
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]