Downloads · 30 days
0
Utkarsh524/codellama_utests_full_new_ver2
codellama_utests_full_new_ver2 is a text generation model from Utkarsh524. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
This is a merged model that combines codellama/CodeLlama-7b-hf with a LoRA adapter fine-tuned on embedded C/C++ code and high-quality unit tests using GoogleTest and CppUTest. This version includes enhanced formatting…
Downloads · 30 days
0
Access
Public
Updated Jun 22, 2025
Parameters
6.7B
13.5 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors13.5 GB · 100%
From the Hugging Face model README
This is a merged model that combines codellama/CodeLlama-7b-hf with a LoRA adapter
fine-tuned on embedded C/C++ code and high-quality unit tests using GoogleTest and CppUTest. This version includes enhanced formatting, stop tokens,
and test cleanup mechanisms.
codellama/CodeLlama-7b-hf<|system|>, <|user|>, <|assistant|>, // END_OF_TESTSathrv/Embedded_Unittest2from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "Utkarsh524/codellama_utests_full_new_ver2"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.float16, device_map="auto")
prompt = """<|system|>
Generate comprehensive unit tests for C/C++ code. Cover all edge cases, boundary conditions, and error scenarios.
Output Constraints:
1. ONLY include test code (no explanations, headers, or main functions)
2. Start directly with TEST(...)
3. End after last test case
4. Never include framework boilerplate
<|user|>
Create tests for:
int add(int a, int b) { return a + b; }
<|assistant|>
"""
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=512, eos_token_id=tokenizer.convert_tokens_to_ids("// END_OF_TESTS"))
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
| Step | Description |
|---|---|
| Dataset | athrv/Embedded_Unittest2 (filtered for valid code-test pairs) |
| Preprocessing | Token length filtering (≤4096), special token injection |
| Quantization | 8-bit (BitsAndBytesConfig), llm_int8_threshold=6.0 |
| LoRA Config | r=64, alpha=32, dropout=0.1 on q_proj/v_proj/k_proj/o_proj |
| Training | 4 epochs, batch=4 (effective 8), lr=2e-4, FP16 |
| Optimization | Paged AdamW 8-bit, gradient checkpointing, custom data collator |
| Special Tokens | Added `< |
Dataset Credit: athrv/Embedded_Unittest2
Report Issues: Model's Hugging Face page