Downloads · 30 days
6
11% of all-time downloads
ngwkhai/CodeLlama-7b-NL2PPlanner
CodeLlama-7b-NL2PPlanner is a machine learning model from ngwkhai. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This model is a fine-tuned version of codellama/CodeLlama-7b-Instruct-hf designed specifically to translate natural language project proposals into concrete, hierarchical directory structures.
Downloads · 30 days
6
11% of all-time downloads
All-time downloads
53
Public
Parameters
6.7B
13.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors13.5 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of codellama/CodeLlama-7b-Instruct-hf designed specifically to translate natural language project proposals into concrete, hierarchical directory structures.
It was trained using a custom Hybrid Loss engine that optimizes for both directory depth correctness and semantic relevance of file/folder names.
The model generates linearized tree sequences using the following special tokens:
<TREE_START> / <TREE_END>: Bounds the entire project structure.<DIR_START> / <DIR_END>: Bounds a directory/folder.<FILE>: Indicates a file.Below is a complete script to load the model, format the prompt correctly, generate the structure, and parse the tokenized output back into a nested Python Dictionary/JSON.
pip install transformers torch
import torch
import re
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "your_hf_username/CodeLlama-7b-NL2PPlanner"
# Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
# Function to parse generated tokens back to a dictionary tree
def regex_parser(text):
tree = {}
stack = [tree]
token_pattern = re.compile(r"(<FILE>|<DIR_START>|<DIR_END>)\s+([^\s<]+)?")
matches = token_pattern.findall(text)
for tag, name in matches:
if tag == "<FILE>" and name:
stack[-1][name] = "file"
elif tag == "<DIR_START>" and name:
new_dir = {}
stack[-1][name] = new_dir
stack.append(new_dir)
elif tag == "<DIR_END>" and len(stack) > 1:
stack.pop()
return tree
# System & User Prompt formatting
system_prompt = """You are a Principal Software Architect. Design the directory structure.
[GRAMMAR RULES]
1. Start with <TREE_START> and end with <TREE_END>.
2. Folders: <DIR_START> name ... <DIR_END>
3. Files: <FILE> name
4. NO JSON. ONLY TOKENS."""
user_input = """Design structure for: "E-Commerce Backend"
[CONTEXT]
- Domain: E-commerce (Web API)
- Desc: A robust backend for handling users, products, and orders.
- Stack: Python, PostgreSQL, Docker
[ARCH]
- **AuthModule**: `/src/auth` (Handles JWT authentication).
- **OrderModule**: `/src/orders` (Processes checkout logic).
[ENTRIES]
['/src/main.py']
[COMMAND]
Generate Linearized Token Sequence."""
formatted_prompt = f"<s>[INST] <<SYS>>\n{system_prompt}\n<</SYS>>\n\n<|user|>{user_input}<|end|> [/INST]"
# Generate
inputs = tokenizer(formatted_prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=1024,
do_sample=False
)
output_text = tokenizer.decode(outputs[0], skip_special_tokens=False)
generated_sequence = output_text.split("[/INST]")[1]
# Parse to JSON
directory_tree = regex_parser(generated_sequence)
print(directory_tree)