Downloads · 30 days
0
jahnavi0803/llama3-8b-code-judge
llama3-8b-code-judge is a text generation model from jahnavi0803. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
A QLoRA fine-tuned adapter for Llama 3.1 8B Instruct, trained to deliver deadpan comedic verdicts on code, as the "Judge" persona in the AI Code Court project.
Downloads · 30 days
0
Access
Public
Updated Jul 26, 2026
Repo size
71.8 MB
Likes
0
Public
Click a slice to open those files.
.safetensors54.6 MB · 76%
From the Hugging Face model README
A QLoRA fine-tuned adapter for Llama 3.1 8B Instruct, trained to deliver deadpan comedic verdicts on code, as the "Judge" persona in the AI Code Court project.
This is a LoRA adapter fine-tuned on top of meta-llama/Llama-3.1-8B-Instruct, trained to act as a courtroom judge that delivers short, dry, comedic verdicts on code quality, given a code snippet and short "prosecutor" and "defense" arguments. It is part of AI Code Court, a project where multiple LLMs argue about the quality of user-submitted code.
Generating a short, stylized, comedic "verdict" on a piece of code, given a code snippet and two short arguments (a "prosecutor" case and a "defense" case), in the specific prompt format described in "How to Get Started" below.
Intended to be plugged into the AI Code Court app as the "Judge" role in a three-LLM courtroom simulation (GPT-4o as prosecutor, Gemini as defense, this model as judge).
Not intended for real code review, security auditing, or any decision with real technical or business consequences. Its "reasoning" output is stylistic and comedic, not a substitute for actual static analysis, linting, or human code review.
Trained on only 50 examples, all written or generated by one person in a single sitting, so its "judgment" reflects one person's sense of humor and coding opinions rather than a broad or vetted standard. It has not been evaluated on code outside common, simple Python patterns (the training set skewed toward classic bugs: SQL injection, bare excepts, hardcoded secrets, off-by-one errors, global state misuse). Performance on unusual, large, or multi-file code, or languages other than Python, is unknown.
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use for entertainment and portfolio/demo purposes only — do not use its verdicts as a substitute for real code review, linting, or security scanning.
Use the code below to get started with the model.
Confirmed working. Given a code snippet and courtroom arguments in the correct prompt format, it reliably generates a structured verdict:
VERDICT: [Guilty of Bad Code / Not Guilty / Guilty with Mitigating Circumstances] REASONING: [one dry, sarcastic sentence] SENTENCE: [a punchy label like "Refactor Immediately"] ONE-LINER: [a quotable roast or compliment]
Example real output from this model:
VERDICT: Guilty of Bad Code REASONING: Because who needs error messages, anyway? SENTENCE: Refactor Immediately ONE-LINER: "This code is so trusting, it's practically a hostage situation."
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
)
base_model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B-Instruct",
quantization_config=bnb_config,
device_map="auto",
)
model = PeftModel.from_pretrained(base_model, "jahnavi0803/llama3-8b-code-judge")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
Prompt format matters: this model was trained on a specific structure (system prompt + code wrapped in triple backticks + labeled PROSECUTOR/DEFENSE arguments + an explicit instruction to use the exact format). See the training code for the exact template used.
50 courtroom transcripts (code snippet + prosecutor argument + defense argument → verdict) collected live from the AI Code Court app itself, using GPT-4o as prosecutor and Gemini as defense, then converted into prompt/completion pairs for supervised fine-tuning. Dataset prep script: prepare_dataset.py.
QLoRA: base model loaded in 4-bit (NF4 quantization), LoRA adapter (rank 16, alpha 32) applied to the attention projection layers (q_proj, k_proj, v_proj, o_proj), trained with supervised fine-tuning (SFTTrainer).
Each training example was formatted as: system prompt (judge persona + output format instructions) + code wrapped in triple backticks + labeled PROSECUTOR/DEFENSE arguments + instruction to deliver the verdict, followed by the target completion (the verdict itself) and an end-of-sequence token.
Trained on a single free-tier Google Colab T4 GPU. Training runtime: approximately 27-30 minutes for 3 epochs over 50 examples. Adapter size: approximately 55MB (LoRA adapter only, not the full ~16GB base model).
No separate held-out test set was used. The dataset is small (50 examples) and the goal was stylistic fine-tuning rather than benchmark performance, so evaluation here is training-time metrics only.
Not disaggregated by any factor — training metrics are reported as an overall average across the full training set.
Training loss (cross-entropy) and mean token accuracy, tracked during training, used to confirm the model was successfully learning the target output format and style.
Training loss: 0.71 → 0.55 → 0.62 (final, across 3 epochs). Mean token accuracy: reached 87% by the final epoch.
The model successfully learned to produce the target four-section verdict format and shows a clear, comedic, consistent judge persona, confirmed via manual generation testing after training.
No formal interpretability analysis was performed. Manual spot-checking of generations confirmed the model reliably follows the trained output format when prompted with the exact training template.
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Causal language model (Llama 3.1 architecture) with a LoRA adapter, trained with a next-token prediction (language modeling) objective, fine-tuned to produce a specific structured text output (verdict format) conditioned on code and courtroom argument context.
Google Colab, free tier.
Single NVIDIA T4 GPU (16GB VRAM)
transformers, peft, trl, bitsandbytes, accelerate, torch, datasets
No formal paper or citation exists for this project.
BibTeX:
None.
APA:
Reddy, J. (2026). llama3-8b-code-judge [Model]. Hugging Face. https://huggingface.co/jahnavi0803/llama3-8b-code-judge
Part of a larger project, AI Code Court — a three-LLM courtroom simulation where GPT-4o prosecutes code, Gemini defends it, and this model delivers the verdict.
Jahnavi Reddy