Downloads · 30 days
289
7% of all-time downloads
josephmayo/Holo-3.1-9B-Coder
Holo-3.1-9B-Coder is a text generation model from josephmayo. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Code-specialized adaptation of Hcompany/Holo-3.1-9B.
Downloads · 30 days
289
7% of all-time downloads
All-time downloads
4.4K
Public
Parameters
9B
55.3 GB on disk
Likes
4
Public
Click a slice to open those files.
.gguf19.5 GB · 52%
From the Hugging Face model README
Code-specialized adaptation of Hcompany/Holo-3.1-9B.
| Model | HumanEval+ pass@1 | LiveCodeBench v2 pass@1 |
|---|---|---|
| Holo-3.1-9B (base) | 52.4% | 31.5% |
| Holo-3.1-9B-Coder (this model) | 65.2% (+12.8) | 37.8% (+6.3) |
LiveCodeBench evaluated with official codegen_metrics, greedy decoding, 6s timeout. Proof: eval/lcb_v2_official.json
| Path | Description |
|---|---|
adapter/ | LoRA adapter (r=8, alpha=16, q/v targets). |
model.safetensors, config.json, tokenizer files | Merged base + adapter model. |
holo-9b-coder-Q4_K_M.gguf | llama.cpp GGUF, Q4_K_M quantization. |
holo-9b-coder-Q5_K_M.gguf | llama.cpp GGUF, Q5_K_M quantization. |
holo-9b-coder-Q6_K.gguf | llama.cpp GGUF, Q6_K quantization. |
Quantization was performed with llama.cpp after merging the adapter into the base model. The GGUF files use K-quant mixture schemes.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = "Hcompany/Holo-3.1-9B"
model = AutoModelForCausalLM.from_pretrained(base, trust_remote_code=True, torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, "josephmayo/Holo-3.1-9B-Coder", subfolder="adapter")
model = model.merge_and_unload()
tok = AutoTokenizer.from_pretrained(base, trust_remote_code=True)
./llama-cli -m holo-9b-coder-Q4_K_M.gguf -p "Return only executable Python code.\n\ndef factorial(n):" -n 256