Downloads · 30 days
57
100% of all-time downloads
p-p-n/Huvm
Huvm is a text generation model from p-p-n. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Huvm is a compact, instruction-tuned language model fine-tuned from Qwen/Qwen2.5-3B-Instruct using LoRA + DPO. It is designed to be fast, precise, and carry a distinct personality — a blend of 50% Grok (direct, lightl…
Downloads · 30 days
57
100% of all-time downloads
All-time downloads
57
Public
Repo size
1.9 GB
Likes
1
Public
Click a slice to open those files.
.gguf1.9 GB · 100%
From the Hugging Face model README
Huvm is a compact, instruction-tuned language model fine-tuned from Qwen/Qwen2.5-3B-Instruct using LoRA + DPO. It is designed to be fast, precise, and carry a distinct personality — a blend of 50% Grok (direct, lightly sarcastic, no-nonsense) and 50% Claude (articulate, deep, careful).
The name comes from "humm" (the sound of thinking) plus the letter V for speed and truth.
f16, q4_k_m)Huvm is built for local, offline use — desktop CPU or mobile via GGUF. It handles:
It is not intended for high-stakes, medical, legal, or safety-critical applications, nor for real-time news, statistics, or link retrieval.
| Trait | Behavior |
|---|---|
| Directness | Gets to the point, minimal filler |
| Sarcasm | Light and playful, never insulting |
| Honesty | Refuses to invent facts, links, or stats |
| Language | Replies in the user's language (pt/en/es) |
Important limitations — read before use.
| File | Size | Note |
|---|---|---|
huvm-q4_k_m.gguf | ~1.93 GB | Recommended — balanced quality/space |
Run with llama.cpp:
llama-cli -m huvm-q4_k_m.gguf \
-p "<|im_start|>system\nEu sou o Huvm...<|im_end|>\n<|im_start|>user\nQuem e voce?<|im_end|>\n<|im_start|>assistant\n" \
-n 200 --temp 0.0
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("your-username/huvm")
tokenizer = AutoTokenizer.from_pretrained("your-username/huvm")
SYSTEM = "Eu sou o Huvm, um assistente de IA pessoal, rapido, preciso e com personalidade. Metade Grok (direto, sarcasmo leve e bem-humorado), metade Claude (articulado, profundo, cuidadoso). Data de referencia: 3 de setembro de 2026. Respondo no idioma do usuario. No codigo eu explico o por que; na matematica mostro o raciocinio; na criatividade fujo do generico. NUNCA invento fatos, fontes, links, citacoes, estatisticas ou noticias. Respondo SEMPRE diretamente, sem bloco de raciocinio, seja objetivo."
messages = [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": "Faca uma funcao em Python que inverte uma string"},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(out[0], skip_special_tokens=True))
Built on Qwen 2.5 (Alibaba), TRL (Hugging Face), and llama.cpp.