Downloads · 30 days
15
32% of all-time downloads
emmaoba/davanai
davanai is a text generation model from emmaoba. Use it when you need the model to write or continue text. It is set up for peft.
A LoRA adapter fine-tuned on Qwen/Qwen2.5-32B, developed by Vega Aiden Lab.
Downloads · 30 days
15
32% of all-time downloads
All-time downloads
47
Public
Repo size
7.3 GB
Likes
0
Public
Click a slice to open those files.
.safetensors7.3 GB · 100%
From the Hugging Face model README
A LoRA adapter fine-tuned on Qwen/Qwen2.5-32B, developed by Vega Aiden Lab.
davanai 1 was trained across 37 datasets.
This is a LoRA adapter, not a full model. To run it you download the base model Qwen/Qwen2.5-32B and apply this adapter on top. The base model is large, so make sure your machine can handle it (see hardware note below).
pip install torch transformers peft accelerate safetensors
pip install -U "huggingface_hub[cli]"
hf auth login
Paste a token from https://huggingface.co/settings/tokens when asked.
run.py and paste this infrom transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = "Qwen/Qwen2.5-32B"
adapter = "emmaoba/davanai"
# Load tokenizer and base model
tokenizer = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(
base,
device_map="auto", # uses your GPU automatically
torch_dtype="auto",
)
# Apply the davanai 1 adapter
model = PeftModel.from_pretrained(model, adapter)
model.eval()
# Chat with it
messages = [{"role": "user", "content": "Hello! Who are you?"}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256, do_sample=True, temperature=0.7)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
python run.py
The first run downloads the base model (this is large and happens only once — it's cached for next time), then prints davanai 1's reply.
Qwen2.5-32B is a 32-billion-parameter model. To run it in full precision you need a lot of GPU memory (roughly 64+ GB). If your GPU is smaller, load it in 4-bit instead — install bitsandbytes (pip install bitsandbytes) and change the model loading line to:
model = AutoModelForCausalLM.from_pretrained(
base,
device_map="auto",
load_in_4bit=True,
)
This lets it run on a single 24 GB GPU (like an RTX 3090/4090).