Downloads · 30 days
0
cdreetz/audio-llama
audio-llama is a machine learning model from cdreetz. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<img src="https://cdn-lfs-us-1.hf.co/repos/64/30/6430a20288b1b07674f3ab60cdfa7cb9483852464dfbac93f3b96b3a002959f8/948b61345b8d64ce4edff3df13d4e3a6e373e58759f20cb34ce565c2b389d27e?response-content-disposition=inline%3B…
Downloads · 30 days
0
Access
Public
Updated Mar 18, 2025
Repo size
173 MB
Likes
0
Public
Click a slice to open those files.
.bin173 MB · 100%
From the Hugging Face model README
This is a PEFT (LoRA) adapter that needs to be combined with the base Llama model to work:
import torch
from peft import PeftModel, PeftConfig
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load the LoRA configuration
config = PeftConfig.from_pretrained("cdreetz/audio-llama")
# Load the base model
model = AutoModelForCausalLM.from_pretrained(
config.base_model_name_or_path,
torch_dtype=torch.float16,
device_map="auto"
)
# Load the tokenizer
tokenizer = AutoTokenizer.from_pretrained(config.base_model_name_or_path)
# Load the LoRA adapter
model = PeftModel.from_pretrained(model, "cdreetz/audio-llama")
# Run inference
prompt = "Transcribe this audio:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=100)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
This model was fine-tuned using LoRA on audio transcription tasks. It starts with a Llama 3 base model and uses Whisper-processed audio features for audio understanding.
This model requires special code for audio processing with Whisper before passing to the Llama model. See the Audio-LLaMA repository for full usage instructions.