Downloads · 30 days
10
5% of all-time downloads
ndebuhr/Gemma-2-27B-Technical-Tutorial-Summarization-QLoRA
Gemma-2-27B-Technical-Tutorial-Summarization-QLoRA is a text generation model from ndebuhr. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
- Max Sequence Length: Trained at 16384 (via RoPE Scaling) - Data Type: Auto detection, with options for Float16 and Bfloat16 - Quantization: 4bit, to reduce memory usage
Downloads · 30 days
10
5% of all-time downloads
All-time downloads
222
Public
Parameters
27.2B
54.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors54.5 GB · 100%
From the Hugging Face model README
Used a private dataset with hundreds of technical tutorials and associated summaries.
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
import torch
input_text = ""
# Set device based on CUDA availability
device = "cuda" if torch.cuda.is_available() else "cpu"
# Load the model and tokenizer
model_name = "ndebuhr/Gemma-2-27B-Technical-Tutorial-Summarization-QLoRA"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name).to(device)
instruction = "Clarify and summarize this tutorial transcript"
prompt = """{}
### Raw Transcript:
{}
### Summary:
"""
# Tokenize the input text
inputs = tokenizer(
prompt.format(instruction, input_text),
return_tensors="pt",
truncation=True,
max_length=16384
).to(device)
# Generate outputs
outputs = model.generate(
**inputs,
max_length=16384,
num_return_sequences=1,
use_cache=True
)
# Decode the generated text
generated_text = tokenizer.batch_decode(outputs, skip_special_tokens=True)
This gemma2 model was trained 2x faster with Unsloth and Huggingface's TRL library.