Downloads · 30 days
13
19% of all-time downloads
rabimba/gemma2racer
gemma2racer is a text generation model from rabimba. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as gemma.
gemma2racer is a specialized optimization of Google's Gemma 2 architecture. This model is fine-tuned and configured specifically for "racing" performance—prioritizing high-speed token generation and low-memory overhea…
Downloads · 30 days
13
19% of all-time downloads
All-time downloads
68
Public
Parameters
5B
10 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.9 GB · 100%
From the Hugging Face model README
gemma2racer is a specialized optimization of Google's Gemma 2 architecture. This model is fine-tuned and configured specifically for "racing" performance—prioritizing high-speed token generation and low-memory overhead for local LLM deployment.
The following table outlines the core technical specifications for the Gemma-2-Racer model.
| Feature | Details |
|---|---|
| Developed by | Rabimba Karanjai |
| Model Type | Causal Language Model (Transformer-based) |
| Base Model | google/gemma-2-2b |
| Architecture | Gemma-2 |
| Optimization Strategy | 4-bit Quantization, torch.compile, and BitsAndBytes |
| Primary Language | English |
| License | Gemma Terms of Use |
This model is designed for developers and researchers who require state-of-the-art performance on consumer-grade hardware. It is specifically optimized for:
To get the model running with the "Racer" performance presets, follow these steps:
Install Requirements: Update your environment with the necessary libraries for quantization and acceleration.
pip install -U transformers accelerate bitsandbytes
Login to Hugging Face: Ensure you have accepted the Gemma license on the official Google repository and authenticate locally.
huggingface-cli login
Python Implementation: Use the following code snippet to load the model in its optimized 4-bit state.
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "rabimba/gemma2racer"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
load_in_4bit=True,
torch_dtype=torch.bfloat16
)
prompt = "Explain quantum physics like I'm a race car driver."
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=150)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
The "Racer" moniker refers to the model's ability to be tuned for different hardware constraints:
model = torch.compile(model) to utilize kernel fusion for significantly higher throughput.accelerate for systems with limited or no dedicated graphics memory.If you use this model in your research or commercial projects, please cite it as follows:
@misc{gemma2racer2024,
author = {Rabimba Karanjai},
title = {Gemma-2-Racer: Optimized Local Inference},
year = {2024},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/rabimba/gemma2racer}}
}