Downloads · 30 days
9
6% of all-time downloads
waqasm86/llcuda-models
llcuda-models is a text generation model from waqasm86. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Optimized GGUF models for llcuda - Zero-config CUDA-accelerated LLM inference.
Downloads · 30 days
9
6% of all-time downloads
All-time downloads
148
Public
Repo size
806 MB
Likes
0
Public
Click a slice to open those files.
.gguf806 MB · 100%
From the Hugging Face model README
Optimized GGUF models for llcuda - Zero-config CUDA-accelerated LLM inference.
Performance:
pip install llcuda
import llcuda
engine = llcuda.InferenceEngine()
engine.load_model("gemma-3-1b-Q4_K_M")
result = engine.infer("What is AI?")
print(result.text)
# Download model
huggingface-cli download waqasm86/llcuda-models google_gemma-3-1b-it-Q4_K_M.gguf --local-dir ./models
# Run with llama.cpp
./llama-server -m ./models/google_gemma-3-1b-it-Q4_K_M.gguf -ngl 26
Apache 2.0 - Models are provided as-is for educational and research purposes.