Downloads · 30 days
19
8% of all-time downloads
BallisticAI/Ballistic-CodeLlama-34B-v1-AWQ
Ballistic-CodeLlama-34B-v1-AWQ is a text generation model from BallisticAI. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama2.
- Model creator: BallisticAI - Based on: CodeLlama 34B hf - Merged with: CodeLlama 34B v2 && speechless-codellama-34b-v2
Downloads · 30 days
19
8% of all-time downloads
All-time downloads
224
Public
Repo size
18.3 GB
Likes
0
Public
Click a slice to open those files.
.safetensors18.3 GB · 100%
From the Hugging Face model README
This repo contains GGUF format model files for Ballistic-CodeLlama-34B-v1.
<!-- description end -->AWQ is an efficient, accurate and blazing-fast low-bit weight quantization method, currently supporting 4-bit quantization. Compared to GPTQ, it offers faster Transformers-based inference.
It is also now supported by continuous batching server vLLM, allowing use of AWQ models for high-throughput concurrent inference in multi-user server scenarios. Note that, at the time of writing, overall throughput is still lower than running vLLM with unquantised models, however using AWQ enables using much smaller GPUs which can lead to easier deployment and overall cost savings. For example, a 70B model can be run on 1 x 48GB GPU instead of 2 x 80GB.
<!-- description end --> <!-- repositories-available start -->This model accepts the Alpaca/Vicuna instruction format.
For example:
### System Prompt
You are an intelligent programming assistant.
### User Message
Implement a linked list in C++
### Assistant
...
<!-- prompt-template end -->
This model has undergone very limited testing. Additional safety testing should be performed before any real-world deployments.
Thanks to: