Downloads · 30 days
22
15% of all-time downloads
pshashid/llama3.1B_8B_SQL_Finetuned_model
llama3.1B_8B_SQL_Finetuned_model is a text generation model from pshashid. Use it when you need the model to write or continue text. It is set up for transformers.
SQL generation model fine-tuned on text-to-SQL tasks, quantized for NVIDIA Blackwell (RTX 50-series) using llm-compressor.
Downloads · 30 days
22
15% of all-time downloads
All-time downloads
145
Public
Parameters
5B
6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors6 GB · 100%
How the weights are stored.
U83.5B · 70%
From the Hugging Face model README
SQL generation model fine-tuned on text-to-SQL tasks, quantized for NVIDIA Blackwell (RTX 50-series) using llm-compressor.
| Component | Format | Notes |
|---|---|---|
| Weights | NVFP4 | ~4.5GB — Blackwell 5th-gen Tensor Core native |
| KV-Cache | FP8 | 50% memory vs FP16 — configured via vLLM |
| Activations | FP16 | lm_head kept in FP16 for output quality |
vllm serve pshashid/llama3.1B_8B_SQL_Finetuned_model \
--dtype float16 \
--quantization fp4 \
--kv-cache-dtype fp8 \
--max-model-len 131072 \
--gpu-memory-utilization 0.85 \
--enable-prefix-caching \
--port 8000
| Metric | Target |
|---|---|
| Time to First Token | < 15ms |
| Throughput (1 replica) | ~200 tok/s |
| Aggregate (8 replicas) | 1,500+ tok/s |
| Max Concurrency | 100+ users |
from vllm import LLM, SamplingParams
llm = LLM(
model = "pshashid/llama3.1B_8B_SQL_Finetuned_model",
quantization = "fp4",
kv_cache_dtype = "fp8",
max_model_len = 131072,
enable_prefix_caching = True,
)
sampling = SamplingParams(temperature=0, max_tokens=200)
outputs = llm.generate(["SELECT"], sampling)
print(outputs[0].outputs[0].text)