Downloads · 30 days
21
4% of all-time downloads
Inferless/deciLM-7B-GPTQ
deciLM-7B-GPTQ is a text generation model from Inferless. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
<div style="width: auto; margin-left: auto; margin-right: auto" <img src="https://pbs.twimg.com/profilebanners/1633782755669708804/1678359514/1500x500" alt="Inferless" style="width: 100%; min-width: 400px; display: bl…
Downloads · 30 days
21
4% of all-time downloads
All-time downloads
489
Public
Repo size
4.1 GB
Likes
1
Public
Click a slice to open those files.
.safetensors4.1 GB · 100%
From the Hugging Face model README
This repo contains GPTQ model files for Deci's DeciLM-7B.
GPTQ is a method that compresses the model size and accelerates inference by quantizing weights based on a calibration dataset, aiming to minimize mean squared error in a single post-quantization step. GPTQ achieves both memory efficiency and faster inference.
It is supported by:
Models are released as sharded safetensors files.
| Branch | Bits | GS | AWQ Dataset | Seq Len | Size |
|---|---|---|---|---|---|
| main | 4 | 128 | VMware Open Instruct | 4096 | 5.96 GB |
You will need the following software packages and python libraries:
build:
cuda_version: "12.1.1"
system_packages:
- "libssl-dev"
python_packages:
- "torch==2.1.2"
- "vllm==0.2.6"
- "transformers==4.36.2"
- "accelerate==0.25.0"
Here is the code for <b>app.py</b>
from vllm import LLM, SamplingParams
class InferlessPythonModel:
def initialize(self):
self.sampling_params = SamplingParams(temperature=0.7, top_p=0.95,max_tokens=256)
self.llm = LLM(model="Inferless/deciLM-7B-GPTQ", quantization="gptq", dtype="float16")
def infer(self, inputs):
prompts = inputs["prompt"]
result = self.llm.generate(prompts, self.sampling_params)
result_output = [[[output.outputs[0].text,output.outputs[0].token_ids] for output in result]
return {'generated_result': result_output[0]}
def finalize(self):
pass