Downloads · 30 days
170
0% of all-time downloads
pyrymikko/nomic-embed-code-W4A16-AWQ
nomic-embed-code-W4A16-AWQ is a machine learning model from pyrymikko. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This is a W4A16 quantized version of nomic-ai/nomic-embed-code.
Downloads · 30 days
170
0% of all-time downloads
All-time downloads
394K
Public
Parameters
7.1B
4.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.5 GB · 100%
How the weights are stored.
I326.5B · 92%
From the Hugging Face model README
This is a W4A16 quantized version of nomic-ai/nomic-embed-code.
Quantized using AWQ (Activation-aware Weight Quantization) with llm-compressor!
from transformers import AutoModel, AutoTokenizer
# Load quantized model
model = AutoModel.from_pretrained(
"nomic-embed-code-W4A16-AWQ",
trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(
"nomic-embed-code-W4A16-AWQ",
trust_remote_code=True
)
# Generate embeddings
texts = ["Hello world", "Example text"]
inputs = tokenizer(texts, padding=True, return_tensors="pt")
embeddings = model(**inputs).last_hidden_state.mean(dim=1)
print(embeddings.shape)
AWQ (Activation-aware Weight Quantization) is a one-shot weight quantization method that:
This quantized model is based on nomic-ai/nomic-embed-code.
If you use this model, please cite the original model and llmcompressor:
@software{llmcompressor,
title = {LLM Compressor},
author = {Neural Magic},
url = {https://github.com/vllm-project/llm-compressor},
year = {2024}
}