Downloads · 30 days
31
0% of all-time downloads
Cheng98/llama-160m
llama-160m is a text generation model from Cheng98. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A toy Llama adapted from JackFram/llama-160m with special tokens added.
Downloads · 30 days
31
0% of all-time downloads
All-time downloads
23.5K
Public
Repo size
1.3 GB
Likes
0
Public
Click a slice to open those files.
.bin650 MB · 100%
From the Hugging Face model README
A toy Llama adapted from JackFram/llama-160m with special tokens added.
This checkpoint can be loaded into MASE's LlamaQuantized
from transformers.models.llama import LlamaTokenizer
from chop.models.manual.llama_quantized import (
LlamaQuantizedConfig,
LlamaQuantizedForCausalLM,
)
name="Cheng98/llama-160m"
tokenizer = LlamaTokenizer.from_pretrained(name)
# override the quant_config to quantized the model
# default does not quantize llama
config = LlamaQuantizedConfig.from_pretrained(
name,
# quant_config="./quant_config_na.toml"
)
llama = LlamaQuantizedForCausalLM.from_pretrained(
name,
config=config,
)