Downloads · 30 days
5
33% of all-time downloads
LenSch/torchao
torchao is a text-to-image model from LenSch. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as other.
This model was created using the pruna library. Pruna is a model optimization framework built for developers, enabling you to deliver more efficient models with minimal implementation overhead.
Downloads · 30 days
5
33% of all-time downloads
All-time downloads
15
Public
Parameters
11.9B
33.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors33.7 GB · 100%
From the Hugging Face model README
This model was created using the pruna library. Pruna is a model optimization framework built for developers, enabling you to deliver more efficient models with minimal implementation overhead.
First things first, you need to install the pruna library:
pip install pruna
You can use the library_name library to load the model but this might not include all optimizations by default.
To ensure that all optimizations are applied, use the pruna library to load the model using the following code:
from pruna import PrunaModel
loaded_model = PrunaModel.from_pretrained(
"LenSch/torchao"
)
# we can then run inference using the methods supported by the base model
Alternatively, you can visit the Pruna documentation for more information.
The compression configuration of the model is stored in the smash_config.json file, which describes the optimization methods that were applied to the model.
{
"batcher": null,
"cacher": null,
"compiler": "torch_compile",
"factorizer": null,
"kernel": "flash_attn3",
"pruner": null,
"quantizer": "torchao",
"torch_compile_backend": "inductor",
"torch_compile_dynamic": null,
"torch_compile_fullgraph": false,
"torch_compile_make_portable": false,
"torch_compile_max_kv_cache_size": 400,
"torch_compile_mode": "max-autotune-no-cudagraphs",
"torch_compile_seqlen_manual_cuda_graph": 100,
"torch_compile_target": "model",
"torchao_excluded_modules": "none",
"torchao_quant_type": "int8dq",
"batch_size": 1,
"device": "cuda:0",
"device_map": null,
"save_fns": [
"save_before_apply",
"save_before_apply"
],
"load_fns": [
"diffusers"
],
"reapply_after_load": {
"factorizer": null,
"pruner": null,
"quantizer": "torchao",
"kernel": "flash_attn3",
"cacher": null,
"compiler": "torch_compile",
"batcher": null
}
}