Downloads · 30 days
0
KBlueLeaf/Stable-Cascade-FP16-fixed
Stable-Cascade-FP16-fixed is a machine learning model from KBlueLeaf. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-4.0.
A modified version of Stable-Cascade which is compatibile with fp16 inference
Downloads · 30 days
0
Access
Public
Updated Feb 19, 2024
Repo size
21.5 GB
Likes
43
Public
Click a slice to open those files.
.safetensors21.5 GB · 100%
From the Hugging Face model README
A modified version of Stable-Cascade which is compatibile with fp16 inference
In theory, you don't need to actually download this model file. It is possible to do onfly modification. This model is for experiments.
| FP16 | BF16 |
|---|---|
![]() | ![]() |
LPIPS difference: 0.088
| FP16 | BF16 |
|---|---|
![]() | ![]() |
LPIPS difference: 0.012
After doing some check to the L1 norm of each hidden state. I found the last block group(8, 24, 24, 8 <- this one) make the hiddens states become bigger and bigger.
So I just apply some transformation on the TimestepBlock to directly modify the scale of hidden state. (Since it is not a residual block, so this is possible)
How the transformation be done is written in the modified "stable_cascade.py", you can put the file into kohya-ss/sd-scripts' stable-cascade branch and uncomment things to check weights or doing the conversion by yourselve.
Some people may know the FP8 quant for inference SDXL with lowvram cards. The technique can be applied to this model too.<br> But since the last block group is basically ruined, so it is recommend to ignore the last block group:<br>
for name, module in generator_c.named_modules():
if "up_blocks.1" in name: continue
if isinstance(module, torch.nn.Linear):
module.to(torch.float8_e5m2)
elif isinstance(module, torch.nn.Conv2d):
module.to(torch.float8_e5m2)
elif isinstance(module, torch.nn.MultiheadAttention):
module.to(torch.float8_e5m2)
This sample code should transform 70% of weight into fp8. (Use FP8 weight with scale is better solution, it is recommended to implement that)
I have tried different transform settings which is more friendly for FP8 but the differences between original model is more significant.
FP8 Demo (Same Seed):

The modified version of model will not be compatibile with the lora/lycoris trained on original weight. <br> (actually it can, just do the same transformation, I'm considering to rewrite a version to use key name to determine what to do.)
Also the ControlNets will not be compatible too. Unless you also apply the needed transformation to them.
I don't want to do all of these by myself so hope some others will do that.
Stable-Cascade is published with a non-commercial lisence so I use CC-BY-NC 4.0 to publish this model. The source code to make this model is published with apache-2.0 license