Downloads · 30 days
29
7% of all-time downloads
aliRafik/glm_flash_finetuned_16bit
glm_flash_finetuned_16bit is a text generation model from aliRafik. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
- Developed by: aliRafik - License: apache-2.0 - Finetuned from model : unsloth/GLM-4.7-Flash
Downloads · 30 days
29
7% of all-time downloads
All-time downloads
419
Public
Parameters
31.2B
62.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors62.4 GB · 100%
How the weights are stored.
BF1631.2B · 100%
From the Hugging Face model README
This glm4_moe_lite model was trained 2x faster with Unsloth and Huggingface's TRL library.
from transformers import AutoModelForCausalLM, AutoTokenizer import torch
model = AutoModelForCausalLM.from_pretrained( "aliRafik/glm_flash_finetuned_16bit", torch_dtype=torch.float16, device_map="auto" )
tokenizer = AutoTokenizer.from_pretrained("aliRafik/glm_flash_finetuned_16bit")
inputs = tokenizer("Hello, how are you?", return_tensors="pt").to("cuda") outputs = model.generate(**inputs, max_new_tokens=100) print(tokenizer.decode(outputs[0]))
