Downloads · 30 days
18
10% of all-time downloads
hassanshka/Biomni-R0-32B-FP8
Biomni-R0-32B-FP8 is a text generation model from hassanshka. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This is an FP8 quantized version of Biomni-R0-32B-Preview, optimized for NVIDIA H100 and L40S hardware acceleration.
Downloads · 30 days
18
10% of all-time downloads
All-time downloads
182
Public
Parameters
32.8B
37.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors37.4 GB · 100%
How the weights are stored.
F8_E4M331.2B · 95%
From the Hugging Face model README
This is an FP8 quantized version of Biomni-R0-32B-Preview, optimized for NVIDIA H100 and L40S hardware acceleration.
| Parameter | Value |
|---|---|
| Scheme | FP8 (8-bit floating point) |
| Method | LLM Compressor QuantizationModifier |
| Calibration | Custom biomedical dataset |
| Hardware | Optimized for H100/L40S (FP8 Tensor Cores) |
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"hassanshka/Biomni-R0-32B-FP8",
device_map="auto",
torch_dtype="auto",
trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained("hassanshka/Biomni-R0-32B-FP8")
# Inference
messages = [{"role": "user", "content": "Your medical question here"}]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0]))
from llmcompressor.modifiers.quantization import QuantizationModifier
from llmcompressor import oneshot
recipe = QuantizationModifier(
targets="Linear",
scheme="FP8",
ignore=["lm_head"]
)
oneshot(
model=model,
dataset=calibration_data,
recipe=recipe,
max_seq_length=4096,
num_calibration_samples=len(calibration_data),
)
⚠️ Requires NVIDIA H100, L40S, or Ada Lovelace GPUs for optimal FP8 performance.
Apache 2.0 (same as base model)
If you use this model, please cite the original Biomni model.