Downloads · 30 days
0
atulkrs/opt-mlops-merged
opt-mlops-merged is a machine learning model from atulkrs. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as mit.
facebook/opt-125m fine-tuned with a LoRA adapter (atulkrs/opt-mlops-lora) and fully merged into base weights via PeftModel.mergeandunload().
Downloads · 30 days
0
Access
Public
Updated Jun 14, 2026
Parameters
125M
501 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors501 MB · 99%
From the Hugging Face model README
facebook/opt-125m fine-tuned with a LoRA adapter
(atulkrs/opt-mlops-lora)
and fully merged into base weights via PeftModel.merge_and_unload().
The adapter deltas are baked in — no PEFT dependency needed at inference time.
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("atulkrs/opt-mlops-merged")
tokenizer = AutoTokenizer.from_pretrained("atulkrs/opt-mlops-merged")
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4", # NormalFloat4 — from QLoRA paper
bnb_4bit_compute_dtype=torch.float16,
)
model = AutoModelForCausalLM.from_pretrained(
"atulkrs/opt-mlops-merged",
quantization_config=bnb_config,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("atulkrs/opt-mlops-merged")
Tip: swap
torch.float16fortorch.bfloat16on Ampere+ GPUs (A100, RTX 30xx+) for better numerical stability at no speed cost.
| Format | Size | Notes |
|---|---|---|
| FP32 (merged) | 477.8 MB | measured via param_size_mb() |
| 4-bit NF4 (est.) | 59.7 MB | approx fp32 / 8 |
| Reduction | ~8x |
4-bit load time benchmark requires Linux + CUDA + bitsandbytes; estimated load time on GPU is typically 2–5s for a 125M model.
| Field | Value |
|---|---|
| Base model | facebook/opt-125m |
| Adapter | atulkrs/opt-mlops-lora |
| Merge method | PeftModel.merge_and_unload() |
| Saved format | PyTorch bin (fp32) |
Merging removes the adapter overhead entirely — no extra matrix multiplications at
inference, no PEFT dependency, and the weights load like any standard
transformers checkpoint. The only trade-off is that you can no longer swap
adapters without re-loading the base model.