Downloads · 30 days
21
4% of all-time downloads
shuff57/lfm2-24b-phase1-reasoning
lfm2-24b-phase1-reasoning is a text generation model from shuff57. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This is Phase 1 of the OGRE (OllamaGradingRubricEvaluator) fine-tuning pipeline.
Downloads · 30 days
21
4% of all-time downloads
All-time downloads
498
Public
Parameters
23.8B
47.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors47.7 GB · 100%
How the weights are stored.
BF1623.8B · 100%
From the Hugging Face model README
This is Phase 1 of the OGRE (OllamaGradingRubricEvaluator) fine-tuning pipeline.
LFM2-24B-A2B fine-tuned on 13,201 synthetic chain-of-thought reasoning examples to improve structured analytical reasoning before domain-specific grading fine-tuning in Phase 2.
| Field | Value |
|---|---|
| Base model | LiquidAI/LFM2-24B-A2B |
| Architecture | LFM2 MoE (Mixture of Experts), 24B total / ~2.3B active params per token |
| Fine-tune method | LoRA (rank 16, alpha 16) via Unsloth |
| Training data | shuff57/ogre-phase1-synth — 13,201 synthetic reasoning examples |
| Training steps | 1,568 steps (1 epoch) |
| Hardware | Google Colab A100 40GB + High-RAM (80GB system RAM) |
| Export format | Merged 16-bit safetensors |
| Phase | Phase 1 of 2 — reasoning pre-training before OGRE grading fine-tune |
# LoRA config
LORA_RANK = 16
LORA_ALPHA = 16
LORA_DROPOUT = 0.0
LORA_TARGET_MODULES = [
"q_proj", "k_proj", "v_proj",
"out_proj", "in_proj", # LFM2 attention projections
"w1", "w2", "w3", # MoE expert MLP layers
]
# SFTConfig
per_device_train_batch_size = 1
gradient_accumulation_steps = 8 # effective batch size = 8
num_train_epochs = 1
learning_rate = 2e-5
lr_scheduler_type = "cosine"
warmup_ratio = 0.1
optim = "adamw_8bit"
bf16 = True
packing = False # disabled for MoE routing stability
train_on_responses_onlyLFM2-24B-A2B requires specific package versions:
pip install unsloth_zoo unsloth
pip install transformers==5.3.0 # >=5.0.0 required for lfm2_moe arch
pip install trl==0.22.2 datasets==4.3.0
Important: device_map="cpu" required during loading to avoid VRAM OOM during MoE expert weight conversion. Use a runtime with ≥50GB system RAM.
from unsloth import FastLanguageModel
import torch
model, tokenizer = FastLanguageModel.from_pretrained(
model_name = "shuff57/lfm2-24b-phase1-reasoning",
max_seq_length = 8192,
load_in_4bit = True,
dtype = torch.bfloat16,
device_map = "cpu",
)
FastLanguageModel.for_inference(model)
messages = [
{"role": "system", "content": "You are a careful analytical reasoner."},
{"role": "user", "content": "Your question here."},
]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(
input_ids = inputs,
max_new_tokens = 1024,
temperature = 0.1,
top_k = 50,
repetition_penalty = 1.05,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
This model is Phase 1 of a two-phase fine-tuning pipeline:
| Phase | Model | Dataset | Purpose |
|---|---|---|---|
| Phase 1 | This model | 13,201 synthetic reasoning examples | Reasoning capability |
| Phase 2 | shuff57/lfm2-24b-grader (coming soon) | 239 OGRE grading examples | Domain-specific grading |
Phase 2 loads this model and fine-tunes further on OGRE statistics grading data, then exports as GGUF Q4_K_M for local Ollama deployment as lfm2-stat-grader.
Trained with Unsloth and HuggingFace TRL.