Downloads · 30 days
10
53% of all-time downloads
eswarankrishnamurthy/murli-assistant-distilgpt2-maximum
murli-assistant-distilgpt2-maximum is a machine learning model from eswarankrishnamurthy. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as mit.
⚠️ WARNING: EXPERIMENTAL MODEL - NOT FOR PRODUCTION USE ⚠️
Downloads · 30 days
10
53% of all-time downloads
All-time downloads
19
Public
Repo size
85.1 MB
Likes
0
Public
Click a slice to open those files.
.pt56.7 MB · 50%
From the Hugging Face model README
⚠️ WARNING: EXPERIMENTAL MODEL - NOT FOR PRODUCTION USE ⚠️
This model represents the absolute maximum possible training for DistilGPT-2 on murli content, but quality remains insufficient for spiritual guidance. Deployed for research, comparison, and educational purposes only.
Use Phi-2 (2.7B params) instead - proven quality for murli chatbot.
LoRA Configuration (MAXIMUM):
Training Data (MAXIMUM):
Training Configuration (MAXIMUM):
Final Training Loss: 1.609 (66% improvement over standard 4.77)
| Version | LoRA Rank | Epochs | Murlis | Loss | Quality |
|---|---|---|---|---|---|
| Standard | 4 | 3 | 150 | 4.77 | ❌ Poor |
| Enhanced | 16 | 10 | 300 | 2.07 | ❌ Poor |
| MAXIMUM | 32 | 15 | 500 | 1.61 | ❌ Still Poor |
Key Finding: Loss improvement does NOT guarantee quality improvement for small models in specialized domains.
✅ LoRA Rank: 32 (8x from standard, 2x from enhanced) ✅ LoRA Alpha: 64 (8x from standard, 2x from enhanced) ✅ Target Modules: c_attn + c_proj + c_fc (ALL layers) ✅ Epochs: 15 (5x from standard, 1.5x from enhanced) ✅ Murlis: 500 (3.3x from standard, 1.67x from enhanced) ✅ Context: 512 tokens (2x from standard, 1.33x from enhanced) ✅ 15 detailed spiritual concepts with full explanations ✅ 7 different formats per murli for comprehensive learning ✅ Ultra-careful learning rate (5e-5) ✅ Maximum warmup (200 steps) ✅ Larger effective batch (16) ✅ Stronger regularization (0.02 weight decay)
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
import torch
# Load base model
tokenizer = AutoTokenizer.from_pretrained("distilgpt2")
base_model = AutoModelForCausalLM.from_pretrained(
"distilgpt2",
torch_dtype=torch.float16,
device_map="auto"
)
# Load MAXIMUM adapter
model = PeftModel.from_pretrained(
base_model,
"eswarankrishnamurthy/murli-assistant-distilgpt2-maximum"
)
# Chat function
def chat(message):
prompt = f"Question: {message}\nAnswer:"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(
**inputs,
max_new_tokens=150,
temperature=0.7,
top_p=0.9,
top_k=50,
repetition_penalty=1.2,
no_repeat_ngram_size=3
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
return response.split("Answer:", 1)[1].strip() if "Answer:" in response else response
# Test (expect mixed quality)
print(chat("What is soul consciousness?"))
Inference Speed (CPU):
Resource Usage:
Compared to Production Models:
Architecture:
Training Process:
What Went Right:
What Went Wrong:
This model proves important insights for AI/ML research:
Before deploying any AI model:
Completed: 2025-10-03T12:25:52.051354
Loss Progression:
Gradient Norms: Stable (0.72 - 1.72)
Technical Success: ✅ Perfect training, lowest loss achieved
Functional Success: ❌ Quality insufficient for spiritual guidance
Research Value: ✅ Invaluable insights for model selection
For production murli chatbot, use Phi-2 fine-tuned on murli data.
This MAXIMUM model demonstrates that small models cannot reliably handle specialized spiritual domains, regardless of training optimization.
@misc{murli-distilgpt2-maximum,
author = {eswarankrishnamurthy},
title = {Murli Assistant - DistilGPT-2 MAXIMUM (Experimental)},
year = {2025},
publisher = {HuggingFace},
note = {Experimental model demonstrating small model limitations},
url = {https://huggingface.co/eswarankrishnamurthy/murli-assistant-distilgpt2-maximum}
}
For questions about this research or the production Phi-2 model, please open an issue.
This model is provided for research and educational purposes only.
For reliable murli assistance, consult:
Om Shanti! 🙏
Maximum training doesn't overcome fundamental capacity limits.
Sometimes you just need a bigger model.
Model Type: Experimental Research Model
Quality Rating: ⭐ (Insufficient for production)
Speed Rating: ⭐⭐⭐⭐⭐ (Excellent)
Recommended Alternative: Phi-2 (⭐⭐⭐⭐⭐ quality)