Downloads · 30 days
23
16% of all-time downloads
jyotirmoy05/VivekanandaAI
VivekanandaAI is a machine learning model from jyotirmoy05. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
An AI system embodying the wisdom and teachings of Swami Vivekananda, built with pure logic, no hardcoding, and fully config-driven architecture.
Downloads · 30 days
23
16% of all-time downloads
All-time downloads
143
Public
Repo size
4.4 GB
Likes
1
Public
Click a slice to open those files.
.gguf4.4 GB · 100%
From the Hugging Face model README
An AI system embodying the wisdom and teachings of Swami Vivekananda, built with pure logic, no hardcoding, and fully config-driven architecture.
The full implementation, scripts, and instructions to build the vector store and run inference (including RAG support) are available on GitHub: 👉 [GitHub: https://github.com/JyotirmoyDutta05/humanoid-AI] please copy the GitHub repository link and paster it on the search bar on your web browser to access the repository.
config.yaml# Clone repository
git clone <your-repo-url>
cd vivekananda-ai
# Run automated setup
chmod +x setup.sh
./setup.sh
# Add PDF files
cp /path/to/pdfs/* data/raw/complete_works/
# Add your JSON dataset
cp vivekananda_dataset_1.json data/processed/
# Add Mistral model
cp mistral-7b-instruct-v0.1.Q4_K_M.gguf models/base/
Edit config.yaml to customize:
source venv/bin/activate
python scripts/01_embed_data.py
Expected output:
======================================================================
VIVEKANANDA AI - VECTOR DATABASE CREATION
======================================================================
STEP 1: LOADING DOCUMENTS
Found 9 PDF files
Loading: SWAMI-VIVEKANANDA-VOL-1.pdf
Loaded 234 pages
...
STEP 2: PROCESSING WITH NLP
Processing documents with NLP...
...
SUCCESS! VECTOR DATABASE READY
Total chunks embedded: 8,469
python scripts/02_query_rag.py
streamlit run app.py
vivekananda-ai/
│
├── config.yaml # Central configuration
├── requirements.txt # Python dependencies
├── setup.sh # Automated setup script
│
├── utils.py # Core utilities
├── nlp_processor.py # NLP processing (spaCy + NLTK)
│
├── scripts/
│ ├── 01_embed_data.py # Create vector database
│ ├── 02_query_rag.py # Test RAG retrieval
│ ├── 03_test_mistral.py # Test Mistral model
│ └── 04_prepare_finetune.py # Prepare fine-tuning data
│
├── data/
│ ├── raw/ # Original PDFs
│ ├── processed/ # JSON datasets
│ └── extracted_text/ # Extracted text files
│
├── vectorstore/
│ └── vivekananda_db/ # FAISS vector database
│
├── models/
│ ├── base/ # Base models (GGUF)
│ └── fine_tuned/ # Fine-tuned adapters
│
├── outputs/
│ ├── logs/ # Application logs
│ └── results/ # Results & analysis
│
└── app.py # Streamlit app
All settings in config.yaml:
model:
generation:
max_tokens: 512
temperature: 0.7
top_p: 0.9
embeddings:
model_name: "sentence-transformers/all-mpnet-base-v2"
chunk:
size: 500
overlap: 50
nlp:
spacy:
model: "en_core_web_sm"
nltk:
tokenizer: "punkt"
preprocessing:
lemmatize: true
remove_stopwords: false
rag:
retrieval:
top_k: 5
similarity_threshold: 0.5
Your JSON dataset structure:
{
"instruction": "What is Karma Yoga?",
"response": "Karma Yoga is the path of selfless action...",
"source": "Karma-Yoga.pdf",
"work_type": "Karma-Yoga",
"topic": "philosophy"
}
Edit config.yaml:
embeddings:
model_name: "BAAI/bge-small-en-v1.5" # Faster alternative
embeddings:
chunk:
size: 300 # Smaller chunks
overlap: 30
nlp:
preprocessing:
lowercase: true
remove_stopwords: true
lemmatize: true
rag:
retrieval:
top_k: 10 # More context
similarity_threshold: 0.3 # Lower threshold
python scripts/01_embed_data.py
python scripts/02_query_rag.py "What is Karma Yoga?"
python scripts/03_test_mistral.py
streamlit run app.py
On Mac Mini M4:
| Task | Time | Memory |
|---|---|---|
| Embedding 9 volumes | 10-15 min | 4-6 GB |
| Query retrieval | 50-200 ms | 2-3 GB |
| Model inference | 1-3 sec | 6-8 GB |
python -m spacy download en_core_web_sm
import nltk
nltk.download('punkt')
nltk.download('stopwords')
Check config.yaml:
hardware:
device: "cpu" # Fallback to CPU
Reduce batch size in config.yaml:
embeddings:
batch_size: 16 # Smaller batches
MIT License - See LICENSE file
Built with for spreading Vivekananda's wisdom
For questions: Create an issue
This repository includes a production-ready pipeline to build a high-fidelity, persona-aligned AI system in the voice of Swami Vivekananda using a Mistral architecture.
Features
Repository Additions
training/data_preprocess.py: Convert processed dataset to JSONL instruction formattraining/sft_train.py: SFT training with LoRA/QLoRAtraining/dpo_train.py: DPO preference optimizationtraining/export_adapter.py: (optional) merge/unload adapters – to be added if neededinference/generate.py: Persona inference with sampling controls and token probabilitiesrag/build_index.py, rag/retrieve.py: Optional retrieval augmentationconfigs/train_sft.yaml, configs/train_dpo.yaml: Training configsprompts/system_vivekananda.txt: Persona system promptdocker/Dockerfile: GPU training environmentQuick Start
Preprocess dataset to JSONL
python training/data_preprocess.py --source data/processed/vivekananda_dataset_1.json --out data/datasets/sft_train.jsonl --val data/datasets/sft_val.jsonl --val-ratio 0.05SFT training (LoRA/QLoRA)
accelerate launch training/sft_train.py --config configs/train_sft.yaml --train data/datasets/sft_train.jsonl --val data/datasets/sft_val.jsonl --output runs/sft_vivekanandaDPO preference optimization
{prompt, chosen, rejected}accelerate launch training/dpo_train.py --config configs/train_dpo.yaml --prefs data/datasets/dpo_prefs.jsonl --sft-checkpoint runs/sft_vivekananda/step_XXXX --output runs/dpo_vivekanandaOptional RAG
python rag/build_index.py --source data/processed/vivekananda_dataset_1.json --out vectorstore/faiss_indexpython rag/retrieve.py --index vectorstore/faiss_index --query "Your question"--rag-contextInference with persona and probabilistic decoding
python inference/generate.py --model mistralai/Mistral-7B-Instruct-v0.2 --adapter runs/sft_vivekananda/step_XXXX --prompt "Question" --temperature 0.3 --top_p 0.85 --top_k 40 --max_tokens 512 --rag-context "<CONTEXT>..."Docker (GPU Training)
docker build -t vivekananda-ai:train -f docker/Dockerfile .docker run --gpus all -it -v $(pwd):/workspace vivekananda-ai:train bashNotes
bitsandbytes detects GPU; use CUDA-compatible base.prompts/system_vivekananda.txt.