Downloads ยท 30 days
0
Vedisasi/UltraThinking-LLM-Training
UltraThinking-LLM-Training is a machine learning model from Vedisasi. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
<p align="center" <img src="docs/images/pp.jpg" alt="ULTRATHINK Logo" width="250" / </p
Downloads ยท 30 days
0
Access
Public
Updated Oct 14, 2025
Repo size
1.7 MB
Likes
1
Public
Click a slice to open those files.
.png1.7 MB ยท 65%
From the Hugging Face model README
ULTRATHINK provides a complete, modular stack for training custom LLMs with state-of-the-art architectures, distributed training, and comprehensive monitoring.
Train state-of-the-art LLMs in 10 lines of code - From prototype to production in minutes, not days.
python train_ultrathink.py \
--dataset c4 --streaming \
--hidden_size 768 --num_layers 12 \
--enable_moe --enable_dre \
--use_amp --gradient_checkpointing
| Feature | ULTRATHINK | Others |
|---|---|---|
| Setup Time | โก 5 minutes | 30-120 minutes |
| Lines to Train | ๐ ~10 | 50-100+ |
| MoE Support | โ Native | โ or Limited |
| Dynamic Reasoning | โ Unique | โ None |
| Constitutional AI | โ Built-in | โ None |
| Documentation | ๐ Comprehensive | Varies |
View benchmarks and performance metrics โ
# Clone repository
git clone https://github.com/vediyappanm/UltraThinking-LLM-Training.git
cd UltraThinking-LLM-Training/deep
# Install dependencies
pip install -r requirements.txt
Tiny Model (CPU-friendly, for testing):
python train_ultrathink.py \
--dataset wikitext \
--hidden_size 256 --num_layers 2 --num_heads 4 \
--batch_size 2 --max_samples 1000 \
--num_epochs 1
Small Model (GPU recommended):
python train_advanced.py --config configs/train_small.yaml
With Advanced Features:
python train_ultrathink.py \
--dataset c4 --streaming \
--hidden_size 768 --num_layers 12 --num_heads 12 \
--enable_moe --enable_dre --enable_constitutional \
--use_amp --gradient_checkpointing \
--use_mlflow
# Run Gradio web interface
docker compose up
# Or build and run manually
docker build -t ultrathink:latest .
docker run -p 7860:7860 ultrathink:latest
# Run all tests
pytest
# Run with coverage
pytest --cov=src --cov-report=html
# Quick smoke test
python tests/smoke_test.py
deep/
โโโ train_ultrathink.py # Main training script
โโโ train_advanced.py # YAML config-based training
โโโ app_gradio.py # Web UI for inference
โโโ src/
โ โโโ models/ # UltraThink, MoE, DRE, architecture
โ โโโ data/ # Datasets, tokenization, validation
โ โโโ training/ # Optimizers, distributed, RLHF
โ โโโ monitoring/ # Metrics and system monitoring
โ โโโ security/ # Input validation and safety
โ โโโ evaluation/ # Benchmarks and metrics
โโโ tests/ # Unit and integration tests
โโโ configs/ # YAML configuration files
โโโ scripts/ # Utilities (profiling, inference)
โโโ docs/ # Documentation and guides
See PROJECT_STRUCTURE.md for detailed explanations.
# WikiText-2 (fast iteration)
python train_ultrathink.py \
--dataset wikitext \
--hidden_size 512 --num_layers 6 --num_heads 8 \
--batch_size 4 --num_epochs 3 \
--use_mlflow
# Streaming C4 with all optimizations
python train_ultrathink.py \
--dataset c4 --dataset_subset en --streaming \
--hidden_size 768 --num_layers 12 --num_heads 12 \
--batch_size 2 --gradient_accumulation_steps 64 \
--learning_rate 3e-4 --warmup_steps 5000 \
--use_amp --gradient_checkpointing \
--max_seq_length 1024 \
--output_dir ./outputs/c4_production
# Small model (4-8GB GPU)
python train_advanced.py --config configs/train_small.yaml
# Medium model (16-32GB GPU)
python train_advanced.py --config configs/train_medium.yaml
# Large model (40GB+ GPU)
python train_advanced.py --config configs/train_large.yaml
Web Interface (Gradio):
docker compose up
# Visit http://localhost:7860
Custom Training:
docker run -v $(pwd)/outputs:/app/outputs ultrathink:latest \
python train_ultrathink.py \
--dataset wikitext \
--hidden_size 256 --num_layers 2 \
--output_dir /app/outputs/my_model
GPU Training:
docker run --gpus all \
-v $(pwd)/outputs:/app/outputs \
ultrathink:latest \
python train_ultrathink.py --use_amp
We welcome contributions! Please see:
If you find ULTRATHINK useful, please consider giving us a star! โญ
| Size | Parameters | Layers | Hidden | Context | Min GPU |
|---|---|---|---|---|---|
| Tiny | 125M | 12 | 768 | 2048 | 6GB |
| Small | 350M | 24 | 1024 | 4096 | 16GB |
| Medium | 760M | 24 | 1536 | 4096 | 24GB |
| Large | 1.3B | 32 | 2048 | 8192 | 40GB |
See MODEL_CARD.md for complete specifications.
MIT License - see LICENSE for details.
If you use ULTRATHINK in your research or project, please cite:
@software{ultrathink2025,
title={ULTRATHINK: Advanced LLM Training Framework with Mixture-of-Experts and Dynamic Reasoning},
author={ULTRATHINK Team},
year={2025},
url={https://github.com/vediyappanm/UltraThinking-LLM-Training},
version={1.0.0}
}
Built something cool with ULTRATHINK? We'd love to hear about it!