Downloads · 30 days
0
stevensun2024/robotics-llm
robotics-llm is a machine learning model from stevensun2024. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A domain-specific Large Language Model fine-tuned for robotics and无人机 (UAV) engineering — combining Retrieval-Augmented Generation (RAG), speculative decoding, and a complete Teacher-Student distillation pipeline.
Downloads · 30 days
0
Access
Public
Updated Jul 5, 2026
Repo size
1 GB
Likes
1
Public
Click a slice to open those files.
.safetensors1 GB · 99%
From the Hugging Face model README
A domain-specific Large Language Model fine-tuned for robotics and无人机 (UAV) engineering — combining Retrieval-Augmented Generation (RAG), speculative decoding, and a complete Teacher-Student distillation pipeline.
Robotics-LLM is a Qwen3-14B model fine-tuned via RAFT (Retrieval-Augmented Fine-Tuning) distillation from a Qwen3-30B-A3B Teacher model. It is purpose-built for:
The system integrates a RAG Proxy Server (OpenAI-compatible API) with hybrid retrieval (BM25 + BGE semantic search) over a curated knowledge base of robotics documentation, academic papers, and engineering guides.
User (Cherry Studio / Open WebUI)
│
▼
┌─────────────────────────────┐
│ RAG Proxy (:5000) │ ← OpenAI-compatible API
│ ┌───────────────────────┐ │
│ │ Hybrid Searcher │ │ ← BM25 + BGE (CPU)
│ │ - BM25 (jieba) │ │
│ │ - BGE-small-zh-v1.5 │ │
│ └───────────────────────┘ │
│ │ │
│ ▼ │
│ ┌───────────────────────┐ │
│ │ vLLM Engine (:8000) │ │ ← 14B main + 1.7B draft
│ │ - tensor-parallel=2 │ │
│ │ - Speculative Decode │ │
│ └───────────────────────┘ │
└─────────────────────────────┘
| Component | Detail |
|---|---|
| Teacher Model | Qwen3-30B-A3B (Mixture of Experts) |
| Student Model | Qwen3-14B (dense) |
| Training Method | RAFT (Retrieval-Augmented Fine-Tuning) |
| Data Sources | 3 independently generated RAFT datasets |
| Cleaning | Deduplication, quality filtering, format normalization |
| Final Dataset | 19,101 high-quality 2-turn QA pairs |
| Parameter | Value |
|---|---|
| Base Model | Qwen3-14B |
| LoRA Rank | 64 |
| LoRA Alpha | 128 |
| LoRA Dropout | 0.05 |
| Target Modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Optimizer | AdamW (8-bit) |
| Learning Rate | 2e-4 (cosine scheduler) |
| Warmup Ratio | 0.03 |
| Max Sequence Length | 4096 |
| Per Device Batch Size | 2 |
| Gradient Accumulation | 4 |
| Effective Batch Size | 16 (2 GPUs × 2 × 4) |
| Training Epochs | 3 |
| Total Steps | 3,405 |
| Precision | bfloat16 |
| Hardware | 2 × H100 80GB |
After SFT, the LoRA adapter is merged into the base model using merge_and_unload() to create the final inference-ready model.
To accelerate generation, speculative decoding pairs the 14B main model with a Qwen3-1.7B draft model:
| Configuration | Value |
|---|---|
| Main Model | Qwen3-14B-SFT-Merged |
| Draft Model | Qwen3-1.7B |
| Speculative Tokens | 5 |
| Tensor Parallel | 2 |
| GPU Memory Utilization | 0.9 |
| Max Model Length | 8192 |
| Mode | Generation Speed |
|---|---|
| With speculative decoding (14B + 1.7B) | ~25 tok/s |
| Without speculative decoding (14B only) | ~60 tok/s |
Measured on 2 × H100 80GB with tensor-parallel-size=2
Note: With RAG retrieval overhead (hybrid search over 1080 documents), end-to-end latency increases by approximately 1–2 seconds per query. Actual throughput varies with prompt length and batch size.
| Source | Chunks |
|---|---|
| PX4 Developer Guide | 170 |
| ROS2 Development Guide | ~120 |
| Computer Vision Guide | ~80 |
| Robotics Academic Papers | ~200 |
| SLAM / Navigation Guides | ~150 |
| Academic Papers (Ego-Planner, VINS, etc.) | ~80 |
| Web Documentation (ExpressLRS, MuJoCo, Gazebo, Isaac Lab, RViz, RQT, ArduPilot, MicoAir) | ~280 |
| Total | 1,080 |
| Component | Parameter | Value |
|---|---|---|
| BM25 | Language | Chinese (jieba) |
| BGE | Model | bge-small-zh-v1.5 |
| BGE | Device | CUDA |
| Hybrid | Alpha (semantic weight) | 0.5 |
| Retrieval | Top-K | 3 |
The proxy exposes an OpenAI-compatible API at /v1/chat/completions, compatible with Cherry Studio, Open WebUI, and any OpenAI SDK client.
System prompt structure:
<think> analysis with critical literature reviewros2 topic echo, jtop, etc.)Monitored metrics (appended to each response):
git clone https://github.com/stevensun2024/robotics-llm
cd robotics-llm/rag_teacher
pip install -r requirements.txt
bash start_vllm.sh
bash start_proxy.sh
Configure your OpenAI-compatible client:
API Endpoint: http://localhost:5000/v1
API Key: sk-rag-prod
Model: rag-14b
robotics-llm/rag_teacher/
├── rag_proxy_server.py # OpenAI-compatible proxy server
├── hybrid_search.py # BM25 + BGE hybrid retriever
├── kb_loader.py # YAML knowledge base loader
├── load_advanced_kb.py # Multi-format knowledge base loader
├── config.py # Central configuration
├── run_lora_sft.py # LoRA SFT training script
├── generate_raft_data.py # RAFT data generation
├── generate_dpo_pairs.py # DPO pair generation
├── clean_raft_dataset.py # Data cleaning and deduplication
├── gpu_burn.py # GPU memory stress test
├── start_vllm.sh # vLLM startup script
├── start_proxy.sh # RAG Proxy startup script
├── deploy.sh # Deployment script
└── requirements.txt # Python dependencies
This project is released under the Apache License 2.0. The base model Qwen3-14B follows its original license terms.
If you use Robotics-LLM in your research, please cite:
@misc{robotics-llm-2026,
author = {Sun, Steven and Contributors},
title = {Robotics-LLM: Domain-Specific Language Model for Robotics Engineering},
year = {2026},
publisher = {Hugging Face / ModelScope},
url = {https://huggingface.co/stevensun2024/robotics-llm}
}