Downloads · 30 days
0
ohmygaugh/agent-mcp-sql
agent-mcp-sql is a machine learning model from ohmygaugh. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This project implements an intelligent, multi-step GraphRAG-powered agent that uses LangChain to orchestrate complex queries against a federated life sciences dataset. The agent leverages a Neo4j graph database to und…
Downloads · 30 days
0
Access
Public
Updated Oct 8, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.py86.1 KB · 47%
From the Hugging Face model README
This project implements an intelligent, multi-step GraphRAG-powered agent that uses LangChain to orchestrate complex queries against a federated life sciences dataset. The agent leverages a Neo4j graph database to understand the relationships between disparate SQLite databases, constructs SQL queries, and returns unified results through a conversational UI.
🤖 LangChain Agent: Orchestrates tools for schema discovery, pathfinding, and query execution.
🕸️ GraphRAG Enabled: Uses a Neo4j knowledge graph of database schemas for intelligent query planning.
🔬 Life Sciences Dataset: Comes with a rich dataset across clinical trials, drug discovery, and lab results.
conversational Conversational UI: A Streamlit-based chat interface for interacting with the agent.
🔌 RESTful MCP Server: All core logic is exposed via a secure and scalable FastAPI server.
┌─────────────────┐ ┌───────────────┐ ┌─────────────────┐
│ Streamlit Chat │──────│ Agent │ │ MCP Server │
│ (UI) │ │ (LangChain) │ │ (FastAPI) │
└─────────────────┘ └───────────────┘ └─────────────────┘
│
┌───────────────────────┼───────────────────────┐
│ │ │
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Neo4j │ │ clinical_ │ │ laboratory │
│ (Schema KG) │ │ trials.db │ │ .db │
└─────────────┘ └─────────────┘ └─────────────┘
│
┌─────────────┐
│ drug_ │
│ discovery.db│
└─────────────┘
clinical_trials.db, drug_discovery.db, laboratory.db) that serve as the federated data sources.Clone and configure:
git clone <repository-url>
cd <repository-name>
touch .env
Add your LLM API key to the .env file.
LLM_API_KEY="sk-your-llm-api-key-here"
Start the system:
make up
Seed the databases and ingest schema:
make seed-db
make ingest
Open the interface:
Once the system is running, open the Streamlit UI and ask a question about the life sciences data, for example:
The agent will then:
SchemaSearchTool to find relevant tables.JoinPathFinderTool to determine how to join them.QueryExecutorTool.To test the agent's logic directly without the full Docker stack, you can run it from your terminal.
Set up the environment:
Make sure the MCP and Neo4j services are running (make up).
Create a Python virtual environment and install dependencies:
python -m venv venv
source venv/bin/activate
pip install -r agent/requirements.txt
Set your API key:
export LLM_API_KEY="sk-your-llm-api-key-here"
Run the agent:
python agent/main.py
The agent will run with the hardcoded example question and print the execution trace and final answer to your console.
├── agent/ # The LangChain agent and its tools
├── streamlit/ # The Streamlit conversational UI
├── mcp/ # FastAPI server with core logic
├── neo4j/ # Neo4j configuration and data
├── data/ # SQLite databases
├── ops/ # Operational scripts (seeding, ingestion, etc.)
├── docker-compose.yml
├── Makefile
└── README.md