Downloads · 30 days
0
ai-singer/four-agent
four-agent is a machine learning model from ai-singer. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
基于 LangGraph 的 Planner(规划)- RuleMaker(标准制定)- Executor(执行)- Evaluator(评估) 通用四智能体协作框架,支持持久化、人工审批、模型分层、错误降级、可观测性、外部工具调用与多场景验证。
Downloads · 30 days
0
Access
Public
Updated Jul 8, 2026
Repo size
214 MB
Likes
0
Public
Click a slice to open those files.
.pyd117 MB · 38%
From the Hugging Face model README
基于 LangGraph 的 Planner(规划)- RuleMaker(标准制定)- Executor(执行)- Evaluator(评估) 通用四智能体协作框架,支持持久化、人工审批、模型分层、错误降级、可观测性、外部工具调用与多场景验证。
本项目将单一 LLM 调用逐步演进为一个结构化的多智能体协作系统:
整个流程由 LangGraph 的 StateGraph 编排,支持重试循环、人工审批、状态持久化与模型分层策略。框架已在研报生成、代码审查、课程设计、客户投诉处理、竞品分析、合同审查六个垂直场景下验证通用性。
| 特性 | 说明 |
|---|---|
| 四智能体闭环 | planner → rule_maker → executor → evaluator,评估失败后自动重规划。 |
| LangGraph 编排 | 使用 StateGraph、conditional_edges 与 interrupt 实现工作流。 |
| 模型分层策略 | 规划/标准制定使用强模型,执行/评估使用轻量模型,兼顾质量与成本。 |
| 持久化 | 支持 MemorySaver(内存)、SqliteSaver(SQLite)与 PostgresSaver,可断点续跑。 |
| 人工审批(HITL) | evaluator 后可插入人工审批节点,默认使用 LangGraph functional interrupt。 |
| 错误处理与降级 | 每个节点被 _robust_node_wrapper 包裹,异常时返回降级状态并记录日志。 |
| 可观测性 | 内置本地 GraphRunObserver 追踪节点事件;支持 JSON/simple 结构化日志。 |
| 外部工具调用 | ExecutorAgent 支持 web_search、file_reader 等工具,可扩展自定义工具。 |
| 多场景验证 | 提供研报生成、代码审查、课程设计、客户投诉处理、竞品分析、合同审查 6 个垂直场景示例。 |
| 可交互 Demo | 提供 Gradio 可视化界面与 FastAPI HTTP 服务。 |
| RAG 向量检索 | VectorSearchTool 基于 ChromaDB 实现本地知识库检索。 |
| MCP 协议接入 | MCPClientTool 可连接外部 MCP Server,复用其暴露的工具。 |
| A2A 协议接入 | A2AClientTool 支持 Agent 之间互相委托任务,并提供 /a2a 端点。 |
| 效果量化基准 | benchmark.py 对比单 Agent 与四智能体,支持关键词覆盖、ROUGE、BERTScore、LLM-as-judge。 |
| 流式输出 | StreamingExecutor 与 /workflow/stream 接口实现 Agent 工作流 token 级 SSE。 |
| CI / 代码规范 | GitHub Actions 自动运行 pytest + ruff。 |
| Docker 部署 | 支持 Dockerfile 与 docker-compose 一键启动。 |
| 端到端验证 | 提供基于真实 LLM API 的端到端测试,默认跳过,手动开启后验证全流程。 |
┌─────────┐ ┌─────────────┐ ┌──────────┐ ┌───────────┐
│ Planner │───▶│ RuleMaker │───▶│ Executor │───▶│ Evaluator │
└─────────┘ └─────────────┘ └──────────┘ └───────────┘
▲ │
│ 评估失败且未达最大重试次数 │
└──────────────────────────────────────────────────┘
在启用 enable_human_review=True 时,evaluator 后会插入 human_review 节点:
Evaluator ──▶ HumanReview ──(retry/end)──▶ Planner / END
.
├── agents.py # 四个 Agent 的实现
├── state.py # FourAgentState 状态定义
├── four_agent_graph.py # LangGraph StateGraph 编排核心
├── tools.py # 外部工具抽象与默认实现(搜索、文件读取)
├── llm_service.py # 底层 LLM 调用服务(百炼/智谱)
├── template_manager.py # Prompt 模板管理
├── thinking_animation.py # 思考动画组件
├── model_strategy.py # 模型分层策略
├── persistence.py # 持久化工厂(memory/sqlite/postgres)
├── observability.py # 本地观测器与 LangSmith 配置
├── error_handling.py # 异常类型与节点降级包装器
├── human_in_the_loop.py # 人工审批 interrupt 节点
├── logging_config.py # 结构化日志配置
├── conftest.py # 测试公共 fixture(FakeLLM 等)
├── test_agents.py # Agent 真实 API 测试(手动运行)
├── test_llm_service.py # LLMService 真实 API 测试(手动运行)
├── test_graph.py # StateGraph 单元测试
├── test_tools.py # 工具模块单元测试
├── test_model_strategy.py # 模型策略单元测试
├── test_persistence.py # 持久化单元测试
├── test_observability.py # 可观测性单元测试
├── test_error_handling.py # 错误处理单元测试
├── test_human_in_the_loop.py # 人工审批单元测试
├── test_logging.py # 日志单元测试
├── test_e2e.py # 真实 LLM 端到端测试(默认跳过)
├── examples/ # 6 个垂直场景示例
│ ├── research_report/ # 研报生成
│ ├── code_review/ # 代码审查
│ ├── course_design/ # 课程设计
│ ├── customer_service/ # 客户投诉处理
│ ├── competitive_analysis/ # 竞品分析
│ ├── contract_review/ # 合同审查
│ └── a2a_collaboration/ # A2A 跨 Agent 协作示例
├── demo_gradio.py # Gradio 交互式 Demo
├── api_server.py # FastAPI 服务
├── benchmark.py # 效果量化基准
├── benchmark/ # benchmark 评估指标模块
│ └── metrics.py # 关键词覆盖、ROUGE、BERTScore、LLM-as-judge
├── streaming_executor.py # Agent 工作流 token 级流式执行器
├── app.py # Hugging Face Spaces 入口
├── Dockerfile # Docker 构建
├── docker-compose.yml # Docker Compose 启动
├── requirements_hf.txt # Hugging Face Spaces 精简依赖
├── .github/workflows/ci.yml # GitHub Actions CI
├── .env # 本地环境变量(真实 API key,不提交)
├── .env.example # 环境变量模板
├── 开发规划.md # 项目演进路线图
└── 优化方案.md # 竞争力分析与优化路径
pip install \
langgraph==0.3.34 \
langchain-core \
langchain-openai \
python-dotenv \
requests \
pytest \
"langgraph-checkpoint-sqlite<3.0.0"
注意:
langgraph-checkpoint-sqlite需选择与当前langgraph版本兼容的 2.x。工具依赖:
ddgs(网页搜索)、pypdf(PDF)、python-docx(Word)已包含在requirements.txt中。如不需要真实工具调用,可暂不安装。
复制模板并填入真实 key:
cp .env.example .env
编辑 .env:
API_KEY_QWEN=your-bailian-api-key
MODEL_QWEN=qwen-plus
MODEL_QWEN_FAST=qwen-turbo
API_KEY_GLM=your-zhipu-api-key
MODEL_GLM=glm-5.2
from llm_service import LLMService
from four_agent_graph import build_four_agent_graph
from persistence import create_memory_saver
llm = LLMService({
"provider": "bailian",
"model": "qwen-plus",
"api_key": "your-api-key",
"base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
})
app = build_four_agent_graph(
llm,
checkpointer=create_memory_saver(),
enable_human_review=False,
)
final = app.invoke({
"user_goal": "为团队分享准备一份关于 AI Agent 的 PPT 大纲",
"professional_role": "",
"plan": [],
"evaluation_rubric": {},
"execution_result": None,
"execution_trace": [],
"evaluation_result": None,
"current_step": "plan",
"should_react": False,
"retry_count": 0,
"max_retries": 3,
"error_message": None,
"human_decision": None,
"history": [],
})
print(final["plan"])
print(final["execution_result"])
print(final["evaluation_result"])
from llm_service import LLMService
from four_agent_graph import build_four_agent_graph
from model_strategy import create_default_model_strategy
strong = LLMService({"provider": "bailian", "model": "qwen-plus", ...})
fast = LLMService({"provider": "bailian", "model": "qwen-turbo", ...})
strategy = create_default_model_strategy(strong, fast)
app = build_four_agent_graph(model_strategy=strategy)
build_four_agent_graph 默认注入 web_search 与 file_reader。ExecutorAgent 会根据计划决定调用哪些工具,并将工具结果追加到 execution_result。
from four_agent_graph import build_four_agent_graph
app = build_four_agent_graph(llm)
final = app.invoke(initial_state)
print(final["execution_result"]) # 包含工具返回内容
print(final["execution_trace"]) # 工具调用记录
实现 BaseTool 接口即可扩展工具:
from tools import BaseTool
class CalculatorTool(BaseTool):
name = "calculator"
description = "执行简单数学计算"
input_schema = {
"type": "object",
"properties": {"expression": {"type": "string", "description": "数学表达式"}},
"required": ["expression"],
}
def run(self, expression: str, **kwargs) -> str:
try:
return str(eval(expression))
except Exception as exc:
return f"计算失败: {exc}"
app = build_four_agent_graph(llm, tools=[CalculatorTool()])
from tools import VectorSearchTool
# 从文本列表构建索引
rag = VectorSearchTool(
documents=["Python 是一种解释型语言...", "LangGraph 是多智能体编排框架..."],
collection_name="kb",
top_k=2,
)
app = build_four_agent_graph(llm, tools=[rag])
from tools import MCPClientTool
mcp = MCPClientTool(server_command=["python", "my_mcp_server.py"])
app = build_four_agent_graph(llm, tools=[mcp])
from tools import A2AClientTool
a2a = A2AClientTool(agent_url="http://localhost:8000/a2a")
app = build_four_agent_graph(llm, tools=[a2a])
运行自带示例(默认 FakeLLM,不消耗 token):
python examples/a2a_collaboration/run.py
python examples/research_report/run.py
python examples/code_review/run.py
python examples/course_design/run.py
python examples/customer_service/run.py
python examples/competitive_analysis/run.py
python examples/contract_review/run.py
python demo_gradio.py
访问 http://127.0.0.1:7860,选择场景并一键运行,可视化观察每个 Agent 的输出。
uvicorn api_server:app --reload --port 8000
主要接口:
POST /invoke:端到端调用POST /stream:SSE 流式返回节点事件POST /workflow/stream:Agent 工作流 token 级 SSE,实时输出每个智能体的生成 tokenPOST /chat/stream:原始 LLM token 流式输出POST /plan:仅运行 PlannerPOST /execute:仅运行 ExecutorPOST /evaluate:仅运行 Evaluatordocker-compose up --build
将同时启动 FastAPI(http://localhost:8000)与 Gradio(http://localhost:7860)。
项目根目录已提供 app.py 作为 Hugging Face Spaces 入口。手动部署步骤:
API_KEY_QWEN(以及可选的 MODEL_QWEN)。注意:Hugging Face Spaces 免费版为 CPU 环境,启动时可能需要几分钟安装依赖。为减少构建时间,可上传 requirements_hf.txt 并改名为
requirements.txt。
from persistence import create_memory_saver
app = build_four_agent_graph(
llm,
checkpointer=create_memory_saver(),
enable_human_review=True,
)
# 使用 stream 运行,捕获 interrupt 事件
for chunk in app.stream(initial_state, config):
if "__interrupt__" in chunk:
print("等待人工审批:", chunk["__interrupt__"][0]["value"])
break
# 人工决策后恢复
from langgraph.types import Command
list(app.stream(Command(resume={"decision": "retry"}), config))
from persistence import create_checkpointer
saver = create_checkpointer(
"postgres",
conn_string="postgresql://postgres:postgres@localhost:5432/four_agent",
)
app = build_four_agent_graph(llm, checkpointer=saver)
或使用环境变量 DATABASE_URL:
export DATABASE_URL="postgresql://postgres:postgres@localhost:5432/four_agent"
python -c "from persistence import create_checkpointer; create_checkpointer('postgres')"
本项目通过 7 个差异较大的垂直场景与协作模式验证框架通用性:
| 场景 | 目录 | 输入示例 | 验证能力 |
|---|---|---|---|
| 研报生成 | examples/research_report/ | 写一份新能源汽车行业研报 | 长文本规划、Web 搜索、结构化输出 |
| 代码审查 | examples/code_review/ | 审查 sample_code.py | 文件读取、技术评估、安全审计 |
| 课程设计 | examples/course_design/ | 为初中生设计 8 课时 Python 课程 | 教育角色适配、评分维度自定义 |
| 客户投诉处理 | examples/customer_service/ | 处理订单延迟投诉 | 服务角色适配、情绪识别、回复生成 |
| 竞品分析 | examples/competitive_analysis/ | 比亚迪海豹 vs 特斯拉 Model 3 | 对比分析、多维度评分、Web 搜索 |
| 合同审查 | examples/contract_review/ | 审查软件开发合同 | 文件读取、法律风险识别 |
| A2A 协作 | examples/a2a_collaboration/ | 父 Agent 委托子 Agent 完成竞品分析 | A2A 协议、跨 Agent 任务委托 |
每个场景包含 input.json(输入)、output.md(典型输出样例)、run.py(一键运行)。
benchmark.py 提供单 Agent vs 四智能体的对比:
# 快速 FakeLLM 模式
python benchmark.py
# 真实 LLM 模式(消耗 API token)
RUN_BENCHMARK=1 python benchmark.py
# 真实 LLM + LLM-as-judge(额外消耗 token)
RUN_BENCHMARK=1 RUN_BENCHMARK_JUDGE=1 python benchmark.py
评估指标包括:关键词覆盖率、结构得分、长度得分、ROUGE、BERTScore(可选依赖)、LLM-as-judge。报告输出到 benchmark_report.json。
基于 qwen-plus 运行 6 个垂直场景,启用 LLM-as-judge(1-10 分制):
| 场景 | 单 Agent 评分 | 四智能体评分 | 四智能体耗时 | 备注 |
|---|---|---|---|---|
| 研报生成 | 9 | 9 | 102.75s | 两者均覆盖四大模块,四智能体额外产出计划与评分表 |
| 代码审查 | 10 | 10 | 108.19s | 四智能体结构化分类更系统 |
| 课程设计 | 10 | 10 | 111.08s | 四智能体提供课时总表与教师支持项 |
| 客户投诉处理 | 9 | 9 | 55.50s | 两者均生成专业回复 |
| 竞品分析 | 10 | 9 | 148.66s | 四智能体调用搜索并输出三维对比矩阵 |
| 合同审查 | 9 | 4 | 95.64s | 四智能体读取文件后产生幻觉,需改进工具结果合成 |
平均 LLM-as-judge 评分:单 Agent 9.5 / 四智能体 8.5。
说明:四智能体在多数场景下与单 Agent 质量相当,并额外提供可解释的计划、评分表与评估轨迹;合同审查场景暴露出工具调用后未充分基于真实文件内容合成答案的问题,是当前最大改进点。
详细评分理由与原始输出见 benchmark_report.json。
使用 FakeLLM 固定返回 JSON,覆盖图编排、策略、持久化、错误处理、HITL、日志等。
python -m pytest -v
预期:48 passed, 6 skipped(跳过项含 E2E 测试、真实工具/向量/MCP/Postgres 依赖测试)
# Windows
$env:RUN_TOOL_TESTS="1"
python -m pytest test_tools.py -v
# Linux/macOS
RUN_TOOL_TESTS=1 python -m pytest test_tools.py -v
需要启动本地 Postgres 并设置 DATABASE_URL:
# Windows
$env:RUN_POSTGRES_TESTS="1"
python -m pytest test_persistence.py::test_postgres_saver_persists_final_state -v
# Linux/macOS
RUN_POSTGRES_TESTS=1 python -m pytest test_persistence.py::test_postgres_saver_persists_final_state -v
# Windows
$env:RUN_E2E_TESTS="1"
python -m pytest test_e2e.py -v
# Linux/macOS
RUN_E2E_TESTS=1 python -m pytest test_e2e.py -v
预期:2 passed
项目已配置 .github/workflows/ci.yml,每次 push/PR 自动运行 pytest 与 ruff。
python test_agents.py
python test_llm_service.py
本项目按照规划分阶段演进:
| 阶段 | 目标 | 关键交付 |
|---|---|---|
| 阶段一 | 基础重构 | LLMService、TemplateManager、ThinkingAnimation 模块化拆分 |
| 阶段二 | 三智能体原型 | Planner → Executor → Evaluator 基础循环 |
| 阶段三 | 四智能体完善 | 引入独立 RuleMakerAgent,实现公平评估闭环 |
| 阶段四 | 生产级增强 | 持久化、HITL、模型分层、错误降级、可观测性、端到端验证 |
| 阶段五 | 竞争力优化 | 多场景验证、工具层扩展、Gradio/FastAPI/Docker 演示化 |
当前已完成全部五个阶段。
RuleMaker 在任务执行前制定评分表,不知道执行器的能力;Evaluator 只依据评分表打分,不知道执行器的意图。这种结构避免了"既当运动员又当裁判员"的评估偏差。
每个节点通过 _robust_node_wrapper 包装:
GraphRunObserver 记录每个节点的进入/退出事件与关键状态变量。enable_human_review:显式控制是否启用人工审批。RUN_E2E_TESTS:显式控制是否运行真实 API 测试。FOUR_AGENT_LOG_FORMAT:切换日志格式。.env 文件包含真实 key,请勿提交到 Git。.env 已在 .gitignore 中忽略。test_e2e.py 默认跳过,开启后会调用真实 LLM,请注意 token 消耗。WebSearchTool 依赖 ddgs,FileReaderTool 依赖 pypdf 与 python-docx。未安装时工具会抛出友好提示,不会中断整个工作流。langgraph==0.3.34,配套使用 langgraph-checkpoint-sqlite<3.0.0,避免版本冲突。MIT