1. 多 Agent 协作为什么会失控:从 Workflow 到 Orchestration 的真实分水岭
如果你最近在折腾多 Agent 项目,大概率遇到过这种场景:三个 Agent 各干各的,Planner 刚规划完,Coder 已经写完代码,Reviewer 又发现设计有漏洞要求重来,Memory 模块还在不停写入新上下文。最后 token 烧了一大半,任务却没跑通。这不是模型不行,而是缺少一个总指挥——Orchestration(编排)。
Orchestration 是什么?简单说,它是多 Agent 系统的控制平面。它不负责推理,也不直接干活,而是决定“谁在什么时候做什么、做完之后下一步给谁、结果不合格怎么办”。适合谁?适合所有从单 Agent 玩具项目往多 Agent 生产系统迁移的开发者,尤其是用 MCP 接工具、用 A2A 做 Agent 间通信的团队。
我试过用固定 Workflow 串四个 Agent,任务一复杂就死锁。后来把调度逻辑抽出来做成独立的 Orchestration 层,配合 TaoToken 统一 Key 通道,才把最小可用的 AI 团队跑通。下面我把这套可复制的配置和验证流程完整拆给你。
核心检索词先明确:Orchestration 编排、AI Agent 协作、MCP 服务注册、A2A 协议、多 Agent 任务分发。这几个词贯穿全文,你跟着做就能跑通。
早期 Workflow 的问题不是分工不对,而是分工的时机错了。它在代码里写死了执行顺序,Planner → Researcher → Coder → Reviewer,一步接一步。任务少的时候没问题,但一旦 Researcher 搜到新资料需要回头改规划,或者 Reviewer 发现设计漏洞要求重跑,整个链路就乱了。有些 Agent 还在执行旧任务,有些已经切到新目标,还有些因为等别人而长期空闲。时间没花在解决问题上,全耗在等待和重复执行上。
Orchestration 的解法是:在一切还没发生前不要锁定所有决策,根据中间结果和实时状态,边走边定下一个谁上、干什么。它做四件事——规划与策略、执行与控制、状态与知识、质量与运营。这四层缺一不可。只有规划没有质量,就是只管派活不管验收;只有执行没有状态,就是每次做完就忘,每条链路重新开始。
MCP 和 A2A 是 Orchestration 的左膀右臂。MCP 解决 Agent 怎么调用工具,把所有外部能力抽象成 Tools、Resources、Prompts 三类,通过 stdio 或 HTTP+SSE 通信,把对接复杂度从 N×M 降到 N+M。A2A 解决 Agent 之间怎么交流,通过 AgentCard 发现能力、Task 管理生命周期、Message 传递内容。MCP 是垂直层决定能力边界,A2A 是水平层决定协作方式。两条协议叠在一起,Orchestration 才有完整的调度能力。
2. TaoToken 前置准备:统一 Key 与 API 通道配置
在跑多 Agent 协作之前,你需要一个统一的模型调用通道。原因很简单:多 Agent 系统里每个 Agent 可能用不同模型,Planner 用推理强的,Coder 用代码能力好的,Reviewer 用长上下文稳的。如果每个 Agent 各自配一套 Key 和 Base URL,管理成本极高,而且排查问题时根本分不清是哪个通道出的错。
TaoToken 在这里的角色是统一 Key/API 通道。你只需要一个 Key,就能在多个 Agent 之间切换不同模型,Base URL 统一指向https://taotoken.net/api。这样 Orchestration 层在分发任务时,不用关心每个 Agent 底层走的是哪个供应商,只需要指定 Model ID 就行。
前置准备分三步。第一步,拿到 API Key。访问 API Keys 管理页面(https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api_keys&utm_campaign=rewrite),创建一个新 Key,复制保存。注意这个 Key 只在创建时显示一次,丢了只能重建。
第二步,确认你要用的 Model ID。不同 Agent 角色建议用不同模型:Planner 用推理能力强的,Coder 用代码专项的,Reviewer 用长上下文稳定的。你可以在模型对话页面(https://taotoken.net/chat?utm_source=taotoken_aicg_blog_end&utm_content=model_chat&utm_campaign=rewrite)先手动测试每个模型的表现,确认符合预期再写进配置。
第三步,规划 Agent 角色和 MCP 服务。最小可用 AI 团队建议四个角色:Planner 负责拆解任务,Researcher 负责检索资料,Coder 负责写代码,Reviewer 负责验证结果。每个角色可以挂载不同的 MCP Server:Researcher 挂搜索类 MCP,Coder 挂文件系统和代码执行类 MCP,Reviewer 挂测试类 MCP。
这里有个关键点:Orchestration 层本身也需要调用模型来做调度决策。它用的模型建议选推理强、响应快的,因为每次任务分发都要它判断。这个模型和 Agent 用的模型可以不同,但都走同一个 TaoToken Key。
配置完成后,你的目录结构大概是这样:
ai-team/ ├── orchestration/ │ ├── config.json │ └── dispatcher.py ├── agents/ │ ├── planner/ │ │ └── agent-card.json │ ├── researcher/ │ │ └── agent-card.json │ ├── coder/ │ │ └── agent-card.json │ └── reviewer/ │ └── agent-card.json └── mcp-servers/ ├── search-server/ ├── fs-server/ └── test-server/每个 Agent 目录下的agent-card.json是 A2A 协议要求的名片文件,Orchestration 通过它发现 Agent 能力。MCP Server 目录下是各个工具服务的实现。Orchestration 目录下是调度核心。
如果你还没拿到 Key,现在去 API Keys 页面创建一个,后面所有配置都要用到。接入文档在 https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite,遇到参数问题可以先查这里。
3. 可复制配置:Agent 角色定义与 MCP 服务注册
这一章是核心,所有配置都可以直接复制修改。先给 Orchestration 的主配置,再给每个 Agent 的 AgentCard,最后给 MCP 服务注册示例。
3.1 Orchestration 主配置 config.json
{ "orchestration": { "base_url": "https://taotoken.net/api", "api_key": "sk-your-taotoken-key", "scheduler_model": "claude-3-5-sonnet-20241022", "max_retry": 3, "timeout_seconds": 300, "state_store": "./state/checkpoints.json" }, "agents": [ { "id": "planner", "card_url": "http://localhost:8001/.well-known/agent-card.json", "model": "claude-3-5-sonnet-20241022", "mcp_servers": [] }, { "id": "researcher", "card_url": "http://localhost:8002/.well-known/agent-card.json", "model": "gpt-4o", "mcp_servers": ["search-server"] }, { "id": "coder", "card_url": "http://localhost:8003/.well-known/agent-card.json", "model": "claude-3-5-sonnet-20241022", "mcp_servers": ["fs-server", "exec-server"] }, { "id": "reviewer", "card_url": "http://localhost:8004/.well-known/agent-card.json", "model": "gpt-4o", "mcp_servers": ["test-server"] } ], "mcp_servers": { "search-server": { "command": "node", "args": ["./mcp-servers/search-server/index.js"], "transport": "stdio" }, "fs-server": { "command": "node", "args": ["./mcp-servers/fs-server/index.js"], "transport": "stdio" }, "exec-server": { "command": "python", "args": ["./mcp-servers/exec-server/main.py"], "transport": "stdio" }, "test-server": { "command": "node", "args": ["./mcp-servers/test-server/index.js"], "transport": "stdio" } } }这个配置里,base_url和api_key是全局的,所有 Agent 和 Orchestration 都走这个通道。scheduler_model是 Orchestration 自己做调度决策时用的模型。agents数组里每个 Agent 指定自己的 AgentCard 地址、模型和挂载的 MCP Server。mcp_servers定义每个 MCP Server 的启动命令和传输方式。
3.2 AgentCard 示例:planner 的 agent-card.json
{ "name": "planner", "description": "任务规划 Agent,负责将复杂目标拆解为可执行子任务", "url": "http://localhost:8001", "version": "1.0.0", "capabilities": { "streaming": true, "pushNotifications": false }, "skills": [ { "id": "task-decomposition", "name": "任务拆解", "description": "将高层目标拆解为带依赖关系的子任务列表", "inputModes": ["text"], "outputModes": ["text", "json"] }, { "id": "dependency-analysis", "name": "依赖分析", "description": "分析子任务之间的依赖关系,标记可并行任务", "inputModes": ["json"], "outputModes": ["json"] } ], "defaultInputModes": ["text"], "defaultOutputModes": ["text", "json"] }其他三个 Agent 的 AgentCard 结构一样,只是name、description、skills不同。researcher 的 skills 里加web-search和content-extraction,coder 的 skills 里加code-generation和file-operation,reviewer 的 skills 里加code-review和test-execution。
3.3 MCP 服务注册示例:search-server
// mcp-servers/search-server/index.js import { Server } from "@modelcontextprotocol/sdk/server/index.js"; import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js"; const server = new Server( { name: "search-server", version: "1.0.0" }, { capabilities: { tools: {} } } ); server.setRequestHandler("tools/list", async () => ({ tools: [ { name: "web_search", description: "搜索公开网页信息", inputSchema: { type: "object", properties: { query: { type: "string", description: "搜索关键词" }, max_results: { type: "number", default: 5 } }, required: ["query"] } } ] })); server.setRequestHandler("tools/call", async (request) => { if (request.params.name === "web_search") { const { query, max_results = 5 } = request.params.arguments; // 这里接入你的搜索实现 return { content: [ { type: "text", text: `搜索结果:${query},共 ${max_results} 条` } ] }; } throw new Error("Unknown tool"); }); const transport = new StdioServerTransport(); await server.connect(transport);这个 MCP Server 注册了一个web_search工具,researcher Agent 通过 MCP 协议调用它。其他 MCP Server 结构类似,只是工具定义不同。fs-server 注册read_file、write_file、list_dir,exec-server 注册run_command,test-server 注册run_test。
3.4 Orchestration 调度核心 dispatcher.py
import json import httpx from pathlib import Path class Orchestrator: def __init__(self, config_path: str): self.config = json.loads(Path(config_path).read_text()) self.base_url = self.config["orchestration"]["base_url"] self.api_key = self.config["orchestration"]["api_key"] self.scheduler_model = self.config["orchestration"]["scheduler_model"] self.state = {} async def plan(self, goal: str) -> list: """调用调度模型做任务拆解""" async with httpx.AsyncClient() as client: resp = await client.post( f"{self.base_url}/v1/chat/completions", headers={"Authorization": f"Bearer {self.api_key}"}, json={ "model": self.scheduler_model, "messages": [ {"role": "system", "content": "你是任务规划器,将目标拆解为带依赖的子任务JSON数组。"}, {"role": "user", "content": goal} ] }, timeout=60 ) resp.raise_for_status() content = resp.json()["choices"][0]["message"]["content"] return json.loads(content) async def dispatch(self, task: dict, agent_id: str) -> dict: """通过 A2A 协议把任务派给指定 Agent""" agent = next(a for a in self.config["agents"] if a["id"] == agent_id) card_url = agent["card_url"] async with httpx.AsyncClient() as client: card_resp = await client.get(card_url) card = card_resp.json() # 构造 A2A Task 请求 task_resp = await client.post( f"{card['url']}/tasks", json={ "taskId": task["id"], "message": { "role": "user", "parts": [{"type": "text", "text": task["description"]}] } }, timeout=300 ) return task_resp.json() async def run(self, goal: str): tasks = await self.plan(goal) self.state["tasks"] = tasks for task in tasks: agent_id = task.get("assignee", "coder") result = await self.dispatch(task, agent_id) self.state.setdefault("results", []).append(result) # 质量检查 if not result.get("ok"): await self.retry(task) return self.state这段代码是 Orchestration 的最小实现:plan方法调用调度模型拆解任务,dispatch方法通过 A2A 协议把任务派给指定 Agent,run方法串起整个流程并做质量检查。你可以直接复制到项目里,改一下配置路径就能跑。
4. 验证请求:一次完整协作流程的成功结果
配置写完了,现在跑一次完整协作流程验证。目标是让 AI 团队完成一个具体任务,比如“写一个 Python 函数计算斐波那契数列并测试”。
4.1 启动所有服务
先启动四个 Agent 和四个 MCP Server。每个 Agent 是一个独立的 HTTP 服务,监听不同端口。MCP Server 通过 stdio 被 Agent 拉起,不需要单独启动。
# 启动 planner cd agents/planner && python server.py --port 8001 & # 启动 researcher cd agents/researcher && python server.py --port 8002 & # 启动 coder cd agents/coder && python server.py --port 8003 & # 启动 reviewer cd agents/reviewer && python server.py --port 8004 & # 启动 Orchestration cd orchestration && python dispatcher.py --goal "写一个Python函数计算斐波那契数列并测试"4.2 观察任务分发链路
Orchestration 启动后,先调用调度模型做任务拆解。你会看到类似这样的输出:
[ { "id": "task-1", "description": "分析需求:写一个Python函数计算斐波那契数列,需要处理边界条件", "assignee": "planner", "depends_on": [] }, { "id": "task-2", "description": "检索斐波那契数列的标准实现和常见边界条件处理方式", "assignee": "researcher", "depends_on": ["task-1"] }, { "id": "task-3", "description": "根据检索结果编写Python函数,包含输入校验和递归/迭代两种实现", "assignee": "coder", "depends_on": ["task-2"] }, { "id": "task-4", "description": "对生成的代码做测试,验证边界条件和性能", "assignee": "reviewer", "depends_on": ["task-3"] } ]然后 Orchestration 按依赖关系依次派发。task-1 和 task-2 没有依赖关系,可以并行;task-3 等 task-2 完成;task-4 等 task-3 完成。每个任务派发时,Orchestration 通过 AgentCard 找到对应 Agent 的地址,构造 A2A Task 请求发过去。
4.3 验证成功结果
当所有任务完成,Orchestration 的 state 里会记录完整结果。你可以通过以下方式验证:
# 查看状态文件 cat orchestration/state/checkpoints.json # 预期输出 { "tasks": [...], "results": [ {"taskId": "task-1", "ok": true, "output": "需求分析完成..."}, {"taskId": "task-2", "ok": true, "output": "检索到3种实现方式..."}, {"taskId": "task-3", "ok": true, "output": "def fibonacci(n): ..."}, {"taskId": "task-4", "ok": true, "output": "测试通过,边界条件覆盖..."} ] }如果每个 result 的ok都是 true,说明协作流程跑通了。如果某个任务ok为 false,Orchestration 会自动重试,最多重试max_retry次。重试仍失败,会在 state 里标记该任务为 failed,并记录错误信息。
4.4 用模型对话验证通道
在跑协作流程之前,建议先用模型对话页面(https://taotoken.net/chat?utm_source=taotoken_aicg_blog_end&utm_content=model_chat&utm_campaign=rewrite)手动验证一下 Key 和模型是否正常。发一条简单请求,确认返回正常,再跑多 Agent 流程。这样可以排除通道问题,把排查范围缩小到 Orchestration 逻辑本身。
验证请求的完整链路是:Orchestration 调用调度模型 → 调度模型返回任务列表 → Orchestration 通过 A2A 派发任务 → Agent 通过 MCP 调用工具 → Agent 返回结果 → Orchestration 汇总。每一环都有日志,出问题时按链路逐段排查。
5. 本篇常见错排查:401、local proxy failed、reading choices、OAuth
多 Agent 协作跑不起来,90% 的问题集中在四类报错。下面逐个拆解原因和修复方式。
5.1 401 Unauthorized
报错原文:
{"error": {"message": "Invalid API key", "type": "invalid_request_error", "code": 401}}原因:API Key 写错、过期、或者没带上。多 Agent 场景下,常见的是某个 Agent 的配置里 Key 没同步,或者环境变量没读到。
排查步骤:先确认config.json里的api_key和你在 API Keys 页面创建的一致。然后检查每个 Agent 启动时是否读到了这个配置。如果 Agent 是独立进程,确认它启动时的工作目录正确,能读到配置文件。最后用 curl 直接测一下:
curl -X POST https://taotoken.net/api/v1/chat/completions \ -H "Authorization: Bearer sk-your-key" \ -H "Content-Type: application/json" \ -d '{"model": "claude-3-5-sonnet-20241022", "messages": [{"role": "user", "content": "hi"}]}'如果 curl 也 401,说明 Key 本身有问题,去 API Keys 页面重新创建一个。如果 curl 正常但 Agent 报 401,说明 Agent 配置读取有问题。
5.2 local proxy failed
报错原文:
Error: local proxy failed: connection refused原因:Agent 或 MCP Server 启动失败,端口没监听。多 Agent 场景下,常见的是某个 Agent 进程崩了,但 Orchestration 还在往它发请求。
排查步骤:先确认所有 Agent 进程都在跑。用ps aux | grep server.py看进程列表。然后逐个测端口:
curl http://localhost:8001/.well-known/agent-card.json curl http://localhost:8002/.well-known/agent-card.json curl http://localhost:8003/.well-known/agent-card.json curl http://localhost:8004/.well-known/agent-card.json哪个端口不通,就去查对应 Agent 的启动日志。常见原因是端口被占用、依赖没装、或者 AgentCard 文件路径写错。
5.3 reading choices 报错
报错原文:
KeyError: 'choices'原因:模型返回的 JSON 结构不符合预期。多 Agent 场景下,常见的是某个 Agent 用的模型不支持当前请求格式,或者返回了错误信息但代码没处理。
排查步骤:在dispatcher.py的plan方法里,把resp.json()打印出来看实际返回结构。如果返回的是{"error": ...},说明请求本身有问题。如果返回结构正常但没有choices,说明模型 ID 写错了。确认scheduler_model和每个 Agent 的model字段都是有效的 Model ID。
5.4 OAuth 相关报错
报错原文:
OAuth token expired or invalid原因:如果你用的是需要 OAuth 的模型通道,token 过期了。TaoToken 的 API Key 方式不需要 OAuth,但如果你在 Agent 里混用了其他通道,可能出现这个问题。
排查步骤:确认所有 Agent 和 Orchestration 都走https://taotoken.net/api这个 Base URL,并且用 API Key 认证。不要混用 OAuth 通道。如果确实需要 OAuth,单独处理 token 刷新逻辑,不要和 API Key 混在一起。
5.5 配置检查清单
跑协作流程前,对照这个清单检查一遍:
| 检查项 | 正确值 | 常见错误 |
|---|---|---|
| Base URL | https://taotoken.net/api | 写成首页地址 |
| API Key | sk- 开头 | 复制时多了空格 |
| Model ID | 有效模型标识 | 拼写错误 |
| AgentCard 路径 | /.well-known/agent-card.json | 路径拼错 |
| MCP 启动命令 | node/python 绝对路径 | 相对路径找不到 |
| 端口 | 8001-8004 不冲突 | 端口被占用 |
如果四类报错都排查完还是跑不通,去接入文档(https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite)查对应错误码的说明。文档里有完整的参数列表和示例请求。
6. 从最小团队到生产系统:Orchestration 的下一步
最小可用 AI 团队跑通之后,你可以往三个方向扩展。
第一,加 Agent。当前四个角色是最小集,实际项目里可能需要 Memory Agent 管理长期上下文、Security Agent 做输入输出过滤、Human-in-the-loop Agent 处理需要人工确认的环节。每加一个 Agent,就在config.json的agents数组里加一项,写好 AgentCard,挂载对应 MCP Server。
第二,加 MCP 工具。当前每个 Agent 挂了一到两个 MCP Server,实际项目里可能需要数据库查询、API 调用、文件转换等更多工具。每个 MCP Server 独立实现,注册到mcp_servers配置里,Agent 按需挂载。MCP 的好处是工具实现和 Agent 解耦,换 Agent 不用重写工具。
第三,加状态管理。当前 state 存在内存和本地文件里,生产环境需要换成 Redis 或数据库。Orchestration 的state_store配置指向持久化存储,每次任务状态变化都写入,支持断点续跑和故障恢复。
长期编码和 Agent 场景,建议用 Coding Plan(https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding_plan&utm_campaign=rewrite),它有专门的额度策略,适合高频调用的多 Agent 系统。如果你还在选模型阶段,先去模型对话页面(https://taotoken.net/chat?utm_source=taotoken_aicg_blog_end&utm_content=model_chat&utm_campaign=rewrite)对比几个模型的实际表现,再决定每个 Agent 用哪个。
最后说一个实际踩过的坑:Orchestration 的调度模型不要和 Agent 用同一个。调度模型需要快速响应和强推理,Agent 模型需要专项能力。混用会导致调度延迟高,整个协作流程变慢。分开配置,各司其职,系统才跑得顺。