☰
【Agent】【workflow】10.多策略工作流与反思机制分析
2026/10/11 1:39:31 网站建设 项目流程

1. 案例目标

本案例展示了如何使用LlamaIndex的Workflow框架实现一个多策略RAG(检索增强生成)系统,该系统能够并行运行多种检索策略,并通过反思机制选择最佳答案。具体目标包括:

  • 实现三种不同的RAG策略:基础RAG、高Top-K检索和重排序检索
  • 设计查询质量评估机制,自动改进低质量查询
  • 构建结果评估系统,自动选择最佳回答
  • 展示工作流中的并行处理和事件收集机制

2. 技术栈与核心依赖

  • llama-index-core - LlamaIndex核心库,提供基础功能
  • llama-index-llms-openai - OpenAI语言模型集成
  • llama-index-utils-workflow - 工作流可视化工具
  • llama-index-readers-file - 文件读取器,用于处理PDF
  • llama-index-embeddings-openai - OpenAI嵌入模型
  • llama-index.core.postprocessor.rankGPT_rerank - RankGPT重排序器

3. 环境配置

pip install llama-index-core llama-index-llms-openai llama-index-utils-workflow llama-index-readers-file llama-index-embeddings-openai

案例使用了旧金山2016-2018年的年度预算PDF文件作为数据源,需要下载这些文件到data目录:

!mkdir data
!wget "https://www.dropbox.com/scl/fi/xt3squt47djba0j7emmjb/2016-CSF_Budget_Book_2016_FINAL_WEB_with-cover-page.pdf?rlkey=xs064cjs8cb4wma6t5pw2u2bl&dl=0" -O "data/2016-CSF_Budget_Book_2016_FINAL_WEB_with-cover-page.pdf"
!wget "https://www.dropbox.com/scl/fi/jvw59g5nscu1m7f96tjre/2017-Proposed-Budget-FY2017-18-FY2018-19_1.pdf?rlkey=v988oigs2whtcy87ti9wti6od&dl=0" -O "data/2017-Proposed-Budget-FY2017-18-FY2018-19_1.pdf"
!wget "https://www.dropbox.com/scl/fi/izknlwmbs7ia0lbn7zzyx/2018-o0181-18.pdf?rlkey=p5nv2ehtp7272ege3m9diqhei&dl=0" -O "data/2018-o0181-18.pdf"

需要设置OpenAI API密钥:

from google.colab import userdata os.environ["OPENAI_API_KEY"] = userdata.get("openai-key")

4. 案例实现

4.1 定义事件类

工作流中定义了多种事件类型,用于不同步骤间的通信:

class JudgeEvent(Event):
query: str


class BadQueryEvent(Event):
query: str


class NaiveRAGEvent(Event):
query: str


class HighTopKEvent(Event):
query: str


class RerankEvent(Event):
query: str


class ResponseEvent(Event):
query: str
response: str


class SummarizeEvent(Event):
query: str
response: str

4.2 工作流类实现

ComplicatedWorkflow类实现了多策略RAG系统,包含以下关键方法:

4.2.1 索引加载与创建

def load_or_create_index(self, directory_path, persist_dir):
# 检查索引是否已存在
if os.path.exists(persist_dir):
print("Loading existing index...")
# 从磁盘加载索引
storage_context = StorageContext.from_defaults(persist_dir=persist_dir)
index = load_index_from_storage(storage_context)
else:
print("Creating new index...")
# 从指定目录加载文档
documents = SimpleDirectoryReader(directory_path).load_data()
# 从文档创建新索引
index = VectorStoreIndex.from_documents(documents)
# 将索引持久化到磁盘
index.storage_context.persist(persist_dir=persist_dir)
return index

4.2.2 查询质量评估

@step
async def judge_query(
self, ctx: Context, ev: StartEvent | JudgeEvent
) -> BadQueryEvent | NaiveRAGEvent | HighTopKEvent | RerankEvent:
# 初始化
llm = await ctx.store.get("llm", default=None)
if llm is None:
await ctx.store.set("llm", OpenAI(model="gpt-4o", temperature=0.1))
await ctx.store.set("index", self.load_or_create_index("data", "storage"))
# 使用聊天引擎以便记住之前的交互
await ctx.store.set("judge", SimpleChatEngine.from_defaults())

response = await ctx.store.get("judge").chat(
f"""
Given a user query, determine if this is likely to yield good results from a RAG system as-is. If it's good, return 'good', if it's bad, return 'bad'.
Good queries use a lot of relevant keywords and are detailed. Bad queries are vague or ambiguous.
Here is the query: {ev.query}
"""
)

if response == "bad":
# 尝试改进查询
return BadQueryEvent(query=ev.query)
else:
# 将查询发送到所有3种策略
self.send_event(NaiveRAGEvent(query=ev.query))
self.send_event(HighTopKEvent(query=ev.query))
self.send_event(RerankEvent(query=ev.query))

4.2.3 查询改进

@step
async def improve_query(
self, ctx: Context, ev: BadQueryEvent
) -> JudgeEvent:
response = await ctx.store.get("llm").complete(
f"""
This is a query to a RAG system: {ev.query}
The query is bad because it is too vague.
Please provide a more detailed query that includes specific keywords and removes any ambiguity.
"""
)

return JudgeEvent(query=str(response))

# 4.2.4 三种RAG策略实现

@step
async def naive_rag(
self, ctx: Context, ev: NaiveRAGEvent
) -> ResponseEvent:
index = await ctx.store.get("index")
engine = index.as_query_engine(similarity_top_k=5)
response = engine.query(ev.query)
print("Naive response:", response)
return ResponseEvent(
query=ev.query,
source="Naive",
response=str(response)
)


@step
async def high_top_k(
self, ctx: Context, ev: HighTopKEvent
) -> ResponseEvent:
index = await ctx.store.get("index")
engine = index.as_query_engine(similarity_top_k=20)
response = engine.query(ev.query)
print("High top k response:", response)
return ResponseEvent(
query=ev.query,
source="High top k",
response=str(response)
)


@step
async def rerank(self, ctx: Context, ev: RerankEvent) -> ResponseEvent:
index = await ctx.store.get("index")
reranker = RankGPTRerank(top_n=5, llm=await ctx.store.get("llm"))
retriever = index.as_retriever(similarity_top_k=20)
engine = RetrieverQueryEngine.from_args(
retriever=retriever,
node_postprocessors=[reranker],
)
response = engine.query(ev.query)
print("Reranker response:", response)
return ResponseEvent(
query=ev.query,
source="Reranker",
response=str(response)
)

4.2.5 结果评估与选择

@step
async def judge(self, ctx: Context, ev: ResponseEvent) -> StopEvent:
ready = ctx.collect_events(ev, [ResponseEvent] * 3)
if ready is None:
return None

response = await ctx.store.get("judge").chat(
f"""
A user has provided a query and 3 different strategies have been used to try to answer the query. Your job is to decide which strategy best answered the query.
The query was: {ev.query}
Response 1 ({ready[0].source}): {ready[0].response}
Response 2 ({ready[1].source}): {ready[1].response}
Response 3 ({ready[2].source}): {ready[2].response}
Please provide the number of the best response (1, 2, or 3). Just provide the number, with no other text or preamble.
"""
)

best_response = int(str(response))
print(
f"Best response was number {best_response}, which was from {ready[best_response-1].source}"
)

return StopEvent(result=str(ready[best_response - 1].response))

4.3 工作流可视化

使用draw_all_possible_flows函数生成工作流的可视化图表:

draw_all_possible_flows( ComplicatedWorkflow, filename="complicated_workflow.html" )

4.4 工作流运行

c = ComplicatedWorkflow(timeout=120, verbose=True)
result = await c.run(query="How has spending changed?")
print(result)

5. 案例效果

当运行工作流时,系统执行以下步骤:

  1. 查询质量评估:系统首先评估查询质量,对于模糊查询如"How has spending changed?",会判断为低质量查询
  2. 查询改进:对于低质量查询,系统使用LLM生成更详细的查询,如添加具体关键词和上下文
  3. 并行执行三种RAG策略:
    • 基础RAG:使用similarity_top_k=5进行检索
    • 高Top-K检索:使用similarity_top_k=20进行检索
    • 重排序检索:先检索20个结果,然后使用RankGPTRerank重排序为5个
  4. 结果评估:收集所有三种策略的结果,使用LLM评估并选择最佳答案
  5. 输出最佳结果:返回被评估为最佳的答案

示例输出:

"Best response was number 3, which was from High top k"

"Spending has increased over the years, with the total budget showing growth in various areas such as aid assistance/grants, materials & supplies, equipment, debt service, services of other departments, and professional & contractual services. Additionally, there have been new investments in programs like workforce development, economic development, film services, and finance and administration. The budget allocations have been adjusted to accommodate changing needs and priorities, reflecting an overall increase in spending across different departments and programs."

6. 案例实现思路

本案例的核心实现思路是构建一个具有反思能力的多策略RAG系统:

  • 多策略并行处理:通过Workflow框架的事件机制,实现三种RAG策略的并行执行,提高检索效率和覆盖面
  • 查询质量评估与改进:引入查询质量评估步骤,自动识别低质量查询并使用LLM改进,确保输入质量
  • 结果自动评估:使用LLM作为裁判,评估不同策略的结果并选择最佳答案,实现自动化的质量控制
  • 事件驱动架构:通过定义多种事件类型和步骤间的依赖关系,构建灵活的工作流
  • 状态管理:使用Context存储共享资源(如LLM实例、索引等),避免重复初始化

工作流程图

StartEvent → judge_query → (并行) NaiveRAGEvent/HighTopKEvent/RerankEvent → ResponseEvent → judge → StopEvent

对于低质量查询:BadQueryEvent → improve_query → JudgeEvent → (循环回judge_query)

7. 扩展建议

  • 添加更多RAG策略:可以添加更多检索策略,如混合检索、多跳检索等,进一步提高系统多样性
  • 实现查询路由:根据查询类型自动选择最适合的策略组合,而不是总是执行所有策略
  • 增强结果评估:引入更复杂的评估指标,如答案准确性、相关性、完整性等,提高评估质量
  • 添加用户反馈循环:收集用户对结果的反馈,用于改进查询改进和结果评估模型
  • 实现缓存机制:对相似查询和结果进行缓存,提高系统响应速度
  • 添加查询历史:维护查询历史,支持上下文感知的查询改进
  • 实现多模态支持:扩展系统以支持图像、表格等多模态数据的检索

8. 总结

本案例展示了如何使用LlamaIndex的Workflow框架构建一个具有反思能力的多策略RAG系统。该系统通过并行执行多种检索策略,自动评估和选择最佳结果,显著提高了RAG系统的可靠性和答案质量。案例的核心价值在于:

  • 展示了Workflow框架在构建复杂AI系统中的强大能力
  • 提供了一种解决RAG系统不确定性的有效方法
  • 演示了事件驱动架构在AI系统中的应用
  • 实现了查询质量自动评估和改进机制
  • 构建了自动化的结果评估和选择系统

这种多策略与反思机制结合的方法,可以广泛应用于各种需要高质量答案的AI应用场景,如智能问答、文档分析、知识检索等,为构建更加可靠和智能的RAG系统提供了有价值的参考。

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询