
1. 案例目标本案例展示了如何使用LlamaIndex工作流(Workflow)实现一个带有内联引用的RAG(检索增强生成)查询引擎。主要目标包括实现一个能够为生成答案提供精确引用的RAG系统展示如何使用工作流构建多步骤的RAG处理流程演示如何将检索到的节点分割为更小的引用块展示如何在生成的答案中嵌入引用标记提供完整的端到端实现从数据加载到查询响应2. 技术栈与核心依赖from llama_index.core.workflow import ( Event, Context, Workflow, StartEvent, StopEvent, step, ) from llama_index.core import SimpleDirectoryReader, VectorStoreIndex from llama_index.llms.openai import OpenAI from llama_index.embeddings.openai import OpenAIEmbedding from llama_index.core.prompts import PromptTemplate from llama_index.core.response_synthesizers import get_response_synthesizer核心依赖包括LlamaIndex工作流框架用于构建和编排多步骤处理流程向量存储索引用于文档的向量化存储和检索OpenAI模型包括LLM和嵌入模型用于生成和向量化响应合成器用于基于检索到的节点生成最终答案事件系统用于工作流步骤间的数据传递3. 环境配置# 安装依赖 !pip install -U llama-index # 设置OpenAI API密钥 os.environ[OPENAI_API_KEY] sk-... # 下载示例数据 !mkdir -p data/paul_graham/ !wget https://raw.githubusercontent.com/run-llama/llama_index/main/docs/examples/data/paul_graham/paul_graham_essay.txt -O data/paul_graham/paul_graham_essay.txt环境配置包括安装最新版本的LlamaIndex库设置OpenAI API密钥确保能够访问GPT模型下载示例数据(Paul Graham的文章)创建数据存储目录4. 案例实现4.1 工作流设计CitationQueryEngine工作流包含以下步骤索引数据创建向量索引使用索引和查询检索相关节点为检索到的节点添加引用标记合成最终响应包含内联引用4.2 定义事件为了处理这些步骤需要定义几个事件from llama_index.core.workflow import Event from llama_index.core.schema import NodeWithScore class RetrieverEvent(Event): 检索结果事件 nodes: list[NodeWithScore] class CreateCitationsEvent(Event): 添加引用事件 nodes: list[NodeWithScore]4.3 引用提示模板定义用于生成带引用答案的提示模板CITATION_QA_TEMPLATE PromptTemplate( Please provide an answer based solely on the provided sources. When referencing information from a source, cite the appropriate source(s) using their corresponding numbers. Every answer should include at least one source citation. Only cite a source when you are explicitly referencing it. If none of the sources are helpful, you should indicate that. For example:\n Source 1:\n The sky is red in the evening and blue in the morning.\n Source 2:\n Water is wet when the sky is red.\n Query: When is water wet?\n Answer: Water will be wet when the sky is red [2], which occurs in the evening [1].\n Now its your turn. Below are several numbered sources of information: \n------\n {context_str} \n------\n Query: {query_str}\n Answer: )4.4 工作流实现步骤1: 检索节点step async def retrieve( self, ctx: Context, ev: StartEvent ) - Union[RetrieverEvent, None]: RAG的入口点由带有查询的StartEvent触发 query ev.get(query) if not query: return None print(fQuery the database with: {query}) # 将查询存储在全局上下文中 await ctx.store.set(query, query) if ev.index is None: print(Index is empty, load some documents before querying!) return None retriever ev.index.as_retriever(similarity_top_k2) nodes retriever.retrieve(query) print(fRetrieved {len(nodes)} nodes.) return RetrieverEvent(nodesnodes)步骤2: 创建引用节点step async def create_citation_nodes( self, ev: RetrieverEvent ) - CreateCitationsEvent: 修改检索到的节点为引用创建细粒度源。 接受NodeWithScore对象列表并将其内容分割为更小的块 为每个块创建新的NodeWithScore对象。每个新节点都标记为编号源 允许在查询结果中进行更精确的引用。 Args: nodes (List[NodeWithScore]): 要处理的NodeWithScore对象列表。 Returns: List[NodeWithScore]: 新的NodeWithScore对象列表其中每个对象 代表原始节点的较小块标记为源。 nodes ev.nodes new_nodes: List[NodeWithScore] [] text_splitter SentenceSplitter( chunk_sizeDEFAULT_CITATION_CHUNK_SIZE, chunk_overlapDEFAULT_CITATION_CHUNK_OVERLAP, ) for node in nodes: text_chunks text_splitter.split_text( node.node.get_content(metadata_modeMetadataMode.NONE) ) for text_chunk in text_chunks: text fSource {len(new_nodes)1}:\n{text_chunk}\n new_node NodeWithScore( nodeTextNode.parse_obj(node.node), scorenode.score ) new_node.node.text text new_nodes.append(new_node) return CreateCitationsEvent(nodesnew_nodes)步骤3: 合成响应step async def synthesize( self, ctx: Context, ev: CreateCitationsEvent ) - StopEvent: 使用检索到的节点返回流式响应 llm OpenAI(modelgpt-4o-mini) query await ctx.store.get(query, defaultNone) synthesizer get_response_synthesizer( llmllm, text_qa_templateCITATION_QA_TEMPLATE, refine_templateCITATION_REFINE_TEMPLATE, response_modeResponseMode.COMPACT, use_asyncTrue, ) response await synthesizer.asynthesize(query, nodesev.nodes) return StopEvent(resultresponse)4.5 创建索引documents SimpleDirectoryReader(data/paul_graham).load_data() index VectorStoreIndex.from_documents( documentsdocuments, embed_modelOpenAIEmbedding(model_nametext-embedding-3-small), )4.6 运行工作流# 创建工作流实例 w CitationQueryEngineWorkflow() # 运行查询 result await w.run(queryWhat information do you have, indexindex)4.7 查看引用# 显示结果 display(Markdown(f{result})) # 查看引用源 print(result.source_nodes[0].node.get_text()) print(result.source_nodes[1].node.get_text())5. 案例效果本案例实现了以下效果精确引用生成的答案中每个事实都带有对应的引用标记细粒度源分割将原始文档分割为更小的块提供更精确的引用清晰的工作流将RAG过程分解为明确的步骤便于理解和维护可验证性用户可以根据引用标记追溯到原始文档内容示例输出The provided sources contain various insights into Paul Grahams experiences and thoughts on programming, writing, and his educational journey. For instance, he reflects on his early experiences with programming on the IBM 1401, where he struggled to create meaningful programs due to the limitations of the technology at the time [2]. He also describes his transition to using microcomputers, which allowed for more interactive programming experiences [3]. Additionally, Graham shares his initial interest in philosophy during college, which he later found less engaging compared to the fields of artificial intelligence and programming [3]. Overall, the sources highlight his evolution as a writer and programmer, as well as his changing academic interests.关键特性本案例中的引用系统具有以下特点每个引用块都有明确的编号如[1]、[2]等引用块大小可配置默认为512个字符重叠20个字符答案中引用的格式遵循学术标准便于验证引用内容保留原始文档的上下文信息6. 案例实现思路本案例的实现思路如下工作流设计将RAG过程分解为检索、引用创建和响应合成三个主要步骤事件驱动使用事件系统在工作流步骤间传递数据和状态引用分割将检索到的文档分割为更小的块每个块标记为编号源提示工程设计专门的提示模板指导模型在答案中包含引用上下文管理使用工作流上下文存储查询信息在步骤间共享工作流数据流StartEvent(查询, 索引) → retrieve步骤retrieve步骤 → RetrieverEvent(检索到的节点)RetrieverEvent → create_citation_nodes步骤create_citation_nodes步骤 → CreateCitationsEvent(带引用的节点)CreateCitationsEvent → synthesize步骤synthesize步骤 → StopEvent(带引用的响应)7. 扩展建议基于本案例可以考虑以下扩展方向多模态引用扩展支持图像、表格等多模态内容的引用引用质量评估实现引用相关性和准确性的自动评估动态引用调整根据查询复杂度动态调整引用块大小和数量引用排序基于相关性对引用进行排序和筛选交互式引用允许用户点击引用查看完整上下文引用格式定制支持多种引用格式如APA、MLA等引用网络分析分析引用间的关系构建知识图谱跨文档引用支持跨多个文档的引用和关联8. 总结本案例展示了如何使用LlamaIndex工作流实现一个带有内联引用的RAG查询引擎。案例中的关键技术点包括工作流编排使用LlamaIndex工作流框架构建多步骤处理流程引用系统将检索到的文档分割为带编号的引用块提示工程设计专门的提示模板指导模型生成带引用的答案事件驱动使用事件系统在工作流步骤间传递数据和状态通过这种方式生成的答案不仅内容准确而且每个事实都可以追溯到原始文档大大提高了RAG系统的可信度和可验证性。这种引用系统特别适用于学术研究、法律分析、医疗诊断等对准确性和可追溯性要求较高的场景。