ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

DeepSeek-V4-Flash-0731 Prompt 编码格式全解析:多轮对话、工具调用、思考模式与快速指令的完整实现

DeepSeek-V4-Flash-0731 Prompt 编码格式全解析:多轮对话、工具调用、思考模式与快速指令的完整实现 人工智能大模型基础模型DeepSeek【免费下载链接】DeepSeek-V4-Flash-0731项目地址https://ai.gitcode.com/hf_mirrors/deepseek-ai/DeepSeek-V4-Flash-0731点击查看免费下载本指南围绕 DeepSeek-V4-Flash-0731 仓库中 encoding/README.md 所定义的 prompt 编码格式结合 encoding_dsv4.py 参考实现与 test_encoding_dsv4.py 测试用例系统讲解 DeepSeek-V4 系列模型如何编码多轮对话、工具调用DSML 格式、扩展思考reasoning以及快速指令任务。读完本文你将掌握encode_messages/parse_message_from_completion_text两个核心函数的完整语义、每个特殊 token 与参数的真实作用并能直接复制文中代码构建可运行的对话、工具调用与思考模式请求。一、编码格式总览一个自包含的参考实现DeepSeek-V4 的 prompt 编码格式由 encoding/README.md 正式定义处理四类核心场景多轮对话multi-turn conversations、工具调用tool calling、扩展思考extended thinking / reasoning与快速指令任务quick instruction tasks。仓库提供了一份自包含的参考实现 encoding_dsv4.py无需依赖任何外部框架即可完成编码与解码其主要 API 为encode_messages(messages, thinking_mode, contextNone, drop_thinkingTrue, add_default_bos_tokenTrue, reasoning_effortNone)把 OpenAI 风格的消息列表编码为 DeepSeek-V4 的 prompt 字符串是编码的主入口parse_message_from_completion_text(text, thinking_mode)把模型单轮输出解析回结构化 assistant 消息含reasoning_content、content、tool_calls三个字段。在 inference/generate.py 的交互式推理中二者被直接串联使用先用encode_messages(messages, thinking_modechat)编码用户输入并调用模型再用parse_message_from_completion_text(completion, thinking_modechat)把生成的文本回填为 assistant 消息从而支撑多轮对话循环。快速上手示例from encoding_dsv4 import encode_messages, parse_message_from_completion_text # 编码一段对话 messages [ {role: system, content: You are a helpful assistant.}, {role: user, content: What is 22?}, ] prompt encode_messages(messages, thinking_modethinking) # begin▁of▁sentenceYou are a helpful assistant.UserWhat is 22?Assistantthink # 解析模型输出为结构化消息 completion Simple arithmetic./think2 2 4.end▁of▁sentence parsed parse_message_from_completion_text(completion, thinking_modethinking) # {role: assistant, reasoning_content: Simple arithmetic., content: 2 2 4., tool_calls: []}注意parse_message_from_completion_text只处理格式良好的模型输出不尝试修正或恢复模型偶尔产生的畸形文本生产环境建议在此基础上增加额外的错误处理。这一点从实现中可得到印证——encoding_dsv4.py 的解析逻辑对缺失/think、缺失 EOS token、工具调用格式错误、内容中残留特殊 token 等情况一律直接assert或raise ValueError。二、特殊 Token 与消息角色特殊 Token 一览编码格式中定义的全部特殊 token常量定义见 encoding_dsv4.pyTokenPurposebegin▁of▁sentenceBeginning of sequence (BOS)对话最开头插入end▁of▁sentenceEnd of assistant turn (EOS)UserUser turn prefixAssistantAssistant turn prefixlatest_reminderLatest reminder日期、地区、语言环境等think//thinkReasoning block delimiters思考块的起止标记DSMLDSML markup token工具调用标记语言前缀另外还有一组面向内部分类任务的快速指令 token定义于DS_TASK_SP_TOKENS见 encoding_dsv4.pyaction、query、authority、domain、title、read_url将在第五节详述。支持的消息角色编码支持以下角色system、user、assistant、tool、latest_reminder和developer。关于developer角色的说明developer角色仅用于内部搜索 Agent 流水线从 tests/test_input_3.json 可见其典型形态——在 developer 消息上定义search/open/find等工具通用聊天与工具调用任务不需要它官方 API 也不接受带该角色的消息。角色渲染规则源码级在render_messageencoding_dsv4.py中各角色的渲染逻辑为system直接输出content若带tools字段则追加\n\n 工具 schema 块若带response_format字段则追加## Response Format:结构化输出模板developer内容以User开头即被编码为一段用户侧指令同样可携带 tools 与 response_formatuser输出User前缀然后输出content若带content_blocks则逐块渲染text块原样输出tool_result块包装进tool_result.../tool_result未知块类型输出[Unsupported ...]占位latest_reminder输出latest_reminder 内容tool直接raise NotImplementedError——DeepSeek-V4 没有独立的 tool 角色工具结果必须通过merge_tool_messages()预处理合并进 user 消息assistant渲染顺序为{reasoning}{content}{tool_calls}最后附 EOS token可通过wo_eosTrue省略 EOS。三、对话编码chat 模式与 thinking 模式基础多轮对话一个简单的多轮对话被编码为begin▁of▁sentence{system_prompt} User{user_message}Assistant/think{response}end▁of▁sentence User{user_message_2}Assistant/think{response_2}end▁of▁sentenceBOS token 总在对话最开头预置add_default_bos_tokenTrue且无前置 context 时见 encoding_dsv4.py在chat 模式thinking_modechat下/think紧跟在Assistant之后立即关闭思考块模型直接生成内容不产生 reasoning。交错思考模式Interleaved Thinking Mode在thinking 模式thinking_modethinking下模型先输出显式的think.../think推理块再给出最终回答begin▁of▁sentence{system_prompt} User{message}Assistantthink{reasoning}/think{response}end▁of▁sentencedrop_thinking 参数历史推理的裁剪策略drop_thinking参数默认值为True控制是否保留较早轮次的推理内容核心逻辑见 encoding_dsv4.py 与_drop_thinking_messagesencoding_dsv4.py无工具场景drop_thinking生效。最后一个 user 消息之前的所有 assistant 轮次的 reasoning 会被剥离reasoning_content从消息中移除只有最后一个 assistant 轮次保留think.../think块有工具场景system 或 developer 消息上定义了toolsdrop_thinking自动失效所有轮次都保留推理。原因是工具调用对话需要完整上下文模型必须跨多次工具调用追踪多步推理。该行为在 test_case_2 中得到验证两轮对话中第一轮 assistant 的reasoning_contentThe user said hello在编码结果中缺失而最后一轮的推理块完整保留。实际编码示例无工具、thinking 模式输入 tests/test_input_2.json两轮问答均带reasoning_content编码结果见 tests/test_output_2.txtbegin▁of▁sentenceYou are a helpful assistant.UserHelloAssistant/thinkHi there! How can I help you?end▁of▁sentenceUserWhat is the capital of France?AssistantthinkThe user asks about the capital of France. It is Paris./thinkThe capital of France is Paris.end▁of▁sentence注意第一轮 assistant 的输出是/thinkHi there!...——由于它在最后一个 user 消息之前思考块被裁剪只保留了/think关闭符最后一轮则完整包含think.../think。四、工具调用DSML 格式工具定义与 schema 注入工具以 OpenAI 兼容格式定义在system或developer消息的tools字段中。当存在工具时编码器会在 system/user prompt 中注入如下 schema 块模板见 encoding_dsv4.py 的TOOLS_TEMPLATE## Tools You have access to a set of tools to help answer the users question. You can invoke tools by writing a 「DSML」tool_calls block like the following: 「DSML」tool_calls 「DSML」invoke name$TOOL_NAME 「DSML」parameter name$PARAMETER_NAME stringtrue|false$PARAMETER_VALUE「DSML」parameter ... 「DSML」invoke 「DSML」invoke name$TOOL_NAME2 ... 「DSML」invoke 「DSML」tool_calls实际输出中的标记字符为DSML为避免混淆此处用「DSML」代替显示。模板随后给出两条使用约束字符串参数原样指定并设stringtrue数字、布尔、数组、对象等其余类型以 JSON 传递并设stringfalse若思考模式被启用由think触发必须先在think.../think中输出完整推理再进行任何工具调用或最终回复否则在/think之后直接输出工具调用或回复。最后注入### Available Tool Schemas段每个工具 schema 以 JSON 行JSON Lines 风格列出并以 You MUST strictly follow the above defined tool name and parameter schemas to invoke tool calls. 收尾。schema 渲染由render_tools()encoding_dsv4.py完成工具列表先经tools_from_openai_format()抽取tool[function]字段encoding_dsv4.py。一次真实的工具调用assistant 轮次中的实际工具调用长这样见 tests/test_output_1.txtAssistantthinkThe user wants to know the weather in Beijing. I should use the get_weather tool./think 「DSML」tool_calls 「DSML」invoke nameget_weather 「DSML」parameter namelocation stringtrueBeijing「DSML」parameter 「DSML」parameter nameunit stringtruecelsius「DSML」parameter 「DSML」invoke 「DSML」tool_callsend▁of▁sentence「DSML」代指DSML。规则要点stringtrue参数值为原始字符串如Beijing、celsiusstringfalse参数值为 JSONnumber、boolean、array、object如{count: 5}会编码为count参数stringfalse。参数编码实现为encode_arguments_to_dsml()encoding_dsv4.py先把argumentsJSON 字符串json.loads为 dict再逐参数判断isinstance(v, str)来决定string标志解码方向对应decode_dsml_to_arguments()encoding_dsv4.py将(value, is_string_flag)还原为 JSON 参数字符串。OpenAI 格式与内部格式的互转由tool_calls_from_openai_format/tool_calls_to_openai_formatencoding_dsv4.py完成。工具结果的合并与排序工具执行结果以tool_result标签包装在 user 消息中Usertool_result{result_json}/tool_resultAssistantthink...由于 DeepSeek-V4 没有独立的 tool 角色merge_tool_messages()encoding_dsv4.py负责把 OpenAI 格式中独立的tool角色消息转换为带content_blocks的 user 消息tool_result块携带tool_use_id与随后的text块会合并进同一个 user 消息。当同一 user 消息中存在多个工具结果时sort_tool_results_by_call_order()encoding_dsv4.py按照前一条 assistant 消息中tool_calls的原始顺序对结果重新排序保证多工具并行调用时结果与调用一一对应。端到端示例带工具的思考对话输入 tests/test_input_1.json 定义了两个工具get_weather、search与四轮消息system → user → assistant 带工具调用 → tool 结果 → assistant 最终回答编码结果见 tests/test_output_1.txtbegin▁of▁sentenceYou are a helpful assistant. ## Tools ... {name: get_weather, description: Get the weather for a specific location, parameters: {...}} {name: search, description: Search the web for information, parameters: {...}} You MUST strictly follow the above defined tool name and parameter schemas to invoke tool calls. UserWhats the weather in Beijing?AssistantthinkThe user wants to know the weather in Beijing. I should use the get_weather tool./think 「DSML」tool_calls 「DSML」invoke nameget_weather 「DSML」parameter namelocation stringtrueBeijing「DSML」parameter 「DSML」parameter nameunit stringtruecelsius「DSML」parameter 「DSML」invoke 「DSML」tool_callsend▁of▁sentenceUsertool_result{temperature: 22, condition: sunny, humidity: 45}/tool_resultAssistantthinkGot the weather data. Let me format a nice response./thinkThe weather in Beijing is currently sunny with a temperature of 22°C and 45% humidity.end▁of▁sentence「DSML」代指DSML。由于系统消息定义了 toolsdrop_thinking被自动禁用两轮 assistant 的思考块都被保留。test_case_1 同时验证了解析侧从第一个 assistant 轮次中正确提取出 1 个get_weather工具调用参数还原为{location: Beijing, unit: celsius}从最终轮次提取出纯文本回答。五、reasoning_effort思考强度控制在 thinking 模式下reasoning_effort参数选择三个等级之一控制模型在回答前的推理投入程度。该等级仅以文本前缀的形式预置在 prompt 最开头system 消息之前其余编码在各等级之间完全一致。三种等级的取值与对应前缀定义于REASONING_EFFORT_PROMPTSencoding_dsv4.pyreasoning_effortPrompt prefixlow默认nonehighReasoning Effort: Absolute maximum ...maxReasoning Effort: Beyond maximum ...reasoning_effort在 chat 模式thinking_modechat下无效因为该模式下模型根本不产生推理块。从实现看只有index 0 and thinking_mode thinking时前缀才会被写入encoding_dsv4.py且None按low处理传入非法等级会直接触发断言。high的完整前缀文本Reasoning Effort: Absolute maximum with no shortcuts permitted. You MUST be very thorough in your thinking and comprehensively decompose the problem to resolve the root cause, rigorously stress-testing your logic against all potential paths, edge cases, and adversarial scenarios. Explicitly write out your entire deliberation process, documenting every intermediate step, considered alternative, and rejected hypothesis to ensure absolutely no assumption is left unchecked.max的完整前缀文本Reasoning Effort: Beyond maximum — exhaustive, relentless, and uncompromising. You MUST reason with the utmost depth and rigor, leaving absolutely nothing to chance: exhaustively decompose the problem into its most fundamental components, trace every causal chain to its root, and resolve the underlying cause rather than any surface symptom. Do not stop reasoning until you have independently verified the solution from multiple angles and are certain that no assumption remains unchecked and no error remains undiscovered.六、快速指令特殊 TokenQuick Instruction Tasks快速指令 token 用于辅助分类与生成类任务通过消息的task字段追加以触发模型输出单 token 或短格式结果。全部 token 由VALID_TASKS校验encoding_dsv4.py非法 task 名会触发断言。Special TokenDescriptionFormataction判断用户 prompt 是否需要联网搜索或可直接回答...User{prompt}Assistantthinkactiontitle在首个 assistant 回复后生成简短的对话标题...Assistant{response}end▁of▁sentencetitlequery为用户 prompt 生成搜索查询...User{prompt}queryauthority分类用户 prompt 对来源权威性的需求...User{prompt}authoritydomain识别用户 prompt 所属领域...User{prompt}domainextracted_urlread_url判断用户 prompt 中每个 URL 是否应抓取并阅读...User{prompt}extracted_url{url}read_url消息格式中的使用规则实现于 encoding_dsv4.pyaction任务挂在 user 消息上actiontoken 放置在 assistant 前缀与思考 token 之后触发路由决策例如输出 Search 或 Answer其他任务query、authority、domain、read_url挂在 user 消息上任务 token 直接追加在用户内容之后title任务挂在 assistant 消息上titletoken 追加在 assistant 的 EOS 之后下一条 assistant 消息即提供生成的标题。实际示例action 路由任务tests/test_input_4.json 展示了典型场景system 消息 latest_reminder 一轮带长回答的 assistant 一个带task: action的 user 消息。编码结果见 tests/test_output_4.txtbegin▁of▁sentence该助手为DeepSeek-V3由深度求索公司创造。 今天是2025年10月17日星期五。latest_reminder2024-11-15,上海市,App,中文User热海大滚锅是世界著名温泉吗Assistant/think关于热海大滚锅...end▁of▁sentenceUser世界著名温泉有哪些Assistant/thinkactionSearchend▁of▁sentence为便于阅读已省略长回答正文。注意 action 任务的编码特征在thinking_modechat下Assistant之后先输出/think关闭思考块再追加action随后模型直接输出单 token 路由结果Search。七、模型输出解析从文本到结构化消息parse_message_from_completion_text(text, thinking_mode)encoding_dsv4.py按以下顺序解析单轮模型输出thinking 模式下先读取think.../think之间的内容作为reasoning_content缺失/think直接断言失败继续读取直到 EOS token 或工具调用块起始标记\n\n「DSML」tool_calls得到summary_content若命中工具调用块则进入parse_tool_calls()encoding_dsv4.py解析invoke name与每个parameter name / string标志还原为 OpenAI 格式的tool_calls列表工具调用之后不允许再有其他内容校验整段文本已消费完毕且summary_content/reasoning_content中不残留任何特殊 token。返回结构固定为{role: assistant, content: ..., reasoning_content: ..., tool_calls: [...]}tool_calls为 OpenAI 格式。八、运行测试验证编码行为仓库自带完整的测试套件 test_encoding_dsv4.py共 4 个用例逐一覆盖上述能力case 1thinking 模式 工具调用多轮、工具结果合并进 user编码结果与 tests/test_output_1.txt 逐字符比对并验证工具调用与最终回答的解析case 2thinking 模式无工具验证drop_thinking确实剥离早期轮次的 reasoningcase 3交错思考 搜索场景developer 角色带工具 latest_reminder中文内容对照 tests/test_input_3.json 与 tests/test_output_3.txtcase 4chat 模式下的快速指令任务action 路由对照 tests/test_input_4.json 与 tests/test_output_4.txt。在encoding/目录下执行python test_encoding_dsv4.py即可运行全部测试cd encoding python test_encoding_dsv4.py期望输出All 4 tests passed!。这些测试既是格式行为的可执行规范也是自定义对话、工具调用与任务 prompt 时最直接的调试参考。九、小结与使用建议DeepSeek-V4 的编码格式可以概括为三条主线对话与思考BOS 开头、User/Assistant分隔角色think.../think承载推理thinking 模式默认通过drop_thinking裁剪历史推理以节省上下文工具场景自动保留全部推理工具调用OpenAI 兼容工具定义经## Tools模板注入调用以DSMLtool_calls块表达string属性区分字符串参数与 JSON 参数工具结果以tool_result合并进 user 消息并按调用顺序排序任务与强度控制reasoning_effort以 prompt 前缀实现思考强度分级task字段触发action/query/authority/domain/title/read_url等单 token 输出任务。实操建议通用对话与工具调用场景使用system/user/assistant/tool角色即可developer角色与action等任务 token 属于内部搜索 Agent 流水线专用parse_message_from_completion_text仅接受格式良好的输出生产环境需包裹错误处理。更深入的推理调用方式可参考 inference/README.md 与 inference/generate.py编码与推理联合起来即构成完整的 DeepSeek-V4 应用闭环。赞分享人工智能大模型基础模型DeepSeek【免费下载链接】DeepSeek-V4-Flash-0731项目地址https://ai.gitcode.com/hf_mirrors/deepseek-ai/DeepSeek-V4-Flash-0731点击查看免费下载相关推荐DeepSeek-V4-Flash-Vision-Exp 提示词编码完全指南特殊 Token、思考模式与工具调用模板逐行拆解DeepSeek V4 Flash Vision Exp 提示词编码完全指南特殊 Token、思考模式与工具调用模板逐行拆解 本指南带你逐行拆解 DeepSe大模型基础模型多模态计算机视觉DeepSeek逆向Qwen3.8-Flash-Next的Chat Template思考块与XML工具调用格式完整实现详解逆向Qwen3.8 Flash Next的Chat Template思考块与XML工具调用格式完整实现详解 本文逆向解析开源多模态模型 Qwen3.8 Fla人工智能基础模型大模型多模态Llama 3.2 Vision 模型 Prompt 格式完全指南从多模态对话到工具调用Llama 3.2 Vision 模型 Prompt 格式完全指南从多模态对话到工具调用 导读 本文以 llama models 仓库中的 vision_pr人工智能大模型基础模型创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表