ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

GPT-6 Luna Batch API 调用实践笔记

GPT-6 Luna Batch API 调用实践笔记 摘要本文记录 GPT-6 Luna 模型通过 Batch API 进行批量异步推理的完整接入过程涵盖模型能力与规格、开发环境搭建、批量任务提交与轮询、推理强度配置、多模态请求、结构化输出、计费规则与成本估算、任务管理与错误处理并附完整批量文本分类示例帮助读者快速掌握 Batch API 的高效用法。目录① 模型能力与规格能力一览模型规格② 开发环境搭建前置要求安装 SDK初始化客户端Batch API 限流③ Batch API 基本流程准备批量请求文件上传文件并创建批量任务④ 轮询任务状态与获取结果轮询状态下载并解析结果⑤ 推理强度配置推理强度选择参考⑥ 多模态批量请求⑦ 结构化输出⑧ 计费规则与成本估算Batch 定价短上下文 ≤ 272K tokens按 Standard 50% 计长上下文阶梯单请求输入 272K tokens成本估算示例提示词缓存⑨ 任务管理与错误处理取消未完成的任务列出历史任务常见错误排查结果校验⑩ 完整示例批量文本分类参考文档记录 GPT-6 Luna 模型通过 Batch API 进行批量异步推理的接入过程涵盖环境准备、批量任务提交、结果轮询、参数调优等场景的代码示例。① 模型能力与规格GPT-6 Luna 是 GPT-6 系列中面向高吞吐量任务的模型支持文本和图像输入、文本输出具备可配置的推理强度和工具调用能力。能力一览能力说明多模态输入文本、图像、文件PDF推理控制reasoning effort: none / low / medium / high / xhigh / max工具调用Function Calling、Web Search、File Search、Computer Use结构化输出JSON Schema 约束提示词缓存自动缓存最低 1024 token 前缀Batch API异步批量处理50% 折扣模型规格项目参数模型 IDgpt-6-luna上下文窗口1,050,000 tokens最大输出128,000 tokens输入模态文本、图像、文件输出模态文本知识截止2026 年 5 月 18 日默认推理强度medium② 开发环境搭建前置要求Python 3.9OpenAI API Keypip安装 SDKpip install openai初始化客户端import os from openai import OpenAI client OpenAI( api_keyos.environ.get(OPENAI_API_KEY) )Batch API 限流TierRPMTPMBatch 队列上限Free不支持——Tier 1500500,0005,000,000Tier 25,0002,000,00020,000,000Tier 35,0004,000,00040,000,000Tier 410,00010,000,0001,000,000,000Tier 530,000180,000,00015,000,000,000③ Batch API 基本流程Batch API 用于异步处理大量请求按 Standard 费率的50%计费24 小时内完成。流程分三步提交 → 轮询 → 取结果。准备批量请求文件Batch API 要求先上传一个 JSONL 文件每行一个请求import json 构造批量请求数据 requests [ { custom_id: task-001, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: user, content: 请用一句话总结人工智能在医疗领域的应用} ], max_completion_tokens: 200 } }, { custom_id: task-002, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: user, content: 请用一句话总结区块链技术的核心思想} ], max_completion_tokens: 200 } }, { custom_id: task-003, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: user, content: 请用一句话总结边缘计算与云计算的区别} ], max_completion_tokens: 200 } } ] 写入 JSONL 文件 with open(batch_requests.jsonl, w, encodingutf-8) as f: for req in requests: f.write(json.dumps(req, ensure_asciiFalse) \n) print(f已生成 {len(requests)} 条请求)上传文件并创建批量任务# 1. 上传 JSONL 文件 batch_file client.files.create( fileopen(batch_requests.jsonl, rb), purposebatch ) 2. 创建批量任务 batch_job client.batches.create( input_file_idbatch_file.id, endpoint/v1/chat/completions, completion_window24h ) print(f批量任务 ID: {batch_job.id}) print(f状态: {batch_job.status})④ 轮询任务状态与获取结果轮询状态import time batch_id batch_job.id while True: batch client.batches.retrieve(batch_id) print(f状态: {batch.status} | f已完成: {batch.request_counts.completed} | f失败: {batch.request_counts.failed} | f总计: {batch.request_counts.total}) if batch.status in (completed, failed, cancelled, expired): break time.sleep(30) # 每 30 秒检查一次下载并解析结果if batch.status completed: # 下载结果文件 result_content client.files.content(batch.output_file_id) result_text result_content.text # 逐行解析 JSONL for line in result_text.strip().split(\n): result json.loads(line) custom_id result[custom_id] response result[response] if response[status_code] 200: content response[body][choices][0][message][content] print(f[{custom_id}] {content}) else: print(f[{custom_id}] 错误: {response}) # 如果有失败记录下载错误文件 if batch.error_file_id: error_content client.files.content(batch.error_file_id) print(f错误详情:\n{error_content.text})Batch 输入和结果文件保留30 天过期自动删除。⑤ 推理强度配置GPT-6 Luna 支持 6 档推理强度Batch 请求中通过reasoning_effort参数控制# 构造不同推理强度的请求 requests [ { custom_id: ftask-{i:03d}, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: user, content: f分析以下代码的时间复杂度并说明理由{code_snippet}} ], reasoning_effort: effort, # none / low / medium / high / xhigh / max max_completion_tokens: 1024 } } for i, (effort, code_snippet) in enumerate([ (none, def f(n): return n * 2), (low, def f(n): return sum(range(n))), (medium, def f(n): return [i*j for i in range(n) for j in range(n)]), (high, def f(n): return sorted([randint(0,n) for _ in range(n*n)])), ]) ] with open(batch_reasoning.jsonl, w, encodingutf-8) as f: for req in requests: f.write(json.dumps(req, ensure_asciiFalse) \n)推理强度选择参考effort适用场景token 消耗none简单分类、格式转换、提取最低low基础问答、短文本摘要低medium默认值通用任务中high代码分析、多步推理高xhigh复杂逻辑、数学证明较高max最深推理耗时最长最高注意Chat Completions API 中使用 Function Calling 时reasoning_effort只能设为none。需要同时使用工具和推理时改用 Responses API。⑥ 多模态批量请求Batch 请求同样支持图像输入将图片转为 base64 后放入消息体import base64 def image_to_data_url(image_path): with open(image_path, rb) as f: b64 base64.b64encode(f.read()).decode(utf-8) return fdata:image/jpeg;base64,{b64} 批量图片分类请求 image_files [img1.jpg, img2.jpg, img3.jpg] requests [] for i, img_path in enumerate(image_files): data_url image_to_data_url(img_path) requests.append({ custom_id: fimg-{i:03d}, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [{ role: user, content: [ {type: text, text: 请用 JSON 格式输出这张图片的类别和置信度格式: {category: ..., confidence: 0.0}}, {type: image_url, image_url: {url: data_url}} ] }], max_completion_tokens: 200, response_format: {type: json_object} } }) with open(batch_vision.jsonl, w, encodingutf-8) as f: for req in requests: f.write(json.dumps(req, ensure_asciiFalse) \n)⑦ 结构化输出通过response_format约束输出为 JSON Schemarequest { custom_id: extract-001, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: user, content: 从以下文本提取人名、日期和金额2026年9月22日张三向李四转账5000元。} ], max_completion_tokens: 256, response_format: { type: json_schema, json_schema: { name: extraction, strict: True, schema: { type: object, properties: { people: {type: array, items: {type: string}}, date: {type: string}, amount: {type: number} }, required: [people, date, amount] } } } } }⑧ 计费规则与成本估算Batch 定价短上下文 ≤ 272K tokens按 Standard 50% 计计费项Batch 单价/百万 tokens输入$0.05缓存命中输入$0.005缓存写入$0.0625输出$0.25长上下文阶梯单请求输入 272K tokens计费项Batch 长上下文单价输入$0.10缓存命中输入$0.01缓存写入$0.125输出$0.375阶梯按单次请求的输入 token 数判定不是 batch 内所有请求的合计。成本估算示例假设 1000 条请求每条输入 2000 tokens、输出 300 tokens无缓存num_requests 1000 input_tokens_per_req 2000 output_tokens_per_req 300 total_input num_requests * input_tokens_per_req # 2,000,000 total_output num_requests * output_tokens_per_req # 300,000 Batch 短上下文 input_cost (total_input / 1_000_000) * 0.05 # $0.10 output_cost (total_output / 1_000_000) * 0.25 # $0.075 total_cost input_cost output_cost print(f输入费用: ${input_cost:.4f}) print(f输出费用: ${output_cost:.4f}) print(f总费用: ${total_cost:.4f}) 输入费用: $0.1000 输出费用: $0.0750 总费用: $0.1750提示词缓存缓存自动生效无需改代码。命中条件前缀 ≥ 1024 tokensTTL 5–10 分钟最长 1 小时缓存命中按 $0.005/M 计费Batch 短上下文适合重复使用相同系统提示词的场景# 所有请求共享相同的系统 prompt长前缀只有 user 内容不同 system_prompt 你是一个专业的文本分类助手。请将输入文本分类到以下类别之一科技、财经、体育、娱乐、教育、健康。只输出类别名称不要解释。 * 20 # 确保超过 1024 tokens requests [ { custom_id: fcls-{i:04d}, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: system, content: system_prompt}, {role: user, content: text} ], max_completion_tokens: 10, reasoning_effort: none } } for i, text in enumerate(texts_to_classify) ] 首次请求写入缓存后续命中缓存输入费用降至 1/10⑨ 任务管理与错误处理取消未完成的任务# 取消批量任务 cancelled client.batches.cancel(batch_id) print(f取消后状态: {cancelled.status})列出历史任务# 列出最近的批量任务 batches client.batches.list(limit10) for b in batches.data: print(fID: {b.id} | 状态: {b.status} | f完成: {b.request_counts.completed}/{b.request_counts.total} | f创建时间: {b.created_at})常见错误排查错误原因解决方案invalid_api_keyAPI Key 错误检查环境变量file_too_largeJSONL 文件超限拆分为多个 batchrate_limit_exceeded超出 Batch 队列上限升级 tier 或分批提交batch_expired超过 24h 窗口未完成检查队列负载减少单次请求量invalid_request请求体格式错误校验 JSONL 每行的body结构结果校验# 下载结果后逐条校验 results [] for line in result_text.strip().split(\n): entry json.loads(line) if entry[response][status_code] 200: results.append(entry) else: # 记录失败请求后续重试 print(f失败: {entry[custom_id]} - {entry[response]}) 校验 JSON 结构化输出 for r in results: try: data json.loads(r[response][body][choices][0][message][content]) except json.JSONDecodeError: print(fJSON 解析失败: {r[custom_id]})⑩ 完整示例批量文本分类将以上步骤整合为一个完整的批量分类流程import os import json import time from openai import OpenAI client OpenAI(api_keyos.environ.get(OPENAI_API_KEY)) 1. 准备数据 texts [ 苹果发布新款 MacBook Pro搭载 M5 芯片, 美联储宣布降息 50 个基点, 中国队获得乒乓球世锦赛团体冠军, 某明星宣布退出娱乐圈, 教育部发布新一轮课程改革方案, 研究发现每天步行 8000 步可显著降低心血管风险, ] system_prompt 你是一个新闻分类助手。请将输入文本分类到以下类别之一科技、财经、体育、娱乐、教育、健康。只输出类别名称。 2. 构造 JSONL requests [ { custom_id: fnews-{i:03d}, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: system, content: system_prompt}, {role: user, content: text} ], max_completion_tokens: 10, reasoning_effort: none } } for i, text in enumerate(texts) ] with open(classify_batch.jsonl, w, encodingutf-8) as f: for req in requests: f.write(json.dumps(req, ensure_asciiFalse) \n) 3. 上传 创建任务 batch_file client.files.create( fileopen(classify_batch.jsonl, rb), purposebatch ) batch_job client.batches.create( input_file_idbatch_file.id, endpoint/v1/chat/completions, completion_window24h ) print(f任务已提交: {batch_job.id}) 4. 轮询 while True: batch client.batches.retrieve(batch_job.id) print(f[{batch.status}] {batch.request_counts.completed}/{batch.request_counts.total}) if batch.status in (completed, failed, cancelled, expired): break time.sleep(15) 5. 输出结果 if batch.status completed: result_text client.files.content(batch.output_file_id).text for line in result_text.strip().split(\n): entry json.loads(line) content entry[response][body][choices][0][message][content] print(f{entry[custom_id]}: {content})参考文档模型文档developers.openai.com/api/docs/models/gpt-6-lunaBatch APIdevelopers.openai.com/api/docs/batch定价developers.openai.com/api/docs/pricingGPT-6 使用指南developers.openai.com/api/docs/guides/latest-model以上内容基于 2026 年 9 月的 API 版本整理具体参数以官方文档为准。
返回列表