**摘要:**本文记录 GPT-6 Luna 模型通过 Batch API 进行批量异步推理的完整接入过程,涵盖模型能力与规格、开发环境搭建、批量任务提交与轮询、推理强度配置、多模态请求、结构化输出、计费规则与成本估算、任务管理与错误处理,并附完整批量文本分类示例,帮助读者快速掌握 Batch API 的高效用法。
目录
[① 模型能力与规格](#① 模型能力与规格)
[② 开发环境搭建](#② 开发环境搭建)
[安装 SDK](#安装 SDK)
[Batch API 限流](#Batch API 限流)
[③ Batch API 基本流程](#③ Batch API 基本流程)
[④ 轮询任务状态与获取结果](#④ 轮询任务状态与获取结果)
[⑤ 推理强度配置](#⑤ 推理强度配置)
[⑥ 多模态批量请求](#⑥ 多模态批量请求)
[⑦ 结构化输出](#⑦ 结构化输出)
[⑧ 计费规则与成本估算](#⑧ 计费规则与成本估算)
[Batch 定价(短上下文 ≤ 272K tokens,按 Standard 50% 计)](#Batch 定价(短上下文 ≤ 272K tokens,按 Standard 50% 计))
[长上下文阶梯(单请求输入 > 272K tokens)](#长上下文阶梯(单请求输入 > 272K tokens))
[⑨ 任务管理与错误处理](#⑨ 任务管理与错误处理)
[⑩ 完整示例:批量文本分类](#⑩ 完整示例:批量文本分类)
记录 GPT-6 Luna 模型通过 Batch API 进行批量异步推理的接入过程,涵盖环境准备、批量任务提交、结果轮询、参数调优等场景的代码示例。
① 模型能力与规格
GPT-6 Luna 是 GPT-6 系列中面向高吞吐量任务的模型,支持文本和图像输入、文本输出,具备可配置的推理强度和工具调用能力。
能力一览
| 能力 | 说明 |
|---|---|
| 多模态输入 | 文本、图像、文件(PDF) |
| 推理控制 | reasoning effort: none / low / medium / high / xhigh / max |
| 工具调用 | Function Calling、Web Search、File Search、Computer Use |
| 结构化输出 | JSON Schema 约束 |
| 提示词缓存 | 自动缓存,最低 1024 token 前缀 |
| Batch API | 异步批量处理,50% 折扣 |
模型规格
| 项目 | 参数 |
|---|---|
| 模型 ID | gpt-6-luna |
| 上下文窗口 | 1,050,000 tokens |
| 最大输出 | 128,000 tokens |
| 输入模态 | 文本、图像、文件 |
| 输出模态 | 文本 |
| 知识截止 | 2026 年 5 月 18 日 |
| 默认推理强度 | medium |
② 开发环境搭建
前置要求
- Python 3.9+
- OpenAI API Key
- pip
安装 SDK
pip install openai
初始化客户端
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get("OPENAI_API_KEY")
)
Batch API 限流
| Tier | RPM | TPM | Batch 队列上限 |
|---|---|---|---|
| Free | 不支持 | --- | --- |
| Tier 1 | 500 | 500,000 | 5,000,000 |
| Tier 2 | 5,000 | 2,000,000 | 20,000,000 |
| Tier 3 | 5,000 | 4,000,000 | 40,000,000 |
| Tier 4 | 10,000 | 10,000,000 | 1,000,000,000 |
| Tier 5 | 30,000 | 180,000,000 | 15,000,000,000 |
③ Batch API 基本流程
Batch API 用于异步处理大量请求,按 Standard 费率的 50% 计费,24 小时内完成。流程分三步:提交 → 轮询 → 取结果。
准备批量请求文件
Batch API 要求先上传一个 JSONL 文件,每行一个请求:
import json
构造批量请求数据
requests = [
{
"custom_id": "task-001",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "gpt-6-luna",
"messages": [
{"role": "user", "content": "请用一句话总结:人工智能在医疗领域的应用"}
],
"max_completion_tokens": 200
}
},
{
"custom_id": "task-002",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "gpt-6-luna",
"messages": [
{"role": "user", "content": "请用一句话总结:区块链技术的核心思想"}
],
"max_completion_tokens": 200
}
},
{
"custom_id": "task-003",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "gpt-6-luna",
"messages": [
{"role": "user", "content": "请用一句话总结:边缘计算与云计算的区别"}
],
"max_completion_tokens": 200
}
}
]
写入 JSONL 文件
with open("batch_requests.jsonl", "w", encoding="utf-8") as f:
for req in requests:
f.write(json.dumps(req, ensure_ascii=False) + "\n")
print(f"已生成 {len(requests)} 条请求")
上传文件并创建批量任务
# 1. 上传 JSONL 文件
batch_file = client.files.create(
file=open("batch_requests.jsonl", "rb"),
purpose="batch"
)
2. 创建批量任务
batch_job = client.batches.create(
input_file_id=batch_file.id,
endpoint="/v1/chat/completions",
completion_window="24h"
)
print(f"批量任务 ID: {batch_job.id}")
print(f"状态: {batch_job.status}")
④ 轮询任务状态与获取结果
轮询状态
import time
batch_id = batch_job.id
while True:
batch = client.batches.retrieve(batch_id)
print(f"状态: {batch.status} | "
f"已完成: {batch.request_counts.completed} | "
f"失败: {batch.request_counts.failed} | "
f"总计: {batch.request_counts.total}")
if batch.status in ("completed", "failed", "cancelled", "expired"):
break
time.sleep(30) # 每 30 秒检查一次
下载并解析结果
if batch.status == "completed":
# 下载结果文件
result_content = client.files.content(batch.output_file_id)
result_text = result_content.text
# 逐行解析 JSONL
for line in result_text.strip().split("\n"):
result = json.loads(line)
custom_id = result["custom_id"]
response = result["response"]
if response["status_code"] == 200:
content = response["body"]["choices"][0]["message"]["content"]
print(f"[{custom_id}] {content}")
else:
print(f"[{custom_id}] 错误: {response}")
# 如果有失败记录,下载错误文件
if batch.error_file_id:
error_content = client.files.content(batch.error_file_id)
print(f"错误详情:\n{error_content.text}")
Batch 输入和结果文件保留 30 天,过期自动删除。
⑤ 推理强度配置
GPT-6 Luna 支持 6 档推理强度,Batch 请求中通过 reasoning_effort 参数控制:
# 构造不同推理强度的请求
requests = [
{
"custom_id": f"task-{i:03d}",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "gpt-6-luna",
"messages": [
{"role": "user", "content": f"分析以下代码的时间复杂度并说明理由:{code_snippet}"}
],
"reasoning_effort": effort, # none / low / medium / high / xhigh / max
"max_completion_tokens": 1024
}
}
for i, (effort, code_snippet) in enumerate([
("none", "def f(n): return n * 2"),
("low", "def f(n): return sum(range(n))"),
("medium", "def f(n): return [i*j for i in range(n) for j in range(n)]"),
("high", "def f(n): return sorted([randint(0,n) for _ in range(n*n)])"),
])
]
with open("batch_reasoning.jsonl", "w", encoding="utf-8") as f:
for req in requests:
f.write(json.dumps(req, ensure_ascii=False) + "\n")
推理强度选择参考
| effort | 适用场景 | token 消耗 |
|---|---|---|
none |
简单分类、格式转换、提取 | 最低 |
low |
基础问答、短文本摘要 | 低 |
medium |
默认值,通用任务 | 中 |
high |
代码分析、多步推理 | 高 |
xhigh |
复杂逻辑、数学证明 | 较高 |
max |
最深推理,耗时最长 | 最高 |
注意 :Chat Completions API 中使用 Function Calling 时,
reasoning_effort只能设为none。需要同时使用工具和推理时,改用 Responses API。
⑥ 多模态批量请求
Batch 请求同样支持图像输入,将图片转为 base64 后放入消息体:
import base64
def image_to_data_url(image_path):
with open(image_path, "rb") as f:
b64 = base64.b64encode(f.read()).decode("utf-8")
return f"data:image/jpeg;base64,{b64}"
批量图片分类请求
image_files = ["img1.jpg", "img2.jpg", "img3.jpg"]
requests = []
for i, img_path in enumerate(image_files):
data_url = image_to_data_url(img_path)
requests.append({
"custom_id": f"img-{i:03d}",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "gpt-6-luna",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "请用 JSON 格式输出这张图片的类别和置信度,格式: {"category": "...", "confidence": 0.0}"},
{"type": "image_url", "image_url": {"url": data_url}}
]
}],
"max_completion_tokens": 200,
"response_format": {"type": "json_object"}
}
})
with open("batch_vision.jsonl", "w", encoding="utf-8") as f:
for req in requests:
f.write(json.dumps(req, ensure_ascii=False) + "\n")
⑦ 结构化输出
通过 response_format 约束输出为 JSON Schema:
request = {
"custom_id": "extract-001",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "gpt-6-luna",
"messages": [
{"role": "user", "content": "从以下文本提取人名、日期和金额:2026年9月22日,张三向李四转账5000元。"}
],
"max_completion_tokens": 256,
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "extraction",
"strict": True,
"schema": {
"type": "object",
"properties": {
"people": {"type": "array", "items": {"type": "string"}},
"date": {"type": "string"},
"amount": {"type": "number"}
},
"required": ["people", "date", "amount"]
}
}
}
}
}
⑧ 计费规则与成本估算
Batch 定价(短上下文 ≤ 272K tokens,按 Standard 50% 计)
| 计费项 | Batch 单价(/百万 tokens) |
|---|---|
| 输入 | $0.05 |
| 缓存命中输入 | $0.005 |
| 缓存写入 | $0.0625 |
| 输出 | $0.25 |
长上下文阶梯(单请求输入 > 272K tokens)
| 计费项 | Batch 长上下文单价 |
|---|---|
| 输入 | $0.10 |
| 缓存命中输入 | $0.01 |
| 缓存写入 | $0.125 |
| 输出 | $0.375 |
阶梯按单次请求的输入 token 数判定,不是 batch 内所有请求的合计。
成本估算示例
假设 1000 条请求,每条输入 2000 tokens、输出 300 tokens,无缓存:
num_requests = 1000
input_tokens_per_req = 2000
output_tokens_per_req = 300
total_input = num_requests * input_tokens_per_req # 2,000,000
total_output = num_requests * output_tokens_per_req # 300,000
Batch 短上下文
input_cost = (total_input / 1_000_000) * 0.05 # $0.10
output_cost = (total_output / 1_000_000) * 0.25 # $0.075
total_cost = input_cost + output_cost
print(f"输入费用: ${input_cost:.4f}")
print(f"输出费用: ${output_cost:.4f}")
print(f"总费用: ${total_cost:.4f}")
输入费用: $0.1000
输出费用: $0.0750
总费用: $0.1750
提示词缓存
缓存自动生效,无需改代码。命中条件:
- 前缀 ≥ 1024 tokens
- TTL 5--10 分钟(最长 1 小时)
- 缓存命中按 $0.005/M 计费(Batch 短上下文)
适合重复使用相同系统提示词的场景:
# 所有请求共享相同的系统 prompt(长前缀),只有 user 内容不同
system_prompt = "你是一个专业的文本分类助手。请将输入文本分类到以下类别之一:科技、财经、体育、娱乐、教育、健康。只输出类别名称,不要解释。" * 20 # 确保超过 1024 tokens
requests = [
{
"custom_id": f"cls-{i:04d}",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "gpt-6-luna",
"messages": [
{"role": "system", "content": system_prompt},
{"role": "user", "content": text}
],
"max_completion_tokens": 10,
"reasoning_effort": "none"
}
}
for i, text in enumerate(texts_to_classify)
]
首次请求写入缓存,后续命中缓存,输入费用降至 1/10
⑨ 任务管理与错误处理
取消未完成的任务
# 取消批量任务
cancelled = client.batches.cancel(batch_id)
print(f"取消后状态: {cancelled.status}")
列出历史任务
# 列出最近的批量任务
batches = client.batches.list(limit=10)
for b in batches.data:
print(f"ID: {b.id} | 状态: {b.status} | "
f"完成: {b.request_counts.completed}/{b.request_counts.total} | "
f"创建时间: {b.created_at}")
常见错误排查
| 错误 | 原因 | 解决方案 |
|---|---|---|
invalid_api_key |
API Key 错误 | 检查环境变量 |
file_too_large |
JSONL 文件超限 | 拆分为多个 batch |
rate_limit_exceeded |
超出 Batch 队列上限 | 升级 tier 或分批提交 |
batch_expired |
超过 24h 窗口未完成 | 检查队列负载,减少单次请求量 |
invalid_request |
请求体格式错误 | 校验 JSONL 每行的 body 结构 |
结果校验
# 下载结果后逐条校验
results = []
for line in result_text.strip().split("\n"):
entry = json.loads(line)
if entry["response"]["status_code"] == 200:
results.append(entry)
else:
# 记录失败请求,后续重试
print(f"失败: {entry['custom_id']} - {entry['response']}")
校验 JSON 结构化输出
for r in results:
try:
data = json.loads(r["response"]["body"]["choices"][0]["message"]["content"])
except json.JSONDecodeError:
print(f"JSON 解析失败: {r['custom_id']}")
⑩ 完整示例:批量文本分类
将以上步骤整合为一个完整的批量分类流程:
import os
import json
import time
from openai import OpenAI
client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
1. 准备数据
texts = [
"苹果发布新款 MacBook Pro,搭载 M5 芯片",
"美联储宣布降息 50 个基点",
"中国队获得乒乓球世锦赛团体冠军",
"某明星宣布退出娱乐圈",
"教育部发布新一轮课程改革方案",
"研究发现每天步行 8000 步可显著降低心血管风险",
]
system_prompt = "你是一个新闻分类助手。请将输入文本分类到以下类别之一:科技、财经、体育、娱乐、教育、健康。只输出类别名称。"
2. 构造 JSONL
requests = [
{
"custom_id": f"news-{i:03d}",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "gpt-6-luna",
"messages": [
{"role": "system", "content": system_prompt},
{"role": "user", "content": text}
],
"max_completion_tokens": 10,
"reasoning_effort": "none"
}
}
for i, text in enumerate(texts)
]
with open("classify_batch.jsonl", "w", encoding="utf-8") as f:
for req in requests:
f.write(json.dumps(req, ensure_ascii=False) + "\n")
3. 上传 + 创建任务
batch_file = client.files.create(
file=open("classify_batch.jsonl", "rb"),
purpose="batch"
)
batch_job = client.batches.create(
input_file_id=batch_file.id,
endpoint="/v1/chat/completions",
completion_window="24h"
)
print(f"任务已提交: {batch_job.id}")
4. 轮询
while True:
batch = client.batches.retrieve(batch_job.id)
print(f"[{batch.status}] {batch.request_counts.completed}/{batch.request_counts.total}")
if batch.status in ("completed", "failed", "cancelled", "expired"):
break
time.sleep(15)
5. 输出结果
if batch.status == "completed":
result_text = client.files.content(batch.output_file_id).text
for line in result_text.strip().split("\n"):
entry = json.loads(line)
content = entry["response"]["body"]["choices"][0]["message"]["content"]
print(f"{entry['custom_id']}: {content}")
参考文档
- 模型文档:
developers.openai.com/api/docs/models/gpt-6-luna - Batch API:
developers.openai.com/api/docs/batch - 定价:
developers.openai.com/api/docs/pricing - GPT-6 使用指南:
developers.openai.com/api/docs/guides/latest-model
以上内容基于 2026 年 9 月的 API 版本整理,具体参数以官方文档为准。