LLM 思考模式的开启、关闭与响应判断
前言
现在越来越多的大模型支持思考模式。同一个模型可能既可以直接回答问题,也可以先进行推理,再生成最终答案。
不同平台对思考模式的控制参数并不完全一致。例如,阿里云百炼使用 enable_thinking,火山方舟使用 thinking.type,本地模型则可能通过 chat_template_kwargs 控制。
本文通过阿里云百炼、火山方舟、DeepSeek、Kimi 和本地部署模型的调用示例,介绍如何开启和关闭思考模式,以及如何从响应中判断模型是否返回了思考内容。
文中的
123只是示例占位符,不能作为真实 API Key 使用。实际开发时请通过环境变量读取密钥。
一、思考模式与响应判断
什么是思考模式
支持思考能力的模型,通常会先生成一段内部推理,再生成最终回答。不同平台对这部分内容的命名不完全一致,常见字段包括:
| 概念 | 常见字段或参数 |
|---|---|
| 开启/关闭思考 | enable_thinking、thinking |
| 思考文本 | reasoning_content |
| 思考 Token 数 | reasoning_tokens |
| 最终回答 | content |
需要注意:**请求参数只能表示期望的模式,最终是否返回思考内容仍应以响应为准。**有些模型会进行内部推理,但不会向客户端公开完整思考文本。
三类模型
从默认行为和是否允许切换来看,可以把模型分成三类:
默认开启、支持切换
这类模型不传参数时默认进入思考模式,但可以在请求中关闭思考。部分 Qwen、DeepSeek、GLM、豆包和 Kimi 模型属于这一类。
python
extra_body={"enable_thinking": True} # 开启
extra_body={"enable_thinking": False} # 关闭
默认关闭、支持切换
这类模型不传参数时直接生成最终回答,但可以通过请求参数主动开启思考。部分商业版 Qwen 模型属于这一类。
仅支持思考、不能关闭
Thinking 专用模型可能不支持关闭思考。即使传入关闭参数,平台也可能忽略参数或直接返回错误。因此使用前应查看对应模型的官方参数说明。
如何判断 response 中是否有思考内容
OpenAI Python SDK 的 message 对象不一定把第三方字段声明成标准属性,但百炼、火山等兼容接口返回的 reasoning_content 通常可以通过 getattr 读取。
python
def parse_message(response):
"""提取最终回答和思考内容。"""
message = response.choices[0].message
content = getattr(message, "content", None)
reasoning_content = getattr(message, "reasoning_content", None)
has_reasoning = (
isinstance(reasoning_content, str)
and bool(reasoning_content.strip())
)
return {
"content": content,
"reasoning_content": reasoning_content,
"has_reasoning": has_reasoning,
}
调用:
python
result = parse_message(response)
if result["has_reasoning"]:
print("模型返回了思考内容:")
print(result["reasoning_content"])
else:
print("没有返回有效的思考内容")
print("最终回答:")
print(result["content"])
如果还要区分"字段不存在"和"字段为空",可以使用 model_dump:
python
message_dict = response.choices[0].message.model_dump(exclude_none=False)
if "reasoning_content" not in message_dict:
print("响应中没有 reasoning_content 字段")
elif message_dict["reasoning_content"]:
print("响应中包含思考内容")
else:
print("字段存在,但思考内容为空")
二、云平台调用
云平台模型范围总结
从本文当前整理的官方 API 调用方式来看,四个平台的模型覆盖范围有所不同:
| 平台 | 模型覆盖特点 | 本文代码示例 |
|---|---|---|
| 阿里云百炼 | 模型种类最多,既有 Qwen 系列,也接入了 DeepSeek、GLM、Kimi 等多类模型 | Qwen |
| 火山方舟 | 以豆包模型为主,同时提供 DeepSeek、GLM 等主流模型 | 豆包 |
| DeepSeek | 官方 API 主要面向 DeepSeek 自有模型,接口和参数相对专一 | DeepSeek |
| Kimi | 官方 API 主要面向 Kimi 自有模型,接口和参数相对专一 | Kimi |
简单来说:
- 阿里云百炼更像一个综合模型平台,模型选择最多;
- 火山方舟以豆包为核心,同时提供部分其他主流模型;
- DeepSeek 官方 API 主要服务 DeepSeek 自有模型;
- Kimi 官方 API 主要服务 Kimi 自有模型。
模型列表和平台能力会持续变化,以上内容是本文整理时的情况,实际调用前应以各平台模型广场和官方文档为准。
2.1 阿里云百炼
官方入口
- 阿里云百炼官网:https://www.aliyun.com/product/bailian
- 阿里云百炼控制台:https://bailian.console.aliyun.com/
- 百炼模型广场:https://bailian.console.aliyun.com/cn-beijing/?tab=model#/model-market/all
- 百炼官方文档:https://help.aliyun.com/zh/model-studio/
模型的默认思考状态
阿里云百炼中的模型,不能只按"是否支持思考"来理解,还要区分三种情况:默认开启、默认关闭、只能思考不能关闭。
默认开启思考
不传 enable_thinking 时,这些模型通常默认进入思考模式:
| 系列 | 典型模型 | 默认状态 |
|---|---|---|
| Qwen3.8 | qwen3.8-max、qwen3.8-flash、qwen3.8-27b、qwen3.8-omni-flash |
✅ 开启 |
| Qwen3.7 | qwen3.7-max、qwen3.7-plus、qwen3.7-flash |
✅ 开启 |
| Qwen3.6 | qwen3.6-max-preview、qwen3.6-plus、qwen3.6-flash |
✅ 开启 |
| Qwen3.5 | qwen3.5-plus、qwen3.5-flash、qwen3.5-397b-a17b、qwen3.5-27b 等 |
✅ 开启 |
| Qwen3 开源版 | qwen3-235b-a22b、qwen3-32b、qwen3-30b-a3b、qwen3-14b、qwen3-8b |
✅ 开启 |
默认关闭思考
这些商业模型不传参数时通常直接生成最终答案:
| 系列 | 典型模型 | 默认状态 |
|---|---|---|
| Qwen3 商业 Max | qwen3-max、qwen3-max-preview |
❌ 关闭 |
| Qwen3 商业 Plus | qwen-plus、qwen-plus-latest |
❌ 关闭 |
| Qwen3 商业 Flash | qwen-flash |
❌ 关闭 |
| Qwen3 商业 Turbo | qwen-turbo |
❌ 关闭 |
只能思考,不能关闭
Thinking 专用模型始终以思考模式运行,传入关闭参数也不能切换成普通模式:
| 模型 | 默认状态 | 是否支持关闭 |
|---|---|---|
qwen3.8-2.4t-a95b |
✅ 开启 | ❌ 不支持 |
qwen3.7-max-preview |
✅ 开启 | ❌ 不支持 |
qwen3-next-80b-a3b-thinking |
✅ 开启 | ❌ 不支持 |
qwen3-235b-a22b-thinking-2507 |
✅ 开启 | ❌ 不支持 |
qwen3-30b-a3b-thinking-2507 |
✅ 开启 | ❌ 不支持 |
qwq-plus |
✅ 开启 | ❌ 不支持 |
模型列表和默认状态可能随平台版本变化,发布文章或上线代码前,应再次核对百炼官方模型说明。
开启和关闭思考到底有什么区别
这里要区分"默认状态"和"主动设置":
| 调用方式 | 含义 |
|---|---|
不传enable_thinking |
使用模型自己的默认状态;默认开启的模型会思考,默认关闭的模型不会思考 |
enable_thinking=True |
对支持切换的模型,主动要求本次调用开启思考 |
enable_thinking=False |
对支持切换的模型,主动要求本次调用关闭思考 |
Thinking 专用模型传False |
通常不能关闭,可能被忽略或返回参数错误 |
Python SDK 中的写法是:
python
# 主动开启本次请求的思考模式
extra_body={"enable_thinking": True}
# 主动关闭本次请求的思考模式
extra_body={"enable_thinking": False}
也就是说,True/False 是本次请求的控制开关 ,而"默认开启/默认关闭"是模型不传参数时的初始行为 。最终是否返回思考内容,仍然要检查响应中的 reasoning_content。
OpenAI 兼容 SDK
使用 OpenAI Python SDK 时,百炼的非标准参数要放到 extra_body 中。
python
import asyncio
import os
from dotenv import load_dotenv
from openai import AsyncOpenAI
load_dotenv()
client = AsyncOpenAI(
api_key=os.getenv("DASHSCOPE_API_KEY", "123"),
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)
def parse_message(response):
"""从百炼响应中提取最终回答和思考内容。"""
message = response.choices[0].message
content = getattr(message, "content", None)
reasoning_content = getattr(message, "reasoning_content", None)
return {
"content": content,
"reasoning_content": reasoning_content,
"has_reasoning": (
isinstance(reasoning_content, str)
and bool(reasoning_content.strip())
),
}
async def call_qwen(enable_thinking: bool):
response = await client.chat.completions.create(
model="qwen3.6-plus",
messages=[{"role": "user", "content": "请计算 12345 × 67890,并说明过程"}],
# True:开启思考;False:关闭思考
extra_body={"enable_thinking": enable_thinking},
)
result = parse_message(response)
print("最终回答:")
print(result["content"])
print("=" * 150)
if result["has_reasoning"]:
print("思考内容:")
print(result["reasoning_content"])
else:
print("没有返回思考内容")
async def main():
await call_qwen(enable_thinking=True)
# await call_qwen(enable_thinking=False)
if __name__ == "__main__":
asyncio.run(main())
HTTP 调用
HTTP 请求不需要 extra_body,可以直接把 enable_thinking 放在请求体中。
开启思考
bash
curl --location 'https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions' \
--header 'Authorization: Bearer 123' \
--header 'Content-Type: application/json' \
--data '{
"model": "qwen3.8-max",
"messages": [{"role": "user", "content": "你是谁?"}],
"enable_thinking": true
}'
关闭思考
bash
curl --location 'https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions' \
--header 'Authorization: Bearer 123' \
--header 'Content-Type: application/json' \
--data '{
"model": "qwen3.8-max",
"messages": [{"role": "user", "content": "你是谁?"}],
"enable_thinking": false
}'
2.2 火山方舟
官方入口与本文覆盖范围
常用入口:
- 火山方舟控制台:https://ark.volcengine.com/
- 模型广场:https://console.volcengine.com/ark/region:cn-beijing/model?view=DEFAULT_VIEW&groupType=ModelGroups
- 官方文档:https://docs.volcengine.com/docs/ark?lang=zh
- OpenAI SDK 兼容调用说明:https://docs.volcengine.com/docs/ark/compatible-with-openai-sdk?lang=en
火山方舟平台搭载的模型会随时间更新。本篇笔记说明豆包、DeepSeek、GLM 三类主流模型的覆盖情况,但 SDK 代码只以豆包为例:
| 模型类别 | 平台模型范围 | 本文代码示例 |
|---|---|---|
| 豆包(Doubao) | 方舟核心模型,使用doubao-* 模型 ID |
✅ 有 |
| DeepSeek | 使用模型广场中可用的 DeepSeek 模型 ID | ❌ 无,仅作范围说明 |
| GLM | 使用模型广场中可用的 GLM 模型 ID | ❌ 无,仅作范围说明 |
| Kimi | 本文暂不纳入火山方舟的模型范围 | ❌ 无 |
| Qwen3.8 | 本文暂不纳入火山方舟的模型范围 | ❌ 无 |
因此,下面的火山方舟 SDK 示例只演示豆包模型的思考开关。调用其他模型时,需要先在模型广场确认模型 ID、地域、计费方式和参数支持情况。
思考模式参数
火山方舟使用 thinking.type 控制思考模式:
| 目标 | 参数 |
|---|---|
| 强制开启 | {"thinking": {"type": "enabled"}} |
| 强制关闭 | {"thinking": {"type": "disabled"}} |
| 自动决定 | {"thinking": {"type": "auto"}} |
Python SDK 中的精简写法是:
python
# 主动开启本次请求的思考模式
extra_body={"thinking": {"type": "enabled"}}
# 主动关闭本次请求的思考模式
extra_body={"thinking": {"type": "disabled"}}
OpenAI 兼容 SDK 调用
python
import asyncio
import os
from openai import AsyncOpenAI
client = AsyncOpenAI(
api_key=os.getenv("ARK_API_KEY", "123"),
base_url="https://ark.cn-beijing.volces.com/api/v3",
)
def parse_message(response):
"""从火山方舟响应中提取最终回答和思考内容。"""
message = response.choices[0].message
content = getattr(message, "content", None)
reasoning_content = getattr(message, "reasoning_content", None)
return {
"content": content,
"reasoning_content": reasoning_content,
"has_reasoning": (
isinstance(reasoning_content, str)
and bool(reasoning_content.strip())
),
}
async def call_doubao(enable_thinking: bool):
thinking_type = "enabled" if enable_thinking else "disabled"
response = await client.chat.completions.create(
model="doubao-seed-2-1-pro-260915",
messages=[{"role": "user", "content": "请分析 9.9 和 9.11 哪个更大"}],
# True:enabled;False:disabled
extra_body={"thinking": {"type": thinking_type}},
)
result = parse_message(response)
print("最终回答:")
print(result["content"])
print("=" * 150)
if result["has_reasoning"]:
print("思考内容:")
print(result["reasoning_content"])
else:
print("没有返回思考内容")
async def main():
await call_doubao(enable_thinking=True)
# await call_doubao(enable_thinking=False)
if __name__ == "__main__":
asyncio.run(main())
HTTP 调用
HTTP 请求中不需要 extra_body,直接把 thinking 放在请求体顶层。
开启思考
bash
curl --location 'https://ark.cn-beijing.volces.com/api/v3/chat/completions' \
--header 'Authorization: Bearer 123' \
--header 'Content-Type: application/json' \
--data '{
"model": "doubao-seed-2-1-pro-260915",
"messages": [
{"role": "user", "content": "请分析 9.9 和 9.11 哪个更大"}
],
"thinking": {"type": "enabled"}
}'
关闭思考
bash
curl --location 'https://ark.cn-beijing.volces.com/api/v3/chat/completions' \
--header 'Authorization: Bearer 123' \
--header 'Content-Type: application/json' \
--data '{
"model": "doubao-seed-2-1-pro-260915",
"messages": [
{"role": "user", "content": "请分析 9.9 和 9.11 哪个更大"}
],
"thinking": {"type": "disabled"}
}'
2.3 DeepSeek
官方入口与本文范围
- DeepSeek 官网:https://deepseek.com/
- DeepSeek API 平台:https://www.deepseek.com/platform/
- DeepSeek API 文档:https://api-docs.deepseek.com/api/create-chat-completion/
本节只讨论 DeepSeek 官方 API 中的 DeepSeek 模型,不混用阿里云百炼或火山方舟上的同名模型。当前示例按"模型默认开启思考"来展示:不额外传关闭参数时,模型使用自己的默认思考行为;是否返回可见的 reasoning_content,仍以实际响应为准。
Python SDK 中的精简写法是:
python
# 主动开启本次请求的思考模式
extra_body={"thinking": {"type": "enabled"}}
# 主动关闭本次请求的思考模式
extra_body={"thinking": {"type": "disabled"}}
SDK 调用
下面示例将响应解析函数一并写出,复制代码块即可运行。
python
import asyncio
import os
from openai import AsyncOpenAI
client = AsyncOpenAI(
api_key=os.getenv("DEEPSEEK_API_KEY", "123"),
base_url="https://api.deepseek.com",
)
def parse_message(response):
"""从 DeepSeek 响应中提取最终回答和思考内容。"""
message = response.choices[0].message
content = getattr(message, "content", None)
reasoning_content = getattr(message, "reasoning_content", None)
return {
"content": content,
"reasoning_content": reasoning_content,
"has_reasoning": (
isinstance(reasoning_content, str)
and bool(reasoning_content.strip())
),
}
async def main():
response = await client.chat.completions.create(
model="deepseek-flash",
messages=[{"role": "user", "content": "请解释什么是思考模型"}],
# 不传 thinking 参数,使用 DeepSeek 模型的默认思考模式
)
result = parse_message(response)
print("最终回答:")
print(result["content"])
print("=" * 150)
if result["has_reasoning"]:
print("思考内容:")
print(result["reasoning_content"])
else:
print("没有返回思考内容")
if __name__ == "__main__":
asyncio.run(main())
HTTP 调用
DeepSeek 的 thinking.type 支持 enabled 和 disabled,默认值为 enabled。
默认模式
不传 thinking 时,使用模型默认的思考模式:
bash
curl --location 'https://api.deepseek.com/chat/completions' \
--header 'Authorization: Bearer 123' \
--header 'Content-Type: application/json' \
--data '{
"model": "deepseek-flash",
"messages": [
{"role": "user", "content": "请解释什么是思考模型"}
]
}'
关闭思考
bash
curl --location 'https://api.deepseek.com/chat/completions' \
--header 'Authorization: Bearer 123' \
--header 'Content-Type: application/json' \
--data '{
"model": "deepseek-flash",
"messages": [
{"role": "user", "content": "请解释什么是思考模型"}
],
"thinking": {"type": "disabled"}
}'
开启思考
bash
curl --location 'https://api.deepseek.com/chat/completions' \
--header 'Authorization: Bearer 123' \
--header 'Content-Type: application/json' \
--data '{
"model": "deepseek-flash",
"messages": [
{"role": "user", "content": "请解释什么是思考模型"}
],
"thinking": {"type": "enabled"}
}'
2.4 Kimi
官方入口与本文范围
- Kimi 官网:https://www.kimi.com/
- Kimi API 开放平台:https://platform.kimi.com/
- Kimi API 文档:https://www.kimi.com/help/kimi-api/api-overview
- Kimi 模型选择:https://www.kimi.com/help/kimi-api/api-model-selection
本节只讨论 Kimi API 提供的 Kimi 模型,不混用其他平台的 Kimi 代理模型。当前示例以 kimi-k3 为例,按默认开启思考模式进行说明;如果平台或模型版本发生变化,应以 Kimi 模型文档中的默认值和参数说明为准。
Python SDK 中的精简写法是:
python
# 主动开启本次请求的思考模式
extra_body={"thinking": {"type": "enabled"}}
# 主动关闭本次请求的思考模式
extra_body={"thinking": {"type": "disabled"}}
OpenAI 兼容 SDK 调用
Kimi 的 OpenAI 兼容接口可以直接使用 AsyncOpenAI。由于 thinking 不是 OpenAI SDK 的标准参数,因此通过 extra_body 传递。
python
import asyncio
import os
from dotenv import load_dotenv
from openai import AsyncOpenAI
load_dotenv()
client = AsyncOpenAI(
api_key=os.getenv("KIMI_API_KEY", "123"),
base_url="https://api.moonshot.cn/v1",
)
def parse_message(response):
"""从 Kimi 响应中提取最终回答和思考内容。"""
message = response.choices[0].message
content = getattr(message, "content", None)
reasoning_content = getattr(message, "reasoning_content", None)
return {
"content": content,
"reasoning_content": reasoning_content,
"has_reasoning": (
isinstance(reasoning_content, str)
and bool(reasoning_content.strip())
),
}
async def call_kimi(enable_thinking: bool = True):
thinking_type = "enabled" if enable_thinking else "disabled"
response = await client.chat.completions.create(
model="kimi-k3",
messages=[
{"role": "user", "content": "请分析 9.9 和 9.11 哪个更大"}
],
# True:开启思考;False:关闭思考
extra_body={"thinking": {"type": thinking_type}},
)
result = parse_message(response)
print("最终回答:")
print(result["content"])
print("=" * 150)
if result["has_reasoning"]:
print("思考内容:")
print(result["reasoning_content"])
else:
print("没有返回思考内容")
async def main():
# Kimi 示例默认开启思考
await call_kimi(enable_thinking=True)
# await call_kimi(enable_thinking=False)
if __name__ == "__main__":
asyncio.run(main())
HTTP 调用
Kimi 的接口使用 thinking.type 表示思考状态。下面三个 HTTP 示例分别展示默认、关闭和开启,避免把三个示例误写成同一个参数。
默认模式(使用模型默认思考状态)
bash
curl --location 'https://api.moonshot.cn/v1/chat/completions' \
--header 'Authorization: Bearer 123' \
--header 'Content-Type: application/json' \
--data '{
"model": "kimi-k3",
"messages": [{"role": "user", "content": "你是谁?"}]
}'
关闭思考
bash
curl --location 'https://api.moonshot.cn/v1/chat/completions' \
--header 'Authorization: Bearer 123' \
--header 'Content-Type: application/json' \
--data '{
"model": "kimi-k3",
"messages": [{"role": "user", "content": "你是谁?"}],
"thinking": {"type": "disabled"}
}'
开启思考
bash
curl --location 'https://api.moonshot.cn/v1/chat/completions' \
--header 'Authorization: Bearer 123' \
--header 'Content-Type: application/json' \
--data '{
"model": "kimi-k3",
"messages": [{"role": "user", "content": "请分析 9.9 和 9.11 哪个更大"}],
"thinking": {"type": "enabled"}
}'
三、本地部署模型
本地部署时,思考开关通常由推理框架的聊天模板控制。它和百炼、火山方舟的云端参数不是一回事,也不是 OpenAI 标准参数。
部分本地推理服务会通过 chat_template_kwargs 把参数传给模型的聊天模板。Python SDK 中的精简写法是:
python
# 主动开启本次请求的思考模式
extra_body={
"chat_template_kwargs": {
"enable_thinking": True,
"thinking": True,
}
}
# 主动关闭本次请求的思考模式
extra_body={
"chat_template_kwargs": {
"enable_thinking": False,
"thinking": False,
}
}
这段配置可以分成三层理解:
| 层级 | 作用 |
|---|---|
extra_body |
OpenAI Python SDK 用来透传非标准参数的容器 |
chat_template_kwargs |
传递给本地推理框架聊天模板的参数集合 |
enable_thinking / thinking |
具体控制模板是否生成思考内容 |
两个字段都设置为 True 表示尝试开启思考模式;都设置为 False 表示尝试关闭思考模式。其中,一些 Qwen 模板识别 enable_thinking,一些模板或推理框架识别 thinking。
需要注意:enable_thinking 和 thinking 并不是所有本地服务都同时支持。有的服务只识别 enable_thinking,有的服务只识别 thinking,也有的服务完全依赖启动参数或服务端默认配置。因此,实际使用时应查看当前推理框架的聊天模板说明;如果确认只支持一个字段,就只保留对应字段即可。
OpenAI 兼容 SDK 调用
如果本地推理服务提供 OpenAI 兼容接口,可以继续使用 AsyncOpenAI。下面示例默认连接 http://127.0.0.1:8000/v1,请根据实际服务地址和模型名称修改。
python
import asyncio
import os
from dotenv import load_dotenv
from openai import AsyncOpenAI
load_dotenv()
client = AsyncOpenAI(
api_key=os.getenv("LOCAL_LLM_API_KEY", "123"),
base_url=os.getenv("LOCAL_LLM_BASE_URL", "http://127.0.0.1:8000/v1"),
)
def parse_message(response):
"""从本地模型响应中提取最终回答和思考内容。"""
message = response.choices[0].message
content = getattr(message, "content", None)
reasoning_content = getattr(message, "reasoning_content", None)
return {
"content": content,
"reasoning_content": reasoning_content,
"has_reasoning": (
isinstance(reasoning_content, str)
and bool(reasoning_content.strip())
),
}
async def call_local_model(enable_thinking: bool):
response = await client.chat.completions.create(
model=os.getenv("LOCAL_LLM_MODEL", "Qwen3.6-35B-A3B-AWQ"),
messages=[
{"role": "user", "content": "请分析 9.9 和 9.11 哪个更大"}
],
extra_body={
"chat_template_kwargs": {
"enable_thinking": enable_thinking,
"thinking": enable_thinking,
}
},
)
result = parse_message(response)
print("最终回答:")
print(result["content"])
print("=" * 150)
if result["has_reasoning"]:
print("思考内容:")
print(result["reasoning_content"])
else:
print("没有返回思考内容")
async def main():
await call_local_model(enable_thinking=True)
# await call_local_model(enable_thinking=False)
if __name__ == "__main__":
asyncio.run(main())
HTTP 调用
对应的 HTTP 请求示例:
bash
curl --location 'http://127.0.0.1:8000/v1/chat/completions' \
--header 'Authorization: Bearer 123' \
--header 'Content-Type: application/json' \
--data '{
"model": "Qwen3.6-35B-A3B-AWQ",
"messages": [{"role": "user", "content": "法国的首都是哪里?"}],
"chat_template_kwargs": {
"enable_thinking": false,
"thinking": false
}
}'
本地部署的参数不是 OpenAI 标准参数,必须以所使用的推理框架和聊天模板文档为准。
四、常见问题
1. enable_thinking=True,但没有 reasoning_content
可能原因包括:
- 模型不支持公开思考文本;
- 当前平台只返回思考 Token 统计;
- 参数没有按平台要求传递;
- 使用了流式调用,但只读取了
delta.content; - 模型实际忽略了该参数。
2. 如何确认模型真的产生了思考 Token
可以同时检查响应中的 usage:
python
usage = response.usage
details = getattr(usage, "completion_tokens_details", None)
reasoning_tokens = getattr(details, "reasoning_tokens", None)
print("reasoning_tokens:", reasoning_tokens)
但不同平台的 usage 字段也不完全一致,仍然应以平台文档为准。
3. 流式调用如何读取
流式响应中,思考文本一般位于 chunk.choices[0].delta.reasoning_content,最终回答位于 delta.content。需要在循环中分别累加两个字段,不能只读取 content。
五、总结
- 思考模式的控制参数由平台决定,常见形式是
enable_thinking或thinking.type。 - OpenAI 兼容 SDK 的第三方参数通常通过
extra_body传递。 - 判断是否返回思考内容,应检查
reasoning_content是否为非空字符串。 - 判断字段是否存在,应使用
model_dump()后检查键名。 - 思考参数、模型默认行为和响应字段都可能随平台或模型版本变化,发布前要以官方文档为准。
参考资料
- 阿里云百炼官网:https://www.aliyun.com/product/bailian
- 阿里云百炼深度思考模型:https://help.aliyun.com/zh/model-studio/deep-thinking
- 阿里云百炼 OpenAI 兼容 Chat Completions:https://help.aliyun.com/zh/model-studio/qwen-api-via-openai-chat-completions
- 火山方舟官方文档:https://docs.volcengine.com/docs/ark?lang=zh
- 火山方舟 OpenAI SDK 兼容调用:https://docs.volcengine.com/docs/ark/compatible-with-openai-sdk?lang=zh
- DeepSeek Chat Completions:https://api-docs.deepseek.com/zh-cn/api/create-chat-completion
- Kimi API:https://www.kimi.com/help/kimi-api/api-overview
声明
本文为个人学习笔记,内容基于撰写时的官方文档和实际测试整理。由于模型版本、平台能力和接口参数可能持续更新,具体使用方式请以各平台最新官方文档为准。如有错误,欢迎指正。