为什么 vLLM 要支持自定义对话模板?
把大语言模型接入线上服务时,我们习惯以 system、user、assistant 这样的结构发送消息:
json
[
{ "role": "system", "content": "你是一个专业的助手。" },
{ "role": "user", "content": "请解释什么是 KV Cache。" }
]
但对模型而言,这些结构化字段并不能直接被理解。模型最终接收到的,始终是一段经过分词后的文本和 token 序列。问题在于,不同模型在训练时使用的"对话文本格式"并不相同:有的使用 ChatML,有的使用 [INST]...[/INST],有的使用 Llama 3 风格的角色标记,也有模型采用自己定义的特殊 token。
这正是对话模板(Chat Template)存在的意义。它负责将接口层传入的标准消息,转换为目标模型在训练阶段最熟悉的 prompt 格式。例如,同样一句用户问题,可能需要被拼接为:
text
<|system|>
你是一个专业的助手。
<|user|>
请解释什么是 KV Cache。
<|assistant|>
也可能需要采用完全不同的格式:
text
<s>[INST] <<SYS>>
你是一个专业的助手。
<</SYS>>
请解释什么是 KV Cache。 [/INST]
比如说如果我们使用LLaMA Factory进行微调,LLaMA Factory微调使用的模板是自己写好的,在下面的代码里面
但是使用vLLM这些推理框架的时候,他们使用的是模型配置文件里面的模板

如果模板与模型训练时的格式不匹配,轻则回答质量下降,重则出现角色混乱、系统提示词失效、重复输出标签、多轮对话错位,甚至工具调用无法正常工作。
因此,vLLM 提供自定义对话模板的能力,并不只是为了"灵活配置提示词"。更重要的是,它让服务端能够准确适配不同模型、不同微调数据格式,以及工具调用、多模态等更复杂的推理场景。理解这一点,是正确部署和使用 vLLM 对话服务的第一步。
导出LLaMA Factory的模板
在LLaMA Factory的template.py文件中有很多内部方法可以生成模板,我们可以使用这些方法来生成我们的模板

创建导出文件
在LLaMA-Factory/src/llamafactory目录下创建一个export_template.py的文件
内容如下
py
import sys
import os
# 将项目根目录添加到Python路径
root_dir = os.path.dirname(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))))
sys.path.append(root_dir)
from llamafactory.data.template import TEMPLATES
from transformers import AutoTokenizer
# 初始化分词器(任意支持的分词器均可)
tokenizer = AutoTokenizer.from_pretrained("/home/gillbert/code/hugging_face_test/modelscope_test/llm/models/Qwen--Qwen3.5-2B/snapshots/master")
# 获取模板对象
template_name = "qwen3" # 该名称是LLaMA Factory template.py中有的模型名称
template = TEMPLATES[template_name]
# 修复分词器的Jinja模板
template.fix_jinja_template(tokenizer)
# 输出模板的Jinja格式
print("=" * 60)
print(tokenizer.chat_template)
执行该脚本,会得到如下输出
sh
(llamafactory) ➜ data git:(main) ✗ python export_template.py
============================================================
{%- set image_count = namespace(value=0) %}
{%- set video_count = namespace(value=0) %}
{%- macro render_content(content, do_vision_count, is_system_content=false) %}
{%- if content is string %}
{{- content }}
{%- elif content is iterable and content is not mapping %}
{%- for item in content %}
{%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
{%- if is_system_content %}
{{- raise_exception('System message cannot contain images.') }}
{%- endif %}
{%- if do_vision_count %}
{%- set image_count.value = image_count.value + 1 %}
{%- endif %}
{%- if add_vision_id %}
{{- 'Picture ' ~ image_count.value ~ ': ' }}
{%- endif %}
{{- '<|vision_start|><|image_pad|><|vision_end|>' }}
{%- elif 'video' in item or item.type == 'video' %}
{%- if is_system_content %}
{{- raise_exception('System message cannot contain videos.') }}
{%- endif %}
{%- if do_vision_count %}
{%- set video_count.value = video_count.value + 1 %}
{%- endif %}
{%- if add_vision_id %}
{{- 'Video ' ~ video_count.value ~ ': ' }}
{%- endif %}
{{- '<|vision_start|><|video_pad|><|vision_end|>' }}
{%- elif 'text' in item %}
{{- item.text }}
{%- else %}
{{- raise_exception('Unexpected item type in content.') }}
{%- endif %}
{%- endfor %}
{%- elif content is none or content is undefined %}
{{- '' }}
{%- else %}
{{- raise_exception('Unexpected content type.') }}
{%- endif %}
{%- endmacro %}
{%- if not messages %}
{{- raise_exception('No messages provided.') }}
{%- endif %}
{%- if tools and tools is iterable and tools is not mapping %}
{{- '<|im_start|>system\n' }}
{{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n</tools>" }}
{{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
{%- if messages[0].role == 'system' %}
{%- set content = render_content(messages[0].content, false, true)|trim %}
{%- if content %}
{{- '\n\n' + content }}
{%- endif %}
{%- endif %}
{{- '<|im_end|>\n' }}
{%- else %}
{%- if messages[0].role == 'system' %}
{%- set content = render_content(messages[0].content, false, true)|trim %}
{{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
{%- for message in messages[::-1] %}
{%- set index = (messages|length - 1) - loop.index0 %}
{%- if ns.multi_step_tool and message.role == "user" %}
{%- set content = render_content(message.content, false)|trim %}
{%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
{%- set ns.multi_step_tool = false %}
{%- set ns.last_query_index = index %}
{%- endif %}
{%- endif %}
{%- endfor %}
{%- if ns.multi_step_tool %}
{{- raise_exception('No user query found in messages.') }}
{%- endif %}
{%- for message in messages %}
{%- set content = render_content(message.content, true)|trim %}
{%- if message.role == "system" %}
{%- if not loop.first %}
{{- raise_exception('System message must be at the beginning.') }}
{%- endif %}
{%- elif message.role == "user" %}
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
{%- elif message.role == "assistant" %}
{%- set reasoning_content = '' %}
{%- if message.reasoning_content is string %}
{%- set reasoning_content = message.reasoning_content %}
{%- else %}
{%- if '</think>' in content %}
{%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
{%- set content = content.split('</think>')[-1].lstrip('\n') %}
{%- endif %}
{%- endif %}
{%- set reasoning_content = reasoning_content|trim %}
{%- if loop.index0 > ns.last_query_index %}
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
{%- else %}
{{- '<|im_start|>' + message.role + '\n' + content }}
{%- endif %}
{%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
{%- for tool_call in message.tool_calls %}
{%- if tool_call.function is defined %}
{%- set tool_call = tool_call.function %}
{%- endif %}
{%- if loop.first %}
{%- if content|trim %}
{{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
{%- else %}
{{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
{%- endif %}
{%- else %}
{{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
{%- endif %}
{%- if tool_call.arguments is defined %}
{%- for args_name, args_value in tool_call.arguments|items %}
{{- '<parameter=' + args_name + '>\n' }}
{%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
{{- args_value }}
{{- '\n</parameter>\n' }}
{%- endfor %}
{%- endif %}
{{- '</function>\n</tool_call>' }}
{%- endfor %}
{%- endif %}
{{- '<|im_end|>\n' }}
{%- elif message.role == "tool" %}
{%- if loop.previtem and loop.previtem.role != "tool" %}
{{- '<|im_start|>user' }}
{%- endif %}
{{- '\n<tool_response>\n' }}
{{- content }}
{{- '\n</tool_response>' }}
{%- if not loop.last and loop.nextitem.role != "tool" %}
{{- '<|im_end|>\n' }}
{%- elif loop.last %}
{{- '<|im_end|>\n' }}
{%- endif %}
{%- else %}
{{- raise_exception('Unexpected message role.') }}
{%- endif %}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n' }}
{%- if enable_thinking is defined and enable_thinking is true %}
{{- '<think>\n' }}
{%- else %}
{{- '<think>\n\n</think>\n\n' }}
{%- endif %}
{%- endif %}
将刚才生成的jinja模板保存到文件qwen.jinja中,我们可以在vLLM启动模型的时候使用该对话模板
sh
vllm serve /home/gillbert/code/vllm_test/llm/models/Qwen--Qwen3.5-0.8B/snapshots/master \
--tensor-parallel-size 1 \
--gpu-memory-utilization 0.9 \
--max-model-len 8192 \
--host 0.0.0.0 \
--port 8000 \
--api-key 123456 \
--enable-auto-tool-choice \
--tool-call-parser hermes \
--chat-template ./qwen.jinja
我们可以使用Open WebUI测试一下
