LangChain —— 多模态大模型的 prompt template

文章目录


一、如何直接将多模态数据传输给模型

 在这里,我们演示了如何将多模式输入直接传递给模型。对于其他的支持多模态输入的模型提供者,langchain 在类中提供了内在逻辑来转化为期待的格式。

 传入图像最常用的方法是将其作为字节字符串传入。这应该适用于大多数模型集成。

python 复制代码
import base64
import httpx

image_url = "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
image_data = base64.b64encode(httpx.get(image_url).content).decode("utf-8")

message = HumanMessage(
    content=[
        {"type": "text", "text": "describe the weather in this image"},
        {
            "type": "image_url",
            "image_url": {"url": f"data:image/jpeg;base64,{image_data}"},
        },
    ],
)
response = model.invoke([message])
print(response.content)

 我们可以直接在"image_URL"类型的内容块中提供图像URL。但是注意,只有一些模型提供程序支持此功能。

python 复制代码
message = HumanMessage(
    content=[
        {"type": "text", "text": "describe the weather in this image"},
        {"type": "image_url", "image_url": {"url": image_url}},
    ],
)
response = model.invoke([message])
print(response.content)

 我们也可以传多个图片。

python 复制代码
message = HumanMessage(
    content=[
        {"type": "text", "text": "are these two images the same?"},
        {"type": "image_url", "image_url": {"url": image_url}},
        {"type": "image_url", "image_url": {"url": image_url}},
    ],
)
response = model.invoke([message])
print(response.content)

二、如何使用 mutimodal prompts

 在这里,我们将描述一下怎么使用 prompt templates 来为模型格式化 multimodal imputs。

python 复制代码
import base64
import httpx

image_url = "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
image_data = base64.b64encode(httpx.get(image_url).content).decode("utf-8")

prompt = ChatPromptTemplate.from_messages(
    [
        ("system", "Describe the image provided"),
        (
            "user",
            [
                {
                    "type": "image_url",
                    "image_url": {"url": "data:image/jpeg;base64,{image_data}"},
                }
            ],
        ),
    ]
)

 我们也可以给模型传入多个图片。

python 复制代码
prompt = ChatPromptTemplate.from_messages(
    [
        ("system", "compare the two pictures provided"),
        (
            "user",
            [
                {
                    "type": "image_url",
                    "image_url": {"url": "data:image/jpeg;base64,{image_data1}"},
                },
                {
                    "type": "image_url",
                    "image_url": {"url": "data:image/jpeg;base64,{image_data2}"},
                },
            ],
        ),
    ]
)

chain = prompt | model

response = chain.invoke({"image_data1": image_data, "image_data2": image_data})
print(response.content)
相关推荐
BD_Marathon4 小时前
调用DeesSeek官网的DeepSeek模型
langchain
测试运维日常笔记8 小时前
大模型应用专项测试指南:从 Prompt、RAG 到 Agent 的完整链路
前端·prompt
91刘仁德10 小时前
RAG实战-LangChain
langchain
BD_Marathon11 小时前
调用智谱和阿里云百炼平台的大模型
langchain
梦想不只是梦与想12 小时前
LangGraph图:State、Node、Edge(一)
langchain·大模型·langgraph
王国强200915 小时前
Deep Agents Code 源码阅读(一):怎么避免在一个 6000 多行的 main.py 中迷失
langchain
YIAN17 小时前
从 SSE 流式到结构化输出:LangChain 三大 OutputParser 与 ToolCall 方案全实战
前端·langchain·node.js
YIAN17 小时前
从 SSE 流式原理到 LangChain 结构化输出:打字机效果与 JSON 解析全方案实战
前端·langchain
10年前端老司机19 小时前
Next.js+LangGraph.js+ 简历工具AI Agent完整落地
前端·langchain·agent
BreezeJiang1 天前
从 LangChain 到 LangGraph:多 Agent 不是玄学,是 token 账本和干扰问题
langchain·agent