Windows 环境下 llama.cpp 编译运行指南

在大模型落地场景中,本地轻量化部署因低延迟、高隐私性、无需依赖云端算力等优势,成为开发者与 AI 爱好者的热门需求。本文聚焦 Windows 环境,详细拆解 llama.cpp 工具的编译流程,并指导如何通过 modelscope 下载 GGUF 格式的模型,最终实现模型本地启动与 API 服务搭建。

1、编译源码之前需要自行安装 Visual Studio 18 2026 开发工具,并打开VS2026 x64 兼容工具命令提示,执行如下命令克隆仓库。

bash 复制代码
Microsoft Visual Studio 18.0> git clone https://github.com/ggml-org/llama.cpp
Microsoft Visual Studio 18.0> mkdir build
Microsoft Visual Studio 18.0> cd build

2、基础编译(仅 CPU 支持)或者选用GPU 加速编译(已安装 CUDA Toolkit)版。

如果只使用 CPU 模式则执行如下配置

bash 复制代码
Microsoft Visual Studio 18.0> cmake .. -G "Visual Studio 18 2026" -A x64 -DLLAMA_CURL=OFF
Microsoft Visual Studio 18.0> cmake --build . --config Release

如果已安装 CUDA Toolkit 工具包,则添加 -DLLAMA_CUDA=ON 开启 GPU 加速支持

bash 复制代码
Microsoft Visual Studio 18.0> cmake .. -G "Visual Studio 18 2026" -A x64 -DLLAMA_CUDA=ON
Microsoft Visual Studio 18.0> cmake --build . --config Release

编译完成后则会在build目录生成对应的可执行文件,其中的llama-server.exe则为主程序

3、下载 GGUF 格式的模型权重文件,模型权重已同步上线两大主流渠道,开发者任选其一即可。

以魔搭社区为例,直接使用官方PIP包拉取仓库内容到本地,下载后的保存位置为 C:\Users\Admin\dir 执行以下命令完成模型下载。

bash 复制代码
# 安装魔搭
C:> pip install -i https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple modelscope

# 下载 Qwen2.5-1.5B-Instruct-GGUF 完整仓库模型文件
C:> modelscope download --model Qwen/Qwen2.5-1.5B-Instruct-GGUF

# 下载 Qwen2.5-1.5B-Instruct-GGUF 仓库下 qwen2.5-1.5b-instruct-q4_k_m.gguf 模型文件
C:> modelscope download --model Qwen/Qwen2.5-1.5B-Instruct-GGUF qwen2.5-1.5b-instruct-q4_k_m.gguf --local_dir ./dir

4、将下载好的qwen2.5-1.5b-instruct-q4_k_m.gguf模型文件,放在llama-server.exe同级目录下,同时打开命令行工具,执行启动命令:

bash 复制代码
# 执行命令行启动
C:> chcp 65001
C:> llama-cli.exe -m qwen2.5-1.5b-instruct-q4_k_m.gguf -i -c 4096

# 执行启动CPU版本服务端
C:> llama-server.exe -m qwen2.5-1.5b-instruct-q4_k_m.gguf --host 127.0.0.1 --port 11433 -c 4096

# 执行驱动GPU加速版服务端
C:> llama-server.exe -m qwen2.5-1.5b-instruct-q4_k_m.gguf --host 127.0.0.1 --port 11433 -c 1024 --n-gpu-layers 32

5、直接调用 /completion 接口,原生Python标准库,无需额外第三方包,验证服务可用性。

python 复制代码
import json
from urllib import request, error

url = "http://127.0.0.1:11433/completion"
headers = {"Content-Type": "application/json"}

prompt = """<|im_start|>user
你好,简单介绍一下自己<|im_end|>
<|im_start|>assistant
"""

data = {
    "model": "qwen2.5-1.5b-instruct-q4_k_m.gguf",
    "prompt": prompt,
    "temperature": 0.7,
    "max_tokens": 512,
    "ctx_size": 4096,
    "stop": ["<|im_end|>"],
    "stream": False
}

try:
    data_json = json.dumps(data).encode("utf-8")
    req = request.Request(url, data=data_json, headers=headers, method="POST")
    with request.urlopen(req, timeout=60) as response:
        result = json.loads(response.read().decode("utf-8"))

    print("生成结果:")
    print(result["content"].strip())

except error.HTTPError as e:
    print(f"调用失败(HTTP错误):{e.code} - {e.reason}")
except error.URLError as e:
    print(f"调用失败(连接/网络错误):{e.reason}")
except Exception as e:
    print(f"调用失败(其他异常):{e}")

运行main.py,正常输出示例如下,代表本地大模型服务部署完成,可以对外提供推理服务。

bash 复制代码
C:> main.py
生成结果:
您好!我是来自阿里巴巴的AI助手,我叫通义千问。
相关推荐
柳鲲鹏2 小时前
安装WIN11(26H2)时,忽略跳过TPM和SecureBoot检查
windows
weixin1997010801613 小时前
《1688物流API接口边界:logistics.trace.get 与 freight.template.list 的隐藏约束》(附Python源码)
windows·python·list
爱编程的Zion20 小时前
【踩坑记录】同一个 WiFi 下,A 电脑能访问 192.168.0.115,B 电脑死活不行 ——Windows 路由表缺失引发的故障排查
windows·电脑
Qwier1 天前
Win Srv 2019 安装补丁后重启
运维·服务器·windows
lusklusklusk1 天前
Python 四大数据结构详解:元组、列表、集合、字典的特点与增删改查
数据结构·windows·python
今年下半年1 天前
【运维】windows虚拟机安装arm(aarch64)服务器
运维·windows·arm·虚拟机·aarch64
浪潮IT馆1 天前
Windows 10 安装 PostgreSQL 9.6.24 完整教程
数据库·windows·postgresql
sukalot1 天前
Windows 驱动实例分析系列:libwdi 驱动分析 - libwdi 篇(一)
windows
奇树谦1 天前
Windows 下使用 PowerShell 和 CMD 排查端口占用
windows
万年咸鱼2 天前
Qt Windows 使用管理员权限运行 Cmd
windows