安装时遇到的一些问题记录在文末了
使用的环境信息
- Windows11 专业版 23H2
- Python 3.12.15
- Conda 25.11.0
- PyTorch 版本:2.10.0+cu128
- PyTorch 使用的 CUDA 版本:12.8
- GPU 型号:NVIDIA Geforce RTX 3070 Laptop
- NVIDIA CUDA Toolkit: release 12.8, V12.8.61
以上环境为部署时使用的环境信息
前置环境
Anaconda / Miniconda
需要使用 conda 环境,可安装 Miniconda 或 Anaconda。因为我已经有 Anaconda 所以用的是 Anaconda,Anaconda 相比 Miniconda 会更大一些,如果追求轻量或暂时不需要完整依赖可以使用 Miniconda。
Anaconda 与 Miniconda 下载地址 www.anaconda.com/download/su...
CUDA
下载地址 developer.nvidia.com/cuda-downlo...
模型部署流程
相关链接:
Meta Sam3模型官方网站: ai.meta.com/research/sa...
GitHub Sam3 仓库:github.com/facebookres...
ModelScope Sam3 仓库:www.modelscope.cn/models/face...
1. 创建 Conda 环境
powershell
conda create -n sam3 python=3.12
conda deactivate
conda activate sam3
2. 安装带有 CUDA 支持的 PyTorch
powershell
pip install torch==2.10.0 torchvision --index-url https://download.pytorch.org/whl/cu128
3. 克隆仓库并安装软件包
powershell
git clone https://github.com/facebookresearch/sam3.git
cd sam3
pip install -e .
4. 下载模型权重
通常是在 Huggingface 上的 sam3 项目中填写申请表,申请通过后可下载权重,但通过概率小的可怜,所以这里使用 ModelScope 下载权重
在 Conda 的 sam3 虚拟环境中执行
python -m pip install modelscope
新建权重目录
arduino
mkdir checkpoints
下载 Sam3 权重
css
modelscope download --model facebook/sam3 sam3.pt --local_dir checkpoints
5. 我的项目结构
lua
D:\xxx\xxx\
|
|--sam3(从 git 克隆的源码项目)
| |---assets
| |---examples
| |---sam3
| |---scripts
| |---test
| |---此处省略
|
|--myproject(我创建的pycharm项目)
| |---checkpoints(模型权重)
| |---images(图片素材)
| |---outputs(输出结果目录)
| |---segment_image.py
6. Demo 代码内容
python
from pathlib import Path
from PIL import Image, ImageDraw
from sam3.model_builder import build_sam3_image_model
from sam3.model.sam3_image_processor import Sam3Processor
from contextlib import nullcontext
import numpy as np
import torch
# ===== 配置 =====
ROOT = Path(__file__).resolve().parent
CHECKPOINT = ROOT / "checkpoints" / "sam3.pt"
IMAGE_PATH = ROOT / "images" / "desktop.jpg"
OUTPUT_DIR = ROOT / "outputs"
PROMPT = "crumpled napkin" # 提示词,可以是 car、cat等图片中存在的、希望找出的内容,使用英文单词
THRESHOLD = 0.5
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
def main():
device = "cuda" if torch.cuda.is_available() else "cpu"
print(f"运行设备: {device}")
if not CHECKPOINT.exists():
raise FileNotFoundError(f"找不到权重: {CHECKPOINT}")
if not IMAGE_PATH.exists():
raise FileNotFoundError(f"找不到图片: {IMAGE_PATH}")
# 1. 加载模型
print("正在加载 SAM3...")
model = build_sam3_image_model(
checkpoint_path=str(CHECKPOINT),
load_from_HF=False,
device=device,
)
model.eval()
processor = Sam3Processor(model)
processor.set_confidence_threshold(THRESHOLD)
# 2. 读取图片
image = Image.open(IMAGE_PATH).convert("RGB")
# 3. 推理
print(f"正在识别: {PROMPT}")
amp_context = (
torch.autocast(device_type="cuda", dtype=torch.bfloat16)
if device == "cuda"
else nullcontext()
)
with torch.inference_mode(), amp_context:
state = processor.set_image(image)
output = processor.set_text_prompt(
state=state,
prompt=PROMPT,
)
masks = output["masks"].detach().float().cpu().numpy()
boxes = output["boxes"].detach().float().cpu().numpy()
scores = output["scores"].detach().float().cpu().numpy()
count = len(scores)
print(f"找到 {len(scores)} 个目标")
# 4. 绘制半透明掩码
result = image.convert("RGBA")
combined_mask = np.zeros(
(image.height, image.width), dtype=bool
)
colors = [
(255, 80, 80, 110),
(50, 200, 120, 110),
(70, 130, 255, 110),
(255, 180, 50, 110),
]
from PIL import ImageFilter
MASK_THRESHOLD = 0.4 # 掩码阈值
FEATHER_RADIUS = 1.5 # 羽化半径,可尝试 1.0 ~ 2.5
for i, (mask, box, score) in enumerate(
zip(masks, boxes, scores)
):
# 1. 原始 mask
mask = np.squeeze(mask)
# 2. 阈值二值化(不要直接 astype(bool))
mask_bin = (mask > MASK_THRESHOLD).astype(np.uint8) * 255
# 3. 转成 PIL 灰度图,并轻微羽化
mask_img = Image.fromarray(mask_bin, mode="L")
mask_img = mask_img.filter(ImageFilter.GaussianBlur(radius=FEATHER_RADIUS))
# 4. 为 combined_mask 保留一个硬掩码版本(用于最终合并统计)
hard_mask = np.array(mask_img) > 127
combined_mask |= hard_mask
color = colors[i % len(colors)]
# 5. 用羽化后的 mask 作为 alpha 叠加
overlay = Image.new("RGBA", image.size, (0, 0, 0, 0))
overlay.paste(
Image.new("RGBA", image.size, color),
(0, 0),
mask_img,
)
result = Image.alpha_composite(result, overlay)
# 5. 绘制边界框、编号、置信度及总数量
from PIL import ImageFont
draw = ImageDraw.Draw(result)
# 根据图片宽度设置字体大小
font_size = max(18, image.width // 65)
# Windows 系统字体
font_path = r"C:\Windows\Fonts\arial.ttf"
font = ImageFont.truetype(font_path, font_size)
for index, (box, score) in enumerate(
zip(boxes, scores), start=1
):
x1, y1, x2, y2 = map(float, box)
# 绘制目标边框
draw.rectangle(
[x1, y1, x2, y2],
outline=(255, 255, 0, 255),
width=3,
)
label = f"#{index} {PROMPT} {score:.1%}"
# 测量文字大小
bbox = draw.textbbox(
(0, 0), label, font=font
)
text_width = bbox[2] - bbox[0]
text_height = bbox[3] - bbox[1]
padding = 6
# 优先显示在边框上方,空间不够则显示在内部
label_x = max(
0, min(x1, image.width - text_width - padding * 2)
)
if y1 >= text_height + padding * 2:
label_y = y1 - text_height - padding * 2
else:
label_y = y1
label_y = max(
0, min(label_y, image.height - text_height - padding * 2)
)
# 绘制文字背景
draw.rectangle(
[
label_x,
label_y,
label_x + text_width + padding * 2,
label_y + text_height + padding * 2,
],
fill=(255, 255, 0, 240),
)
# 绘制黑色文字
draw.text(
(label_x + padding, label_y + padding - bbox[1]),
label,
font=font,
fill=(0, 0, 0, 255),
)
# 注意:总数量绘制放在 for 循环外面
count_text = f"{PROMPT} Count: {count}"
count_font = ImageFont.truetype(
font_path, max(22, image.width // 45)
)
count_bbox = draw.textbbox(
(0, 0), count_text, font=count_font
)
count_width = count_bbox[2] - count_bbox[0]
count_height = count_bbox[3] - count_bbox[1]
draw.rectangle(
[0, 0, count_width + 24, count_height + 24],
fill=(0, 0, 0, 220),
)
draw.text(
(12, 12 - count_bbox[1]),
count_text,
font=count_font,
fill=(255, 255, 255, 255),
)
# 6. 保存标注图
result.save(OUTPUT_DIR / "result.png")
# 7. 保存黑白掩码
mask_image = Image.fromarray(
(combined_mask.astype(np.uint8) * 255),
mode="L",
)
mask_image.save(OUTPUT_DIR / "mask.png")
# 8. 保存透明背景抠图
cutout = image.convert("RGBA")
cutout.putalpha(mask_image)
cutout.save(OUTPUT_DIR / "cutout.png")
print("处理完成!")
print(f"结果目录: {OUTPUT_DIR}")
if __name__ == "__main__":
main()
Demo 中如有欠缺或不足之处还请见谅
遇到的问题记录
以下内容部分由 AI 辅助整理,如有不足请见谅
1. 移动项目后,Python 无法导入 SAM3
问题描述
此前使用 Conda 的 sam3 环境能够正常运行代码,但移动项目目录并重新配置 PyCharm 后,出现 ModuleNotFoundError: No module named 'sam3',报错位置:from sam3.model_builder import build_sam3_image_model。主要原因是 SAM3 源码位置发生变化,而 Python 环境没有正确关联新的源码位置。
解决方式
在包含 pyproject.toml 的 SAM3 官方仓库根目录执行可编辑安装:
bash
conda activate sam3
cd /d D:\CHMMCH\sam3\sam3
python -m pip install -e .
安装后验证
scss
python -c "import sam3; print(sam3.__file__)"
预期能够输出 SAM3 源码中 __init__.py 的实际路径。
补充说明
pip install -e . 属于 Editable Installation,Python 会关联本地源码目录。
因此,如果以后再次移动或删除这个目录,可能需要重新执行安装。
2. 安装 SAM3 后出现 NumPy 依赖冲突
问题描述
在官方仓库执行:
erlang
python -m pip install -e .
安装成功后出现以下冲突:
arduino
opencv-python 4.12.0.88 requires numpy<2.3.0,>=2
but you have numpy 1.26.4
reme-ai 0.3.0.5 requires numpy>=2.2.6
but you have numpy 1.26.4
原因是 SAM3 0.1.0 声明:
shell
numpy>=1.26,<2
但当前可见环境中的 OpenCV 4.12 和 reme-ai 需要 NumPy 2.x。
不同软件包的依赖要求无法同时满足。
解决方式
优先保持 SAM3 所需的 NumPy 版本:
arduino
python -m pip install "numpy==1.26.4"
如果需要 OpenCV,可选择支持 NumPy 1.x 的版本,例如:
arduino
python -m pip install "opencv-python==4.11.0.86"
更重要的是,先检查是否混入了其他 Python 环境中的包,而不是立即卸载所有冲突软件。
对于只在其他项目中使用的依赖,建议放在对应的独立环境中。
3. Conda 环境意外使用了用户级 Python 软件包
问题描述
安装日志中多次出现:
makefile
Requirement already satisfied in
C:\Users\user\AppData\Roaming\Python\Python312\site-packages
但当前激活的 Conda 环境位于:
makefile
C:\Users\user.conda\envs\sam3
说明 Conda 环境能够读取用户级 site-packages。
这导致 pip 认为部分依赖已经安装,实际上这些依赖并不在 sam3 环境自己的目录中。
后续也造成了依赖检查和软件卸载方面的混乱。
解决方式
先在当前 Windows CMD 窗口中临时禁用用户级包:
ini
set PYTHONNOUSERSITE=1
检查是否生效:
scss
python -c "import site; print(site.ENABLE_USER_SITE)"
预期:
python
False
确认没有问题后,可以给 Conda 环境单独设置环境变量:
arduino
conda env config vars set -n sam3 PYTHONNOUSERSITE=1
conda deactivate
conda activate sam3
这样在以后重新激活 sam3 环境时,也会禁用用户级第三方包目录。
备注: 禁用用户级目录只是隔离包,并不会把这些包自动安装到 Conda 环境中。
4. 卸载 reme-ai 后,pip check 出现其他项目依赖错误
问题描述
执行:
python -m pip uninstall reme-ai
后,再运行:
sql
python -m pip check
出现:
erlang
copaw 0.0.5 requires reme-ai, which is not installed.
andriller 3.6.3 has requirement Jinja2<3,>=2.11.3,
but you have jinja2 3.1.6.
andriller 3.6.3 has requirement MarkupSafe==2.0.1,
but you have markupsafe 3.0.3.
从卸载日志来看,reme-ai 原先位于用户级 Python 包目录,而非独立 Conda 环境。
因此,虽然执行命令时激活的是 sam3,实际卸载的仍是其他 Python 项目可能使用的用户级软件包。
解决方式
停止继续卸载无关依赖。
先隔离用户级包:
ini
set PYTHONNOUSERSITE=1
再检查当前 Conda 环境:
sql
python -m pip check
对于 CoPaw 等其他项目,应在其对应环境中单独检查或恢复依赖。
小结
在执行 pip uninstall 前,可以先检查包的实际位置:
sql
python -m pip show reme-ai
重点查看 Location 字段。
5. 禁用用户级包后,突然出现大量依赖缺失
问题描述
执行:
ini
set PYTHONNOUSERSITE=1
python -m pip check
后,出现大量报错,例如:
arduino
ftfy 6.1.1 requires wcwidth, which is not installed.
iopath 0.1.10 requires portalocker, which is not installed.
sam3 0.1.0 requires huggingface-hub, which is not installed.
sam3 0.1.0 requires regex, which is not installed.
torch 2.10.0+cu128 requires filelock, which is not installed.
torchvision 0.25.0+cu128 requires pillow, which is not installed.
这并不是禁用用户级包导致依赖被删除,而是这些依赖原先只存在于用户级 Python 目录中。
隔离后,Conda 环境便无法继续访问它们。
解决方式
确保已经激活 sam3,并保持用户级包隔离:
ini
conda activate sam3
set PYTHONNOUSERSITE=1
回到 SAM3 源码目录:
bash
cd /d D:\CHMMCH\sam3\sam3
重新安装项目依赖:
erlang
python -m pip install -e .
然后检查:
sql
python -m pip check
如果还有缺失依赖,再针对性补装。
这一步的核心目的是:让 SAM3 的依赖真正安装在它自己的 Conda 环境中,而不是依赖其他 Python 环境。
6. 成功安装 SAM3 后,运行时仍然缺少 einops
问题描述
运行:
python test_sam3.py
后出现:
arduino
File "sam3\sam\rope.py", line 15
from einops import rearrange, repeat
ModuleNotFoundError: No module named 'einops'
说明 SAM3 在导入模型模块时,引用了 einops 库,但当前 Conda 环境没有安装它。
即使 pip install -e . 已经成功,仍可能出现此类问题,因为项目的基础依赖声明并不一定覆盖所有运行路径。
解决方式
在 sam3 环境中安装:
python -m pip install einops
验证:
scss
python -c "import einops; print(einops.__version__)"
然后重新运行:
python test_sam3.py
如果还需要使用官方 Notebook 示例,也可以考虑安装对应的可选依赖:
arduino
python -m pip install -e ".[notebooks]"
但这个命令会引入更多第三方软件包,并非仅执行图像推理所必需。
官方 GitHub 也有用户反馈 Windows 下安装成功后仍缺少部分模块的问题。
7. Windows 环境无法通过 pip 安装 Triton
问题描述
执行:
python -m pip install triton
出现:
yaml
ERROR: Could not find a version that satisfies
the requirement triton (from versions: none)
ERROR: No matching distribution found for triton
原因是标准 Triton 在 PyPI 上没有适用于当前 Windows 环境的发行包。
Triton 与 PyTorch 的部分 GPU 计算、内核编译及优化功能有关。
解决方式
在 Windows 原生环境下使用 Windows 适配版本:
arduino
python -m pip install "triton-windows>=3.6,<3.7"
其中版本范围是针对当时使用的 PyTorch 2.10 环境选择的,其他 PyTorch 版本需要重新检查兼容性。
安装完成后验证:
scss
python -c "import triton; print(triton.__version__)"
再次运行模型:
python test_sam3.py
后续已经成功进入实际图像推理阶段,说明 Triton 导入问题已解决。
相关项目: triton-windows GitHub
备注: triton-windows 是安装包名称,但在 Python 中仍使用:
arduino
import triton
tips:安装成功不代表所有 Triton GPU 内核都一定能正常编译和运行,还需要结合显卡、驱动和 PyTorch 版本验证。