.shot 文件格式规范
概述
.shot 文件是 JSON 格式的镜头信息文件,通过 FFmpeg 场景检测自动生成,包含视频中每个镜头的帧号范围、时间信息和语义信息。
设计理念
- 零侵入 - 标准 SRT 保持纯净,所有语义信息集中在 .shot 文件
- 轻量级 - 只存储元数据(帧号、时间、语义),不存储实际图像
- 按需抽取 - AI 可以根据帧号按需抽取特定帧进行分析
- 与视频同名 - 便于管理和关联
- 机器可读 - JSON 格式,易于程序解析
文件命名
.shot 文件与视频文件同名,仅扩展名不同:
- 视频文件:
video.mp4 - 镜头文件:
video.shot
格式结构
json
{
"video": "视频文件路径",
"fps": 帧率,
"created_at": "视频创建时间(ISO 8601格式)",
"modified_at": "视频修改时间(ISO 8601格式)",
"fingerprint": "视频指纹(SHA256哈希)",
"shot_created_at": ".shot文件创建时间(ISO 8601格式)",
"shot_modified_at": ".shot文件最后修改时间(ISO 8601格式)",
"properties": {
"overall_emotion": "整体情感",
"overall_pace": "整体节奏",
"color_tone": "色调",
"style": "风格",
"sound_type": "声音类型",
"genre": "类型",
"theme": "主题",
"summary": "摘要",
"tags": ["标签1", "标签2"]
},
"shots": [
{
"id": 镜头序号,
"start_frame": 起始帧号,
"end_frame": 结束帧号,
"start_time": "起始时间",
"end_time": "结束时间",
"duration": "持续时间",
"notes": "人工注释",
"semantic": {
"type": "类型",
"emotion": "情感",
"pace": "节奏",
"scene": "场景",
"description": "描述",
"tags": ["标签1", "标签2"]
}
}
],
"subtitles": [
{
"index": 字幕序号,
"start_time": "起始时间",
"end_time": "结束时间",
"text": "字幕文本",
"semantic": {
"type": "类型",
"emotion": "情感",
"pace": "节奏",
"scene": "场景",
"description": "描述",
"tags": ["标签1", "标签2"]
}
}
]
}
字段说明
顶层字段
video(string): 视频文件路径fps(float): 视频帧率created_at(string): 视频创建时间,ISO 8601格式(如:2026-10-02T12:00:00Z)modified_at(string): 视频修改时间,ISO 8601格式fingerprint(string): 视频指纹,SHA256哈希值,用于检测视频是否被修改shot_created_at(string): .shot文件创建时间,ISO 8601格式shot_modified_at(string): .shot文件最后修改时间,ISO 8601格式properties(object): 影片整体属性shots(array): 镜头列表subtitles(array): 字幕列表(带语义信息)
影片整体属性 (properties)
overall_emotion(string): 整体情感,可选值:positive, negative, neutral, mixed 等overall_pace(string): 整体节奏,可选值:fast, slow, medium, variable 等color_tone(string): 色调,可选值:warm, cool, neutral, high_contrast, low_contrast 等style(string): 风格,可选值:cinematic, documentary, vlog, commercial, tutorial 等sound_type(string): 声音类型,可选值:music, dialogue, ambient, music+dialogue, silent 等genre(string): 类型,可选值:travel, food, product, lifestyle, tutorial, review 等theme(string): 主题(自由文本)summary(string): 影片摘要(1-2句话)tags(array): 自定义标签
镜头字段 (Shot)
id(int): 镜头序号,从 1 开始start_frame(int): 起始帧号(包含)end_frame(int): 结束帧号(包含)start_time(string): 起始时间,格式:HH:MM:SS,mmmend_time(string): 结束时间,格式:HH:MM:SS,mmmduration(string): 持续时间,格式:HH:MM:SS,mmmnotes(string): 人工注释(可选)semantic(object): 语义信息(可选)type(string): 类型,可选值:opening, transition, highlight, explanation, closing, dialogue, monologue, action, scenery, product, otheremotion(string): 情感,可选值:positive, negative, neutral, excited, calm, tense, sad, happy, surprised, otherpace(string): 节奏,可选值:fast, slow, mediumscene(string): 场景,可选值:outdoor, indoor, product, person, landscape, city, nature, food, travel, otherdescription(string): 详细描述(1-2句话)tags(array): 自定义标签
字幕字段 (Subtitle)
index(int): 字幕序号,从 1 开始start_time(string): 起始时间,格式:HH:MM:SS,mmmend_time(string): 结束时间,格式:HH:MM:SS,mmmtext(string): 字幕文本semantic(object): 语义信息(可选)type(string): 类型,可选值:opening, transition, highlight, explanation, closing, dialogue, monologue, action, scenery, product, otheremotion(string): 情感,可选值:positive, negative, neutral, excited, calm, tense, sad, happy, surprised, otherpace(string): 节奏,可选值:fast, slow, mediumscene(string): 场景,可选值:outdoor, indoor, product, person, landscape, city, nature, food, travel, otherdescription(string): 详细描述(1-2句话)tags(array): 自定义标签
示例
json
{
"video": "video.mp4",
"fps": 30.0,
"properties": {
"overall_emotion": "positive",
"overall_pace": "medium",
"color_tone": "warm",
"style": "vlog",
"sound_type": "music+dialogue",
"genre": "travel",
"theme": "海滨旅行",
"summary": "记录了一次海滨旅行的美好时光",
"tags": ["海滩", "阳光", "美食", "休闲"]
},
"shots": [
{
"id": 1,
"start_frame": 0,
"end_frame": 900,
"start_time": "00:00:00,000",
"end_time": "00:00:30,000",
"duration": "00:00:30,000",
"notes": "开场镜头,展示酒店外观",
"semantic": {
"type": "opening",
"emotion": "positive",
"pace": "slow",
"scene": "outdoor",
"description": "开场镜头,表达轻松愉悦的心情",
"tags": ["阳光", "蓝天", "轻松"]
}
},
{
"id": 2,
"start_frame": 901,
"end_frame": 1800,
"start_time": "00:00:30,033",
"end_time": "00:01:00,000",
"duration": "00:00:29,967",
"notes": "海滩全景,海浪声",
"semantic": {
"type": "scenery",
"emotion": "positive",
"pace": "medium",
"scene": "outdoor",
"description": "展示海滩全景,海浪声背景",
"tags": ["海滩", "海浪", "风景"]
}
},
{
"id": 3,
"start_frame": 1801,
"end_frame": 2700,
"start_time": "00:01:00,033",
"end_time": "00:01:30,000",
"duration": "00:00:29,967",
"notes": "品尝当地美食",
"semantic": {
"type": "product",
"emotion": "positive",
"pace": "medium",
"scene": "food",
"description": "产品特写镜头,展示细节工艺",
"tags": ["产品特写", "细节展示"]
}
}
],
"subtitles": [
{
"index": 1,
"start_time": "00:00:02,000",
"end_time": "00:00:05,000",
"text": "今天天气真好",
"semantic": {
"type": "opening",
"emotion": "positive",
"pace": "slow",
"scene": "outdoor",
"description": "开场台词,表达轻松愉悦的心情",
"tags": ["阳光", "蓝天", "轻松"]
}
},
{
"index": 2,
"start_time": "00:00:05,500",
"end_time": "00:00:08,000",
"text": "我们来到了美丽的海滩",
"semantic": {
"type": "scenery",
"emotion": "positive",
"pace": "medium",
"scene": "outdoor",
"description": "介绍场景,展示海滩美景",
"tags": ["海滩", "风景", "旅行"]
}
}
]
}
与 SRT 的关联
标准 SRT 文件保持纯净,不包含任何语义信息。字幕内容会被读取并存储到 .shot 文件的 subtitles 字段中,同时添加每条字幕的语义信息。
工作流程
- 读取标准 SRT - 解析纯净的 SRT 字幕文件
- 存储到 .shot - 将字幕文本和时间戳存储到
subtitles字段 - 添加语义标签 - 为每条字幕添加语义信息(type, emotion, pace, scene, description, tags)
- AI 使用 - AI 可以直接从 .shot 文件读取字幕及其语义信息
优势
- ✅ 零侵入 - 标准 SRT 保持纯净,可被任何播放器/工具使用
- ✅ 集中管理 - 所有语义信息统一在 .shot 文件中
- ✅ 细粒度控制 - 每条字幕都有独立的语义标签
- ✅ 便于 AI 理解 - AI 可以精确理解每个时间片段的内容和情感
按需抽取帧
当 AI 需要查看某个镜头的画面时,可以根据帧号使用 FFmpeg 抽取帧:
bash
# 抽取镜头的中间帧
ffmpeg -i video.mp4 -vf "select='eq(n,450)'" -vframes 1 frame_450.jpg
# 根据时间抽取帧
ffmpeg -ss 00:00:15 -i video.mp4 -vframes 1 frame.jpg
生成方式
使用 video-ai-cutter detect-shots 命令生成:
bash
video-ai-cutter detect-shots -i video.mp4 -o video.shot
命令参数:
-i, --input: 输入视频文件路径-o, --output: 输出 .shot 文件路径(可选,默认与视频同名)-t, --threshold: 场景变化阈值 0-1(可选,默认 0.3)
生成 .shot 文件后,使用 video-ai-cutter enhance-shot 命令为镜头添加语义信息:
bash
video-ai-cutter enhance-shot -s video.shot -i video.srt
场景检测原理
使用 FFmpeg 的 select 滤镜检测场景变化:
bash
ffmpeg -i video.mp4 -vf "select='gt(scene,0.3)'" -f null -
scene计算相邻帧之间的差异gt(scene,0.3)表示差异超过 0.3 时视为场景切换- 阈值越低,检测越敏感(可能产生更多镜头)
- 阈值越高,检测越保守(可能遗漏细微切换)
视频指纹和去重
指纹作用
fingerprint字段存储视频的 SHA256 哈希值- 用于检测视频是否被修改或替换
- 避免对同一视频重复进行语义分析
工作流程
-
首次分析:
- 计算视频文件的 SHA256 哈希值,存储到
fingerprint字段 - 记录视频的创建时间和修改时间
- 生成 .shot 文件并添加语义标签
- 计算视频文件的 SHA256 哈希值,存储到
-
后续使用:
- 检查 .shot 文件是否存在
- 如果存在,比较当前视频的哈希值与 .shot 文件中的
fingerprint - 如果哈希值匹配,直接使用现有的 .shot 文件,跳过分析
- 如果哈希值不匹配,说明视频已修改,需要重新分析
-
避免重复分析:
- 即使视频文件被移动或重命名,只要内容不变,指纹就会匹配
- 大幅节省算力和时间成本
- 实现真正的"一次标注,永久复用"
时间戳管理
created_at/modified_at:视频文件的原始时间戳shot_created_at/shot_modified_at:.shot 文件的时间戳- 用于追踪文件的历史和版本
示例场景
场景1:视频未修改
less
原视频:video.mp4 (fingerprint: abc123...)
.shot文件:video.shot (fingerprint: abc123...)
结果:直接使用现有 .shot 文件,无需重新分析
场景2:视频已修改
less
原视频:video.mp4 (fingerprint: abc123...)
修改后:video.mp4 (fingerprint: def456...)
.shot文件:video.shot (fingerprint: abc123...)
结果:指纹不匹配,需要重新生成 .shot 文件
场景3:视频移动位置
less
原位置:/old/path/video.mp4 (fingerprint: abc123...)
新位置:/new/path/video.mp4 (fingerprint: abc123...)
.shot文件:/new/path/video.shot (fingerprint: abc123...)
结果:指纹匹配,可以直接使用 .shot 文件