文献:LocoVLM: Grounding Vision and Language for Adapting Versatile Legged Locomotion Policies
作者:I Made Aswin Nahrendra, Seunghyun Lee, Dongkyu Lee, Hyun Myung
单位:KAIST(Korea Advanced Institute of Science and Technology)
会议版本:ICRA 2025 Workshop on Safe Vision-Language Models(SafeVLMs)
项目主页:https://locovlm.github.io/
四足平台:Unitree Go1 ;人形泛化验证:Unitree H1(MuJoCo)
高层 Foundation Model:GPT-4o + BLIP-2
低层运动控制:PPO + Asymmetric Actor-Critic + Style-Conditioned Locomotion Policy
仿真平台:Isaac Gym(Go1)/ MuJoCo(H1 泛化实验)
核心关键词:Vision-Language Grounding、Legged Locomotion、Skill Retrieval、Prompted Reasoning、Compliant Contact Tracking、Sim-to-Real
1. 核心结论
LocoVLM 解决的不是"让一个大模型直接输出四足机器人的 12 个关节动作",而是更实际的层级式问题:
如何把视觉或自然语言中的高层语义,实时、稳定地转换为足式机器人能够执行的运动风格参数,并且不把云端大语言模型放进高频控制环。
它的完整思路可以概括为:
Vision / Language Query → BLIP-2 Retrieval → M = ( T , ψ , v x l i m i t ) → Style-Conditioned RL Policy → q d e s → Motor Controller \text{Vision / Language Query} \rightarrow \text{BLIP-2 Retrieval} \rightarrow \mathbf{M}=(T,\psi,v_x^{\mathrm{limit}}) \rightarrow \text{Style-Conditioned RL Policy} \rightarrow q^{\mathrm{des}} \rightarrow \text{Motor Controller} Vision / Language Query→BLIP-2 Retrieval→M=(T,ψ,vxlimit)→Style-Conditioned RL Policy→qdes→Motor Controller
其中:
- T T T:步态周期;
- ψ \psi ψ:四条腿的相位偏移;
- v x l i m i t v_x^{\mathrm{limit}} vxlimit:前向速度上限。
GPT-4o 并不在真机在线控制环中持续调用,而是在离线阶段生成大量"语言指令---推理---运动描述符"数据,形成 Skill Database(技能数据库)。部署时由轻量得多的预训练 VLM------BLIP-2------完成视觉/语言查询与技能数据库之间的实时匹配,再把运动参数发送给 PPO 训练的运动控制器。
因此这篇论文真正的系统结构是:
LLM 离线知识生成 + VLM 在线语义检索 + RL 高频运动控制 \boxed{ \text{LLM 离线知识生成} + \text{VLM 在线语义检索} + \text{RL 高频运动控制} } LLM 离线知识生成+VLM 在线语义检索+RL 高频运动控制
这也是它和 SayTap、传统速度控制策略以及端到端 VLA 的根本区别。
2. 研究背景:论文到底解决了什么问题
2.1 传统足式运动控制主要理解"几何",但不理解"语义"
传统四足运动策略通常依据:
- 高程图;
- 深度图;
- 点云;
- 地形法向量;
- 本体状态;
- 目标速度;
来判断"哪里能踩""哪里有障碍物"。
这类表示对于几何避障非常有效,但无法直接理解:
- "这里是图书馆,安静一点";
- "地面结冰了,谨慎走";
- "像兔子一样跳";
- "前方环境拥挤,慢一点";
- "像鹿一样轻快地走"。
换句话说,几何感知可以表示:
Obstacle Geometry \text{Obstacle Geometry} Obstacle Geometry
却不天然包含:
Affordance + Social Context + Task Semantics + Commonsense \text{Affordance} + \text{Social Context} + \text{Task Semantics} + \text{Commonsense} Affordance+Social Context+Task Semantics+Commonsense
LocoVLM 的第一个目标,就是把这类高层语义信息真正作用到低层步态和速度上。
2.2 大语言模型很强,但不适合直接进入 50--200 Hz 控制环
如果直接让 GPT-4o 在线生成机器人动作,会遇到:
- 网络依赖;
- 云端 API 延迟;
- 推理时间不可确定;
- 成本随控制时间持续增加;
- 断网时无法工作;
- LLM 输出不适合直接驱动连续动力学系统。
因此论文没有设计:
LLM → Joint Torque \text{LLM} \rightarrow \text{Joint Torque} LLM→Joint Torque
而是使用 Foundation Model Knowledge Distillation(基础模型知识蒸馏):
GPT-4o Knowledge → Offline Skill Database → BLIP-2 Retrieval \text{GPT-4o Knowledge} \rightarrow \text{Offline Skill Database} \rightarrow \text{BLIP-2 Retrieval} GPT-4o Knowledge→Offline Skill Database→BLIP-2 Retrieval
GPT-4o 只承担离线的"老师"角色。
2.3 步态风格跟踪与鲁棒性之间存在冲突
如果强化学习控制器被强制严格遵循预定义足端接触时序,那么在:
- 台阶;
- 离散落脚点;
- 粗糙地面;
- 突发扰动;
中,机器人可能为了"守住规定步态"而牺牲稳定性。
论文因此提出 Compliant Contact Tracking(柔顺接触跟踪):
当实际接触与目标接触只存在一定范围内的误差时,不强制处罚;只有偏差超过阈值后才恢复严格约束。
这让控制器可以在遇到扰动时暂时违反理想步态,优先维持稳定。
2.4 技能数据库扩大后,如何兼顾检索速度与语义准确率
BLIP-2 有两类可用于检索的信息:
- Embedding Cosine Similarity:速度快,但语义区分能力有限;
- Image-Text Matching(ITM)Head:语义匹配更精确,但对整个数据库逐项计算代价太高。
论文提出 Mixed-Precision Retrieval(混合精度检索):
Fast Coarse Retrieval → Accurate Fine Re-ranking \text{Fast Coarse Retrieval} \rightarrow \text{Accurate Fine Re-ranking} Fast Coarse Retrieval→Accurate Fine Re-ranking
这里的"Mixed-Precision"不是 FP16/FP32 混合浮点精度训练,而是指:
使用低成本的粗粒度相似度先筛选,再使用高精度 ITM 对少量候选重新排序。
3. 整体架构

图 1 LocoVLM 总体框架。来源:Nahrendra et al., LocoVLM, Figure 1,官方项目主页。LLM 离线生成技能数据库;部署阶段 VLM 通过混合精度检索得到运动描述符;低层 Style-Conditioned Locomotion Policy 根据运动描述符、速度命令和本体状态执行动作。
论文把系统理解为一个高低层分离的 Teacher--Student(教师---学生)结构。
3.1 高层 Teacher:GPT-4o
GPT-4o 用于离线构造数据库:
D = { ( I 1 , M 1 ) , ( I 2 , M 2 ) , ... } \mathcal{D} =\left\{ (\mathcal{I}_1,\mathbf{M}_1), (\mathcal{I}_2,\mathbf{M}_2), \dots \right\} D={(I1,M1),(I2,M2),...}
其中:
- I \mathcal{I} I:语言指令;
- M \mathbf{M} M:可执行运动描述符。
GPT-4o 不负责真机实时推理。
3.2 高层 Student:BLIP-2
部署时,BLIP-2 输入可以是:
I q u e r y = { text , RGB image . \mathcal{I}_{\mathrm{query}} =\begin{cases} \text{text}, \\ \text{RGB image}. \end{cases} Iquery={text,RGB image.
VLM 根据语义从数据库中取回最匹配的运动描述符。
3.3 低层控制器:Style-Conditioned RL Policy
低层控制器输入:
Proprioception + Velocity Command + Gait Style Parameters \text{Proprioception} + \text{Velocity Command} + \text{Gait Style Parameters} Proprioception+Velocity Command+Gait Style Parameters
输出:
q d e s q^{\mathrm{des}} qdes
即关节目标角,由机器人电机控制器进一步转成力矩。
4. Style-Conditioned Locomotion Policy
4.1 为什么不直接使用离散 Gait Label
如果只给策略输入:
text
TROT
PACE
BOUND
它只能在有限离散动作之间切换。
LocoVLM 将步态参数化为:
( T , ψ ) (T,\psi) (T,ψ)
其中:
T ∈ R T\in\mathbb{R} T∈R
表示 Gait Cycle Duration(步态周期),而
ψ = ψ F L , ψ F R , ψ R L , ψ R R ∈ R 4 \psi = \\psi_{\\mathrm{FL}}, \\psi_{\\mathrm{FR}}, \\psi_{\\mathrm{RL}}, \\psi_{\\mathrm{RR}} \in\mathbb{R}^{4} ψ=ψFL,ψFR,ψRL,ψRR∈R4
表示四条腿的相位偏移。
同一种 gait 还可以通过改变 T T T 连续调节步频:
f = 1 T f=\frac{1}{T} f=T1
所以:
T ↓ ⇒ f ↑ T\downarrow \quad\Rightarrow\quad f\uparrow T↓⇒f↑
即周期越短,步频越高。
4.2 Gait Phase Encoding
论文使用二维 Clock Input:
ϕ ( t ) = sin ( 2 π t T ) , cos ( 2 π t T ) \phi(t) =\left \\sin\\left(2\\pi\\frac{t}{T}\\right), \\cos\\left(2\\pi\\frac{t}{T}\\right) \\right ϕ(t)=sin(2πTt),cos(2πTt)
再结合每条腿的:
ψ i \psi_i ψi
获得不同腿的接触相位关系。

图 2 步态相位编码与 Compliance Zone。来源:Nahrendra et al., LocoVLM, Figure 2。图中以 T = 1 T=1 T=1 为例,橙色正弦/余弦曲线表示周期相位,绿色区域表示允许策略偏离理想接触时序的柔顺区域。
这类时钟信号的作用是让策略明确知道:
当前处于一个 gait cycle 的什么位置。
而不是让策略完全依赖历史状态自行推断步态相位。
5. 五种基础步态及相位偏移
表 1 原论文 Table III:五种步态的 Foot Phase Offsets。
| Gait | FL | FR | RL | RR |
|---|---|---|---|---|
| Pronk | 0.0 | 0.0 | 0.0 | 0.0 |
| Trot | 0.0 | 0.5 | 0.5 | 0.0 |
| Pace | 0.0 | 0.5 | 0.0 | 0.5 |
| Bound | 0.0 | 0.0 | 0.5 | 0.5 |
| Rotary gallop | 0.0 | 0.2 | 0.7 | 0.5 |
其中:
- FL:Front Left;
- FR:Front Right;
- RL:Rear Left;
- RR:Rear Right。
5.1 Pronk
ψ = 0 , 0 , 0 , 0 \psi=0,0,0,0 ψ=0,0,0,0
四腿同相:
F L = F R = R L = R R \mathrm{FL} =\mathrm{FR} =\mathrm{RL} =\mathrm{RR} FL=FR=RL=RR
表现为四脚同时起跳和落地。
5.2 Trot
ψ = 0 , 0.5 , 0.5 , 0 \psi=0,0.5,0.5,0 ψ=0,0.5,0.5,0
因此:
F L ∼ R R \mathrm{FL}\sim\mathrm{RR} FL∼RR
F R ∼ R L \mathrm{FR}\sim\mathrm{RL} FR∼RL
是典型对角小跑。
5.3 Pace
ψ = 0 , 0.5 , 0 , 0.5 \psi=0,0.5,0,0.5 ψ=0,0.5,0,0.5
即:
F L ∼ R L \mathrm{FL}\sim\mathrm{RL} FL∼RL
F R ∼ R R \mathrm{FR}\sim\mathrm{RR} FR∼RR
同侧腿同步。
5.4 Bound
ψ = 0 , 0 , 0.5 , 0.5 \psi=0,0,0.5,0.5 ψ=0,0,0.5,0.5
前腿同步,后腿同步。
5.5 Rotary Gallop
ψ = 0 , 0.2 , 0.7 , 0.5 \psi=0,0.2,0.7,0.5 ψ=0,0.2,0.7,0.5
四腿依次错相,形成更复杂的高速旋转式 Gallop 接触序列。
6. Compliant Contact Tracking:论文最重要的低层控制创新
传统 Gait-Conditioned Policy 往往加入严格接触奖励:
c i ( t ) ≈ c ^ i ( t ) c_i(t)\approx\hat{c}_i(t) ci(t)≈c^i(t)
其中:
- c i ( t ) c_i(t) ci(t):真实足端接触状态;
- c ^ i ( t ) \hat{c}_i(t) c^i(t):目标接触状态。
严格跟踪会产生一个问题:
当前环境要求临时提前落脚,但奖励函数却强迫机器人继续保持 Swing Phase。
这样可能导致机器人跌倒。
论文将误差重新定义为:
ϕ e r r o r c o m p l y = { 0 , ϕ e r r o r ≤ δ , ϕ e r r o r , otherwise . \phi_{\mathrm{error}}^{\mathrm{comply}} =\begin{cases} 0, & \phi_{\mathrm{error}}\le\delta, \\ \phi_{\mathrm{error}}, & \text{otherwise}. \end{cases} ϕerrorcomply={0,ϕerror,ϕerror≤δ,otherwise.
其中:
δ \delta δ
为 Compliance Threshold(柔顺阈值)。
实验统一采用:
δ = 0.5 \delta=0.5 δ=0.5
直观理解是:
在步态周期约 50% 范围内允许实际接触偏离理想时刻,不立即施加接触跟踪惩罚。
这不是放弃步态跟踪,而是把目标从:
Strict Contact Tracking \boxed{\text{Strict Contact Tracking}} Strict Contact Tracking
改为:
Track When Possible, Deviate When Necessary \boxed{\text{Track When Possible, Deviate When Necessary}} Track When Possible, Deviate When Necessary
即"能按计划走时按计划走,需要保命时允许违反计划"。
6.1 Contact Reward
附录进一步给出:
r c o n t a c t = exp ( − ϕ e r r o r σ ) r_{\mathrm{contact}} =\exp\left( -\frac{\phi_{\mathrm{error}}}{\sigma} \right) rcontact=exp(−σϕerror)
其中:
σ = 0.25 \sigma=0.25 σ=0.25
为 Smoothing Factor。
指数核使接触误差从"硬开关"变成连续平滑的奖励变化,有利于 PPO 训练。
7. 为什么 Compliant Tracking 能提高鲁棒性
机器人在粗糙地形上真正需要满足的是:
Dynamic Feasibility \text{Dynamic Feasibility} Dynamic Feasibility
而不只是:
Reference Gait Fidelity \text{Reference Gait Fidelity} Reference Gait Fidelity
如果严格接触调度:
c ^ ( t ) \hat{c}(t) c^(t)
与当前地形动力学冲突,那么最优行为实际上可能是:
c ( t ) ≠ c ^ ( t ) c(t)\neq\hat{c}(t) c(t)=c^(t)
例如:
- 提前触地;
- 延迟离地;
- 某一步缩短 Swing;
- 临时延长 Stance。
因此柔顺阈值相当于给 Gait Schedule 加入:
Slack \text{Slack} Slack
其作用和优化控制中软约束的思想类似:
风格是软约束,稳定性是更高优先级目标。
8. 五种步态跟踪结果

图 5 五种步态的接触状态与机器人姿态。来源:Nahrendra et al., LocoVLM, Figure 5。上排是四足接触状态,下排是真实/仿真机器人姿态,分别对应 (a) pronk、(b) trot、© pace、(d) bound、(e) rotary gallop。
论文通过这个实验验证:
( T , ψ ) (T,\psi) (T,ψ)
确实足以让同一个 Locomotion Policy 表达多种基础 gait,而不是为每个步态单独训练一个 policy。
9. Gait Tracking 定量实验
论文设置:
v x c m d = 1.2 m / s v_x^{\mathrm{cmd}} =1.2\ \mathrm{m/s} vxcmd=1.2 m/s
T = 0.4 s T=0.4\ \mathrm{s} T=0.4 s
最大 Episode:
20 s 20\ \mathrm{s} 20 s
每种配置:
1000 rollouts 1000\ \text{rollouts} 1000 rollouts
评价指标是机器人在 episode 内平均能够走多远。
表 2 原论文 Table I:不同 Compliance Threshold 下的平均行走距离。
| Gait | δ \delta δ | Rough / m | Discrete / m | Stairs / m |
|---|---|---|---|---|
| Pronk | 0 | 15.13 ± 3.61 15.13\pm3.61 15.13±3.61 | 17.34 ± 6.41 17.34\pm6.41 17.34±6.41 | 13.56 ± 2.53 13.56\pm2.53 13.56±2.53 |
| Pronk | 0.25 | 15.82 ± 3.43 15.82\pm3.43 15.82±3.43 | 16.13 ± 6.65 16.13\pm6.65 16.13±6.65 | 14.78 ± 2.69 14.78\pm2.69 14.78±2.69 |
| Pronk | 0.5 | 15.59 ± 3.38 15.59\pm3.38 15.59±3.38 | 18.42 ± 6.29 18.42\pm6.29 18.42±6.29 | 14.06 ± 2.74 14.06\pm2.74 14.06±2.74 |
| Pronk | 0.75 | 16.98 ± 4.05 16.98\pm4.05 16.98±4.05 | 16.79 ± 6.64 16.79\pm6.64 16.79±6.64 | 14.07 ± 5.13 14.07\pm5.13 14.07±5.13 |
| Trot | 0 | 16.72 ± 3.53 16.72\pm3.53 16.72±3.53 | 16.62 ± 6.50 16.62\pm6.50 16.62±6.50 | 14.92 ± 2.82 14.92\pm2.82 14.92±2.82 |
| Trot | 0.25 | 14.98 ± 4.03 14.98\pm4.03 14.98±4.03 | 18.96 ± 6.14 18.96\pm6.14 18.96±6.14 | 15.94 ± 2.86 15.94\pm2.86 15.94±2.86 |
| Trot | 0.5 | 16.84 ± 3.72 16.84\pm3.72 16.84±3.72 | 18.29 ± 6.18 18.29\pm6.18 18.29±6.18 | 16.34 ± 2.41 16.34\pm2.41 16.34±2.41 |
| Trot | 0.75 | 15.32 ± 3.04 15.32\pm3.04 15.32±3.04 | 15.59 ± 5.89 15.59\pm5.89 15.59±5.89 | 15.70 ± 2.32 15.70\pm2.32 15.70±2.32 |
| Pace | 0 | 17.11 ± 3.51 17.11\pm3.51 17.11±3.51 | 17.48 ± 5.89 17.48\pm5.89 17.48±5.89 | 16.83 ± 2.35 16.83\pm2.35 16.83±2.35 |
| Pace | 0.25 | 18.48 ± 4.45 18.48\pm4.45 18.48±4.45 | 18.43 ± 6.58 18.43\pm6.58 18.43±6.58 | 16.71 ± 3.40 16.71\pm3.40 16.71±3.40 |
| Pace | 0.5 | 16.31 ± 3.95 16.31\pm3.95 16.31±3.95 | 19.49 ± 5.77 19.49\pm5.77 19.49±5.77 | 18.18 ± 2.49 18.18\pm2.49 18.18±2.49 |
| Pace | 0.75 | 17.23 ± 4.19 17.23\pm4.19 17.23±4.19 | 18.05 ± 5.67 18.05\pm5.67 18.05±5.67 | 16.07 ± 4.15 16.07\pm4.15 16.07±4.15 |
| Bound | 0 | 14.57 ± 2.91 14.57\pm2.91 14.57±2.91 | 14.53 ± 5.64 14.53\pm5.64 14.53±5.64 | 14.17 ± 2.29 14.17\pm2.29 14.17±2.29 |
| Bound | 0.25 | 16.99 ± 3.98 16.99\pm3.98 16.99±3.98 | 16.74 ± 5.44 16.74\pm5.44 16.74±5.44 | 15.53 ± 2.55 15.53\pm2.55 15.53±2.55 |
| Bound | 0.5 | 16.04 ± 3.57 16.04\pm3.57 16.04±3.57 | 17.43 ± 5.16 17.43\pm5.16 17.43±5.16 | 15.48 ± 2.92 15.48\pm2.92 15.48±2.92 |
| Bound | 0.75 | 16.66 ± 4.40 16.66\pm4.40 16.66±4.40 | 17.47 ± 6.33 17.47\pm6.33 17.47±6.33 | 14.39 ± 4.13 14.39\pm4.13 14.39±4.13 |
| Rotary gallop | 0 | 16.67 ± 3.56 16.67\pm3.56 16.67±3.56 | 16.71 ± 6.20 16.71\pm6.20 16.71±6.20 | 16.40 ± 2.88 16.40\pm2.88 16.40±2.88 |
| Rotary gallop | 0.25 | 17.89 ± 4.15 17.89\pm4.15 17.89±4.15 | 17.77 ± 5.69 17.77\pm5.69 17.77±5.69 | 17.17 ± 3.02 17.17\pm3.02 17.17±3.02 |
| Rotary gallop | 0.5 | 18.18 ± 3.91 18.18\pm3.91 18.18±3.91 | 18.61 ± 5.08 18.61\pm5.08 18.61±5.08 | 18.04 ± 2.10 18.04\pm2.10 18.04±2.10 |
| Rotary gallop | 0.75 | 16.97 ± 4.68 16.97\pm4.68 16.97±4.68 | 18.55 ± 5.88 18.55\pm5.88 18.55±5.88 | 16.38 ± 3.94 16.38\pm3.94 16.38±3.94 |
加粗值是同一 gait、同一 terrain 下论文表格中的最佳结果。
9.1 实验说明了什么
并不是 δ \delta δ 越大越好,而是存在:
Style Fidelity ↔ Robustness \text{Style Fidelity} \leftrightarrow \text{Robustness} Style Fidelity↔Robustness
之间的折中。
过小:
δ → 0 \delta\rightarrow0 δ→0
接近严格接触跟踪,鲁棒性降低。
过大:
δ → 1 \delta\rightarrow1 δ→1
又可能使 gait constraint 过弱。
作者最终在全部实验中统一选择:
δ = 0.5 \boxed{\delta=0.5} δ=0.5
作为整体折中,而不是针对每一种 gait 和地形单独调最优参数。
10. 离线 Skill Database:如何把 GPT-4o 的知识"蒸馏"下来

图 3 Offline Skill Database Generation Pipeline。来源:Nahrendra et al., LocoVLM, Figure 3。第一阶段先生成语言指令,第二阶段通过 Meta-Prompt 将指令转换成结构化 Skill Database 条目。
作者没有人工写几千条:
text
language → gait parameters
而是利用 GPT-4o 自动扩展。
生成分为两步。
10.1 Stage 1:Instruction Description Generation
GPT-4o 先生成三类 instruction。
Mimicking Behaviors
例如:
text
let's hop like a rabbit
run beautifully like a horse
you are a kangaroo
Scene Responses
例如:
text
there is an icy patch
this is a library
the snow is slippery
Direct Instructions
例如:
text
trot slowly
bound quickly
reduce your noise
每次要求模型产生:
n = 100 n=100 n=100
条指令。
采用分类生成而不是把全部类型一次性混在一起,可以降低重复和无结构输出。
10.2 Stage 2:Motion Descriptor Generation
每条语言指令被转换成:
M = ( T , ψ , v x l i m i t ) \mathbf{M} =(T,\psi,v_x^{\mathrm{limit}}) M=(T,ψ,vxlimit)
完整数据库条目可以表示为:
d = ( I , M ) d =(\mathcal{I},\mathbf{M}) d=(I,M)
实际 JSON 还包含:
text
instruction
reasoning
T
gait_phase_offsets
vel_lim
11. 为什么要加入 Prompted Reasoning
如果输入:
text
trot quickly
LLM 很容易映射到:
- Trot Phase Offsets;
- 较小 T T T;
- 较大 v x l i m i t v_x^{\mathrm{limit}} vxlimit。
但输入:
text
trundle along like a hippo
LLM 必须先理解:
河马 → 沉重、缓慢 → 步频低、速度低。
所以作者要求 GPT-4o 在给出数值前先生成中间 reasoning。
例如:
"trundle along like a hippo" \text{"trundle along like a hippo"} "trundle along like a hippo"
先推理成:
slow and heavy trot \text{slow and heavy trot} slow and heavy trot
再得到:
T ↑ , v x l i m i t ↓ T\uparrow, \qquad v_x^{\mathrm{limit}}\downarrow T↑,vxlimit↓
这相当于:
High-Level Semantics → Technical Motion Semantics → Numerical Descriptor \text{High-Level Semantics} \rightarrow \text{Technical Motion Semantics} \rightarrow \text{Numerical Descriptor} High-Level Semantics→Technical Motion Semantics→Numerical Descriptor
比直接:
Language → Numbers \text{Language} \rightarrow \text{Numbers} Language→Numbers
更加稳定。
12. Prompted Reasoning 对数据库质量的影响

图 6 三种技能数据库生成方式的统计分布。来源:Nahrendra et al., LocoVLM, Figure 6。左:类似 SayTap 的逐条生成 Baseline;中:LocoVLM Batch Generation、无 Prompted Reasoning;右:LocoVLM + Prompted Reasoning。
实验统一生成:
300 300 300
条 Motion Descriptors。
12.1 Gait 分布
Baseline 中:
45.7 % 45.7\% 45.7%
被归类为 Others,即不属于五种稳定标准 gait 的非结构化 phase offsets。
LocoVLM 无 reasoning:
25.7 % 25.7\% 25.7%
LocoVLM + Prompted Reasoning:
5.3 % 5.3\% 5.3%
说明显式推理显著减少了无结构步态参数。
表 3 Figure 6(a) 的 Gait Category Distribution。
| 方法 | Trot | Bound | Pace | Pronk | Rotary gallop | Others |
|---|---|---|---|---|---|---|
| Baseline | 32.3% | 4.0% | 5.0% | 4.0% | 9.0% | 45.7% |
| LocoVLM w/o reasoning | 44.0% | 9.3% | 12.7% | 2.3% | 6.0% | 25.7% |
| LocoVLM + reasoning | 53.7% | 10.7% | 10.0% | 8.0% | 12.3% | 5.3% |
12.2 Gait Cycle 分布
Prompted Reasoning 后, T T T 主要分布在:
0.2 ∼ 0.7 s 0.2\sim0.7\ \mathrm{s} 0.2∼0.7 s
而 Baseline 与无 reasoning 版本更集中在约:
0.5 s 0.5\ \mathrm{s} 0.5 s
并出现:
T ≈ 1.0 s T\approx1.0\ \mathrm{s} T≈1.0 s
的异常/不稳定长周期。
12.3 Velocity Limit
在:
v x l i m i t v_x^{\mathrm{limit}} vxlimit
分布上三种方法差异没有 gait phase 和 cycle 那么明显。
论文认为原因是速度语义相对容易:
text
quickly → high velocity
slowly → low velocity
但"像某种动物""在某种社会环境下如何走"到 gait phase 的映射更复杂,因此 reasoning 的帮助更明显。
13. 生成成本
生成 300 个 Motion Descriptors:
| 方法 | GPT-4o 生成成本 |
|---|---|
| Baseline | $$1.16$ |
| LocoVLM without prompted reasoning | $$0.21$ |
| LocoVLM with prompted reasoning | $$0.25$ |
批量生成的成本只有逐条查询的一小部分。
同时 Batch Generation 相当于给模型提供了一个局部"上下文记忆",可以减少同一批次中重复描述符的产生。
14. Motion Descriptor 为什么只选三个参数
论文定义:
M = ( T , ψ , v x l i m i t ) \boxed{ \mathbf{M} =(T,\psi,v_x^{\mathrm{limit}}) } M=(T,ψ,vxlimit)
这是一个很重要的工程取舍。
如果让 VLM 输出:
- 12 个关节角;
- 足端轨迹;
- GRF;
- Body Pose;
- Torque;
高层语义与动作空间之间会变得非常难对齐。
而:
( T , ψ , v x l i m i t ) (T,\psi,v_x^{\mathrm{limit}}) (T,ψ,vxlimit)
恰好具有三个特点:
- 维度低;
- 人类可解释;
- 足够控制多种 gait 和速度。
因此它相当于机器人运动系统的一个 Semantic Control Interface(语义控制接口)。
15. Vision-Language Grounding:Skill Retrieval

图 4 Skill Database Retrieval。来源:Nahrendra et al., LocoVLM, Figure 4。文本或图像 Query 被编码到 BLIP-2 的共享语义空间,与技能数据库中的 instruction 表示进行匹配,最近的 instruction 对应一个 reasoning 和 motion descriptor。
数据库:
D = { d i } \mathcal{D} =\{d_i\} D={di}
其中:
d i = ( I i , M i ) d_i=(\mathcal{I}_i,\mathbf{M}_i) di=(Ii,Mi)
给定 Query:
I q u e r y \mathcal{I}_{\mathrm{query}} Iquery
目标是寻找:
I ∗ = arg max I ∈ D sim ( I q u e r y , I ) \mathcal{I}^{*} =\arg\max_{\mathcal{I}\in\mathcal{D}} \operatorname{sim} \left( \mathcal{I}_{\mathrm{query}}, \mathcal{I} \right) I∗=argI∈Dmaxsim(Iquery,I)
然后:
I ∗ → M ∗ \mathcal{I}^{*} \rightarrow \mathbf{M}^{*} I∗→M∗
将:
M ∗ \mathbf{M}^{*} M∗
发送到低层运动策略。
16. Mixed-Precision Retrieval 算法
16.1 Stage 1:Cosine Similarity 进行快速 Top-K 检索
设:
f B L I P ( ⋅ ) f_{\mathrm{BLIP}}(\cdot) fBLIP(⋅)
是 BLIP-2 Encoder。
先求:
I K = TopK I ∈ D cossim ( f B L I P ( I q u e r y ) , f B L I P ( I ) ) \mathbf{I}^{K} =\operatorname{TopK}{\mathcal{I}\in\mathcal{D}} \operatorname{cossim} \left( f{\mathrm{BLIP}}(\mathcal{I}{\mathrm{query}}), f{\mathrm{BLIP}}(\mathcal{I}) \right) IK=TopKI∈Dcossim(fBLIP(Iquery),fBLIP(I))
只保留最相似的 K K K 个候选。
然后将这些候选的 Cosine Similarity 转成:
p 1 ( I K ) = softmax ( cossim ( f B L I P ( I q u e r y ) , f B L I P ( I K ) ) ) p_1(\mathbf{I}^{K}) =\operatorname{softmax} \left( \operatorname{cossim} \left( f_{\mathrm{BLIP}}(\mathcal{I}{\mathrm{query}}), f{\mathrm{BLIP}}(\mathbf{I}^{K}) \right) \right) p1(IK)=softmax(cossim(fBLIP(Iquery),fBLIP(IK)))
16.2 Stage 2:ITM Head 重排序
对每个:
I k ∈ I K \mathcal{I}_k\in\mathbf{I}^{K} Ik∈IK
利用 BLIP-2 Image-Text Matching Head:
p 2 ( I k ) = softmax ( f I T M ( I q u e r y , I k ) ) p_2(\mathcal{I}k) =\operatorname{softmax} \left( f{\mathrm{ITM}} ( \mathcal{I}_{\mathrm{query}}, \mathcal{I}_k ) \right) p2(Ik)=softmax(fITM(Iquery,Ik))
最后:
I ∗ = arg max I k p 1 ( I k ) + p 2 ( I k ) \mathcal{I}^{*} =\arg\max_{\mathcal{I}_k} \left p_1(\\mathcal{I}_k) + p_2(\\mathcal{I}_k) \\right I∗=argIkmaxp1(Ik)+p2(Ik)
得到最终技能。
16.3 Algorithm 1
text
Input:
Query I_query
Database D
BLIP encoder f_BLIP
ITM head f_ITM
Top-K size K
1. 使用 BLIP embedding + cosine similarity 从 D 中取 Top-K
2. 对 Top-K cosine similarity 做 softmax,得到 p1
3. 对 Top-K 中每一条候选执行 BLIP-2 ITM
4. 得到 ITM matching probability p2
5. 将 p1 与 p2 组合
6. 选择综合概率最大的 instruction I*
7. 读取其 motion descriptor M*
时间复杂度直观上由:
O ( N d ) + O ( K C I T M ) O(Nd) + O(KC_{\mathrm{ITM}}) O(Nd)+O(KCITM)
组成。
其中:
K ≪ N K\ll N K≪N
所以避免了:
O ( N C I T M ) O(NC_{\mathrm{ITM}}) O(NCITM)
式的全数据库 ITM 暴力比较。
17. Text-as-Image:一个很反直觉但有效的技巧
BLIP-2 本质上主要学习:
Image ↔ Text \text{Image} \leftrightarrow \text{Text} Image↔Text
的对齐。
它不一定特别擅长:
Text ↔ Text \text{Text} \leftrightarrow \text{Text} Text↔Text
的句子级匹配。
作者于是把用户文本:
text
shh! the baby is sleeping
绘制为:
白色背景 + 黑色文字
得到一张"文字图片",再送进 VLM Image Encoder。
于是原本的:
Text Query → Text Retrieval \text{Text Query} \rightarrow \text{Text Retrieval} Text Query→Text Retrieval
变成:
Rendered Text Image → Image-Text Matching \text{Rendered Text Image} \rightarrow \text{Image-Text Matching} Rendered Text Image→Image-Text Matching
相当于把问题重新投影到 BLIP-2 最擅长的训练分布。
这是论文中一个很有工程价值的小技巧。
18. Retrieval Accuracy
作者人工标注:
100 100 100
条 instruction 作为检索评估集。
表 4 原论文 Table II:不同检索方法的准确率。
| Retrieval Metric | Text as String | Text as Image | Average |
|---|---|---|---|
| Cosine similarity | 21/100 | 30/100 | 20.5%(原论文数值) |
| Top-K similarity | 27/100 | 48/100 | 37.5% |
| Top-K to ITM | 51/100 | 57/100 | 54.0% |
| Mixed-Precision | 72/100 | 87/100 | 79.5% |
最终最佳组合:
Text-as-Image + Mixed-Precision Retrieval = 87 % \boxed{ \text{Text-as-Image} + \text{Mixed-Precision Retrieval} =87\% } Text-as-Image+Mixed-Precision Retrieval=87%
论文摘要所说的最高 87% Instruction-Following Accuracy,需要结合这个实验口径理解:它主要对应这里的 100 条人工标注 instruction 上的检索/匹配准确率,而不是"任意开放世界指令 87% 成功率"。
19. 超出数据库的语义泛化
LocoVLM 不要求用户输入与数据库字符串完全一致。
例如数据库可能存在:
text
let's hop like a rabbit
用户输入:
text
you are a kangaroo
BLIP-2 会因为:
kangaroo ∼ jump ∼ rabbit hop \text{kangaroo} \sim \text{jump} \sim \text{rabbit hop} kangaroo∼jump∼rabbit hop
在语义 embedding 中找到相近技能。
同样:
text
this is a library
可以检索到:
text
move quietly
这种能力并不是 LocoVLM 在线重新推理出新的控制器,而是:
Novel Query → Nearest Semantically Compatible Skill \text{Novel Query} \rightarrow \text{Nearest Semantically Compatible Skill} Novel Query→Nearest Semantically Compatible Skill
因此它属于 Retrieval-Based Semantic Grounding(基于检索的语义落地)。
20. Robot-Centric Vision:真正把视觉环境语义作用到 Gait

图 7 Robot-Centric RGB Scene Interpretation。来源:Nahrendra et al., LocoVLM, Figure 7。机器人从普通路面进入雪地后,VLM 根据视觉场景检索出不同的 Motion Descriptor。
20.1 普通路面
VLM 检索到类似:
text
traipse lightly like a deer
输出:
T = 0.5 s T=0.5\ \mathrm{s} T=0.5 s
v x l i m i t = 0.6 m / s v_x^{\mathrm{limit}} =0.6\ \mathrm{m/s} vxlimit=0.6 m/s
并使用 Trot。
20.2 冰雪区域
VLM 输出:
text
a field of ice, walk light-footed
或:
text
skulk with stealth like a lynx
速度降低到:
0.2 ∼ 0.3 m / s 0.2\sim0.3\ \mathrm{m/s} 0.2∼0.3 m/s
周期增加到:
0.6 ∼ 0.7 s 0.6\sim0.7\ \mathrm{s} 0.6∼0.7 s
即:
Snow / Ice → Cautious Semantics → v x l i m i t ↓ , T ↑ \text{Snow / Ice} \rightarrow \text{Cautious Semantics} \rightarrow v_x^{\mathrm{limit}}\downarrow, \quad T\uparrow Snow / Ice→Cautious Semantics→vxlimit↓,T↑
这就是论文标题中 Grounding Vision 最直接的体现。
关键点不是"识别出这是雪",而是:
视觉场景语义最终改变了机器人能够执行的 gait style 和速度限制。
21. 真机系统配置
21.1 Unitree Go1 Locomotion Controller
训练算法:
PPO \text{PPO} PPO
结构:
Asymmetric Actor-Critic \text{Asymmetric Actor-Critic} Asymmetric Actor-Critic
并结合 State Estimation。
训练仿真器:
Isaac Gym \text{Isaac Gym} Isaac Gym
为了 Sim-to-Real,随机化:
- Robot Mass;
- Center of Mass;
- Motor Stiffness;
- Motor Damping;
- Terrain Friction;
- System Delay。
21.2 真机执行频率
RL Policy:
50 H z 50\ \mathrm{Hz} 50 Hz
输出:
q d e s q^{\mathrm{des}} qdes
Go1 Motor Controller:
200 H z 200\ \mathrm{Hz} 200 Hz
将目标关节角转换为电机力矩。
板载计算:
Jetson Xavier NX \text{Jetson Xavier NX} Jetson Xavier NX
21.3 VLM 模块
单独计算机:
NVIDIA RTX 3070 Ti \text{NVIDIA RTX 3070 Ti} NVIDIA RTX 3070 Ti
BLIP-2 通过 ROS 与机器人通信。
VLM 推理 + 数据通信:
< 100 m s <100\ \mathrm{ms} <100 ms
VLM Advisor 与 Locomotion Policy:
Asynchronous \boxed{\text{Asynchronous}} Asynchronous
因此低层 50 Hz Policy 不需要等待每一次 VLM 推理完成。
这一点对真机非常重要:
Semantic Loop Frequency ≪ Locomotion Control Frequency \text{Semantic Loop Frequency} \ll \text{Locomotion Control Frequency} Semantic Loop Frequency≪Locomotion Control Frequency
22. 为什么异步架构比在线 LLM 控制更合理
可以把系统拆成三个时间尺度。
高频闭环
50 ∼ 200 H z 50\sim200\ \mathrm{Hz} 50∼200 Hz
负责:
- 平衡;
- 关节动作;
- 接触恢复。
中低频语义层
∼ 100 m s \sim100\ \mathrm{ms} ∼100 ms
负责:
- Scene Interpretation;
- Language Grounding;
- Skill Retrieval。
极低频离线知识生成
GPT-4o 只在数据库生成阶段运行。
因此即使 Foundation Model 有延迟,也不会直接破坏:
Dynamic Stability Loop \text{Dynamic Stability Loop} Dynamic Stability Loop
这种设计对四足/人形机器人比让一个大型 VLM 直接生成 50 Hz 关节动作更工程化。
23. Zero-Shot Cross-Embodiment Generalization

图 8 LocoVLM 在 Unitree H1 人形机器人上的 Zero-Shot Skill Database Transfer。来源:Nahrendra et al., LocoVLM, Figure 8。相同的四足语言技能数据库被复用于 H1,执行 "go quickly""shh! the baby is sleeping""you are a kangaroo"等命令。
作者没有把 Go1 Policy 直接放到 H1 上。
真正复用的是:
Skill Database \boxed{\text{Skill Database}} Skill Database
H1 重新训练了自己的 Style-Conditioned Locomotion Policy,但只使用:
ψ l e f t , ψ r i g h t \psi_{\mathrm{left}}, \psi_{\mathrm{right}} ψleft,ψright
两个腿部 Phase Offsets。
H1 只训练:
- Trot-Like Alternating Gait;
- Pronk-Like Hopping Gait。
然后直接使用原先为 Quadruped 生成的数据库中的前两个 Phase Offsets。
所以"Zero-Shot Across Embodiments"准确含义是:
VLM 与 Skill Database 不需要重新训练/重新生成,而目标 embodiment 仍然需要自己的低层 Locomotion Policy。
不能误解成:
"Go1 的 RL Policy 零样本直接控制 H1。"
24. Appendix Figure 9:技能数据库生成实现细节

图 9 Appendix 中再次给出的 Offline Skill Database Generation Pipeline。来源:Nahrendra et al., LocoVLM, Figure 9。该图与正文 Figure 3 对应同一数据生成流程,附录结合 Prompt Listings 给出更多实现细节。
Appendix 进一步说明:
- 三种 instruction category 分开生成;
- 每类明确要求 LLM 输出 n = 100 n=100 n=100 条;
- 再把 instruction 送入统一 Meta-Prompt;
- 输入列表会 shuffle;
- 防止 LLM 记忆固定 input-output 顺序;
- 输出保存成结构化 JSON。
25. 原论文 Prompt 的关键信息
Skill Prompt 告诉 GPT-4o:
T T T
0.2 < T ≤ 1 0.2<T\le1 0.2<T≤1
并特别提示:
0.3 ≤ T ≤ 0.6 0.3\le T\le0.6 0.3≤T≤0.6
通常更稳定。
Gait Phase Offsets
使用五种基础 gait 作为参考:
text
trot [0.0, 0.5, 0.5, 0.0]
rotary_gallop [0.0, 0.2, 0.7, 0.5]
pace [0.0, 0.5, 0.0, 0.5]
pronk [0.0, 0.0, 0.0, 0.0]
bound [0.0, 0.0, 0.5, 0.5]
但 Prompt 同时强调不要只复制 gait dictionary,而要根据语义创造新的描述。
Velocity Limit
Prompt 告诉模型:
v x l i m i t ≤ 1.5 m / s v_x^{\mathrm{limit}} \le 1.5\ \mathrm{m/s} vxlimit≤1.5 m/s
并通过:
- Slow;
- Fast;
- Stop;
等语义调整数值。
这相当于在 Foundation Model 上增加一个机器人可执行空间的人工先验边界。
26. Reasoning 示例
表 5 原论文 Table IV:语言指令与 GPT-4o Reasoning。
| Instruction | Reasoning |
|---|---|
| trundle along like a hippo | slow and heavy trot, lower vel_lim and increase T |
| oh no! catch that thief running! | fast and aggressive gait, low T and high vel_lim |
| the sound of a human voice, stay hidden. | use a trot with high T for stealthy movement, low vel_lim for quietness |
| a busy marketplace, navigate through the crowd. | slow pace with moderate T for careful navigation |
可以看到 Prompted Reasoning 实际完成了:
Semantic Concept → Locomotion Concept \text{Semantic Concept} \rightarrow \text{Locomotion Concept} Semantic Concept→Locomotion Concept
例如:
stealth → slow gait + high T + low velocity \text{stealth} \rightarrow \text{slow gait} + \text{high } T + \text{low velocity} stealth→slow gait+high T+low velocity
这一步是 LLM 知识真正进入机器人 Gait Parameter Space 的桥梁。
27. Appendix:将视觉语义作为导航约束
这部分很值得注意,因为它说明 LocoVLM 不只能改变 gait,还可以给传统 Navigation Stack 提供语义速度约束。
系统仅取:
v x l i m i t v_x^{\mathrm{limit}} vxlimit
然后把它作为 Local Planner 的最大速度约束。
27.1 从宽阔区域进入狭窄区域

图 10 LocoVLM 输出的速度上限对局部规划器的约束。来源:Nahrendra et al., LocoVLM, Figure 10。机器人进入狭窄、拥挤区域后,VLM 降低 v x l i m i t v_x^{\mathrm{limit}} vxlimit,局部规划器仍自行计算具体速度,但必须满足该上限。
这里:
v c m d ≤ v x l i m i t v_{\mathrm{cmd}} \le v_x^{\mathrm{limit}} vcmd≤vxlimit
VLM 不替代 Local Planner,而是在它外面增加 Semantic Constraint。
论文将 VLM Inference Period 设置为:
5 s 5\ \mathrm{s} 5 s
避免速度上限频繁变化导致 jitter。
27.2 轨迹安全性

图 11 有无 LocoVLM 语义速度约束时的导航轨迹对比。来源:Nahrendra et al., LocoVLM, Figure 11。绿色为加入 LocoVLM Constraint,粉色为无该约束;加入语义速度限制后,机器人在狭窄障碍区域保持更大的 Obstacle Clearance。
这个实验的意义在于:
Foundation Model \text{Foundation Model} Foundation Model
不一定非要直接输出机器人 Action。
它也可以输出:
Constraint \text{Constraint} Constraint
再让传统 Planner / Controller 在约束内工作。
这是一种非常实用的:
Semantic Constraint + Classical/RL Control \boxed{ \text{Semantic Constraint} + \text{Classical/RL Control} } Semantic Constraint+Classical/RL Control
架构。
28. LocoVLM 完整数据流
text
OFFLINE
┌─────────────────────────────────────────────────────────────┐
│ │
│ Prompt ──► GPT-4o ──► Instruction Set │
│ │ │
│ ▼ │
│ Meta-Prompt │
│ │ │
│ ▼ │
│ {instruction, reasoning, T, phase offsets, vel_lim} │
│ │ │
│ ▼ │
│ Skill Database D │
│ │
└─────────────────────────────────────────────────────────────┘
ONLINE
┌─────────────────────────────────────────────────────────────┐
│ │
│ Text Query ─────┐ │
│ ├──► BLIP-2 Encoder ─► Cosine Top-K │
│ RGB Image ──────┘ │ │
│ ▼ │
│ BLIP-2 ITM Re-rank │
│ │ │
│ ▼ │
│ Motion Descriptor M* │
│ (T, ψ, vel_lim) │
│ │ │
│ Proprioception ────────────────┤ │
│ Velocity Command ──────────────┤ │
│ ▼ │
│ Style-Conditioned PPO Policy │
│ │ │
│ ▼ │
│ Desired Joint Angles │
│ │ │
│ ▼ │
│ Unitree Go1 │
│ │
└─────────────────────────────────────────────────────────────┘
29. PPO 在这篇论文里负责什么
PPO 并不负责理解语言。
语言/视觉语义部分由:
GPT-4o + BLIP-2 \text{GPT-4o} + \text{BLIP-2} GPT-4o+BLIP-2
完成。
PPO 解决的是:
( s t , T , ψ , v x ) → a t (s_t,T,\psi,v_x) \rightarrow a_t (st,T,ψ,vx)→at
即:
给定机器人本体状态、步态周期、相位关系和速度命令,生成稳定可执行的运动。
典型 PPO Probability Ratio:
r t ( θ ) = π θ ( a t ∣ s t ) π θ o l d ( a t ∣ s t ) r_t(\theta) =\frac{ \pi_\theta(a_t\mid s_t) }{ \pi_{\theta_{\mathrm{old}}}(a_t\mid s_t) } rt(θ)=πθold(at∣st)πθ(at∣st)
Clipped Objective(裁剪目标函数):
L C L I P = E t min ( r t ( θ ) A \^ t , clip ( r t ( θ ) , 1 − ϵ , 1 + ϵ ) A \^ t ) L^{\mathrm{CLIP}} =\mathbb{E}_t \left \\min \\left( r_t(\\theta)\\hat{A}_t, \\operatorname{clip} \\left( r_t(\\theta), 1-\\epsilon, 1+\\epsilon \\right) \\hat{A}_t \\right) \\right LCLIP=Etmin(rt(θ)A\^t,clip(rt(θ),1−ϵ,1+ϵ)A\^t)
PPO 的价值是让 Locomotion Policy 在高维连续控制空间中稳定迭代,而 LocoVLM 的新意并不在 PPO 算法本身,而在:
Style Conditioning + Compliant Contact Reward + Semantic Skill Interface \text{Style Conditioning} + \text{Compliant Contact Reward} + \text{Semantic Skill Interface} Style Conditioning+Compliant Contact Reward+Semantic Skill Interface
30. Asymmetric Actor-Critic 的含义
论文采用 Asymmetric Actor-Critic。
其一般形式是:
a t ∼ π θ ( a t ∣ o t ) a_t \sim \pi_\theta(a_t\mid o_t) at∼πθ(at∣ot)
Actor 只使用真机可获得的 Observation:
o t o_t ot
而 Critic 在训练阶段允许使用更完整的 Privileged State:
V ϕ ( s t ) V_\phi(s_t) Vϕ(st)
其中:
s t ⊇ o t s_t\supseteq o_t st⊇ot
这样训练时 Critic 能获得更准确的 Value Estimate,但部署时 Actor 不依赖仿真器特权信息。
这是强化学习 Sim-to-Real Locomotion 中常见的设计:
Train with Privileged Information → Deploy with Realistic Observations \boxed{ \text{Train with Privileged Information} \rightarrow \text{Deploy with Realistic Observations} } Train with Privileged Information→Deploy with Realistic Observations
需要注意:LocoVLM 正文没有进一步完整列出所有 Actor/Critic 网络层尺寸,因此不能根据其他 Locomotion 项目臆造具体 MLP 结构。
31. Sim-to-Real 设计
论文明确随机化:
m r o b o t m_{\mathrm{robot}} mrobot
C o M \mathrm{CoM} CoM
k p / k d -related motor properties k_p/k_d\text{-related motor properties} kp/kd-related motor properties
μ t e r r a i n \mu_{\mathrm{terrain}} μterrain
以及:
System Delay \text{System Delay} System Delay
目的都是扩大训练环境分布:
p s i m ( θ d y n ) p_{\mathrm{sim}}(\theta_{\mathrm{dyn}}) psim(θdyn)
让真机真实动力学:
θ r e a l \theta_{\mathrm{real}} θreal
更可能落在训练分布覆盖范围内。
因此:
Domain Randomization + State Estimation + Compliant Contact Tracking \text{Domain Randomization} + \text{State Estimation} + \text{Compliant Contact Tracking} Domain Randomization+State Estimation+Compliant Contact Tracking
共同支撑 Go1 真机部署。
32. LocoVLM 与 SayTap 的核心区别
| 对比维度 | SayTap | LocoVLM |
|---|---|---|
| 高层模型 | GPT-4 | GPT-4o + BLIP-2 |
| Vision | 无 | 有 |
| LLM 是否在线 | 是,高层在线生成 | GPT-4o 仅离线生成数据库 |
| 在线语义模块 | GPT-4 | BLIP-2 Retrieval |
| 中间表示 | Foot Contact Pattern | ( T , ψ , v x l i m i t ) (T,\psi,v_x^{\mathrm{limit}}) (T,ψ,vxlimit) |
| 数据生成 | Prompt Few-Shot | 两阶段批量生成 + Prompted Reasoning |
| 检索 | 无 | Mixed-Precision Retrieval |
| Text-as-Image | 无 | 有 |
| 低层 | RL Contact-Conditioned Policy | Style-Conditioned PPO Policy |
| 鲁棒性机制 | RL + Contact Pattern | Compliant Contact Tracking |
| 图像场景适应 | 无 | 有 |
| Cross-Embodiment | 非核心实验 | Go1 → H1 Skill Database Transfer |
可以把两者的思想演化写成:
SayTap : Language → Contact Pattern → RL \text{SayTap}: \quad \text{Language} \rightarrow \text{Contact Pattern} \rightarrow \text{RL} SayTap:Language→Contact Pattern→RL
LocoVLM : Vision/Language → Semantic Retrieval → ( T , ψ , v l i m ) → RL \text{LocoVLM}: \quad \text{Vision/Language} \rightarrow \text{Semantic Retrieval} \rightarrow (T,\psi,v_{\mathrm{lim}}) \rightarrow \text{RL} LocoVLM:Vision/Language→Semantic Retrieval→(T,ψ,vlim)→RL
LocoVLM 很明显吸收了 SayTap 的"Foundation Model 不直接控制关节,而输出运动中间表示"的思想,但进一步解决了:
- Vision Grounding;
- 在线 LLM 依赖;
- 数据扩展成本;
- 检索效率;
- 严格 Gait Tracking 的鲁棒性问题。
33. LocoVLM 是不是严格意义上的 VLA
严格按照当前主流 Vision-Language-Action 模型定义,例如:
( Image , Language ) → Single Learned Policy Action Chunk (\text{Image},\text{Language}) \xrightarrow{\text{Single Learned Policy}} \text{Action Chunk} (Image,Language)Single Learned Policy Action Chunk
LocoVLM 不是典型端到端 VLA。
它更准确的结构是:
( Image or Language ) → B L I P - 2 Retrieved Motion Descriptor → R L Joint Action (\text{Image or Language}) \xrightarrow{\mathrm{BLIP\text{-}2}} \text{Retrieved Motion Descriptor} \xrightarrow{\mathrm{RL}} \text{Joint Action} (Image or Language)BLIP-2 Retrieved Motion DescriptorRL Joint Action
因此它属于:
Hierarchical Vision-Language-Guided Locomotion Framework(分层视觉语言引导运动框架)
而不是 OpenVLA、RT-2、 π 0 \pi_0 π0 那类统一模型直接 Decode Robot Action。
但从机器人系统工程角度,它有一个很重要的启发:
VLM/VLA 不一定需要进入最内层动态控制环 \boxed{ \text{VLM/VLA 不一定需要进入最内层动态控制环} } VLM/VLA 不一定需要进入最内层动态控制环
对于四足/人形这种高速动态系统,分层设计常常更实际:
Foundation Model → Skill / Constraint → Specialized Locomotion Policy \text{Foundation Model} \rightarrow \text{Skill / Constraint} \rightarrow \text{Specialized Locomotion Policy} Foundation Model→Skill / Constraint→Specialized Locomotion Policy
34. 论文的主要创新点
34.1 Vision + Language 到 Locomotion Style 的实时 Grounding
它不是简单给机器人一句"走快点",而是把:
Semantic Scene \text{Semantic Scene} Semantic Scene
转换成:
( T , ψ , v x l i m i t ) (T,\psi,v_x^{\mathrm{limit}}) (T,ψ,vxlimit)
真正改变步态。
34.2 LLM 离线知识蒸馏
GPT-4o:
Teacher \text{Teacher} Teacher
BLIP-2 + Skill Database:
Student / Knowledge Carrier \text{Student / Knowledge Carrier} Student / Knowledge Carrier
避免云端 LLM 在线控制。
34.3 两阶段 Instruction / Motion Descriptor 数据生成
将:
Instruction Generation \text{Instruction Generation} Instruction Generation
和:
Motion Parameter Generation \text{Motion Parameter Generation} Motion Parameter Generation
分开,使数据规模扩展更便宜、更结构化。
34.4 Prompted Reasoning 提高可执行描述符质量
把:
Vague Semantic Instruction \text{Vague Semantic Instruction} Vague Semantic Instruction
先翻译成:
Technical Locomotion Reasoning \text{Technical Locomotion Reasoning} Technical Locomotion Reasoning
再输出数值参数。
34.5 Mixed-Precision Retrieval
将:
Fast Embedding Search + Accurate ITM Matching \text{Fast Embedding Search} + \text{Accurate ITM Matching} Fast Embedding Search+Accurate ITM Matching
结合,避免全库高成本 ITM。
34.6 Text-as-Image
利用 VLM 原生 Image-Text Alignment 能力弥补 Text-Text Retrieval 较弱的问题。
34.7 Compliant Contact Tracking
把 Gait Tracking 从硬约束改成柔顺目标,在粗糙地形中获得更好的 Stability / Style Trade-Off。
34.8 Cross-Embodiment Skill Database Reuse
同一个高层语义数据库可以映射到新的 embodiment,只要目标机器人拥有兼容的低层 Style-Conditioned Controller。
35. 学术贡献的本质
这篇论文最值得注意的并不是某个单一网络,而是提出了一种明确的机器人 Foundation Model 系统划分:
Semantic Intelligence ≠ Dynamic Control Intelligence \boxed{ \text{Semantic Intelligence} \neq \text{Dynamic Control Intelligence} } Semantic Intelligence=Dynamic Control Intelligence
前者交给:
LLM / VLM \text{LLM / VLM} LLM / VLM
后者交给:
Specialized RL Locomotion Policy \text{Specialized RL Locomotion Policy} Specialized RL Locomotion Policy
中间通过:
M = ( T , ψ , v x l i m i t ) \mathbf{M} =(T,\psi,v_x^{\mathrm{limit}}) M=(T,ψ,vxlimit)
连接。
这种做法避免要求一个大模型同时学习:
- 世界知识;
- 图像理解;
- 社会语义;
- Gait Planning;
- Contact Dynamics;
- 50 Hz Joint Control。
论文实际上提出了一个非常清晰的 Semantic-to-Dynamics Interface。
36. 局限性
36.1 Image 和 Text 不能真正联合输入
作者在结论中明确指出,当前 LocoVLM 分别处理:
Image \text{Image} Image
或:
Text \text{Text} Text
但不能自然完成:
( Image , Text ) (\text{Image},\text{Text}) (Image,Text)
联合条件推理。
例如:
"看到冰面以后慢走,但如果我说可以跑就恢复高速。"
当前框架无法像真正多模态 VLA 那样联合解释视觉和语言约束。
36.2 Retrieval 不是生成式 Motion Planning
LocoVLM 最终选择的是数据库中的已有技能:
I ∗ ∈ D \mathcal{I}^{*} \in \mathcal{D} I∗∈D
即使 Query 是新的,其实际行为仍来自:
Nearest Skill Entry \text{Nearest Skill Entry} Nearest Skill Entry
所以表达能力上限受到数据库覆盖度约束。
36.3 Motion Descriptor 维度仍较低
当前只有:
( T , ψ , v x l i m i t ) (T,\psi,v_x^{\mathrm{limit}}) (T,ψ,vxlimit)
没有显式描述:
- Footstep Position;
- Foothold;
- Body Height;
- Body Orientation;
- Swing Height;
- Contact Force;
- Terrain-Specific Foothold;
- Yaw Rate Limit;
- Lateral Velocity。
因此它更适合"运动风格自适应",而不是完整复杂地形足端规划。
36.4 VLM 仍依赖外部 GPU
真机 Locomotion Policy 在 Jetson Xavier NX 上运行,但 BLIP-2 放在 RTX 3070 Ti Laptop。
所以:
"无云端 LLM"
并不等于:
"所有模型都完全在 Go1 板载 Xavier NX 上运行"。
36.5 87% 不是开放世界任务成功率
87% 来源于:
100 100 100
条人工标注 Instruction 的 Retrieval Benchmark。
因此不能把它解释为:
"在任何视觉语言场景下机器人都有 87% 成功率"。
36.6 H1 泛化并非完整控制策略 Zero-Shot
复用的是:
Skill Database + VLM Mapping \text{Skill Database} + \text{VLM Mapping} Skill Database+VLM Mapping
而 H1 的 Locomotion Controller 仍然重新训练。
36.7 Prompt 中仍有较强人工机器人先验
GPT-4o 并不是自由发现所有可行 gait。
Prompt 已经提供:
- 五种 Gait Dictionary;
- T T T 合理范围;
- Velocity Limit;
- Quadruped Gait 知识。
所以本质是:
Foundation Model Commonsense + Human-Specified Locomotion Prior \text{Foundation Model Commonsense} + \text{Human-Specified Locomotion Prior} Foundation Model Commonsense+Human-Specified Locomotion Prior
37. 对四足 VLA 研究的启发
如果在此基础上做真正的四足 VLA,可以将 LocoVLM 的中间层继续升级。
37.1 从 Retrieval Descriptor 到 Learned Skill Token
当前:
z = ( T , ψ , v x l i m i t ) z =(T,\psi,v_x^{\mathrm{limit}}) z=(T,ψ,vxlimit)
未来可以学习:
z m o t i o n ∈ R d z_{\mathrm{motion}} \in \mathbb{R}^{d} zmotion∈Rd
由大量四足 Motion / RL Trajectory 自动提取。
于是:
VLM → z m o t i o n → Low-Level Policy \text{VLM} \rightarrow z_{\mathrm{motion}} \rightarrow \text{Low-Level Policy} VLM→zmotion→Low-Level Policy
减少手工 Gait Dictionary。
37.2 加入 Footstep-Level Planning
进一步输出:
M = ( T , ψ , v , p f o o t , h s w i n g ) \mathbf{M} =( T, \psi, v, p_{\mathrm{foot}}, h_{\mathrm{swing}} ) M=(T,ψ,v,pfoot,hswing)
这样不仅能说:
"雪地慢点走。"
还能确定:
"下一步脚具体踩哪里。"
37.3 Vision + Language Joint Conditioning
当前:
Image OR Language \text{Image OR Language} Image OR Language
可以升级为:
Image AND Language \text{Image AND Language} Image AND Language
例如:
图像看到楼梯 + 用户说"快速上楼"。
统一 VLM 输出 Gait / Foothold / Velocity Constraint。
37.4 加入 Safety / Feasibility Critic
VLM 输出:
M V L M \mathbf{M}_{\mathrm{VLM}} MVLM
以后先经过:
Dynamics Feasibility + Safety Critic \text{Dynamics Feasibility} + \text{Safety Critic} Dynamics Feasibility+Safety Critic
再送低层 Policy:
M s a f e \mathbf{M}_{\mathrm{safe}} Msafe
可以进一步降低 Hallucinated Gait Parameters 的风险。
38. 和端到端 VLA 的两种路线
路线 A:End-to-End
( I , L , P ) → Transformer / Flow / Diffusion → a t : t + H (I,L,P) \rightarrow \text{Transformer / Flow / Diffusion} \rightarrow a_{t:t+H} (I,L,P)→Transformer / Flow / Diffusion→at:t+H
优点:
- 统一;
- 可以自动学习中间表示。
问题:
- 四足高频控制难;
- 动力学稳定性要求高;
- 训练数据需求巨大。
路线 B:LocoVLM 式 Hierarchical VLA
( I , L ) → z s k i l l → π l o c o m o t i o n → a t (I,L) \rightarrow z_{\mathrm{skill}} \rightarrow \pi_{\mathrm{locomotion}} \rightarrow a_t (I,L)→zskill→πlocomotion→at
优点:
- 可解释;
- 低层已有成熟 RL Policy;
- 语义层低频即可;
- 真机安全边界更清晰。
对于当前四足机器人 Foundation Model 研究,第二条路线往往更容易先完成可工作的 Sim-to-Real 系统。
39. 关键参数汇总
表 6 LocoVLM 关键系统参数。
| 项目 | 设置 |
|---|---|
| 四足平台 | Unitree Go1 |
| 人形验证平台 | Unitree H1 |
| LLM | GPT-4o |
| VLM | BLIP-2 |
| 四足 RL | PPO |
| Actor-Critic | Asymmetric Actor-Critic |
| 四足仿真 | Isaac Gym |
| H1 泛化仿真 | MuJoCo |
| 四足 Policy 频率 | 50 Hz |
| Motor Controller | 200 Hz |
| 四足板载计算 | Jetson Xavier NX |
| VLM GPU | RTX 3070 Ti Laptop |
| VLM 推理 + 通信 | < 100 m s <100\ \mathrm{ms} <100 ms |
| Skill Descriptor | M = ( T , ψ , v x l i m i t ) \mathbf{M}=(T,\psi,v_x^{\mathrm{limit}}) M=(T,ψ,vxlimit) |
| 基础 Gait 数量 | 5 |
| Compliance Threshold | δ = 0.5 \delta=0.5 δ=0.5 |
| Contact Smoothing | σ = 0.25 \sigma=0.25 σ=0.25 |
| Gait Test Velocity | 1.2 m / s 1.2\ \mathrm{m/s} 1.2 m/s |
| Gait Test Period | 0.4 s 0.4\ \mathrm{s} 0.4 s |
| Gait Test Rollouts | 1000 |
| Retrieval Benchmark | 100 Instructions |
| 最佳 Retrieval | 87% |
| Navigation VLM Period | 5 s 5\ \mathrm{s} 5 s |
| GPT-4o 是否在线控制 | 否 |
40. 论文价值总结
LocoVLM 的研究价值可以浓缩成三个层次。
第一层是语言/视觉语义到运动参数的 Grounding:
Semantic Understanding → Locomotion Parameters \text{Semantic Understanding} \rightarrow \text{Locomotion Parameters} Semantic Understanding→Locomotion Parameters
第二层是Foundation Model 与实时控制解耦:
Offline LLM + Online VLM + High-Frequency RL \text{Offline LLM} + \text{Online VLM} + \text{High-Frequency RL} Offline LLM+Online VLM+High-Frequency RL
第三层是把 Gait Style 设计为柔顺目标而不是绝对约束:
Style Fidelity + Disturbance Recovery \text{Style Fidelity} + \text{Disturbance Recovery} Style Fidelity+Disturbance Recovery
它不是一个典型的端到端 VLA,但对于四足机器人而言,这种层级式结构非常有代表性:
VLM/LLM 负责"应该怎么动" + RL Policy 负责"如何稳定地动" \boxed{ \text{VLM/LLM 负责"应该怎么动"} + \text{RL Policy 负责"如何稳定地动"} } VLM/LLM 负责"应该怎么动"+RL Policy 负责"如何稳定地动"
如果研究目标是将 Vision-Language Foundation Model 真正部署到 Go1、Go2、ANYmal 等动态足式平台,LocoVLM 展示了一条很清晰的工程路线:不必让大模型直接承担 50--200 Hz 的动力学控制,而可以利用一个低维、可解释、动力学相关的中间技能空间连接两端。
41. 参考资料
-
LocoVLM Project Page
-
LocoVLM: Grounding Vision and Language for Adapting Versatile Legged Locomotion Policies
-
LocoVLM Full Paper PDF
-
ICRA 2025 SafeVLMs Workshop Paper
-
LocoVLM Demo Video
-
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
-
SayTap: Language to Quadrupedal Locomotion