ICRA 2025|LocoVLM:视觉与语言驱动的通用足式运动策略自适应【文献解读】

文献:LocoVLM: Grounding Vision and Language for Adapting Versatile Legged Locomotion Policies

作者:I Made Aswin Nahrendra, Seunghyun Lee, Dongkyu Lee, Hyun Myung

单位:KAIST(Korea Advanced Institute of Science and Technology)

会议版本:ICRA 2025 Workshop on Safe Vision-Language Models(SafeVLMs)

项目主页:https://locovlm.github.io/

四足平台:Unitree Go1 ;人形泛化验证:Unitree H1(MuJoCo)

高层 Foundation Model:GPT-4o + BLIP-2

低层运动控制:PPO + Asymmetric Actor-Critic + Style-Conditioned Locomotion Policy

仿真平台:Isaac Gym(Go1)/ MuJoCo(H1 泛化实验)

核心关键词:Vision-Language Grounding、Legged Locomotion、Skill Retrieval、Prompted Reasoning、Compliant Contact Tracking、Sim-to-Real


1. 核心结论

LocoVLM 解决的不是"让一个大模型直接输出四足机器人的 12 个关节动作",而是更实际的层级式问题:

如何把视觉或自然语言中的高层语义,实时、稳定地转换为足式机器人能够执行的运动风格参数,并且不把云端大语言模型放进高频控制环。

它的完整思路可以概括为:

Vision / Language Query → BLIP-2 Retrieval → M = ( T , ψ , v x l i m i t ) → Style-Conditioned RL Policy → q d e s → Motor Controller \text{Vision / Language Query} \rightarrow \text{BLIP-2 Retrieval} \rightarrow \mathbf{M}=(T,\psi,v_x^{\mathrm{limit}}) \rightarrow \text{Style-Conditioned RL Policy} \rightarrow q^{\mathrm{des}} \rightarrow \text{Motor Controller} Vision / Language Query→BLIP-2 Retrieval→M=(T,ψ,vxlimit)→Style-Conditioned RL Policy→qdes→Motor Controller

其中:

  • T T T:步态周期;
  • ψ \psi ψ:四条腿的相位偏移;
  • v x l i m i t v_x^{\mathrm{limit}} vxlimit:前向速度上限。

GPT-4o 并不在真机在线控制环中持续调用,而是在离线阶段生成大量"语言指令---推理---运动描述符"数据,形成 Skill Database(技能数据库)。部署时由轻量得多的预训练 VLM------BLIP-2------完成视觉/语言查询与技能数据库之间的实时匹配,再把运动参数发送给 PPO 训练的运动控制器。

因此这篇论文真正的系统结构是:

LLM 离线知识生成 + VLM 在线语义检索 + RL 高频运动控制 \boxed{ \text{LLM 离线知识生成} + \text{VLM 在线语义检索} + \text{RL 高频运动控制} } LLM 离线知识生成+VLM 在线语义检索+RL 高频运动控制

这也是它和 SayTap、传统速度控制策略以及端到端 VLA 的根本区别。


2. 研究背景:论文到底解决了什么问题

2.1 传统足式运动控制主要理解"几何",但不理解"语义"

传统四足运动策略通常依据:

  • 高程图;
  • 深度图;
  • 点云;
  • 地形法向量;
  • 本体状态;
  • 目标速度;

来判断"哪里能踩""哪里有障碍物"。

这类表示对于几何避障非常有效,但无法直接理解:

  • "这里是图书馆,安静一点";
  • "地面结冰了,谨慎走";
  • "像兔子一样跳";
  • "前方环境拥挤,慢一点";
  • "像鹿一样轻快地走"。

换句话说,几何感知可以表示:

Obstacle Geometry \text{Obstacle Geometry} Obstacle Geometry

却不天然包含:

Affordance + Social Context + Task Semantics + Commonsense \text{Affordance} + \text{Social Context} + \text{Task Semantics} + \text{Commonsense} Affordance+Social Context+Task Semantics+Commonsense

LocoVLM 的第一个目标,就是把这类高层语义信息真正作用到低层步态和速度上。

2.2 大语言模型很强,但不适合直接进入 50--200 Hz 控制环

如果直接让 GPT-4o 在线生成机器人动作,会遇到:

  • 网络依赖;
  • 云端 API 延迟;
  • 推理时间不可确定;
  • 成本随控制时间持续增加;
  • 断网时无法工作;
  • LLM 输出不适合直接驱动连续动力学系统。

因此论文没有设计:

LLM → Joint Torque \text{LLM} \rightarrow \text{Joint Torque} LLM→Joint Torque

而是使用 Foundation Model Knowledge Distillation(基础模型知识蒸馏):

GPT-4o Knowledge → Offline Skill Database → BLIP-2 Retrieval \text{GPT-4o Knowledge} \rightarrow \text{Offline Skill Database} \rightarrow \text{BLIP-2 Retrieval} GPT-4o Knowledge→Offline Skill Database→BLIP-2 Retrieval

GPT-4o 只承担离线的"老师"角色。

2.3 步态风格跟踪与鲁棒性之间存在冲突

如果强化学习控制器被强制严格遵循预定义足端接触时序,那么在:

  • 台阶;
  • 离散落脚点;
  • 粗糙地面;
  • 突发扰动;

中,机器人可能为了"守住规定步态"而牺牲稳定性。

论文因此提出 Compliant Contact Tracking(柔顺接触跟踪):

当实际接触与目标接触只存在一定范围内的误差时,不强制处罚;只有偏差超过阈值后才恢复严格约束。

这让控制器可以在遇到扰动时暂时违反理想步态,优先维持稳定。

2.4 技能数据库扩大后,如何兼顾检索速度与语义准确率

BLIP-2 有两类可用于检索的信息:

  1. Embedding Cosine Similarity:速度快,但语义区分能力有限;
  2. Image-Text Matching(ITM)Head:语义匹配更精确,但对整个数据库逐项计算代价太高。

论文提出 Mixed-Precision Retrieval(混合精度检索):

Fast Coarse Retrieval → Accurate Fine Re-ranking \text{Fast Coarse Retrieval} \rightarrow \text{Accurate Fine Re-ranking} Fast Coarse Retrieval→Accurate Fine Re-ranking

这里的"Mixed-Precision"不是 FP16/FP32 混合浮点精度训练,而是指:

使用低成本的粗粒度相似度先筛选,再使用高精度 ITM 对少量候选重新排序。


3. 整体架构

图 1 LocoVLM 总体框架。来源:Nahrendra et al., LocoVLM, Figure 1,官方项目主页。LLM 离线生成技能数据库;部署阶段 VLM 通过混合精度检索得到运动描述符;低层 Style-Conditioned Locomotion Policy 根据运动描述符、速度命令和本体状态执行动作。

论文把系统理解为一个高低层分离的 Teacher--Student(教师---学生)结构。

3.1 高层 Teacher:GPT-4o

GPT-4o 用于离线构造数据库:

D = { ( I 1 , M 1 ) , ( I 2 , M 2 ) , ...   } \mathcal{D} =\left\{ (\mathcal{I}_1,\mathbf{M}_1), (\mathcal{I}_2,\mathbf{M}_2), \dots \right\} D={(I1,M1),(I2,M2),...}

其中:

  • I \mathcal{I} I:语言指令;
  • M \mathbf{M} M:可执行运动描述符。

GPT-4o 不负责真机实时推理。

3.2 高层 Student:BLIP-2

部署时,BLIP-2 输入可以是:

I q u e r y = { text , RGB image . \mathcal{I}_{\mathrm{query}} =\begin{cases} \text{text}, \\ \text{RGB image}. \end{cases} Iquery={text,RGB image.

VLM 根据语义从数据库中取回最匹配的运动描述符。

3.3 低层控制器:Style-Conditioned RL Policy

低层控制器输入:

Proprioception + Velocity Command + Gait Style Parameters \text{Proprioception} + \text{Velocity Command} + \text{Gait Style Parameters} Proprioception+Velocity Command+Gait Style Parameters

输出:

q d e s q^{\mathrm{des}} qdes

即关节目标角,由机器人电机控制器进一步转成力矩。


4. Style-Conditioned Locomotion Policy

4.1 为什么不直接使用离散 Gait Label

如果只给策略输入:

text 复制代码
TROT
PACE
BOUND

它只能在有限离散动作之间切换。

LocoVLM 将步态参数化为:

( T , ψ ) (T,\psi) (T,ψ)

其中:

T ∈ R T\in\mathbb{R} T∈R

表示 Gait Cycle Duration(步态周期),而

ψ = ψ F L , ψ F R , ψ R L , ψ R R ∈ R 4 \psi = \\psi_{\\mathrm{FL}}, \\psi_{\\mathrm{FR}}, \\psi_{\\mathrm{RL}}, \\psi_{\\mathrm{RR}} \in\mathbb{R}^{4} ψ=ψFL,ψFR,ψRL,ψRR∈R4

表示四条腿的相位偏移。

同一种 gait 还可以通过改变 T T T 连续调节步频:

f = 1 T f=\frac{1}{T} f=T1

所以:

T ↓ ⇒ f ↑ T\downarrow \quad\Rightarrow\quad f\uparrow T↓⇒f↑

即周期越短,步频越高。

4.2 Gait Phase Encoding

论文使用二维 Clock Input:

ϕ ( t ) = sin ⁡ ( 2 π t T ) , cos ⁡ ( 2 π t T ) \phi(t) =\left \\sin\\left(2\\pi\\frac{t}{T}\\right), \\cos\\left(2\\pi\\frac{t}{T}\\right) \\right ϕ(t)=sin(2πTt),cos(2πTt)

再结合每条腿的:

ψ i \psi_i ψi

获得不同腿的接触相位关系。

图 2 步态相位编码与 Compliance Zone。来源:Nahrendra et al., LocoVLM, Figure 2。图中以 T = 1 T=1 T=1 为例,橙色正弦/余弦曲线表示周期相位,绿色区域表示允许策略偏离理想接触时序的柔顺区域。

这类时钟信号的作用是让策略明确知道:

当前处于一个 gait cycle 的什么位置。

而不是让策略完全依赖历史状态自行推断步态相位。


5. 五种基础步态及相位偏移

表 1 原论文 Table III:五种步态的 Foot Phase Offsets。

Gait FL FR RL RR
Pronk 0.0 0.0 0.0 0.0
Trot 0.0 0.5 0.5 0.0
Pace 0.0 0.5 0.0 0.5
Bound 0.0 0.0 0.5 0.5
Rotary gallop 0.0 0.2 0.7 0.5

其中:

  • FL:Front Left;
  • FR:Front Right;
  • RL:Rear Left;
  • RR:Rear Right。

5.1 Pronk

ψ = 0 , 0 , 0 , 0 \psi=0,0,0,0 ψ=0,0,0,0

四腿同相:

F L = F R = R L = R R \mathrm{FL} =\mathrm{FR} =\mathrm{RL} =\mathrm{RR} FL=FR=RL=RR

表现为四脚同时起跳和落地。

5.2 Trot

ψ = 0 , 0.5 , 0.5 , 0 \psi=0,0.5,0.5,0 ψ=0,0.5,0.5,0

因此:

F L ∼ R R \mathrm{FL}\sim\mathrm{RR} FL∼RR

F R ∼ R L \mathrm{FR}\sim\mathrm{RL} FR∼RL

是典型对角小跑。

5.3 Pace

ψ = 0 , 0.5 , 0 , 0.5 \psi=0,0.5,0,0.5 ψ=0,0.5,0,0.5

即:

F L ∼ R L \mathrm{FL}\sim\mathrm{RL} FL∼RL

F R ∼ R R \mathrm{FR}\sim\mathrm{RR} FR∼RR

同侧腿同步。

5.4 Bound

ψ = 0 , 0 , 0.5 , 0.5 \psi=0,0,0.5,0.5 ψ=0,0,0.5,0.5

前腿同步,后腿同步。

5.5 Rotary Gallop

ψ = 0 , 0.2 , 0.7 , 0.5 \psi=0,0.2,0.7,0.5 ψ=0,0.2,0.7,0.5

四腿依次错相,形成更复杂的高速旋转式 Gallop 接触序列。


6. Compliant Contact Tracking:论文最重要的低层控制创新

传统 Gait-Conditioned Policy 往往加入严格接触奖励:

c i ( t ) ≈ c ^ i ( t ) c_i(t)\approx\hat{c}_i(t) ci(t)≈c^i(t)

其中:

  • c i ( t ) c_i(t) ci(t):真实足端接触状态;
  • c ^ i ( t ) \hat{c}_i(t) c^i(t):目标接触状态。

严格跟踪会产生一个问题:

当前环境要求临时提前落脚,但奖励函数却强迫机器人继续保持 Swing Phase。

这样可能导致机器人跌倒。

论文将误差重新定义为:

ϕ e r r o r c o m p l y = { 0 , ϕ e r r o r ≤ δ , ϕ e r r o r , otherwise . \phi_{\mathrm{error}}^{\mathrm{comply}} =\begin{cases} 0, & \phi_{\mathrm{error}}\le\delta, \\ \phi_{\mathrm{error}}, & \text{otherwise}. \end{cases} ϕerrorcomply={0,ϕerror,ϕerror≤δ,otherwise.

其中:

δ \delta δ

为 Compliance Threshold(柔顺阈值)。

实验统一采用:

δ = 0.5 \delta=0.5 δ=0.5

直观理解是:

在步态周期约 50% 范围内允许实际接触偏离理想时刻,不立即施加接触跟踪惩罚。

这不是放弃步态跟踪,而是把目标从:

Strict Contact Tracking \boxed{\text{Strict Contact Tracking}} Strict Contact Tracking

改为:

Track When Possible, Deviate When Necessary \boxed{\text{Track When Possible, Deviate When Necessary}} Track When Possible, Deviate When Necessary

即"能按计划走时按计划走,需要保命时允许违反计划"。

6.1 Contact Reward

附录进一步给出:

r c o n t a c t = exp ⁡ ( − ϕ e r r o r σ ) r_{\mathrm{contact}} =\exp\left( -\frac{\phi_{\mathrm{error}}}{\sigma} \right) rcontact=exp(−σϕerror)

其中:

σ = 0.25 \sigma=0.25 σ=0.25

为 Smoothing Factor。

指数核使接触误差从"硬开关"变成连续平滑的奖励变化,有利于 PPO 训练。


7. 为什么 Compliant Tracking 能提高鲁棒性

机器人在粗糙地形上真正需要满足的是:

Dynamic Feasibility \text{Dynamic Feasibility} Dynamic Feasibility

而不只是:

Reference Gait Fidelity \text{Reference Gait Fidelity} Reference Gait Fidelity

如果严格接触调度:

c ^ ( t ) \hat{c}(t) c^(t)

与当前地形动力学冲突,那么最优行为实际上可能是:

c ( t ) ≠ c ^ ( t ) c(t)\neq\hat{c}(t) c(t)=c^(t)

例如:

  • 提前触地;
  • 延迟离地;
  • 某一步缩短 Swing;
  • 临时延长 Stance。

因此柔顺阈值相当于给 Gait Schedule 加入:

Slack \text{Slack} Slack

其作用和优化控制中软约束的思想类似:

风格是软约束,稳定性是更高优先级目标。


8. 五种步态跟踪结果

图 5 五种步态的接触状态与机器人姿态。来源:Nahrendra et al., LocoVLM, Figure 5。上排是四足接触状态,下排是真实/仿真机器人姿态,分别对应 (a) pronk、(b) trot、© pace、(d) bound、(e) rotary gallop。

论文通过这个实验验证:

( T , ψ ) (T,\psi) (T,ψ)

确实足以让同一个 Locomotion Policy 表达多种基础 gait,而不是为每个步态单独训练一个 policy。


9. Gait Tracking 定量实验

论文设置:

v x c m d = 1.2 m / s v_x^{\mathrm{cmd}} =1.2\ \mathrm{m/s} vxcmd=1.2 m/s

T = 0.4 s T=0.4\ \mathrm{s} T=0.4 s

最大 Episode:

20 s 20\ \mathrm{s} 20 s

每种配置:

1000 rollouts 1000\ \text{rollouts} 1000 rollouts

评价指标是机器人在 episode 内平均能够走多远。

表 2 原论文 Table I:不同 Compliance Threshold 下的平均行走距离。

Gait δ \delta δ Rough / m Discrete / m Stairs / m
Pronk 0 15.13 ± 3.61 15.13\pm3.61 15.13±3.61 17.34 ± 6.41 17.34\pm6.41 17.34±6.41 13.56 ± 2.53 13.56\pm2.53 13.56±2.53
Pronk 0.25 15.82 ± 3.43 15.82\pm3.43 15.82±3.43 16.13 ± 6.65 16.13\pm6.65 16.13±6.65 14.78 ± 2.69 14.78\pm2.69 14.78±2.69
Pronk 0.5 15.59 ± 3.38 15.59\pm3.38 15.59±3.38 18.42 ± 6.29 18.42\pm6.29 18.42±6.29 14.06 ± 2.74 14.06\pm2.74 14.06±2.74
Pronk 0.75 16.98 ± 4.05 16.98\pm4.05 16.98±4.05 16.79 ± 6.64 16.79\pm6.64 16.79±6.64 14.07 ± 5.13 14.07\pm5.13 14.07±5.13
Trot 0 16.72 ± 3.53 16.72\pm3.53 16.72±3.53 16.62 ± 6.50 16.62\pm6.50 16.62±6.50 14.92 ± 2.82 14.92\pm2.82 14.92±2.82
Trot 0.25 14.98 ± 4.03 14.98\pm4.03 14.98±4.03 18.96 ± 6.14 18.96\pm6.14 18.96±6.14 15.94 ± 2.86 15.94\pm2.86 15.94±2.86
Trot 0.5 16.84 ± 3.72 16.84\pm3.72 16.84±3.72 18.29 ± 6.18 18.29\pm6.18 18.29±6.18 16.34 ± 2.41 16.34\pm2.41 16.34±2.41
Trot 0.75 15.32 ± 3.04 15.32\pm3.04 15.32±3.04 15.59 ± 5.89 15.59\pm5.89 15.59±5.89 15.70 ± 2.32 15.70\pm2.32 15.70±2.32
Pace 0 17.11 ± 3.51 17.11\pm3.51 17.11±3.51 17.48 ± 5.89 17.48\pm5.89 17.48±5.89 16.83 ± 2.35 16.83\pm2.35 16.83±2.35
Pace 0.25 18.48 ± 4.45 18.48\pm4.45 18.48±4.45 18.43 ± 6.58 18.43\pm6.58 18.43±6.58 16.71 ± 3.40 16.71\pm3.40 16.71±3.40
Pace 0.5 16.31 ± 3.95 16.31\pm3.95 16.31±3.95 19.49 ± 5.77 19.49\pm5.77 19.49±5.77 18.18 ± 2.49 18.18\pm2.49 18.18±2.49
Pace 0.75 17.23 ± 4.19 17.23\pm4.19 17.23±4.19 18.05 ± 5.67 18.05\pm5.67 18.05±5.67 16.07 ± 4.15 16.07\pm4.15 16.07±4.15
Bound 0 14.57 ± 2.91 14.57\pm2.91 14.57±2.91 14.53 ± 5.64 14.53\pm5.64 14.53±5.64 14.17 ± 2.29 14.17\pm2.29 14.17±2.29
Bound 0.25 16.99 ± 3.98 16.99\pm3.98 16.99±3.98 16.74 ± 5.44 16.74\pm5.44 16.74±5.44 15.53 ± 2.55 15.53\pm2.55 15.53±2.55
Bound 0.5 16.04 ± 3.57 16.04\pm3.57 16.04±3.57 17.43 ± 5.16 17.43\pm5.16 17.43±5.16 15.48 ± 2.92 15.48\pm2.92 15.48±2.92
Bound 0.75 16.66 ± 4.40 16.66\pm4.40 16.66±4.40 17.47 ± 6.33 17.47\pm6.33 17.47±6.33 14.39 ± 4.13 14.39\pm4.13 14.39±4.13
Rotary gallop 0 16.67 ± 3.56 16.67\pm3.56 16.67±3.56 16.71 ± 6.20 16.71\pm6.20 16.71±6.20 16.40 ± 2.88 16.40\pm2.88 16.40±2.88
Rotary gallop 0.25 17.89 ± 4.15 17.89\pm4.15 17.89±4.15 17.77 ± 5.69 17.77\pm5.69 17.77±5.69 17.17 ± 3.02 17.17\pm3.02 17.17±3.02
Rotary gallop 0.5 18.18 ± 3.91 18.18\pm3.91 18.18±3.91 18.61 ± 5.08 18.61\pm5.08 18.61±5.08 18.04 ± 2.10 18.04\pm2.10 18.04±2.10
Rotary gallop 0.75 16.97 ± 4.68 16.97\pm4.68 16.97±4.68 18.55 ± 5.88 18.55\pm5.88 18.55±5.88 16.38 ± 3.94 16.38\pm3.94 16.38±3.94

加粗值是同一 gait、同一 terrain 下论文表格中的最佳结果。

9.1 实验说明了什么

并不是 δ \delta δ 越大越好,而是存在:

Style Fidelity ↔ Robustness \text{Style Fidelity} \leftrightarrow \text{Robustness} Style Fidelity↔Robustness

之间的折中。

过小:

δ → 0 \delta\rightarrow0 δ→0

接近严格接触跟踪,鲁棒性降低。

过大:

δ → 1 \delta\rightarrow1 δ→1

又可能使 gait constraint 过弱。

作者最终在全部实验中统一选择:

δ = 0.5 \boxed{\delta=0.5} δ=0.5

作为整体折中,而不是针对每一种 gait 和地形单独调最优参数。


10. 离线 Skill Database:如何把 GPT-4o 的知识"蒸馏"下来

图 3 Offline Skill Database Generation Pipeline。来源:Nahrendra et al., LocoVLM, Figure 3。第一阶段先生成语言指令,第二阶段通过 Meta-Prompt 将指令转换成结构化 Skill Database 条目。

作者没有人工写几千条:

text 复制代码
language → gait parameters

而是利用 GPT-4o 自动扩展。

生成分为两步。

10.1 Stage 1:Instruction Description Generation

GPT-4o 先生成三类 instruction。

Mimicking Behaviors

例如:

text 复制代码
let's hop like a rabbit
run beautifully like a horse
you are a kangaroo
Scene Responses

例如:

text 复制代码
there is an icy patch
this is a library
the snow is slippery
Direct Instructions

例如:

text 复制代码
trot slowly
bound quickly
reduce your noise

每次要求模型产生:

n = 100 n=100 n=100

条指令。

采用分类生成而不是把全部类型一次性混在一起,可以降低重复和无结构输出。

10.2 Stage 2:Motion Descriptor Generation

每条语言指令被转换成:

M = ( T , ψ , v x l i m i t ) \mathbf{M} =(T,\psi,v_x^{\mathrm{limit}}) M=(T,ψ,vxlimit)

完整数据库条目可以表示为:

d = ( I , M ) d =(\mathcal{I},\mathbf{M}) d=(I,M)

实际 JSON 还包含:

text 复制代码
instruction
reasoning
T
gait_phase_offsets
vel_lim

11. 为什么要加入 Prompted Reasoning

如果输入:

text 复制代码
trot quickly

LLM 很容易映射到:

  • Trot Phase Offsets;
  • 较小 T T T;
  • 较大 v x l i m i t v_x^{\mathrm{limit}} vxlimit。

但输入:

text 复制代码
trundle along like a hippo

LLM 必须先理解:

河马 → 沉重、缓慢 → 步频低、速度低。

所以作者要求 GPT-4o 在给出数值前先生成中间 reasoning。

例如:

"trundle along like a hippo" \text{"trundle along like a hippo"} "trundle along like a hippo"

先推理成:

slow and heavy trot \text{slow and heavy trot} slow and heavy trot

再得到:

T ↑ , v x l i m i t ↓ T\uparrow, \qquad v_x^{\mathrm{limit}}\downarrow T↑,vxlimit↓

这相当于:

High-Level Semantics → Technical Motion Semantics → Numerical Descriptor \text{High-Level Semantics} \rightarrow \text{Technical Motion Semantics} \rightarrow \text{Numerical Descriptor} High-Level Semantics→Technical Motion Semantics→Numerical Descriptor

比直接:

Language → Numbers \text{Language} \rightarrow \text{Numbers} Language→Numbers

更加稳定。


12. Prompted Reasoning 对数据库质量的影响

图 6 三种技能数据库生成方式的统计分布。来源:Nahrendra et al., LocoVLM, Figure 6。左:类似 SayTap 的逐条生成 Baseline;中:LocoVLM Batch Generation、无 Prompted Reasoning;右:LocoVLM + Prompted Reasoning。

实验统一生成:

300 300 300

条 Motion Descriptors。

12.1 Gait 分布

Baseline 中:

45.7 % 45.7\% 45.7%

被归类为 Others,即不属于五种稳定标准 gait 的非结构化 phase offsets。

LocoVLM 无 reasoning:

25.7 % 25.7\% 25.7%

LocoVLM + Prompted Reasoning:

5.3 % 5.3\% 5.3%

说明显式推理显著减少了无结构步态参数。

表 3 Figure 6(a) 的 Gait Category Distribution。

方法 Trot Bound Pace Pronk Rotary gallop Others
Baseline 32.3% 4.0% 5.0% 4.0% 9.0% 45.7%
LocoVLM w/o reasoning 44.0% 9.3% 12.7% 2.3% 6.0% 25.7%
LocoVLM + reasoning 53.7% 10.7% 10.0% 8.0% 12.3% 5.3%

12.2 Gait Cycle 分布

Prompted Reasoning 后, T T T 主要分布在:

0.2 ∼ 0.7 s 0.2\sim0.7\ \mathrm{s} 0.2∼0.7 s

而 Baseline 与无 reasoning 版本更集中在约:

0.5 s 0.5\ \mathrm{s} 0.5 s

并出现:

T ≈ 1.0 s T\approx1.0\ \mathrm{s} T≈1.0 s

的异常/不稳定长周期。

12.3 Velocity Limit

在:

v x l i m i t v_x^{\mathrm{limit}} vxlimit

分布上三种方法差异没有 gait phase 和 cycle 那么明显。

论文认为原因是速度语义相对容易:

text 复制代码
quickly → high velocity
slowly → low velocity

但"像某种动物""在某种社会环境下如何走"到 gait phase 的映射更复杂,因此 reasoning 的帮助更明显。


13. 生成成本

生成 300 个 Motion Descriptors:

方法 GPT-4o 生成成本
Baseline $$1.16$
LocoVLM without prompted reasoning $$0.21$
LocoVLM with prompted reasoning $$0.25$

批量生成的成本只有逐条查询的一小部分。

同时 Batch Generation 相当于给模型提供了一个局部"上下文记忆",可以减少同一批次中重复描述符的产生。


14. Motion Descriptor 为什么只选三个参数

论文定义:

M = ( T , ψ , v x l i m i t ) \boxed{ \mathbf{M} =(T,\psi,v_x^{\mathrm{limit}}) } M=(T,ψ,vxlimit)

这是一个很重要的工程取舍。

如果让 VLM 输出:

  • 12 个关节角;
  • 足端轨迹;
  • GRF;
  • Body Pose;
  • Torque;

高层语义与动作空间之间会变得非常难对齐。

而:

( T , ψ , v x l i m i t ) (T,\psi,v_x^{\mathrm{limit}}) (T,ψ,vxlimit)

恰好具有三个特点:

  1. 维度低;
  2. 人类可解释;
  3. 足够控制多种 gait 和速度。

因此它相当于机器人运动系统的一个 Semantic Control Interface(语义控制接口)。


15. Vision-Language Grounding:Skill Retrieval

图 4 Skill Database Retrieval。来源:Nahrendra et al., LocoVLM, Figure 4。文本或图像 Query 被编码到 BLIP-2 的共享语义空间,与技能数据库中的 instruction 表示进行匹配,最近的 instruction 对应一个 reasoning 和 motion descriptor。

数据库:

D = { d i } \mathcal{D} =\{d_i\} D={di}

其中:

d i = ( I i , M i ) d_i=(\mathcal{I}_i,\mathbf{M}_i) di=(Ii,Mi)

给定 Query:

I q u e r y \mathcal{I}_{\mathrm{query}} Iquery

目标是寻找:

I ∗ = arg ⁡ max ⁡ I ∈ D sim ⁡ ( I q u e r y , I ) \mathcal{I}^{*} =\arg\max_{\mathcal{I}\in\mathcal{D}} \operatorname{sim} \left( \mathcal{I}_{\mathrm{query}}, \mathcal{I} \right) I∗=argI∈Dmaxsim(Iquery,I)

然后:

I ∗ → M ∗ \mathcal{I}^{*} \rightarrow \mathbf{M}^{*} I∗→M∗

将:

M ∗ \mathbf{M}^{*} M∗

发送到低层运动策略。


16. Mixed-Precision Retrieval 算法

16.1 Stage 1:Cosine Similarity 进行快速 Top-K 检索

设:

f B L I P ( ⋅ ) f_{\mathrm{BLIP}}(\cdot) fBLIP(⋅)

是 BLIP-2 Encoder。

先求:

I K = TopK ⁡ I ∈ D cossim ⁡ ( f B L I P ( I q u e r y ) , f B L I P ( I ) ) \mathbf{I}^{K} =\operatorname{TopK}{\mathcal{I}\in\mathcal{D}} \operatorname{cossim} \left( f{\mathrm{BLIP}}(\mathcal{I}{\mathrm{query}}), f{\mathrm{BLIP}}(\mathcal{I}) \right) IK=TopKI∈Dcossim(fBLIP(Iquery),fBLIP(I))

只保留最相似的 K K K 个候选。

然后将这些候选的 Cosine Similarity 转成:

p 1 ( I K ) = softmax ⁡ ( cossim ⁡ ( f B L I P ( I q u e r y ) , f B L I P ( I K ) ) ) p_1(\mathbf{I}^{K}) =\operatorname{softmax} \left( \operatorname{cossim} \left( f_{\mathrm{BLIP}}(\mathcal{I}{\mathrm{query}}), f{\mathrm{BLIP}}(\mathbf{I}^{K}) \right) \right) p1(IK)=softmax(cossim(fBLIP(Iquery),fBLIP(IK)))

16.2 Stage 2:ITM Head 重排序

对每个:

I k ∈ I K \mathcal{I}_k\in\mathbf{I}^{K} Ik∈IK

利用 BLIP-2 Image-Text Matching Head:

p 2 ( I k ) = softmax ⁡ ( f I T M ( I q u e r y , I k ) ) p_2(\mathcal{I}k) =\operatorname{softmax} \left( f{\mathrm{ITM}} ( \mathcal{I}_{\mathrm{query}}, \mathcal{I}_k ) \right) p2(Ik)=softmax(fITM(Iquery,Ik))

最后:

I ∗ = arg ⁡ max ⁡ I k p 1 ( I k ) + p 2 ( I k ) \mathcal{I}^{*} =\arg\max_{\mathcal{I}_k} \left p_1(\\mathcal{I}_k) + p_2(\\mathcal{I}_k) \\right I∗=argIkmaxp1(Ik)+p2(Ik)

得到最终技能。

16.3 Algorithm 1

text 复制代码
Input:
    Query I_query
    Database D
    BLIP encoder f_BLIP
    ITM head f_ITM
    Top-K size K

1. 使用 BLIP embedding + cosine similarity 从 D 中取 Top-K
2. 对 Top-K cosine similarity 做 softmax,得到 p1
3. 对 Top-K 中每一条候选执行 BLIP-2 ITM
4. 得到 ITM matching probability p2
5. 将 p1 与 p2 组合
6. 选择综合概率最大的 instruction I*
7. 读取其 motion descriptor M*

时间复杂度直观上由:

O ( N d ) + O ( K C I T M ) O(Nd) + O(KC_{\mathrm{ITM}}) O(Nd)+O(KCITM)

组成。

其中:

K ≪ N K\ll N K≪N

所以避免了:

O ( N C I T M ) O(NC_{\mathrm{ITM}}) O(NCITM)

式的全数据库 ITM 暴力比较。


17. Text-as-Image:一个很反直觉但有效的技巧

BLIP-2 本质上主要学习:

Image ↔ Text \text{Image} \leftrightarrow \text{Text} Image↔Text

的对齐。

它不一定特别擅长:

Text ↔ Text \text{Text} \leftrightarrow \text{Text} Text↔Text

的句子级匹配。

作者于是把用户文本:

text 复制代码
shh! the baby is sleeping

绘制为:

白色背景 + 黑色文字

得到一张"文字图片",再送进 VLM Image Encoder。

于是原本的:

Text Query → Text Retrieval \text{Text Query} \rightarrow \text{Text Retrieval} Text Query→Text Retrieval

变成:

Rendered Text Image → Image-Text Matching \text{Rendered Text Image} \rightarrow \text{Image-Text Matching} Rendered Text Image→Image-Text Matching

相当于把问题重新投影到 BLIP-2 最擅长的训练分布。

这是论文中一个很有工程价值的小技巧。


18. Retrieval Accuracy

作者人工标注:

100 100 100

条 instruction 作为检索评估集。

表 4 原论文 Table II:不同检索方法的准确率。

Retrieval Metric Text as String Text as Image Average
Cosine similarity 21/100 30/100 20.5%(原论文数值)
Top-K similarity 27/100 48/100 37.5%
Top-K to ITM 51/100 57/100 54.0%
Mixed-Precision 72/100 87/100 79.5%

最终最佳组合:

Text-as-Image + Mixed-Precision Retrieval = 87 % \boxed{ \text{Text-as-Image} + \text{Mixed-Precision Retrieval} =87\% } Text-as-Image+Mixed-Precision Retrieval=87%

论文摘要所说的最高 87% Instruction-Following Accuracy,需要结合这个实验口径理解:它主要对应这里的 100 条人工标注 instruction 上的检索/匹配准确率,而不是"任意开放世界指令 87% 成功率"。


19. 超出数据库的语义泛化

LocoVLM 不要求用户输入与数据库字符串完全一致。

例如数据库可能存在:

text 复制代码
let's hop like a rabbit

用户输入:

text 复制代码
you are a kangaroo

BLIP-2 会因为:

kangaroo ∼ jump ∼ rabbit hop \text{kangaroo} \sim \text{jump} \sim \text{rabbit hop} kangaroo∼jump∼rabbit hop

在语义 embedding 中找到相近技能。

同样:

text 复制代码
this is a library

可以检索到:

text 复制代码
move quietly

这种能力并不是 LocoVLM 在线重新推理出新的控制器,而是:

Novel Query → Nearest Semantically Compatible Skill \text{Novel Query} \rightarrow \text{Nearest Semantically Compatible Skill} Novel Query→Nearest Semantically Compatible Skill

因此它属于 Retrieval-Based Semantic Grounding(基于检索的语义落地)。


20. Robot-Centric Vision:真正把视觉环境语义作用到 Gait

图 7 Robot-Centric RGB Scene Interpretation。来源:Nahrendra et al., LocoVLM, Figure 7。机器人从普通路面进入雪地后,VLM 根据视觉场景检索出不同的 Motion Descriptor。

20.1 普通路面

VLM 检索到类似:

text 复制代码
traipse lightly like a deer

输出:

T = 0.5 s T=0.5\ \mathrm{s} T=0.5 s

v x l i m i t = 0.6 m / s v_x^{\mathrm{limit}} =0.6\ \mathrm{m/s} vxlimit=0.6 m/s

并使用 Trot。

20.2 冰雪区域

VLM 输出:

text 复制代码
a field of ice, walk light-footed

或:

text 复制代码
skulk with stealth like a lynx

速度降低到:

0.2 ∼ 0.3 m / s 0.2\sim0.3\ \mathrm{m/s} 0.2∼0.3 m/s

周期增加到:

0.6 ∼ 0.7 s 0.6\sim0.7\ \mathrm{s} 0.6∼0.7 s

即:

Snow / Ice → Cautious Semantics → v x l i m i t ↓ , T ↑ \text{Snow / Ice} \rightarrow \text{Cautious Semantics} \rightarrow v_x^{\mathrm{limit}}\downarrow, \quad T\uparrow Snow / Ice→Cautious Semantics→vxlimit↓,T↑

这就是论文标题中 Grounding Vision 最直接的体现。

关键点不是"识别出这是雪",而是:

视觉场景语义最终改变了机器人能够执行的 gait style 和速度限制。


21. 真机系统配置

21.1 Unitree Go1 Locomotion Controller

训练算法:

PPO \text{PPO} PPO

结构:

Asymmetric Actor-Critic \text{Asymmetric Actor-Critic} Asymmetric Actor-Critic

并结合 State Estimation。

训练仿真器:

Isaac Gym \text{Isaac Gym} Isaac Gym

为了 Sim-to-Real,随机化:

  • Robot Mass;
  • Center of Mass;
  • Motor Stiffness;
  • Motor Damping;
  • Terrain Friction;
  • System Delay。

21.2 真机执行频率

RL Policy:

50 H z 50\ \mathrm{Hz} 50 Hz

输出:

q d e s q^{\mathrm{des}} qdes

Go1 Motor Controller:

200 H z 200\ \mathrm{Hz} 200 Hz

将目标关节角转换为电机力矩。

板载计算:

Jetson Xavier NX \text{Jetson Xavier NX} Jetson Xavier NX

21.3 VLM 模块

单独计算机:

NVIDIA RTX 3070 Ti \text{NVIDIA RTX 3070 Ti} NVIDIA RTX 3070 Ti

BLIP-2 通过 ROS 与机器人通信。

VLM 推理 + 数据通信:

< 100 m s <100\ \mathrm{ms} <100 ms

VLM Advisor 与 Locomotion Policy:

Asynchronous \boxed{\text{Asynchronous}} Asynchronous

因此低层 50 Hz Policy 不需要等待每一次 VLM 推理完成。

这一点对真机非常重要:

Semantic Loop Frequency ≪ Locomotion Control Frequency \text{Semantic Loop Frequency} \ll \text{Locomotion Control Frequency} Semantic Loop Frequency≪Locomotion Control Frequency


22. 为什么异步架构比在线 LLM 控制更合理

可以把系统拆成三个时间尺度。

高频闭环

50 ∼ 200 H z 50\sim200\ \mathrm{Hz} 50∼200 Hz

负责:

  • 平衡;
  • 关节动作;
  • 接触恢复。

中低频语义层

∼ 100 m s \sim100\ \mathrm{ms} ∼100 ms

负责:

  • Scene Interpretation;
  • Language Grounding;
  • Skill Retrieval。

极低频离线知识生成

GPT-4o 只在数据库生成阶段运行。

因此即使 Foundation Model 有延迟,也不会直接破坏:

Dynamic Stability Loop \text{Dynamic Stability Loop} Dynamic Stability Loop

这种设计对四足/人形机器人比让一个大型 VLM 直接生成 50 Hz 关节动作更工程化。


23. Zero-Shot Cross-Embodiment Generalization

图 8 LocoVLM 在 Unitree H1 人形机器人上的 Zero-Shot Skill Database Transfer。来源:Nahrendra et al., LocoVLM, Figure 8。相同的四足语言技能数据库被复用于 H1,执行 "go quickly""shh! the baby is sleeping""you are a kangaroo"等命令。

作者没有把 Go1 Policy 直接放到 H1 上。

真正复用的是:

Skill Database \boxed{\text{Skill Database}} Skill Database

H1 重新训练了自己的 Style-Conditioned Locomotion Policy,但只使用:

ψ l e f t , ψ r i g h t \psi_{\mathrm{left}}, \psi_{\mathrm{right}} ψleft,ψright

两个腿部 Phase Offsets。

H1 只训练:

  • Trot-Like Alternating Gait;
  • Pronk-Like Hopping Gait。

然后直接使用原先为 Quadruped 生成的数据库中的前两个 Phase Offsets。

所以"Zero-Shot Across Embodiments"准确含义是:

VLM 与 Skill Database 不需要重新训练/重新生成,而目标 embodiment 仍然需要自己的低层 Locomotion Policy。

不能误解成:

"Go1 的 RL Policy 零样本直接控制 H1。"


24. Appendix Figure 9:技能数据库生成实现细节

图 9 Appendix 中再次给出的 Offline Skill Database Generation Pipeline。来源:Nahrendra et al., LocoVLM, Figure 9。该图与正文 Figure 3 对应同一数据生成流程,附录结合 Prompt Listings 给出更多实现细节。

Appendix 进一步说明:

  1. 三种 instruction category 分开生成;
  2. 每类明确要求 LLM 输出 n = 100 n=100 n=100 条;
  3. 再把 instruction 送入统一 Meta-Prompt;
  4. 输入列表会 shuffle;
  5. 防止 LLM 记忆固定 input-output 顺序;
  6. 输出保存成结构化 JSON。

25. 原论文 Prompt 的关键信息

Skill Prompt 告诉 GPT-4o:

T T T

0.2 < T ≤ 1 0.2<T\le1 0.2<T≤1

并特别提示:

0.3 ≤ T ≤ 0.6 0.3\le T\le0.6 0.3≤T≤0.6

通常更稳定。

Gait Phase Offsets

使用五种基础 gait 作为参考:

text 复制代码
trot          [0.0, 0.5, 0.5, 0.0]
rotary_gallop [0.0, 0.2, 0.7, 0.5]
pace          [0.0, 0.5, 0.0, 0.5]
pronk         [0.0, 0.0, 0.0, 0.0]
bound         [0.0, 0.0, 0.5, 0.5]

但 Prompt 同时强调不要只复制 gait dictionary,而要根据语义创造新的描述。

Velocity Limit

Prompt 告诉模型:

v x l i m i t ≤ 1.5 m / s v_x^{\mathrm{limit}} \le 1.5\ \mathrm{m/s} vxlimit≤1.5 m/s

并通过:

  • Slow;
  • Fast;
  • Stop;

等语义调整数值。

这相当于在 Foundation Model 上增加一个机器人可执行空间的人工先验边界。


26. Reasoning 示例

表 5 原论文 Table IV:语言指令与 GPT-4o Reasoning。

Instruction Reasoning
trundle along like a hippo slow and heavy trot, lower vel_lim and increase T
oh no! catch that thief running! fast and aggressive gait, low T and high vel_lim
the sound of a human voice, stay hidden. use a trot with high T for stealthy movement, low vel_lim for quietness
a busy marketplace, navigate through the crowd. slow pace with moderate T for careful navigation

可以看到 Prompted Reasoning 实际完成了:

Semantic Concept → Locomotion Concept \text{Semantic Concept} \rightarrow \text{Locomotion Concept} Semantic Concept→Locomotion Concept

例如:

stealth → slow gait + high T + low velocity \text{stealth} \rightarrow \text{slow gait} + \text{high } T + \text{low velocity} stealth→slow gait+high T+low velocity

这一步是 LLM 知识真正进入机器人 Gait Parameter Space 的桥梁。


27. Appendix:将视觉语义作为导航约束

这部分很值得注意,因为它说明 LocoVLM 不只能改变 gait,还可以给传统 Navigation Stack 提供语义速度约束。

系统仅取:

v x l i m i t v_x^{\mathrm{limit}} vxlimit

然后把它作为 Local Planner 的最大速度约束。

27.1 从宽阔区域进入狭窄区域

图 10 LocoVLM 输出的速度上限对局部规划器的约束。来源:Nahrendra et al., LocoVLM, Figure 10。机器人进入狭窄、拥挤区域后,VLM 降低 v x l i m i t v_x^{\mathrm{limit}} vxlimit,局部规划器仍自行计算具体速度,但必须满足该上限。

这里:

v c m d ≤ v x l i m i t v_{\mathrm{cmd}} \le v_x^{\mathrm{limit}} vcmd≤vxlimit

VLM 不替代 Local Planner,而是在它外面增加 Semantic Constraint。

论文将 VLM Inference Period 设置为:

5 s 5\ \mathrm{s} 5 s

避免速度上限频繁变化导致 jitter。

27.2 轨迹安全性

图 11 有无 LocoVLM 语义速度约束时的导航轨迹对比。来源:Nahrendra et al., LocoVLM, Figure 11。绿色为加入 LocoVLM Constraint,粉色为无该约束;加入语义速度限制后,机器人在狭窄障碍区域保持更大的 Obstacle Clearance。

这个实验的意义在于:

Foundation Model \text{Foundation Model} Foundation Model

不一定非要直接输出机器人 Action。

它也可以输出:

Constraint \text{Constraint} Constraint

再让传统 Planner / Controller 在约束内工作。

这是一种非常实用的:

Semantic Constraint + Classical/RL Control \boxed{ \text{Semantic Constraint} + \text{Classical/RL Control} } Semantic Constraint+Classical/RL Control

架构。


28. LocoVLM 完整数据流

text 复制代码
                         OFFLINE
┌─────────────────────────────────────────────────────────────┐
│                                                             │
│  Prompt ──► GPT-4o ──► Instruction Set                     │
│                         │                                   │
│                         ▼                                   │
│                    Meta-Prompt                              │
│                         │                                   │
│                         ▼                                   │
│  {instruction, reasoning, T, phase offsets, vel_lim}        │
│                         │                                   │
│                         ▼                                   │
│                   Skill Database D                          │
│                                                             │
└─────────────────────────────────────────────────────────────┘

                         ONLINE
┌─────────────────────────────────────────────────────────────┐
│                                                             │
│  Text Query ─────┐                                         │
│                  ├──► BLIP-2 Encoder ─► Cosine Top-K        │
│  RGB Image ──────┘                         │                │
│                                            ▼                │
│                                     BLIP-2 ITM Re-rank      │
│                                            │                │
│                                            ▼                │
│                             Motion Descriptor M*             │
│                               (T, ψ, vel_lim)               │
│                                            │                │
│             Proprioception ────────────────┤                │
│             Velocity Command ──────────────┤                │
│                                            ▼                │
│                          Style-Conditioned PPO Policy        │
│                                            │                │
│                                            ▼                │
│                               Desired Joint Angles          │
│                                            │                │
│                                            ▼                │
│                                   Unitree Go1               │
│                                                             │
└─────────────────────────────────────────────────────────────┘

29. PPO 在这篇论文里负责什么

PPO 并不负责理解语言。

语言/视觉语义部分由:

GPT-4o + BLIP-2 \text{GPT-4o} + \text{BLIP-2} GPT-4o+BLIP-2

完成。

PPO 解决的是:

( s t , T , ψ , v x ) → a t (s_t,T,\psi,v_x) \rightarrow a_t (st,T,ψ,vx)→at

即:

给定机器人本体状态、步态周期、相位关系和速度命令,生成稳定可执行的运动。

典型 PPO Probability Ratio:

r t ( θ ) = π θ ( a t ∣ s t ) π θ o l d ( a t ∣ s t ) r_t(\theta) =\frac{ \pi_\theta(a_t\mid s_t) }{ \pi_{\theta_{\mathrm{old}}}(a_t\mid s_t) } rt(θ)=πθold(at∣st)πθ(at∣st)

Clipped Objective(裁剪目标函数):

L C L I P = E t min ⁡ ( r t ( θ ) A \^ t , clip ⁡ ( r t ( θ ) , 1 − ϵ , 1 + ϵ ) A \^ t ) L^{\mathrm{CLIP}} =\mathbb{E}_t \left \\min \\left( r_t(\\theta)\\hat{A}_t, \\operatorname{clip} \\left( r_t(\\theta), 1-\\epsilon, 1+\\epsilon \\right) \\hat{A}_t \\right) \\right LCLIP=Etmin(rt(θ)A\^t,clip(rt(θ),1−ϵ,1+ϵ)A\^t)

PPO 的价值是让 Locomotion Policy 在高维连续控制空间中稳定迭代,而 LocoVLM 的新意并不在 PPO 算法本身,而在:

Style Conditioning + Compliant Contact Reward + Semantic Skill Interface \text{Style Conditioning} + \text{Compliant Contact Reward} + \text{Semantic Skill Interface} Style Conditioning+Compliant Contact Reward+Semantic Skill Interface


30. Asymmetric Actor-Critic 的含义

论文采用 Asymmetric Actor-Critic。

其一般形式是:

a t ∼ π θ ( a t ∣ o t ) a_t \sim \pi_\theta(a_t\mid o_t) at∼πθ(at∣ot)

Actor 只使用真机可获得的 Observation:

o t o_t ot

而 Critic 在训练阶段允许使用更完整的 Privileged State:

V ϕ ( s t ) V_\phi(s_t) Vϕ(st)

其中:

s t ⊇ o t s_t\supseteq o_t st⊇ot

这样训练时 Critic 能获得更准确的 Value Estimate,但部署时 Actor 不依赖仿真器特权信息。

这是强化学习 Sim-to-Real Locomotion 中常见的设计:

Train with Privileged Information → Deploy with Realistic Observations \boxed{ \text{Train with Privileged Information} \rightarrow \text{Deploy with Realistic Observations} } Train with Privileged Information→Deploy with Realistic Observations

需要注意:LocoVLM 正文没有进一步完整列出所有 Actor/Critic 网络层尺寸,因此不能根据其他 Locomotion 项目臆造具体 MLP 结构。


31. Sim-to-Real 设计

论文明确随机化:

m r o b o t m_{\mathrm{robot}} mrobot

C o M \mathrm{CoM} CoM

k p / k d -related motor properties k_p/k_d\text{-related motor properties} kp/kd-related motor properties

μ t e r r a i n \mu_{\mathrm{terrain}} μterrain

以及:

System Delay \text{System Delay} System Delay

目的都是扩大训练环境分布:

p s i m ( θ d y n ) p_{\mathrm{sim}}(\theta_{\mathrm{dyn}}) psim(θdyn)

让真机真实动力学:

θ r e a l \theta_{\mathrm{real}} θreal

更可能落在训练分布覆盖范围内。

因此:

Domain Randomization + State Estimation + Compliant Contact Tracking \text{Domain Randomization} + \text{State Estimation} + \text{Compliant Contact Tracking} Domain Randomization+State Estimation+Compliant Contact Tracking

共同支撑 Go1 真机部署。


32. LocoVLM 与 SayTap 的核心区别

对比维度 SayTap LocoVLM
高层模型 GPT-4 GPT-4o + BLIP-2
Vision 无 有
LLM 是否在线 是,高层在线生成 GPT-4o 仅离线生成数据库
在线语义模块 GPT-4 BLIP-2 Retrieval
中间表示 Foot Contact Pattern ( T , ψ , v x l i m i t ) (T,\psi,v_x^{\mathrm{limit}}) (T,ψ,vxlimit)
数据生成 Prompt Few-Shot 两阶段批量生成 + Prompted Reasoning
检索 无 Mixed-Precision Retrieval
Text-as-Image 无 有
低层 RL Contact-Conditioned Policy Style-Conditioned PPO Policy
鲁棒性机制 RL + Contact Pattern Compliant Contact Tracking
图像场景适应 无 有
Cross-Embodiment 非核心实验 Go1 → H1 Skill Database Transfer

可以把两者的思想演化写成:

SayTap : Language → Contact Pattern → RL \text{SayTap}: \quad \text{Language} \rightarrow \text{Contact Pattern} \rightarrow \text{RL} SayTap:Language→Contact Pattern→RL

LocoVLM : Vision/Language → Semantic Retrieval → ( T , ψ , v l i m ) → RL \text{LocoVLM}: \quad \text{Vision/Language} \rightarrow \text{Semantic Retrieval} \rightarrow (T,\psi,v_{\mathrm{lim}}) \rightarrow \text{RL} LocoVLM:Vision/Language→Semantic Retrieval→(T,ψ,vlim)→RL

LocoVLM 很明显吸收了 SayTap 的"Foundation Model 不直接控制关节,而输出运动中间表示"的思想,但进一步解决了:

  • Vision Grounding;
  • 在线 LLM 依赖;
  • 数据扩展成本;
  • 检索效率;
  • 严格 Gait Tracking 的鲁棒性问题。

33. LocoVLM 是不是严格意义上的 VLA

严格按照当前主流 Vision-Language-Action 模型定义,例如:

( Image , Language ) → Single Learned Policy Action Chunk (\text{Image},\text{Language}) \xrightarrow{\text{Single Learned Policy}} \text{Action Chunk} (Image,Language)Single Learned Policy Action Chunk

LocoVLM 不是典型端到端 VLA。

它更准确的结构是:

( Image or Language ) → B L I P - 2 Retrieved Motion Descriptor → R L Joint Action (\text{Image or Language}) \xrightarrow{\mathrm{BLIP\text{-}2}} \text{Retrieved Motion Descriptor} \xrightarrow{\mathrm{RL}} \text{Joint Action} (Image or Language)BLIP-2 Retrieved Motion DescriptorRL Joint Action

因此它属于:

Hierarchical Vision-Language-Guided Locomotion Framework(分层视觉语言引导运动框架)

而不是 OpenVLA、RT-2、 π 0 \pi_0 π0 那类统一模型直接 Decode Robot Action。

但从机器人系统工程角度,它有一个很重要的启发:

VLM/VLA 不一定需要进入最内层动态控制环 \boxed{ \text{VLM/VLA 不一定需要进入最内层动态控制环} } VLM/VLA 不一定需要进入最内层动态控制环

对于四足/人形这种高速动态系统,分层设计常常更实际:

Foundation Model → Skill / Constraint → Specialized Locomotion Policy \text{Foundation Model} \rightarrow \text{Skill / Constraint} \rightarrow \text{Specialized Locomotion Policy} Foundation Model→Skill / Constraint→Specialized Locomotion Policy


34. 论文的主要创新点

34.1 Vision + Language 到 Locomotion Style 的实时 Grounding

它不是简单给机器人一句"走快点",而是把:

Semantic Scene \text{Semantic Scene} Semantic Scene

转换成:

( T , ψ , v x l i m i t ) (T,\psi,v_x^{\mathrm{limit}}) (T,ψ,vxlimit)

真正改变步态。

34.2 LLM 离线知识蒸馏

GPT-4o:

Teacher \text{Teacher} Teacher

BLIP-2 + Skill Database:

Student / Knowledge Carrier \text{Student / Knowledge Carrier} Student / Knowledge Carrier

避免云端 LLM 在线控制。

34.3 两阶段 Instruction / Motion Descriptor 数据生成

将:

Instruction Generation \text{Instruction Generation} Instruction Generation

和:

Motion Parameter Generation \text{Motion Parameter Generation} Motion Parameter Generation

分开,使数据规模扩展更便宜、更结构化。

34.4 Prompted Reasoning 提高可执行描述符质量

把:

Vague Semantic Instruction \text{Vague Semantic Instruction} Vague Semantic Instruction

先翻译成:

Technical Locomotion Reasoning \text{Technical Locomotion Reasoning} Technical Locomotion Reasoning

再输出数值参数。

34.5 Mixed-Precision Retrieval

将:

Fast Embedding Search + Accurate ITM Matching \text{Fast Embedding Search} + \text{Accurate ITM Matching} Fast Embedding Search+Accurate ITM Matching

结合,避免全库高成本 ITM。

34.6 Text-as-Image

利用 VLM 原生 Image-Text Alignment 能力弥补 Text-Text Retrieval 较弱的问题。

34.7 Compliant Contact Tracking

把 Gait Tracking 从硬约束改成柔顺目标,在粗糙地形中获得更好的 Stability / Style Trade-Off。

34.8 Cross-Embodiment Skill Database Reuse

同一个高层语义数据库可以映射到新的 embodiment,只要目标机器人拥有兼容的低层 Style-Conditioned Controller。


35. 学术贡献的本质

这篇论文最值得注意的并不是某个单一网络,而是提出了一种明确的机器人 Foundation Model 系统划分:

Semantic Intelligence ≠ Dynamic Control Intelligence \boxed{ \text{Semantic Intelligence} \neq \text{Dynamic Control Intelligence} } Semantic Intelligence=Dynamic Control Intelligence

前者交给:

LLM / VLM \text{LLM / VLM} LLM / VLM

后者交给:

Specialized RL Locomotion Policy \text{Specialized RL Locomotion Policy} Specialized RL Locomotion Policy

中间通过:

M = ( T , ψ , v x l i m i t ) \mathbf{M} =(T,\psi,v_x^{\mathrm{limit}}) M=(T,ψ,vxlimit)

连接。

这种做法避免要求一个大模型同时学习:

  • 世界知识;
  • 图像理解;
  • 社会语义;
  • Gait Planning;
  • Contact Dynamics;
  • 50 Hz Joint Control。

论文实际上提出了一个非常清晰的 Semantic-to-Dynamics Interface。


36. 局限性

36.1 Image 和 Text 不能真正联合输入

作者在结论中明确指出,当前 LocoVLM 分别处理:

Image \text{Image} Image

或:

Text \text{Text} Text

但不能自然完成:

( Image , Text ) (\text{Image},\text{Text}) (Image,Text)

联合条件推理。

例如:

"看到冰面以后慢走,但如果我说可以跑就恢复高速。"

当前框架无法像真正多模态 VLA 那样联合解释视觉和语言约束。

36.2 Retrieval 不是生成式 Motion Planning

LocoVLM 最终选择的是数据库中的已有技能:

I ∗ ∈ D \mathcal{I}^{*} \in \mathcal{D} I∗∈D

即使 Query 是新的,其实际行为仍来自:

Nearest Skill Entry \text{Nearest Skill Entry} Nearest Skill Entry

所以表达能力上限受到数据库覆盖度约束。

36.3 Motion Descriptor 维度仍较低

当前只有:

( T , ψ , v x l i m i t ) (T,\psi,v_x^{\mathrm{limit}}) (T,ψ,vxlimit)

没有显式描述:

  • Footstep Position;
  • Foothold;
  • Body Height;
  • Body Orientation;
  • Swing Height;
  • Contact Force;
  • Terrain-Specific Foothold;
  • Yaw Rate Limit;
  • Lateral Velocity。

因此它更适合"运动风格自适应",而不是完整复杂地形足端规划。

36.4 VLM 仍依赖外部 GPU

真机 Locomotion Policy 在 Jetson Xavier NX 上运行,但 BLIP-2 放在 RTX 3070 Ti Laptop。

所以:

"无云端 LLM"

并不等于:

"所有模型都完全在 Go1 板载 Xavier NX 上运行"。

36.5 87% 不是开放世界任务成功率

87% 来源于:

100 100 100

条人工标注 Instruction 的 Retrieval Benchmark。

因此不能把它解释为:

"在任何视觉语言场景下机器人都有 87% 成功率"。

36.6 H1 泛化并非完整控制策略 Zero-Shot

复用的是:

Skill Database + VLM Mapping \text{Skill Database} + \text{VLM Mapping} Skill Database+VLM Mapping

而 H1 的 Locomotion Controller 仍然重新训练。

36.7 Prompt 中仍有较强人工机器人先验

GPT-4o 并不是自由发现所有可行 gait。

Prompt 已经提供:

  • 五种 Gait Dictionary;
  • T T T 合理范围;
  • Velocity Limit;
  • Quadruped Gait 知识。

所以本质是:

Foundation Model Commonsense + Human-Specified Locomotion Prior \text{Foundation Model Commonsense} + \text{Human-Specified Locomotion Prior} Foundation Model Commonsense+Human-Specified Locomotion Prior


37. 对四足 VLA 研究的启发

如果在此基础上做真正的四足 VLA,可以将 LocoVLM 的中间层继续升级。

37.1 从 Retrieval Descriptor 到 Learned Skill Token

当前:

z = ( T , ψ , v x l i m i t ) z =(T,\psi,v_x^{\mathrm{limit}}) z=(T,ψ,vxlimit)

未来可以学习:

z m o t i o n ∈ R d z_{\mathrm{motion}} \in \mathbb{R}^{d} zmotion∈Rd

由大量四足 Motion / RL Trajectory 自动提取。

于是:

VLM → z m o t i o n → Low-Level Policy \text{VLM} \rightarrow z_{\mathrm{motion}} \rightarrow \text{Low-Level Policy} VLM→zmotion→Low-Level Policy

减少手工 Gait Dictionary。

37.2 加入 Footstep-Level Planning

进一步输出:

M = ( T , ψ , v , p f o o t , h s w i n g ) \mathbf{M} =( T, \psi, v, p_{\mathrm{foot}}, h_{\mathrm{swing}} ) M=(T,ψ,v,pfoot,hswing)

这样不仅能说:

"雪地慢点走。"

还能确定:

"下一步脚具体踩哪里。"

37.3 Vision + Language Joint Conditioning

当前:

Image OR Language \text{Image OR Language} Image OR Language

可以升级为:

Image AND Language \text{Image AND Language} Image AND Language

例如:

图像看到楼梯 + 用户说"快速上楼"。

统一 VLM 输出 Gait / Foothold / Velocity Constraint。

37.4 加入 Safety / Feasibility Critic

VLM 输出:

M V L M \mathbf{M}_{\mathrm{VLM}} MVLM

以后先经过:

Dynamics Feasibility + Safety Critic \text{Dynamics Feasibility} + \text{Safety Critic} Dynamics Feasibility+Safety Critic

再送低层 Policy:

M s a f e \mathbf{M}_{\mathrm{safe}} Msafe

可以进一步降低 Hallucinated Gait Parameters 的风险。


38. 和端到端 VLA 的两种路线

路线 A:End-to-End

( I , L , P ) → Transformer / Flow / Diffusion → a t : t + H (I,L,P) \rightarrow \text{Transformer / Flow / Diffusion} \rightarrow a_{t:t+H} (I,L,P)→Transformer / Flow / Diffusion→at:t+H

优点:

  • 统一;
  • 可以自动学习中间表示。

问题:

  • 四足高频控制难;
  • 动力学稳定性要求高;
  • 训练数据需求巨大。

路线 B:LocoVLM 式 Hierarchical VLA

( I , L ) → z s k i l l → π l o c o m o t i o n → a t (I,L) \rightarrow z_{\mathrm{skill}} \rightarrow \pi_{\mathrm{locomotion}} \rightarrow a_t (I,L)→zskill→πlocomotion→at

优点:

  • 可解释;
  • 低层已有成熟 RL Policy;
  • 语义层低频即可;
  • 真机安全边界更清晰。

对于当前四足机器人 Foundation Model 研究,第二条路线往往更容易先完成可工作的 Sim-to-Real 系统。


39. 关键参数汇总

表 6 LocoVLM 关键系统参数。

项目 设置
四足平台 Unitree Go1
人形验证平台 Unitree H1
LLM GPT-4o
VLM BLIP-2
四足 RL PPO
Actor-Critic Asymmetric Actor-Critic
四足仿真 Isaac Gym
H1 泛化仿真 MuJoCo
四足 Policy 频率 50 Hz
Motor Controller 200 Hz
四足板载计算 Jetson Xavier NX
VLM GPU RTX 3070 Ti Laptop
VLM 推理 + 通信 < 100 m s <100\ \mathrm{ms} <100 ms
Skill Descriptor M = ( T , ψ , v x l i m i t ) \mathbf{M}=(T,\psi,v_x^{\mathrm{limit}}) M=(T,ψ,vxlimit)
基础 Gait 数量 5
Compliance Threshold δ = 0.5 \delta=0.5 δ=0.5
Contact Smoothing σ = 0.25 \sigma=0.25 σ=0.25
Gait Test Velocity 1.2 m / s 1.2\ \mathrm{m/s} 1.2 m/s
Gait Test Period 0.4 s 0.4\ \mathrm{s} 0.4 s
Gait Test Rollouts 1000
Retrieval Benchmark 100 Instructions
最佳 Retrieval 87%
Navigation VLM Period 5 s 5\ \mathrm{s} 5 s
GPT-4o 是否在线控制 否

40. 论文价值总结

LocoVLM 的研究价值可以浓缩成三个层次。

第一层是语言/视觉语义到运动参数的 Grounding:

Semantic Understanding → Locomotion Parameters \text{Semantic Understanding} \rightarrow \text{Locomotion Parameters} Semantic Understanding→Locomotion Parameters

第二层是Foundation Model 与实时控制解耦:

Offline LLM + Online VLM + High-Frequency RL \text{Offline LLM} + \text{Online VLM} + \text{High-Frequency RL} Offline LLM+Online VLM+High-Frequency RL

第三层是把 Gait Style 设计为柔顺目标而不是绝对约束:

Style Fidelity + Disturbance Recovery \text{Style Fidelity} + \text{Disturbance Recovery} Style Fidelity+Disturbance Recovery

它不是一个典型的端到端 VLA,但对于四足机器人而言,这种层级式结构非常有代表性:

VLM/LLM 负责"应该怎么动" + RL Policy 负责"如何稳定地动" \boxed{ \text{VLM/LLM 负责"应该怎么动"} + \text{RL Policy 负责"如何稳定地动"} } VLM/LLM 负责"应该怎么动"+RL Policy 负责"如何稳定地动"

如果研究目标是将 Vision-Language Foundation Model 真正部署到 Go1、Go2、ANYmal 等动态足式平台,LocoVLM 展示了一条很清晰的工程路线:不必让大模型直接承担 50--200 Hz 的动力学控制,而可以利用一个低维、可解释、动力学相关的中间技能空间连接两端。


41. 参考资料

  1. LocoVLM Project Page

    https://locovlm.github.io/

  2. LocoVLM: Grounding Vision and Language for Adapting Versatile Legged Locomotion Policies

    https://arxiv.org/abs/2602.10399

  3. LocoVLM Full Paper PDF

    https://locovlm.github.io/static/images/locovlm_paper.pdf

  4. ICRA 2025 SafeVLMs Workshop Paper

    https://locovlm.github.io/static/images/workshop_final.pdf

  5. LocoVLM Demo Video

    https://www.youtube.com/watch?v=smXBsOTeXrE

  6. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

    https://arxiv.org/abs/2301.12597

  7. SayTap: Language to Quadrupedal Locomotion

    https://arxiv.org/abs/2306.07580


相关推荐
一直在努力的小宁4 小时前
【阅读笔记】具身操作的数采方案概览
后端·json·restful·具身智能·vla·vlm
音视频牛哥2 天前
四足机器人如何做好智慧电力巡检?从低延迟图传到可信巡检数据闭环
音视频·四足机器人·智慧电力巡检·whip whep·机器人巡检·工业音视频巡检·智能电力巡检技术方案
一颗小树x10 天前
vLLM大模型推理:Jetson AGX Thor 的 GPU 共享内存清理实战
jetson·vllm·vlm·gpu内存清理
山顶夕景11 天前
【VLM】Qwen3.8-Omni-Flash长音视频理解Agentic模型
音视频·vlm·视频理解·agentic·全模态
程序猿编码12 天前
不用 GPU!RK3588 跑视觉大模型,Qwen3.5-2B VLM 端侧落地
大模型·llm·rk3588·qwen·推理·vlm
Robot_Nav14 天前
前沿 | VLA 演进:从动作 Token 到分层具身智能体
具身智能·vla·vlm
一直在努力的小宁15 天前
[特殊字符] 具身智能Agent开发调研|零基础超详细笔记(万字长文)
agent·vla·vlm·vln·harness
Robot_Nav23 天前
T-RO 2026 | 一个滤波器适配所有控制器:未知环境下四足机器人鲁棒安全导航【文献解读】
四足机器人·未知环境·鲁棒安全控制
山顶夕景24 天前
【Omni】OmniGAIA: Towards Native Omni-Modal AI Agents
agent·多模态·vlm·omni·全模态