← Back to Index
Daily Research Digest

arXiv Papers

2026-09-15
115
Papers
1
Categories
115
Translated
收藏清单 0
机器人学 (Robotics)
115
cs.RO / 1 / 2609.13224

Seeing What the Vehicle Sees: Video-Augmented Virtual Reality for Physical Autonomous Vehicles

看见车辆所见:面向实体自动驾驶车辆的视频增强虚拟现实
Islam, Md Tanjemul, Shafin, Mohammad, Kabir, Md Rafiul
Abstract
Autonomous vehicles are expected to improve road safety and efficiency, but passengers often remain uncertain about what the vehicle perceives and why it acts as it does. Virtual reality (VR) offers a safe and repeatable medium for presenting this information, yet most passenger-facing VR studies rely on fully simulated vehicles or pre-scripted scenarios, so the motion and perception shown to the user do not originate from a physically operating autonomous system. This paper presents a video-augmented VR framework that couples a physical ROS 2 autonomous robot vehicle to a Unity 6 application deployed on a Meta Quest 3S headset. The vehicle state and live onboard camera stream are transmitted over two independent communication channels, allowing the virtual vehicle to mirror the physical robot's motion while the passenger simultaneously views the vehicle's first-person camera feed and its navigation decisions through an in-vehicle dashboard interface. We evaluate the framework over 20 repeated closed-loop navigation trials. The system achieves a mean state-update latency of 29.63 ms, a mean relative route-progress error of 2.28% between the physical and virtual vehicles, and video delivery at 10.006 frames per second with 0.25% frame loss. All monitored navigation decisions were correctly reflected in the VR interface with no missed or incorrect notifications. The results indicate that the framework can support temporally synchronized, semantically consistent, and accurate route-progress representation for immersive observation of physical autonomous-vehicle behavior.
Chinese Translation
自动驾驶车辆有望提升道路安全与效率,但乘客往往仍不确定车辆感知到了什么,以及它为何如此行动。虚拟现实(VR)为呈现这些信息提供了一种安全且可重复的媒介,然而大多数面向乘客的VR研究依赖于完全仿真的车辆或预先脚本化的场景,因此呈现给用户的运动与感知并非源自实际运行的自动驾驶系统。本文提出了一种视频增强的VR框架,将一台实体ROS 2自主机器人车辆与部署在Meta Quest 3S头显上的Unity 6应用相耦合。车辆状态和实时车载摄像头视频流通过两条独立通信信道传输,使虚拟车辆能够镜像物理机器人的运动,同时乘客可通过车内仪表盘界面同时查看车辆的第一人称摄像头画面及其导航决策。我们在20次重复闭环导航试验中评估了该框架。系统实现了29.63 ms的平均状态更新延迟,实体车辆与虚拟车辆之间的平均相对路线进度误差为2.28%,视频以10.006帧/秒传输且帧丢失率为0.25%。所有监测到的导航决策均在VR界面中正确反映,无遗漏或错误通知。结果表明,该框架能够支持时间同步、语义一致且准确的路线进度表示,用于对实体自动驾驶车辆行为的沉浸式观察。
cs.RO / 2 / 2609.13231

ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models

ShieldVLA:面向视觉-语言-动作模型的可行性感知安全对齐
Tayal, Manan, Nambi, Akshay
Abstract
Vision-Language-Action (VLA) models demonstrate strong generalization in robotic manipulation and navigation, but existing fine-tuning methods provide limited safety guarantees. Current approaches primarily rely on Lagrangian optimization that enforces safety through soft penalties on expected cumulative cost, often resulting in residual constraint violations or overly conservative behavior. Moreover, learning safety in visual domains is challenging due to the absence of dense per-step safety annotations. We propose ShieldVLA, a safety-aligned fine-tuning framework for VLA models based on Hamilton-Jacobi (HJ) reachability. ShieldVLA learns a model-free approximation of the HJ reachability value function directly from visual observations to estimate the safe operating region. The learned safety critic gates policy optimization by separating reward maximization within feasible regions from recovery near unsafe states, avoiding persistent reward-cost trade-offs. To enable scalable supervision in visual environments, we introduce rubric-based VLM safety scores that convert semantic safety feedback into structured critic targets without requiring manual cost labels. Across five navigation and manipulation benchmarks spanning multiple VLA backbones, ShieldVLA reduces cumulative safety cost by 57% on average and improves task success rate by +0.13 over SafeVLA.
Chinese Translation
视觉-语言-动作(VLA)模型在机器人操作和导航中展现出强大的泛化能力,但现有的微调方法提供的安全保障有限。当前方法主要依赖拉格朗日优化,通过对期望累积代价的软惩罚来强制安全,往往导致残余约束违反或过度保守的行为。此外,由于缺乏密集的逐步安全标注,在视觉领域学习安全性具有挑战性。我们提出了ShieldVLA,一种基于汉密尔顿-雅可比(HJ)可达性的VLA模型安全对齐微调框架。ShieldVLA直接从视觉观察中学习HJ可达性值函数的无模型近似,以估计安全操作区域。学习到的安全评价器通过将可行区域内的奖励最大化与不安全状态附近的恢复分开,来门控策略优化,避免了持续的奖励-成本权衡。为了在视觉环境中实现可扩展的监督,我们引入了基于评分标准的VLM安全评分,将语义安全反馈转换为结构化的评价器目标,而无需手动成本标签。在跨越多个VLA主干的五个导航和操作基准测试中,ShieldVLA平均将累积安全成本降低了57%,并将任务成功率比SafeVLA提高了+0.13。
cs.RO / 3 / 2609.13234

Harnessing human expertise for high-precision robotic assembly in industrialized construction: A sample-efficient installer-in-the-loop interactive reinforcement learning framework

利用人类专业知识实现工业化建筑中的高精度机器人装配:一种样本高效的安装工在环交互式强化学习框架
Jin, Zekai, Wang, Huiguang, Sun, Xiaoning, Shao, Yi
Abstract
Industrialized construction imposes stringent precision requirements on robotic assembly of modular components such as prefabricated window units. In tolerance-critical operations, the central bottleneck is not only mechanical clearance but also converting tacit installer expertise into data-efficient autonomy under sparse acceptance feedback, contact variability, and millimeter-scale constraints. We present an installer-in-the-loop interactive reinforcement learning framework that acquires expertise through offline teleoperated demonstrations, sparse event-driven binary takeovers at contact-failure boundaries, and acceptance-aligned terminal rewards, logged under a unified schema for traceable offline-to-online adaptation. A temporally abstract action-sequence policy built on Q-chunking with Flow Q-Learning captures multimodal recovery maneuvers under sparse terminal rewards, while a non-updating warm-start phase stabilizes the offline-to-online transition. The framework is evaluated in MuJoCo across the workflow from suction acquisition through clearance-limited seating, under structured staging and end-to-end randomized placement. Within a defined stress-test regime with 2 mm per-side clearance, bounded pose perturbations, and friction randomization, the pipeline attains 100\% autonomous seating with 12--15 min of cumulative installer supervision over 3.0 h of online training, and reaches the 95\% success milestone in approximately 0.5 h and 1.5 h in the two experiments. We also report wall-clock adaptation time, cumulative takeover minutes, intervention-rate decay, and stage-wise failure attribution to inform supervision budgeting. Ablations isolate the complementary contributions of temporal abstraction, installer intervention, and warm-start value calibration.
Chinese Translation
工业化建筑对模块化构件(如预制窗单元)的机器人装配提出了严格的精度要求。在公差关键操作中,核心瓶颈不仅是机械间隙,还包括在稀疏验收反馈、接触变异性和毫米级约束下,将隐性的安装工专业知识转化为数据高效的自主性。我们提出了一种安装工在环交互式强化学习框架,该框架通过离线遥操作演示、接触失败边界处的稀疏事件驱动二元接管以及与验收对齐的终端奖励来获取专业知识,并在统一模式下记录,以实现可追溯的离线到在线自适应。基于Q-chunking与Flow Q-Learning构建的时间抽象动作序列策略,能够在稀疏终端奖励下捕获多模态恢复动作,而非更新热启动阶段则稳定了离线到在线的过渡。该框架在MuJoCo中进行评估,涵盖从吸盘获取到间隙受限就位的整个工作流程,分别在结构化分段和端到端随机放置条件下进行。在定义的应力测试机制下(每侧间隙2 mm、有界位姿扰动和摩擦随机化),该流程在3.0小时的在线训练中,累计安装工监督12-15分钟,实现了100%的自主就位,并在两个实验中分别在约0.5小时和1.5小时达到95%的成功里程碑。我们还报告了实际时钟自适应时间、累计接管分钟数、干预率衰减以及分阶段故障归因,以指导监督预算。消融实验分离了时间抽象、安装工干预和热启动价值校准的互补贡献。
cs.RO / 4 / 2609.13235

Learning Manipulation-Sufficient Representations via Outcome Bottlenecks

通过结果瓶颈学习操作充分的表示
Sarowar, Md Selim, Kim, Sungho
Abstract
Networked manipulation endpoints couple perception to actuation across compute- and bandwidth-limited links, yet commonly exchange dense geometric states optimized for fidelity rather than action outcomes. A stochastic representation is learned with a policy-free, action-conditioned outcome bottleneck: marginal outcome log-loss supplies distortion and a KL term regularizes rate. The construction is motivated by the minimal statistic that preserves the outcome distribution of every admissible action, while the implemented finite model is evaluated as a rate-regularized mixture predictor. The same encoder and outcome head support grasp selection, singleton conformal filtering, active viewpoint selection, and latent test-time adaptation. A finite-probe theorem identifies the local level-set tangent space with the null space of an outcome Jacobian. The synthetic oracle verifies this result; on scanned objects, an analytic surrogate agrees with measured simulator invariances within \(0.56^\circ\). Across 11,979 simulated grasps on 13 objects, a reconstructed-geometry wrench score attains 0.542 AUC against lift success and falls below chance on curved objects, while our representation attains 0.876. At 25\% commitment, executed-grasp success is 0.503 versus 0.984. The 512-byte interface is \(288\times\) smaller than one RGB-D frame and runs at 16\,ms per CPU decision. Independent synthetic points track the tested conformal levels; scanned-object all-pair coverage is reported as a clustered empirical diagnostic. On unseen objects, within-scene AUC falls to 0.569, and a full-feedback update raises empirical mean pairwise coverage from 0.728 to 0.883.
Chinese Translation
网络化操作端点通过计算和带宽受限的链路将感知与执行耦合,但通常交换的是为保真度而非动作结果优化的密集几何状态。学习一种随机表示,采用无策略的、以动作为条件的结果瓶颈:边缘结果对数损失提供失真,KL项正则化速率。该构造的动机是保留每个允许动作的结果分布的最小统计量,而实现的有限模型则作为速率正则化的混合预测器进行评估。相同的编码器和结果头支持抓取选择、单例共形过滤、主动视角选择和潜在测试时适应。一个有限探测定理将局部水平集切空间与结果雅可比矩阵的零空间等同起来。合成预言机验证了该结果;在扫描物体上,解析代理与测量的模拟器不变性在0.56°内一致。在13个物体上的11,979次模拟抓取中,重建几何力旋量分数针对提升成功达到0.542 AUC,并在弯曲物体上低于随机水平,而我们的表示达到0.876。在25%承诺率下,执行抓取成功率为0.503,而对比为0.984。512字节接口比一帧RGB-D小288倍,并以每次CPU决策16 ms运行。独立合成点跟踪测试的共形水平;扫描物体全对覆盖率作为聚类经验诊断报告。在未见物体上,场景内AUC降至0.569,而全反馈更新将经验平均成对覆盖率从0.728提升至0.883。
cs.RO / 5 / 2609.13236

Self-Evolving AI for Humanoids: Mechanisms, Safety, and Evaluation of Post-Deployment Self-Improvement

面向人形机器人的自演化AI:部署后自我改进的机制、安全与评估
Nguyen, Loc X., Raha, Avi Deb, Le, Huy Q., Huh, Eui-Nam, Niyato, Dusit, Hong, Choong Seon
Abstract
Humanoid robots are becoming an important part of embodied artificial intelligence, driven by advances in reinforcement learning for locomotion, world models for prediction, and vision-language-action models for general control. However, most of these systems remain static after deployment. A policy is trained offline for a fixed objective and then frozen, even though the tasks, environments, and robot bodies keep drifting over time. An emerging paradigm of self-evolving agents aims to address this problem by allowing systems to improve from their own post-deployment experience. Since most existing studies focus on disembodied software agents, this survey examines how self-evolution changes when an agent has a physical body. We first define self-evolution for humanoids and represent a deployed robot using a state tuple that includes its policy, perception, memory, workflow, and body. This state is updated by an evolution operator in a slow outer loop with a lifelong objective. We then organize the literature into four complementary mechanisms of self-evolution, presented in increasing order of autonomy: self-learning, self-adaptation, self-optimization, and self-generation. Since changes to a humanoid can introduce physical hazards, we treat safety and uncertainty as key design dimensions of the evolution operator, and further formulate admissible evolution as a constraint enforced by a world-model verification gate within a human-oversight envelope. Finally, we present that evaluation should track the robot's evolving trajectory rather than a fixed checkpoint, and we identify the lack of a benchmark designed specifically for self-evolving humanoids. Moreover, we outline open challenges spanning AI algorithms, on-board systems, and governance.
Chinese Translation
人形机器人正在成为具身人工智能的重要组成部分,其推动力来自面向运动的强化学习、用于预测的世界模型以及用于通用控制的视觉-语言-动作模型等方面的进展。然而,这些系统在部署后大多保持静态。一个策略针对固定目标离线训练后被冻结,尽管任务、环境和机器人本体随时间不断漂移。新兴的自演化智能体范式旨在通过允许系统从自身部署后的经验中改进来解决这一问题。由于现有研究大多集中于无实体的软件智能体,本综述考察当智能体具有物理实体时,自演化会发生怎样的变化。我们首先为人形机器人定义自演化,并用一个状态元组表示已部署机器人,该元组包括其策略、感知、记忆、工作流和本体。该状态由一个演化算子以慢速外环和终身目标进行更新。然后,我们将文献组织为四种互补的自演化机制,并按其自主性递增的顺序呈现:自学习、自适应、自优化和自生成。由于对人形机器人的改变可能带来物理危害,我们将安全性和不确定性视为演化算子的关键设计维度,并进一步将可允许演化形式化为一个约束,该约束由人类监督范围内的世界模型验证门来强制执行。最后,我们提出评估应跟踪机器人不断演化的轨迹,而不是固定的检查点,并指出目前缺乏专门为自演化人形机器人设计的基准。此外,我们概述了跨越AI算法、机载系统和治理的开放挑战。
cs.RO / 6 / 2609.13243

GzDRL: Reproducible and Scalable Deep Reinforcement Learning with Gazebo

GzDRL:基于Gazebo的可复现且可扩展的深度强化学习
Haridevan, Amal Dev, Kang, Junjie, Shan, Jinjun
Abstract
We present GzDRL, a novel single-process reinforcement learning (RL) framework for Gazebo that overcomes longstanding bottlenecks in scalable, reproducible robotics experimentation. Unlike conventional middleware-based RL-Gazebo integrations that suffer from nondeterminism and irreproducibility, GzDRL introduces a systematic, middleware-free environment-stepping mechanism that directly synchronizes agent actions and physics updates. This design enables deterministic, high-throughput data collection, efficient vectorization, and reproducible RL training and evaluation. Comprehensive benchmarks demonstrate that GzDRL achieves the highest workstation throughput among the evaluated frameworks while remaining competitive with GPU-accelerated simulators on laptop hardware, and maintains precise agent-environment synchronization, multi-agent scalability, and experiment-level reproducibility. We further validate sim-to-real transfer by deploying learned policies directly onto a physical quadrotor, without fine-tuning. Our results establish GzDRL as an accessible and reproducible platform for advancing RL in robotics and automation.
Chinese Translation
我们提出了GzDRL,一种新颖的用于Gazebo的单进程强化学习(RL)框架,它克服了可扩展、可复现机器人实验中长期存在的瓶颈。与传统的基于中间件的RL-Gazebo集成不同(这些集成受到非确定性和不可复现性的困扰),GzDRL引入了一种系统性的、无中间件的环境步进机制,直接同步智能体动作与物理更新。该设计能够实现确定性、高通量的数据收集、高效向量化以及可复现的RL训练与评估。综合基准测试表明,GzDRL在评估框架中实现了最高的工作站吞吐量,同时在笔记本电脑硬件上与GPU加速模拟器相比仍具竞争力,并保持了精确的智能体-环境同步、多智能体可扩展性和实验级可复现性。我们进一步通过将学习到的策略直接部署到物理四旋翼飞行器上(无需微调)来验证仿真到现实迁移。我们的结果确立了GzDRL作为一个易用且可复现的平台,用于推进机器人和自动化中的RL。
cs.RO / 7 / 2609.13244

Physical Kernel: Structured Visual Latents for Dark Manipulation

物理内核:用于暗操作的结构化视觉潜变量
Hang, Jinting, Li, Hong, Cai, Zhenhui, Zhao, Zhihao, He, Jian
Abstract
We study dark manipulation: after a brief lit Write encodes z0 = Enc(rgb), a policy pi(z) and open-loop dynamics f(z,a) complete contact-rich skills without further pixels (dark_f). On ManiSkill StackCube (n=160; seed packs 0/1000), dark_f attains 68.1% stacked on the five-rung chain (near_A -> grasped -> lifted -> on_B -> stacked), compared with 35.6% for per-step lit_reenc and 0% for freeze/encode_black. On a shared Write->HOLD protocol (n=40), occlusion and camera-aligned GT contact-neighbor masks drive lit lift from 43% to 0%, while dark_f holds 82.5%; shuffling actions inflates dynamics MSE by ~9.4x; write-time appearance shifts break encoding (night: 0% stacked), yet the same shifts during HOLD leave dark_f lift unchanged; Write length Tw is flat once the stop phase is reached, while earlier stops and write-time blur/JPEG sharply cut stacked. A dedicated pi_write reaches 35% vision-budget stacked (n=80); matched Dreamer-style/pixel nulls without privileged geom stay at 0%. Privileged state-RSSM MPC reaches ~35% stacked with 9D dark observations -- a stronger-observation null, not a matched visual baseline.
Chinese Translation
我们研究暗操作:在短暂的亮光写入将 z0 = Enc(rgb) 编码后,一个策略 pi(z) 和开环动力学 f(z,a) 无需进一步的像素即可完成接触丰富的技能(dark_f)。在 ManiSkill StackCube(n=160;种子包 0/1000)上,dark_f 在五级链(near_A -> grasped -> lifted -> on_B -> stacked)上达到 68.1% 的堆叠成功率,而逐步 lit_reenc 为 35.6%,freeze/encode_black 为 0%。在共享的 Write->HOLD 协议(n=40)上,遮挡和相机对齐的 GT 接触邻域掩码将 lit lift 从 43% 降至 0%,而 dark_f 保持 82.5%;打乱动作会使动力学 MSE 膨胀约 9.4 倍;写入时的外观变化会破坏编码(夜间:0% 堆叠),但在 HOLD 期间的相同变化使 dark_f lift 保持不变;一旦到达停止阶段,写入长度 Tw 是平坦的,而较早停止和写入时模糊/JPEG 会急剧降低堆叠成功率。专用的 pi_write 达到 35% 的视觉预算堆叠(n=80);没有特权几何信息的匹配 Dreamer 风格/像素零模型保持 0%。特权状态-RSSM MPC 在 9D 暗观测下达到约 35% 的堆叠——这是一个更强观测的零模型,而非匹配的视觉基线。
cs.RO / 8 / 2609.13270

Conflict-Predictive Variable Horizons in Multi-Drone Distributed Model Predictive Control

多无人机分布式模型预测控制中的冲突预测可变时域
Mümken, Linda, Schwung, Michael, Lier, Stefan, Schwung, Andreas
Abstract
In distributed model predictive control for multi-drone collision avoidance, a fixed prediction horizon forces a compromise: a short horizon is inexpensive but reacts late to approaching neighbors, whereas a long one anticipates conflicts at a per-step cost that grows superlinearly with its length. We propose a conflict-predictive variable horizon that each drone sets locally, leaving the distributed model predictive control itself unchanged. From a short history of observed positions, a drone extrapolates the flight lines of its neighbors, tests each against its own using confidence funnels that narrow with prediction range, and obtains each time to conflict in closed form. The horizon is then the smallest admissible value whose planning window covers the farthest predicted conflict. It collapses to its minimum in clear airspace and grows only when a conflict lies ahead. Provided this minimum meets a single computable feasibility bound, we prove that recursive feasibility and asymptotic stability are preserved for every horizon the policy can select. These guarantees hold for a linear model, and a cascaded inner loop reduces each quadrotor's translational dynamics to a perturbed double integrator, so they carry over to the linearized quadrotor model and, as practical stability, to the full nonlinear one. In simulation on dense antipodal-swap benchmarks, the variable horizon reduces both per-step solver cost and total computation well below those of a long fixed horizon, and it maintains separation in every run, which a short fixed horizon of comparable per-step cost does not.
Chinese Translation
在多无人机避撞的分布式模型预测控制中,固定预测时域迫使做出折衷:短时域成本低,但对接近的邻居反应滞后;而长时域能预测冲突,但其每步成本随长度超线性增长。我们提出一种冲突预测可变时域,由每架无人机局部设定,保持分布式模型预测控制本身不变。根据观测位置的短历史,无人机外推其邻居的飞行路线,使用随预测范围缩小的置信漏斗将每个邻居的飞行路线与自己的进行对比测试,并以闭合形式获得每个冲突时间。然后,时域是规划窗口覆盖最远预测冲突的最小允许值。在空旷空域中它收缩到最小值,仅当存在冲突时才增长。只要该最小值满足单个可计算的可行性界,我们证明对于策略能够选择的每个时域,递归可行性和渐近稳定性都得以保持。这些保证对线性模型成立,并且级联内环将每个四旋翼的平移动力学简化为一个受扰双积分器,因此它们可以推广到线性化四旋翼模型,并且作为实际稳定性,推广到完整的非线性模型。在密集对跖交换基准的仿真中,可变时域将每步求解器成本和总计算量都降低到远低于长固定时域的水平,并且在每次运行中都保持间隔,而具有相当每步成本的短固定时域则不能做到这一点。
cs.RO / 9 / 2609.13289

Extending the Speed Limit of Quadrupedal Locomotion via Refined Actuator Modeling and Adaptive Command Scheduling

通过精细化执行器建模与自适应指令调度拓展四足运动的速度极限
Tao, Yucheng, Cheng, Shaowen, Lan, Guorong, Yuan, Yanyan, Jin, Yongbin, Wang, Hongtao
Abstract
Achieving high-speed locomotion in quadrupedal robots remains highly challenging, as actuators operate near their physical limits and exhibit pronounced nonlinearities. However, many existing methods neglect actuator nonlinearities and physical constraints during training, leading to a significant sim-to-real gap under highly dynamic motions and limiting achievable performance. To address this issue, we propose a high-speed locomotion framework that reduces sim-to-real discrepancies and stabilizes learning over a wide command distribution. A refined actuator model explicitly captures high-speed voltage coupling and magnetic saturation, enabling a more accurate representation of the torque-speed envelope. In addition, a reinforcement learning framework incorporating a two-stage curriculum and adaptive command scheduling (ACS) ensures stable training. Experiments on the 36.5 kg quadruped BlackPanther2 (BP2) demonstrate speeds of up to 13.2 m/s on a treadmill and 11.65 m/s outdoors, establishing a new state-of-the-art and, to the best of our knowledge, a world record for quadrupedal robot locomotion. The results further highlight the importance of accurate actuator modeling in preventing non-physical policy exploitation, and show that ACS improves robustness without sacrificing performance.
Chinese Translation
在四足机器人中实现高速运动仍然极具挑战,因为执行器在其物理极限附近运行并表现出显著的非线性。然而,许多现有方法在训练中忽略了执行器的非线性和物理约束,导致在高动态运动下出现显著的仿真到现实(sim-to-real)差距,并限制了可达到的性能。为了解决这个问题,我们提出了一个高速运动框架,该框架减少了仿真到现实的差异,并在广泛的指令分布上稳定学习。一个精细化的执行器模型显式地捕捉了高速电压耦合和磁饱和,从而能够更准确地表示转矩-速度包络。此外,一个结合了两阶段课程学习和自适应指令调度(ACS)的强化学习框架确保了稳定的训练。在36.5公斤的四足机器人BlackPanther2(BP2)上的实验表明,在跑步机上速度可达13.2 m/s,户外可达11.65 m/s,创造了新的最先进水平,并且据我们所知,创造了四足机器人运动的世界纪录。结果进一步强调了准确执行器建模在防止非物理策略利用方面的重要性,并表明ACS在不牺牲性能的情况下提高了鲁棒性。
cs.RO / 10 / 2609.13290

Breaking speed scaling in quadrupedal robots via Huygens' coupled-pendulum dynamics

通过惠更斯耦合摆动力学打破四足机器人的速度标度限制
Tao, Yucheng, Jin, Yongbin, Cheng, Shaowen, Liu, Xianwei, Yuan, Yanyan, Liang, Yanhong, Su, Chengkai, Fu, Chaojie, Lan, Guorong, Yang, Wei, Wang, Hongtao
Abstract
Achieving biological-level running speeds has largely been pursued through advances in control algorithms, which improve the utilization of existing hardware. However, the ultimate speed limits remain governed by the underlying force and torque requirements of rapid locomotion, which are typically addressed through increased actuator capacity. Inspired by Huygens' coupled pendulums, we demonstrate that superior locomotion can emerge from principled exploitation of intrinsic dynamics rather than brute-force hardware scaling. Inter-limb inertial coupling redistributes energy across the gait cycle and reduces peak joint torque required for rapid periodic motion, thereby expanding the achievable speed without proportional increases in actuator capability. Incorporating hardware parameters as additional design variables further extends this analysis into a co-optimization framework, enabling the systematic utilization of inertial coupling in robot design. Guided by this framework, a quadruped robot achieves a running speed of 10.74 m/s (Froude number 21.4) and completes a 100-meter sprint in 12.2 seconds, representing the first legged robot to surpass 10 m/s. These results establish inertial coupling as an underlying mechanism governing high-speed legged locomotion and highlight its role in reducing force requirements, offering new insights into the design of agile robotic systems.
Chinese Translation
实现生物级别的奔跑速度主要通过对控制算法的改进,从而提高现有硬件的利用率。然而,最终速度极限仍受快速运动所需力和力矩要求的制约,这些要求通常通过增加执行器能力来解决。受惠更斯耦合摆的启发,我们证明,优异的运动性能可以通过对内在动力学的原理性利用而实现,而非通过蛮力式的硬件缩放。肢间惯性耦合在步态周期内重新分配能量,并降低快速周期性运动所需的峰值关节扭矩,从而在不按比例增加执行器能力的情况下扩大可达速度。将硬件参数作为额外设计变量纳入,进一步将这一分析扩展为协同优化框架,使得在机器人设计中能够系统性地利用惯性耦合。在该框架指导下,一台四足机器人实现了10.74 m/s的奔跑速度(弗劳德数21.4),并在12.2秒内完成100米冲刺,成为首个速度超过10 m/s的腿式机器人。这些结果确立了惯性耦合作为支配高速腿式运动的潜在机制,并凸显了其在降低力需求方面的作用,为敏捷机器人系统的设计提供了新见解。
cs.RO / 11 / 2609.13295

Ergodic Control and Controlled Diffusion for Robot Learning: Review and Tutorial

面向机器人学习的遍历控制与受控扩散:综述与教程
Sun, Max Muchen, Bilaloglu, Cem, Rao, Ananya, Ivic, Stefan, Sartoretti, Guillaume, Fitzsimons, Kathleen, Abraham, Ian, Calinon, Sylvain, Murphey, Todd
Abstract
Diffusion learning leverages the statistical mechanism of diffusion processes for learning, reasoning, and inferring complex distributions from data. Recent advances in diffusion learning have been transformative, with robot learning emerging as a key opportunity area, with applications spanning perception, control, and decision-making. At the same time, the statistical mechanism of diffusion processes can be controlled to shape the temporal evolution of the state distribution underlying robot trajectories, inducing ergodic behavior in robotic systems. The frameworks of controlled diffusion and ergodic control were developed around the same time as diffusion learning, and their theories and algorithms have increasingly converged. Ergodicity induced by controlled diffusion has several significant implications for robot learning, distinct from applying diffusion learning to robotics problems: it formally enforces statistical properties required to ensure optimality of robot learning, enables non-myopic search over uncertain information landscapes for data collection, and enables behavior specification based on spatial rather than temporal characteristics of trajectories. This survey introduces the intuition behind controlled diffusion for robot learning, explores its connection to diffusion learning, presents theoretical foundations and numerical tutorials for solving controlled diffusion and ergodic control problems, and reviews applications across robotics. Finally, we discuss key challenges and future opportunities in leveraging controlled diffusion for robot learning.
Chinese Translation
扩散学习利用扩散过程的统计机制,从数据中学习、推理和推断复杂分布。扩散学习的最新进展具有变革性,机器人学习正成为一个关键的机会领域,其应用涵盖感知、控制和决策。与此同时,扩散过程的统计机制可以被控制以塑造机器人轨迹背后的状态分布的时间演化,从而在机器人系统中诱导遍历行为。受控扩散和遍历控制的框架与扩散学习大约同时发展起来,它们的理论和算法日益融合。由受控扩散诱导的遍历性对机器人学习具有若干重要意义,这不同于将扩散学习应用于机器人问题:它正式地强制要求确保机器人学习最优性所需的统计属性,使非短视搜索能够在不确定的信息景观上进行数据收集,并使得基于轨迹的空间而非时间特性来指定行为成为可能。本综述介绍了面向机器人学习的受控扩散背后的直觉,探讨了其与扩散学习的联系,提出了解决受控扩散和遍历控制问题的理论基础和数值教程,并回顾了机器人领域的应用。最后,我们讨论了利用受控扩散进行机器人学习的关键挑战和未来机遇。
cs.RO / 12 / 2609.13307

IMM-based Multiple Object Tracking using a State Prediction Neural Network

基于IMM与状态预测神经网络的多目标跟踪
Lim, Chan-Bin, Paek, Dong-Hee, Kong, Seung-Hyun
Abstract
Object tracking is essential for autonomous vehicles to avoid obstacles and plan routes. Radar maintains detection performance even in adverse weather and can measure relative velocity through the Doppler effect, making it well suited for object tracking. In this paper, we propose a data-driven state PRedictor-based Interacting Multiple Model tracking method (PR-IMM) that improves nonlinear object-motion representation while preserving the stability and interpretability of physics-based motion models. The proposed method employs a transformer-based PRediction model (PR) that incorporates radar Doppler measurements to predict object displacement. The PR model is integrated into the IMM as a mode alongside the CV, CA, and CT motion models, and their prior positions are dynamically combined according to the mode probabilities. Experimental results show that PR-IMM reduces position-estimation error by 57.3% over the IMM and by 16.5% over the PR, while reducing ID switches by 25.3% and improving IDF1 by 9.6% over the IMM.
Chinese Translation
目标跟踪对于自动驾驶汽车避障和路径规划至关重要。雷达即使在恶劣天气下也能保持检测性能,并且可以通过多普勒效应测量相对速度,因此非常适合目标跟踪。本文提出了一种数据驱动的、基于状态预测器的交互式多模型跟踪方法(PR-IMM),该方法在保持基于物理的运动模型的稳定性和可解释性的同时,提升了对非线性目标运动的表征能力。所提方法采用基于Transformer的预测模型(PR),并融合雷达多普勒测量来预测目标位移。PR模型作为与CV、CA和CT运动模型并列的一种模式被集成到IMM中,并根据模式概率动态组合它们的先验位置。实验结果表明,PR-IMM相较于IMM将位置估计误差降低了57.3%,相较于PR降低了16.5%,同时相较于IMM将ID切换降低了25.3%,并将IDF1提高了9.6%。
cs.RO / 13 / 2609.13318

Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning

Attention-DP3:通过几何对齐注意力条件实现空间对象感知的3D扩散策略
Yan, Changbo, Zhang, Zhongbo, Zhang, Zaibin, Wang, Yifan, Wang, Lijun, Lu, Huchuan
Abstract
3D point-cloud observations are inherently ambiguous in complex, cluttered manipulation scenes, where target objects may be partially occluded or tightly intermingled with visually similar distractors. As a result, standard 3D diffusion policies often struggle to localize and exploit task-relevant geometry as scene complexity grows. We propose \textbf{Attention-DP3}, a spatially object-aware 3D diffusion policy that injects object-level geometric cues via attention while keeping the DP3 diffusion backbone unchanged. Our pipeline performs open-vocabulary 2D segmentation on RGB images, then lifts predicted target masks into 3D using calibrated camera geometry to obtain object-centric geometric priors. We incorporate these cues through Tri-field Attentional Conditioning, which constructs three complementary fields: (i) a targetness field to anchor the target object, (ii) an intra-target saliency field to emphasize task-relevant geometry within the target, and (iii) a backgroundness field to suppress distractors and clutter. Experiments on Adroit, DexArt, MetaWorld, and the real-world SO101 platform show consistent improvements over DP3, achieving state-of-the-art performance across benchmarks. Notably, as distractor objects increase, DP3 drops sharply, whereas Attention-DP3 remains stable and outperforms DP3 by up to 31\% under heavy clutter. The code is publicly available at https://github.com/zhangzhongbo2213/Attention-DP3.
Chinese Translation
三维点云观测在复杂、杂乱的操纵场景中本质上是模糊的,其中目标物体可能被部分遮挡或与视觉相似的干扰物紧密混杂。因此,随着场景复杂度的增加,标准3D扩散策略往往难以定位和利用任务相关的几何信息。我们提出 Attention-DP3,一种空间对象感知的3D扩散策略,通过注意力注入对象级几何线索,同时保持 DP3 扩散骨干不变。我们的流程在RGB图像上执行开放词汇2D分割,然后利用标定的相机几何将预测的目标掩码提升到3D,以获得以对象为中心的几何先验。我们通过三场注意力条件(Tri-field Attentional Conditioning)整合这些线索,构造三个互补场:(i) 目标性场以锚定目标对象;(ii) 目标内显著性场以强调目标内任务相关几何;(iii) 背景性场以抑制干扰物和杂乱。在 Adroit、DexArt、MetaWorld 和真实世界 SO101 平台上的实验表明,相比 DP3 有持续改进,并在基准测试中达到最先进性能。值得注意的是,随着干扰物增加,DP3 急剧下降,而 Attention-DP3 保持稳定,在严重杂乱下比 DP3 高出多达31%。代码公开在 https://github.com/zhangzhongbo2213/Attention-DP3。
cs.RO / 14 / 2609.13335

Bridging Thought and Action: Taming Long-Horizon Instability in Open-Source LLM Agents with a MetaTool-Enhanced ROS Framework

连接思维与行动:利用 MetaTool 增强的 ROS 框架驯服开源 LLM 智能体中的长时程不稳定性
Mahmud, Kazi Abrar, Dhurubo, Nilotpaul Kundu, Kirttonia, Tamal, Ujjal, Sabbir Hossain, Haque, Mohammad Ariful
Abstract
Large Language Models (LLMs) have enabled more natural human-robot interaction, but open-source models often exhibit unstable long-horizon reasoning and inefficient action execution when deployed in agentic robotic frameworks. This paper presents an enhanced ROS-Agent based architecture that improves task reliability and execution efficiency for agentic robotic systems using open-source LLMs. The proposed system introduces a novel intermediate mechanism, termed the MetaTool, which enforces structured planning prior to action execution. Given a natural-language command, the MetaTool induces the LLM to generate a pseudo-code plan of intended tool invocations, which is stored in the ROS-Agent's scratchpad and persists throughout execution. By explicitly separating planning from execution, the proposed approach reduces execution loops and improves deterministic behavior. The architecture is validated on a custom mobile robotic platform with multimodal perception and motion control capabilities. Experimental results on real-world interactive tasks demonstrate improved task completion and contextual consistency, with up to ~24% gains on complex tasks compared to the baseline framework.
Chinese Translation
大型语言模型(LLMs)使得更自然的人机交互成为可能,但开源模型在部署于智能体机器人框架中时,往往表现出不稳定的长时程推理和低效的动作执行。本文提出了一种增强的基于 ROS-Agent 的架构,该架构利用开源 LLM 提高了智能体机器人系统的任务可靠性和执行效率。所提出的系统引入了一种新颖的中间机制,称为 MetaTool,该机制在动作执行之前强制进行结构化规划。给定自然语言命令,MetaTool 引导 LLM 生成预期工具调用的伪代码计划,该计划存储在 ROS-Agent 的暂存区中,并在整个执行过程中持续存在。通过明确地将规划与执行分离,所提出的方法减少了执行循环并提高了确定性行为。该架构在具有多模态感知和运动控制能力的定制移动机器人平台上进行了验证。在真实世界交互任务上的实验结果表明,与基线框架相比,任务完成度和上下文一致性有所提高,在复杂任务上增益高达约 24%。
cs.RO / 15 / 2609.13428

Chance-Constrained Belief-Space Maneuver Planning for Autonomous Collision Avoidance Under Uncertainty

不确定性下自主避碰的机会约束信念空间机动规划
Kim, Grace Ra, Eddy, Duncan, Kochenderfer, Mykel J.
Abstract
Increasing conjunction frequency in low Earth orbit places growing pressure on spacecraft operators to determine not only whether an encounter requires mitigation, but whether sufficient information is available to commit to a maneuver. This work formulates this information-action tradeoff as a belief-space planning problem for conjunctions between a maneuverable spacecraft and an unmaneuverable secondary object. The planner represents the uncertain orbital states as Gaussian beliefs and uses a chance-constrained belief-space Monte Carlo tree search framework to reason over possible future tracking updates before time of closest approach (TCA). A terminal chance constraint limits the probability of reaching TCA above a prescribed collision-risk threshold, allowing the planner to wait for informative tracking while intervening when deferral becomes too risky. We evaluate the approach on eight historical conjunctions from NASA's Conjunction Assessment Risk Analysis dataset. By varying the secondary-object measurement quality and tracking cadence, we generate a total of 96 distinct evaluation scenarios. Across the evaluated conditions, the planner reaches TCA without maneuvering in approximately 40% of episodes while maintaining no terminal collision-risk violations. In contrast, fixed-time rule-based maneuver policies resolve more encounters without maneuvering when intervention is deferred closer to TCA, but at the expense of increasing terminal risk violations. The fraction of episodes reaching TCA without maneuvering depends strongly on tracking quality and measurement cadence, ranging from 76% under accurate, frequent measurements to approximately 18%-20% under the poorest tracking conditions. These results show that tracking quality and frequency are not only inputs to collision-risk estimation: they can determine when intervention becomes necessary.
Chinese Translation
低地球轨道中不断增加的接近事件频率给航天器操作者带来越来越大的压力,要求他们不仅判断一次交会是否需要采取缓解措施,还要判断是否有足够信息来实施机动。本工作将这种信息-行动权衡形式化为可机动航天器与不可机动次级目标之间交会事件的信念空间规划问题。规划器将不确定的轨道状态表示为高斯信念,并使用机会约束信念空间蒙特卡洛树搜索框架,在最近接近时间(TCA)之前对可能的未来跟踪更新进行推理。终端机会约束将到达 TCA 时超过规定碰撞风险阈值的概率限制在一定范围内,使规划器能够等待有信息量的跟踪,同时在推迟变得风险过高时进行干预。我们在 NASA 交会评估风险分析数据集中的八次历史交会事件上评估该方法。通过改变次级目标的测量质量和跟踪节奏,我们共生成 96 个不同的评估场景。在所有评估条件下,规划器在约 40% 的回合中不进行机动即到达 TCA,同时没有终端碰撞风险违规。相比之下,固定时间基于规则的机动策略在将干预推迟到更接近 TCA 时,能够在不机动的情况下化解更多交会事件,但代价是终端风险违规增加。不机动到达 TCA 的回合比例强烈依赖于跟踪质量和测量节奏,在准确且频繁的测量下为 76%,在最差跟踪条件下约为 18%-20%。这些结果表明,跟踪质量和频率不仅是碰撞风险评估的输入:它们还可以决定何时必须进行干预。
cs.RO / 16 / 2609.13458

STAGE: Diagnosing Semantic Transfer at Grounded Execution in Embodied Agents

STAGE:诊断具身智能体接地执行中的语义迁移
Jin, Baosheng, Liang, Yushen, Shen, Hua
Abstract
Embodied language grounding requires more than identifying the referent of an instruction: recovered semantics must also control the action an agent exposes. We study this missing link as a semantic-action gap, where instruction semantics are recoverable but weakly expressed in native continuous actions. We introduce SAT-Bench, a fixed-observation counterfactual benchmark that holds the visual scene and agent state fixed while changing only instruction semantics. On LIBERO target-name and pixel-grounded relation swaps, target recovery reaches 100.0% and 95.8%, whereas OpenVLA action sensitivity remains only 6.8% and 7.7%. The gap persists across 1,000 additional compositional and temporal/procedural counterfactuals, with overall action sensitivity of 6.1%. Hidden-state, threshold-free, cross-policy, and rollout diagnostics further support this semantic-action transfer failure. We introduce VISA, a lightweight execution-time interface that converts recovered semantics into ALLOW, DEFER, target-consistency, and verified-selection decisions. VISA reduces invalid-instruction blind execution from 92.7% to 2.8% while preserving 94.0% of normal commands, and verified selection further improves target-consistent action exposure without updating the underlying policy. Overall, embodied language evaluation should measure semantic-action transfer, not semantic parsing alone.
Chinese Translation
具身语言接地不仅需要识别指令的指称对象:恢复的语义还必须控制智能体所输出的动作。我们将这一缺失环节视为语义-动作鸿沟,即指令语义可恢复但在原生连续动作中表达微弱。我们引入了SAT-Bench,一个固定观察的反事实基准,它保持视觉场景和智能体状态固定,仅改变指令语义。在LIBERO目标名称和基于像素的关系交换上,目标恢复率达到100.0%和95.8%,而OpenVLA的动作敏感度仅为6.8%和7.7%。这一鸿沟在1,000个额外的组合和时间/过程反事实中持续存在,整体动作敏感度为6.1%。隐藏状态、无阈值、跨策略和推演诊断进一步支持这种语义-动作迁移失败。我们引入了VISA,一个轻量级执行时接口,将恢复的语义转换为ALLOW、DEFER、目标一致性和验证选择决策。VISA将无效指令的盲目执行从92.7%降低到2.8%,同时保留94.0%的正常指令,并且验证选择进一步改善了目标一致的动作输出,而无需更新底层策略。总体而言,具身语言评估应衡量语义-动作迁移,而不仅仅是语义解析。
cs.RO / 17 / 2609.13516

Constraint-Grounded Reinforcement Learning for Variable Impedance Control in Contact-Rich Robotic Insertion

接触丰富机器人插入中变阻抗控制的约束锚定强化学习
He, Lin, Deng, Min
Abstract
In robotic insertion under uncertain contact, the axial force limit and the appropriate controller gain vary across tasks. As a result, a single fixed gain is unlikely to remain suitable across different task conditions, making conventional impedance controllers reliant on manual retuning. To eliminate manual retuning, we propose Constraint-Grounded Reinforcement Learning (CG-RL), a variable impedance framework for online gain adaptation. Conditioned on the force limit and contact feedback, the policy outputs a residual motion, an insertion rate, and a requested gain. The controller projects this gain into the admissible range without exposing the range itself to the policy. This separation allows a single policy to operate under different force limits without retraining or manual retuning. We evaluate CG-RL on simulated oblique insertion across five training seeds. CG-RL achieves an $85.8\pm7.7\%$ (mean $\pm$ SD) success rate of insertions without violating the force limit, while keeping the applied gain within the admissible range. As a comparison, a fixed-gain baseline using the midpoint gain achieves a success rate of $50.1\%$. The policy adapts its insertion rate continuously to the specified force limit and further generalizes to more permissive force limits above the training range. In contrast, the same actor without force-limit input does not exhibit this adaptation. The applied gain is guaranteed to remain within the admissible range, while force-limit satisfaction is validated empirically rather than guaranteed formally.
Chinese Translation
在不确定接触下的机器人插入中,轴向力限制和适当的控制器增益因任务而异。因此,单一固定增益不太可能在不同任务条件下保持适用,使得传统阻抗控制器依赖手动重新调整。为了消除手动重新调整,我们提出了约束锚定强化学习(Constraint-Grounded Reinforcement Learning, CG-RL),一种用于在线增益自适应的变阻抗框架。以力限制和接触反馈为条件,策略输出残差运动、插入速率和请求增益。控制器将该增益投影到允许范围内,而不将范围本身暴露给策略。这种分离允许单一策略在不同力限制下运行,无需重新训练或手动重新调整。我们在五个训练种子上评估CG-RL在模拟倾斜插入中的表现。CG-RL在不违反力限制的情况下实现了85.8±7.7%(均值±标准差)的插入成功率,同时将施加的增益保持在允许范围内。作为比较,使用中点增益的固定增益基线达到了50.1%的成功率。该策略根据指定的力限制连续调整其插入速率,并进一步泛化到训练范围以上的更宽容的力限制。相比之下,没有力限制输入的相同actor不表现出这种自适应。施加的增益保证保持在允许范围内,而力限制的满足是通过经验验证而非形式化保证。
cs.RO / 18 / 2609.13545

Runtime-Incremental Transformer for Reinforcement-Learning-Based Adaptive Control

面向基于强化学习自适应控制的运行时增量Transformer
Cirrincione, Giansalvo, Fagiolini, Adriano
Abstract
Learning-based adaptive control of robotic manipulators with non-observable friction memory has been addressed by attention- based meta-controllers whose number of attention heads is fixed before training and is tuned by costly offline search. At long memory horizons, such fixed-capacity controllers are prone to catastrophic failures on a sizeable fraction of training seeds. The present paper introduces a runtime mechanism that grows and prunes the heads of the attention block during reinforcement learning, governed by two signals: the effective rank of the on-policy context distribution, which triggers growth when representational capacity becomes insufficient, and the per-head output magnitude, which flags redundant heads for removal. Policy continuity at growth events and a quantitative bound at prune events are established analytically. On a two- link manipulator with Stribeck friction, the proposed mechanism attains full success across all memory regimes, eliminating the long-horizon failure mode and removing the need for offline tuning of the head count.
Chinese Translation
针对具有不可观测摩擦记忆的机器人机械臂,基于学习的自适应控制已通过基于注意力的元控制器得到解决,这些控制器的注意力头数量在训练前固定,并通过昂贵的离线搜索进行调优。在长记忆时域下,这种固定容量的控制器在相当大一部分训练随机种子上容易出现灾难性失败。本文提出了一种运行时机制,在强化学习过程中对注意力块的头进行增长和剪枝,该机制由两个信号控制:同策略上下文分布的有效秩,当表示容量不足时触发增长;以及每个头的输出幅度,用于标记冗余头以移除。增长事件处的策略连续性和剪枝事件处的定量界限通过解析方式建立。在具有Stribeck摩擦的两连杆机械臂上,所提机制在所有记忆模式下均取得完全成功,消除了长时域失败模式,并消除了对头数量进行离线调优的需要。
cs.RO / 19 / 2609.13572

An MRI-Guided Robotic System to Improve Hippocampal Access for Epilepsy Interventions

一种用于改善癫痫干预中海马体入路的MRI引导机器人系统
Peters, John E., Grillo, Abby M., Mansouri, Mahshid, Esser, Daniel S., Garrow, Sarah, Kumar, Nithin S., Englot, Dario J., Neimat, Joseph S., Grissom, William A., Webster III, Robert J., Barth, Eric J.
Abstract
This paper presents an MRI-guided robotic system that improves hippocampal access by delivering a curved needle-like laser ablator through the foramen ovale, a natural opening in the base of the skull. Both the delivery path and the curved nature of the needle improve upon current clinical straight-line laser interstitial thermal therapy (LITT) as measured by hippocampal cannulation percentage. We describe the design of the robotic system which includes positioning, aiming, and curved needle deployment stages, followed by experimental results assessing cannulation percentages using MRI images in phantoms. In three curvilinear, single-insertion experiments, we achieved cannulation percentages of 93.5%, 96.5%, and 60.4%, exceeding reported clinical averages of 50-60%. Seizure control is believed by physicians to be a function of hippocampal volume treated, and our system provides a means of treating a greater volume of the hippocampus with LITT-based interventions.
Chinese Translation
本文提出了一种MRI引导的机器人系统,通过将弯曲的针状激光消融器经卵圆孔(foramen ovale)——颅底的自然开口——输送,来改善海马体入路。以海马体插管百分比衡量,输送路径和针的弯曲特性均优于当前临床直线激光间质热疗(LITT)。我们描述了该机器人系统的设计,包括定位、瞄准和弯曲针部署阶段,随后是在体模中使用MRI图像评估插管百分比的实验结果。在三次曲线单次插入实验中,我们分别实现了93.5%、96.5%和60.4%的插管百分比,超过了报道的50-60%的临床平均值。医生认为癫痫控制与所治疗的海马体体积有关,我们的系统提供了一种通过基于LITT的干预治疗更大海马体体积的方法。
cs.RO / 20 / 2609.13588

Differential Realizability of Static Control Allocation in Multirotors: An Impossibility under Nonredundant Full Actuation and a Pseudoinverse Obstruction under Redundant Actuation

多旋翼静态控制分配的微分可实现性:非冗余全驱动下的不可能性与冗余驱动下的伪逆障碍
Franchi, Antonio, Mizzoni, Mirko
Abstract
Control allocation for multirotors with bidirectional propellers is commonly formulated in signed-thrust variables, where the wrench map is linear. The signed-quadratic map from physical rotor speed to thrust, however, is not a local diffeomorphism at zero speed. This work derives two distinct consequences. Under nonredundant full actuation, a single-propeller reversal removes one instantaneous task direction; hence, no global continuously differentiable exact static allocator exists over the complete task space. Under redundant actuation, the physical task map may remain regular, yet a transverse pseudoinverse zero crossing requires an unbounded rotor-speed derivative. We define differential realizability as regularity of the physical lift of an actuator-output section, derive exact and first-order validity conditions, and distinguish structural rank loss from an allocator- induced rate singularity. A local nullspace deformation repairs isolated pseudoinverse reversals, while a global fixed-orthant construction establishes existence of regular sections at the cost of persistent task-preserving internal actuation.
Chinese Translation
对于采用双向螺旋桨的多旋翼飞行器,其控制分配通常在有符号推力变量中表述,此时力旋量映射是线性的。然而,从物理转子转速到推力的有符号二次映射在零转速处不是局部微分同胚。本文推导出两个不同的结果。在非冗余全驱动下,单个螺旋桨的反向会消除一个瞬时任务方向;因此,在整个任务空间上不存在全局连续可微的精确静态分配器。在冗余驱动下,物理任务映射可能保持正则,但横向伪逆零穿越需要无界的转子转速导数。我们将微分可实现性定义为执行器输出截面的物理提升的正则性,推导了精确和一阶有效性条件,并区分了结构性秩亏损与分配器引起的速率奇异性。局部零空间变形可以修复孤立的伪逆反向,而全局固定象限构造则以持续的任务保持内部驱动为代价,建立了正则截面的存在性。
cs.RO / 21 / 2609.13606

From Vision to Harvest: Benchmarking Vision-Language Models for Multi-Arm Robotic Fruit Harvesting

从视觉到收获:多臂机器人水果采摘的视觉语言模型基准测试
Inukollu, Vrishan, Zaman, Adyan, Kudaraya, Anvi, Lazcano, Carlos, Zhu, Yuankai, Vougioukas, Stavros, Yu, Xiaofan
Abstract
Multi-arm robotic harvesting offers a promising path to improve harvesting efficiency and reduce reliance on manual labor. However, practical deployment remains challenging because the system must generalize across diverse environments while efficiently coordinating multiple arms in a shared workspace. Existing methods often require substantial data collection in target environments or rely on simplifying assumptions that limit planning quality. In this work, we introduce the first comprehensive benchmark for evaluating pretrained Vision-Language Models (VLMs) on zero-shot multi-arm fruit harvesting planning. Our benchmark uses real-world apple and citrus orchard images and compares a VLM-based planning pipeline with a traditional perception-and-planning pipeline. The VLM pipeline directly generates harvesting sequences and waypoints for each arm, while a lightweight trajectory verifier checks for collisions. Our results show that frontier VLMs can generate effective multi-arm harvesting plans zero-shot, but a practical deployment remains limited by accurate 3D waypoint generation and collision-aware coordination. These results highlight both the promise and current limitations of pretrained VLMs for multi-arm robotic harvesting.
Chinese Translation
多臂机器人采摘为提高采摘效率和减少对人工劳动的依赖提供了一条有前景的路径。然而,实际部署仍然具有挑战性,因为系统必须在不同环境之间泛化,同时在共享工作空间中高效协调多个机械臂。现有方法通常需要在目标环境中收集大量数据,或者依赖于限制规划质量的简化假设。在这项工作中,我们引入了第一个综合基准,用于评估预训练视觉语言模型(VLMs)在零样本多臂水果采摘规划上的表现。我们的基准使用真实世界的苹果和柑橘园图像,并将基于VLM的规划流程与传统的感知和规划流程进行比较。VLM流程直接为每个机械臂生成采摘序列和路点,而轻量级轨迹验证器检查碰撞。我们的结果表明,前沿VLMs可以零样本生成有效的多臂采摘计划,但实际部署仍然受限于准确的3D路点生成和碰撞感知协调。这些结果突出了预训练VLMs在多臂机器人采摘中的前景和当前局限性。
cs.RO / 22 / 2609.13627

How Well Do Pseudo-Rigid Body Models Capture Real Plants? In-Field Validation of Simulated Blueberry Canes

伪刚体模型能多好地捕捉真实植物?——模拟蓝莓茎秆的田间验证
Kolano, Hannah, VanAtter, Chelse, Yang, Wei, Grimm, Cindy, Davidson, Joseph R.
Abstract
Many labor-intensive tasks in fruit production such as pruning require physically interacting with the plant (e.g., pushing, pulling, bending limbs, etc.). Due to increasing labor shortages, there is widespread interest in the adoption of robotics in this area. When the robot must physically interact with the system, a deformable model of the plant is beneficial. Coupled, spring-loaded beams have been used in prior work to build deformable plant models in simulation for learning, planning, and control, but this approach has rarely been validated against the mechanical properties of real-world plants. In this paper, we use this rigid-body model, grounded in real material properties and beam bending theory, to simulate the deformation of blueberry canes under loading. To validate the model, we simulate the canes in MuJoCo and compare their behavior with real-world data collected from probing canes at a commercial farm with a custom testbed. We perform sensitivity analysis on several key modeling variables and show that this approach is highly sensitive to the plant diameter calculations and a priori flexural modulus estimation, which is dependent on season and blueberry variety. We also share our dataset of live blueberry cane deformation, including RGB-D images and measured forces and displacements.
Chinese Translation
水果生产中的许多劳动密集型任务(如修剪)需要与植物进行物理交互(例如,推、拉、弯曲枝条等)。由于劳动力短缺日益严重,在该领域采用机器人技术引起了广泛兴趣。当机器人必须与系统进行物理交互时,植物的可变形模型是有益的。在先前的工作中,耦合弹簧加载梁已被用于在仿真中构建可变形植物模型,以用于学习、规划和控制,但这种方法很少针对真实世界植物的力学特性进行验证。在本文中,我们使用这种基于真实材料属性和梁弯曲理论的刚体模型,来模拟蓝莓茎秆在负载下的变形。为了验证该模型,我们在MuJoCo中模拟茎秆,并将其行为与在商业农场使用定制测试平台探测茎秆收集的真实世界数据进行比较。我们对几个关键建模变量进行了敏感性分析,并表明该方法对植物直径计算和先验弯曲模量估计高度敏感,而后者取决于季节和蓝莓品种。我们还分享了我们的活体蓝莓茎秆变形数据集,包括RGB-D图像以及测量的力和位移。
cs.RO / 23 / 2609.13651

Compositional Shift Algebra: Extrapolating Mixed Robot Shifts Without Mixed Finetuning

组合偏移代数:无需混合微调即可外推混合机器人偏移
Hang, Jinting, Cai, Zhenhui
Abstract
Robot deployments rarely change one mechanism at a time: cameras, action interfaces, and physical dynamics often shift together. Prior adaptation recipes either finetune a new model for every mix or attempt to select which module to update. We instead learn shift operators on a modular stack z{=}E(o), a{=}g(z,u), z'{=}f(z,a) and compose them. Compositional Shift Algebra (CSA) fits single-factor observation, policy, and dynamics operators from exact-reset probes, then extrapolates held-out mixed shifts by operator composition---without mixed-shift finetuning. On ManiSkill StackCube, residual CSA matches an oracle mixed inverse on held-out mixes (success 1.0 over 10 seeds) while beating best-single / zero-shot / parameter-average baselines by { approx}67 pp. RGB-D vision-in-the-loop composition remains near oracle and far above non-compositional arms; a delay commutator stress shows ordered necessity for policy timesdelay. On a second task (PickCube), residual CSA again reaches compose 1.0 vs. 0.33 non-compositional (n{=}10), and an L1 vision controller without privileged cube/goal poses or grasp flags in the control loop retains compose 0.95 vs. 0.00. Main-track upgrades freeze PushCube (+33 pp), PegInsertion joint8 / pose7 EE (+67 pp each), and thin BC under frozen CSA (+67 pp); deeper BC and fair adapt baselines still need compose (+67 pp each), vision-localized BC needs compose (+56 pp), and delay favors ordered/few-shot deploy. We report Intervention-Gated Adaptation as a negative control.
Chinese Translation
机器人部署很少一次只改变一个机制:相机、动作接口和物理动力学经常同时发生变化。先前的适应方法要么为每种混合微调一个新模型,要么尝试选择要更新哪个模块。我们转而在模块化堆栈 z=E(o), a=g(z,u), z'=f(z,a) 上学习偏移算子,并将它们组合起来。组合偏移代数 (CSA) 从精确重置探针中拟合单因素观察、策略和动力学算子,然后通过算子组合外推留出的混合偏移——无需混合偏移微调。在 ManiSkill StackCube 上,残差 CSA 在留出的混合上与 oracle 混合逆相匹配(10 个种子成功率 1.0),同时比最佳单一/零样本/参数平均基线高出约 67 个百分点。RGB-D 视觉在环组合仍然接近 oracle,且远高于非组合分支;延迟交换子压力测试显示了策略×延迟的有序必要性。在第二个任务 (PickCube) 上,残差 CSA 再次达到组合成功率 1.0,而非组合为 0.33(n=10),并且一个 L1 视觉控制器在控制回路中没有特权立方体/目标位姿或抓取标志的情况下,保持组合成功率 0.95 对比 0.00。主赛道升级冻结 PushCube (+33 pp),PegInsertion joint8 / pose7 EE (+67 pp 每个),以及冻结 CSA 下的薄 BC (+67 pp);更深的 BC 和公平的自适应基线仍然需要组合(各 +67 pp),视觉定位的 BC 需要组合(+56 pp),而延迟有利于有序/少样本部署。我们将干预门控适应(Intervention-Gated Adaptation)报告为阴性对照。
cs.RO / 24 / 2609.13675

Does Online Gravity Estimation Matter? Revisiting a Silent Design Split in LiDAR-Inertial Odometry

在线重力估计重要吗?重新审视激光雷达惯性里程计中的一个隐性设计分歧
Xu, Jie, Jin, Ziyi, Yu, Kangjin, Jiang, Can, Huang, Hongjun, Jin, Tongxing, Luo, Hongkun, Xia, Zhongpu
Abstract
LiDAR-inertial odometry (LIO) systems differ in whether they continue estimating gravity after initialization. We compare four gravity-bias state configurations in each of FAST-LIO2 and LIO-SAM, then separately test a gravity-direction factor. Across 12 dataset sequences evaluated with FAST-LIO2, fixing gravity under continuous LiDAR correction produces mean paired changes in vertical and 3D position errors with 90% confidence intervals within $\pm 2\%$. Tests on 4 sequences with LIO-SAM likewise show no consistent benefit from online gravity. Multi-second LiDAR outages, unlike reduced range or field of view, reveal trajectory-dependent costs of fixing gravity. A history-matched 23D-to-21D switch places the repeatable 3D error increase after LiDAR updates resume. Under 5-s outages, a direction factor from the same IMU used for preintegration improves accuracy on Hall05 but worsens both errors with online gravity on TUHH. Dynamic-start tests also show fixed-bias failures at particular starting phases. We recommend keeping gravity and accelerometer bias online for robustness; use a direction factor only after verifying vertical and 3D accuracy gains under the intended operating conditions.
Chinese Translation
激光雷达惯性里程计(LIO)系统在初始化后是否继续估计重力方面存在差异。我们比较了FAST-LIO2和LIO-SAM中四种重力-偏置状态配置,然后单独测试了一个重力方向因子。在使用FAST-LIO2评估的12个数据集序列上,在连续激光雷达校正下固定重力产生的垂直和3D位置误差的平均配对变化,其90%置信区间在±2%以内。在LIO-SAM上对4个序列的测试同样表明在线重力没有一致的好处。与距离或视场减小不同,多秒激光雷达中断揭示了固定重力的轨迹相关代价。一个历史匹配的23D到21D切换将可重复的3D误差增加置于激光雷达更新恢复之后。在5秒中断下,来自用于预积分的同一IMU的方向因子在Hall05上提高了精度,但在TUHH上在线重力下使两个误差都恶化。动态起始测试还显示在特定起始阶段出现固定偏置失败。我们建议为了鲁棒性保持重力和加速度计偏置在线;仅在验证了预期操作条件下的垂直和3D精度提升后,才使用方向因子。
cs.RO / 25 / 2609.13679

How to Better Train VLAs: Lessons Learned From the REAL-I Challenge at ICRA 2026

如何更好地训练VLA:从ICRA 2026的REAL-I挑战赛中获得的经验教训
Wang, Jiaming, Chen, Jizhuo, Liu, Diwen, Song, Wang, Wang, Qiang, Ren, Jie, Fu, Chao, Zhu, Dingkun, Ruan, Minchi, Li, Hongtong, Jiang, Yuhua, Xue, Zhiwei, Pan, Yongping, Soh, Harold
Abstract
How can robot policies learn more effectively from a fixed demonstration budget? The first Real-world Embodied AI Learning (REAL-I) Challenge at ICRA 2026 examined this question through simulation, real-robot evaluation, and an on-site final on a shared dual-arm humanoid platform. We describe the challenge tasks, data and deployment interfaces, and competition results, then compare the approaches contributed by NUS-CLEAR, RCL-Lab, and Deeptouch.ai. Their systems combined pretrained vision-language-action models and task-specific imitation policies with different strategies for data curation, staged adaptation, checkpoint selection, and action-space design. The team reports highlight the importance of adapting to the deployment environment while retaining prior capabilities, treating demonstration quality at an appropriate temporal scale, and suppressing errors in inactive robot components. They also expose the limitations of offline action-prediction metrics for forecasting closed-loop success. These observations motivate a view of fixed-data robot learning that integrates data, adaptation, evaluation, and deployment.
Chinese Translation
机器人策略如何从固定的演示预算中更有效地学习?ICRA 2026的首届真实世界具身AI学习(REAL-I)挑战赛通过仿真、真实机器人评估以及在共享双臂人形平台上的现场决赛来研究这一问题。我们描述了挑战赛的任务、数据和部署接口以及竞赛结果,然后比较了NUS-CLEAR、RCL-Lab和Deeptouch.ai所贡献的方法。他们的系统将预训练的视觉-语言-动作模型和特定任务的模仿策略与不同的数据管理、分阶段适应、检查点选择和动作空间设计策略相结合。团队报告强调了在保留先前能力的同时适应部署环境、在适当的时间尺度上处理演示质量以及抑制非活动机器人组件中的错误的重要性。他们还揭示了离线动作预测指标在预测闭环成功方面的局限性。这些观察结果促使我们形成一种整合数据、适应、评估和部署的固定数据机器人学习视角。
cs.RO / 26 / 2609.13695

GROOVE: Geometry-Guided Reduction of Operational-Space Jerk in VLA Execution

GROOVE:几何引导的VLA执行中操作空间加加速度降低
Yun, Sangho, Kim, Minsoo, Cho, Minwoo, Yu, Hwanjo
Abstract
Chunked vision language action (VLA) policies execute several commands per query, but jerk within chunks and across replanning boundaries can induce oscillatory motion and sharp actuator transients. We present GROOVE, an online regulator that searches directional correction regions around the raw three dimensional end effector (EEF) path, without retraining or additional VLA inference. It optimizes the new chunk using delivered commands as boundary conditions, reducing boundary and within chunk jerk while bounding cumulative translation and local axis angle deviation from the raw plan after every command. Using quadratic programs (QPs), GROOVE generates a cube reference and thirteen directional candidates, then selects the one with the lowest command space jerk under a reference relative deviation cap. On a held out LIBERO benchmark, GROOVE achieves the largest reductions among the evaluated methods, reducing translational and rotational EEF jerk by 33.02% and 43.42%, respectively, with task success of 95.75% versus 93.75% for raw execution. Across 50 matched UR5e pairs with measured execution timing, it reduces translational and rotational tool center point (TCP) jerk by 16.39% and 19.49% and joint current slew by 29.09%.
Chinese Translation
分块视觉语言动作(VLA)策略每次查询执行多个命令,但块内以及重规划边界处的加加速度可能导致振荡运动和急剧的执行器瞬态。我们提出GROOVE,一种在线调节器,它在原始三维末端执行器(EEF)路径周围搜索方向性校正区域,无需重新训练或额外的VLA推理。它使用已传递的命令作为边界条件来优化新的块,减少边界和块内加加速度,同时限制每个命令后相对于原始计划的累积平移和局部轴角度偏差。使用二次规划(QP),GROOVE生成一个立方体参考和十三个方向候选,然后在参考相对偏差上限下选择命令空间加加速度最低的那个。在留出的LIBERO基准测试上,GROOVE在所评估的方法中实现了最大的降低,分别将平移和旋转EEF加加速度降低了33.02%和43.42%,任务成功率为95.75%,而原始执行的成功率为93.75%。在具有测量执行时间的50对匹配UR5e上,它将平移和旋转工具中心点(TCP)加加速度降低了16.39%和19.49%,并将关节电流变化率降低了29.09%。
cs.RO / 27 / 2609.13711

Decentralized Multi-Robot Task Allocation Under Degraded Communication: A Benchmark of Performance, Reliability, and Computation

通信退化下的分散式多机器人任务分配:性能、可靠性与计算基准
Lott, James, Honary, Vahraz
Abstract
Selecting a decentralized Multi-Robot Task Allocation (MRTA) method for embedded deployment on autonomous platforms requires considering more than route performance alone. We benchmark six decentralized MRTA allocators (CBAA, ACBBA, PI, HIPC, DMCHBA, and DGA) in the Collaborative Visit (CV) scenario to characterize tradeoffs among MinMax and MinSum travel, communication robustness and demand, allocation reliability, computational burden, and scale sensitivity. The core study uses 500 paired ten-target instances across 25 ideal and degraded communication conditions spanning Bernoulli loss, Gilbert--Elliott loss, and Rayleigh fading, with additional campaigns examining pre-allocation, execution-integrated computation, and sensitivity to grid size, robot density, and target load. Across the 24 impaired core conditions, DGA and DMCHBA achieved the lowest mean MinMax travel at 24.49 and 24.78 steps, respectively. HIPC narrowly led mean MinSum travel at 66.95 steps, followed by DGA at 67.22, with both methods occupying the top two in every impaired condition. DMCHBA had the lowest publication intensity at 2.08 publications per team step. In ten-target pre-allocation, HIPC and DMCHBA remained viable and stable in every tested condition, while ACBBA, PI, and DGA lost stability or viability as communication degraded. Under ideal delivery, median full-protocol computation $\Cterm$ in the primary ten-target comparison ranged from 4.88 ms for DMCHBA to 1.346 s for DGA. Static route quality preserved DGA and DMCHBA as the leading MinMax methods, while DGA led MinSum at three of four target loads and HIPC led at 50 targets. Static and execution-integrated computation rankings diverged as task load increased. The results identify distinct allocator operating regions across route objective, communication behavior, reliability, and computational constraints.
Chinese Translation
选择用于自主平台嵌入式部署的分散式多机器人任务分配(MRTA)方法,需要考虑的不仅仅是路径性能。我们在协作访问(CV)场景中对六种分散式 MRTA 分配器(CBAA、ACBBA、PI、HIPC、DMCHBA 和 DGA)进行基准测试,以刻画 MinMax 和 MinSum 行程、通信鲁棒性与需求、分配可靠性、计算负担以及规模敏感性之间的权衡。核心研究使用了 500 个配对的十目标实例,覆盖 25 种理想和退化通信条件,包括伯努利损失、Gilbert-Elliott 损失和瑞利衰落,并附加实验考察了预分配、执行集成计算以及对网格大小、机器人密度和目标负载的敏感性。在 24 种受损核心条件下,DGA 和 DMCHBA 分别实现了最低的平均 MinMax 行程,为 24.49 和 24.78 步。HIPC 以 66.95 步略微领先平均 MinSum 行程,其次是 DGA 的 67.22 步,这两种方法在每个受损条件下都占据前两名。DMCHBA 的发布强度最低,为每团队步 2.08 次发布。在十目标预分配中,HIPC 和 DMCHBA 在每个测试条件下都保持可行且稳定,而 ACBBA、PI 和 DGA 随着通信退化失去稳定性或可行性。在理想交付下,主要十目标比较中的中位完整协议计算 $\Cterm$ 从 DMCHBA 的 4.88 毫秒到 DGA 的 1.346 秒不等。静态路径质量保持了 DGA 和 DMCHBA 作为领先的 MinMax 方法,而 DGA 在四种目标负载中的三种中领先 MinSum,HIPC 在 50 个目标时领先。随着任务负载增加,静态和执行集成计算的排名出现分歧。结果确定了分配器在路径目标、通信行为、可靠性和计算约束方面的不同操作区域。
cs.RO / 28 / 2609.13761

Learning In-Hand Object Reaching to General 6D Poses

学习手内物体到达通用6D位姿
Lin, Junxiao, Wu, Tianyue, Yin, Jie, Pan, Jia, Zhang, Kaifeng, Zhi, Weiming
Abstract
In-hand manipulation allows multi-fingered dexterous hands to reconfigure grasped objects without releasing and regrasping them. This improves manipulation efficiency by reducing repeated grasp acquisition and large arm motions. However, most learning-based methods focus on reorientation, continuous rotation, or translation, whereas many tasks require joint control of object position and orientation. We formulate this capability as in-hand 6D object pose reaching: starting from an existing grasp, coordinated finger motions move the object to a palm-relative target pose. We present POISE (Palm-relative Object reaching In SE(3)), a sim-to-real reinforcement learning framework for this task. POISE combines diverse stable-grasp initialization, goal- and geometry-conditioned control, an adaptive 6D goal curriculum, and a compact reward scheme for pose reaching and grasp preservation. In simulation, diverse initialization raises held-out-grasp success from 40.1% to 51.5% and post-drop recovery from 33.8% to 72.9%; the curriculum raises full-range success from 6.2% to 59.5%. On hardware, the grasp-maintenance reward improves three-target sequence success from 20% to 80%. In real-world experiments, POISE reaches successive 6D targets without manual reset across multiple object geometries and wrist orientations, and recovers from external disturbances. To support further research in dexterous manipulation, we will release our code at https://junxiaolin.github.io/poise-website/.
Chinese Translation
手内操作使多指灵巧手能够在无需释放并重新抓取物体的情况下重新配置被抓取的物体。这通过减少重复抓取获取和大幅手臂运动来提高操作效率。然而,大多数基于学习的方法关注重新定向、连续旋转或平移,而许多任务需要联合控制物体的位置和姿态。我们将这一能力形式化为手内6D物体位姿到达:从现有抓取出发,协调的手指运动将物体移动到相对于手掌的目标位姿。我们提出了POISE(Palm-relative Object reaching In SE(3)),一个用于该任务的仿真到现实的强化学习框架。POISE结合了多样化的稳定抓取初始化、目标和几何条件控制、自适应6D目标课程学习,以及用于位姿到达和抓取保持的紧凑奖励方案。在仿真中,多样化初始化将留出抓取成功率从40.1%提高到51.5%,将掉落恢复率从33.8%提高到72.9%;课程学习将全范围成功率从6.2%提高到59.5%。在硬件上,抓取维持奖励将三目标序列成功率从20%提高到80%。在真实世界实验中,POISE能够在多种物体几何形状和手腕朝向下连续到达6D目标而无需手动重置,并且能够从外部扰动中恢复。为支持灵巧操作领域的进一步研究,我们将在https://junxiaolin.github.io/poise-website/发布我们的代码。
cs.RO / 29 / 2609.13777

When Do Learned Priors Help Visual Inertial Estimation? A Controlled Study of Prior Integration, Calibration, Initialization, and Backend Consistency

学习先验何时有助于视觉惯性估计?关于先验集成、标定、初始化和后端一致性的受控研究
Zhang, Jinchang, Lu, Guoyu
Abstract
Learned components are increasingly integrated into geometric visual--inertial estimators to provide motion, depth, bias, uncertainty, or confidence cues. Yet it remains unclear whether gains arise from useful learned priors or from changes in the backend, calibration, initialization, temporal association, or evaluation gauge. We present a controlled framework for learning-augmented visual--inertial estimation that separates fusion gain from the incremental value of a learned prior and evaluates four evidence layers: local motion consistency, global trajectory accuracy, physical-state correctness, and numerical consistency. We instantiate the framework with a MonoViT-based monocular motion prior added as a local relative-motion factor to an unchanged VINS backend. We compare Original VINS and learned-prior VINS under matched sensor streams, timestamps, initialization, frontend/backend settings, and camera--IMU extrinsics, while probing calibration, initialization, state coupling, scale, bundle adjustment, and loop closure. On KITTI, with fixed reference extrinsics, translation APE RMSE is 31.4 m for Original VINS and 31.8 m with the learned prior. Across four recordings, the prior changes mean APE by only -0.2%, while a five-times-higher weight worsens it by 8.2%. Online extrinsic updates increase mean APE by 45.7% and 52.1%, respectively, while mean RPE changes by less than 2%. These results show that fusion performance alone cannot establish the value of learned priors. Reliable evaluation requires same-backend controls and joint analysis of prior compatibility, calibration, initialization, global drift, physical-state error, and backend consistency.
Chinese Translation
学习组件正越来越多地被集成到几何视觉惯性估计器中,以提供运动、深度、偏差、不确定性或置信度线索。然而,目前尚不清楚增益是来自有用的学习先验,还是来自后端、标定、初始化、时间关联或评估标准的变化。我们提出了一个用于学习增强视觉惯性估计的受控框架,该框架将融合增益与学习先验的增量价值分离开来,并评估四个证据层:局部运动一致性、全局轨迹精度、物理状态正确性和数值一致性。我们使用基于MonoViT的单目运动先验作为局部相对运动因子添加到未更改的VINS后端,以此实例化该框架。我们在匹配的传感器流、时间戳、初始化、前端/后端设置和相机-IMU外参下,比较原始VINS和带学习先验的VINS,同时探究标定、初始化、状态耦合、尺度、光束法平差和回环检测。在KITTI上,使用固定参考外参,原始VINS的平移APE RMSE为31.4米,带学习先验的为31.8米。在四个记录序列上,先验仅使平均APE变化了-0.2%,而五倍高的权重则使其恶化了8.2%。在线外参更新分别使平均APE增加了45.7%和52.1%,而平均RPE的变化小于2%。这些结果表明,仅靠融合性能无法确立学习先验的价值。可靠的评估需要相同后端的对照,并联合分析先验兼容性、标定、初始化、全局漂移、物理状态误差和后端一致性。
cs.RO / 30 / 2609.13779

Force-Aware Reinforcement Learning with Hybrid Sensorless Force Estimation for Wheeled-Legged Loco-Manipulation

面向轮腿式移动操作的结合混合无传感器力估计的力感知强化学习
Zeng, Xuanqi, Wang, Jiaming, Zhang, Tianlin, Zhang, Lingwei, Xu, Botian, Xia, Weipeng, Li, Zhongyu, Liu, Yun-Hui
Abstract
Force-controlled loco-manipulation requires a whole-body policy to coordinate locomotion and arm motion while regulating end-effector interaction forces. This is challenging under floating-base dynamics and changing support contacts, particularly when end-effector force/torque sensing is unavailable for control. This paper presents a force-aware reinforcement learning approach with hybrid sensorless force estimation for wheeled-legged loco-manipulation. The proposed method provides a structured estimate of the end-effector force as an explicit policy observation, enabling force-guided contact behavior without using an end-effector force/torque sensor for control. The force estimate is obtained by combining generalized momentum observation, contact-constrained wrench projection, and temporal residual learning: the model-based components extract the physically structured part of the whole-body disturbance, while the residual network compensates the remaining motion-dependent bias. The estimated force is integrated into a mode-conditioned whole-body policy with an axis-wise force/position selector, allowing free-space motion, pure force regulation, and hybrid force/position control within one controller. Simulation results demonstrate improved sensorless force estimation and force-control performance. Hardware experiments further validate the proposed controller through quantitative valve-rotation and hybrid wiping evaluations, together with force-guided door opening and zero-force human-guided motion on a real wheeled-legged platform.
Chinese Translation
力控制的移动操作需要全身策略来协调移动与机械臂运动,同时调节末端执行器交互力。在浮动基座动力学和变化的支撑接触下,这具有挑战性,尤其是当控制中无法使用末端执行器力/力矩传感时。本文提出一种面向轮腿式移动操作的、结合混合无传感器力估计的力感知强化学习方法。所提方法将末端执行器力的结构化估计作为显式策略观测,从而无需使用末端执行器力/力矩传感器进行控制即可实现力引导的接触行为。该力估计通过结合广义动量观测、接触约束力旋量投影和时间残差学习获得:基于模型的组件提取全身扰动的物理结构化部分,而残差网络补偿剩余的运动相关偏差。估计力被集成到带有逐轴力/位置选择器的模式条件化全身策略中,使一个控制器内可实现自由空间运动、纯力调节以及混合力/位置控制。仿真结果证明了改进的无传感器力估计和力控制性能。硬件实验进一步在真实轮腿式平台上通过定量阀门旋转和混合擦拭评估,以及力引导开门和零力人类引导运动,验证了所提控制器。
cs.RO / 31 / 2609.13812

GeomVLA: Unifying Scene, Motion, and Action in 3D

GeomVLA:在 3D 中统一场景、运动与动作
Xiong, Ziyin, Gkanatsios, Nikos, Reuss, Moritz, Fragkiadaki, Katerina
Abstract
We present GeomVLA, a Vision-Language-Action (VLA) model that unifies perception, latent scene motion prediction, and action generation within a shared robot-centric 3D coordinate frame. Our approach lifts pretrained VLM features into spatially grounded 3D scene tokens using depth and camera calibration, while retaining the semantic representations learned during VLM pretraining. We further introduce a 3D Scene Trajectory Denoiser, a task-conditioned module that learns a latent representation of how scene points are expected to move in 3D. Rather than executing the predicted trajectory as an open-loop plan, GeomVLA extracts intermediate motion tokens from the trajectory denoiser and uses them to condition a 3D flow-based action denoiser through geometry-aware attention. GeomVLA achieves state-of-the-art performance on CALVIN, competitive performance on LIBERO and RoboTwin2.0, and outperforms strong baselines in real-world manipulation settings without robot-action pretraining. Extensive ablations show that future-motion reasoning alone is insufficient: the primary gains are associated with maintaining geometric consistency among scene representation, motion prediction, and robot actions throughout the perception-to-action pipeline.
Chinese Translation
我们提出 GeomVLA,一种视觉-语言-动作(Vision-Language-Action, VLA)模型,其在共享的以机器人为中心的三维坐标系内统一了感知、潜在场景运动预测与动作生成。我们的方法利用深度和相机标定,将预训练 VLM 特征提升为具有空间基础的 3D 场景 token,同时保留 VLM 预训练期间学到的语义表示。我们进一步引入 3D Scene Trajectory Denoiser(三维场景轨迹去噪器),这是一个以任务为条件的模块,学习场景点在三维空间中预期如何运动的潜在表示。GeomVLA 并不将预测轨迹作为开环计划来执行,而是从轨迹去噪器中提取中间运动 token,并通过几何感知注意力将其作为条件输入一个基于 3D 流的动作去噪器。GeomVLA 在 CALVIN 上取得了最先进性能,在 LIBERO 和 RoboTwin2.0 上具有竞争力,并在无需机器人动作预训练的真实世界操作设置中优于强基线。大量消融实验表明,仅靠未来运动推理是不够的:主要收益与在从感知到动作的整个流程中保持场景表示、运动预测和机器人动作之间的几何一致性有关。
cs.RO / 32 / 2609.13845

LePlanner: An Iterative Amortized Controller For World Models

LePlanner:一种用于世界模型的迭代摊销控制器
Bansal, Saksham, Naphade, Om, Aggarwal, Chayan, M, Vrishin
Abstract
World models trained with joint-embedding predictive architectures learn compact, structured latent representations from physical interaction, yet planning in these latent spaces typically relies on one of two costly approaches. Search-based planners such as CEM, MPPI, and iCEM optimize action sequences through many predictor rollouts, achieving strong performance at the cost of high per-decision compute and latency. Policy-based methods amortize inference into a single forward pass but can degrade on contact-rich tasks where the demonstration distribution is multimodal. We propose LePlanner, an amortized iterative controller that learns to construct and refine latent action sequences through a frozen world-model predictor. LePlanner is trained with an arrival-and-hold objective that encourages the controller to reach the goal at the earliest feasible horizon and remain there. This addresses horizon-reset procrastination, a failure mode in which repeated receding-horizon replanning continually postpones goal arrival. An additional action-Gaussian loss keeps generated actions near the support of the offline dataset. Across navigation, contact-rich manipulation, and continuous-control environments, LePlanner matches or exceeds search-based planners while requiring an order of magnitude fewer predictor evaluations and 3-49x lower wall-clock time per decision. It achieves success rates of 98% on PushT, 100% on Reacher, 100% on TwoRooms, and 92% on the OGBench Cube task. These results show that much of the structure discovered through online search can be amortized into a lightweight iterative policy, enabling fast, horizon-aware, nonlinear physical control without online optimization.
Chinese Translation
使用联合嵌入预测架构训练的世界模型从物理交互中学习紧凑、结构化的潜在表示,然而在这些潜在空间中进行规划通常依赖于两种代价高昂的方法之一。基于搜索的规划器(如 CEM、MPPI 和 iCEM)通过大量预测器展开来优化动作序列,以高每次决策计算量和延迟为代价实现了强大性能。基于策略的方法将推理摊销到单次前向传播中,但在演示分布为多模态的接触丰富任务上可能会性能下降。我们提出 LePlanner,一种摊销迭代控制器,它学习通过冻结的世界模型预测器来构造和精炼潜在动作序列。LePlanner 使用到达并保持目标进行训练,该目标鼓励控制器在最早可行的时域到达目标并停留在那里。这解决了时域重置拖延问题,这是一种失败模式,其中重复的滚动时域重规划不断推迟目标到达。额外的动作高斯损失使生成的动作保持在离线数据集的支撑集附近。在导航、接触丰富操作和连续控制环境中,LePlanner 匹配或超越基于搜索的规划器,同时需要的预测器评估次数少一个数量级,每次决策的挂钟时间降低 3-49 倍。它在 PushT 上达到 98% 的成功率,在 Reacher 上达到 100%,在 TwoRooms 上达到 100%,在 OGBench Cube 任务上达到 92%。这些结果表明,通过在线搜索发现的大部分结构可以摊销到一个轻量级迭代策略中,从而实现快速、时域感知的非线性物理控制,而无需在线优化。
cs.RO / 33 / 2609.13851

ReWeight: Leveraging Human Data for VLA Post-Training via Demonstration Retrieval and Sample Weighting

ReWeight:利用人类数据,通过演示检索与样本加权实现VLA后训练
Wang, Chenwei, Huang, Dianye, Ko, Match W. L., Bai, Chenjia, Jiang, Zhongliang
Abstract
Post-training vision-language-action (VLA) models for specific robots and tasks requires in-domain demonstrations, yet collecting diverse robot data is costly. Egocentric human demonstrations provide a scalable alternative, but directly mixing human and robot data can introduce cross-embodiment discrepancies and degrade policy performance. To address this challenge, we introduce ReWeight, a framework that incorporates human data into VLA post-training through demonstration-level retrieval and sample-level weighting. ReWeight learns a cross-embodiment visuomotor representation that combines visual observations with future actions to measure behavioral similarity between human and robot demonstrations. Based on optimal transport, it retrieves human demonstrations relevant to the target robot data and assigns larger weights to samples with smaller cross-embodiment discrepancies. We evaluate ReWeight using $\pi_{0.5}$ across eight simulation tasks and four real-world tasks under both clean and randomized settings. In simulation, ReWeight improves the average success rate of post-trained $\pi_{0.5}$ from 39% with only robot data and 44% with randomly mixed human-robot data to 57%. In the physical experimental setting, it achieves an average success rate of 68.8%, outperforming the baselines by 28.8% and 13.8%, respectively. Overall, ReWeight provides an effective paradigm for transforming abundant egocentric human experience into transferable supervision for robot learning. (Project webpage: https://reweight-vla.github.io/)
Chinese Translation
针对特定机器人和任务对视觉-语言-动作(VLA)模型进行后训练需要域内演示数据,但采集多样化机器人数据成本高昂。第一人称视角的人类演示提供了一种可扩展的替代方案,但直接混合人类与机器人数据会引入跨具身差异,并降低策略性能。为应对这一挑战,我们提出ReWeight,一个通过演示级检索和样本级加权将人类数据纳入VLA后训练的框架。ReWeight学习一种跨具身视觉运动表征,将视觉观测与未来动作相结合,以度量人类与机器人演示之间的行为相似性。基于最优传输,它检索与目标机器人数据相关的人类演示,并为跨具身差异较小的样本分配更大权重。我们在干净和随机化设置下,使用$\pi_{0.5}$在八个仿真任务和四个真实世界任务上评估ReWeight。在仿真中,ReWeight将后训练的$\pi_{0.5}$平均成功率从仅使用机器人数据时的39%、随机混合人-机器人数据时的44%提升至57%。在物理实验设置中,它取得了68.8%的平均成功率,分别比基线高出28.8%和13.8%。总体而言,ReWeight提供了一种有效范式,可将丰富的第一人称人类经验转化为机器人学习的可迁移监督。(项目网页:https://reweight-vla.github.io/)
cs.RO / 34 / 2609.13876

Co-Speech with You: Training-Free Personalization of Robot Co-Speech Gestures

与你共语:机器人共语手势的免训练个性化
Ding, Bosong, Ancel, Selma, Spigler, Giacomo, Kirtay, Murat
Abstract
Personal robots should adapt their co-speech gesture style to a new user without requiring model retraining. We present a training-free personalization pipeline that combines a frozen audio-conditioned diffusion prior with a gesture style encoder and lightweight conditioning adapters. The encoder is first trained to discriminate speaker identities and then jointly refined with the adapters using the diffusion objective, enabling a reusable style embedding to be extracted from approximately 10 seconds of enrollment motion through a single forward pass. To support this setting, we also release a Quest~3 capture application and a dataset of spontaneous co-speech motion from ten participants. We evaluate the system on held-out speakers using Style Recognition Accuracy (SRA) and Fr'echet Gesture Distance (FGD) to measure personalization and motion quality. Our approach improves SRA from 27.6\% for the frozen prior to 69.5\% while preserving motion quality (FGD 34.2 versus 34.8), and replacing the enrollment embedding with another person's reduces SRA to 11.4\%. The generated gestures are retargeted to a physical NAO robot, and this improvement also transfers to the robot deployment setting, where speech is synthesized using five TTS voices, retaining 67.3\% SRA.
Chinese Translation
个人机器人应能在无需重新训练模型的情况下,将其共语手势风格适配到新用户。我们提出一种免训练个性化流程,将冻结的音频条件扩散先验与手势风格编码器及轻量条件适配器相结合。该编码器首先被训练用于区分说话人身份,随后使用扩散目标与适配器联合优化,从而能够通过单次前向传播从约10秒的注册动作中提取可复用的风格嵌入。为支持这一设置,我们还发布了一个 Quest 3 采集应用以及一个包含十名参与者的自发共语动作数据集。我们在留出说话人上使用风格识别准确率(SRA)和 Fréchet 手势距离(FGD)评估系统,以衡量个性化程度和动作质量。我们的方法将 SRA 从冻结先验的 27.6% 提升至 69.5%,同时保持动作质量(FGD 34.2 对 34.8);将注册嵌入替换为另一个人的嵌入会使 SRA 降至 11.4%。生成的手势被重定向到实体 NAO 机器人,且这一改进也迁移到机器人部署场景,其中语音使用五种 TTS 声音合成,仍保持 67.3% 的 SRA。
cs.RO / 35 / 2609.13982

Rethinking the Implications of Human Feedback for Preference Learning in Human-Robot Collaboration

重新思考人类反馈对人机协作中偏好学习的隐含意义
Zhang, Qiping, Candon, Kate, Ghose, Debasmita, Vázquez, Marynel
Abstract
In Human-Robot Interaction, the standard approach to learn a reward model that represents human preferences for robot behavior consists of three steps. First, the robot collects limited direct evidence from human feedback (e.g., positive or negative binary feedback). Then, the robot utilizes the direct evidence to derive accepted or rejected labels to feasible but unchosen actions using fixed implication rules. Finally, the robot updates the reward model with both the direct and derived evidence. Unfortunately, the fixed rule can hinder preference learning: in a user study with two collaborative simulation environments, human-provided implication labels often differed from the standard fixed rule, and using the human labels substantially improved reward learning with the Preference Learning from Implicit and Explicit Feedback (PIE) algorithm. Consequently, we propose IMPLIED, an implication modeling method that treats fixed-rule implications as an initial guide while learning to infer and revise accepted and rejected action labels over time. Across evaluations on recorded human-robot interaction trajectories and a physical robot pizza-making study, IMPLIED predicts human implications more accurately than the fixed rule approach and LLM baselines, approaching the performance of a human-label oracle. In turn, IMPLIED reduces preference-estimation error and leads to robot actions that are more often rational with respect to a combined reward (which includes the true preference reward and a task-specific reward) compared to baselines. By learning to reason about the implications of human feedback, this work enables more faithful and efficient robot behavior adaptation during human-robot collaboration.
Chinese Translation
在人机交互中,学习一个表示人类对机器人行为偏好的奖励模型的标准方法包含三个步骤。首先,机器人从人类反馈中收集有限的直接证据(例如,正面或负面的二元反馈)。然后,机器人利用固定推断规则,将这些直接证据用于为可行但未被选择的动作推导出“接受”或“拒绝”标签。最后,机器人同时使用直接证据和推导出的证据来更新奖励模型。遗憾的是,固定规则可能阻碍偏好学习:在一项包含两个协作仿真环境的用户研究中,人类提供的推断标签常常不同于标准固定规则,而使用这些人类标签,通过从隐式和显式反馈中进行偏好学习(PIE)算法,显著改善了奖励学习。因此,我们提出 IMPLIED,一种推断建模方法;它将固定规则推断作为初始指导,同时学习随时间推断并修正接受和拒绝的动作标签。在记录的人机交互轨迹和一项实体机器人制作披萨研究的评估中,IMPLIED 预测人类推断的准确率高于固定规则方法和 LLM 基线,接近人类标签 oracle 的性能。反过来,与基线相比,IMPLIED 降低了偏好估计误差,并使机器人动作在组合奖励(包含真实偏好奖励和任务特定奖励)下更常是理性的。通过学习推理人类反馈的隐含意义,这项工作使人机协作过程中的机器人行为适应更加忠实且高效。
cs.RO / 36 / 2609.13984

What Makes an Efficient VLA? Navigating Action-Head Design, Scaling, and Latency

什么造就了高效的VLA?探索动作头设计、扩展与延迟
Sun, Luoyang, Xia, Guoyang, Li, Fengfa, Ren, Lei, Cui, Xinyu, Zhang, Haifeng, Feng, Fangxiang, Zhang, Kaike, Zhan, Kun, Xie, Yan, Wang, Jun, Deng, Cheng
Abstract
Vision-Language-Action (VLA) models combine a pretrained vision encoder, a language backbone, and an action head, but their relative contribution has not been established under controlled, latency-paired conditions. We fix the backbone families (SigLIP2 and Qwen2.5) and the training pipeline, sweep action-head design and module scale, and pair each configuration with measured on-device latency. The study yields three findings. First, action-head performance is governed primarily by initialization rather than decoder architecture, loss, or inference budget: copying the last transformer layers of the language backbone into the head is the single largest lever, at no latency cost, and the only axis that helps at every module scale. Alignment also explains the other axes: flow matching and a heavier decoder pay off only while the head is misaligned and reverse once it is aligned, and extra inference passes give no measurable benefit; expressiveness appears to substitute for missing alignment. We read this as representation transfer: the aligned head keeps attending to the instruction's object nouns and stays close to the backbone in weight space rather than relearning to act from scratch. Because we reach alignment only through initialization, we offer this as the account that best organizes the measurements, not a demonstrated cause, and name the control that would settle it. Second, capacity pays only after alignment: the aligned action head is the highest-return module to scale. Third, those returns diminish sharply near the size today's $\pi$-series VLAs already use, so further growth buys little in-domain accuracy for its latency. These specify EffVLA, a compact model matching the strongest open-source VLAs on standard LIBERO, leading on most LIBERO-Plus perturbation axes at lower latency, and transferring to a real SO-ARM101 arm with the recipe unchanged.
Chinese Translation
视觉-语言-动作(VLA)模型结合了预训练的视觉编码器、语言主干网络和动作头,但在受控且延迟配对的条件下,它们的相对贡献尚未确定。我们固定主干网络系列(SigLIP2和Qwen2.5)和训练流程,遍历动作头设计和模块规模,并将每种配置与实测的设备端延迟配对。该研究得出三项发现。第一,动作头的性能主要受初始化而非解码器架构、损失或推理预算的影响:将语言主干网络的最后几层Transformer层复制到动作头中是最大的杠杆,且不增加延迟成本,并且是唯一在所有模块规模下都有帮助的维度。对齐也解释了其他维度:流匹配和更重的解码器仅在动作头未对齐时有效,而一旦对齐则效果反转,额外的推理步骤没有可测量的收益;表达能力似乎替代了缺失的对齐。我们将此解读为表示迁移:对齐后的动作头持续关注指令中的物体名词,并在权重空间中保持与主干网络接近,而非从头重新学习动作。由于我们仅通过初始化达到对齐,我们将其作为最能组织这些测量结果的解释,而非已证实的因果,并指出可以解决这一问题的对照实验。第二,容量仅在对齐后才有回报:对齐后的动作头是扩展回报最高的模块。第三,在当今$\pi$系列VLA已经使用的规模附近,这些回报急剧递减,因此进一步增长在域内精度上收益甚微,却要付出延迟代价。这些结果具体化为EffVLA,这是一个紧凑模型,在标准LIBERO上匹配最强的开源VLA,在大多数LIBERO-Plus扰动维度上以更低延迟领先,并且在不改变方案的情况下迁移到真实的SO-ARM101机械臂。
cs.RO / 37 / 2609.14058

Mobile Multi-Robot Navigation under Runtime Uncertainty via Koopman Operator Learning and Nonlinear Model Predictive Control

基于Koopman算子学习和非线性模型预测控制的运行时不确定性下的移动多机器人导航
Zhang, Xiaobin, Karydis, Konstantinos
Abstract
In this work, we developed a nonlinear model predictive control (NMPC) framework that employs learned dynamics via the Koopman Operator theory for mobile multi-robot navigation. We formulated and solved NMPC problems using a lifted bilinear Koopman-based model that accurately predicts affine input systems affected by perturbations and uncertainties. Two exemplary multi-robot navigation problems are considered: target reaching and formation control. The output of our method enables closed-loop multi-robot navigation and formation control in environments populated with obstacles, whereby the Koopman operator-based model used in the NMPC formulation addresses runtime uncertainties, namely, various degrees of random wheel slipping. We validated the effectiveness of our method for both problems via extensive numerical simulations in different environments with wheeled robots affected by different amounts of slip and without knowledge of their true dynamic models.
Chinese Translation
在本工作中,我们开发了一种非线性模型预测控制(NMPC)框架,该框架利用通过Koopman算子理论学习到的动力学来实现移动多机器人导航。我们使用一个提升的双线性Koopman模型来构建和求解NMPC问题,该模型能够准确预测受扰动和不确定性影响的仿射输入系统。我们考虑了两种示例性的多机器人导航问题:目标到达和编队控制。我们方法的输出能够在布满障碍物的环境中实现闭环多机器人导航和编队控制,其中NMPC公式中使用的基于Koopman算子的模型解决了运行时不确定性,即不同程度的随机车轮打滑。我们通过在不同环境中进行的广泛数值仿真验证了我们方法对这两个问题的有效性,这些环境中的轮式机器人受到不同程度的打滑影响,且不知道其真实动力学模型。
cs.RO / 38 / 2609.14087

Gradient-Free Neural Hamilton-Jacobi Reachability for Scalable Safety-Critical Control

面向可扩展安全关键控制的无梯度神经Hamilton-Jacobi可达性
Feng, Zeyuan, Sahin, Ali Fuat, Thorup, Santiago, Bansal, Somil
Abstract
Hamilton-Jacobi (HJ) reachability provides a principled framework for synthesizing safety certificates and robust controllers for safety-critical robotic systems. However, applying reachability analysis to high-dimensional nonlinear systems remains challenging: classical grid-based solvers suffer from the curse of dimensionality, continuous-time neural solvers require accurate spatial value gradients, and reinforcement-learning-based approaches often suffer from weak boundary anchoring and non-stationary adversarial policy optimization. We propose a discrete-time neural reachability framework for control-disturbance-affine systems that learns backward reachable tubes (BRTs) and backward reach-avoid tubes (BRATs) through Bellman-Isaacs value propagation. Our key idea is to combine equation-driven self-supervision with structured policy learning: rather than computing explicit PDE-gradients, we exploit the bang-bang structure of optimal safety interventions to construct approximate teacher actions from gradient-free value probes, converting adversarial actor learning into supervised policy learning. To stabilize long-horizon value propagation, we leverage the learned actor to train the value function backward from the terminal boundary using a windowed temporal curriculum, where each window is used as the boundary condition for the next window. Across benchmark problems up to 80 dimensions, our method learns accurate reachability value functions while improving stability over existing learning-based solvers. We further demonstrate observation-space scalability on F1-tenth racing with over 16,000-dimensional egocentric inputs. The learned safety filter generalizes zero-shot to unseen tracks and transfers to a physical RC car, achieving real-time robust collision avoidance.
Chinese Translation
Hamilton-Jacobi (HJ)可达性为安全关键机器人系统合成安全证书和鲁棒控制器提供了有原则的框架。然而,将可达性分析应用于高维非线性系统仍然具有挑战性:经典的基于网格的求解器面临维数灾难,连续时间神经求解器需要精确的空间价值梯度,而基于强化学习的方法通常存在边界锚定薄弱和非平稳对抗策略优化的问题。我们提出了一种针对控制-扰动-仿射系统的离散时间神经可达性框架,通过Bellman-Isaacs价值传播学习后向可达管(BRTs)和后向可达-避免管(BRATs)。我们的关键思想是将方程驱动的自监督与结构化策略学习相结合:不是计算显式的PDE梯度,而是利用最优安全干预的bang-bang结构,从无梯度价值探测中构造近似教师动作,将对抗式策略学习转化为监督式策略学习。为了稳定长时程价值传播,我们利用学习到的策略从终端边界开始,使用窗口化时间课程向后训练价值函数,其中每个窗口用作下一个窗口的边界条件。在高达80维的基准问题上,我们的方法学习了准确的可达性价值函数,同时相比现有的基于学习的求解器提高了稳定性。我们进一步在F1-tenth竞速中展示了观测空间的可扩展性,其具有超过16,000维的自我中心输入。学习到的安全滤波器能够零样本泛化到未见过的赛道,并迁移到物理RC小车,实现实时鲁棒避碰。
cs.RO / 39 / 2609.14133

Vision-Force Admittance Learning for Peg Insertion into a Movable Hole

视觉-力导纳学习用于将轴插入可移动孔洞
Chen, Yuzhong, Liang, Yongqing, Xu, Yunzhi, Fang, Irving, Kidder, Chase, Wang, Hui-ping, Haque, Raihan, Zhang, Yubiao, Feng, Chen
Abstract
Precise manipulation in dynamic environments, whether induced by a mobile robot base or a target with unknown motion, remains a major challenge in robotics. Manipulation in dynamic environments introduces substantial uncertainty, which fundamentally conflicts with the tight precision requirement of precise tasks such as peg-in-the-hole. We propose a Vision-Force Admittance Learning (VFAL) framework that fuses asynchronous visual feedback with a high-frequency force-based model, using visual pose estimations as a regularization term. VFAL adapts insertion strategies online to dynamic motion while maintaining millimeter-level precision. To obtain robust, low-frequency pose information, we employ state-of-the-art vision foundation models for visual pose estimation. Additionally, we incorporate failure recovery mechanisms to enhance overall robustness. We validate our approach in real-world experiments, demonstrating high success rates and strong adaptability to various pegs and dynamic environments.
Chinese Translation
动态环境中的精确操作,无论是源于移动机器人基座还是具有未知运动的目标,仍然是机器人学中的一个重大挑战。动态环境中的操作引入了巨大的不确定性,这与轴孔插入等精密任务的高精度要求存在根本性冲突。我们提出了一种视觉-力导纳学习(VFAL)框架,该框架融合了异步视觉反馈与基于力的高频模型,并将视觉位姿估计用作正则化项。VFAL 能够在线调整插入策略以适应动态运动,同时保持毫米级精度。为了获得鲁棒的低频位姿信息,我们采用了最先进的视觉基础模型进行视觉位姿估计。此外,我们引入了故障恢复机制以增强整体鲁棒性。我们在真实世界实验中验证了我们的方法,展示了高成功率以及对各种轴和动态环境的强大适应性。
cs.RO / 40 / 2609.14146

When Faster VLA Deployment Changes Closed-Loop Behavior: Task Success-Latency Analysis of SmolVLA Across PyTorch and ONNX Variants

当更快的 VLA 部署改变闭环行为:SmolVLA 在 PyTorch 与 ONNX 变体上的任务成功率-延迟分析
Islam, Rafiqul
Abstract
Vision-language-action (VLA) deployment can reduce inference latency while changing closed-loop task behavior. We evaluate HuggingFaceVLA/smolvla_libero on an RTX 2060 (6 GB) in LIBERO Spatial and Object (MuJoCo 3.3.2, LeRobot 0.6.1, seed 42), comparing PyTorch+AMP with ONNX Runtime CUDA Execution Provider (CUDA EP). The main evaluation uses 100 episodes/suite; a paired rollout uses 300 episodes/suite. PyTorch+AMP reaches 70.0%/88.0% Spatial/Object success at 1181 ms p99. Requested-FP16 and requested-INT8 ONNX reduce tether-inspect p99 to 601 ms and 532 ms, while Spatial success falls to 41.0% and 40.0% and Object remains at 89.0%. A graph audit shows those artifacts are byte-identical FP32 graphs, so the requested-INT8 row is not operator-level INT8 quantization. A static language-width ablation (16/24/32 tokens) yields Spatial success of 41.0%, 75.0%, and 71.0%; widths 24 and 32 recover much of the Spatial drop while Object success and uniform-bench latency stay approximately stable. Width-24 ONNX Spatial success is comparable to the PyTorch+AMP baseline at roughly half the latency (Wilson intervals overlap; two-proportion chi-squared p=0.53). Context width is an important contributor in this stack; it does not account for every PyTorch-vs-ONNX difference. Deployment evaluation should jointly report latency, artifact inspection, interface constraints, and closed-loop success. Code: https://github.com/rafiqul713/smolvla-libero-onnx.
Chinese Translation
视觉-语言-动作(VLA)部署可以降低推理延迟,同时改变闭环任务行为。我们在 RTX 2060(6 GB)上于 LIBERO Spatial 和 Object(MuJoCo 3.3.2,LeRobot 0.6.1,种子 42)中评估 HuggingFaceVLA/smolvla_libero,比较 PyTorch+AMP 与 ONNX Runtime CUDA Execution Provider(CUDA EP)。主评估每个套件使用 100 个回合;配对 rollout 每个套件使用 300 个回合。PyTorch+AMP 在 p99 为 1181 ms 时达到 70.0%/88.0% 的 Spatial/Object 成功率。Requested-FP16 和 Requested-INT8 ONNX 将 tether-inspect 的 p99 降至 601 ms 和 532 ms,而 Spatial 成功率降至 41.0% 和 40.0%,Object 保持 89.0%。图审计显示这些产物是字节完全相同的 FP32 图,因此 requested-INT8 行并非算子级 INT8 量化。静态语言宽度消融(16/24/32 个 token)得到的 Spatial 成功率为 41.0%、75.0% 和 71.0%;宽度 24 和 32 恢复了 Spatial 的大部分下降,而 Object 成功率和 uniform-bench 延迟大致保持稳定。宽度 24 的 ONNX 在 Spatial 成功率上与 PyTorch+AMP 基线相当,而延迟约为其一半(Wilson 区间重叠;两比例卡方检验 p=0.53)。上下文宽度是该技术栈中的一个重要影响因素;但它并不能解释 PyTorch 与 ONNX 之间的所有差异。部署评估应联合报告延迟、产物检查、接口约束和闭环成功率。代码:https://github.com/rafiqul713/smolvla-libero-onnx。
cs.RO / 41 / 2609.14156

Visible Touch: Rendering Contact for Visuomotor Policies

Visible Touch:为视觉运动策略渲染接触信息
Dogan, Metin Alp, Sun, Edward, Xu, Feng, Wu, Daniel, Peng, Allen, Hong, Dennis, Cui, Yuchen
Abstract
Integrating contact information into visuomotor policies remains an open problem. Touch is essential to robust manipulation, yet most modern policies, including pretrained vision-language-action (VLA) models, operate from vision and proprioception alone. Existing approaches to closing this gap require specialized tactile hardware, add separate tactile encoders, or commit to non-image policy backbones, all incompatible with the modern paradigm of image-conditioned policies built on pretrained 2D visual representations. Our key insight is that the bottleneck is not the contact information itself, but how it is delivered: when contact signals are exposed in the same spatial frame as the scene the policy already attends to, they become directly usable by any image-conditioned policy without architectural changes. We operationalize this insight in Visible Touch, paired with a custom low-cost magnetic contact sensor that is open-sourced and fabricated from off-the-shelf parts via a parametric CAD-to-mold pipeline. Across the LIBERO benchmark, Visible Touch improves BC-Transformer success by 15.7 percentage points on average in the 2-view setting, with similar gains in the 1-view setting; controlled comparisons show that the contact-integration strategy strongly affects how effectively tactile information is used. The pattern holds when fine-tuning pretrained VLAs: miniVLA on LIBERO gains 25 percentage points on average, and $\pi_{0.5}$ on four real-world contact-rich tasks gains 30 percentage points with our custom sensor.
Chinese Translation
将接触信息集成到视觉运动策略中仍然是一个未解决的问题。触觉对于稳健的操作至关重要,然而大多数现代策略,包括预训练的视觉-语言-动作(VLA)模型,仅依靠视觉和本体感觉运行。现有的弥补这一差距的方法需要专门的触觉硬件,添加单独的触觉编码器,或采用非图像策略主干,这些都与基于预训练2D视觉表示的图像条件策略的现代范式不兼容。我们的关键见解是,瓶颈不是接触信息本身,而是它如何被传递:当接触信号暴露在策略已经关注的场景相同的空间框架中时,它们可以被任何图像条件策略直接使用,而无需架构改变。我们在Visible Touch中实现了这一见解,并搭配了一个定制的低成本磁性接触传感器,该传感器是开源的,通过参数化CAD到模具的流程由现成部件制造。在LIBERO基准测试中,Visible Touch在2视图设置下将BC-Transformer的成功率平均提高了15.7个百分点,在1视图设置下也有类似的提升;受控比较表明,接触集成策略强烈影响触觉信息的使用效果。这一模式在微调预训练VLA时也成立:在LIBERO上,miniVLA平均提高了25个百分点,而在四个真实世界的接触丰富任务中,使用我们的定制传感器,$\pi_{0.5}$提高了30个百分点。
cs.RO / 42 / 2609.14173

GIFT: Glove-Inferred Force Transfer: Force-Aware Human-to-Robot Skill Transfer from a Wearable Sensing Glove to a Robot Hand Without Tactile Sensors

GIFT:手套推断力传递——从可穿戴传感手套到无触觉传感器机器人手的力感知人类到机器人技能迁移
Sarusi, Tzah
Abstract
Human-to-robot skill transfer from sensing gloves has so far relied on shared hardware: the same tactile glove worn by the demonstrator and the robot, or a learned alignment between two tactile sensors. We present GIFT (Glove-Inferred Force Transfer), a pipeline in which the interface between human and robot is a physical unit rather than a shared sensor: fingertip force is measured in newtons on the human side and estimated in newtons on the robot side. A wearable glove records finger flexion, calibrated fingertip force, and wrist orientation, while a head-mounted camera records the demonstration; no robot is present. At deployment, the robot estimates force from actuator-current residuals relative to a free-space baseline, through a calibrated mapping to newtons, so any position-controlled hand that reports motor current can serve as the deployment platform. The policy uses a glove-space state and predicts finger-position targets; the robot enters only through two calibrated adapters, a retargeting decoder and a force estimator. We evaluate GIFT on a cup grasp-and-hold task with two action-chunking policies trained on the same demonstrations, with fingertip-force inputs retained in one and zeroed in the other. In a 50-rollout evaluation with sample size and metrics fixed before scoring, both policies succeeded in all 25 rollouts. The median of the per-rollout hold-phase grip-force estimates was 53% lower with force inputs: 1.20 N versus 2.55 N (one-sided Mann-Whitney U, p<0.0001). In an observation ablation, a vision-only policy achieved 0/15 grasps, policies given hand-command state acquired the grasp, and the force inputs determined how hard the policy held. A force channel measured on the human hand thus transfers to a robot hand with no tactile hardware, through a retargeting map from five glove channels to seven robot actuators, with no sensor shared between the two.
Chinese Translation
迄今为止,从传感手套进行的人到机器人技能迁移一直依赖于共享硬件:演示者和机器人佩戴同一只触觉手套,或两个触觉传感器之间学习得到的对齐。我们提出 GIFT(Glove-Inferred Force Transfer,手套推断力传递),其中人与机器人之间的接口是一个物理单位而非共享传感器:人类侧以牛顿为单位测量指尖力,机器人侧以牛顿为单位估计指尖力。可穿戴手套记录手指弯曲、校准后的指尖力和手腕朝向,头戴式相机记录演示过程;过程中没有机器人。部署时,机器人根据相对于自由空间基线的执行器电流残差,通过校准映射到牛顿来估计力,因此任何能报告电机电流的位置控制手都可以作为部署平台。该策略使用手套空间状态并预测手指位置目标;机器人仅通过两个校准适配器接入,即重定向解码器和力估计器。我们在一个杯子抓取并保持任务上评估 GIFT,使用两个在同一组演示上训练的动作分块策略,其中一个保留指尖力输入,另一个将指尖力输入置零。在一次包含50次 rollout 的评估中,样本量和指标在评分前固定,两个策略在各自的全部25次 rollout 中均成功。每次 rollout 保持阶段握力估计的中位数在使用力输入时低53%:1.20 N 对 2.55 N(单侧 Mann-Whitney U 检验,p<0.0001)。在一项观测消融中,仅视觉策略实现了 0/15 次抓取,给定手部命令状态的策略成功抓取,而力输入决定了策略握持的力度。因此,在人类手上测量的力通道通过一个从五个手套通道到七个机器人执行器的重定向映射,传递到无触觉硬件的机器人手,二者之间不共享任何传感器。
cs.RO / 43 / 2609.14183

DreamSat-Bench: Development and Initial Testing of a Testbed for AI-Based Pose Estimation from 3D Reconstruction

DreamSat-Bench:用于基于3D重建的AI位姿估计的测试平台的开发与初步测试
Posadas-Nava, Alex, Berne, August, Lavezzi, Giovanni, Shah, Kareena, Carrasco, Alejandro, Uwumukiza, Josiane, Battaglia, Giacomo, Panicucci, Paolo, Wijayatunga, Minduli C., Rodriguez-Fernandez, Victor, Linares, Richard
Abstract
This paper presents the development and initial testing of DreamSat-Bench, a modular rendezvous and proximity operation testbed designed to benchmark AI-based relative navigation techniques. By integrating a software- and hardware-in-the-loop robotic pipeline, the platform enables a seamless transition from digital simulation to physical reality. DreamSat-Bench unifies state-of-the-art robotic learning tools such as MuJoCo, Isaac Lab, and LeRobot into a single benchmarking platform, utilizing robotic arms to trace 3D trajectories. The platform allows for extensive customization of orbital environments and lighting to evaluate the simulation-to-reality gap. We demonstrate the testbed's utility by evaluating an end-to-end vision-based navigation pipeline that pairs DreamSat, a generative AI framework for single-view 3D reconstruction, with FoundationPose for zero-shot 6-DoF tracking of unseen spacecraft. Initial testing explores mission-representative orbital segments, including fixed-point station-keeping and fly-around characterization. Through a series of parametric studies, we quantify the impacts of reconstruction latency, mesh resolution, orbital range, and illumination geometry on pose estimation accuracy. Finally, a preliminary hardware-in-the-loop campaign qualitatively validates the physical deployment of the pipeline, identifying target symmetry and accumulated tracking drift as critical factors for robust navigation. DreamSat-Bench provides a rigorous framework for maturing autonomous navigation with unprepared space assets in the absence of prior geometric models.
Chinese Translation
本文介绍了DreamSat-Bench的开发与初步测试,这是一个模块化的交会与近距离操作测试平台,旨在对基于AI的相对导航技术进行基准测试。通过集成软件在环和硬件在环的机器人流程,该平台实现了从数字仿真到物理现实的无缝过渡。DreamSat-Bench将MuJoCo、Isaac Lab和LeRobot等最先进的机器人学习工具统一到一个单一的基准测试平台中,利用机械臂来追踪3D轨迹。该平台允许对轨道环境和光照进行广泛定制,以评估仿真到现实的差距。我们通过评估一个端到端的基于视觉的导航流程来展示该测试平台的实用性,该流程将DreamSat(一种用于单视图3D重建的生成式AI框架)与FoundationPose(用于未见航天器的零样本6自由度跟踪)相结合。初步测试探索了具有任务代表性的轨道段,包括定点站位保持和绕飞表征。通过一系列参数研究,我们量化了重建延迟、网格分辨率、轨道范围和光照几何对位姿估计精度的影响。最后,初步的硬件在环实验定性地验证了该流程的物理部署,并将目标对称性和累积跟踪漂移确定为鲁棒导航的关键因素。DreamSat-Bench提供了一个严格的框架,用于在缺乏先验几何模型的情况下,使自主导航与未准备的太空资产日趋成熟。
cs.RO / 44 / 2609.14198

Novel Ex-vivo Calf Brain Model with Integrated Sub-Skull Force Sensors to Access Simulated Neurosurgical Procedures

新型离体小牛脑模型集成颅骨下力传感器用于评估模拟神经外科手术
Binhammad, Hamad, Ballestero, Matheus, Babgi, Mohammed, Shaka, Seana, Hemati, Nima, Saeedi, Rothaina, Mazidi, Aiden, Giglio, Bianca, Dou, Rukun, Gueziri, Houssem-Eddine, Hooshiar, Amir, Del Maestro, Rolando F.
Abstract
Surgical tissue manipulation demands precision; however, tool-tissue manipulation force magnitudes under realistic conditions are rarely quantified. To address this gap, we proposed and validated a portable ex-vivo force-sensing platform that measures tool-tissue interaction forces across the skull-brain interface during simulated neurosurgery. The system involves fresh calf brain tissue, used as a biological surrogate for brain parenchyma, placed in a 3D-printed human skull model equipped with a 6 degree-of-freedom force/torque sensor and a real-time data acquisition system. Five validation protocols assessed the accuracy and dynamic fidelity of the platform against ground-truth measurement, static accuracy and linearity using calibrated weights (0.5-50 g), minimum detectable force, spatial consistency across different anatomical regions, effect of surgical draping, and long-duration stability. Across protocols, measured forces showed excellent agreement with reference loads (correlation R = 0.9997), with RMSE < 0.005 N and mean relative error under 2%. The platform reliably detected low-magnitude forces down to 1 g (9.8 mN), while surgical drapes introduced no meaningful signal distortion and prolonged recordings exhibited minimal drift. Overall, the proposed framework provides objective, high-fidelity force quantification for skill training and performance assessment using fresh calf brain tissue and may serve as a foundation for force-based evaluation across other surgical procedures. Future work will integrate clinically used surgical instruments to increase procedural realism and will progress toward clinical trials to evaluate usability, educational impact, and translational relevance in practice-adjacent settings.
Chinese Translation
手术组织操作要求精确;然而,在真实条件下,器械-组织操作力的大小很少被量化。为了弥补这一空白,我们提出并验证了一种便携式离体力传感平台,用于测量模拟神经外科手术过程中跨颅骨-脑界面的器械-组织相互作用力。该系统涉及新鲜小牛脑组织,作为脑实质的生物替代物,置于配备有6自由度力/扭矩传感器和实时数据采集系统的3D打印人体颅骨模型中。五种验证方案评估了平台的准确性和动态保真度与真实值测量、使用校准重量(0.5-50克)的静态准确性和线性度、最小可检测力、不同解剖区域的空间一致性、手术铺巾的影响以及长时间稳定性。在所有方案中,测量力与参考载荷显示出极好的一致性(相关系数R = 0.9997),RMSE < 0.005 N,平均相对误差低于2%。该平台可靠地检测到低至1克(9.8毫牛)的低幅值力,而手术铺巾未引入有意义的信号失真,长时间记录表现出最小漂移。总体而言,所提出的框架使用新鲜小牛脑组织为技能培训和性能评估提供了客观、高保真的力量化,并可能作为其他外科手术中基于力的评估的基础。未来的工作将整合临床使用的手术器械以提高手术的真实感,并将推进临床试验,以评估在实践相关环境中的可用性、教育影响和转化相关性。
cs.RO / 45 / 2609.14208

ITA-LaCAM: A Complete and Scalable TAPF Solver via Assignment-Aware Configuration-Space Search

ITA-LaCAM:一种基于分配感知配置空间搜索的完整且可扩展的TAPF求解器
Tang, Yimin, Zhang, Han, Chan, Shao-Hung, Kim, Junsoo, Bıyık, Erdem, Koenig, Sven, Chen, Jingkai
Abstract
Combined Target Assignment and Path Finding (TAPF) requires assigning targets for agents while simultaneously planning collision-free paths. We present ITA-LaCAM, a complete and scalable TAPF solver inspired by LaCAM and ITA-CBS. In ITA-LaCAM, each joint-configuration node carries an agent-to-target matching. When a successor is generated, ITA-LaCAM incrementally repairs the matching for the agents that moved and uses the targets to guide PIBT successor generation. This design enables adaptive reassignment without explicitly enumerating the combinatorial assignment space, while preserving LaCAM's completeness and scalability. Across 9,760 benchmark instances on eight maps with 5--200 agents, ITA-LaCAM solved 100% of the instances, compared with 95.6% for IR-TAPF configured with DBS-Hungarian. ITA-LaCAM found an initial solution faster in 84.0% of the comparisons and achieved a lower sum of costs in 65.0% of the instances solved by both methods.
Chinese Translation
组合目标分配与路径规划(TAPF)需要为智能体分配目标,同时规划无碰撞路径。我们提出了ITA-LaCAM,一种受LaCAM和ITA-CBS启发的完整且可扩展的TAPF求解器。在ITA-LaCAM中,每个联合配置节点都携带一个智能体到目标的匹配。当生成后继节点时,ITA-LaCAM增量地修复移动智能体的匹配,并使用目标来指导PIBT后继生成。这种设计能够在不显式枚举组合分配空间的情况下进行自适应重新分配,同时保持LaCAM的完整性和可扩展性。在八张地图上使用5至200个智能体的9,760个基准实例中,ITA-LaCAM解决了100%的实例,而配置了DBS-Hungarian的IR-TAPF解决了95.6%。在84.0%的比较中,ITA-LaCAM更快地找到初始解,并且在两种方法都解决的实例中,65.0%的实例实现了更低的总成本。
cs.RO / 46 / 2609.14219

Task-Specified Active Metrological Inspection with Measurement-Steered VLA Manipulation and Deterministic Evidence Gating

任务指定的主动计量检测:测量引导的 VLA 操作与确定性证据门控
Chen, Zhiling, Ge, Jingzhan, Chen, Ruimin, Castanier, Matthew P., Gorsich, David, Imani, Farhad
Abstract
High-mix low-volume (HMLV) manufacturing requires inspection systems to adapt to changing parts, specifications, and work orders without repeated task-specific programming. Existing inspection automation typically assumes predefined sensing sequences, while general purpose robot agents optimize task completion rather than the completeness and validity of metrological evidence. We formulate task-specified active metrological inspection and propose From Requirements to Admissible Metrological Evidence (FRAME), a hierarchical dual-arm framework that converts an inspection instruction and structured specification into traceable conformance evidence. FRAME coordinates learned manipulation with calibrated laser profilometry: a task manager grounds and schedules requirements, active surface correspondence verifies physical-to-specification localization, and evidence memory tracks measurement provenance, admissibility, and coverage. Learned components may propose inspection targets and physical access actions, but deterministic datum-grounded measurement, admissibility checks, coverage auditing, and conformance evaluation prevent incomplete or unverified evidence from authorizing PASS. A series of physical experiments shows that FRAME achieves higher end-to-end inspection reliability, fewer false accepts, and shorter task completion time.
Chinese Translation
高混合低产量(HMLV)制造要求检测系统能够适应变化的零件、规格和工作订单,而无需重复进行任务特定的编程。现有的检测自动化通常假定预定义的传感序列,而通用机器人代理优化任务完成,而非计量证据的完整性和有效性。我们提出了任务指定的主动计量检测问题,并提出了一种从需求到可接受计量证据(FRAME)的分层双臂框架,该框架将检测指令和结构化规范转换为可追溯的符合性证据。FRAME 协调学习到的操作与校准的激光轮廓测量:任务管理器锚定并调度需求,主动表面对应验证物理到规范的定位,证据记忆跟踪测量来源、可接受性和覆盖范围。学习组件可以提出检测目标和物理访问动作,但确定性的基于基准的测量、可接受性检查、覆盖审计和符合性评估防止不完整或未验证的证据授权通过(PASS)。一系列物理实验表明,FRAME 实现了更高的端到端检测可靠性、更少的错误接受和更短的任务完成时间。
cs.RO / 47 / 2609.14261

VGFM: Expressive Robot Policies via Dense Value Guidance in Flow Matching

VGFM:通过流匹配中的密集价值引导实现富有表现力的机器人策略
Koirala, Prajwal, Campbell, Mark
Abstract
Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative models enable rich and multimodal action representations, expanding the capability of this paradigm for complex robotic control. However, policy improvement with multi-step generative actors remains challenging. In offline reinforcement learning (RL), incorporating value-based objectives along generative trajectories often introduces substantial training complexity, including backpropagation through time (BPTT), auxiliary architectures, or distillation losses. We propose Value-Guided Flow Matching (VGFM), a scalable offline RL framework that enables dense value-guided shaping within a flow-based policy while avoiding BPTT and additional algorithmic overhead. VGFM parameterizes the policy as a conditional flow-matching model in action (x-prediction) space, ensuring that each intermediate flow step produces a valid robot action that can be directly evaluated by a standard offline RL critic. This design allows value guidance to be applied at randomly sampled flow times without differentiating through the entire generative trajectory, while preserving inference-time flexibility by varying the discretization of the underlying flow ODE without retraining. Evaluated on robotic locomotion and manipulation tasks in OGBench, VGFM achieves strong performance across a wide range of tasks under rigorous evaluation protocols. With minimal hyperparameter tuning, these results demonstrate that VGFM provides a simple, scalable, and effective approach for expressive policy learning in long-horizon, goal-oriented robotic control.
Chinese Translation
近期机器人学习范式越来越依赖大规模的机器人交互离线数据集来训练控制策略。富有表现力的生成模型能够实现丰富且多模态的动作表示,扩展了该范式在复杂机器人控制中的能力。然而,使用多步生成式执行器进行策略改进仍然具有挑战性。在离线强化学习(RL)中,沿生成轨迹引入基于价值的目标通常会导致大量的训练复杂性,包括时间反向传播(BPTT)、辅助架构或蒸馏损失。我们提出价值引导流匹配(VGFM),一种可扩展的离线 RL 框架,能够在基于流的策略中实现密集价值引导的塑形,同时避免 BPTT 和额外的算法开销。VGFM 将策略参数化为动作(x-预测)空间中的条件流匹配模型,确保每个中间流步骤产生有效的机器人动作,该动作可由标准离线 RL 评论家直接评估。这种设计允许在随机采样的流时间上应用价值引导,而无需对整个生成轨迹进行微分,同时通过改变底层流 ODE 的离散化来保持推理时的灵活性,而无需重新训练。在 OGBench 中的机器人运动和操作任务上进行评估,VGFM 在严格的评估协议下在广泛的任务中取得了强大的性能。通过最少的超参数调整,这些结果表明 VGFM 为长时域、目标导向的机器人控制中的富有表现力的策略学习提供了一种简单、可扩展且有效的方法。
cs.RO / 48 / 2609.14268

Learning Communication-Conditioned Generative Policies for Decentralized Multi-Agent Collision Avoidance

学习面向去中心化多智能体避碰的通信条件生成策略
Koirala, Prajwal, Campbell, Mark
Abstract
In this work, we propose a decentralized communication-conditioned generative framework for multi-agent collision avoidance. Agents generate short-horizon action sequences using a flow-matching policy trained from privileged offline demonstrations with access to global state. The demonstrations do not include explicit communication signals; instead, agents learn to exchange and aggregate latent messages that encode interaction-relevant intent under partial observability. This formulation supports flexible inference at test time, where unconditioned generation corresponds to independent behavior and communication-conditioned generation enables coordinated interaction without centralized planning. The resulting policies operate in a fully decentralized manner at execution time, relying only on local observations and learned messages. Combined with a receding-horizon inference scheme, the proposed approach enables efficient single-step inference of short-horizon action sequences and degrades gracefully under communication dropouts. Extensive simulation results demonstrate near-expert collision avoidance performance and strong generalization to denser, unseen multi-agent scenarios, along with zero-shot transfer to real-robot experiments.
Chinese Translation
本文提出了一种面向多智能体避碰的去中心化通信条件生成框架。智能体使用流匹配策略生成短时域动作序列,该策略从可访问全局状态的特权离线示范中训练得到。这些示范不包含显式的通信信号;相反,智能体学习交换和聚合潜在消息,这些消息在部分可观测性下编码了与交互相关的意图。这种形式支持测试时的灵活推理,其中无条件生成对应于独立行为,而通信条件生成则无需集中式规划即可实现协调交互。所得策略在执行时以完全去中心化的方式运行,仅依赖局部观测和学到的消息。结合滚动时域推理方案,所提方法能够高效地单步推理短时域动作序列,并在通信中断时优雅降级。大量仿真结果表明,该方法具有接近专家的避碰性能,对更密集、未见过的多智能体场景具有很强的泛化能力,并且能够零样本迁移到真实机器人实验。
cs.RO / 49 / 2609.14297

Multi-Task Visual Perception Network with LLM Conditioning for Autonomous Navigation

面向自主导航的LLM条件化多任务视觉感知网络
Kumar, Praveen, Guruprasad, K. R., Sandhan, Tushar
Abstract
Long-term navigation for service robots faces crit- ical challenges like the accumulation of odometry drift and sensor error, which progressively degrade 2D maps and renders traditional path planning algorithms (e.g., A*, RRT*, DiPPer, ViT-A*) ineffective over time. To address this, we propose a user-friendly, interactive framework that eliminates the reliance on globally consistent maps. Our approach integrates visual perception with Large Language Models (LLM) to interpret user commands via text or voice. Instead of relying on a drift- prone global map, the system generates a sequential action plan based on local visual cues and egocentric geometric instructions. These action plans are executed sequentially, allowing the robot to navigate known and unknown environments safely. By reset- ting localization relative to immediate targets, our framework effectively works with a minimum accumulation drift strategy, ensuring accurate, efficient, and collision-free navigation without the maintenance overhead of traditional mapping. Experiments on real-world and simulated data have shown significant improve- ments over other methods. Our source code is publicly accessible at https://github.com/PraveenSingh24/VL-Navigation.
Chinese Translation
服务机器人的长期导航面临里程计漂移和传感器误差累积等关键挑战,这会使二维地图逐渐退化,并导致传统路径规划算法(如 A*、RRT*、DiPPer、ViT-A*)随时间推移失效。为解决该问题,我们提出了一种用户友好的交互式框架,消除了对全局一致地图的依赖。我们的方法将视觉感知与大语言模型(LLM)相结合,以通过文本或语音解析用户指令。该系统不依赖易产生漂移的全局地图,而是基于局部视觉线索和以自我为中心的几何指令生成顺序动作计划。这些动作计划被依次执行,使机器人能够安全地在已知和未知环境中导航。通过相对于即时目标重置定位,我们的框架能够以最小累积漂移策略有效工作,确保准确、高效且无碰撞的导航,同时无需传统建图的维护开销。在真实世界和仿真数据上的实验表明,该方法相较其他方法有显著提升。我们的源代码公开于 https://github.com/PraveenSingh24/VL-Navigation。
cs.RO / 50 / 2609.14310

VLBiMan++: Expanding the Generalization Boundary of Vision-Language Anchored One-Shot Bimanual Manipulation

VLBiMan++:拓展视觉语言锚定的单样本双臂操作的泛化边界
Zhou, Huayi, Gao, Wei, Han, Yiyang, Jia, Kui, Huang, Hui
Abstract
Generalizable bimanual robotic manipulation requires a reusable task prior that can persist across increasingly diverse tasks, objects, scenes, embodiments, and execution conditions, thus avoiding the prohibitive cost of large-scale teleoperated demonstrations and policy retraining. In this work, we present VLBiMan++, an extended framework that expands the generalization boundary of vision-language anchored one-shot bimanual manipulation. Starting from a single human demonstration, VLBiMan++ performs task-aware decomposition to identify reusable and adaptable skill components, and employs vision-language grounded geometric adaptation to transfer these skills to novel configurations without retraining. Building on this foundation, we systematically extend generalization along five dimensions: task generalization through diverse and long-horizon skill compositions; object generalization across unseen categories, varying geometries, and more complex articulated or deformable objects; scene generalization under clutter, occlusion, and dynamic interference; embodiment generalization across heterogeneous dual-arm robotic platforms; and deployment generalization through prolonged closed-loop execution under repeated external perturbations. To support this broader scope, we further introduce object-state-aware adaptation and lightweight trajectory optimization mechanisms that accommodate changes beyond simple rigid 6-DoF pose variations while preserving reliable bimanual coordination. Extensive real-world experiments demonstrate that VLBiMan++ maintains strong task success and adaptation capability across these increasingly challenging settings. Overall, VLBiMan++ advances one-shot bimanual manipulation from demonstrating isolated transferability toward a more systematic and scalable framework for generalization across tasks, objects, scenes, embodiments, and long-term deployment conditions.
Chinese Translation
可泛化的双臂机器人操作需要一种可复用的任务先验,能够在日益多样的任务、物体、场景、具身形态和执行条件中持续适用,从而避免大规模遥操作演示和策略重训练的高昂成本。在这项工作中,我们提出了 VLBiMan++,一个扩展框架,旨在拓展视觉语言锚定的单样本双臂操作的泛化边界。从单个人类演示出发,VLBiMan++ 执行任务感知分解,以识别可复用且可适应的技能组件,并采用基于视觉语言的几何适配,将这些技能迁移到新配置中而无需重新训练。在此基础上,我们沿五个维度系统性地扩展泛化能力:通过多样且长时程的技能组合实现任务泛化;跨未见类别、变化几何形状以及更复杂的铰接或可变形物体实现物体泛化;在杂乱、遮挡和动态干扰下实现场景泛化;跨异构双臂机器人平台实现具身泛化;通过重复外部扰动下的长时间闭环执行实现部署泛化。为支持这一更广泛的范围,我们进一步引入物体状态感知的自适应和轻量级轨迹优化机制,以应对超越简单刚性六自由度位姿变化的变化,同时保持可靠的双臂协调。大量真实世界实验表明,VLBiMan++ 在这些日益具有挑战性的设置中保持较高的任务成功率和适应能力。总体而言,VLBiMan++ 将单样本双臂操作从展示孤立的迁移能力推进为一种更系统、可扩展的框架,用于跨任务、物体、场景、具身形态和长期部署条件的泛化。
cs.RO / 51 / 2609.14343

BIG-CBF: Behavior-Imagination-Guided Control Barrier Function with Shared Uncertainty for Mobile Robot Navigation

BIG-CBF:行为想象引导的共享不确定性控制屏障函数用于移动机器人导航
Li, Shibo, Wang, Zhongcheng, Cao, Jiahe, Yang, Jianhua, Wu, Ke
Abstract
Control barrier functions (CBFs) provide a mathematically grounded framework for enforcing local collision-avoidance constraints in autonomous mobile robots, commonly through optimization-based safety filters. However, a minimum-intervention CBF filter lacks task-level maneuver awareness and may fail to select a productive avoidance direction when multiple distinct maneuvers are locally viable, leading to safe but stalled behavior in geometrically ambiguous environments. This paper presents BIG-CBF, Behavior-Imagination-Guided Control Barrier Function with shared uncertainty, a two-rate navigation architecture that separates low-rate maneuver selection from high-rate safety filtering. Over a short horizon, six closed-loop feedback behaviors are imagined and evaluated using analytic CBF compatibility together with a lightweight objective accounting for task progress, freezing, smoothness, and switching. To reduce planning-execution mismatch, the imagination and execution layers share consistent uncertainty sources for relative-motion delay, obstacle prediction, zero-order-hold motion, and command-execution residuals, while a hard CBF remains the final safety authority. In a 3,600-episode comparative benchmark across nine scenarios, BIG-CBF achieves the highest overall task success rate of 99.78% while substantially reducing downstream CBF intervention. On a physical omnidirectional robot with onboard Jetson Orin Nano computation, BIG-CBF completes all 15 evaluation runs without a recorded contact event. Matched hardware comparisons against the non-shared variant further show lower CBF intervention energy and activation frequency, supporting improved consistency between maneuver selection and safety-critical execution.
Chinese Translation
控制屏障函数(CBFs)为在自主移动机器人中强制执行局部避碰约束提供了一个有数学依据的框架,通常通过基于优化的安全滤波器来实现。然而,最小干预的CBF滤波器缺乏任务级机动意识,并且当多个不同的机动在局部可行时,可能无法选择有效的避让方向,导致在几何模糊环境中出现安全但停滞的行为。本文提出了BIG-CBF,即行为想象引导的共享不确定性控制屏障函数,一种双速率导航架构,将低速率机动选择与高速率安全滤波分离。在短时域内,想象六种闭环反馈行为,并使用解析CBF兼容性以及考虑任务进度、冻结、平滑度和切换的轻量级目标函数进行评估。为了减少规划-执行不匹配,想象层和执行层共享一致的不确定性来源,包括相对运动延迟、障碍物预测、零阶保持运动和命令执行残差,同时硬CBF仍然是最终的安全权威。在跨越九个场景的3600回合对比基准测试中,BIG-CBF实现了最高的总体任务成功率为99.78%,同时大幅减少了下游CBF干预。在配备机载Jetson Orin Nano计算的全向物理机器人上,BIG-CBF完成了所有15次评估运行,没有记录到接触事件。与非共享变体进行的匹配硬件比较进一步表明,CBF干预能量和激活频率更低,支持了机动选择与安全关键执行之间一致性的提高。
cs.RO / 52 / 2609.14376

Orientation Control of Soft Robots via Adiabatic Spectral Submanifolds

基于绝热谱子流形的软体机器人方向控制
Karakai, Aron, Kaundinya, Roshan S., Michelis, Mike Yan, Katzschmann, Robert, Haller, George
Abstract
Soft robots are commonly sought for safety-critical interactions in delicate environments, where accurate position and orientation control is imperative. Model predictive control (MPC) offers a solution, but it requires a model of the robot's infinite-dimensional nonlinear dynamics that is at once accurate and computationally cheap. Recent theory on adiabatic spectral submanifolds (aSSMs) and their applications to soft robots provide data-driven model-reduction methods to construct such models. Here, we extend these methods to identify aSSMs from enlarged observable datasets and upgrade the currently available aSSM-MPC schemes. Evaluated on a high-fidelity finite-element simulation of a pressure-actuated soft arm, our controller reduces position and orientation tracking error by more than 60% compared to existing data-driven baselines.
Chinese Translation
软体机器人通常被用于精细环境中的安全关键交互,其中准确的位置和方向控制至关重要。模型预测控制(MPC)提供了一种解决方案,但它需要机器人的无限维非线性动力学模型,该模型既要准确又要计算成本低。最近关于绝热谱子流形(aSSMs)的理论及其在软体机器人中的应用提供了数据驱动的模型降阶方法来构建此类模型。在此,我们扩展了这些方法,从扩大的可观测数据集中识别aSSMs,并升级了当前可用的aSSM-MPC方案。在压力驱动的软臂的高保真有限元仿真上评估,我们的控制器与现有的数据驱动基线相比,将位置和方向跟踪误差降低了60%以上。
cs.RO / 53 / 2609.14426

Learning-Based Dynamic Obstacle Avoidance for a UAV Using Only Three Range Sensors

基于学习的仅使用三个测距传感器的无人机动态避障
Divkoti, Mohammad Reza Ranjbar, Aguiar, A. Pedro
Abstract
We present a learning-based approach to kinodynamic online motion planning for an Unmanned Aerial Vehicle (UAV) operating at a fixed altitude in unknown dynamic environments, where real-time avoidance of both static and dynamic obstacles must be achieved under conditions of extreme partial observability. The UAV is controlled with a single degree of freedom (yaw only), resulting in constrained, nonholonomic motion similar to fixed-wing platforms. The proposed framework integrates a behavior grid map representation with Deep Reinforcement Learning (DRL), using Proximal Policy Optimization (PPO) for stable policy learning in continuous control. The key idea is the co-design of a state representation and control policy that enables reliable navigation using only three low-cost directional range sensors, without reliance on dense sensing modalities such as LiDAR or vision-based systems. The behavior grid map dynamically aggregates sparse measurements into a structured local representation that supports real-time decision-making for obstacle avoidance and target reaching. Extensive simulations across environments of varying sizes and obstacle densities demonstrate that the proposed standard and enhanced methods achieve higher success rates than PPO variants and Model Predictive Control (MPC) (94\% vs. 79--90\% in small-scale high-congestion scenarios, and 83\% vs. 62--71\% in large-scale high-congestion scenarios), while maintaining real-time performance. Real-world experiments across four scenarios further confirm practical feasibility, with consistent target-reaching behaviour and no collisions under the tested conditions.
Chinese Translation
我们提出了一种基于学习的方法,用于在未知动态环境中以固定高度运行的无人机(UAV)的运动动力学在线运动规划,其中需要在极端部分可观测的条件下实现静态和动态障碍物的实时避障。无人机仅使用一个自由度(仅偏航)进行控制,导致类似固定翼平台的受限非完整运动。提出的框架将行为网格地图表示与深度强化学习(DRL)相结合,使用近端策略优化(PPO)在连续控制中进行稳定的策略学习。关键思想是状态表示和控制策略的协同设计,仅使用三个低成本定向测距传感器即可实现可靠导航,而不依赖于激光雷达或基于视觉的系统等密集传感模式。行为网格地图将稀疏测量动态聚合成结构化的局部表示,支持用于避障和目标到达的实时决策。在不同规模和障碍物密度的环境中进行的广泛仿真表明,提出的标准方法和增强方法比PPO变体和模型预测控制(MPC)具有更高的成功率(在小规模高拥堵场景中为94% vs. 79-90%,在大规模高拥堵场景中为83% vs. 62-71%),同时保持实时性能。在四个场景中的真实世界实验进一步证实了实际可行性,在测试条件下具有一致的目标到达行为且无碰撞。
cs.RO / 54 / 2609.14432

EMoG: Emotion-Modulated Gait Generation for Expressive Humanoid Locomotion

EMoG:用于富有表现力的人形机器人运动的情绪调制步态生成
Lu, Yi, Jiang, Tianhao, Tian, Honglong, Zhang, Yumeng, Zhao, Qingrui, Wang, Zhengtao, Long, Xiao-Xiao, Shen, Qiu, Cao, Xun
Abstract
Existing humanoid locomotion systems primarily focus on stability and task execution, while integrating expressiveness with explicit locomotion control remains challenging. We propose EMoG, an emotion-modulated gait generation framework for expressive humanoid locomotion. EMoG introduces an emotional-style code with continuously adjustable intensity. Conditioned on this code and physical commands, a lightweight MLP generates expressive, command-consistent periodic gait trajectories in real time, which are tracked by a unified reinforcement learning policy for physical execution. To support training, we collect a large-scale emotion-annotated gait dataset from professional performers and develop an automated pipeline to extract physically consistent periodic gait cycles. EMoG also integrates an LLM-based parser that converts free-form language into emotional style and motion parameters for interactive control. Experiments demonstrate continuous gait-style modulation with perceptible expressive cues while maintaining command tracking. EMoG provides a practical approach to parameterized emotional-style walking for human-robot interaction.
Chinese Translation
现有的人形机器人运动系统主要关注稳定性和任务执行,而将表现力与显式运动控制相结合仍然具有挑战性。我们提出了EMoG,一种用于富有表现力的人形机器人运动的情绪调制步态生成框架。EMoG引入了一种强度连续可调的情绪风格编码。以该编码和物理命令为条件,一个轻量级MLP实时生成富有表现力且与命令一致的周期性步态轨迹,这些轨迹由统一的强化学习策略进行跟踪以执行物理动作。为了支持训练,我们从专业表演者那里收集了一个大规模的情绪标注步态数据集,并开发了一个自动化流水线来提取物理一致的周期性步态周期。EMoG还集成了一个基于LLM的解析器,可将自由形式的语言转换为情绪风格和运动参数,以实现交互式控制。实验表明,在保持命令跟踪的同时,实现了连续的步态风格调制,并具有可感知的表现力线索。EMoG为面向人机交互的参数化情绪风格行走提供了一种实用方法。
cs.RO / 55 / 2609.14450

EcoBoat: Design and Experimental Validation of an Autonomous Body-Board Boat For Cleaning Water Bodies

EcoBoat:用于清理水体的自主板体船的设计与实验验证
Ansari, M. Aman, Khan, Saifullah, Kulkarni, Rahul, Sujit, PB
Abstract
Cleaning water bodies such as swimming pools and lakes typically demands significant manual effort or reliance on costly, sensor-intensive robotic systems. This paper presents EcoBoat, a low-cost autonomous surface vehicle built on a modified hull, designed to collect floating debris in both indoor and outdoor water bodies. For indoor environments, EcoBoat uses ultrasonic sensors to detect boundaries and obstacles, combining random-walk motion with boundary-following behavior. For outdoor environments, it relies on GPS-based geofencing paired with a random-walk strategy for area coverage. A key design feature is a Tesla-valve-inspired collection basket that allows debris intake during forward motion while preventing its escape during turning or braking. Field experiments in pools and lakes validated the design through iterative refinement, demonstrating that a minimalist sensing and computation approach can achieve effective, versatile debris collection across diverse water bodies.
Chinese Translation
清理游泳池、湖泊等水体通常需要大量人工投入,或依赖成本高昂、传感器密集的机器人系统。本文提出了 EcoBoat,一种基于改装船体构建的低成本自主水面航行器,旨在收集室内和室外水体中的漂浮垃圾。在室内环境中,EcoBoat 使用超声波传感器检测边界和障碍物,将随机游走运动与边界跟随行为相结合。在室外环境中,它依靠基于 GPS 的地理围栏,并结合随机游走策略以实现区域覆盖。一个关键设计特征是受特斯拉阀启发的收集篮,其允许在向前运动时吸入垃圾,同时在转向或制动时防止垃圾逃逸。在游泳池和湖泊中进行的现场实验通过迭代改进验证了该设计,表明极简的传感与计算方法能够在不同水体中实现有效且多功能的垃圾收集。
cs.RO / 56 / 2609.14539

Fault Diagnosis for Underwater Vehicles using Moving Horizon Estimation and Gaussian Processes

基于移动 horizon 估计与高斯过程的水下航行器故障诊断
Panetsos, Fotis, Kyriakopoulos, Kostas J.
Abstract
This work proposes a model-based fault detection and diagnosis framework for underwater vehicles subject to actuator faults that explicitly accounts for the presence of unmodeled dynamics. To this end, a Moving Horizon Estimator (MHE) is developed to estimate the lumped disturbance, capturing both unmodeled and fault effects. Gaussian Processes (GPs) are employed to approximate the unmodeled dynamics, providing predictions of the corresponding mean and uncertainty across diverse operating conditions. During online operation, the residual between the MHE lumped disturbance estimate and the GP prediction is evaluated using a Generalized Likelihood Ratio Test. By incorporating GP-based predictions within the diagnostic framework, robustness to unmodeled dynamics is achieved, enabling effective fault detection and isolation as well as accurate quantitative estimation of fault magnitude. The proposed methodology is experimentally validated in a laboratory water tank, demonstrating reliable diagnostic performance under both open-loop and closed-loop control.
Chinese Translation
本文提出了一种基于模型的水下航行器故障检测与诊断框架,针对执行器故障,并明确考虑了未建模动态的存在。为此,开发了一个移动 horizon 估计器 (MHE),用于估计集总扰动,同时捕获未建模动态和故障的影响。采用高斯过程 (GPs) 来近似未建模动态,提供在不同操作条件下相应的均值和不确定性的预测。在线运行期间,使用广义似然比检验评估 MHE 集总扰动估计与 GP 预测之间的残差。通过在诊断框架中融入基于 GP 的预测,实现了对未建模动态的鲁棒性,从而能够进行有效的故障检测与隔离,并准确量化估计故障幅值。所提出的方法在实验室水箱中进行了实验验证,在开环和闭环控制下均展示了可靠的诊断性能。
cs.RO / 57 / 2609.14543

NavPatch: Evidence-Guided Object-Level Costmap Correction with Vision-Language Models

NavPatch:基于视觉语言模型的证据引导对象级代价地图校正
Sun, Shiji, Tao, Xingyu, Wang, Hao, Wang, Ling, Chen, Zhengyi
Abstract
Mobile robots typically rely on geometric maps for obstacle avoidance and path planning, but the resulting obstacle representation does not always match how an object should affect navigation. A low lying cable may be missed, a flexible curtain may create spurious blockage, and a traffic cone may require an exclusion region larger than its observed footprint. We present NavPatch, an object level correction layer that assigns ADD, REMOVE, or EXTEND to navigation relevant object categories through periodic scene understanding with a vision-language model. Open vocabulary grounding localizes object instances, and LiDAR and RGB-D observations provide 3D support. Observation quality filtering and cross frame maintenance determine when each correction patch is committed, replaced, or revoked. In 50 real robot trials across five layouts, NavPatch achieves an overall success rate of 86.0%. An ablation study of four configurations with 200 runs in total shows that NavPatch improves the success rate from 70.0% to 86.0% and reduces the false commit rate from 68.4% to 40.7% compared with updates based only on the current observation.
Chinese Translation
移动机器人通常依赖几何地图进行避障和路径规划,但由此产生的障碍物表示并不总是与物体应如何影响导航相匹配。低矮的线缆可能被遗漏,柔性帘布可能产生虚假的阻塞,而交通锥可能需要比其观测足迹更大的排除区域。我们提出NavPatch,一个对象级校正层,通过使用视觉语言模型进行周期性场景理解,为导航相关的对象类别分配ADD、REMOVE或EXTEND操作。开放词汇定位对物体实例进行定位,而LiDAR和RGB-D观测提供3D支持。观测质量过滤和跨帧维护决定每个校正补丁何时被提交、替换或撤销。在跨越五种布局的50次真实机器人试验中,NavPatch实现了86.0%的总体成功率。对四种配置共200次运行的消融研究表明,与仅基于当前观测的更新相比,NavPatch将成功率从70.0%提高到86.0%,并将错误提交率从68.4%降低到40.7%。
cs.RO / 58 / 2609.14558

Language-Grounded Semantic Target Navigation for Autonomous Surface Vehicles

面向自主水面艇的语言接地语义目标导航
Lin, Yuqing, Kim, Youngroung
Abstract
Autonomous Surface Vehicles (ASVs) are increasingly expected to operate in ports and harbour environments, where operators may specify navigation targets through language-based descriptions rather than predefined coordinates or fixed target identifiers. However, existing ASV navigation methods mainly execute predefined geometric goals or task-specific objectives and give limited attention to language-grounded target specification. This study proposes Semantically Grounded Navigation (SGNav), a framework that enables an ASV to identify and approach a maritime target from an operator-provided description. SGNav integrates text-guided semantic grounding, harbour-aware candidate filtering, CLIP-based semantic verification, grounded target control-state construction, and Proximal Policy Optimisation-based closed-loop control. It grounds the target description in onboard RGB observations, suppresses visually or semantically irrelevant distractors, and converts the selected target into a compact control-oriented representation for policy execution. Experiments in simulated port environments show that SGNav achieves success rates of $97.0\pm1.2\%$, $92.0\pm1.5\%$, and $90.0\pm1.8\%$ across three representative target-reaching tasks, with semantic target accuracy above $97\%$ and wrong-target rates below $3\%$. SGNav also maintains $97.7$--$98.7\%$ success rates across held-out port layouts. In the Task~3 ablation study, removing harbour-aware filtering or semantic consistency reduces the success rate to $40.4\pm2.6\%$ and $50.4\pm3.1\%$, respectively. These findings demonstrate the importance of semantic grounding, harbour-aware filtering, and semantic verification for reliable language-grounded ASV navigation. These results indicate that the proposed perception-to-control framework can support language-grounded target approach manoeuvres of ASV under the complex port environments.
Chinese Translation
自主水面艇(ASVs)越来越多地被期望在港口和港湾环境中运行,在这些环境中,操作员可能通过基于语言的描述来指定导航目标,而不是预定义的坐标或固定的目标标识符。然而,现有的ASV导航方法主要执行预定义的几何目标或特定任务目标,对语言接地的目标指定关注有限。本研究提出了语义接地导航(SGNav),一个使ASV能够根据操作员提供的描述识别并接近海上目标的框架。SGNav集成了文本引导的语义接地、港口感知的候选过滤、基于CLIP的语义验证、接地的目标控制状态构建以及基于近端策略优化的闭环控制。它将目标描述接地于船载RGB观测,抑制视觉或语义上无关的干扰物,并将所选目标转换为紧凑的面向控制的表示以用于策略执行。在模拟港口环境中的实验表明,SGNav在三个代表性的目标到达任务中分别达到了97.0±1.2%、92.0±1.5%和90.0±1.8%的成功率,语义目标准确率高于97%,错误目标率低于3%。SGNav在留出的港口布局中也保持了97.7%--98.7%的成功率。在任务3的消融研究中,移除港口感知过滤或语义一致性分别将成功率降低至40.4±2.6%和50.4±3.1%。这些发现证明了语义接地、港口感知过滤和语义验证对于可靠的语言接地ASV导航的重要性。这些结果表明,所提出的从感知到控制的框架能够支持ASV在复杂港口环境下的语言接地目标接近机动。
cs.RO / 59 / 2609.14561

GLAM: Training a latent world model over global spatiotemporal memory for active exploration and navigation

GLAM:基于全局时空记忆训练用于主动探索与导航的潜在世界模型
Ieong, I-Tak, Feng, Ruizhi, Lu, Zhaoyang, Cao, Yifei, Zhao, Jiayao, Li, Leon, Zhu, Senhua, Ding, Wenbo
Abstract
Active exploration and semantic navigation require an embodied agent to build memory from partial observations, predict how the evolution of observed spatial memory may support future motion, and convert that prediction into actionable plans. We present GLAM, a goal-conditioned latent world model trained over global spatiotemporal memory, and GLAM NAV, the complete navigation system built around it. Given historical map tokens, a navigation goal, and the current robot pose, GLAM jointly predicts future map representations and robot-centric waypoint latents, allowing future spatial context and navigation intent to be inferred in a shared representation space. The model follows a JEPA-like latent prediction paradigm, operates directly on map-level latent tokens rather than RGB reconstruction, and uses a pretrained waypoint encoder-decoder to supervise and decode navigation plans within GLAM NAV. Training data are collected by replaying ObjectNav expert trajectories in Habitat over HM3D v0.2 scene assets and slicing them into multi-timescale prediction samples. On a controlled HM3D-ObjectNav subset reproduction setting, GLAM NAV improves over a reproduced BSC-Nav baseline in both success rate and success weighted by path length.
Chinese Translation
主动探索和语义导航要求具身智能体从部分观测中构建记忆,预测所观测到的空间记忆的演化如何支持未来运动,并将该预测转化为可执行的计划。我们提出了 GLAM,一个在全局时空记忆上训练的目标条件潜在世界模型,以及 GLAM NAV,围绕它构建的完整导航系统。给定历史地图 token、导航目标和当前机器人位姿,GLAM 联合预测未来地图表示和以机器人为中心的路点潜在表示,从而能够在共享表示空间中推断未来的空间上下文和导航意图。该模型遵循类 JEPA 的潜在预测范式,直接在地图级潜在 token 上操作而非 RGB 重建,并使用预训练的路点编码器-解码器在 GLAM NAV 内监督和解码导航计划。训练数据通过在 Habitat 中回放基于 HM3D v0.2 场景资产的 ObjectNav 专家轨迹,并将其切分为多时间尺度预测样本而收集。在受控的 HM3D-ObjectNav 子集复现设置中,GLAM NAV 在成功率和路径长度加权成功率上均优于复现的 BSC-Nav 基线。
cs.RO / 60 / 2609.14567

Learning Multi-Agent Task Assignment and Navigation in the Factory: from Simulation to Real Robots

在工厂中学习多智能体任务分配与导航:从仿真到真实机器人
Abdalwhab, Abdalwhab Bakheet Mohamed, Beltrame, Giovanni, St-Onge, David
Abstract
Reinforcement learning (RL) has shown considerable promise for robotic decision-making, yet deploying multi-agent RL (MARL) on physical multi-robot systems in industrial environments remains challenging. This paper investigates the real-world applicability of decentralized MARL for multi-robot multi-machine tending. We propose Feature-fusion Multi-Agent Proximal Policy Optimization (FMAPPO), which fuses 2D LiDAR measurements with task-specific state information to enable safe decentralized multi-robot task assignment and navigation. A complete simulation-to-reality pipeline was developed using high-fidelity robotic simulation and ROS2 and deployed on physical mobile-manipulator platforms operating under realistic real-world conditions, with the robotic arms disabled during the experiments. We further investigate the sensitivity of the learned policy to command update frequency, an important consideration for real-world deployment. Comparative evaluation in simulation demonstrated that FMAPPO significantly outperformed state-of-the-art baselines with a large effect size, achieving improvements of 106\% and 21\% in parts delivery and 48\% and 11\% in parts collection over MAPPO and SMAPPO, respectively. FMAPPO also increased machine utilization by 31 and 10 percentage points, respectively, while reducing collisions by 18\% and 15\% and increasing the safety score by 14 and 6 percentage points compared with MAPPO and SMAPPO, respectively. Furthermore, real-world experiments demonstrated that the learned decentralized policies can coordinate multiple robots to service multiple machines while maintaining safe operation under real-world sensing and control constraints. Videos of the real-world experiment are available online https://anonymouspapers123.github.io/FMAPPO/.
Chinese Translation
强化学习(RL)在机器人决策方面已展现出巨大潜力,然而在工业环境中将多智能体强化学习(MARL)部署到物理多机器人系统上仍然具有挑战性。本文研究了去中心化多智能体强化学习在多机器人多机器看护中的实际适用性。我们提出了特征融合多智能体近端策略优化(FMAPPO),它将2D激光雷达测量与任务特定状态信息融合,以实现安全的去中心化多机器人任务分配与导航。利用高保真机器人仿真和ROS2开发了一套完整的从仿真到现实的流程,并部署在真实世界条件下运行的物理移动操作机器人平台上,实验期间机械臂被禁用。我们进一步研究了学习策略对指令更新频率的敏感性,这是实际部署中的一个重要考虑因素。仿真中的对比评估表明,FMAPPO以较大的效应量显著优于最先进的基线,在零件交付方面比MAPPO和SMAPPO分别提高了106%和21%,在零件收集方面分别提高了48%和11%。与MAPPO和SMAPPO相比,FMAPPO还将机器利用率分别提高了31和10个百分点,同时将碰撞分别减少了18%和15%,并将安全分数分别提高了14和6个百分点。此外,真实世界实验表明,学习到的去中心化策略能够协调多个机器人服务多台机器,同时在真实世界的感知和控制约束下保持安全运行。真实世界实验的视频可在线获取:https://anonymouspapers123.github.io/FMAPPO/。
cs.RO / 61 / 2609.14589

Robotic Servo Tracking of Moving Targets with Dynamic Imitation Constraints

动态模仿约束下的机器人移动目标伺服跟踪
Luo, Yazhe, Ruan, Sipu, Li, Yifei, Chen, Diansheng
Abstract
Imposing explicit trajectory constraints in robot visual servoing remains challenging. Existing tracking methods achieve fast responses by mapping visual residuals to control velocities, but they have weak constraints on the intermediate motion process, which lead to trajectory discontinuity, oscillation, or conservative behaviors. To enable constrained tracking for moving targets, this paper proposes a servo tracking method based on imitation trajectory constraints. A dynamic model describing the robot approaching a moving target is formulated and analyzed for convergence. A time-scalable deformation mechanism and a trajectory modulation incorporating shape and amplitude components are introduced to generate a series of trajectories in real time, from which tracking points are adaptively determined to form dynamic constraints. The robot velocity is then computed from target pose differentials or tracked key features to follow the constrained trajectory. Simulation and real-world experiments demonstrate that the proposed method can achieve dynamic obstacle avoidance and high-precision convergence compared with several state-of-the-art methods in complex environments.
Chinese Translation
在机器人视觉伺服中施加显式的轨迹约束仍然具有挑战性。现有的跟踪方法通过将视觉残差映射到控制速度来实现快速响应,但它们在中间运动过程上的约束较弱,这导致轨迹不连续、振荡或保守行为。为了实现移动目标的约束跟踪,本文提出了一种基于模仿轨迹约束的伺服跟踪方法。建立了一个描述机器人接近移动目标的动态模型,并分析了其收敛性。引入了时间可伸缩的变形机制和融合形状与幅度分量的轨迹调制,以实时生成一系列轨迹,从中自适应地确定跟踪点以形成动态约束。然后,根据目标位姿微分或跟踪的关键特征计算机器人速度,以跟随受约束的轨迹。仿真和真实世界实验表明,与几种最先进的方法相比,所提方法可以在复杂环境中实现动态避障和高精度收敛。
cs.RO / 62 / 2609.14633

REVOLVE: An Automated Closed-Loop Framework for Evolving Robot Manipulation with Minimal Human Intervention

REVOLVE:一种在最小人工干预下进化机器人操作的自动闭环框架
Liu, Hanyu, Li, Qian, Ding, Yizhu, Wen, Jiayi, Ren, Keqiang, Ma, Yunsheng, Jian, Tao, Wang, Zhihua, Yu, Zhuofan, Li, Xinran, Song, Zhigong
Abstract
Recent advances in data-driven robot manipulation policies have substantially improved task execution and generalization. However, real-world deployment still relies heavily on humans for failure assessment, correction, and environment reset, while models often fail to continually learn from failures and corrective experience. We present REVOLVE (Robot Evolving via Orchestrated Loops, Verification, and Experience), an automated closed-loop framework for evolving robot manipulation with minimal human intervention. Built on a unified software platform, REVOLVE integrates data collection, policy training and deployment, failure recovery, and continual learning into a single closed-loop workflow. Its Automated Reset and Collection (ARC) architecture automatically resets the environment and intervenes to correct policy failures. Dual-Loop Evolution (DLE) continually improves the manipulation policy and agent by feeding real-world interaction and failure--correction data back into policy learning and using an external mismatch memory to refine agent judgments. Experiments across four real-world manipulation tasks show that, after five iterations, REVOLVE improves average policy success rate by 18.5% and agent judgment accuracy by 8.5%, while reducing human effort in data collection and deployment testing by 94.4% and 95.1%, respectively. These results demonstrate that REVOLVE transforms real-world deployment into a closed-loop learning process that continually accumulates and uses execution experience, enabling continual evolution of both the policy and supervisory model with substantially less human intervention.
Chinese Translation
近年来,数据驱动的机器人操作策略取得了显著进展,大幅提升了任务执行和泛化能力。然而,真实世界部署仍然严重依赖人工进行失败评估、纠正和环境重置,而模型通常无法持续从失败和纠正经验中学习。我们提出了REVOLVE(Robot Evolving via Orchestrated Loops, Verification, and Experience),一种以最小人工干预进化机器人操作的自动闭环框架。REVOLVE构建在统一软件平台上,将数据收集、策略训练与部署、失败恢复和持续学习集成到一个闭环工作流中。其自动重置与收集(ARC)架构自动重置环境并进行干预以纠正策略失败。双环进化(DLE)通过将真实世界交互和失败-纠正数据反馈到策略学习,并使用外部不匹配记忆来优化智能体判断,从而持续改进操作策略和智能体。在四个真实世界操作任务上的实验表明,经过五次迭代后,REVOLVE将平均策略成功率提高了18.5%,智能体判断准确率提高了8.5%,同时将数据收集和部署测试中的人工工作量分别减少了94.4%和95.1%。这些结果表明,REVOLVE将真实世界部署转变为闭环学习过程,持续积累并利用执行经验,以显著减少的人工干预实现策略和监督模型的持续进化。
cs.RO / 63 / 2609.14647

Skill Composition for Legged Robot Reinforcement Learning

腿式机器人强化学习的技能组合
Gigliotti, Daniel, Maiorana, Flavio, Patrizi, Fabio, Iocchi, Luca
Abstract
Robots, and humanoid robots in particular, are increasingly competent at individual behaviors, each obtained by training a specialized controller. A specialized skill is quick to train, converges reliably because the problem it faces is narrow, and can be validated on its own, none of which is true of a single end-to-end policy asked to cover everything. What remains fragile is the transition between them. We argue that the composition of independent sub-policies deserves to be treated as a research problem in its own right, rather than as an implementation detail left to whatever mechanism happens to be at hand. Reliable composition is what turns a collection of separate skills into a repertoire that can be used, extended and shared. More fundamentally, if control can be passed between specialized policies safely, and at any moment, the choice of what the robot should do next can be delegated to a component of an entirely different nature, such as a planner, an automaton or a symbolic controller, whose behavior can be inspected in advance. The policies would then only ever have to act, and what the robot can be trusted to do would become verifiable.
Chinese Translation
机器人,尤其是人形机器人,在单个行为上越来越有能力,每个行为都是通过训练一个专用控制器获得的。专用技能训练速度快,收敛可靠,因为它面临的问题范围狭窄,并且可以独立验证;而一个被要求覆盖所有任务的单一端到端策略则不具备这些优点。仍然脆弱的是它们之间的过渡。我们认为,独立子策略的组合本身值得作为一个研究问题来对待,而不是作为实现细节,留给手头恰好可用的任何机制。可靠的组合才能将一组分离的技能转变为一个可以使用、扩展和共享的技能库。更根本地说,如果控制可以在专用策略之间安全地、随时传递,那么机器人下一步应该做什么的选择就可以委托给一个性质完全不同的组件,例如规划器、自动机或符号控制器,其行为可以事先检查。然后策略只需执行动作,机器人可以被信任做什么将变得可验证。
cs.RO / 64 / 2609.14710

A Novel Robot-Assisted Learning Pedagogy for Children with ASD

一种针对ASD儿童的新型机器人辅助学习教学法
Boccanfuso, Laura, Barney, Erin, Mademtzi, Marilena, Foster, Claire, Wang, Quan, Torres, Colette, Chen, Lisa, Scassellati, Brian, Ventola, Pamela, Shic, Frederick
Abstract
Interaction paradigms used in robot-assisted autism intervention have historically employed robots as teachers, clinical assistants, or more-abled peers to promote a variety of social skills. These modalities often leverage the expertise of trained practitioners to ensure that child-robot interactions are productive or clinically grounded to yield positive therapeutic benefits for children across the autism spectrum. Yet, despite the fact that the majority of children with autism spectrum disorder (ASD) attend mainstream schools and spend 80% or more of their time in the general classroom [27], there is a paucity of research incorporating validated classroom teaching pedagogies into robot-assisted autism interventions. In this work, we introduce a novel teaching methodology for advancing social skills in school-aged children with ASD. We evaluate the effectiveness of a novel robot-assisted autism intervention which incorporates the learning-by-teaching pedagogy and explores the comparative benefits of employing a robot versus a human confederate for improved performance on a set of social skills tasks. Results show that 80% of study participants performed better in the robot condition (mean performance in the robot condition=63%, mean performance in the confederate condition=37%), irrespective of the scenario order. Further, 90% of all participants were significantly more engaged in the robot condition (mean engagement: robot=61%, confederate=32%) and, while the effect did not result in the confederate condition, analyses indicate that overall engagement in the robot condition contributed to improved performance. These results suggest that robots employed in a learning-by-teaching context may help enhance engagement and improve performance on a simple social skills task for children with ASD.
Chinese Translation
在机器人辅助自闭症干预中使用的交互范式,历史上一直将机器人用作教师、临床助理或能力更强的同伴,以促进各种社交技能。这些模式通常利用训练有素的从业者的专业知识,确保儿童与机器人的互动富有成效或有临床依据,从而为整个自闭症谱系的儿童带来积极的治疗效果。然而,尽管大多数自闭症谱系障碍(ASD)儿童就读于主流学校,并在普通教室中度过80%或更多的时间[27],但将经过验证的课堂教学法纳入机器人辅助自闭症干预的研究却很少。在这项工作中,我们介绍了一种新颖的教学方法,用于提高学龄ASD儿童的社交技能。我们评估了一种新颖的机器人辅助自闭症干预的有效性,该干预结合了“通过教学学习”的教学法,并探讨了使用机器人对比人类同盟者在一系列社交技能任务中提高表现的相对优势。结果表明,80%的研究参与者在机器人条件下表现更好(机器人条件下的平均表现=63%,同盟者条件下的平均表现=37%),无论场景顺序如何。此外,90%的所有参与者在机器人条件下明显更投入(平均投入度:机器人=61%,同盟者=32%),虽然这种效应在同盟者条件下没有出现,但分析表明,机器人条件下的整体投入有助于提高表现。这些结果表明,在“通过教学学习”的情境中使用机器人可能有助于增强ASD儿童的投入度,并改善他们在简单社交技能任务中的表现。
cs.RO / 65 / 2609.14748

Beyond Dead Reckoning: A Point of View on Camera--DAS--GNSS Continuity in Road Tunnels

超越航位推算:道路隧道中摄像头--DAS--GNSS连续性的观点
Saoud, Lyes Saad, Rakha, Hesham A., Jaber, Mona, Ayyash, Moussa
Abstract
Road tunnels remove satellite visibility where connected and automated vehicles still require continuous, attributable, and integrity-bounded positioning. This Point of View argues that tunnel localization should be treated as infrastructure-assisted cross-modal track continuity, not as extrapolation from the last trusted satellite fix. A trusted portal satellite solution provides the global anchor, distributed acoustic sensing (DAS) continuous motion evidence, cameras sparse identity and lane anchors, onboard sensing short-term dynamics, and edge computing association, fusion, integrity monitoring, and guarded reacquisition. Ground truth is reserved for offline calibration and validation. The article develops a falsifiable research and deployment agenda for progressing from synchronized multimodal evidence to validated, integrity-aware tunnel positioning continuity.
Chinese Translation
道路隧道消除了卫星可见性,而网联自动驾驶汽车仍然需要连续、可归因且完整性有界的定位。这篇观点文章认为,隧道定位应被视为基础设施辅助的跨模态轨迹连续性,而不是从最后一次可信卫星定位进行外推。一个可信的隧道口卫星解提供全局锚点,分布式声学传感(DAS)提供连续运动证据,摄像头提供稀疏的身份和车道锚点,车载传感提供短期动态,边缘计算进行关联、融合、完整性监测和受保护的重捕获。真值保留用于离线校准和验证。本文提出了一个可证伪的研究与部署议程,以从同步多模态证据推进到经过验证的、完整性感知的隧道定位连续性。
cs.RO / 66 / 2609.14756

An Adaptive Fixed-Time Line-of-Sight Guidance Scheme for 3D Path Following of Underwater Vehicles: Theory and Experiment

水下航行器三维路径跟踪的自适应固定时间视线制导方案:理论与实验
Yang, Hanzhi, Chavez-Galaviz, Jalil, Mahmoudian, Nina
Abstract
Reliable path tracking is crucial for autonomous underwater vehicles (AUVs) operating in dynamic and uncertain marine environments. However, traditional line-of-sight (LOS) guidance methods rely on asymptotic convergence, resulting in slow disturbance recovery and unpredictable tracking performance. Existing robust control methods typically require modifications to the underlying vehicle controller, limiting their practical application on commercial AUV platforms. This paper proposes a robust fixed-time adaptive LOS guidance framework for 3D path tracking for AUVs. By combining fixed-time stability theory with LOS guidance, this method guarantees path tracking convergence within a preset time range, with the convergence time independent of initial conditions. Furthermore, a fixed-time adaptive estimator is developed to rapidly compensate for time-varying sideslip disturbances caused by ocean currents. A time-varying look-ahead mechanism is also introduced to improve tracking performance on curved paths. Lyapunov analysis proves the fixed-time stability of the proposed framework, and numerical simulations and physical experiments demonstrate that, compared to state-of-the-art adaptive LOS methods, this framework exhibits superior tracking accuracy, convergence speed, and anti-interference capability. In simulation, the time-varying look-ahead variant reduced cross-track and vertical-track RMSE by 69.37\% and 67.46\%, respectively, during curved-path tracking. In field experiments with an Iver 3 AUV, the proposed fixed-time guidance reduced average tracking error by 56.35\% in straight-path evaluation and 27.59\% in curved-path evaluation compared with conventional adaptive LOS guidance.The proposed method provides a practical guidance-level solution for achieving reliable autonomous navigation of AUVs in complex marine environments.
Chinese Translation
可靠的路径跟踪对于在动态和不确定的海洋环境中运行的自主水下航行器(AUV)至关重要。然而,传统的视线(LOS)制导方法依赖于渐近收敛,导致扰动恢复缓慢且跟踪性能不可预测。现有的鲁棒控制方法通常需要修改底层航行器控制器,限制了其在商业AUV平台上的实际应用。本文提出了一种用于AUV三维路径跟踪的鲁棒固定时间自适应LOS制导框架。通过将固定时间稳定性理论与LOS制导相结合,该方法保证了在预设时间范围内实现路径跟踪收敛,且收敛时间与初始条件无关。此外,开发了一种固定时间自适应估计器,以快速补偿由洋流引起的时变侧滑扰动。还引入了时变前视机制以提高曲线路径上的跟踪性能。李雅普诺夫分析证明了所提出框架的固定时间稳定性,数值仿真和物理实验表明,与最先进的自适应LOS方法相比,该框架具有更优越的跟踪精度、收敛速度和抗干扰能力。在仿真中,时变前视变体在曲线路径跟踪期间将横向和垂直方向的RMSE分别降低了69.37%和67.46%。在Iver 3 AUV的现场实验中,与传统的自适应LOS制导相比,所提出的固定时间制导在直线路径评估中将平均跟踪误差降低了56.35%,在曲线路径评估中降低了27.59%。所提出的方法为实现AUV在复杂海洋环境中的可靠自主导航提供了一种实用的制导级解决方案。
cs.RO / 67 / 2609.14765

A Personalized Dynamic Balance Evaluation Paradigm for Hip Exoskeleton-Assisted Walking under Unexpected Ground Perturbations

突发地面扰动下髋关节外骨骼辅助行走的个性化动态平衡评估范式
Chen, Yun, Akinniyi, Oluwasegun T., Zhang, Qiang
Abstract
Hip exoskeletons may improve recovery from unexpected gait perturbations, yet personalizing assistance remains difficult because balance is multidimensional and human-in-the-loop experiments are small-sample and noisy. We present a participant-specific composite balance cost that integrates seven biomechanical sub-metrics spanning margin of stability, center-of-mass dynamics, and whole-body angular momentum. The sub-metrics are converted to direction-aligned, dimensionless cost features, and nonnegative fusion weights are learned on the simplex. Coupled with an empirical-Bayes hierarchical model, the learned-composite selector estimates each tested condition's posterior probability of being best, P(best), and a high-probability candidate set with size $K_{0.8}$. The framework was evaluated with three participants walking at 1.1 m/s during unilateral belt-slip perturbations across 46 hip-assistance conditions. In the full-budget analysis (B = 4 repeats per condition), the selector concentrated 80% of the posterior probability within 1 to 5 of 46 conditions, compared with 2 to 12 for equal-weight fusion and 4 to 37 for principal component analysis fusion. This smaller candidate set could shorten personalization experiments and limit participants' exposure to repeated perturbations in future studies. Selected-condition trials showed lower observed composite costs than no-torque trials, with nominal p < 0.05 for P2 and P3. Leave-one-repeat-out refits yielded positive mean held-out rank correlations for all participants and moderate stability of the learned weights and candidate sets. These proof-of-concept results support participant-specific composite balance evaluation for candidate selection in perturbation-based human-in-the-loop experiments.
Chinese Translation
髋关节外骨骼可能改善对意外步态扰动的恢复,但个性化辅助仍然困难,因为平衡是多维的,且人在回路实验样本小且噪声大。我们提出了一种参与者特定的复合平衡成本,它整合了七个生物力学子指标,涵盖稳定裕度、质心动力学和全身角动量。这些子指标被转换为方向对齐的无量纲成本特征,并在单纯形上学习非负融合权重。结合经验贝叶斯分层模型,学习复合选择器估计每个测试条件为最佳的后验概率,P(best),以及大小为 $K_{0.8}$ 的高概率候选集。该框架在3名参与者以1.1 m/s行走时,在单侧皮带滑移扰动下,跨越46种髋关节辅助条件进行了评估。在全预算分析(B = 4,每种条件重复4次)中,选择器将80%的后验概率集中在46种条件中的1到5种,而等权融合为2到12种,主成分分析融合为4到37种。这种更小的候选集可以缩短个性化实验,并在未来研究中限制参与者暴露于重复扰动。选定条件试验显示,观察到的复合成本低于无扭矩试验,P2和P3的名义p < 0.05。留一次重复重新拟合对所有参与者产生了正的留出秩相关均值,并且学习到的权重和候选集具有中等稳定性。这些概念验证结果支持在基于扰动的人在回路实验中,用于候选选择的参与者特定复合平衡评估。
cs.RO / 68 / 2609.14783

Language-Guided Representation Learning for Robust Cross-Sensor Material Recognition

面向鲁棒跨传感器材料识别的语言引导表示学习
Mohsan, Mashood M., Din, Muhayy Ud, Xu, Binzhao, Abubakar, Ahmad, Hussain, Irfan
Abstract
Robots need touch to manipulate objects safely and reliably, as many properties, such as softness, texture, and contact stability, are hard to infer from vision alone. However, vision-based tactile sensors yield different observations of the same material due to variations in optics, elastomer properties, and illumination, leading to poor generalization when trained on a single or multiple sensors. We propose a language-guided distillation framework for learning sensor-robust tactile representations. Language encodes high-level semantic properties of touch (e.g., rough, soft, slippery) that remain invariant across sensing hardware, providing a natural sensor-agnostic supervisory signal. We construct a 39K-sample touch-language dataset with human-annotated material labels and train a tactile encoder to align sensor-specific tactile images with language embeddings in a shared semantic space. We evaluate our approach for few-shot learning and cross-sensor transfer and benchmark it on six existing tactile datasets. Our method achieves 95% accuracy in the 100-shot setting, improves cross-sensor transfer by an average of 13.3% accuracy, and yields up to 19% accuracy gains across six existing tactile datasets. These results demonstrate that language-guided distillation enables scalable and hardware-agnostic tactile representation learning. Code and dataset are available at https://mashood3624.github.io/Language_Tactile/
Chinese Translation
机器人需要触觉来安全可靠地操作物体,因为许多属性,如柔软度、纹理和接触稳定性,仅从视觉难以推断。然而,基于视觉的触觉传感器由于光学、弹性体特性和光照的变化,对相同材料产生不同的观察结果,导致在单个或多个传感器上训练时泛化能力差。我们提出了一种语言引导的蒸馏框架,用于学习传感器鲁棒的触觉表示。语言编码了触觉的高层语义属性(例如,粗糙、柔软、光滑),这些属性在不同传感硬件之间保持不变,提供了一种天然的传感器无关的监督信号。我们构建了一个包含39K样本的触觉-语言数据集,带有人工标注的材料标签,并训练触觉编码器将传感器特定的触觉图像与共享语义空间中的语言嵌入对齐。我们在小样本学习和跨传感器迁移任务上评估了我们的方法,并在六个现有的触觉数据集上进行了基准测试。我们的方法在100-shot设置下达到了95%的准确率,将跨传感器迁移的平均准确率提高了13.3%,并在六个现有触觉数据集上获得了高达19%的准确率提升。这些结果表明,语言引导的蒸馏能够实现可扩展且硬件无关的触觉表示学习。代码和数据集可在 https://mashood3624.github.io/Language_Tactile/ 获取。
cs.RO / 69 / 2609.14806

Belief-Adaptive Online Autonomy for Quadrotor UAV Navigation under GNSS Degradation in Urban Environments

城市环境下GNSS退化时四旋翼无人机导航的信念自适应在线自主性
Panda, Deepak Kumar, Guo, Weisi
Abstract
Reliable online autonomy is critical for quadrotor operation in urban airspaces, where global navigation satellite systems (GNSS) measurements suffer from multipath, blockage, and latency issues, introducing non-stationary, temporally correlated errors that degrade conventional GNSS-IMU fusion. This paper presents a belief-adaptive online autonomy framework that augments an extended Kalman filter (EKF) with explicit GNSS trust modelling, second-order online belief adaptation, and latency-aware out-of-sequence measurement handling. GNSS trust is represented as a latent belief state that modulates measurement weighting and multipath bias uncertainty, and is updated online using EKF consistency signals. Unlike reactive covariance tuning, the proposed approach enables proactive and stable sensor trust adaptation without prior environmental knowledge or offline training. Evaluation in simulated urban air mobility scenarios with correlated multipath, stochastic latency, and obstacle constraints demonstrates improved belief convergence, smoother trajectories, and reduced estimation and tracking errors compared to naive, adaptive, and first-order baselines. The framework preserves classical GNSS-IMU fusion structure and can be integrated directly into existing flight control pipelines, supporting robust online autonomy in GNSS degraded environments.
Chinese Translation
可靠的在线自主性对于城市空域中的四旋翼运行至关重要,其中全球导航卫星系统(GNSS)测量受到多径、遮挡和延迟问题的影响,引入非平稳、时间相关的误差,从而降低传统GNSS-IMU融合的性能。本文提出了一种信念自适应在线自主框架,该框架通过显式GNSS信任建模、二阶在线信念自适应和延迟感知的乱序测量处理来增强扩展卡尔曼滤波器(EKF)。GNSS信任被表示为一个潜在信念状态,该状态调制测量权重和多径偏差不确定性,并使用EKF一致性信号进行在线更新。与反应式协方差调整不同,所提方法无需先验环境知识或离线训练,即可实现主动且稳定的传感器信任自适应。在具有相关多径、随机延迟和障碍物约束的模拟城市空中交通场景中的评估表明,与朴素、自适应和一阶基线相比,所提方法具有改善的信念收敛、更平滑的轨迹以及减少的估计和跟踪误差。该框架保留了经典GNSS-IMU融合结构,可直接集成到现有飞行控制流程中,支持GNSS退化环境中的鲁棒在线自主性。
cs.RO / 70 / 2609.14868

Primitive-Informed Sampling-Based MPC for Multi-Fingered Dexterous Manipulation

面向多指灵巧操作的基元引导采样型MPC
Küçüktabak, Emek Barış, Patel, Karankumar, Cui, Jinda, Yang, Zhaodong, Sasabuchi, Kazuhiro, Takamatsu, Jun
Abstract
We present a primitive-informed sampling-based model predictive control (MPC) framework for multi-fingered dexterous manipulation. Sampling-based MPC avoids the need for gradients through complex contact dynamics, but direct exploration of the high-dimensional joint space is inefficient and makes performance strongly dependent on the sampling distribution. Our framework biases sampling using low-dimensional manipulation primitives that encode coordinated finger motions, while simultaneously optimizing joint-level residuals to adapt these motions to the current hand-object configuration. Task-related rollout constraints reject infeasible trajectories during forward simulation, improving the effective use of the sampling budget. We evaluate the approach on a physical Allegro hand using a synchronized MuJoCo digital twin. Ablations show that both the primitive and residual are necessary for reliable continuous in-hand rotation, that increasing the sampling budget alone does not recover this coordination, and that rollout constraints substantially improve success rate. A primitive extracted from a single object remains effective across object sizes and under model mismatch, and the framework further supports grasping, object reorientation, and coordinated arm-hand reach-grasp-transport, using primitives extracted from both a simulation-trained policy and human hand-motion data.
Chinese Translation
我们提出了一个面向多指灵巧操作的基元引导采样型模型预测控制(MPC)框架。基于采样的MPC避免了对复杂接触动力学求梯度的需求,但直接探索高维关节空间效率低下,且使性能强烈依赖于采样分布。我们的框架利用编码协调手指运动的低维操作基元来引导采样,同时优化关节级残差,以使这些运动适应当前的手-物配置。任务相关的推演约束在前向仿真过程中拒绝不可行轨迹,提高了采样预算的有效利用率。我们在物理Allegro手上使用同步的MuJoCo数字孪生评估了该方法。消融实验表明,基元和残差对于可靠的连续手内旋转都是必要的;仅增加采样预算并不能恢复这种协调性;推演约束显著提高了成功率。从单个物体提取的基元在不同物体尺寸和模型失配下仍然有效;该框架还支持抓取、物体重定向以及协调的手臂-手部伸手-抓取-运输,使用的基元来自仿真训练的策略和人类手部运动数据。
cs.RO / 71 / 2609.14878

Real-World Reinforcement Learning with MPC Scaffolding for Dexterous Manipulation

基于MPC脚手架的真实世界强化学习用于灵巧操作
Küçüktabak, Emek Barış, Patel, Karankumar, Yang, Zhaodong, Cui, Jinda, Sasabuchi, Kazuhiro, Takamatsu, Jun
Abstract
Real-world reinforcement learning (RL) offers a promising route to dexterous manipulation policies that can adapt directly from physical interaction, but learning is hindered by inefficient early exploration and costly failures. We propose a framework that uses sampling-based model predictive control (MPC) as scaffolding for real-world dexterous RL, providing structured prior experience and task-directed guidance during learning without human demonstrations or corrective actions. A small set of MPC trajectories is first used to populate an offline replay buffer and to pretrain the actor and critic. During online learning, MPC intermittently guides data collection while an off-policy Soft Actor-Critic learner trains from both prior MPC experience and newly collected physical interaction, with control gradually transitioning to the learned policy. On continuous in-hand rotation with a 16-DoF Allegro hand, the method reaches 100\% success in policy-only evaluation (5/5 trials) after 7 minutes of online RL, following initialization with 20 MPC trajectories collected on hardware in 12 minutes. Online training incurs about three object drops on average. After 20 minutes of online learning, the policy achieves more than five times the rotation speed of the MPC controller. It completes 1000 consecutive rotations over more than 110 minutes without a drop. Ablations show complementary benefits from MPC-based pretraining, retained MPC experience, and online MPC guidance. We further demonstrate rapid adaptation to different object geometries and successful goal-conditioned reorientation, showing that the framework enables efficient, low-intervention, real-world dexterous RL.
Chinese Translation
真实世界强化学习(RL)为能够直接从物理交互中适应的灵巧操作策略提供了一条有前景的途径,但学习受到低效早期探索和代价高昂的失败的阻碍。我们提出了一个框架,使用基于采样的模型预测控制(MPC)作为真实世界灵巧强化学习的脚手架,在学习过程中提供结构化先验经验和任务导向的指导,而无需人类演示或纠正动作。首先使用一小组MPC轨迹填充离线回放缓冲区,并预训练行动者(actor)和评论家(critic)。在在线学习过程中,MPC间歇性地指导数据收集,同时一个离策略Soft Actor-Critic学习器从先前的MPC经验和新收集的物理交互中训练,控制逐渐过渡到学习到的策略。在16自由度Allegro手上的连续手内旋转任务上,该方法在7分钟的在线强化学习后,在仅策略评估中达到100%的成功率(5/5次试验),此前使用在硬件上12分钟内收集的20条MPC轨迹进行初始化。在线训练平均导致约三次物体掉落。经过20分钟的在线学习,策略达到MPC控制器旋转速度的五倍以上。它在超过110分钟内完成1000次连续旋转而无掉落。消融实验显示了基于MPC的预训练、保留的MPC经验和在线MPC指导的互补优势。我们进一步展示了针对不同物体几何形状的快速适应以及成功的目标条件重定向,表明该框架实现了高效、低干预的真实世界灵巧强化学习。
cs.RO / 72 / 2609.14903

HydroMap: Probabilistic Water Surface Elevation Mapping for Semantic Scene Representation in Inland Waterways

HydroMap:面向内河航道语义场景表示的概率水面高程建图
Luo, Zhongbi, Wang, Yunjia, Bruyninckx, Herman, Slaets, Peter
Abstract
Autonomous surface vehicles operating in inland waterways require a persistent representation of both surrounding structures and the water surface. LiDAR-based simultaneous localization and mapping often produces sparse or missing water returns, leaving this operational surface absent from the reconstructed scene. We propose HydroMap, an odometry-decoupled framework that reconstructs water surface elevation from stereo observations and integrates it with the structural map. Per-frame water points form joint cell observations with propagated stereo and pose uncertainty, and successive observations are fused into a persistent probabilistic elevation map. Semantic map conversion then combines the elevation map with structural geometry in a unified 2.5D representation of water, boundaries, structures, and overhead regions. On the Pohang Canal and Leuven Vaart datasets, the elevation RMSE remains below 5 cm relative to LiDAR references expressed in the same map frame. The elevation and semantic maps are published at 2 Hz and 1 Hz, respectively. HydroMap thereby complements LiDAR maps with a persistent representation of the water surface for downstream navigation in inland waterways.
Chinese Translation
在内河航道中运行的自主水面艇需要对周围结构物和水面进行持久表示。基于激光雷达的同时定位与建图通常会产生稀疏或缺失的水面回波,导致该作业表面在重建场景中缺失。我们提出 HydroMap,一种与里程计解耦的框架,从立体观测中重建水面高程,并将其与结构地图集成。每帧水点与传播的立体和位姿不确定性形成联合单元观测,连续观测被融合为持久的概率高程图。随后,语义地图转换将高程图与结构几何结合,形成水面、边界、结构物和上方区域的统一 2.5D 表示。在 Pohang Canal 和 Leuven Vaart 数据集上,相对于同一地图坐标系中表示的激光雷达参考,高程 RMSE 保持在 5 cm 以下。高程图和语义图分别以 2 Hz 和 1 Hz 发布。因此,HydroMap 通过水面的持久表示补充了激光雷达地图,用于内河航道中的下游导航。
cs.RO / 73 / 2609.14935

Exact Feasibility Certification and Optimal Responsibility Allocation for Multi-Robot CBF Safety Filters

多机器人CBF安全滤波器的精确可行性认证与最优责任分配
Sah, Chandan Kumar, Keshavan, Jishnu
Abstract
Multi-robot Control Barrier Function (CBF) safety filters can become infeasible, but a failed quadratic program (QP) does not indicate why the conflict occurred or how to resolve it. To address this, we develop an exact feasibility certificate for multi-agent CBF filters with heterogeneous control-affine dynamics and convex input sets. The certificate quantifies a feasibility reserve by separating the demand imposed by safety constraints from the available actuator supply. This decomposition shows when CBF gain tuning or increased actuation can and cannot resolve infeasibility, and identifies the agents and interactions responsible for the conflict. We further propose an algorithm to optimally allocate shared safety constraints by maximizing the worst local feasibility margin, yielding a linear program for polyhedral input sets. In $320$ paired closed-loop simulations, the proposed allocation reduces infeasible control steps from roughly $50\%$ to $6.2\%$, and reduces safety-violating runs from $118/160$ to $24/160$. In addition, across $52$ infeasibility events, the certificate identifies an interaction whose relaxation restores feasibility in $94\%$ of cases.
Chinese Translation
多机器人控制屏障函数(CBF)安全滤波器可能变得不可行,但二次规划(QP)失败并不能说明冲突发生的原因或如何解决。为了解决这个问题,我们为具有异构控制仿射动力学和凸输入集的多智能体CBF滤波器开发了一种精确可行性认证。该认证通过将安全约束施加的需求与可用执行器供给分离,量化了可行性储备。这种分解显示了CBF增益调整或增加执行器何时能够以及何时不能解决不可行性,并识别了导致冲突的智能体和交互。我们进一步提出了一种通过最大化最坏局部可行性裕度来最优分配共享安全约束的算法,对于多面体输入集产生一个线性规划。在320对闭环仿真中,所提出的分配将不可行控制步骤从大约50%减少到6.2%,并将违反安全的运行从118/160减少到24/160。此外,在52次不可行事件中,该认证识别出一种交互,其松弛在94%的情况下恢复了可行性。
cs.RO / 74 / 2609.14936

Comparing Trajectories from Positions Alone: Curvature-Based Time Alignment and Drift Error Metric

仅从位置比较轨迹:基于曲率的时间对齐与漂移误差度量
Daum, Effie, De Martini, Daniele, Dune, Claire, Pomerleau, François
Abstract
In field robotics, acquiring independent large-scale reference trajectories more accurate than the evaluated estimates remains an open challenge. The domain is widely reliant on Absolute Trajectory Error (ATE) and Relative Pose Error (RPE), computed with automated tools, that rest on assumptions and evaluation parameters rarely made explicit. When unreported, the errors can be misleading and hinder fair comparisons. This paper introduces a trajectory-evaluation protocol for standardized and reliable accuracy assessment in state estimation, localization, and Simultaneous Localization And Mapping (SLAM). The approach combines a novel temporal alignment method based on curvature signals with an error metric normalized by travelled distance. We explicitly account for temporal synchronization, sampling alignment, and extrinsic calibration, quantifying their influence through a sensitivity analysis. The proposed protocol contributes to more rigorous, reproducible, and standardized trajectory evaluation.
Chinese Translation
在野外机器人中,获取比被评估估计更准确的独立大规模参考轨迹仍然是一个未决挑战。该领域广泛依赖绝对轨迹误差(Absolute Trajectory Error, ATE)和相对位姿误差(Relative Pose Error, RPE),这些指标由自动化工具计算,但其所依赖的假设和评估参数很少被明确说明。当未报告时,这些误差可能产生误导,并妨碍公平比较。本文提出了一种轨迹评估协议,用于状态估计、定位以及同步定位与建图(Simultaneous Localization And Mapping, SLAM)中标准化且可靠的精度评估。该方法将一种基于曲率信号的新型时间对齐方法与一种按行驶距离归一化的误差度量相结合。我们明确考虑了时间同步、采样对齐和外参标定,并通过敏感性分析量化它们的影响。所提出的协议有助于实现更严格、可重复且标准化的轨迹评估。
cs.RO / 75 / 2609.14984

Two-Stage Personalized Gait Phase Estimation in Stroke Survivors During Exoskeleton-Assisted Walking: An Offline Feasibility Study

外骨骼辅助行走期间卒中幸存者的两阶段个性化步态相位估计:一项离线可行性研究
Ryu, Hyungseok, Hur, Pilwon
Abstract
This study evaluated personalized gait phase estimation for stroke survivors using functional inertial measurement unit (IMU) alignment and two-stage sequential adaptation of models pre-trained on healthy gait. The estimator used signals from a thigh-mounted IMU. Heel force-sensitive resistor measurements provided reference phase labels for offline adaptation and evaluation. Stage 1 established a distillation-regularized participant-specific model, and Stage 2 performed conditional refinement using low-rank adaptation. Long Short-Term Memory (LSTM), Temporal Convolutional Network (TCN), and Transformer models were evaluated in five stroke survivors walking with a powered knee exoskeleton using leave-one-subject-out hyperparameter selection and sequential test-then-adapt Stage 2 replay. Relative to the non-adapted baselines, Stage 1+2 reduced the mean participant-wise phase root mean square error by 84.2%, 77.0%, and 60.7%, respectively. The Transformer achieved the lowest final error (2.90 +- 1.13$% of the gait cycle) and heel-strike timing error (23.7 +- 4.5ms). Policy-specific ablations showed that every-cycle updates generally produced the lowest or near-lowest error, whereas conditional updating reduced the update frequency with small accuracy differences. After personalization, alignment produced model-dependent changes in phase error while preserving or improving heel-strike detection and reducing heel-strike timing error for the LSTM and Transformer. Concurrent embedded tests showed that the TCN and Transformer maintained 100-Hz inference during Stage 2 updates without deadline misses, whereas the LSTM missed the 10-ms deadline in 6.6% of inferences. All updates completed within 0.8s. These results support the offline feasibility and embedded computational timing of the proposed framework for exoskeleton-assisted walking.
Chinese Translation
本研究评估了卒中幸存者的个性化步态相位估计,该方法使用功能性惯性测量单元(IMU)对齐以及在健康步态上预训练模型的两阶段顺序自适应。该估计器使用来自大腿安装的IMU的信号。足跟力敏电阻测量提供了用于离线自适应和评估的参考相位标签。第一阶段建立了蒸馏正则化的参与者特定模型,第二阶段使用低秩自适应进行条件细化。在五名使用动力膝关节外骨骼行走的卒中幸存者中,使用留一受试者超参数选择和顺序测试后自适应第二阶段回放,对长短期记忆(LSTM)、时间卷积网络(TCN)和Transformer模型进行了评估。相对于非自适应基线,阶段1+2将平均参与者相位均方根误差分别降低了84.2%、77.0%和60.7%。Transformer实现了最低的最终误差(步态周期的2.90 ± 1.13%)和足跟触地时间误差(23.7 ± 4.5 ms)。策略特定消融表明,每周期更新通常产生最低或接近最低的误差,而条件更新在精度差异较小的情况下降低了更新频率。个性化后,对齐产生了相位误差的模型依赖性变化,同时保持或改善了足跟触地检测,并降低了LSTM和Transformer的足跟触地时间误差。并发嵌入式测试表明,TCN和Transformer在第二阶段更新期间保持了100 Hz的推理,没有错过截止时间,而LSTM在6.6%的推理中错过了10 ms的截止时间。所有更新均在0.8秒内完成。这些结果支持所提出的外骨骼辅助行走框架的离线可行性和嵌入式计算时序。
cs.RO / 76 / 2609.14997

From Learned-Mode AV-Traffic Pairing to Planner Decisions: A Marginal-Preserving Study on Argoverse 2

从学习模式AV-交通配对到规划器决策:一项在Argoverse 2上的边缘保持研究
Wang, Jingyu
Abstract
Joint motion forecasts pair each autonomous-vehicle (AV) future with surrounding traffic, but actor-level metrics do not show whether that structure matters to a planner. We study this question with a marginal-preserving product control that removes AV-traffic pairing among the learned modes while retaining the fixed constant-velocity pair and holding trajectories, actor-level marginals before planner conditioning, candidates, the cost terms and weights, and fallback fixed. The intervention also changes candidate-conditioned concentration. Across twelve runs on 1,400 held-out Argoverse 2 scenarios, the intervention changes 3.0% of route-level offline selections at $\tau=4$ m. Control-minus-joint recorded-trajectory regret is $-0.026$ and $-0.118$ at the two training sizes; crossed and seed-$t$ intervals span zero. At $\tau=1$ m, relative costs change in 87.9% of route evaluations and route-level offline selections in 8.1%. Before concentration matching, descriptive outcome estimates favor the control. Most of this gap disappears along an approximate concentration-matching path; the remaining contrasts are $+0.112$ and $-0.047$, and both crossed intervals span zero. Actor-level forecast metrics remain identical. Pairing-strength and temperature sweeps show that the decision contrast grows with pairing removal and sharper conditioning. The intervention changes planner decisions even though actor-level metrics remain unchanged. The matching analysis, however, cannot separate any recorded-outcome effect of learned-mode AV-traffic pairing from the accompanying change in conditioned concentration.
Chinese Translation
联合运动预测将每个自动驾驶车辆(AV)的未来与周围交通配对,但参与者级别的指标并不能表明这种结构对规划器是否重要。我们通过一个边缘保持的乘积控制来研究这个问题,该控制在学习模式中移除了AV-交通配对,同时保持固定的恒定速度对,并保持轨迹、规划器条件化之前的参与者级别边缘、候选、成本项和权重以及回退固定不变。该干预还改变了候选条件化的集中度。在1400个留出的Argoverse 2场景上进行的十二次运行中,该干预在τ=4 m时改变了3.0%的路线级离线选择。在两种训练规模下,控制组减去联合组的记录轨迹遗憾分别为-0.026和-0.118;交叉区间和种子t区间均跨越零。在τ=1 m时,相对成本在87.9%的路线评估中发生变化,路线级离线选择在8.1%中发生变化。在集中度匹配之前,描述性结果估计有利于控制组。大部分差距沿着近似集中度匹配路径消失;剩余对比为+0.112和-0.047,且两个交叉区间均跨越零。参与者级别的预测指标保持不变。配对强度和温度扫描表明,决策对比随着配对移除和条件化变得更尖锐而增长。即使参与者级别指标保持不变,该干预也改变了规划器决策。然而,匹配分析无法将学习模式AV-交通配对的任何记录结果效应与伴随的条件化集中度变化区分开来。
cs.RO / 77 / 2609.15005

IMPACT-VLA: Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies

IMPACT-VLA:通过反事实轨迹实现视觉-语言-动作策略的交互感知多模态传播归因
Kim, Jinwoong, Park, Sangjin
Abstract
Vision-Language-Action (VLA) policies perform robot manipulation tasks using multimodal inputs such as visual observations, proprioceptive states, and language instructions. However, it remains unclear at which execution stages each modality contributes to final task success and how input interventions propagate through subsequent states, observations, and actions. Existing attribution approaches primarily measure local sensitivity or temporally aggregated importance, limiting their ability to capture phase-dependent contributions and cross-phase dependencies. We propose Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies (IMPACT-VLA). IMPACT-VLA constructs behavioral phases from action transitions in a successful reference rollout, aligns them with policy query boundaries, and defines phase-modality blocks as attribution units. It then performs closed-loop counterfactual re-execution to quantify each block's contribution to final task success. We further analyze cross-phase non-additive interactions and trajectory propagation while distinguishing behavioral from functional recovery. Across 30 LIBERO robot manipulation tasks using OpenVLA-OFT, dominant-modality transitions occurred in 25 tasks (83.3%), and closed-loop attribution identified task-critical information more faithfully than Static Action Perturbation. Later-block marginal gains for negatively interacting pairs increased by approximately 3.3x under early-phase input replacement, while functional recovery could occur without behavioral recovery. These results reveal when multimodal inputs support task success and how their contributions become conditionally coupled during closed-loop execution.
Chinese Translation
视觉-语言-动作(VLA)策略利用多模态输入(如视觉观察、本体感觉状态和语言指令)执行机器人操作任务。然而,目前尚不清楚每个模态在哪些执行阶段对最终任务成功有贡献,以及输入干预如何通过后续状态、观察和动作进行传播。现有的归因方法主要衡量局部敏感性或时间聚合的重要性,限制了它们捕捉阶段依赖贡献和跨阶段依赖的能力。我们提出了通过反事实轨迹的交互感知多模态传播归因(IMPACT-VLA),用于视觉-语言-动作策略。IMPACT-VLA从成功参考轨迹中的动作转换构建行为阶段,将其与策略查询边界对齐,并定义阶段-模态块作为归因单元。然后,它执行闭环反事实重新执行,以量化每个块对最终任务成功的贡献。我们进一步分析了跨阶段的非加性交互和轨迹传播,同时区分行为恢复和功能恢复。在30个使用OpenVLA-OFT的LIBERO机器人操作任务中,主导模态转换在25个任务(83.3%)中发生,并且闭环归因比静态动作扰动更忠实地识别了任务关键信息。在早期阶段输入替换下,负交互对的后期块边际增益增加了约3.3倍,而功能恢复可能在没有行为恢复的情况下发生。这些结果揭示了多模态输入何时支持任务成功,以及它们的贡献如何在闭环执行过程中有条件地耦合。
cs.RO / 78 / 2609.15012

Atomic Motion Coordinate for Language-Steerable and Force-Responsive Manipulation

Atomic Motion Coordinate:面向语言可引导与力响应操作
Zhai, Jiaqi, Zhao, Jingkai, Yang, Chen, Ma, Siyuan, Zhang, Yutian, Yang, Liwen, Wu, Qinglian, Fan, Weiqi, Wang, Yifei, Zheng, Yi, Gu, Chenxi, Wei, Dong, Zhang, Wei
Abstract
Can changing only the language instruction redirect a VLA policy's end effector, or does the visually driven motion prior dominate? We present Atomic Motion Coordinate, a geometry-grounded coordinate for steerable and force-responsive manipulation. Each arm owns thirteen signed translation, rotation, and hold atoms grounded from text and forward kinematics with vision withheld, and the coordinate is injected into every action-expert block via weighted codebook alignment. Contact history modulates the same coordinate through a bounded spherical residual that is recomputed from a fixed nominal latent to regenerate only the unexecuted horizon suffix. Across 7,520 offline horizon interventions, opposite-atom separation reaches 92.5/83.1% (single/dual) versus 39.1/24.0% for LA4VLA-style. Across 50 real-robot trials per task, AMC raises OOD fruit progress from 60.5% to 87.8%; force adaptation raises Plug/Vase from 59.0/71.5% to 78.5/75.2%.
Chinese Translation
仅改变语言指令能否重定向 VLA 策略的末端执行器,还是视觉驱动的运动先验占主导?我们提出 Atomic Motion Coordinate,一种面向可引导与力响应操作的基于几何的坐标。每条机械臂拥有十三个带符号的平移、旋转和保持原子,这些原子在隐去视觉的情况下由文本和正向运动学接地;该坐标通过加权码本对齐注入每个动作专家模块。接触历史通过一个有界球面残差调制同一坐标,该残差从固定的标称隐变量重新计算,以仅重新生成未执行的时间视野后缀。在 7,520 次离线时间视野干预中,相反原子分离度达到 92.5/83.1%(单/双),而 LA4VLA 风格为 39.1/24.0%。在每个任务 50 次真实机器人试验中,AMC 将 OOD 水果任务进度从 60.5% 提升至 87.8%;力自适应将 Plug/Vase 从 59.0/71.5% 提升至 78.5/75.2%。
cs.RO / 79 / 2609.15014

Steering Generative Robot Policies with Lexicographic Preferences

用字典序偏好引导生成式机器人策略
Jia, Yixuan, How, Jonathan P.
Abstract
Pretrained generative robot policies can produce effective behaviors across diverse environments, but deployment can lead to requirements and preferences that may not have been represented during training. Furthermore, at deployment, an operator, user, or application may assign these requirements and preferences a priority order that can vary across deployments. For example, embodiment-specific feasibility constraints may need to be satisfied first, while user-specific preferences guide behavior among the feasible options. We show that a frozen generative robot policy---based on either diffusion or flow matching---can be steered at inference time to respect such lexicographically ordered deployment objectives. To achieve this, we introduce two modifications to the sampler. First, we apply dynamic-barrier guidance to sampled trajectories, constraining lower-priority updates so that higher-priority costs do not increase (up to first order). Second, we select the executed sample using a cascade that successively filters candidate samples according to each priority level. The policy weights remain unchanged. On a navigation benchmark, we demonstrate that our method improves success, traversability, and preference compliance over the frozen policy, and achieves substantially better compliance than tuned weighted-sum baselines. The same method transfers to a flow-matching manipulation policy on LIBERO, where it improves compliance without reducing task success. A controlled manipulation study further shows that, in settings where a fixed weight can match the desired ordering, the dynamic barrier reaches comparable best performance over a substantially wider range of parameter settings.
Chinese Translation
预训练的生成式机器人策略能在多样环境中产生有效行为,但部署时可能带来训练期间未体现的需求和偏好。此外,在部署时,操作员、用户或应用可能为这些需求和偏好指定优先级顺序,而该顺序可能因部署而异。例如,特定具身的可行性约束可能需要首先满足,而用户特定偏好则在可行选项中引导行为。我们表明,一个冻结的生成式机器人策略——基于扩散模型或流匹配——可以在推理时被引导,以遵循这种按字典序排列的部署目标。为此,我们引入对采样器的两项修改。首先,我们对采样轨迹应用动态屏障引导,约束低优先级更新,使高优先级成本不增加(至一阶)。其次,我们使用级联选择执行的样本,该级联根据每个优先级依次过滤候选样本。策略权重保持不变。在导航基准上,我们证明我们的方法相比冻结策略提高了成功率、可通行性和偏好符合度,并且比调优的加权和基线取得显著更好的符合度。该方法同样迁移到 LIBERO 上的流匹配操作策略,在不降低任务成功率的情况下提高符合度。一项受控操作研究进一步表明,在固定权重能够匹配所需排序的设置中,动态屏障在显著更宽的参数设置范围内达到相当的最佳性能。
cs.RO / 80 / 2609.15082

Task-Distribution-Aware Counterweight Synthesis and Constrained Co-Design for Serial Manipulators

任务分布感知的配重合成与串联机械臂约束协同设计
Abbadi, Mohammad
Abstract
Passive counterweights are simple gravity compensators, but a counterweight selected from a single pose is not generally optimal for the configurations and tasks a manipulator actually executes. This paper develops a task-distribution-aware synthesis framework in which the operating distribution $\rho(q)$ enters the design explicitly. For a counterweight moment $p=m_c r_c$ with gravity torque $-gp\phi(q)$, the weighted mean-square residual gravity torque has the closed-form minimizer $p^*=E_\rho[\tau_g\phi]/(gE_\rho[\phi^2])$. If payload gravity torque is affine in payload mass, the optimum is also affine: $p^*(m_p,\rho)=p_0^*(\rho)+m_pK_p(\rho)$. For fixed static moment, added counterweight inertia is $I_c=pr_c$ while mass is $m_c=p/r_c$, so mass-radius selection is underdetermined unless physical constraints are specified. A recovered three-link manipulator is used as a case study. At $r_c=0.20$ m, zero-payload equivalent optima are 0.672 kg for uniform joint-space operation, 0.683 kg for approximately uniform task-space operation, 0.713 kg for a representative pick-and-place family, and 0.952 kg for a high-gravity-biased distribution, a change of more than 40% caused solely by the operating distribution. Nondominated fronts show that preferred mass-radius pairs depend on declared engineering bounds. A rated-torque-referenced all-joint screen increases zero-payload feasible task-space coverage from 78.1% without compensation to 93.7% for the uniform-distribution design. A lumped point-mass trajectory study gives a provisional crossover from no counterweight at very aggressive motion to stronger compensation as motion slows. These actuator and dynamic results are engineering consequence studies rather than physical validation.
Chinese Translation
被动配重是简单的重力补偿器,但从单一位姿选择的配重对于机械臂实际执行的构型和任务通常并非最优。本文提出了一个任务分布感知的综合框架,其中操作分布 $\rho(q)$ 显式地参与设计。对于配重力矩 $p=m_c r_c$,其重力矩为 $-gp\phi(q)$,加权均方残余重力矩具有闭式最小化器 $p^*=E_\rho[\tau_g\phi]/(gE_\rho[\phi^2])$。如果负载重力矩与负载质量成仿射关系,则最优值也是仿射的:$p^*(m_p,\rho)=p_0^*(\rho)+m_pK_p(\rho)$。对于固定的静力矩,附加配重惯量为 $I_c=pr_c$,而质量为 $m_c=p/r_c$,因此除非指定物理约束,否则质量-半径选择是不确定的。以一个恢复的三连杆机械臂作为案例研究。在 $r_c=0.20$ m 时,零负载等效最优值分别为:均匀关节空间操作为 0.672 kg,近似均匀任务空间操作为 0.683 kg,代表性拾放任务族为 0.713 kg,高重力偏置分布为 0.952 kg,仅由操作分布引起的改变超过 40%。非支配前沿表明,优选的质量-半径对取决于所声明的工程界限。以额定扭矩为参考的全关节筛选将零负载可行任务空间覆盖率从无补偿的 78.1% 提高到均匀分布设计的 93.7%。集总质点轨迹研究给出了一个暂定的交叉点:在非常剧烈的运动下无需配重,而随着运动减慢则需要更强的补偿。这些执行器和动力学结果是工程后果研究,而非物理验证。
cs.RO / 81 / 2609.15113

Legislating World-Model-Based Planning with Legal Reasoning

基于法律推理的世界模型规划立法
Waldner, Dylan, Kantaros, Yiannis, Governatori, Guido, Miikkulainen, Risto, Banifatemi, Amir
Abstract
As robotic systems grow more general, legal norms are needed to integrate them into society. This paper extends the isomorphism problem of aligning legal source texts with their encodings, and measures two key challenges to robot normative control: (1) the \textit{grounding isomorphism gap}, where perception error grounds false atoms for legal reasoning, and (2) the \textit{ontological isomorphism gap}, where one legal conclusion admits many faithful translations into planning constraints. The paper introduces a legal planning stack that employs Defeasible Deontic Logic (DDL) to constrain a motion planner. The stack leverages learned world models to plan and to provide legal context, enabling \textit{ex ante} governance that intervenes before an illegal action is executed. It was deployed on a simulated robot arm pushing a cube across a $3\times3$ grid. The findings were (1) the legislated agent abided substantially more often than the non-legislated one, and modeling perception uncertainty lifted abidance even further, (2) the legal reasoning ran efficiently at runtime and its verdicts were auditable, and (3) the stack adapted to exogenous signals and endogenous rule changes. Both gaps were measured: (4) world model and probe error corrupted the factual input for the DDL reasoner, and (5) a single law admitted several faithful metric interpretations yielding drastically different abidance. Thus, \textit{ex ante} legislation functions as intended, and closing these gaps with a standardized mapping from the law to runtime constraints and improved fact grounding from perception will yield robust laws that align robot behavior with society's norms.
Chinese Translation
随着机器人系统变得越来越通用,需要法律规范将其融入社会。本文扩展了将法律源文本与其编码对齐的同构问题,并衡量了机器人规范控制的两个关键挑战:(1)基础同构差距,其中感知错误为法律推理提供了错误的基础原子命题,以及(2)本体论同构差距,其中一个法律结论允许多种忠实的转换成为规划约束。本文介绍了一个法律规划栈,它采用可废止道义逻辑(DDL)来约束运动规划器。该栈利用学习到的世界模型进行规划并提供法律背景,从而实现事前治理,在非法行动执行之前进行干预。它被部署在一个模拟机械臂上,该机械臂在3×3网格上推动一个立方体。研究结果如下:(1)受立法约束的智能体比未受立法约束的智能体遵守频率显著更高,并且对感知不确定性进行建模进一步提高了遵守程度,(2)法律推理在运行时高效运行,其裁决可审计,(3)该栈适应外源信号和内源规则变化。两个差距都被测量:(4)世界模型和探测误差破坏了DDL推理器的事实输入,以及(5)单一法律允许多种忠实的度量解释,导致截然不同的遵守程度。因此,事前立法按预期发挥作用,而通过标准化的法律到运行时约束的映射以及改进的感知事实基础来弥合这些差距,将产生稳健的法律,使机器人行为与社会规范保持一致。
cs.RO / 82 / 2609.15142

C$^2$Nav: Compare Before You Commit for Zero-Shot Vision-and-Language Navigation

C$^2$Nav:用于零样本视觉-语言导航的先比较后承诺
Zheng, Runtian, Zhang, Congpeng, Liu, Ying
Abstract
Zero-shot vision-and-language navigation in continuous environments (VLN-CE) increasingly places foundation vision-language models (VLMs) inside the navigation loop. Existing systems commonly request cardinal outputs such as waypoints, pixels, headings, progress values, or absolute arrival decisions, coupling a generative response to geometric magnitude or an irreversible commitment. We study a complementary model-robot interface: the VLM compares controller-constructed alternatives, while geometry, thresholds, action magnitude, and execution remain on the physical side. We instantiate this idea in C2Nav, a training-free framework with three coordinated faculties. Seeing performs ordinal Gaze Election over physically vetted candidate views; Remembering maintains a compact route sketch and compares adjacent instruction-leg hypotheses; and Arriving combines a hesitation ladder, look-back comparison, and revocable walk-back for reliable stopping. On the public OpenNav R2R-CE 100 protocol, C2Nav with Qwen3-VL-8B-Instruct obtains 41.0% OSR, 31.0% SR, and 16.7% SPL, while the same interface with the standard GPT-5.5 model reaches 54.0% OSR, 44.0% SR, and 29.0% SPL. Whole-faculty ablations reduce SR to 14.0% without Seeing, 25.0% without Remembering, and 29.0% without Arriving. Matched role inversions that replace only the comparative answer form with cardinal/absolute questions reduce SR to 12.0%, 28.0%, and 21.0% in the spatial, transition, and terminal slots, respectively. The results indicate that a constrained decision interface and stronger VLM reasoning are complementary rather than interchangeable.
Chinese Translation
连续环境中的零样本视觉-语言导航(VLN-CE)越来越多地将基础视觉-语言模型(VLM)置于导航回路中。现有系统通常请求基数输出,如路点、像素、航向、进度值或绝对到达决策,将生成式响应与几何量级或不可逆承诺耦合。我们研究一种互补的模型-机器人接口:VLM比较控制器构建的备选方案,而几何、阈值、动作幅度和执行仍留在物理侧。我们在C2Nav中实例化了这一想法,这是一个无需训练的框架,具有三个协调的能力。Seeing在物理审查的候选视图上执行序数凝视选举;Remembering维护紧凑的路线草图并比较相邻指令段假设;Arriving结合犹豫阶梯、回看比较和可撤销的走回,以实现可靠停止。在公开的OpenNav R2R-CE 100协议上,使用Qwen3-VL-8B-Instruct的C2Nav获得41.0%的OSR、31.0%的SR和16.7%的SPL,而相同接口使用标准GPT-5.5模型达到54.0%的OSR、44.0%的SR和29.0%的SPL。全能力消融实验将SR降低至:无Seeing时14.0%,无Remembering时25.0%,无Arriving时29.0%。匹配的角色反转,仅将比较性答案形式替换为基数/绝对问题,在空间、过渡和终端槽中分别将SR降低至12.0%、28.0%和21.0%。结果表明,受约束的决策接口和更强的VLM推理是互补的,而不是可互换的。
cs.RO / 83 / 2609.15162

LieSpline-DP: Lie-Group B-Spline Diffusion Policy for Smooth Robot Manipulation

LieSpline-DP:用于平滑机器人操作的李群B样条扩散策略
Xie, Erxuan, Liu, Bang, Nie, Pingyun, Liu, Xingkai, Fu, Zhuang, Zhang, Bo
Abstract
Diffusion Policy (DP) is a powerful Learning from Demonstration (LfD) method for robotic manipulation, yet it suffers from discontinuous and non-smooth trajectories. Spline-based action representations promote smooth motion within individual action chunks, but existing spline-based methods neither guarantee cross-chunk $C^2$ continuity nor account for the group structure of $\mathrm{SE}(3)$. We therefore propose LieSpline-DP, a Lie-group B-spline diffusion policy that generates end-effector trajectories directly on $\mathrm{SE}(3)$ and couples consecutive plans by sharing their boundary control poses, ensuring $C^2$ continuity throughout the entire planned trajectory. Across three real-robot tasks, LieSpline-DP produces lower trajectory jerk and higher task success rates than the DP baseline. The gains are particularly pronounced in real-world tasks involving liquids and flexible objects: in our real-robot experiments, LieSpline-DP achieved a 100% success rate on both pouring and bucket hooking, whereas the DP baseline achieved only 10% and 30%, respectively.
Chinese Translation
扩散策略(Diffusion Policy, DP)是一种强大的机器人操作示范学习(Learning from Demonstration, LfD)方法,但其轨迹存在不连续和不平滑的问题。基于样条的动作表示能够促进单个动作块(action chunk)内的平滑运动,然而现有基于样条的方法既不能保证跨动作块的 C^2 连续性,也没有考虑 SE(3) 的群结构。因此,我们提出 LieSpline-DP,一种李群 B 样条扩散策略,它直接在 SE(3) 上生成末端执行器轨迹,并通过共享相邻规划的边界控制位姿来耦合连续规划,从而保证整个规划轨迹上的 C^2 连续性。在三个真实机器人任务中,LieSpline-DP 相比 DP 基线产生了更低的轨迹加加速度(jerk)和更高的任务成功率。这些优势在涉及液体和柔性物体的真实世界任务中尤为明显:在我们的真实机器人实验中,LieSpline-DP 在倒水和钩桶任务上均达到了 100% 的成功率,而 DP 基线分别仅达到 10% 和 30%。
cs.RO / 84 / 2609.15195

HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

HarnessVLN:通过 Agent Harness 统一免训练具身导航
Chen, Yang, Che, Lirong, Huang, Zhenyu, Fu, Wenbo, Wang, Chuang, Cao, Xu, Liu, Daqi, Yang, Yuzhe, Su, Jian, Guo, Lan-Zhe
Abstract
Embodied navigation requires agents to interpret visual observations, accumulate spatial knowledge, and execute actions to follow instructions or locate objects. Training-based methods face generalization challenges, while training-free methods exploit multimodal large language models (MLLMs) but often lack mechanisms to reconcile proposed actions with spatial evidence, task progress, and execution failures. We present HarnessVLN, a zero-shot, training-free framework whose Agent Harness coordinates perception, retrieval, grounding, navigation, recovery, and termination through a unified tool interface. The Harness validates planner proposals against spatial evidence, geometric feasibility, and subgoal consistency, incorporating structured tool feedback into subsequent decisions. Hierarchical event memory tracks task progress and execution history, while a persistent Spatiotemporal Graph maintains reusable spatial evidence and failure annotations for verification and recovery. A replaceable Navigation Executor converts validated targets into executable motions, allowing the same Harness protocol to support instruction-following and object-goal navigation. HarnessVLN achieves success rates of 60.8%, 53.9%, 76.0%, and 59.3% on R2R, RxR, HM3D-v2, and HM3D-OVON, respectively, surpassing prior training-free SOTA results. Humanoid deployment further demonstrates its applicability to both tasks in real-world environments. The project page is: https://harnessvln.netlify.app/.
Chinese Translation
具身导航要求智能体解读视觉观察、积累空间知识,并执行动作以遵循指令或定位目标物体。基于训练的方法面临泛化挑战,而免训练方法利用多模态大语言模型(MLLMs),但往往缺乏机制来协调所提出的动作与空间证据、任务进展和执行失败。我们提出 HarnessVLN,一个零样本、免训练的框架,其 Agent Harness 通过统一的工具接口协调感知、检索、定位、导航、恢复和终止。Harness 根据空间证据、几何可行性和子目标一致性验证规划器提出的方案,并将结构化的工具反馈纳入后续决策。分层事件记忆跟踪任务进展和执行历史,而持久化的时空图(Spatiotemporal Graph)维护可复用的空间证据和失败标注,用于验证和恢复。可替换的导航执行器(Navigation Executor)将已验证的目标转换为可执行运动,使同一 Harness 协议能够支持指令遵循和物体目标导航。HarnessVLN 在 R2R、RxR、HM3D-v2 和 HM3D-OVON 上分别达到 60.8%、53.9%、76.0% 和 59.3% 的成功率,超越了此前免训练的 SOTA 结果。人形机器人部署进一步证明了其在真实世界环境中对这两类任务的适用性。项目页面:https://harnessvln.netlify.app/。
cs.RO / 85 / 2609.15198

PredTac: Learning Contact-Rich Manipulation with Predicted Touch

PredTac:利用预测触觉学习富接触操作
Fan, Weijia, Guo, Daqiang
Abstract
Contact-rich manipulation benefits from tactile feedback, yet physical tactile sensors introduce hardware, calibration, synchronization, and maintenance costs that complicate policy learning and deployment. We formulate predicted touch as an alternative to measured tactile input and present PredTac, a framework that learns to infer tactile states from causal visual observations and robot states and uses the predicted touch as an explicit interface for policy learning and execution. A tactile predictor is first trained with tactile supervision and then used to provide contact information without requiring measured tactile input during downstream policy training or execution. We evaluate PredTac across three contact-rich manipulation tasks in simulation and on a real robot, and further examine how policy performance depends on the predicted contact content. In simulation goal-offset evaluations, predicted-touch policies achieve 27.0%, 52.0%, and 44.7% success on USB, Barbed-spike, and Valve, respectively, improving over the visual baseline by 8.0-13.7 percentage points. On the real robot, predicted-touch ACT achieves 70.0%, 50.0%, and 90.0% success on USB insertion, Barbed extraction, and Valve rotation, respectively, with a three-task mean of 70.0%, approaching measured-touch ACT at 72.2% and substantially outperforming visual ACT at 21.1%. Fixed-policy interventions further show that performance is sensitive to the spatial structure of predicted contact, with spatial rearrangement at fixed value distributions reducing Valve success by 10.7 percentage points. These results demonstrate that predicted touch can provide useful contact information for contact-rich manipulation without requiring tactile sensing as a policy input.
Chinese Translation
富接触操作受益于触觉反馈,然而物理触觉传感器引入了硬件、校准、同步和维护成本,使得策略学习和部署复杂化。我们提出将预测触觉作为测量触觉输入的替代方案,并提出了PredTac框架,该框架学习从因果视觉观察和机器人状态推断触觉状态,并将预测触觉作为策略学习和执行的显式接口。首先用触觉监督训练触觉预测器,然后用于提供接触信息,而无需在下游策略训练或执行期间要求测量触觉输入。我们在仿真和真实机器人上对三个富接触操作任务评估PredTac,并进一步考察策略性能如何依赖于预测接触内容。在仿真目标偏移评估中,预测触觉策略在USB、Barbed-spike和Valve上分别达到27.0%、52.0%和44.7%的成功率,比视觉基线提高了8.0-13.7个百分点。在真实机器人上,预测触觉ACT在USB插入、Barbed提取和Valve旋转上分别达到70.0%、50.0%和90.0%的成功率,三个任务平均为70.0%,接近测量触觉ACT的72.2%,并大幅优于视觉ACT的21.1%。固定策略干预进一步表明,性能对预测接触的空间结构敏感,在固定值分布下进行空间重排使Valve成功率降低了10.7个百分点。这些结果表明,预测触觉可以为富接触操作提供有用的接触信息,而无需将触觉传感作为策略输入。
cs.RO / 86 / 2609.15213

X-WBC: A Cross-Embodiment Foundation Model for Humanoid Whole-Body Control

X-WBC:一种用于人形机器人全身控制的跨本体基础模型
Zhang, Juntong, Gu, Chun, Zhang, Li
Abstract
Scaling humanoid whole-body control toward general-purpose deployment requires large human motion corpora and training experience shared across robot bodies. Existing methods usually train one policy per robot, leaving motion experience isolated across embodiments. We introduce X-WBC, a cross-embodiment foundation framework that separates relatively shared human motion semantics from embodiment-specific physical execution. Human-centered command tokens align full human motion, robot reference motion, and sparse VR observations. A causal Transformer learns reusable temporal structure from mixed multi-robot rollouts, while lightweight robot-specific modules map the shared representation to each robot's proprioception and action space. Across nine simulated embodiments, external motions, and four real robots, experiments show that joint training improves tracking, the aligned representation supports consistent control across command sources, and the learned policy remains competitive beyond the training corpus. These results support heterogeneous humanoids as joint data sources and establish cross-embodiment joint training as a practical route toward whole-body control foundation models.
Chinese Translation
将人形机器人全身控制扩展到通用部署,需要大规模人体运动语料库以及跨机器人本体共享的训练经验。现有方法通常为每个机器人训练一个策略,导致运动经验在不同本体之间相互隔离。我们提出 X-WBC,一个跨本体基础框架,将相对共享的人体运动语义与特定本体的物理执行分离。以人为中心的命令 token 对齐完整人体运动、机器人参考运动和稀疏 VR 观测。因果 Transformer 从混合多机器人 rollout 中学习可复用的时序结构,而轻量化的机器人特定模块将共享表示映射到每个机器人的本体感知和动作空间。在九个仿真本体、外部运动和四台真实机器人上,实验表明联合训练提升了跟踪性能,对齐表示支持跨命令源的一致控制,且学习到的策略在训练语料之外仍具有竞争力。这些结果支持将异构人形机器人作为联合数据源,并确立了跨本体联合训练作为通向全身控制基础模型的一条实用路径。
cs.RO / 87 / 2609.15232

A Vision Based Framework Integrating Attention and Action Cues for Interpretable Cognitive Workload Assessment in Human Robot Collaborative Assembly

一种基于视觉的框架,融合注意力与动作线索,用于人机协作装配中可解释的认知工作负荷评估
Xionga, Junyan, Feng, Naiyi, Xia, Xingke, Fan, Qihang, Chen, Suchang, Guo, Daqiang
Abstract
The introduction of human-robot collaboration (HRC) in industrial assembly operations is revolutionizing the manufacturing landscape. In this evolving environment, operators are required to seamlessly coordinate their manual tasks with real-time task information and robotic behaviors. These demands fluctuate during operation, yet conventional workload assessments depend on body-worn physiological sensors that complicate practical deployment. Here, we present a vision-based attention--action framework for continuous and interpretable workload-related assessment in HRC assembly. The framework combines RGB-D observations with robot states and calibrated task-related areas to construct a temporally confirmed representation of operator behavior. This representation identifies where task demand is concentrated and explains how it develops when attention and action diverge, the task context changes, or the operator hesitates. We evaluated the framework in a three-level collaborative gearbox assembly experiment with ten participants, using subjective ratings and synchronized physiological signals as independent references. Raw NASA-TLX ratings confirmed increasing perceived workload across conditions, with significant effects on overall workload and its mental and temporal dimensions. The vision-derived HRC-CWL output was significantly associated with ECG-derived features in seven of nine participants with complete correlation data. Synchronized interaction episodes further showed temporal correspondence between detected hesitation and physiological activity. Real-time deployment demonstrated that the framework can operate without requiring operators to wear additional sensors. These findings support HRC-CWL as an interpretable behavioral proxy for cognitive ergonomics analysis and adaptive robot assistance, rather than a direct psychophysiological measure of workload.
Chinese Translation
在工业装配操作中引入人机协作(HRC)正在彻底改变制造业的格局。在这种不断变化的环境中,操作员需要将手动任务与实时任务信息和机器人行为无缝协调。这些需求在操作过程中波动,而传统的工作负荷评估依赖于穿戴式生理传感器,这使实际部署变得复杂。在此,我们提出一个基于视觉的注意力-动作框架,用于HRC装配中连续且可解释的工作负荷相关评估。该框架将RGB-D观测与机器人状态和校准的任务相关区域相结合,构建操作员行为的时序确认表示。该表示识别任务需求集中的位置,并解释当注意力与动作偏离、任务情境变化或操作员犹豫时,任务需求如何发展。我们在一个三级协作变速箱装配实验中评估了该框架,有十名参与者,使用主观评分和同步生理信号作为独立参考。原始NASA-TLX评分证实,不同条件下感知工作负荷增加,对整体工作负荷及其心理和时间维度有显著影响。在九名具有完整相关数据的参与者中,有七名参与者的视觉衍生HRC-CWL输出与ECG衍生特征显著相关。同步的交互片段进一步显示了检测到的犹豫与生理活动之间的时间对应关系。实时部署表明,该框架可以在不需要操作员穿戴额外传感器的情况下运行。这些发现支持HRC-CWL作为认知工效学分析和自适应机器人辅助的可解释行为代理,而不是工作负荷的直接心理生理测量。
cs.RO / 88 / 2609.15266

Distributed Safe Cooperative Vector Field for Trajectory Curvature Constrained Multi-Robot Systems

面向轨迹曲率约束多机器人系统的分布式安全协同向量场
Xiao, Zhouru, Teng, Tao, Yao, Weijia, Zhang, Hui, Wang, Xiangke, Wang, Yaonan
Abstract
Trajectory curvature constraints are inherent in practical multi-robot systems due to the limited turning capabilities of the robots. Without properly accounting for these constraints, robots may fail to accomplish assigned tasks, and their trajectories may diverge from the intended paths. This paper proposes a distributed safe cooperative vector field approach for multi-robot systems subject to trajectory curvature constraints. The proposed approach is composed of a cooperative vector field and a safety-oriented collision avoidance vector field, aiming to address the problems of cooperative motion and safe collision avoidance in multi-robot path-following tasks. A safety-oriented collision avoidance vector field with adaptively adjustable reactive boundary is developed to accommodate the kinematic curvature constraints of robots, thereby ensuring the physical feasibility of collision avoidance maneuvers. The proposed vector field requires only a single virtual variable from each neighboring robot to achieve cooperative motion and ensure both obstacle avoidance and inter-robot collision avoidance. The effectiveness of the proposed approach is validated through both simulations and real-world experiments on an actual multi-robot platform.
Chinese Translation
由于机器人的转向能力有限,轨迹曲率约束在实际多机器人系统中是固有的。如果没有适当考虑这些约束,机器人可能无法完成分配的任务,并且其轨迹可能会偏离预期路径。本文针对受轨迹曲率约束的多机器人系统,提出了一种分布式安全协同向量场方法。所提方法由协同向量场和面向安全的避碰向量场组成,旨在解决多机器人路径跟踪任务中的协同运动和安全避碰问题。开发了一种具有自适应可调反应边界的安全导向避碰向量场,以适应机器人的运动学曲率约束,从而确保避碰机动的物理可行性。所提出的向量场仅需从每个相邻机器人获取一个虚拟变量,即可实现协同运动,并确保避障和机器人间避碰。通过仿真和在实际多机器人平台上的真实世界实验,验证了所提方法的有效性。
cs.RO / 89 / 2609.15276

Low Clearance Hinge Joint Mechanism Based on 3D Printing on Sheet Fabrication Methodology

基于片材上3D打印制造方法的低间隙铰链关节机构
Jang, Jaehyung, Shin, Euibin, Okamura, Allison M., Ryu, Jee-Hwan
Abstract
This paper presents a low-clearance hinge joint mechanism based on the 3D printing on sheet fabrication method. This approach simplifies the fabrication of hinge mechanisms and overcomes limitations of conventional origami manufacturing by eliminating the need for adhesives commonly used during assembly, making it suitable for robots at the tens-of-centimeters scale. The advantages and disadvantages of three types of hinge joint mechanisms are compared, and a hinge joint that can be designed with low clearance for various facet thicknesses is selected. Based on the selected hinge joint, the twisting angle and bending force are analyzed, leading to the implementation of a clearance of 0.1 mm. Torsional resistance is experimentally evaluated to measure the torque required for twisting caused by plastic deformation and clearance. The results show that the torque associated with plastic deformation is sufficient to constrain the undesired degrees of freedom of the hinge joint, while the torque required for twisting due to clearance is minimal. Based on the analyzed data, the proposed hinge joint mechanism is applied to a 3-degree-of-freedom delta robot manipulator, demonstrating precise motion with low clearance.
Chinese Translation
本文提出了一种基于片材上3D打印制造方法的低间隙铰链关节机构。该方法简化了铰链机构的制造,并通过消除装配过程中常用的粘合剂,克服了传统折纸制造的局限性,使其适用于几十厘米尺度的机器人。比较了三种铰链关节机构的优缺点,并选择了一种可针对不同面片厚度设计为低间隙的铰链关节。基于所选铰链关节,分析了扭转角度和弯曲力,从而实现了0.1 mm的间隙。通过实验评估抗扭阻力,以测量由塑性变形和间隙引起的扭转所需扭矩。结果表明,与塑性变形相关的扭矩足以约束铰链关节的非期望自由度,而由间隙引起的扭转所需扭矩极小。基于分析数据,将所提出的铰链关节机构应用于三自由度Delta机器人机械臂,展示了低间隙下的精确运动。
cs.RO / 90 / 2609.15322

Planning in the Backbone: DiffAdapterVLA for Native Continuous Trajectory Generation with Driving VLMs

在骨干网络中规划:DiffAdapterVLA 实现驾驶 VLM 的原生连续轨迹生成
Lu, Changxin, Meng, Xiaoliang, Wu, Yu, Huang, Rui, Li, Honglin, Chen, Tao, Zhou, Kaixuan, Shao, Yadong
Abstract
Pretrained driving vision-language models (VLMs) integrate visual, route, language, and driving context into rich driving priors, yet their representation objectives remain separated from continuous driving planning. Existing methods typically begin trajectory generation only after the VLM has formed a final condition, leaving depth-wise condition computation outside the stepwise formation of trajectory state. We introduce DiffAdapterVLA, which realizes Planning in the Backbone: it injects explicit trajectory tokens into selected VLM late layers, bringing trajectory state into backbone forward computation, where it co-evolves with driving conditions at different depths. Lightweight layer-wise DiffAdapters organize this computation into recursive trajectory refinement, while asymmetric joint attention preserves directed guidance from the condition stream to trajectory planning. By placing planning within existing backbone computation rather than relying on an independent trajectory planner, DiffAdapterVLA adapts only lightweight trajectory modules to turn existing driving priors into efficient continuous planning capability. NAVSIM results show that it achieves high-quality closed-loop planning with low end-to-end latency using few trainable parameters, and demonstrate that jointly evolving trajectory state and depth-wise driving conditions in VLM late-layer computation effectively realizes continuous trajectory planning.
Chinese Translation
预训练的驾驶视觉语言模型(VLM)将视觉、路线、语言和驾驶上下文整合为丰富的驾驶先验,但其表示目标仍与连续驾驶规划相分离。现有方法通常仅在 VLM 形成最终条件后才开始轨迹生成,使逐深度条件计算处于轨迹状态逐步形成过程之外。我们提出 DiffAdapterVLA,实现了“在骨干网络中规划”:它将显式轨迹 token 注入选定的 VLM 后层,将轨迹状态带入骨干网络前向计算,使其在不同深度上与驾驶条件共同演化。轻量级的逐层 DiffAdapter 将该计算组织为递归轨迹细化,而非对称联合注意力保持条件流对轨迹规划的有向引导。通过将规划置于现有骨干网络计算中,而不是依赖独立的轨迹规划器,DiffAdapterVLA 仅适配轻量级轨迹模块,就将现有驾驶先验转化为高效的连续规划能力。NAVSIM 结果表明,它使用少量可训练参数即可实现高质量闭环规划并具有低端到端延迟,并证明在 VLM 后层计算中联合演化轨迹状态与逐深度驾驶条件能够有效实现连续轨迹规划。
cs.RO / 91 / 2609.15352

Assistance Torque Estimation via Dynamics-Aware Optimization for Lower-Limb Exoskeleton in Complex Environments

复杂环境下下肢外骨骼的动力学感知优化辅助力矩估计
Liu, Xiao-Yin, Li, Guotao, Wang, Weiqun, Hou, Zeng-Guang
Abstract
Ground-truth human joint torque estimation relies on motion capture systems, which suffer from limited outdoor usability and significant deployment expenses. Furthermore, direct scaling of ground-truth joint torques to obtain motor torque commands is not necessarily the optimal strategy. To address the aforementioned limitations, inspired by the human motion generation process, this paper proposes a novel assistance torque estimation method based on the dynamic model. From an optimization perspective, the proposed method directly generates motor-assist torque and lowers the cost of data acquisition. Then, a data-driven assistance torque prediction network is trained to enable accurate real-time prediction under complex outdoor environments. Experimental results demonstrate that optimized (estimated) assistance torque exhibits better phase consistency with gait trajectories and better alignment with task characteristics. Relative to the Zero torque condition, the predicted torque can decrease metabolic rate by 11.8%-17.7%, heart rate by 8.9%-14.3%, and peak muscle activation levels by 28.2%-54.0%, respectively. This provides a new perspective for low-cost adaptive exoskeleton assistance.
Chinese Translation
真实人体关节力矩估计依赖于运动捕捉系统,而这类系统户外可用性有限且部署成本高昂。此外,直接缩放真实关节力矩以获得电机力矩指令并不一定是最优策略。针对上述局限性,受人体运动生成过程的启发,本文提出了一种基于动力学模型的新型辅助力矩估计方法。从优化的角度,该方法直接生成电机辅助力矩,并降低了数据采集成本。随后,训练了一个数据驱动的辅助力矩预测网络,以在复杂的户外环境中实现准确的实时预测。实验结果表明,优化(估计)的辅助力矩与步态轨迹具有更好的相位一致性,并与任务特征更好地对齐。相对于零力矩条件,预测力矩可分别降低代谢率11.8%-17.7%、心率8.9%-14.3%以及峰值肌肉激活水平28.2%-54.0%。这为低成本自适应外骨骼辅助提供了新视角。
cs.RO / 92 / 2609.15362

Understanding User Preferences of a Slope-Aware Variable-Admittance Filter for a Robot Guide Dog

理解用户对用于机器人导盲犬的坡度感知可变导纳滤波器的偏好
Esposito, Federico, Selvaggio, Mario, Link, Aaron, Ruggiero, Fabio
Abstract
This letter investigates how the parameters of a slope-aware variable-admittance filter influence user preferences in force-based interaction with a robotic guide dog for visually impaired individuals. The proposed system consists of a quadruped robot equipped with a sensor-free rigid handle for physical guidance. The framework combines path following, momentum-based interaction-wrench estimation, and a variable-admittance filter whose stiffness and damping are adapted online from slope information extracted by the robot's depth camera. The adaptation policies are evaluated through high-fidelity simulations and a human-subject study involving blindfolded sighted participants. Multiple strategies are compared using a Taguchi L9 design of experiments. Preliminary main-effect results suggest that increasing stiffness uphill and decreasing it downhill improves both objective and subjective metrics, whereas damping shows no significant main effect.
Chinese Translation
本文研究了坡度感知可变导纳滤波器的参数如何影响视障人士与机器人导盲犬基于力的交互中的用户偏好。所提出的系统由配备无传感器刚性手柄的四足机器人组成,用于物理引导。该框架结合了路径跟踪、基于动量的交互力旋量估计以及可变导纳滤波器,其刚度和阻尼根据机器人深度相机提取的坡度信息在线调整。通过高保真仿真和涉及蒙眼视力正常参与者的真人研究评估了自适应策略。采用田口L9实验设计比较了多种策略。初步主效应结果表明,上坡时增加刚度、下坡时减小刚度可改善客观和主观指标,而阻尼则无显著主效应。
cs.RO / 93 / 2609.15382

From Prediction to Decision: World-Model-Guided Action Selection for Continuous Pile Excavation

从预测到决策:世界模型引导的连续料堆挖掘动作选择
Zhang, Ailing, Gao, Fan, Zhang, Song, Leong, Kawa, Wu, Ziyu, Wang, Yafei
Abstract
Wheel-loader excavation is a sequential decision problem in which every scoop changes the terrain available to subsequent actions. A practical world model must predict action consequences accurately, rank candidates in real time, and operate inside the closed loop of a full-size machine. We present the World-Action Model (WAM), which proposes multiple scoops, rejects geometrically inadmissible candidates, jointly predicts signed terrain change and loaded volume, executes the candidate with the largest predicted load, and replans from the newly observed terrain. On 32 geometry-disjoint MinSlope test episodes, adding world-model ranking to matched diffusion proposals reduces the mean scoop count from 651.8 to 540.6 (17.1%), preserves 32/32 completion, and improves every paired episode. In a complete-system comparison, WAM completes 32/32 episodes versus 29/32 for an independently trained soft actor-critic policy. Comparisons of input representations, spatial support, and five architectures identify an accurate and efficient physics-structured predictor. We further evaluate the interface on event-disjoint full-size-loader data and deploy the complete perception-proposal-prediction-selection-execution loop for autonomous excavation. The ROS2/TensorRT implementation processes five candidates in 72.4 ms on a Jetson AGX Orin. The simulation results establish decision-level gains, while the physical experiments demonstrate real-world closed-loop feasibility.
Chinese Translation
轮式装载机挖掘是一个序贯决策问题,其中每次铲挖都会改变后续动作可用的地形。实用的世界模型必须准确预测动作后果,实时对候选动作进行排序,并在全尺寸机器的闭环内运行。我们提出了世界-动作模型(WAM),它提出多个铲挖动作,拒绝几何上不可行的候选动作,联合预测带符号的地形变化和装载体积,执行预测负载最大的候选动作,并根据新观测到的地形重新规划。在32个几何不相交的MinSlope测试回合中,将世界模型排序添加到匹配的扩散提议中,将平均铲挖次数从651.8减少到540.6(17.1%),保持32/32的完成率,并改善每个配对回合。在完整系统比较中,WAM完成了32/32个回合,而独立训练的软演员-评论家策略完成了29/32个回合。通过比较输入表示、空间支持和五种架构,确定了一个准确且高效的物理结构化预测器。我们进一步在事件不相交的全尺寸装载机数据上评估该接口,并部署完整的感知-提议-预测-选择-执行循环用于自主挖掘。ROS2/TensorRT实现可在Jetson AGX Orin上以72.4毫秒处理五个候选动作。仿真结果确立了决策层面的增益,而物理实验证明了真实世界闭环的可行性。
cs.RO / 94 / 2609.15399

Dynamics-Informed Reinforcement Learning for Agile and Energy-Efficient Locomotion of a Monopedal Hopping Quadcopter

动力学信息引导的强化学习用于单足跳跃四旋翼的敏捷与节能运动
Chen, Ruigang, Zhang, Qi, Zhong, Zhicheng, Yun, Zhuorui, Or, Yizhar, Liu, Mingyi
Abstract
Although aerial-legged robots offer combined agility and efficiency, controlling high-speed hopping under complex hybrid dynamics is challenging. Reinforcement Learning (RL) is promising but prone to energy-inefficient "reward hacking". We propose a Dynamics-Informed RL framework for a monopedal hopping quadcopter. By embedding a target Specific Energy into the reward, we constrain the optimization to a physically viable energy manifold, ensuring stable hopping behaviour. By rewarding the phase-consistent behavior, it can encourage bio-inspired stance-phase impulse. Furthermore, penalizing the electro-mechanical power waste induces the motors generate an efficient impulse. This enables the policy to inject energy strictly during spring restitution without heuristic state machines. MuJoCo simulations validate robust height regulation and forward velocity tracking up to 2.0 m/s despite severe attitude-contact coupling. Ultimately, our approach yields a highly agile hopping gait, reducing energy consumption by 82% and 73% compared to hovering baselines and inefficiency baseline, respectively.
Chinese Translation
尽管空中腿式机器人兼具敏捷性和高效性,但在复杂混合动力学下控制高速跳跃仍具有挑战性。强化学习(RL)具有前景,但容易陷入能源效率低下的“奖励黑客”(reward hacking)问题。我们提出了一种用于单足跳跃四旋翼的动力学信息引导强化学习框架。通过将目标比能量(Specific Energy)嵌入奖励中,我们将优化约束在物理可行的能量流形上,确保稳定的跳跃行为。通过奖励相位一致的行为,它可以鼓励仿生支撑相冲量。此外,惩罚机电功率浪费促使电机产生高效冲量。这使得策略能够在弹簧回弹期间严格注入能量,而无需启发式状态机。MuJoCo 仿真验证了在严重的姿态-接触耦合下,仍具有鲁棒的高度调节和最高 2.0 m/s 的前向速度跟踪能力。最终,我们的方法产生了一种高度敏捷的跳跃步态,与悬停基线和不节能基线相比,能耗分别降低了 82% 和 73%。
cs.RO / 95 / 2609.15447

Learning to Exploit Passive Dynamics for Energy-Efficient Target Hopping of a Spring-Legged Quadcopter

学习利用被动动力学实现弹簧腿四旋翼飞行器的能量高效目标跳跃
Chen, Ruigang, Zhang, Qi, Zhong, Zhicheng, Yun, Zhuorui, Or, Yizhar, Liu, Mingyi
Abstract
Combining aerial thrust with spring-loaded hopping makes monopedal quadcopters promising for locomotion over complex terrain, but heuristic proportional-integral-derivative (PID) tuning limits coordination between active thrust and passive contact dynamics. We present a direct estimated-state-to-motor Proximal Policy Optimization (PPO) policy that commands four motors without an explicit hopping state machine or low-level attitude PID. Its reward combines Energy-Manifold Shaping for mass-normalized vertical-energy tracking and apex-state anchoring with Efficiency Shaping, which uses a history-aware power estimator to penalize general power use, impose an additional airborne-power cost, and penalize airborne near-stationarity. In representative hardware runs, the PPO-based control stack reduced cycle-averaged measured electrical power by 30.7% and mean total normalized thrust by 49.8% relative to the tuned PID-based control stack, while retaining repeatable commanded-height hopping and more concentrated landings. These observations are consistent with improved use of passive dynamics and reduced measured electrical demand.
Chinese Translation
将空中推力与弹簧加载跳跃相结合,使单腿四旋翼飞行器在复杂地形上的运动具有前景,但启发式比例-积分-微分(PID)调参限制了主动推力与被动接触动力学之间的协调。我们提出了一种直接从估计状态到电机的近端策略优化(PPO)策略,该策略命令四个电机,而无需显式的跳跃状态机或低级姿态PID。其奖励函数结合了能量流形塑造(Energy-Manifold Shaping)与效率塑造(Efficiency Shaping),前者用于质量归一化的垂直能量跟踪和顶点状态锚定,后者使用历史感知功率估计器来惩罚一般功率使用、施加额外的空中功率成本,并惩罚空中近静止状态。在代表性硬件运行中,基于PPO的控制栈相对于调优的基于PID的控制栈,将周期平均测量电功率降低了30.7%,平均总归一化推力降低了49.8%,同时保持了可重复的命令高度跳跃和更集中的着陆。这些观察结果与改进的被动动力学利用和降低的测量电力需求一致。
cs.RO / 96 / 2609.15455

InterSocialBench: Benchmarking Human and LLM Preferences for Companion-Robot Social Behavior

InterSocialBench:人类与LLM对陪伴机器人社交行为偏好的基准测试
Xu, Yaodan, Guo, Boyang, Gu, Yuqing, Zhang, Qingxin, Deng, Yiwen, Liu, Meng, Li, Lintian
Abstract
Companion robots face everyday situations in which several feasible behaviors may be appropriate, yet different people prefer different responses. We introduce InterSocialBench, a benchmark of 210 domestic scenarios and 18 high-level behaviors, pairing judgments from 100 human participants with 23,520 responses from seven large language models under 16 personality conditions. Each human annotation preserves a preferred action alongside explicitly appropriate and inappropriate candidates. A structured construction pipeline covers behavioral alternatives, competing situational cues, and relevant history and future tasks. Evaluation distinguishes preferred-choice agreement from explicit rejection, using scenario-grouped splits for trainable predictors. Simple frequency and persona-voting baselines illustrate these objectives. Across the tested prompts, model and human behavior distributions differ, and the diversity gap remains after matching response counts: humans exhibit 4.68 distinct choices per scenario, compared with 2.06--3.46 for the models. Human scenario-level plurality agreement is 51.5%, describing disagreement rather than a universal prediction ceiling. InterSocialBench supports evaluating social behavior selection without replacing individual judgments with a single consensus label.
Chinese Translation
陪伴机器人面临日常情境,其中多种可行行为可能都合适,但不同的人偏好不同的回应。我们提出了InterSocialBench,一个包含210个家庭场景和18种高层行为的基准,将100名人类参与者的判断与七种大型语言模型在16种人格条件下的23,520个回应配对。每个人类标注保留了一个偏好动作以及明确合适和不合适的候选动作。一个结构化的构建流程涵盖了行为替代方案、竞争性情境线索以及相关的历史和未来任务。评估区分了偏好选择一致性与明确拒绝,并对可训练预测器使用了按场景分组的划分。简单的频率和人格投票基线说明了这些目标。在测试的提示中,模型和人类的行为分布不同,并且在匹配回应数量后多样性差距依然存在:人类在每个场景中表现出4.68种不同的选择,而模型为2.06--3.46种。人类在场景层面的多数一致性为51.5%,这描述的是分歧,而非普遍预测上限。InterSocialBench支持评估社交行为选择,而无需用单一共识标签替代个体判断。
cs.RO / 97 / 2609.15475

P-POSEMEM: Projective Semantic Memory for Consistent Language Grounding under Pose-Graph Rewrites

P-POSEMEM:用于位姿图重写下一致语言定位的投影语义记忆
Sier, Ha, Salmasi, Ali, Xu, Mengya, Zhang, Haizhou, Lu, Jie, Zou, Zhuo, Yu, Xianjia, Westerlund, Tomi
Abstract
A robot following language instructions needs its semantic memory to keep naming the same physical object while the SLAM pose graph underneath is optimized, loop-closed and compressed. Maps committing each detection to a world coordinate cannot: a closure moves the anchor it was measured from, or the solver marginalizes that anchor, and the query then selects a different object although both graphs represent the same posterior. P-POSEMEM stores each observation as an immutable event at its birth keyframe, retains the Bayes-tree elimination conditional of every marginalized keyframe, and integrates the semantic likelihood over the reconstructed joint posterior of poses, anchors and identities. Dproj, the total-variation defect between the language-goal distributions of inference-equivalent full and marginalized graphs, measures this directly. Over 40 HM3DSem scenes and 112,000 queries, P-POSEMEM reproduces the full-graph oracle (Dproj = 0) and reduces goal flips against every memory-reducing baseline. On an eight-run campaign whose 761 closures rewrote the map by up to 47 m, Dproj stays below 10^-13 with 0/288 goal flips when elimination follows the closures, where every ablation and a coordinate committed at insertion flip goals it does not; under a live bounded solver the same memory flips 23/288 against 53 for that frozen coordinate. A pre-registered negative control is detected by Dproj while leaving calibration error and navigation success unchanged, indicating that these measures capture distinct failure modes. Retrieval is held fixed by a shared frozen detector, isolating the gain to memory consistency. Code and data: https://anonymous.4open.science/r/posemem-2328/.
Chinese Translation
机器人在遵循语言指令时,需要其语义记忆在底层的 SLAM 位姿图被优化、回环闭合和压缩时,仍能对同一物理对象保持相同的命名。将每次检测都提交到世界坐标的地图无法做到这一点:一次回环会移动其测量所依据的锚点,或者求解器边缘化该锚点,然后查询会选择不同的对象,尽管两个图表示相同的后验。P-POSEMEM 将每个观测存储为其出生关键帧处的一个不可变事件,保留每个被边缘化关键帧的 Bayes-tree 消元条件,并在位姿、锚点和标识的重建联合后验上集成语义似然。Dproj,即推理等价的完整图和边缘化图的语言目标分布之间的总变差缺陷,直接度量了这一点。在 40 个 HM3DSem 场景和 112,000 次查询上,P-POSEMEM 复现了完整图 oracle(Dproj = 0),并相对于每个减少记忆的基线减少了目标翻转。在一个八次运行的活动中,其 761 次回环最多将地图重写了 47 米,当消元跟随回环时,Dproj 保持在 10^-13 以下,目标翻转为 0/288,而每个消融实验和在插入时提交的坐标都会使其目标翻转;在实时有界求解器下,同一记忆翻转 23/288,而该冻结坐标翻转 53 次。Dproj 检测到一个预注册的阴性对照,同时保持校准误差和导航成功不变,表明这些度量捕获了不同的失败模式。检索由共享的冻结检测器固定,从而将增益隔离到记忆一致性。代码和数据:https://anonymous.4open.science/r/posemem-2328/。
cs.RO / 98 / 2609.15509

StereoPatch: Patch-Aligned RGB-Depth Fusion for Spatial Perception in Robot Manipulation

StereoPatch:面向机器人操作空间感知的图块对齐RGB-深度融合
Zhou, Yanan, Qian, Zhaoyan, Zhao, James, Zhi, Weiming
Abstract
Recent advances in robot imitation learning have produced visuomotor policies that predict actions directly from visual observations. Yet visually similar scenes can require different actions as target position, object height, or contact geometry changes. Pretrained RGB features may map these geometrically distinct states to similar policy inputs, while simply adding depth requires the policy to learn RGB-depth correspondence from the same limited demonstrations used to learn control. We introduce StereoPatch, a patch-aligned RGB-depth representation that binds registered metric geometry directly to the RGB patches used for action prediction. On a shared 2-D patch grid, asymmetric cross-attention incorporates depth information into the corresponding RGB features before action decoding. The resulting StereoPatch Tokens provide a geometry-aware visual representation that can condition general visuomotor policies without changing their underlying learning objectives. Across six real-robot tasks, StereoPatch achieves higher closed-loop success than appearance-only, geometry-only, raw RGB-D, and late-fusion baselines. Additional experiments across three simulation suites evaluate compatibility across visuomotor policy architectures, spatial generalization, and operating limits. Results suggest that resolving control-relevant geometric ambiguity benefits from aligning depth directly with the visual features used for action prediction, rather than supplying it as an independent modality. Project page: https://aus.bot/research/stereopatch/.
Chinese Translation
最近机器人模仿学习的进展产生了直接从视觉观察预测动作的视觉运动策略。然而,当目标位置、物体高度或接触几何发生变化时,视觉上相似的场景可能需要不同的动作。预训练的RGB特征可能将这些几何上不同的状态映射为相似的策略输入,而简单添加深度则要求策略从用于学习控制的同样有限的演示中学习RGB-深度对应关系。我们提出了StereoPatch,一种图块对齐的RGB-深度表示,它将配准的度量几何直接绑定到用于动作预测的RGB图块上。在共享的二维图块网格上,非对称交叉注意力在动作解码之前将深度信息融入对应的RGB特征中。所得的StereoPatch Tokens提供了一种几何感知的视觉表示,可以在不改变其基本学习目标的情况下调节通用视觉运动策略。在六个真实机器人任务中,StereoPatch实现了比仅外观、仅几何、原始RGB-D和后期融合基线更高的闭环成功率。在三个仿真套件上的额外实验评估了视觉运动策略架构间的兼容性、空间泛化和操作限制。结果表明,解决与控制相关的几何模糊性,受益于将深度直接与用于动作预测的视觉特征对齐,而不是将其作为独立模态提供。项目页面:https://aus.bot/research/stereopatch/。
cs.RO / 99 / 2609.15570

DIDO: Distilling Interaction-Centric Dynamics into One-Step Denoising for World Action Models

DIDO:将交互中心动力学蒸馏到世界动作模型的一步去噪中
Lyu, Jing, Bai, Shuanghao, Xiao, Runze, Liao, Zhenyu, Tan, Wenxing, Tang, Zihan, Shi, Ruochuan, Peng, Cheng, Ji, Yuheng, Wang, Yihao, Chen, Badong, Wang, Pengwei, Wang, Zhongyuan, Zhao, Xiaoguang
Abstract
World Action Models (WAMs) use video generation models to predict future visual dynamics for robotic manipulation, but iterative denoising introduces additional latency for closed-loop control. We empirically find that visual content converges at different rates during denoising. Static background structure forms early, whereas the gripper and manipulated object remain blurry after the first step, with their interaction dynamics emerging only through subsequent denoising. Consequently, naively truncating a multi-step video model to one step preserves scene structure but loses the interaction-centric dynamics most critical for manipulation. To address this issue, we propose DIDO, which distills the converged dynamics of a multi-step video model into a single denoising step. DIDO combines distribution matching distillation with interaction-centric representation guidance. Beyond compressing multi-step generation into one forward pass, DIDO explicitly models the gripper, manipulated object, and their interaction using supervised bounding-box visual reasoning tokens. Additionally, DIDO aligns the target object's representations across multiple model layers with features from a pretrained DINOv3 encoder. This interaction-centric guidance helps the distilled model preserve both the relevant entities and their future dynamics in a single step, while substantially reducing inference latency. DIDO achieves an average success rate of 99.0\% on LIBERO, 76.6\% on LIBERO-Plus, and 92.0\% on RoboTwin, while also demonstrating effective transfer to long-horizon and generalization tasks in real-world robotic manipulation.
Chinese Translation
世界动作模型(WAMs)利用视频生成模型预测机器人操作的未来视觉动力学,但迭代去噪会给闭环控制带来额外延迟。我们通过实验发现,视觉内容在去噪过程中以不同速率收敛:静态背景结构很早就形成,而夹爪和被操作物体在第一步之后仍然模糊,它们的交互动力学只有通过后续去噪才逐渐显现。因此,单纯地将多步视频模型截断为一步,虽能保留场景结构,却会丢失对操作最为关键的、以交互为中心的动力学。为解决这一问题,我们提出 DIDO,它将多步视频模型已收敛的动力学蒸馏到单步去噪中。DIDO 将分布匹配蒸馏与以交互为中心的表示引导相结合。除了将多步生成压缩为一次前向传播外,DIDO 还使用有监督的边界框视觉推理 token 显式建模夹爪、被操作物体及其交互。此外,DIDO 将目标物体在多个模型层中的表示与预训练 DINOv3 编码器的特征对齐。这种以交互为中心的引导有助于蒸馏模型在单步中同时保留相关实体及其未来动力学,同时大幅降低推理延迟。DIDO 在 LIBERO 上达到 99.0% 的平均成功率,在 LIBERO-Plus 上达到 76.6%,在 RoboTwin 上达到 92.0%,同时还展示了向真实世界机器人操作中的长时程任务和泛化任务的有效迁移。
cs.RO / 100 / 2609.15587

An Information-Space Perspective to Scene Graph Sufficiency for Robotic Task Planning

从信息空间视角看机器人任务规划中的场景图充分性
Sakçak, Başak, Verdoja, Francesco
Abstract
Planning in complex environments requires task specifications grounded in representations that capture objects, relations, and affordances; scene graphs meet this need, but their size in large environments hinders efficient planning. While task-aware pruning and hierarchical abstractions have been explored, a general, task-centric formalization of what constitutes a sufficient scene graph for planning remains open. This paper provides such a formalization by modeling planning over scene graphs within an information-spaces framework through the definition of scene graph transition systems and relevant action semantics for navigation and manipulation. We then introduce derived scene graphs via information mappings that merge and prune nodes and induce quotient transition systems augmented with motion primitives to capture higher-level actions over merged graph nodes. Sufficiency is characterized by two conditions: (i) the information mapping yields a deterministic quotient, and (ii) the task is well-posed over derived traces, ensuring plans found on the derived model are feasible on the maximal system. We illustrate the framework using a task over an example environment, showing both sufficient and insufficient reduced scene graphs.
Chinese Translation
复杂环境中的规划需要基于能够捕获对象、关系和可供性的表示来制定任务规范;场景图满足了这一需求,但在大型环境中其规模阻碍了高效规划。尽管任务感知的剪枝和分层抽象已被探索,但关于什么构成用于规划的充分场景图的通用、以任务为中心的形式化仍然悬而未决。本文通过定义场景图转移系统以及导航和操作的相关动作语义,在信息空间框架内对基于场景图的规划进行建模,提供了这样的形式化。然后,我们通过信息映射引入派生场景图,这些映射合并和剪枝节点,并诱导出用运动原语增强的商转移系统,以捕获合并图节点上的高层动作。充分性由两个条件刻画:(i) 信息映射产生确定性商,(ii) 任务在派生轨迹上是适定的,确保在派生模型上找到的计划在最大系统上是可行的。我们用一个示例环境中的任务来说明该框架,展示了充分和不充分的约简场景图。
cs.RO / 101 / 2609.15631

Flow-Matched Motion Priors: Online Optimal-Transport Rewards for Imitation Learning

流匹配运动先验(Flow-Matched Motion Priors):面向模仿学习的在线最优传输奖励
Zou, Yilin, Liu, Chenghua, Wu, Chenglong, Jiang, Fanghua
Abstract
Learning a motion prior requires a reward that guides a policy from its current behavior toward demonstrated motion. Adversarial Motion Priors (AMP) provide such a reward with a discriminator. However, adversarial objectives can become uninformative when policy and expert supports are far apart. A naive use of optimal transport (OT) averages matched expert successors into a barycentric target. Averaging across gait phases can weaken the target's joint motion. We introduce Flow-Matched Motion Priors (FMP), an online scalar reward learned from paths connecting current rollout histories to an expert motion bank. Entropic OT supplies the coupling. Before each policy update, we train a neural potential with flow matching (FM) along the rollout-to-expert paths, endpoint-gradient supervision, and relative-value calibration. The actor receives only physical observations and the reward remains a scalar, as in AMP. Controlled reward-model experiments show substantially better generalization beyond the fitting rollout than value-only or endpoint-only fitting. On Unitree G1, matched 50-million-transition experiments compare FMP with AMP, a barycentric OT reward, and nested ablations under demonstration and fixed-pose initialization. FMP produces stable forward walking at 0.727 m/s from demonstration resets and 0.338 m/s from a fixed default pose. In the fixed-pose condition, it incurs 129 falls versus 243 for the endpoint-only control. Against a static score-gradient teacher, dynamic FM reduces score-increment error at interpolation fractions 0.25 and 0.50 while using 29% less offline fitting time.
Chinese Translation
学习运动先验需要一个奖励,该奖励引导策略从其当前行为朝向演示运动。对抗运动先验(AMP)通过判别器提供这样的奖励。然而,当策略和专家的支撑集相距甚远时,对抗目标可能变得无信息量。朴素地使用最优传输(OT)将匹配的专家后继平均为重心目标。跨步态阶段平均会削弱目标的关节运动。我们提出流匹配运动先验(FMP),一种从连接当前推演历史到专家运动库的路径中学习的在线标量奖励。熵最优传输提供耦合。在每次策略更新之前,我们沿着推演到专家的路径,使用流匹配(FM)、端点梯度监督和相对值校准来训练神经势函数。与AMP一样,演员只接收物理观测,奖励保持为标量。受控奖励模型实验表明,与仅值拟合或仅端点拟合相比,其在拟合推演之外的泛化能力显著更好。在Unitree G1上,匹配的5000万次转换实验在演示和固定姿态初始化下将FMP与AMP、重心OT奖励以及嵌套消融进行比较。FMP在演示重置下产生0.727 m/s的稳定前向行走,在固定默认姿态下产生0.338 m/s的稳定前向行走。在固定姿态条件下,它导致129次摔倒,而仅端点控制组为243次。与静态分数梯度教师相比,动态FM在插值分数0.25和0.50处降低了分数增量误差,同时减少29%的离线拟合时间。
cs.RO / 102 / 2609.15667

Tracking the Ground: Online Lidar Identification of Robot-Induced Soil Deformation in Agricultural Environments

追踪地面:农业环境中机器人诱发土壤变形的在线激光雷达识别
Montagnon, Tom, Laconte, Johann, Thuilot, Benoit, Cho, Wonjae, Lenain, Roland
Abstract
Agriculture faces many challenges, and robotic systems can play an important role in addressing them by improving the efficiency and sustainability of field operations. Among these challenges, preserving soil health is a critical concern, as vehicle-soil interactions can degrade the soil structure and produce unwanted surface deformation. A key step toward soil-aware robotics is to explicitly account for how vehicle traffic deforms the ground, yet soil state is typically not treated as a variable. We address this gap by proposing a framework to quantify traffic-induced soil deformation and estimate its evolution online from lidar observations. The method relies on a reduced-order parametric model that represents the soil behavior via physically interpretable parameters, yielding a continuously updated and observable representation of soil state. Experiments conducted in different soil conditions demonstrate the ability of the approach to capture deformation induced by the robot. By making soil response measurable and interpretable during operation, the proposed framework establishes a basis for soil-aware robotic operation, in which the estimated state can be exploited to adapt robotic behaviors in order to reduce soil degradation.
Chinese Translation
农业面临诸多挑战,机器人系统可以通过提高田间作业的效率和可持续性,在应对这些挑战中发挥重要作用。在这些挑战中,保护土壤健康是一个关键问题,因为车辆-土壤相互作用会破坏土壤结构并产生不期望的地表变形。迈向土壤感知机器人技术的关键一步是明确考虑车辆通行如何使地面变形,然而土壤状态通常并未被视为一个变量。为弥补这一空白,我们提出一个框架,用于量化通行引起的土壤变形,并基于激光雷达观测在线估计其演化。该方法依赖于一个降阶参数模型,通过物理上可解释的参数表征土壤行为,从而得到土壤状态的持续更新且可观测的表示。在不同土壤条件下进行的实验表明,该方法能够捕捉机器人引起的变形。通过使土壤响应在作业过程中可测量且可解释,所提出的框架为土壤感知机器人作业奠定了基础,其中估计的状态可用于调整机器人行为,以减少土壤退化。
cs.RO / 103 / 2609.15680

Volumetric Harmonic Field Navigation for Quadrotors

四旋翼飞行器的体谐波场导航
Jia, Shuxiu, Mukherjee, Amartya, Yuan, Yating, Liu, Jun
Abstract
Quadrotor navigation in cluttered 3-D environments requires global guidance while local motion remains subject to collision and motion limits. Harmonic potentials provide dense guidance from a global boundary value problem, but coupling a volumetric harmonic field to constrained physical quadrotor motion remains an open experimental problem. We couple a precomputed volumetric harmonic field with a constrained predictive planner that queries the field at predicted positions instead of extracting a global reference path. In Structured 3-D tests, harmonic guidance yields larger minimum clearance and lower RMS jerk than matched Dijkstra guidance, at the cost of longer paths; the same pattern remains when both methods use the same passage. Long maze tests span routes far beyond one prediction horizon, and Crazyflie trials validate physical execution. To the best of our knowledge, this is the first physical quadrotor demonstration of volumetric harmonic field navigation. The results show that globally constructed harmonic guidance can directly support local constrained motion generation on a physical quadrotor.
Chinese Translation
在杂乱的三维环境中进行四旋翼导航需要全局引导,同时局部运动仍受碰撞和运动限制的约束。谐波势从全局边值问题提供密集引导,但将体谐波场与受约束的物理四旋翼运动耦合仍是一个开放的实验问题。我们将预计算的体谐波场与受约束的预测规划器耦合,该规划器在预测位置查询场,而不是提取全局参考路径。在结构化三维测试中,谐波引导比匹配的 Dijkstra 引导产生更大的最小间隙和更低的均方根加加速度,代价是路径更长;当两种方法使用相同通道时,这种模式仍然存在。长迷宫测试的路线远远超过一个预测时域,Crazyflie 试验验证了物理执行。据我们所知,这是体谐波场导航的首次物理四旋翼演示。结果表明,全局构建的谐波引导可以直接支持物理四旋翼上的局部约束运动生成。
cs.RO / 104 / 2609.15716

Continuous Manifold-Decomposed Impedance Retargeting for Contact-Rich Imitation Learning

面向富接触模仿学习的连续流形分解阻抗重定向
Liu, Jiahao, Kawaharazuka, Kento, Makabe, Tasuku, Okada, Kei
Abstract
CMDIR extends Manifold-Decomposed Impedance Retargeting (MDIR) to transform fixed-impedance demonstrations into continuous variable-impedance controllers, which can also serve as structured supervision for imitation learning. Continuous Task-Manifold Impedance Representation (TMIR) pairs an evolving task frame with controller instructions. Demo-relative Compromise dynamics retain moving-basis transport and control/physical metric mismatch, yielding displacement, reaction-impulse, and perturbation-sensitivity criteria. Quality-to-Fast automatically compiles a solver structure from development paths within a predefined finite space, re-instantiates that structure for each demonstration, and certifies the resulting candidate by multi-resolution evaluation. Across 225 retargeted-controller trials in three real contact tasks, full CMDIR improves mean task-proxy retention and reduces mean pose deviation, force fluctuation, and peak force relative to discrete MDIR. FastMPO achieves a $5.8$--$9.4\times$ speedup over C-MPO with comparable closed-loop outcomes. Downstream experiments demonstrate learnability of the complete TMIR supervision interface; lower force fluctuation and peak force are observed among successful executions, while completion reliability remains uneven across tasks and environments.
Chinese Translation
CMDIR扩展了流形分解阻抗重定向(MDIR),将固定阻抗演示转化为连续变阻抗控制器,这也可以作为模仿学习的结构化监督。连续任务流形阻抗表示(TMIR)将演进的任务框架与控制器指令配对。演示相对折衷动力学保留了移动基座传输和控制/物理度量失配,产生了位移、反作用冲量和扰动敏感性准则。Quality-to-Fast在预定义的有限空间内从开发路径自动编译求解器结构,为每个演示重新实例化该结构,并通过多分辨率评估认证生成的候选方案。在三个真实接触任务的225次重定向控制器试验中,完整CMDIR相对于离散MDIR提高了平均任务代理保持率,并降低了平均位姿偏差、力波动和峰值力。FastMPO相比C-MPO实现了5.8–9.4倍的加速,且闭环结果相当。下游实验证明了完整TMIR监督接口的可学习性;在成功执行中观察到较低的力波动和峰值力,而完成可靠性在不同任务和环境间仍不均衡。
cs.RO / 105 / 2609.15726

Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands

Bench2Dex:跨灵巧手的视触觉双手灵巧操作基准测试
Yang, Zhenjie, Zhang, Yideng, Zhang, Dongjie, Jiang, Chenyu, Liu, Xianshuai, Li, Yufeng, Ge, Zuhao, Jiao, Xingyu, Zhang, Zheng, He, Kaiyu, Wang, He, Zhong, Yuwen, Deng, Yi, Jiang, Muyun, Huang, Xianliang, Su, Haisheng, Zhang, Donghang, Zhang, Jian, Yang, Xue, Li, Hongyang, Wu, Zuxuan, Jiang, Yu-Gang, Jia, Xiaosong, Yan, Junchi
Abstract
Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not converged to a common design. Dexterous hands differ in finger structure, contact surfaces, and sensor layouts, while simulated tactile signals still differ from measurements produced by physical sensors. These factors make it difficult to study visuo-tactile manipulation across diverse dexterous hands within a consistent experimental setting. We present Bench2Dex, a simulation benchmark for visuo-tactile bimanual manipulation across 12 dexterous hands. We adapt existing robot models with a shared simulated tactile interface that converts local contact geometry into image-like tactile observations. The interface provides a consistent observation format across different hand morphologies without attempting to reproduce the output of a specific physical tactile sensor. Bench2Dex includes 26 bimanual manipulation tasks that involve tool use, articulated-object interaction, and multi-stage manipulation, together with about 1.3K human-teleoperated demonstrations. The benchmark provides synchronized visual, tactile, proprioceptive, action, and object-state observations, together with executable task metrics. For robustness, we group seven perturbation types into invariance axis, where the correct action does not change, and equivariance axis, where the correct action changes together with the perturbation. We evaluate ACT, Diffusion Policy, pi0.5, and GR00T N1.5 on Bench2Dex and report their performance and failure modes. Bench2Dex is meant as a platform for studying visuo-tactile learning across dexterous hands. It does not assume that simulated tactile observations can replace real tactile sensing; it offers a shared setting for algorithm development while tactile hardware and simulation models are still evolving.
Chinese Translation
触觉感知提供了仅靠视觉难以推断的接触信息,但灵巧手的触觉硬件尚未统一为通用设计。灵巧手在手指结构、接触表面和传感器布局上各不相同,而模拟触觉信号仍与物理传感器产生的测量结果存在差异。这些因素使得在一致的实验设置下研究跨不同灵巧手的视触觉操作变得困难。我们提出了 Bench2Dex,一个用于跨 12 种灵巧手的视触觉双手操作的仿真基准。我们通过一个共享的模拟触觉接口改编现有机器人模型,该接口将局部接触几何转换为类图像触觉观测。该接口在不同手部形态之间提供一致的观测格式,而不试图复现特定物理触觉传感器的输出。Bench2Dex 包含 26 项双手操作任务,涉及工具使用、铰接物体交互和多阶段操作,以及约 1.3K 人类遥操作演示。该基准提供同步的视觉、触觉、本体感觉、动作和物体状态观测,以及可执行的任务指标。为了评估鲁棒性,我们将七种扰动类型分为不变轴和等变轴,其中不变轴下正确动作不改变,等变轴下正确动作随扰动一起改变。我们在 Bench2Dex 上评估了 ACT、Diffusion Policy、pi0.5 和 GR00T N1.5,并报告了它们的性能和失败模式。Bench2Dex 旨在作为一个研究跨灵巧手的视触觉学习的平台。它并不假设模拟触觉观测可以替代真实触觉感知;它为触觉硬件和仿真模型仍在演进的过程中提供了一个用于算法开发的共享设置。
cs.RO / 106 / 2609.15770

JEPLO: Joint-Embedding Predictive Learning for LiDAR-Based Legged Locomotion

JEPLO:面向基于LiDAR的腿足式运动的联合嵌入预测学习
Yuan, Qihao, Qiu, Yixuan, Cao, Ziyu, Cao, Ming, Li, Kailai
Abstract
Light detection and ranging (LiDAR) remains less explored than RGB-D sensing for perceptive legged locomotion, and existing LiDAR-based approaches often rely on explicit mapping. We present JEPLO (Joint-Embedding Predictive learning for legged LOcomotion), a single-stage learning framework for mapping-free, LiDAR-based perceptive locomotion for legged robots. We introduce a proprio-exteroceptive JEPA (PE-JEPA) world model to learn predictive egocentric terrain representations from onboard observations, including raw LiDAR scans. A concurrent JEPA-teacher-student (CJTS) pipeline is further proposed to train a locomotion policy informed by JEPA latent representations in simulation using deep reinforcement learning with a simple reward formulation. The framework achieves successful sim-to-real transfer, enabling omnidirectional traversal of diverse terrains, including long staircases and high boxes, with lightweight onboard computation. Evaluations demonstrate greater robustness than existing perceptive locomotion frameworks, particularly under degraded perception caused by occlusion, sparsity and noise. Further analysis validates JEPLO's ability to retain task-relevant information under these challenging conditions. We open-source our implementation, experimental datasets, and hardware setup designs https://github.com/ASIG-X/JEPLO.
Chinese Translation
在感知腿足式运动方面,光探测与测距(LiDAR)相较于RGB-D感知仍较少被探索,并且现有的基于LiDAR的方法通常依赖于显式建图。我们提出JEPLO(面向腿足式运动的联合嵌入预测学习),一种用于腿足式机器人的无建图、基于LiDAR的感知运动学习的单阶段学习框架。我们引入本体-外界感知JEPA(PE-JEPA)世界模型,从包括原始LiDAR扫描在内的机载观测中学习预测性的以自我为中心的地形表示。进一步提出了一种并发的JEPA-教师-学生(CJTS)流程,在仿真中使用深度强化学习和简单的奖励公式,训练由JEPA潜在表示引导的运动策略。该框架实现了成功的仿真到现实迁移,能够以轻量级机载计算实现多种地形的全向穿越,包括长楼梯和高箱子。评估表明,该框架比现有的感知运动框架具有更强的鲁棒性,特别是在由遮挡、稀疏性和噪声引起的感知退化情况下。进一步分析验证了JEPLO在这些具有挑战性的条件下保留任务相关信息的能力。我们开源了我们的实现、实验数据集和硬件设置设计:https://github.com/ASIG-X/JEPLO。
cs.RO / 107 / 2609.15840

Uncertainty-Guided Sparse Refinement for Action Chunking Transformer Policies

面向动作分块Transformer策略的不确定性引导稀疏细化
Wang, Chenyang, Wang, Yuntian, Yang, Xiaoxiong, Jiang, Dingde, Liu, Siao, Liu, Yang
Abstract
Learning chunk-based visuomotor policies for long-horizon robot manipulation remains challenging. Recent action-chunking methods have shown promising performance by predicting temporally extended action sequences. However, their failures are often dominated by prediction errors at a small number of critical timesteps rather than uniformly poor predictions across the entire action chunk, making uniform refinement inefficient and insufficiently targeted. To address this bottleneck, we propose Uncertainty-Guided Refinement (UGR), a sparse refinement framework for chunk-based visuomotor policies. Specifically, UGR follows a coarse-to-refine design: it first predicts a full action chunk, estimates per-step temporal uncertainty from the coarse hidden states, and applies residual correction only to the most uncertain timesteps selected by a binary mask. The uncertainty branch is decoupled from the coarse action predictor, enabling clean attribution of the refinement gains to uncertainty-guided correction rather than additional predictor capacity. Extensive experiments on five dual-arm manipulation tasks from the RoboTwin benchmark show that UGR achieves the best success rate on four tasks, improves over the ACT baseline by up to 13% absolute, and outperforms both full-chunk and position-agnostic block refinement in ablation studies.
Chinese Translation
学习基于分块的视觉运动策略用于长时程机器人操作仍然具有挑战性。近期的动作分块方法通过预测时间上扩展的动作序列展现了有前景的性能。然而,它们的失败通常主要由少数关键时间步的预测误差主导,而不是整个动作分块中均匀的较差预测,这使得均匀细化效率低下且针对性不足。为了解决这一瓶颈,我们提出了不确定性引导细化(UGR),一种用于基于分块的视觉运动策略的稀疏细化框架。具体来说,UGR遵循从粗到细的设计:它首先预测完整的动作分块,从粗粒度隐藏状态估计每个时间步的时间不确定性,并仅对由二值掩码选择的最不确定的时间步应用残差校正。不确定性分支与粗动作预测器解耦,使得细化增益可以清晰地归因于不确定性引导的校正,而非额外的预测器容量。在RoboTwin基准的五个双臂操作任务上进行的大量实验表明,UGR在四个任务上取得了最佳成功率,相对于ACT基线提升了高达13%的绝对值,并且在消融研究中优于全分块和位置无关的块细化。
cs.RO / 108 / 2609.15861

DuctAM: A Duct-Assisted Quadrotor-Based Aerial Manipulator Enabling High-Force Push-and-Pull Interactions

DuctAM:一种基于四旋翼的涵道辅助空中机械臂,可实现高力推拉交互
Wang, Yi, Jin, Rui, Xu, Xinhang, Jin, Haotian, Liu, Ruiyang, Yang, Yizhuo, Xie, Lihua
Abstract
Uncrewed Aerial Manipulators (UAMs) extend the capabilities of Uncrewed Aerial Vehicles (UAVs) from perception to physical interaction. Among various aerial interactions, push-and-pull operations are fundamental manipulation primitives that require sustained horizontal forces while maintaining stable flight. In this paper, we propose DuctAM, a compact aerial manipulation platform that enhances horizontal force capability for push-and-pull interactions using two ducted fans integrated along the quadrotor interaction axis. An attitude-force decoupled control scheme enables controllable horizontal forces without requiring large attitude changes. Extensive real-world experiments are conducted to validate the DuctAM. Figure-eight trajectory tracking experiments demonstrate stable flight and accurate motion control in both quad and duct modes. Force-measurement experiments quantify the decoupled longitudinal force capability of DuctAM. Finally, representative push-and-pull interaction tasks, including cart pushing, door closing, and drawer opening, verify the practical effectiveness of DuctAM. The results show that DuctAM achieves significantly improved horizontal interaction force capability while maintaining stable flight compared with conventional UAVs.
Chinese Translation
无人空中机械臂(UAMs)将无人飞行器(UAVs)的能力从感知扩展到物理交互。在各种空中交互中,推拉操作是基本的操作基元,需要在保持稳定飞行的同时施加持续的水平力。本文提出 DuctAM,这是一种紧凑型空中操作平台,通过沿四旋翼交互轴集成的两个涵道风扇,增强推拉交互中的水平力能力。一种姿态-力解耦控制方案无需大幅姿态变化即可实现可控水平力。我们进行了大量真实世界实验来验证 DuctAM。8字形轨迹跟踪实验表明,DuctAM 在四旋翼模式和涵道模式下均具有稳定飞行和精确运动控制能力。力测量实验量化了 DuctAM 的解耦纵向力能力。最后,典型推拉交互任务(包括推车、关门和开抽屉)验证了 DuctAM 的实际有效性。结果表明,与传统无人机相比,DuctAM 在保持稳定飞行的同时显著提升了水平交互力能力。
cs.RO / 109 / 2609.15870

WLA$^3$: World Latent Action Modeling for Semantics, Dynamics, and Kinematics

WLA$^3$:面向语义、动力学与运动学的世界潜在动作建模
Liu, Peidong, Xiang, Zhiyuan, Li, Mingyang, Li, Wenhao, Zhang, Jiale, Sun, Jiahao, Li, Jiawei
Abstract
Scaling generalist policy models with heterogeneous data is limited by the lack of unified, low-noise action supervision. Human egocentric videos are abundant, but only a small fraction comes with high-quality hand-action labels. Observed world transitions offer a common source of action-related supervision across data sources. We introduce WLA$^3$ (World Latent Action Modeling for Semantics, Dynamics, and Kinematics), a unified generalist policy model framework built around representations learned by a World Latent Action Model (WLAM). WLAM first learns how multimodal world states change over a local interval, encoding synchronized camera views and available embodiment-state changes into a compact local latent action and a richer transition feature. Reconstruction from partial modalities and consistency across overlapping windows encourage robust transition representations. WLA$^3$ reuses them across semantics, dynamics, and kinematics: local latent actions support action-sensitive physical-dynamics modeling, segment-level features directly supervise the VLM through a Semantic Latent Aggregate (SLA), and an action expert jointly predicts latent actions together with embodiment-specific robot controls. Human videos provide scalable transition supervision, while robot trajectories ground the shared representation in executable native controls. On LARYBench, the final 32D latent action reaches 67.89\% average classification accuracy. WLA$^3$ achieves 81.9% average success across six real-robot tasks versus 66.2% for $\pi_{0.5}$. Performance improves as generalist policy model mid-training data scales, and human videos support human-to-robot transfer. Project page can be found at https://wla-3.github.io/.
Chinese Translation
利用异质数据扩展通用策略模型受限于缺乏统一、低噪声的动作监督。人类第一人称视角视频丰富,但只有一小部分具有高质量的手部动作标签。观察到的世界状态转换提供了跨数据源的动作相关监督的通用来源。我们引入 WLA$^3$(面向语义、动力学与运动学的世界潜在动作建模),这是一个统一的通用策略模型框架,其核心是由世界潜在动作模型(WLAM)学习到的表示。WLAM 首先学习多模态世界状态在局部时间间隔内如何变化,将同步的相机视图和可用的具身状态变化编码为紧凑的局部潜在动作和更丰富的转换特征。从部分模态重建以及重叠窗口间的一致性促进了鲁棒的转换表示。WLA$^3$ 在语义、动力学和运动学中复用这些表示:局部潜在动作支持动作敏感物理动力学建模,片段级特征通过语义潜在聚合(SLA)直接监督 VLM,动作专家联合预测潜在动作和具身特定的机器人控制。人类视频提供可扩展的转换监督,而机器人轨迹将共享表示锚定在可执行的原生控制中。在 LARYBench 上,最终 32 维潜在动作达到 67.89% 的平均分类准确率。WLA$^3$ 在六个真实机器人任务中达到 81.9% 的平均成功率,而 $\pi_{0.5}$ 为 66.2%。随着通用策略模型中期训练数据规模的扩大,性能得到提升,且人类视频支持从人到机器人的迁移。项目页面见 https://wla-3.github.io/。
cs.RO / 110 / 2609.15895

Goal-Oriented Communications for Physical AI: Design and Testbed

面向Physical AI的目标导向通信:设计与测试床
Chen, Shutong, Zhang, Wenkai, Aijaz, Adnan, Guo, Miao, Deng, Yansha
Abstract
Physical AI relies on frequently-updated, latency-sensitive video stream to perceive, reason, and interact with the physical world, resulting in strict latency requirements with much higher data volumes that existing 5G networks cannot support. Goal-oriented communication (GoC) offers as a promising approach to solve this challenge by transmitting only task-relevant semantic representations. However, existing GoC frameworks were mainly evaluated in the simulations while their effectiveness has never been validated in a practical deployment of physical AI application. In this work, we develop an end-to-end GoC testbed for Physical AI, which connects a PiPER robot arm equipped with an RGB-D camera and a 5G modem to an NVIDIA Jetson AGX Orin edge server through a 5G OpenAirInterface network. We propose and implement three GoC frameworks that transmit 3D bounding boxes, 2D scene graphs, and 3D scene graphs, as three types of semantic representations, respectively. They share the common functional modules designed for closed-loop Physical AI applications, including semantic extraction, full stack 5G transmission, language model inference, digital twin validation, and robotic control. Extensive experiments on our testbed show that our GoC frameworks reduce the task completion time by up to 52.6% and improve task success probability by up to 45%, compared to the traditional framework that periodically transmits the raw image data. These results validate the practical effectiveness of our GoC framework and pave the way for efficient and reliable Physical AI applications over future 6G networks. Project website: https://sites.google.com/view/goc-physical-ai-testbed.
Chinese Translation
Physical AI依赖于频繁更新、对延迟敏感的视频流来感知、推理并与物理世界交互,这导致严格的延迟要求以及现有5G网络无法支持的更高数据量。目标导向通信(GoC)通过仅传输任务相关的语义表示,为解决这一挑战提供了一种有前景的方法。然而,现有的GoC框架主要在仿真中进行评估,而其有效性从未在实际部署的Physical AI应用中得到验证。在这项工作中,我们开发了一个面向Physical AI的端到端GoC测试床,它通过5G OpenAirInterface网络将配备RGB-D相机和5G调制解调器的PiPER机械臂连接到NVIDIA Jetson AGX Orin边缘服务器。我们提出并实现了三种GoC框架,分别传输3D边界框、2D场景图和3D场景图作为三种类型的语义表示。它们共享为闭环Physical AI应用设计的通用功能模块,包括语义提取、全栈5G传输、语言模型推理、数字孪生验证和机器人控制。在我们的测试床上进行的大量实验表明,与定期传输原始图像数据的传统框架相比,我们的GoC框架将任务完成时间减少了高达52.6%,并将任务成功概率提高了高达45%。这些结果验证了我们GoC框架的实际有效性,并为未来6G网络上的高效可靠Physical AI应用铺平了道路。项目网站:https://sites.google.com/view/goc-physical-ai-testbed.
cs.RO / 111 / 2609.15910

SlipSense: Multimodal Tactile Learning for Low-Latency and Generalized Slip Detection

SlipSense:面向低延迟与泛化滑动检测的多模态触觉学习
Jian, Tong, Kumar, Aditya Thurvas Senthil, Li, Xinyi, Chen, Ziling, Dai, Tianyu, Sengul, Ali, Grimaldi, Matteo, Lu, Wenjie, Nabi, Saleh, Yu, Tao
Abstract
Slip detection is fundamental to dexterous manipulation, yet existing systems often lack precise characterization of detection latency and cross-platform generalization. We present SlipSense, a multimodal tactile slip-detection framework built on TacV5, a compact sensor integrating a $32 \times 32$ piezoresistive array operating at 240 Hz and a 3-axis MEMS accelerometer operating at 8 kHz. The piezoresistive array captures spatial pressure distributions, while the accelerometer captures friction-induced vibrations, providing complementary slip cues. The framework performs modality-specific encoding, intra-sensor fusion, and cross-modal attention with causal temporal prediction at 240 Hz. Experiments on a dataset of 1.4 million frames spanning 37 objects demonstrate the complementarity of the two modalities. SlipSense achieves 96.7% Macro F1 with a false-positive rate below 1.6%, detecting 76% of slip events within 23.1 ms. When trained solely on UMI data, SlipSense generalizes zero-shot to a Tesollo dexterous hand, transferring across unseen objects, distinct sensor units, and robotic platforms without retraining.
Chinese Translation
滑动检测是灵巧操作的基础,但现有系统往往缺乏对检测延迟和跨平台泛化能力的精确表征。我们提出SlipSense,一个基于TacV5构建的多模态触觉滑动检测框架。TacV5是一种紧凑型传感器,集成了以240 Hz工作的32×32压阻阵列和以8 kHz工作的三轴MEMS加速度计。压阻阵列捕捉空间压力分布,而加速度计捕捉摩擦引起的振动,提供互补的滑动线索。该框架执行模态特定编码、传感器内融合和跨模态注意力,并具有240 Hz的因果时间预测。在包含37个物体的140万帧数据集上的实验证明了两种模态的互补性。SlipSense实现了96.7%的Macro F1分数,假阳性率低于1.6%,在23.1毫秒内检测到76%的滑动事件。当仅在UMI数据上训练时,SlipSense能够零样本泛化到Tesollo灵巧手,无需重新训练即可跨未见物体、不同传感器单元和机器人平台迁移。
cs.RO / 112 / 2609.15921

Touch2Trace: Tactile-Driven Imitation Learning for Dexterous Cable Tracing

Touch2Trace:触觉驱动的灵巧线缆跟踪模仿学习
Grimaldi, Matteo, Klee, David, Chen, Ziling, Jian, Tong, Lee, Wonju, Lu, Wenjie, Yu, Tao, Nabi, Saleh
Abstract
Dexterous manipulation of deformable objects demands continuous fingertip-level regulation of pressure, friction, and incipient slip. We study one of the most challenging cases: dexterous cable tracing, feeding a cable through the hand with repeated pinch-and-curl motions of the thumb and index finger. We introduce Touch2Trace, a tactile-driven imitation-learning system for this task, and provide, to our knowledge, the first systematic real-world characterization of how encoder pretraining, control rate, temporal context, and spatial resolution each shape policy performance. The winning learning recipe combines a tactile encoder pretrained for a custom 32 x 32 piezoresistive sensor (TacV5) via self-supervised learning with a lightweight transformer policy trained on teleoperated demonstrations via behavior cloning, deployed at 60 Hz on a Tesollo DG-5F hand. Tactile feedback without vision or explicit cable-state estimation significantly improves tracing performance versus a proprioception-only baseline: from 0.2 cm to 20.1 cm mean distance and 0% to 93% success rate, with zero-shot transfer to unseen cables and routing conditions. The results quantify the influence of key parameters in tactile-driven systems for reliable dexterous deformable object manipulation.
Chinese Translation
可变形物体的灵巧操作要求对压力、摩擦和初始滑移进行连续的指尖级调控。我们研究最具挑战性的场景之一:灵巧线缆跟踪,即通过拇指和食指重复的捏合和卷曲动作将线缆穿过手部。我们提出了Touch2Trace,一种针对该任务的触觉驱动模仿学习系统,并据我们所知,首次在真实世界中系统地表征了编码器预训练、控制频率、时间上下文和空间分辨率各自如何影响策略性能。最优学习方案结合了通过自监督学习为定制的32x32压阻传感器(TacV5)预训练的触觉编码器,以及通过行为克隆在遥操作演示上训练的轻量级transformer策略,部署在Tesollo DG-5F手上,以60 Hz运行。在没有视觉或显式线缆状态估计的情况下,触觉反馈显著提高了跟踪性能,与仅依靠本体感觉的基线相比:平均距离从0.2厘米增加到20.1厘米,成功率从0%提高到93%,并且能够零样本迁移到未见过的线缆和布线条件。这些结果量化了触觉驱动系统中关键参数对于可靠的灵巧可变形物体操作的影响。
cs.RO / 113 / 2609.15940

Beyond Single-Axis Testing: Paired Evaluation of Compound Robustness in Vision-Language-Action Policies

超越单轴测试:视觉-语言-动作策略中复合鲁棒性的配对评估
Sawada, Hiroki, Kasahara, Shunichi
Abstract
Vision-language-action policies are typically evaluated one perturbation at a time, providing a useful diagnosis of their sensitivity to individual distribution shifts. Real-world deployment, however, may involve several shifts simultaneously, and it remains unclear how these individual robustness measurements compose. We ask whether compound robustness can be inferred from single-axis evaluations. We introduce LIBERO-CTRL, a six-axis benchmark that pairs each initial state across single-axis conditions and a matched simultaneous condition. This design reveals two opposing outcome changes that aggregate success rates cannot distinguish: emergent failures, where all single-axis rollouts succeed but the simultaneous rollout fails, and compensated successes, where at least one single-axis rollout fails but the simultaneous rollout succeeds. Because one transition decreases compound success while the other increases it, they can cancel, making aggregate compound performance appear consistent with single-axis measurements even when individual outcomes differ substantially. These opposing transitions can largely cancel in aggregate: even when the difference between the two transition rates is not statistically distinguishable from zero, as many as 29.0% of matched initial states still change outcome. Across six policies and three severity levels, such outcome changes reach 34.5% in the most affected condition. The relative prevalence of the two transitions varies across policies and severities, while the transition rates remain similar under independent re-evaluation of stochastic policies. Compound robustness therefore cannot be characterized from aggregate single-axis success rates alone; matched per-instance evaluation is needed to reveal how joint perturbations alter behavior.
Chinese Translation
视觉-语言-动作策略通常一次评估一个扰动,为它们对单个分布偏移的敏感性提供了有用的诊断。然而,现实世界的部署可能同时涉及多个偏移,并且这些单独的鲁棒性度量如何组合仍不清楚。我们探究复合鲁棒性是否可以从单轴评估中推断出来。我们引入 LIBERO-CTRL,一个六轴基准,它将每个初始状态在单轴条件下与匹配的同时条件进行配对。这种设计揭示了两种相反的结果变化,而聚合成功率无法区分它们:涌现失败,即所有单轴 rollout 都成功但同时 rollout 失败;补偿成功,即至少一个单轴 rollout 失败但同时 rollout 成功。因为一种转变降低复合成功,而另一种增加它,它们可以相互抵消,使得聚合的复合性能看起来与单轴测量一致,即使个体结果差异很大。这些相反的转变在总体上可以大部分抵消:即使两种转变率之间的差异在统计上与零无法区分,多达 29.0% 的匹配初始状态仍然改变了结果。在六个策略和三个严重程度等级上,这种结果变化在最受影响的情况下达到 34.5%。两种转变的相对普遍性因策略和严重程度而异,而在随机策略的独立重新评估下,转变率保持相似。因此,复合鲁棒性不能仅从聚合的单轴成功率来表征;需要匹配的逐实例评估来揭示联合扰动如何改变行为。
cs.RO / 114 / 2609.15976

MessyMem: Learning-from-Doing Memory for Mobile Manipulation

MessyMem:面向移动操作的从做中学记忆
Banwasi, Anuva, Muckelroy III, William, Sundaresan, Priya, Zhao, Linfeng, Bohg, Jeannette, Ho, Cherie
Abstract
Mobile manipulators deployed across many rooms and visits should improve with experience: after discovering that a cabinet is locked or finding an object in a drawer, the robot should reuse that knowledge rather than start each task from scratch. Yet today's robots often treat each task as new: compact scene representations omit interaction-derived knowledge, raw video histories are difficult to query, and VLM planners reason at inference time without persistently updating what the robot knows. We present MessyMem, a persistent memory system that enables mobile manipulators to learn from experience and reuse that knowledge across future tasks. It maintains a spatially grounded 3D scene graph of objects and locations, augments it with properties and outcomes learned through interaction, and links visual observations for fine-grained recall. We evaluate MessyMem in simulation and on a real mobile manipulator. In a continuous 25-task simulation spanning over 3 hours, MessyMem achieves 80.0% task progress, outperforming the strongest ablation by 14.8 percentage points and the strongest external baseline by 28.9 points, while retrieving task-relevant evidence from thousands of stored keyframes and over an hour into the past.
Chinese Translation
部署在多个房间并多次访问的移动操作机器人应能随经验提升:在发现柜子锁着或在抽屉中找到物体后,机器人应复用该知识,而非每个任务都从头开始。然而,当今的机器人常将每个任务视为新任务:紧凑的场景表示忽略了交互衍生的知识,原始视频历史难以查询,VLM规划器在推理时进行推理,却不持久更新机器人已知的内容。我们提出MessyMem,一种持久记忆系统,使移动操作机器人能够从经验中学习,并在未来任务中复用这些知识。它维护一个空间锚定的物体与位置3D场景图,用通过交互学到的属性和结果对其进行增强,并链接视觉观察以实现细粒度回忆。我们在仿真和真实移动操作机器人上评估MessyMem。在持续超过3小时的连续25任务仿真中,MessyMem达到80.0%任务进度,比最强消融高14.8个百分点,比最强外部基线高28.9个百分点,同时能从数千个存储关键帧以及回溯超过一小时的过去中检索任务相关证据。
cs.RO / 115 / 2609.15988

ResSafe: Learning Safety Filtering with Residual Reinforcement Learning for Humanoids

ResSafe:学习使用残差强化学习为人形机器人进行安全过滤
Qu, Gechen, Zhang, Tong, Zhang, Bike, Wang, Yen-Jen, Sreenath, Koushil, Tomlin, Claire, Choi, Jason Jangho
Abstract
Safe control of humanoid robots remains challenging due to their high-dimensional dynamics, contact-rich interactions, and sensitivity to disturbances. Although reinforcement learning has enabled effective locomotion and motion tracking, learned policies can still generate unsafe actions that lead to instability or falls. In this work, we propose residual reinforcement learning as an implicit safety-filtering mechanism for safe humanoid control. Instead of relying on a single nominal policy to simultaneously balance performance, safety, and robustness, we decouple performance and safety. The nominal policy focuses solely on task performance, while a residual policy learns safety corrections. This decoupling leads to a better performance--safety Pareto trade-off and avoids the need for careful tuning of multiple competing reward terms within a single policy training. We show that the residual policy can act as an implicit safety filter.
Chinese Translation
人形机器人的安全控制仍然具有挑战性,因为其高维动力学、富接触交互以及对扰动的敏感性。尽管强化学习已实现了有效的运动与动作跟踪,但学习到的策略仍可能生成导致不稳定或摔倒的不安全动作。在这项工作中,我们提出将残差强化学习作为一种隐式安全过滤机制,用于安全的人形机器人控制。我们不再依赖单一的名义策略同时平衡性能、安全性和鲁棒性,而是将性能与安全解耦。名义策略仅关注任务性能,而残差策略学习安全修正。这种解耦带来了更好的性能-安全帕累托权衡,并避免了在单个策略训练中仔细调整多个相互竞争的奖励项的需要。我们表明,残差策略可以充当隐式安全过滤器。