← Back to Index
Daily Research Digest

arXiv Papers

2026-09-17
95
Papers
1
Categories
95
Translated
收藏清单 0
机器人学 (Robotics)
95
cs.RO / 1 / 2609.17540

Human-Centric Grasp State Assessment: Toward Transferring Subjective Evaluation to Robots

以人为中心的抓取状态评估:将主观评价迁移至机器人
Kobayashi, Ryohei, Isomoto, Kosei, Yano, Yuga, Tanaka, Yuichiro, Tamukoh, Hakaru
Abstract
We propose a framework that transfers tacit human subjective criteria to robotic systems for the appropriate grasping of deformable objects. Achieving such behavior is challenging because a semantic gap exists between qualitative human expectations and quantitative robotic measurements. Conventional deep learning approaches for bridging this gap also require prohibitive amounts of manually annotated data for each newly encountered object. To address these challenges, our framework integrates a Vision-Language Model (VLM)-based semi-automated supervisor generator with a lightweight grasp state predictor, using a minimal set of human-annotated trials as contextual anchors to propagate subjective criteria to unannotated data. The prediction model then enables rapid online adaptation by sequentially estimating the grasp state from time-series tactile and grasping force measurements. Through experiments on three representative deformable objects and a human evaluation study with 25 participants, we demonstrate the feasibility of the proposed framework for adjusting grasping force according to human-perceived grasp appropriateness in the evaluated task setting.
Chinese Translation
我们提出了一种框架,将人类隐性的主观标准迁移到机器人系统中,以实现对可变形物体的恰当抓取。实现这种行为具有挑战性,因为人类的定性期望与机器人的定量测量之间存在语义鸿沟。用于弥合这一鸿沟的传统深度学习方法还需要为每个新遇到的物体提供数量庞大的人工标注数据。为应对这些挑战,我们的框架将基于视觉-语言模型(Vision-Language Model, VLM)的半自动化监督标签生成器与轻量级抓取状态预测器相结合,利用少量人工标注的试验作为上下文锚点,将主观标准传播到未标注的数据中。随后,预测模型通过从时序触觉和抓取力测量中依次估计抓取状态,实现快速的在线自适应。通过在三种代表性可变形物体上的实验,以及一项有25名参与者参与的人类评价研究,我们证明了所提框架在所评估的任务场景中能够根据人类感知的抓取恰当性来调整抓取力的可行性。
cs.RO / 2 / 2609.17541

Real-Time Service Robot Replanning via Simple Button Interaction for Improved Task Success and User Experience

基于简单按钮交互的实时服务机器人重规划:提升任务成功率与用户体验
Terashima, Ryo, Yano, Yuga, Arimura, Koshun, Tamukoh, Hakaru
Abstract
Service robots must respond to unexpected instructions in real-world environments. However, robots cannot detect all failures and exceptions during a task. To address these issues, we propose a real-time feedback function that enables robots to modify their behavior based on human feedback. In this system, users can intuitively send feedback to the robot by pressing a single button on a tablet when the robot fails to act correctly. Robots use this feedback to consider their failures and replan appropriate actions to complete the task. We conducted experiments with and without the feedback function to verify the following hypothesis: "Simple interactions do not cause a negative user experience." All questionnaire responses are evaluated on a five-point Likert scale. After adding the feedback function, the response score for the question "Did you feel that the robot's behavior was unexpected?" improved by 0.5 points, and that for "Did you feel anxious about the robot's behavior at times?" improved by 0.9 points. These results support the study's hypothesis and indicate that incorporating this real-time feedback function can simultaneously improve task success and the user experience.
Chinese Translation
服务机器人在现实环境中必须能够响应意外指令。然而,机器人无法检测到任务执行过程中出现的所有失败和异常。为解决这些问题,我们提出了一种实时反馈功能,使机器人能够根据人类反馈调整其行为。在该系统中,当机器人未能正确执行动作时,用户可以通过按下平板电脑上的单个按钮,直观地向机器人发送反馈。机器人利用该反馈分析自身的失败原因,并重新规划合适的动作以完成任务。我们通过有反馈功能和无反馈功能的对照实验验证了以下假设:"简单交互不会导致负面的用户体验"。所有问卷回答均采用五级李克特量表进行评估。加入反馈功能后,问题"你是否觉得机器人的行为出乎意料?"的回答得分提高了0.5分,问题"你是否有时对机器人的行为感到焦虑?"的回答得分提高了0.9分。这些结果支持了本研究的假设,并表明引入这种实时反馈功能能够同时提升任务成功率和用户体验。
cs.RO / 3 / 2609.17625

Predictive Varanus: Combining CSP Conformance Monitoring with Predictive LTL Runtime Verification

预测式 Varanus:结合 CSP 一致性监测与预测式 LTL 运行时验证
Ferrando, Angelo, Luckcuck, Matt, Ribeiro, Pedro
Abstract
Runtime Verification is well suited to autonomous and robotic systems because it checks the behaviour that is actually observed during execution. Its main limitation, however, is that it is usually reactive: the monitor detects a violation only after the system has already performed a bad event. This can be too late in domains where failures are costly or unsafe. In this paper we present PREDICTIVE VARANUS, a two-stage verification pipeline that combines VARANUS, a runtime verifier that uses models written in the process algebra Communicating Sequential Processes (CSP), with predictive runtime verification for LTL. A CSP model is first used as a conformance gate over the observed event trace; the same model is then translated into a Buchi automaton that constrains the futures explored by a predictive LTL monitor. In this way, out-of-model behaviour is rejected immediately, while model-consistent prefixes can be classified as already guaranteeing satisfaction, already forcing violation, or still being inconclusive for the monitored temporal property. We formalise the combined monitor, explain its implementation, and illustrate the approach on a robotic rover for nuclear-store inspection. The case study shows how the combination of CSP validation and predictive LTL can provide earlier verdicts than standard runtime monitoring while reusing an existing design-time CSP model.
Chinese Translation
运行时验证非常适合自主系统和机器人系统,因为它检查的是执行过程中实际观察到的行为。然而,其主要局限在于通常是反应式的:监视器只有在系统已经执行了不良事件之后才能检测到违规。在故障代价高昂或不安全的领域,这可能为时已晚。本文提出了 PREDICTIVE VARANUS,一个两阶段验证流水线,它将 VARANUS(一个使用进程代数通信顺序进程(CSP)编写的模型进行验证的运行时验证器)与针对 LTL 的预测式运行时验证相结合。CSP 模型首先作为对所观察事件轨迹的一致性门控;然后将同一模型转换为 Büchi 自动机,以约束预测式 LTL 监视器所探索的未来路径。这样,超出模型的行为会被立即拒绝,而与模型一致的前缀则可以被分类为已经保证满足、已经必然导致违规,或对于所监测的时序性质仍无定论。我们对组合监视器进行了形式化描述,说明了其实现,并在一个用于核存储设施巡检的月球车机器人上展示了该方法。案例研究表明,CSP 验证与预测式 LTL 的结合可以在复用现有设计时 CSP 模型的同时,比标准运行时监测更早地给出判定结论。
cs.RO / 4 / 2609.17627

Flexible-body Modeling, Kinematic Identification, and Assembly Accuracy of Overconstrained Spatial Linkages

过约束空间连杆机构的柔性体建模、运动学参数辨识与装配精度研究
Huczala, Daniel, Pieber, Michael, Gerstmayr, Johannes, Mair, Andreas, Schulte, Frederik, Glas, Silvia, Postulka, Tomas, Vysocky, Ales, Pfurner, Martin
Abstract
Overconstrained rational single-loop linkages are efficient, compact, and low-cost custom mechanisms, yet their deployment in industrial settings is limited. In simulations, rigid body formulations fail due to redundant constraints. This study presents a flexible multibody modeling framework based on the floating frame of reference formulation, and delivers an overall accuracy analysis of assembled linkages prototypes. The approach is validated against 3D-printed PLA prototypes of a Bennett four-bar mechanism, including variants with intentional joint-axis misalignment, which theoretically, from the rigid body point of view, cannot be assembled. A supplementary contribution is delivered in the form of a kinematic parameter identification methodology suited for this type of mechanism with ill-conditioned Jacobian. The experimental and simulation results are compared and reveal that these overconstrained mechanisms exhibit a self-assembling tendency -- structural compliance drives the assembly toward the ideal geometric configuration, distributing constraint stress throughout the structure. Additional qualitative demonstrations using cardboard tubes and bamboo sticks as link building blocks confirm that functional mechanisms can be realized from low-cost, unconventional materials with limited manufacturing accuracy. The proposed modeling pipeline is fully algorithmic and enables design optimization in the future.
Chinese Translation
过约束单环连杆机构是一种高效、紧凑且低成本的特殊机构,然而其在工业领域的应用仍然有限。在仿真中,刚体建模会因冗余约束而失效。本研究提出了一种基于浮动坐标系(floating frame of reference)方法的柔性多体建模框架,并对装配后的机构原型进行了整体精度分析。该方法通过Bennett四杆机构的3D打印PLA原型进行了验证,其中包括具有故意引入的关节轴错位的变体——从刚体角度看,这些变体在理论上无法装配。本研究的补充贡献是提出了一种适用于此类雅可比矩阵病态机构的运动学参数辨识方法。通过实验与仿真结果的对比发现,这类过约束机构表现出自装配趋势——结构柔性使装配趋向理想的几何构型,并将约束应力分散到整个结构中。此外,使用纸筒和竹签作为连杆构建单元的定性演示进一步证实,利用制造精度有限的低成本、非常规材料也可以实现功能正常的机构。所提出的建模流程完全算法化,为未来的设计优化奠定了基础。
cs.RO / 5 / 2609.17628

LEAP: Learning Emergent Active Perception for Quadruped Navigation

LEAP:面向四足导航的学习型涌现主动感知
Gökbakan, Ü. Bora, Caron, Stéphane, Souères, Philippe
Abstract
Active perception allows autonomous agents to select their viewpoints rather than passively process the viewpoints given to them, enabling them to target where to reduce uncertainty about their environment. Learned systems typically encourage this behavior with hand-designed proxy objectives, such as coverage or curiosity bonuses, that may conflict with the task. In this work, we propose a method to learn emergent active perception (LEAP) without augmentation of the task objective. We formulate the problem of goal-oriented navigation over hazardous terrains with goals that must be discovered visually. We then propose an architecture for navigation policies with active perception, and train them on a terrain curriculum where task pressure alone leads to the emergence of gaze control. Key to this emergence, LEAP works on a gaze-invariant representation that integrates depth images into egocentric belief maps. We validate its performance in held-out evaluation scenarios, where it achieves a 92.7% success rate, compared to 74.2% for scripted or 34.5% for passive perception, and comes within 4.6 points of a privileged oracle. We validate that LEAP navigation policies, unchanged, can be directly applied to steering quadrupedal locomotion policies in physics simulation.
Chinese Translation
主动感知使自主智能体能够自主选择视角,而非被动地处理给定的视角,从而使其能够有针对性地减少对环境的不确定性。基于学习的系统通常使用手工设计的代理目标(如覆盖度或好奇心奖励)来鼓励这种行为,但这些目标可能与任务本身相冲突。在本工作中,我们提出了一种无需对任务目标进行增广即可学习涌现式主动感知的方法(LEAP)。我们将面向目标的危险地形导航问题形式化,其中目标必须通过视觉发现。随后,我们提出了一种具备主动感知能力的导航策略架构,并在地形课程学习中进行训练,仅凭任务压力即可促使注视控制(gaze control)行为的涌现。这种涌现的关键在于,LEAP基于一种注视不变的表征,该表征将深度图像融合到以自我为中心的信念地图中。我们在保留的评估场景中验证了其性能:成功率达到92.7%,而脚本化感知为74.2%,被动感知仅为34.5%,并且与具备特权信息的先知(oracle)仅相差4.6个百分点。我们进一步验证了未经修改的LEAP导航策略可以直接应用于物理仿真中对四足运动策略的转向控制。
cs.RO / 6 / 2609.17714

Vision-Language Grounded Task-Context-Aware Imitation Learning for Robotic Disassembly

面向机器人拆卸的视觉-语言接地任务-上下文感知模仿学习
Kang, Jeon Ho, Tamarkin, Igal, Niu, Ethan, Novales, Ian, Gupta, Satyandra K.
Abstract
Real-world robotic disassembly requires long-horizon execution, where robots must perform ordered sequences of manipulation tasks across multiple parts within a single scene. Multiple valid task goals and diverse assembly configurations make it difficult for imitation policies to infer the intended skill from raw observations alone, particularly when training data cannot cover the combinatorial diversity of real-world configurations and part geometries. We show that incorporating task context through language alleviates these challenges by providing explicit structure for skill selection and associating language-specified tasks with their corresponding manipulation targets in the visual scene. The proposed framework combines hierarchical task selection with task-context-aware imitation learning to ground language instructions in spatial visual representations for robotic disassembly. The resulting framework generalizes across diverse connector geometries and assembly configurations without requiring explicit object annotations. Our method improves end-to-end task success by 35 percentage points over the baseline diffusion policy and by 75 percentage points over the previous task-context-aware baseline.
Chinese Translation
现实世界中的机器人拆卸任务需要长时程执行,机器人必须在同一场景中对多个部件按顺序执行一系列操作任务。多个有效的任务目标以及多样化的装配结构使得模仿策略难以仅凭原始观测推断出预期的技能,尤其是在训练数据无法覆盖现实装配结构和部件几何形状的组合多样性时。我们证明,通过语言引入任务上下文可以缓解这些挑战,它为技能选择提供了显式的结构,并将语言指定的任务与视觉场景中相应的操作目标关联起来。所提出的框架将层次化任务选择与任务-上下文感知的模仿学习相结合,将语言指令接地到用于机器人拆卸的空间视觉表示中。该框架能够泛化到多样的连接器几何形状和装配结构,而无需显式的物体标注。与基线扩散策略(diffusion policy)相比,我们的方法将端到端任务成功率提高了35个百分点;与先前的任务-上下文感知基线相比,提高了75个百分点。
cs.RO / 7 / 2609.17728

RAF-VLA: Representation Alignment with the Future for End-to-End Autonomous Driving

RAF-VLA:面向端到端自动驾驶的未来表示对齐
Kim, Dogun, Lee, Yongjae, Lim, Joonhee, Lee, Yeina, Park, Junhyeok, Park, Moogeun, Kum, Dongsuk
Abstract
Recent Vision-Language-Action (VLA) models for autonomous driving have incorporated world modeling by predicting future driving scenes alongside driving actions, demonstrating strong planning performance. Future driving scenes are utilized as dense supervision, encouraging the policy to learn rich internal representations useful for planning. However, these World-Modeling VLAs rely on explicit future generation to learn such representations, thereby introducing two key limitations: additional training burden and inference latency. To address these limitations, we propose RAF-VLA (Representation Alignment with the Future), a VLA-based autonomous driving framework that shapes planning-relevant internal representations through direct guidance from future-frame representations. RAF-VLA employs Future-Aligned Supervised Fine-Tuning, in which a straightforward regularization aligns the policy's hidden states with future-frame representations obtained from a pretrained world encoder while learning driving actions. This simple alignment allows RAF-VLA to avoid the training burden and inference latency associated with future generation. Extensive experiments on the NAVSIM benchmark show that RAF-VLA achieves competitive planning performance against state-of-the-art VLA planners with substantially fewer training samples seen. Moreover, RAF-VLA incurs only 3.8% training overhead and a negligible 1 ms inference overhead.
Chinese Translation
近年来,面向自动驾驶的视觉-语言-动作(Vision-Language-Action, VLA)模型在预测驾驶动作的同时引入了世界建模,通过预测未来驾驶场景展现了强大的规划性能。未来驾驶场景被用作稠密监督信号,促使策略学习到对规划有用的丰富内部表示。然而,这些基于世界建模的VLA模型依赖显式的未来生成来学习此类表示,从而带来两个关键局限:额外的训练负担和推理延迟。为解决这些局限,我们提出了RAF-VLA(Representation Alignment with the Future,与未来对齐的表示),这是一个基于VLA的自动驾驶框架,通过未来帧表示的直接引导来塑造与规划相关的内部表示。RAF-VLA采用未来对齐的监督微调(Future-Aligned Supervised Fine-Tuning),其中一种简单的正则化方法在学习驾驶动作的同时,将策略的隐藏状态与从预训练世界编码器获得的未来帧表示对齐。这一简单的对齐方式使RAF-VLA无需承担与未来生成相关的训练负担和推理延迟。在NAVSIM基准上的大量实验表明,RAF-VLA在训练样本量显著更少的情况下,取得了与最先进VLA规划器相当的规划性能。此外,RAF-VLA仅引入3.8%的训练开销和可忽略的1毫秒推理开销。
cs.RO / 8 / 2609.17758

CALOS: Control-Affine Lyapunov On-manifold Safety Layer for Safe Deep Reinforcement Learning for Quadrotors

CALOS:面向四旋翼安全深度强化学习的控制仿射李雅普诺夫流形上安全层
Cesareo, Fabrizio, Mengozzi, Sebastiano, Mimmo, Nicola, Acquaviva, Andrea
Abstract
Deep Reinforcement Learning has demonstrated remarkable capability in quadrotor control, yet learned policies offer no guarantee of respecting safety constraints during training or deployment. We present CALOS (Control-Affine Lyapunov On-manifold Safety), a runtime safety layer that enforces attitude constraints on a quadrotor without modifying the underlying learning algorithm. CALOS formulates four tilt-angle inequalities and a Lyapunov descent condition as a single quadratic program whose solution is the minimum-norm correction to the nominal torque output of the policy. The quadratic program is solved exactly via active-set enumeration over the three-dimensional torque space, with a computational cost low enough to enforce constraints in real time across thousands of parallel simulation environments, as required by modern massively parallel Deep Reinforcement Learning training. Evaluated on trajectory-tracking tasks in NVIDIA Isaac Lab, CALOS reduces lateral tracking error by 55-60% relative to an unconstrained Proximal Policy Optimization baseline while achieving zero attitude-constraint violations on the training trajectory. By restricting exploration to safe regions of the state space, the safety layer also accelerates training convergence and improves data efficiency without producing suboptimal policies.
Chinese Translation
深度强化学习在四旋翼控制中已展现出卓越能力,但学习得到的策略无法保证在训练或部署过程中遵守安全约束。我们提出了CALOS(控制仿射李雅普诺夫流形上安全,Control-Affine Lyapunov On-manifold Safety),这是一种运行时安全层,能够在不修改底层学习算法的前提下对四旋翼施加姿态约束。CALOS将四个倾斜角不等式和一个李雅普诺夫下降条件表述为一个二次规划问题,其解是对策略标称力矩输出的最小范数修正。该二次规划通过在三维力矩空间上进行有效集枚举精确求解,其计算成本足够低,可在现代大规模并行深度强化学习训练所需的数千个并行仿真环境中实时执行约束。在NVIDIA Isaac Lab中的轨迹跟踪任务上的评估表明,相对于无约束的近端策略优化(Proximal Policy Optimization)基线,CALOS将横向跟踪误差降低了55-60%,同时在训练轨迹上实现了零姿态约束违反。通过将探索限制在状态空间的安全区域内,该安全层还加速了训练收敛并提高了数据效率,且不会产生次优策略。
cs.RO / 9 / 2609.17766

DRT&R: Direct Radar Teach & Repeat

DRT&R:直接雷达示教与复现导航
Krawciw, Alexander, Lisus, Daniil, Gentil, Cedric Le, Barfoot, Timothy D.
Abstract
Radar-based navigation is appealing for its robustness to adverse conditions involving airborne particles, such as precipitation, dust, fog, and smoke, that can cause lidar-based systems to fail. Recently, direct methods that retain and use the entire radar scan rather than sparse points have improved on-road global localization performance. However, they have yet to be deployed in off-road environments or in closed-loop systems. Additionally, even direct global maps may lose information: their global nature leads to a smoothing out of viewpoint-dependent radar artifacts, which can provide pose information when mapping and localization occur along similar trajectories. This paper introduces Direct Radar Teach & Repeat (DRT&R): a direct spinning radar-based navigation stack that maximizes the amount of retained information by combining direct radar processing with local mapping. DRT&R yields state-of-the-art (SOTA) localization performance in both on-road and off-road environments. Using 344 km of on-road data and 20 km of off-road data, DRT&R is able to localize to within 4 cm in most on-road and off-road conditions, and 12 cm in geometrically degenerate and sparse environments. DRT&R is also evaluated autonomously in closed loop with an MPC controller for more than 10 km using a Clearpath Warthog off-road vehicle, demonstrating that it runs in real time and achieves SOTA tracking performance for off-road radar navigation.
Chinese Translation
基于雷达的导航因其在存在降水、灰尘、雾、烟等空气悬浮颗粒物的恶劣条件下依然稳健而备受关注,而这类条件往往会导致基于激光雷达的系统失效。近年来,保留并利用完整雷达扫描数据而非稀疏点云的直接法显著提升了道路场景下的全局定位性能。然而,这类方法尚未被部署到越野环境或闭环系统中。此外,即使是直接法构建的全局地图也可能丢失信息:其全局性会导致依赖视角的雷达伪影被平滑掉,而这些伪影在建图与定位沿相似轨迹进行时能够提供位姿信息。本文提出了直接雷达示教与复现导航方法(Direct Radar Teach & Repeat, DRT&R):一种基于旋转雷达的直接导航系统,通过将直接雷达处理与局部建图相结合,最大化保留信息的数量。DRT&R 在道路和越野环境中均取得了最先进(SOTA)的定位性能。基于 344 公里的道路数据和 20 公里的越野数据,DRT&R 在大多数道路和越野条件下可将定位误差控制在 4 厘米以内,在几何退化且稀疏的环境中误差为 12 厘米。DRT&R 还结合 MPC 控制器在 Clearpath Warthog 越野车上进行了超过 10 公里的自主闭环评估,结果表明其能够实时运行,并在越野雷达导航中实现了最先进的跟踪性能。
cs.RO / 10 / 2609.17771

HINT-Plan: Human Intention-Aware Robot Task Planning in Context-Rich Environments using Vision Language Models

HINT-Plan:基于视觉语言模型的情境丰富环境中人类意图感知机器人任务规划
Liu, Yuchen, Palmieri, Luigi, Li, Lujun, State, Radu, Georgievski, Ilche, Aiello, Marco
Abstract
Approaches to incorporating human awareness into mobile robot decision-making mainly focus on collision avoidance in low-level motion planning, often overlooking the challenges posed by human presence and high-level behavior. To address this vacancy, we present HINT-Plan, a novel approach to integrate human intention prediction into robot task planning. HINT-Plan employs Vision Language Models (VLMs) to anticipate high-level human intentions from third-person image observations, convert them into goal states, and solve joint task-planning problems. To effectively enable scene awareness in context-rich environments, we use hierarchical Scene Graphs (SGs) as high-level representations of the environment, and translate environmental topology and actionable knowledge into formal planning language to ensure executable plans. Evaluated in a photorealistic simulation, HINT-Plan achieves an overall success rate of 69.71% in joint human-robot task planning, substantially outperforming the baselines by up to 35.29%, while also reducing functional conflicts. The results show the effectiveness of explicitly incorporating inferred human intentions into formal multi-agent task planning for proactive human-aware robot decision-making.
Chinese Translation
将人类感知融入移动机器人决策的方法主要集中于底层运动规划中的避碰,往往忽视了人类存在及高层行为所带来的挑战。为填补这一空白,我们提出了HINT-Plan,一种将人类意图预测融入机器人任务规划的新方法。HINT-Plan利用视觉语言模型(Vision Language Models, VLMs)从第三人称图像观测中预测高层人类意图,将其转化为目标状态,并求解联合任务规划问题。为了在情境丰富的环境中有效实现场景感知,我们采用层次化场景图(Scene Graphs, SGs)作为环境的高层表示,并将环境拓扑与可操作知识转化为形式化规划语言,以确保规划的可执行性。在逼真仿真环境中的评估表明,HINT-Plan在人机联合任务规划中实现了69.71%的总体成功率,相较基线方法最高提升35.29%,同时减少了功能性冲突。结果表明,将推断出的人类意图显式地融入形式化多智能体任务规划,对于实现主动的人类感知机器人决策是有效的。
cs.RO / 11 / 2609.17781

SPROUT: The Open-Source Soft Growing Robot for Search and Rescue

SPROUT:用于搜救的开源软体生长机器人
Valdivia, Antonio Alvarez, McFarland, Ciera, Reeve, Robert, Dhawan, Ankush, Council, Chad, Richardson, Megan, McGuinness, Margaret, Hanson, Nathaniel
Abstract
Soft robotic systems have long been theorized as ideal candidates for use in search and rescue operations; however, there have been significant barriers to entry in graduating soft robotic systems from the laboratory to the field. To address this gap in replicable, reliable soft robot systems, we present the designs for SPROUT, the Soft Pathfinding Robotic Observation Unit. The system has matured over years of interaction with professional urban search and rescue communities, with the goal of operating in dusty, wet, and isolated conditions. This manuscript contains supplementary material, including code, parts manifests, CAD assemblies, and build instructions to allow the broader robotics research community to build their own SPROUT systems. The modular hardware and ROS 2-based software stack support task-specific payloads and control functions, allowing SPROUT to be adapted for applications beyond search and rescue, including infrastructure inspection and archaeology. Here, we demonstrate SPROUT performing a variety of challenging inspection and traversal tasks in collapsed structure training sites used by first responders. The main project page is available online at https://sprout-mitll.github.io/sprout/.
Chinese Translation
软体机器人系统长期以来一直被认为是在搜救行动中应用的理想候选者;然而,将软体机器人系统从实验室推向实际应用领域仍存在重大障碍。为解决可复制、可靠的软体机器人系统方面的这一空白,我们提出了SPROUT(软体路径规划机器人观测单元,Soft Pathfinding Robotic Observation Unit)的设计方案。该系统经过与专业城市搜救社区多年的互动交流而不断成熟,目标是能够在多尘、潮湿和隔离的环境中运行。本文稿包含补充材料,包括代码、零件清单、CAD装配体和构建说明,使更广泛的机器人学研究社区能够构建自己的SPROUT系统。模块化硬件和基于ROS 2的软件栈支持特定任务的有效载荷和控制功能,使SPROUT能够应用于搜救之外的领域,包括基础设施检测和考古。本文展示了SPROUT在救援人员使用的倒塌结构训练场地中执行各种具有挑战性的检测和通行任务。项目主页可在线访问:https://sprout-mitll.github.io/sprout/。
cs.RO / 12 / 2609.17824

Learning Multi-Humanoid Pickup and Transport via Decentralized Object-Centric Control

基于去中心化物体中心控制的多仿人机器人拾取与搬运学习
Pandit, Bikram, Gadde, Mohitvishnu S., Shrestha, Aayam Kumar, Fern, Alan
Abstract
We study cooperative multi-humanoid pickup and transport of objects with varying size, weight, and geometry, requiring robot teams of different sizes. Our approach uses decentralized object-centric control, where each humanoid is assigned a local attachment region on the shared object and learns to realize pickup and transport through gripperless bimanual pinching. This attachment-based interface provides a common control abstraction spanning single-robot pickup, cooperative multi-robot transport, and robot-to-robot handover, without per-task redesign. We find that policies trained only on single-robot pickup already transfer nontrivially to cooperative settings, suggesting that this abstraction captures much of the structure needed for coordination. At the same time, explicit multi-robot training further improves performance, showing that shared-object coupling introduces coordination dynamics that are beneficial to learn directly. We validate the approach in simulation across varying team sizes and object geometries, and demonstrate sim-to-real transfer on hardware, where the learned controllers enable real humanoids to perform cooperative manipulation tasks.
Chinese Translation
我们研究了多仿人机器人协作拾取和搬运尺寸、重量和几何形状各异的物体,这需要不同规模的机器人团队。我们的方法采用去中心化的物体中心控制,为每个仿人机器人在共享物体上分配一个局部附着区域,并学习通过无夹爪的双臂捏取来实现拾取和搬运。这种基于附着的接口提供了一种通用的控制抽象,涵盖单机器人拾取、多机器人协作搬运以及机器人间交接,无需针对每个任务进行重新设计。我们发现,仅在单机器人拾取上训练的策略已经能够非平凡地迁移到协作场景中,这表明该抽象捕捉了协调所需的大部分结构。同时,显式的多机器人训练进一步提升了性能,表明共享物体的耦合引入了值得直接学习的协调动力学。我们在仿真中针对不同的团队规模和物体几何形状对该方法进行了验证,并在硬件上展示了从仿真到现实的迁移,所学习的控制器使真实的仿人机器人能够执行协作操作任务。
cs.RO / 13 / 2609.17832

Adaptive-MHE : A Sampling-Based Adaptive MPC for Legged Loco-Manipulation via Moving Horizon Estimation

Adaptive-MHE:一种基于采样与移动时域估计的腿式机器人移动操作自适应MPC方法
Keshavarz, Hossein, Ramirez-Serrano, Alejandro, Khadiv, Majid
Abstract
Legged robots have demonstrated a remarkable ability to traverse various terrains, yet generating effective loco-manipulation behaviors remains challenging. A key difficulty is that object and terrain parameters are typically unknown to the robot, and mismatches between these parameters and their simulated counterparts introduce a sim-to-real gap that degrades control performance. Classical system identification (Sys-ID) methods often assume differentiable dynamics, an assumption that does not hold for contact-rich legged systems. Sampling-based Sys-ID avoids this restriction by directly matching simulated and recorded state trajectories through massively parallel rollouts, but existing approaches are typically applied offline and do not adapt as environmental conditions change. We present Adaptive-MHE an online sampling-based Sys-ID framework, based on moving horizon estimation (MHE), that estimates the physical parameters of objects and terrain in the environment (e.g., mass, friction) and couples this estimate with a sampling-based model predictive controller, enabling adaptive loco-manipulation in changing and uncertain environments. In simulation and hardware experiments, our framework consistently outperforms baselines and matches the performance of a controller with access to ground-truth parameters.
Chinese Translation
腿式机器人已展现出穿越多种地形的卓越能力,然而生成有效的移动操作(loco-manipulation)行为仍然具有挑战性。一个关键难点在于,物体和地形参数对机器人而言通常是未知的,而这些参数与仿真中对应参数之间的失配会造成仿真到现实(sim-to-real)的差距,从而降低控制性能。经典的系统辨识(Sys-ID)方法通常假设动力学可微,这一假设对于接触丰富的腿式系统并不成立。基于采样的系统辨识通过大规模并行滚动仿真直接匹配仿真与记录的状态轨迹,从而避免了这一限制,但现有方法通常在离线状态下应用,无法随环境条件的变化进行自适应调整。我们提出了Adaptive-MHE,一个基于移动时域估计(MHE)的在线采样系统辨识框架,能够估计环境中物体和地形的物理参数(如质量、摩擦系数),并将该估计与基于采样的模型预测控制器(MPC)相结合,从而在变化和不确定的环境中实现自适应的移动操作。在仿真和硬件实验中,我们的框架始终优于基线方法,并达到了拥有真实参数的控制器所能实现的性能水平。
cs.RO / 14 / 2609.17834

VCTP: Vehicle-Conditioned Terrain Planning for Off-Road Navigation

VCTP:面向越野导航的车辆条件化地形规划
Naik, Akshay, Sreenivas, Ramavarapu S., Nottage, Dustin, Soylemezoglu, Ahmet
Abstract
A vehicle's heading affects both the surfaces beneath its tires and its pitch and roll. We present Vehicle-Conditioned Terrain Planning (VCTP), which retains these relationships by evaluating shared elevation and surface-ID layers at eight headings. Body geometry constrains admissibility, while loaded wheel contacts determine modeled surface cost and predicted pitch and roll. Established D* Lite and vehicle-state search use these evaluations to plan routes with forward and reverse motion. When observations change, VCTP recomputes every affected body or contact query. In fully observed two-track simulations, sampling at wheel contacts rather than at the vehicle center lowers modeled surface cost by 39.3% while shortening the route. In offline planning on RGator recordings, VCTP also reduces modeled costs over identified surfaces and observed support relative to distance-focused planning, although incomplete coverage leaves full-route rankings unresolved. Selective updates match full recomputation in all 474 comparisons using recorded map changes. These results identify when wheel-contact placement and vehicle heading affect route choice.
Chinese Translation
车辆的朝向不仅影响其轮胎下方的地表,也影响其俯仰角和侧倾角。我们提出了车辆条件化地形规划(Vehicle-Conditioned Terrain Planning,VCTP),通过在八个朝向上评估共享的高程与地表ID图层来保留这些关系。车体几何结构约束了可通行性,而承载轮的接触点决定了建模的地表代价以及预测的俯仰角和侧倾角。已有的D* Lite与车辆状态搜索利用这些评估结果来规划包含前进和倒车运动的路径。当观测发生变化时,VCTP会重新计算所有受影响的车体或接触点查询。在全观测的双履迹仿真中,以轮子接触点而非车辆中心进行采样,使建模地表代价降低39.3%,同时缩短了路径。在基于RGator记录数据的离线规划中,相较于侧重距离的规划方法,VCTP在已识别地表和观测支撑方面同样降低了建模代价,尽管覆盖不完整使得完整路径的排序仍无法确定。在全部474次使用记录地图变化的比较中,选择性更新与完全重计算的结果一致。这些结果揭示了轮接触点位置和车辆朝向何时会影响路径选择。
cs.RO / 15 / 2609.17904

Timely Activation of Safety Filters via One-Step Reachability Expansion

基于单步可达性扩展的安全过滤器及时激活
Borquez, Javier
Abstract
Least-restrictive safety filters based on Hamilton-Jacobi reachability provide strong safety guarantees by overriding a nominal controller only when the system reaches the boundary of the set of unsafe states defined as a Backward Reachable Tube (BRT). These guarantees, however, rely on the continuous-time nature of the underlying formulation. In practice, robotic systems apply control at discrete sampling intervals, which creates a mismatch where the system may jump into the unsafe BRT between updates, allowing failures that are theoretically avoidable. This work introduces a principled solution based on a one-step expanded BRT that predicts all states capable of reaching the true BRT within a single timestep. By using this expanded boundary as the activation condition for the safety filter, safety interventions occur early enough to ensure correctness under discrete-time execution. We formulate this expanded set as a modified reachability problem and compute it using standard continuous-time solvers.
Chinese Translation
基于哈密顿-雅可比可达性(Hamilton-Jacobi reachability)的最小限制安全过滤器提供了强大的安全保障:仅当系统到达定义为后向可达管(Backward Reachable Tube, BRT)的不安全状态集合的边界时,才会覆盖标称控制器。然而,这些保障依赖于底层公式的连续时间特性。在实际中,机器人系统以离散采样间隔施加控制,这产生了不匹配问题:系统可能在两次更新之间跳入不安全的BRT,从而导致理论上可避免的失效。本工作提出了一种基于单步扩展BRT的原理性解决方案,该方案预测所有能够在单个时间步内到达真实BRT的状态。通过将此扩展边界作为安全过滤器的激活条件,安全干预能够足够早地发生,从而保证离散时间执行下的正确性。我们将该扩展集合表述为一个改进的可达性问题,并使用标准的连续时间求解器进行计算。
cs.RO / 16 / 2609.17910

Map2Route: Benchmarking Compositional Language-Grounded Route Planning over Semantic Maps

Map2Route:基于语义地图的组合式语言引导路径规划基准测试
Bao, Muyi, Xu, Hang, Tang, Jingfan, Liu, Zihan, Cai, Yuxin, Lv, Chen, Wang, Wenshan, Zhang, Ji
Abstract
We introduce Map2Route, a human-curated benchmark for compositional language-grounded route planning over pre-built semantic maps. Map2Route contains 1,000 episodes across 40 scenes, where instructions use relational, comparative, and nested descriptions to identify route-relevant objects and regions, while specifying ordered must-pass regions, must-avoid requirements, five categories of soft preferences, and spatial and route-stage scopes, which is partially tested by existing works. Alongside Map2Route, we propose Grounding2Route, which combines executable code-as-grounding with verification-guided repair and scope-aware planning.Across seven representative adapted baselines, Grounding2Route substantially outperforms existing methods in all metrics. Despite these gains, a substantial gap to human demonstrations remains, highlighting the difficulty of Map2Route and the considerable headroom for future progress. Additional qualitative results and resources are available on https://anonymous.4open.science/w/Map2Route-F05F/.
Chinese Translation
我们提出了Map2Route,一个由人工精心构建的基准数据集,用于在预构建的语义地图上进行组合式语言引导的路径规划。Map2Route包含40个场景中的1,000个回合,其指令使用关系性、比较性和嵌套式描述来识别与路径相关的物体和区域,同时指定有序的必经区域、必避要求、五类软偏好以及空间和路径阶段范围,而这些内容仅被现有工作部分测试过。与Map2Route一同,我们提出了Grounding2Route,该方法将可执行的代码作为语言落地(code-as-grounding)与验证引导的修复以及范围感知规划相结合。在七个具有代表性的改编基线方法中,Grounding2Route在所有指标上均大幅优于现有方法。尽管取得了这些进展,与人类演示相比仍存在显著差距,这凸显了Map2Route的难度以及未来进步的巨大空间。更多定性结果和资源可在https://anonymous.4open.science/w/Map2Route-F05F/获取。
cs.RO / 17 / 2609.17929

Multi-Session Multimodal Underwater Mapping with Acoustic and Optical Imaging

基于声学与光学成像的多批次多模态水下地图构建
Philip-Ifabiyi, Precious, Franchi, Valerio, Ferreira, Fausto, Gracias, Nuno
Abstract
Accurate seafloor mapping is essential for marine science, archaeology, and environmental monitoring. However, integrating data from different sensors, such as side-scan sonar and optical cameras, collected across separate survey sessions, remains challenging due to positioning drift and sensor offsets. This paper presents a multi-session, multimodal underwater mapping framework based on factor graph optimization. The method jointly optimizes vehicle trajectories, 3D landmark positions, sensor extrinsics, and per-session global alignment transformations. By combining rigid inter-session corrections with local trajectory deformations, it compensates for both inter-session offsets and intra-session distortions from accumulated navigation errors. The proposed methodology was validated on real-world datasets collected along the Catalan coast. Results show measurable improvements in map consistency over both unoptimized and rigid-alignment baselines across all metrics, including Pixel Accuracy and mean Intersection over Union. The method achieves a 3.4% improvement in pixel accuracy over the unoptimized baseline, corresponding to improved semantic labelling across approximately 14700 $\text{m}^2$ of mapped area. Qualitative results further show consistent co-registration between sonar and optical maps, even in the presence of significant trajectory distortions and inter-session misalignments. These findings demonstrate the potential of the proposed framework to generate coherent multimodal seafloor maps from heterogeneous underwater surveys.
Chinese Translation
精确的海底测绘对海洋科学、考古学和环境监测至关重要。然而,整合在不同勘测批次中采集的多传感器数据(如侧扫声呐和光学相机)仍面临挑战,其原因在于定位漂移和传感器偏移。本文提出了一种基于因子图优化的多批次、多模态水下建图框架。该方法联合优化载体轨迹、三维地标位置、传感器外参以及各批次的全局对齐变换。通过将批次间的刚性校正与局部轨迹形变相结合,该方法能够同时补偿批次间偏移以及由导航误差累积引起的批次内畸变。所提出的方法在加泰罗尼亚沿海采集的真实数据集上得到了验证。结果表明,在包括像素精度(Pixel Accuracy)和平均交并比(mean Intersection over Union)在内的所有指标上,地图一致性相较于未优化基线和刚性对齐基线均有可衡量的提升。该方法相较于未优化基线在像素精度上提升了3.4%,对应约14700平方米测绘区域的语义标注得到改善。定性结果进一步表明,即使在存在显著轨迹畸变和批次间错位的情况下,声呐地图与光学地图之间仍保持一致的空间配准。这些发现证明了所提出框架从异构水下勘测数据中生成一致多模态海底地图的潜力。
cs.RO / 18 / 2609.17946

Feedback-Modulated Harmonic Policies for Quadruped Locomotion

面向四足机器人运动的反馈调制谐波策略
Jia, Yixuan, Roche, Steven, How, Jonathan P.
Abstract
Learned quadruped locomotion policies commonly map observations directly to joint-level actions, leaving the periodic structure of locomotion implicit in the policy. We investigate an alternative representation in which each joint trajectory is expressed as a command-conditioned Fourier series and modified online using feedback from the robot state. A context network generates the Fourier coefficients and the weights of a per-step feedback network, whose outputs adjust joint offsets, harmonic gains, frequency, and phase during execution. In simulation, we examine this explicit frequency structure alongside the hidden activations of an MLP policy that directly outputs joint targets. The harmonic waveforms change frequency and shape with commanded speed. Dynamic mode decomposition of selected MLP rollouts reveals dominant activation modes near the foot-height oscillation frequency and its second harmonic, showing periodic structure without an explicit Fourier generator. On a Unitree Go2, the simulation-trained harmonic controller records a provisional onboard-estimated peak speed of 3.67 meter per second and carries added loads up to 5.883 kilogram in separate trials.
Chinese Translation
通过学习获得的四足机器人运动策略通常将观测直接映射到关节级动作,使运动的周期性结构隐含在策略之中。我们研究了一种替代性表示方法,将每个关节轨迹表示为以指令为条件的傅里叶级数,并利用机器人状态的反馈进行在线修正。一个上下文网络生成傅里叶系数以及逐步反馈网络的权重,反馈网络的输出在执行过程中调整关节偏移量、谐波增益、频率和相位。在仿真中,我们将这种显式频率结构与一个直接输出关节目标的MLP策略的隐层激活进行对比分析。谐波波形随指令速度的变化而改变频率和形状。对若干MLP轨迹回放进行动态模态分解后发现,主导激活模态出现在足部高度振荡频率及其二次谐波附近,表明即使没有显式的傅里叶生成器,策略中也存在周期性结构。在Unitree Go2机器人上,经仿真训练的谐波控制器实现的车载估计峰值速度为3.67米/秒,并在单独测试中承载了最高5.883千克的附加负载。
cs.RO / 19 / 2609.18016

Causal-History Test-Time Scaling for Failure Recovery in Autoregressive World-Action Models

面向自回归世界-动作模型失败恢复的因果历史测试时扩展
Li, Lin, Chen, Long, Kwunhang, Wong, Lei, Jiaming, Jin, Song, Du, Shucheng, Zhang, Chuhan, Ma, Songchen, Zhang, Weihao, Xiao, Jun, Kwang-Ting, Cheng
Abstract
World-action models (WAMs) have emerged as a promising paradigm for robot manipulation by jointly modeling future visual dynamics and robot actions. However, existing WAMs are trained predominantly on successful trajectories, making them prone to failure when real-world execution diverges from the learned dynamics. This issue is amplified in autoregressive WAMs, where execution errors become part of the causal history and continue to influence subsequent predictions. To this end, we introduce \method{}, a training-free framework that reformulates failure recovery as \emph{test-time scaling over causal histories}. This formulation decomposes recovery into three coupled decisions: \emph{when} to revise the causal history, \emph{where} to recover a reliable history prefix, and \emph{which} history configuration best supports subsequent execution. Specifically, \method{} realizes these decisions through three stages: 1) \textbf{Progress-Aware Recovery Trigger} detects persistent non-progress and triggers recovery only when the current execution state permits intervention; 2) \textbf{History-Prefix Recovery} identifies the unreliable history suffix, retrieves a historical anchor matching the current physical state, and reconstructs the causal KV state from the retained prefix while conditioning on the latest real observation; and 3) \textbf{Hypothesis Verification} compares the future continuations induced by complete-history, recovered-prefix, and full-reset hypotheses, and commits the best-supported hypothesis. Experiments in both simulated and real-world manipulation settings demonstrate consistent improvements in task success, while ablations confirm the contribution of each recovery stage.
Chinese Translation
世界-动作模型(World-Action Models, WAMs)通过联合建模未来视觉动态与机器人动作,已成为机器人操作领域一种极具前景的范式。然而,现有 WAMs 主要在成功轨迹上进行训练,当真实世界执行过程偏离所学动态时容易失败。该问题在自回归 WAMs 中被进一步放大,因为执行错误会成为因果历史的一部分,并持续影响后续预测。为此,我们提出了 \method{},一个无需训练的框架,将失败恢复重新表述为“基于因果历史的测试时扩展”。该表述将恢复分解为三个相互耦合的决策:何时修订因果历史、在哪里恢复可靠的历史前缀,以及哪种历史配置最能支持后续执行。具体而言,\method{} 通过三个阶段实现这些决策:1)进展感知恢复触发器检测持续的无进展状态,并仅在当前执行状态允许干预时触发恢复;2)历史前缀恢复识别不可靠的历史后缀,检索与当前物理状态匹配的历史锚点,在保留前缀的基础上以最新真实观测为条件重构因果 KV 状态;3)假设验证比较由完整历史、恢复前缀和完全重置三种假设导出的未来延续,并采纳支持度最高的假设。在仿真和真实世界操作环境中的实验表明,该方法在任务成功率上取得了一致的提升,消融实验也验证了各恢复阶段的贡献。
cs.RO / 20 / 2609.18031

Online Multimodal Workload Assessment in Contact-Rich Physical Human-Robot Interaction

接触丰富的物理人机交互中的在线多模态工作负荷评估
Chen, Yanyi, Yang, Fan, Deng, Min
Abstract
Contact-rich physical human--robot interaction (pHRI) imposes time-varying demands associated with physical interaction, motor regulation, and physiological response, motivating continuous assessment of interaction workload. This paper presents an online multimodal assessment framework that integrates interaction wrench, planar tool-center-point (TCP) kinematics, and skin conductance level (SCL) into four interpretable workload-related factors. Their relative contributions are adjusted using path curvature to reflect changes in motion demand and task progression to account for gradual physiological variation over time. The framework was evaluated with 24 participants across 18 controlled combinations of temperature, acoustic noise, and illuminance under two admittance-control modes. Strict leave-one-subject-out (LOSO) evaluation used standardized pupil diameter ($\mathrm{PD}_z$) as an independent physiological reference and included comparisons with static variants and representative state-of-the-art learning-based baselines. The proposed framework achieves a cohort-mean $30\,\mathrm{s}$ block-wise Spearman correlation of $\rho_{30}=0.308$ with the physiological reference, with positive subject-level correspondence in 23 of 24 participants. Its overall performance is comparable to the state-of-the-art learning-based baseline. At the same time, our framework keeps the assessment process transparent through explicit workload-related factors and defined weighting rules, while outperforming the corresponding fixed-weight formulation. The framework also maintains consistent performance across the two tested admittance-control modes. These results support a transparent and interpretable approach to continuous interaction workload assessment in contact-rich pHRI.
Chinese Translation
接触丰富的物理人机交互(pHRI)会带来与物理交互、运动调节和生理响应相关的时变需求,因而需要对交互工作负荷进行连续评估。本文提出了一种在线多模态评估框架,将交互力/力矩(interaction wrench)、平面工具中心点(TCP)运动学以及皮肤电导水平(SCL)整合为四个可解释的与工作负荷相关的因素。通过路径曲率调整各因素的相对权重,以反映运动需求的变化;同时结合任务进程,以考虑生理状态随时间的渐变。该框架在两种导纳控制模式下,由24名参与者在温度、声学噪声和照度的18种受控组合条件下进行了评估。采用严格的留一被试交叉验证(LOSO),以标准化瞳孔直径($\mathrm{PD}_z$)作为独立的生理参考,并与静态变体及具有代表性的最先进基于学习的基线方法进行了比较。所提出的框架与生理参考在30秒分块Spearman相关上达到群体均值$\rho_{30}=0.308$,且在24名参与者中有23人呈现正向的个体层面对应关系。其整体性能与最先进的基于学习的基线方法相当。同时,我们的框架通过显式的工作负荷相关因素和明确的加权规则保持了评估过程的透明性,并优于相应的固定权重方法。该框架在两种测试的导纳控制模式下也保持了一致的性能。这些结果为接触丰富的pHRI中连续交互工作负荷评估提供了一种透明且可解释的方法。
cs.RO / 21 / 2609.18050

Geometric Shortcuts for Complex Trunk Postures: Dual-Helicity Coupling Enables Low-Dimensional Control

复杂象鼻姿态的几何捷径:双螺旋耦合实现低维控制
Huang, Huishi, Chen, Danlu, Preti, Matteo Lo, Liu, Jun, Ang Jr., Marcelo H., Laschi, Cecilia
Abstract
How do elephant trunks generate complex postures without relying solely on fine segmental activation? We propose that part of this complexity arises from a low-dimensional geometric shortcut: dual-helicity coupling between opposite-handed oblique muscles. In a simplified soft-robotic prototype, varying only two geometric parameters generates a broad library of elephant-like postures, suggesting a dual-layer control architecture with implications for continuum robot design and biological hypotheses.
Chinese Translation
象鼻如何在并不完全依赖精细分段肌肉激活的情况下产生复杂的姿态?我们提出,这种复杂性部分源于一种低维的几何捷径:相反手性斜向肌肉之间的双螺旋耦合。在一个简化的软体机器人原型中,仅改变两个几何参数即可生成丰富多样的类象鼻姿态库。这暗示了一种双层控制架构,对连续体机器人设计及生物学假说均具有启发意义。
cs.RO / 22 / 2609.18051

CLASP: A Cluster-Level Autonomous Selective Picking Robot with a Soft Rolling-Band Gripper for Fresh-Market Blueberry Harvesting

CLASP:一种采用软体滚动带式夹持器的簇级自主选择性采摘机器人,用于鲜食蓝莓采收
Xia, Yixuan, Cai, Yilin, Espinoza, Natalia Belen, Li, Changying, Ames, Zilfina Rubio, Zhang, Xin, Chen, Yue
Abstract
Fresh-market blueberries require selective, gentle picking, which is labor-intensive and expensive. Over-the-row machine harvesters are fast but non-selective, bruising mixed-ripeness fruit and limiting yield to the processing market. Selective robotic harvesters typically target individual fruits rather than fruit clusters, which limits harvesting efficiency for small, densely clustered blueberries. This paper presents CLASP, a Cluster-Level Autonomous Selective Picking robot with a Soft Active Rolling-Band Gripper (SARB-Gripper). Two compliant bands envelop the cluster and roll against the fruit, drawing mature berries off in sequence, while closed-loop regulation of the pulling force keeps the applied load below the immature detachment threshold. A global-to-local perception pipeline pairs an eye-to-hand camera for global cluster detection and target selection with an eye-in-hand camera for local localization and cluster orientation estimation. Field measurements confirm a clear detachment-force separation between mature and immature fruit, and the SARB-Gripper reproduces a commanded pulling force to within \SI{3.7}{\percent}, enabling selective harvesting at the cluster level. In end-to-end field trials, CLASP autonomously grasped 23 of 25 presented clusters (\SI{92}{\percent}). With the component cost of approximately \$3326 per unit, CLASP offers a scalable approach to selective cluster-level harvesting for fresh-market blueberries.
Chinese Translation
鲜食蓝莓需要进行选择性、轻柔的采摘,而这一过程劳动强度大且成本高昂。跨行式机械采收机速度快,但缺乏选择性,会碰伤成熟度混杂的果实,使产量只能进入加工市场。选择性机器人采收机通常以单个果实而非果簇为采摘目标,这限制了其采摘体积小、簇生密集的蓝莓的采收效率。本文提出了CLASP——一种具备软体主动滚动带式夹持器(SARB-Gripper)的簇级自主选择性采摘机器人。两条柔性带包裹果簇并与果实产生滚动摩擦,依次将成熟浆果摘下,同时对拉拔力进行闭环调节,使所施加的载荷保持在未成熟果实的脱离阈值以下。一个由全局到局部的感知流水线将用于全局果簇检测与目标选择的眼在手外(eye-to-hand)相机和用于局部定位与果簇姿态估计的眼在手(eye-in-hand)相机相结合。田间实测证实成熟果实与未成熟果实的脱离力存在明显区分,且SARB-Gripper能够以不超过3.7%的误差复现指令拉拔力,从而实现簇级选择性采摘。在端到端田间试验中,CLASP自主抓取了25个目标果簇中的23个(92%)。每台组件成本约为3326美元,CLASP为鲜食蓝莓的选择性簇级采收提供了一种可扩展的方法。
cs.RO / 23 / 2609.18070

An Efficient Algorithm for Minimum-Pressure Growth Planning of Vine Robots

藤蔓机器人最小压力生长规划的高效算法
Torres, Andres C., Marcucci, Tobia, Hawkes, Elliot W.
Abstract
Vine robots navigate cluttered environments by extending from their tip. Although their ability to operate in such environments has been extensively demonstrated, little work has addressed growth planning, i.e., finding optimal growth paths. Moreover, existing planners do not account for the growth pressure necessary to follow a given path, which can cause the robot to burst when it is too high. In this paper, we address the problem of finding minimum-pressure paths for vine robots growing around polytopic obstacles. We propose an efficient algorithm that is guaranteed to find globally optimal solutions in 2D and approximate solutions in 3D, with an error that vanishes as a discretization parameter approaches zero. First, we derive a growth pressure equation for vine robots of arbitrary shape, which we use to show that there always exists a minimum-pressure path that is piecewise-linear and can bend only at specific points on the obstacles. We then leverage this observation to reduce the growth-planning problem to a shortest-path problem with time-dependent weights, which we efficiently solve using a modified Dijkstra's algorithm. We demonstrate the speed and scalability of our approach through numerical simulations. We also validate our algorithm with hardware experiments and provide an open-source and high-performance implementation in the Python package, VinePlanner: https://github.com/Ahsoka/VinePlanner.
Chinese Translation
藤蔓机器人通过从其尖端延伸来在杂乱环境中导航。尽管其在这种环境中运行的能力已被广泛验证,但关于生长规划(即寻找最优生长路径)的研究却很少。此外,现有的规划器未考虑沿给定路径生长所需的生长压力,当压力过高时可能导致机器人爆裂。本文研究了藤蔓机器人绕多面体障碍物生长时寻找最小压力路径的问题。我们提出了一种高效算法,该算法在二维空间中保证找到全局最优解,在三维空间中找到近似解,且其误差随着离散化参数趋近于零而消失。首先,我们推导了任意形状藤蔓机器人的生长压力方程,并利用该方程证明:总是存在一条分段线性的最小压力路径,且该路径仅在障碍物上的特定点处弯曲。随后,我们利用这一观察结果,将生长规划问题转化为具有时变权重的最短路径问题,并使用改进的Dijkstra算法高效求解。我们通过数值仿真展示了该方法的速度和可扩展性,并通过硬件实验验证了该算法,同时在Python包VinePlanner中提供了开源且高性能的实现:https://github.com/Ahsoka/VinePlanner。
cs.RO / 24 / 2609.18073

Characterizing Refraction-Induced Ranging Bias in Underwater Collaborative Localization

水下协同定位中折射引起的测距偏差表征
Kogucki, Timothy, Papalia, Alan
Abstract
This work studies how refraction-induced bias on acoustic ranging affects multi-agent collaborative localization in a range of oceanographic conditions and spatial scales. While multi-agent range-aided navigation, which uses range measurements to either fixed infrastructure or other agents, is a promising solution to the challenges of large-scale underwater localization, its accuracy depends strongly on the quality of range measurements. Sound speed variability induces refraction (bending) of acoustic rays, yet, for algorithmic tractability, standard sensor fusion pipelines assume straight-line propagation. This refraction systematically biases range measurements to be longer than the straight-line assumption predicts. However, the effects of this bias on multi-agent collaborative localization on kilometer scales remains unexplored. We present a series of simulated experiments with several agents operating over kilometer scales. The simulation uses HYCOM reanalysis data to recreate realistic oceanographic conditions, ray tracing to generate refraction-informed ranges, and a centralized multi-agent factor graph estimator to quantify the resulting measurement bias on estimated trajectories. Preliminary results indicate that refraction-induced bias can induce significant degradation of estimated trajectories, particularly in regions with sharp sound-speed gradients. We also share the simulation environment to support further studies https://github.com/UMich-RobotExploration/manta-ray.
Chinese Translation
本研究探讨了折射引起的声学测距偏差如何影响一系列海洋条件和空间尺度下的多智能体协同定位。多智能体测距辅助导航利用固定基础设施或其他智能体之间的距离测量,是应对大规模水下定位挑战的一种有前景的解决方案,但其精度在很大程度上取决于距离测量的质量。声速变化会引起声线的折射(弯曲),然而出于算法可实现性的考虑,标准的传感器融合流程假设声波沿直线传播。这种折射会系统性地使距离测量值偏长,超出直线传播假设的预测。然而,这种偏差对公里尺度上多智能体协同定位的影响仍未被探索。我们进行了一系列仿真实验,其中多个智能体在公里尺度上运行。该仿真使用HYCOM再分析数据再现真实的海洋环境条件,采用射线追踪生成考虑折射的距离测量,并利用集中式多智能体因子图估计器来量化由此产生的测量偏差对估计轨迹的影响。初步结果表明,折射引起的偏差会导致估计轨迹的显著退化,尤其是在声速梯度剧烈的区域。我们还公开了仿真环境以支持进一步研究:https://github.com/UMich-RobotExploration/manta-ray。
cs.RO / 25 / 2609.18084

Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models

并非所有层都需要微调:视觉-语言-动作模型中适应的诊断与引导
Syed, Shahram Najam, Jakobsson, Arthur, Sachdev, Prayuj, Ichnowski, Jeffrey
Abstract
Fine-tuning a Vision-Language-Action (VLA) model for a new deployment environment is expensive, yet most methods apply uniform-capacity adapters to every network region as if every region requires equal adjustment. This paper tests that assumption on five architecturally diverse VLAs (OpenVLA-OFT, $\pi_0$, SmolVLA, DTP, Octo; 93M-7B parameters). Measuring per-region adaptation cost as normalized parameter displacement under region-isolated fine-tuning reveals an adaptation spectrum in which appearance shifts concentrate cost in the vision encoder, instruction shifts in the language backbone, and novel-object shifts in the vision encoder together with the action head, across all five architectures. To exploit this structure, we introduce a pipeline that observes, diagnoses, allocates, and adapts. From ten unlabeled target observations and without fine-tuning, the diagnostic estimates per-region cost by combining reference-free gradient and Monte Carlo Dropout signals with a Centered Kernel Alignment score against a cached source reference; the allocator converts the estimates into variable-rank LoRA adapters under a parameter budget and freezes well-calibrated regions; and standard LoRA fine-tuning trains the resulting adapters. The diagnostic ranks regions within each deployment at a median Spearman of 0.91, and the allocation matches or exceeds uniform LoRA at every budget we tested on LIBERO and CALVIN. On a physical xArm-7, the pipeline matches full fine-tuning under an instruction-wording shift with 0.04% of its trainable parameters, and on five held-out scenes evaluated without retraining it leads every baseline, with 11-23 successes of 30 rollouts against 8-18 for the strongest parameter-efficient baseline at equal or larger budgets and 2-11 for full fine-tuning. These results suggest that adaptation cost in VLAs is structured enough to measure before fine-tuning begins.
Chinese Translation
针对新的部署环境对视觉-语言-动作(Vision-Language-Action, VLA)模型进行微调代价高昂,然而大多数方法都将等容量的适配器(adapter)应用于网络的每个区域,仿佛每个区域都需要同等的调整。本文在五个架构各异的VLA模型(OpenVLA-OFT、$\pi_0$、SmolVLA、DTP、Octo;参数量93M至7B)上检验了这一假设。通过在区域隔离微调下以归一化参数位移来测量各区域的适应成本,研究发现存在一个适应谱系:外观变化使成本集中于视觉编码器,指令变化集中于语言主干,而新物体变化则集中于视觉编码器与动作头,且这一规律在全部五种架构中均成立。为利用这一结构,我们提出了一个包含观察、诊断、分配和适应四个阶段的流水线。仅利用十条无标注的目标观测数据且无需微调,诊断模块通过将免参考梯度信号与蒙特卡洛Dropout信号,以及相对缓存源参考的中心化核对齐(Centered Kernel Alignment, CKA)分数相结合,估计各区域的适应成本;分配模块在给定参数预算下将这些估计转化为可变秩的LoRA适配器,并冻结校准良好的区域;随后采用标准LoRA微调训练所得适配器。该诊断模块在每个部署场景内对区域排序的中位数Spearman相关系数达0.91,且在我们于LIBERO和CALVIN上测试的每一档预算下,该分配方法均能达到或超过均匀LoRA的表现。在物理xArm-7机器人上,该流水线在指令措辞变化下以0.04%的可训练参数量达到与全参数微调相当的效果;在五个未经再训练的预留测试场景中,其表现优于所有基线方法——在30次执行中取得11至23次成功,而参数预算相同或更高的最强参数高效基线仅为8至18次,全参数微调仅为2至11次。这些结果表明,VLA模型的适应成本具有足够的结构化特性,可以在微调开始之前加以测量。
cs.RO / 26 / 2609.18092

ReRadar: Robust Radar Global Localization via Rotation-Equivariant Descriptor Learning

ReRadar:基于旋转等变描述子学习的鲁棒雷达全局定位
Nguyen, Duc Manh, Dao, Truong Giang, Luong, Gia Nghiem, Hoang, Viet Trung, Nguyen, Anh Quang
Abstract
Global localization with scanning millimeter-wave radar remains challenging because place-recognition descriptors often discard spatial structure needed for accurate pose retrieval. We present ReRadar, a radar global localization pipeline that extracts rotation-equivariant intermediate features using steerable convolutional neural networks, forms rotation-invariant descriptors through group pooling and NetVLAD aggregation, and combines descriptor retrieval with landmark-based matching to estimate the robot's three-degree-of-freedom (3-DoF) pose. Across fixed database-query evaluations, ReRadar with target-dataset adaptation achieves 99.37% Recall@1 on OORD Bellmouth, 91.44% Recall@1 with 80.99% F1_max on Mulran DCC01, and 99.38% Recall@1 on falling-snow Boreas sequence. Without target-dataset data, the cross-dataset model reaches 98.07% Recall@1 on OORD, performing comparably to the evaluated state-of-the-art methods.
Chinese Translation
利用扫描式毫米波雷达进行全局定位仍然具有挑战性,因为位置识别描述子通常会丢弃实现精确位姿恢复所需的空间结构信息。我们提出了ReRadar,一种雷达全局定位流程,该流程利用可操纵卷积神经网络(steerable convolutional neural networks)提取旋转等变的中间特征,通过群池化(group pooling)和NetVLAD聚合构建旋转不变的描述子,并将描述子检索与基于地标(landmark)的匹配相结合以估计机器人的三自由度(3-DoF)位姿。在固定数据库-查询的评估中,经过目标数据集自适应的ReRadar在OORD Bellmouth数据集上达到99.37%的Recall@1,在Mulran DCC01数据集上达到91.44%的Recall@1和80.99%的F1_max,在下雪场景的Boreas序列上达到99.38%的Recall@1。在不使用目标数据集数据的情况下,跨数据集模型在OORD上达到98.07%的Recall@1,性能与所评估的最先进方法相当。
cs.RO / 27 / 2609.18100

Beyond Pixel Similarity: Task-Aware Evaluation of GAN-Based Synthetic Sonar Data for Robotic Perception

超越像素相似性:面向机器人感知的基于GAN的合成声呐数据的任务感知评估
Keen, Hannan Ejaz, Fraz, Muhammad Moazam, Berns, Karsten
Abstract
Synthetic data can reduce the cost of collecting and annotating training data for robotic perception, but generating sensor observations that preserve the characteristics relevant to downstream perception remains challenging, particularly for sonar imagery. In this work, we investigate whether conventional image-fidelity metrics adequately reflect the downstream perception performance of GAN-generated synthetic sonar data. We employ a Pix2Pix conditional generative adversarial network with four discriminator configurations characterized by different receptive fields: PixelGAN, PatchGAN-16, PatchGAN-70, and ImageGAN. The models are trained using sonar imagery from two datasets and evaluated using conventional image-fidelity metrics, including Structural Similarity Index (SSIM), Peak Signal-to-Noise Ratio (PSNR), and Mean Squared Error (MSE). To complement these pixel-level measures with task-oriented evaluation, YOLOX-S, YOLOX-L, and Faster R-CNN detectors are trained exclusively on real sonar imagery and subsequently evaluated on the GAN-generated images using identical test samples and annotations across all discriminator configurations. The results reveal a discrepancy between image-fidelity and downstream object-detection performance: the configuration achieving the best SSIM, PSNR, and MSE does not consistently yield the best detection performance. In particular, PatchGAN configurations achieve strong downstream detection results despite not achieving the highest pixel-level similarity scores. These findings suggest, for the datasets and models considered, pixel-level image-fidelity metrics alone may not consistently capture the task-relevant realism of synthetic sonar observations and motivate the use of task-aware evaluation for synthetic sensor data intended for robotic perception.
Chinese Translation
合成数据可以降低机器人感知训练数据采集与标注的成本,但生成能够保留与下游感知相关特性的传感器观测数据仍然具有挑战性,尤其是对于声呐图像而言。在本工作中,我们研究了传统的图像保真度指标是否能够充分反映GAN生成的合成声呐数据的下游感知性能。我们采用Pix2Pix条件生成对抗网络,并配置了四种具有不同感受野的判别器:PixelGAN、PatchGAN-16、PatchGAN-70和ImageGAN。模型使用来自两个数据集的声呐图像进行训练,并采用传统图像保真度指标进行评估,包括结构相似性指数(SSIM)、峰值信噪比(PSNR)和均方误差(MSE)。为了以面向任务的评估补充这些像素级度量,我们仅在真实声呐图像上训练YOLOX-S、YOLOX-L和Faster R-CNN检测器,随后使用相同的测试样本和标注,在所有判别器配置下对GAN生成的图像进行评估。结果显示,图像保真度与下游目标检测性能之间存在差异:在SSIM、PSNR和MSE上表现最优的配置并不一定能持续带来最佳的检测性能。特别是,PatchGAN配置尽管未获得最高的像素级相似性得分,却取得了较强的下游检测结果。这些发现表明,就所研究的数据集和模型而言,仅依靠像素级图像保真度指标可能无法持续捕捉合成声呐观测数据中与任务相关的真实性,因此对于面向机器人感知的合成传感器数据,有必要采用任务感知的评估方法。
cs.RO / 28 / 2609.18108

Technical Report: One-Step Drifting Action Heads for GR00T N1.7

技术报告:面向GR00T N1.7的单步漂移动作头
Shao, Xihe
Abstract
One-step action generation can substantially reduce the inference cost of vision-language-action (VLA) policies, but its effect on closed-loop task success remains an open question. This technical report studies a GR00T N1.7 variant in which the iterative diffusion-transformer action head is replaced by a one-step drifting action head, together with an overlap-conditioned extension for asynchronous chunk replacement. All multi-seed drifting runs were trained on two NVIDIA A800 GPUs. On LIBERO, the action head reduces the mean model-forward time of the action head from approximately $45.3\,\mathrm{ms}$ to $5.0\,\mathrm{ms}$, while the measured backbone-plus-head time falls from approximately $70.0\,\mathrm{ms}$ to $30.6\,\mathrm{ms}$. However, this speedup is accompanied by a systematic reduction in task success. Across three drifting seeds, success is $64.0\pm4.0\%$ on LIBERO-Spatial, $52.0\pm1.0\%$ on LIBERO-Goal, and $26.0\pm2.6\%$ on LIBERO-Long. The low seed variance indicates that the degradation is not explained by random initialization alone. We report the result as a speed--success trade-off rather than an overall improvement, and discuss likely contributing factors including deterministic one-step mode averaging, batch-dependent geometry estimation, long open-loop chunk execution, and the fact that synchronous LIBERO evaluation does not exercise the asynchronous overlap path.
Chinese Translation
单步动作生成可以显著降低视觉-语言-动作(VLA)策略的推理成本,但其对闭环任务成功率的影响仍是一个悬而未决的问题。本技术报告研究了GR00T N1.7的一个变体,其中将迭代式扩散Transformer动作头替换为单步漂移(drifting)动作头,并引入了一种基于重叠条件的异步动作块(chunk)替换扩展方法。所有多种子漂移训练均在两块NVIDIA A800 GPU上完成。在LIBERO基准上,该动作头将动作头的平均前向推理时间从约 $45.3\,\mathrm{ms}$ 降至 $5.0\,\mathrm{ms}$,实测的主干网络加动作头总时间从约 $70.0\,\mathrm{ms}$ 降至 $30.6\,\mathrm{ms}$。然而,这种加速伴随着任务成功率的系统性下降。在三个漂移随机种子下,LIBERO-Spatial的成功率为 $64.0\pm4.0\%$,LIBERO-Goal为 $52.0\pm1.0\%$,LIBERO-Long为 $26.0\pm2.6\%$。较低的种子间方差表明,性能下降不能仅归因于随机初始化。我们将该结果报告为一种速度与成功率之间的权衡,而非整体性能提升,并讨论了可能的成因,包括确定性的单步模式平均、依赖批大小的几何估计、长时开环动作块执行,以及同步式LIBERO评估未能触发异步重叠路径这一事实。
cs.RO / 29 / 2609.18111

A Comprehensive Review of Generative Physical Artificial Intelligence

生成式物理人工智能综述
Gaba, Satyam, Rana, Krutiksinh, Sai, Siva, Chamola, Vinay, Niyato, Dusit
Abstract
The integration of large-scale foundation models with physical embodiments has led to significant advancements in robotics known as Generative Physical Artificial Intelligence (GPAI). These agentic AI systems autonomously perceive, reason, and act in complex real-world situations. This survey comprehensively analyzes GPAI systems, focusing on their architectural foundations, current applications, and key limitations. We introduce a taxonomy of five distinct approaches: Robot Foundation Models (RFMs) for cross-platform skill transfer; Vision-Language Action (VLA) models for end-to-end multi-modal perception and control; Large Behavior Models (LBMs) for human-like movement generation; Diffusion Policy Models (DPMs) for diffusion model-based temporally coherent action generation; and World Foundation Models (WFMs) for physics-compliant simulation and data generation. We examine how these approaches complement each other: WFMs generate training data for VLAs and DPMs, RFMs enable cross-platform deployment of learned policies, while LBMs provide motion priors for natural behavior. Through examples across autonomous vehicles, industrial automation, healthcare robotics, and humanoid systems, we identify significant performance improvements and summarize promising research directions in data-efficient learning, sim-to-real transfer, edge-compatible architectures, and safety frameworks. These insights advance embodied AI for IoT-connected environments where intelligent agents interact with networked sensors, actuators, and edge devices.
Chinese Translation
大规模基础模型与物理实体的结合推动了机器人领域的重大进展,这一领域被称为生成式物理人工智能(GPAI)。这些智能体AI系统能够在复杂的现实场景中自主感知、推理和行动。本综述全面分析了GPAI系统,重点关注其架构基础、当前应用及关键局限。我们提出了一个包含五种不同方法的分类体系:用于跨平台技能迁移的机器人基础模型(RFM);用于端到端多模态感知与控制的视觉-语言-动作(VLA)模型;用于类人运动生成的大行为模型(LBM);基于扩散模型的时间连贯动作生成的扩散策略模型(DPM);以及用于符合物理规律仿真与数据生成的世界基础模型(WFM)。我们考察了这些方法之间的互补关系:WFM为VLA和DPM生成训练数据,RFM实现已学习策略的跨平台部署,而LBM则为自然行为提供运动先验。通过自动驾驶、工业自动化、医疗机器人和人形系统等领域的实例,我们识别了显著的性能提升,并总结了在数据高效学习、仿真到现实迁移、边缘兼容架构和安全框架等方向上具有前景的研究方向。这些洞见推进了面向物联网连接环境的具身智能发展,使智能体能够与网络化传感器、执行器和边缘设备进行交互。
cs.RO / 30 / 2609.18117

OpenDexGrasp: Open-vocabulary Task-Oriented Dexterous Grasping

OpenDexGrasp:开放词汇的任务导向灵巧抓取
Zhang, Jiyao, Wang, Junhan, Wang, Tianyu, Chen, Zeyuan, Bolten, Anthony, Peng, Yitong, Dong, Hao
Abstract
Dexterous grasp synthesis has advanced rapidly in generating stable and physically plausible hand poses, but real-world manipulation requires grasps that preserve the function implied by the task. We study open-vocabulary task-oriented dexterous grasp generation, where a robot must infer functional intent from free-form language, ground it in multi-view visual observations and object geometry, and generate an executable high-degree-of-freedom grasp. We present OpenDexGrasp, a unified data and generative modeling framework for this setting. OpenDexVerse provides dual-source supervision organized by the Coverage-to-Alignment (C2A) Recipe: OpenDex-Scale offers large-scale semantic and geometric coverage through automatic grasp synthesis and vision-language annotation, while OpenDex-Align supplies high-quality embodied alignment through human teleoperation and category-level transfer. OpenDexGrasp learns a shared perception-action latent representation that couples open-vocabulary vision-language context with dexterous action generation. Affordance grounding and grasp generation provide complementary supervision over this latent space, enabling direct generation of task-consistent dexterous grasps without a separate affordance-to-pose inference stage. Extensive simulation and real-robot experiments demonstrate improved functional alignment, physical feasibility, generalization to unseen categories, and real-world execution success. Additional details and videos are available at https://opendexgrasp.github.io/.
Chinese Translation
灵巧抓取合成在生成稳定且物理合理的手部姿态方面进展迅速,但现实世界的操作要求抓取能够保留任务所隐含的功能。我们研究了开放词汇的任务导向灵巧抓取生成问题,即机器人需要从自由形式的语言中推断功能性意图,将其与多视角视觉观测和物体几何相联系,并生成可执行的高自由度抓取。我们提出了OpenDexGrasp,一个面向该场景的统一数据与生成建模框架。OpenDexVerse提供了由“覆盖到对齐”(Coverage-to-Alignment,C2A)方案组织的双源监督:OpenDex-Scale通过自动抓取合成与视觉-语言标注提供大规模的语义与几何覆盖;OpenDex-Align则通过人类遥操作和类别级迁移提供高质量的具身对齐数据。OpenDexGrasp学习一个共享的感知-动作潜在表示,将开放词汇的视觉-语言上下文与灵巧动作生成相耦合。可供性(affordance)对齐与抓取生成在该潜在空间上提供互补监督,从而能够直接生成任务一致的灵巧抓取,无需单独的“可供性到姿态”推理阶段。大量仿真与真实机器人实验表明,本方法在功能对齐、物理可行性、对未见类别的泛化以及真实世界执行成功率方面均有提升。更多细节与视频请访问 https://opendexgrasp.github.io/。
cs.RO / 31 / 2609.18119

Fetch My Beer: Synthetic-to-real Hierarchical Policy for Smooth Pick-and-place

Fetch My Beer:面向平滑抓取放置的合成到真实分层策略
Li, Yingyue, Zhang, Chenyangguang, Zhang, Ruida, Fu, Bowen, Zhai, Guangyao, Ji, Xiangyang
Abstract
Many real-world robotic applications require dynamically sensitive manipulation, where success depends not only on reaching a target state but on maintaining stable object dynamics throughout execution. We study the stable transport of liquid-filled containers, where a robot must move objects to target locations while suppressing sloshing and preventing spillage. Unlike conventional pick-and-place, this task imposes stringent requirements on motion smoothness and trajectory-level stability, exposing clear limitations in existing systems. Specifically, fluid simulation remains too costly for online reinforcement learning; human teleoperation introduces unintended accelerations that induce sloshing during imitation learning; and current policy pipelines optimize for task completion rather than dynamic stability. We propose a synthetic-to-real framework coupling physically validated data generation with a hierarchical, diffusion-based controller. The scalable data pipeline synthesizes grasps, filters unstable poses via a vision-language model, and validates transport trajectories through fluid simulation. The policy is organized with a high-level module that translates language and visual observations into SE(3) control targets, and a latent diffusion controller that first plans efficiently in a compact latent space and then decodes dense action chunks, enabling the high control frequency needed for smooth and stable motion. Extensive experiments show our system outperforms state-of-the-art manipulation policies in transport smoothness and dynamic stability. Our project page: https://fetch-my-beer.github.io/
Chinese Translation
许多真实世界的机器人应用需要动态敏感的操作,其成功不仅取决于到达目标状态,还取决于在整个执行过程中保持稳定的物体动力学。我们研究了液体容器的稳定运输任务,机器人必须在将物体移动到目标位置的同时抑制液体晃动并防止溢出。与传统的抓取放置任务不同,该任务对运动平滑性和轨迹级稳定性提出了严格要求,暴露了现有系统的明显局限性。具体而言,流体仿真对于在线强化学习而言成本过高;人类遥操作会在模仿学习中引入意外的加速度从而诱发液体晃动;而现有的策略流程则只优化任务完成而非动态稳定性。我们提出了一种合成到真实的框架,将经过物理验证的数据生成与基于扩散模型的分层控制器相结合。该可扩展的数据流水线综合生成抓取姿态,通过视觉语言模型过滤不稳定的姿态,并通过流体仿真验证运输轨迹。该策略由高层模块和潜在扩散控制器组成:高层模块将语言和视觉观测转化为SE(3)控制目标;潜在扩散控制器首先在紧凑的潜在空间中高效规划,然后解码密集的动作块,从而实现平滑稳定运动所需的高控制频率。大量实验表明,我们的系统在运输平滑性和动态稳定性方面优于最先进的操作策略。项目页面:https://fetch-my-beer.github.io/
cs.RO / 32 / 2609.18153

Prior Evolution and Task Alignment for Aerial Grasping

面向空中抓取的先验演化与任务对齐
Deng, Weiliang, Dang, Zhengyang, Mu, Yao, Lyu, Ximin
Abstract
Aerial grasping is a remarkable capability exhibited by predatory birds, allowing them to capture prey through highly coordinated maneuvers in flight. Inspired by this capability, researchers have developed various formulations to reproduce such maneuvers through trajectory optimization. However, two limitations remain in practice. First, the resulting optimization problem is highly nonconvex and sensitive to initialization, making high-quality solutions difficult to obtain under a limited computational budget. Second, prescribed numerical objectives are human-designed abstractions that describe successful grasping through a limited set of mathematically tractable quantities and may not fully capture what determines task success. We investigate how learning can address these limitations within an analytical planner. Accordingly, a trajectory prior is first learned from optimized motions and then evolved through a CEM-based process that evaluates sampled initializations with the deployed optimizer and retains favorable ones as new supervision. An Execution-Aware Critic learns from contact, lift, and completion outcomes to assess whether the optimized trajectories are likely to succeed in physical execution. Its frozen energy can further serve as a differentiable grasping cost, allowing execution data to directly shape trajectory generation. Simulation and real-world experiments demonstrate improved optimization reliability, trajectory consistency, and grasping performance.
Chinese Translation
空中抓取是猛禽所展现出的一种卓越能力,使它们能够通过高度协调的飞行机动来捕获猎物。受这一能力启发,研究者们已提出多种方法,通过轨迹优化来复现此类机动动作。然而,实践中仍存在两个局限。其一,所得到的优化问题高度非凸且对初始化敏感,使得在有限的计算预算下难以获得高质量的解。其二,预设的数值目标是由人类设计的抽象描述,仅通过有限的数学上易处理的量来刻画成功抓取,可能无法完全捕捉决定任务成功的因素。我们研究了如何通过学习在解析式规划器中解决这些局限。为此,首先从优化后的运动中学习轨迹先验,然后通过基于CEM(交叉熵方法)的过程对其进行演化:利用部署的优化器评估采样得到的初始值,并保留有利的样本作为新的监督信号。一个执行感知评论家模型(Execution-Aware Critic)从接触、提升和完成等结果中学习,以评估优化得到的轨迹在物理执行中能否成功。其冻结的能量函数可进一步作为可微的抓取代价,使执行数据能够直接塑造轨迹生成过程。仿真与真实世界实验表明,该方法提升优化可靠性、轨迹一致性和抓取性能。
cs.RO / 33 / 2609.18164

Energy-Regularized Imitation Learning for Force- and Work-Aware Robotic Manipulation

面向力与功感知机器人操作的能量正则化模仿学习
Otani, Toshiki, Taketsugu, Hiromu, Ukita, Norimichi
Abstract
This paper studies energy-aware manipulation as a physically grounded learning problem. We define a joint-space mechanical-work proxy from joint torque and angular displacement, and train a differentiable energy predictor that estimates this work from robot states and actions. The predictor converts a non-differentiable simulator-side physical quantity into a differentiable regularizer for fine-tuning a pretrained manipulation policy. We instantiate the framework with RVT-2 on RLBench and evaluate 12 manipulation tasks involving object contact, articulated motion, placement, pushing, and sweeping. The proposed fine-tuning reduces the average mechanical work from 208.8J to 204.4J (i.e., 2.1% reduction), while the mean task success rate also increases slightly from 86.2% to 86.9%. These results show that work-aware policy optimization can suppress physically inefficient motion without requiring an explicit differentiable dynamics model.
Chinese Translation
本文将能量感知的操作研究作为一个具有物理基础的学习问题。我们基于关节力矩与角位移定义了一个关节空间机械功代理量,并训练一个可微的能量预测器,从机器人状态和动作中估计该功。该预测器将模拟器端不可微的物理量转换为可微的正则项,用于微调预训练的操作策略。我们在RLBench上以RVT-2为例实例化该框架,并评估了涉及物体接触、关节运动、放置、推动和清扫的12个操作任务。所提出的微调方法将平均机械功从208.8焦耳降低到204.4焦耳(即降低2.1%),同时平均任务成功率也从86.2%略有提升至86.9%。这些结果表明,功感知的策略优化无需显式的可微动力学模型即可抑制物理上低效的运动。
cs.RO / 34 / 2609.18165

LUMO: Designing Luminous Contact Morphology for Repeatable Whole-Finger Contact Observation

LUMO:设计发光接触形态以实现可重复的全指接触观测
Kang, Dong Ho, Ko, Youngsu, Sentis, Luis
Abstract
A low-impedance robot finger reports through joint torque how strongly it is loaded, but the same torque can arise from a small force near the fingertip or a large force near the joint. Resolving the force therefore requires knowing where along the finger contact occurred. LUMO makes that location externally observable. Embedded LEDs illuminate a compliant silicone pad, and contact deforms the pad so that light emerging from the finger's side changes in a pattern set by where the load acts. Because the same structure also carries the contact load, we optimize its cross-section, including the pad profile, rigid carrier, and lateral void, for two behaviors at once. Mechanically, the pad conforms under low preload while the carrier increasingly restricts further deformation as load rises. Optically, different contact locations produce separated responses on the finger's side. The search uses rigid--soft contact simulation, ray tracing, and multi-objective Bayesian optimization. Across two silicones, six contact locations, and 10- and 30-mm spherical indenters, the optimized morphologies improve neighboring-location separation relative to variation from re-establishing contact by \(15\)--\(59\%\). Estimating contact location from the optical response using the known LED spacing and combining it with joint torque gives \(1.44~\mathrm{N}\) normal-force MAE over 931 samples. In a two-finger hand, localized side responses appear on several links simultaneously during grasps.
Chinese Translation
低阻抗机器人手指通过关节力矩报告其受力大小,但相同的力矩可能来自指尖附近的小力或靠近关节处的大力。因此,要确定力的大小就需要知道接触发生在手指的哪个位置。LUMO使该位置可从外部观测。嵌入的LED照亮柔性硅胶垫,接触使硅胶垫变形,从而改变从手指侧面射出的光线,其变化模式由负载作用位置决定。由于同一结构还需承载接触负载,我们对其横截面(包括垫片轮廓、刚性载体和侧向空腔)进行了优化,以同时满足两种行为需求。在力学上,垫片在低预载下发生贴合变形,而随着负载增大,载体逐渐限制进一步变形。在光学上,不同的接触位置在手指侧面产生可分离的响应。搜索过程采用了刚-柔接触仿真、光线追踪和多目标贝叶斯优化。在两种硅胶、六个接触位置以及10毫米和30毫米球形压头的条件下,优化后的形态相对于重新建立接触引起的变异,将相邻位置的区分度提升了15%至59%。利用已知的LED间距从光学响应中估计接触位置,并将其与关节力矩相结合,在931个样本上实现了1.44 N的法向力平均绝对误差(MAE)。在一个双指灵巧手中,抓取过程中多个连杆上同时出现了局部化的侧面响应。
cs.RO / 35 / 2609.18167

Characterizing Replay Retention Under Dynamics Shift in Model-Based Reinforcement Learning

基于模型的强化学习中动态变化下的回放保留特性研究
Yang, Everest, Thompson, Skye, Konidaris, George D.
Abstract
Adapting to changes in robot dynamics requires learning from new data without discarding experience that may still be useful. In continual model-based reinforcement learning (RL), replay collected before a dynamics change can slow adaptation, while removing it unnecessarily reduces available training data and can be especially costly if earlier dynamics return. We study when recent transitions are preferable to the full replay history. Two quantities characterize this trade-off: change magnitude and age-staleness area under the curve (AUC), measuring how well transition age separates stale from fresh data. Forgetting stale data helps after large permanent shifts but hurts when dynamics recur and older data becomes useful again. Choosing a replay strategy therefore depends on predicting when older data will help or hurt. We test these effects across two locomotion morphologies, two model-based RL algorithms, and Real-World RL benchmark perturbations. Because ground-truth staleness labels are unavailable on deployed robots, we evaluate whether an estimator built from interaction data can still provide the quantities needed to choose a replay strategy after permanent changes. Our results show that replay retention depends on change magnitude and on how the dynamics evolve.
Chinese Translation
适应机器人动力学(dynamics)的变化需要从新数据中学习,同时不丢弃可能仍有用的经验。在持续基于模型的强化学习(RL)中,在动力学变化之前收集的回放数据可能会减慢适应速度,而移除这些数据又会不必要地减少可用的训练数据,并且在早期动力学再次出现时代价尤其高昂。我们研究了何时近期转移样本优于完整的回放历史。两个量可以刻画这一权衡:变化幅度和转移样本年龄的时效性曲线下面积(age-staleness AUC),后者衡量转移样本的年龄能在多大程度上区分陈旧数据与新鲜数据。在发生较大的永久性变化后,遗忘陈旧数据有所帮助,但当动力学再次出现、旧数据重新变得有用时,遗忘则会带来损害。因此,选择回放策略取决于预测旧数据何时有益、何时有害。我们在两种运动形态、两种基于模型的RL算法以及Real-World RL基准扰动下测试了这些效应。由于在部署的机器人上无法获得真实(ground-truth)的陈旧性标签,我们评估了由交互数据构建的估计器是否仍能在永久性变化后提供选择回放策略所需的量。我们的结果表明,回放保留策略取决于变化幅度以及动力学的演化方式。
cs.RO / 36 / 2609.18169

Approximating High Dimensional Self-Motion Manifolds via Deep Generative Models

基于深度生成模型的高维自运动流形近似
Gao, Haitao, Song, Yang, Wu, Liao
Abstract
Self-motion manifold (SMM) characterizes the geometric structure of the infinite inverse kinematic solutions set of a redundant manipulator at a fixed end-effector pose, and its efficient recovery underpins feasible and global optimal motion planning. Existing methods such as null-space continuation and learning-based methods are formulated around the assumption that an SMM is a curve, and do not extend to higher redundancy orders. We instead adopt a probabilistic view: SMMs are the support of the conditional posterior over configurations given a target pose, so that recovering it reduces to sampling from a learned distribution and separating its disjoint components by clustering. The formulation is independent of the manifold dimension and requires no architectural change as the redundancy order grows. In this work, we demonstrate that our method can approximate 1-D SMMs with performance comparable to the latest null-space continuation and learning-based approach, and that it is the first method capable of approximating highly redundant 4-D SMMs in a 7R manipulator for position tasks. Project website: \href{https://github.com/accuracy-maker/high-dimenstional-self-motion-manifold-approximation}{https://github.com/accuracy-maker/high-dimenstional-self-motion-manifold-approximation}
Chinese Translation
自运动流形(Self-Motion Manifold, SMM)刻画了冗余机械臂在固定末端位姿下无穷多逆运动学解集的几何结构,其高效恢复是实现可行且全局最优运动规划的基础。现有方法(如零空间延拓法和基于学习的方法)均建立在SMM为一条曲线的假设之上,无法扩展到更高冗余阶数。我们转而采用概率视角:SMM是给定目标位姿时构型条件后验分布的支撑集,因此恢复SMM可转化为从学习到的分布中采样,并通过聚类将其不连通的分量分离开来。该公式化与流形维度无关,且随着冗余阶数的增加无需对网络结构做任何改动。在本工作中,我们证明了所提方法能够以与最新的零空间延拓法和基于学习的方法相当的性能近似一维SMM,并且是首个能够近似7R机械臂在位置任务中高度冗余的四维SMM的方法。项目网站:\href{https://github.com/accuracy-maker/high-dimenstional-self-motion-manifold-approximation}{https://github.com/accuracy-maker/high-dimenstional-self-motion-manifold-approximation}
cs.RO / 37 / 2609.18174

TacBPM: A Tactile-conditioned Behavior Prior Model for Dexterous Reorientation

TacBPM:一种用于灵巧重定向的触觉条件化行为先验模型
Yin, Jie, Xing, Wanli, Zhao, Zeyuan, Zhu, Xuezhou, Deng, Zhijie, Zhang, Kaifeng
Abstract
Dexterous in-hand manipulation requires policies that coordinate high-DoF hand joints through intermittent, contact-rich interaction. Beyond target-orientation tracking, such policies must discover finger gaits that preserve object stability while adapting to geometry, anisotropy, pose, contact, and sensing changes. We propose \method, a tactile-conditioned behavior prior model for dexterous reorientation. \method distills multi-scale sphere specialists into a latent controller and lets downstream policies reuse the fixed tactile prior through residual latent actions, reducing renewed exploration from raw joint commands. The prior conditions on tactile-proprioceptive history so latent behavior reflects the current hand-object interaction. We evaluate arbitrary-pose transfer across anisotropic objects, commanded-axis rotation, and an arm-hand Grasp-to-AnyPose task in which the robot must grasp, lift, transport, and reach goal poses for novel tool geometries and generalized placements. Extensive experiments demonstrate that the proposed method accelerates training and enables stable policies where matched raw-action PPO remains near failure, with successful sim-to-real transfer in in-hand and arm-hand tasks.
Chinese Translation
灵巧的手内操作需要策略能够通过间歇性的、富接触的交互来协调高自由度的手部关节。除了目标朝向跟踪之外,此类策略还必须发现能够保持物体稳定性的手指步态,同时适应几何形状、各向异性、姿态、接触和感知的变化。我们提出了 TacBPM,一种面向灵巧重定向的触觉条件化行为先验模型。TacBPM 将多尺度球体专家策略蒸馏为一个潜在控制器,并让下游策略通过残差潜在动作复用固定的触觉先验,从而减少从原始关节指令出发的重复探索。该先验以触觉-本体感知历史为条件,使潜在行为能够反映当前的手-物体交互状态。我们在多项任务中进行了评估,包括各向异性物体的任意姿态迁移、指定轴旋转,以及一个臂-手协作的抓取到任意姿态(Grasp-to-AnyPose)任务——在该任务中,机器人必须针对新颖的工具几何形状和泛化的摆放位置完成抓取、抬起、搬运并到达目标姿态。大量实验表明,所提出的方法能够加速训练,并在匹配的原始动作 PPO 接近失败的场景中实现稳定策略,且在手内操作与臂-手协作任务中均实现了成功的从仿真到现实(sim-to-real)的迁移。
cs.RO / 38 / 2609.18191

OmniRisk: Omnidirectional Trajectory-Risk Learning for Agile Quadrotor Dynamic Avoidance

OmniRisk:面向敏捷四旋翼动态避障的全向轨迹-风险学习
He, Yifan, Liu, Yang, Zhao, Wenhao, Lin, Hai, Zhang, Deping, Ma, Mingze, Gao, Fei, Yu, Huan, Dai, Zipeng, Ding, Ziming
Abstract
Agile quadrotor avoidance of fast-moving obstacles requires anticipating collisions and selecting feasible maneuvers within short reaction windows. Reliable predictive avoidance remains challenging because sparse range observations do not directly reveal obstacle motion, while online trajectory optimizers either scale poorly with obstacle count or remain efficient at the expense of reliability in dense, high-speed encounters. We present OmniRisk, an omnidirectional planning framework that learns trajectory-level risk offline for efficient onboard evasion. A fixed-dimensional tensor combines LiDAR range panoramas, dynamic masks, and Cartesian surface velocities to represent geometry and motion jointly. We formulate an asymmetric risk field aligned with obstacle velocity that emphasizes approaching interactions and attenuates receding ones. Accumulating this risk along predicted relative trajectories provides dense supervision and discourages unnecessary hesitation after obstacles pass. A dual-branch circular convolutional network predicts terminal boundary states and dynamic risks for candidate primitives over an omnidirectional anchor lattice in a single forward pass, followed by selection and closed-form reconstruction of the selected candidate primitive. This formulation removes online risk accumulation along trajectories and makes risk-inference cost independent of obstacle count. OmniRisk enables efficient onboard avoidance, with real-world flights demonstrating consecutive evasive maneuvers at relative encounter speeds up to 15 m/s without fine-tuning. Code is available at https://github.com/VANdexj/OmniRisk.
Chinese Translation
敏捷四旋翼避让快速移动的障碍物需要在短反应窗口内预测碰撞并选择可行的机动动作。可靠的预测性避障仍然具有挑战性,因为稀疏的距离观测无法直接揭示障碍物的运动,而在线轨迹优化器要么随障碍物数量增加而扩展性变差,要么以牺牲密集、高速遭遇情形下的可靠性为代价来保持计算效率。我们提出OmniRisk,一个全向规划框架,通过离线学习轨迹级风险来实现高效的机载规避。该框架使用固定维度的张量融合LiDAR距离全景图、动态掩码以及笛卡尔坐标系下的表面速度,以联合表示障碍物的几何与运动信息。我们构建了一个与障碍物速度对齐的非对称风险场,该风险场强调接近中的交互并衰减远离中的交互。沿预测的相对轨迹累积该风险可提供密集的监督信号,并在障碍物通过后避免不必要的犹豫。一个双分支循环卷积网络在单次前向传播中,于全向锚点栅格上预测候选基元的终端边界状态与动态风险,随后进行候选基元的选择和闭式重构。该方案免除了在线沿轨迹进行的风险累积,使风险推理的开销与障碍物数量无关。OmniRisk实现了高效的机载避障,真实飞行实验表明,无需微调即可在相对遭遇速度高达15 m/s的情况下连续执行规避机动。代码可在 https://github.com/VANdexj/OmniRisk 获取。
cs.RO / 39 / 2609.18193

WAVE-Go: World-Model Navigation with Adaptive Execution for Wheel-Legged Robots

WAVE-Go:面向轮腿机器人的世界模型导航与自适应执行
Li, Mingyi, Li, Ji, Ouyang, Zhihao, He, Yage, Karlsson, Börje F.
Abstract
World models can anticipate the consequences of navigation actions, but predicted action sequences may become invalid during execution, especially when wheel-legged robots encounter dynamic obstacles or change locomotion modes. We propose WAVE-Go, an image-goal navigation framework that separates world-action prediction from interruptible command execution. Its executor adaptively selects an action prefix and cancels pending commands when updated observations invalidate execution. A conditional-risk formulation specifies prefix selection under an estimated cumulative failure budget, while posture and locomotion-mode transitions require clearance, stability, and task-evidence checks. In the reported navigation evaluation, WAVE-Go achieves 74.1% in-distribution success and 63.3% dynamic out-of-distribution success, exceeding the strongest baseline by 4.7 and 7.7 percentage points, respectively, while reducing collisions from 4.4 to 2.9 per 100 m. Compared with interruptible fixed four-command execution, WAVE-Go raises success by 4.0 percentage points while reducing replanning frequency by 51.2% and collision rate by 6.5%. Execution ablations also show that runtime interruption improves success, collision rate, and reaction latency at the cost of additional replanning. These results support adaptive, interruptible execution as a means of balancing navigation performance and planning overhead. Code is available at https://github.com/vigorlee/wave-go.
Chinese Translation
世界模型能够预判导航动作的后果,但预测的动作序列在执行过程中可能失效,尤其是当轮腿机器人遇到动态障碍物或改变运动模式时。我们提出了 WAVE-Go,一个将世界-动作预测与可中断的指令执行相分离的图像目标导航框架。其执行器能够自适应地选择动作前缀,并在更新后的观测使执行失效时取消待执行的指令。通过条件风险建模,在估计的累积失败预算约束下指定前缀选择,而姿态和运动模式的切换则需通过间隙性、稳定性和任务证据检查。在所报告的导航评估中,WAVE-Go 在分布内场景中达到 74.1% 的成功率,在动态分布外场景中达到 63.3% 的成功率,分别比最强基线高出 4.7 和 7.7 个百分点,同时将每 100 米的碰撞次数从 4.4 次降至 2.9 次。与可中断的固定四指令执行相比,WAVE-Go 将成功率提高了 4.0 个百分点,同时将重规划频率降低 51.2%、碰撞率降低 6.5%。执行消融实验还表明,运行时中断能够以额外的重规划为代价,提升成功率、降低碰撞率并改善反应延迟。这些结果支持将自适应、可中断的执行作为平衡导航性能与规划开销的一种手段。代码发布于 https://github.com/vigorlee/wave-go。
cs.RO / 40 / 2609.18197

WholeBodyWAM: Learning Whole-Body World Action Models with Scalable Motion Priors

WholeBodyWAM:利用可扩展运动先验学习全身世界动作模型
Zhang, Bowei, Zhang, Qiyao, Bai, Shuanghao, Wang, Xinhua, Li, Meng, Wang, Yilei, Zhang, Leiwang, Tang, Jian, Zhou, Lu, Sun, Lei, Che, Zhengping
Abstract
Humanoid whole-body manipulation requires coordinated whole-body dynamics, yet large-scale trajectories from a target robot are expensive to collect and difficult to scale. In contrast, whole-body motion from human and humanoid sources is abundantly available, although such data cannot be directly used as embodiment-specific robot actions. This work asks whether these scalable motion resources can instead provide a transferable predictive prior for humanoid world-action modeling. We introduce WholeBodyWAM, a humanoid world-action model that learns whole-body dynamics from large-scale heterogeneous motion before target-robot training. We curate UniMotion-4K, a motion corpus spanning more than 4K hours from human videos, native 3D motion datasets, and heterogeneous humanoid platforms, and canonicalize these diverse sources into a unified motion space. A language-conditioned Motion Expert is then pretrained to predict future whole-body motion without target-robot action supervision. During robot post-training, the pretrained Motion Expert is integrated with Video and Action Experts through asymmetric Mixture-of-Transformers (MoT) attention, enabling predictive scene dynamics and whole-body motion to jointly inform embodiment-specific action generation. Experiments show that WholeBodyWAM consistently benefits from increased motion-pretraining scale, improves future-motion prediction and downstream task performance, and transfers effectively to real-world humanoid manipulation. Moreover, the pretrained motion prior substantially improves data efficiency under limited target-robot demonstrations.
Chinese Translation
人形机器人全身操作需要协调的全身动力学,然而从目标机器人收集大规模轨迹数据成本高昂且难以扩展。相比之下,来自人类和人形机器人来源的全身运动数据十分丰富,但此类数据无法直接用作特定本体(embodiment)的机器人动作。本工作探讨这些可扩展的运动资源能否转而为构建人形机器人的世界-动作建模提供可迁移的预测先验。我们提出 WholeBodyWAM,这是一个在人形机器人的世界-动作模型,它先从大规模异构运动数据中学习全身动力学,再进行目标机器人训练。我们构建了 UniMotion-4K 运动语料库,涵盖来自人类视频、原生 3D 运动数据集以及异构人形机器人平台的超过 4000 小时数据,并将这些多样化的数据源规范化到统一的运动空间中。随后,我们预训练一个以语言为条件的运动专家(Motion Expert),使其在无需目标机器人动作监督的情况下预测未来的全身运动。在机器人后训练阶段,预训练的运动专家通过非对称的 Mixture-of-Transformers(MoT)注意力机制与视频专家和动作专家集成,使预测性场景动力学与全身运动共同指导特定本体的动作生成。实验表明,WholeBodyWAM 持续受益于运动预训练规模的扩大,提升了未来运动预测能力和下游任务性能,并能有效迁移到真实世界的人形机器人操作任务。此外,预训练的运动先验在目标机器人演示数据有限的情况下显著提升了数据效率。
cs.RO / 41 / 2609.18207

Reinforcement Learning for Real-Time Vision-Language-Action Policies

面向实时视觉-语言-动作策略的强化学习
Dong, Perry, Hung, Kuo-Han, Sadigh, Dorsa, Finn, Chelsea
Abstract
Reinforcement learning fine-tuning on top of large, pretrained Vision-Language-Action (VLA) models offers promise for highly reliable robot deployment. However, because of their scale, modern VLA models suffer from high inference latency, so the observation used to select an action is often stale by execution time, creating a distribution shift that can substantially degrade reliability and performance. Prior work has explored asynchronous policy execution to reduce the effect of latency, but these methods are mostly built on imitation learning and offer no mechanism for moving beyond the training distribution toward higher reliability. We close this gap by enabling RL fine-tuning that meets the real-time control requirements of dynamic real-world manipulation. Our approach builds on EXPO-FT, a framework for sample-efficient, reliable VLA fine-tuning with reinforcement learning, and decouples slow, expressive action generation from fast, reactive action edits: a large pretrained VLA proposes action chunks using its strong behavior prior, while a lightweight edit policy performs fast, reactive decision-making by editing actions in response to changes in state, conditioned on the latest observation. We instantiate this as Real-Time EXPO-FT, an RL framework for finetuning real-time VLA policies. On the Kinetix benchmark, Real-Time EXPO-FT enables a delayed policy to achieve the best performance among delayed and non-delayed methods in 10 out of 10 environments. On four dynamic real-world tasks, robot object passing, ball balancing, table soccer kicking, and dynamic object picking, with online robot data capped at 10 minutes, Real-Time EXPO-FT improves average policy performance from 42% to 97%, all without human intervention, demonstrating rapid, sample-efficient adaptation to challenging real-world dynamics. Website: https://pd-perry.github.io/real-time-expo-ft
Chinese Translation
在大规模预训练视觉-语言-动作(Vision-Language-Action,VLA)模型之上进行强化学习微调,为高度可靠的机器人部署带来了希望。然而,由于其模型规模庞大,现代VLA模型的推理延迟较高,导致用于选择动作的观测信息在执行时往往已经过时,由此产生的分布偏移会显著降低可靠性和性能。已有研究探索了异步策略执行来降低延迟的影响,但这些方法大多基于模仿学习,缺乏超越训练分布以提升可靠性的机制。我们通过使强化学习微调能够满足动态真实世界操作的实时控制需求来填补这一空白。我们的方法建立在EXPO-FT——一个样本高效、可靠的VLA强化学习微调框架——之上,并将缓慢而富有表达力的动作生成与快速的反应式动作编辑解耦:大型预训练VLA模型凭借其强大的行为先验提出动作块,而一个轻量级的编辑策略则以最新观测为条件,通过响应状态变化来编辑动作,从而实现快速的反应式决策。我们将其实现为Real-Time EXPO-FT,一个用于微调实时VLA策略的强化学习框架。在Kinetix基准测试中,Real-Time EXPO-FT使带延迟的策略在全部10个环境中的10个里取得了优于延迟与非延迟方法的最佳性能。在四项动态真实世界任务——机器人传递物体、球体平衡、桌上足球踢球和动态物体抓取——中,在在线机器人数据限制为10分钟的条件下,Real-Time EXPO-FT将平均策略性能从42%提升至97%,且全程无需人工干预,展示了对具有挑战性的真实世界动态的快速、样本高效的适应能力。网站:https://pd-perry.github.io/real-time-expo-ft
cs.RO / 42 / 2609.18213

CANTABILE: Learning Expressive Dynamics for Robotic Piano Performance

CANTABILE:面向机器人钢琴演奏表现性力度的学习
Kim, Woosik, Choi, Wonhyeok, Im, Sunghoon
Abstract
Robotic piano playing has emerged as a standard benchmark for dexterous bimanual manipulation, yet progress on it has been measured almost entirely by note accuracy -- which keys are pressed (pitch) and when (onset) -- leaving the musical dynamics essential for expressive performance neither rewarded nor evaluated. We propose CANTABILE, a dynamics-aware framework for robotic piano performance that (i) closes the score-to-contact loop by conditioning the policy on upcoming velocity goals and mapping each key's angular velocity at onset back to MIDI velocity, (ii) couples a velocity-fidelity reward with an onset-coverage reward, so that dynamics cannot be improved by omitting difficult notes, and (iii) refines a frozen dynamics-aware base policy with an alpha-scaled, finger-only residual that localizes strike-intensity adaptation away from nominal note execution. On EXPRESSIVE-51, a dynamics-rich 51-song subset of RoboPianist, CANTABILE raises Velocity F1 -- jointly measuring pitch, onset, and intensity within a +/-8 MIDI-velocity tolerance -- from 0.06 to 0.34 over the RoboPianist baseline, improves all 51 songs, more than halves matched-note velocity error, and reduces log-mel distance to reference audio by 8%. Intensity-randomized training further enables runtime control of performance intensity without retraining.
Chinese Translation
机器人钢琴演奏已成为灵巧双手操作的标准测试基准,然而其进展几乎完全以音符准确率来衡量——即按下哪些琴键(音高)以及何时按下(时间点)——而对表现性演奏至关重要的音乐力度既未被奖励也未被评估。我们提出 CANTABILE,一个力度感知的机器人钢琴演奏框架,其特点是:(i)通过将策略条件化于即将到来的速度目标,并将每个琴键在按下瞬间的角速度映射回 MIDI 力度值,从而闭合“乐谱到接触”的循环;(ii)将力度保真度奖励与音符覆盖奖励相结合,使得无法通过省略困难音符来提升力度表现;(iii)在冻结的力度感知基础策略之上,采用经过 alpha 缩放的、仅作用于手指的残差策略进行微调,将击键强度的自适应与名义音符执行相分离。在 EXPRESSIVE-51(RoboPianist 数据集中包含丰富力度信息的 51 首乐曲子集)上,CANTABILE 将 Velocity F1——在 +/-8 个 MIDI 力度容差范围内联合衡量音高、时间点和强度的指标——相较 RoboPianist 基线从 0.06 提升至 0.34,在全部 51 首乐曲上均有改进,将匹配音符的力度误差降低一半以上,并将与参考音频的对数梅尔距离降低 8%。基于随机强度扰动的训练进一步使系统能够在运行时控制演奏强度而无需重新训练。
cs.RO / 43 / 2609.18232

UMI-Bridge: Action-Anchored Latent Alignment across Human and Robot Manipulation Data

UMI-Bridge:基于动作锚定的人类与机器人操作数据间潜在表征对齐
Liu, Haiyi, Ma, Jingming, Rui, Ke, Wei, Yuteng, Ma, Yuan, Zuo, Yushen, Tian, Honglong, Jia, Haoran, Zhou, Weitao, Wang, Jiawei, Li, Minglei, Chen, Shiyi, Mao, Haiyan, Zhang, Jiaqi, Zhang, Chun
Abstract
Real-robot demonstrations are limited, motivating the use of human manipulation data collected without robots, including egocentric videos and handheld Universal Manipulation Interface (UMI) demonstrations. However, differences in viewpoint, embodiment, and available action supervision make it difficult to align representations across these sources according to manipulation motion rather than visual appearance. We introduce UMI-Bridge, which uses UMI as an intermediate domain to align representations according to action equivalence rather than pixel similarity. UMI action supervision anchors the latent representation to end-effector motion and gripper behavior, while synchronized head-wrist observations and paired ego-UMI clips support alignment across views and domains. We train a dual-view latent action model (LAM) on human manipulation data without robot demonstrations, then freeze its wrist teacher and dynamics model to regularize vision-language-action (VLA) post-training on UMI and robot data. The shared wrist interface enables this training-time supervision across both domains while preserving the policy's standard inference architecture. Across three real-robot tasks, UMI-Bridge achieves 91.7% mean success versus 73.3% for Naive Co-training with matched UMI and robot data. On two data-efficiency tasks, it surpasses a full-data Robot-only baseline using 25% of the robot demonstrations together with UMI data. It also achieves 85% and 90% success on two additional tasks learned from UMI demonstrations without task-specific robot demonstrations. These results support action-anchored latent alignment for data-efficient robot learning and UMI-to-robot task transfer.
Chinese Translation
真实机器人演示数据有限,因此人们开始利用无需机器人即可采集的人类操作数据,包括第一人称视角(egocentric)视频和手持式通用操作接口(Universal Manipulation Interface, UMI)演示。然而,视角、具身形态以及可用动作监督的差异,使得难以依据操作动作而非视觉外观来对齐这些数据源之间的表征。我们提出UMI-Bridge,它以UMI作为中间域,根据动作等价性而非像素相似性来对齐表征。UMI的动作监督将潜在表征锚定到末端执行器运动和夹爪行为上,同时同步的头戴-腕部观测以及配对的第一人称-UMI视频片段支持跨视角和跨域的对齐。我们在无需机器人演示的人类操作数据上训练一个双视角潜在动作模型(Latent Action Model, LAM),然后冻结其腕部教师模型和动力学模型,以正则化在UMI和机器人数据上的视觉-语言-动作(VLA)后训练。共享的腕部接口使得这种训练时监督能够同时适用于两个域,同时保持策略的标准推理架构。在三项真实机器人任务上,UMI-Bridge在使用等量UMI和机器人数据的情况下达到91.7%的平均成功率,而朴素协同训练(Naive Co-training)仅为73.3%。在两个数据效率任务上,该方法仅使用25%的机器人演示并结合UMI数据,即超越了使用全量数据的纯机器人基线。此外,在两个仅由UMI演示学习、无任务特定机器人演示的额外任务上,该方法分别达到85%和90%的成功率。这些结果验证了基于动作锚定的潜在表征对齐对于数据高效的机器人学习以及UMI到机器人的任务迁移的有效性。
cs.RO / 44 / 2609.18242

ForceDelta-VLA: Distilling Force-Conditioned ActionCorrections for Contact-Rich Manipulation

ForceDelta-VLA:面向接触密集型操作的力条件动作校正蒸馏方法
Dong, Ju, Fu, Yu, Chen, Jian, Liu, Yimeng, Zhao, Haocheng, Zhang, Lei, Bai, Kaixin, Zhang, Liding, Zheng, Diwen, Knoll, Alois Christian, Schoellig, Angela P., Zhang, Jianwei
Abstract
Force-aware Vision-Language-Action (VLA) policies improve contact-rich manipulation, but typically combine task-level motion and contact-dependent adjustment in a single action prediction. Demonstrations provide no explicit labels for decomposing that prediction into a reusable reference action and a correction. We present ForceDelta-VLA, a correction-distillation framework that constructs an explicit force-correction target using paired predictions from a frozen teacher's force-conditioned and learned force-agnostic modes. A separate delay-correction target accounts for reference-action mismatch and the change in reference state. Training uses asynchronous schedule replay with the cached task context available during execution. The resulting lightweight policy adjusts the reference actions using recent force history and robot state, responding to contact changes between reference-action updates without regenerating complete action chunks. Across nine single-arm and bimanual contact-rich tasks, ForceDelta-VLA achieves an 82.2% mean success rate, compared with 54.4% for the original ForceVLA baseline. Direct execution of our Stage-1 Temporal Teacher achieves 70.6%. Relative to ForceVLA, the complete system reduces mean peak contact force over successful trials by approximately 26% on both platforms.
Chinese Translation
力感知的视觉-语言-动作(Vision-Language-Action, VLA)策略能够改善接触密集型操作,但通常在单一动作预测中同时结合任务级运动和依赖接触的调整。演示数据中并没有显式标签将该预测分解为可复用的参考动作和校正量。我们提出 ForceDelta-VLA,这是一个校正蒸馏框架,利用冻结教师模型的力条件模式和习得的力无关模式的成对预测,构建显式的力校正目标。此外,单独的延迟校正目标用于处理参考动作的不匹配以及参考状态的变化。训练采用异步调度回放,并使用执行期间缓存的上下文信息。所得到的轻量级策略利用最近的力历史和机器人状态对参考动作进行调整,能够在参考动作更新之间响应接触变化,而无需重新生成完整的动作块。在九项单臂和双臂接触密集型任务上,ForceDelta-VLA 的平均成功率达到 82.2%,而原始基线 ForceVLA 仅为 54.4%。直接执行我们的第一阶段时序教师模型(Temporal Teacher)可达到 70.6%。与 ForceVLA 相比,完整系统在两个平台上均将成功试验中的平均峰值接触力降低约 26%。
cs.RO / 45 / 2609.18243

Acting in Meters: Learning Metric Interactions for Precise Robotic Manipulation

以米为单位进行操作:学习度量交互以实现精确的机器人操控
Wang, Lijie, Lu, Zheng, Wang, Yiming, Yu, Heyang, Hoi, Kenghou, Hu, Bowen, Cui, Di, Xin, Tianyu, Liao, Haoran, Zhong, Wanqi, Fan, Xingjie, Xu, Yizhao, Wang, Ziliang, Gao, Fei, Li, Yiming
Abstract
Vision-Language-Action models and World-Action Models have advanced language-conditioned robotic manipulation, yet often leave metric relations among actions, objects, and scene geometry implicit. Human manipulation combines semantic understanding of task-relevant objects with spatial feedback that guides hand motion relative to objects and their surroundings. Inspired by this, we introduce a metric interaction framework that models object-level and scene-level interactions in physical Cartesian space at a shared metric scale. At the object level, Interaction-Centric Tokens (ICTs) explicitly represent end-effector pose trajectories relative to manipulated objects and are jointly denoised with actions, providing physically grounded interaction supervision. At the scene level, the Metric Action Interaction Field (MAIF) uses action and ICT queries to attend to metric scene point-cloud features and learns geometry-conditioned action corrections. Through two-stage adaptation, our framework improves diverse VLA and WAM baselines with a small number of additional parameters and training steps. Experiments demonstrate average success-rate gains of 0.80 and 3.59 percentage points on LIBERO and RoboTwin~2.0, respectively, alongside gains of 6.80 percentage points on real-world tasks and 7.45 percentage points on their out-of-distribution variants.
Chinese Translation
视觉-语言-动作模型和世界-动作模型推动了语言条件下的机器人操控的发展,但它们往往将动作、物体和场景几何之间的度量关系隐式化。人类的操控将针对任务相关物体的语义理解与引导手部相对于物体及其周围环境运动的空间反馈相结合。受此启发,我们提出了一个度量交互框架,该框架在共享的度量尺度下于物理笛卡尔空间中建模物体级和场景级交互。在物体层面,以交互为中心的标记显式地表示相对于被操作物体的末端执行器位姿轨迹,并与动作进行联合去噪,从而提供具有物理依据的交互监督。在场景层面,度量动作交互场利用动作和ICT查询来关注度量的场景点云特征,并学习以几何为条件的动作修正。通过两阶段适配,我们的框架仅需少量额外参数和训练步骤即可提升多种VLA和WAM基线模型的性能。实验表明,该方法在LIBERO和RoboTwin 2.0上的平均成功率分别提升了0.80和3.59个百分点,在真实世界任务上提升了6.80个百分点,在其分布外变体任务上提升了7.45个百分点。
cs.RO / 46 / 2609.18245

A3P5 NEMESIS Integrated Rover Design for Environmental Reconnaissance and Robotic Sampling with Reproducible Mobility Analysis and an External Data Machine Learning Calibration Benchmark

A3P5 NEMESIS综合漫游车设计:用于环境侦察与机器人采样,包含可复现的移动性分析及外部数据机器学习校准基准
Sultan, Shafi Bin, Sultan, Sabik Bin, Sadad, Safwan
Abstract
A3P5 NEMESIS is a four-wheel rover intended to combine remote inspection, environmental observation and lightweight manipulation within one serviceable platform. This study develops a photo-constrained geometric reconstruction, a subsystem architecture and a reproducible analytical assessment while distinguishing physical prototype evidence from proposed functions. An exploratory search retrieved 5,000 bibliographic records across ten queries, yielding 4,897 distinct DOI records and 1,212 metadata candidates; selected primary studies and technical documents informed the design. The reconstructed configuration retains the carbon-pattern enclosure, independently steered wheel assemblies, folded manipulator, inclined camera mast and side sampling equipment. A declared 24 kg scenario predicts 3.28 newton-metres of gearbox-output torque per wheel on a 20-degree grade under equal load sharing; a separate static model shows how a 2 kg forward payload reduces the geometric front-tipping bound from 38.1 degrees to 32.7 degrees. These are design screens, not measured operating limits. A public-data calibration benchmark uses 7,344 eligible hourly observations, eight sensor/environmental predictors and chronological training, validation and test partitions. Validation-selected ridge regression achieves a held-out CO root-mean-square error of 0.502 milligrams per cubic metre, with a 95% daily-block bootstrap interval of 0.435-0.569 milligrams per cubic metre. This result concerns an external sensor array and cannot establish NEMESIS accuracy. The combined analysis identifies priority measurements, proposed control interfaces and mission-specific validation requirements. The contribution is a traceable engineering design study and evaluation framework for a prototype whose integrated field performance remains to be established.
Chinese Translation
A3P5 NEMESIS是一种四轮漫游车,旨在将远程检查、环境观测和轻量级操作集成于一个可维护的平台之中。本研究开发了一种基于照片约束的几何重建方法、一种子系统架构以及可复现的分析评估,同时将物理原型证据与拟议功能加以区分。一项探索性检索在十项查询中获取了5000条文献记录,得到4897条不同的DOI记录和1212条元数据候选记录;所选的主要研究和技术文档为设计提供了依据。重建后的构型保留了碳纤维纹路外壳、独立转向轮组件、折叠式机械臂、倾斜相机桅杆和侧面采样设备。在等载荷分配条件下,标称24 kg场景预测在20度坡度上每轮减速器输出扭矩为3.28牛顿·米;另一独立静态模型表明,2 kg的前向载荷会使几何前倾失稳界限从38.1度降至32.7度。这些结果是设计筛选指标,而非实测运行极限。一项基于公开数据的校准基准采用了7344条符合条件的逐时观测数据、八个传感器/环境预测变量,并按时间顺序划分训练集、验证集和测试集。经验证集筛选的岭回归(ridge regression)在留出集上实现了0.502毫克/立方米的CO均方根误差,其95%逐日块自助(bootstrap)置信区间为0.435-0.569毫克/立方米。该结果针对的是外部传感器阵列,无法确立NEMESIS自身的精度。综合分析确定了优先测量项、拟议的控制接口以及任务专属的验证要求。本文的贡献在于为一台综合现场性能尚待验证的原型机提供了可追溯的工程设计研究与评估框架。
cs.RO / 47 / 2609.18259

${M}^2$Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models

${M}^2$Tok:面向视觉-语言-动作模型的多头多码本离散动作分词器
Xu, Chunpu, Liang, Zhixuan, Zhang, Yuhao, Chan, Chi-Min, Wang, Jessie, Xiao, Yang, Hu, Mengkang, Yang, Xiaokang, Mu, Yao
Abstract
Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions. Since raw action signals are continuous, effective tokenization is essential to map high-dimensional inputs into compact discrete tokens for autoregressive processing. However, existing discrete action tokenizers often suffer from high reconstruction loss, failing to preserve the fine-grained dynamics required for precise control. This ``discretization bottleneck'' significantly limits the performance ceiling of downstream Vision-Language-Action (VLA) models. To address this, we propose $\mathcal{M}^2$Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance. Our approach introduces two key structural innovations: (1) we decompose the latent action features into multiple heads, enabling the model to implicitly align specific heads with distinct action dimensions; (2) we assign independent codebooks to each head for quantization. By leveraging the combinatorial nature of multiple codebooks, we significantly expand the representational expressivity of the tokenizer, leading to substantially lower reconstruction loss compared to previous methods. We evaluate the $\mathcal{M}^2$Tok-based VLA on the RoboTwin, Simpler-Env, and 3 zero-shot real-world tasks. Experimental results demonstrate our method not only achieves superior reconstruction fidelity but also significantly boosts the success rate of VLA models. Comprehensive ablation studies further confirm the effectiveness of the multi-head and multi-codebook mechanisms. Code is available at \href{https://github.com/cpaaax/M2Tok}{https://github.com/cpaaax/M2Tok}.
Chinese Translation
近年来的研究进展已成功使自回归语言模型能够处理图像和动作等多模态信号。由于原始动作信号是连续的,因此需要有效的分词化方法,将高维输入映射为紧凑的离散词元以进行自回归处理。然而,现有的离散动作分词器往往存在较高的重建损失,无法保留精确控制所需的细粒度动态信息。这一“离散化瓶颈”显著限制了下游视觉-语言-动作(VLA)模型的性能上限。为解决这一问题,我们提出了 $\mathcal{M}^2$Tok,一种旨在最小化重建误差并提升策略性能的多头多码本动作分词器。我们的方法引入了两项关键结构创新:(1)将潜在动作特征分解为多个头,使模型能够隐式地将特定头与不同的动作维度对齐;(2)为每个头分配独立的码本进行量化。通过利用多个码本的组合特性,我们显著扩展了分词器的表示能力,使重建损失相比以往方法大幅降低。我们在 RoboTwin、Simpler-Env 以及三个零样本真实世界任务上对基于 $\mathcal{M}^2$Tok 的 VLA 模型进行了评估。实验结果表明,我们的方法不仅实现了更优的重建保真度,还显著提升了 VLA 模型的成功率。全面的消融研究进一步验证了多头与多码本机制的有效性。代码已发布于 \href{https://github.com/cpaaax/M2Tok}{https://github.com/cpaaax/M2Tok}。
cs.RO / 48 / 2609.18293

Function-Preserving Data Generation for Zero-Shot Real-to-Sim-to-Real Manipulation

面向零样本Real-to-Sim-to-Real操作的保函数数据生成
Xiang, Tianyi, Xie, Xupeng, Cao, Jiahang, Luo, Andrew F., Li, Haoang, Ma, Jun
Abstract
Robotic data generation is a promising paradigm for scaling robot learning without collecting large-scale real-world data. However, generating geometrically diverse yet physically valid data for contact-rich tasks remains challenging, especially when success depends on precise geometric interfaces. Standard shape augmentation methods often distort task-critical interfaces, resulting in invalid contact relationships, e.g., fit mismatches or interpenetration, rendering downstream interactions infeasible. To address these limitations, we propose a function-preserving Real-to-Sim-to-Real framework that generates synthetic demonstrations from reconstructed assets without teleoperated source trajectories. Our method augments task-relevant object geometries through constraint-guided mesh deformation, together with physically consistent transfer of task poses and collision proxies. Visual domain randomization is further applied during simulation rollouts, enabling robust zero-shot policy deployment without real-world fine-tuning. Extensive experiments in both real-world and simulation settings demonstrate that our method enables robust generalization across unseen object geometries and diverse visual conditions in contact-rich and long-horizon tasks. Our method provides a practical path toward scalable robot learning for contact-rich tasks via shape deformation.
Chinese Translation
机器人数据生成是一种无需收集大规模真实世界数据即可扩展机器人学习的有前景范式。然而,为接触密集型任务生成几何多样且物理有效的数据仍然具有挑战性,尤其是当任务成功依赖于精确的几何接口时。标准的形状增强方法往往会破坏任务关键接口,导致无效的接触关系(例如配合失配或相互穿透),使下游交互不可行。为解决这些局限,我们提出了一种保函数的Real-to-Sim-to-Real框架,能够从重建资产中生成合成示范,而无需遥操作源轨迹。我们的方法通过约束引导的网格变形增强任务相关物体的几何形状,并结合物理一致的任务位姿与碰撞代理迁移。此外,在仿真回放过程中进一步应用视觉域随机化,从而实现无需真实世界微调的鲁棒零样本策略部署。在真实世界和仿真环境中的大量实验表明,我们的方法能够在接触密集且长时序的任务中,对未见过的物体几何形状和多样的视觉条件实现鲁棒泛化。我们的方法为通过形状变形实现接触密集型任务的可扩展机器人学习提供了一条实用途径。
cs.RO / 49 / 2609.18324

RAFAIL: Relationship-Aware Failure Detection for Robotic Manipulation

RAFAIL:面向机器人操作的关系感知失败检测
Schneider, Loris, Welte, Edgar, Rayyes, Rania
Abstract
Detecting failures during execution is essential for reliable robotic manipulation. Vision-language models (VLMs) can assess task outcomes semantically but add runtime computation, whereas out-of-distribution (OOD) detectors may respond to harmless scene variations rather than failure-relevant deviations. We introduce RAFAIL, a framework for detecting execution failures during robotic manipulation. RAFAIL identifies failures by detecting anomalies in task-relevant relationships between entities, such as a gripper and an object or an object and its target. By focusing OOD detection on relevant parts of the observation, RAFAIL reduces sensitivity to task-irrelevant scene variation. Offline, a VLM annotates successful demonstrations with task progress and relationship importance, which are used to learn point-cloud-based relationship representations without relying on policy-internal features. At runtime, relationship-specific OOD detectors evaluate these representations while relationship importance and task progress are predicted without VLM inference. RAFAIL requires no failure data and achieves 73.4% balanced accuracy across three real-world robotic manipulation tasks, outperforming the strongest evaluated OOD- and uncertainty-based baselines.
Chinese Translation
在执行过程中检测失败对于实现可靠的机器人操作至关重要。视觉-语言模型(VLM)能够从语义层面评估任务结果,但会增加运行时计算开销;而分布外(OOD)检测器可能对无害的场景变化做出响应,而非与失败相关的偏差。我们提出了RAFAIL,一个用于检测机器人操作执行失败的框架。RAFAIL通过检测实体之间任务相关关系中的异常来识别失败,例如夹爪与物体之间、或物体与其目标之间的关系。通过将OOD检测聚焦于观测中相关的部分,RAFAIL降低了对任务无关场景变化的敏感性。在离线阶段,VLM为成功演示标注任务进度和关系重要性,用于学习基于点云的关系表示,而无需依赖策略内部特征。在运行时,针对特定关系的OOD检测器对这些表示进行评估,同时无需VLM推理即可预测关系重要性和任务进度。RAFAIL不需要失败数据,在三个真实世界机器人操作任务上实现了73.4%的平衡准确率,优于所评估的最强OOD基线和基于不确定性的基线方法。
cs.RO / 50 / 2609.18326

UAVs Meet Embodied Intelligence: Bridging Human Intents and Flying Dynamics Via Harnessing Physical-Digital AI Agents

无人机与具身智能的相遇:通过驾驭物理-数字智能体连接人类意图与飞行动力学
Tian, Yonglin, Wang, Weiyi, Lu, Houhua, Li, Xinyi, Wu, Yihao, Chen, Jingyang, Sun, Jianli, Li, Chengxiang, Chen, Yinuo, Lin, Fei, Zhang, Tengchao, Yang, Jing, Ji, Deyi, Di, Jian, Wu, Naiqi, Lv, Yisheng
Abstract
Unmanned aerial vehicles (UAVs) extend embodied intelligence into continuous three-dimensional space, where perception, reasoning, physical embodiment, and action are tightly coupled through flight and environmental interaction. Recent advances in foundation models, world models, and AI agents are shifting UAV autonomy from task-specific perception and control toward systems that can interpret human intent, understand open environments, reason about physical consequences, and organize complex behaviors under embodiment and flight-dynamic constraints. We characterize this emerging paradigm as UAV embodied intelligence (UAV EI) and distinguish it from its system realization, the embodied-intelligent UAV (EI UAV). To provide a unified view of the field, we introduce a 5+5 framework that describes UAV EI through five capability dimensions and EI UAVs through five architectural layers spanning physical embodiment, general cognition, embodied skills, external interaction, and system harnessing. Based on this framework, we systematically review recent progress in embodied morphology, embodied perception, world models, embodied planning, vision-language navigation, embodied manipulation, and embodied collaboration. We further identify long-horizon autonomy, predictive physical reasoning, test-time skill acquisition, and autonomous capability evolution as key challenges toward more general aerial embodied intelligence. Finally, we argue that harnessing physical-digital AI agents, through persistent coupling of digital intelligence with physical sensing, dynamics, action, and feedback, provides a system-level pathway toward adaptive and continuously evolving UAV autonomy. Project resources are available at our project website and GitHub repository.
Chinese Translation
无人驾驶飞行器(UAV)将具身智能拓展至连续的三维空间,其中感知、推理、物理载体与行动通过飞行和环境交互紧密耦合。基础模型、世界模型和AI智能体(AI agents)的最新进展,正在推动无人机自主性从特定任务的感知与控制,转向能够理解人类意图、认知开放环境、推理物理后果,并在具身与飞行动力学约束下组织复杂行为的系统。我们将这一新兴范式刻画为无人机具身智能(UAV EI),并将其与其系统实现形式——具身智能无人机(EI UAV)相区分。为提供该领域的统一视角,我们提出了一个5+5框架,通过五个能力维度描述UAV EI,并通过涵盖物理载体、通用认知、具身技能、外部交互和系统驾驭五个架构层描述EI UAV。基于该框架,我们系统综述了具身形态、具身感知、世界模型、具身规划、视觉语言导航、具身操作和具身协作方面的最新进展。我们进一步指出,长时程自主性、预测性物理推理、测试时技能获取和自主能力演化,是实现更通用空中具身智能的关键挑战。最后,我们认为,通过数字智能与物理感知、动力学、行动和反馈的持续耦合,驾驭物理-数字AI智能体为实现自适应且持续演化的无人机自主性提供了一条系统级路径。项目资源可在我们的项目网站和GitHub代码库中获取。
cs.RO / 51 / 2609.18358

GraphPoint: Semantic Entity Graphs and Point Trajectories for Compositional Robot Manipulation

GraphPoint:用于组合式机器人操作的语义实体图与点轨迹
Luo, Kang, Wang, Hesheng
Abstract
Robot manipulation policies often struggle to generalize beyond their demonstrations, even when new instructions involve familiar objects and behaviors. When language and scenes are strongly correlated during training, a policy can learn a fixed visual-action mapping rather than respond to the requested behavior. We investigate compositional reuse at two levels: within a subtask, combining familiar entities, action types, and action modifiers; and across subtasks, reusing learned subtasks in unseen long-horizon tasks. We introduce CoMani, a benchmark with controlled splits for evaluating both capabilities. Matched initial scenes and controlled changes to a single semantic factor encourage reliance on language rather than visual shortcuts. We further propose GraphPoint, which connects semantic entity graphs to geometric control by predicting future gripper point trajectories and converting them into actions using robot geometry. The framework organizes the gripper and objects by semantic roles and conditions their interactions on action types and modifiers, while predicted progress guides transitions during execution. Experiments and ablations on CoMani validate the effectiveness of our method for instruction-dependent generalization at both levels. Code will be released at GraphPoint.
Chinese Translation
机器人操作策略往往难以泛化到演示之外的场景,即使新指令涉及熟悉的物体和行为也是如此。当训练中语言与场景高度相关时,策略可能学习到固定的视觉-动作映射,而非响应所请求的行为。我们在两个层面研究组合式复用:在子任务内部,组合熟悉的实体、动作类型和动作修饰符;在子任务之间,将已学习的子任务复用于未见过的长时程任务。我们提出了CoMani基准,通过受控的数据划分来评估这两种能力。匹配的初始场景以及对单一语义因素的可控变化促使策略依赖语言而非视觉捷径。我们进一步提出GraphPoint,通过预测未来夹爪点轨迹并将其利用机器人几何信息转换为动作,将语义实体图与几何控制联系起来。该框架按语义角色组织夹爪和物体,并以动作类型和修饰符为条件约束它们之间的交互,同时预测的进度在执行过程中引导状态转换。在CoMani上的实验和消融研究验证了我们的方法在上述两个层面实现依赖指令泛化的有效性。代码将发布于GraphPoint。
cs.RO / 52 / 2609.18359

RecMorph: Topology-Guided Spatial Recurrence for Generalized Morphology Control

RecMorph:面向广义形态控制的拓扑引导空间递归方法
Rao, Quanrui, Liu, Yong, Xiao, Xueming, Luo, Yingbo, Wu, Kun, Xu, Zhenyu, Yao, Meibao
Abstract
Generalized morphology control requires a single policy to transform information across limbs with different physical roles, coordinate whole-body motion, and remain efficient as body size grows. Existing communication mechanisms address these requirements only partially. We introduce RecMorph, a topology-guided spatial recurrent architecture that uses recurrent sequence computation to jointly perform cross-limb communication and representation transformation. A depth-first traversal converts the kinematic tree into a morphology-derived sequence, along which shared bidirectional transitions progressively transform limb information before action decoding. Residual preservation, RMS normalization, and input-dependent channel modulation stabilize this repeated spatial transformation, yielding linear token complexity at fixed model width and depth. Across five UNIMAL tasks, RecMorph achieves the strongest mean final training performance among the evaluated generalized morphology controllers and the highest measured inference throughput on FT, while generalizing to unseen variations and bodies with up to 30 limbs. We further migrate representative generalized controllers from UNIMAL benchmarks to a four-platform quadruped setting. RecMorph achieves the best macro-averaged performance under nominal and high friction, reduces nominal velocity RMSE by 43.5% relative to specialist MLPs, and one shared policy completes 40 physical Go1/Go2 trials without falls. These results show that topology-guided recurrent transformation provides an effective and efficient communication mechanism for Generalized Morphology Control and remains effective when transferred from procedural bodies to physical robot platforms. Code and experimental resources are publicly available at https://github.com/quanruirao/RecMorph.
Chinese Translation
广义形态控制要求单一策略能够在具有不同物理角色的肢体之间传递信息、协调全身运动,并随着体型增长保持高效。现有的通信机制仅能部分满足这些要求。我们提出RecMorph,一种拓扑引导的空间递归架构,利用递归序列计算联合实现跨肢体通信与表示变换。深度优先遍历将运动学树转换为源自形态的序列,共享的双向转移沿该序列在动作解码前逐步变换肢体信息。残差保持、RMS归一化以及依赖输入的通道调制稳定了这种重复的空间变换,在固定模型宽度和深度下实现了线性token复杂度。在五个UNIMAL任务中,RecMorph在所评估的广义形态控制器中取得了最高的平均最终训练性能,并在FT任务上达到最高的实测推理吞吐量,同时能够泛化到未见过的形态变体以及多达30个肢体的身体。我们进一步将代表性的广义控制器从UNIMAL基准迁移到四平台四足机器人设置。RecMorph在标称和高摩擦条件下取得了最佳的宏平均性能,相对于专用MLP将标称速度RMSE降低了43.5%,并且单一共享策略在40次物理Go1/Go2实验中全部完成且未发生跌倒。这些结果表明,拓扑引导的递归变换为广义形态控制提供了一种有效且高效的通信机制,并且在从程序化生成的身体迁移到物理机器人平台时仍然有效。代码和实验资源已在 https://github.com/quanruirao/RecMorph 公开。
cs.RO / 53 / 2609.18374

Decoupling Vision, Language, and Action for Efficient Multi-Task Robot Policies

解耦视觉、语言与动作以实现高效的多任务机器人策略
Sun, Xiatao, Liang, Chen, Zeng, Ziyao, Wang, Qian, Zhang, Haoyang, Sun, Yue, Li, Qiucheng, Rakita, Daniel
Abstract
Vision-Language-Action (VLA) models attach an action module to a Vision-Language Model (VLM) with billions of parameters and pay for that backbone at every control step. For a low-level manipulation policy, this cost may be unnecessary: the VLM supplies vision and language embeddings, and recent standalone vision encoders and encoder-only language models now match or exceed large VLMs on visual embedding and language understanding benchmarks. We study this question with a controlled experiment. Holding the demonstrations, the training budget, the tasks, and the measurement platform fixed, we vary the vision encoder, the language encoder, and the action head of a decoupled policy and compare against seven VLA baselines. The study yields the Decoupled Embodiment Model (DEM), which pairs a fine-tuned DINOv3 encoder and a frozen NeoBERT encoder with a MeanFlow head that generates each action chunk in a single forward pass. On 18 simulated manipulation tasks with held-out language paraphrases and randomized scenes, and on three real-robot tasks, DEM achieves observed success comparable to state-of-the-art VLM-backbone policies under our evaluation protocol, while running at eight to seventeen times their inference frequency and drawing six to fifteen times less energy per inference. Within this task scope, modern decoupled components offer a better success--latency--energy trade-off.
Chinese Translation
视觉-语言-动作(VLA)模型在拥有数十亿参数的视觉-语言模型(VLM)上附加一个动作模块,并在每个控制步骤中为该主干网络付出计算代价。对于低层操作策略而言,这一代价可能是不必要的:VLM 提供视觉和语言嵌入,而近期的独立视觉编码器和仅编码器语言模型在视觉嵌入和语言理解基准上已经达到甚至超越大型 VLM。我们通过一项受控实验来研究这一问题。在固定演示数据、训练预算、任务和测量平台的前提下,我们改变解耦策略的视觉编码器、语言编码器和动作头,并与七个 VLA 基线进行比较。该研究产出了去耦合具身模型(Decoupled Embodiment Model, DEM),它将经过微调的 DINOv3 编码器与冻结的 NeoBERT 编码器相结合,并搭配一个可通过单次前向传播生成每个动作块的 MeanFlow 动作头。在 18 个包含保留语言改述和随机化场景的仿真操作任务以及三个真实机器人任务上,在我们的评估协议下,DEM 取得了与最先进的基于 VLM 主干策略相当的观测成功率,同时其推理频率达到后者的八到十七倍,每次推理的能耗降低六到十五倍。在该任务范围内,现代解耦组件提供了更优的成功率-延迟-能耗权衡。
cs.RO / 54 / 2609.18392

DistAL: Distance-based Advantage Learning for VLA Fine-Tuning

DistAL:面向VLA微调的基于距离的优势学习
O'Mahoney, Reece, Havoutis, Ioannis
Abstract
Vision-language-action models (VLAs) have trans- formed the field of robotic manipulation in recent years by combining the semantic understanding of LLMs with the precise control of flow-matching policies. Advantage conditioning is a recent technique that iteratively improves VLAs by training a value function on deployment data and using this to train an advantage-conditioned policy. Previous works have only applied simple, low-information success/failure rewards, which leave the value function unable to distinguish states of differing quality beyond how far along the task they appear. Motivated by an exploration of out-of-distribution (OOD) detection methods, we introduce Distance-based Advantage Learning (DistAL), which, by using an embedding space distance as a reward, produces a more informative value function and subsequently a higher downstream task success rate. We validate our method on a series of simulation benchmarks and dexterous bi-manual manipulation tasks on real hardware.
Chinese Translation
视觉-语言-动作模型(Vision-Language-Action Models,VLAs)近年来通过将大语言模型(LLMs)的语义理解能力与流匹配策略的精确控制相结合,变革了机器人操作领域。优势条件化(Advantage Conditioning)是一种新兴技术,其通过在部署数据上训练价值函数,并利用该价值函数训练优势条件化策略,从而迭代地改进VLA模型。以往的工作仅采用简单的、低信息量的成功/失败奖励,使得价值函数除了判断任务进度之外,无法区分质量不同的状态。受分布外(OOD)检测方法探索的启发,我们提出了基于距离的优势学习(Distance-based Advantage Learning,DistAL),该方法通过使用嵌入空间距离作为奖励,产生了信息量更大的价值函数,进而获得更高的下游任务成功率。我们在一系列仿真基准以及真实硬件上的灵巧双臂操作任务上验证了所提出的方法。
cs.RO / 55 / 2609.18395

DetAug: Obstacle-Blind Trajectory Augmentation for Zero-shot Obstacle Avoidance

DetAug:面向零样本避障的障碍盲轨迹增强方法
O'Mahoney, Reece, Zoellner, Moritz, Havoutis, Ioannis
Abstract
Policies for robotic manipulation are produced by training on large teleoperated datasets. These datasets typically consist of free-space trajectories, making them difficult to transfer to test-time environments with obstacles. Previous methods for closing this gap have largely fallen into two groups. Dataset augmentation addresses it at training time but needs obstacle geometry in advance, whereas steering an existing checkpoint at inference time avoids that requirement but is limited in flexibility. Our method draws from both areas without inheriting either drawback. DetAug applies an obstacle-blind augmentation scheme to the transit phases of a free-space dataset, leaving object interactions untouched, and records the augmentation parameters as an explicit conditioning label. At inference it samples a batch of labels and executes the trajectory with the lowest collision cost. On the SafeLIBERO benchmark DetAug achieves a collision-free success rate more than 20pp above the next best method, and selecting over the label space outperforms guidance on the same policy by 26pp. On real hardware, inference-time steering methods collapse on tasks requiring large detours, while DetAug matches or exceeds an obstacle-conditioned baseline without ever seeing obstacles in training.
Chinese Translation
机器人操作策略通常是通过在大型遥操作数据集上训练得到的。这些数据集一般由自由空间轨迹构成,因此难以迁移到测试时存在障碍物的环境中。此前弥合这一差距的方法主要分为两类:数据集增强方法在训练阶段解决该问题,但需要预先获知障碍物几何信息;而在推理阶段对已有模型检查点进行引导的方法虽然避免了这一要求,但灵活性有限。本文方法兼顾两者的优势,同时不继承各自的缺点。DetAug 对自由空间数据集的转移阶段施加一种障碍盲的增强方案,保持物体交互部分不变,并将增强参数记录为显式的条件标签。在推理时,该方法采样一批标签,并执行碰撞代价最低的轨迹。在 SafeLIBERO 基准上,DetAug 的无碰撞成功率比次优方法高出 20 个百分点以上;在同一策略上,通过在标签空间中进行选择,其性能比引导方法高出 26 个百分点。在真实硬件上,推理时引导方法在需要大范围绕行的任务上表现崩溃,而 DetAug 在训练中从未见过障碍物的情况下,达到甚至超过了障碍物条件化基线方法的性能。
cs.RO / 56 / 2609.18434

Hardware-Free Robotics Laboratories in Mixed Reality

基于混合现实的无硬件机器人实验室
Berrezueta-Guzman, Santiago, Khalil, Habiba-Loai, Koshelev, Andrei, Metaj, Vanesa, Wagner, Stefan
Abstract
Teaching robotics relies on screen-based simulation, showing robot motion in an abstract coordinate frame rather than at real scale in the learner's own space, while access to physical hardware is limited by cost, safety, and scheduling constraints. We present MR-Robotics LAB, a mixed-reality (MR) platform that replays MATLAB-generated robot trajectories at real scale within the learner's physical environment. A browser-based service validates a MATLAB workspace file (.mat), normalizes units, and publishes a versioned JSON trajectory; a Unity application on a Meta Quest 3 then reproduces the authored joint configurations under position control and replays them at the declared frame rate within a physics-enabled scene that supports collision detection and end-effector grasping. A formative single-group evaluation with engineering students found that participants reported low setup effort (M = 4.67 on a 5-point scale) and perceived support for workspace understanding from multi-viewpoint inspection (M = 4.56), and 83% of participants affirmed their willingness to use the platform in an introductory robotics course. The evaluation instrument records only perceived outcomes, without counterbalancing or a learning measure, so no comparative advantage over desktop simulation is claimed. The contribution is a reusable simulation-to-MR trajectory pathway and design guidance for hardware-free robot visualization in engineering education.
Chinese Translation
机器人学教学依赖于基于屏幕的仿真,在抽象坐标系中展示机器人运动,而非在学习者自身的真实空间中以真实比例呈现,同时实体硬件的使用又受成本、安全和排期等条件限制。我们提出了 MR-Robotics LAB,一个混合现实(MR)平台,可在学习者的物理环境中以真实比例回放由 MATLAB 生成的机器人轨迹。基于浏览器的服务验证 MATLAB 工作区文件(.mat)、统一单位并发布带版本号的 JSON 轨迹;随后,运行于 Meta Quest 3 上的 Unity 应用在位置控制模式下复现所编排的关节构型,并按设定的帧率在支持碰撞检测与末端执行器抓取的物理场景中进行回放。面向工科学生的单组形成性评估表明,参与者报告了较低的配置工作量(5 分量表中 M = 4.67),并认为多视角观察有助于对工作空间的理解(M = 4.56),83% 的参与者表示愿意在机器人学导论课程中使用该平台。该评估工具仅记录感知结果,未进行平衡处理,也未测量学习效果,因此不主张其相较于桌面仿真具有比较优势。本研究的贡献在于提供了一条可复用的“仿真到混合现实”的轨迹通路,以及面向工程教育中无硬件机器人可视化的设计指导。
cs.RO / 57 / 2609.18451

VLM-MPPI: Grounding Natural Language in Behaviorally Diverse Trajectories for Aerial Navigation

VLM-MPPI:将自然语言映射到行为多样性轨迹的空中导航方法
Zhang, Hanbing, Zhao, Fangguo, Li, Zerui, Guan, Xin, Cheng, Peng, Li, Shuo
Abstract
We present a hierarchical UAV navigation framework that aligns natural-language intent with dynamically feasible flight behaviors in cluttered indoor environments. To bridge the gap between abstract semantics and low-level control, we employ a parallelized ensemble of six behavior-conditioned Model Predictive Path Integral (MPPI) planners. Crucially, by designing mode-specific guiding costs and sampling biases, we induce distinct trajectory modes that converge to unique behavioral means, yielding a compact set of intentionally diverse candidates rather than mere stochastic variations. We project these 3D candidates onto the onboard first-person-view RGB stream, turning language grounding into a visual action selection problem. A pretrained vision--language model (VLM) asynchronously selects the candidate index given the overlaid FPV image and a natural-language prompt, while MPPI replans at 20Hz and a PID-based low-level controller tracks the selected trajectory. We implement the full pipeline in NVIDIA Isaac Sim and on a real-world quadrotor platform equipped with LiDAR and RGB sensing. Experiments in both simulation and real-world flights show semantically meaningful behavior diversity, robust language alignment despite VLM latency, and safe, repeatable flight across all modes, achieving 100% task success in our evaluated scenarios.
Chinese Translation
我们提出了一种分层无人机导航框架,能够在杂乱的室内环境中将自然语言意图与动态可行的飞行行为对齐。为了弥合抽象语义与底层控制之间的鸿沟,我们采用了六个行为条件化的模型预测路径积分(MPPI)规划器的并行化集成。关键在于,通过设计针对不同模式的引导代价和采样偏置,我们诱导出不同的轨迹模式并收敛到独特的行为均值,从而得到一组紧凑且有意多样化的候选轨迹,而非仅仅是随机变化。我们将这些3D候选轨迹投影到机载第一人称视角(FPV)RGB图像流上,将语言接地转化为视觉动作选择问题。预训练的视觉-语言模型(VLM)根据叠加后的FPV图像和自然语言提示,异步地选择候选轨迹的索引,同时MPPI以20Hz的频率进行重规划,基于PID的底层控制器跟踪所选轨迹。我们在NVIDIA Isaac Sim以及配备LiDAR和RGB传感器的真实四旋翼平台上实现了完整的系统流程。仿真和真实飞行实验均表明,该系统具有语义上有意义的行为多样性、尽管存在VLM延迟仍保持稳健的语言对齐能力,以及所有模式下安全且可重复的飞行,在我们评估的场景中实现了100%的任务成功率。
cs.RO / 58 / 2609.18455

ForwardDLO: Model-Based Bimanual Shape Matching of Unconstrained Deformable Linear Objects

ForwardDLO:基于模型的非约束可变形线性物体双臂形状匹配
Missal, Tim, Guler, Berk, Domingues, Lucas, Manschitz, Simon, Peters, Jan, Costa, Paula Dornhofer Paro
Abstract
Ropes, cables, and other deformable linear objects appear in tasks from untangling to cable routing and suturing, yet controlling their shape remains a challenge in robot manipulation. We study model-based shape control in a general setting: the object lies unfixated on a support surface and two arms may grasp and move it anywhere along its length. Because each arm chooses a grasp point, direction, and magnitude, the joint action space is combinatorially large, and the dynamics model's per-prediction cost bounds how much of it a planner can search. We present ForwardDLO, a recurrent latent dynamics model for this unfixated bimanual setting that predicts per-segment displacements grounded in the observed rope state at every step. Our model reaches accuracy comparable to more expensive baselines while containing no explicit segment-to-segment operations, which makes batched evaluation of candidate actions cheap. On open-loop prediction of real rope motion it reaches the lowest error of the learned models we evaluate, 13% below the strongest baseline. Within a fixed time budget it scores 8 to 22 times more candidate actions than models of comparable accuracy while matching them in real-world shape matching; and on a simulated routing task at a 30Hz control rate, this throughput converts into 98% task success versus at most 30% for the baselines at their own budgets. We release the model, code, and a dataset of 2.42 million simulated and 14,107 real rope transitions at https://anonymous.4open.science/r/ForwardDLO/
Chinese Translation
绳子、线缆及其他可变形线性物体出现在从解缠、线缆布线到缝合等多种任务中,然而在机器人操作中控制其形状仍然是一个挑战。我们研究一般场景下基于模型的形状控制:物体未被固定地放置在支撑面上,两只机械臂可沿其长度任意位置抓取并移动它。由于每条机械臂需选择抓取点、方向和力度,联合动作空间呈组合级规模,而动力学模型每次预测的计算成本限制了规划器所能搜索的范围。我们提出ForwardDLO,一种面向这种非固定双臂场景的循环隐变量动力学模型,可在每一步基于观测到的绳子状态预测各段的位移。我们的模型在不包含任何显式段间操作的情况下达到了与更昂贵的基线方法相当的精度,这使得对候选动作进行批量评估变得十分高效。在真实绳子运动的开环预测中,它在我们评估的学习型模型中取得了最低误差,比最强基线低13%。在固定时间预算内,它对候选动作的评估数量是精度相当模型的8至22倍,同时在真实世界形状匹配任务中与之表现相当;在30Hz控制频率的仿真布线任务中,这种吞吐量优势转化为98%的任务成功率,而基线方法在其各自预算下最高仅为30%。我们在 https://anonymous.4open.science/r/ForwardDLO/ 发布了模型、代码,以及包含242万条仿真和14,107条真实绳子状态转移的数据集。
cs.RO / 59 / 2609.18482

Real-Time Bounded Catenary Solver for UAV Tether Modeling

用于无人机系留缆绳建模的实时有界悬链线求解器
Beffert, Max, Zell, Andreas
Abstract
For non-stationary tethered multirotor UAVs in real-world conditions, simulating the forces imposed on the drone by the aerodynamic drag of the tether becomes crucial, with online use cases placing a hard bound on the maximum solve time. In previous work, a quasi-analytical catenary tether model reached a mean solve time of 0.51 ms using a general-purpose root finder, but without any worst-case guarantees or proven convergence. In this work, we reformulate the inner solver by reducing the catenary boundary-value problem to a single transcendental equation in one well-conditioned unknown. We derive a closed-form bracket and prove monotonicity and convexity as well as existence and uniqueness of the root, which together guarantee convergence of the solver. We further propose a two-regime initial guess which approximates the true root within 3.4% and reduces the mean iteration count by 68.0% to 2.36 compared to the textbook initialization. Building on the hybrid root-finding method rtsafe (Newton-Raphson with bisection fallback giving bounded iteration counts), we implement a specialized variant that exploits the problem structure to omit unnecessary checks while retaining correctness, which gives up to 1.3 times speedup. With the proposed solver the full tether model achieves a nearly constant solve time of 6.9 us on average and 7.7 us at worst, a 40 times speedup over an optimized re-implementation of the previous method, while agreeing with it to a relative deviation of 8.7e-9. Because the reformulation leaves the underlying physical model untouched, the experimental validation of the previous work carries over unchanged. We further demonstrate its suitability for embedded, resource-constrained platforms with a Lua implementation running directly in ArduPilot on a drone's flight controller, where it stays well inside the scheduling budget with a mean solve time of 0.74 ms.
Chinese Translation
对于真实环境下的非静止系留多旋翼无人机而言,模拟系留缆绳的气动阻力对无人机所施加的力至关重要,而在线应用场景对最大求解时间提出了严格的上界要求。在先前的工作中,一种准解析悬链线缆绳模型使用通用求根器达到了0.51 ms的平均求解时间,但缺乏任何最坏情况的保证或收敛性证明。在本工作中,我们通过将悬链线边值问题简化为关于单个良态未知量的超越方程,重新构建了内部求解器。我们推导出了闭式区间(bracket),并证明了该方程的单调性、凸性以及根的存在性和唯一性,这些性质共同保证了求解器的收敛性。我们进一步提出了一种双区制(two-regime)初始猜测方法,可将真实根近似至3.4%以内,并将平均迭代次数相比教科书式的初始化方法降低了68.0%,降至2.36次。基于混合求根方法rtsafe(即带二分法回退的牛顿-拉夫逊法,可提供有界的迭代次数),我们实现了一个专门化的变体,该变体利用问题结构省去了不必要的检查同时保持正确性,从而获得了最高1.3倍的加速。采用所提出的求解器,完整的缆绳模型实现了近乎恒定的求解时间——平均6.9微秒,最坏情况7.7微秒,相比先前方法的一个优化重实现获得了40倍的加速,同时与其结果的相对偏差仅为8.7e-9。由于该重新构建并未改动底层的物理模型,先前工作的实验验证结果可直接沿用。我们进一步通过一个直接运行于无人机飞控ArduPilot中的Lua实现,展示了该方法在嵌入式资源受限平台上的适用性,其中其平均求解时间为0.74 ms,完全满足调度预算要求。
cs.RO / 60 / 2609.18487

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

ActionPiece:重新思考自回归视觉-语言-动作模型的动作分词方法
Lian, Shijie, Yu, Bin, Shen, Zhaolong, Lin, Xiaopeng, Du, Yichao, Zhang, Zhirui, Yang, Laurence T., Chen, Kai
Abstract
Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.
Chinese Translation
动作分词器(action tokenizer)在自回归视觉-语言-动作(VLA)模型中扮演着核心角色,它既决定了策略训练的目标,也决定了从预测词元中恢复的可执行指令。其保真度通常采用逐点重建指标(如均方误差 MSE)进行评估,然而微小的个体误差并不能完全刻画跨演示样本的动作调整被保留的程度。压缩之后,相似的动作可能仍然聚集在某个代表性运动周围,而不同情境下所需的调整却被削弱、扭曲甚至反转。我们提出物理秩一致性(physical rank consistency, PRC)来衡量分词方法在重建后对局部物理距离排序的保留程度。对解码后的动作进行评估为不同词表和不同解码器架构提供了统一的参考基准,在逐点精度之外补充了对关系保真度的度量。我们进一步提出 ActionPiece,通过表征学习与量化的联合监督来保留物理动作关系。物理秩保持监督编码器与量化特征距离中的远近排序,而量化正则化则将同样的排序应用于码字分配分布。这两个目标均是对重建的增广,所产生的离散动作词元可用于标准的自回归策略学习,并通过冻结的解码器执行。在相同的 Qwen3-VL-4B 策略训练设置下,ActionPiece 在 LIBERO 上取得 94.8%、在未见过的 LIBERO-Plus 上取得 68.8% 的成绩,此外在 SimplerEnv 上达到 71.9%,在 VLA-Arena L0-L2 上达到 51.5%。组件消融实验表明,这两个目标共同提升了 PRC 与策略成功率,证明了物理关系监督对动作分词的价值。
cs.RO / 61 / 2609.18497

TAO-Force: Unifying Force-Aware Perception and Fast-Slow Control for Contact-Rich Manipulation

TAO-Force:面向接触密集型操作的力感知感知与快慢控制统一框架
Gan, Bohan, Wen, Xuanzhang, Zhao, Yongsheng, Cheng, Baoping, Jia, Wenhe, Wang, Ye, Yao, Gongxin, Gao, Han, Tang, Jingyao, Zhao, Lei, Ge, Ji
Abstract
Vision-Language-Action (VLA) models have demonstrated strong performance across diverse robotic manipulation tasks, yet their predominantly vision-centric perception and position-controlled execution remain insufficient for contact-rich manipulation. Visual observations alone often provide limited evidence of contact onset and interaction magnitude, while position-control policies cannot respond compliantly to rapidly changing contact dynamics. To bridge both the perception and control gaps, we propose TAO-Force, a force-conditioned VLA framework that combines force-aware policy learning with contact-regulated execution. For force-aware perception, TAO-Force introduces Force-conditioned Feature-wise Linear Modulation (F-FiLM) to inject encoded force feedback into the representations of a frozen pretrained visual-language backbone while preserving its semantic priors. For responsive control, it employs a contact-gated fast-slow architecture, with a slow position-control branch tracking nominal trajectories during non-contact phases and a fast admittance-control branch regulating physical interaction during contact phases. Detailed analyses on a force-perception task and real-world evaluations across four contact-rich manipulation tasks validate the effectiveness and robustness of TAO-Force.
Chinese Translation
视觉-语言-动作(VLA)模型在多种机器人操作任务中展现出强大的性能,但其以视觉为主的感知方式和基于位置控制的执行方式仍不足以应对接触密集型操作。仅依靠视觉观测往往难以提供足够的关于接触发生时刻与交互强度的信息,而位置控制策略也无法柔顺地响应快速变化的接触动力学。为同时弥合感知与控制两方面的差距,我们提出了 TAO-Force,一个将力感知策略学习与接触调节执行相结合的力条件化 VLA 框架。在力感知方面,TAO-Force 引入了力条件化特征级线性调制(Force-conditioned Feature-wise Linear Modulation, F-FiLM),将编码后的力反馈注入冻结的预训练视觉-语言骨干网络的表征中,同时保留其语义先验。在响应式控制方面,它采用接触门控的快慢架构:在非接触阶段,慢速位置控制分支跟踪标称轨迹;在接触阶段,快速导纳控制分支调节物理交互。针对力感知任务的详细分析以及在四项接触密集型操作任务上的真实世界实验,验证了 TAO-Force 的有效性与鲁棒性。
cs.RO / 62 / 2609.18504

InterMASH: A Unified Geometric Representation for Grasp Synthesis

InterMASH:一种用于抓取合成的统一几何表示
Yang, Xuanze, Liu, Yumeng, Xin, Haiyang, Li, Changhao, Shen, Haowei, Xu, Kai, Liu, Ligang, Hu, Ruizhen
Abstract
Grasp synthesis aims to generate stable and physically plausible hand--object interactions, and has become a fundamental problem in both human hand modeling and robotic manipulation. However, a unified representation across human and robotic hands is still lacking, mainly due to differences in hand morphology and surface modeling. Prior methods typically rely on either contact maps or dense implicit descriptors to represent interaction, but these representations are often incomplete or computationally expensive and redundant. We propose InterMASH, a unified geometric representation that establishes cross-embodiment correspondence using sphere-fixed anchors. At each anchor, low-degree spherical harmonics compactly encode local hand geometry, object geometry, and contact, forming an explicit and interpretable token sequence. Building on this natively tokenized structure, we introduce a conditional Diffusion Transformer that operates directly in the proposed InterMASH representation space and jointly generates hand geometry and contact, improving consistency and physical plausibility. Our method achieves competitive performance with state-of-the-art methods on key physical feasibility metrics in a large-scale ShadowHand benchmark, supports joint training across multiple hands, and shows that cross-embodiment fine-tuning with human grasp data can improve robotic grasp success and diversity. Project page is available at https://inter-mash.github.io/.
Chinese Translation
抓取合成旨在生成稳定且物理合理的手—物体交互,已成为人手建模与机器人操作中的一个基础性问题。然而,由于人手与机器手在手部形态和表面建模上的差异,目前仍缺乏一种跨两者的统一表示。已有方法通常依赖接触图或稠密隐式描述符来表示交互,但这些表示往往不完整,或计算开销大且存在冗余。我们提出 InterMASH,一种统一的几何表示,利用固定于球面的锚点建立跨具身(cross-embodiment)对应关系。在每个锚点处,低阶球谐函数紧凑地编码局部手部几何、物体几何以及接触信息,形成一个显式且可解释的 token 序列。基于这种天然 token 化的结构,我们引入条件扩散Transformer(Diffusion Transformer),直接在所提出的 InterMASH 表示空间中运行,联合生成手部几何与接触信息,提升了交互的一致性和物理合理性。在大规模 ShadowHand 基准上,我们的方法在关键物理可行性指标上取得了与最先进方法相当的性能;支持多手部联合训练,并表明利用人手抓取数据进行跨具身微调可以提升机器人抓取的成功率与多样性。项目页面见 https://inter-mash.github.io/。
cs.RO / 63 / 2609.18514

ActiveScale: Scaling Active Perception for Robots across Model, Data, and Hardware

ActiveScale:跨模型、数据与硬件的机器人主动感知规模化
Zhou, Shuai, Pang, Kaisheng, Song, Wenxuan, Zhang, Wenjie, Zheng, Xinhu, Li, Haoang
Abstract
Active perception is essential for robotic manipulation when fixed viewpoints leave task-relevant information occluded or unobserved. However, enabling vision-language-action (VLA) models to reason across changing viewpoints and actively acquire informative observations remains challenging. We present ActiveScale, a framework that advances active perception through coordinated model, data, and hardware designs. Our model augments a VLA with historical video observations and explicit camera-pose supervision, using per-frame pose tokens and a lightweight prediction head to associate observations across viewpoints and support a coherent understanding of the scene. To learn from the camera motion naturally present in human activity, we introduce a scalable human--robot mid-training recipe using 1000 hours of egocentric and robotic data, adapting the model to temporal inputs and pose supervision. We further introduce Active-perception Mobile-manipulation Platform (AMP), a robotic platform that supports active perception and mobile manipulation through single-operator teleoperation, enabling scalable collection of demonstrations that coordinate viewpoint changes and manipulation. Experiments demonstrate improved success rates on active-perception tasks, while ablation studies validate the contributions of camera-pose-aware modeling and egocentric mid-training. Together, these components provide an integrated foundation for studying and developing active perception in robotic manipulation.
Chinese Translation
主动感知对于机器人操作至关重要,因为固定视角会导致任务相关信息被遮挡或无法观测。然而,使视觉-语言-动作(VLA)模型能够在变化的视角间进行推理并主动获取有信息量的观测仍然具有挑战性。我们提出了ActiveScale,一个通过模型、数据和硬件协同设计来推进主动感知的框架。我们的模型在VLA基础上增加了历史视频观测和显式的相机位姿监督,利用逐帧位姿标记(pose tokens)和一个轻量级预测头在不同视角间关联观测,从而支持对场景的连贯理解。为了从人类活动中自然存在的相机运动中学习,我们引入了一种可扩展的人-机器人中期训练(mid-training)方法,使用1000小时的第一人称视角(egocentric)和机器人数据,使模型适应时序输入和位姿监督。我们进一步提出了主动感知移动操作平台(Active-perception Mobile-manipulation Platform, AMP),该机器人平台通过单人遥操作支持主动感知和移动操作,实现了视角变化与操作相协调的大规模示范数据采集。实验表明,该方法在主动感知任务上提升了成功率,消融研究验证了相机位姿感知建模和第一人称中期训练的贡献。这些组件共同为研究和开发机器人操作中的主动感知提供了一个集成化基础。
cs.RO / 64 / 2609.18549

DynoFluxBench: Benchmarking Kinodynamic Space-Time Planners in Dynamic Environments

DynoFluxBench:动态环境中动力学可行时空规划器的基准测试框架
Queißner, Franz, Orthey, Andreas, Hönig, Wolfgang
Abstract
Robots that leave structured, static environments must plan motions that are kinodynamically feasible and safe among moving obstacles. However, there are no dedicated benchmark frameworks that combine both aspects. To overcome this, we present DynoFluxBench, a framework to compare kinodynamic planners in known, dynamic environments with unbounded arrival time. To demonstrate its utility and establish strong baselines, we develop three dedicated planners, named ST-Db-RRT, ST-GBRRT, and KIST, that fuse kinodynamic and space-time methods, covering different kinodynamic search paradigms: ST-Db-RRT expands with randomly selected discontinuity-bounded motion primitives using trajectory optimization, whereas KIST and ST-GBRRT maintain a kinodynamically feasible tree with different heuristic guidance. We analyze the probabilistic completeness guarantees of those new planners in dynamic environments. Finally, we evaluate ST-Db-RRT, ST-GBRRT, and KIST using DynoFluxBench, showing that ST-Db-RRT reaches a first solution up to 32 times faster, while KIST and ST-GBRRT remain valuable where trajectory optimization is fragile. Videos and further analysis can be found at https://dynofluxbench.github.io/dynofluxbench/.
Chinese Translation
离开结构化静态环境的机器人必须在移动障碍物中规划运动学上可行且安全的运动。然而,目前尚无将这两个方面结合起来的专用基准测试框架。为克服这一问题,我们提出了 DynoFluxBench,一个用于在已知动态环境(且到达时间不受限制)中比较动力学可行(kinodynamic)规划器的框架。为展示其实用性并建立强基线,我们开发了三个专用规划器,分别命名为 ST-Db-RRT、ST-GBRRT 和 KIST,它们融合了动力学可行方法与时空方法,涵盖不同的动力学可行搜索范式:ST-Db-RRT 借助轨迹优化,以随机选择的带不连续性约束的运动基元进行扩展;而 KIST 和 ST-GBRRT 则采用不同的启发式引导来维护一棵运动学可行的搜索树。我们分析了这些新规划器在动态环境中的概率完备性保证。最后,我们使用 DynoFluxBench 对 ST-Db-RRT、ST-GBRRT 和 KIST 进行了评估,结果表明 ST-Db-RRT 获得首个解的速度最高可快 32 倍,而在轨迹优化方法较为脆弱的场景中,KIST 和 ST-GBRRT 仍具有重要价值。视频及进一步分析请参见 https://dynofluxbench.github.io/dynofluxbench/。
cs.RO / 65 / 2609.18581

GroundingVLN: Reasoning and Acting with Grounding for Vision-Language Navigation

GroundingVLN:基于视觉定位的推理与行动视觉语言导航
Li, Kailing, Han, Yu, Qian, Tianwen, Fu, Yuqian, Gong, Jingyu, Shi, Jiangming, Wang, Xiaoling
Abstract
Although vision-language models (VLMs) possess strong visual understanding and reasoning capabilities, existing vision-and-language navigation (VLN) agents struggle to connect semantic reasoning with spatial execution. Two coupled gaps remain in this connection, as intermediate reasoning is not explicitly anchored to visual evidence and high-level decisions lack precise spatial goals to guide low-level motion. Cognitive science suggests that human navigation bridges these levels hierarchically by anchoring cognition to relevant landmarks and guiding locomotion toward spatial goals. Motivated by this principle, we propose GroundingVLN, which uses visual grounding as a shared interface between reasoning and action. GroundingVLN first reasons with grounding by anchoring task-relevant visual evidence to precise image locations throughout structured reasoning. It then acts through grounding by predicting a progress-aligned pixel goal that a geometric planner translates into primitive actions. To learn these capabilities, we construct GroundingCOTVLN-188K, a dataset of temporally aligned grounded reasoning traces, and introduce Grounded and Execution-Aware Reinforcement Learning (GEAR), which aligns grounded reasoning and spatial decisions with downstream execution. Experiments demonstrate that GroundingVLN achieves state-of-the-art performance (69.9% SR on R2R-CE and 75.1% SR on RxR-CE) with high sample efficiency, using just 0.9% as much training data as the strongest baseline. It also generalizes strongly across datasets, attaining 59.9% SR on RxR-CE when trained solely on R2R, a gain of 20.1% over the strongest baseline.
Chinese Translation
尽管视觉语言模型(VLM)具备强大的视觉理解与推理能力,现有的视觉语言导航(VLN)智能体仍难以将语义推理与空间执行相连接。这种连接中存在两个相互耦合的缺口:中间推理过程未被显式地锚定到视觉证据上,且高层决策缺乏精确的空间目标来引导底层运动。认知科学研究表明,人类导航通过将认知锚定于相关地标并引导身体朝空间目标移动,从而以层次化的方式衔接这两个层面。受此原理启发,我们提出了GroundingVLN,将视觉定位(visual grounding)作为推理与行动之间的共享接口。GroundingVLN首先进行带定位的推理:在整个结构化推理过程中,将任务相关的视觉证据锚定到精确的图像位置。随后通过定位进行行动:预测一个与导航进度对齐的像素级目标,并由几何规划器将其转换为底层动作。为了学习这些能力,我们构建了GroundingCOTVLN-188K,一个时间对齐的带定位推理轨迹数据集,并提出带定位与执行感知的强化学习方法(Grounded and Execution-Aware Reinforcement Learning, GEAR),使带定位的推理与空间决策与下游执行相一致。实验表明,GroundingVLN以极高的样本效率实现了最先进的性能(在R2R-CE上达到69.9% SR,在RxR-CE上达到75.1% SR),训练数据量仅为最强基线的0.9%。同时,该方法在跨数据集上展现出强大的泛化能力:仅在R2R上训练即可在RxR-CE上达到59.9% SR,比最强基线提升了20.1%。
cs.RO / 66 / 2609.18620

DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation

DeformSmith:面向机器人操作的物理测试框架引导的可形变资产分层生成方法
Li, Can, Gu, Jie, Deng, Zishun, Chen, Jingmin, Sun, Lei
Abstract
Creating deformable assets for robot manipulation requires jointly specifying their geometry, appearance, and physical properties. This is especially challenging for deformable objects, since text and images provide limited evidence about how they deform and respond to contact, yet these responses directly affect their suitability for interaction. Automated generation therefore needs to resolve coupled physical requirements and use interaction evidence to guide construction and refinement. We present DeformSmith, a framework that enables automated generation of interactive, physically credible deformable assets from text or a single image. Through hierarchical agentic construction and a shared physics-grounded harness, it progressively builds, tests, and refines geometry, physical models, material behavior, and robot interaction until the resulting asset is ready for simulation and manipulation. Robot interaction closes the generation loop through manipulation feedback and replayable interaction data. Results show that DeformSmith generates assets with better visual quality and physical plausibility than state-of-the-art baselines, including PhysGen3D, PhysGM, and PhysX-Omni, while supporting the synthesis of data for robotic manipulation of deformable objects. Project page: https://can-lee.github.io/deformsmith-web/
Chinese Translation
为机器人操作创建可形变资产需要同时确定其几何形状、外观和物理属性。这对可形变物体而言尤其具有挑战性,因为文本和图像只能提供关于物体如何形变及响应接触的有限信息,而这些响应直接决定了物体是否适合交互。因此,自动化生成需要解决相互耦合的物理需求,并利用交互证据来引导资产的构建与精化。我们提出了DeformSmith,一个能够从文本或单张图像自动生成可交互、物理上合理的可形变资产的框架。通过分层智能体构建和共享的物理约束测试框架(harness),该框架逐步构建、测试并精化几何形状、物理模型、材料行为以及机器人交互,直至生成的资产可用于仿真与操作。机器人交互通过操作反馈和可回放的交互数据来闭合生成回路。实验结果表明,与PhysGen3D、PhysGM和PhysX-Omni等最先进的基线方法相比,DeformSmith生成的资产具有更好的视觉质量和物理合理性,同时支持为可形变物体的机器人操作合成数据。项目页面:https://can-lee.github.io/deformsmith-web/
cs.RO / 67 / 2609.18628

Benchmarking Visual-Inertial Odometry in Subterranean Environments Under Sensor Degradation, Miscalibration, and Dynamic Occlusion

传感器退化、标定误差与动态遮挡条件下地下环境中视觉惯性里程计的基准测试
Zhu, Yueying, Li, Xiang, Nguyen, Thien-Minh, Wang, Xuehe, Yuan, Shenghai
Abstract
Visual-inertial odometry (VIO) is a core capability for autonomous operation in GPS-denied subterranean environments, yet its reliability can degrade sharply under sensor drift, calibration errors, and dynamic occlusion. Existing evaluations mainly emphasize nominal-condition accuracy, offering limited insight into when practical deployment failures occur. In this work, we present a failure-centric stress-test benchmark for VIO in underground environments using the CERBERUS dataset. We systematically evaluate four representative VIO systems spanning filtering-, optimization-, and learning-based paradigms under nine practical perturbation settings, including IMU bias and noise variation, camera intrinsic and extrinsic drift, and dynamic scene occlusion. Beyond conventional trajectory error, we analyze robustness limits through coverage ratio and failure thresholds, revealing breakdown behaviors that are not captured by nominal-condition performance alone. Our study shows distinct vulnerability patterns across VIO paradigms: some methods are more sensitive to inertial degradation, while others are more affected by geometric miscalibration or dynamic interference. These results provide deployment-oriented guidance for VIO selection, calibration prioritization, and reliable operation in challenging underground scenarios. To support reproducible evaluation and future extensions, we will release the full benchmark scripts and evaluation pipeline.
Chinese Translation
视觉惯性里程计(VIO)是在GPS拒止的地下环境中实现自主运行的核心能力,然而其在传感器漂移、标定误差和动态遮挡条件下的可靠性可能急剧下降。现有评估主要强调理想条件下的精度,对实际部署中故障何时发生缺乏深入的洞察。在本工作中,我们基于CERBERUS数据集,提出了一套以故障为中心的地下环境VIO压力测试基准。我们系统地评估了涵盖滤波、优化和基于学习三类范式的四个代表性VIO系统,在九种实际扰动设置下的表现,包括IMU偏置与噪声变化、相机内参与外参漂移以及动态场景遮挡。除传统轨迹误差外,我们通过覆盖率与故障阈值分析鲁棒性极限,揭示了仅凭理想条件性能无法捕捉的失效行为。我们的研究表明,不同VIO范式具有不同的脆弱性模式:某些方法对惯性退化更为敏感,而另一些方法则更易受几何标定误差或动态干扰的影响。这些结果为VIO的选择、标定优先级排序以及在具有挑战性的地下场景中的可靠运行提供了面向部署的指导。为支持可复现的评估与后续扩展,我们将发布完整的基准测试脚本与评估流程。
cs.RO / 68 / 2609.18650

From Gameplay to Policy: Towards Scalable Robot Data Collection via Gamified Robot-Free Interaction

从游戏玩法到策略:通过游戏化的无机器人交互实现可扩展的机器人数据采集
Li, Zheng, Zhu, Liang, Wang, Junzhe, Chen, Huayuan, Liu, Ziyun, Cao, Jiahang, Sheng, Xinyu, Qu, Pei, Jia, Yufei, Zhang, Ximeng, Xie, Jiarui, Yuan, Zizhao, Li, Haoang, Cai, Yi, Zhou, Jinni, Ma, Jun
Abstract
Learning generalizable robot manipulation policies requires large-scale and diverse interaction data, yet collecting real-world demonstrations remains costly and difficult to scale. Existing approaches to data collection are either dependent on specific robot hardware that limits crowdsourcing and transferability, or suffer from incomplete annotation and limited behavioral diversity. Inspired by how games sustain long-term human engagement, we explore an alternative paradigm that turns data collection into an engaging gameplay experience and transfers the resulting human manipulation experience to real robots. We present Project Kitchen, a VR-based gamified egocentric data collection platform that elicits diverse, goal-directed manipulation while remaining independent of specific robot embodiments and hardware, making it applicable to broader and potentially large-scale deployment. To bridge the game-to-real gap, we further introduce Game2Policy, which extracts embodiment-invariant affordance cues, including contact points and sub-goal states, from gameplay trajectories. An affordance model is pre-trained on game-collected data and then jointly fine-tuned with downstream policies using only a handful of real-robot demonstrations. Experiments show that Game2Policy improves average success rates by 10.0 points in simulation and 18.3 points on real robots in the few-shot setting. User studies and quantitative analyses further show that Project Kitchen promotes diverse manipulation behaviors and provides an engaging data collection experience. These results demonstrate the potential of gamified virtual environments as a scalable source of manipulation knowledge. The platform and code will be released upon acceptance.
Chinese Translation
学习可泛化的机器人操作策略需要大规模且多样化的交互数据,然而采集真实世界的演示数据仍然成本高昂且难以规模化。现有的数据采集方法要么依赖于特定的机器人硬件,从而限制了众包和可迁移性;要么存在标注不完整和行为多样性受限的问题。受游戏如何维持人类长期参与度的启发,我们探索了一种替代范式,即将数据采集转变为引人入胜的游戏体验,并将由此产生的人类操作经验迁移到真实机器人上。我们提出了Project Kitchen,一个基于VR的游戏化自我中心(egocentric)数据采集平台,它能够激发多样化、目标导向的操作行为,同时不依赖于特定的机器人形态和硬件,从而适用于更广泛且潜在大规模的部署。为弥合从游戏到现实的差距,我们进一步提出了Game2Policy,它从游戏轨迹中提取形态无关(embodiment-invariant)的可供性线索,包括接触点和子目标状态。在游戏采集的数据上预训练一个可供性模型,然后仅使用少量真实机器人演示与下游策略进行联合微调。实验表明,在少样本设置下,Game2Policy在仿真中将平均成功率提升了10.0个百分点,在真实机器人上提升了18.3个百分点。用户研究和定量分析进一步表明,Project Kitchen能够促进多样化的操作行为,并提供引人入胜的数据采集体验。这些结果证明了游戏化虚拟环境作为可扩展操作知识来源的潜力。平台和代码将在论文被接收后发布。
cs.RO / 69 / 2609.18651

FIERCE: From Generalist Robot Policies to Fast Specialists via Progress-Failure Feedback

FIERCE:基于进展-失败反馈从通用机器人策略到快速专用策略
Tan, Runjia, Tu, Yuang, Yan, Yujie, Yu, Lan, Tian, Xuesong, Lv, Chen
Abstract
Generalist robot policies offer useful initialization, but refining compact specialists through limited physical interaction requires informative learning feedback. We present FIERCE, a generalist-initialized reinforcement learning framework centered on a unified, task-adaptive progress-failure evaluator. Its architecture shares an observation-language representation between an observed-progress head and an action-conditioned latent predictor whose past and current predictions feed a causal sequence head for task-failure estimation. Joint supervision from progress and preference labels, synchronized commands and observations, and terminal outcomes trains the evaluator; target-task rollouts support adaptation and calibration. Fixed evaluator snapshots provide progress shaping and failure-risk penalties alongside independently verified terminal rewards, while evaluator and policy updates alternate as new experience is collected. Refinement requires neither continued generalist action queries nor a dedicated target-task simulator or manually annotated dense rewards. Only the compact specialist is retained at deployment. The evaluation separates feedback quality, policy-learning efficiency, and deployment cost across simulation and two contact-rich real tasks. Code, model weights, and data-restoration tools are released at https://github.com/ar-mine/FIERCE.
Chinese Translation
通用机器人策略提供了有用的初始模型,但通过有限的物理交互将其精炼为紧凑的专用策略,需要信息丰富的学习反馈。我们提出FIERCE,一个以通用策略初始化的强化学习框架,其核心是一个统一的、任务自适应的进展-失败评估器。该架构在观测进展头与动作条件隐变量预测器之间共享观测-语言表示,预测器的过去与当前预测被输入因果序列头以进行任务失败估计。评估器通过进展标签与偏好标签的联合监督、同步的命令与观测以及终止结果进行训练;目标任务的实际执行数据用于支持适应与校准。固定的评估器快照与独立验证的终止奖励一同提供进展塑形和失败风险惩罚,同时评估器与策略的更新在收集新经验的过程中交替进行。该精炼过程既不需要持续的通用策略动作查询,也不需要专门的目标任务模拟器或人工标注的密集奖励。部署时仅保留紧凑的专用策略。评估在仿真以及两个富接触的真实任务上分别考察了反馈质量、策略学习效率和部署成本。代码、模型权重和数据恢复工具已在 https://github.com/ar-mine/FIERCE 发布。
cs.RO / 70 / 2609.18663

VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge

VLA-ULAP:在边缘端以超轻量本地动作预测与云端VLA调用交错执行
Cao, Deyu, Oi, Ryuji, Matsushima, Kosuke, Pan, Yuxuan, Wang, Ziheng, Fujiki, Daichi, Kosuge, Atsutake
Abstract
Billion-parameter vision--language--action (VLA) policies demand substantial onboard power, while communication delays in remote inference hinder timely responses. We propose VLA-ULAP, which interleaves remote VLA calls with an Ultra-Lightweight Local Action Predictor (ULAP). With approximately 7.4M parameters including the frozen vision encoder, ULAP combines current views, proprioception, and executed action history to predict chunks in one pass. Trained independently, it requires no VLA hidden states, online verification, or server round trips. On Jetson Orin Nano, ULAP takes 19.9 ms and 0.183 J per inference, compared with 284.3 ms and 50.55 J for GR00T on RTX A6000. Across three simulated base-policy/benchmark pairs, selected operating points remove 48.8--76.7\% of VLA calls while retaining 95.0--97.5\% of the baseline success rate. Against local VLA-acceleration alternatives on VLA-JEPA, ULAP uses an estimated 49.2\% less inference time and 51.0\% less GPU energy per successful episode than ACT at comparable success rates, and 77.1\% less time and 79.9\% less energy than SP-VLA at equal success rates. Physical SO-101 experiments retain 95.2--100\% of the baseline success rate across seen and held-out placements while reducing inference time by an estimated 47.9--58.0\% and inference-device energy by 52.1--62.5\%, based on successful-episode call counts and measured device costs. Faster responses also improve dynamic-task success rates: in latency-aware LIBERO-Safety simulation, VLA-ULAP exceeds $\pi_{0.5}$ by 11.0 and 15.5 percentage points on two tasks while approximately halving VLA calls.
Chinese Translation
数十亿参数的视觉-语言-动作(VLA)策略需要大量机载算力,而远程推理中的通信延迟会阻碍及时响应。我们提出VLA-ULAP,将远程VLA调用与超轻量本地动作预测器(Ultra-Lightweight Local Action Predictor,ULAP)交错执行。ULAP包含约740万参数(含冻结的视觉编码器),融合当前视角、本体感觉与已执行动作历史,单次前向即可预测动作块。其训练独立进行,无需VLA隐藏状态、在线验证或与服务器的往返通信。在Jetson Orin Nano上,ULAP每次推理仅需19.9毫秒和0.183焦耳,而GR00T在RTX A6000上需284.3毫秒和50.55焦耳。在三组仿真基座策略/基准测试组合中,所选运行点在移除48.8–76.7%的VLA调用的同时,保留了基线成功率95.0–97.5%。与VLA-JEPA上的本地VLA加速替代方案相比,在相近成功率下,ULAP每次成功回合的估计推理时间比ACT少49.2%、GPU能耗少51.0%;在相同成功率下,比SP-VLA少77.1%的时间和79.9%的能量。物理SO-101实验在已见与保留摆放位置下均保留基线成功率的95.2–100%,同时基于成功回合的调用次数和实测设备成本,估计减少推理时间47.9–58.0%、推理设备能耗52.1–62.5%。更快的响应还提升了动态任务成功率:在延迟感知的LIBERO-Safety仿真中,VLA-ULAP在两个任务上分别超过π₀.₅ 11.0和15.5个百分点,同时将VLA调用次数约减少一半。
cs.RO / 71 / 2609.18669

M$^3$P-R1: Reinforcement Learning for Large Language Model Guided Multi-Modal Motion Planning via MIP Code Generation

M$^3$P-R1:基于强化学习的大语言模型引导多模态运动规划——通过MIP代码生成
Sun, Xingpeng, Pan, Zherong, Cheng, Kai, Tang, Xindi, Bukhari, Syed Talha, Bera, Aniket
Abstract
Multi-Modal Motion Planning (M$^3$P) requires joint reasoning over continuous motions and discrete mode transitions, making it difficult to solve efficiently. For instance, a bipedal robot may walk to a target location and then use its arms to grasp an object. This scenario captures both mode transitions and continuous dynamics, yielding feasible paths that neither purely discrete nor continuous planners can handle. While Mixed-Integer Programming (MIP) offers a principled framework, constructing tractable formulations for non-convex problems is typically manual and domain-specific, especially in the approximate, discretization-based MIP regime needed for non-convex robotic tasks. We propose M$^3$P-R1, a reinforcement learning method that fine-tunes large language models (LLMs) to decompose M$^3$P tasks into MIP variables, constraints, and objectives. Instead of directly outputting answers, which are often prone to hallucination, the model generates executable Python code using MIP optimization libraries and constraint interfaces. This enables solver-backed execution for robust and verifiable solutions. Trained with an outcome-driven reward against the solver, M$^3$P-R1 learns to compose modality-level discretization primitives and synthesize cross-modal coupling constraints, producing executable MIP programs for complex M$^3$P tasks.
Chinese Translation
多模态运动规划(M$^3$P)需要对连续运动和离散模式转换进行联合推理,因此难以高效求解。例如,双足机器人可能需要行走至目标位置,然后用手臂抓取物体。这一场景同时涉及模式转换和连续动力学,其可行路径既无法由纯离散规划器处理,也无法由纯连续规划器求解。尽管混合整数规划(MIP)提供了一个有原则的框架,但为非凸问题构建易处理的公式通常需要人工完成且依赖特定领域知识,尤其是在非凸机器人任务所需的基于离散化的近似MIP方法中。我们提出M$^3$P-R1,这是一种强化学习方法,通过微调大语言模型(LLM),将M$^3$P任务分解为MIP变量、约束和目标。模型并不直接输出容易产生幻觉的答案,而是利用MIP优化库和约束接口生成可执行的Python代码,从而实现基于求解器的执行,以获得鲁棒且可验证的解。通过以求解器结果为导向的奖励进行训练,M$^3$P-R1学会组合模态级离散化原语并合成跨模态耦合约束,从而为复杂的M$^3$P任务生成可执行的MIP程序。
cs.RO / 72 / 2609.18685

WeaveRL: Weaving Reconstruction into Scene-Aware Fabrics for Perceptive Reinforcement Learning

WeaveRL:将重构编织进场景感知织物以实现感知强化学习
Steiner, Remo, Ramasamy, Vikram, Tingdahl, David, Mady, Sam, Van Wyk, Karl, Ratliff, Nathan, Lafuente, David Recasens, Pouya, Soha, Stuyck, Tuur, Millane, Alex
Abstract
Reinforcement learning allows robots to acquire complex skills, but producing policies for geometrically complex manipulation remains difficult. A promising approach is to learn on top of collision-avoidant controllers, such as geometric fabrics. However, these approaches have relied on static, hand-specified representations of the scene. Integrating active, online 3D perception into massively parallel RL training has so far been inaccessible. We introduce a GPU-accelerated method that reconstructs the scene as a collection of surfels across thousands of parallel simulation instances during active rollouts. This lets policies operate over sensor-derived, rather than hand-specified, geometry. On a suite of collision-dense manipulation tasks, our surfel fabrics enable policies to tackle geometrically complex scenes where primitive-based baselines fail, while maintaining sim-to-real transfer. Furthermore, policies learned with a scene-aware fabric are more robust to the introduction of novel geometry at test time, improving collision-free task completion under unseen obstacles from 35% to 61%. We release our reconstruction system, training code and test dataset to spur research in this direction.
Chinese Translation
强化学习使机器人能够习得复杂技能,但针对几何复杂操作生成策略仍然困难。一种有前景的方法是在避碰控制器(如几何织物/geometric fabrics)的基础上进行学习。然而,这些方法依赖于静态的、人工指定的场景表示。迄今为止,将主动式在线3D感知集成到大规模并行强化学习训练中仍难以实现。我们提出了一种GPU加速的方法,在主动滚动执行(active rollouts)过程中,跨数千个并行仿真实例将场景重构为面元(surfels)集合。这使策略能够基于传感器获取的几何信息而非人工指定的几何进行操作。在一组碰撞密集的操作任务套件上,我们的面元织物(surfel fabrics)使策略能够应对基于图元的基线方法失效的几何复杂场景,同时保持仿真到现实的迁移能力。此外,使用场景感知织物学习的策略在测试时对新增几何体具有更强的鲁棒性,将在未见障碍物下的无碰撞任务完成率从35%提升至61%。我们公开发布了重构系统、训练代码和测试数据集,以推动该方向的研究。
cs.RO / 73 / 2609.18700

Toward 3D Printable Non-Planar Electroadhesive Structures for Active Anchoring

面向主动锚定的三维可打印非平面电粘附结构研究
Atalla, Mostafa A., Derla, Carol R., Pericàs, Pep Canyelles, Krijnen, Gijs
Abstract
This paper investigates multi-material 3D printing as a method to fabricate non-planar structures with 3D-printed electrode patterns for electroadhesion. We printed flat electroadhesion pads as a planar benchmark and cylindrical pads as a non-planar demonstration, with conductive interdigitated electrodes 3D printed as part of the structure. Normal-force measurements showed voltage-controlled modulation in both geometries. At 3 kV, the flat pads generated approximately 0.11-0.13 N, while the cylindrical pads reached approximately 0.16 N at 5 N preload. These results demonstrate a step toward printed parts with built-in, electrically controlled adhesion and friction, beyond conventional planar electroadhesion pads.
Chinese Translation
本文研究了多材料3D打印作为一种制造非平面结构的方法,该结构带有3D打印的电极图案,用于电粘附。我们打印了平板电粘附垫作为平面基准,并打印了圆柱形垫作为非平面演示,其中的导电叉指电极通过3D打印直接集成于结构之中。法向力测量结果表明,两种几何形状均实现了电压可控的调制。在3 kV电压下,平板垫产生约0.11-0.13 N的力,而圆柱形垫在5 N预紧力下达到约0.16 N。这些结果表明,我们在实现内置电控粘附与摩擦的打印部件方面迈出了一步,超越了传统的平面电粘附垫。
cs.RO / 74 / 2609.18718

Calibrated Probabilistic Obstruction Reasoning with Vision-Language Models for Grasping in Clutter

基于视觉-语言模型的校准概率式遮挡推理用于杂乱场景抓取
Tran, Thanh-Tuan, Chu, Ngoc-Chien, Canh, Thanh Nguyen, Chong, Nak Young, Ha, Nguyen-Viet, HoangVan, Xiem
Abstract
Retrieving a target from clutter requires deciding whether to grasp the target, remove a blocker, or defer. Existing methods typically commit to a single obstruction graph or removal strategy, ignoring uncertainty across alternative scene interpretations. They also rely on miscalibrated vision-language model (VLM) predictions and can produce pairwise obstruction relations that are jointly inconsistent. Moreover, current approximations provide no guarantees about the impact of discarded hypotheses on the final decision. We propose CPOR-Grasp, a calibrated probabilistic obstruction-reasoning framework that propagates uncertainty from pairwise evidence to action decisions. CPOR-Grasp calibrates and fuses VLM, depth, and amodal-mask cues to estimate obstruction probabilities, induces a distribution over valid obstruction graphs, and marginalizes over these graphs to compute the likelihood that the target is accessible or that a given blocker should be removed. To make inference tractable, it retains only the highest-probability graphs and derives a total-variation bound on the discarded probability mass, enabling certified decisions, adaptive stopping, and principled deferral. On synthetic and real UNOBench scenes, CPOR-Grasp outperforms state-of-the-art baselines. Calibration error decreases from 0.1416 to 0.0185 on the Gemini Robotics backbone, while graph truncation matches exact inference on 99.74\% of decisions using 56 times fewer graphs. In real-world experiments, CPOR-Grasp achieves a 77.8\% average success rate, surpassing SOTA baselines.
Chinese Translation
从杂乱场景中检索目标需要决定是直接抓取目标、移除遮挡物,还是暂缓行动。现有方法通常仅依赖于单一的遮挡图或移除策略,忽略了不同场景解释之间的不确定性。这些方法还依赖于校准不足的视觉-语言模型(VLM)预测,并且可能产生联合不一致的成对遮挡关系。此外,当前的近似方法无法对被舍弃假设对最终决策的影响提供任何保证。我们提出了CPOR-Grasp,一个校准的概率式遮挡推理框架,能够将不确定性从成对证据传播到动作决策。CPOR-Grasp对VLM、深度和幅度掩码线索进行校准与融合,以估计遮挡概率,在有效的遮挡图上诱导出一个分布,并对这些图进行边缘化,从而计算目标可被抓取或应移除某个遮挡物的可能性。为使推理易于处理,该方法仅保留概率最高的若干遮挡图,并对被舍弃的概率质量推导出全变差界,从而实现可认证的决策、自适应停止和有原则的暂缓机制。在合成与真实UNOBench场景上,CPOR-Grasp优于最先进的基线方法。在Gemini Robotics骨干网络上,校准误差从0.1416降至0.0185,同时图截断在99.74%的决策上与精确推理结果一致,且所使用的图数量减少了56倍。在真实世界实验中,CPOR-Grasp取得了77.8%的平均成功率,超越了最先进的基线方法。
cs.RO / 75 / 2609.18732

PASSAGE: Scaling Scene-Aligned Motion Learning for Perceptive Humanoid Traversal in Cluttered Environments

PASSAGE:面向杂乱环境中感知型人形机器人通行的场景对齐运动学习扩展方法
Ma, Yuxuan, Zeng, Zicheng, Peng, Chunlin, Li, Zhoujian, Zhao, Zetong, Zhang, Zhikai, Lian, Yunrui, Xue, Han, Liang, Sikai, Zhu, Weiyi, Chen, Mulin, Lin, Chenghuai, Zeng, Jiayu, An, Yanwei, Zhang, Songan, Gu, Jiayuan, Wang, Jilong, Wang, Jingbo, Wang, He, Yi, Li
Abstract
Humanoid robots can step over, squeeze past, and duck under obstacles, but learning to select and coordinate these behaviors from onboard perception remains challenging. Many existing approaches rely on task-specific reinforcement-learning objectives or curated motion libraries, making broad behavioral coverage costly. We present PASSAGE, a perception-conditioned planner--tracker framework for humanoid traversal. Using virtual reality and inertial motion capture, we collect 100 h of scene-aligned human motion across 1,500 cluttered scenes. A conditional flow-matching planner generates short-horizon references from motion history, a local destination, and a robot-centric multi-layer elevation map, while a perceptive whole-body tracker executes them at 50 Hz with geometric feedback. Real-time chunking promotes inter-chunk consistency, and planner-side RL post-training under the frozen tracker further improves closed-loop performance. Without skill annotations or obstacle-specific policies, one planner--tracker pair selects and composes traversal behaviors across unseen geometries. In simulation, component ablations quantify the contribution of each stage. Across three independent training seeds, scaling captured data from 6 to 100 h increases mean contact-free success from 48.1% to 68.9% on held-out scenes, while the final model with validated scene augmentation reaches 70.3%. The fully onboard system integrates egocentric 3D LiDAR perception, online occupancy mapping, 6.25 Hz planning, and 50 Hz control on a Jetson AGX Orin; tests across 50 unseen physical layouts demonstrate traversal without prebuilt maps or offboard computation.
Chinese Translation
人形机器人可以跨过、挤过和俯身钻过障碍物,但如何基于机载感知来选择并协调这些行为仍然具有挑战性。许多现有方法依赖特定任务的强化学习目标或精选的运动库,使得实现广泛的行为覆盖代价高昂。我们提出了PASSAGE,一个面向人形机器人通行的感知条件化规划器—跟踪器框架。利用虚拟现实和惯性动作捕捉,我们在1,500个杂乱场景中采集了100小时与场景对齐的人体运动数据。一个条件流匹配规划器根据运动历史、局部目标点以及以机器人为中心的多层高程图,生成短时域参考轨迹;一个具备感知能力的全身跟踪器以50 Hz频率结合几何反馈执行这些轨迹。实时分块机制促进了各分块之间的一致性,在冻结跟踪器的前提下对规划器进行强化学习后训练,进一步提升了闭环性能。在无需技能标注或针对障碍物的专用策略的情况下,同一组规划器—跟踪器能够在未见过的几何结构中选择并组合通行行为。在仿真中,各组件的消融实验量化了每个阶段的贡献。在三个独立训练种子下,将采集数据从6小时扩展到100小时,使留出场景上的平均无接触成功率从48.1%提升至68.9%,而经过验证的场景增强后的最终模型达到了70.3%。完全机载的系统在Jetson AGX Orin上集成了第一视角3D激光雷达感知、在线占据栅格建图、6.25 Hz规划和50 Hz控制;在50个未见过的真实布局上的测试表明,该系统无需预建地图或离线计算即可完成通行。
cs.RO / 76 / 2609.18738

Active perception for robotic harvesting: 3D reconstruction and localisation of tomatoes hidden within clusters in a Mediterranean greenhouse

面向机器人采摘的主动感知:地中海温室中隐藏于果串内的番茄三维重建与定位
Cañadas-Aránega, Fernando, Border, Rowan, Moreno, José C., Blanco-Claraco, José L.
Abstract
Automating robotic harvesting in intensive agriculture within Mediterranean greenhouses requires overcoming significant challenges related to the geometric complexity of plants and occluded fruits. Although existing literature offers solutions targeting crops that grow in isolation (e.g., apples, sweet peppers, or peaches), the fundamental challenge lies in cluster-growing vegetables, where fixed sensors mounted on robotic systems fail to detect fruits hidden behind the visible surface. To address this limitation, this study presents a comprehensive pipeline for the 3D reconstruction and precise localization of each fruit within a cluster, including heavily occluded instances. The proposed methodology is structured into five sequential stages: i) point cloud acquisition using the AgriSEE Next Best View (NBV) active planner; ii) stochastic noise filtering via Statistical Outlier Removal (SOR); iii) surface classification and segmentation using Region Growing (RG); iv) isolation and recovery of occluded fruits through Density-Based Spatial Clustering of Applications with Noise (DBSCAN); and v) 3D pose estimation (position and orientation). This approach extracts the complete cluster geometry, ensuring the reliable identification of partially hidden tomatoes. Evaluated across multiple scenarios with varying occlusion levels within a simulation framework rigorously validated against real-world conditions, the system achieves a precision exceeding 90\%, an average recall of 82.8\%, and a mean Intersection over Union (mIoU) of 80.7\%. Furthermore, it demonstrates high repeatability in centroid estimation with a Root Mean Square Error (RMSE) of merely 4.2~mm, verifying its technical feasibility and high accuracy for autonomous harvesting operations.
Chinese Translation
在地中海温室的集约化农业中实现机器人采摘自动化,需要克服与植物几何复杂性和果实遮挡相关的重大挑战。尽管现有文献提供了针对孤立生长作物(如苹果、甜椒或桃子)的解决方案,但根本性挑战在于串生蔬菜,因为安装在机器人系统上的固定传感器无法检测到隐藏在可见表面之后的果实。为解决这一局限,本研究提出了一套完整的流水线,用于对果串中每个果实(包括严重遮挡的果实)进行三维重建与精确定位。所提出的方法分为五个连续阶段:i) 使用 AgriSEE 下一最优视点(Next Best View, NBV)主动规划器进行点云采集;ii) 通过统计离群点剔除(Statistical Outlier Removal, SOR)进行随机噪声滤波;iii) 使用区域生长(Region Growing, RG)进行表面分类与分割;iv) 通过基于密度的带噪声应用空间聚类(DBSCAN)实现被遮挡果实的隔离与恢复;v) 三维位姿估计(位置与朝向)。该方法能够提取完整的果串几何结构,确保对部分隐藏番茄的可靠识别。在一个经过真实环境条件严格验证的仿真框架中,在多种不同遮挡程度的场景下进行评估,该系统实现了超过90%的精确率、82.8%的平均召回率以及80.7%的平均交并比(mIoU)。此外,质心估计的均方根误差(RMSE)仅为4.2毫米,表现出高度的可重复性,验证了该技术在自主采摘作业中的技术可行性与高精度。
cs.RO / 77 / 2609.18752

QMSR: Query-Conditioned Mask-wise Expert Routing for Robust Open-Vocabulary Underwater Object Retrieval

QMSR:面向鲁棒开放词汇水下目标检索的查询条件掩码级专家路由
Zhang, Fuming, Huang, Dongyue, Wen, Junjie, Xie, Lihua
Abstract
Open-vocabulary object retrieval remains challenging in complex underwater environments. Although underwater image enhancement (UIE) can improve visual quality, fixed UIE strategies may even underperform the Raw representation in retrieval, indicating that enhancement should not be applied as a uniform preprocessing step. To address this problem, we propose \textbf{QMSR}, a query-conditioned mask-wise expert routing framework for underwater open-vocabulary retrieval. Specifically, QMSR selects one pretrained UIE expert for each query--candidate pair and predicts a continuous Raw--Expert fusion strength, enabling adaptive enhancement while preserving useful Raw semantics. During training, a privileged ranking oracle provides expert-selection and fusion-strength supervision, while an annealed soft-routing relaxation facilitates optimization of the hard Top-1 routing policy. Experiments show that QMSR improves NDCG@10 by 17.6\% over an image--query shared router, while consistently outperforming fixed UIE strategies and remaining effective on held-out query categories. These results demonstrate the effectiveness of query-conditioned and candidate-specific enhancement routing for underwater open-vocabulary retrieval.
Chinese Translation
在复杂的水下环境中,开放词汇目标检索仍然极具挑战性。尽管水下图像增强(UIE)可以改善视觉质量,但固定的UIE策略在检索任务中甚至可能不及原始(Raw)表示,这表明增强不应被作为统一的预处理步骤来应用。为解决这一问题,我们提出了QMSR,一个用于水下开放词汇检索的查询条件掩码级专家路由框架。具体而言,QMSR为每个查询—候选对选择一个预训练的UIE专家,并预测一个连续的Raw—Expert融合强度,从而在保留有用Raw语义的同时实现自适应增强。在训练过程中,特权排序监督信号为专家选择和融合强度提供监督,同时采用退火软路由松弛策略来促进硬Top-1路由策略的优化。实验表明,QMSR相比图像—查询共享路由器将NDCG@10提升了17.6%,并且持续优于固定的UIE策略,同时在保留的查询类别上依然有效。这些结果证明了查询条件和候选特定增强路由在水下开放词汇检索中的有效性。
cs.RO / 78 / 2609.18763

Gated Residual Body-Hand Coordination for Whole-Body Humanoid Teleoperation

面向全身人形机器人遥操作的门控残差身体-手部协同方法
Wu, Ruiming, Li, Shuang, Zhang, Liding, Knoll, Alois, Chen, Zhaopeng
Abstract
Whole-body humanoid teleoperation commonly combines a motion-tracking policy with a separate dexterous-hand retargeter. However, independently generated commands do not explicitly preserve body-hand geometric relations, leading to mismatches in relative wrist poses and fingertip positions during bimanual interaction. We present a gated residual coordination framework that keeps both modules frozen and applies bounded corrections to their outputs. A motion-conditioned action gate allocates correction authority across joint groups, while reference-geometry-dependent reward gates emphasize relevant interaction objectives during training. To establish the nominal body controller on Agile One, we introduce multi-pose morphology calibration that jointly estimates triaxial scales and effector-local offsets, together with staged motion dataset curation for training a SONIC-based tracker. The residual policy uses human motion references, initial commands, and robot proprioception without explicit object or contact observations. In simulation, it reduces wrist and fingertip geometry errors by 39.2-56.3% over direct composition on held-out GRAB motions, while preserving whole-body tracking on AMASS, with success rates of 89.03% without residual coordination and 89.29% with it. Ablations characterize the contributions of reward gating, adaptive correction authority, and separate body and hand correction heads.
Chinese Translation
全身人形机器人遥操作通常将运动跟踪策略与独立的灵巧手重定向器相结合。然而,独立生成的指令无法显式保持身体与手部之间的几何关系,导致双手交互过程中手腕相对位姿与指尖位置出现失配。我们提出了一种门控残差协同框架,该框架保持两个模块冻结不变,并对其输出施加有界修正。一个由运动条件驱动的动作门控在关节组之间分配修正权限,而依赖于参考几何的奖励门控则在训练过程中强调相关的交互目标。为了在 Agile One 平台上建立标称身体控制器,我们引入了多姿态形态标定方法,联合估计三轴缩放系数与末端执行器局部偏移量,并采用分阶段的运动数据集筛选来训练基于 SONIC 的跟踪器。残差策略仅需人体运动参考、初始指令和机器人本体感知,无需显式的物体或接触观测。在仿真中,在留出的 GRAB 动作上,该方法相比直接组合方式将手腕与指尖几何误差降低了 39.2%–56.3%,同时在 AMASS 数据集上保持了全身跟踪性能,不使用残差协同时的成功率为 89.03%,使用后为 89.29%。消融实验刻画了奖励门控、自适应修正权限以及身体与手部分离修正头各自的贡献。
cs.RO / 79 / 2609.18776

TRACER: Adaptive Multi-Robot Social Navigation via Joint Human-Response Prediction and Interaction-Aware Replanning

TRACER:基于联合人类响应预测与交互感知重规划的适应性多机器人社会导航
Hu, Lan, Liwang, Minghui, Zhu, Wenbo, Yi, Xinlei, Gong, Wei, Hong, Yiguang, Hosseinalipour, Seyyedali
Abstract
Multi-robot navigation in human-shared spaces is inherently interactive: coordinated robot motions influence how nearby entities respond, while those responses provide valuable information for subsequent robot decisions. However, existing methods typically address action-conditioned prediction, multi-robot planning, or online adaptation separately, and therefore lack a unified mechanism for modeling joint robot-entity interactions and adapting future decisions from executed interaction outcomes. To address this gap, we propose TRACER, a bi-directional receding-horizon framework that closes the loop between prediction and adaptation. TRACER evaluates candidate (i.e., alternative feasible future motion plans for the robot team) trajectories using a per-entity probabilistic response model that separates individual-robot effects from non-additive pairwise interactions; after executing the selected trajectory prefix, it updates persistent identity-bound beliefs over latent response modes using the synchronized observed responses. These updated beliefs then guide subsequent candidate evaluation under probabilistic safety and response-aware cost criteria. Experiments show that (i) TRACER more accurately captures non-additive multi-robot interaction effects than a capacity-matched additive predictor, (ii) persistent identity-consistent evidence improves response prediction and downstream replanning, and (iii) the complete TRACER framework improves collision-free completion over an independent-robot baseline on the SocialGym2 multi-robot social-navigation benchmark.
Chinese Translation
人机共享空间中的多机器人导航本质上是交互性的:机器人之间的协调运动会影响附近实体的响应,而这些响应又为机器人后续决策提供了宝贵的信息。然而,现有方法通常将动作条件预测、多机器人规划或在线适应分开处理,因此缺乏一个统一的机制来建模机器人与实体之间的联合交互,并根据已执行的交互结果调整未来决策。为填补这一空白,我们提出了TRACER,一个双向滚动时域框架,实现了预测与适应之间的闭环。TRACER使用一个按实体划分的概率响应模型来评估候选轨迹(即机器人团队备选的可行未来运动方案),该模型将单个机器人的效应与非可加性的成对交互效应分离;在执行所选轨迹的前缀后,它利用同步观测到的响应更新关于潜在响应模式的持久性、绑定身份的信念。这些更新后的信念随后在概率安全性和响应感知代价准则下指导后续候选轨迹的评估。实验表明:(i) 与容量匹配的可加性预测器相比,TRACER能更准确地捕捉非可加性的多机器人交互效应;(ii) 持久的身份一致性证据改进了响应预测和下游重规划;(iii) 在SocialGym2多机器人社会导航基准上,完整的TRACER框架相比独立机器人基线提升了无碰撞完成率。
cs.RO / 80 / 2609.18789

AdaGeoVLN: Selective Geometry Across Representation Depth and Navigation Time for Vision-Language Navigation

AdaGeoVLN:面向视觉语言导航的跨表示深度与导航时间的自适应选择性几何方法
Pham, Quan-Dung, Dao, Anh, Le, Danh Vinh, Pham, Nguyen Viet Tri, Nguyen, The-Anh, Dai, Zhirui, Chen, Yiyu, Le, Tuyen P., Nguyen, Truong, Nguyen, Quan
Abstract
Vision-language navigation requires aligning language with visual observations while maintaining spatial understanding over time. Geometry foundation models (GFMs) expose intermediate representations throughout their hierarchy, but how navigation policies should use these features and retain historical geometric evidence remains unresolved. We introduce \method{}, a streaming VLN framework that addresses these questions across \textbf{representation depth} and \textbf{navigation time}. Hierarchical GFM--VLM fusion couples earlier, intermediate, and later GFM representations to successive policy stages instead of repeatedly injecting a terminal feature. Navigation-aware GFM memory retains historical VGGT global-attention KV states according to instruction relevance, geometric confidence, and transition novelty under a bounded per-layer budget. Retained states provide geometric context for subsequent observations before fusion with the policy. Across R2R-CE and RxR-CE, \method{} achieves strong performance using a single RGB stream without additional navigation-specific external data. Controlled ablations show that multi-depth coupling substantially outperforms repeated terminal-feature injection at matched fusion locations. Bounded navigation-aware retention preserves navigation performance while considerably reducing GFM-KV memory relative to larger-memory temporal retention. These findings support jointly examining the geometric representations exposed to the policy and the historical evidence retained for future inference. Code will be released upon acceptance at https://humanoid-research.github.io/adageovln/.
Chinese Translation
视觉语言导航(Vision-Language Navigation, VLN)要求在将语言与视觉观测对齐的同时,随时间保持空间理解能力。几何基础模型(Geometry Foundation Models, GFMs)在其层级结构中暴露出多种中间表示,但导航策略应如何利用这些特征并保留历史几何证据仍悬而未决。我们提出了 AdaGeoVLN,一个跨表示深度和导航时间解决上述问题的流式 VLN 框架。层级化 GFM-VLM 融合将早期、中期和后期的 GFM 表示与相继的策略阶段相耦合,而非反复注入末端特征。导航感知的 GFM 记忆在每层有界预算下,根据指令相关性、几何置信度和转移新颖性保留历史的 VGGT 全局注意力 KV 状态。所保留的状态在后续观测与策略融合之前为其提供几何上下文。在 R2R-CE 和 RxR-CE 数据集上,AdaGeoVLN 仅使用单一 RGB 流,且无需额外的导航专用外部数据,即取得了优异的性能。受控消融实验表明,在相同融合位置下,多深度耦合显著优于反复的末端特征注入。有界的导航感知保留在保持导航性能的同时,相比更大内存的时间保留方式大幅减少了 GFM-KV 内存。这些发现支持对暴露给策略的几何表示与为未来推理保留的历史证据进行联合考察。代码将于论文录用后在 https://humanoid-research.github.io/adageovln/ 公布。
cs.RO / 81 / 2609.18813

Asymptotically Optimal Multi-Robot Task and Motion Planning

渐近最优的多机器人任务与运动规划
Duong, Thi Thuy Ngan, Yiu, Cheuk Tung Shadow, Shome, Rahul, Sung, Yoonchang
Abstract
Multi-robot task and motion planning (MR-TAMP) requires jointly reasoning about discrete task decisions and continuous collision-free motions of multiple interacting robots. Although asymptotically optimal algorithms have been developed for task and motion planning, extending these guarantees to the multi-robot setting introduces an important challenge: different task transitions may involve different subsets of robots and therefore impose constraints of different dimensions on the composite configuration space. Consequently, an asymptotically optimal planner must not only optimize motion within each task mode, but also ensure sufficient exploration of the different types of transitions connecting them. We characterize this transition structure and establish sufficient conditions for global asymptotic optimality in MR-TAMP, requiring persistent coverage of relevant transitions and asymptotically improving motion planning within connected feasible regions. Based on these conditions, we develop an efficient asymptotically optimal MR-TAMP algorithm that combines evolving individual-robot roadmaps with implicit tensor-product search, avoiding explicit construction of the composite roadmap. The planner further employs conditional transition sampling, lazy collision checking, and mode- and solution-level guidance to improve finite-time planning efficiency while retaining persistent exploration. The resulting framework provides asymptotic optimality guarantees for multi-robot manipulation while efficiently exploiting the structure of individual-robot motion planning.
Chinese Translation
多机器人任务与运动规划(MR-TAMP)需要联合推理离散任务决策与多个交互机器人的连续无碰撞运动。尽管任务与运动规划已有渐近最优算法,但将这些保证扩展到多机器人场景带来一个重要挑战:不同的任务转移可能涉及不同的机器人子集,因此在复合构型空间上施加不同维度的约束。因此,渐近最优的规划器不仅要在每个任务模式内优化运动,还必须确保对连接各模式的不同类型转移进行充分探索。我们刻画了这种转移结构,并建立了MR-TAMP全局渐近最优性的充分条件,即要求对相关转移的持续覆盖以及连通可行区域内运动规划的渐近改进。基于这些条件,我们提出了一种高效的渐近最优MR-TAMP算法,该算法将不断演化的单个机器人路线图与隐式张量积搜索相结合,避免了显式构建复合路线图。该规划器还采用条件转移采样、懒碰撞检测以及模式级和解决方案级引导来提升有限时间内的规划效率,同时保持持续探索。所得框架在充分利用单个机器人运动规划结构的同时,为多机器人操作提供了渐近最优性保证。
cs.RO / 82 / 2609.18819

SEAM: Submap-Anchored Evidence for Lifelong LiDAR Mapping under Trajectory Deformation

SEAM:面向轨迹形变下终身激光雷达建图的子地图锚定证据方法
Kim, Kyuwon
Abstract
We propose SEAM, a LiDAR-based lifelong mapping framework. Instead of relying on a single anchor spanning the entire session, SEAM generates evidence based on a trajectory optimized with submap-level anchors, and performs dynamic object removal and change detection. Through submap-level reprojection, the generated evidence remains usable even if the trajectory is subsequently modified by a new session, eliminating the need to recompute the entire process from scratch. SEAM suppresses geometrically unreliable inter-session loop edges using a DOP-based confidence measure. Suppressing unreliable loop edges prevents alignment errors. SEAM also uses a directional voxel-wise evidence model. The model accounts for occupancy patterns that vary with ray direction. Direction-aware evidence separates dynamic objects from environmental changes more precisely. Experiments on a real construction-site dataset and a long-term multi-session dataset show that SEAM achieves higher accuracy and faster processing than existing methods.
Chinese Translation
我们提出了SEAM,一种基于激光雷达的终身建图框架。SEAM不依赖于覆盖整个采集过程的单一锚点,而是基于由子地图级锚点优化得到的轨迹生成证据,并执行动态物体去除与变化检测。通过子地图级的重投影,即使轨迹随后被新的采集过程修改,所生成的证据仍然可用,从而无需从头重新计算整个流程。SEAM利用基于DOP(动态占据概率)的置信度度量来抑制几何上不可靠的跨会话回环边。抑制不可靠的回环边可防止配准误差。SEAM还采用了一种方向性的体素级证据模型,该模型能够考虑随光线方向变化的占据模式。方向感知的证据能够更精确地区分动态物体与环境变化。在真实建筑工地数据集和长期多会话数据集上的实验表明,SEAM比现有方法具有更高的精度和更快的处理速度。
cs.RO / 83 / 2609.18869

KINO: A Keyframe Interface for VLM Planning and Whole-Body Control in Humanoid Loco-Manipulation

KINO:一种用于人形机器人移动操作中VLM规划与全身控制的关键帧接口
Chen, Sitong, Zargarbashi, Fatemeh, Cheng, Jin, An, Tianxu, Coros, Stelian
Abstract
Humanoid loco-manipulation requires robots to interpret task instructions and scene semantics while executing coordinated whole-body motions. We propose a hierarchical framework that uses motion keyframes as an intermediate representation between Vision-Language Model (VLM) planning and Reinforcement Learning (RL) control. Each keyframe specifies a target whole-body robot pose and, when applicable, an object pose. Given a language instruction, scene observations, and execution feedback, the VLM selects successive task-relevant keyframes from a predefined library. The selected keyframes are retargeted to the current scene to account for object poses and dimensions. A keyframe-conditioned whole-body policy then generates joint-level actions to reach these goals. We introduce a saliency-based keyframe sampling strategy for low-level policy training that improves end-to-end task success rate from 44% to 92% when using sparse VLM keyframes. We evaluate our framework on object pickup, transport, and placement tasks in simulation and on a Unitree G1 humanoid. The system successfully performs both one- and two-handed manipulation and generalises to placement locations beyond the training reference data.
Chinese Translation
人形机器人的移动操作(loco-manipulation)要求机器人在执行协调的全身运动的同时理解任务指令和场景语义。我们提出了一个分层框架,该框架使用运动关键帧作为视觉语言模型(VLM)规划与强化学习(RL)控制之间的中间表示。每个关键帧指定一个目标全身机器人姿态,以及(在适用时)一个物体姿态。给定语言指令、场景观测和执行反馈,VLM从预定义的关键帧库中选择连续的任务相关关键帧。所选关键帧会根据当前场景中的物体姿态和尺寸进行重定向(retarget)。随后,一个以关键帧为条件的全身策略生成关节级别的动作以达成这些目标。我们提出了一种基于显著性的关键帧采样策略用于底层策略训练,在使用稀疏VLM关键帧时,将端到端任务成功率从44%提升至92%。我们在仿真环境中以及Unitree G1人形机器人上,针对物体抓取、搬运和放置任务对该框架进行了评估。该系统能够成功完成单手和双手操作,并能泛化到训练参考数据之外的放置位置。
cs.RO / 84 / 2609.18881

Body-Motion Control of a Simulated Aerial Swarm from a First-Person View

第一人称视角下仿真空中集群的身体运动控制
Chen, Yang, Giannoli, Darius, Floreano, Dario
Abstract
First-person-view (FPV) teleoperation of aerial swarms requires an operator to coordinate collective translation, viewing direction, and formation spacing. We present an upper-body interface that maps torso inclination, hand position, and head rotation to five continuous command dimensions. Neutral postures and motion ranges are calibrated for each participant. In a within-subject study, 14 participants navigated a simulated 15-agent swarm through three-dimensional obstacle courses using this interface and a conventional transmitter. Body-motion control reduced completion time by 19.4% and centroid path length by 7.0%, and increased path directness. Delivered-command variation was 88.8% lower, and concurrent command changes were more frequent. These command measures characterize the complete interfaces, which differed in calibration and filtering. No differences were detected in gate-centering error, collection yield, crash or disconnection counts, overall workload, or usability. All participants reported higher physical demand with body-motion control. The implemented interface therefore improved FPV navigation efficiency at the cost of greater physical demand.
Chinese Translation
空中集群的第一人称视角(FPV)遥操作要求操作员协调集群的整体平移、观察方向和编队间距。我们提出了一种上半身交互界面,将躯干倾斜、手部位置和头部旋转映射到五个连续的指令维度。每位参与者的中性姿态和运动范围均经过校准。在一项被试内研究中,14名参与者使用该界面和传统遥控器,操控由15个智能体组成的仿真集群穿越三维障碍赛道。身体运动控制使完成时间缩短了19.4%,质心路径长度减少了7.0%,并提高了路径的直线性。指令传递的变化量降低了88.8%,并发指令变更更为频繁。这些指令指标表征的是完整界面的表现,二者在校准和滤波方面存在差异。在穿越门居中误差、收集收益、碰撞或断连次数、总体工作负荷以及可用性方面均未检测到差异。所有参与者均报告身体运动控制带来更高的体力需求。因此,所实现的界面以更高的体力需求为代价,提升了FPV导航效率。
cs.RO / 85 / 2609.18893

SOL-SLAM: Inverse Compositional Gauss-Newton Direct Registration for Fast Sonar-Only Local SLAM

SOL-SLAM:用于快速纯声呐局部SLAM的逆向组合高斯-牛顿直接配准方法
Jakkala, Kalvik, O'Kane, Jason
Abstract
Autonomous underwater navigation typically relies on complex and expensive multi-modal sensor suites designed to prioritize global Simultaneous Localization and Mapping (SLAM) accuracy. However, local reactive behaviors such as coarse navigation and obstacle avoidance require only local consistency---a capability that should be feasible using only a Forward-Looking Sonar (FLS), yet remains largely unaddressed, leaving a critical gap in FLS-only local SLAM. Moreover, existing acoustic SLAM frameworks predominantly rely on sparse feature extraction methods that discard substantial portions of the already information-sparse acoustic returns. To overcome these limitations, this work introduces a dense direct registration approach that aligns full acoustic intensity scans to a recursively updated local map. Real-time execution is achieved via an Inverse Compositional Gauss-Newton optimization strategy that minimizes computational overhead. Experimental evaluations show that this dense method yields significant improvements on translation error compared to sparse keypoint baselines, maintaining stable sub-meter tracking precision over wide displacement gaps. Moreover, this approach delivers odometry performance comparable to multi-sensor fusion pipelines (FLS, DVL, and IMU), bypassing expensive payload dependencies in feature-rich environments. We validate real-world applicability through AUV field trials, running the full local SLAM approach onboard an embedded, resource-constrained computer.
Chinese Translation
自主水下导航通常依赖于复杂且昂贵的多模态传感器套件,其设计以全局同步定位与建图(SLAM)的精度为优先。然而,诸如粗导航和避障等局部反应式行为只需要局部一致性——这一能力本应仅凭前视声呐(FLS)即可实现,但目前仍基本未被研究,在纯FLS局部SLAM领域留下了关键空白。此外,现有的声学SLAM框架主要依赖稀疏特征提取方法,丢弃了本已信息稀疏的声学回波中的大量信息。为克服这些局限,本工作提出了一种稠密直接配准方法,将完整的声学强度扫描与递归更新的局部地图进行对齐。通过采用逆组合高斯-牛顿优化策略,最小化计算开销,实现了实时执行。实验评估表明,与稀疏关键点基线相比,该稠密方法在平移误差方面取得了显著改进,并在较大位移间隔下保持稳定的亚米级跟踪精度。此外,该方法在特征丰富的环境中可提供与多传感器融合管线(FLS、DVL和IMU)相当的里程计性能,从而避免了对昂贵载荷的依赖。我们通过AUV实海试验验证了该方法的实际适用性,在嵌入式资源受限计算机上运行了完整的局部SLAM方案。
cs.RO / 86 / 2609.18900

Examining the Difference in Human Behavior Between Virtual and Real-World Human-Robot Teaming

考察虚拟环境与真实环境中人机协作行为差异
Dallas, Sean, Getachew, Absalat, AbuHijleh, Motaz, Macklem-Zabel, Andrea, Zytko, Douglas, Brudnak, Mark, Louie, Wing-Yue Geoffrey
Abstract
Prototyping and evaluating human-robot teaming (HRT) scenarios in the real-world is costly. Virtual simulation of HRT scenarios has been adopted as an alternative to conducting user studies in the real-world to investigate user perceptions, behaviors, and performance during human-robot interactions. The consistency of human behavior between the real and virtual-worlds is integral to the validity of utilizing such virtual experimentation. This paper presents a user study examining the difference in human behavior between an HRT scenario conducted in a virtual vs real environment. We employed mixed-methods to examine team performance and human factors during the HRT. Our quantitative results showed a significant difference in workload between the two modalities. Our qualitative analysis expanded on the quantitative results and found differences in participants' strategy, their mental model of robots, and the type of trust they had for robots between modalities.
Chinese Translation
在真实环境中构建和评估人机协作(HRT)场景的成本很高。虚拟仿真的人机协作场景已被采用作为真实环境用户研究的替代方案,用于研究人机交互过程中用户的感知、行为和表现。真实世界与虚拟世界中人类行为的一致性对于此类虚拟实验的有效性至关重要。本文提出了一项用户研究,考察在虚拟环境与真实环境中进行人机协作场景时人类行为的差异。我们采用混合方法来考察人机协作过程中的团队绩效和人为因素。定量结果显示,两种模态之间的工作负荷存在显著差异。定性分析进一步拓展了定量结果,发现参与者在策略、对机器人的心理模型以及对机器人的信任类型方面,在不同模态之间存在差异。
cs.RO / 87 / 2609.18910

CaSCo: Cascade-Aware Soft-Collision Motion Planning

CaSCo:级联感知的软碰撞运动规划
Kumar, Shivaram, Liu, Gaoyuan, Sung, Yoonchang
Abstract
Conventional motion planning treats collision as a binary constraint, although contact with different objects can have drastically different consequences. A robot may safely brush against a cardboard box while even minor contact with a glass, laptop, or unstable object may be undesirable. Moreover, a direct robot--object collision can move the contacted object and trigger secondary object--object collisions, making the risk of a motion depend on the physical evolution of the scene rather than only on the robot's geometric path. We present CaSCo, a cascade-aware soft-collision motion planning framework in which a vision-language or language model assigns semantic risk to objects and a physics simulator predicts the consequences of candidate robot motions. CaSCo searches for a path that minimizes the total semantic risk of the unique objects displaced either directly by the robot or indirectly through cascaded collisions. Because collisions change the environment, we augment roadmap states with the predicted object arrangement and the set of objects whose risk has already been incurred. We develop an optimal graph-search algorithm with an admissible and consistent cascade-relaxed heuristic and caching and pruning mechanisms for efficient search. Experiments in cluttered manipulation environments evaluate semantic risk, cascade reasoning, planning efficiency, and real-robot operation.
Chinese Translation
传统运动规划将碰撞视为二元约束,然而与不同物体发生接触可能产生截然不同的后果。机器人可以安全地擦过一个纸箱,而即使与玻璃杯、笔记本电脑或不稳定物体的轻微接触也可能是不可取的。此外,机器人与物体的直接碰撞可能移动被接触的物体并触发次级的物体间碰撞,使得一次运动的风险取决于场景的物理演化,而不仅仅是机器人的几何路径。我们提出了CaSCo,一个级联感知的软碰撞运动规划框架,其中视觉-语言模型或语言模型为物体分配语义风险,物理仿真器预测候选机器人运动的后果。CaSCo搜索一条路径,使直接被机器人移动或通过级联碰撞间接移动的不同物体所引起的总语义风险最小化。由于碰撞会改变环境,我们用预测的物体排列以及风险已经产生的物体集合来增强路线图状态。我们开发了一种具有可采纳且一致的级联松弛启发式的最优图搜索算法,并辅以缓存和剪枝机制以实现高效搜索。在杂乱操作环境中的实验评估了语义风险、级联推理、规划效率以及真实机器人运行。
cs.RO / 88 / 2609.18930

Learning Holistic Whole-Body Loco-Manipulation with a Bipedal Mobile Manipulator

基于双足移动操作机器人的全身整体移动-操作学习
Chen, Zhongyu, Nai, Yuxuan, Chen, Qian, Zhu, Yidong, Jing, Chen, Wang, Qihan, Li, Xudong, Li, Zhizhan, Chang, Leixin, Yang, Liangjing, Chen, Hua
Abstract
Bipedal loco-manipulation enables robots to interact with objects beyond the nominal workspace of their arms by coordinating locomotion and manipulation. Realizing this capability requires a low-level whole-body controller that translates task-level manipulation goals into coordinated arm and leg motions while maintaining balance. We present a unified whole-body controller trained with reinforcement learning that directly maps 6-DoF end-effector targets to coordinated actions for the bipedal base and robotic arm. Given only an end-effector target, the learned controller autonomously coordinates reaching, postural adaptation, and stepping without explicit base-velocity or footstep commands. A reward-gating strategy regulates the trade-offs among end-effector tracking, locomotion, and balance during training, while a temporal context estimator combines windowed Transformer encoding, recurrent GRU memory, and auxiliary dynamics prediction to extract dynamics-relevant information from observation history. Real-robot experiments demonstrate that the same controller supports reaching, postural adaptation, and stepping under commands from VR teleoperation, a learned diffusion policy, and scripted trajectories, providing a common end-effector interface for diverse manipulation tasks.
Chinese Translation
双足移动-操作使机器人能够通过协调运动与操作,与手臂标称工作空间之外的物体进行交互。实现这一能力需要一个低层全身控制器,将任务级操作目标转化为协调的手臂与腿部运动,同时保持平衡。我们提出了一种通过强化学习训练的统一全身控制器,可直接将6自由度末端执行器目标映射为双足底座与机械臂的协调动作。仅给定末端执行器目标,学习得到的控制器即可自主协调够取、姿态调整与迈步,无需显式的底座速度或落脚点指令。奖励门控策略在训练过程中调节末端执行器跟踪、运动与平衡之间的权衡,同时时序上下文估计器结合窗口化Transformer编码、循环GRU记忆和辅助动力学预测,从观测历史中提取与动力学相关的信息。真机实验表明,同一控制器可在VR遥操作、学习得到的扩散策略以及脚本轨迹的指令下支持够取、姿态调整与迈步,为多种操作任务提供了统一的末端执行器接口。
cs.RO / 89 / 2609.18970

"What's going to happen after I'm gone?": Parent Perspectives on Technology in Supporting Independent Living for Adults with Intellectual Disabilities

“我离开之后会发生什么?”:家长对利用技术支持智力障碍成年人独立生活的看法
Tyshka, Alexander, Macklem-Zabel, Andrea, Getachew, Absalat, Chen, Foong Ling, Louie, Wing-Yue Geoffrey
Abstract
Adults with intellectual and developmental disabilities (IDD) are increasingly transitioning from family homes towards semi-inde\-pendent living. As parents hand off the role of primary caregiver, they face numerous challenges in arranging consistent and quality support. Our research centers on understanding these caregiving routines. By focusing on the unique lived experience of parents, who possess extensive explicit and tacit knowledge of their adult child's requirements, we aim to map the management of care they provide. This foundational understanding is essential to identifying how assistive technologies can effectively serve a key role in supporting adults with IDD in this transition. In this work, we interviewed 16 parents of adults with IDD beginning this transition to understand: 1) how they currently provide support for daily living and what makes their support effective; 2) what their experiences and perceptions are regarding the use of technology; and 3) how they envision assistive technology successfully integrating with their adult child's new home or with other supports to promote independence. Our thematic analysis produced four themes: common modes of support, inside the routine to support growth, the fragility of continuity of care across transitions, and participants' perceptions and experiences of assistive technology. Building on these findings, we derive four design principles for growth-oriented assistive technology. These include pre-transition onboarding to capture caregiver tacit knowledge; structured scaffolding towards long-term growth; adaptive sensing that responds to day-to-day variability; and customization for the individual balanced with consistency for the care network.
Chinese Translation
患有智力与发展性障碍(IDD)的成年人正日益从家庭生活过渡到半独立生活。随着父母移交主要照护者的角色,他们在安排持续且高质量的照护支持方面面临诸多挑战。我们的研究旨在理解这些照护惯例。通过关注父母这一群体的独特生活经验——他们对成年子女的需求拥有大量显性和隐性知识——我们致力于梳理他们所提供的照护管理。这一基础性理解对于识别辅助技术如何在这一过渡中有效发挥关键作用、支持IDD成年人至关重要。在本研究中,我们访谈了16位正处于这一过渡阶段的IDD成年人的父母,以了解:1)他们目前如何为子女的日常生活提供支持,以及其支持为何有效;2)他们对技术使用的经验和看法;3)他们如何设想辅助技术能够成功地与成年子女的新居或其他支持资源相整合,以促进其独立性。我们的主题分析得出了四个主题:常见的支持模式、促进成长的日常惯例内部机制、过渡期照护连续性的脆弱性,以及参与者对辅助技术的看法与经验。基于这些发现,我们提出了面向成长的辅助技术的四项设计原则,包括:过渡前的入职引导以获取照护者的隐性知识;面向长期成长的结构化支架;响应日常变化的自适应感知;以及兼顾个体定制化与照护网络一致性的设计。
cs.RO / 90 / 2609.19012

Information-Based Trajectory Planning for Spacecraft-to-Spacecraft Tracking and Navigation in Cislunar Space

基于信息的地月空间航天器间跟踪与导航轨迹规划
Wolf, Trevor N., Jones, Brandon A., McMahon, Jay W.
Abstract
We present a trajectory planning method that balances information collection and control effort to improve cislunar spacecraft-to-spacecraft absolute tracking. Expanding use of cislunar space requires alternative navigation and tracking procedures that minimize reliance on ground-based support. Among efforts to address this need, spacecraft-to-spacecraft tracking exploits nonlinear dynamical tracers encoded in a series of relative measurements to infer absolute states of both an observer and a target. The geometry between spacecraft operating under this mode can significantly influence tracking performance. This work considers this geometrical impact by designing observer trajectories that jointly balance control effort and an information-theoretic quantification of the expected spacecraft-to-spacecraft tracking performance. Our methods are designed for multiple low-thrust observation platforms of various sensing modalities and can incorporate multiple space object targets in planning. By leveraging information gain in the optimal control problem, we report almost an order of magnitude improvement in the expected navigation and tracking errors for an optical observer operating in a Distant Retrograde Orbit (DRO). This work demonstrates the feasibility and value of information-optimal low-thrust spacecraft trajectory design for current and upcoming cislunar missions.
Chinese Translation
我们提出了一种在信息采集与控制消耗之间进行权衡的轨迹规划方法,以改进地月空间中航天器间的绝对跟踪。地月空间应用的不断扩展需要尽量减少对地面支持依赖的替代性导航与跟踪方案。在应对这一需求的相关研究中,航天器间跟踪利用一系列相对测量中编码的非线性动力学示踪特性来推断观测者与目标航天器的绝对状态。在该模式下运行的航天器之间的几何构型会显著影响跟踪性能。本工作通过设计观测者轨迹来考虑这种几何影响,同时联合权衡控制消耗与对航天器间预期跟踪性能的信息论量化度量。我们的方法面向采用多种传感方式的多个低推力观测平台,并可在规划中纳入多个空间目标。通过在最优控制问题中利用信息增益,我们报告了在远距离逆行轨道(Distant Retrograde Orbit, DRO)中运行的光学观测者的预期导航与跟踪误差近乎一个数量级的改善。这项工作展示了信息最优低推力航天器轨迹设计对于当前及未来地月空间任务的可行性与价值。
cs.RO / 91 / 2609.19040

Learning to Stack: Cube-Stacking Imitation Learning from Virtual Reality Demonstrations

学习堆叠:基于虚拟现实示范的立方体堆叠模仿学习
Reizian, Gryffin, Dowdy, Jordan, Vaz, Jean Chagas
Abstract
Imitation learning is attractive for robot manipulation, but collecting demonstrations remains a bottleneck for multi-stage tasks requiring repeated scene resets. This work presents a virtual-reality data-collection pipeline for cube-stacking with a custom 5-DoF arm in NVIDIA Isaac Sim and Isaac Lab. Using an HTC Vive Pro 2, Manus Quantum gloves, and OpenXR, an operator provides SE(3) end-effector commands to generate task demonstrations. The proposed framework separates demonstration collection from dataset construction by replaying recorded trajectories, converting task-space commands into joint-space actions, and re-rendering demonstrations with updated sensor or state configurations. This allows previously collected demonstrations to be reused for new observation and action spaces without repeating human teleoperation. The task requires stacking the red cube on the blue cube and the green cube on the red cube, with randomized cube placement. In 30 minutes, 200 virtual demonstrations were collected, compared with 45 real-world demonstrations, and Isaac Mimic generated 100 additional samples. A behavior-cloning policy was trained from the virtual demonstrations using LeRobot-style dual-camera observations and evaluated in simulation.
Chinese Translation
模仿学习对机器人操作极具吸引力,但对于需要反复重置场景的多阶段任务,收集示范数据仍然是瓶颈。本工作在 NVIDIA Isaac Sim 和 Isaac Lab 中,针对自定义的5自由度机械臂,提出了一套面向立方体堆叠任务的虚拟现实数据采集流程。操作者借助 HTC Vive Pro 2、Manus Quantum 数据手套和 OpenXR 提供 SE(3) 末端执行器指令,从而生成任务示范。所提出的框架将示范采集与数据集构建分离,通过回放已记录的轨迹、将任务空间指令转换为关节空间动作,并以更新后的传感器或状态配置重新渲染示范,使先前收集的示范能够被复用于新的观测和动作空间,而无需重复人工遥操作。任务要求将红色立方体堆叠在蓝色立方体上,再将绿色立方体堆叠在红色立方体上,且立方体的放置位置是随机的。在30分钟内共收集了200条虚拟示范,相比之下真实环境仅收集了45条,此外 Isaac Mimic 还额外生成了100个样本。基于虚拟示范,使用 LeRobot 风格的双相机观测训练了行为克隆(behavior cloning)策略,并在仿真环境中进行了评估。
cs.RO / 92 / 2609.19041

Loco-Loco-RL: Low-Cost Terrain Mapping for Humanoid Locomotion with Reinforcement Learning

Loco-Loco-RL:基于强化学习的低成本人形机器人运动地形感知与建图
Dowdy, Jordan, Reizian, Gryffin, Vaz, Jean Chagas
Abstract
Informative terrain perception is important for robust reinforcement learning policies in humanoid locomotion. Still, common sensors such as depth cameras and LiDARs incur high cost, power, and processing overhead while often producing redundant, high-resolution data. This work uses a low-cost time-of-flight sensor to provide a compact 3D local terrain representation for humanoid locomotion. To efficiently use this sparse exteroceptive input, we introduce a token-compressed temporal transformer policy. Proprioceptive and terrain observations are tokenized and processed by a self-attention multi-head transformer to capture within-timestep relationships between observation terms. The attended tokens are then compressed through an MLP-based latent-space token compression module before being stored in a rolling 15-timestep history. A second cross-attention multi-head transformer extracts temporal locomotion features from this compact history for policy learning. By compressing tokens before temporal aggregation, the architecture preserves important terrain-observation structure while limiting the dimensional growth of attention over observation histories. We validate our method through sim-to-real transfer on physical hardware using a terrain-based locomotion benchmark, demonstrating robust humanoid terrain walking with low-cost local terrain sensing.
Chinese Translation
丰富的地形感知对于人形机器人运动中鲁棒的强化学习策略至关重要。然而,深度相机和激光雷达(LiDAR)等常见传感器成本高、功耗大、处理开销高,且往往产生冗余的高分辨率数据。本工作使用一种低成本飞行时间(time-of-flight)传感器,为人形机器人运动提供紧凑的三维局部地形表征。为了高效利用这种稀疏的外部感知输入,我们提出了一种令牌压缩的时序Transformer策略。本体感知和地形观测被令牌化后,由自注意力多头Transformer处理,以捕获单个时间步内各观测项之间的关系。随后,经过注意力处理的令牌通过一个基于MLP的潜空间令牌压缩模块进行压缩,再存入一个长度为15个时间步的滚动历史中。第二个交叉注意力多头Transformer从这个紧凑的历史中提取时序运动特征,用于策略学习。通过在时序聚合之前压缩令牌,该架构在保留重要的地形观测结构的同时,限制了注意力在观测历史上随时间的维度增长。我们通过基于地形的人形运动基准测试,在真实硬件上进行了仿真到现实的迁移验证,展示了利用低成本局部地形感知实现鲁棒的人形机器人地形行走。
cs.RO / 93 / 2609.19080

ElastiQP: An Always-Feasible QP Solver for Constrained Robot Control

ElastiQP:一种用于约束机器人控制的始终可行的二次规划求解器
Morton, Daniel, Arrizabalaga, Jon, Manchester, Zachary, Pavone, Marco
Abstract
As robot capabilities increase, quadratic programming (QP)-based controllers must account for a similarly increasing number of constraints to ensure safe, reliable operation. Yet, with each added constraint, this introduces more chances of momentary conflict: in which case, a QP solver that returns an "infeasible" status leaves the controller with nothing to execute. To address this, we introduce ElastiQP, a modified dual active-set QP solver that relaxes every inequality constraint with an exact, per-constraint l1 penalty while keeping equality constraints (dynamics) hard. Notably, ElastiQP does so by folding the slack variables into the solver analytically, maintaining a constant size of the condensed linear system. On a suite of robot control benchmarks, ElastiQP achieves microsecond-level performance, matching or outperforming leading modern solvers on feasible problems. On infeasible problems, ElastiQP handles these gracefully, confining violations to strictly the conflicting inequality terms, returning a usable solution up to 40x faster than the best alternative solvers. ElastiQP is available as an open-source C++ header-only library, with Python and JAX interfaces, at https://github.com/StanfordASL/elastiqp.
Chinese Translation
随着机器人能力的不断提升,基于二次规划(QP)的控制器必须考虑数量同样日益增长的约束,以确保安全可靠的运行。然而,每增加一个约束,都会带来更多出现瞬时冲突的可能性:在这种情况下,返回“不可行”状态的QP求解器将使控制器无指令可执行。为解决这一问题,我们提出了ElastiQP,一种改进的对偶有效集QP求解器,它通过精确的逐约束l1惩罚来松弛所有不等式约束,同时保持等式约束(动力学)为硬约束。值得注意的是,ElastiQP通过解析方式将松弛变量折叠进求解器,从而保持压缩后线性系统的规模恒定。在一套机器人控制基准测试中,ElastiQP实现了微秒级的性能,在可行问题上与领先的主流求解器相当甚至更优。在不可行问题上,ElastiQP能够优雅地处理这些情况,将违反严格限制在冲突的不等式项上,并以比最佳替代求解器快达40倍的速度返回可用解。ElastiQP以开源C++仅头文件库的形式提供,并附带Python和JAX接口,网址为https://github.com/StanfordASL/elastiqp。
cs.RO / 94 / 2609.19104

rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference

rMuscle:面向高效视觉-语言-动作模型推理的机器人肌肉记忆机制
Zhou, Kaijun, Li, Zhiyang, Chen, Le, Gu, Jinyu
Abstract
Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs tractable for current policies. Vision-Language-Action (VLA) models now dominate as the policy paradigm for these robots. The inference latency of VLA models directly affects robot responsiveness and motion smoothness. However, existing VLA inference frameworks do not fully exploit the characteristics of embodied workloads or account for the distinct bottlenecks across different stages of VLA inference. In this paper, we first characterize embodied workloads and identify substantial task similarity across repeated robot executions. We further find that such similarity extends beyond observations and action trajectories to internal model states. Drawing on these observations, we present rMuscle, a real-time VLA inference framework inspired by human muscle memory. It exploits cross-execution similarity through a dual-phase muscle-memory cache. The Context Cache reuses visual-token outputs to reduce computation, while the Action Cache reuses neuron activation patterns to reduce weight accesses. We keep both the cache memory footprint and access overhead low through online cache recomputation, sliding-window cache retrieval, and mask sharing across consecutive denoising steps. rMuscle achieves 1.29-1.42X speedup on RTX 4090 and Jetson Thor across LIBERO, RoboTwin, and physical manipulation tasks, while maintaining the original success rates on real-world robots.
Chinese Translation
工厂作业是具身智能(Embodied AI)一个颇具前景的早期应用场景:将重复性手工工作分配给机器人具有明确的经济回报,而结构化的工作站也使这些任务对当前策略而言更易处理。视觉-语言-动作(Vision-Language-Action, VLA)模型目前已成为此类机器人策略范式的主流。VLA模型的推理延迟直接影响机器人的响应速度和运动平滑性。然而,现有的VLA推理框架既未充分利用具身工作负载的特性,也未考虑VLA推理不同阶段各自的瓶颈。在本文中,我们首先对具身工作负载进行了特性分析,发现机器人重复执行中存在显著的任务相似性。我们进一步发现,这种相似性不仅体现在观测和动作轨迹上,还延伸至模型内部状态。基于这些观察,我们提出了rMuscle,一个受人类肌肉记忆启发的实时VLA推理框架。它通过双阶段肌肉记忆缓存来利用跨执行相似性:上下文缓存(Context Cache)复用视觉token输出以减少计算量,而动作缓存(Action Cache)复用神经元激活模式以减少权重访问。我们通过在线缓存重计算、滑动窗口缓存检索以及跨连续去噪步骤的掩码共享,将缓存的内存占用和访问开销保持在较低水平。rMuscle在RTX 4090和Jetson Thor上,于LIBERO、RoboTwin及真实物理操作任务中实现了1.29-1.42倍的加速,同时在真实机器人上保持了原有的成功率。
cs.RO / 95 / 2609.19137

Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation

畅想接触之声:利用视频与音频生成实现零样本力感知操作与数据生成
Ji, Guanhua, Li, Tianyu, Suh, Dayoon, Zhang, Yuqian, Zhang, Boyan, Figueroa, Nadia
Abstract
Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories that lack force information, causing failures in contact-rich tasks where appropriate contact forces are essential for success. In this work, we explore augmenting generated video with audio to shape a bounded, time-varying desired-force profile using the loudness of generated contact sounds. We present a pipeline that jointly leverages generated video and audio to derive motion trajectories and corresponding desired-force profiles from a structured natural-language task prompt. We execute these force-aware trajectories on a Franka Panda robot using a closed-loop force regulator that tracks the audio-shaped force profile during contact. We evaluate our pipeline on multiple tasks that require making contact and demonstrate successful manipulation where a kinematic-only baseline fails. We also use the pipeline as a data generation engine to train policies that achieve the tasks in a closed-loop manner. Project website, videos, and dataset: https://dreamingcontactsound.github.io/
Chinese Translation
视频生成领域的最新进展使机器人能够从生成的视频中学习操作轨迹。然而,这些方法产生的轨迹仅包含运动学信息,缺乏力信息,导致在依赖合适接触力才能成功的接触密集型任务中表现不佳。在本工作中,我们探索在生成视频的基础上引入音频,利用生成的接触声音的响度来构建有界的、随时间变化的期望力曲线。我们提出了一个流程,从结构化的自然语言任务提示出发,联合利用生成的视频和音频来推导运动轨迹及相应的期望力曲线。我们在 Franka Panda 机器人上执行这些力感知轨迹,采用闭环力调节器在接触过程中跟踪由音频塑造的力曲线。我们在多个需要建立接触的任务上对该流程进行了评估,结果表明其能成功完成操作,而仅基于运动学的基线方法则失败。我们还将该流程用作数据生成引擎,训练能够以闭环方式完成任务的策略。项目网站、视频与数据集:https://dreamingcontactsound.github.io/