← Back to Index
Daily Research Digest

arXiv Papers

2026-09-16
62
Papers
1
Categories
62
Translated
收藏清单 0
机器人学 (Robotics)
62
cs.RO / 1 / 2609.16040

Bi-MoDe: Bilateral Control-based Imitation Learning via Modifier-Conditioned Decoding for Modulation of Execution Speed and Contact Intensity

Bi-MoDe:基于双边控制的模仿学习,通过修饰符条件化解码实现执行速度与接触强度的调节
Kobayashi, Takumi, Kobayashi, Masato, Uranishi, Yuki
Abstract
Bilateral control-based imitation learning captures both position and force information, making it well suited to contact-rich manipulation. However, existing approaches provide limited means for an operator to specify how a learned task should be executed at inference time, such as slowly or quickly, gently or firmly. We propose Bi-MoDe, a modifier-conditioned decoding framework that injects a constrained latent into every layer of the Transformer action decoder via adaLN-Zero, allowing behavioral directives to directly influence action-chunk generation. We evaluate the method on a real-world whiteboard wiping task with combinations of temporal and physical modifiers. Bi-MoDe improves physical directive following over the action-chunking baseline while maintaining comparable temporal control. An ablation further shows that decoder conditioning and latent-space composition interact, and that their combination is important for accurate physical directive following. Additional material is available at the https://mertcookimg.github.io/bi-mode/
Chinese Translation
基于双边控制(bilateral control)的模仿学习能够同时捕获位置和力信息,因而非常适合接触丰富的操作任务。然而,现有方法在推理阶段为操作者提供的手段有限,难以指定已学习任务应如何执行,例如慢速或快速、轻柔或用力。我们提出Bi-MoDe,一种修饰符条件化(modifier-conditioned)的解码框架,通过adaLN-Zero将受约束的潜变量注入Transformer动作解码器的每一层,使行为指令能够直接影响动作块(action chunk)的生成。我们在一个真实世界的白板擦拭任务上,结合时间性与物理性修饰符的组合对该方法进行了评估。结果表明,Bi-MoDe在保持相当的时间控制能力的同时,相比动作分块基线提升了对物理指令的遵循能力。消融实验进一步表明,解码器条件化与潜空间组合之间存在交互作用,二者的结合对于准确遵循物理指令至关重要。更多资料请访问 https://mertcookimg.github.io/bi-mode/
cs.RO / 2 / 2609.16041

MR-GLi: Mixed Reality-Based Gripper-Linked Overlays for Underwater Robot Arm Teleoperation via Bilateral Control

MR-GLi:基于混合现实的夹爪联动叠加界面,用于基于双边控制的水下机械臂遥操作
Sasago, Masashi, Kobayashi, Masato, Uranishi, Yuki
Abstract
Visual torque feedback supports underwater bilateral teleoperation, but the benefit of mixed reality (MR) over conventional monitor presentation remains unclear. We present MR-GLi, an MR interface that spatially registers a reaction torque indicator and wrist-camera image to the robot gripper. Twenty participants performed lift and pick-and-place tasks with rigid and compliant objects in a counterbalanced within-subject comparison with a 2D monitor, using identical visual-feedback content and four-channel bilateral control. MR-GLi provided gripper-linked access to visual feedback while maintaining a similar level of torque-regulation performance to the 2D monitor. Subjective evaluation further indicated reduced perceived burden associated with shifting attention between the workspace and visual feedback. These results demonstrate the feasibility of gripper-linked MR overlays for underwater bilateral teleoperation and highlight the importance of considering information access in addition to task performance. Additional material: https://mertcookimg.github.io/mr-gli/
Chinese Translation
视觉力矩反馈可支持水下双边遥操作,但混合现实(MR)相比传统监视器呈现的优势尚不明确。我们提出了MR-GLi,这是一种将反作用力矩指示器和腕部相机图像空间配准到机器人夹爪上的混合现实界面。二十名参与者在刚性和柔性物体上执行提升以及抓取放置任务,采用被试内平衡比较方法与2D监视器进行对比,两者使用完全相同的视觉反馈内容和四通道双边控制。结果表明,MR-GLi在提供夹爪联动的视觉反馈访问的同时,保持了与2D监视器相当的力矩调节性能。主观评价进一步表明,其降低了参与者在工作空间与视觉反馈之间转移注意力所感知的负担。这些结果证明了夹爪联动混合现实叠加界面用于水下双边遥操作的可行性,并强调了除任务性能之外考虑信息获取方式的重要性。附加材料:https://mertcookimg.github.io/mr-gli/
cs.RO / 3 / 2609.16074

World-Action Models for Robot Learning and Control: A Survey

面向机器人学习与控制的世界-动作模型:综述
Lu, Zuxing, Zhai, Hongjia, Wang, Guanzhi, Zeng, Huajian, Yang, Jiaqi, Liu, Jingyu, Cheng, Lei, Zhang, Yuantai, Qiu, Yuheng, Cheng, Zezhou, Laptev, Ivan, Xu, Danfei, Riviere, Benjamin, Loianno, Giuseppe, Xing, Eric, Zuo, Xingxing
Abstract
Robots operating in open environments act under partial observability, physical constraints, and dynamic task contexts. Beyond mapping observations and language instructions to actions, they must anticipate how candidate actions may affect future states and task-relevant outcomes. Recent advances in world models, video generation, and Vision-Language-Action (VLA) policies have motivated the development of World-Action Models (WAMs), which couple future world prediction with executable action generation. This survey provides a robotics-oriented review of WAMs. We clarify their scope relative to conventional world models, model-based reinforcement learning, action-conditioned video generation, and reactive VLA policies, and organize existing methods through a unified taxonomy covering representations, transition modeling, action interfaces, architectures, training pipelines, data modalities, and scaling strategies. We further review applications of WAMs in manipulation, navigation, and autonomous driving, and we summarize the datasets, benchmarks, metrics, and protocols used to evaluate WAM systems. Finally, we discuss key challenges in action alignment, world-action factorization, spatial and multi-view consistency, long-horizon memory, neural simulation for closed-loop policy learning, and efficient inference. Taken together, this survey aims to provide a concise technical foundation for integrating predictive world modeling with action generation, toward more reliable embodied robot intelligence. Project page: https://rcl-robotics.github.io/Awesome-World-Action-Models.
Chinese Translation
在开放环境中运行的机器人需要在部分可观测性、物理约束和动态任务情境下行动。除了将观测和语言指令映射为动作之外,机器人还必须预判候选动作可能如何影响未来状态和与任务相关的结果。世界模型、视频生成以及视觉-语言-动作(VLA)策略的最新进展推动了世界-动作模型(World-Action Models,WAMs)的发展,该模型将未来世界预测与可执行的动作生成相结合。本综述从机器人学角度对WAMs进行了系统回顾。我们阐明了WAMs相对于传统世界模型、基于模型的强化学习、动作条件视频生成以及反应式VLA策略的范围与区别,并通过一个统一的分类体系对现有方法进行组织,涵盖表征、状态转移建模、动作接口、架构、训练流程、数据模态和扩展策略。我们进一步综述了WAMs在操作、导航和自动驾驶中的应用,并总结了用于评估WAM系统的数据集、基准、指标和评测协议。最后,我们讨论了动作对齐、世界-动作因式分解、空间与多视角一致性、长时程记忆、用于闭环策略学习的神经仿真以及高效推理等关键挑战。总体而言,本综述旨在为将预测性世界建模与动作生成相集成提供一个简明的技术基础,以迈向更可靠的具体智能。项目页面:https://rcl-robotics.github.io/Awesome-World-Action-Models。
cs.RO / 4 / 2609.16075

AssemblyGrid v1: A Benchmark for Multi-Robot Production with Temporary Coalitions, Local Information, and Geometric Constraints

AssemblyGrid v1:面向临时联盟、局部信息与几何约束的多机器人生产基准测试
Bahrpeyma, Fouad, Heik, David, Reichelt, Dirk
Abstract
Flexible robotic production requires joint decisions on process progression, material routing, resource assignment, temporary cooperation, and simultaneous execution, since each decision can affect the feasibility of the others. The challenge is greater under decentralized control, where each robot acts from bounded local information while system progress depends on collective decisions, shared resources, material state, and workspace compatibility. These properties closely match cooperative multi-agent decision making under partial observability and resource contention. This paper introduces AssemblyGrid v1, a reproducible benchmark for repeated multi-robot production that combines explicit process progression, decentralized observations, material transfer, temporary multi-robot coalitions, productive concurrency, and geometry-dependent feasibility within one task-level formulation. The benchmark includes Flow, Coalition, and Concurrency workload families, each with three scenario levels. Task success and evaluation measures are defined independently of learning reward and solution method, allowing learning-based and non-learning methods to address the same production problem. AssemblyGrid v1 is evaluated through executable conformance checks, mechanism studies, and algorithmic experiments using a privileged centralized reference, structured decentralized controllers, and MARL methods including IPPO, MAPPO, and QMIX. Results demonstrate productive execution under centralized and decentralized control. The MARL experiments further show that decentralized policies can learn effective production behavior from local observations and actions, supporting AssemblyGrid as a controlled benchmark for studying cooperative decision making in flexible robotic production.
Chinese Translation
柔性机器人生产需要对工艺推进、物料路径规划、资源分配、临时协作以及并行执行进行联合决策,因为每一项决策都可能影响其他决策的可行性。在去中心化控制下,这一挑战更为严峻:每个机器人仅基于有界的局部信息行动,而系统进展依赖于集体决策、共享资源、物料状态和工作空间兼容性。这些特性与部分可观测性和资源竞争下的协作式多智能体决策高度吻合。本文提出 AssemblyGrid v1,一个可复现的重复多机器人生产基准,它在单一任务级形式化框架中结合了显式工艺推进、去中心化观测、物料传输、临时多机器人联盟、有效并发以及依赖几何的可行性判定。该基准包含 Flow、Coalition 和 Concurrency 三个工作负载族,每个族具有三个场景级别。任务成功与评估指标独立于学习奖励和求解方法定义,使基于学习的方法与非学习方法能够解决相同的生产问题。AssemblyGrid v1 通过可执行的一致性检查、机制研究以及算法实验进行评估,实验采用了具有特权信息的集中式参考方法、结构化去中心化控制器,以及包括 IPPO、MAPPO 和 QMIX 在内的多智能体强化学习(MARL)方法。结果表明,在集中式和去中心化控制下均能实现有效生产执行。MARL 实验进一步表明,去中心化策略能够从局部观测和动作中学习到有效的生产行为,从而支持将 AssemblyGrid 作为研究柔性机器人生产中协作决策的受控基准。
cs.RO / 5 / 2609.16089

Structure-Preserving Quantum Circuit Architectures for Robot Kinematics

面向机器人运动学的结构保持量子电路架构
Morghen, Andrea, Arpenti, Pierluigi, Schiattarella, Roberto, Acampora, Giovanni, Siciliano, Bruno
Abstract
Structured spatial data require quantum encodings that preserve geometric relations, expose measurable observables, and remain implementable on finite-depth hardware. This work introduces a quantum representation and circuit architecture for rigid-body transformations and specializes it to Denavit--Hartenberg kinematics of serial open-chain manipulators. Each translational contribution is factorized into a classical metric magnitude and a signed unit direction encoded by a single-qubit Bloch vector, while parameterized rotations reproduce the ordered propagation of frame directions. A selector register prepares probabilities proportional to the contribution magnitudes, and the reduced state of a designated readout qubit encodes their normalized weighted sum. The retained classical scale then reconstructs the metric end-effector position. Two additional readout qubits encode terminal-frame axes, providing a compact and geometrically interpretable pose interface. At the ideal expectation-value level, measured Pauli observables reproduce the corresponding classical kinematic quantities. Alternative circuit architectures realize the same representation with different tradeoffs in qubit count, circuit depth, controlled operations, and measurement requirements. Validation on a serial manipulator yields numerically negligible position and orientation reconstruction errors under ideal simulation. Finite-shot simulations, noisy executions, transpilation analysis, and a hardware demonstration further characterize statistical error, noise sensitivity, and implementation overhead without asserting computational advantage.
Chinese Translation
结构化空间数据需要能够保持几何关系、提供可测量观测量、并可在有限深度硬件上实现的量子编码。本文为刚体变换提出了一种量子表示与电路架构,并将其专门应用于串联开链机械臂的Denavit–Hartenberg(D–H)运动学。每个平移分量被分解为一个经典度量幅值和一个由单量子比特Bloch矢量编码的带符号单位方向,而参数化旋转则复现坐标系方向的有序传播。一个选择寄存器制备与分量幅值成比例的概率分布,指定读出量子比特的约化态编码它们的归一化加权和。保留的经典尺度随后重建度量化的末端执行器位置。另外两个读出量子比特编码终端坐标系的轴,从而提供了一个紧凑且几何上可解释的位姿接口。在理想期望值层面,测得的Pauli观测量可复现相应的经典运动学量。不同的电路架构以量子比特数、电路深度、受控操作和测量需求方面的不同权衡实现同一表示。在串联机械臂上的验证表明,理想仿真下位置与姿态重建误差在数值上可忽略不计。有限采样仿真、含噪执行、编译转换分析以及硬件演示进一步刻画了统计误差、噪声敏感性和实现开销,但本文并未声称具有计算优势。
cs.RO / 6 / 2609.16186

Occupancy Network-Guided Autonomous Robotic Partial Nephrectomy

占用网络引导的自主机器人部分肾切除术
Kilmer, Ethan, Henrich, Pit, Ge, Jiawei, Scheikl, Paul M., Connolly, Laura, Lokeshwar, Soum D., Chen, Joseph, Opfermann, Justin D., Kumar, Kaitlyn, Shepard, Lauren, Ghazi, Ahmed, Singla, Nirmish, Cha, Richard J., Cleary, Kevin, Mathis-Ullrich, Franziska, Krieger, Axel
Abstract
Autonomous soft-tissue cancer surgery has been limited to interventions on organ surfaces, because current systems cannot perceive and adapt to anatomy once it deforms or is cut. We introduce the first vision-guided autonomous system capable of performing complete tumor resections for partial nephrectomy. Our system integrates conditional occupancy networks, trained entirely in a physics-based simulation, that infer full 3-D anatomy (tumor, margin tissue, and kidney) from single-view partial point clouds. These occupancy networks maintain intraoperative tracking even as tissue is cut and deformed, enabling adaptive planning and execution. The surgical platform combines a depth camera for capturing surface point clouds, dual robotic arms for electrosurgical cutting and vacuum-based tissue manipulation, and an autonomous control strategy for tumor resection. In patient-derived hydrogel phantoms under an open partial nephrectomy setting, the robot performed eight consecutive autonomous tumor resections comprising 77 electrosurgical cuts, with all cuts achieving negative surgical margins and 1.61 $\pm$ 0.48 mm mean absolute margin error. This work demonstrates, for the first time, a foundation for supervised autonomous closed-loop, imaging-driven, margin-negative tumor removal in phantoms.
Chinese Translation
自主软组织癌症手术一直局限于器官表面的干预操作,因为现有系统无法在解剖结构发生形变或被切割后进行感知和适应。我们提出了首个能够完成部分肾切除术肿瘤完整切除的视觉引导自主系统。该系统集成了完全基于物理仿真训练的条件占用网络,可从单视角部分点云推断完整的三维解剖结构(肿瘤、切缘组织和肾脏)。这些占用网络即使在组织被切割和形变的情况下也能保持术中跟踪,从而实现自适应的规划与执行。该手术平台结合了用于采集表面点云的深度相机、用于电外科切割和基于真空的组织操作的双机械臂,以及用于肿瘤切除的自主控制策略。在开放部分肾切除术 setting 下的患者来源水凝胶仿体实验中,机器人连续完成了八次自主肿瘤切除,共包含77次电外科切割,所有切割均达到阴性手术切缘,平均绝对切缘误差为 1.61 ± 0.48 mm。这项工作首次在仿体中证明了有监督自主闭环、影像驱动、切缘阴性的肿瘤切除的基础。
cs.RO / 7 / 2609.16256

Tendon-Driven Continuum Robot with Modular Stiffness and In-Situ Self Pose Estimation

具有模块化刚度与原位自姿态估计的腱驱动连续体机器人
Ning, Guo, Sue, Cao, Zheng, Hu, Junzhe, Bu, Xiangyun, Quinn, David, Wu, Tiancheng, Erickson, Zackory, Majidi, Carmel
Abstract
Continuum robots enable smooth shape morphing and safe interaction in confined environments. However, most existing systems are task-specific and depend on external sensing infrastructure, limiting their adaptability and real-world deployment. This paper presents a self-contained modular continuum robotic platform that combines mechanical reconfigurability with onboard pose estimation. The robot is constructed from interchangeable continuum joints with analytically precomputed stiffness, allowing rapid assembly and direct programming of the robot shape. Proprioceptive sensing is achieved using magnetic sensors and a modular learning-based framework, where a single model is trained per joint and reused across configurations. The system is experimentally validated in real world, demonstrating self-sensing capabilities and adaptation without external tracking.
Chinese Translation
连续体机器人能够在受限环境中实现平滑的形变和安全的交互。然而,现有的大多数系统都是针对特定任务的,并且依赖于外部感知基础设施,这限制了其适应性和实际部署能力。本文提出了一种集成了机械可重构性与机载姿态估计的独立式模块化连续体机器人平台。该机器人由具有解析预计算刚度的可互换连续关节构成,支持快速组装并可直接对机器人形状进行编程。本体感知通过磁传感器和模块化的基于学习的框架实现,其中每个关节仅需训练一个模型,并可在不同配置中复用。该系统在真实环境中进行了实验验证,展示了在无外部跟踪的情况下自感知与自适应的能力。
cs.RO / 8 / 2609.16319

ConGraspXL: Controllable Constraint-Conditioned Dexterous Grasping Motion Synthesis

ConGraspXL:可控的约束条件驱动的灵巧抓取动作合成
Zhang, Hui, Meboldt, Mirko, Song, Jie
Abstract
Dexterous grasping is usually conducted for specific tasks, leading to heterogeneous constraints such as specific approach directions, desired contact regions, specified wrist trajectories, and functional hand poses. Our previous work, GraspXL, achieves scalable grasping motion synthesis for diverse objects and hand morphologies, while lacking controllability for synthesis under such various task-driven constraints. In this paper, we propose ConGraspXL, which extends GraspXL with controllable constraint-conditioned grasp motion synthesis that accommodates diverse task-driven constraints and their combinations. We introduce a hierarchical constraint formulation, enable flexible constraint composition with a masked residual interface, and improve control precision with dynamic hand centers and feed-forward wrist guidance. Without losing the strong generalization capabilities of GraspXL, ConGraspXL enables precise and flexible controllability for various individual constraints and their combinations, providing a plug-and-play low-level grasp controller for downstream applications such as whole-body grasp completion, functional grasping, and human-motion imitation.
Chinese Translation
灵巧抓取通常针对特定任务执行,因此会产生异构约束,例如特定的接近方向、期望的接触区域、指定的腕部轨迹以及功能性手部姿态。我们先前的工作 GraspXL 实现了面向多样物体和手部形态的可扩展抓取动作合成,但在上述各类任务驱动约束下缺乏合成的可控性。本文提出 ConGraspXL,在 GraspXL 的基础上扩展了可控的约束条件化抓取动作合成,能够适应多样化的任务驱动约束及其组合。我们引入了一种层次化的约束建模方法,通过掩码残差接口实现灵活的约束组合,并利用动态手部中心和前馈腕部引导提升控制精度。在保留 GraspXL 强泛化能力的同时,ConGraspXL 实现了对各种单独约束及其组合的精确且灵活的可控性,可为全身抓取完成、功能性抓取和人体动作模仿等下游应用提供一个即插即用的底层抓取控制器。
cs.RO / 9 / 2609.16331

ManiSkillFormer: Demonstration-Free Compositional Manipulation via Task-Conditioned Geometric Contracts

ManiSkillFormer:基于任务条件几何契约的无示范组合式机器人操作
Yu, Peiqi, Dabhi, Mosam, Li, Shangtao, Li, Bowei, Jeni, Laszlo, Liu, Changliu
Abstract
We present ManiSkillFormer, a neuro-symbolic framework for demonstration-free and compositional robotic manipulation. Instead of learning end-to-end visuomotor policies, ManiSkillFormer introduces task-conditioned geometric contracts that explicitly structure the interface between perception and action. Each manipulation skill declares the semantic geometric primitives required for execution, such as object keypoints and surface normals. Building on human-defined skill structures, LLM agents generate these contracts and corresponding motion templates for different objects and task contexts. These contracts guide the perception module to ground task-relevant 3D primitives from observations, which are then used to instantiate reusable motion templates stored in a skill library. We evaluate ManiSkillFormer on Galaxea R1-Lite dual-arm robot across three settings: zero-shot pick-and-place over 8 object categories with 30 different instances, functional manipulation tasks including unscrewing, pouring, pressing, and folding, and 3 long-horizon tasks. ManiSkillFormer achieves higher average success rates than the evaluated baselines and two ablated pipelines: 88.24% for demonstration-free pick-and-place, 75.00% average success on functional manipulation and 50--80% completion rates across the long-horizon tasks. These results show that our design enables composable and reusable manipulation across objects and tasks without per-object policy fine-tuning or additional robot demonstrations.
Chinese Translation
我们提出了ManiSkillFormer,一个用于无示范、组合式机器人操作的神经符号框架。ManiSkillFormer并非学习端到端的视觉运动策略,而是引入了任务条件几何契约,显式地构建感知与动作之间的接口结构。每个操作技能都声明了执行所需的语义几何基元,例如物体关键点和表面法向量。在人类定义的技能结构基础上,大语言模型(LLM)智能体针对不同物体和任务上下文生成这些契约及相应的运动模板。这些契约引导感知模块从观测中提取与任务相关的3D基元,随后利用这些基元实例化存储在技能库中的可复用运动模板。我们在Galaxea R1-Lite双臂机器人上对ManiSkillFormer进行了评估,涵盖三种场景:覆盖8个物体类别、30个不同实例的零样本抓取与放置,包括拧开、倾倒、按压和折叠等功能性操作任务,以及3个长时程任务。ManiSkillFormer取得了高于所评估基线和两个消融流程的平均成功率:无示范抓取与放置任务达88.24%,功能性操作平均成功率为75.00%,长时程任务的完成率为50%–80%。这些结果表明,我们的设计无需针对每个物体进行策略微调或额外的机器人演示,即可实现跨物体、跨任务的可组合、可复用操作。
cs.RO / 10 / 2609.16346

Auto-HSI: Personalized human control of a robot swarm on demand by using LLMs for online automatic code generation

Auto-HSI:利用大语言模型(LLM)进行在线自动代码生成,实现按需求的个性化人机群交互控制
Nazzari, Alessandro, Cerisara, Nathan, Tonnis, Dorian, Zakir, Raina, Labarile, Lorenzo, Zhu, Weixu, Dorigo, Marco, Heinrich, Mary Katherine
Abstract
This paper presents Auto-HSI, a method for generating personalized human-swarm interaction (HSI) interfaces on demand. The objective is to enable untrained operators to use natural language descriptions and gesture demonstrations to explain how they want the robots to collectively behave in response to their gestures. Based on these inputs, the code should automatically be generated for personalized state machines that will control the robots as desired, in response to the desired gesture inputs. In the developed Auto-HSI prototype, the generated code produces a personalized interface for centralized control using one- and two-handed gestures, enabling a user to teleoperate the robots' motion, formation shape, and shape deformation. We test the gesture tracking and code generation components of Auto-HSI against performance benchmarks. We then test the full Auto-HSI prototype in ``live'' operation experiments, in which real human operators centrally control 50 simulated robots in a physics-based simulator, under nominal and noisy conditions. In these experiments, robots are teleoperated to: score a goal, traverse a maze that requires shape deformation, and score two simultaneous goals by splitting into two groups. We also demonstrate a real human operator making live updates to their personalized Auto-HSI interface during operation (in simulation). Finally, we demonstrate live operation of real robots.
Chinese Translation
本文提出了Auto-HSI,一种按需生成个性化人-群交互(HSI)界面的方法。其目标是使未经训练的操作者能够通过自然语言描述和手势演示,说明他们希望机器人如何根据其手势作出集体行为响应。基于这些输入,系统应自动生成个性化的状态机代码,从而按照操作者的意愿,响应指定的手势输入来控制机器人。在所开发的Auto-HSI原型中,生成的代码产生了一个使用单手和双手手势进行集中控制的个性化界面,使用户能够遥操作机器人的运动、队形形状及形状变形。我们针对性能基准测试了Auto-HSI的手势跟踪和代码生成组件,随后在“实时”运行实验中测试了完整的Auto-HSI原型,其中真实的人类操作者在基于物理的仿真器中、在标称和噪声条件下集中控制50个仿真机器人。在这些实验中,机器人被遥操作完成以下任务:射门得分、穿越需要形状变形的迷宫,以及通过分组同时完成两个射门任务。我们还展示了真实操作者在运行过程中(仿真环境下)实时更新其个性化Auto-HSI界面。最后,我们演示了真实机器人的实时运行。
cs.RO / 11 / 2609.16368

UDAV: Uncertainty-Driven Adaptive VLM Waypoint Planner

UDAV:不确定性驱动的自适应VLM航点规划器
Farhani, Ghazal, Shabani, Shabnam
Abstract
Vision-language models (VLMs) can generate routes directly from aerial imagery for off-road navigation, but their predictions provide no indication of reliability. We present UDAV, an Uncertainty-Driven Adaptive VLM Waypoint Planner for UAV-guided UGV navigation. UDAV draws multiple stochastic trajectory predictions, selects their medoid as a self-consistent nominal route, and estimates predictive uncertainty from their spatial dispersion. When the maximum uncertainty across interior waypoints exceeds a threshold, UDAV invokes a reconsideration stage; otherwise, it returns the medoid directly. We evaluate UDAV on 400 held-out trajectory queries from two UAV flights. Stochastic medoid selection reduces the mean average displacement error (ADE) from 147.4 pixels for a deterministic prediction to 115.9 pixels. The complete planner achieves a mean ADE of 110.4 pixels, a 25.1% reduction relative to deterministic planning, while producing valid trajectories for all queries. UDAV also yields the lowest 90th- and 95th-percentile errors among all evaluated configurations, including a higher-budget K=10 consensus baseline. Relative to the K=5 medoid, UDAV reduces these errors from 225.3 and 326.0 pixels to 199.0 and 290.8 pixels, respectively. These results demonstrate that stochastic VLM predictions provide both a stronger nominal route and an actionable uncertainty signal for selectively mitigating large planning errors.
Chinese Translation
视觉-语言模型(VLM)可以直接从航空影像生成用于越野导航的路线,但其预测无法提供可靠性指示。我们提出了UDAV,一种用于无人机(UAV)引导无人地面车辆(UGV)导航的不确定性驱动自适应VLM航点规划器。UDAV进行多次随机轨迹预测,选取其中中心点作为自洽的名义路线,并根据其空间离散度估计预测不确定性。当内部航点中的最大不确定性超过阈值时,UDAV会启动重新考虑阶段;否则,直接返回该中心点。我们在来自两次无人机飞行的400个保留轨迹查询上评估了UDAV。随机中心点选择将平均绝对位移误差(ADE)从确定性预测的147.4像素降低到115.9像素。完整规划器实现了110.4像素的平均ADE,相比确定性规划降低了25.1%,同时为所有查询生成了有效轨迹。在所有评估配置中,包括预算更高的K=10共识基线,UDAV还取得了最低的第90和第95百分位误差。相对于K=5中心点方法,UDAV将这些误差分别从225.3和326.0像素降低到199.0和290.8像素。这些结果表明,随机VLM预测既能提供更强的名义路线,又能提供可操作的不确定性信号,从而有选择地缓解较大的规划误差。
cs.RO / 12 / 2609.16378

Geometry vs Structure: Graph-Based Diagnostics for LiDAR Point-Cloud Simulation Fidelity

几何与结构之辨:基于图的LiDAR点云仿真保真度诊断方法
Farhani, Ghazal, Rahman, Taufiq
Abstract
Digital twins provide a scalable and cost-effective complement to real-world testing for validating autonomous-driving and advanced driver-assistance system (ADAS) sensor pipelines. However, quantifying their fidelity remains challenging, particularly for 3D LiDAR point clouds, where conventional geometric metrics may overlook important structural discrepancies. We present a graph-based framework for evaluating the structural fidelity of simulated LiDAR point clouds against real-world scans. While scan-level metrics such as Chamfer distance capture point-wise geometric similarity, they do not explicitly represent connectivity, topology, or object-level organization. Our framework constructs graphs from real and simulated point clouds, applies Louvain community detection to identify spatially coherent subgraphs, and matches corresponding communities using centroid proximity. For each matched pair, we compute $r_\lambda$, a bounded graph-spectral metric motivated by Weyl's inequality, and compare it with density-aware Chamfer distance (CDC) as a geometric baseline. Controlled perturbation experiments demonstrate that $r_\lambda$ is invariant to rigid transformations and robust to sensor noise while remaining sensitive to structural deformation. We evaluate the framework on 50 paired real and simulated LiDAR scans acquired using a Velodyne VLP-32C sensor and CARLA, respectively. The dataset contains more than 1,000 matched communities across four representative classes: vehicles, vegetation, trees, and building walls. The results show that geometric and structural measures capture complementary aspects of simulation fidelity, supporting graph-spectral analysis as an additional diagnostic layer for validating digital twins in ADAS and autonomous-driving applications.
Chinese Translation
数字孪生为自动驾驶和高级驾驶辅助系统(ADAS)传感器流水线的验证提供了一种可扩展且经济高效的现实世界测试补充手段。然而,对其保真度的量化仍然具有挑战性,尤其是对于3D LiDAR点云而言,传统的几何度量可能会忽略重要的结构性差异。我们提出了一种基于图的框架,用于评估仿真LiDAR点云相对于真实世界扫描的结构保真度。尽管Chamfer距离等扫描级度量能够捕捉逐点的几何相似性,但它们并未显式地表示连通性、拓扑结构或对象级组织。我们的框架从真实和仿真点云构建图,应用Louvain社区发现算法识别空间上连贯的子图,并通过质心邻近性匹配相应的社区。对于每一对匹配的社区,我们计算$r_\lambda$——一种受外尔(Weyl)不等式启发、具有界的图谱度量——并将其与密度感知Chamfer距离(CDC)这一几何基线进行比较。受控扰动实验表明,$r_\lambda$对刚性变换保持不变,对传感器噪声具有鲁棒性,同时对结构形变保持敏感。我们在50对分别由Velodyne VLP-32C传感器和CARLA采集的配对真实与仿真LiDAR扫描数据上对该框架进行了评估。该数据集涵盖车辆、植被、树木和建筑墙面四个代表性类别,包含1000多个匹配社区。结果表明,几何度量与结构度量捕捉了仿真保真度的互补方面,支持将图谱分析作为验证ADAS和自动驾驶应用中数字孪生的附加诊断层。
cs.RO / 13 / 2609.16405

Collision-Aware Humanoid Whole-Body Control under Imperfect Tracking Targets

不完美跟踪目标下的碰撞感知人形机器人全身控制
Gadde, Mohitvishnu S., Malik, Ashish, Dugar, Pranay, Shrestha, Aayam Kumar, Fern, Alan
Abstract
Humanoid robots often execute motion commands through whole-body controllers (WBCs) that track targets while maintaining balance and stability. However, most WBCs are blind to scene geometry, which can lead to collisions from imperfect target motions that are geometrically unsafe due to perception, planning, or teleoperation errors. We propose RECAL, a Robot--Environment Cross-Attention Layer that wraps a blind WBC to trade off target tracking against collision avoidance using external scene geometry. RECAL supports collision-aware tracking of floating-base and end-effector commands, including collision avoidance for held objects. It represents the robot, held objects, and environment as point clouds, using cross-attention between robot/object points and the environment to produce geometry-aware control features. In simulation, RECAL improves collision avoidance while preserving target-tracking performance across frozen-arm and adaptive-arm locomotion, object-carrying, and standing-manipulation scenarios relative to alternative geometry-aware WBC architectures. We further demonstrate the controller on a real Digit V3 humanoid robot.
Chinese Translation
人形机器人通常通过全身控制器(WBC)执行运动指令,在跟踪目标的同时保持平衡与稳定。然而,大多数全身控制器对场景几何信息一无所知,当目标运动因感知、规划或遥操作错误而在几何上不安全时,可能导致碰撞。我们提出了RECAL(机器人—环境交叉注意力层,Robot–Environment Cross-Attention Layer),它包裹在一个“盲”的全身控制器之外,利用外部场景几何信息在目标跟踪与碰撞避免之间进行权衡。RECAL支持对浮动基座和末端执行器指令的碰撞感知跟踪,包括对所持物体的碰撞避免。它将机器人、所持物体和环境表示为点云,通过机器人/物体点与环境之间的交叉注意力生成几何感知的控制特征。在仿真中,相较于其他几何感知的全身控制器架构,RECAL在冻结臂行走、自适应臂行走、搬运物体以及站立操作等场景下,在保持目标跟踪性能的同时提升了碰撞避免能力。我们还在真实的Digit V3人形机器人上演示了该控制器。
cs.RO / 14 / 2609.16437

XRoboToolKit-T: Teleoperation with High Stability and Precision with Tactile Sensing for Contact-rich Manipulation

XRoboToolKit-T:面向接触密集型操作的高稳定性、高精度触觉感知遥操作系统
Dengxiong, Xiwen, Wang, Xueting, Jing, Ke, Li, Rui, Zhang, Yunbo
Abstract
Collecting high-quality robot data for contact-rich manipulation tasks is essential for enabling robots to acquire real-world skills. However, existing data collection solutions often lack the capability to obtain stable and high-frequency tactile feedback, limiting their effectiveness in contact-rich manipulation scenarios. In this work, we propose a versatile teleoperation system with tactile-driven assistance to enable high-frequency and stable contact-rich manipulation. The proposed XRoboToolKit-T teleoperation system incorporates a tactile-informed force control architecture, designed to ensure both stable and precise force control in contact-rich manipulation during teleoperation. The stabilizer haptic module rapidly analyzes the normal force distribution and infers pseudo shear force, enabling real-time tactile-based assistance during manipulation. The refiner haptic module integrates a vision-language-action model to predict and refine manipulation actions based on tactile sensing data and task descriptions. We apply the proposed teleoperation system to challenging contact-rich manipulation tasks, including grasping a deformable rubber pipette for liquid transfer and inserting a medical syringe into a vascular training pad, to demonstrate the effectiveness of tactile-informed force control. Furthermore, the system achieves higher data collection efficiency and improved manipulation stability compared to state-of-the-art teleoperation without tactile assistance.
Chinese Translation
为接触密集型操作任务采集高质量机器人数据,对于使机器人获得真实世界技能至关重要。然而,现有的数据采集方案通常缺乏获取稳定且高频触觉反馈的能力,限制了其在接触密集型操作场景中的有效性。在本工作中,我们提出了一种具有触觉驱动辅助功能的通用遥操作系统,以实现高频且稳定的接触密集型操作。所提出的 XRoboToolKit-T 遥操作系统融合了触觉感知的力控制架构,旨在确保遥操作过程中接触密集型操作的力控制既稳定又精确。稳定器(stabilizer)触觉模块可快速分析法向力分布并推断伪剪切力,从而在操作过程中提供基于触觉的实时辅助。精调器(refiner)触觉模块集成了视觉-语言-动作模型,基于触觉感知数据和任务描述来预测并优化操作动作。我们将所提出的遥操作系统应用于具有挑战性的接触密集型操作任务,包括抓取可变形橡胶移液管进行液体转移,以及将医用注射器插入血管训练垫,以展示触觉感知力控制的有效性。此外,与最先进的无触觉辅助遥操作相比,该系统实现了更高的数据采集效率和更优的操作稳定性。
cs.RO / 15 / 2609.16443

The Neverwhere Visual Parkour Benchmark Suite

The Neverwhere 视觉跑酷基准测试套件
Chen, Ziyu, Bao, Henghui, Chang, Haoran, Yu, Alan, Choi, Ran, McClennen, Kai, Huh, Gio, Yang, Kevin, Qiu, Ri-Zhao, Ravan, Yajvan, Leonard, John J., Wang, Xiaolong, Isola, Phillip, Yang, Ge, Wang, Yue
Abstract
State-of-the-art visual locomotion controllers are increasingly capable at handling complex visual environments, making evaluating their real-world performance before deployment increasingly difficult. This work intends to narrow this train/evaluation gap by developing a collection of hyper-photo-realistic, closed-loop evaluation environments - The Neverwhere Benchmark Suite - comprised of over sixty 3D Gaussian Splatting reconstructions of urban indoor and outdoor scenes. Our goal is to encourage large-scale and reproducible robot evaluation by making it easier to create and integrate Gaussian splats-based reconstructions into simulated continuous testing setups. We also underscore the potential pitfalls of relying exclusively on 3D Gaussian-generated data for training, by providing policy checkpoints trained over multiple Neverwhere scenes and their performance when evaluated in novel scenes. Our analysis illustrates the necessity of sourcing diverse data to ensure performance. Code and data are available on the project page: https://ziyc.github.io/neverwhere-bench/.
Chinese Translation
最先进的视觉运动控制器在处理复杂视觉环境方面的能力日益增强,这使得在部署前评估其真实世界性能变得愈发困难。本工作旨在缩小这种训练/评估差距,开发了一组超照片级真实感、闭环的评估环境——The Neverwhere 基准测试套件,其中包含六十多个基于 3D Gaussian Splatting 重建的城市室内外场景。我们的目标是通过让基于 Gaussian splats 的重建结果更易于创建并集成到模拟的持续测试环境中,来促进大规模、可复现的机器人评估。我们还通过提供在多个 Neverwhere 场景上训练的策略检查点及其在新场景中的评估表现,强调了仅依赖 3D Gaussian 生成数据进行训练的潜在陷阱。我们的分析说明了获取多样化数据以确保性能的必要性。代码与数据可在项目页面获取:https://ziyc.github.io/neverwhere-bench/。
cs.RO / 16 / 2609.16503

Dense to MoE Adaptation for Compact Vision Language Action Policies

面向紧凑型视觉-语言-动作策略的稠密模型到混合专家模型的自适应转换
Niu, Muchun, Chen, Shuang, Wu, Yuzhou, Zhang, Linfeng
Abstract
Vision language action (VLA) policies continue to grow in parameter count, making deployment on resource-constrained robot platforms difficult. The central goal is to reduce the number of LLM-side parameters retained in the deployed policy while preserving downstream task performance. Our approach, AdaDE, adapts selected dense feed forward blocks into mixture of experts (MoE) layers and derives expert retention masks from router statistics during fine tuning. The Dense2MoE conversion preserves the original dense FFN function at initialization, so expert deactivation can start without a separate recovery stage. Instead of using a fixed shutdown rule, expert masks are updated dynamically from router usage statistics, with staged training and expert protection to avoid early collapse. With 40% of the LLM parameters deactivated, AdaDE retains 95.1% average success in LIBERO and 42.0% average success across all 50 RobotWin2.0 tasks. These results suggest that dense to MoE adaptation with dynamic expert deactivation is a practical direction for reducing active VLA model size without severe performance loss.
Chinese Translation
视觉-语言-动作(Vision language action, VLA)策略的参数量持续增长,使其难以部署在资源受限的机器人平台上。本文的核心目标是在保持下游任务性能的前提下,减少部署策略中保留的LLM侧参数数量。我们提出的方法AdaDE将选定的稠密前馈模块转换为混合专家层,并在微调过程中根据路由器统计信息推导专家保留掩码。这种Dense2MoE转换在初始化时保持了原始稠密FFN函数,因此专家的停用可以立即开始,无需单独的恢复阶段。与使用固定关闭规则不同,专家掩码根据路由器使用统计信息进行动态更新,并通过分阶段训练和专家保护机制来避免早期的性能崩溃。在停用40%的LLM参数的情况下,AdaDE在LIBERO上保持了95.1%的平均成功率,并在全部50个RobotWin2.0任务上达到42.0%的平均成功率。这些结果表明,结合动态专家停用的稠密到MoE自适应转换,是在不造成严重性能损失的情况下减小VLA模型有效规模的实用方向。
cs.RO / 17 / 2609.16504

UniDex-ViTac: Learning Unified Visuo-Tactile Dexterous Manipulation Policy from Human Video Data

UniDex-ViTac:从人类视频数据中学习统一的视觉-触觉灵巧操作策略
Lee, Hyesung, Heo, Si-Hwan, Yang, Sungwook
Abstract
Human videos provide demonstrations of dexterous manipulation but lack robot-executable actions and tactile measurements. We present UniDex-ViTac, a framework that uses human-video-guided simulation to generate robot demonstrations paired with fingertip contact observations for training a deployable visuo-tactile policy. Object-specific residual reinforcement learning specialists adapt annotated human-object interaction references to a robotic arm-hand system. Their successful rollouts pair final robot action targets with robot-side fingertip contact observations. From 50 human demonstrations across ten objects, we collect 10,000 simulated trajectories to train a single Action Chunking with Transformers (ACT) based generalist. The policy combines point clouds, proprioception, and four binary contact signals encoded through fingertip labels and a separate token, without requiring human references or privileged object identity and pose at deployment. The contact-augmented configuration achieves 68.3% macro-average success in simulation, compared with 55.5% for the point-cloud-only baseline. Without real-robot demonstrations or policy fine-tuning, it succeeds in 73/110 physical trials (66.4%) across six seen and five unseen objects, compared with 60/110 (54.5%) for the baseline, an increase of 11.8 percentage points. These results support the feasibility of learning a unified visuo-tactile dexterous manipulation policy from video-guided simulated interactions. Project page: https://unidex-vitac.github.io/
Chinese Translation
人类视频提供了灵巧操作的示范,但缺乏机器人可执行的动作和触觉测量数据。我们提出了UniDex-ViTac,这是一个利用人类视频引导的仿真来生成机器人演示的框架,并将其与指尖接触观测配对,用于训练可部署的视觉-触觉策略。针对特定物体的残差强化学习专家策略将标注的人类-物体交互参考轨迹适配到机械臂-灵巧手系统上。其成功的执行轨迹将最终的机器人动作目标与机器人侧的指尖接触观测进行配对。基于十个物体的50个人类演示,我们收集了10,000条仿真轨迹,用于训练一个基于动作分块Transformer(Action Chunking with Transformers,ACT)的单一通用策略。该策略融合了点云、本体感觉以及通过指尖标签和独立token编码的四个二值接触信号,在部署时无需人类参考轨迹或特权物体身份与位姿信息。接触信息增强的配置在仿真中实现了68.3%的宏平均成功率,而仅使用点云的基线为55.5%。在没有真实机器人演示或策略微调的情况下,该策略在六个已见物体和五个未见物体上的110次物理实验中成功了73次(66.4%),而基线为60/110(54.5%),提升了11.8个百分点。这些结果验证了从视频引导的仿真交互中学习统一视觉-触觉灵巧操作策略的可行性。项目页面:https://unidex-vitac.github.io/
cs.RO / 18 / 2609.16586

ProxiDex: Learning Dynamics-Guided Proximity Policy for Dexterous Manipulation

ProxiDex:面向灵巧操作的动力学引导接近策略学习
Bai, Yushan, Zheng, Boyu, Mao, Zhiyang, Sun, Hongzheng, Tong, Yuchuang, Li, En, Zhang, Zhengtao
Abstract
Multi-finger dexterous manipulation relies on stable hand-object interactions, yet these interactions are partially observable in practice. Visual observations are often occluded by the hand, tactile sensors introduce hardware-specific modalities and calibration burdens, and existing policies rarely model how these cues evolve under actions, making them brittle under contact uncertainty. To address these, we present ProxiDex, a dynamics-guided proximity policy framework that treats hand-object proximity as an interaction state for dexterous manipulation. ProxiDex reconstructs interaction point clouds and converts geometric distances into proximity cues, forming a hardware-agnostic contact representation that provides immersive feedback during VR teleoperation. Built on this representation, ProxiDex learns action-conditioned proximity dynamics with a coupled forward-inverse design: future observation latents are predicted from actions, while proximity variations are decoded from latent changes. Leveraging these dynamics, ProxiDex adaptively reweights proximity tokens across manipulation phases and uses dynamics-consistency supervision to guide policy inference, stabilizing action generation under unreliable visual feedback. Simulation and real-world experiments demonstrate improved success rates and robustness over representative baselines across standard, unseen objects, and perturbation scenarios. Additional visualizations are available at https://proxidex.github.io/.
Chinese Translation
多指灵巧操作依赖于稳定的手-物体交互,然而在实际场景中这些交互是部分可观测的。视觉观测常常被手部遮挡,触觉传感器引入了硬件相关的模态和标定负担,且现有策略很少建模这些线索在动作作用下的演化过程,导致其在接触不确定性下表现脆弱。为解决这些问题,我们提出了ProxiDex,一个将手-物体接近度作为灵巧操作交互状态的动力学引导接近策略框架。ProxiDex重建交互点云,并将几何距离转换为接近度线索,形成一种与硬件无关的接触表示,可在VR遥操作过程中提供沉浸式反馈。基于该表示,ProxiDex通过耦合的前向-逆向设计学习动作条件下的接近度动力学:从动作预测未来观测的潜在表征,同时从潜在变化中解码接近度变化。借助这些动力学,ProxiDex在操作的不同阶段自适应地重新加权接近度标记(token),并利用动力学一致性监督来引导策略推理,在视觉反馈不可靠的情况下稳定动作生成。仿真与真实世界实验表明,在标准场景、未见物体以及扰动场景下,本方法相比代表性基线在成功率和鲁棒性上均有提升。更多可视化结果请访问 https://proxidex.github.io/。
cs.RO / 19 / 2609.16629

Learning to Optimize UAV Path Planning for Data Sensing in Wireless Sensor Networks

面向无线传感器网络数据感知的无人机路径规划学习优化方法
Ma, Sijie, Ma, Zeyuan, Cao, Weijia, Gong, Yue-Jiao, Ma, Lingling, Huang, Zhiyang, Zhang, Jun
Abstract
UAVs have emerged as highly flexible platforms for data sensing in Wireless Sensor Networks (WSNs). Path planning for UAVs in such tasks plays a key role to assure remote sensing effectiveness and friendly energy consumption. However, existing approaches show two key limitations: i) they are primarily hand-crafted with certain design biases that harm adaptation on unseen tasks. ii) they predominantly assume idealized spatial complexities of actual environments through simplified simulation, causing them to underperform during real-world deployment. In this paper, we propose a novel learning-assisted planning framework, termed Landscape-Aware Meta Differential Evolution (LAMDE), to tackle the mentioned limitations. The major contributions come from the following aspects. We first re-formulate such UAV path planning problem to embrace challenging constraints. To efficiently navigate this highly constrained space, we propose a bi-level learning to optimize approach, where the meta-level is a trainable algorithm configuration policy that meta-learns an adaptable planning strategy for low-level planning algorithm. To address the potential training data scarcity and distribution shift in real-world environments, we introduce a landscape-aware automatic augmentation scheme that enriches training data. At the low-level, a Differential Evolution algorithm is deployed for solving the path planning tasks. To enhance the solving flexibility, we further design a variable-length encoding strategy that dynamically prunes redundant hover points and optimizes continuous flight parameters concurrently within a unified search space. Based on all proposed designs, we meta-train LAMDE and compare it with representative baselines. Comprehensive experiments demonstrate that LAMDE achieves state-of-the-art performance on the tested complex UAV path planning tasks in WSN data collection scenarios.
Chinese Translation
无人机(UAV)已成为无线传感器网络(WSN)中数据感知的高度灵活平台。在此类任务中,无人机路径规划对保障远程感知的有效性和友好的能耗起着关键作用。然而,现有方法存在两个关键局限:i)它们主要依赖人工设计,带有一定的设计偏差,损害了对未见任务的适应性;ii)它们大多通过简化仿真假设实际环境的理想化空间复杂性,导致在真实部署中表现不佳。本文提出一种新颖的辅助学习规划框架,称为景观感知元差分进化(Landscape-Aware Meta Differential Evolution, LAMDE),以解决上述局限。主要贡献包括以下方面:首先,我们重新表述了该无人机路径规划问题,以涵盖具有挑战性的约束条件。为了高效地在这一高度受限的空间中进行导航,我们提出了一种双层学习优化(bi-level learning to optimize)方法,其中元层是一个可训练的算法配置策略,通过元学习为底层规划算法提供可适应的规划策略。为解决真实环境中潜在的训练数据稀缺和分布偏移问题,我们引入了一种景观感知的自动数据增强方案来丰富训练数据。在底层,采用差分进化(Differential Evolution)算法求解路径规划任务。为增强求解灵活性,我们进一步设计了一种变长编码策略,可在统一搜索空间中动态剪除冗余悬停点并同时优化连续飞行参数。基于上述所有设计,我们对LAMDE进行了元训练,并与代表性基线方法进行了比较。大量实验表明,LAMDE在WSN数据收集场景下测试的复杂无人机路径规划任务上取得了最先进的性能。
cs.RO / 20 / 2609.16641

SAVLA: Symmetry-Aware Vision-Language-Action Models for Robotic Manipulation

SAVLA:面向机器人操作的对称性感知视觉-语言-动作模型
Li, Junle, Li, Weixian Waylon, Wu, Fuxiang, Hao, Fusheng, He, Fengxiang
Abstract
Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the range of scene poses that the demonstrations cover. We propose SAVLA, an end-to-end symmetry-aware VLA model for robust and data-efficient policy learning. Our approach keeps the pretrained vision-language backbone entirely frozen while combining it with an equivariant flow-matching action head and a learned canonicalizer. The head decomposes its state, action, and conditioning inputs into invariant and equivariant channels, and preserves this typing throughout all of its layers. The canonicalizer transforms oblique-view images into a canonical frame and rotates the geometric conditions consistently. We evaluate our model on LIBERO. Compared with the GR00T N1.5 baseline, SAVLA improves the success rate averaged over all four LIBERO suites by 5.1 points and increases the mean success rate under rotation on LIBERO-Goal from 41.5% to 90.4%.
Chinese Translation
视觉-语言-动作(VLA)模型已成为语言条件机器人操作的主流范式。然而,尽管图像和语言指令本身蕴含几何信息,VLA模型的空间能力却完全从示范数据中习得。因此,它们仅在示范数据所覆盖的场景姿态范围内才可靠。我们提出了SAVLA,一种端到端的对称性感知VLA模型,用于鲁棒且数据高效的政策学习。我们的方法完全冻结预训练的视觉-语言主干网络,同时将其与等变流匹配(equivariant flow-matching)动作头和一个可学习的规范化器(canonicalizer)相结合。该动作头将其状态、动作和条件输入分解为不变通道和等变通道,并在所有网络层中保持这一类型划分。规范化器将斜视角图像变换到规范坐标系中,并一致地旋转几何条件。我们在LIBERO基准上评估了我们的模型。与GR00T N1.5基线相比,SAVLA在全部四个LIBERO任务集上的平均成功率提升了5.1个百分点,并将LIBERO-Goal上旋转扰动下的平均成功率从41.5%提升至90.4%。
cs.RO / 21 / 2609.16644

WholeBodyWAM: Generalizing Pre-trained World-Action Priors to Humanoid Loco-Manipulation via WBC-Grounded Coordination

WholeBodyWAM:通过WBC接地的协调将预训练的世界-动作先验泛化到人形机器人移动操作
Li, Zhuo, Yao, Yiming, Tan, Jim, Jing, Mengjie, Dong, Zhipeng, Chen, Fei
Abstract
World Action Models (WAMs) offer a promising approach to general-purpose robot manipulation by jointly modeling visual dynamics and actions. However, most WAM studies focus on tabletop or arm-centric manipulation, while humanoid loco-manipulation remains less explored. To address this gap, we introduce WholeBodyWAM, which jointly predicts future visual dynamics, manipulation actions, and whole-body control intents for generalizable humanoid loco-manipulation. It preserves pre-trained world-action priors while grounding heterogeneous whole-body controller (WBC) semantics and coordinating whole-body behavior. Extensive experiments show that WholeBodyWAM achieves an overall simulation task success rate of 91.9%, with a 0.23 improvement in real-world out-of-distribution task progress and a 70% reduction in success-rate variance across WBCs relative to the respective baselines. These results suggest a path toward scalable humanoid whole-body intelligence by extending pre-trained world-action priors through structured WBC grounding and coordination, rather than relearning whole-body behavior from scratch. Project page: https://wholebodywam.github.io/.
Chinese Translation
世界动作模型(World Action Models, WAMs)通过联合建模视觉动态与动作,为通用机器人操作提供了一种有前景的方法。然而,大多数WAM研究聚焦于桌面或以机械臂为中心的操作,而人形机器人的移动操作仍较少被探索。为填补这一空白,我们提出了WholeBodyWAM,它联合预测未来的视觉动态、操作动作以及全身控制意图,以实现可泛化的人形机器人移动操作。该方法在保留预训练世界-动作先验的同时,将异构的全身控制器(WBC)语义接地,并协调全身行为。大量实验表明,WholeBodyWAM在仿真任务中实现了91.9%的总体成功率,相较于各自的基线方法,在真实世界分布外任务进度上提升了0.23,且在不同WBC之间的成功率方差降低了70%。这些结果表明,通过结构化的WBC接地与协调来扩展预训练的世界-动作先验,而非从零重新学习全身行为,为实现可扩展的人形机器人全身智能提供了一条可行路径。项目页面:https://wholebodywam.github.io/。
cs.RO / 22 / 2609.16683

Weave: Learning Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions

Weave:从人-物交互中学习全身灵巧的移动操作
Cao, Liu, Wu, Xingze, Cui, Jingzhi, Xu, Botian, Pei, Mingzhi, Chen, Ruoqu, Xu, Mengdi
Abstract
Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of coordinated interaction, but transferring these behaviors to humanoid robots requires learning how to establish and maintain effective contacts under different embodiments and dynamics. We present Weave, a unified framework for learning whole-body dexterous humanoid-object interaction from captured human demonstrations. Weave first converts captured human-object interactions into executable robot-object references through contact-aware retargeting and approach-motion completion. At its core is a contact- and geometry-aware policy that jointly commands 29 body joints and 12 actuated finger joints across multiple objects and interaction sequences. Evaluation across nine objects yields a 92.5% success rate on trained interactions and, without any additional training, 65.0% on sequences never seen during training. We additionally release ~9,000 physically executed rollouts spanning ~23 hours, providing robot-object trajectories with contact annotations for downstream interaction-policy learning and physically consistent HOI motion generation. Project website: https://xiaohu-art.github.io/Weave/
Chinese Translation
学习人形机器人与物体的交互需要协调全身平衡、移动以及灵巧手的接触,以同时控制机器人和物体的运动。人类演示提供了协调交互的范例,但将这些行为迁移到人形机器人上,需要学习如何在不同具身形态和动力学条件下建立并维持有效接触。我们提出了Weave,一个从捕捉的人类演示中学习人形机器人全身灵巧人-物交互的统一框架。Weave首先通过接触感知的重定向(retargeting)和接近动作补全,将捕捉到的人-物交互转换为可执行的机器人-物体参考。其核心是一个具备接触与几何感知的策略,能够在多种物体和交互序列上联合控制29个身体关节和12个被驱动的手指关节。在九种物体上的评估显示,该框架在训练过的交互上达到92.5%的成功率,且无需任何额外训练,在训练中从未见过的序列上达到65.0%的成功率。我们还额外发布了约9,000次物理执行、总时长约23小时的滚动(rollout)数据,提供了带接触标注的机器人-物体轨迹,用于下游交互策略学习和物理一致的人-物交互(HOI)动作生成。项目网站:https://xiaohu-art.github.io/Weave/
cs.RO / 23 / 2609.16686

Differentiable Mesh State Estimation via Factor Graph Inference for Deformable Object Reconstruction

基于因子图推理的可微分网格状态估计方法用于可变形物体重建
Al-Zogbi, Lidia, Li, Fangjie, Tobin, Samuel, Ferguson, James, Kumar, Nithesh, Chara, Alejandro, Chung, Kuan-I, Rao, Mingxing, Acar, Ayberk, Stern, Susheela Sharma, Webster, Robert, Moyer, Daniel, Kuntz, Alan, Rucker, Caleb, Hermans, Tucker, Wu, Jie Ying
Abstract
Estimating deformable object states remains a fundamental challenge in robotics and simulation. We propose a novel factor graph-based framework for probabilistic mesh state estimation of deformable objects. The method directly updates a tetrahedral mesh, a rich and physically-grounded representation of an environment, by combining physics priors, noisy sensor measurements, and temporal smoothness constraints within a unified probabilistic formulation. The estimation problem is posed as a nonlinear least-squares optimization and solved using Levenberg-Marquardt. Ex vivo central-airway obstruction experiments and simulations on deforming cube models demonstrate reliable and accurate reconstruction under both rigid motion and deformation, highlighting the potential of this probabilistic approach for principled, measurement-driven mesh state estimation in deformable object reconstruction.
Chinese Translation
可变形物体的状态估计仍然是机器人学与仿真领域的一项基础性挑战。我们提出了一种新颖的基于因子图的概率框架,用于可变形物体的网格状态估计。该方法将物理先验、含噪传感器测量和时间平滑约束统一到一个概率表达式中,直接对四面体网格——一种丰富且具有物理依据的环境表示形式——进行更新。该估计问题被表述为非线性最小二乘优化问题,并采用Levenberg-Marquardt算法求解。离体中央气道阻塞实验以及可变形立方体模型的仿真结果表明,该方法在刚体运动和形变两种情况下均能实现可靠且精确的重建,凸显了这种概率化方法在可变形物体重建中进行有原则的、测量驱动的网格状态估计方面的潜力。
cs.RO / 24 / 2609.16696

IL-ACT: Imitation Learning with Adaptive Cartesian Tracking Control for a 30-ton Excavator

IL-ACT:面向30吨级挖掘机的基于自适应笛卡尔跟踪控制的模仿学习
Shahna, Mehdi Heydari, Kim, Seihun, Jung, Soyi, Park, Soohyun, Mattila, Jouni, Kim, Joongheon
Abstract
Autonomous excavator control is challenged by coupled kinematics, actuation lag, and uncertainty. We propose imitation learning and adaptive Cartesian tracking (IL-ACT), a novel motion control framework for a 30-ton-class excavator. An anchored, 14-input imitation policy pretrained on operator demonstrations generates nominal joint rates; adaptive Cartesian feedback and gated gain/bias estimation correct these commands before a stopping-distance governor constrains joint-reference generation. Simscape evaluation covers 100 sequential goals and spiral, figure-eight, and rounded-raster tracking, including 88 additional runs across three training seeds, two initializations, and speeds, under hydraulic response and sensing conditions. Compared with Teacher+ACT, IL-ACT completes all goals with shorter duration and lower terminal errors under both response conditions. Telemetry-initialized IL-ACT lowers RMSE in all 24 figure-eight and rounded-raster seed comparisons and lowers additional-load spiral mean RMSE by approximately 29%. Original spiral RMSE also improves over IL-only and PID. Under a shared sensor-noise realization, telemetry-initialized IL-ACT achieves 27.67% lower mean RMSE than Teacher+ACT; enabling estimation reduces mean RMSE by $22.44\%$ relative to the frozen estimator. Pretrained-weight effects remain mixed, and the original teacher comparison exhibits a spiral RMSE--maximum-error tradeoff. Analysis establishes bounded adaptive states and Cartesian feedback, with reference admissibility conditional on governor feasibility.
Chinese Translation
自主挖掘机控制面临着运动学耦合、执行滞后和不确定性等挑战。我们提出了一种将模仿学习与自适应笛卡尔跟踪相结合的方法(IL-ACT),这是一种面向30吨级挖掘机的新型运动控制框架。一个基于操作员演示预训练的、具有14个输入的锚定模仿策略生成名义关节速率;自适应笛卡尔反馈与门控增益/偏置估计在停止距离调节器约束关节参考生成之前对这些指令进行修正。Simscape仿真评估涵盖100个连续目标以及螺旋线、八字形和圆角光栅轨迹跟踪任务,其中包括在液压响应与感知条件下、跨三个训练种子、两种初始化方式和多种速度的88次附加运行。与Teacher+ACT相比,IL-ACT在两种响应条件下均能以更短的时长和更低的终端误差完成全部目标。采用遥测初始化的IL-ACT在全部24组八字形和圆角光栅种子对比中均降低了RMSE,并在附加负载螺旋线任务中将平均RMSE降低约29%。原始螺旋线任务的RMSE也优于仅模仿学习(IL-only)和PID方法。在共享的传感器噪声实现下,遥测初始化的IL-ACT相比Teacher+ACT实现了27.67%的平均RMSE降低;启用估计相比冻结估计器使平均RMSE相对降低22.44%。预训练权重的影响仍然好坏参半,且原始教师对比表现出螺旋线RMSE与最大误差之间的权衡。分析建立了自适应状态与笛卡尔反馈的有界性,参考轨迹的可行性则取决于调节器的可行性。
cs.RO / 25 / 2609.16697

World Models for Embodied Intelligence: From Plausible to Controllable to Actionable

面向具身智能的世界模型:从可信到可控再到可行动
Yao, Nanjie, Wang, Hao, Cheng, Chong, Chen, Zhikang, Li, Wenzhe, Lyu, Jiafei, Shen, Li, Zhao, Peilin, Lu, Zongqing, Huang, Gao, Hoi, Steven, Tao, Dacheng, Ye, Deheng
Abstract
World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interventions, and adapting when execution departs from expectations. Although progress is often measured by visual fidelity, their value lies in improving behavior. Before reaching for a cup, a person anticipates its weight and resistance to grasping, shaping the hand before contact. Such anticipation is coarse and rarely pictorial, yet it guides action. This raises a central question: which predictive capabilities improve behavior? Existing surveys, organized by architecture, output modality, or application domain, leave this question implicit. We introduce three progressively stronger capability levels: Plausible models preserve task-relevant temporal, geometric, or physical structure; Controllable models additionally predict how interventions alter that structure; and Actionable models translate predictions into measurable gains in planning, action, learning, evaluation, verification, recovery, or data selection. We complement this hierarchy with a 3 x 4 matrix crossing geometry, physics, and action grounding with improvement loops centered on data, rewards, policies, and the model itself. Using this framework, we survey manipulation, navigation, locomotion, autonomous driving, and general embodied learning, tracing technical progressions, clarifying capability requirements, and examining datasets, benchmarks, and evaluation protocols. We identify challenges in long-horizon consistency, uncertainty calibration, causal intervention testing, latency, verification and recovery, and cross-embodiment transfer. This perspective shifts evaluation from visual plausibility toward whether predictions capture task-relevant state, reflect intervention effects, and improve the closed-loop behavior of embodied agents.
Chinese Translation
世界模型通过维护隐状态、预判后果、比较干预效果并在执行偏离预期时进行调整,从而连接具身智能中的感知与决策。尽管其进展常以视觉保真度来衡量,但其真正价值在于改进行为。在伸手拿杯子之前,人会预先估计其重量和抓握阻力,并在接触之前就调整好手的形态。这种预判是粗略的且很少以图像形式呈现,却能有效引导行动。这引出一个核心问题:哪些预测能力真正能改进行为?现有的综述按架构、输出模态或应用领域组织,对这一问题未能给出明确回答。我们提出了三个逐级增强的能力层次:可信模型保留与任务相关的时序、几何或物理结构;可控模型进一步预测干预如何改变该结构;可行动模型将预测转化为在规划、行动、学习、评估、验证、恢复或数据选择方面可度量的收益。我们在这一层级体系之上补充了一个 3×4 矩阵,将几何、物理和动作落地三个维度,与以数据、奖励、策略和模型自身为核心的四类改进环路相交叉。基于这一框架,我们综述了操作、导航、运动控制、自动驾驶和通用具身学习领域,梳理技术演进脉络,阐明能力需求,并考察相关数据集、基准和评估协议。我们识别了长时程一致性、不确定性校准、因果干预测试、延迟、验证与恢复以及跨具身迁移等方面的挑战。本视角将评估重心从视觉可信性转向预测是否能捕捉任务相关状态、反映干预效果,并改善具身智能体的闭环行为。
cs.RO / 26 / 2609.16705

The Robot Data Factory

机器人数据工厂
Haddadin, Sami, Laptev, Ivan, Reid, Ian, Song, Dezhen, Stefanini, Cesare, Swikir, Abdalla, Zuo, Xingxing, Saoud, Lyes Saad, Hamandi, Mahmoud, Heshmat, Mohamed, Doukhi, Oualid, Naceri, Abdeldjallil, Bashar, Attique, Mohamed, Abdelrahim, Tomic, Teodor, Peng, Yue, Schneider, Samuel, Lee, Cheng-Chung, Guo, Janine, Zhang, Qinghao, Jeffery, Kim
Abstract
Physical AI requires more than increasingly large robot datasets: intelligent robots acquire knowledge through continuous interaction with the physical world. We argue that the defining scientific resource of Physical AI is therefore not raw robot data alone, but robot experience - physically grounded interaction whose observations, actions, embodiment, context, and outcomes preserve the perception-action-consequence loop. We introduce the Robot Data Factory (RDF), a mission-driven infrastructure and methodology for continuously generating, validating, benchmarking, and reusing such experience. RDF organizes heterogeneous robots and environment-specific training grounds through reproducible missions, skill curricula, synchronized multimodal sensing, external ground truth, an agentic robot network, data pipelines, and living benchmarks. Rather than treating datasets as static end products, RDF implements a closed Deploy-Measure-Learn-Repeat cycle in which validated physical experience supports world models, vision-language-action models, embodied policies, digital twins, and subsequent robot deployment. We further formalize robot experience and its quality, introduce a mission-task-skill-episode-dataset-benchmark-capability hierarchy, and derive quantitative scaling laws and an algorithmic synthesis procedure connecting robot fleet size, sensor rates, storage, learning representations, tokenization, training compute, inference, and latency to Embodied-AI cluster requirements. The framework is instantiated in three complementary physical training grounds for domestic, environmental, and energy applications. RDF thus reframes robot data generation as a continuous scientific production process and provides a pathway toward reproducible, scalable, and eventually federated infrastructure for Physical AI.
Chinese Translation
物理人工智能(Physical AI)需要的不仅仅是日益庞大的机器人数据集:智能机器人通过与物理世界的持续交互来获取知识。因此,我们认为物理人工智能的关键科学资源不仅仅是原始的机器人数据,而是机器人经验——即物理扎根的交互,其观测、动作、具身形态、情境与结果共同保留了感知-动作-后果的闭环。我们提出机器人数据工厂(Robot Data Factory, RDF),这是一种以任务为驱动的基础设施与方法论,用于持续地生成、验证、基准测试和复用此类经验。RDF 通过可复现的任务(mission)、技能课程、同步多模态感知、外部真值参照(ground truth)、智能体化机器人网络、数据管线以及动态演进的基准测试,来组织异构机器人与特定环境的训练场。RDF 并不将数据集视为静态的最终产品,而是实现了一个闭环的“部署-测量-学习-重复”(Deploy-Measure-Learn-Repeat)循环,其中经过验证的物理经验支持世界模型、视觉-语言-动作(vision-language-action)模型、具身策略、数字孪生以及后续的机器人部署。我们进一步形式化了机器人经验及其质量定义,提出了“任务-子任务-技能-回合-数据集-基准-能力”(mission-task-skill-episode-dataset-benchmark-capability)层级体系,并推导出连接机器人集群规模、传感器采样率、存储、学习表示、标记化(tokenization)、训练算力、推理与延迟到具身智能(Embodied-AI)集群需求的定量扩展定律(scaling laws)与算法综合流程。该框架已在面向家庭、环境和能源应用的三个互补的物理训练场中得到实例化。由此,RDF 将机器人数据生成重新构建为一个持续的科学生产过程,并为物理人工智能提供了一条通向可复现、可扩展乃至联邦化基础设施的路径。
cs.RO / 27 / 2609.16724

CorrRisk-WM: Corridor-Conditioned Risk World Modeling for Safety-Critical Trajectory Planning

CorrRisk-WM:面向安全关键轨迹规划的走廊条件化风险世界建模
Guo, Tingyu, Langari, Reza
Abstract
Safe local planning requires forecasting surrounding-agent motion and evaluating candidate-specific risks, since identical agent motion can pose different risks to different ego trajectories. We present CorrRisk-WM, a planning-oriented partial world model coupling environment evolution with supervised intrusion and near-miss prediction over bounded candidate-trajectory corridors. A latent environment model recursively predicts agent states and updates agent-agent and agent-map interactions. Each candidate queries the evolving environment through footprint- aware geometry and learned agent-corridor representations. A lightweight recurrent risk module uses temporal context to estimate per-slice hazards; survival aggregation yields first-entry and horizon-level event probabilities. On 29,176 scenarios from 100 Waymo validation shards, CorrRisk-WM achieves intrusion average precision (AP) of 0.8567 and 1-m near-miss first-entry AP of 0.8671. In baseline comparisons, it attains the highest near-miss AP at all three distance thresholds and the lowest observed open-loop collision rate (4.88%), with route progress of 15.35 m. Across three seeds, removing dynamic environment modeling or candidate-conditioned geometric interaction reduces mean intrusion AP from 0.8590 to 0.7624 and 0.7252, respectively. These results support coupling environment evolution with candidate-conditioned geometric reasoning for risk prediction and safety-oriented candidate selection.
Chinese Translation
安全的局部规划需要预测周围智能体的运动并评估针对特定候选轨迹的风险,因为相同的智能体运动对不同自车轨迹可能构成不同的风险。我们提出了CorrRisk-WM,这是一个面向规划的部分世界模型,将有界候选轨迹走廊上的环境演化与有监督的侵入和险情预测相耦合。潜在环境模型递归地预测智能体状态,并更新智能体-智能体以及智能体-地图之间的交互。每个候选轨迹通过考虑占用轮廓的几何方法和学习到的智能体-走廊表示来查询演化中的环境。一个轻量级循环风险模块利用时序上下文估计每一时间片段的危险程度;通过生存聚合得到首次侵入和时域级别的事件概率。在来自100个Waymo验证数据分片的29,176个场景上,CorrRisk-WM实现了0.8567的侵入平均精度(AP)和0.8671的1米险情首次侵入AP。在基线比较中,它在所有三个距离阈值下均取得最高的险情AP,并取得最低的开环碰撞率(4.88%),路线行进距离为15.35米。在三个随机种子上的消融实验表明,移除动态环境建模或候选条件化的几何交互,会使平均侵入AP分别从0.8590降至0.7624和0.7252。这些结果支持将环境演化与候选条件化几何推理相耦合,以实现风险预测和面向安全的候选轨迹选择。
cs.RO / 28 / 2609.16737

Seeing What Matters: Visual Cue Guided Video Planning for Generalizable Robot Navigation

看见关键所在:视觉线索引导的视频规划实现可泛化的机器人导航
Lee, Hojin, Li, Sizhe Lester, Hilger, Maximilian, Lu, Susie, Lilienthal, Achim J., Sitzmann, Vincent, Duecker, Daniel A.
Abstract
Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches often condition video planning on short-horizon guidance and recover geometric waypoints through scene reconstruction, leaving longer-horizon planning and precise video-to-action translation less explored. We present CueNav, a video model-based navigation framework combining visual cue guided video planning with an embodiment-specific Inverse-Dynamics Model (IDM). As visual cues, we use a Bird's-Eye View (BEV) map to convey global task context and retain part of the robot body in the egocentric observation to expose embodiment context. These cues guide the video planner, while the IDM translates dense flow fields extracted from the video plan into robot actions. With the visual cue encoding global task context, CueNav achieves nearly 2x higher success in maze navigation than planning without the cue. The body-aware view with the IDM enables precise navigation with 70% success in a narrow passage where comparison methods largely fail to complete the task. We further demonstrate zero-shot semantic-conditioned navigation and deployment of the same video planner across different robot platforms. Our results show that visual cue-guided video planning with embodiment-specific action grounding paves the way toward a generalizable navigation framework for longer-horizon planning and embodiment-aware control. Additional results and code are available on our project website: https://cuenav.github.io.
Chinese Translation
生成式视频模型通过预测未来观测作为视频规划,可以成为机器人导航的有前景的骨干框架。现有方法通常基于短时程引导进行视频规划,并通过场景重建恢复几何路径点,而对更长时程的规划以及精确的视频到动作转换仍探索不足。我们提出了CueNav,一个基于视频模型的导航框架,它将视觉线索引导的视频规划与具身特定的逆动力学模型(Inverse-Dynamics Model, IDM)相结合。作为视觉线索,我们使用鸟瞰图(Bird's-Eye View, BEV)地图来传达全局任务上下文,并在自我中心观测中保留部分机器人本体以暴露具身上下文。这些线索引导视频规划器,而IDM将视频中提取的稠密光流场转换为机器人动作。借助编码全局任务上下文的视觉线索,CueNav在迷宫导航中的成功率比不使用线索的规划方法提高了近2倍。结合具备本体感知的视角与IDM,机器人在狭窄通道中实现了70%成功率的精确导航,而对比方法大多无法完成该任务。我们进一步展示了零样本的语义条件导航,以及同一视频规划器在不同机器人平台上的部署。我们的结果表明,视觉线索引导的视频规划结合具身特定的动作接地,为面向更长时程规划和具身感知控制的可泛化导航框架铺平了道路。更多结果和代码请访问我们的项目网站:https://cuenav.github.io。
cs.RO / 29 / 2609.16745

The Latent That Never Was: A Forensic Re-run of the CVAE Ablation in Action Chunking Transformer

从未存在的潜变量:对动作分块Transformer中CVAE消融实验的复现核查
Kang, Bo
Abstract
Action Chunking Transformers (ACT) are widely used to learn robot manipulation from demonstrations. Their conditional variational autoencoder includes an encoder meant to capture differences between demonstrations during training. The original ACT paper reported that encoder removal dropped the mean success rate from 35% to 2% on two simulated tasks with human demonstrations. We re-ran this ablation in the original code and checked whether the findings depend on the implementation or training data. The published drop does not reappear in our tests, although smaller gains or losses in success rate remain uncertain. To investigate the discrepancy, we varied training length and how checkpoints are selected for evaluation. Both can reverse which policy scores higher, but the published drop's cause remains unknown. Success rates alone leave open whether the encoder provides information that helps the policy reconstruct demonstrated actions. On the tested ACT benchmark, the sampled latent provides little reconstruction benefit at every tested nonzero weight of the penalty on latent information. At inference, ACT leaves this latent unused and sets it to zero. Skipping the encoder increases training throughput in both implementations we timed. We release code, evaluation tools and results so others can repeat the comparisons and test the encoder on other tasks.
Chinese Translation
动作分块Transformer(Action Chunking Transformer, ACT)被广泛用于从人类演示中学习机器人操作。其条件变分自编码器(CVAE)包含一个编码器,旨在训练过程中捕捉不同演示之间的差异。原ACT论文报告称,移除该编码器会使两个基于人类演示的仿真任务的平均成功率从35%降至2%。我们在原始代码中重新运行了该消融实验,并检验这些结论是否依赖于具体实现或训练数据。在我们的测试中,已发表的这种性能下降并未重现,尽管成功率的较小增益或损失仍存在不确定性。为探究这一差异,我们改变了训练时长以及用于评估的检查点选择方式。两者均可能逆转哪个策略得分更高,但已发表性能下降的原因仍不清楚。仅凭成功率无法判断编码器是否提供了有助于策略重构演示动作的信息。在所测试的ACT基准上,在所有测试的非零潜变量信息惩罚权重下,采样的潜变量几乎不带来重构收益。在推理阶段,ACT并不使用该潜变量,而是将其置零。在我们计时的两种实现中,跳过编码器均可提高训练吞吐量。我们公开了代码、评估工具和结果,以便他人重复这些比较并在其他任务上测试该编码器。
cs.RO / 30 / 2609.16786

Optimal Excitation Trajectories for System Identification of Underwater Vehicles

用于水下航行器系统辨识的最优激励轨迹
Panetsos, Fotis, Kyriakopoulos, Kostas J.
Abstract
In this work, we propose a structured methodology for the system identification of underwater vehicles through the design of optimal excitation trajectories. To this end, the trajectories are parameterized using Bezier curves, which ensure smooth and differentiable motion profiles while facilitating the enforcement of constraints through appropriate manipulation of the control points. An optimization problem is formulated to determine a dynamically feasible excitation trajectory that respects safety limits and maximizes the quality of the collected data, thereby enabling reliable estimation of the vehicle's dynamic parameters using least squares. The proposed methodology is experimentally validated in a laboratory water tank, where the dynamic parameters, identified from the optimized trajectory, are evaluated by predicting the vehicle's velocity through forward simulation on previously unseen trajectories.
Chinese Translation
在这项工作中,我们提出了一种结构化的方法,通过设计最优激励轨迹来实现水下航行器的系统辨识。为此,轨迹采用贝塞尔曲线(Bezier curves)进行参数化,这不仅保证了运动轨迹的平滑性和可微性,还便于通过对控制点的适当操作来施加约束。我们构建了一个优化问题,以确定一条动力学可行且满足安全限制的激励轨迹,同时最大化所采集数据的质量,从而能够利用最小二乘法可靠地估计航行器的动力学参数。该方法在实验室水池中进行了实验验证:由优化轨迹辨识出的动力学参数,通过在先前未见过的轨迹上进行前向仿真来预测航行器的速度,并以此进行评估。
cs.RO / 31 / 2609.16810

Motion planning in high dimensional spaces hybridizing RRT and HAR via position-direction decoupling

基于位置-方向解耦的高维空间运动规划:RRT与HAR的混合方法
Cazals, Frederic, Feyeux, Nelson
Abstract
The exploration of high-dimensional spaces remains a challenging problem, in particular in the presence of narrow passages and small clearances. We propose novel sampling-based path-planning methods for high-dimensional spaces combining Rapidly-exploring Random Trees (RRT) and Hit-and-Run (HAR) random walks by decoupling the point being extended from the direction of extension. We also show that RRT and HAR appear as special cases of a generic algorithm coupling the biases used for the point and direction extension, respectively. We further study a sparse-move strategy in which only a fraction p_r of the robots is moved at each step, helping both RRT and the proposed HAR algorithms handle cluttered instances. Tests are presented for two families of models: classical piano mover problems in 3D, and complex molecular systems involving tens of rigid domains moving relatively to one another -- the latter viewed as independent robots exploring the motion space SE(3)N . Within seconds on a standard laptop, our algorithms solve instances with up to 64 robots and 384 degrees of freedom. We conclude by suggesting one of our methods, HARF, as the method of choice for complex multi-robot planning problems, being up to two orders of magnitude faster than the classical RRT moving all robots at each step--when it succeeds at all, and still up to 2.4 fold faster on most instances when both use their best p_r.
Chinese Translation
高维空间的探索仍然是一个具有挑战性的问题,尤其是在存在狭窄通道和小间隙的情况下。我们提出了新颖的基于采样的高维空间路径规划方法,该方法通过将待扩展点与扩展方向解耦,将快速扩展随机树(RRT)与Hit-and-Run(HAR)随机游走相结合。我们还证明了RRT和HAR可视为一个通用算法的特例,该算法分别耦合用于点和方向扩展的偏置。我们进一步研究了一种稀疏移动策略,即每一步仅移动一部分比例为p_r的机器人,这有助于RRT和所提出的HAR算法处理杂乱的场景。我们在两类模型上进行了测试:经典的3D钢琴搬运工问题,以及涉及数十个刚体域相对运动的复杂分子系统——后者被视为在运动空间SE(3)^N中探索的独立机器人。在标准笔记本电脑上,我们的算法可在数秒内求解多达64个机器人和384个自由度的实例。最后,我们建议将其中一种方法HARF作为复杂多机器人规划问题的首选方法:在经典RRT(每步移动所有机器人)能够成功求解的情况下,HARF最多快两个数量级;而在两者均采用各自最优p_r的大多数实例上,HARF仍最多快2.4倍。
cs.RO / 32 / 2609.16815

Rethinking Visual Embodiment Dependence in Visuomotor Policies

重新思考视觉运动策略中的视觉具身依赖性
Fang, Hongjie, Lu, Yuxuan, Wang, Chenxi, Qin, Haoxiang, Tang, Shirun, He, Zihao, Xia, Shangning, Chen, Jingjing, Liu, Wanxi, Wang, Shiquan, Lu, Cewu
Abstract
Visuomotor policies observe both the task scene and the acting embodiment, allowing embodiment-specific visual cues to influence action prediction. We study this phenomenon as visual embodiment dependence (VED) and show, through cue-conflict interventions across representative policies, that visible robot configuration can become a shortcut to task progress. Rather than eliminating VED, we argue that it should be structured around embodiment information that supports control and generalization. We realize this through embodiment canonicalization in 3D point clouds, replacing the original embodiment with a canonical end-effector representation (CER) that preserves control-relevant geometry while abstracting embodiment-specific morphology. Its editable form further enables configuration-decorrelation augmentation for unfamiliar robot configurations. Experiments show that embodiment canonicalization substantially improves human-to-robot policy transfer without robot demonstrations, while simply removing the embodiment is insufficient without preserving control-relevant geometry. We further find that CER itself can become a configuration shortcut when robot configuration becomes decoupled from task progress; configuration-decorrelation augmentation mitigates this failure mode and restores robust recovery without sacrificing performance on seen configurations. Together, these results show that robust visuomotor learning benefits from structuring, rather than removing, visual embodiment information. Project website: https://tonyfang.net/ved
Chinese Translation
视觉运动策略同时观察任务场景和执行动作的具身体,使具身特有的视觉线索能够影响动作预测。我们将这一现象研究为视觉具身依赖性(Visual Embodiment Dependence, VED),并通过在多种代表性策略上进行线索冲突干预,表明可见的机器人构型可能成为通往任务进展的捷径。我们主张不应消除VED,而应围绕有利于控制与泛化的具身信息来对其进行结构化组织。为此,我们在三维点云中实现了具身规范化(embodiment canonicalization),用一种规范的末端执行器表示(Canonical End-effector Representation, CER)替代原始具身体,该表示在保留与控制相关的几何信息的同时,抽象掉了具身特有的形态。其可编辑的形式进一步支持针对陌生机器人构型的构型去相关增强(configuration-decorrelation augmentation)。实验表明,具身规范化在无需机器人示教的情况下显著提升了人到机器人的策略迁移能力,而仅简单移除具身体而保留与控制相关的几何信息则不足以达到这一效果。我们进一步发现,当机器人构型与任务进展解耦时,CER本身也可能成为构型捷径;构型去相关增强可缓解这一失败模式,在不牺牲已见构型性能的前提下恢复稳健性。综上,这些结果表明,稳健的视觉运动学习受益于对视觉具身信息进行结构化组织,而非将其移除。项目网站:https://tonyfang.net/ved
cs.RO / 33 / 2609.16864

TEMPO: Learning Temporal Context for Dynamic Robot Manipulation

TEMPO:面向动态机器人操作的时间上下文学习
Feng, Zhenyang, Heo, Jimin, Sudderth, Erik B., Jain, Unnat
Abstract
Vision-language-action (VLA) models have achieved impressive performance in quasi-static manipulation, but struggle in dynamic manipulation tasks because they operate on a single observation at inference time. We identify two representational failures that underlie this limitation. The first is motion ambiguity, where a single observation does not include scene dynamics and therefore cannot anticipate the future state of moving objects. The second is state aliasing, where visually similar observations from different points in a task require different actions. We argue that these failures persist regardless of model scale and inference latency, showing that the bottleneck is missing temporal context rather than model capacity. Based on this insight, we propose TEMPO, which augments a pretrained VLA with two temporal inputs: a motion summary extracted from a frozen video foundation model to resolve motion ambiguity and a compact proprioceptive history to resolve state aliasing. TEMPO requires no modification to the backbone and adds minimal compute overhead at training or deployment. Across four dynamic manipulation tasks, it improves Bottle Handover success from 44% to 74% and is the only method that solves state aliasing. Probing and ablation studies confirm that each temporal signal independently addresses its corresponding failure. We further release TEMPO-Bench, a benchmark of over 50k annotated frames for evaluating motion-aware robot perception in both regression and multiple-choice formats. Project Website: https://tempo-robot.github.io/
Chinese Translation
视觉-语言-动作(VLA)模型在准静态操作任务中取得了令人瞩目的性能,但由于在推理时仅依赖单一观测,它们在动态操作任务中表现不佳。我们识别出导致这一局限性的两类表征缺陷:一是运动模糊性(motion ambiguity),即单一观测不包含场景动态信息,因此无法预判运动物体的未来状态;二是状态混叠(state aliasing),即任务中不同时间点视觉上相似的观测需要采取不同的动作。我们论证了无论模型规模和推理延迟如何,这些缺陷都会持续存在,这表明瓶颈在于时间上下文的缺失,而非模型容量。基于这一洞察,我们提出了TEMPO,它在预训练的VLA模型上引入两种时间输入:从冻结的视频基础模型中提取的运动摘要(用于解决运动模糊性),以及紧凑的本体感知历史信息(用于解决状态混叠)。TEMPO无需对主干网络进行任何修改,且在训练或部署时仅增加极小的计算开销。在四项动态操作任务中,TEMPO将“瓶子传递”任务的成功率从44%提升至74%,并且是唯一能够解决状态混叠问题的方法。探测实验和消融研究证实,每种时间信号都独立地解决了其对应的问题。我们进一步发布了TEMPO-Bench,这是一个包含超过5万标注帧的基准数据集,用于在回归和多项选择两种形式下评估运动感知的机器人感知能力。项目网站:https://tempo-robot.github.io/
cs.RO / 34 / 2609.16880

Artificial Intelligence-Enabled Space Robot Operations: Technologies, Challenges and Prospects

人工智能赋能的空间机器人操作:技术、挑战与展望
Huang, Zeyuan, Chen, Gang, Hao, Zixuan, Tang, Guoqin, Zong, Junyi, Ban, Guoyou, Wang, Jiale, Lv, Haoyang, Ren, Chaoqian, Liu, Sitong
Abstract
Space robots are increasingly expected to perform long-duration, contact-rich, and multi-stage operations with limited human intervention. Recent advances in artificial intelligence (AI), robot learning, and embodied foundation models provide new opportunities to improve the autonomy and adaptability of such systems, but their transfer to space is constrained by scarce mission data, space-specific dynamics and sensing conditions, limited onboard resources, and stringent safety requirements. This article reviews artificial intelligence-enabled space robot operations (AI-SRO) from a capability-building perspective. We first summarize representative operational scenarios, autonomy trends, and space-specific constraints. We then establish a three-layer technical framework comprising capability foundations, capability formation, and capability deployment/evolution. Within this framework, we review simulation environments, datasets and benchmarks; task and environment understanding, state perception, decision-making and planning, and action execution; and onboard deployment, ground-to-space adaptation, continual learning, and capability transfer. Finally, we propose key research directions toward trustworthy simulation and data, open-world multimodal cognition, long-horizon safe decision-making, physically constrained policy learning, and space computing infrastructures.
Chinese Translation
空间机器人正日益被期望在有限的人工干预下执行长时程、接触密集的多阶段操作。人工智能(AI)、机器人学习以及具身基础模型的最新进展为提升此类系统的自主性与适应性提供了新的机遇,但将其迁移到空间环境仍受到任务数据稀缺、空间特有的动力学与感知条件、在轨资源有限以及严格安全要求等因素的制约。本文从能力构建的视角对人工智能赋能的空间机器人操作(AI-SRO)进行了综述。我们首先总结了代表性的操作场景、自主化发展趋势以及空间特有的约束条件。随后,我们建立了一个三层技术框架,包括能力基础、能力形成和能力部署/演进。在该框架下,我们综述了仿真环境、数据集与基准;任务与环境理解、状态感知、决策与规划以及动作执行;以及在轨部署、地-空适配、持续学习与能力迁移。最后,我们提出了若干关键研究方向,包括可信仿真与数据、开放世界多模态认知、长时程安全决策、物理约束下的策略学习以及空间计算基础设施。
cs.RO / 35 / 2609.16958

Waggle Dance Inspired Motion Communication for Multiple UAVs in MuJoCo

基于摇摆舞启发的MuJoCo多无人机运动通信方法
Nengbo, Zhang
Abstract
The honeybee waggle dance motivates a communication mechanism in which one agent's movement conveys spatial information that guides other agents' actions. This paper presents a MuJoCo system that extends the point-to-point motion communication setting of MoCom to one performer and multiple observers. A performer broadcasts a six-bit navigation payload using four flight primitives and explicit null signals. Each of one to five observers processes its own onboard RGB images, extracts optical-flow trajectories, recognizes symbols, parses the message, and starts navigation only after confirming its own complete frame. Reception states and execution triggers are separate across observers, while simulation control and safety checks use shared ground truth. With stationary observers, 25 Hz image input, and ideal state-feedback control, a fixed standard suite yielded 44 correct complete messages from 53 receiver exposures across 17 nominal broadcasts; 13 broadcasts passed all group-level decoding and execution checks. Three additional no-message or input-fault controls met their expected outcomes. A separately reported supplemental suite, using the same frozen code at the default geometry, achieved 14 successful receiver exposures across three broadcasts. Near-range and wide-angle configurations exposed tracking and recognition failures, while unsuccessful receivers remained stationary. These finite simulation results support the feasibility of a waggle-dance-inspired broadcast-to-action mechanism under the tested conditions and identify the present perceptual and protocol limits.
Chinese Translation
蜜蜂的摇摆舞启发了一种通信机制,即通过一个智能体的运动传递空间信息,从而引导其他智能体的行动。本文提出了一个MuJoCo系统,将MoCom的点对点运动通信设置扩展为一名执行者与多个观察者的场景。执行者使用四种飞行动作基元和显式空信号广播一个六位导航载荷。一至五个观察者各自处理自身的机载RGB图像,提取光流轨迹,识别符号,解析消息,并仅在确认收到自己完整的帧后才开始导航。各观察者的接收状态和执行触发相互独立,而仿真控制和安全检查则使用共享的真实状态(ground truth)。在观察者静止、图像输入频率为25 Hz以及理想状态反馈控制的条件下,一个固定的标准测试套件在17次名义广播中,53次接收器暴露产生了44条正确的完整消息;13次广播通过了所有组级解码和执行检查。另外三个无消息或输入故障的对照实验均达到了预期结果。另外一个单独报告的补充套件,在使用相同冻结代码和默认几何配置的情况下,在三次广播中实现了14次成功的接收器暴露。近距和广角配置暴露了跟踪和识别方面的失败,而未成功的接收器保持静止。这些有限的仿真结果验证了在测试条件下摇摆舞启发的广播-行动机制的可行性,并指出了当前感知和协议方面的局限性。
cs.RO / 36 / 2609.16996

Overcoming technical adoption barriers for mobile service robots in rehabilitation

克服康复领域移动服务机器人的技术采用障碍
Sternitzke, Christian, Blumenthal, Sebastian, Kleedoerfer, Lukas, Deserno, Verena, Mayfarth, Anke
Abstract
Many publications on robotic systems in healthcare describe early-stage work on low technology readiness levels. This paper describes how a mobile service robot approved as a medical device reaches higher technology readiness levels by adding peripheral functions and smaller improvements, which are pivotal for user acceptance in clinical environments and which often cannot be elicited by questioning users ex-ante as certain aspects only come in mind from testing the systems in clinical settings or operational environments. Especially developers of service robots in healthcare are advised to plan with such downstream developments, which can take significant implementation time, to obtain user acceptance and achieve widespread adoption of their robotic systems.
Chinese Translation
许多关于医疗保健机器人系统的研究论文描述的是处于较低技术就绪水平的早期阶段工作。本文描述了一款已获批准作为医疗器械的移动服务机器人如何通过添加外围功能和进行较小的改进来达到更高的技术就绪水平。这些改进对于用户在临床环境中接受机器人至关重要,而这类需求往往难以通过事先询问用户获得,因为某些问题只有在临床环境或实际运行环境中对系统进行测试时才会显现。特别建议医疗保健领域的服务机器人开发人员在规划时预留此类下游开发工作的时间,这些工作可能需要较长的实施时间,但对于获得用户认可并实现机器人系统的广泛应用至关重要。
cs.RO / 37 / 2609.17035

SWIM: Vision-Language-Grounded Soft Whole-Body Interactive Manipulation

SWIM:基于视觉-语言的软体全身交互操作
Liu, Tingcong, Aung, Aye Phyu Phyu, Xiong, Junjie, Ma, Siyi, An, Bo, Wu, Ke, Jayavelu, Senthilnath
Abstract
Soft and continuum robots enable manipulation through distributed body deformation and contact, yet translating language and visual context into executable whole-body actuation remains a fundamental challenge. We present SWIM, a framework that maps an initial RGB observation and a language instruction to a complete actuation-command sequence. Its vision-language-action (VLA) policy, SWIM-VLA, combines a diffusion action head with Visual Soft Proprioception (VSP) through a shared representation of RGB observations, language instructions, and tendon states. The diffusion head models conditional distributions of expert command chunks, while VSP supervises ordered body-anchor predictions using simulation ground truth, encouraging the representation to retain body geometry when learning from limited demonstrations. Embodied mechanical intelligence supports physical execution of command sequences generated through iterative virtual rollout from evolving simulated observations, with intrinsic compliance providing local contact adaptation without online policy queries. We evaluate SWIM on packing, reaching, and grasping on a planar tendon-driven soft robot, with grasping targets anchored. In simulation, SWIM-VLA achieves success rates of 100\%, 96\%, and 88\%, respectively, outperforming an adapted OpenVLA-OFT baseline and controlled ablations. On hardware, SWIM achieves success rates of 100\%, 80\%, and 75\%, compared with 75\%, 40\%, and 25\% for direct online deployment of the same policy checkpoint.
Chinese Translation
软体与连续体机器人能够通过分布式的身体形变与接触实现操作,然而如何将语言与视觉上下文转化为可执行的全身驱动指令仍是一项根本性挑战。我们提出了SWIM,一个将初始RGB观测与语言指令映射为完整驱动指令序列的框架。其视觉-语言-动作(VLA)策略SWIM-VLA将扩散动作头与视觉软体本体感知(Visual Soft Proprioception, VSP)相结合,通过RGB观测、语言指令与腱绳状态的共享表征实现。扩散动作头对专家指令片段的条件分布进行建模,而VSP利用仿真真值监督有序的身体锚点预测,促使表征在有限示范数据学习中保留身体几何信息。具身机械智能支持物理执行通过从不断演化的仿真观测中进行迭代虚拟滚动所生成的指令序列,其内在柔顺性提供了局部接触自适应能力,而无需在线策略查询。我们在一个平面腱绳驱动软体机器人上,针对打包、到达与抓取(抓取目标固定锚定)任务对SWIM进行评估。在仿真中,SWIM-VLA分别达到100%、96%和88%的成功率,优于改造后的OpenVLA-OFT基线以及受控消融实验。在硬件上,SWIM的成功率分别为100%、80%和75%,而直接在线部署同一策略检查点的成功率仅为75%、40%和25%。
cs.RO / 38 / 2609.17106

BRAVE-6D: Benchmark for Robotic Active Vision in 6DOF Pose Estimation

BRAVE-6D:面向6自由度位姿估计的机器人主动视觉基准
Ausserlechner, Philipp, Neuberger, Bernhard, Scherl, Alessandro, Schebek, Michael, Thalhammer, Stefan, Vincze, Markus
Abstract
Detecting and grasping small objects remains a significant challenge in robotics. Active vision, where the robot moves closer to the object, is an intuitive solution, yet comparing approaches on common ground is difficult since identical physical scene setups are required. Hence, we introduce BRAVE-6D, a benchmark designed to evaluate robotic active vision systems for object pose estimation, a crucial first step in grasping objects. BRAVE-6D leverages view synthesis based on Gaussian Splats (3DGS) to provide scenes and tools for benchmarking active vision systems. We show baseline solutions performing visual servoing within the scene and accurately estimating the poses of small objects.
Chinese Translation
小物体的检测与抓取在机器人学中仍然是一项重大挑战。主动视觉——即机器人移动至更接近物体的位置——是一种直观的解决方案,但由于需要完全相同的物理场景设置,在共同的基础上比较不同方法十分困难。因此,我们提出了BRAVE-6D,一个旨在评估用于物体位姿估计的机器人主动视觉系统的基准,而位姿估计是抓取物体的关键第一步。BRAVE-6D利用基于高斯泼溅(3D Gaussian Splatting, 3DGS)的视图合成技术,为主动视觉系统的基准测试提供场景和工具。我们展示了在场景内执行视觉伺服并准确估计小物体位姿的基线解决方案。
cs.RO / 39 / 2609.17115

Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement

内在机器人奖励:复用VLA表征以实现自主评估与策略改进
Schaffer, Tobias, Elkhayat, Mohab, Nicklas, Daniela, Almohamad, Mustafa, Al-Fuqara, Elham
Abstract
Vision-language-action (VLA) systems already bring together two valuable resources for robot learning: rich visual representations and demonstrations of successful task execution. Intrinsic Robot Rewarding (IRR) proposes to use these resources for a second, complementary purpose: evaluating the robot's own outcomes and providing feedback for policy improvement. Successful demonstration endpoints define task-specific references, and the policy's frozen visual encoder provides the feature space in which new outcomes are assessed. The core reward mechanism adds a reference bank and a scoring operation to the existing pipeline, without requiring a separate learned evaluator or an additional perception backbone. Our position is that this reuse offers a promising route to lower integration effort, efficient reward computation, and reduced recurring human outcome scoring. Building on established research in visual rewards and learning from experience, IRR brings these ideas into the robot's existing perception and demonstration pipeline. An operational COMAU Racer 3 demonstrator is available at technology readiness level 4 (TRL 4). This laboratory foundation supports the next research step: connecting internal outcome evaluation to physical policy improvement. We present the reward formulation, central research questions, and an evaluation methodology linking reward reliability to task success and supervision effort. The intended contribution is a reusable approach to learn and improve from the data and experience already available in industrial robot systems.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)系统已经为机器人学习整合了两项宝贵资源:丰富的视觉表征和成功任务执行的示范。内在机器人奖励(Intrinsic Robot Rewarding, IRR)提出将这些资源用于第二个互补目的:评估机器人自身的执行结果,并为策略改进提供反馈。成功的示范终点定义了任务特定的参考基准,而策略中冻结的视觉编码器提供了评估新结果的特征空间。核心奖励机制在现有流程中增加了一个参考库和一个评分操作,无需单独训练的评估器或额外的感知骨干网络。我们的观点是,这种复用为降低集成成本、实现高效的奖励计算以及减少重复性的人工结果评分提供了一条有前景的路径。基于视觉奖励和经验学习领域的既有研究,IRR将这些思想融入机器人现有的感知与示范流程中。一台可运行的COMAU Racer 3演示系统已达到技术就绪水平4(TRL 4)。这一实验室基础支撑了下一步研究:将内部结果评估与物理策略改进相连接。我们提出了奖励的数学表述、核心研究问题,以及一种将奖励可靠性与任务成功率和监督工作量相关联的评估方法。本工作的预期贡献是提供一种可复用的方法,使工业机器人系统能够从已有的数据和经验中进行学习与改进。
cs.RO / 40 / 2609.17124

LOTUSim-Energy: A Maritime Simulator for Human-Drone Interaction in Autonomous Offshore Operation \& Maintenance

LOTUSim-Energy:面向海上自主运维中人机交互的海事仿真器
Grosset, Juliette, Dubromel, Marie, Lechêne, Hélène, Arzel, Quentin, Buche, Cédric
Abstract
Offshore maintenance requires operations in the air, the surface, and the subsea domain and include human supervision. This paper presents LOTUSim-Energy, a real-time maritime simulator designed for multi-domain human--drone interaction for offshore operation and maintenance. The plat- form unifies heterogeneous unmanned vehicles (Unmanned Aerial Vehicles: UAVs, Unmanned Surface Vehicles: USVs, Autonomous Underwater Vehicles: AUVs, Remotely Operated Vehicles: ROVs) within a distributed architecture coupling environment forcing (wind, waves, currents) and provides immersive user interfaces for supervision (desktop and virtual reality). A structured offshore task library enables repeatable evaluation of autonomy stacks under realistic metocean disturbances. The simulator supports realistic physics, energy-aware battery modeling, and fault-detection pipelines as modular validation tools. System-level performance is demonstrated on a multi-domain inspection scenario for monopile and transition piece structure, where we evaluate the reliability of integrated waypoint-follower plugin and Automatic Identification System (AIS)-referenced trajectory tracking under real-time energy monitoring. By combining unified environmental physics, heterogeneous vehicle simulation, and immersive supervision, LOTUSim-Energy provides an integration testbed for prototyping and rehearsing offshore human--robot collaboration workflows, as a step toward de-risking sea deployment.
Chinese Translation
海上维护作业需要空中、海面和水下多个领域的协同操作,并包含人员监督。本文提出了LOTUSim-Energy,一个面向海上运维多领域人-无人机交互的实时海事仿真平台。该平台在分布式架构中统一集成各类异构无人载具(无人机UAV、无人水面艇USV、自主水下航行器AUV、遥控水下机器人ROV),耦合风、浪、流等环境驱动力,并提供桌面端和虚拟现实的沉浸式监督用户界面。结构化的海上任务库支持在真实海洋气象扰动下对自主系统进行可重复的评估。该仿真器支持真实物理模拟、能耗感知的电池建模以及故障检测流程,作为模块化的验证工具。我们在单桩基础与过渡段结构的多领域巡检场景中展示了系统级性能,评估了集成航点跟踪插件以及基于船舶自动识别系统(AIS)参考的轨迹跟踪在实时能耗监测下的可靠性。通过结合统一的环境物理、异构载具仿真和沉浸式监督,LOTUSim-Energy为海上人机协作工作流的原型设计与演练提供了一个集成测试平台,为降低海上部署风险迈出了一步。
cs.RO / 41 / 2609.17141

Continual Learning for Traversability Prediction with Uncertainty-Aware Adaptation

基于不确定性感知自适应的可通行性预测持续学习
Lee, Hojin, Lee, Yunho, Duecker, Daniel A, Kwon, Cheolhyeon
Abstract
Traversability prediction is a critical component of autonomous navigation in unstructured environments, where complex and uncertain robot-terrain interactions pose significant challenges such as traction loss and dynamic instability. Despite recent progress in learning-based traversability prediction, these methods often fail to adapt to novel terrains. Even when adaptation is achieved, retaining experience from previously trained environments remains a challenge, a problem known as catastrophic forgetting. To address this challenge, we propose a continual learning framework for traversability prediction that incrementally adapts to new terrains using a generative experience recall model. A key virtue of the proposed framework is two folds: i) retain prior experience without storing past data; and ii) incorporate the uncertainty of the generated samples from the recall model, enabling uncertainty-aware adaptation. Real-world experiments with a skid-steering robot validate the effectiveness of the proposed framework, demonstrating its ability to adapt across a series of diverse environments while mitigating catastrophic forgetting.
Chinese Translation
可通行性预测是无人化导航在非结构化环境中的关键组成部分,复杂且不确定的机器人-地形交互会带来牵引力损失和动态失稳等重大挑战。尽管基于学习的可通行性预测近年来取得了进展,这些方法往往难以适应新地形。即使实现了适应,如何保留先前训练环境中的经验仍然是一个难题,即所谓的灾难性遗忘问题。为应对这一挑战,我们提出了一个用于可通行性预测的持续学习框架,该框架利用生成式经验回忆模型逐步适应新地形。所提框架具有两大优点:i) 无需存储历史数据即可保留先前经验;ii) 融合了来自回忆模型的生成样本的不确定性,从而实现不确定性感知的自适应。基于滑移转向机器人的真实世界实验验证了所提框架的有效性,表明其能够在一系列多样化环境中进行适应,同时缓解灾难性遗忘。
cs.RO / 42 / 2609.17145

LiLi: Lie Theory Based 3D LiDAR Scan Alignment Degeneracy Detection

LiLi:基于李理论的3D激光雷达扫描对齐退化检测
Hulchuk, Vsevolod, Bayer, Jan, Faigl, Jan
Abstract
In this paper, we study 3D LiDAR scan alignment in challenging scenarios with degeneracies, such as straight corridors or flat fields, where the alignment solution is not unique and compromises localization and mapping accuracy. Existing degeneracy detection methods that neglect the potential for reassociating data points are prone to being sensitive to noise and complex degeneracies. Therefore, we propose LiLi - a novel method that leverages Lie theory to identify the full set of degenerate transformations within the SE(3) Lie group of rigid transformations. The method employs perturbations of the optimized solution and compares the resulting optimized poses to ensure robust detection of degeneracies. By leveraging generators from the Lie algebra se(3), the method provides a systematic approach to describing the set of degenerate transformations. Quantitative evaluations on synthetic data show significant improvement over the state-of-the-art Hessian-based method, reducing alignment error by 50%, with more significant improvements for datasets featuring noise. In the real-world degenerate datasets, the proposed method integrated into LiDAR-based odometry yields superior localization performance compared to the reference solution based on the Hessian-based degeneracy detector on a 260 m long trajectory, and succeeds on a 430 m long round-trip tunnel trajectory where the reference fails.
Chinese Translation
本文研究存在退化的挑战性场景(如笔直走廊或平坦开阔地)中的3D激光雷达(LiDAR)扫描对齐问题。在这些场景中,对齐解不唯一,会降低定位与建图的精度。现有退化检测方法忽略了数据点重新关联的可能性,因而容易对噪声和复杂退化情况敏感。为此,我们提出LiLi——一种利用李理论(Lie theory)在刚体变换的SE(3)李群中识别完整退化变换集合的新方法。该方法通过对优化解施加扰动并比较所得优化位姿,实现对退化的鲁棒检测。通过利用李代数se(3)中的生成元,该方法为描述退化变换集合提供了一种系统化途径。在合成数据上的定量评估表明,该方法相较于最先进的基于Hessian的方法有显著提升,将对齐误差降低了50%,且在含噪声的数据集上提升更为明显。在真实世界的退化数据集中,将所提方法集成到基于激光雷达的里程计后,在260米长的轨迹上取得了优于基于Hessian退化检测器的参考方案的定位性能,并且在参考方案失败的430米长隧道往返轨迹上成功完成定位。
cs.RO / 43 / 2609.17147

Kernel-Based Metrics Learning for Uncertain Opponent Vehicle Trajectory Prediction in Autonomous Racing

基于核度量学习的自动驾驶赛车中不确定对手车辆轨迹预测
Lee, Hojin, Nam, Youngim, Lee, Sanghun, Kwon, Cheolhyeon
Abstract
Autonomous racing confronts significant challenges in safely overtaking Opponent Vehicles (OVs) that exhibit uncertain trajectories, stemming from unknown driving policies. To address these challenges, this study proposes heterogeneous kernel metrics for Deep Kernel Learning (DKL), designed to robustly capture the diverse driving policies of OVs, and carry out precise trajectory predictions along with the associated uncertainties. A key virtue of the proposed kernel metrics lies in their ability to align similar driving policies and disjoin dissimilar ones in an unsupervised manner, given the observed interactions between the Ego Vehicle (EV) and OVs. The efficacy of the proposed method is substantiated through experimental studies on a 1/10th scale racecar platform, demonstrating improved prediction accuracy and thereby safely overtaking against OVs. Furthermore, our method is computationally efficient for onboard computing units, affirming its viability in fast-paced racing environments. The video and source code can be found at https://github.com/HMCL-UNIST/OpponentPredictionWithKMDKL.git.
Chinese Translation
自动驾驶赛车在安全超越行驶轨迹不确定的对手车辆(Opponent Vehicles, OVs)时面临重大挑战,这种不确定性源于未知的驾驶策略。为应对这些挑战,本研究提出了一种用于深度核学习(Deep Kernel Learning, DKL)的异构核度量,旨在稳健地捕捉对手车辆多样化的驾驶策略,并实现精确的轨迹预测及相应的不确定性估计。所提核度量的一个关键优势在于,能够基于自车(Ego Vehicle, EV)与对手车辆之间观测到的交互,以无监督方式对相似的驾驶策略进行对齐、对不同者进行分离。通过在1/10比例赛车平台上的实验研究验证了所提方法的有效性,结果表明该方法提升了预测精度,从而能够安全超越对手车辆。此外,本方法计算效率高,适用于车载计算单元,证明了其在快节奏赛车环境中的可行性。视频与源代码可在 https://github.com/HMCL-UNIST/OpponentPredictionWithKMDKL.git 获取。
cs.RO / 44 / 2609.17168

HuMemSLAM: Efficient Human-Inspired Semantic Place Recognition for Robust Visual SLAM

HuMemSLAM:面向鲁棒视觉SLAM的高效类人语义位置识别
Adebambo, Mayowa, Donnelly, Sebastian, Amaritei, Armand, Bradley, Andrew, Rast, Alexander
Abstract
Autonomous systems require reliable place recognition for efficient and effective simultaneous localisation and mapping (SLAM). Traditional geometric visual SLAM approaches rely on low-level features and geometric consistency, but remain vulnerable to perceptual aliasing, where different places appear similar, and perceptual variation, where the same place appears different. Although semantic SLAM and modern learned visual place recognition (VPR) methods improve robustness under challenging perceptual conditions, real-time deployment requires both high retrieval accuracy and low latency. Inspired by human memory and perception, we propose HuMem-VPR, which exploits the bidirectional relationship between bottom-up perceptual evidence and top-down contextual reasoning to achieve high-level place understanding. We further introduce HuMemSLAM, the integration of HuMem-VPR with ORB-SLAM3. HuMem VPR achieved the highest aggregate retrieval accuracy on the real-image benchmark, competitive accuracy on the CARLA benchmark, and approximately two to three times lower latency than the evaluated state-of-the-art VPR methods. Across the evaluated dataset families and online experiments, HuMemSLAM substantially improved integrated Recall @1 over ORB-SLAM3's native retrieval while reducing the proposals submitted to its geometric backend.
Chinese Translation
自主系统需要可靠的位置识别,以实现高效有效的同步定位与建图(SLAM)。传统几何视觉SLAM方法依赖底层特征和几何一致性,但仍然容易受到感知混淆(不同地点看起来相似)和感知变化(同一地点看起来不同)的影响。尽管语义SLAM和现代基于学习的视觉位置识别(VPR)方法在具有挑战性的感知条件下提高了鲁棒性,但实时部署同时需要高检索精度和低延迟。受人类记忆与感知的启发,我们提出了HuMem-VPR,该方法利用自下而上的感知证据与自上而下的上下文推理之间的双向关系来实现高层次的位置理解。我们进一步提出了HuMemSLAM,即HuMem-VPR与ORB-SLAM3的集成。HuMem-VPR在真实图像基准上取得了最高的总体检索精度,在CARLA基准上取得了具有竞争力的精度,且延迟比所评估的最先进VPR方法低约两到三倍。在所评估的数据集系列和在线实验中,HuMemSLAM相比ORB-SLAM3的原生检索显著提高了集成Recall@1,同时减少了提交给其几何后端的候选提议数量。
cs.RO / 45 / 2609.17172

Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand

手指即腿:利用仿人机械手学习自支撑运动与操作
Kazemipour, Amirhossein, Zheng, Hehui, Katzschmann, Robert
Abstract
A walking robotic hand must use the same fingers to move its body, support its weight, and interact with the environment. We show how an anthropomorphic hand can learn these skills while retaining its finger design and position controller. Onboard power and computation make the platform self-contained. Our reinforcement learning approach accounts for the hand's unequal fingers, with training in a simulator calibrated from hardware measurements. In simulation, the hand moves faster with our reward formulation than with tuned rewards originally designed for quadrupeds. On hardware, task-specific policies enable untethered crawling, steering, and fall recovery. While supporting its own weight, the hand also executes successive keyboard commands without vision and pushes an object to targets using overhead visual feedback. These results demonstrate a compact mobile manipulator that reuses its fingers for locomotion and interaction, without a separate locomotion mechanism.
Chinese Translation
一只行走的机器人手必须使用相同的手指来移动身体、支撑自身重量并与环境交互。我们展示了仿人机械手(anthropomorphic hand)如何在保留其手指设计和位置控制器的同时学习这些技能。机载电源与计算使该平台完全独立运行。我们的强化学习方法考虑了该手部各手指不均衡的特点,并在基于硬件测量标定的仿真器中进行训练。在仿真中,采用我们的奖励设计时,手部的移动速度快于为四足机器人调优设计的原始奖励。在真实硬件上,任务特定的策略实现了无缆爬行、转向和摔倒恢复。在支撑自身重量的同时,该手部还能在没有视觉的情况下连续执行键盘指令,并借助顶部视觉反馈将物体推向目标位置。这些结果展示了一种紧凑的移动操作平台,它复用手指来完成运动与交互,而无需单独的运动机构。
cs.RO / 46 / 2609.17187

Fleet-To-Lab: A Transfer Learning Framework For Lunar Rover Slippage Estimation Via Model Fusion

Fleet-to-Lab:一种通过模型融合实现月球车滑移估计的迁移学习框架
Viviano, Riccardo, Omi, Saki, Orsula, Andrej, Olivares-Mendez, Miguel
Abstract
Accurate wheel slip estimation is essential for autonomous lunar rover mobility and navigation. Machine Learning models trained on terrestrial data generalize poorly to lunar terrain, and real lunar datasets are scarce due to the limited number of missions and costly data acquisition. We present Fleet-to-Lab, a transfer learning framework that leverages proprioceptive data collected by previously deployed heterogeneous lunar rovers to mitigate the Earth-Moon domain gap in slip estimation for a future deployable unit. We fuse several heterogeneous expert models into a single architecture, using a modest dataset collected after the rover deployment. We propose AcoMerge, a new hybrid swarm-intelligence algorithm that performs model fusion by searching for an optimal combi- nation of expert parameters. Experiments conducted in a high- fidelity physics simulation show balanced accuracy and macro- F1 improvements compared to deep model fusion baselines. AcoMerge exhibits competitive performance with joint training on deep architectures, while achieving higher macro-F1 and balanced accuracy on a smaller model. Overall, our framework shows model fusion as a possible transfer learning alternative for slippage estimation in space robotic missions with limited data.
Chinese Translation
精确的车轮滑移估计对自主月球车的移动与导航至关重要。基于地球数据训练的机器学习模型对月球地形的泛化能力较差,而由于任务数量有限且数据获取成本高昂,真实的月球数据集十分稀缺。我们提出了Fleet-to-Lab,这是一种迁移学习框架,它利用先前部署的异构月球车所采集的本体感知数据,来缓解未来可部署单元在滑移估计中的地月域差距。我们将多个异构专家模型融合到单一架构中,仅使用月球车部署后采集的少量数据集。我们提出了AcoMerge,一种新的混合群体智能算法,通过搜索专家参数的最优组合来实现模型融合。在高保真物理仿真中进行的实验表明,与深度模型融合基线相比,该方法在平衡准确率和宏F1分数上均有提升。AcoMerge的性能可与在深度架构上的联合训练相媲美,同时在更小的模型上取得了更高的宏F1分数和平衡准确率。总体而言,我们的框架表明,模型融合可以成为数据有限的空间机器人任务中滑移估计的一种可行的迁移学习替代方案。
cs.RO / 47 / 2609.17198

TIO-Former: Ultra-Lightweight 6-Directional ToF-Inertial Odometry for Nano-UAVs via a Streaming Causal Transformer

TIO-Former:基于流式因果Transformer的超轻量级六向ToF-惯性里程计,面向微型无人机
Liu, Yang, He, Yifan, Zhao, Wenhao, Mo, Xiangyu, Xu, Yang, Wei, Hao, Ma, Mingze, Li, Huan, Wu, Yifan, Dai, Zipeng, Zhou, Xin, Gao, Fei
Abstract
Autonomous nano-UAV navigation requires accurate ego-motion estimation under stringent size, weight, power, and computing (SWaP-C) constraints, where visual sensors and LiDARs exceed payload limits, optical flow degrades in low-texture scenes, and inertial-only state estimation is susceptible to accumulated drift. While multi-zone time-of-flight (ToF) arrays provide a lightweight metric complement, 6-DoF estimation from merely 384 ranges per frame is challenged by invalid returns, anisotropic observability, and temporal computational scaling. We propose TIO-FORMER, a camera-free, optical-flow-free, and mapless range-inertial odometry framework driven by an IMU and an ultra-lightweight (15 g) payload of six orthogonal 8 x 8 ToF arrays. Our frontend pairs consecutive range grids with a bilateral gated difference, while IMU-guided cross-attention dynamically routes directional features conditioned on platform kinematics. A Streaming Causal Transformer couples an uncompressed Local KV cache with compressed Chunk-FIFO memory, maintaining bounded inference cost and memory footprint independent of flight duration. In real-flight evaluations, TIO-FORMER reduces open-loop position error by 54.4% compared to nano-UAV optical flow and by 66.4%-89.1% over learned inertial baselines. We also evaluate performance across multiple environments and robustness under severe sensing degradation. Deployed on an edge RISC-V companion computer, TIO-FORMER achieves a P95 latency of 10.466 ms and peak resident memory of 6.324 MiB (less than 5 percent system RAM), demonstrating that sparse range sensing provides practical geometric anchoring for resource-constrained micro-aerial robots. Code is available at https://github.com/Ly041021/TIO-Former.
Chinese Translation
自主微型无人机(nano-UAV)导航需要在严格的尺寸、重量、功耗与计算(SWaP-C)约束下实现精确的自运动估计。在这些约束下,视觉传感器与激光雷达超出了载荷限制,光流法在低纹理场景中性能下降,而仅依赖惯性的状态估计容易产生累积漂移。多区域飞行时间(ToF)传感器阵列提供了一种轻量化的度量补充,但仅凭每帧384个距离测量值进行六自由度(6-DoF)估计面临无效回波、各向异性可观测性以及时间计算规模扩展等挑战。我们提出TIO-FORMER,这是一种无相机、无光流、无需地图的距离-惯性里程计框架,由一个惯性测量单元(IMU)和一个超轻量级(15克)的六组正交8×8 ToF阵列载荷驱动。我们的前端通过双边门控差分对相邻的距离网格进行配对,同时由IMU引导的交叉注意力根据平台运动学状态动态分配方向特征。流式因果Transformer(Streaming Causal Transformer)将未压缩的局部KV缓存与压缩的Chunk-FIFO记忆相结合,使推理成本和内存占用保持有界,且与飞行时长无关。在实际飞行评估中,TIO-FORMER相比微型无人机光流法将开环位置误差降低了54.4%,相比学习型惯性基线降低了66.4%–89.1%。我们还在多种环境中评估了其性能,以及严重感知退化下的鲁棒性。部署于边缘RISC-V协处理器上时,TIO-FORMER实现了10.466毫秒的P95延迟和6.324 MiB的峰值常驻内存(不足系统内存的5%),表明稀疏距离感知可为资源受限的微型飞行机器人提供实用的几何锚定。代码发布于 https://github.com/Ly041021/TIO-Former。
cs.RO / 48 / 2609.17210

FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence

FluxVLA引擎:面向具身智能的一站式VLA工程平台
Li, Yinhao, Mao, Weixin, Lan, Zihan, Rong, Jikun, Hu, Qirui, Zhang, Yiming, Deng, Weipeng, Shen, Bowen, Zhu, Minzhao, Mao, Yiming, Yang, Yan, Cui, Chenguang, Chen, Hongyuan, Huang, Xu, Zhao, Zheyi, Shen, Pinxi, He, Bozhen, Fu, Zhen, Wang, Yifan, Zhang, Zexin, Gao, Ang, Chen, Haoyu, Shi, Chengqi, Chen, Hua
Abstract
Vision-language-action (VLA) models, world-action models (WAMs), and offline reinforcement learning methods are rapidly expanding the design space of embodied policies, yet turning these algorithms into reliable robot systems remains constrained by fragmented data formats, training stacks, evaluation protocols, inference runtimes, and embodiment-specific interfaces. We present $\mathrm{FluxVLA}$ Engine, an open, configuration-driven platform that turns heterogeneous embodied-policy components into a reproducible data-to-deployment workflow. Rather than introducing another policy model, $\mathrm{FluxVLA}$ standardizes interfaces for datasets, visual-language and world models, action heads, reward- or advantage-weighted learning, distributed training, simulation evaluation, optimized inference, and robot operators. The engine further integrates compositional dual-arm simulation, scalable automatic data generation, and model-decoupled human-in-the-loop rollout, takeover, correction collection, and reward annotation. For responsive physical execution, it combines Real-Time Chunking (RTC) with accelerated inference backends, lightweight remote GPU serving, and configurable trajectory post-processing. Together, these capabilities connect offline learning, simulation validation, online correction, and real-robot execution through shared and auditable contracts. $\mathrm{FluxVLA}$ therefore targets the engineering bottlenecks separating promising embodied-learning algorithms from reproducible evaluation and dependable deployment. Code is available at https://github.com/FluxVLA/FluxVLA
Chinese Translation
视觉-语言-动作(VLA)模型、世界-动作模型(WAM)以及离线强化学习方法正在迅速拓展具身策略的设计空间,然而将这些算法转化为可靠的机器人系统仍受到碎片化的数据格式、训练框架、评估协议、推理运行时以及具身专属接口的制约。我们提出了FluxVLA引擎,这是一个开放的、配置驱动的平台,能将异构的具身策略组件转化为可复现的从数据到部署的工作流程。FluxVLA并未引入新的策略模型,而是为数据集、视觉-语言模型与世界模型、动作头、基于奖励或优势加权的学习、分布式训练、仿真评估、优化推理以及机器人操作器等环节标准化了接口。该引擎还进一步集成了可组合的双臂仿真、可扩展的自动数据生成,以及与模型解耦的人在回路(human-in-the-loop)的轨迹执行、接管、纠正数据采集和奖励标注。为实现响应灵敏的物理执行,它将实时分块(Real-Time Chunking, RTC)与加速推理后端、轻量级远程GPU服务以及可配置的轨迹后处理相结合。这些能力通过共享且可审计的契约,将离线学习、仿真验证、在线纠正和真机执行连接在一起。因此,FluxVLA旨在解决将有前景的具身学习算法与可复现评估及可靠部署分隔开来的工程瓶颈。代码可在 https://github.com/FluxVLA/FluxVLA 获取。
cs.RO / 49 / 2609.17240

Swim-and-Breach at Palm Scale: A Rudder-Steered Two-Propeller Underwater Robot Platform with Differential-Thrust Pitch Control

掌上尺度的游动与跃出:一种舵控双螺旋桨水下机器人平台及差动推力俯仰控制
Choi, Daehyun, Bergerson, Ian, Zhu, Hengjia, Lan, Tianjun, Bhamla, Saad
Abstract
We present a palm-scale (65 mm, 34 g) swim-and-breach robot platform. Two vertically stacked propellers provide both propulsion and differential-thrust pitch control under a proportional-integral-derivative (PID) loop, and a tail rudder adds yaw control. The hull, evaluated by flow simulation, reduces the drag five-fold relative to an equivalent cuboid, and the propellers are optimized using B-series modeling validated by dynamometer measurements. The current robot swims at 13.9 body lengths per second and turns at 209 deg per second, corresponding to the upper limits reported for underwater robots. In free swimming, the pitch loop turns the body to any commanded nose-up pitch angle, and, with the rudder stabilizing the exit, the current robot leaps 1.6 body lengths high and 3.7 long in a seamless cruise-leap-cruise sequence. The platform can be used to build small-scale robots that cross barriers and dry gaps between pools for inspection in streams, flooded structures, and industrial systems.
Chinese Translation
我们提出了一种掌上尺度(65毫米、34克)的游动与跃出机器人平台。两个垂直堆叠的螺旋桨在比例-积分-微分(PID)控制回路下同时提供推进力和差动推力俯仰控制,尾舵则提供偏航控制。通过流场仿真评估,该壳体相比等效长方体可将阻力降低五倍;螺旋桨采用B系列(B-series)建模进行优化,并通过测力计测量加以验证。当前机器人的游动速度达到13.9体长/秒,转向速度达到209度/秒,达到了水下机器人已报道性能的上限水平。在自由游动中,俯仰控制回路可使机体转向任意指定的抬头俯仰角,并且在尾舵稳定出水姿态的条件下,该机器人能以“巡航-跃出-巡航”的无缝动作序列跳出1.6倍体长的高度和3.7倍体长的距离。该平台可用于构建小型机器人,使其跨越障碍物并跃出水域之间的干燥间隙,从而在溪流、被淹没的建筑和工业系统中执行巡检任务。
cs.RO / 50 / 2609.17247

DriveMCP: An Agentic AI framework for Advanced Driver Assistance System

DriveMCP:面向高级驾驶辅助系统的智能体AI框架
Nadiri, Farzad, Cina, Mehdi, Rad, Ahmad B.
Abstract
An agentic AI driver-assistance framework that integrates perception, compliance reasoning, vehicle-state interpretation, and safety arbitration into a modular and auditable pipeline. The architecture, referred to as DriveMCP, incorporates a sensor-like perception stack alongside DriveLM as the vision-language front end to generate a graph-structured scene understanding (Graph Visual Question Answering) and language-grounded driving information. Key compliance elements in world_state, including posted speed limits and jurisdiction cues, are derived from DriveLM outputs through a structured parsing layer rather than being injected as simulator ground truth. A stateful orchestration layer coordinates specialized experts exposed as Model Context Protocol (MCP) servers: (i) a Rules server that performs retrieval-augmented compliance reasoning over jurisdiction-specific traffic codes and sign conventions, (ii) a Weather server that estimates traction risk and contextual speed advisories, and (iii) an MCP-CAN server that surfaces Controller Area Network (CAN)/On-Board Diagnostics (OBD) telemetry and diagnostic context for health-aware risk shaping. These outputs are fused to generate a structured decision that prompts a recommended course of action. The outcome is then further filtered by a Responsibility-Sensitive Safety (RSS)-inspired guardrail that arbitrates speak versus act decisions under bounded online adaptation. In CARLA simulation across multilingual, cross-border, and dynamic speed-limit scenarios, DriveMCP reduces traffic infractions and overspeed relative to the VLM-Direct, VLM-Direct+RAG, and VLM-Tools-NoArbiter baselines, while improving hazard response time and maintaining sub-second advisory latency.
Chinese Translation
本文提出一个智能体AI驾驶辅助框架,将感知、合规推理、车辆状态解读和安全仲裁整合为模块化且可审计的处理流水线。该架构称为DriveMCP,它结合了类似传感器的感知堆栈,并以DriveLM作为视觉-语言前端,生成图结构化的场景理解(Graph Visual Question Answering,图视觉问答)和基于语言的驾驶信息。world_state中的关键合规要素(如限速标志和管辖区提示)通过结构化解析层从DriveLM输出中提取,而非作为仿真器的真值注入。有状态编排层协调以模型上下文协议(MCP)服务器形式暴露的多个专用专家:(i)规则服务器,对特定管辖区的交通法规和标志规范执行检索增强的合规推理;(ii)天气服务器,估计牵引力风险并提供情境化的速度建议;(iii)MCP-CAN服务器,提供控制器局域网(CAN)/车载诊断(OBD)遥测数据和诊断上下文,用于健康感知的风险调整。这些输出被融合以生成结构化决策,提示推荐行动方案。随后,受责任敏感安全(RSS)启发的护栏对结果进行进一步过滤,在有界在线自适应条件下仲裁“提示”与“执行”决策。在CARLA仿真中的多语言、跨境和动态限速场景实验表明,相较于VLM-Direct、VLM-Direct+RAG和VLM-Tools-NoArbiter基线,DriveMCP减少了交通违规和超速行为,同时缩短了危险响应时间,并将建议延迟保持在亚秒级。
cs.RO / 51 / 2609.17249

Port-Hamiltonian Koopman Operator Synthesis for Mechanical Systems

面向机械系统的端口哈密顿Koopman算子综合
Singh, Rajpal, Singh, Aditya, Keshavan, Jishnu
Abstract
Finite-dimensional Koopman models enable efficient linear prediction and control of nonlinear robotic systems. However, models learned purely from trajectory data may violate the energetic structure of the underlying mechanics, producing predictions that exhibit artificial energy growth and diverge under recursive propagation. This work presents a structure-preserving Koopman framework for Euler-Lagrange systems built on generalized-momentum coordinates. The momentum transformation exposes the mechanical actuation as a known, state-independent port, which is preserved explicitly in the lifted dynamics. A structure-constrained neural architecture is developed to jointly learn the lifting functions and a port-Hamiltonian Koopman generator, rendering the learned dynamics passive by construction rather than through penalty terms or post-hoc projection. A Cayley-midpoint discretization further preserves the corresponding storage-dissipation balance exactly in discrete time. These properties are established analytically by deriving the discrete storage balance and associated stability guarantees of the learned predictor. Simulation and experimental studies demonstrate improved prediction accuracy, data efficiency, and closed-loop tracking over Koopman baselines, with increasing gains for higher-dimensional systems.
Chinese Translation
有限维Koopman模型能够对非线性机器人系统进行高效的线性预测与控制。然而,纯粹从轨迹数据学习得到的模型可能违背底层力学的能量结构,导致预测结果出现人为能量增长,并在递归传播下发散。本工作基于广义动量坐标,为欧拉-拉格朗日系统提出了一种保持结构的Koopman框架。该动量变换将机械执行输入显式地表示为一个已知的、与状态无关的端口,并在提升后的动力学中对其进行显式保持。我们开发了一种结构约束的神经网络架构,以联合学习提升函数与端口哈密顿Koopman生成器,使所学到的动力学在构造上即为无源的,而无需依赖惩罚项或事后投影。此外,采用Cayley中点离散化方法在离散时间下精确保持相应的存储-耗散平衡。本文通过推导所学预测器的离散存储平衡及相关稳定性保证,对这些性质进行了分析证明。仿真与实验研究表明,与Koopman基线方法相比,该方法在预测精度、数据效率和闭环跟踪性能上均有提升,且对于更高维系统增益更为显著。
cs.RO / 52 / 2609.17263

CAD-Based Relation Learning and Geometric-Symbolic Planning for Robotic Assembly

基于CAD的关系学习与几何-符号规划在机器人装配中的应用
Harlacher, Fabian, Friedrich, Christian
Abstract
Assembly Sequence Planning (ASP) remains a challenging problem due to its combinatorial nature, making exhaustive planning approaches impractical for complex industrial assemblies. Furthermore, many CAD models lack reliable semantic contact information or require extensive manual preprocessing, limiting the applicability of existing methods. This paper presents a hybrid ASP framework combining learning-based relation extraction with geometric-symbolic reasoning to generate feasible robotic disassembly sequences from imperfect CAD data. A neural network predicts semantic geometric relations from point clouds, while human-in-the-loop verification enables correction of uncertain predictions and planning failures. Extracted relations are transformed into a symbolic assembly graph, enabling a geometric-symbolic planner to efficiently compute locally valid sets of robotic manipulation primitives. A visibility-based ray-casting strategy guides the search for feasible disassembly directions without requiring an exhaustive combinatorial search, while the local solution space enables efficient sequence optimization. The framework is evaluated on an introduced assembly dataset and on the ASAP test dataset. On the ASAP test dataset, the proposed planner achieves an 85.83% planning success rate while reducing the median planning time by more than one order of magnitude across all assembly sizes and by more than a factor of 50 for assemblies with more than 30 components compared to the baseline. The results demonstrate that the proposed hybrid framework enables efficient robotic assembly sequence planning from imperfect CAD data while substantially reducing planning time. By combining learning-based feature segmentation, human-in-the-loop verification, and geometric-symbolic reasoning, the framework provides a practical foundation for scalable and adaptable robotic assembly and disassembly planning.
Chinese Translation
装配序列规划(Assembly Sequence Planning, ASP)由于其组合特性,仍然是一个具有挑战性的问题,使得穷举式规划方法难以应用于复杂的工业装配。此外,许多CAD模型缺乏可靠的语义接触信息,或需要大量人工预处理,限制了现有方法的适用性。本文提出了一种混合式ASP框架,将基于学习的关系提取与几何-符号推理相结合,从不完善的CAD数据中生成可行的机器人拆卸序列。神经网络从点云中预测语义几何关系,同时通过人机协同(human-in-the-loop)验证机制可以对不确定的预测和规划失败进行修正。提取的关系被转换为符号化装配图,使几何-符号规划器能够高效计算局部有效的机器人操作基元集合。基于可视性的射线投射策略引导搜索可行的拆卸方向,无需进行穷举的组合搜索,同时局部解空间支持高效的序列优化。该框架在一个新引入的装配数据集以及ASAP测试数据集上进行了评估。在ASAP测试数据集上,与基线方法相比,所提出的规划器达到了85.83%的规划成功率,并在所有装配规模下将中位规划时间缩短了一个数量级以上,对于超过30个部件的装配体更是缩短了50倍以上。结果表明,所提出的混合框架能够从不完善的CAD数据中高效地完成机器人装配序列规划,并显著降低规划时间。通过将基于学习的特征分割、人机协同验证与几何-符号推理相结合,该框架为可扩展、可适应的机器人装配与拆卸规划提供了实用的基础。
cs.RO / 53 / 2609.17292

Escape-Aware Control Barrier Functions for Quadrotor Safety under Body-Rate Limits

机体角速率约束下四旋翼安全的逃生感知控制障碍函数
Shi, Lei, Wen, Haosong, Liu, Qichao
Abstract
Control barrier functions for input-constrained systems place the admissible input set inside the definition of the safe set, yet the resulting barrier is almost always a function of the state alone; On a quadrotor this is not cosmetic: because the thrust vector must be reoriented before it can decelerate an approach, and reorientation is limited by the attainable body rate, a state-only barrier certifies states from which no escape is reachable in time; We characterize the certification gap in closed form and show its width is proportional to closing speed and inversely proportional to the body-rate limit; We then define an escape barrier on the augmented pair of state and previously applied input, with escape authority measured over the one-step reachable thrust cap; It admits a closed form and an analytic inverse for the maximum certifiable closing speed, and embeds in a predictive controller at no additional state cost; Across 550 paired closed-loop episodes on a 13-state quadrotor, the proposed controller completes every tested scenario, whereas the stopping-distance barrier enforced over the same horizon fails 15% and 25% of episodes in exactly the two scenarios that enter the predicted gap; Against an online backup-CBF baseline enforcing the same escape condition at the reached state, it holds a 29-74 degree larger directional margin and 3-18 times the clearance, and an independent conservative rollout referee finds no certified state from which escape fails.
Chinese Translation
针对输入受限系统的控制障碍函数通常将可容许输入集纳入安全集的定义之中,然而由此得到的障碍函数几乎总是仅依赖于状态;在四旋翼飞行器上,这并非无关紧要:由于推力矢量必须先重新定向才能实现减速,而重新定向受限于可达到的机体角速率,仅依赖状态的障碍函数所认证的某些状态实际上可能无法及时实现逃生。我们以闭式形式刻画了这一认证间隙,并证明其宽度与接近速度成正比、与机体角速率上限成反比。随后,我们在状态与先前施加输入的增广对上定义了逃生障碍,其逃生能力通过单步可达推力上限来度量。该障碍函数具有闭式表达以及最大可认证接近速度的解析逆,且可以嵌入预测控制器而不增加任何额外的状态代价。在四旋翼13状态模型上进行的550组成对闭环实验中,所提出的控制器完成了所有测试场景,而在相同预测时域内施加的停车距离障碍在恰好落入预测间隙的两个场景中分别有15%和25%的失败率。与在已达状态处施加相同逃生条件的在线备份CBF基线相比,所提方法保持了29至74度更大的方向裕度和3至18倍的间隙,并且独立的保守滚动评估器未发现任何认证状态存在逃生失败的情况。
cs.RO / 54 / 2609.17302

Online Geometric Change Detection via Scene Decomposition

基于场景分解的在线几何变化检测
Thorne, David, Chua, Samuel Jia Cong, Joshi, Nakul, Wong, Aiden, Robison, Christa S., Osteen, Philip, Lopez, Brett T.
Abstract
Autonomous robots are increasingly deployed on long duration single- and multi-session missions in dynamic environments, where the ability to identify environmental changes such as fallen trees or opened doors provides important contextual information for online planning. We propose a framework called Change Detection via Scene Decomposition (CDSD) for accurate online geometric change detection using LiDAR or RGB-D sensors. Recent advances in geometric SLAM have made it possible to generate dense, tightly aligned maps without post processing, but comparing global maps across entire sessions is computationally expensive and does not allow for single-session online change detection. CDSD instead spatially decomposes mapped environments into unique scenes where changes can be found efficiently by comparing dense, local subsets of the global map called submaps. As the first submap-based approach for geometric change detection, we identify and address the following core challenges: 1) identifying appropriate scenes for change detection that require minimal redundant information; 2) generating dense and representative submaps for each scene; 3) detecting changes between submaps with differing fields of view; and 4) processing detected changes for real-time map reconstruction. Results demonstrate our algorithm on custom datasets collected at the Army Research Laboratory facility in Graces Quarters, Maryland, and on open-source multi-session change detection datasets.
Chinese Translation
自主机器人越来越多地被部署于动态环境中的长时单次或多次任务中,在这些环境中,识别诸如倒下的树木或打开的门等环境变化的能力能够为在线规划提供重要的上下文信息。我们提出了一种名为基于场景分解的变化检测(Change Detection via Scene Decomposition, CDSD)的框架,用于利用激光雷达(LiDAR)或RGB-D传感器进行精确的在线几何变化检测。几何SLAM的最新进展使得无需后处理即可生成稠密且紧密对齐的地图成为可能,但对整个任务的全局地图进行比较在计算上代价高昂,且无法实现单次任务的在线变化检测。CDSD转而将已建图的环境在空间上分解为多个独立场景,通过比较全局地图中称为子图(submap)的稠密局部子集,从而高效地发现变化。作为首个基于子图的几何变化检测方法,我们识别并解决了以下核心挑战:1)识别适用于变化检测且冗余信息最少的场景;2)为每个场景生成稠密且具有代表性的子图;3)检测具有不同视野的子图之间的变化;4)处理检测到的变化以实现实时地图重建。我们在马里兰州Graces Quarters陆军研究实验室设施采集的自定义数据集以及开源的多次任务变化检测数据集上演示了我们的算法。
cs.RO / 55 / 2609.17372

XPACE: Joint World and Action Modeling from Heterogeneous Experience

XPACE:基于异构经验的世界与动作联合建模
Wei, Jiacheng, Bai, Jerry, Yue, Xiaoyu, Wang, Zidong, Guo, Xiaoyang, Chen, Cheng, Pu, Fanqi, Wu, Fan, Yue, Zhixu, Li, Yizhuo, Qiu, Feng, Liu, Bo, Ge, Yuying, Zhou, Hui, Chen, Chenyi, Ge, Yixiao
Abstract
A general-purpose robot needs to draw on diverse experience, choose actions, and anticipate how those actions will change the world. We introduce XPACE, a unified embodied world model that serves as both a world action model, jointly predicting executable robot actions and future video, and a world simulator, predicting the visual consequences of prescribed actions. Our key insight is that video prediction can both connect heterogeneous experience to action learning and generate new experience for policy improvement. With a shared video backbone between the policy and simulator, we use action-unlabeled video to learn visual dynamics and action-labeled human and robot demonstrations to jointly learn video and action prediction. Building on this architecture, a coarse-to-fine training curriculum progressively emphasizes robot control while retaining human experience, allowing the policy to learn behaviors beyond those covered by robot demonstrations. Beyond learning from recorded experience, XPACE uses its simulator to create additional recovery supervision for the policy. Specifically, we adapt the simulator to its own generated context, synthesize deviation-recovery trajectories around expert demonstrations, and fine-tune the policy on filtered recovery examples. Experiments on XPENG's IRON humanoid robot show that heterogeneous training improves robustness and enables transfer of human-observed skills to tasks absent from robot demonstrations, while recovery data generated by the model's own simulator further improves real-world task completion. Together, these results demonstrate how joint world and action modeling connects learning from heterogeneous experience with simulation-driven policy self-improvement.
Chinese Translation
通用机器人需要利用多样化的经验、选择动作并预判这些动作将如何改变世界。我们提出了XPACE,一个统一的具身世界模型,它既可以作为世界-动作模型,联合预测可执行的机器人动作和未来视频,也可以作为世界模拟器,预测给定动作的视觉后果。我们的核心洞察是:视频预测既能将异构经验与动作学习相连接,又能为策略改进生成新经验。通过在策略和模拟器之间共享视频骨干网络,我们利用无动作标签的视频学习视觉动力学,并利用带动作标签的人类和机器人示教数据联合学习视频预测与动作预测。基于该架构,我们采用由粗到细的训练课程,在保留人类经验的同时逐步强化机器人控制,使策略能够学习超出机器人示教覆盖范围的行为。除了从已有经验中学习外,XPACE还利用其模拟器为策略生成额外的恢复监督信号。具体而言,我们将模拟器适配到其自身生成的上下文中,围绕专家示教合成偏差-恢复轨迹,并在筛选后的恢复样本上微调策略。在小鹏(XPENG)IRON人形机器人上的实验表明,异构训练提升了鲁棒性,并使人类观察到的技能能够迁移到机器人示教中不存在的任务上;而由模型自身模拟器生成的恢复数据进一步提高了真实世界的任务完成率。这些结果共同展示了世界与动作联合建模如何将异构经验学习与仿真驱动的策略自我改进相结合。
cs.RO / 56 / 2609.17384

Exact Fusion and Coordinated Exploration in Multi-Robot Active Inference

多机器人主动推理中的精确融合与协调探索
Wu, Peng, Imani, Mohsen, Kamara, Amidu, Islam, Md Tamzeed, Ghoreishi, Seyede Fatemeh, Imani, Mahdi
Abstract
Robot teams that learn a common environment model exchange belief summaries and plan by the expected information gain of their actions. Under conjugate exponential-family beliefs the shared belief is counted once per robot at two points: at fusion, the product of local posteriors counts the common prior $n$ times, and at planning, every robot scores its plan under the same belief and the team converges on the same unknown. Both errors are removed by adding evidence increments to the shared natural parameter, realized increments at fusion and expected increments at planning. The expected increment of a committed teammate gives the next robot its conditional gain; corrected gains sum to the joint gain, the redundancy removed equals the total correlation of the planned observation streams, and sequential commitment keeps the $1/2$ greedy guarantee. The expected increment is exact for Gaussian beliefs with fixed sampling paths and for Dirichlet beliefs under the novelty approximation of discrete active inference, whose team objective has a closed concave form within an explicit bound of the exact mutual information, and fails for finite hypothesis classes, where a short exact enumeration replaces it. Experiments on cooperative RockSample, foraging, and field monitoring show that fusion correction leaves exploration redundancy unchanged, anticipated evidence removes it, and sequential commitment recovers most of the value of centralized joint planning at cost linear in the team size.
Chinese Translation
学习共同环境模型的机器人团队通过交换信念摘要并依据动作的期望信息增益进行规划。在共轭指数族信念下,共享信念在每个机器人处被重复计算两次:在融合时,局部后验的乘积将共同先验重复计算了 $n$ 次;在规划时,每个机器人都基于同一信念评估其计划,导致团队收敛到同一个未知目标。通过向共享的自然参数中加入证据增量可以消除这两类错误——融合时使用已实现的增量,规划时使用期望增量。已承诺队友的期望增量为下一个机器人提供了其条件增益;校正后的增益之和等于联合增益,所去除的冗余等于计划观测流的总相关性(total correlation),且顺序承诺机制保持了 $1/2$ 的贪心保证。期望增量对于固定采样路径的高斯信念是精确的,对于离散主动推理中在新颖性近似下的狄利克雷信念也是精确的——该团队目标具有闭式凹形式,且处于精确互信息的显式界内;而对于有限假设类,期望增量不再精确,此时可用简短的精确枚举来替代。在合作式 RockSample、觅食和野外监测任务上的实验表明:融合校正不会改变探索冗余,预期证据可消除该冗余,而顺序承诺能以与团队规模呈线性关系的代价恢复集中式联合规划的大部分价值。
cs.RO / 57 / 2609.17404

Residual Fault Adaptation for Dexterous In-Hand Manipulation Under Runtime Joint Faults

面向运行时关节故障的灵巧手内操作残差故障自适应方法
Deng, Linan, Liu, Xing, Hong, Lin, Hua, Feng, Ma, Guijun, Yue, Zuogong, Zhang, Fumin
Abstract
Dexterous in-hand manipulation requires coordinated control of multiple actuated joints, and a runtime joint fault can abruptly disrupt the contact configuration required for successful manipulation. In this work, we propose residual fault adaptation (RFA), a teacher-anchored framework for compensating for hidden command-channel faults. RFA retains a frozen healthy teacher to provide nominal behavior and trains a recurrent residual policy to infer corrective actions from proprioceptive and command-response history. During training, fault-injection domain randomization (FIDR) varies the fault mode, affected joint, severity, and onset time, while adaptive sampling increases the frequency of fault modes associated with lower recent performance. A frozen Direct FIDR policy provides a distributional reference only on fault-active training samples and is absent from deployment. The deployed controller receives neither fault labels nor controller-switching signals. Simulation experiments on the dexterous hand indicate that RFA can improve manipulation performance relative to the healthy policy under a fixed mixed-fault protocol. Real-robot experiments with software-injected faults further demonstrate zero-shot deployment of the learned adaptation policy.
Chinese Translation
灵巧手内操作需要对多个驱动关节进行协调控制,而运行时关节故障会突然破坏成功操作所需的接触构型。在本工作中,我们提出残差故障自适应(Residual Fault Adaptation, RFA),这是一种以教师为锚点的框架,用于补偿隐蔽的指令通道故障。RFA保留一个冻结的健康教师策略以提供标称行为,并训练一个循环残差策略,从本体感觉和指令响应历史中推断纠正动作。在训练过程中,故障注入域随机化(Fault-Injection Domain Randomization, FIDR)改变故障模式、受影响关节、严重程度和发生时间,同时自适应采样提高与近期较低性能相关的故障模式的采样频率。冻结的Direct FIDR策略仅在故障激活的训练样本上提供分布参考,且在部署阶段不使用。部署的控制器既不接收故障标签,也不接收控制器切换信号。在灵巧手上的仿真实验表明,在固定的混合故障协议下,RFA相较于健康策略能够提升操作性能。通过软件注入故障的真机实验进一步证明了所学习的自适应策略的零样本(zero-shot)部署能力。
cs.RO / 58 / 2609.17405

Optimized Wrench Polytope Analysis for Real-Time Stability Control of Legged Robots in Complex Multi-Contact Configurations

面向复杂多接触构型下腿式机器人实时稳定性控制的优化 wrench 多面体分析
Graaf, Friedrich, Birkefeld, Elias, Eichmann, Christian, Hofele, Elias, Schnell, Tristan, Heppner, Georg, Roennau, Arne, Dillmann, Rüdiger
Abstract
Legged robots offer a variety of automation applications in real-world scenarios. But areas that are difficult to traverse, like slopes, caves, or scaffolding, still pose a great challenge for traversal. To tackle this problem, we propose an optimized algorithm for evaluating the full actuatable wrench polytope for arbitrary contact scenarios. With our improved analysis algorithm, the torques for each joint of the robot can be calculated within a control frequency of 49 Hz. The achieved speedup allows for deployment within a regular control loop for actuating robot poses for different contact scenarios. We evaluated our stability controller extensively in simulation scenarios and validated its applicability by deploying it on actual walking robot hardware. The proposed controller achieved stability in very complex scenarios that are currently not achievable by any other controller.
Chinese Translation
腿式机器人在现实场景中具有多种自动化应用。但难以通行的区域,如斜坡、洞穴或脚手架,仍然对通行构成巨大挑战。为解决这一问题,我们提出了一种优化算法,用于评估任意接触场景下完整可驱动的 wrench 多面体(wrench polytope)。借助我们改进的分析算法,可在 49 Hz 的控制频率内计算机器人各关节的力矩。所实现的加速使得该算法能够部署于常规控制回路中,以针对不同接触场景驱动机器人姿态。我们在仿真场景中对所提出的稳定性控制器进行了广泛评估,并通过在实际行走机器人硬件上部署验证了其适用性。该控制器在目前任何其他控制器均无法应对的极复杂场景中实现了稳定性。
cs.RO / 59 / 2609.17430

Hamilton-Jacobi Reachability for Hybrid Systems: Unified Goal-Driven Control with Safety Guarantees

混合系统的哈密顿-雅可比可达性分析:具有安全性保证的统一目标驱动控制
Borquez, Javier, Peng, Shuang, Bansal, Somil
Abstract
Hybrid dynamical systems provide a powerful modeling framework for robotic systems, particularly in contact-rich environments. However, ensuring safety and performance in such systems remains challenging due to the intricate coupling between continuous dynamics and discrete mode transitions. In this work, we extend classical Hamilton-Jacobi (HJ) reachability analysis, a formal verification method for continuous-time nonlinear systems, to hybrid dynamical systems. Our framework characterizes safe sets for hybrid systems through a generalized value function defined over both discrete and continuous states while accounting for control constraints and model uncertainty. We additionally provide a numerical algorithm to compute this value function. Building on these safe sets, we propose two different mechanisms to integrate performance objectives. First, we introduce a hybrid least-restrictive safety filter that intervenes on both the discrete and continuous components of a nominal controller only when necessary to avoid unsafe states, thereby preserving nominal behavior whenever possible. Second, we formulate and compute hybrid backward reach-avoid tubes, enabling the simultaneous enforcement of safety and goal-reaching behavior, an extension not previously addressed within hybrid HJ reachability. This enables the synthesis of continuous and discrete control policies that guarantee both safety and task completion. We validate our framework through simulation studies and real-world experiments on a quadrupedal robot, demonstrating its effectiveness in hybrid mode planning and safety-critical applications.
Chinese Translation
混合动力系统为机器人系统提供了一个强大的建模框架,尤其适用于接触丰富的环境。然而,由于连续动力学与离散模式转换之间复杂的耦合关系,确保此类系统的安全性和性能仍然具有挑战性。在本工作中,我们将经典哈密顿-雅可比(Hamilton-Jacobi, HJ)可达性分析——一种针对连续时间非线性系统的形式化验证方法——扩展至混合动力系统。我们的框架通过定义在离散状态与连续状态之上的广义值函数来刻画混合系统的安全集,同时考虑控制约束与模型不确定性。此外,我们提供了一种用于计算该值函数的数值算法。基于这些安全集,我们提出了两种不同的机制来融合性能目标。首先,我们引入了一种混合最小限制安全滤波器,仅在必要时才对标称控制器的离散和连续分量进行干预以避免不安全状态,从而在可能的情况下尽可能保留标称行为。其次,我们提出并计算了混合后向可达避免管道(backward reach-avoid tubes),使得安全性与目标到达行为能够同时得到保证,这是混合HJ可达性分析中此前未曾解决的扩展。这使得能够综合出可同时保证安全性和任务完成的连续与离散控制策略。我们通过四足机器人的仿真研究和真实世界实验验证了该框架的有效性,展示了其在混合模式规划和安全关键应用中的有效性。
cs.RO / 60 / 2609.17463

Gaussian Processes for Modelling Spatial Fields with Robot Swarms

基于高斯过程的机器人集群空间场建模
Herranz, Guillermo Legarda, Francesca, Gianpiero, Birattari, Mauro
Abstract
Robot swarms, by virtue of their decentralised architecture, are a natural tool for scalable, robust modelling of spatial fields, such as water temperature, wind velocity, or terrain elevation. However, existing methods rely on external positioning systems that allow each robot to determine its own position in space. Here, we introduce location-unaware Gaussian process regression (LU-GPR) as a solution to the modelling of spatial fields in the absence of such positioning systems. LU-GPR allows each robot to infer the posterior mean and variance of the field in space, while simultaneously agreeing on a common frame of reference with its peers, using only local sensing and communication. We propose an online algorithm that allows each robot to consistently infer local estimates as its local frame of reference converges to the common one. By means of a product of experts model, each robot also combines the estimates of its peers with its own to obtain a global model. Our results show that LU-GPR scales well with the number of robots and is robust to limited communication ranges. We also demonstrate how it can be used in real-world monitoring scenarios to estimate the flow of an evacuating crowd.
Chinese Translation
机器人集群凭借其去中心化架构,是对水温、风速或地形海拔等空间场进行可扩展、鲁棒建模的天然工具。然而,现有方法依赖于外部定位系统,使每个机器人能够确定自身在空间中的位置。本文提出了位置未知高斯过程回归(LU-GPR),作为在缺乏此类定位系统情况下进行空间场建模的解决方案。LU-GPR 使每个机器人仅利用局部感知和通信,即可推断场在空间中的后验均值和方差,同时与其同伴就共同的参考坐标系达成一致。我们提出了一种在线算法,使每个机器人在其局部参考坐标系收敛到共同坐标系的过程中,能够一致地推断局部估计。通过专家乘积(product of experts)模型,每个机器人还将同伴的估计与自身的估计相结合,以获得全局模型。结果表明,LU-GPR 随机器人数量增加具有良好的可扩展性,并对有限的通信范围具有鲁棒性。我们还演示了如何在现实世界的监测场景中利用该方法估计疏散人群的流动。
cs.RO / 61 / 2609.17484

Dissecting Motion-Prior Regularization for Data-Scarce Robotic Insertion

面向数据稀缺机器人插入任务的运动先验正则化剖析
Hu, Ning, Li, Shuai, Tan, Jindong
Abstract
This study asks whether training-time motion-prior regularization can improve insertion success when a diffusion policy is learned from only 15 demonstrations. Minimum jerk discourages abrupt changes in predicted translational acceleration; speed-curvature regularization instead couples movement speed to path geometry. These are candidate mechanisms for task completion, not safety guarantees. We compare the priors individually and jointly, neither prior, and generic smoothness, with 80 real-robot trials per setting pooled over four recorded condition classes. Joint and minimum-jerk-only settings each achieved 70/80 successes (87.5%), versus 69/80 (86.3%) for speed-curvature only, 66/80 (82.5%) for neither prior, and 67/80 (83.8%) for generic smoothness. Success rates and Wilson 95% confidence intervals are visualized for direct comparison. Joint regularization exceeded neither by 5.0 percentage points but provided no observed gain over minimum jerk alone. The results motivate minimum jerk as the simpler candidate for replication, without establishing synergy, biomechanical specificity, improved safety, or distribution-shift robustness.
Chinese Translation
本研究探讨训练时运动先验正则化能否在扩散策略仅基于15条示范数据进行学习的情况下提升插入任务成功率。最小急动度(minimum jerk)先验抑制预测平移加速度的突变;速度-曲率正则化则将运动速度与路径几何耦合。这些是任务完成的候选机制,而非安全保障。我们在四种记录条件类别下,对每个设置各进行80次真实机器人试验,比较了各先验单独使用、联合使用、均不使用以及通用平滑正则化的效果。联合正则化与仅最小急动度设置均达到70/80次成功(87.5%),仅速度-曲率为69/80(86.3%),无先验为66/80(82.5%),通用平滑为67/80(83.8%)。我们通过可视化展示了各设置的成功率及Wilson 95%置信区间以便直接比较。联合正则化比无先验设置高出5.0个百分点,但相比仅使用最小急动度未见观测增益。结果表明最小急动度因其更简单而值得优先复现验证,但本研究并未确立两者的协同效应、生物力学特异性、安全性提升或分布偏移鲁棒性。
cs.RO / 62 / 2609.17524

Modality-Autoregressive World-Action Models

模态自回归世界-动作模型
Hung, Adam, Duisterhof, Bardienus P., Ramanan, Deva, Ichnowski, Jeffrey
Abstract
World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images. Other visual modalities such as depth, pretrained visual features, and point tracks can more efficiently capture geometric, semantic, and motion features. However, how best to combine these modalities within WAMs remains an open question. We introduce ModAR, the first WAM to autoregressively denoise multiple future modalities before predicting actions. This allows each prediction to condition on previously generated modalities. We train from scratch to systematically study how training-data mixtures, predicted modalities, and WAM formulations affect performance. In our evaluations, WAMs benefit from predicting point tracks, DINO features, and depth maps, while additionally predicting future RGB does not provide a consistent benefit. We also find that ModAR's sequential generation outperforms existing WAM formulations, with the highest average success rate at all evaluated data scales. We also fine-tune the video-model-initialized WAM Flex-$\pi$ on the same data; ModAR achieves a slightly higher observed average success rate (75% vs. 72%) while using approximately $20\times$ fewer training FLOPs and no pretraining. On three real-world bimanual tasks, ModAR outperforms baselines and improves with human videos.
Chinese Translation
世界-动作模型(World-Action Models, WAMs)联合建模未来观测与动作,通常以RGB图像的形式预测未来。其他视觉模态(如深度图、预训练视觉特征和点轨迹)能够更高效地捕捉几何、语义和运动特征。然而,如何在这些模态组合下最优地构建WAM仍是一个开放问题。我们提出了ModAR,这是首个在预测动作之前对多种未来模态进行自回归去噪的WAM。这使得每个预测都能以先前生成的模态为条件。我们从零开始训练,系统地研究训练数据混合、预测模态以及WAM构建方式如何影响性能。在评估中,WAM受益于预测点轨迹、DINO特征和深度图,而额外预测未来RGB并不能带来一致的性能提升。我们还发现,ModAR的顺序生成方式优于现有的WAM构建方法,在所有评估的数据规模上均取得了最高的平均成功率。我们还在相同数据上微调了由视频模型初始化的WAM Flex-$\pi$;ModAR在观察到略高的平均成功率(75% 对 72%)的同时,训练FLOPs减少约$20\times$,且无需预训练。在三个真实世界的双手操作任务上,ModAR优于基线方法,并能通过人类视频进一步提升性能。