← Back to Index
Daily Research Digest

arXiv Papers

2026-09-14
53
Papers
1
Categories
53
Translated
收藏清单 0
机器人学 (Robotics)
53
cs.RO / 1 / 2609.11947

Scenario-Independent Criticality Assessment and Prediction for Vulnerable Road Users in Autonomous Driving

自动驾驶中弱势道路使用者的场景独立关键性评估与预测
Gamerdinger, Jörg, Schwarzenberger, Victor, Schmid, Philipp, Teufel, Sven, Bringmann, Oliver
Abstract
Increasing safety is the primary objective of automated vehicles. Achieving this goal requires reliable safety metrics that incorporate safety-relevant factors such as object type, velocity, and criticality. A key capability of such metrics is the distinction between critical and non-critical objects, which is addressed through criticality or relevance estimation. Existing criticality metrics are typically designed for specific scenarios and primarily focus on vehicle-to-vehicle interactions. In this paper, we therefore propose a novel criticality metric tailored to vulnerable road users (VRUs), which require special consideration due to their less predictable motion behavior. Furthermore, to avoid the complexity introduced by scenario-specific metrics, we introduce a scenario-independent criticality prediction framework applicable to all traffic participant classes. The effectiveness of both the proposed VRU-centric criticality metric and the criticality prediction framework is evaluated using the DeepAccident dataset, which contains a diverse set of safety-critical traffic scenarios. The proposed VRU-centric criticality metric improves pedestrian criticality classification performance by up to 50 %. In addition, the proposed criticality prediction framework outperforms state-of-the-art metrics by 275 %, achieving an F1-score of 0.96 and enabling scenario-independent criticality assessment across all object classes. These results demonstrate the strong potential of the proposed approaches to enhance criticality assessment for safety evaluation in automated driving systems.
Chinese Translation
提高安全性是自动驾驶汽车的首要目标。实现这一目标需要可靠的安全指标,这些指标应纳入与安全相关的因素,如物体类型、速度和关键性。此类指标的一个关键能力是区分关键物体和非关键物体,这通过关键性或相关性估计来实现。现有的关键性指标通常针对特定场景设计,并且主要关注车对车交互。因此,在本文中,我们提出了一种针对弱势道路使用者(VRUs)的新型关键性指标,由于其运动行为较难预测,需要特别考虑。此外,为了避免特定场景指标带来的复杂性,我们引入了一个适用于所有交通参与者类别的场景独立关键性预测框架。我们使用DeepAccident数据集评估了所提出的以VRU为中心的关键性指标和关键性预测框架的有效性,该数据集包含多样化的安全关键交通场景。所提出的以VRU为中心的关键性指标将行人关键性分类性能提高了高达50%。此外,所提出的关键性预测框架比现有最先进的指标高出275%,达到了0.96的F1分数,并实现了跨所有物体类别的场景独立关键性评估。这些结果表明,所提出的方法在增强自动驾驶系统安全评估的关键性评估方面具有巨大潜力。
cs.RO / 2 / 2609.12036

Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence

Pelican-Sim 1.0:面向具身智能的通用世界模型模拟器
Zou, Shilong, Zhang, Shilin, Zhang, Yingji, Huang, Yuhang, Zhang, Yi, Ding, Zeyuan, Dong, Han, Liao, Junwei, Dai, Yong, Tang, Jian, Ju, Xiaozhu
Abstract
In this technical report, we propose Pelican-Sim 1.0, a general world model simulator for embodied intelligence that predicts future observations from visual context and robot actions to support downstream learning and decision making. The model incorporates four key design features: (1) Unified action representation: a 28-dimensional action value space covering most mainstream embodiments, keeping one model valid across heterogeneous devices. (2) Action-visual injection: URDF- and camera-rendered action videos bridge actions and pixels, giving markedly better controllability across embodiments, scenes, and tasks (PSNR +0.904 over alternative fusion baselines). (3) Sparse mixture-of-experts (MoE): sparse MoE layers add capacity for heterogeneous dynamics and absorb the action modality while reducing inter-modality conflict (FVD -6.530 vs. the dense backbone). (4) Efficient rollout generation: causal adaptation and few-step distillation yield a four-step autoregressive simulator, achieving a 5.67-fold speedup over the 35-step model. Benefiting from these designs, we train on approximately one million real-world and simulated trajectories and obtain large gains in action controllability and video quality: PSNR improves over the strongest evaluated baselines by 4.636 on AgiBotWorld Beta, 2.080 on RoboMIND, and 10.343 on RoboTwin, with the adapted EWMBench DYN score up 0.426 on RoboTwin. Relying on this, four downstream applications on RoboTwin succeed: 500 generated trajectories added to 50 demonstrations per task raise policy success from 70% to 93%; policy evaluation reaches a Pearson correlation of 0.994 across five checkpoints; and relative success gains reach 47.7% for action selection and 20.3% for policy improvement. Qualitative generalization across trajectory, scene, object, embodiment, and viewpoint shifts highlights its potential as a general-purpose world model simulator.
Chinese Translation
在本技术报告中,我们提出了 Pelican-Sim 1.0,一种面向具身智能的通用世界模型模拟器,它根据视觉上下文和机器人动作预测未来观测,以支持下游学习和决策。该模型包含四个关键设计特性:(1)统一动作表示:一个28维的动作值空间,覆盖大多数主流具身形态,使单一模型在异构设备上保持有效。(2)动作-视觉注入:URDF和相机渲染的动作视频桥接了动作与像素,在具身形态、场景和任务上显著提升了可控性(相较于其他融合基线,PSNR提升0.904)。(3)稀疏混合专家(MoE):稀疏MoE层为异构动力学增加了容量,并吸收了动作模态,同时减少了模态间冲突(相较于稠密骨干网络,FVD降低6.530)。(4)高效推演生成:因果适配和少步蒸馏产生了四步自回归模拟器,相比35步模型实现了5.67倍的加速。得益于这些设计,我们在约一百万个真实世界和模拟轨迹上进行训练,并在动作可控性和视频质量方面获得了大幅提升:在 AgiBotWorld Beta 上,PSNR 比最强的评估基线提高了4.636,在 RoboMIND 上提高了2.080,在 RoboTwin 上提高了10.343;在 RoboTwin 上,适配后的 EWMBench DYN 分数提高了0.426。基于此,RoboTwin 上的四个下游应用取得了成功:每个任务中,将500条生成的轨迹添加到50个演示中,使策略成功率从70%提升到93%;策略评估在五个检查点上达到0.994的皮尔逊相关系数;动作选择的相对成功率增益达到47.7%,策略改进的相对成功率增益达到20.3%。跨轨迹、场景、物体、具身形态和视角变化的定性泛化,凸显了其作为通用世界模型模拟器的潜力。
cs.RO / 3 / 2609.12081

MoPA: Coordinated Mobile Manipulation via Subsystem-Specific Perception Alignment

MoPA:通过子系统特定感知对齐的协调移动操作
Chen, Guangyu, Liang, Qiwei, Zhu, Shaolong, Chen, Tianxing, Xiao, Zikuan, Xie, Yifan, Zhang, Lingfeng, Luo, Ping, Xu, Renjing, Ding, Wenbo
Abstract
Mobile manipulation requires perceptual evidence at different spatial scales for base motion and arm control, while the two action modalities remain kinematically coupled. Existing policies often employ specialized action generation for different subsystems but condition heterogeneous action branches on a shared perceptual representation, leaving subsystem-specific perception-action correspondence implicit. We present MoPA, a framework that aligns perceptual conditioning with mobility and manipulation while preserving coordination at the action level. Dual Perceptual Streams employ two mutually masked query banks to extract separate perceptual representations from a shared vision-language context. Perception2Action Adaptation jointly updates each query bank and its corresponding action stream at every layer of a structured Mixture-of-Transformers decoder, while enabling information exchange between the two action streams. Coupled conditional flow matching learns a joint vector field for coordinated generation of both action chunks. On the ManiSkill-HAB benchmark, MoPA achieves state-of-the-art performance across all three task suites. Across four real-world tasks, MoPA achieves a mean full-task success rate of 76.3%, outperforming the best baseline by 12.5 percentage points. Ablation studies and further analyses validate the effectiveness of the proposed design. Website is available at: https://mopa-policy.github.io/.
Chinese Translation
移动操作需要为基座运动和机械臂控制提供不同空间尺度上的感知证据,同时这两种动作模态在运动学上保持耦合。现有策略通常为不同子系统采用专门的动作生成,但将异构动作分支条件于共享的感知表示上,使得子系统特定的感知-动作对应关系隐式化。我们提出MoPA,一个在动作层面保持协调的同时,将感知条件与移动和操作对齐的框架。双感知流采用两个相互掩码的查询库,从共享的视觉-语言上下文中提取独立的感知表示。Perception2Action Adaptation在结构化Mixture-of-Transformers解码器的每一层联合更新每个查询库及其对应的动作流,同时实现两个动作流之间的信息交换。耦合条件流匹配学习一个联合向量场,用于协调生成两个动作块。在ManiSkill-HAB基准上,MoPA在所有三个任务套件上均达到最先进性能。在四个真实世界任务中,MoPA实现了平均76.3%的全任务成功率,比最佳基线高出12.5个百分点。消融研究和进一步分析验证了所提出设计的有效性。网站:https://mopa-policy.github.io/。
cs.RO / 4 / 2609.12095

Uncertainty-Aware Conflict Detection Against Operator-Conditioned Weather Hazards

面向操作员条件化天气危险的不确定性感知冲突检测
Kandoria, Balram, Kim, Seulki, Samyal, Aryaman Singh
Abstract
Strategic flight plan validation in Advanced Air Mobility (AAM) environments requires robust methods for predicting aircraft state uncertainty and detecting potential conflicts with dynamic airspace hazards. This paper presents a novel framework for uncertainty-conditioned trajectory prediction combined with polyhedra hazard representation for pre-flight conflict detection. We introduce a closed-form uncertainty estimation method that couples non-uniform rational B-spline (NURBS) curve fitting for kinematic trajectory generation with a Kalman Filter for state covariance propagation. Drawing from the Light Propagation Algorithm (LPA) paradigm, we employ a sigmoid-blended measurement noise model that captures the uncertainty reduction behavior of flight management systems approaching the required time of arrival (RTA) for waypoints. The resulting temporal uncertainty bounds are derived through a velocity-to-time variance transformation, enabling probabilistic assessment of arrival time deviations along the flight path. For hazard representation, we develop an operator-conditioned classification scheme that transforms gridded environmental data, specifically weather phenomena, into three-dimensional polyhedra volumes with intensity-based stratification. These hazard polyhedra incorporate aircraft-specific safety buffers computed from vehicle performance characteristics. Conflict detection is performed through mesh intersection algorithms operating on the spatial uncertainty tube surrounding the mean trajectory against the hazard polyhedra and temporal overlap. The framework enables the continuous strategic validation of flight plans throughout the pre-flight planning time horizon as environmental conditions evolve.
Chinese Translation
先进空中交通(AAM)环境中的战略飞行计划验证需要鲁棒的方法来预测飞机状态不确定性并检测与动态空域危险的潜在冲突。本文提出了一种新颖的框架,用于不确定性条件化的轨迹预测,结合多面体危险表示,用于飞行前冲突检测。我们引入了一种闭式不确定性估计方法,该方法将用于运动学轨迹生成的非均匀有理B样条(NURBS)曲线拟合与用于状态协方差传播的卡尔曼滤波器相结合。借鉴光传播算法(LPA)范式,我们采用了一种S形混合测量噪声模型,该模型捕捉了飞行管理系统在接近航路点所需到达时间(RTA)时的不确定性降低行为。由此产生的时间不确定性边界通过速度-时间方差变换导出,使得能够对沿飞行路径的到达时间偏差进行概率评估。对于危险表示,我们开发了一种操作员条件化的分类方案,将网格化环境数据(特别是天气现象)转换为基于强度分层的三维多面体体积。这些危险多面体纳入了根据飞行器性能特征计算的飞机特定安全缓冲。冲突检测通过网格相交算法执行,该算法针对围绕平均轨迹的空间不确定性管与危险多面体及时间重叠进行操作。该框架能够在环境条件演变过程中,在整个飞行前规划时间范围内对飞行计划进行连续的战略验证。
cs.RO / 5 / 2609.12103

RodForesight: A World Model Enhanced Diffusion Policy for Slender and Material Agnostic Rod Insertion

RodForesight:一种世界模型增强的扩散策略用于细长且材料无关的杆插入
Yu, Chuanbo, Yue, Mingyu, Lyu, Yan, Song, Chuhan, Wang, Peng
Abstract
Slender rod insertion arises in precision manufacturing, where millimetre scale diameter and tight clearances demand accurate perception and control. Conventional peg-in-hole methods assume a rigid object whose tip pose is fixed relative to the gripper. This assumption breaks down for a high aspect ratio rod, which can bend during manipulation, making its tip motion dependent on the rod configuration, grasp, material properties, and contact. We present RodForesight, a learning framework that factorises the task into two stages: 1) coarse approaching, which uses visual servoing to map diverse initial configurations into a compact near hole hand-off region; and 2) predictive insertion, which performs fine alignment and completes the insertion. It is worth noting that the two stages can be wrapped into an end-to-end design. During insertion, a diffusion policy generates candidate action chunks, while an action conditioned world model predicts their effects on rod-hole alignment. This pre-execution evaluation enables RodForesight to select the best action chunk based on predicted tilt and radial errors before execution. Experiments investigate the performance of different stages and the end-to-end setting, where RodForesight improves the success rate from 88.9% to 96.7%, compared to baseline methods such as diffusion policy.
Chinese Translation
细长杆插入出现在精密制造中,其中毫米级直径和紧公差要求精确的感知和控制。传统的轴孔装配方法假设刚性物体,其末端位姿相对于夹爪固定。对于高长径比杆,这一假设失效,因为杆在操作过程中可能弯曲,使其末端运动取决于杆的构型、抓取、材料属性和接触。我们提出RodForesight,一个学习框架,将任务分解为两个阶段:1)粗接近,使用视觉伺服将多样的初始构型映射到紧凑的孔附近交接区域;2)预测插入,执行精对齐并完成插入。值得注意的是,这两个阶段可以封装为端到端设计。在插入过程中,扩散策略生成候选动作块,而动作条件世界模型预测它们对杆孔对齐的影响。这种执行前评估使RodForesight能够在执行前基于预测的倾斜和径向误差选择最佳动作块。实验研究了不同阶段和端到端设置的性能,其中RodForesight将成功率从88.9%提高到96.7%,与扩散策略等基线方法相比。
cs.RO / 6 / 2609.12108

Multi-Objective Agent-Based Model Predictive Controller for Plug-and-Play Vehicle Control

用于即插即用车辆控制的多目标基于智能体的模型预测控制器
Zhong, Jiaming, Khoshnevisan, Ladan, Huang, Shucheng, Pirani, Mohammad, Pant, Yash Vardhan, Khajepour, Amir
Abstract
Functional integration is a growing trend in vehicle control, often involving the coordination of multiple controllers to achieve various objectives simultaneously. The need for flexibility and reliability has led to a "plug-and-play" approach in control system design, which presents challenges for traditional integrated model predictive control (MPC). Agent-based model predictive control (AMPC) has recently emerged as a distributed solution that treats controllers as agents, creating a collaborative framework among them to reach a common goal. However, this approach struggles to manage distributed conflicting objectives when agents are coupled or interdependent. To address this, we propose a novel, practical distributed control scheme called multi-objective AMPC, which adapts the alternating direction method of multipliers (ADMM) into a general control strategy that approximates global optimization while decoupling objectives. We systematically develop three formulations that maintain convergence while addressing control regularization and inequality constraints, applying them to complex vehicle control systems for the first time. The proposed method has been tested on two vehicle control scenarios with a multi-objective topology. Different formulations are compared through simulations, and the most computationally efficient one was implemented on an electric vehicle for real-world evaluations. The results demonstrate that the proposed multi-objective AMPC can converge approximately to the same global optimum as integrated MPC with greater flexibility and the potential to reduce computational costs.
Chinese Translation
功能集成是车辆控制中日益增长的趋势,通常涉及多个控制器的协调以同时实现各种目标。对灵活性和可靠性的需求导致了控制系统设计中的“即插即用”方法,这给传统的集成模型预测控制(MPC)带来了挑战。基于智能体的模型预测控制(AMPC)最近作为一种分布式解决方案出现,它将控制器视为智能体,在它们之间创建一个协作框架以实现共同目标。然而,当智能体耦合或相互依赖时,这种方法难以管理分布式冲突目标。为了解决这个问题,我们提出了一种新颖、实用的分布式控制方案,称为多目标AMPC,它将交替方向乘子法(ADMM)适配为一种通用控制策略,在解耦目标的同时逼近全局优化。我们系统地开发了三种形式,在保持收敛性的同时处理控制正则化和不等式约束,并首次将其应用于复杂的车辆控制系统。所提出的方法已在两个具有多目标拓扑的车辆控制场景中进行了测试。通过仿真比较了不同的形式,并将计算效率最高的一个实现在电动汽车上进行实际评估。结果表明,所提出的多目标AMPC可以近似收敛到与集成MPC相同的全局最优解,具有更大的灵活性并有可能降低计算成本。
cs.RO / 7 / 2609.12142

A Data-Driven Distributed Control Scheme: Learning Multi-Objective Agent-Based MPC for Path-Tracking

数据驱动的分布式控制方案:学习多目标基于智能体的MPC用于路径跟踪
Zhong, Jiaming, Mehrizi, Reza Valiollahi, Pant, Yash Vardhan, Khajepour, Amir
Abstract
Agent-based model predictive control (AMPC) has recently been proposed for vehicle systems with various controllers, such as differential braking and torque vectoring, where controllers are regarded as distributed agents contributing to the same objective. However, this scheme is challenging in handling multiple conflicting objectives with coupled agents. A common approach for such tasks is the integrated MPC, where all objectives and agents are stacked together in one optimization. Nevertheless, as more agents and objectives are involved, the integrated MPC will face challenges like computational burdens and maintenance difficulties in practice. To this end, this paper proposes a learning multi-objective AMPC that can improve design flexibility and computing efficiency. First, under the assumption of information exchange, a multi-objective AMPC tailored from the alternating direction method of multipliers (ADMM) is proposed to decouple the system and achieve the same performance as the integrated scheme iteratively. Second, a learning-based method for initializing iterations is proposed to accelerate convergence. In addition, a data management method is proposed for real-time efficiency, and an authentication module is designed for learning reliability. We compare the proposed scheme against the integrated scheme via a combined path-tracking simulation for autonomous vehicles with various controllers. The proposed scheme achieves the same control performance as the integrated one while reducing the computational time by 43.5%. Furthermore, the learning-based method saves 88.6% more computational time than without learning, making it suitable for real-time implementation.
Chinese Translation
基于智能体的模型预测控制(AMPC)最近被提出用于具有各种控制器的车辆系统,例如差动制动和扭矩矢量分配,其中控制器被视为分布式智能体,为同一目标做出贡献。然而,该方案在处理具有耦合智能体的多个冲突目标时具有挑战性。此类任务的常见方法是集成式MPC,其中所有目标和智能体被堆叠在一个优化中。然而,随着更多智能体和目标的引入,集成式MPC在实际中将面临计算负担和维护困难等挑战。为此,本文提出了一种学习型多目标AMPC,可以提高设计灵活性和计算效率。首先,在信息交换的假设下,提出了一种根据交替方向乘子法(ADMM)定制的多目标AMPC,以解耦系统并迭代地达到与集成方案相同的性能。其次,提出了一种基于学习的迭代初始化方法以加速收敛。此外,提出了一种数据管理方法以提高实时效率,并设计了一个认证模块以保证学习可靠性。我们通过针对具有各种控制器的自动驾驶车辆的组合路径跟踪仿真,将所提方案与集成方案进行比较。所提方案在达到与集成方案相同控制性能的同时,将计算时间减少了43.5%。此外,基于学习的方法比不使用学习的方法节省了88.6%的计算时间,使其适合实时实现。
cs.RO / 8 / 2609.12159

A Physics-Based Closed-Loop Robotic Bioprinting Framework Towards Volumetric Muscle Loss Treatment

面向体积肌肉损失治疗的基于物理闭环机器人生物打印框架
Rezayof, Omid, Andrews, Jerin T., Zobeidi, Ehsan, Ghasemkhani, Ali, Kamaraj, Meenakshi, Tilton, Maryam, John, Johnson V., Alambeigi, Farshid
Abstract
Robotic bioprinting and Direct Ink Writing (DIW) are being explored towards the treatment of Volumetric Muscle Loss (VML). While previous studies have shown the importance of proper parameter selection on the print outcome, existing approaches often rely on time- and material-intensive design of experiments methods, or require large, well-curated datasets for training machine learning models. In this paper, we propose a physics-based closed-loop robotic bioprinting system capable of near real-time parameter adaptation. The system integrates a 3D point cloud camera and fully autonomous vision-based algorithms to provide quantitative evaluation of printed constructs. This evaluation is fed into a controller that adjusts printing parameters to achieve a desired bead thickness. To assess the framework's performance, four experimental configurations were tested, each repeated three times. In these tests, printing began from an arbitrary initial parameter value, and the controller was tasked with adjusting the parameters to reach the desired thickness. The system converged in all trials, achieving a tracking error below 0.5 mm within an average of 5.2 seconds from the start of printing. The low standard deviation of the converged pressure over different tests (0.04 bar on average) demonstrates robustness and repeatability. Additional experiments were conducted with the controller turned off, enabling direct comparison with open-loop DIW bioprinting, further confirming the effectiveness of the proposed closed-loop framework in achieving the desired bead geometry.
Chinese Translation
机器人生物打印和直接墨水书写(DIW)正被探索用于体积肌肉损失(VML)的治疗。虽然先前的研究已经表明适当的参数选择对打印结果的重要性,但现有方法通常依赖于耗时且耗材的实验设计方法,或者需要大型、精心整理的数据集来训练机器学习模型。在本文中,我们提出了一种基于物理的闭环机器人生物打印系统,能够进行近乎实时的参数自适应。该系统集成了3D点云相机和完全自主的基于视觉的算法,以提供对打印结构的定量评估。该评估结果被馈送到控制器中,控制器调整打印参数以达到所需的珠丝厚度。为了评估该框架的性能,测试了四种实验配置,每种重复三次。在这些测试中,打印从任意初始参数值开始,控制器的任务是调整参数以达到所需厚度。系统在所有试验中都收敛,从打印开始平均在5.2秒内实现低于0.5毫米的跟踪误差。不同测试中收敛压力的低标准差(平均0.04巴)证明了鲁棒性和可重复性。在关闭控制器的情况下进行了额外的实验,从而能够与开环DIW生物打印进行直接比较,进一步证实了所提出的闭环框架在实现所需珠丝几何形状方面的有效性。
cs.RO / 9 / 2609.12188

Battery-Aware Predictive Trajectory Planning and Control for Multirotors Under Disturbances

多旋翼在扰动下的电池感知预测轨迹规划与控制
Kidambi, Krishna Bhavithavya
Abstract
This paper presents a battery-aware predictive trajectory-planning and control framework for multirotors operating under spatially localized disturbances. Candidate trajectories are evaluated through closed-loop vehicle--motor--battery propagation, allowing disturbance-induced control demand, electrical energy, battery evolution, and terminal-voltage-dependent actuator capability to enter the planning process. % A reduced-order battery model is numerically benchmarked against an independently implemented Simscape equivalent-circuit reference, with a power NRMSE of $0.64\%$ and a cumulative-energy discrepancy below $0.7\%$. % In a $150$-s, $640$-m mission containing three disturbance regions, the selected trajectory reduces electrical energy consumption by $7.46\%$ and position-tracking RMSE by approximately $72\%$ relative to the disturbance-aware fixed-reference baseline. % Planner ablations show that battery-dependent terms are nonbinding at nominal SOC but alter the selected trajectory under a depleted-battery stress condition. % Execution with multiple feedback controllers further demonstrates that controller selection changes the tradeoff among tracking accuracy, energy consumption, and actuator utilization. % The results demonstrate the benefit of accounting for predicted closed-loop energetic and battery--actuator consequences during trajectory selection.
Chinese Translation
本文提出了一种面向在空间局部扰动下运行的多旋翼的电池感知预测轨迹规划与控制框架。候选轨迹通过闭环飞行器-电机-电池传播进行评估,使得扰动引起的控制需求、电能、电池演化以及端电压相关的执行器能力能够进入规划过程。降阶电池模型与独立实现的 Simscape 等效电路参考模型进行了数值基准测试对比,功率归一化均方根误差(NRMSE)为 0.64%,累积能量差异低于 0.7%。在包含三个扰动区域的 150 秒、640 米任务中,与扰动感知的固定参考基线相比,所选轨迹将电能消耗降低了 7.46%,位置跟踪均方根误差(RMSE)降低了约 72%。规划器消融实验表明,电池相关项在标称 SOC 下不起约束作用,但在电池耗尽压力条件下会改变所选轨迹。使用多个反馈控制器执行进一步表明,控制器选择改变了跟踪精度、能量消耗和执行器利用率之间的权衡。结果表明,在轨迹选择过程中考虑预测的闭环能量以及电池-执行器后果是有益的。
cs.RO / 10 / 2609.12206

An Automated Thickness Evaluation Procedure Using an Integrated Structured Light 3D Camera in a Robotic Bioprinting Framework

一种在机器人生物打印框架中使用集成结构光3D相机的自动化厚度评估流程
Zobeidi, Ehsan, Rezayof, Omid, Alambeigi, Farshid
Abstract
Bioprinting is emerging as a tissue engineering technique to replace common treatment methods for large scale injuries. While thickness of the BioPrinted Constructs (BPCs) have shown to be of importance in the cell maturation and integration, the literature lacks a robust, automated, and quantitative method for measuring these metrics. In this paper, we propose a fully automated vision-based method for measuring the thickness of the BPCs with complex geometries. Leveraging the point cloud and RGB images of a structured light 3D camera, our proposed method performs an image-based segmentation for delineating the BPCs from the RGB images, accompanied by novel geometry-based thickness measurement algorithms performed on the point cloud scans. These algorithms combine the segmentation mask with the robot's forward kinematics data and a 3D point cloud scan to precisely measure the aforementioned metrics for complex-shaped BPCs. The proposed method was evaluated in simulation and experimental studies. In simulation studies, the algorithms were used to measure the thickness of some virtually created BPCs with known thickness. The comparison between the measured and true thicknesses demonstrates the high accuracy of the proposed method, achieving mean absolute errors between 0.025 mm and 0.057 mm in simulation at a spatial resolution of 0.1 mm x 0.1 mm per pixel. Furthermore, we successfully deployed the algorithms on our robotic bioprinting setup utilizing a structure light 3D camera, where complex patterns were printed and the developed methods utilized to accurately measure the thickness of printed BPCs.
Chinese Translation
生物打印正逐渐成为一种组织工程技术,用于替代大面积损伤的常见治疗方法。尽管生物打印构建体(BPCs)的厚度已被证明在细胞成熟和整合中具有重要性,但文献中仍缺乏一种稳健、自动化且定量的测量这些指标的方法。在本文中,我们提出了一种全自动的基于视觉的方法,用于测量具有复杂几何形状的BPC的厚度。利用结构光3D相机的点云和RGB图像,我们提出的方法执行基于图像的分割,以从RGB图像中勾勒出BPC,并结合在点云扫描上执行的新的基于几何的厚度测量算法。这些算法将分割掩膜与机器人的正运动学数据和3D点云扫描相结合,以精确测量复杂形状BPC的上述指标。所提出的方法在仿真和实验研究中进行了评估。在仿真研究中,这些算法被用于测量一些已知厚度的虚拟创建的BPC的厚度。测量厚度与真实厚度的比较表明,所提出的方法具有高精度,在每像素0.1 mm x 0.1 mm的空间分辨率下,仿真中平均绝对误差在0.025 mm至0.057 mm之间。此外,我们成功地将这些算法部署到我们的机器人生物打印装置上,使用结构光3D相机,打印了复杂图案,并使用所开发的方法准确测量了打印的BPC的厚度。
cs.RO / 11 / 2609.12216

Guardrailed Meta-Agent Loops: Stress-Testing Policy Pinning, Budget Bounds, and Crash Recovery

有护栏的元智能体循环:压力测试策略固定、预算界限与崩溃恢复
Ma, Qinzhen, Wu, Jialin
Abstract
Self-improving agent workflows create an audit problem when the same controller can change both its behavior and the conditions under which that behavior is judged. We present GuardrailLoop, a simulation-based testbed that makes three operational contracts jointly testable: preservation of human-defined policy, compute accounting at every recorded execution prefix, and recovery of a specified scientific state after crashes. A hash-pinned policy fixes goals, scope, evaluation identity, budget, and release conditions; machine-directed evolution is restricted to a code-owned feature catalog and bounded knobs. The contribution is an executable boundary and an evaluation protocol that separates useful adaptation, state recovery, and repeated execution. In a paired 50-seed 2 x 2 study, round-stage growth changes target attainment by +1.00 and restricted mean compute to target by -56.97 simulated GPU-hours (95% paired-bootstrap interval [-58.91,-54.70]); idle growth has zero measured utility effect. Across 240 enumerated crash injections, all runs recover the defined outcome, but only 210 preserve the normalized trace: 30 pre-commit crashes repeat a planner call. Resource-drift, kill-switch, integrity, and output-guard matrices satisfy their specified checks. These findings show why successful outcome recovery is insufficient evidence of exactly-once execution. They establish conformance within one calibrated deterministic testbed, rather than general safety or real-world self-improvement.
Chinese Translation
自我改进的智能体工作流会创造一个审计问题:当同一个控制器可以改变其行为以及评判该行为的条件时。我们提出了 GuardrailLoop,一个基于模拟的试验台,使三个操作契约可以联合测试:人类定义策略的保持、在每个记录的执行前缀处的计算核算,以及崩溃后恢复指定的科学状态。一个哈希固定的策略固定了目标、范围、评估标识、预算和发布条件;机器导向的演化被限制在一个代码拥有的特征目录和有界旋钮内。贡献是一个可执行的边界和一个评估协议,该协议将有用的适应、状态恢复和重复执行分开。在一项配对的 50 种子 2x2 研究中,轮次阶段增长将目标达成改变了 +1.00,并将到达目标的受限平均计算量改变了 -56.97 模拟 GPU 小时(95% 配对自举区间 [-58.91,-54.70]);空闲增长对效用没有可测量的影响。在 240 次枚举的崩溃注入中,所有运行都恢复了定义的结果,但只有 210 次保持了归一化轨迹:30 次预提交崩溃重复了一次规划器调用。资源漂移、终止开关、完整性和输出防护矩阵满足其指定的检查。这些发现说明了为什么成功的结果恢复不足以证明恰好一次执行。它们在一个校准的确定性试验台内建立了符合性,而不是一般安全性或真实世界的自我改进。
cs.RO / 12 / 2609.12221

Chain-SLAM: Globally Consistent Backend for Multi-Session LiDAR SLAM via Chained Loop Closure

Chain-SLAM:通过链式回环实现多会话 LiDAR SLAM 的全局一致后端
Li, Zhiheng, Liu, Xinhao, Zhang, Juexiao, Liang, Yongqing, Feng, Chen
Abstract
Maintaining consistency over long spatial and temporal horizons remains a fundamental challenge in large-scale LiDAR SLAM, particularly when integrating maps collected across multiple sessions. We present Chain-SLAM, a LiDAR SLAM backend enabling online multi-session map alignment and reuse with global consistency at large scale. We implement a chained loop closure mechanism that efficiently propagates geometric constraints across inter-session keyframes through an adjacency graph, enabling robust long-horizon consistency triggered by reliable short-horizon loop closures. The system initializes inter-session alignment with GNSS-proximity place recognition, then performs on-the-fly loop closure detections and joint optimization of loaded maps and newly acquired trajectories within a unified factor graph, maintaining both inter- and intra-session geometric consistency without dynamic object removal, and cross-platform robustness with minimal hyperparameter tuning. Experimental results show improved trajectory accuracy and robust multi-session integration on large-scale datasets. We release our source code to support reproducible research in large-scale multi-session LiDAR SLAM. Project site: https://ai4ce.github.io/Chain-SLAM/
Chinese Translation
在大规模 LiDAR SLAM 中,保持长时间和空间范围的一致性仍然是一个基本挑战,特别是在整合跨多个会话收集的地图时。我们提出了 Chain-SLAM,一个 LiDAR SLAM 后端,能够在大规模下实现具有全局一致性的在线多会话地图对齐和重用。我们实现了一种链式回环机制,通过邻接图在跨会话关键帧之间高效传播几何约束,从而通过可靠的短时回环触发稳健的长时一致性。该系统使用 GNSS 邻近地点识别初始化会话间对齐,然后在统一因子图内实时执行回环检测,并对加载的地图和新采集的轨迹进行联合优化,无需动态物体去除即可保持会话间和会话内的几何一致性,并以最少的超参数调整实现跨平台鲁棒性。实验结果表明,在大规模数据集上,轨迹精度得到提高,多会话集成具有鲁棒性。我们发布了源代码,以支持大规模多会话 LiDAR SLAM 的可重复研究。项目网站:https://ai4ce.github.io/Chain-SLAM/
cs.RO / 13 / 2609.12245

DIA: Denoising Intermediate Advantage for Diffusion Policy Optimization

DIA:面向扩散策略优化的去噪中间优势
Sohal, Arjun, Zhao, Yuchi, Bogdanovic, Miroslav, Aspuru-Guzik, Alan
Abstract
Diffusion-based robot policies have become widely used in robotic manipulation, where they are typically trained with behavior cloning. However, policies trained purely from demonstrations are limited by the quality and coverage of the available data. Reinforcement learning can further improve the performance of these pretrained policies through interaction. A common approach is to use policy-gradient methods that formulate diffusion-policy fine-tuning as an outer environment MDP together with an inner denoising MDP. However, existing methods typically assign the same environment-level credit to all denoising steps used to construct an action chunk, without distinguishing which intermediate decisions contributed most to the final return. We introduce Denoising Intermediate Advantage (DIA), a policy-gradient method that learns a value function over partially denoised actions and uses it to construct a denoising- level advantage for each step of the generative process. DIA com- bines this inner credit signal with the standard environment-level PPO advantage, providing state-dependent credit throughout the denoising chain. Across Robomimic, FurnitureBench, Franka Kitchen, and D3IL, DIA consistently improves final performance over existing diffusion-policy fine-tuning methods. Beyond final reward, DIA reaches successful states more efficiently and can shift farther from the pretrained behavior distribution, enabling it to discover more effective and efficient task-level strategies and subtask sequences that baseline methods fail to reach.
Chinese Translation
基于扩散的机器人策略已在机器人操作中广泛使用,通常通过行为克隆进行训练。然而,仅从演示中训练的策略受限于可用数据的质量和覆盖范围。强化学习可以通过交互进一步提升这些预训练策略的性能。一种常见方法是使用策略梯度方法,将扩散策略微调形式化为一个外层环境MDP和一个内层去噪MDP。然而,现有方法通常对用于构建动作块的所有去噪步骤赋予相同的环境级信用,而不区分哪些中间决策对最终回报贡献最大。我们提出去噪中间优势(DIA),一种策略梯度方法,它学习部分去噪动作上的价值函数,并利用其为生成过程的每一步构造去噪级优势。DIA将这种内部信用信号与标准的环境级PPO优势相结合,在整个去噪链中提供状态相关的信用。在Robomimic、FurnitureBench、Franka Kitchen和D3IL上,DIA相较于现有扩散策略微调方法持续提升最终性能。除了最终奖励之外,DIA能更高效地到达成功状态,并且可以更远地偏离预训练行为分布,从而发现基线方法无法达到的更有效、更高效的任务级策略和子任务序列。
cs.RO / 14 / 2609.12248

Trajectory Bundle Method in SE(3) for Black-Box Fixed-Wing Aircraft Trajectory Optimization

SE(3)中的轨迹束方法用于黑箱固定翼飞机轨迹优化
Osburn, Matthew D., Peterson, Cameron K., Salmon, John L.
Abstract
Dynamically feasible trajectory optimization for rigid-body systems is naturally formulated on the special Euclidean group SE(3) but is challenging when dynamics are available only as black-box computations without derivatives. This paper formulates the Trajectory Bundle Method (TBM) for motion planning implicitly on SE(3). Bundles are constructed in the Lie algebra and propagated through nonlinear rigid-body dynamics using exponential and logarithmic maps, enabling derivative-free planning of non-Euclidean trajectories. We show that Euclidean TBM interpolation error is bounded quadratically by bundle diameter and extend this result to SE(3), where the bound additionally depends on a local Lipschitz constant of the Log map. Numerical experiments corroborate these bounds. Finally, we demonstrate SE(3) TBM by optimizing an acrobatic, collision-free fixed-wing maneuver through a rotated aperture without explicit models or derivatives of the vehicle dynamics, aerodynamics, or collision model.
Chinese Translation
刚体系统的动态可行轨迹优化自然地在特殊欧几里得群SE(3)上表述,但当动力学仅以无导数的黑箱计算形式可用时,这具有挑战性。本文隐式地在SE(3)上为运动规划制定了轨迹束方法(TBM)。束在李代数中构造,并通过指数和对数映射在非线性刚体动力学中传播,从而实现了非欧几里得轨迹的无导数规划。我们证明了欧几里得TBM插值误差由束直径二次界定,并将该结果推广到SE(3),其中该界还依赖于Log映射的局部利普希茨常数。数值实验证实了这些界。最后,我们通过优化一个穿过旋转孔径的特技、无碰撞固定翼机动来演示SE(3) TBM,而无需车辆动力学、空气动力学或碰撞模型的显式模型或导数。
cs.RO / 15 / 2609.12258

Pneumatic neurons for soft robots enable inflate-and-fire networks for rhythmic motion

用于软体机器人的气动神经元实现节律运动的充气-激发网络
Li, Dongting, Tolley, Michael, Gravish, Nick
Abstract
Animals coordinate their movements through distributed neural circuits, but soft robots still typically depend on external, centralized electronics for control. Building soft robots that operate without centralized electronic controllers while remaining responsive to their environment remains a frontier challenge in soft robotics. In this work we introduce a soft-robot control architecture inspired by leaky integrate-and-fire models of biological neural circuits. The Pneumatic neuron (Pneu-ron) is a soft actuator that unifies energy conversion, logic, and actuation in one component. Each module combines a low-boiling-point fluid (LBF), a heater, and a mechanical switch into a self-excitable unit. Boiling the LBF inflates the module and triggers excitation and inhibition of adjacent modules in a process we call "inflate-and-fire". When interconnected into excitatory-inhibitory rings, Pneu-rons generate stable, sequential oscillations whose frequency emerges from the material dynamics and environmental conditions. By harnessing the inflation of Pneu-rons for actuation these networks can drive oscillatory locomotion of soft robots. Pneu-ron networks sustain oscillation under mechanical load and thermal variations, adapting through material physics rather than computation. Dynamical modeling of these networks reveals a dimensionless bifurcation diagram that dictates the network's oscillatory behavior. Encoding logic and actuation into material-level modules presents a new avenue for adaptive, electronics controller-free, soft robots.
Chinese Translation
动物通过分布式神经回路协调其运动,但软体机器人通常仍依赖外部集中式电子设备进行控制。构建无需集中式电子控制器、同时能对环境保持响应的软体机器人,仍然是软体机器人领域的前沿挑战。在这项工作中,我们引入了一种受生物神经回路泄漏整合-发放模型启发的软体机器人控制架构。气动神经元(Pneu-ron)是一种软致动器,将能量转换、逻辑和致动统一于单一组件中。每个模块将低沸点流体(LBF)、加热器和机械开关组合成一个自激发单元。使LBF沸腾会使模块膨胀,并触发相邻模块的兴奋和抑制,我们将这一过程称为“充气-激发”(inflate-and-fire)。当互连成兴奋-抑制环时,Pneu-ron产生稳定的、顺序振荡,其频率由材料动力学和环境条件决定。通过利用Pneu-ron的膨胀进行致动,这些网络可以驱动软体机器人的振荡运动。Pneu-ron网络在机械负载和热变化下维持振荡,通过材料物理而非计算进行适应。这些网络的动力学建模揭示了一个无量纲分岔图,它决定了网络的振荡行为。将逻辑和致动编码到材料级模块中,为自适应、无电子控制器的软体机器人提供了一条新途径。
cs.RO / 16 / 2609.12285

AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation

AnchorVLN:面向开放词汇导航的几何锚定视觉-语言接地推理
Vu, Long Giang, Yao, Chengkai, Liu, Yuxin, Aryan, FNU, Aralikatti, Rajath Chandrashekar
Abstract
Vision-Language Navigation (VLN) in unseen indoor environments is useful in real-world robotics, where an agent must follow natural-language instructions, locate objects, and answer spatial questions without a pre-built map or fixed object vocabulary. Multimodal vision-language models (VLMs) provide strong open-vocabulary grounding and zero-shot reasoning, but struggle to emit reliable metric quantities such as range, bearing, and comparative spatial relations directly from images. Existing approaches address this by folding geometry into hand-engineered pipelines or asking models to output waypoints, requiring changes to the control stack for different robots, tasks, or vocabularies. We introduce AnchorVLN, an open-vocabulary VLN system built on a simple rule: the VLM proposes semantics; geometry decides metrics. It is realised as EMBODIED-NAV-MCP, a Model Context Protocol (MCP) server driven by a VLM agent through a compact set of callable tools. Since no tool accepts distance in metres or bearing in radians, the schema enforces the semantic-geometry boundary without modifying the downstream autonomy stack. We benchmark both tasks of the CMU Vision-Language Navigation Challenge 2026: 30 instruction-following questions over 15 scenes and a frozen 45-question object-reference set. The full system achieves 64.4 percent on instruction following, dropping by 13.3 percentage points without controller modeling (t = 2.77). On object reference, geometric anchoring clears the challenge overlap threshold on 10 of 45 questions, versus 0 of 45 for direct coordinate estimation, reducing median center error from 3.37 m to 2.48 m.
Chinese Translation
在未见过的室内环境中的视觉语言导航(VLN)在现实世界机器人中很有用,其中智能体必须遵循自然语言指令、定位物体并回答空间问题,而无需预先构建的地图或固定的物体词汇表。多模态视觉语言模型(VLM)提供了强大的开放词汇接地和零样本推理,但难以直接从图像中输出可靠的度量量,如距离、方位和比较空间关系。现有方法通过将几何信息融入手工设计的流程或要求模型输出路径点来解决这个问题,但这需要针对不同的机器人、任务或词汇表更改控制栈。我们介绍了AnchorVLN,一个开放词汇VLN系统,基于一个简单的规则:VLM提出语义;几何决定度量。它实现为EMBODIED-NAV-MCP,一个由VLM智能体通过一组紧凑的可调用工具驱动的模型上下文协议(MCP)服务器。由于没有工具接受以米为单位的距离或以弧度表示的方位,该模式强制了语义-几何边界,而无需修改下游自主栈。我们对CMU视觉语言导航挑战赛2026的两个任务进行了基准测试:15个场景中的30个指令跟随问题和固定的45个问题对象参考集。完整系统在指令跟随上达到64.4%,在没有控制器建模的情况下下降了13.3个百分点(t = 2.77)。在对象参考上,几何锚定在45个问题中的10个上通过了挑战重叠阈值,而直接坐标估计为0个,将中位中心误差从3.37米降低到2.48米。
cs.RO / 17 / 2609.12292

Mission Performance: Automatic and Adaptive Race Pace Progression for Autonomous Racing

任务性能:自动驾驶赛车的自动自适应比赛节奏推进
Lambertini, Giovanni, Pini, Matteo, Musiu, Nicola, Raji, Ayoub, Iacovacci, Francesco, Bertogna, Marko
Abstract
In this paper, we describe the Mission Performance module implemented for a fully autonomous racing car to automatically manage the longitudinal, lateral, and combined performances, aiming to speedup the laptime progression while assuring safety. Motivated by the difficulty and risks of applying the real-time estimation of the grip to critical modules like the motion planner and controller, the Mission Performance guides these modules adapting their target performance instead of changing the vehicle model parameters. The module is formed by pre-defined progressions to warm up the tires at the beginning of a run. Then, the system continuously monitors safety and vehicle dynamics metrics on a per-sector basis to adaptively reduce, maintain, or increase the performance levels for each sector, progressively converging toward the maximum allowed value. The solution's effectiveness is demonstrated on the EAV-25, a fully autonomous Dallara Superformula, at the Yas Marina Circuit during the Abu Dhabi Autonomous Racing League (A2RL) Season 2.
Chinese Translation
在本文中,我们描述了为全自动驾驶赛车实现的Mission Performance模块,该模块自动管理纵向、横向和组合性能,旨在加快圈速提升的同时确保安全。由于将抓地力的实时估计应用于运动规划器和控制器等关键模块存在困难和风险,Mission Performance通过调整这些模块的目标性能来引导它们,而不是改变车辆模型参数。该模块由预定义的渐进方案组成,用于在运行开始时预热轮胎。然后,系统基于每个路段持续监控安全性和车辆动力学指标,以自适应地降低、保持或提高每个路段的性能水平,逐步收敛到最大允许值。该解决方案的有效性在阿布扎比自动驾驶赛车联盟(A2RL)第二赛季期间,于亚斯码头赛道(Yas Marina Circuit)上,通过EAV-25——一辆全自动驾驶的Dallara Superformula赛车——得到了验证。
cs.RO / 18 / 2609.12316

DATAFARM: Distribution-Aligned Task and Motion Planning for Fine-Tuning Vision-Language-Action Models

DATAFARM:面向微调视觉-语言-动作模型的分布对齐任务与运动规划
Sahoo, Samrat, Huang, Yixuan, Silver, Tom
Abstract
Collecting high-quality robot data remains a fundamental challenge for training robot foundation models. Task and motion planning (TAMP) offers a scalable way to generate demonstrations, but our experiments show that raw TAMP trajectories provide surprisingly little benefit when used to fine-tune pretrained vision-language-action (VLA) models, despite successfully solving the target tasks. We hypothesize that this failure arises from a behavioral distribution mismatch between planner-generated trajectories and the data used to pretrain the VLA. To address this mismatch, we introduce DATAFARM: Distribution-Aligned Task And motion planning for Fine-tuning A Robot foundation Model, an approach that incorporates the pretraining distribution directly into TAMP trajectory generation. DATAFARM aligns generated trajectories with the pretraining data in robot joint configurations, motion style, and temporal execution profiles. We evaluate DATAFARM on three tabletop manipulation tasks that TAMP can perform and a cloth-folding task beyond the capability of TAMP. DATAFARM achieves an average success rate of 56.7%, substantially outperforming raw TAMP (8.3%) while approaching human teleoperation (61.7%). On Deformable Object Manipulation, which is outside the fine-tuning distribution, the fine-tuned model retains 85% success, compared with 90% for the pretrained model. These results show that aligning planner-generated demonstrations with the pretraining distribution can make TAMP an effective source of data for VLA fine-tuning. Website and code: https://prpl-group.com/datafarm/
Chinese Translation
收集高质量的机器人数据仍然是训练机器人基础模型的一个根本挑战。任务与运动规划(TAMP)提供了一种可扩展的方式来生成演示,但我们的实验表明,尽管原始TAMP轨迹成功解决了目标任务,但在用于微调预训练的视觉-语言-动作(VLA)模型时,其带来的收益出人意料地小。我们假设这种失败源于规划器生成的轨迹与用于预训练VLA的数据之间的行为分布不匹配。为了解决这种不匹配,我们提出了DATAFARM:分布对齐的任务与运动规划用于微调机器人基础模型(Distribution-Aligned Task And motion planning for Fine-tuning A Robot foundation Model),一种将预训练分布直接纳入TAMP轨迹生成的方法。DATAFARM在机器人关节配置、运动风格和时间执行轮廓方面将生成的轨迹与预训练数据对齐。我们在三个TAMP可以执行的桌面操作任务和一个超出TAMP能力的叠布任务上评估了DATAFARM。DATAFARM达到了56.7%的平均成功率,显著优于原始TAMP(8.3%),并接近人类遥操作(61.7%)。在微调分布之外的可变形物体操作任务上,微调后的模型保留了85%的成功率,而预训练模型为90%。这些结果表明,将规划器生成的演示与预训练分布对齐,可以使TAMP成为VLA微调的有效数据来源。网站和代码:https://prpl-group.com/datafarm/
cs.RO / 19 / 2609.12347

DWMP: Leveraging Dual World Models for Humanoid Obstacle Traversal

DWMP:利用双世界模型实现人形机器人障碍物穿越
Jin, Rongjun, Ma, Jianming, Gao, Yue
Abstract
Humanoid robots must traverse cluttered obstacle fields using onboard proprioceptive and visual observations, yet existing methods usually process multimodal observations without explicitly considering their different characteristics: proprioceptive observations are low-dimensional but governed by highly nonlinear robot dynamics, while egocentric visual observations are high-dimensional, noisy, and redundant. We propose DWMP (Dual World Model Policy), a framework that provides the actor with separate but complementary world-model representations for humanoid obstacle traversal. A Koopman-based dynamics world model lifts proprioceptive observations into a latent space where their temporal evolution is approximately linear, making the dynamics features easier for the actor to learn from. An RSSM-based visual world model compresses egocentric depth observations into compact stochastic states while preserving obstacle-related geometry. The student policy receives the fused latent representation for action generation, combining linearized proprioceptive dynamics with compressed visual perception. Experiments in simulation and on a Unitree G1 humanoid robot show that DWMP improves obstacle traversal performance over baselines and supports real-world deployment under randomized obstacle layouts.
Chinese Translation
人形机器人必须利用机载本体感觉和视觉观测来穿越杂乱障碍场,然而现有方法通常处理多模态观测时没有明确考虑它们的不同特性:本体感觉观测是低维的,但受高度非线性的机器人动力学支配,而以自我为中心的视觉观测是高维、嘈杂和冗余的。我们提出DWMP(双世界模型策略),一个为actor提供分离但互补的世界模型表示以用于人形机器人障碍穿越的框架。一个基于Koopman的动力学世界模型将本体感觉观测提升到一个潜在空间,其中它们的时间演化近似线性,从而使动力学特征更容易被actor学习。一个基于RSSM的视觉世界模型将以自我为中心的深度观测压缩成紧凑的随机状态,同时保留与障碍物相关的几何信息。学生策略接收融合的潜在表示以生成动作,结合了线性化的本体感觉动力学与压缩的视觉感知。在仿真和Unitree G1人形机器人上的实验表明,DWMP在障碍物穿越性能上优于基线,并支持在随机障碍布局下的真实世界部署。
cs.RO / 20 / 2609.12349

A Deployable Architecture for Robot-Mediated Tasks (DART): Evaluation in Socially Assistive Robot-Guided Cognitive Behavioral Therapy Exercises

可部署的机器人中介任务架构(DART):在社交辅助机器人引导的认知行为疗法练习中的评估
Kian, Mina, Ignatova, Lydia, Wang, Jiong, Lee, Ji Min, Li, Jiancheng, Guo, Qianwei, Weiss, Emily, O'Connell, Amy, Zareno, Kaitlin, Li, Jiani, Patel, Reyna, George, Leyaa, Huang, Minyu, Yang, Justin, Matarić, Maja J.
Abstract
Socially assistive robots (SARs) can support structured health and well-being interventions, but hardware and cost constraints limit interaction complexity and longitudinal real-world deployments. We present DART: Deployable Architecture for Robot-Mediated Tasks, an architecture that extends SARs through a web application and cloud infrastructure, enabling visual content, user input, remote computation, and persistent data storage synergistically with the robot's physical embodiment, speech, and movement. We evaluated DART by instantiating it in an interatively-developed full-stack HRI system for helping university students with elevated generalized anxiety to complete cognitive behavioral therapy (CBT) homework exercises. The resulting system, which used the low-cost open-source Blossom robot platform, was refined and evaluated through a participatory design process and multiple user studies, and finally evaluated in an in-lab study with 103 participants, and then a six-week in-home deployment with four participants. In the in-lab evaluation, participants showed significant within-session reductions in stress, state anxiety, and negative affect, and gave the platform a mean System Usability Scale score of 78.89. In the home deployment, the mean System Usability Scale score was 87.5, with positive qualitative feedback on usability. Participants across both groups identified speech input, visual presentation, and web-robot synchronization as priorities for improvement. These findings validate DART as an effective architecture for extending the capabilities of a low-cost SAR in both in-lab single-session and in real-world longitudinal deployments.
Chinese Translation
社交辅助机器人(SARs)可以支持结构化的健康与福祉干预,但硬件和成本限制限制了交互复杂性和纵向真实世界部署。我们提出了DART:可部署的机器人中介任务架构,这是一种通过Web应用程序和云基础设施扩展SAR的架构,能够与机器人的物理实体、语音和运动协同实现视觉内容、用户输入、远程计算和持久数据存储。我们通过在一个迭代开发的全栈HRI系统中实例化DART来评估它,该系统用于帮助有广泛性焦虑的大学生完成认知行为疗法(CBT)家庭作业练习。由此产生的系统使用了低成本开源Blossom机器人平台,通过参与式设计过程和多项用户研究进行了改进和评估,最后在一项有103名参与者的实验室研究中进行了评估,然后进行了为期六周的四名参与者的家庭部署。在实验室评估中,参与者在会话内压力、状态焦虑和负面情绪方面表现出显著减少,并给予该平台平均78.89的系统可用性量表分数。在家庭部署中,平均系统可用性量表分数为87.5,对可用性的定性反馈积极。两组参与者都指出语音输入、视觉呈现和Web-机器人同步是改进的优先事项。这些发现验证了DART作为一种有效架构,可以在实验室单次会话和真实世界纵向部署中扩展低成本SAR的能力。
cs.RO / 21 / 2609.12371

READ: Learning Risk-Informed Fields for End-to-End Autonomous Driving

READ:为端到端自动驾驶学习风险信息场
Liu, Zhiyuan, Tian, Yuanxin, Ke, Zehong, Li, Jinhao, Cheng, Hao, Xu, Zhenhua, Yu, Wenhao, Wang, Jianqiang
Abstract
Autonomous driving requires more than recognizing what is present in a scene: a planner must determine how road structure, surrounding agents, and their motion states should influence a future maneuver. Existing learning-based planners can capture these influences through latent scene features and trajectory decoders, but the relationship between environmental factors and candidate actions often remains implicit. This limits the ability to inspect, diagnose, or refine how scene context affects the safety of a predicted trajectory. Classical safety fields provide an explicit spatial representation of this relationship, but their risk shapes and relative weights are prescribed in advance and do not adapt to each scene. We introduce READ, a framework that learns an explicit, planning-aligned risk representation from complementary geometric and behavioral constraints. READ instantiates this representation as a continuous spatiotemporal field, enabling differentiable queries along candidate trajectories. The learned field connects scene understanding with action selection by encouraging predicted trajectories to align with low-risk regions, while retaining a differentiable interface for trajectory evaluation and refinement. READ integrates with both end-to-end planners and Vision-Language-Action models. Experiments on NAVSIM show consistent gains across matched end-to-end backbones and strong performance in a VLA setting; READ also achieves competitive results on NAVSIM v2. These results establish learned spatial risk as an explicit, adaptable representation for safe planning.
Chinese Translation
自动驾驶不仅需要识别场景中存在什么:规划器还必须确定道路结构、周围智能体及其运动状态应如何影响未来的机动。现有的基于学习的规划器能够通过潜在场景特征和轨迹解码器捕捉这些影响,但环境因素与候选动作之间的关系往往仍然是隐式的。这限制了检查、诊断或改进场景上下文如何影响预测轨迹安全性的能力。经典安全场提供了这种关系的显式空间表示,但其风险形状和相对权重是预先规定的,无法适应每个场景。我们引入了READ,一个从互补的几何和行为约束中学习显式的、与规划对齐的风险表示的框架。READ将该表示实例化为一个连续的时空场,从而能够沿候选轨迹进行可微查询。学习到的场通过鼓励预测轨迹与低风险区域对齐,将场景理解与动作选择联系起来,同时保留用于轨迹评估和细化的可微接口。READ既能与端到端规划器集成,也能与视觉-语言-动作(VLA)模型集成。在NAVSIM上的实验表明,在匹配的端到端骨干网络上获得了一致的增益,并在VLA设置中表现出强劲性能;READ在NAVSIM v2上也取得了有竞争力的结果。这些结果确立了学习到的空间风险作为安全规划的一种显式、可适应的表示。
cs.RO / 22 / 2609.12384

Understanding Whole-Body Robot Teleoperation Strategies Under Diverse Task Objectives and Constraints

理解多样任务目标与约束下的全身机器人遥操作策略
Lin, Tsung-Chi, Chen, Juo-Tung, Huang, Chien-Ming
Abstract
This work investigates the control strategies of complex whole-body robot teleoperation that coordinate active perception, bimanual manipulation, and navigation. We developed a hybrid control framework, combining the free-form and constrained control, for the whole-body teleoperation of the TIAGo mobile manipulator. We conducted a user study to explore people's control strategies under different task constraints such as limited time and low tolerance of errors. Our results highlight the effective use of coordinated control in improving task efficiency and reducing the risk of reaching individual joint limits. We discuss our results and their implications for designing future whole-body robot teleoperation systems.
Chinese Translation
本文研究了复杂全身机器人遥操作的控制策略,该策略协调主动感知、双臂操作和导航。我们开发了一个混合控制框架,结合了自由形式和约束控制,用于TIAGo移动操作机器人的全身遥操作。我们进行了一项用户研究,以探索在不同任务约束(如有限时间和低错误容忍度)下人们的控制策略。我们的结果强调了协调控制在提高任务效率和降低达到单个关节极限风险方面的有效使用。我们讨论了我们的结果及其对设计未来全身机器人遥操作系统的意义。
cs.RO / 23 / 2609.12433

FoldNet++: a Large-Scale Synthetic Dataset for Robotic T-Shirt Folding and Unfolding

FoldNet++:一个用于机器人T恤折叠与展开的大规模合成数据集
Chen, Yuxing, Wei, Zhiyuan, Xiao, Bowen, Zhang, Zhizheng, Wang, He
Abstract
Due to the highly deformable nature of garments, training a generalizable policy for robotic T-shirt folding and unfolding remains a significant challenge. In this work, we present a large-scale synthetic dataset for robotic T-shirt folding and unfolding, covering 6 robotic embodiments, 1K T-shirts, 1K environmental assets, and 120K episodes with rich annotations, which can be used to train a wide range of manipulation policies. We first follow the FoldNet pipeline to generate a large-scale dataset of physically simulatable T-shirts with diverse appearances and annotated semantic keypoints. Based on these semantic keypoints, we then generate manipulation demonstrations for different robotic embodiments through a unified rule-based framework. We use these demonstrations to train visuomotor policies, and experimental results demonstrate that models trained solely on our synthetic data can achieve over 90\% end-to-end task success rates when directly deployed to unseen real-world environments and previously unseen T-shirts from arbitrary initial configurations. Project URL: https://pku-epic.github.io/FoldNetXX/.
Chinese Translation
由于衣物具有高度可变形性,训练用于机器人T恤折叠与展开的可泛化策略仍然是一项重大挑战。在本工作中,我们提出了一个用于机器人T恤折叠与展开的大规模合成数据集,涵盖6种机器人本体、1K件T恤、1K个环境资产以及120K个带有丰富标注的回合,可用于训练各种操作策略。我们首先遵循FoldNet流程,生成一个大规模数据集,包含外观多样且带有标注语义关键点的可物理模拟T恤。基于这些语义关键点,我们随后通过一个统一的基于规则的框架,为不同的机器人本体生成操作演示。我们使用这些演示来训练视觉运动策略,实验结果表明,仅在我们的合成数据上训练的模型,在直接部署到未见过的真实环境以及来自任意初始配置的先前未见过的T恤时,能够达到超过90%的端到端任务成功率。项目网址:https://pku-epic.github.io/FoldNetXX/。
cs.RO / 24 / 2609.12441

IMPLY: Physically Anchored Consistency for World-Model Rollouts

IMPLY:世界模型推演的物理锚定一致性
Mehta, Aman, Baviskar, Riya
Abstract
A world model asked what happens if an object is pushed at several speeds produces several futures. If the model has the object in mind, those futures agree about it: each implies the same mass and friction. The consistency checks now used to vet world-action models ask whether a model's futures agree with each other, and none of them knows any physics. We show that this is not enough, and what to do instead. IMPLY reads the physics each rollout implies by inverting a simulator and scores a set of rollouts by how well one object explains all of them, anchored to two calibration pushes the model has observed. In a controlled setting, self-consistency gives a perfect score to a model that ignores the object and always predicts a typical push; anchoring exposes it (AUROC 0.70 versus 1.00). On a real model, V-JEPA 2-AC adapted to the scene, the same thing happens. Given its own calibration pushes the model tracks the object (per-object correlation with the truth 0.91); given another object's, it does not (0.05). Self-consistency cannot tell these apart, preferring the right evidence on 52% of objects, chance level, while anchored disagreement prefers it on 73% and correlates 0.92-0.99 with the rollouts' error. Used to choose among candidate rollout sets, it comes within 0.003 of an oracle that sees the truth. A model that has internalised the wrong object is exactly as self-consistent as one that has internalised the right one; consistency has to be anchored to evidence.
Chinese Translation
当被问及以不同速度推动物体会发生什么时,世界模型会产生多个未来。如果模型正确地把握了该物体,这些未来关于它是一致的:每个都隐含相同的质量和摩擦。目前用于评估世界-动作模型的一致性检查,关注的是模型的各个未来是否彼此一致,而这些检查都不涉及任何物理知识。我们表明这还不够,并给出了替代方案。IMPLY 通过反演模拟器来读取每个推演所隐含的物理,并通过一个物体能多好地解释所有推演来对一组推演评分,其锚定于模型观察到的两次校准推动。在受控设置中,自一致性给一个忽略物体且总是预测典型推动的模型打出完美分数;锚定则揭露了它(AUROC 0.70 对 1.00)。在真实模型 V-JEPA 2-AC 上,适配到该场景后,同样的事情发生了。给定其自身的校准推动,模型能跟踪该物体(每个物体与真值的相关性为 0.91);给定另一个物体的校准推动,则不能(0.05)。自一致性无法区分这些,在 52% 的物体上偏好正确的证据,即随机水平;而锚定不一致性在 73% 的物体上偏好正确证据,并与推演的误差相关 0.92-0.99。用于在候选推演集之间进行选择时,其与能看到真相的 oracle 相差在 0.003 以内。一个内化了错误物体的模型,其自一致性与内化了正确物体的模型完全相同;一致性必须锚定于证据。
cs.RO / 25 / 2609.12456

PATH: Continuous Target Sensing among Autonomous Cooperative Drones

PATH:自主协同无人机间的连续目标感知
Kim, Heegyeong, James, Alice, Seth, Avishkar, Kuantama, Endrowednes, Williamson, Jane, Feng, Yimeng, Han, Richard
Abstract
Continuous target sensing by uncrewed aerial vehicles (UAVs) is constrained by limited flight endurance, motivating the transfer of tracking responsibility between cooperating UAVs. Such a handoff requires the receiver to identify the same physical target currently tracked by the sender despite differences in viewpoint, scale, and target appearance. Existing approaches based on global target localization or appearance-based cross-view association are limited by positioning uncertainty or ambiguous visual features. This paper presents Perspective Alignment \& Tracking Handoff (\textbf{PATH}), a platform-agnostic, geometry-assisted sensing and verification framework for target handoff between two moving UAVs. The sender reconstructs the tracked target as a metric 3D point using RGB-D sensing, while the receiver estimates its relative pose from a fiducial observation and projects the transmitted target point into its own image as a spatial prior for target acquisition. The receiver-generated candidate is then returned to the sender and verified through a cross-view Mutual Agreement Handshake before tracking responsibility is transferred. Real-world UAV experiments show mean relative-position and target-position errors of 0.047~m and 0.030~m, respectively. Under visually ambiguous conditions, PATH achieves 96.0\% frame-level receiver-side target acquisition accuracy, with 2.0\% false-positive and 2.0\% false-negative rates. A sensor-error sensitivity analysis shows that relative-pose uncertainty is the dominant contributor to receiver-view projection error. The implementation operates at video rate with compact inter-UAV communication below 16~kB/s at 60~Hz, demonstrating the feasibility of lightweight geometry-assisted target handoff on resource-constrained UAV platforms.
Chinese Translation
无人机(UAV)的连续目标感知受限于有限的飞行续航能力,这促使在协同无人机之间转移跟踪任务。这种交接要求接收方识别发送方当前跟踪的同一物理目标,尽管存在视角、尺度和目标外观的差异。基于全局目标定位或基于外观的跨视角关联的现有方法受到定位不确定性或模糊视觉特征的限制。本文提出了透视对齐与跟踪交接(PATH),一种平台无关的、几何辅助的感知与验证框架,用于两个移动无人机之间的目标交接。发送方使用 RGB-D 感知将跟踪目标重建为度量 3D 点,而接收方从基准观测中估计其相对位姿,并将传输的目标点投影到其自身图像中,作为目标获取的空间先验。然后,接收方生成的候选目标被返回给发送方,并通过跨视角相互一致握手进行验证,之后才转移跟踪责任。真实世界无人机实验显示,平均相对位置误差和目标位置误差分别为 0.047 m 和 0.030 m。在视觉模糊条件下,PATH 实现了 96.0% 的帧级接收方目标捕获准确率,误报率和漏报率分别为 2.0% 和 2.0%。传感器误差敏感性分析表明,相对位姿不确定性是接收方视角投影误差的主要贡献因素。该实现在视频速率下运行,无人机间通信紧凑,在 60 Hz 下低于 16 kB/s,证明了在资源受限的无人机平台上进行轻量级几何辅助目标交接的可行性。
cs.RO / 26 / 2609.12498

ArtManip: Category-Level Articulated In-Hand Manipulation

ArtManip:类别级关节物体手内操作
Yang, Yang, Liu, Tengyu, Li, Puhao, Chen, Zeyuan, Li, Yuyang, Wang, Xingwan, Wu, Yingying, Cui, Zhaopeng, Huang, Siyuan
Abstract
Category-level in-hand manipulation of articulated objects is a formidable yet underexplored challenge for dexterous robotic hands. This difficulty stems from two core bottlenecks: first, controlling an object's internal degrees of freedom is tightly coupled with maintaining grasp stability on a free-floating base; second, acquiring diverse object models and functional grasps at scale is highly labor-intensive, yet vital for generalization given the system's sensitivity to initial configurations. In this work, we present ArtManip, the first category-level articulated in-hand manipulation method that generalizes across object instances and diverse initial grasps. For initial configuration construction, we develop an automated pipeline that procedurally generates diverse articulated objects and synthesizes task-oriented functional grasps. For policy learning, we propose a robust two-stage training strategy that incorporates articulation physics randomization, reward curriculum, and latent representation distillation to handle complex contact and joint dynamics during deployment. Extensive experiments across four object categories demonstrate that our policy generalizes to unseen instances and varied configurations in simulation, and achieves zero-shot transfer to 12 real-world objects featuring diverse shapes and joint mechanics.
Chinese Translation
针对关节物体的类别级手内操作对灵巧机械手而言是一项艰巨且尚未充分探索的挑战。这一困难源于两个核心瓶颈:其一,控制物体的内部自由度与在自由漂浮基座上保持抓取稳定性紧密耦合;其二,大规模获取多样化的物体模型和功能性抓取高度依赖人力,但考虑到系统对初始配置的敏感性,这对泛化能力至关重要。在这项工作中,我们提出了 ArtManip,这是首个能够泛化到不同物体实例和多种初始抓取的类别级关节物体手内操作方法。对于初始配置构建,我们开发了一个自动化流程,能够程序化生成多样化的关节物体并合成面向任务的功能性抓取。对于策略学习,我们提出了一种鲁棒的两阶段训练策略,融合了关节物理随机化、奖励课程学习和潜在表示蒸馏,以处理部署过程中复杂的接触和关节动力学。在四个物体类别上的大量实验表明,我们的策略能够泛化到仿真中未见过的实例和多样化的配置,并实现了对 12 个具有不同形状和关节力学的真实世界物体的零样本迁移。
cs.RO / 27 / 2609.12502

Communication-Constrained Multi-Robot Exploration With Adaptive Communication Windows

基于自适应通信窗口的通信受限多机器人探索
Rossano, Ben, Lim, Jaein, How, Jonathan P.
Abstract
Exploring unknown environments with multi-robot teams can improve efficiency by allowing robots to explore in parallel. However, realizing these gains requires effective information sharing. When communication is intermittent, robots must balance the benefits of sharing information against the cost of diverting from exploration to establish communication. This paper introduces MACE, a decentralized exploration framework that actively evaluates whether establishing communication is worthwhile. At scheduled communication windows, robots estimate the cost of reaching previously identified communication locations. By formulating this decision as a variant of the Vehicle Orienteering Problem, robots evaluate routes based on the travel required to establish communication and the exploration that can be completed along the way. This approach enables robots to communicate more frequently than under purely opportunistic strategies while reducing the unnecessary travel associated with fixed rendezvous strategies. Across a set of simulated environments with varying size and geometry, we demonstrate that MACE reduces the total exploration time by up to 23% compared to existing communication-constrained exploration strategies.
Chinese Translation
使用多机器人团队探索未知环境可以通过允许机器人并行探索来提高效率。然而,实现这些收益需要有效的信息共享。当通信是间歇性的时,机器人必须在共享信息的好处与从探索转向建立通信的成本之间进行权衡。本文介绍了MACE,一种去中心化探索框架,它主动评估建立通信是否值得。在预定的通信窗口,机器人估计到达先前识别的通信位置的成本。通过将该决策表述为车辆定向问题(Vehicle Orienteering Problem)的一个变体,机器人根据建立通信所需的行程以及沿途可完成的探索来评估路线。这种方法使机器人能够比纯机会主义策略更频繁地通信,同时减少与固定会合策略相关的不必要行程。在一组具有不同大小和几何形状的模拟环境中,我们证明MACE与现有的通信受限探索策略相比,将总探索时间减少了高达23%。
cs.RO / 28 / 2609.12530

Autonomous Precision Milling of Biological Structures via Generic Anatomical Priors and Active Boundary Perception

基于通用解剖先验与主动边界感知的生物结构自主精密铣削
Zhao, Enduo, Lin, Xiaofeng, Wang, Yifan, Song, Yuhan, Li, Weihan, Perez, Saul Alexis Heredia, Harada, Kanako
Abstract
Autonomous precision milling of biological structures is challenged by incomplete knowledge of target geometry, local material thickness, and critical internal boundaries. Subject-specific preoperative models can address geometric and thickness variations, but static models cannot determine boundary status encountered during execution, while repeated target-specific imaging limits scalability. This article presents an uncertainty-aware autonomous milling framework that assigns complementary roles to generic anatomical priors and active boundary perception. A generic anatomical prior provides conservative global guidance and is transformed through semantic-guided registration and hybrid vision-force calibration into robot-executable guidance for individual targets. As milling approaches uncertain boundaries, the robot actively probes the remaining structure and uses relative stiffness changes to estimate boundary status and structural detachability. A state-adaptive controller governs transitions between active perception and spatially selective incremental refinement, repeating this cycle until the termination criterion is satisfied. Hierarchical experiments on biological surrogates and in vivo mouse cranial window creation demonstrate accurate anatomical prior transfer, reliable boundary adaptation, and autonomous precision milling of biological structures.
Chinese Translation
生物结构的自主精密铣削受到目标几何形状、局部材料厚度和关键内部边界知识不完整的挑战。患者特异性术前模型可以处理几何形状和厚度变化,但静态模型无法确定执行过程中遇到的边界状态,而重复的目标特异性成像限制了可扩展性。本文提出了一种不确定性感知的自主铣削框架,该框架为通用解剖先验和主动边界感知分配了互补的角色。通用解剖先验提供保守的全局引导,并通过语义引导配准和视觉-力混合校准转化为针对个体目标的机器人可执行引导。当铣削接近不确定边界时,机器人主动探测剩余结构,并利用相对刚度变化来估计边界状态和结构可分离性。状态自适应控制器管理主动感知和空间选择性增量细化之间的转换,重复此循环直到满足终止准则。在生物替代物和活体小鼠颅窗建立上的分层实验证明了准确的解剖先验迁移、可靠的边界适应以及生物结构的自主精密铣削。
cs.RO / 29 / 2609.12549

STAR: Sparse Tactile Representation Learning in Vision-Tactile-Language-Action Models for Dexterous Manipulation

STAR:面向灵巧操作的视觉-触觉-语言-动作模型中的稀疏触觉表征学习
Liu, Xiangcheng, Wu, Tianhao, Zheng, Le, Wang, Yidong, Jiang, Bowen, Pan, Mingjie, Ren, Xinlin, Liu, Yi, Luo, Jianlan
Abstract
Dexterous manipulation requires coordinated multi-finger control and effective tactile feedback, yet learning these capabilities remains challenging due to the lack of large-scale real-world data and the difficulty of extracting effective representations from sparse tactile signals. We build a robot platform and teleoperation system to collect a 200-hour bimanual dexterous manipulation dataset with synchronized visual, tactile, and language annotations, comprising 10,576 trajectories across 65 tasks, 69.5% of which involve dexterous multi-finger manipulation. We further propose STAR, an integrated training recipe for vision-tactile-language-action (VTLA) models that addresses the spatial, temporal, and informational sparsity of tactile signals through visual-tactile joint pre-training, sparse-global tactile token representation, and sparse future tactile prediction. Trained on this dataset, STAR achieves a 61% average success rate across four real-world tasks with 100 post-training trajectories per task, demonstrating dexterous performance under task-specific post-training.
Chinese Translation
灵巧操作需要协调的多指控制和有效的触觉反馈,然而由于缺乏大规模真实世界数据以及难以从稀疏触觉信号中提取有效表征,学习这些能力仍然具有挑战性。我们构建了一个机器人平台和遥操作系统,收集了一个200小时的双臂灵巧操作数据集,其中包含同步的视觉、触觉和语言标注,涵盖65个任务中的10,576条轨迹,其中69.5%涉及灵巧多指操作。我们进一步提出了STAR,一种用于视觉-触觉-语言-动作(VTLA)模型的集成训练方案,通过视觉-触觉联合预训练、稀疏全局触觉token表示和稀疏未来触觉预测,解决了触觉信号的空间、时间和信息稀疏性问题。在该数据集上训练后,STAR在四个真实世界任务上实现了61%的平均成功率,每个任务有100条后训练轨迹,展示了在任务特定后训练下的灵巧性能。
cs.RO / 30 / 2609.12595

A Hierarchical Coverage Path Planning Algorithm for Unknown Environments

未知环境下的分层覆盖路径规划算法
Shen, Zongyuan, Liu, Haodong, Wang, Gao, Ma, Hongbin, Ou, Yaming, Zhao, Shancheng, Zhou, Dehua
Abstract
This paper presents an online coverage path planning algorithm for unknown environments. During navigation, the initially unknown search area is progressively decomposed into disconnected subareas as new obstacle information is acquired and coverage proceeds. These subareas are organized in an incrementally constructed decomposition tree that preserves their hierarchical parent-child relationships. Based on this tree, a global coverage tour is maintained and updated online by prioritizing newly generated child subareas according to their exploration states and distances from the robot. A local planner then generates coverage motions within each selected subarea, allowing the robot to adapt its trajectory as the environment is gradually revealed. Its performance is evaluated via high-fidelity simulations in complex scenarios. The results show improved coverage efficiency in terms of path length and overlap ratio in comparison to three baseline algorithms.
Chinese Translation
本文提出了一种面向未知环境的在线覆盖路径规划算法。在导航过程中,随着新障碍信息的获取和覆盖的进行,初始未知的搜索区域被逐步分解为不连通的子区域。这些子区域被组织在一棵增量构建的分解树中,该树保留了它们的分层父子关系。基于这棵树,通过根据新生成的子区域的探索状态和与机器人的距离对其进行优先级排序,在线维护和更新全局覆盖路径。然后,局部规划器在每个选定的子区域内生成覆盖运动,使机器人能够在环境逐渐揭示时调整其轨迹。其性能通过复杂场景中的高保真仿真进行评估。结果表明,与三种基线算法相比,在路径长度和重叠率方面,覆盖效率有所提高。
cs.RO / 31 / 2609.12609

Quantifying Spectral Differences in Vehicle Between Production Autonomous and Human-Driven Vehicles Across Driving Scenarios

量化不同驾驶场景下量产自动驾驶车辆与人类驾驶车辆之间的车辆频谱差异
Fang, Peiyi, Li, Xiangyu, Weng, Yonglin, Ma, Ke
Abstract
Differences in vehicle kinematic characteristics between production autonomous vehicles (PAVs) and human-driven vehicles (HVs) have been limitedly investigated by empirical studies. Most recent studies rely on simulation-based models, while some further investigate low-level adaptive cruise control (ACC) systems in controlled experiments. These methods commonly adapt some time-domain metrics to characterize PAV-HV differences across limited driving conditions. However, current PAVs equipped with high-level autonomous driving systems generate driving behaviors in a black box using data-driven models. These fundamentally different mechanisms for generating behaviors may produce distinct kinematic characteristics in traffic. More importantly, these time-domain metrics cannot reflect frequency-related traffic dynamics across different driving scenarios. Thus, this study adapted a real-world PAV dataset with four PAV platforms and developed a frequency-domain framework to quantify kinematic differences between PAVs and HVs across diverse driving scenarios, including varying driving states, lighting, weather, and vehicle densities. The framework transforms kinematic signals into the frequency domain and extracts spectral features, and then compares these features between PAVs and HVs based on kernel density estimation and Wasserstein distance. The results reveal clear scenario-dependent PAV-HV spectral differences. Specifically, speed-related differences were consistently smaller during car-following than cruising, while rainy conditions consistently enlarged acceleration-related differences compared with clear conditions. These findings highlight the necessity of multi-scenario evaluations and demonstrate the value of frequency-domain analysis for characterizing PAV-HV kinematic differences under real-world conditions.
Chinese Translation
关于量产自动驾驶车辆(PAVs)与人类驾驶车辆(HVs)之间车辆运动学特征的差异,实证研究开展得有限。最近大多数研究依赖于基于仿真的模型,而一些研究则在受控实验中进一步研究低级别的自适应巡航控制(ACC)系统。这些方法通常采用一些时域指标来刻画有限驾驶条件下PAV-HV的差异。然而,当前配备高级自动驾驶系统的PAV使用数据驱动模型以黑箱方式生成驾驶行为。这些根本不同的行为生成机制可能在交通中产生不同的运动学特征。更重要的是,这些时域指标无法反映不同驾驶场景中与频率相关的交通动态。因此,本研究采用了一个包含四个PAV平台的真实世界PAV数据集,并开发了一个频域框架,以量化PAV与HV在多种驾驶场景(包括不同的驾驶状态、光照、天气和车辆密度)下的运动学差异。该框架将运动学信号转换到频域并提取频谱特征,然后基于核密度估计和Wasserstein距离比较PAV和HV之间的这些特征。结果揭示了明显的依赖于场景的PAV-HV频谱差异。具体而言,与巡航相比,跟驰过程中与速度相关的差异始终较小;而与晴天条件相比,雨天条件始终会放大与加速度相关的差异。这些发现凸显了多场景评估的必要性,并证明了频域分析在表征真实世界条件下PAV-HV运动学差异方面的价值。
cs.RO / 32 / 2609.12614

ProClosure: Hierarchical Room-Object Assignment using Progressive Boundary Closure from Monocular Video

ProClosure:基于渐进边界闭合的单目视频分层房间-物体分配
Muthuraj, Vinoth Kumar, Banik, Soumyadeep, Sharma, Kushal, Jain, Hardik
Abstract
A 3D scene graph groups objects into rooms. When a robot is asked to fetch an object from the kitchen, that grouping is what tells it where to look. An object recorded in the wrong room is not retrievable by a query naming the correct room. We introduce Progressive Boundary Closure, which recovers room layer from a monocular RGB video. A SLAM front end and an open-vocabulary segmenter supply a structural point cloud, camera trajectory and object tracks. The cloud is rasterised into a top-down map, rooms are recovered from it, and each object takes the room holding most of its extent. The difficulty lies in the map itself. Walls are recorded only where the camera looked, so a gap in the boundary may be a doorway or a stretch of wall that was never observed; nothing distinguishes the two. Prior methods treat both as passages, merging rooms that should remain separate. We observe that both require the same treatment: a room should not extend across either, so both are closed and need not be distinguished. Such an opening closes under a small amount of boundary growth, and few sightlines cross it, so points in different rooms rarely see one another. We use the first to recover rooms and the second to assign objects to them. Rooms are obtained by Progressively thickening the boundary inward and freezing each free-space region once it becomes enclosed, so every opening seals at its own scale rather than at a radius fixed in advance. Camera poses are used as seeds, which removes the sampling heuristic and makes the segmentation deterministic. Over 10 floors of 6 HM3D-Semantics scenes, scored against HOV-SG on identical top-down maps, we recover 74 rooms for 72 annotated regions (HOV-SG: 44), raising room F_1 from 0.741 to 0.890 at IoU 0.25 at some cost in precision, and object-to-room ARI from 0.488 to 0.696 (p=0.002, ahead on every floor).
Chinese Translation
3D场景图将物体分组到房间中。当机器人被要求从厨房取一个物体时,这种分组告诉它该去哪里找。记录在错误房间中的物体无法通过命名正确房间的查询检索到。我们引入渐进边界闭合(Progressive Boundary Closure),它从单目RGB视频中恢复房间层。SLAM前端和开放词汇分割器提供结构化点云、相机轨迹和物体轨迹。点云被栅格化为俯视地图,从中恢复房间,每个物体分配到包含其大部分范围的房间。困难在于地图本身。墙壁仅在相机看到的地方被记录,因此边界上的缺口可能是门口,也可能是一段从未被观察到的墙壁;两者无法区分。先前的方法将两者都视为通道,合并了本应保持分离的房间。我们观察到两者需要相同的处理:房间不应跨越任一者,因此两者都被闭合,无需区分。这样的开口在少量边界增长下闭合,且很少有视线穿过它,因此不同房间中的点很少互相看见。我们使用第一种来恢复房间,第二种来将物体分配到房间。房间通过渐进地向内加厚边界,并在每个自由空间区域被包围时将其冻结来获得,因此每个开口在其自身的尺度上封闭,而不是在预先固定的半径上。相机姿态被用作种子,这消除了采样启发式,并使分割具有确定性。在6个HM3D-Semantics场景的10个楼层上,在相同的俯视地图上与HOV-SG进行评分对比,我们为72个标注区域恢复了74个房间(HOV-SG:44个),在IoU 0.25下将房间F_1从0.741提升到0.890,代价是精度有所下降,并将物体到房间的ARI从0.488提升到0.696(p=0.002,在每个楼层上都领先)。
cs.RO / 33 / 2609.12634

Online Material Estimation for Conditioned Diffusion Policy in Shaping Deformable Linear Objects

用于塑造可变形线性物体的条件扩散策略的在线材料估计
Yamada, Ryunosuke, Motoda, Tomohiro, Domae, Yukiyasu, Tsuji, Tokuo
Abstract
Shape control of deformable linear objects (DLOs) is challenging for imitation learning because deformation behavior varies with material properties such as stiffness and elasticity, so a single policy must generate different action sequences for different objects even when the goal shape is identical. We propose a diffusion policy conditioned on material labels that are estimated online during manipulation. A recurrent estimation network predicts the material label of the grasped object from the time series of multi-view images and robot joint states, and the predicted label conditions the diffusion policy at every inference step. We collected 480 real-robot demonstrations covering four DLO materials and three groove-placement tasks, and compared per-material specialist policies, a task-conditioned policy without material labels, a policy conditioned on ground-truth material labels, and the proposed policy. Conditioning on ground-truth material labels improved the average success rate from 45.8% to 60.0% over the task-only policy, and the proposed policy reached 60.8% without any prior material information, matching the policy given ground-truth labels. A post-hoc analysis shows that the estimator extracts material-related information from the manipulation observations and that the diffusion policy responds to the resulting conditioning signal, while the one pronounced failure case is associated with persistent confusion between two similar materials.
Chinese Translation
可变形线性物体(DLOs)的形状控制对模仿学习具有挑战性,因为变形行为随材料属性(如刚度和弹性)而变化,因此即使目标形状相同,单一策略也必须为不同物体生成不同的动作序列。我们提出了一种以在操作过程中在线估计的材料标签为条件的扩散策略。一个循环估计网络从多视角图像和机器人关节状态的时间序列中预测被抓取物体的材料标签,预测的标签在每个推理步骤对扩散策略进行条件化。我们收集了480个真实机器人演示,涵盖四种DLO材料和三种凹槽放置任务,并比较了每种材料的专家策略、没有材料标签的任务条件策略、以真实材料标签为条件的策略以及所提出的策略。以真实材料标签为条件将平均成功率从仅任务策略的45.8%提高到60.0%,而所提出的策略在没有任何先验材料信息的情况下达到60.8%,与给定真实标签的策略相当。事后分析表明,估计器从操作观察中提取与材料相关的信息,并且扩散策略对所产生的条件信号做出响应,而一个明显的失败案例与两种相似材料之间的持续混淆有关。
cs.RO / 34 / 2609.12641

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

打破视觉-动作捷径:面向可泛化机器人基础模型的潜在接口训练
Lin, Jianman, Shailesh, Shailesh, Luo, Zhongyi, Duan, Jiafei
Abstract
Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.
Chinese Translation
机器人基础模型在分布内性能表现强劲,但在视觉分布偏移下往往性能下降。当学习从预训练视觉表征生成动作时,模型可能利用与训练分布中示范动作相关的任务无关视觉线索。当这些相关性在分布偏移下发生变化时,此类视觉-动作捷径会损害泛化能力。缓解这些捷径需要约束视觉信息用于动作生成的方式,同时保留任务相关的空间信息。我们提出潜在接口训练(Latent Interface Training, LIT),一种框架无关的两阶段策略:首先在不使用图像的情况下建立以空间目标为条件的动作先验,然后通过姿态监督的潜在接口约束视觉条件化。第一阶段训练动作专家,以语言、机器人状态以及每个示范动作块的末端 SE(3) 末端执行器位姿为条件生成动作块,从而独立于视觉线索学习目标导向的动作生成。第二阶段引入一个潜在接口,该接口聚合视觉与语义表征,并作为预训练动作专家唯一的视觉条件化通路。对该接口进行监督以重建先前用于第一阶段条件化的终端位姿,从而促使其保留动作生成所需的目标相关空间信息。在四种视觉-语言-动作与世界-动作架构(Pi0.5、MolmoAct2、FAST-WAM 和 ImageWAM)上,LIT 将 LIBERO-Plus 的总体成功率提升 3.87-10.70 个百分点,同时保持或提升平均 LIBERO 成功率。真实世界评估表明,在未见过的相机配置、光照变化和干扰物下,跨三个任务汇总的成功率提升 13.30-16.70 个百分点。
cs.RO / 35 / 2609.12660

Driving Context-guided Model Predictive Planning and Control for Autonomous Car Racing at the Limit and Beyond

面向极限及超越极限自动驾驶赛车的驾驶情境引导模型预测规划与控制
Raji, Ayoub, Sacco, Federico, Musiu, Nicola, Bertogna, Marko
Abstract
This paper presents a Model Predictive Control-based motion planning and control pipeline for autonomous car racing capable of adapting to different driving contexts, such as overtaking, nominal driving, and countersteering. A Cost Blending state machine manages the identification of different driving contexts and the selection of their predefined weights to be applied to the Model Predictive Planning (MPP) and Control (MPC) modules. The two optimization-based solutions share the same problem formulation and model, differing only in horizon length, rate, tuning, and in their open-loop versus closed-loop approach to maximize the effectiveness of their interaction. The work is validated on the fully autonomous open-wheel racecar Superformula EAV-25, with a lap time achieved that is within 2% of the best human driver reference. The results demonstrate the capability of the solution in driving at the limit of handling, smoothly executing overtaking maneuvers, and quickly reacting to high oversteering conditions to recover the vehicle stability.
Chinese Translation
本文提出了一种基于模型预测控制的运动规划与控制流程,用于自动驾驶赛车,能够适应不同的驾驶情境,例如超车、常规驾驶和反打方向。一个代价混合状态机负责管理不同驾驶情境的识别以及其预定义权重的选择,以应用于模型预测规划(MPP)和控制(MPC)模块。这两个基于优化的解决方案共享相同的问题表述和模型,仅在预测时域长度、频率、调参以及开环与闭环方法上有所不同,以最大化其交互的有效性。该工作在完全自动驾驶的开轮式赛车Superformula EAV-25上进行了验证,实现的圈速与最佳人类驾驶员参考圈速的差距在2%以内。结果表明,该解决方案能够在操控极限下驾驶,平顺地执行超车动作,并对严重转向过度情况做出快速反应以恢复车辆稳定性。
cs.RO / 36 / 2609.12677

Size Doesn't Matter: Material-State Reinforcement Learning for Excavator Transferable Soil Manipulation

大小无关:面向挖掘机可迁移土壤操作的材料状态强化学习
Werner, Lennart, Eyschen, Pol, Costello, Sean, Micarelli, Pierluigi, Cramariuc, Andrei, Hutter, Marco
Abstract
Earthmoving tasks such as excavation, backfilling, or embankment construction require deliberate repositioning of deformable soil. For these tasks, human operators use all shovel faces, while autonomous systems so far are limited to excavation and dumping. Current methods often rely on heuristic models but do not incorporate soil mechanics. We address this shortcoming by using Reinforcement Learning in a GPU-parallelized Material Point Method particle simulation. Our controllers are conditioned on material state such as shape and compactness, enabling skills that use multiple contact faces of the tool and displace material both inside and outside of the shovel. To use the same learned weights across machines, our policies operate in a normalized end-effector space and are deployed through a calibrated machine interface. We evaluate this calibrated transfer on an 11.5t hydraulic excavator and a 500g tabletop robot. We validate performance through autonomous construction of a 42m long, 2.1m high embankment in 45min, executing 201 individual policy strokes without failure, retry, or operator intervention. In a direct comparison, the autonomous controller matches an expert operator's progression speed and produces a higher, more consistent embankment. Additional qualitative backfilling and compaction experiments demonstrate the material-state awareness and calibrated transfer across machines.
Chinese Translation
土方工程任务,如挖掘、回填或筑堤,需要对可变形土壤进行有意的重新定位。对于这些任务,人类操作员使用铲斗的所有面,而迄今为止的自主系统仅限于挖掘和倾倒。当前方法通常依赖于启发式模型,但并未纳入土壤力学。我们通过在GPU并行化的物质点法(Material Point Method)粒子模拟中使用强化学习来解决这一不足。我们的控制器以材料状态(如形状和密实度)为条件,使得技能能够利用工具的多个接触面,并在铲斗内外移动材料。为了在不同机器上使用相同的学习权重,我们的策略在归一化的末端执行器空间中运行,并通过校准的机器接口进行部署。我们在11.5吨液压挖掘机和500克桌面机器人上评估了这种校准迁移。我们通过自主建造一条42米长、2.1米高的堤坝,在45分钟内执行了201次单独的策略行程,无失败、重试或操作员干预,从而验证了性能。在直接比较中,自主控制器匹配了专家操作员的推进速度,并产生了更高、更一致的堤坝。额外的定性回填和压实实验展示了材料状态感知以及跨机器的校准迁移。
cs.RO / 37 / 2609.12721

Improving Imitation Learning Efficiency for Manipulation through Geometric Prior Pretraining

通过几何先验预训练提升操作任务的模仿学习效率
Iwakata, Shogo, Motoda, Tomohiro, Yamada, Ryosuke, Makihara, Koshi, Nakajo, Ryoichi, Tanaka, Keitaro, Murooka, Masaki, Mykhailyshyn, Roman, Kataoka, Hirokatsu, Morishima, Shigeo, Domae, Yukiyasu
Abstract
Applying an imitation learning policy to a new manipulation task usually requires collecting new demonstrations and retraining the model, which makes sample efficiency a practical concern. Pretraining on large-scale robot datasets is effective in this respect, but such datasets are costly to collect and train on, while data augmentation techniques typically require a new round of data generation and retraining for each task. A complementary question is what useful prior can be provided to a policy at negligible cost before any task-specific data are collected. In this study, we construct a geometric visual pretraining dataset in which each scene contains only a plane, an object, and a hand, and trajectories are generated automatically. The scenes contain neither textures nor backgrounds; pretraining primarily exposes the policy to the geometric relationship between the hand and the object. Furthermore, representing the hand as a cube avoids tailoring the dataset to a specific robot morphology. We evaluate this geometric prior using ACT on three simulated robots across five manipulation tasks each, as well as on three real-world robot tasks. Across many of these robot--task combinations, fine-tuning from the geometric prior achieves higher success rates in the early stages of training than training from scratch while using only a small number of task demonstrations. These results suggest that even highly simplified geometric scenes can provide a useful initialization that transfers across robots and to real-world tasks when task data are limited.
Chinese Translation
将模仿学习策略应用于新的操作任务通常需要收集新的演示并重新训练模型,这使得样本效率成为一个实际关注的问题。在这方面,基于大规模机器人数据集的预训练是有效的,但此类数据集的收集和训练成本高昂,而数据增强技术通常需要为每个任务重新进行一轮数据生成和重训练。一个补充性问题是:在收集任何任务特定数据之前,能够以可忽略的成本为策略提供什么有用的先验。在本研究中,我们构建了一个几何视觉预训练数据集,其中每个场景仅包含一个平面、一个物体和一只手,并且轨迹是自动生成的。这些场景既没有纹理也没有背景;预训练主要让策略接触手与物体之间的几何关系。此外,将手表示为一个立方体,避免了使数据集针对特定机器人形态进行定制。我们使用 ACT 在三个仿真机器人上评估该几何先验,每个机器人涵盖五项操作任务,并在三个真实世界机器人任务上也进行了评估。在许多这样的机器人-任务组合中,与从头训练相比,从几何先验进行微调在仅使用少量任务演示的情况下,在训练早期阶段取得了更高的成功率。这些结果表明,即使高度简化的几何场景也能提供有用的初始化,当任务数据有限时,这种初始化可以跨机器人迁移并迁移到真实世界任务。
cs.RO / 38 / 2609.12737

Control Architecture for Safe Grasping of Fragile Objects Using a Coarse Position-Controlled Gripper

基于粗位置控制夹爪的易碎物体安全抓取控制架构
Pavlic, Marko, Geier, Moritz, Markert, Timo, Burschka, Darius
Abstract
Robots are increasingly used in unstructured environments. The need for them to safely grasp unknown objects without damaging them becomes crucial. Humans achieve this by sensing and quickly responding by adjusting their grasping force. Similarly, effective grasp acquisition in robots requires compliant interaction strategies that can adapt to uncertain object properties and adjust to any instabilities during manipulation. We present a geometry-aware force/torque-based contact estimation method for a coarse position-controlled gripper, combined with an adaptive admittance controller for safe grasp acquisition. The desired contact forces are estimated online to keep stable contact with objects of unknown properties. This enables compliant and stable grasps while avoiding excessive forces. Experiments with objects of different sizes, shapes, stiffnesses, and weights show that the proposed algorithm not only prevents slippage but also applies minimal force to safely grasp an object without causing excessive deformation.
Chinese Translation
机器人越来越多地用于非结构化环境。它们安全抓取未知物体而不造成损坏的需求变得至关重要。人类通过感知并迅速响应调整其抓取力来实现这一点。类似地,机器人有效抓取需要柔顺的交互策略,能够适应不确定的物体属性并调整操作过程中的任何不稳定。我们提出了一种针对粗位置控制夹爪的几何感知的基于力/力矩的接触估计方法,结合自适应导纳控制器以实现安全抓取。期望接触力在线估计,以保持与未知属性物体的稳定接触。这实现了柔顺且稳定的抓取,同时避免过大的力。对不同尺寸、形状、刚度和重量的物体进行的实验表明,所提算法不仅能防止滑动,还能以最小的力安全抓取物体,而不会造成过度变形。
cs.RO / 39 / 2609.12795

High-Fidelity Multi-Body Simulator for Autonomous Racing

用于自动驾驶赛车的高保真多体仿真器
Musiu, Nicola, Iacovacci, Francesco, Lupo, Fausto, Pini, Matteo, Scapicchi, Giovanni, Moretti, Francesco, Mascaro, Eugenio, Musso, Pietro, Raji, Ayoub, Bertogna, Marko, Arricale, Vincenzo Maria, Sapio, Angelo Lo, Piccarelli, Alessandro, Fish, Garron
Abstract
We present a custom high-fidelity vehicle dynamics simulation environment for testing and validation of Autonomous Racing software. The digital twin of the autonomous vehicle is developed in Dymola, using racecar dynamics modeling libraries to build a complete multi-body model. A 3D road surface, including elevation profiles and curbs, is implemented using the Curved Regular Grid (CRG) standard. The model is exported from Dymola as a Functional Mock-up Unit (FMU) and integrated into a custom software-in-the-loop simulator, where communication interfaces with the autonomous racing stack were developed in C++. A calibration procedure based on experimental data is also presented, along with a validation study to further support the quality of the proposed framework. The simulator runs in real time on a portable computer and provides reliable ground truth for algorithms validation prior to real-world deployment.
Chinese Translation
我们提出了一种定制的高保真车辆动力学仿真环境,用于自动驾驶赛车软件的测试与验证。自动驾驶车辆的数字孪生是在Dymola中开发的,使用赛车动力学建模库构建完整的多体模型。使用弯曲规则网格(Curved Regular Grid, CRG)标准实现了包含高程剖面和路缘的3D路面。该模型从Dymola导出为功能样机单元(Functional Mock-up Unit, FMU),并集成到定制的软件在环仿真器中,其中与自动驾驶赛车软件栈的通信接口是用C++开发的。还提出了基于实验数据的校准程序,以及一项验证研究,以进一步支持所提出框架的质量。该仿真器在便携式计算机上实时运行,并为实际部署前的算法验证提供可靠的真值。
cs.RO / 40 / 2609.12831

VertexCBF: Improving Neural Control Barrier Functions via Vertex-Restricted Control Search

VertexCBF:通过顶点受限控制搜索改进神经控制屏障函数
Derajić, Bojan, Bernhard, Sebastian, Hönig, Wolfgang
Abstract
As the number of autonomous robots continues to grow, safety becomes increasingly important. Control barrier functions (CBFs) provide a theoretically grounded framework for ensuring safety, but existing design methods often face limitations in effectiveness, scalability, or interpretability, and may result in overly conservative safe sets. In this paper, we propose \emph{VertexCBF}, a framework for learning neural CBFs in a scalable, systematic, and explainable way. We approximate the stationary Hamilton--Jacobi value function using a neural network trained via a combination of physics-informed and sparsely supervised learning. By exploiting control-affine dynamics and a convex polytope control set, under which the Hamiltonian is maximized at the control vertices, we efficiently generate supervision points via GPU-parallel vertex-restricted tree search, while a residual architecture guarantees that the learned CBF is never larger than the specified constraint function. We evaluate the method on 15 systems and compare it against relevant baselines, showing that it reliably recovers large safe sets where the baselines are conservative or fail completely. In addition, we perform a hardware experiment in which a mobile robot safely avoids pedestrians using a neural CBF trained with our method.
Chinese Translation
随着自主机器人数量持续增长,安全性变得愈发重要。控制屏障函数(CBFs)为确保安全提供了具有理论基础的框架,但现有设计方法往往在有效性、可扩展性或可解释性方面存在局限,并可能导致过于保守的安全集。在本文中,我们提出了VertexCBF,一个以可扩展、系统化和可解释的方式学习神经CBF的框架。我们通过结合物理信息学习和稀疏监督学习训练的神经网络来近似平稳Hamilton-Jacobi值函数。利用控制仿射动力学和凸多面体控制集,在此情况下哈密顿量在控制顶点处最大化,我们通过GPU并行顶点受限树搜索高效生成监督点,而残差架构保证学习到的CBF永远不会大于指定的约束函数。我们在15个系统上评估了该方法,并与相关基线进行比较,表明它能够可靠地恢复大安全集,而基线则保守或完全失败。此外,我们还进行了一项硬件实验,其中移动机器人使用我们方法训练的神经CBF安全地避开行人。
cs.RO / 41 / 2609.12837

Parameter Sensitivity Analysis for Aerial LiDAR-Inertial Odometries in low-altitude flights

低空飞行中空中 LiDAR-惯性里程计的参数敏感性分析
Milijas, Robert, Dios, Jose Ramiro Martinez-de, Bogdan, Stjepan
Abstract
LiDAR-based SLAM (Simultaneous Localization and Mapping) and LIO (LiDAR-inertial odometry) algorithms are often used for precise navigation of unmanned aerial vehicles, especially during interactions with the aerial robot's environment. However, the performance of these algorithms is greatly dependent on the scenario, LiDAR, and robot motion characteristics, often requiring an intensive tuning process to achieve the desired performance. To aid these tuning efforts, this paper analyzes the influence on performance of the parameters of an EKF-based LIO algorithm (FAST-LIO2) and the LIO module of a graph-based SLAM algorithm (Cartographer) on aerial LiDAR SLAM datasets recorded using different LiDARs in low-to-moderate-altitude flights in diverse environments. The analysis is conducted on the absolute trajectory error (ATE) resulting from processing the datasets with the LIO algorithms configured with each combination of parameters obtained in an exhaustive grid search. The relationship between individual parameters and the ATE results is assessed using Pearson's correlation, while the influence of each parameter is assessed using random forest permutation importance analyses with random forest models trained to predict the resulting ATE values based on the choice of parameters. The performed analysis obtains for Cartographer and FAST-LIO2: i) the identification of parameters with stronger influence in performance, ii) a simplified tuning procedure, and iii) tuning recommendations. Using the proposed tuning recommendations, both algorithms obtain on the analyzed datasets ATE values within 5 cm to the optimal performance found in the grid search procedure in 94% of the analyzed cases.
Chinese Translation
基于 LiDAR 的 SLAM(同步定位与建图)和 LIO(LiDAR-惯性里程计)算法通常用于无人机的精确导航,尤其是在与空中机器人环境交互期间。然而,这些算法的性能在很大程度上取决于场景、LiDAR 和机器人运动特性,通常需要大量的调参过程才能达到所需性能。为了辅助这些调参工作,本文分析了基于 EKF 的 LIO 算法(FAST-LIO2)和基于图的 SLAM 算法(Cartographer)的 LIO 模块的参数对性能的影响,这些分析是在不同环境下使用不同 LiDAR 在低至中等高度飞行中记录的空中 LiDAR SLAM 数据集上进行的。该分析基于绝对轨迹误差(ATE)进行,该误差是通过使用在穷举网格搜索中获得的每种参数组合配置的 LIO 算法处理数据集而产生的。使用 Pearson 相关性评估单个参数与 ATE 结果之间的关系,而每个参数的影响则通过随机森林排列重要性分析进行评估,其中随机森林模型经过训练以根据参数选择预测所得的 ATE 值。所进行的分析为 Cartographer 和 FAST-LIO2 获得了:i)识别出对性能影响更强的参数,ii)简化的调参流程,以及 iii)调参建议。使用所提出的调参建议,两种算法在分析的数据集上,在 94% 的分析案例中,获得了与网格搜索过程中找到的最优性能相差 5 cm 以内的 ATE 值。
cs.RO / 42 / 2609.12871

A Multi-Vehicle Dataset with Camera, LiDAR, and Radar Sensors and Scanned 3D Models for Custom Auto-Annotation using RTK-GNSS

一个包含相机、LiDAR和雷达传感器以及用于使用RTK-GNSS进行自定义自动标注的扫描3D模型的多车辆数据集
Berthold, Philipp, Forkel, Bianca, Maehlisch, Mirko
Abstract
Datasets are a crucial element in the development of perception algorithms. They relate sensor measurement data to annotated reference information and allow for the deduction of sensor and object characteristics. In autonomous driving, the reference data commonly consist of semantic image segmentation, point-wise associations, or bounding box annotations. The dataset proposed in this work, however, aims to dig deeper into the evaluation of measurement principles and provides scanned 3D models of all vehicles together with a pose and continuous kinematics reference obtained by RTK-GNSS. Combined, the state of the complete dynamic surrounding of the sensor vehicle is known for any point in time. Subsequent reference formats can be easily computed in user-defined granularity. This dataset involves single-object and multi-object recordings with seven target vehicles. In particular, measurement effects such as occlusion, as well as reflections, can be evaluated, as the normals of the shape of the target vehicles are known. We describe the dataset, discuss the technical background of its development, and briefly present exemplary evaluations.
Chinese Translation
数据集是感知算法开发中的关键要素。它们将传感器测量数据与标注的参考信息关联起来,并允许推导传感器和目标的特性。在自动驾驶中,参考数据通常包括语义图像分割、逐点关联或边界框标注。然而,本工作提出的数据集旨在更深入地评估测量原理,并提供所有车辆的扫描3D模型,以及通过RTK-GNSS获得的位姿和连续运动学参考。结合起来,传感器车辆周围完整的动态环境状态在任意时刻都是已知的。后续的参考格式可以按照用户定义的粒度轻松计算。该数据集包含针对七辆目标车辆的单目标和多目标记录。特别是,由于目标车辆形状的法向量已知,可以评估遮挡以及反射等测量效应。我们描述了该数据集,讨论了其开发的技术背景,并简要展示了示例性评估。
cs.RO / 43 / 2609.12883

From Transportation to Manipulation: Enabling Grasping in Magnetic Robotics

从运输到操作:在磁机器人中实现抓取
Bergmann, Lara, Greis, Noah, Grothues, Cedric, Weigelt, Lisa-Marie, Neumann, Klaus
Abstract
Magnetic levitation (MagLev) systems have great potential for application in high-mix, low-volume manufacturing due to their scalability and flexibility, enabling highly reconfigurable in-machine material flow. However, their manipulation capabilities remain largely unexploited, as current applications almost exclusively focus on transportation. To enable grasping and manipulation directly on MagLev systems without requiring additional costly handling equipment, such as industrial robot arms, we present the Gripper MagBot, a low-cost parallel 6-DoF manipulator with an integrated 1-DoF gripper that mechanically couples three MagLev movers. The Gripper MagBot supports two operating configurations: a default mode and a single-track mode, selectable depending on the required stability and workspace footprint. To reconfigure a machine, the MagBot can be autonomously dropped off and picked up using a docking station. We showcase pick-and-place examples in simulation, as well as with the real Gripper MagBot using our inverse kinematics controller. CAD files, assembly instructions, a component list, and videos are available at https://sites.google.com/view/gripper-magbot.
Chinese Translation
磁悬浮(MagLev)系统由于其可扩展性和灵活性,在高混合、低产量制造中具有巨大应用潜力,能够实现高度可重构的机内物料流。然而,其操作能力在很大程度上仍未被充分利用,因为当前应用几乎只关注运输。为了无需额外昂贵的搬运设备(如工业机械臂)即可直接在 MagLev 系统上实现抓取和操作,我们提出了 Gripper MagBot,一种低成本并联 6-DoF 机械手,集成了 1-DoF 夹爪,并以机械方式耦合三个 MagLev 动子。Gripper MagBot 支持两种操作配置:默认模式和单轨模式,可根据所需稳定性和工作空间占地面积进行选择。为了重新配置机器,MagBot 可以借助对接站自主完成停放与拾取。我们展示了仿真中的拾放示例,以及使用我们的逆运动学控制器在真实 Gripper MagBot 上的拾放示例。CAD 文件、组装说明、组件清单和视频可在 https://sites.google.com/view/gripper-magbot 获取。
cs.RO / 44 / 2609.12894

Before the Tipping Point: Force-Guided Active Perception for Shape-Agnostic Estimation of 3D Centers of Mass

在倾覆点之前:力引导主动感知用于形状无关的3D质心估计
Hyland, Steven M., Xiao, Jing, Onal, Cagdas D.
Abstract
Estimating the 3D center of mass of unknown objects is challenging when grasping is infeasible, geometry is irregular, or mass distribution is uneven. We present a force-based method that estimates CoM height and mass from a single sub-critical tipping experiment by a robot manipulator. The robot applies a quasistatic elevated push and retract motion, using force-angle measurements recorded during tipping to identify parameters from the object trajectory. Our proposed push-retract cycle mitigates frictional bias, enabling generalized fitting. We experimentally validate our method using a robot manipulator with a six-axis force torque sensor on varying types of objects without prior shape information and without specific models. We also propose a method to prevent toppling, keeping the object in a sub-critical tipping regime by leveraging a safety margin. In experimental studies, our method recovers mass, CoM height, and toppling angle with relative errors below 5.0 percent across all unknown objects. This work demonstrates reliable 3D inertial parameter estimation under proper safety thresholds in tipping. Our proposed method informs and enables reliable non-prehensile manipulation and robotic grasping of challenging objects that were previously infeasible.
Chinese Translation
当抓取不可行、几何形状不规则或质量分布不均匀时,估计未知物体的3D质心具有挑战性。我们提出了一种基于力的方法,通过机器人机械臂的单次亚临界倾覆实验来估计质心高度和质量。机器人施加准静态的抬升推拉运动,利用倾覆过程中记录的力-角度测量值从物体轨迹中识别参数。我们提出的推拉循环减轻了摩擦偏差,实现了广义拟合。我们使用配备六轴力/力矩传感器的机器人机械臂,在没有先验形状信息且没有特定模型的情况下,对不同类型物体进行了实验验证。我们还提出了一种防止翻倒的方法,通过利用安全裕度使物体保持在亚临界倾覆状态。在实验研究中,我们的方法对所有未知物体恢复质量、质心高度和倾覆角,相对误差低于5.0%。这项工作展示了在倾覆过程中适当安全阈值下可靠的3D惯性参数估计。我们提出的方法为以前不可行的挑战性物体的可靠非抓取操作和机器人抓取提供了信息并使之成为可能。
cs.RO / 45 / 2609.12927

Robust Underwater Grasping of Sloped Objects with a Waterproof Passive Adaptive Gripper

使用防水被动自适应夹爪的倾斜物体鲁棒水下抓取
Hong, Jooyoung, Hong, Daewon, Kim, Joohyung
Abstract
Robust grasping of everyday objects remains challenging for parallel-jaw grippers, particularly when handling sloped or asymmetric items that induce torque-driven rolling and shear slip. These challenges become even more severe in domestic environments such as kitchens, where objects are often wet or submerged, drastically reducing friction between the gripper and the object. To address these issues, we present a waterproof passive adaptive gripper that combines local and global adaptability for high-performance grasping in the submerged environments. The proposed gripper features a fully waterproof design that integrates a passive rotational joint for global self-alignment on sloped surfaces and passive variable-stiffness pads for local surface adaptation. Experiments in both dry and underwater conditions on various cylindrical and conical objects demonstrate superior holding capability and grasp robustness compared to rigid-pad and fixed-joint baselines. The proposed design offers a practical and robust solution that successfully enables stable underwater manipulation of diverse everyday objects, effectively addressing a critical gap in current underwater robotic grasping for daily-life applications.
Chinese Translation
对于平行夹爪而言,日常物体的鲁棒抓取仍然具有挑战性,尤其是在处理会引发力矩驱动的滚动和剪切滑移的倾斜或非对称物体时。这些挑战在厨房等家庭环境中更加严峻,因为物体通常潮湿或浸没,大幅降低了夹爪与物体之间的摩擦力。为了解决这些问题,我们提出了一种防水被动自适应夹爪,它结合了局部和全局适应性,可在浸没环境中实现高性能抓取。所提出的夹爪采用全防水设计,集成了一个被动旋转关节以实现倾斜表面上的全局自对准,以及被动变刚度垫以实现局部表面适应。在干燥和水下条件下对各种圆柱形和圆锥形物体进行的实验表明,与刚性垫和固定关节基线相比,该夹爪具有更优越的保持能力和抓取鲁棒性。所提出的设计提供了一种实用且鲁棒的解决方案,成功实现了对多种日常物体的稳定水下操作,有效弥补了当前水下机器人抓取在日常应用中的关键空白。
cs.RO / 46 / 2609.12932

ARC: Autonomous Robotics Compliance A Three-Layer Governance Architecture for Deployed Autonomous Systems

ARC:自主机器人合规——面向已部署自主系统的三层治理架构
Eide, Tord, Holt, Einar
Abstract
Proposed governance framework for autonomous robotic systems, introducing a three-layer compliance architecture (ARC) instantiated through model safety validation, cognitive certification benchmarks, and operational authorization standards.
Chinese Translation
针对自主机器人系统提出的治理框架,引入了一种通过模型安全验证、认知认证基准和运行授权标准实例化的三层合规架构(ARC)。
cs.RO / 47 / 2609.12937

A Robot Among People:From Social Imitation to the Social Becoming of Human Groups

人群中机器人:从社会模仿到人类群体的社会性生成
Pham, Victor Tuan Vu, Dörrenbächer, Judith, Weisswange, Thomas H., Hassenzahl, Marc
Abstract
Robots designed to mediate human groups often fall into the solutionist trap: they are framed as sociable agents that fix problems such as conflict, disengagement, or lack of coordination. We suggest a different way of thinking. Rather than discrete agents, robots can be understood as situated elements of shared environments; catalysts and carriers of group experience whose meaning emerges through how people position, interpret, and interact with them. From this perspective, robots are not there to repair some ostensible dysfunctionality, but to enable group-level sense-making around care, norms, and identity. Our prior work on robotic street furniture suggests that this does not happen by imitating human sociality but by taking the shape of deliberately constrained, group-facing entities that happen and act for \textit{us} without being socially entangled as one of us. We thus understand robots in public spaces not in terms of autonomy or intelligence, but as a relational capacity. This implies designing robots not in our image or for our utility, but grounded in our needs in being and becoming together.
Chinese Translation
旨在调解人类群体的机器人常常陷入解决主义陷阱:它们被设定为社交代理,用以修复冲突、脱离或缺乏协调等问题。我们提出一种不同的思考方式。机器人不是离散的代理,而是共享环境中的情境化要素;是群体经验的催化剂和载体,其意义通过人们如何定位、解读和与之互动而显现。从这个角度来看,机器人不是用来修复某种表面上的功能失调,而是围绕关怀、规范和身份促成群体层面的意义建构。我们先前关于机器人街道家具的研究表明,这并非通过模仿人类的社会性来实现,而是通过采取刻意受限的、面向群体的实体形态,这些实体为我们而存在和行动,却不在社会关系上纠缠为我们中的一员。因此,我们不是从自主性或智能的角度来理解公共空间中的机器人,而是将其视为一种关系能力。这意味着设计机器人不是按照我们的形象或为了我们的效用,而是植根于我们共同存在和生成的需求。
cs.RO / 48 / 2609.12959

Distributed Stochastic Optimal Control for Pattern-Oriented Swarms

面向图案集群的分布式随机最优控制
Zhang, Qingrui, Yu, Chenghao, Xue, Feng, Wang, Xintong
Abstract
While offering significant promise for diverse applications, pattern-oriented swarms encounter multifaceted challenges in geometric control, self-organization, and safe navigation through dynamic environments. In this paper, we present a GRF-based stochastic optimal control framework to address these challenges within a unified probabilistic architecture. By extending the GRF into the temporal domain, the proposed framework casts collective coordination as a Bayesian inference task, enabling swarms to accommodate environmental uncertainty, satisfy non-convex constraints, and reconcile heterogeneous dynamics across diverse platforms. We develop an uncertainty- and safety-aware collision avoidance module for navigation in the presence of stochastic obstacle motion. The unscented transform is employed to propagate state uncertainty for both dynamic obstacles and neighboring agents, yielding principled confidence bounds for collision avoidance. In addition, density-guided pattern control is introduced, which encodes geometric patterns as implicit density fields. This representation decouples pattern specification from explicit agent-to-target assignments, thereby facilitating intrinsic self-healing and elastic reconfiguration in a distributed manner. The proposed framework is extensively evaluated through Monte Carlo simulations across diverse scenarios. Its model-agnostic nature is demonstrated on both quadrotor and fixed-wing UAV swarms, highlighting its generalizability across platforms with heterogeneous dynamics. Finally, the efficacy and robustness of the proposed method are validated through indoor experiments with a 15-quadrotor swarm and outdoor deployments involving 4 custom-built autonomous quadrotors. These experiments substantiate the proposed framework's capacity to maintain reliable geometric pattern transitions and safety-aware navigation within real-world environments.
Chinese Translation
尽管面向图案的集群为多样化的应用提供了巨大的前景,但在几何控制、自组织以及在动态环境中的安全导航方面面临着多方面的挑战。在本文中,我们提出了一个基于GRF的随机最优控制框架,以在一个统一的概率架构内解决这些挑战。通过将GRF扩展到时间域,所提出的框架将集体协调转化为贝叶斯推理任务,使集群能够适应环境不确定性、满足非凸约束,并协调不同平台间的异构动力学。我们开发了一个不确定性和安全性感知的避碰模块,用于在随机障碍物运动情况下的导航。采用无迹变换来传播动态障碍物和相邻智能体的状态不确定性,从而为避碰生成有原则的置信边界。此外,引入了密度引导的图案控制,它将几何图案编码为隐式密度场。这种表示将模式指定与显式的智能体到目标的分配解耦,从而以分布式方式促进内在的自愈和弹性重构。所提出的框架通过蒙特卡洛仿真在多种场景下进行了广泛评估。其模型无关的特性在四旋翼和固定翼无人机集群上得到了验证,突显了其在具有异构动力学的平台上的泛化能力。最后,通过15架四旋翼集群的室内实验和涉及4架定制自主四旋翼的室外部署,验证了所提方法的有效性和鲁棒性。这些实验证实了所提框架在真实环境中保持可靠的几何图案转换和安全性感知导航的能力。
cs.RO / 49 / 2609.12971

Tuning ROS 2 for Energy-Efficient Navigation: Empirical Insights from Costmap 2D Configurations

面向节能导航的 ROS 2 调优:来自 Costmap 2D 配置的实证见解
Albonico, Michel, Wortmann, Andreas, Malavolta, Ivano
Abstract
Robots are increasingly used in diverse application areas, where autonomous navigation plays a central role. As these systems become more widespread, improving their energy efficiency is critical to extending operational time and reducing environmental impact. The Robot Operating System (ROS) is a widely adopted middleware for robotics, offering a rich set of configurable packages. However, this flexibility can result in suboptimal software configurations in dynamic environments, negatively affecting both performance and energy consumption. This paper investigates the impact of ROS 2 package reconfigurations on the energy efficiency of mobile robot navigation. We conduct a controlled experiment in two warehouse-like scenarios (small and large) with varying obstacle layouts and Costmap 2D configurations (essential to the Nav2 stack). Through repeated trials, we measure energy usage, power profile, CPU load, memory consumption, and navigation performance. Results show that configurations must be carefully chosen for the specific robotic environment, and we were able to identify critical settings that lead to good and poor performance and energy consumption.
Chinese Translation
机器人正越来越多地应用于各种领域,其中自主导航发挥着核心作用。随着这些系统日益普及,提高其能效对于延长运行时间和减少环境影响至关重要。机器人操作系统(ROS)是一种被广泛采用的机器人中间件,提供丰富的可配置软件包。然而,这种灵活性可能导致动态环境中的软件配置欠佳,从而对性能和能耗产生负面影响。本文研究 ROS 2 软件包重配置对移动机器人导航能效的影响。我们在两种类仓库场景(小型和大型)中开展受控实验,改变障碍物布局和 Costmap 2D 配置(对 Nav2 栈至关重要)。通过重复试验,我们测量能耗、功率曲线、CPU 负载、内存消耗和导航性能。结果表明,必须针对特定机器人环境仔细选择配置,并且我们能够识别出导致良好和较差性能与能耗的关键设置。
cs.RO / 50 / 2609.13011

Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies

构造即舒适:面向学习驾驶策略的自适应、舒适度约束动作空间
Rothenhäusler, Anna, Jost, Daniel, Rajan, Raghu, Janjos, Faris, Scheel, Oliver, Look, Andreas, Boedecker, Joschka
Abstract
Data-driven driving simulators command accelerations and steering rates from a fixed grid without constraining the realized accelerations and jerks. As a result, reinforcement-learning policies inflate safety metrics through abrupt, last-second maneuvers that lie far outside the range of human driving and would be unacceptable to occupants of a real vehicle, so the metrics measure simulator permissiveness rather than policy quality. Enforcing comfort bounds naively is not enough: lateral limits shrink quadratically with speed, so clamping a static grid saturates it and destroys fine-grained control ("grid collapse"). We propose an adaptive action parameterization that rediscretizes the grid at every step to span exactly the per-step feasible control set, via closed-form inversion of the lateral-jerk constraint. We further present PufferDrive-Editor, a browser-based tool to audit realized kinematics and author kinematically challenging scenes. On the Waymo Open Motion Dataset and a hand-authored slalom, our adaptive model holds comfort violations below 1% while outperforming clipped-grid and direct-jerk baselines in navigability.
Chinese Translation
数据驱动的驾驶模拟器从固定网格中指令加速度和转向速率,而不约束实际产生的加速度和加加速度。因此,强化学习策略通过突然的、最后一秒的机动来夸大安全指标,这些机动远远超出人类驾驶的范围,并且对真实车辆的乘员来说是不可接受的,所以这些指标衡量的是模拟器的宽容度而非策略质量。简单地强制舒适度边界是不够的:横向限制随速度呈二次方缩小,因此钳制静态网格会使其饱和并破坏细粒度控制(“网格崩溃”)。我们提出了一种自适应动作参数化方法,该方法通过横向加加速度约束的闭式反演,在每一步重新离散化网格,以精确覆盖每步可行的控制集。我们进一步提出了 PufferDrive-Editor,一个基于浏览器的工具,用于审计实际运动学并创作运动学上具有挑战性的场景。在 Waymo 开放运动数据集和手工编写的蛇形测试中,我们的自适应模型将舒适度违规保持在 1% 以下,同时在可导航性方面优于裁剪网格和直接加加速度基线。
cs.RO / 51 / 2609.13015

Global Path Planner with Multi-Model Switching

多模型切换的全局路径规划器
Gori, Pietro, Iotti, Francesco, Zelenay, Eduard, Marko, Rastislav, Pierallini, Michele, Angelini, Franco, Pannocchia, Gabriele, Garabini, Manolo
Abstract
This work enhances global path planning via a pure-pursuit controller with multi-model kinematic switching that sustains plan fidelity across diverse terrains. The system includes a traversability graph for terrain analysis, a Heading-Aware A* algorithm for generating feasible paths, and a multi-model Pure Pursuit controller for dynamic tracking. A core innovation is adaptive kinematic modeling, enabling real-time switching between kinematic models based on terrain features and robot states. This adaptability optimizes path efficiency and energy use in challenging scenarios. We validate the approach in simulation on different platforms, namely the Artaban quadruped and the X3 quadrotor drone, showcasing improved performance, robustness, and adaptability over standard baselines.
Chinese Translation
本工作通过一种带有多种运动学模型切换的纯跟踪控制器来增强全局路径规划,该控制器能够在多样地形中保持规划保真度。系统包括用于地形分析的可通行性图、用于生成可行路径的 Heading-Aware A* 算法,以及用于动态跟踪的多模型 Pure Pursuit 控制器。核心创新是自适应运动学建模,能够根据地形特征和机器人状态在运动学模型之间实时切换。这种自适应性在具有挑战性的场景中优化了路径效率和能量使用。我们在不同平台上通过仿真验证了该方法,即 Artaban 四足机器人和 X3 四旋翼无人机,展示了相较标准基线在性能、鲁棒性和适应性方面的提升。
cs.RO / 52 / 2609.13053

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

Dynin-Robotics:全模态统一扩散视觉-语言-动作模型
Lee, Hoeun, Kim, Jaeik, Oh, Jusang, Kim, Jinhyeok, Choi, Geon, Kim, Hyeonggeun, Do, Jaeyoung
Abstract
Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation and selection through a shared trajectory model. Dynin-Robotics implements this formulation on Dynin-Omni, an omnimodal masked-diffusion backbone, representing language, visual observations, goals, and actions as discrete tokens. By varying conditioning and target spans, the same model learns action prediction, action-conditioned next-observation prediction, terminal goal-state prediction, and trajectory-to-instruction reconstruction. These interfaces support test-time scaling through goal prediction, action-candidate evaluation, and joint refinement of action and future-state predictions. We continually pretrain the model on approximately 1.33 million trajectories from 48 Open X-Embodiment datasets and adapt it separately to downstream domains. On two VLABench tasks, robot pretraining improves adaptation within a fixed Stage-2 step budget, and the full objective mixture improves shifted-instruction success over Policy-only post-training under the same coupled decoder. Combining goal guidance with joint action-next-state denoising further improves shifted-instruction success over action-only decoding; the benefit depends on how the predictions are composed. Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot. An optimized block-parallel implementation accelerates model-side action decoding by up to 29.2x relative to the base implementation under the reported profiling setup. These results support shared trajectory modeling as a common interface for learning complementary robot objectives and composing their predictions during control.
Chinese Translation
视觉目标和动态预测可以为语言条件化的机器人策略提供目标结果以及动作依赖的场景变化的表示。我们通过共享轨迹模型将这些预测引入动作生成和选择中。Dynin-Robotics 在 Dynin-Omni 上实现了这一公式,Dynin-Omni 是一个全模态掩码扩散骨干网络,将语言、视觉观察、目标和动作表示为离散 token。通过改变条件和目标跨度,同一模型学习动作预测、动作条件下的下一观察预测、终端目标状态预测以及轨迹到指令的重建。这些接口通过目标预测、动作候选评估以及对动作和未来状态预测的联合细化来支持测试时扩展。我们在来自 48 个 Open X-Embodiment 数据集的约 133 万条轨迹上持续预训练该模型,并将其分别适配到下游领域。在两个 VLABench 任务上,机器人预训练在固定的 Stage-2 步骤预算内改善了适配,并且完整的目标混合在相同的耦合解码器下,相比于仅策略后训练,提高了偏移指令的成功率。将目标引导与联合动作-下一状态去噪相结合,进一步提高了偏移指令的成功率,优于仅动作解码;其收益取决于预测如何组合。Dynin-Robotics 在 LIBERO 和零样本 LIBERO-Plus 上取得了有竞争力的性能,并在 Franka Research 3 机器人上四种操作条件下平均成功率达到 78.4%。在报告的性能分析设置下,优化的块并行实现相对于基础实现将模型侧动作解码加速了高达 29.2 倍。这些结果支持共享轨迹建模作为学习互补机器人目标并在控制期间组合其预测的通用接口。
cs.RO / 53 / 2609.13083

ASTRIL-MPC: Autonomous Traversal Framework of Articulated Tracked Robots with Language-Guided Neural-Kinematic MPC

ASTRIL-MPC:基于语言引导神经运动学模型预测控制的关节式履带机器人自主通行框架
Gan, Zhenfeng, Chen, Yanbo, Che, Lirong, Tan, Junbo, Wang, Xueqian
Abstract
In urban search and rescue, articulated tracked robots (ATRs) must traverse structured but contact-rich environments such as stairwells and cluttered building interiors. Reliable autonomy remains challenging because robot-terrain interaction (RTI) is hybrid and discontinuous, and effective flipper-track coordination is difficult to model analytically. We present ASTRIL-MPC, a language-guided neural kinematics model predictive control (MPC) framework for autonomous traversal. A learned kinematics model predicts short-horizon task-state increments from a height sequence and recent trajectories; NMPC plans with multi-objective costs and strict feasibility constraints; and a large language model (LLM) proposes bounded updates to selected weights and bounds through a safety-checked interface with range clipping, rate limiting, and consistency checks. The compiled predictor enables a full control cycle within 100 ms. Across three traversal tasks and a multi-height generalization setting, ASTRIL-MPC improves an aggregate traversal-quality score by up to 71% over a non-adaptive NMPC and by 67% over a PPO baseline, while eliminating measurable collision impacts during descent. These results indicate that combining learned kinematics, optimization-based planning, and language-guided retuning yields data-efficient and robust autonomy for articulated tracked robots.
Chinese Translation
在城市搜救中,关节式履带机器人(ATRs)必须穿行于结构化但接触密集的环境,例如楼梯间和杂乱的建筑内部。可靠的自主性仍然具有挑战性,因为机器人-地形交互(RTI)是混合且不连续的,并且有效的摆臂-履带协调难以通过解析方法建模。我们提出了 ASTRIL-MPC,一种用于自主通行的语言引导神经运动学模型预测控制(MPC)框架。一个学习到的运动学模型从高度序列和近期轨迹中预测短时域任务状态增量;非线性模型预测控制(NMPC)使用多目标代价和严格可行性约束进行规划;大语言模型(LLM)通过具有范围裁剪、速率限制和一致性检查的安全检查接口,对选定的权重和边界提出有界更新。编译后的预测器可在 100 毫秒内实现完整控制周期。在三个通行任务和一个多高度泛化设置中,ASTRIL-MPC 相比非自适应 NMPC 将综合通行质量分数提高了最多 71%,相比 PPO 基线提高了 67%,同时消除了下降过程中可测量的碰撞冲击。这些结果表明,结合学习到的运动学、基于优化的规划和语言引导的重新调参,可为关节式履带机器人带来数据高效且鲁棒的自主性。