Scenario-Independent Criticality Assessment and Prediction for Vulnerable Road Users in Autonomous Driving
自动驾驶中弱势道路使用者的场景独立关键性评估与预测
Gamerdinger, Jörg, Schwarzenberger, Victor, Schmid, Philipp, Teufel, Sven, Bringmann, Oliver
Abstract
Increasing safety is the primary objective of automated vehicles. Achieving this goal requires reliable safety metrics that incorporate safety-relevant factors such as object type, velocity, and criticality. A key capability of such metrics is the distinction between critical and non-critical objects, which is addressed through criticality or relevance estimation. Existing criticality metrics are typically designed for specific scenarios and primarily focus on vehicle-to-vehicle interactions. In this paper, we therefore propose a novel criticality metric tailored to vulnerable road users (VRUs), which require special consideration due to their less predictable motion behavior. Furthermore, to avoid the complexity introduced by scenario-specific metrics, we introduce a scenario-independent criticality prediction framework applicable to all traffic participant classes. The effectiveness of both the proposed VRU-centric criticality metric and the criticality prediction framework is evaluated using the DeepAccident dataset, which contains a diverse set of safety-critical traffic scenarios. The proposed VRU-centric criticality metric improves pedestrian criticality classification performance by up to 50 %. In addition, the proposed criticality prediction framework outperforms state-of-the-art metrics by 275 %, achieving an F1-score of 0.96 and enabling scenario-independent criticality assessment across all object classes. These results demonstrate the strong potential of the proposed approaches to enhance criticality assessment for safety evaluation in automated driving systems.
In this technical report, we propose Pelican-Sim 1.0, a general world model simulator for embodied intelligence that predicts future observations from visual context and robot actions to support downstream learning and decision making. The model incorporates four key design features: (1) Unified action representation: a 28-dimensional action value space covering most mainstream embodiments, keeping one model valid across heterogeneous devices. (2) Action-visual injection: URDF- and camera-rendered action videos bridge actions and pixels, giving markedly better controllability across embodiments, scenes, and tasks (PSNR +0.904 over alternative fusion baselines). (3) Sparse mixture-of-experts (MoE): sparse MoE layers add capacity for heterogeneous dynamics and absorb the action modality while reducing inter-modality conflict (FVD -6.530 vs. the dense backbone). (4) Efficient rollout generation: causal adaptation and few-step distillation yield a four-step autoregressive simulator, achieving a 5.67-fold speedup over the 35-step model. Benefiting from these designs, we train on approximately one million real-world and simulated trajectories and obtain large gains in action controllability and video quality: PSNR improves over the strongest evaluated baselines by 4.636 on AgiBotWorld Beta, 2.080 on RoboMIND, and 10.343 on RoboTwin, with the adapted EWMBench DYN score up 0.426 on RoboTwin. Relying on this, four downstream applications on RoboTwin succeed: 500 generated trajectories added to 50 demonstrations per task raise policy success from 70% to 93%; policy evaluation reaches a Pearson correlation of 0.994 across five checkpoints; and relative success gains reach 47.7% for action selection and 20.3% for policy improvement. Qualitative generalization across trajectory, scene, object, embodiment, and viewpoint shifts highlights its potential as a general-purpose world model simulator.
Mobile manipulation requires perceptual evidence at different spatial scales for base motion and arm control, while the two action modalities remain kinematically coupled. Existing policies often employ specialized action generation for different subsystems but condition heterogeneous action branches on a shared perceptual representation, leaving subsystem-specific perception-action correspondence implicit. We present MoPA, a framework that aligns perceptual conditioning with mobility and manipulation while preserving coordination at the action level. Dual Perceptual Streams employ two mutually masked query banks to extract separate perceptual representations from a shared vision-language context. Perception2Action Adaptation jointly updates each query bank and its corresponding action stream at every layer of a structured Mixture-of-Transformers decoder, while enabling information exchange between the two action streams. Coupled conditional flow matching learns a joint vector field for coordinated generation of both action chunks. On the ManiSkill-HAB benchmark, MoPA achieves state-of-the-art performance across all three task suites. Across four real-world tasks, MoPA achieves a mean full-task success rate of 76.3%, outperforming the best baseline by 12.5 percentage points. Ablation studies and further analyses validate the effectiveness of the proposed design. Website is available at: https://mopa-policy.github.io/.
Uncertainty-Aware Conflict Detection Against Operator-Conditioned Weather Hazards
面向操作员条件化天气危险的不确定性感知冲突检测
Kandoria, Balram, Kim, Seulki, Samyal, Aryaman Singh
Abstract
Strategic flight plan validation in Advanced Air Mobility (AAM) environments requires robust methods for predicting aircraft state uncertainty and detecting potential conflicts with dynamic airspace hazards. This paper presents a novel framework for uncertainty-conditioned trajectory prediction combined with polyhedra hazard representation for pre-flight conflict detection. We introduce a closed-form uncertainty estimation method that couples non-uniform rational B-spline (NURBS) curve fitting for kinematic trajectory generation with a Kalman Filter for state covariance propagation. Drawing from the Light Propagation Algorithm (LPA) paradigm, we employ a sigmoid-blended measurement noise model that captures the uncertainty reduction behavior of flight management systems approaching the required time of arrival (RTA) for waypoints. The resulting temporal uncertainty bounds are derived through a velocity-to-time variance transformation, enabling probabilistic assessment of arrival time deviations along the flight path. For hazard representation, we develop an operator-conditioned classification scheme that transforms gridded environmental data, specifically weather phenomena, into three-dimensional polyhedra volumes with intensity-based stratification. These hazard polyhedra incorporate aircraft-specific safety buffers computed from vehicle performance characteristics. Conflict detection is performed through mesh intersection algorithms operating on the spatial uncertainty tube surrounding the mean trajectory against the hazard polyhedra and temporal overlap. The framework enables the continuous strategic validation of flight plans throughout the pre-flight planning time horizon as environmental conditions evolve.
Slender rod insertion arises in precision manufacturing, where millimetre scale diameter and tight clearances demand accurate perception and control. Conventional peg-in-hole methods assume a rigid object whose tip pose is fixed relative to the gripper. This assumption breaks down for a high aspect ratio rod, which can bend during manipulation, making its tip motion dependent on the rod configuration, grasp, material properties, and contact. We present RodForesight, a learning framework that factorises the task into two stages: 1) coarse approaching, which uses visual servoing to map diverse initial configurations into a compact near hole hand-off region; and 2) predictive insertion, which performs fine alignment and completes the insertion. It is worth noting that the two stages can be wrapped into an end-to-end design. During insertion, a diffusion policy generates candidate action chunks, while an action conditioned world model predicts their effects on rod-hole alignment. This pre-execution evaluation enables RodForesight to select the best action chunk based on predicted tilt and radial errors before execution. Experiments investigate the performance of different stages and the end-to-end setting, where RodForesight improves the success rate from 88.9% to 96.7%, compared to baseline methods such as diffusion policy.
Functional integration is a growing trend in vehicle control, often involving the coordination of multiple controllers to achieve various objectives simultaneously. The need for flexibility and reliability has led to a "plug-and-play" approach in control system design, which presents challenges for traditional integrated model predictive control (MPC). Agent-based model predictive control (AMPC) has recently emerged as a distributed solution that treats controllers as agents, creating a collaborative framework among them to reach a common goal. However, this approach struggles to manage distributed conflicting objectives when agents are coupled or interdependent. To address this, we propose a novel, practical distributed control scheme called multi-objective AMPC, which adapts the alternating direction method of multipliers (ADMM) into a general control strategy that approximates global optimization while decoupling objectives. We systematically develop three formulations that maintain convergence while addressing control regularization and inequality constraints, applying them to complex vehicle control systems for the first time. The proposed method has been tested on two vehicle control scenarios with a multi-objective topology. Different formulations are compared through simulations, and the most computationally efficient one was implemented on an electric vehicle for real-world evaluations. The results demonstrate that the proposed multi-objective AMPC can converge approximately to the same global optimum as integrated MPC with greater flexibility and the potential to reduce computational costs.
A Data-Driven Distributed Control Scheme: Learning Multi-Objective Agent-Based MPC for Path-Tracking
数据驱动的分布式控制方案:学习多目标基于智能体的MPC用于路径跟踪
Zhong, Jiaming, Mehrizi, Reza Valiollahi, Pant, Yash Vardhan, Khajepour, Amir
Abstract
Agent-based model predictive control (AMPC) has recently been proposed for vehicle systems with various controllers, such as differential braking and torque vectoring, where controllers are regarded as distributed agents contributing to the same objective. However, this scheme is challenging in handling multiple conflicting objectives with coupled agents. A common approach for such tasks is the integrated MPC, where all objectives and agents are stacked together in one optimization. Nevertheless, as more agents and objectives are involved, the integrated MPC will face challenges like computational burdens and maintenance difficulties in practice. To this end, this paper proposes a learning multi-objective AMPC that can improve design flexibility and computing efficiency. First, under the assumption of information exchange, a multi-objective AMPC tailored from the alternating direction method of multipliers (ADMM) is proposed to decouple the system and achieve the same performance as the integrated scheme iteratively. Second, a learning-based method for initializing iterations is proposed to accelerate convergence. In addition, a data management method is proposed for real-time efficiency, and an authentication module is designed for learning reliability. We compare the proposed scheme against the integrated scheme via a combined path-tracking simulation for autonomous vehicles with various controllers. The proposed scheme achieves the same control performance as the integrated one while reducing the computational time by 43.5%. Furthermore, the learning-based method saves 88.6% more computational time than without learning, making it suitable for real-time implementation.
Robotic bioprinting and Direct Ink Writing (DIW) are being explored towards the treatment of Volumetric Muscle Loss (VML). While previous studies have shown the importance of proper parameter selection on the print outcome, existing approaches often rely on time- and material-intensive design of experiments methods, or require large, well-curated datasets for training machine learning models. In this paper, we propose a physics-based closed-loop robotic bioprinting system capable of near real-time parameter adaptation. The system integrates a 3D point cloud camera and fully autonomous vision-based algorithms to provide quantitative evaluation of printed constructs. This evaluation is fed into a controller that adjusts printing parameters to achieve a desired bead thickness. To assess the framework's performance, four experimental configurations were tested, each repeated three times. In these tests, printing began from an arbitrary initial parameter value, and the controller was tasked with adjusting the parameters to reach the desired thickness. The system converged in all trials, achieving a tracking error below 0.5 mm within an average of 5.2 seconds from the start of printing. The low standard deviation of the converged pressure over different tests (0.04 bar on average) demonstrates robustness and repeatability. Additional experiments were conducted with the controller turned off, enabling direct comparison with open-loop DIW bioprinting, further confirming the effectiveness of the proposed closed-loop framework in achieving the desired bead geometry.
Battery-Aware Predictive Trajectory Planning and Control for Multirotors Under Disturbances
多旋翼在扰动下的电池感知预测轨迹规划与控制
Kidambi, Krishna Bhavithavya
Abstract
This paper presents a battery-aware predictive trajectory-planning and control framework for multirotors operating under spatially localized disturbances. Candidate trajectories are evaluated through closed-loop vehicle--motor--battery propagation, allowing disturbance-induced control demand, electrical energy, battery evolution, and terminal-voltage-dependent actuator capability to enter the planning process. % A reduced-order battery model is numerically benchmarked against an independently implemented Simscape equivalent-circuit reference, with a power NRMSE of $0.64\%$ and a cumulative-energy discrepancy below $0.7\%$. % In a $150$-s, $640$-m mission containing three disturbance regions, the selected trajectory reduces electrical energy consumption by $7.46\%$ and position-tracking RMSE by approximately $72\%$ relative to the disturbance-aware fixed-reference baseline. % Planner ablations show that battery-dependent terms are nonbinding at nominal SOC but alter the selected trajectory under a depleted-battery stress condition. % Execution with multiple feedback controllers further demonstrates that controller selection changes the tradeoff among tracking accuracy, energy consumption, and actuator utilization. % The results demonstrate the benefit of accounting for predicted closed-loop energetic and battery--actuator consequences during trajectory selection.
An Automated Thickness Evaluation Procedure Using an Integrated Structured Light 3D Camera in a Robotic Bioprinting Framework
一种在机器人生物打印框架中使用集成结构光3D相机的自动化厚度评估流程
Zobeidi, Ehsan, Rezayof, Omid, Alambeigi, Farshid
Abstract
Bioprinting is emerging as a tissue engineering technique to replace common treatment methods for large scale injuries. While thickness of the BioPrinted Constructs (BPCs) have shown to be of importance in the cell maturation and integration, the literature lacks a robust, automated, and quantitative method for measuring these metrics. In this paper, we propose a fully automated vision-based method for measuring the thickness of the BPCs with complex geometries. Leveraging the point cloud and RGB images of a structured light 3D camera, our proposed method performs an image-based segmentation for delineating the BPCs from the RGB images, accompanied by novel geometry-based thickness measurement algorithms performed on the point cloud scans. These algorithms combine the segmentation mask with the robot's forward kinematics data and a 3D point cloud scan to precisely measure the aforementioned metrics for complex-shaped BPCs. The proposed method was evaluated in simulation and experimental studies. In simulation studies, the algorithms were used to measure the thickness of some virtually created BPCs with known thickness. The comparison between the measured and true thicknesses demonstrates the high accuracy of the proposed method, achieving mean absolute errors between 0.025 mm and 0.057 mm in simulation at a spatial resolution of 0.1 mm x 0.1 mm per pixel. Furthermore, we successfully deployed the algorithms on our robotic bioprinting setup utilizing a structure light 3D camera, where complex patterns were printed and the developed methods utilized to accurately measure the thickness of printed BPCs.
Chinese Translation
生物打印正逐渐成为一种组织工程技术,用于替代大面积损伤的常见治疗方法。尽管生物打印构建体(BPCs)的厚度已被证明在细胞成熟和整合中具有重要性,但文献中仍缺乏一种稳健、自动化且定量的测量这些指标的方法。在本文中,我们提出了一种全自动的基于视觉的方法,用于测量具有复杂几何形状的BPC的厚度。利用结构光3D相机的点云和RGB图像,我们提出的方法执行基于图像的分割,以从RGB图像中勾勒出BPC,并结合在点云扫描上执行的新的基于几何的厚度测量算法。这些算法将分割掩膜与机器人的正运动学数据和3D点云扫描相结合,以精确测量复杂形状BPC的上述指标。所提出的方法在仿真和实验研究中进行了评估。在仿真研究中,这些算法被用于测量一些已知厚度的虚拟创建的BPC的厚度。测量厚度与真实厚度的比较表明,所提出的方法具有高精度,在每像素0.1 mm x 0.1 mm的空间分辨率下,仿真中平均绝对误差在0.025 mm至0.057 mm之间。此外,我们成功地将这些算法部署到我们的机器人生物打印装置上,使用结构光3D相机,打印了复杂图案,并使用所开发的方法准确测量了打印的BPC的厚度。
Self-improving agent workflows create an audit problem when the same controller can change both its behavior and the conditions under which that behavior is judged. We present GuardrailLoop, a simulation-based testbed that makes three operational contracts jointly testable: preservation of human-defined policy, compute accounting at every recorded execution prefix, and recovery of a specified scientific state after crashes. A hash-pinned policy fixes goals, scope, evaluation identity, budget, and release conditions; machine-directed evolution is restricted to a code-owned feature catalog and bounded knobs. The contribution is an executable boundary and an evaluation protocol that separates useful adaptation, state recovery, and repeated execution. In a paired 50-seed 2 x 2 study, round-stage growth changes target attainment by +1.00 and restricted mean compute to target by -56.97 simulated GPU-hours (95% paired-bootstrap interval [-58.91,-54.70]); idle growth has zero measured utility effect. Across 240 enumerated crash injections, all runs recover the defined outcome, but only 210 preserve the normalized trace: 30 pre-commit crashes repeat a planner call. Resource-drift, kill-switch, integrity, and output-guard matrices satisfy their specified checks. These findings show why successful outcome recovery is insufficient evidence of exactly-once execution. They establish conformance within one calibrated deterministic testbed, rather than general safety or real-world self-improvement.
Maintaining consistency over long spatial and temporal horizons remains a fundamental challenge in large-scale LiDAR SLAM, particularly when integrating maps collected across multiple sessions. We present Chain-SLAM, a LiDAR SLAM backend enabling online multi-session map alignment and reuse with global consistency at large scale. We implement a chained loop closure mechanism that efficiently propagates geometric constraints across inter-session keyframes through an adjacency graph, enabling robust long-horizon consistency triggered by reliable short-horizon loop closures. The system initializes inter-session alignment with GNSS-proximity place recognition, then performs on-the-fly loop closure detections and joint optimization of loaded maps and newly acquired trajectories within a unified factor graph, maintaining both inter- and intra-session geometric consistency without dynamic object removal, and cross-platform robustness with minimal hyperparameter tuning. Experimental results show improved trajectory accuracy and robust multi-session integration on large-scale datasets. We release our source code to support reproducible research in large-scale multi-session LiDAR SLAM. Project site: https://ai4ce.github.io/Chain-SLAM/
Chinese Translation
在大规模 LiDAR SLAM 中,保持长时间和空间范围的一致性仍然是一个基本挑战,特别是在整合跨多个会话收集的地图时。我们提出了 Chain-SLAM,一个 LiDAR SLAM 后端,能够在大规模下实现具有全局一致性的在线多会话地图对齐和重用。我们实现了一种链式回环机制,通过邻接图在跨会话关键帧之间高效传播几何约束,从而通过可靠的短时回环触发稳健的长时一致性。该系统使用 GNSS 邻近地点识别初始化会话间对齐,然后在统一因子图内实时执行回环检测,并对加载的地图和新采集的轨迹进行联合优化,无需动态物体去除即可保持会话间和会话内的几何一致性,并以最少的超参数调整实现跨平台鲁棒性。实验结果表明,在大规模数据集上,轨迹精度得到提高,多会话集成具有鲁棒性。我们发布了源代码,以支持大规模多会话 LiDAR SLAM 的可重复研究。项目网站:https://ai4ce.github.io/Chain-SLAM/
DIA: Denoising Intermediate Advantage for Diffusion Policy Optimization
DIA:面向扩散策略优化的去噪中间优势
Sohal, Arjun, Zhao, Yuchi, Bogdanovic, Miroslav, Aspuru-Guzik, Alan
Abstract
Diffusion-based robot policies have become widely used in robotic manipulation, where they are typically trained with behavior cloning. However, policies trained purely from demonstrations are limited by the quality and coverage of the available data. Reinforcement learning can further improve the performance of these pretrained policies through interaction. A common approach is to use policy-gradient methods that formulate diffusion-policy fine-tuning as an outer environment MDP together with an inner denoising MDP. However, existing methods typically assign the same environment-level credit to all denoising steps used to construct an action chunk, without distinguishing which intermediate decisions contributed most to the final return. We introduce Denoising Intermediate Advantage (DIA), a policy-gradient method that learns a value function over partially denoised actions and uses it to construct a denoising- level advantage for each step of the generative process. DIA com- bines this inner credit signal with the standard environment-level PPO advantage, providing state-dependent credit throughout the denoising chain. Across Robomimic, FurnitureBench, Franka Kitchen, and D3IL, DIA consistently improves final performance over existing diffusion-policy fine-tuning methods. Beyond final reward, DIA reaches successful states more efficiently and can shift farther from the pretrained behavior distribution, enabling it to discover more effective and efficient task-level strategies and subtask sequences that baseline methods fail to reach.
Trajectory Bundle Method in SE(3) for Black-Box Fixed-Wing Aircraft Trajectory Optimization
SE(3)中的轨迹束方法用于黑箱固定翼飞机轨迹优化
Osburn, Matthew D., Peterson, Cameron K., Salmon, John L.
Abstract
Dynamically feasible trajectory optimization for rigid-body systems is naturally formulated on the special Euclidean group SE(3) but is challenging when dynamics are available only as black-box computations without derivatives. This paper formulates the Trajectory Bundle Method (TBM) for motion planning implicitly on SE(3). Bundles are constructed in the Lie algebra and propagated through nonlinear rigid-body dynamics using exponential and logarithmic maps, enabling derivative-free planning of non-Euclidean trajectories. We show that Euclidean TBM interpolation error is bounded quadratically by bundle diameter and extend this result to SE(3), where the bound additionally depends on a local Lipschitz constant of the Log map. Numerical experiments corroborate these bounds. Finally, we demonstrate SE(3) TBM by optimizing an acrobatic, collision-free fixed-wing maneuver through a rotated aperture without explicit models or derivatives of the vehicle dynamics, aerodynamics, or collision model.
Pneumatic neurons for soft robots enable inflate-and-fire networks for rhythmic motion
用于软体机器人的气动神经元实现节律运动的充气-激发网络
Li, Dongting, Tolley, Michael, Gravish, Nick
Abstract
Animals coordinate their movements through distributed neural circuits, but soft robots still typically depend on external, centralized electronics for control. Building soft robots that operate without centralized electronic controllers while remaining responsive to their environment remains a frontier challenge in soft robotics. In this work we introduce a soft-robot control architecture inspired by leaky integrate-and-fire models of biological neural circuits. The Pneumatic neuron (Pneu-ron) is a soft actuator that unifies energy conversion, logic, and actuation in one component. Each module combines a low-boiling-point fluid (LBF), a heater, and a mechanical switch into a self-excitable unit. Boiling the LBF inflates the module and triggers excitation and inhibition of adjacent modules in a process we call "inflate-and-fire". When interconnected into excitatory-inhibitory rings, Pneu-rons generate stable, sequential oscillations whose frequency emerges from the material dynamics and environmental conditions. By harnessing the inflation of Pneu-rons for actuation these networks can drive oscillatory locomotion of soft robots. Pneu-ron networks sustain oscillation under mechanical load and thermal variations, adapting through material physics rather than computation. Dynamical modeling of these networks reveals a dimensionless bifurcation diagram that dictates the network's oscillatory behavior. Encoding logic and actuation into material-level modules presents a new avenue for adaptive, electronics controller-free, soft robots.
Vision-Language Navigation (VLN) in unseen indoor environments is useful in real-world robotics, where an agent must follow natural-language instructions, locate objects, and answer spatial questions without a pre-built map or fixed object vocabulary. Multimodal vision-language models (VLMs) provide strong open-vocabulary grounding and zero-shot reasoning, but struggle to emit reliable metric quantities such as range, bearing, and comparative spatial relations directly from images. Existing approaches address this by folding geometry into hand-engineered pipelines or asking models to output waypoints, requiring changes to the control stack for different robots, tasks, or vocabularies. We introduce AnchorVLN, an open-vocabulary VLN system built on a simple rule: the VLM proposes semantics; geometry decides metrics. It is realised as EMBODIED-NAV-MCP, a Model Context Protocol (MCP) server driven by a VLM agent through a compact set of callable tools. Since no tool accepts distance in metres or bearing in radians, the schema enforces the semantic-geometry boundary without modifying the downstream autonomy stack. We benchmark both tasks of the CMU Vision-Language Navigation Challenge 2026: 30 instruction-following questions over 15 scenes and a frozen 45-question object-reference set. The full system achieves 64.4 percent on instruction following, dropping by 13.3 percentage points without controller modeling (t = 2.77). On object reference, geometric anchoring clears the challenge overlap threshold on 10 of 45 questions, versus 0 of 45 for direct coordinate estimation, reducing median center error from 3.37 m to 2.48 m.
In this paper, we describe the Mission Performance module implemented for a fully autonomous racing car to automatically manage the longitudinal, lateral, and combined performances, aiming to speedup the laptime progression while assuring safety. Motivated by the difficulty and risks of applying the real-time estimation of the grip to critical modules like the motion planner and controller, the Mission Performance guides these modules adapting their target performance instead of changing the vehicle model parameters. The module is formed by pre-defined progressions to warm up the tires at the beginning of a run. Then, the system continuously monitors safety and vehicle dynamics metrics on a per-sector basis to adaptively reduce, maintain, or increase the performance levels for each sector, progressively converging toward the maximum allowed value. The solution's effectiveness is demonstrated on the EAV-25, a fully autonomous Dallara Superformula, at the Yas Marina Circuit during the Abu Dhabi Autonomous Racing League (A2RL) Season 2.
DATAFARM: Distribution-Aligned Task and Motion Planning for Fine-Tuning Vision-Language-Action Models
DATAFARM:面向微调视觉-语言-动作模型的分布对齐任务与运动规划
Sahoo, Samrat, Huang, Yixuan, Silver, Tom
Abstract
Collecting high-quality robot data remains a fundamental challenge for training robot foundation models. Task and motion planning (TAMP) offers a scalable way to generate demonstrations, but our experiments show that raw TAMP trajectories provide surprisingly little benefit when used to fine-tune pretrained vision-language-action (VLA) models, despite successfully solving the target tasks. We hypothesize that this failure arises from a behavioral distribution mismatch between planner-generated trajectories and the data used to pretrain the VLA. To address this mismatch, we introduce DATAFARM: Distribution-Aligned Task And motion planning for Fine-tuning A Robot foundation Model, an approach that incorporates the pretraining distribution directly into TAMP trajectory generation. DATAFARM aligns generated trajectories with the pretraining data in robot joint configurations, motion style, and temporal execution profiles. We evaluate DATAFARM on three tabletop manipulation tasks that TAMP can perform and a cloth-folding task beyond the capability of TAMP. DATAFARM achieves an average success rate of 56.7%, substantially outperforming raw TAMP (8.3%) while approaching human teleoperation (61.7%). On Deformable Object Manipulation, which is outside the fine-tuning distribution, the fine-tuned model retains 85% success, compared with 90% for the pretrained model. These results show that aligning planner-generated demonstrations with the pretraining distribution can make TAMP an effective source of data for VLA fine-tuning. Website and code: https://prpl-group.com/datafarm/
Chinese Translation
收集高质量的机器人数据仍然是训练机器人基础模型的一个根本挑战。任务与运动规划(TAMP)提供了一种可扩展的方式来生成演示,但我们的实验表明,尽管原始TAMP轨迹成功解决了目标任务,但在用于微调预训练的视觉-语言-动作(VLA)模型时,其带来的收益出人意料地小。我们假设这种失败源于规划器生成的轨迹与用于预训练VLA的数据之间的行为分布不匹配。为了解决这种不匹配,我们提出了DATAFARM:分布对齐的任务与运动规划用于微调机器人基础模型(Distribution-Aligned Task And motion planning for Fine-tuning A Robot foundation Model),一种将预训练分布直接纳入TAMP轨迹生成的方法。DATAFARM在机器人关节配置、运动风格和时间执行轮廓方面将生成的轨迹与预训练数据对齐。我们在三个TAMP可以执行的桌面操作任务和一个超出TAMP能力的叠布任务上评估了DATAFARM。DATAFARM达到了56.7%的平均成功率,显著优于原始TAMP(8.3%),并接近人类遥操作(61.7%)。在微调分布之外的可变形物体操作任务上,微调后的模型保留了85%的成功率,而预训练模型为90%。这些结果表明,将规划器生成的演示与预训练分布对齐,可以使TAMP成为VLA微调的有效数据来源。网站和代码:https://prpl-group.com/datafarm/
DWMP: Leveraging Dual World Models for Humanoid Obstacle Traversal
DWMP:利用双世界模型实现人形机器人障碍物穿越
Jin, Rongjun, Ma, Jianming, Gao, Yue
Abstract
Humanoid robots must traverse cluttered obstacle fields using onboard proprioceptive and visual observations, yet existing methods usually process multimodal observations without explicitly considering their different characteristics: proprioceptive observations are low-dimensional but governed by highly nonlinear robot dynamics, while egocentric visual observations are high-dimensional, noisy, and redundant. We propose DWMP (Dual World Model Policy), a framework that provides the actor with separate but complementary world-model representations for humanoid obstacle traversal. A Koopman-based dynamics world model lifts proprioceptive observations into a latent space where their temporal evolution is approximately linear, making the dynamics features easier for the actor to learn from. An RSSM-based visual world model compresses egocentric depth observations into compact stochastic states while preserving obstacle-related geometry. The student policy receives the fused latent representation for action generation, combining linearized proprioceptive dynamics with compressed visual perception. Experiments in simulation and on a Unitree G1 humanoid robot show that DWMP improves obstacle traversal performance over baselines and supports real-world deployment under randomized obstacle layouts.
Socially assistive robots (SARs) can support structured health and well-being interventions, but hardware and cost constraints limit interaction complexity and longitudinal real-world deployments. We present DART: Deployable Architecture for Robot-Mediated Tasks, an architecture that extends SARs through a web application and cloud infrastructure, enabling visual content, user input, remote computation, and persistent data storage synergistically with the robot's physical embodiment, speech, and movement. We evaluated DART by instantiating it in an interatively-developed full-stack HRI system for helping university students with elevated generalized anxiety to complete cognitive behavioral therapy (CBT) homework exercises. The resulting system, which used the low-cost open-source Blossom robot platform, was refined and evaluated through a participatory design process and multiple user studies, and finally evaluated in an in-lab study with 103 participants, and then a six-week in-home deployment with four participants. In the in-lab evaluation, participants showed significant within-session reductions in stress, state anxiety, and negative affect, and gave the platform a mean System Usability Scale score of 78.89. In the home deployment, the mean System Usability Scale score was 87.5, with positive qualitative feedback on usability. Participants across both groups identified speech input, visual presentation, and web-robot synchronization as priorities for improvement. These findings validate DART as an effective architecture for extending the capabilities of a low-cost SAR in both in-lab single-session and in real-world longitudinal deployments.
Autonomous driving requires more than recognizing what is present in a scene: a planner must determine how road structure, surrounding agents, and their motion states should influence a future maneuver. Existing learning-based planners can capture these influences through latent scene features and trajectory decoders, but the relationship between environmental factors and candidate actions often remains implicit. This limits the ability to inspect, diagnose, or refine how scene context affects the safety of a predicted trajectory. Classical safety fields provide an explicit spatial representation of this relationship, but their risk shapes and relative weights are prescribed in advance and do not adapt to each scene. We introduce READ, a framework that learns an explicit, planning-aligned risk representation from complementary geometric and behavioral constraints. READ instantiates this representation as a continuous spatiotemporal field, enabling differentiable queries along candidate trajectories. The learned field connects scene understanding with action selection by encouraging predicted trajectories to align with low-risk regions, while retaining a differentiable interface for trajectory evaluation and refinement. READ integrates with both end-to-end planners and Vision-Language-Action models. Experiments on NAVSIM show consistent gains across matched end-to-end backbones and strong performance in a VLA setting; READ also achieves competitive results on NAVSIM v2. These results establish learned spatial risk as an explicit, adaptable representation for safe planning.
Understanding Whole-Body Robot Teleoperation Strategies Under Diverse Task Objectives and Constraints
理解多样任务目标与约束下的全身机器人遥操作策略
Lin, Tsung-Chi, Chen, Juo-Tung, Huang, Chien-Ming
Abstract
This work investigates the control strategies of complex whole-body robot teleoperation that coordinate active perception, bimanual manipulation, and navigation. We developed a hybrid control framework, combining the free-form and constrained control, for the whole-body teleoperation of the TIAGo mobile manipulator. We conducted a user study to explore people's control strategies under different task constraints such as limited time and low tolerance of errors. Our results highlight the effective use of coordinated control in improving task efficiency and reducing the risk of reaching individual joint limits. We discuss our results and their implications for designing future whole-body robot teleoperation systems.
FoldNet++: a Large-Scale Synthetic Dataset for Robotic T-Shirt Folding and Unfolding
FoldNet++:一个用于机器人T恤折叠与展开的大规模合成数据集
Chen, Yuxing, Wei, Zhiyuan, Xiao, Bowen, Zhang, Zhizheng, Wang, He
Abstract
Due to the highly deformable nature of garments, training a generalizable policy for robotic T-shirt folding and unfolding remains a significant challenge. In this work, we present a large-scale synthetic dataset for robotic T-shirt folding and unfolding, covering 6 robotic embodiments, 1K T-shirts, 1K environmental assets, and 120K episodes with rich annotations, which can be used to train a wide range of manipulation policies. We first follow the FoldNet pipeline to generate a large-scale dataset of physically simulatable T-shirts with diverse appearances and annotated semantic keypoints. Based on these semantic keypoints, we then generate manipulation demonstrations for different robotic embodiments through a unified rule-based framework. We use these demonstrations to train visuomotor policies, and experimental results demonstrate that models trained solely on our synthetic data can achieve over 90\% end-to-end task success rates when directly deployed to unseen real-world environments and previously unseen T-shirts from arbitrary initial configurations. Project URL: https://pku-epic.github.io/FoldNetXX/.
IMPLY: Physically Anchored Consistency for World-Model Rollouts
IMPLY:世界模型推演的物理锚定一致性
Mehta, Aman, Baviskar, Riya
Abstract
A world model asked what happens if an object is pushed at several speeds produces several futures. If the model has the object in mind, those futures agree about it: each implies the same mass and friction. The consistency checks now used to vet world-action models ask whether a model's futures agree with each other, and none of them knows any physics. We show that this is not enough, and what to do instead. IMPLY reads the physics each rollout implies by inverting a simulator and scores a set of rollouts by how well one object explains all of them, anchored to two calibration pushes the model has observed. In a controlled setting, self-consistency gives a perfect score to a model that ignores the object and always predicts a typical push; anchoring exposes it (AUROC 0.70 versus 1.00). On a real model, V-JEPA 2-AC adapted to the scene, the same thing happens. Given its own calibration pushes the model tracks the object (per-object correlation with the truth 0.91); given another object's, it does not (0.05). Self-consistency cannot tell these apart, preferring the right evidence on 52% of objects, chance level, while anchored disagreement prefers it on 73% and correlates 0.92-0.99 with the rollouts' error. Used to choose among candidate rollout sets, it comes within 0.003 of an oracle that sees the truth. A model that has internalised the wrong object is exactly as self-consistent as one that has internalised the right one; consistency has to be anchored to evidence.
PATH: Continuous Target Sensing among Autonomous Cooperative Drones
PATH:自主协同无人机间的连续目标感知
Kim, Heegyeong, James, Alice, Seth, Avishkar, Kuantama, Endrowednes, Williamson, Jane, Feng, Yimeng, Han, Richard
Abstract
Continuous target sensing by uncrewed aerial vehicles (UAVs) is constrained by limited flight endurance, motivating the transfer of tracking responsibility between cooperating UAVs. Such a handoff requires the receiver to identify the same physical target currently tracked by the sender despite differences in viewpoint, scale, and target appearance. Existing approaches based on global target localization or appearance-based cross-view association are limited by positioning uncertainty or ambiguous visual features. This paper presents Perspective Alignment \& Tracking Handoff (\textbf{PATH}), a platform-agnostic, geometry-assisted sensing and verification framework for target handoff between two moving UAVs. The sender reconstructs the tracked target as a metric 3D point using RGB-D sensing, while the receiver estimates its relative pose from a fiducial observation and projects the transmitted target point into its own image as a spatial prior for target acquisition. The receiver-generated candidate is then returned to the sender and verified through a cross-view Mutual Agreement Handshake before tracking responsibility is transferred. Real-world UAV experiments show mean relative-position and target-position errors of 0.047~m and 0.030~m, respectively. Under visually ambiguous conditions, PATH achieves 96.0\% frame-level receiver-side target acquisition accuracy, with 2.0\% false-positive and 2.0\% false-negative rates. A sensor-error sensitivity analysis shows that relative-pose uncertainty is the dominant contributor to receiver-view projection error. The implementation operates at video rate with compact inter-UAV communication below 16~kB/s at 60~Hz, demonstrating the feasibility of lightweight geometry-assisted target handoff on resource-constrained UAV platforms.
Category-level in-hand manipulation of articulated objects is a formidable yet underexplored challenge for dexterous robotic hands. This difficulty stems from two core bottlenecks: first, controlling an object's internal degrees of freedom is tightly coupled with maintaining grasp stability on a free-floating base; second, acquiring diverse object models and functional grasps at scale is highly labor-intensive, yet vital for generalization given the system's sensitivity to initial configurations. In this work, we present ArtManip, the first category-level articulated in-hand manipulation method that generalizes across object instances and diverse initial grasps. For initial configuration construction, we develop an automated pipeline that procedurally generates diverse articulated objects and synthesizes task-oriented functional grasps. For policy learning, we propose a robust two-stage training strategy that incorporates articulation physics randomization, reward curriculum, and latent representation distillation to handle complex contact and joint dynamics during deployment. Extensive experiments across four object categories demonstrate that our policy generalizes to unseen instances and varied configurations in simulation, and achieves zero-shot transfer to 12 real-world objects featuring diverse shapes and joint mechanics.
Communication-Constrained Multi-Robot Exploration With Adaptive Communication Windows
基于自适应通信窗口的通信受限多机器人探索
Rossano, Ben, Lim, Jaein, How, Jonathan P.
Abstract
Exploring unknown environments with multi-robot teams can improve efficiency by allowing robots to explore in parallel. However, realizing these gains requires effective information sharing. When communication is intermittent, robots must balance the benefits of sharing information against the cost of diverting from exploration to establish communication. This paper introduces MACE, a decentralized exploration framework that actively evaluates whether establishing communication is worthwhile. At scheduled communication windows, robots estimate the cost of reaching previously identified communication locations. By formulating this decision as a variant of the Vehicle Orienteering Problem, robots evaluate routes based on the travel required to establish communication and the exploration that can be completed along the way. This approach enables robots to communicate more frequently than under purely opportunistic strategies while reducing the unnecessary travel associated with fixed rendezvous strategies. Across a set of simulated environments with varying size and geometry, we demonstrate that MACE reduces the total exploration time by up to 23% compared to existing communication-constrained exploration strategies.
Autonomous precision milling of biological structures is challenged by incomplete knowledge of target geometry, local material thickness, and critical internal boundaries. Subject-specific preoperative models can address geometric and thickness variations, but static models cannot determine boundary status encountered during execution, while repeated target-specific imaging limits scalability. This article presents an uncertainty-aware autonomous milling framework that assigns complementary roles to generic anatomical priors and active boundary perception. A generic anatomical prior provides conservative global guidance and is transformed through semantic-guided registration and hybrid vision-force calibration into robot-executable guidance for individual targets. As milling approaches uncertain boundaries, the robot actively probes the remaining structure and uses relative stiffness changes to estimate boundary status and structural detachability. A state-adaptive controller governs transitions between active perception and spatially selective incremental refinement, repeating this cycle until the termination criterion is satisfied. Hierarchical experiments on biological surrogates and in vivo mouse cranial window creation demonstrate accurate anatomical prior transfer, reliable boundary adaptation, and autonomous precision milling of biological structures.
Dexterous manipulation requires coordinated multi-finger control and effective tactile feedback, yet learning these capabilities remains challenging due to the lack of large-scale real-world data and the difficulty of extracting effective representations from sparse tactile signals. We build a robot platform and teleoperation system to collect a 200-hour bimanual dexterous manipulation dataset with synchronized visual, tactile, and language annotations, comprising 10,576 trajectories across 65 tasks, 69.5% of which involve dexterous multi-finger manipulation. We further propose STAR, an integrated training recipe for vision-tactile-language-action (VTLA) models that addresses the spatial, temporal, and informational sparsity of tactile signals through visual-tactile joint pre-training, sparse-global tactile token representation, and sparse future tactile prediction. Trained on this dataset, STAR achieves a 61% average success rate across four real-world tasks with 100 post-training trajectories per task, demonstrating dexterous performance under task-specific post-training.
This paper presents an online coverage path planning algorithm for unknown environments. During navigation, the initially unknown search area is progressively decomposed into disconnected subareas as new obstacle information is acquired and coverage proceeds. These subareas are organized in an incrementally constructed decomposition tree that preserves their hierarchical parent-child relationships. Based on this tree, a global coverage tour is maintained and updated online by prioritizing newly generated child subareas according to their exploration states and distances from the robot. A local planner then generates coverage motions within each selected subarea, allowing the robot to adapt its trajectory as the environment is gradually revealed. Its performance is evaluated via high-fidelity simulations in complex scenarios. The results show improved coverage efficiency in terms of path length and overlap ratio in comparison to three baseline algorithms.
Quantifying Spectral Differences in Vehicle Between Production Autonomous and Human-Driven Vehicles Across Driving Scenarios
量化不同驾驶场景下量产自动驾驶车辆与人类驾驶车辆之间的车辆频谱差异
Fang, Peiyi, Li, Xiangyu, Weng, Yonglin, Ma, Ke
Abstract
Differences in vehicle kinematic characteristics between production autonomous vehicles (PAVs) and human-driven vehicles (HVs) have been limitedly investigated by empirical studies. Most recent studies rely on simulation-based models, while some further investigate low-level adaptive cruise control (ACC) systems in controlled experiments. These methods commonly adapt some time-domain metrics to characterize PAV-HV differences across limited driving conditions. However, current PAVs equipped with high-level autonomous driving systems generate driving behaviors in a black box using data-driven models. These fundamentally different mechanisms for generating behaviors may produce distinct kinematic characteristics in traffic. More importantly, these time-domain metrics cannot reflect frequency-related traffic dynamics across different driving scenarios. Thus, this study adapted a real-world PAV dataset with four PAV platforms and developed a frequency-domain framework to quantify kinematic differences between PAVs and HVs across diverse driving scenarios, including varying driving states, lighting, weather, and vehicle densities. The framework transforms kinematic signals into the frequency domain and extracts spectral features, and then compares these features between PAVs and HVs based on kernel density estimation and Wasserstein distance. The results reveal clear scenario-dependent PAV-HV spectral differences. Specifically, speed-related differences were consistently smaller during car-following than cruising, while rainy conditions consistently enlarged acceleration-related differences compared with clear conditions. These findings highlight the necessity of multi-scenario evaluations and demonstrate the value of frequency-domain analysis for characterizing PAV-HV kinematic differences under real-world conditions.
A 3D scene graph groups objects into rooms. When a robot is asked to fetch an object from the kitchen, that grouping is what tells it where to look. An object recorded in the wrong room is not retrievable by a query naming the correct room. We introduce Progressive Boundary Closure, which recovers room layer from a monocular RGB video. A SLAM front end and an open-vocabulary segmenter supply a structural point cloud, camera trajectory and object tracks. The cloud is rasterised into a top-down map, rooms are recovered from it, and each object takes the room holding most of its extent. The difficulty lies in the map itself. Walls are recorded only where the camera looked, so a gap in the boundary may be a doorway or a stretch of wall that was never observed; nothing distinguishes the two. Prior methods treat both as passages, merging rooms that should remain separate. We observe that both require the same treatment: a room should not extend across either, so both are closed and need not be distinguished. Such an opening closes under a small amount of boundary growth, and few sightlines cross it, so points in different rooms rarely see one another. We use the first to recover rooms and the second to assign objects to them. Rooms are obtained by Progressively thickening the boundary inward and freezing each free-space region once it becomes enclosed, so every opening seals at its own scale rather than at a radius fixed in advance. Camera poses are used as seeds, which removes the sampling heuristic and makes the segmentation deterministic. Over 10 floors of 6 HM3D-Semantics scenes, scored against HOV-SG on identical top-down maps, we recover 74 rooms for 72 annotated regions (HOV-SG: 44), raising room F_1 from 0.741 to 0.890 at IoU 0.25 at some cost in precision, and object-to-room ARI from 0.488 to 0.696 (p=0.002, ahead on every floor).
Shape control of deformable linear objects (DLOs) is challenging for imitation learning because deformation behavior varies with material properties such as stiffness and elasticity, so a single policy must generate different action sequences for different objects even when the goal shape is identical. We propose a diffusion policy conditioned on material labels that are estimated online during manipulation. A recurrent estimation network predicts the material label of the grasped object from the time series of multi-view images and robot joint states, and the predicted label conditions the diffusion policy at every inference step. We collected 480 real-robot demonstrations covering four DLO materials and three groove-placement tasks, and compared per-material specialist policies, a task-conditioned policy without material labels, a policy conditioned on ground-truth material labels, and the proposed policy. Conditioning on ground-truth material labels improved the average success rate from 45.8% to 60.0% over the task-only policy, and the proposed policy reached 60.8% without any prior material information, matching the policy given ground-truth labels. A post-hoc analysis shows that the estimator extracts material-related information from the manipulation observations and that the diffusion policy responds to the resulting conditioning signal, while the one pronounced failure case is associated with persistent confusion between two similar materials.
Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.
Driving Context-guided Model Predictive Planning and Control for Autonomous Car Racing at the Limit and Beyond
面向极限及超越极限自动驾驶赛车的驾驶情境引导模型预测规划与控制
Raji, Ayoub, Sacco, Federico, Musiu, Nicola, Bertogna, Marko
Abstract
This paper presents a Model Predictive Control-based motion planning and control pipeline for autonomous car racing capable of adapting to different driving contexts, such as overtaking, nominal driving, and countersteering. A Cost Blending state machine manages the identification of different driving contexts and the selection of their predefined weights to be applied to the Model Predictive Planning (MPP) and Control (MPC) modules. The two optimization-based solutions share the same problem formulation and model, differing only in horizon length, rate, tuning, and in their open-loop versus closed-loop approach to maximize the effectiveness of their interaction. The work is validated on the fully autonomous open-wheel racecar Superformula EAV-25, with a lap time achieved that is within 2% of the best human driver reference. The results demonstrate the capability of the solution in driving at the limit of handling, smoothly executing overtaking maneuvers, and quickly reacting to high oversteering conditions to recover the vehicle stability.
Earthmoving tasks such as excavation, backfilling, or embankment construction require deliberate repositioning of deformable soil. For these tasks, human operators use all shovel faces, while autonomous systems so far are limited to excavation and dumping. Current methods often rely on heuristic models but do not incorporate soil mechanics. We address this shortcoming by using Reinforcement Learning in a GPU-parallelized Material Point Method particle simulation. Our controllers are conditioned on material state such as shape and compactness, enabling skills that use multiple contact faces of the tool and displace material both inside and outside of the shovel. To use the same learned weights across machines, our policies operate in a normalized end-effector space and are deployed through a calibrated machine interface. We evaluate this calibrated transfer on an 11.5t hydraulic excavator and a 500g tabletop robot. We validate performance through autonomous construction of a 42m long, 2.1m high embankment in 45min, executing 201 individual policy strokes without failure, retry, or operator intervention. In a direct comparison, the autonomous controller matches an expert operator's progression speed and produces a higher, more consistent embankment. Additional qualitative backfilling and compaction experiments demonstrate the material-state awareness and calibrated transfer across machines.
Chinese Translation
土方工程任务,如挖掘、回填或筑堤,需要对可变形土壤进行有意的重新定位。对于这些任务,人类操作员使用铲斗的所有面,而迄今为止的自主系统仅限于挖掘和倾倒。当前方法通常依赖于启发式模型,但并未纳入土壤力学。我们通过在GPU并行化的物质点法(Material Point Method)粒子模拟中使用强化学习来解决这一不足。我们的控制器以材料状态(如形状和密实度)为条件,使得技能能够利用工具的多个接触面,并在铲斗内外移动材料。为了在不同机器上使用相同的学习权重,我们的策略在归一化的末端执行器空间中运行,并通过校准的机器接口进行部署。我们在11.5吨液压挖掘机和500克桌面机器人上评估了这种校准迁移。我们通过自主建造一条42米长、2.1米高的堤坝,在45分钟内执行了201次单独的策略行程,无失败、重试或操作员干预,从而验证了性能。在直接比较中,自主控制器匹配了专家操作员的推进速度,并产生了更高、更一致的堤坝。额外的定性回填和压实实验展示了材料状态感知以及跨机器的校准迁移。
Applying an imitation learning policy to a new manipulation task usually requires collecting new demonstrations and retraining the model, which makes sample efficiency a practical concern. Pretraining on large-scale robot datasets is effective in this respect, but such datasets are costly to collect and train on, while data augmentation techniques typically require a new round of data generation and retraining for each task. A complementary question is what useful prior can be provided to a policy at negligible cost before any task-specific data are collected. In this study, we construct a geometric visual pretraining dataset in which each scene contains only a plane, an object, and a hand, and trajectories are generated automatically. The scenes contain neither textures nor backgrounds; pretraining primarily exposes the policy to the geometric relationship between the hand and the object. Furthermore, representing the hand as a cube avoids tailoring the dataset to a specific robot morphology. We evaluate this geometric prior using ACT on three simulated robots across five manipulation tasks each, as well as on three real-world robot tasks. Across many of these robot--task combinations, fine-tuning from the geometric prior achieves higher success rates in the early stages of training than training from scratch while using only a small number of task demonstrations. These results suggest that even highly simplified geometric scenes can provide a useful initialization that transfers across robots and to real-world tasks when task data are limited.
Control Architecture for Safe Grasping of Fragile Objects Using a Coarse Position-Controlled Gripper
基于粗位置控制夹爪的易碎物体安全抓取控制架构
Pavlic, Marko, Geier, Moritz, Markert, Timo, Burschka, Darius
Abstract
Robots are increasingly used in unstructured environments. The need for them to safely grasp unknown objects without damaging them becomes crucial. Humans achieve this by sensing and quickly responding by adjusting their grasping force. Similarly, effective grasp acquisition in robots requires compliant interaction strategies that can adapt to uncertain object properties and adjust to any instabilities during manipulation. We present a geometry-aware force/torque-based contact estimation method for a coarse position-controlled gripper, combined with an adaptive admittance controller for safe grasp acquisition. The desired contact forces are estimated online to keep stable contact with objects of unknown properties. This enables compliant and stable grasps while avoiding excessive forces. Experiments with objects of different sizes, shapes, stiffnesses, and weights show that the proposed algorithm not only prevents slippage but also applies minimal force to safely grasp an object without causing excessive deformation.
We present a custom high-fidelity vehicle dynamics simulation environment for testing and validation of Autonomous Racing software. The digital twin of the autonomous vehicle is developed in Dymola, using racecar dynamics modeling libraries to build a complete multi-body model. A 3D road surface, including elevation profiles and curbs, is implemented using the Curved Regular Grid (CRG) standard. The model is exported from Dymola as a Functional Mock-up Unit (FMU) and integrated into a custom software-in-the-loop simulator, where communication interfaces with the autonomous racing stack were developed in C++. A calibration procedure based on experimental data is also presented, along with a validation study to further support the quality of the proposed framework. The simulator runs in real time on a portable computer and provides reliable ground truth for algorithms validation prior to real-world deployment.
Chinese Translation
我们提出了一种定制的高保真车辆动力学仿真环境,用于自动驾驶赛车软件的测试与验证。自动驾驶车辆的数字孪生是在Dymola中开发的,使用赛车动力学建模库构建完整的多体模型。使用弯曲规则网格(Curved Regular Grid, CRG)标准实现了包含高程剖面和路缘的3D路面。该模型从Dymola导出为功能样机单元(Functional Mock-up Unit, FMU),并集成到定制的软件在环仿真器中,其中与自动驾驶赛车软件栈的通信接口是用C++开发的。还提出了基于实验数据的校准程序,以及一项验证研究,以进一步支持所提出框架的质量。该仿真器在便携式计算机上实时运行,并为实际部署前的算法验证提供可靠的真值。
VertexCBF: Improving Neural Control Barrier Functions via Vertex-Restricted Control Search
VertexCBF:通过顶点受限控制搜索改进神经控制屏障函数
Derajić, Bojan, Bernhard, Sebastian, Hönig, Wolfgang
Abstract
As the number of autonomous robots continues to grow, safety becomes increasingly important. Control barrier functions (CBFs) provide a theoretically grounded framework for ensuring safety, but existing design methods often face limitations in effectiveness, scalability, or interpretability, and may result in overly conservative safe sets. In this paper, we propose \emph{VertexCBF}, a framework for learning neural CBFs in a scalable, systematic, and explainable way. We approximate the stationary Hamilton--Jacobi value function using a neural network trained via a combination of physics-informed and sparsely supervised learning. By exploiting control-affine dynamics and a convex polytope control set, under which the Hamiltonian is maximized at the control vertices, we efficiently generate supervision points via GPU-parallel vertex-restricted tree search, while a residual architecture guarantees that the learned CBF is never larger than the specified constraint function. We evaluate the method on 15 systems and compare it against relevant baselines, showing that it reliably recovers large safe sets where the baselines are conservative or fail completely. In addition, we perform a hardware experiment in which a mobile robot safely avoids pedestrians using a neural CBF trained with our method.
Parameter Sensitivity Analysis for Aerial LiDAR-Inertial Odometries in low-altitude flights
低空飞行中空中 LiDAR-惯性里程计的参数敏感性分析
Milijas, Robert, Dios, Jose Ramiro Martinez-de, Bogdan, Stjepan
Abstract
LiDAR-based SLAM (Simultaneous Localization and Mapping) and LIO (LiDAR-inertial odometry) algorithms are often used for precise navigation of unmanned aerial vehicles, especially during interactions with the aerial robot's environment. However, the performance of these algorithms is greatly dependent on the scenario, LiDAR, and robot motion characteristics, often requiring an intensive tuning process to achieve the desired performance. To aid these tuning efforts, this paper analyzes the influence on performance of the parameters of an EKF-based LIO algorithm (FAST-LIO2) and the LIO module of a graph-based SLAM algorithm (Cartographer) on aerial LiDAR SLAM datasets recorded using different LiDARs in low-to-moderate-altitude flights in diverse environments. The analysis is conducted on the absolute trajectory error (ATE) resulting from processing the datasets with the LIO algorithms configured with each combination of parameters obtained in an exhaustive grid search. The relationship between individual parameters and the ATE results is assessed using Pearson's correlation, while the influence of each parameter is assessed using random forest permutation importance analyses with random forest models trained to predict the resulting ATE values based on the choice of parameters. The performed analysis obtains for Cartographer and FAST-LIO2: i) the identification of parameters with stronger influence in performance, ii) a simplified tuning procedure, and iii) tuning recommendations. Using the proposed tuning recommendations, both algorithms obtain on the analyzed datasets ATE values within 5 cm to the optimal performance found in the grid search procedure in 94% of the analyzed cases.
Chinese Translation
基于 LiDAR 的 SLAM(同步定位与建图)和 LIO(LiDAR-惯性里程计)算法通常用于无人机的精确导航,尤其是在与空中机器人环境交互期间。然而,这些算法的性能在很大程度上取决于场景、LiDAR 和机器人运动特性,通常需要大量的调参过程才能达到所需性能。为了辅助这些调参工作,本文分析了基于 EKF 的 LIO 算法(FAST-LIO2)和基于图的 SLAM 算法(Cartographer)的 LIO 模块的参数对性能的影响,这些分析是在不同环境下使用不同 LiDAR 在低至中等高度飞行中记录的空中 LiDAR SLAM 数据集上进行的。该分析基于绝对轨迹误差(ATE)进行,该误差是通过使用在穷举网格搜索中获得的每种参数组合配置的 LIO 算法处理数据集而产生的。使用 Pearson 相关性评估单个参数与 ATE 结果之间的关系,而每个参数的影响则通过随机森林排列重要性分析进行评估,其中随机森林模型经过训练以根据参数选择预测所得的 ATE 值。所进行的分析为 Cartographer 和 FAST-LIO2 获得了:i)识别出对性能影响更强的参数,ii)简化的调参流程,以及 iii)调参建议。使用所提出的调参建议,两种算法在分析的数据集上,在 94% 的分析案例中,获得了与网格搜索过程中找到的最优性能相差 5 cm 以内的 ATE 值。
Datasets are a crucial element in the development of perception algorithms. They relate sensor measurement data to annotated reference information and allow for the deduction of sensor and object characteristics. In autonomous driving, the reference data commonly consist of semantic image segmentation, point-wise associations, or bounding box annotations. The dataset proposed in this work, however, aims to dig deeper into the evaluation of measurement principles and provides scanned 3D models of all vehicles together with a pose and continuous kinematics reference obtained by RTK-GNSS. Combined, the state of the complete dynamic surrounding of the sensor vehicle is known for any point in time. Subsequent reference formats can be easily computed in user-defined granularity. This dataset involves single-object and multi-object recordings with seven target vehicles. In particular, measurement effects such as occlusion, as well as reflections, can be evaluated, as the normals of the shape of the target vehicles are known. We describe the dataset, discuss the technical background of its development, and briefly present exemplary evaluations.
From Transportation to Manipulation: Enabling Grasping in Magnetic Robotics
从运输到操作:在磁机器人中实现抓取
Bergmann, Lara, Greis, Noah, Grothues, Cedric, Weigelt, Lisa-Marie, Neumann, Klaus
Abstract
Magnetic levitation (MagLev) systems have great potential for application in high-mix, low-volume manufacturing due to their scalability and flexibility, enabling highly reconfigurable in-machine material flow. However, their manipulation capabilities remain largely unexploited, as current applications almost exclusively focus on transportation. To enable grasping and manipulation directly on MagLev systems without requiring additional costly handling equipment, such as industrial robot arms, we present the Gripper MagBot, a low-cost parallel 6-DoF manipulator with an integrated 1-DoF gripper that mechanically couples three MagLev movers. The Gripper MagBot supports two operating configurations: a default mode and a single-track mode, selectable depending on the required stability and workspace footprint. To reconfigure a machine, the MagBot can be autonomously dropped off and picked up using a docking station. We showcase pick-and-place examples in simulation, as well as with the real Gripper MagBot using our inverse kinematics controller. CAD files, assembly instructions, a component list, and videos are available at https://sites.google.com/view/gripper-magbot.
Before the Tipping Point: Force-Guided Active Perception for Shape-Agnostic Estimation of 3D Centers of Mass
在倾覆点之前:力引导主动感知用于形状无关的3D质心估计
Hyland, Steven M., Xiao, Jing, Onal, Cagdas D.
Abstract
Estimating the 3D center of mass of unknown objects is challenging when grasping is infeasible, geometry is irregular, or mass distribution is uneven. We present a force-based method that estimates CoM height and mass from a single sub-critical tipping experiment by a robot manipulator. The robot applies a quasistatic elevated push and retract motion, using force-angle measurements recorded during tipping to identify parameters from the object trajectory. Our proposed push-retract cycle mitigates frictional bias, enabling generalized fitting. We experimentally validate our method using a robot manipulator with a six-axis force torque sensor on varying types of objects without prior shape information and without specific models. We also propose a method to prevent toppling, keeping the object in a sub-critical tipping regime by leveraging a safety margin. In experimental studies, our method recovers mass, CoM height, and toppling angle with relative errors below 5.0 percent across all unknown objects. This work demonstrates reliable 3D inertial parameter estimation under proper safety thresholds in tipping. Our proposed method informs and enables reliable non-prehensile manipulation and robotic grasping of challenging objects that were previously infeasible.
Robust Underwater Grasping of Sloped Objects with a Waterproof Passive Adaptive Gripper
使用防水被动自适应夹爪的倾斜物体鲁棒水下抓取
Hong, Jooyoung, Hong, Daewon, Kim, Joohyung
Abstract
Robust grasping of everyday objects remains challenging for parallel-jaw grippers, particularly when handling sloped or asymmetric items that induce torque-driven rolling and shear slip. These challenges become even more severe in domestic environments such as kitchens, where objects are often wet or submerged, drastically reducing friction between the gripper and the object. To address these issues, we present a waterproof passive adaptive gripper that combines local and global adaptability for high-performance grasping in the submerged environments. The proposed gripper features a fully waterproof design that integrates a passive rotational joint for global self-alignment on sloped surfaces and passive variable-stiffness pads for local surface adaptation. Experiments in both dry and underwater conditions on various cylindrical and conical objects demonstrate superior holding capability and grasp robustness compared to rigid-pad and fixed-joint baselines. The proposed design offers a practical and robust solution that successfully enables stable underwater manipulation of diverse everyday objects, effectively addressing a critical gap in current underwater robotic grasping for daily-life applications.
ARC: Autonomous Robotics Compliance A Three-Layer Governance Architecture for Deployed Autonomous Systems
ARC:自主机器人合规——面向已部署自主系统的三层治理架构
Eide, Tord, Holt, Einar
Abstract
Proposed governance framework for autonomous robotic systems, introducing a three-layer compliance architecture (ARC) instantiated through model safety validation, cognitive certification benchmarks, and operational authorization standards.
A Robot Among People:From Social Imitation to the Social Becoming of Human Groups
人群中机器人:从社会模仿到人类群体的社会性生成
Pham, Victor Tuan Vu, Dörrenbächer, Judith, Weisswange, Thomas H., Hassenzahl, Marc
Abstract
Robots designed to mediate human groups often fall into the solutionist trap: they are framed as sociable agents that fix problems such as conflict, disengagement, or lack of coordination. We suggest a different way of thinking. Rather than discrete agents, robots can be understood as situated elements of shared environments; catalysts and carriers of group experience whose meaning emerges through how people position, interpret, and interact with them. From this perspective, robots are not there to repair some ostensible dysfunctionality, but to enable group-level sense-making around care, norms, and identity. Our prior work on robotic street furniture suggests that this does not happen by imitating human sociality but by taking the shape of deliberately constrained, group-facing entities that happen and act for \textit{us} without being socially entangled as one of us. We thus understand robots in public spaces not in terms of autonomy or intelligence, but as a relational capacity. This implies designing robots not in our image or for our utility, but grounded in our needs in being and becoming together.
While offering significant promise for diverse applications, pattern-oriented swarms encounter multifaceted challenges in geometric control, self-organization, and safe navigation through dynamic environments. In this paper, we present a GRF-based stochastic optimal control framework to address these challenges within a unified probabilistic architecture. By extending the GRF into the temporal domain, the proposed framework casts collective coordination as a Bayesian inference task, enabling swarms to accommodate environmental uncertainty, satisfy non-convex constraints, and reconcile heterogeneous dynamics across diverse platforms. We develop an uncertainty- and safety-aware collision avoidance module for navigation in the presence of stochastic obstacle motion. The unscented transform is employed to propagate state uncertainty for both dynamic obstacles and neighboring agents, yielding principled confidence bounds for collision avoidance. In addition, density-guided pattern control is introduced, which encodes geometric patterns as implicit density fields. This representation decouples pattern specification from explicit agent-to-target assignments, thereby facilitating intrinsic self-healing and elastic reconfiguration in a distributed manner. The proposed framework is extensively evaluated through Monte Carlo simulations across diverse scenarios. Its model-agnostic nature is demonstrated on both quadrotor and fixed-wing UAV swarms, highlighting its generalizability across platforms with heterogeneous dynamics. Finally, the efficacy and robustness of the proposed method are validated through indoor experiments with a 15-quadrotor swarm and outdoor deployments involving 4 custom-built autonomous quadrotors. These experiments substantiate the proposed framework's capacity to maintain reliable geometric pattern transitions and safety-aware navigation within real-world environments.
Robots are increasingly used in diverse application areas, where autonomous navigation plays a central role. As these systems become more widespread, improving their energy efficiency is critical to extending operational time and reducing environmental impact. The Robot Operating System (ROS) is a widely adopted middleware for robotics, offering a rich set of configurable packages. However, this flexibility can result in suboptimal software configurations in dynamic environments, negatively affecting both performance and energy consumption. This paper investigates the impact of ROS 2 package reconfigurations on the energy efficiency of mobile robot navigation. We conduct a controlled experiment in two warehouse-like scenarios (small and large) with varying obstacle layouts and Costmap 2D configurations (essential to the Nav2 stack). Through repeated trials, we measure energy usage, power profile, CPU load, memory consumption, and navigation performance. Results show that configurations must be carefully chosen for the specific robotic environment, and we were able to identify critical settings that lead to good and poor performance and energy consumption.
Data-driven driving simulators command accelerations and steering rates from a fixed grid without constraining the realized accelerations and jerks. As a result, reinforcement-learning policies inflate safety metrics through abrupt, last-second maneuvers that lie far outside the range of human driving and would be unacceptable to occupants of a real vehicle, so the metrics measure simulator permissiveness rather than policy quality. Enforcing comfort bounds naively is not enough: lateral limits shrink quadratically with speed, so clamping a static grid saturates it and destroys fine-grained control ("grid collapse"). We propose an adaptive action parameterization that rediscretizes the grid at every step to span exactly the per-step feasible control set, via closed-form inversion of the lateral-jerk constraint. We further present PufferDrive-Editor, a browser-based tool to audit realized kinematics and author kinematically challenging scenes. On the Waymo Open Motion Dataset and a hand-authored slalom, our adaptive model holds comfort violations below 1% while outperforming clipped-grid and direct-jerk baselines in navigability.
This work enhances global path planning via a pure-pursuit controller with multi-model kinematic switching that sustains plan fidelity across diverse terrains. The system includes a traversability graph for terrain analysis, a Heading-Aware A* algorithm for generating feasible paths, and a multi-model Pure Pursuit controller for dynamic tracking. A core innovation is adaptive kinematic modeling, enabling real-time switching between kinematic models based on terrain features and robot states. This adaptability optimizes path efficiency and energy use in challenging scenarios. We validate the approach in simulation on different platforms, namely the Artaban quadruped and the X3 quadrotor drone, showcasing improved performance, robustness, and adaptability over standard baselines.
Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
Dynin-Robotics:全模态统一扩散视觉-语言-动作模型
Lee, Hoeun, Kim, Jaeik, Oh, Jusang, Kim, Jinhyeok, Choi, Geon, Kim, Hyeonggeun, Do, Jaeyoung
Abstract
Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation and selection through a shared trajectory model. Dynin-Robotics implements this formulation on Dynin-Omni, an omnimodal masked-diffusion backbone, representing language, visual observations, goals, and actions as discrete tokens. By varying conditioning and target spans, the same model learns action prediction, action-conditioned next-observation prediction, terminal goal-state prediction, and trajectory-to-instruction reconstruction. These interfaces support test-time scaling through goal prediction, action-candidate evaluation, and joint refinement of action and future-state predictions. We continually pretrain the model on approximately 1.33 million trajectories from 48 Open X-Embodiment datasets and adapt it separately to downstream domains. On two VLABench tasks, robot pretraining improves adaptation within a fixed Stage-2 step budget, and the full objective mixture improves shifted-instruction success over Policy-only post-training under the same coupled decoder. Combining goal guidance with joint action-next-state denoising further improves shifted-instruction success over action-only decoding; the benefit depends on how the predictions are composed. Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot. An optimized block-parallel implementation accelerates model-side action decoding by up to 29.2x relative to the base implementation under the reported profiling setup. These results support shared trajectory modeling as a common interface for learning complementary robot objectives and composing their predictions during control.
ASTRIL-MPC: Autonomous Traversal Framework of Articulated Tracked Robots with Language-Guided Neural-Kinematic MPC
ASTRIL-MPC:基于语言引导神经运动学模型预测控制的关节式履带机器人自主通行框架
Gan, Zhenfeng, Chen, Yanbo, Che, Lirong, Tan, Junbo, Wang, Xueqian
Abstract
In urban search and rescue, articulated tracked robots (ATRs) must traverse structured but contact-rich environments such as stairwells and cluttered building interiors. Reliable autonomy remains challenging because robot-terrain interaction (RTI) is hybrid and discontinuous, and effective flipper-track coordination is difficult to model analytically. We present ASTRIL-MPC, a language-guided neural kinematics model predictive control (MPC) framework for autonomous traversal. A learned kinematics model predicts short-horizon task-state increments from a height sequence and recent trajectories; NMPC plans with multi-objective costs and strict feasibility constraints; and a large language model (LLM) proposes bounded updates to selected weights and bounds through a safety-checked interface with range clipping, rate limiting, and consistency checks. The compiled predictor enables a full control cycle within 100 ms. Across three traversal tasks and a multi-height generalization setting, ASTRIL-MPC improves an aggregate traversal-quality score by up to 71% over a non-adaptive NMPC and by 67% over a PPO baseline, while eliminating measurable collision impacts during descent. These results indicate that combining learned kinematics, optimization-based planning, and language-guided retuning yields data-efficient and robust autonomy for articulated tracked robots.