Autonomous vehicles are expected to improve road safety and efficiency, but passengers often remain uncertain about what the vehicle perceives and why it acts as it does. Virtual reality (VR) offers a safe and repeatable medium for presenting this information, yet most passenger-facing VR studies rely on fully simulated vehicles or pre-scripted scenarios, so the motion and perception shown to the user do not originate from a physically operating autonomous system. This paper presents a video-augmented VR framework that couples a physical ROS 2 autonomous robot vehicle to a Unity 6 application deployed on a Meta Quest 3S headset. The vehicle state and live onboard camera stream are transmitted over two independent communication channels, allowing the virtual vehicle to mirror the physical robot's motion while the passenger simultaneously views the vehicle's first-person camera feed and its navigation decisions through an in-vehicle dashboard interface. We evaluate the framework over 20 repeated closed-loop navigation trials. The system achieves a mean state-update latency of 29.63 ms, a mean relative route-progress error of 2.28% between the physical and virtual vehicles, and video delivery at 10.006 frames per second with 0.25% frame loss. All monitored navigation decisions were correctly reflected in the VR interface with no missed or incorrect notifications. The results indicate that the framework can support temporally synchronized, semantically consistent, and accurate route-progress representation for immersive observation of physical autonomous-vehicle behavior.
ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models
ShieldVLA:面向视觉-语言-动作模型的可行性感知安全对齐
Tayal, Manan, Nambi, Akshay
Abstract
Vision-Language-Action (VLA) models demonstrate strong generalization in robotic manipulation and navigation, but existing fine-tuning methods provide limited safety guarantees. Current approaches primarily rely on Lagrangian optimization that enforces safety through soft penalties on expected cumulative cost, often resulting in residual constraint violations or overly conservative behavior. Moreover, learning safety in visual domains is challenging due to the absence of dense per-step safety annotations. We propose ShieldVLA, a safety-aligned fine-tuning framework for VLA models based on Hamilton-Jacobi (HJ) reachability. ShieldVLA learns a model-free approximation of the HJ reachability value function directly from visual observations to estimate the safe operating region. The learned safety critic gates policy optimization by separating reward maximization within feasible regions from recovery near unsafe states, avoiding persistent reward-cost trade-offs. To enable scalable supervision in visual environments, we introduce rubric-based VLM safety scores that convert semantic safety feedback into structured critic targets without requiring manual cost labels. Across five navigation and manipulation benchmarks spanning multiple VLA backbones, ShieldVLA reduces cumulative safety cost by 57% on average and improves task success rate by +0.13 over SafeVLA.
Harnessing human expertise for high-precision robotic assembly in industrialized construction: A sample-efficient installer-in-the-loop interactive reinforcement learning framework
利用人类专业知识实现工业化建筑中的高精度机器人装配:一种样本高效的安装工在环交互式强化学习框架
Jin, Zekai, Wang, Huiguang, Sun, Xiaoning, Shao, Yi
Abstract
Industrialized construction imposes stringent precision requirements on robotic assembly of modular components such as prefabricated window units. In tolerance-critical operations, the central bottleneck is not only mechanical clearance but also converting tacit installer expertise into data-efficient autonomy under sparse acceptance feedback, contact variability, and millimeter-scale constraints. We present an installer-in-the-loop interactive reinforcement learning framework that acquires expertise through offline teleoperated demonstrations, sparse event-driven binary takeovers at contact-failure boundaries, and acceptance-aligned terminal rewards, logged under a unified schema for traceable offline-to-online adaptation. A temporally abstract action-sequence policy built on Q-chunking with Flow Q-Learning captures multimodal recovery maneuvers under sparse terminal rewards, while a non-updating warm-start phase stabilizes the offline-to-online transition. The framework is evaluated in MuJoCo across the workflow from suction acquisition through clearance-limited seating, under structured staging and end-to-end randomized placement. Within a defined stress-test regime with 2 mm per-side clearance, bounded pose perturbations, and friction randomization, the pipeline attains 100\% autonomous seating with 12--15 min of cumulative installer supervision over 3.0 h of online training, and reaches the 95\% success milestone in approximately 0.5 h and 1.5 h in the two experiments. We also report wall-clock adaptation time, cumulative takeover minutes, intervention-rate decay, and stage-wise failure attribution to inform supervision budgeting. Ablations isolate the complementary contributions of temporal abstraction, installer intervention, and warm-start value calibration.
Learning Manipulation-Sufficient Representations via Outcome Bottlenecks
通过结果瓶颈学习操作充分的表示
Sarowar, Md Selim, Kim, Sungho
Abstract
Networked manipulation endpoints couple perception to actuation across compute- and bandwidth-limited links, yet commonly exchange dense geometric states optimized for fidelity rather than action outcomes. A stochastic representation is learned with a policy-free, action-conditioned outcome bottleneck: marginal outcome log-loss supplies distortion and a KL term regularizes rate. The construction is motivated by the minimal statistic that preserves the outcome distribution of every admissible action, while the implemented finite model is evaluated as a rate-regularized mixture predictor. The same encoder and outcome head support grasp selection, singleton conformal filtering, active viewpoint selection, and latent test-time adaptation. A finite-probe theorem identifies the local level-set tangent space with the null space of an outcome Jacobian. The synthetic oracle verifies this result; on scanned objects, an analytic surrogate agrees with measured simulator invariances within \(0.56^\circ\). Across 11,979 simulated grasps on 13 objects, a reconstructed-geometry wrench score attains 0.542 AUC against lift success and falls below chance on curved objects, while our representation attains 0.876. At 25\% commitment, executed-grasp success is 0.503 versus 0.984. The 512-byte interface is \(288\times\) smaller than one RGB-D frame and runs at 16\,ms per CPU decision. Independent synthetic points track the tested conformal levels; scanned-object all-pair coverage is reported as a clustered empirical diagnostic. On unseen objects, within-scene AUC falls to 0.569, and a full-feedback update raises empirical mean pairwise coverage from 0.728 to 0.883.
Self-Evolving AI for Humanoids: Mechanisms, Safety, and Evaluation of Post-Deployment Self-Improvement
面向人形机器人的自演化AI:部署后自我改进的机制、安全与评估
Nguyen, Loc X., Raha, Avi Deb, Le, Huy Q., Huh, Eui-Nam, Niyato, Dusit, Hong, Choong Seon
Abstract
Humanoid robots are becoming an important part of embodied artificial intelligence, driven by advances in reinforcement learning for locomotion, world models for prediction, and vision-language-action models for general control. However, most of these systems remain static after deployment. A policy is trained offline for a fixed objective and then frozen, even though the tasks, environments, and robot bodies keep drifting over time. An emerging paradigm of self-evolving agents aims to address this problem by allowing systems to improve from their own post-deployment experience. Since most existing studies focus on disembodied software agents, this survey examines how self-evolution changes when an agent has a physical body. We first define self-evolution for humanoids and represent a deployed robot using a state tuple that includes its policy, perception, memory, workflow, and body. This state is updated by an evolution operator in a slow outer loop with a lifelong objective. We then organize the literature into four complementary mechanisms of self-evolution, presented in increasing order of autonomy: self-learning, self-adaptation, self-optimization, and self-generation. Since changes to a humanoid can introduce physical hazards, we treat safety and uncertainty as key design dimensions of the evolution operator, and further formulate admissible evolution as a constraint enforced by a world-model verification gate within a human-oversight envelope. Finally, we present that evaluation should track the robot's evolving trajectory rather than a fixed checkpoint, and we identify the lack of a benchmark designed specifically for self-evolving humanoids. Moreover, we outline open challenges spanning AI algorithms, on-board systems, and governance.
GzDRL: Reproducible and Scalable Deep Reinforcement Learning with Gazebo
GzDRL:基于Gazebo的可复现且可扩展的深度强化学习
Haridevan, Amal Dev, Kang, Junjie, Shan, Jinjun
Abstract
We present GzDRL, a novel single-process reinforcement learning (RL) framework for Gazebo that overcomes longstanding bottlenecks in scalable, reproducible robotics experimentation. Unlike conventional middleware-based RL-Gazebo integrations that suffer from nondeterminism and irreproducibility, GzDRL introduces a systematic, middleware-free environment-stepping mechanism that directly synchronizes agent actions and physics updates. This design enables deterministic, high-throughput data collection, efficient vectorization, and reproducible RL training and evaluation. Comprehensive benchmarks demonstrate that GzDRL achieves the highest workstation throughput among the evaluated frameworks while remaining competitive with GPU-accelerated simulators on laptop hardware, and maintains precise agent-environment synchronization, multi-agent scalability, and experiment-level reproducibility. We further validate sim-to-real transfer by deploying learned policies directly onto a physical quadrotor, without fine-tuning. Our results establish GzDRL as an accessible and reproducible platform for advancing RL in robotics and automation.
We study dark manipulation: after a brief lit Write encodes z0 = Enc(rgb), a policy pi(z) and open-loop dynamics f(z,a) complete contact-rich skills without further pixels (dark_f). On ManiSkill StackCube (n=160; seed packs 0/1000), dark_f attains 68.1% stacked on the five-rung chain (near_A -> grasped -> lifted -> on_B -> stacked), compared with 35.6% for per-step lit_reenc and 0% for freeze/encode_black. On a shared Write->HOLD protocol (n=40), occlusion and camera-aligned GT contact-neighbor masks drive lit lift from 43% to 0%, while dark_f holds 82.5%; shuffling actions inflates dynamics MSE by ~9.4x; write-time appearance shifts break encoding (night: 0% stacked), yet the same shifts during HOLD leave dark_f lift unchanged; Write length Tw is flat once the stop phase is reached, while earlier stops and write-time blur/JPEG sharply cut stacked. A dedicated pi_write reaches 35% vision-budget stacked (n=80); matched Dreamer-style/pixel nulls without privileged geom stay at 0%. Privileged state-RSSM MPC reaches ~35% stacked with 9D dark observations -- a stronger-observation null, not a matched visual baseline.
Conflict-Predictive Variable Horizons in Multi-Drone Distributed Model Predictive Control
多无人机分布式模型预测控制中的冲突预测可变时域
Mümken, Linda, Schwung, Michael, Lier, Stefan, Schwung, Andreas
Abstract
In distributed model predictive control for multi-drone collision avoidance, a fixed prediction horizon forces a compromise: a short horizon is inexpensive but reacts late to approaching neighbors, whereas a long one anticipates conflicts at a per-step cost that grows superlinearly with its length. We propose a conflict-predictive variable horizon that each drone sets locally, leaving the distributed model predictive control itself unchanged. From a short history of observed positions, a drone extrapolates the flight lines of its neighbors, tests each against its own using confidence funnels that narrow with prediction range, and obtains each time to conflict in closed form. The horizon is then the smallest admissible value whose planning window covers the farthest predicted conflict. It collapses to its minimum in clear airspace and grows only when a conflict lies ahead. Provided this minimum meets a single computable feasibility bound, we prove that recursive feasibility and asymptotic stability are preserved for every horizon the policy can select. These guarantees hold for a linear model, and a cascaded inner loop reduces each quadrotor's translational dynamics to a perturbed double integrator, so they carry over to the linearized quadrotor model and, as practical stability, to the full nonlinear one. In simulation on dense antipodal-swap benchmarks, the variable horizon reduces both per-step solver cost and total computation well below those of a long fixed horizon, and it maintains separation in every run, which a short fixed horizon of comparable per-step cost does not.
Achieving high-speed locomotion in quadrupedal robots remains highly challenging, as actuators operate near their physical limits and exhibit pronounced nonlinearities. However, many existing methods neglect actuator nonlinearities and physical constraints during training, leading to a significant sim-to-real gap under highly dynamic motions and limiting achievable performance. To address this issue, we propose a high-speed locomotion framework that reduces sim-to-real discrepancies and stabilizes learning over a wide command distribution. A refined actuator model explicitly captures high-speed voltage coupling and magnetic saturation, enabling a more accurate representation of the torque-speed envelope. In addition, a reinforcement learning framework incorporating a two-stage curriculum and adaptive command scheduling (ACS) ensures stable training. Experiments on the 36.5 kg quadruped BlackPanther2 (BP2) demonstrate speeds of up to 13.2 m/s on a treadmill and 11.65 m/s outdoors, establishing a new state-of-the-art and, to the best of our knowledge, a world record for quadrupedal robot locomotion. The results further highlight the importance of accurate actuator modeling in preventing non-physical policy exploitation, and show that ACS improves robustness without sacrificing performance.
Achieving biological-level running speeds has largely been pursued through advances in control algorithms, which improve the utilization of existing hardware. However, the ultimate speed limits remain governed by the underlying force and torque requirements of rapid locomotion, which are typically addressed through increased actuator capacity. Inspired by Huygens' coupled pendulums, we demonstrate that superior locomotion can emerge from principled exploitation of intrinsic dynamics rather than brute-force hardware scaling. Inter-limb inertial coupling redistributes energy across the gait cycle and reduces peak joint torque required for rapid periodic motion, thereby expanding the achievable speed without proportional increases in actuator capability. Incorporating hardware parameters as additional design variables further extends this analysis into a co-optimization framework, enabling the systematic utilization of inertial coupling in robot design. Guided by this framework, a quadruped robot achieves a running speed of 10.74 m/s (Froude number 21.4) and completes a 100-meter sprint in 12.2 seconds, representing the first legged robot to surpass 10 m/s. These results establish inertial coupling as an underlying mechanism governing high-speed legged locomotion and highlight its role in reducing force requirements, offering new insights into the design of agile robotic systems.
Diffusion learning leverages the statistical mechanism of diffusion processes for learning, reasoning, and inferring complex distributions from data. Recent advances in diffusion learning have been transformative, with robot learning emerging as a key opportunity area, with applications spanning perception, control, and decision-making. At the same time, the statistical mechanism of diffusion processes can be controlled to shape the temporal evolution of the state distribution underlying robot trajectories, inducing ergodic behavior in robotic systems. The frameworks of controlled diffusion and ergodic control were developed around the same time as diffusion learning, and their theories and algorithms have increasingly converged. Ergodicity induced by controlled diffusion has several significant implications for robot learning, distinct from applying diffusion learning to robotics problems: it formally enforces statistical properties required to ensure optimality of robot learning, enables non-myopic search over uncertain information landscapes for data collection, and enables behavior specification based on spatial rather than temporal characteristics of trajectories. This survey introduces the intuition behind controlled diffusion for robot learning, explores its connection to diffusion learning, presents theoretical foundations and numerical tutorials for solving controlled diffusion and ergodic control problems, and reviews applications across robotics. Finally, we discuss key challenges and future opportunities in leveraging controlled diffusion for robot learning.
IMM-based Multiple Object Tracking using a State Prediction Neural Network
基于IMM与状态预测神经网络的多目标跟踪
Lim, Chan-Bin, Paek, Dong-Hee, Kong, Seung-Hyun
Abstract
Object tracking is essential for autonomous vehicles to avoid obstacles and plan routes. Radar maintains detection performance even in adverse weather and can measure relative velocity through the Doppler effect, making it well suited for object tracking. In this paper, we propose a data-driven state PRedictor-based Interacting Multiple Model tracking method (PR-IMM) that improves nonlinear object-motion representation while preserving the stability and interpretability of physics-based motion models. The proposed method employs a transformer-based PRediction model (PR) that incorporates radar Doppler measurements to predict object displacement. The PR model is integrated into the IMM as a mode alongside the CV, CA, and CT motion models, and their prior positions are dynamically combined according to the mode probabilities. Experimental results show that PR-IMM reduces position-estimation error by 57.3% over the IMM and by 16.5% over the PR, while reducing ID switches by 25.3% and improving IDF1 by 9.6% over the IMM.
3D point-cloud observations are inherently ambiguous in complex, cluttered manipulation scenes, where target objects may be partially occluded or tightly intermingled with visually similar distractors. As a result, standard 3D diffusion policies often struggle to localize and exploit task-relevant geometry as scene complexity grows. We propose \textbf{Attention-DP3}, a spatially object-aware 3D diffusion policy that injects object-level geometric cues via attention while keeping the DP3 diffusion backbone unchanged. Our pipeline performs open-vocabulary 2D segmentation on RGB images, then lifts predicted target masks into 3D using calibrated camera geometry to obtain object-centric geometric priors. We incorporate these cues through Tri-field Attentional Conditioning, which constructs three complementary fields: (i) a targetness field to anchor the target object, (ii) an intra-target saliency field to emphasize task-relevant geometry within the target, and (iii) a backgroundness field to suppress distractors and clutter. Experiments on Adroit, DexArt, MetaWorld, and the real-world SO101 platform show consistent improvements over DP3, achieving state-of-the-art performance across benchmarks. Notably, as distractor objects increase, DP3 drops sharply, whereas Attention-DP3 remains stable and outperforms DP3 by up to 31\% under heavy clutter. The code is publicly available at https://github.com/zhangzhongbo2213/Attention-DP3.
Mahmud, Kazi Abrar, Dhurubo, Nilotpaul Kundu, Kirttonia, Tamal, Ujjal, Sabbir Hossain, Haque, Mohammad Ariful
Abstract
Large Language Models (LLMs) have enabled more natural human-robot interaction, but open-source models often exhibit unstable long-horizon reasoning and inefficient action execution when deployed in agentic robotic frameworks. This paper presents an enhanced ROS-Agent based architecture that improves task reliability and execution efficiency for agentic robotic systems using open-source LLMs. The proposed system introduces a novel intermediate mechanism, termed the MetaTool, which enforces structured planning prior to action execution. Given a natural-language command, the MetaTool induces the LLM to generate a pseudo-code plan of intended tool invocations, which is stored in the ROS-Agent's scratchpad and persists throughout execution. By explicitly separating planning from execution, the proposed approach reduces execution loops and improves deterministic behavior. The architecture is validated on a custom mobile robotic platform with multimodal perception and motion control capabilities. Experimental results on real-world interactive tasks demonstrate improved task completion and contextual consistency, with up to ~24% gains on complex tasks compared to the baseline framework.
Chance-Constrained Belief-Space Maneuver Planning for Autonomous Collision Avoidance Under Uncertainty
不确定性下自主避碰的机会约束信念空间机动规划
Kim, Grace Ra, Eddy, Duncan, Kochenderfer, Mykel J.
Abstract
Increasing conjunction frequency in low Earth orbit places growing pressure on spacecraft operators to determine not only whether an encounter requires mitigation, but whether sufficient information is available to commit to a maneuver. This work formulates this information-action tradeoff as a belief-space planning problem for conjunctions between a maneuverable spacecraft and an unmaneuverable secondary object. The planner represents the uncertain orbital states as Gaussian beliefs and uses a chance-constrained belief-space Monte Carlo tree search framework to reason over possible future tracking updates before time of closest approach (TCA). A terminal chance constraint limits the probability of reaching TCA above a prescribed collision-risk threshold, allowing the planner to wait for informative tracking while intervening when deferral becomes too risky. We evaluate the approach on eight historical conjunctions from NASA's Conjunction Assessment Risk Analysis dataset. By varying the secondary-object measurement quality and tracking cadence, we generate a total of 96 distinct evaluation scenarios. Across the evaluated conditions, the planner reaches TCA without maneuvering in approximately 40% of episodes while maintaining no terminal collision-risk violations. In contrast, fixed-time rule-based maneuver policies resolve more encounters without maneuvering when intervention is deferred closer to TCA, but at the expense of increasing terminal risk violations. The fraction of episodes reaching TCA without maneuvering depends strongly on tracking quality and measurement cadence, ranging from 76% under accurate, frequent measurements to approximately 18%-20% under the poorest tracking conditions. These results show that tracking quality and frequency are not only inputs to collision-risk estimation: they can determine when intervention becomes necessary.
STAGE: Diagnosing Semantic Transfer at Grounded Execution in Embodied Agents
STAGE:诊断具身智能体接地执行中的语义迁移
Jin, Baosheng, Liang, Yushen, Shen, Hua
Abstract
Embodied language grounding requires more than identifying the referent of an instruction: recovered semantics must also control the action an agent exposes. We study this missing link as a semantic-action gap, where instruction semantics are recoverable but weakly expressed in native continuous actions. We introduce SAT-Bench, a fixed-observation counterfactual benchmark that holds the visual scene and agent state fixed while changing only instruction semantics. On LIBERO target-name and pixel-grounded relation swaps, target recovery reaches 100.0% and 95.8%, whereas OpenVLA action sensitivity remains only 6.8% and 7.7%. The gap persists across 1,000 additional compositional and temporal/procedural counterfactuals, with overall action sensitivity of 6.1%. Hidden-state, threshold-free, cross-policy, and rollout diagnostics further support this semantic-action transfer failure. We introduce VISA, a lightweight execution-time interface that converts recovered semantics into ALLOW, DEFER, target-consistency, and verified-selection decisions. VISA reduces invalid-instruction blind execution from 92.7% to 2.8% while preserving 94.0% of normal commands, and verified selection further improves target-consistent action exposure without updating the underlying policy. Overall, embodied language evaluation should measure semantic-action transfer, not semantic parsing alone.
Constraint-Grounded Reinforcement Learning for Variable Impedance Control in Contact-Rich Robotic Insertion
接触丰富机器人插入中变阻抗控制的约束锚定强化学习
He, Lin, Deng, Min
Abstract
In robotic insertion under uncertain contact, the axial force limit and the appropriate controller gain vary across tasks. As a result, a single fixed gain is unlikely to remain suitable across different task conditions, making conventional impedance controllers reliant on manual retuning. To eliminate manual retuning, we propose Constraint-Grounded Reinforcement Learning (CG-RL), a variable impedance framework for online gain adaptation. Conditioned on the force limit and contact feedback, the policy outputs a residual motion, an insertion rate, and a requested gain. The controller projects this gain into the admissible range without exposing the range itself to the policy. This separation allows a single policy to operate under different force limits without retraining or manual retuning. We evaluate CG-RL on simulated oblique insertion across five training seeds. CG-RL achieves an $85.8\pm7.7\%$ (mean $\pm$ SD) success rate of insertions without violating the force limit, while keeping the applied gain within the admissible range. As a comparison, a fixed-gain baseline using the midpoint gain achieves a success rate of $50.1\%$. The policy adapts its insertion rate continuously to the specified force limit and further generalizes to more permissive force limits above the training range. In contrast, the same actor without force-limit input does not exhibit this adaptation. The applied gain is guaranteed to remain within the admissible range, while force-limit satisfaction is validated empirically rather than guaranteed formally.
Runtime-Incremental Transformer for Reinforcement-Learning-Based Adaptive Control
面向基于强化学习自适应控制的运行时增量Transformer
Cirrincione, Giansalvo, Fagiolini, Adriano
Abstract
Learning-based adaptive control of robotic manipulators with non-observable friction memory has been addressed by attention- based meta-controllers whose number of attention heads is fixed before training and is tuned by costly offline search. At long memory horizons, such fixed-capacity controllers are prone to catastrophic failures on a sizeable fraction of training seeds. The present paper introduces a runtime mechanism that grows and prunes the heads of the attention block during reinforcement learning, governed by two signals: the effective rank of the on-policy context distribution, which triggers growth when representational capacity becomes insufficient, and the per-head output magnitude, which flags redundant heads for removal. Policy continuity at growth events and a quantitative bound at prune events are established analytically. On a two- link manipulator with Stribeck friction, the proposed mechanism attains full success across all memory regimes, eliminating the long-horizon failure mode and removing the need for offline tuning of the head count.
An MRI-Guided Robotic System to Improve Hippocampal Access for Epilepsy Interventions
一种用于改善癫痫干预中海马体入路的MRI引导机器人系统
Peters, John E., Grillo, Abby M., Mansouri, Mahshid, Esser, Daniel S., Garrow, Sarah, Kumar, Nithin S., Englot, Dario J., Neimat, Joseph S., Grissom, William A., Webster III, Robert J., Barth, Eric J.
Abstract
This paper presents an MRI-guided robotic system that improves hippocampal access by delivering a curved needle-like laser ablator through the foramen ovale, a natural opening in the base of the skull. Both the delivery path and the curved nature of the needle improve upon current clinical straight-line laser interstitial thermal therapy (LITT) as measured by hippocampal cannulation percentage. We describe the design of the robotic system which includes positioning, aiming, and curved needle deployment stages, followed by experimental results assessing cannulation percentages using MRI images in phantoms. In three curvilinear, single-insertion experiments, we achieved cannulation percentages of 93.5%, 96.5%, and 60.4%, exceeding reported clinical averages of 50-60%. Seizure control is believed by physicians to be a function of hippocampal volume treated, and our system provides a means of treating a greater volume of the hippocampus with LITT-based interventions.
Differential Realizability of Static Control Allocation in Multirotors: An Impossibility under Nonredundant Full Actuation and a Pseudoinverse Obstruction under Redundant Actuation
多旋翼静态控制分配的微分可实现性:非冗余全驱动下的不可能性与冗余驱动下的伪逆障碍
Franchi, Antonio, Mizzoni, Mirko
Abstract
Control allocation for multirotors with bidirectional propellers is commonly formulated in signed-thrust variables, where the wrench map is linear. The signed-quadratic map from physical rotor speed to thrust, however, is not a local diffeomorphism at zero speed. This work derives two distinct consequences. Under nonredundant full actuation, a single-propeller reversal removes one instantaneous task direction; hence, no global continuously differentiable exact static allocator exists over the complete task space. Under redundant actuation, the physical task map may remain regular, yet a transverse pseudoinverse zero crossing requires an unbounded rotor-speed derivative. We define differential realizability as regularity of the physical lift of an actuator-output section, derive exact and first-order validity conditions, and distinguish structural rank loss from an allocator- induced rate singularity. A local nullspace deformation repairs isolated pseudoinverse reversals, while a global fixed-orthant construction establishes existence of regular sections at the cost of persistent task-preserving internal actuation.
Multi-arm robotic harvesting offers a promising path to improve harvesting efficiency and reduce reliance on manual labor. However, practical deployment remains challenging because the system must generalize across diverse environments while efficiently coordinating multiple arms in a shared workspace. Existing methods often require substantial data collection in target environments or rely on simplifying assumptions that limit planning quality. In this work, we introduce the first comprehensive benchmark for evaluating pretrained Vision-Language Models (VLMs) on zero-shot multi-arm fruit harvesting planning. Our benchmark uses real-world apple and citrus orchard images and compares a VLM-based planning pipeline with a traditional perception-and-planning pipeline. The VLM pipeline directly generates harvesting sequences and waypoints for each arm, while a lightweight trajectory verifier checks for collisions. Our results show that frontier VLMs can generate effective multi-arm harvesting plans zero-shot, but a practical deployment remains limited by accurate 3D waypoint generation and collision-aware coordination. These results highlight both the promise and current limitations of pretrained VLMs for multi-arm robotic harvesting.
How Well Do Pseudo-Rigid Body Models Capture Real Plants? In-Field Validation of Simulated Blueberry Canes
伪刚体模型能多好地捕捉真实植物?——模拟蓝莓茎秆的田间验证
Kolano, Hannah, VanAtter, Chelse, Yang, Wei, Grimm, Cindy, Davidson, Joseph R.
Abstract
Many labor-intensive tasks in fruit production such as pruning require physically interacting with the plant (e.g., pushing, pulling, bending limbs, etc.). Due to increasing labor shortages, there is widespread interest in the adoption of robotics in this area. When the robot must physically interact with the system, a deformable model of the plant is beneficial. Coupled, spring-loaded beams have been used in prior work to build deformable plant models in simulation for learning, planning, and control, but this approach has rarely been validated against the mechanical properties of real-world plants. In this paper, we use this rigid-body model, grounded in real material properties and beam bending theory, to simulate the deformation of blueberry canes under loading. To validate the model, we simulate the canes in MuJoCo and compare their behavior with real-world data collected from probing canes at a commercial farm with a custom testbed. We perform sensitivity analysis on several key modeling variables and show that this approach is highly sensitive to the plant diameter calculations and a priori flexural modulus estimation, which is dependent on season and blueberry variety. We also share our dataset of live blueberry cane deformation, including RGB-D images and measured forces and displacements.
Compositional Shift Algebra: Extrapolating Mixed Robot Shifts Without Mixed Finetuning
组合偏移代数:无需混合微调即可外推混合机器人偏移
Hang, Jinting, Cai, Zhenhui
Abstract
Robot deployments rarely change one mechanism at a time: cameras, action interfaces, and physical dynamics often shift together. Prior adaptation recipes either finetune a new model for every mix or attempt to select which module to update. We instead learn shift operators on a modular stack z{=}E(o), a{=}g(z,u), z'{=}f(z,a) and compose them. Compositional Shift Algebra (CSA) fits single-factor observation, policy, and dynamics operators from exact-reset probes, then extrapolates held-out mixed shifts by operator composition---without mixed-shift finetuning. On ManiSkill StackCube, residual CSA matches an oracle mixed inverse on held-out mixes (success 1.0 over 10 seeds) while beating best-single / zero-shot / parameter-average baselines by { approx}67 pp. RGB-D vision-in-the-loop composition remains near oracle and far above non-compositional arms; a delay commutator stress shows ordered necessity for policy timesdelay. On a second task (PickCube), residual CSA again reaches compose 1.0 vs. 0.33 non-compositional (n{=}10), and an L1 vision controller without privileged cube/goal poses or grasp flags in the control loop retains compose 0.95 vs. 0.00. Main-track upgrades freeze PushCube (+33 pp), PegInsertion joint8 / pose7 EE (+67 pp each), and thin BC under frozen CSA (+67 pp); deeper BC and fair adapt baselines still need compose (+67 pp each), vision-localized BC needs compose (+56 pp), and delay favors ordered/few-shot deploy. We report Intervention-Gated Adaptation as a negative control.
LiDAR-inertial odometry (LIO) systems differ in whether they continue estimating gravity after initialization. We compare four gravity-bias state configurations in each of FAST-LIO2 and LIO-SAM, then separately test a gravity-direction factor. Across 12 dataset sequences evaluated with FAST-LIO2, fixing gravity under continuous LiDAR correction produces mean paired changes in vertical and 3D position errors with 90% confidence intervals within $\pm 2\%$. Tests on 4 sequences with LIO-SAM likewise show no consistent benefit from online gravity. Multi-second LiDAR outages, unlike reduced range or field of view, reveal trajectory-dependent costs of fixing gravity. A history-matched 23D-to-21D switch places the repeatable 3D error increase after LiDAR updates resume. Under 5-s outages, a direction factor from the same IMU used for preintegration improves accuracy on Hall05 but worsens both errors with online gravity on TUHH. Dynamic-start tests also show fixed-bias failures at particular starting phases. We recommend keeping gravity and accelerometer bias online for robustness; use a direction factor only after verifying vertical and 3D accuracy gains under the intended operating conditions.
How can robot policies learn more effectively from a fixed demonstration budget? The first Real-world Embodied AI Learning (REAL-I) Challenge at ICRA 2026 examined this question through simulation, real-robot evaluation, and an on-site final on a shared dual-arm humanoid platform. We describe the challenge tasks, data and deployment interfaces, and competition results, then compare the approaches contributed by NUS-CLEAR, RCL-Lab, and Deeptouch.ai. Their systems combined pretrained vision-language-action models and task-specific imitation policies with different strategies for data curation, staged adaptation, checkpoint selection, and action-space design. The team reports highlight the importance of adapting to the deployment environment while retaining prior capabilities, treating demonstration quality at an appropriate temporal scale, and suppressing errors in inactive robot components. They also expose the limitations of offline action-prediction metrics for forecasting closed-loop success. These observations motivate a view of fixed-data robot learning that integrates data, adaptation, evaluation, and deployment.
GROOVE: Geometry-Guided Reduction of Operational-Space Jerk in VLA Execution
GROOVE:几何引导的VLA执行中操作空间加加速度降低
Yun, Sangho, Kim, Minsoo, Cho, Minwoo, Yu, Hwanjo
Abstract
Chunked vision language action (VLA) policies execute several commands per query, but jerk within chunks and across replanning boundaries can induce oscillatory motion and sharp actuator transients. We present GROOVE, an online regulator that searches directional correction regions around the raw three dimensional end effector (EEF) path, without retraining or additional VLA inference. It optimizes the new chunk using delivered commands as boundary conditions, reducing boundary and within chunk jerk while bounding cumulative translation and local axis angle deviation from the raw plan after every command. Using quadratic programs (QPs), GROOVE generates a cube reference and thirteen directional candidates, then selects the one with the lowest command space jerk under a reference relative deviation cap. On a held out LIBERO benchmark, GROOVE achieves the largest reductions among the evaluated methods, reducing translational and rotational EEF jerk by 33.02% and 43.42%, respectively, with task success of 95.75% versus 93.75% for raw execution. Across 50 matched UR5e pairs with measured execution timing, it reduces translational and rotational tool center point (TCP) jerk by 16.39% and 19.49% and joint current slew by 29.09%.
Decentralized Multi-Robot Task Allocation Under Degraded Communication: A Benchmark of Performance, Reliability, and Computation
通信退化下的分散式多机器人任务分配:性能、可靠性与计算基准
Lott, James, Honary, Vahraz
Abstract
Selecting a decentralized Multi-Robot Task Allocation (MRTA) method for embedded deployment on autonomous platforms requires considering more than route performance alone. We benchmark six decentralized MRTA allocators (CBAA, ACBBA, PI, HIPC, DMCHBA, and DGA) in the Collaborative Visit (CV) scenario to characterize tradeoffs among MinMax and MinSum travel, communication robustness and demand, allocation reliability, computational burden, and scale sensitivity. The core study uses 500 paired ten-target instances across 25 ideal and degraded communication conditions spanning Bernoulli loss, Gilbert--Elliott loss, and Rayleigh fading, with additional campaigns examining pre-allocation, execution-integrated computation, and sensitivity to grid size, robot density, and target load. Across the 24 impaired core conditions, DGA and DMCHBA achieved the lowest mean MinMax travel at 24.49 and 24.78 steps, respectively. HIPC narrowly led mean MinSum travel at 66.95 steps, followed by DGA at 67.22, with both methods occupying the top two in every impaired condition. DMCHBA had the lowest publication intensity at 2.08 publications per team step. In ten-target pre-allocation, HIPC and DMCHBA remained viable and stable in every tested condition, while ACBBA, PI, and DGA lost stability or viability as communication degraded. Under ideal delivery, median full-protocol computation $\Cterm$ in the primary ten-target comparison ranged from 4.88 ms for DMCHBA to 1.346 s for DGA. Static route quality preserved DGA and DMCHBA as the leading MinMax methods, while DGA led MinSum at three of four target loads and HIPC led at 50 targets. Static and execution-integrated computation rankings diverged as task load increased. The results identify distinct allocator operating regions across route objective, communication behavior, reliability, and computational constraints.
In-hand manipulation allows multi-fingered dexterous hands to reconfigure grasped objects without releasing and regrasping them. This improves manipulation efficiency by reducing repeated grasp acquisition and large arm motions. However, most learning-based methods focus on reorientation, continuous rotation, or translation, whereas many tasks require joint control of object position and orientation. We formulate this capability as in-hand 6D object pose reaching: starting from an existing grasp, coordinated finger motions move the object to a palm-relative target pose. We present POISE (Palm-relative Object reaching In SE(3)), a sim-to-real reinforcement learning framework for this task. POISE combines diverse stable-grasp initialization, goal- and geometry-conditioned control, an adaptive 6D goal curriculum, and a compact reward scheme for pose reaching and grasp preservation. In simulation, diverse initialization raises held-out-grasp success from 40.1% to 51.5% and post-drop recovery from 33.8% to 72.9%; the curriculum raises full-range success from 6.2% to 59.5%. On hardware, the grasp-maintenance reward improves three-target sequence success from 20% to 80%. In real-world experiments, POISE reaches successive 6D targets without manual reset across multiple object geometries and wrist orientations, and recovers from external disturbances. To support further research in dexterous manipulation, we will release our code at https://junxiaolin.github.io/poise-website/.
Chinese Translation
手内操作使多指灵巧手能够在无需释放并重新抓取物体的情况下重新配置被抓取的物体。这通过减少重复抓取获取和大幅手臂运动来提高操作效率。然而,大多数基于学习的方法关注重新定向、连续旋转或平移,而许多任务需要联合控制物体的位置和姿态。我们将这一能力形式化为手内6D物体位姿到达:从现有抓取出发,协调的手指运动将物体移动到相对于手掌的目标位姿。我们提出了POISE(Palm-relative Object reaching In SE(3)),一个用于该任务的仿真到现实的强化学习框架。POISE结合了多样化的稳定抓取初始化、目标和几何条件控制、自适应6D目标课程学习,以及用于位姿到达和抓取保持的紧凑奖励方案。在仿真中,多样化初始化将留出抓取成功率从40.1%提高到51.5%,将掉落恢复率从33.8%提高到72.9%;课程学习将全范围成功率从6.2%提高到59.5%。在硬件上,抓取维持奖励将三目标序列成功率从20%提高到80%。在真实世界实验中,POISE能够在多种物体几何形状和手腕朝向下连续到达6D目标而无需手动重置,并且能够从外部扰动中恢复。为支持灵巧操作领域的进一步研究,我们将在https://junxiaolin.github.io/poise-website/发布我们的代码。
When Do Learned Priors Help Visual Inertial Estimation? A Controlled Study of Prior Integration, Calibration, Initialization, and Backend Consistency
学习先验何时有助于视觉惯性估计?关于先验集成、标定、初始化和后端一致性的受控研究
Zhang, Jinchang, Lu, Guoyu
Abstract
Learned components are increasingly integrated into geometric visual--inertial estimators to provide motion, depth, bias, uncertainty, or confidence cues. Yet it remains unclear whether gains arise from useful learned priors or from changes in the backend, calibration, initialization, temporal association, or evaluation gauge. We present a controlled framework for learning-augmented visual--inertial estimation that separates fusion gain from the incremental value of a learned prior and evaluates four evidence layers: local motion consistency, global trajectory accuracy, physical-state correctness, and numerical consistency. We instantiate the framework with a MonoViT-based monocular motion prior added as a local relative-motion factor to an unchanged VINS backend. We compare Original VINS and learned-prior VINS under matched sensor streams, timestamps, initialization, frontend/backend settings, and camera--IMU extrinsics, while probing calibration, initialization, state coupling, scale, bundle adjustment, and loop closure. On KITTI, with fixed reference extrinsics, translation APE RMSE is 31.4 m for Original VINS and 31.8 m with the learned prior. Across four recordings, the prior changes mean APE by only -0.2%, while a five-times-higher weight worsens it by 8.2%. Online extrinsic updates increase mean APE by 45.7% and 52.1%, respectively, while mean RPE changes by less than 2%. These results show that fusion performance alone cannot establish the value of learned priors. Reliable evaluation requires same-backend controls and joint analysis of prior compatibility, calibration, initialization, global drift, physical-state error, and backend consistency.
Force-controlled loco-manipulation requires a whole-body policy to coordinate locomotion and arm motion while regulating end-effector interaction forces. This is challenging under floating-base dynamics and changing support contacts, particularly when end-effector force/torque sensing is unavailable for control. This paper presents a force-aware reinforcement learning approach with hybrid sensorless force estimation for wheeled-legged loco-manipulation. The proposed method provides a structured estimate of the end-effector force as an explicit policy observation, enabling force-guided contact behavior without using an end-effector force/torque sensor for control. The force estimate is obtained by combining generalized momentum observation, contact-constrained wrench projection, and temporal residual learning: the model-based components extract the physically structured part of the whole-body disturbance, while the residual network compensates the remaining motion-dependent bias. The estimated force is integrated into a mode-conditioned whole-body policy with an axis-wise force/position selector, allowing free-space motion, pure force regulation, and hybrid force/position control within one controller. Simulation results demonstrate improved sensorless force estimation and force-control performance. Hardware experiments further validate the proposed controller through quantitative valve-rotation and hybrid wiping evaluations, together with force-guided door opening and zero-force human-guided motion on a real wheeled-legged platform.
We present GeomVLA, a Vision-Language-Action (VLA) model that unifies perception, latent scene motion prediction, and action generation within a shared robot-centric 3D coordinate frame. Our approach lifts pretrained VLM features into spatially grounded 3D scene tokens using depth and camera calibration, while retaining the semantic representations learned during VLM pretraining. We further introduce a 3D Scene Trajectory Denoiser, a task-conditioned module that learns a latent representation of how scene points are expected to move in 3D. Rather than executing the predicted trajectory as an open-loop plan, GeomVLA extracts intermediate motion tokens from the trajectory denoiser and uses them to condition a 3D flow-based action denoiser through geometry-aware attention. GeomVLA achieves state-of-the-art performance on CALVIN, competitive performance on LIBERO and RoboTwin2.0, and outperforms strong baselines in real-world manipulation settings without robot-action pretraining. Extensive ablations show that future-motion reasoning alone is insufficient: the primary gains are associated with maintaining geometric consistency among scene representation, motion prediction, and robot actions throughout the perception-to-action pipeline.
Chinese Translation
我们提出 GeomVLA,一种视觉-语言-动作(Vision-Language-Action, VLA)模型,其在共享的以机器人为中心的三维坐标系内统一了感知、潜在场景运动预测与动作生成。我们的方法利用深度和相机标定,将预训练 VLM 特征提升为具有空间基础的 3D 场景 token,同时保留 VLM 预训练期间学到的语义表示。我们进一步引入 3D Scene Trajectory Denoiser(三维场景轨迹去噪器),这是一个以任务为条件的模块,学习场景点在三维空间中预期如何运动的潜在表示。GeomVLA 并不将预测轨迹作为开环计划来执行,而是从轨迹去噪器中提取中间运动 token,并通过几何感知注意力将其作为条件输入一个基于 3D 流的动作去噪器。GeomVLA 在 CALVIN 上取得了最先进性能,在 LIBERO 和 RoboTwin2.0 上具有竞争力,并在无需机器人动作预训练的真实世界操作设置中优于强基线。大量消融实验表明,仅靠未来运动推理是不够的:主要收益与在从感知到动作的整个流程中保持场景表示、运动预测和机器人动作之间的几何一致性有关。
LePlanner: An Iterative Amortized Controller For World Models
LePlanner:一种用于世界模型的迭代摊销控制器
Bansal, Saksham, Naphade, Om, Aggarwal, Chayan, M, Vrishin
Abstract
World models trained with joint-embedding predictive architectures learn compact, structured latent representations from physical interaction, yet planning in these latent spaces typically relies on one of two costly approaches. Search-based planners such as CEM, MPPI, and iCEM optimize action sequences through many predictor rollouts, achieving strong performance at the cost of high per-decision compute and latency. Policy-based methods amortize inference into a single forward pass but can degrade on contact-rich tasks where the demonstration distribution is multimodal. We propose LePlanner, an amortized iterative controller that learns to construct and refine latent action sequences through a frozen world-model predictor. LePlanner is trained with an arrival-and-hold objective that encourages the controller to reach the goal at the earliest feasible horizon and remain there. This addresses horizon-reset procrastination, a failure mode in which repeated receding-horizon replanning continually postpones goal arrival. An additional action-Gaussian loss keeps generated actions near the support of the offline dataset. Across navigation, contact-rich manipulation, and continuous-control environments, LePlanner matches or exceeds search-based planners while requiring an order of magnitude fewer predictor evaluations and 3-49x lower wall-clock time per decision. It achieves success rates of 98% on PushT, 100% on Reacher, 100% on TwoRooms, and 92% on the OGBench Cube task. These results show that much of the structure discovered through online search can be amortized into a lightweight iterative policy, enabling fast, horizon-aware, nonlinear physical control without online optimization.
ReWeight: Leveraging Human Data for VLA Post-Training via Demonstration Retrieval and Sample Weighting
ReWeight:利用人类数据,通过演示检索与样本加权实现VLA后训练
Wang, Chenwei, Huang, Dianye, Ko, Match W. L., Bai, Chenjia, Jiang, Zhongliang
Abstract
Post-training vision-language-action (VLA) models for specific robots and tasks requires in-domain demonstrations, yet collecting diverse robot data is costly. Egocentric human demonstrations provide a scalable alternative, but directly mixing human and robot data can introduce cross-embodiment discrepancies and degrade policy performance. To address this challenge, we introduce ReWeight, a framework that incorporates human data into VLA post-training through demonstration-level retrieval and sample-level weighting. ReWeight learns a cross-embodiment visuomotor representation that combines visual observations with future actions to measure behavioral similarity between human and robot demonstrations. Based on optimal transport, it retrieves human demonstrations relevant to the target robot data and assigns larger weights to samples with smaller cross-embodiment discrepancies. We evaluate ReWeight using $\pi_{0.5}$ across eight simulation tasks and four real-world tasks under both clean and randomized settings. In simulation, ReWeight improves the average success rate of post-trained $\pi_{0.5}$ from 39% with only robot data and 44% with randomly mixed human-robot data to 57%. In the physical experimental setting, it achieves an average success rate of 68.8%, outperforming the baselines by 28.8% and 13.8%, respectively. Overall, ReWeight provides an effective paradigm for transforming abundant egocentric human experience into transferable supervision for robot learning. (Project webpage: https://reweight-vla.github.io/)
Co-Speech with You: Training-Free Personalization of Robot Co-Speech Gestures
与你共语:机器人共语手势的免训练个性化
Ding, Bosong, Ancel, Selma, Spigler, Giacomo, Kirtay, Murat
Abstract
Personal robots should adapt their co-speech gesture style to a new user without requiring model retraining. We present a training-free personalization pipeline that combines a frozen audio-conditioned diffusion prior with a gesture style encoder and lightweight conditioning adapters. The encoder is first trained to discriminate speaker identities and then jointly refined with the adapters using the diffusion objective, enabling a reusable style embedding to be extracted from approximately 10 seconds of enrollment motion through a single forward pass. To support this setting, we also release a Quest~3 capture application and a dataset of spontaneous co-speech motion from ten participants. We evaluate the system on held-out speakers using Style Recognition Accuracy (SRA) and Fr'echet Gesture Distance (FGD) to measure personalization and motion quality. Our approach improves SRA from 27.6\% for the frozen prior to 69.5\% while preserving motion quality (FGD 34.2 versus 34.8), and replacing the enrollment embedding with another person's reduces SRA to 11.4\%. The generated gestures are retargeted to a physical NAO robot, and this improvement also transfers to the robot deployment setting, where speech is synthesized using five TTS voices, retaining 67.3\% SRA.
In Human-Robot Interaction, the standard approach to learn a reward model that represents human preferences for robot behavior consists of three steps. First, the robot collects limited direct evidence from human feedback (e.g., positive or negative binary feedback). Then, the robot utilizes the direct evidence to derive accepted or rejected labels to feasible but unchosen actions using fixed implication rules. Finally, the robot updates the reward model with both the direct and derived evidence. Unfortunately, the fixed rule can hinder preference learning: in a user study with two collaborative simulation environments, human-provided implication labels often differed from the standard fixed rule, and using the human labels substantially improved reward learning with the Preference Learning from Implicit and Explicit Feedback (PIE) algorithm. Consequently, we propose IMPLIED, an implication modeling method that treats fixed-rule implications as an initial guide while learning to infer and revise accepted and rejected action labels over time. Across evaluations on recorded human-robot interaction trajectories and a physical robot pizza-making study, IMPLIED predicts human implications more accurately than the fixed rule approach and LLM baselines, approaching the performance of a human-label oracle. In turn, IMPLIED reduces preference-estimation error and leads to robot actions that are more often rational with respect to a combined reward (which includes the true preference reward and a task-specific reward) compared to baselines. By learning to reason about the implications of human feedback, this work enables more faithful and efficient robot behavior adaptation during human-robot collaboration.
Vision-Language-Action (VLA) models combine a pretrained vision encoder, a language backbone, and an action head, but their relative contribution has not been established under controlled, latency-paired conditions. We fix the backbone families (SigLIP2 and Qwen2.5) and the training pipeline, sweep action-head design and module scale, and pair each configuration with measured on-device latency. The study yields three findings. First, action-head performance is governed primarily by initialization rather than decoder architecture, loss, or inference budget: copying the last transformer layers of the language backbone into the head is the single largest lever, at no latency cost, and the only axis that helps at every module scale. Alignment also explains the other axes: flow matching and a heavier decoder pay off only while the head is misaligned and reverse once it is aligned, and extra inference passes give no measurable benefit; expressiveness appears to substitute for missing alignment. We read this as representation transfer: the aligned head keeps attending to the instruction's object nouns and stays close to the backbone in weight space rather than relearning to act from scratch. Because we reach alignment only through initialization, we offer this as the account that best organizes the measurements, not a demonstrated cause, and name the control that would settle it. Second, capacity pays only after alignment: the aligned action head is the highest-return module to scale. Third, those returns diminish sharply near the size today's $\pi$-series VLAs already use, so further growth buys little in-domain accuracy for its latency. These specify EffVLA, a compact model matching the strongest open-source VLAs on standard LIBERO, leading on most LIBERO-Plus perturbation axes at lower latency, and transferring to a real SO-ARM101 arm with the recipe unchanged.
Mobile Multi-Robot Navigation under Runtime Uncertainty via Koopman Operator Learning and Nonlinear Model Predictive Control
基于Koopman算子学习和非线性模型预测控制的运行时不确定性下的移动多机器人导航
Zhang, Xiaobin, Karydis, Konstantinos
Abstract
In this work, we developed a nonlinear model predictive control (NMPC) framework that employs learned dynamics via the Koopman Operator theory for mobile multi-robot navigation. We formulated and solved NMPC problems using a lifted bilinear Koopman-based model that accurately predicts affine input systems affected by perturbations and uncertainties. Two exemplary multi-robot navigation problems are considered: target reaching and formation control. The output of our method enables closed-loop multi-robot navigation and formation control in environments populated with obstacles, whereby the Koopman operator-based model used in the NMPC formulation addresses runtime uncertainties, namely, various degrees of random wheel slipping. We validated the effectiveness of our method for both problems via extensive numerical simulations in different environments with wheeled robots affected by different amounts of slip and without knowledge of their true dynamic models.
Gradient-Free Neural Hamilton-Jacobi Reachability for Scalable Safety-Critical Control
面向可扩展安全关键控制的无梯度神经Hamilton-Jacobi可达性
Feng, Zeyuan, Sahin, Ali Fuat, Thorup, Santiago, Bansal, Somil
Abstract
Hamilton-Jacobi (HJ) reachability provides a principled framework for synthesizing safety certificates and robust controllers for safety-critical robotic systems. However, applying reachability analysis to high-dimensional nonlinear systems remains challenging: classical grid-based solvers suffer from the curse of dimensionality, continuous-time neural solvers require accurate spatial value gradients, and reinforcement-learning-based approaches often suffer from weak boundary anchoring and non-stationary adversarial policy optimization. We propose a discrete-time neural reachability framework for control-disturbance-affine systems that learns backward reachable tubes (BRTs) and backward reach-avoid tubes (BRATs) through Bellman-Isaacs value propagation. Our key idea is to combine equation-driven self-supervision with structured policy learning: rather than computing explicit PDE-gradients, we exploit the bang-bang structure of optimal safety interventions to construct approximate teacher actions from gradient-free value probes, converting adversarial actor learning into supervised policy learning. To stabilize long-horizon value propagation, we leverage the learned actor to train the value function backward from the terminal boundary using a windowed temporal curriculum, where each window is used as the boundary condition for the next window. Across benchmark problems up to 80 dimensions, our method learns accurate reachability value functions while improving stability over existing learning-based solvers. We further demonstrate observation-space scalability on F1-tenth racing with over 16,000-dimensional egocentric inputs. The learned safety filter generalizes zero-shot to unseen tracks and transfers to a physical RC car, achieving real-time robust collision avoidance.
Precise manipulation in dynamic environments, whether induced by a mobile robot base or a target with unknown motion, remains a major challenge in robotics. Manipulation in dynamic environments introduces substantial uncertainty, which fundamentally conflicts with the tight precision requirement of precise tasks such as peg-in-the-hole. We propose a Vision-Force Admittance Learning (VFAL) framework that fuses asynchronous visual feedback with a high-frequency force-based model, using visual pose estimations as a regularization term. VFAL adapts insertion strategies online to dynamic motion while maintaining millimeter-level precision. To obtain robust, low-frequency pose information, we employ state-of-the-art vision foundation models for visual pose estimation. Additionally, we incorporate failure recovery mechanisms to enhance overall robustness. We validate our approach in real-world experiments, demonstrating high success rates and strong adaptability to various pegs and dynamic environments.
Vision-language-action (VLA) deployment can reduce inference latency while changing closed-loop task behavior. We evaluate HuggingFaceVLA/smolvla_libero on an RTX 2060 (6 GB) in LIBERO Spatial and Object (MuJoCo 3.3.2, LeRobot 0.6.1, seed 42), comparing PyTorch+AMP with ONNX Runtime CUDA Execution Provider (CUDA EP). The main evaluation uses 100 episodes/suite; a paired rollout uses 300 episodes/suite. PyTorch+AMP reaches 70.0%/88.0% Spatial/Object success at 1181 ms p99. Requested-FP16 and requested-INT8 ONNX reduce tether-inspect p99 to 601 ms and 532 ms, while Spatial success falls to 41.0% and 40.0% and Object remains at 89.0%. A graph audit shows those artifacts are byte-identical FP32 graphs, so the requested-INT8 row is not operator-level INT8 quantization. A static language-width ablation (16/24/32 tokens) yields Spatial success of 41.0%, 75.0%, and 71.0%; widths 24 and 32 recover much of the Spatial drop while Object success and uniform-bench latency stay approximately stable. Width-24 ONNX Spatial success is comparable to the PyTorch+AMP baseline at roughly half the latency (Wilson intervals overlap; two-proportion chi-squared p=0.53). Context width is an important contributor in this stack; it does not account for every PyTorch-vs-ONNX difference. Deployment evaluation should jointly report latency, artifact inspection, interface constraints, and closed-loop success. Code: https://github.com/rafiqul713/smolvla-libero-onnx.
Integrating contact information into visuomotor policies remains an open problem. Touch is essential to robust manipulation, yet most modern policies, including pretrained vision-language-action (VLA) models, operate from vision and proprioception alone. Existing approaches to closing this gap require specialized tactile hardware, add separate tactile encoders, or commit to non-image policy backbones, all incompatible with the modern paradigm of image-conditioned policies built on pretrained 2D visual representations. Our key insight is that the bottleneck is not the contact information itself, but how it is delivered: when contact signals are exposed in the same spatial frame as the scene the policy already attends to, they become directly usable by any image-conditioned policy without architectural changes. We operationalize this insight in Visible Touch, paired with a custom low-cost magnetic contact sensor that is open-sourced and fabricated from off-the-shelf parts via a parametric CAD-to-mold pipeline. Across the LIBERO benchmark, Visible Touch improves BC-Transformer success by 15.7 percentage points on average in the 2-view setting, with similar gains in the 1-view setting; controlled comparisons show that the contact-integration strategy strongly affects how effectively tactile information is used. The pattern holds when fine-tuning pretrained VLAs: miniVLA on LIBERO gains 25 percentage points on average, and $\pi_{0.5}$ on four real-world contact-rich tasks gains 30 percentage points with our custom sensor.
GIFT: Glove-Inferred Force Transfer: Force-Aware Human-to-Robot Skill Transfer from a Wearable Sensing Glove to a Robot Hand Without Tactile Sensors
GIFT:手套推断力传递——从可穿戴传感手套到无触觉传感器机器人手的力感知人类到机器人技能迁移
Sarusi, Tzah
Abstract
Human-to-robot skill transfer from sensing gloves has so far relied on shared hardware: the same tactile glove worn by the demonstrator and the robot, or a learned alignment between two tactile sensors. We present GIFT (Glove-Inferred Force Transfer), a pipeline in which the interface between human and robot is a physical unit rather than a shared sensor: fingertip force is measured in newtons on the human side and estimated in newtons on the robot side. A wearable glove records finger flexion, calibrated fingertip force, and wrist orientation, while a head-mounted camera records the demonstration; no robot is present. At deployment, the robot estimates force from actuator-current residuals relative to a free-space baseline, through a calibrated mapping to newtons, so any position-controlled hand that reports motor current can serve as the deployment platform. The policy uses a glove-space state and predicts finger-position targets; the robot enters only through two calibrated adapters, a retargeting decoder and a force estimator. We evaluate GIFT on a cup grasp-and-hold task with two action-chunking policies trained on the same demonstrations, with fingertip-force inputs retained in one and zeroed in the other. In a 50-rollout evaluation with sample size and metrics fixed before scoring, both policies succeeded in all 25 rollouts. The median of the per-rollout hold-phase grip-force estimates was 53% lower with force inputs: 1.20 N versus 2.55 N (one-sided Mann-Whitney U, p<0.0001). In an observation ablation, a vision-only policy achieved 0/15 grasps, policies given hand-command state acquired the grasp, and the force inputs determined how hard the policy held. A force channel measured on the human hand thus transfers to a robot hand with no tactile hardware, through a retargeting map from five glove channels to seven robot actuators, with no sensor shared between the two.
Chinese Translation
迄今为止,从传感手套进行的人到机器人技能迁移一直依赖于共享硬件:演示者和机器人佩戴同一只触觉手套,或两个触觉传感器之间学习得到的对齐。我们提出 GIFT(Glove-Inferred Force Transfer,手套推断力传递),其中人与机器人之间的接口是一个物理单位而非共享传感器:人类侧以牛顿为单位测量指尖力,机器人侧以牛顿为单位估计指尖力。可穿戴手套记录手指弯曲、校准后的指尖力和手腕朝向,头戴式相机记录演示过程;过程中没有机器人。部署时,机器人根据相对于自由空间基线的执行器电流残差,通过校准映射到牛顿来估计力,因此任何能报告电机电流的位置控制手都可以作为部署平台。该策略使用手套空间状态并预测手指位置目标;机器人仅通过两个校准适配器接入,即重定向解码器和力估计器。我们在一个杯子抓取并保持任务上评估 GIFT,使用两个在同一组演示上训练的动作分块策略,其中一个保留指尖力输入,另一个将指尖力输入置零。在一次包含50次 rollout 的评估中,样本量和指标在评分前固定,两个策略在各自的全部25次 rollout 中均成功。每次 rollout 保持阶段握力估计的中位数在使用力输入时低53%:1.20 N 对 2.55 N(单侧 Mann-Whitney U 检验,p<0.0001)。在一项观测消融中,仅视觉策略实现了 0/15 次抓取,给定手部命令状态的策略成功抓取,而力输入决定了策略握持的力度。因此,在人类手上测量的力通道通过一个从五个手套通道到七个机器人执行器的重定向映射,传递到无触觉硬件的机器人手,二者之间不共享任何传感器。
This paper presents the development and initial testing of DreamSat-Bench, a modular rendezvous and proximity operation testbed designed to benchmark AI-based relative navigation techniques. By integrating a software- and hardware-in-the-loop robotic pipeline, the platform enables a seamless transition from digital simulation to physical reality. DreamSat-Bench unifies state-of-the-art robotic learning tools such as MuJoCo, Isaac Lab, and LeRobot into a single benchmarking platform, utilizing robotic arms to trace 3D trajectories. The platform allows for extensive customization of orbital environments and lighting to evaluate the simulation-to-reality gap. We demonstrate the testbed's utility by evaluating an end-to-end vision-based navigation pipeline that pairs DreamSat, a generative AI framework for single-view 3D reconstruction, with FoundationPose for zero-shot 6-DoF tracking of unseen spacecraft. Initial testing explores mission-representative orbital segments, including fixed-point station-keeping and fly-around characterization. Through a series of parametric studies, we quantify the impacts of reconstruction latency, mesh resolution, orbital range, and illumination geometry on pose estimation accuracy. Finally, a preliminary hardware-in-the-loop campaign qualitatively validates the physical deployment of the pipeline, identifying target symmetry and accumulated tracking drift as critical factors for robust navigation. DreamSat-Bench provides a rigorous framework for maturing autonomous navigation with unprepared space assets in the absence of prior geometric models.
Surgical tissue manipulation demands precision; however, tool-tissue manipulation force magnitudes under realistic conditions are rarely quantified. To address this gap, we proposed and validated a portable ex-vivo force-sensing platform that measures tool-tissue interaction forces across the skull-brain interface during simulated neurosurgery. The system involves fresh calf brain tissue, used as a biological surrogate for brain parenchyma, placed in a 3D-printed human skull model equipped with a 6 degree-of-freedom force/torque sensor and a real-time data acquisition system. Five validation protocols assessed the accuracy and dynamic fidelity of the platform against ground-truth measurement, static accuracy and linearity using calibrated weights (0.5-50 g), minimum detectable force, spatial consistency across different anatomical regions, effect of surgical draping, and long-duration stability. Across protocols, measured forces showed excellent agreement with reference loads (correlation R = 0.9997), with RMSE < 0.005 N and mean relative error under 2%. The platform reliably detected low-magnitude forces down to 1 g (9.8 mN), while surgical drapes introduced no meaningful signal distortion and prolonged recordings exhibited minimal drift. Overall, the proposed framework provides objective, high-fidelity force quantification for skill training and performance assessment using fresh calf brain tissue and may serve as a foundation for force-based evaluation across other surgical procedures. Future work will integrate clinically used surgical instruments to increase procedural realism and will progress toward clinical trials to evaluate usability, educational impact, and translational relevance in practice-adjacent settings.
Combined Target Assignment and Path Finding (TAPF) requires assigning targets for agents while simultaneously planning collision-free paths. We present ITA-LaCAM, a complete and scalable TAPF solver inspired by LaCAM and ITA-CBS. In ITA-LaCAM, each joint-configuration node carries an agent-to-target matching. When a successor is generated, ITA-LaCAM incrementally repairs the matching for the agents that moved and uses the targets to guide PIBT successor generation. This design enables adaptive reassignment without explicitly enumerating the combinatorial assignment space, while preserving LaCAM's completeness and scalability. Across 9,760 benchmark instances on eight maps with 5--200 agents, ITA-LaCAM solved 100% of the instances, compared with 95.6% for IR-TAPF configured with DBS-Hungarian. ITA-LaCAM found an initial solution faster in 84.0% of the comparisons and achieved a lower sum of costs in 65.0% of the instances solved by both methods.
High-mix low-volume (HMLV) manufacturing requires inspection systems to adapt to changing parts, specifications, and work orders without repeated task-specific programming. Existing inspection automation typically assumes predefined sensing sequences, while general purpose robot agents optimize task completion rather than the completeness and validity of metrological evidence. We formulate task-specified active metrological inspection and propose From Requirements to Admissible Metrological Evidence (FRAME), a hierarchical dual-arm framework that converts an inspection instruction and structured specification into traceable conformance evidence. FRAME coordinates learned manipulation with calibrated laser profilometry: a task manager grounds and schedules requirements, active surface correspondence verifies physical-to-specification localization, and evidence memory tracks measurement provenance, admissibility, and coverage. Learned components may propose inspection targets and physical access actions, but deterministic datum-grounded measurement, admissibility checks, coverage auditing, and conformance evaluation prevent incomplete or unverified evidence from authorizing PASS. A series of physical experiments shows that FRAME achieves higher end-to-end inspection reliability, fewer false accepts, and shorter task completion time.
VGFM: Expressive Robot Policies via Dense Value Guidance in Flow Matching
VGFM:通过流匹配中的密集价值引导实现富有表现力的机器人策略
Koirala, Prajwal, Campbell, Mark
Abstract
Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative models enable rich and multimodal action representations, expanding the capability of this paradigm for complex robotic control. However, policy improvement with multi-step generative actors remains challenging. In offline reinforcement learning (RL), incorporating value-based objectives along generative trajectories often introduces substantial training complexity, including backpropagation through time (BPTT), auxiliary architectures, or distillation losses. We propose Value-Guided Flow Matching (VGFM), a scalable offline RL framework that enables dense value-guided shaping within a flow-based policy while avoiding BPTT and additional algorithmic overhead. VGFM parameterizes the policy as a conditional flow-matching model in action (x-prediction) space, ensuring that each intermediate flow step produces a valid robot action that can be directly evaluated by a standard offline RL critic. This design allows value guidance to be applied at randomly sampled flow times without differentiating through the entire generative trajectory, while preserving inference-time flexibility by varying the discretization of the underlying flow ODE without retraining. Evaluated on robotic locomotion and manipulation tasks in OGBench, VGFM achieves strong performance across a wide range of tasks under rigorous evaluation protocols. With minimal hyperparameter tuning, these results demonstrate that VGFM provides a simple, scalable, and effective approach for expressive policy learning in long-horizon, goal-oriented robotic control.
Learning Communication-Conditioned Generative Policies for Decentralized Multi-Agent Collision Avoidance
学习面向去中心化多智能体避碰的通信条件生成策略
Koirala, Prajwal, Campbell, Mark
Abstract
In this work, we propose a decentralized communication-conditioned generative framework for multi-agent collision avoidance. Agents generate short-horizon action sequences using a flow-matching policy trained from privileged offline demonstrations with access to global state. The demonstrations do not include explicit communication signals; instead, agents learn to exchange and aggregate latent messages that encode interaction-relevant intent under partial observability. This formulation supports flexible inference at test time, where unconditioned generation corresponds to independent behavior and communication-conditioned generation enables coordinated interaction without centralized planning. The resulting policies operate in a fully decentralized manner at execution time, relying only on local observations and learned messages. Combined with a receding-horizon inference scheme, the proposed approach enables efficient single-step inference of short-horizon action sequences and degrades gracefully under communication dropouts. Extensive simulation results demonstrate near-expert collision avoidance performance and strong generalization to denser, unseen multi-agent scenarios, along with zero-shot transfer to real-robot experiments.
Multi-Task Visual Perception Network with LLM Conditioning for Autonomous Navigation
面向自主导航的LLM条件化多任务视觉感知网络
Kumar, Praveen, Guruprasad, K. R., Sandhan, Tushar
Abstract
Long-term navigation for service robots faces crit- ical challenges like the accumulation of odometry drift and sensor error, which progressively degrade 2D maps and renders traditional path planning algorithms (e.g., A*, RRT*, DiPPer, ViT-A*) ineffective over time. To address this, we propose a user-friendly, interactive framework that eliminates the reliance on globally consistent maps. Our approach integrates visual perception with Large Language Models (LLM) to interpret user commands via text or voice. Instead of relying on a drift- prone global map, the system generates a sequential action plan based on local visual cues and egocentric geometric instructions. These action plans are executed sequentially, allowing the robot to navigate known and unknown environments safely. By reset- ting localization relative to immediate targets, our framework effectively works with a minimum accumulation drift strategy, ensuring accurate, efficient, and collision-free navigation without the maintenance overhead of traditional mapping. Experiments on real-world and simulated data have shown significant improve- ments over other methods. Our source code is publicly accessible at https://github.com/PraveenSingh24/VL-Navigation.
Generalizable bimanual robotic manipulation requires a reusable task prior that can persist across increasingly diverse tasks, objects, scenes, embodiments, and execution conditions, thus avoiding the prohibitive cost of large-scale teleoperated demonstrations and policy retraining. In this work, we present VLBiMan++, an extended framework that expands the generalization boundary of vision-language anchored one-shot bimanual manipulation. Starting from a single human demonstration, VLBiMan++ performs task-aware decomposition to identify reusable and adaptable skill components, and employs vision-language grounded geometric adaptation to transfer these skills to novel configurations without retraining. Building on this foundation, we systematically extend generalization along five dimensions: task generalization through diverse and long-horizon skill compositions; object generalization across unseen categories, varying geometries, and more complex articulated or deformable objects; scene generalization under clutter, occlusion, and dynamic interference; embodiment generalization across heterogeneous dual-arm robotic platforms; and deployment generalization through prolonged closed-loop execution under repeated external perturbations. To support this broader scope, we further introduce object-state-aware adaptation and lightweight trajectory optimization mechanisms that accommodate changes beyond simple rigid 6-DoF pose variations while preserving reliable bimanual coordination. Extensive real-world experiments demonstrate that VLBiMan++ maintains strong task success and adaptation capability across these increasingly challenging settings. Overall, VLBiMan++ advances one-shot bimanual manipulation from demonstrating isolated transferability toward a more systematic and scalable framework for generalization across tasks, objects, scenes, embodiments, and long-term deployment conditions.
BIG-CBF: Behavior-Imagination-Guided Control Barrier Function with Shared Uncertainty for Mobile Robot Navigation
BIG-CBF:行为想象引导的共享不确定性控制屏障函数用于移动机器人导航
Li, Shibo, Wang, Zhongcheng, Cao, Jiahe, Yang, Jianhua, Wu, Ke
Abstract
Control barrier functions (CBFs) provide a mathematically grounded framework for enforcing local collision-avoidance constraints in autonomous mobile robots, commonly through optimization-based safety filters. However, a minimum-intervention CBF filter lacks task-level maneuver awareness and may fail to select a productive avoidance direction when multiple distinct maneuvers are locally viable, leading to safe but stalled behavior in geometrically ambiguous environments. This paper presents BIG-CBF, Behavior-Imagination-Guided Control Barrier Function with shared uncertainty, a two-rate navigation architecture that separates low-rate maneuver selection from high-rate safety filtering. Over a short horizon, six closed-loop feedback behaviors are imagined and evaluated using analytic CBF compatibility together with a lightweight objective accounting for task progress, freezing, smoothness, and switching. To reduce planning-execution mismatch, the imagination and execution layers share consistent uncertainty sources for relative-motion delay, obstacle prediction, zero-order-hold motion, and command-execution residuals, while a hard CBF remains the final safety authority. In a 3,600-episode comparative benchmark across nine scenarios, BIG-CBF achieves the highest overall task success rate of 99.78% while substantially reducing downstream CBF intervention. On a physical omnidirectional robot with onboard Jetson Orin Nano computation, BIG-CBF completes all 15 evaluation runs without a recorded contact event. Matched hardware comparisons against the non-shared variant further show lower CBF intervention energy and activation frequency, supporting improved consistency between maneuver selection and safety-critical execution.
Chinese Translation
控制屏障函数(CBFs)为在自主移动机器人中强制执行局部避碰约束提供了一个有数学依据的框架,通常通过基于优化的安全滤波器来实现。然而,最小干预的CBF滤波器缺乏任务级机动意识,并且当多个不同的机动在局部可行时,可能无法选择有效的避让方向,导致在几何模糊环境中出现安全但停滞的行为。本文提出了BIG-CBF,即行为想象引导的共享不确定性控制屏障函数,一种双速率导航架构,将低速率机动选择与高速率安全滤波分离。在短时域内,想象六种闭环反馈行为,并使用解析CBF兼容性以及考虑任务进度、冻结、平滑度和切换的轻量级目标函数进行评估。为了减少规划-执行不匹配,想象层和执行层共享一致的不确定性来源,包括相对运动延迟、障碍物预测、零阶保持运动和命令执行残差,同时硬CBF仍然是最终的安全权威。在跨越九个场景的3600回合对比基准测试中,BIG-CBF实现了最高的总体任务成功率为99.78%,同时大幅减少了下游CBF干预。在配备机载Jetson Orin Nano计算的全向物理机器人上,BIG-CBF完成了所有15次评估运行,没有记录到接触事件。与非共享变体进行的匹配硬件比较进一步表明,CBF干预能量和激活频率更低,支持了机动选择与安全关键执行之间一致性的提高。
Orientation Control of Soft Robots via Adiabatic Spectral Submanifolds
基于绝热谱子流形的软体机器人方向控制
Karakai, Aron, Kaundinya, Roshan S., Michelis, Mike Yan, Katzschmann, Robert, Haller, George
Abstract
Soft robots are commonly sought for safety-critical interactions in delicate environments, where accurate position and orientation control is imperative. Model predictive control (MPC) offers a solution, but it requires a model of the robot's infinite-dimensional nonlinear dynamics that is at once accurate and computationally cheap. Recent theory on adiabatic spectral submanifolds (aSSMs) and their applications to soft robots provide data-driven model-reduction methods to construct such models. Here, we extend these methods to identify aSSMs from enlarged observable datasets and upgrade the currently available aSSM-MPC schemes. Evaluated on a high-fidelity finite-element simulation of a pressure-actuated soft arm, our controller reduces position and orientation tracking error by more than 60% compared to existing data-driven baselines.
Learning-Based Dynamic Obstacle Avoidance for a UAV Using Only Three Range Sensors
基于学习的仅使用三个测距传感器的无人机动态避障
Divkoti, Mohammad Reza Ranjbar, Aguiar, A. Pedro
Abstract
We present a learning-based approach to kinodynamic online motion planning for an Unmanned Aerial Vehicle (UAV) operating at a fixed altitude in unknown dynamic environments, where real-time avoidance of both static and dynamic obstacles must be achieved under conditions of extreme partial observability. The UAV is controlled with a single degree of freedom (yaw only), resulting in constrained, nonholonomic motion similar to fixed-wing platforms. The proposed framework integrates a behavior grid map representation with Deep Reinforcement Learning (DRL), using Proximal Policy Optimization (PPO) for stable policy learning in continuous control. The key idea is the co-design of a state representation and control policy that enables reliable navigation using only three low-cost directional range sensors, without reliance on dense sensing modalities such as LiDAR or vision-based systems. The behavior grid map dynamically aggregates sparse measurements into a structured local representation that supports real-time decision-making for obstacle avoidance and target reaching. Extensive simulations across environments of varying sizes and obstacle densities demonstrate that the proposed standard and enhanced methods achieve higher success rates than PPO variants and Model Predictive Control (MPC) (94\% vs. 79--90\% in small-scale high-congestion scenarios, and 83\% vs. 62--71\% in large-scale high-congestion scenarios), while maintaining real-time performance. Real-world experiments across four scenarios further confirm practical feasibility, with consistent target-reaching behaviour and no collisions under the tested conditions.
Chinese Translation
我们提出了一种基于学习的方法,用于在未知动态环境中以固定高度运行的无人机(UAV)的运动动力学在线运动规划,其中需要在极端部分可观测的条件下实现静态和动态障碍物的实时避障。无人机仅使用一个自由度(仅偏航)进行控制,导致类似固定翼平台的受限非完整运动。提出的框架将行为网格地图表示与深度强化学习(DRL)相结合,使用近端策略优化(PPO)在连续控制中进行稳定的策略学习。关键思想是状态表示和控制策略的协同设计,仅使用三个低成本定向测距传感器即可实现可靠导航,而不依赖于激光雷达或基于视觉的系统等密集传感模式。行为网格地图将稀疏测量动态聚合成结构化的局部表示,支持用于避障和目标到达的实时决策。在不同规模和障碍物密度的环境中进行的广泛仿真表明,提出的标准方法和增强方法比PPO变体和模型预测控制(MPC)具有更高的成功率(在小规模高拥堵场景中为94% vs. 79-90%,在大规模高拥堵场景中为83% vs. 62-71%),同时保持实时性能。在四个场景中的真实世界实验进一步证实了实际可行性,在测试条件下具有一致的目标到达行为且无碰撞。
Existing humanoid locomotion systems primarily focus on stability and task execution, while integrating expressiveness with explicit locomotion control remains challenging. We propose EMoG, an emotion-modulated gait generation framework for expressive humanoid locomotion. EMoG introduces an emotional-style code with continuously adjustable intensity. Conditioned on this code and physical commands, a lightweight MLP generates expressive, command-consistent periodic gait trajectories in real time, which are tracked by a unified reinforcement learning policy for physical execution. To support training, we collect a large-scale emotion-annotated gait dataset from professional performers and develop an automated pipeline to extract physically consistent periodic gait cycles. EMoG also integrates an LLM-based parser that converts free-form language into emotional style and motion parameters for interactive control. Experiments demonstrate continuous gait-style modulation with perceptible expressive cues while maintaining command tracking. EMoG provides a practical approach to parameterized emotional-style walking for human-robot interaction.
EcoBoat: Design and Experimental Validation of an Autonomous Body-Board Boat For Cleaning Water Bodies
EcoBoat:用于清理水体的自主板体船的设计与实验验证
Ansari, M. Aman, Khan, Saifullah, Kulkarni, Rahul, Sujit, PB
Abstract
Cleaning water bodies such as swimming pools and lakes typically demands significant manual effort or reliance on costly, sensor-intensive robotic systems. This paper presents EcoBoat, a low-cost autonomous surface vehicle built on a modified hull, designed to collect floating debris in both indoor and outdoor water bodies. For indoor environments, EcoBoat uses ultrasonic sensors to detect boundaries and obstacles, combining random-walk motion with boundary-following behavior. For outdoor environments, it relies on GPS-based geofencing paired with a random-walk strategy for area coverage. A key design feature is a Tesla-valve-inspired collection basket that allows debris intake during forward motion while preventing its escape during turning or braking. Field experiments in pools and lakes validated the design through iterative refinement, demonstrating that a minimalist sensing and computation approach can achieve effective, versatile debris collection across diverse water bodies.
Fault Diagnosis for Underwater Vehicles using Moving Horizon Estimation and Gaussian Processes
基于移动 horizon 估计与高斯过程的水下航行器故障诊断
Panetsos, Fotis, Kyriakopoulos, Kostas J.
Abstract
This work proposes a model-based fault detection and diagnosis framework for underwater vehicles subject to actuator faults that explicitly accounts for the presence of unmodeled dynamics. To this end, a Moving Horizon Estimator (MHE) is developed to estimate the lumped disturbance, capturing both unmodeled and fault effects. Gaussian Processes (GPs) are employed to approximate the unmodeled dynamics, providing predictions of the corresponding mean and uncertainty across diverse operating conditions. During online operation, the residual between the MHE lumped disturbance estimate and the GP prediction is evaluated using a Generalized Likelihood Ratio Test. By incorporating GP-based predictions within the diagnostic framework, robustness to unmodeled dynamics is achieved, enabling effective fault detection and isolation as well as accurate quantitative estimation of fault magnitude. The proposed methodology is experimentally validated in a laboratory water tank, demonstrating reliable diagnostic performance under both open-loop and closed-loop control.
Chinese Translation
本文提出了一种基于模型的水下航行器故障检测与诊断框架,针对执行器故障,并明确考虑了未建模动态的存在。为此,开发了一个移动 horizon 估计器 (MHE),用于估计集总扰动,同时捕获未建模动态和故障的影响。采用高斯过程 (GPs) 来近似未建模动态,提供在不同操作条件下相应的均值和不确定性的预测。在线运行期间,使用广义似然比检验评估 MHE 集总扰动估计与 GP 预测之间的残差。通过在诊断框架中融入基于 GP 的预测,实现了对未建模动态的鲁棒性,从而能够进行有效的故障检测与隔离,并准确量化估计故障幅值。所提出的方法在实验室水箱中进行了实验验证,在开环和闭环控制下均展示了可靠的诊断性能。
Mobile robots typically rely on geometric maps for obstacle avoidance and path planning, but the resulting obstacle representation does not always match how an object should affect navigation. A low lying cable may be missed, a flexible curtain may create spurious blockage, and a traffic cone may require an exclusion region larger than its observed footprint. We present NavPatch, an object level correction layer that assigns ADD, REMOVE, or EXTEND to navigation relevant object categories through periodic scene understanding with a vision-language model. Open vocabulary grounding localizes object instances, and LiDAR and RGB-D observations provide 3D support. Observation quality filtering and cross frame maintenance determine when each correction patch is committed, replaced, or revoked. In 50 real robot trials across five layouts, NavPatch achieves an overall success rate of 86.0%. An ablation study of four configurations with 200 runs in total shows that NavPatch improves the success rate from 70.0% to 86.0% and reduces the false commit rate from 68.4% to 40.7% compared with updates based only on the current observation.
Language-Grounded Semantic Target Navigation for Autonomous Surface Vehicles
面向自主水面艇的语言接地语义目标导航
Lin, Yuqing, Kim, Youngroung
Abstract
Autonomous Surface Vehicles (ASVs) are increasingly expected to operate in ports and harbour environments, where operators may specify navigation targets through language-based descriptions rather than predefined coordinates or fixed target identifiers. However, existing ASV navigation methods mainly execute predefined geometric goals or task-specific objectives and give limited attention to language-grounded target specification. This study proposes Semantically Grounded Navigation (SGNav), a framework that enables an ASV to identify and approach a maritime target from an operator-provided description. SGNav integrates text-guided semantic grounding, harbour-aware candidate filtering, CLIP-based semantic verification, grounded target control-state construction, and Proximal Policy Optimisation-based closed-loop control. It grounds the target description in onboard RGB observations, suppresses visually or semantically irrelevant distractors, and converts the selected target into a compact control-oriented representation for policy execution. Experiments in simulated port environments show that SGNav achieves success rates of $97.0\pm1.2\%$, $92.0\pm1.5\%$, and $90.0\pm1.8\%$ across three representative target-reaching tasks, with semantic target accuracy above $97\%$ and wrong-target rates below $3\%$. SGNav also maintains $97.7$--$98.7\%$ success rates across held-out port layouts. In the Task~3 ablation study, removing harbour-aware filtering or semantic consistency reduces the success rate to $40.4\pm2.6\%$ and $50.4\pm3.1\%$, respectively. These findings demonstrate the importance of semantic grounding, harbour-aware filtering, and semantic verification for reliable language-grounded ASV navigation. These results indicate that the proposed perception-to-control framework can support language-grounded target approach manoeuvres of ASV under the complex port environments.
Active exploration and semantic navigation require an embodied agent to build memory from partial observations, predict how the evolution of observed spatial memory may support future motion, and convert that prediction into actionable plans. We present GLAM, a goal-conditioned latent world model trained over global spatiotemporal memory, and GLAM NAV, the complete navigation system built around it. Given historical map tokens, a navigation goal, and the current robot pose, GLAM jointly predicts future map representations and robot-centric waypoint latents, allowing future spatial context and navigation intent to be inferred in a shared representation space. The model follows a JEPA-like latent prediction paradigm, operates directly on map-level latent tokens rather than RGB reconstruction, and uses a pretrained waypoint encoder-decoder to supervise and decode navigation plans within GLAM NAV. Training data are collected by replaying ObjectNav expert trajectories in Habitat over HM3D v0.2 scene assets and slicing them into multi-timescale prediction samples. On a controlled HM3D-ObjectNav subset reproduction setting, GLAM NAV improves over a reproduced BSC-Nav baseline in both success rate and success weighted by path length.
Learning Multi-Agent Task Assignment and Navigation in the Factory: from Simulation to Real Robots
在工厂中学习多智能体任务分配与导航:从仿真到真实机器人
Abdalwhab, Abdalwhab Bakheet Mohamed, Beltrame, Giovanni, St-Onge, David
Abstract
Reinforcement learning (RL) has shown considerable promise for robotic decision-making, yet deploying multi-agent RL (MARL) on physical multi-robot systems in industrial environments remains challenging. This paper investigates the real-world applicability of decentralized MARL for multi-robot multi-machine tending. We propose Feature-fusion Multi-Agent Proximal Policy Optimization (FMAPPO), which fuses 2D LiDAR measurements with task-specific state information to enable safe decentralized multi-robot task assignment and navigation. A complete simulation-to-reality pipeline was developed using high-fidelity robotic simulation and ROS2 and deployed on physical mobile-manipulator platforms operating under realistic real-world conditions, with the robotic arms disabled during the experiments. We further investigate the sensitivity of the learned policy to command update frequency, an important consideration for real-world deployment. Comparative evaluation in simulation demonstrated that FMAPPO significantly outperformed state-of-the-art baselines with a large effect size, achieving improvements of 106\% and 21\% in parts delivery and 48\% and 11\% in parts collection over MAPPO and SMAPPO, respectively. FMAPPO also increased machine utilization by 31 and 10 percentage points, respectively, while reducing collisions by 18\% and 15\% and increasing the safety score by 14 and 6 percentage points compared with MAPPO and SMAPPO, respectively. Furthermore, real-world experiments demonstrated that the learned decentralized policies can coordinate multiple robots to service multiple machines while maintaining safe operation under real-world sensing and control constraints. Videos of the real-world experiment are available online https://anonymouspapers123.github.io/FMAPPO/.
Imposing explicit trajectory constraints in robot visual servoing remains challenging. Existing tracking methods achieve fast responses by mapping visual residuals to control velocities, but they have weak constraints on the intermediate motion process, which lead to trajectory discontinuity, oscillation, or conservative behaviors. To enable constrained tracking for moving targets, this paper proposes a servo tracking method based on imitation trajectory constraints. A dynamic model describing the robot approaching a moving target is formulated and analyzed for convergence. A time-scalable deformation mechanism and a trajectory modulation incorporating shape and amplitude components are introduced to generate a series of trajectories in real time, from which tracking points are adaptively determined to form dynamic constraints. The robot velocity is then computed from target pose differentials or tracked key features to follow the constrained trajectory. Simulation and real-world experiments demonstrate that the proposed method can achieve dynamic obstacle avoidance and high-precision convergence compared with several state-of-the-art methods in complex environments.
Recent advances in data-driven robot manipulation policies have substantially improved task execution and generalization. However, real-world deployment still relies heavily on humans for failure assessment, correction, and environment reset, while models often fail to continually learn from failures and corrective experience. We present REVOLVE (Robot Evolving via Orchestrated Loops, Verification, and Experience), an automated closed-loop framework for evolving robot manipulation with minimal human intervention. Built on a unified software platform, REVOLVE integrates data collection, policy training and deployment, failure recovery, and continual learning into a single closed-loop workflow. Its Automated Reset and Collection (ARC) architecture automatically resets the environment and intervenes to correct policy failures. Dual-Loop Evolution (DLE) continually improves the manipulation policy and agent by feeding real-world interaction and failure--correction data back into policy learning and using an external mismatch memory to refine agent judgments. Experiments across four real-world manipulation tasks show that, after five iterations, REVOLVE improves average policy success rate by 18.5% and agent judgment accuracy by 8.5%, while reducing human effort in data collection and deployment testing by 94.4% and 95.1%, respectively. These results demonstrate that REVOLVE transforms real-world deployment into a closed-loop learning process that continually accumulates and uses execution experience, enabling continual evolution of both the policy and supervisory model with substantially less human intervention.
Chinese Translation
近年来,数据驱动的机器人操作策略取得了显著进展,大幅提升了任务执行和泛化能力。然而,真实世界部署仍然严重依赖人工进行失败评估、纠正和环境重置,而模型通常无法持续从失败和纠正经验中学习。我们提出了REVOLVE(Robot Evolving via Orchestrated Loops, Verification, and Experience),一种以最小人工干预进化机器人操作的自动闭环框架。REVOLVE构建在统一软件平台上,将数据收集、策略训练与部署、失败恢复和持续学习集成到一个闭环工作流中。其自动重置与收集(ARC)架构自动重置环境并进行干预以纠正策略失败。双环进化(DLE)通过将真实世界交互和失败-纠正数据反馈到策略学习,并使用外部不匹配记忆来优化智能体判断,从而持续改进操作策略和智能体。在四个真实世界操作任务上的实验表明,经过五次迭代后,REVOLVE将平均策略成功率提高了18.5%,智能体判断准确率提高了8.5%,同时将数据收集和部署测试中的人工工作量分别减少了94.4%和95.1%。这些结果表明,REVOLVE将真实世界部署转变为闭环学习过程,持续积累并利用执行经验,以显著减少的人工干预实现策略和监督模型的持续进化。
Robots, and humanoid robots in particular, are increasingly competent at individual behaviors, each obtained by training a specialized controller. A specialized skill is quick to train, converges reliably because the problem it faces is narrow, and can be validated on its own, none of which is true of a single end-to-end policy asked to cover everything. What remains fragile is the transition between them. We argue that the composition of independent sub-policies deserves to be treated as a research problem in its own right, rather than as an implementation detail left to whatever mechanism happens to be at hand. Reliable composition is what turns a collection of separate skills into a repertoire that can be used, extended and shared. More fundamentally, if control can be passed between specialized policies safely, and at any moment, the choice of what the robot should do next can be delegated to a component of an entirely different nature, such as a planner, an automaton or a symbolic controller, whose behavior can be inspected in advance. The policies would then only ever have to act, and what the robot can be trusted to do would become verifiable.
Interaction paradigms used in robot-assisted autism intervention have historically employed robots as teachers, clinical assistants, or more-abled peers to promote a variety of social skills. These modalities often leverage the expertise of trained practitioners to ensure that child-robot interactions are productive or clinically grounded to yield positive therapeutic benefits for children across the autism spectrum. Yet, despite the fact that the majority of children with autism spectrum disorder (ASD) attend mainstream schools and spend 80% or more of their time in the general classroom [27], there is a paucity of research incorporating validated classroom teaching pedagogies into robot-assisted autism interventions. In this work, we introduce a novel teaching methodology for advancing social skills in school-aged children with ASD. We evaluate the effectiveness of a novel robot-assisted autism intervention which incorporates the learning-by-teaching pedagogy and explores the comparative benefits of employing a robot versus a human confederate for improved performance on a set of social skills tasks. Results show that 80% of study participants performed better in the robot condition (mean performance in the robot condition=63%, mean performance in the confederate condition=37%), irrespective of the scenario order. Further, 90% of all participants were significantly more engaged in the robot condition (mean engagement: robot=61%, confederate=32%) and, while the effect did not result in the confederate condition, analyses indicate that overall engagement in the robot condition contributed to improved performance. These results suggest that robots employed in a learning-by-teaching context may help enhance engagement and improve performance on a simple social skills task for children with ASD.
Beyond Dead Reckoning: A Point of View on Camera--DAS--GNSS Continuity in Road Tunnels
超越航位推算:道路隧道中摄像头--DAS--GNSS连续性的观点
Saoud, Lyes Saad, Rakha, Hesham A., Jaber, Mona, Ayyash, Moussa
Abstract
Road tunnels remove satellite visibility where connected and automated vehicles still require continuous, attributable, and integrity-bounded positioning. This Point of View argues that tunnel localization should be treated as infrastructure-assisted cross-modal track continuity, not as extrapolation from the last trusted satellite fix. A trusted portal satellite solution provides the global anchor, distributed acoustic sensing (DAS) continuous motion evidence, cameras sparse identity and lane anchors, onboard sensing short-term dynamics, and edge computing association, fusion, integrity monitoring, and guarded reacquisition. Ground truth is reserved for offline calibration and validation. The article develops a falsifiable research and deployment agenda for progressing from synchronized multimodal evidence to validated, integrity-aware tunnel positioning continuity.
Reliable path tracking is crucial for autonomous underwater vehicles (AUVs) operating in dynamic and uncertain marine environments. However, traditional line-of-sight (LOS) guidance methods rely on asymptotic convergence, resulting in slow disturbance recovery and unpredictable tracking performance. Existing robust control methods typically require modifications to the underlying vehicle controller, limiting their practical application on commercial AUV platforms. This paper proposes a robust fixed-time adaptive LOS guidance framework for 3D path tracking for AUVs. By combining fixed-time stability theory with LOS guidance, this method guarantees path tracking convergence within a preset time range, with the convergence time independent of initial conditions. Furthermore, a fixed-time adaptive estimator is developed to rapidly compensate for time-varying sideslip disturbances caused by ocean currents. A time-varying look-ahead mechanism is also introduced to improve tracking performance on curved paths. Lyapunov analysis proves the fixed-time stability of the proposed framework, and numerical simulations and physical experiments demonstrate that, compared to state-of-the-art adaptive LOS methods, this framework exhibits superior tracking accuracy, convergence speed, and anti-interference capability. In simulation, the time-varying look-ahead variant reduced cross-track and vertical-track RMSE by 69.37\% and 67.46\%, respectively, during curved-path tracking. In field experiments with an Iver 3 AUV, the proposed fixed-time guidance reduced average tracking error by 56.35\% in straight-path evaluation and 27.59\% in curved-path evaluation compared with conventional adaptive LOS guidance.The proposed method provides a practical guidance-level solution for achieving reliable autonomous navigation of AUVs in complex marine environments.
A Personalized Dynamic Balance Evaluation Paradigm for Hip Exoskeleton-Assisted Walking under Unexpected Ground Perturbations
突发地面扰动下髋关节外骨骼辅助行走的个性化动态平衡评估范式
Chen, Yun, Akinniyi, Oluwasegun T., Zhang, Qiang
Abstract
Hip exoskeletons may improve recovery from unexpected gait perturbations, yet personalizing assistance remains difficult because balance is multidimensional and human-in-the-loop experiments are small-sample and noisy. We present a participant-specific composite balance cost that integrates seven biomechanical sub-metrics spanning margin of stability, center-of-mass dynamics, and whole-body angular momentum. The sub-metrics are converted to direction-aligned, dimensionless cost features, and nonnegative fusion weights are learned on the simplex. Coupled with an empirical-Bayes hierarchical model, the learned-composite selector estimates each tested condition's posterior probability of being best, P(best), and a high-probability candidate set with size $K_{0.8}$. The framework was evaluated with three participants walking at 1.1 m/s during unilateral belt-slip perturbations across 46 hip-assistance conditions. In the full-budget analysis (B = 4 repeats per condition), the selector concentrated 80% of the posterior probability within 1 to 5 of 46 conditions, compared with 2 to 12 for equal-weight fusion and 4 to 37 for principal component analysis fusion. This smaller candidate set could shorten personalization experiments and limit participants' exposure to repeated perturbations in future studies. Selected-condition trials showed lower observed composite costs than no-torque trials, with nominal p < 0.05 for P2 and P3. Leave-one-repeat-out refits yielded positive mean held-out rank correlations for all participants and moderate stability of the learned weights and candidate sets. These proof-of-concept results support participant-specific composite balance evaluation for candidate selection in perturbation-based human-in-the-loop experiments.
Robots need touch to manipulate objects safely and reliably, as many properties, such as softness, texture, and contact stability, are hard to infer from vision alone. However, vision-based tactile sensors yield different observations of the same material due to variations in optics, elastomer properties, and illumination, leading to poor generalization when trained on a single or multiple sensors. We propose a language-guided distillation framework for learning sensor-robust tactile representations. Language encodes high-level semantic properties of touch (e.g., rough, soft, slippery) that remain invariant across sensing hardware, providing a natural sensor-agnostic supervisory signal. We construct a 39K-sample touch-language dataset with human-annotated material labels and train a tactile encoder to align sensor-specific tactile images with language embeddings in a shared semantic space. We evaluate our approach for few-shot learning and cross-sensor transfer and benchmark it on six existing tactile datasets. Our method achieves 95% accuracy in the 100-shot setting, improves cross-sensor transfer by an average of 13.3% accuracy, and yields up to 19% accuracy gains across six existing tactile datasets. These results demonstrate that language-guided distillation enables scalable and hardware-agnostic tactile representation learning. Code and dataset are available at https://mashood3624.github.io/Language_Tactile/
Belief-Adaptive Online Autonomy for Quadrotor UAV Navigation under GNSS Degradation in Urban Environments
城市环境下GNSS退化时四旋翼无人机导航的信念自适应在线自主性
Panda, Deepak Kumar, Guo, Weisi
Abstract
Reliable online autonomy is critical for quadrotor operation in urban airspaces, where global navigation satellite systems (GNSS) measurements suffer from multipath, blockage, and latency issues, introducing non-stationary, temporally correlated errors that degrade conventional GNSS-IMU fusion. This paper presents a belief-adaptive online autonomy framework that augments an extended Kalman filter (EKF) with explicit GNSS trust modelling, second-order online belief adaptation, and latency-aware out-of-sequence measurement handling. GNSS trust is represented as a latent belief state that modulates measurement weighting and multipath bias uncertainty, and is updated online using EKF consistency signals. Unlike reactive covariance tuning, the proposed approach enables proactive and stable sensor trust adaptation without prior environmental knowledge or offline training. Evaluation in simulated urban air mobility scenarios with correlated multipath, stochastic latency, and obstacle constraints demonstrates improved belief convergence, smoother trajectories, and reduced estimation and tracking errors compared to naive, adaptive, and first-order baselines. The framework preserves classical GNSS-IMU fusion structure and can be integrated directly into existing flight control pipelines, supporting robust online autonomy in GNSS degraded environments.
We present a primitive-informed sampling-based model predictive control (MPC) framework for multi-fingered dexterous manipulation. Sampling-based MPC avoids the need for gradients through complex contact dynamics, but direct exploration of the high-dimensional joint space is inefficient and makes performance strongly dependent on the sampling distribution. Our framework biases sampling using low-dimensional manipulation primitives that encode coordinated finger motions, while simultaneously optimizing joint-level residuals to adapt these motions to the current hand-object configuration. Task-related rollout constraints reject infeasible trajectories during forward simulation, improving the effective use of the sampling budget. We evaluate the approach on a physical Allegro hand using a synchronized MuJoCo digital twin. Ablations show that both the primitive and residual are necessary for reliable continuous in-hand rotation, that increasing the sampling budget alone does not recover this coordination, and that rollout constraints substantially improve success rate. A primitive extracted from a single object remains effective across object sizes and under model mismatch, and the framework further supports grasping, object reorientation, and coordinated arm-hand reach-grasp-transport, using primitives extracted from both a simulation-trained policy and human hand-motion data.
Real-world reinforcement learning (RL) offers a promising route to dexterous manipulation policies that can adapt directly from physical interaction, but learning is hindered by inefficient early exploration and costly failures. We propose a framework that uses sampling-based model predictive control (MPC) as scaffolding for real-world dexterous RL, providing structured prior experience and task-directed guidance during learning without human demonstrations or corrective actions. A small set of MPC trajectories is first used to populate an offline replay buffer and to pretrain the actor and critic. During online learning, MPC intermittently guides data collection while an off-policy Soft Actor-Critic learner trains from both prior MPC experience and newly collected physical interaction, with control gradually transitioning to the learned policy. On continuous in-hand rotation with a 16-DoF Allegro hand, the method reaches 100\% success in policy-only evaluation (5/5 trials) after 7 minutes of online RL, following initialization with 20 MPC trajectories collected on hardware in 12 minutes. Online training incurs about three object drops on average. After 20 minutes of online learning, the policy achieves more than five times the rotation speed of the MPC controller. It completes 1000 consecutive rotations over more than 110 minutes without a drop. Ablations show complementary benefits from MPC-based pretraining, retained MPC experience, and online MPC guidance. We further demonstrate rapid adaptation to different object geometries and successful goal-conditioned reorientation, showing that the framework enables efficient, low-intervention, real-world dexterous RL.
HydroMap: Probabilistic Water Surface Elevation Mapping for Semantic Scene Representation in Inland Waterways
HydroMap:面向内河航道语义场景表示的概率水面高程建图
Luo, Zhongbi, Wang, Yunjia, Bruyninckx, Herman, Slaets, Peter
Abstract
Autonomous surface vehicles operating in inland waterways require a persistent representation of both surrounding structures and the water surface. LiDAR-based simultaneous localization and mapping often produces sparse or missing water returns, leaving this operational surface absent from the reconstructed scene. We propose HydroMap, an odometry-decoupled framework that reconstructs water surface elevation from stereo observations and integrates it with the structural map. Per-frame water points form joint cell observations with propagated stereo and pose uncertainty, and successive observations are fused into a persistent probabilistic elevation map. Semantic map conversion then combines the elevation map with structural geometry in a unified 2.5D representation of water, boundaries, structures, and overhead regions. On the Pohang Canal and Leuven Vaart datasets, the elevation RMSE remains below 5 cm relative to LiDAR references expressed in the same map frame. The elevation and semantic maps are published at 2 Hz and 1 Hz, respectively. HydroMap thereby complements LiDAR maps with a persistent representation of the water surface for downstream navigation in inland waterways.
Exact Feasibility Certification and Optimal Responsibility Allocation for Multi-Robot CBF Safety Filters
多机器人CBF安全滤波器的精确可行性认证与最优责任分配
Sah, Chandan Kumar, Keshavan, Jishnu
Abstract
Multi-robot Control Barrier Function (CBF) safety filters can become infeasible, but a failed quadratic program (QP) does not indicate why the conflict occurred or how to resolve it. To address this, we develop an exact feasibility certificate for multi-agent CBF filters with heterogeneous control-affine dynamics and convex input sets. The certificate quantifies a feasibility reserve by separating the demand imposed by safety constraints from the available actuator supply. This decomposition shows when CBF gain tuning or increased actuation can and cannot resolve infeasibility, and identifies the agents and interactions responsible for the conflict. We further propose an algorithm to optimally allocate shared safety constraints by maximizing the worst local feasibility margin, yielding a linear program for polyhedral input sets. In $320$ paired closed-loop simulations, the proposed allocation reduces infeasible control steps from roughly $50\%$ to $6.2\%$, and reduces safety-violating runs from $118/160$ to $24/160$. In addition, across $52$ infeasibility events, the certificate identifies an interaction whose relaxation restores feasibility in $94\%$ of cases.
Comparing Trajectories from Positions Alone: Curvature-Based Time Alignment and Drift Error Metric
仅从位置比较轨迹:基于曲率的时间对齐与漂移误差度量
Daum, Effie, De Martini, Daniele, Dune, Claire, Pomerleau, François
Abstract
In field robotics, acquiring independent large-scale reference trajectories more accurate than the evaluated estimates remains an open challenge. The domain is widely reliant on Absolute Trajectory Error (ATE) and Relative Pose Error (RPE), computed with automated tools, that rest on assumptions and evaluation parameters rarely made explicit. When unreported, the errors can be misleading and hinder fair comparisons. This paper introduces a trajectory-evaluation protocol for standardized and reliable accuracy assessment in state estimation, localization, and Simultaneous Localization And Mapping (SLAM). The approach combines a novel temporal alignment method based on curvature signals with an error metric normalized by travelled distance. We explicitly account for temporal synchronization, sampling alignment, and extrinsic calibration, quantifying their influence through a sensitivity analysis. The proposed protocol contributes to more rigorous, reproducible, and standardized trajectory evaluation.
Two-Stage Personalized Gait Phase Estimation in Stroke Survivors During Exoskeleton-Assisted Walking: An Offline Feasibility Study
外骨骼辅助行走期间卒中幸存者的两阶段个性化步态相位估计:一项离线可行性研究
Ryu, Hyungseok, Hur, Pilwon
Abstract
This study evaluated personalized gait phase estimation for stroke survivors using functional inertial measurement unit (IMU) alignment and two-stage sequential adaptation of models pre-trained on healthy gait. The estimator used signals from a thigh-mounted IMU. Heel force-sensitive resistor measurements provided reference phase labels for offline adaptation and evaluation. Stage 1 established a distillation-regularized participant-specific model, and Stage 2 performed conditional refinement using low-rank adaptation. Long Short-Term Memory (LSTM), Temporal Convolutional Network (TCN), and Transformer models were evaluated in five stroke survivors walking with a powered knee exoskeleton using leave-one-subject-out hyperparameter selection and sequential test-then-adapt Stage 2 replay. Relative to the non-adapted baselines, Stage 1+2 reduced the mean participant-wise phase root mean square error by 84.2%, 77.0%, and 60.7%, respectively. The Transformer achieved the lowest final error (2.90 +- 1.13$% of the gait cycle) and heel-strike timing error (23.7 +- 4.5ms). Policy-specific ablations showed that every-cycle updates generally produced the lowest or near-lowest error, whereas conditional updating reduced the update frequency with small accuracy differences. After personalization, alignment produced model-dependent changes in phase error while preserving or improving heel-strike detection and reducing heel-strike timing error for the LSTM and Transformer. Concurrent embedded tests showed that the TCN and Transformer maintained 100-Hz inference during Stage 2 updates without deadline misses, whereas the LSTM missed the 10-ms deadline in 6.6% of inferences. All updates completed within 0.8s. These results support the offline feasibility and embedded computational timing of the proposed framework for exoskeleton-assisted walking.
From Learned-Mode AV-Traffic Pairing to Planner Decisions: A Marginal-Preserving Study on Argoverse 2
从学习模式AV-交通配对到规划器决策:一项在Argoverse 2上的边缘保持研究
Wang, Jingyu
Abstract
Joint motion forecasts pair each autonomous-vehicle (AV) future with surrounding traffic, but actor-level metrics do not show whether that structure matters to a planner. We study this question with a marginal-preserving product control that removes AV-traffic pairing among the learned modes while retaining the fixed constant-velocity pair and holding trajectories, actor-level marginals before planner conditioning, candidates, the cost terms and weights, and fallback fixed. The intervention also changes candidate-conditioned concentration. Across twelve runs on 1,400 held-out Argoverse 2 scenarios, the intervention changes 3.0% of route-level offline selections at $\tau=4$ m. Control-minus-joint recorded-trajectory regret is $-0.026$ and $-0.118$ at the two training sizes; crossed and seed-$t$ intervals span zero. At $\tau=1$ m, relative costs change in 87.9% of route evaluations and route-level offline selections in 8.1%. Before concentration matching, descriptive outcome estimates favor the control. Most of this gap disappears along an approximate concentration-matching path; the remaining contrasts are $+0.112$ and $-0.047$, and both crossed intervals span zero. Actor-level forecast metrics remain identical. Pairing-strength and temperature sweeps show that the decision contrast grows with pairing removal and sharper conditioning. The intervention changes planner decisions even though actor-level metrics remain unchanged. The matching analysis, however, cannot separate any recorded-outcome effect of learned-mode AV-traffic pairing from the accompanying change in conditioned concentration.
IMPACT-VLA: Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies
IMPACT-VLA:通过反事实轨迹实现视觉-语言-动作策略的交互感知多模态传播归因
Kim, Jinwoong, Park, Sangjin
Abstract
Vision-Language-Action (VLA) policies perform robot manipulation tasks using multimodal inputs such as visual observations, proprioceptive states, and language instructions. However, it remains unclear at which execution stages each modality contributes to final task success and how input interventions propagate through subsequent states, observations, and actions. Existing attribution approaches primarily measure local sensitivity or temporally aggregated importance, limiting their ability to capture phase-dependent contributions and cross-phase dependencies. We propose Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies (IMPACT-VLA). IMPACT-VLA constructs behavioral phases from action transitions in a successful reference rollout, aligns them with policy query boundaries, and defines phase-modality blocks as attribution units. It then performs closed-loop counterfactual re-execution to quantify each block's contribution to final task success. We further analyze cross-phase non-additive interactions and trajectory propagation while distinguishing behavioral from functional recovery. Across 30 LIBERO robot manipulation tasks using OpenVLA-OFT, dominant-modality transitions occurred in 25 tasks (83.3%), and closed-loop attribution identified task-critical information more faithfully than Static Action Perturbation. Later-block marginal gains for negatively interacting pairs increased by approximately 3.3x under early-phase input replacement, while functional recovery could occur without behavioral recovery. These results reveal when multimodal inputs support task success and how their contributions become conditionally coupled during closed-loop execution.
Can changing only the language instruction redirect a VLA policy's end effector, or does the visually driven motion prior dominate? We present Atomic Motion Coordinate, a geometry-grounded coordinate for steerable and force-responsive manipulation. Each arm owns thirteen signed translation, rotation, and hold atoms grounded from text and forward kinematics with vision withheld, and the coordinate is injected into every action-expert block via weighted codebook alignment. Contact history modulates the same coordinate through a bounded spherical residual that is recomputed from a fixed nominal latent to regenerate only the unexecuted horizon suffix. Across 7,520 offline horizon interventions, opposite-atom separation reaches 92.5/83.1% (single/dual) versus 39.1/24.0% for LA4VLA-style. Across 50 real-robot trials per task, AMC raises OOD fruit progress from 60.5% to 87.8%; force adaptation raises Plug/Vase from 59.0/71.5% to 78.5/75.2%.
Steering Generative Robot Policies with Lexicographic Preferences
用字典序偏好引导生成式机器人策略
Jia, Yixuan, How, Jonathan P.
Abstract
Pretrained generative robot policies can produce effective behaviors across diverse environments, but deployment can lead to requirements and preferences that may not have been represented during training. Furthermore, at deployment, an operator, user, or application may assign these requirements and preferences a priority order that can vary across deployments. For example, embodiment-specific feasibility constraints may need to be satisfied first, while user-specific preferences guide behavior among the feasible options. We show that a frozen generative robot policy---based on either diffusion or flow matching---can be steered at inference time to respect such lexicographically ordered deployment objectives. To achieve this, we introduce two modifications to the sampler. First, we apply dynamic-barrier guidance to sampled trajectories, constraining lower-priority updates so that higher-priority costs do not increase (up to first order). Second, we select the executed sample using a cascade that successively filters candidate samples according to each priority level. The policy weights remain unchanged. On a navigation benchmark, we demonstrate that our method improves success, traversability, and preference compliance over the frozen policy, and achieves substantially better compliance than tuned weighted-sum baselines. The same method transfers to a flow-matching manipulation policy on LIBERO, where it improves compliance without reducing task success. A controlled manipulation study further shows that, in settings where a fixed weight can match the desired ordering, the dynamic barrier reaches comparable best performance over a substantially wider range of parameter settings.
Task-Distribution-Aware Counterweight Synthesis and Constrained Co-Design for Serial Manipulators
任务分布感知的配重合成与串联机械臂约束协同设计
Abbadi, Mohammad
Abstract
Passive counterweights are simple gravity compensators, but a counterweight selected from a single pose is not generally optimal for the configurations and tasks a manipulator actually executes. This paper develops a task-distribution-aware synthesis framework in which the operating distribution $\rho(q)$ enters the design explicitly. For a counterweight moment $p=m_c r_c$ with gravity torque $-gp\phi(q)$, the weighted mean-square residual gravity torque has the closed-form minimizer $p^*=E_\rho[\tau_g\phi]/(gE_\rho[\phi^2])$. If payload gravity torque is affine in payload mass, the optimum is also affine: $p^*(m_p,\rho)=p_0^*(\rho)+m_pK_p(\rho)$. For fixed static moment, added counterweight inertia is $I_c=pr_c$ while mass is $m_c=p/r_c$, so mass-radius selection is underdetermined unless physical constraints are specified. A recovered three-link manipulator is used as a case study. At $r_c=0.20$ m, zero-payload equivalent optima are 0.672 kg for uniform joint-space operation, 0.683 kg for approximately uniform task-space operation, 0.713 kg for a representative pick-and-place family, and 0.952 kg for a high-gravity-biased distribution, a change of more than 40% caused solely by the operating distribution. Nondominated fronts show that preferred mass-radius pairs depend on declared engineering bounds. A rated-torque-referenced all-joint screen increases zero-payload feasible task-space coverage from 78.1% without compensation to 93.7% for the uniform-distribution design. A lumped point-mass trajectory study gives a provisional crossover from no counterweight at very aggressive motion to stronger compensation as motion slows. These actuator and dynamic results are engineering consequence studies rather than physical validation.
Legislating World-Model-Based Planning with Legal Reasoning
基于法律推理的世界模型规划立法
Waldner, Dylan, Kantaros, Yiannis, Governatori, Guido, Miikkulainen, Risto, Banifatemi, Amir
Abstract
As robotic systems grow more general, legal norms are needed to integrate them into society. This paper extends the isomorphism problem of aligning legal source texts with their encodings, and measures two key challenges to robot normative control: (1) the \textit{grounding isomorphism gap}, where perception error grounds false atoms for legal reasoning, and (2) the \textit{ontological isomorphism gap}, where one legal conclusion admits many faithful translations into planning constraints. The paper introduces a legal planning stack that employs Defeasible Deontic Logic (DDL) to constrain a motion planner. The stack leverages learned world models to plan and to provide legal context, enabling \textit{ex ante} governance that intervenes before an illegal action is executed. It was deployed on a simulated robot arm pushing a cube across a $3\times3$ grid. The findings were (1) the legislated agent abided substantially more often than the non-legislated one, and modeling perception uncertainty lifted abidance even further, (2) the legal reasoning ran efficiently at runtime and its verdicts were auditable, and (3) the stack adapted to exogenous signals and endogenous rule changes. Both gaps were measured: (4) world model and probe error corrupted the factual input for the DDL reasoner, and (5) a single law admitted several faithful metric interpretations yielding drastically different abidance. Thus, \textit{ex ante} legislation functions as intended, and closing these gaps with a standardized mapping from the law to runtime constraints and improved fact grounding from perception will yield robust laws that align robot behavior with society's norms.
C$^2$Nav: Compare Before You Commit for Zero-Shot Vision-and-Language Navigation
C$^2$Nav:用于零样本视觉-语言导航的先比较后承诺
Zheng, Runtian, Zhang, Congpeng, Liu, Ying
Abstract
Zero-shot vision-and-language navigation in continuous environments (VLN-CE) increasingly places foundation vision-language models (VLMs) inside the navigation loop. Existing systems commonly request cardinal outputs such as waypoints, pixels, headings, progress values, or absolute arrival decisions, coupling a generative response to geometric magnitude or an irreversible commitment. We study a complementary model-robot interface: the VLM compares controller-constructed alternatives, while geometry, thresholds, action magnitude, and execution remain on the physical side. We instantiate this idea in C2Nav, a training-free framework with three coordinated faculties. Seeing performs ordinal Gaze Election over physically vetted candidate views; Remembering maintains a compact route sketch and compares adjacent instruction-leg hypotheses; and Arriving combines a hesitation ladder, look-back comparison, and revocable walk-back for reliable stopping. On the public OpenNav R2R-CE 100 protocol, C2Nav with Qwen3-VL-8B-Instruct obtains 41.0% OSR, 31.0% SR, and 16.7% SPL, while the same interface with the standard GPT-5.5 model reaches 54.0% OSR, 44.0% SR, and 29.0% SPL. Whole-faculty ablations reduce SR to 14.0% without Seeing, 25.0% without Remembering, and 29.0% without Arriving. Matched role inversions that replace only the comparative answer form with cardinal/absolute questions reduce SR to 12.0%, 28.0%, and 21.0% in the spatial, transition, and terminal slots, respectively. The results indicate that a constrained decision interface and stronger VLM reasoning are complementary rather than interchangeable.
LieSpline-DP: Lie-Group B-Spline Diffusion Policy for Smooth Robot Manipulation
LieSpline-DP:用于平滑机器人操作的李群B样条扩散策略
Xie, Erxuan, Liu, Bang, Nie, Pingyun, Liu, Xingkai, Fu, Zhuang, Zhang, Bo
Abstract
Diffusion Policy (DP) is a powerful Learning from Demonstration (LfD) method for robotic manipulation, yet it suffers from discontinuous and non-smooth trajectories. Spline-based action representations promote smooth motion within individual action chunks, but existing spline-based methods neither guarantee cross-chunk $C^2$ continuity nor account for the group structure of $\mathrm{SE}(3)$. We therefore propose LieSpline-DP, a Lie-group B-spline diffusion policy that generates end-effector trajectories directly on $\mathrm{SE}(3)$ and couples consecutive plans by sharing their boundary control poses, ensuring $C^2$ continuity throughout the entire planned trajectory. Across three real-robot tasks, LieSpline-DP produces lower trajectory jerk and higher task success rates than the DP baseline. The gains are particularly pronounced in real-world tasks involving liquids and flexible objects: in our real-robot experiments, LieSpline-DP achieved a 100% success rate on both pouring and bucket hooking, whereas the DP baseline achieved only 10% and 30%, respectively.
HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness
HarnessVLN:通过 Agent Harness 统一免训练具身导航
Chen, Yang, Che, Lirong, Huang, Zhenyu, Fu, Wenbo, Wang, Chuang, Cao, Xu, Liu, Daqi, Yang, Yuzhe, Su, Jian, Guo, Lan-Zhe
Abstract
Embodied navigation requires agents to interpret visual observations, accumulate spatial knowledge, and execute actions to follow instructions or locate objects. Training-based methods face generalization challenges, while training-free methods exploit multimodal large language models (MLLMs) but often lack mechanisms to reconcile proposed actions with spatial evidence, task progress, and execution failures. We present HarnessVLN, a zero-shot, training-free framework whose Agent Harness coordinates perception, retrieval, grounding, navigation, recovery, and termination through a unified tool interface. The Harness validates planner proposals against spatial evidence, geometric feasibility, and subgoal consistency, incorporating structured tool feedback into subsequent decisions. Hierarchical event memory tracks task progress and execution history, while a persistent Spatiotemporal Graph maintains reusable spatial evidence and failure annotations for verification and recovery. A replaceable Navigation Executor converts validated targets into executable motions, allowing the same Harness protocol to support instruction-following and object-goal navigation. HarnessVLN achieves success rates of 60.8%, 53.9%, 76.0%, and 59.3% on R2R, RxR, HM3D-v2, and HM3D-OVON, respectively, surpassing prior training-free SOTA results. Humanoid deployment further demonstrates its applicability to both tasks in real-world environments. The project page is: https://harnessvln.netlify.app/.
PredTac: Learning Contact-Rich Manipulation with Predicted Touch
PredTac:利用预测触觉学习富接触操作
Fan, Weijia, Guo, Daqiang
Abstract
Contact-rich manipulation benefits from tactile feedback, yet physical tactile sensors introduce hardware, calibration, synchronization, and maintenance costs that complicate policy learning and deployment. We formulate predicted touch as an alternative to measured tactile input and present PredTac, a framework that learns to infer tactile states from causal visual observations and robot states and uses the predicted touch as an explicit interface for policy learning and execution. A tactile predictor is first trained with tactile supervision and then used to provide contact information without requiring measured tactile input during downstream policy training or execution. We evaluate PredTac across three contact-rich manipulation tasks in simulation and on a real robot, and further examine how policy performance depends on the predicted contact content. In simulation goal-offset evaluations, predicted-touch policies achieve 27.0%, 52.0%, and 44.7% success on USB, Barbed-spike, and Valve, respectively, improving over the visual baseline by 8.0-13.7 percentage points. On the real robot, predicted-touch ACT achieves 70.0%, 50.0%, and 90.0% success on USB insertion, Barbed extraction, and Valve rotation, respectively, with a three-task mean of 70.0%, approaching measured-touch ACT at 72.2% and substantially outperforming visual ACT at 21.1%. Fixed-policy interventions further show that performance is sensitive to the spatial structure of predicted contact, with spatial rearrangement at fixed value distributions reducing Valve success by 10.7 percentage points. These results demonstrate that predicted touch can provide useful contact information for contact-rich manipulation without requiring tactile sensing as a policy input.
X-WBC: A Cross-Embodiment Foundation Model for Humanoid Whole-Body Control
X-WBC:一种用于人形机器人全身控制的跨本体基础模型
Zhang, Juntong, Gu, Chun, Zhang, Li
Abstract
Scaling humanoid whole-body control toward general-purpose deployment requires large human motion corpora and training experience shared across robot bodies. Existing methods usually train one policy per robot, leaving motion experience isolated across embodiments. We introduce X-WBC, a cross-embodiment foundation framework that separates relatively shared human motion semantics from embodiment-specific physical execution. Human-centered command tokens align full human motion, robot reference motion, and sparse VR observations. A causal Transformer learns reusable temporal structure from mixed multi-robot rollouts, while lightweight robot-specific modules map the shared representation to each robot's proprioception and action space. Across nine simulated embodiments, external motions, and four real robots, experiments show that joint training improves tracking, the aligned representation supports consistent control across command sources, and the learned policy remains competitive beyond the training corpus. These results support heterogeneous humanoids as joint data sources and establish cross-embodiment joint training as a practical route toward whole-body control foundation models.
The introduction of human-robot collaboration (HRC) in industrial assembly operations is revolutionizing the manufacturing landscape. In this evolving environment, operators are required to seamlessly coordinate their manual tasks with real-time task information and robotic behaviors. These demands fluctuate during operation, yet conventional workload assessments depend on body-worn physiological sensors that complicate practical deployment. Here, we present a vision-based attention--action framework for continuous and interpretable workload-related assessment in HRC assembly. The framework combines RGB-D observations with robot states and calibrated task-related areas to construct a temporally confirmed representation of operator behavior. This representation identifies where task demand is concentrated and explains how it develops when attention and action diverge, the task context changes, or the operator hesitates. We evaluated the framework in a three-level collaborative gearbox assembly experiment with ten participants, using subjective ratings and synchronized physiological signals as independent references. Raw NASA-TLX ratings confirmed increasing perceived workload across conditions, with significant effects on overall workload and its mental and temporal dimensions. The vision-derived HRC-CWL output was significantly associated with ECG-derived features in seven of nine participants with complete correlation data. Synchronized interaction episodes further showed temporal correspondence between detected hesitation and physiological activity. Real-time deployment demonstrated that the framework can operate without requiring operators to wear additional sensors. These findings support HRC-CWL as an interpretable behavioral proxy for cognitive ergonomics analysis and adaptive robot assistance, rather than a direct psychophysiological measure of workload.
Trajectory curvature constraints are inherent in practical multi-robot systems due to the limited turning capabilities of the robots. Without properly accounting for these constraints, robots may fail to accomplish assigned tasks, and their trajectories may diverge from the intended paths. This paper proposes a distributed safe cooperative vector field approach for multi-robot systems subject to trajectory curvature constraints. The proposed approach is composed of a cooperative vector field and a safety-oriented collision avoidance vector field, aiming to address the problems of cooperative motion and safe collision avoidance in multi-robot path-following tasks. A safety-oriented collision avoidance vector field with adaptively adjustable reactive boundary is developed to accommodate the kinematic curvature constraints of robots, thereby ensuring the physical feasibility of collision avoidance maneuvers. The proposed vector field requires only a single virtual variable from each neighboring robot to achieve cooperative motion and ensure both obstacle avoidance and inter-robot collision avoidance. The effectiveness of the proposed approach is validated through both simulations and real-world experiments on an actual multi-robot platform.
Low Clearance Hinge Joint Mechanism Based on 3D Printing on Sheet Fabrication Methodology
基于片材上3D打印制造方法的低间隙铰链关节机构
Jang, Jaehyung, Shin, Euibin, Okamura, Allison M., Ryu, Jee-Hwan
Abstract
This paper presents a low-clearance hinge joint mechanism based on the 3D printing on sheet fabrication method. This approach simplifies the fabrication of hinge mechanisms and overcomes limitations of conventional origami manufacturing by eliminating the need for adhesives commonly used during assembly, making it suitable for robots at the tens-of-centimeters scale. The advantages and disadvantages of three types of hinge joint mechanisms are compared, and a hinge joint that can be designed with low clearance for various facet thicknesses is selected. Based on the selected hinge joint, the twisting angle and bending force are analyzed, leading to the implementation of a clearance of 0.1 mm. Torsional resistance is experimentally evaluated to measure the torque required for twisting caused by plastic deformation and clearance. The results show that the torque associated with plastic deformation is sufficient to constrain the undesired degrees of freedom of the hinge joint, while the torque required for twisting due to clearance is minimal. Based on the analyzed data, the proposed hinge joint mechanism is applied to a 3-degree-of-freedom delta robot manipulator, demonstrating precise motion with low clearance.
Pretrained driving vision-language models (VLMs) integrate visual, route, language, and driving context into rich driving priors, yet their representation objectives remain separated from continuous driving planning. Existing methods typically begin trajectory generation only after the VLM has formed a final condition, leaving depth-wise condition computation outside the stepwise formation of trajectory state. We introduce DiffAdapterVLA, which realizes Planning in the Backbone: it injects explicit trajectory tokens into selected VLM late layers, bringing trajectory state into backbone forward computation, where it co-evolves with driving conditions at different depths. Lightweight layer-wise DiffAdapters organize this computation into recursive trajectory refinement, while asymmetric joint attention preserves directed guidance from the condition stream to trajectory planning. By placing planning within existing backbone computation rather than relying on an independent trajectory planner, DiffAdapterVLA adapts only lightweight trajectory modules to turn existing driving priors into efficient continuous planning capability. NAVSIM results show that it achieves high-quality closed-loop planning with low end-to-end latency using few trainable parameters, and demonstrate that jointly evolving trajectory state and depth-wise driving conditions in VLM late-layer computation effectively realizes continuous trajectory planning.
Ground-truth human joint torque estimation relies on motion capture systems, which suffer from limited outdoor usability and significant deployment expenses. Furthermore, direct scaling of ground-truth joint torques to obtain motor torque commands is not necessarily the optimal strategy. To address the aforementioned limitations, inspired by the human motion generation process, this paper proposes a novel assistance torque estimation method based on the dynamic model. From an optimization perspective, the proposed method directly generates motor-assist torque and lowers the cost of data acquisition. Then, a data-driven assistance torque prediction network is trained to enable accurate real-time prediction under complex outdoor environments. Experimental results demonstrate that optimized (estimated) assistance torque exhibits better phase consistency with gait trajectories and better alignment with task characteristics. Relative to the Zero torque condition, the predicted torque can decrease metabolic rate by 11.8%-17.7%, heart rate by 8.9%-14.3%, and peak muscle activation levels by 28.2%-54.0%, respectively. This provides a new perspective for low-cost adaptive exoskeleton assistance.
This letter investigates how the parameters of a slope-aware variable-admittance filter influence user preferences in force-based interaction with a robotic guide dog for visually impaired individuals. The proposed system consists of a quadruped robot equipped with a sensor-free rigid handle for physical guidance. The framework combines path following, momentum-based interaction-wrench estimation, and a variable-admittance filter whose stiffness and damping are adapted online from slope information extracted by the robot's depth camera. The adaptation policies are evaluated through high-fidelity simulations and a human-subject study involving blindfolded sighted participants. Multiple strategies are compared using a Taguchi L9 design of experiments. Preliminary main-effect results suggest that increasing stiffness uphill and decreasing it downhill improves both objective and subjective metrics, whereas damping shows no significant main effect.
Wheel-loader excavation is a sequential decision problem in which every scoop changes the terrain available to subsequent actions. A practical world model must predict action consequences accurately, rank candidates in real time, and operate inside the closed loop of a full-size machine. We present the World-Action Model (WAM), which proposes multiple scoops, rejects geometrically inadmissible candidates, jointly predicts signed terrain change and loaded volume, executes the candidate with the largest predicted load, and replans from the newly observed terrain. On 32 geometry-disjoint MinSlope test episodes, adding world-model ranking to matched diffusion proposals reduces the mean scoop count from 651.8 to 540.6 (17.1%), preserves 32/32 completion, and improves every paired episode. In a complete-system comparison, WAM completes 32/32 episodes versus 29/32 for an independently trained soft actor-critic policy. Comparisons of input representations, spatial support, and five architectures identify an accurate and efficient physics-structured predictor. We further evaluate the interface on event-disjoint full-size-loader data and deploy the complete perception-proposal-prediction-selection-execution loop for autonomous excavation. The ROS2/TensorRT implementation processes five candidates in 72.4 ms on a Jetson AGX Orin. The simulation results establish decision-level gains, while the physical experiments demonstrate real-world closed-loop feasibility.
Although aerial-legged robots offer combined agility and efficiency, controlling high-speed hopping under complex hybrid dynamics is challenging. Reinforcement Learning (RL) is promising but prone to energy-inefficient "reward hacking". We propose a Dynamics-Informed RL framework for a monopedal hopping quadcopter. By embedding a target Specific Energy into the reward, we constrain the optimization to a physically viable energy manifold, ensuring stable hopping behaviour. By rewarding the phase-consistent behavior, it can encourage bio-inspired stance-phase impulse. Furthermore, penalizing the electro-mechanical power waste induces the motors generate an efficient impulse. This enables the policy to inject energy strictly during spring restitution without heuristic state machines. MuJoCo simulations validate robust height regulation and forward velocity tracking up to 2.0 m/s despite severe attitude-contact coupling. Ultimately, our approach yields a highly agile hopping gait, reducing energy consumption by 82% and 73% compared to hovering baselines and inefficiency baseline, respectively.
Chinese Translation
尽管空中腿式机器人兼具敏捷性和高效性,但在复杂混合动力学下控制高速跳跃仍具有挑战性。强化学习(RL)具有前景,但容易陷入能源效率低下的“奖励黑客”(reward hacking)问题。我们提出了一种用于单足跳跃四旋翼的动力学信息引导强化学习框架。通过将目标比能量(Specific Energy)嵌入奖励中,我们将优化约束在物理可行的能量流形上,确保稳定的跳跃行为。通过奖励相位一致的行为,它可以鼓励仿生支撑相冲量。此外,惩罚机电功率浪费促使电机产生高效冲量。这使得策略能够在弹簧回弹期间严格注入能量,而无需启发式状态机。MuJoCo 仿真验证了在严重的姿态-接触耦合下,仍具有鲁棒的高度调节和最高 2.0 m/s 的前向速度跟踪能力。最终,我们的方法产生了一种高度敏捷的跳跃步态,与悬停基线和不节能基线相比,能耗分别降低了 82% 和 73%。
Combining aerial thrust with spring-loaded hopping makes monopedal quadcopters promising for locomotion over complex terrain, but heuristic proportional-integral-derivative (PID) tuning limits coordination between active thrust and passive contact dynamics. We present a direct estimated-state-to-motor Proximal Policy Optimization (PPO) policy that commands four motors without an explicit hopping state machine or low-level attitude PID. Its reward combines Energy-Manifold Shaping for mass-normalized vertical-energy tracking and apex-state anchoring with Efficiency Shaping, which uses a history-aware power estimator to penalize general power use, impose an additional airborne-power cost, and penalize airborne near-stationarity. In representative hardware runs, the PPO-based control stack reduced cycle-averaged measured electrical power by 30.7% and mean total normalized thrust by 49.8% relative to the tuned PID-based control stack, while retaining repeatable commanded-height hopping and more concentrated landings. These observations are consistent with improved use of passive dynamics and reduced measured electrical demand.
Companion robots face everyday situations in which several feasible behaviors may be appropriate, yet different people prefer different responses. We introduce InterSocialBench, a benchmark of 210 domestic scenarios and 18 high-level behaviors, pairing judgments from 100 human participants with 23,520 responses from seven large language models under 16 personality conditions. Each human annotation preserves a preferred action alongside explicitly appropriate and inappropriate candidates. A structured construction pipeline covers behavioral alternatives, competing situational cues, and relevant history and future tasks. Evaluation distinguishes preferred-choice agreement from explicit rejection, using scenario-grouped splits for trainable predictors. Simple frequency and persona-voting baselines illustrate these objectives. Across the tested prompts, model and human behavior distributions differ, and the diversity gap remains after matching response counts: humans exhibit 4.68 distinct choices per scenario, compared with 2.06--3.46 for the models. Human scenario-level plurality agreement is 51.5%, describing disagreement rather than a universal prediction ceiling. InterSocialBench supports evaluating social behavior selection without replacing individual judgments with a single consensus label.
P-POSEMEM: Projective Semantic Memory for Consistent Language Grounding under Pose-Graph Rewrites
P-POSEMEM:用于位姿图重写下一致语言定位的投影语义记忆
Sier, Ha, Salmasi, Ali, Xu, Mengya, Zhang, Haizhou, Lu, Jie, Zou, Zhuo, Yu, Xianjia, Westerlund, Tomi
Abstract
A robot following language instructions needs its semantic memory to keep naming the same physical object while the SLAM pose graph underneath is optimized, loop-closed and compressed. Maps committing each detection to a world coordinate cannot: a closure moves the anchor it was measured from, or the solver marginalizes that anchor, and the query then selects a different object although both graphs represent the same posterior. P-POSEMEM stores each observation as an immutable event at its birth keyframe, retains the Bayes-tree elimination conditional of every marginalized keyframe, and integrates the semantic likelihood over the reconstructed joint posterior of poses, anchors and identities. Dproj, the total-variation defect between the language-goal distributions of inference-equivalent full and marginalized graphs, measures this directly. Over 40 HM3DSem scenes and 112,000 queries, P-POSEMEM reproduces the full-graph oracle (Dproj = 0) and reduces goal flips against every memory-reducing baseline. On an eight-run campaign whose 761 closures rewrote the map by up to 47 m, Dproj stays below 10^-13 with 0/288 goal flips when elimination follows the closures, where every ablation and a coordinate committed at insertion flip goals it does not; under a live bounded solver the same memory flips 23/288 against 53 for that frozen coordinate. A pre-registered negative control is detected by Dproj while leaving calibration error and navigation success unchanged, indicating that these measures capture distinct failure modes. Retrieval is held fixed by a shared frozen detector, isolating the gain to memory consistency. Code and data: https://anonymous.4open.science/r/posemem-2328/.
Recent advances in robot imitation learning have produced visuomotor policies that predict actions directly from visual observations. Yet visually similar scenes can require different actions as target position, object height, or contact geometry changes. Pretrained RGB features may map these geometrically distinct states to similar policy inputs, while simply adding depth requires the policy to learn RGB-depth correspondence from the same limited demonstrations used to learn control. We introduce StereoPatch, a patch-aligned RGB-depth representation that binds registered metric geometry directly to the RGB patches used for action prediction. On a shared 2-D patch grid, asymmetric cross-attention incorporates depth information into the corresponding RGB features before action decoding. The resulting StereoPatch Tokens provide a geometry-aware visual representation that can condition general visuomotor policies without changing their underlying learning objectives. Across six real-robot tasks, StereoPatch achieves higher closed-loop success than appearance-only, geometry-only, raw RGB-D, and late-fusion baselines. Additional experiments across three simulation suites evaluate compatibility across visuomotor policy architectures, spatial generalization, and operating limits. Results suggest that resolving control-relevant geometric ambiguity benefits from aligning depth directly with the visual features used for action prediction, rather than supplying it as an independent modality. Project page: https://aus.bot/research/stereopatch/.
World Action Models (WAMs) use video generation models to predict future visual dynamics for robotic manipulation, but iterative denoising introduces additional latency for closed-loop control. We empirically find that visual content converges at different rates during denoising. Static background structure forms early, whereas the gripper and manipulated object remain blurry after the first step, with their interaction dynamics emerging only through subsequent denoising. Consequently, naively truncating a multi-step video model to one step preserves scene structure but loses the interaction-centric dynamics most critical for manipulation. To address this issue, we propose DIDO, which distills the converged dynamics of a multi-step video model into a single denoising step. DIDO combines distribution matching distillation with interaction-centric representation guidance. Beyond compressing multi-step generation into one forward pass, DIDO explicitly models the gripper, manipulated object, and their interaction using supervised bounding-box visual reasoning tokens. Additionally, DIDO aligns the target object's representations across multiple model layers with features from a pretrained DINOv3 encoder. This interaction-centric guidance helps the distilled model preserve both the relevant entities and their future dynamics in a single step, while substantially reducing inference latency. DIDO achieves an average success rate of 99.0\% on LIBERO, 76.6\% on LIBERO-Plus, and 92.0\% on RoboTwin, while also demonstrating effective transfer to long-horizon and generalization tasks in real-world robotic manipulation.
An Information-Space Perspective to Scene Graph Sufficiency for Robotic Task Planning
从信息空间视角看机器人任务规划中的场景图充分性
Sakçak, Başak, Verdoja, Francesco
Abstract
Planning in complex environments requires task specifications grounded in representations that capture objects, relations, and affordances; scene graphs meet this need, but their size in large environments hinders efficient planning. While task-aware pruning and hierarchical abstractions have been explored, a general, task-centric formalization of what constitutes a sufficient scene graph for planning remains open. This paper provides such a formalization by modeling planning over scene graphs within an information-spaces framework through the definition of scene graph transition systems and relevant action semantics for navigation and manipulation. We then introduce derived scene graphs via information mappings that merge and prune nodes and induce quotient transition systems augmented with motion primitives to capture higher-level actions over merged graph nodes. Sufficiency is characterized by two conditions: (i) the information mapping yields a deterministic quotient, and (ii) the task is well-posed over derived traces, ensuring plans found on the derived model are feasible on the maximal system. We illustrate the framework using a task over an example environment, showing both sufficient and insufficient reduced scene graphs.
Learning a motion prior requires a reward that guides a policy from its current behavior toward demonstrated motion. Adversarial Motion Priors (AMP) provide such a reward with a discriminator. However, adversarial objectives can become uninformative when policy and expert supports are far apart. A naive use of optimal transport (OT) averages matched expert successors into a barycentric target. Averaging across gait phases can weaken the target's joint motion. We introduce Flow-Matched Motion Priors (FMP), an online scalar reward learned from paths connecting current rollout histories to an expert motion bank. Entropic OT supplies the coupling. Before each policy update, we train a neural potential with flow matching (FM) along the rollout-to-expert paths, endpoint-gradient supervision, and relative-value calibration. The actor receives only physical observations and the reward remains a scalar, as in AMP. Controlled reward-model experiments show substantially better generalization beyond the fitting rollout than value-only or endpoint-only fitting. On Unitree G1, matched 50-million-transition experiments compare FMP with AMP, a barycentric OT reward, and nested ablations under demonstration and fixed-pose initialization. FMP produces stable forward walking at 0.727 m/s from demonstration resets and 0.338 m/s from a fixed default pose. In the fixed-pose condition, it incurs 129 falls versus 243 for the endpoint-only control. Against a static score-gradient teacher, dynamic FM reduces score-increment error at interpolation fractions 0.25 and 0.50 while using 29% less offline fitting time.
Tracking the Ground: Online Lidar Identification of Robot-Induced Soil Deformation in Agricultural Environments
追踪地面:农业环境中机器人诱发土壤变形的在线激光雷达识别
Montagnon, Tom, Laconte, Johann, Thuilot, Benoit, Cho, Wonjae, Lenain, Roland
Abstract
Agriculture faces many challenges, and robotic systems can play an important role in addressing them by improving the efficiency and sustainability of field operations. Among these challenges, preserving soil health is a critical concern, as vehicle-soil interactions can degrade the soil structure and produce unwanted surface deformation. A key step toward soil-aware robotics is to explicitly account for how vehicle traffic deforms the ground, yet soil state is typically not treated as a variable. We address this gap by proposing a framework to quantify traffic-induced soil deformation and estimate its evolution online from lidar observations. The method relies on a reduced-order parametric model that represents the soil behavior via physically interpretable parameters, yielding a continuously updated and observable representation of soil state. Experiments conducted in different soil conditions demonstrate the ability of the approach to capture deformation induced by the robot. By making soil response measurable and interpretable during operation, the proposed framework establishes a basis for soil-aware robotic operation, in which the estimated state can be exploited to adapt robotic behaviors in order to reduce soil degradation.
Volumetric Harmonic Field Navigation for Quadrotors
四旋翼飞行器的体谐波场导航
Jia, Shuxiu, Mukherjee, Amartya, Yuan, Yating, Liu, Jun
Abstract
Quadrotor navigation in cluttered 3-D environments requires global guidance while local motion remains subject to collision and motion limits. Harmonic potentials provide dense guidance from a global boundary value problem, but coupling a volumetric harmonic field to constrained physical quadrotor motion remains an open experimental problem. We couple a precomputed volumetric harmonic field with a constrained predictive planner that queries the field at predicted positions instead of extracting a global reference path. In Structured 3-D tests, harmonic guidance yields larger minimum clearance and lower RMS jerk than matched Dijkstra guidance, at the cost of longer paths; the same pattern remains when both methods use the same passage. Long maze tests span routes far beyond one prediction horizon, and Crazyflie trials validate physical execution. To the best of our knowledge, this is the first physical quadrotor demonstration of volumetric harmonic field navigation. The results show that globally constructed harmonic guidance can directly support local constrained motion generation on a physical quadrotor.
CMDIR extends Manifold-Decomposed Impedance Retargeting (MDIR) to transform fixed-impedance demonstrations into continuous variable-impedance controllers, which can also serve as structured supervision for imitation learning. Continuous Task-Manifold Impedance Representation (TMIR) pairs an evolving task frame with controller instructions. Demo-relative Compromise dynamics retain moving-basis transport and control/physical metric mismatch, yielding displacement, reaction-impulse, and perturbation-sensitivity criteria. Quality-to-Fast automatically compiles a solver structure from development paths within a predefined finite space, re-instantiates that structure for each demonstration, and certifies the resulting candidate by multi-resolution evaluation. Across 225 retargeted-controller trials in three real contact tasks, full CMDIR improves mean task-proxy retention and reduces mean pose deviation, force fluctuation, and peak force relative to discrete MDIR. FastMPO achieves a $5.8$--$9.4\times$ speedup over C-MPO with comparable closed-loop outcomes. Downstream experiments demonstrate learnability of the complete TMIR supervision interface; lower force fluctuation and peak force are observed among successful executions, while completion reliability remains uneven across tasks and environments.
Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not converged to a common design. Dexterous hands differ in finger structure, contact surfaces, and sensor layouts, while simulated tactile signals still differ from measurements produced by physical sensors. These factors make it difficult to study visuo-tactile manipulation across diverse dexterous hands within a consistent experimental setting. We present Bench2Dex, a simulation benchmark for visuo-tactile bimanual manipulation across 12 dexterous hands. We adapt existing robot models with a shared simulated tactile interface that converts local contact geometry into image-like tactile observations. The interface provides a consistent observation format across different hand morphologies without attempting to reproduce the output of a specific physical tactile sensor. Bench2Dex includes 26 bimanual manipulation tasks that involve tool use, articulated-object interaction, and multi-stage manipulation, together with about 1.3K human-teleoperated demonstrations. The benchmark provides synchronized visual, tactile, proprioceptive, action, and object-state observations, together with executable task metrics. For robustness, we group seven perturbation types into invariance axis, where the correct action does not change, and equivariance axis, where the correct action changes together with the perturbation. We evaluate ACT, Diffusion Policy, pi0.5, and GR00T N1.5 on Bench2Dex and report their performance and failure modes. Bench2Dex is meant as a platform for studying visuo-tactile learning across dexterous hands. It does not assume that simulated tactile observations can replace real tactile sensing; it offers a shared setting for algorithm development while tactile hardware and simulation models are still evolving.
Light detection and ranging (LiDAR) remains less explored than RGB-D sensing for perceptive legged locomotion, and existing LiDAR-based approaches often rely on explicit mapping. We present JEPLO (Joint-Embedding Predictive learning for legged LOcomotion), a single-stage learning framework for mapping-free, LiDAR-based perceptive locomotion for legged robots. We introduce a proprio-exteroceptive JEPA (PE-JEPA) world model to learn predictive egocentric terrain representations from onboard observations, including raw LiDAR scans. A concurrent JEPA-teacher-student (CJTS) pipeline is further proposed to train a locomotion policy informed by JEPA latent representations in simulation using deep reinforcement learning with a simple reward formulation. The framework achieves successful sim-to-real transfer, enabling omnidirectional traversal of diverse terrains, including long staircases and high boxes, with lightweight onboard computation. Evaluations demonstrate greater robustness than existing perceptive locomotion frameworks, particularly under degraded perception caused by occlusion, sparsity and noise. Further analysis validates JEPLO's ability to retain task-relevant information under these challenging conditions. We open-source our implementation, experimental datasets, and hardware setup designs https://github.com/ASIG-X/JEPLO.
Learning chunk-based visuomotor policies for long-horizon robot manipulation remains challenging. Recent action-chunking methods have shown promising performance by predicting temporally extended action sequences. However, their failures are often dominated by prediction errors at a small number of critical timesteps rather than uniformly poor predictions across the entire action chunk, making uniform refinement inefficient and insufficiently targeted. To address this bottleneck, we propose Uncertainty-Guided Refinement (UGR), a sparse refinement framework for chunk-based visuomotor policies. Specifically, UGR follows a coarse-to-refine design: it first predicts a full action chunk, estimates per-step temporal uncertainty from the coarse hidden states, and applies residual correction only to the most uncertain timesteps selected by a binary mask. The uncertainty branch is decoupled from the coarse action predictor, enabling clean attribution of the refinement gains to uncertainty-guided correction rather than additional predictor capacity. Extensive experiments on five dual-arm manipulation tasks from the RoboTwin benchmark show that UGR achieves the best success rate on four tasks, improves over the ACT baseline by up to 13% absolute, and outperforms both full-chunk and position-agnostic block refinement in ablation studies.
Uncrewed Aerial Manipulators (UAMs) extend the capabilities of Uncrewed Aerial Vehicles (UAVs) from perception to physical interaction. Among various aerial interactions, push-and-pull operations are fundamental manipulation primitives that require sustained horizontal forces while maintaining stable flight. In this paper, we propose DuctAM, a compact aerial manipulation platform that enhances horizontal force capability for push-and-pull interactions using two ducted fans integrated along the quadrotor interaction axis. An attitude-force decoupled control scheme enables controllable horizontal forces without requiring large attitude changes. Extensive real-world experiments are conducted to validate the DuctAM. Figure-eight trajectory tracking experiments demonstrate stable flight and accurate motion control in both quad and duct modes. Force-measurement experiments quantify the decoupled longitudinal force capability of DuctAM. Finally, representative push-and-pull interaction tasks, including cart pushing, door closing, and drawer opening, verify the practical effectiveness of DuctAM. The results show that DuctAM achieves significantly improved horizontal interaction force capability while maintaining stable flight compared with conventional UAVs.
Scaling generalist policy models with heterogeneous data is limited by the lack of unified, low-noise action supervision. Human egocentric videos are abundant, but only a small fraction comes with high-quality hand-action labels. Observed world transitions offer a common source of action-related supervision across data sources. We introduce WLA$^3$ (World Latent Action Modeling for Semantics, Dynamics, and Kinematics), a unified generalist policy model framework built around representations learned by a World Latent Action Model (WLAM). WLAM first learns how multimodal world states change over a local interval, encoding synchronized camera views and available embodiment-state changes into a compact local latent action and a richer transition feature. Reconstruction from partial modalities and consistency across overlapping windows encourage robust transition representations. WLA$^3$ reuses them across semantics, dynamics, and kinematics: local latent actions support action-sensitive physical-dynamics modeling, segment-level features directly supervise the VLM through a Semantic Latent Aggregate (SLA), and an action expert jointly predicts latent actions together with embodiment-specific robot controls. Human videos provide scalable transition supervision, while robot trajectories ground the shared representation in executable native controls. On LARYBench, the final 32D latent action reaches 67.89\% average classification accuracy. WLA$^3$ achieves 81.9% average success across six real-robot tasks versus 66.2% for $\pi_{0.5}$. Performance improves as generalist policy model mid-training data scales, and human videos support human-to-robot transfer. Project page can be found at https://wla-3.github.io/.
Physical AI relies on frequently-updated, latency-sensitive video stream to perceive, reason, and interact with the physical world, resulting in strict latency requirements with much higher data volumes that existing 5G networks cannot support. Goal-oriented communication (GoC) offers as a promising approach to solve this challenge by transmitting only task-relevant semantic representations. However, existing GoC frameworks were mainly evaluated in the simulations while their effectiveness has never been validated in a practical deployment of physical AI application. In this work, we develop an end-to-end GoC testbed for Physical AI, which connects a PiPER robot arm equipped with an RGB-D camera and a 5G modem to an NVIDIA Jetson AGX Orin edge server through a 5G OpenAirInterface network. We propose and implement three GoC frameworks that transmit 3D bounding boxes, 2D scene graphs, and 3D scene graphs, as three types of semantic representations, respectively. They share the common functional modules designed for closed-loop Physical AI applications, including semantic extraction, full stack 5G transmission, language model inference, digital twin validation, and robotic control. Extensive experiments on our testbed show that our GoC frameworks reduce the task completion time by up to 52.6% and improve task success probability by up to 45%, compared to the traditional framework that periodically transmits the raw image data. These results validate the practical effectiveness of our GoC framework and pave the way for efficient and reliable Physical AI applications over future 6G networks. Project website: https://sites.google.com/view/goc-physical-ai-testbed.
Slip detection is fundamental to dexterous manipulation, yet existing systems often lack precise characterization of detection latency and cross-platform generalization. We present SlipSense, a multimodal tactile slip-detection framework built on TacV5, a compact sensor integrating a $32 \times 32$ piezoresistive array operating at 240 Hz and a 3-axis MEMS accelerometer operating at 8 kHz. The piezoresistive array captures spatial pressure distributions, while the accelerometer captures friction-induced vibrations, providing complementary slip cues. The framework performs modality-specific encoding, intra-sensor fusion, and cross-modal attention with causal temporal prediction at 240 Hz. Experiments on a dataset of 1.4 million frames spanning 37 objects demonstrate the complementarity of the two modalities. SlipSense achieves 96.7% Macro F1 with a false-positive rate below 1.6%, detecting 76% of slip events within 23.1 ms. When trained solely on UMI data, SlipSense generalizes zero-shot to a Tesollo dexterous hand, transferring across unseen objects, distinct sensor units, and robotic platforms without retraining.
Dexterous manipulation of deformable objects demands continuous fingertip-level regulation of pressure, friction, and incipient slip. We study one of the most challenging cases: dexterous cable tracing, feeding a cable through the hand with repeated pinch-and-curl motions of the thumb and index finger. We introduce Touch2Trace, a tactile-driven imitation-learning system for this task, and provide, to our knowledge, the first systematic real-world characterization of how encoder pretraining, control rate, temporal context, and spatial resolution each shape policy performance. The winning learning recipe combines a tactile encoder pretrained for a custom 32 x 32 piezoresistive sensor (TacV5) via self-supervised learning with a lightweight transformer policy trained on teleoperated demonstrations via behavior cloning, deployed at 60 Hz on a Tesollo DG-5F hand. Tactile feedback without vision or explicit cable-state estimation significantly improves tracing performance versus a proprioception-only baseline: from 0.2 cm to 20.1 cm mean distance and 0% to 93% success rate, with zero-shot transfer to unseen cables and routing conditions. The results quantify the influence of key parameters in tactile-driven systems for reliable dexterous deformable object manipulation.
Beyond Single-Axis Testing: Paired Evaluation of Compound Robustness in Vision-Language-Action Policies
超越单轴测试:视觉-语言-动作策略中复合鲁棒性的配对评估
Sawada, Hiroki, Kasahara, Shunichi
Abstract
Vision-language-action policies are typically evaluated one perturbation at a time, providing a useful diagnosis of their sensitivity to individual distribution shifts. Real-world deployment, however, may involve several shifts simultaneously, and it remains unclear how these individual robustness measurements compose. We ask whether compound robustness can be inferred from single-axis evaluations. We introduce LIBERO-CTRL, a six-axis benchmark that pairs each initial state across single-axis conditions and a matched simultaneous condition. This design reveals two opposing outcome changes that aggregate success rates cannot distinguish: emergent failures, where all single-axis rollouts succeed but the simultaneous rollout fails, and compensated successes, where at least one single-axis rollout fails but the simultaneous rollout succeeds. Because one transition decreases compound success while the other increases it, they can cancel, making aggregate compound performance appear consistent with single-axis measurements even when individual outcomes differ substantially. These opposing transitions can largely cancel in aggregate: even when the difference between the two transition rates is not statistically distinguishable from zero, as many as 29.0% of matched initial states still change outcome. Across six policies and three severity levels, such outcome changes reach 34.5% in the most affected condition. The relative prevalence of the two transitions varies across policies and severities, while the transition rates remain similar under independent re-evaluation of stochastic policies. Compound robustness therefore cannot be characterized from aggregate single-axis success rates alone; matched per-instance evaluation is needed to reveal how joint perturbations alter behavior.
Mobile manipulators deployed across many rooms and visits should improve with experience: after discovering that a cabinet is locked or finding an object in a drawer, the robot should reuse that knowledge rather than start each task from scratch. Yet today's robots often treat each task as new: compact scene representations omit interaction-derived knowledge, raw video histories are difficult to query, and VLM planners reason at inference time without persistently updating what the robot knows. We present MessyMem, a persistent memory system that enables mobile manipulators to learn from experience and reuse that knowledge across future tasks. It maintains a spatially grounded 3D scene graph of objects and locations, augments it with properties and outcomes learned through interaction, and links visual observations for fine-grained recall. We evaluate MessyMem in simulation and on a real mobile manipulator. In a continuous 25-task simulation spanning over 3 hours, MessyMem achieves 80.0% task progress, outperforming the strongest ablation by 14.8 percentage points and the strongest external baseline by 28.9 points, while retrieving task-relevant evidence from thousands of stored keyframes and over an hour into the past.
ResSafe: Learning Safety Filtering with Residual Reinforcement Learning for Humanoids
ResSafe:学习使用残差强化学习为人形机器人进行安全过滤
Qu, Gechen, Zhang, Tong, Zhang, Bike, Wang, Yen-Jen, Sreenath, Koushil, Tomlin, Claire, Choi, Jason Jangho
Abstract
Safe control of humanoid robots remains challenging due to their high-dimensional dynamics, contact-rich interactions, and sensitivity to disturbances. Although reinforcement learning has enabled effective locomotion and motion tracking, learned policies can still generate unsafe actions that lead to instability or falls. In this work, we propose residual reinforcement learning as an implicit safety-filtering mechanism for safe humanoid control. Instead of relying on a single nominal policy to simultaneously balance performance, safety, and robustness, we decouple performance and safety. The nominal policy focuses solely on task performance, while a residual policy learns safety corrections. This decoupling leads to a better performance--safety Pareto trade-off and avoids the need for careful tuning of multiple competing reward terms within a single policy training. We show that the residual policy can act as an implicit safety filter.