← Back to Index
Daily Research Digest

arXiv Papers

2026-09-18
413
Papers
4
Categories
413
Translated
收藏清单 0
机器人学 (Robotics)
116
cs.RO / 1 / 2609.19194

AthenaZero: A low-inertia, bimanual robot for dynamic manipulation

AthenaZero:一种用于动态操作的低惯性双臂机器人
Morgan, Andrew S., Xie, Gregory, Bass, Capprin, Thomasson, Rachel, Wang, Chunpeng, Shahriari, Erfan, Busa, Harrison, Araromi, Oluwaseun, Aronov, Joseph, Burgess, Michael, Dimitrov, Velin D., Estrada, Matthew A., Hamdi, Faris, Kendig, Samuel, Lee, Taeyoon, Mur-Miranda, Jose Oscar, Sommers, Emma, Titchener, Paul, Wang, Margaret, Wilson, Achu, Yeatman, Mark, Yirmibesoglu, Osman Dogan, Rizzi, Alfred A., Mozeika, Annan, Rojas, Nicolas, Odhner, Lael
Abstract
AthenaZero is a bimanual manipulator designed to minimize inertia without compromising control authority. By utilizing quasi-direct drive actuation and transmission remotization techniques, the system achieves an effective endpoint mass comparable to that of a human---about an order of magnitude less than conventional robot manipulators. This characteristic, combined with its inherent torque transparency, makes AthenaZero exceptionally well-suited for dynamic manipulation. We describe the methodology} that led to this design and demonstrate the robot's capabilities on three baseball-inspired tasks: throwing, catching, and batting, which showcase complex interactions on human-comparable timescales where milliseconds matter. AthenaZero was capable of throwing at speeds in excess of 30 m/s, with catching and batting at speeds in excess of 14 m/s over a short 7.3 m distance. Batting practice and a game of catch were subsequently performed in robot-to-robot and human-to-robot variations, showcasing the efficacy and adaptability of our system in tasks that require high acceleration.
Chinese Translation
AthenaZero是一款旨在在不牺牲控制能力的前提下最小化惯性的双臂机械臂。通过采用准直驱(quasi-direct drive)驱动和传动远程化技术,该系统的有效末端质量达到了与人类相当的水平——比传统机器人机械臂低约一个数量级。这一特性结合其固有的力矩透明性,使AthenaZero特别适合执行动态操作任务。我们阐述了实现这一设计的方法论,并在三个受棒球启发的任务上展示了机器人的能力:投掷、接球和击球,这些任务展示了在毫秒级至关重要的、与人类相当的时间尺度上的复杂交互。AthenaZero能够以超过30 m/s的速度投掷,并在仅7.3米的短距离内以超过14 m/s的速度完成接球和击球。随后,我们进行了机器人对机器人和人对机器人两种形式的击球练习和传接球游戏,展示了我们的系统在需要高加速度任务中的有效性和适应性。
cs.RO / 2 / 2609.19196

DITTO: Dexterous Interface for Transparent TeleOperation

DITTO:面向透明遥操作的灵巧接口
Palacios, Joaquin, Lee, Katelyn, Zhang, Cheng, He, Zhanpeng, Ciocarlie, Matei
Abstract
Collecting data for manipulation with high-DOF hands is challenging, as interfaces must capture rich hand motion while rendering the contact interactions essential for precise manipulation. Existing data collection approaches face a trade-off: teleoperation ensures deployment consistency but lacks force feedback, while handheld (in-the-wild) systems provide natural force transparency but introduce a visual embodiment gap at deployment. We present DITTO, a Dexterous Interface for Transparent TeleOperation, which resolves this through the anatomically informed co-design of a dexterous 7-DOF robotic hand and a kinematically equivalent motorized exoskeleton. A 1-to-1 actuator mapping between the exoskeleton and robotic hand enables handheld (in-the-wild) data collection and bilateral teleoperation with joint-level force feedback unified in a single platform. We demonstrate that the DITTO exoskeleton spans the operator's natural index-to-thumb workspace, and showcase DITTO's dexterous capabilities via learned policies on contact-rich tasks.
Chinese Translation
为高自由度灵巧手采集操作数据极具挑战性,因为数据采集接口既要捕捉丰富的手部运动,又要呈现对精确操作至关重要的接触交互。现有的数据采集方法面临权衡:遥操作能保证与部署环境的一致性,但缺乏力反馈;而手持式(in-the-wild)系统提供自然的力透明性,但在部署时引入了视觉本体差异。我们提出DITTO——一种面向透明遥操作的灵巧接口,通过对7自由度灵巧机械手与运动学等效的电动外骨骼进行基于人体解剖学知识的协同设计来解决这一问题。外骨骼与机械手之间的一对一执行器映射,使手持式(in-the-wild)数据采集与具备关节级力反馈的双边遥操作统一在同一个平台上。我们证明DITTO外骨骼覆盖了操作者食指到拇指的自然工作空间,并通过在接触密集型任务上学习的策略展示了DITTO的灵巧操作能力。
cs.RO / 3 / 2609.19200

ULOHA: An Underwater Bimanual Robot System for Robot Learning

ULOHA:一个面向机器人学习的水下双臂机器人系统
Kobayashi, Masato, Tsunoori, Takeru
Abstract
Underwater visuomotor policy learning has focused primarily on single manipulators, while bimanual imitation learning has been studied largely in air. We present ULOHA, an underwater bimanual robot learning platform that combines custom-designed leader--follower hardware with software extensions to LeRobot, integrating teleoperation, multi-view sensing, demonstration collection, policy training, and autonomous deployment. Real-robot experiments demonstrate a range of coordinated underwater bimanual behaviors, including inter-arm transfer, shared-object manipulation, and buoyancy-driven interception. We evaluate ACT, Diffusion Policy, and the vision--language--action model SmolVLA on the platform. We investigate how learning methods and execution strategies developed for manipulation in air perform underwater, examining bubble disturbances, buoyancy-driven object motion, action-execution horizons, and real-time chunking. A separate single-arm study examines policy transfer between air and water and shows that demonstrations spanning both media support execution in both under the tested conditions. ULOHA provides a unified experimental platform for studying underwater bimanual robot learning under the coupled perceptual and physical effects of underwater environments. Additional material: https://mertcookimg.github.io/uloha/
Chinese Translation
水下视觉运动策略学习的研究主要集中在单机械臂上,而双臂模仿学习则大多在空气环境中进行研究。我们提出了ULOHA,一个水下双臂机器人学习平台,它将自主设计的领导-跟随(leader-follower)硬件与基于LeRobot的软件扩展相结合,集成了遥操作、多视角感知、演示数据采集、策略训练和自主部署等功能。真实机器人实验展示了多种协调的水下双臂行为,包括双臂间物体传递、共享物体操作以及浮力驱动的拦截。我们在该平台上评估了ACT、Diffusion Policy以及视觉-语言-动作模型SmolVLA。我们研究了为空气中操作设计的学习方法与执行策略在水下环境中的表现,考察了气泡扰动、浮力驱动的物体运动、动作执行时域以及实时分块(real-time chunking)等因素。此外,一项独立的单臂研究考察了策略在空气与水中之间的迁移能力,结果表明在所测试的条件下,跨越两种介质的演示数据能够支持在两种介质中的策略执行。ULOHA为研究水下环境感知与物理效应耦合作用下的水下双臂机器人学习提供了一个统一的实验平台。附加材料:https://mertcookimg.github.io/uloha/
cs.RO / 4 / 2609.19204

REACT: A Fully Spiking State-Space Model for Real-Time Event-Driven Temporal Perception

REACT:一种用于实时事件驱动时序感知的全脉冲状态空间模型
Keime, Geoffroy, Cuperlier, Nicolas, Cottereau, Benoit R.
Abstract
Robotic systems operating in dynamic environments require visual perception that evolves continuously with the incoming sensory stream. Event cameras provide microsecond temporal resolution and asynchronous sensing, but most learning-based methods accumulate events into frames or temporal bins, introducing an integration delay that can limit fast reaction. Here we propose REACT, a fully spiking state-space model for event-driven temporal perception that processes raw events one by one, without temporal accumulation. REACT uses a complex-valued spiking neuron, C-SiLIF, whose continuous-time dynamics are driven by the physical inter-event interval, allowing its internal state to evolve at the temporal resolution of individual events. We evaluate REACT on gesture recognition and time-to-collision (TTC) estimation from full-field event streams, without a target bounding box or localization input. On EvTTC, REACT achieves a 9.59% relative TTC error with 4.6 ms end-to-end inference latency, within 0.15 percentage points of the best learned method while requiring no target prior. At the dataset's mean approach speed, this latency corresponds to only 4 cm of vehicle motion, compared with 1 m for the fastest competing learned method. REACT further supports anytime TTC prediction, zero-shot transfer to a different driving sequence, and INT8 quantization, reducing the estimated energy consumption from 18.5 to 2.8 mJ per 32,768 events. These results show that event-driven spiking state-space dynamics can provide low-latency, continuously updated temporal perception for reactive robotic systems.
Chinese Translation
在动态环境中工作的机器人系统需要随输入感知流持续演化的视觉感知。事件相机提供微秒级的时间分辨率和异步感知能力,但大多数基于学习的方法将事件累积为帧或时间片段,引入了可能限制快速反应的积分延迟。本文提出REACT,一种用于事件驱动时序感知的全脉冲状态空间模型,它逐个处理原始事件,无需时间累积。REACT采用复值脉冲神经元C-SiLIF,其连续时间动力学由物理事件间隔驱动,使其内部状态能够以单个事件的时间分辨率演化。我们在手势识别和基于全场事件流的时间到碰撞(TTC)估计任务上评估REACT,无需目标边界框或定位输入。在EvTTC数据集上,REACT实现了9.59%的相对TTC误差和4.6 ms的端到端推理延迟,与最佳学习方法差距在0.15个百分点以内,且无需目标先验信息。在数据集的平均接近速度下,该延迟仅对应4厘米的车辆运动距离,而最快的竞争学习方法对应1米。REACT还支持随时(anytime)TTC预测、向不同驾驶序列的零样本迁移以及INT8量化,将每32,768个事件的估计能耗从18.5 mJ降至2.8 mJ。这些结果表明,事件驱动的脉冲状态空间动力学能够为反应式机器人系统提供低延迟、持续更新的时序感知。
cs.RO / 5 / 2609.19216

4D Radar Perception Algorithms for Autonomous Driving: A Review

面向自动驾驶的4D雷达感知算法:综述
Wu, Xumin, Zhou, Jun, Mei, Jilin, Min, Chen, Hu, Yu
Abstract
Research on 4D millimeter-wave radar perception algorithms has flourished in recent years, extending from signal processing and object detection to semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. This review organizes the field according to the evolution of perception tasks and algorithms. It first introduces radar fundamentals, data representations, and quality-enhancement methods, and then reviews object-level perception, motion and localization, local and dense spatial perception, and dynamic scene understanding. Across these directions, we compare radar-only learning, multimodal fusion, and cross-modal supervision and knowledge distillation. Particular attention is paid to how elevation, Doppler measurements, and radar physical priors are exploited across tasks. We further summarize the task coverage, input data, annotations, and evaluation protocols of existing datasets, clarifying the empirical support for different research directions. Finally, we discuss the common challenges and future directions of 4D radar perception for autonomous driving. This review provides a task-oriented perspective on the transition from sparse object perception to dynamic spatial understanding.
Chinese Translation
近年来,4D毫米波雷达感知算法的研究蓬勃发展,研究范围已从信号处理和目标检测扩展到语义分割、运动估计、占据预测和动态场景重建。本综述按照感知任务与算法的演进脉络对该领域进行梳理。首先介绍雷达基础知识、数据表示与质量增强方法,随后综述目标级感知、运动与定位、局部与稠密空间感知以及动态场景理解等方向。在这些方向中,我们比较了仅基于雷达的学习、多模态融合以及跨模态监督与知识蒸馏方法。本文特别关注高程、多普勒测量以及雷达物理先验在各任务中的利用方式。我们进一步总结了现有数据集的任务覆盖范围、输入数据、标注和评估协议,阐明了不同研究方向的实证支持。最后,我们讨论了面向自动驾驶的4D雷达感知所面临的共同挑战与未来方向。本综述为从稀疏目标感知向动态空间理解的转变提供了面向任务的视角。
cs.RO / 6 / 2609.19228

Grasping by interconnection: robust closing motions from coarse object templates

基于互联的抓取:从粗略物体模板获得鲁棒的合拢运动
Vanderheyden, Julien, Drion, Guillaume, Forni, Fulvio, Sacré, Pierre
Abstract
Dexterous robot hands must often grasp objects whose shape, size, and pose are known only approximately. Grasp planners typically require accurate object models or correct errors with feedback, but how much inaccuracy a closing motion can tolerate on its own remains unclear. To address this question, we designed a motion planner based on four principles: a coarse template of the object, human grasp types, an object-centric interaction, and compliant, sliding contacts instead of prescribed contact points. This paper presents the planner, implemented through virtual model control, and its evaluation on a Shadow Dexterous Hand. Without feedback, the planned closing motions tolerated size errors of about 1cm and pose errors of several centimeters and tens of degrees, a wider range than a state-of-the-art data-driven planner in 25 of 27 tested conditions. They also grasped 82.5% of 80 everyday objects and succeeded within an autonomous pipeline. Robustness can thus be designed into the closing motion itself, rather than left only to feedback. This planner opens a path toward reliable manipulation in uncertain settings, which we will pursue by combining it with adaptive feedback control on the physical hand.
Chinese Translation
灵巧机器人手常常需要抓取形状、尺寸和位姿仅被近似已知的物体。抓取规划器通常需要精确的物体模型,或通过反馈来纠正误差,但合拢运动本身究竟能容忍多大的误差仍不清楚。为解答这一问题,我们基于四项原则设计了一个运动规划器:物体的粗略模板、人类抓取类型、以物体为中心的交互,以及柔顺的滑动接触而非预设的接触点。本文介绍了通过虚拟模型控制(virtual model control)实现的该规划器,以及其在Shadow灵巧手(Shadow Dexterous Hand)上的评估结果。在没有反馈的情况下,所规划的合拢运动能够容忍约1厘米的尺寸误差以及数厘米和数十度的位姿误差,在27个测试条件中的25个里,其容忍范围均超过了一项最先进的数据驱动规划器。该规划器还成功抓取了80件日常物品中的82.5%,并在自主流程中取得了成功。因此,鲁棒性可以被设计到合拢运动本身之中,而不仅仅依赖于反馈。该规划器为在不确定环境中实现可靠操作开辟了道路,我们将通过在实体机械手上结合自适应反馈控制来继续推进这一方向。
cs.RO / 7 / 2609.19272

Learning Safe Humanoid Navigation from Reduced Order Models

基于降阶模型学习安全的人形机器人导航
Compton, William D., Olkin, Zachary, Bena, Ryan, Ames, Aaron D.
Abstract
Research in humanoid robotics has achieved rapid progress in locomotion, and recent results have pushed the boundary on autonomous navigation. We demonstrate that a standard single-stage RL navigation pipeline struggles to scale to multi-level and multi-story terrain, limited by the difficulty of complex humanoid terrain interactions such as stairs. To overcome this challenge, we decompose the navigation problem into two pieces. First, we train a policy operating on the reduced order dynamics but with full 3D LiDAR observations to navigate complex, multi-story terrain. We then utilize this navigation knowledge to kickstart a policy operating on the full-order humanoid dynamics, with a frozen locomotion policy in the loop. Additionally, we demonstrate that applying a Poisson safety filter to the navigation policy output recovers safety in the presence of out-of-distribution obstacles, without dropping navigation success rate. We demonstrate the resulting RoM-Nav policy on a Unitree G1, accomplishing mapless multi-floor navigation covering trials with over 10m of vertical displacement and over 100m of path length. Project page with videos https://wdc3iii.github.io/rom-nav/ .
Chinese Translation
人形机器人研究在运动控制方面取得了快速进展,最近的研究成果进一步推动了自主导航的边界。我们证明,标准的单阶段强化学习(RL)导航流水线难以扩展到多层和多楼层地形,其限制因素在于人形机器人与复杂地形交互(如楼梯)的难度。为克服这一挑战,我们将导航问题分解为两个部分。首先,我们在降阶动力学上训练一个策略,但使用完整的3D激光雷达(LiDAR)观测来导航复杂的多楼层地形。随后,我们利用这一导航知识来启动一个在完整人形动力学上运行的策略,并将一个冻结的运动控制策略嵌入环路中。此外,我们证明,对导航策略输出应用泊松(Poisson)安全滤波器可以在存在分布外障碍物的情况下恢复安全性,同时不降低导航成功率。我们在Unitree G1机器人上展示了所得到的RoM-Nav策略,完成了无地图的多楼层导航实验,其中包含垂直位移超过10米、路径长度超过100米的测试。项目页面(含视频):https://wdc3iii.github.io/rom-nav/。
cs.RO / 8 / 2609.19302

OHRID-Retail: An Open Multimodal Dataset of Human Activity in Retail Environments

OHRID-Retail:一个关于零售环境中人类活动的开放多模态数据集
Wang, Xiangrui, Wu, Yuetong, Beeman, Jalen, Cook, Robert, Gu, Yu, Pearson, Nathanial, Smith, Trevor, Hayes, Read, Hu, Boyi
Abstract
Open datasets describing human behavior in environments shared with mobile robots remain limited, particularly for retail activities that combine locomotion, reaching, object handling, and robot guided movement. This paper introduces OHRID Retail, an open, human centered multimodal dataset collected from 16 healthy adults performing a simulated shelf picking task under three within participant conditions: no robot, low speed robot guidance, and high speed robot guidance. Each participant completed two trials per condition. Whole body kinematics were recorded using 17 Xsens Awinda inertial sensors and muscle activity was measured at 10 locations using Delsys Trigno surface electromyography sensors. Descriptive analyses demonstrate variation in whole body movement intensity and muscle activation across robot interaction conditions and body locations. OHRID Retail provides openly available raw recordings, processed measures, documentation, and reproducible analysis resources. The dataset can support research in human activity recognition, multimodal sensor fusion, occupational biomechanics, ergonomics, human aware robot navigation, and human robot interaction in retail and related shared environments.
Chinese Translation
描述人类在与移动机器人共享环境中的行为特征的开放数据集仍然有限,尤其是对于结合了行走、伸手取物、物体搬运以及机器人引导移动等动作的零售活动。本文介绍了 OHRID-Retail,这是一个以人为中心的开放多模态数据集,采集自16名健康成年人在三种被试内条件下执行模拟货架拣选任务的数据,这三种条件分别为:无机器人、低速机器人引导和高速机器人引导。每名参与者在每种条件下各完成两次试验。全身运动学数据通过17个 Xsens Awinda 惯性传感器记录,肌肉活动通过 Delsys Trigno 表面肌电传感器在10个部位进行测量。描述性分析表明,全身运动强度和肌肉激活程度在不同机器人交互条件及身体部位间存在差异。OHRID-Retail 提供了公开可用的原始记录、处理后数据、文档以及可复现的分析资源。该数据集可支持人类活动识别、多模态传感器融合、职业生物力学、人体工效学、人机感知机器人导航以及零售及相关共享环境中人机交互等领域的研究。
cs.RO / 9 / 2609.19315

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

GAVEL:用于经验证且高效的长时程LLM任务规划的图世界模型
Wang, Ruiyang, Hsu, Hao-Lun, Mehta, Swarajh, Kim, Jiwoo, Dou, Zhihao, Pajic, Miroslav
Abstract
Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fail to respect embodiment constraints, recover from planning errors, or reason effectively under partial observability. We present GAVEL, a framework for verifying and repairing long-horizon LLM planning built around an explicit graph world model. The graph represents relevant object-relations, action pre-conditions and effects, and probabilistic beliefs over unobserved object locations. This model can predict the consequences of LLM-generated actions before execution, detect violations, and repair those whose corrections follow directly from the world model. This method also reserves LLM replanning solely for errors requiring semantic reasoning. For multi-task instructions, GAVEL reasons over distributions of possible object locations to reorder remaining subtasks and minimize expected search cost. We evaluate GAVEL on BEHAVIOR-1K across 100 single long-horizon tasks and 500 multi-task instructions. With Qwen3-8B, GAVEL improves single-task success from 41.2% to 91.8% and multi-task success from 19.9% to 92.6%. Distributional belief reasoning also reduces travel distance by approximately 5.4% compared with a static variant. These improvements show that an explicit graph world model harness can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.
Chinese Translation
大语言模型(LLM)为长时程机器人规划提供了灵活的接口,但生成的计划往往无法遵循具身约束、无法从规划错误中恢复,或无法在部分可观测条件下进行有效推理。我们提出了GAVEL,一个围绕显式图世界模型构建的用于验证和修复长时程LLM规划的框架。该图表示相关的对象关系、动作的前提条件与效果,以及对未观测对象位置的概率信念。该模型可以在执行前预测LLM生成动作的后果,检测违规情况,并修复那些可直接依据世界模型进行纠正的错误。该方法还将LLM重新规划仅保留给需要语义推理的错误。对于多任务指令,GAVEL通过对可能对象位置分布进行推理,来重新排序剩余子任务并最小化期望搜索成本。我们在BEHAVIOR-1K上对GAVEL进行了评估,涵盖100个单长时程任务和500个多任务指令。使用Qwen3-8B,GAVEL将单任务成功率从41.2%提升至91.8%,多任务成功率从19.9%提升至92.6%。与静态变体相比,分布式信念推理还将移动距离减少了约5.4%。这些改进表明,显式图世界模型框架能够显著提升长时程具身规划在紧凑型与前沿托管LLM能力下的可靠性与效率。
cs.RO / 10 / 2609.19328

A Morphing Aerial Robot With Thruster-Integrated Flexible Continuum Links for Shape Adaptive Aerial Manipulation

一种带有推力器一体化柔性连续链节的形变空中机器人,用于形状自适应空中操作
Sawada, Eri, Sugihara, Kazuki, Miyamichi, Ayano, Kojima, Kunio, Okada, Kei
Abstract
In recent years, aerial manipulation has attracted increasing attention as a key to expand the application of aerial robots. In this work, we focus on two major research directions for achieving versatile aerial manipulation: (i) acquiring high environmental adaptability using soft manipulators, and (ii) expanding the feasible wrench space by distributing thrusters along the manipulator. However, no aerial robot has simultaneously satisfied these two requirements. Therefore, in this paper, we propose a morphing rotor-distributed aerial robot with flexible continuum links that achieves both high shape adaptability and an expanded wrench space. The flexible continuum links function as soft manipulators, passively conforming to the shape of the environment, while the thrusters distributed along the continuum links expand the feasible thrust wrench space and enable the end-effector to exert large interaction forces. To realize the proposed robot, it is essential to suppress vibrations of the lightweight continuum links. Thus, we develop a composite leaf-spring structure that provides both high torsional and vertical stiffness, and vibration-suppressing control methods. Using these implementations, we demonstrate stable flight and a variety of aerial manipulation tasks. To the best of our knowledge, this is the first work to realize aerial manipulations using flexible links with an integrated thruster.
Chinese Translation
近年来,空中操作作为拓展空中机器人应用的关键技术受到了越来越多的关注。本工作聚焦于实现多样化空中操作的两个主要研究方向:(i)利用软体机械臂获得高环境适应性,以及(ii)通过沿机械臂分布推力器来扩展可行力旋量空间。然而,目前尚无空中机器人能同时满足这两个要求。因此,本文提出一种具有柔性连续链节的推力器分布式形变旋翼空中机器人,同时实现了高形状自适应能力和扩展的力旋量空间。柔性连续链节充当软体机械臂,能够被动地顺应环境形状,而分布在连续链节上的推力器扩展了可行的推力力旋量空间,使末端执行器能够施加较大的交互力。为实现所提出的机器人,抑制轻质连续链节的振动至关重要。为此,我们开发了一种兼具高扭转刚度和高垂直刚度的复合板弹簧结构,以及振动抑制控制方法。借助这些实现手段,我们演示了稳定飞行以及多种空中操作任务。据我们所知,这是首个利用一体化推力器的柔性链节实现空中操作的工作。
cs.RO / 11 / 2609.19330

SemSafe-3DGS: Semantic Risk-Aware Active Navigation in Uncertain 3D Gaussian Splatting Maps

SemSafe-3DGS:不确定三维高斯泼溅地图中的语义风险感知主动导航
Khass, Amirhossein Mollaei, Cosse, Athanasios, Motee, Nader
Abstract
Autonomous robots operating in partially observed environments must navigate safely while acquiring observations that improve future planning. Existing safety formulations generally reason primarily about geometry. Consequently, geometrically similar scene elements may induce comparable control responses despite having different semantic consequences. We present a semantic risk aware safe-active perception framework for navigation in attributed 3D Gaussian maps. Semantic attributes modulate an Average Value-at-Risk collision clearance model through class dependent risk weights, allowing safety-critical Gaussian primitives to receive greater influence in the composite barrier. The resulting weighted clearances are aggregated into a control barrier function, while a trajectory-relevant active perception barrier promotes observations that reduce geometric map uncertainty along the robot's anticipated motion. Both objectives are integrated in a unified CBF-QP that enforces semantic risk-aware collision avoidance as a hard constraint while relaxing information acquisition when it conflicts with safety or task progress. Experiments demonstrate efficient safety constraint, improved navigation through active perception, semantic dependent trajectory adaptation, and real-robot execution under Ackermann dynamics.
Chinese Translation
在部分观测环境中运行的自主机器人必须在安全导航的同时获取能够改善未来规划的观测信息。现有的安全性表述通常主要从几何角度进行推理。因此,几何上相似的场景元素即使具有不同的语义后果,也可能引发相近的控制响应。我们提出了一种用于在带属性三维高斯地图中导航的语义风险感知安全主动感知框架。语义属性通过类别相关的风险权重来调制平均风险价值(Average Value-at-Risk)碰撞安全裕度模型,使安全关键的高斯基元在复合障碍函数中获得更大的影响。所得的加权安全裕度被聚合到控制障碍函数(CBF)中,同时一个与轨迹相关的主动感知障碍函数促进获取能够降低机器人预期运动路径上几何地图不确定性的观测。这两个目标被集成到统一的CBF-QP中,将语义风险感知的碰撞避免作为硬约束加以强制执行,而当信息获取与安全性或任务进展冲突时则予以松弛。实验结果表明,该框架实现了高效的安全约束、通过主动感知改善的导航性能、依赖语义的轨迹自适应,以及在阿克曼(Ackermann)动力学下的真机实验验证。
cs.RO / 12 / 2609.19336

Dynamic-LIVO: A Dynamic-Aware LiDAR-Inertial-Visual Odometry System Using Spatio-Temporal Normals

Dynamic-LIVO:一种基于时空法向量的动态感知激光雷达-惯性-视觉里程计系统
Zhang, Zhixin, Ahiwe, Samuel, Hale, Matthew, Zhao, Liang, Ladosz, Pawel
Abstract
This paper proposes Dynamic-LIVO, a dynamic-aware LiDAR-Inertial-Visual Odometry (LIVO) system for robust state estimation and static colored mapping in dynamic environments. Dynamic-LIVO employs Spatio-Temporal (S-T) normal analysis to identify dynamic LiDAR points and propagates the resulting classification to both LiDAR-inertial and visual-inertial updates, preventing dynamic LiDAR measurements and their associated visual observations from affecting state estimation and mapping. However, S-T normal estimation can be unreliable in newly observed and spatially sparse regions due to insufficient spatio-temporal observations. To address this issue, we introduce a time-delayed S-T normal estimation strategy that defers the classification of insufficiently constrained points and re-evaluates them as additional observations become available. This strategy improves dynamic classification reliability while preserving valid static points for map construction. Extensive experiments on public and self-collected datasets with diverse sensor configurations demonstrate that Dynamic-LIVO improves localization accuracy and produces cleaner static colored maps in challenging dynamic environments. The source code and self-collected dataset will be publicly released upon acceptance.
Chinese Translation
本文提出了Dynamic-LIVO,一种面向动态环境的动态感知激光雷达-惯性-视觉里程计(LIVO)系统,可实现鲁棒的状态估计和静态彩色建图。Dynamic-LIVO采用时空(Spatio-Temporal, S-T)法向量分析来识别动态激光雷达点,并将分类结果传播至激光雷达-惯性以及视觉-惯性更新中,从而防止动态激光雷达测量及其相关视觉观测对状态估计和建图造成影响。然而,在新观测和空间稀疏的区域,由于时空观测不足,S-T法向量估计可能不可靠。为解决该问题,我们引入了一种时延S-T法向量估计策略,将约束不足的点分类结果进行延迟处理,并在获得更多观测后重新评估。该策略在提高动态分类可靠性的同时,保留了有效的静态点用于地图构建。在具有多种传感器配置的公开数据集和自采数据集上的大量实验表明,Dynamic-LIVO提升了动态挑战环境下的定位精度,并生成更干净的静态彩色地图。源代码和自采数据集将在论文被接收后公开发布。
cs.RO / 13 / 2609.19340

ViLoMan: Learning Visual-Proprioceptive Whole-Body Loco-Manipulation Skills for Humanoid Robots

ViLoMan:为人形机器人学习视觉-本体感觉驱动的全身移动操作技能
Tian, Zejie, Hou, Ruibing, Ma, Bingpeng, Karlsson, Börje F., Shan, Shiguang
Abstract
Humanoid loco-manipulation requires adaptive whole-body coordination to seamlessly integrate locomotion and physical interaction. Despite recent advances, learning autonomous loco-manipulation remains challenging due to the scarcity of diverse, physically executable robot-object interaction data and the difficulty of learning unified whole-body control directly from onboard observations. We present ViLoMan, a scalable framework for autonomous humanoid loco-manipulation. ViLoMan first transforms partial kinematic demonstrations of human-object interactions into complete, physically executable robot trajectories. It then leverages these trajectories within a teacher-student distillation framework to learn a unified policy that maps egocentric depth observations and proprioceptive measurements directly to joint-level whole-body actions. During deployment, the policy requires neither reference motions nor intermediate commands. We evaluate ViLoMan on door-closing tasks across diverse door configurations and robot initial conditions in both simulation and the real world. Experimental results demonstrate that a single policy enables a Unitree G1 humanoid to complete the full task using only onboard depth sensing and proprioception, while generalizing robustly across task variations and transferring effectively from simulation to reality. Project page: viloman-anonymous.pages.dev.
Chinese Translation
人形机器人移动操作需要自适应的全身协调,以无缝整合运动与物理交互。尽管近期已取得诸多进展,自主移动操作的学习仍然具有挑战性,其原因在于:多样化且物理上可执行的机器人-物体交互数据稀缺,以及难以直接从机载观测中学习统一的全身控制。我们提出了ViLoMan,一个可扩展的自主人形机器人移动操作框架。ViLoMan首先将部分的人-物体交互运动学演示转化为完整的、物理上可执行的机器人轨迹;然后在教师-学生蒸馏框架中利用这些轨迹,学习一个统一策略,该策略将自我中心的深度观测和本体感觉测量直接映射为关节级的全身动作。在部署过程中,该策略既不需要参考动作,也不需要中间命令。我们在关门任务上评估了ViLoMan,涵盖多样的门体配置和机器人初始条件,并同时在仿真和真实世界中进行了测试。实验结果表明,单一策略即可使Unitree G1人形机器人仅依靠机载深度传感和本体感觉完成完整任务,同时在任务变化中表现出鲁棒的泛化能力,并有效地实现了从仿真到现实的迁移。项目页面:viloman-anonymous.pages.dev。
cs.RO / 14 / 2609.19347

Kinematics-Grounded Agentic AI for Robotic Additive Manufacturing Process Planning

面向运动学基础的山地智能体AI用于机器人增材制造工艺规划
Ge, Jingzhan, Chen, Ruimin, Haghighi, Azadeh, Tang, Jiong, Imani, Farhad
Abstract
Robotic additive manufacturing (AM) extends material-extrusion printing beyond gantry kinematics but makes process planning robot-dependent. A slicer-generated plan that appears favorable in part coordinates can become infeasible or robotically unfavorable on a manipulator because slicer-process decisions and part orientation determine the generated path, while part orientation and workspace placement affect its kinematic realization. Existing AM tools, large language model (LLM)-based decision-support methods, and digital-shadow systems do not provide integrated pre-execution evaluation of these coupled decisions. This paper presents agentic robotic additive manufacturing (A-RAM), an agent-specialist-tool framework that converts user intent and a part file into traceable, execution-ready plans. The LLM interprets manufacturing objectives and constraints, identifies prescribed and searchable planning variables, and encodes this reasoning in a schema-constrained request; a deterministic Planning Agent instantiates the corresponding search workflow, while domain tools compute quantitative evidence for slicing, placement, inverse kinematics, trajectory timing, Joint-6 jerk, and extrusion. The framework is evaluated on a six-axis robotic-arm AM cell through three case studies covering expert-specified planning, goal-only planning, objective-dependent infill screening, and geometry-dependent orientation-placement selection. Across the evaluated candidate sets, selected plans achieve up to 53.5% lower maximum Joint-6 jerk and 48.3% lower mean absolute Joint-6 jerk than the least favorable valid candidates, while objective-specific infill screening yields motion-plan completion times up to 40.1% shorter and extrusion paths up to 12.7% shorter than the corresponding least favorable screened patterns.
Chinese Translation
机器人增材制造(AM)将材料挤出打印扩展到了龙门式运动学之外,但使工艺规划依赖于机器人。在零件坐标系中看似有利的切片器生成的规划,在机械臂上可能变得不可行或在机器人运动学上不利,因为切片工艺决策和零件方向决定了生成的路径,而零件方向和工作空间布置则影响其运动学实现。现有的增材制造工具、基于大语言模型(LLM)的决策支持方法以及数字孪生系统都无法对这些耦合决策提供集成的执行前评估。本文提出了智能体机器人增材制造(A-RAM),这是一种智能体-专家-工具框架,可将用户意图和零件文件转换为可追溯、可执行的规划。LLM解释制造目标和约束,识别规定的和可搜索的规划变量,并将这一推理编码在模式约束的请求中;确定性的规划智能体实例化相应的搜索工作流,同时领域工具为切片、放置、逆运动学、轨迹时间、第六轴关节加加速度(Joint-6 jerk)和挤出计算定量证据。该框架在六轴机械臂增材制造单元上通过三个案例研究进行评估,涵盖专家指定的规划、仅目标规划、目标相关的填充筛选以及几何相关的方向-放置选择。在评估的候选集合中,所选规划相比最不利的有效候选方案,最大第六轴关节加加速度最多降低53.5%,平均绝对第六轴关节加加速度最多降低48.3%;同时,目标相关的填充筛选相比对应的最低效筛选模式,运动规划完成时间最多缩短40.1%,挤出路径最多缩短12.7%。
cs.RO / 15 / 2609.19367

Design and Experimental Validation of a 3D Printed Torsional Series Elastic Actuator for Safe Human Robot Interaction

面向安全人机交互的3D打印扭转串联弹性执行器的设计与实验验证
Pisco, Joel Hidalgo, Condo, Melissa Cobos, Miranda, Luigi, Paillacho, Dennys
Abstract
Ensuring intrinsic safety in physical human robot interaction (pHRI) is a critical requirement for social and service robots. While Series Elastic Actuators (SEAs) offer hardware based compliance, traditional metallic designs often require complex, multi part assemblies. This paper presents the design, finite element analysis (FEA), and experimental validation of a low stiffness, torsional spring for SEAs, manufactured via 3D printed thermoplastic polyurethane (TPU). The compliant element exhibits a highly linear torque deformation response (Ks = 0.066 Nm/degree), matching numerical predictions with under 3% deviation, a variance attributed to FDM structural anisotropy. To accommodate external interactions using standard position limited servomotors, a hybrid position controller with torque threshold switching was implemented. Experimental evaluations demonstrate the system ability to accurately track non stationary trajectories and safely yield to external disturbances. Furthermore, the inherent material damping of the TPU acts as a passive low pass filter, preventing high frequency oscillations during control mode transitions. The proposed architecture offers a cost effective, reliable, and easily manufacturable solution for safe pHRI.
Chinese Translation
确保物理人机交互(pHRI)中的内在安全性是社交机器人和服务机器人的关键需求。虽然串联弹性执行器(Series Elastic Actuators, SEAs)提供了基于硬件的柔顺性,但传统的金属设计通常需要复杂的多部件装配。本文提出了一种用于SEA的低刚度扭转弹簧的设计、有限元分析(FEA)及实验验证,该弹簧采用3D打印热塑性聚氨酯(TPU)制造。该柔性元件表现出高度线性的扭矩-变形响应(Ks = 0.066 Nm/度),与数值预测的偏差小于3%,该偏差归因于FDM打印的结构各向异性。为了在使用标准位置限制型伺服电机的情况下应对外部交互,实现了一种带扭矩阈值切换的混合位置控制器。实验评估表明,该系统能够准确跟踪非平稳轨迹,并能安全地屈服于外部干扰。此外,TPU固有的材料阻尼起到无源低通滤波器的作用,防止了控制模式切换过程中的高频振荡。所提出的架构为安全的pHRI提供了一种经济、可靠且易于制造的解决方案。
cs.RO / 16 / 2609.19375

Enhancing the Perception of Safety and Comfort during Physical Human-Robot Handshake Interactions by Integrating Flexible Elements into a Robotic Arm

通过在机械臂中集成柔性元件提升人机物理握手交互中的安全感与舒适度感知
Hidalgo, Joel, Paillacho, Dennys, Cobos, Melissa, Miranda, Luigi
Abstract
Safety and comfort in human-robot physical in-teractions are essential aspects in the development of social technologies, where natural gestures, such as handshakes, represent a challenge due to their direct physical contact. The implementation of series elastic actuators (SEA) to absorb impacts is proposed as a design strategy that favors safer interactions. This paper presents an experimental study aimed at evaluating how the incorporation of SEAs in robotic arms influences perceived safety and the interaction experience dur-ing handshaking. The design allows a direct comparison of the effect of rigidity versus the incorporation of elastic elements, in order to identify the advantages of SEAs in improving the physical safety and social acceptance of robotic systems in everyday contexts. The experiment was carried out with 10 volunteers (6 men and 4 women), who performed two interactions with each robotic arm: one with rigid joints and the other with flexible joints using SEA. During testing, objective data on end-effector trajectories were collected, as well as subjective information through a perception survey focused on safety, naturalness, and confidence during the handshake. The survey results show increased perceptions of safety and comfort with the SEA-equipped arm, supporting its potential to facilitate safer and more socially accepted human-robot interactions.
Chinese Translation
人机物理交互中的安全性与舒适性是社交技术发展中的关键要素,而握手等自然手势由于其直接的物理接触而构成一项挑战。本文提出采用串联弹性驱动器(SEA)来吸收冲击,作为一种有利于实现更安全交互的设计策略。本研究开展了一项实验,旨在评估在机械臂中引入SEA如何影响握手过程中的安全感知及交互体验。该设计能够直接比较刚性结构与弹性元件引入的效果,从而识别SEA在提升机器人系统日常情境中的物理安全性和社会接受度方面的优势。实验由10名志愿者(6男4女)参与,每位志愿者分别与两种机械臂进行交互:一种为刚性关节,另一种为采用SEA的柔性关节。测试过程中收集了末端执行器轨迹的客观数据,并通过一项聚焦于握手过程中安全性、自然度和信任度的感知问卷调查获取了主观信息。调查结果显示,配备SEA的机械臂提升了安全感和舒适感感知,支持了其在促进更安全、更易被社会接受的人机交互方面的潜力。
cs.RO / 17 / 2609.19378

Task-Oriented Active Learning of Residual Dynamics for Model Predictive Path Integral Control

面向任务的残差动力学主动学习用于模型预测路径积分控制
Aoki, Nobuaki, Lee, Hojin, Sosnowski, Stefan, Hirche, Sandra
Abstract
Online residual learning can reduce model mismatch in predictive control, but passive data collection may fail to adequately cover states that become important later in the task. Task-agnostic active learning targets uncertain or informative regions, but information acquired in such regions does not necessarily improve task performance. This paper introduces Task-Oriented Information Acquisition (ToIA), an active-learning criterion for model predictive path integral control (MPPI) with online Gaussian process (GP) residual learning. For each sampled control sequence, ToIA estimates how much an observation obtained early in the rollout would reduce predictive uncertainty at later states on the same rollout, and weights this reduction by the rollout's relevance to the task. The score is evaluated over the existing MPPI rollout batch without sampling future observations or re-optimizing control under hypothetical posterior updates. In simulated off-road navigation across held-out maps with heterogeneous terrain, ToIA improved the goal-reaching success rate over passive GP learning by 19.3 and 27.4 percentage points and outperformed task-agnostic active-learning baselines across dense and sparse online-learning intervals. An ablation study indicates that task relevance is particularly important under sparse model updates. The implementation supports online control at 20 Hz on an NVIDIA RTX 2080 Ti.
Chinese Translation
在线残差学习可以减少预测控制中的模型失配,但被动式数据采集可能无法充分覆盖任务后期变得重要的状态区域。与任务无关的主动学习以不确定或信息丰富的区域为目标,但在这些区域获取的信息并不一定能提升任务性能。本文提出了面向任务的信息采集方法(Task-Oriented Information Acquisition, ToIA),这是一种用于带在线高斯过程(GP)残差学习的模型预测路径积分控制(MPPI)的主动学习准则。对于每条采样的控制序列,ToIA 估计在轨迹展开早期获得的观测能在多大程度上降低同一轨迹展开后期状态的预测不确定性,并根据该轨迹展开与任务的相关性对这一不确定性降低量进行加权。该评分直接在现有的 MPPI 轨迹展开批次上计算,无需采样未来观测,也无需在假设后验更新下重新优化控制。在不同地形异构的留出地图上进行的越野导航仿真中,ToIA 相较于被动 GP 学习将目标到达成功率分别提高了 19.3 和 27.4 个百分点,并在稠密与稀疏的在线学习间隔设置下均优于与任务无关的主动学习基线。消融实验表明,任务相关性在模型更新稀疏的情况下尤为重要。该实现可在 NVIDIA RTX 2080 Ti 上以 20 Hz 的频率支持在线控制。
cs.RO / 18 / 2609.19413

From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

从轨迹执行到场景重置:面向自主长时程操作评估的图基框架
Jiang, Jing, Yang, Yue, Jiang, Xinkai, Bertasius, Gedas, Szafir, Daniel J., Lioutikov, Rudolf
Abstract
Robot manipulation policies are improving quickly, and real-robot evaluation remains the standard evidence for that progress. It still relies on a human to reset the scene between rollouts, which consumes operator time and leaves the initial state distribution unspecified, so results reproduce poorly. A recent system, AutoEval, automates both reset and scoring, but only for single-step tasks, because a long-horizon rollout can terminate in combinatorially many configurations that no single learned reset policy covers. We present HALTER, a Harness for Autonomous Long-horizon Task Evaluation and Reset, which restores the scene by planning over a library of learned atomic reset skills, so demonstration cost scales with the size of that library rather than with the number of terminal states. HALTER builds a spatial scene graph online from point clouds and vision foundation models, and an LLM reasons over this graph to score the rollout, plan the reset, and verify that the reset succeeded, without collecting labeled success images for any task. On four long-horizon tasks on a Franka arm, HALTER restores the scene in 76% of episodes, against 52% for AutoEval and 65% for a motion-planning reset, and it estimates the completed-skill fraction correctly in 90% of episodes, against 76%. Its reset-verification verdict is correct in 91% of episodes, compared with 78% for AutoEval. It also cuts the operator time of an evaluation campaign by 72% relative to manual reset. We further measure compositional generalization on three held-out tasks, where HALTER resets 74.7% of episodes against 1.3% for a per-task reset policy, and we ablate the scene representation and the graph update rate.
Chinese Translation
机器人操作策略正在快速进步,而真机评估仍是衡量这一进展的标准证据。然而,现有评估仍依赖人工在每次轨迹执行之间重置场景,这不仅耗费操作员时间,还导致初始状态分布不明确,使实验结果难以复现。近期提出的 AutoEval 系统实现了重置与评分的自动化,但仅适用于单步任务,因为长时程轨迹执行可能终止于组合级数量众多的配置,而单一的学习型重置策略无法覆盖所有情况。我们提出 HALTER(自主长时程任务评估与重置框架),通过在预训练的原子重置技能库上进行规划来恢复场景,因此演示数据的收集成本随技能库规模而非终止状态数量增长。HALTER 基于点云和视觉基础模型在线构建空间场景图,并由大语言模型(LLM)对该图进行推理,以对轨迹执行进行评分、规划重置并验证重置是否成功,且无需为任何任务收集带标签的成功图像。在 Franka 机械臂上的四个长时程任务中,HALTER 在 76% 的回合中成功恢复场景,而 AutoEval 为 52%,运动规划重置方法为 65%;HALTER 在 90% 的回合中正确估计了已完成技能的比例,而对比方法为 76%。其重置验证判定在 91% 的回合中正确,相比之下 AutoEval 为 78%。与人工重置相比,HALTER 还将评估活动中的操作员时间减少了 72%。我们进一步在三个保留任务上测量了组合泛化能力,HALTER 重置了 74.7% 的回合,而针对单一任务训练的重置策略仅为 1.3%;此外,我们还对场景表示和图更新频率进行了消融实验。
cs.RO / 19 / 2609.19441

Predict Before You Deploy: Offline Prediction of Quantization-Induced Task Degradation for World Action Models

部署前预测:面向世界动作模型的量化引发任务退化离线预测
Xu, Jiuyi, Guo, Jinjia, Chen, Meida, Du, Jing, Shi, Yangming
Abstract
World action models (WAMs) rely on video-generation backbones, requiring substantial memory and compute for deployment. Post-training quantization reduces memory and can accelerate inference, but bit width, grouping, and quantizer choice define a large configuration space. Identifying configurations that preserve task performance through exhaustive closed-loop evaluation is costly. We propose PreDE (Predict Before You Deploy), a policy-calibrated framework for predicting quantization-induced task degradation from offline action deviations. Using closed-loop outcomes from a small development set, PreDE calibrates two thresholds and accepts, rejects, or defers new configurations using a fixed observation log. Under a within-setting label-ordering hypothesis, the rule issues decisions where all thresholds consistent with the development labels agree. Across five WAMs and four benchmark settings, quantization produces configuration-dependent task losses that cannot be explained by bit width alone or a shared deviation threshold. Across 28 held-out configurations from two policies, PreDE issued 21 decisions before observing closed-loop outcomes (75% coverage), all matching the observed acceptable or degraded labels. Deferred candidates included both acceptable outcomes and a 33-percentage-point loss. In 450 Franka Research 3 trials across two independently fine-tuned policies, all configurations assigned to high-deviation groups before testing showed significant degradation, while low-deviation comparisons showed no statistically significant degradation. On the real robot, W4A4 achieved a 1.37x action-query speedup and approximately 44% lower peak memory. These results support policy-specific behavioral calibration for quantization configuration selection while identifying candidates that require closed-loop evaluation. The code is available at https://github.com/jiuyixu25/PreDE.
Chinese Translation
世界动作模型(World Action Models, WAMs)依赖于视频生成骨干网络,部署时需要大量的内存和计算资源。训练后量化(Post-training quantization)可以减少内存占用并加速推理,但位宽、分组和量化器的选择构成了一个庞大的配置空间。通过穷举式闭环评估来识别能保持任务性能的配置,其代价高昂。我们提出 PreDE(Predict Before You Deploy),一个策略校准框架,可基于离线动作偏差预测量化引发的任务退化。PreDE 利用一个小规模开发集的闭环结果校准两个阈值,并基于固定的观测日志对新配置做出接受、拒绝或延后决策。在“设定内标签排序”假设下,该规则仅在所有与开发集标签一致的阈值均得出相同结论时才给出决策。在五个 WAM 和四个基准设定上的实验表明,量化导致的任务退化依赖于具体配置,无法仅用位宽或统一的偏差阈值来解释。在来自两个策略的 28 个保留配置中,PreDE 在观察闭环结果之前对其中 21 个做出了决策(75% 的覆盖率),且全部与实际观测到的“可接受”或“退化”标签一致。被延后的候选配置中既包含可接受的结果,也包含一个退化达 33 个百分点的案例。在两个独立微调策略共 450 次 Franka Research 3 机器人试验中,所有在测试前被归入高偏差组的配置均表现出显著退化,而低偏差组的对比则未表现出统计上显著的退化。在真实机器人上,W4A4 实现了 1.37 倍的动作查询加速,峰值内存降低约 44%。这些结果支持在量化配置选择中采用策略特定的行为校准,同时识别出需要闭环评估的候选配置。代码发布于 https://github.com/jiuyixu25/PreDE。
cs.RO / 20 / 2609.19447

From Wizard-of-Oz Human-Robot Dialogue Collection to a Taxonomy of Robot Response Decisions: A Retrospective Analysis of Assistive Pilot Interactions

从巫师奥兹(Wizard-of-Oz)人机对话收集到机器人响应决策分类体系:辅助性试点交互的回顾性分析
Liu, Guangping, Hawkins, Nicholas, Sultan, Tipu, Esposito, Flavio, Dian, Madi
Abstract
Robots that follow natural-language instructions in everyday indoor environments must act on incomplete human utterances. Instructions often omit essential information, such as the identity of an out-of-view object, an intended destination, or the user's goal. Existing datasets contain little real-world situated dialogue and provide few practice-grounded criteria for deciding when a robot should act, confirm, clarify, or refuse. We retrospectively analyze a pilot Wizard-of-Oz study in which five participants performed everyday indoor tasks, including door opening, drawer opening, feeding, drinking, and cleaning, with a wheelchair-mounted mobile manipulator while the wizard responded without a formal communication policy. This preserved authentic user behavior but produced inconsistent robot-side decisions, motivating an explicit decision scheme. From 40 episodes, we derived a hierarchical taxonomy of six response modes (ANSWER, REPORT_DONE, REFUSE, CONFIRM, CLARIFY, ACT) and four ambiguity types (intent, referential, spatial, intelligibility). Two human annotators and an AI annotator applied the scheme to the pilot data. Clean-label rates were 91% and 89%, and Cohen's ranged from 0.72 to 0.95 across decision-point, mode, and ambiguity levels for both human-human and human-AI comparisons. Fine-tuning LLaVA-1.6-7B on taxonomy-derived labels for ACT and CLARIFY indicates the feasibility of training vision-language models using annotations from our taxonomy. Remaining boundary cases in decision-point identification and REPORT_DONE motivate a constrained protocol for more consistent dialogue collection.
Chinese Translation
在日常室内环境中遵循自然语言指令的机器人必须能够针对不完整的人类话语采取行动。指令往往遗漏关键信息,例如视野外物体的身份、预期目的地或用户的目标。现有数据集几乎不包含真实世界的情境化对话,也缺乏基于实践的准则来判断机器人何时应当行动、确认、澄清或拒绝。我们对一项试点性巫师奥兹(Wizard-of-Oz)研究进行了回顾性分析:五名参与者使用安装在轮椅上的移动机械臂执行日常室内任务,包括开门、开抽屉、喂食、饮水和清洁,而巫师在无正式沟通策略的情况下进行响应。这种设置保留了真实的用户行为,但导致了机器人侧决策的不一致,因而需要一个明确的决策方案。基于40个任务片段,我们构建了一个分层分类体系,包括六种响应模式(ANSWER、REPORT_DONE、REFUSE、CONFIRM、CLARIFY、ACT)和四种歧义类型(意图歧义、指代歧义、空间歧义、可理解性歧义)。两名人类标注者和一名AI标注者将该方案应用于试点数据。在决策点、模式和歧义三个层面,人-人及人-AI比较的标注准确率分别为91%和89%,Cohen's κ系数介于0.72至0.95之间。在分类体系导出的ACT和CLARIFY标签上对LLaVA-1.6-7B进行微调,表明利用我们所提出分类体系的标注训练视觉-语言模型是可行的。决策点识别和REPORT_DONE中遗留的边界案例促使我们设计一种受约束的协议,以实现更一致的对话收集。
cs.RO / 21 / 2609.19449

Winning a Won Game: Strict Reach-Avoid-Stay Control Barrier Functions for High-Dimensional Black-Box Systems

胜局必胜:面向高维黑盒系统的严格到达-避障-驻留控制屏障函数
Oh, Donggeon David, Nguyen, Duy P., Yuan, Gongkai, Li, Qingchen, Fisac, Jaime Fernández, Hu, Haimin
Abstract
Robots must complete their tasks and maintain the achieved outcomes while avoiding safety failures at all times. Strict reach-avoid-stay (sRAS) formalizes this requirement: safely reaching a target and remaining there indefinitely after first entry. We propose an sRAS Q-control barrier function (CBF) safety filter for high-dimensional black-box systems under bounded uncertainty. Our construction combines a stay value encoding safe permanent residence in a target subset with a reach-avoid value encoding safe reachability of this subset while avoiding target states from which safe permanent residence cannot be guaranteed. We prove that these values jointly yield a valid robust discrete-time CBF and lift them to state-action Q-functions for runtime intervention. For exact values and under a measure-zero condition, our filter preserves sRAS feasibility from almost every winnable initial state and keeps the system safely within the target after first entry, against all admissible uncertainty realizations. We adopt reachability-based adversarial reinforcement learning for scalable value approximation using only black-box interactions. Notably, neither synthesis nor deployment of our filter requires known dynamics, affine structure, value derivatives, or hand-designed barriers. We validate our framework in quadruped gap jumping in simulation and hardware, where the robot crosses the gap, lands safely, and remains safe afterward. Simulated F1TENTH races further demonstrate safe overtaking and lead retention.
Chinese Translation
机器人必须完成任务并保持已取得的成果,同时始终避免发生安全性失效。严格到达-避障-驻留(strict reach-avoid-stay, sRAS)形式化了这一要求:安全地到达目标区域,并在首次进入后无限期地驻留其中。我们提出了一种针对有界不确定性下高维黑盒系统的 sRAS Q-控制屏障函数(CBF)安全滤波器。我们的构造将编码在目标子集中安全永久驻留的驻留值(stay value),与编码在避开那些无法保证安全永久驻留的目标状态的同时安全到达该子集的到达-避障值(reach-avoid value)相结合。我们证明这两个值共同构成一个有效的鲁棒离散时间 CBF,并将其提升为状态-动作 Q 函数以用于运行时干预。对于精确值,在测度为零的条件下,我们的滤波器能从几乎所有可赢的初始状态保持 sRAS 可行性,并在首次进入后将系统安全地保持在目标区域内,抵御所有可容许的不确定性实现。我们采用基于可达性的对抗强化学习,仅利用黑盒交互实现可扩展的价值近似。值得注意的是,我们的滤波器在综合与部署阶段均不需要已知的动力学模型、仿射结构、价值函数的导数或人工设计的屏障函数。我们在仿真和硬件中通过四足机器人跨越间隙实验验证了该框架:机器人成功跨越间隙、安全着陆并在之后保持安全。F1TENTH 赛车仿真进一步展示了安全超车和领先保持能力。
cs.RO / 22 / 2609.19452

GLAMDRING: Gait Learning And Morphology co-Design via Reinforcement LearnING of CPGs

GLAMDRING:基于强化学习的中枢模式发生器步态学习与形态协同设计
Joshi, Amogh, Roy, Kaushik
Abstract
Robots are moving out of the structured factory floor and into unstructured environments such as disaster sites, planetary surfaces, and agricultural fields, for which the right robot often does not yet exist. We present GLAMDRING, a framework that synthesizes the optimal robot for a locomotion task and, jointly, learns the controller that drives it. For the given specifications of forward-velocity bounds, a per-actuator power budget, an actuator library, and a payload requirement, GLAMDRING returns a matched quadruped morphology (link geometry and per-joint actuators) and a Hopf-oscillator Central Pattern Generator (CPG) gait policy. We rank feasible designs against a target design objective, viz., maximum speed, minimum Cost of Transport (CoT), or max Payload Margin. Because body and locomotion are coupled, the optimal morphology dictates how a robot is driven, while optimal gait depends on the physical body. We train a small number of CPG policies by reinforcement learning across the space of candidate morphologies, co-learning the gait with the underlying robot hardware. Link lengths and actuators are then resolved post-hoc from the policy's logged operating envelope, reducing synthesis cost to a small, fixed number of reinforcement-learning runs instead of one per candidate. Our experiments show three key findings: co-designing body and gait is necessary to satisfy locomotion constraints; actuator-envelope feasibility, rather than locomotion success alone, determines realizable payload capacity; and canonical animal gaits emerge naturally in most designs from morphology and constraints alone. A real-world demonstration further highlights the efficacy of our work.
Chinese Translation
机器人正从结构化的工厂车间走向非结构化环境,如灾难现场、行星表面和农业用地,而针对这些环境的合适机器人往往尚未存在。我们提出了GLAMDRING,一个能够为运动任务综合出最优机器人并同时学习其控制器的框架。对于给定的前向速度界限、单执行器功率预算、执行器库和有效载荷要求等规格,GLAMDRING返回一个匹配的四足机器人形态(连杆几何结构和各关节执行器)以及一个基于Hopf振荡器的中枢模式发生器(CPG)步态策略。我们根据目标设计目标对可行设计进行排序,目标包括最大速度、最小运输成本或最大有效载荷余量。由于机体与运动相互耦合,最优形态决定了机器人的驱动方式,而最优步态又依赖于物理机体。我们在候选形态空间中通过强化学习训练少量CPG策略,在习得步态的同时协同学习底层机器人硬件。随后根据策略记录的工作包络事后确定连杆长度和执行器,从而将综合成本降低为少量固定的强化学习运行次数,而无需为每个候选设计单独运行。我们的实验展示了三个关键发现:机体与步态的协同设计是满足运动约束的必要条件;决定可实现有效载荷能力的是执行器包络的可行性,而非仅仅是运动成功率;且在大多数设计中,典型的动物步态仅凭形态和约束即可自然涌现。真实世界的演示进一步验证了我们工作的有效性。
cs.RO / 23 / 2609.19460

Pose-aware Legged Robot Semantic Exploration with Omnidirectional Perception in Confined Unknown Environments

受限未知环境中基于全向感知的姿态感知腿式机器人语义探索
Zhan, Xiaoyang, Chen, Shiyu, Shimada, Kenji
Abstract
Semantic exploration in confined environments requires both environment mapping and detailed observation of target objects. For ground robots, limited sensor vertical fields of view and restricted standoff distances can leave upper object surfaces unobserved from planar viewpoints. Body tilting can improve coverage, but additional observations and posture transitions increase mission time. To address this trade-off, we present POSE, a pose-aware semantic exploration system that exploits a legged robot's intrinsic body pitch and roll with omnidirectional camera-LiDAR perception. The proposed pose-aware viewpoint sampling module selects body postures from partial object maps according to expected coverage gain, while aim-aligned execution reduces unnecessary body reorientation. Further, we introduce an object-centric viewpoint pruning strategy assisted by a vision-language model (VLM), which uses persistent observation history and bird's-eye-view (BEV) maps to reduce redundant inspection visits. The resulting semantic viewpoints are combined with geometric exploration viewpoints in a global exploration planner. Simulations show that POSE improves final target-surface coverage by 8-10 percentage points over the planar planning baseline while reducing exploration time by 17-32%, and achieves the highest mean object coverage AUC among the evaluated baselines. Real-world experiments with a legged robot carrying an omnidirectional camera-LiDAR suite in a machine shop further demonstrate the system's applicability. These results support adaptive body-posture planning for improving the coverage-efficiency trade-off in legged robot semantic exploration. We plan to release the code for community benefit in the future.
Chinese Translation
受限环境中的语义探索既需要环境建图,也需要对目标物体进行细致观测。对于地面机器人而言,传感器垂直视场受限以及观测距离受限,导致从平面视角无法观测到物体的上表面。通过身体倾斜可以提升覆盖范围,但额外的观测和姿态转换会增加任务时间。为解决这一权衡问题,我们提出了POSE,一种姿态感知的语义探索系统,该系统利用腿式机器人固有的身体俯仰和翻滚能力,结合全向相机-激光雷达(camera-LiDAR)感知。所提出的姿态感知视点采样模块根据预期覆盖增益从部分物体地图中选择身体姿态,同时目标对准的执行方式减少了不必要的身体重定向。此外,我们引入了一种由视觉-语言模型(VLM)辅助的以物体为中心的视点剪枝策略,该策略利用持久化的观测历史和鸟瞰图(BEV)来减少冗余的检测访问。所得到的语义视点与几何探索视点在全局探索规划器中相结合。仿真结果表明,与平面规划基线相比,POSE将最终目标表面覆盖率提高了8-10个百分点,同时将探索时间缩短了17-32%,并在所有评估基线中取得了最高的平均物体覆盖AUC。在机加工车间中,搭载全向相机-激光雷达套件的腿式机器人真实世界实验进一步验证了该系统的实用性。这些结果支持通过自适应身体姿态规划来改善腿式机器人语义探索中覆盖范围与效率的权衡。我们计划在未来发布代码以供社区使用。
cs.RO / 24 / 2609.19475

FASA: Feedback-Aware Sampling Adaptation for Efficient Diffusion-Based VLA Models

FASA:面向高效扩散式视觉-语言-动作模型的反馈感知采样自适应方法
Han, Yuchen, Wu, Jianhan, Qu, Xiaoyang, Kong, Lingwei, Li, Shiyi, Wang, Jianzong
Abstract
Diffusion-based Vision-Language-Action (VLA) models achieve strong performance in embodied tasks, but their iterative sampling imposes heavy computational and memory-access cost, blocking real-time deployment on edge platforms. Existing acceleration methods either require expensive training (e.g., distillation, flow matching) or degrade perception via statically scheduled pruning and caching, ignoring the dynamic workload variance of robotic interactions. This paper presents FASA (Feedback-Aware Sampling Adaptation), a training-free runtime framework that treats real-time multimodal feedback as a control signal for the denoising pipeline: an interaction-driven range adaptor modulates the global sampling-step budget based on visual and gripper-force feedback, and a proprioception-aware step adaptor pinpoints the optimized step within the adapted range. This co-designed framework allows the underlying hardware architecture to adaptively match the workload demands of different execution phases. Comparative evaluations across several benchmarks show that the inference speed can be increased by up to 1.45$\times$ while maintaining competitive success rates, providing a novel dynamic runtime architecture paradigm for deploying heavy generative embodied AI workloads onto resource-constrained computing platforms.
Chinese Translation
基于扩散模型的视觉-语言-动作(Vision-Language-Action, VLA)模型在具身任务中表现出色,但其迭代采样过程带来高昂的计算和内存访问开销,阻碍了在边缘平台上的实时部署。现有的加速方法要么需要代价高昂的训练(如蒸馏、流匹配),要么通过静态调度的剪枝和缓存降低感知性能,忽略了机器人交互中动态的工作负载变化。本文提出FASA(反馈感知采样自适应,Feedback-Aware Sampling Adaptation),一种无需训练的运行时框架,它将实时多模态反馈视为去噪流水线的控制信号:交互驱动的范围调节器根据视觉和夹爪力反馈调制全局采样步数预算,而本体感知驱动的步数调节器则在适配后的范围内精确定位优化的步数。这一协同设计的框架使底层硬件架构能够自适应地匹配不同执行阶段的工作负载需求。在多个基准上的对比评估表明,推理速度可提升至1.45倍,同时保持具有竞争力的成功率,为将重型生成式具身AI工作负载部署到资源受限的计算平台提供了一种新颖的动态运行时架构范式。
cs.RO / 25 / 2609.19480

Self-excited actuation enables adaptive and resilient flapping-wing flight

自激驱动实现自适应且具韧性的扑翼飞行
Yang, Rundong, Wold, Ethan S., Liu, Ellen, Lynch, James, Zhou, Wei, Jankauski, Mark, Sponberg, Simon, Gravish, Nick
Abstract
The muscles that power insect flight fall into one of two categories: 1) synchronous muscles that contract under direct control from the nervous system, and 2) asynchronous muscles which have an intrinsic stretch activation response that spontaneously generates wingbeats without the need for signaling from the brain. It is thought that the emergent nature of asynchronous wingbeats provides both adaptive and responsive capabilities for flight control. To date, most flying robots use synchronous actuation. In this paper we develop the first flight-capable flapping wing robot that uses asynchronous actuation. We demonstrate that asynchronous actuation allows wings to respond to changes in the resonant mechanics of the body without control input, and wings can react instantaneously to collisions with obstacles with no extrinsic sensing needed. Flight tests within cluttered environments demonstrate that asynchronous actuation significantly improves stability and performance when compared to synchronous actuation. In total this work demonstrates that a flapping wing robot actuation strategy that emulates the asynchronous muscles of flying insects can provide fast, reactive actuation responses before a control system would need to intervene. This partitioning of embodied control to both the low-level actuation dynamics and and high-level sensorimotor system provides a compelling blueprint for new flying robots.
Chinese Translation
为昆虫飞行提供动力的肌肉可分为两类:1)同步肌肉,在神经系统的直接控制下收缩;2)异步肌肉,具有内在的拉伸激活响应,能够在无需大脑信号的情况下自发产生翅膀拍动。人们认为,异步拍翅的涌现特性为飞行控制提供了自适应和快速响应的能力。迄今为止,大多数飞行机器人采用同步驱动。本文研制了首个采用异步驱动且具备飞行能力的扑翼机器人。我们证明,异步驱动使翅膀能够在无控制输入的情况下响应机体共振力学的变化,并且翅膀能够在无需外部传感的情况下对障碍物碰撞做出瞬时反应。在复杂环境中的飞行测试表明,与同步驱动相比,异步驱动显著提高了稳定性和性能。总体而言,这项工作表明,模仿飞行昆虫异步肌肉的扑翼机器人驱动策略,能够在控制系统介入之前提供快速的反应性驱动响应。将这种具身控制分配给低层驱动动力学和高层感觉运动系统,为新型飞行机器人提供了一个引人注目的设计蓝图。
cs.RO / 26 / 2609.19510

PIVOT: Perception-aware Independent Viewpoint Online Optimization

PIVOT:感知感知的独立视点在线优化
Chen, Yuyang, Sadeghi, Shekoufeh, Adhivarahan, Charuvahan, Lemos, Elton, Wang, Chen, Koppal, Sanjeev J., Dantu, Karthik
Abstract
A fundamental assumption in robotic perception is that the sensor's field of view (FoV) is fixed relative to the robot body. Motion-decoupled sensors, such as gimbal-mounted cameras and MEMS-based LiDARs, instead allow sensing direction to be controlled independently at runtime. This freedom creates a computational challenge: efficiently selecting useful viewing directions online in feature-dense environments. We propose PIVOT, a lightweight iterative method that optimizes sensor viewing direction along a fixed translation trajectory to maximize feature visibility. Under a conical FoV model, visibility depends only on the optical axis, yielding a two-degree-of-freedom optimization on the viewing sphere $S^2$. Coordinate-free $SO(3)$ exponential-map updates enable efficient continuous optimization without explicit angular parameterizations or exhaustive viewing-sphere search. Monte Carlo evaluations retain 98.1--99.6% of brute-force visibility with a 76--85x speedup. Photorealistic simulation and real-world experiments further demonstrate improved visual localization robustness and practical viewpoint control on a quadruped robot.
Chinese Translation
机器人感知中的一个基本假设是传感器的视场角(FoV)相对于机器人本体是固定的。而运动解耦传感器,如云台相机和基于MEMS的激光雷达,允许在运行时独立控制感知方向。这种自由度带来了一个计算挑战:如何在特征密集的环境中在线高效地选择有用的观察方向。我们提出PIVOT,一种轻量级的迭代方法,通过优化传感器观察方向(沿固定平移轨迹)以最大化特征可见性。在锥形视场模型下,可见性仅取决于光轴,从而将问题转化为观察球面 $S^2$ 上的二自由度优化。基于无坐标的 $SO(3)$ 指数映射更新,实现了高效的连续优化,无需显式角度参数化或对观察球面进行穷举搜索。蒙特卡洛评估表明,该方法保留了98.1%至99.6%的穷举搜索可见性,同时实现了76至85倍的速度提升。照片级逼真仿真和真实世界实验进一步证明,该方法提高了四足机器人上视觉定位的鲁棒性,并实现了实用的视点控制。
cs.RO / 27 / 2609.19512

CoreSense: Traceable Failure Recall and Conflict-Aware Belief Gating for Auditable Robot Decisions

CoreSense:面向可审计机器人决策的可追溯失败回忆与冲突感知信念门控
Li, Zoe
Abstract
Robots can recall prior failures without knowing whether recalled evidence remains valid, conflicts with current observations, or is sufficient to guide a decision. We present CoreSense, a robot-system integration architecture that combines traceable episodic evidence with a conflict-aware belief gate and bounded, auditable recommendations. The gate checks scope, provenance, time, contradiction, and support before it permits PROCEED, requests re-observation, abstains, or escalates. Evaluation follows three complementary layers without commanding a physical robot: offline public real-robot data, a frozen signal-level simulation, and a live cloud deployment path. On CableTrace-120 and BotFails-200, belief gating reduces protocol-defined unsafe proceeds from 20% and 40% to 0%. A disjointly calibrated raw-video policy also reaches 0% unsafe proceed, but overblocks every nominal episode. On public data, a ViFailback-BotFails visual detector reaches 0.778 AUROC yet remains all-blocking, whereas cycle-disjoint UR3 telemetry for protective stops yields 0% unsafe proceed, 36.1% overblocking, and 61.9% coverage; grip-loss transfer remains a negative result. Controlled physical corroboration yields 3.3%, 0%, and 42.0%, while conflict-aware fusion yields 4.7%, 0%, and 42.8%. Finally, 20/20 cloud recalls validate a CockroachDB Cloud-Amazon Bedrock deployment path. The evidence supports an auditable integration pattern, not autonomous recovery or certified safety.
Chinese Translation
机器人可以回忆起先前的失败经历,但无法确定所回忆的证据是否仍然有效、是否与当前观测相冲突,或是否足以指导决策。我们提出了CoreSense,一种机器人系统集成架构,它将可追溯的情景证据与冲突感知信念门控及有界、可审计的推荐相结合。该门控在允许PROCEED(继续执行)、请求重新观测、弃权或上报之前,会检查范围、来源、时间、矛盾性和支持度。评估遵循三个互补层面,而不指挥物理机器人:离线公开真实机器人数据、冻结的信号级仿真,以及实时云端部署路径。在CableTrace-120和BotFails-200数据集上,信念门控将协议定义的不安全继续执行比例从20%和40%降至0%。一个分离校准的原始视频策略同样达到0%的不安全继续执行,但在所有正常场景中均过度阻断。在公开数据上,ViFailback-BotFails视觉检测器达到0.778的AUROC,但仍然全阻断;而针对保护性停机的周期分离UR3遥测数据实现了0%的不安全继续执行、36.1%的过度阻断和61.9%的覆盖率;抓取丢失的迁移仍是一个负面结果。受控的物理验证得到3.3%、0%和42.0%的结果,而冲突感知融合得到4.7%、0%和42.8%的结果。最后,20/20的云端回忆验证了CockroachDB Cloud-Amazon Bedrock部署路径。这些证据支持的是一种可审计的集成模式,而非自主恢复或经认证的安全性。
cs.RO / 28 / 2609.19527

AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation

AURORA:一个自然语言驱动的智能体框架,用于理解、推理与编排可靠的空地协同仿真
Wu, Keshu, Zhang, Hao, Gan, Rui, Gao, Xiangbo, Li, Xiaopeng, Tu, Zhengzhong, Zhou, Yang
Abstract
Air-ground transportation research increasingly relies on co-simulation, yet constructing scenarios remains labor-intensive and difficult to validate. More importantly, a generated scenario may execute successfully while failing to realize the spatial, temporal, communication, or behavioral relationships requested by the user. This paper presents AURORA, a natural-language-driven agentic framework that treats air-ground scenario generation as a process of compilation with verification. Central to AURORA is the Air-Ground Scenario Graph (AGSG), a typed intermediate representation that explicitly connects agents, aerial missions, events, communication links, success conditions, and their cross-domain dependencies. This shared representation enables simulator-grounded parsing, joint road-airspace grounding, temporal planning, pre-execution feasibility checking, trace-based runtime verification, failure localization, and bounded repair within a unified workflow. We further introduce AURORA-Bench to evaluate not only whether generated scenarios execute, but whether they faithfully realize the requested interactions. Experiments across multiple language models show that structured execution substantially improves reliability, while runtime verification exposes silent failures that completion-based evaluation overlooks. Localized repair further resolves many violations without regenerating the entire scenario. The results show that reliable scenario generation requires verifying realized behavior, not merely executable code, and demonstrate the value of explicit intermediate representations for verifiable and repairable language-driven co-simulation.
Chinese Translation
空地交通研究日益依赖协同仿真,然而场景构建仍然费时费力且难以验证。更重要的是,生成的场景可能执行成功,却未能实现用户所要求的空间、时间、通信或行为关系。本文提出AURORA,一个自然语言驱动的智能体框架,将空地场景生成视为一个带验证的编译过程。AURORA的核心是空地场景图(Air-Ground Scenario Graph, AGSG),这是一种类型化的中间表示,显式地连接智能体、空中任务、事件、通信链路、成功条件及其跨域依赖关系。这一共享表示使得基于仿真器的解析、道路与空域的联合接地、时间规划、执行前可行性检查、基于轨迹的运行时验证、故障定位以及有界修复能够在统一的工作流中完成。我们进一步引入AURORA-Bench,不仅评估生成的场景能否执行,还评估其是否忠实地实现了所要求的交互。在多个语言模型上的实验表明,结构化执行显著提高了可靠性,而运行时验证能够揭示基于完成度的评估所忽略的静默失败。局部化修复进一步在无需重新生成整个场景的情况下解决了许多违约问题。结果表明,可靠的场景生成需要验证已实现的行为,而不仅仅是可执行的代码,并证明了显式中间表示对于可验证、可修复的语言驱动协同仿真的价值。
cs.RO / 29 / 2609.19533

SLAMSqueezeBench: Comparing SLAM Systems under Resource Constraints

SLAMSqueezeBench:资源受限条件下的SLAM系统比较
Hefny, Mohamed, Dantu, Karthik, Ko, Steven Y.
Abstract
Simultaneous localization and mapping (SLAM) is one of the services running on an autonomous robot. It is typically run to assist other tasks such as planning, manipulation, etc. All these tasks are run on edge hardware and are subject to severe resource constraints. However, most SLAM systems are built and tested in isolation, and their performance is reported as if they are the only task running on a system. We observe that existing benchmarks lack a common mechanism for comparing SLAM systems under realistic resource constraints. To address this limitation, we have developed SLAMSqueezeBench, a framework that allows testing of SLAM systems under realistic workloads on edge hardware. It does so by imposing constraints on compute and memory resources available for the SLAM system during execution. It also simulates realistic camera frame acquisition with frame drops when a finite buffer is full. Using SLAMSqueezeBench, we compare nine SLAM systems spanning classical systems, learning-based systems, and approaches for Gaussian splatting. Our testing framework will be available for use by the community upon publication.
Chinese Translation
同时定位与建图(SLAM)是自主机器人上运行的服务之一,通常用于辅助规划、操作等其他任务。所有这些任务都运行在边缘硬件上,并受到严格的资源限制。然而,大多数SLAM系统都是在隔离环境下构建和测试的,其性能报告仿佛它们是系统中唯一运行的任务。我们观察到,现有基准测试缺乏在现实资源约束下比较SLAM系统的通用机制。为解决这一局限,我们开发了SLAMSqueezeBench,一个允许在边缘硬件上以真实工作负载测试SLAM系统的框架。它通过在SLAM系统执行期间对其可用的计算和内存资源施加约束来实现这一目标。该框架还模拟了现实的相机帧采集过程,包括有限缓冲区满时的丢帧情况。利用SLAMSqueezeBench,我们比较了九个SLAM系统,涵盖经典系统、基于学习的系统以及高斯泼溅(Gaussian Splatting)方法。我们的测试框架将在论文发表后向社区开放使用。
cs.RO / 30 / 2609.19541

Navigate or Relocate? Planning Among Movable Obstacles in Unknown Environments

导航还是搬移?未知环境中可移动障碍物间的规划问题
Zhang, Yuqing, Zhu, Haoyu, Kantaros, Yiannis
Abstract
Conventional robot planning methods seek collision-free paths to a goal but fail when all paths are blocked. In these cases, the robot must determine which objects to relocate, in what order, and where to place them to clear a path---a problem known as Navigation Among Movable Obstacles (NAMO). Most NAMO planners assume a known environment, while existing approaches for unknown environments typically reason locally about relocations and cannot plan interdependent relocation sequences. We consider NAMO in unknown environments revealed through onboard sensing, where the robot must decide whether a blocked route requires relocation or a feasible path may exist through unexplored space. We propose an online framework that addresses this ambiguity by selecting between navigation and relocation using shortest paths that treat discovered movable objects as obstacles or as removable. Navigation relies on existing motion planners, while relocation uses a sampling-based approach that, unlike existing approaches for unknown environments, searches over \textit{interdependent} relocation sequences and uses an LLM to bias sampling. Numerical experiments demonstrate scalability to cluttered environments requiring interdependent relocations and improved plan quality over existing baselines.
Chinese Translation
传统的机器人规划方法寻求通往目标的无碰撞路径,但当所有路径都被阻挡时便会失效。在这些情况下,机器人必须确定需要搬移哪些物体、以何种顺序搬移以及将其放置在何处以清理出一条路径——这一问题被称为可移动障碍物间导航(Navigation Among Movable Obstacles,NAMO)。大多数NAMO规划器假设环境已知,而现有的面向未知环境的方法通常只能对搬移操作进行局部推理,无法规划相互依赖的搬移序列。我们研究通过机载传感器逐步揭示的未知环境中的NAMO问题,此时机器人必须判断被阻挡的路线是否需要搬移物体,还是通过未探索的空间可能存在可行路径。我们提出一种在线框架,通过选择导航或搬移来解决这种歧义:该方法利用最短路径进行决策,将已发现的可移动物体视为障碍物或可移除对象。导航依赖现有的运动规划器,而搬移则采用基于采样的方法,与现有的面向未知环境的方法不同,它能够搜索相互依赖的搬移序列,并利用大语言模型(LLM)对采样进行引导。数值实验表明,该方法可扩展至需要相互依赖搬移的杂乱环境,且规划质量优于现有基线方法。
cs.RO / 31 / 2609.19554

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

VABench:通过视觉演示、主动感知与度量控制评估具身空间智能
Zhang, Zhongbo, Jin, Jiayi, Wang, Yifan, Zhang, Zaibin, Diao, Haiwen, Wang, Lijun, Lu, Huchuan
Abstract
Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedural context from RGB-only demonstrations, actively select camera viewpoints, issue metric Cartesian commands, and revise them from execution feedback. Models receive no privileged object poses, oracle trajectories, or learned action heads. A fixed model-agnostic controller executes only model-specified targets. VA-Bench contains 14 base task families (11 single-arm and three dual-arm), seven held-out geometry/layout variants, and a long-horizon five-object composition track. We evaluate 12 primary model conditions in three independent runs over the same 20 physically verified seeds per base task, reporting terminal success, nine trajectory-level behavioral diagnostics, and subtask progress. First, the best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run. Its three-run macro-average task success is only 53.93+/-3.17%. Second, active camera control significantly improves task success over passive multi-view observation. In one matched comparison, success rises from 27.86% to 57.50%. Third, held-out geometric transfer can reduce task success by over 30 percentage points. No model completes a strict long-horizon episode, despite substantial partial progress. VA-Bench thus tests whether general-purpose MLLMs can turn visual demonstrations and actively acquired evidence into successful embodied action.
Chinese Translation
空间智能不仅仅需要描述物体的位置。在不完整观测条件下,模型必须识别并获取缺失的证据,在统一的空间坐标系中对其进行解释,并据此采取行动。我们提出了VA-Bench,用于评估完整的“观察-推理-行动-修正”闭环。通用多模态大语言模型(MLLM)从纯RGB演示中学习程序性上下文,主动选择相机视角,发出度量化的笛卡尔坐标指令,并根据执行反馈进行修正。模型不会获得任何特权的物体位姿、理想轨迹或预训练的动作头。一个固定的、与模型无关的控制器仅执行模型指定的目标。VA-Bench包含14个基础任务族(11个单臂任务和3个双臂任务)、七个保留的几何/布局变体,以及一个长时程五物体组合任务轨道。我们在每个基础任务的相同20个物理验证种子上,对12个主要模型条件进行三次独立运行评估,报告终端成功率、九项轨迹级行为诊断指标以及子任务进度。首先,表现最好的模型在标注运行中的目标定位得分达到100.0%,空间关系得分为78.9%,但其三次运行的宏平均任务成功率仅为53.93±3.17%。其次,主动相机控制相比被动多视角观测能显著提升任务成功率:在一项匹配对比实验中,成功率从27.86%提升至57.50%。第三,保留的几何迁移可使任务成功率下降超过30个百分点。尽管模型取得了可观的局部进展,但没有任何模型能够完成严格的长时程任务。因此,VA-Bench检验的是通用MLLM能否将视觉演示和主动获取的证据转化为成功的具身行动。
cs.RO / 32 / 2609.19577

Tele-Traversability: Rethinking Traversability for Teleoperated Ground Robots in Terrain Navigation

遥操作可通行性:重新思考地形导航中遥操作地面机器人的可通行性
Feng, Lewei, Chen, Qi, Wang, Wenshuo, Meng, Xianghao, Kieu, Minh, Guan, Haijie, Wu, Jin, Sun, Fuchun, Xi, Junqiang
Abstract
Teleoperation, a human-in-the-loop control scheme, allows a human operator to remotely command and guide a mobile robot to navigate in off-road environments, yet fluent and user-friendly tele-navigation requires an alignment of traversability evaluation between human and robot. In the teleoperation system, the human operator typically utilizes off-site incomplete and delayed feedback via a human-machine interface to make a judgment of traversability, while the robot makes such an evaluation based on in situ onboard sensory information, which could cause divergent traversability estimation and thus generate mismatched decisions and actions. Existing approaches for traversability modeling, estimation, and prediction are mainly derived from the view of robots, i.e., robot-centric, and are practically suitable for fully autonomous mobile robots, but neglect the influence of human operators. To address this problem, this paper extends the concept of traversability from robot-centric to human-centric by accounting for the operator's cognitive states, such as attention, workload, and risk tolerance or awareness, termed tele-traversability. We first revisit the definitions and roles of traversability in robotics and then extend them to teleoperation settings. Finally, we highlight future trends and open challenges of tele-traversability toward human-centric teleoperation systems.
Chinese Translation
遥操作(Teleoperation)作为一种人在回路的控制方案,允许人类操作员远程指挥和引导移动机器人在越野环境中导航,然而流畅且用户友好的遥导航需要在人与机器人之间实现可通行性(traversability)评估的一致性。在遥操作系统中,人类操作员通常通过人机界面利用场外不完整且延迟的反馈来判断可通行性,而机器人则基于原位的机载传感信息进行此类评估,这可能导致可通行性估计的分歧,从而产生不匹配的决策和行动。现有的可通行性建模、估计和预测方法主要从机器人的视角出发(即以机器人为中心),实际上适用于完全自主的移动机器人,但忽略了人类操作员的影响。为解决这一问题,本文通过考虑操作员的认知状态(如注意力、工作负荷以及风险容忍度或风险意识),将可通行性概念从以机器人为中心扩展到以人为中心,称为遥操作可通行性(tele-traversability)。我们首先回顾了可通行性在机器人学中的定义和作用,然后将其扩展到遥操作场景。最后,我们展望了面向以人为中心的遥操作系统中遥操作可通行性的未来趋势和开放性挑战。
cs.RO / 33 / 2609.19579

Recovering Aggressively Pruned Vision-Language-Action Models with Offline Hidden-State Distillation

利用离线隐状态蒸馏恢复激进剪枝的视觉-语言-动作模型
Kim, Chiyoung, Choi, Sanghyuk Roy, Lee, Minhyeok
Abstract
Vision-language-action (VLA) models let robots follow language instructions, but their language backbones of several billion parameters are the main obstacle to running them on robot hardware. Structured pruning reduces that backbone, and removing 63% of it from OpenVLA-OFT drops LIBERO-Long success from 93.2% to 0.8%. A recent approach restores such a model with supervised fine-tuning followed by reinforcement learning, which needs online rollouts and hundreds of GPU-hours. We recover most of the lost success entirely offline. Width pruning narrows the blocks but keeps the residual stream at its original size, so teacher and student hidden states have the same shape and are matched directly, without a projector. Training against a cache built in one teacher pass lifts the 63%-reduced student to within 3.5 points of the teacher in about 8 GPU-hours. A sweep over nine ratios locates where the recovery objective starts to matter. Up to 45% reduction the two do not differ significantly on OpenVLA-OFT. Hidden-state distillation then adds +2.1 to +4.5 points there between 63% and 87%, and +9.4 to +22.1 points on CogACT from 63% onward. At 81% on CogACT, a tripled recovery budget narrows the distilled student's gap to the teacher to 3.9 points on average, while supervised recovery stays more than 20 points below. At matched compression, width pruning yields higher success and depth pruning lower latency. On a 6-DoF manipulator, the distilled student at 72% reduction reaches 77.5% success against 59.5% for supervised recovery, runs 2.23x faster on-board than the teacher, and uses 62% less memory.
Chinese Translation
视觉-语言-动作(Vision-language-action, VLA)模型使机器人能够遵循语言指令,但其数十亿参数的语言骨干网络是将其部署到机器人硬件上的主要障碍。结构化剪枝可以缩减该骨干网络,例如从 OpenVLA-OFT 中移除 63% 的参数会使 LIBERO-Long 成功率从 93.2% 骤降至 0.8%。近期的一种方法通过监督微调结合强化学习来恢复此类模型,但这需要在线交互采样(rollout)并消耗数百个 GPU 小时。我们则完全离线地恢复了大部分损失的性能。宽度剪枝缩小了各个模块的规模,但保持残差流(residual stream)的原始尺寸,因此教师模型与学生模型的隐状态形状相同,可以直接匹配,无需投影器。基于一次教师前向传播构建的缓存进行训练,可将 63% 剪枝后的学生模型在大约 8 个 GPU 小时内提升至距教师模型 3.5 个百分点以内。通过对九个剪枝比例的扫描,我们确定了恢复目标开始发挥作用的阈值。在剪枝不超过 45% 时,二者在 OpenVLA-OFT 上差异不显著;在 63% 至 87% 区间,隐状态蒸馏可带来 +2.1 至 +4.5 个百分点的提升;在 CogACT 上,从 63% 起可带来 +9.4 至 +22.1 个百分点的提升。在 CogACT 的 81% 剪枝率下,将恢复预算增至三倍可使蒸馏学生模型与教师模型的平均差距缩小至 3.9 个百分点,而监督恢复仍低 20 多个百分点。在相同压缩率下,宽度剪枝带来更高的成功率,而深度剪枝具有更低的延迟。在 6 自由度机械臂上,剪枝 72% 的蒸馏学生模型达到 77.5% 的成功率(监督恢复为 59.5%),在机载运行速度比教师模型快 2.23 倍,内存占用减少 62%。
cs.RO / 34 / 2609.19582

OmniCalib: Target-Free, Task-Structured Self-Calibration for Humanoid Robots

OmniCalib:面向人形机器人的无目标、任务结构化自标定方法
Lu, Kaixiang, Lan, Haiyu, Qiao, Chunxiao, Li, You, Li, Enyu, Lu, Yehao, Yang, Jiarui, Lin, Peiwen, Wang, Chuang
Abstract
Assembly, wear, and component replacement perturb the sensor extrinsics and joint zeros encoded by a humanoid CAD model. Existing procedures calibrate one sensor pair or require external fiducials. Using only robot-native motion and onboard sensing, we present OmniCalib, a target-free workflow that calibrates the full upper limbs---all 14 arm joint zeros and the extrinsics of both wrist and chest cameras---as well as lower limbs and the multi-camera head rig. Each module matches a robot-native task to a parameter block, checks observability, and writes only supported corrections to the CAD model. Our depth ICP method recovers all 14 arm joint zeros and calibrates all RGB-D camera extrinsics without any calibration target. Relative to CAD, the estimated extrinsic corrections are 10.56 mm and 1.74 degrees for the left wrist, 6.33 mm and 1.25 degrees for the right wrist, and 9.81 mm and 0.929 degrees for the chest RGB-D camera. ICP point-to-plane residual is 2.09 mm. On the same injected offsets, ICP and ArUco recover all 14 joint zeros below the 0.1-degree encoder-resolution reference. On an AGIBOT A3 Ultra humanoid, four static double-support stances recover all 12 lower-limb joint-zero offsets injected with an RMS error of 0.063 degrees. The head module combines multi-camera visual odometry with legged odometry and dynamic compensation through the live ROS transform tree. Using only planar walking, it attains a mean SO(3) error of 1.061 degrees across three sequences. The best sequence reaches 0.775 degrees, competitive with iKalibr at 0.902 degrees from rich 6-DOF excitation. Rig-relative angles repeat within 0.140 degrees. Injection recovery and held-out tests validate each observable block.
Chinese Translation
装配、磨损和部件更换会扰动人形机器人CAD模型中编码的传感器外参和关节零位。现有标定流程只能标定单个传感器对,或需要外部基准标记。我们提出了OmniCalib,一种仅利用机器人自身运动和机载传感的无目标标定流程,可标定完整上肢——包括全部14个手臂关节零位、双腕部相机与胸部相机的外参——以及下肢和多相机头部装置。每个模块将一个机器人原生任务匹配到一个参数块,检验可观测性,并仅将受支持的修正写入CAD模型。我们的深度ICP方法无需任何标定目标即可恢复全部14个手臂关节零位并标定所有RGB-D相机外参。相对于CAD,估计的外参修正为:左腕10.56毫米和1.74度,右腕6.33毫米和1.25度,胸部RGB-D相机9.81毫米和0.929度。ICP点到面残差为2.09毫米。在相同注入偏移下,ICP与ArUco方法恢复的全部14个关节零位误差均低于0.1度的编码器分辨率参考值。在AGIBOT A3 Ultra人形机器人上,通过四个静态双脚支撑姿态,以0.063度的均方根误差恢复了全部12个注入的下肢关节零位偏移。头部模块将多相机视觉里程计与腿式里程计结合,并通过实时ROS变换树进行动态补偿。仅使用平面行走,该方法在三个序列上的平均SO(3)误差为1.061度,最佳序列达到0.775度,与基于丰富六自由度激励的iKalibr(0.902度)具有竞争力。装置相对角度的重复性在0.140度以内。注入恢复实验和保留测试验证了每个可观测参数块的有效性。
cs.RO / 35 / 2609.19588

Quantifying Mechanical Intelligence in Legged Robots with Information Theory

基于信息论的腿式机器人机械智能量化研究
Patterson, Zach J.
Abstract
Mechanical intelligence, loosely defined as the reduction in control burden afforded by a robot's physical form, has become a prominent concept in robotics, with instantiations in bioinspired robotics, soft robotics, robotic swarms, and many other areas. However, rigorous theoretical understanding and quantitative measures of mechanical intelligence have lagged behind the engineering systems that the community has developed. In this work, using modern legged robots as a benchmark and exemplar, we propose several information-theoretic metrics for quantifying mechanical intelligence. By viewing body dynamics as both a computational process and a communication channel, we show that several prior insights in legged-robot engineering can be described using information theory, and we quantify how bits are processed by mechanical modes and across robot coordinates. Specifically, we examine the trade-off between explicitly incorporating compliance through series-elastic actuation and using so-called proprioceptive, low-gear-ratio transmissions, and we explore how these mechanisms interact with control policies during locomotion. We develop these results on systems of increasing complexity: a simplified linear model of a robot-leg transmission, a nonlinear single-leg simulation, and simulated quadruped robots controlled by a learned policy while navigating challenging terrain. These results lay the groundwork for broader study of robot mechanisms and their role in embodied computation.
Chinese Translation
机械智能(Mechanical Intelligence)泛指由机器人物理形态本身所减轻的控制负担,这一概念在机器人学中日益突出,广泛应用于仿生机器人、软体机器人、机器人群集等诸多领域。然而,对机械智能的严谨理论理解与定量度量手段,仍滞后于学界所研发的各类工程系统。在本工作中,我们以现代腿式机器人为基准和范例,提出了几种用于量化机械智能的信息论指标。通过将身体动力学同时视为一个计算过程和一个通信信道,我们证明腿式机器人工程中的若干已有洞见可以用信息论加以描述,并量化了信息比特在机械模态内以及机器人各坐标之间被处理的程度。具体而言,我们研究了通过串联弹性驱动显式引入柔性与采用所谓本体感知式(proprioceptive)低减速比传动之间的权衡,并探讨了这些机制在运动过程中如何与控制策略相互作用。我们在复杂度递增的系统上发展了这些结果:简化的机器人腿部传动线性模型、非线性单腿仿真,以及由学习策略控制、在复杂地形中导航的仿真四足机器人。这些结果为更广泛地研究机器人机构及其在具身计算中的作用奠定了基础。
cs.RO / 36 / 2609.19600

WorldContact: A Contact-Centric World Model for Scalable Robot Learning

WorldContact:一种面向可扩展机器人学习的以接触为中心的世界模型
Wang, Caoliwen, Wang, Mengdi, Zhang, Heng, Huang, Shixun, Chen, Siyuan, Liu, Chao, Chen, Anpei, Wang, Zhendong, Chen, Peter Yichen, Wang, Huamin
Abstract
Adapting robots to new objects and tasks requires interaction experience that can be costly to obtain. We present WorldContact, a contact-centric world model for deformable-object manipulation, constructed from a limited set of high-quality trajectories to generate additional training data efficiently. It predicts object dynamics using larger time steps than the source numerical simulator, which requires small integration steps to resolve rapid motion and prevent interpenetration. We evaluate WorldContact across 16 shopping-bag manipulation tasks. State-rollout measurements on a single H100 GPU show a $10\times$ speedup over the source simulator, excluding rendering and disk I/O. We use the generated data to fine-tune an existing vision-language-action policy and deploy it directly on a real robot. In bag lifting, the same policy achieves 65% single-attempt success when fine-tuned on source simulation data alone, compared with 95% when fine-tuned on the dataset expanded with WorldContact. These results support efficient data generation with WorldContact for robot policy adaptation.
Chinese Translation
使机器人适应新物体和新任务需要获取代价高昂的交互经验。我们提出了WorldContact,一种面向可变形物体操作的以接触为中心的世界模型,它由少量高质量轨迹构建,能够高效地生成额外的训练数据。该模型使用比源数值仿真器更大的时间步长来预测物体动力学,而源仿真器需要较小的积分步长来处理快速运动并防止物体相互穿透。我们在16个购物袋操作任务上对WorldContact进行了评估。在单块H100 GPU上的状态轨迹推演测量显示,其速度比源仿真器提升10倍(不包括渲染和磁盘I/O)。我们利用生成的数据对现有的视觉-语言-动作(VLA)策略进行微调,并直接部署到真实机器人上。在提袋任务中,仅使用源仿真数据微调的同一策略单次尝试成功率为65%,而使用经WorldContact扩充的数据集微调后成功率达95%。这些结果验证了WorldContact在机器人策略适配中高效数据生成的有效性。
cs.RO / 37 / 2609.19613

TacSushi: Tactile-Grounded World-Action Modeling for Dexterous Sushi Manipulation

TacSushi:面向灵巧寿司操作的触觉接地世界-动作建模
Hu, Haodi, Kogashi, Kaen, Koike-Akino, Toshiaki
Abstract
Dexterous food manipulation requires control under deformation, occlusion, and uncertain contact. We present TacSushi, a tactile-grounded, Cosmos3-based world-action policy that learns from recorded future consequences while acting on current observations. The backbone encodes current RGB, language, and hand state, and feature-wise gated fusion incorporates fingertip tactile features into the action representation. During training, a decoder conditioned on demonstrated action chunks predicts logged future visual observations, task progress, relative contact risk, and tactile summaries; this decoder is removed at deployment. Failed trials provide consequence supervision, but their actions are excluded from imitation. We train TacSushi on 340 successful and 50 failed real-robot trials and compare six methods in 600 separate rollouts across three in-distribution tasks and two out-of-distribution ingredient variants. To assess food quality beyond a single geometric threshold, we score terminal outcomes using an anchored visual-quality protocol that equally weights five human ratings and three vision-language-model ratings per rollout. Full TacSushi achieves 68.3% average in-distribution success and 37.5% out-of-distribution success, compared with 36.7%/10.0% without future-consequence supervision and 25.0%/17.5% with direct tactile concatenation in place of gated fusion. These comparisons support complementary benefits of feature-wise gated tactile fusion and training-only predictive supervision.
Chinese Translation
灵巧的食物操作需要在形变、遮挡和不确定接触条件下的控制。我们提出了TacSushi,一种基于触觉接地、以Cosmos3为基础的世界-动作策略,它在基于当前观测进行动作的同时,从记录的未来后果中学习。其主干网络编码当前的RGB图像、语言和手部状态,并通过特征级门控融合将指尖触觉特征融入动作表示。在训练过程中,一个以演示动作块为条件的解码器预测记录的未来视觉观测、任务进度、相对接触风险和触觉摘要;该解码器在部署时被移除。失败的试验提供后果监督,但其动作被排除在模仿学习之外。我们在340次成功和50次失败的真实机器人试验上训练TacSushi,并在三个分布内任务和两个分布外食材变体上,通过600次独立测试 rollout 比较了六种方法。为了超越单一几何阈值评估食物质量,我们采用锚定的视觉质量评分协议对最终结果进行评分,该协议对每次 rollout 的五项人类评分和三项视觉-语言模型评分赋予相同权重。完整的TacSushi在分布内任务中平均成功率为68.3%,分布外成功率为37.5%,相比之下,没有未来后果监督的版本为36.7%/10.0%,而用直接触觉拼接替代门控融合的版本为25.0%/17.5%。这些对比结果支持了特征级门控触觉融合与仅训练阶段的预测监督所带来的互补优势。
cs.RO / 38 / 2609.19659

EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence

EmbodiedMind:面向高效具身智能的自适应数据筛选与前缀树强化学习
Wang, Feifan, Zhang, Zongbing, Zhang, Yu, Wang, Lingfeng, Zhu, Yurui, Deng, Jin, Zhang, Mingliang, Gao, Zhengguang, Wang, Yongcheng, Xu, Jin, Yang, Ri
Abstract
Training embodied foundation models typically requires massive-scale datasets and extensive computational resources, yet often suffers from three critical limitations: (1) inefficient sample utilization due to low-informative samples; (2) imbalanced gradient contributions across heterogeneous tasks; and (3) severe credit assignment problem in long-horizon planning, where trajectory-level rewards indiscriminately penalize all tokens. To address these issues, we propose an efficient training paradigm that achieves state-of-the-art average performance through strategic data selection and hierarchical policy optimization. Our approach consists of three synergistic stages. First, Rejection Sampling-based Fine-Tuning (RSFT) filters out low-informative samples to establish robust behavioral priors while preventing distributional collapse. Second, Iterative Rejection GRPO (IR-GRPO) employs task-specific queues stratified by difficulty to keep datasets balanced across reinforcement learning iterations, coupled with a hybrid reward mechanism for precise cross-task feedback. Third, to enhance long-horizon task planning, we introduce Trie-GRPO, a novel reinforcement learning algorithm based on action prefix trees, which enables step-level advantage estimation. This resolves the credit assignment problem by isolating intermediate correct decisions from downstream errors, while effectively balancing exploration efficiency and depth compared to conventional search trees. As a result, EmbodiedMind achieves a state-of-the-art average performance of 70.02% across 18 benchmarks, and significantly outperforms other embodied foundation models in long-horizon task planning accuracy. Our project will be released for reproducibility.
Chinese Translation
训练具身基础模型通常需要大规模数据集和大量计算资源,然而往往面临三个关键局限:(1)由于低信息量样本的存在导致样本利用效率低下;(2)异构任务间梯度贡献不均衡;(3)长时程规划中存在严重的信用分配问题,即轨迹级别的奖励会不加区分地惩罚所有token。为解决这些问题,我们提出了一种高效的训练范式,通过策略性的数据选择和分层策略优化实现了最先进的平均性能。我们的方法包含三个协同阶段。首先,基于拒绝采样的微调(Rejection Sampling-based Fine-Tuning, RSFT)过滤掉低信息量样本,在防止分布坍缩的同时建立稳健的行为先验。其次,迭代拒绝GRPO(Iterative Rejection GRPO, IR-GRPO)采用按难度分层的任务专用队列,以在强化学习迭代过程中保持数据集均衡,并结合混合奖励机制实现精确的跨任务反馈。第三,为增强长时程任务规划能力,我们提出了Trie-GRPO,一种基于动作前缀树的新型强化学习算法,可实现步骤级的优势估计。该方法通过将中间的正确决策与下游错误隔离,解决了信用分配问题,同时相较于传统搜索树,能够有效平衡探索效率与深度。最终,EmbodiedMind在18个基准测试中取得了70.02%的最先进平均性能,并在长时程任务规划准确率上显著优于其他具身基础模型。为保证可复现性,我们的项目将被公开发布。
cs.RO / 39 / 2609.19661

ReShoot: Generative Visual Domain Randomization of Recorded Robot Demonstrations for Visuomotor Policy Learning

ReShoot:基于已录制机器人演示的生成式视觉域随机化用于视觉运动策略学习
Kim, Chiyoung, Choi, Min Sung, Ju, Jinho, Gu, Chanhoe, Hwang, Donghwan, Choi, Wonseok, Jeon, Woongsun, Lee, Minhyeok
Abstract
Imitation-learned robot policies are frequently overfit to the visual conditions present in their training demonstrations. Consequently, variations in object color or background appearance often induce substantial performance degradation. A common mitigation strategy is to acquire additional demonstrations in each novel visual context; however, this approach is resource-intensive, requiring repeated access to a robot, a controlled environment, and human operation for every appearance condition to be covered. We introduce ReShoot, a framework that synthesizes visual diversity by re-rendering previously recorded demonstrations under altered appearances, thereby shifting the burden from data collection to generation. A vision-language model captions the scene, edits a targeted attribute (e.g., background, object color, or material), and an edge-conditioned video generator re-renders both camera views to match. The instruction is updated accordingly. The action sequence and proprioceptive trajectory are copied verbatim without relabeling, so each generated episode retains the recorded action and proprioceptive labels. On LIBERO, a policy trained on an equal mixture of recorded and re-rendered demonstrations matches the performance of recorded-only training (96.5% vs. 96.9%). Moreover, the mixed training set improves robustness to scene perturbations on LIBERO-Plus (85.5% vs. 82.3%). Across two physical robotic platforms, deploying ReShoot with 43 and 100 pre-collected demonstrations increased the success rate on recolored objects from 0.0% to 42.9% and 47.5%, respectively, while maintaining performance under the original recorded appearance.
Chinese Translation
通过模仿学习获得的机器人策略往往会对训练演示中出现的视觉条件产生过拟合。因此,物体颜色或背景外观的变化常常导致性能显著下降。一种常见的缓解策略是在每个新的视觉环境下采集额外的演示数据;然而,这种方法资源消耗巨大,需要针对每一种待覆盖的外观条件反复使用机器人、受控环境和人工操作。我们提出了ReShoot,一个通过在改变外观后重新渲染先前录制的演示来合成视觉多样性的框架,从而将负担从数据采集转移到数据生成。具体而言,视觉-语言模型(VLM)对场景进行描述并编辑目标属性(如背景、物体颜色或材质),然后由边缘条件控制的视频生成器重新渲染两个相机视角以匹配编辑后的场景,并相应更新指令文本。动作序列和本体感觉轨迹则被原样复制而无需重新标注,因此每条生成的轨迹都保留了录制时的动作和本体感觉标签。在LIBERO基准上,使用录制演示与重渲染演示等比例混合训练的策略达到了与仅用录制数据训练相当的性能(96.5% 对比 96.9%)。此外,混合训练集提升了对场景扰动的鲁棒性(LIBERO-Plus上85.5% 对比 82.3%)。在两个物理机器人平台上,利用43条和100条预先采集的演示部署ReShoot后,策略在变色物体上的成功率分别从0.0%提升至42.9%和47.5%,同时在原始录制外观下保持了原有性能。
cs.RO / 40 / 2609.19665

Runtime Safety Filtering for Two-Terminal Hazards in Robotic Battery Recycling

面向机器人电池回收中双端子危险的运行时安全过滤
Cao, Yuxin, Song, Wei, Yang, Xianglin, Guo, Fusen, Li, Lin, Cheng, Xiao, Dong, Jin Song
Abstract
Runtime safety filters for learned manipulation policies typically define unsafe states as unions of object-wise keep-out regions. This representation can be unnecessarily restrictive for hazards that depend on a joint spatial relation, such as battery recycling, where a conductive payload can short a charged cell only when it approaches both terminals simultaneously. We study runtime filtering for this two-terminal hazard in LIBERO using frozen OpenVLA policies. We factor a runtime filter into three design choices: the predicate structure, its geometric margin, and the fallback action applied when a commanded action is rejected. We compare a conjunctive predicate, a conventional two-site keep-out, and a composite of the two. For each predicate, we vary its margin to obtain a frontier between task success and residual hazard. We then compare four fallback strategies at matched operating points: holding, retreat, sampled search, and a continuous-action barrier projection. Across three workcells, the three predicate families trace nearly identical safety--utility frontiers once each is evaluated over its own margin. In contrast, the fallback strategy has a substantially larger effect: holding reduces task success by up to 0.302 relative to retreat without reducing hazard, while both minimally invasive fallbacks leave substantially more residual hazard. This ordering transfers to a second policy and task suite, while retreat-based filtering remains effective under standing errors in the clearances available to the filter, although correlated error in the estimated payload size is more damaging than larger independent errors in terminal position. These results show that, for proximity-defined manipulation hazards, margin selection and fallback strategy can matter more than predicate structure in determining the safety--utility trade-off of a runtime filter.
Chinese Translation
面向学习型操作策略的运行时安全过滤器通常将不安全状态定义为各物体禁入区域的并集。对于依赖于联合空间关系的危险而言,这种表示可能过于保守,例如电池回收场景:导电载荷只有在同时接近两个端子时才会使带电电芯短路。我们在 LIBERO 基准上,针对冻结的 OpenVLA 策略,研究了此类双端子危险的运行时过滤问题。我们将运行时过滤器分解为三个设计选择:谓词结构、其几何安全裕度,以及在拒绝指令动作时所采取的回退动作。我们比较了合取谓词、传统的双点位禁入区域,以及二者组合。对于每种谓词,我们通过调整其裕度,得到任务成功率与残余危险之间的权衡前沿。随后,在匹配的工作点上比较四种回退策略:保持、后退、采样搜索以及连续动作屏障投影。在三个工作单元上的实验表明,若在各自裕度范围内评估,三类谓词族描绘出的安全—效用前沿几乎完全相同。相比之下,回退策略的影响显著更大:与后退策略相比,保持策略使任务成功率最多降低 0.302,却并未减少危险;而两种最小侵入式回退策略则会留下明显更多的残余危险。该排序在第二个策略与任务套件上同样成立;基于后退的过滤在过滤器可用间隙存在系统性误差时依然有效,尽管载荷尺寸估计中的相关误差比端子位置上更大的独立误差更具破坏性。这些结果表明,对于由邻近关系定义的操作危险,在决定运行时过滤器的安全—效用权衡时,裕度选择和回退策略可能比谓词结构更为重要。
cs.RO / 41 / 2609.19666

Towards High-DoF Dexterous Manipulation through VLA Post-Training

通过VLA后训练实现高自由度灵巧操作
Zhu, Junlei, Yao, Shenzhe, Huang, Chaogui, Zhu, Wenkai, Peng, Jingwei, He, Guanqi, Schwertfeger, Soren, Chen, Jiahao, Liu, Yide
Abstract
Imitation-learned vision--language--action (VLA) foundation models acquire broad manipulation capabilities by scaling robot data across tasks and embodiments, but reliable deployment on a specific downstream task and hardware platform still requires post-training. Dexterous hands make this adaptation particularly difficult: their broad behavioural repertoire and high degree of freedom create a large and structured action space. Three obstacles are central: open-source VLAs do not natively provide an action interface for high-DoF hands; gesture mismatch during human-gated DAgger takeover creates command discontinuities and contaminates corrective trajectories; and reinforcement learning in the raw joint space is sample-inefficient. We present a unified four-step post-training pipeline comprising a learned temporal hand-action codec, supervised fine-tuning, DAgger, and real-world residual reinforcement learning. The codec adapts a pretrained VLA to absolute dexterous-hand commands. Buffered rollback, pose alignment, and smooth command blending enable continuous, task-relevant DAgger corrections, while latent residual RL confines exploration to coordinated hand motions captured by the codec. We evaluate the pipeline on five diverse real-world tasks spanning bimanual transfer, in-hand reorientation, and tool use. Within the reported post-training budgets, the resulting policies achieve 100\% success on every evaluated task over 20 trials per task. These results provide a practical path for adapting VLA foundation models to reliable real-world dexterous manipulation.
Chinese Translation
通过模仿学习训练的视觉-语言-动作(VLA)基础模型,通过跨任务和跨机器人本体的数据扩展获得了广泛的操作能力,但要在特定的下游任务和硬件平台上实现可靠部署,仍需要进行后训练。灵巧手使这一适应过程尤为困难:其广泛的行为 repertoire 和高自由度构成了一个庞大且结构化的动作空间。本文聚焦于三个核心障碍:开源VLA模型本身不提供面向高自由度灵巧手的动作接口;在人工介入的DAgger接管过程中,手势不匹配会造成指令不连续并污染纠正轨迹;而在原始关节空间中进行强化学习的样本效率低下。我们提出了一个统一的四步后训练流程,包括学习型时序手部动作编解码器、监督微调、DAgger以及真实世界残差强化学习。该编解码器将预训练的VLA适配为绝对灵巧手指令。通过缓冲回滚、姿态对齐和平滑指令混合,实现了连续且与任务相关的DAgger纠正;同时,潜在残差强化学习将探索范围限制在编解码器所捕捉的协调性手部动作上。我们在五项多样化的真实世界任务上对该流程进行了评估,涵盖双手迁移、手内重定向和工具使用。在所报告的后训练预算内,所得策略在每项任务的20次试验中均达到100%的成功率。这些结果为将VLA基础模型适配至可靠的真实世界灵巧操作提供了一条切实可行的路径。
cs.RO / 42 / 2609.19681

VAST: V2X/Dynamic Map-Aware Autonomous Driving Systems Validation Toolchain

VAST:面向V2X/动态地图感知的自动驾驶系统验证工具链
Ito, Shunsuke, Azumi, Takuya
Abstract
Cooperative autonomous driving in the IoT-to-Edge-to-Cloud continuum requires system-level validation across vehicles, infrastructure sensors, edge-side Dynamic Map services, and in-vehicle autonomous-driving stacks. This paper presents VAST, a V2X/Dynamic Map-aware validation toolchain that connects Scenic, Scenario Simulator v2, AWSIM, Autoware, and SIM-LDM. VAST does not introduce a new search algorithm; instead, it addresses interoperability challenges, including Lanelet2-to-Scenic mapping, ROS 2-based co-simulation through SS2, Dynamic Map object injection into Autoware, and collection of TTC, PET, collision, timeout, and performance measurements. In occluded-intersection scenarios, Lanelet2-compatible constrained sampling increases the edge-case discovery rate from 40.0% to 80.0% and reduces the average time per discovered edge case from 259.7 s to 110.4 s. Under the same generated scenario distribution, Dynamic Map availability reduces the collision rate from 78.0% to 40.0% and increases non-collision outcomes from 22.0% to 60.0%, with statistically significant TTC/PET shifts. A throughput study with 1-16 NPCs shows that sampling remains below 0.1 s, whereas AWSIM/Autoware execution and restart overhead dominate runtime. These results position VAST as a practical validation infrastructure for cooperative autonomous-driving CPSs.
Chinese Translation
从物联网(IoT)到边缘再到云端的连续体中的协同自动驾驶,需要在车辆、基础设施传感器、边缘侧动态地图服务以及车载自动驾驶软件栈之间进行系统级验证。本文提出了VAST,一种V2X/动态地图感知的验证工具链,它将Scenic、Scenario Simulator v2、AWSIM、Autoware和SIM-LDM连接起来。VAST并未引入新的搜索算法,而是解决互操作性挑战,包括Lanelet2到Scenic的映射、通过SS2实现基于ROS 2的联合仿真、向Autoware注入动态地图对象,以及收集TTC、PET、碰撞、超时和性能测量数据。在遮挡交叉路口场景中,与Lanelet2兼容的约束采样将边缘用例发现率从40.0%提升至80.0%,并将每个被发现边缘用例的平均时间从259.7秒降低至110.4秒。在相同的生成场景分布下,动态地图的可用性将碰撞率从78.0%降低至40.0%,并将无碰撞结果从22.0%提升至60.0%,且TTC/PET的变化具有统计显著性。一项针对1-16个NPC的吞吐量研究表明,采样耗时保持在0.1秒以下,而AWSIM/Autoware的执行与重启开销主导了整体运行时间。这些结果表明,VAST可作为协同自动驾驶信息物理系统(CPS)的实用验证基础设施。
cs.RO / 43 / 2609.19688

LYRIC: Language-Driven Physics-Based Character Control for Contact-Rich Whole-Body Object Interaction

LYRIC:基于语言驱动的物理仿真角色控制,实现富含接触的全身物体交互
Han, Zeyu, Meng, Zichong, Tanke, Julian, Matsumoto, Minami, Bashkirov, Sergey, Fan, Yingruo, Engin, Selim, Shim, Dongseok, Shibuya, Takashi, Mitsufuji, Yuki, Jiang, Huaizu
Abstract
We present LYRIC, a generative flow-matching controller for language-driven physics-based contact-rich interaction control, that enables simulated characters to perform contact-rich whole-body object interactions from a free-form language instruction and a sparse terminal object goal. To obtain reliable expert trajectories from imperfect motion-capture references, a single tracking policy is trained using geometry-conditioned interaction rewards and relaxed reference tracking near hand-object contact. To guide interaction progress without prescribing a full-body kinematic reference, we factorize the controller into a task-level planner that predicts short-horizon object and humanoid-root trajectories, and an action generator that resolves whole-body motion and contacts in closed loop. After behavior cloning, we freeze the planner and post-tune the action generator on policy using the planner's predictions as stable supervision for intermediate task progression. In a controlled OMOMO evaluation, our tracker achieves 64.3% success compared with 53.2% for an InterMimic reimplementation, while a unified policy achieves 76.5% on the full OMOMO dataset. On the held-out split, LYRIC achieves 90.3% task success, compared with 74.2% for the strongest matched kinematic-planner baseline, with better semantic alignment and motion quality. Without retraining, the controller also supports test-time object-waypoint guidance. Qualitative results further demonstrate robust, natural contact-rich interactions and zero-shot transfer to novel object shapes. The webpage is available at https://neu-vi.github.io/LYRIC/
Chinese Translation
我们提出了 LYRIC,一种基于语言驱动、面向物理仿真中富含接触交互控制的生成式流匹配(flow-matching)控制器,使仿真角色能够根据自由形式的语言指令和稀疏的终端物体目标执行富含接触的全身物体交互。为了从并不完美的动作捕捉参考中获取可靠的专家轨迹,我们使用基于几何条件的交互奖励以及在手-物体接触附近放松的参考跟踪来训练单一跟踪策略。为了在不规定全身运动学参考的情况下引导交互进程,我们将控制器分解为一个任务级规划器(用于预测短时程的物体与人形根部轨迹)和一个动作生成器(以闭环方式求解全身运动与接触)。在行为克隆之后,我们冻结规划器,并以规划器的预测作为中间任务进展的稳定监督,对动作生成器进行在策略(on-policy)的后调优。在受控的 OMOMO 评估中,我们的跟踪器达到 64.3% 的成功率,而 InterMimic 的复现版本为 53.2%;统一策略在完整 OMOMO 数据集上达到 76.5%。在留出测试集上,LYRIC 达到 90.3% 的任务成功率,而最强的匹配运动学规划器基线为 74.2%,且具有更好的语义对齐与运动质量。无需重新训练,该控制器还支持测试时的物体路径点引导。定性结果进一步展示了鲁棒、自然的富含接触交互以及对新物体形状的零样本迁移。项目网页见 https://neu-vi.github.io/LYRIC/
cs.RO / 44 / 2609.19690

UniExo: Unified Multi-Skill Policies for Musculoskeletal Locomotion and Co-Adaptive Exoskeleton Control

UniExo:面向肌肉骨骼运动与协同自适应外骨骼控制的统一多技能策略
Yuan, Yifei, Wolf, Jakob, Androwis, Ghaith, Zhou, Xianlian
Abstract
Daily locomotion encompasses diverse activities and frequent transitions between them, yet most exoskeleton controllers are designed for a single activity or a narrow set of related movements. Changes in activity therefore typically require explicit mode switching and separately tuned or retrained controllers. Simulation-based learning reduces the need for hardware-based tuning but generally retains this limitation. Here we present UniExo, a framework that first constructs a multi-skill musculoskeletal human policy and then jointly trains an exoskeleton control policy with it. Four single-skill imitation experts for walking, turning, running and backward walking are distilled into a single network structured by a skill latent and subsequently fine-tuned through reinforcement learning on transition sequences. The resultant unified human policy achieves a mean tracking success rate of 94.7% on unseen clips of the four skills and exhibits greater robustness to perturbations than its constituent experts. A single hip exoskeleton controller (UniExo) is initialized from hip moment prediction of the human policy and co-adapted with it through multi-agent reinforcement learning across the four skills. This co-adaptation shifts the timing of the assistance torque and raises the fraction of positive work delivered to the hip. When deployed on a custom hip exoskeleton, the controller generalizes across four treadmill speeds in six participants and assists one participant through a continuous route of all four skills and their transitions, without skill labels or explicit mode switching. UniExo thus provides a step towards replacing activity-specific controllers with unified, user-specific controllers that support diverse locomotor activities and the transitions between them.
Chinese Translation
日常运动包含多种多样的活动以及活动之间的频繁切换,然而大多数外骨骼控制器仅针对单一活动或一组狭窄的相关动作设计。因此,活动变化通常需要显式的模式切换以及单独调参或重新训练的控制器。基于仿真的学习方法减少了对硬件调参的依赖,但通常仍保留了这一局限性。本文提出UniExo,该框架首先构建一个多技能的肌肉骨骼人体策略,然后与其联合训练外骨骼控制策略。针对行走、转身、跑步和倒走四种单一技能的模仿专家被蒸馏到一个由技能隐变量组织的单一网络中,随后通过强化学习在过渡序列上进行微调。所得的统一人体策略在四种技能的未见动作片段上实现了94.7%的平均跟踪成功率,并且对扰动的鲁棒性优于其各组成专家。单一髋关节外骨骼控制器(UniExo)以人体策略的髋关节力矩预测进行初始化,并通过多智能体强化学习在四种技能上与之协同自适应。这种协同自适应调整了助力力矩的时机,并提高了传递到髋关节的正功比例。在定制髋关节外骨骼上部署时,该控制器在六名受试者中可泛化于四种跑步机速度,并辅助一名受试者完成包含全部四种技能及其过渡的连续路线,无需技能标签或显式模式切换。因此,UniExo为用支持多样运动活动及其过渡的统一化、用户定制化控制器取代针对特定活动的控制器迈出了一步。
cs.RO / 45 / 2609.19726

Decoupling Physical Speed from Path Parameterization in Singularity-Free Guiding Vector Fields

无奇异导引向量场中物理速度与路径参数化的解耦
Xiao, Zhouru, Luo, Sha, Lu, Yang, Xiao, Mingliang, Yao, Weijia, Lin, Bohuan, Cheng, Xianzhe, Wang, Yaonan
Abstract
The existing singularity-free guiding vector field (SF-GVF) with an additional virtual coordinate can eliminate singular points (i.e., points where the vector field vanishes) inherent in conventional GVFs and guarantee global convergence of robot trajectories to closed and self-intersecting desired paths. However, the desired speed given by the GVF along the desired path in the original lower-dimensional space cannot be arbitrarily specified but depends on path parameterizations. One possible workaround is to partially normalize the physical projection of the SF-GVF and assign a user-designed speed. However, we show that this workaround may introduce new singularities since the normalization denominator can become zero. To address this issue, we propose a new SF-GVF with prescribed physical speed (PPS). The integral curves of the new SF-GVF converge exponentially to the desired path from any initial condition in the higher-dimensional space (including virtual dimension); more importantly, the robot's physical speed converges to the PPS, while the path-error dynamics remain invariant under regular reparameterizations of the desired path. We further develop a saturated acceleration control law for second-order kinematic models. Finally, comparative simulations and 3D path-following experiments with a quadrotor under different PPS profiles validate the theoretical results and demonstrate the effectiveness of the proposed approach.
Chinese Translation
现有的带有附加虚拟坐标的无奇异导引向量场(Singularity-Free Guiding Vector Field, SF-GVF)能够消除传统导引向量场(GVF)中固有的奇异点(即向量场为零的点),并保证机器人轨迹全局收敛到闭合且自相交的期望路径。然而,由GVF在原始低维空间中沿期望路径给出的期望速度无法任意指定,而是依赖于路径参数化。一种可行的变通方法是对SF-GVF的物理投影进行部分归一化,并赋予用户设计的速度。然而,我们证明这种变通方法可能引入新的奇异点,因为归一化分母可能变为零。为解决该问题,我们提出了一种具有指定物理速度(Prescribed Physical Speed, PPS)的新SF-GVF。该新SF-GVF的积分曲线从更高维空间(包括虚拟维度)中的任意初始条件指数收敛到期望路径;更重要的是,机器人的物理速度收敛到指定的物理速度(PPS),同时路径误差动力学在期望路径的正则重参数化下保持不变。我们进一步针对二阶运动学模型设计了一种饱和加速度控制律。最后,通过对比仿真以及四旋翼无人机在不同PPS曲线下的三维路径跟踪实验,验证了理论结果并证明了所提方法的有效性。
cs.RO / 46 / 2609.19742

Equivariant Filter Design for Acoustic and Depth Aided Inertial Navigation Systems

面向声学与深度辅助惯性导航系统的等变滤波器设计
Lunawat, Arihant, van Goor, Pieter, Dellaert, Frank, Williams, Stefan B.
Abstract
Autonomous Underwater Vehicles (AUVs) navigating without GPS typically fuse inertial measurements with acoustic Doppler Velocity Log (DVL) velocities and pressure-derived depth. Posing the navigation state on a Lie group improves accuracy and consistency. However, state-of-the-art filters based on the Invariant Extended Kalman Filter (IEKF) append the Inertial Measurement Unit (IMU) biases as a Euclidean extension, which breaks the group-affine structure required for exact log-linear error dynamics, causing the reported covariance to degrade alongside the estimate. We apply the Tangent-Group (TG) symmetry, which carries the biases within the geometry of the state space, to derive an Equivariant Filter (EqF) for this system, leaving zero linearization error in the navigation states and second-order error only in the biases. We develop an equivariant output model for the DVL, whose update incurs only third-order linearization error, together with a direct pressure output. Monte Carlo simulations benchmark the TG-EqF against a Two-Frame-Group IEKF and a Multiplicative EKF. The TG-EqF reduces error by 18--25\% against both alternatives in each of attitude, velocity, and position. The main benefit is in the covariance it estimates: its Average Normalized Estimation Error Squared (ANEES) stays closer to its nominal value of one than that of the others. Offline analysis on AUV field data corroborates the findings of the simulations, demonstrating reduced position drift.
Chinese Translation
无GPS环境下自主水下航行器(AUV)的导航通常将惯性测量与声学多普勒速度计(DVL)速度以及由压力推导的深度相融合。将导航状态置于李群上可以提高精度和一致性。然而,基于不变扩展卡尔曼滤波器(IEKF)的最先进滤波器将惯性测量单元(IMU)偏置作为欧几里得扩展附加处理,这破坏了实现精确对数线性误差动力学所需的群仿射结构,导致所报告的协方差随估计值一同退化。我们应用切群(Tangent-Group, TG)对称性——它将偏置纳入状态空间的几何结构之中——为该系统推导了一种等变滤波器(Equivariant Filter, EqF),使导航状态的线性化误差为零,偏置仅有二阶误差。我们为DVL开发了一种等变输出模型,其更新仅产生三阶线性化误差,并同时提供直接的压力输出。蒙特卡洛仿真将TG-EqF与双帧群(Two-Frame-Group)IEKF以及乘性EKF进行了对比基准测试。TG-EqF在姿态、速度和位置各方面的误差均比两种替代方法降低了18--25%。其主要优势体现在所估计的协方差上:其平均归一化估计误差平方(ANEES)比其他方法更接近其标称值1。基于AUV实地数据的离线分析验证了仿真结果,表明位置漂移有所减小。
cs.RO / 47 / 2609.19796

LIFD: Anchored Diffusion for 3D-Aware Scene Memory in Robotic Manipulation

LIFD:面向机器人操作的锚定扩散3D感知场景记忆
Li, Wenbo, Chen, Yiteng, Li, Wenhao, Wu, Qingyao
Abstract
Robotic manipulation under partial observability requires spatial information that extends beyond the current view. Geometry-aware RGB features describe visible structure, but previously observed regions may disappear as the robot or scene moves. Maintaining a useful scene representation therefore requires retaining observation history while inferring missing content without losing its connection to visible evidence. We introduce LIFD (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory. LIFD learns a scene-token representation from multi-view agreement and completes it from a single RGB view and recurrent memory. A rectified-flow model generates the tokens while Anchor-Guided Cross-Attention conditions completion on current geometric features. Compact slot features connect this representation to a manipulation policy. Multi-view and geometric supervision are used during representation learning; deployment requires one RGB camera, proprioception, and a task instruction. LIFD (Staged) reaches 91.6% average success on LIBERO and 79.8% on MetaWorld, improving LIBERO average success by 3.1 percentage points over Joint training. On four UR5e task families with ten demonstrations per family, it achieves 56.0% mean success, compared with 40.5% for OpenVLA-7B.
Chinese Translation
部分可观测条件下的机器人操作需要超越当前视野的空间信息。几何感知的RGB特征能够描述可见结构,但随着机器人或场景的移动,先前观测到的区域可能会消失。因此,维护一个有用的场景表示需要在保留观测历史的同时,推断缺失内容且不失去其与可见证据的联系。我们提出了LIFD(Look, Imagine, Focus, and Do,即观察、想象、聚焦与执行),一个持久化、3D感知的场景记忆框架。LIFD通过多视角一致性学习场景令牌(scene-token)表示,并基于单个RGB视图和循环记忆对其进行补全。一个矫正流(rectified-flow)模型生成这些令牌,同时锚定引导交叉注意力(Anchor-Guided Cross-Attention)使补全过程以当前几何特征为条件。紧凑的槽(slot)特征将该表示与操作策略相连接。表示学习阶段使用多视角和几何监督;部署时仅需一个RGB相机、本体感知和任务指令。LIFD(分阶段训练,Staged)在LIBERO上达到91.6%的平均成功率,在MetaWorld上达到79.8%,相比联合(Joint)训练将LIBERO平均成功率提升了3.1个百分点。在四个UR5e任务族(每个任务族十个演示样本)上,其平均成功率为56.0%,而OpenVLA-7B为40.5%。
cs.RO / 48 / 2609.19802

Affective Shared Autonomy: Temporal Affect Dynamics and Subjective Evaluation in Bimanual Teleoperation Tasks

情感共享自主:双手遥操作任务中的时序情感动态与主观评价
Liang, Zhengji, Tian, Guiyin, Qu, Sijin, Liu, Hainan, Hu, Shiyan
Abstract
Physical teleoperation integrates human cognitive flexibility with robotic precision, yet demanding manipulation tasks frequently induce severe cognitive workload, acute frustration, and execution breakdown. Conventional shared autonomy paradigms rely primarily on task-based rules, such as spatial error boundaries, which disregard the operator's transient affective state and risk misaligned control interventions. To address this limitation, we propose an affect-aware shared autonomy teleoperation framework that dynamically modulates robotic assistance based on real-time operator state estimation. The system estimates operator affective states from synchronized facial video, cardiac signals, and bilateral arm kinematics, outputting a seven-state affective distribution and a three-category operational abstraction (neutral, productive, adverse). Affect-aware assistance is selectively triggered when the user is detected in a continuous adverse state, preserving task-positive engagement without unnecessary disruption. The empirical user study ($N = 30$) confirms that the proposed affective assistance increases the productive states by up to 39.7% without compromising user agency. The collected dataset represents the first multimodal dataset that provides continuous visual, physiological, and operator's bilateral motion tracking of temporal affective state shifts during bimanual teleoperation. Our multimodal fusion model outperforms zero-shot baselines (Qwen, MiniCPM-V) in tracking temporal state dynamics. This real-world deployment offers a new human-centric framework that integrates visual, physiological, and motion tracking for physical human-robot interaction.
Chinese Translation
物理遥操作将人类的认知灵活性与机器人的精确性相结合,然而高要求的操作任务常常引发严重的认知负荷、强烈的挫败感以及执行失败。传统的共享自主范式主要依赖于基于任务的规则,例如空间误差边界,这类方法忽略了操作者的瞬时情感状态,并可能导致控制干预与实际需求不匹配。为解决这一局限,我们提出了一种情感感知的共享自主遥操作框架,该框架基于对操作者状态的实时估计来动态调节机器人辅助。系统通过同步的面部视频、心电信号以及双臂运动学数据估计操作者的情感状态,输出一个七类情感分布和一个三类操作抽象(中性、有效、不利)。当检测到用户处于持续不利状态时,才选择性地触发情感感知辅助,从而在不造成不必要干扰的前提下保持对任务有利的参与度。实证用户研究(N = 30)证实,所提出的情感辅助可将有效状态最多提升39.7%,且不损害用户的自主性。所收集的数据集是首个多模态数据集,提供了双手遥操作过程中情感状态时序变化的连续视觉、生理以及操作者双侧运动追踪数据。我们的多模态融合模型在追踪时序状态动态方面优于零样本基线模型(Qwen、MiniCPM-V)。这一真实场景部署提供了一个新的以人为中心的框架,整合了视觉、生理与运动追踪信息,服务于物理人机交互。
cs.RO / 49 / 2609.19803

HEROIC: Heterogeneous Evidential Reasoning for Open-Vocabulary Identification and Cross-Robot Collaboration

HEROIC:面向开放词汇识别与跨机器人协作的异构证据推理方法
Chauhan, Mihir, Jain, Aarav, Zucek, Addison, Dang, Manmeet, Conover, Damon, Bera, Aniket
Abstract
Multi-agent heterogeneous air-ground robot teams are attractive for open world search, with applications for reconnaissance, urban search and rescue missions (USAR), disaster response and recovery, and hazardous environments. These two platforms have different failure modes: aerial robots cover ground quickly but cannot resolve small or occluded targets from altitude, while ground robots can identify objects-of-interest, such as people or hazardous objects, at close range but cover less area. Existing language-tasked teams either have roles fixed prior, or have a language model assign them from hand-written capability tags, so the team is unable to know when within a mission an asset is no longer useful. We present HEROIC, a decentralized heterogeneous multi-agent open-vocabulary search coordination framework that requires agents to communicate in natural language only. HEROIC's initial agent role assignment is derived from sensor properties and a scale law to determine whether targets can be detected with a high confidence. From the mission's natural language prompt alone, this law assigns aerial flight altitudes and sweep spacing. When this calculated height falls below the altitude for safe flight, aerial agents re-task themselves from searcher to aerial triage, escort, and route guide for ground agents. Both robots maintain an evidential belief over the search area (bearing rays for positive evidence, a log-odds posterior for negative evidence) and gate any arrival on close-range verification. In full-stack experiments, HEROIC reaches the target 84% of the time across all 6 scenes, compares to 35-54% for vision-language frontier baselines, frontier-based search, lawnmower, and random-walk running the same perception, all while being 2-4x sooner to arrive at the target.
Chinese Translation
多智能体异构空地机器人团队在开放世界搜索中极具吸引力,可应用于侦察、城市搜索与救援任务(USAR)、灾害响应与恢复以及危险环境作业。这两类平台具有不同的失效模式:空中机器人可快速覆盖地面,但无法从高空分辨小型或被遮挡的目标;而地面机器人能够在近距离识别感兴趣目标(如人员或危险物体),但覆盖范围较小。现有的语言任务驱动机器人团队要么在任务前预先固定角色,要么由语言模型根据人工编写的能力标签进行分配,因此团队无法在任务执行过程中判断某个资产何时不再有用。我们提出了HEROIC,一个去中心化的异构多智能体开放词汇搜索协作框架,仅要求智能体之间以自然语言进行通信。HEROIC的初始智能体角色分配基于传感器特性和一个尺度定律,用于确定能否以高置信度检测到目标。仅凭任务的自然语言提示,该定律即可为空中智能体分配飞行高度和扫描间距。当计算出的高度低于安全飞行高度时,空中智能体会自行从搜索者重新分配为空中分诊、护航员或地面智能体的路线引导员。两类机器人均在搜索区域上维护一个证据信念模型(以方位射线表示正证据,以对数几率后验表示负证据),并对任何抵达事件进行近距离验证的门控。在全栈实验中,HEROIC在全部6个场景中的目标抵达率为84%,而在相同感知条件下,视觉语言前沿基线、基于前沿的搜索、割草式搜索和随机游走方法的抵达率仅为35-54%,同时HEROIC抵达目标的速度快2-4倍。
cs.RO / 50 / 2609.19813

Vehicle Trajectory Prediction via Neural Fusion of Multiple EKF-Based Trajectory Candidates

基于多个EKF轨迹候选的神经融合方法进行车辆轨迹预测
Kim, Seong-Jun, Kong, Seung-Hyun
Abstract
Predicting the future trajectories of surrounding vehicles in autonomous driving is important for collision risk assessment and safe ego-vehicle path planning. Conventional neural network-based trajectory predictors typically achieve strong prediction performance by exploiting agent history, dynamic scene graphs, and semantic maps. However, in specific motion regimes such as acceleration, deceleration, and turning, these predictors may fail to reflect physically feasible trajectories. To address this issue, this study proposes a framework that fuses the output of Trajectron++, a neural network-based trajectory predictor, with extended Kalman filter (EKF)-based multiple trajectory candidates at a late stage. On the nuScenes dataset, the proposed method reduces the average displacement error and final displacement error of the Trajectron++ robot baseline by 13.7% and 14.6%, respectively, without modifying the baseline architecture. These results indicate that EKF-based trajectory candidates can effectively complement neural trajectory prediction through learned fusion.
Chinese Translation
在自动驾驶中,预测周围车辆的未来轨迹对于碰撞风险评估和自车安全路径规划至关重要。传统的基于神经网络的轨迹预测器通常通过利用智能体历史信息、动态场景图和语义地图来实现较强的预测性能。然而,在加速、减速和转弯等特定运动状态下,这些预测器可能无法反映物理上可行的轨迹。为解决这一问题,本研究提出了一种在后期阶段将基于神经网络的轨迹预测器Trajectron++的输出与基于扩展卡尔曼滤波器(EKF)的多个轨迹候选进行融合的框架。在nuScenes数据集上,所提出的方法在不修改基线架构的情况下,将Trajectron++机器人基线的平均位移误差和最终位移误差分别降低了13.7%和14.6%。这些结果表明,基于EKF的轨迹候选可以通过学习到的融合有效补充神经轨迹预测。
cs.RO / 51 / 2609.19817

RotateIt! Fast and Reliable Single-Arm Garment Unfolding via Online-Adaptive Dynamic Rotation

RotateIt!:基于在线自适应动态旋转的快速可靠单臂衣物展开方法
Zhang, Zeqing, Xie, Zuokun, Fang, Ao, Dai, Bin, Shu, Zhengjie, Tang, Yifeng, Wang, Ziwei
Abstract
Robotic garment unfolding is essential for downstream tasks, yet quasi-static methods require repeated actions, while existing dynamic approaches predominantly rely on bimanual flinging. We present RotateIt!, a single-arm framework that uses adaptive axial rotation for dynamic garment unfolding. To the best of our knowledge, it is the first unfolding framework to employ dynamic axial rotation as its primary manipulation primitive. From a randomly initialized tabletop configuration, the robot selects a rotation-effective grasp and rotates the lifted garment about an approximately fixed anchor, generating inertial tension that separates overlapping layers within a compact workspace. A grasp ranker selects the anchor, while an online residual policy adapts the rotation extent and speed, thereby determining the release timing. Across seen and unseen simulated garments and eight unseen real garments, RotateIt! improves success within three attempts by 44.0-61.0 percentage points over quasi-static pick-and-place. The simulation-trained policies transfer zero-shot to the real world, achieving 75.6% success, 41% higher first-attempt coverage, and 26% higher final coverage. The resulting states further enable autonomous robotic folding without manual rearrangement.
Chinese Translation
机器人衣物展开是后续任务的关键环节,然而准静态方法需要反复执行动作,而现有的动态方法主要依赖双臂甩动。我们提出了RotateIt!,一种利用自适应轴向旋转实现动态衣物展开的单臂框架。据我们所知,这是首个以动态轴向旋转作为主要操作基元的展开框架。从随机初始化的桌面配置出发,机器人选择一个具有旋转效益的抓取点,并绕一个近似固定的锚点旋转被提起的衣物,产生惯性张力,从而在紧凑的工作空间内分离重叠的衣物层。系统采用抓取排序器选择锚点,同时利用在线残差策略自适应调整旋转的幅度和速度,从而确定释放时机。在已见和未见过的仿真衣物以及八件未见过的真实衣物上,RotateIt! 在三次尝试内的成功率相比准静态抓取放置方法提升了44.0至61.0个百分点。仿真训练的策略可零样本迁移至真实世界,实现了75.6%的成功率、高出41%的首试展开覆盖率以及高出26%的最终覆盖率。所得到的状态进一步支持了无需人工整理的自主机器人折叠。
cs.RO / 52 / 2609.19824

TADreamer: Zero-Shot Language-Guided 3D Navigation for Terrestrial-Aerial Bimodal Robots via Video Imagination

TADreamer:基于视频想象的陆空双模态机器人零样本语言引导三维导航
Li, Xiangyu, Lai, Tiancheng, Huang, Xijie, Pang, Ruitian, Shen, Siqi, Chen, Juncheng, Pan, Zaisheng, Xu, Chao, Gao, Fei, Cao, Yanjun
Abstract
Language-guided navigation for terrestrial-aerial bimodal robots requires selecting routes and locomotion modes that match scene context and task intent. Generated videos can represent such motion sequences, but recovering metrically consistent navigation references from them is challenging because of scale ambiguity and axis-dependent geometric distortions. We present TADreamer, a zero-shot framework that grounds video-imagined navigation in measured geometry without task-specific training or fine-tuning. A vision-language model translates onboard observations and instructions into navigation prompts, selects valid generated videos, and provides corrective feedback when regeneration is needed. The selected video is reconstructed into 3D waypoints annotated with terrestrial or aerial modes. A two-stage calibration procedure uses field-of-view constraints to initialize scale estimation, then refines axis-dependent scales, rotation, and translation by registering the reconstructed point cloud to measured geometry. The calibrated waypoints and mode labels guide a planner that incorporates measured geometry for robot execution. Real-world experiments demonstrate navigation across seven indoor and outdoor scenarios. With five candidates per round, usable videos are obtained within two rounds in all seven scenarios. On the calibration observations, our method reduces mean absolute depth error by 87.7% and mean absolute relative depth error by 86.3% compared with NavDreamer.
Chinese Translation
面向陆空双模态机器人的语言引导导航需要选择与场景上下文和任务意图相匹配的路线和运动模式。生成的视频可以表示此类运动序列,但由于尺度不确定性和依赖于坐标轴的几何畸变,从中恢复度量一致导航参考极具挑战性。我们提出了TADreamer,一个零样本框架,无需任务特定训练或微调,即可将视频想象的导航锚定在测量几何之上。视觉-语言模型将机载观测与指令转化为导航提示,筛选有效的生成视频,并在需要重新生成时提供纠正性反馈。被选中的视频被重建为标注有陆行或飞行模式的三维航路点。一个两阶段标定流程首先利用视场约束初始化尺度估计,然后通过将重建点云配准到测量几何,进一步优化依赖于坐标轴的尺度、旋转和平移。标定后的航路点与模式标签引导规划器结合测量几何以供机器人执行。真实世界实验验证了在七个室内外场景中的导航能力。每轮使用五个候选视频,在所有七个场景中均于两轮内获得可用视频。在标定观测上,与NavDreamer相比,我们的方法将平均绝对深度误差降低87.7%,平均绝对相对深度误差降低86.3%。
cs.RO / 53 / 2609.19846

Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision

基于动作相似性监督改进潜在动作模型中的跨本体迁移
Alvarez, Maxime, Caballero, Renzo, Matsushima, Tatsuya, Iwasawa, Yusuke, Matsuo, Yutaka
Abstract
As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain costly to collect and tied to the robot that recorded them. Latent action models (LAMs) address both by learning latent actions from action-free videos that can be shared across embodiments, however, in practice, LAMs are sensitive to background visual noise, and the same motion from two different robots may be encoded with different latents. One solution to the background visual noise is to add an auxiliary loss predicting the robot action from the latent action, further associating the latent action space to the embodiment specific robot action space. We study a different use of the same labels, through action-similarity supervision. The similarity between any two latent actions is trained to match the similarity of the two ground-truth robot action sequences. The ground-truth actions are never predicted by the LAM, so the latent action does not need to encode embodiment specifics. We evaluate cross-embodiment transfer on RoboTwin 2.0 in a controlled setup, two bimanual robots demonstrate disjoint task sets, a policy is trained on all the demonstrations, and each robot is evaluated closed-loop on the tasks only the other demonstrated. With the policy architecture and its hyperparameters, the dataset, and the evaluation protocol fixed, predicting latent actions instead of ground-truth actions more than doubles cross-embodiment success. Given the same ground-truth actions, similarity supervision transfers better than an auxiliary loss that predicts the ground-truth action during the LAM training. Computing the similarities on end-effector motion rather than joint-space motion, and letting the loss compare latent actions across the two robots, gives the best approach of the study.
Chinese Translation
随着通用机器人策略通过网络规模预训练获得视觉和语言能力,演示数据的收集成本依然高昂,且与采集它们的机器人本体紧密绑定。潜在动作模型(Latent Action Models, LAMs)通过从无动作标签的视频中学习可跨机器人本体共享的潜在动作,同时解决了这两个问题。然而在实际应用中,LAMs 对背景视觉噪声较为敏感,且来自两个不同机器人的相同运动可能被编码为不同的潜在表示。应对背景视觉噪声的一种方案是添加一个从潜在动作预测机器人动作的辅助损失,从而将潜在动作空间与特定本体的机器人动作空间进一步关联。我们研究了这些标签的另一种使用方式——动作相似性监督:训练任意两个潜在动作之间的相似性,使其匹配两个真实机器人动作序列之间的相似性。真实动作从不被 LAM 直接预测,因此潜在动作无需编码本体特定信息。我们在 RoboTwin 2.0 基准上于受控设置下评估跨本体迁移:两个双臂机器人分别演示互不重叠的任务集,在所有演示数据上训练一个策略,然后让每个机器人仅在对方演示过的任务上进行闭环评估。在策略架构及其超参数、数据集和评估协议均固定的条件下,预测潜在动作而非真实动作,使跨本体成功率提升了一倍以上。在使用相同真实动作的情况下,相似性监督比在 LAM 训练中预测真实动作的辅助损失具有更好的迁移效果。基于末端执行器运动而非关节空间运动计算相似性,并让损失在两个机器人的潜在动作之间进行比较,构成了本研究中表现最优的方法。
cs.RO / 54 / 2609.19850

GR2PO: Group Relative Return Policy Optimization for Continuous Robot Control

GR2PO:面向连续机器人控制的组相对回报策略优化
Wang, Pengqin, Zhang, Qiming, Shen, Shaojie, Ma, Jun
Abstract
Actor-critic architecture has been widely used in continuous robot control. However, they rely on learning a value network, introducing additional computational overhead during training. Moreover, policy learning may also be affected by the approximation error of value estimation. Critic-free group relative policy optimization methods provide a simpler training approach by removing the need for a critic. However, they fail to learn long-term action outcomes when directly applying immediate rewards to policy optimization in dense-reward environments. To address these problems, we propose Group Relative Return Policy Optimization (GR2PO), a critic-free reinforcement learning framework for continuous robot control. GR2PO estimates the discounted returns from the parallelly collected trajectories, performs group normalization at each rollout time index, and uses relative advantages and clipped targets to update the policy. To evaluate the effectiveness of the proposed framework, we instantiate it on robot control simulation environments and deploy the model to a real-world edge device. The results show that GR2PO significantly outperforms critic-free baselines that use immediate rewards and performs competitively against state-of-the-art actor-critic methods. Furthermore, GR2PO demonstrates competitive training efficiency. Inference tests on NVIDIA Jetson TX2 demonstrate the feasibility of deploying the learned policies on edge platforms. Further ablation experiments analyze the effects of parallel group size, return estimation methods, and target clipping ratio on learning performance. To support follow-up research, we will make the complete code publicly available after the paper is accepted, including the framework implementation, experimental configuration, and training and evaluation scripts.
Chinese Translation
Actor-critic(演员-评论家)架构已被广泛应用于连续机器人控制。然而,这类方法依赖于学习一个价值网络,在训练过程中引入了额外的计算开销。此外,策略学习也可能受到价值估计近似误差的影响。无评论家(critic-free)的组相对策略优化方法通过移除对评论家网络的需求,提供了一种更简单的训练方式。然而,在稠密奖励环境中直接将即时奖励用于策略优化时,这类方法无法学习到动作的长期结果。为解决这些问题,我们提出了组相对回报策略优化(Group Relative Return Policy Optimization,GR2PO),这是一种面向连续机器人控制的无评论家强化学习框架。GR2PO从并行采集的轨迹中估计折扣回报,在每个rollout时间步进行组归一化,并利用相对优势和截断目标来更新策略。为评估所提框架的有效性,我们将其应用于机器人控制仿真环境,并将模型部署到真实世界的边缘设备上。结果表明,GR2PO显著优于使用即时奖励的无评论家基线方法,并与最先进的actor-critic方法具有相当的性能。此外,GR2PO展现出具有竞争力的训练效率。在NVIDIA Jetson TX2上的推理测试证明了将所学策略部署于边缘平台的可行性。进一步的消融实验分析了并行组大小、回报估计方法和目标截断比例对学习性能的影响。为支持后续研究,我们将在论文被接收后公开完整代码,包括框架实现、实验配置以及训练和评估脚本。
cs.RO / 55 / 2609.19863

Feeling Terrain Before Crossing: World Models for Off-Road Navigation

跨越地形之前先感知地形:面向越野导航的世界模型
Son, E-In, Kim, Dong-Wook, Hwang, Ji-Hoon, Lee, Kangsun, Bae, Jisung, Kim, Jung-Taak, Seo, Seung-Woo
Abstract
Navigation world models plan by foresight, predicting the future that each candidate action sequence produces and selecting the best, rather than mapping observations to actions directly. Unlike urban settings where a predicted scene is a sufficient proxy, off-road navigation hinges on the robot--terrain interaction, so the prediction must cover not only what the camera will see but what the robot will feel. However, existing scene-focused models do not predict how much the robot will slip, tilt or shake along a planned trajectory. Proprioception captures these dynamics directly and, when used as input, improves the prediction of the physical future. We present Feel-WM, the first off-road navigation world model that conditions on proprioception and predicts what the robot will feel alongside what the camera will see. The physical future takes the form of a future proprioceptive state and a failure risk, both learned from the robot's own experience without human labels. The planner rolls out the physical future alongside the scene and weighs the predicted failure risk against goal similarity in a separable score. Experiments on real off-road data and in simulation demonstrate that Feel-WM outperforms visual-only navigation world models in open-loop planning and closed-loop rough-terrain navigation across wheeled and legged platforms. Deployed on a Husky on mountain trails, Feel-WM plans onboard, predicts rough ground ahead and steers around it, completing courses that an end-to-end policy fails.
Chinese Translation
导航世界模型通过前瞻进行规划:预测每个候选动作序列所产生的未来,并选择最优者,而不是直接将观测映射为动作。在城市环境中,预测的场景足以作为代理,但越野导航的关键在于机器人与地形的交互,因此预测不仅要涵盖相机将看到的内容,还要涵盖机器人将感受到的内容。然而,现有的以场景为中心的模型无法预测机器人沿规划轨迹将产生多大的打滑、倾斜或震动。本体感受(proprioception)能够直接捕捉这些动力学特性,作为输入使用时可改善对物理未来的预测。我们提出了Feel-WM,这是首个以本体感受为条件的越野导航世界模型,在预测相机将看到的场景的同时,预测机器人将感受到的物理状态。物理未来以未来的本体感受状态和失败风险的形式呈现,二者均从机器人自身经验中学习,无需人工标注。规划器同时推演物理未来与场景,并在一个可分离的评分中将预测的失败风险与目标相似度进行权衡。在真实越野数据和仿真中的实验表明,Feel-WM在开环规划和闭环崎岖地形导航中均优于仅依赖视觉的导航世界模型,并在轮式和足式平台上均得到验证。部署于山地路径上的Husky机器人时,Feel-WM能够在机载端完成规划,预测前方的崎岖地面并绕行,成功完成了端到端策略无法完成的路线。
cs.RO / 56 / 2609.19894

Learning Reliable Parking Policies via Offline Reinforcement Learning with Quantized Action Representations

基于量化动作表示的离线强化学习学习可靠的泊车策略
Yang, Zewei, Peng, Zengqi, Ma, Jun
Abstract
Parking is a routine yet safety-critical task for autonomous vehicles operating in urban environments. However, cluttered and weakly structured parking spaces, compounded by the interactive uncertainty from surrounding vehicles, hinder reliable maneuver generation. To address these challenges, we develop a waypoint-level offline reinforcement learning framework for interaction-aware autonomous parking. Specifically, a dedicated parking dataset is constructed from hierarchical expert rollouts with rotational waypoint augmentation, covering both non-interactive scenarios and interactive ones. The policy is then conditioned on a compact state representation, in which LiDAR-based obstacle features are adapted to the target pose via feature-wise linear modulation. A state-conditioned tokenizer further quantizes continuous waypoint sequences into discrete action tokens, over which conservative Q-learning is performed to suppress value overestimation on poorly supported actions. Extensive closed-loop experiments are conducted in the high-fidelity CARLA simulator. The proposed framework attains the highest parking success rate among all baselines and transfers reliably to unseen parking slots.
Chinese Translation
泊车是自动驾驶汽车在城市环境中运行时一项常规但关乎安全的任务。然而,杂乱且弱结构化的泊车空间,加之周围车辆带来的交互不确定性,阻碍了可靠操纵轨迹的生成。为应对这些挑战,我们开发了一个用于交互感知自动驾驶泊车的路径点级离线强化学习框架。具体而言,我们基于分层专家轨迹并结合旋转路径点增强构建了专门的泊车数据集,涵盖非交互式和交互式两类场景。策略以紧凑的状态表示为条件,其中基于激光雷达(LiDAR)的障碍物特征通过特征级线性调制适配到目标位姿。状态条件化的分词器进一步将连续的路径点序列量化为离散动作令牌,并在其上执行保守Q学习(Conservative Q-Learning),以抑制对支撑不足动作的价值高估。我们在高保真的CARLA仿真器中进行了大量闭环实验。所提出的框架在所有基线方法中取得了最高的泊车成功率,并能可靠地迁移到未见过的泊车位。
cs.RO / 57 / 2609.19906

Learning and Transferring Closed-Loop Robot Software

闭环机器人软件的学习与迁移
Kuroki, So, Tang, Yujin
Abstract
Closed-loop robot policies require observation processing, state management, and situation-dependent branching, making them costly to design and tune manually. Although coding agents increasingly support control-code generation and optimization, it remains unclear whether implementations improved on source tasks also support policy acquisition for new tasks. We study this question by treating complete closed-loop implementations as reusable execution experience. For each source task, a coding agent generates policy code from a few successful demonstrations and iteratively improves it using simulation feedback. The validation-selected implementations are retained in a software archive. For new tasks, the agent generates and improves policies using archived implementations, target demonstrations, and execution feedback. The resulting policy is then frozen and executes without further model calls. Across four source tasks in RoboCasa, iterative optimization increases mean success from 28.3% to 64.2%. Across nine target tasks and three independent runs, mean success is 45.2% without references, 41.5% with initial source code, and 57.0% with optimized source code. Optimized references outperform initial references in all three runs on the nine-task average, with a mean gain of 15.6 percentage points. These results demonstrate the value of execution-improved software as a resource for acquiring new policies in this setting, although initial references remain better on two target tasks when averaged across runs.
Chinese Translation
闭环机器人策略需要观测处理、状态管理以及依据情境的分支控制,这使得其人工设计与调优成本高昂。尽管代码智能体日益支持控制代码的生成与优化,但在源任务上改进的实现是否也能支持新任务的策略获取,仍不清楚。我们通过将完整的闭环实现视为可复用的执行经验来研究这一问题。对于每个源任务,代码智能体从少量成功演示中生成策略代码,并利用仿真反馈对其进行迭代改进。经验证筛选的实现被保存在软件存档中。对于新任务,智能体利用存档中的实现、目标演示以及执行反馈来生成并改进策略。最终策略被冻结后无需进一步的模型调用即可执行。在RoboCasa的四个源任务上,迭代优化将平均成功率从28.3%提升至64.2%。在九个目标任务、三次独立运行中,无参考时的平均成功率为45.2%,使用初始源代码时为41.5%,使用优化后的源代码时为57.0%。在九任务平均值上,优化后的参考在所有三次运行中均优于初始参考,平均增益为15.6个百分点。这些结果表明,在此场景下,经过执行改进的软件作为获取新策略的资源具有价值,尽管在按运行取平均时,初始参考仍在两个目标任务上表现更优。
cs.RO / 58 / 2609.19912

Distributed Model Predictive Control with Connectivity-based Contracts

基于连通性约束的分布式模型预测控制
Geurts, Jorit, Saccani, Danilo, Zeilinger, Melanie N., Carron, Andrea
Abstract
Teams of mobile robots rely on continuous communication with their neighbors for coordination, yet most distributed model predictive control (DMPC) schemes assume the communication network stays connected rather than actively enforcing it. Adding such a guarantee is hard since the usual mathematical condition for connectivity is nonconvex and links every agent to every other, which is incompatible with a scalable distributed real-time controller. We propose a DMPC framework in which each agent is assigned a connectivity contract: a local region prescribing where its predicted positions may lie over the prediction horizon. The contracts are designed so that, as long as every agent stays within its own contract, the team is guaranteed to remain connected. Given the maintained contract graph, an agent builds its contract from a single exchange with its immediate neighbors, after which every agent solves its own optimization problem independently. We prove that the resulting closed-loop system maintains connectivity, avoids collisions, and respects local state and input constraints. Simulation and hardware experiments on miniature autonomous car-like robots demonstrate the approach.
Chinese Translation
移动机器人团队依赖与邻居的持续通信来实现协调,然而大多数分布式模型预测控制(DMPC)方案假设通信网络始终保持连通,而非主动保障其连通性。引入这种保证十分困难,因为连通性的常规数学条件是非凸的,且将每个智能体与其他所有智能体相关联,这与可扩展的分布式实时控制器不兼容。我们提出了一种DMPC框架,为每个智能体分配一个连通性契约(connectivity contract):即在预测时域内规定其预测位置所处范围的局部区域。契约的设计使得只要每个智能体保持在自身契约范围内,就能保证整个团队保持连通。在维持契约图的前提下,每个智能体仅需与其直接邻居进行一次信息交换即可构建自身契约,之后每个智能体独立求解各自的优化问题。我们证明了所得到的闭环系统能够保持连通性、避免碰撞,并满足局部状态与输入约束。在微型自主轮式机器人上的仿真与硬件实验验证了该方法的有效性。
cs.RO / 59 / 2609.19923

Co-VLA: Consensus-based Federated Training for Vision-Language-Action Models

Co-VLA:基于共识的视觉-语言-动作模型联邦训练方法
Li, Haolong, Er, Guner Dilsad, Muehlebach, Michael, Stueckler, Joerg
Abstract
Vision-language-action models (VLAs) have emerged as a promising paradigm for general-purpose robot learning, with performance improving as models and datasets scale. Scaling robot data collection, however, remains challenging because data are naturally distributed across robots, tasks, and locations, making centralization costly or impractical. Federated learning offers a way to train on decentralized robot data, but applying it to VLAs requires accounting for heterogeneous robot client data distributions. We present Co-VLA, which applies consensus optimization using the Alternating Direction Method of Multipliers~(ADMM) to federated VLA training. We show that the same algorithm supports both full-model training and parameter-efficient fine-tuning with both fixed-rank and rank-adaptive adapters. The name Co-VLA reflects both consensus and collaboration: clients with different local robot datasets collaboratively train a shared model without sharing their data. Our experiments demonstrate that Co-VLA achieves performance comparable to centralized training in both full-model training and parameter-efficient fine-tuning settings.
Chinese Translation
视觉-语言-动作模型(Vision-Language-Action Models, VLAs)已成为通用机器人学习的一种有前景的范式,其性能随着模型和数据集规模的扩大而不断提升。然而,机器人数据采集的规模化仍然具有挑战性,因为数据天然分布在不同的机器人、任务和地点之间,使得集中化收集成本高昂或难以实现。联邦学习为在分布式机器人数据上进行训练提供了一种途径,但将其应用于VLA模型需要考虑各机器人客户端数据分布的异构性。我们提出了Co-VLA,该方法利用乘子交替方向法(Alternating Direction Method of Multipliers, ADMM)进行共识优化,应用于VLA模型的联邦训练。我们证明,同一算法既支持全模型训练,也支持使用固定秩和自适应秩适配器的参数高效微调。Co-VLA这一名称同时体现了共识与协作的含义:拥有不同本地机器人数据集的客户端在不共享数据的前提下协同训练一个共享模型。实验结果表明,Co-VLA在全模型训练和参数高效微调两种设置下均达到了与集中式训练相当的性能。
cs.RO / 60 / 2609.19946

Execution-Aware Pre-Execution Ranking for Grasp-Conditioned Robotic Placement

面向抓取条件化机器人放置的执行感知预执行排序方法
Liu, Tianyuan, Patamia, Rutherford Agbeshi, Champion, Benjamin, Dazeley, Richard, Cosgun, Akansel
Abstract
A geometrically valid placement can still be difficult to execute because the selected grasp changes the required end-effector pose, collision geometry, and transport motion. Placement is formulated as a pre-execution ranking problem in which supplied grasp-placement candidates are scored before planning. The model combines a typed target-conditioned point cloud with three pose descriptors and hierarchical heads for planning success and execution success conditioned on planning. On a 30-object, 1,235-scene dataset with scene-group-held-out splits, three-seed top-1 success on covered test groups reaches 85.63 +/- 1.08% for joint selection and 79.84 +/- 0.16% for fixed-target ranking. For the designated frozen seed-42 checkpoint, top-1 success improves from 72.84% to 85.78% over full-pool cuMotion for joint ranking and from 59.65% to 79.67% for fixed-target ranking. Frozen transfer to xArm7/MoveIt requires no xArm-specific retraining. Across 27 locked cases, 13 complete end to end (48.15%). Of the 16 cases that pass Top-5 preflight and begin execution, 13 succeed (81.25%). Candidate-level deployment-feasibility prediction reaches 81.25% recall, 85.20% specificity, and 83.23% balanced accuracy.
Chinese Translation
几何上有效的放置方案仍可能难以执行,因为所选抓取方式会改变所需的末端执行器位姿、碰撞几何以及搬运运动。本文将放置问题建模为一个预执行排序问题,即在规划之前对给定的抓取-放置候选方案进行评分。该模型将带有类型标注的目标条件点云与三个位姿描述子相结合,并采用分层预测头,分别预测规划成功率以及在规划成功条件下的执行成功率。在一个包含30个物体、1,235个场景的数据集上(采用按场景组划分的留出测试集),三次随机种子实验中,覆盖测试组上的top-1成功率为:联合选择达85.63 ± 1.08%,固定目标排序达79.84 ± 0.16%。对于指定的冻结seed-42检查点,与全候选池cuMotion相比,联合排序的top-1成功率从72.84%提升至85.78%,固定目标排序从59.65%提升至79.67%。冻结模型向xArm7/MoveIt的迁移无需针对xArm进行再训练。在27个锁定测试用例中,13个(48.15%)完整实现了端到端执行。在通过Top-5预检并开始执行的16个用例中,13个(81.25%)成功完成。候选级别的部署可行性预测达到81.25%的召回率、85.20%的特异性和83.23%的平衡准确率。
cs.RO / 61 / 2609.19954

LapaTrack-3D: 6 DoF pre-operative shape tracking for laparoscopic surgery

LapaTrack-3D:面向腹腔镜手术的术前形状六自由度跟踪
Song, Jingwei, Jakir, Javid Hussain, Zhang, Ray, Zhang, Wenwei, Zhou, Hao, Xian, Xiaomeng, Ghaffari, Maani
Abstract
This work proposes a real-time 6 Degree-of-Freedom (6 DoF) tracking algorithm for monocular laparoscopic surgery. It provides alignment between intra-operative video and pre-operative data (e.g., CT). The 6 DoF tracking offers a solution for accurately locating the internal anatomy of the target organ despite the lack of tactile feedback and transparency. The ORB-SLAM2 framework is adopted and modified for prior-based 3D tracking with four major modifications. First, the primitive 3D shape is used for fast initialization of the ORB-SLAM2 monocular mode. Second, a pseudo-segmentation strategy is employed to separate the target organ from the background for tracking. Third, the 3D shape is incorporated as a geometric prior in its pose graph optimization. Fourth, the Multi-Scale Retinex with Chromaticity Preservation (MSRCP) algorithm is leveraged and modified for image enhancement in challenging illumination scenarios. In-vivo and ex-vivo experiments validate that LapaTrack-3D provides robust 3D tracking and effectively handles typical challenges such as poor illumination, fast motion, out-of-field-of-view scenarios, partial visibility, and ``organ-background'' relative motion. LapaTrack-3D achieves a processing rate of 13 Hz for 1280*720 pixel video.
Chinese Translation
本工作提出了一种用于单目腹腔镜手术的实时六自由度(6 DoF)跟踪算法。该算法实现了术中视频与术前数据(如CT)之间的对齐。六自由度跟踪在缺乏触觉反馈和透明度信息的情况下,为精确定位目标器官的内部解剖结构提供了解决方案。本研究采用并改进了ORB-SLAM2框架,以实现基于先验的三维跟踪,主要包含四项改进:第一,利用原始三维形状实现ORB-SLAM2单目模式的快速初始化;第二,采用伪分割策略将目标器官从背景中分离出来进行跟踪;第三,将三维形状作为几何先验纳入位姿图优化中;第四,采用并改进了保留色度的多尺度Retinex(MSRCP)算法,以在具有挑战性的光照场景下进行图像增强。体内和体外实验验证了LapaTrack-3D能够提供稳健的三维跟踪,并有效应对诸如光照不佳、快速运动、目标离开视野、部分可见以及“器官-背景”相对运动等典型挑战。LapaTrack-3D对1280*720像素的视频可达到13 Hz的处理速率。
cs.RO / 62 / 2609.19962

Hybrid Residual Reinforcement Learning for Contact-Rich Robotic Book Insertion

面向接触密集型机器人插书任务的混合残差强化学习方法
Liu, Tianyuan, Patamia, Rutherford Agbeshi, Champion, Benjamin, Cosgun, Akansel, Dazeley, Richard
Abstract
Placing a grasped book into a tight shelf is a compact but difficult contact-rich control problem: millimetre-scale pose error can turn a geometrically valid approach into jamming, failed release, or incomplete seating. We study this final phase after grasp acquisition and global approach, and ask how control authority should be divided between known geometry and learned behaviour. Our method retains a nominal task-space controller for structured insertion and seating, while residual PPO supplies bounded local corrections and decides when to release. Only the brief open-retreat-reclose transition is scripted. For the final policy used on hardware, a deployment-matched simulation evaluation over 512 fixed conditions yields 98.50 percent mean success (0.23 percentage-point sample SD) across three independent training runs, compared with 37.89 percent for nominal control. On the physical xArm7, 60 trials over 30 matched conditions show the same qualitative advantage: residual control raises success from 26.7 percent to 63.3 percent, reduces failures from 22 to 11, and wins 13 of the 15 matched conditions in which the two controllers differ. Robustness tests show that performance remains above 87 percent under initialization perturbations up to 1.5x, while very tight clearances expose the geometric limit of local correction. These results support a hybrid design in which geometry preserves reliable task structure and learning is concentrated on the contact-sensitive behaviour that fixed rules handle poorly.
Chinese Translation
将抓取的书籍放入紧凑的书架是一个简洁但困难的接触密集型控制问题:毫米级的位姿误差就可能使几何上可行的接近过程变成卡死、释放失败或未完全插入。我们研究了抓取和全局接近之后的这一最终阶段,并探讨控制权限应如何在已知几何信息与学习到的行为之间进行划分。我们的方法保留了名义任务空间控制器用于结构化的插入和就位,同时残差PPO(Proximal Policy Optimization)提供有界的局部修正并决定何时释放。只有简短的开爪-后退-再合拢过渡过程是脚本化的。对于部署在硬件上的最终策略,在512个固定条件下进行与部署匹配的仿真评估,三次独立训练的平均成功率为98.50%(样本标准差0.23个百分点),而名义控制仅为37.89%。在物理xArm7机械臂上,针对30个匹配条件进行的60次试验显示出同样的定性优势:残差控制将成功率从26.7%提高到63.3%,失败次数从22次减少到11次,并在两个控制器表现不同的15个匹配条件中赢得13个。鲁棒性测试表明,在高达1.5倍的初始化扰动下,性能仍保持在87%以上,而极小的间隙则暴露了局部修正的几何极限。这些结果支持一种混合设计:几何信息保证可靠的任务结构,而学习集中于固定规则难以处理的接触敏感行为。
cs.RO / 63 / 2609.19974

MaskHarness-WAM: Instance-Grounded Harnessing for Long-Horizon Robot Manipulation

MaskHarness-WAM:面向长时程机器人操作的实例接续控制框架
Huang, Zitai, Su, Taiyi, Zhu, Jian, Zhang, Jianjun, Ma, Chong, Liu, Tianbin, Lu, Weiyi, Xu, Yi, Wang, Hanli
Abstract
Long-horizon robot manipulation requires not only stable local visuomotor control, but also continuous target tracking and reliable task progress assessment throughout execution. This challenge becomes particularly critical when multiple objects share identical appearances and must be manipulated in a prescribed order. In such scenarios, relying solely on a limited-horizon manipulation policy is often insufficient to determine which instance should be operated on and when the task should transition to the next stage. To address this challenge, we propose MaskHarness-WAM, an instance-grounded harness for long-horizon manipulation. The proposed system connects high-level task planning with low-level manipulation policies through target masks, while leveraging visual feedback for subtask scheduling and continuous execution. Since each subtask corresponds to a different target instance, the low-level policy requires a newly established initial target mask under the updated scene at each subtask transition. The harness continuously re-observes the environment, generates, and verifies the target mask at subtask boundaries, thereby updating the instance-level spatial condition provided to the low-level policy. Furthermore, the system advances the manipulation process by switching target instances according to the verified completion status of each subtask. Experiments on a real robot platform demonstrate that MaskHarness-WAM substantially outperforms limited-horizon policies on sequential multi-object manipulation, showing its effectiveness in extending local manipulation skills to reliable long-horizon execution.
Chinese Translation
长时程机器人操作不仅需要稳定的局部视觉运动控制,还需要在整个执行过程中进行持续的目标跟踪和可靠的任务进度评估。当多个物体外观完全相同且必须按指定顺序进行操作时,这一挑战尤为突出。在此类场景中,仅依靠有限时程的操作策略往往不足以确定应当操作哪一个实例以及何时将任务转换到下一阶段。为应对这一挑战,我们提出了MaskHarness-WAM,一种面向长时程操作的实例接续控制框架。该系统通过目标掩码将高层任务规划与底层操作策略相连接,同时利用视觉反馈进行子任务调度和持续执行。由于每个子任务对应不同的目标实例,底层策略需要在每次子任务转换时,依据更新后的场景重新建立初始目标掩码。该框架在子任务边界处持续重新观测环境,生成并验证目标掩码,从而更新提供给底层策略的实例级空间条件。此外,系统根据每个子任务经验证的完成状态切换目标实例,从而推进操作进程。在真实机器人平台上的实验表明,MaskHarness-WAM在序贯多物体操作任务上显著优于有限时程策略,证明了其在将局部操作技能扩展为可靠的长时程执行方面的有效性。
cs.RO / 64 / 2609.19976

Compliance for Free: Learning Identifiable Impedance via Bilateral Teleoperation

零成本实现柔顺性:通过双边遥操作学习可辨识的阻抗
Guda, Harsha, Colomé, Adrià, Torras, Carme
Abstract
Vision-language-action models tell a robot where to move, but not how hard to push. Contact-rich tasks depend on that second quantity, compliance, yet no widely used demonstration interface records it. The obstacle is identifiability as realized pose and measured force cannot separate the operator's intended equilibrium from their stiffness, so VR controllers, SpaceMouse and handheld grippers cannot supply compliance supervision even in principle. Prior compliance-output policies work around this with hand-specified task structure, privileged simulation contact state, or dedicated force and tactile hardware. Four-channel bilateral teleoperation removes the ambiguity directly by using the leader arm as a separate measurement of the intended equilibrium, making per-axis stiffness identifiable by regression using only the joint-torque sensing already on the manipulator. This yields per-timestep, direction-dependent compliance labels at zero annotation cost, which we use to fine-tune a VLA to emit stiffness alongside pose. On a Franka Research 3 wiping task, ours is the only policy of five whose contact force changes when the instruction asks for a firm wipe rather than a normal one (6.4N (normal) to 9.1N (firm) RMS, Cohen's d = 0.89, p = 0.023
Chinese Translation
视觉-语言-动作(Vision-Language-Action,VLA)模型告诉机器人往哪里移动,但无法告诉它该用多大的力去推。接触丰富的任务依赖于后者——柔顺性(compliance),然而目前尚无广泛使用的示教数据采集界面能够记录这一信息。其障碍在于可辨识性问题:由于实际位姿与测量力无法将操作者期望的平衡位置与其刚度分离开来,因此即使是VR控制器、SpaceMouse和手持夹爪,原则上也无法提供柔顺性监督信号。已有的柔顺性输出策略通过人工指定任务结构、使用特权仿真接触状态,或依赖专用的力/触觉硬件来绕过这一问题。四通道双边遥操作则直接消除了这一歧义:它将主臂(leader arm)作为对期望平衡位置的独立测量,使得各轴刚度可以通过仅使用机械臂上已有的关节力矩传感器进行回归而辨识出来。由此,我们能够以零标注成本获得逐时间步、与方向相关的柔顺性标签,并用其对VLA模型进行微调,使其在输出位姿的同时输出刚度。在Franka Research 3机械臂的擦拭任务上,在五个策略中,只有我们的策略在指令要求“用力擦拭”而非“正常擦拭”时接触力会发生变化(RMS力从6.4N(正常)变为9.1N(用力),Cohen's d = 0.89,p = 0.023)。
cs.RO / 65 / 2609.20035

DR-MPC: Fast and Feasible Dynamics-Relaxed Model-Predictive Control for Legged Locomotion

DR-MPC:面向足式运动的高效可行动力学松弛模型预测控制
Wang, Run, Tuerxun, Alapati, Liu, Shuo, Xiao, Wei, Drgoňa, Ján, Mo, Yilin, Wu, Liang
Abstract
This paper presents dynamics-relaxed model predictive control (DR-MPC), a novel MPC formulation for legged locomotion, and a tailored interior-point method (IPM) solver. The formulation combines online optimization feasibility by construction with a contact-aware input parameterization. DR-MPC moves the dynamics equality and affine input constraints into quadratic penalties and retains only nonempty box constraints. The resulting box-constrained quadratic program (QP) has a block-arrow Hessian that enables the state and affine-output directions to be eliminated through a Schur complement. The solver factors only the reduced control system after swing-force elimination and contact-aligned move blocking. For the evaluated implementations using the same DR-MPC formulation, our method achieves median end-to-end MPC speedups of $16.0\times$ over HPIPM and $4.4\times$ over OSQP, with comparable locomotion performance in simulation. DR-MPC achieves a median onboard MPC end-to-end time of $4.4$ ms and is validated on a Unitree Go1 quadruped. Open-source code will be made available after publication.
Chinese Translation
本文提出了一种面向足式运动的新型模型预测控制(MPC)方法——动力学松弛模型预测控制(DR-MPC),以及一个为其量身定制的内点法(IPM)求解器。该方法在构建上保证了在线优化的可行性,并结合了接触感知的输入参数化。DR-MPC 将动力学等式约束和仿射输入约束移入二次惩罚项中,仅保留非空的箱式约束。由此得到的箱式约束二次规划(QP)具有块箭头结构的海森矩阵(Hessian),可通过舒尔补(Schur complement)消去状态和仿射输出方向。经过摆动力的消除和与接触对齐的移动阻塞处理后,该求解器仅需对降阶后的控制系统进行因式分解。在采用相同 DR-MPC 公式的评估实现中,我们的方法相比 HPIPM 实现了中位数 $16.0\times$ 的端到端 MPC 加速,相比 OSQP 实现了 $4.4\times$ 的加速,且在仿真中具有相当的运动性能。DR-MPC 的中位数机载 MPC 端到端时间为 $4.4$ 毫秒,并在 Unitree Go1 四足机器人上进行了验证。开源代码将于论文发表后公开。
cs.RO / 66 / 2609.20048

Mechanical Precision Weeding with a Quadruped Robot

基于四足机器人的机械式精准除草
Beumer, Ruben, Janssen, Tom, van de Molengraft, René, Antunes, Duarte
Abstract
Herbicide-based weed control is increasingly unsustainable due to rising weed resistance and the adverse environmental impacts of chemical use. While mechanical weed control avoids these drawbacks, it is typically implemented using large machines that cause soil compaction. We propose a novel alternative based on small mobile robots for mechanical weeding. Compared with existing automated mechanical weeding approaches, the proposed method offers reduced soil compaction, simpler automation, and improved scalability. Our solution involves a Boston Dynamics Spot quadruped robot equipped with a custom weed removal tool featuring a milling bit at its end. The tool is rigidly attached to the robot and uses the degrees of freedom of the robot base by actuating the legs, while keeping the feet stationary. We develop a software architecture that enables autonomous weed removal and integrate this system with all other required components. We analyze the accuracy and efficiency of the current proof of concept both in an indoor and outdoor environment and provide recommendations for future work to make the system more accurate and efficient.
Chinese Translation
由于杂草抗药性不断增强以及化学品使用对环境的不利影响,基于除草剂的杂草控制方式正变得越来越不可持续。虽然机械除草可以避免这些缺点,但通常采用大型机械实施,会造成土壤压实。我们提出了一种基于小型移动机器人进行机械除草的新型替代方案。与现有的自动化机械除草方法相比,所提出的方法具有土壤压实更少、自动化更简单以及可扩展性更好等优势。我们的方案使用一台波士顿动力公司(Boston Dynamics)的Spot四足机器人,其上配备了一个定制的除草工具,该工具末端装有一个铣削头。该工具刚性地固定在机器人上,并通过驱动腿部运动利用机器人基座的自由度,同时保持足端静止。我们开发了一套能够实现自主除草的软件架构,并将该系统与所有其他所需组件集成在一起。我们在室内和室外环境中分析了当前概念验证系统的准确性和效率,并为进一步提高系统的精度和效率提出了未来工作的建议。
cs.RO / 67 / 2609.20078

FlipToSee: A Probabilistic Stable Placement Prior for Active Visual Exploration via Regrasping

FlipToSee:一种基于重新抓取的主动视觉探索的概率性稳定放置先验
Shu, Chang, Dinesh, Sushil Samuel, Park, Shinkyu
Abstract
Active visual exploration of tabletop objects often requires reorienting an unknown resting object onto a different stable support face to expose occluded surfaces. To identify such placements without exhaustive physical search, we learn a probabilistic placement prior from a single-view point cloud. Stable placement prediction is inherently multimodal, and conventional 6-DoF regression introduces further ambiguity by modeling translation and in-plane yaw. We therefore propose FlipToSee, a probabilistic framework that removes this representational ambiguity by parameterizing placements as unit support normals on $S^2$ while modeling their multimodal conditional distribution via a von Mises--Fisher mixture density network. To decouple mode diversity from physical robustness, FlipToSee deterministically extracts a compact candidate set from the mixture components and applies robustness-aware reranking using an auxiliary head trained with candidate-aligned supervision. In simulation, FlipToSee achieves $98.4\%$ first-proposal success on in-distribution objects, $95.3\%$ on out-of-distribution shapes, and $90.0\%$ under zero-shot transfer to household YCB objects. We further demonstrate the learned placement prior on a physical robot by integrating it with grasp and motion planning for exploratory regrasping.
Chinese Translation
对桌面物体的主动视觉探索通常需要将一个未知放置姿态的物体重新定向到不同的稳定支撑面上,以暴露被遮挡的表面。为了避免穷举式的物理搜索来识别此类放置方式,我们从单视角点云中学习一个概率性放置先验。稳定放置预测本质上是多模态的,而传统的六自由度(6-DoF)回归由于同时对平移和平面内偏航角进行建模,会引入额外的歧义。因此,我们提出 FlipToSee,一个通过将放置参数化为单位球面 $S^2$ 上的单位支撑法向量来消除这种表示歧义的概率框架,并利用冯·米泽斯–费舍尔(von Mises–Fisher)混合密度网络对其多模态条件分布进行建模。为了将模式多样性与物理鲁棒性解耦,FlipToSee 从混合分量中确定性地提取一个紧凑的候选集合,并利用一个通过候选对齐监督训练的辅助头来执行鲁棒性感知的重排序。在仿真中,FlipToSee 在分布内物体上达到 $98.4\%$ 的首次提议成功率,在分布外形状上达到 $95.3\%$,在零样本迁移到日常 YCB 物体上达到 $90.0\%$。我们进一步将该学习到的放置先验与抓取和运动规划相结合,在实际机器人上实现了探索性重新抓取。
cs.RO / 68 / 2609.20087

Graph-Based Design of Soft Grippers with Multi-Objective Quality-Diversity Optimisation

基于图结构设计与多目标质量-多样性优化的软体夹持器设计
Farinha, Andre, Shi, Ge, Bowman, Harry, Tidd, Brendan, Howard, David, Pinskier, Josh
Abstract
Effective manipulation across diverse objects is critical for applications ranging from agricultural harvesting to laboratory and domestic automation. While the inherent compliance of soft robotics is well suited to this challenge, designing grippers that generalize across tasks remains difficult due to the vast design space of continuum mechanics and the risk of overfitting to specific scenarios. We propose a graph-based design space for representing soft structures and mechanisms, coupled with a multi-objective, diversity-driven genetic optimization framework that explicitly promotes solution variety throughout the design process. Using multiple grasping scenarios during optimization, we study how task diversity influences the emergence of generalization to unseen objects and contact conditions. Our results show that optimization over a sufficiently diverse set of grasping cases leads to designs with emergent generalization, exhibiting improved robustness compared to task-specific solutions on novel scenarios. These findings suggest that diversity-driven optimization offers a principled pathway toward general-purpose soft grippers, aligned with the adaptable nature of soft robotics.
Chinese Translation
对多样化物体的有效抓取操控对于从农业采摘到实验室及家庭自动化等应用至关重要。尽管软体机器人固有的柔顺性非常适合应对这一挑战,但由于连续介质力学设计空间庞大且存在针对特定场景过拟合的风险,设计能够跨任务泛化的夹持器仍然十分困难。我们提出了一种基于图结构的软体结构与机构设计空间表示方法,并结合多目标、多样性驱动的遗传优化框架,在整个设计过程中显式地促进解的多样性。通过在优化过程中使用多种抓取场景,我们研究了任务多样性如何影响对未见物体和接触条件的泛化能力的涌现。结果表明,在足够多样化的抓取案例集合上进行优化,可以得到具有涌现泛化能力的设计,在新场景下相比任务特定解决方案表现出更强的鲁棒性。这些发现表明,多样性驱动的优化为设计通用软体夹持器提供了一条有原则的路径,并与软体机器人自适应的天然特性相契合。
cs.RO / 69 / 2609.20103

Safety-Critical Scenanrio Emerges from Initial Scene

安全关键场景从初始场景中涌现
Wu, Yin, Wei, Jiarong, Esselborn, Carl, Phoolari, Shubham, Abouelazm, Ahmed, Slieter, Daniel, Zöllner, J. Marius
Abstract
Safety-critical driving scenario generation has largely focused on manipulating the behavior of surrounding agents while starting from an initial scene from driving data. This assumption can limit the space of discoverable failures, since driving data can provide little opportunity for meaningful interaction. For example, in the Waymo Open Motion Dataset, 20.44% of recorded slices feature a stationary ego vehicle that never moves, and 30.39% of initial frames contain no nearby traffic participants within 10 meters. We instead study safety-critical scenario generation as an initialization problem: given agnostic black-box driving policies, we learn to generate realistic initial scenes that are more likely to evolve into critical interactions. We propose AdvScene, a conditional latent diffusion model that is trained in two stages. Starting from pretraining on naturalistic driving data, we post-train the adversarial-agent generation branch using reinforcement learning with feedback from closed-loop simulator rollouts. Conditioning on ego driving displacement prevents the ego from remaining static, and RL finetuning induces criticality directly with non-differentiable safety-critical metrics. Experiments on the Waymo dataset across 12 combinations of ego and traffic policies show that our AdvScene substantially increases the rate of ego-fault collision events and TTC<3s events.
Chinese Translation
安全关键驾驶场景生成的研究大多集中于在从驾驶数据获得的初始场景基础上操纵周围智能体的行为。这一假设会限制可发现失效情形的空间,因为驾驶数据可能几乎无法提供有意义的交互机会。例如,在Waymo开放运动数据集中,20.44%的记录片段中自车(ego vehicle)始终静止不动,30.39%的初始帧在10米范围内不存在任何附近交通参与者。我们转而将安全关键场景生成视为一个初始化问题:给定不可知的黑盒驾驶策略,我们学习生成更可能演变为关键交互的真实初始场景。我们提出AdvScene,一种分两阶段训练的条件潜在扩散模型。在自然驾驶数据上进行预训练后,我们利用闭环仿真器回放反馈的强化学习对对抗智能体生成分支进行后训练。以自车行驶位移作为条件可防止自车保持静止,而强化学习微调则通过不可微的安全关键指标直接诱导关键性。在Waymo数据集上跨12种自车与交通策略组合的实验表明,我们的AdvScene显著提高了自车责任碰撞事件和TTC<3秒事件的发生率。
cs.RO / 70 / 2609.20107

AnyViewDex: View-Invariant Dexterous Manipulation from RGB Observations

AnyViewDex:基于RGB观测的视角不变灵巧操作
Patil, Soham, Gunjal, Om Sanjay, Bhosale, Sourabh, Chavare, Arhan, Hora, Ramandeep Singh, Roy, Spandan
Abstract
Visuomotor policies for multi-fingered dexterous manipulation are highly sensitive to camera viewpoint shifts. To achieve view invariance, recent methods increasingly rely on explicit 3D modalities like RGB-D or point clouds, which can introduce hardware dependencies, calibration requirements, and vulnerability to sensor noise during real-world deployment. In this work, we show that view-invariant control can be achieved without explicit test-time 3D sensing by encoding geometric knowledge into the visual representation during simulation. We present AnyViewDex, an asymmetric training pipeline that combines multi-view contrastive alignment with privileged 3D geometric supervision. By regressing absolute 3D object coordinates during simulated training, this auxiliary objective provides a geometric grounding signal that mitigates the spatial collapse of the globally pooled contrastive embedding. At deployment, the policy operates zero-shot using only uncalibrated monocular RGB and proprioception. We validate this approach across both reinforcement learning and student-teacher distillation. In hardware evaluation on an xArm7 with a 16-DoF LEAP Hand, AnyViewDex reaches 76.7% grasping success across eight unseen objects and six uncalibrated viewpoints (480 trials; 2,400 across all ablation conditions), indicating that geometrically grounded monocular policies transfer zero-shot without test-time depth. Project Page: https://anyviewdex.github.io/
Chinese Translation
面向多指灵巧操作的视觉运动策略对相机视角变化高度敏感。为实现视角不变性,近期方法日益依赖RGB-D或点云等显式3D模态,但这会引入硬件依赖、标定需求,以及在真实部署中对传感器噪声的脆弱性。在本工作中,我们证明了无需测试时显式3D感知即可实现视角不变控制,其方法是在仿真训练阶段将几何知识编码到视觉表示中。我们提出AnyViewDex,一种将多视角对比对齐与特权3D几何监督相结合的非对称训练流水线。通过在仿真训练中对物体的绝对3D坐标进行回归,该辅助目标提供了几何锚定信号,缓解了全局池化对比嵌入的空间塌缩问题。在部署时,该策略仅需未标定的单目RGB和本体感觉即可实现零样本运行。我们在强化学习和师生蒸馏两种框架下验证了该方法。在配备16自由度LEAP Hand的xArm7硬件评估中,AnyViewDex在八个未见物体和六个未标定视角下达到76.7%的抓取成功率(480次试验;所有消融条件共计2,400次),表明具有几何锚定的单目策略无需测试时深度信息即可实现零样本迁移。项目页面:https://anyviewdex.github.io/
cs.RO / 71 / 2609.20114

Universal Navigation Interface: Robot-Free Data for Wheeled Robot Navigation

通用导航接口(UNI):面向轮式机器人导航的无机器人数据
Prajapati, Sarvesh, Trivedi, Ananya, Genua, Lorena Maria, Moore, Drake, Maxwell, Bruce, Padir, Taskin
Abstract
Collecting real-world navigation data for mobile robots typically requires platform-specific teleoperation, making large-scale collection expensive and difficult to scale. We introduce Universal Navigation Interface (UNI), a robot-free data collection paradigm that uses a four-wheeled rollator walker (rollator) and smartphone to collect physically constrained human demonstrations. Because the rollator cannot climb stairs, negotiate uncut curbs, or pass through narrow gaps, demonstrations are naturally biased toward wheeled-feasible routes. Using UNI, we collect 37.2 km of real-world navigation data and recover metric trajectories that directly supervise goal-conditioned navigation models. Fine-tuning visual-navigation models on UNI reduces trajectory prediction error by 17.4-24.8% on held-out UNI demonstrations. Evaluation on other navigation datasets shows benefits that vary by dataset and metric. We further demonstrate closed-loop transfer to a powered wheelchair in curb, staircase, and curb-cut scenarios. These results support low-cost physical proxies as a practical source of navigation supervision collected without the target robot.
Chinese Translation
为移动机器人采集真实世界导航数据通常需要针对特定平台的遥操作,使得大规模数据采集成本高昂且难以扩展。我们提出了通用导航接口(Universal Navigation Interface, UNI),一种无机器人的数据采集范式,利用四轮助行器(rollator)和智能手机采集具有物理约束的人类示范。由于助行器无法爬楼梯、越过未修整的路缘或通过狭窄缝隙,所采集的示范天然偏向于轮式可行的路线。借助 UNI,我们采集了 37.2 公里的真实世界导航数据,并恢复出度量轨迹,可直接用于监督目标条件导航模型。在 UNI 上微调视觉导航模型,可在留出的 UNI 示范上将轨迹预测误差降低 17.4%–24.8%。在其他导航数据集上的评估表明,其收益因数据集和指标而异。我们进一步展示了模型在电动轮椅上于路缘、楼梯和缘石坡道场景中的闭环迁移能力。这些结果支持以低成本的物理替代装置作为导航监督数据的实用来源,且无需目标机器人即可完成采集。
cs.RO / 72 / 2609.20116

GPT-6-Astra in a Navigation Workflow: Behavioral Analysis in Zero-Shot Vision-and-Language Navigation in Continuous Environments

GPT-6-Astra在导航工作流中的应用:连续环境中零样本视觉语言导航的行为分析
Dai, Guangzhao, Wu, Qi, Zhu, Bin
Abstract
We study GPT-6-Astra in a zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) system, where it interprets instructions, assesses its surroundings, and proposes actions. The system uses a common observation--decision--execution workflow with direct model API calls, without a packaged agent harness or navigation-specific fine-tuning. In this workflow, each request receives selected observations, execution feedback, and retained progress records. Evaluation covers the complete system, including context management and action control. We evaluate the system on 50 of the 100 R2R-CE val-unseen episodes used by Open-Nav. It achieves a success rate of 52.0\%, an SPL of 48.9\%, and an nDTW of 70.8\%. Our analysis highlights three findings. First, recorded responses link landmarks and earlier actions to instructions using observations and supplied history. Second, reviews include requests for additional views and revisions of uncertain judgments. Third, the results suggest a gap between task understanding and autonomous completion: an unfinished crossing is recognized while rotation continues. At termination, 36.0\% of episodes succeed with a workflow-accepted STOP, while another 16.0\% meet the distance criterion at the step limit. These results highlight a central challenge: translating correct local judgments into sustained progress and appropriate stopping.
Chinese Translation
我们在零样本连续环境视觉语言导航(VLN-CE)系统中研究了GPT-6-Astra,该模型负责理解指令、评估周围环境并提出动作。该系统采用通用的“观察—决策—执行”工作流,直接调用模型API,无需封装的智能体框架,也无需导航相关的微调。在此工作流中,每次请求都会接收所选的观察信息、执行反馈以及保留的进度记录。评估覆盖完整系统,包括上下文管理和动作控制。我们在Open-Nav所使用的100个R2R-CE val-unseen片段中的50个上评估该系统,取得了52.0%的成功率、48.9%的SPL和70.8%的nDTW。我们的分析得出三点发现:第一,记录的响应利用观察结果和提供的历史信息,将地标和先前动作与指令关联起来;第二,系统在回顾过程中会请求额外的视角并修正不确定的判断;第三,结果表明任务理解与自主完成任务之间存在差距:系统能识别出未完成的穿越,但旋转动作仍在继续。在终止时,36.0%的片段通过工作流接受的STOP指令成功完成,另有16.0%的片段在步数上限处满足距离标准。这些结果凸显了一个核心挑战:如何将正确的局部判断转化为持续的进展和恰当的停止。
cs.RO / 73 / 2609.20191

VLN on the Fly: An Onboard Vision-Language Navigation Stack for Aerial Robots

即时视觉语言导航:面向空中机器人的机载视觉语言导航系统
Tayar, Marco S., Tommaselli, Felipe, Capezutto, Gianluca, Saraiva, Pedro Antonio Rabelo, de Freitas, Pedro H. V., Kido, Lucas, Sonego, Guilherme, Godoy, Ricardo V., Becker, Marcelo
Abstract
Running vision-language navigation fully onboard an aerial robot is hard, since grounding, planning, and control must share limited compute and a single-stage error is difficult to isolate in flight. End-to-end aerial policies fuse these stages into one network, giving up the observability and safety checks a modular stack keeps available. We propose VLN on the Fly, an onboard stack that keeps grounding, planning, and control as separate, inspectable stages. A quantized VLM grounds an instruction to a coarse image cell, depth lifts it to a 3D goal, a fast B-spline planner returns a feasible trajectory, and a pretrained reinforcement learning policy tracks it to motor commands across quadrotors. Across 15 onboard flights over three everyday referents in a controlled indoor volume, the stack reaches the target in 13 of 15 trials with 5.72 cm mean goal error and 39.3% average GPU utilization. In 6 additional cluttered-environment trials, the stack tracks collision-free trajectories under onboard perception gating.
Chinese Translation
在空中机器人上完全机载运行视觉语言导航十分困难,因为视觉定位、规划与控制必须共享有限的计算资源,且单级误差在飞行过程中难以隔离排查。端到端空中策略将这些阶段融合进单一网络,放弃了模块化系统所能保持的可观测性与安全检查。我们提出VLN on the Fly,一个将视觉定位、规划和控制保持为独立、可检查阶段的机载系统。量化后的视觉语言模型(VLM)将指令定位到粗粒度的图像单元,深度信息将其提升为三维目标点,快速B样条规划器生成可行轨迹,预训练的强化学习策略将其跟踪为电机指令,适用于四旋翼无人机。在受控室内空间中针对三个日常指代物进行的15次机载飞行实验中,该系统在15次尝试中有13次到达目标,平均目标误差为5.72厘米,平均GPU利用率为39.3%。在另外6次复杂环境试验中,该系统在机载感知门控下实现了无碰撞轨迹跟踪。
cs.RO / 74 / 2609.20227

Strategic Transformer for Resource-Constrained Multi-Object Navigation in Ultra-Large-Scale Environments

面向超大规模环境资源受限多目标导航的战略Transformer
Iwata, Daiki, Tanaka, Kanji, Hishida, Senta
Abstract
Resource-constrained multi-object navigation in vast indoor environments ($>2,000\text{ m}^2$) poses significant challenges for efficiency and strategic planning. To tackle this, we reformulate the task as a Set Orienteering Problem (SOP), providing an optimization framework under resource constraints where exploitation is governed by the SOP model and exploration is managed by a separate heuristic switcher. Conventional baselines suffer from either rigid planning or myopic behaviors. To overcome these limitations and resolve the NP-hard computational challenges of SOP for real-time navigation, we develop the Strategic Transformer. This lightweight architecture functions as a priority planner that internalizes expert combinatorial logic into a predictable $41.03\text{ ms}$ forward pass while reducing teacher-student information asymmetry. Incorporating geometric attention biases allows the network to effectively model long-range structural dependencies. By coupling the Transformer's macro-plan with a bounded iterative 2-opt local refinement on a capped candidate graph, our framework achieves a $94\times$ speedup compared to heavy meta-heuristics, ensuring bounded-latency inference suitable for onboard deployment. Experiments on ProcTHOR validate that our method successfully bridges the gap between exploration and exploitation, outperforming carefully re-implemented baselines under Progress weighted by Path Length (PPL) and establishing a new benchmark for scalable, resource-constrained navigation.
Chinese Translation
在超大面积室内环境($>2,000\text{ m}^2$)中进行资源受限的多目标导航,对效率和战略规划提出了重大挑战。为解决这一问题,我们将该任务重新表述为集合定向问题,提供了一个资源约束下的优化框架:其中开发由SOP模型控制,探索则由独立的启发式切换器管理。传统基线方法要么规划僵化,要么行为短视。为克服这些局限并解决SOP用于实时导航时的NP难计算挑战,我们开发了战略Transformer。这一轻量级架构作为优先级规划器,将专家组合逻辑内化为可预测的$41.03\text{ ms}$前向传播,同时降低了师生之间的信息不对称。通过引入几何注意力偏置,网络能够有效建模长程结构依赖关系。通过将Transformer的宏观规划与有界候选图上的有界迭代2-opt局部精炼相结合,我们的框架相比重量级元启发式算法实现了$94\times$的加速,确保了适合机载部署的有界延迟推理。在ProcTHOR上的实验验证了我们的方法成功弥合了探索与开发之间的鸿沟,在以路径长度加权的进度(PPL)指标上优于经过精心复现的基线方法,为可扩展的资源受限导航建立了新的基准。
cs.RO / 75 / 2609.20318

LLM-Guided Transformation of Non-Critical Driving Scenes into Safety-Critical Scenarios Using Augmented Reality

基于大语言模型引导的利用增强现实将非关键驾驶场景转化为安全关键场景
Fady, Noura, Khaled, Farah, Elias, Catherine M.
Abstract
Testing Autonomous Driving Systems (ADS) requires realistic safety-critical scenarios, but collecting such data from real-world driving is costly and unsafe. This paper presents an automated pipeline that transforms safe driving scenes into safety-critical scenarios by combining computer vision, Large Language Models (LLMs), and Augmented Reality (AR). The system detects and tracks road users, extracts safety features including distance, velocity, motion direction, and Time-to-Collision (TTC), and assesses scene criticality. Safe scenes are modified by an LLM, which generates realistic collision-inducing objects and behaviors that are integrated into the original scene using AR. The proposed pipeline was evaluated on the nuScenes dataset, achieving 97.52% safety classification accuracy and successfully generating realistic scenarios such as pedestrian crossings, rear overtaking vehicles, and sudden-stop events. The results demonstrate an effective and flexible approach for automated generation of safety-critical scenarios to support the testing and validation of autonomous driving systems.
Chinese Translation
自动驾驶系统(ADS)的测试需要真实的安全关键场景,但从真实驾驶中采集此类数据成本高昂且不安全。本文提出了一种自动化流程,结合计算机视觉、大语言模型(LLM)和增强现实(AR),将安全驾驶场景转化为安全关键场景。该系统检测并跟踪道路使用者,提取包括距离、速度、运动方向和碰撞时间(TTC)在内的安全特征,并评估场景的临界程度。安全场景由大语言模型进行修改,生成真实可信的诱发碰撞的物体和行为,并利用增强现实技术将其融入原始场景。该流程在 nuScenes 数据集上进行了评估,取得了 97.52% 的安全分类准确率,并成功生成了行人横穿、后方超车车辆和急停事件等真实场景。结果表明,这是一种有效且灵活的方法,可用于自动化生成安全关键场景,以支持自动驾驶系统的测试与验证。
cs.RO / 76 / 2609.20321

Implementation of Tightly-Coupled SLAM Fusion of GPS, IMU, and LiDAR for Autonomous Vehicles

面向自动驾驶车辆的GPS、IMU与激光雷达紧耦合SLAM融合实现
Elmehrath, Amr O., Khaled, Farah, Nahas, Rana, Elias, Catherine M.
Abstract
Autonomous vehicles depend entirely on Simultaneous Localization and Mapping (SLAM) to navigate safely in unknown environments. However, relying on a single sensory modality introduces critical failure points: LiDAR systems degrade in featureless corridors, Inertial Measurement Units (IMUs) accumulate mathematical drift, and GPS drops frequently in urban canyons. This paper presents the implementation of a tightly-coupled SLAM fusion architecture that integrates a Velodyne 3D LiDAR, a high-frequency IMU, and GPS to achieve continuous spatial awareness. Utilizing a phased development methodology, we establish a 2D baseline to validate hardware synchronization and transform geometries before upgrading to a full 3D architecture driven by FAST-LIO2. This advanced approach uses an Iterated Error-State Kalman Filter (IESKF) to process dense 3D laser points alongside continuous inertial data, eliminating motion blur at high speeds. To eradicate long-term drift, a GTSAM pose-graph optimization back-end executes multi-modal loop closures. Evaluated across simulated environments and physical deployments, the results demonstrate that tightly-coupled 3D fusion effectively overcomes individual sensor blind spots to generate highly detailed point clouds, providing the foundational High-Definition (HD) maps required for advanced downstream autonomous planners.
Chinese Translation
自动驾驶车辆完全依赖同步定位与建图(SLAM)在未知环境中实现安全导航。然而,依赖单一传感器模态会引入关键的失效点:激光雷达(LiDAR)系统在无特征走廊中性能退化,惯性测量单元(IMU)会产生累积性数学漂移,而GPS在城市峡谷中频繁失效。本文实现了一种紧耦合SLAM融合架构,集成Velodyne三维激光雷达、高频IMU和GPS,以实现连续的空间感知。采用分阶段开发方法,我们首先建立二维基线以验证硬件同步与变换几何关系,随后升级为基于FAST-LIO2驱动的完整三维架构。该先进方法使用迭代误差状态卡尔曼滤波器(IESKF)处理稠密三维激光点云与连续惯性数据,消除了高速运动下的运动模糊。为消除长期漂移,后端采用GTSAM位姿图优化执行多模态回环检测。在仿真环境和实际部署中的评估结果表明,紧耦合三维融合有效克服了单一传感器的盲区,生成高度精细的点云,为高级下游自动驾驶规划器提供了所需的高精(HD)地图基础。
cs.RO / 77 / 2609.20330

RoboFind: Multi-Agent Personalized Object Search for People Who Are Blind or Have Low Vision

RoboFind:面向盲人或低视力人群的多智能体个性化物体搜索
Liu, Ruiping, Quan, Shaofang, Yin, Qian, Zhang, Jingqi, Zheng, Junwei, Chen, Yufan, Wen, Di, Fan, Weijia, Yang, Kailun, Sarfraz, M. Saquib, Asfour, Tamim, Peng, Kunyu, Stiefelhagen, Rainer
Abstract
Blind and low-vision users often need to locate a specific personal object rather than an arbitrary instance of the same category. The task calls for a robot that can move through the space and reach viewpoints the user cannot, and for an accessible interface where the user says which object is meant and learns whether the right one was found. We present RoboFind, a multi-agent framework in which a smartphone teaches the target and a quadruped robot carries out the search. A Target Teaching Agent converts guided smartphone recordings into a semantic target profile and a reusable multi-view reference bank through an accessible capture flow with AR guidance, speech and haptic feedback, and screen-reader support, so later missions refer to a stored object without repeating the teaching process. At runtime, a Navigation Agent explores the environment and proposes candidate targets, a Verification Agent checks each candidate against the stored references, and a Coordination and Recovery Agent completes the mission or triggers recovery and continued search. Across 32 real-robot missions, RoboFind reaches 85.0% success against 25.0% for a reconstructed sequential first-stop baseline over 20 trials with ten targets, and reduces false success from 75.0% to 5.0%. On six shared targets it succeeds in 10/12 trials, against 5/12 for 12 independently executed GPT-6 Astra-only trials. These results show that the multi-agent design fits the demands of personalized object search, where verifying object identity before declaring completion is what makes the outcome something a user can rely on.
Chinese Translation
盲人和低视力用户通常需要找到某件特定的个人物品,而非同一类别的任意实例。该任务要求机器人能够在空间中移动并到达用户无法到达的观察点,同时需要一个无障碍界面,让用户可以说明指的是哪件物品,并得知是否找到了正确的物品。我们提出了RoboFind,这是一个多智能体框架,其中智能手机用于教授目标信息,四足机器人执行搜索。目标教学智能体(Target Teaching Agent)通过带有AR引导、语音和触觉反馈以及读屏器支持的无障碍采集流程,将引导式的智能手机录制内容转换为语义目标档案和可复用的多视角参考库,从而使后续任务可以直接引用已存储的物品,无需重复教学过程。在运行阶段,导航智能体(Navigation Agent)探索环境并提出候选目标,验证智能体(Verification Agent)将每个候选目标与存储的参考库进行比对,协调与恢复智能体(Coordination and Recovery Agent)负责完成任务或触发恢复与继续搜索。在32次真实机器人任务中,RoboFind在20次试验、10个目标的条件下达到85.0%的成功率,而重构的顺序式首站基线仅为25.0%,并将误报成功率从75.0%降低至5.0%。在6个共享目标上,RoboFind在12次试验中成功10次,而12次独立执行的仅使用GPT-6 Astra的试验仅成功5次。这些结果表明,多智能体设计契合个性化物体搜索的需求——在宣告完成之前验证物体身份,正是让用户能够信赖结果的关键所在。
cs.RO / 78 / 2609.20388

Navi-Agent: Unlocalized Monocular Navigation Agent

Navi-Agent:无定位单目导航智能体
Xie, Wenyuan, Hong, Mengyang, Wang, Yongzhong, Ji, Yanbiao, Zhou, Yijin, Wu, Shaokai, Sirejiding, Shalayiding, Zhou, Huayi, Chen, Yi-Chao, Ling, Ma, Ding, Yue, Lu, Hongtao
Abstract
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to execute long-horizon instructions in unknown environments. Existing zero-shot VLN-CE systems typically maintain spatial states through geometric localization or coordinate-based representations. Recent geometry-constrained navigation removes depth and globally consistent coordinates, but maintaining persistent spatial awareness for place confirmation, progress verification, and recovery remains challenging. We present Navi-Agent, a zero-shot VLN-CE agent that constructs a coordinate-free spatial state from visual observations and executed motion histories. Navi-Agent organizes this state as a navigation topology, where nodes represent visual places and edges represent motion transitions. This representation enables observation-based approximate self-localization, task progress verification, and visual revisitation-based recovery. Navi-Agent performs closed-loop navigation by decomposing instructions into sub-goals, executing local visual navigation, and verifying visited places through the constructed spatial state. Experiments on zero-shot VLN-CE benchmark and real-world robot platforms show that Navi-Agent achieves state-of-the-art performance among geometry-constrained methods while remaining competitive with approaches relying on geometric localization.
Chinese Translation
连续环境中的视觉语言导航(Vision-Language Navigation in Continuous Environments, VLN-CE)要求具身智能体在未知环境中执行长时程指令。现有的零样本VLN-CE系统通常通过几何定位或基于坐标的表示来维护空间状态。近期基于几何约束的导航方法虽然去除了深度信息和全局一致的坐标,但维持持续的空间感知以进行地点确认、进度验证和恢复仍具有挑战性。我们提出了Navi-Agent,这是一个零样本VLN-CE智能体,它能够根据视觉观测和已执行的运动历史构建一种无坐标的空间状态。Navi-Agent将该状态组织为一种导航拓扑结构,其中节点表示视觉地点,边表示运动转换。这种表示方法支持基于观测的近似自定位、任务进度验证以及基于视觉重访的恢复机制。Navi-Agent通过将指令分解为子目标、执行局部视觉导航,并借助所构建的空间状态验证已访问地点,从而实现闭环导航。在零样本VLN-CE基准测试和真实机器人平台上的实验表明,Navi-Agent在基于几何约束的方法中取得了最先进的性能,同时与依赖几何定位的方法相比也具有竞争力。
cs.RO / 79 / 2609.20396

Imagine-TAMP: Imagination-Guided Task and Motion Planning in Partial Observability

Imagine-TAMP:部分可观测性下基于想象的任务与运动规划
Singha, Antareep, Kumar, Shivaram, Kim, Yoonwoo, Sung, Yoonchang
Abstract
Robots operating in cluttered environments must often manipulate objects whose locations are only partially observable. A central challenge is deciding whether to acquire another observation or to first manipulate objects that may occlude the target. Conventional task and motion planning (TAMP) approaches typically make this decision using symbolic action costs or expensive geometric planning, neither of which adequately captures how likely an observation is to reveal an occluded target. We introduce Imagine-TAMP, an interleaved planning and execution framework that uses semantic and geometric imagination to compare alternative task-level strategies under partial observability before committing to expensive motion planning. A vision-language model shapes a particle belief over target locations using commonsense relationships between the target and visible objects, while a generative scene model estimates plausible geometry in unobserved regions. Given a target hypothesis and imagined scene, Imagine-TAMP generates multiple symbolic plan skeletons and assigns non-unit costs that approximate both manipulation effort and target visibility from sensing actions, distinguishing a short but poorly informative observation strategy from a longer strategy that first manipulates an occluder to better expose the target. The selected skeleton is then refined into a feasible continuous plan and executed, with new observations updating the belief and triggering replanning when necessary. Experiments show that imagination-guided evaluation improves observation-versus-manipulation decisions: in viewpoint-constrained shelf scenes, non-unit geometric evaluation increases success from 46.0% to 84.0%, while semantic belief shaping further reduces manipulation and replanning. On a real robot, the complete system reduces planning time by 32% relative to a geometry-only ablation.
Chinese Translation
在杂乱环境中作业的机器人常常需要操控位置仅部分可观测的物体。一个核心挑战是决定是获取新的观测,还是先移动可能遮挡目标的物体。传统的任务与运动规划(TAMP)方法通常依据符号动作代价或昂贵的几何规划来做出这一决策,但两者都无法充分刻画一次观测能够揭示被遮挡目标的可能性。我们提出 Imagine-TAMP,一个规划与执行交替进行的框架,它在部分可观测条件下、在投入昂贵的运动规划之前,利用语义想象与几何想象来比较不同的任务级策略。视觉-语言模型利用目标与可见物体之间的常识关系,构建目标位置上的粒子信念,同时生成式场景模型估计未观测区域中合理的几何结构。给定目标假设和想象场景后,Imagine-TAMP 生成多个符号规划骨架,并为其分配非单位代价,以近似操控动作的开销与感知动作对目标可见性的影响,从而区分“耗时短但信息量少的观测策略”与“先移动遮挡物以更好地暴露目标的较长策略”。随后,所选骨架被细化为可行的连续规划并执行,新的观测会更新信念并在必要时触发重规划。实验表明,基于想象的评估改善了观测与操控之间的决策:在视角受限的货架场景中,非单位几何评估将成功率从 46.0% 提升至 84.0%,而语义信念构建进一步减少了操控次数与重规划次数。在真实机器人上,完整系统相比仅使用几何信息的消融版本将规划时间减少了 32%。
cs.RO / 80 / 2609.20407

Resilient Motion Planning for Free-Flying Space Robots under Actuator Failures

执行器故障下自由飞行空间机器人的弹性运动规划
de Maddalena, Nicolas, Verhagen, Joris, Tumova, Jana
Abstract
Free-flying robots rely on multiple thrusters to maneuver in space. If one or more of these thrusters fail, the robot may lose control authority and risk mission failure. At the same time, their free-flying nature implies that, even in the absence of actuation, they continue along (locally) straight-line trajectories. In this work we present a probabilistic, proactive, motion planning framework that explicitly accounts for actuator failures in space. We model actuator failure modes as a Markov chain and propagate the probability of successfully reaching the goal along the planning horizon. Precomputed reachable sets evaluate the robot's capabilities of reaching waypoints under potential failures and an RRT$^*$-based planner concatenates these waypoints. The resulting algorithm maximizes the overall target-reaching probability, providing maximally resilient motion plans utilizing free-flying properties. We validate our approach experimentally on a physical free-flyer platform with injected actuator failures.
Chinese Translation
自由飞行机器人依靠多个推进器在空间中进行机动。如果一个或多个推进器发生故障,机器人可能失去控制能力,面临任务失败的风险。同时,其自由飞行特性意味着,即使没有推进力,机器人仍会沿(局部)直线轨迹继续运动。本文提出了一种概率性的、主动式的运动规划框架,显式地考虑了空间中的执行器故障。我们将执行器故障模式建模为马尔可夫链,并在规划时域内传播成功到达目标的概率。通过预计算的可达集评估机器人在潜在故障下到达路径点的能力,并采用基于RRT$^*$的规划器将这些路径点连接起来。所提出的算法最大化了整体目标到达概率,利用自由飞行特性提供了最大程度弹性的运动规划。我们在一个注入执行器故障的物理自由飞行平台上对该方法进行了实验验证。
cs.RO / 81 / 2609.20435

Time-Efficient Iterative Learning Planning for Safety-Critical Dynamic Obstacle Avoidance

面向安全关键动态避障的时间高效迭代学习规划
Chen, Zhiyi, Lv, Shuli, Min, Chen, Xu, Yong, Sun, Jian, Quan, Quan
Abstract
Autonomous mobile robots require timeefficient planning and safety-critical dynamic obstacle avoidance under constrained onboard computation. While Iterative Learning Planning (ILP) offers lightweight and efficient traversal planning, it lacks explicit mechanisms for dynamic obstacle perception and avoidance. This article extends ILP to safety-critical navigation in dynamic environments by integrating an anticipatory risk-blended control barrier function (ARB-CBF). The extended ILP learns traversal-speed and steering-bias profiles via a fractionalpower update based on local obstacle risk, generating nominal control commands that ARB-CBF modifies at runtime for real-time safety guarantees. Algorithmic analysis demonstrates that the ILP replanning stage scales at O(kN) for k iterations and N waypoints, while ARB-CBF executes with linear complexity. Comprehensive simulations and real-world experiments validate the framework, demonstrating superior temporal efficiency and safety with lower computational overhead compared to optimizationbased baselines, making it highly suitable for resourceconstrained platforms.
Chinese Translation
自主移动机器人需要在受限的机载计算条件下实现时间高效的规划与安全关键的动态避障。虽然迭代学习规划(ILP)能够提供轻量且高效的遍历规划,但其缺乏针对动态障碍物感知与规避的显式机制。本文通过集成前瞻性风险混合控制障碍函数(ARB-CBF),将ILP扩展至动态环境中的安全关键导航。扩展后的ILP通过基于局部障碍物风险的分数幂更新来学习遍历速度与转向偏置曲线,生成标称控制指令;ARB-CBF在运行时对该指令进行修正,以提供实时安全保障。算法分析表明,ILP的重规划阶段在k次迭代与N个路径点下的复杂度为O(kN),而ARB-CBF以线性复杂度运行。全面的仿真与真实世界实验验证了该框架的有效性,与基于优化的基线方法相比,其展现出更优的时间效率与安全性,同时计算开销更低,非常适合资源受限的平台。
cs.RO / 82 / 2609.20443

Spatial-Semantic Uncertainty in VLM-Based Target Search: Balancing Exploration and Identification

基于视觉语言模型目标搜索中的空间-语义不确定性:探索与识别的平衡
Srivastava, Alkesh K., Diller, Jonathan, Kumar, Vijay, Dames, Philip
Abstract
Robots searching for a target from a natural-language description must determine not only where to search, but also which observed candidate is the desired target. These decisions reflect two distinct sources of uncertainty - spatial uncertainty over candidate locations and semantic uncertainty over target identity - that are often conflated in VLM-based search systems. We introduce a spatial-semantic uncertainty formulation that maintains separate beliefs over each component and integrates probabilistic VLM evidence into a global target-identity posterior, including probability mass for undiscovered targets. This decomposition allows an information-theoretic planner to independently value candidate discovery and target disambiguation through spatial and semantic expected information gain (EIG), providing an explicit mechanism for trading broader exploration against earlier identification. We evaluate six VLM uncertainty-elicitation interfaces on 500 synthetic targets and show that similar recognition accuracy can conceal substantial differences in calibration and false confidence. In degraded-observation search-and-identify experiments, EIG-based planners reach confident decisions in 75.0%-92.5% of trials, compared with 20.0% for Random search, while different spatial-semantic weightings achieve comparable identification accuracy once confidence is attained. Increasing semantic emphasis reduces unnecessary exploration and VLM queries, demonstrating that explicitly planning over semantic uncertainty can accelerate target resolution without sacrificing decision quality. These results highlight the distinct roles of uncertainty representation and uncertainty-driven planning in embodied VLM systems.
Chinese Translation
机器人根据自然语言描述搜索目标时,不仅需要确定搜索位置,还需要判断观察到的候选物体中哪个是目标。这些决策反映了两种不同的不确定性来源——候选位置的空间不确定性和目标身份的语义不确定性——而在基于视觉语言模型(VLM)的搜索系统中,这两者往往被混为一谈。我们提出了一种空间-语义不确定性建模方法,对每个组成部分分别维护独立的信念,并将概率化的VLM证据融入全局目标身份后验分布中,其中包含尚未发现目标的概率质量。这种分解使得基于信息论的规划器能够通过空间和语义期望信息增益(EIG)分别评估候选发现与目标消歧的价值,从而提供了一种在更广泛探索与更早识别之间进行权衡的显式机制。我们在500个合成目标上评估了六种VLM不确定性引出接口,结果表明相似的识别准确率可能掩盖校准度和虚假置信度方面的显著差异。在观测退化条件下的搜索与识别实验中,基于EIG的规划器在75.0%–92.5%的试验中做出了高置信度决策,而随机搜索仅为20.0%;且在达到置信度后,不同的空间-语义权重设置均能取得相当的识别准确率。提高语义权重可减少不必要的探索和VLM查询,表明对语义不确定性进行显式规划可以在不牺牲决策质量的前提下加速目标解析。这些结果凸显了不确定性表示与不确定性驱动规划在具身VLM系统中的不同作用。
cs.RO / 83 / 2609.20477

Visual Sim-to-Real Learning for Robotic Insertion under Geometric Variations: Application to Rebar Installation

几何变化下机器人插入任务的视觉Sim-to-Real学习:以钢筋安装为例
Sun, Tao, Han, Beining, Yin, Patrick, Xu, Rui, He, Harry, Gupta, Abhishek, Rusinkiewicz, Szymon, Shao, Yi
Abstract
Rebar insertion is among the most repetitive and physically demanding tasks on construction sites, and a contact-rich problem at 1.4 mm clearance. The parts, however, vary at two levels: a nominal design per structural member, and fabrication tolerance around each nominal design. Real-world data therefore has to be re-collected as designs and batches change. We present RebarSim, a visual sim-to-real system trained entirely in simulation. A privileged state-based teacher is trained with reinforcement learning over procedurally generated rebar geometries, then distilled into a multi-view student that maps raw RGB and proprioception directly to actions under extensive domain randomization. The student transfers to the real world zero-shot, seating rebars taken from a real factory production run in 91.3% of real-robot rollouts. Underlying that result, geometry diversity and pretraining both bring benefits. Training across a diverse set of nominal designs rather than one lifts the zero-shot success of both the teacher and the student on unseen designs, and the student policy outperforms a single-design specialist on that specialist's own design. A pretrained student then adapts to a new design with 4--6x fewer distillation samples than one trained from scratch. Visual sim-to-real transfer depends on appearance randomization and the DAgger mixture: removing either one sharply lowers success. Videos, code, and task assets are available at https://rebarsim.github.io.
Chinese Translation
钢筋插入是建筑工地上最重复且体力要求最高的任务之一,也是一个在1.4毫米间隙下的富接触问题。然而,部件在两个层面存在变化:每个结构构件的名义设计,以及围绕每个名义设计的制造公差。因此,随着设计和批次的变化,必须重新收集真实世界数据。我们提出了RebarSim,一个完全在仿真中训练的视觉sim-to-real系统。首先通过强化学习在程序化生成的钢筋几何形状上训练一个基于特权状态信息的教师策略,然后将其蒸馏为一个多视角学生策略,该学生在广泛的域随机化下将原始RGB图像和本体感觉直接映射为动作。学生策略可零样本迁移到真实世界,在91.3%的真实机器人测试中成功插入来自真实工厂生产线的钢筋。在这一结果背后,几何多样性和预训练均带来收益。在多样化的名义设计集合上训练而非单一设计,可提升教师和学生策略在未见设计上的零样本成功率,且学生策略在该专家策略自身的设计上也能超越单一设计的专家。经过预训练的学生策略适应新设计所需的蒸馏样本比从零开始训练的少4至6倍。视觉sim-to-real迁移依赖于外观随机化和DAgger混合训练:移除其中任一项都会显著降低成功率。视频、代码和任务资源可在https://rebarsim.github.io获取。
cs.RO / 84 / 2609.20480

Worst-Case Hidden-Vehicle Trajectory Search in Spatiotemporal Occlusion Regions

时空遮挡区域中最坏情况的隐藏车辆轨迹搜索
Tan, Ruichen, Lei, Zengxiang, Ukkusuri, Satish
Abstract
Occlusion creates fundamental uncertainty in autonomous driving. Existing methods often propagate frame-wise hypotheses or optimize ego behavior against prescribed hidden-agent predictions, leaving the worst history-consistent interaction unexplored. We introduce History-Conditioned Minimax Trajectory Search (HC-MTS), which combines temporal occlusion reasoning with response-aware search. First, HC-MTS constructs finite hidden-state modes, each certified by a backward witness satisfying multi-frame visibility, occupancy, semantic-map support, and class-specific kinematic constraints. It then solves a bilevel minimax problem: an inner finite oracle maximizes the ego driving score over destination attainment and ride comfort, while the outer search selects the legal hidden-vehicle trajectory that minimizes this best-response value. Across eight Waymo Open Motion Dataset scenarios, increasing the visibility-memory horizon from K=1 to K=20 reduces the mean per-scenario vehicle, pedestrian, and total retained hidden-seed counts by 18.12%, 21.67%, and 18.45%, respectively. HC-MTS identifies six avoidable counterexamples, while no legal collision-producing attacker is found in the remaining two scenes within the finite search budget.
Chinese Translation
遮挡给自动驾驶带来了根本性的不确定性。现有方法通常传播逐帧假设,或针对预设的隐藏智能体预测来优化自车行为,而未能探索符合历史的最坏交互情况。我们提出了历史条件极小极大轨迹搜索(History-Conditioned Minimax Trajectory Search,HC-MTS),该方法将时序遮挡推理与响应感知搜索相结合。首先,HC-MTS构建有限个隐藏状态模式,每个模式均由满足多帧可见性、占据、语义地图支持以及类别特定运动学约束的反向见证(backward witness)所验证。随后,该方法求解一个双层极小极大问题:内层有限oracle在到达目的地与乘坐舒适性方面最大化自车驾驶评分,而外层搜索选择使该最优响应值最小化的合法隐藏车辆轨迹。在八个Waymo Open Motion Dataset场景中,将可见性记忆时长从K=1增加到K=20,使每个场景平均保留的车辆、行人及总隐藏种子数分别减少了18.12%、21.67%和18.45%。HC-MTS识别出六个可避免的反例,而在剩余两个场景中,在有限搜索预算内未发现任何能够产生碰撞的合法攻击者。
cs.RO / 85 / 2609.20499

Towards AI-enhanced control: a numerical technique for trajectory smoothing of a parallel robot for pancreatic surgery

迈向AI增强控制:一种用于胰腺外科手术并联机器人轨迹平滑的数值方法
Birlescu, Iosif, Pusca, Alexandru, Gherman, Bogdan, Vaida, Calin, Zima, Ionut, Chablat, Damien, Pisla, Doina
Abstract
The paper presents a numerical approach for the end-effector trajectory smoothing of a parallel robot designed for minimally invasive pancreatic surgery. The approach is tailored for real-time master-slave control architecture and uses a 3D space mouse for command input for velocity control. The trajectory smoothing is achieved by generating S-curves in the end-effector velocity fields, thus controlling the accelerations, which in turn reduces tissue trauma in the minimally invasive procedures. Real-time control is enabled by segmenting the S-curves based on the command inputs from the 3D space mouse. A special case is considered where the acceleration time is constant for all command inputs. Numeric results demonstrate stable transitions (without abrupt changes) in both the end-effector parameter space and in the active joints parameters, thereby validating the proposed approach. Further work aims to test the approach on an experimental model and integrate it into AI-based training modules.
Chinese Translation
本文提出了一种数值方法,用于为微创胰腺外科手术设计的并联机器人末端执行器的轨迹平滑。该方法专为实时主从控制架构而设计,使用3D空间鼠标作为速度控制的指令输入。轨迹平滑通过在末端执行器速度场中生成S形曲线来实现,从而控制加速度,进而减少微创手术中的组织创伤。实时控制通过基于3D空间鼠标的指令输入对S形曲线进行分段来实现。文中考虑了一种特殊情况,即所有指令输入的加速时间均为常数。数值结果表明,末端执行器参数空间和主动关节参数中均实现了平稳过渡(无突变),从而验证了所提出的方法。后续工作旨在将该方法在实验模型上进行测试,并将其集成到基于AI的训练模块中。
cs.RO / 86 / 2609.20540

Integrated Guidance and Control of a Mother-Child UAV-UGV System for Cooperative Missions

面向协同任务的子母式无人机-无人地面车辆系统一体化制导与控制
Sahu, Aashish, Kumar, R. Prasanth
Abstract
Autonomous recovery of a small multirotor onto a hovering multirotor carrier differs from recovery onto ground or shipborne platforms because the recovery surface is itself an actively controlled, thrust-limited aerial vehicle. This paper presents a field-validated autonomy framework for a heterogeneous rover-mothership-child system executing rover supervision, mothership transit, child deployment and sortie, autonomous return, aerial recovery, and synchronized descent. The recovery stack combines jerk-bounded reference generation, disturbance-observer-augmented planar tracking, feasibility-aware vertical control, a discrete-time barrier-based safety filter for relative vertical geometry, and communication-aware carrier-state prediction. The contribution is the coordinated system-level integration of these methods for recovery onto a hovering multirotor and its full-scale outdoor validation. The framework is implemented on a PX4-ROS 2 architecture using RTK-enabled GNSS, IMU, and barometric fusion, with mothership-side 1D lidar used only as an auxiliary near-contact cue. RTK-fixed positioning was maintained throughout testing. Across 20 outdoor cooperative missions, 17 successfully completed deployment, sortie, and recovery, giving an observed mission success rate of 85%. For successful recoveries, mean terminal-alignment time was 6.3 s, mean planar alignment error at acceptance was 0.18 m, maximum terminal planar deviation was 0.32 m within a 0.40 m capture radius, and minimum logged relative vertical separation during coupled descent was 0.41 m. Mothership planar station-keeping RMS error was 0.25 m. The three unsuccessful trials occurred at different mission stages and are analyzed separately. Results demonstrate practical autonomous aerial recovery within the tested outdoor operating envelope.
Chinese Translation
小型多旋翼无人机在悬停多旋翼载机上的自主回收,与在地面或舰船平台上的回收不同,因为回收面本身是一架主动控制且推力受限的飞行器。本文提出了一种经过外场验证的自主系统框架,用于异构的地面机器人-载机-子机系统,执行地面机器人监控、载机转场、子机部署与出击、自主返航、空中回收以及同步下降等任务。回收技术栈结合了加加速度受限的参考轨迹生成、扰动观测器增强的平面跟踪、考虑可行性的垂直控制、用于相对垂直几何关系的基于障碍函数的离散时间安全滤波器,以及通信感知的载机状态预测。本文的贡献在于将这些方法进行系统级协同集成,实现向悬停多旋翼载机的回收,并完成了全尺寸外场验证。该框架在PX4-ROS 2架构上实现,采用RTK增强的GNSS、IMU与气压计融合定位,载机侧的1D激光雷达仅用作近接触阶段的辅助提示。整个测试过程中始终保持RTK固定解定位。在20次外场协同任务中,17次成功完成了部署、出击与回收,观测到的任务成功率为85%。在成功的回收中,平均终端对准时间为6.3秒,验收时平均平面对准误差为0.18米,在0.40米捕获半径内最大终端平面偏差为0.32米,耦合下降期间记录的最小相对垂直间距为0.41米。载机平面定点保持的RMS误差为0.25米。三次失败试验发生在不同的任务阶段,并分别进行了分析。结果表明,在所测试的室外运行包线内,自主空中回收是切实可行的。
cs.RO / 87 / 2609.20558

Learning Slope-Adaptive Whole-Body Locomotion for Humanoid Robots in Roofing Construction

面向屋面施工的人形机器人坡度自适应全身运动学习
Liu, Songyang, Li, Shuai
Abstract
Roofing requires workers to coordinate locomotion, balance, and work-related body motions on pitched surfaces, creating a challenging application for humanoid robots. Directly retargeted human demonstrations, however, may preserve motion appearance while placing the robot's feet or hands incorrectly relative to the roof. This study presents a task-semantic scene-grounded framework for learning roofer-style whole-body motions on a Unitree G1. Human demonstrations are captured using a tracking system and retargeted to the robot, while a metric roof model supplies the spatial reference unavailable from the tracking system. A trajectory-level optimization grounds inferred support contacts and annotated work relations to the roof, and execution-aware reinforcement learning encourages the resulting policy to preserve these relations under dynamic tracking errors. The framework is evaluated through a multi-motion tracking study, a roof-pitch coverage matrix, a five-way nailgun ablation, cross-task experiments on hammering and lateral pushing, and comparisons with pure reinforcement learning and zero-shot teleoperation. Our method enables the robot to satisfy support, work-clearance, and nonpenetration criteria across all evaluated seeds. Across nailgun, hammering, and pushing, it achieves work-clearance errors between 0.256 and 0.531 cm and 3/3 successful evaluations per task. Physical experiments reproduce uphill walking, nailgun, hammering, and bending motions with mean base-frame motion errors below 80 mm. These findings establish scene-grounded human motion learning as a promising basis for construction-oriented humanoid motion primitives.
Chinese Translation
屋面作业要求工人在倾斜表面上协调运动、平衡与工作相关的身体动作,这对人形机器人而言是一项具有挑战性的应用。然而,直接重定向的人类示范可能在保留运动外观的同时,使机器人的脚或手相对于屋顶的位置不正确。本研究提出了一个任务语义、场景锚定的框架,用于在宇树G1机器人上学习屋面工人风格的全身运动。人类示范通过动作捕捉系统采集并重定向到机器人,同时一个精确的屋顶模型提供了动作捕捉系统无法获得的空间参考。轨迹级优化将推断出的支撑接触点和标注的工作关系锚定到屋顶上,而执行感知的强化学习则促使所得策略在动态跟踪误差下保持这些关系。该框架通过多动作跟踪研究、屋面坡度覆盖矩阵、五项钉枪消融实验、锤击与侧向推动的跨任务实验,以及与纯强化学习和零样本遥操作的比较进行了评估。我们的方法使机器人在所有评估种子下均满足支撑、工作间隙和无穿透准则。在钉枪、锤击和推动任务中,其工作间隙误差介于0.256至0.531厘米之间,每项任务均实现3/3的成功评估。物理实验复现了上坡行走、钉枪、锤击和弯腰动作,平均基座坐标系运动误差低于80毫米。这些发现确立了场景锚定的人类运动学习作为面向建筑的人形机器人运动基元的有前景基础。
cs.RO / 88 / 2609.20566

OmniMimic: Dynamics-completed Motion Augmentation for Multi-style Omnidirectional Quadruped Locomotion

OmniMimic:面向多风格全向四足运动的动力学补全运动增强方法
Wu, Sheng, Zhao, Guoqiang, Yang, Zhe, Teng, Fei, Zhou, Zhikun, Yang, Yanlin, Fang, Zheng, Zheng, Hong, Wang, Yaonan, Yang, Kailun
Abstract
Animal demonstrations provide quadruped robots with natural and distinctive gait styles that are difficult to specify through hand-crafted rewards. However, their narrow directional coverage leaves little style-consistent supervision for backward, lateral, and turning commands. We present OmniMimic, a training framework that turns directionally limited animal demonstrations into a single multi-gait policy over target per-axis velocity ranges. OmniMimic first combines temporal reversal, constrained dynamics completion, and sagittal reflection to construct robot-specific kinematic and physical supervision beyond the observed directions. It then expands commands progressively from the demonstrated velocity distribution toward the target per-axis bounds, and uses a shared actor with soft-gated, gait-specialized residual experts to balance reusable locomotion skills with gait-specific corrections. Across four gaits in simulation, OmniMimic reduces mean foot-position RMSE at forward and backward reference velocities by 12.9% and velocity-tracking RMSE on a uniform Cartesian command grid by 63.1%, compared with the matched APEX baseline. The project page is at https://OmniMimic.github.io.
Chinese Translation
动物示范为四足机器人提供了难以通过手工设计奖励函数来指定的自然且独特的步态风格。然而,动物示范的方向覆盖范围狭窄,难以为后退、侧移和转向指令提供风格一致的大量监督信号。我们提出了OmniMimic,这是一个训练框架,能够将方向受限的动物示范转化为一个可在目标各轴速度范围内运行的单一多步态策略。OmniMimic首先结合时间反转、受限动力学补全和矢状面镜像,构建超出观测方向范围的机器人专属运动学与物理监督信号。随后,它将指令从示范速度分布逐步扩展至目标各轴速度边界,并采用带有软门控(soft-gated)步态专用残差专家的共享执行器(actor),在可复用的运动技能与步态特定的修正之间取得平衡。在涵盖四种步态的仿真实验中,与匹配的APEX基线相比,OmniMimic在前向和后向参考速度下的平均足端位置RMSE降低了12.9%,在均匀笛卡尔指令网格上的速度跟踪RMSE降低了63.1%。项目页面见 https://OmniMimic.github.io。
cs.RO / 89 / 2609.20570

Walking on the Slope: Stable Bipedal Gaits with Genetic-Algorithm-Optimized Trajectories

斜坡行走:基于遗传算法优化轨迹的稳定双足步态
Rijal, Madhav
Abstract
This paper presents the kinematic and dynamic modeling, trajectory generation, and stability analysis of an 8-degree-of-freedom (DOF) biped robot walking on flat and inclined terrain. Denavit-Hartenberg (DH) parameters and homogeneous transformations are used to derive the forward kinematics, while closed-form inverse kinematics maps the desired hip and swing-foot Cartesian trajectories, generated with cubic splines, to joint angles. Joint torques are computed using the Newton-Euler iterative algorithm, and dynamic stability is evaluated using the zero moment point (ZMP) criterion. A genetic algorithm (GA) optimizes the hip height, maximum swing-foot lift, and frontal-plane tilt angle by minimizing the work done by the joints subject to a ZMP feasibility penalty. Simulation results in MATLAB show that the nominal 8-DOF model remains ZMP-stable for step completion times down to 0.5 s and for slope inclinations up to 22.5 degrees with the given foot geometry. Beyond these limits, the ZMP leaves the support polygon, and either the foot dimensions or the trajectory parameters must be modified. The results also show that ZMP stability is governed by the mass distribution among the links rather than the total mass of the robot.
Chinese Translation
本文提出了一种8自由度(DOF)双足机器人在平坦和倾斜地形上行走的运动学与动力学建模、轨迹生成及稳定性分析。采用Denavit-Hartenberg(DH)参数和齐次变换推导正向运动学,并利用封闭形式的逆运动学将所需的髋部与摆动腿足部的笛卡尔轨迹(由三次样条生成)映射为关节角度。关节力矩通过牛顿-欧拉(Newton-Euler)迭代算法计算,动态稳定性采用零力矩点(ZMP)准则进行评估。遗传算法(GA)通过最小化关节做功(以ZMP可行性作为惩罚项)来优化髋部高度、摆动腿最大抬脚高度和额状面倾角。MATLAB仿真结果表明,在给定的足部几何尺寸下,标称8自由度模型在步态完成时间低至0.5秒、坡度高达22.5度时仍能保持ZMP稳定。超出这些限制时,ZMP将离开支撑多边形,此时必须修改足部尺寸或轨迹参数。研究结果还表明,ZMP稳定性取决于连杆间的质量分布,而非机器人的总质量。
cs.RO / 90 / 2609.20575

Accelerating Visual Policy Learning with Sampling-Based Model Predictive Control

基于采样的模型预测控制加速视觉策略学习
Liu, Yilang, You, Haoxiang, Wang, Qian, Rakita, Daniel, Abraham, Ian
Abstract
Learning visual policies for locomotion and manipulation requires coordinating contact with the environment and can incur substantial computation and GPU memory costs. First-order policy gradients (FoPG) reduce training cost through differentiable simulation, but local optimization can converge to unintended contact patterns. To address this shortfall, we propose Sampling-Guided Policy Search (SGPS), which couples recurring action-target refinement by sampling-based model-predictive control with first-order policy optimization. Behavior cloning initializes the policy from sampled actions; training then alternates sampling-based refinement with short-horizon FoPG updates under perturbed initial states and randomized dynamics. For visual policy training, we use a decoupled FoPG formulation that excludes rendering from the computation graph, enabling direct learning from depth observations without a state-policy teacher. On a single GPU, SGPS learns policies for locomotion, obstacle traversal, crate pushing, and bimanual carrying on simulated Unitree Go2 and G1 robots. Our experiments further show that refinement improves policy learning beyond initialization and tracking alone. For hardware deployment, the distilled policy transfers zero-shot to a real Go2 and uses onboard depth to autonomously trot, crawl, clear hurdles, and switch between these behaviors.
Chinese Translation
面向运动与操作任务的视觉策略学习需要与环境的接触进行协调,并可能带来巨大的计算开销和GPU显存成本。一阶策略梯度利用可微仿真降低训练成本,但局部优化可能收敛到非预期的接触模式。为解决这一不足,我们提出采样引导策略搜索,它将基于采样的模型预测控制进行的循环动作目标细化与一阶策略优化相结合。行为克隆利用采样动作对策略进行初始化;随后,训练在扰动初始状态和随机化动力学条件下,交替进行基于采样的细化与短时域一阶策略梯度更新。对于视觉策略训练,我们采用一种解耦的一阶策略梯度形式,将渲染排除在计算图之外,从而无需状态-策略教师模型即可直接从深度观测中学习。在单块GPU上,SGPS为仿真中的Unitree Go2和G1机器人学习了对角小跑、越障、推箱和双臂搬运等策略。实验进一步表明,这种细化改进超越了仅靠初始化和跟踪所带来的策略学习效果。在硬件部署方面,经蒸馏的策略可零样本迁移到真实的Go2机器人上,利用机载深度相机自主实现对角小跑、匍匐、跨越障碍以及在这些行为之间切换。
cs.RO / 91 / 2609.20582

V2-STRep: VLM-Grounded Structured Task Representations for Reusable Robot Skills Acquired from Generated Videos

V2-STRep:基于视觉语言模型的从生成视频获取可复用机器人技能的结构化任务表示
Hu, Yexin, Lee, Dongheui
Abstract
Human manipulation videos provide rich motion and interaction cues for acquiring robot skills without robot demonstrations. Video generation models synthesize such demonstrations from an initial scene image and task instruction, avoiding the need to record demonstrations for each task. However, the recovered motion captures only one scene-specific realization, leaving task structure, geometric relations, and constraints implicit. We present V2-STRep, a zero-shot framework that converts generated video motion into reusable robot skills through VLM-grounded structured task representations. The representation specifies motion phases, references, and task-relevant constraints, with targets described by minimal geometric structures: points, point-normals, axes, planes, and full 6D poses. VLM-provided 2D image-space cues are lifted into 3D using RGB-D observations to reconstruct task geometry and candidate grasp poses. Geometry-specific rules transfer motion to new scenes, while task-constrained trajectory optimization couples grasp selection with complete robot motion planning. It preserves task requirements while using remaining rotational freedom to accommodate joint limits. Updating deployment grounding and constraints enables reuse under new compatible instructions without generating another video. Experiments on six real-world manipulation tasks demonstrate improved execution success over baselines, reliable cross-scene transfer of successfully acquired skills, and adaptation to changed deployment instructions.
Chinese Translation
人类操作视频为无需机器人示教即可习得机器人技能提供了丰富的运动与交互线索。视频生成模型可以从初始场景图像和任务指令合成此类演示,从而避免了为每个任务专门录制示教的需要。然而,恢复出的运动仅捕捉了特定场景下的一种实现方式,使任务结构、几何关系和约束保持隐式状态。我们提出了V2-STRep,这是一个零样本框架,通过基于视觉语言模型(VLM)的结构化任务表示,将生成视频中的运动转化为可复用的机器人技能。该表示明确了运动阶段、参考对象和任务相关约束,其目标由最小几何结构描述:点、点-法线、轴、平面以及完整的6D位姿。VLM提供的2D图像空间线索通过RGB-D观测提升至3D空间,以重建任务几何和候选抓取位姿。针对特定几何的规则将运动迁移到新场景,而任务约束的轨迹优化则将抓取选择与完整的机器人运动规划相结合。该方法在保留任务要求的同时,利用剩余的旋转自由度来适应关节限位。通过更新部署时的定位与约束,可以在新的兼容指令下复用技能,而无需重新生成视频。在六个真实世界操作任务上的实验表明,该方法相较于基线方法具有更高的执行成功率、已成功习得技能的可靠跨场景迁移能力,以及对部署指令变化的适应能力。
cs.RO / 92 / 2609.20586

CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding

CoRef-GS:面向多智能体场景理解的协同指代高斯泼溅方法
Zhou, Zhikun, Peng, Kunyu, Yang, Runyi, Cai, Junhao, Wen, Di, Liu, Ruiping, Paudel, Danda Pani, Zhou, Yi, Van Gool, Luc, Yang, Kailun
Abstract
Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such grounding within one agent's observations, cooperative settings require this ability to remain effective after independently reconstructed maps are aligned and fused. In this setting, the referred target or its contextual landmark may come from another agent's observations, while spatial relations must still be interpreted from the querying robot's viewpoint. We formulate this problem as cooperative referring Gaussian grounding over fused maps, which requires geometric alignability, instance-level semantic comparability, and view-conditioned relation reasoning. Existing language-aware Gaussian methods mainly focus on single-map querying, whereas Gaussian registration methods optimize geometric or photometric alignment without preserving language-grounding-oriented semantic compatibility. We propose CoRef-GS, a cooperative referring Gaussian splatting framework. CoRef-GS constructs local open-vocabulary instance-aware Gaussian maps, then aligns partially overlapping maps with a cross-agent alignment module by geometric and semantic consistency, and grounds queries using a view-conditioned mask relation graph. We further introduce CoQuad-Ref, a dual-quadruped benchmark spanning both real-world and simulated indoor scenes. Experiments show that, on simulated scenes, CoRef-GS reduces the rotation error from 2.58{\deg} after coarse initialization to 0.15{\deg} after refinement, and improves real-world referring mIoU over ReferSplat from 52.6% to 68.8%. The established benchmark and source code will be publicly released at https://github.com/ruojiruoli17/CoRef-GS.git.
Chinese Translation
具身机器人的指代场景理解需要在指定视角下将面向对象和关系的语言查询进行定位。局部语义高斯地图可以在单个智能体的观测范围内支持此类定位,但在协同场景中,该能力需在独立重建的地图被对齐和融合后仍然有效。在此场景下,被指代的目标或其上下文地标可能来自其他智能体的观测,而空间关系仍必须基于查询机器人的视角来解释。我们将该问题形式化为融合地图上的协同指代高斯定位问题,其要求具备几何可对齐性、实例级语义可比性以及基于视角的关系推理能力。现有的语言感知高斯方法主要集中于单地图查询,而高斯配准方法仅优化几何或光度对齐,未能保持面向语言定位的语义兼容性。我们提出 CoRef-GS,一种协同指代高斯泼溅框架。CoRef-GS 首先构建局部开放词汇的实例感知高斯地图,然后通过跨智能体对齐模块基于几何和语义一致性对部分重叠的地图进行对齐,并利用基于视角的掩码关系图进行查询定位。我们进一步提出 CoQuad-Ref,一个涵盖真实世界和仿真室内场景的双四足机器人基准数据集。实验表明,在仿真场景中,CoRef-GS 将旋转误差从粗初始化后的 2.58° 降低到精化后的 0.15°,并将真实世界指代 mIoU 相较于 ReferSplat 从 52.6% 提升至 68.8%。所建立的基准数据集和源代码将在 https://github.com/ruojiruoli17/CoRef-GS.git 公开发布。
cs.RO / 93 / 2609.20604

Semantic SLAM in Precision Agriculture using Bayesian Inference

基于贝叶斯推断的精准农业语义SLAM
Beumer, Ruben, Doodeman, Sander, van de Molengraft, René, Antunes, Duarte
Abstract
This paper presents a real-time semantic world modeling framework specialized for precision agriculture using autonomous robots. The framework combines probabilistic mapping of objects and their semantic attributes, updated through Bayesian inference, with a graph-based Simultaneous Localization and Mapping (SLAM) approach implemented using $g^2o$, a general framework for graph optimization. This integration enables accurate mapping and localization without relying solely on GPS. By leveraging semantic information such as plant type, size, and health, the robot can perform tasks while mapping and localizing itself within a field of crops. The proposed framework was validated through Gazebo simulations and physical experiments on an indoor field with artificial plants using Boston Dynamics' robot dog Spot. A YOLOv8n object detection model was trained to extract object and semantic data from depth camera observations. These simulations and experiments demonstrate that the system can successfully perform real-time mapping of up to at least 400 plants.
Chinese Translation
本文提出了一种专用于自主机器人精准农业的实时语义世界建模框架。该框架将物体及其语义属性的概率地图构建(通过贝叶斯推断进行更新)与基于图的同步定位与地图构建(SLAM)方法相结合,后者使用图优化通用框架 $g^2o$ 实现。这种集成使机器人能够在不完全依赖GPS的情况下实现精确的地图构建与定位。通过利用植物类型、大小和健康状况等语义信息,机器人可以在作物田间进行地图构建和自定位的同时执行任务。该框架通过Gazebo仿真实验以及使用波士顿动力公司(Boston Dynamics)机器狗Spot在含人工植物的室内场地进行的物理实验得到了验证。研究训练了一个YOLOv8n目标检测模型,用于从深度相机观测中提取物体和语义数据。仿真和实验结果表明,该系统能够成功实现对至少400株植物的实时地图构建。
cs.RO / 94 / 2609.20605

Bayesian Continuum Robot Dynamics and State Estimation

贝叶斯连续体机器人动力学与状态估计
Ferguson, James M., Hermans, Tucker, Kuntz, Alan
Abstract
Recent factor graph approaches to continuum robot state estimation have been successful for quasi-static applications and spatiotemporal estimation using white-noise kinematic motion priors. However, when inertial effects are significant, these approximations may fail to capture the underlying physics, limiting accuracy during dynamic motions. In contrast, our approach approximates the Cosserat rod dynamics of continuum robots. We write inertia and damping as equivalent applied loads, so that the dynamic balance retains the algebraic form of the static one from prior work with quasi-static robots. Without backbone observations, the framework reduces to a stochastic forward simulation of the robot's motion. Given observations, it jointly refines kinematic and dynamic states and infers external loads, among other states. We validate the approach through simulation and experiments, demonstrating stochastic forward simulation as well as state estimation on tendon-driven continuum robots.
Chinese Translation
近年来基于因子图的连续体机器人状态估计方法在准静态应用以及使用白噪声运动学运动先验的时空估计中取得了成功。然而,当惯性效应显著时,这些近似可能无法捕捉潜在的物理特性,从而限制了动态运动过程中的精度。相比之下,我们的方法对连续体机器人的Cosserat杆动力学进行了近似。我们将惯性和阻尼表示为等效施加载荷,使动力学平衡保持了先前工作中准静态机器人静力学平衡的代数形式。在没有骨干线观测的情况下,该框架退化为对机器人运动的随机前向仿真;在给定观测的情况下,该框架可联合优化运动学与动力学状态,并推断外部载荷等其他状态。我们通过仿真和实验验证了该方法,演示了随机前向仿真以及在肌腱驱动连续体机器人上的状态估计。
cs.RO / 95 / 2609.20615

INSPECT: Learning Robot View Selection from Assistant Use

INSPECT:从助手使用中学习机器人视角选择
Wen, Di, Yang, Kailun, Guo, Wenhao, Shi, Yitian, Zheng, Junwei, Chen, Yufan, Liu, Ruiping, Wei, Jiale, Rayyes, Rania, Peng, Kunyu
Abstract
Robots inspecting an assembly must determine which parts are present and whether they are correctly installed. During egocentric assembly assistance, head motion and workpiece handling reveal evidence for these checks, while spoken state confirmations link observations to procedural outcomes. We introduce INSPECT, which learns robot view preferences from records of a smart-glasses assistant that answers part queries and provides next-step guidance. Presence-Invariant TwinSwap (PI-TwinSwap) calibrates object evidence through paired identity interventions. Claim-indexed supervision separates evidence requirements from camera-reproducible observation changes. Object-centered calibration adapts relative view preferences to robot poses, while clause-level screening checks predicted evidence. The robot selects views using only its current observation and known poses, without candidate images. Evaluation uses annotated assistant-video replay to simulate state feedback, without target-domain view labels for policy training. On images of physical gearbox assemblies, INSPECT achieves the highest view utility among the compared non-oracle policies and raises human-rated full verifiability from 34.8% to 41.7% compared with keeping the current view. On commercial angle-grinder recordings in IMPACT, the transferred relative-view selector increases the correct decision rate from 50.6% to 54.3% with a frozen perception head. The source code is available at https://github.com/Kratos-Wen/INSPECT.
Chinese Translation
检查装配体的机器人必须确定哪些零件存在以及它们是否被正确安装。在以自我为中心的装配辅助过程中,头部运动和工件操作为这些检查提供了证据,而口头的状态确认则将观察结果与程序化结果联系起来。我们提出INSPECT,它从一副智能眼镜助手(该助手回答零件查询并提供下一步指导)的记录中学习机器人视角偏好。存在不变孪生交换(Presence-Invariant TwinSwap, PI-TwinSwap)通过成对的同一性干预来校准物体证据。以声明为索引的监督将证据需求与可通过相机复现的观察变化分离开来。以物体为中心的校准将相对视角偏好适配到机器人位姿,而子句级筛选则检查预测的证据。机器人仅使用其当前观察和已知位姿来选择视角,无需候选图像。评估使用带标注的助手视频回放来模拟状态反馈,且策略训练不依赖目标域的视角标签。在实体齿轮箱装配体的图像上,INSPECT在所比较的非理想(non-oracle)策略中实现了最高的视角效用,与保持当前视角相比,将人工评定完全可验证的比例从34.8%提升至41.7%。在IMPACT数据集中的商用角磨机录制视频上,迁移的相对视角选择器在使用冻结感知头的情况下将正确决策率从50.6%提升至54.3%。源代码可在 https://github.com/Kratos-Wen/INSPECT 获取。
cs.RO / 96 / 2609.20620

A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies

面向AUV故障恢复的仿真平台:探索基于大语言模型的诊断策略
Halba, Khalid, Cooper, Kylie, Bellingham, James G.
Abstract
Autonomous underwater vehicles (AUVs) operating beyond reliable communications must recover from failures without human intervention. We investigate an architecture in which conventional deterministic layered control autonomy manages normal operations, while an invokable large language model (LLM) serves as a diagnostic and recovery planner when onboard anomaly detection identifies performance outside expected limits. Because language models are stochastic, rigorous evaluation requires ensemble testing rather than individual demonstrations. We present a closed-loop simulation architecture that couples real-time C vehicle software with a higher-level orchestration layer for physics-based fault injection, structured prompting, language-model interaction, mission file generation, validation, execution, and LLM-judge scoring. The framework, which we call SPAR (Simulation Platform for AUV Recovery), supports evaluation across fault realizations, prompt structures, reasoning models, and mission conditions. We vary these for a mass-shift fault over 480 SPAR trials, evaluating a frontier model and three off-the-shelf locally deployable LLMs. Model choice dominates diagnosis: the frontier model places the CG-shift mechanism in its top three hypotheses in 85-90% of trials, versus 60-78% for the best local model. Reasoning analysis indicates that local-model success is associated with following the complete diagnostic procedure, whereas weaker models often commit prematurely to elevator failure even though the actuator tracks its command. Diagnosis and operational decision performance do not appear to be coupled in this dataset. The contributions are an architecture extending unanticipated-fault recovery from detection to mitigation and an ensemble methodology for evaluating LLM-assisted mission management on low-power AUVs.
Chinese Translation
自主水下航行器(AUV)在无法保证可靠通信的海域作业时,必须在无人工干预的情况下从故障中恢复。我们研究了一种体系架构:常规的确定性分层控制自主系统负责正常操作,而当机载异常检测发现性能超出预期范围时,可调用大语言模型(LLM)作为诊断与恢复规划器。由于语言模型具有随机性,严格的评估需要集成测试而非单个演示。我们提出了一种闭环仿真架构,将实时C语言车辆软件与上层编排层相耦合,实现基于物理的故障注入、结构化提示、语言模型交互、任务文件生成、验证、执行以及LLM裁判评分。该框架称为SPAR(AUV恢复仿真平台),支持跨故障实现、提示结构、推理模型和任务条件的评估。针对重心偏移故障,我们在480次SPAR试验中改变这些变量,评估了一个前沿模型和三个可本地部署的现成LLM。模型选择对诊断起主导作用:前沿模型在85–90%的试验中将重心偏移机制列入其前三位假设,而最好的本地模型为60–78%。推理分析表明,本地模型的成功与遵循完整的诊断流程相关,而较弱的模型往往过早地断定是升降舵故障,尽管该执行器实际上在跟踪其指令。在该数据集中,诊断性能与操作决策性能之间似乎不存在耦合关系。本文的贡献包括:一种将意外故障恢复从检测扩展到缓解的体系架构,以及一种用于评估低功耗AUV上LLM辅助任务管理的集成评估方法。
cs.RO / 97 / 2609.20624

SmellDiffusion: Diffusion-Based Quadruped Navigation with Olfactory Scene Graphs

SmellDiffusion:基于扩散模型与嗅觉场景图的四足机器人导航
Ogunwoye, Faith, Zhura, Iana, Amjad, Hajira, Kozlov, Timofei, Seyidov, Didar, Plotnikov, Dmitrii, Fedorov, Fedor, Tsetserukou, Dzmitry
Abstract
A robot sent to a named gas leak must preserve gas identity, estimate the source, and navigate to the resulting goal. We present SmellDiffusion, a simulation pipeline that represents species-specific gas zones in an open-vocabulary olfactory scene graph and shares the selected goal between classical and diffusion planners. Its key components are a peak-local geometric gate for selective source correction and diffusion-based, gas-guided trajectory generation. Among 424 unique source-wind configurations in solved flow, 28 have a concentration peak displaced more than 0.5m from the source. A source-independent geometric gate, calibrated only on the training split and evaluated at the observed peak, detects 9 of 10 held-out displacements at 0.64 precision. Gating a precomputed forward-matching correction reduces mean error on the displaced cases from 1.468m to 0.592m (60%), using matching for only 14/204 cases. All-case mean error falls from 0.205m to 0.180m. All planners receive the same scene-graph source estimate as their goal. In a controlled comparison, best-of-ten diffusion achieves mean gas exposure comparable to gas-guided A* (0.0476 versus 0.0455). A single diffusion proposal takes 41.7ms, compared with 72.3ms for gas-guided A*, although best-of-ten sequential sampling increases total runtime. Plain A* also reaches the same goal and remains the fastest and shortest-path method. Six matched Gazebo runs give mean robot-to-source errors of 0.39m for A* and 0.31m for diffusion.
Chinese Translation
被派往指定气体泄漏点的机器人必须保持对气体种类的识别、估计泄漏源位置,并导航至由此确定的目标点。我们提出了 SmellDiffusion,一个仿真流水线,它在开放词汇的嗅觉场景图中表示特定气体种类的气体区域,并在经典规划器与扩散规划器之间共享所选目标点。其关键组件包括一个用于选择性源位置校正的峰值局部几何门控(peak-local geometric gate),以及基于扩散模型的气体引导轨迹生成。在已求解流场的424个独立“源-风”配置中,有28个配置的浓度峰值偏离源位置超过0.5米。一个与源位置无关的几何门控(仅在训练集上校准,并在观测峰值处评估)在精度为0.64的情况下检测出10个留出集位移中的9个。对预计算的前向匹配校正进行门控,使位移情况下的平均误差从1.468米降至0.592米(降低60%),且仅在204个案例中的14个使用了匹配。所有案例的平均误差从0.205米降至0.180米。所有规划器均接收相同的场景图源估计作为其目标点。在受控比较中,十选一(best-of-ten)的扩散方法实现了与气体引导A*相当的气体暴露均值(0.0476 对 0.0455)。单次扩散提议耗时41.7毫秒,而气体引导A*耗时72.3毫秒,尽管十次顺序采样会增加总运行时间。普通A*同样能到达相同目标,且仍然是最快和路径最短的方法。六组匹配的 Gazebo 运行结果显示,机器人到源位置的平均误差为:A*为0.39米,扩散方法为0.31米。
cs.RO / 98 / 2609.20629

RTK-Vision PPO for Autonomous Micro UAV Recovery on an Airborne Carrier

基于RTK视觉PPO的机载母机自主微型无人机回收
Sahu, Aashish, Kumar, R Prasanth
Abstract
Autonomous recovery of a micro unmanned aerial vehicle (UAV) onto a moving airborne carrier enables reusable deploy-mission-recover operation, but couples long-range rendezvous, close-range perception, carrier motion, aerodynamic interaction, and a discontinuous contact event. This paper presents an RTK-vision-guided reinforcement-learning framework in which a child UAV is physically transported by a larger carrier, takes off from the carrier while airborne, executes an independent sortie, returns to the carrier's current position, redocks, and subsequently descends with the carrier. Both vehicles carry RTK-GNSS, and the carrier continuously shares its navigation state with the child. Near the recovery deck, RTK remains active while a downward-facing camera with a fiducial marker detector provides marker-relative alignment cues. A proximal policy optimization (PPO) policy governing the terminal recovery phase is trained in a physics-based MuJoCo simulation environment with explicit sensor noise models, an aerodynamic disturbance surrogate, and marker-latency randomization, then transferred to hardware. PX4 retains low-level stabilization, and a deterministic safety gate authorizes descent independently of the learned policy. The PPO checkpoint achieves 99.55% success over 2,000 held-out randomized terminal episodes, compared with 78.4% for a tuned PD baseline under identical conditions, with a median planar terminal error of 6.62 cm. Across 14 outdoor trials, the full mission succeeds in 13 trials (92.9%), spanning both near-region recovery and recovery after the carrier translates away from the release point. The results demonstrate a complete autonomous aerial deployment-and-recovery cycle rather than an isolated landing maneuver, establishing a practical basis for reusable carrier-child operation in inspection, surveillance, and mobile-logistics applications.
Chinese Translation
微型无人机(UAV)在移动机载母机上的自主回收可以实现可复用的部署-任务-回收作业模式,但同时涉及远距离会合、近距离感知、母机运动、气动干扰以及非连续接触事件等多个耦合难题。本文提出了一种RTK视觉引导的强化学习框架:子无人机由较大的母机物理搭载,在空中从母机上起飞,执行独立 sortie(出动任务),返回母机当前位置,重新对接,并随后随母机一起下降。两架无人机均搭载RTK-GNSS,母机持续向子机共享其导航状态。在回收甲板附近,RTK保持有效,同时一个朝下的相机配合基准标记检测器提供相对标记的对准信息。控制终端回收阶段的近端策略优化(PPO)策略在基于物理的MuJoCo仿真环境中训练,其中包含显式传感器噪声模型、气动扰动代理模型以及标记延迟随机化,随后迁移至硬件平台。PX4保留底层稳定控制,一个确定性安全门控独立于学习策略授权下降动作。在2000个留出的随机化终端回合中,PPO策略达到99.55%的成功率,而相同条件下经过调优的PD基线仅为78.4%,平面终端误差中位数为6.62厘米。在14次室外试验中,完整任务在13次试验中成功(92.9%),涵盖近区域回收以及母机移离释放点后的回收场景。研究结果展示了一个完整的自主空中部署-回收循环,而非孤立的着陆机动,为巡检、监视和移动物流应用中可复用的母机-子机作业奠定了实用基础。
cs.RO / 99 / 2609.20646

TraceFlow: Guiding Frozen Flow-Matching Robot Policies with Success and Failure Traces

TraceFlow:利用成功与失败轨迹引导冻结的流匹配机器人策略
Zhang, Jiaxuan, Liu, Ruizhe, Zhang, Yu, Yang, Yanchao
Abstract
A vision-language-action (VLA) policy with a flow-matching action expert generates each action chunk (a short command sequence) by integrating a learned velocity field; once its weights are fixed, the success or failure of an earlier rollout cannot change the chunk generated now. Concurrent test-time methods give a frozen policy such an input from retrieved successes, a learned critic, a verifier, or a dynamics model, but none uses the robot's own failed rollouts as negative evidence with nothing but a terminal outcome bit. We introduce TraceFlow, a progress-aligned guidance field that turns the action densities of retrieved successful and failed rollouts into a bounded correction to a frozen flow-matching action expert, using one terminal outcome bit per rollout and no other label. Its TraceBank stores traces, time-ordered state-action records with a terminal label, starts from the target-task training traces, and later admits the deployed robot's own rollouts. On an ordered real-robot packing task the base completes 21 of 50 trials in order, TraceFlow 39, and one stacking round without any weight update 47, with wrong-sequence episodes falling from 20 to 0. In simulation the gain is selective: with per-suite selected settings, TraceFlow raises RoboMemArena Sequence from 78.92\% to 91.50\% task success and Transferring from 54.41\% to 62.00\% at stacking round 2, leaves the 26-task aggregate unchanged, lowers Counting and Occlusion by 1.12 and 1.42 points, and changes LIBERO-Plus (Long) by +1.27 points (p = 0.0733). Stacking gains are finite, every branch peaking before round ten, and the bank's success-to-failure ratio predicts no retrieval allocation.
Chinese Translation
具有流匹配动作专家的视觉-语言-动作(VLA)策略通过积分学习到的速度场来生成每个动作块(一段短指令序列);一旦其权重固定,早期执行的成功或失败便无法改变当前生成的动作块。现有的测试时方法为冻结策略提供来自检索到的成功案例、学习的评论者、验证器或动力学模型的输入,但没有任何方法仅凭一个终止结果比特,将机器人自身失败的执行轨迹用作负面证据。我们提出TraceFlow,一个与进度对齐的引导场,它将检索到的成功与失败执行轨迹的动作密度转化为对冻结流匹配动作专家的有界修正,每条轨迹仅需一个终止结果比特,无需其他标签。其TraceBank存储轨迹(带终止标签的按时间排序的状态-动作记录),初始为目标任务的训练轨迹,随后接纳部署机器人自身的执行轨迹。在一个有序的真实机器人装箱任务中,基线策略在50次试验中按序完成21次,TraceFlow完成39次,而仅一轮堆叠(无需任何权重更新)即完成47次,且错误序列的回合数从20降至0。在仿真中,增益具有选择性:在按套件选定的设置下,TraceFlow将RoboMemArena Sequence的任务成功率从78.92%提升至91.50%,在第2轮堆叠时将Transferring从54.41%提升至62.00%,26项任务的总体成绩保持不变,Counting和Occlusion分别下降1.12和1.42分,LIBERO-Plus (Long)变化为+1.27分(p = 0.0733)。堆叠带来的增益是有限的,每个分支都在第十轮之前达到峰值,且库中成功与失败轨迹之比对检索分配没有预测作用。
cs.RO / 100 / 2609.20648

SkipVLA: Skipping VLA Steps with Classical Planning for Fast Robot Manipulation

SkipVLA:利用经典规划跳过VLA步骤以实现快速机器人操作
Agrawal, Kaivalya, Rahman, Md Ashiqur, Yeh, Raymond A., Kingston, Zachary
Abstract
Vision-Language-Action (VLA) models are a class of generalist robot policies that map camera images and language instructions directly to robot actions. While promising, these models remain slow at test time, particularly for long-horizon tasks that require many queries to the policy. Recent efforts reduce VLA latency by distilling smaller models, overlapping asynchronous action chunks, or pairing the VLA with a fast low-level policy, but still run a learned policy for the entire task. In contrast to VLA, classical motion planners quickly find collision-free motions, but require an explicit goal and have no semantic understanding of the task. In this work, we present SkipVLA, a hybrid policy that combines a pretrained VLA with a classical motion planner, using the planner for free-space motion and querying the VLA only for contact-rich skills such as grasping and placing. SkipVLA reuses the frozen vision-language backbone of the VLA to predict a target pose for each planned motion, and learns this predictor without additional demonstrations introduced into the system by using what was already learnt by the large VLA. We evaluate SkipVLA with three VLAs on 13 LIBERO tasks in simulation and three pick-and-place tasks on a physical 6-DoF YAM arm, demonstrating up to 2.5x faster task completion and significantly lower energy consumption while achieving the same task success rate.
Chinese Translation
视觉-语言-动作(VLA)模型是一类将相机图像和语言指令直接映射为机器人动作的通用机器人策略。尽管前景可观,这类模型在测试时仍然较慢,尤其是对于需要多次查询策略的长时程任务。近期的研究通过蒸馏更小的模型、重叠异步动作块或将VLA与快速的底层策略配对来降低VLA延迟,但仍需在整个任务中运行学习到的策略。与VLA不同,经典运动规划器能够快速找到无碰撞运动,但需要显式的目标,且不具备任务的语义理解能力。在本工作中,我们提出了SkipVLA,一种将预训练VLA与经典运动规划器相结合的混合策略:利用规划器完成自由空间运动,仅在抓取和放置等需要接触丰富的技能时查询VLA。SkipVLA复用VLA冻结的视觉-语言骨干网络来预测每次规划运动的目标位姿,并利用大型VLA已经学到的内容来训练该预测器,无需引入额外的演示数据。我们在仿真的13个LIBERO任务上使用三个VLA评估了SkipVLA,并在物理6自由度YAM机械臂上评估了三个抓放任务,结果显示任务完成速度最高提升2.5倍,能耗显著降低,同时保持相同的任务成功率。
cs.RO / 101 / 2609.20649

DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation

DexTouch-WM:从人类触觉中学习动作条件化的触觉世界模型以实现灵巧机器人操作
Qin, Yan, Chen, Yue, Lin, Wenwei, Liu, Shujia, Lyu, Chuqiao, Su, Kailun, Yu, Chenze, Luo, Ping, Ding, Wenbo, Chen, Tianxing, Xu, Renjing
Abstract
Learning predictive models of contact-rich dexterous manipulation requires dense tactile interaction, but such data are costly to scale on real robots and remain tied to embodiment-specific sensors. We introduce DexTouch-WM, an action-conditioned world model that learns from scalable human touch to jointly predict future RGB observations and bilateral tactile dynamics. Our insight is that human and robot manipulation share transferable contact dynamics when their tactile observations and action spaces are made compatible. We deploy flexible piezoresistive arrays with a shared sensing layout on both human and dexterous robot hands, and retarget human motion into the robot action space so that human interaction can supervise the same dynamics model used for real-robot prediction. DexTouch-WM couples a pretrained video expert with a lightweight tactile expert using anatomy-aware tactile tokens and aligned action conditioning. In human-to-robot scaling experiments, we keep five hours of real-robot supervision fixed while increasing human interaction from 0 to 100 hours, and observe substantial improvements in held-out robot-domain visual, geometric, and contact prediction despite disjoint human and robot task sets. Beyond prediction, we evaluate the world models as surrogate environments for policy evaluation and as generators of synthetic trajectories for real-robot policy learning, showing that scalable human interaction provides a complementary data axis for learning dexterous robot world models.
Chinese Translation
学习接触密集型灵巧操作的预测模型需要大量的触觉交互数据,但此类数据在真实机器人上难以低成本规模化采集,且往往绑定于特定本体的传感器。我们提出了DexTouch-WM,一个动作条件化的世界模型,它从可规模化采集的人类触觉中学习,联合预测未来的RGB观测和双侧触觉动力学。我们的核心洞察是:当触觉观测与动作空间被设计为兼容时,人类操作与机器人操作共享可迁移的接触动力学。我们在人手与灵巧机器人手上部署了具有相同传感布局的柔性压阻式触觉阵列,并将人类运动重定向到机器人动作空间,从而使人类交互数据能够监督与真实机器人预测所使用的同一个动力学模型。DexTouch-WM通过解剖结构感知的触觉token和动作条件对齐,将预训练的视频专家模型与轻量级触觉专家模型耦合在一起。在人到机器人的规模化实验中,我们保持5小时真实机器人监督数据不变,将人类交互数据从0小时增加到100小时,尽管人类与机器人的任务集互不重叠,仍观察到模型在机器人域留出集上的视觉、几何和接触预测能力显著提升。除了预测之外,我们还将该世界模型作为策略评估的替代环境和用于真实机器人策略学习的合成轨迹生成器进行评估,结果表明可规模化的人类交互数据为学习灵巧机器人世界模型提供了一个互补的数据维度。
cs.RO / 102 / 2609.20659

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

HIL-UMI:将视觉-语言-动作模型的人类在环后训练引入通用操作接口
Han, Zimu, Zeng, Yiming, Zhang, Jiyao, Zhao, Zihao, Wang, Yuanfei, Jin, Yixiang, Li, Shiqi, Chen, Shuangben, Huang, Wei, Li, Ruodai, Shen, Hui, Dong, Hao
Abstract
Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limited coverage of out-of-distribution states, and standard imitation objectives do not distinguish progressing behavior from less useful data. Interactive post-training can address these limitations, but typically requires repeated policy execution and human intervention on a physical robot. We introduce HIL-UMI, a policy-guided Universal Manipulation Interface (UMI) framework for robot-free human-in-the-loop VLA post-training. During handheld UMI demonstrations, HIL-UMI queries the current policy on the same observation stream without executing its predictions. The Energy Score compares the human action trajectory with policy inference and triggers collection when their discrepancy indicates an out-of-distribution region. In a separate feedback loop, low online advantage predictions identify essential segments for refining a progress-based advantage estimator. The updated estimator then guides advantage-conditioned behavioral cloning using a balanced mixture of base demonstrations and new policy data. This design preserves the iterative and policy-aware nature of human-in-the-loop learning while decoupling data collection from robot deployment. Experiments on four real-world tasks spanning long-horizon and precise manipulation show that HIL-UMI achieves consistent improvement over SFT and benefits from both targeted collection and advantage refinement. Moreover, HIL-UMI outperforms HG-DAgger on Clean Up Table with lower per-frame collection time, suggesting a scalable path for VLA post-training across operators and locations.
Chinese Translation
大规模视觉-语言-动作(VLA)模型为机器人操作提供了强大的先验,然而将其适配到特定部署场景仍然具有挑战性。在任务特定演示上进行监督微调(SFT)是迈向部署的一步,但面临两个持久的局限:静态数据对分布外状态的覆盖有限,且标准的模仿学习目标无法区分有进展的行为与较无用的数据。交互式后训练可以解决这些局限,但通常需要在物理机器人上反复执行策略并进行人工干预。我们提出HIL-UMI,这是一个策略引导的通用操作接口(Universal Manipulation Interface, UMI)框架,用于无机器人的人类在环VLA后训练。在手持UMI演示过程中,HIL-UMI在相同的观测流上查询当前策略,而不执行其预测。能量分数(Energy Score)将人类动作轨迹与策略推理进行比较,当二者差异表明处于分布外区域时触发数据收集。在另一个单独的反馈回路中,低在线优势预测用于识别关键片段,以改进基于进度的优势估计器。更新后的估计器随后引导基于优势条件化的行为克隆,使用基础演示与新策略数据的均衡混合。该设计在将数据收集与机器人部署解耦的同时,保留了人类在环学习的迭代性和策略感知特性。在涵盖长时程和精细操作的四项真实世界任务上的实验表明,HIL-UMI相比SFT取得了一致的提升,并从针对性收集和优势改进中双双受益。此外,HIL-UMI在Clean Up Table任务上以更低的每帧收集时间优于HG-DAgger,为跨操作者和跨地点的VLA后训练提供了一条可扩展的路径。
cs.RO / 103 / 2609.20669

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

无需显式轨迹的三维扩散策略前瞻学习能力
Zhang, Zhongbo, Zhang, Zaibin, Wang, Yifan, Yan, Changbo, Wang, Lijun, Lu, Huchuan
Abstract
3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasible now, but also anticipating where the interaction is heading. Existing policies largely leave such foresight to emerge implicitly from action learning. We introduce Movement Trend Guidance, a simple but effective way to provide this foresight without introducing an explicit plan. From a short observation history, the policy learns a compact latent representation of interaction evolution. During training, sparse future gripper states supervise this representation; at inference, only the latent is retained as future-oriented conditioning alongside the current observation. The latent provides global conditioning for action generation, while an additional gated FiLM branch is used only at the UNet bottleneck. Despite adding only 3.52% more parameters to DP3, our method preserves the original dense-action and receding-horizon formulation and consistently improves upon DP3 across RoboTwin2.0, LIBERO-40, and DexArt. It reaches 62.8% vs. 56.1% in 50-task RoboTwin2.0 mixed training, 71.93% vs. 37.08% on LIBERO-40, and 72.0% vs. 49.0% on five real-robot tasks. These results show that a diffusion policy can benefit substantially from knowing where an interaction is heading, without being told exactly where to move.
Chinese Translation
3D扩散策略擅长从当前观测中生成具有几何依据的动作,但成功的操作不仅需要知道当前哪种运动是可行的,还需要预判交互的发展方向。现有策略大多将这种前瞻能力留给动作学习过程隐式地涌现。我们提出了运动趋势引导,这是一种简单而有效的方法,能够在不引入显式规划的情况下提供这种前瞻能力。策略从短时观测历史中学习一个紧凑的潜在表示,用以刻画交互的演化。在训练阶段,稀疏的未来夹爪状态对该表示进行监督;在推理阶段,仅保留该潜在表示作为面向未来的条件,与当前观测共同使用。该潜在表示为动作生成提供全局条件,而附加的门控FiLM分支仅在UNet瓶颈层使用。尽管仅为DP3增加了3.52%的参数量,我们的方法保留了原有的密集动作和滚动时域框架,并在RoboTwin2.0、LIBERO-40和DexArt上持续优于DP3。在包含50个任务的RoboTwin2.0混合训练中,该方法达到62.8%(对比DP3的56.1%);在LIBERO-40上达到71.93%(对比37.08%);在五项真实机器人任务上达到72.0%(对比49.0%)。这些结果表明,扩散策略可以在无需被告知确切移动位置的情况下,通过了解交互的发展方向而获得显著收益。
cs.RO / 104 / 2609.20670

MAGNETAR: Multipath-Guided Spatial Posteriors for Transmitter Pose Inference in the Upper Mid-Band

MAGNETAR:基于多径引导的空间后验分布推断上部中频段发射机位姿
Lei, Haozhe, Chen, Ruibin, Jiang, Yuhan, Rasteh, Ali, Dhananjay, Aditya, Rangan, Sundeep
Abstract
Robots that localize a radio transmitter need more than a point estimate: in cluttered rooms, one measurement is often consistent with several transmitter locations and, because upper-mid-band antennas are directional, several headings. We present MAGNETAR, which infers a joint posterior over planar transmitter position and heading from a single asynchronous radio-frequency (RF) multipath snapshot, represented by angle-of-arrival and signal-to-noise-ratio estimates, given the room layout and receiver pose. Among our five neural scorers, MAGNETAR adopts a shared 2D U-Net conditioned on each candidate heading, jointly normalizing scores over a discretized position-heading grid. Training uses real-to-sim-calibrated 10 GHz simulations and a small measured subset. Grid-based joint posteriors outperform parametric ones on held-out simulations, the heading-conditioned scorer transfers best to robotic measurements, and fusing joint posteriors improves on fusing position-only marginals.
Chinese Translation
对无线电发射机进行定位的机器人不仅需要点估计:在杂乱的房间中,单次测量往往与多个发射机位置相容,而且由于上部中频段(upper-mid-band)天线具有方向性,还可能与多个朝向相容。我们提出了MAGNETAR,它在给定房间布局和接收机位姿的条件下,仅利用单次异步射频(RF)多径快照(由到达角和信噪比估计表示),推断发射机平面位置与朝向的联合后验分布。在我们提出的五种神经评分器中,MAGNETAR采用了以每个候选朝向为条件的共享2D U-Net,并在离散化的位置-朝向网格上对评分进行联合归一化。训练使用了经过实况到仿真校准的10 GHz仿真数据以及少量实测数据子集。在留出仿真数据上,基于网格的联合后验优于参数化后验;以朝向为条件的评分器在机器人实测数据上的迁移效果最好;融合联合后验的效果优于仅融合位置边缘分布。
cs.RO / 105 / 2609.20680

Towards Scaling Marine Perception with Synthetic Data

基于合成数据实现海洋感知的可扩展性
Ma, Haoyu, Bagoren, Onur, Sheppard, Anja, Fandi, Elias, Edukulla, Ashrith, Aslan, Tanner, Sieh, Natasha, Song, Jingyu, Skinner, Katherine A.
Abstract
Scalable machine learning in challenging underwater environments is strongly limited by the lack of labeled real-world training data. This data is often expensive and laborious to gather, making large-scale real-world data challenging to gather and curate. However, simulated data can help close the gap, enabling many learning-based tasks for underwater perception. In this work, we extend OceanSim, an IsaacSim-based underwater perception simulator, with a Synthetic Data Generation (SDG) pipeline for training models to be used in underwater scenarios. The proposed pipeline enables users to generate large, automatically labeled, photorealistic datasets with configurable scene appearance, structure, and sensor settings. We evaluate the pipeline on a real-world sea urchin detection task and study how different forms of synthetic scene variation affect sim-to-real performance. Based on these experiments, we discuss findings on our results, main limitations of the current pipeline and identify future directions for improving underwater rendering fidelity, scene diversity, and the evaluation of sim-to-real generalization. The open-source code can be found at https://github.com/umfieldrobotics/OceanSim.
Chinese Translation
在具有挑战性的水下环境中,可扩展的机器学习受到真实世界标注训练数据匮乏的严重限制。此类数据的采集往往成本高昂且费时费力,使得大规模真实世界数据的收集与整理极具挑战性。然而,仿真数据可以帮助弥合这一差距,为水下感知中的众多学习任务提供支持。在本工作中,我们扩展了OceanSim——一个基于IsaacSim的水下感知仿真器,构建了一个用于训练水下场景模型的合成数据生成(Synthetic Data Generation, SDG)流水线。所提出的流水线使用户能够生成大规模、自动标注且具有照片级真实感的数据集,并支持对场景外观、结构和传感器设置进行可配置的调整。我们在一个真实世界海胆检测任务上对该流水线进行了评估,并研究了不同形式的合成场景变化对仿真到现实(sim-to-real)迁移性能的影响。基于这些实验,我们讨论了所得结果的发现、当前流水线的主要局限性,并指出了在提升水下渲染真实感、场景多样性以及仿真到现实泛化能力评估方面的未来研究方向。开源代码可在 https://github.com/umfieldrobotics/OceanSim 获取。
cs.RO / 106 / 2609.20691

Custom PX4 firmware for autonomous hybrid aerial-marine missions

面向自主空海混合任务的PX4定制固件
Capuozzo, Andrea, Ruggiero, Fabio, Lippiello, Vincenzo
Abstract
Mapping and monitoring aquatic environments can benefit from hybrid aerial-amphibious drones able to combine flight and water-surface navigation within the same mission. This paper presents a PX4 firmware extension for such platforms, introducing manual and autonomous marine navigation modes integrated with the standard PX4 mission pipeline and QGroundControl interface. The proposed framework preserves existing flight functionalities and safety mechanisms while enabling unified planning and execution of hybrid aerial-marine missions with differentiated aerial and marine waypoints. Simulated case studies validate the implementation and demonstrate stable surface navigation under calm and wavy conditions.
Chinese Translation
水生环境的测绘与监测可以受益于能够在同一任务中结合飞行与水面航行的空-两栖混合无人机。本文针对此类平台提出了一种PX4固件扩展,引入了与标准PX4任务流程及QGroundControl接口集成的手动和自主海洋导航模式。所提出的框架在保留现有飞行功能与安全机制的同时,实现了具有差异化空中与海洋航点的空海混合任务的统一规划与执行。仿真案例研究验证了该实现,并证明了在平静和有波浪条件下水面导航的稳定性。
cs.RO / 107 / 2609.20694

HOPHY: A Hierarchical Hypergraph Representation for Off-Road Path and Mission Planning

HOPHY:一种用于越野路径与任务规划的分层超图表示
Meshram, Pranay, Adhivarahan, Charuvahan, Poddar, Prithvi, Esfahani, Ehsan Tarkesh, Wang, Chen, Chowdhury, Souma, Dantu, Karthik
Abstract
Mission-level autonomy for disaster response, search and rescue, and tactical UGV operations requires repeated path and mission planning as terrain conditions, agent types, and objectives change. Pixel-grid search is costly for repeated kilometer-scale queries, while semantic abstractions must maintain valid costs and connectivity as conditions change. We present HOPHY (Hierarchical Off-Road Planning using Hypergraphs), a reusable hierarchical terrain representation that organizes map-scale terrain into geometrically connected semantic regions (GSNodes), connectivity-preserving critical regions (Coarse Regions), and typed hyperedges for terrain, agent, and weather context. Hyperedge intersections select affected regions and incident edges for state updates without rebuilding the hierarchy. Across real off-road maps spanning kilometer-scale areas, HOPHY achieves 100% planning success and less than 0.01% median cost deviation from the oracle (pixel A*), with substantially lower query and replanning latency than the evaluated pixel and abstraction baselines. Applied to a multi-robot task-allocation (MRTA) problem, these gains reduce total computation by 79x over pixel A* and 7.2x over the fastest abstraction baseline, with mission makespan comparable to pixel A*. Finally, we demonstrate HOPHY on a physical Clearpath Jackal that successfully executes a 1.5-km, eight-task mission across mixed-surface outdoor terrain and a blockage-triggered replanned route.
Chinese Translation
面向灾难响应、搜索救援和战术无人地面车辆(UGV)行动的任务级自主性要求在地形条件、智能体类型和目标发生变化时反复进行路径与任务规划。像素网格搜索在重复的公里级查询中代价高昂,而语义抽象则必须在条件变化时保持有效的代价和连通性。我们提出了HOPHY(基于超图的分层越野规划),这是一种可复用的分层地形表示方法,它将地图尺度的地形组织为几何连通的语义区域(GSNodes)、保持连通性的关键区域(Coarse Regions),以及用于地形、智能体和天气上下文的带类型超边。超边交集用于选择受影响的区域和相关联的边以进行状态更新,而无需重建整个层次结构。在覆盖公里级区域的真实越野地图上,HOPHY实现了100%的规划成功率,且相对最优解(像素A*)的中位代价偏差小于0.01%,同时查询和重规划延迟显著低于所评估的像素与抽象基线方法。应用于多机器人任务分配(MRTA)问题时,这些优势使总计算量相比像素A*降低79倍,相比最快的抽象基线降低7.2倍,而任务完成时间与像素A*相当。最后,我们在一台物理Clearpath Jackal机器人上演示了HOPHY,成功执行了跨越混合路面户外地形的1.5公里、八项任务,并完成了由障碍触发的重规划路线。
cs.RO / 108 / 2609.20709

MoWAM: Explicit Future Motion Prediction for Efficient World Action Models

MoWAM:面向高效世界动作模型的显式未来运动预测
Wang, Jiayu, Zhu, Bin, Yu, Yue, Chen, Jingjing
Abstract
World Action Models (WAMs) improve robot policy learning by incorporating future dynamics, yet explicitly generating future videos at inference introduces substantial computational overhead. Removing future generation improves efficiency, but leaves future dynamics only implicitly encoded in observation features, which can limit robustness under distribution shifts. We propose MoWAM, an efficient WAM that replaces future video generation with explicit future motion prediction. Instead of reconstructing the complete future scene, MoWAM models structured robot motion as a compact abstraction of the future, capturing how the robot is expected to evolve under the current scene and interaction constraints. A Mixture-of-Transformer architecture learns future visual dynamics during training while jointly predicting motion and action, allowing video generation to be removed entirely at inference while retaining an explicit representation of the future. The compact motion representation further enables efficient inference-time scaling by sampling multiple candidates of motion and action pairs and selecting among them with a motion-aware task-progress verifier. Experiments on LIBERO, LIBERO-Plus, and real-world manipulation tasks demonstrate that MoWAM achieves strong in-distribution performance, improved out-of-distribution robustness, and higher average real-world success than representative WAM baselines. In addition, performance improves as more candidates are explored, demonstrating that explicit future motion provides an effective and efficient basis for inference-time scaling.
Chinese Translation
世界动作模型(World Action Models, WAMs)通过引入未来动态来改进机器人策略学习,但在推理阶段显式生成未来视频会带来显著的计算开销。去除未来生成虽然能提升效率,但未来动态仅被隐式编码在观测特征中,这可能限制模型在分布偏移下的鲁棒性。我们提出 MoWAM,一种用显式未来运动预测替代未来视频生成的高效世界动作模型。MoWAM 不再重构完整的未来场景,而是将结构化的机器人运动建模为未来的紧凑抽象,以刻画机器人在当前场景与交互约束下预期如何演化。我们采用混合Transformer(Mixture-of-Transformer)架构,在训练过程中学习未来视觉动态,并同时预测运动与动作,使得推理阶段可以完全去除视频生成,同时保留对未来的显式表示。这种紧凑的运动表示还支持高效的推理时扩展:通过采样多个运动与动作候选对,并利用运动感知的任务进度验证器(motion-aware task-progress verifier)从中进行选择。在 LIBERO、LIBERO-Plus 以及真实世界操作任务上的实验表明,MoWAM 在分布内具有强劲性能,在分布外具有更好的鲁棒性,且真实世界的平均成功率高于代表性的 WAM 基线。此外,随着探索更多候选,性能持续提升,证明显式未来运动为推理时扩展提供了有效且高效的基础。
cs.RO / 109 / 2609.20731

Underwater Visual Target Tracking with Target-Specific Depth Estimation and Adaptive Model-Fusion Predictive Control

基于目标特定深度估计与自适应模型融合预测控制的水下视觉目标跟踪
Zhou, Yuheng, Cheng, Haiyang, Feng, Yanqi, Fong, Pangkit, Lee, Mei Xuan, Gee, Marcus, Fang, Chongrong, He, Jianping
Abstract
Vision-based underwater target tracking is challenged by unreliable depth measurements and unknown target motion. This paper proposes a stereo visual-servoing framework for an autonomous underwater vehicle (AUV). For perception, the framework derives a stable 3D relative state from stereo images through target-specific depth extraction and Kalman filtering. It constructs a target-depth mask from color, disparity, and temporal cues to select reliable target pixels, and then filters the resulting depth measurement and detected image center separately. For control, the framework decouples yaw regulation from translational control, avoiding computationally expensive coupled multi-DOF optimization and enabling real-time translational MPC. The translational controller employs adaptive model-fusion predictive control, combining constant-velocity and zero-velocity target models to accommodate different target-motion patterns. It updates the model weights using historical prediction errors and computes translational commands subject to actuation, following-distance, and field-of-view constraints. Through simulations and real-world experiments, we validate the effectiveness of the proposed framework and show it has better performance than existing frameworks.
Chinese Translation
基于视觉的水下目标跟踪面临深度测量不可靠和目标运动未知等挑战。本文提出了一种面向自主水下航行器(AUV)的双目视觉伺服框架。在感知方面,该框架通过目标特定的深度提取和卡尔曼滤波,从双目图像中获得稳定的三维相对状态。它利用颜色、视差和时间线索构建目标深度掩膜,以选择可靠的目标像素,然后分别对所得的深度测量值和检测到的图像中心进行滤波。在控制方面,该框架将偏航调节与平移控制解耦,避免了计算开销高昂的多自由度耦合优化,并实现了实时的平移模型预测控制(MPC)。平移控制器采用自适应模型融合预测控制,结合匀速和零速目标模型以适应不同的目标运动模式。该方法利用历史预测误差更新模型权重,并在执行器、跟随距离和视场约束下计算平移指令。通过仿真和真实实验,我们验证了所提框架的有效性,并表明其性能优于现有框架。
cs.RO / 110 / 2609.20747

MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving

MILER:面向非结构化自动驾驶中Sim-to-Real强化学习的语义中层表示
Steinecker, Thomas, Trescher, Denis, Bienemann, Alexander, Luettel, Thorsten, Maehlisch, Mirko
Abstract
Reinforcement learning constitutes a promising approach owing to its potential for superhuman performance and self-learned policies. However, its application to real-world autonomous driving remains scarce, particularly in unstructured environments, because of the challenges associated with sim-to-real transfer for unstructured environments. In this work, we present MILER, an end-to-end policy framework with zero-shot sim-to-real transfer. During offline training, we employ a custom semantic mid-level representation (MLR) simulator and train the policy network using reinforcement learning, with its control outputs applied directly to a bicycle model. During deployment on the real vehicle, camera and LiDAR data are processed by BEVFusion to generate a semantic bird's-eye-view representation consistent with that of the MLR simulator. The actions generated by the policy network are not applied directly to the real vehicle. Instead, we employ a trajectory-alignment strategy that enables zero-shot sim-to-real transfer of both perception and control. We extensively evaluate the proposed framework on a diverse test track comprising numerous challenges, including various obstacles, hairpin curves, velocities of up to 33.6 km/h, and off-road sections. In total, we drove 17.3 km with two different vehicles on a 3.0 km test track without human intervention, thereby demonstrating the effectiveness of our approach. Furthermore, the entire software stack runs on a Jetson AGX Orin.
Chinese Translation
强化学习因其超越人类性能的潜力和自学习策略而成为一种有前景的方法。然而,由于非结构化环境中sim-to-real迁移的挑战,其在真实世界自动驾驶中的应用仍然稀少,尤其是在非结构化环境中。在本工作中,我们提出了MILER,一个具备零样本sim-to-real迁移能力的端到端策略框架。在离线训练阶段,我们采用自定义的语义中层表示(MLR)仿真器,并使用强化学习训练策略网络,其控制输出直接应用于自行车模型(bicycle model)。在真实车辆部署阶段,相机和LiDAR数据经BEVFusion处理,生成与MLR仿真器一致的语义鸟瞰图表示。策略网络生成的动作并非直接作用于真实车辆,而是采用一种轨迹对齐策略,从而实现感知与控制的零样本sim-to-real迁移。我们在一条包含多种挑战的多样化测试赛道上对该框架进行了广泛评估,挑战包括各类障碍物、发夹弯、最高达33.6 km/h的车速以及非铺装路段。最终,我们驾驶两辆不同的车辆在3.0 km的测试赛道上无人工干预地行驶了总计17.3 km,从而证明了该方法的有效性。此外,整个软件栈运行在Jetson AGX Orin上。
cs.RO / 111 / 2609.20756

OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

OPTED:基于免渲染教师的端到端驾驶在策略微调方法
Da Col, Damiano, Igl, Maximilian, Karkus, Peter, Chitta, Kashyap, Ivanovic, Boris, Pavone, Marco, Schindler, Konrad, Sakaridis, Christos
Abstract
As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop deployment can take the vehicle outside the training data distribution, increasing the risk of safety-critical incidents. Closed-loop post-training can mitigate this risk but requires costly simulation for sensor-based policies. We propose OPTED (on-policy fine-tuning for end-to-end driving) which decouples reinforcement learning from the post-training of the end-to-end policy: a privileged teacher is trained using RL on vectorized inputs (HD-map and bounding boxes). This teacher then provides supervision to the pre-trained student during closed-loop post-training. We apply OPTED to two camera-based models, TransFuser and VaVAM, and fine-tune them in AlpaSim, using neural reconstructions (3DGS) of real driving logs. Driving scores increase by factors of 1.6$\times$ and 9.5$\times$, respectively. In controlled experiments OPTED matches closed-loop performance with approximately three orders of magnitude fewer simulator interactions than direct RL post-training, while staying closer to the human prior. Project page: https://01dami23.github.io/opted/
Chinese Translation
随着单纯扩展预训练数据的收益递减,后训练在自动驾驶等物理AI领域正变得越来越重要。端到端驾驶策略通常通过行为克隆在人类演示数据上进行开环预训练。然而,在闭环部署过程中,误差累积会使车辆偏离训练数据分布,增加安全关键事故的风险。闭环后训练可以缓解这一风险,但对于基于传感器的策略而言,需要代价高昂的仿真。我们提出了OPTED(面向端到端驾驶的在策略微调),该方法将强化学习与端到端策略的后训练解耦:一个特权教师使用强化学习在向量化输入(高精地图和边界框)上进行训练。随后,该教师在闭环后训练过程中为预训练的学生模型提供监督。我们将OPTED应用于两个基于相机的模型TransFuser和VaVAM,并在AlpaSim中利用真实驾驶日志的神经重建(3DGS)进行微调。两者的驾驶得分分别提升了1.6倍和9.5倍。在受控实验中,OPTED在达到与直接强化学习后训练相当的闭环性能的同时,所需的仿真交互次数减少约三个数量级,并且更接近人类先验。项目主页:https://01dami23.github.io/opted/
cs.RO / 112 / 2609.20761

Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control

Agile-WAM:一种用于接触密集型机器人控制的敏捷触觉世界动作模型
Zhou, Hanchu, Lynch, Brendan, Goyal, Raman, Gao, Dechen, Kasap, Begum, Zhao, Boqi, Zhang, Junshan
Abstract
World Action Models (WAMs) advance beyond conventional visuomotor policies by jointly predicting future world states and robot actions, enabling the policy to learn physical dynamics that support effective control. However, recent tactile WAMs often rely on large-scale pretrained generative backbones to capture contact-rich physical dynamics, which limit their inference efficiency and flexible deployment. In this paper, we present \ABBR{}, an agile tactile World Action Model for contact-rich robot control. \ABBR{} encodes visual and tactile observations into a shared latent that serves as the source of a direct vision-tactile-to-action flow-matching process, which can jointly generate latent representations of action chunks and future visual/tactile latents. A key observation is that vision and tactile signals evolve at inherently different timescales: adjacent visual frames are often highly similar, whereas tactile signals can change abruptly upon contact. We therefore introduce multi-horizon multimodal prediction in \ABBR{}, which provides supervision for visual latent at a larger temporal offset while predicting the tactile latent in the next frame to capture fine-grained contact dynamics. Across nine simulated and five real-world contact-rich manipulation tasks, \ABBR{} demonstrates strong and robust performance, outperforming the strongest baseline in success rate while maintaining low inference latency. In particular, in five real-world experiments, \ABBR{} yields a relative gain of $\textbf{29.4\%}$ in overall success rates while achieving inference latency of $\textbf{11.9 ms}$. These results demonstrate that multimodal WAM can be achieved with an agile architecture suitable for precise and high-frequency robot control. More details are available on our project page: https://hanchuzhou.github.io/TARO_project_page/.
Chinese Translation
世界动作模型通过联合预测未来世界状态和机器人动作,超越了传统的视觉运动策略,使策略能够学习支持有效控制的物理动力学。然而,近期的触觉世界动作模型通常依赖大规模预训练的生成式骨干网络来捕捉接触密集型的物理动力学,这限制了其推理效率和灵活部署。在本文中,我们提出了 Agile-WAM,一种用于接触密集型机器人控制的敏捷触觉世界动作模型。Agile-WAM 将视觉和触觉观测编码到一个共享的潜在空间中,作为直接的“视觉-触觉到动作”流匹配过程的源头,从而能够联合生成动作块的未来潜在表示以及未来视觉/触觉潜在表示。一个关键观察是,视觉和触觉信号以本质不同的时间尺度演化:相邻视觉帧通常高度相似,而触觉信号在接触时可能发生突变。因此,我们在 Agile-WAM 中引入了多时域多模态预测,在较大的时间偏移上为视觉潜在表示提供监督,同时预测下一帧的触觉潜在表示以捕捉细粒度的接触动力学。在九个仿真任务和五个真实世界接触密集型操作任务中,Agile-WAM 展现出强大且稳健的性能,在成功率上超越了最强的基线方法,同时保持了较低的推理延迟。特别是在五个真实世界实验中,Agile-WAM 的总体成功率取得了 29.4% 的相对提升,同时推理延迟仅为 11.9 毫秒。这些结果表明,多模态世界动作模型可以采用适用于精确、高频机器人控制的敏捷架构来实现。更多详情请访问我们的项目页面:https://hanchuzhou.github.io/TARO_project_page/。
cs.RO / 113 / 2609.20776

GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies

GeoAAC:基于几何的VLA策略去噪轨迹自适应动作分块方法
Chen, Xin, Chen, Sen, Ding, Yujuan, Liu, Jian, Wang, Guoqing, Ye, Wei, Shen, Heng Tao, Bin, Yi
Abstract
Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different task stages may require different levels of action continuity, control precision, and closed-loop feedback, making a fixed horizon unable to accommodate changing control requirements. We propose \textbf{GeoAAC}, a geometry-based adaptive action chunking method for flow-based VLA policies that adjusts the action horizon according to the reliability of the current action prediction. We show that the geometry of Flow Matching denoising trajectories provides process-level information for characterizing prediction reliability, with geometric variation across action prefixes remaining positively correlated with predictive uncertainty. GeoAAC uses this prefix-wise geometry to construct a horizon-wise geometric profile and adaptively determine the action horizon from a single generation without additional training. Experiments with GR00T N1.5 and {\pi}0.5 on LIBERO, LIBERO-Pro, RoboCasa365, and real-world manipulation tasks show consistent improvements over fixed-action-horizon baselines and existing adaptive methods, including up to 8.7 percentage points in simulation and an increase in average real-world success rate from 53.3\% to 74.4\%.
Chinese Translation
动作分块(Action Chunking)被广泛应用于视觉-语言-动作(VLA)策略中的动作生成与执行,但现有方法通常采用固定的动作时域(action horizon)。在滚动执行(rollout)过程中,不同的任务阶段可能需要不同水平的动作连续性、控制精度和闭环反馈,固定时域无法适应不断变化的控制需求。我们提出了GeoAAC,一种基于几何的自适应动作分块方法,适用于基于流(flow-based)的VLA策略,能够根据当前动作预测的可靠性调整动作时域。我们证明了流匹配(Flow Matching)去噪轨迹的几何特性为刻画预测可靠性提供了过程级信息,且动作前缀上的几何变化与预测不确定性保持正相关。GeoAAC利用这种前缀级几何特性构建时域级的几何剖面,并通过单次生成自适应地确定动作时域,无需额外训练。在LIBERO、LIBERO-Pro、RoboCasa365以及真实世界操作任务上使用GR00T N1.5和{\pi}0.5进行的实验表明,该方法相较于固定动作时域基线和现有自适应方法取得了一致的性能提升,在仿真中提升高达8.7个百分点,真实世界平均成功率从53.3%提升至74.4%。
cs.RO / 114 / 2609.20791

StageGuard: Learning Stage Transitions for Long-Horizon Robot Tasks via Agentic Distillation

StageGuard:通过智能体蒸馏学习长时程机器人任务的阶段转换
Huang, Jinbang, Hu, Yuanzhao, Li, Zhiyuan, Qi, Ran, Xiao, Yixin, Wu, Yangzheng, Ba, Tengyue, Zhang, Zhanguang, Zhang, Yingxue
Abstract
Hierarchical planning frameworks combine skills from multiple robot control policies for long-horizon task execution, where determining when to terminate the current skill and advance to the next subtask is essential. Existing approaches often rely on pre-designed completion signal checkers that are hard to obtain in real-world execution. Large-scale vision-language models (VLMs) offer strong reasoning capabilities, but their decision boundaries are not inherently aligned with task completion criteria, while cloud deployment and lengthy reasoning introduce substantial latency, limiting real-time monitoring. We propose StageGuard, an agentic distillation framework for accurate and efficient stage-transition decisions. StageGuard combines teacher-model reasoning with demonstration trajectories to generate structured explanations of subtask completion and policy switching. A lightweight student VLM uses these explanations to generate compact self-explanations, which are used for supervised fine-tuning. We evaluate stage-transition prediction on trajectories from two benchmarks and assess closed-loop task success through integration into hierarchical robot control on BEHAVIOR-1K, with further validation on real robots. Results show substantial improvements in stage-transition prediction while supporting efficient online monitoring.
Chinese Translation
层次化规划框架通过组合来自多个机器人控制策略的技能来执行长时程任务,其中确定何时终止当前技能并推进到下一子任务至关重要。现有方法通常依赖预先设计的完成信号检测器,而这些检测器在真实世界执行中难以获得。大规模视觉-语言模型(VLM)具备强大的推理能力,但其决策边界与任务完成标准并不天然对齐,同时云端部署和冗长的推理会引入显著的延迟,限制了实时监控。我们提出StageGuard,一个用于准确且高效的阶段转换决策的智能体蒸馏框架。StageGuard将教师模型的推理与示范轨迹相结合,生成关于子任务完成和策略切换的结构化解释。一个轻量级的学生VLM利用这些解释生成简洁的自解释内容,用于监督微调。我们在两个基准的轨迹上评估阶段转换预测,并通过将其集成到BEHAVIOR-1K上的层次化机器人控制中评估闭环任务成功率,同时在真实机器人上进行了进一步验证。结果表明,阶段转换预测性能有显著提升,同时支持高效的在线监控。
cs.RO / 115 / 2609.20820

Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision

工作空间模型:基于显著性驱动监督的轻量级机器人记忆
Dashora, Nitish, Chen, Douglas, Shenfeld, Idan, Marangola, John, Agrawal, Pulkit, Simchowitz, Max
Abstract
Complex robotic manipulation tasks frequently require a long-term memory of past events and actions. As conditioning on full histories renders policies prone to spurious correlations and degrades performance, many approaches to policy memory involve compressing historical information through expensive VLM queries in-the-loop to process only task-salient information. In this paper, we propose an alternative approach in which computationally intensive VLM queries are made during train-time to learn a lightweight latent memory that can be efficiently queried at deployment time. Our representation, which we call the \textbf{workspace token}, is trained by (1) using a VLM to identify current and historical information necessary for completing a task, then (2) distilling these into the workspace token using a set-reconstruction decoder loss. In both simulation and hardware, we show that the workspace token can be used as a drop-in replacement for observations during deployment, enabling policies to solve memory-intensive tasks without the need for VLM reasoning in-the-loop. Interestingly, we found that workspace tokens are not only more lightweight but also lead to better policy performance.
Chinese Translation
复杂的机器人操作任务通常需要对过去事件和动作的长期记忆。由于以完整历史信息为条件会使策略容易受到虚假相关性的影响并降低性能,许多策略记忆方法通过在控制回路中进行昂贵的视觉-语言模型(VLM)查询来压缩历史信息,从而仅处理与任务相关的显著信息。本文提出了一种替代方案:在训练阶段进行计算密集型的VLM查询,以学习一个可在部署时高效查询的轻量级潜在记忆。我们将这种表示称为工作空间标记(workspace token),其训练方式为:(1)使用VLM识别完成任务所需的当前和历史信息,然后(2)通过集合重建解码器损失将这些信息蒸馏到工作空间标记中。在仿真和硬件实验中,我们证明了工作空间标记可以作为部署时观测信息的即插即用替代品,使策略无需在控制回路中进行VLM推理即可解决需要记忆的任务。有趣的是,我们发现工作空间标记不仅更加轻量,还能带来更好的策略性能。
cs.RO / 116 / 2609.20822

Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation

具有障碍感知安全框架的编程智能体用于安全机器人操作
Xu, Bingxin, Shang, Yuzhang, Dong, Zhen, Ferrara, Emilio
Abstract
Coding agents have emerged as a promising paradigm for robot manipulation: a language model writes the robot controller as a program, and agents built in this way now operate robots without robot-specific training.Whether this paradigm is also safe, however, has not been asked. We evaluate coding agent under a safety constraint, where each task pairs a manipulation goal with an obstacle the robot must not touch. The agent pursues the goal but collides with the obstacle in most cases, treating task completion as its sole objective while neglecting safety. The agent reasons about the obstacle in its traces, and the prompt already forbids touching it, so neither perception nor instruction is at fault; the fault lies in the planning, where the stated constraint never becomes a priority. By decomposing manipulation into a route phase and a contact-rich moment, we locate the source of the failure. Along the route, the model cannot prioritize the safety constraint, having no notion of a clearing route and none of replanning once a chosen route becomes infeasible. At the contact, it is unaware that contact execution is bounded by the same constraint. To close this gap, we present SafeHarness, which equips the model with two obstacle-aware harnesses that enable it to prioritize the safety constraint. Obstacle-aware route planning grounds the objects as bounding boxes and draws candidate routes over them as sequences of waypoints. The agent then plans a route in advance, verifies it, replans when necessary, and only then executes it. Obstacle-aware contact execution instead selects the contact position so that the contact itself avoids the obstacle. SafeHarness attains 71.9% task success and 87.5% collision avoidance, surpassing the previous SOTA by 6.5% and 27.0%, respectively. These results are $2.3\times$ and $1.5\times$ those of the same agent without harnesses.
Chinese Translation
编程智能体已成为机器人操作的一种有前景的范式:语言模型将机器人控制器编写为程序,以此方式构建的智能体现在无需针对特定机器人的训练即可操作机器人。然而,这一范式是否同样安全,此前尚未被探讨。我们在安全约束下评估编程智能体,其中每个任务在操作目标之外还包含一个机器人不得触碰的障碍物。该智能体追求目标,但在大多数情况下会与障碍物发生碰撞,将任务完成视为唯一目标而忽视安全。智能体在推理轨迹中会对障碍物进行思考,且提示词中已经禁止触碰障碍物,因此问题既不在于感知也不在于指令;错误出在规划环节,即所陈述的约束从未成为优先事项。通过将操作分解为路径阶段和富接触时刻,我们定位了失败的根源。在路径阶段,模型无法优先考虑安全约束,既没有避让路径的概念,也没有在所选路径不可行时进行重规划的能力。在接触阶段,模型没有意识到接触执行同样受该约束的限制。为弥合这一差距,我们提出了SafeHarness,它为模型配备了两个障碍感知的安全框架,使其能够优先考虑安全约束。障碍感知路径规划将物体定位为边界框,并在其上绘制由路径点序列组成的候选路径。智能体随后预先规划路径、进行验证、必要时重新规划,最后才执行。障碍感知接触执行则通过选择接触位置,使接触本身避开障碍物。SafeHarness达到了71.9%的任务成功率和87.5%的碰撞规避率,分别超越此前最优方法6.5%和27.0%。这些结果分别是无安全框架的同一智能体的2.3倍和1.5倍。
人工智能 (Artificial Intelligence)
89
cs.AI / 1 / 2609.19170

Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes

正则化强调型时序差分学习:常步长下的稳定性
Chen, Xingguo, Wu, Zhaohui, Ye, Jinguo, Li, Chao, Yang, Shangdong, Yang, Guang, Liang, Skylar, Wang, Wenhao
Abstract
Emphatic temporal-difference learning (ETD) stabilizes the expected off-policy TD update and changes its projection geometry, but neither property determines constant-stepsize sampled dynamics. We construct an ergodic two-state counterexample in which the ETD mean map contracts while the sampled product has a positive top Lyapunov exponent. Regenerative-cycle analysis separates this sign from the infinite variance of the follow-on trace. We introduce regularized emphatic TD (RETD), a normalized first-order post-shock repair that leaves the trace and importance ratios unchanged, stores the emphatic TD signal in a leaky scalar state, and releases a delayed correction. RETD's raw equilibrium is an affine shift of the ETD equilibrium; single- and two-regularization readouts recover the ETD fixed point exactly. We prove almost-sure convergence for harmonic diminishing stepsizes and a conditional constant-stepsize moment-contraction result from a Markovian random-product bound. RETD has certified negative exponents on the two-state construction and one Baird point, whereas the positive Baird ETD sign remains numerical. Paired 10,000-run experiments validate both separations, fixed-point recovery, a nonmonotone stability region, and task dependence. RETD changes post-shock dynamics; it does not reduce the shared follow-on-trace variance.
Chinese Translation
强调型时序差分学习(ETD)能够稳定期望离策略TD更新并改变其投影几何,但这两个性质均不能决定常步长下的采样动态。我们构造了一个遍历的两状态反例:其中ETD的均值映射是压缩的,而采样的乘积却具有正的李雅普诺夫指数。再生周期分析将这一符号与后续迹(follow-on trace)的无穷方差分离开来。我们提出了正则化强调型TD(RETD),这是一种归一化的一阶冲击后修复方法:它保持迹和重要性比率不变,将强调型TD信号存储在一个带泄漏的标量状态中,并释放延迟校正。RETD的原始平衡点是ETD平衡点的一个仿射平移;单正则化和双正则化读出可以精确恢复ETD的不动点。我们证明了在调和递减步长下RETD的几乎必然收敛性,并基于马尔可夫随机乘积界给出了条件性的常步长矩压缩结果。在两状态构造和一个Baird问题上,RETD具有经认证的负指数,而Baird问题上ETD的正号仍仅为数值结果。成对的10,000次运行实验验证了上述两种分离、不动点恢复、一个非单调稳定区域以及任务依赖性。RETD改变的是冲击后动态,并不能降低共享的后续迹方差。
cs.AI / 2 / 2609.19180

BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research

BioPhys-Bridge:面向物理驱动的生物学研究中跨学科科学推理的基准测试
Xu, Qingyang
Abstract
Language models face unique challenges in analyzing interdisciplinary scientific research literature. In biophysics research, faithful answers require grounding observed data in source evidence, interpreting it through a quantitative physics model, and linking it to a biological mechanism. To address this challenge, we introduce BioPhys-Bridge, a novel benchmark dataset for evidence-grounded scientific reasoning over biophysical literature. Each case contains evidence blocks, stable evidence IDs, quantitative values, units, equations, assumptions, mechanisms, and next decisions as grounding targets for question answering (QA) and retrieval-augmented generation (RAG). The initial release contains 500 cases, 1,517 agent-facing tasks, and covers six biological domains and nine physical model families, including three sparse families reserved for future expansion. We enforce strict quality gates for all cases in schema, evidence-integrity, quantitative-grounding, source-license, duplicate, unit-normalization, with domain expert review and annotation for 81 cases. Preliminary evaluations show that DeepSeek-V4-Flash obtain the highest evidence-ID $F_1$ score (0.360), followed by Qwen3.7-Max (0.316) and GPT-4o-mini (0.294). BioPhys-Bridge is an interdisciplinary benchmark for evaluating attribution, faithfulness, hallucination reduction, and biological experiment design with complex, multi-step scientific reasoning. Future works will increase the size and complexity of the dataset and perform comprehensive evaluations. Code and data are available in the GitHub repository and on Hugging Face.
Chinese Translation
语言模型在分析跨学科科学研究文献时面临独特挑战。在生物物理学研究中,给出可靠的答案需要将观测数据溯源至原始证据,通过定量物理模型对其进行解释,并将其与生物学机制联系起来。为应对这一挑战,我们提出了 BioPhys-Bridge,一个用于生物物理文献中基于证据的科学推理的新型基准数据集。每个案例包含证据块、稳定的证据 ID、定量数值、单位、方程、假设、机制以及下一步决策,作为问答(QA)和检索增强生成(RAG)的 grounding 目标。初始版本包含 500 个案例和 1,517 个面向智能体的任务,覆盖六个生物学领域和九个物理模型家族,其中包括三个预留用于未来扩展的稀疏家族。我们对所有案例在模式校验、证据完整性、定量 grounding、来源许可、查重和单位归一化等方面实施严格的质量控制,并对 81 个案例进行了领域专家评审与标注。初步评估结果显示,DeepSeek-V4-Flash 获得了最高的证据 ID $F_1$ 分数(0.360),其次是 Qwen3.7-Max(0.316)和 GPT-4o-mini(0.294)。BioPhys-Bridge 是一个跨学科基准,用于评估归因能力、忠实性、幻觉抑制以及涉及复杂多步科学推理的生物学实验设计。未来工作将扩大数据集的规模和复杂性,并开展全面的评估。代码和数据已在 GitHub 仓库和 Hugging Face 上发布。
cs.AI / 3 / 2609.19182

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

我们对大语言模型(LLM)有何期待?大语言模型基准测试设计图谱
Wang, Chao
Abstract
Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms. The collection shows growing emphasis on action, interaction, and professional applications, while established and newer design elements frequently coexist. Model participation also develops unevenly: LLM-based scoring grows within both agent and non-agent groups, whereas model-generated materials show no comparable sustained increase in recent cohorts. These findings illuminate how public research translates capability expectations into concrete tests and criteria for success. As AI participates in constructing tests, performing tasks, and judging responses, they also raise a question: does expanding evaluation provide more independent evidence, or risk reproducing the preferences and blind spots of its participating models?
Chinese Translation
基准测试是评估和传达大语言模型(LLM)进展的核心手段。然而,仅凭模型排名难以揭示评估需求本身的变化。不断扩展的基准测试种类提供了另一个视角:研究者期望大语言模型做什么,以及他们将什么视为成功的表现。我们对2022年1月至2026年8月间arXiv提交的论文进行系统性梳理,共纳入14,767篇引入或更新评估资源的论文。通过分阶段筛选和自动化全文编码,我们考察了目标系统与领域、评估材料与条件以及评分机制的变化。这一文献集合显示,研究重心日益向行动、交互和专业应用倾斜,同时既有的设计元素与较新的设计元素经常并存。模型参与的发展也不均衡:基于LLM的评分在智能体(agent)与非智能体两类群体中均在增长,而模型生成的评估材料在近期的论文队列中并未呈现相应的持续增长。这些发现阐明了公开研究如何将能力期望转化为具体的测试和成功标准。随着AI参与构建测试、执行任务和评判回答,这也引出一个问题:评估的扩展是在提供更多独立的证据,还是有风险复制参与其中的模型的偏好与盲点?
cs.AI / 4 / 2609.19203

Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer

观点:是时候用自进化操作系统层来虚拟化基础模型了
Bhattacharya, Suparna, Kumar, Tarun, Xu, Cong, Mopur, Satish Kumar, Li, Jiahao, Mishra, Ashish, Tripathy, Aalap, Koomthanam, Annmary Justine, Foltin, Martin, Foster, Ian
Abstract
AI applications have shifted from single, monolithic foundation models (FM) to compound agentic systems. Yet today's stacks remain fragmented: even as protocols (e.g., MCP, A2A) ease tool/agent connectivity, each framework embeds an implicit runtime for state, memory, budgets, and guardrails, making behavior non-portable and governance brittle. It mirrors computing before operating systems, when every program re-implemented basic services. This position paper argues that the field now needs a Foundation Model Operating System (FMOS) -- a system layer that virtualizes FM interactions analogous to how virtual machines abstract physical hardware, giving applications the illusion of dedicated, trustworthy FM instances with effectively unbounded capabilities. Internally, the FMOS orchestrates knowledge across memory tiers, model selection and resource allocation, and verification and policy enforcement. Like the human brain switching between fast intuition and slow deliberation, the FMOS learns when to intervene and when to let inference proceed directly and continuously adapting its policies based on operational experience.
Chinese Translation
AI应用已从单一、整体式的基础模型(Foundation Model, FM)转向复合型智能体系统。然而,当今的技术栈仍然碎片化:尽管各类协议(如MCP、A2A)简化了工具与智能体之间的连接,但每个框架都内嵌了一套隐式的运行时来管理状态、记忆、预算和防护机制,导致模型行为不可移植且治理脆弱。这类似于操作系统出现之前的计算形态——当时每个程序都需要重新实现基础服务。本立场论文认为,该领域现在需要一个基础模型操作系统(Foundation Model Operating System, FMOS)——一个将FM交互虚拟化的系统层,正如虚拟机抽象物理硬件一样,为应用程序提供专属、可信且能力实际上无限制的FM实例的错觉。在内部,FMOS负责跨记忆层级进行知识编排、模型选择与资源分配,以及验证与策略执行。就像人脑在快速直觉与慢速深思之间切换一样,FMOS学习何时进行干预、何时让推理直接进行,并基于运行经验持续调整其策略。
cs.AI / 5 / 2609.19212

What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis

当前系统性泛化任务遗漏了什么?一种以推理为中心的分析
Qi, Chengwen, Ye, Deheng, Bian, Yatao
Abstract
Systematic generalization, the ability to solve novel problems by recombining known atomic elements, is central to human intelligence but difficult to study rigorously under controlled settings. Existing studies therefore rely on simplifications such as approximately linear action composition, productivity-based tests, and action-explicit goals, which make systematic generalization easier to study but omit some essential aspects of this capability. To characterize what these simplifications miss, we adopt a reasoning-centered lens and introduce TranSGrid, a testbed that brings deductive, inductive, and abductive reasoning together within a unified task. Experiments with seven Transformers on 4,800 TranSGrid instances show that all models perform much worse on TranSGrid than on a held-out test set: the largest model solves 79.6% of the test set, but only 55.3% of TranSGrid and 15.8% of the hardest subset. The gap remains within the training length range, showing that productivity alone is not sufficient to evaluate systematic generalization. Additionally, we reintroduce the other two simplifications into TranSGrid: one variant makes actions compose almost linearly (reducing the inductive demand), the other makes goals action-explicit (reducing the abductive one). In both, solve rates return to roughly the test set level, showing that either simplification alone is enough to reduce TranSGrid to an ordinary held-out test set. Together, our results show that existing tasks reduce either or both of the inductive and abductive demands, and that comprehensively measuring systematic generalization requires a task that involves all three forms of reasoning.
Chinese Translation
系统性泛化是指通过重组已知的原子元素来解决新问题的能力,它是人类智能的核心,但在受控环境下难以进行严格研究。因此,现有研究依赖于一些简化设定,例如近似线性的动作组合、基于生成性(productivity)的测试以及动作显式的目标,这些简化使系统性泛化更易于研究,却忽略了该能力的某些本质方面。为了刻画这些简化所遗漏的内容,我们采用以推理为中心的视角,提出了TranSGrid——一个在统一任务中融合演绎、归纳与溯因推理的测试平台。在4,800个TranSGrid实例上对七个Transformer模型进行的实验表明,所有模型在TranSGrid上的表现都远逊于留出测试集:最大的模型在测试集上达到79.6%的解决率,但在TranSGrid上仅为55.3%,在最难子集上仅为15.8%。该差距在训练长度范围内依然存在,说明仅凭生成性不足以评估系统性泛化。此外,我们将另外两种简化重新引入TranSGrid:一个变体使动作几乎线性组合(降低了归纳需求),另一个变体使目标动作显式(降低了溯因需求)。在两种情况下,解决率都回升到大致与测试集相当的水平,表明任一简化都足以将TranSGrid降格为普通的留出测试集。综上所述,我们的结果表明,现有任务降低了归纳或溯因需求(或二者兼有),而要全面衡量系统性泛化,需要一个涉及全部三种推理形式的任务。
cs.AI / 6 / 2609.19244

Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

对话式大语言模型智能体的网络搜索特征分析:从搜索决策与策略到搜索结果与响应
Amani, Mahsa, Lee, Seungeon, Dash, Abhisek, Fraihi, Asmaa El, Jang, Yunah, Kirsten, Elisabeth, Wu, Qinyuan, Gummadi, Krishna P., Gupta, Manish, Ravichander, Abhilasha, Zafar, Muhammad Bilal, Das, Soumi
Abstract
Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-world user interactions (invivo) with controlled experiments using the same platform's models by their APIs (invitro). We investigate the quality of agentic decisions to invoke Web search, their strategies to formulate queries, the potential domain preferences in the search results they receive, and the choices they make when transforming search results into grounded responses. We find that Web-search decisions vary substantially across platforms and models, while more frequent Web-search invocation does not necessarily yield better response quality. We further show that conversational agents employ different complex querying strategies and that platform specific search engines return search results from their preferred domains. Finally, although responses are largely grounded in search results, some claims rely on uncited search results, raising concerns about attribution and reliability. Our findings have important implications for the design of future AI agents and Web search tools optimized for conversational retrieval.
Chinese Translation
对话式大语言模型(LLM)智能体日益依赖网络搜索,然而智能体搜索的端到端生命周期仍缺乏深入理解。我们首次针对四大主流对话平台(ChatGPT、Claude、Grok 和 DeepSeek)的网络搜索行为展开研究,结合真实用户交互(in vivo,体内实验)与通过 API 使用相同平台模型进行的受控实验(in vitro,体外实验)。我们考察了智能体调用网络搜索的决策质量、其构建查询的策略、所获搜索结果中潜在的域名偏好,以及其将搜索结果转化为有据可依的响应时的选择方式。我们发现,网络搜索的调用决策在不同平台和模型之间差异显著,且更频繁地调用网络搜索并不一定能带来更好的响应质量。我们进一步发现,对话式智能体采用不同的复杂查询策略,且各平台特定的搜索引擎会返回来自其偏好域名的搜索结果。最后,尽管响应在很大程度上基于搜索结果,但部分论断依赖于未经引用的搜索结果,这引发了关于来源归属和可靠性的担忧。我们的研究结果对未来 AI 智能体的设计以及面向对话式检索优化的网络搜索工具具有重要意义。
cs.AI / 7 / 2609.19387

Do AI Agents Understand Computer Architecture?

AI智能体理解计算机体系结构吗?
Sharan, Ambika, Chirkov, Grigory, Abbasloo, Soheil
Abstract
Agents are increasingly asked to design hardware, and increasingly reported to succeed. Such reports establish that a design improved; they cannot establish why. An agent that improves an accelerator may be reasoning about the machine, or may be searching competently over knobs whose meaning it never recovers -- and only the first transfers to the next architecture. Existing evaluations cannot tell the two apart, because they vary the agent while holding the framing of the problem fixed. We do the opposite. AutoTuring hands the same agent the same 15-dimensional accelerator space twice: once as named architectural knobs with simulator counters, once as anonymous variables on [0,1], with the evaluator, the legal space and the reachable optima held identical, so that the only thing that varies is whether the problem means anything. The gap between the two is the measurement. On a nine-kernel FP16 GEMM basket, meaning pays: the architect beats a modeled H200 by 5.4% and its blind counterpart by 12.3% on average, with 70.1% fewer simulator calls. It does not pay uniquely: a critic loop recovers most of that gap for the blind agent and buys the architect nothing, so architectural knowledge and structured critique behave as substitutes rather than as complements. We report these as preliminary findings -- five to six runs per condition on a single modeled accelerator -- and take the comparison itself, not the accelerator, to be the contribution.
Chinese Translation
智能体越来越多地被要求设计硬件,且越来越多地被报告取得成功。此类报告只能证明某个设计得到了改进,却无法证明改进的原因。一个改进了加速器的智能体,可能是在对机器进行推理,也可能只是在其从未理解含义的诸多参数上进行高效搜索——只有前者才能迁移到下一代体系结构上。现有评估无法区分这两种情况,因为它们在保持问题框架不变的同时改变了智能体。我们反其道而行之。AutoTuring让同一个智能体面对同一个15维加速器设计空间两次:一次是以带有模拟器计数器的具名体系结构参数形式呈现,一次是以[0,1]区间上的匿名变量形式呈现,同时保持评估器、合法设计空间和可达最优解完全一致,从而使唯一变化的因素是问题本身是否具有意义。两者之间的差距即为测量结果。在一个包含九个核心的FP16 GEMM基准任务上,语义确实带来收益:该架构师智能体平均比建模的H200高出5.4%,比其盲化对照版本高出12.3%,且模拟器调用次数减少70.1%。但这种收益并非其独有:一个批评者循环能为盲化智能体挽回大部分差距,而对架构师智能体却毫无增益,这表明体系结构知识与结构化批评行为上是替代品而非互补品。我们将这些报告为初步发现——在单个建模加速器上每个条件仅进行五到六次运行——并认为这一比较本身,而非加速器,才是本文的贡献。
cs.AI / 8 / 2609.19391

MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

MAGS:基于多智能体自动形式化保障智能体输出安全性的框架
Wu, Albert, Roberts, Nicholas, Huang, Tzu-Heng, Lin, Haoran, Friedman, Gil, Cho, Sungjun, Orlanski, Gabriel, Sala, Frederic
Abstract
LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-Verifier, can detect many failures but struggle to cover all possible edge cases. Formal verification addresses this by providing machine-checkable guarantees over specified properties, but traditionally demands substantial manual specification and proof engineering. We introduce a unified multi-agent framework, MAGS, that generates executable programs with formal safety guarantees, using Dafny as a verification-aware intermediate representation where safety properties can be mechanically checked. MAGS formalizes and freezes human-audited APIs and safety requirements, translates generated code into Dafny, repairs violations using verifier feedback, and compiles verified programs back into executable code. We evaluate MAGS on 100 CUDA kernels, 100 terminal scripts, and 20 robotic-arm tasks. Across all 220 examples, it achieves a 100% success rate in producing programs with non-trivial safety guarantees against frozen specifications. Independent safety and functional evaluations further show strong performance across all three domains, while revealing failures when the auto-formalized semantics do not fully capture the target behavior.
Chinese Translation
LLM编程智能体如今生成的复杂程序规模日益庞大,使得全面的人工审查愈发困难,从而增加了安全(safety)与信息安全(security)失败的风险。常见方法,包括模糊测试、静态分析以及LLM即验证器(LLM-as-a-Verifier),虽然能够检测许多故障,但难以覆盖所有可能的边界情况。形式化验证通过对指定属性提供机器可验证的保证来解决这一问题,但传统上需要大量的人工规范编写与证明工程。我们提出了一个统一的多智能体框架MAGS,该框架能够生成具有形式化安全保证的可执行程序,并以Dafny作为支持验证的中间表示,使安全属性可被机械地检验。MAGS将经人工审计的API和安全需求进行形式化并冻结,将生成的代码翻译为Dafny,利用验证器反馈修复违规之处,并将经验证的程序编译回可执行代码。我们在100个CUDA内核、100个终端脚本和20个机械臂任务上对MAGS进行了评估。在全部220个示例中,MAGS在针对冻结规范生成具有非平凡安全保证的程序方面均达到100%的成功率。独立的安全性与功能性评估进一步表明其在所有三个领域均有出色表现,同时揭示了当自动形式化的语义未能完整刻画目标行为时的失败情形。
cs.AI / 9 / 2609.19425

Closed-World Resolution Against Tool Hallucination in LLM Agents

面向LLM智能体工具幻觉的封闭世界解决机制
Iyer, Laxmipriya Ganesh
Abstract
Tool-augmented large language model (LLM) agents fail in a way no tool-selection or tool-security method addresses: they call tools that do not exist and pass arguments no schema declares. Existing defenses either pick the right tool (selection) or constrain what an agent may do with real tools (gating), both of which presuppose the emitted call refers to a real tool at all. We show this is a structural blind spot: a hallucinated call is by construction not a decision any gate made, so no gate can reject it. This paper is primarily a measurement and benchmark study. We give a five-class taxonomy of tool hallucination (H1-H5) and, as a reference point, the Resolution Rung: a training-free, closed-world resolver (registry membership plus a signature check) whose interest is where it must sit, not what it computes. We prove hallucination defense must precede any causal gate, and characterize the one irreducible residue (borrowed arguments schema-indistinguishable from a valid call). Across ten hosted models under two invocation surfaces we measure 322 genuine hallucinations; fabricated-tool calls concentrate on the unconstrained raw-JSON surface (34 vs. 3), and model scale does not help (a 675B model matches a 7-8B one). We then extend to the Model Context Protocol, where merging several servers into one namespace creates hallucination surfaces a single registry cannot express (a second taxonomy, M1-M5); on the live MCP surface we measure 154 hallucinations, including from frontier models that were clean on the single-registry surface, because collisions and shadowing are structural to the merge. We release the versioned Hallucinated-Tools Benchmark (HTB) so any resolver is comparable across submissions.
Chinese Translation
工具增强的大语言模型(LLM)智能体会以一种现有工具选择或工具安全方法均无法应对的方式失败:它们调用并不存在的工具,并传递模式(schema)中从未声明的参数。现有防御机制要么选择正确的工具(工具选择),要么限制智能体对真实工具的操作(访问控制),但二者都预设了所发出的调用至少指向一个真实存在的工具。我们证明这是一个结构性盲点:幻觉调用从构造上就不是任何访问控制机制所做的决策,因此任何访问控制都无法将其拒绝。本文主要是一项测量与基准研究。我们提出了工具幻觉的五分类体系(H1–H5),并作为参照点提出了解决层(Resolution Rung):一个免训练的封闭世界解决器(包括注册表成员资格检查加签名校验),其意义在于它必须所处的位置,而非它计算的内容。我们证明幻觉防御必须先于任何因果性访问控制,并刻画了唯一不可消除的残余(即与有效调用在模式上无法区分的“借用参数”)。在两种调用接口下的十个托管模型上,我们测量到322次真实幻觉;伪造工具调用集中在无约束的原始JSON接口上(34次对比3次),且模型规模无济于事(一个675B模型与7–8B模型表现相当)。随后,我们将研究扩展到模型上下文协议(Model Context Protocol,MCP),其中将多个服务器合并到一个命名空间会产生单一注册表无法表达的幻觉面(第二个分类体系M1–M5);在实际MCP接口上我们测量到154次幻觉,其中包括在单一注册表接口上表现干净的前沿模型,因为冲突与遮蔽是服务器合并的结构性特征。我们发布了带版本管理的幻觉工具基准(Hallucinated-Tools Benchmark,HTB),使任何解决器的结果都可以在不同提交之间进行比较。
cs.AI / 10 / 2609.19448

The syntax and semantics of goals

目标的句法与语义
Abel, David M., Ho, Mark K.
Abstract
In both cognitive science and computer science, goals are conceptualized as cognitive states that flexibly combine with world knowledge to organize and specify purposeful behavior. In this way, goals are compositional representations whose content relates to rational behavior. We here draw attention to goals as representations and their content because it highlights a parallel with other areas in cognitive science - in particular, the syntax-semantics interface in linguistics and logic - while also foregrounding foundational questions about the expressivity, design, and efficiency of different goal representations. For example, goals are typically taken as fixed and imposing constraints on desirable behaviors, but we can also identify constraints on goal representations themselves, such as whether a particular goal language is sufficiently expressive to capture behaviors of interest, or whether different goal representations capture the same behavior. Here, we synthesize work that aims to characterize the properties of different goal representations and suggest these are points of a broader design space. We close by discussing how distinguishing the form and meaning of goals can elucidate the implicit assumptions we make about goals, inform the study of interactions between higher-level cognition and motivation, and isolate axes of variation for different conceptions of goals.
Chinese Translation
在认知科学和计算机科学中,目标被概念化为一种认知状态,它能够灵活地与世界知识相结合,以组织和明确有目的的行为。由此,目标是组合性的表征,其内容与理性行为相关。我们在此关注目标作为表征及其内容的问题,因为这凸显了认知科学中其他领域的相似性——尤其是语言学和逻辑学中的句法-语义接口——同时也凸显了关于不同目标表征的表达能力、设计和效率的基础性问题。例如,目标通常被视为固定的、并对期望行为施加约束的,但我们也可以识别目标表征本身所受的约束,例如某个特定的目标语言是否具有足够强的表达能力以刻画所关注的行为,或者不同的目标表征是否刻画了相同的行为。在此,我们综合了旨在刻画不同目标表征属性的研究工作,并指出这些表征是一个更广泛的设计空间中的若干节点。最后,我们讨论了如何通过区分目标的形式与意义,来阐明我们对目标所持有的隐含假设,为高层认知与动机之间交互作用的研究提供启示,并为不同目标概念分离出变化的维度。
cs.AI / 11 / 2609.19465

Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

强化学习后训练下语言模型的组合推理
He, Yu, Li, Yingxi, Wang, Yifei, Vitercik, Ellen
Abstract
Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learning (RL) have substantially improved the reasoning abilities of language models (LMs), their effects on compositional reasoning remain less well understood. We propose a dependency-graph framework to formalize compositional reasoning, yielding three levels of compositionality with increasing complexity. Empirically, we instantiate this framework with data-structure tasks, which provide deterministic reward computation and clear compositional structure. We find a consistent decomposed-to-composed asymmetry: decomposed-skill training does not reliably transfer to composed tasks, whereas composed-task training transfers more readily back to decomposed tasks. We provide theoretical explanation for this asymmetry, and further evaluate compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. Finally, we present a pilot study on real-world tool-calling benchmarks, showing preliminary evidence that the decomposed-to-composed asymmetry can extend to practical settings.
Chinese Translation
组合推理对现实世界的问题解决至关重要:由于训练数据必然有限,模型必须通过以新方式组合已学习的技能来实现泛化。尽管强化学习(RL)等后训练方法已显著提升了语言模型(LM)的推理能力,但其对组合推理的影响仍不甚明晰。我们提出一个依赖图框架来形式化组合推理,划分出复杂度递增的三个组合性层次。在实证方面,我们以数据结构任务实例化该框架,这些任务提供确定性的奖励计算和清晰的组合结构。我们发现一种一致的“分解到组合”不对称性:分解技能训练不能可靠地迁移到组合任务,而组合任务训练则更容易反向迁移到分解任务。我们为这一不对称性提供了理论解释,并进一步在长度外推、结构分布偏移以及向需要未见技能的任务迁移等场景下评估组合泛化能力。最后,我们针对现实世界的工具调用基准进行了初步研究,提供了初步证据表明“分解到组合”不对称性可以延伸到实际应用场景。
cs.AI / 12 / 2609.19472

Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

超越界面层的安全性:基于大语言模型潜在状态的有害内容检测
Khatri, Alizishaan, Prabhu, Chiquita, Neogi, Omkar
Abstract
Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model's internal workings, creating a fundamental assurance gap. We ask: does the model already know when the content is harmful? We extract activations from LLaMA-3.1-8B and train lightweight MLP classifier probes (12.6M parameters) to detect harmful prompts. Evaluated on WildJailbreak, Beavertails, and AEGIS 2.0, our probes achieve F1 scores of 99%, 83%, and 84%, respectively competitive with 1000x larger guard models while cutting latency and compute costs.
Chinese Translation
自主系统日益依赖大语言模型(LLM),但围绕这些模型构建的安全基础设施会引入延迟和计算开销,限制了其在资源受限、时间敏感场景中的实用性。现有的外部护栏(guardrail)模型无法感知模型的内部运作机制,造成了根本性的保障缺口。我们提出这样一个问题:模型本身是否已经知道内容是有害的?我们从LLaMA-3.1-8B中提取激活值,并训练轻量级MLP分类器探针(1260万参数)来检测有害提示词。在WildJailbreak、Beavertails和AEGIS 2.0数据集上的评估显示,我们的探针分别取得了99%、83%和84%的F1分数,与规模大1000倍的护栏模型相当,同时显著降低了延迟和计算成本。
cs.AI / 13 / 2609.19513

QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training

QVAC Genesis III:一个用于高效语言模型预训练的大规模高质量开放合成STEM语料库
Vitabile, Davide, Ranjan, N., Nambiar, Akshay, Gupta, Kamal K., Nazir, Amril
Abstract
High-quality pre-training data is a critical bottleneck for educational and STEM-specific language models targeting edge AI and on-device deployment where token budgets are tightly constrained. While major organizations train ever-larger models on private corpora, the open ecosystem lacks STEM-focused synthetic datasets that deliver high per-token learning value efficiently for small models. To address this gap, we introduce QVAC Genesis III, a 191.43B-token, STEM-focused multi-domain synthetic corpus covering 19 domains across several difficulty levels and different educational styles. QVAC Genesis III is built via a dual generation strategy that performs targeted teacher distillation using a weak edge-scale student model as signal: the student's failures are converted into corrective explanations, while its successes are expanded into contrastive option-level reasoning over all answer choices. We further introduce an LLM-as-a-parser evaluation protocol that extracts final answers from free-form outputs and tracks both accuracy and answer validity. To validate the effectiveness of our QVAC Genesis III data, we conduct controlled from-scratch ablations with 1.7B-parameter models, showing that models trained with QVAC Genesis III consistently outperform both models trained with the open-source synthetic corpus Cosmopedia-v2 and the publicly released Cosmo-1B model across ARC, GPQA Diamond, and MMLU STEM benchmarks, achieving up to +28.57% on ARC-E and +21.35% on ARC-C, while reaching a Valid Answer Rate of up to 99.45%.
Chinese Translation
高质量的预训练数据是面向边缘AI和端侧部署的教育及STEM专用语言模型的关键瓶颈,因为这些场景下token预算受到严格限制。尽管主要机构在私有语料上训练越来越大的模型,开放生态系统仍缺乏能够为小模型高效提供高单位token学习价值的STEM导向合成数据集。为填补这一空白,我们提出了QVAC Genesis III,一个包含1914.3亿token、聚焦STEM的多领域合成语料库,覆盖19个领域、多个难度层级和不同的教学风格。QVAC Genesis III采用双生成策略构建,该策略以一个弱小型的边缘规模学生模型为信号进行针对性教师蒸馏:将学生模型的失败转化为纠正性讲解,同时将其成功之处扩展为对所有答案选项的对比式选项级推理。我们进一步引入了LLM-as-a-parser(大语言模型作为解析器)评估协议,从自由格式的输出中提取最终答案,并同时追踪准确率和答案有效率。为验证QVAC Genesis III数据的有效性,我们使用17亿参数模型进行了受控的从零训练消融实验,结果表明使用QVAC Genesis III训练的模型在ARC、GPQA Diamond和MMLU STEM基准上始终优于使用开源合成语料库Cosmopedia-v2训练的模型以及公开发布的Cosmo-1B模型,在ARC-E上最高提升28.57%,在ARC-C上最高提升21.35%,同时有效答案率(Valid Answer Rate)最高达到99.45%。
cs.AI / 14 / 2609.19515

LLM-as-an-Improver: Turning Verification into Better Candidates

LLM作为改进者:将验证转化为更优候选解
Tomihari, Akiyoshi, Ichikawa, Yuma
Abstract
Verifier-based selection improves LLM performance by generating multiple candidate solutions and using a verifier to select the most promising one. However, existing methods typically treat verification only as a ranking step and discard its feedback once a fixed candidate pool has been evaluated. In this paper, we ask whether verification can also improve the candidate set itself. To this end, we introduce LLM-as-an-Improver and propose Verify--Repair--Reselect (VRR), which uses verification feedback to generate and reselect improved candidates. VRR retains the initial winner while conditionally generating three complementary alternatives: repaired versions of the winner and runner-up, and a solution based on a new approach. It filters invalid and duplicate candidates using only inference-time information and then reselects the final answer under the original evaluation criteria. Across diverse models and code-generation and reasoning benchmarks, VRR improves over fixed-pool verifier-based selection in many settings and can recover correct solutions even when all candidates in the initial pool are incorrect. These results highlight a broader role for LLMs as improvers: verification feedback can not only select among existing solutions but also construct stronger candidates beyond the initial pool.
Chinese Translation
基于验证器的选择方法通过生成多个候选解并使用验证器挑选最有希望的解来提升大语言模型(LLM)的性能。然而,现有方法通常仅将验证作为排序步骤,一旦对固定的候选池完成评估便丢弃其反馈。本文探讨验证是否也能改进候选集本身。为此,我们提出了 LLM-as-an-Improver 框架,并设计了 Verify--Repair--Reselect(VRR)方法,利用验证反馈生成并重新选择改进后的候选解。VRR 在保留初始最优解的同时,有条件地生成三个互补的替代方案:最优解和次优解的修复版本,以及基于新思路的一个解。它仅利用推理时的信息过滤无效和重复的候选解,然后按照原始评估标准重新选择最终答案。在多种模型以及代码生成和推理基准上,VRR 在许多设定下优于基于固定候选池的验证器选择方法,并且即使初始候选池中的所有候选解都不正确时,也能恢复出正确的解。这些结果凸显了 LLM 作为改进者的更广泛作用:验证反馈不仅可以在现有解中进行选择,还能构建超越初始候选池的更强候选解。
cs.AI / 15 / 2609.19519

An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence

长时程智能体的架构:层级、节拍与级联智能
Nijkamp, Erik, Koul, Anurag, Pakhomov, Egor, Pang, Bo
Abstract
Language-model agents are increasingly asked to carry out work spanning days or weeks, such as an operations remediation or a research programme. Such a task outlives any context window, any process and any interval at which a person can attend. In this paper, we argue that a long-horizon agent must run continually without forgetting before it can learn continually. This ability lies in the harness around the model rather than in the model itself. We derive seven bottlenecks from the long-horizon setting and answer them with a hierarchical architecture of three parts: (i) levels indexed by time scale, each keeping a bounded file summarising the level below; (ii) a clocked tick as the unit of autonomous action; and (iii) cascaded intelligence, where work is escalated to a more capable model only after failing review. We report on a ten-day campaign in which an agent built on this architecture reproduced a published reinforcement-learning result with a human attending once a day, and show (1) the agent kept the thread across every context reset and session boundary of the campaign, (2) operating knowledge written early changed later behaviour with no change to model weights, and (3) where learned components would enter such a system. Overall, our experience suggests continual learning for these agents needs a substrate outliving every context and process, and the checks the harness already runs are where a learner belongs.
Chinese Translation
语言模型智能体正日益被要求执行跨越数天或数周的工作,例如运维修复或研究项目。这类任务的持续时间超出了任何上下文窗口、任何进程以及任何人能够关注的间隔。在本文中,我们论证了一个长时程智能体必须能够在持续运行中不遗忘,然后才能实现持续学习。这种能力的关键在于模型周围的运行框架(harness),而非模型本身。我们从长时程设定中推导出七个瓶颈,并用一个由三部分组成的分层架构加以解决:(i) 按时间尺度索引的层级(levels),每层维护一个有界的文件以总结下一层内容;(ii) 作为自主行动单元的定时节拍(tick);(iii) 级联智能(cascaded intelligence),即只有在工作未通过审查时,才将其升级到更强的模型。我们报告了一个为期十天的实验:基于该架构构建的智能体在人类每天仅需关注一次的情况下,复现了一项已发表的强化学习结果,并展示了:(1) 智能体在实验的每一次上下文重置与会话边界中都保持了任务线索的连续性;(2) 早期写入的操作知识在不改变模型权重的情况下改变了后续行为;(3) 学习型组件应进入此类系统的位置。总体而言,我们的经验表明,这类智能体的持续学习需要一个超越所有上下文和进程的底层载体,而运行框架已经执行的检查正是学习器应当所在之处。
cs.AI / 16 / 2609.19523

EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data

EconSkills:研究网页智能体在实时经济数据上的技能迁移与检索
Quan, Yinzhu, Liu, Zefang
Abstract
Web agents often revisit the same sites, yet most evaluations discard the procedures learned in earlier successful interactions. We introduce EconSkills, a skill library and evaluation framework that distills verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data. Each skill records its scope, navigation procedure, site-specific guidance, verification checks, and recovery steps while replacing source-instance values with placeholders. EconSkills separates two questions: whether a known relevant procedure transfers to a held-out task, and whether an agent can retain that benefit when selecting from a library. In controlled transfer, matched skills improve success over no-skill prompting and require fewer steps on paired successes, while abstraction is substantially more effective than replaying raw trajectories. At library scale, retrieval is competitive with the no-skill baseline overall and performs best on directly covered tasks; coverage-stratified outcomes show that approximate matches on uncovered tasks offset these gains. Browser trajectories further identify when procedural guidance shortens portal-specific navigation and when semantic verification remains necessary. These results establish that reusable economic web procedures can transfer across task instances and provide a concrete design target for coverage-aware selection and context delivery.
Chinese Translation
网页智能体经常会重复访问相同的网站,但大多数评估方法会丢弃在早期成功交互中学到的操作流程。我们提出了EconSkills,这是一个技能库和评估框架,它将经过验证的EconWebArena轨迹提炼为参数化的标准操作流程,用于检索实时经济数据。每个技能记录其适用范围、导航流程、特定网站的指导信息、验证检查和恢复步骤,同时用占位符替换源实例中的具体数值。EconSkills区分了两个问题:已知的相关流程能否迁移到留出的任务上,以及智能体在从技能库中选择时能否保持这一优势。在受控迁移实验中,匹配的技能相比无技能提示提升了成功率,并且在配对成功案例中所需步骤更少,而抽象化的技能比直接重放原始轨迹有效得多。在技能库规模上,检索的总体表现与无技能基线相当,并在直接覆盖的任务上表现最佳;按覆盖率分层的结果显示,在未覆盖任务上的近似匹配抵消了这些收益。浏览器轨迹进一步揭示了程序性指导何时能缩短特定门户的导航过程,以及何时仍然需要语义验证。这些结果表明,可复用的经济网页操作流程能够跨任务实例迁移,并为覆盖感知的选择与上下文交付提供了具体的设计目标。
cs.AI / 17 / 2609.19524

A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

面向可信大语言模型、智能体AI与多模态系统的统一评估框架
Raza, Shaina, Radwan, Ahmed Y., Liaquat, Imran, Hume, Kathryn
Abstract
Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems. Large language models (LLMs), agentic systems, and multimodal models (MLLMs) require different forms of assessment, yet their evaluation evidence must remain interpretable for development and oversight. We propose a unified framework that connects output-level, trajectory-level, and cross-modal assessment through eight trustworthiness dimensions: capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency. The framework preserves system-specific metrics while mapping native measurements to common performance bands, accompanied by uncertainty estimates and traceable evidence. A meta-evaluation layer examines the validity, reliability, and reproducibility of the evaluation itself. Multidimensional profiles expose strengths and weaknesses, while safety-critical overrides prevent aggregate scores from masking critical failures. Mappings to governance frameworks, international standards, and European Union regulatory requirements connect technical assessment with oversight needs. The framework provides a structured basis for assessing both system performance and the credibility of the evidence supporting it, with empirical validation across deployment contexts remaining an essential next step.
Chinese Translation
仅凭基准测试分数不足以全面评估现代人工智能系统的可信赖性。大语言模型(LLM)、智能体系统和多模态模型(MLLM)需要不同形式的评估,但其评估证据必须对开发与监督保持可解释性。我们提出了一个统一框架,通过八个可信赖性维度——能力、鲁棒性、安全性、公平性、透明性、治理、监督和效率——将输出级、轨迹级和跨模态评估联系起来。该框架在保留系统特定指标的同时,将原生测量结果映射到统一的性能区间,并辅以不确定性估计和可追溯的证据。元评估层(meta-evaluation layer)用于检验评估本身的有效性、可靠性和可复现性。多维度画像能够揭示系统的优势与不足,而安全关键型覆盖机制可防止综合评分掩盖关键性失败。该框架与治理框架、国际标准以及欧盟监管要求建立了映射,从而将技术评估与监督需求相衔接。该框架为评估系统性能及其支撑证据的可信度提供了结构化基础,而在实际部署场景中的实证验证仍是必不可少的下一步工作。
cs.AI / 18 / 2609.19526

Self Improvement via Fast Tree-search

基于快速树搜索的自我改进
Fu, Xinghong, Kulanthaivelu, Aravinth, Yamada, Yutaro
Abstract
Coding agents can recursively modify their own implementations, forming a loop of self-improvement. While prior work shows this can boost performance on coding benchmarks, existing approaches are costly and compute-intensive. We introduce a simple, sample-efficient self-improvement framework that significantly improves coding performance under strict budget constraints. We identify evaluation of candidate self-modifications as the main runtime bottleneck since prior approaches estimate their effectiveness by re-running a subset of benchmark tasks with the modified agent, which is time-consuming. We introduce Recursive Self Improvement via Fast Tree-search (SIFT), which augments these downstream task evaluations with an LLM-as-a-judge signal that performs pairwise comparisons between candidate patches, where the win-loss record is aggregated with a regularized Bradley-Terry model, and the resulting strength scores drive rank-based parent sampling inside a lightweight disaggregated tree search. Expensive downstream task evaluations are reserved only for the most promising nodes. Using a fully disaggregated tree search pipeline, the judge scores provide intermediate signal to guide exploration on promising candidate patches without being bottlenecked by slow evaluation runs. SIFT outperforms existing tree-search based self-evolution frameworks on the full Polyglot benchmark with significantly lower resource requirements in terms of CPU hours, wall clock time, and API cost.
Chinese Translation
编程智能体(coding agents)可以递归地修改自身的实现,从而形成自我改进的循环。虽然先前的研究表明这可以提升在编程基准测试上的表现,但现有方法成本高昂且计算密集。我们提出了一种简单且样本高效的自我改进框架,能够在严格的预算约束下显著提升编程性能。我们发现,候选自我修改的评估是主要的运行时瓶颈,因为先前的方法需要通过让修改后的智能体重新运行部分基准任务来估计其有效性,这非常耗时。我们提出了基于快速树搜索的递归自我改进方法(SIFT),该方法在下游任务评估之外引入了一种“大语言模型作为裁判”(LLM-as-a-judge)的信号,用于在候选补丁之间进行成对比较,其胜负记录通过正则化的Bradley-Terry模型进行聚合,所得到的实力分数用于驱动轻量级非聚合树搜索中基于排名的父节点采样。昂贵的下游任务评估仅保留给最有前景的节点。通过完全非聚合的树搜索流水线,裁判分数提供了中间信号,以引导对有前景的候选补丁的探索,而不会因缓慢的评估运行而受到瓶颈限制。在完整的Polyglot基准测试中,SIFT在CPU小时、实际运行时间和API成本等资源需求显著更低的情况下,超越了现有的基于树搜索的自我进化框架。
cs.AI / 19 / 2609.19530

When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent R\'esum\'e Screening

当招聘成为代理中介化过程:双代理简历筛选中的通过率与复现性评估
Gao, Jian, Jiang, Hang
Abstract
Hiring is bilateral: employers assess fit, while candidates present and defend evidence of their qualifications. Yet r\'esum\'e screening, the first gate, is commonly automated as a static, one-call judgment over a r\'esum\'e-job pair. We study a two-agent alternative in which employer-side and candidate-side agents represent these roles, exchange evidence, and update their judgments before deciding who advances. We compare procedures on 600 constructed r\'esum\'e-job pairs using GPT-5.5 and Claude Opus 4.7. Two-agent screening advances more applications (33.3% to 39.3% for GPT-5.5; 34.0% to 35.5% for Opus 4.7). Across three runs on the common 191-pair borderline pool, pass-instance rates rise from 4.5% to 26.2% and from 6.5% to 16.1%, respectively. This is not a uniform relaxation: two-agent screening rejects applications one-call advances, changing decisions in both directions. At similar pass volumes, the procedures advance different applications, and no one-call threshold recovers applications consistently selected by two-agent screening. Among discovery-selected cases re-executed in fresh runs, two-agent-only selections recur less often than shared selections, clearly under GPT-5.5 and less certainly under Opus 4.7, while a separate one-call follow-up shows no comparable decline. As hiring becomes agent-mediated on both sides, the screening procedure, not only the model behind it, shapes who reaches human review and how reliably that access recurs.
Chinese Translation
招聘是双向的:雇主评估匹配度,而候选人则呈现并论证其资质的证据。然而,作为第一道关口的简历筛选,通常被自动化为对简历-职位对的静态、单次调用判断。我们研究了一种双代理替代方案,其中雇主方代理与候选人方代理分别代表这两种角色,交换证据,并在决定谁进入下一轮之前更新各自的判断。我们使用GPT-5.5和Claude Opus 4.7,在600组构建的简历-职位对上比较了不同流程。双代理筛选推进了更多的申请(GPT-5.5从33.3%提升至39.3%;Opus 4.7从34.0%提升至35.5%)。在共享的191组边缘申请池上进行的三次运行中,通过实例率分别从4.5%上升至26.2%,以及从6.5%上升至16.1%。这并非简单的普遍放宽:双代理筛选会拒绝单次调用流程所推进的申请,决策在两个方向上都发生了改变。在相近的通过数量下,两种流程推进的申请并不相同,且没有任何单次调用阈值能够复现双代理筛选始终选中的申请。在独立新运行中重新执行的被发现选中的案例中,仅由双代理筛选选中的案例的复现率低于两者共同选中的案例,这一现象在GPT-5.5下十分明显,而在Opus 4.7下则不那么确定;与此同时,单独的单次调用后续实验未显示出类似的下降。当招聘在双方都变为代理中介化时,决定谁能进入人工审核以及这种通过机会复现可靠程度的,是筛选流程本身,而不仅仅是其背后的模型。
cs.AI / 20 / 2609.19538

Agentic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks

面向低空无线网络中异构无人机系统的智能体式人工智能组网
Quang, Nguyen Duc Minh, Liu, Chang, Li, Shuangyang, Ng, Derrick Wing Kwan
Abstract
Low-altitude wireless networks (LAWNs) are emerging as a key infrastructure for heterogeneous unmanned aerial systems that support concurrent services within a shared three-dimensional airspace. Their coexistence creates strong coupling among mobility, connectivity, and shared network resources, while heterogeneous services impose distinct and time-varying requirements. These interactions naturally form a dynamic non-cooperative game in which both operating conditions and coordination objectives evolve over time. Conventional optimization and learning-based controllers typically rely on predefined objectives, limiting their ability to adapt autonomously to changing service requirements and resource priorities. To address this challenge, we propose a hierarchical hybrid large language model (LLM)- multi-agent reinforcement learning (MARL) architecture organized as a dual-loop structure. Specifically, an outer adaptation loop employs LLM-assisted game orchestration to interpret service requirements and operator intent, and reconfigure objectives and resource priorities, while an inner loop executes decentralized, parameter-conditioned MARL policies under the configured game. A logistics-monitoring case study illustrates how the proposed framework facilitates coordinated coexistence among heterogeneous services, adapting to evolving operating conditions without retraining the underlying MARL policies. Finally, we discuss key challenges and research directions toward scalable, trustworthy, and adaptive agentic LAWNs.
Chinese Translation
低空无线网络(LAWNs)正在成为支持共享三维空域内并发业务的异构无人机系统的关键基础设施。多系统的共存使移动性、连通性与共享网络资源之间形成了强耦合,而异构业务又提出了各不相同且随时间变化的需求。这些交互自然地构成了一个动态非合作博弈,其运行条件与协调目标均随时间演化。传统的优化方法及基于学习的控制器通常依赖于预定义的目标,限制了其自主适应不断变化的业务需求和资源优先级的能力。为应对这一挑战,我们提出了一种分层混合大语言模型(LLM)-多智能体强化学习(MARL)架构,采用双环路结构组织。具体而言,外层适应环路利用LLM辅助的博弈编排来解析业务需求和运营方意图,并重新配置目标与资源优先级;内层环路则在所配置的博弈下执行去中心化的、参数条件化的MARL策略。一个物流监控案例研究展示了所提框架如何促进异构业务之间的协调共存,能够在无需对底层MARL策略进行重新训练的情况下适应不断变化的运行条件。最后,我们讨论了迈向可扩展、可信赖且自适应的智能体式低空无线网络的关键挑战与研究方向。
cs.AI / 21 / 2609.19551

Continual Enterprise World Model Discovery in Dynamic Systems

动态系统中的持续企业世界模型发现
Mishra, Shambhavi, Vazquez, David, Taslakian, Perouz, Pedersoli, Marco, Dolz, Jose, Laradji, Issam H.
Abstract
In an enterprise system, updating one field can set another, create a record, or start an approval. These effects are produced by business rules that are not built into the platform but written by each organization and revised over time. An agent working in such a system cannot predict the result of its own actions without knowing these rules. We study continual enterprise world model discovery, where an agent starts without knowledge of these business rules and discovers them by interacting with records and observing the outcomes. From those observations it builds a world model, which it revises as the rules change. To evaluate this, we introduce EnterpriseWorldShift, built on a live ServiceNow environment with nine tables, 25 hidden rules and 600 evaluation actions. It presents four versions of the same enterprise world, with the tables and records held fixed while a rule is modified, then added, then removed, so that discovery, revision, extension and retirement are each tested in turn. Our Continual Discovery Agent (CDA) builds such a model and carries it from one world to the next. It predicts the effects of the hidden rules more accurately than looking them up for each question, the approach taken by prior work, by up to 8.98 IoU points, and it answers from its own model without querying the running system.
Chinese Translation
在企业系统中,更新某一字段可能会设置另一个字段、创建一条记录或启动一个审批流程。这些效应由业务规则产生,而这些规则并非内置于平台之中,而是由各组织自行编写并随时间不断修订。在这样的系统中工作的智能体(agent),若不了解这些规则,就无法预测自身操作的结果。我们研究持续企业世界模型发现问题:智能体最初不了解这些业务规则,通过与记录交互并观察结果来发现它们,并基于这些观察构建世界模型,且随着规则的变化不断修订该模型。为评估这一任务,我们提出了EnterpriseWorldShift,它构建于一个真实的ServiceNow环境之上,包含九张表、25条隐藏规则和600个评估操作。该基准呈现同一企业世界的四个版本:在保持表和记录不变的前提下,先修改一条规则,再新增,再删除,从而依次测试发现、修订、扩展和退役能力。我们的持续发现智能体(Continual Discovery Agent, CDA)构建了这样的世界模型,并将其从一个世界带向下一个世界。相比先前工作所采用的针对每个问题进行查询的方法,它在预测隐藏规则的效果上准确率最高可提升8.98个IoU点,且无需查询运行中的系统,即可依靠自身模型作答。
cs.AI / 22 / 2609.19610

SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership

SIMLIFE:面向长时程人-智能体协作的模式理解
Peng, Run, Nie, Zinnia, Ding, Jing, Dai, Yinpei, Zhang, Yichi, Wu, Zengqing, Fu, Yao, Ma, Ziqiao, Mao, Jiayuan, Chai, Joyce
Abstract
Understanding humans over long horizons requires agents to infer not only what people need in the moment, but also how routines form, why they repeat, and when they change. We introduce SimLife, a scalable platform for simulating long-term household life with rich visual observations, ground-truth action logs, and synthetic dialogues with audio. Built on SimLife, SimLife-BP evaluates long-context pattern understanding: the ability to infer latent behavioral rules from weeks or months of everyday observations. The benchmark contains 106 episodes averaging 15.49 hours and 38.57 in-game days, and 1,439 question-answer pairs. Each task probes direct, counterfactual, noisy, and inverse reasoning under different levels of rule hints. Evaluating frontier models and architectures, we find that current models often achieve surface-level prediction without comprehensive rule understanding, rely on frequency-based heuristics rather than if-then reasoning over evidence, and struggle to adapt when behavioral patterns change. These findings suggest that long-context pattern understanding remains a major bottleneck for future embodied agents, while SimLife opens a broader space for studying memory, personalization, adaptation, and long-horizon planning in everyday human-AI interaction.
Chinese Translation
要在长时程上理解人类,智能体不仅需要推断人们当下的需求,还需要理解日常习惯如何形成、为何重复以及何时发生改变。我们提出了SimLife,一个可扩展的长期家庭生活仿真平台,提供丰富的视觉观测、真实动作日志以及带音频的合成对话。基于SimLife构建的SimLife-BP用于评估长上下文模式理解能力,即从数周或数月的日常观察中推断潜在行为规则的能力。该基准包含106个片段,平均时长15.49小时、相当于游戏内38.57天,以及1,439个问答对。每个任务在不同程度的规则提示下,考察直接推理、反事实推理、含噪推理和逆向推理。通过对前沿模型与架构的评估,我们发现当前模型往往只能实现表面层面的预测,而缺乏对规则的全面理解;它们依赖于基于频率的启发式方法,而非基于证据的"如果-那么"推理;并且在行为模式发生变化时难以适应。这些发现表明,长上下文模式理解仍然是未来具身智能体的一大瓶颈,而SimLife为研究日常人机交互中的记忆、个性化、适应以及长时程规划开辟了更广阔的空间。
cs.AI / 23 / 2609.19630

From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization

从意图到行动:车辆语音命令授权中大模型安全性的基准测试
Afroze, Diba, Zhang, Xingli, Tu, Yazhou, Hei, Xiali
Abstract
Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking natural-language requests to vehicle functions creates a safety-critical authorization problem. Before executing a command, the system must choose whether to execute, refuse, clarify, require confirmation, defer to manual control, trigger an emergency response, or make no tool call. To our knowledge, prior evaluations do not isolate this pre-action decision across speaker role, authentication status, vehicle state, and tool availability. We introduce a 202-scenario benchmark with Reference Decisions under a seven-class taxonomy. We evaluate two local open-weight models and three API-based LLMs using Decision Alignment and safety-specific error metrics. Alignment ranges from 40.1% for Llama 3.2 3B to 89.1% for Gemini 3.1 Pro Preview. The API-based models score between 83.2% and 89.1%, with no statistically significant differences among them. Even these models produce two to three False Executes among 161 non-execution scenarios, and persistent errors remain in confirmation and manual-control decisions. A controlled Llama 3.2 3B ablation increases alignment to 40.1% under the structured authorization policy, versus 28.2-29.2% under schema-only and generic-safety baselines, but it does not eliminate False Executes. Structured LLM decisions are therefore insufficient as a standalone safety mechanism, and deployment requires an independent enforcement layer that verifies tool permissions and vehicle-state constraints before invoking any vehicle function.
Chinese Translation
大语言模型(LLM)正日益被集成到车辆语音助手中。但将自然语言请求与车辆功能关联起来,带来了一个安全攸关的授权问题。在执行命令之前,系统必须决定是执行、拒绝、澄清、要求确认、转由人工控制、触发紧急响应,还是不进行任何工具调用。据我们所知,先前的评估并未在说话者角色、认证状态、车辆状态和工具可用性等维度上隔离这一行动前决策。我们引入了一个包含202个场景、并在七类分类体系下附带参考决策的基准测试。我们使用决策一致性(Decision Alignment)和针对安全性的错误指标评估了两个本地开源权重模型和三个基于API的LLM。一致性得分从Llama 3.2 3B的40.1%到Gemini 3.1 Pro Preview的89.1%不等。基于API的模型得分介于83.2%至89.1%之间,且彼此之间无统计学显著差异。即便是这些模型,在161个非执行场景中也出现了两到三次误执行(False Executes),且在确认和人工控制决策方面仍存在持续性错误。对Llama 3.2 3B的受控消融实验表明,在结构化授权策略下,一致性提升至40.1%,而在仅使用模式(schema-only)和通用安全基线时仅为28.2%至29.2%,但并未消除误执行。因此,结构化的LLM决策作为独立的安全机制是不够的,部署时需要一个独立的执行层,在调用任何车辆功能之前验证工具权限和车辆状态约束。
cs.AI / 24 / 2609.19636

Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs

到达还是解决?基于检查点交接的智能体强化学习收益归因
Liu, Xuan, Qian, Jingbin
Abstract
Reinforcement learning now trains language-model agents that act over dozens of steps in live environments. The gains are large, and they are read as better decision-making. An agent in a closed loop writes its own inputs. Each observation follows from its own earlier actions, so the states it meets late in an episode are partly of its own making. An SFT checkpoint and an RL checkpoint are then scored from different states, even on identical tasks. Endpoint success mixes two changes: where the agent arrives, and what it does once it is there. Restricting the comparison to states both policies reach does not separate them. That restriction selects on an outcome, and in our data it flips the sign of the effect. We introduce checkpoint handoff, an evaluation protocol that clones a state one released checkpoint reached and hands it to another, with no retraining. Crossing a reacher role and a solver role over SFT and RL splits an endpoint gain into REACH and SOLVE. REACH is how often a policy arrives at a state the environment confirms is a fixed number of actions from success. SOLVE is how often it finishes from an identical cloned state. Across two benchmarks and two independently released pipelines, the reacher by solver interaction is positive in all five conditions. An RL history is worth more to an RL solver than the same history is to an SFT solver. On ALFWorld, RL improves both terms, and the SFT solver never succeeds where the RL solver fails. Independent REACH and SOLVE gaps predict the aggregate interaction. Handoff asks only that one checkpoint's history can be replayed under another, so long-horizon evaluation can report arrival and completion beside endpoint success.
Chinese Translation
强化学习如今可以训练在真实环境中执行数十个步骤的语言模型智能体。其收益巨大,且通常被解读为决策能力的提升。处于闭环中的智能体会自己生成输入,每个观测都源于其先前的动作,因此其在回合后期所处的状态部分是由自身造成的。这样,即使面对相同的任务,SFT检查点与RL检查点也是从不同的状态被评估的。终点成功率混淆了两类变化:智能体到达了何处,以及到达后做了什么。仅将比较限定在两种策略都能到达的状态上并不能将二者分离,因为这种限定是基于结果的选择,而在我们的数据中它会翻转效应的正负号。我们提出检查点交接这一评估协议:克隆某个已发布检查点所到达的状态,并将其交给另一个检查点,且无需重新训练。通过在SFT和RL之间交叉配置“到达者”和“解决者”两种角色,可将终点收益分解为REACH和SOLVE两部分。REACH衡量策略到达某一状态的频率,该状态被环境确认为距成功仅剩固定步数;SOLVE衡量其从相同的克隆状态完成任务的频率。在两个基准和两条独立发布的训练流程上,到达者与解决者的交互效应在全部五个条件下均为正:同样的历史状态对RL解决者的价值高于对SFT解决者的价值。在ALFWorld上,RL同时提升了两个指标,且凡RL解决者能成功的状态,SFT解决者也从不失败。独立的REACH和SOLVE差距可以预测总体交互效应。交接仅需一个检查点的历史状态可在另一个检查点下重放,因此长程评估可以在终点成功率之外同时报告到达与完成情况。
cs.AI / 25 / 2609.19644

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

ScientistTwo:以自主人工智能开拓人类知识前沿
Nam, Jaehyun, Yoon, Jinsung, Pan, Yanzhou, Wang, Yubo, Meng, Rui, Ranganathan, Parthasarathy, Pfister, Tomas
Abstract
Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory. The ultimate vision for AI in science is problem-driven autonomous research: given a fundamental challenge by a human expert, the AI independently navigates the scientific landscape, uncovers theoretical and empirical bottlenecks, and systematically expands the frontier of knowledge. In this paper, we introduce ScientistTwo, a fully autonomous multi-agent framework designed to realize this vision. Specifically, ScientistTwo takes an initial problem as input, establishes state-of-the-art baselines, formulates novel hypotheses, and coordinates specialized agents to orchestrate an end-to-end discovery cycle without human intervention. Moreover, the framework rigorously conducts experiments using diverse datasets and metrics, refines methodologies through automated ablation studies, and validates research findings via a closed-loop simulated peer-review rebuttal engine. To evaluate ScientistTwo's capabilities against the highest standards of human scientific achievement, we benchmark it across papers accepted at top-tier conferences such as ICLR, ICML, and NeurIPS. As a result, ScientistTwo autonomously generates expert-level, publishable papers and fully verified, executable codebases. Its solutions consistently outperform human state-of-the-art models, and achieve higher average review ratings than human-authored papers under automated AI review agents. These results show that ScientistTwo is not merely an assistive tool but an autonomous scientific pioneer capable of pushing the frontiers of human discovery. Project website: https://scientist-two.github.io/
Chinese Translation
科学发现的本质在于识别现有知识的边界并探索未知领域。人工智能在科学领域的终极愿景是问题驱动的自主研究:在人类专家提出一个基础性挑战后,AI能够独立地探索科学图景,发现理论与实证瓶颈,并系统性地拓展知识前沿。本文提出ScientistTwo,一个旨在实现这一愿景的全自主多智能体框架。具体而言,ScientistTwo以一个初始问题作为输入,建立最先进的基线,提出新颖的假设,并协调专门的智能体以在无人干预的情况下编排端到端的发现循环。此外,该框架能够严格地使用多样化的数据集和指标进行实验,通过自动化消融研究改进方法,并通过闭环模拟同行评审答辩引擎验证研究发现。为了以人类科学成就的最高标准评估ScientistTwo的能力,我们在ICLR、ICML和NeurIPS等顶级会议录用的论文上对其进行基准测试。结果表明,ScientistTwo能够自主生成专家级、可发表的论文以及经过完整验证的可执行代码库。其解决方案始终优于人类最先进的模型,并且在自动化AI评审智能体下获得了比人类撰写的论文更高的平均评审评分。这些结果表明,ScientistTwo不仅仅是一个辅助工具,而是一个能够推动人类发现前沿的自主科学开拓者。项目网站:https://scientist-two.github.io/
cs.AI / 26 / 2609.19654

Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions

重新规划、修复还是编辑?资源中断下行程修订旅行智能体的统一实证评估
Yuan, Xiaofei, Zhang, Yan, Qiao, Shaobo, He, Huangleshuai, Ni, Leyan, Ju, Mingchen, Yang, Lujia, Xu, Sijia, Tang, Yifu, Yang, Zhengyi
Abstract
Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures. Revising these itineraries involves full replanning, classical plan repair, and LLM-based travel-agent revision, whose differing task formulations and evaluation protocols hinder comparison. We conduct a systematic empirical study using two TREK-derived benchmark sets: 500 single-disruption cases, including feasible and infeasible instances, and 200 feasible simultaneous compound-disruption cases. We compare LLM-Z3 full replanning, IPyHOPPER hierarchical repair, and an iTIMO local-revision adapter across effectiveness, plan stability, and computational cost. LLM-Z3 with Gemini achieved the highest observed compound-disruption success. IPyHOPPER nearly matched that configuration's single-disruption overall success, while preserving substantially more of the accepted itinerary on successful repairs. Successful hierarchical and local repairs made fewer edits and retained more accepted commitments than full replanning. Computational profiles differed: IPyHOPPER used no LLM inference, the evaluated LLM-Z3 adapter used compact one-call inference, and the iTIMO adapter consumed substantially more tokens. The study provides practical guidelines for balancing feasibility recovery, commitment preservation, and computational cost within evaluated settings.
Chinese Translation
旅行规划智能体生成的行程在被接受后,可能因航班取消、酒店不可用或景点关闭而变得不可行。对这些行程进行修订涉及完全重新规划、经典计划修复以及基于大语言模型(LLM)的旅行智能体修订,但三者不同的任务形式化定义和评估协议阻碍了相互比较。我们使用两个源自 TREK 的基准数据集开展了系统性实证研究:包含可行与不可行实例的 500 个单次中断案例,以及 200 个可行的同时复合中断案例。我们在有效性、计划稳定性和计算成本三个维度上比较了 LLM-Z3 完全重新规划、IPyHOPPER 分层修复和 iTIMO 局部修订适配器。基于 Gemini 的 LLM-Z3 在复合中断中取得了最高的成功率。IPyHOPPER 在单次中断总体成功率上几乎与该配置持平,同时在成功修复时保留了更多已接受的行程内容。与完全重新规划相比,成功的分层修复和局部修复进行了更少的修改,并保留了更多已接受的承诺。计算特征各不相同:IPyHOPPER 不使用 LLM 推理,所评估的 LLM-Z3 适配器采用紧凑的单次调用推理,而 iTIMO 适配器则消耗了多得多的 token。本研究为在所评估的设置中权衡可行性恢复、承诺保留和计算成本提供了实用指南。
cs.AI / 27 / 2609.19671

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

When2Think:面向高效混合推理模型的难度感知长度控制学习方法
Shim, Jaejun, Kim, HyunJin, Kim, Young Jin, Bak, JinYeong
Abstract
Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rigid routing incur an efficiency tax, trading reduced computation on easy instances for accuracy loss on hard instances. We formulate efficient reasoning as an instance-adaptive computation allocation problem and propose When2Think, a post-training framework for hybrid reasoning that dynamically allocates computation based on problem difficulty. Our method introduces Instance-level Difficulty-Aware Control (IDAC), a reward-shaping mechanism that leverages pre-computed reference statistics (accuracy and token usage) to regulate reasoning depth. Combined with verifier-based rewards and batch-wise standardized advantages, IDAC enables stable critic-free optimization without learned reward models or online reference-model queries. When2Think encourages direct answering on easy instances while preserving extended reasoning on hard instances, thereby learning when to use System 1 (NoThink) versus System 2 (Think). Experiments on mathematical benchmarks demonstrate improved accuracy-efficiency trade-offs: on AIME24, Pass@3 increases by 10.0% while token usage is reduced by 27.9% relative to the base model, and on AIME25, When2Think achieves 40.0% Pass@3, outperforming compression and routing-only baselines.
Chinese Translation
大型推理模型(Large Reasoning Models, LRMs)在复杂任务上表现出色,但存在系统性的低效问题:它们往往在简单问题上过度思考,而在困难问题上思考不足。现有基于统一长度惩罚或刚性路由的方法会带来效率代价,即以困难实例的准确率损失换取简单实例上的计算量减少。我们将高效推理形式化为一个实例自适应的计算资源分配问题,并提出 When2Think,一个基于问题难度动态分配计算资源的混合推理后训练框架。我们的方法引入了实例级难度感知控制(Instance-level Difficulty-Aware Control, IDAC),这是一种奖励塑形机制,利用预先计算的参考统计信息(准确率和词元使用量)来调节推理深度。结合基于验证器的奖励和批次内标准化优势,IDAC 实现了无需学习奖励模型或在线参考模型查询的稳定无评论家(critic-free)优化。When2Think 鼓励模型在简单实例上直接作答,同时在困难实例上保留扩展推理,从而学习何时使用系统 1(NoThink)与系统 2(Think)。在数学基准上的实验表明,该方法改善了准确率与效率的权衡:在 AIME24 上,Pass@3 相对于基础模型提升了 10.0%,同时词元使用量减少了 27.9%;在 AIME25 上,When2Think 达到了 40.0% 的 Pass@3,优于仅压缩和仅路由的基线方法。
cs.AI / 28 / 2609.19680

FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

FINSKILLOPS:一个用于SEC文件问答的自进化多智能体系统
Ma, Yanzhang, Tai, Zhenghan, Wu, Hanwei, Guan, Sizhe, Lei, Jianliang, He, Hailin, Jiang, Chaolong, Chi, Jijun, Kwok, Tung Sum Thomas, Xiao, Bohuai, Tian, Jingrui, Wu, Xinlu, Zhan, Xingao, Lu, Peng, Li, Muzhi, Wu, Yihong, Ma, Liheng, Lyu, Sicheng, Yan, Tianshuo, Zhu, Junhao, Xu, Yaqian, Ding, Lei, Cui, Yufei, Liu, Ziquan, Han, Boyu, Liu, Hengli, Zhou, Ling, Wang, Xinyu
Abstract
Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter. In practice, new SEC-filing questions repeatedly expose heterogeneous errors in period, entity, evidence use, and calculation. Existing self-improvement methods can turn failures into new behaviors, but offer limited control over where a correction should apply or which previously correct answers it may break. We therefore frame post-deployment improvement as controlled behavioral maintenance: recurring failures should become scoped skill patches, and each patch should earn deployment with- out introducing regressions. We instantiate this view in FINSKILLOPS, a multi-agent system for SEC filing QA. FINSKILLOPS derives reusable skills from evidence-grounded, typed failure diagnoses and governs them through targeted validation, protected-case regression checks, negative controls, and versioned replacement or retirement. Across six financial QA benchmarks, a single frozen skill registry achieves the highest verdict-weighted correctness and reference consistency among the evaluated systems. Evolved skills raise correctness from 3.70 to 4.55 on our enhanced benchmark. In a separate 12-round operational study, only six of 33 proposed skills are promoted, while the monitoring non-correct rate falls from 20.0% to 12.5%. These results establish controlled skill scope, admission, and lifecycle management as the foundation for reliable self-improvement.
Chinese Translation
金融问答系统通常在部署前通过改进检索、提示或智能体协作来提升性能,此后其可靠性表现便固定不变。在实践中,新的SEC文件问题会反复暴露出在时间期间、实体、证据使用和计算方面的异构错误。现有的自我改进方法虽然能够将失败转化为新行为,但对于修正应应用于何处、以及可能破坏哪些原本正确的回答,控制能力有限。因此,我们将部署后的改进视为受控行为维护:重复出现的失败应转化为范围明确(scoped)的技能补丁,且每个补丁都应在通过验证后才能部署,同时不引入回归。我们将这一理念具体化为FINSKILLOPS,一个用于SEC文件问答的多智能体系统。FINSKILLOPS从基于证据、带类型标注的失败诊断中提取可复用的技能,并通过针对性验证、保护用例回归检查、阴性对照以及版本化的替换或退役机制对其进行治理。在六个金融问答基准上,单一冻结技能注册表在所评估的系统中取得了最高的加权判定正确率和参考一致性。演化出的技能将我们在增强基准上的正确率从3.70提升至4.55。在一项独立的12轮运行研究中,33个候选技能中仅有6个获得晋升,同时监控的非正确率从20.0%下降至12.5%。这些结果确立了受控的技能范围界定、准入和生命周期管理是可靠自我改进的基础。
cs.AI / 29 / 2609.19721

LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents

LearnActCoder:面向自适应临床编码智能体的角色感知错误记忆机制
Ghaffari, Meysam, Sen, Bhaskar, Sabetpour, Nasim, Fatehi, Nina, Agarwal, Animesh, Morato, Carlos
Abstract
Clinical coding agents repeatedly encounter the same failure modes, including unsupported codes, missed documented conditions, specificity errors, and procedure-coding convention mismatches. We introduce Learn-Then-Act, an inference-time adaptation framework that converts errors from a small labeled LEARN batch into a structured Mistake Knowledge Database (MistakeKDB). False-negative lessons are routed to a recall-oriented Coder, while false-positive lessons are routed to a precision-oriented Judge. We instantiate the framework in LearnActCoder, a Coder-Judge clinical coding pipeline with lookup-table grounding where available. On 150 matched MIMIC-III notes, structured MistakeKDB improves CPT F1 by 5.9 percentage points, while raw-example and reflection-style memories remain near the no-memory baseline; the ICD-9 improvement is not significant. On a matched MIMIC-IV cohort, memory shifts ICD-10 coding toward higher precision at a recall cost, leaving F1 statistically unchanged. Applying the same memory to 1,000 held-out MIMIC-III notes maintains a stable ICD operating point, providing scale/stability evidence. Overall, the results are consistent with structured, feedback-derived error memory being useful for adapting clinical coding behavior across cases without weight updates or changes to the underlying workflow. Absolute CPT/HCPCS performance remains low, and the system is evaluated retrospectively rather than in clinical deployment.
Chinese Translation
临床编码智能体反复遇到相同的失败模式,包括不支持的编码、遗漏已记录的疾病诊断、特异性错误以及操作编码规范不匹配等问题。我们提出“学习-行动”(Learn-Then-Act)框架,这是一种推理时自适应框架,能够将一个小的标注LEARN批次中的错误转化为结构化的错误知识库(Mistake Knowledge Database,MistakeKDB)。假阴性(漏报)的经验教训被路由至面向召回率的Coder(编码器),而假阳性(误报)的经验教训则被路由至面向精确率的Judge(判定器)。我们将该框架实例化为LearnActCoder——一个Coder-Judge临床编码流水线,并在可用时利用查找表进行接地(grounding)。在150份匹配的MIMIC-III病历上,结构化的MistakeKDB将CPT F1分数提升了5.9个百分点,而原始样例记忆和反思式记忆几乎与无记忆基线持平;ICD-9的提升则不显著。在一个匹配的MIMIC-IV队列上,记忆机制使ICD-10编码向更高精确率偏移,但以召回率为代价,F1在统计上无显著变化。将相同的记忆应用于1,000份留出的MIMIC-III病历时,系统保持了稳定的ICD运行点,提供了规模与稳定性方面的证据。总体而言,这些结果表明,结构化的、由反馈推导的错误记忆能够在无需权重更新或更改底层工作流程的情况下,有效地跨病例自适应调整临床编码行为。需要指出的是,CPT/HCPCS的绝对性能仍然较低,且该系统是在回顾性数据上评估的,尚未在临床部署环境中进行验证。
cs.AI / 30 / 2609.19754

AutoData: Agentic Search for Pre-training Data Selection

AutoData:面向预训练数据选择的智能体搜索
Meng, Yan, Srikanth, Dhruv, Zhao, Bingchen, Jiang, Zhengyao, Wu, Yuxiang
Abstract
LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-training data selection as heuristic engineering over per-document features, i.e., lexical statistics, categorical labels, and perplexity. We introduce AutoData, an agent that searches directly over executable selection algorithms. Unlike prior data mixture methods that optimise weights over a fixed set of domains, AutoData searches a richer program space of scoring, stratification, and stochastic selection rules, discovering feature interactions automatically by iteratively refining algorithms with validation feedback from a proxy model. Within an overnight search, AutoData discovers a selection algorithm that outperforms existing human-designed curation pipelines. Despite being searched only on this small proxy, the discovered recipe transfers to larger scales and improves the downstream metric CORE. These results suggest that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.
Chinese Translation
大语言模型(LLM)智能体近来展现出通过在执行反馈下编辑模型和训练代码来自动化机器学习工程的潜力。然而,数据在很大程度上仍处于这一智能体优化循环之外。我们将预训练数据选择界定为针对文档级别特征的启发式工程,即词法统计、类别标签和困惑度(perplexity)。我们提出了 AutoData,一个直接在可执行选择算法上进行搜索的智能体。与以往在固定领域集合上优化权重的数据配比方法不同,AutoData 搜索一个更丰富的程序空间,包括打分、分层和随机选择规则,并通过代理模型(proxy model)的验证反馈迭代地改进算法,从而自动发现特征之间的交互。在一次通宵搜索中,AutoData 发现的选择算法优于现有的人工设计的数据筛选流程。尽管仅在这个小型代理模型上进行搜索,所发现的方案仍可迁移到更大规模,并提升了下游指标 CORE。这些结果表明,数据工程可以被 treated 为一个智能体机器学习问题,从而将自主研究从模型和训练代码优化扩展到数据领域。
cs.AI / 31 / 2609.19759

Rethinking Multi-Agent Collaboration: When More Is Less

重新思考多智能体协作:何时多即是少
Yuan, Yishuo, Wu, Yibo, Zhang, Yihan, Sun, Minyuan, Li, Shenliang, Ma, Xinkai, Li, Yifan, Liu, Jiaheng
Abstract
The rapid advancement of large language models and single-agent harnesses has reshaped the landscape of autonomous systems, raising a critical question of when multi-agent collaboration offers genuine value. As individual agent capabilities continue to scale, multi-agent collaboration faces diminishing returns while incurring growing context overhead. Through systematic analysis, we delineate the capability boundaries of multi-agent collaboration relative to single-agent alternatives, showing that it confers systematic benefits specifically in long-horizon tasks with sparse dependencies, while single-agent harnesses remain superior in tightly coupled, sequential workflows. Building on these insights, we propose SAIGE, a lightweight multi-agent collaboration mechanism based on Semantic-Aware Incremental Graph Evolution. SAIGE models collaboration as a dynamically evolving graph, where nodes are agent instances spawned on demand and edges encode semantic dependencies established through content-based information retrieval. Experiments on long-horizon, complex task benchmarks show that SAIGE achieves a favorable trade-off between context efficiency and task performance, and that scaling the agent pool or deepening the recursion level does not consistently improve outcomes. Our findings suggest that multi-agent superiority is bounded by task structure rather than universal, and that more agents do not necessarily make a system more intelligent.
Chinese Translation
大型语言模型与单智能体框架的快速发展重塑了自主系统的格局,也引出了一个关键问题:多智能体协作何时才能真正提供价值。随着单个智能体能力的持续提升,多智能体协作面临收益递减,同时带来日益增长的开销。通过系统性分析,我们划定了多智能体协作相对于单智能体方案的能力边界,表明其仅在依赖关系稀疏的长时程任务中具有系统性优势,而在紧密耦合的顺序化工作流中,单智能体框架仍然更优。基于这些洞察,我们提出了SAIGE,一种基于语义感知增量图演化的轻量级多智能体协作机制。SAIGE将协作建模为动态演化的图,其中节点为按需生成的智能体实例,边则通过基于内容的信息检索编码语义依赖关系。在长时程复杂任务基准上的实验表明,SAIGE在上下文效率与任务性能之间取得了良好的权衡,且扩大智能体规模或加深递归层级并不能持续提升结果。我们的研究结果表明,多智能体的优越性受任务结构约束而非普遍成立,更多的智能体并不一定使系统更智能。
cs.AI / 32 / 2609.19770

TorchCraft: Unified binder design by inverting an all-atom structure predictor

TorchCraft:通过逆向利用全原子结构预测器实现统一的结合剂设计
TorchCraft Team, Liu, Yu, Shen, Zhouhanyu, Li, Zhengyi, Huang, Xikun, Liu, Jiaqi, Gao, Shuxian, Yu, Qilin, Qin, Xiayan, Zhang, Yucheng, Chen, Mingchen
Abstract
All-atom structure predictors model diverse molecular interactions, but using their learned structural priors for binder design remains challenging. Here we present TorchCraft, a unified binder-design framework that optimizes sequence logits through a frozen all-atom predictor. Implemented in TorchFold, TorchCraft combines confidence, contact, geometric, and sequence-prior objectives within a shared optimization procedure for minibinders, framework-conditioned VHHs, cyclic peptides, and ligand-binding proteins. Using pretrained AlphaFold 3 weights, TorchCraft generated representative minibinders and VHHs with experimentally measured binding across four targets in each format, without post hoc sequence redesign. Computational benchmarks further demonstrated the framework's applicability to cyclic peptides and ligand-conditioned pocket design. TorchCraft extends predictor inversion to multiple binder formats and molecular contexts, providing a common framework for reusing all-atom structural priors in design.
Chinese Translation
全原子结构预测器能够建模多样化的分子相互作用,但如何利用其学习到的结构先验进行结合剂(binder)设计仍具有挑战性。在此,我们提出TorchCraft,一个统一的结合剂设计框架,通过冻结的全原子预测器对序列logits进行优化。TorchCraft基于TorchFold实现,在共享的优化流程中结合了置信度、接触、几何和序列先验等多重目标,适用于迷你结合剂(minibinders)、框架条件化VHH、环肽以及配体结合蛋白。利用预训练的AlphaFold 3权重,TorchCraft生成了具有代表性的迷你结合剂和VHH,且每种格式均在四个靶点上通过实验验证了结合活性,无需事后的序列重设计。计算基准测试进一步证明了该框架在环肽和配体条件化口袋设计方面的适用性。TorchCraft将预测器逆向利用扩展至多种结合剂格式和分子情境,为在全原子结构先验基础上开展设计提供了一个通用框架。
cs.AI / 33 / 2609.19775

Integrating knowledge from case reports: a medical ontology based multimodal information system with structured summary

整合病例报告知识:一种基于医学本体的多模态信息系统及结构化摘要
Guo, Shuyu, Huang, Lan, Liu, Yichen, Ma, Hanbin, Bai, Tian
Abstract
Published medical case reports serve as a crucial medical information carrier, documenting discoveries in rare diseases, diagnostic methods, and innovative treatments. Despite the wealth of clinical knowledge in millions of case reports in the public medicine literature database (PubMed), accessing relevant information efficiently is hindered by the limitations of traditional keyword-based retrieval tools on unstructured and diverse case reports. To address the above issues, we introduce a comprehensive multimodal information system for case reports integrating structured clinical summaries of patients including medical images and biomedical named entities from 52949 open-access case reports published from 2000 to 2021. The multimodal essential information is organized in a well-structured medical ontology. Also, a powerful interface for searching and browsing case reports is designed to assist junior clinicians in retrieving cases effectively and improving the identification and diagnosis of rare diseases.
Chinese Translation
已发表的医学病例报告是重要的医学信息载体,记录了罕见疾病、诊断方法和创新治疗方案的发现。尽管公共医学文献数据库(PubMed)中数百万份病例报告蕴含丰富的临床知识,但传统的基于关键词的检索工具在处理非结构化且形式多样的病例报告时存在局限,阻碍了相关信息的高效获取。为解决上述问题,我们构建了一个面向病例报告的综合多模态信息系统,该系统整合了患者的结构化临床摘要,包括医学图像和生物医学命名实体,数据来源于2000年至2021年间发表的52949份开放获取病例报告。多模态的关键信息以结构良好的医学本体(medical ontology)形式组织。此外,我们设计了一个功能强大的病例报告搜索与浏览界面,以协助初级临床医生高效检索病例,提升对罕见疾病的识别与诊断能力。
cs.AI / 34 / 2609.19789

Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems

交易大厅中的传染:对抗性信号如何在多智能体交易系统中传播
Sua, Qi Rong, Dong, Junhao, Thai, Nguyen Duc, Wen, Yuqing, Tan, Cheston, Ong, Yew-Soon
Abstract
Multi-agent trading systems built on large language models (LLMs) are beginning to appear in quantitative finance, yet their robustness to adversarial inputs is largely unknown. We study the vulnerability of LLM trading stacks to black-box, input-only attacks that enter solely via admissible social-media feeds. We introduce the Generic Multi-Agent Trading System (GMATS), a framework that captures modern multiagent trading architectures and instantiate a class of black-box poisoning attackers that treat an LLM as a post generator and inject budget-constrained, plausibly benign social-media content into the analyst's evidence stream. We define contagion metrics that trace how adversarial content propagates through the stack, including belief-shift scores at analyst and coordinator layers and attack-clean deltas on standard backtest metrics. Experiments on a safe offline benchmark with historical market and social data show that even simple input-only attackers can materially degrade risk-return profiles, sharply reducing Sharpe ratios. At the same time, we find that suitably designed multi-agent topologies and coordinator prompts can dampen adversarial shocks and improve average robustness under identical poisoning budgets.
Chinese Translation
基于大语言模型(LLM)构建的多智能体交易系统正开始出现在量化金融领域,然而其对对抗性输入的鲁棒性在很大程度上仍属未知。我们研究了LLM交易系统对黑盒、仅输入攻击的脆弱性,这类攻击仅通过可准入的社交媒体信息流进入系统。我们提出了通用多智能体交易系统(Generic Multi-Agent Trading System, GMATS),该框架刻画了现代多智能体交易架构,并实例化了一类黑盒投毒攻击者:它们将LLM视为帖子生成器,向分析师的证据流中注入预算受限且看似良性的社交媒体内容。我们定义了追踪对抗性内容如何在系统中传播的传染度量指标,包括分析师层和协调器层的信念偏移分数,以及标准回测指标上的攻击-干净差值。在一个使用历史市场与社交数据的安全离线基准上的实验表明,即使是简单的仅输入攻击者也能实质性恶化风险-收益特征,显著降低夏普比率。与此同时,我们发现经过合理设计的多智能体拓扑结构和协调器提示词能够在相同的投毒预算下抑制对抗性冲击,并提升平均鲁棒性。
cs.AI / 35 / 2609.19820

Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy

通过参考策略在正则化自博弈中引导均衡选择
Leal, Luis
Abstract
Regularized self-play -- the family behind DeepNash's Stratego play -- drives a two-player zero-sum policy to a Nash equilibrium by best-responding to a slowly moving, entropy-regularized reference policy $\rho$. When the game has a polytope of value-equivalent equilibria, the regularizer silently breaks the tie: with a uniform reference it selects the maximum-entropy member, the I-projection of $\rho$ onto the Nash set. Can the reference be used to choose the equilibrium on purpose? On five exactly solvable games plus a 2-D polytope, with exact best responses and equivalence tests over independent seeds, anchoring the reference at a target member and refining steers self-play to that member with mean coordinate error 0.007 at median exploitability $5\times10^{-5}$, TOST-equivalent to the request within $\pm0.05$; the anchoring persists through refinement and follows the reference, not the initialization. Selection follows the reach-weighted I-projection (slope 0.969 [0.950, 0.987]). We report with equal emphasis where the story breaks: fixed off-manifold references cost 0.08-0.25 exploitability; stiff or flat families require a smaller mirror step, set by a pre-registered rule; boundary targets undershoot; curvature predicts where boundary saturation bites (rank correlation 0.90, p=0.037) while interior precision is curvature-independent. Table and MLP steering maps are equivalent within $\pm0.03$ at every target (30 seeds); matched control arms show attention's robust signature is excess seed variance, any systematic shift bounded at 0.018 and not significant. Against a best response the selection-robustness trade-off is degenerate: steering matters only against fixed, non-equilibrium opponents. The recipe -- anchor the reference at the desired member and refine -- reinterprets the KL anchor of RLHF-style RL as a selection knob, not only a stability leash.
Chinese Translation
正则化自博弈——DeepNash 玩 Stratego 背后的方法家族——通过对一个缓慢移动的熵正则化参考策略 $\rho$ 进行最优响应,将双人零和策略驱动至纳什均衡。当博弈存在由价值等价均衡构成的多面体时,正则项会隐式地打破平局:均匀参考策略会选择最大熵成员,即 $\rho$ 在纳什集上的 I-投影。那么,参考策略能否被有意地用于选择均衡?在五个可精确求解的博弈以及一个二维多面体上,采用精确最优响应并基于独立随机种子进行等价性检验,我们将参考策略锚定于目标成员并加以细化,使自博弈收敛到该成员,平均坐标误差为 0.007,可利用性中位数为 $5\times10^{-5}$,TOST 等价于 $\pm0.05$ 范围内的要求;锚定效应在细化过程中得以保持,且跟随参考策略而非初始化。选择遵循可达性加权 I-投影(斜率 0.969 [0.950, 0.987])。我们同样重点报告了该结论失效的情形:固定的流形外参考策略会带来 0.08–0.25 的可利用性代价;陡峭或平坦的策略族需要更小的镜像步长(由预先注册的规则设定);边界目标会出现欠冲;曲率可预测边界饱和发生的位置(秩相关系数 0.90,p=0.037),而内部精度与曲率无关。表格与 MLP 引导映射在每个目标点上的差异在 $\pm0.03$ 以内(30 个种子);匹配的对照组表明注意力机制(attention)的稳健特征是过大的种子方差,任何系统性偏移均被限制在 0.018 以内且不显著。面对最优响应时,选择-鲁棒性权衡是退化的:引导仅在对抗固定的非均衡对手时才有意义。这一方法——将参考策略锚定于期望成员并加以细化——将 RLHF 风格强化学习中的 KL 锚定重新诠释为一个选择旋钮,而不仅仅是稳定性约束。
cs.AI / 36 / 2609.19830

Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization

面向LLM智能体的双轴策略优化:贝叶斯反馈归因与轨迹质量归一化
Zhuang, Yingxuan, Yu, Binhe, Yang, Jingxiao, Sun, Ruopei, Li, Ziting, Tan, Cheng, Zhang, Xuhong, Yin, Jianwei, Chen, Jintao
Abstract
Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework. BATON instantiates the first axis with Bayesian Feedback Attribution, which constructs a feedback-conditioned posterior over sampled actions, and the second with Trajec- tory Mass Normalization (TMN), which assigns equal optimization mass to com- plete trajectories. Experiments with GRPO and GiGPO on ALFWorld, WebShop, and SearchQA show that both axes provide independent gains and that their combi- nation consistently achieves the strongest overall performance across model scales.
Chinese Translation
面向大语言模型(LLM)智能体的强化学习涉及两个不同的优化维度:如何在轨迹内利用环境反馈,以及如何在一个批次(batch)内聚合完整轨迹。我们将这两个维度形式化为“轨迹内反馈归因”(Intra-Trajectory Feedback Attribution)和“轨迹间目标聚合”(Inter-Trajectory Objective Aggregation),并提出BATON(Bayesian Attribution and Trajectory Objective Normalization,贝叶斯归因与轨迹目标归一化),一个双轴策略优化框架。BATON在第一个轴上采用贝叶斯反馈归因(Bayesian Feedback Attribution),构建以反馈为条件的采样动作后验分布;在第二个轴上采用轨迹质量归一化(Trajectory Mass Normalization, TMN),为完整轨迹分配相等的优化权重。在ALFWorld、WebShop和SearchQA数据集上使用GRPO和GiGPO进行的实验表明,两个轴均能带来独立的性能提升,且两者的结合在不同模型规模下均持续取得最强的整体性能。
cs.AI / 37 / 2609.19832

MetaRTL: Meta-path Attention Enhanced Relational Table Learning

MetaRTL:元路径注意力增强的关系型表格学习
Zhong, Ken, Li, Weichen, Wang, Zheng
Abstract
Relational table learning has gained increasing attention with the widespread use of relational databases. Existing methods typically rely on deep GNN or HGNN stacks, leading to high computational costs and limited performance on large real-world databases. We propose MetaRTL, a two-stage framework for scalable and expressive relational table learning. In the first stage, MetaRTL obtains initial table embeddings via lightweight pre-training. In the second stage, it performs non-parametric message passing to derive meta-path features, which are then aggregated by an attention module, MetaAttn. By shifting computation from deep message passing to efficient meta-path aggregation, MetaRTL captures rich relational semantics while maintaining high efficiency. Experiments on 10 real-world datasets across 24 tasks demonstrate the effectiveness of the proposed method.
Chinese Translation
随着关系型数据库的广泛使用,关系型表格学习(Relational Table Learning)受到越来越多的关注。现有方法通常依赖深层图神经网络(GNN)或异构图神经网络(HGNN)堆叠结构,导致计算成本高,且在大型真实数据库上的性能受限。我们提出MetaRTL,一个兼具可扩展性与表达能力的关系型表格学习两阶段框架。在第一阶段,MetaRTL通过轻量级预训练获得初始表格嵌入。在第二阶段,它执行非参数化消息传递以获取元路径特征,再由注意力模块MetaAttn进行聚合。通过将计算从深层消息传递转移到高效的元路径聚合,MetaRTL在保持高效率的同时捕获了丰富的关系语义。在10个真实世界数据集上跨越24个任务的实验证明了所提方法的有效性。
cs.AI / 38 / 2609.19843

A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents

基于LLM的GUI代理对数字助推敏感性的双系统视角研究
Halimeh, Haya, Kaltenpoth, Sascha, Bösch, Kevin, Müller, Oliver
Abstract
LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users. While behavioural biases in the textual outputs of LLMs are well-documented, far less is known about how such influence operates when models act as agents that perceive interfaces and execute decisions---and, in particular, whether the reasoning capabilities increasingly built into these agents make them more robust to it. Drawing on Dual-Process Theory, we empirically investigate whether LLM-based GUI agents are susceptible to automatic (Type 1) and reflective (Type 2) digital nudges, and how their reasoning configuration moderates this susceptibility. In a randomized online shopping experiment with 3,600 agents and a total of 21,600 simulations across six frontier models from three providers, we found that agents were vulnerable to both nudge types. Crucially, the reasoning configuration moderated these effects in opposing directions, reducing susceptibility to automatic default nudges while heightening it to reflective social influence nudges. Extensive reasoning therefore did not make agents more robust but redirected the route through which choice architecture takes effect. Exploratory analysis further showed this redirection to be systematically structured by model scale. Beyond establishing nudge susceptibility as a behavioural property of agentic AI, the study positions interface design as a governance concern for organizations that delegate decisions to autonomous agents.
Chinese Translation
基于大语言模型(LLM)的GUI代理越来越多地代表用户在为人类用户设计的数字环境中执行操作。这些图形用户界面在设计上既支持用户的行为与决策,也对其进行有意的引导。尽管LLM文本输出中的行为偏差已有大量文献记录,但当模型作为感知界面并执行决策的代理时,这种影响如何发挥作用却鲜有研究——尤其是,这些代理中日益增强的推理能力是否使其对此类影响更具抵抗力。本研究借鉴双系统理论(Dual-Process Theory),实证考察了基于LLM的GUI代理是否易受自动化(类型1)与反思性(类型2)数字助推的影响,以及其推理配置如何调节这种敏感性。在一项包含3,600个代理、共21,600次模拟、涵盖三家提供商六个前沿模型的随机化在线购物实验中,我们发现代理对两类助推均表现出易感性。关键的是,推理配置以相反方向调节了这些效应:降低了代理对自动化默认助推的敏感性,却提高了其对反思性社会影响助推的敏感性。因此,扩展推理并未使代理更加稳健,而是改变了选择架构发挥作用的路径。探索性分析进一步表明,这种路径转移受到模型规模的系统性调节。除了将助推敏感性确立为代理式AI(Agentic AI)的一种行为属性外,本研究还将界面设计定位为将决策委托给自主代理的组织所面临的治理问题。
cs.AI / 39 / 2609.19848

Constraint-Safe Graph-Context Scoring for Stable Point-Feature Labels Under Text-Width and Accessibility-Inspired Profiles

基于文本宽度与无障碍启发式约束的约束安全图上下文评分方法,用于稳定的点要素标注
Ahmad, Taimoor
Abstract
Point-feature label placement on interactive maps must reconcile geometric validity, display yield, local placement utility, and stability across camera motion. Accessibility and multilingual requirements further change label dimensions, yet algorithmic evaluations often collapse these concerns into overlap counts. We present LABELSENSE-Pilot, a reproducible prototype that generates eight compass candidates per feature, scores candidates with a multilayer perceptron over graph-context summaries, adds a previous-placement bonus, and selects a layout through mixed-integer optimization. The executed scorer is deliberately not described as a graph transformer. Every returned layout is checked for viewport containment, per-feature uniqueness, and pairwise clearance. Experiments use 2,500 airport coordinates and names spanning 155 countries, with country-grouped splits and generated density, camera, text-suffix, preference, and enlarged-font stressors. Across five seeds, LABELSENSE-Pilot displayed 85.62 percent of labels with 2.09 percent flicker and zero collisions. Versus a handcrafted-utility integer program, LABELSENSE-Pilot sacrificed 1.43 percentage points of display while reducing flicker by 12.04 points. Enlarged-box-aware layouts produced zero proxy violations, whereas standard geometry reevaluated at 1.5x violated 52.57 percent of selected placements. These results establish an auditable engineering trade-off, not human accessibility, multilingual usability, or preference. Official recent baselines and participant evidence remain required before submission.
Chinese Translation
交互式地图上的点要素标注放置必须协调几何有效性、显示产出、局部放置效用以及相机运动下的稳定性。无障碍与多语言需求还会进一步改变标注尺寸,然而算法评估往往将这些考量简化为重叠计数。我们提出了LABELSENSE-Pilot,一个可复现的原型系统:它为每个要素生成八个罗盘方向候选位置,利用多层感知机对图上下文摘要进行候选评分,加入先前位置奖励,并通过混合整数优化选择布局。已执行的评分器刻意不被描述为图 Transformer。每个返回的布局都会经过视口包含性、要素唯一性和两两间距的检查。实验使用覆盖155个国家的2,500个机场坐标与名称,采用按国家分组的划分方式,并生成了密度、相机、文本后缀、偏好和放大字号等压力因子。在五个随机种子下,LABELSENSE-Pilot显示了85.62%的标注,闪烁率为2.09%,且零碰撞。与手工设计效用函数的整数规划相比,LABELSENSE-Pilot牺牲了1.43个百分点的显示量,但将闪烁率降低了12.04个百分点。感知放大框的布局产生了零代理违规,而标准几何布局在1.5倍缩放重新评估时违规了52.57%的选定位置。这些结果确立了一种可审计的工程权衡,而非人类无障碍性、多语言可用性或偏好。在提交之前,仍需要官方的最新基线和参与者证据。
cs.AI / 40 / 2609.19866

Reproducibility is not construct validity: LLM measurement of institutionally situated communication

可复现性不等于构念效度:对制度情境化传播的大语言模型测量
Batzdorfer, Veronika, Santagiustina, Carlo Romano Marcello Alessandro
Abstract
High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({\=g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.
Chinese Translation
高标注可复现性并不必然意味着大语言模型(LLM)推断的测量能够捕捉其旨在测量的构念。我们利用欧盟委员会《人工智能法案》公开征求意见的数据集检验了这一区别,将结构化问卷回答与同一利益相关方提交的自由文本意见相互关联。对征求意见文本的LLM标注具有极高的可复现性(组内相关系数 > 0.99),但与其名义上要近似的构念在问卷报告测量之间的收敛性有限。问卷测量与LLM推断的文本测量之间的分歧在利益相关方群体间呈现系统性差异:商业协会在文本征求意见中对人工智能风险表达了比问卷回答中更高的担忧(g = +1.0),而公共机构和若干非商业群体则表现出较小或负向的分歧。评分间的分歧显示欧洲各国之间存在正向空间自相关(Moran's I = 0.347, p = 0.036),表明来自邻近国家的利益相关方在人工智能安全问题上倾向于表达更为相似的文本立场。尽管存在分歧,问卷报告的担忧在所有分歧水平上仍与对可解释性的支持密切相关。这些结果表明,LLM标注的可复现性可能与糟糕的构念对应性并存,因此在将大语言模型用作测量工具时,需要建立能够区分可复现性、构念效度和传播情境差异的验证程序。
cs.AI / 41 / 2609.19871

Physical knowledge on historical data matters more than enforcing physical constraints on the forecast

历史数据中的物理知识比对预测施加物理约束更为重要
Lehembre, Etienne, Audigane, Pascal, Nguyen, Vincent, Vrain, Christel, Dao, Thi-Bich-Hanh
Abstract
Time series forecasting has seen signicant advancements with the emergence of new deep learning models. However, forecasting time series in applications involving physical processes remains a major challenge. Despite the apparition of Physics Informed Neural Networks (PINN), recent models do not estimate unobservable intermediate physical variables, which are important for domain experts to understand the target behavior. To this end, we propose a Physics Informed Recurrent Neural Network (PIRNN) which predicts, along the target, unobservable variables on both historic data and forecast target. This approach enhances the model robustness and results interpretation using domain knowledge. Our method is easily adaptable to any physical model using several equations, each having its own set of unobservable variables, to describe it-self. As a case study, we incorporate physical equations used for groundwater levels predictions by the physical model called Gardenia. This model uses transfers equations between reservoirs, optimized with data assimilation, to simulate the evolution of groundwater levels. Evaluation includes several well known neural network models and the Gardenia model compared on twelve real world datasets. In addition, we study the impact of each component through an ablation study. Our model outperforms other models on ve out of the twelve datasets and our ablation study underlines the importance of having a physical background in our time series forecasting task. Finally, the coherence of the physical variables predicted by our neural network is assessed by a domain expert.
Chinese Translation
随着新型深度学习模型的出现,时间序列预测取得了显著进展。然而,对涉及物理过程的应用中的时间序列进行预测仍然是一个重大挑战。尽管出现了物理信息神经网络(Physics Informed Neural Networks, PINN),但近期的模型并未估计不可观测的中间物理变量,而这些变量对于领域专家理解目标行为非常重要。为此,我们提出了一种物理信息循环神经网络(Physics Informed Recurrent Neural Network, PIRNN),它在预测目标的同时,还能对历史数据和预测目标上的不可观测变量进行预测。该方法利用领域知识增强了模型的鲁棒性和结果的可解释性。我们的方法易于适配任何物理模型,可使用多个方程(每个方程拥有自己的一组不可观测变量)来描述模型自身。作为案例研究,我们引入了名为Gardenia的物理模型中用于地下水位预测的物理方程。该模型使用经数据同化优化的水库间水量传递方程来模拟地下水位的变化。评估工作在十二个真实世界数据集上比较了多种著名的神经网络模型与Gardenia模型。此外,我们通过消融实验研究了各组成部分的影响。我们的模型在十二个数据集中的五个上优于其他模型,消融实验强调了在时间序列预测任务中具备物理背景的重要性。最后,领域专家评估了我们神经网络所预测物理变量的合理性。
cs.AI / 42 / 2609.19897

TRACE: Accountable Agentic Retrieval for Source Discovery in Digital Archives

TRACE:面向数字档案来源发现的可问责智能体检索框架
Bian, Donghan, Puren, Marie, Cafiero, Florian
Abstract
Historical archives pose a difficult retrieval problem for retrievalaugmented generation systems: documents are OCR-degraded, heterogeneous across genres and sources, and require strong source traceability for scholarly and institutional use. We introduce TRACE, a training-free agentic retrieval framework designed for accountable source discovery over historical corpora. The system was developed in the context of DECIDON, an interdisciplinary project on the circulation of political discourse between parliamentary debates and the press during the French Third Republic, involving digitised historical collections and institutional use cases. The prototype is currently deployed internally within the project and accessible to 24 researchers across six partner institutions. We evaluate TRACE on HistoriQA-ThirdRepublic, a benchmark of 1,752 French historical questions over parliamentary debates and newspapers from 1887, with documents derived from Biblioth{\`e}que nationale de France digitised collections. TRACE achieves R@10 = 0.856 and MRR = 0.653, outperforming sparse, dense, graph-based, and agentic RAG baselines, with the largest gains on multi-hop and cross-corpus questions. At approximately $0.02 per question under the default hosted inference configuration, TRACE also remains economically feasible for heritage institutions, laboratories or companies that cannot rely on costly local GPU infrastructure. These results suggest that, for large digital libraries and archives, retrieval accountability and corpus-aware agent design can provide a practical alternative to heavier training-based or graph-construction approaches.
Chinese Translation
历史档案为检索增强生成系统带来了棘手的检索难题:文档存在OCR质量退化、体裁与来源异构等问题,并且对学术和机构用途而言需要强来源可追溯性。我们提出了TRACE,一个免训练的智能体检索框架,专为历史语料库上的可问责来源发现而设计。该系统是在DECIDON项目背景下开发的,该项目是一个跨学科项目,研究法兰西第三共和国时期议会辩论与新闻界之间政治话语的流通,涉及数字化历史馆藏和机构使用场景。该原型目前部署于项目内部,可供六个合作机构的24名研究人员使用。我们在HistoriQA-ThirdRepublic上对TRACE进行了评估,该基准包含1,752个关于1887年议会辩论和报纸的法语历史问题,文档来源于法国国家图书馆(Bibliothèque nationale de France)数字化馆藏。TRACE实现了R@10 = 0.856和MRR = 0.653,优于稀疏检索、稠密检索、基于图的方法以及智能体RAG基线,在多跳和跨语料库问题上取得最大提升。在默认托管推理配置下,每个问题成本约为0.02美元,对于无法依赖昂贵本地GPU基础设施的遗产机构、实验室或企业而言,TRACE在经济上仍然可行。这些结果表明,对于大型数字图书馆和档案馆,检索可问责性和语料库感知的智能体设计可以成为基于训练或图构建的更重方法的一种实用替代方案。
cs.AI / 43 / 2609.19928

From "Who Is This User?" to "What Does This Purchase Mean?": A Deployed Pipeline for Semantic User Profiling at Bank Scale

从“这位用户是谁?”到“这笔购买意味着什么?”:银行级规模的语义化用户画像部署流水线
Mitsuhashi, Ryota, Morimura, Tetsuro, Ito, Hirotake
Abstract
Per-user LLM inference on transaction histories binds the inference budget linearly to user count, which becomes prohibitive at applied scale. We re-cast attribute inference from per-user to per-transaction-pattern. The pipeline runs in three phases: Resolve abstracts item names with optional web grounding, Profile infers attributes for each frequent pattern, and Tag clusters free-text attributes into a queryable database. In Profile, a single LLM call per pattern emits predefined categorical labels, free-text attributes, and per-attribute prevalence estimates. Because inference runs over patterns rather than users, the budget grows with the pattern count rather than the user count. On the public Open e-commerce corpus, the database is statistically indistinguishable from an LLM that reads each user's raw history directly in AUC across the evaluated attributes, and the prevalence estimates carry discriminative signal between positive and negative users. The pipeline is deployed at a major Japanese bank profiling on the order of tens of millions of users, with close to a three-order-of-magnitude reduction in LLM inference targets versus a per-user pipeline. The code is publicly available on https://github.com/CyberAgentAILab/profiling-agent-open-ecommerce.
Chinese Translation
针对每位用户在交易历史上运行大语言模型(LLM)推理会使推理预算与用户数量呈线性关系增长,在实际应用规模下变得不可承受。我们将属性推断从“每用户”重新构造为“每交易模式”。该流水线分三个阶段运行:Resolve(解析)利用可选的网络信息对商品名称进行消歧;Profile(画像)为每种高频模式推断属性;Tag(标记)将自由文本属性聚类为可查询的数据库。在 Profile 阶段,每种模式仅需一次 LLM 调用,即可输出预定义的类别标签、自由文本属性以及每个属性的流行度估计。由于推理针对模式而非用户运行,预算随模式数量而非用户数量增长。在公开的 Open e-commerce 语料库上,该数据库在所评估属性的 AUC 上与直接读取每位用户原始历史的 LLM 在统计上不可区分,且流行度估计在正负用户之间具有判别信号。该流水线已部署于一家日本大型银行,画像规模达数千万用户,与每用户流水线相比,LLM 推理目标数量减少近三个数量级。代码已在 https://github.com/CyberAgentAILab/profiling-agent-open-ecommerce 公开发布。
cs.AI / 44 / 2609.19934

Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models

超越深度截断:递归语言模型中深度利用的可控评估
Van Dau, Ha, Khuat, Thanh Tung, Dung, Nguyen Thanh
Abstract
Depth-recurrent language models iteratively apply a small layer stack, decoupling per-token compute from distinct parameter count. To determine whether such a model genuinely utilizes its depth, both recurrence and layer-pruning literatures rely on a shared evaluation: truncating depth at inference time, plotting quality against retained depth fraction, and reading off the slope. While cheap and training-free, this metric suffers from an unexamined flaw: it extracts a single scalar from an intervention that alters multiple model properties simultaneously. Depth truncation concurrently reduces the number of block applications, decreases the volume of distinct computation performed, and pushes the readout head onto an out-of-distribution residual stream. The observed slope conflates all three factors, yet is conventionally interpreted as reflecting solely the second. We propose the Depth Control Protocol (DCP), a diagnostic suite that disentangles these three quantities. DCP comprises three positive controls that isolate each factor while varying the others, a negative control applying the identical interventions to dense transformers to ensure the effect is not an artifact of the measurement protocol, and a controlled training intervention to verify causality. The linchpin control, running the full budget of block applications while executing only a single distinct iteration, is strictly realizable only in depth-wise weight-sharing architectures, since in a dense network repeating a layer yields an entirely different model rather than the same model in an alternative configuration.
Chinese Translation
深度递归语言模型通过迭代应用一个小型层堆栈,将每token的计算量与独立参数数量解耦。为了判断此类模型是否真正利用了其深度,递归研究与层剪枝文献均依赖一种共享的评估方法:在推理时截断深度,绘制性能随保留深度比例变化的曲线,并读取其斜率。这种方法虽然成本低廉且无需训练,但存在一个未被审视的缺陷:它从一种会同时改变模型多个属性的干预中提取单一标量。深度截断同时减少了块应用次数、降低了执行的独立计算量,并将读取头推入分布外的残差流。所观测到的斜率混淆了这三个因素,却通常被解释为仅反映第二个因素。我们提出深度控制协议(Depth Control Protocol, DCP),一套用于解耦这三个量的诊断套件。DCP包含三个阳性对照,在改变其他因素的同时隔离每一单个因素;一个阴性对照,将相同的干预应用于稠密Transformer,以确保该效应不是测量协议本身的伪影;以及一个受控训练干预,以验证因果关系。其中关键的对照——在完整块应用预算下仅执行一次独立迭代——只有在按深度共享权重的架构中才能严格实现,因为在稠密网络中重复某层得到的是一个完全不同的模型,而非同一模型的另一种配置。
cs.AI / 45 / 2609.19944

MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation

MaSCoD:一种结构上下文引导的候选因果图生成多智能体框架
Nakada, Yudai, Nishiura, Yuichiro, Splichal, Jin Michael
Abstract
Large language models (LLMs) have been applied to causal discovery, but candidate-graph generation rarely treats premature omission of potentially relevant causal relations as an explicit design objective. We propose MaSCoD, a multi-agent framework that organizes candidate third variables and local structural patterns before direct-edge judgment. We evaluate MaSCoD on Auto-MPG, DWD, and Sachs using GPT-5.4 as the primary backbone and GPT-4o for replication. MaSCoD exhibits a dataset- and backbone-dependent retention-selectivity profile rather than uniform superiority. Across all six dataset-backbone settings, Full, which supplies structural hypotheses before direct-edge judgment, achieved higher mean Recall and F1 than No Phase 1, which instead constructs them within the judgment procedure, while also increasing false-positive rates. Additional reference-edge retention over all evaluated baselines was observed on DWD with GPT-5.4 and on Sachs with GPT-4o, rather than uniformly across settings. Partial ablations showed that supplying both information components did not always outperform supplying only one. For GPT-5.4, stage-wise analysis showed that the Full-No Phase 1 retention gap was already present after direct-edge judgment, while reconciliation introduced additional reference-edge loss for Full on Sachs. These findings support structural pre-organization as an explicit design and evaluation target for omission control and motivate evaluating context construction jointly with its utilization in judgment.
Chinese Translation
大语言模型(LLM)已被应用于因果发现,但在候选图生成中,过早遗漏潜在相关因果关系这一问题很少被作为明确的设计目标。我们提出MaSCoD,一个多智能体框架,它在进行直接边判断之前先组织候选第三变量和局部结构模式。我们以GPT-5.4作为主要骨干模型、GPT-4o用于复现验证,在Auto-MPG、DWD和Sachs数据集上对MaSCoD进行了评估。结果显示,MaSCoD表现出依赖于数据集和骨干模型的保留—选择性特征,而非全面的优越性。在全部六种数据集—骨干模型组合中,Full(在直接边判断之前提供结构假设)相比No Phase 1(在判断过程内部构建这些假设)取得了更高的平均召回率(Recall)和F1值,但同时也提高了假阳性率。相比所有评估基线,额外的参考边保留仅在GPT-5.4的DWD和GPT-4o的Sachs上被观察到,而非在所有设置中普遍存在。部分消融实验表明,同时提供两种信息组件并不总是优于仅提供其中一种。对于GPT-5.4,分阶段分析显示,Full与No Phase 1之间的保留差距在直接边判断阶段即已存在,而在Sachs上,协调阶段为Full引入了额外的参考边损失。这些发现支持将结构预组织作为遗漏控制的明确设计与评估目标,并提示应在判断过程中将上下文构建与其利用情况联合评估。
cs.AI / 46 / 2609.19947

Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

并非所有AI智能体都相同:资源与性能动态特性分析
Choi, Wonmi, Park, Minuk, Niu, Zhixiong, Xiong, Yongqiang, Yoo, Chuck, Yang, Gyeongsik
Abstract
LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving the invocation of remote LLM APIs with local tool containers. This execution model can make the optimization of agent serving difficult because latency, local resource demand, and container bottlenecks inter-mix across requests. However, the current agent ecosystem runs without much consideration of resource dynamics, which results in significant waste of the precious resources. This paper analyzes the resource inter-mix of AI agents for three representative tasks: retrieval-augmented question answering, web search, and software coding. To this end, we characterize the latency with respect to the resource dynamics of processing multiple requests and tasks concurrently. Our measurements show that agents have a wide range of behaviors depending on tasks, so that even the same tool can differ substantially in resource dynamics. We also find that running multiple requests concurrently exposes task-dependent bottlenecks in resource dynamics such as CPU, disk I/O, and memory. Furthermore, we uncover that faster LLM responses or more CPU cores do not always accelerate agents. Based on these observations, we demonstrate new optimization opportunities that exploit the resource dynamics of tasks: CPU-aware tool admission and task-aware CPU allocation. Our results show that the latency of CPU-sensitive agent tasks improves $\sim$5.4$\times$, and the average latency across multiple tasks is reduced $\sim$32% compared to native agents.
Chinese Translation
基于LLM的AI智能体通过迭代推理和工具执行来处理用户请求,通常涉及调用远程LLM API以及本地工具容器。这种执行模型使得智能体服务优化变得困难,因为延迟、本地资源需求和容器瓶颈在请求之间相互交织。然而,当前的智能体生态系统在运行时并未充分考虑资源动态特性,导致宝贵资源的严重浪费。本文针对三种代表性任务分析了AI智能体的资源交织情况:检索增强问答、网页搜索和软件编码。为此,我们刻画了并发处理多个请求和任务时,延迟与资源动态之间的关系。我们的测量表明,智能体的行为因任务而异,范围广泛,即使相同的工具在资源动态上也可能存在显著差异。我们还发现,并发运行多个请求会暴露出任务相关的资源动态瓶颈,如CPU、磁盘I/O和内存。此外,我们揭示出更快的LLM响应或更多的CPU核心并不总能加速智能体。基于这些观察,我们展示了利用任务资源动态特性的新优化机会:CPU感知的工具准入和任务感知的CPU分配。结果表明,CPU敏感的智能体任务延迟提升了约5.4倍,与原生智能体相比,多个任务的平均延迟降低了约32%。
cs.AI / 47 / 2609.19961

Neuro-Symbolic Agentic AI for Networked Low-Altitude UAVs

面向网络化低空无人机的神经-符号智能体人工智能
Ping, Yuqi, Liang, Tianhao, Su, Nanchi, Lei, Guangyu, Wu, Junwei, Zhang, Qinyu, Zhang, Tingting
Abstract
Networked low-altitude unmanned aerial vehicles (UAVs) need reliable and adaptive decision-making capabilities to operate under uncertain observations, dynamic environments, and intermittent connectivity, while many existing agentic systems remain limited by hallucination risks, data dependence, and weak generalization. This article investigates neuro-symbolic agentic AI (NSAAI) as a framework for combining neural grounding, symbolic reasoning, and closed-loop agentic interaction to support more reliable and adaptive UAV autonomy. We first examine its capability foundations in data efficiency, compositional generalization, continual learning, and zero-shot transfer, and then develop a reference architecture integrating task and goal management, neuro-symbolic planning, verification and metacognition, skill execution and network interaction, and shared knowledge and memory. An urban fire-inspection case implemented in LAESim illustrates how a UAV can coordinate sensing and cloud access under intermittent connectivity, reuse a verified image-delivery skill, and satisfy explicit evidence conditions before completing the mission. The results illustrate the potential of NSAAI to support reusable skills, evidence-grounded decision-making, and adaptive mission execution in networked UAV systems. We further discuss key research directions in uncertainty-aware reasoning, knowledge and skill expansion, adaptive self-monitoring, and standardized evaluation.
Chinese Translation
网络化低空无人机(UAV)需要在不确定观测、动态环境和间歇性连接条件下运行时具备可靠且自适应的决策能力,而许多现有智能体系统仍受幻觉风险、数据依赖和泛化能力弱的限制。本文研究神经-符号智能体人工智能(NSAAI)作为一个框架,将神经感知接地、符号推理与闭环智能体交互相结合,以支持更可靠和自适应的无人机自主性。我们首先考察其在数据效率、组合泛化、持续学习和零样本迁移方面的能力基础,然后提出一种参考架构,整合任务与目标管理、神经-符号规划、验证与元认知、技能执行与网络交互,以及共享知识与记忆。在LAESim中实现的城市火灾巡检案例展示了无人机如何在间歇性连接下协调感知与云端访问、复用经过验证的图像传输技能,并在完成任务前满足明确的证据条件。结果说明了NSAAI在网络化无人机系统中支持可复用技能、基于证据的决策和自适应任务执行的潜力。我们进一步讨论了不确定性感知推理、知识与技能扩展、自适应自我监控以及标准化评估等关键研究方向。
cs.AI / 48 / 2609.19996

Customizable and Jointly Optimized Route Planning: A Deep Architecture Enabling Differentiable Shortest-Path Search

可定制与联合优化的路径规划:一种支持可微分最短路径搜索的深度架构
Zhao, Rui, Chen, Chao, Xu, Longfei, Ji, Chenguang, Cui, Hengbin, Liu, Kaikui, Li, Xiaolong
Abstract
With the widespread use of online navigation and ride-hailing services, achieving optimal route planning for diverse user preferences has recently attracted increasing attention. Classic graph algorithms for pathfinding use heuristic cost functions to define edge weight, thus providing no optimality guarantee of route quality. Prior data-driven approaches equating ground truth of the optimal route with user trajectory, which is however moderately influenced by the navigation service, suffers from the feedback loop problem. To address these issues, we propose a deep architecture that is able to jointly optimize cost functions and route-ranking model towards any route preference. First, we run a multi-objective Dijkstra algorithm offline to collect the set of Pareto optimal routes, deeming it as the complete candidate set. Exploiting the property of such a set, we design a neural network structure that emulates shortest-path search and route ranking in an end-to-end differentiable manner. Second, we define route preference as a task of constrained optimization of route attributes, and propose a novel loss function that optimizes a single-objective variable, with other variables strictly under constraints. We conduct extensive experiments on real-world datasets. The results show that our architecture significantly outperforms state-of-the-art methods in route quality and customizability.
Chinese Translation
随着在线导航和网约车服务的广泛使用,如何针对多样化用户偏好实现最优路径规划近期受到越来越多的关注。经典的图搜索算法使用启发式代价函数来定义边权重,因而无法保证路径质量的最优性。以往的数据驱动方法将最优路径的真实标准等同于用户轨迹,然而用户轨迹受导航服务的影响较大,存在反馈循环问题。为解决这些问题,我们提出了一种深度架构,能够针对任意路径偏好联合优化代价函数与路径排序模型。首先,我们离线运行多目标 Dijkstra 算法收集 Pareto 最优路径集合,并将其视为完整的候选集合。利用该集合的特性,我们设计了一种神经网络结构,以端到端可微分的方式模拟最短路径搜索和路径排序。其次,我们将路径偏好定义为路径属性约束优化问题,并提出一种新颖的损失函数,对单一目标变量进行优化,同时使其他变量严格满足约束。我们在真实世界数据集上进行了大量实验,结果表明,我们的架构在路径质量和可定制性方面显著优于现有最先进方法。
cs.AI / 49 / 2609.20001

E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews

E-AVI:面向自动化视频面试的基于证据的多模态评估
Wang, Haoshen, Che, Dongbo, Xie, Zeyi, Du, Yuanjie, Hua, Shicheng, Wang, Xingyu
Abstract
Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an evidence-grounded framework that extracts timestamped multimodal evidence and integrates dimension-conditioned evidence attention with source-level embeddings for scoring. A shared evidence pool further supports natural-language feedback and follow-up question answering. On RecruitView and a private hospitality dataset, E-AVI consistently outperforms fine-tuned multimodal baselines in rank correlation. Ablation, evidence-deletion, bootstrap, human-audit, and QA analyses characterize the predictive contribution, grounding, and practical utility of the evidence pathway. Together, these results demonstrate that our proposed E-AVI framework improves predictive performance while providing inspectable support for assessment, feedback, and interactive analysis.
Chinese Translation
自动化视频面试评估需要整合语言内容、声音表现和视觉行为,然而仅凭数值预测所能提供的可检验依据十分有限。我们提出了E-AVI,一个基于证据的框架,该框架提取带时间戳的多模态证据,并将维度条件化的证据注意力机制与源级嵌入相融合以进行评分。共享证据池还进一步支持自然语言反馈生成和追问问答任务。在RecruitView和一个私密的酒店行业数据集上,E-AVI在秩相关性指标上持续优于经过微调的多模态基线模型。消融实验、证据删除、自助法(bootstrap)、人工审核以及问答分析共同刻画了证据通路在预测贡献、依据支撑和实际效用方面的特性。综合来看,这些结果表明我们所提出的E-AVI框架不仅提升了预测性能,还能为评估、反馈和交互式分析提供可检验的支撑依据。
cs.AI / 50 / 2609.20005

Geopolitical Divisions Across Languages in Large Language Models

大型语言模型中跨语言的地缘政治分歧
Chupilkin, Maxim
Abstract
People increasingly turn to AI chatbots for news and explanations of world events. But do they receive the same political answers when they ask in different languages? Here we show that the language of a question can change how the same AI systems assess the war in Ukraine. We ask GPT, Claude and Gemini to evaluate twenty statements about the war in 112 languages, collecting 67,200 responses. The balance between Russia-leaning and Ukraine-leaning responses differs across languages. When we group responses by countries' official languages, they follow a pattern resembling worldwide political divisions: relatively more Russia-leaning answers correspond to more favourable public views of Russia, less support for Ukraine in United Nations votes, and less aid to Ukraine. The broad pattern recurs across all three models and remains when individual statement pairs are removed. Our findings suggest a possible route through which information warfare may shape the text used to train AI models, which may in turn spread geopolitical biases.
Chinese Translation
人们越来越多地借助AI聊天机器人获取新闻和了解世界事件的解释。但当用户以不同语言提问时,会得到相同的政治性回答吗?我们的研究表明,提问所用的语言会改变同一AI系统对乌克兰战争的评估。我们让GPT、Claude和Gemini以112种语言评估关于这场战争的二十项陈述,共收集了67,200条回应。倾向俄罗斯与倾向乌克兰的回应之间的平衡在不同语言间存在差异。当我们按各国的官方语言对回应进行分组时,其呈现出与全球政治格局相似的模式:相对更多倾向俄罗斯的回答,对应着民众对俄罗斯更正面的看法、在联合国投票中对乌克兰更少的支持,以及对乌克兰更少的援助。这一总体模式在三个模型中均反复出现,且在剔除个别陈述对后依然保持。我们的发现揭示了一条可能的路径:信息战可能通过影响用于训练AI模型的文本,进而传播地缘政治偏见。
cs.AI / 51 / 2609.20026

FedeRICo: Federated Region-Influenced Coupling for Traffic Flow Prediction

FedeRICo:面向交通流预测的联邦区域影响耦合方法
Orozco, Fermin, Luo, Man, Wahlström, Johan
Abstract
Urban traffic forecasting often relies on information distributed across stakeholders who may be unable to share raw data due to privacy or commercial constraints, motivating federated spatial-temporal approaches. In such federated settings, each client observes traffic over a distinct sensor subgraph with its own spatial topology and temporal dynamics, leading to significant heterogeneity across clients. Existing federated spatial-temporal methods typically rely on model parameter aggregation and provide limited mechanisms for recovering spatial dependencies across client boundaries. This introduces two key limitations. Specifically, parameter aggregation across heterogeneous graph domains tends to dilute client-specific representations, while road network partitioning breaks the propagation of traffic dynamics across client boundaries. To address these challenges, we propose FedeRICo, a federated traffic forecasting framework that combines gradient-level collaboration with boundary-aware residual communication. FedeRICo employs a dual-branch forecasting architecture in which a globally guided branch captures transferable forecasting structure, while a private residual branch preserves client-specific corrections and incorporates boundary residual signals. The global branch is coordinated through gradient alignment across all clients, enabling collaborative optimisation without destructive parameter interference. To recover cross-client spatial dependencies, boundary messages are extracted through a trend-residual decomposition that suppresses periodic structure and communicates only transient spatial-temporal residual signals between physically adjacent clients. Experiments across four real-world traffic forecasting benchmarks demonstrate that FedeRICo consistently outperforms state-of-the-art federated spatial-temporal baselines while maintaining competitive training runtime.
Chinese Translation
城市交通预测通常依赖于分布在多个利益相关方之间的信息,而由于隐私或商业限制,这些相关方可能无法共享原始数据,这推动了联邦时空学习方法的发展。在此类联邦环境中,每个客户端观测到的交通数据来自一个具有自身空间拓扑和时间动态的独立传感器子图,导致客户端之间存在显著的异质性。现有的联邦时空方法通常依赖于模型参数聚合,缺乏恢复客户端边界之外空间依赖关系的有效机制。这带来了两个关键局限:具体而言,跨异构图域的参数聚合容易稀释客户端特定的表征,而路网划分则阻断了交通动态在客户端边界之间的传播。为应对这些挑战,我们提出了FedeRICo,一个将梯度级协作与边界感知残差通信相结合的联邦交通预测框架。FedeRICo采用双分支预测架构:全局引导分支捕捉可迁移的预测结构,而私有残差分支保留客户端特定的修正并融合边界残差信号。全局分支通过所有客户端之间的梯度对齐进行协调,实现协作优化且不产生破坏性的参数干扰。为恢复跨客户端的空间依赖关系,边界消息通过趋势-残差分解提取,该分解抑制周期性结构,仅在物理相邻的客户端之间传递瞬态时空残差信号。在四个真实世界交通预测基准数据集上的实验表明,FedeRICo在保持具有竞争力的训练运行时间的同时,始终优于最先进的联邦时空基线方法。
cs.AI / 52 / 2609.20027

Can Data Attribution Filter Out Subliminal Learning? Not Reliably

数据归因能否过滤掉潜意识学习?并不可靠
Weckbecker, Moritz, Jena, Sweta, Müller, Jonas, Kumaraguru, Ponnurangam, Lapuschkin, Sebastian, Samek, Wojciech, Jaburi, Louis, Paulo, Gonçalo
Abstract
Subliminal learning allows language models to transmit behavioral traits through training data with no obvious semantic relationship to those traits, undermining content-based data filtering as a safety intervention. Training data attribution offers an alternative: it identifies the training examples responsible for a given model behavior, independent of their semantic content, and so may apply in exactly the cases where semantic inspection fails. We evaluate three gradient-based attribution methods (GradCos, a contrastive GradCos variant, and EK-FAC) across three models, comparing them against divergence tokens, a strong baseline previously shown to localize subliminal learning (albeit one that requires access to counterfactual teacher models). Filtering at the token level, EK-FAC mitigates a significant part of the effect, the other methods provide little benefit, and all mostly fall short of divergence tokens. Filtering entire samples is less effective for every method, though EK-FAC often gives a stronger signal than divergence tokens in this setting. Success is inconsistent across methods and settings: variants that work well for some model-preference combinations fail for others, and we do not identify a consistent explanation for these differences. Our results suggest that gradient-based attribution can identify data responsible for subliminal learning in some settings, but that some approximations are more reliable than others.
Chinese Translation
潜意识学习(Subliminal Learning)使语言模型能够通过与这些行为特征没有明显语义关系的训练数据来传递行为特征,这削弱了基于内容的数据过滤作为一种安全干预手段的作用。训练数据归因(Training data attribution)提供了一种替代方案:它能够识别导致特定模型行为的训练样本,而不依赖于这些样本的语义内容,因此可能恰好适用于语义检查失效的情形。我们在三个模型上评估了三种基于梯度的归因方法(GradCos、一种对比式GradCos变体以及EK-FAC),并将其与发散词元(divergence tokens)进行比较——后者是一种强基线方法,此前已被证明能够定位潜意识学习(尽管它需要访问反事实教师模型)。在词元级别进行过滤时,EK-FAC能够缓解该效应的相当一部分,其他方法几乎没有益处,且所有方法大多不及发散词元的效果。过滤整个样本对每种方法而言效果都较差,不过在这一设置下EK-FAC往往比发散词元给出更强的信号。各种方法在不同设置下的成功并不一致:对某些模型-偏好组合有效的变体在其他组合上会失效,而我们未能为这些差异找到一致的解释。我们的结果表明,基于梯度的归因在某些设置下可以识别导致潜意识学习的数据,但其中某些近似方法比其他方法更可靠。
cs.AI / 53 / 2609.20051

DART: Distillation-Aware Reparameterization for Training-Free LoRA Reuse in Few-Step Video Diffusion Models

DART:面向少步视频扩散模型中免训练LoRA复用的蒸馏感知重参数化方法
Li, Shihong, Xu, Juntao, JinCao, Tang, Maowen, Huang, Jun, Li, Jintao
Abstract
Step distillation reduces the cost of video generation, but reusing a LoRA trained for a longer trajectory can alter its functional effect or degrade target quality. Static parameter compatibility offers one perspective on this problem; our observations show that similar measured geometry can coexist with different adapter behavior under a shortened denoising schedule. We propose DART, a training-free method that combines low-rank coordinate transport with target-schedule response calibration using forward evaluations and no source training videos. On a four-step Wan2.2 target, DART-F improves the joint quality score from 0.9029 to 0.9227 and changes macro functional retention from -0.4644 to +0.1349. Component analysis shows that calibration accounts for most of the quality improvement, while coordinate transport provides complementary gains when combined with calibration. Adapter-level results reveal positive functional effects for some adapters and strong attenuation with reduced negative functional effects for others. Evaluations on two additional targets show the same aggregate trend. These results motivate evaluating distilled-model LoRA reuse jointly through functional preservation and negative-transfer avoidance, without assuming recovery for every adapter.
Chinese Translation
步数蒸馏降低了视频生成的成本,但复用为较长去噪轨迹训练的LoRA可能会改变其功能效果或降低目标质量。静态参数兼容性为该问题提供了一个视角;我们的观察表明,在缩短的去噪调度下,相似的测量几何特性可能与不同的适配器行为并存。我们提出了DART,这是一种免训练方法,它将低秩坐标传输与基于前向评估的目标调度响应校准相结合,且无需源训练视频。在四步的Wan2.2目标上,DART-F将联合质量分数从0.9029提升至0.9227,并将宏观功能保持度从-0.4644改变为+0.1349。组件分析表明,校准贡献了大部分质量提升,而坐标传输在与校准结合时提供了互补增益。适配器层面的结果显示,部分适配器获得了正向功能效果,而其他适配器则表现出强衰减且负向功能效应有所降低。在另外两个目标上的评估显示出相同的总体趋势。这些结果促使我们通过功能保持与负迁移规避的联合视角来评估蒸馏模型的LoRA复用,而不假设每个适配器都能恢复原有功能。
cs.AI / 54 / 2609.20056

MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution

MAGMA-GEN:通过反事实重执行从模糊故障中获取经验证的恢复监督
Bernat, Loan, Grard, Matthieu, Herbulot, Ariane, Lamiraux, Florent
Abstract
Hierarchical robotic systems executing long-horizon manipulation tasks must make high-level semantic decisions that orchestrate stochastic low-level skills. In this setting, failed rollouts are ambiguous: a poor downstream state may reflect an invalid high-level decision, partial observation, or a valid decision whose physical execution failed. Traditional supervised learning lacks data for such recovery states, while reinforcement learning struggles with sparse rewards and non-local credit assignment. We propose MAGMA-GEN, an on-policy data-generation pipeline that converts ambiguous failed rollouts into validated recovery supervision. MAGMA-GEN first uses a privileged coach to hypothesize an early decision-level error and propose localized correction or recovery actions. Because this diagnosis is fallible, candidates are retained only if re-execution from the same state under matched conditions improves downstream progress. This produces supervised examples from the agent's own failure distribution without per-step human demonstrations. Evaluated on interactive long-horizon manipulation tasks, MAGMA-GEN improves task success and recovery capabilities, against distillation and trajectory-repair baselines under evolving task constraints in both simulation and real-robot execution.
Chinese Translation
执行长时程操作任务的分层机器人系统需要做出高层语义决策,以协调随机的低层技能。在这种设定下,失败的执行轨迹(rollout)具有模糊性:较差的下游状态可能源于无效的高层决策、部分可观测性,或是一个有效决策的物理执行失败。传统监督学习缺乏此类恢复状态的数据,而强化学习则受困于稀疏奖励和非局部信用分配问题。我们提出 MAGMA-GEN,一种在线策略(on-policy)数据生成流水线,能够将模糊的失败轨迹转化为经验证的恢复监督数据。MAGMA-GEN 首先利用一个具有特权信息的教练(coach)来推测早期决策层面的错误,并提出局部化的纠正或恢复动作。由于这种诊断可能出错,只有当在相同条件下从同一状态重执行能够改善下游进展时,候选方案才会被保留。由此可在无需逐步人类示教的情况下,从智能体自身的失败分布中生成监督样本。在交互式长时程操作任务上的评估表明,在任务约束不断演变的仿真和真实机器人执行环境中,MAGMA-GEN 相较于蒸馏和轨迹修复基线方法,提升了任务成功率和恢复能力。
cs.AI / 55 / 2609.20057

WiCleanData: Guaranteeing the Type Consistency of Wikidata by Taxonomy Refinement and Constraint Enforcement

WiCleanData:通过分类体系细化与类型约束执行保障Wikidata的类型一致性
Peng, Yiwen, Jeanmougin, Marc, Bonald, Thomas
Abstract
Because of its collaborative nature, Wikidata suffers from errors, in- consistencies, and excessive complexity, such as redundant classes, ambiguity between instances and classes, wrong taxonomic paths, and type constraint violations. The manual curation of these issues is infeasible at scale. To address these challenges, we introduce WiCleanData, a refined version of Wikidata with a consistent tax- onomy and free from type constraint violations. Specifically, we have designed an automated pipeline that first cleans the taxonomy with language model assistance, then simplifies type constraints by hierarchical aggregation, and finally filters facts accordingly. The resulting knowledge graph, free from any type violation, is made publicly available via a Web interface, enabling easy exploration and downstream applications.
Chinese Translation
由于其协作性质,Wikidata存在错误、不一致和过度复杂等问题,例如冗余的类、实例与类之间的歧义、错误的分类路径以及类型约束违反。人工对这些问题的处理在大规模情况下是不可行的。为了应对这些挑战,我们提出了WiCleanData,这是Wikidata的一个改进版本,具有一致的分类体系且不存在类型约束违反。具体而言,我们设计了一个自动化流水线:首先借助语言模型清理分类体系,然后通过层次聚合简化类型约束,最后据此对事实进行过滤。所得到的知识图不存在任何类型违反,已通过Web界面公开发布,便于探索和下游应用。
cs.AI / 56 / 2609.20067

FCA-Guided Counterfactual Explanations for Multi-Modal Breast Cancer Diagnosis: A Framework Achieving Perfect Validity with Emergent Sparsity

FCA引导的多模态乳腺癌诊断反事实解释:一个实现完全有效性并具有涌现稀疏性的框架
Isa, Abdullahi, Boukari, Souley, Aliyu, Muhammad
Abstract
Deep learning models for multi-modal breast cancer diagnosis achieve high predictive accuracy but remain clinically unacceptable without actionable, counterfactual explanations. Attribution-based methods (LIME, SHAP) are categorically inapplicable to this purpose, as they generate no alternative instances and thus cannot be evaluated on counterfactual quality metrics. This investigation provides empirical evidence that FCA-Guided Counterfactual (FCA-CF) framework that uses a Formal Concept Analysis (FCA) concept lattice as a hard structural constraint on counterfactual search, operating over a multi-modal TCGA-BRCA dataset. We benchmark against four genuine counterfactual methods: Wachter-style CF, DiCE, FACE, and NICE, evaluated on 60 benign-predicted TCGA-BRCA instances. The FCA-CF framework achieves Validity = 1.0000 (100% of counterfactuals successfully flip the prediction), Sparsity = 2.37 features changed (best among all valid methods), and Proximity = 0.900 (normalised L2-based, matching NICE as joint best). The classifier achieves Accuracy = 0.980, F1 = 0.976, ROC-AUC = 0.9947. Ablation analysis confirms that the FCA lattice constraint is the primary sparsity driver (removing it increases sparsity by +40%, p < 0.001, Cohen's d = 0.78), while Phase C greedy refinement accounts for the largest individual contribution (+113% sparsity increase when disabled, p < 0.001, d = 5.01). FCA-guided counterfactual generation achieves a clinically important Pareto-dominant outcome; it is simultaneously the sparsest and among the most proximate of all valid methods, with perfect validity. The emergent sparsity property arising from lattice topology rather than numerical penalty terms constitutes a structurally novel contribution to the counterfactual explanation literature.
Chinese Translation
用于多模态乳腺癌诊断的深度学习模型虽然具有很高的预测准确率,但在缺乏可操作的反事实解释的情况下,仍无法在临床上被接受。基于归因的方法(如LIME、SHAP)从根本上不适用于此目的,因为它们不生成替代实例,因此无法基于反事实质量指标进行评估。本研究提供了实证证据,证明FCA引导的反事实框架(FCA-CF)使用形式概念分析(FCA)概念格作为反事实搜索的硬性结构约束,并在多模态TCGA-BRCA数据集上运行。我们与四种真正的反事实方法进行了基准比较:Wachter式CF、DiCE、FACE和NICE,并在60个被预测为良性的TCGA-BRCA实例上进行评估。FCA-CF框架实现了有效性(Validity)= 1.0000(100%的反事实成功翻转预测结果)、稀疏性(Sparsity)= 2.37个特征变化(在所有有效方法中最佳),以及邻近度(Proximity)= 0.900(基于归一化L2,与NICE并列最佳)。该分类器的准确率(Accuracy)= 0.980,F1 = 0.976,ROC-AUC = 0.9947。消融分析证实,FCA格约束是稀疏性的主要驱动因素(移除该约束会使稀疏性增加40%,p < 0.001,Cohen's d = 0.78),而阶段C的贪心优化则贡献了最大的单项影响(禁用时稀疏性增加113%,p < 0.001,d = 5.01)。FCA引导的反事实生成在临床上实现了重要的帕累托占优结果:它在所有有效方法中同时具有最稀疏性和最高的邻近度之一,并具有完全的有效性。这种由格拓扑结构而非数值惩罚项所产生的涌现稀疏性特性,构成了对反事实解释文献在结构上的新颖贡献。
cs.AI / 57 / 2609.20068

Marginal utility, matrix factorization, and the Key-Value (KV) cache: a unified information-economic framework for sovereign geo-mining inference

边际效用、矩阵分解与键值(KV)缓存:面向主权地矿信息抽取推理的统一信息经济学框架
Combe, Caroline Gans
Abstract
This paper builds a theoretical bridge between the economic notion of marginal utility and two machine-learning constructs, matrix factorization and the Key--Value cache of transformer language models. The singular value spectrum of a rating matrix is shown to be a diminishing marginal utility schedule for latent factors, the eigenvalue spectrum of the projected covariance operator to be the marginal utility schedule of a model's learned representation, and cache eviction and low-rank cache compression to be instances of constrained utility maximization under a memory budget. The three collapse into a single allocation rule: retain the top dimensions whose eigenvalue exceeds the shadow price of the binding constraint. The framework is applied to the automated extraction of structured information from geo-mining documents, where it motivates a multi-pass inference protocol, a layer-wise TIES model merging procedure, and a selection policy combining extraction quality, localization drift and energy, scalarized with a Conditional Value-at-Risk term on drift. Two empirical contributions are reported. An 11.2-million-parameter hierarchical classifier, trained in about five minutes on a single GPU, reaches 90.0 per cent level-1 accuracy on a held-out test set from a 973-document uranium-exploration corpus, against 92.0 per cent for a proprietary model on a fifty-document human audit of the same corpus, at a latency of 2.62 ms per card against approximately 2,000 ms for the API and at negligible cost. A diagnostic of uniform-density TIES merging exposes a reproducible degenerate mode in which the merged model returns token-identical outputs across five geographically distinct districts while declaring high confidence; re-executing the merge under layer-wise calibrated densities removes that signature on the diagnostic sample. The full-scale extraction benchmark, including LoRA fine-tuning, is reported as projected rather than measured and remains an empirical extension of this work.
Chinese Translation
本文在经济学中的边际效用概念与两个机器学习构造——矩阵分解与Transformer语言模型的键值(KV)缓存——之间建立了一座理论桥梁。研究表明:评分矩阵的奇异值谱可视为潜在因子递减的边际效用表;投影协方差算子的特征值谱可视为模型已学表示的边际效用表;而缓存驱逐与低秩缓存压缩则可视为内存预算约束下效用最大化问题的实例。三者可归结为同一条分配规则:保留特征值超过约束条件的影子价格的头部维度。该框架被应用于地矿文档中结构化信息的自动化抽取,由此启发了一种多轮(multi-pass)推理协议、一种逐层的TIES模型合并方法,以及一项综合考虑抽取质量、定位漂移与能耗的选择策略,并通过条件风险价值(Conditional Value-at-Risk)项对漂移进行标量化。本文报告了两项实证贡献:其一,一个1120万参数的层级分类器,在单块GPU上约五分钟即可完成训练,在包含973份铀矿勘探文档的语料库的保留测试集上达到90.0%的一级准确率;相较之下,某专有模型在同一语料库的50份文档人工审计中达到92.0%,但每张卡片的推理延迟约为2000毫秒,而前者仅为2.62毫秒,且成本可忽略不计。其二,对均匀密度TIES合并的诊断揭示了一种可复现的退化模式:合并后的模型在五个地理位置不同的区域返回逐token完全相同的输出,同时声称高置信度;在采用逐层校准密度重新执行合并后,该特征在诊断样本上被消除。包括LoRA微调在内的全规模抽取基准是以预估而非实测方式报告的,仍是本工作待开展的实证扩展。
cs.AI / 58 / 2609.20077

Tailored to you: longitudinal effects of personalising language models

为你量身定制:个性化语言模型的纵向效应研究
Akbulut, Canfer, Breuch, Justine, Manzini, Arianna, Ibrahim, Lujain, Franklin, Matija, Patel, Roma, Gabriel, Iason, Lum, Kristian, Weidinger, Laura
Abstract
Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people's perception of and behaviour toward AI remain poorly understood. Most critically, downstream consequences outside the immediate human--AI interaction loop, such as effects on users' self-perceptions and interpersonal relationships, remain largely unexamined. In this study, we recruited 992 participants to complete daily advice-seeking interactions with language models over the course of five days, comparing outcomes from a non-personalised baseline against two personalisation approaches: memory-based (conditioned on prior conversational history) and survey-based (conditioned on information collected through a pre-study intake survey). We find that several changes in human-AI interaction over time are driven primarily by repeated exposure rather than personalisation itself. However, participants interacting with personalised models experienced differences in advice-seeking and information-sharing attitudes and behaviours: participants in the memory-based condition engaged in greater self-disclosure and rated the model as less creepy, while participants in the survey-based condition reported higher regret about having shared personal information with the AI. We conclude by highlighting the nuanced effects of different personalisation approaches on interaction outcomes, and discussing the implications of these findings for the responsible design and deployment of personalised AI systems.
Chinese Translation
人们对开发个性化语言模型的兴趣正迅速增长。虽然个性化通常被视为更好地满足多样化用户需求的机制,但与个性化模型的持续交互对人们认知和行为的影响仍知之甚少。最关键的是,在即时人机交互循环之外的下游后果,例如对用户自我认知和人际关系的影响,在很大程度上尚未被研究。在本研究中,我们招募了992名参与者,在五天内每天与语言模型进行寻求建议的交互,将非个性化基线与两种个性化方法的结果进行比较:基于记忆的个性化(以先前对话历史为条件)和基于问卷的个性化(以通过研究前入库问卷收集的信息为条件)。我们发现,人机交互随时间发生的若干变化主要由重复暴露驱动,而非个性化本身。然而,与个性化模型交互的参与者在寻求建议和信息共享的态度与行为上表现出差异:基于记忆条件下的参与者表现出更多的自我表露,并认为模型较不可怕;而基于问卷条件下的参与者则报告了对与AI分享个人信息的更高后悔感。最后,我们强调了不同个性化方法对交互结果的细微影响,并讨论了这些发现对负责任地设计和部署个性化AI系统的意义。
cs.AI / 59 / 2609.20080

A Proposal for an Agentic AI Architecture to Support Multi-Domain Decision-Making in the Brazilian Armed Forces

一种支持巴西武装力量多域决策的智能体AI架构提案
Braga, Gioliano de Oliveira, Barbieri, Sidnei, Ferraz, Ágney Lopes Roth, Sonaglio, Wagner Comin, Pereira Jr, Henrique Curi de Miranda e Lourenço Alves
Abstract
The growing complexity of multi-domain operational environments (land, aerospace, naval, cyber, and electromagnetic spectrum) has increased the volume and velocity of data reaching command-and-control (C2) centers, straining the observe-orient-decide-act (OODA) decision cycle. Artificial Intelligence (AI) systems currently employed in defense are, in general, reactive and isolated tools that still rely heavily on human operators to integrate information, assess scenarios, and formulate courses of action. This paper proposes a conceptual Agentic AI architecture for AI systems that can plan, access data sources, execute tools, and act autonomously and audibly, aimed at supporting decision-making across the three Brazilian Armed Forces (Navy, Army, and Air Force). Four application fronts are discussed (decision support, situational analysis, feasibility studies, and countermeasure suggestion), as well as the data and sensor access requirements and the security and permission safeguards necessary for responsible employment across administrative, strategic, operational, and tactical contexts.
Chinese Translation
多域作战环境(陆域、空天、海域、网络和电磁频谱)日益复杂,使得到达指挥控制(C2)中心的数据量和速度不断增加,给观察-判断-决策-行动(OODA)决策循环带来了压力。目前国防领域应用的人工智能(AI)系统大多是被动且孤立的工具,仍然严重依赖人类操作员来整合信息、评估态势并制定行动方案。本文提出一种概念性的智能体AI(Agentic AI)架构,该AI系统能够进行规划、访问数据源、执行工具,并以自主且可审计的方式行动,旨在支持巴西三军(海军、陆军和空军)的决策活动。文中讨论了四个应用方向(决策支持、态势分析、可行性研究以及对抗措施建议),以及数据与传感器访问需求和在行政、战略、作战和战术层面负责任应用所必需的安全与权限保障措施。
cs.AI / 60 / 2609.20089

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers:增强智能体强化学习中的工具集成推理
Liao, Wenjie, Zhao, Liangjie, Cao, Zehong
Abstract
Self-evolving methods reduce the need for human-annotated trajectories by allowing tool-using agents to generate their own training data. Yet existing methods typically separate trajectory generation from evaluation, relying on static verifiers that cannot adapt to emerging failure modes or self-consistency signals that may reinforce errors shared across trajectories. Jointly adapting planning, execution, and evaluation offers a promising alternative, but introduces a fundamental coordination challenge: each component continuously changes the data or feedback used to train the others. We address this challenge with \textbf{UnifiedPlayers}, a cooperative framework comprising a Planning Player that generates tasks, an Execution Player that produces multi-turn trajectories with Python tool calls, and an Evaluation Player that constructs executable verifiers. We design role-specific rewards that coordinate the three players toward a shared learning objective under GRPO. Across two model backbones and twelve reasoning benchmarks, UnifiedPlayers outperforms the strongest prior baseline by at least 3.5\% on mathematical reasoning and 3.9\% on general reasoning tasks. Moreover, the learned verifier achieves 84.2\% adversarial detection accuracy, while its reward signal exhibits 2.03$\times$ higher per-question variance than a self-consistency baseline, providing more discriminative verifications. These results highlight cooperation among specialized players as a promising path toward self-enhanced tool-integrated agents.
Chinese Translation
自进化方法通过允许使用工具的智能体自行生成训练数据,减少了对人工标注轨迹的需求。然而,现有方法通常将轨迹生成与评估分离,依赖于无法适应新出现失败模式的静态验证器,或依赖可能强化跨轨迹共有错误的自一致性信号。联合调整规划、执行与评估提供了一种有前景的替代方案,但引入了一个根本性的协调挑战:每个组件会持续改变用于训练其他组件的数据或反馈。我们通过 UnifiedPlayers 应对这一挑战,这是一个协作框架,包含生成任务的规划玩家(Planning Player)、通过调用 Python 工具产生多轮轨迹的执行玩家(Execution Player),以及构建可执行验证器的评估玩家(Evaluation Player)。我们设计了角色特定的奖励,在 GRPO 下将三个玩家协调到共同的学习目标上。在两个模型骨干和十二个推理基准上,UnifiedPlayers 在数学推理任务上超越最强先前基线至少 3.5%,在通用推理任务上超越至少 3.9%。此外,学习到的验证器达到了 84.2% 的对抗检测准确率,其奖励信号的每题方差是自一致性基线的 2.03 倍,提供了更具区分性的验证。这些结果表明,专业化玩家之间的协作是实现自增强工具集成智能体的一条有前景的路径。
cs.AI / 61 / 2609.20091

Solving Minimum Span Antibandwidth and Cyclic Antibandwidth Labeling Problems

求解最小跨度反带宽与循环反带宽标号问题
Xuan, Hieu Truong, Van, Khanh To
Abstract
The Antibandwidth and Cyclic Antibandwidth problems are NP-hard graph labeling problems that aim to maximize the minimum (cyclic) distance between labels assigned to adjacent vertices. Extensive research on these problems has resulted in a variety of mathematical formulations and computational approaches. However, their minimum span perspective, in which a prescribed minimum (cyclic) distance is fixed and the objective is to minimize the label span, has received comparatively little attention. In this paper, we consider this complementary perspective by introducing the Minimum Span Antibandwidth/Cyclic Antibandwidth Labeling (MSABL/MSCABL) problems and developing a unified Boolean Satisfiability (SAT)-based framework for solving them. The SAT-based framework formulates MSABL/MSCABL as a sequence of decision problems and exploits their monotonicity to accelerate the search process. We also consider two SAT solving strategies, parallel and incremental SAT solving: the former examines multiple candidate spans concurrently, while the latter reuses a single SAT instance while progressively restricting the label domain. The proposed approaches are evaluated on benchmark instances from the Harwell-Boeing Sparse Matrix Collection and compared with CPLEXCP, CPLEXMIP, and Gurobi. The results show that SAT-based approaches are highly competitive in solution quality, with the parallel approach performing best overall for MSCABL and the incremental approach for MSABL. With the no-hole constraint, they remain competitive with CPLEXCP and significantly outperform CPLEXMIP and Gurobi, particularly for MSCABL. These results demonstrate the effectiveness of SAT solving as an exact approach for MSABL and MSCABL.
Chinese Translation
反带宽(Antibandwidth)与循环反带宽(Cyclic Antibandwidth)问题是NP难的图标号问题,其目标是最大化相邻顶点所分配标号之间的最小(循环)距离。针对这些问题的大量研究产生了多种数学模型和计算方法。然而,其最小跨度视角——即固定一个给定的最小(循环)距离,目标是最小化标号跨度——却较少受到关注。本文从这一互补视角出发,提出了最小跨度反带宽/循环反带宽标号(Minimum Span Antibandwidth/Cyclic Antibandwidth Labeling,MSABL/MSCABL)问题,并开发了一个统一的基于布尔可满足性(SAT)的求解框架。该框架将MSABL/MSCABL表述为一系列判定问题,并利用其单调性加速搜索过程。我们还考虑了两种SAT求解策略:并行SAT求解和增量SAT求解。前者同时检查多个候选跨度,后者则复用单个SAT实例并逐步限制标号域。所提出的方法在Harwell-Boeing稀疏矩阵集合的基准实例上进行了评估,并与CPLEXCP、CPLEXMIP和Gurobi进行了比较。结果表明,基于SAT的方法在解的质量上极具竞争力:对于MSCABL,并行方法总体表现最佳;对于MSABL,增量方法表现最佳。在无孔约束下,这些方法与CPLEXCP保持相当的水平,并显著优于CPLEXMIP和Gurobi,尤其是对于MSCABL。这些结果证明了SAT求解作为MSABL和MSCABL精确求解方法的有效性。
cs.AI / 62 / 2609.20110

Perception, Layout, and Validation: Calibrated Confidence for Reliable Straight-Through Processing of Financial Documents

感知、布局与验证:用于财务文档可靠直通处理的校准置信度方法
Jin, Yichao, Wang, Yushuo, Han, Yuxuan, Sonia, Kwan Ching Yee, Song, Weiyang, Kent, Chiu Jin-Chun, Hwee, Wong Chong, Kiat, Wong Tiong, Ke, Kenneth Zhu, Zhao, Jingyuan
Abstract
Straight-through processing (STP) on extracted key-value fields from financial documents without human review requires a calibrated probability together with a bounded guarantee on the residual error of the auto-approved tier. The emergence of modern Vision Language Models (VLMs) provides an out-of-the-box capability for extracting the key-values, but their verbalized confidence signals are unreliable and weakly track field correctness. This paper introduces a decomposed confidence layer along three interpretable channels, including perception, layout, and validation. Together with a final conformal risk control, the score can be used for reliable STP of financial documents. The method is validated on three public datasets covering real invoices, synthetic invoices, and ad-buy forms, using two different VLM families (Qwen3.6-27B and Gemini-3.1-Flash-Lite). Our decomposed score consistently improves the separation of correct from incorrect extractions, substantially raising the AUROC from 0.54-0.74 for VLM verbalized signals to 0.90-0.99 with contributions from all three designed channels. Crucially for industrial deployment, this enables usable STP. The native VLM confidence signals could clear only 0.1%-7.0% of fields under risk control at a target error of <10%. In contrast, the proposed method auto-approves 49-72% of fields while holding the empirical error of the accepted tier at or below the target.
Chinese Translation
要在无需人工审核的情况下对从财务文档中提取的键值字段实现直通处理(Straight-Through Processing, STP),需要校准的概率以及自动批准层级残差误差的有界保证。现代视觉语言模型(Vision Language Models, VLMs)的出现为键值提取提供了开箱即用的能力,但其口头表达的置信度信号不可靠,且与字段正确性的关联较弱。本文引入了一种沿三个可解释通道分解的置信度层,包括感知、布局和验证。结合最终的一致性风险控制(conformal risk control),该评分可用于财务文档的可靠直通处理。该方法在三个公开数据集上进行了验证,涵盖真实发票、合成发票和广告投放表单,并使用了两种不同的VLM家族(Qwen3.6-27B和Gemini-3.1-Flash-Lite)。我们提出的分解评分持续改善了对正确与错误提取的区分能力,将AUROC从VLM口头表达信号的0.54-0.74大幅提升至0.90-0.99,且三个设计的通道均有贡献。对于工业部署而言,关键的是这使得可用的直通处理成为可能。在目标误差<10%的风险控制下,VLM原生的置信度信号仅能自动批准0.1%-7.0%的字段。相比之下,所提出的方法可自动批准49%-72%的字段,同时将已批准层级的实际误差保持在目标误差或以下。
cs.AI / 63 / 2609.20152

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

MTVA-Bench:评估级联语音代理中的语言模型
Mishra, Pritish, Kumar, Ishaan, Mandoli, Akshat, Kamath, Sudarshan
Abstract
Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations measure it either too broadly or too narrowly. End-to-end voice benchmarks score the full pipeline, so recognition errors and model errors mix into a single number. LLM benchmarks isolate the model but they do not evaluate what makes real phone calls hard, such as transcription issues, caller's voice being split across messages and the requirement that replies follow the language and script specified. We introduce the Multi-Turn Voice Agent Benchmark (MTVA-Bench), which evaluates the language model on the same conditions it faces inside a cascaded system. The caller is played by an LLM following a set of rubrics and tool calls are answered by a mock backend which responds to the arguments the model actually sent. The benchmark contains 49 agents working across 490 reviewed scenarios and supports 7 languages. Scoring is a combination of deterministic checks on tool calls with two LLM judges, one that scores scenario specific rules and one that grades conversation quality without access to the task. Both judges must cite specific messages from the transcript. Task and conversation scores are weighted equally, since a call can complete its task and still go badly for the caller. In a seven-model study, six of the models select the correct tool within 6.4 points of one another, but their overall scores span 24.4 points. Most of the gap comes from argument values, action ordering, rule compliance, and what the model says around its tool calls.
Chinese Translation
通常,大多数语音代理都是级联系统,即由ASR模型将呼叫者的音频转录为文本,由语言模型读取转录文本并决定说什么以及调用哪些后端工具,再由TTS模型播报回复。几乎所有决策都发生在语言模型中,但现有评估的测量方式要么过于宽泛,要么过于狭窄。端到端语音基准测试对整个流水线进行评分,因此识别错误和模型错误混杂在同一个数字中。LLM基准测试虽能隔离语言模型本身,却没有评估真实电话通话的难点,例如转录问题、呼叫者的语音被拆分在多条消息中,以及回复必须遵循指定的语言和文字系统。我们提出了多轮语音代理基准测试(MTVA-Bench),它在语言模型在级联系统中所面临的相同条件下对其进行评估。呼叫者由遵循一组评分标准的LLM扮演,工具调用则由模拟后端响应,该后端会针对模型实际发送的参数作出回应。该基准测试包含49个代理,涵盖490个经人工审核的场景,并支持7种语言。评分由确定性的工具调用检查与两个LLM评判相结合:其中一个评判场景特定的规则,另一个则在不知晓任务的情况下对对话质量进行评分。两个评判都必须引用转录中的具体消息。任务得分与对话得分权重相等,因为一次通话即使完成了任务,对呼叫者而言仍可能体验糟糕。在一项涉及七个模型的研究中,六个模型在正确工具选择上的差距在6.4分以内,但它们的总得分差距达24.4分。这一差距主要来源于参数取值、动作顺序、规则遵循情况,以及模型在工具调用前后的表述。
cs.AI / 64 / 2609.20177

PaGNet: A Panel-Aware GBDT--Neural Network for Multi-Target Corporate Tax Avoidance Proxy Forecasting

PaGNet:一种面向面板数据的GBDT—神经网络混合模型,用于企业税收规避代理指标的多目标预测
Song, Wonho, Kim, Hyungjoon
Abstract
Forecasting corporate tax avoidance proxies from firm--year panel data is challenging because predictive signals are distributed across short firm histories and related targets, while screening-oriented use requires transparent model behavior. We propose PaGNet (Panel-Aware GBDT--Neural Network), a two-branch hybrid that combines a LightGBM branch using panel-temporal summaries with a Panel-MLP branch using attention-pooled temporal aggregation and shared-trunk multi-task learning. A per-target validation-optimal blender produces both the final prediction and a compact branch-reliance diagnostic without trainable fusion parameters. On the KoTaP panel of 1{,}754 Korean listed firms from 2011--2024, PaGNet is evaluated under a leakage-free, shared-hyperparameter protocol across four feature regimes. In the direct-proxy-lag-excluded FS1 regime and the tax-history-augmented FS2 regime, accrual targets (TSTA, TSDA) route stably to the LightGBM branch, where PaGNet raises explained variance over the strongest of six baselines by roughly $0.08$--$0.11$ on the primary split. GETR often leans toward the neural branch, while CETR exposes a validation--test branch-selection mismatch rather than a stable branch assignment. A panel-flatten control shows that most accrual gains come from observed multi-year base-panel values, with PaGNet's panel-aware representation adding a smaller but directionally consistent refinement. Rolling-origin analysis confirms stable accrual routing, bounds ETR diagnostics to split-specific behavior, and identifies a far-horizon split where supervised models underperform naive persistence. PaGNet is therefore best viewed not as a universally superior tabular learner, but as a proxy-aware panel model that combines competitive forecasting with explicit per-target branch-reliance reporting.
Chinese Translation
基于公司—年度面板数据预测企业税收规避代理指标面临诸多挑战:预测信号分散在较短的公司历史和相关目标之中,而面向筛选的应用场景要求模型行为透明。我们提出了PaGNet(面板感知的GBDT—神经网络混合模型),这是一个双分支混合架构,将使用面板时间摘要的LightGBM分支与采用注意力池化时间聚合和共享主干多任务学习的Panel-MLP分支相结合。基于每个目标验证集最优的融合器在不引入可训练融合参数的情况下,同时给出最终预测和简洁的分支依赖诊断。在涵盖2011—2024年1754家韩国上市公司的KoTaP面板上,PaGNet在无泄漏、共享超参数的协议下、于四种特征体系中进行评估。在不包含直接代理指标滞后项的FS1体系和增强税收历史的FS2体系中,权责发生制类目标(TSTA、TSDA)稳定地路由至LightGBM分支,PaGNet在主要划分上相对于六个基线中最强者将解释方差提高了约0.08—0.11。GETR往往倾向于神经网络分支,而CETR则暴露出验证集—测试集之间的分支选择不一致问题,而非稳定的分支归属。面板展平对照实验表明,权责发生制目标的大部分提升来自可观测的多年基础面板数值,PaGNet的面板感知表示仅带来较小但方向一致的进一步改进。滚动起点分析证实了权责发生制目标的路由稳定性,将ETR诊断限定为特定划分的行为,并识别出一个远期预测划分,在该划分上监督模型表现逊于简单的持续性预测。因此,PaGNet不应被视为普遍优越的表格数据学习器,而应被视为一种代理指标感知的面板模型,它在提供有竞争力的预测能力的同时,能够显式报告每个目标的分支依赖情况。
cs.AI / 65 / 2609.20179

Sequential Contextual Fit Predicts Human Behavioural and Neural Dynamics Across Domains

序列上下文契合度可跨领域预测人类行为与神经动态
Sun, Kun, Wang, Rong
Abstract
Human perception, action and decision making unfold in sequences, but computational predictors are often domain-specific. This study computes and tests sequential contextual fit (SCF), an embedding-based measure of how well a current information state matches its recent context. The metric uses a simple recency-weighted similarity kernel and can be applied to words, sounds, visual scenes, affective states, choices, actions and neural representations. Across language processing, music-evoked emotion, a subset of audiovisual emotion EEG data, gambling decisions, human activity recognition and decision-related EEG, lower contextual fit predicted longer processing times, larger affective or behavioural transitions and stronger neural-state changes. These effects remained after controlling for established predictors including surprisal, reinforcement-learning prediction error, acoustic change, visual change and sensor change. SCF therefore provides a computational measurement layer for relating contextual compatibility to behavioural processing and cognitive/neural state-transition dynamics.
Chinese Translation
人类的感知、行动和决策以序列方式展开,但现有计算预测指标往往局限于特定领域。本研究计算并检验了序列上下文契合度(Sequential Contextual Fit, SCF),这是一种基于嵌入(embedding)的度量,用于衡量当前信息状态与其近期上下文的匹配程度。该指标采用简单的近因加权相似性核,可应用于词语、声音、视觉场景、情感状态、选择、动作以及神经表征。在语言加工、音乐诱发情绪、部分视听情绪脑电(EEG)数据、赌博决策、人类活动识别以及决策相关脑电等多类任务中,较低的上下文契合度预示着更长的加工时间、更大的情感或行为转变以及更强的神经状态变化。在控制了包括惊讶度(surprisal)、强化学习预测误差、声学变化、视觉变化和传感器变化等既有预测因子之后,这些效应依然显著。因此,SCF 提供了一个计算测量层面,可用于将上下文兼容性与行为加工以及认知/神经状态转变动态联系起来。
cs.AI / 66 / 2609.20200

JointMatch: A Unified Heterogeneous Graph Neural Solver for Large-Scale Ride-Sharing Matching

JointMatch:面向大规模拼车匹配的统一异构图神经求解器
Zhao, Kun, Chen, Xu
Abstract
Ride-sharing platforms must continuously decide which open requests to bundle into shared trips and which idle vehicles should serve them. The dominant academic approach decomposes this into two sequential matching problems -- request pairing first, then vehicle assignment -- and applies a separate solver to each. This decomposition is convenient computationally but loses revenue and scales poorly because the first stage commits to ride bundles before the available vehicles are known. We propose JointMatch, a learning-based framework that handles request pairing and vehicle assignment together on a single graph. The graph is sparsified by spatial proximity so that its size grows linearly rather than quadratically with the number of vehicles and requests, and a graph neural network scores all candidate decisions in one forward pass. On the New York City Yellow Taxi data, the framework already exceeds both the classical Blossom heuristic and a faithfully-trained two-stage GNN baseline -- often by a wide margin -- and at city scale (fleet 10000) it runs more than $20\times$ faster per dispatch epoch than either. A supervised training stage closes most of the remaining revenue gap, and a policy-gradient fine-tune aligns the trained model with realised revenue.
Chinese Translation
拼车平台必须持续决定将哪些未完成的请求组合为共享行程,以及调度哪些空闲车辆来服务这些行程。学术界的主流方法将其分解为两个顺序匹配问题——先进行请求配对,再进行车辆分配——并对每个阶段分别使用独立求解器。这种分解在计算上较为便捷,但由于第一阶段在可用车辆尚未确定的情况下就确定了行程组合,因此会损失收益且扩展性差。我们提出 JointMatch,一个基于学习的框架,可在单一图上同时处理请求配对与车辆分配。该图通过空间邻近性进行稀疏化,使其规模随车辆和请求数量线性而非二次增长;图神经网络(GNN)在前向传播中一次性对所有候选决策进行打分。在纽约市黄色出租车数据上,该框架已超越经典的 Blossom 启发式算法和经过忠实训练的两阶段 GNN 基线——且往往大幅领先——并在城市规模下(车队规模 10000)每个调度轮次的运行速度比两者快 20 倍以上。监督训练阶段弥补了大部分剩余的收益差距,而策略梯度微调则使训练后的模型与实际收益保持一致。
cs.AI / 67 / 2609.20261

When AI Agents Commit: Cognitive Serializability Across Data, Evidence, Policy, and Authority

当AI智能体提交时:跨数据、证据、策略与权威的认知可串行化
He, Jun, Yu, Deying
Abstract
Autonomous agents derive concrete mutations from database reads, retrieved evidence, policy, beliefs, and delegated authority. Those inputs may change while reasoning is in progress. Database isolation orders the submitted transaction; agentic transaction processing determines whether a proposal satisfies an executable contract. Neither guarantee establishes a common valid point for the mutation and its derivation inputs unless the contract represents the relevant predicates. Typed dependency tokens distinguish content integrity from applicability, and trusted mediation captures the values exposed to reasoning. Under strict Cognitive Serializability, committed effects admit a serial order and a logical event at which every value exposed to derivation is unchanged. The fences last until the runtime event that realizes the sealed durability domain. The weaker Effect-Compatible Cognitive Admission recertifies an effect against a simultaneously held current dependency vector and current policy without claiming to serialize the original stochastic derivation. TCT combines immutable versioned executable definitions, registry-derived authority plans, sealed envelopes, guard-first commit transactions, post-seal envelope- and witness-bound grants, co-committed receipts, idempotent grant finalization, and receipt-driven epistemic reconciliation. Complete registered footprints and a single growing phase induce an acyclic lock-point order over local guards and incompatible external reservations. The corresponding results give serializability conditions and an observational-equivalence boundary for zero-error soundness and positive progress. A falsification suite tests the implementation obligations: the prototype prevented all injected anomalies and added 3.22 ms mean commit overhead.
Chinese Translation
自主智能体从数据库读取、检索到的证据、策略、信念以及被委托的权威中推导出具体的变更操作。这些输入可能在推理进行过程中发生变化。数据库隔离机制对已提交的事务进行排序;而智能体事务处理则判定某个提议是否满足可执行的契约。除非该契约能够表达相关的谓词,否则这两种保证都无法为变更及其推导输入建立一个共同的有效时间点。类型化依赖令牌将内容完整性与适用性区分开来,可信中介则捕获暴露给推理过程的值。在严格的认知可串行性下,已提交的效果存在一个串行顺序和一个逻辑事件,在该事件处所有暴露给推导过程的值均保持不变。这些栅栏一直持续到实现密封持久性域的运行时事件为止。较弱的“效果兼容认知准入”则依据同时持有的当前依赖向量和当前策略对效果进行重新认证,而不声称对原始的随机推导过程进行串行化。TCT(可信提交事务)结合了不可变的版本化可执行定义、基于注册表的权威计划、密封信封、防护优先的提交事务、封印后的信封与见证绑定的授权、共同提交的回执、幂等的授权终结以及由回执驱动的认知调和。完整的注册足迹和单一的增长阶段在本地防护与不兼容的外部预留之上诱导出一个无环的锁点顺序。相应结果给出了零错误可靠性与正向进展的可串行化条件和观察等价边界。一个证伪测试套件检验了实现义务:原型系统阻止了所有注入的异常,并增加了3.22毫秒的平均提交开销。
cs.AI / 68 / 2609.20271

AI-Driven Real-Time Relay Optimisation in Smart Urban NR-V2X Networks via Learning-to-Optimise Graph Neural Networks

基于学习优化图神经网络的智能城市NR-V2X网络中AI驱动的实时中继优化
Amati, Giambattista, Mangiatordi, Federica, Pallotti, Emiliano, Angelini, Simone
Abstract
Reliable and low-latency communication is a fundamental requirement for smart city services and Industry 4.0 applications enabled by NR-V2X networks. However, limited Road-Side Unit (RSU) deployment and complex urban propagation conditions often prevent Connected and Automated Vehicles (CAVs) from maintaining stable connectivity. This paper proposes an AI-driven Learning-to-Optimise (L2O) framework based on Graph Neural Networks (GNNs) for real-time multi-hop relay selection in NR-V2X systems. The vehicular network is modelled as a graph, where nodes represent CAVs and RSUs, and edges encode radio-link characteristics. An offline Mixed-Integer Linear Programming (MILP) formulation provides optimal relay decisions used as supervision for training an edge-aware Graph Isomorphism Network with Edge Features (GINE). Extensive experiments on realistic urban datasets demonstrate that the proposed approach achieves near-optimal connectivity performance, recovering up to 11.3% connectivity gain, while reducing execution time by orders of magnitude (up to 100 x speed-up) compared to MILP. The framework enables scalable and real-time network control, making it suitable for smart city and Industry 4.0 deployments.
Chinese Translation
可靠且低时延的通信是由NR-V2X网络赋能的智慧城市服务和工业4.0应用的基本需求。然而,路侧单元(RSU)部署有限以及复杂的城市传播环境,往往使网联自动驾驶汽车(CAV)难以保持稳定的连接。本文提出了一种基于图神经网络(GNN)的AI驱动学习优化(Learning-to-Optimise, L2O)框架,用于NR-V2X系统中的实时多跳中继选择。车辆网络被建模为图,其中节点表示CAV和RSU,边编码无线链路特性。通过离线的混合整数线性规划(MILP)公式获得最优中继决策,作为监督信号训练具有边特征的图同构网络(GINE)。在真实城市数据集上的大量实验表明,所提方法可实现接近最优的连接性能,最多可恢复11.3%的连接增益,同时与MILP相比,执行时间降低数个数量级(最高加速100倍)。该框架支持可扩展的实时网络控制,适用于智慧城市和工业4.0部署场景。
cs.AI / 69 / 2609.20277

JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations

JEPA-WAM:通过JEPA潜在表示将生成的视觉指令与世界动作模型相连接
Liu, Tianbin, Zhu, Jian, Su, Taiyi, Zhang, Jianjun, Ma, Chong, Huang, Zitai, Xu, Yi
Abstract
World Action Models (WAMs) have demonstrated strong robotic manipulation capabilities by augmenting pretrained video generative models with action experts. However, current WAMs still show limited instruction-following ability when conditioned solely on text instructions. We argue that this limitation stems in part from a structural imbalance in robot-learning data: rich visual-action trajectories are often paired with sparse and repetitive language annotations, allowing policies to identify tasks from visual context and motion regularities rather than grounding the instruction itself. To address this limitation, we introduce JEPA-WAM, which augments each text instruction with a bank of stochastically generated visual instructions, providing diverse visual cues for instruction following. Specifically, JEPA-WAM uses an off-the-shelf text-to-image generator to sample multiple task-completion images conditioned on the text instruction, without training the generator. Although these generated images may differ from the current visual scene in appearance and layout, they remain semantically aligned with the instruction and serve as visual goal references. To focus on task-level semantics beyond appearance, we encode these references with a frozen V-JEPA 2.1 encoder. The resulting dense goal representations are compressed into compact goal tokens that condition both the video and action experts through cross-attention. We further construct a real-robot instruction-following benchmark covering in-distribution, out-of-distribution scene, and out-of-distribution instruction settings. On this benchmark, JEPA-WAM achieves success rates of 87.3%, 74.5%, and 80.9% in these three settings, outperforming {\pi}0 and Fast-WAM by at least 10.0, 27.3, and 14.5 percentage points, respectively.
Chinese Translation
世界动作模型(World Action Models, WAMs)通过在预训练视频生成模型基础上增加动作专家,展现出了强大的机器人操作能力。然而,当前的世界动作模型在仅以文本指令为条件时,指令遵循能力仍然有限。我们认为,这一限制部分源于机器人学习数据的结构性失衡:丰富的视觉-动作轨迹往往与稀疏且重复的语言标注相配对,使得策略能够从视觉上下文和运动规律中识别任务,而无需真正将指令本身作为依据。为了解决这一限制,我们提出了JEPA-WAM,该方法为每条文本指令配以一组随机生成的视觉指令,为指令遵循提供多样化的视觉线索。具体而言,JEPA-WAM使用现成的文本到图像生成器,以文本指令为条件采样多张任务完成图像,而无需训练该生成器。尽管这些生成的图像在外观和布局上可能与当前视觉场景不同,但它们在语义上与指令保持一致,可作为视觉目标参考。为了关注超越外观层面的任务级语义,我们使用冻结的V-JEPA 2.1编码器对这些参考进行编码。所得到的稠密目标表示被压缩为紧凑的目标token,通过交叉注意力同时为视频专家和动作专家提供条件。我们还构建了一个真实机器人指令遵循基准,涵盖分布内、分布外场景和分布外指令三种设置。在该基准上,JEPA-WAM在这三种设置中分别取得了87.3%、74.5%和80.9%的成功率,分别比{\pi}0和Fast-WAM至少高出10.0、27.3和14.5个百分点。
cs.AI / 70 / 2609.20301

AgentPProf: Semantic Profiler for Long Horizon AI Agents

AgentPProf:面向长时程AI智能体的语义剖析器
Zheng, Yusheng, Chang, Chaokun, Mao, Yu, Wu, Tianyuan, Huang, Yuxi, Ma, Tao, Mao, Wenan, Cheng, Shuyi, Quinn, Andi, Wang, Wei
Abstract
AI agents increasingly orchestrate long-running activities with users, tools, and system resources for days and weeks. To improve agent quality, safety, and cost efficiency, developers need to determine where failures happen, what triggers unsafe effects, and which tasks consume the most budget, then optimize those tasks. In systems software, profiling answers similar questions by aggregating resource consumption and attributing it to responsible code paths to identify hotspots. Yet existing agent observability tools focus on per-execution debugging and tracing rather than cross-run, long term profiling, making these questions difficult to answer at scale. Agent observability needs profiling, not only debugging, but profiling agents is challenging: the responsible entities are task intent like diagnose authentication, compare branches rather than code paths, and lack stable identifiers for aggregation. We propose a semantic operation stack model that adapts profiling to agent trajectories. Uniform operations represent all activities, and operation stacks replace the runtime call stack, enabling hierarchical attribution at different granularities. We observe that an agent's task occupies a contiguous span and decomposes into subtasks, so we introduce recursive operation segmentation, which recursively splits trajectories at task boundaries. AgentPProf is a profiler that aggregates agent trajectories into pprof-compatible profiles, enabling flame graph visualization and analysis. AgentPProf reaches 0.764 $B^3$ F1 against human annotations on CodeTraceBench. On three problem-localization benchmarks, the profile raises MAP by up to 56%, demonstrating that it effectively attributes resources, locates problems, and helps optimize token cost at practical profiling cost. AgentPProf is available at https://github.com/eunomia-bpf/agentsight.
Chinese Translation
AI智能体日益需要与用户、工具和系统资源协同编排持续数天乃至数周的长时运行活动。为了提升智能体的质量、安全性和成本效率,开发者需要确定故障发生在何处、什么触发了不安全效应,以及哪些任务消耗了最多的预算,进而对这些任务进行优化。在系统软件领域,剖析通过聚合资源消耗并将其归因到相应的代码路径以识别热点,从而回答类似的问题。然而,现有的智能体可观测性工具侧重于单次执行的调试与追踪,而非跨运行、长期的分析,使得这些问题难以大规模地回答。智能体可观测性需要的是剖析而不仅仅是调试,但对智能体进行剖析充满挑战:其责任实体是任务意图(如“诊断认证问题”“比较分支”)而非代码路径,且缺乏用于聚合的稳定标识符。我们提出了一种语义操作栈模型,将剖析方法适配到智能体轨迹上。统一的操作表示所有活动,操作栈取代运行时调用栈,从而支持在不同粒度上进行层次化归因。我们观察到智能体的任务占据一段连续的时间跨度,并可分解为子任务,因此引入了递归操作分割方法,在任务边界处递归地切分轨迹。AgentPProf 是一个剖析器,它将智能体轨迹聚合为兼容 pprof 的性能剖析文件,支持火焰图可视化与分析。在 CodeTraceBench 上,AgentPProf 相对于人工标注达到了 0.764 的 $B^3$ F1 分数。在三个问题定位基准上,该剖析文件将 MAP 提升了最多 56%,表明它能有效进行资源归因、定位问题,并以实际的剖析成本帮助优化 token 开销。AgentPProf 已发布于 https://github.com/eunomia-bpf/agentsight。
cs.AI / 71 / 2609.20304

Diagnose, Recover, Certify: Task Readiness under Hidden Dynamics Changes

诊断、恢复、认证:隐藏动力学变化下的任务就绪性
Kiet, Nguyen Viet Tuan, Binh, Huynh Thi Thanh
Abstract
A deployed control policy can conceal consequential dynamics changes: an actuator may lose effectiveness without affecting the current task when the policy rarely excites it, despite being critical for a future task that has not yet been specified. We introduce task readiness under dormant dynamics drift, a decision problem that unifies active change diagnosis and post-change control recovery under a limited, task-agnostic interaction budget. An agent must identify whether and where local dynamics have changed, use a small number of informative interactions to characterize the change before downstream task identity is revealed, and subsequently provide each candidate task with either a recovered policy and a calibrated lower bound on its achievable return or an abstention decision to a safe fallback. We propose Evidence-Gated Matched-Pulse Transport, an intervention-based Bayesian procedure that couples fault localization with estimation of actuator effectiveness through a shared matched-response representation, thereby preserving diagnostic reliability while converting localized evidence into recovery-relevant uncertainty. This uncertainty is propagated to task-conditioned policy selection and readiness certification, enabling deployment decisions that explicitly trade off expected performance, confidence, and fallback use. We evaluate the resulting framework on a diverse suite of dormant-actuator benchmarks spanning multiple simulators, under a protocol that separates diagnosis from capability recovery, scores deployment by readiness coverage, selective risk, and interaction cost as well as return, and identifies the fault regimes in which transported evidence is decisive.
Chinese Translation
已部署的控制策略可能掩盖重大的动力学变化:当策略很少激励某个执行器时,该执行器可能在不影响当前任务的情况下失去有效性,尽管它对于尚未确定的未来任务至关重要。我们提出了休眠动力学漂移下的任务就绪性(task readiness under dormant dynamics drift)这一决策问题,它将主动变化诊断与变化后控制恢复统一在一个有限的、与任务无关的交互预算之下。智能体必须识别局部动力学是否以及在哪里发生了变化,在下游任务身份被揭示之前,使用少量信息性交互来刻画该变化,并随后为每个候选任务提供以下二者之一:一个恢复的策略以及其可达回报的校准下界,或者一个转向安全回退的弃权决策。我们提出证据门控匹配脉冲传输(Evidence-Gated Matched-Pulse Transport),这是一种基于干预的贝叶斯方法,通过共享的匹配响应表示将故障定位与执行器有效性估计相耦合,从而在保持诊断可靠性的同时,将局部化证据转化为与恢复相关的不确定性。该不确定性被传播到任务条件化的策略选择与就绪性认证中,使得部署决策能够显式地在期望性能、置信度与回退使用之间进行权衡。我们在跨多个模拟器的多样化休眠执行器基准套件上评估了所提出的框架,该协议将诊断与能力恢复分离,以就绪覆盖率、选择性风险、交互成本以及回报对部署进行评分,并识别出传输证据起决定性作用的故障情形。
cs.AI / 72 / 2609.20323

NeuSOGA3D: A Neuro-Symbolic Framework for Explainable 3D Geometric Reconstruction

NeuSOGA3D:一个面向可解释三维几何重建的神经-符号框架
Li, Qingde, Hong, Qingqi, Li, Zihan, Tian, Jie
Abstract
Three-dimensional reconstruction from unorganized point clouds remains a challenging problem in computer vision, geometric modeling, and computer-aided design. While neural implicit methods achieve impressive reconstruction accuracy, geometry is typically encoded in latent representations that limit interpretability and reuse within engineering workflows. We present NeuSOGA3D (Neuro-Symbolic Geometric Abstraction in 3D), a hybrid framework that combines learned perceptual priors inherited from NeuSOGA with explicit symbolic geometric reasoning. The method projects point clouds onto principal orthographic planes, constructs symbolic implicit spline representations from the resulting observations, and fuses them through shape-preserving constructive solid geometry operations to generate a coarse visual hull. Additional geometric detail is recovered through cross-sectional decomposition and volumetric reconstruction using Partial Shape-Preserving Splines. Unlike conventional neural implicit approaches, NeuSOGA3D progressively transforms observations into explicit symbolic entities, including control polygons, implicit spline fields, cross-sections, and volumetric lofts. Experiments on all forty categories of the ModelNet40 benchmark demonstrate the ability of the framework to recover structurally meaningful and CAD-compatible geometric representations from diverse point-cloud observations. The results highlight the potential of combining learned perception with symbolic geometric reasoning for explainable geometric intelligence.
Chinese Translation
从无序点云中进行三维重建仍然是计算机视觉、几何建模和计算机辅助设计中的一个具有挑战性的问题。尽管神经隐式方法实现了令人瞩目的重建精度,但几何信息通常被编码在潜在表示中,这限制了其可解释性以及在工程工作流程中的复用性。我们提出了NeuSOGA3D(三维神经-符号几何抽象,Neuro-Symbolic Geometric Abstraction in 3D),这是一个混合框架,它将继承自NeuSOGA的学习感知先验与显式的符号几何推理相结合。该方法将点云投影到主正交平面上,根据所得观测结果构建符号化的隐式样条表示,并通过保持形状的构造实体几何(CSG)操作将其融合,以生成粗略的视觉外壳(visual hull)。随后,通过横截面分解以及基于部分形状保持样条(Partial Shape-Preserving Splines)的体积重建来恢复更多几何细节。与传统的神经隐式方法不同,NeuSOGA3D将观测结果逐步转化为显式的符号实体,包括控制多边形、隐式样条场、横截面以及体积放样。在ModelNet40基准全部四十个类别上的实验表明,该框架能够从多样的点云观测中恢复出具有结构意义且与CAD兼容的几何表示。这些结果凸显了将学习感知与符号几何推理相结合以实现可解释几何智能的潜力。
cs.AI / 73 / 2609.20334

Structured Four-Stage Legal Translation: From Natural-Language Traffic Rules to PROLOG

结构化四阶段法律翻译:从自然语言交通规则到 PROLOG
Zin, May Myo, Fungwacharakorn, Wachara, Satoh, Ken, Nitta, Katsumi
Abstract
Traffic regulations are written for human interpretation and therefore rely on shared background knowledge and flexible phrasing, which inherently introduce ambiguity, context dependence, and semantic underspecification. These linguistic characteristics conflict with the precision required by computational reasoning engines such as Prolog, which demand explicit logical structure. This study evaluates two baseline translation approaches, Natural Language to Prolog ($NL\rightarrow Prolog$) and Logical English to Prolog ($LE\rightarrow Prolog$), and introduces a new reasoning-guided translation framework called Structured Four-Stage Legal Translation ($S4L\rightarrow Prolog$). The proposed S4L framework performs semantic role extraction, scene completion, logical mapping, and Prolog rule generation within a single guided prompt, enabling direct translation of raw traffic rules into executable logic without human intervention. A benchmark consisting of twenty real-world traffic rules was used to evaluate each approach in terms of syntactic validity, semantic correctness, and logical completeness. $S4L\rightarrow Prolog$ achieves the highest accuracy, correctly formalizing 75 percent of the rules, while $NL\rightarrow Prolog$ reaches 60 percent and $LE\rightarrow Prolog$ reaches 55 percent. Qualitative analysis further shows that S4L captures implicit causal relations, deontic modality, and exception structure more reliably than the baselines. These results demonstrate that structured reasoning prompts can substantially improve the reliability of natural-language-to-logic translation for legal and safety-critical applications.
Chinese Translation
交通法规是为人类解读而编写的,因此依赖于共享的背景知识和灵活的表述方式,这天然地引入了歧义、上下文依赖性和语义欠明确性。这些语言特性与 Prolog 等计算推理引擎所要求的精确性相冲突,因为后者需要显式的逻辑结构。本研究评估了两种基线翻译方法,即自然语言到 Prolog($NL\rightarrow Prolog$)和逻辑英语到 Prolog($LE\rightarrow Prolog$),并提出了一种新的推理引导的翻译框架,称为结构化四阶段法律翻译($S4L\rightarrow Prolog$)。所提出的 S4L 框架在单个引导式提示中完成语义角色提取、场景补全、逻辑映射和 Prolog 规则生成,能够将原始交通规则直接翻译为可执行逻辑,无需人工干预。研究使用由二十条真实交通规则组成的基准测试,从句法有效性、语义正确性和逻辑完备性三个方面评估每种方法。$S4L\rightarrow Prolog$ 取得了最高的准确率,正确形式化了 75% 的规则,而 $NL\rightarrow Prolog$ 达到 60%,$LE\rightarrow Prolog$ 达到 55%。定性分析进一步表明,与基线方法相比,S4L 能更可靠地捕捉隐含的因果关系、道义情态和例外结构。这些结果表明,结构化推理提示能够显著提升自然语言到逻辑翻译的可靠性,适用于法律和安全关键应用。
cs.AI / 74 / 2609.20349

A Qualitative Model for Reasoning about Path and Support

一种用于路径与支撑推理的定性模型
Jaiswal, Abhishek, Falomir, Zoe
Abstract
Spatial reasoning abilities correlate strongly with performance in STEM fields. Games offer a compelling medium for training these critical skills in developing children who have a natural proclivity for play. However, to facilitate human-like tutoring and player guidance, these games require an AI agent capable of making commonsense inferences from spatial events. Qualitative reasoning (QR) models appear to be a suitable framework for these application domains. As these models reason in symbolic representations, they can seamlessly translate game states into interpretable feedback for human-like player guidance. This paper introduces a hybrid qualitative model designed for Camelot Jr., a block-puzzle game that requires constructing multi-level bridges to connect two avatars stationed on separate towers. The game poses a challenge for the player, who must make platforms stable, plan their path, and ensure they use all the provided blocks. To handle the precise physics required by the domain, we integrate a mathematical center-of-mass stability logic to guide our qualitative solver. Our work facilitates spatial skill training in Camelot Jr. and contributes to the development of human-centric, explainable game-playing agents.
Chinese Translation
空间推理能力与STEM领域的表现密切相关。游戏为培养儿童的这些关键技能提供了一种极具吸引力的媒介,因为处于发展阶段的儿童天生喜欢玩耍。然而,为了实现类人的辅导和玩家引导,这些游戏需要一个能够从空间事件中进行常识性推理的AI智能体。定性推理(Qualitative Reasoning, QR)模型似乎是适用于这些应用领域的框架。由于这些模型基于符号表示进行推理,它们可以将游戏状态无缝地转化为可解释的反馈,从而实现类人的玩家引导。本文介绍了一个为Camelot Jr.设计的混合定性模型,这是一款积木解谜游戏,要求搭建多层桥梁以连接驻留在不同塔楼上的两个角色。该游戏对玩家提出了挑战:玩家必须使平台保持稳定、规划路径,并确保使用所有提供的积木块。为了处理该领域所需的精确物理特性,我们集成了一种数学上的质心稳定性逻辑来指导我们的定性求解器。我们的工作促进了Camelot Jr.中空间技能的训练,并为开发以人为中心的可解释游戏智能体做出了贡献。
cs.AI / 75 / 2609.20358

Generating Heterogeneous 3D Geological Microstructures from 2D Images via a Stable Diffusion-Adversarial Model

基于稳定扩散-对抗模型从二维图像生成非均质三维地质微观结构
Aouf, Ali, Laloy, Eric, Rogiers, Bart, De Vleeschouwer, Christophe
Abstract
Characterizing the physical properties of clay and cementitious materials matters across many fields, from materials science to geological waste disposal. Property simulation typically calls for 3D imaging, which is expensive, not always accessible, and technically limited for certain materials. Recent progress in deep generative models offers a way around this, reconstructing 3D volumes from the more easily acquired 2D images. Among GAN-based methods for 3D microstructure generation, SliceGAN has shown strong results for homogeneous isotropic and anisotropic systems. It struggles, however, to capture the finer detail of more complex heterogeneous microstructures, which motivates alternative generative frameworks. We introduce a hybrid approach that draws on the stability and generation quality of denoising diffusion models. Since no 3D ground truth is available, we replace the standard denoising loss with an adversarial loss, which yields a stable training process in our experiments. We show that the resulting model generates microstructures of varying complexity with minimal slice artefacts and close agreement with ground-truth phase fractions and structural descriptors.
Chinese Translation
表征黏土和胶凝材料的物理性质在从材料科学到地质废物处置的众多领域都至关重要。性质模拟通常需要三维成像,但三维成像成本高昂、并非总是可行,且对某些材料存在技术限制。深度生成模型的最新进展为此提供了一条解决途径,即从更易获取的二维图像重建三维体数据。在基于生成对抗网络(GAN)的三维微观结构生成方法中,SliceGAN在均质各向同性和各向异性体系中表现优异。然而,它在捕捉更复杂的非均质微观结构的精细细节方面存在困难,这促使人们探索其他生成框架。我们提出了一种混合方法,借鉴了去噪扩散模型的稳定性和生成质量。由于缺乏三维真实数据作为基准,我们用对抗损失替代标准的去噪损失,从而在实验中获得了稳定的训练过程。结果表明,所得模型能够生成不同复杂度的微观结构,其切片伪影极少,且在相分数和结构描述符方面与真实数据高度吻合。
cs.AI / 76 / 2609.20449

The Organization of Inference: Information, Resource Constraints, and AI Production

推理的组织:信息、资源约束与AI生产
Zhang, Yukun, Xu, Kemu, Chen, Yishen
Abstract
The economic value of inference depends on how capacity and task information are distributed across stages of AI production. We study these organizational margins using controlled workflow experiments on externally verified software-engineering tasks. In two matched resource panels, direct execution records the same success rate of 59.6 percent at logical-token ceilings of 12,000 and 24,000, while success under information-constrained planning rises from 36.2 to 51.2 percent. The planning disadvantage narrows by 15.0 percentage points (95 percent task-cluster bootstrap interval: 4.2 to 25.8). A strict read-only planning campaign varies whether the planner sees the task issue. At 12,000 tokens, issue access raises success by about 16 percentage points over issue-hidden planning. Compared with direct execution, task-informed planning is about 10 points lower at 12,000 tokens; at 24,000 tokens, it shows a 29.6-point advantage. In the resource panels, direct execution uses substantially less than either ceiling, while the planning workflow's binding rate falls from 46.2 to 0.8 percent and downstream execution accounts for 89.9 percent of the increase in total use. Scale determines the capacity available to a system; workflow and information structure shape the productive value
Chinese Translation
推理的经济价值取决于能力和任务信息在AI生产各阶段之间的分布方式。我们通过在外部验证的软件工程任务上进行受控工作流实验来研究这些组织性边际。在两个匹配的资源面板中,直接执行在逻辑token上限分别为12,000和24,000时保持59.6%的相同成功率,而信息受限规划下的成功率则从36.2%提升至51.2%。规划的劣势缩小了15.0个百分点(95%任务聚类bootstrap置信区间:4.2至25.8)。在一项严格的只读规划活动中,我们变化规划者是否能看到任务问题(issue)。在12,000 token时,与问题隐藏规划相比,问题可见使成功率提高约16个百分点。与直接执行相比,任务知情规划在12,000 token时低约10个百分点;而在24,000 token时,它展现出29.6个百分点的优势。在资源面板中,直接执行的使用量大幅低于任一上限,而规划工作流的约束触顶率从46.2%降至0.8%,且下游执行占总使用量增长的89.9%。规模决定了系统可用的能力;而工作流与信息结构则塑造了生产的有效价值。
cs.AI / 77 / 2609.20455

SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback

SkillAA:基于归因引导的技能图更新方法,结合针对性验证与回滚机制
Shang, Ziqiao, Ge, Ling-Yue, Guo, Lan-Zhe
Abstract
External skills provide domain procedures without parameter updates, but existing methods often edit skills directly from failed rollouts without structured routing from an observed failure to an editable location; existing skill graphs also underuse semantic boundaries, object addresses, and topological dependencies for skill retrieval, targeted updating, and scoped validation. We introduce SkillAA (Skill Abductive Attribution), a structured skill-optimization framework for frozen language models. It represents skill applicability, execution, and composition in a unified graph, allowing the same structure to support skill selection, attribution-guided repair, and update validation. SkillAA contrasts successful and failed executions to route candidate repairs to specific graph objects, updates only the selected local structure, and uses Local and Big Gates to screen candidate changes before commitment. With gpt-5.6-sol, SkillAA reaches 81.5%, 66.7%, and 91.2% on SearchQA, LiveMath, and DocVQA, respectively, and attains the highest observed mean in every main setting. These results support the utility of attribution-guided graph editing and graph-scoped validation.
Chinese Translation
外部技能能够在无需参数更新的情况下提供领域操作流程,但现有方法通常直接根据失败的执行结果对技能进行修改,缺乏从观察到的失败到可编辑位置的结构化路由;现有的技能图也未能充分利用语义边界、对象地址和拓扑依赖关系来实现技能检索、针对性更新和范围化验证。我们提出了 SkillAA(Skill Abductive Attribution,技能溯因归因),这是一个面向冻结语言模型的结构化技能优化框架。它将技能的适用性、执行和组合表示在一个统一的图中,使同一结构能够支持技能选择、归因引导的修复以及更新验证。SkillAA 通过对比成功与失败的执行,将候选修复路由到图中特定的对象,仅更新所选定的局部结构,并使用局部门(Local Gate)和大门(Big Gate)在提交变更前对候选修改进行筛选。基于 gpt-5.6-sol,SkillAA 在 SearchQA、LiveMath 和 DocVQA 上分别达到 81.5%、66.7% 和 91.2%,并在所有主要设置中取得观测到的最高平均成绩。这些结果支持了归因引导的图编辑和图范围化验证的有效性。
cs.AI / 78 / 2609.20474

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

智能体框架如何创造价值?有状态LLM智能体中的规划信息与发布控制
Zhang, Yukun, Xu, Kemu, Chen, Yishen
Abstract
Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $\tau^2$-bench. The primary comparison pairs prewritten task-specific plans (Fixed) with shuffled policy text matched in word count (Sham), isolating the contribution of guidance content. Across 265 matched cells, Fixed improves oracle-verified success by 7.17 percentage points (90\% task-clustered bootstrap interval, 1.15--13.36 points), with gains concentrated in higher-complexity tasks. A read-only terminal verifier rejects 61\% of Retail oracle-invalid episodes while withholding 17\% of correct ones, at less than one cent of additional cost per episode. Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate---and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.
Chinese Translation
智能体框架(agent harness)提供规划指导、组织执行并检查任务完成情况。我们在τ²-bench的两个Retail实验和一个Airline试点实验中研究了这些组件如何影响成功率、错误接受率和成本。核心对比是将预先编写的任务特定规划(Fixed)与字数匹配但打乱的策略文本(Sham)进行比较,以隔离指导内容的贡献。在265个匹配单元格中,Fixed将经过oracle验证的成功率提高了7.17个百分点(90%任务聚类自助法置信区间为1.15–13.36个百分点),且提升主要集中在复杂度较高的任务上。一个只读的终端验证器可拒绝61%的Retail oracle无效回合,同时误拦17%的正确回合,每个回合的额外成本不足一美分。哪个组件更重要取决于错误接受所造成的损失:在低责任情形下,规划带来的收益占主导;在高责任情形下,验证器避免的误通过占主导——而且独立验证器以极小的成本就能获得规划加验证完整技术栈中几乎全部的误通过收益。
cs.AI / 79 / 2609.20519

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

SoL-Pi:递归扩展自动研究循环以实现高效的Agent框架
Liu, Haozhe, Ye, Tian, Gao, Sensen, Cao, Qihang, Li, Yitong, Zhuge, Mingchen, Wang, Duomin, Zhang, Ruihua, Luo, Ping, Bian, Jiawang, Zhu, Lei, Zhu, Ligeng, Xie, Enze, Han, Song
Abstract
As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement. We take an RSI-inspired approach at the harness layer, scaling auto-research loops across increasingly numerous and diverse environments for harness rollouts. At this scale, the process yields reusable improvements that transfer beyond their development setting, moving automated harness discovery toward production-level outcomes. Four mechanisms survive selection and form SoL-Pi, spanning action execution, context compaction, observation handling, and delegated reading. On the 51-task EdgeBench evaluation, SoL-Pi achieves performance comparable to Pi across GPT-5.6 Sol and Opus 5 while reducing recorded token traffic by 44.7-49.0% and API cost by about one third. In other words, estimated hourly savings are \$8.75-\$13.50 relative to native Codex and Claude Code harnesses, and \$4.36-\$5.71 relative to Pi.
Chinese Translation
随着编码智能体(coding agents)从受监督的代码补全转向无人值守的全天候探索,其工作范围从孤立的预测扩展到涵盖推理、工具使用和反馈的长轨迹。因此,Token效率成为扩展递归自我改进(RSI)的关键。我们在框架(harness)层采用受RSI启发的方法,在日益众多和多样的环境中扩展用于框架测试的自动研究循环。在这一规模下,该过程产生了可复用的改进,这些改进能够迁移到其开发环境之外,推动自动化框架发现走向生产级成果。四种机制在筛选中存留下来并构成SoL-Pi,涵盖动作执行、上下文压缩、观测处理和委托阅读。在包含51个任务的EdgeBench评估中,SoL-Pi在GPT-5.6 Sol和Opus 5上实现了与Pi相当的性能,同时将记录的Token流量降低44.7%–49.0%,API成本降低约三分之一。换言之,相对于原生的Codex和Claude Code框架,估计每小时可节省8.75–13.50美元,相对于Pi每小时可节省4.36–5.71美元。
cs.AI / 80 / 2609.20535

FreqCondNorm: Towards Cross-domain Predictive Maintenance through a Frequency-Conditioned Transformer Foundation Model

FreqCondNorm:基于频率条件化Transformer基础模型实现跨域预测性维护
Raounak, Zaynab, LHermine, Camille, Zeng, Zhiguo
Abstract
Deep learning predictive maintenance models suffer from poor transferability across machines and operating conditions, especially when labelled data are scarce and signals span five orders of magnitude in sampling frequency (1 Hz to ~100 kHz). We propose FreqCondNorm, a Transformer-based architecture that introduces a FiLM-style frequency-conditioned normalization layer to unify heterogeneous time-series within a single model. The architecture is pretrained on five public predictive maintenance datasets (CWRU, MFPT, UOC18, PRONOSTIA, CMAPSS) using masked auto-encoding and contrastive learning with balanced domain sampling. On fault diagnosis, the model achieves 99.2% accuracy on CWRU (+6.4 pp over CNN) and 82.1% zero-shot accuracy on MFPT, demonstrating strong transfer across sampling frequencies. However, the approach does not improve remaining useful life prediction, suggesting a mismatch between pretraining and RUL objectives that warrants future investigation.
Chinese Translation
深度学习预测性维护模型在跨机器和跨运行工况下的迁移能力较差,尤其是在标注数据稀缺且信号采样频率跨越五个数量级(1 Hz至约100 kHz)的情况下。我们提出了FreqCondNorm,这是一种基于Transformer的架构,引入了FiLM风格的频率条件化归一化层,从而在单一模型内统一异构时间序列。该架构在五个公开的预测性维护数据集(CWRU、MFPT、UOC18、PRONOSTIA、CMAPSS)上进行预训练,采用掩码自编码和对比学习,并配合均衡的域采样。在故障诊断任务中,该模型在CWRU数据集上达到99.2%的准确率(比CNN提升6.4个百分点),在MFPT数据集上达到82.1%的零样本准确率,展现出跨采样频率的强大迁移能力。然而,该方法并未提升剩余使用寿命(RUL)预测性能,这表明预训练目标与RUL目标之间存在不匹配,值得未来进一步研究。
cs.AI / 81 / 2609.20538

Refuse, Decompose, Refresh: A Claim-Safe Protocol for Closed-Loop AI Evaluation

拒绝、分解、刷新:一种面向闭环AI评估的声明安全协议
Zhu, Peiying, Chang, Sidi
Abstract
An AI evaluation can be perfectly reproducible and still support the wrong claim. This risk is acute in closed-loop systems: policy determines visited states, observable components, and which failures leave a measurable trace. We propose a claim-safe protocol with three actions. Refuse: abstain when a clean reference stream or matched runtime comparison lacks support. Decompose: report protocol execution, operational false admission, and structural hypotheses separately rather than as one PASS/FAIL label. Refresh: treat distribution-shift alarms as requests to invalidate and recompute a reference map, not as fault evidence. We instantiate the protocol in an aggregate-only simulator with 24 policy components, three demand regimes, two fault-mask families, and independent development and heldout seeds. The preregistered heldout contains 1,440 cases and 21,600 partition rows. Only 55/72 regime-component units were reference-admitted and 54/55 remained runtime-admitted, making abstention part of the result. Stable false admission was 0/20 represented components, with a one-sided exact 95% upper bound of 0.1391 under a frozen 0.20 rule. Within admitted units, affected clean traffic outpredicted nominal fault-cell fraction: across 540 unit-arm rows nested in 20 component clusters, the cell-minus-traffic negative-log-likelihood difference was 0.1264 nats per row, with a 95% component-cluster interval of [0.0593, 0.1918]. A drift log shows why "null" must be reference-relative: clean fault-null streams triggered 15/15, 0/15, and 14/15 alarms across three regimes, while only the middle regime matched the frozen detector reference. Rather than a universal threshold, we contribute an executable contract linking observable support, statistical calibration, and justified claims.
Chinese Translation
一项AI评估即使完全可复现,仍可能支持错误的结论。这一风险在闭环系统中尤为突出:策略决定了所访问的状态、可观测的组件,以及哪些故障会留下可测量的痕迹。我们提出了一种包含三种操作的声明安全协议。拒绝(Refuse):当缺乏干净的参考数据流或匹配的运行时比较支持时,选择弃权。分解(Decompose):将协议执行情况、操作性错误接纳以及结构性假设分别报告,而非合并为单一的通过/失败(PASS/FAIL)标签。刷新(Refresh):将分布偏移警报视为重新无效化并重算参考映射的请求,而非故障证据。我们在一个仅聚合的模拟器中实例化该协议,其中包含24个策略组件、三种需求体制、两类故障掩码族,以及独立的开发种子和留出种子。预注册的留出集包含1,440个案例和21,600个分区行。仅有72个体制-组件单元中的55个获得参考接纳,其中54个在运行时仍保持接纳,使弃权本身成为结果的一部分。稳定的错误接纳为所表示的20个组件中的0个,在冻结的0.20规则下单侧精确95%置信上界为0.1391。在被接纳的单元内,受影响的干净流量的预测表现优于名义故障单元比例:在嵌套于20个组件簇中的540个单元-臂行中,单元减去流量的负对数似然差为每行0.1264 nats,其95%组件簇区间为[0.0593, 0.1918]。漂移日志表明为何“零假设”必须相对于参考而言:干净的故障-零流在三种体制下分别触发了15/15、0/15和14/15的警报,而只有中间体制与冻结的检测器参考相匹配。我们并未提出一个通用阈值,而是贡献了一个可执行的契约,将可观测支持、统计校准与有依据的声明联系起来。
cs.AI / 82 / 2609.20543

Language-model groups overstate consensus when replaying human deliberation on a reasoning task

在重演人类推理任务讨论时,语言模型群体会高估共识程度
Shao, Tengfei
Abstract
Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.
Chinese Translation
完全共识率常被视为集体认知的指标,但其结果取决于如何界定参与和最终状态。我们使用匹配的大语言模型(LLM)智能体组重演了100个留出的Wason任务人类小组,为每位参与者讨论前的答案各设定一个信念锚定智能体,并用同一编码标准对智能体和人类进行评分。在不同的人类评分定义下,共识率估计介于24.0%至57.0%之间;约五分之一的参与者从未发言,而智能体几乎总会发言。在去盲后的两项敏感性分析中,智能体组仍表现出更高的共识度:基于提交的比较(n = 98)显示聊天模式和推理模式的差距分别为34.0和43.9个百分点,参与匹配的比较(n = 45)显示差距分别为34.1和44.4个百分点。这两条互补的分析路径减少了不同的测量不对称性,结果却在0.5个百分点内趋同。在取消提前停止以及采用去除可记忆答案的重新参数化后,该差距依然存在;此时推理模式组几乎达成完全一致,且多数情况下一致于错误答案。模拟共识并不能反映集体准确性,在此情境下,信念锚定的智能体组是人类群体结果分布的有偏估计器。这些分析为评估人类审议结果的模拟群体估计提供了评分明确的基础。
cs.AI / 83 / 2609.20581

Limits of Confidence in Diffusion

扩散模型中置信度的局限性
Webb, Russ, Shidani, Amitis, Bizeul, Alice, Busbridge, Dan
Abstract
Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write from those same distributions. For domains of general interest (pixels, phonemes, or words) there are inherent dependencies between tokens. We show that a step matches the training distribution only when the positions it writes are conditionally independent given the tokens already fixed, that no product of per-position distributions can match a dependent group, and that per-position distributions do not determine whether a group is dependent: two joint distributions can have identical per-position marginals while differing in which combinations of values occur. On ScanAndAdd, a synthetic task whose joint distribution is available in closed form, we verify that every group of two or more undetermined positions a confidence ranking writes is dependent, and measure the generated distribution to be $29\times$ the sampling-noise floor total variation while per-sample metrics are $1.0$.
Chinese Translation
离散扩散模型,包括重掩码(remasking)和均匀状态采样器,通过每步写入多个token位置来生成序列:每个位置的token从各自的位置分布中抽取,并依据这些相同的分布选择要写入哪些位置。对于一般感兴趣的领域(像素、音素或词语),token之间存在着固有的依赖关系。我们证明,仅当一步所写入的位置在给定已固定token的条件下相互独立时,该步才与训练分布相匹配;任何按位置分布的乘积都无法匹配具有依赖关系的组;并且按位置分布并不能决定一组位置是否相互依赖——两个联合分布可以拥有完全相同的按位置边缘分布,却在哪些取值组合会出现上有所不同。在ScanAndAdd这一具有闭式解联合分布的合成任务上,我们验证了置信度排序所写入的每一组包含两个及以上未确定位置的位置都是相互依赖的,并测得生成分布的总变差为采样噪声底限的29倍,而逐样本指标则为1.0。
cs.AI / 84 / 2609.20634

PAA: The Probabilistic Allen Algebra: A Generative and Complete Probabilistic Extension of Allen's Interval Relations

PAA:概率艾伦代数——艾伦区间关系的一个生成式且完备的概率扩展
Eggert, Julian
Abstract
Allen's interval algebra is a qualitative calculus for temporal relations, but its thirteen base relations are crisp predicates over exact interval boundaries. This is inadequate for temporal information from language, perception, databases, or uncertain histories, where times, durations, and boundaries are uncertain and expressions such as "just before" or "roughly during" have graded meaning. We develop the probabilistic Allen algebra (PAA): a generative and complete extension in which relation probabilities are derived from distributions over interval boundaries rather than assigned as scores. Time points are Gaussian; intervals have Gaussian midpoints and truncated-Gaussian durations. Every relation is a boundary-ordering predicate in one common probability space: point-point relations reduce to error functions, and point-interval and interval-interval relations to multivariate Gaussian orthant probabilities induced by linear inequalities. Contact relations (meets, starts, finishes, equals) receive positive measure through a tolerance band, and under a single tolerance the thirteen relations form a true partition that recovers crisp Allen as the tolerance vanishes. The construction derives Allen's taxonomy rather than positing it: coarse predicates such as precedence, overlap, and containment are unions of leaves whose probabilities are leaf sums, and this hierarchy is preserved as intervals collapse to points and thirteen relations reduce to five and then three. Each relation further decomposes into correlation-aware temporal primitives in the spirit of CIDOC CRM. The algebra is scale-invariant and separates graded expressions such as "shortly before" from contact relations. All results are Monte-Carlo validated and shipped as an open, tested Python package.
Chinese Translation
艾伦区间代数是一种用于时间关系的定性演算,但其十三种基本关系是基于精确区间边界的清晰谓词。这对于来自语言、感知、数据库或不确定历史记录的时间信息而言是不够的,因为在这些情形下,时间、时长和边界均具有不确定性,且诸如“就在之前”或“大致在……期间”之类的表达具有渐进的语义。我们提出了概率艾伦代数:一个生成式且完备的扩展,其中关系概率由区间边界上的分布推导得出,而非作为分数被直接赋值。时间点服从高斯分布;区间具有高斯分布的中点和截断高斯分布的时长。每种关系都是同一概率空间中的边界排序谓词:点—点关系可归约为误差函数,点—区间和区间—区间关系则可归约为由线性不等式诱导的多元高斯象限概率。接触关系(相接、开始、结束、相等)通过一个容差带获得正测度,并且在单一容差下,这十三种关系构成一个真正的划分,当容差趋近于零时可恢复为清晰艾伦代数。这一构造推导出了艾伦的分类体系而非将其作为公设:诸如先后、重叠和包含等粗粒度谓词是叶关系的并集,其概率为叶概率之和,并且当区间坍缩为点、十三种关系依次约减为五种再到三种时,这一层级结构依然保持。此外,每种关系在CIDOC CRM的精神下进一步分解为考虑相关性的时间基元。该代数具有尺度不变性,并将“不久之前”等渐进表达与接触关系区分开来。所有结果均经过蒙特卡洛验证,并以一个开放的、经过测试的Python软件包形式发布。
cs.AI / 85 / 2609.20658

Ownership in AI-Assisted Everyday Tasks

AI辅助日常任务中的所有权感
Wei, Megan, Subbiah, Melanie, Lee, Audrey, Dahmani, Annya, Edwards, Dave, Edwards, Helen, Pavlick, Ellie
Abstract
When does work done with AI still feel like ours? As AI becomes woven into everyday tasks, we must examine what happens to our sense of ownership and contribution when a machine shares in producing what we make. We report an exploratory qualitative survey in which participants were asked to describe two recent, self-selected tasks completed with AI: one that felt like their own and one that did not. We find that felt ownership depends on the process of collaboration: people disown work when they merely approve AI's suggestions, but retain ownership when they lead, iterate, or rewrite. Ownership can also extend to settings where people own the vision for a project but not the execution; respondents reported high ownership on tasks they could not have completed without AI. Loss of personal voice and a lack of comprehension of the output both erode ownership. Finally, willingness to disclose AI use is often decoupled from actual pride or ownership, and instead shaped by community norms and fear of credit erasure. We propose several research directions as a result of these findings to promote AI development that supports people's sense of authorship over their own lives.
Chinese Translation
借助AI完成的工作何时仍会让人感觉是属于自己的?随着AI日益融入日常任务,我们必须审视当机器参与共同创造我们所做的成果时,我们的所有权感和贡献感会发生什么变化。我们开展了一项探索性定性调查,要求参与者描述最近自选的两个借助AI完成的任务:一个让他们感觉属于自己的任务和一个不属于自己的任务。研究发现,所有权感取决于协作过程:当人们仅仅认可AI的建议时,会否认对工作的归属,而当他们主导、迭代或重写时,则会保留所有权感。所有权感还可以延伸到人们拥有项目愿景但不负责具体执行的情形;受访者报告称,即使在没有AI就无法完成的任务上,他们的所有权感依然很高。个人声音的丧失以及对产出内容缺乏理解都会削弱所有权感。最后,披露AI使用意愿往往与实际的自豪感或所有权感脱钩,而是受到社群规范和对功劳被抹去的担忧的影响。基于这些发现,我们提出了若干研究方向,以促进支持人们对自身生活拥有作者感的AI发展。
cs.AI / 86 / 2609.20722

Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models

Deep Noir:基于Transformer模型架构计时学的自主转向参数发现
Bobe III, Frank E., Vetaw, Gregory D., Bryner, Darshan W., Cook, Matthew G., Salas-Vernis, Jose L.
Abstract
Activation steering modifies LLM behavior at inference time, but identifying where and how strongly to steer remains manual. We introduce Deep Noir, a framework that uses Logit Lens convergence and causal head-level attribution to autonomously discover optimal steering parameters. Across three scales (1B x 3, 2-3B x 2, and 7-9B x 4), our engine achieves 16.7 percentage-point improvement on spam at 1B (standard deviation 4.7; 39 runs), with gains increasing to 21 to 42 percentage points at 7-9B across four architectures. On SST-2 sentiment, it achieves a 13.1 percentage-point improvement with zero code changes. Mechanistic grounding enables automated discovery of intervention points that generalize across tasks and architectures. On sentiment, RepE without head masking fails to improve over baseline, while Deep Noir improves all models (p less than 0.01). We further show that steering creates a predictable prompt-injection attack surface whose vulnerability increases monotonically with steering magnitude. This finding is relevant to agent systems deploying steered classifiers.
Chinese Translation
激活转向(Activation Steering)可在推理阶段修改大语言模型(LLM)的行为,但确定在何处转向以及转向强度的大小仍需人工完成。我们提出了Deep Noir,一个利用Logit Lens收敛性和因果性注意力头级别归因来自主发现最优转向参数的框架。在三个模型规模上(1B×3、2-3B×2 和 7-9B×4),我们的引擎在1B模型的垃圾检测任务上实现了16.7个百分点的提升(标准差4.7;共39次运行),在7-9B规模的四种架构上提升幅度进一步扩大至21至42个百分点。在SST-2情感分类任务上,该方法在零代码修改的情况下实现了13.1个百分点的提升。基于机制层面的解释,能够自动发现可跨任务和架构泛化的干预点。在情感分类任务中,未使用注意力头掩蔽的RepE方法无法超越基线,而Deep Noir改进了所有模型(p小于0.01)。我们进一步证明,转向操作会产生一个可预测的提示注入(prompt-injection)攻击面,其脆弱性随转向强度的增加而单调上升。这一发现对部署转向分类器的智能体系统具有重要意义。
cs.AI / 87 / 2609.20732

Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure

对任意电子表格进行问答需要理解其网格结构
Smoleń, Zofia
Abstract
Semantic cell annotation improves chunking interpretability for spreadsheets in LLM-driven RAG systems, aiding answer generation through enriched context rather than improved retrieval accuracy. We propose a novel framework of splitting any spreadsheet into interpretable chunks using cell role annotation. Our framework beats the state of the art, yet it faces a hard ceiling. Spreadsheets are fundamentally two-dimensional unstructured data with continuous relationships and infinite potential cell roles. Because classification models are restricted to finite, pre-defined classes, they cannot perfectly capture this structural nuance, even with human-level annotation. We show that addressing the spreadsheet-to-LLM bottleneck requires moving beyond discrete cell classification. Instead, the field must develop dimensionality-reduction techniques to directly flatten 2D unstructured spreadsheets into 1D unstructured text. Text chunks would be easier for downstream RAG to interpret and generate from.
Chinese Translation
语义单元格标注提升了LLM驱动的RAG系统中电子表格分块的可解释性,通过丰富上下文来辅助答案生成,而非提升检索准确率。我们提出了一种新颖的框架,利用单元格角色标注将任意电子表格切分为可解释的分块。我们的框架超越了当前最先进的方法,但仍面临一个难以逾越的上限。电子表格本质上是具有连续关系和无限潜在单元格角色的二维非结构化数据。由于分类模型受限于有限的、预定义的类别,即使具备人类水平的标注,也无法完美捕捉这种结构上的细微差异。我们证明,要解决电子表格到LLM的瓶颈,必须超越离散的单元格分类。相反,该领域必须开发降维技术,直接将二维非结构化电子表格展平为一维非结构化文本。这样的文本分块将更易于下游RAG系统进行解释和生成。
cs.AI / 88 / 2609.20754

RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents

RAFT:面向故障排除智能体的有状态检索增强框架
Zhang, Mingxuan, Wang, Xiaowen, Sharan, Anupma, Chen, Zhengyi, Zhang, Chenyu Diana, Yang, Shanshan, Pacharu, Chittibabu
Abstract
Effective troubleshooting agents in enterprise customer support depend on retrieving actionable guidance from similar historical cases, yet existing retrieval-augmented generation (RAG) systems treat support cases as static documents and overlook their multi-stage, stateful nature. We introduce RAFT (Retrieval-Augmented Framework for Troubleshooting Agents), a stateful RAG framework that abstracts each closed historical case into a directed chain of timeline entries and retrieves at the entry level, surfacing cases whose intermediate states match the active case and returning the parent-case trajectory anchored at the matched state; an optional case-level graph links cases through a configurable similarity representation. We evaluate this retrieval layer directly, which, unlike evaluating a full agent system, requires no production deployment. Because public multi-stage troubleshooting data is extremely rare, we pair a synthetic benchmark built from Microsoft Learn Windows Server documentation with real Apache Jira issues carrying human-created duplicate labels. RAFT improves Case Hit over vanilla RAG and GraphRAG baselines at every stage of case progress, with statistically significant gains over the strongest baseline; the Jira results provide directional evidence that the advantage transfers to real case histories. We release our benchmark, implementation, and the Apache Jira evaluation set.
Chinese Translation
企业客户支持中有效的故障排除智能体依赖于从相似的历史案例中检索可操作的指导信息,然而现有的检索增强生成(RAG)系统将支持案例视为静态文档,忽略了其多阶段、有状态的特性。我们提出了RAFT(面向故障排除智能体的检索增强框架,Retrieval-Augmented Framework for Troubleshooting Agents),这是一个有状态的RAG框架,它将每个已关闭的历史案例抽象为时间线条目的有向链,并在条目级别进行检索,从而找出与当前案例中间状态相匹配的案例,并返回以匹配状态为锚点的父案例轨迹;此外,一个可选的案例级图通过可配置的相似性表示将案例相互关联。我们直接对该检索层进行评估,与评估完整的智能体系统不同,这无需生产环境部署。由于公开的多阶段故障排除数据极为稀缺,我们配套构建了一个基于Microsoft Learn Windows Server文档的合成基准,并结合带有由人工标注的重复标签的真实Apache Jira问题。在案例进展的每个阶段,RAFT的Case Hit指标均优于朴素RAG和GraphRAG基线,且相对于最强基线的提升具有统计学显著性;Jira上的结果提供了方向性证据,表明该优势可迁移至真实的案例历史。我们公开了我们的基准、实现以及Apache Jira评估集。
cs.AI / 89 / 2609.20804

An Empirical Study of Harness Design for Coding Agents

关于编码智能体工具框架设计的实证研究
Fan, Run-Ze, Zhang, Zihao, Ma, Simin, Hu, Yebowen, Wang, Shouju, Song, Kaiqiang, Liu, Fei, Zamani, Hamed, Wang, Xiaoyang
Abstract
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management. Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 176 matched settings spanning five context-management strategies, four context-window budgets, and targeted ablations of planning and action space. We find that: (1) Context management becomes increasingly valuable as the context-window budget tightens, with most of its benefit coming from preventing context-overflow failures. (2) Staging rule-based elision before LLM-based summarization provides the strongest overall efficiency among the context-management strategies, whereas making elided content recoverable adds machinery that models rarely use and yields no accuracy gain. (3) Planning shifts from an accuracy scaffold for weaker models to a cost saver for stronger models, with little change in accuracy. (4) Predefined tools improve performance for models with weaker bash proficiency, whereas bash-capable models can operate effectively with a bash-only interface and achieve substantially lower cost, especially on command-line-centric tasks. Trajectory-level analysis explains these effects: context management extends execution trajectories without substantially altering agent behavior, planning changes where trajectories stop, and the action space changes the granularity at which code is written. These findings inform model- and budget-aware harness design and provide a modular framework for evaluating future harness components.
Chinese Translation
编码工具框架(coding harness)决定了自主编码智能体如何将模型能力转化为长时程软件工程性能,然而现有工作通常将工具框架作为整体系统进行评估,导致各独立组件的有效性尚不明确。为实现组件层面的比较,我们基于一个轻量级编码工具框架开展研究,其执行循环保持固定,而对三个组件进行变化:规划(planning)、动作空间(action space)和上下文管理(context management)。我们在 SWE-Bench Verified 和 Terminal-Bench 2.1 上对四个模型进行评估,共涵盖 176 个匹配设置,包括五种上下文管理策略、四种上下文窗口预算以及针对规划和动作空间的定向消融实验。我们发现:(1)随着上下文窗口预算收紧,上下文管理的价值日益凸显,其收益主要来自防止上下文溢出失败。(2)在上下文管理策略中,先执行基于规则的删减、再进行基于 LLM 的摘要,可获得最佳的整体效率;而使被删减内容可恢复所引入的机制却很少被模型使用,且未带来准确率提升。(3)规划的作用从对较弱模型的准确率支撑,转变为对较强模型的成本节省手段,而准确率几乎不变。(4)预定义工具可提升 bash 熟练度较弱的模型的性能,而具备 bash 能力的模型仅凭 bash 接口即可高效运行,并实现显著更低的成本,尤其在以命令行为中心的任务上。轨迹层面的分析解释了这些效应:上下文管理在不显著改变智能体行为的前提下延长了执行轨迹;规划改变了轨迹停止的位置;动作空间则改变了代码编写的粒度。这些发现为面向模型与预算感知的工具框架设计提供了依据,并为评估未来的工具框架组件提供了一个模块化框架。
计算机视觉 (Computer Vision)
101
cs.CV / 1 / 2609.19230

Open ultrasound foundation model for robust segmentation and clinical measurement across heterogeneous settings

面向异质环境下鲁棒分割与临床测量的开放超声基础模型
Qin, Chao, Khan, Fahad Shahbaz, Khan, Salman, Ather, Sarim, Anwar, Siddiq, Anwer, Rao Muhammad, Khan, Shadab
Abstract
Ultrasound is the most widely deployed imaging modality worldwide, yet clinical AI remains fragmented into narrow single-task models that fail when device, operator, or anatomy changes. Here we present SonoCorpus, an open resource unifying 456,963 images and 1,626,085 expert masks from 53 public datasets spanning 24 clinical applications and 17 countries, and SonoBase, an interactive segmentation foundation model pretrained on it. Across fifteen evaluation datasets introducing new organs, devices, operators, and geographies, SonoBase outperforms SAM2, MedSAM2, and the concept-promptable MedSAM3 on every dataset and matches per-dataset specialist models trained on the same data; on fully external data it exceeds the accuracy these baselines achieve on their own in-distribution benchmarks. Ejection fraction derived from its segmentations falls within inter-observer variability (6.63\% error), with fewer misclassifications at the defibrillator-candidacy threshold than either promptable baseline (13\% versus 18--42\%); fetal head-circumference (1.81~mm) and gestational-age (1.2 days) errors fall below inter-observer variability. Where a baseline fails outright, one in four test cases, SonoBase recovers a usable segmentation in 81\% of them, including on handheld probes operated by minimally trained users in two low- and middle-income countries (Sierra Leone and Tanzania). Five labeled examples can help the model adapt to a new setting, and the identical training protocol transfers well to newer models such as SAM3, locating the advantage in ultrasound-specific pretraining rather than any single architecture. To ensure reproducibility and enable the community to build on SonoBase as a platform, we release all checkpoints, optimizer states, data-split indices, deduplication hashes, and starter code.
Chinese Translation
超声是全球部署最广泛的成像方式,然而临床人工智能仍局限于狭窄的单任务模型,当设备、操作者或解剖部位发生变化时便会失效。本文提出SonoCorpus,一个整合了来自53个公开数据集、覆盖24个临床应用和17个国家、共456,963张图像和1,626,085个专家标注掩码的开放资源;以及SonoBase,一个在其上进行预训练的交互式分割基础模型。在引入新器官、新设备、新操作者和新地域的十五个评估数据集上,SonoBase在所有数据集上均优于SAM2、MedSAM2以及概念可提示的MedSAM3,并与在同一数据上训练的各数据集专家模型相当;在完全外部数据上,其精度甚至超过了这些基线模型在各自分布内基准上所达到的水平。由其分割结果推导的射血分数误差处于观察者间变异范围内(6.63%),在除颤器适用阈值处的误分类少于任一可提示基线模型(13%对18–42%);胎儿头围(1.81毫米)和孕周(1.2天)误差亦低于观察者间变异。当基线模型完全失败时(约占四分之一的测试案例),SonoBase在其中81%的情况下仍能恢复出可用的分割结果,包括在两个中低收入国家(塞拉利昂和坦桑尼亚)由仅受极少培训的用户操作的手持探头上的应用。仅需五个标注样本即可帮助模型适应新环境,且相同的训练协议可良好迁移至SAM3等更新的模型,这表明其优势在于超声特异性预训练而非某一特定架构。为确保可复现性并使社区能够将SonoBase作为平台进行构建,我们发布了所有检查点、优化器状态、数据划分索引、去重哈希和启动代码。
cs.CV / 2 / 2609.19236

RAUL: Reference-Assisted Ureteroscopy Localization for Skill Assessment

RAUL:基于参考辅助的输尿管镜定位技能评估方法
Li, Fangjie, Bui, Mai, Mohan, Charan, Miga, Michael, Chabanas, Matthieu, Kavoussi, Nicholas, Wu, Jie Ying
Abstract
Objective: Incomplete navigation of anatomy during ureteroscopic kidney stone surgeries can contribute to repeat interventions. While skilled surgeons have lower reintervention rates, there are no objective metrics to quantify scope-navigation performance to evaluate when a trainee becomes skilled. This work aims to recover ureteroscope trajectories from endoscopic video and derive navigation metrics to quantify differences in skill. Methods: We propose RAUL, a reference-assisted reconstruction framework for recovering ureteroscope trajectories from ureteroscope videos only in phantoms. For each phantom, we use a slow, high-quality reference exploration video to generate a reference reconstruction. We localize subsequent exploration videos against this reference. We evaluate localization accuracy against electromagnetically tracked scope pose. We compute navigation metrics from phantom exploration trajectories to compare surgical residents across experience levels. Results: The proposed reference-assisted framework achieves a mean translation root mean square error of $0.5 \pm 0.1$ mm across 9 phantoms. Compared to standard Structure-from-Motion (SfM), the proposed pipeline increases frame-wise localization coverage from $50.5 \pm 14.9\%$ to $86.1 \pm 7.2\%$ of all video frames. The reconstructed trajectories revealed significant differences between high- and low-experience trainees in established navigation metrics. Conclusion: RAUL enables substantially more complete recovery of ureteroscope trajectories from videos compared to standard SfM pipelines, enabling trajectory-based skill assessment without additional tracking equipment. Significance: To the best of our knowledge, this is the first use of video-only recovery of ureteroscope trajectories without external tracking sensors for skill assessment, supporting scalable automated assessment of ureteroscopy navigation skill.
Chinese Translation
目的:在输尿管镜肾结石手术中,解剖结构导航的不完整可能导致重复干预。虽然熟练外科医生的再干预率较低,但目前尚缺乏客观指标来量化镜体导航表现,以评估受训者何时达到熟练水平。本工作旨在从内窥镜视频中恢复输尿管镜的运动轨迹,并推导导航指标以量化技能差异。方法:我们提出RAUL,一种仅在体模中仅利用输尿管镜视频恢复输尿管镜轨迹的参考辅助重建框架。对于每个体模,我们使用一段缓慢、高质量的参考探查视频生成参考重建,然后将后续的探查视频相对于该参考进行定位。我们以电磁跟踪的镜体位姿为基准评估定位精度,并从体模探查轨迹中计算导航指标,以比较不同经验水平的住院医师。结果:所提出的参考辅助框架在9个体模上实现了平均0.5±0.1毫米的平移均方根误差。与标准的运动恢复结构(Structure-from-Motion, SfM)相比,所提出的流程将逐帧定位覆盖率从所有视频帧的50.5±14.9%提高到86.1±7.2%。重建的轨迹在既有的导航指标上显示出高经验与低经验受训者之间的显著差异。结论:与标准SfM流程相比,RAUL能够从视频中更完整地恢复输尿管镜轨迹,从而无需额外跟踪设备即可实现基于轨迹的技能评估。意义:据我们所知,这是首次在不使用外部跟踪传感器的情况下,仅通过视频恢复输尿管镜轨迹用于技能评估,支持输尿管镜导航技能的可扩展自动化评估。
cs.CV / 3 / 2609.19354

Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment

视觉-语言模型能评判奥运会跳水吗:零样本动作质量评估中从推理到打分
Velesaca, Henry O., Freire-Obregon, David, Miranda, Luigi, Reyes-Angulo, Abel
Abstract
Automated action quality assessment (AQA) in Olympic sports remains a challenging task due to the complexity of human motion and the subjectivity inherent in expert judging. This work evaluates the capability of open-source Vision-Language Models (VLMs) to perform zero-shot action quality assessment on Olympic diving videos using the AQA-7 benchmark dataset. In this regard, a regression-based framework is pro-posed to leverage both the semantic reasoning and phase-level sub-scores generated by the VLMs, combining TF-IDF vectorization, dimensionality reduction, and ensemble learning to predict final competition scores. Experimental results show that standalone VLMs achieve moderate Spearman correlations below 0.32, while the proposed ensemble regression framework substantially improves performance in the reported evaluation, reaching a Spearman correlation of 0.67 with a four-model configuration. Textual reasoning features con-sistently outperformed raw numerical sub-scores, highlighting the richness of VLM-generated explanations for action quality analysis. These findings suggest that VLMs hold strong potential as assistive tools for explainable and semi-automated sports performance evaluation. The code is publicly available on GitHub https://github.com/hvelesaca/olympic diving judge vlm
Chinese Translation
奥运会项目中自动化动作质量评估(AQA)仍然是一项具有挑战性的任务,其原因在于人体运动的复杂性以及专家评分中固有的主观性。本研究利用AQA-7基准数据集,评估开源视觉-语言模型(VLM)在奥运会跳水视频上进行零样本动作质量评估的能力。为此,我们提出了一种基于回归的框架,利用VLM生成的语义推理和阶段级子评分,结合TF-IDF向量化、降维和集成学习来预测最终比赛得分。实验结果表明,单独使用VLM仅能达到低于0.32的中等Spearman相关系数,而所提出的集成回归框架在报告的评估中显著提升了性能,在四模型配置下达到0.67的Spearman相关系数。文本推理特征始终优于原始数值子评分,凸显了VLM生成的解释在动作质量分析中的丰富性。这些发现表明,VLM作为可解释和半自动化体育表现评估的辅助工具具有巨大潜力。代码已在GitHub上公开:https://github.com/hvelesaca/olympic diving judge vlm
cs.CV / 4 / 2609.19358

Open-vocabulary 3D object detection with promptable segmentation

基于可提示分割的开放词汇三维目标检测
Deniz, Ömer Faruk, Koçyiğit, Mustafa Taha
Abstract
Three-dimensional object detection for autonomous driving is dominated by detectors trained on large corpora of human-annotated 3D boxes. Such a detector learns a fixed category list, and everything outside it is invisible. This paper asks whether the task can be solved training-free and open-vocabulary. A promptable segmentation model (SAM3), queried with class names as text prompts, supplies instance masks in the vehicle's six surround-view cameras, and the masks are turned into metric 3D boxes using the geometry of the scene. The core is a controlled three-stage comparison on nuScenes in which 2D detection is held fixed and only the source of 3D geometry changes. Geometry predicted from images alone reaches 0.183 mean average precision (mAP) under the official protocol; fitting boxes from raw LiDAR points inside the same masks with training-free rules reaches 0.298 mAP / 0.348 nuScenes detection score (NDS) at zero labeling cost; borrowing supervised box geometry at inference time lifts the same detections to 0.413 mAP / 0.555 NDS, which locates the pipeline's largest deficit in measurement precision rather than 2D detection, while class confusion and confidence calibration survive that substitution. Reversing the direction, a three-state camera-witness rule built from the same masks improves a supervised LiDAR-only detector from 0.596 to 0.630 mAP, roughly half the gain of fully supervised camera fusion, with no training. A coverage analysis shows that SAM3 finds 84% of in-range objects with a correctly named mask; the classes that fail in the official metric are misnamed or geometrically unforgiving, not unseen.
Chinese Translation
自动驾驶的三维目标检测目前主要依赖在大规模人工标注三维框数据上训练的检测器。这类检测器只能学习固定的类别列表,列表之外的一切目标均不可见。本文探讨该任务能否以免训练(training-free)和开放词汇的方式解决。我们使用可提示分割模型 SAM3,以类别名称作为文本提示进行查询,在车辆环视的六个相机中获取实例掩码,并利用场景几何将掩码转换为度量三维框。核心是在 nuScenes 数据集上进行的一项受控三阶段对比实验:保持二维检测结果不变,仅改变三维几何的来源。仅凭图像预测的几何在官方评测协议下达到 0.183 的平均精度均值(mAP);在相同掩码内使用免训练规则拟合原始 LiDAR 点得到三维框,在零标注成本下达到 0.298 mAP / 0.348 nuScenes 检测分数(NDS);在推理时借用有监督的框几何可将同样的检测提升至 0.413 mAP / 0.555 NDS。这表明该流程的最大短板在于测量精度而非二维检测,而类别混淆与置信度校准问题在替换几何来源后依然存在。反方向实验中,基于相同掩码构建的三态相机见证规则(three-state camera-witness rule)在无训练的情况下,将有监督的纯 LiDAR 检测器的 mAP 从 0.596 提升至 0.630,约为完全有监督相机融合所带来增益的一半。覆盖分析显示,SAM3 能够以正确命名的掩码找到 84% 范围内的目标;在官方指标中失败的类别属于命名错误或几何上难以处理的目标,而非未见类别。
cs.CV / 5 / 2609.19377

LinePilot Digitizer: Line-Plot Recovery with Manual and Automatic Calibration

LinePilot 数字化工具:基于手动与自动校准的线图恢复
Ma, Fengbo, Akhtar, Rayan, Joshi, Aakash H., Li, Xiaoting, Sun, Haijian, Xiang, Zhen, Chen, Xianyan, Zhao, Yiping
Abstract
Recovering numerical series from line plots requires accurate axis calibration and reliable curve extraction. We present LinePilot Digitizer (LinePilot), which combines continuous color-based curve recovery with three calibration modes: LinePilot (standard), LinePilot (enhanced), and LinePilot (OCR). We also introduce DigitizerBench, the first dedicated benchmark for systematically evaluating digitizer performance, using an orthogonal design spanning signal, rendering, and plot-structure factors with complementary automatic and human-guided evaluations. We evaluate performance using failure-penalized capped normalized root-mean-square error (FPC-NRMSE), which assigns unit loss to missing, unusable, or catastrophically inaccurate outputs. On DigitizerBench-Full, LinePilot (OCR) achieves the lowest mean FPC-NRMSE (0.672) and highest trusted usability (38.2%) among the tested automatic pipelines. On DigitizerBench-Lite, LinePilot (enhanced) achieves the lowest mean FPC-NRMSE (0.081), 100% output success, and highest trusted usability (93.3%). The orthogonal benchmark design further enables factor analysis to identify the factors that most significantly affect digitizer performance. Together, the three calibration modes provide a practical trade-off between automation, user control, and accuracy within a shared curve-recovery workflow.
Chinese Translation
从线图中恢复数值序列需要精确的坐标轴校准和可靠的曲线提取。我们提出了 LinePilot 数字化工具(LinePilot),它将基于颜色的连续曲线恢复方法与三种校准模式相结合:LinePilot(标准)、LinePilot(增强)和 LinePilot(OCR)。我们还介绍了 DigitizerBench,这是首个用于系统性评估数字化工具性能的专用基准,采用涵盖信号、渲染和图形结构因素的正交设计,并辅以自动化与人工引导相结合的评估方式。我们使用带失败惩罚的截断归一化均方根误差(FPC-NRMSE)来评估性能,该指标对缺失、不可用或严重不准确的输出赋予单位损失。在 DigitizerBench-Full 上,LinePilot(OCR)在所测试的自动化流程中取得了最低的平均 FPC-NRMSE(0.672)和最高的可信可用性(38.2%)。在 DigitizerBench-Lite 上,LinePilot(增强)取得了最低的平均 FPC-NRMSE(0.081)、100% 的输出成功率以及最高的可信可用性(93.3%)。正交基准设计还支持因素分析,以识别对数字化工具性能影响最显著的因素。总体而言,这三种校准模式在共享的曲线恢复流程中,为自动化程度、用户控制与准确性之间提供了实用的权衡。
cs.CV / 6 / 2609.19384

Riemannian--Lorentz Fusion of Vision Transformers and State-Space Models

视觉Transformer与状态空间模型的黎曼—洛伦兹融合
Patro, Badri N., Agneeswaran, Vijay S.
Abstract
Scaling deep learning faces critical bottlenecks: data exhaustion, exponential training costs, and resource concentration. Model merging combines pre-trained checkpoints without gradient descent, offering orders-of-magnitude savings versus retraining. Combining independently trained vision models is difficult when their architectures and parameter shapes differ. Existing weight-space merging methods generally assume aligned, shape-compatible checkpoints, whereas a Vision Transformer (ViT) and a state-space model (SSM) implement token mixing with different operators. We study a hybrid Heterogeneous merging setting that retains both architectures while aligning parameter groups by semantic role. Our proposed Riemannian--Lorentz Parameter Fusion (RLPF) method projects aligned groups to common coordinates, lifts selected coordinates to the Lorentz hyperboloid model of hyperbolic space, computes a regularized geodesic barycenter, and decodes the result into the two branches. A learned gate then combines branch logits for each input. Component groups use fixed curvature values, with normalization parameters treated as Euclidean. In the results available in this manuscript, the fine-tuned system obtains 82.37\% on CIFAR-10, 75.04\% on Oxford-IIIT Pet, and 78.58\% top-1 accuracy on ImageNet-1K; the corresponding best-parent accuracies are 76.54\%, 71.42\%, and 76.42\%. On ImageNet-1K, the reported pre-fine-tuning initialization reaches 77.80\%. These results support further study of geometry-aware heterogeneous fusion, but not a training-free single-checkpoint merge: RLPF is a two-branch hybrid whose gate and reported final models are trained.
Chinese Translation
深度学习的扩展面临关键瓶颈:数据枯竭、指数级增长的训练成本以及资源集中。模型合并无需梯度下降即可结合预训练检查点,相比重新训练可节省数个数量级的开销。然而,当独立训练的视觉模型具有不同架构和参数形状时,合并它们十分困难。现有的权重空间合并方法通常假设检查点是对齐且形状兼容的,而视觉Transformer(Vision Transformer, ViT)与状态空间模型(State-Space Model, SSM)采用不同的算子实现词元混合。我们研究一种混合异构合并设定,即同时保留两种架构,并按语义角色对参数组进行对齐。我们提出的黎曼—洛伦兹参数融合(Riemannian–Lorentz Parameter Fusion, RLPF)方法将已对齐的参数组投影到公共坐标系,将选定的坐标提升至双曲空间的洛伦兹双曲面模型,计算正则化的测地重心,并将结果解码回两个分支。随后通过一个可学习的门控机制针对每个输入组合两个分支的输出逻辑值(logits)。各组件组使用固定的曲率值,归一化参数则按欧几里得空间处理。在本文获得的结果中,微调后的系统在CIFAR-10上取得82.37%,在Oxford-IIIT Pet上取得75.04%,在ImageNet-1K上取得78.58%的top-1准确率;相应的最优父模型准确率分别为76.54%、71.42%和76.42%。在ImageNet-1K上,所报告的微调前初始化达到77.80%。这些结果支持对几何感知异构融合的进一步研究,但并不支持免训练的单检查点合并:RLPF是一个双分支混合模型,其门控及所报告的最终模型均经过训练。
cs.CV / 7 / 2609.19393

WZPlanner: Safe End-to-End Path Planning for Autonomous Driving in Work Zones

WZPlanner:面向作业区自动驾驶的安全端到端路径规划
Sahu, Nishad, Qian, Changzhong, Cai, Guangzhou, Sural, Shounak, Ragunathan, Rajkumar
Abstract
Work zones alter lane geometry through temporary traffic controls and closures that may be absent from on-board maps, challenging autonomous vehicle (AV) perception and planning. Generalization is also limited by scarce public datasets with structured geometric supervision. We present WorkZonePlan, a dataset comprising 149K+ synthetic and 5K+ real-world multimodal samples with 3D annotations for lane boundaries, work zone boundaries, and driving trajectory options. It also provides 76 closed-loop CARLA scenarios replayed under three weather conditions, yielding 228 Bench2Drive-format evaluation routes. We introduce WAVE (Work-zone-focused AV data generation in Virtual and rEal Environments), a semi-automated pipeline for creating the dataset, and BoundaryFormer (BF), a transformer-based model that jointly predicts lane and work zone boundary polynomials and driving trajectories. BF uses slot attention for boundary prediction. Ablations show that a separate trajectory decoder using boundary slot features substantially improves trajectory prediction over a slot-attention-only approach. Building on this finding, BF++ offers Camera and Camera+LiDAR variants with metric ground-plane encoding, typed boundary/trajectory queries, long-range point anchors, image-space curve refinement, and conservative gated LiDAR fusion. On the 211 routes common to all four models at the evaluation freeze, BF++-Camera and BF++-Camera+LiDAR achieve Driving Scores of 63.0 and 64.4, respectively, compared with 59.3 for SimLingo and 26.1 for TransFuser++ (TF++). BF++ is 40 times smaller than SimLingo and more than 10 times smaller than TF++, while achieving higher Driving Scores. These results support jointly predicting lane boundaries, work zone boundaries, and driving trajectories as a promising direction toward safer AV operation in work zones. Code and dataset: https://github.com/Nishad-Sahu/WZPlanner.
Chinese Translation
作业区(Work Zones)通过临时交通管制和道路封闭改变车道几何结构,而这些变化可能未包含在车载地图中,给自动驾驶车辆(Autonomous Vehicle, AV)的感知与规划带来挑战。同时,缺乏带有结构化几何监督的公开数据集也限制了模型的泛化能力。我们提出了 WorkZonePlan 数据集,包含超过 14.9 万条合成样本和超过 5 千条真实世界多模态样本,并附带针对车道边界、作业区边界和驾驶轨迹选项的三维标注。该数据集还提供了 76 个闭环 CARLA 场景,并在三种天气条件下进行回放,生成 228 条 Bench2Drive 格式的评估路线。我们提出了 WAVE(Work-zone-focused AV data generation in Virtual and rEal Environments),一种用于构建该数据集的半自动化流水线,并提出了 BoundaryFormer(BF)——一个基于 Transformer 的模型,可联合预测车道与作业区边界多项式以及驾驶轨迹。BF 使用槽注意力机制(slot attention)进行边界预测。消融实验表明,利用边界槽特征(boundary slot features)的独立轨迹解码器相比仅使用槽注意力的方法能显著提升轨迹预测性能。基于这一发现,BF++ 提供了 Camera 和 Camera+LiDAR 两种变体,具有度量地面编码、类型化边界/轨迹查询、远程点锚点、图像空间曲线细化以及保守门控 LiDAR 融合。在评估冻结时所有四个模型共有的 211 条路线上,BF++-Camera 和 BF++-Camera+LiDAR 分别取得 63.0 和 64.4 的驾驶分数(Driving Score),而 SimLingo 为 59.3,TransFuser++(TF++)为 26.1。BF++ 的模型规模比 SimLingo 小 40 倍,比 TF++ 小 10 倍以上,同时取得了更高的驾驶分数。这些结果表明,联合预测车道边界、作业区边界和驾驶轨迹,是实现作业区内更安全自动驾驶运营的一个有前景的方向。代码与数据集:https://github.com/Nishad-Sahu/WZPlanner。
cs.CV / 8 / 2609.19421

RGS: Reflection-aware Gaussian Splatting via Learning Geometry Continuity for Reflective Objects

RGS:基于学习几何连续性的反射感知高斯泼溅方法,用于反光物体
Du, Xiaobiao, Wang, Yida, Bi, Cheng, Zhan, Kun, Yu, Xin
Abstract
Gaussian Splatting has significantly improved the quality of novel view synthesis with explicit Gaussian representation. However, we observed that existing 3D Gaussian Splatting methods (3DGS) often suffer from surface collapse issues on reflective regions, and thus produce inferior geometry and low-quality specular. In this work, we propose a physically-based deferred rendering framework, named Reflection-aware Gaussian Splatting (RGS), that can accurately model specular regions and improve novel view synthesis performance. Specifically, we found that a powerful 3D foundation model can provide a strong 3D geometric prior to foster correct geometric modeling. Based on this, we propose a cross-view shape consistency regularization to regularize the geometry surface with the large model prior and cross-view constraints. In this manner, our RGS can produce smoother geometric surfaces on reflective regions while reducing geometric hollows. To further improve rendering results on reflective regions, we present a reflection-aware densification strategy that is designed to capture specular variations across various views. With this strategy, our RGS is able to render novel views of objects in higher quality. Extensive experiments demonstrate our method consistently renders high-quality reflective objects, achieving state-of-the-art performance.
Chinese Translation
Gaussian Splatting(高斯泼溅)通过显式高斯表示显著提升了新视角合成的质量。然而,我们观察到现有的3D Gaussian Splatting方法(3DGS)在反光区域常常出现表面塌陷问题,从而导致较差的几何效果和低质量的高光渲染。在本工作中,我们提出了一种基于物理的延迟渲染框架,命名为反射感知高斯泼溅(Reflection-aware Gaussian Splatting, RGS),能够准确建模镜面反射区域并提升新视角合成性能。具体而言,我们发现强大的3D基础模型可以提供强有力的3D几何先验,以促进正确的几何建模。基于此,我们提出了一种跨视角形状一致性正则化方法,利用大模型先验和跨视角约束来正则化几何表面。通过这种方式,我们的RGS能够在反光区域生成更平滑的几何表面,同时减少几何空洞。为了进一步提升反光区域的渲染效果,我们提出了一种反射感知的致密化策略,旨在捕捉不同视角间的高光变化。借助该策略,我们的RGS能够以更高质量渲染物体的新视角。大量实验表明,我们的方法能够持续渲染高质量的反光物体,达到了最先进的性能。
cs.CV / 9 / 2609.19444

Seeing Abnormal from Normal: Glomerular Abnormality in Representations of Normal Renal Morphology

从正常中识别异常:正常肾脏形态学表征中的肾小球异常
Hasko, Greta, Saluja, Rachit, Shi, Tianyu, Zhao, Leiyue, Yang, Yuechen, Reisenbuechler, Daniel, Yao, Tianyuan, Guo, Zhenhao, Cannon, John, Yang, Haichun, Huo, Yuankai, Chi, Yuling, Gudas, Lorraine, Sabuncu, Mert R., Yang, Yihe, Deng, Ruining
Abstract
Fine-grained evaluation of glomerular pathology must distinguish normal glomeruli from abnormalities such as global and segmental glomerulosclerosis, obsolescent, ischemic, solidified, disappearing, and atubular glomeruli. Supervised classification requires labeled examples of every category, which is impractical when subtypes are rare or absent from the training cohort. One-class anomaly detection offers an alternative by modeling normal data and scoring deviations, allowing previously unseen abnormalities to be detected. We use the frozen residual U-Net backbone of Omni-Seg, pretrained to segment structurally normal renal primitives without abnormal-subtype labels. We propose NoRDeC (Normal-Reference Detection and Characterization), a framework combining Mahalanobis normal-reference scoring with layer-wise representation analysis to determine whether and where glomerular pathology is encoded, how spatial aggregation affects detection, and whether abnormalities alter inter-layer relationships differently. Using glomerular images from two institutions, we evaluate backbone layers and aggregation strategies, compare NoRDeC with PaDiM and PatchCore, and analyze representations using centered kernel alignment (CKA). Layer 4 with Center-70 aggregation achieved a pooled AUROC of $0.926\pm0.013$. NoRDeC achieved the highest AUROC in six of seven abnormality categories and in the pooled analysis, while CKA suggested subtype-dependent changes in inter-layer relationships not captured by anomaly scores alone. The normal-reference model is fitted using only normal glomeruli; abnormality labels are used for configuration selection, evaluation, and grouping in the representation analysis. These results show that a frozen renal feature extractor can support both detection and representation-level characterization of glomerular abnormalities without using abnormal examples to fit the detector.
Chinese Translation
肾小球病理的细粒度评估必须将正常肾小球与球性及节段性肾小球硬化、废弃性、缺血性、实性、消失性和无小管性肾小球等异常区分开来。有监督分类需要每个类别的标注样本,当某些亚型在训练队列中罕见或缺失时,这一要求难以实现。单类异常检测通过建模正常数据并对偏离程度进行评分,提供了一种替代方案,从而能够检测出先前未见的异常。我们使用 Omni-Seg 的冻结残差 U-Net 骨干网络,该网络经过预训练以分割结构正常的肾脏基本结构,无需异常亚型标签。我们提出了 NoRDeC(Normal-Reference Detection and Characterization,正常参考检测与表征)框架,该框架将马氏距离正常参考评分与逐层表征分析相结合,以确定肾小球病理是否以及在哪里被编码、空间聚合如何影响检测结果,以及异常是否以不同方式改变层间关系。使用来自两家机构的肾小球图像,我们评估了骨干网络的各层及聚合策略,将 NoRDeC 与 PaDiM 和 PatchCore 进行了比较,并使用中心核对齐(CKA)分析表征。第4层结合 Center-70 聚合策略取得了 $0.926\pm0.013$ 的合并 AUROC。NoRDeC 在七个异常类别中的六个以及合并分析中取得了最高的 AUROC,同时 CKA 表明层间关系存在依赖于亚型的变化,这是仅凭异常评分无法捕捉到的。正常参考模型仅使用正常肾小球进行拟合;异常标签仅用于配置选择、评估以及表征分析中的分组。这些结果表明,冻结的肾脏特征提取器能够在不使用异常样本拟合检测器的情况下,同时支持肾小球异常的检测和表征层面的刻画。
cs.CV / 10 / 2609.19451

Efficient Unified Multimodal Understanding (EUMU): Winning Solution for the MUMU Track at the 8th LSVOS Challenge

高效统一多模态理解(EUMU):第八届LSVOS挑战赛MUMU赛道的冠军方案
Kil, Dayoung, Kim, Seong-heum
Abstract
The Mobile Unified Multimodal Understanding (MUMU) Challenge requires a single efficient model to jointly perform multi-concept image tagging, open-vocabulary object detection, and image captioning. We present Efficient Unified Multimodal Understanding (EUMU), the winning solution for the MUMU Track of the 8th LSVOS Challenge. EUMU builds on a shared pretrained multimodal model, using its prompt-based capabilities for detection and captioning and training lightweight heads on shared visual features to predict quality, scene, and event tags. Rather than treating the three tasks independently, EUMU applies task-aware inference refinement by reusing task outputs as cross-task cues. For detection, caption cues help recover objects missed by the initial detection. For captioning, detection cues help refine the caption to better reflect the detected objects. For tagging, image statistics refine quality predictions, while caption and detection cues refine scene and event predictions. This design unifies all three tasks within a single model while satisfying the challenge's resource constraints. EUMU contains 239.169M parameters, requires 23.947 GFLOPs, uses 4.5 GB of peak inference memory, and achieves a final challenge score of 17.3409. Code and models are available at https://github.com/Dayoung-Kil/EUMU.
Chinese Translation
移动端统一多模态理解(MUMU)挑战赛要求单一高效模型联合执行多概念图像标注、开放词汇目标检测和图像描述生成三项任务。我们提出高效统一多模态理解(Efficient Unified Multimodal Understanding,EUMU),即第八届LSVOS挑战赛MUMU赛道的冠军方案。EUMU基于一个共享的预训练多模态模型构建,利用其基于提示(prompt)的能力执行检测和描述任务,并在共享视觉特征上训练轻量级预测头来预测质量、场景和事件标签。EUMU并未将三项任务独立处理,而是通过将任务输出复用为跨任务线索来实现任务感知的推理精化。对于检测任务,描述线索有助于找回初始检测遗漏的目标;对于描述任务,检测线索有助于精化描述内容,使其更好地反映检测到的目标;对于标注任务,图像统计信息用于精化质量预测,而描述和检测线索用于精化场景和事件预测。该设计在满足挑战赛资源限制的同时,将三项任务统一于单一模型之中。EUMU包含2.39169亿(239.169M)参数,计算量为23.947 GFLOPs,推理峰值内存为4.5 GB,最终挑战赛得分为17.3409。代码和模型已发布于 https://github.com/Dayoung-Kil/EUMU。
cs.CV / 11 / 2609.19463

ParticleSplat: Self-supervised Object-centric Latent Particle Splatting

ParticleSplat:自监督对象中心潜在粒子溅射
He, Lyuxing, Guo, Daniel, Terveen, Elizabeth, Pathak, Deepak, Held, David, Daniel, Tal
Abstract
We present ParticleSplat, a self-supervised object-centric learning method that decomposes scenes into a set of latent ''particles'' representing semantic entities through feedforward 3D Gaussian Splatting. Building on the Deep Latent Particles (DLP) framework, which represents images as a set of particles with attributes such as position, scale, and visual appearance, we address a key limitation of DLP: its inherently 2D nature, which prevents explicit 3D spatial and geometric reasoning that are critical for downstream tasks such as robotic manipulation. Leveraging the structural similarity between latent particles and 3D Gaussian primitives, we introduce a 3D latent particle space trained with a novel view synthesis objective. Our model jointly encodes multiple views with camera poses into a shared 3D object-centric latent space, then transforms particles into particle-aligned 3D Gaussians whose composition reconstructs the full scene. On simulated and real-world datasets, we show that this formulation inherently learns object masks without supervision and supports controllable 3D scene editing, such as moving objects by modifying particles in the latent space. We further establish that the learned 3D representation improves downstream performance on robotic manipulation tasks.
Chinese Translation
我们提出了ParticleSplat,一种自监督的对象中心学习方法,通过前馈3D高斯溅射(3D Gaussian Splatting)将场景分解为一组表示语义实体的潜在“粒子”。Deep Latent Particles(DLP)框架将图像表示为一组具有位置、尺度和视觉外观等属性的粒子。在此基础上,我们解决了DLP的一个关键局限:其固有的2D特性阻碍了对机器人操作等下游任务至关重要的显式3D空间与几何推理。利用潜在粒子与3D高斯基元之间的结构相似性,我们引入了一个以新视角合成目标进行训练的3D潜在粒子空间。我们的模型将带相机位姿的多个视角共同编码到一个共享的3D对象中心潜在空间中,然后将粒子变换为与粒子对齐的3D高斯,其组合可重建完整场景。在仿真和真实世界数据集上,我们证明该公式能够在无监督的情况下内在地学习对象掩码,并支持可控的3D场景编辑,例如通过修改潜在空间中的粒子来移动物体。我们进一步证明,所学习到的3D表示能够提升机器人操作任务的下游性能。
cs.CV / 12 / 2609.19483

SCOUT: Sim-to-Real Text-Based Person Retrieval by Embedding-Space Prediction over Frozen Video Features

SCOUT:基于冻结视频特征的嵌入空间预测实现模拟到真实场景的文本行人检索
Traoré, Abdarahmane, Couturier, Andy, Hervet, Éric
Abstract
Text-based person retrieval under a sim-to-real gap (synthetic training data, a real-image gallery) is usually tackled with costly fine-tuned cross-encoders. We ask whether a frozen-encoder system can compete. We present SCOUT, which casts cross-modal retrieval as prediction in embedding space. A trainable predictor maps the patch tokens of a frozen video encoder into the embedding space of a frozen text encoder under a bidirectional InfoNCE objective, and no encoder is fine-tuned in the base model. The video encoder is V-JEPA, the text encoder is EmbeddingGemma, and the predictor is initialized from a Qwen3.5-0.8B decoder. We make three findings. First, the best frozen text encoder is simply the one whose geometry best matches the video features. A training-free alignment score ranks three candidate text encoders in the same order as their retrieval accuracy on our held-out split (Spearman $\rho = 1.0$); a fourth, LLM-based encoder shows the rule is metric-dependent, holding for a neighborhood-overlap score ($\rho = 0.8$) but not for a linear probe ($\rho = -0.2$). Second, two precision-targeted levers, parameter-efficient ExPLoRA adaptation of the video encoder and a training-free attribute-decomposed reranker built on a vision-language model, improve the top-rank precision that otherwise limits the frozen system, adding 2.2 points of leaderboard R@1. Third, a local-versus-public calibration study explains which interventions transfer to the real domain. On AI City Challenge 2026 Track 4 the full retrieve-fuse-rerank system reaches 84.25 mAP@10 on the final leaderboard, while a single frozen model submitted alone reaches 60.63. Our trained components cost about 95 GPU-hours. CMP, the dataset authors' fine-tuned cross-encoder that trains for sixteen GPU-days, is one fusion member of the full system, not an alternative. Code and annotations: https://github.com/abtraore/SCOUT-ECCV
Chinese Translation
在模拟到真实差距(合成训练数据、真实图像检索库)下的文本行人检索通常依赖代价高昂的微调交叉编码器。我们探讨冻结编码器系统能否与之竞争。我们提出 SCOUT,将跨模态检索转化为嵌入空间中的预测。一个可训练的预测器在双向 InfoNCE 目标下,将冻结视频编码器的 patch token 映射到冻结文本编码器的嵌入空间中,基础模型不对任何编码器进行微调。视频编码器为 V-JEPA,文本编码器为 EmbeddingGemma,预测器由 Qwen3.5-0.8B 解码器初始化。我们有三项发现。第一,最佳的冻结文本编码器就是其几何结构与视频特征最匹配的那一个。一种免训练的对齐分数对三个候选文本编码器的排序与其在我们留出划分上的检索精度排序一致(Spearman ρ = 1.0);第四个基于 LLM 的编码器表明该规律依赖于度量方式:邻域重叠分数成立(ρ = 0.8),而线性探针不成立(ρ = -0.2)。第二,两个针对精度的手段——对视频编码器进行参数高效的 ExPLoRA 适配,以及构建在视觉-语言模型上的免训练属性分解重排器——提升了原本限制冻结系统的前列排序精度,在排行榜 R@1 上增加 2.2 个百分点。第三,一项本地与公开排行榜的校准研究解释了哪些干预措施能迁移到真实领域。在 AI City Challenge 2026 Track 4 上,完整的“检索-融合-重排”系统在最终排行榜上达到 84.25 mAP@10,而单独提交的单一冻结模型达到 60.63。我们训练的组件仅耗费约 95 GPU 小时。CMP(数据集作者微调、训练耗时十六个 GPU 天的交叉编码器)是完整系统中的一个融合成员,而非替代方案。代码与标注:https://github.com/abtraore/SCOUT-ECCV
cs.CV / 13 / 2609.19518

AMB3R-SLAM: Kilometer-scale SLAM with Hierarchical Backend

AMB3R-SLAM:基于分层后端的公里级SLAM系统
Wang, Hengyi, Agapito, Lourdes
Abstract
We present AMB3R-SLAM, a real-time monocular SLAM system capable of reconstructing kilometer-scale trajectories over 10k frames on a single consumer-grade GPU. Our model couples a lightweight front-end for low-latency online tracking with a hierarchical backend that progressively enforces local, mid-level, and global consistency. By avoiding bundle adjustment that relies on the static world assumption, our system naturally handles complex dynamic scenes out of the box. Furthermore, we demonstrate that our method can be extended to leverage stereo, RGB-D, and LiDAR as additional inputs. AMB3R-SLAM achieves strong camera tracking performance across 9 datasets, reducing the absolute trajectory error (ATE) of previous state-of-the-art methods on VBR and Oxford Spires by over 70%. With additional LiDAR input, our model further reduces ATE to sub-meter level on KITTI and VBR datasets.
Chinese Translation
我们提出了AMB3R-SLAM,这是一个实时单目SLAM系统,能够在单块消费级GPU上重建超过1万帧的公里级轨迹。我们的模型将低延迟在线跟踪的轻量级前端与渐进式实施局部、中观和全局一致性的分层后端相结合。通过避免依赖静态世界假设的束调整(bundle adjustment),我们的系统能够开箱即用地处理复杂的动态场景。此外,我们证明该方法可以扩展以利用双目相机、RGB-D相机和LiDAR作为额外输入。AMB3R-SLAM在9个数据集上实现了出色的相机跟踪性能,在VBR和Oxford Spires数据集上将以往最先进方法的绝对轨迹误差(ATE)降低了70%以上。通过额外的LiDAR输入,我们的模型在KITTI和VBR数据集上进一步将ATE降低至亚米级。
cs.CV / 14 / 2609.19542

PerSeM: Persistent Semantic Memory for Long-Horizon Open-Vocabulary UAV Mapping

PerSeM:面向长时程开放词汇无人机建图的持久语义记忆
Jamwal, Saurbh Singh, Ramakrishnan, Ganesh
Abstract
Open-vocabulary segmentation enables rich semantic perception for UAVs, but frame-wise predictions can remain temporally inconsistent across repeated observations and changing viewpoints. We present PerSeM, a training-free persistent semantic memory framework for long-horizon open-vocabulary UAV mapping. PerSeM associates frame-wise semantic observations with persistent world-space voxels and constructs a majority-based semantic memory, which is conservatively refined through history-preserving spatial refinement, trust-aware replay, and context-guided verification. Experiments on the Forest and UAVScenes benchmarks show that persistent 3D memory provides substantial gains in semantic correctness and temporal stability over frame-wise predictions. Beyond this strong persistent-memory baseline, PerSeM provides consistent additional improvements, improving both semantic accuracy and temporal stability across all five evaluated UAVScenes sequences. Analysis using regions identified independently of the final PerSeM predictions further shows that these gains are concentrated in semantically difficult and temporally unstable regions, where majority-based memory is most likely to remain uncertain. These results demonstrate that persistent 3D aggregation provides a strong foundation for long-horizon semantic mapping, while conservative refinement of uncertain memory states can provide additional improvements without retraining or additional neural-network inference.
Chinese Translation
开放词汇分割为无人机提供了丰富的语义感知能力,但在重复观测和视角变化的情况下,逐帧预测可能在时间上保持不一致。我们提出了PerSeM,一个无需训练的持久语义记忆框架,用于长时程开放词汇无人机建图。PerSeM将逐帧语义观测与持久的世界空间体素相关联,构建基于多数投票的语义记忆,并通过保留历史的空间精化、信任感知回放以及上下文引导的验证对其进行保守精化。在Forest和UAVScenes基准上的实验表明,持久的3D记忆相比逐帧预测在语义正确性和时间稳定性方面带来了显著提升。在这一强大的持久记忆基线之上,PerSeM还提供了一致的额外改进,在所有五个评估的UAVScenes序列上均提升了语义精度和时间稳定性。使用独立于PerSeM最终预测所识别区域的分析进一步表明,这些增益集中在语义困难且时间不稳定的区域,而这些正是基于多数投票的记忆最可能保持不确定的区域。这些结果表明,持久的3D聚合为长时程语义建图提供了坚实的基础,而对不确定记忆状态的保守精化可以在无需重新训练或额外神经网络推理的情况下提供额外的改进。
cs.CV / 15 / 2609.19555

A Multi-Modal Generative Model for Tomato Disease Leaves Understanding

一种用于番茄病叶理解的多模态生成模型
Quoc, Khang Nguyen, Tran, Minh-Phuoc, Truong, Gia-Han, Quach, Luyl-Da
Abstract
Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information. However, practical deployment in precision agriculture remains limited because most existing approaches treat disease understanding as isolated prediction tasks, failing to capture the complementary relationships among symptom recognition, severity assessment, and question-driven diagnostic reasoning. In tomato pathology, accurate interpretation of diseased leaves requires more than label prediction; it demands integrating visual symptoms with semantic context to support a comprehensive and explainable understanding. Here, we present SOLAR, a multimodal generative model that understands tomato disease spanning six question-answering tasks. SOLAR learns to align visual features with task-aware language representations by Fusion Expert module based on mixture-of-expert, enabling it to generate contextually relevant answers across diverse diagnostic tasks. By formulating tomato disease analysis as a generative Visual Question Answering (VQA) task, SOLAR provides a flexible framework that supports multi-task inference within a single model while improving performance and cross-task knowledge sharing. We evaluate SOLAR on $41,677$ images, including $216,209$ Question-Answering (QA) pairs to understand tomato leaf disease under both closed and open-ended QA settings. Experimental results show that SOLAR consistently outperforms state-of-the-art vision-only, vision-language, and task-specific models across all tasks, demonstrating superior accuracy, robustness, and multimodal reasoning. These findings highlight the potential of generative multimodal modeling as an effective direction for understanding of plant disease. The code for this study is available at https://github.com/EnalisUs/SOLAR.
Chinese Translation
植物病害分析的人工智能已从特定任务的分类器发展到能够联合解释视觉与文本信息的多模态模型。然而,由于现有方法大多将病害理解视为孤立的预测任务,未能捕捉症状识别、严重程度评估与问题驱动诊断推理之间的互补关系,其在精准农业中的实际部署仍然有限。在番茄病理学中,对病叶的准确解读不仅仅是标签预测,还需要将视觉症状与语义语境相整合,以支持全面且可解释的理解。本文提出了SOLAR,一个能够理解番茄病害的多模态生成模型,涵盖六项问答任务。SOLAR通过基于混合专家(mixture-of-expert)的融合专家模块(Fusion Expert module)学习将视觉特征与任务感知的语言表示对齐,从而能够在多样的诊断任务中生成符合语境的回答。通过将番茄病害分析构建为生成式视觉问答(Visual Question Answering, VQA)任务,SOLAR提供了一个灵活的框架,支持在单一模型内进行多任务推理,同时提升了性能和跨任务知识共享。我们在41,677张图像(包含216,209组问答对)上对SOLAR进行评估,以在封闭式和开放式问答设置下理解番茄叶片病害。实验结果表明,SOLAR在所有任务上均持续优于最先进的纯视觉、视觉-语言及特定任务模型,展现出卓越的准确性、鲁棒性和多模态推理能力。这些发现凸显了生成式多模态建模作为植物病害理解有效方向的潜力。本研究的代码可在 https://github.com/EnalisUs/SOLAR 获取。
cs.CV / 16 / 2609.19592

Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions

面向机器人采摘的选择性棉铃定位:田间条件下深度学习视觉模型的评估
Thayananthan, Thevathayarajh, Zhang, Xin, Badu, Isuru Laddusinghe, Harjono, Jonathan, Rains, Glen C., Li, Beiwen, Bastos, Leonardo M., Wijewardane, Nuwan K., Martins, Vitor S.
Abstract
This study developed and evaluated a deep-learning-based perception framework for selective robotic cotton picking. The dataset contained 1,008 annotated field images collected using three cameras under varying natural lighting and weather conditions. Object-detection models from the YOLOv8 through YOLOv13 families were evaluated using their default configurations, while segmentation performance was assessed using YOLOv8-seg, YOLOv11-seg, YOLOv12-seg, the Segment Anything Model (SAM), SAMv2.1, FastSAM, and Grounded-SAM with the Recognize Anything Model (RAM). Among the detection models, GELAN-s achieved the most favorable balance between mean average precision (mAP) and inference speed, obtaining an mAP of 86.1%, precision of 81.6%, recall of 76.6%, and an F1-score of 79.0%, with an average inference time of 42.3 ms per image. Among the direct segmentation models, YOLOv12-m-seg provided the most favorable balance between AP@0.5 and FPS, achieving a segmentation AP@0.5 of 83.7% with an inference time of 20.4 ms per image. In the detection-prompted segmentation approach, bounding-box prompts generated by GELAN-s improved the localization of cotton bolls for SAM and SAMv2.1, while SAMv2.1 Tiny consistently outperformed FastSAM and Grounded-SAM with RAM. In the area-based evaluation against manually annotated segmentation masks, YOLOv12-m-seg achieved an $R^2$ value of 0.966, compared with 0.860 for GELAN-s + SAMv2.1 Tiny. Field experiments conducted using a UR5e robotic manipulator, a custom end-effector, and a ZED2i stereo camera further validated the effectiveness of the YOLOv12-m-seg model for real-time cotton boll detection, segmentation, and selective picking under varying confidence levels. These results demonstrate that YOLOv12-m-seg provides an efficient perception model for robotic cotton harvesting and has strong potential for field deployment.
Chinese Translation
本研究开发并评估了一种基于深度学习的选择性机器人棉花采摘感知框架。数据集包含1,008张标注的田间图像,这些图像使用三台相机在不同自然光照和天气条件下采集。目标检测模型评估了YOLOv8至YOLOv13系列模型的默认配置,分割性能则使用YOLOv8-seg、YOLOv11-seg、YOLOv12-seg、Segment Anything Model(SAM)、SAMv2.1、FastSAM以及结合Recognize Anything Model(RAM)的Grounded-SAM进行评估。在检测模型中,GELAN-s在平均精度均值(mAP)与推理速度之间取得了最佳平衡,mAP达到86.1%,精确率为81.6%,召回率为76.6%,F1分数为79.0%,平均每张图像推理时间为42.3毫秒。在直接分割模型中,YOLOv12-m-seg在AP@0.5与FPS之间取得了最佳平衡,分割AP@0.5达到83.7%,每张图像推理时间为20.4毫秒。在检测引导的分割方法中,由GELAN-s生成的边界框提示改善了SAM和SAMv2.1对棉铃的定位效果,且SAMv2.1 Tiny始终优于FastSAM和结合RAM的Grounded-SAM。在与人工标注分割掩码的基于面积的评估中,YOLOv12-m-seg的$R^2$值达到0.966,而GELAN-s + SAMv2.1 Tiny为0.860。使用UR5e机器人机械臂、定制末端执行器和ZED2i双目相机进行的田间实验进一步验证了YOLOv12-m-seg模型在不同置信度水平下对棉铃实时检测、分割和选择性采摘的有效性。这些结果表明,YOLOv12-m-seg为机器人棉花采摘提供了一种高效的感知模型,具有很强的田间部署潜力。
cs.CV / 17 / 2609.19628

VGGT-GS SLAM: Uncalibrated Monocular Gaussian Splatting SLAM with Feed-Forward Priors

VGGT-GS SLAM:基于前馈先验的免标定单目高斯泼溅SLAM
Han, Yuhang, Wang, Hao, Cao, Jiaxi, Liu, Xingyu
Abstract
We present VGGT-GS SLAM, a monocular 3D Gaussian Splatting SLAM system designed for uncalibrated videos. Starting from feed-forward VGGT pose and depth priors, our system performs submap differentiable bundle adjustment that jointly refines camera poses and a 3D Gaussian map, while optimizing submap-shared intrinsics and radial--tangential distortion through analytic calibration Jacobians. To improve global consistency, we introduce Gaussian-native alignment (GNA) for camera-anchored scale refinement between sequential submaps and verification of loop-closure candidates. Extensive experiments on standard indoor benchmarks show consistent improvements in localization accuracy and strong rendering quality under uncalibrated settings, establishing a strong baseline for uncalibrated Gaussian SLAM.
Chinese Translation
我们提出了VGGT-GS SLAM,这是一个面向免标定视频的单目3D高斯泼溅(Gaussian Splatting)SLAM系统。该系统以前馈式VGGT位姿和深度先验为起点,执行子地图可微分光束法平差(bundle adjustment),联合优化相机位姿和3D高斯地图,同时通过解析标定雅可比矩阵优化子地图共享的内参和径向-切向畸变。为提升全局一致性,我们引入了高斯原生对齐(Gaussian-native Alignment, GNA),用于序列子地图间基于相机锚定的尺度细化以及回环候选的验证。在标准室内基准数据集上的大量实验表明,该系统在免标定设置下能够持续提升定位精度并具有出色的渲染质量,为免标定高斯SLAM建立了一个强有力的基线。
cs.CV / 18 / 2609.19631

Instance Segmentation and Fine-grained Classification for Urban Buildings with Adaptive Region Dividing and Spatially-Supervised Contrastive Learning

基于自适应区域划分与空间监督对比学习的城市建筑实例分割与细粒度分类
Zhang, Weiyuan, Zhang, Qi, Huang, Hui
Abstract
Accurate instance-level and functional understanding of urban buildings in large-scale point clouds is essential for digital city modeling and urban analysis. However, the extensive spatial coverage of urban scenes leads most existing methods to rely on predefined blocks for training and evaluation, although such partitions are rarely available in real-world applications and introduce additional preprocessing while fragmenting complete building structures. To address this issue, we propose an adaptive region-dividing strategy with unified scene-level evaluation. Specifically, the 3D point cloud is projected onto a bird's-eye-view (BEV) plane, where a pretrained segmentation model is used to detect building regions. The detected bounding boxes are then back-projected to the original point cloud to construct structure-aligned adaptive training blocks, enabling semantically guided dynamic partitioning without manual design. Furthermore, beyond instance-level understanding, few methods have explored fine-grained classification for urban buildings, and thus we also put forward a fine-grained classification model for urban buildings with a spatially-supervised contrastive loss. First, for each segmented building, a point transformer classifier jointly encodes its body and local context using geometric, color, and core-context information. Then, the class-balanced weighted cross-entropy is used to alleviate severe class imbalance. The proposed spatially-supervised contrastive loss further enhances inter-class discriminability by assigning greater weight to spatially proximate, same-category buildings, encouraging compact functional representations while separating easily confused categories. Extensive experiments on UrbanBIS and STPLS3D demonstrate the advantages of the proposed method in building instance segmentation and fine-grained classification compared to existing SOTA methods.
Chinese Translation
在大规模点云中对城市建筑进行精确的实例级与功能性理解,对于数字城市建模和城市分析至关重要。然而,城市场景的空间覆盖范围广,导致大多数现有方法依赖于预定义的分块进行训练与评估,尽管此类划分在实际应用中很少可用,且引入了额外的预处理,同时割裂了完整的建筑结构。为解决这一问题,我们提出了一种自适应区域划分策略并采用统一的场景级评估。具体而言,将三维点云投影到鸟瞰图(BEV)平面上,利用预训练的分割模型检测建筑区域,然后将检测到的边界框反投影回原始点云,构建与建筑结构对齐的自适应训练分块,实现无需人工设计的语义引导动态划分。此外,在实例级理解之外,很少有方法探索城市建筑的细粒度分类,因此我们还提出了一种采用空间监督对比损失的城市建筑细粒度分类模型。首先,针对每个分割出的建筑,点Transformer分类器(Point Transformer)联合编码其本体与局部上下文,利用几何、颜色及上下文核心信息。然后,采用类别平衡的加权交叉熵以缓解严重的类别不平衡。所提出的空间监督对比损失进一步增强了类间可区分性,其通过为空间上邻近的同类别建筑赋予更大权重,在促进紧凑的功能表示的同时分离易混淆的类别。在UrbanBIS和STPLS3D数据集上的大量实验表明,与现有最先进(SOTA)方法相比,所提方法在建筑实例分割与细粒度分类方面具有优势。
cs.CV / 19 / 2609.19634

Scientific Image Quality Assessment via Multi-modal Retrieval-Augmented Generation

基于多模态检索增强生成的科学图像质量评估
Zhang, Yinuo, Liu, Bingshuo, Tu, Zhiying, Chu, Dianhui, Liu, Qingbin, Chen, Xi, Bian, Jiang, Yu, Xiaoyan, Sui, Dianbo
Abstract
This paper proposes a Retrieval-Augmented Generation (RAG) framework for scientific image quality assessment, designed to simultaneously address both the understanding track (SIQA-U) and the scoring track (SIQA-S) of the SIQA challenge. We construct a multimodal index that integrates textual semantics with fine-grained visual features, and develop a multi-route retrieval and fusion mechanism to provide large language models with highly relevant reference cases, thereby enhancing their capability to evaluate complex scientific images. Experimental results demonstrate that the proposed framework effectively aligns with the judgment criteria of human experts. Ultimately, our method achieves 1st place in the SIQA-U track of the SIQA challenge at the ICME 2026 Grand Challenges.
Chinese Translation
本文提出了一种用于科学图像质量评估的检索增强生成(Retrieval-Augmented Generation, RAG)框架,旨在同时应对SIQA挑战赛的理解赛道(SIQA-U)和评分赛道(SIQA-S)。我们构建了一个将文本语义与细粒度视觉特征相融合的多模态索引,并设计了一种多路检索与融合机制,为大语言模型提供高度相关的参考案例,从而增强其评估复杂科学图像的能力。实验结果表明,所提出的框架能够有效契合人类专家的评判标准。最终,我们的方法在ICME 2026 Grand Challenges的SIQA挑战赛SIQA-U赛道中获得第一名。
cs.CV / 20 / 2609.19662

Towards Active Cross-View Object Geo-Localization

迈向主动式跨视角目标地理定位
Yao, Shunyu, Zhang, Xiaohan, Yang, Zhuoran, Lai, Haoqi, Ming, Qi, Hu, Xiaoxi, Shen, Hui-Liang, Cao, Si-Yuan
Abstract
Cross-view object geo-localization (CVOGL) typically assumes a fixed query image, overlooking the ability of mobile agents to actively acquire more informative observations. To address this limitation, we introduce Active Cross-View Object Geo-Localization (ActiveGeo), where an agent sequentially selects new viewpoints and determines when to stop, aiming to improve localization with minimal observations. We further propose ActiveMoPT, an ActiveGeo framework with three-stage training. First, Multi-View Prompt-Preserving Adaptation enables the model to aggregate multiple query views while reusing the initial prompt. Second, Trajectory-Guided Policy Initialization uses supervised agent trajectories to learn viewpoint selection and initial stopping behavior. Third, Cost-Aware Policy Refinement employs GRPO with a gain-cost reward to jointly optimize localization accuracy and observation efficiency. We also construct ActiveGeo-858, a zero-shot test set containing 858 scenes and 1,716 target annotations. Experiments show that ActiveMoPT achieves state-of-the-art performance on MoP-UAV using only 1.45 query views on average, and substantially outperforms previous CVOGL approaches under zero-shot evaluation on ActiveGeo-858.
Chinese Translation
跨视角目标地理定位(Cross-View Object Geo-Localization, CVOGL)通常假设查询图像是固定的,忽略了移动智能体主动获取更具信息量的观测的能力。为解决这一局限,我们提出了主动式跨视角目标地理定位(Active Cross-View Object Geo-Localization, ActiveGeo),其中智能体依次选择新的视点并决定何时停止观测,旨在以尽可能少的观测提升定位精度。我们进一步提出ActiveMoPT,一个包含三阶段训练的ActiveGeo框架。首先,多视角提示保留适配(Multi-View Prompt-Preserving Adaptation)使模型能够聚合多个查询视角,同时复用初始提示。其次,轨迹引导的策略初始化(Trajectory-Guided Policy Initialization)利用有监督的智能体轨迹来学习视点选择和初始停止行为。第三,代价感知的策略精化(Cost-Aware Policy Refinement)采用带收益-代价奖励的GRPO来联合优化定位精度与观测效率。我们还构建了ActiveGeo-858,一个包含858个场景和1,716个目标标注的零样本测试集。实验表明,ActiveMoPT在MoP-UAV上仅平均使用1.45个查询视角即取得最先进的性能,并在ActiveGeo-858的零样本评测中显著优于以往CVOGL方法。
cs.CV / 21 / 2609.19664

VideoResearcher: Self-Improving Tool Design for Long-Video Understanding

VideoResearcher:面向长视频理解的自改进工具设计
Ye, Dingqiang, Zhao, Dongdi, Wang, Kaishen, Hu, Qingqiao, Sun, Jingchen, Liang, Yijun, Jia, Yuqi, Huang, Yiqiao, Tian, Yunjie, Zhang, Jiaxing, Jin, Chuanyang, Zhang, Ke, Patel, Vishal M., Fu, Di
Abstract
Video agents have made substantial progress in long-video understanding. Yet effective video-agent systems require costly, time-consuming manual design and trial and error. Current self-improvement methods either refine low-impact prompts, recombine predefined micro-tools, or struggle with convergence in harness optimization. To bridge this gap, we target high-impact video-tool with VideoResearcher, a training-free multi-agent framework that autonomously designs, tests, and refines tools for video understanding, like a human researcher. VideoResearcher operates through dual Solving and Evolving loops: it analyzes tool-use trajectories to identify capability gaps, coordinates specialized agents to develop and validate executable tools, and reuses evolved tools to strengthen evidence acquisition in subsequent video reasoning. Through iterative tool refinement and validation, it progressively strengthens evidence acquisition without updating model parameters. VideoResearcher achieves state-of-the-art performance among self-improving agents and approaches the human-designed upper bound, demonstrating a training-free paradigm for long-video understanding that expands agent capabilities through autonomous tool development while reducing costly manual engineering.
Chinese Translation
视频智能体(video agent)在长视频理解方面已取得长足进展。然而,构建有效的视频智能体系统需要成本高昂、耗时费力的手动设计与反复试错。现有的自改进方法要么只对低影响的提示词进行优化,要么重新组合预定义的微型工具,要么在框架优化中难以收敛。为弥合这一差距,我们提出了 VideoResearcher,聚焦于高影响的视频工具。这是一个无需训练的多智能体框架,能够像人类研究者一样,自主地为视频理解设计、测试和改进工具。VideoResearcher 通过“求解(Solving)”与“演化(Evolving)”双循环运行:它分析工具使用轨迹以识别能力缺口,协调专门的智能体来开发并验证可执行工具,并复用演化后的工具以强化后续视频推理中的证据获取。通过迭代式的工具改进与验证,它在不更新模型参数的情况下逐步增强证据获取能力。VideoResearcher 在自改进智能体中取得了最先进的性能,并接近人工设计的上限,展示了一种无需训练的长视频理解范式——通过自主的工具开发扩展智能体能力,同时减少昂贵的人工工程投入。
cs.CV / 22 / 2609.19669

Beyond Patch Removal: Persistent Adversarial Effects in Vision-Language-Action Policies

超越补丁移除:视觉-语言-动作策略中的持续性对抗效应
Wu, Enhao, Guo, Fusen, Cao, Yuxin, Lyu, Ziyang, Li, Lin, Song, Wei
Abstract
Adversarial patches to Vision-Language-Action (VLA) policies can cause both immediate action corruption and persistent state effects that remain after the patch is removed. Existing evaluations largely focus on continuous attacks and do not separate these two effects. We introduce a state-restoration protocol that removes the patch at matched action-chunk boundaries and measures subsequent recoverability under the same remaining step budget. Clean, random-patch, deviation-matched, and fixed-direction controls distinguish adversarial effects from occlusion, action-error magnitude, and directional persistence. We also evaluate a recovery adapter trained on attack-induced states under controlled intervention latency. On OpenVLA-OFT with EDPA attacks, only 36.2% of LIBERO-Long episodes remain recoverable after five chunks, compared with 89.9% and 87.0% for the deviation-matched and fixed-direction controls. Similar persistent effects are observed on autoregressive OpenVLA. The recovery adapter improves recovery from 7.7% to 47.4% at one-chunk latency, but its benefit decreases substantially with delayed intervention. These results show that adversarial effects can persist after patch removal and that timely intervention is critical for recovery.
Chinese Translation
针对视觉-语言-动作(VLA)策略的对抗补丁既能造成即时的动作破坏,也会产生在补丁移除后仍然持续的状态效应。现有评估大多关注持续攻击,未能区分这两种效应。我们提出一种状态恢复协议,在匹配的动作块(action-chunk)边界处移除补丁,并在相同的剩余步数预算下测量后续可恢复性。通过干净、随机补丁、偏差匹配和固定方向等对照实验,将对抗效应与遮挡、动作误差幅度以及方向持续性区分开来。我们还在受控干预延迟条件下,评估了一个在攻击诱导状态下训练的恢复适配器(recovery adapter)。在采用EDPA攻击的OpenVLA-OFT上,五个动作块后仅36.2%的LIBERO-Long回合仍可恢复,而偏差匹配和固定方向对照分别为89.9%和87.0%。在自回归OpenVLA上也观察到类似的持续效应。恢复适配器在一个动作块延迟下将恢复率从7.7%提升至47.4%,但其收益随干预延迟增加而显著下降。这些结果表明,对抗效应可在补丁移除后持续存在,且及时干预对恢复至关重要。
cs.CV / 23 / 2609.19693

IMFD: End-to-end Multi-Face Forgery Detection through Instruction-based Large Vision-Language Models

IMFD:基于指令型大型视觉语言模型的端到端多人脸伪造检测
Choi, Dasom, Moon, Sangjun, Im, Hyeongchan, Park, Jaeeon, Kwon, Jingun, Kamigaito, Hidetaka, Watanabe, Taro, Okumura, Manabu
Abstract
The rapid increase of deepfakes has raised significant concerns due to their spread on social media. Traditional multi-face forgery detectors crop and verify each face independently, ignoring background context and inter-face relationships, which often yields suboptimal performance. To overcome these limitations, we leverage instruction-based Large Vision-Language Models (LVLMs), which can interpret entire images and follow complex textual instructions. We propose a simple yet effective single-stage multi-face forgery detector, called IMFD (Instruction-based Multi-face Forgery Detector), which is trained end-to-end to jointly localize faces and predict per-face forgery labels. Rather than treating face box prediction only as a joint objective, IMFD explicitly integrates predicted face bounding boxes into the instruction as visual cues that enhance instruction grounding and forgery detection. To support the training and evaluation of IMFD, we convert existing multi-face forgery datasets into an instruction-based format. Experimental results and analyses show that IMFD improves multi-face forgery detection by integrating face bounding boxes into the instruction, and consistently outperforms various state-of-the-art methods.
Chinese Translation
深度伪造内容的快速增加及其在社交媒体上的传播引发了严重关切。传统的多人脸伪造检测器独立地裁剪并验证每个人脸,忽略了背景上下文和人脸之间的关系,因此往往性能欠佳。为克服这些局限,我们利用指令型大型视觉语言模型(Large Vision-Language Models,LVLMs),该模型能够理解整幅图像并遵循复杂的文本指令。我们提出了一种简单而有效的单阶段多人脸伪造检测器,称为IMFD(Instruction-based Multi-face Forgery Detector,基于指令的多人脸伪造检测器),它以端到端方式训练,可联合定位人脸并预测每个人脸的伪造标签。IMFD不仅将人脸框预测作为联合目标,还将预测的人脸边界框显式地整合到指令中,作为增强指令对齐和伪造检测的视觉线索。为支持IMFD的训练和评估,我们将现有多人脸伪造数据集转换为指令型格式。实验结果与分析表明,IMFD通过将人脸边界框整合到指令中提升了多人脸伪造检测性能,并持续超越多种最先进的方法。
cs.CV / 24 / 2609.19702

Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation

理解并利用自回归图像生成中的对角注意力稀疏性
Kim, Daeun, Hong, Junwha, Oh, Changhun, Kim, Yoonsung, Lee, Yoonhyeong, Park, Jongse
Abstract
Autoregressive image generation has emerged as a paradigm for multimodal AI systems due to its compatibility with transformer-based LLM serving infrastructures. However, generating thousands of visual tokens per request makes decoding increasingly bottlenecked by KV cache accesses during attention computation. Sparse attention is particularly attractive for this workload because many visual generation applications tolerate moderate quality degradation in exchange for improved performance and efficiency. While sparse attention has been extensively explored for text-based LLM inference, it remains unclear whether its sparsity assumptions generalize effectively to autoregressive image generation. We present the first systematic characterization of attention sparsity in autoregressive image generation across diverse workloads and representative open-source models. Our analysis reveals several distinguishing properties, including a pronounced prefill-decode asymmetry, strong attention concentration on prompt and local tokens, and a unique diagonal attention sparsity pattern arising from the spatial locality of visual tokens. Motivated by these observations, we propose a diagonal-aware sparse attention mechanism that selectively skips KV entries along the diagonal attention direction within a recent window. Implemented on top of a GPU-based serving system using FlexGen, FlashAttention-2, and custom kernels, our approach achieves up to 3.1x throughput and 1.19x latency improvements with less than 2% quality degradation compared to dense inference.
Chinese Translation
自回归图像生成因其与基于Transformer的大语言模型(LLM)服务基础设施的兼容性,已成为多模态AI系统的一种重要范式。然而,每次请求需要生成数千个视觉token,使得解码过程日益受到注意力计算中KV缓存访问的瓶颈制约。稀疏注意力对这类工作负载尤其具有吸引力,因为许多视觉生成应用能够容忍一定程度的质量下降,以换取性能和效率的提升。尽管稀疏注意力在基于文本的LLM推理中已被广泛研究,但其稀疏性假设能否有效推广到自回归图像生成仍不清楚。我们首次对自回归图像生成中的注意力稀疏性进行了系统性的刻画,涵盖了多样化的工作负载和具有代表性的开源模型。我们的分析揭示了若干独特性质,包括显著的预填充-解码不对称性、注意力高度集中于提示词(prompt)和局部token,以及由视觉token空间局部性产生的独特对角注意力稀疏模式。基于这些观察,我们提出了一种对角感知的稀疏注意力机制,可在近期窗口内沿对角注意力方向选择性地跳过KV条目。该机制实现于基于GPU的服务系统之上,结合了FlexGen、FlashAttention-2以及自定义内核。与稠密推理相比,我们的方法在质量下降不超过2%的情况下,实现了最高3.1倍的吞吐量提升和1.19倍的延迟改善。
cs.CV / 25 / 2609.19716

GAPrompt++: Multi-Granular Geometry-Aware Point Cloud Prompt for 3D Vision Model

GAPrompt++:面向3D视觉模型的多粒度几何感知点云提示方法
Ai, Zixiang, Cui, Zhenyu, Guo, Yufei, Qiang, Wenwen, Chen, Lei, Lu, Jiwen, Zhou, Jiahuan
Abstract
Pre-trained 3D vision models have substantially advanced point cloud analysis, yet adapting them to downstream tasks via full fine-tuning is computationally expensive and storage-intensive. Parameter-Efficient Fine-Tuning (PEFT) offers a promising alternative by reducing both adaptation cost and storage burden. However, existing prompting-based approaches ignore the intrinsic geometric structures of point clouds, thereby limiting their adaptation capability. This limitation stems from their inability to encode both fine-grained geometric cues and coarse-grained structural semantics, as well as failing to propagate such information effectively through the model hierarchy. To address these challenges, we propose GAPrompt++, a multi-granular geometry-aware prompting method that provides richer geometric guidance for efficient 3D task adaptation. Specifically, we introduce a Point Shift Prompter that extracts multi-granular geometric features across different scales, enabling instance-specific geometric adjustments during adaptation. Next, a Keypoint Prompter adaptively generates point-level prompts to highlight local geometric saliency and fine-grained structural details. Furthermore, a Prompt Propagation mechanism injects these multi-granular geometric cues throughout the feature extraction hierarchy, strengthening the ability to capture essential geometric characteristics. Extensive experiments show that GAPrompt++ achieves state-of-the-art performance among prompting-based PEFT methods and even surpasses full fine-tuning across diverse benchmarks, while requiring less than 2\% trainable parameters. In addition, to address the saturation of existing evaluation datasets, we construct two more challenging benchmarks derived from 3D Gaussian Splatting and Multi-View Stereo reconstruction, offering diverse and realistic point cloud scenarios to promote future research.
Chinese Translation
预训练3D视觉模型极大地推动了点云分析的发展,然而通过全量微调将其适配到下游任务在计算和存储上都十分昂贵。参数高效微调通过降低适配成本和存储负担提供了一种有前景的替代方案。然而,现有基于提示的方法忽略了点云的内在几何结构,从而限制了其适配能力。这一局限源于它们既无法编码细粒度几何线索与粗粒度结构语义,也无法在模型层级中有效传播这些信息。为应对这些挑战,我们提出了GAPrompt++,一种多粒度几何感知提示方法,可为高效的3D任务适配提供更丰富的几何引导。具体而言,我们引入了点偏移提示器,用于提取不同尺度下的多粒度几何特征,从而在适配过程中实现针对特定实例的几何调整。其次,关键点提示器自适应地生成点级提示,以突出局部几何显著性及细粒度结构细节。此外,提示传播机制将这些多粒度几何线索注入整个特征提取层级,增强了捕捉关键几何特性的能力。大量实验表明,GAPrompt++在基于提示的PEFT方法中取得了最先进的性能,甚至在多个基准测试中超越了全量微调,同时可训练参数量不足2%。此外,为解决现有评估数据集趋于饱和的问题,我们基于3D高斯泼溅(3D Gaussian Splatting)和多视图立体重建构建了两个更具挑战性的基准,提供了多样且真实的点云场景,以促进未来研究。
cs.CV / 26 / 2609.19719

SeetaPsych v1.0: An Open-source Computer Vision Toolkit for Behavior-based Psychological Measurement

SeetaPsych v1.0:一个面向基于行为的心理测量的开源计算机视觉工具包
Zeng, Jiabei, Li, Chiqin, Li, Kaizhou, Chang, Fei, Li, Yong, Zhao, Yuanhao, Han, Dan, Yang, Wenqiang, Chen, Xilin, Shan, Shiguang
Abstract
Automated visual analysis opens new avenues for behavior--based psychological measurement. Nevertheless, existing technological modules are typically scattered across task specific systems with heterogeneous interfaces and disparate deployment requirements. In this work, we present SeetaPsych v1.0, an open source, unified and extensible computer vision toolkit designed to extract psychologically relevant signals from facial images and/or face based videos. The current release encompasses four major core modules aiming at behavior--based physiological perception: unified face based emotion analysis (simultaneous facial expression recognition, facial action unit detection, and valence--arousal estimation), camera based heart rate estimation, screen point--of--gaze estimation, and scene gaze following. A suite of auxiliary preprocessing modules for human centric visual analysis is also included, comprising face detection, facial landmark detection, and head detection. These functionalities are encapsulated within a modular Pipeline/Runner architecture that automatically resolves attribute dependencies, constructs computation graphs, and support intermediate result sharing among modules. SeetaPsych provides standardized Python APIs to facilitate reproducible, large scale analyses, alongside an interactive WebUI for rapid, code--free method evaluation. Overall, SeetaPsych offers an integrated and accessible visual measurement platform for research in psychology, behavioral science, human computer interaction, and related fields.
Chinese Translation
自动化视觉分析为基于行为的心理测量开辟了新的途径。然而,现有的技术模块通常分散于面向特定任务的系统中,具有异构的接口和不同的部署要求。在本工作中,我们提出了SeetaPsych v1.0,一个开源、统一且可扩展的计算机视觉工具包,旨在从面部图像和/或基于人脸的视频中提取与心理相关的信号。当前版本包含四个面向基于行为的生理感知的核心模块:统一的人脸情感分析(同时进行面部表情识别、面部动作单元检测和效价-唤醒度估计)、基于摄像头的心率估计、基于屏幕的注视点估计以及场景注视跟随。此外,该工具包还包含一套面向以人为中心的视觉分析的辅助预处理模块,包括人脸检测、面部关键点检测和头部检测。这些功能被封装在一个模块化的Pipeline/Runner架构中,该架构能够自动解析属性依赖关系、构建计算图,并支持模块间的中间结果共享。SeetaPsych提供了标准化的Python API,以便于可复现的大规模分析,同时还提供一个交互式WebUI,用于快速、无需编写代码的方法评估。总体而言,SeetaPsych为心理学、行为科学、人机交互及相关领域的研究提供了一个集成且易于使用的视觉测量平台。
cs.CV / 27 / 2609.19729

Recency Forcing: Bridging the Long-Horizon Gap in Autoregressive Video Generation

近因强制(Recency Forcing):弥合自回归视频生成中的长时程差距
Cao, Tri, Nguyen, Hung, Nguyen, Phong, Nguyen, Khoi
Abstract
Autoregressive (AR) video generation degrades over long horizons due to an overlooked train-inference discrepancy we term KV eviction mismatch: models train on short clips where all context frames reside in the KV cache, but at inference, memory constraints force distant frames to be evicted from the KV cache - removing context the model was conditioned on. Rather than simulating eviction via context truncation - which discards temporal information the model still needs and degrades motion coherence - we keep the context but while progressively reducing the influence of distant frames, making their eventual eviction negligible. To guide this design, we introduce the positional response $R( \Delta, \, t_{\text{denoise}})$, a perturbation-based sensitivity measure revealing that context influence decays steeply with temporal distance and varies systematically across denoising steps. Motivated by this analysis, we propose Recency Forcing, which applies a non-positive, timestep-dependent bias, termed Temporal Response Bias (TRB), on pre-softmax attention logits derived directly from $R$, closing the train-inference gap without modifying context length or training objectives. We further introduce Biased Attention Reparameterization (BAR), an exact reformulation that moves the bias outside the softmax, making TRB a standard FlashAttention call at zero overhead. Recency Forcing operates in both training-free mode and training-based mode. Experiments on VBench and VBench-Long demonstrate state-of-the-art long-horizon generation quality at no additional inference cost.
Chinese Translation
自回归(AR)视频生成长时程性能下降的原因在于一种被忽视的训练-推理差异,我们将其命名为KV缓存淘汰失配(KV eviction mismatch):模型在短片段上训练时,所有上下文帧都保留在KV缓存中;但在推理时,内存限制迫使远处的帧从KV缓存中被淘汰——这移除了模型曾依赖的上下文信息。以往方法通过截断上下文来模拟淘汰,但这会丢弃模型仍然需要的时间信息并降低运动连贯性。与之不同,我们保留上下文,同时逐步降低远处帧的影响,使其最终被淘汰的影响可以忽略不计。为指导这一设计,我们引入位置响应 $R(\Delta, t_{\text{denoise}})$,这是一种基于扰动的敏感度度量,揭示了上下文影响力随时间距离急剧衰减,并在不同去噪步骤间呈现系统性变化。基于这一分析,我们提出Recency Forcing,直接由 $R$ 导出一种非正的、依赖时间步的偏置项——称为时间响应偏置(Temporal Response Bias, TRB)——应用于softmax前的注意力logits上,从而在不修改上下文长度或训练目标的情况下弥合训练-推理差距。我们进一步提出偏置注意力重参数化(Biased Attention Reparameterization, BAR),这是一种精确的等价重构,将偏置移到softmax之外,使TRB可以通过标准的FlashAttention调用实现,且零额外开销。Recency Forcing既支持无训练模式,也支持基于训练的模式。在VBench和VBench-Long上的实验表明,该方法在无额外推理成本的情况下实现了最先进的长时程生成质量。
cs.CV / 28 / 2609.19740

Federated Learning Framework for Privacy-Preserving Kidney Stone Detection

面向隐私保护肾结石检测的联邦学习框架
Younas, Najiyya, Abdulkader, Omar, Shah, Yaser Ali, Ikram, Muhammad Jawad, Khan, Jebran, Khalil, Amaad
Abstract
Recent innovations in deep learning have significantly enhanced the diagnosis of medical images, although they are based on the use of centralized data storage that pose severe threats to patient privacy and medical data security. To address this issue, this research proposes a Federated Learning (FL) model that is coupled with an optimized YOLOv8 network to detect the kidney stones on a computed tomography (CT) image and at the same time, protect privacy of the patients. The suggested system can help various medical organizations to jointly train a common model without exchanging the information about the patients. This is to ensure that data protection laws like GDPR and HIPAA are adhered to. The residual feature fusion and DropBlock regularization among other architectural improvements are also included in YOLOv8 to enhance detection robustness and minimize overfitting. Experimental analysis carried out on a distributed CT dataset demonstrated that the federated YOLOv8 model has a mAP at 50 of 0.733 and is able to keep the data confidential. Moreover, its lean design facilitates fast edge deployment and real-time inference across a clinical setting. Altogether, these findings indicate that Federated Learning is a safe and efficient solution to AI-assisted diagnosis in contemporary healthcare when combined with the use of sophisticated object detection models.
Chinese Translation
深度学习的最新创新显著提升了医学图像的诊断水平,但其依赖于集中式数据存储,这对患者隐私和医疗数据安全构成了严重威胁。为解决这一问题,本研究提出了一种联邦学习(Federated Learning, FL)模型,并结合优化的YOLOv8网络,在计算机断层扫描(CT)图像中检测肾结石,同时保护患者隐私。所提出的系统可使多个医疗机构在不交换患者信息的情况下联合训练一个共同模型,从而确保符合GDPR和HIPAA等数据保护法规。此外,YOLOv8还引入了残差特征融合和DropBlock正则化等架构改进,以增强检测的鲁棒性并减少过拟合。在分布式CT数据集上进行的实验分析表明,联邦YOLOv8模型的mAP@50达到0.733,同时能够保证数据的机密性。而且,其精简的设计便于在临床环境中快速进行边缘部署和实时推理。综上所述,这些结果表明,联邦学习与先进的目标检测模型相结合,是当代医疗中AI辅助诊断的一种安全且高效的解决方案。
cs.CV / 29 / 2609.19745

Region-Level Policy Optimization for Fine-grained MLLM Perception

面向细粒度多模态大模型感知的区域级策略优化
Shi, Yuheng, Pei, Xiaohuan, Dong, Minjing, Xu, Chang
Abstract
Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have different resolution requirements. In a controlled diagnostic, localization tolerates roughly 3 to 4 times stronger token compression than recognition, which motivates localizing from a coarse view and concentrating resolution on the selected evidence. Decoding coordinates with the MLLM can be trained end-to-end from answers, but costs a full model pass per query and depends on grounding ability. A lightweight proposal network distilled from the model's attention is fast, but inherits the noise of its attention targets. The RoI from the proposal network reaches the answer through a discrete region choice, so its faithfulness to the answer cannot supervise the network. We therefore optimize the proposal network with region-level reinforcement learning, which we call Vision-RL2. It treats coherent regions as actions, and a frozen MLLM reader scores each one by how its removal changes the answer likelihood. Complementary subtractive and additive objectives suppress distracting proposals and recover missing evidence, updating only the predictor without region annotations, response sampling, or reasoning trajectories. The refined proposal further enables a sparse encoding that magnifies evidence and excludes background tokens. Across six fine-grained benchmarks and four MLLM backbones, Vision-RL2 improves accuracy over the base model at every token budget and surpasses its largest-budget accuracy with about 4 times fewer visual tokens. Code is available at https://github.com/YuHengsss/VisionRL2 .
Chinese Translation
多模态大语言模型(MLLM)的细粒度视觉感知通常通过提高分辨率来改进,但新增的视觉令牌会增加视觉编码和语言模型预填充(prefilling)的开销。我们发现,细粒度感知所依赖的两个操作——定位感兴趣区域(RoI)和识别其内容——对分辨率的需求不同。在一项受控诊断实验中,定位任务能够容忍比识别任务强约3至4倍的令牌压缩,这启发我们先从粗粒度视图进行定位,再将分辨率集中于所选证据。通过MLLM解码坐标可以从答案端到端地训练,但每次查询需要完整的前向计算,且依赖模型的定位能力。从模型注意力中蒸馏得到的轻量级候选区域生成网络速度快,但会继承注意力目标的噪声。候选网络生成的RoI通过一次离散的区域选择到达答案,因此其与答案的忠实性无法用于监督该网络。为此,我们采用区域级强化学习来优化候选网络,称之为Vision-RL2。该方法将连贯区域视为动作,由一个冻结的MLLM阅读器根据移除该区域后答案似然的变化对每个区域进行打分。互补的减法目标与加法目标分别抑制干扰性候选区域并恢复缺失的证据,仅更新预测器本身,无需区域标注、响应采样或推理轨迹。优化后的候选区域进一步支持一种稀疏编码方式,放大证据并排除背景令牌。在六个细粒度基准和四个MLLM骨干模型上,Vision-RL2在每个令牌预算下都优于基础模型的准确率,并以约4倍更少的视觉令牌超越了其最大预算下的准确率。代码已发布于 https://github.com/YuHengsss/VisionRL2 。
cs.CV / 30 / 2609.19747

STAR: Structure-aware Test-time Adaptation for diffusion-based light field Reconstruction

STAR:面向基于扩散模型的光场重建的结构感知测试时自适应方法
Choi, Wontae, Moon, Ki Ryum, Lee, Jae Young, Yun, Hyung Sup, Chun, Il Yong
Abstract
Light field (LF) reconstruction from limited and noisy focal stack (FS) measurements is a highly ill-posed inverse problem. Although the LF-to-FS imaging geometry is fixed for a given optical setup, LF spatial-angular structure---including within-view spatial details, cross-view angular dependencies, and disparity across views---varies across scenes. Consequently, a fixed pre-trained prior may not optimally capture the spatial-angular structure of each test LF. We propose Structure-aware Test-time Adaptation for diffusion-based light field Reconstruction (STAR), the first test-time adaptation framework for reconstructing an LF from FS. For each test LF, STAR freezes a pre-trained diffusion prior and fits three lightweight adapters to the observed FS to jointly adapt the three components of the LF's spatial-angular structure. STAR outperforms existing state-of-the-art methods in both two- and three-focal-sheet settings, with shorter inference times than those with test-time parameter updates.
Chinese Translation
从有限且含噪的焦栈(Focal Stack, FS)测量数据进行光场(Light Field, LF)重建是一个高度不适定的逆问题。尽管对于给定的光学系统,LF到FS的成像几何是固定的,但LF的空-角结构——包括视图内空间细节、跨视图角度依赖性以及视图间视差——会随场景而变化。因此,固定的预训练先验可能无法最优地捕捉每个测试LF的空-角结构。我们提出了结构感知的测试时自适应扩散光场重建方法(STAR),这是首个从FS重建LF的测试时自适应框架。对于每个测试LF,STAR冻结预训练的扩散先验,并对观测到的FS拟合三个轻量级适配器,以联合自适应LF空-角结构的三个组成部分。在双焦面和三焦面设置下,STAR均优于现有最先进的方法,且推理时间短于其他需要测试时参数更新的方法。
cs.CV / 31 / 2609.19767

Benchmarking MLLMs via Cognitive Expected Scene Graph for Safety-Critical Visual Negation Understanding

基于认知期望场景图的安全关键视觉否定理解多模态大语言模型基准测试
Jiang, Zhiyun, Wang, Hanyong, Liang, Binbin, Xie, Yu, Yang, Menglong, Li, Wei
Abstract
True machine intelligence requires transcending passive pixel registration to master top-down functional reasoning over absent information via visual negation understanding. However, unconstrained visual negation paradigms remain overly open-ended, and pervasive affirmation bias causes both existing Multi-Modal Large Language Models (MLLMs) and evaluation metrics to fail under negative semantics. To solve these intertwined challenges systematically, we first anchor the boundaries of negation reasoning within specific cognitive goals. Specifically, by focusing on safety as a highly pragmatic and critical cognitive dimension, we define the task of \textbf{S}cene \textbf{N}egation \textbf{U}nderstanding under \textbf{S}afety Cognition (\textbf{SNUS}). Under this framework, we construct a high-fidelity negative caption dataset mapping dense assertions of localized hazards. Concurrently, we propose the Cognitive Expected Scene Graph (CESG) Score, a structure-grounded, polarity-aware evaluation metric. Extensive experiments demonstrate that while current models struggle on the task, traditional metrics completely collapse under semantic reversals. Conversely, our framework delivers a solid benchmark for SNUS, providing a rigorous foundation to advance risk-aware situational comprehension and counterfactual cognition.
Chinese Translation
真正的机器智能需要超越被动的像素记录,通过视觉否定理解掌握对缺失信息进行自上而下功能推理的能力。然而,无约束的视觉否定范式过于开放,且普遍存在的肯定偏差导致现有多模态大语言模型(MLLMs)和评测指标在负向语义下均告失效。为系统地解决这些相互交织的挑战,我们首先将否定推理的边界锚定在特定的认知目标之内。具体而言,通过聚焦于安全这一高度实用且关键的认知维度,我们定义了安全认知下的场景否定理解任务( extbf{S}cene extbf{N}egation extbf{U}nderstanding under extbf{S}afety Cognition, extbf{SNUS})。在该框架下,我们构建了一个高保真的负向描述数据集,将密集的断言映射到局部化危险。同时,我们提出了认知期望场景图(CESG)分数——一种基于结构、具有极性感知能力的评测指标。大量实验表明,尽管当前模型在该任务上表现不佳,传统指标在语义反转下更是完全失效。相反,我们的框架为SNUS提供了一个坚实的基准,为推进风险感知的场景理解和反事实认知奠定了严谨的基础。
cs.CV / 32 / 2609.19793

AI Smart Glasses for Wearable Intelligence: From Egocentric Sensing to Agentic Personalization

面向可穿戴智能的AI智能眼镜:从自我中心感知到智能体化个性化
Yuan, Xu, Wang, Yi, Jiang, Zhuohang, Qu, Haohao, Ding, Yujuan, Lin, Shanru, Xing, Guoliang, Yang, Hongxia, Cao, Jiannong, Li, Qing, Fan, Wenqi
Abstract
Recent advances in artificial intelligence (AI) are reshaping smart glasses from egocentric capture and display devices into platforms for wearable intelligence. Smart glasses increasingly serve as wearable AI systems that connect first-person observation with real-time assistance under strict form-factor constraints. We frame this transition through the lens of \emph{AI smart glasses} and define them as a system-level concept in which egocentric sensing, resource-aware computing, intelligent reasoning, multimodal interaction, and real-world application constraints are co-designed for personalized assistance in the physical world. To systematically study this perspective, we organize the survey around four connected dimensions. First, we examine the hardware foundation that bounds sensing, computation, feedback delivery, and sustained deployment. Second, we study wearable intelligence, where egocentric signals are transformed into perceptual, contextual, and agentic capabilities. Third, we discuss interaction design, through which users request, receive, correct, and regulate assistance during ongoing activity. Fourth, we analyze application scenarios across healthcare, accessibility, situated learning, daily life assistance, cultural tourism, and industrial support, showing how domain requirements reshape system design and evaluation. We further identify five cross-cutting research challenges for future AI smart glasses: next-generation hardware, trustworthy egocentric intelligence, lifelong personalized memory, proactive intelligence, and embodied foundation models. By centering smart glasses as wearable-intelligence platforms, this survey provides a unified framework for organizing technologies, applications, and open challenges in this emerging area.
Chinese Translation
人工智能(AI)的最新进展正在将智能眼镜从自我中心的采集与显示设备重塑为可穿戴智能平台。智能眼镜日益成为在严格的形态约束下将第一人称观察与实时辅助相连接的可穿戴AI系统。我们以“AI智能眼镜”为视角来审视这一转变,并将其定义为一个系统级概念:其中自我中心感知、资源感知计算、智能推理、多模态交互以及真实世界应用约束被协同设计,以在物理世界中实现个性化辅助。为系统性地研究这一视角,本文围绕四个相互关联的维度组织综述。首先,我们考察界定感知、计算、反馈交付与持续部署的硬件基础。其次,我们研究可穿戴智能,即自我中心信号如何被转化为感知、情境与智能体能力。第三,我们讨论交互设计,即用户在持续活动中如何请求、接收、纠正和调节辅助。第四,我们分析医疗健康、无障碍辅助、情境化学习、日常生活辅助、文化旅游和工业支持等应用场景,展示领域需求如何重塑系统设计与评估。我们进一步识别出未来AI智能眼镜的五大交叉研究挑战:下一代硬件、可信赖的自我中心智能、终身个性化记忆、主动式智能以及具身基础模型。通过将智能眼镜定位为可穿戴智能平台,本综述为组织这一新兴领域的技术、应用与开放挑战提供了统一框架。
cs.CV / 33 / 2609.19812

Absence is Presence: Understanding Visual Scene Negative Events Under Safety Cognitive Constraint

缺席即存在:安全认知约束下的视觉场景负事件理解
Jiang, Zhiyun, Wang, Hanyong, Liang, Binbin, Xie, Yu, Yang, Menglong, Li, Wei
Abstract
Traditional scene understanding focuses on affirmative information objectively present in images. However, in safety-critical domains, comprehending key information that should exist but is actually absent is vital for risk mitigation. To bridge this gap, we focus on visual scene negative captioning with safety as the cognitive constraint. The core challenge is to convert physical absence into semantic negative events. Existing vision-language models (VLMs) struggle with this process because affirmation bias suppresses negative reasoning, while limited mental filling capability and representation bias further hinder the inference of absent information. To address these challenges, we propose a negative captioning framework based on counterfactual reconstruction and contrastive decoding (CRCD). Inspired by human cognition, CRCD reformulates the task as counterfactual latent change captioning to bypass affirmation bias. It contrasts a synthesized safe expectation with reality to identify semantic omissions. To address limited mental filling, we design a dual-branch counterfactual reconstruction architecture. The amodal completion branch restores defective objects, while the functional association branch infers completely absent safety objects. Concurrently, a multi-condition representation learning mechanism is integrated to mitigate representation bias by projecting universal features onto predefined safety criteria subspaces, thereby capturing information across more dimensions. By decoding feature-level semantic residuals between the reconstructed scene prototype and raw input, CRCD bounds the non-existence search space and activates the decoder's negative logic. Extensive experiments validate the effectiveness of CRCD, establishing a high-performance baseline for this pioneering task.
Chinese Translation
传统的场景理解关注图像中客观存在的肯定性信息。然而,在安全关键领域中,理解那些应当存在但实际上缺失的关键信息对于风险缓解至关重要。为弥合这一差距,我们聚焦于以安全为认知约束的视觉场景负描述任务。其核心挑战在于将物理上的缺席转化为语义上的负事件。现有的视觉语言模型(VLMs)难以完成这一过程,原因在于肯定性偏差抑制了负向推理,而有限的心理填充能力和表征偏差进一步阻碍了对缺失信息的推断。为应对这些挑战,我们提出了一种基于反事实重构与对比解码的负描述框架(CRCD)。受人类认知启发,CRCD将该任务重新表述为反事实潜在变化描述,以绕过肯定性偏差。它通过对比合成的安全期望与现实来识别语义上的遗漏。为解决心理填充能力有限的问题,我们设计了一种双分支反事实重构架构:失觉补全分支用于恢复有缺陷的物体,而功能关联分支用于推断完全缺失的安全物体。同时,我们集成了多条件表征学习机制,通过将通用特征投影到预定义的安全标准子空间来缓解表征偏差,从而从更多维度捕获信息。通过解码重构的场景原型与原始输入之间的特征级语义残差,CRCD界定了非存在搜索空间并激活了解码器的负向逻辑。大量实验验证了CRCD的有效性,为这一开创性任务建立了高性能基线。
cs.CV / 34 / 2609.19815

SnapPhysics: A Physics-Aware Scene Graph from a Single View for Interactive Mixed Reality Scenes

SnapPhysics:一种面向交互式混合现实场景的物理感知单视图场景图
Kang, Suji, Kim, Seok-Young, Kim, Young Bin, Ha, Taewook, Schmalstieg, Dieter, Mori, Shohei, Woo, Woontack
Abstract
We propose SnapPhysics, a training-free framework that reconstructs 3D objects and estimates their physical properties such as mass, friction, and center of gravity from a single image. For physically coherent interactions in mixed reality (MR), such properties are as important as geometry. Prior approaches infer them by analyzing object dynamics in video, which is computationally costly, or by querying vision-language models (VLMs) on single images, which lacks geometric grounding and inter-object relationships. We address these limitations by combining instance-level 3D reconstruction and spatial alignment with a physics-aware scene graph that encodes these relationships and per-object metric geometry as structured context for VLM-based property reasoning. Experiments on 3D-FRONT show that SnapPhysics improves scene-level F-Score by 18.6% over the best learning-based method, and on real captured scenes with ground-truth mass, it reduces the mean absolute log difference error (mALDE) by up to 20.5% and improves log-scale correlation ($r^2_{\mathrm{ls}}$) by up to 19.6% over VLM-only estimation. SnapPhysics enables physically interactive MR experiences without manual parameter tuning. Project page: https://snapphysics-ismar2026.github.io/.
Chinese Translation
我们提出了SnapPhysics,一个无需训练的框架,能够从单张图像重建三维物体并估计其物理属性,如质量、摩擦力和重心。为了在混合现实中实现物理上连贯的交互,这些属性与几何信息同等重要。已有方法通过分析视频中物体的动态来推断这些属性,但计算代价高昂;或者通过在单张图像上查询视觉语言模型(VLM)来推断,但缺乏几何基础和物体间关系。我们通过将实例级三维重建与空间对齐相结合,并构建一个物理感知的场景图来编码这些关系以及每个物体的度量几何信息,作为VLM属性推理的结构化上下文,从而解决上述局限。在3D-FRONT数据集上的实验表明,SnapPhysics相较于最优的基于学习的方法将场景级F-Score提升了18.6%;在具有真实质量标注的真实采集场景上,相较于仅使用VLM的估计,其平均绝对对数差误差(mALDE)最多降低20.5%,对数尺度相关性($r^2_{\mathrm{ls}}$)最多提升19.6%。SnapPhysics无需手动参数调节即可实现物理交互式的MR体验。项目页面:https://snapphysics-ismar2026.github.io/。
cs.CV / 35 / 2609.19840

KoUniTalk: A Lightweight Articulation-Centered Korean-English 3D Talking Face Benchmark

KoUniTalk:一个轻量级以发音为中心的韩英双语3D说话人脸基准数据集
Chung, Hyunjung, Park, Unsang
Abstract
High-quality 3D talking face datasets remain largely English- centric, and Korean 3D facial motion data are difficult to combine with standard English benchmarks because of differences in mesh topology, spatial scale, coordinate system, and temporal sampling. We present KoUniTalk, a lightweight articulation-centered Korean-English 3D talk- ing face benchmark that retargets VOCASET and the released Korean speech-based 3D talking face data to a shared mesh topology using de- formation transfer. Rather than proposing a new deformation-transfer algorithm or a full-head identity-preserving avatar dataset, KoUniTalk provides an identity-neutral canonical output space for controlled speech- driven facial articulation training and evaluation across English and Ko- rean. The unified template contains 1,176 vertices and focuses on the mouth and adjacent lower- and mid-face regions, reducing the output dimensionality from 15,069 and 72,147 dimensions to 3,528 dimensions, corresponding to 4.27-fold and 20.45-fold reductions compared with VO- CASET/FLAME and the original Korean mesh, respectively. To exam- ine whether retargeting preserves speech-relevant motion, we evaluate semantic mouth-landmark trajectories, including mouth opening, mouth width, aperture ratio, and mouth-opening dynamics. Since the official test set of the Korean dataset is not publicly released, we additionally define a subject-disjoint Korean benchmark split. The processed matched benchmark contains 22 speakers, 4,978 sequences, and 642,781 frames, enabling Korean-English cross-dataset evaluation of speech-driven 3D fa- cial animation models in a single compact articulation-template space. Source-reported inventory counts are listed separately from these pro- cessed counts
Chinese Translation
高质量的3D说话人脸数据集在很大程度上仍以英语为中心,而韩语3D面部运动数据由于网格拓扑、空间尺度、坐标系和时间采样等方面的差异,难以与标准英语基准数据集结合使用。我们提出了KoUniTalk,一个轻量级的以发音为中心的韩英双语3D说话人脸基准数据集,它利用变形迁移(deformation transfer)技术将VOCASET和已发布的基于韩语语音的3D说话人脸数据重定向到共享的网格拓扑上。KoUniTalk并不提出新的变形迁移算法或完整的保身份头模数据集,而是为英语和韩语的受控语音驱动面部发音训练与评估提供一个身份无关的规范化输出空间。该统一模板包含1,176个顶点,专注于嘴部及相邻的下面部和面中部区域,将输出维度从15,069维和72,147维分别降至3,528维,相比VOCASET/FLAME和原始韩语网格分别实现了4.27倍和20.45倍的降维。为检验重定向是否保留了与语音相关的运动,我们评估了语义嘴部关键点轨迹,包括嘴部张开程度、嘴部宽度、开口比例和开口动态。由于韩语数据集的官方测试集未公开发布,我们额外定义了一个按说话人不重叠划分的韩语基准分割。处理后的匹配基准数据集包含22名说话人、4,978个序列和642,781帧,使得语音驱动的3D面部动画模型能够在单一紧凑的发音模板空间中进行韩英跨数据集评估。数据来源报告的原始统计数量与这些处理后数量分别列出。
cs.CV / 36 / 2609.19853

PACE: Precise AI Cinematic Expression: A Typed Specification for Script-Grounded Previsualization and Geometric Conformance

PACE:精确的AI电影化表达——一种面向剧本锚定预可视化和几何一致性的类型化规范
Duan, Bing, Guo, Qiang, Li, Linpu, Mao, Zhijian, Zhu, Min, Ren, Zhirui, Yan, Yiwei, Chu, Xi, Li, Xiaoding
Abstract
Between a screenplay and a film sits a planning problem that is spatial first: who stands where, and what a camera sees from where it stands. An image diffusion model asked for a shot in free text settles that plan by its own defaults. We present PACE (Precise AI Cinematic Expression), a typed representation for the plan: the screenplay evidence, the characters, props and locations it needs, where each subject stands, and what the camera does. A value is written once at the level it belongs to (script, scene, shot or panel) and inherited below it. A compiler turns the result into both the prompt sent to the diffusion model and a 3D scene built in metres, and a camera solver places the camera so that the declared framing is the framing built. Where a declared value becomes geometry, PACE measures, field by field, how far the compiled camera and the staged render sit from the declaration, rather than asking a model to judge. On the 11-scene Automatic Drive screenplay, every staged single-subject panel places its subject within 1.2% of frame width of its declared position; with two or three subjects one camera pose cannot satisfy every position, and the residual is reported rather than absorbed. On 204 external director-storyboard shots, delivered head height is 1.906 times the staged target from the director's words, 1.733 from the compiled prompt, and 0.955 with the greybox control; the condition that holds framing best draws the described action least. Declaring the pose on 30 shots raises the action drawn from 58.9% to 74.4% without moving the framing. Transitions, fitted motion and human review of the generated panels remain open. Code: https://github.com/StudioPiLabs/pace-core
Chinese Translation
在剧本与影片之间存在一个以空间为首要问题的规划问题:谁站在哪里,以及摄像机从其所处位置能看到什么。当仅以自由文本向图像扩散模型请求一个镜头时,该规划由模型自身的默认设置来确定。我们提出PACE(Precise AI Cinematic Expression,精确AI电影化表达),一种对该规划的类型化表示:包括剧本依据、其所需的角色、道具与场景、每个主体的站位,以及摄像机的运动。每个值只在其所属层级(剧本、场景、镜头或分镜画面)上书写一次,并在下层自动继承。一个编译器将该表示转化为发送给扩散模型的提示词,同时构建一个以米为单位的3D场景,并由摄像机求解器放置摄像机,使所声明的取景即为实际构建的取景。在声明值转化为几何的地方,PACE逐字段地度量编译后的摄像机与舞台化渲染结果偏离声明的程度,而不是让模型去评判。在包含11个场景的Automatic Drive剧本上,每个舞台化的单主体分镜画面中,主体位置与声明位置的偏差均在画面宽度的1.2%以内;当有两到三个主体时,单个摄像机位无法满足所有位置要求,此时报告残差而非将其吸收。在204个外部导演分镜镜头上,交付的人物头部高度是舞台化目标的1.906倍(依据导演文字)、1.733倍(依据编译后的提示词),而在灰盒控制下为0.955倍;最能保持取景的条件反而对所描述动作的呈现最少。在30个镜头上声明姿态,可将所呈现动作的比例从58.9%提升至74.4%,且不改变取景。转场、拟合运动以及对生成画面的真人审阅仍为待解决的问题。代码:https://github.com/StudioPiLabs/pace-core
cs.CV / 37 / 2609.19867

Socialized UAV Cross-Task Learning: Towards Cross-Granularity Collaboration through Hierarchical Interaction

社交化无人机跨任务学习:通过层级交互实现跨粒度协作
Yao, Xinjie, Zhao, Ruipu, Zhu, Yunqi, Fan, Zhihe, Guo, Zhoupeng, Li, Weihao, Wang, Zhen, Wang, Qilong, Zhu, Pengfei
Abstract
Joint learning across heterogeneous tasks is often treated as task coupling through feature sharing, distillation, or auxiliary supervision. However, in cross-task learning, mismatched representational and supervisory granularities make such coupling prone to interference, teacher bias, or unidirectional collapse. We argue that cross-granularity learning is fundamentally a problem of hierarchical interaction regulation rather than simple task coupling. This issue is particularly evident in UAV perception, where visual shifts and detection--segmentation objectives naturally form coarse- and fine-grained knowledge sources. To systematically study this problem, we introduce CrossUAV, a UAV benchmark for joint object detection and instance segmentation that provides a unified evaluation platform for cross-granularity task collaboration. To address these challenges, we propose Cross-Granularity Socialized Collaboration (CGSC), a progressive and adaptive framework that regulates when, where, and how tasks exchange information across network hierarchies. CGSC progressively activates cross-task interactions and adaptively adjusts the strength according to task contribution, suppressing harmful interference while exploiting complementary coarse- and fine-grained structures. Extensive experiments demonstrate consistent improvements on both tasks, validating hierarchical dynamic interaction as an effective mechanism for cross-granularity collaboration.
Chinese Translation
异构任务之间的联合学习通常被视为通过特征共享、蒸馏或辅助监督实现的任务耦合。然而,在跨任务学习中,表示与监督粒度的不匹配使得这种耦合容易受到干扰、教师偏差或单向坍塌的影响。我们认为,跨粒度学习本质上是一个层级交互调控问题,而非简单的任务耦合。这一问题在无人机感知中尤为明显,其中视觉偏移以及检测与分割目标天然形成了粗粒度和细粒度的知识来源。为了系统地研究这一问题,我们提出了CrossUAV,一个面向联合目标检测与实例分割的无人机基准数据集,为跨粒度任务协作提供了统一的评估平台。为应对上述挑战,我们提出了跨粒度社交化协作框架(Cross-Granularity Socialized Collaboration, CGSC),这是一个渐进式自适应框架,用于调控任务在何时、何地以及如何跨网络层级交换信息。CGSC渐进地激活跨任务交互,并根据任务贡献自适应地调整交互强度,在抑制有害干扰的同时利用互补的粗粒度与细粒度结构。大量实验表明,该框架在两个任务上均取得了一致的性能提升,验证了层级动态交互作为跨粒度协作有效机制的价值。
cs.CV / 38 / 2609.19872

PART: Learning 3D Part Assembly and Retrieval with Transformers

PART:基于Transformer的3D部件组装与检索学习方法
Bao, Ruchao, Wu, Wenzheng, Xiang, Chucheng, Liu, Zhongyuan, Liu, Yuan, Dong, Jinxin, Liu, Ligang, Wang, Ziqi
Abstract
3D assembly is fundamental to modern manufacturing and digital content creation. In this paper, we present PART, a unified transformer-based framework for 3D part retrieval and assembly: given a target shape and a part library, PART automatically selects the appropriate parts and predicts their 6-DoF poses to reconstruct the target. While prior work has achieved impressive progress on assembling a pre-defined set of parts, this more practical retrieval-based setting remains largely unexplored. The task faces three key challenges: (i) a combinatorially explosive search space that grows exponentially with library size; (ii) variable-length outputs, as different targets require different numbers of parts; and (iii) continuous 6-DoF pose estimation for part assembly. To address these, we formulate retrieval and assembly as a set prediction problem and design a novel transformer-based framework that retrieves parts and regresses their poses with variable-length output. Additionally, we exploit the duality between part pose estimation and target segmentation through joint training and a novel segmentation-enhanced optimization module. Finally, We curate a large-scale dataset of 80K+ shapes, and the results show that PART generalizes to scene layouts, image targets, and real-world scans. Project Page: https://iambrc.github.io/PART-project-page/.
Chinese Translation
3D组装是现代制造和数字内容创作的基础。在本文中,我们提出了PART,一个统一的基于Transformer的3D部件检索与组装框架:给定目标形状和部件库,PART自动选择合适的部件并预测其6自由度(6-DoF)位姿以重建目标。尽管先前的工作在组装预定义部件集合方面取得了令人瞩目的进展,但这一更实用的基于检索的设置在很大程度上仍未被探索。该任务面临三个关键挑战:(i) 随部件库规模呈指数级增长的组合爆炸式搜索空间;(ii) 可变长度输出,因为不同目标需要不同数量的部件;(iii) 部件组装的连续6自由度位姿估计。为解决这些问题,我们将检索与组装形式化为一个集合预测问题,并设计了一种新颖的基于Transformer的框架,以可变长度输出检索部件并回归其位姿。此外,我们通过联合训练和一个新颖的分割增强优化模块,利用部件位姿估计与目标分割之间的对偶性。最后,我们构建了一个包含8万多个形状的大规模数据集,结果表明PART能够泛化到场景布局、图像目标和真实世界扫描。项目页面:https://iambrc.github.io/PART-project-page/.
cs.CV / 39 / 2609.19875

BINDER: A Latent Variable Model for Probabilistic Medical Image Registration

BINDER:用于概率性医学图像配准的潜变量模型
Cerri, Stefano, Hassankhani, Amirhossein, Balbastre, Yaël, Van Leemput, Koen
Abstract
We propose a new probabilistic model for general-purpose medical image registration that builds upon the mutual information registration criterion. It centers around a spatial interpolation technique that assumes latent voxel-wise correspondences between the images being registered. By exploiting these latent variables, we derive dedicated optimization and MCMC sampling techniques that only involve closed-form iterative updates. When applied to nonlinear registration, an efficient demons-like optimization algorithm is obtained that shows robust out-of-the-box performance across a variety of monomodal and multimodal registration tasks. We also demonstrate a corresponding sampler that can quantify, for the first time, uncertainty in multimodal registration scenarios with very high-dimensional 3D deformations. Our code, which we call BINDER (Bayesian INference for DEformable Registration), is freely available at https://github.com/ste93ste/BINDER.
Chinese Translation
我们提出了一种新的通用医学图像配准概率模型,该模型建立在互信息配准准则之上。其核心是一种空间插值技术,该技术假设待配准图像之间存在潜在的体素级对应关系。通过利用这些潜变量,我们推导出了专门的优化方法和MCMC采样技术,它们仅涉及闭式迭代更新。当应用于非线性配准时,可得到一种类似demons的高效优化算法,在各种单模态和多模态配准任务中均表现出开箱即用的鲁棒性能。我们还展示了一种相应的采样器,首次能够在具有非常高维3D形变的多模态配准场景中量化不确定性。我们的代码称为BINDER(Bayesian INference for DEformable Registration,可变形配准的贝叶斯推断),可在 https://github.com/ste93ste/BINDER 免费获取。
cs.CV / 40 / 2609.19876

SlugTrails: An Egocentric Benchmark for Floor Plan Localization in Large Buildings

SlugTrails:面向大型建筑平面图定位的第一视角基准数据集
Cheng, Yunqian, Manduchi, Roberto
Abstract
Floor-plan-based indoor visual localization enables infrastructure-free positioning, but most methods are developed and evaluated in small residential environments unlike the large public buildings of real deployment. We introduce SlugTrails, a floor plan localization benchmark for large indoor spaces under realistic egocentric sensing: $30$ Hz Aria glasses recordings across three campus buildings and six floors ($22089$ m$^2$ of floor plan outline), CAD-derived floor plans with semantic classes and circulation space masks, and trajectories aligned into the floor plan frame using laser-surveyed anchors. One protocol covers three practical ways of gathering geometry under a limited field of view -- a single walking frame, a stationary multi-view sweep, and a walking stream with odometry -- so methods designed for different regimes are compared on the same buildings and ground truth. Evaluating five representative geometric and learned systems under their native sensing configurations, we find that stock checkpoints (official released weights) are near zero on SlugTrails (at most $0.004$ R@1m30$^{\circ}$ on walking single frames), while fine-tuning on SlugTrails improves every trainable family on all three tasks (e.g., F$^3$Loc $0.0 \rightarrow 0.141$ single-frame and $0.03 \rightarrow 0.66$ sequential), with gains compounding as observations accumulate. The same fine-tuned weights also improve cross-dataset generalization on LaMAR with no LaMAR training (sequential R@1m $0.048 \rightarrow 0.143$ for F$^3$Loc and $0.063 \rightarrow 0.127$ for UnLoc), whereas train-from-scratch on SlugTrails alone stays far below fine-tuning from stock weights -- evidence that floor plan localization is currently limited by indoor data rather than by architecture. We release the dataset, protocols, and tools at https://github.com/Head-inthe-Cloud/SlugTrails.
Chinese Translation
基于平面图的室内视觉定位能够实现无需基础设施的定位,但大多数方法都是在小型住宅环境中开发和评估的,这与实际部署中的大型公共建筑截然不同。我们提出了SlugTrails,一个面向大型室内空间、基于现实第一视角(egocentric)感知的平面图定位基准:包含跨三栋校园建筑、六个楼层(平面图轮廓面积达22089平方米)的30 Hz Aria眼镜录像,由CAD导出并带有语义类别和流通空间掩膜的平面图,以及通过激光测量锚点对齐到平面图坐标系中的轨迹。我们的评测协议涵盖了在有限视场下获取几何信息的三种实用方式——单步行帧、静止多视角扫描以及带里程计的步行视频流——从而使针对不同模式设计的方法能够在相同的建筑和真值(ground truth)上进行比较。我们在五种代表性几何方法和学习方法的原始感知配置下进行评估,发现官方预训练权重在SlugTrails上表现接近于零(步行单帧上的R@1m30°至多仅0.004),而在SlugTrails上微调可以提升所有三个任务中所有可训练方法族的性能(例如,F³Loc单帧性能从0.0提升至0.141,序列性能从0.03提升至0.66),且随着观测的累积,性能增益不断叠加。相同微调后的权重在无需任何LaMAR训练的情况下,也能提升在LaMAR数据集上的跨数据集泛化能力(F³Loc的序列R@1m从0.048提升至0.143,UnLoc从0.063提升至0.127),而仅从零开始在SlugTrails上训练的效果仍远低于基于预训练权重的微调——这表明平面图定位目前主要受限于室内数据的匮乏,而非模型架构。我们在 https://github.com/Head-inthe-Cloud/SlugTrails 上发布了该数据集、评测协议和工具。
cs.CV / 41 / 2609.19881

BinoGen: Scaling egocentric binocular data for embodied visual perception and learning

BinoGen:面向具身视觉感知与学习的自我中心双目数据扩展
Li, Chunpeng, Li, Ya-tang
Abstract
Embodied visual perception relies on temporally coherent visual experience accumulated through continuous engagement with the environment. However, collecting large-scale egocentric binocular observations together with dense annotations remains costly and difficult. Moreover, visual experience is shaped not only by the environment but also by the embodiment of the observer, including viewing height, field of view, binocular geometry, and motion through the scene. To address these challenges, we present BinoGen, an automated framework for generating large-scale, embodiment-aware egocentric binocular visual experiences in indoor environments. BinoGen jointly models environmental and observer variation through generative scene synthesis, probabilistic object instantiation, appearance randomization, stochastic trajectory generation, and configurable binocular camera setups. The framework produces synchronized binocular videos together with dense multimodal supervision, including depth maps, optical flow, surface normals, semantic maps, object coordinates, and camera poses. Using BinoGen, we construct a dataset comprising more than 20 million annotated images for supervised learning. We demonstrate two complementary utilities of BinoGen. First, incorporating BinoGen data consistently improves real-world visual perception, including depth estimation, object detection, and video object tracking. Second, paired human-inspired and mouse-inspired observations from the same environments enable controlled investigation of how observer embodiment affects perceptual learning. Embodiment-specific adaptation substantially improves performance, while joint training enables a single model to perform competitively across both embodiments. Together, these results demonstrate that large-scale, controllable visual experience can improve embodied perception...
Chinese Translation
具身视觉感知依赖于通过与环境的持续互动所积累的时间上连贯的视觉经验。然而,收集大规模的自我中心双目观测数据及稠密标注仍然成本高昂且困难重重。此外,视觉经验不仅受环境影响,还受观察者身体形态的影响,包括观察高度、视场范围、双目几何结构以及在场景中的运动方式。为应对这些挑战,我们提出了 BinoGen,一个用于在室内环境中生成大规模、具身感知的自我中心双目视觉经验的自动化框架。BinoGen 通过生成式场景合成、概率化物体实例化、外观随机化、随机轨迹生成以及可配置的双目相机设置,联合建模环境与观察者的变化。该框架能够生成同步的双目视频以及稠密的多模态监督信号,包括深度图、光流、表面法线、语义图、物体坐标和相机位姿。利用 BinoGen,我们构建了一个包含超过 2000 万张标注图像的数据集,用于监督学习。我们展示了 BinoGen 的两种互补用途:第一,引入 BinoGen 数据能够持续提升真实世界视觉感知任务的性能,包括深度估计、物体检测和视频目标跟踪;第二,来自相同环境的仿人视角与仿鼠视角配对观测,使得可以控制变量地研究观察者身体形态如何影响感知学习。针对特定身体形态的适配能显著提升性能,而联合训练则使单一模型在两种身体形态下均能取得有竞争力的表现。这些结果共同表明,大规模、可控的视觉经验能够提升具身感知性能……
cs.CV / 42 / 2609.19907

GS-PI: An Optimization-Decoupled Appearance Decomposition Approach for Generating PBR Gaussian Assets

GS-PI:一种用于生成PBR高斯资产的优化解耦外观分解方法
Xu, Jieting, Xie, Rengan, Huang, Zijian, Jin, Zehui, Wang, Rui, Huo, Yuchi
Abstract
Gaussian Splatting (GS) excels at novel-view synthesis but encodes baked-in radiance, tightly entangling illumination with geometry and preventing seamless integration into physically based rendering (PBR) pipelines. Existing inverse-rendering methods attempt to disentangle materials via joint optimization, but often suffer from competing objectives that cause severe ambiguities and residual lighting artifacts. To overcome this, we present GS-PI, a novel optimization-decoupled framework that casts PBR material generation as a geometry-conditioned diffusion process on 3D point clouds. By operating directly in the 3D domain, our method inherently guarantees multi-view consistency, sidestepping the severe pixel correspondence issues that challenge 2D diffusion approaches. We introduce a multi-scale cross-view conditioning mechanism that integrates three complementary components: a global semantic prior, source-anchored photometric cues, and an absolute spatial learned view-direction conditioning signal. This design efficiently compresses complex multi-view evidence, mitigating cross-view projection misalignment and successfully preventing specular highlights from baking into intrinsic colors. By extracting a point cloud from a pre-trained Gaussian model, predicting PBR attributes via conditional diffusion, and distilling them back through differentiable rasterisation, we yield a fully relightable PBR-GS asset. GS-PI outperforms recent inverse-rendering baselines while replacing per-scene joint illumination/BRDF optimization with a learned diffusion pass followed by a short target-driven distillation, without requiring proxy meshes.
Chinese Translation
三维高斯泼溅(Gaussian Splatting, GS)在新视角合成方面表现出色,但其编码的是烘焙进模型中的辐射值,导致光照与几何紧密耦合,无法无缝集成到基于物理的渲染(PBR)流程中。现有的逆渲染方法试图通过联合优化来解耦材质,但常常因目标相互竞争而产生严重的歧义性和残留的光照伪影。为克服这一问题,我们提出了GS-PI,这是一种新颖的优化解耦框架,它将PBR材质生成转化为3D点云上的几何条件扩散过程。通过直接在3D域中操作,我们的方法天然保证了多视角一致性,规避了困扰2D扩散方法的严重像素对应问题。我们引入了一种多尺度跨视角条件机制,它整合了三个互补的组件:全局语义先验、以源为锚点的光度线索,以及绝对空间的学习型视角方向条件信号。这一设计高效地压缩了复杂的多视角证据,缓解了跨视角投影错位问题,并成功防止高光被烘焙进固有颜色中。通过从预训练的高斯模型中提取点云、经条件扩散预测PBR属性,再通过可微光栅化将其蒸馏回模型,我们得到了完全可重光照的PBR-GS资产。GS-PI优于近期的逆渲染基线方法,同时以一次学习型扩散过程加上短时的目标驱动蒸馏,取代了逐场景的联合光照/BRDF优化,且无需代理网格。
cs.CV / 43 / 2609.19911

CitySTAR: Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding

CitySTAR:面向开放词汇城市三维定位的结构化与拓扑感知推理
Zhang, Shuai, Hou, Hongye, Liu, Qinghe, Li, Zhuoxiao, Wu, Dongli, Ou, Jing, Liu, Yuan, Zhao, Wufan
Abstract
3D grounding aims to localize target entities in complex scenes from natural language and plays a fundamental role in embodied perception and spatial reasoning. However, existing approaches mostly rely on feature similarity or direct matching, making it difficult to connect natural-language intent with the implicit semantic and geometric structures hidden in billion-scale urban point clouds. We reformulate city-scale 3D grounding as structured constraint reasoning, where description semantics are organized into computable cross-modal constraints over open-vocabulary 3D entities, attributes, and spatial relations. We present CitySTAR, a training-free framework for reasoning-driven urban 3D grounding. CitySTAR lifts raw billion-scale urban point clouds into a query-ready scene graph of open-vocabulary 3D instances, with CodeLLM-driven tools supplying multimodal evidence for node attributes and 3D spatial relations. It then models target-context topology with paired hypergraphs and performs bidirectional topology verification for structural disambiguation. Finally, a Reflective Cross-modal Grounding module integrates topology consistency and candidate-centered 2D visual evidence to make decisions over a metric-aware 3D context graph. To further support this setting, we introduce CitySTAR-3D, an enhanced benchmark that improves semantic coverage, instance completeness, bounding-box fidelity, and spatial-relation complexity in city-scale 3D grounding. Extensive experiments show that CitySTAR consistently improves open-world urban 3D grounding while maintaining strong interpretability and generalization.
Chinese Translation
三维定位旨在根据自然语言在复杂场景中定位目标实体,是具身感知与空间推理的基础任务。然而,现有方法大多依赖特征相似度或直接匹配,难以将自然语言意图与隐藏在十亿级规模城市点云中的隐式语义和几何结构相连接。我们将城市尺度的三维定位重新表述为结构化约束推理问题,其中描述语义被组织为可在开放词汇的三维实体、属性和空间关系上进行计算化的跨模态约束。我们提出CitySTAR,一个无需训练的、由推理驱动的城市三维定位框架。CitySTAR将原始的十亿级规模城市点云提升为可直接查询的开放词汇三维实例场景图,并由CodeLLM驱动的工具为节点属性和三维空间关系提供多模态证据。随后,该方法利用成对超图建模目标-上下文拓扑,并执行双向拓扑验证以实现结构性消歧。最后,反思式跨模态定位模块整合拓扑一致性与以候选对象为中心的二维视觉证据,在一个度量感知的三维上下文图上做出决策。为进一步支持该任务设定,我们引入了CitySTAR-3D基准数据集,提升了城市尺度三维定位中的语义覆盖度、实例完整性、边界框保真度和空间关系复杂度。大量实验表明,CitySTAR持续提升开放世界城市三维定位性能,同时保持较强的可解释性与泛化能力。
cs.CV / 44 / 2609.19927

DirtyMoCap: Robust Motion Capture from Unconstrained Markers

DirtyMoCap:基于非受限标记点的鲁棒动作捕捉
Wang, Long, Zhao, Shuting, Yan, Shen, Yu, Siyuan, Li, Xiaoben, Cai, Zeyu, Hou, Yumeng, Xiu, Yuliang
Abstract
Optical motion capture delivers high-fidelity human motion, but its reliance on strict marker layouts and clean trajectories severely limits its real-world applicability. In practice, tracking systems frequently output unconstrained markers: sparse, noisy, and unordered point clouds with unknown or varying configurations. To bridge the gap between corrupted raw markers and parametric human models, we introduce DirtyMoCap, a robust, marker-layout-free framework. Our core insight is to map unordered marker observations to a fixed set of "proxy anchors" comprising skeletal joints and body surface points, which serve as a stable intermediate representation. We first initialize and track these anchors over long sequences using a recurrent sliding-window architecture. Then, a custom differentiable Gauss-Newton solver fits the SMPL-H model to the tracked anchors to recover full-body pose, translation, and shape. By explicitly deriving geometric residuals, our solver learns adaptive observation confidence, smoothness, and prior weights end-to-end, adapting dynamically to the reliability of the input data. Extensive experiments on diverse, noisy marker configurations demonstrate that DirtyMoCap successfully generalizes across arbitrary layouts using only a single trained model. It consistently outperforms state-of-the-art configuration-specific baselines in both joint and vertex reconstruction accuracy, while our custom CUDA solver achieves up to a 100x speedup over standard PyTorch implementations. We further apply DirtyMoCap to heterogeneous raw optical MoCap recordings of traditional Chinese martial arts, yielding a Kung Fu motion dataset of temporally coherent SMPL-H reconstructions. Code and data are available at https://wanglongzju.github.io/DirtyMoCap-Project-Page.
Chinese Translation
光学动作捕捉能够提供高保真的人体运动数据,但其对严格标记点布局和干净轨迹的依赖严重限制了其在现实场景中的适用性。在实际应用中,追踪系统经常输出非受限的标记点:即稀疏、含噪且无序的点云,其配置未知或动态变化。为了在受损的原始标记点与参数化人体模型之间架起桥梁,我们提出了DirtyMoCap——一个鲁棒的、无需标记点布局约束的框架。我们的核心思想是将无序的标记点观测映射到一组固定的"代理锚点"上,包括骨骼关节点和身体表面点,这些锚点充当稳定的中间表示。我们首先使用循环滑动窗口架构在长序列上初始化并追踪这些锚点。然后,一个定制的可微高斯-牛顿(Gauss-Newton)求解器将SMPL-H模型拟合到追踪到的锚点上,以恢复全身姿态、平移和形状。通过显式推导几何残差,我们的求解器能够端到端地学习自适应的观测置信度、平滑度和先验权重,从而动态适应输入数据的可靠性。在多样化、含噪标记点配置上的大量实验表明,DirtyMoCap仅使用单个训练好的模型即可成功泛化到任意布局。在关节和顶点重建精度方面,它始终优于最先进的特定配置基线方法,同时我们的定制CUDA求解器相比标准PyTorch实现实现了高达100倍的加速。我们进一步将DirtyMoCap应用于传统中华武术的异构原始光学动作捕捉记录,生成了一个包含时间连贯的SMPL-H重建结果的功夫运动数据集。代码和数据可在https://wanglongzju.github.io/DirtyMoCap-Project-Page获取。
cs.CV / 45 / 2609.19964

Enhanced Knowledge Distillation for Detection Transformer via Teacher Prediction Refinement

基于教师预测精炼的检测Transformer增强知识蒸馏
Xing, Yitong, Cheng, Yuhao, Li, Yanping, Yan, Yichao
Abstract
Detection Transformers (DETRs) achieve strong performance in object detection but remain challenging to deploy on edge devices due to their high computational cost. Existing DETR distillation methods mainly focus on aligning distillation points, while largely overlooking the quality of the teacher's supervision itself. We observe that due to stage-wise non-monotonic prediction behavior in DETRs, well-localized or correctly classified predictions from earlier stages may degrade in later ones, and some negative predictions become increasingly overconfident. As a result, relying solely on the current stage's predictions yields inaccurate and inconsistent supervision. To address this issue, we propose Teacher Prediction Refinement Distillation (TPRD), a plug-and-play module that refines teacher predictions before distillation by exploiting stage-wise prediction information. TPRD improves supervision quality through Positive Prediction Correction (PPC), which corrects degraded positive predictions by restoring more accurate ones from earlier stages, ensuring reliable localization and classification signals, and Negative Prediction Suppression (NPS) suppresses the influence of overconfident negatives, preventing them from providing misleading supervision to the student. To preserve informative dark knowledge, we further introduce Maximum Dark Knowledge Preservation (MDKP), which selectively refines target-class logits while retaining non-target relations. Extensive experiments on MS COCO and PASCAL VOC demonstrate the effectiveness and robustness of the proposed method. Our code is available at https://github.com/xingyitong1/TPRD.
Chinese Translation
检测Transformer(DETR)在目标检测中取得了优异的性能,但由于其高昂的计算成本,在边缘设备上的部署仍然充满挑战。现有的DETR蒸馏方法主要关注对齐蒸馏点,而在很大程度上忽视了教师监督信号本身的质量。我们观察到,由于DETR中存在分阶段的非单调预测行为,来自较早阶段的定位良好或分类正确的预测可能在较晚阶段出现退化,且一些负样本预测变得过度自信。因此,仅依赖当前阶段的预测会产生不准确且不一致的监督信号。为解决这一问题,我们提出了教师预测精炼蒸馏(Teacher Prediction Refinement Distillation, TPRD),这是一个即插即用模块,通过利用分阶段预测信息在蒸馏前对教师预测进行精炼。TPRD通过正预测校正(Positive Prediction Correction, PPC)和负预测抑制(Negative Prediction Suppression, NPS)来提升监督质量:PPC通过恢复较早阶段更准确的预测来校正退化的正预测,确保可靠的定位和分类信号;NPS则抑制过度自信负样本的影响,防止其为学生提供误导性的监督。为保留有价值的暗知识,我们进一步引入了最大暗知识保留(Maximum Dark Knowledge Preservation, MDKP),在保留非目标类别关系的同时,有选择地精炼目标类别的logits。在MS COCO和PASCAL VOC数据集上的大量实验证明了所提方法的有效性和鲁棒性。我们的代码发布于 https://github.com/xingyitong1/TPRD。
cs.CV / 46 / 2609.19966

Beyond the Foreground: FOV-Aware Polyp Image Synthesis via Lesion-Guided Adaptive Mucosal Context Propagation

超越前景:基于病灶引导的自适应黏膜上下文传播的视场感知息肉图像合成
Wang, Tong, He, Yuting, Ren, Bin, Xie, Yutong, Yang, Guanyu
Abstract
Synthetic image and mask pairs can alleviate scarce colonoscopy annotations, but realistic synthesis requires preserving the supplied lesion while generating compatible mucosa. Existing foreground-guided methods treat all non-foreground pixels as background and rely mainly on local integration. Directly applying them to colonoscopy causes two problems: non-mucosal black regions contaminate generated tissue, and local reasoning produces inconsistent mucosal texture and illumination. We propose LAMP, the first foreground-guided framework for polyp image synthesis based on lesion-guided adaptive mucosal context propagation. LAMP explicitly separates the lesion, valid mucosa, and camera exterior using a field-of-view (FOV) mask. Lesion-to-Mucosa cross-attention extracts lesion appearance conditions for valid-mucosa locations, while FOV-constrained multidirectional Vision Receptance Weighted Key Value propagates them over legal tissue support. An adaptive gate then controls their residual fusion into the diffusion U-Net. Extensive experiments on five polyp datasets demonstrate that LAMP substantially outperforms existing methods in overall generation quality and consistently improves five downstream segmentation models. Our code will be released at https://github.com/wangtong627/LAMP.
Chinese Translation
合成的图像与掩码对可以缓解结肠镜检查标注稀缺的问题,但逼真的合成需要在保留给定病灶的同时生成与之兼容的黏膜组织。现有的前景引导方法将所有非前景像素视为背景,并且主要依赖局部整合。将这些方法直接应用于结肠镜图像会带来两个问题:非黏膜的黑色区域会污染生成的组织,而局部推理会产生不一致的黏膜纹理和光照。我们提出了LAMP,这是首个基于病灶引导的自适应黏膜上下文传播的息肉图像合成前景引导框架。LAMP利用视场(FOV)掩码显式地区分病灶、有效黏膜和相机外部区域。病灶到黏膜的交叉注意力为有效黏膜位置提取病灶外观条件,而FOV约束的多方向Vision Receptance Weighted Key Value则在合法的组织支撑域上传播这些条件。随后,自适应门控控制这些条件以残差融合的方式融入扩散U-Net。在五个息肉数据集上的大量实验表明,LAMP在整体生成质量上显著优于现有方法,并持续提升五个下游分割模型的性能。我们的代码将发布于 https://github.com/wangtong627/LAMP。
cs.CV / 47 / 2609.19973

An Event Preserving Velocity Invariant Representation for Event Cameras

一种面向事件相机的保事件速度不变表示方法
Ikura, Mikihiro, Gava, Luna, Wu, Jiahang, Bartolozzi, Chiara, Glover, Arren
Abstract
Event cameras provide low-latency, high temporal resolution perception for real-time vision tasks such as robotics.The novel circuitry (i.e. asynchronous, independent pixels) that enables these advantages also introduces new algorithmic challenges. Velocity-invariant representations alleviate missing observations under slow motion and motion blur under fast motion, but most discard temporal information by converting events into image-like representations. We propose Set of Centre Active Receptive Fields (SCARF), a real-time velocity-invariant representation that preserves raw events while consistently handling fast motion, stationary scenes, and independently moving objects. SCARF achieves state-of-the-art performance in both computational efficiency and representation quality.
Chinese Translation
事件相机为机器人等实时视觉任务提供了低延迟、高时间分辨率的感知能力。实现这些优势的新型电路结构(即异步、独立的像素)也带来了新的算法挑战。速度不变表示能够缓解慢速运动下的观测缺失问题和快速运动下的运动模糊问题,但大多数方法通过将事件转换为类图像表示而丢弃了时间信息。我们提出了中心活跃感受野集合(Set of Centre Active Receptive Fields, SCARF),这是一种实时的速度不变表示方法,在保留原始事件的同时,能够一致地处理快速运动、静止场景以及独立运动的物体。SCARF 在计算效率和表示质量方面均取得了最先进的性能。
cs.CV / 48 / 2609.19990

QCPruner: Query-Conditioned Population Coverage for Visual Token Pruning

QCPruner:面向视觉令牌剪枝的查询条件化群体覆盖方法
He, Shengli, Liang, Yongchao, He, Roumeng, Zeng, Junjie, He, Jiyuan, Wu, Can, Zheng, Li
Abstract
The high visual-token load in multimodal large language models (MLLMs) motivates training-free pruning to reduce later-layer computation, but under a fixed budget, pruning must preserve query-relevant evidence while avoiding redundancy. Existing methods rank tokens, diversify selected subsets, or optimize coverage without using a shared per-visual query utility to weight both visual targets and candidate representatives. We introduce QCPruner, which makes both roles query-conditioned through bilateral utility weighting. Using keyword-matched query anchors, QCPruner fuses two cross-modal cues into utility and applies it to both visual targets and candidate representatives within visual-affinity-based coverage. The resulting nonnegative facility-location objective is monotone and submodular, retains the standard (1-1/e) greedy guarantee, and requires no model training or parameter updates. Across LLaVA-1.5, LLaVA-NeXT, LLaVA-Video, and Qwen2.5-VL, QCPruner achieves the highest average relative performance among evaluated complete-system pruning methods at every reported token budget. At 32 of 576 tokens on LLaVA-1.5-7B, it retains 96.1% of unpruned performance, versus 93.9% for the strongest evaluated baseline. At 256 of 1296 tokens on Qwen2.5-VL-7B, the corresponding values are 96.7% and 92.5%.
Chinese Translation
多模态大语言模型(MLLM)中视觉令牌的高负载促使研究者采用无需训练的剪枝方法来减少后层计算,但在固定预算下,剪枝必须保留与查询相关的证据,同时避免冗余。现有方法对令牌进行排序、增加所选子集的多样性,或优化覆盖率,但均未使用共享的逐视觉查询效用(utility)来同时对视觉目标和候选代表进行加权。我们提出QCPruner,通过双边效用加权使这两种角色都以查询为条件。QCPruner利用关键词匹配的查询锚点,将两个跨模态线索融合为效用,并在基于视觉亲和度的覆盖机制中将其应用于视觉目标和候选代表。由此得到的非负设施选址(facility-location)目标函数是单调且次模的,保留了标准的(1-1/e)贪心保证,且无需模型训练或参数更新。在LLaVA-1.5、LLaVA-NeXT、LLaVA-Video和Qwen2.5-VL上,在所有报告的令牌预算下,QCPruner在所评估的完整系统剪枝方法中均取得了最高的平均相对性能。在LLaVA-1.5-7B上仅保留576个令牌中的32个时,它保持了96.1%的未剪枝性能,而所评估的最强基线为93.9%。在Qwen2.5-VL-7B上保留1296个令牌中的256个时,对应数值分别为96.7%和92.5%。
cs.CV / 49 / 2609.19991

AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models

AVTrace:诊断全模态模型中的视听时间推理能力
Zhang, Longyin, Mahendra, Parth Sakhare, Wei, Chengwei, Zhang, Ning, Chong, Lim Ming, He, Sirui, Aw, Ai Ti
Abstract
Omni models can describe video content, but can they locate events in time, preserve event order, and judge audio-visual synchronization? We introduce AVTrace (Audio-Visual Temporal Reasoning Assessment and Capability Evaluation), a silver-standard diagnostic suite spanning onset and span grounding, synchronization, next-step prediction, cross-modal localization, chain parsing, and event-conditioned comprehension. It contains 34,114 training examples and category-balanced development and test splits of 3,500 and 7,000 examples. We evaluate five open omni models under their respective input configurations using reference-blind response normalization followed by deterministic scoring. All five off-the-shelf systems score below the test split's majority-label baseline of 0.556 on synchronization verification, and obtain low scores on chain parsing and event-conditioned grounding and comprehension. Development-set perturbations reveal task-dependent sensitivity in Qwen3-Omni-30B to modality removal and changes in visual input processing, without isolating their underlying causes. Parameter-efficient temporal post-training improves Gemma4-E4B-it on several benchmark metrics. On three external image benchmarks, task metrics change modestly, including some degradations, while teacher-forcing perplexity decreases. Together, these findings show that semantic reference-text overlap should not be treated as a proxy for temporal localization, and that AVTrace can identify task-specific weaknesses while providing a testbed for temporal post-training.
Chinese Translation
全模态(Omni)模型能够描述视频内容,但它们能否定位事件发生的时间、保持事件顺序,并判断视听同步性?我们提出了 AVTrace(视听时间推理评估与能力评测),这是一个银标准(silver-standard)诊断套件,涵盖起始时间与时间跨度定位、同步性、下一步预测、跨模态定位、链条解析以及事件条件理解等任务。该套件包含 34,114 个训练样本,以及类别均衡的开发集和测试集(分别为 3,500 和 7,000 个样本)。我们在各模型相应的输入配置下评估了五个开源全模态模型,采用先进行与参考答案无关的响应归一化、再进行确定性评分的方法。全部五个现成系统在同步性验证上的得分均低于测试集多数标签基线 0.556,且在链条解析、事件条件定位与理解任务上得分较低。开发集上的扰动实验显示,Qwen3-Omni-30B 对模态移除和视觉输入处理方式的变化表现出任务相关的敏感性,但并未分离出其根本原因。参数高效的时间后训练在多个基准指标上提升了 Gemma4-E4B-it 的表现。在三个外部图像基准上,任务指标变化不大(包括某些性能下降),而教师强制困惑度(teacher-forcing perplexity)则有所降低。综上,这些发现表明,语义参考文本重叠不应被视为时间定位能力的替代指标,而 AVTrace 能够识别任务特定的弱点,同时为时间后训练提供了一个测试平台。
cs.CV / 50 / 2609.20012

GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction

GRF-Recon:面向长序列前馈重建的全局射线-场优化
Li, Enpeng, Zhang, Yunzhou, Zhang, Zhiyao, Lyu, Dexuan, Wang, Chenyu, Cui, Chiyuan, Cheng, Cheng
Abstract
Feed-forward 3D reconstruction provides an efficient paradigm for scene modeling from image sequences. Scaling these models to large monocular scenarios are constrained by excessive GPU memory footprint, degraded local geometry, and long-term trajectory drift. Existing chunk-based optimization strategies provide limited geometric constraints and fail to maintain global consistency over extended trajectories. We present a unified framework for stable and scalable feed-forward 3D reconstruction from long monocular sequences. Our approach builds on coarse-to-fine trajectory alignment augmented by lightweight geometric prior injection. Distilling monocular geometric cues into the feed-forward backbone via LoRA adaptation improves depth accuracy on fine structures while preserving inference efficiency. We introduce a hybrid-weight sparse ray-field optimization that leverages high-frequency geometric features to guide local point-cloud refinement and enforce consistent inter-frame ray constraints. Unlike prior chunk-based methods, this establishes strong cross-frame geometric coupling while maintaining scalability. Finally, an efficient trajectory stitching strategy with joint ray-error optimization explicitly reduces accumulated drift. Extensive experiments show that our approach achieves competitive trajectory accuracy compared with representative SLAM systems, while maintaining globally consistent 3D reconstruction in large-scale scenarios.
Chinese Translation
前馈式三维重建为基于图像序列的场景建模提供了一种高效范式。然而,将这类模型扩展至大规模单目场景时,受到GPU显存占用过高、局部几何质量下降以及长程轨迹漂移等问题的制约。现有的基于分块的优化策略提供的几何约束有限,难以在长轨迹上保持全局一致性。我们提出了一个统一框架,用于实现从长单目序列中进行稳定且可扩展的前馈三维重建。我们的方法建立在由轻量级几何先验注入增强的由粗到精轨迹对齐之上。通过LoRA适配将单目几何线索蒸馏到前馈主干网络中,在保持推理效率的同时提高了精细结构上的深度精度。我们提出了一种混合权重稀疏射线场优化方法,利用高频几何特征引导局部点云细化,并施加一致的帧间射线约束。与以往基于分块的方法不同,该方法在保持可扩展性的同时建立了强大的跨帧几何耦合。最后,一种结合联合射线误差优化的高效轨迹拼接策略显式地减少了累积漂移。大量实验表明,我们的方法在轨迹精度上可与代表性SLAM系统相媲美,同时能在大规模场景中保持全局一致的三维重建。
cs.CV / 51 / 2609.20034

Astronex-World 1.0: Real-Time Interactive World Model Foundation

Astronex-World 1.0:实时交互式世界模型基础
Zhou, Xin, Miao, Cong
Abstract
We present Astronex-World 1.0, an open controllable video world-model foundation. Given a text prompt (text-to-video) or an initial observation (image-to-video), the model predicts future visual states under frame-aligned camera trajectories, continuous actions, and an embodiment identifier, and accepts text events inserted at a specified position of a rollout. The family provides a bidirectional model for full-context generation and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior. PRoPE injects camera intrinsics and extrinsics, while a 64-dimensional action stream modulates every Transformer layer. A five-stage training path develops bidirectional camera and action control, converts the backbone to block-causal generation, distills a few-step student, restores mixed-domain dynamics, and applies asymmetric DMD/DMD2 distribution matching. The causal model generates 832x480 video at 24 fps. All five training stages run on two NVIDIA L20 48 GB GPUs, and the causal model streams in real time on one. It scores 73.5 on WBench Navi and 70.0 on WBench Full. On Full, this 5B model is above the 13.6B LongCat-Video and the 14B Helios, within one point of the 22B LTX-2.3, and above YUME 1.5, which is post-trained from the same 5B prior on NVIDIA A100 GPUs. The reserved action input and output interfaces allow post-training for embodied intelligence and autonomous driving.
Chinese Translation
我们提出了 Astronex-World 1.0,一个开放的可控视频世界模型基础。给定文本提示(文生视频)或初始观测(图生视频),该模型在帧对齐的相机轨迹、连续动作和具身标识符的条件下预测未来视觉状态,并接受插入到生成序列(rollout)指定位置的文本事件。该模型系列包含一个用于全上下文生成的双向模型,以及一个具有块因果注意力(block-causal attention)和跨块 KV 缓存以支持持续生成的因果模型,二者均基于 Wan2.2-TI2V-5B 先验构建。PRoPE 注入相机内参和外参,同时一个 64 维的动作流调制每一个 Transformer 层。五阶段训练路径依次实现:开发双向相机与动作控制、将骨干网络转换为块因果生成、蒸馏出少步数学生模型、恢复混合域动态,以及应用非对称的 DMD/DMD2 分布匹配。该因果模型能够以 24 fps 生成 832x480 分辨率的视频。全部五个训练阶段仅需两块 NVIDIA L20 48GB GPU,而因果模型可在单块 GPU 上实现实时流式生成。它在 WBench Navi 上得分为 73.5,在 WBench Full 上得分为 70.0。在 WBench Full 上,该 5B 模型的表现优于 13.6B 的 LongCat-Video 和 14B 的 Helios,与 22B 的 LTX-2.3 相差不到一分,并优于在相同 5B 先验上使用 NVIDIA A100 GPU 进行后训练的 YUME 1.5。预留的动作输入与输出接口使其可通过后训练应用于具身智能和自动驾驶领域。
cs.CV / 52 / 2609.20064

A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition

免费的午餐?将 PP-OCRv6 适配于历史文本识别
Kiessling, Benjamin
Abstract
Despite impressive reported scores, large vision-language models have seen limited practical uptake in historical automatic text recognition because of their computational cost, dependence on large-scale pretraining, and hallucination. Historical ATR therefore continues to rely largely on compact CRNN line recognizers, which are visually grounded and trainable on modest data. Lightweight recurrence-free recognizers promise the accuracy of larger models with the practical advantages of CRNNs, yet have not been comprehensively evaluated on historical writing. We adapt PP-OCRv6, a recent compact text recognizer without strong language modeling, for historical line recognition and compare it with a conventional CRNN across generalized pretraining, domain-specific training, corpus-level fine-tuning, and manuscript-specific few-shot adaptation on multilingual Latin- and Arabic-script material. While PP-OCRv6 does not consistently outperform the baseline when trained from scratch, heterogeneous pretraining produces markedly better generalization. Comparisons with the Qwen3.5-based Medusa recognizer further show that fine-tuned PP-OCRv6 can outperform a large VLM tailored towards historical Latin-script HTR.
Chinese Translation
尽管大型视觉-语言模型(VLM)在报告中取得了令人瞩目的成绩,但由于其计算成本高、依赖大规模预训练以及存在幻觉问题,它们在历史自动文本识别(ATR)中的实际应用仍然有限。因此,历史 ATR 目前在很大程度上仍依赖于紧凑型 CRNN 行识别器,这类模型具有视觉基础,且可在适度数据量下进行训练。轻量级无循环结构(recurrence-free)识别器有望兼具大型模型的准确率和 CRNN 的实用优势,然而尚未在历史手写文本上得到全面评估。我们将 PP-OCRv6——一个近期出现的、不依赖强大语言建模的紧凑型文本识别器——适配于历史行识别,并在多语言拉丁文字和阿拉伯文字材料上,将其与传统 CRNN 在通用预训练、领域特定训练、语料库级微调以及手稿特定少样本适配等多个层面进行了比较。尽管 PP-OCRv6 在从零训练时并未持续优于基线模型,但异构预训练带来了显著更好的泛化能力。与基于 Qwen3.5 的 Medusa 识别器的比较进一步表明,经过微调的 PP-OCRv6 能够超越专为历史拉丁文字 HTR 打造的大型视觉-语言模型。
cs.CV / 53 / 2609.20066

PointEvent: Rethinking Event-based Tiny Object Detection via Serialized Motion Evidence Accumulation

PointEvent:基于序列化运动证据积累的事件相机微小目标检测方法再思考
Wu, Zongze, Jia, Baofeng, Yan, Weiqi, Zhang, Jingyuan, Zang, Yu, Chen, Xiaoyu, Han, Jing
Abstract
Event cameras offer high temporal resolution and motion sensitivity for tiny UAV detection, yet distant targets generate sparse and fragmented events that are easily overwhelmed by clutter and ego-motion. Existing methods mainly rely on dense event representations or local sparse spatiotemporal modeling, resulting in redundant computation or fragmented modeling of motion continuity across distant asynchronous events. To address this limitation, we introduce serialized motion evidence accumulation, which treats motion continuity as an ordered evidence propagation process. Specifically, the same event stream is organized into locality-preserving spatiotemporal paths and chronology-preserving temporal paths through the latent complementary serializations. Based on this principle, we propose PointEvent, a lightweight event-wise state-space framework that alternates serialized scans across the complementary orders, progressively consolidating fragmented motion evidence beyond fixed local neighborhoods. A high-resolution event branch preserves fine-grained target responses, while compact context modulation suppresses interference. Experiments demonstrate that PointEvent achieves SOTA with the fewest parameters and fastest measured inference among the compared methods. Code: https://github.com/wzz-z/PointEvent
Chinese Translation
事件相机为微小无人机目标检测提供了高时间分辨率和运动敏感性,然而远距离目标产生的事件稀疏且碎片化,容易被杂波和自身运动所淹没。现有方法主要依赖稠密事件表示或局部的稀疏时空建模,导致计算冗余,或对远距离异步事件间运动连续性的建模碎片化。为解决这一局限,我们引入序列化运动证据积累,将运动连续性视为一种有序的证据传播过程。具体而言,同一事件流通过隐式的互补序列化被组织为保持局部性的时空路径和保持时序性的时间路径。基于该原理,我们提出 PointEvent——一个轻量级的事件级状态空间框架,它在互补顺序之间交替进行序列化扫描,在固定局部邻域之外逐步整合碎片化的运动证据。高分辨率事件分支保留了细粒度的目标响应,而紧凑的上下文调制则抑制干扰。实验表明,PointEvent 在对比方法中以最少的参数量和最快的实测推理速度达到了当前最优(SOTA)性能。代码:https://github.com/wzz-z/PointEvent
cs.CV / 54 / 2609.20088

G^2RA-NET: Graph-based Cross-Slice Relation Modeling with Attention Gating for Medical Image Segmentation

G^2RA-NET:基于图的跨层关系建模与注意力门控的医学图像分割
Wang, Shengye, Wu, Zonglin, Fan, Liang, Xue, Yule, Zhao, Haozhe
Abstract
Medical image segmentation supports quantitative clinical analysis and computer-aided diagnosis. Recent methods for medical image segmentation have improved both local feature representation and volumetric context modeling. However, existing methods still strug- gle to efficiently model cross-slice relations in anisotropic volumet- ric images, limiting segmentation consistency and accuracy. This pa- per proposes G^2RA-Net, a medical image segmentation framework that combines graph-based cross-slice relation modeling with atten- tion gating. Graph-Based Slice Relationship Modeling (GSRM) cap- tures anatomical dependencies across consecutive slices by repre- senting each slice as a graph node and propagating semantic con- text through graph message passing. The Cross-Slice Attention Gate (CSAG) then selects relevant neighboring context and emphasizes target anatomical regions through attention-guided feature modula- tion. Experiments on brain MRI and abdominal CT datasets demon- strate that G^2RA-Net outperforms representative methods in seg- mentation accuracy and boundary quality. Ablation studies further validate the proposed design.
Chinese Translation
医学图像分割为定量临床分析和计算机辅助诊断提供支持。近期的医学图像分割方法在局部特征表示和体积上下文建模方面均有所改进。然而,现有方法仍难以高效地建模各向异性体积图像中的跨层关系,限制了分割的一致性和准确性。本文提出G^2RA-Net,一种将基于图的跨层关系建模与注意力门控相结合的医学图像分割框架。基于图的层间关系建模模块(GSRM)将每个切片表示为图节点,并通过图消息传递传播语义上下文,从而捕捉连续切片之间的解剖学依赖关系。随后,跨层注意力门(CSAG)通过注意力引导的特征调制,选择相关的相邻上下文并突出目标解剖区域。在脑部MRI和腹部CT数据集上的实验表明,G^2RA-Net在分割精度和边界质量方面优于代表性方法。消融实验进一步验证了所提设计的有效性。
cs.CV / 55 / 2609.20100

A Smaller Transformer in Your Transformer

Transformer中的小型Transformer
Tomar, Dhananjay, Aasan, Marius, Kleppe, Andreas, Rivera, Adín Ramírez
Abstract
Recent findings indicate that Vision Transformers settle into locally similar computational phases, implying a level of depthwise computational redundancy. However, existing methods to exploit this redundancy either fail to reduce inference compute or severely degrade model expressivity. In this work, we formalise a unified view of block redundancy that decouples the geometry from specific surrogate interventions. We then introduce Transformer-Within-Transformer (TWT), a post-hoc method that fuses contiguous groups of redundant layers into a single learned surrogate layer. TWT reduces parameter count and inference compute while remaining competitive with original models using half the depth on natural images, and in several downstream histopathology settings, TWT matches or even improves on the original baseline.
Chinese Translation
近期研究发现,视觉Transformer(Vision Transformers)会进入局部相似的计算阶段,这暗示了存在一定程度的深度方向计算冗余。然而,现有的利用这种冗余的方法要么无法降低推理计算量,要么会严重损害模型的表达能力。在本工作中,我们对块冗余(block redundancy)进行了统一的形式化描述,将几何结构与特定的代理干预(surrogate interventions)解耦。随后,我们提出了Transformer内嵌Transformer(Transformer-Within-Transformer,TWT),这是一种事后(post-hoc)方法,可将连续的冗余层组融合为一个学习得到的代理层。TWT在降低参数量和推理计算量的同时,在自然图像上仍与深度减半的原始模型保持竞争力;在若干下游组织病理学任务中,TWT能够达到甚至超越原始基线模型的性能。
cs.CV / 56 / 2609.20106

AnyviewMeter: Adapting Robotic Reward Models with Camera Geometry and Multi-View Attention

AnyviewMeter:利用相机几何与多视角注意力机制适配机器人奖励模型
Tu, Yuang, Tan, Runjia, Yan, Yujie, Hu, Jinghan, Lv, Chen
Abstract
Robotic reward models evaluate task execution from visual observations, but their predictions can change with camera viewpoint and occlusion even when the underlying task state is unchanged. Adapting a pretrained reward model to a local task therefore requires accounting for how that task is observed. We introduce AnyviewMeter, a geometry-conditioned adaptation framework for robotic reward models that represent task progress as a scalar reward signal. It combines low-rank fine-tuning with token-aligned Plucker rays and synchronous block attention: ray conditioning incorporates camera geometry into visual features and attention queries and keys, while block attention fuses synchronized views inside the pretrained decoder. The framework supports both single-view reward prediction and joint multi-view evaluation through parameter-efficient adaptation of a pretrained Robometer model. On PickCube, single-view adaptation improves progress prediction in every camera group and reduces mean absolute error under a changed field of view by approximately 21% relative to RGB fine-tuning. Across simulated manipulation tasks, joint multi-view prediction reduces progress error by 41-69% compared with averaging single-view RGB predictions and improves temporal ordering in approximately 88% of task-camera groups. On real tasks with fixed and wrist-mounted cameras, mean absolute error decreases by approximately 21% relative to averaged RGB fine-tuning. These results support camera geometry and joint visual evidence as useful components of task-specific robotic reward adaptation.
Chinese Translation
机器人奖励模型通过视觉观测来评估任务执行情况,但即使底层任务状态未发生变化,其预测也可能随相机视角和遮挡情况而改变。因此,将预训练的奖励模型适配到特定本地任务时,需要考虑该任务的观测方式。我们提出了AnyviewMeter,一种面向机器人奖励模型的几何条件化适配框架,该模型以标量奖励信号表示任务进度。该框架将低秩微调与token对齐的Plücker射线(Plücker rays)及同步块注意力相结合:射线条件化将相机几何信息融入视觉特征以及注意力的查询和键中,而块注意力则在预训练解码器内部融合同步的多视角信息。该框架通过对预训练的Robometer模型进行参数高效适配,同时支持单视角奖励预测和联合多视角评估。在PickCube任务上,单视角适配在所有相机组中均提升了进度预测性能,在视野变化情况下,相比RGB微调将平均绝对误差降低约21%。在多个仿真操作任务中,联合多视角预测相比单视角RGB预测取平均的方式将进度误差降低41%-69%,并在约88%的任务-相机组中改善了时序排序。在使用固定相机和腕戴式相机的真实任务上,相比RGB微调取平均的方式,平均绝对误差降低约21%。这些结果表明,相机几何信息和联合视觉证据是任务特定的机器人奖励适配的有效组成部分。
cs.CV / 57 / 2609.20139

Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models

跨模态注意力扮演频率滤波器角色:为何冗长提示词能提升视觉-语言模型的鲁棒性
Wani, Farooq Ahmad, Bucarelli, Maria Sofia, Mirza, Mujtaba Hussain, Pryymak, Oleksandr, Gema, Aryo Pradipta, Masi, Iacopo, Minervini, Pasquale, Silvestri, Fabrizio
Abstract
Vision-language models (VLMs) are fragile under image corruption. We find that the wording of the question affects VLMs in two opposite ways. Verbose questions make VLMs substantially more robust---e.g., rephrasing "Is there a cat?" into "Please look carefully and answer: is there a cat?". Conversely, VLMs become more fragile under corruption when the question is semantically complex or finer-grained, e.g., "what colour is the cup left of the chair?" instead of "is there a cup?". Both effects stem from question-conditioned cross-modal attention, which induces a spectral filter over image patches: verbose questions broaden its frequency support, while fine-grained questions concentrate it onto fewer visual scales. The model's answer drifts most when this filter and the corruption sit on the same spatial frequencies. We test the filter view on Qwen3-VL and LLaVA-OneVision across GQA and CLEVR; verbose paraphrasing reduces drift variance by 70--81% on the 8B models. The practical recipe---pad the prompt---further yields measurable gains in accuracy, even under image corruption.
Chinese Translation
视觉-语言模型(VLM)在图像损坏下表现脆弱。我们发现提问的措辞会以两种截然相反的方式影响VLM:冗长的问题能显著提升VLM的鲁棒性——例如,将“Is there a cat?”(有一只猫吗?)改写为“Please look carefully and answer: is there a cat?”(请仔细观察并回答:有一只猫吗?)。相反,当问题在语义上更复杂或更细粒度时,VLM在损坏下会变得更脆弱,例如用“what colour is the cup left of the chair?”(椅子左边的杯子是什么颜色?)替代“is there a cup?”(有杯子吗?)。这两种效应都源于以问题为条件的跨模态注意力,它对图像块构成一种频谱滤波器:冗长的问题拓宽了该滤波器的频率支持范围,而细粒度的问题则将其集中于更少的视觉尺度上。当该滤波器与图像损坏处于相同的空间频率时,模型答案的漂移最为严重。我们在Qwen3-VL和LLaVA-OneVision上针对GQA和CLEVR数据集验证了这一滤波器观点:冗长改写在8B模型上使漂移方差降低了70%–81%。这一实用方法——扩充提示词——即使存在图像损坏,也能带来可测量的准确率提升。
cs.CV / 58 / 2609.20147

Bridging Modalities on the Cortex: Surface-based MRI to PET Translation with a Diffusion Bridge

在皮层上连接不同模态:基于表面的MRI到PET转换的扩散桥方法
Li, Yitong, Samoylova, Alexandra, Bongratz, Fabian, Grimmer, Timo, Hedderich, Dennis M., Yakushev, Igor, Wachinger, Christian
Abstract
Cortical hypometabolism measured by Fluorodeoxyglucose Positron Emission Tomography (FDG-PET) is a highly sensitive biomarker for dementia diagnosis. However, high costs, radiation exposure, and limited accessibility constrain its clinical utility. While cross-modal synthesis from Magnetic Resonance Imaging (MRI) offers a promising alternative, existing volumetric generation methods do not explicitly account for the highly folded cortical geometry, where disease-related patterns predominantly reside. To address this, we introduce a novel surface-based diffusion bridge framework DB-SUiT for MRI-to-PET translation that operates natively on the cortical manifold. A conditional Spherical U-shaped vision Transformer (SUiT) is specifically designed to model the intricate cross-modal relationships while preserving surface topology. It combines spherical convolutional encoders for multi-scale surface feature extraction with bottleneck Transformers to capture long-range spatial dependencies, while incorporating demographic and subcortical conditions to refine the synthesis. Evaluated on two datasets, including subjects with different dementia types, DB-SUiT demonstrates high-fidelity synthesis that substantially outperforms other baselines. In automated dementia classification, synthesized PET surfaces improve performance over MRI by 14.2% and PET volumes by 11.3%, approaching the performance of real PET surfaces. In a blinded reader study, synthetic PET achieved 85.5% diagnostic accuracy, compared with 75.8% for MRI and 95.2% for real PET. This further demonstrates cross-cohort and cross-pathology generalization, as the model was evaluated without retraining on an external cohort that included a dementia subtype not represented during training. Our code is available at https://github.com/ai-med/DB-SUiT.
Chinese Translation
通过氟脱氧葡萄糖正电子发射断层扫描(FDG-PET)测量的皮层低代谢是痴呆症诊断的高度敏感生物标志物。然而,高昂的成本、辐射暴露以及有限的可及性制约了其临床应用。虽然从磁共振成像(MRI)进行跨模态合成提供了一种有前景的替代方案,但现有的体积生成方法并未显式考虑疾病相关模式主要所在的高度折叠的皮层几何结构。为解决这一问题,我们提出了一种新颖的基于表面的扩散桥框架DB-SUiT,用于在皮层流形上原生运行的MRI到PET转换。我们专门设计了条件球面U形视觉Transformer(SUiT)来建模复杂的跨模态关系,同时保持表面拓扑结构。该模型结合了用于多尺度表面特征提取的球面卷积编码器与捕获长程空间依赖关系的瓶颈Transformer,并融入人口统计学信息和皮层下条件以优化合成结果。在包含不同痴呆类型受试者的两个数据集上的评估表明,DB-SUiT实现了高保真合成,显著优于其他基线方法。在自动化痴呆分类任务中,合成PET表面相比MRI将性能提升了14.2%,相比PET体积提升了11.3%,接近真实PET表面的性能。在一项盲法阅片研究中,合成PET的诊断准确率达到85.5%,而MRI为75.8%,真实PET为95.2%。该模型还在一个外部队列上无需重新训练进行评估,该队列包含训练中未出现过的痴呆亚型,进一步证明了跨队列和跨病理的泛化能力。我们的代码已发布于https://github.com/ai-med/DB-SUiT。
cs.CV / 59 / 2609.20150

Task-Oriented Semantic Feature Transmission for Multi-Task Satellite Remote Sensing over Low-SNR Channels

面向任务的语义特征传输:面向低信噪比信道下的多任务卫星遥感
Sun, Shuoyuan, Wang, Hongyu, Peng, Mugen, Xu, Wenjia
Abstract
Conventional satellite remote sensing transmission follows a reconstruct-then-infer paradigm that optimizes pixel-level fidelity, creating an objective mismatch with downstream tasks such as classification and detection, especially at low SNR. This paper investigates a task-oriented framework that bypasses image reconstruction and directly transmits semantic features extracted by a multitask-pretrained backbone. A lightweight channel adaptation module (CAM) compresses feature dimensionality for bandwidth reduction, and a feature restorer recovers task-relevant structure after channel corruption. With the backbone frozen, the CAM and task-specific downstream heads are jointly optimized with task and feature-level supervision under random-SNR training. Under the adopted AWGN setting, experiments on scene classification and object detection show consistent gains over reconstruction-oriented JSCC baselines across different SNR conditions, with the largest improvements in the low-SNR regime.
Chinese Translation
传统的卫星遥感传输遵循“先重建后推理”的范式,优化的是像素级保真度,这与分类和检测等下游任务的目标不匹配,尤其是在低信噪比(SNR)条件下。本文研究了一种面向任务的框架,该框架跳过图像重建,直接传输由多任务预训练主干网络提取的语义特征。一个轻量级的信道自适应模块(CAM)压缩特征维度以降低带宽占用,特征恢复器则在信道失真后恢复与任务相关的结构信息。在冻结主干网络的情况下,CAM与各任务特定的下游头在随机信噪比训练条件下,通过任务级与特征级监督进行联合优化。在所采用的加性高斯白噪声(AWGN)信道设置下,场景分类和目标检测实验表明,该框架在不同信噪比条件下均持续优于面向重建的联合信源信道编码(JSCC)基线方法,且在低信噪比区间改进最为显著。
cs.CV / 60 / 2609.20151

Ischemic Stroke Segmentation and Net Water Uptake Quantification on Multicenter Non-Contrast CT Using Supervised Target-Domain Adaptation

基于监督式目标域自适应的多中心非增强CT缺血性脑卒中分割及净摄水量定量
Britt, Linus, Nielsen, Maximilian, Klapproth, Susan, Kemmling, Andre, Lev, Michael H., Broocks, Gabriel, Werner, Rene, Sentker, Thilo
Abstract
Objectives: Quantitative assessment of infarct hypodensity on non-contrast computed tomography (NCCT), including net water uptake (NWU), requires manual or semi-manual lesion delineation, often guided by CT perfusion or diffusion-weighted MRI, limiting clinical applicability. Automated segmentation on NCCT could enable efficient biomarker extraction such as NWU but remains challenging across heterogeneous multicenter data. This study aimed to develop and externally test a domain-aware deep learning framework for ischemic stroke segmentation on NCCT and assess its suitability for NWU quantification. Materials & Methods: In this retrospective multicenter study of 801 patients from four datasets, an nnU-Net-based model was trained on NCCT scans from the University Medical Center Hamburg-Eppendorf and the Acute Ischemic Stroke Dataset. To adapt to new domains, the model was fine-tuned on target-domain subsets from Boston (n=11) and ISLES (n=75), with evaluation on held-out cases not used for fine-tuning. Automated segmentations and NWU values were compared with expert references. Results: For lesions $\geq$ 30 mL, median Dice was 0.68 (Boston) and 0.56 (ISLES). Including smaller lesions, which predominated in ISLES, median Dice was 0.54 (interquartile range [IQR] 0.30-0.70) for acute lesion segmentation (Boston dataset) and 0.20 (IQR 0.03-0.41) for NCCT lesion segmentations when compared to post-treatment infarct (primary target of the ISLES challenge). Automated NWU mean absolute error was 1.37 percentage points (SD 1.61, Boston). Conclusion: Target-domain adaptation supported NCCT-only infarct segmentation across heterogeneous external cohorts, although performance varied across domains. The approach enabled low-error NWU quantification from baseline NCCT without advanced imaging, supporting further prospective clinical evaluation.
Chinese Translation
目的:对非增强计算机断层扫描(NCCT)上梗死低密度区的定量评估,包括净摄水量(NWU),通常需要手动或半手动勾画病灶,且常需借助CT灌注或弥散加权MRI进行引导,限制了其临床适用性。基于NCCT的自动分割可高效提取NWU等生物标志物,但在异质性的多中心数据中仍具挑战性。本研究旨在开发并外部验证一种域感知的深度学习框架,用于NCCT上的缺血性脑卒中分割,并评估其在NWU定量中的适用性。材料与方法:在这项纳入来自四个数据集的801例患者的回顾性多中心研究中,基于nnU-Net的模型使用汉堡-埃彭多夫大学医学中心和急性缺血性脑卒中数据集的NCCT扫描进行训练。为适应新域,模型在来自Boston(n=11)和ISLES(n=75)的目标域子集上进行微调,并在未用于微调的保留病例上进行评估。将自动分割结果和NWU值与专家参考标准进行比较。结果:对于体积≥30 mL的病灶,中位Dice系数为0.68(Boston)和0.56(ISLES)。纳入较小的病灶(在ISLES中占主导)后,急性病灶分割的中位Dice系数为0.54(四分位距[IQR] 0.30-0.70,Boston数据集),与治疗后梗死灶(ISLES挑战赛的主要目标)相比,NCCT病灶分割的中位Dice系数为0.20(IQR 0.03-0.41)。自动NWU的平均绝对误差为1.37个百分点(标准差1.61,Boston)。结论:目标域自适应支持在异质性外部队列中仅基于NCCT的梗死分割,尽管性能在不同域间存在差异。该方法无需高级影像检查即可从基线NCCT实现低误差的NWU定量,为后续前瞻性临床评估提供了支持。
cs.CV / 61 / 2609.20160

Needles in a Raystack: Ultra-Sparse LiDAR Occupancy Detection for Bat Tracks

光线堆叠中的大海捞针:面向蝙蝠飞行轨迹的超稀疏激光雷达占据检测
Klar, Nico, Rana, Pankaj, Gifary, Nizam, Traub, Jakob, Ahmad, Aamir
Abstract
Monitoring flying animals is important for understanding and protecting biodiversity, but nocturnal species such as bats are difficult to observe in the field. Using LiDAR, bat movements at night result in ultra-sparse 3D spatio-temporal data in which standard reconstruction losses tend to predict only background and miss real flight paths. We study this problem as voxel-wise occupancy detection in sensor-centric LiDAR raystacks. A lightweight 3D U-Net is proposed that preserves temporal resolution, uses skip connections for spatial detail, and combines weighted binary cross-entropy with Dice loss to handle the strong class imbalance. In real LiDAR recordings of bats over open fields, cross-checked with acoustic monitoring, a reconstruction-based 3D convolutional autoencoder baseline fails to recover foreground trajectories. In contrast, the proposed U-Net recovers sparse foreground occupancy in diagnostic experiments and produces coherent occupancy patterns along bat flight trajectories, providing a practical basis for validation-scale experiments, later clustering of flight tracks, and future integration of bat activity information into biodiversity-aware turbine curtailment strategies.
Chinese Translation
监测飞行中的动物对于理解和保护生物多样性至关重要,但蝙蝠等夜行物种在野外难以观察。利用激光雷达(LiDAR),蝙蝠在夜间的运动会产生超稀疏的三维时空数据,标准的重建损失往往只预测背景而遗漏真实的飞行路径。我们将该问题转化为以传感器为中心的LiDAR光线堆叠(raystack)中的体素级占据检测任务。本文提出一种轻量级3D U-Net,该网络保持了时间分辨率,利用跳跃连接保留空间细节,并将加权二元交叉熵与Dice损失相结合以应对严重的类别不平衡。在经声学监测交叉验证的野外蝙蝠真实LiDAR记录中,基于重建的3D卷积自编码器基线无法恢复前景轨迹;相比之下,所提出的U-Net在诊断性实验中能够恢复稀疏的前景占据,并沿蝙蝠飞行轨迹生成连贯的占据模式,为验证规模的实验、后续的飞行轨迹聚类以及未来将蝙蝠活动信息整合到兼顾生物多样性的风机降载策略中提供了切实可行的基础。
cs.CV / 62 / 2609.20178

MTF-Net: Multi-Modal Temporal Feature Fusion Network for Pedestrian Intention Prediction

MTF-Net:用于行人意图预测的多模态时序特征融合网络
Rahman, Md Mahfuzur, Zhou, Pengzhan, Noor, A. F. M. Abdun, Ahasan, Md Imam, Rahman, Md Mustafizur, Qu, Fang
Abstract
Accurately predicting pedestrian intentions is crucial for ensuring safe and proactive interaction between autonomous vehicles and pedestrians. However, existing approaches often depend on architectures that either model temporal dependencies within individual modalities or fuse modalities only at coarse semantic levels. To address these limitations, we propose MTF-Net, a novel Multi-Modal Temporal Feature Fusion Network that jointly models kinematic, appearance, and contextual cues for pedestrian intention prediction. MTF-Net integrates four complementary modalities-bounding-box dynamics, human pose keypoints, local context, and scene-level semantics within a recurrent fusion framework enhanced by gated linear units (GLUs). These GLU-based modules adaptively regulate cross-modal information flow, enabling interpretable and efficient feature interaction across temporal scales. Through three dedicated temporal encoding branches and an attention-guided fusion head, the proposed model robustly anticipates pedestrian crossing intentions several frames before they occur. Extensive evaluations on the PIE and JAAD benchmarks demonstrate that MTF-Net surpasses recent transformer- and graph-based models, achieving up to 0.95 AUC on PIE and 0.94 AUC on JAAD, while maintaining real-time performance. The results highlight that reliable pedestrian intention prediction arises from principled multi-modal fusion rather than excessive architectural complexity.
Chinese Translation
准确预测行人意图对于确保自动驾驶车辆与行人之间安全、主动的交互至关重要。然而,现有方法通常依赖于以下两类架构:要么仅建模单一模态内部的时序依赖关系,要么仅在粗粒度语义层面进行模态融合。为解决这些局限性,我们提出了MTF-Net,一种新颖的多模态时序特征融合网络,它联合建模运动学、外观和上下文线索以进行行人意图预测。MTF-Net在由门控线性单元(GLU)增强的循环融合框架中集成了四种互补的模态——边界框动态、人体姿态关键点、局部上下文和场景级语义。这些基于GLU的模块能够自适应地调节跨模态信息流,实现跨时间尺度上可解释且高效的特征交互。通过三个专用的时序编码分支和一个注意力引导的融合头,所提出的模型能够在行人过街意图发生前若干帧对其进行稳健的预测。在PIE和JAAD基准数据集上的大量评估表明,MTF-Net超越了近期基于Transformer和图网络的模型,在PIE上达到0.95的AUC,在JAAD上达到0.94的AUC,同时保持了实时性能。这些结果表明,可靠的行人意图预测源于有原则的多模态融合,而非过度的架构复杂性。
cs.CV / 63 / 2609.20181

MoSSGate: Memory-Modulated State-Space Gating for Skin Lesion Segmentation

MoSSGate:用于皮肤病变分割的记忆调制状态空间门控机制
Awan, Anum, Buriro, Mahnoor, Khan, Muhammad Younas, Ahasan, Md Imam
Abstract
Accurate skin lesion segmentation is crucial for reliable computer-aided dermatological diagnosis, yet existing convolutional and transformer-based models often struggle to jointly capture long-range spatial dependencies and fine boundary details under limited computational budgets. This trade-off between global context modeling and boundary-aware localization frequently leads to over-segmentation, fragmented predictions, or missing thin peripheral structures. To address this challenge, we propose MoSSGate, a plug-and-play module for U-Net that integrates (i) boundary-aware spatial gating to restrict long-range propagation to informative regions, (ii) an external memory modulator that provides sample-adaptive dynamic control, and (iii) parallel 2D state-space modeling for efficient global context aggregation with linear complexity. The proposed design enables adaptive, context-aware information propagation while preserving sharp and accurate lesion boundaries. Extensive experiments on the ISIC 2017 and ISIC 2018 benchmarks demonstrate state-of-the-art accuracy with strong efficiency, achieving 86.3% and 85.9% mIoU and 92.6% and 90.6% Dice, respectively, while requiring substantially fewer FLOPs than most competing CNN-based methods. These results highlight a favorable accuracy efficiency trade-off for high-resolution medical image segmentation.
Chinese Translation
精确的皮肤病变分割对于可靠的计算机辅助皮肤病诊断至关重要,然而现有的卷积模型和基于Transformer的模型在有限计算预算下往往难以同时捕获长程空间依赖和精细的边界细节。全局上下文建模与边界感知定位之间的这种权衡常常导致过度分割、预测碎片化或细小外周结构缺失。为应对这一挑战,我们提出了MoSSGate,这是一种可用于U-Net的即插即用模块,它集成了:(i)边界感知空间门控,将长程信息传播限制在信息丰富的区域;(ii)提供样本自适应动态控制的外部记忆调制器;以及(iii)以线性复杂度实现高效全局上下文聚合的并行2D状态空间建模。所提出的设计实现了自适应的、上下文感知的信息传播,同时保持了锐利准确的病变边界。在ISIC 2017和ISIC 2018基准数据集上的大量实验表明,该方法以高效率达到了最先进的精度,mIoU分别为86.3%和85.9%,Dice分别为92.6%和90.6%,且所需的FLOPs远低于大多数基于CNN的竞争方法。这些结果突显了该方法在高分辨率医学图像分割中良好的精度-效率权衡。
cs.CV / 64 / 2609.20222

A Two-Stage Multi-Scale Attention-Based Network for Weakly Supervised Cataract Fundus Image Enhancement

一种用于弱监督白内障眼底图像增强的两阶段多尺度注意力网络
Fang, Xiaoyong, Wang, Yue, Li, Xiangyu, Fan, Wanshu, Zhou, Dongsheng
Abstract
Cataract is a major cause of vision loss and hinders further diagnosis. However, cataract fundus image enhancement often grapples with challenges such as limited paired cataract retinal images and insufficient recovery of fine details in the retinal images. To mitigate these challenges, we in this paper propose a two-stage multi-scale attention-based network (TSMSA-Net) for weakly supervised cataract fundus image enhancement. Our TSMSA-Net leverages the domain transformation to synthesis paired real-like cataract images, solving the problem of difficult acquisition of paired images. To further extract detailed information from fundus images and reduce the generation of artifacts during the enhancement process, we propose a multi-scale attention-based stage to learn more useful features for cataract image enhancement. Experimental results on Kaggle and ODIR-5K demonstrate that our TSMSA-Net outperforms current state-of-the-art cataract fundus images enhancement even without paired images and exhibits certain generalization ability. Experimental results on Kaggle and ODIR-5K datasets indicate that our TSMSA-Net outperforms the current state-of-the-art methods for cataract fundus image enhancement, even in the absence of paired images. Additionally, it demonstrates a certain level of generalization capability. The enhancement also can improve the performance of vessel segmentation and classification in cataract images.
Chinese Translation
白内障是导致视力下降的主要原因之一,并阻碍进一步的眼底诊断。然而,白内障眼底图像增强常常面临配对白内障视网膜图像稀缺以及视网膜图像细微细节恢复不足等挑战。为缓解这些问题,本文提出了一种用于弱监督白内障眼底图像增强的两阶段多尺度注意力网络(TSMSA-Net)。我们的TSMSA-Net利用域变换来合成成对的类真实白内障图像,解决了配对图像难以获取的问题。为了进一步从眼底图像中提取细节信息并减少增强过程中伪影的生成,我们提出了一个基于多尺度注意力的阶段,以学习对白内障图像增强更有用的特征。在Kaggle和ODIR-5K数据集上的实验结果表明,即使在缺乏配对图像的情况下,我们的TSMSA-Net仍优于当前最先进的白内障眼底图像增强方法,并表现出一定的泛化能力。此外,该增强方法还能提升白内障图像中血管分割和分类的性能。
cs.CV / 65 / 2609.20235

Scene-Q: Confidence-Aware Coarse-to-Fine Querying of 3D Scenes with Selective VLM Reasoning

Scene-Q:基于选择性VLM推理的置信度感知三维场景由粗到细查询方法
Kim, Juno, Park, Yesol, Yoon, Hye-Jung, Zhang, Byoung-Tak
Abstract
Indoor mobile robots require open-vocabulary scene understanding that grounds natural-language queries in a consistent 3D map. Many existing systems ultimately rely on cosine-similarity retrieval with contrastive image--text encoders, which is efficient but brittle when labels are near-synonymous or multiple similar instances appear. We present Scene-Q, a confidence-aware coarse-to-fine querying framework that normalizes encoder scores with temperature scaling and selectively invokes a reasoning VLM only for low-confidence cases. High-confidence queries are answered by fast retrieval, while ambiguous ones are reranked over a small top-K candidate set using the original multi-view images and instance bounding boxes, enabling context-aware disambiguation at low cost. Scene-Q improves open-vocabulary 3D instance segmentation on ScanNet200 and natural-language 3D instance retrieval on real-world reconstructions, with the largest gains on spatial and relational queries while keeping a substantial fraction of queries on the fast path.
Chinese Translation
室内移动机器人需要开放词汇的场景理解能力,即将自然语言查询与一致的三维地图相关联。许多现有系统最终依赖于对比图文编码器的余弦相似度检索,这种方法虽然高效,但在标签近似同义或出现多个相似实例时表现脆弱。我们提出了Scene-Q,这是一个置信度感知的由粗到细查询框架,它通过温度缩放对编码器分数进行归一化,并仅针对低置信度的情况选择性地调用推理型视觉语言模型(VLM)。高置信度查询通过快速检索直接回答,而模糊查询则利用原始多视角图像和实例边界框,在较小的top-K候选集上进行重排序,从而以较低成本实现上下文感知的消歧。Scene-Q在ScanNet200上的开放词汇三维实例分割以及真实世界重建场景的自然语言三维实例检索任务上均取得提升,其中在空间和关系类查询上收益最大,同时将相当大比例的查询保留在快速路径上。
cs.CV / 66 / 2609.20245

SAGE-Yoga: Multi-Cue Learning for Yoga Pose Classification and Joint-Level Correction

SAGE-Yoga:用于瑜伽体式分类与关节级纠正的多线索学习方法
Chi, Hung Le, Huynh, Khanh Minh, Pham, Long Nghia Tran, Huynh, Tan Phuc, Nguyen, Trong-Thuan, Tran, Minh-Triet
Abstract
Automated yoga analysis requires both accurate pose classification and interpretable feedback on pose execution. However, existing methods often rely on a single visual prediction, struggle to distinguish visually similar poses, and treat pose classification and correction as separate tasks. To address these limitations, we propose SAGE-Yoga, a unified coarse-to-fine framework for yoga pose classification and joint-level correction from a single RGB image. Inspired by how yoga instructors assess posture using multiple complementary cues, SAGE-Yoga first employs a bagging-based ensemble of complementary visual backbones to generate a ranked set of candidate pose classes. Additionally, a margin-based gating mechanism preserves confident visual predictions while invoking geometric verification only for ambiguous cases. Moreover, once the final pose class is determined, SAGE-Yoga retrieves a medoid reference pose and compares the observed joint angles with class-specific distributions to identify misaligned joints. Finally, these deviations are translated into actionable corrective feedback. Empirically, experiments on the Yoga-82 dataset show that the visual ensemble achieves 89.0% Top-1 accuracy, while the complete framework improves performance to 90.7% Top-1 accuracy and 90.1% Macro-F1. These results demonstrate that combining complementary visual evidence with selective geometric verification improves fine-grained pose classification while enabling interpretable, joint-level correction.
Chinese Translation
自动化瑜伽分析既需要准确的体式分类,也需要对体式执行情况提供可解释的反馈。然而,现有方法通常依赖于单一的视觉预测,难以区分视觉上相似的体式,并将体式分类与纠正视为相互独立的任务。为解决这些局限性,我们提出了SAGE-Yoga,这是一个统一的由粗到细的框架,可从单张RGB图像实现瑜伽体式分类和关节级纠正。受瑜伽教练利用多种互补线索评估姿势的启发,SAGE-Yoga首先采用基于装袋(bagging)的互补视觉骨干网络集成,生成一组排序后的候选体式类别。此外,一种基于间隔的门控机制在保留高置信度视觉预测的同时,仅对模糊的情况调用几何验证。而且,一旦确定了最终的体式类别,SAGE-Yoga会检索一个中心(medoid)参考体式,并将观察到的关节角度与特定类别的分布进行比较,以识别未对齐的关节。最后,将这些偏差转化为可操作的纠正反馈。实验表明,在Yoga-82数据集上,视觉集成模型达到89.0%的Top-1准确率,而完整框架将性能提升至90.7%的Top-1准确率和90.1%的Macro-F1。这些结果表明,将互补的视觉证据与选择性几何验证相结合,不仅能够提升细粒度体式分类性能,还能实现可解释的关节级纠正。
cs.CV / 67 / 2609.20248

Distance to Class Prototypes: Active Learning for Object Detection

到类别原型的距离:面向目标检测的主动学习
Zhang, Licheng, Gong, Zheng
Abstract
Deploying a deep object detector in a new setting is limited less by architecture than by the cost of annotating data from that setting. Active learning lowers the cost by choosing which images to label, and the choice is only as good as the signal used to score an unlabeled image. That signal is usually the class posterior, which is cheap but poorly calibrated, or the disagreement across several models or several stochastic passes, which is better but multiplies inference over a pool far larger than the labeled set. We propose a signal richer than the posterior yet still read from one forward pass of one network. A supervised contrastive term added to the training objective shapes a per-object embedding space in which distance encodes class membership, and an unlabeled detection is scored by how far it lies from the region occupied by its predicted category, weighted by its confidence. The criterion needs no ensemble, no auxiliary predictor and no repeated inference, and its entire cost is 2.89M parameters, an increase of 8.3% over a bare detector. On PASCAL VOC and MS-COCO it beats the posterior of the same detector in every round in which a selection is made, by up to 1.08% mAP50 against run to run deviations of 0.02% to 0.18%, and it stays competitive with ensemble and Monte Carlo dropout criteria costing three to fifty forward passes per unlabeled image. Experiments use the single-stage detector under which the compared criteria report their results, so that the selection decision is isolated from the strength of the detector.
Chinese Translation
在新的场景中部署深度目标检测器,其限制更多来自于该场景下数据标注的成本,而非模型架构本身。主动学习通过选择需要标注的图像来降低成本,而选择的好坏取决于用于给未标注图像打分的信号。该信号通常是类别后验概率,虽然计算廉价但校准性差;或者是多个模型或多次随机前向传播之间的不一致性,效果更好但需要对远大于已标注集的图像池进行多次推理。我们提出一种比后验概率信息更丰富、且仅需单网络单次前向传播即可获得的信号。通过在训练目标中加入监督对比学习项,构建一个逐目标的嵌入空间,其中距离编码了类别归属关系;对未标注检测结果的打分则依据其到其预测类别所占据区域之间的距离,并以置信度加权。该准则无需集成模型、无需辅助预测器、无需重复推理,其全部开销仅为289万参数,相较于裸检测器仅增加8.3%。在PASCAL VOC和MS-COCO数据集上,在每一轮需要进行选择的迭代中,该方法的性能均优于同一检测器的后验概率方法,mAP50最高提升1.08%,而逐次运行偏差仅为0.02%至0.18%;同时该方法与每张未标注图像需要三到五十次前向传播的集成(ensemble)及蒙特卡洛 dropout(Monte Carlo dropout)准则相比仍保持竞争力。实验采用了所比较的各准则报告结果时所使用的单阶段检测器,从而使选择决策与检测器本身的强弱相互独立。
cs.CV / 68 / 2609.20262

Generative Verification: Rethinking the Uncertainty Signal for Active Learning of Object Detection

生成式验证:重新思考目标检测主动学习中的不确定性信号
Zhang, Licheng, Gong, Zheng
Abstract
Nearly every acquisition function for active object detection shares one arrangement, in that the model being improved is also the model being interrogated. We depart from it. In generative verification an independent generative model re-derives the label of a detection from the pixels inside its predicted box, and the disagreement between the two becomes the acquisition signal. Two properties follow from the arrangement itself rather than from any tuning. A displaced box, a box on background and a correct box carrying the wrong label all yield a crop that fails verification, so the failure modes arrive already combined in one scalar and the hand-weighted classification and localization terms of existing criteria are no longer needed. And because the verifier never observes the detector confidence, confidently wrong detections score highest, although a self-derived signal reads them as uninteresting and they are the costliest to leave unlabeled. We build the verifier as a conditional diffusion model whose diffusion target is a label representation rather than an image. Its reverse process is stochastic, so repeated generations return a distribution whose concentration reports how firmly the evidence determines the label, where a classifier returns a single point estimate. On PASCAL VOC and MS-COCO the signal outperforms output-uncertainty, feature-geometry, perturbation and ensemble criteria, gaining about one mAP50 point per round on MS-COCO, with its largest margins in the early rounds where confident detector errors are most common.
Chinese Translation
几乎所有用于主动目标检测的采集函数都共享一种安排:被改进的模型同时也是被查询的模型。我们突破了这一惯例。在生成式验证中,一个独立的生成式模型根据预测框内的像素重新推导该检测的标签,两者之间的分歧即作为采集信号。该安排本身(而非任何调参)带来两个特性。其一,偏移的框、位于背景上的框以及框正确但标签错误的检测,其裁剪区域都无法通过验证,因此各类失败模式已经以组合形式呈现在一个标量中,无需再像现有准则那样对分类项和定位项进行手工加权。其二,由于验证器从不接触检测器的置信度,高置信度的错误检测会获得最高评分,而基于自身的信号会将这些错误视为无趣样本,且将它们留在未标注状态恰恰代价最高。我们将验证器构建为一个条件扩散模型,其扩散目标是标签表示而非图像。其逆向过程是随机的,因此重复生成会返回一个分布,其集中程度反映了证据对标签的确定强度,而分类器只能返回单点估计。在 PASCAL VOC 和 MS-COCO 上,该信号优于输出不确定性、特征几何、扰动和集成等准则,在 MS-COCO 上每轮约提升一个 mAP50 点,且在检测器高置信度错误最常见的前几轮中优势最大。
cs.CV / 69 / 2609.20263

AI or Real: Detecting Partially Altered Videos Under Resource-Constrained Environments

AI还是真实:资源受限环境下部分篡改视频的检测
Chakraborty, Tamoghna, Absur, Md Nurul, Saha, Sourya, Debroy, Saptarshi
Abstract
The proliferation of generative video models has shifted the practical detection threat from fully fabricated clips to partially manipulated footages. Although modern detectors achieve strong accuracy using foundation backbones of 400M+ parameters, their resource footprint precludes edge deployment. In this paper, we present a lightweight full-frame detector for partially manipulated AI-generated video, designed for deployment on edge hardware without face-detection preprocessing. The system distills a DINOv2-Base teacher into a frozen MobileNetV3-Small student through a pipeline that combines temperature-annealed soft-label transfer, attention-diversity regularization, frame-level supervision, and a residual feature adapter that conditions ImageNet features for artifact detection. We additionally target two failure modes specific to the partial-manipulation regime: false positives on legitimate scene cuts, addressed through within-video temporal hard negatives; and threshold-level miscalibration on the dominant pure-real class, addressed through calibration-aware sampling. Evaluation on a 55,393-sample spliced test set across fake-frame ratios from 6.2% to 31.2% demonstrates the student model closing 58% of the gap to the DINOv2-Base teacher (AUC 0.766) while running at 3.65 ms per 16-frame clip on RTX A4000 with a 150.4 MB checkpoint compatible with edge-device memory and latency budgets.
Chinese Translation
生成式视频模型的迅速普及,使实际检测威胁从完全伪造的片段转变为部分篡改的视频素材。尽管现代检测器借助4亿以上参数的基础骨干网络取得了很高的准确率,但其资源开销使其无法部署在边缘设备上。本文提出一种面向部分篡改AI生成视频的轻量级全帧检测器,专为边缘硬件部署设计,且无需人脸检测预处理。该系统通过结合温度退火软标签迁移、注意力多样性正则化、帧级监督以及用于将ImageNet特征调整为伪影检测的残差特征适配器,将DINOv2-Base教师模型蒸馏到冻结的MobileNetV3-Small学生模型中。此外,我们针对部分篡改场景下的两种典型失效模式:一是对合法场景切换的误报,通过视频内时间维度的困难负样本加以解决;二是占主导地位的纯真实类别在阈值层面的校准偏差,通过感知校准的采样策略加以解决。在一个包含55,393个样本、伪造帧比例从6.2%到31.2%的拼接测试集上的评估表明,学生模型弥合了与DINOv2-Base教师模型58%的性能差距(AUC为0.766),同时在RTX A4000上每16帧片段仅需3.65毫秒,且150.4 MB的模型体积符合边缘设备的内存与延迟预算。
cs.CV / 70 / 2609.20267

CleanVideo: Adaptive Concept Erasure for Text-to-Video Diffusion Models

CleanVideo:面向文本到视频扩散模型的自适应概念擦除方法
Liao, Junchi, Li, Hongji, Zhou, Wenrui, Hu, Lijie
Abstract
Concept erasure aims to selectively eliminate undesired visual semantics from pre-trained generative models without compromising their general utility. Extending concept erasure from images to video is nontrivial. Target concepts emerge gradually and vary across frames and denoising steps. As a result, fixed interventions may miss the target or introduce blurring, jitter, and content distortion. We propose CleanVideo, a selective erasure framework that performs low-dimensional subspace intervention controlled by a tri-modal gating mechanism. By jointly processing spatiotemporal visual features, timestep signals, and textual semantics, CleanVideo determines where, when, and whether to intervene, steering erased content toward natural surrogate concepts when such surrogates can be clearly defined while preserving non-target content. Experiments on three video diffusion models show that CleanVideo effectively erases target concepts while maintaining visual fidelity and temporal coherence, outperforming existing baselines under frame-level and video-level evaluations and under concept-recovery attacks when the protected pipeline remains intact.
Chinese Translation
概念擦除旨在选择性地消除预训练生成模型中不需要的视觉语义,同时不损害其通用能力。将概念擦除从图像扩展到视频并非易事:目标概念在帧与去噪步骤之间逐渐显现且不断变化。因此,固定的干预方式可能会错过目标概念,或引入模糊、抖动和内容失真。我们提出了 CleanVideo,一种选择性擦除框架,通过三模态门控机制控制低维子空间干预。通过联合处理时空视觉特征、时间步信号和文本语义,CleanVideo 决定在何处、何时以及是否进行干预,在能够明确定义替代概念时将被擦除内容引导至自然的替代概念,同时保留非目标内容。在三个视频扩散模型上的实验表明,CleanVideo 能够有效擦除目标概念,同时保持视觉保真度和时间一致性,在帧级和视频级评估以及受保护流水线完整时的概念恢复攻击下均优于现有基线方法。
cs.CV / 71 / 2609.20283

Queries Knew More Than We Thought: Uncovering Latent Knowledge in Segmentation Models

查询所知远超我们预期:揭示分割模型中的潜在知识
De la Jara, Ignacio M., Rodriguez-Opazo, Cristian, Ranasinghe, Damith
Abstract
Modern segmenters often fail after the expensive computation has already been done: a useful mask is present among the model's query-conditioned candidates, but the deployed selection rule does not expose it. We study this output-selection bottleneck in frozen DETR-family models. A ground-truth-only oracle first shows substantial hidden headroom in already-computed mask proposals. This raises a simple question: How can we better use the masks a segmenter has already computed but does not expose? We then ask whether that headroom can be recovered without adding queries, generating new masks, rerunning the backbone, or updating weights. HYDRA is a small selector trained only on cached frozen outputs. At inference time, it scores the cached candidates against an explicit keep-baseline option and acts only when a held-out calibrated margin indicates the selected candidate is sufficiently better. Trained on training-split caches and calibrated on held-out data, HYDRA improves Mask2Former, MaskDINO, and OneFormer by up to +7.41 dataset mIoU points on ADE20k and COCO, and improves SAM 3 by +9.4 class-macro prompt-IoU points on average across eight domains while preserving useful predictions through calibration. Paired LoRA controls show that lightweight weight adaptation does not remove the bottleneck: exposed predictions are often flat or worse, while routing over the adapted candidates still recovers accuracy. Finally, we connect the effect to query specialization under bipartite matching and verify it in a controlled TinyDETR study. These results show that frozen segmenters should be evaluated not only by the masks they expose, but also by the useful candidates they suppress.
Chinese Translation
现代分割模型常常在昂贵的计算已经完成后仍会失败:一个有用的掩码其实存在于模型的查询条件化候选结果之中,但部署时的选择规则未能将其暴露出来。我们在冻结的 DETR 系列模型中研究这一输出选择瓶颈。首先,一个仅基于真值的理想选择器表明,在已计算的掩码候选中存在大量隐藏的提升空间。由此引出一个简单的问题:如何更好地利用分割模型已经计算但未暴露的掩码?我们进一步追问,是否可以在不增加查询、不生成新掩码、不重新运行骨干网络、也不更新权重的前提下,恢复这一提升空间。HYDRA 是一个仅基于缓存冻结输出训练的小型选择器。在推理时,它对缓存的候选掩码进行打分,并将其与一个显式的保留基线选项比较,只有当留出校准的边际分数表明所选候选足够优于基线时才采取替换动作。HYDRA 在训练集缓存上训练并在留出数据上校准,使 Mask2Former、MaskDINO 和 OneFormer 在 ADE20k 和 COCO 数据集上的 mIoU 最高提升 +7.41 个百分点,并使 SAM 3 在八个领域上平均提升 +9.4 个类宏平均提示 IoU 点,同时通过校准保留了原有的有效预测。与 LoRA 的配对对照实验表明,轻量级权重适配并不能消除该瓶颈:直接暴露的预测常常持平或更差,而对适配后的候选进行路由选择仍能恢复精度。最后,我们将这一效应与二分匹配下的查询专业化联系起来,并通过受控的 TinyDETR 实验加以验证。这些结果表明,对冻结分割模型的评估不应仅基于其暴露的掩码,还应考虑其抑制的有效候选。
cs.CV / 72 / 2609.20290

TinyCNN: A 193K-Parameter Network for On-Device Plant Disease Detection, with a Cross-Dataset Robustness Diagnosis

TinyCNN:一个用于设备端植物病害检测的193K参数网络,附跨数据集鲁棒性诊断
Ho-Lam, Ngoc-Bao, Nguyen, Thai-Anh
Abstract
Detecting crop disease early is central to sustainable agriculture and food security under United Nations Sustainable Development Goal 2 (Zero Hunger), and is especially urgent in resource-constrained regions where expert diagnosis is scarce but low-cost mobile devices are widespread. This paper presents TinyCNN, a lightweight convolutional neural network for on-device plant disease classification. TinyCNN uses depthwise separable convolution blocks and contains only 193,190 trainable parameters with 110.05M MACs for a 224x224 input image. On the 38-class PlantVillage benchmark, TinyCNN achieves 98.88% test accuracy and 98.03% macro-F1 while being approximately 58x smaller than ResNet18 and 11.8x smaller than a MobileNetV2 teacher, directly reducing the energy, memory, and cost footprint of inference in line with Green AI principles. The paper further analyzes vanilla knowledge distillation as a sustainable model-compression strategy; an ablation over alpha in {0.3, 0.5, 0.7} and T in {2, 4} selects alpha=0.3, T=4, producing a distilled TinyCNN with 98.81% test accuracy. Finally, cross-dataset evaluation from PlantVillage to PlantDoc reveals a substantial robustness gap under real-world conditions, which a Grad-CAM analysis attributes to off-leaf, background-driven attention consistent with shortcut learning. TinyCNN is thus an energy-efficient, deployable building block for sustainable agricultural intelligence, while field robustness remains the key barrier to durable real-world impact.
Chinese Translation
作物病害的早期检测是联合国可持续发展目标2(零饥饿)下可持续农业和粮食安全的核心,在资源受限地区尤为紧迫——这些地区缺乏专家诊断,但低成本移动设备已广泛普及。本文提出了TinyCNN,一种用于设备端植物病害分类的轻量级卷积神经网络。TinyCNN采用深度可分离卷积模块,仅包含193,190个可训练参数,对224x224输入图像的计算量为110.05M MACs。在包含38个类别的PlantVillage基准数据集上,TinyCNN达到98.88%的测试准确率和98.03%的宏F1值,同时其规模约为ResNet18的1/58、MobileNetV2教师模型的1/11.8,直接降低了推理的能耗、内存和成本足迹,符合绿色AI(Green AI)原则。本文进一步将朴素知识蒸馏分析为一种可持续的模型压缩策略;通过对alpha∈{0.3, 0.5, 0.7}和T∈{2, 4}的消融实验,选定alpha=0.3、T=4,得到测试准确率为98.81%的蒸馏版TinyCNN。最后,从PlantVillage到PlantDoc的跨数据集评估揭示了在真实世界条件下的显著鲁棒性差距,Grad-CAM分析将其归因于偏离叶片、由背景驱动的注意力机制,与捷径学习(shortcut learning)现象一致。因此,TinyCNN是一种节能、可部署的可持续农业智能基础构件,而现场鲁棒性仍是实现持久现实影响的关键障碍。
cs.CV / 73 / 2609.20299

Not All Layers Are Equal: Dynamic Layer Routing for Reliable CLIP OOD Detection

并非所有层都同等重要:用于可靠CLIP分布外检测的动态层路由方法
De la Jara, Ignacio M., Rodriguez-Opazo, Cristian, Ranasinghe, Damith
Abstract
Information aggregation across model layers are revealed to improve OOD detection. In contrast to crafting a method for layer-wise information aggregation in recent work, we investigate if layer selection is a learnable problem. In other words, we transpose the question from how to fuse layers to one asking which layers to trust for an input. Using a generalizable, weak, out of distribution context crafting approach for supervision, shown to be more effective than state of the art methods' mechanisms, we formulate learning a lightweight router to select a sparse, final-layer-anchored expert over CLIP's layer depth for OOD detection. Across three diverse benchmarks we demonstrate our learnable routing method dubbed Voyager improves OOD detection. On ImageNet-1K, Voyager achieves an average FPR@95 of 18.86, outperforming the strongest, comparable, prompt-learning method by 8.8 points. These gains persist across multiple supervision sources, including those used by existing state-of-the-art prompt-learning methods, demonstrating that, whilst our weak OOD supervision context is highly effective, the key advantage is realized from the learnable router component rather than the supervision source. Significantly, Voyager is highly practical; router learning takes approximately two minutes using less than 1 GB of memory, making it approximately 20x more efficient than current prompt-learning approaches. Anonymized Code: https://anonymous.4open.science/r/Voyager/
Chinese Translation
跨模型层的信息聚合被证明可以改善分布外(OOD)检测。与近期工作构造逐层信息聚合方法不同,我们研究层选择是否是一个可学习的问题。换言之,我们将问题从“如何融合各层”转换为“对于给定输入应信任哪些层”。通过一种可泛化的、弱监督的分布外上下文构造方法进行监督(该方法已被证明比现有最先进方法的机制更有效),我们提出学习一个轻量级路由器,在CLIP的层深度上选择一个稀疏的、以最后一层为锚点的专家,用于OOD检测。在三个不同的基准上,我们证明了这种可学习路由方法(命名为Voyager)能够改善OOD检测。在ImageNet-1K上,Voyager实现了18.86的平均FPR@95,比最强的可比较提示学习方法高出8.8个百分点。这些增益在多种监督来源下均得以保持,包括现有最先进提示学习方法所使用的监督来源,这表明虽然我们的弱OOD监督上下文非常有效,但关键优势来自可学习的路由器组件而非监督来源。值得注意的是,Voyager具有很高的实用性:路由器学习仅需约两分钟且占用不到1 GB的内存,效率比当前的提示学习方法高约20倍。匿名代码:https://anonymous.4open.science/r/Voyager/
cs.CV / 74 / 2609.20325

AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images

AgriScope:面向农业图像的像素级多模态理解
Boudiaf, Abderrahmene, Alanssari, Mohamad, Hussain, Irfan, Javed, Sajid
Abstract
Agricultural image understanding requires fine-grained recognition of plant diseases, pests, crop structures, and botanical species under complex real-world conditions. Despite recent advances in Multimodal Large Language Models (MLLMs), existing models remain limited to text-only outputs and lack pixel-level visual grounding capabilities. In this work, we introduce AgriScope, a unified pixel-grounded multimodal framework for agricultural image understanding. AgriScope jointly supports image-level, region-level, and pixel-level understanding within a unified framework, enabling tasks such as grounded caption generation, referring expression segmentation, and multi-turn multimodal interaction for agricultural imagery. AgriScope integrates biologically specialized semantic representations with dense spatial grounding through biological-semantic encoding, dense spatial representations, and pixel decoding. To support large-scale grounded learning, we introduce AgriGround, a large-scale pixel-grounded agricultural multimodal instruction-tuning dataset containing over 500K images and 11M instruction-following samples spanning plant disease analysis, crop and weed identification, insect pest recognition, and fine-grained botanical understanding. AgriGround is constructed through a multi-stage automatic annotation pipeline that integrates multimodal caption generation, phrase-level grounding, segmentation mask generation, and task-oriented instruction synthesis to produce densely grounded supervision. Extensive experiments across multiple agricultural vision-language tasks demonstrate the effectiveness of AgriScope in pixel-grounded multimodal understanding, establishing a strong benchmark for agricultural vision-language learning and visual grounding. The dataset and code will be made publicly available at (https://github.com/boudiafA/AgriScope)
Chinese Translation
农业图像理解需要在复杂真实世界条件下对植物病害、害虫、作物结构和植物物种进行细粒度识别。尽管多模态大语言模型(MLLMs)近期取得了显著进展,但现有模型仍局限于纯文本输出,缺乏像素级的视觉定位能力。在本工作中,我们提出了AgriScope,一个面向农业图像理解的统一像素级多模态框架。AgriScope在统一框架内同时支持图像级、区域级和像素级理解,能够实现诸如定位描述生成、指代表达分割以及面向农业图像的多轮多模态交互等任务。AgriScope通过生物语义编码、密集空间表示和像素解码,将生物学专用的语义表示与密集空间定位相融合。为支持大规模定位学习,我们构建了AgriGround,一个大规模像素级定位的农业多模态指令微调数据集,包含超过50万张图像和1100万个指令跟随样本,涵盖植物病害分析、作物与杂草识别、害虫识别以及细粒度植物学理解。AgriGround通过多阶段自动标注流水线构建,该流水线集成了多模态描述生成、短语级定位、分割掩码生成和面向任务的指令合成,以生成密集的定位监督信号。在多个农业视觉-语言任务上的大量实验表明,AgriScope在像素级定位多模态理解方面的有效性,为农业视觉-语言学习和视觉定位建立了强有力的基准。数据集和代码将在(https://github.com/boudiafA/AgriScope)公开发布。
cs.CV / 75 / 2609.20340

FreqDINO++: A Frequency-Guided Multi-Task Routing Vision Foundation Model for Universal Ultrasound Analysis

FreqDINO++:一种频率引导的多任务路由视觉基础模型,用于通用超声分析
Xu, Qing, Zhang, Yixuan, Li, Yue, He, Xiangjian, Zhang, Qian, Haque, Mainul, Qu, Rong, Duan, Wenting, Bai, Jieyun, Chen, Zhen
Abstract
Ultrasound image analysis plays a crucial role in cancer screening and prenatal diagnosis, yet comprehensive assessment requires jointly addressing tasks such as lesion segmentation and benign-malignant classification. While recent vision foundation models have shown remarkable universal representations, unlocking their potential for ultrasound is bottlenecked by the considerable domain gap from natural images. Existing methods typically fine-tune heavy vision encoders for isolated tasks, incurring substantial computational overhead while overlooking the underlying commonalities across heterogeneous tasks. In this work, we propose FreqDINO++, a frequency-guided multi-task routing vision foundation model for universal ultrasound analysis. We first introduce a Multi-task Routing Adapter (MR-Adapter) to support parameter-efficient integration of task-common and task-specific knowledge, a Frequency-aware Feature Enhancer (F$^2$-Enhancer) is then designed to capture the rich multi-scale frequency characteristics of ultrasound images, and a Task-aligned Collaborative Decoder (TC-Decoder) is devised to promote collaboration between dense and global prediction tasks through global-local token interaction. Extensive experiments on large-scale multi-task and external single-task ultrasound benchmarks demonstrate that FreqDINO++ consistently outperforms strong baselines and recent foundation models across 27 diverse clinical task scenarios, while also showing promising generalization to unseen data. The code is at https://github.com/MingLang-FD/FreqDINO-Plus.
Chinese Translation
超声图像分析在癌症筛查和产前诊断中发挥着至关重要的作用,然而全面的评估需要同时处理病灶分割和良恶性分类等多种任务。尽管近期的视觉基础模型展现出卓越的通用表征能力,但其应用于超声领域的潜力受到自然图像与超声图像之间巨大领域差距的制约。现有方法通常针对孤立任务微调大型视觉编码器,这不仅带来了高昂的计算开销,还忽略了异构任务之间潜在的共性。在本工作中,我们提出了FreqDINO++,一种面向通用超声分析的频率引导多任务路由视觉基础模型。我们首先引入多任务路由适配器(Multi-task Routing Adapter, MR-Adapter),以参数高效的方式整合任务共性与任务特定知识;随后设计了频率感知特征增强器(Frequency-aware Feature Enhancer, F$^2$-Enhancer),用于捕捉超声图像丰富的多尺度频率特征;并设计了任务对齐协同解码器(Task-aligned Collaborative Decoder, TC-Decoder),通过全局-局部token交互促进稠密预测任务与全局预测任务之间的协作。在大规模多任务和外部单任务超声基准上的大量实验表明,FreqDINO++在27种不同的临床任务场景中始终优于强基线方法和近期的基础模型,同时对未见数据也展现出良好的泛化能力。代码位于 https://github.com/MingLang-FD/FreqDINO-Plus。
cs.CV / 76 / 2609.20341

Fast Cross-Strength Multi-Contrast Brain MRI Translation using Latent Bridge Matching

基于潜在桥匹配的快速跨场强多对比度脑部MRI翻译
Srivastava, Siddharth, Bretschneider, Till
Abstract
Magnetic Resonance Imaging (MRI) acquired at different field strengths exhibits pronounced variation in noise, resolution, homogeneity, and contrast, which limits comparability across acquisition settings and complicates downstream analysis. We address this with a unified conditional model for controllable field-to-field synthesis, built on the framework of conditional latent bridge matching. Our single model achieves highly competitive results across the validation phase for all three tasks of the MRIxFields2026 challenge without task-specific architectures or training. We achieve fast generation with only a single inference step, producing all modality and field-strength combinations for $30$ axial slices in under $90$ seconds, as well as cross-modality-strength translation for a full volume in under $70$ seconds, on a single NVIDIA A5000 GPU. We further provide extensive ablations regarding different components of our solution. Code: https://gitlab.com/siddharthsrivastava/mrixfields-2026
Chinese Translation
在不同场强下采集的磁共振成像(MRI)在噪声、分辨率、均匀性和对比度方面表现出显著差异,这限制了不同采集设置之间的可比性,并使下游分析变得复杂。我们基于条件潜在桥匹配(conditional latent bridge matching)框架,构建了一个用于可控场强间合成的统一条件模型。在无需任务专用架构或训练的情况下,我们的单一模型在MRIxFields2026挑战赛验证阶段的全部三项任务中均取得了极具竞争力的结果。我们仅需单步推理即可实现快速生成:在单块NVIDIA A5000 GPU上,90秒内即可生成30个轴位切片的所有模态与场强组合,70秒内即可完成整个体积的跨模态-场强翻译。此外,我们针对解决方案的不同组件进行了大量消融实验。代码:https://gitlab.com/siddharthsrivastava/mrixfields-2026
cs.CV / 77 / 2609.20348

EliGSiR: Continual RGB-D Mapping with Gaussian Splatting under Bounded Compute

EliGSiR:有限计算资源下基于高斯泼溅的持续RGB-D建图
Ellensohn, Björn, Rueckert, Elmar, Rauch, Christian
Abstract
Conventional 3D Gaussian Splatting assumes a closed set of observations and long optimization schedules. Continual RGB-D mapping in contrast poses the problem that new observations arrive online, while previously reconstructed regions must be preserved. We present EliGSiR (Evidence-guided Load-adaptive Incremental Gaussian Splatting with Image Replay), a continual Gaussian mapper that controls how the available optimization budget is used as the reconstruction evolves. Map-Guided View Scheduling filters redundant incoming views and reconsiders retained views according to the current state of the map. Load-Adaptive Fidelity adjusts supervision resolution to the current mapping load instead of following a fixed resolution schedule. Targeted Geometry Growth separates depth supervision from Gaussian creation and adds geometric capacity only where repeated RGB-D observations indicate missing or misplaced structure. Together, these mechanisms adapt which views are optimized, how much image detail is used, and where the representation grows while mapping remains active. We evaluate EliGSiR on Replica, TUM RGB-D, ScanNet++, and real RGB-D sensor sequences, considering both the final reconstruction and the map available throughout acquisition. On TUM RGB-D fr3/long_office_household, EliGSiR reaches 21.52 dB with the same ground-truth mapping poses used by the controlled baselines, compared with 19.42 dB for SplaTAM. In the tracked-pose comparison, EliGSiR with live ORB-SLAM3 poses reaches 23.02 dB in 155.5 s, compared with 20.10 dB in 230.9 s for CaRtGS using its native tracker. We further evaluate reconstruction throughout acquisition and show how EliGSiR adaptive view scheduling, supervision fidelity, and geometry growth improve the use of the available mapping budget.
Chinese Translation
传统三维高斯泼溅(3D Gaussian Splatting)假设观测数据是封闭集合,并且可以进行长时间的优化。而持续RGB-D建图所面临的问题是:新观测在线到达,同时必须保留先前已重建的区域。我们提出了EliGSiR(基于证据引导的负载自适应图像回放增量式高斯泼溅),这是一种持续高斯建图方法,能够根据重建过程的演进控制可用优化预算的使用方式。地图引导的视角调度(Map-Guided View Scheduling)过滤冗余的输入视角,并根据地图当前状态重新考虑保留的视角。负载自适应保真度(Load-Adaptive Fidelity)根据当前建图负载调整监督分辨率,而非遵循固定的分辨率调度。有针对性的几何增长(Targeted Geometry Growth)将深度监督与高斯基元的创建分离,仅在重复RGB-D观测表明存在缺失或位置错误的结构时才增加几何容量。这些机制共同作用,在建图持续进行的同时,自适应地决定优化哪些视角、使用多少图像细节以及表示在哪里增长。我们在Replica、TUM RGB-D、ScanNet++以及真实RGB-D传感器序列上评估了EliGSiR,同时考虑了最终重建结果以及采集中可用的地图。在TUM RGB-D fr3/long_office_household序列上,使用与受控基线相同的真值建图位姿,EliGSiR达到21.52 dB,而SplaTAM为19.42 dB。在跟踪位姿对比中,采用实时ORB-SLAM3位姿的EliGSiR在155.5秒内达到23.02 dB,而使用其自带跟踪器的CaRtGS在230.9秒内为20.10 dB。我们还进一步评估了采集中各个阶段的重建质量,展示了EliGSiR的自适应视角调度、监督保真度和几何增长如何改进对可用建图预算的利用。
cs.CV / 78 / 2609.20377

MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

MM-Future:面向自动驾驶的多模式联合世界-动作建模
Liu, Shuai, Gong, Hechangle, Jiang, Hao, He, Runlin, Zhan, Junxiang, Huang, Kai, Yang, Sheng, Ren, Shaoqing
Abstract
Autonomous driving involves coupled decision-making and scene evolution under multi-mode uncertainty. To capture this coupling and uncertainty, we introduce MM-Future, a world-action model that generates multiple paired scene-action hypotheses and models bidirectional interaction within each pair. Each hypothesis is initialized from a structured action prior and an independent future scene source, which are then co-evolved through a modality-aware diffusion Transformer. To support efficient multi-mode rollout, MM-Future compresses multi-view video into planning-oriented representations, dubbed MM-Tokens. Finally, a future-conditioned proposal scorer ranks trajectory candidates by shared history context and their paired predicted future. On NAVSIM navtest, MM-Future achieves 94.0 PDMS and 91.5 EPDMS, while attaining a 32.3 HD-Score in zero-shot closed-loop evaluation on HUGSIM. Ablations show consistent improvements over both single-mode and action-only variants, validating the benefit of multi-mode joint world-action modeling.
Chinese Translation
自动驾驶涉及多模式不确定性下耦合的决策与场景演化。为了捕捉这种耦合性与不确定性,我们提出了MM-Future,一种能够生成多个成对的场景-动作假设并对每对内双向交互进行建模的世界-动作模型。每个假设由结构化动作先验和一个独立的未来场景源初始化,随后通过模态感知的扩散Transformer(diffusion Transformer)进行协同演化。为支持高效的多模式推演,MM-Future将多视角视频压缩为面向规划的表示,称为MM-Tokens。最后,一个未来条件化的候选提案评分器基于共享的历史上下文及其配对的预测未来对轨迹候选进行排序。在NAVSIM navtest上,MM-Future达到了94.0 PDMS和91.5 EPDMS,并在HUGSIM上零样本闭环评估中取得了32.3的HD-Score。消融实验表明,该方法相较于单模式和仅动作的变体均取得了一致的改进,验证了多模式联合世界-动作建模的有效性。
cs.CV / 79 / 2609.20386

Compact Vision Models for Iris Presentation Attack Detection under Presentation Attack Instrument Shift and Environmental Degradation

面向呈现攻击仪器变化与环境退化的紧凑型虹膜呈现攻击检测视觉模型
Angelakis, Athanasios, Gomez-Barrero, Marta
Abstract
Iris presentation attack detection (PAD) is security-critical when a subsystem that appears reliable during development encounters presentation attack instruments (PAIs) or acquisition conditions absent from validation data. We benchmark three compact scratch-trained computer-vision models, each with at most approximately 0.26 million trainable parameters, on the Notre Dame subset of LivDet-Iris 2017 under PAI-driven domain shift and environmental degradation. All models are trained without external pretraining or data augmentation and evaluated over five seeds. A validation-selected threshold is transferred unchanged to the known-attack, unknown-attack, corrupted, and pooled test partitions. From known to unknown attack presentations, Attack Presentation Classification Error Rate (APCER) increases by 17.11-30.47 percentage points and Detection Equal Error Rate (D-EER) increases by 7.38-12.73 percentage points. At the validation-selected threshold, ZACH-ViT obtains the lowest unknown-attack APCER (47.69 +/- 4.84%) and D-EER (38.87 +/- 0.93%), while Compact-TransMIL obtains the lowest Bona Fide Presentation Classification Error Rate (BPCER). ZACH-ViT also gives the lowest unknown-attack BPCER at an APCER limit of 10% (81.29 +/- 1.95%). The high absolute errors show that the comparative advantage of the best compact model does not constitute deployment readiness under unknown PAIs.
Chinese Translation
当一个在开发阶段看似可靠的子系统遇到验证数据中未出现的呈现攻击仪器(PAI)或采集条件时,虹膜呈现攻击检测(PAD)的安全性至关重要。我们在LivDet-Iris 2017的Notre Dame子集上,针对PAI驱动的领域偏移和环境退化,对三个紧凑的从零训练的计算机视觉模型进行了基准测试,每个模型的可训练参数不超过约26万。所有模型均在不使用外部预训练或数据增强的情况下训练,并在五个随机种子上进行评估。通过验证集选择的阈值未经更改地迁移到已知攻击、未知攻击、退化以及合并测试分区。从已知攻击到未知攻击呈现,攻击呈现分类错误率(APCER)增加了17.11-30.47个百分点,检测等错误率(D-EER)增加了7.38-12.73个百分点。在验证集选择的阈值下,ZACH-ViT取得了最低的未知攻击APCER(47.69 ± 4.84%)和D-EER(38.87 ± 0.93%),而Compact-TransMIL取得了最低的真实呈现分类错误率(BPCER)。在APCER上限为10%的条件下,ZACH-ViT同样取得了最低的未知攻击BPCER(81.29 ± 1.95%)。较高的绝对误差表明,最优紧凑模型的比较优势并不代表其在未知PAI条件下具备部署就绪性。
cs.CV / 80 / 2609.20414

TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation

TouchSight:基于生成式视觉增强的第一视角视频裸手触觉预测
Zhou, Danyan, Lu, Jinxuan, Lin, Jiawei, Chen, Tianxing, Lyu, Chuqiao, Ding, Wenbo
Abstract
Tactile signals provide direct contact and force measurements that are essential for understanding physical interactions and enabling dexterous robotic manipulation. However, tactile sensing requires direct measurement at contact interfaces, making large-scale data collection reliant on intrusive, costly, and restrictive instrumentation. We present TouchSight, a monocular egocentric vision framework for dense full-hand contact force prediction that leverages 500 hours of pressure-glove recordings and extensive hand-object interaction (HOI) data. To address the appearance gap between gloved training data and bare-hand real-world scenarios, we construct TwinTouch-20H: 20 hours of paired visual data in which generative video models re-render gloved recordings as bare-hand observations against new backgrounds while preserving the original measured tactile labels. TouchSight predicts dense force from both gloved and generated bare-hand videos, outperforms prior contact prediction methods on OakInk2, qualitatively generalizes to natural bare-hand egocentric videos from unseen datasets, and improves consistently as glove supervision scales. These results demonstrate that dense tactile signals can be recovered from egocentric vision alone, without tactile instrumentation at capture time.
Chinese Translation
触觉信号能够提供直接的接触和力测量,这对于理解物理交互以及实现灵巧的机器人操作至关重要。然而,触觉感知需要在接触界面处进行直接测量,使得大规模数据采集依赖于侵入式、昂贵且受限的仪器设备。我们提出了TouchSight,这是一个用于稠密全手接触力预测的单目第一视角(egocentric)视觉框架,该框架利用了500小时的压力手套记录数据以及大量手-物体交互(HOI)数据。为了解决戴手套训练数据与裸手真实场景之间的外观差异,我们构建了TwinTouch-20H:包含20小时的成对视觉数据,其中生成式视频模型将戴手套的记录重新渲染为在新背景下的裸手观测,同时保留原始测量的触觉标签。TouchSight能够从戴手套视频和生成的裸手视频中预测稠密的力分布,在OakInk2数据集上的表现优于现有的接触预测方法,能够定性地泛化到来自未见数据集的自然裸手第一视角视频,并且随着手套监督数据规模的增加而持续提升。这些结果表明,稠密触觉信号可以仅通过第一视角视觉来恢复,而无需在采集时使用触觉仪器设备。
cs.CV / 81 / 2609.20423

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

WeVisDoc:从覆盖率到能力,实现稳健的端到端文档解析
Yu, Hao, Liu, Kang, Zhao, Linnan, Zhan, Jiabo, Sun, Chong, Li, Chen, Lyu, Jing
Abstract
Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address a parser's remaining weaknesses. We present WeVisDoc, a two-stage data-centric framework for robust end-to-end document parsing. Stage I broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis. Stage II uses a held-out probe to measure the Stage I parser's residual errors within fixed visual-structural clusters. These diagnostics guide targeted data construction and reallocation of the target-token budget. WeVisDoc-4B achieves an Overall score of 95.38 on OmniDocBench v1.6 and a mean Overall score of 75.54 across the three PureDocBench tracks, ranking first among the compared end-to-end parsers in all four settings. Compared with Stage I, Stage II improves Overall scores for the 2B and 4B models on both benchmarks, with larger gains on the degraded PureDocBench tracks, including a 4.03-point gain for the 4B model on the Real Degraded track.
Chinese Translation
文档解析旨在将文档图像转换为结构化内容,并要求在多样化的版式与采集条件下具备可靠性能。然而,训练语料偏向常见文档类型和干净的数字页面,且仅扩大覆盖率并不能说明如何解决解析器剩余的薄弱环节。我们提出了WeVisDoc,一个面向稳健端到端文档解析的两阶段数据中心框架。第一阶段通过异构数据和保结构的退化合成,拓展语义、结构和外观方面的覆盖率。第二阶段使用预留的探测集(held-out probe)在固定的视觉-结构聚类内度量第一阶段解析器的残余误差。这些诊断结果指导针对性的数据构建以及目标词元(target-token)预算的重新分配。WeVisDoc-4B在OmniDocBench v1.6上取得了95.38的总分(Overall),并在PureDocBench三条测试轨道上取得75.54的平均总分,在所有四个设置中均位列所比较的端到端解析器之首。与第一阶段相比,第二阶段提升了2B和4B模型在两个基准上的总分,且在退化的PureDocBench轨道上提升更为显著,其中4B模型在Real Degraded轨道上获得了4.03分的提升。
cs.CV / 82 / 2609.20427

When Do Language-Grounded Explanations Help? A Graph-Bottleneck for Farm Monitoring Interpretable Sheep Facial Pain

基于语言的解释何时有效?用于农场监测可解释绵羊面部疼痛识别的图瓶颈方法
Noor, Alam, Gait'an, Miguel Guti'errez
Abstract
Automated pain recognition from facial expression could make continuous welfare assessment practical in sheep, but adoption depends on trust: a stockperson cannot act on a score that arrives without justification. We ground a model in the Sheep Pain Facial Expression Scale (SPFES) by letting each detected facial region attend over text embeddings of the clinical descriptors and then test whether the resulting explanations mean anything. They do not. Ablating an entire descriptor changes the predicted logit by about $10^{-4}$, and the most-attended cue agrees with the predicted pain level in only $32.6\%$ of regions, although the attention maps, the learned gate, and the generated text all proposed otherwise. We therefore remove the appearance bypass with a concept bottleneck whose classifier reads only SPFES concept scores, supervised by per-region state annotations that image-level pipelines discard. This costs $0.05$--$0.10$ in Cohen's $\kappa$ but yields concepts that are demonstrably learned: minority pain-indicating states are recovered at $3.5$--$8.3\times$ their base rates, and the ear and eye severity orderings emerge without severity supervision. Removing the supervision alone leaves $\kappa$ unchanged while concept accuracy falls to $0.109$, showing that architectural necessity does not imply semantic validity. We also show that pooled concept accuracy is misleading under clinical imbalance and provide a cross-validated, protocol-matched benchmark of seven methods on this dataset.
Chinese Translation
从面部表情自动识别疼痛可以使对绵羊的持续性福利评估变得切实可行,但其实际应用取决于信任:牧民无法在没有依据的情况下依据一个评分采取行动。我们以绵羊疼痛面部表情量表(Sheep Pain Facial Expression Scale, SPFES)为基础构建模型,让每个检测到的面部区域关注临床描述文本的文本嵌入,然后检验由此产生的解释是否具有实际意义。结果表明并无意义:移除某个完整描述符仅使预测logit变化约$10^{-4}$,且注意力最高的线索仅在32.6%的区域中与预测疼痛水平一致,尽管注意力图、学习到的门控机制和生成的文本都提出了相反的暗示。因此,我们通过概念瓶颈(concept bottleneck)消除了外观旁路,其分类器仅读取SPFES概念得分,并由图像级流程通常丢弃的逐区域状态标注进行监督。这使Cohen's $\kappa$降低了0.05--0.10,但得到了可证明已被学习的概念:少数疼痛指示状态以其基率的3.5--8.3倍被识别出来,且耳部和眼部的严重程度排序在无严重程度监督的情况下自发涌现。仅移除监督时,$\kappa$保持不变,而概念准确率降至0.109,这表明架构上的必要性并不意味着语义上的有效性。我们还证明,在临床类别不平衡的情况下,汇总的概念准确率具有误导性,并在该数据集上提供了经交叉验证、协议匹配的七种方法基准测试。
cs.CV / 83 / 2609.20441

Cross-Architecture Foundation-Model Distillation for Edge Flood Segmentation

面向边缘洪水分割的跨架构基础模型蒸馏
Schmalstieg, Fabian, Mueller, Karsten, Samek, Wojciech
Abstract
Geospatial foundation models can provide strong flood-segmentation performance, but their size limits deployment on memory-constrained edge hardware. We distill a 300-million-parameter Prithvi-EO-2.0 teacher, fine-tuned on the 252 manually labeled Sen1Floods11 training scenes, into a 0.7-million-parameter EfficientViT-B0 student. The teacher supervises additional unlabeled Sentinel-2 imagery, allowing the student training set to grow without new manual annotations. At the matched budget of 252 scenes, teacher-supervised training is competitive with direct training and improves STURM-Flood performance across tested configurations; a geometry-matched control shows that label source alone does not explain the difference. Scaling the teacher-supervised pool to 2,500 scenes narrows the remaining student--teacher gap: the float student reaches 0.787 water intersection over union on the Sen1Floods11 test split against 0.822 for the teacher, matches the teacher on STURM-Flood under our evaluation protocol, and remains below it on WorldFloods-v2. After activation replacement and quantization-aware training, the student runs as a 1.5-megabyte 8-bit integer (INT8) TensorRT engine on a Jetson Xavier NX at 5.57 milliseconds of graphics processing unit (GPU) compute per 512-by-512 image, with approximately 14 megabytes of runtime device memory. A fixed modified normalized difference water index (MNDWI) threshold is competitive with both models on the two clean external benchmarks, so we interpret those benchmarks as generalization tests rather than as evidence of learned-model superiority over a spectral rule. The results support the conclusion: foundation-model supervision can amplify a fixed manual annotation budget into a substantially larger training set and yield a compact, deployable edge model.
Chinese Translation
地理空间基础模型能够提供强大的洪水分割性能,但其模型规模限制了在内存受限的边缘硬件上的部署。我们将一个在252个经人工标注的Sen1Floods11训练场景上微调过的3亿参数Prithvi-EO-2.0教师模型,蒸馏为0.7百万(70万)参数的EfficientViT-B0学生模型。教师模型对额外的无标注Sentinel-2影像进行监督,使学生的训练集无需新的人工标注即可扩展。在252个场景的同等数据预算下,教师监督训练与直接训练具有竞争力,并在所测试的配置中提升了STURM-Flood的性能;几何匹配的对照实验表明,仅标签来源并不能解释这一差异。将教师监督数据池扩展至2500个场景后,学生与教师之间剩余的差距被缩小:在Sen1Floods11测试集上,浮点学生模型的水体交并比(IoU)达到0.787,而教师模型为0.822;在我们的评估协议下,学生在STURM-Flood上与教师持平,但在WorldFloods-v2上仍低于教师。经过激活替换和量化感知训练后,该学生模型可作为一个1.5兆字节的8位整数(INT8)TensorRT引擎,在Jetson Xavier NX上以每张512×512图像5.57毫秒的图形处理器(GPU)计算时间运行,运行时设备内存占用约为14兆字节。在两个干净的外部基准上,固定的改进归一化差异水体指数(MNDWI)阈值方法与这两个模型均具有竞争力,因此我们将这些基准视为泛化性测试,而非学习型模型优于光谱规则的证据。上述结果支持如下结论:基础模型监督可以将固定的人工标注预算放大为规模大得多的训练集,并产出紧凑、可部署的边缘模型。
cs.CV / 84 / 2609.20475

SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation

SenseFuse:用于开放词汇3D实例分割的无标签图像与形状编码器融合方法
Han, Euiseok, Ton, Tri, Kim, Hwanhee, Ryu, Seungyeon, Yoo, Chang D.
Abstract
Open-vocabulary scene understanding is fundamental for robotics, laying the groundwork for spatial reasoning and object manipulation. While closed-vocabulary 3D instance segmentation heavily leverages 3D shape information, state-of-the-art open-vocabulary methods remain predominantly restricted to 2D image features or image-distilled representations during mask labeling. In this paper, we propose SenseFuse, a label-free fusion method that balances 2D image and 3D shape encoders for robust open-vocabulary 3D instance segmentation, refining only the mask-labeling stage of existing pipelines. We reveal that 2D image and 3D shape encoders exhibit largely disjoint failure patterns and rarely share identical wrong labels, whereas two 2D image encoders frequently repeat the same errors. This distinct behavior makes the 2D and 3D pair inherently complementary. We introduce an adaptive mechanism that selects a scene-level fusion weight to maximize a label-free sensitivity measure, estimated directly from a single scene's unlabeled proposals in milliseconds. SenseFuse improves labeling accuracy in every evaluated setting across ScanNet200, Replica, and ScanNet++, recovering 67-100% (median 93%) of the gain achievable with an oracle weight, and it raises instance AP in 21 of 22 reported settings. Code is available at https://github.com/hanes1207/SenseFuse.
Chinese Translation
开放词汇场景理解是机器人技术的基础,为空间推理与物体操作奠定基础。尽管封闭词汇的3D实例分割大量利用3D形状信息,但最先进的开放词汇方法在掩码标注阶段仍主要局限于2D图像特征或由图像蒸馏得到的表示。本文提出SenseFuse,一种无标签的融合方法,它平衡2D图像编码器与3D形状编码器,以实现鲁棒的开放词汇3D实例分割,且仅对现有流程中的掩码标注阶段进行优化。我们发现,2D图像编码器与3D形状编码器的失败模式在很大程度上互不相交,很少产生相同的错误标签;而两个2D图像编码器则经常重复相同的错误。这种独特的行为使得2D与3D编码器的组合具有天然的互补性。我们引入一种自适应机制,通过选择场景级融合权重来最大化一个无标签的敏感度度量,该度量可直接在毫秒级时间内从单个场景的无标注提案中估计得出。SenseFuse在ScanNet200、Replica和ScanNet++的所有评估设置中均提升了标注准确率,恢复了使用预言机(oracle)权重可获得收益的67-100%(中位数为93%),并在22个报告设置中的21个提高了实例AP。代码发布于 https://github.com/hanes1207/SenseFuse。
cs.CV / 85 / 2609.20508

Grounded Product Understanding in Livestream Videos

直播视频中的落地化产品理解
Zhang, Xinyu, Chen, Junjie, Ge, Jiawei, Li, Qianlong, Ma, Libin, Pan, Baokun, Luo, Yahui
Abstract
E-commerce livestreams have emerged as an important channel for presenting products to online consumers, containing multiple products whose information is scattered in different moments. This poses significant challenges for downstream product understanding applications, such as product-centric livestream clipping, where models need to identify the product and its relevant segments for information gathering. However, existing benchmarks for general product understanding typically evaluate product retrieval and temporal localization in isolation, leaving the critical correspondence between product identity and temporal evidence largely unassessed. To address this limitation, we introduce GPUB, a large-scale benchmark comprising 3,000 livestream instances with quality-controlled multi-moment temporal annotations and a catalog of over 31K fashion products. GPUB supports three evaluation tasks: the main task Grounded Product Understanding (GPrU) requires jointly identifying the target product and localizing its supporting moments from a livestream video and a candidate product set; Product Retrieval and Product Moment Localization serve as two complementary subtasks. Evaluation of existing multimodal models shows that GPrU remains highly challenging, with the best-performing baseline achieving only 10.13% Pair mAP@.3. To narrow the performance gap, we further develop UniPro, a unified product understanding model that derives product-aligned and temporally structured representations from shared multimodal encoding, improving Pair mAP@.3 to 21.53% while achieving 37.23% Joint R@1@.3 on GPrU.
Chinese Translation
电商直播已成为向在线消费者展示产品的重要渠道,其中包含多个产品,且产品信息分散于不同时间段。这给下游产品理解应用带来了重大挑战,例如以产品为中心的直播剪辑,模型需要识别产品并定位其相关片段以进行信息收集。然而,现有的通用产品理解基准通常将产品检索和时间定位分开评估,未能充分考察产品身份与时间证据之间的关键对应关系。为弥补这一不足,我们提出了GPUB,一个大规模基准数据集,包含3,000个直播实例、经过质量控制的多样本时间标注,以及超过31,000个时尚产品的产品目录。GPUB支持三项评估任务:主任务落地化产品理解要求在给定直播视频和候选产品集的情况下,联合识别目标产品并定位其支撑时刻;产品检索与产品时刻定位作为两个互补的子任务。对现有多模态模型的评估表明,GPrU任务仍极具挑战性,表现最佳的基线模型仅达到10.13%的Pair mAP@.3。为缩小性能差距,我们进一步开发了UniPro,一个统一的产品理解模型,通过共享多模态编码获得产品对齐且具有时间结构的表示,将Pair mAP@.3提升至21.53%,并在GPrU上取得37.23%的Joint R@1@.3。
cs.CV / 86 / 2609.20509

Automated Goldsmith's Mark Retrieval in Silverware

银器上金匠标识的自动检索
Tiwari, Atmik, Christlein, Vincent, Fichtner, Mark, Gohlke, Freya, Schübel, Birgit, Witting, Theresa, Zech, Heike, Zinnen, Mathias
Abstract
For art historians, goldsmith marks play a critical role in the identification and dating of artifacts. In practice, experts must manually compare a query mark against hundreds of documented examples, a process that is both tedious and highly dependent on specialist knowledge. To address this, we present an AI-assisted retrieval pipeline that combines mark localization with metric-learning fine-tuning across three backbone architectures: an ImageNet-pretrained ResNet-50, a supervised ViT-S/16, and a self-supervised DINOv2 ViT-S/14. We conduct a systematic evaluation of cropping strategies, where we measure the impact of no cropping, manual ground-truth cropping, and learned detection-based cropping, and assess their interaction with each backbone. Our strongest configuration, DINOv2 ViT-S/14 with manual crop and metric-learning fine-tuning, achieves an mAP of 62.63% and a Top-1 accuracy of 73.74%. Our experiments show that self-supervised pretraining and mark localization are the two most impactful factors, with learned cropping recovering the majority of the gain from manual cropping without requiring ground-truth annotations at inference time. To enable reproducibility and adoption in the digital humanities, we release our manually annotated dataset and codebase, and deploy the system via a public web interface.
Chinese Translation
对于艺术史学家而言,金匠标识(goldsmith marks)在文物的鉴定与断代中起着关键作用。在实际工作中,专家必须人工将待查询标识与数百个已记录的范例进行比较,这一过程既繁琐又高度依赖专业知识。为解决这一问题,我们提出了一种AI辅助的检索流水线,该流水线将标识定位与度量学习微调相结合,并在三种骨干架构上进行实验:ImageNet预训练的ResNet-50、有监督的ViT-S/16以及自监督的DINOv2 ViT-S/14。我们对裁剪策略进行了系统性评估,衡量了不裁剪、人工真实标注(ground-truth)裁剪以及基于学习检测的裁剪的影响,并评估了它们与各骨干架构之间的交互作用。我们最强的配置——DINOv2 ViT-S/14结合人工裁剪与度量学习微调——达到了62.63%的mAP和73.74%的Top-1准确率。实验表明,自监督预训练和标识定位是影响最大的两个因素,而基于学习的裁剪无需在推理阶段使用真实标注即可获得接近人工裁剪的大部分增益。为促进数字人文领域的可复现性与应用,我们发布了人工标注的数据集和代码库,并通过公开的网页界面部署了该系统。
cs.CV / 87 / 2609.20562

A Dual-Stream Regulated Reconstruction and Segmentation Network with Hierarchical Artifact-Prior Modeling for Ultra-Low-Field Pediatric Neuroimaging

一种基于分层伪影先验建模的双流调控重建与分割网络,用于超低场儿童神经影像
Jafrasteh, Bahram, Milecki, Leo, Zhao, Qingyu
Abstract
Automated quality assessment, enhancement, and segmentation of multiple structures in $0.064\,\mathrm{T}$ ultra-low-field pediatric MRI are limited by a low signal-to-noise ratio, weak anatomical boundaries, and frequent artifacts. We present a unified framework for the LISA 2026 Challenge that performs all three tasks together within one inference pipeline. A network with two coupled streams, built on a 3D U-Net, first reconstructs an enhanced uLF volume and then combines the original and enhanced images for subcortical segmentation. To improve boundary stability, we add an auxiliary class covering brain tissue outside the target structures, derived from whole brain masks. A head conditioned on an artifact graph predicts the seven artifact ratings from reconstruction residuals and frozen segmentation features. We address the scarcity of dense annotations using diffeomorphic registration from atlas to target for label propagation and to regularize anatomical reconstruction. We report validation results across all three tasks.
Chinese Translation
0.064 T超低场儿童MRI中多结构的自动化质量评估、增强与分割受到低信噪比、微弱解剖边界以及频繁伪影的限制。我们提出了一个面向LISA 2026挑战赛的统一框架,在单一推理流程中同时完成这三项任务。该网络构建于3D U-Net之上,包含两条相互耦合的流:首先重建增强的超低场(uLF)体数据,然后将原始图像与增强图像结合用于皮层下结构分割。为提高边界稳定性,我们增加了一个辅助类别,覆盖目标结构之外的脑组织,该类别由全脑掩膜推导得到。一个以伪影关系图为条件的预测头,从重建残差和冻结的分割特征中预测七项伪影评分。针对密集标注稀缺的问题,我们采用从图谱到目标图像的微分同胚配准进行标签传播,并对解剖重建进行正则化。我们报告了在全部三项任务上的验证结果。
cs.CV / 88 / 2609.20574

DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering

DocAttriBench:文档视觉问答中答案定位的基准测试
De Grandis, Luca, Cappelletti, Silvia, Raccagni, William, Cornia, Marcella, Baraldi, Lorenzo, Cucchiara, Rita
Abstract
Answer grounding in document visual question answering remains an open challenge: most benchmarks lack grounding annotations or provide limited-quality labels, while constructing grounded datasets still requires costly manual effort. We introduce DocAttriBench (DAB), a large-scale benchmark for fine-grained, element-level source attribution in Document VQA, grounding answers to specific layout elements such as text blocks, tables, and images. To build DAB, we propose a Mask-based Perplexity-Derived Attribution method (MAPPET) that combines document layout and language modeling to identify the most informative element for each answer. MAPPET measures the increase in perplexity after masking candidate elements and attributes the answer to the element contributing most to model confidence. Applying MAPPET to multiple existing Document VQA datasets yields DAB, with 237k documents and 296k question-answer pairs with element-level grounding. We benchmark grounding-capable multimodal LLMs on DAB, evaluating answer accuracy, attribution accuracy, and overall answer quality. Results show that while larger models generally achieve higher answer accuracy, even the strongest models often fail to localize the supporting elements. DAB provides a scalable benchmark for developing grounded, verifiable, and trustworthy Document VQA models. Dataset and code are available at https://aimagelab.github.io/DocAttriBench/.
Chinese Translation
文档视觉问答中的答案定位仍然是一个开放的挑战:大多数基准测试缺乏定位标注,或仅提供质量有限的标签,而构建带定位信息的数据集仍需高昂的人工成本。我们提出了DocAttriBench(DAB),一个面向文档视觉问答(Document VQA)中细粒度、元素级来源归因的大规模基准,将答案定位到文本块、表格和图像等特定版面元素。为构建DAB,我们提出了一种基于掩码的困惑度推导归因方法(Mask-based Perplexity-Derived Attribution,MAPPET),该方法结合文档版面分析与语言建模,为每个答案识别信息量最大的元素。MAPPET通过测量掩蔽候选元素后困惑度的增加量,将答案归因于对模型置信度贡献最大的元素。将MAPPET应用于多个现有的Document VQA数据集,构建了包含23.7万篇文档和29.6万个具有元素级定位标注的问答对的DAB基准。我们在DAB上对具备定位能力的多模态大语言模型进行了基准测试,评估其答案准确率、归因准确率和整体答案质量。结果表明,尽管更大的模型通常能获得更高的答案准确率,但即使是最强的模型也常常无法定位支持答案的元素。DAB为开发可定位、可验证且可信的Document VQA模型提供了一个可扩展的基准。数据集和代码可在https://aimagelab.github.io/DocAttriBench/获取。
cs.CV / 89 / 2609.20589

RawSLAM: Online HDR Gaussian SLAM from Linear Radiance

RawSLAM:基于线性辐射度的在线HDR高斯SLAM
González, Marina Orozco, Merino, Luis
Abstract
Current dense visual SLAM systems rely almost exclusively on 8-bit tonemapped Low Dynamic Range (LDR) inputs, limiting their robustness in extreme lighting where shadows and highlights trigger tracking drift and mapping collapse. Conversely, existing raw and High Dynamic Range (HDR) reconstruction pipelines operate strictly offline. They depend on Structure-from-Motion preprocessing and are not suited for large inter-frame motion. We present, to the best of our knowledge, the first online Gaussian SLAM framework that tracks and maps directly on single-exposure 16-bit linear HDR imagery. Our method rests on three core components: an architecture-agnostic HDR Gaussian Splatting module featuring an MLP-free logarithmic parameterization of Gaussian color features; a Reinhard range-compressed photometric objective; and structure-guided spatial gradient weighting. Combined, these components allow our approach to outperform a direct HDR adaptation of MonoGS in both trajectory and reconstruction accuracy, while rendering natively in linear scene radiance for post-rendering processing. The same formulation runs unchanged on standard 8-bit inputs, roughly halving the MonoGS baseline error. Furthermore, our HDR Gaussian module transfers seamlessly to SplaTAM, Gaussian SLAM, and DROID-W, eliminating all tracking failures these systems suffer on challenging illumination sequences. To enable this research, we introduce RawSLAM: a dataset of 10 real-world indoor sequences featuring 16-bit RAW imagery, aligned depth, IMU measurements, and external OptiTrack poses. Code and dataset will be made publicly available soon.
Chinese Translation
当前的稠密视觉SLAM系统几乎完全依赖于8位色调映射的低动态范围(LDR)输入,这限制了它们在极端光照条件下的鲁棒性——阴影和高光会引发跟踪漂移和建图崩溃。与此相反,现有的RAW和高动态范围(HDR)重建流程严格限于离线运行,它们依赖运动恢复结构(Structure-from-Motion)预处理,且不适用于较大的帧间运动。据我们所知,我们提出了首个可直接在单次曝光的16位线性HDR图像上进行跟踪与建图的在线高斯SLAM框架。我们的方法基于三个核心组件:一个与架构无关的HDR高斯泼溅(Gaussian Splatting)模块,其特点是对高斯颜色特征采用无需MLP的对数参数化;一个经过Reinhard范围压缩的光度目标函数;以及结构引导的空间梯度加权。这些组件相结合,使我们的方法在轨迹精度和重建精度上均优于MonoGS的直接HDR适配版本,同时能够以线性场景辐射度进行原生渲染,便于后渲染处理。相同的框架无需修改即可在标准8位输入上运行,与MonoGS基线相比大约将误差减半。此外,我们的HDR高斯模块可无缝迁移至SplaTAM、Gaussian SLAM和DROID-W,消除了这些系统在挑战性光照序列上出现的所有跟踪失败。为推动该领域研究,我们推出了RawSLAM数据集:包含10个真实世界室内序列,具有16位RAW图像、对齐的深度图、IMU测量数据以及外部OptiTrack位姿。代码和数据集将很快公开发布。
cs.CV / 90 / 2609.20623

PhGS: Post-Hoc Pruning and Refinement of Single-View Feed-Forward 3D Gaussian Reconstructions

PhGS:单视图前馈3D高斯重建的事后剪枝与优化
Yagawa, Rinto, Cheng, Han, Schmalstieg, Dieter, Saito, Hideo, Mori, Shohei
Abstract
Recent single-view feed-forward 3D Gaussian Splatting (3DGS) generation predicts a fixed number of Gaussians per camera ray, introducing severe spatial redundancy. Most existing compaction strategies target multi-view setups to exploit cross-view consistency and are incompatible with single-image models. Instead of retraining the base feed-forward network to directly output compact representations, our insight is to keep the base models frozen and apply post-hoc pruning and recurrent refinement to the generated Gaussians. Consequently, we propose a backbone-agnostic compaction pipeline for single-view feed-forward 3DGS that couples an importance-score-based pruning mechanism with a trainable, lightweight recurrent refinement module, which iteratively updates the surviving primitives to restore image quality. Our results demonstrate seamless integration with existing baselines while preserving novel-view rendering fidelity and achieving high memory reduction. Furthermore, our method supports flexible inference-time keep ratios for application needs.
Chinese Translation
近期的单视图前馈3D高斯泼溅(3D Gaussian Splatting, 3DGS)生成方法为每条相机光线预测固定数量的高斯,引入了严重的空间冗余。现有的大多数压缩策略面向多视图设置,以利用跨视图一致性,因而与单图像模型不兼容。我们并不通过重新训练基础前馈网络来直接输出紧凑表示,而是保持基础模型冻结,对生成的高斯施加事后(post-hoc)剪枝与循环精化。据此,我们提出了一种与骨干网络无关的单视图前馈3DGS压缩流水线,该流水线将基于重要性得分的剪枝机制与一个可训练的轻量级循环精化模块相结合,后者迭代地更新保留下来的图元以恢复图像质量。实验结果表明,本方法可与现有基线模型无缝集成,在保持新视角渲染保真度的同时实现大幅内存缩减。此外,我们的方法支持灵活的推理时保留比例,以适应不同的应用需求。
cs.CV / 91 / 2609.20633

Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network

精炼天然可编辑:基于生成式精炼网络的无训练提示到提示图像编辑
Chen, Yulong, Zhang, Ziqian, Zhang, Haoyu, He, Ao, Li, Senmao, Wang, Kai
Abstract
Text-guided image editing must introduce the requested changes while preserving unrelated source content. Diffusion-based editors rely on spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal autoregressive editors face a further constraint: their fixed decoding order limits revision of earlier decisions. We introduce RefineEdit, a training-free prompt-to-prompt image editing framework built on a Generative Refinement Network. Our key idea is to couple edit localization with content generation through the global refinement of binary image codes, allowing editing evidence to be reassessed as the image evolves. RefineEdit initializes an editing branch from an intermediate source state, reusing the emerging layout. We compare the probabilities assigned by the two branches to the same source-sampled bits, using their signed differences to select editable positions and bits. Selected bits follow editing refinement, while the remaining bits copy the evolving source state. To stabilize editing across refinement steps, adaptive spatial freezing limits unnecessary mask expansion, while finite bit locking keeps recently selected bits editable. The framework requires no additional training, external masks, or attention control. Across nine editing categories of PIE-Bench, RefineEdit achieves the best background-preservation scores in PSNR, LPIPS, MSE and SSIM, together with the highest whole-image and edited-region CLIP scores among the evaluated methods.
Chinese Translation
文本引导的图像编辑需要在实现所需修改的同时保留无关的源内容。基于扩散模型的编辑器依赖空间控制,其不精确性可能导致编辑不完整或改动无关区域。因果自回归编辑器则面临进一步的限制:其固定的解码顺序限制了对先前决策的修正。我们提出了 RefineEdit,一个基于生成式精炼网络(Generative Refinement Network)构建的无训练提示到提示图像编辑框架。我们的核心思想是通过二值图像码的全局精炼,将编辑定位与内容生成相耦合,使编辑证据能够随着图像的演化而被重新评估。RefineEdit 从源图像的中间状态初始化一个编辑分支,复用正在形成的布局。我们比较两个分支对相同源采样比特所赋予的概率,利用其带符号差值来选择可编辑的位置和比特。被选中的比特遵循编辑精炼,而其余比特则复制不断演化的源状态。为了在精炼步骤间稳定编辑过程,自适应空间冻结限制了不必要的掩码扩张,同时有限比特锁定保持最近选中的比特可被编辑。该框架无需额外训练、外部掩码或注意力控制。在 PIE-Bench 的九个编辑类别上,RefineEdit 在 PSNR、LPIPS、MSE 和 SSIM 上取得了最佳的背景保持分数,并在所有被评估方法中获得了最高的全图和编辑区域 CLIP 分数。
cs.CV / 92 / 2609.20638

PROVIA: Procedure State Tracking for Online Mistake Detection in Egocentric Videos

PROVIA:面向第一视角视频在线错误检测的程序状态跟踪
Wen, Di, Yang, Kailun, Weissert, Jimmy, Scherrer, Luc Maria, Zöllner, Cedric, Liu, Ruiping, Chen, Yufan, Wei, Jiale, Zheng, Junwei, Peng, Kunyu
Abstract
An assistant watching egocentric video should notice a mistake from past frames alone, before the next step begins, and keep working once the person recovers. A mistake changes the state of the work, so every later step has to be read against what was done rather than against the plan. The first-mistake protocol that current online methods report on cuts each recording at its first mistake, so a fixed-time rule that never looks at the video is right on every case. We evaluate on complete trials, where mistakes and recoveries arise naturally, under a validation false-alarm budget and against controls that use timing alone. PROVIA keeps two records apart: a factual state, a learned summary of the steps each actor performed, mistakes included, and the accepted progress, an exact posterior over the state of an automaton induced from correct demonstrations by Bayesian state merging and over the execution status of each actor. Procedure-state transitions occur only in the correct-status branch; the mistake and correction branches retain the source state. A sequential test turns the per-frame mistake probability into alarms. With one filter and one optimization rule, PROVIA ranks mistakes best among the evaluated controlled baselines on CaptainCook4D, IndustReal, HoloAssist and IMPACT-ego. At a validation budget of 0.1 false alarms per minute it recalls .154 against .128 on CaptainCook4D and .034 against .015 on HoloAssist, where it leads at every budget. The pipeline runs at 58-70 frames per second. The source code is available at https://github.com/Kratos-Wen/PROVIA.
Chinese Translation
一个观看第一视角(egocentric)视频的助手应当仅凭过去的帧就能在下一步开始之前察觉到错误,并在操作者恢复后继续正常工作。错误会改变工作的状态,因此后续每一步都必须对照已完成的实际情况而非原计划来解读。当前在线方法所报告的“首次错误”评测协议会在每个录像的首次错误处截断,这使得一个从不观看视频、仅按固定时间规则的简单方法也能在所有案例上表现正确。我们在完整的试验上进行评估,其中错误与恢复自然出现,并在验证误报预算下以及与仅使用时间信息的对照方法进行比较。PROVIA 将两类记录区分开来:一是事实状态(factual state),即对每个操作者(含错误在内)所执行步骤的学习式摘要;二是已接受进度(accepted progress),即通过贝叶斯状态合并(Bayesian state merging)从正确示范中归纳出的自动机状态的精确后验分布,以及每个操作者执行状态的精确后验。程序状态的转换仅发生在正确状态分支中;错误和纠正分支则保留源状态。一个序贯检验将逐帧错误概率转化为报警信号。仅需一个滤波器和一个优化规则,PROVIA 在 CaptainCook4D、IndustReal、HoloAssist 和 IMPACT-ego 数据集上的错误排序表现优于所有受控基线。在每分钟 0.1 次误报的验证预算下,其在 CaptainCook4D 上的召回率为 0.154(对比 0.128),在 HoloAssist 上为 0.034(对比 0.015),并在所有预算水平上均处于领先地位。该流程的运行速度为每秒 58–70 帧。源代码可在 https://github.com/Kratos-Wen/PROVIA 获取。
cs.CV / 93 / 2609.20662

Earth Surface Immune System for Rapid Monitoring of Unknown Anomalies

用于快速监测未知异常的地球表面免疫系统
Li, Jingtao, Zhu, Qian, Wang, Xinyu, Li, Deren, Zhang, Liangpei, Zhong, Yanfei
Abstract
Earth surface anomalies, driven by escalating climate change, and expanding human activities, are increasing in both frequency and diversity, yet their limited historical data and unpredictability make them fundamentally different from conventional remote sensing targets. Existing methods address specific anomaly categories or stop at localization, leaving a gap between detection and actionable information. Here we present ESIA, an Earth Surface Immune System whose architecture is constrained by three principles from the biological immune system, refined over millions of years against equally diverse and uncertain threats. A non-specific innate immune stage treats anomalies as unobserved changes in time-series satellite imagery, generating binary localization maps at 14.51 km2/s without assuming any anomaly category, surpassing the strongest general baseline by 37% in F1. A specific adaptive immune stage applies negative selection to filter text prompts and matches surviving prompts with localized image patches through a multi-modal foundation model, enabling open-vocabulary recognition of unknown anomaly attributes including category, affected area, and damage severity, with recognition F1 exceeding 80%. A mutation mechanism tunes minimal embeddings at test time, adapting to each scene in 3.26s using a single reference image pair. We validate ESIA on a global-scale dataset covering 19,801.60 km2 across six anomaly categories, comparing against 22 models, and further apply it to quantify degraded farmland in the Dnipro Delta following the Kakhovka Dam collapse and assess burn severity from 2025 Palisades Fire in Los Angeles. This unprecedented flexibility in handling unknown anomalies opens new avenues for real-time disaster response and environmental surveillance.
Chinese Translation
受气候变化加剧和人类活动不断扩张的驱动,地球表面异常在频率和多样性上均呈上升趋势。然而,其历史数据有限且不可预测,使其与常规遥感目标存在本质差异。现有方法仅针对特定异常类别,或止步于定位,在检测与可操作信息之间留下了空白。本文提出ESIA(Earth Surface Immune System,地球表面免疫系统),其架构受生物免疫系统三条原则的约束——该系统历经数百万年演化,以应对同样多样且不确定的威胁。非特异性先天免疫阶段将异常视为时序卫星影像中的未观测变化,在不对异常类别做任何假设的情况下,以14.51平方公里/秒的速度生成二值定位图,在F1分数上超越最强的通用基线方法37%。特异性适应性免疫阶段利用负选择机制筛选文本提示,并通过多模态基础模型将存活的提示与定位的影像斑块进行匹配,实现对未知异常属性(包括类别、受影响面积和损害严重程度)的开放词汇识别,识别F1分数超过80%。突变机制在测试时调整最小嵌入,仅需一对参考影像即可在3.26秒内适应每个场景。我们在覆盖六类异常、总面积达19,801.60平方公里的全球尺度数据集上验证了ESIA,并与22个模型进行了对比;我们进一步将其应用于量化卡霍夫卡大坝溃决后第聂伯河三角洲的农田退化情况,并评估了2025年洛杉矶帕利塞兹大火(Palisades Fire)的燃烧严重程度。这种在处理未知异常方面前所未有的灵活性,为实时灾害响应和环境监测开辟了新的途径。
cs.CV / 94 / 2609.20673

FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents

FunArt:从生成式三维潜在表示中解码功能结构与关节连接
Rotondi, Dennis, Werby, Abdelrhman, Arras, Kai O.
Abstract
To operate effectively in human environments, robots must identify articulated objects, segment their movable and interactive parts, and estimate their kinematic models. Existing articulated scene representations typically recover kinematics from observed interactions, while methods operating on static scans often decouple articulation from functional interactive elements. We present FunArt, a framework that constructs articulation-aware functional 3D scene graphs from posed RGB-D observations captured in a single static configuration. FunArt reconstructs object instances, converts their fused geometry directly into the O-Voxel representation of TRELLIS.2, and exploits its frozen, sparse-compression VAE as a structural prior. A lightweight query-based decoder combines compact object-level latents with dense, surface-aligned features to jointly segment movable parts and functional interactive elements while estimating motion type, axis, origin, and range. On the Articulate3D dataset, FunArt achieves state-of-the-art performance across movable-part segmentation, articulation estimation, and functional-element segmentation, both with and without ground-truth object input. In the end-to-end setting, it outperforms the strongest baselines by 1.5 AP_{50} points for movable parts, 2.8 AP_{50} points under joint origin-and-axis constraints, and 6.7 AP_{50} points for functional elements. These results demonstrate that generative 3D latents encode actionable structural cues that can initialize robotic perception and planning before physical interaction.
Chinese Translation
为了在人类环境中有效运作,机器人必须识别可动关节物体、分割其可移动和可交互的部分,并估计其运动学模型。现有的可动关节场景表示通常从观察到的交互中恢复运动学信息,而基于静态扫描的方法往往将关节连接与功能性交互元素相分离。我们提出了FunArt,一个从单一静态配置下采集的带位姿RGB-D观测中构建关节感知的功能性三维场景图的框架。FunArt重建物体实例,将其融合的几何体直接转换为TRELLIS.2的O-Voxel表示,并利用其冻结的稀疏压缩VAE作为结构先验。一个轻量级的基于查询的解码器将紧凑的物体级潜在表示与稠密的表面对齐特征相结合,在联合估计运动类型、旋转轴、轴心和运动范围的同时,共同分割可移动部分和功能性交互元素。在Articulate3D数据集上,无论是否提供真实物体输入,FunArt在可动部分分割、关节估计和功能元素分割方面均取得了最先进的性能。在端到端设置下,它相较于最强的基线方法,在可动部分上提升了1.5个AP_{50}点,在联合轴心与旋转轴约束下提升了2.8个AP_{50}点,在功能元素上提升了6.7个AP_{50}点。这些结果表明,生成式三维潜在表示编码了可操作的结构线索,能够在物理交互之前初始化机器人感知与规划。
cs.CV / 95 / 2609.20700

Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation

该病例是否应进行适应?预测碎片化控制测试时适应
Wang, Lili, Li, Jing, Sun, Xiaowen, Hu, Xiangyu, Gu, Zhuangzhuang, Liu, Jian, Nelakuditi, Srihari, Tong, Yan
Abstract
Episodic test-time adaptation resets a frozen segmenter to source weights $M_0$ on each case and adapts for a fixed step count. A fixed horizon conflates a cohort-level question, how far to adapt, with an irreducibly per-case one, whether this case should be adapted at all. Cohort means hide that decision: on cross-vendor cardiac MRI the mean $\Delta$Dice from adaptation is statistically indistinguishable from zero while 58.7% of cases are individually made worse. We quantify this harm as harmful accepted area (HA), the harmful fraction of the edited area a controller deploys. Held-out tuning gives a stronger baseline than a fixed horizon, but the budget it selects transfers on neither of the two main medical benchmarks, and no global budget can condition on the case. We show that prediction fragmentation---the disagreement geometry between $M_0$ and the adapted mask $M_k$---predicts HA with no labels or extra backward passes at decision time, comparably on three benchmarks (Spearman $\rho$ 0.50--0.60), at a quarter of gradient-norm's latency. A case-level router built on it cuts HA from 0.228 to 0.139 on a benchmark that took no part in its design, with the design frozen and only cut-points recalibrated there. On the cardiac benchmark the design was selected on, the router cuts HA from 0.129 to 0.013 at matched Dice and 1.10 deployed updates, against the retrospective-best budget found post hoc on evaluation labels, and reduces that 58.7% to 20.0%, an upper bound we quantify. Where the retained cases are not net-helped (as on prostate), the router still cuts HA but concedes accuracy, a boundary we report. Thresholds are fit once on a labeled split disjoint from evaluation; decisions use no labels or gradients. The template ports across architecture and domain (nnU-Net$\to$SegFormer, Cityscapes$\to$ACDC) with coordinate, thresholds and per-bucket actions instantiated per domain.
Chinese Translation
情景式测试时适应在每个病例上将冻结的分割器重置为源权重 $M_0$,并进行固定步数的适应。固定的适应步数将一个群体层面的问题——适应到什么程度——与一个不可化约的病例层面的问题——该病例是否应该进行适应——混为一谈。群体平均值掩盖了这一决策:在跨厂商心脏MRI数据上,适应带来的平均 $\Delta$Dice 在统计上与零无显著差异,而58.7%的病例却被单独恶化。我们将这种损害量化为有害接受面积(HA),即控制器部署的编辑区域中有害部分的比例。在保留集上调优得到的基线比固定步数更强,但其所选的预算在两个主要医学基准上均不能迁移,且任何全局预算都无法针对具体病例进行条件判断。我们证明,预测碎片化——即 $M_0$ 与适应后掩码 $M_k$ 之间的分歧几何结构——能够在决策时无需标签或额外反向传播的情况下预测HA,在三个基准上表现相当(Spearman $\rho$ 为0.50–0.60),且延迟仅为梯度范数方法的四分之一。基于此构建的病例级路由器,在一个未参与其设计的基准上将HA从0.228降至0.139,其中设计保持冻结,仅在该基准上重新校准了切割点。在设计所选的心脏基准上,路由器在匹配Dice和1.10次部署更新的条件下,将HA从0.129降至0.013,优于在评估标签上事后找到的回顾性最佳预算,并将58.7%的恶化病例比例降至20.0%(我们量化了这一上界)。当保留病例未能获得净收益时(如前列腺数据集),路由器仍能降低HA,但会牺牲一定精度,我们报告了这一边界。阈值仅在评估集之外的带标签数据划分上拟合一次;决策过程不使用任何标签或梯度。该模板可跨架构和领域移植(nnU-Net$\to$SegFormer,Cityscapes$\to$ACDC),其中坐标、阈值和各桶动作均按领域实例化。
cs.CV / 96 / 2609.20769

FlowSGS: Improving Flow Matching Priors for Inverse Imaging with Stochastic Interpolants

FlowSGS:利用随机插值改进逆问题成像中的流匹配先验
Li, Tianao, Qian, Xinhui, Alexander, Emma
Abstract
Flow matching has emerged as the state-of-the-art generative model and has been used for plug-and-play (PnP) priors to solve inverse problems in computational imaging. However, existing flow-based inverse solvers assume linear forward models and/or make simplifying approximations in posterior sampling. To circumvent these problems, we introduce FlowSGS, a flow-based posterior sampling method using Split Gibbs Sampling (SGS) to decompose the posterior into a likelihood step and a prior step. Specifically, we sample from the likelihood step using Langevin dynamics and leverage the Stochastic Interpolants (SI) framework to integrate a pretrained flow model into the prior step. We provide a form for the prior step that uses SI's reverse-time SDE, and show connections to previous PnP methods. Moreover, with the aid of the flow prior's straight probability paths and a novel timestep correction technique for the reverse-time SDE, FlowSGS requires fewer network evaluations in its prior step than plug-and-play diffusion samplers. Our experiments show state-of-the-art performance on a range of inverse problems. For the first time, we provide an experiment on a nonlinear inverse problem (Fourier phase retrieval) for flow-based inverse solvers.
Chinese Translation
流匹配(Flow Matching)已成为最先进的生成模型,并被用作即插即用(Plug-and-Play, PnP)先验以求解计算成像中的逆问题。然而,现有的基于流的逆问题求解器通常假设正向模型为线性模型,和/或在后验采样中采用简化近似。为规避这些问题,我们提出了FlowSGS,一种基于流的后验采样方法,利用分裂吉布斯采样(Split Gibbs Sampling, SGS)将后验分解为似然步骤和先验步骤。具体而言,我们使用朗之万动力学(Langevin dynamics)从似然步骤中采样,并借助随机插值(Stochastic Interpolants, SI)框架将预训练的流模型融入先验步骤。我们给出了使用SI逆时间随机微分方程(SDE)的先验步骤形式,并展示了其与以往PnP方法的联系。此外,借助流先验的直线路径以及一种针对逆时间SDE的新型时间步校正技术,FlowSGS在先验步骤中所需的网络评估次数少于即插即用扩散采样器。我们的实验表明,该方法在一系列逆问题上取得了最先进的性能。我们首次为基于流的逆问题求解器提供了一个非线性逆问题(傅里叶相位恢复)的实验。
cs.CV / 97 / 2609.20815

ERCPMP-Gx: Endoscopic Image and Video Dataset for Morphological, Histopathological, and Genomic Characterization of Colorectal Polyposis

ERCPMP-Gx:用于结直肠息肉病形态学、组织病理学和基因组学特征描述的内镜图像与视频数据集
Ghaffari, Zahra, Bahar, Massih, Forootan, Mojgan, Darvishi, Ali, Bolhasani, Hamidreza
Abstract
Hereditary polyposis syndromes can be precursor lesions to colorectal cancer and are associated with a broad spectrum of extracolonic tumors. Early identification and accurate classification of these syndromes are essential for timely diagnosis, individualized patient management, and targeted surveillance strategies for affected families. However, public endoscopic datasets are largely organized around the individual sporadic polyp, and none links the polyposis phenotype to histopathology and germline findings at the patient level. Here, we present ERCPMP-Gx, an endoscopic, histopathological, and genomic dataset developed to support the application of artificial intelligence (AI) in the recognition, characterization, and classification of colorectal polyposis. Most procedures were performed using the Olympus EVIS X1 system with white-light endoscopy (WLE), narrow-band imaging (NBI), magnifying NBI (M-NBI), and NBI with near focus modes, yielding 160 images and accompanying video clips. Approximately eighty percent of cases represent clinically and/or genetically confirmed hereditary polyposis syndromes (PG), including familial adenomatous polyposis (FAP), Peutz-Jeghers syndrome (PJS), juvenile polyposis syndrome (JPS), and ganglioneuroma syndrome (GNS), while the remaining twenty percent comprise non-hereditary polyps and polyp-mimicking lesions with overlapping morphological features (Non-PG), included to support differential classification. Each released record is linked, where available, to standardized endoscopic annotations, representative histopathology, and clinically reported germline findings, forming an AI-ready, patient-level annotation framework. The dataset is publicly accessible at Mendeley (https://doi.org/10.17632/nzyfc544bx.2). For the latest updates and further information, readers are referred to the DataBioX website: https://databiox.com.
Chinese Translation
遗传性息肉病综合征可能是结直肠癌的癌前病变,并与多种结肠外肿瘤相关。对这些综合征进行早期识别和准确分类,对于及时诊断、个体化患者管理以及针对受累家族的靶向监测策略至关重要。然而,现有公开内镜数据集大多围绕单个散发性息肉组织,尚无数据集在患者层面将息肉病表型与组织病理学和胚系检测结果相关联。本文提出ERCPMP-Gx,这是一个涵盖内镜、组织病理学和基因组学的数据集,旨在支持人工智能(AI)在结直肠息肉病的识别、特征描述和分类中的应用。大多数检查采用Olympus EVIS X1系统,在白光内镜(WLE)、窄带成像(NBI)、放大NBI(M-NBI)及近聚焦NBI模式下进行,共获得160张图像及配套视频片段。约80%的病例为临床和/或基因确诊的遗传性息肉病综合征(PG),包括家族性腺瘤性息肉病(FAP)、黑斑息肉综合征(PJS)、幼年性息肉病综合征(JPS)和节细胞神经瘤综合征(GNS);其余20%为具有相似形态学特征的非遗传性息肉及息肉样仿似病变(Non-PG),以支持鉴别分类。每个发布的数据记录在可行的情况下均与标准化内镜标注、代表性组织病理学结果以及临床报告的胚系检测结果相关联,形成了面向AI的、患者层面的标注框架。该数据集可通过Mendeley公开访问(https://doi.org/10.17632/nzyfc544bx.2)。有关最新更新和更多信息,请参阅DataBioX网站:https://databiox.com。
cs.CV / 98 / 2609.20816

Paint-Anything: Unified Any-Color Control for Image Generation and Editing

Paint-Anything:面向图像生成与编辑的统一任意颜色控制
Xie, Ji, Zhou, Dewei, Huang, Xinyu, Chen, Zhennan, Wang, Xun
Abstract
Professional design requires any-color control: the ability to specify an object's target color with any 24-bit hex value for image generation and editing. Prior work has explored color generation, editing, and colorization, but often relies on dedicated color representations or specialized inference procedures. Advances in large language models offer a simpler starting point: even compact models can associate hex values with color semantics. We present Paint-Anything, which learns a shared hex-prompt interface for generation and editing through object-level color supervision. We develop a data pipeline that constructs Paint-500K from real images through object grounding, perceptual color labeling, and editing-pair synthesis. Since shadows make real-image labels only approximate colors, we complement this supervision with pure-color anchors whose pixels exactly match their paired hex values. These anchors are used only at high-noise timesteps, leaving low-noise training to natural images. We further introduce Any Color Benchmark (ACBench), comprising ACBench-T2I and ACBench-Edit, to measure object-level hex color fidelity across both tasks. On FLUX.2-4B, Paint-Anything improves ACBench-T2I and ACBench-Edit scores by 85.3% and 28.3%, respectively, relative to the base model, with ablations supporting the training recipe. It also achieves the highest average CompColor score among the compared methods.
Chinese Translation
专业设计需要任意颜色控制能力,即能够以任意24位十六进制(hex)值指定对象的目标颜色,用于图像生成与编辑。已有研究探索了颜色生成、编辑和着色,但通常依赖专门的颜色表示或特定的推理流程。大语言模型的进展提供了一个更简单的起点:即使是紧凑型模型也能将hex值与颜色语义关联起来。我们提出Paint-Anything,通过对象级颜色监督学习一个共享的hex提示接口,统一服务于生成与编辑任务。我们构建了一条数据流水线,通过对象定位、感知颜色标注和编辑对合成,从真实图像构建了Paint-500K数据集。由于阴影导致真实图像的颜色标签只能是近似值,我们用纯色锚点补充这一监督,这些锚点的像素与其配对的hex值完全匹配。这些锚点仅在高噪声时间步使用,低噪声阶段的训练则留给自然图像。我们进一步提出任意颜色基准(Any Color Benchmark,ACBench),包括ACBench-T2I和ACBench-Edit,用于衡量两项任务中对象级hex颜色的保真度。在FLUX.2-4B上,相对于基础模型,Paint-Anything将ACBench-T2I和ACBench-Edit分数分别提升了85.3%和28.3%,消融实验验证了该训练方案的有效性。此外,它在所比较的方法中取得了最高的平均CompColor分数。
cs.CV / 99 / 2609.20817

FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

FAMOS:基于稀疏观测的前馈式三维铰接物体建模
Qu, Kevin, Sun, Tao, Viola, Massimiliano, Zhu, Liyuan, Zhou, Zhizhuo, Sarkar, Sayan Deb, Schindler, Konrad, Armeni, Iro
Abstract
Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a single observation and therefore rely heavily on learned category-level shape priors. We present FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds. Our model jointly reasons over multiple observations and naturally supports a variable number of inputs, including a single view. To aggregate articulation cues across observations, we introduce a Multi-state Articulation Transformer with alternating state-wise and global attention. We further propose an observed articulation span objective that supervises the motion range each part exhibits across the input observations, encouraging the model to leverage the full observation set. To overcome the limited scale and diversity of existing datasets, we introduce a procedural data generator that synthesizes self-annotated assets during training. Experiments on PartNet-Mobility, ACD, and ArtiCraft-10K demonstrate consistent improvements over both feed-forward and optimization-based baselines. Project page: https://kevinqu7.github.io/famos
Chinese Translation
从稀疏的单目视角对铰接物体进行建模极具挑战性,因为每次观测仅能揭示部分几何与运动信息。大多数前馈方法仅从单一观测推断铰接结构,因而严重依赖于学习到的类别级形状先验。我们提出了FAMOS,一个能够从稀疏、无序的部分点云集合中预测可动部分分割与关节参数的前馈模型。我们的模型对多次观测进行联合推理,并自然支持可变数量的输入,包括单一视角。为了跨观测聚合铰接线索,我们提出了一种交替使用状态级注意力与全局注意力的多状态铰接Transformer(Multi-state Articulation Transformer)。此外,我们提出了一种观测铰接范围目标函数,用于监督各部分在输入观测中所呈现的运动范围,从而鼓励模型充分利用全部观测集合。为克服现有数据集在规模和多样性上的不足,我们引入了一个程序化数据生成器,可在训练过程中合成自带标注的资产。在PartNet-Mobility、ACD和ArtiCraft-10K上的实验表明,我们的方法相较于前馈式和基于优化的基线方法均取得了一致的性能提升。项目页面:https://kevinqu7.github.io/famos
cs.CV / 100 / 2609.20818

SplashSplat: Reconstructing Splashing Liquids from Real-World Multi-View Videos

SplashSplat:从真实世界多视角视频重建飞溅液体
Liu, Peiyu, Zhang, Dingxi, Tombari, Federico, Pollefeys, Marc, Tsalicoglou, Christina, Barath, Daniel
Abstract
A splash lives for a fraction of a second: sheets tear into ligaments and droplets, appearance is view-dependent and nearly textureless, and little persists long enough to track. Reconstruction research has consequently focused on smoke, synthetic liquids, or gently deforming surfaces. To our knowledge, no synchronized multi-view dataset of splashing liquids exists. We therefore introduce a benchmark of 20 real scenes, from coherent streams to violent splashes, captured by seven synchronized, calibrated 4K cameras at 60 fps, with manually refined per-view liquid and container masks and fixed evaluation splits. We further present SplashSplat, built on a single principle: impose physical structure only where the observations can constrain it. Per-frame liquid SDFs fused from the masks provide the geometry, level-set transport between consecutive SDFs yields a coarse velocity field, and Lagrangian carriers advected along this flow, corrected against each new observation and reseeded where coverage is lost, decode local Gaussians for differentiable rendering. SplashSplat outperforms state-of-the-art dynamic Gaussian splatting methods on our real captures and on a synthetic benchmark, with physically more plausible motion and a lower training cost. The same representation supports temporal interpolation and style transfer without re-optimization.
Chinese Translation
飞溅现象仅持续不到一秒:液膜撕裂成液柱和液滴,其外观依赖视角且几乎没有纹理,几乎没有物质能存留足够长的时间以供跟踪。因此,重建研究一直聚焦于烟雾、合成液体或缓慢变形的表面。据我们所知,目前尚不存在飞溅液体的同步多视角数据集。为此,我们引入了一个包含20个真实场景的基准数据集,涵盖从连贯水流到剧烈飞溅的多种情形,由七台同步标定的4K相机以60帧每秒拍摄,并提供人工精细修正的各视角液体与容器掩膜以及固定的评估划分。我们进一步提出了SplashSplat,其构建基于单一原则:仅在观测数据能够约束之处施加物理结构。由掩膜融合得到的逐帧液体符号距离场(SDF)提供几何信息,相邻SDF之间的水平集输运产生粗略速度场,沿该流动平流传播的拉格朗日载体针对每次新观测进行校正,并在覆盖丢失处重新播种,解码局部高斯以实现可微渲染。SplashSplat在我们的真实采集数据和合成基准上均优于最先进的动态高斯泼溅方法,其运动在物理上更合理且训练成本更低。同一表示还支持时间插值和风格迁移,无需重新优化。
cs.CV / 101 / 2609.20819

Can 4D Foundation Models Remember?

4D基础模型能够记忆吗?
He, Guangzhao, Averbuch-Elor, Hadar, Ma, Wei-Chiu
Abstract
Perceiving and remembering the visual world is fundamental to navigating and interacting with our environment. Current 4D foundation models, such as camera-controllable video models or 4D reconstruction models, can perceive and reconstruct dynamic environments, but how well they remember what they have perceived remains an open question. Existing benchmarks largely rely on pixel-level metrics and lack ground truth for objects once they leave the field of view, making them unable to evaluate visual memory in an object-centric manner against references. To fill this gap, we introduce PersistBench, a dataset and metric suite that leverages 360{\deg} videos as omniscient ground truth and proposes three evaluation aspects: object permanence, motion continuity, and appearance preservation. Evaluating various models across diverse categories reveals that current models can only maintain short-term consistency that degrades significantly once objects leave the field of view. Our findings highlight the gap between current model capabilities and robust visual memory ("seeing is not remembering"), providing guidance for future development of 4D foundation models. Dataset and code are available on the project page: https://guangzhaohe.com/persistbench.
Chinese Translation
感知并记忆视觉世界是导航和与环境交互的基础。当前的4D基础模型,例如相机可控的视频模型或4D重建模型,能够感知和重建动态环境,但它们对所感知内容的记忆能力如何仍是一个开放性问题。现有基准测试主要依赖像素级指标,并且缺乏物体离开视野后的真值(ground truth),因而无法以物体为中心的方式对照参考来评估视觉记忆。为填补这一空白,我们提出了PersistBench,一个利用360度视频作为全知真值的数据集与指标体系,并提出了三个评估维度:物体持久性(object permanence)、运动连续性(motion continuity)和外观保持性(appearance preservation)。对多种模型在不同类别上的评估表明,当前模型只能维持短期一致性,一旦物体离开视野,一致性便会显著下降。我们的发现揭示了当前模型能力与鲁棒视觉记忆之间的差距(“看见并不等于记住”),为4D基础模型的未来发展提供了指导。数据集和代码已发布在项目页面:https://guangzhaohe.com/persistbench。
机器学习 (Machine Learning)
107
cs.LG / 1 / 2609.19209

Generative Query Suggestion via Intent Coverage and Query-Level Credit Assignment

基于意图覆盖与查询级信用分配的生成式查询推荐
Liu, Xinpeng, Ma, Lu, Qiao, Jiayi, Zhou, Mengyu, Li, Linglong, Bian, Xiaofeng, Chen, Haonan, Jiang, Xiaoxi, Jiang, Guanjun
Abstract
Generative query suggestion aims to enhance user engagement by anticipating user intents and recommending relevant follow-up queries. A central challenge is to generate slates whose individual queries are useful while the slate covers distinct intents. We propose an Intent-Driven Query Suggestion Framework with dual-stage optimization. First, intent-aware diversity modeling constructs intent-aligned supervised fine-tuning (SFT) data and uses an Intent-Aware Diversity Reward to optimize intent coverage. Second, query-level credit assignment routes individual quality signals to the corresponding query tokens while sharing a slate-level diversity signal across the slate. Experiments on a large-scale production dataset, including online A/B testing and offline evaluation, show improvements in click-through rate, query quality, and intent coverage.
Chinese Translation
生成式查询推荐旨在通过预测用户意图并推荐相关的后续查询来提升用户参与度。一个核心挑战是生成的查询列表既要使其中每个查询都有用,又要使整个列表覆盖不同的意图。我们提出了一种意图驱动的查询推荐框架,采用双阶段优化。首先,意图感知的多样性建模构建与意图对齐的监督微调(SFT)数据,并使用意图感知多样性奖励(Intent-Aware Diversity Reward)来优化意图覆盖。其次,查询级信用分配将个体质量信号路由至相应的查询词元,同时在查询列表内共享列表级的多样性信号。在大规模生产数据集上的实验,包括在线A/B测试和离线评估,表明该方法在点击率、查询质量和意图覆盖方面均取得了提升。
cs.LG / 2 / 2609.19213

Layer-wise Curriculum Learning for Efficient LLM Compression

基于逐层课程学习的高效大语言模型压缩
Lee, Donggeon, Na, Dooyeon, Oh, Seungmin, Ryu, Jongbin
Abstract
In this paper, we introduce layer-wise curriculum learning for efficient LLM compression. The proposed method facilitates the knowledge transfer from the teacher model to the student model, utilizing a curriculum learning approach that begins with easier optimization tasks and progressively tackles harder ones. In order to adopt the layer-wise learning in LLM compression, we partition the whole model into multiple segments consisting of layers, thereby enabling more computationally efficient knowledge transfer for LLMs. Based on our theoretical analysis of cumulative error phenomenon, layer-wise curriculum learning accelerates convergence while stabilizing the knowledge transfer process. In addition, we present a feature caching method with a multi-threading strategy to efficiently address feature misalignment across layers, maximizing GPU utilization. Consequently, our method exhibits advanced model compression performance, as well as high computational efficiency in terms of minimized memory usage and short training hours. Experiments on multiple datasets show that the proposed method achieves state-of-the-art performance while reducing GPU memory usage and training hours by more than 50\% on BERT and GPT-2. Moreover, it outperforms the other pruning methods on LLaMA-family and Qwen models under the same training hours, with a lower GPU memory footprint.
Chinese Translation
本文提出了一种用于高效压缩大语言模型(LLM)的逐层课程学习方法。该方法采用课程学习策略,从较易的优化任务开始,逐步过渡到更难的任务,从而促进知识从教师模型向学生模型的迁移。为了在LLM压缩中采用逐层学习,我们将整个模型划分为多个由若干层组成的片段,从而实现计算效率更高的知识迁移。基于对累积误差现象的理论分析,逐层课程学习在加速收敛的同时,稳定了知识迁移过程。此外,我们提出了一种结合多线程策略的特征缓存方法,以高效解决跨层特征错位问题,最大化GPU利用率。因此,我们的方法展现出先进的模型压缩性能,同时在最小化内存占用和缩短训练时间方面具有很高的计算效率。在多个数据集上的实验表明,所提出的方法在BERT和GPT-2上将GPU内存占用和训练时间减少了50%以上,同时达到了最先进的性能。此外,在相同训练时间下,该方法在LLaMA系列和Qwen模型上的剪枝效果优于其他方法,且GPU内存占用更低。
cs.LG / 3 / 2609.19242

Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training

面向高效分布式长上下文扩散语言模型训练的块并行方法
Suresh, Tarun, Chaturvedi, Pranshu, Kang, Hangoo, Shroff, Parth, Khare, Ishan S., Kumbong, Hermann, Mirhoseini, Azalia
Abstract
Block diffusion language models (BDLMs) combine autoregressive dependencies across blocks with parallel denoising within blocks, but long-context training is constrained by distributed attention communication and activation memory. Conventional context parallelism (CP) shards the combined clean-plus-corrupted sequence by position, communicating shared clean K/V together with block-specific corrupted K/V and their gradients. We observe that the BDLM objective separates over target blocks. We introduce block parallelism (BP), a new distributed parallelism dimension that assigns each corrupted-block computation to one rank. To scale BP to long contexts, we introduce context-sharded block parallelism (CSBP), which also shards the shared clean sequence across those ranks. CSBP keeps corrupted K/V and gradients local, avoids replicated clean prefixes, and preserves BDLM training semantics. On 16 H200 GPUs at 256K context, CSBP improves throughput over the best baseline by 1.18-1.45x for supervised fine-tuning and 1.27-1.33x for conversion of autoregressive models to BDLMs, while matching or reducing peak HBM. Full-model speedup reaches 1.61x at 512K. On eight H100 GPUs, CSBP accelerates DFlash2 speculative-decoder training by 2.48x at 512K and 7.59x at 1M. In matched 12-hour DiffusionGemma 26B-A4B SFT runs, CSBP achieves higher pass rates at every trained checkpoint on SWE-bench Verified and Terminal-Bench Lite. Code: https://github.com/ScalingIntelligence/Turbo-dLLM
Chinese Translation
块扩散语言模型(Block Diffusion Language Models, BDLMs)将跨块的自回归依赖与块内并行去噪相结合,但长上下文训练受到分布式注意力通信和激活内存的限制。传统的上下文并行(Context Parallelism, CP)按位置对干净序列与加噪序列的合并序列进行分片,需要同时传递共享的干净K/V以及各块特有的加噪K/V及其梯度。我们观察到BDLM的目标函数在目标块之间是可分离的。据此,我们提出块并行(Block Parallelism, BP),一种新的分布式并行维度,将每个加噪块的计算分配给一个计算节点(rank)。为了将BP扩展到长上下文,我们进一步提出上下文分片块并行(Context-Sharded Block Parallelism, CSBP),它同时将共享的干净序列在这些节点之间进行分片。CSBP使加噪K/V及其梯度保持在本地,避免了干净前缀的重复存储,并保持了BDLM的训练语义。在16块H200 GPU上、256K上下文长度下,CSBP相较于最优基线,在有监督微调(SFT)中将吞吐量提升1.18-1.45倍,在自回归模型转换为BDLM的任务中提升1.27-1.33倍,同时持平或降低了峰值HBM内存占用。在512K上下文下,完整模型加速比达到1.61倍。在8块H100 GPU上,CSBP在512K上下文下将DFlash2推测解码器训练加速2.48倍,在1M上下文下加速7.59倍。在同等12小时的DiffusionGemma 26B-A4B SFT运行中,CSBP在SWE-bench Verified和Terminal-Bench Lite上的每个训练检查点均取得了更高的通过率。代码:https://github.com/ScalingIntelligence/Turbo-dLLM
cs.LG / 4 / 2609.19243

Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices

用于词-文档矩阵谱共聚类的随机化SVD近似方法
Mazdarani, Fateme, Toxtli, Carlos
Abstract
Spectral co-clustering is a useful tool for discovering latent structure in word-document matrices, but its reliance on singular value decomposition (SVD) can make standard formulations expensive on high-dimensional data. This paper presents two randomized approximations for normalized spectral co-clustering of bipartite text data when the numbers of document and word clusters may differ. The first method uses randomized SVD through random projection, while the second combines partial SVD with element-wise random sampling. Across real-world and synthetic datasets, both methods reduce runtime relative to the full-SVD baseline, but their behavior depends on matrix sparsity. The random projection method is the more reliable approximation across the tested settings, whereas the sampling-based method is most useful on denser matrices and provides limited benefit on already sparse text data. These results show that randomized approximations for spectral co-clustering should be selected according to the underlying structure of the data.
Chinese Translation
谱共聚类是发现词-文档矩阵潜在结构的有效工具,但其对奇异值分解(SVD)的依赖使得标准方法在高维数据上计算成本较高。本文针对文档聚类数与词聚类数可能不同的情形,提出了两种用于二部文本数据归一化谱共聚类的随机化近似方法。第一种方法通过随机投影实现随机化SVD,第二种方法则将部分SVD与逐元素随机采样相结合。在真实数据集和合成数据集上的实验表明,与全SVD基线相比,两种方法均能缩短运行时间,但其表现依赖于矩阵的稀疏性。在所有测试设置中,随机投影方法是更可靠的近似方法,而基于采样的方法更适用于较稠密的矩阵,对于本身已很稀疏的文本数据收益有限。这些结果表明,谱共聚类的随机化近似方法应根据数据的底层结构进行选择。
cs.LG / 5 / 2609.19279

Radio-Frequency Convolutional Neural Networks

射频卷积神经网络
Gao, Zhihui, Ma, Shi-Yuan, Chen, Yiran, Englund, Dirk, Chen, Tingjun
Abstract
Running artificial intelligence (AI) models directly on edge devices such as smartphones, wearables, and drones offers low latency, pervasive scalability, and data privacy, but these devices rarely carry the computing capability that modern neural networks demand. Edge accelerators have been developed in response, yet each adds computing hardware to devices already constrained in size, weight, power, and cost (SWaP-C). An alternative lies in what these devices already carry: the frequency mixer in every wireless radio multiplies signals in time, natively performing convolution in the frequency domain. Here we introduce radio-frequency convolutional neural networks (RF-CNNs), which repurpose existing communication hardware for CNN inference. Multi-channel convolutions are mapped onto frequency tones for a passive mixer to execute in a single pass. We experimentally demonstrate that RF-CNN runs deep CNNs up to 26.4 million parameters and nine layers from classification of wireless signals and images to controllable image generation, close to full-precision performance. Because the weights arrive over the air and the analog hardware is shared with communication, the edge device spends energy only on data preparation and readout-down to 0.72 femtojoules per multiply-accumulate, two orders of magnitude less than it would cost on an added digital processor. These results suggest that deployed wireless infrastructure can bring efficient, state-of-the-art AI inference to the billions of devices it already connects.
Chinese Translation
在智能手机、可穿戴设备和无人机等边缘设备上直接运行人工智能(AI)模型,具有低延迟、普遍可扩展性和数据隐私等优势,但这些设备通常不具备现代神经网络所需的计算能力。为应对这一问题,人们开发了边缘加速器,然而每一种加速器都在本已在尺寸、重量、功耗和成本(SWaP-C)方面受限的设备上增加了计算硬件。另一种思路则在于利用这些设备已有的部件:每个无线射频中的频率混频器可以在时域中对信号进行相乘,从而天然地在频域中执行卷积运算。本文提出了射频卷积神经网络(RF-CNN),将现有通信硬件重新用于CNN推理。多通道卷积被映射到多个频率音(frequency tones)上,由无源混频器一次性执行。我们通过实验证明,RF-CNN可以运行参数量高达2640万、层数达九层的深度CNN,应用涵盖无线信号与图像分类以及可控图像生成,性能接近全精度水平。由于权重通过空中传输、且模拟硬件与通信功能共享,边缘设备仅需在数据准备和读出上消耗能量——低至每次乘累加运算0.72飞焦耳,比使用额外数字处理器所需的能耗低两个数量级。这些结果表明,已部署的无线基础设施可以为其已连接的数十亿设备带来高效、最先进的AI推理能力。
cs.LG / 6 / 2609.19288

Learning-Induced Dynamical Transition in Recurrent Neural Networks

循环神经网络中由学习诱发的动力学转变
Vaidya, Varun
Abstract
Learning in recurrent neural networks can fundamentally reshape their underlying dynamics, transforming initially chaotic activity into stable task-dependent behavior. We develop a non-equilibrium dynamical mean-field theory(DMFT) to describe this transition during learning. We show that a slow feedback-driven learning process generates an evolving effective feedback strength that drives the network through a transition from chaotic to stable dynamics defined by a bifurcation of the DMFT solution. By deriving the two-time correlation function throughout learning, we identify a critical feedback strength and a corresponding learning rate dependent critical time separating these regimes. The transition arises from the progressive deformation of an effective dynamical landscape by the growing learned feedback structure. Starting from the untrained state, the theory predicts the time evolution of the network output during training and shows quantitative agreement with numerical simulations.
Chinese Translation
循环神经网络中的学习能够从根本上重塑其潜在的动力学,将最初混沌的活动转化为稳定的任务相关行为。我们发展了一种非平衡动力学平均场理论(DMFT)来描述学习过程中的这一转变。我们证明,缓慢的反馈驱动学习过程会产生一个不断演化的有效反馈强度,驱动网络经历一个由DMFT解的分岔所定义的、从混沌到稳定动力学的转变。通过推导整个学习过程中的双时关联函数,我们确定了一个临界反馈强度以及相应的依赖于学习速率的临界时间,用以区分这两种动力学状态。该转变源于不断增长的学习反馈结构对有效动力学景观的逐步形变。从初始未训练状态出发,该理论预测了训练期间网络输出的时间演化,并与数值模拟结果在定量上吻合。
cs.LG / 7 / 2609.19337

Personalized Federated Hierarchical Gaussian Processes for Privacy-Preserving Modeling of Heterogeneous Distributed Systems

用于异构分布式系统隐私保护建模的个性化联邦层次高斯过程
Xie, Xianjian, Yan, Hao
Abstract
We present Personalized Federated Hierarchical Gaussian Processes (pFedHGP) for probabilistic regression and classification when data are distributed across heterogeneous clients. Each client's latent function decomposes into (i) a shared global component, (ii) a client-specific deviation that shares the global kernel structure, and (iii) a flexible local residual. Sparse inducing-variable approximations and federated variational inference keep raw data local while the server synchronizes only low-dimensional statistics for the shared component. Full predictive distributions support uncertainty-aware decisions. In application studies, pFedHGP attains perfect fault classification in press tonnage monitoring using 13.77% of labeled cycles and recovers geographic zones in federated air-quality modeling without centralizing station-level time series. An Instantaneous Linear Mixing Model viewpoint links the hierarchy to multi-output Gaussian processes for correlated sensors.
Chinese Translation
我们提出了个性化联邦层次高斯过程(Personalized Federated Hierarchical Gaussian Processes,pFedHGP),用于数据分布在异构客户端上时的概率回归与分类任务。每个客户端的潜在函数分解为三部分:(i)共享的全局分量,(ii)共享全局核结构的客户端特定偏差,以及(iii)灵活的局部残差。稀疏诱导变量近似和联邦变分推断使原始数据保留在本地,服务器仅同步共享分量的低维统计量。完整的预测分布支持具有不确定性感知的决策。在应用研究中,pFedHGP 仅使用13.77%的标注周期便在压力机吨位监测中实现了完美的故障分类,并在联邦空气质量建模中无需集中站点级时间序列即可恢复地理分区。瞬时线性混合模型(Instantaneous Linear Mixing Model)的视角将该层次结构与面向相关传感器多输出的高斯过程联系起来。
cs.LG / 8 / 2609.19356

How to Guide Your Language Flow

如何引导你的语言流
Dilip, Rohit, Chen, Tianrong, Wang, Yuyang, Van Valen, David, Susskind, Joshua, Bautista, Miguel Angel
Abstract
We introduce a new method to guide flow matching models. Our approach, which we call probe guidance, uses the frozen internal states of an existing diffusion model to construct a guidance signal. This works using a similar principle as autoguidance, but eliminates the need for an additional forward pass at inference time and provides a reliable path to ensure that the weak and strong model share similar dynamics. We apply and benchmark this method on continuous diffusion language models, where probe guidance sets a new state-of-the-art performance on unconditional generation. When applied to a 1.7B diffusion language model, probe guidance consistently improves on multiple choice question answering benchmarks. Using our probes, we study the traditional autoguidance setting where the strong model is a weak checkpoint, and find that the weak model must come from a low-entropy region of training. These findings both provide a practical way to improve diffusion language models and shed light on the actual mechanism behind autoguidance, which is currently poorly understood.
Chinese Translation
我们提出了一种引导流匹配模型的新方法。我们将该方法称为探针引导,它利用现有扩散模型的冻结内部状态来构建引导信号。其工作原理与自引导类似,但无需在推理时进行额外的前向传播,并提供了一条可靠的途径来确保弱模型与强模型具有相似的动力学特性。我们将该方法应用于连续扩散语言模型并进行基准测试,探针引导在无条件生成任务上创造了新的最先进性能。当应用于一个1.7B参数的扩散语言模型时,探针引导在多项选择问答基准上取得了一致的提升。利用我们的探针,我们研究了传统自引导设置(其中强模型是一个弱检查点)的情形,发现弱模型必须来自训练的低熵区域。这些发现既为改进扩散语言模型提供了一种实用方法,也揭示了目前理解甚少的自引导背后的真实机制。
cs.LG / 9 / 2609.19359

Smart Insole Human Activity Recognition for Continuous Monitoring in Elderly Care

用于老年人护理持续监测的智能鞋垫人体活动识别
Rios, Edwin, Garcia, Antony, Yuan, Fengpei, Huang, Xinming
Abstract
Falls in older adults are often preceded by changes in mobility, balance, and postural transitions. This paper presents a wireless smart insole platform and machine-learning workflow for recognizing sitting, standing, walking, and unstable walking from plantar-pressure and inertial signals. Each insole integrates 16 active pressure-sensing locations and a six-dimensional IMU stream consisting of tri-axial acceleration and angular velocity. Data were collected from 15 healthy adults at 80~Hz and segmented into overlapping windows. Window length and candidate model families were first screened with stratified 10-fold cross-validation; the primary performance estimate was then obtained with participant-independent 5-fold Stratified Group cross-validation, ensuring that all windows from a participant remained in a single fold. Under this protocol, Histogram-Based Gradient Boosting (HGB) achieved macro-F1 scores of 0.954 and 0.959 for the left and right feet, respectively, and 0.980 with bilateral sensing. A compact 1D-CNN evaluated with the same participant-independent folds did not significantly outperform HGB ($p=0.0625$). The results show that low-profile footwear sensing can infer activity state from pressure and IMU measurements for participants unseen during training, establishing a basis for activity monitoring and fall prevention in elderly care.
Chinese Translation
老年人的跌倒往往发生在行动能力、平衡能力和姿势转换发生变化之前。本文提出了一种无线智能鞋垫平台及机器学习工作流程,用于根据足底压力和惯性信号识别坐姿、站姿、行走和不稳定行走。每个鞋垫集成了16个主动压力传感点和一个六维IMU数据流(包括三轴加速度和角速度)。数据采集自15名健康成年人,采样频率为80 Hz,并被分割为重叠的窗口。首先通过分层10折交叉验证对窗口长度和候选模型系列进行筛选,随后采用与参与者无关的5折分层分组交叉验证获得主要性能估计,以确保同一参与者的所有窗口都保留在同一折中。在该协议下,基于直方图的梯度提升模型(HGB)在左脚和右脚上分别取得了0.954和0.959的宏平均F1分数,采用双脚传感时达到0.980。在相同的与参与者无关的折数划分下评估的紧凑型一维卷积神经网络(1D-CNN)并未显著优于HGB(p=0.0625)。结果表明,低轮廓鞋类传感能够针对训练中未见过的参与者,从压力和IMU测量数据中推断活动状态,为老年人护理中的活动监测和跌倒预防奠定了基础。
cs.LG / 10 / 2609.19363

Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not

Stiefel注意力:Transformer投影矩阵的几何何时主导优化器选择——以及何时不主导
Guerrero, Rubén Darío
Abstract
The query and key projections $\WQ,\WK$ in attention are almost always trained by Euclidean optimizers with no constraint on their geometry. We constrain them to the Stiefel manifold and optimize them there with a Riemannian Adam that carries one scalar second moment per frame, caps its step by a trust region, and retracts polarly. Four propositions prove this update is steepest descent in the embedded metric, independent of gradient scale, well conditioned, and exactly $\mathrm{O}(d)$-equivariant, each certified numerically in \texttt{float64}. A fifth supplies the mechanism: weight decay has \emph{identically zero} Riemannian gradient on $\St(d,r)$, since $W = W I_r$ lies in the normal space, so the learned attention geometry survives the collapse cycles that decay drives through the rest of the model. On modular arithmetic grokking, a single run holds $97.0\%$ validation accuracy at epoch 20\,000 against the baseline's $61.1\%$---an unstable endpoint we report as evidence for the mechanism rather than as an effect size. On CIFAR-10 patches the same rule gains $\mathbf{+8.98}$\,pp over 12 paired starts ($t{=}60.6$, $12/12$), and the gap widens with data rather than eroding. The step rule earns this: a fixed-step Riemannian update is degree one in the gradient, so it moves $24$--$40\times$ less per step than an identically shaped AdamW matrix---its frames barely leave their initialization, and freezing them outright costs only $0.28$\,pp. An ablation credits the whole gain to making the step scale free, and nothing measurable to the projector or to equivariance. A negative result sharpens the account: gauge removal cannot motivate the method, because a direction along which the loss is invariant carries no gradient at all.
Chinese Translation
注意力机制中的查询和键投影矩阵 $W_Q, W_K$ 几乎总是由无几何约束的欧氏优化器进行训练。我们将其约束到 Stiefel 流形上,并用黎曼 Adam 在流形上对其进行优化,该优化器为每帧携带一个标量二阶矩,通过信赖域限制步长,并采用极分解进行收缩(retraction)。四个命题证明该更新在嵌入度量下是最速下降的,与梯度尺度无关,条件数良好,且精确满足 $\mathrm{O}(d)$-等变性,每个命题均在 \texttt{float64} 精度下通过数值验证。第五个命题提供了机制解释:权重衰减在 $\mathrm{St}(d,r)$ 上的黎曼梯度恒等于零,因为 $W = W I_r$ 位于法空间中,因此所学到的注意力几何能够在衰减驱动模型其余部分经历的坍缩周期中存活下来。在模运算涌现(grokking)任务中,单次运行在第 20000 轮保持 97.0% 的验证准确率,而基线仅为 61.1%——我们将这一不稳定的终点报告为机制的证据,而非效应量。在 CIFAR-10 图像块上,同一规则在 12 组配对启动中取得 +8.98 个百分点的提升($t=60.6$,12/12),且该差距随数据量增加而扩大而非缩小。这一步长规则功不可没:固定步长的黎曼更新关于梯度是一次齐次的,因此其每步移动量比形状完全相同的 AdamW 矩阵少 24–40 倍——其帧几乎不离开初始化位置,而将它们完全冻结也仅损失 0.28 个百分点。消融实验将全部增益归因于使步长尺度自由,而对投影器或等变性没有任何可测量的贡献。一个阴性结果进一步澄清了这一论述:规范去除无法为该方法提供动机,因为损失函数沿其不变的方向根本不携带梯度。
cs.LG / 11 / 2609.19374

Machine-Learning Assessment of the Predictive Value of Inflammatory Biomarkers for Cognitive Impairment in an Older Hispanic Adult Cohort

机器学习评估老年西班牙裔成年人队列中炎症生物标志物对认知障碍的预测价值
Garcia, Antony, Britton, Gabrielle, Villarreal, Alcibiades, Oviedo, Diana, Rangel, Giselle, Huang, Xinming
Abstract
Small clinical tabular datasets require interpretable machine learning because deep learning is often impractical and ensemble models can be difficult to inspect. A key pitfall is that statistical significance does not necessarily imply predictive utility. Using data from the Panama Aging Research Initiative--Health Disparities (PARI-HD) cohort (n=165), we implemented a leakage-safe threshold-likelihood Bernoulli/Categorical Naive Bayes (BNB/CNB) classifier. Within every training fold, each continuous predictor was reduced to a supervised chi-square-derived state, while income entered the model through a categorical likelihood. All data-dependent steps were performed within repeated stratified 10-fold cross-validation with 30 repeats. The demographic baseline achieved a ROC-AUC of 0.630 +/- 0.017. I-309 (CCL1) was the dominant incremental feature, increasing AUC by 0.110, with paired DeLong tests yielding p<0.05 in 100% of repeats. In the pre-specified primary analysis, I-309 produced a fixed-partition DeLong p=0.0018, with robustness assessed across 200 random partitions, where the median p-value was 0.0011. Within the exploratory family of 18 candidate markers, I-309 achieved a Benjamini-Hochberg-adjusted q=0.032 on the frozen partition and satisfied q<0.05 in 85% of random partitions, whereas no other marker demonstrated reliable incremental predictive value. Because the fitted model is an inspectable table of thresholds and class-conditional probabilities, these results identify I-309/CCL1 as an interpretable candidate feature for tabular prediction of cognitive impairment, pending external validation.
Chinese Translation
小型临床表格数据集需要可解释的机器学习方法,因为深度学习通常不切实际,且集成模型往往难以审查。一个关键的陷阱在于,统计显著性并不必然意味着预测效用。利用巴拿马老龄化研究计划——健康差异(PARI-HD)队列的数据(n=165),我们实现了一种防泄漏的阈值似然Bernoulli/Categorical朴素贝叶斯(BNB/CNB)分类器。在每个训练折中,每个连续型预测变量均被转化为基于监督式卡方检验推导的状态,而收入则通过类别似然进入模型。所有依赖数据的步骤均在30次重复的分层10折交叉验证内完成。人口统计学基线模型的ROC-AUC为0.630±0.017。I-309(CCL1)是最主要的增量特征,可将AUC提高0.110,且配对DeLong检验在100%的重复中均得到p<0.05。在预先设定的主分析中,I-309在固定划分下DeLong检验p=0.0018,并在200个随机划分中评估了稳健性,其中位p值为0.0011。在包含18个候选标志物的探索性分析中,I-309在固定划分下经Benjamini-Hochberg校正后q=0.032,并在85%的随机划分中满足q<0.05,而其他标志物均未表现出可靠的增量预测价值。由于所拟合的模型是一个由阈值和类条件概率组成的可审查表格,这些结果将I-309/CCL1确定为用于认知障碍表格化预测的可解释候选特征,但仍需外部验证。
cs.LG / 12 / 2609.19383

FCx: An algorithm for finding Feasible Counterfactual Explanations

FCx:一种寻找可行反事实解释的算法
Markou, Kleopatra, Kalogeraki, Vana, Gunopulos, Dimitrios
Abstract
Counterfactual (CF) explanations identify changes that alter an input's classification. While existing methods produce realistic and low-cost CFs, they often fail to ensure feasibility, by suggesting non-constructive modifications or incompatible with future changes (e.g., changing an individual's race to secure a job offer). We introduce a refinement of CF explanations that explicitly enforces feasibility. Our approach is the first to efficiently generate CFs that are realistic, low-cost and feasible. We accommodate both hard feasible constraints, specified by domain knowledge users, and soft feasible constraints, inferred automatically via causal inference from the dataset. Our method, Feasible Counterfactual Explanations (FCx), is based on a modified Variational Autoencoder (VAE) optimized with a multi-factor loss function. We measure the cost of a change based on the absolute change in values (proximity) as well as the number of features changed (sparsity) while realism is measured based on the LOF for density estimation, guaranteeing that CFs reside in densely populated regions. Extensive experiments on four public datasets show that our approach matches state-of-the-art performance across multiple metrics while guaranteeing feasibility.
Chinese Translation
反事实(Counterfactual,CF)解释旨在识别能够改变输入分类结果的变化。尽管现有方法能够生成真实且低代价的反事实解释,但它们往往无法保证可行性,例如给出非建设性的修改建议或与未来变化不相容的修改(如通过改变个人的种族来获得工作机会)。我们提出了一种反事实解释的改进方法,显式地强制保证可行性。我们的方法是首个能够高效生成兼具真实性、低代价和可行性的反事实解释的方法。本方法同时支持硬可行性约束(由用户根据领域知识指定)和软可行性约束(通过因果推断从数据集中自动推断得到)。我们的方法即可行反事实解释(Feasible Counterfactual Explanations,FCx),基于一个改进的变分自编码器(Variational Autoencoder,VAE),并使用多因子损失函数进行优化。我们基于数值的绝对变化量(邻近性)以及被改变特征的数量(稀疏性)来衡量变化代价,同时基于局部离群因子(LOF)进行密度估计来衡量真实性,从而保证反事实解释位于数据密集分布的区域。在四个公开数据集上的大量实验表明,我们的方法在多项指标上达到了最先进的性能,同时保证了可行性。
cs.LG / 13 / 2609.19414

Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation

通过选择性奖励刺激改进离线目标条件强化学习
Zhang, Jing
Abstract
Goal-conditioned reinforcement learning aims to learn policies that reach specified goals, but remains challenging in offline settings with sparse rewards and long-horizon dependencies. In such settings, goal-completion information can be temporally distant from the early decisions that enable success, while offline value estimation introduces additional error. We study this issue from a reward-propagation perspective and show, in a stylized delayed-goal setting, how goal-directed value separation can become small relative to local estimation error. Motivated by this analysis, we propose Reward Stimulation Implicit Q-Learning (RSIQL), a simple non-hierarchical method that introduces additional reward signals at progress-making intermediate states in offline trajectories. RSIQL uses an auxiliary goal-conditioned value function to identify intermediate states estimated to make progress toward the goal and applies reward stimulation to provide less-delayed training supervision. Unlike hierarchical methods, RSIQL does not learn a separate high-level subgoal policy. Experiments on D4RL goal-reaching benchmarks and OGBench show that RSIQL improves over goal-conditioned IQL on average and achieves performance competitive with hierarchical offline goal-conditioned methods, while retaining a simple flat policy structure.
Chinese Translation
目标条件强化学习旨在学习能够到达指定目标的策略,但在具有稀疏奖励和长时程依赖的离线设置中仍然具有挑战性。在此类设置中,目标完成信息在时间上可能与促成成功的早期决策相距甚远,而离线值估计又会引入额外的误差。我们从奖励传播的角度研究这一问题,并在一个简化的延迟目标设置中展示了目标导向的值分离如何相对于局部估计误差变得很小。受此分析的启发,我们提出了奖励刺激隐式Q学习(Reward Stimulation Implicit Q-Learning,RSIQL),这是一种简单的非分层方法,它在离线轨迹中取得进展的中间状态处引入额外的奖励信号。RSIQL使用一个辅助的目标条件值函数来识别估计能向目标推进的中间状态,并施加奖励刺激以提供延迟更少的训练监督。与分层方法不同,RSIQL不学习单独的高层子目标策略。在D4RL目标到达基准和OGBench上的实验表明,RSIQL在平均性能上优于目标条件IQL,并取得了与分层离线目标条件方法相当的性能,同时保持了简单的扁平策略结构。
cs.LG / 14 / 2609.19437

Bayesian Optimization with Rich Auxiliary Information via LLMs

利用大语言模型融合丰富辅助信息的贝叶斯优化
Gupta, Tejus, Karagözlü, Efe Mert, Sonker, Rohit, Póczos, Barnabás, Schnieder, Jeff
Abstract
Bayesian Optimization (BO) is widely used for optimizing expensive black-box functions, yet many real-world optimization problems contain substantially richer information than function evaluations alone. Examples include training curves in hyperparameter optimization, expert notes and images in scientific experimentation, and prior knowledge about where optima may lie. We show that large language models (LLMs) can effectively leverage such rich auxiliary information to guide optimization. Motivated by these findings, we develop three methods for incorporating auxiliary information into BO using LLMs. Across hyperparameter optimization benchmarks and a real-world nuclear fusion optimization task, our methods consistently outperform both standard BO and existing LLM-based optimization approaches. Our results demonstrate the effectiveness of LLMs for leveraging rich auxiliary information in BO.
Chinese Translation
贝叶斯优化(Bayesian Optimization, BO)被广泛用于优化昂贵的黑盒函数,然而许多现实世界的优化问题所包含的信息远比单纯的函数评估丰富得多。例如,超参数优化中的训练曲线、科学实验中的专家笔记和图像,以及关于最优解可能位置先验知识。我们证明了大语言模型(LLMs)能够有效利用这类丰富的辅助信息来引导优化。基于这些发现,我们提出了三种利用大语言模型将辅助信息融入贝叶斯优化的方法。在超参数优化基准测试和真实世界的核聚变优化任务中,我们的方法均一致优于标准贝叶斯优化方法及现有的基于大语言模型的优化方法。我们的结果证明了大语言模型在贝叶斯优化中利用丰富辅助信息的有效性。
cs.LG / 15 / 2609.19453

Sharpness-Aware Minimization (SAM) Improves Classification Accuracy of Bacterial Raman Spectral Data Enabling Portable Diagnostics

锐度感知最小化(SAM)提升细菌拉曼光谱数据的分类精度,助力便携式诊断
Zareno, Kaitlin, Dewbury, Jarett, Sorooshyari, Siamak K., Mobahi, Hossein, Tadesse, Loza F.
Abstract
Antimicrobial resistance is expected to claim 10 million lives per year by 2050, and resource-limited regions are most affected. Raman spectroscopy is a novel pathogen diagnostic approach promising rapid and portable antibiotic resistance testing within a few hours, compared to days when using gold standard methods. However, current algorithms for Raman spectra analysis 1) are unable to generalize well on limited datasets across diverse patient populations and 2) require increased complexity due to the necessity of non-trivial pre-processing steps, such as feature extraction, which are essential to mitigate the low-quality nature of Raman spectral data. In this work, we address these limitations using Sharpness-Aware Minimization (SAM) to enhance model generalization across a diverse array of hyperparameters in clinical bacterial isolate classification tasks. We demonstrate that SAM achieves accuracy improvements of up to 10.5% on a single split, and an increase in average accuracy of 2.7% across all splits in spectral classification tasks over the traditional optimizer, Adam. These results display the capability of SAM to advance the clinical application of AI-powered Raman spectroscopy tools.
Chinese Translation
预计到2050年,抗微生物药物耐药性每年将夺去1000万人的生命,而资源有限的地区受影响最为严重。拉曼光谱是一种新型病原体诊断方法,与需要数天的金标准方法相比,有望在几小时内实现快速、便携的抗生素耐药性检测。然而,当前的拉曼光谱分析算法存在以下问题:1)在跨不同患者群体的有限数据集上泛化能力不足;2)由于需要进行复杂的预处理步骤(如特征提取)而增加了复杂度,而这些步骤对于缓解拉曼光谱数据质量较低的问题至关重要。在本研究中,我们采用锐度感知最小化(Sharpness-Aware Minimization, SAM)来增强模型在临床细菌分离株分类任务中跨多种超参数设置的泛化能力,从而解决上述局限性。结果表明,在光谱分类任务中,SAM 在单一数据划分上的准确率提升高达10.5%,在所有数据划分上的平均准确率较传统优化器 Adam 提升2.7%。这些结果展示了 SAM 在推动人工智能驱动的拉曼光谱工具临床应用方面的能力。
cs.LG / 16 / 2609.19466

Enhanced Agriculture-informed Neural Network by Domain Knowledge

基于领域知识增强的农业信息神经网络
Lin, Ci, Li, Futong, Chong-Wu, Rose, Yeap, Tet, Kiringa, Iluju
Abstract
Accurate prediction of nitrous oxide (N2O) emissions from agriculture is important for assessing environmental impacts and supporting sustainable farming. However, prediction remains difficult because N2O emissions result from complex interactions among soil properties, climate, biochemical processes, and management practices, while high-quality observations are limited. Deep learning models can capture nonlinear relationships but often lack physical interpretability and may generalize poorly across environmental conditions. We propose the Knowledge-enhanced Agriculture-informed Neural Network (KAINN), a hybrid neural-mechanistic framework that extends the Agriculture-informed Neural Network by incorporating domain knowledge about fertilizer diffusion, soil respiration, and water-filled porosity. We evaluate KAINN using CNN, LSTM, and Transformer architectures across multiple growing seasons and input-feature configurations. The results show that KAINN generally provides lower root mean square error and mean absolute error and higher R-squared values than purely data-driven models and the original AINN. Analysis of the learned interfaces also shows smoother and more physically consistent parameter trajectories with reduced uncertainty. These findings demonstrate that incorporating environmental knowledge into neural networks can improve the reliability, interpretability, and generalization of agricultural N2O-emission predictions.
Chinese Translation
准确预测农业氧化亚氮(N2O)排放对于评估环境影响和支持可持续农业至关重要。然而,由于N2O排放源于土壤性质、气候、生化过程和管理措施之间的复杂相互作用,且高质量观测数据有限,预测仍然十分困难。深度学习模型能够捕捉非线性关系,但往往缺乏物理可解释性,且在不同环境条件下的泛化能力可能较差。我们提出了知识增强的农业信息神经网络(Knowledge-enhanced Agriculture-informed Neural Network, KAINN),这是一种神经-机理混合框架,通过融合关于肥料扩散、土壤呼吸和充水孔隙度的领域知识,对农业信息神经网络(Agriculture-informed Neural Network, AINN)进行了扩展。我们采用CNN、LSTM和Transformer架构,在多个生长季和多种输入特征配置下对KAINN进行了评估。结果表明,与纯数据驱动模型和原始AINN相比,KAINN总体上具有更低的均方根误差和平均绝对误差,以及更高的R²值。对所学习接口的分析还表明,其参数轨迹更加平滑、物理一致性更强,且不确定性更低。这些发现证明,将环境知识融入神经网络可以提高农业N2O排放预测的可靠性、可解释性和泛化能力。
cs.LG / 17 / 2609.19476

Search at the Cost of Sampling: Nearly-Instant Latent Space Bayesian Optimization

以采样换搜索:近即时的潜在空间贝叶斯优化
Fan, Donney, Doumont, Colin, Kalisz, Aleksandra, Duckworth, Paul, Gardner, Jacob R., Moss, Henry, Pleiss, Geoff
Abstract
Generative models are increasingly central to many de novo discovery pipelines, in which designs are generated at scale and filtered through virtual screens to determine a set of candidates to experimentally validate. While Bayesian optimization (BO) is a natural fit for this setting, as it uses past evaluations to guide future proposals, the computational overhead required for its sequential decision-making becomes a bottleneck when virtual screens are relatively cheap. We make BO practical in this regime by exploiting the unique combination of a linear model constrained to a spherical domain where high-dimensional latents concentrate. We build off recent work justifying the use of linear surrogates, while deriving nearly closed-form solutions to the surrogate modelling and acquisition problems that exploit spherical symmetry. The result is at least a 100x speedup over state-of-the art baselines, with matching or improved performance across molecular and image generation benchmarks. Altogether, our method makes BO a practical drop-in for de novo pipelines where it was previously too slow to consider.
Chinese Translation
生成模型在许多从头发现(de novo)流程中日益居于核心地位,在这些流程中,设计被大规模生成,并通过虚拟筛选过滤出一组候选方案以进行实验验证。贝叶斯优化(Bayesian Optimization, BO)由于能够利用过去的评估来指导未来的提议,因而非常适合这一场景;但当虚拟筛选的成本相对较低时,其顺序决策所需的计算开销便成为瓶颈。我们通过利用线性模型与球形域约束这一独特组合——高维潜在表示恰好集中于球形域——使贝叶斯优化在该场景下变得实用。我们在近期关于线性代理模型合理性的研究基础上,推导出利用球对称性的近乎闭式解,以解决代理建模与采集函数两个问题。其结果是相比最先进的基线方法至少获得100倍的加速,同时在分子生成和图像生成基准上保持相当或更优的性能。总之,我们的方法使贝叶斯优化成为从头发现流程中一种实用的即插即用方案,而此前它因速度过慢而难以被纳入考虑。
cs.LG / 18 / 2609.19499

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

样本数量并不足够:候选生成策略决定大语言模型测试时扩展的能耗与性能
Kashaniyan, Mobina, Jannesari, Ali
Abstract
Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget is often described by the number of generated candidates, N. However, N tells us how many candidates are generated, not how they are executed. The same candidate budget can be produced in one batched generation call or split across several sequential calls with smaller batch sizes. We first study the effect of increasing N on reasoning accuracy using Phi-3-mini and Qwen2.5-1.5B on 500 GSM8K prompts. As expected, increasing N from 1 to 8 improves accuracy by 8.4 percentage points for Phi-3-mini and 18.4 points for Qwen2.5-1.5B. However, accuracy alone does not show the systems cost of using a larger candidate budget. We therefore fix N = 8 and compare four generation schedules: 1x8, 2x4, 4x2, and 8x1, where axb denotes a generation calls with b candidates per call. We measure latency, throughput, GPU-hours, and gross GPU-device energy while keeping the total candidate count fixed. On A100 GPUs, eight serial calls use 4.64-4.86x as much gross GPU-device energy and have 5.77-6.12x the P95 latency of one batched call with eight candidates. The same pattern appears across three independently scheduled A100 nodes per model and in short-output SciQ/V100 experiments. These results show that candidate count alone is not enough to describe the systems cost of multi-candidate test-time scaling. When candidates are independent and memory allows it, fewer generation calls with larger batch sizes are more efficient. Evaluations should therefore report not only candidate count and accuracy, but also generation schedule and GPU-level systems metrics.
Chinese Translation
测试时扩展(test-time scaling)可以通过生成并组合多个候选响应来提升大语言模型的推理能力。在基于采样的方法中,推理预算通常由生成的候选数量N来描述。然而,N只说明了生成了多少个候选,并未说明它们是如何被执行的。相同的候选预算可以在一次批量生成调用中完成,也可以拆分为多次具有更小批量的串行调用。我们首先使用Phi-3-mini和Qwen2.5-1.5B在500个GSM8K提示上研究了增大N对推理准确率的影响。正如预期,将N从1增加到8使Phi-3-mini的准确率提升8.4个百分点,使Qwen2.5-1.5B提升18.4个百分点。然而,仅凭准确率无法体现使用更大候选预算的系统成本。因此,我们固定N=8,比较四种生成调度方案:1x8、2x4、4x2和8x1,其中axb表示a次生成调用、每次调用生成b个候选。在保持候选总数固定的情况下,我们测量了延迟、吞吐量、GPU小时数以及总GPU设备能耗。在A100 GPU上,八次串行调用的总GPU设备能耗是一次生成八个候选的批量调用的4.64-4.86倍,P95延迟是其5.77-6.12倍。同样的模式在每个模型的三个独立调度的A100节点上以及短输出的SciQ/V100实验中均得到验证。这些结果表明,仅凭候选数量不足以描述多候选测试时扩展的系统成本。当候选相互独立且内存允许时,使用更少的生成调用和更大的批量更为高效。因此,评估不仅应报告候选数量和准确率,还应报告生成调度方案以及GPU层面的系统指标。
cs.LG / 19 / 2609.19521

LSTM-UT and Recurrent-Depth Transformers on Cellular Automata

LSTM-UT与循环深度Transformer在元胞自动机上的研究
Kavuncu, Aras
Abstract
Recurrent-depth Transformers apply shared computation repeatedly, but differ in how they retain information across steps. We compare a Block Universal Transformer (BUT), which carries only its current hidden state; CoTFormer, which also retains an expanding attention cache; and a new LSTM Universal Transformer (LSTM-UT) with bounded gated memory. On Rule 30 cellular automata, BUT extrapolates to unseen recurrent depths more reliably than CoTFormer, although its accuracy eventually degrades. State and cache interventions show that CoTFormer's failure depends on their interaction: correcting the current state can temporarily restore accuracy, while retained history can undermine that correction. In a delayed-recall task, BUT also outperforms CoTFormer despite lacking direct access to past states; CoTFormer does not reliably select the requested cached representation. LSTM-UT improves both depth extrapolation and delayed recall over these baselines. The results support bounded gated memory as an effective inductive bias for repeated computation and later retrieval in these tasks.
Chinese Translation
循环深度Transformer通过重复应用共享计算来工作,但它们在跨步骤保留信息的方式上有所不同。我们比较了三种模型:块通用Transformer(Block Universal Transformer, BUT),它仅携带当前隐藏状态;CoTFormer,它还保留一个不断扩展的注意力缓存;以及一种新的带有有界门控记忆的LSTM通用Transformer(LSTM-UT)。在Rule 30元胞自动机上,BUT向未见过的循环深度的外推能力比CoTFormer更可靠,尽管其精度最终会下降。对状态和缓存的干预实验表明,CoTFormer的失败取决于二者的交互作用:修正当前状态可以暂时恢复精度,而保留的历史信息可能会破坏这种修正。在一个延迟回忆任务中,尽管BUT无法直接访问过去的状态,它仍然优于CoTFormer;CoTFormer无法可靠地选择所请求的缓存表示。LSTM-UT在深度外推和延迟回忆两方面均优于这些基线模型。结果表明,有界门控记忆是这些任务中支持重复计算和后续检索的有效归纳偏置。
cs.LG / 20 / 2609.19539

Compressed Active Subspaces for Scalable Bayesian Inference

面向可扩展贝叶斯推断的压缩主动子空间方法
Flynn, Thomas, Jantre, Sanket, Yoon, Byung-Jun, Kim, Kibaek
Abstract
Active subspace methods provide a framework for quantifying predictive uncertainty in high-dimensional models by identifying and performing inference along parameter directions that have the greatest influence on the model output. However, the construction of active subspaces requires storing many full-dimensional model gradients, which becomes prohibitive as model size increases. We address this limitation by proposing Compressed Active Subspaces (CAS), a scalable approach that first maps the model parameters to a compressed space using a structured isometric embedding and then constructs the active subspace within this reduced parameterization. Our approach substantially reduces the memory required for active subspace construction and enables Bayesian inference for large models where standard active subspace methods become impractical. We demonstrate the scalability of CAS on neural networks of increasing size while maintaining predictive performance and robust uncertainty estimates.
Chinese Translation
主动子空间方法通过识别对模型输出影响最大的参数方向并沿这些方向进行推断,为高维模型的预测不确定性量化提供了一个框架。然而,主动子空间的构建需要存储大量全维度的模型梯度,随着模型规模的增大,这一要求变得难以承受。为解决这一局限,我们提出了压缩主动子空间(Compressed Active Subspaces, CAS),这是一种可扩展的方法:首先利用结构化等距嵌入将模型参数映射到一个压缩空间,然后在该降维参数化中构建主动子空间。我们的方法大幅降低了主动子空间构建所需的内存开销,使得在标准主动子空间方法变得不切实际的大型模型上进行贝叶斯推断成为可能。我们在规模不断增大的神经网络上演示了CAS的可扩展性,同时保持了预测性能和稳健的不确定性估计。
cs.LG / 21 / 2609.19559

FedFIbOS: Fisher Importance based Optimal Submodelling for Heterogeneous Federated Learning

FedFIbOS:基于Fisher重要性的异构联邦学习最优子模型方法
Afzal, Yasmeen, Deng, Jeremiah D., Zhang, Haibo
Abstract
Heterogeneous federated learning requires clients with diverse computational capacities to collaboratively train a global model, where each client trains a capacity-constrained submodel. Existing methods select submodel parameters using heuristic importance measures---most prominently parameter magnitude---without theoretical justification for why these measures support convergence. We identify a fundamental gap: existing parameter selection criteria lack theoretical grounding in the convergence framework, partial client participation introduces additional estimation effects in the Fisher scores. We propose \textbf{FedFIbOS}: Fisher Importance-based Optimal Submodelling for heterogeneous federated learning, using Fisher Information in a principled criterion derived from minimizing submodel masking error. %We formally establish when magnitude selection is equivalent to Fisher selection fail under non-IID heterogeneous federated learning. We theoretically formulate submodel selection through a Fisher-weighted quadratic masking surrogate and show that the raw Fisher top-$k$ rule implemented by FedFIbOS solves this surrogate under a Fisher-dominant ranking condition. The resulting method retains the convergence structure of the underlying masked federated optimization bound. Fisher scores are efficiently estimated from empirical diagonal Fisher information using squared gradients, enabling stable and adaptive parameter selection without additional optimization overhead. Experiments on CIFAR-10, CIFAR-100, and AGNews under pathological and Dirichlet non-IID settings show FedFIbOS achieves ${\approx}10\%$ higher accuracy than the state of the art, with improvements becoming more pronounced under stronger heterogeneity.
Chinese Translation
异构联邦学习要求具有不同计算能力的客户端协同训练一个全局模型,其中每个客户端训练一个受容量约束的子模型。现有方法采用启发式重要性度量来选择子模型参数——最典型的是参数幅值——但缺乏这些度量为何能支持收敛的理论依据。我们发现了一个根本性的缺口:现有参数选择准则缺乏收敛理论框架下的理论支撑,且部分客户端参与会在Fisher得分中引入额外的估计效应。我们提出FedFIbOS:一种用于异构联邦学习的基于Fisher重要性的最优子模型方法,采用由最小化子模型掩码误差推导出的有原则的Fisher信息准则。我们通过Fisher加权的二次掩码代理问题从理论上形式化子模型选择问题,并证明FedFIbOS所实现的原始Fisher top-$k$规则在Fisher主导排序条件下可求解该代理问题。由此得到的方法保留了底层掩码联邦优化界所具有的收敛结构。Fisher得分通过平方梯度从经验对角Fisher信息中高效估计,从而实现稳定且自适应的参数选择,且不引入额外的优化开销。在CIFAR-10、CIFAR-100和AGNews数据集上、在病态(pathological)和Dirichlet非独立同分布(non-IID)设置下的实验表明,FedFIbOS比现有最先进方法的准确率高出约10%,且在异构性更强时提升更为显著。
cs.LG / 22 / 2609.19616

The Complexity Kink: A Prompt-Side Structural Complexity Index for Code-Generation Reliability

复杂性拐点:一种用于代码生成可靠性的提示侧结构复杂性指数
Hernandez, Michael, Zhao, Tian
Abstract
Complexity measured from generated code is failure-dependent: a difficult prompt can yield a short failing program and be assigned low output complexity. We introduce a six-dimension prompt-side structural-complexity index scored before generation and kept separate from correctness. We select 5,000 Python prompts across six bands of a preliminary single-rater rubric. Four out-of-panel LLM raters rescore the locked prompts, giving 19,997 score rows; composite inter-rater reliability is ICC = 0.872 on the 4,998 prompts with all four ratings. We evaluate 21 models per prompt, yielding 105,000 generations. In the unadjusted mean-pooled analysis, pass rate has a nonmonotone breakpoint at composite 13.75, with 79.9% at or below and 87.6% above. This is not a universal failure cutoff. Task-type fixed effects shift the breakpoint to 10.75 and cut the regime gap from 7.6 to 2.1 points. A construction-frame control shifts it to 8.50 with a raw gap of -3.5 points, and neither frame alone reproduces the pooled +7.6-point change. Model-specific fits include 16 upward and five downward changes. A 365-prompt audit-clean extension matches the original five-model estimates at bins 15 and 16 but adds only 14 prompts above bin 16. Among zero-pass generations with computable Lizard complexity, 28.5% pair a prompt composite above 8 with output complexity at most 10. Human agreement is moderate and rater-dependent on a disagreement-enriched calibration set; paraphrase and cross-language rescoring preserve score ordering. Overidentification tests reject the joint restrictions on the six dimensions, so we treat the composite as an index and make no causal interpretation of the 2SLS estimates. The contribution is a pre-generation measurement framework and a bounded observational analysis of reliability regimes.
Chinese Translation
从生成代码中测得的复杂性依赖于失败状态:一个困难的提示可能产生一个简短的失败程序,从而被赋予较低的输出复杂性。我们提出了一种六维度的提示侧结构复杂性指数,该指数在生成前进行评分,并与正确性保持分离。我们根据初步的单评分者评分标准选取了六个区间内共5,000个Python提示。四位面板外的LLM评分者对锁定的提示进行重新评分,得到19,997条评分记录;在所有四位评分者均给出评分的4,998个提示上,综合评分者间信度为ICC = 0.872。我们对每个提示评估21个模型,共产生105,000个生成结果。在未经调整的均值汇总分析中,通过率在综合得分13.75处出现非单调断点,该点及以下的通过率为79.9%,该点以上为87.6%。这并非一个通用的失败分界线。任务类型固定效应将断点移至10.75,并将区间差距从7.6个百分点缩小至2.1个百分点。构建框架控制将断点移至8.50,原始差距为-3.5个百分点,且任一框架单独都无法重现汇总分析中+7.6个百分点的变化。针对具体模型的拟合结果包括16个向上变化和5个向下变化。一个365个提示的审计清洗扩展在第15和第16个分箱处与原先五个模型的估计相吻合,但在第16分箱以上仅增加了14个提示。在具有可计算Lizard复杂度的零通过生成结果中,28.5%呈现提示综合得分高于8而输出复杂度不超过10的组合。在一个富含分歧的校准集上,人类一致性为中等水平且依赖于评分者;改写和跨语言重新评分能够保持得分排序。过度识别检验拒绝了六个维度上的联合约束,因此我们将该综合得分视为一个指数,不对2SLS估计作因果解释。本研究的贡献在于提供了一个生成前的测量框架,以及对可靠性区间的有边界观测分析。
cs.LG / 23 / 2609.19640

A Policy Profile for Croissant: Refusal as a Property of the Dataset

Croissant 的策略配置文件:作为数据集属性的拒绝机制
Chernov, Alexander
Abstract
Croissant is the de facto machine-readable descriptor for ML datasets: JSON-LD over schema.org. Since version 1.1 it also carries data use conditions, recommending DUO and ODRL for them. What no version specifies is how any of them is evaluated: no decision procedure, no bound on evaluation cost, no outcome for a condition an implementation cannot evaluate, no record of what was checked, and nothing on composition with caller-side authority. We supply that half. An additive profile lets a dataset declare the operations it admits and the conditions under which it admits them, over a closed set of five operators whose decision procedure is given in full, so a gate decides from the descriptor alone and records what it checked. Two corpora evaluate it and their evidence is kept apart. Three descriptors that gated a real nf-core pipeline give the deployment result: decisions from a profile document match the gate's native descriptor record for record, stripping the layer leaves a valid Croissant document, and the added cost is 11.7 $\mu$s against a 119 $\mu$s decision. A corpus generated from the profile's grammar gives the breadth, covering every operator, refusal class and conformance clause. Across its valid cases, 552 complete decision records agree three ways -- native descriptor, profile terms, and the same policy as ODRL in usageInfo. The carrier is therefore not the contribution; the evaluation semantics is. Finally, caller-bound and data-bound policies range over non-overlapping state spaces, so neither permit set contains the other.
Chinese Translation
Croissant 是机器学习数据集事实上的机器可读描述符:基于 schema.org 的 JSON-LD。自 1.1 版本起,它还承载数据使用条件,并推荐使用 DUO 和 ODRL 来表达这些条件。然而,任何版本都没有说明如何评估这些条件:没有决策程序、没有对评估成本的约束、没有针对实现无法评估的条件的结果规定、没有对已检查内容的记录,也没有关于与调用方权限组合的说明。我们补齐了这缺失的一半。一个可叠加的配置文件使数据集能够声明其允许的操作以及允许这些操作的条件,基于一个由五个操作符构成的封闭集合,其决策程序被完整给出,从而使门控(gate)仅凭描述符即可做出决策并记录其检查内容。两个语料库对该方法进行了评估,且其证据被分开保留。三个用于门控真实 nf-core 流水线的描述符给出了部署结果:来自配置文件文档的决策与门控原生描述符的记录逐条一致,剥离该层后仍得到有效的 Croissant 文档,且新增成本为 11.7 微秒,而决策本身耗时 119 微秒。由该配置文件的文法生成的一个语料库提供了广度覆盖,涵盖了每个操作符、拒绝类别和一致性条款。在其所有有效用例中,552 份完整决策记录在三个方面相互一致——原生描述符、配置文件条款,以及 usageInfo 中以 ODRL 表示的相同策略。因此,载体并非贡献所在,评估语义才是。最后,调用方绑定的策略与数据绑定的策略的取值状态空间互不重叠,因此二者的许可集合互不包含。
cs.LG / 24 / 2609.19670

CoRe: Coherence and Relational Alignment for Multivariate Time Series Forecasting

CoRe:面向多元时间序列预测的相干性与关系对齐
Lin, Xiaoyu, Duan, Huiran, Liu, Yining, Wu, Zhixiang, Lin, Chu, Lu, Lin
Abstract
Direct forecasting has become a standard paradigm for multivariate time-series forecasting because it predicts the full future horizon in a single pass. However, its training objective is often still decomposed into pointwise errors such as MSE. Such objectives provide stable supervision, but they do not explicitly preserve the structure of the future trajectory: temporal coherence within each variable and relational consistency across variables can both be weakened. We propose CoRe, a model-agnostic learning objective for direct multivariate forecasting. CoRe replaces pointwise supervision with two output-space constraints: a frequency coherence loss that aligns predicted and target spectra, and a low-rank relational graph loss that matches sampled pairwise differences in a target-derived PCA subspace. The resulting objective introduces no trainable parameters and can be applied to existing forecasting backbones by changing only the loss. Experiments on standard benchmarks show that CoRe improves strong baselines, compares favorably with recent forecasting objectives, and remains effective across different backbones, datasets, and hyperparameter settings overall consistently.
Chinese Translation
直接预测已成为多元时间序列预测的标准范式,因为它能够一次性预测完整的未来时域。然而,其训练目标通常仍被分解为均方误差(MSE)等逐点误差。此类目标虽然提供了稳定的监督信号,但并未显式地保留未来轨迹的结构:每个变量内部的时间相干性以及变量之间的关系一致性都可能被削弱。我们提出了CoRe,一种面向直接多元预测的模型无关学习目标。CoRe用两个输出空间约束取代逐点监督:一是频率相干性损失,用于对齐预测谱与目标谱;二是低秩关系图损失,在由目标导出的PCA子空间中匹配采样得到的成对差异。由此得到的目标不引入任何可训练参数,只需修改损失函数即可应用于现有的预测骨干网络。在标准基准上的实验表明,CoRe能够改进多个强基线模型,与近期的预测目标相比表现更优,并且在不同骨干网络、数据集和超参数设置下总体上保持了一致的有效性。
cs.LG / 25 / 2609.19674

Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models

守恒性为物理世界模型带来稳定性,因式分解带来反事实性
Wang, Yufeng, Priye, Parivesh, Wei, Lu, Ling, Haibin
Abstract
A learned simulator can reproduce its training conditions accurately yet fail in two distinct ways once those conditions change. Over long rollouts, small errors accumulate until the trajectory drifts away from physically plausible behavior; under an intervention on a physical parameter, the model may continue to follow the law seen during training rather than the intervened one. We show that these two failures require different structural remedies. Evolving a learned energy with a symplectic integrator preserves the geometry of the conservative dynamics and keeps rollouts bounded and physically meaningful for up to $100\times$ the training horizon, while equal-capacity predictors, an energy-regularized predictor, and a tuned neural ODE diverge. By contrast, encoding the physical coupling through an explicit linear factorization enables the model to follow a never-seen sign of that coupling, whereas an unrestricted parameterization remains locked to the training law. Crucially, the two mechanisms are separable: removing the structure responsible for long-horizon stability leaves counterfactual transfer intact, while removing the factorized coupling destroys counterfactual transfer without eliminating stability. This double dissociation, established with matched controls that remove or replace one structural component at a time, persists beyond the headline three-body system and remains visible when the physical state must be inferred from pixels rather than provided directly. The result is a concrete design principle for physical world models: long-horizon stability and changed-law generalization arise from distinct structural commitments, and each can be imposed deliberately without requiring the other.
Chinese Translation
一个学习得到的模拟器能够准确复现其训练条件,但一旦这些条件发生变化,便可能以两种截然不同的方式失效。在长时间的滚动推演(rollout)中,微小误差不断累积,直至轨迹偏离物理上合理的行为;而在对某个物理参数进行干预时,模型可能继续遵循训练中看到的规律,而非被干预后的规律。我们证明这两种失效需要不同的结构性补救措施。利用辛积分器演化学习得到的能量函数,可以保持保守动力学的几何结构,使滚动推演在长达训练范围 $100\times$ 的时域内保持有界且物理上有意义;而同等容量的预测器、能量正则化预测器以及经过调优的神经 ODE(neural ODE)则会发散。相比之下,通过显式的线性因式分解来编码物理耦合,使模型能够遵循从未见过的耦合符号方向,而无约束的参数化则被锁定在训练规律上。至关重要的是,这两种机制是可分离的:移除负责长时程稳定性的结构不会破坏反事实迁移,而移除因式分解的耦合则会摧毁反事实迁移却不消除稳定性。这种双重分离(double dissociation)通过匹配的对照实验得以确立——每次仅移除或替换一个结构组件——其结论不仅限于核心的三体系统,并且在必须从像素而非直接给定的物理状态进行推断时依然成立。由此得到一个针对物理世界模型的具体设计原则:长时程稳定性与规律变化下的泛化能力源自不同的结构性承诺,二者都可以被刻意施加,而无需依赖另一方。
cs.LG / 26 / 2609.19695

Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems

基于意见动力学的异构物联网系统联邦学习联盟形成方法
Hanjri, Mohammed El, Abouaomar, Anas, Tembine, Hamidou, Kobbane, Abdellatif
Abstract
Federated learning (FL) enables privacy-preserving, on-device training across heterogeneous Internet-of-Things (IoT) deployments such as smart-city water-metering networks, where each smart meter observes a household-specific consumption time series. Under such statistical heterogeneity, the standard Federated Averaging (FedAvg) aggregation averages dissimilar local models into a single global model that may fail to capture client-specific patterns. We address this by forming client coalitions directly in the local-weight space and aggregating at the coalition level. Extending a prior weight-driven coalition-formation scheme, we model coalition formation as a Hegselmann-Krause (HK) bounded-confidence opinion-dynamics process acting on the local weights, and develop variants of the HK interaction based on Euclidean-distance and cosine-similarity confidence criteria. The framework is applied to short-term water-consumption forecasting with local Long Short-Term Memory (LSTM) models and evaluated against FedAvg, Per-FedAvg, FedProx, and FedAvg with Euclidean-distance or cosine-similarity coalition formation. Experiments on a real smart-metering dataset of water consumption show that the proposed HK-based coalition formation produces stable, endogenous coalition structures within at most ten inner iterations, incurs no additional client-side computation or communication compared to FedAvg, and reduces the average MAE by up to 54% relative to FedAvg, 39% relative to FedProx, and 24% relative to Per-FedAvg, while achieving the highest global accuracy (83-85%).
Chinese Translation
联邦学习(FL)能够在异构物联网部署(如智慧城市水表计量网络)中实现隐私保护的设备端训练,其中每个智能水表观测特定家庭的用水时间序列。在这种统计异构性下,标准的联邦平均算法将差异较大的本地模型平均为单一全局模型,可能无法捕捉客户端特有的模式。为此,我们直接在本地权重空间中构建客户端联盟,并在联盟层面进行聚合。在已有的基于权重的联盟形成方案基础上,我们将联盟形成建模为作用于本地权重的 Hegselmann-Krause(HK)有界置信意见动力学过程,并开发了基于欧氏距离和余弦相似度置信准则的HK交互变体。该框架被应用于基于本地长短期记忆网络(LSTM)模型的短期用水量预测,并与FedAvg、Per-FedAvg、FedProx以及采用欧氏距离或余弦相似度联盟形成的FedAvg进行对比评估。在真实智能水表用水数据集上的实验表明,所提出的基于HK的联盟形成方法最多在十次内部迭代内即可产生稳定、内生的联盟结构,与FedAvg相比不增加任何客户端计算或通信开销,平均MAE相对FedAvg降低最多54%,相对FedProx降低39%,相对Per-FedAvg降低24%,同时取得最高的全局准确率(83–85%)。
cs.LG / 27 / 2609.19709

Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits

比值比汤普森采样:基于对比的多臂老虎机算法的规范与设计指南
Kim, Sulgi
Abstract
Batched multi-armed bandits update on a service's own schedule, and the usual implementation carries each arm's absolute reward rate from one update to the next. When the shared level moves between batches, that memory goes stale even though the comparisons between arms may not have. Odds-Ratio Thompson Sampling (OR-TS) instead carries the joint posterior over log-odds contrasts and fits the common level afresh in every batch, marginalizing it out. This paper specifies that update, places it inside a Bayesian bandit agent with two controls, decay for how much past evidence survives an update and aggressiveness for how sharply belief becomes allocation, and evaluates it against absolute-rate memory. Across 86 public A/B series the level varies about twenty-five times more than the contrast. In prespecified synthetic environments a moving level costs absolute-rate memory five times the regret and leaves the best arm below a majority of traffic in 7 of 20 runs, against none for OR-TS. In a policy simulation built from 71 real experiments, where the contrasts are too small to resolve, expected-click differences stay within 0.1% for 58 of them, yet contrast memory still ends on the better arm more than twice as often. Where the contrasts themselves move, the bet fails, and that case is reported too.
Chinese Translation
分批多臂老虎机(batched multi-armed bandits)按照服务自身的日程进行更新,通常的实现方式会将每个臂的绝对奖励率从一次更新延续到下一次更新。当共享水平(shared level)在批次之间发生移动时,这种记忆便会过时,尽管各臂之间的对比关系可能并未改变。比值比汤普森采样(Odds-Ratio Thompson Sampling, OR-TS)则改为保留对数比值比对比(log-odds contrasts)的联合后验分布,并在每个批次中重新拟合共同水平并将其边缘化。本文对这一更新过程进行了规范化说明,将其嵌入一个具有两个控制参数的贝叶斯老虎机智能体中——衰减(decay)控制过去证据在每次更新后存留的比例,激进程度(aggressiveness)控制信念转化为流量分配的陡峭程度——并将其与绝对率记忆方法进行对比评估。在86个公开A/B测试序列中,水平的波动幅度约为对比关系的二十五倍。在预先设定的合成环境中,移动的水平使绝对率记忆产生的遗憾(regret)增加至五倍,且在20次运行中有7次最优臂获得的流量低于半数,而OR-TS则无一发生。在一个基于71个真实实验构建的策略模拟中,由于对比值过小而难以分辨,其中58个实验的期望点击差异保持在0.1%以内,但对比记忆算法最终停留在更优臂上的频率仍是绝对率记忆的两倍以上。当对比关系本身发生移动时,这种策略便会失效,本文也报告了该情形。
cs.LG / 28 / 2609.19717

Learn Your Own Thoughts: Abstract Token Curriculum

学习你自己的思考:抽象标记课程学习
Gatmiry, Khashayar, Ghosh, Avrajit, Mirtaheri, Parsa, Lee, Jason D., Haghtalab, Nika, Abbe, Emmanuel, Bartlett, Peter
Abstract
Large Language Models (LLMs) have achieved remarkable reasoning capabilities by utilizing chain-of-thought (CoT) as a scratchpad for intermediate stages of thinking. However, CoT techniques require explicit supervision on thinking tokens, which requires rich, task-specific data. In this work, we propose Abstract Token Curriculum (ATC), a novel curriculum learning framework that elicits effective continuous intermediate representations without direct supervision or manual scratchpad design. ATC gradually increases problem complexity through a sequence of distributions, training the model to develop internal abstract ``thoughts'' in the continuous representation space. This paper provides both theoretical and experimental evidence for the benefits of ATC and its advantages over previous methods for training continuous thoughts. Theoretically, we show that for learning parity functions with single-layer softmax attention using ATC, attention naturally focuses on the CoT tokens in the context that provide the ``easiest path'' to predicting the next token. Experimentally, we show ATC's effectiveness on graph reachability and arithmetic learning tasks.
Chinese Translation
大型语言模型(LLMs)通过利用思维链(Chain-of-Thought, CoT)作为中间思考阶段的草稿板,已经展现出卓越的推理能力。然而,CoT技术需要对思考标记进行显式监督,这需要丰富的、特定任务的数据。在本工作中,我们提出了抽象标记课程学习(Abstract Token Curriculum, ATC),这是一种新颖的课程学习框架,能够在无需直接监督或人工设计草稿板的情况下,引出有效的连续中间表示。ATC通过一系列分布逐步增加问题的复杂度,训练模型在连续表示空间中发展出内部的抽象“思维”。本文从理论和实验两方面为ATC的优势及其相较于以往训练连续思维方法的长处提供了证据。在理论上,我们证明了在使用ATC训练单层softmax注意力模型学习奇偶性函数时,注意力会自然地聚焦于上下文中能为预测下一个标记提供“最易路径”的CoT标记。在实验上,我们展示了ATC在图可达性和算术学习任务上的有效性。
cs.LG / 29 / 2609.19748

Alliance Beats Isolation: Unifying Heterogeneous Allied Datasets Improves Classifier Performance

联盟胜过孤立:统一异构联盟数据集提升分类器性能
Palshikar, Girish Keshav
Abstract
In many application domains, such as student dropout, insurance fraud, loan approval, and machine failures, several labelled public datasets are available where (i) data is about the same type of objects but the set of actual underlying objects are disjoint; and (ii) the class labels are same; and (iii) the feature spaces of the datasets are largely distinct (heterogeneous), with a few shared features. We call such datasets as allied. A single classifier cannot be trained on both datasets together, and one classifier trained on one dataset cannot be tested on the other. In this paper, we propose a method to merge the feature-spaces into a single feature-space for a pair of given allied heterogeneous datasets. We then use a matrix completion method to create a unified dataset based on the merged feature-space. The hypothesis is that the merged representation facilitates the transfer of classification knowledge from one dataset to another. We conduct experiments on several pairs of allied, heterogeneous datasets and several classifiers to demonstrate that any classifier trained on the unified representation always outperforms classifiers separately trained on the constituent allied datasets on several pairs of allied datasets. This work provides an easy way to substantially improve classifier performance by unifying and using multiple allied datasets together.
Chinese Translation
在许多应用领域中,例如学生辍学预测、保险欺诈、贷款审批和机器故障检测,存在多个带标签的公开数据集,这些数据集具有以下特点:(i)数据涉及相同类型的对象,但实际底层对象集合互不相交;(ii)类别标签相同;(iii)数据集的特征空间在很大程度上是不同的(异构的),仅有少量共享特征。我们将此类数据集称为联盟数据集(allied)。单个分类器无法同时在两个数据集上联合训练,而在一个数据集上训练的分类器也无法在另一个数据集上测试。在本文中,我们提出了一种方法,将给定的一对联盟异构数据集的特征空间合并为单一特征空间。然后,我们使用矩阵补全方法基于合并后的特征空间创建一个统一的数据集。我们的假设是,合并后的表示能够促进分类知识从一个数据集向另一个数据集的迁移。我们在多对联盟异构数据集和多种分类器上进行了实验,结果表明,在统一表示上训练的任何分类器在多对联盟数据集上的表现始终优于分别在各个成员数据集上单独训练的分类器。这项工作提供了一种简便方法,通过统一并联合使用多个联盟数据集来显著提升分类器性能。
cs.LG / 30 / 2609.19768

OceanMoE: Structured Conditional Sparse Computation for Long-Horizon Multivariate Ocean Forecasting

OceanMoE:面向长时程多变量海洋预报的结构化条件稀疏计算
Zhu, Yishun, Wang, Jian
Abstract
Multivariate ocean forecasting must exploit shared evolution in a coupled ocean system while adapting to the heterogeneous statistical and dynamical characteristics of different prediction variables and locations. Fully shared models may lack the flexibility to handle this heterogeneity, whereas fully independent models discard the common ocean context shared across variables. The key question is how to retain shared context in a unified model while allowing computation to specialize according to the prediction target and local state. We propose OceanMoE, a structured conditional sparse Mixture-of-Experts framework that combines sharing and specialization for multivariate ocean forecasting. OceanMoE fuses cross-variable information to construct target-specific local representations and uses them to perform content-conditioned sparse routing at each spatial location, with the number of active experts adapted to router confidence. In the decoder, routing is augmented with a learned geographic bias parameterized by spherical-harmonic spatial bases, while shared residual and seasonal pathways provide common cross-variable and month-dependent context. Experiments on long-horizon autoregressive ORAS5 forecasting show that OceanMoE lowers aggregate forecasting error in both evaluated settings and maintains lower geometric-mean normalized RMSE than the corresponding baselines over most later rollout months. Routing analyses further show that expert allocation varies with prediction targets and spatial locations. These results support structured conditional computation as a modeling strategy for balancing shared ocean context with adaptive specialization.
Chinese Translation
多变量海洋预报既要利用耦合海洋系统中的共同演变规律,又要适应不同预报变量和位置所具有的异质统计与动力特征。完全共享的模型可能缺乏处理这种异质性的灵活性,而完全独立的模型则丢弃了各变量之间共有的海洋背景信息。关键问题在于如何在统一模型中保留共享背景的同时,使计算能够根据预报目标和局部状态进行专门化。我们提出OceanMoE,一种结合共享与专门化的结构化条件稀疏混合专家框架,用于多变量海洋预报。OceanMoE融合跨变量信息以构建特定目标的局部表示,并利用这些表示在每个空间位置执行基于内容的稀疏路由,激活专家的数量可根据路由器的置信度自适应调整。在解码器中,路由通过由球谐空间基参数化的可学习地理偏置得到增强,同时共享残差通路和季节通路提供共同的跨变量和随月份变化的背景信息。在长时程自回归ORAS5预报实验中,OceanMoE在两种评估设置下均降低了总体预报误差,并在大多数后续滚动预报月份中保持比相应基线更低的几何平均归一化RMSE。路由分析进一步表明,专家分配随预报目标和空间位置而变化。这些结果支持将结构化条件计算作为平衡共享海洋背景与自适应专门化的建模策略。
cs.LG / 31 / 2609.19776

PhyRestore: Physics-Structured Latent-Factor Restoration

PhyRestore:物理结构化潜在因子恢复
Shafee, Ahmed, Lahiri, Chayan
Abstract
Estimating temporal soil-loss change is challenging when physically meaningful input factors are noisy or corrupted, particularly because substantial changes are rare relative to the large number of locations exhibiting little change. We study this problem through the Revised Universal Soil Loss Equation (RUSLE) and introduce PhyRestore, a physics-structured latent-factor restoration framework. Rather than directly predicting soil-loss change or correcting a degraded physical estimate, PhyRestore restores corrupted physical factors and reconstructs temporal change through the known physical relationship. We evaluate PhyRestore in a watershed-scale bitemporal raster setting under isolated and simultaneous corruption of rainfall erosivity and cover management, comparing it with the degraded RUSLE estimate and Direct RF, XGBoost, MLP, and CNN models. Factor restoration improves high-magnitude recovery when the corrupted factors remain identifiable, but its advantage weakens under joint corruption, sparse positive extremes, and factor values outside the training support.
Chinese Translation
当具有物理意义的输入因子存在噪声或损坏时,估计土壤流失的时间变化极具挑战性,尤其是因为相对于大量几乎无变化的位置而言,显著变化的情况十分罕见。我们基于修正通用土壤流失方程(RUSLE)研究这一问题,并提出PhyRestore——一个物理结构化的潜在因子恢复框架。PhyRestore并非直接预测土壤流失变化或修正退化的物理估计,而是恢复损坏的物理因子,并通过已知的物理关系重建时间变化。我们在流域尺度、双时相栅格的设定下评估PhyRestore,分别考虑降雨侵蚀力与覆盖管理因子被单独损坏和同时损坏的情形,并将其与退化的RUSLE估计以及Direct RF、XGBoost、MLP和CNN等模型进行比较。当损坏因子仍可辨识时,因子恢复能改善大幅值变化的恢复效果;但在联合损坏、正极端值稀疏以及因子值超出训练支撑范围的情况下,其优势会减弱。
cs.LG / 32 / 2609.19786

Learning-Based Reconstruction of Optical Properties in Bilayered Media from Single-distance Time-Resolved Reflectance Measurements

基于学习的双层介质光学性质重建:基于单距离时间分辨反射测量
Amendola, Caterina, Maffeis, Giulia, Buffoni, Lorenzo, Chicchi, Lorenzo, Coghi, Francesco, Fanelli, Duccio, Marino, Raffaele, Martelli, Fabrizio, Paoli, Riccardo, Pattelli, Lorenzo, Spinelli, Lorenzo
Abstract
The inverse problem of reconstructing optical properties, specifically absorption and scattering coefficients, in layered biological media from time-domain reflectance measurements remains a significant challenge for traditional analytical models. Inverse solvers based on the diffusion equation often struggle with structural heterogeneity, frequently yielding poor accuracy for superficial absorption and deep-layers scattering. In this work, we propose a machine learning framework as an alternative approach to reconstruct the optical properties of a bilayered medium, benchmarking its efficiency and accuracy against model-based algorithms. To overcome the intrinsic approximations of diffusion theory and inverse reconstruction, we generated a robust synthetic dataset of forward DTOF using exact Monte Carlo simulations at multiple source-detector distances. A machine learning pipeline was then trained on this dataset and validated against state-of-the-art model-based reconstruction methods. Besides the significant reconstruction speed-up, the machine learning approach achieves higher accuracy than model-based inverse solvers, further providing an estimate of the parameter space dimensionality without requiring any a priori information about the number of layers in the investigated geometry. Further enhancements in the reconstruction accuracy can be expected in future extensions of this work, by training the pipeline over multiple DTOF curves from the same medium, in a joint multi-distance reconstruction approach.
Chinese Translation
从时域反射测量中重建层状生物介质的光学性质(特别是吸收系数和散射系数)这一逆问题,对传统解析模型而言仍是一项重大挑战。基于扩散方程的逆求解器往往难以处理结构非均匀性,对于表层吸收和深层散射的重建精度常常较差。在本工作中,我们提出了一种机器学习框架作为重建双层介质光学性质的替代方法,并将其效率与精度与基于模型的算法进行了基准比较。为克服扩散理论和逆重建的固有近似,我们利用精确的蒙特卡洛模拟在多个光源-探测器距离下生成了鲁棒的前向DTOF(分布时间飞行)合成数据集。随后,在该数据集上训练了机器学习流程,并与最先进的基于模型的重建方法进行了验证比较。除了显著加快重建速度外,机器学习方法还取得了高于基于模型的逆求解器的重建精度,并且无需任何关于所研究几何结构层数的先验信息,即可对参数空间维度进行估计。在未来工作中,通过在同一介质的多条DTOF曲线上训练该流程,采用联合多距离重建方法,有望进一步提升重建精度。
cs.LG / 33 / 2609.19801

DeliveryGym: An RL Environment for Long-Horizon Embodied Agent Planning with Adaptive Curriculum

DeliveryGym:一个具有自适应课程的长时程具身智能体规划强化学习环境
Kang, Haoqiang, Zhang, Yiming, Guo, Yiyang, Li, Chuying, Shen, Jianzhi, Xu, Tianruo Rose, Ye, Xiaokang, Qin, Lianhui
Abstract
Executable environments enable LLM agents to learn from the consequences of their actions. For embodied agents, those consequences extend beyond whether the current task succeeds: completing a delivery can consume the time, energy, or money needed for later work. Learning to plan therefore requires environments that preserve these dependencies and turn them into feedback across a complete trajectory. We introduce DeliveryGym, a 3D environment for evaluating and training agents on continuous courier shifts. It couples multimodal tool interaction with persistent world dynamics and computes trajectory rewards from simulator events, making the costs of an agent's decisions available for reinforcement learning (RL). The environment also adapts future training shifts to the policy's observed weaknesses while keeping evaluation fixed. Across six models and 13 city maps, evaluation exposes a gap between reliably executing assigned deliveries and choosing and sequencing work over a shift. On the fixed test suite, RL improves Qwen3-VL-4B's net income by 54.3%, showing that learning from complete shifts improves performance under these coupled constraints. Adapting the training environment improves test income by 16.5% over uniform sampling at the same rollout budget, indicating that which situations an agent practices also matters. DeliveryGym provides an executable setting for studying how agents learn to coordinate deliveries and preserve resources for later orders within an episode.
Chinese Translation
可执行环境使大语言模型(LLM)智能体能够从其行动的后果中学习。对于具身智能体而言,这些后果不仅限于当前任务是否成功:完成一次配送可能消耗后续工作所需的时间、能量或金钱。因此,学习规划需要能够保留这些依赖关系、并将其转化为贯穿完整轨迹的反馈的环境。我们提出了 DeliveryGym,一个用于在连续快递员班次上评估和训练智能体的三维环境。它将多模态工具交互与持久的世界动态相结合,并基于模拟器事件计算轨迹奖励,使智能体决策的代价可用于强化学习(RL)。该环境还能根据策略所观察到的弱点自适应地调整未来的训练班次,同时保持评估固定。在六个模型和 13 张城市地图上的评估揭示了可靠执行已分配配送任务与在班次中选择和排序工作之间的差距。在固定测试集上,强化学习使 Qwen3-VL-4B 的净收入提升了 54.3%,表明从完整班次中学习能够提升智能体在这些耦合约束下的表现。在相同的 rollout 预算下,自适应调整训练环境比均匀采样进一步将测试收入提升了 16.5%,说明智能体练习何种情境同样重要。DeliveryGym 为研究智能体如何在单个回合内学会协调配送并为后续订单保留资源提供了一个可执行的研究环境。
cs.LG / 34 / 2609.19842

Beyond Flattened Tokens: Structure-Preserving EEG Decoding with Reusable TriDim Blocks

超越扁平化词元:基于可复用TriDim模块的结构保持式脑电解码
Su, Shiyue, Wang, Song, Zhan, Zekai, Zeng, Junjie, Lu, Ziling, Li, Zongsheng, Ye, Xinyuan, Ma, Zhiyuan, Shen, Xinke, Liu, Quanying
Abstract
Effective EEG decoding requires representations that preserve organization among channels, local waveform dynamics, and long-range temporal context. Existing EEG architectures often capture these structures using separate specialized modules or collapse them into a single token sequence, making it difficult to maintain their distinct roles and coordinate their interactions throughout the backbone. We propose TriDim, a reusable block that preserves the representation shape and keeps three EEG axes explicit: channel, sample position within each patch, and patch position across the recording. These axes correspond to spatial, short-term temporal, and long-term temporal information, respectively. Each TriDim block applies feed-forward transformations along individual axes and cross-axis attention to coordinate information exchange among them. By stacking TriDim blocks with a multi-level tri-axis readout, we construct TriDimEEG, a standalone EEG decoder. Under strict cross-subject evaluation on eight datasets spanning clinical diagnosis, sleep staging, motor imagery, and emotion recognition, TriDimEEG achieves the best overall performance among fifteen evaluated models, with a 4.3% relative improvement in average accuracy over the second-best model. Replacing Transformer blocks in three EEG foundation models with TriDim blocks yields an average relative improvement of 7.4% in downstream accuracy while reducing parameter counts by 17.0% to 47.3%. These results establish TriDim as an effective and reusable building block and TriDimEEG as a strong standalone EEG decoder. Code and parameters of TriDimEEG are available at https://github.com/ncclab-sustech/TriDim_model.
Chinese Translation
有效的脑电(EEG)解码需要能够保持通道间组织关系、局部波形动态以及长程时间上下文的表征。现有的EEG架构通常采用各自独立的专用模块来捕捉这些结构,或将它们压缩为单一词元序列,这使得在整个骨干网络中难以维持各结构的独特作用并协调其交互。我们提出TriDim,一种可复用的模块,它保持表征的形状不变,并显式地保留EEG的三个维度:通道、每个片段内的采样位置,以及整段记录中的片段位置。这三个维度分别对应空间信息、短时时间信息和长时时间信息。每个TriDim模块沿各个维度独立施加前馈变换,并通过跨维度注意力机制协调它们之间的信息交换。通过堆叠TriDim模块并采用多层级三维度读出机制,我们构建了TriDimEEG,一个独立的EEG解码器。在涵盖临床诊断、睡眠分期、运动想象和情绪识别的八个数据集上的严格跨被试评估中,TriDimEEG在十五个评估模型中取得了最佳的整体性能,平均准确率相对第二名模型提升了4.3%。将三个EEG基础模型中的Transformer模块替换为TriDim模块后,下游任务准确率平均相对提升7.4%,同时参数量减少17.0%至47.3%。这些结果确立了TriDim作为一种有效且可复用的基础构建模块,以及TriDimEEG作为一种强大的独立EEG解码器的地位。TriDimEEG的代码和参数可在 https://github.com/ncclab-sustech/TriDim_model 获取。
cs.LG / 35 / 2609.19858

Expected Hypervolume Maximization for Multiobjective Optimization under Uncertainties

不确定性条件下多目标优化的期望超体积最大化
Trappler, Victor
Abstract
The problem of multiobjective optimization under uncertainties is often approached by taking the expectation of each objective. In this work, we propose instead to formulate this as a Bayesian decision problem and to rely on the expected value of the hypervolume, which is to be maximized with respect to a finite set of input points. We show that this can be performed using methods based on gradients in a stochastic optimization framework, provided that care is taken with respect to dominated points. Moreover, in the absence of readily available differentiable code, we propose to use Gaussian Processes as differentiable surrogate models, in order to perform the optimization. An additional contribution in this work are some active learning strategies, through acquisition functions which helps construct a surrogate model well-designed for the multiobjective optimization problem at stake. These strategies are compared on simple analytical problems to assess their performances.
Chinese Translation
不确定性条件下的多目标优化问题通常通过计算每个目标的期望值来求解。在本工作中,我们提出将其表述为一个贝叶斯决策问题,并依赖于超体积(hypervolume)的期望值,即关于有限输入点集对该期望值进行最大化。我们证明,在随机优化框架下,只要对被支配点加以注意,这一过程可以借助基于梯度的方法来完成。此外,在缺乏现成可微分代码的情况下,我们提出使用高斯过程(Gaussian Processes)作为可微分的代理模型来执行优化。本工作的另一项贡献是提出了一些主动学习策略,即通过采集函数(acquisition functions)帮助构建针对所研究多目标优化问题而精心设计的代理模型。这些策略在简单的解析问题上进行了比较,以评估其性能。
cs.LG / 36 / 2609.19865

Pretrained Medical Representations for the Practical Screening of Drug Repositioning Candidates

基于预训练医学表示的药物重定位候选药物实用筛选方法
Fujioka, Yuhei, Misawa, Daitaro, Fukuma, Shingo
Abstract
Representation learning from medical code sequences in electronic health records and medical claims data has been successful in various clinical applications, such as those regarding disease prediction. However, significant challenges remain in extending this approach to the discovery of scientific hypotheses. One reason is that many existing BERT-based models fail to adequately capture the hierarchical structure of medical codes and the complex interactions between diagnoses and treatments. To address these limitations, we propose a new unified pre-training framework that explicitly integrates hierarchical sub-token aggregation, partial masking, and cross-reference mechanisms. The proposed model consistently outperformed existing methods on both pre-training objectives and downstream clinical event prediction tasks, including the onset of dementia and hospitalization. We also conducted an in silico drug repositioning case study targeting Alzheimer's disease. In the hypothesis generation step, our approach successfully rediscovered known promising drugs in a data-driven manner without relying on such external knowledge sources as the literature. Subsequently, in the hypothesis prioritization step, we introduced a Task-Adaptive Representation Approach to alleviate the over-encoding of historical prescription information within diagnostic vectors, enabling the robust prioritization of generated hypotheses. This study establishes an exploratory screening workflow for hypothesis generation and prioritization based on observational associations. Importantly, this framework is not intended to provide causal evidence, but rather to identify promising candidates for subsequent rigorous causal inference. Overall, this study demonstrates that domain-informed representation learning combined with task-adaptive representation control can enable a practical hypothesis discovery workflow.
Chinese Translation
从电子健康档案和医疗理赔数据中的医学编码序列进行表示学习,已在疾病预测等多种临床应用中取得成功。然而,将该方法扩展至科学假设的发现仍面临重大挑战。原因之一在于,许多现有的基于BERT的模型未能充分捕捉医学编码的层次结构以及诊断与治疗之间的复杂交互。为解决这些局限性,我们提出了一种新的统一预训练框架,显式地整合了层次化子词聚合、部分掩码和交叉引用机制。所提出的模型在预训练目标和下游临床事件预测任务(包括痴呆发病和住院预测)上均持续优于现有方法。我们还开展了一项针对阿尔茨海默病的计算机模拟药物重定位案例研究。在假设生成步骤中,我们的方法以数据驱动的方式成功重新发现了已知具有潜力的药物,且不依赖于文献等外部知识来源。随后,在假设优先级排序步骤中,我们引入了一种任务自适应表示方法(Task-Adaptive Representation Approach),以缓解诊断向量中对历史处方信息的过度编码,从而实现对所生成假设的稳健排序。本研究建立了一种基于观察性关联的假设生成与优先级排序的探索性筛选工作流程。需要强调的是,该框架旨在识别有前景的候选对象以供后续严格的因果推断,而非提供因果证据。总体而言,本研究表明,结合任务自适应表示控制的领域知识驱动的表示学习能够支持实用的假设发现工作流程。
cs.LG / 37 / 2609.19873

AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection

AURA:用于电子邮件威胁检测的自适应不确定性路由分析
Berjawi, Omran, fahs, Walid, Khatoun, Rida
Abstract
Email spam and phishing attacks remain a critical security threat. Adversaries increasingly exploit large language models to craft contextually convincing malicious messages, and existing spam detection systems often struggle to keep pace. Generalization across diverse and evolving attack scenarios is limited, which reduces effectiveness once these systems are deployed in practice. This paper introduces Adaptive Uncertainty-Routed Analysis (AURA), a multimodal email threat detection system that analyzes both the content of an email and its embedded URLs. AURA is built around two layers: the first quantifies prediction uncertainty from a URL classifier, and only ambiguous messages are escalated to a fine-tuned transformer encoder for semantic analysis. The system is evaluated on eight heterogeneous training corpora together with two held-out real-world corpora spanning a decade of adversarial campaigns. AURA reaches a macro F1-score of 0.9858 in-distribution, and on NazPhish-Eval and GuenterTrap-Eval it maintains 0.9502 and 0.9436, respectively, which is evidence of robust generalization under genuine distribution shift.
Chinese Translation
垃圾邮件和钓鱼攻击仍然是一个严重的安全威胁。攻击者越来越多地利用大语言模型来生成具有上下文说服力的恶意信息,而现有的垃圾邮件检测系统往往难以跟上这一步伐。这些系统在多样且不断演变的攻击场景中的泛化能力有限,导致其在实际部署后效力下降。本文提出了自适应不确定性路由分析(Adaptive Uncertainty-Routed Analysis,AURA),这是一种多模态电子邮件威胁检测系统,可同时分析邮件内容及其嵌入的URL。AURA由两个层级构成:第一层量化URL分类器的预测不确定性,仅将语义模糊的邮件进一步交由微调的Transformer编码器进行语义分析。该系统在八个异构训练语料库以及两个跨越十年对抗攻击活动的预留真实世界语料库上进行了评估。AURA在分布内达到了0.9858的宏平均F1分数,在NazPhish-Eval和GuenterTrap-Eval上分别保持了0.9502和0.9436的分数,这证明了其在真实分布偏移下具有稳健的泛化能力。
cs.LG / 38 / 2609.19878

Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning

Uni-LaDiR:潜在扩散统一多模态推理
Kang, Haoqiang, Zhang, Yizhe, Kuang, Nikki Lijing, Ma, Yian, Qin, Lianhui
Abstract
Multimodal reasoning requires models to draw on information from multiple modalities throughout the reasoning process. Yet existing methods often concatenate modality-specific thought tokens in a single sequence, leaving the model to bridge representational differences as it reasons across modalities. We introduce Uni-LaDiR (Unified Latent Diffusion Reasoner), a framework that brings these thoughts into a shared latent space for reasoning. A unified encoder maps teacher reasoning steps from different modalities into shared thought tokens, trained to preserve the information needed for later reasoning steps and the final answer or action. Because the same context can support multiple valid next steps, we use diffusion to predict the next block of thought tokens from the input and preceding blocks. Jointly training the encoder and diffusion reasoner with shared model weights encourages thought tokens to be both useful for the task and predictable from the available context. At inference, the model generates these tokens without teacher observations. Across eleven vision-language model (VLM) benchmarks and two vision-language-action (VLA) suites, Uni-LaDiR achieves relative gains over the strongest evaluated baselines of 7.3% on visual reasoning tasks and 6.1% on robot manipulation tasks.
Chinese Translation
多模态推理要求模型在推理过程中充分利用来自多个模态的信息。然而,现有方法通常在单一序列中拼接特定模态的思维标记,使模型在跨模态推理时需要自行弥合表示差异。我们提出了Uni-LaDiR(Unified Latent Diffusion Reasoner,统一潜在扩散推理器),这是一个将各类思维引入共享潜在空间进行推理的框架。统一编码器将来自不同模态的教师推理步骤映射为共享的思维标记,并通过训练使其保留后续推理步骤及最终答案或动作所需的信息。由于同一上下文可以支持多个有效的下一步,我们采用扩散模型根据输入和前序思维块来预测下一块思维标记。通过共享模型权重联合训练编码器和扩散推理器,促使思维标记既对任务有用,又可从可用上下文中预测得到。在推理阶段,模型无需教师观测即可生成这些标记。在十一个视觉语言模型(VLM)基准和两个视觉语言动作(VLA)测试套件上,Uni-LaDiR相较最强评估基线在视觉推理任务上取得7.3%的相对提升,在机器人操作任务上取得6.1%的相对提升。
cs.LG / 39 / 2609.19891

Online Adaptive Kernel Mixing for Gaussian Process Decision Making

面向高斯过程决策的在线自适应核混合方法
Aravindan, Kavin, Sriram, Mani Tej, Dasarathy, Gautam, Bodas, Tejas
Abstract
Gaussian Processes (GPs) are widely used as surrogates for black-box functions in sequential decision-making problems such as Bayesian optimization (BO), level set estimation (LSE), and Bayesian active learning (BAL). GP performance critically depends on kernels, and standard kernels can lead to suboptimal decisions under misspecification. To address this, we introduce HACK GPs (Hedge Adaptive Cumulative Kernels), a method that views kernel selection as an online learning with expert advice problem. HACK treats each candidate kernel as a GP "expert" and updates a distribution over experts online using AdaHedge, based on a loss received as a proxy for their ability to fit the function and align with the task objective. We provide two variants of HACK: (i) Mixture of Gaussians (MoG) and (ii) categorical sampling. We establish general guarantees showing that, under a loss-gap condition, the weight concentrates on the best kernel and the resulting acquisition function is close to that of the best expert. Empirically, we observe robust performance across BO, LSE, and BAL compared to standard kernels such as Squared Exponential and Matern-5/2, as well as simple ensemble baselines.
Chinese Translation
高斯过程(Gaussian Processes, GPs)被广泛用作黑盒函数的代理模型,应用于贝叶斯优化(BO)、水平集估计(LSE)和贝叶斯主动学习(BAL)等序贯决策问题中。高斯过程的性能关键取决于核函数的选择,而标准核函数在核设定错误时可能导致次优决策。为解决这一问题,我们提出了HACK GPs(Hedge Adaptive Cumulative Kernels),该方法将核选择视为一个专家建议下的在线学习问题。HACK将每个候选核视为一个高斯过程“专家”,并基于一个作为其拟合函数能力和任务目标一致性的代理的损失,使用AdaHedge在线更新专家分布。我们提出了HACK的两种变体:(i) 高斯混合和 (ii) 类别采样。我们建立了一般性理论保证,表明在损失差距条件下,权重会集中于最优核,且由此得到的采集函数接近最优专家的采集函数。在实验中,与平方指数(Squared Exponential)和Matern-5/2等标准核以及简单的集成基线相比,HACK在BO、LSE和BAL任务中均表现出稳健的性能。
cs.LG / 40 / 2609.19903

REARL: A Closed-loop Autonomous Driving Simulation Enhancement Framework with Real Traffic Data and Large Language Models

REARL:一种融合真实交通数据与大语言模型的闭环自动驾驶仿真增强框架
Bi, Xiaojun, Jiang, Jun, Sun, Yiwen, Ou, Quanyi, Cheng, Ke, Bi, Mingjie, Li, Yexin
Abstract
Accurate simulation is crucial for autonomous driving development, yet capturing real-world traffic complexity remains challenging. Existing simulators that rely on predefined rules or static data playback struggle with dynamic traffic. CRITICAL uses real traffic data and a large language model (LLM) to adjust the initial simulation configuration, but the simulated distribution still diverges from real traffic as the rollout evolves. We propose REARL, a closed-loop simulation enhancement framework that integrates real traffic data with LLMs. Real traffic data are clustered, and each cluster center is used as a representative scenario that provides typical real-world traffic patterns for the LLM. A timed sliding-window detector then monitors discrepancies in vehicle speed distribution and mean spacing between pairs of vehicles. If a metric exceeds a threshold, the LLM adjusts vehicle decision-making; otherwise the existing controller is kept. The LLM also selects a matching real vehicle from a traffic snapshot and modulates the simulated vehicle with reference to that real action. In a controlled HighD highway setting, compared with the CRITICAL baseline and a PPO-based learning baseline, REARL reduces the Hellinger distance for speed distributions to 0.3067 and the MAPE for mean spacing to 0.8371, while achieving a time headway (THW) of 22.8575 and a lane change rate of 0.0708.
Chinese Translation
精确的仿真对自动驾驶技术的发展至关重要,然而捕捉真实世界的交通复杂性仍然充满挑战。现有依赖预定义规则或静态数据回放的仿真器难以应对动态交通。CRITICAL 利用真实交通数据和大语言模型(LLM)来调整初始仿真配置,但随着仿真过程的推进,其仿真分布仍会偏离真实交通。我们提出了 REARL,一个将真实交通数据与大语言模型相融合的闭环仿真增强框架。真实交通数据经过聚类处理,每个聚类中心被用作代表性场景,为大语言模型提供典型的真实世界交通模式。随后,一个定时滑动窗口检测器监测车辆速度分布以及车辆对之间平均间距的差异。若某项指标超过阈值,则由大语言模型调整车辆决策;否则保留现有控制器。大语言模型还会从交通快照中选择一辆匹配的真实车辆,并参照该真实车辆的动作对仿真车辆进行调控。在受控的 HighD 高速公路场景中,与 CRITICAL 基线和基于 PPO 的学习基线相比,REARL 将速度分布的 Hellinger 距离降低至 0.3067,平均间距的平均绝对百分比误差(MAPE)降低至 0.8371,同时实现了 22.8575 的时间车头时距(THW)和 0.0708 的换道率。
cs.LG / 41 / 2609.19913

Digital Twins for Opinion Dynamics: A Generative LLM Framework for Social Networks

面向舆论动力学的数字孪生:一种面向社交网络的生成式大语言模型框架
Berjawi, Omran, Fenza, Giuseppe, Khatoun, Rida, Zeadally, Sherali
Abstract
The study of opinion dynamics in social networks is one of the key challenges in computational social science with direct relevance to understanding political polarization, misinformation, and health responses. Current approaches focus on simplified mathematical models that ignore linguistic and contextual factors related to belief updates or use Large Language Model (LLM)-based simulations that have not been validated against real data. We present a framework based on the concept of a digital twin to simulate opinion dynamics in social networks. The approach fills the gap by cloning a real-world Twitter network, assigns a set of attributes for agents (such as persona, emotions, centrality, stubbornness, and influence), and employs Mistral-7B to perform opinion update based on memory and social exposure. To evaluate the proposed approach, we validate it against two real Twitter datasets (COVID-19 discourse and U.S elections 2020). The results show that the capability of the proposed framework reproduces opinion trajectories and reduces individual prediction error by more than 50% compared to the best-performing classical baseline (Mistral-7B achieves Mean Absolute Error (MAE) = 0.150 and 0.121 on the COVID-19 and US Election 2020 datasets, respectively). We observe similar improvements in structural alignment (Delta_r = 0.120 and 0.180) and polarization dynamics (Delta_Var = 0.106 and 0.115) on the two datasets, respectively. Additionally, the ablation studies confirm that agent attributes, memory, and social exposure all contribute to the framework's predictive fidelity in reproducing opinion trajectories, with agent attributes being the most critical contributor. Overall, our results demonstrate that grounding Mistral-7B within empirically cloned interaction networks produces a realistic simulation framework capable of reproducing complex social dynamics.
Chinese Translation
社交网络中舆论动力学的研究是计算社会科学的关键挑战之一,与理解政治极化、错误信息和健康响应直接相关。现有方法要么聚焦于忽略与信念更新相关的语言和情境因素的简化数学模型,要么采用尚未经过真实数据验证的基于大语言模型(LLM)的仿真。我们提出一个基于数字孪生概念的框架,用于模拟社交网络中的舆论动力学。该方法通过克隆真实的Twitter网络来填补这一空白,为智能体分配一组属性(如人设、情绪、中心性、顽固性和影响力),并利用Mistral-7B基于记忆和社会暴露执行观点更新。为评估所提方法,我们在两个真实的Twitter数据集(COVID-19讨论和美国2020年大选)上进行了验证。结果表明,与表现最好的经典基线相比,该框架能够再现舆论轨迹并将个体预测误差降低50%以上(Mistral-7B在COVID-19和美国2020年大选数据集上的平均绝对误差(MAE)分别为0.150和0.121)。我们还观察到,在两个数据集上结构对齐(Delta_r分别为0.120和0.180)和极化动力学(Delta_Var分别为0.106和0.115)也有类似的改进。此外,消融实验证实,智能体属性、记忆和社会暴露均有助于框架在再现舆论轨迹方面的预测保真度,其中智能体属性是最关键的贡献因素。总体而言,我们的结果表明,将Mistral-7B嵌入到实证克隆的交互网络中,能够产生一个可再现复杂社会动力学的现实仿真框架。
cs.LG / 42 / 2609.19915

Amortizing Physics-Informed Neural Solvers via Graph Hypernetworks

基于图超网络的摊销式物理信息神经求解器
Jing, Cheng, Verma, Abhishek, Bera, Kallol, He, Yixuan, Lee, Kookjin
Abstract
Amortizing physics-informed neural networks (PINNs) across related PDEs requires describing each equation to a reusable solver. Coefficient vectors encode numerical parameters in predefined slots, leaving operator and cross-field assignments implicit. We make these relationships explicit in an operator graph, with nodes for fields, derivatives, terms, and residuals and coefficients retained as term attributes. A graph hypernetwork generates diagonal codes that initialize a meta-trained factorized PINN for each target equation. Meta-training and target-specific adaptation use governing equations and prescribed conditions without solution labels. We compare coefficient-vector, DeepSets-based term-set, and graph conditioning by solution accuracy within a fixed adaptation budget. In scalar convection-diffusion-reaction problems, both term-based descriptors improve high-reaction accuracy, with similar performance. In two-field Fisher-KPP, meta-training sees uncoupled and one-way systems; after 3,000 adaptation steps on unseen two-way coupling, the graph's mean final error is 35.7% below the term set and 67.7% below the coefficient vector. In a fixed-structure capacitively coupled plasma model, the coefficient vector performs best. These results support extending coefficient conditioning with explicit equation relationships for physics-based solver adaptation.
Chinese Translation
在相关偏微分方程(PDE)之间摊销物理信息神经网络(PINN)需要向一个可复用的求解器描述每个方程。系数向量在预定义的槽位中编码数值参数,而算子关系和跨场分配则保持隐式。我们在一个算子图中将这些关系显式化,图中包含表示场、导数、项和残差的节点,而系数则作为项的属性保留。图超网络生成对角码,用于为每个目标方程初始化一个经过元训练的分解式PINN。元训练和目标特定的自适应均仅使用控制方程和给定条件,无需解的标签。我们在固定的自适应预算下,通过求解精度比较了系数向量、基于DeepSets的项集合和图条件化三种方式。在标量对流-扩散-反应问题中,两种基于项的描述方式都提高了高反应率下的精度,且性能相近。在双场Fisher-KPP问题中,元训练仅见过非耦合和单向耦合系统;在未见过的双向耦合上进行3,000步自适应后,图条件化的平均最终误差比项集合低35.7%,比系数向量低67.7%。在一个固定结构的电容耦合等离子体模型中,系数向量表现最佳。这些结果支持在系数条件化的基础上扩展显式的方程关系,以实现基于物理的求解器自适应。
cs.LG / 43 / 2609.19955

One Intervention per Component is Enough: Towards Identifiability in Linear Stochastic Dynamics from Steady State

每个连通分量一次干预即可:从稳态实现线性随机动力学的可辨识性
Salehkaleybar, Saber
Abstract
We study the problem of recovering the parameters of a multivariate Ornstein-Uhlenbeck (OU) process from steady-state observational and interventional data. In many applications, such as large-scale gene perturbation experiments, only stationary "snapshot" measurements are available, making standard stochastic differential equation estimation methods that rely on time-series trajectories inapplicable. We first establish an identifiability result: one intervention per strongly connected component (SCC) of the drift graph suffices to recover all OU process parameters generically up to a global scaling factor. This holds provided that the SCC condensation graph is connected with a single root and certain spectral nondegeneracy assumptions hold. We propose a recursive learning algorithm that orders SCCs topologically and, for each component, isolates its marginal dynamics and solves a linear system derived from the steady-state moment equations, leveraging parameters recovered for upstream components. Building on this theoretical foundation, we propose a regularized least-squares estimator that jointly minimizes residuals of the steady-state mean and covariance equations across observational and interventional data. Experimental results validate our theoretical findings in recovering parameters of the underlying OU process.
Chinese Translation
我们研究从稳态的观测数据与干预数据中恢复多元Ornstein-Uhlenbeck(OU)过程参数的问题。在许多应用中,例如大规模基因扰动实验,只能获得平稳的“快照”式测量数据,使得依赖时间序列轨迹的标准随机微分方程估计方法不再适用。我们首先建立了一个可辨识性结果:在一般情形下,对漂移图(drift graph)的每个强连通分量(SCC)施加一次干预,即可在全局缩放因子的意义下恢复OU过程的全部参数。该结果成立的前提是SCC的凝聚图连通且具有单一根节点,并满足一定的谱非退化假设。我们提出一种递归学习算法:按拓扑顺序排列各SCC,对每个分量分离其边际动力学,并利用已恢复的上游分量参数,求解由稳态矩方程导出的线性方程组。在此理论基础之上,我们进一步提出一种正则化最小二乘估计器,其在观测数据与干预数据上联合最小化稳态均值方程和协方差方程的残差。实验结果验证了我们在恢复底层OU过程参数方面的理论发现。
cs.LG / 44 / 2609.19956

Graph-Based Stochastic Power-UCT: Monte-Carlo Graph Search with Power Mean Estimation

基于图的随机Power-UCT:带幂均值估计的蒙特卡洛图搜索
Tran, Tung, Mai, Viet Bao, Ta, Hoang, Dam, Tuan
Abstract
Tree-based Monte-Carlo Tree Search (MCTS) duplicates the same state when it is reached through different trajectories, which can waste simulations in stochastic MDPs. We introduce Graph-Based Stochastic-Power-UCT (GS-Power-UCT), which shares states reached at the same planning depth while keeping separate values for states reached at different depths. This design applies to general stochastic MDPs, including problems with cycles. We prove that for a fixed planning horizon, the root estimate converges to the finite-horizon value at rate $O(n^{-1/2})$, matching tree-based Stochastic-Power-UCT while reusing samples across shared states. We also study two full-state variants: GS-Power-UCT-F, which stores one node per physical state to increase sample sharing but may mix values from different remaining horizons, and GS-Power-UCT-F$^+$, which uses an adaptive horizon to control this bias. The latter converges to $V^{\star}(s_0)$, the optimal infinite-horizon discounted value at the root state $s_0$, when the remaining cross-depth gap vanishes. Experiments on stochastic planning benchmarks show improved sample efficiency over tree-based and graph-based baselines.
Chinese Translation
基于树的蒙特卡洛树搜索(MCTS)在通过不同轨迹到达同一状态时会重复该状态,这在随机马尔可夫决策过程(MDP)中可能浪费模拟次数。我们提出了基于图的随机Power-UCT(Graph-Based Stochastic-Power-UCT,简称GS-Power-UCT),该方法在相同规划深度下共享所到达的状态,同时对不同深度到达的状态保持各自独立的值函数。这一设计适用于一般的随机MDP,包括含循环的问题。我们证明,对于固定的规划时域,根节点估计值以 $O(n^{-1/2})$ 的速率收敛到有限时域值,与基于树的Stochastic-Power-UCT相当,同时通过共享状态复用了样本。我们还研究了两个全状态变体:GS-Power-UCT-F,它为每个物理状态存储一个节点以增加样本共享,但可能混合不同剩余时域的值;以及GS-Power-UCT-F$^+$,它采用自适应时域来控制这种偏差。当剩余的跨深度间隙消失时,后者收敛到 $V^{\star}(s_0)$,即根状态 $s_0$ 处的最优无限时域折扣值。在随机规划基准上的实验表明,相较于基于树和基于图的基线方法,该方法具有更优的样本效率。
cs.LG / 45 / 2609.19970

CellRFT: Reinforcement Fine-Tuning for Single-Cell Perturbation Modeling

CellRFT:面向单细胞扰动建模的强化微调方法
Yan, Jie, Liu, Li, Guo, Hanze, Hu, Jiaxin, He, Houxin, Qi, Xiaoning, Wang, Haoran, Li, Cong, Zhang, Zhong-Yuan, Wang, Yong
Abstract
Predicting cellular responses to perturbations supports the study of gene function, disease mechanisms, and therapeutic strategies. Despite advances in single-cell perturbation modeling, existing models typically optimize surrogate losses that do not directly reflect the biological criteria used for evaluation, so better data fitting need not yield better biological predictions. To address this mismatch, we introduce \textbf{CellRFT}, a reinforcement fine-tuning framework that uses biological evaluation as direct training feedback. CellRFT uses policy-gradient optimization to learn from non-differentiable evaluations of generated cell populations and integrates multiple biological rewards through hierarchical reward aggregation. Comprehensive experiments demonstrate CellRFT's applicability across different pretrained models and effectiveness in improving perturbation prediction, reveal that optimizing one biological criterion can help or hinder others, and show that complementary rewards can improve criteria beyond those directly optimized, offering a way to probe how biological metrics shape model behavior, with the potential to inform evaluation design. Code will be made available.
Chinese Translation
预测细胞对扰动的响应有助于研究基因功能、疾病机制和治疗策略。尽管单细胞扰动建模已取得进展,但现有模型通常优化的是不能直接反映评估所用生物学标准的替代损失,因此更好的数据拟合未必带来更好的生物学预测。为解决这一不匹配问题,我们提出了 CellRFT,一个以生物学评估作为直接训练反馈的强化微调框架。CellRFT 采用策略梯度优化,从生成的细胞群落的不可微评估中学习,并通过层次化奖励聚合整合多种生物学奖励。全面的实验表明,CellRFT 可适用于不同的预训练模型,并能有效提升扰动预测性能;实验还揭示了对某一生物学准则的优化可能促进或阻碍其他准则,且互补的奖励可以改善直接优化目标之外的准则。这为探究生物学指标如何塑造模型行为提供了途径,并有望为评估设计提供参考。代码将会开源。
cs.LG / 46 / 2609.19985

Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification

过去、未来、一步兼顾:基于事后JANUS校正缓解稳定性-可塑性困境
Zheng, Zhilong, Tao, Letian, Guan, Yang, Yang, Yujie, Xiong, Wei, Sheng, Kehua, Zhang, Bo, Duan, Jingliang, Li, Keqiang, Li, Shengbo Eben
Abstract
Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While existing works attempt to mitigate this on the basis of parameter-efficient fine-tuning methods, they adopted an overly restrictive Subspace Orthogonality condition. In this paper, we introduce a purely post-hoc and tuning-agnostic weight rectification framework that achieves Parameter Space Orthogonality, which is the necessary and sufficient condition for preserving historical performance to the first order. By projecting parameter updates into the JAcobian NUll Space (JANUS), our method significantly recovers compromised historical knowledge without interfering with the underlying fine-tuning process. To overcome the local validity of the Jacobian approximation, we further propose a Multi-step Adaptive Rectification mechanism that utilizes the JANUS shift to dynamically verify the valid trust region and adjust step sizes. Coupled with our proposed ghost projection, ghost orientation comparison, and sequence-level singular value decomposition compression techniques, JANUS also achieves great temporal and spatial efficiency. Experiments demonstrate that JANUS seamlessly integrates with various fine-tuning methods, significantly mitigating the stability-plasticity dilemma by recovering historical knowledge while preserving downstream task adaptation.
Chinese Translation
在新任务上对基础模型进行微调不可避免地会遭受灾难性遗忘。尽管现有工作试图在参数高效微调方法的基础上缓解这一问题,但它们采用了过于严格的子空间正交性条件。在本文中,我们提出了一个纯事后且与微调方式无关的权重校正框架,以实现参数空间正交性,这是一阶意义上保持历史性能的充要条件。通过将参数更新投影到雅可比零空间,我们的方法在不干扰底层微调过程的前提下,显著恢复了受损的历史知识。为克服雅可比近似的局部有效性限制,我们进一步提出了一种多步自适应校正机制,利用JANUS偏移动态验证有效信赖域并调整步长。结合我们提出的幽灵投影、幽灵方向比较以及序列级奇异值分解压缩技术,JANUS还实现了极高的时间和空间效率。实验表明,JANUS能够与各种微调方法无缝集成,通过恢复历史知识同时保留下游任务适应能力,显著缓解了稳定性-可塑性困境。
cs.LG / 47 / 2609.20004

EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning

EPIG-Tree:面向梯度高效强化学习的计算最优分支策略
Khomich, Nikita, Hermansson, Leopold, Hakimi, Ido
Abstract
Reward-based reinforcement learning for language models, exemplified by Group Relative Policy Optimization (GRPO), collapses an entire stochastic trajectory into a single scalar reward. This is clean and scalable, but it explores and allocates reward inefficiently: a trajectory may contain many causal decisions, recovery attempts, and environment-randomness events, yet every token or action inherits one trajectory-level advantage. We study tree-based rollout construction as a compute-allocation problem for policy-gradient estimation. Our central claim is that branches should be placed not where the policy is merely uncertain, but where an additional branch most reduces uncertainty about the policy gradient per unit of compute. From a law-of-total-variance decomposition of the local policy-gradient random variable, we derive two allocation laws: new branches reduce decision uncertainty, while repeated suffix rollouts reduce continuation uncertainty. The resulting EPIG-Tree score allocates branches using the already computed rollouts. It estimates occupancy- and score-weighted value uncertainty, along with a suffix law $n_e \propto w_e \|\nabla_\theta \log \pi(a_e|h_e)\| \sigma_e / \sqrt{c_e}$. Empirically, EPIG reduces gradient MSE in cloned-state control, winning in all nine dense continuous-control environments of a 13-environment sweep and recovering the reference gradient direction near-perfectly, and it improves frozen-LLM gradient calibration relative to entropy branching. In online single-turn math, tree-local credit beats flat GRPO, while branch placement is secondary to token-level credit assignment. In online multi-turn Wordle, EPIG attains the highest final win rate (0.850), overtaking flat GRPO, which saturates early at 0.790, and entropy branching as training proceeds, confirming that the gradient-estimation advantage transfers to a stateful, large-action setting.
Chinese Translation
以群体相对策略优化(GRPO)为代表的基于奖励的语言模型强化学习,将整条随机轨迹压缩为单一标量奖励。这种方式简洁且易于扩展,但其在探索和奖励分配上效率低下:一条轨迹可能包含众多因果决策、恢复尝试和环境随机性事件,但每个词元(token)或动作都继承同一个轨迹级别的优势值。我们将基于树的 rollout 构建视为策略梯度估计中的计算分配问题。我们的核心主张是:分支不应仅仅放在策略不确定的地方,而应放在每单位计算能最大程度降低策略梯度不确定性的地方。基于对局部策略梯度随机变量的全方差定律分解,我们推导出两条分配定律:新分支用于降低决策不确定性,而重复的后缀 rollout 用于降低延续不确定性。由此得到的 EPIG-Tree 分数利用已计算完成的 rollout 来分配分支。它估计基于占据率和分数加权的价值不确定性,并遵循后缀定律 $n_e \propto w_e \|\nabla_\theta \log \pi(a_e|h_e)\| \sigma_e / \sqrt{c_e}$。实验表明,在克隆状态控制任务中,EPIG 降低了梯度均方误差(MSE),在包含13个环境的评测中于全部九个稠密连续控制环境中胜出,并近乎完美地恢复了参考梯度方向;相对于熵分支策略,它还改善了冻结大语言模型(LLM)的梯度校准。在在线单轮数学任务中,树局部信用分配优于扁平的 GRPO,而分支位置的重要性则次于词元级信用分配。在在线多轮 Wordle 任务中,EPIG 取得了最高的最终胜率(0.850),随着训练进行超越了早期收敛于 0.790 的扁平 GRPO 以及熵分支策略,证实了这种梯度估计优势能够迁移到有状态、大动作空间的场景中。
cs.LG / 48 / 2609.20008

Dynamic Generalized Gromov-Wasserstein Optimal Transport

动态广义Gromov-Wasserstein最优传输
Ying, Junda, Zeng, Zhiwei, Zhou, Peijie, Zhang, Lei
Abstract
Gromov--Wasserstein optimal transport (GW-OT) extends classical optimal transport by introducing structure-aware transport cost. This is particularly relevant for spatial transcriptomics, where dynamical reconstruction should preserve tissue structure in addition to matching expression patterns. While static formulations have been widely used for such structure-aware alignment, a general dynamic formulation for reconstructing continuous trajectories is still missing. We introduce Travelling Pair Dynamical Alignment and Trajectory Estimation (TP-DATE), a theoretical and computational framework to generalize GW-OT dynamically in a simulation-free manner. We formulate a broad class of static and dynamic Quadratic-form OT (QOT) through path actions and prove the static dynamic equivalence. We further develop travelling-pair flow matching, which allows interacting conditional paths and marginalizes their interactions into a single vector field. On synthetic and real spatial transcriptomics data, TP-DATE better preserves spatial structure and improves continuous 3D dynamics reconstruction.
Chinese Translation
Gromov-Wasserstein最优传输(GW-OT)通过引入结构感知的传输代价扩展了经典最优传输。这对于空间转录组学尤为重要,因为动态重建除了匹配表达模式外,还应保留组织结构。尽管静态形式已被广泛用于此类结构感知的对齐,但目前仍缺乏用于重建连续轨迹的通用动态形式。我们提出了 Travelling Pair Dynamical Alignment and Trajectory Estimation(TP-DATE),一个以无需仿真的方式动态推广GW-OT的理论与计算框架。我们通过路径作用量构造了一类广泛的静态与动态二次形式最优传输(Quadratic-form OT, QOT),并证明了静态与动态的等价性。我们进一步发展了 travelling-pair flow matching,它允许条件路径之间相互作用,并将其相互作用边缘化为单一向量场。在合成数据和真实空间转录组学数据上,TP-DATE 更好地保留了空间结构,并提升了连续三维动力学重建的效果。
cs.LG / 49 / 2609.20045

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

现在正确,日后不足:上下文压缩中的更新充分性审计
Zhang, Guangzhe
Abstract
A memory can answer a current query correctly while discarding distinctions required by a later update. We investigate this failure with a paired-history audit: two histories have the same current answer, receive a shared future update, and require different subsequent answers. A pilot evaluates 24 history pairs across six synthetic mechanisms, 12 memory conditions, two repeats, and two model backends. A deterministic frontier selector obtains strict reveal accuracy of 96/96 on DeepSeek and 82/96 on GLM; a structured writer obtains 62 successes with one unresolved outcome and 56/96. The configured four-outcome joint contrast has finite-sample identification intervals of [0.521, 0.542] and [0.292, 0.313], not confidence intervals. A record-level audit distinguishes retained-state adequacy, response delivery, and answer-schema compliance without changing those original scores. It finds 26 and 25 well-formed but semantically wrong structured reveal memories, while all 14 GLM frontier reveal failures contain correct values in the wrong wrapper. Tombstone removal produces 16/16 exact replay failures in the targeted mechanism. Identifier renaming then exposes a separate flaw: original frontier late-reference adequacy falls from 8/8 to 94/320 transformed instances. We provide and test a label-equivariant repair, but it preserves only 2/8 original late-reference answers: eliminating a naming shortcut does not solve unknown future relevance. These results support a scoped evaluation methodology and reproducible failure analysis, not general superiority of the repaired algorithm. Paid pilot evidence, retrospective diagnostics, and new offline tests are reported separately; no independent held-out or natural-task validation is claimed.
Chinese Translation
一段记忆能够正确回答当前查询,同时却丢弃了后续更新所需要的区分信息。我们通过一种配对历史审计方法来研究这种失效:两条历史具有相同的当前答案,接收到相同的未来更新后,却需要给出不同的后续答案。一项先导实验评估了24对历史,涵盖六种合成机制、12种记忆条件、两次重复以及两个模型后端。一个确定性的前沿选择器在DeepSeek上获得96/96的严格揭示准确率,在GLM上获得82/96;一个结构化写入器获得62次成功、一次未决结果以及56/96的成绩。所配置的四结果联合对比的有限样本识别区间分别为[0.521, 0.542]和[0.292, 0.313],而非置信区间。记录级审计在不改变上述原始分数的情况下,区分了保留状态充分性、响应传递和答案模式合规性。审计发现分别有26个和25个格式良好但语义错误的结构化揭示记忆,而GLM前沿选择器的全部14次揭示失败均为数值正确但封装错误。墓碑(tombstone)删除机制在目标场景中产生了16/16的精确重放失败。随后的标识符重命名暴露了另一个缺陷:原始前沿选择器的后期引用充分性从8/8下降至变换后320个实例中的94个。我们提供并测试了一种标签等变修复方法,但它仅保留了2/8的原始后期引用答案:消除命名捷径并不能解决未知的未来相关性问题。这些结果支持一种范围受限的评估方法和可复现的失效分析,而非修复算法的普遍优越性。付费先导实验证据、回溯性诊断和新的离线测试分别报告;本文不声称任何独立的留存验证或自然任务验证。
cs.LG / 50 / 2609.20058

Evaluating Explanation Methods by the Predictors They Induce

通过解释方法所诱导的预测器来评估解释方法
Selbæk, Jacob, Hammer, Hugo L.
Abstract
Explanations of machine learning models are usually judged by criteria that are hard to compare. We propose a simpler test: if an explanation really describes how a model uses its features, it should be possible to rebuild the model's predictions from it. We turn each explanation into a predictor by reading each feature's effect and adding them up, and measure how well that predictor reproduces the model on unseen data. Nothing is fitted, so the score reflects the explanation itself. The test applies to any explanation that can be written as a function of the features; we demonstrate it on partial dependence plots (PDP), accumulated local effects (ALE), SHAP and LIME. We prove that summing partial dependence curves gives the best possible additive summary of a model when its features are independent, and that this fails when they are dependent. Across 13 real datasets and 9 synthetic designs and four model families, which method scores best depends entirely on feature dependence: where features are independent SHAP is slightly worse than PDP, exactly as the theory predicts; on dependent real data SHAP leads. Some widely used quality metrics even prefer a damaged explanation to an intact one.
Chinese Translation
机器学习模型的解释通常依据难以相互比较的标准来评判。我们提出一种更简单的检验方法:如果某种解释真正描述了模型如何使用其特征,那么应当能够依据该解释重建模型的预测。我们将每种解释转化为一个预测器——读取每个特征的效应并将其相加——并衡量该预测器在未见数据上重现模型预测的能力。由于这一过程不涉及任何拟合,因此得分纯粹反映解释本身。该检验适用于任何可以写成特征函数的解释;我们在部分依赖图(PDP)、累积局部效应(ALE)、SHAP 和 LIME 上进行了演示。我们证明,当特征相互独立时,对部分依赖曲线求和可获得模型的最优可加性摘要;而当特征存在依赖关系时,这一方法会失效。在13个真实数据集、9个合成设计以及四类模型族上的实验表明,哪种方法得分最高完全取决于特征依赖性:在特征独立的情形下,SHAP 略逊于 PDP,与理论预测完全一致;而在存在特征依赖的真实数据上,SHAP 表现领先。一些被广泛使用的质量指标甚至会更偏好受损的解释而非完好的解释。
cs.LG / 51 / 2609.20082

MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards

MATCH:基于课程调度与分层门控奖励的模型感知工具学习
Liu, Shihao, Yin, Hao, Liu, Lijun, Chen, Zhengzong, Zhao, Yuanyuan, Huang, Fei
Abstract
Tool learning enables large language models (LLMs) to use external tools for tasks beyond parametric knowledge. Reinforcement learning can optimize tool-call behavior from feedback, but current methods still face two problems: fixed-threshold curricula can become misaligned with the policy's evolving capability boundary, and additive rewards can leak argument-level credit when the predicted tool is wrong. To address these problems, we propose MATCH, a closed-loop framework for model-aware tool learning with curriculum scheduling and hierarchically gated rewards. Model-Aware Curriculum Learning (MACL) maintains reward-derived sample difficulty that co-evolves with the policy, and each epoch selects samples near the current capability boundary together with a top-k pool of harder cases. Hierarchical Tool-call Gated Reward (HTGR) scores tool name, argument key, and argument value as a gated chain, granting credit at each level only when prerequisites hold. The same HTGR rewards drive both GRPO updates and MACL's difficulty refresh, closing the loop between policy optimization and sample scheduling. On API-Bank and BFCL V3, MATCH reaches 72.19% and 62.87% overall accuracy, outperforming the main supervised and RL-based baselines. Backbone experiments further show consistent improvements across four backbones from two model families.
Chinese Translation
工具学习使大语言模型(LLMs)能够使用外部工具完成超出参数化知识范围的任务。强化学习可以基于反馈优化工具调用行为,但现有方法仍面临两个问题:固定阈值的课程设置可能与策略不断演进的能力边界脱节;加性奖励在预测工具错误时仍可能泄漏参数级别的信用分配。为解决这些问题,我们提出了MATCH,一个结合课程调度与分层门控奖励的模型感知工具学习闭环框架。模型感知课程学习(MACL)维护由奖励推导出的样本难度,使其与策略共同演化,并在每个epoch选择当前能力边界附近的样本,同时辅以一个难度更高的top-k样本池。分层工具调用门控奖励(HTGR)将工具名称、参数键和参数值作为门控链进行评分,仅在前置条件满足时才在相应层级给予信用。相同的HTGR奖励同时驱动GRPO更新和MACL的难度刷新,从而在策略优化与样本调度之间形成闭环。在API-Bank和BFCL V3基准上,MATCH分别达到72.19%和62.87%的总体准确率,优于主要的有监督和基于强化学习的基线方法。骨干模型实验进一步表明,MATCH在来自两个模型家族的四个骨干模型上均取得了一致的提升。
cs.LG / 52 / 2609.20086

SETTer: Sparse-Encoder Transformer for Long-term Multivariate Time Series Forecasting

SETTer:用于长期多变量时间序列预测的稀疏编码器Transformer
Ezema, Abraham, Eze, Chijioke, Ponci, Ferdinanda, Monti, Antonello
Abstract
Long-term multivariate time series plays a significant role in many application areas such as power systems, trading, etc. However, their accurate prediction is quite difficult for conventional forecasting methods as they often exhibit high dimensionality and complex relationships. Recent works show that transformer-based approaches are quite effective for long-term forecasting thanks to their attention mechanism. However, in the presence of complex high-dimensional inputs, they show evidence of oversmoothing, limited capacity, and opacity. To this end, this paper introduces SETTer, a transformer-based model that addresses these challenges by incorporating novel techniques for decoupled self-attention and hybrid masking. The proposed techniques enable SETTer to effectively capture the dominant short- and long-term patterns across the temporal and channel dimensions. In addition, we enrich the model layers with simple explainable structures that indicate the discriminative pattern of SETTer. We show that with a single-layer transformer architecture, SETTer can effectively model long-term dependencies in the presence of varying data complexities. Extensive experiments on real-word benchmark datasets for long-term multivariate time series forecasting demonstrate that SETTer outperforms state-of-the-art models in 88% of the scenarios.
Chinese Translation
长期多变量时间序列在电力系统、交易等许多应用领域中都发挥着重要作用。然而,由于其通常具有高维度和复杂的关联关系,传统的预测方法难以对其进行准确预测。近期研究表明,基于Transformer的方法凭借其注意力机制在长期预测中非常有效。然而,在面对复杂的高维输入时,这些方法会出现过度平滑、容量受限和不透明等问题。为此,本文提出了SETTer,这是一种基于Transformer的模型,通过引入解耦自注意力(decoupled self-attention)和混合掩码(hybrid masking)等新技术来应对这些挑战。所提出的技术使SETTer能够有效捕捉时间维度和通道维度上的主要短期与长期模式。此外,我们在模型层中加入了简单的可解释结构,以揭示SETTer的判别模式。研究表明,仅需单层Transformer架构,SETTer便能在不同数据复杂度下有效建模长期依赖关系。在真实世界长期多变量时间序列预测基准数据集上的大量实验表明,SETTer在88%的场景中优于当前最先进的模型。
cs.LG / 53 / 2609.20098

CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning

CARE-VI:面向离策略Actor-Critic学习中价值改进的保守自适应可靠性估计
Zou, Xiang, Shi, Shengzhu, Gao, Junqi, Guo, Zhichang
Abstract
Reliable temporal-difference targets are central to off-policy actor-critic learning. Direct value improvement refines the next-state target with alternative actions, but the reliability of this refinement depends on how candidate actions are ranked, reviewed, and weighted. Noisy rankings may force premature candidate commitment, reusing selection scores may bias target valuation, and fixed enhancement weights may amplify weak evidence. To address these risks, we develop Conservative Adaptive Ranking and Screening (CARS), which retains an ordered candidate prefix within a preset budget and narrows it only when the observed boundary gap exceeds a disagreement-scaled uncertainty radius. Selector-Evaluator Value Assessment (SEVA) uses selector critics to order candidates and a separately parameterized evaluator critic to review the selected value, then caps the reviewed value at the selector reference. Dynamic Adaptive Risk-aware Enhancement (DARE) then regulates each residual correction using candidate reliability, the gap between selector and evaluator signals, and a finite stage factor. Together, CARS, SEVA, and DARE form CARE-VI, an evidence-regulated target construction framework that preserves the backbone interfaces for critic regression and actor updates. The analysis bounds the CARS boundary error, the SEVA selected-value overestimation, and the one-sided deviation of the DARE residual displacement from its population counterpart, and establishes fixed-policy recovery after the finite-stage perturbation ends. Experiments with SAC, TD3, and TD7 on four MuJoCo tasks show that CARE-VI achieves the highest mean return in all twelve settings. Grouped ablations and scalar diagnostics support the roles of the three components in improving target reliability.
Chinese Translation
可靠的时序差分目标是离策略actor-critic学习的核心。直接价值改进利用备选动作对下一状态目标进行精炼,但这种精炼的可靠性取决于候选动作的排序、审查和加权方式。噪声化的排序可能导致过早的候选动作确定,复用选择评分可能使目标价值估计产生偏差,而固定的增强权重可能放大弱证据。为应对这些风险,我们提出了保守自适应排序与筛选(Conservative Adaptive Ranking and Screening, CARS),它在预设预算内保留有序的候选前缀,仅当观测到的边界差距超过由分歧缩放的不确定性半径时才缩小该前缀。选择器-评估器价值评估(Selector-Evaluator Value Assessment, SEVA)使用选择器评论家对候选动作排序,并使用单独参数化的评估器评论家审查所选价值,随后将审查后的价值截断在选择器参考值上。动态自适应风险感知增强(Dynamic Adaptive Risk-aware Enhancement, DARE)则依据候选动作可靠性、选择器与评估器信号之间的差距以及有限阶段因子来调节每个残差修正量。CARS、SEVA和DARE共同构成CARE-VI,一个由证据调节的目标构造框架,并保留用于评论家回归和行动者更新的主干接口。本文的分析对CARS边界误差、SEVA所选价值的高估以及DARE残差位移相对于其总体对应量单侧偏差给出了界定,并证明了有限阶段扰动结束后固定策略的可恢复性。在四个MuJoCo任务上使用SAC、TD3和TD7进行的实验表明,CARE-VI在全部十二种设置中均取得了最高的平均回报。分组消融实验和标量诊断结果支持了三个组件在提升目标可靠性方面的作用。
cs.LG / 54 / 2609.20123

QoS-Aware Federated Learning for Multimodal In-Cabin Interaction in Smart Vehicles

面向智能汽车多模态座舱交互的QoS感知联邦学习
Gül, Baran Can, Nakıp, Mert, Jazdi, Nasser, Weyrich, Michael
Abstract
Modern smart vehicles leverage multimodal sensors, ranging from high-bandwidth vision systems to low-rate physiological monitors, to provide personalized in-cabin services. However, integrating high-fidelity multimodal fusion with collaborative training is often hindered by the heterogeneous and time-varying Quality of Service (QoS) constraints of vehicular networks. Standard Federated Learning (FL) approaches enforce rigid synchronous rounds that fail to account for these resource asymmetries, leading to safety-critical timing violations and energy exhaustion. In this paper, we propose FedQoS, a novel asynchronous, event-triggered FL framework that decouples local computation from global communication via a two-phase gating mechanism. First, we introduce a resource-aware training gate that initializes local learning only when sensing buffers and energy reserves meet safety thresholds, preventing ML tasks from compromising core vehicle mobility. Second, a QoS-aware transmission policy gates uplink updates based on an efficiency score that balances model novelty against instantaneous latency and energy costs. Locally, clients optimize an objective featuring a staleness-aware proximal term that dynamically adjusts the global anchor strength based on update age. Extensive experiments on multimodal vehicular datasets demonstrate that FedQoS achieves competitive personalized accuracy with only marginal performance loss compared to FedAvg, while substantially reducing QoS violations, cutting communication overhead by 76.7\%, and lowering latency cost by 26.0\%, demonstrating a highly favorable accuracy and efficiency balance for real-world vehicular deployments.
Chinese Translation
现代智能汽车利用多模态传感器(从高带宽视觉系统到低速率生理监测设备)来提供个性化座舱服务。然而,将高保真多模态融合与协同训练相结合,往往受到车辆网络异构且时变的服务质量(QoS)约束的阻碍。标准的联邦学习(FL)方法强制执行僵化的同步轮次,无法考虑这些资源不对称性,从而导致安全关键的时序违规和能量耗尽。本文提出了一种新颖的异步、事件触发的联邦学习框架FedQoS,通过两阶段门控机制将本地计算与全局通信解耦。首先,我们引入资源感知训练门控,仅在感知缓冲区和能量储备满足安全阈值时才初始化本地学习,防止机器学习任务损害核心的车辆移动性。其次,QoS感知传输策略基于一个效率评分对上行更新进行门控,该评分权衡模型新颖性与瞬时延迟和能量成本。在本地,客户端优化一个包含陈旧性感知近端项的目标函数,该近端项根据更新时间动态调整全局锚点强度。在多模态车辆数据集上的大量实验表明,FedQoS在仅相比FedAvg有微小性能损失的情况下取得了具有竞争力的个性化精度,同时显著减少了QoS违规,将通信开销降低了76.7%,并将延迟成本降低了26.0%,展现出面向真实车辆部署的高度有利的精度与效率平衡。
cs.LG / 55 / 2609.20129

Local Sparsity Enables Unsupervised LLM Safety Detection

局部稀疏性使无监督的大语言模型安全检测成为可能
Chen, Xin, Kur, Gil, Shevchenko, Alexander, Krause, Andreas
Abstract
Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising concerns about whether anomaly detection is statistically feasible. We show that, under the linear representation hypothesis (LRH), there may indeed be hope. In the LRH concept space, which is typically recovered via a sparse autoencoder (SAE), nearby points share a small common active support. Using this local sparsity insight, we propose a framework for locally masked SAE-based anomaly detection, supported by theoretical justifications. We validate it on various architectures and datasets, including both capability-testing datasets and safety-specific datasets. Finally, when we allow algorithms to use 1% out-of-distribution data for calibration, locally sparse methods achieve near-optimal performance, demonstrating their ability to capture meaningful safety information while using only 1-2% of SAE neurons for computation.
Chinese Translation
大语言模型(LLM)部署时的安全方法主要是有监督的,并假定可以访问不安全训练数据。然而,新的攻击和伤害类别不断出现,无法被以这种有监督方式训练的模型所捕捉。另一种方法是从异常检测的角度看待这一问题,即仅依靠对安全数据建模并标记分布外输入。然而,LLM 的激活值位于高维空间中,这引发了对异常检测在统计上是否可行的担忧。我们证明,在线性表示假设(LRH)下,确实存在希望。在通常通过稀疏自编码器(SAE)恢复的 LRH 概念空间中,相近的点共享一个较小的共同激活支持集。利用这一局部稀疏性洞察,我们提出了一个基于局部掩码 SAE 的异常检测框架,并提供了理论依据。我们在多种架构和数据集上对该方法进行了验证,包括能力测试数据集和特定的安全数据集。最后,当允许算法使用 1% 的分布外数据进行校准时,局部稀疏方法达到了接近最优的性能,表明其能够捕捉有意义的安全信息,同时仅使用 1-2% 的 SAE 神经元进行计算。
cs.LG / 56 / 2609.20138

Fast-varying Natural Frequencies and Damping Ratio Identification for Linear Time-Varying System

线性时变系统快变固有频率与阻尼比辨识
Bozaci, Melisa, Cicirello, Alice
Abstract
This work proposes a physics-enhanced machine learning approach for the system identification of Linear Time-Varying (LTV) systems under time-varying operating conditions in terms of fast-varying natural frequencies and damping ratios by combining a long short-term memory network with an Extended Kalman Filter (EKF). The proposed approach uses vibration data (displacement and velocity measurements), domain knowledge of modal damping ratios, and a physics-based model that can yield an approximate natural frequencies time-dependency model. The approach is validated using synthetic data generated from a finite element model of a 2-blade offshore wind turbine under realistic environmental and operating conditions. This system displays fast time-varying frequencies due to operating conditions, whose identification is particularly challenging because of the wind and wave loading. The robustness of the proposed approach is assessed under assumed incorrect system information (e.g. damping ratio). The proposed approach is evaluated across different environmental and operating conditions to show its applicability to different operating regimes. The results show the approach can accurately identify the selected fast-varying natural frequency, 1st Fore-Aft (FA-1) mode, with a maximum root mean square error of 0.0012 Hz. The results demonstrate that the model trained on EKF estimates depends on accurate damping values, whereas the model trained on physics-based data exhibits robustness to incorrect damping assumptions. The approach is extended to damping ratio identification for the selected mode by estimating the root mean square error between models trained on EKF estimates and physics-based data. The results show that the approach can yield a good approximation of the FA-1 mode damping ratio using grid search, offering an improvement over covariance-driven stochastic subspace identification.
Chinese Translation
本文提出了一种物理增强的机器学习方法,通过将长短期记忆网络(LSTM)与扩展卡尔曼滤波器(EKF)相结合,对快变固有频率和阻尼比等时变工况下的线性时变(LTV)系统进行系统辨识。该方法利用振动数据(位移和速度测量值)、模态阻尼比的领域知识,以及一个能够给出固有频率随时间近似变化规律的物理模型。该方法使用基于双叶片海上风力机有限元模型在真实环境与运行工况下生成的合成数据进行验证。该系统由于运行工况的影响呈现出快时变频率特性,且由于风浪载荷的作用,其辨识尤为困难。在假设系统信息(如阻尼比)不准确的情况下,评估了所提方法的鲁棒性。该方法在不同环境和运行工况下进行了评估,以展示其对不同运行状态的适用性。结果表明,该方法能够准确辨识所选的快变固有频率,即一阶前后方向(FA-1)模态,最大均方根误差为0.0012 Hz。结果表明,基于EKF估计值训练的模型依赖于准确的阻尼值,而基于物理数据训练的模型对不正确的阻尼假设表现出鲁棒性。该方法还通过估计基于EKF估计值训练的模型与基于物理数据训练的模型之间的均方根误差,扩展至所选模态阻尼比的辨识。结果表明,采用网格搜索可以较好地近似FA-1模态的阻尼比,相比协方差驱动的随机子空间辨识方法有所改进。
cs.LG / 57 / 2609.20156

QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization

QUALS:基于模式量化与可学习性同步的通用预测语料库均衡方法
Li, Yujie, Shao, Zezhi, Yu, Chengqing, Fu, Yisong, Zhu, Weijie, Du, Yifan, Hu, Jilin, Yang, Bin, Xu, Yongjun, Wang, Fei
Abstract
Ubiquitous time series data across diverse domains enables critical applications in areas such as transportation systems and power grids. Recently, training foundation models on massive datasets to achieve accurate zero-shot forecasting has emerged as a major research focus. However, current studies predominantly prioritize architectural innovations while insufficiently addressing data diversity, often relying on simple data sampling strategies that fail to manage complex data distributions effectively, leading to inefficient use of training data and suboptimal performance. To address this, we propose QUALS, a large-scale time series corpus equilibrium framework. QUALS significantly enhances data efficiency, i.e., enabling existing models to achieve superior performance using only a small fraction of the original training data. Specifically, QUALS operates through two core mechanisms. First, a pattern quantization framework systematically decodes heterogeneous patterns from mixed corpora via vector quantization and uniform binning. Second, a learnability synchronization framework calibrates sampling weights for heterogeneous patterns, bridging the optimization gap between simple and complex motifs to maximize overall training efficiency. Extensive benchmarks demonstrate that pre-training on QUALS consistently achieves superior zero-shot performance, even under substantially reduced training budgets.
Chinese Translation
遍布不同领域的广泛时间序列数据为交通系统和电网等领域的关键应用提供了支撑。近年来,在海量数据集上训练基础模型以实现准确的零样本预测已成为重要的研究方向。然而,现有研究主要侧重于架构创新,而对数据多样性的关注不足,往往依赖简单的数据采样策略,无法有效管理复杂的数据分布,导致训练数据利用效率低下且性能欠佳。为此,我们提出了QUALS,一个大规模时间序列语料库均衡框架。QUALS显著提升了数据效率,即让现有模型仅使用原始训练数据的一小部分即可取得更优性能。具体而言,QUALS通过两个核心机制运作:其一,模式量化框架通过向量量化与均匀分箱系统地解码混合语料库中的异构模式;其二,可学习性同步框架为异构模式校准采样权重,弥合简单模式与复杂模式之间的优化差距,以最大化整体训练效率。大量基准实验表明,在QUALS上进行预训练始终能取得优越的零样本性能,即使训练预算大幅削减的情况下依然如此。
cs.LG / 58 / 2609.20162

A Noise Optimum in Rehearsal-Free Continual Learning: Isolation, Mechanism, and Scope

无重放持续学习中的噪声最优点:分离、机制与适用范围
Howe, Gunner Levi
Abstract
Injecting stochastic noise into a consolidation rule can improve a network's retention of earlier tasks up to an optimal level, then degrade it -- an inverted-U in retention vs. noise. This paper isolates what produces that optimum and maps where it holds, entirely in simulation. (1) Phenomenon: the retention inverted-U appears on several related-task continual-learning benchmarks (Split-MNIST, FashionMNIST, continual Yin-Yang). (2) Isolation: a magnitude-matched ladder shows the effect requires coherent restoring toward the consolidated weights -- a random-direction force of identical magnitude produces no optimum, and a coherent force toward the wrong target actively hurts. (3) Active ingredient: most of the optimum is recovered by coupling the anchor gain to the injected-noise variance sigma^2 -- a one-line rule that neither Ornstein-Uhlenbeck Adaptation (fixed gain) nor MESU (posterior-variance gain) implements. A forced Ornstein-Uhlenbeck calculation derives the rising flank and predicts that the optimal noise rises with per-task interference g -- confirmed out-of-sample in direction against pre-existing measurements (the exponent is unresolved at our grid). The barrier-conditioning of the originating Doob h-transform is a low-sigma safety net that bounds forgetting where the coupled gain is too weak. (4) Scope: the optimum requires shared task structure -- it is absent on permuted-MNIST, and a controlled rotated-vs-permuted comparison localizes the boundary to task structure; the precise governing quantity is left open. (5) Length: at matched severity the advantage persists but attenuates with task count, and we show no rotation family can attribute the trend (a compact-group identity). A single-seed BrainScaleS-2 demonstration of the originating rule is reported separately (Howe, arXiv:2607.06924); this paper makes no hardware claim.
Chinese Translation
向巩固规则中注入随机噪声可以将网络对早期任务的记忆保持能力提升至某个最优水平,超过该水平后则会使其下降——即保持能力与噪声之间呈倒U形关系。本文完全在仿真环境中分离出产生该最优点的因素,并绘制了其适用范围。(1)现象:该保持能力的倒U形出现在若干相关任务的持续学习基准上(Split-MNIST、FashionMNIST、continual Yin-Yang)。(2)分离:幅度匹配的对照阶梯实验表明,该效应要求对已巩固权重施加相干的恢复力——相同幅度的随机方向力不产生最优点,而指向错误目标的相干力则会主动造成损害。(3)有效成分:该最优点的绝大部分可以通过将锚定增益与注入噪声方差 sigma^2 相耦合来恢复——这是一条既非固定增益的 Ornstein-Uhlenbeck Adaptation、也非后验方差增益的 MESU 所实现的单行规则。通过一个受迫 Ornstein-Uhlenbeck 计算,本文推导出倒U形的上升侧,并预测最优噪声随每任务干扰 g 上升——该预测在方向上得到了与已有测量结果(超出训练样本)的证实(在我们的网格上指数仍未确定)。源自 Doob h-变换的势垒条件化是一个低 sigma 的安全网,在耦合增益过弱时限制遗忘。(4)适用范围:该最优点要求任务间存在共享结构——在 permuted-MNIST 上它不存在,且一项受控的旋转与置换对比实验将边界定位于任务结构;其确切的支配量仍待研究。(5)长度:在匹配严重程度的条件下,该优势随任务数量增加而持续存在但逐渐减弱,并且我们证明任何旋转任务族都无法归因于该趋势(一个紧群恒等式)。原始规则的单种子 BrainScaleS-2 硬件演示另行报道(Howe, arXiv:2607.06924);本文不提出任何硬件层面的主张。
cs.LG / 59 / 2609.20166

Small Enough to Know Everything: The Fully-Enumerable Transformer as an Instrument for the Science of Delayed Generalization

小到足以洞悉一切:全可枚举Transformer作为延迟泛化科学的研究工具
Ootani, Yoshiyuki
Abstract
Tiny transformers trained on fully-enumerable tasks occupy an unusual position in the study of grokking: every input can be evaluated, every generalization ceiling can be computed exactly, and hundreds of seeds cost minutes. We argue this regime is a scientific instrument with four capabilities that approximate settings cannot offer: (a) exact, falsifiable generalization ceilings; (b) task surgery that manipulates one structural variable while provably fixing all others; (c) direct observation of every weight; and (d) survival-time statistics over many seeds that recast "does not grok" as a censored observation. The obvious objection is that laws characterized at 10^4 parameters may not mean anything beyond them. We answer it with a preregistered conservation study: three task-side laws established at 12K parameters -- a recoverability-ceiling law, a role-conflict delay law, and a weight-decay response law -- are re-measured under an identical from-scratch protocol at 12K, 1M, and 50M parameters (a 4,000x span; 360 runs plus a 44-run control arm). The ceiling law and the delay law are conserved (0/144 Holm-corrected ceiling violations; Spearman rho >= 0.75 at every scale, permutation p < 1e-4), while the weight-decay law deforms systematically, steepening with scale. Preregistered controls show the 50M role-conflict deficit survives learning-rate adjustment and a tripled budget. Conservation was tested against criteria frozen before data collection, and one law's deformation shows the test could have failed. These results license the fully-enumerable transformer as a model organism for the task-side laws of delayed generalization: what it measures exactly, larger models largely obey -- and where they deviate, the deviation is itself lawful and measurable.
Chinese Translation
在全可枚举任务上训练的微型Transformer在grokking(顿悟)现象研究中占据着独特地位:每个输入都可以被评估,每个泛化上限都可以被精确计算,而数百个随机种子的实验仅需数分钟。我们论证这一研究尺度是一种科学仪器,具备四种近似设置无法提供的能力:(a) 精确且可证伪的泛化上限;(b) 任务手术,即操纵单一结构变量同时可证明地固定其他所有变量;(c) 对每个权重的直接观测;(d) 跨多个随机种子的生存时间统计,将'未发生grokking'重新表述为删失观测数据。显而易见的质疑是:在10^4参数规模下刻画的规律在更大规模下可能毫无意义。我们通过一项预注册的守恒研究来回应这一质疑:在1.2万参数规模下确立的三条任务侧规律——可恢复性上限规律、角色冲突延迟规律和权重衰减响应规律——在从零训练的相同协议下于1.2万、100万和5000万参数规模(跨4000倍范围;360次运行加上44次运行的对照组)被重新测量。上限规律和延迟规律得以守恒(0/144次经Holm校正的上限违反;每个尺度上Spearman rho >= 0.75,置换检验p < 1e-4),而权重衰减规律则发生系统性形变,随规模增大而变陡。预注册对照实验表明,5000万参数下的角色冲突缺陷在学习率调整和三倍训练预算后依然存在。守恒性检验的标准在数据收集前即被冻结,而其中一条规律的形变表明该检验本可能失败。这些结果证明了全可枚举Transformer可作为延迟泛化任务侧规律的模式生物:它能精确测量到的规律,更大的模型在很大程度上予以遵循——而在偏离之处,偏离本身也是有规律且可测量的。
cs.LG / 60 / 2609.20171

Support Thresholds, Not Algorithms, Limit Rare-Association Recovery in Co-Purchase Networks

限制共购买网络中罕见关联恢复的是支持度阈值,而非算法
Han, Xiao, Zhang, Zhen, Zhao, Xin, Lei, Jiechun, Zheng, Moxuan, Wang, Youting
Abstract
The support threshold of the Apriori algorithm involves a trade-off in conducting market basket analysis: the associations that occur frequently are noted with high threshold; however, the low ones lead to generating the large amount of rules. The paper compares five methods for co-purchase edge filtration on two grocery datasets: i.e., Instacart (3.2 million baskets) and Dunnhumby (208 thousand baskets), including Apriori, Apriori + lift post-filtering, top-$K$ ranking based on lift, and two methods based on networks, noise-corrected (NC) and disparity filter (DF). The top-$K$ method ensures the maximum average lift, while the NC achieves similar lift level by means of a single value of the significance parameter ($\alpha$). These two methods recover substantially more rare high-lift associations than Apriori (80-100% against 22-28%). NC and top-$K$ select meaningfully different edges (18-29% non-overlapping): NC retains statistically validated pairs, while top-$K$ retains rare pairs with high lift but low statistical significance. A rolling-origin holdout evaluation shows that top-$K$ edges recur at higher rates at every split, but NC edges are ~12 pp more likely to remain statistically significant in the held-out network.
Chinese Translation
Apriori算法的支持度阈值在市场篮子分析中存在权衡:较高的阈值能够识别频繁出现的关联;然而,较低的阈值会产生大量的规则。本文在两个食品杂货数据集(即Instacart,320万个购物篮;Dunnhumby,20.8万个购物篮)上比较了五种共购买边过滤方法,包括Apriori、Apriori加提升度后过滤、基于提升度的top-$K$排序,以及两种基于网络的方法:噪声校正过滤(NC)和差异过滤(DF)。top-$K$方法确保了最大的平均提升度,而NC方法通过单一显著性参数值($\alpha$)即可达到相近的提升度水平。与Apriori相比,这两种方法能够恢复多得多的罕见高提升度关联(80-100% 对 22-28%)。NC和top-$K$所选择的边存在实质性差异(18-29%不重叠):NC保留统计上经过验证的商品对,而top-$K$保留提升度高但统计显著性低的罕见商品对。基于滚动起点的保留集评估显示,top-$K$的边在每个划分中重复出现的比率更高,但NC的边在保留网络中保持统计显著性的可能性高出约12个百分点。
cs.LG / 61 / 2609.20174

Robust Federated Q-Learning with Almost No Communication

几乎无需通信的鲁棒联邦Q学习
Maity, Sreejeet, Mitra, Aritra
Abstract
We consider a federated reinforcement learning setting involving $M$ agents, all of whom interact with a common Markov Decision Process (MDP). The agents exchange information via a central server to learn the optimal value function. Our goal is to understand to what extent one can hope for collaborative sample-complexity speedups in such a setting, when a small fraction of the agents are adversarial and can act arbitrarily. To that end, we propose Robust Fed-Q}, a federated Q-learning algorithm that blends ideas from both model-based and model-free RL, along with the median-of-means device from robust statistics. We prove that despite corruption, with high-probability, Robust Fed-Q (i) guarantees exact convergence to the optimal value function in the limit of infinite samples, and (ii) enjoys near-optimal finite-time rates that benefit from collaboration. In addition, our approach requires just $\tilde{O}(1)$ rounds of communication to achieve each of the above guarantees, a feature of independent interest in FL where communication is the major bottleneck.
Chinese Translation
我们考虑一个包含 $M$ 个智能体的联邦强化学习场景,其中所有智能体都与同一个马尔可夫决策过程(MDP)进行交互。智能体通过中央服务器交换信息以学习最优值函数。我们的目标是:当一小部分智能体是对抗性的且可以任意行动时,在这种设定下协作式样本复杂度加速究竟能达到何种程度。为此,我们提出了 Robust Fed-Q,这是一种联邦Q学习算法,它融合了基于模型的强化学习与无模型强化学习的思想,并结合了鲁棒统计中的均值中位数方法。我们证明,尽管存在数据污染,Robust Fed-Q 仍能以高概率(i)在样本数趋于无穷时保证精确收敛到最优值函数,并且(ii)获得得益于协作的接近最优的有限时间收敛速率。此外,我们的方法仅需 $ ilde{O}(1)$ 轮通信即可实现上述各项保证,这一特性在通信为主要瓶颈的联邦学习(FL)中具有独立的研究价值。
cs.LG / 62 / 2609.20193

When Does Retrieval Help Time-Series Forecasting?

检索何时有助于时间序列预测?
Cakiroglu, Mert Onur, Buxton, Elham, Dalkilic, Mehmet, Kurban, Hasan
Abstract
Retrieval plug-ins supply a deep forecaster with information its lookback window cannot carry. Published evaluations report consistent gains, and each credits its own mechanism. We show that the benefit belongs instead to the operating point: the relation between window length $S$ and dominant seasonal period $L$, an axis the standard protocol never varies. Stratifying the evaluation by that relation exposes the regime. At $S{=}12$, a simple control that repeats the last observed period beats the six standard backbones, in aggregate, on four of seven benchmarks by $8\%$ to $44\%$ of MSE. It beats the strongest plug-in we run on ETTm1 and matches it on ECL. It is worse by up to $25\%$ on the three datasets whose training-split spectra lack a concentrated, shared period. A controlled synthetic sweep of horizon, period, and window shows the benefit boundary tracks the period (correlation $+0.71$), not the horizon ($-0.23$). A paired control with no phase to recover nearly erases the effect, consistent with phase starvation. Zero-shot pretraining does not escape it: a foundation model trails trained backbones by $22\%$ to $50\%$ on the periodic benchmarks. Within our instrument, exact lookup matches graph diffusion: the payoff is consulting the record, not the machinery on top. Two interpretable statistics, a trend test and a staleness rate, predict the sign of the per-cell benefit at $0.76$ accuracy under leave-one-dataset-out evaluation, a suggestive margin over the $0.69$ majority rule, where a 22-feature stack manages $0.57$. We propose no new plug-in. The contribution is the regime map, the protocol that reveals it, and two statistics that screen it before deployment. Code: https://github.com/KurbanIntelligenceLab/retrieval-regime.
Chinese Translation
检索插件为深度预测器提供其回看窗口无法承载的信息。已发表的评价报告了一致的收益,且各自将其归功于自身的机制。我们证明这一收益实际上属于工作点:即窗口长度 $S$ 与主导季节周期 $L$ 之间的关系——这是标准评测协议从未变动的一个维度。按该关系对评测进行分层可以揭示收益的适用区域。在 $S{=}12$ 时,一个简单地重复最近观测周期的对照方法,在七个基准中的四个上以相当于 MSE 的 8% 到 44% 的幅度总体优于六种标准主干模型;在 ETTm1 上优于我们运行的最强插件,在 ECL 上与之相当;而在三个训练集频谱缺乏集中且共享周期的数据集上,它最多落后 25%。对预测时域、周期和窗口的受控合成扫描显示,收益边界跟踪的是周期(相关性 $+0.71$),而非预测时域($-0.23$)。一个无相位可恢复的配对对照几乎完全消除了该效应,这与“相位匮乏”假说一致。零样本预训练也无法摆脱这一规律:一个基础模型在周期性基准上比训练的主干模型落后 22% 至 50%。在我们的测试框架内,精确查找与图扩散效果相当:收益来自于查阅历史记录,而非其上的复杂机制。两个可解释的统计量——趋势检验和陈旧率——在留一数据集评估下能以 0.76 的准确率预测每个单元收益的正负号,相对于 0.69 的多数规则基线是一个有启发性的优势,而一个 22 特征的组合模型仅达到 0.57。我们不提出新的插件。本文的贡献在于揭示收益适用区域的图谱、揭示该图谱的评测协议,以及可在部署前对其进行筛选的两个统计量。代码:https://github.com/KurbanIntelligenceLab/retrieval-regime。
cs.LG / 63 / 2609.20194

SoftTri: Smooth Triangular Membership Functions for Adaptive Fuzzy Inference Systems

SoftTri:用于自适应模糊推理系统的平滑三角隶属函数
Sarani, Babak, Ardakanian, Rahman, Mousavi, Ali
Abstract
Triangular membership functions (MFs) are widely used in fuzzy systems because of their interpretability, low parameterization complexity, and strong locality properties. However, their inherent nondifferentiability at knot points limits the effectiveness of gradient-based optimization in adaptive neuro-fuzzy architectures, often necessitating subgradient approximations or heuristic smoothing techniques. In this paper, we propose \emph{SoftTri}, a differentiable triangular membership function constructed using a smooth soft-hinge mechanism inspired by Swish-type activations. The proposed formulation preserves the geometric structure and localized behavior of classical triangular MFs while providing $C^\infty$ smoothness with respect to both the input variable and the membership parameters $(a,b,c)$ for any finite sharpness parameter $\beta>0$. Closed-form analytical gradients are derived to enable efficient and fully differentiable backpropagation-based learning. SoftTri is integrated into a Takagi--Sugeno fuzzy neural network with grid-partitioned rules and evaluated on multiple one-dimensional and two-dimensional nonlinear approximation benchmarks as well as a real-world regression task using the Airfoil Self-Noise dataset. Experimental results demonstrate that SoftTri consistently improves optimization stability and approximation accuracy compared with classical triangular membership functions, while achieving performance comparable to or better than Gaussian MFs under identical rule structures and training settings. The proposed approach provides an effective compromise between interpretability and differentiable optimization in modern neuro-fuzzy learning systems.
Chinese Translation
三角隶属函数(Membership Functions, MFs)因其可解释性强、参数化复杂度低以及良好的局部性特性而被广泛应用于模糊系统中。然而,其在节点处固有的不可微性限制了基于梯度的优化方法在自适应神经模糊架构中的有效性,通常需要采用次梯度近似或启发式平滑技术。本文提出SoftTri,一种受Swish类激活函数启发、采用平滑软铰链机制构建的可微三角隶属函数。所提出的公式在保持经典三角隶属函数的几何结构和局部化行为的同时,对于任意有限的锐度参数β>0,均能对输入变量和隶属参数(a,b,c)实现C∞光滑性。我们推导了闭式解析梯度,以实现高效且完全可微的基于反向传播的学习。将SoftTri集成到采用网格划分规则的Takagi-Sugeno模糊神经网络中,并在多个一维和二维非线性逼近基准任务以及使用Airfoil Self-Noise数据集的实际回归任务上进行评估。实验结果表明,与经典三角隶属函数相比,SoftTri在优化稳定性和逼近精度方面均有一致的提升;在相同的规则结构和训练设置下,其性能达到甚至优于高斯隶属函数。所提出的方法在现代神经模糊学习系统中为可解释性与可微优化之间提供了有效的折中方案。
cs.LG / 64 / 2609.20198

Evaluating Financial Sentiment in the Age of AI

人工智能时代的金融情绪评估
Bisharat, Arslan, Hean, Oudom
Abstract
Financial sentiment measures are widely used in empirical finance, but it remains unclear whether general-purpose large language models (LLMs) improve on existing finance-specific methods. This paper evaluates twelve sentiment models, including dictionary-based methods, finance-specific transformers, and open-source LLMs, using two criteria: linguistic validity and economic validity. We find that general-purpose LLMs achieve classification performance comparable to finance-specific transformer models without task-specific fine-tuning. However, higher classification accuracy does not translate into stronger economic relationships. Several models produce sentiment measures that are significantly associated with earnings surprises, but none is significantly associated with next-day stock returns. Model performance is strongest for announcements with large earnings beats or misses and substantially weaker for announcements with more moderate earnings surprises. These findings suggest that financial sentiment captures information about firms' economic performance but has limited ability to explain short-run market reactions
Chinese Translation
金融情绪指标在实证金融研究中被广泛使用,但通用大语言模型(LLMs)是否优于现有的金融专用方法仍不明确。本文从语言学有效性和经济学有效性两个标准出发,评估了十二种情绪模型,包括基于词典的方法、金融专用的transformer模型以及开源大语言模型。我们发现,通用大语言模型无需针对特定任务进行微调,即可实现与金融专用transformer模型相当的分类性能。然而,更高的分类准确率并不转化为更强的经济关联性。若干模型产生的情绪指标与盈余意外显著相关,但没有一个模型与次日股票收益显著相关。模型在盈余大幅超预期或大幅低于预期的公告上表现最佳,而在盈余意外较为温和的公告上表现明显较弱。这些发现表明,金融情绪能够捕捉企业经济表现的相关信息,但在解释短期市场反应方面能力有限。
cs.LG / 65 / 2609.20199

Towards a Unified Modality-Agnostic Multimodal Framework for Cognitive Workload Assessment

迈向统一的模态无关认知负荷评估多模态框架
Gkikas, Stefanos, Cruz, Christian Arzate, Joseph, Calvin, Giannakakis, Giorgos, Rojas, Raul Fernandez
Abstract
Cognitive workload reflects the mental effort required during task performance and is central to the design of adaptive human-machine systems. The use of biosignals to measure cognitive workload has been extensively researched and documented; however, studies examining the effects of combining heterogeneous biosignal modalities for this purpose remain limited. To provide insight into this area, we developed a unified, modality-agnostic, hierarchical Transformer-based architecture to process heterogeneous biosignal modalities within a single model. We use this framework in a pilot study evaluating all $31$ possible combinations of five modalities: Electrocardiogram (ECG), Electrodermal Activity (EDA), Respiration (RESP), Peripheral Oxygen Saturation (SpO$_2$), and Electroencephalogram (EEG), under leave-one-subject-out validation across three cognitively distinct tasks: abstract reasoning (IQ), arithmetic problem solving (MATH), and a game task (GAME). In this pilot setting, the results suggest that: (i) EEG is the strongest single modality, ranking highest in IQ, GAME, and the pooled ALL setting, where samples from all three tasks are combined; (ii) adding more modalities does not consistently improve performance; (iii) the full five-modality combination achieves the highest \textit{Average} score of $73.02%$ on IQ and $68.08%$ when the \textit{Average} scores are averaged over the four evaluation settings: IQ, MATH, GAME, and ALL; and (iv) the proposed method reduces model size by approximately $50%$ compared with late-fusion alternatives while maintaining a lower inference time.
Chinese Translation
认知负荷反映任务执行过程中所需的脑力投入,是自适应人机系统设计的核心。利用生物信号测量认知负荷已被广泛研究和记录;然而,关于组合异构生物信号模态以实现该目的的研究仍然有限。为深入探讨这一领域,我们开发了一个统一的、模态无关的、基于层次化Transformer(Transformer)的架构,可在单一模型内处理异构生物信号模态。我们在一项初步研究中使用该框架,评估了五种模态——心电图(ECG)、皮肤电活动(EDA)、呼吸信号(RESP)、外周血氧饱和度(SpO$_2$)和脑电图(EEG)——的全部$31$种可能组合,并在三个认知上不同的任务下采用留一被试交叉验证:抽象推理(IQ)、算术问题求解(MATH)和游戏任务(GAME)。在该初步设置中,结果表明:(i)EEG是最强的单一模态,在IQ、GAME以及汇总三种任务样本的ALL设置中排名最高;(ii)增加更多模态并不总能提升性能;(iii)完整的五模态组合在IQ任务上取得了最高的Average得分$73.02%$,且在IQ、MATH、GAME和ALL四种评估设置的Average得分的平均值上达到$68.08%$;(iv)与后期融合替代方案相比,所提出的方法在保持更低推理时间的同时,将模型规模减少了约$50%$。
cs.LG / 66 / 2609.20209

Scene-Conditioned Relation Routing for urban cellular activity forecasting

面向城市蜂窝活动预测的场景条件关系路由方法
Li, Qingzhong, Lin, Jingye, Ma, Hui, Zhang, Yajun, Pei, Xinjun, Yan, Ming, Xing, Fei
Abstract
Urban cellular activity forecasting requires jointly modeling heterogeneous spatiotemporal signals, including SMS usage, mobile network traffic, and call activity. Existing methods often separate temporal modeling, spatial relation learning, and multi-signal prediction, relying on fixed graph structures or static multi-task learning schemes, which limits their adaptability to changing urban scenes. We propose SCRR-Net, a scene-conditioned spatial relation routing framework in which urban contextual information jointly controls spatial dependency selection and cross-task knowledge transfer. SCRR-Net includes a context encoder, a spatial graph expert routing module, a temporal Transformer encoder, and a task knowledge routing module. Experiments on the Milano and Trento datasets demonstrate that SCRR-Net consistently outperforms competing methods on SMS, network traffic, and call activity forecasting, while providing interpretable routing behaviors.
Chinese Translation
城市蜂窝活动预测需要联合建模异构时空信号,包括短信使用量、移动网络流量和通话活动。现有方法通常将时间建模、空间关系学习和多信号预测分离,依赖固定的图结构或静态多任务学习方案,这限制了其对不断变化的城市场景的适应能力。我们提出SCRR-Net,一种场景条件空间关系路由框架,其中城市上下文信息联合控制空间依赖选择和跨任务知识迁移。SCRR-Net包括上下文编码器、空间图专家路由模块、时间Transformer编码器和任务知识路由模块。在Milano和Trento数据集上的实验表明,SCRR-Net在短信、网络流量和通话活动预测上持续优于竞争方法,同时提供可解释的路由行为。
cs.LG / 67 / 2609.20213

Subdomain-aware representation compression for pretrained image embeddings

面向子域的预训练图像嵌入表示压缩方法
Niamluang, Poowanut, Fakcharoenphol, Jittat
Abstract
Dimensionality reduction is a well-known technique for improving space efficiency, typically applied uniformly across an entire dataset. This paper investigates the possibilities of using dimensionality reduction techniques for subdomain representation compression. We explore standard techniques such as Principal Component Analysis (PCA) and Linear discriminant analysis (LDA) in image domains. The results not only demonstrate the expected improvements in space and computation complexity crucial for edge-device ML applications but also show improvements in accuracy over direct full-embedding procedure. One possible explanation is that dimensionality reduction effectively extracts subdomain features. We also performed experiments to demonstrate transfer learning capabilities using the compressed representations.
Chinese Translation
降维是一种众所周知的提升空间效率的技术,通常在整个数据集上统一应用。本文研究了利用降维技术进行子域表示压缩的可能性。我们在图像领域中探索了主成分分析(PCA)和线性判别分析(LDA)等标准技术。结果不仅展示了在空间和计算复杂度方面的预期改进——这对边缘设备机器学习应用至关重要——而且相较于直接使用完整嵌入的方法,准确率也有所提升。一种可能的解释是降维有效地提取了子域特征。我们还进行了实验,以验证基于压缩表示的迁移学习能力。
cs.LG / 68 / 2609.20217

Explaining spatial information flow in short-term traffic forecasting models using a gated graph attention network

使用门控图注意力网络解释短期交通预测模型中的空间信息流
Li, Yue, Chen, Shujuan, Jin, Ying
Abstract
Short-term traffic forecasting supports real-time monitoring and control of road networks, and graph attention networks (GAT) are the standard means of representing spatial dependence in these models. GAT layers are widely described as capturing the influence of neighbouring locations, but this is seldom verified, because the attention weights offered in support cannot be compared against any measured quantity. That leaves two questions open, how the model should be explained and which of its components are necessary. We address this by adding a gate to the GAT layer which learns, at every sensor and every time step, what share of a sensor's updated state is drawn from its neighbours rather than from itself. Regularising the gate withdraws neighbour information progressively and thereby provides a graded form of ablation. We apply the gated GAT to ST-MetaNet, whose encoder and decoder each place one GAT layer between two recurrent layers, and train it on one calendar year of records from 498 loop detectors on the strategic road network of England. The gate assigns a larger share of neighbour information to sensors carrying heavier traffic and follows the daily and weekly cycle of travel, consistent with adjacent locations being more strongly coupled when busy. Mild regularisation improves accuracy slightly, and accuracy declines at higher strengths as the penalty withdraws information the model needs. The encoder gate closes before the decoder gate, but direct ablation qualifies that ordering. Removing either GAT layer alone leaves accuracy at least as good as keeping both, whereas removing both degrades it substantially, so the two layers are largely redundant rather than either being indispensable. The gated GAT therefore yields a modest accuracy gain, an explanation of where and when spatial information flows, and evidence on which layers the architecture requires.
Chinese Translation
短期交通预测支持路网的实时监测与控制,而图注意力网络是此类模型中表征空间依赖性的标准手段。GAT层通常被描述为能够捕捉邻近位置的影响,但这一说法很少得到验证,因为模型提供的注意力权重无法与任何实测量进行比较。这留下了两个悬而未决的问题:模型应如何被解释,以及其哪些组件是必要的。为此,我们在GAT层中增加了一个门控,该门控在每个传感器和每个时间步上学习传感器更新状态中来自邻居而非自身的比例。对门控进行正则化可以逐步撤回邻居信息,从而提供一种分级的消融形式。我们将门控GAT应用于ST-MetaNet——其编码器和解码器各自在两个循环层之间放置一个GAT层——并使用英格兰战略路网上498个环形检测器一整个日历年的记录进行训练。门控为交通负荷较重的传感器分配了更大比例的邻居信息,并遵循出行的日周期和周周期,这与拥挤时相邻位置之间耦合更强的现象相一致。轻度正则化略微提升了精度,而在更强的正则化下,惩罚项撤回了模型所需的信息,精度随之下降。编码器门控在解码器门控之前关闭,但直接消融实验对这一先后顺序提出了修正。单独移除任一GAT层后的精度至少与保留两层时相当,而同时移除两层则使精度大幅下降,因此这两个层在很大程度上是冗余的,而非缺一不可。因此,门控GAT带来了适度的精度提升、对空间信息流在何时何地发生的解释,以及关于该架构需要哪些层的证据。
cs.LG / 69 / 2609.20218

Is It Still Worth Training a Classical Model in the Era of LLMs? A Crossover Benchmark on Tabular Data

在大语言模型时代,训练经典模型还值得吗?一个关于表格数据的交叉对比基准研究
Ding, Kaihua
Abstract
Large language models can label a tabular row from a plain-English description with no training - a capability now shipping in mainstream spreadsheet tools such as Microsoft Copilot in Excel and Anthropic's Claude for Excel - raising a practical question for the many business prediction problems where labels are expensive: should you prompt a frozen LLM, or collect data and train a model - and if so, how much data? We quantify the answer with the labeled-data crossover N*, the training-set size at which a trained classical model's learning curve overtakes a frozen LLM's training-free (and therefore flat) error. Aggregating 126 independent student evaluations of small GPT models under eight prompting configurations across 18 tabular datasets, paired with authoritative power-law learning curves for six classical model families, we find that training wins fast: even given an oracle choice of its best prompt configuration, a trained classical model beats the small frozen LLM using no more labeled data than is already on hand in 86% of cases, and wins by the smallest labeled subset we evaluate in 40%, with the observed crossover at a median of ~6% of the training set. In-context few-shot examples do not behave like training - error versus shot count does not follow a power law - and the same protocol re-run by independent implementers varies with a coefficient of variation of 0.148. A controlled probe indicates the LLM depends on recognizable feature-name semantics, which plausibly makes our crossover a conservative estimate (we do not claim memorization). For a typical business table, the evidence is clear: collect a few hundred labels and train a gradient-boosted model.
Chinese Translation
大语言模型能够根据纯英文描述直接为表格数据行打标签,无需任何训练——这一能力现已集成于主流电子表格工具中,例如 Microsoft Copilot in Excel 和 Anthropic 的 Claude for Excel——这为许多标签成本高昂的业务预测问题提出了一个现实问题:应该提示一个冻结的大语言模型(LLM),还是收集数据并训练一个模型?如果是后者,需要多少数据?我们用标注数据交叉点 N* 来量化这一答案:N* 是训练后的经典模型的学习曲线超越冻结 LLM 的免训练(因此为平坦)误差曲线所需的训练集规模。通过汇总 18 个表格数据集上八种提示配置下对小规模 GPT 模型的 126 次独立学生评估,并结合六个经典模型家族的权威幂律学习曲线,我们发现训练模型很快便能胜出:即使在给予最优提示配置的神谕选择下,在 86% 的情形中,训练后的经典模型所击败小规模冻结 LLM 所需的标注数据不超过已掌握的数据量,并在 40% 的情形中以我们所评估的最小标注子集即获胜,观测到的交叉点中位数约为训练集的 6%。上下文中的少样本示例并不等同于训练——误差与样本数之间不遵循幂律关系——且独立实现者重复相同协议的变异系数为 0.148。一项受控探测实验表明,LLM 依赖于可识别的特征名称语义,这很可能使我们的交叉点估计偏于保守(我们并不声称存在记忆现象)。对于典型的业务表格,证据是明确的:收集几百条标签并训练一个梯度提升模型即可。
cs.LG / 70 / 2609.20249

Accuracy Is Not Enough: A Cross-Architecture Audit of Demographic Bias in Deep Knowledge Tracing

准确率并不足够:深度知识追踪中人口群体偏差的跨架构审计
Minh, Dang Quang, Son, Nguyen Dung, Loi, Nguyen Huu, Vu, Truong Viet, Anh, Nguyen Thai
Abstract
Deep knowledge tracing (DKT) models implicitly decide which students an adaptive system believes have mastered a skill, yet almost all evidence on their demographic fairness comes from Bayesian knowledge tracing; the deep models that power modern systems have received no comparable cross-architecture audit. We close this gap: four architectures (DKT, DKVMN, SAKT, AKT) trained under three regimes (standard, reweighting, adversarial) on two public datasets with demographic metadata, Eedi (15.9M interactions) and OULAD (167k after preprocessing), evaluated with ABROCA, student-level bootstrap confidence intervals, and permutation tests addressing recent critiques of fairness-metric instability. Three findings emerge. (i) Bias is real but context-dependent: every architecture shows a significant socioeconomic ABROCA on Eedi (0.018-0.023, $p<0.005$), with per-group AUC lower for economically disadvantaged students, while gender bias is significant on OULAD for three of four architectures after multiplicity correction yet negligible on Eedi. (ii) The most accurate architecture is the most biased: AKT gains about 4 AUC points from item-level Rasch embeddings and shows the largest socioeconomic ABROCA, exceeding every other architecture under a paired bootstrap ($p\leq0.002$); ablating only the Rasch embeddings removes the accuracy gain and the excess bias together. (iii) Standard mitigation is unreliable: reweighting and adversarial debiasing leave ABROCA essentially unchanged in every configuration that preserves accuracy, even though the adversary is pinned at chance at full reversal strength and a weak-strength positive control rules out a dead probe.
Chinese Translation
深度知识追踪(DKT)模型在隐性地决定自适应系统认为哪些学生已掌握某项技能,然而关于其人口群体公平性的证据几乎全部来自贝叶斯知识追踪;驱动现代系统的深度模型尚未接受过类似的跨架构审计。我们填补了这一空白:在两个包含人口统计元信息的公开数据集——Eedi(1590万次交互)和OULAD(预处理后16.7万条记录)——上,训练四种架构(DKT、DKVMN、SAKT、AKT),采用三种训练方式(标准、重加权、对抗),并使用ABROCA、学生层面的自助法(bootstrap)置信区间以及针对公平性指标不稳定性近期批评的置换检验进行评估。研究得出三点发现。(i)偏差真实存在但依赖于具体情境:所有架构在Eedi数据集上均表现出显著的社会经济地位ABROCA(0.018–0.023,p<0.005),经济困难学生的分组AUC更低;而在经多重性校正后,四种架构中有三种在OULAD上表现出显著的性别偏差,但在Eedi上性别偏差可忽略不计。(ii)准确率最高的架构偏差最大:AKT通过题目层面的Rasch嵌入获得约4个AUC点的提升,并表现出最大的社会经济地位ABROCA,在配对自助法检验下超过其他所有架构(p≤0.002);仅消融Rasch嵌入即可同时消除准确率提升和额外偏差。(iii)标准缓解手段不可靠:在所有保持准确率的配置中,重加权和对抗去偏几乎未改变ABROCA,即使对抗器在完全反转强度下被固定在随机水平,且弱强度阳性对照排除了探测失败的可能性。
cs.LG / 71 / 2609.20250

How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU?

小于3B参数的开源语言模型在8GB消费级GPU上进行零样本作文评分能走多远?
Son, Nguyen Dung, Minh, Dang Quang, Loi, Nguyen Huu, Vu, Truong Viet, Anh, Nguyen Thai
Abstract
Zero-shot essay scoring with large language models is usually demonstrated with proprietary API models, yet the settings where automated scoring is most needed, such as public schools grading thousands of essays under strict privacy rules, are often those where sending student writing to a third-party API is unacceptable. We ask how much capability survives when the model must be a sub-3B open model running fully locally in FP16, with a controlled study of four instruction-tuned models from two families (Qwen2.5 at 0.5B/1.5B/3B, SmolLM2 at 1.7B) on all eight ASAP-AES prompts on a single 8 GB consumer GPU, with bootstrap confidence intervals, Holm-corrected paired tests, and deployment-realistic variants of the key design choices. Three findings emerge. (i) Rubric-decomposed prompting beats holistic prompting for every model under batch min-max aggregation (though Qwen2.5-3B drops significantly on one prompt), and under mean aggregation two unrelated families land within 0.01 at the 1.5-1.7B scale. (ii) Mapping trait scores into the prompt range is fragile to grader calibration: one model compresses traits into a narrow low band (2-4 on 0-10) and naive mean aggregation collapses, while the min-max normalization of Multi-Trait Specialization repairs it (macro QWK 0.204 to 0.388) and stays within 0.03 when its statistics are frozen on 30 held-out essays. (iii) Signed error falls with essay length in eleven of twelve configurations, opposite to the verbosity bias reported for large LLM judges; normalized rubric decomposition largely flattens this slope for well-calibrated models. We anchor results honestly: the best local configuration (0.388) remains far below both the human inter-rater ceiling (0.769) and a length-only baseline (0.523), so we position sub-3B local models strictly for formative, human-supervised feedback.
Chinese Translation
基于大语言模型的零样本作文评分通常以专有API模型进行演示,然而最需要自动评分的场景——例如公立学校在严格隐私规定下批改数千篇作文——往往无法接受将学生作文发送至第三方API。我们探究当模型必须是完全在本地以FP16运行的小于3B参数的开源模型时,还能保留多少能力。我们对来自两个家族的四个指令微调模型(Qwen2.5的0.5B/1.5B/3B版本和SmolLM2的1.7B版本)进行了控制实验,在单个8GB消费级GPU上对全部八个ASAP-AES作文题进行评测,采用自助法置信区间、经Holm校正的配对检验,并测试了关键设计选择在部署现实条件下的变体。研究得出三点发现。(i)在批处理最小-最大聚合下,量规分解式提示在所有模型上都优于整体式提示(尽管Qwen2.5-3B在某一道题上显著下降);而在均值聚合下,两个不相关家族的模型在1.5-1.7B规模上表现相差不超过0.01。(ii)将维度分数映射到提示所要求的分数范围对评分者校准非常脆弱:一个模型将分数压缩到狭窄的低分段(0-10分制中的2-4分),导致朴素的均值聚合失效,而多维度特化(Multi-Trait Specialization)的最小-最大归一化方法修复了这一问题(宏观QWK从0.204提升至0.388),且当其统计量在30篇保留作文上冻结时,性能仍保持在0.03以内。(iii)在十二种配置中的十一种里,带符号误差随作文长度增加而下降,这与已有报道中大模型评判者的冗长偏差相反;归一化量规分解对校准良好的模型在很大程度上消除了这种斜率。我们诚实地锚定结果:最佳本地配置(0.388)仍远低于人类评分者间一致性上限(0.769)和仅基于长度的基线(0.523),因此我们严格将小于3B参数的本地模型定位于形成性评估中由人工监督的反馈场景。
cs.LG / 72 / 2609.20268

Improving Online Reinforcement Learning via Bidirectional Behavior Prior Distillation

基于双向行为先验蒸馏的在线强化学习改进方法
Gao, Gong, Lai, Xiao, Shen, Jiaji, Jia, Ning, Liu, Xianhui, Zhao, Weidong
Abstract
Online reinforcement learning (RL) algorithms frequently exhibit poor sample efficiency and unstable learning dynamics, stemming from systematic critic estimation errors that are exacerbated by greedy policy updates. Existing behavior-prior reinforcement learning methods attempt to alleviate this issue by relying on offline pre-training to learn behavior models from fixed datasets and using policy priors to constrain online policy updates. However, the limited quality of offline datasets often hinders the ability to provide high-value policies that can effectively guide policy updates. The absence of expert trajectories significantly impairs online policy learning, leading to low sample efficiency and suboptimal performance. To address these challenges, we depart from conventional behavior prior approaches and propose a Bidirectional Behavior Prior Distillation (B2PD) algorithm. B2PD leverages action-value priors to guide a conditional variational autoencoder (CVAE) in generating a high-value behavior support set. The resulting expert behavior priors are further distilled into the agent, effectively reducing inefficient exploration and enabling stable policy optimization, while establishing a bidirectional knowledge flow mechanism. Empirical evaluations on both state- and pixel-based tasks verify that B2PD substantially improves sample efficiency while maintaining stable policy optimization. More broadly, this work shows that enforcing high-quality behavioral support during online learning effectively mitigates critic-induced error amplification, enabling structured behavior priors to guide policy updates in a principled and sample-efficient manner.
Chinese Translation
在线强化学习(RL)算法通常表现出较差的样本效率和不稳定的学习动态,其原因在于系统性的评论家(critic)估计误差在贪婪策略更新的作用下被进一步放大。现有的行为先验强化学习方法试图通过离线预训练从固定数据集中学习行为模型,并利用策略先验来约束在线策略更新,从而缓解这一问题。然而,离线数据集的质量有限,往往难以提供能够有效指导策略更新的高价值策略。专家轨迹的缺失严重阻碍了在线策略学习,导致样本效率低下和性能次优。为应对这些挑战,我们突破传统行为先验方法的框架,提出了一种双向行为先验蒸馏(Bidirectional Behavior Prior Distillation, B2PD)算法。B2PD 利用动作价值先验引导条件变分自编码器(CVAE)生成高价值行为支撑集,并将所得到的专家行为先验进一步蒸馏到智能体中,从而有效减少低效探索、实现稳定的策略优化,同时建立了双向知识流动机制。在基于状态和基于像素的任务上的实验评估表明,B2PD 在保持稳定策略优化的同时,显著提升了样本效率。更广泛地说,这项工作表明,在在线学习过程中引入高质量的行为支撑,能够有效缓解由评论家引起的误差放大问题,使结构化行为先验能够以有原则且样本高效的方式指导策略更新。
cs.LG / 73 / 2609.20269

Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks

放置是自由的,构成则不然:拉丁方作为异构序列混合器堆叠的可证明均衡构造
Kim, Taebong, Hong, Youngsik, Kim, Minsik, Choi, Sunyoung, Jang, Jaewon, Kim, Minseo
Abstract
Since GPT, most Transformers have repeated the same attention mechanism at every layer. Yet this design is largely a convention rather than a tested conclusion. When multiple sequence mixers are combined in one stack, improvements may arise from mechanism choice, placement, or both, making causal attribution difficult. We introduce Aether-7B-5Attn, a 6.59B-parameter mixture-of-experts model ($\approx$2.98B active) whose 49 layers contain seven sequence-mixing mechanisms arranged as a $7\times7$ Latin square. Because each mechanism appears exactly once in every row and column, the design guarantees balanced exposure across depth while eliminating placement confounds. To evaluate this principle, we build a parameter-matched proxy with four mechanisms arranged as a $4\times4$ Latin square over sixteen layers, matched to 700.9M parameters and trained with eight seeds per arm. The results reveal a clear dissociation. Rearranging a distributed heterogeneous stack into a balanced periodic cycle changes validation loss by only 0.16\%, indicating that exact placement has little effect. In contrast, clustering the same mechanisms into contiguous depth bands incurs a 0.59\% penalty, while replacing the heterogeneous stack with a homogeneous one incurs a 1.68\% penalty. These results indicate that performance depends primarily on heterogeneous composition distributed across depth rather than on any particular permutation. We confirm this finding at 2.16$\times$ larger scale (1.514B parameters), where the homogeneous-stack penalty increases to 2.63\% and removing the SSM-family mechanism produces a 3.20\% degradation. We further report per-mechanism cost profiles, English and Korean evaluations, and a causal-safety audit of all 49 layers. We release model weights, training recipes, training code, logs, and architecture source code.
Chinese Translation
自GPT以来,大多数Transformer在每一层都重复使用相同的注意力机制。然而,这种设计在很大程度上是一种惯例,而非经过检验的结论。当多种序列混合器(sequence mixer)组合在同一个堆叠中时,性能提升可能来自机制的选择、放置位置,或两者兼有,这使得因果归因变得困难。我们提出了Aether-7B-5Attn,一个参数量为65.9亿(约29.8亿激活参数)的混合专家模型(mixture-of-experts),其49层中包含七种序列混合机制,按照7×7拉丁方(Latin square)排列。由于每种机制在每一行和每一列中恰好出现一次,该设计保证了跨深度的均衡暴露,同时消除了放置位置带来的混淆因素。为评估这一原则,我们构建了一个参数匹配的代理模型,将四种机制按4×4拉丁方排列于十六层中,参数量匹配至7.009亿,每种配置使用八个随机种子进行训练。结果揭示出明显的分离现象:将分布式异构堆叠重排为均衡的周期性循环仅使验证损失变化0.16%,表明精确的放置位置几乎没有影响。相比之下,将相同机制聚集为连续的深度区段会带来0.59%的损失,而用同构堆叠替换异构堆叠则会带来1.68%的损失。这些结果表明,性能主要取决于分布于整个深度的异构构成,而非任何特定的排列方式。我们在2.16倍更大规模(15.14亿参数)上证实了这一发现:同构堆叠的损失增加至2.63%,移除SSM(状态空间模型)族机制则导致3.20%的性能下降。我们还报告了各机制的成本画像、英语和韩语评估,以及对全部49层的因果安全性审计。我们公开了模型权重、训练方案、训练代码、日志和架构源代码。
cs.LG / 74 / 2609.20278

Labeled Incidence Structures for Native Transformer Modeling of Text, Knowledge Graphs, and Hypergraphs

用于文本、知识图谱与超图的原生Transformer建模的标记关联结构
Godavarti, Mahesh
Abstract
Text, knowledge graphs, and hypergraphs all have elements that play distinct roles within relation instances, structure that is lost when data is flattened into token sequences. We introduce labeled incidence structures (LIS), a uniform representation that encodes each endpoint as $(x_d, s, e)$: content $x_d$, a role or slot $s$, and the relation instance $e$ in which that role appears. Because every data type maps to the same $(x_d, s, e)$ representation without flattening, a single standard transformer can process them all natively, structural differences are carried entirely by the operators, not the architecture. LIS assigns a structural address to each endpoint by composing a slot operator and an instance operator, $A(s,e) = R_s R_e$. We characterize when this factorization gives every token a unique, path-independent address. When it does, the natural operator comparing endpoint $j$ to endpoint $i$ is the relative transport $P_{j\to i} = A_i^{-1} A_j$, which gives attention a role- and relation-aware inductive bias without imposing an arbitrary sequence order. Additive encodings of the form "position term plus relation term" can miss information that depends jointly on $s$ and $e$. We prove this in a controlled example family: when the journey operator is approximated by the sum of a position-only term and a relation-only term, the approximation cannot capture how position and relation combine, only their separate effects. We also analyze persistent knowledge repositories. Identifiers tied to storage locations make models sensitive to storage order, while freely learned identifiers can become harder to control as the repository size $M$ grows relative to the sample size $n$. Computing relation-instance operators from content avoids this storage-order issue and yields a capacity bound independent of $M$, under fixed architectural and Lipschitz assumptions.
Chinese Translation
文本、知识图谱和超图中的元素在关系实例中扮演着不同的角色,而当数据被展平为词元序列时,这种结构信息便会丢失。我们提出标记关联结构(Labeled Incidence Structures, LIS),这是一种统一表示,将每个端点编码为 $(x_d, s, e)$:内容 $x_d$、角色或槽位 $s$,以及该角色所在的关系实例 $e$。由于每种数据类型都无需展平即可映射到相同的 $(x_d, s, e)$ 表示,单个标准Transformer即可原生地处理所有数据类型,结构差异完全由算子而非架构承载。LIS通过组合槽位算子与实例算子,即 $A(s,e) = R_s R_e$,为每个端点分配一个结构地址。我们刻画了该分解何时能为每个词元提供唯一的、与路径无关的地址。当满足该条件时,将端点 $j$ 与端点 $i$ 进行比较的自然算子是相对传输算子 $P_{j\to i} = A_i^{-1} A_j$,它为注意力机制提供了角色与关系感知的归纳偏置,同时不施加任意的序列顺序。形如“位置项加关系项”的加性编码可能遗漏同时依赖于 $s$ 和 $e$ 的信息。我们在一个可控的示例族中证明了这一点:当传输算子被近似为仅位置项与仅关系项之和时,该近似无法捕捉位置与关系的组合方式,只能捕捉它们的各自效应。我们还分析了持久知识库。与存储位置绑定的标识符使模型对存储顺序敏感,而自由学习的标识符随着知识库规模 $M$ 相对于样本量 $n$ 的增长可能变得更难控制。在固定的架构与Lipschitz假设下,从内容计算关系实例算子可避免这一存储顺序问题,并得到与 $M$ 无关的容量界。
cs.LG / 75 / 2609.20296

Personalising a Cross-User Surface Electromyography Encoder Under a Small Calibration Budget

在小校准预算下对跨用户表面肌电编码器进行个性化
Odeyemi, Jethro, Zhang, W. J.
Abstract
A myoelectric interface needs calibration from the user before it will function. Earlier work has treated calibration as a quantity, but has not asked the question of what a device should do with the calibration repetitions once they have been collected. This paper views personalizing the cross-user encoder as a design decision with a cost. Four alternative approaches to using exactly the same labeled repetitions were tested from a single cross-user encoder per held-out subject. Prototypical adaptation, linear probes, scaled fine-tuning and full fine-tuning were tested at every budget up to the maximum each database allows, five repetitions on DB1 and four on DB2 and DB5. Comparing four ways to spend a small calibration budget across 77 subjects, full fine-tuning is the most accurate at every budget, consistently enough that there is no exception among subsets of subjects. The result which impacts how one might make an engineering decision however is that a gradient free prototypical rule recovers 52 to 78 per cent of its benefit with no optimiser and no per-user copy of the weights, which makes personalisation something a worn device can do at donning time. The widespread intuition that a good representation only needs a fresh classifier is incorrect here. How well each method may perform relative to a per-user classifier that would be fitted by a clinic will depend on the specific database.
Chinese Translation
肌电接口在正常工作前需要用户进行校准。以往的研究将校准视为一个数量问题,但未曾探讨设备在收集校准重复数据后应如何利用它们。本文将从个性化跨用户编码器视为一项带有成本的设计决策。针对每个留出被试,从单个跨用户编码器出发,测试了四种使用完全相同的带标签重复数据的替代方法。在每一步校准预算(直至各数据库允许的最大值——DB1为五次、DB2和DB5为四次)上,测试了原型自适应(prototypical adaptation)、线性探针(linear probes)、缩放微调(scaled fine-tuning)和全量微调(full fine-tuning)。通过在77名被试上比较花费小额校准预算的四种方式,结果表明全量微调在每一预算水平下都最为准确,且在各被试子集中无一例外。然而,影响工程决策的关键结果是:一种无需梯度的原型规则能够获得全量微调收益的52%至78%,且无需优化器、也无需为每个用户保存模型权重副本,这使得个性化成为可穿戴设备在穿戴时即可完成的任务。“良好表征只需训练一个新分类器”这一普遍直觉在此并不成立。各方法相对于临床拟合的每用户分类器的表现,将取决于具体数据库。
cs.LG / 76 / 2609.20297

Intact-to-Amputee Transfer in Surface-EMG Gesture Decoding: Training Source and Calibration Budget

表面肌电图手势解码中的健全人-截肢者迁移:训练源与校准预算
Odeyemi, Jethro, Zhang, W. J.
Abstract
A recogniser trained on one person rarely transfers to the next, and useful performance usually demands a fresh round of labelled calibration from the end user. A systematic review of 1077 studies quantifies where the evidence is thin: amputees appear in about one in six. Here a montage-agnostic cross-user encoder is carried to eleven transradial amputees on a protocol matched to its intact-limb training data. Zero-shot cross-population transfer fails outright: the encoder requires labeled data from the new user before it begins decoding, and it then exceeds the per-user classifier a clinic would fit by 0.190 macro F1 at three repetitions and for every subject in the cohort. Given three labelled repetitions it reaches 0.779 macro-F1 against 0.589 for the per-user pipeline. Training on forty intact subjects produces better transfers to a new amputee than training on ten other amputees, and combining the two produces better transfers than either individually. The prediction pre-registered for this study, which extends the encoder's baseline-strength account with the premise that amputee EMG is less separable, holds true only after a few repetitions become available and after enriching the source pool with additional amputees. At a single repetition, and at every budget under a source matched to the intact-limb comparison, it fails. Thus, it locates the boundary of the proposed account.
Chinese Translation
在一个人身上训练的识别器很难迁移到另一个人身上,而有效的性能通常需要终端用户重新进行一轮标注校准。对1077项研究的系统性综述量化了证据薄弱之处:截肢者仅约占六分之一。本研究将一个与电极排布无关的跨用户编码器应用于十一名经桡骨截肢者,实验方案与其健全肢体训练数据相匹配。零样本跨人群迁移完全失败:编码器需要新用户的标注数据才能开始解码,随后其宏平均F1在三次重复动作及全部受试者上均超过临床会为单个用户拟合的分类器0.190。在三次标注重复动作下,其宏平均F1达到0.779,而单用户流程仅为0.589。在四十名健全受试者上训练比在十名其他截肢者上训练能产生更好的向新截肢者的迁移,而两者结合又优于任一单独使用。本研究预先注册的预测——在编码器的基线强度解释基础上加入截肢者肌电信号可分性较低这一前提——仅在获得若干次重复动作数据并向源数据池中补充更多截肢者后成立。在单次重复动作时,以及在源数据量低于与健全肢体对照相匹配的每种预算下,该预测均不成立。因此,本研究界定了所提解释的适用边界。
cs.LG / 77 / 2609.20298

A Learning Algorithm for Threshold Boolean Networks with Prescribed Fixed Points

一种具有指定不动点的阈值布尔网络学习算法
Ruz, Gonzalo A.
Abstract
We present a learning algorithm for inferring threshold Boolean networks (TBNs) with a prescribed set of fixed points. The proposed method employs a custom differentiable loss function that jointly enforces fixed point preservation, penalizes spurious attractors, encourages binary outputs, and promotes sparsity through L1 regularization. Applied to the FOS-GRN model of Arabidopsis thaliana, the approach achieved perfect reconstruction (i.e., all 10 desired fixed points and no spurious ones) in 5 out of 30 independent runs, recovering on average 8.53 $\pm$ 0.90 correct fixed points with no spurious attractors. In contrast, standard methods such as the Perceptron and Logistic Regression recovered up to 10 fixed points but introduced between 8 and 31 spurious ones. An additional analysis varying the sparsity coefficient ($\lambda$) confirmed that the method's performance and the structural properties of the inferred networks remain robust within a practical range (up to 0.01) of regularization strengths. Overall, the results demonstrate the effectiveness and stability of the proposed algorithm in capturing meaningful network dynamics under prescribed dynamical constraints.
Chinese Translation
我们提出了一种学习算法,用于推断具有指定不动点集合的阈值布尔网络(Threshold Boolean Networks,TBNs)。该方法采用一种自定义的可微损失函数,联合实现以下目标:保持指定不动点、惩罚伪吸引子、促进二值化输出,并通过L1正则化促进网络稀疏性。将该算法应用于拟南芥(Arabidopsis thaliana)的FOS-GRN模型,在30次独立运行中有5次实现了完美重构(即恢复了全部10个期望的不动点且没有伪吸引子),平均恢复了8.53 ± 0.90个正确的不动点,且无伪吸引子。相比之下,感知机(Perceptron)和逻辑回归(Logistic Regression)等标准方法虽然最多能恢复10个不动点,但会引入8至31个伪吸引子。通过改变稀疏性系数(λ)的补充分析证实,在实际正则化强度范围内(最高至0.01),该方法的性能及所推断网络的结构性质保持稳健。总体而言,实验结果表明,所提出的算法能够在指定动力学约束下有效且稳定地捕捉有意义的网络动力学。
cs.LG / 78 / 2609.20300

Improving Generalization and Robustness in Offline Reinforcement Learning via Boundary-Aware Data Augmentation

基于边界感知数据增强的离线强化学习泛化性与鲁棒性提升方法
Gao, Gong, Zhao, Weidong, Liu, Xianhui
Abstract
Current offline reinforcement learning (ORL) algorithms tend to overfit the training dataset and exhibit poor in-distribution generalization and robustness performance when deployed to real environments, thus compromising their effectiveness. Existing methods typically enhance in-distribution generalization and robustness by leveraging regularization techniques widely used in computer vision. However, due to the high sensitivity of low-level physical signals to distributional shifts, these methods still suffer from notable limitations in in-distribution generalization and robustness, making it difficult to achieve stable performance in complex environments. To address this issue, we theoretically analyze the error bounds of the behavior policy and action-value function trained with random episode interpolation, revealing that the error scales positively correlated with the distance between states. Based on this insight, we propose a method called $\bf{B}$oundary-$\bf{A}$ware $\bf{D}$ata $\bf{A}$ugmentation (BADA), which leverages neighboring states to construct interpolation boundaries, enabling the generation of synthetic data that more faithfully preserves the original data distribution. We first conduct qualitative studies in a toy environment, showing that BADA generates mixed samples that preserve desirable policy smoothness while accurately reconstructing multimodal value distributions. Extensive experiments on limited offline datasets further demonstrate that BADA attains state-of-the-art performance across diverse benchmarks.
Chinese Translation
当前的离线强化学习(Offline Reinforcement Learning, ORL)算法容易过拟合训练数据集,在部署到真实环境时表现出较差的分布内泛化能力和鲁棒性,从而影响其有效性。现有方法通常借助计算机视觉领域广泛使用的正则化技术来提升分布内泛化性和鲁棒性。然而,由于底层物理信号对分布偏移的高度敏感性,这些方法在分布内泛化与鲁棒性方面仍存在明显局限,难以在复杂环境中实现稳定性能。为解决这一问题,我们从理论上分析了基于随机片段插值训练的行为策略和动作价值函数的误差界,揭示了误差规模与状态间距离呈正相关。基于这一发现,我们提出了一种称为边界感知数据增强(Boundary-Aware Data Augmentation, BADA)的方法,该方法利用相邻状态构建插值边界,从而生成能够更忠实保留原始数据分布的合成数据。我们首先在一个玩具环境中进行了定性研究,结果表明 BADA 生成的混合样本在保持良好策略平滑性的同时,能够准确重建多模态价值分布。在受限离线数据集上的大量实验进一步证明,BADA 在多种基准测试中达到了最先进的性能。
cs.LG / 79 / 2609.20302

SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption

SAGG:面向异构损坏下鲁棒多模态学习的样本自适应梯度门控方法
Zhang, Wentao, Zhu, Yifan, Zhang, Yutong, Mo, Wentao
Abstract
Multimodal gradient balancing methods modulate encoder gradients with a shared scalar per modality, implicitly assuming that corruption is uniform across the training batch. In practice, corruption is sample-heterogeneous: within a single mini-batch, different samples may have different modalities corrupted. We prove that under this heterogeneous corruption model, any batch-level sample-agnostic linear estimator with a shared modulation parameter incurs an irreducible bias with respect to the clean-data gradient, and that sample-level all-or-nothing gating is the unique unbiased strategy within a natural distribution-free estimator class. Motivated by this result, we propose Sample-Adaptive Gradient Gating (SAGG), which makes a binary retain-or-discard decision per sample via an online feature-norm quality test and incorporates a truncation mechanism for variance control. We prove that SAGG-based SGD converges at the standard O(1/sqrt(T)) rate to stationary points of the clean loss without a corruption-dependent error floor, and derive a certified robustness radius for the independent-encoder architecture that connects per-modality Lipschitz constants to the classification margin. Experiments on Kinetics-Sounds and UCF-101 under Gaussian noise injection, partial modality missing, and natural contribution imbalance show that SAGG consistently outperforms ten existing methods, with the largest gains in high-corruption regimes where batch-level bias is most severe.
Chinese Translation
多模态梯度平衡方法通过为每个模态分配一个共享标量来调节编码器梯度,这隐含地假设损坏在整个训练批次中是均匀分布的。在实际中,损坏是样本异构的:在同一小批次(mini-batch)内,不同样本可能存在不同模态的损坏。我们证明,在这种异构损坏模型下,任何采用共享调节参数的批次级、样本无关的线性估计器,都会相对干净数据梯度产生不可消除的偏差;并且在一个自然的无分布估计器类中,样本级的“全有或全无”门控是唯一无偏策略。基于这一结果,我们提出了样本自适应梯度门控(Sample-Adaptive Gradient Gating, SAGG),该方法通过在线特征范数质量测试对每个样本做出保留或丢弃的二元决策,并引入截断机制以控制方差。我们证明,基于SAGG的SGD以标准O(1/sqrt(T))速率收敛至干净损失的驻点,且不存在依赖损坏程度的误差下限;同时,我们为独立编码器架构推导了一个可认证的鲁棒半径,将各模态的Lipschitz常数与分类间隔联系起来。在Kinetics-Sounds和UCF-101数据集上,于高斯噪声注入、部分模态缺失以及自然贡献不平衡等条件下的实验表明,SAGG持续优于十种现有方法,且在批次级偏差最为严重的高损坏场景中增益最大。
cs.LG / 80 / 2609.20309

Hypernetwork-Parameterized Spatially Adaptive Neural Operators for PDE Learning

基于超网络参数化的空间自适应神经算子用于偏微分方程学习
Zhang, Jiaquan, Zhang, Chaoning, Chen, Shuxu, Ye, Meng, Zhang, Xiaofeng, He, Qiang, Huang, Weifeng, Wang, Guoqing, Yang, Yang, Qin, Caiyan
Abstract
Spatially heterogeneous partial differential equations (PDEs) exhibit location-dependent dynamics arising from variations in geometry and physical coefficients. Existing neural operators improve localized modeling through multiscale features, attention mechanisms, or domain decomposition, yet their update rules often remain spatially shared. Hypernetwork-based methods adapt parameters across PDE instances but typically generate only one global parameterization per instance. Consequently, shared operators may underfit boundaries and high-gradient regions, with these localized errors accumulating during autoregressive rollout. We propose a spatially adaptive neural operator (SANO), which replaces this spatially shared parameterization with a spatially continuous field of location-dependent operator parameters. SANO uses Fourier-encoded coordinates and a coordinate-conditioned hypernetwork to generate spatial operator-conditioning codes at sampling points. A Hyper-Neural Element (HNE) mechanism interpolates these codes within local subregions, coupling neighboring operators while allowing their update rules to vary across space, and partition-of-unity weights assemble the overlapping local predictions. Experiments on one-, two-, and three-dimensional PDEs and two perforated-domain elliptic benchmarks show that SANO consistently outperforms competitive neural-operator, hypernetwork-based, and physics-informed baselines.
Chinese Translation
空间异质偏微分方程(PDE)由于几何形状和物理系数的变化而表现出依赖于位置的动力学特性。现有的神经算子通过多尺度特征、注意力机制或区域分解来改进局部建模,但其更新规则往往仍在空间上共享。基于超网络的方法能够在不同PDE实例间自适应调整参数,但通常每个实例仅生成一个全局参数化。因此,共享算子可能在边界和高梯度区域出现欠拟合,并且这些局部误差会在自回归推演过程中不断累积。我们提出一种空间自适应神经算子(SANO),用空间连续的、依赖于位置的算子参数场取代这种空间共享的参数化。SANO使用傅里叶编码坐标和坐标条件超网络在采样点处生成空间算子调节码。超神经单元(Hyper-Neural Element, HNE)机制在局部子区域内对这些调节码进行插值,在耦合相邻算子的同时允许其更新规则随空间变化,并利用单位分解权重对重叠的局部预测进行组装。在一维、二维和三维PDE以及两个多孔域椭圆基准上的实验表明,SANO持续优于具有竞争力的神经算子、基于超网络以及物理信息驱动的基线方法。
cs.LG / 81 / 2609.20310

ZeroHAT: Behavior-Conditioned Zero-Shot Human Activity Trace Generation

ZeroHAT:基于行为条件的零样本人类活动轨迹生成
Xu, Rongchao, Yu, Dahai, Jiang, Lin, Wang, Guang
Abstract
Human activity traces record individuals' timestamped visits to points of interest and are essential for applications such as mobility prediction and urban simulation. However, accessing large-scale HATs is challenging due to high collection costs and privacy concerns. Synthetic HAT generation offers a promising way to make such data available and has attracted growing interest from both industry and academia. Although many efforts have been devoted to this topic, most of them rely on real data from a region to generate synthetic data for the same region, which is infeasible for the many regions where real HATs are unavailable. To fill this gap, we propose ZeroHAT, a behavior-conditioned framework that generates synthetic HATs for a target region in a zero-shot manner by transferring behavioral patterns learned from real HATs in source regions and adapting them with publicly available contextual information about the target region. ZeroHAT has three key novel components: (i) a multidimensional consistency-aware intent extractor; (ii) a cross-region behavioral cloning module; and (iii) a behavior-conditioned activity realization module. We evaluate ZeroHAT on a ten-city benchmark, where extensive experiments show that ZeroHAT achieves 4.5-6.4x the normalized downstream utility of the strongest baseline and improves average fidelity by 15.6%-40.8% across target regions.
Chinese Translation
人类活动轨迹(Human Activity Traces, HATs)记录了个体在时间戳标注下对兴趣点的访问,对于移动性预测和城市模拟等应用至关重要。然而,由于高昂的采集成本和隐私问题,获取大规模人类活动轨迹十分困难。合成人类活动轨迹生成为此类数据的可用性提供了一条有前景的途径,并日益受到工业界和学术界的关注。尽管已有很多相关工作,但它们大多依赖某一地区的真实数据来生成同一地区的合成数据,这对于许多无法获得真实人类活动轨迹的地区并不可行。为填补这一空白,我们提出了ZeroHAT,这是一个基于行为条件的框架,通过迁移从源地区真实人类活动轨迹中学习到的行为模式,并利用目标地区公开可获取的上下文信息进行适配,以零样本方式为目标地区生成合成人类活动轨迹。ZeroHAT包含三个关键创新组件:(i)多维一致性感知的意图提取器;(ii)跨地区行为克隆模块;(iii)基于行为条件的活动实现模块。我们在包含十个城市的基准数据集上评估ZeroHAT,大量实验表明,ZeroHAT的归一化下游任务效用达到最强基线的4.5-6.4倍,并在各目标地区将平均保真度提升了15.6%-40.8%。
cs.LG / 82 / 2609.20313

EviRec: Continual Evidence Learning for Dual Cold-Start POI Recommendation

EviRec:面向双重冷启动兴趣点推荐的持续证据学习
Xu, Rongchao, Jiang, Lin, Wang, Guang
Abstract
Point-of-Interest (POI) recommendation is a core task in location-based services, yet most existing methods assume a fixed user population and POI catalog. Through a large-scale data-driven analysis of 10 U.S. cities, we identify substantial POI churn, user turnover, category drift, and decay in static POI memory, motivating the study of continual dual cold-start POI recommendation. To address this setting, we propose EviRec, a continual evidence-learning framework that estimates how much historical evidence should be trusted separately for each candidate POI. EviRec scores each visible candidate from three complementary views: a matching view based on the user's recent mobility profile, a transition-memory view that captures repeated mobility routines, and a lifecycle view that reflects candidate maturity. Because a near-zero transition score may indicate either irrelevance or insufficient observation, EviRec qualifies the evidence using each candidate's observation state and applies a reliability gate to adaptively route between transition-memory and lifecycle evidence. We evaluate EviRec on a full-year, five-city POI check-in dataset containing more than 30,000 users and 684,200 trajectories. Experimental results show that EviRec consistently outperforms state-of-the-art baselines, with the largest gains concentrated on cold-start queries. In particular, EviRec improves NDCG@10 by 20.4\% on Dual-New cases over the strongest baseline. In-depth analyses further confirm that these gains arise primarily from candidate-specific reliability gating while largely preserving previously learned mobility routines.
Chinese Translation
兴趣点(POI)推荐是基于位置服务的核心任务,然而现有方法大多假设用户群体和兴趣点目录是固定不变的。通过对美国10个城市的大规模数据驱动分析,我们发现了显著的兴趣点更替、用户流失、类别漂移以及静态兴趣点记忆的衰减,这促使我们研究持续的双重冷启动兴趣点推荐问题。为应对这一场景,我们提出了EviRec,一个持续证据学习框架,它能够针对每个候选兴趣点分别估计应信任多少历史证据。EviRec从三个互补的视角对每个可见候选进行评分:基于用户近期移动画像的匹配视角、捕捉重复移动习惯的转移记忆视角,以及反映候选成熟度的生命周期视角。由于接近零的转移得分既可能表示不相关,也可能表示观测不足,EviRec利用每个候选的观测状态对证据进行评估,并通过可靠性门控机制在转移记忆证据与生命周期证据之间进行自适应路由。我们在一个涵盖全年、五个城市、包含超过30,000名用户和684,200条轨迹的兴趣点签到数据集上对EviRec进行了评估。实验结果表明,EviRec始终优于最先进的基线方法,且提升最大的部分集中于冷启动查询。特别地,在最强的基线之上,EviRec在双重新增(Dual-New)场景下将NDCG@10提升了20.4%。深入分析进一步证实,这些提升主要来自候选特定的可靠性门控机制,同时在很大程度上保留了先前学习到的移动习惯。
cs.LG / 83 / 2609.20333

Sharp Reconstruction Bounds for Autoencoders Using the Same Forward Map

使用相同前向映射的自编码器的尖锐重构界
Medina, Patricia, Lam, Hy P. G.
Abstract
We study reconstruction in autoencoders that apply the same forward map before and after setting the observed coordinates to zero. For equal odd input and hidden dimensions $d\geq 3$, among orientation-preserving diffeomorphisms whose Jacobian singular values lie in $[m,M]$, we show that the least uniform reconstruction-derivative error is $\max\{1-M(M-m)/2,0\}$, with affine maps attaining this sharp bound at every prescribed depth. A translated radial rotation can nevertheless reconstruct any prescribed ball exactly with singular values arbitrarily close to one, motivating additional conditions for a finite-data bound. We test this prediction on a 798,452-point terrestrial LiDAR forest scan. At input scale $0.05$, the mean theoretical bound is $0.155$, about $84\%$ of the mean normalized training error $0.185$ across four spatial regions, two depths, and three seeds. At this scale, adding one hidden coordinate reduces the mean reconstruction error below $6\times10^{-6}$.
Chinese Translation
我们研究了自编码器中的重构问题,该类自编码器在将观测坐标置零前后应用相同的前向映射。当输入与隐藏维度相等且为奇数 d≥3 时,在雅可比矩阵奇异值位于区间 [m,M] 内的保定向微分同胚中,我们证明了最小一致重构导数误差为 max{1-M(M-m)/2, 0},且仿射映射能够在任意给定深度处达到该尖锐界。然而,一个平移径向旋转可以精确重构任意给定的球体,且其奇异值可任意接近于一,这促使我们为有限数据界引入附加条件。我们在一个包含 798,452 个点的地面激光雷达(LiDAR)森林扫描数据上验证了这一预测。在输入尺度为 0.05 时,平均理论界为 0.155,约为四个空间区域、两个深度和三个随机种子下平均归一化训练误差 0.185 的 84%。在该尺度下,增加一个隐藏坐标可将平均重构误差降至 6×10^{-6} 以下。
cs.LG / 84 / 2609.20352

COMPASS: Ordered Clustered Routing at 100K Scale

COMPASS:10万规模的有序聚类路径规划
Greenberg, Ido, Linsenmaier, Hugo, Sielski, Piotr, Mannor, Shie, Fender, Alex, Chechik, Gal, Meirom, Eli
Abstract
Large-scale routing often requires visiting clusters of nodes in a prescribed order, giving rise to the Ordered Clustered Traveling Salesman Problem (OCTSP). Optimizing each cluster independently seems natural, but misses non-local dependencies. We introduce the COMPASS algorithm for OCTSP, which combines search with learning-accelerated routing by orchestrating parallel sub-solvers. COMPASS has no quality ceiling and its solutions keep improving with compute. It exploits the clustered structure, and can reach exact solutions in time exponential in cluster size rather than instance size. Empirically, COMPASS consistently outperforms alternative methods. Unlike common large-scale routing solvers, COMPASS consumes general distance matrices and is not limited to coordinate inputs. We demonstrate scaling to 100K synthetic nodes and to 28.5K real e-commerce nodes. To our knowledge, the latter is the largest reported routing solution over asymmetric distances, 9x beyond established ATSP benchmarks.
Chinese Translation
大规模路径规划通常需要按指定顺序访问节点簇,由此产生了有序聚类旅行商问题(Ordered Clustered Traveling Salesman Problem, OCTSP)。独立优化每个簇看似自然,但会忽略非局部依赖关系。我们提出了面向OCTSP的COMPASS算法,该算法通过协调并行子求解器,将搜索与学习加速的路径规划相结合。COMPASS没有质量上限,其解会随计算资源的增加持续改进。它利用聚类结构,可以在与簇规模而非实例规模呈指数关系的时间内达到精确解。实验表明,COMPASS始终优于其他替代方法。与常见的大规模路径规划求解器不同,COMPASS可以直接使用一般距离矩阵,不局限于坐标输入。我们在10万个合成节点和2.85万个真实电商节点上演示了其可扩展性。据我们所知,后者是已报道的基于非对称距离的最大规模路径规划求解结果,比现有的ATSP(非对称旅行商问题)基准高出9倍。
cs.LG / 85 / 2609.20353

Minimax-Optimal Online Contract Design with Unrestricted Bounded Contracts

具有无约束有界契约的极小极大最优在线契约设计
Ai, Rui, Simchi-Levi, David, Zhong, Han
Abstract
We study repeated contract design when a principal observes outcomes but not the actions that generate them. The principal may use any bounded outcome-contingent payment vector, and the agent's best response can make expected profit discontinuous in those payments. For every fixed number $m\ge2$ of outcomes, the minimax regret over $T$ rounds is of order $T^{m/(m+1)}$, up to logarithmic factors. The upper bound allows arbitrary action spaces and agent heterogeneity, without smoothness or monotone-surplus assumptions. Its key is an effective-dimension reduction that the benchmark can be normalized even when fixed tie-breaking is not shift invariant, after which revealed preference yields a monotone response map in payment-difference coordinates. A learning policy built on a Lipschitz parametrization of this map attains the rate using only observed outcome categories. The lower-bound construction accounts for how incentive losses accumulate across outcome dimensions. It shows that each additional contractible outcome creates a precise and unavoidable increase in the worst-case cost of learning.
Chinese Translation
我们研究了委托人只能观察到结果而无法观察到产生这些结果的行动的重复契约设计问题。委托人可以使用任意有界的结果依赖支付向量,而代理人的最优反应可能使其期望利润关于这些支付不连续。对于每一个固定的结果数 $m\ge2$,$T$ 轮上的极小极大遗憾(在不考虑对数因子的情况下)为 $T^{m/(m+1)}$ 阶。该上界允许任意的行动空间和代理人异质性,无需平滑性或单调剩余假设。其关键在于一种有效维数约简技术:即使固定的平局打破规则不具有平移不变性,基准仍可被规范化,此后显示偏好原理在支付差坐标下给出一个单调的反应映射。基于该映射的 Lipschitz 参数化构建的学习策略仅利用观察到的结果类别即可达到该速率。下界构造刻画了激励损失在各结果维度上的累积方式,表明每一个新增的可契约化结果都会导致最坏情况学习成本的精确且不可避免的增加。
cs.LG / 86 / 2609.20404

Learning Principal-Agent Contracts for Equitable Smallholder Carbon Farming under Moral Hazard and Adverse Selection

在道德风险与逆向选择下学习面向公平小农碳农业的委托-代理合约
Bharadwaj, Rishi, Narahari, Yadati
Abstract
Agricultural soils are a major untapped carbon sink. Carbon farming is emerging as a promising practice for tapping this potential. Smallholder farmers, who dominate agriculture across South Asia and sub-Saharan Africa, are key to scaling climate mitigation via carbon farming. It is ironic that real-world carbon programs largely fail to reach them. We study this important gap through the lens of contract design. An aggregator offers a single pooled contract to a heterogeneous population of smallholder farmers who have private adoption costs (adverse selection) and exert unobserved effort (moral hazard), with agronomic outcomes evolving over multiple seasons. We formulate this evolving contracting problem as a POMDP and use reinforcement learning to learn a dynamic profit-maximising contract. We analyse the performance of the aggregator under various conditions. We find that a profit-maximising aggregator does not merely inherit the exclusion of smallholders, it amplifies it. On large farms the aggregator realises 87.7% of achievable adoption, against only 8.2% on smallholdings. Per-hectare Measurement, Reporting and Verification (MRV) costs fall as farm size rises, and the aggregator's pooling contract compounds this gradient rather than offsetting it. A counterfactual that makes MRV costs purely area-proportional eliminates this disparity. Our results and simulation can guide contract and policy design that opens carbon income to smallholders while enabling agricultural soils to contribute to climate mitigation at scale.
Chinese Translation
农业土壤是一个尚未开发的重要碳汇,碳农业正成为挖掘这一潜力的有前景的做法。主导南亚和撒哈拉以南非洲农业的小农户是通过碳农业扩大气候减缓规模的关键。然而具有讽刺意味的是,现实中的碳项目大多未能惠及他们。我们通过合约设计的视角研究这一重要差距。聚合商向异质性的小农群体提供一份统一的 pooled 合约,这些农户拥有私有的采纳成本(逆向选择)并付出不可观测的努力(道德风险),且农艺结果随多个季节演化。我们将这一演化的合约问题建模为部分可观测马尔可夫决策过程(POMDP),并利用强化学习学习动态的利润最大化合约。我们分析了聚合商在各种条件下的表现。我们发现,利润最大化的聚合商不仅继承了现有项目对小农户的排斥,还放大了这种排斥。在大农场上,聚合商实现了可达采纳水平的87.7%,而在小农场上仅有8.2%。每公顷的测量、报告与核查(MRV)成本随农场规模增大而下降,而聚合商的统一合约加剧而非缓解了这一梯度差异。一个使MRV成本纯粹与面积成比例的反事实情形可消除这种差距。我们的结果与模拟可以为合约和政策设计提供指导,使小农户能够获得碳收入,同时让农业土壤能够大规模地为气候减缓作出贡献。
cs.LG / 87 / 2609.20409

The Bias of Nonlinear Two-Time-scale Stochastic Approximation under Constant Step-Sizes

常数步长下非线性两时间尺度随机逼近的偏差
Lamouri, Djamel Rassem, Baudry, Dorian, Gast, Nicolas
Abstract
Two-timescale stochastic approximation (TTSA) is a fundamental tool for analyzing coupled iterative algorithms in reinforcement learning, optimization, and stochastic control. However, finite-time guarantees for nonlinear two-timescale schemes remain difficult to obtain, especially under constant step-sizes. In this paper, we study nonlinear TTSA with step-sizes $\alpha\gg\beta$. Under standard stability, regularity, and Markovian noise assumptions, we upper bound the mean-squared error and the bias of both iterates around their limiting equilibria. Our bounds scale as $O(\alpha+\beta^2/\alpha^2)$, which we prove to be tight when $\beta\le\alpha^{3/2}$. The analysis separates the contributions of initial conditions, fast-timescale tracking error, Markovian dependence, and timescale coupling, thereby clarifying the origin of the $\beta^2/\alpha^2$ term. Our results reveal qualitative differences from the linear TTSA setting previously studied, showing that nonlinear dynamics introduce additional finite-time effects that are absent in the linear case.
Chinese Translation
两时间尺度随机逼近(Two-Timescale Stochastic Approximation, TTSA)是分析强化学习、优化和随机控制中耦合迭代算法的基本工具。然而,非线性两时间尺度方案的有限时间保证仍然难以获得,尤其是在常数步长的情形下。本文研究步长满足 $\alpha\gg\beta$ 的非线性TTSA。在标准的稳定性、正则性和马尔可夫噪声假设下,我们对两条迭代序列围绕其极限平衡点的均方误差和偏差给出了上界。我们的界为 $O(\alpha+\beta^2/\alpha^2)$ 量级,并且我们证明了当 $\beta\le\alpha^{3/2}$ 时该界是紧的。该分析将初始条件、快时间尺度跟踪误差、马尔可夫依赖性以及时间尺度耦合的贡献分离开来,从而阐明了 $\beta^2/\alpha^2$ 项的来源。我们的结果揭示了与先前研究的线性TTSA情形在性质上的差异,表明非线性动力学引入了线性情形中不存在的额外有限时间效应。
cs.LG / 88 / 2609.20419

SCGFM-ART: Amortized Relational Transport for Structure-Centric Graph Foundation Models

SCGFM-ART:面向结构中心的图基础模型的摊销关系传输方法
He, Xiaodong, Wang, Xincheng, Kang, Zhao
Abstract
Graph foundation models (GFMs) aim to learn transferable representations across severely heterogeneous graph domains. However, severe domain shifts in topology, graph scale, and feature semantics impede the construction of a unified, domain-agnostic representation space. To address this, we propose SCGFM-ART, a structure-centric GFM framework that aligns arbitrary graphs onto a shared relational atlas via Amortized Relational Transport (ART). The relational atlas serves as a universal coordinate system defined by a finite set of relational landmarks (bases), while ART directly predicts reusable, end-to-end graph-to-base transport plans, bypassing costly runtime Gromov-Wasserstein optimizations. Under this formulation, SCGFM-ART decomposes a graph into a unified representation: globally via its relational response coordinates relative to the atlas, and locally via its node-to-role structural correspondences. These correspondences project disparate node attributes into a canonical role space, resolving structural and semantic heterogeneity within a singular alignment interface. Rigorously modeling graphs and atlas bases as finite measured relational spaces, we establish coordinate fidelity bounds, prove stability under predicted transport plans, and derive an amortized coverage bound that guarantees our learning objective tightly surrogates ideal relational coverage. Benchmarked across 14 cross-domain graph- and node-level classification tasks, SCGFM-ART achieves state-of-the-art transferability, securing superior average ranks of 2.29 and 1.14, respectively. Topological perturbation analyses demonstrate that node-role transport retains fine-grained structural nuances beyond global coordinates. On real-world benchmarks, the amortized formulation yields 44.2 to 85.1 times faster frozen target-domain inference by avoiding iterative alignment at test time.
Chinese Translation
图基础模型(Graph Foundation Models, GFMs)旨在跨严重异构的图域学习可迁移的表示。然而,拓扑结构、图规模和特征语义方面的剧烈域偏移阻碍了统一的、与领域无关的表示空间的构建。为解决这一问题,我们提出了SCGFM-ART,一种以结构为中心的GFM框架,通过摊销关系传输(Amortized Relational Transport, ART)将任意图对齐到一个共享的关系图谱上。该关系图谱作为一个由有限的关系基(基向量)定义的通用坐标系,而ART则直接预测可重用的、端到端的图到基的传输计划,从而避免了代价高昂的运行时Gromov-Wasserstein优化。在此公式化框架下,SCGFM-ART将图分解为统一的表示:在全局层面,通过其相对于图谱的关系响应坐标;在局部层面,通过其节点到角色的结构对应关系。这些对应关系将不同的节点属性投影到一个规范的角色空间中,在单一的对齐接口内解决了结构与语义的异构性。我们将图与图谱基严格建模为有限测度的关系空间,建立了坐标保真度界,证明了在预测传输计划下的稳定性,并推导出摊销覆盖界,保证我们的学习目标能够紧密逼近理想的关系覆盖。在14个跨域的图级和节点级分类任务上进行基准测试,SCGFM-ART实现了最先进的可迁移性,平均排名分别达到优异的2.29和1.14。拓扑扰动分析表明,节点-角色传输在全局坐标之外保留了细粒度的结构细节。在真实世界基准上,摊销公式通过避免测试时的迭代对齐,使冻结的目标域推理速度提升44.2至85.1倍。
cs.LG / 89 / 2609.20451

Seismic Site Response Prediction from Sparse Observations Using Finite-Element-Pretrained Latent Dynamics

基于有限元预训练潜在动力学的稀疏观测场地地震反应预测
Zhu, Yi, Chen, Su, Li, Xiaojun
Abstract
Numerical site-response predictions often deviate from observations, yet correcting these discrepancies is difficult because records are limited in both sensor coverage and number of events. This study proposes the Transfer-Enabled Forced Latent Autoencoder for Response Equations (FLARE-T) to improve these predictions by learning and calibrating low-dimensional latent dynamics that connect the base acceleration input to acceleration outputs at multiple depths. FLARE-T learns a low-dimensional response manifold and input-driven dynamics from dense finite-element simulations. It then trains a sparse encoder to map simulated sensor responses into the learned coordinates and uses limited records to calibrate the dynamics within them. A short response window initializes each prediction, while the complete base motion drives the response. The framework was evaluated using a layered-soil centrifuge test and the Lotung field vertical array. Test-set results show that FLARE-T improved multi-depth acceleration histories and 5%-damped pseudoacceleration response spectra relative to the original finite-element models, reducing errors at every evaluated sensor for motions of different intensities and, at Lotung, for both horizontal components. Two Lotung source models with different constitutive parameters achieved comparable test-set accuracy, indicating reduced dependence on precise prior calibration. FLARE-T therefore provides a data-efficient means of combining dense numerical response information with limited field records to improve future site-response predictions.
Chinese Translation
数值场地地震反应预测结果常与实测记录存在偏差,但由于传感器覆盖范围和地震事件数量均有限,校正这些偏差十分困难。本研究提出了一种响应方程转移使能强迫潜在自编码器(Transfer-Enabled Forced Latent Autoencoder for Response Equations,FLARE-T),通过学习并校准连接基底加速度输入与多个深度处加速度输出的低维潜在动力学,来改进数值预测。FLARE-T首先从稠密的有限元模拟中学习低维响应流形及输入驱动的动力学;随后训练一个稀疏编码器,将模拟的传感器响应映射到所学坐标系中,并利用有限的实测记录校准其中的动力学。预测时仅需一个较短的响应窗口进行初始化,而完整的基底运动则驱动整个响应过程。该框架通过分层土体离心机试验和Lotung场地垂直阵列数据进行了评估。测试集结果表明,与原始有限元模型相比,FLARE-T改进了多深度加速度时程和5%阻尼比的伪加速度反应谱,在不同强度地震动下的每个评估传感器处均降低了误差,在Lotung场地对两个水平分量同样有效。采用不同本构参数的两个Lotung震源模型获得了相当的测试集精度,表明该方法降低了对精确先验校准的依赖。因此,FLARE-T提供了一种数据高效的方式,将稠密的数值响应信息与有限的现场记录相结合,以改进未来的场地地震反应预测。
cs.LG / 90 / 2609.20465

Training Neural Networks to Approach the Optimum Bayes Estimator in Dense Multi-Emitter Localization

训练神经网络以逼近密集多发射源定位中的最优贝叶斯估计器
Sun, Yi, Sharifi, Mona, Yumman, Muzna
Abstract
We train neural networks on synthesized frames to approach the optimum Bayes estimator for dense emitter localization. The result justifies the future work on training neural networks to achieve high-throughput large-FOV super spatiotemporal resolution SMLM.
Chinese Translation
我们在合成图像帧上训练神经网络,以逼近密集发射源定位中的最优贝叶斯估计器。该结果为未来训练神经网络以实现高吞吐量、大视场、超高时空分辨率的单分子定位显微镜(SMLM)的研究奠定了基础。
cs.LG / 91 / 2609.20467

Deep Learning-Based Classification of Cognitive and Resting States Using Electroencephalography Signals

基于深度学习的脑电信号认知状态与静息状态分类
Fernando, K. A. Januka S., Srivastava, Harshit
Abstract
The categorization of cognitive and resting states derived from electroencephalography (EEG) signals is crucial for comprehending fluctuations in brain activity linked to various mental states. EEG provides a non-intrusive approach for documenting brain function in both resting and task-oriented cognitive conditions, whilst deep learning techniques enable the automatic extraction of significant patterns from intricate EEG data. This study presents a deep learning framework to distinguish between resting and cognitive states through EEG records. The proposed framework integrates a Convolutional Neural Network (CNN) stacked with a Gated Recurrent Unit (GRU) for the extraction of features from EEG signals. Time-frequency analysis is conducted to explore the salient aspects of signals, and the derived features are then assessed utilizing conventional deep learning and machine learning classifiers, including the suggested 2D-Net architecture. The proposed approach and feature extraction strategy outperform the evaluated comparative methods, achieving accuracies of 83.177% for resting-versus-mathematical task classification, 76.107% for resting-versus-memory task classification, and 83.432% for resting-versus-music task classification. The findings illustrate the efficacy of integrating signal processing with deep learning methodologies to discriminate resting from cognitive states utilizing EEG signals.
Chinese Translation
基于脑电图(EEG)信号的认知状态与静息状态分类,对于理解与各种心理状态相关的大脑活动波动至关重要。EEG提供了一种非侵入式的方法,可在静息和任务导向的认知条件下记录大脑功能,而深度学习技术则能够从复杂的EEG数据中自动提取有意义的模式。本研究提出了一种深度学习框架,通过EEG记录区分静息状态与认知状态。该框架将卷积神经网络(CNN)与门控循环单元(GRU)堆叠,用于从EEG信号中提取特征。研究进行了时频分析以探索信号的显著特征,并利用传统的深度学习和机器学习分类器(包括所提出的2D-Net架构)对所提取的特征进行评估。所提出的方法和特征提取策略优于所评估的对比方法,在静息与数学任务分类中达到83.177%的准确率,在静息与记忆任务分类中达到76.107%的准确率,在静息与音乐任务分类中达到83.432%的准确率。研究结果证明了将信号处理与深度学习方法相结合,利用EEG信号区分静息状态与认知状态的有效性。
cs.LG / 92 / 2609.20501

Distributionally Robust Federated Learning with Multi-Source Data

基于多源数据的分布鲁棒联邦学习
Liu, Yingzhu, Li, Zhongkui, You, Pengcheng, Cherukuri, Ashish
Abstract
Federated learning trains a shared model from private client data. In practice, data-generating distributions may differ, and the true mixture across clients is often unknown, making the underlying group distribution difficult to specify. Existing approaches address cross-client mixture uncertainty by optimizing against the worst-case mixture, yet assume accurate client-wise distribution estimates. However, these estimates can be unreliable when based on finite samples. To handle both cross-client mixture uncertainty and within-client distributional ambiguity, we construct a global ambiguity set as the union of admissible mixtures of local ambiguity sets. The construction allows client-specific ambiguity radii and admits a client-wise separable reformulation. Leveraging this structure, we establish a high-probability out-of-sample performance guarantee. We further develop a federated algorithm for a penalty-based reformulation and prove its convergence under milder regularity conditions. Simulations validate the algorithm's effectiveness.
Chinese Translation
联邦学习利用客户端的私有数据训练一个共享模型。在实际应用中,各客户端的数据生成分布可能不同,且客户端之间的真实混合比例通常是未知的,这使得底层的组分布难以确定。现有方法通过针对最坏情况下的混合分布进行优化来处理客户端间混合比例的不确定性,但其假设各客户端的分布估计是准确的。然而,基于有限样本的分布估计可能是不可靠的。为同时应对客户端间混合比例的不确定性与客户端内部的分布模糊性,我们构建了一个全局模糊集,即由各局部模糊集的可容许混合所构成的并集。该构建方式允许各客户端具有特定的模糊半径,并可转化为客户端可分离的重构形式。利用这一结构,我们建立了高概率的样本外性能保证。我们进一步针对基于惩罚项的重构形式开发了联邦学习算法,并在较温和的正则性条件下证明了其收敛性。仿真实验验证了该算法的有效性。
cs.LG / 93 / 2609.20507

Radio Frequency Detection and Classification of Microplastics in Water

水中微塑料的射频检测与分类
Tolbert, Jaden, Islam, Md Saiful, Wang, Pingshan
Abstract
Micro- and nano-plastic particles (MPs/NPs) are ubiquitous environmental contaminants whose increasing abundance and potential health impacts have created an urgent need for rapid, label-free detection methods. As particle size decreases to the low-micrometer range, conventional optical and spectroscopic techniques become increasingly challenging because of limited throughput and/or complex sample preparation. In this work, we present a machine learning (ML)-assisted radio-frequency (RF) dielectric spectroscopic cytometry (DiSC) platform for the label-free detection and classification of MPs. Eight types of $ 10 $ {\mu}m nominal-diameter MP particles suspended in deionized (DI) water were characterized at four frequencies spanning $ 0.2\text{-}9\text{ GHz} $. The measured alterations in RF scattering parameters (S-parameters), referenced to the carrier medium, were used to train supervised ML models for material classification, including the identification of MPs in mixed samples and saline-water environments. For eight MP classes suspended in DI water, the proposed method achieved macro-average F1-score, precision, and recall values exceeding $ 0.71 $. Furthermore, PET classification performance was largely maintained in saline carrier media containing $3.3\% $ and $ 6.6\% $ sea salt. These results demonstrate the feasibility of ML-assisted RF DiSC for rapid, single-particle MP classification in aqueous environments. Future work will focus on improving classification performance through enhanced RF calibration, increased spectral coverage, larger training datasets, and validation using environmentally aged and biologically contaminated microplastics.
Chinese Translation
微塑料和纳米塑料颗粒(MPs/NPs)是普遍存在的环境污染物,其丰度不断增加且对健康具有潜在影响,因此迫切需要快速、无标记的检测方法。当颗粒尺寸减小至低微米范围时,传统光学和光谱技术由于通量有限和/或样品制备复杂而面临越来越大的挑战。本工作提出了一种机器学习(ML)辅助的射频(RF)介电谱流式细胞术(DiSC)平台,用于微塑料的无标记检测与分类。我们表征了悬浮于去离子(DI)水中、标称直径为10 μm的八种微塑料颗粒,测量频率覆盖0.2-9 GHz范围内的四个频率点。以载液介质为参考,将测得的射频散射参数(S参数)变化用于训练监督式机器学习模型,实现材料分类,包括混合样品和盐水环境中微塑料的识别。对于悬浮于去离子水中的八类微塑料,所提方法的宏平均F1分数、精确率和召回率均超过0.71。此外,PET(聚对苯二甲酸乙二醇酯)的分类性能在含有3.3%和6.6%海盐的盐水载液介质中基本得以保持。这些结果证明了ML辅助射频DiSC技术用于水环境中快速单颗粒微塑料分类的可行性。未来的工作将致力于通过改进射频校准、扩大频谱覆盖范围、增大训练数据集以及使用环境老化和生物污染的微塑料进行验证,来进一步提升分类性能。
cs.LG / 94 / 2609.20511

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

当EOS标记不一致时:理解在线策略蒸馏中的长度膨胀现象
Yang, Yuxiao, Yu, Tianrun, Li, Shangzhe, Zhao, Kaixiang, Zhang, Xuchao, Bansal, Chetan, Yao, Huaxiu, Killian, Taylor W., Zhang, Weitong
Abstract
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify \emph{termination-token mismatch} between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, the two models can place their stopping probability on different EOS tokens, even when their declared stopping sets are identical. This mismatch can suppress the student's preferred termination action without reliably transferring the teacher-preferred alternative. We show that aligning the decoding stopping set alone is insufficient, while treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigates mismatch-induced length inflation across all three model families. To further understand how termination behavior evolves over training, we study OPD across different K2-Horizon training stages. This stage-wise analysis shows that termination preferences can shift substantially during training, while also revealing a distinct length inflation late in the OPD run that persists beyond termination alignment. Together, these results identify termination mismatch as an important, but not exhaustive, source of OPD length dynamics. We release an implementation incorporating the proposed termination-handling corrections.
Chinese Translation
我们研究了在线策略蒸馏(On-Policy Distillation, OPD)中的长度膨胀问题,即学生模型的回复可能变得过长,甚至会耗尽生成预算。我们发现基础学生模型与后训练教师模型之间的终止标记(termination-token)不匹配是导致这一行为的重要来源。在Qwen3、Llama和Gemma三个模型家族中,即使两个模型声明的停止集合完全相同,它们也可能将停止概率置于不同的EOS标记上。这种不匹配可能会抑制学生模型偏好的终止动作,同时又无法可靠地迁移教师模型偏好的替代动作。我们证明,仅对齐解码停止集合是不够的,而将功能等价的EOS标记视为共享的语义停止动作,能够显著缓解所有三个模型家族中由不匹配引起的长度膨胀。为了进一步理解终止行为在训练过程中的演变,我们研究了不同K2-Horizon训练阶段的OPD表现。这一阶段性分析表明,终止偏好在训练过程中可能发生显著变化,同时还揭示了OPD运行后期出现的一种独特的长度膨胀现象,该现象在终止对齐之后仍然持续存在。综上,这些结果将终止不匹配确定为OPD长度动态的一个重要但非唯一的来源。我们发布了包含所提出的终止处理修正的实现代码。
cs.LG / 95 / 2609.20539

Parallelism, critical windows, and separations among diffusion language models

扩散语言模型中的并行性、临界窗口与相互分离
Chen, Sitan, Wang, Liye
Abstract
A popular selling point of diffusion large language models (dLLMs) is their capacity for parallelism: the ability to generate sequences of text far more efficiently than autoregressive models, which require one forward pass per token. Yet among the many competing paradigms for dLLMs, from masked to uniform to Gaussian diffusion, principled understanding of how these different proposals compare in parallelism remains limited. In this work, we initiate a fine-grained comparison of the capacity for parallelism among these three leading approaches and prove the following: - Uniform and Gaussian diffusion can sample in a number of forward passes which scales with the dual total correlation of the underlying distribution, a measure of intrinsic complexity which can be much smaller than the context length. Previously, it was only known how to achieve this using masked diffusion. - For a certain family of random empirical measures, we show that $\widetilde{\Theta}(\sqrt{d})$ forward passes are necessary and sufficient to sample using uniform or Gaussian diffusion, yet there exist approximate score oracles for which $\widetilde{\Omega}(d)$ forward passes are needed for masked diffusion. This establishes the first provable separation in parallelism between the three prevailing dLLM paradigms. Contrary to popular intuition that masked diffusions are harder to parallelize because they must commit to token values, the latter separation instead comes from the fact that the critical windows in masked diffusion sampling are asymptotically narrower than those in uniform and Gaussian diffusion sampling.
Chinese Translation
扩散大语言模型(dLLM)的一个广受欢迎的卖点是其并行能力:即能够以远高于自回归模型的效率生成文本序列,而自回归模型每个词元都需要一次前向传播。然而,在从掩码扩散到均匀扩散再到高斯扩散的众多竞争性dLLM范式中,对于这些不同方案在并行性方面如何比较,原则性的理解仍然有限。在本工作中,我们对这三种主流方法的并行能力进行了细粒度比较的开创性研究,并证明了以下结论: - 均匀扩散和高斯扩散可以在与前向传播次数按底层分布的对偶总相关性(dual total correlation)规模扩展的次数内完成采样。该对偶总相关性是一种内在复杂度度量,可以远小于上下文长度。此前,仅已知掩码扩散可以实现这一点。 - 对于某一类随机的经验测度族,我们证明使用均匀扩散或高斯扩散进行采样,$\widetilde{\Theta}(\sqrt{d})$ 次前向传播是必要且充分的;然而对于某些近似分数预言机(score oracle),掩码扩散需要 $\widetilde{\Omega}(d)$ 次前向传播。这建立了三种主流dLLM范式之间在并行性上的首个可证明分离。 与流行的直觉——即掩码扩散因其必须确定词元的取值而更难并行化——相反,后一种分离实际上源于掩码扩散采样中的临界窗口在渐近意义上比均匀扩散和高斯扩散采样中的临界窗口更窄这一事实。
cs.LG / 96 / 2609.20548

Mitigating Retaliatory Algorithmic Collusion in Repeated Games

缓解重复博弈中的报复性算法合谋
Sivachandran, Karthik, Paleja, Rohan
Abstract
Reinforcement learning agents trained to maximize their own reward in repeated interactions can converge to supra-competitive outcomes resembling explicit collusion, without communication or shared design. Existing mitigation approaches are largely tied to specific economic settings, like two-sided platforms and auctions, leaving open how to design interventions for general repeated games. We address this gap by formalizing the connection between empirical observations from prior work on Q-learning collusion and classical theory of Simple Penal Codes (SPCs). We show any non-trivial SPC induces a quantifiable conditional dependence in agents' policies, detectable via the total variation distance between an agent's action distributions across cooperation and defection histories. Building on this connection, we propose CURB (Collusion Unwinding via Reward shaping and Belief injection), a reward-shaping framework that penalizes this Total Variation (TV) distance signal during Q-learning and is guaranteed to convert any SPC fixed point of the dynamics into a trivial one, thus precluding collusive equilibria sustained by punishment threats. Empirically, CURB substantially reduces collusion by Q-learning agents in both Bertrand and Cournot Competition Repeated Games. We further demonstrate that CURB extends to deep Q-network agents in Bertrand competition, suggesting the mechanism generalizes beyond tabular Q-learning.
Chinese Translation
在重复交互中训练以最大化自身奖励的强化学习智能体,可以在没有通信或共享设计的情况下收敛到类似明示合谋的超竞争结果。现有的缓解方法大多依赖于特定的经济场景,如双边平台和拍卖,因而如何为一般重复博弈设计干预手段仍是开放问题。我们通过形式化先前工作中关于Q学习合谋的实证观察与经典简单惩罚码(Simple Penal Codes, SPCs)理论之间的联系来填补这一空白。我们证明,任何非平凡的SPC都会在智能体的策略中引入一种可量化的条件依赖性,该依赖性可通过智能体在不同合作与背叛历史下动作分布之间的总变差距离(total variation distance)检测出来。基于这一联系,我们提出了CURB(Collusion Unwinding via Reward shaping and Belief injection,通过奖励塑形与信念注入实现合谋瓦解),这是一个奖励塑形框架,在Q学习过程中对该总变差(TV)距离信号进行惩罚,并保证将动力学中任何SPC不动点转化为平凡不动点,从而排除由惩罚威胁所维持的合谋均衡。实验表明,CURB在Bertrand和Cournot重复竞争博弈中均能显著减少Q学习智能体的合谋行为。我们进一步证明CURB可扩展至Bertrand竞争中的深度Q网络(DQN)智能体,表明该机制可推广到表格型Q学习之外。
cs.LG / 97 / 2609.20592

CrystalMO-TuRBO: Multi-Objective Trust-Region Bayesian Optimization for High-precision Joint Crystal Structure Refinement

CrystalMO-TuRBO:用于高精度晶体结构联合精修的多目标信赖域贝叶斯优化
Agada, Joseph, Wang, Yishu, Biswas, Arpan
Abstract
Crystal structure refinement is a fundamental inverse problem in materials characterization, where structural parameters are optimized to reproduce experimental diffraction data. Conventional approaches, such as least-squares and likelihood-based optimization, rely on local search and often struggle with non-convex, noisy, and highly correlated parameter landscapes, particularly when integrating multiple diffraction modalities. Joint refinement of X-ray and neutron data is especially challenging due to their complementary but competing sensitivities, which are typically combined through scalarized objectives requiring manual weighting and leading to suboptimal solutions. We propose CrystalMO-TuRBO, a multi-objective trust region Bayesian optimization architecture for joint crystal structure refinement. The method models X-ray and neutron discrepancies as separate objectives and transforms the problem into a normalized maximization setting. A two-phase optimization strategy is introduced: Phase 1 performs global exploration using parallel trust-region Bayesian optimization across multiple scalarizations to identify promising regions of the parameter space, while Phase 2 conducts localized refinement within a shrinking region to achieve high-precision solutions. This design explicitly separates global search from fine-grained optimization, addressing the unique accuracy requirements of refinement tasks. We evaluate the proposed method on experimentally collected X-ray and neutron diffraction data from single-crystal Ho2Ti2O7. Results demonstrate improved convergence, robustness, and parameter precision compared to classical refinement methods and Bayesian optimization baselines on refinement of a single-crystal pyrochlore material system.
Chinese Translation
晶体结构精修是材料表征中的一个基本逆问题,其目标是通过优化结构参数以再现实验衍射数据。传统方法(如最小二乘法和基于似然的优化)依赖于局部搜索,在处理非凸、含噪声且参数高度关联的参数空间时常常表现不佳,尤其是在融合多种衍射模态时。由于X射线和中子数据具有互补但相互竞争的敏感性,其联合精修尤其具有挑战性;通常需要通过标量化目标进行组合,这要求人工设定权重,容易导致次优解。我们提出CrystalMO-TuRBO,一种用于晶体结构联合精修的多目标信赖域贝叶斯优化架构。该方法将X射线和中子数据的偏差建模为独立目标,并将问题转化为归一化的最大化形式。我们引入了一种两阶段优化策略:第一阶段利用并行信赖域贝叶斯优化在多个标量化任务上进行全局探索,以识别参数空间中有希望的区域;第二阶段在不断收缩的区域内进行局部精修,以获得高精度解。这一设计明确地将全局搜索与精细优化分离开来,从而满足精修任务对精度的独特要求。我们在实验采集的单晶Ho2Ti2O7的X射线和中子衍射数据上评估了所提出的方法。结果表明,与经典精修方法和贝叶斯优化基线相比,该方法在单晶烧绿石材料体系精修中展现出更优的收敛性、鲁棒性和参数精度。
cs.LG / 98 / 2609.20594

Recursive Quantum Long Short-Term Memory for Stable Short-Horizon Temperature Forecasting

用于稳定短时温度预测的递归量子长短期记忆网络
Lee, Mu-En, Liu, Yen-Ku, Chen, Samuel Yen-Chi, Tsai, Yun-Cheng
Abstract
Quantum long short-term memory (QLSTM) models extend recurrent sequence learning with variational quantum circuits, but their optimization behavior can vary substantially across random initializations and temporal contexts. This paper evaluates a recursive QLSTM architecture against a standard QLSTM for one-step-ahead prediction of daily minimum and maximum temperature. Using daily weather observations from Toronto and identical training settings, we compare convergence, predictive accuracy, and generalization across input windows of 8, 16, and 32 days over 20 random seeds. The recursive model consistently reaches a near-optimal test loss earlier, reduces mean absolute error and root mean squared error, and exhibits a smaller generalization gap. These results indicate that recursive quantum feature transformations can improve stability and out-of-sample performance for compact hybrid quantum--classical temporal models.
Chinese Translation
量子长短期记忆网络(QLSTM)通过变分量子电路扩展了循环序列学习,但其优化行为在不同随机初始化和时间情境下可能有显著差异。本文针对每日最低和最高温度的单步预测任务,评估了一种递归QLSTM架构与标准QLSTM的对比表现。基于多伦多的每日气象观测数据以及完全相同的训练设置,我们在8天、16天和32天的输入窗口下,通过20个随机种子比较了二者的收敛性、预测精度和泛化能力。递归模型始终能更早达到接近最优的测试损失,降低了平均绝对误差和均方根误差,并表现出更小的泛化差距。这些结果表明,递归量子特征变换能够提升紧凑型混合量子-经典时序模型的稳定性与样本外性能。
cs.LG / 99 / 2609.20598

COIN-GP: Cooperative Online Learning in Networked Distributed Systems with Partial Measurements via Gaussian Process Regression

COIN-GP:基于高斯过程回归的网络化分布式系统中具有部分测量的协作在线学习
Yang, Zewen, Dai, Xiaobing, Yin, Zhenxiao, Zhao, Hang, Li, Zhijun, Chan, C. C.
Abstract
In this paper, we tackle the problem of jointly estimating the system states and partially unknown dynamics within distributed sensor-equipped networks, particularly in scenarios where only partial state observations are available. To address this issue, we propose an observer-based dynamic cooperative learning framework incorporating online distributed Gaussian Process (GP) regression, which enables accurate estimation despite incomplete in measurements and deficient GP models. In addition, a novel data collection strategy is introduced, with theoretical conditions ensuring feasible data acquisition. Moreover, we also derive an error upper bound encompassing state estimation and model estimation, leveraging the deterministic error bounds of GPs. Empirical simulations demonstrate the superiority of our approach compared to existing distributed GP-based methods.
Chinese Translation
本文研究在分布式传感器网络中联合估计系统状态与部分未知动力学的问题,特别是在仅有部分状态观测可用的场景下。为解决该问题,我们提出了一种基于观测器的动态协作学习框架,该框架结合了在线分布式高斯过程(GP)回归,能够在测量不完整和GP模型有缺陷的情况下实现准确估计。此外,我们还引入了一种新颖的数据采集策略,并给出了确保数据采集可行的理论条件。同时,利用GP的确定性误差界,我们推导了涵盖状态估计与模型估计的误差上界。仿真实验表明,与现有的基于分布式GP的方法相比,我们的方法具有优越性。
cs.LG / 100 / 2609.20650

Multi-center Medical Data Mining with FL-Net - A One-stop Shop for Federated Learning

基于 FL-Net 的多中心医学数据挖掘——一站式联邦学习平台
Süwer, Simon, Klemm, Julian, Acitelli, Elisa, Almeida, Mathieu, Altucci, Lucia, Bagyura, Zsolt, Barbieri, Michelangela, Bedő, Zsolt-Zoltán, Benedetti, Rosaria, Bihari, Béla, Csalóka, Csongor, Dicunta, Lucia, Ehrlich, Stanislav, Eskofier, Bjoern M., Fejér, Sándor-József, Fröwis, Georg, Hötzendorfer, Walter, Kautzky-Willer, Alexandra, Lohmann, Jens Johann Georg, Maranghi, Marianna, Marconi, Lorenzo, Mayer, Rudolf, Megchelenbrink, Wouter Leonard, Moga, Monika, Mottalib, Adham, Mehta, Sanjeev, Müller, Madeleine, Nyström, Thomas, Orbán, Balázs-Attila, O'Toole, Paul, Paolisso, Giuseppe, Parini, Paolo, Pedrelli, Matteo, Petrillo, Enrico, Poindl, Philipp, Probul, Niklas, Pustozerova, Anastasia, Šarčević, Tanja, Weilguny, Lukas, Baumbach, Jan, Maier, Andreas
Abstract
Federated learning enables collaborative training without sharing patient-level data, but most studies remain simulations. Based on five requirements derived from the literature, we analyzed 14 FL frameworks and found that none fully satisfied these requirements. We present FL-Net, a novel federated clinical research framework to fulfill all requirements. It integrates modular data harmonization, data discovery, disclosure control, securely built versioned FL-Net-Tools and containerized federated workflow execution into a persistent network. It enables the re-use of harmonized data and workflows across studies. FL-Net's end-to-end capabilities were evaluated through harmonization, cross-study patient discovery across MIMIC and US-130, and reproducible, audited federated workflows with up to 50 concurrent clients. FL-Net is being developed within the dAIbetes and Microb-AI-ome EU projects and will cover over 800,000 patients across 10 hospitals in 9 countries covering longitudinal and single point in time data, FL-Net provides a practical foundation for interoperable, reproducible, and privacy-preserving multicenter clinical research.
Chinese Translation
联邦学习使协同训练成为可能,而无需共享患者级别的数据,但大多数研究仍停留在仿真阶段。基于从文献中得出的五项需求,我们分析了14个联邦学习框架,发现没有一个框架能完全满足这些需求。我们提出了 FL-Net,一个能够满足全部需求的新型联邦临床研究框架。它将模块化的数据标准化、数据发现、披露控制、安全构建的版本化 FL-Net-Tools 以及容器化的联邦工作流执行集成到一个持久化网络中。它支持跨研究复用标准化后的数据和工作流。通过数据标准化、跨 MIMIC 与 US-130 数据集的跨研究患者发现,以及支持多达50个并发客户端的可复现、经审计的联邦工作流,我们对 FL-Net 的端到端能力进行了评估。FL-Net 正在欧盟 dAIbetes 和 Microb-AI-ome 项目中开发,将覆盖9个国家10家医院的超过80万名患者,涵盖纵向数据与单时点数据。FL-Net 为可实现互操作、可复现且保护隐私的多中心临床研究提供了实用的基础。
cs.LG / 101 / 2609.20676

Epidemiological Causal Graph Identification: Challenges, Identifiability and Algorithms

流行病学因果图识别:挑战、可识别性与算法
Mishra, Sambit, Wang, Yingying, Johnson, Christine K., Mitra, Urbashi
Abstract
Causal discovery from observational data is fundamental to statistics and machine learning, yet determining causal direction without interventions necessitates structural assumptions. Existing identifiability research primarily focuses on continuous variables under additive noise models, often neglecting mixed datasets containing ordinal scales, counts, and continuous measurements. This paper investigates causal discovery in Directed Acyclic Graphs (DAGs) where nodes follow either an ordinal distribution (via an ordered logit model) or a regular one-parameter exponential family distribution. We prove that the edge direction between an ordinal and an exponential family node is distributionally identifiable for generic parameter values. Our findings generalize previous Ordinal-Poisson results to the broader exponential family. Computationally, we introduce a score-based exhaustive search and a masked continuous optimization framework using DAGMA for larger graphs. Numerical results validate the theory, recovering edge orientations within a Markov equivalence class that are unidentifiable under classical structural equation models.
Chinese Translation
从观测数据进行因果发现是统计学和机器学习的基础,然而在不进行干预的情况下确定因果方向需要结构性假设。现有的可识别性研究主要集中于加性噪声模型下的连续变量,往往忽略了包含有序尺度、计数和连续测量的混合数据集。本文研究了有向无环图(DAG)中的因果发现问题,其中节点服从有序分布(通过有序logit模型)或常规的单参数指数族分布。我们证明,对于一般参数值,有序节点与指数族节点之间的边方向是分布可识别的。我们的发现将先前关于有序-泊松(Ordinal-Poisson)的结果推广到更广泛的指数族。在计算方面,我们提出了一种基于评分的穷举搜索方法,以及一种利用DAGMA的掩码连续优化框架以处理更大规模的图。数值结果验证了理论,恢复了马尔可夫等价类中在经典结构方程模型下无法识别的边方向。
cs.LG / 102 / 2609.20677

RISC-V and machine learning: a survey

RISC-V与机器学习:综述
Keshri, Shriman, Singh, Apparna, Palo, Chinmaya Kumar, Adya, Shreya, Mishra, Subhankar
Abstract
The intersection of open-source processor architectures and machine learning is driving the demand for customizable, efficient, and accessible hardware. This survey examines the state of the RISC-V ISA in machine learning applications, analyzing current capabilities, challenges, and future directions based on recent research. The analysis covers academic and commercial implementations, software frameworks, and real-world applications. The RISC-V machine learning ecosystem is evaluated, from instruction set extensions and core implementations to compiler optimizations and deployment strategies. Key contributions include a unified taxonomy of RISC-V ML implementations, a comparative analysis of performance and design trade-offs, an evaluation of software toolchain maturity, and the identification of emerging trends in instruction set extensions and specialized accelerators. Findings reveal progress in energy efficiency, specialized instruction development, and framework integration, while highlighting challenges in standardization, verification complexity, and ecosystem fragmentation. The analysis proposes four research directions to address current limitations: specialized neural processing extensions, adaptive and modular processor architectures, security frameworks, and energy-efficient multi-domain architectures. These directions provide a roadmap for advancing RISC-V as a foundational platform for next-generation machine learning systems.
Chinese Translation
开源处理器架构与机器学习的交叉融合正推动着对可定制、高效且易于获取的硬件的需求。本综述考察了RISC-V指令集架构(ISA)在机器学习应用中的现状,基于近期研究分析了当前的能力、挑战与未来方向。分析涵盖学术与商业实现、软件框架以及实际应用。本文对RISC-V机器学习生态系统进行了评估,内容涵盖指令集扩展、核心实现、编译器优化以及部署策略。主要贡献包括:提出了RISC-V机器学习实现的统一分类法,对性能与设计权衡进行了比较分析,评估了软件工具链的成熟度,并识别了指令集扩展与专用加速器领域的新兴趋势。研究结果显示,在能效、专用指令开发和框架集成方面取得了进展,同时也凸显了标准化、验证复杂性和生态系统碎片化方面的挑战。本分析提出了四个应对当前局限的研究方向:专用神经网络处理扩展、自适应与模块化处理器架构、安全框架以及高能效的多领域架构。这些方向为推进RISC-V成为下一代机器学习系统的基础平台提供了路线图。
cs.LG / 103 / 2609.20715

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

不要掩盖环境:观察监督改变了智能体在强化学习下的探索方式
Zhang, Juzheng, Makhija, Disha, Arivazhagan, Manoj Ghuhan, Kumar, Vinayshekhar Bannihatti, Gangadharaiah, Rashmi
Abstract
Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters, sequence tokens, or forward passes. The methods perform similarly after SFT but diverge after GRPO. On Qwen3-4B, GRPO from ActObs achieves higher pass@k at every evaluated sampling budget than its action-only counterpart on Terminal-Bench 2.0. On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks. The advantage extends to cross-domain code editing on aider-polyglot (+4.2 pp at pass@1 at 4B), whose tasks are unseen during SFT and RL. ActObs retains more entropy during RL while requiring less policy movement, leaving the final policy closer to its SFT initialization. Our analysis traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model. Joint supervision prevents this one-sided specialization, preserving consequence prediction and preparing the policy for downstream exploration.
Chinese Translation
智能体轨迹记录了智能体的行为以及随后发生的事情。然而,标准的监督微调(SFT)仅对智能体生成的动作词元施加损失,将环境观察仅作为上下文使用,而不作为预测目标。我们探究这一惯例是否能为后续的强化学习提供最佳初始化。我们提出了 ActObs,它还对每条轨迹中已有的观察词元进行监督。尽管部署时的智能体从不生成观察,但学习预测观察能够促使策略对动作后果进行建模,且无需增加数据、参数、序列词元或前向传播。两种方法在 SFT 之后表现相近,但在 GRPO 之后出现分化。在 Qwen3-4B 上,基于 ActObs 的 GRPO 在 Terminal-Bench 2.0 的每个评估采样预算下均比仅监督动作的对应方法取得更高的 pass@k。在 Qwen3-8B 上,它以一定的 pass@1 可靠性换取了更高的 pass@k(pass@16 提升 3.4 个百分点),并解决了更多不同的任务。这一优势还扩展到跨域代码编辑任务 aider-polyglot 上(4B 模型在 pass@1 上提升 4.2 个百分点),该任务在 SFT 和 RL 阶段均未出现过。ActObs 在 RL 过程中保留了更多的熵,同时所需的政策移动更小,使最终策略更接近其 SFT 初始化。我们的分析将这一差异追溯至 SFT 阶段:动作和观察的梯度迅速变得正交,而仅训练动作会使大量残余的观察梯度留存,并导致环境预测能力退化至低于基础模型的水平。联合监督防止了这种单边特化,保留了对后果的预测能力,并为策略的下游探索做好了准备。
cs.LG / 104 / 2609.20744

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Video DeltaNet:一种面向直播视频生成的视频原生混合注意力机制
Xi, Haocheng, Xie, Yiming, Zhao, Hexu, Zhang, Yiwen, Liu, Michael, Creavin, Thomas, Keutzer, Kurt, Li, Xiuyu, Lv, Zhaoyang, Xu, Chenfeng, Feng, Haiwen
Abstract
Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for long-range video context. Its linear branch introduces Video Delta Attention (VDA), which updates memory once per frame by jointly incorporating its spatial tokens. Separate output projections and learnable gates calibrate the two branches, while a staged teacher-alignment recipe progressively introduces the new pathway into pretrained models. We instantiate VDN on MiniMax H3, applying the hybrid to video-to-video interactions while retaining Softmax for interactions involving text or audio. With eight-step distillation and an optimized SGLang serving stack, VDN-H3 completes DiT denoising for a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs, corresponding to a 14.5x speedup over the 50-step dense H3 baseline on the same GPU count.
Chinese Translation
视频扩散模型在去噪过程中需要反复处理漫长的时空标记序列,使得注意力机制成为主要的计算瓶颈。线性注意力提供了一种有吸引力的替代方案,并已在近期的大型语言模型中被广泛采用,但将其直接应用于视频模型往往无法保留高质量生成所需的细粒度交互。我们提出了 Video DeltaNet(VDN),它将局部 Softmax 注意力与双向线性记忆相结合,以捕获长程视频上下文。其线性分支引入了 Video Delta Attention(VDA),通过联合融合每帧的空间标记,实现每帧一次的记忆更新。分离的输出投影和可学习的门控机制用于校准两个分支,同时采用分阶段的教师对齐策略,将新通路逐步引入预训练模型。我们在 MiniMax H3 上实现了 VDN,将该混合架构应用于视频到视频的交互,同时在涉及文本或音频的交互中保留 Softmax 注意力。借助八步蒸馏和优化的 SGLang 推理服务栈,VDN-H3 在八块 NVIDIA B200 GPU 上仅需 6.70 秒即可完成一段 14.3 秒、768p 视频的 DiT 去噪,相比相同 GPU 数量下 50 步稠密 H3 基线实现了 14.5 倍的加速。
cs.LG / 105 / 2609.20765

Calibrated RF-Fingerprinting Under Interference With Heterogeneous Transmission Protocols

异构传输协议干扰下的校准射频指纹识别
Abdul-Quddoos, Tariq, Li, Xiangfang, Qian, Lijun
Abstract
Radio Frequency(RF)-Fingerprinting is a spectrum monitoring technique that identifies specific transmitters based on hardware impairments imprinted within the emitted signal. Although widely researched, studies almost exclusively consider scenarios where only one transmitter is emitting at a time, limiting real world applicability. In this work, we further the study of RF-Fingerprinting by considering co-channel interference, with multiple emitted signals interfering with each other, overlapping in time and frequency. Specifically, we formulate this problem as a multi-label classification problem and employ a 1D convolutional neural network (CNN). Furthermore, the models are calibrated such that the confidence thresholds for the label probabilities are derived, with guarantees on the upper bound on the average number of False Negatives, providing a degree of confidence in not missing a true spectrum policy violation. The proposed method is validated using real world data from the POWDER 5G testbed on devices transmitting 802.11a(Wi-Fi), 4G LTE, and 5G NR waveforms. The results show accuracy as high as 97% and as low as 73% after calibration depending on channel conditions. Also calibrating for various average false negatives upper bounds achieves micro recall scores of approximately (1 - calibrated false negatives) with the calibration robust to out-of-distribution interference, demonstrating the potential of the proposed method in a realistic high contention wireless environment
Chinese Translation
射频指纹识别(RF-Fingerprinting)是一种频谱监测技术,它基于发射信号中蕴含的硬件损伤痕迹来识别特定的发射机。尽管该技术已被广泛研究,但现有研究几乎仅考虑同一时刻只有一个发射机发射信号的场景,限制了其在现实世界中的适用性。在本工作中,我们通过考虑同频干扰进一步拓展了射频指纹识别的研究,即多个发射信号相互干扰,并在时间和频率上重叠。具体而言,我们将该问题形式化为一个多标签分类问题,并采用一维卷积神经网络(CNN)。此外,我们对模型进行了校准,从而推导出标签概率的置信度阈值,并对假阴性(False Negatives)平均数量的上界提供保证,从而为不漏检真实的频谱政策违规行为提供一定程度的置信度。所提出的方法利用来自 POWDER 5G 试验台的真实数据进行了验证,测试设备发射 802.11a(Wi-Fi)、4G LTE 和 5G NR 波形。结果表明,根据信道条件的不同,校准后的准确率最高可达 97%,最低为 73%。此外,针对不同的假阴性平均上界进行校准后,微平均召回率(micro recall)约为(1 - 校准假阴性率),且该校准对分布外干扰具有鲁棒性,展示了所提方法在现实高竞争无线环境中的潜力。
cs.LG / 106 / 2609.20794

PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers

PosteriorBench:从点估计到后验匹配的生成式逆问题求解器评估
Yao, Jiachen, Hsu, Zi-Siang, Deng, Xi, Gupta, Aditi, Ju, Xin, Benson, Sally M, Wen, Gege, Anandkumar, Anima
Abstract
Generative models are increasingly used to solve scientific inverse problems, but existing evaluations still focus primarily on whether a method can produce a single plausible reconstruction. This is insufficient for ill-posed problems, where multiple solutions may be consistent with the same sparse or noisy observations. In these settings, a method can achieve strong pointwise accuracy while still failing to capture the true posterior through mode collapse, overconfident uncertainty, or averaging incompatible solutions. We introduce PosteriorBench, a benchmark for evaluating the distributional accuracy of generative inverse solvers. PosteriorBench evaluates four physics-based inverse problems: Darcy flow inversion, Poisson source recovery, carbon capture and storage, and light transport material inference. For each task, we construct high-fidelity reference posteriors using computationally heavy but established procedures such as rejection sampling and Markov chain Monte Carlo, enabling direct assessment of whether solvers recover the full set of solutions rather than the single best sample. We pair these references with a five-metric posterior evaluation suite: posterior-mean error, posterior-standard-deviation error, maximum mean discrepancy, sliced Wasserstein distance, and radially averaged power-spectrum error. These metrics assess pointwise accuracy, marginal uncertainty, distributional alignment, and global frequency fidelity. The benchmark spans sparse sensing, low-resolution observations, nonlinear forward models, varying noise levels, and multimodal priors, with a unified pipeline for distribution matching and uncertainty quantification. Our experiments reveal substantial distribution-matching gaps across current solvers, while showing that neural operators improve resolution robustness, and guidance weights and generation noise are key to posterior-variance calibration.
Chinese Translation
生成模型正被越来越多地用于求解科学逆问题,但现有的评估仍主要关注方法能否产生单一合理的重建结果。这对于不适定问题而言是不够的,因为在此类问题中,多个解可能与相同稀疏或含噪的观测一致。在这些场景下,一种方法即使在逐点精度上表现出色,仍可能因模式坍缩、过度自信的不确定性估计,或对互不兼容的解进行平均,而无法捕捉真实后验分布。我们提出了PosteriorBench,一个用于评估生成式逆问题求解器分布精度的基准。PosteriorBench评估了四个基于物理的逆问题:Darcy流反演、泊松源恢复、碳捕集与封存,以及光传输材质推断。对于每个任务,我们使用计算开销较大但成熟可靠的方法(如拒绝采样和马尔可夫链蒙特卡洛)构建高保真参考后验,从而能够直接评估求解器是否恢复了完整的解集合,而不仅仅是单一最优样本。我们将这些参考与一个包含五种度量的后验评估套件相结合:后验均值误差、后验标准差误差、最大均值差异(MMD)、切片Wasserstein距离,以及径向平均功率谱误差。这些指标分别评估逐点精度、边缘不确定性、分布对齐性和全局频率保真度。该基准涵盖稀疏感知、低分辨率观测、非线性正向模型、不同噪声水平以及多模态先验,并提供了一个用于分布匹配和不确定性量化的统一流程。我们的实验揭示了当前求解器在分布匹配方面的显著差距,同时表明神经算子提升了分辨率鲁棒性,而引导权重和生成噪声是后验方差校准的关键。
cs.LG / 107 / 2609.20807

Score Centering Stabilizes Off-policy Reinforcement Learning

分数中心化稳定离策略强化学习
Marek, Martin, Ryabinin, Max
Abstract
Reinforcement learning (RL) of large language models is notoriously sensitive to small differences between training and inference engines, often referred to as the training-inference mismatch (TIM). However, completely eliminating TIM is impractical, as it would come at a major cost to rollout efficiency. In this paper, we show that the instability of RL under TIM is primarily caused by drift: a persistent bias between training and inference engines that accumulates with every training step. We derive an additive "score centering" correction term that stabilizes RL under TIM by canceling drift. When training models from 0.6B to 30B parameters, score centering alone matches or outperforms methods based on importance sampling under quantization, with the gap growing as the mismatch becomes more severe. Because the correction is additive, score centering also composes with importance sampling -- their composition outperforms pure importance-sampling baselines in our staleness experiments.
Chinese Translation
大语言模型的强化学习(RL)对训练引擎与推理引擎之间的微小差异极为敏感,这种现象通常被称为训练-推理失配(training-inference mismatch, TIM)。然而,完全消除TIM并不现实,因为这会严重牺牲 rollout 效率。在本文中,我们证明TIM下强化学习的不稳定性主要由漂移(drift)引起:即训练与推理引擎之间随每个训练步不断累积的持续性偏差。我们推导出一个加性的“分数中心化”(score centering)修正项,通过消除漂移来稳定TIM下的强化学习。在训练参数量从0.6B到30B的模型时,仅使用分数中心化即可达到或超越基于量化下重要性采样的方法,且失配越严重,优势越明显。由于该修正是加性的,分数中心化还可以与重要性采样相结合——在过时性(staleness)实验中,二者的结合优于纯重要性采样基线。