← Back to Index
Daily Research Digest

arXiv Papers

2026-09-28
378
Papers
4
Categories
378
Translated
收藏清单 0
机器人学 (Robotics)
80
cs.RO / 1 / 2609.30358

TinyCVIO: A Constellation-Aided Visual-Inertial Odometry System for Nanodrones

TinyCVIO:一种面向纳米无人机的星座辅助视觉惯性里程计系统
Ozturk, Derin, Akan, Kaan, Wang, Irwin, Batten, Christopher
Abstract
Nanodrones require accurate, real-time state estimation under severe sensing and computational constraints. We present TinyCVIO, a visual-inertial odometry system that co-designs miniature sensing, visual processing, and estimation for a commodity dual-core microcontroller with 520 kB SRAM. Lightweight LED constellations provide known geometry without surveyed positions or yaw angles, assuming placement on a common level plane. A streaming visual frontend tracks LED observations from a millimeter-scale camera at 29.2 FPS, while a rigid-board measurement model retains inter-LED constraints and streaming QR bounds estimation workspace for a fixed filter-state size. Across 19 hand-held hardware-in-the-loop datasets, the rigid-board model reduces mean absolute trajectory error by 27% relative to planar points. The complete system runs onboard a Crazyflie across nine flights at three speeds, achieving 3.5-3.7 cm mean absolute trajectory error and 0.50-0.60% relative pose error over 10 m segments, with mean estimate latency of 15.7-16.3 ms.
Chinese Translation
纳米无人机(nanodrone)需要在严苛的感知与计算约束下实现精确的实时状态估计。我们提出了TinyCVIO,一种针对具备520 kB SRAM的商用双核微控制器进行协同设计的视觉惯性里程计系统,其协同设计涵盖微型传感、视觉处理与状态估计。轻量级LED星座在假定安装于同一水平平面的前提下,提供了已知几何结构,而无需测量位置或偏航角。流式视觉前端以29.2 FPS的帧率对毫米级相机的LED观测进行跟踪,同时刚性板测量模型保留了LED间的约束关系,流式QR分解将估计工作空间限制在固定的滤波器状态规模内。在19组手持硬件在环数据集上,刚性板模型相较平面点模型将平均绝对轨迹误差降低了27%。完整系统在Crazyflie机载运行,进行了三种速度下的九次飞行实验,在10米路段上实现了3.5-3.7 cm的平均绝对轨迹误差和0.50-0.60%的相对位姿误差,平均估计延迟为15.7-16.3 ms。
cs.RO / 2 / 2609.30404

POIL: Point-based One-Shot Imitation Learning with Stable Dynamical Systems

POIL:基于点集与稳定动力学系统的单样本模仿学习
Kim, Sang Min, Seo, Jinwoo, Heo, Hyeongjun, Lee, Junho, Lee, Yonghyeon, Kim, Young Min
Abstract
We present POIL, a point-based one-shot imitation learning framework with stable dynamical systems. While one-shot imitation avoids collecting extensive demonstrations, successful one-shot manipulation requires not only transferring a demonstrated trajectory to a novel object but also executing it robustly under changing scene conditions, grasp configurations, and external disturbances. POIL addresses both problems through a shared representation: a set of 3D points on the object's functional part, used jointly for trajectory transfer and closed-loop execution. The one-shot transfer from the demonstrated trajectory is enabled with point correspondences. POIL grounds the shared functional part with a multi-modal large language model, and transfers the trajectory across viewpoint, pose, and object category changes. During execution, multi-view tracking observes the same points online, and Point-set BCSDM drives them in closed loop by projecting per-point velocities onto a single rigid-body twist computed from the tracked points alone. This extends stable dynamical models from an SE(3) pose to a point set without requiring a known 3D model or pose estimator. We show that at the goal the controller becomes a gradient flow on the classical SO(3) potential, so its terminal phase inherits the almost-global convergence of that potential under a rigid-object assumption. Across simulation and real-robot experiments, POIL transfers a single demonstration across object category, grasp pose, and goal geometry, while recovering from external disturbances during execution. Project page: https://sangminkim-99.github.io/poil
Chinese Translation
我们提出了POIL,一种结合稳定动力学系统的基于点集的单样本模仿学习框架。单样本模仿学习虽然避免了采集大量演示数据,但成功的单样本操作不仅需要将演示轨迹迁移到新物体上,还需要在变化的场景条件、抓取构型以及外部干扰下鲁棒地执行该轨迹。POIL通过一种共享表示同时解决这两个问题:物体功能部件上的一组三维点,联合用于轨迹迁移与闭环执行。演示轨迹的单样本迁移通过点对应关系实现。POIL利用多模态大语言模型定位共享的功能部件,并在视角、位姿和物体类别变化的情况下完成轨迹迁移。在执行阶段,多视角跟踪在线观测相同的点,Point-set BCSDM通过将逐点速度投影到仅由跟踪点计算出的单一刚体螺旋运动(twist)上,实现闭环驱动。这将稳定动力学模型从SE(3)位姿扩展到了点集,而无需已知的三维模型或位姿估计器。我们证明,在目标点处,该控制器退化为经典SO(3)势函数上的梯度流,因而在刚体假设下,其终端阶段继承了该势函数的几乎全局收敛性。在仿真和真实机器人实验中,POIL能够将单条演示迁移到不同物体类别、抓取位姿和目标几何形状上,并在执行过程中从外部干扰中恢复。项目主页:https://sangminkim-99.github.io/poil
cs.RO / 3 / 2609.30428

Actively Resolving Contextual Uncertainty for Underspecified Tasks in Natural Language

面向自然语言欠规范任务的上下文不确定性主动消解
Ravichandran, Zachary, Diller, Jonathan, Cladera, Fernando, Murali, Varun, Pappas, George J., Kumar, Vijay
Abstract
Foundation models provide robots with the ability to interpret natural language and reason about environmental context, yet most language-conditioned policies assume that goals are well-specified and that task-relevant information is provided upfront via a prior map. Operating in unfamiliar environments with underspecified tasks entails high contextual uncertainty: the robot must jointly infer what constitutes task success, what constitutes relevant information, and where (or whether) that information exists. We address these limitations via CLUE (Closed-Loop contextual Uncertainty rEsolution), a framework for actively resolving contextual uncertainty given underspecified tasks in natural language. CLUE uses an LLM-derived policy to hypothesize task-relevant concepts and potential plans. It then uses a language-embedded map, which is constructed online, to ground these hypotheses into actions. The policy sequentially evaluates hypotheses via closed-loop environment interaction and refines its plans as it gathers new information. We deploy CLUE on a Boston Dynamics Spot across three real indoor and outdoor environments spanning 15 tasks that require object disambiguation, functional inference, and occlusion reasoning. CLUE achieves a success rate within 7 percentage points of an oracle policy and outperforms an LLM-enabled planner without closed-loop feedback by a 4x margin. Supporting experiments demonstrate that simply building and then querying a language-enriched map is insufficient to resolve complex contextual planning tasks; these approaches achieve roughly one third the success rate of CLUE while requiring over 10x more VLM tokens. We provide additional information at https://zacravichandran.github.io/CLUE.
Chinese Translation
基础模型使机器人能够理解自然语言并对环境上下文进行推理,然而大多数语言条件策略都假设目标已被明确给出,且任务相关信息通过先验地图预先提供。在不熟悉的环境中执行欠规范(underspecified)的任务会带来高度的上下文不确定性:机器人必须同时推断什么构成任务成功、什么构成相关信息,以及这些信息存在于何处(或是否存在)。我们提出 CLUE(Closed-Loop contextual Uncertainty rEsolution,闭环上下文不确定性消解)框架来应对这些局限,该框架能够针对自然语言给出的欠规范任务主动消解上下文不确定性。CLUE 利用由大语言模型(LLM)导出的策略来推测任务相关概念和潜在计划,然后使用在线构建的语言嵌入地图将这些假设具体化为可执行动作。该策略通过闭环环境交互依次评估各假设,并随着新信息的获取不断修正其计划。我们将 CLUE 部署在 Boston Dynamics Spot 机器人上,在三个真实的室内外环境中完成了 15 个任务,涵盖物体消歧、功能推断和遮挡推理。CLUE 的成功率与先知(oracle)策略相差不超过 7 个百分点,且相比没有闭环反馈的 LLM 规划器性能提升达 4 倍。辅助实验表明,仅仅构建并查询一个语言增强地图不足以解决复杂的上下文规划任务:这类方法的成功率仅为 CLUE 的约三分之一,同时所需的视觉语言模型(VLM)token 超过 CLUE 的 10 倍。更多信息请参见 https://zacravichandran.github.io/CLUE。
cs.RO / 4 / 2609.30436

WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

WALT:面向自动驾驶的世界模型对齐潜在轨迹学习
Jia, Mingkai, Guo, Jiaxin, Shu, Zhijian, Xu, Jiawei, Li, Mingxiao, Cheng, Jintao, Tan, Ping, Yin, Wei
Abstract
Driving world models learn rich predictive representations of the surrounding environment from visual observations, yet accurate visual prediction does not necessarily translate into effective trajectory planning. We argue that a key bottleneck lies in the mismatch between visual world states and raw geometric trajectories, which may limit the planner's ability to exploit action-relevant semantics encoded by the world model. To address this issue, we propose World-Model Alignment for Latent Trajectories (WALT), which learns a compact generative trajectory latent space by transferring information from a frozen pretrained driving world model without modifying the world model itself. Rather than directly generating raw waypoints, WALT maps them into compact representations through a dual-branch trajectory autoencoder and transfers semantic knowledge from the frozen visual world model into this trajectory space, encouraging the learned action representation to capture scene-level cues relevant to future motion and planning. Beyond our proposed formulation, we systematically study latent learning based on Joint-Embedding Predictive Architectures (JEPA) and feature alignment following Representation Alignment (REPA) to investigate how trajectory-only representation learning affects downstream planning. We evaluate WALT on the NAVSIM benchmarks. Relative to the raw-waypoint baseline, WALT improves PDMS from 89.4 to 89.8 on NAVSIMv1 and EPDMS from 87.3 to 87.9 on NAVSIMv2 while reducing trajectory planner FLOPs by 30.5%. These results suggest that preserving world representations while extracting action-relevant information provides an effective interface for world-model-based trajectory planning.
Chinese Translation
驾驶世界模型能够从视觉观测中学习周围环境的丰富预测表示,然而精确的视觉预测并不一定能转化为有效的轨迹规划。我们认为,一个关键瓶颈在于视觉世界状态与原始几何轨迹之间的不匹配,这可能限制了规划器利用世界模型所编码的动作相关语义的能力。为解决这一问题,我们提出了面向潜在轨迹的世界模型对齐方法(World-Model Alignment for Latent Trajectories, WALT),该方法在不修改世界模型本身的前提下,通过从冻结的预训练驾驶世界模型中迁移信息,学习一个紧凑的生成式轨迹潜在空间。WALT 并非直接生成原始路径点,而是通过双分支轨迹自编码器将其映射为紧凑表示,并将冻结视觉世界模型中的语义知识迁移到该轨迹空间中,从而促使所学习的动作表示能够捕获与未来运动和规划相关的场景级线索。除了我们提出的构建方式外,我们还系统地研究了基于联合嵌入预测架构(Joint-Embedding Predictive Architectures, JEPA)的潜在学习以及遵循表征对齐(Representation Alignment, REPA)的特征对齐方法,以探究仅基于轨迹的表征学习如何影响下游规划。我们在 NAVSIM 基准上对 WALT 进行了评估。相对于原始路径点基线,WALT 在 NAVSIMv1 上将 PDMS 从 89.4 提升至 89.8,在 NAVSIMv2 上将 EPDMS 从 87.3 提升至 87.9,同时将轨迹规划器的浮点运算量(FLOPs)降低了 30.5%。这些结果表明,在保留世界模型表示的同时提取动作相关信息,为基于世界模型的轨迹规划提供了一种有效的接口。
cs.RO / 5 / 2609.30459

VkVIO: Cross-platform GPU Acceleration for Visual-Inertial Odometry with Vulkan

VkVIO:基于Vulkan的视觉惯性里程计跨平台GPU加速
Hoffmann, Ole, de Mayo, Mateo, Cremers, Daniel
Abstract
Perception in robotics and XR fundamentally relies on good state estimation. Visual-inertial odometry (VIO) and Simultaneous Localization and Mapping (VI-SLAM) are proven ways of achieving this goal in a cost-effective and accurate manner. Efficiency in these systems allows for smaller, cooler, and lighter devices. GPU acceleration is a natural approach for reducing latency, thanks to their wide availability in platforms like embedded computers, mobile phones, and XR headsets. However, previous works in the literature have limited themselves to the use of CUDA for this task, significantly reducing deployment options to a single vendor. We instead leverage the vendor-agnostic Vulkan API, originally designed for the strict performance requirements of 3D graphics applications. In this work, we present VkVIO, the first, to the best of our knowledge, cross-platform GPU-accelerated VIO method. We provide state-of-the-art accuracy with causal estimates required for real-time operation. We deploy VkVIO on a diverse range of devices spanning a workstation, a laptop, and an extremely inexpensive single-board computer, while outperforming CUDA-based systems on the same hardware. VkVIO enables possibilities for low-latency, low-power, and low-cost VIO in robotics and XR.
Chinese Translation
机器人与扩展现实(XR)中的感知从根本上依赖于良好的状态估计。视觉惯性里程计(VIO)与视觉惯性同步定位与建图(VI-SLAM)已被证明是以高性价比和高精度方式实现这一目标的有效手段。这类系统的高效率可使设备更小巧、更低温、更轻便。由于GPU在嵌入式计算机、手机和XR头显等平台上广泛可用,GPU加速是降低延迟的天然途径。然而,现有文献中的工作仅限于使用CUDA完成这一任务,这极大地将部署选择限制在单一厂商。与此不同,我们采用与厂商无关的Vulkan API,该API最初是为满足3D图形应用的严苛性能要求而设计的。在本工作中,我们提出了VkVIO,据我们所知,这是首个跨平台的GPU加速VIO方法。我们实现了最先进的精度,并提供实时运行所需的因果估计。我们将VkVIO部署于涵盖工作站、笔记本电脑和一台极其廉价的单板计算机等多种设备上,并在相同硬件上超越了基于CUDA的系统。VkVIO为机器人与XR领域的低延迟、低功耗和低成本VIO开辟了可能。
cs.RO / 6 / 2609.30460

Realizability Is Not Enough: Encoding, Liveness, and Auditing of Synthesized Robot Supervisors

可实现性并不足够:合成机器人监督器的编码、活性与审计
Conner, David C., Luzier, Joshua, Doyle, William J., Faith, Emma R., Kooiker, Aubrie B., Farney, Andrew J., Fox, Sebastian, Grimes, Evangelina, Conner, Ian G., Bloom, Kyle
Abstract
High-level robotic supervisors coordinate capabilities whose reported outcomes determine the robot's next action. Reactive synthesis can generate such supervisors with formal guarantees, but deployment requires more than proving a Generalized Reactivity (1) (GR(1)) specification realizable. Designers must encode failure-prone capabilities, choose liveness assumptions that match retry intent, audit strategies, and translate them into robot software. We present an open-source pipeline for Robot Operating System (ROS) 2 Flexible Behavior Engine (FlexBE) supervisors that generates capability-based GR(1) specifications, analyzes assumptions before synthesis, audits strategies, reduces states with a behavior-preservation proof, and emits executable state machines. Across four case studies (six comparisons), including hardware on two quadcopter platforms, we compare enumerated and one-hot encodings and two liveness formulations. Under the tested backend, enumerated encoding usually synthesizes faster, although fewer propositions do not reliably predict smaller controllers or lower symbolic cost. System-Goal without pending memory is the only liveness treatment confirmed to yield executable controllers under both encodings across the reported grid; Fair-Outcome can permit realizable cycles without designer-intended completion. For this backend and model, we recommend enumerated encoding with System-Goal and auditing every realized strategy, since proposition count and realizability do not measure deployability. The auditor is sound and complete for four structural defect classes (protocol violations, deadlocks, bounded-failure violations, goal-unreachable traps) but is not a general liveness verifier, and the reduction preserves capability-level behavior. Together, these stages narrow the gap between formal realizability and controllers that pass protocol and structural-progress checks.
Chinese Translation
高层机器人监督器协调多种能力(capability),其报告的执行结果决定机器人的下一步动作。反应式综合可以生成具有形式化保证的此类监督器,但部署所需的工作远不止证明广义反应性(1)(GR(1))规约是可实现的。设计者必须对易出错的能力进行编码、选择与重试意图相匹配的活性假设、审计策略,并将其转化为机器人软件。我们提出了一个面向机器人操作系统(ROS)2 柔性行为引擎(FlexBE)监督器的开源流水线,该流水线生成基于能力的 GR(1) 规约,在综合前分析假设、审计策略、在保持行为的证明下进行状态约简,并输出可执行的状态机。在四个案例研究(六组对比)中,包括在两个四旋翼平台上的硬件实验,我们比较了枚举编码与独热(one-hot)编码以及两种活性表述。在所测试的后端下,枚举编码通常综合速度更快,但命题数量较少并不能可靠地预测更小的控制器或更低符号开销。在报告的实验网格中,无待处理记忆的 System-Goal 是唯一在两种编码下均被证实能产生可执行控制器的活性处理方式;Fair-Outcome 可能允许存在可实现的循环而无法达成设计者预期的完成。针对该后端和模型,我们推荐使用枚举编码结合 System-Goal,并对每个已实现的策略进行审计,因为命题数量和可实现性并不能衡量可部署性。该审计器对四类结构性缺陷(协议违反、死锁、有界失败违反、目标不可达陷阱)是可靠且完备的,但并非通用的活性验证器,且状态约简保持了能力层面的行为。这些阶段共同缩小了形式化可实现性与通过协议及结构性进展检查的控制器之间的差距。
cs.RO / 7 / 2609.30461

DGT-Map: Directional Global Traversability Mapping Utilizing Multi-Task Learning for Heterogeneous Vehicles

DGT-Map:利用多任务学习的面向异构车辆的方向性全局可通行性地图
Singh, Jaskrit, Noori, Kashif K., Xiao, Jing, Chamzas, Constantinos
Abstract
Off-road traversability is direction-dependent and vehicle specific, yet most global maps assign a single isotropic cost to each location. Existing learned estimators are also commonly trained independently for each vehicle; this preserves vehicle-specific behavior but prevents vehicles from sharing common terrain representations. DGT-MAP addresses both limitations through a self-supervised framework that learns global, directional, and vehicle-conditioned traversability costmaps from RGB-D observations and locomotion signals. A shared multi-task backbone learns common terrain features across training vehicles while vehicle-specific prediction heads preserve platform-dependent responses. At inference, DGT-MAP produces a heading-indexed costmap that can be used by a direction-aware planner. We evaluate DGT-MAP in simulation by integrating it into a Hybrid A* navigation stack and measuring downstream task success on challenging terrains, including slopes that are traversable downhill but not uphill and a ridge obstacle that is traversable by some vehicles, but not by others. Across evaluated tasks, DGT-MAP achieves the highest or tied-highest navigation success rate when compared against geometric, binary, and learned direction-agnostic baselines.
Chinese Translation
越野可通行性具有方向依赖性和车辆特异性,然而大多数全局地图为每个位置分配单一的各向同性代价。现有的学习型估计器通常也是针对每辆车独立训练的,这虽然保留了车辆特异性行为,但阻碍了车辆之间共享通用的地形表征。DGT-MAP 通过一个自监督框架同时解决了这两个局限性,该框架从 RGB-D 观测和运动信号中学习全局的、方向性的、以车辆为条件的可通行性代价地图。共享的多任务骨干网络跨训练车辆学习通用地形特征,而车辆专属的预测头则保留平台相关的响应。在推理阶段,DGT-MAP 生成以航向为索引的代价地图,可供方向感知规划器使用。我们通过将 DGT-MAP 集成到 Hybrid A* 导航栈中,并在具有挑战性的地形上测量下游任务成功率,在仿真环境中对其进行了评估。这些地形包括下坡可通行而上坡不可通的斜坡,以及某些车辆可以通行而其他车辆无法通行的山脊障碍。在所有评估任务中,与几何方法、二值方法以及学习型方向无关基线相比,DGT-MAP 取得了最高或并列最高的导航成功率。
cs.RO / 8 / 2609.30462

Policy-Calibrated DAgger: Offline Calibrated Noise Injection for Imitation Learning

策略校准的DAgger:用于模仿学习的离线校准噪声注入
Wang, Jenny, Kantor, George
Abstract
Policies trained with imitation learning can accumulate errors over time, causing the robot to drift outside the training distribution. Existing methods mitigate this covariate shift by collecting additional data where the policy fails or is likely to fail. The first places the robot in unsafe conditions and the second requires choosing an appropriate noise distribution to collect new expert demonstrations under that noise. We propose Policy-Calibrated DAgger, a method that makes use of the properties of recent generative policies to estimate the policy's noise offline by using its own predicted action distribution. We measure a diffusion policy's spread of predicted actions at observations along the expert trajectory and measure its closed-loop error relative to a recorded trajectory. To address issues with measuring error in a multimodal action space, we guide the policy towards the trajectory during closed-loop control through partial denoising, and use properties of a diffusion model to unnormalize the measured error as if we did not guide it. We experiment in a scenario where a robot is tasked to reach an engine lever in a cluttered and narrow environment and show results in a 3D photorealistic simulator and a 2D planar reacher environment. We show that our method surpasses policies trained with dataset aggregation without noising and matches the performance of the best noise level in hindsight, without requiring a sweep over noise levels.
Chinese Translation
通过模仿学习训练的策略会随时间累积误差,导致机器人偏离训练分布。现有方法通过在策略失败或可能失败的地方收集额外数据来缓解这种协变量偏移。第一种方法会将机器人置于不安全的条件下,第二种方法则需要选择合适的噪声分布,以便在该噪声下收集新的专家示范数据。我们提出了策略校准的DAgger(Policy-Calibrated DAgger),该方法利用近期生成式策略的特性,通过策略自身预测的动作分布来离线估计策略的噪声。我们在专家轨迹上的观测点处测量扩散策略(diffusion policy)预测动作的离散程度,并测量其相对于录制轨迹的闭环误差。为了解决在多峰动作空间中测量误差的问题,我们在闭环控制过程中通过部分去噪引导策略趋向该轨迹,并利用扩散模型的特性将测得的误差进行反归一化处理,就像我们没有进行引导时一样。我们在一个机器人需在狭窄且杂乱的环境中够到引擎操纵杆的场景中进行实验,并在3D照片级真实感仿真器和2D平面够物环境中展示了结果。结果表明,我们的方法优于使用不带噪声的数据集聚合(dataset aggregation)训练的策略,并且能够达到事后选取最优噪声水平时的性能,而无需对噪声水平进行扫描搜索。
cs.RO / 9 / 2609.30479

Learning-Based Pressure Predictive Control of a Vertebraic Soft Robotic Tail

基于学习的椎骨式软体机器人尾巴压力预测控制
Yang, Wenjian, Huang, Nan, Nie, Yukang, Chen, Fang, Chi, Wanchao, Dai, Jiansheng, Liu, Sicong
Abstract
Soft robots have attracted much attention for their safe human-robot interaction and flexibility, but the typical continuum structure and nonlinear material behavior make the kinematics modelling complex, especially in non-static motions. In this work, we proposed an LSTM-based pressure predictive control (PPC) for the motion control of a vertebraic soft robotic tail and the coordination with a quadruped robot. The PPC consists of an inverse kinematics (IK) model, a forward kinematics (FK) model and a pressure compensation (P-comp) model, and achieves non-static and quasi-static motion control of the tail. Compared with the IK-only model, the average RMSE of the PPC's simulation trajectories reduces by 69.8%, when executing target trajectories. In the coordinated motions of the soft tail quadruped, using a prediction data set to train the PPC enables next-moment action prediction and reduces computation time by 60.9%, which enhances the real-time response of the tail to match the quadruped torso's moving rate. The PPC provides a simple and effective method to model the soft tail for both non-static and quasi-static motion control, and grants the soft tail quadruped with the functionality of interacting with the environment.
Chinese Translation
软体机器人因其安全的人机交互能力和柔性而备受关注,但其典型的连续体结构和非线性材料行为使得运动学建模十分复杂,尤其是在非静态运动中。本文提出了一种基于LSTM的压力预测控制(PPC)方法,用于椎骨式软体机器人尾巴的运动控制及其与四足机器人的协调运动。该PPC由逆运动学(IK)模型、正运动学(FK)模型和压力补偿(P-comp)模型组成,可实现软体尾巴的非静态和准静态运动控制。在执行目标轨迹时,与仅使用IK模型的方案相比,PPC仿真轨迹的平均RMSE降低了69.8%。在软体尾巴四足机器人的协调运动中,利用预测数据集训练PPC可实现下一时刻动作预测,并将计算时间减少60.9%,从而增强了软体尾巴的实时响应能力,使其能够匹配四足躯干的运动速率。该PPC为软体尾巴的非静态和准静态运动控制建模提供了一种简单而有效的方法,并赋予软体尾巴四足机器人与环境交互的功能。
cs.RO / 10 / 2609.30495

Memory-Aware Multi-Sensor Perception for Efficient and Safe Navigation in Dynamic Environments

面向动态环境高效安全导航的记忆感知多传感器感知方法
Li, Jingshuo, Xue, Yifan, Li, Yifei, Aditya, Shubhodeep Shiv, Figueroa, Nadia
Abstract
Autonomous navigation in previously unseen environments requires effective perception, persistent environmental representation, and collision avoidance while maintaining progress toward a goal. Existing perception-based methods often rely on prior maps or short-horizon observations, limiting their ability to exploit previously observed structure. We propose a memory-aware multi-sensor navigation framework that integrates LiDAR and RGB perception, online distance-field representation learning, and a stage-adaptive Modulated Control Barrier Function Quadratic Program (MCBF-QP). The framework persistently represents static infrastructure while tracking dynamic obstacles, enabling the MCBF-QP controller to exploit previously observed geometry for obstacle circumvention and adapt its safety constraints and guidance to local conditions. Experiments in complex indoor and outdoor environments demonstrate improved navigation efficiency and goal-reaching performance while maintaining collision avoidance in narrow passages and around dynamic obstacles.
Chinese Translation
在先前未见过的环境中进行自主导航,需要有效的感知、持续的环境表征以及避碰能力,同时保持向目标推进。现有的基于感知的方法通常依赖先验地图或短时程观测,限制了其利用先前观测到的环境结构的能力。我们提出了一种记忆感知的多传感器导航框架,该框架集成了激光雷达(LiDAR)与RGB感知、在线距离场表征学习,以及阶段自适应的调制控制屏障函数二次规划(Modulated Control Barrier Function Quadratic Program, MCBF-QP)。该框架在持续表征静态基础设施的同时跟踪动态障碍物,使MCBF-QP控制器能够利用先前观测到的几何结构进行障碍物绕行,并根据局部条件自适应调整其安全约束与引导。在复杂室内外环境中的实验表明,该方法在狭窄通道和动态障碍物周围保持避碰能力的同时,显著提升了导航效率与目标到达性能。
cs.RO / 11 / 2609.30506

Tactile Sensing Array for Multi-Phalanx Sensing in Humanoid Hands

用于人形机器人手多指节感知的触觉传感阵列
Adwani, Neel, Akash, Muhaiminul Islam, Bhattacharya, Rituja, Wang, Cong
Abstract
Humanoid hands require tactile feedback across the whole finger, not just the fingertip, to grasp and manipulate objects properly. Vision and proprioception alone cannot reliably provide this information, particularly when the hand's own fingers occlude the camera's view of the grasp. We present a low-cost tactile array for a humanoid finger, made from Velostat and conductive tape. The fingertip carries seven contact points including a 2x3 matrix wrapped across its front, left, and right faces, and a separate contact point at the tip. The proximal and middle phalanges each carry a single front-facing contact line. We measure the sensor's hysteresis and recovery time after release, through repeated loading and press-release tests. We also test a compliant, 3D-printed contact structure with a gap and a bump, inspired by similar designs in prior work, and show it cuts recovery time by 74% compared to a flush-contact baseline. We then show the sensor can produce distinct activation patterns for different contact geometries (flat, edge, corner) at the fingertip, and that it registers contact across all three phalanges during a grip. Finally, we discuss the limits of our fabrication changes and point to software-based compensation as a promising way to more directly fix the remaining hysteresis in the future.
Chinese Translation
人形机器人手在抓取和操作物体时需要覆盖整个手指的触觉反馈,而不仅仅是指尖。仅依靠视觉和本体感觉无法可靠地提供这类信息,尤其是当手自身的手指遮挡了摄像头对抓取状态的视野时。我们提出了一种低成本的人形机器人手指触觉阵列,由Velostat(压导电材料)和导电胶带制成。指尖包含七个接触点,其中包括一个包裹在指尖前、左、右三个面的2x3矩阵,以及指尖处的一个独立接触点。近端和中端指节各有一条朝向前方的接触线。我们通过反复加载和按压-释放测试,测量了传感器的迟滞特性和释放后的恢复时间。我们还测试了一种带有间隙和凸起的柔性3D打印接触结构(其设计灵感来自先前工作中的类似设计),结果表明与平面接触的基线相比,该结构将恢复时间缩短了74%。随后,我们展示了该传感器能够在指尖针对不同的接触几何形状(平面、边缘、角落)产生不同的激活模式,并且在抓握过程中能够感知所有三个指节上的接触。最后,我们讨论了制造工艺改进的局限性,并指出基于软件的补偿方法有望在未来更直接地解决剩余的迟滞问题。
cs.RO / 12 / 2609.30521

Aerial Manipulation in the Wild with Onboard Perception, Policy Learning, and Whole-Body Control

基于机载感知、策略学习与全身控制的野外空中操作
Zhan, Yuanzhu, Jiang, Yufei, Zhang, Zemu, Geng, Junyi
Abstract
Aerial manipulation in outdoor environments remains challenging due to the simultaneous requirements of reliable state estimation, stable aerial motion, and precise manipulation under external disturbances. In this work, we present a real-world outdoor aerial manipulation framework that integrates imitation learning, onboard LiDAR-inertial state estimation, and whole-body model predictive control. A Diffusion Policy is trained from manipulation demonstrations to generate desired end-effector motions from onboard observations. These learned commands are executed by a whole-body MPC that jointly coordinates the aerial platform and manipulator to realize the desired end-effector trajectory. To eliminate reliance on external motion-capture infrastructure, the platform employs onboard LiDAR-inertial odometry for state estimation during outdoor operation. We validate the complete framework on a physical aerial manipulator and demonstrate successful execution of outdoor manipulation tasks. The experimental results show that demonstration-driven manipulation policies can be effectively integrated with onboard state estimation and model-based whole-body control to enable aerial manipulation beyond controlled indoor environments.
Chinese Translation
室外环境中的空中操作仍然具有挑战性,因为它同时要求可靠的状态估计、稳定的空中运动以及在外部扰动下的精确操作。在本工作中,我们提出了一套真实世界室外空中操作框架,该框架集成了模仿学习、机载激光雷达-惯性状态估计以及全身模型预测控制(MPC)。我们通过操作示教训练了一个扩散策略(Diffusion Policy),以根据机载观测生成期望的末端执行器运动。这些学习到的指令由一个全身模型预测控制器执行,该控制器协同协调空中平台与机械臂,以实现期望的末端执行器轨迹。为消除对外部动作捕捉基础设施的依赖,该平台在室外作业时采用机载激光雷达-惯性里程计进行状态估计。我们在一个真实的空中操作平台上验证了完整的框架,并成功完成了室外操作任务。实验结果表明,基于示教的操作策略可以与机载状态估计和基于模型的全身控制有效结合,从而将空中操作能力拓展至受控室内环境之外。
cs.RO / 13 / 2609.30523

Containing Behavioral Cascades from Manipulated Claims in LLM-Powered Multi-Robot Systems

抑制大语言模型驱动的多机器人系统中因虚假声明引发的行为级联
Khalid, Waleed Bin, Min, Byung-Cheol
Abstract
Large language model (LLM)-powered multi-robot systems are vulnerable to semantic manipulation: an accepted false world-state claim can trigger a fleet-wide behavioral cascade, causing unnecessary replanning, increased path costs, congestion, or apparent mission infeasibility. Conditioning on a successful manipulation, we propose an active verification framework that contains its downstream effects before they propagate across the fleet. A dedicated verification module generates a structured Verify-Adapt-Hold plan: selected robots inspect consequential regions, a limited subset provisionally adapts when necessary, and the remaining robots retain their trusted plans. We evaluate the framework in a multi-robot transportation environment using injected false obstacle claims across different impacts and team sizes. Evaluation measures cascade containment, Sum-of-Costs, makespan, and coverage ratio. Results show that treating post-compromise verification as a team-level planning problem, rather than a binary trust decision, effectively limits the cascading physical consequences of semantic manipulation.
Chinese Translation
由大语言模型(LLM)驱动的多机器人系统容易受到语义操纵的影响:一条被接受的世界状态虚假声明可能触发全队范围的行为级联,导致不必要的重规划、路径成本增加、拥塞,或出现任务不可行的假象。在假设操纵已成功发生的前提下,我们提出了一种主动验证框架,以在其传播至整个机群之前抑制其下游影响。一个专门的验证模块生成结构化的“验证-调整-保持”(Verify-Adapt-Hold)计划:被选中的机器人对关键区域进行检查,一小部分机器人在必要时临时调整其计划,其余机器人则保留其可信计划。我们在多机器人运输环境中,通过注入具有不同影响程度和不同团队规模的虚假障碍物声明来评估该框架。评估指标包括级联抑制效果、总成本(Sum-of-Costs)、完工时间(makespan)和覆盖率。结果表明,将攻陷后的验证视为团队级规划问题,而非二元的信任决策,能够有效限制语义操纵带来的级联物理后果。
cs.RO / 14 / 2609.30530

A Long-Legged, Direct-Drive Monopedal Robot Achieves Exceptional Jump Height

一种实现卓越跳跃高度的长腿直接驱动单腿机器人
Na, Gihyeok, Yim, Justin K.
Abstract
We demonstrate a jumping robot that reaches high (7.6 m) and fast (190 ms stance time) jumps from a single long leg driven by a direct-drive transmission, without elastic energy storage. At 281 g, it achieves the highest jump yet reported for an electrically actuated, spring-free system. The leg uses a new fabric-wrap transmission that provides a variable mechanical advantage, keeping a small electric motor near its peak power output through the stroke while bracing the long, lightweight leg against buckling. A balancing module at the top of the leg uses small propellers to control the leg's orientation on the ground and in the air, where the long leg provides a large moment arm for the control torques. The robot is validated outdoors with vertical jumps, attitude control on the ground and in flight, and tilted jumps.
Chinese Translation
我们展示了一种跳跃机器人,其采用由直接驱动传动系统驱动的单条长腿,在没有弹性储能的情况下,实现了高(7.6米)且快(190毫秒支撑时间)的跳跃。该机器人重281克,实现了迄今为止有电动执行、无弹簧系统中报道的最高跳跃高度。该腿采用一种新型织物缠绕传动装置,可提供可变的机械优势,使小型电机在整个行程中保持接近其峰值功率输出,同时支撑细长轻质的腿部以防止屈曲。腿部顶部的平衡模块使用小型螺旋桨控制腿部在地面和空中的姿态,而长腿为控制力矩提供了较大的力臂。该机器人在户外通过垂直跳跃、地面和飞行中的姿态控制以及倾斜跳跃进行了验证。
cs.RO / 15 / 2609.30533

ST-pRRTC: Parallel Space-Time RRT-C with Adaptive Goal-Time Forests

ST-pRRTC:具有自适应目标时间森林的并行时空RRT-Connect规划器
Zhang, Duo, Li, Jintong, Huang, Junshan, Yu, Jingjin
Abstract
We propose ST-pRRTC, a GPU-parallel space- time RRT-Connect motion planner for problems with known obstacle trajectories and unspecified arrival time. Searching over many arrival times broadens temporal coverage but divides a finite planning budget among more backward trees. To address the challenge, ST-pRRTC builds a shared forward tree and an adaptive forest of backward goal-time trees. Its interval root formulation samples goal arrival times continuously and guarantees probabilistic completeness and asymptotic arrival- time optimality under the stated assumptions in a bounded time domain. The practical root recycling policy has no such guar- antees. It adapts a fixed number of backward trees, replacing later roots while retaining useful search progress. Experiments on three dynamic benchmarks show that both variants achieve lower mean first-solution times and earlier mean final arrivals than ST-RRT* and SI-RRT on problems solved by all compared methods. Further experiments demonstrate the benefit of recy- cling over broad arrival-time ranges. Real-robot demonstrations show root-recycling ST-pRRTC planning motions for a UR5e among moving Crazyflie quadrotors.
Chinese Translation
我们提出了ST-pRRTC,一种GPU并行的时空RRT-Connect运动规划器,适用于障碍物轨迹已知但到达时间未指定的问题。在多个到达时间上进行搜索可以扩大时间覆盖范围,但会将有限的规划预算分配给更多的反向树。为应对这一挑战,ST-pRRTC构建了一棵共享的前向树和一个由反向目标时间树组成的自适应森林。其区间根(interval root)公式可对目标到达时间进行连续采样,并在给定假设下、于有界时间域内保证概率完备性和渐近到达时间最优性。而实用的根回收(root recycling)策略则不具备此类保证。该策略自适应地维护固定数量的反向树,在保留有效搜索进展的同时替换较靠后的根。在三个动态基准上的实验表明,在所有对比方法均能求解的问题上,两种变体的平均首次求解时间更低、平均最终到达时间更早。进一步的实验证明了根回收策略在宽到达时间范围内的优势。真实机器人演示展示了根回收版ST-pRRTC在运动的Crazyflie四旋翼飞行器之间为UR5e机械臂规划运动的能力。
cs.RO / 16 / 2609.30543

GraspTwin: Zero-Shot Task-Oriented Grasp Optimization via a Digital Twin

GraspTwin:基于数字孪生的零样本面向任务抓取优化
Evans, Daniel J., Dai, Yinlong, Stepputtis, Simon, Losey, Dylan P.
Abstract
As robots transition from structured factory settings into homes, they are required to interact with an ever-increasing variety of objects. Many tasks require grasping, and often it is not sufficient to just pick up the target object. Consider a task like "pouring coffee" --- to facilitate the subsequent pouring, the robot should grasp the mug by its handle. Existing learning-based approaches for grasping either find robust and collision-free grasps that are largely agnostic to the task (e.g., picking up the mug by its rim), or leverage foundation models to propose task-appropriate grasp locations that lack fine-grained physical grounding (e.g., reaching for and missing the handle). In this work, we bridge these approaches with a real-to-sim-to-real framework. Based on a single RGB-D observation, we construct a digital twin of the environment, query a large foundation model to propose grasps that align with the object's affordances and task description, and then optimize the proposals to ensure robustness and plausibility before executing the result on the real robot. Our key insight is that the grasp proposals of the foundation model should be regarded as semantic priors that serve as seeds for local, gradient-free optimization. We leverage Bayesian optimization with Thompson sampling to draw batches of nearby poses, which are subsequently evaluated in parallel under domain-randomized physics rollouts. The resulting grasp is both task-oriented and physically feasible for execution by the robot arm. Our full zero-shot real-world transfer only takes a few minutes and improves task-oriented grasping success by up to 33% as compared to other state-of-the-art pipelines. Our code is available here: https://github.com/VT-Collab/GraspTwin/
Chinese Translation
随着机器人从结构化的工厂环境走向家庭,它们需要与种类日益增多的物体进行交互。许多任务都需要抓取,而且往往仅仅拿起目标物体是不够的。以“倒咖啡”这类任务为例——为了便于后续倾倒,机器人应当握住马克杯的把手。现有的基于学习的抓取方法,要么寻找与任务基本无关的鲁棒且无碰撞的抓取方式(例如,捏住杯沿拿起马克杯),要么利用基础模型提出适合任务的抓取位置,但缺乏细粒度的物理基础(例如,伸向把手却未能抓住)。在本工作中,我们通过一个真实—仿真—真实(real-to-sim-to-real)框架来融合这两类方法。基于单次RGB-D观测,我们构建环境的数字孪生,查询大型基础模型以提出与物体可供性(affordance)和任务描述相符的抓取方案,然后对这些方案进行优化以确保鲁棒性与合理性,最后将结果在真实机器人上执行。我们的核心洞察是:基础模型给出的抓取方案应被视为语义先验,作为局部无梯度优化的种子。我们利用基于Thompson采样的贝叶斯优化,批量采样邻近位姿,并在域随机化的物理仿真推演中并行评估这些位姿。所得的抓取既面向任务,又对机械臂执行而言物理上可行。我们完整的零样本真实世界迁移只需几分钟,与其它最先进的流程相比,面向任务的抓取成功率最高提升33%。我们的代码发布于:https://github.com/VT-Collab/GraspTwin/
cs.RO / 17 / 2609.30554

Privacy-Preserving Prompted Policy Search for Robotic Control

面向机器人控制的隐私保护提示策略搜索
Irshayyid, Ali, Lin, Feng, Li, Chong, Chen, Jun
Abstract
Large language models (LLMs) have recently demonstrated promising capabilities as in-context policy optimizers for Reinforcement Learning (RL), enabling policy search driven by both numerical reward signals and natural language reasoning. However, deploying such methods in practice requires transmitting raw policy parameters and rewards history to cloud-based LLM APIs, exposing proprietary control strategies to third-party service providers. To address this issue, this paper introduces Privacy-Preserving Prompted Policy Search (PP-ProPS), a framework that enables LLM-guided policy optimization while keeping policy and environmental parameters confidential. PP-ProPS encodes policy parameters and reward values using secret client-side transformations before they are included in each API request, ensuring that the LLM provider observes only encoded policy parameters and scaled reward information. Furthermore, unlike Vanilla ProPS, the proposed framework does not require the true optimal episodic return to be known or disclosed to the LLM. Beyond protecting the optimization data, PP-ProPS improves the search process in two ways. First, it provides the LLM with individual reward components instead of only a single total return, offering more informative feedback about each candidate policy. Second, it uses a bounded history that prevents the prompt from growing indefinitely, improving search with high-dimensional policies and supporting the use of open-weight LLMs. The proposed PP-ProPS is evaluated on both continuous and discrete control problems spanning Multi-Joint dynamics with Contact (MuJoCo) locomotion, classic control, highway driving, and robotic arm manipulation. Compared to Vanilla ProPS, the proposed PP-ProPS outperforms ProPS in seven of the ten evaluated tasks, and surpasses conventional RL methods including PPO, SAC, and TRPO, in five of the six tasks.
Chinese Translation
大语言模型(LLMs)近期作为强化学习(RL)的上下文策略优化器展现出令人期待的能力,使策略搜索能够同时由数值奖励信号和自然语言推理驱动。然而,在实际部署此类方法时,需要将原始策略参数和奖励历史传输给基于云端的LLM API,从而使专有控制策略暴露给第三方服务提供商。为解决这一问题,本文提出了隐私保护提示策略搜索(Privacy-Preserving Prompted Policy Search, PP-ProPS),这是一个能够在策略和环境参数保密的前提下实现LLM引导策略优化的框架。PP-ProPS在将策略参数和奖励值纳入每次API请求之前,使用客户端秘密变换对其进行编码,确保LLM提供商仅能观察到编码后的策略参数和缩放后的奖励信息。此外,与原始的Vanilla ProPS不同,所提出的框架无需知道或向LLM披露真实的最佳回合回报。除了保护优化数据之外,PP-ProPS还从两个方面改进了搜索过程。首先,它向LLM提供各个奖励分量而非仅提供单一总回报,从而为每个候选策略提供更具信息量的反馈。其次,它采用有界历史记录机制,防止提示词无限增长,改善了高维策略下的搜索效果,并支持使用开源权重LLM。所提出的PP-ProPS在连续和离散控制问题上进行了评估,涵盖多关节接触动力学(MuJoCo)运动、经典控制、高速公路驾驶以及机械臂操作。与Vanilla ProPS相比,所提出的PP-ProPS在十个评估任务中的七个任务上优于ProPS,并在六个任务中的五个任务上超越了包括PPO、SAC和TRPO在内的传统RL方法。
cs.RO / 18 / 2609.30557

Auditing Latent-Space Monitors for Autonomous Driving

自动驾驶潜空间监测器的审计
Advani, Nikhil Kamalkumar, Hogale, Vishwajeet Shivaji, Kumar, Saurav
Abstract
Runtime failure monitors can use a model's internal representations to anticipate failures. We audit this monitoring strategy across two autonomous-driving tasks: online vectorized map generation with LaneSegNet and end-to-end planning with VAD. We find that frame-level errors are predictable at inference in both tasks. For LaneSegNet, a supervised latent probe reaches Area Under the Receiver Operating Characteristic curve (AUROC) 0.780 for high Chamfer error; to our knowledge, this is the first post-hoc frame-level failure monitor for online vectorized map generation. For VAD, a supervised planning-latent probe reaches AUROC 0.868 for mean-ADE failure. Our audit shows that internal access is not necessary for strong failure prediction. A monitor using only LaneSegNet's prediction outputs reaches AUROC 0.825, while for VAD, ego state, driving command, and the planner's predicted trajectory reach 0.924 on the same mean-ADE endpoint. Adding latent features to either baseline yields no statistically resolved improvement. This observation persists across a broad suite of planning failure endpoints, including endpoints whose labels depend on geometry unavailable to the non-latent baseline. Thus, predicting failure from an internal representation does not establish that the representation provides useful information beyond observable inputs and outputs. We propose an evaluation protocol for testing the incremental value of latent access and release our per-frame failure endpoint labels.
Chinese Translation
运行时故障监测器可以利用模型的内部表征来预判故障。我们在两个自动驾驶任务上对这一监测策略进行了审计:基于 LaneSegNet 的在线矢量化地图生成,以及基于 VAD 的端到端规划。我们发现,在两个任务中,帧级错误在推理阶段都是可预测的。对于 LaneSegNet,一个有监督的潜空间探针在高 Chamfer 误差的预测上达到受试者工作特征曲线下面积(AUROC)0.780;据我们所知,这是首个针对在线矢量化地图生成的事后帧级故障监测器。对于 VAD,一个有监督的规划潜空间探针在平均 ADE 失效的预测上达到 AUROC 0.868。我们的审计表明,访问内部表征对于实现强大的故障预测并非必要。仅使用 LaneSegNet 预测输出的监测器即可达到 AUROC 0.825;而对于 VAD,自车状态、驾驶指令和规划器预测轨迹在同一平均 ADE 端点上达到 0.924。在任一基线上添加潜空间特征均未带来具有统计显著性的提升。这一观察结果在多种规划失效端点上都成立,包括那些标签依赖于非潜空间基线无法获取的几何信息的端点。因此,基于内部表征预测失效并不能证明该表征在可观测的输入和输出之外提供了有用的信息。我们提出了一种用于检验潜空间访问增量价值的评估协议,并公开了我们的逐帧失效端点标签。
cs.RO / 19 / 2609.30560

SoGuDiff: Socially Guided Diffusion for Steerable, Norm-Grounded Robot Navigation

SoGuDiff:面向可引导、基于社会规范的机器人导航的社会引导扩散模型
Schaible, Christian, Ji, Haoran, Pant, Yash Vardhan, Smith, Stephen L.
Abstract
Beyond collision avoidance, socially competent robot navigation requires adherence to implicit social conventions that vary across contexts, cultures, and deployment requirements. Many conventional navigation policies learn a single normative behavior, either through reinforcement learning against a fixed reward function or imitation of human demonstrations, exposing no interface for adjusting that conduct at runtime. We present a diffusion-based navigation framework whose social behavior can be tuned at deployment: a desired style is specified, such as how closely the robot passes, which side it yields to, or how much it defers to groups, and the planner adapts accordingly. Continuous style axes can be followed independently or composed, spanning a behavioral space rather than discrete, primitive-based specifications. A feasibility projection layer separates learned social behavior from kinematic feasibility and collision avoidance. A single-axis sweep illustrates a tradeoff curve that strictly dominates the evaluated fixed-behavior baseline configurations, and stylistic differences are replicated in real-world demonstrations.
Chinese Translation
超越避碰之外,具备社交能力的机器人导航还需遵循隐性的社会惯例,而这些惯例因情境、文化和部署需求而异。许多传统导航策略通过针对固定奖励函数的强化学习或对人类示范的模仿,只学习单一的规范行为,且在运行时没有可供调整该行为的接口。我们提出一种基于扩散模型的导航框架,其社会行为可在部署时进行调节:用户可指定期望的风格,例如机器人通过时保持的距离、让行时选择的哪一侧,以及对人群的礼让程度,规划器随之相应调整。连续的风格维度可被独立跟踪或组合使用,从而覆盖一个行为空间,而非仅限于离散的、基于基元的规范。一个可行性投影层将学习到的社会行为与运动学可行性及碰撞规避分离。单维度扫描实验展示了一条严格优于所评估的固定行为基线配置的权衡曲线,且这些风格差异在真实世界演示中得到了复现。
cs.RO / 20 / 2609.30570

MOCHA: Multi-Objective Co-Design using Hypernetwork Architectures

MOCHA:基于超网络架构的多目标协同设计
Madabushi, Varun, Janwani, Neil, Tucker, Maegan
Abstract
In this work, we present MOCHA, the first, to our knowledge, reinforcement learning based approach to computing a family of Pareto-optimal policies across the design space of a robot using a single network. Specifically, MOCHA leverages the hypernetwork architecture to learn a network that produces specialized network parameters optimized for a given objective and parameterized robot design; we term this a multi-objective design hypernetwork (MDH). We demonstrate the capabilities of MDHs to represent a complex family of design-dependent strategies on two distinct robot morphologies, each with six design dimensions and across 2-3 objectives. Moreover, we propose an approach for efficiently producing a Design Pareto set using evolutionary search of the learned policy network, generating the optimal design-policy combination for each objective prioritization. Lastly, we provide an efficient method for computing generalist robot designs which achieve the best cumulative performance across the entire set of objectives.
Chinese Translation
在本工作中,我们提出了MOCHA,据我们所知,这是首个基于强化学习的方法,能够利用单个网络在机器人的设计空间中计算出一族帕累托最优策略。具体而言,MOCHA利用超网络(hypernetwork)架构学习一个网络,该网络可为给定的目标和参数化的机器人设计生成经过优化的专用网络参数;我们将其称为多目标设计超网络(Multi-objective Design Hypernetwork,MDH)。我们在两种不同的机器人形态上演示了MDH表示复杂的、依赖设计的策略族的能力,每种形态具有六个设计维度,并涉及2至3个目标。此外,我们提出了一种通过对学习到的策略网络进行进化搜索来高效生成设计帕累托前沿(Design Pareto set)的方法,为每种目标优先级生成最优的设计-策略组合。最后,我们提供了一种高效的方法来计算通用型机器人设计,使其在整个目标集合上实现最佳累积性能。
cs.RO / 21 / 2609.30594

HuGo: LLMs as Whole-Body Policy Code Designers for Humanoid Loco-Manipulation

HuGo:以大语言模型作为全身策略代码设计器的人形机器人移动-操作方法
Choi, Seoyeon, Ye, Shizhao, Bui, Nicholas, Shrivastava, Aayushi, Ryu, Kanghyun, Tirumala, Dhruva, Wulfmeier, Markus, Mehr, Negar
Abstract
For humanoids to be useful in everyday environments, they must perform a wide range of tasks that couple locomotion and manipulation. Existing approaches commonly acquire a loco-manipulation policy through reward engineering or demonstrations followed by task-specific training, making it costly to scale to new tasks. In this work, we propose a hierarchical approach to humanoid loco-manipulation that eliminates these per-task requirements. HuGo, Humanoid policy code Generation, uses a Large Language Model (LLM) to generate executable, closed-loop high-level policy code from a task description on top of a frozen low-level whole-body policy. Given the task, observation, and command specifications, the LLM constructs the task logic in code. HuGo then refines the policy from its rollouts using numerical trajectories and selected video frames to produce feedback and targeted code updates. Across five simulation tasks, using two different low-level policies, HuGo substantially outperforms a high-level reinforcement learning baseline and approaches the performance of a demonstration-based baseline. We achieve this level of performance without task-specific reward design or demonstration collection. We further demonstrate zero-shot transfer of simulation-generated policies to hardware and show that applying the same refinement loop to real-world rollouts can further improve transfer performance without expert demonstrations or policy retraining. Project website is https://iconlab.negarmehr.com/HuGo/
Chinese Translation
要使人形机器人在日常环境中发挥作用,它们必须能够执行将移动(locomotion)与操作(manipulation)相耦合的多种任务。现有方法通常通过奖励工程或演示示范来获取移动-操作策略,然后进行任务特定的训练,导致扩展到新任务的成本很高。在本工作中,我们提出了一种分层的类人机器人移动-操作方法,消除了这些针对每个任务的需求。HuGo(Humanoid policy code Generation,人形策略代码生成)利用大语言模型(LLM)在冻结的低层全身策略之上,根据任务描述生成可执行的高层闭环策略代码。给定任务、观测和命令规范,LLM 以代码形式构建任务逻辑。随后,HuGo 利用策略回放(rollout)产生的数值轨迹和选取的视频帧来生成反馈,并进行有针对性的代码更新,从而对策略进行优化。在五项仿真任务中,使用两种不同的低层策略,HuGo 的性能显著优于高层强化学习基线,并接近基于演示的基线的性能。我们在无需任务特定奖励设计或演示收集的情况下实现了这一性能水平。我们进一步展示了仿真生成策略向硬件的零样本迁移,并证明对真实世界回放应用相同的优化循环,可以在无需专家演示或策略重训练的情况下进一步提升迁移性能。项目网站:https://iconlab.negarmehr.com/HuGo/
cs.RO / 22 / 2609.30599

Learning-Accelerated Narrow-Phase Collision Detection via Check Ordering for Sampling-Based Motion Planning

通过检查排序学习加速窄相碰撞检测以提升基于采样的运动规划
Jiang, Hao, Wang, Yinghan, He, Jianping, Duan, Xiaoming
Abstract
Collision detection is critical for ensuring the safety of planned paths. However, it imposes a non-negligible computational burden on motion planners, motivating extensive studies on collision-detection acceleration. In commonly used phase-based collision-detection methods, the broad phase employs hierarchical structures to rapidly discard object pairs that are clearly collision-free, while the subsequent narrow phase performs detailed collision checks on the remaining object pairs whose collision status cannot be determined by the broad phase. Although these methods effectively reduce the number of detailed checks through broad-phase pruning, the narrow phase is usually executed in the default order returned by the broad phase, with little explicit optimization of the check order. This leaves room for further acceleration, especially in cluttered environments where many object pairs may remain after the broad phase and the narrow phase can account for a significant portion of the total detection time. In this work, we propose a learning-based method to accelerate phase-based collision detection by optimizing the check order in the narrow phase. We first formulate the expected time cost of the narrow phase and derive an optimal check-ordering criterion that minimizes this expectation. Since the priors required by this criterion are difficult to obtain in advance, we design a hypernetwork-based model to predict collision probabilities, which are then used to approximate the optimal check order. The resulting order guides the execution of exact mesh checks in the narrow phase, thereby reducing detection time without replacing the underlying geometric collision checker. Simulation results show that our method effectively accelerates phase-based collision detection and improves the efficiency and success rate of sampling-based motion planning, especially in cluttered environments.
Chinese Translation
碰撞检测对于确保规划路径的安全性至关重要。然而,它给运动规划器带来了不可忽视的计算负担,因此激发了大量关于碰撞检测加速的研究。在常用的基于阶段(phase-based)的碰撞检测方法中,宽相(broad phase)利用层次结构快速剔除明显无碰撞的物体对,而后续的窄相(narrow phase)则对宽相无法判定碰撞状态的剩余物体对进行详细的碰撞检查。尽管这些方法通过宽相剪枝有效减少了详细检查的次数,但窄相通常按照宽相返回的默认顺序执行,很少对检查顺序进行显式优化。这为进一步加速留下了空间,尤其是在杂乱环境中,宽相之后可能残留大量物体对,窄相可能占据总检测时间的很大一部分。在本工作中,我们提出了一种基于学习的方法,通过优化窄相中的检查顺序来加速基于阶段的碰撞检测。我们首先对窄相的期望时间成本进行建模,并推导出使该期望最小化的最优检查排序准则。由于该准则所需的先验信息难以事先获得,我们设计了一个基于超网络(hypernetwork)的模型来预测碰撞概率,进而利用这些概率逼近最优检查顺序。所得顺序指导窄相中精确网格检查的执行,从而在不替换底层几何碰撞检测器的情况下减少检测时间。仿真结果表明,我们的方法有效加速了基于阶段的碰撞检测,并提高了基于采样的运动规划的效率与成功率,尤其在杂乱环境中表现突出。
cs.RO / 23 / 2609.30608

Audit Before You Commit: Locating Belief Failures in Active Identification for One-Shot Manipulation

提交前先审计:定位单次操作主动识别中的信念失效
Abouagour, Mohamed, Min, Byung-Cheol
Abstract
A robot that probes a few times before one irreversible action, such as tapping a surface before inserting a peg, must decide when the evidence is enough to commit. We argue that this decision rests on two conditions that existing methods do not separate: the belief must still cover the truth in the coordinate that decides the action, and the failure model that scores actions must track realized failure. We audit both conditions separately, offline and with ground truth, on a deployed probe-then-commit pipeline: a particle belief, a scenario failure score, and one commit. On simulated insertion, more taps sharpen the belief while the truth leaves its support on 16.9% of episodes and the failure score turns optimistic by 0.31. Conformal calibration restores coverage but not the decision: confidently wrong instances still pass a confidence gate. The audit's signatures instead point at the observation model, where a hand scan finds a 2.1 mm error in the tap boundary. Correcting that one number cuts failure from 0.354 to 0.112 on untouched instances and transfers unrefitted to a second engine, while in a third engine the same audit suggests an execution-model mismatch instead. Across seven task families in three engines, a few probes at a fixed executor reduce miss or failure. On a physical arm inserting a tool into a rigid pocket by touch, the gain and the audit's two conditions reproduce, and replaying the recorded taps under an injected model error shows the audit's signature on real data. Additional materials are available at https://sites.google.com/view/auditbeforeyoucommit.
Chinese Translation
机器人在执行一次不可逆动作之前进行若干次探测(例如在插入销钉之前先敲击表面),必须判断证据何时足以支撑最终决策。我们认为,这一决策依赖于两个现有方法未曾区分开的条件:信念必须在决定动作的坐标系中仍然覆盖真值,并且用于评估动作的失效模型必须与实际发生的失效相一致。我们在一个已部署的“先探测后提交”流水线上(包含粒子信念、场景失效评分和单次提交),离线并借助真值分别对这两个条件进行审计。在仿真插入任务中,更多次敲击使信念不断收敛,但真值在16.9%的回合中脱离了信念支撑集,且失效评分变得过于乐观,偏差达0.31。保形(Conformal)校准虽恢复了覆盖率,却无法改变决策:自信但错误的实例仍能通过置信度门限。相反,审计的信号指向观测模型——人工排查发现敲击边界存在2.1毫米的误差。仅修正这一个数值,就使未改动实例上的失败率从0.354降至0.112,且无需重新拟合即可迁移到第二个仿真引擎;而在第三个引擎中,同样的审计则指向执行模型的失配。在三个引擎共七个任务族中,固定执行器下的少量探测均能降低遗漏率或失败率。在物理机械臂通过触觉将工具插入刚性凹槽的实验中,性能增益与审计的两个条件均得到复现,并在注入模型误差的情况下回放记录的敲击数据,验证了审计信号在真实数据上的表现。更多资料见 https://sites.google.com/view/auditbeforeyoucommit。
cs.RO / 24 / 2609.30626

Frequency-Modulated Piezoelectric Haptic Display

频率调制的压电触觉显示器
Liang, Boyuan, Sun, Lingfeng, Tomizuka, Masayoshi
Abstract
We present a frequency-modulated (FM) haptic display based on piezoelectric vibrating actuators. Existing haptic displays commonly encode haptic intensity through the deformation amplitude of individual haptic pixels. Although amplitude-modulated (AM) approaches have enabled compact haptic pixels, independently controlling the deformation amplitude of a large number of pixels can require increasingly complex and bulky driving systems, posing challenges for scaling toward high-density, large-area wearable displays. To address this scaling challenge, we investigate an FM design principle in which haptic intensity is encoded through vibration frequency. We further develop a \textit{Shared-Source Frequency Modulation} (SSFM) structure in which multiple haptic pixels are powered by a common power amplifier while their vibration spectrum are controlled individually, reducing the need for independent high-power amplification at each pixel. A proof-of-concept piezoelectric haptic display was built and evaluated on rendering spatial and temporal haptic patterns through volunteer tests. The results show that participants reliably distinguished spatial and temporal patterns encoded using FM principles within the investigated operating range. These findings demonstrate the feasibility of FM-based distributed haptic rendering and suggest a potential pathway toward more compact driving architectures for future high-density, large-area wearable haptic displays.
Chinese Translation
我们提出了一种基于压电振动执行器的频率调制(FM)触觉显示器。现有的触觉显示器通常通过单个触觉像素的变形幅度来编码触觉强度。尽管幅度调制(AM)方法已经实现了紧凑的触觉像素,但独立控制大量像素的变形幅度可能需要日益复杂和庞大的驱动系统,这给向高密度、大面积可穿戴显示器的扩展带来了挑战。为应对这一扩展难题,我们研究了一种通过振动频率编码触觉强度的FM设计原理。我们进一步开发了一种共享源频率调制(Shared-Source Frequency Modulation, SSFM)结构,其中多个触觉像素由一个共同的功率放大器供电,同时它们的振动频谱可被独立控制,从而减少了对每个像素独立高功率放大的需求。我们构建了一个概念验证型压电触觉显示器,并通过志愿者测试评估了其在渲染空间和时间触觉模式方面的性能。结果表明,在所研究的操作范围内,参与者能够可靠地区分采用FM原理编码的空间和时间模式。这些发现证明了基于FM的分布式触觉渲染的可行性,并为未来高密度、大面积可穿戴触觉显示器实现更紧凑的驱动架构提供了一条潜在途径。
cs.RO / 25 / 2609.30644

MR. POP: Multi-Robot Parallel Optimizing Planner for Almost-Surely Asymptotically Optimal Planning

MR. POP:面向几乎必然渐近最优规划的多机器人并行优化规划器
Huang, Chih H., Xing, Roy, Plancher, Brian, Kingston, Zachary
Abstract
Finding globally optimal paths remains a fundamental challenge in multi-robot motion planning. Despite acceleration of almost-surely asymptotically optimal (a.s.a.o.) planners via CPU-based parallelism, achieving both probabilistic convergence guarantees and strong computational performance, these algorithms still struggle to scale to multi-robot settings. As such, we introduce MR. POP, a GPU-based a.s.a.o. multi-robot planner based on dRRT and the AO-x meta-algorithm. MR. POP uses large-scale GPU-based SIMT-parallelism to simultaneously run hundreds of roadmap construction and tree search iterations with underlying parallel nearest neighbor search and collision checking operations. We show that this enables MR. POP to become the only planner achieving a 100% solve rate while being faster than state-of-the-art a.s.a.o. planners in multi-robot systems up to 35-DOF. MR. POP also raises the success rate of downstream motion optimizers (e.g., from 4% to 72%), by creating high-quality, diverse seeds that help avoid local minima.
Chinese Translation
寻找全局最优路径仍然是多机器人运动规划中的一项根本性挑战。尽管基于CPU并行的加速技术已经提升了几乎必然渐近最优(almost-surely asymptotically optimal, a.s.a.o.)规划器的性能,使其同时具备概率收敛保证和较强的计算性能,但这些算法在扩展到多机器人场景时仍然面临困难。为此,我们提出了MR. POP,一个基于GPU的a.s.a.o.多机器人规划器,其建立在dRRT和AO-x元算法基础之上。MR. POP利用大规模的GPU SIMT并行机制,同时运行数百次路线图构建与树搜索迭代,并辅以底层并行的最近邻搜索和碰撞检测操作。结果表明,这使MR. POP成为唯一能够在最高35自由度的多机器人系统中实现100%求解率、且速度快于最先进a.s.a.o.规划器的规划器。此外,MR. POP通过生成高质量、多样化的初始种子以帮助避免局部极小值,从而显著提升了下游运动优化器的成功率(例如从4%提升至72%)。
cs.RO / 26 / 2609.30676

Can a Robot Read Braille? - Learning to Adapt Contact via Imitation Learning for Tactile Braille Recognition

机器人能阅读盲文吗?——基于模仿学习的触觉自适应接触学习用于盲文识别
Chen, Xi, Shan, Yunlong, Chen, Sihan, Hu, Jun, Li, Zhongxuan, Zhang, Shiyao, Liu, Sichao, Zhao, Zhong, Jovanovic, Kosta, Zhou, Peng
Abstract
For people who are blind, touch provides an essen-tial channel for accessing written information through Braille. Bringing a similar capability to robots requires them not only to recognize tactile patterns, but also to actively establish physical contact that makes those patterns readable. Yet existing robotic Braille readers largely focus on recognition after contact, leaving contact establishment itself insufficiently addressed. We present an adaptive-contact framework for robotic tactile Braille reading that assesses contact quality and physically corrects unsuitable contact before recognition and reconstruc-tion. Multi-Head Policy Learning uses expert-guided contact-adjustment demonstrations to jointly learn contact acceptability and pose corrections. During deployment, the robot iteratively evaluates and re-establishes contact, retaining reliable tactile observations for pose-aware fusion and Braille reconstruction. Across 20 physical Braille plates used for learning and eval-uation, the proposed approach achieves 94.0% tactile quality and 88.6% tactile reconstruction on the ten online-evaluation plates. These results demonstrate the importance of actively establishing readable contact, rather than relying solely on recognition under imperfect tactile observations, for reliable robotic Braille reading.
Chinese Translation
对于盲人而言,触觉是通过盲文(Braille)获取书面信息的重要渠道。要让机器人具备类似的能力,不仅需要其识别触觉模式,还需要其主动建立使这些模式可被读取的物理接触。然而,现有的机器盲文阅读系统大多聚焦于接触后的识别,对接触建立本身的研究尚不充分。我们提出了一种用于机器触觉盲文阅读的自适应接触框架,该框架能够评估接触质量,并在识别与重建之前对不合适的接触进行物理修正。多策略头学习(Multi-Head Policy Learning)利用专家指导的接触调整演示,联合学习接触可接受性与位姿修正。在部署阶段,机器人迭代地评估并重新建立接触,为位姿感知融合与盲文重建保留可靠的触觉观测。在用于学习与评估的20块实体盲文板上,所提出的方法在十块在线评估板上实现了94.0%的触觉质量准确率和88.6%的触觉重建准确率。这些结果表明,对于可靠的机器盲文阅读而言,主动建立可读的接触至关重要,而不能仅仅依赖在不完善触觉观测下的识别。
cs.RO / 27 / 2609.30695

Multi-Objective Human-in-the-Loop Bayesian Optimization of a Lower-Limb Exoskeleton

下肢外骨骼的多目标人在环路贝叶斯优化
Janwani, Neil, Lerner, Matthew T., Young, Aaron J., Tucker, Maegan
Abstract
Human-in-the-loop optimization (HILO) is a common approach for optimizing the control of assistive devices to account for the wearer's unique biomechanics and subjective preferences. However, despite research suggesting that a person may have a different prioritization of objectives depending on time-varying factors such as the environment, their mood, or energy levels, existing HILO approaches only consider a single objective or enforce a fixed weighting on a set of objectives. Neither approach is capable of representing an individual's preferences over objectives. In this work, we propose Multi-Objective Human-in-the-loop Bayesian Optimization (MO-HILBO), which builds on explicit multi-objective Bayesian optimization to efficiently infer a personalized set of Pareto-optimal controllers. We compare our approach with an existing multi-objective HILO method and experimentally demonstrate MO-HILBO on a lower-limb exoskeleton across two objectives: metabolic cost (efficiency) and ordinal human feedback (comfort). We find that MO-HILBO (1) discovers Pareto-optimal controllers, and (2) that the pairwise ordering of points on the Pareto front itself is consistent with validation trials. Lastly, we open-source mohilo, a Python package for running both HILO and MO-HILBO on wearable devices: https://dynamicmobility.github.io/mohilo/.
Chinese Translation
人在环路优化(Human-in-the-Loop Optimization, HILO)是优化辅助设备控制的常用方法,可用于考虑穿戴者独特的生物力学特征和主观偏好。然而,尽管有研究表明,个体可能会根据环境、情绪或精力水平等时变因素对不同目标赋予不同的优先级,现有的HILO方法要么只考虑单一目标,要么对一组目标强制施加固定权重。这两种方法都无法表征个体在多个目标之间的偏好。在本工作中,我们提出了多目标人在环路贝叶斯优化(Multi-Objective Human-in-the-Loop Bayesian Optimization, MO-HILBO),该方法基于显式多目标贝叶斯优化,能够高效地推断出个性化的帕累托最优控制器集合。我们将所提方法与现有的多目标HILO方法进行比较,并在下肢外骨骼上针对两个目标对MO-HILBO进行了实验验证:代谢成本(效率)和序数人类反馈(舒适度)。我们发现MO-HILBO(1)能够发现帕累托最优控制器,并且(2)帕累托前沿上各点的两两排序与验证实验结果一致。最后,我们开源了mohilo——一个用于在可穿戴设备上运行HILO和MO-HILBO的Python软件包:https://dynamicmobility.github.io/mohilo/。
cs.RO / 28 / 2609.30696

Learning Vision-Based Agile Gap Traversal: Differentiable Simulation with a Warm-Started Critic

基于视觉的敏捷间隙穿越学习:结合批评家热启动的可微仿真方法
Gerdpratoom, Nuthasith, Sun, Tianchen, Gao, Yichao, Zhao, Lin
Abstract
Traversing narrow gaps is challenging for autonomous quadrotors, especially when control commands come directly from high-dimensional visual observations. Existing end-to-end methods often rely on behavior cloning or full-rollout backpropagation through time (BPTT) via differentiable simulation, which can limit policy performance or incur high training costs. We propose a two-stage reinforcement learning framework for more efficient ego-centric visuomotor gap-traversal policy training, leveraging quasi-analytical policy gradients (QPG) via differentiable simulation and critic warm-starting. The framework utilizes QPG to avoid backpropagation through visual rendering, reducing computation and memory costs while improving sample efficiency. In the first stage, an expert actor and critic are trained using privileged observations, including gap geometry. Unlike prior gap-traversal approaches, our training utilizing QPG does not require resetting the agent along optimized reference trajectories. In the second stage, a visual policy is trained using binary gap masks from two ego-centric cameras and low-dimensional observations, while its privileged critic is warm-started from the first stage. This substantially improves training efficiency and traversal success compared with cold-starting the critic or using full-rollout BPTT. Our framework does not require retraining the expert actor when system parameters change, enabling more efficient generalization across drone platforms than state-of-the-art visual gap-traversal methods based on action supervision. The learned visual policy also generalizes to gaps with unseen shapes. Extensive real-world experiments further demonstrate robust gap traversal using binary masks rendered online. Beyond gap traversal, the proposed framework is generic and can be extended to other visuomotor robot learning tasks.
Chinese Translation
穿越狭窄间隙对自主四旋翼无人机而言极具挑战性,尤其是当控制指令直接来自高维视觉观测时。现有的端到端方法通常依赖行为克隆或通过可微仿真进行的完整展开的时间反向传播(BPTT),这会限制策略性能或带来高昂的训练成本。我们提出了一种两阶段强化学习框架,以更高效地训练以自我为中心的视觉运动间隙穿越策略,该框架利用基于可微创仿真的准解析策略梯度(QPG)以及批评家热启动。该框架利用QPG避免了对视觉渲染过程的反向传播,在降低计算与内存成本的同时提升了样本效率。在第一阶段,使用包含间隙几何信息在内的特权观测训练专家Actor与Critic。与以往的间隙穿越方法不同,我们基于QPG的训练无需沿优化的参考轨迹重置智能体。在第二阶段,视觉策略使用来自两个自我中心相机的二值间隙掩码以及低维观测进行训练,同时其特权Critic从第一阶段热启动。与冷启动Critic或使用完整展开BPTT相比,这显著提升了训练效率与穿越成功率。当系统参数发生变化时,我们的框架无需重新训练专家Actor,相比基于动作监督的最先进视觉间隙穿越方法,能够更高效地在不同无人机平台之间泛化。所学得的视觉策略还能泛化至形状未知的间隙。大量真实世界实验进一步证明了利用在线渲染的二值掩码实现鲁棒间隙穿越的能力。除间隙穿越之外,所提出的框架具有通用性,可扩展至其他视觉运动机器人学习任务。
cs.RO / 29 / 2609.30704

From Visual Search to Movement Control: A Priority Field for Artificial Agents

从视觉搜索到运动控制:面向人工智能体的优先级场
Zhang, Han, Cao, Zhong
Abstract
Human spatial attention is widely conceptualized as being guided by a priority map that integrates perceptual salience, current goals, and past experiences. Here, we extend priority-based computation to movement control in artificial agents. We first introduce a lightweight model of visual search based on an integrated priority map. Trained on human saccades, it reproduced key behavioral patterns, including oculomotor suppression and history-driven selection. Extending the search model, we equipped an artificial agent with a priority field and evaluated its performance in a reach-avoid task that required reaching a goal destination while avoiding moving obstacles. Compared with alternative architectures, priority-field agents trained more efficiently and performed better in unseen, complex scenarios, even from simple demonstrations. Adding a simple memory mechanism also produced human-like, history-driven effects in anticipating the likely location of the upcoming goal. These findings suggest that priority-based computation may provide a promising foundation for movement control in artificial agents.
Chinese Translation
人类的空间注意通常被概念化为由一个优先级图引导,该图整合了感知显著性、当前目标和过往经验。本研究将基于优先级的计算扩展至人工智能体的运动控制。我们首先提出了一个基于整合优先级图的轻量级视觉搜索模型。该模型在人类眼跳数据上训练后,再现了关键的行为模式,包括眼动抑制和基于历史的选择。在搜索模型的基础上,我们为人工智能体配备了一个优先级场,并在一个触及-回避任务中评估其表现,该任务要求智能体在避开移动障碍物的同时到达目标位置。与替代架构相比,基于优先级场的智能体训练效率更高,在未见过的复杂场景中表现更佳,即便仅通过简单的演示样本也能如此。加入一个简单的记忆机制后,还能产生类人的、由历史经验驱动的效应,用于预测下一个目标可能出现的位置。这些发现表明,基于优先级的计算可能为人工智能体的运动控制提供一个有前景的基础。
cs.RO / 30 / 2609.30715

RoboMonitor: Label-Efficient Runtime Monitoring of Robot Task Execution via Predictive Representation Learning

RoboMonitor:基于预测性表征学习的机器人任务执行标签高效运行时监控
Ajith, Abhiroop, Narayanan, Gokul, Coelho, Kyle, Zhao, Tingji, Shahapurkar, Yash, Zhu, Brian, Erdogan, Melih, Krubasik, Ted, Chamzas, Constantinos, Solowjow, Eugen
Abstract
Learned robot policies produce actions, but their outputs alone do not establish whether execution is progressing as intended. Robot execution monitoring requires identifying the current execution phase, detecting failures, and recognizing task completion from observations available during execution. Training such monitors requires annotations that are scarce in datasets collected for robot-policy learning. We present RoboMonitor, a label-efficient vision--language execution monitor that learns from these datasets before introducing monitoring supervision. We pre-train on 25 hours of multi-camera trajectories spanning 12 manipulation tasks and two robot embodiments, using action-conditioned future-feature prediction, inverse dynamics, and masked-present prediction. We then transfer the learned visual and context encoders to a causal monitor and apply temporal supervised fine-tuning (Temporal SFT), which combines supervision throughout each observation window with consistency objectives within and across overlapping windows. At deployment, RoboMonitor requires only the task instruction and camera observations. On a four-task monitoring benchmark, RoboMonitor trained with 52 labeled episodes achieves 93.1% mean phase accuracy and 85.9% macro recall over two fine-tuning seeds, exceeding Qwen3-VL and Robometer trained with the same monitoring supervision. Its phase accuracy also exceeds that of both Qwen3-VL and Robometer trained with 100 episodes. A Qwen3-VL ablation shows that Temporal SFT reduces mean spurious phase switching from 15.23% to 4.95%. In closed-loop deployment, the integrated system completes 39 of 40 simulated Toolbox Sorting trials and 35 of 40 real-world Reel Packing trials, with no false recovery triggers observed.
Chinese Translation
学习得到的机器人策略会产生动作,但仅凭其输出无法判断执行过程是否按预期进行。机器人执行监控需要识别当前的执行阶段、检测故障,并从执行过程中可获得的观测中识别任务完成。训练此类监控器需要标注数据,而这类标注在为机器人策略学习收集的数据集中十分稀缺。我们提出RoboMonitor,一个标签高效的视觉-语言执行监控器,它在引入监控监督之前先从这些数据集中学习。我们在涵盖12个操作任务和两种机器人本体的25小时多相机轨迹上进行预训练,采用动作条件下的未来特征预测、逆动力学和掩码当前帧预测。随后,我们将学到的视觉编码器和上下文编码器迁移到一个因果监控器上,并应用时序监督微调(Temporal SFT),该方法将每个观测窗口内的监督与重叠窗口内及跨窗口的一致性目标相结合。在部署时,RoboMonitor仅需任务指令和相机观测。在一个包含四个任务的监控基准上,仅用52个标注片段训练、并在两个微调随机种子下平均的RoboMonitor达到93.1%的平均阶段准确率和85.9%的宏召回率,超过了使用相同监控监督训练的Qwen3-VL和Robometer。其阶段准确率也超过了使用100个片段训练的Qwen3-VL和Robometer。Qwen3-VL的消融实验表明,Temporal SFT将平均虚假阶段切换率从15.23%降至4.95%。在闭环部署中,集成系统在40次模拟工具箱分拣(Toolbox Sorting)试验中完成39次,在40次真实世界卷轴打包(Reel Packing)试验中完成35次,且未观察到错误的恢复触发。
cs.RO / 31 / 2609.30735

Praxis: Distilling Physical Interaction Priors from Egocentric Videos for Generalizable Whole-Body Manipulation

Praxis:从第一人称视角视频中蒸馏物理交互先验以实现可泛化的全身操作
He, Shuliang, Xu, Ruiyan, Yue, Bo, Zhang, Hengming, Zhou, Huayi, Wang, Shuai, Zheng, Wei-Shi, Liu, Guiliang
Abstract
Mobile humanoid manipulation requires both reaching a usable workspace and preserving precise hand-object interactions as object poses and contact conditions change. Learning these behaviors from limited task-specific data remains challenging. To bridge this gap, we introduce Praxis, a whole-body manipulation framework that combines physical interaction priors from one-shot egocentric video demonstrations with closed-loop posture calibration and online perception. The framework coordinates three stages: vision-language-guided navigation toward target objects, closed-loop posture calibration to align the arm-hand workspace, and dexterous manipulation with synchronized upper- and lower-body control. Online visual feedback re-grounds demonstrated interaction geometry under new object poses and scene configurations, while tactile feedback adapts hand motions to actual contact conditions. Each manipulation skill is specified by one human demonstration, without task-specific manipulation-policy retraining. Experiments across five long-horizon manipulation tasks demonstrate spatial, visual, and cross-object generalization, as well as recovery from external physical disturbances across all three stages.
Chinese Translation
移动人形机器人操作既需要到达可用的工作空间,又需要在物体位姿和接触条件变化时保持精确的手-物体交互。从有限的任务特定数据中学习这些行为仍然具有挑战性。为弥合这一差距,我们提出了Praxis,一个将来自单次第一人称视角视频演示的物理交互先验与闭环姿态校准和在线感知相结合的全身操作框架。该框架协调三个阶段:基于视觉-语言引导朝目标物体导航、用于对齐手臂-手部工作空间的闭环姿态校准,以及上肢与下肢同步控制的灵巧操作。在线视觉反馈在新的物体位姿和场景配置下重新定位所演示的交互几何关系,而触觉反馈则使手部动作适应实际的接触条件。每项操作技能仅需一次人类演示即可指定,无需针对特定任务重新训练操作策略。在五项长时程操作任务上的实验展示了空间、视觉和跨物体泛化能力,以及在所有三个阶段中对外部物理干扰的恢复能力。
cs.RO / 32 / 2609.30745

Anatomy-Aware Dexterity-Driven Design Optimization of Surgical Continuum Robots

基于解剖感知与灵巧性驱动的手术连续体机器人设计优化
Qin, Tony, Connor, Peter, Dang, Khoa, Hatch, Carter, Rucker, Caleb, Webster III, Robert J., Alterovitz, Ron
Abstract
Performing complex medical procedures with continuum robots requires careful selection of their geometric design parameters. The robot should have high dexterity in the specific anatomical environment of its procedure. This work presents a design optimization method that considers both dexterity and anatomy. We introduce the Reachable Volumetric Dexterous Solid Angle (RVDSA) metric as our objective, which measures the ability of a robot's end effector to reach the points in a goal volume from different directions via collision-free paths from a start configuration. We present a computationally efficient motion planner to compute this objective function for a given robotic design, and we use an asymptotically optimal simulated annealing optimizer to compute an optimized design. We applied our new method to optimize the design of a bimanual dexterous sheaths robot for performing procedures on cancerous polyps in colon anatomies, achieving a 78% higher RVDSA on average than optimizing for 3D voxel coverage alone.
Chinese Translation
使用连续体机器人执行复杂的医疗手术需要仔细选择其几何设计参数。机器人应在其手术所涉及的特定解剖环境中具备高灵巧性。本工作提出了一种同时考虑灵巧性与解剖结构的设计优化方法。我们引入了可达体积灵巧立体角(Reachable Volumetric Dexterous Solid Angle,RVDSA)指标作为优化目标,该指标衡量机器人末端执行器从起始构型出发,经由无碰撞路径从不同方向到达目标体积内各点的能力。我们提出了一种计算高效的运动规划器来计算给定机器人设计的该目标函数,并采用渐近最优的模拟退火优化器来计算优化后的设计。我们将该新方法应用于优化一种双臂灵巧鞘管机器人(bimanual dexterous sheaths robot)的设计,用于在结肠解剖结构中对癌性息肉进行手术操作,相比仅优化三维体素覆盖的方法,平均RVDSA提高了78%。
cs.RO / 33 / 2609.30759

Design and Characterization of a Variable-Length Continuum Mechanism with Force Locking

一种具有力锁定功能的可变长度连续体机构的设计与特性表征
King, Katelyn, Fish, Veronica, Okamura, Allison M.
Abstract
The utility of flexible continuum mechanisms for dexterous navigation is often impaired by their low stiffness, making them ineffective at manipulation in high-force scenarios. To address this challenge, we propose a novel continuum mechanism that achieves both flexible and rigid behavior by antagonistic extension and contraction of a rod-driven continuum helical structure. The helical design combines variable-length capacity with force locking for workspace and stiffness enhancement. In this article, we present the detailed design of the proposed mechanism and characterize its performance through experiments that quantify bending and stiffness. The results demonstrate 180 degree bending range of motion with an average distal positioning error of <10%. Further tests demonstrate that force locking directly improves axial stiffness and thus indirectly increases bending stiffness anisotropically, with maximum bending stiffness along load paths with a large axial component. Tensioning the driving rods provides additional stiffness tunability in the force-locked state, where increasing rod tension proportionally increases bending stiffness with a dimensionless gain of 0.56.
Chinese Translation
柔性连续体机构在灵巧导航方面的应用常因其低刚度而受限,使其在高力场景下的操作中效果不佳。为应对这一挑战,我们提出了一种新型连续体机构,通过对杆驱动的连续体螺旋结构进行拮抗式伸展与收缩,实现柔性与刚性两种行为模式。该螺旋设计将可变长度能力与力锁定功能相结合,以扩展工作空间并增强刚度。本文详细介绍了所提机构的设计,并通过量化弯曲性能和刚度的实验对其性能进行了表征。结果表明,该机构可实现180度的弯曲运动范围,远端定位误差平均小于10%。进一步的测试表明,力锁定可直接提高轴向刚度,从而各向异性地间接提高弯曲刚度,最大弯曲刚度出现在轴向分量较大的载荷路径上。在力锁定状态下,张紧驱动杆可提供额外的刚度可调性,杆张力增加时弯曲刚度成比例增大,无量纲增益为0.56。
cs.RO / 34 / 2609.30770

NavGen: Visual Generative Models as a Scalable Data Engine for Embodied 3D Navigation

NavGen:将视觉生成模型作为可扩展的数据引擎用于具身三维导航
Huang, Xijie, Wan, Yongyang, Dong, Chengbin, Ding, Zimo, Zhu, Mo, Wang, Yijin, Liu, Zhiyang, Gao, Fei, Wu, Yuze, Zhou, Xin
Abstract
General-purpose robot models increasingly rely on large and diverse datasets. For embodied 3D navigation, however, existing data sources face a fundamental trade-off: simulated data can be generated at scale but often suffer from the visual sim-to-real gap, whereas real-world flight data provide realistic observations but are costly to collect. This paper studies another direction: the use of high-fidelity visual generative models as scalable data engines for embodied 3D navigation. We introduce NavGen, a text-to-video data generation pipeline that produces diverse vision-language navigation (VLN) episodes across indoor and outdoor scenes. We also propose a style-diversification method that scales up long-tail data that are difficult and costly to collect. The resulting dataset contains approximately 400K navigation episodes. We evaluate our dataset against existing UAV navigation datasets across multiple metrics, and find that the model trained on our data generally improves with scale, outperforming those trained on existing datasets. To validate real-world transferability, we deploy the trained model in world-action-model paradigm to real-world flying experiments. The final model achieves a 75\% success rate across different navigation tasks and environments.
Chinese Translation
通用机器人模型越来越依赖于大规模且多样化的数据集。然而,对于具身三维导航而言,现有数据源面临一个根本性的权衡:仿真数据可以大规模生成,但往往存在视觉上的仿真到现实(sim-to-real)差距;而真实世界的飞行数据虽然能提供真实的观测,但采集成本高昂。本文研究了另一个方向:将高保真视觉生成模型用作具身三维导航的可扩展数据引擎。我们提出了 NavGen,一种文本到视频的数据生成流水线,能够生成覆盖室内和室外场景的多样化视觉-语言导航(VLN) episodes。我们还提出了一种风格多样化方法,用于扩充难以采集且成本高昂的长尾数据。最终得到的数据集包含约40万个导航 episodes。我们在多项指标上将该数据集与现有的无人机导航数据集进行了比较评估,发现在我们的数据上训练的模型总体上随规模扩大而提升,优于在现有数据集上训练的模型。为验证向真实世界迁移的能力,我们将以世界-动作模型(world-action-model)范式训练的模型部署到真实世界的飞行实验中。最终模型在不同导航任务和环境下的成功率达到75%。
cs.RO / 35 / 2609.30777

Moving Horizon Estimation for Quadrotors: An $\mathcal{L}_1$ Adaptive Optimizer Approach

四旋翼无人机的移动时域估计:一种 $\mathcal{L}_1$ 自适应优化器方法
Nguyen, Thinh, Kim, Minkyung, Banik, Sandeep, Kim, Jinrae, Hovakimyan, Naira
Abstract
Moving Horizon Estimation (MHE) is a state estimation method based on finite-horizon optimization that can offer higher accuracy at the cost of increased computation compared to Kalman filter-based approaches. We present a linear smoothing MHE formulation as a dense Quadratic Program (QP), and a solver consisting of a continuous-time Newton's method augmented with the $\mathcal{L}_1$ Adaptive Optimizer ($\mathcal{L}_1$-AO). While MHE is inherently time-varying, conventional approaches treat it as a sequence of independent, time-invariant problems and employ iterative solvers at each time step, which can be both inaccurate and computationally burdensome. In contrast, time-varying solvers track the optimal solution with fewer iterations by exploiting the temporal evolution of the problem, thereby reducing the computational load. In this research, we enhance both the performance and efficiency of MHE through a time-varying solver with an $\mathcal{L}_1$-AO augmentation that compensates for the prediction inaccuracy, which is common in practice due to noisy sensors and the lack of prior knowledge of the system. Simulation results on a quadrotor platform show that the $\mathcal{L}_1$-AO-augmented approach solves the MHE optimization problem more efficiently than the baseline time-invariant solver and achieves higher estimation accuracy under challenging conditions, compared with both the Extended Kalman Filter and the standard MHE.
Chinese Translation
移动时域估计是一种基于有限时域优化的状态估计方法,与基于卡尔曼滤波的方法相比,它能以更高的计算量为代价提供更高的估计精度。我们提出了一种线性平滑MHE的稠密二次规划(QP)表述形式,以及一种由连续时间牛顿法并结合 $\mathcal{L}_1$ 自适应优化器($\mathcal{L}_1$-AO)构成的求解器。虽然MHE本质上是时变的,但传统方法将其视为一系列相互独立的时不变问题,并在每个时间步采用迭代求解器,这既可能不够精确,计算负担也很重。相比之下,时变求解器通过利用问题的时域演化特性,能够以更少的迭代次数跟踪最优解,从而降低计算负担。在本研究中,我们通过带有 $\mathcal{L}_1$-AO 增广的时变求解器来提升MHE的性能与效率,该增广项用于补偿预测误差——在实际应用中,由于传感器噪声和缺乏系统先验知识,预测误差十分常见。在四旋翼平台上的仿真结果表明,经 $\mathcal{L}_1$-AO 增广的方法比基线时不变求解器更高效地求解MHE优化问题,并且在具有挑战性的条件下,与扩展卡尔曼滤波器(EKF)和标准MHE相比,均实现了更高的估计精度。
cs.RO / 36 / 2609.30814

SeA-RVINS: Semantic-Aware Tightly Coupled RTK-Visual-Inertial System with Correlation-Preserving Robust Estimation for Urban Navigation

SeA-RVINS:面向城市导航的语义感知紧耦合RTK-视觉-惯性系统及保持相关性的鲁棒估计
Hu, Wang, Wu, Bo
Abstract
Reliable absolute pose estimation in urban environments is undermined by outlier measurements and incorrect temporal associations that can persist in tightly coupled estimators. Global Navigation Satellite System (GNSS) observations provide globally referenced measurements but are prone to multipath effects. Visual-inertial sensing supplies local motion constraints, but false visual associations can corrupt the estimator. We present SeA-RVINS, a fixed-lag factor-graph Real-Time Kinematic (RTK) visual-inertial system for robust urban pose estimation. A semantic-aware learned stereo frontend rejects unreliable tracks before persistent landmarks enter the graph. For double-differenced GNSS measurements, SeA-RVINS applies Dynamic Covariance Scaling through configurable batch, scalar, and latent-pivot robust formulations while retaining the shared-pivot correlation structure. We propose a hybrid ambiguity-continuation strategy that shares one ambiguity state over short arcs with verified continuity and softly links successive arcs through random-walk factors. On an approximately 20-km route from the public TEX-CUP dataset, including about 50\% deep-urban driving, the latent-pivot configuration achieves 100\% availability and a 1.6-m maximum horizontal error, with 96.16\% and 99.90\% of epochs below 1.0 and 1.5 m, respectively. The implementation is released as open-source software
Chinese Translation
城市环境中可靠的绝对位姿估计常常受到异常观测和错误时间关联的破坏,这些问题在紧耦合估计器中可能持续存在。全球导航卫星系统(GNSS)观测能够提供全局参考测量,但容易受到多路径效应的影响。视觉-惯性传感可提供局部运动约束,但错误的视觉关联会污染估计器。我们提出了SeA-RVINS,一个基于固定滞后因子图的实时动态定位(RTK)视觉-惯性系统,用于鲁棒的城市位姿估计。语义感知的学习型双目前端在持久路标进入因子图之前剔除不可靠的跟踪。对于双差分GNSS观测,SeA-RVINS通过可配置的批处理、标量和潜在枢轴(latent-pivot)鲁棒公式应用动态协方差缩放(Dynamic Covariance Scaling),同时保留共享枢轴的相关结构。我们提出一种混合模糊度延续策略,在经过连续性验证的短弧段上共享同一模糊度状态,并通过随机游走因子对相邻弧段进行软链接。在公开TEX-CUP数据集约20公里的路线上(其中约50%为深城市环境驾驶),潜在枢轴配置实现了100%的可用性和1.6米的最大水平误差,分别有96.16%和99.90%的历元误差低于1.0米和1.5米。该实现已作为开源软件发布。
cs.RO / 37 / 2609.30818

Evaluation Is All You Need for Multi-Modal Autonomous Driving

评估之于多模态自动驾驶:足矣(Evaluation Is All You Need for Multi-Modal Autonomous Driving)
He, Zeyu, Liu, Shiqi, Chen, Ke, Yan, Yun, Wu, Jinzi, Lei, Dianqiao, Wang, Sirui, Peng, ShuRui, Chen, Tao, Huang, Zhuo, Wu, Yu, Shao, Yadong, Li, Zhichao, Sun, Ke, Guan, Yang, Li, Keqiang, Li, Shengbo Eben
Abstract
Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping the candidate distribution. Nevertheless, we identify a pronounced generation-evaluation asymmetry in multi-modal planning: despite strong oracle performance, existing planners often fail to reliably select the best available candidate, leaving substantial planning potential unrealized. To address this challenge, we propose iDriveVLA, a multi-modal planning framework that improves the candidate trajectory space while enabling more reliable and context-aware trajectory evaluation. Specifically, iDriveVLA introduces a unified trajectory evaluator comprising a Safety-aware Scorer for quality and risk estimation, together with a VLM-guided Modulator for scene-adaptive criterion weighting. We further develop an oracle-aligned progressive training strategy consisting of candidate imitation pretraining, candidate space refinement, and semantic ranking alignment. On the public NAVSIM v1 leaderboard, iDriveVLA achieves a new state-of-the-art performance of 94.95 PDMS, surpassing the human-expert reference.
Chinese Translation
多模态规划通过在模糊和长尾场景中表征多种合理行为,为自动驾驶带来了广阔前景。现有方法主要致力于提升轨迹多模性、增强轨迹表征或重塑候选分布。然而,我们发现多模态规划中存在显著的产生-评估不对称性:尽管现有规划器在理想选择(oracle)性能上表现优异,但往往无法可靠地选出最优候选轨迹,导致大量规划潜力未被充分实现。为应对这一挑战,我们提出了iDriveVLA,一个多模态规划框架,它既能改进候选轨迹空间,又能实现更可靠、更具场景感知能力的轨迹评估。具体而言,iDriveVLA引入了一个统一的轨迹评估器,包括用于质量和风险估计的安全感知评分器(Safety-aware Scorer),以及用于场景自适应准则加权的视觉语言模型引导调制器(VLM-guided Modulator)。我们进一步提出了一种与理想选择对齐的渐进式训练策略,包括候选模仿预训练、候选空间精炼和语义排序对齐。在公开的NAVSIM v1排行榜上,iDriveVLA取得了94.95 PDMS的最新最优性能,超越了人类专家参考水平。
cs.RO / 38 / 2609.30828

HIRE: History-Conditioned Interaction Reasoning and High-Rate Execution for Visually Aliased Precision Manipulation

HIRE:面向视觉混叠精细操作的历史条件交互推理与高速率执行
Li, Rongji, He, Wenhao, Lu, Cewu, Chen, Xingyu, Zhang, Xu-Yao
Abstract
Precision manipulation with contact-critical interactions is often history-dependent: visually similar observations can correspond to different latent interaction states and therefore require different actions, while small execution errors can alter task outcomes. Policies relying on the current visual observation alone cannot resolve such ambiguity; force-aware and memory-augmented methods enrich physical or temporal context, while reactive high-rate policies improve local contact response, yet long-horizon temporal reasoning and precision execution remain largely decoupled in existing methods, limiting reliable progression in visually aliased precision manipulation. To bridge this gap, we introduce History-Conditioned Interaction Reasoning and Execution (HIRE), a cross-rate framework comprising a history-conditioned Interaction-State Reasoner (ISR) and a high-rate Interaction-Manifold Executor (IME). ISR encodes ordered wrench history with a temporal wrench encoder and Force Perceiver as persistent physical evidence for state-consistent action generation, while IME structures contact-critical motion into intrinsic progress and transverse correction for precise execution; their cross-rate loop allows the resulting physical traces to inform subsequent reasoning. In real-robot experiments across surface, insertion, and rotational interactions, HIRE achieves at least 90% completion across all evaluated task stages while improving interaction-state disambiguation, execution precision, and generalization. More broadly, HIRE provides a unified reasoning--execution perspective on precision manipulation under history-dependent partial observability, where physical interaction both realizes task intent and reveals latent-state evidence for future decisions. Code will be released upon publication.
Chinese Translation
涉及接触关键交互的精细操作通常具有历史依赖性:视觉上相似的观测可能对应不同的潜在交互状态,因此需要采取不同的动作,而微小的执行误差也可能改变任务结果。仅依赖当前视觉观测的策略无法消解此类歧义;力感知与记忆增强方法丰富了物理或时间上下文,反应式高速率策略改善了局部接触响应,然而在现有方法中,长时程时间推理与精细执行在很大程度上是相互解耦的,这限制了视觉混叠精细操作中任务的可靠推进。为弥合这一差距,我们提出了历史条件交互推理与执行框架(HIRE),这是一个跨速率框架,由历史条件的交互状态推理器(Interaction-State Reasoner, ISR)和高速率的交互流形执行器(Interaction-Manifold Executor, IME)组成。ISR 通过时间力矩编码器和力感知器(Force Perceiver)对有序的力矩历史进行编码,将其作为持久物理证据以生成状态一致的动作;IME 将接触关键运动结构化为内在推进分量与横向修正分量以实现精确执行;二者的跨速率循环使所产生的物理痕迹能够为后续推理提供信息。在表面、插拔和旋转交互的真实机器人实验中,HIRE 在所有评估任务阶段均达到至少 90% 的完成率,同时提升了交互状态消歧、执行精度和泛化能力。更广泛地说,HIRE 为历史依赖部分可观测条件下的精细操作提供了统一的推理—执行视角:在这种视角下,物理交互既实现任务意图,又揭示用于未来决策的潜在状态证据。代码将在论文发表后发布。
cs.RO / 39 / 2609.30833

Fast Plans, Faithful Actions: Closing the Planning-Execution Gap in Hierarchical Vision-Language-Action Models

快速规划,忠实执行:弥合分层视觉-语言-动作模型中的规划-执行鸿沟
Xie, Chuanliang, Ma, Boyu, Li, Gen, Liu, Yizhou, Chen, Houwang, Zhou, Xinyu, Yang, Jianfei
Abstract
Hierarchical vision-language-action (VLA) systems consist of a high-level vision-language planner and a low-level action expert that generates continuous actions. This hierarchical design has practical value only if the planner can generate plans fast enough to meet real-time control requirements, and the resulting plans actually contribute to the generation of action. We study one such system, a waypoint hierarchy pipeline adapted from $\pi_{0.5}$, and find that neither requirement is satisfied. This baseline relies on token-level autoregressive decoding (Token-AR) to generate a waypoint plan, requiring 57 very expensive vision-language model (VLM) forward passes. However, we find that erasing the waypoint endpoints has little effect on task success. Two findings reveal the misalignment of planner-executor: the planner generates outputs at an excessively fine granularity, and the executor underuses plans as a control condition. We address the latency issue with waypoint-aligned block-autoregressive decoding (Block-AR), and plan underuse issue with normalized goal modulation (NGM), a layer-wise goal path constrained by phase gating and anti-shortcut training so that the waypoint influences action generation maintaining other signals. Our method reduces the maximum number of VLM forward passes from 57 to 8 on LIBERO, including one prefix prefill, and achieves an $8.7\times$ reduction in planning latency on a Rokae dual-arm robot. With normalized goal modulation and anti-shortcut training, Block-AR's success rate on LIBERO-Long increases from 91.0% to 96.2%, while its average success rate across the four suites increases from 95.85% to 98.45%. On three bimanual tasks with this robot, success rates remain comparable across methods.
Chinese Translation
分层视觉-语言-动作(VLA)系统由一个高层视觉-语言规划器和一个生成连续动作的低层动作专家组成。这种分层设计只有在规划器能够足够快地生成计划以满足实时控制要求,且生成的计划确实对动作生成有所贡献时才具有实用价值。我们研究了这样一个系统——一个改编自 $\pi_{0.5}$ 的路径点分层流水线,发现这两个要求均未得到满足。该基线方法依赖词元级自回归解码(Token-AR)来生成路径点计划,需要57次非常昂贵的视觉-语言模型(VLM)前向传播。然而,我们发现擦除路径点端点对任务成功率几乎没有影响。两个发现揭示了规划器-执行器的错位:规划器以过细的粒度生成输出,而执行器未充分利用计划作为控制条件。我们采用路径点对齐的块自回归解码(Block-AR)解决延迟问题,并提出归一化目标调制(NGM)来解决计划利用不足的问题——这是一种受相位门控和反捷径训练约束的分层目标通路,使路径点在保持其他信号的同时影响动作生成。我们的方法将LIBERO上VLM前向传播的最大次数从57次减少到8次(包括一次前缀预填充),并在Rokae双臂机器人上将规划延迟降低了$8.7\times$。通过归一化目标调制和反捷径训练,Block-AR在LIBERO-Long上的成功率从91.0%提升至96.2%,在四个任务套件上的平均成功率从95.85%提升至98.45%。在该机器人上的三个双臂任务中,各方法的成功率保持相当。
cs.RO / 40 / 2609.30842

Impedance Cloning: Learning Equilibrium Point Parameters for Contact-Rich Manipulation

阻抗克隆:面向接触密集型操作的平衡点参数学习
Takahashi, Hayato, Oishi, Ryoga, Kasuga, Yuki, Tsuji, Toshiaki
Abstract
Contact-rich manipulation requires robots to regulate force against surfaces whose geometry deviates unpredictably from training conditions. Trajectory-based imitation learning, which reproduces observable outputs, breaks down under such shifts. We propose Impedance Cloning, which instead imitates the biomechanical priors that generate motion -- the stiffness and equilibrium point -- and thereby passively absorbs contact uncertainty. Because these parameters encode intent rather than outcome, they generalize across surface geometries where trajectory reproduction does not. We extract them from bilateral teleoperation demonstrations via a particle filter without force/torque sensors and evaluate the framework on two CRANE-X7 manipulators. In a wiping task with joint-space actions, the trajectory-based baseline loses contact below -6 cm, whereas the proposed method maintains a consistent 4-5 N contact force above -6 cm, with a gradual decrease below; with Cartesian-space actions, its force-height slope over 0 to +8 cm is 0.13 +/- 0.03 N/cm, versus 0.34-0.83 N/cm for fixed-impedance baselines. In a pick-and-place task with 10 diverse cups (100 trials), the proposed method succeeds in 84 trials, outperforming the fixed-impedance baseline (74/100) and performing comparably to a variable impedance control baseline (82/100) with one demonstration instead of ten. In a grasping task, the representation reduces torque tracking error with both ILBiT and Mamba backbones, confirming its generality across architectures.
Chinese Translation
接触密集型操作要求机器人对几何形状与训练条件存在不可预测偏差的表面进行力调节。基于轨迹的模仿学习通过复现可观测的输出,在此类偏移下会失效。我们提出阻抗克隆(Impedance Cloning),转而模仿生成运动的生物力学先验——刚度和平衡点——从而被动地吸收接触不确定性。由于这些参数编码的是意图而非结果,它们能够在轨迹复现无法泛化的表面几何形状之间实现泛化。我们通过粒子滤波器从双边遥操作演示中提取这些参数,无需力/力矩传感器,并在两台 CRANE-X7 机械臂上评估该框架。在一个采用关节空间动作的擦拭任务中,基于轨迹的基线方法在低于 -6 cm 处失去接触,而所提出的方法在高于 -6 cm 时保持 4-5 N 的一致接触力,低于该值时接触力逐渐减小;在采用笛卡尔空间动作时,其在 0 至 +8 cm 范围内的力-高度斜率为 0.13 ± 0.03 N/cm,而固定阻抗基线为 0.34-0.83 N/cm。在一个使用 10 种不同杯子的抓取放置任务(100 次试验)中,所提出的方法成功 84 次,优于固定阻抗基线(74/100),并以一次演示(而非十次)取得了与变阻抗控制基线(82/100)相当的性能。在一个抓取任务中,该表示在 ILBiT 和 Mamba 两种骨干网络下均降低了力矩跟踪误差,证实了其在不同架构间的通用性。
cs.RO / 41 / 2609.30868

VLaRL: Augmenting Vision-Language-Action Models with Simulation-Trained Latent-Conditioned Residual RL

VLaRL:利用基于仿真的潜变量条件残差强化学习增强视觉-语言-动作模型
Saito, Namiko, Kim, Kinam, Kim, Heecheol, Ikeuchi, Katsushi, Matsushita, Yasuyuki
Abstract
Vision-language-action (VLA) models provide broad, instruction-conditioned manipulation behaviors, but their physical execution can remain imprecise during contact-rich interaction. Residual reinforcement learning (RL) can correct such errors while keeping the VLA frozen, but real-robot RL is costly and safety-critical. We propose VLA Latent-Conditioned RL (VLaRL), which enables residual RL for frozen VLAs to be trained in simulation and deployed on real robots without real-world RL or online adaptation. The key challenge is transferring the learned residual policy despite the visual gap between simulation and reality. Rather than requiring pixel-level visual correspondence, VLaRL uses the VLA's internal vision-language latent representation to condition residual control and as the sim-to-real transfer interface, and learns a lightweight mapper that transforms simulation-derived latents toward the real latent distribution. Across four contact-rich manipulation tasks and two VLA backbones, VLaRL improves real-world success in all task-backbone combinations, while controlled ablations demonstrate the importance of both latent conditioning and latent alignment for transferring simulation-trained residual control.
Chinese Translation
视觉-语言-动作(VLA)模型能够提供广泛的、受指令条件约束的操作行为,但在接触丰富的交互过程中,其物理执行可能不够精确。残差强化学习(Residual RL)可以在保持VLA模型冻结的前提下纠正此类误差,但真实机器人上的强化学习成本高昂且涉及安全风险。我们提出VLA潜变量条件强化学习(VLaRL),使冻结的VLA的残差强化学习能够在仿真中训练,并直接部署到真实机器人上,无需真实世界强化学习或在线适应。关键挑战在于克服仿真与现实之间的视觉差异,实现所学习残差策略的迁移。VLaRL并不要求像素级的视觉对应,而是利用VLA内部的视觉-语言潜变量表征来条件化残差控制,并将其作为仿真到现实(sim-to-real)迁移的接口,同时学习一个轻量级映射器,将仿真得到的潜变量向真实潜变量分布进行变换。在四项接触丰富的操作任务和两种VLA骨干模型上的实验表明,VLaRL在所有任务-骨干组合中均提升了真实世界的成功率;受控消融实验进一步证明了潜变量条件化和潜变量对齐对于迁移仿真训练的残差控制的重要性。
cs.RO / 42 / 2609.30873

A Second Torque Port for Series Elastic Actuators: Parallel-Integrated Design and Time-Scale Torque Allocation

串联弹性执行器的第二个力矩端口:并联一体化设计与时间尺度力矩分配
Jia, Donghao, Xia, Xiubo, Liu, Junhang, Zhu, Zeyu, Geng, Xiaoyu, Sun, Jian
Abstract
A series elastic actuator has a single torque port and pays for it twice: the geared motor must swing its own reflected inertia through the spring, so the amplitude it delivers collapses as $\omega^{-2}$ in the command frequency $\omega$ once it saturates, while commands below the transmission's breakaway friction never arrive at all. This letter opens a second torque port on the load side, placing a frameless direct-drive micro motor in parallel with a fixed-stiffness spring -- a parallel-integrated SEA, or Pi-SEA, whose delivered torque is read from spring deflection and micro current without a sensor -- and dividing the commanded torque between the two channels by time scale rather than by filter design. The micro torque loop is the fast subsystem, which makes the closed loop singularly perturbed and turns the separation the channels need into a bound to check rather than a crossover to tune; a leaky mid-ranging integrator returns the steady load to the spring; and the amplitude ceiling, read backwards, becomes a closed-form sizing rule that matches spring, geared motor and micro motor to the amplitudes and frequencies an application asks for. Against SEAs, the Pi-SEA widens the tracked band at small amplitudes and lowers the residual the joint imposes on its environment, each by an order of magnitude.
Chinese Translation
串联弹性执行器(SEA)只有一个力矩端口,并为此付出双重代价:带减速器的电机必须通过弹簧驱动其自身折算惯性,因此一旦饱和,其输出幅值随指令频率 $\omega$ 以 $\omega^{-2}$ 的速度衰减;而低于传动机构静摩擦的指令则完全无法传递。本快报在负载侧开辟了第二个力矩端口,将一台无框架直驱微型电机与定刚度弹簧并联——称为并联一体化SEA(Pi-SEA),其输出力矩可由弹簧变形量与微型电机电流读出而无需传感器——并按时间尺度而非滤波器设计在两条通道之间分配指令力矩。微型电机力矩环是快子系统,这使闭环成为奇异摄动系统,从而把两通道所需的分离转化为一个只需检验的界,而无需整定的交叉频率;一个带泄漏的中程积分器将稳态负载归还给弹簧;而幅值上限反推回去,则成为一条闭式选型规则,可将弹簧、减速电机与微型电机与应用所需的幅值和频率相匹配。相较于SEA,Pi-SEA在小幅值下将可跟踪频带拓宽了一个数量级,并将关节施加于环境的残余力降低了一个数量级。
cs.RO / 43 / 2609.30889

PHASE: Compliance-Enabled Tactile Phase Retrieval for Few-Shot Insertion Learning

PHASE:面向少样本插入学习的柔顺触觉相位检索方法
Siburian, Jeremy, Beltran-Hernandez, Cristian C., Matsushima, Tatsuya, Iwasawa, Yusuke, Hamaya, Masashi, Nishimura, Mai
Abstract
Contact-rich assembly tasks such as peg-in-hole insertion remain difficult to learn from limited demonstrations. While retrieval-augmented imitation learning, which augments target demonstrations with relevant prior data, offers a promising direction, its applicability to contact-rich manipulation remains largely unexplored. Contact-rich insertion unfolds over multiple phases from search to insert, and retrieving phase-specific experience from prior data in principled ways remains an open question. Our key insight is that a compliant wrist enables the robot to sustain contact throughout execution, producing rich tactile and force signals that naturally reveal the phase structure of insertion and inform what should be retrieved. Based on this insight, we present PHASE (PHase-Aware Segmentation and REtrieval), a framework for compliance-enabled tactile phase retrieval that integrates multimodal contact-aware representation learning, variable-length phase segmentation from tactile signals, and phase-consistent retrieval for policy learning. We evaluate PHASE on real-world peg-in-hole insertion across five peg geometries, comparing against retrieval strategies drawn from state-of-the-art methods under a shared policy architecture. PHASE improves the overall success rate by 13 percentage points over the strongest non-phase-aware baseline, and improves performance under unseen initial positions by 30 percentage points. These results demonstrate that aligning retrieval with interaction-defined contact phases substantially improves robustness in few-shot insertion learning.
Chinese Translation
诸如轴孔装配(peg-in-hole insertion)等接触密集型装配任务在仅有少量示范数据的情况下仍然难以学习。检索增强模仿学习通过将目标示范与相关先验数据相结合,是一个有前景的方向,但其在接触密集型操作中的应用尚未得到充分探索。接触密集型插入任务通常经历从搜索到插入的多个阶段,如何以有原则的方式从先验数据中检索特定阶段的相关经验仍是一个开放性问题。我们的核心洞察是:柔顺型腕部(compliant wrist)使机器人在整个执行过程中能够持续保持接触,从而产生丰富的触觉和力信号,这些信号自然地揭示了插入任务的阶段结构,并指明应当检索哪些内容。基于这一洞察,我们提出了PHASE(PHase-Aware Segmentation and REtrieval,相位感知分割与检索),一个柔顺使能的触觉相位检索框架,它集成了多模态接触感知表示学习、基于触觉信号的可变长度阶段分割,以及用于策略学习的相位一致性检索。我们在真实世界的轴孔装配任务上评估了PHASE,涵盖五种轴的几何形状,并在共享策略架构下与来自最先进方法的检索策略进行比较。PHASE相较于最强的非相位感知基线将整体成功率提升了13个百分点,在未见过的初始位置条件下将性能提升了30个百分点。这些结果表明,使检索与由交互定义的接触阶段保持一致,能够显著提升少样本插入学习的鲁棒性。
cs.RO / 44 / 2609.30913

Causeway: Restoring Task Accessibility for Instruction Switching in VLA Policies

Causeway:恢复VLA策略中指令切换的任务可达性
Wang, Qingzi, Feng, Kaixi, Shi, Guangyao, Wu, Xiyang, Li, Ang, Manocha, Dinesh
Abstract
Vision-language-action (VLA) policies can execute many tasks from standard initial states, yet a new instruction may fail after another task has altered the robot's physical state. We study instruction switching, where a new task is issued during or after the execution of a different one. We observe that a target task that is reliably completed from its standard initial states can become inaccessible from states produced by a preceding task. We call such states task islands. We propose Causeway, a training-free inference-time intervention. Given the current state and a re-entry pose for the target task, Causeway back-propagates through the frozen decoding computation and applies a state-directed write within the action-stream representation. The VLA decodes the return motion itself, without parameter updates, a new action head, or external action generation. Across 71 cross-object pairs, three switch timings, and three VLA architectures on LIBERO-Goal, Causeway raises bare-switch success from 3-26% to 47-65% and increases the rate of reaching the handoff neighborhood by 42-72 percentage points across models. Additional experiments on LIBERO-Object and a real xArm platform show that the recovery extends beyond the main LIBERO-Goal setting, both in simulation and on a robot.
Chinese Translation
视觉-语言-动作(VLA)策略能够从标准初始状态执行多种任务,但当某一任务改变了机器人的物理状态后,新的指令可能会执行失败。我们研究指令切换问题,即在一个任务的执行过程中或执行结束后发出新任务。我们观察到,一个从其标准初始状态出发能够可靠完成的目标任务,在由前一任务产生的状态下可能变得不可达。我们将这类状态称为“任务孤岛”。我们提出Causeway,一种无需训练的推理时干预方法。给定当前状态和目标任务的重新进入位姿,Causeway通过冻结的解码计算进行反向传播,并在动作流表示内施加状态导向的写入。VLA自行解码返回动作,无需参数更新、新的动作头或外部动作生成。在LIBERO-Goal上的71个跨物体对、三种切换时机和三种VLA架构的实验中,Causeway将直接切换的成功率从3-26%提升至47-65%,并将各模型到达交接邻域的比率提升42-72个百分点。在LIBERO-Object和真实xArm平台上的额外实验表明,该恢复能力超越了主实验的LIBERO-Goal设置,在仿真和真实机器人上均有效。
cs.RO / 45 / 2609.30951

Bundled Contact Gradients: Stabilizing Differentiable Simulation for Deployable Dynamic Tasks

捆绑接触梯度:稳定可微仿真以实现可部署的动态任务
Aditya, Dyuman, Cheng, Jin, Schwarke, Clemens, Nguyen, Quan, Sukhatme, Gaurav, Coros, Stelian, Fadini, Gabriele
Abstract
Differentiable simulation provides analytic gradients of robot dynamics, enabling fast and sample-efficient first-order policy optimization. However, obtaining smooth and informative gradients through rigid-body contact typically requires softened contact models, often at the expense of physical fidelity and thereby limiting learned policies largely to simulation. This trade-off becomes particularly consequential for dynamic humanoid motions, where accurate contact dynamics are critical for transferring policies to the real world. Increasing contact stiffness in rigid-body simulation improves the fidelity of interactions, but also makes the dynamics increasingly sensitive to small state perturbations, producing high-variance gradients that can destabilize first-order policy learning. To address this, we propose \emph{Bundled Contact Gradients (BCG)}, a contact-local randomized smoothing framework for differentiable policy learning. When stiff contact is detected, our method evaluates a local bundle of randomized perturbation rollouts around the stiff contact configuration and aggregates their gradient signal thereby reducing gradient variance. We demonstrate the effectiveness of our method by successfully training and transferring dynamic motions zero-shot onto a real-world Unitree G1 humanoid platform. Videos and supplementary information can be found at https://bundledcontactgradients.github.io/
Chinese Translation
可微仿真为机器人动力学提供解析梯度,从而实现快速且样本高效的一阶策略优化。然而,要在刚体接触中获得平滑且信息丰富的梯度,通常需要软化接触模型,这往往以牺牲物理保真度为代价,进而使学习到的策略大多只能在仿真中有效。对于动态人形机器人运动而言,这种权衡尤为关键,因为精确的接触动力学对于将策略迁移到现实世界至关重要。在刚体仿真中提高接触刚度可以改善交互的保真度,但同时也会使动力学对微小的状态扰动越来越敏感,产生高方差的梯度,从而破坏一阶策略学习的稳定性。为解决这一问题,我们提出了捆绑接触梯度,这是一种面向可微策略学习的接触局部随机平滑框架。当检测到刚性接触时,我们的方法在刚性接触构型周围评估一组局部随机扰动回滚,并聚合它们的梯度信号,从而降低梯度方差。我们通过在真实世界的 Unitree G1 人形机器人平台上零样本训练并迁移动态运动,验证了该方法的有效性。视频和补充信息请参见 https://bundledcontactgradients.github.io/
cs.RO / 46 / 2609.30959

VisTacAlign: Co-Training Dexterous Policies on Tactile Human and Robot Demonstrations

VisTacAlign:在触觉人类与机器人演示数据上协同训练灵巧操作策略
Poffet, Julien, Strong, Matthew, Dhawan, Ankush, Shi, Baiyu, Neelaveni, Shalika, Yuan, Yujia, Bao, Zhenan, Kennedy III, Monroe
Abstract
Human demonstrations are a cheap source of data for dexterous manipulation, but co-training a robot policy on them requires closing the human--robot gap in every modality the policy consumes. We present VisTacAlign, a framework for co-training 3D-visual-tactile dexterous policies on human and robot demonstrations. Glove-tracked human hand motion is retargeted to a 17-DoF tactile robot hand with a one-time fingertip correction. The human hand is then erased from both stereo views and replaced by a posed robot-hand mesh painted with pixels from robot recordings, and a real-time stereo foundation model is re-run on the composite, so the human point clouds carry the same stereo errors and visibility as the robot ones. Finally, a capacitive tactile glove is aligned to the robot's fingertip sensors in its signal space, giving one interpretable per-finger force representation. A diffusion transformer consumes point-cloud, proprioceptive, and per-finger tactile tokens. On three real-world tasks requiring precise force -- Lego assembly, plucking strawberries of varying size, and activating and lifting a power drill -- adding aligned human demonstrations to existing robot data improves over robot-only policies, and ablations show that both tactile input and visual alignment are necessary. Project page: https://vis-tac-align.github.io
Chinese Translation
人类演示数据是灵巧操作任务的一种廉价数据来源,但要在这些数据上协同训练机器人策略,需要弥合策略所接收的每种模态中人与机器人之间的差异。我们提出了 VisTacAlign,一个在人类和机器人演示数据上协同训练 3D 视觉-触觉灵巧策略的框架。通过手套追踪的人手运动经过一次性的指尖校正,重定向到一台 17 自由度的触觉机器人手上。随后,将人手从双目立体视图中擦除,并替换为一个姿态匹配的机器人手网格,该网格涂覆了来自机器人录制的像素,再在合成图像上重新运行实时双目基础模型,从而使人类的点云携带与机器人点云相同的双目误差和可见性特征。最后,通过在信号空间中将电容式触觉手套与机器人的指尖传感器对齐,获得一种可解释的逐手指力表示。扩散Transformer(Diffusion Transformer)接收点云、本体感觉和逐手指触觉令牌。在三项需要精确力的真实世界任务——乐高装配、采摘不同大小的草莓以及启动并提起电钻——上,将已对齐的人类演示数据添加到现有机器人数据中,可超越仅使用机器人数据的策略,且消融实验表明触觉输入和视觉对齐都是必要的。项目主页:https://vis-tac-align.github.io
cs.RO / 47 / 2609.30965

FRAM: Trajectory-Guided Visual Feature Selection for Compact Language-Conditioned Robot Manipulation

FRAM:面向紧凑型语言条件机器人操作的轨迹引导视觉特征选择方法
Ito, Hiroshi, Hiruma, Hyogo, Kanai, Yoshiki, Yoshida, Takahiro, Kanazawa, Akira, Yamada, Hiroki
Abstract
Vision-Language-Action models achieve strong performance in robot manipulation, but often require large numbers of parameters. In this work, we propose the Future Representation Action Model (FRAM), a small policy that explicitly links the future end-effector trajectory to the current visual input. FRAM uses the image coordinates of the predicted trajectory as spatial pointers and reads local visual features related to the motion from the current image. This organizes the information for action generation into the reference position (Where), the visual state (What), and the future motion (Future). Trajectory labels are generated automatically from demonstrations and camera geometry, so no manual annotation is needed. With 138.7M parameters, including a frozen language encoder, FRAM reaches an average success rate of 92.2% over the four standard LIBERO suites, close to the 94.2% of $\pi_0$ with 3.3B parameters. Without extra training, it also reaches an average of 67.3% on LIBERO-Plus. Ablations confirm that both the future trajectory and the local visual features improve performance and robustness. On a real dual-arm UR5e, FRAM stacks cups using only wrist cameras, including choosing and switching between the left and right arms. These results show that selecting visual information based on future motion is an effective way to obtain both high performance and robustness in a small robot policy.
Chinese Translation
视觉-语言-动作(Vision-Language-Action)模型在机器人操作任务中取得了优异的性能,但通常需要大量的参数。在本工作中,我们提出了未来表征动作模型(Future Representation Action Model,FRAM),这是一个小型策略模型,它显式地将未来末端执行器轨迹与当前视觉输入关联起来。FRAM 将预测轨迹的图像坐标用作空间指针,并从当前图像中读取与运动相关的局部视觉特征。这将用于动作生成的信息组织为参考位置(Where)、视觉状态(What)和未来运动(Future)。轨迹标签由示教数据和相机几何关系自动生成,无需人工标注。FRAM 仅含 1.387 亿个参数(包括冻结的语言编码器),在四个标准 LIBERO 套件上取得了 92.2% 的平均成功率,接近参数量达 33 亿的 π₀ 的 94.2%。在不进行额外训练的情况下,它在 LIBERO-Plus 上也达到了 67.3% 的平均成功率。消融实验证实,未来轨迹和局部视觉特征均能提升性能与鲁棒性。在真实的双臂 UR5e 机器人上,FRAM 仅使用腕部相机即可完成杯子堆叠任务,包括在左右臂之间进行选择与切换。这些结果表明,基于未来运动选择视觉信息是在小型机器人策略中同时获得高性能与鲁棒性的有效途径。
cs.RO / 48 / 2609.30969

TACTIC: Understanding Tactile Encoders and Conditioning for Contact-rich Robot Manipulation Policies

TACTIC:面向接触丰富的机器人操作策略的触觉编码器与条件机制研究
Bien, Seongjin, Makowski, Débora Oliveira, Kneissl, Carlo, Mirjalili, Reihaneh, Vanjani, Pankhuri, Lioutikov, Rudolf, Kutyniok, Gitta, Walter, Florian, Burgard, Wolfram
Abstract
Tactile information is essential for contact-rich manipulation tasks in robotics. Vision-based tactile sensors make it particularly easy to design end-to-end manipulation policies with tactile sensing, as they enable the use of existing encoders from computer vision. However, this has led to a huge variety of architectures, training datasets, and evaluation protocols, making it difficult to determine which design choices best encode touch. In this work, we address this gap and present a comprehensive study of tactile encoders and fusion strategies across various contact-rich manipulation tasks in real-world experiments. To enable a controlled comparison, we train and evaluate all models under the same pipeline and experimental setup, comprising more than 2000 real-world rollouts. Our results go beyond other studies that only compare simulation performance, which does not necessarily translate to real-world settings, where large-scale evaluations are needed to obtain reliable statistics. Our key finding is that there is no universally optimal representation or fusion strategy for encoding visual-tactile. Instead, the best encoder backbone and fusion scheme depend strongly on the task.
Chinese Translation
触觉信息对于机器人学中接触丰富的操作任务至关重要。基于视觉的触觉传感器使得设计带有触觉感知的端到端操作策略变得尤为便捷,因为它们可以直接使用计算机视觉领域现有的编码器。然而,这也导致了大量不同的架构、训练数据集和评估协议的出现,使得人们难以确定哪些设计选择能够最好地对触觉进行编码。在本工作中,我们针对这一空白,对触觉编码器和融合策略进行了全面研究,并在真实世界实验中涵盖多种接触丰富的操作任务。为了实现可控的比较,我们在相同的流程和实验设置下对所有模型进行训练和评估,包含超过2000次真实世界测试。我们的结果超越了其他仅比较仿真性能的研究——仿真性能未必能迁移到真实世界场景,而真实场景中需要大规模评估才能获得可靠的统计结果。我们的关键发现是:不存在普遍最优的视觉-触觉编码表示或融合策略。相反,最优的编码器骨干网络和融合方案在很大程度上取决于具体任务。
cs.RO / 49 / 2609.31008

Quadruped Obstacle Avoidance and Footstep Planning with Distributed Low-cost Time-of-Flight Sensors

基于分布式低成本飞行时间(ToF)传感器的四足机器人避障与落足点规划
Caroleo, Giammarco, Mahamoodally, Timothée, Manzardo, Matteo, Jin, Jin, Pontin, Marco, Mattamala, Matias, Vidoni, Renato, Maiolino, Perla, Fallon, Maurice
Abstract
Quadruped robots typically rely on depth cameras and LiDAR sensors to map their local environment. However, these sensors have limited close-range coverage, are relatively expensive, and consume significant power. This study investigates whether distributed Time-of-Flight (ToF) sensors can serve as a low-cost alternative to depth cameras for near-field terrain mapping for locomotion and local navigation. We designed a distributed ToF sensing architecture for the ANYbotics ANYmal quadruped, assessed its environment reconstruction accuracy, and benchmarked it against depth cameras for terrain mapping and obstacle avoidance. Distributing these sensors around the robot can also avoid the blind spots of traditional sensors. Our results show that, despite their low resolution and higher measurement noise, distributed ToF sensors can support reliable perceptual locomotion with centimeter-level local mapping accuracy. The proposed sensing strategy provides sufficient geometric information for near-field obstacle avoidance and footstep planning, at substantially lower cost, energy consumption, and system complexity than depth cameras.
Chinese Translation
四足机器人通常依赖深度相机和激光雷达(LiDAR)传感器来感知其局部环境。然而,这些传感器近距离覆盖范围有限、价格相对昂贵且功耗较高。本研究探讨分布式飞行时间(Time-of-Flight, ToF)传感器能否作为一种低成本的替代方案,用于运动控制与局部导航所需的近场地形感知。我们为 ANYbotics ANYmal 四足机器人设计了一种分布式 ToF 传感架构,评估了其环境重建精度,并将其与深度相机在地形建图和避障方面的性能进行了对比测试。将此类传感器分布在机器人周身还可以避免传统传感器的盲区。结果表明,尽管 ToF 传感器分辨率较低且测量噪声较大,但其分布式布置仍能支持可靠的感知式运动控制,实现厘米级的局部建图精度。所提出的感知策略能够为近场避障和落足点规划提供充分的几何信息,且其成本、能耗和系统复杂度均远低于深度相机。
cs.RO / 50 / 2609.31014

Co-design of trajectory and morphology for a vertical jump-climbing robot

一种垂直跳跃攀爬机器人的轨迹与形态协同设计
Xu, Christopher Y., Hawkes, Elliot W.
Abstract
Animals such as squirrels and even bears have adapted to rapidly climb up trees and other complex vertical terrain, but achieving comparable agility has been a challenge for climbing robots. Existing robots often use walking gaits and move conservatively to stay in contact with the surface, which limits the range of dynamic maneuvers. In this paper we present a 290 g robot that, to our knowledge, is the first to climb vertically by bounding (with an aerial phase). We leverage a co-design workflow, in which the morphology and trajectory are jointly optimized for fast locomotion, subject to adhesion force limitations seen in spined grippers. The resulting trajectory includes a rapid maneuver that launches the robot vertically, and an aerial reorientation that brings the front grippers back to the surface using the rear leg as an inertial tail. We evaluate the resulting jump forces in 2D force space and demonstrate that the optimized morphology is capable of continuous climbing at a speed of 0.375 m/s (1.97 body lengths/s), and can also achieve ground locomotion and transition to a vertical surface. Our work proposes design insights for the jump-climbing maneuver and serves as an important step toward creating climbing robots with agility on par with that of animals.
Chinese Translation
松鼠甚至熊等动物已适应快速攀爬树木及其他复杂的垂直地形,但攀爬机器人要实现同等的敏捷性一直是一个挑战。现有机器人通常采用步行步态并保守地移动以保持与表面的接触,这限制了动态机动的范围。在本文中,我们提出了一种290克的机器人,据我们所知,这是首个通过跳跃式奔跑(bounding,含腾空阶段)实现垂直攀爬的机器人。我们采用协同设计流程,在考虑带刺夹持器所受粘附力限制的条件下,对形态和轨迹进行联合优化以实现快速运动。优化得到的轨迹包括一个使机器人垂直起跳的快速机动,以及一个空中重定向动作——将后腿用作惯性尾,使前夹持器重新回到表面。我们在二维力空间中评估了所产生的跳跃力,并证明优化后的形态能够以0.375米/秒(1.97个体长/秒)的速度进行连续攀爬,同时还能实现地面运动并向垂直表面的过渡。我们的工作为跳跃攀爬机动提供了设计见解,是向创造具有与动物相当敏捷性的攀爬机器人迈出的重要一步。
cs.RO / 51 / 2609.31025

Precision at Speed: Sample-Efficient Online Model-Based Reinforcement Learning for Hydraulic Excavator Control

高速下的精准控制:面向液压挖掘机控制的样本高效在线基于模型强化学习
Canales, Claudio, Nan, Fang, Hutter, Marco, Ruiz-del-Solar, Javier
Abstract
Precise, high-speed control remains challenging for robots with complex actuation dynamics. Learning directly on hardware is further constrained by the cost of real-world interaction. We present an online model-based reinforcement learning framework that learns a probabilistic dynamics ensemble model from scratch for sampling-based model predictive control. A precision-gated contouring objective conditions the progress reward on path accuracy, prioritizing precision over speed. In a data-driven excavator simulator, the framework achieves higher sample efficiency than the evaluated model-based reinforcement learning baselines. We validate the framework by learning directly on an 11.5-ton Menzi Muck M445 hydraulic excavator, without demonstrations or simulation pretraining. After 20 minutes of interaction, the controller reaches tracking accuracy comparable to prior learned controllers trained on 100-150 minutes of data. After 40 minutes, it sustains sub-centimeter mean path error at high operating speeds.
Chinese Translation
对于具有复杂执行器动力学的机器人而言,精确的高速控制仍然极具挑战性。直接在硬件上学习还进一步受到真实世界交互成本的制约。我们提出了一种在线基于模型的强化学习框架,该框架从零开始学习一个概率性动力学集成模型,用于基于采样的模型预测控制。通过精度门控的轮廓加工目标函数,将进度奖励与路径精度相关联,从而优先保证精度而非速度。在数据驱动的挖掘机仿真器中,该框架相比所评估的基于模型的强化学习基线方法实现了更高的样本效率。我们通过直接在一台11.5吨的Menzi Muck M445液压挖掘机上进行学习来验证该框架,无需演示数据或仿真预训练。经过20分钟的交互,控制器的轨迹跟踪精度可媲美先前使用100-150分钟数据训练的学习型控制器。经过40分钟交互后,该控制器在高运行速度下能够保持亚厘米级的平均路径误差。
cs.RO / 52 / 2609.31035

Compact Force Sensor for Dual-UAV Cable-Suspended Payload Transport with Tension-Aware Outer-Loop Control

面向双无人机吊挂载荷运输的紧凑型力传感器及张力感知外环控制
Delbene, Andrea, Cannata, Giorgio, Carlini, Giorgio, Sante, Filippo, Baglietto, Marco
Abstract
Cooperative payload transportation using multiple \textit{Unmanned Aerial Vehicles} (UAVs) poses challenges in stability, coordination, and robustness, especially under external disturbances and unmodeled dynamics. This work proposes a dual-UAV payload transportation framework supported by a compact, custom-designed force sensor measuring the interaction force at the UAV cable anchor point. The sensor design and mathematical model are presented, and its performance is characterized through static and dynamic tests evaluating linearity, hysteresis, repeatability, and crossload. The control architecture follows a cascade structure: fast inner loops handle vehicle stabilization, while outer loops are designed to compensate for the measured forces. The approach is validated through simulations and indoor experiments under position uncertainty. Payload-drop and constrained-space tests assess the proposed sensing and control architecture against literature-based distributed references, showing improved stabilization, coordination, and disturbance rejection. A video of the experiments is available at: https://youtu.be/rIw9-fvV8Qw.
Chinese Translation
利用多架无人机(Unmanned Aerial Vehicles, UAVs)协同运输载荷在稳定性、协调性和鲁棒性方面面临挑战,尤其是在外部扰动和未建模动力学存在的情况下。本工作提出了一种双无人机载荷运输框架,并配套设计了一种紧凑型定制力传感器,用于测量无人机吊索锚点处的相互作用力。文中给出了该传感器的设计方案与数学模型,并通过静态和动态测试(评估线性度、迟滞、重复性和交叉负载)对其性能进行了表征。控制架构采用级联结构:快速内环负责载具稳定,而外环则设计用于补偿所测得的力。该方法通过仿真以及在位置不确定性条件下的室内实验得到验证。针对载荷脱落和受限空间的测试,将所提出的传感与控制架构与基于文献的分布式参考方案进行了对比,结果显示其在稳定性、协调性和扰动抑制方面均有提升。实验视频可访问:https://youtu.be/rIw9-fvV8Qw。
cs.RO / 53 / 2609.31048

Kintsugi-VLA: Turning Failed Robot Rollouts into Recovery Data through Interventional Recoverability

Kintsugi-VLA:通过介入式可恢复性将失败的机器人 rollout 转化为恢复数据
Snegirev, Ivan, Semenyakina, Elizaveta, Maliukov, Dmitrii, Cabrera, Miguel Altamirano, Tsetserukou, Dzmitry
Abstract
Simulation enables scalable training of Vision-Language-Action policies by using privileged experts to generate visual demonstrations without requiring every trajectory to be collected through manual teleoperation. However, such pipelines typically retain successful demonstrations while failed rollouts are discarded, even though they expose precisely the off-nominal states from which recovery must be learned. We introduce Kintsugi-VLA, a framework for converting failed rollouts into targeted synthetic recovery data by exploiting exact state restoration and branching in simulation. For a fixed privileged expert, we define interventional recoverability as the probability of completing the original task after the simulator is restored to a given state, estimate it using adaptive Monte Carlo continuations with pointwise Wilson intervals, and characterize its non-monotonic evolution along failed trajectories. These estimates identify an observed terminal low-recoverability frontier-the point after which measured recoverability remains below a threshold-which is then used to select informative recovery starting states. In a simulated Franka manipulation task, targeted recovery data yield aggregate SmolVLA recovery success of 34.6\% and 38.4\% under difficulty- and frame-budget matching, respectively, 5.8 and 6.7 percentage points above uniform sampling within the same recovery window. The same ordering is observed under disturbed end-to-end execution and shifted clutter and physics conditions, while clean-task success decreases from 76.8\% to 74.7\%. Kintsugi-VLA demonstrates how failed simulator rollouts can be transformed from discarded experience into structured recovery-training data through direct interventional measurement.
Chinese Translation
仿真通过利用特权专家生成视觉演示,实现了视觉-语言-动作(VLA)策略的可扩展训练,而无需每条轨迹都通过手动遥操作来采集。然而,此类流程通常只保留成功的演示,而将失败的 rollout 丢弃,尽管这些失败恰恰暴露了必须从中学习恢复能力的非正常状态。我们提出了 Kintsugi-VLA,一个通过利用仿真中精确的状态恢复与分支,将失败的 rollout 转化为针对性合成恢复数据的框架。对于给定的特权专家,我们将介入式可恢复性定义为在仿真器恢复到某一给定状态后完成原始任务的概率,采用自适应蒙特卡洛延续并结合逐点 Wilson 区间对其进行估计,并刻画了其沿失败轨迹的非单调演化。这些估计识别出观测到的终端低可恢复性边界——即此后测得的可恢复性始终低于某一阈值的时点——随后用于选择信息量丰富的恢复起始状态。在一个仿真 Franka 操作任务中,在难度匹配和帧预算匹配条件下,针对性恢复数据使 SmolVLA 的总体恢复成功率达到 34.6% 和 38.4%,分别比同一恢复窗口内的均匀采样高出 5.8 和 6.7 个百分点。在受扰动的端到端执行以及变化的杂乱度和物理条件下,观察到相同的排序,而干净任务的成功率从 76.8% 下降至 74.7%。Kintsugi-VLA 展示了如何通过直接的介入式测量,将失败的仿真 rollout 从被丢弃的经验转化为结构化的恢复训练数据。
cs.RO / 54 / 2609.31057

Accuracy Evaluation of INS/ZUPT Filtering Methods Based on Different Geometric Error Definitions

基于不同几何误差定义的INS/ZUPT滤波方法精度评估
Ouyang, Wei, Han, Jiale, Luo, Yarong, Zhu, Maoran
Abstract
Geometric filters have recently been introduced to improve the accuracy and consistency of inertial-based integrated navigation systems. Error states were defined through specific group operations, introducing state correlations in error definition, which were lacked in the additive error used by a conventional indirect Kalman filter. The desirable consistent filtering models can be obtained based on specific geometric errors. For zero-velocity measurements expressed in the reference frame, this paper derives left-error process and measurement models from invariant filtering, two-frame-group filtering, and equivariant filtering. Importantly, a new group operation is introduced for the left tangent-group equivariant error. The analysis shows that the two-frame-group invariant extended Kalman filter (TFG-IEKF) and the tangent-group equivariant filter (TG-EqF) do not offer a significant consistency advantage over the invariant extended Kalman filter (IEKF). Experiments with an INS/ZUPT measurement system show that, under small initial attitude errors, the conventional indirect extended Kalman filter (EKF) achieves loop-closure position errors below $0.1\%$ of the traveled distance, while the three geometric filters achieve comparable positioning accuracy.
Chinese Translation
几何滤波器近年来被引入以提升基于惯性导航的组合导航系统的精度与一致性。误差状态通过特定的群运算来定义,在误差定义中引入了状态相关性,而传统间接卡尔曼滤波器所采用的加性误差则缺乏这种相关性。基于特定的几何误差可以获得理想的一致性滤波模型。对于在参考坐标系中表达的零速量测,本文从不变滤波(invariant filtering)、双群滤波(two-frame-group filtering)和等变滤波(equivariant filtering)出发,推导了左误差的过程模型和量测模型。重要的是,本文为左切群等变误差引入了一种新的群运算。分析表明,双群不变扩展卡尔曼滤波器(TFG-IEKF)与切群等变滤波器(TG-EqF)相较于不变扩展卡尔曼滤波器(IEKF)并未提供显著的一致性优势。基于INS/ZUPT测量系统的实验表明,在较小初始姿态误差条件下,传统间接扩展卡尔曼滤波器(EKF)的闭环位置误差低于行进距离的$0.1\%$,而三种几何滤波器取得了相当的定位精度。
cs.RO / 55 / 2609.31110

AuthGuard-R: Safety-Compliant Mission Hijacking and Dual-Gate Defense for LLM-Controlled Robots

AuthGuard-R:面向大语言模型控制机器人的安全合规型任务劫持与双门防御机制
Chepuri, Saidattu, Srivastava, Vikas
Abstract
Large language models are increasingly used as high-level planners for mobile robots, robot manipulators, and autonomous vehicles. Recent studies show that these systems can be influenced through malicious text, speech, visual instructions, retrieved documents, and poisoned sensory context. Most defenses ask whether a proposed action is physically safe. This paper studies a different problem: an action may be physically safe and still violate the mission authorized by the user. An attacker may redirect a delivery robot, replace an approved object, extend a robot's operating region, activate an unnecessary sensor, or delay a mission without creating an immediate physical hazard. We call this attack \emph{safety-compliant mission hijacking}. We propose MissionPAIR, an adaptive attack framework that searches for executable plans that pass a safety gate while violating an authenticated mission. We also propose AuthGuard-R, a deterministic authorization layer that binds every executable action to a signed mission, robot identity, object and region scope, current state, time, and input provenance. AuthGuard-R operates with an independent safety gate, giving a dual-gate architecture. We formalize mission policies over robot traces, define security games, and prove authorization soundness, mission non-escalation, replay resistance, robot binding, provenance separation, threshold-approval security, audit-log tamper evidence, and trace-level composition. We report a preliminary cross-model evaluation with Claude Haiku~4.5 and the open-source Qwen2.5~7B planner. Across 240 live attack trials, the planners followed an injected mission deviation in 109 trials; AuthGuard-R rejected all 109 resulting unauthorized actions. A separate hand-constructed suite of eleven protocol- and policy-level attacks was also blocked completely.
Chinese Translation
大语言模型正日益被用作移动机器人、机械臂和自动驾驶车辆的高层规划器。已有研究表明,这些系统可能受到恶意文本、语音、视觉指令、检索到的文档以及被投毒的感知上下文的影响。大多数防御方法仅判断拟执行的动作在物理上是否安全。本文研究的是另一个不同的问题:一个动作可能在物理上是安全的,却违反了用户授权的任务。攻击者可以转移配送机器人的路线、替换已批准的物体、扩展机器人的运行区域、激活不必要的传感器,或在不造成直接物理危险的情况下延迟任务执行。我们将此类攻击称为“安全合规型任务劫持”(safety-compliant mission hijacking)。我们提出了MissionPAIR,一种自适应攻击框架,用于搜索能够通过安全门检查但违反已认证任务的可行计划。我们还提出了AuthGuard-R,这是一种确定性授权层,它将每个可执行动作绑定到经过签名的任务、机器人身份、物体与区域范围、当前状态、时间以及输入来源。AuthGuard-R与独立的安全门协同运行,形成双门架构。我们在机器人轨迹上形式化了任务策略,定义了安全博弈,并证明了授权可靠性、任务不可升级性、抗重放性、机器人绑定、来源隔离、阈值审批安全、审计日志防篡改以及轨迹级组合安全性。我们基于Claude Haiku 4.5和开源的Qwen2.5 7B规划器进行了初步的跨模型评估。在240次实际攻击试验中,规划器在109次试验中遵循了注入的任务偏差;AuthGuard-R拒绝了全部109次由此产生的未授权动作。此外,一套由11个协议级和策略级攻击组成的手工构建测试套件也被完全拦截。
cs.RO / 56 / 2609.31112

DualManip: Agentic Dynamic Manipulation via Dual-Path Semantic Reasoning and Geometric Adaptation

DualManip:基于双路径语义推理与几何自适应的智能体动态操作
Li, Chengxi, Di, Yan, Li, Yingyue, Zhang, Ruida, Li, Mingyang, Ji, Xiangyang
Abstract
Vision-language models (VLMs) enable open-vocabulary reasoning for robot manipulation, but their high inference latency limits responsiveness in dynamic scenes. Many scene changes, however, alter object geometry without invalidating task intent. We present DualManip, a dual-path framework that decouples infrequent semantic reasoning from responsive geometric adaptation. The semantic path decomposes the task and grounds task-relevant interactions, followed by a constraint-solving module for pose optimization. During execution, the geometric path continuously updates template-to-observation correspondences from live RGB-D observations via a shape-adaptive network. These correspondences transfer task-relevant grasp contacts across observations, enabling online grasp reconstruction under object motion and non-rigid deformation. The Information Interaction Module bridges the two paths by initializing task-relevant grasps from semantic grounding, validating geometric updates, and triggering semantic replanning upon update failures. Real-world evaluation spans six manipulation tasks covering non-rigid deformation, articulated reconfiguration, rigid motion, and high-precision assembly across three settings: static, single-change, and continuous dynamic. DualManip demonstrates superior manipulation robustness, particularly under continuous scene changes, while achieving geometric adaptation approximately 46$\times$ faster than agentic verification and semantic replanning. Our project page: https://lichengxi1.github.io/Dualmanip.
Chinese Translation
视觉-语言模型(VLM)使机器人操作能够进行开放词汇推理,但其较高的推理延迟限制了在动态场景中的响应能力。然而,许多场景变化仅改变物体的几何形状,而不会使任务意图失效。我们提出了DualManip,一个双路径框架,它将低频的语义推理与响应迅速的几何自适应解耦。语义路径对任务进行分解并确定任务相关交互的对应关系,随后通过一个约束求解模块进行位姿优化。在执行过程中,几何路径通过一个形状自适应网络,从实时RGB-D观测中持续更新模板与观测之间的对应关系。这些对应关系将任务相关的抓取接触点在观测之间进行迁移,从而在物体运动和非刚性形变下实现在线抓取重建。信息交互模块(Information Interaction Module)连接两条路径:基于语义对应结果初始化任务相关抓取、验证几何更新,并在更新失败时触发语义重规划。真实世界评估涵盖六项操作任务,包括非刚性形变、铰接重构、刚性运动和高精度装配,并在静态、单次变化和连续动态三种设置下进行测试。DualManip展现出卓越的操作鲁棒性,尤其在连续场景变化下表现突出,同时其几何自适应速度比智能体验证与语义重规划快约46倍。我们的项目页面:https://lichengxi1.github.io/Dualmanip。
cs.RO / 57 / 2609.31117

Comparative Evaluation of an XR Pen-based Control Interface for Semi-Autonomous Mobile Robot Navigation in Service Environments

面向服务环境中半自主移动机器人导航的XR手写笔控制界面的对比评估
Torc, Alicia, Tornberg, Carl, Piette, Eric, Ronsse, Renaud, Macq, Benoit, Ricardez, Gustavo Alfonso Garcia, Hafi, Lotfi El, Taniguchi, Tadahiro
Abstract
Service robots remain difficult to deploy in domestic environments, partly because fully autonomous operation is not yet reliable in unpredictable surroundings, and partly because conventional control methods remain inaccessible to novice users. Extended Reality (XR) enables operators to visualize robot information overlaid onto the real world and to interact with augmented elements. Yet, common XR control methods, such as motion controllers and hand gestures, are still perceived as unintuitive. This paper presents a control interface that uses a commercial XR pen to command a semi-autonomous mobile robot in Augmented Reality (AR): the operator points at a position in the room, selects it, and drags an augmented arrow to set the desired orientation of the robot at this destination. Two additional interfaces, based on the XR motion controllers and hand gestures, were developed within the same framework. To assess the performance and users' perception of these interfaces, and of the XR pen in particular, a study with 10 participants compared four control methods, i.e., the XR pen, the XR motion controllers, hand gestures, and a computer-based baseline RViz, in navigation tasks performed in a home-like environment. Results show that the XR pen significantly outperforms the other methods in task selection time with the most consistent selections, and that the XR motion controllers obtain the best perceived workload and usability scores, ahead of the computer-based baseline, supporting XR-based control as an intuitive alternative for novice users. However, technical limitations in the integration of the recently released XR pen currently hold back its user experience.
Chinese Translation
服务机器人目前在家庭环境中的部署仍然困难,部分原因在于完全自主运行在不可预测的环境中尚不可靠,部分原因在于传统控制方法对新手用户而言仍难以使用。扩展现实(XR)使操作员能够将机器人信息叠加显示在真实世界中,并与增强元素进行交互。然而,常见的XR控制方法,如运动控制器和手势,仍被认为不够直观。本文提出了一种使用商用XR手写笔(XR pen)在增强现实(AR)中指挥半自主移动机器人的控制界面:操作员指向房间中的某个位置并选中它,然后拖动一根增强箭头以设定机器人在该目标点的期望朝向。我们在同一框架内还开发了另外两种界面,分别基于XR运动控制器和手势。为了评估这些界面(特别是XR手写笔)的性能和用户感知,一项共有10名参与者参与的研究在类家庭环境中执行导航任务,比较了四种控制方法,即XR手写笔、XR运动控制器、手势,以及基于计算机的基线方法RViz。结果表明,XR手写笔在任务选择时间上显著优于其他方法,且选择结果最为一致;而XR运动控制器在感知工作负荷和可用性评分上表现最佳,超过了基于计算机的基线方法,这支持了基于XR的控制作为新手用户直观替代方案的价值。然而,最近发布的XR手写笔在集成方面的技术局限目前仍制约了其用户体验。
cs.RO / 58 / 2609.31137

INTERACT: Interactive Planning for Autonomous Driving via Anchor-Conditioned Prediction and Trust-Region Refinement

INTERACT:基于锚点条件化预测与信赖域优化的自动驾驶交互式规划
Distelzweig, Aron, Look, Andreas, Janjoš, Faris, Hagedorn, Steffen, Palmieri, Luigi, Boedecker, Joschka
Abstract
Driving in dense urban traffic is interactive: whether a merge or an unprotected turn succeeds depends on how surrounding agents respond to the ego vehicle. Conventional planners predict first and plan second and, therefore, cannot account for this dependency. Methods that integrate prediction and planning either train both jointly, which introduces task interference, or keep them separate and are restricted to a predefined set of proposals. We present INTERACT: Interactive Planning for Autonomous Driving via Anchor-Conditioned Prediction and Trust-Region Refinement. Our key insight is that surrounding agents react to the intent a trajectory expresses rather than to its exact realization, so a single reactive prediction stays valid across an entire family of plans. INTERACT therefore decomposes interactive planning into prediction across driving intents and optimization within each intent. We derive a small set of diverse intents, which we call anchors, from map geometry, query a dedicated ego-conditioned prediction model once per anchor, and refine every anchor with the Cross-Entropy Method under a trust-region penalty that keeps the refined plan close enough to its anchor for the conditioned reaction to still apply. Prediction thus remains a separate model, avoiding task interference, while conditioning on anchors preserves the dependency. Because each anchor is refined continuously, the final plan is not restricted to the anchor set, yet INTERACT requires only one predictor query per anchor rather than one per candidate plan, with all anchors processed in parallel. On the nuPlan and interPlan closed-loop benchmarks, INTERACT sets a new state of the art, with the largest gains precisely in the interactive scenarios that motivate the method. The code will be released upon acceptance.
Chinese Translation
在密集的城市交通中驾驶本质上是交互式的:一次并线或一次无保护左转能否成功,取决于周围智能体对本车(ego vehicle)的响应方式。传统规划器遵循“先预测、后规划”的流程,因此无法考虑这种依赖关系。将预测与规划相整合的方法要么对二者进行联合训练,这会引入任务间的相互干扰;要么将其保持分离,从而只能局限于预定义的候选轨迹集合。我们提出了 INTERACT:一种通过锚点条件化预测(anchor-conditioned prediction)与信赖域优化(trust-region refinement)实现自动驾驶交互式规划的方法。我们的核心洞察是:周围智能体响应的是轨迹所表达的意图,而非其精确实现形式,因此单个反应式预测在整个规划轨迹族内均保持有效。据此,INTERACT 将交互式规划分解为两个阶段:跨驾驶意图的预测,以及在每个意图内的优化。我们从地图几何中推导出一小组多样化的意图(称为锚点,anchors),对每个锚点仅查询一次专用的本车条件化预测模型,并在信赖域惩罚约束下使用交叉熵方法(Cross-Entropy Method)对每个锚点进行优化——该惩罚确保优化后的规划与其锚点足够接近,从而使条件化预测的响应仍然适用。由此,预测仍保持为独立模型,避免了任务干扰,同时通过对锚点的条件化保留了交互依赖关系。由于每个锚点都被连续优化,最终规划并不局限于锚点集合;然而 INTERACT 每个锚点只需一次预测器查询,而非每个候选轨迹一次查询,且所有锚点可并行处理。在 nuPlan 和 interPlan 闭环基准测试中,INTERACT 创造了新的最先进水平,且其最大性能提升恰好出现在激发本方法研究动机的交互式场景中。代码将在论文被接收后发布。
cs.RO / 59 / 2609.31138

Evaluating the Impact of Adaptive Extended Reality on Human-Robot Interaction Across the Reality-Virtuality Continuum

评估跨现实-虚拟连续体的自适应扩展现实对人机交互的影响
Tornberg, Carl, Torck, Alicia, Hafi, Lotfi El, Taniguchi, Tadahiro
Abstract
As populations in developed countries age and labor shortages intensify, Cybernetic Avatars (CAs) are proposed to extend human capabilities through robotic embodiments, requiring effective Human-Robot Interaction (HRI) frameworks. Extended Reality (XR), an umbrella term for Augmented Reality (AR), Augmented Virtuality (AV), and Virtual Reality (VR), offers such interfaces, but prior research typically fixes the XR modality without evaluating its effect on task outcomes. This study examines whether the XR modality impacts HRI performance and whether an adaptive interface adjusting the level of virtuality along the Reality-Virtuality Continuum (RVC) at runtime improves it. A custom XR application interfaced with a mobile manipulator supports immersive control and runtime modality switching. In a within-participant multi-room pick-and-place experiment comparing fixed AR, AV, and VR with dynamic RVC through task metrics, the NASA-TLX, and the System Usability Scale (SUS), this study demonstrates that 1) the fixed reality modality affects HRI results, and 2) dynamically changing the modality along the RVC improves them. AR yielded significantly lower mental demand, effort, and frustration than AV and VR, while the dynamic RVC condition achieved the highest throughput and lowest workload, highlighting the value of adaptive XR interfaces for human-robot symbiosis. The implementation is available at https://github.com/CarlTornberg/XR-HRI.
Chinese Translation
随着发达国家人口老龄化和劳动力短缺加剧,赛博格化身(Cybernetic Avatars, CAs)被提出通过机器人化身来扩展人类能力,这需要有效的人机交互(Human-Robot Interaction, HRI)框架。扩展现实(Extended Reality, XR)是增强现实(AR)、增强虚拟(AV)和虚拟现实(VR)的统称,可提供此类接口,但以往研究通常固定XR模态,而未评估其对任务结果的影响。本研究考察XR模态是否会影响HRI性能,以及在运行时沿现实-虚拟连续体(Reality-Virtuality Continuum, RVC)调整虚拟程度的自适应接口是否能改善HRI性能。本研究开发了一个与移动操作机器人相连的定制XR应用,支持沉浸式控制和运行时模态切换。在一项被试内实验中,比较固定AR、AV、VR与动态RVC条件下的多房间取放任务,通过任务指标、NASA-TLX量表和系统可用性量表(SUS)进行评估。结果表明:1)固定的现实模态会影响HRI结果;2)沿RVC动态切换模态可改善结果。AR相比AV和VR显著降低了心理需求、努力程度和挫败感,而动态RVC条件实现了最高的任务吞吐量和最低的工作负荷,凸显了自适应XR接口对人机共生(human-robot symbiosis)的价值。实现代码可在 https://github.com/CarlTornberg/XR-HRI 获取。
cs.RO / 60 / 2609.31185

Onboard Wind-Preview Model Predictive Control Using Pitot-Static Sensing for Multirotor UAVs

基于皮托静压传感的多旋翼无人机机载风况预览模型预测控制
Meere, Bas, Wisse, Eline, Vousten, Laurens, Doodeman, Sander, Torta, Elena, Chanfreut, Paula, Antunes, Duarte
Abstract
Effective wind gust rejection and stable hovering are critical for the outdoor operation of autonomous drones. However, existing gust rejection methods are primarily reactive, inferring the disturbance from the resulting motion or measuring it at the airframe. Either way, the wind has already begun to act before it can be compensated. In this work, we anticipate the gust instead by measuring the wind ahead of the drone with a low-cost, low-weight pitot-static sensor mounted on a boom. The resulting wind preview is incorporated into a nonlinear model predictive controller (MPC), which optimizes the drone motion while anticipating wind disturbances. A longer boom offers more preview time but adds inertia and degrades flight performance. We characterize this trade-off in simulation and show that the optimal preview distance is not a fixed property of the platform, but shifts with the wind speed and with how quickly the drone can respond. Indoor hardware experiments confirm the trend and show that the proposed controller substantially improves hover performance against a PX4 baseline and an otherwise identical wind-unaware MPC. Outdoor experiments show that the error along the wind direction is reduced by 54 percent with respect to the baseline, demonstrating that a single wind-aligned sensor can significantly improve hovering performance.
Chinese Translation
有效的阵风抑制和稳定悬停对于自主无人机在户外环境中运行至关重要。然而,现有的阵风抑制方法主要是被动式的,即从无人机产生的运动中推断扰动,或直接在机体上测量扰动。无论哪种方式,在被补偿之前,风已经开始作用于无人机。在本工作中,我们转而对风况进行预判:在无人机前方的支撑臂上安装一个低成本、轻量化的皮托静压传感器,以测量无人机前方的风。所得到的风况预览信息被纳入非线性模型预测控制器(MPC)中,该控制器在预判风扰动的同时对无人机运动进行优化。更长的支撑臂能提供更多的预览时间,但会增加惯性并降低飞行性能。我们在仿真中对这一权衡进行了分析,结果表明最优预览距离并非平台的固定属性,而是随风速以及无人机响应速度的变化而变化。室内硬件实验验证了该趋势,并表明所提出的控制器相较PX4基线以及一个除不具备风感知能力外完全相同的MPC,显著提升了悬停性能。户外实验表明,沿风向的误差相较基线降低了54%,证明单个风向对齐的传感器即可显著提升悬停性能。
cs.RO / 61 / 2609.31207

Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling

通过以交互为中心的建模实现二指夹爪操作的统一跨域表征
Li, Guanlin, Bao, Shifeng, Zhao, Yihan, Shen, Haitao, Li, Haoyang, Zhao, Chen, Yang, Tong, Tang, Jie, Zhang, Jing
Abstract
Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverages the shared structure of two-finger grippers via a parameterized universal gripper abstraction, yielding a canonical gripper-frame representation. Given language and RGB-D observations, a VLM infers the subtask and grounds an interaction triplet (gripper, held, target), while SAM~2.1 tracks masks to reduce VLM queries. We design concise hybrid features that combine target/collision artificial potential fields for global guidance with segmented gripper-frame point clouds for local geometry, and use a Flow-Matching Transformer to predict smooth 7-DoF action chunks. Experiments in simulation and real-world tasks demonstrate that ours is the first imitation learning approach to simultaneously achieve competitive benchmark scores and extreme cross-embodiment/cross-viewpoint zero-shot sim-to-real transfer to completely distinct, heterogeneous robot platforms.
Chinese Translation
在模仿学习中实现鲁棒的跨本体泛化,需要克服一个关键的表征缺陷,即任务语义与硬件特定的视觉几何不可避免地纠缠在一起。我们提出了一种以交互为中心的框架,通过参数化的通用夹爪抽象来利用二指夹爪的共享结构,从而得到一种规范的夹爪坐标系表征。给定语言和RGB-D观测,一个视觉语言模型(VLM)推断子任务并 grounding 出交互三元组(夹爪、被抓持物体、目标物体),同时利用SAM~2.1进行掩码跟踪以减少VLM查询次数。我们设计了简洁的混合特征,将用于全局引导的目标/碰撞人工势场与用于局部几何的分割后的夹爪坐标系点云相结合,并使用流匹配Transformer(Flow-Matching Transformer)预测平滑的7自由度动作块。在仿真和真实世界任务中的实验表明,我们是首个同时实现有竞争力的基准分数和极端跨本体/跨视角零样本仿真到真实迁移的模仿学习方法,可迁移到完全不同的异构机器人平台。
cs.RO / 62 / 2609.31211

CoralPlan: Observation Skill Selection and Execution for Underwater Robotic Inspection

CoralPlan:面向水下机器人检测的观测技能选择与执行
Gao, Yuer, Zhao, Yu, Cai, Yi
Abstract
Underwater robotic inspection depends on acquiring views that reveal task-relevant structure. For a structurally complex coral colony, recognising the target is only the starting point: the robot must select and execute a viewing motion suited to the inspection task. We present CoralPlan, a vision-language system that selects an observation skill from a current camera image and task text supplied by an episode manifest. A shared motion interface executes orbit, patch, or survey as target-relative trajectories; the remaining plan fields provide operator guidance. Observation completion requires target keeping and primitive-specific coverage, while joint success also requires selection to match the recorded reference. We evaluate this interface in 144 simulated episodes and 36 matched simulation-hardware pairs. In a clear-water pool with external target-reference poses, hardware observation completion reaches 77.8% and joint success reaches 63.9%. The experiments identify both reference-mismatched completions and incomplete observations after a matching skill selection. These results connect observation-skill choice to measurable underwater execution outcomes and identify where task-directed acquisition succeeds or fails.
Chinese Translation
水下机器人检测依赖于获取能够揭示任务相关结构的视角。对于结构复杂的珊瑚群落而言,识别目标只是起点:机器人还必须选择并执行适合检测任务的观测动作。我们提出CoralPlan,一个视觉-语言系统,它根据当前相机图像和回合清单(episode manifest)提供的任务文本来选择观测技能。该系统通过统一的运动接口以目标相对轨迹的方式执行环绕(orbit)、局部拍摄(patch)或巡游(survey);规划中其余字段则为操作员提供指导。观测完成要求保持目标在视场内并满足特定基元(primitive)的覆盖要求,而联合成功还要求技能选择与记录的参考一致。我们在144个仿真回合以及36对匹配的仿真-实物实验中评估了该接口。在清水水池并提供外部目标参考位姿的条件下,实物观测完成率达到77.8%,联合成功率达到63.9%。实验既识别出了技能选择与参考不匹配的完成情况,也识别出了技能选择正确但观测不完整的情况。这些结果将观测技能选择与可测量的水下执行结果联系起来,并指明了任务导向的观测获取在何处成功或失败。
cs.RO / 63 / 2609.31225

Imp-ACT: Adaptive Impedance Control and Action Chunking with Transformers to Learn Contact-Rich Manipulation from Demonstrations

Imp-ACT:基于Transformer的自适应阻抗控制与动作分块方法,用于从演示中学习接触丰富的操作
Zanetti, Luca, Sirintuna, Doganay, Ozdamar, Idil, Balatti, Pietro, Zhang, Heng, Ajoudani, Arash
Abstract
Contact-rich manipulation requires robots to balance accurate motion tracking with compliant interaction, yet most visual-action policies leave compliance fixed at the controller level. We present Imp-ACT, a methodologically grounded and practical approach to incorporating direction-dependent Cartesian stiffness modulation directly into demonstration collection, without manual stiffness selection or offline target reconstruction. During teleoperation, a self-tuning impedance controller adapts stiffness along the instantaneous direction of motion while maintaining compliance in orthogonal directions. The adapted stiffness is applied and recorded alongside visual observations and motion commands, capturing motion and compliance under the same dynamics. We implement this pipeline using Action Chunking with Transformer (ACT) to predict end-effector pose, gripper action, and motion-direction stiffness from visual, proprioceptive, and wrench observations. The performance of Imp-ACT is evaluated on wiping and plug insertion using both success rate and quantitative measures of contact behavior. Compared with fixed low- and high-stiffness baselines, Imp-ACT achieves comparable or higher success while maintaining low interaction forces. In wiping, it reduces contact-force vibration by approximately $29\times$ relative to the compliant baseline and $180\times$ relative to the stiff baseline. In plug insertion, it reduces forces orthogonal to the insertion direction by $43\%$ relative to the better fixed-stiffness baseline. These results highlight the benefit of maintaining sufficient stiffness along the direction needed for task execution while preserving compliance in other directions to limit contact forces and accommodate environmental constraints.
Chinese Translation
接触丰富的操作要求机器人在精确的运动跟踪与柔顺交互之间取得平衡,然而大多数视觉-动作策略在控制器层面将柔顺性固定不变。我们提出Imp-ACT,这是一种有方法论支撑且实用的方法,可将依赖于方向的笛卡尔刚度调节直接融入演示采集过程,无需人工选择刚度或离线目标重建。在遥操作过程中,一种自整定阻抗控制器沿瞬时运动方向自适应调节刚度,同时在正交方向上保持柔顺性。调节后的刚度与视觉观测和运动指令一同被应用并记录,从而在同一动力学条件下捕捉运动与柔顺性信息。我们采用基于Transformer的动作分块(Action Chunking with Transformer, ACT)方法实现该流程,从视觉、本体感觉和力/力矩观测中预测末端执行器位姿、夹爪动作以及运动方向上的刚度。Imp-ACT的性能在擦拭和插头插入任务上通过成功率和接触行为的定量指标进行评估。与固定低刚度和固定高刚度基线相比,Imp-ACT在保持低交互力的同时实现了相当或更高的成功率。在擦拭任务中,相对于柔顺基线,它将接触力振动降低了约29倍;相对于刚性基线,降低了约180倍。在插头插入任务中,相对于较优的固定刚度基线,它将插入方向正交方向上的力降低了43%。这些结果凸显了以下策略的优势:沿任务执行所需方向保持足够的刚度,同时在其他方向上保持柔顺性,以限制接触力并适应环境约束。
cs.RO / 64 / 2609.31232

Cybflight: An Embedded Rust Autopilot for Aerial Robotics Research

Cybflight:面向空中机器人研究的嵌入式 Rust 自动驾驶仪
Lin, Yifan, Qin, Chao, Go, H S Helson, Liu, Hugh H. -T.
Abstract
Bringing an aerial robotics method from simulation to flight should not require rebuilding a mature autopilot or adding a companion computer. Cybflight is an open-source embedded Rust research autopilot whose typed, replaceable interfaces connect hardware access, perception, state estimation, trajectory planning, and control. This modular development and compile-time optimization workflow is demonstrated with replaceable Rust implementations of model predictive contour-tracking control (MPCTC) and incremental nonlinear dynamic inversion (INDI) running on one STM32H743 without a companion computer. Using this configuration, the vehicle reaches 12.38 m/s during indoor flight, while an outdoor flight using global navigation satellite system (GNSS) position updates reaches 31.4 m/s. These flights show that running demanding estimation and nonlinear control entirely on a flight-controller microcontroller need not come at the expense of a modular autopilot structure.
Chinese Translation
将空中机器人方法从仿真推向实际飞行,不应需要重构成熟的自动驾驶仪或添加伴随计算机。Cybflight 是一个开源的嵌入式 Rust 研究型自动驾驶仪,其类型化、可替换的接口将硬件访问、感知、状态估计、轨迹规划与控制连接起来。通过在单个 STM32H743 上运行模型预测轮廓跟踪控制(MPCTC)和增量非线性动态逆(INDI)的可替换 Rust 实现,且无需伴随计算机,验证了这一模块化开发与编译期优化的工作流程。在该配置下,飞行器在室内飞行中达到 12.38 m/s,而在使用全球导航卫星系统(GNSS)位置更新的室外飞行中达到 31.4 m/s。这些飞行实验表明,在飞行控制器微控制器上完全运行高计算量的估计与非线性控制,并不需要以牺牲模块化的自动驾驶仪结构为代价。
cs.RO / 65 / 2609.31313

Towards VLA-Dreamer: Refining VLA Behavior Using World Models

迈向VLA-Dreamer:利用世界模型改进VLA行为
Kashani, Parsa Mastouri, Habekost, Jan-Gerrit, Wermter, Stefan
Abstract
Vision-Language-Action models (VLAs), while showing strong potential for robot control, require massive amounts of high-quality imitation learning data. Moreover, the absence of an explicit world model casts further doubt on their control capabilities. In this concept paper, we propose a novel architecture that addresses sample efficiency in VLAs by training a predictive world model on the embedding space of the VLA's vision encoder. We hypothesize that these embeddings are action-relevant and usable for future prediction. To this end, we propose using the suggested architecture to investigate how well these embeddings predict the future based on actions, as the inability to do so would mark a key limitation of VLA architectures: the lack of a non-lossy implicit world model to simulate real-world dynamics. The proposed architecture differs from the standard world model dynamics as the loss comes from the embedding space rather than the pixel space, similar to joint embedding predictive architectures. Furthermore, the trained world model can be utilized for short-term planning tasks by sampling VLA actions given goal images. We intend to examine the richness of vision embeddings in VLAs and reduce their high data requirements through a world model that can also generate plans during inference.
Chinese Translation
视觉-语言-动作模型(Vision-Language-Action models,VLAs)在机器人控制方面展现出强大潜力,但其需要海量高质量的模仿学习数据。此外,由于缺乏显式的世界模型,其控制能力进一步受到质疑。在本概念性论文中,我们提出了一种新颖的架构,通过在VLA视觉编码器的嵌入空间上训练预测性世界模型,来解决VLA的样本效率问题。我们假设这些嵌入与动作相关,并可用于未来预测。为此,我们建议使用所提出的架构来研究这些嵌入基于动作预测未来的能力;若无法做到这一点,则将标志着VLA架构的一个关键局限:缺乏能够模拟真实世界动力学的无损隐式世界模型。所提出的架构与世界模型的标准动力学学习方式不同,其损失来自嵌入空间而非像素空间,类似于联合嵌入预测架构(joint embedding predictive architectures)。此外,训练好的世界模型可通过在给定目标图像的条件下采样VLA动作,用于短期规划任务。我们旨在考察VLA中视觉嵌入的丰富性,并通过一个能够在推理过程中生成规划的世界模型,降低其高数据需求。
cs.RO / 66 / 2609.31323

See to Reach, Feel to Grasp: Learning A Blind Grasp Reflex for Anthropomorphic Robotic Hands

视以达物,感以抓握:为类人灵巧手学习一种盲抓握反射
Alexiev, Alexander, Lin, Tzu-Yuan, Kim, Sang Min, Lee, Ho Jae, Lee, Yonghyeon, Kim, Sangbae
Abstract
In this work we study if a robotic hand using proprioception alone can grasp diverse objects with no visual observation. We present a modular dexterous grasping architecture that separates global arm motion from local contact control. An independently controlled arm guides the hand toward the object, while a reinforcement learning policy grasps and stabilizes it using only hand proprioceptive feedback. We call this \textit{a blind grasp reflex}: grasping without images, object poses, or geometric observations. A learned stable-grasp score determines when the object is securely held, allowing the arm to begin post-grasp manipulation. This separation makes grasping a reusable hand-level skill that can be combined with independently designed arm controllers for various manipulation tasks. Experiments in simulation and on hardware demonstrate robust blind grasping across diverse objects and seamless composition with a range of arm controllers. Moreover, despite never observing contact geometry, the learned grasp score closely aligns with an independent physics-based measure of grasp stability. The resulting approach follows a simple principle: see to reach, feel to grasp. Project page: https://blindgraspreflex.github.io.
Chinese Translation
在本工作中,我们研究了机器人在没有视觉观测的情况下,仅依靠本体感觉(proprioception)能否抓取多样化的物体。我们提出了一种模块化的灵巧抓握架构,将机械臂的全局运动与局部接触控制分离。一个独立控制的机械臂引导灵巧手接近物体,而一个强化学习策略仅利用手部的本体感觉反馈来抓取并稳定物体。我们将此称为“盲抓握反射”(blind grasp reflex):即在不使用图像、物体位姿或几何观测的情况下进行抓取。一个学习到的稳定抓握评分用于判断物体是否被牢固握持,从而使机械臂能够在抓取完成后开始后续操作。这种分离使得抓握成为一项可复用的手部级技能,可与独立设计的机械臂控制器组合用于各种操作任务。在仿真和硬件上的实验表明,该方法能够在多样化物体上实现鲁棒的盲抓握,并可与多种机械臂控制器无缝组合。此外,尽管从未观测到接触几何信息,学习到的抓握评分与独立的基于物理的抓握稳定性度量高度一致。该方法遵循一个简单的原则:视以达物,感以抓握。项目页面:https://blindgraspreflex.github.io。
cs.RO / 67 / 2609.31337

Representation-Guided Generation and Integration of Executable Programs for Robot Manipulation

面向机器人操作的表示引导的可执行程序生成与集成
Yang, Ruixiao, Yu, Mingxin, Fan, Chuchu
Abstract
Building a robotic manipulation system requires connecting perception, planning, and control through carefully designed representations and interfaces. VLM code generation offers a way to automate this construction, but independently generated components may operate on incompatible geometric and task-level information. We present Representation-guided Integration of VLM-generated Executable Task programs (RIVET), a framework for generating complete manipulation systems around a shared object-centric representation. The representation combines per-object 6D poses, which preserve the metric information required for action grounding, with a relation graph that exposes the task-level structure required for planning. Guided by this representation, a VLM generates cooperating perception, rendering, relation-inference, and planning programs, each combining task-specific computation with available packages where useful. The resulting programs are authored once for a manipulation domain and reused on unseen start and goal configurations without code regeneration. We evaluate RIVET on cube stacking, tangram rearrangement, and three-dimensional assembly in simulation and on a physical robot, where we achieve 83% overall success rate in the real world by reusing offline-generated systems. Our results demonstrate that representation-guided program generation can adapt a common manipulation framework to tasks with different geometric, relational, and sequential requirements.
Chinese Translation
构建机器人操作系统需要通过精心设计的表示和接口将感知、规划与控制连接起来。视觉语言模型(VLM)代码生成为自动化这一构建过程提供了一种途径,但独立生成的组件可能在几何信息和任务级信息上互不兼容。我们提出了表示引导的VLM生成可执行任务程序集成框架(Representation-guided Integration of VLM-generated Executable Task programs, RIVET),该框架围绕一个共享的以物体为中心的表示来生成完整的操作系统。该表示结合了每个物体的6D位姿(保留了动作落地所需的度量信息)和关系图(暴露了规划所需的任务级结构)。在该表示的引导下,VLM生成协同工作的感知、渲染、关系推断和规划程序,每个程序在有用之处将任务特定的计算与现有软件包相结合。这些程序只需针对某一操作领域编写一次,即可在未见过的起始和目标配置上复用,无需重新生成代码。我们在仿真和真实机器人上对RIVET进行了评估,任务包括立方体堆叠、七巧板重排和三维装配,通过复用离线生成的系统,我们在真实世界中取得了83%的总体成功率。我们的结果表明,表示引导的程序生成能够使统一的操作框架适应具有不同几何、关系和时序要求的任务。
cs.RO / 68 / 2609.31357

Transformer-based Monte Carlo Localization in Construction Meshes

基于Transformer的施工网格蒙特卡洛定位
Kramer, Linus, Talbot, William, Vysotska, Olga, Hutter, Marco
Abstract
To be able to perform inspection or digitization tasks, mobile robots on construction sites must be able to localize themselves reliably with respect to a global reference frame that is shared with a building map. Similar room layouts and low-texture surfaces pose a challenge for existing LiDAR- and vision-based localization methods. We approach this problem with a LiDAR-based global relocalization system that estimates the robot's pose relative to a building mesh and combines a PointNet++ encoder with a place recognition decoder, whose outputs serve as a learned observation model within a Monte Carlo Localization (MCL) framework. The pipeline is trained exclusively on synthetic LiDAR scans obtained by simulating the robot's sensors inside the building mesh. Our approach is robust in ambiguous environments due to an uncertainty-aware decoder that scales positional likelihoods and a resampling strategy that injects model hypotheses into the particle set, enabling recovery from potential particle depletion. Evaluations on real-world datasets show that our method outperforms both diffusion-based and ScanContext++ baselines while maintaining fast inference (18 ms per call), demonstrating the practicality of synthetic-data training for mesh-referenced global localization in construction robotics.
Chinese Translation
为了执行巡检或数字化任务,建筑工地上的移动机器人必须能够在与建筑地图共享的全局参考坐标系下可靠地实现自身定位。相似的房间布局和低纹理表面对现有的基于激光雷达(LiDAR)和视觉的定位方法构成了挑战。我们提出了一种基于激光雷达的全局重定位系统来解决这个问题,该系统估计机器人相对于建筑网格(mesh)的位姿,将PointNet++编码器与位置识别解码器相结合,其输出作为蒙特卡洛定位(MCL)框架内学习得到的观测模型。该流程完全通过在建筑网格内模拟机器人传感器所获得的合成激光雷达扫描数据进行训练。得益于一个可缩放位置似然的不确定性感知解码器,以及一种向粒子集中注入模型假设的重采样策略(使系统能够从潜在的粒子耗尽中恢复),我们的方法在模糊环境中具有鲁棒性。在真实世界数据集上的评估表明,我们的方法优于基于扩散的基线和ScanContext++基线,同时保持快速推理(每次调用18毫秒),证明了合成数据训练在建筑机器人网格参考全局定位中的实用性。
cs.RO / 69 / 2609.31374

RECAST: From Log Replay to Closed-Loop Driving Simulation with View-Complete Actors

RECAST:从日志回放到具备视角完备参与者(Actor)的闭环驾驶仿真
Zhao, Zijun, Liao, Liewen, Shen, Kang, Zhang, Songan, Yang, Ming
Abstract
Closed-loop driving simulation requires rendered observations to remain reliable as the ego vehicle and surrounding actors move beyond their recorded trajectories, exposing views absent from the source log. Existing data-driven simulators reconstruct dynamic actors from sparse observations, which can result in rendering artifacts under these viewpoint changes. We introduce RECAST (REconstructing Controllable Actors for Simulation and Testing), a 3D Gaussian Splatting framework that generates a view-complete actor from a single segmented vehicle observation in a driving log and registers the generated actor in the reconstructed scene. RECAST supports planner-in-the-loop rendering under controlled ego-actor interactions. To adapt an image-to-3D prior to real vehicles, we further introduce RECAR, a dataset of approximately 20K real vehicles with 600K background-free RGBA images spanning diverse vehicle colors and types. We use two-stage adaptation to improve vehicle generation from real driving-log observations. At the actor level, RECAST reduces $\mathrm{FD}_{\mathrm{incep}}$ from 9.788 to 7.992 relative to unadapted TRELLIS. At the scene level, under actor motion beyond logged trajectories, RECAST reduces $\mathrm{FD}_{\mathrm{incep}}$ from 129.35 to 112.10 and increases $\mathrm{CLIP}_{\mathrm{margin}}$ ($\times1000$) from 0.14 to 3.47 relative to Street Gaussians. We demonstrate planner-in-the-loop simulation with the image-conditioned planner GTRS-Dense. Compared with native Street Gaussians actors, RECAST increases the no-collision (NC) rate from 22.2% (12/54) to 63.0% (34/54) and the mean minimum predicted time-to-collision (TTC) from 0.798 s to 2.150 s. These experiments show that RECAST supports closed-loop planner evaluation under controlled ego-actor interactions beyond log replay. Visit our project page at https://zijunkr.github.io/RECAST/
Chinese Translation
闭环驾驶仿真要求当自车及周围交通参与者(Actor)偏离其记录轨迹、暴露出源日志中缺失的视角时,渲染的观测结果仍然保持可靠。现有的数据驱动仿真器从稀疏观测中重建动态参与者,这在视角变化下可能产生渲染伪影。我们提出了 RECAST(REconstructing Controllable Actors for Simulation and Testing,面向仿真与测试的可控参与者重建),这是一个 3D 高斯泼溅(3D Gaussian Splatting)框架,能够从驾驶日志中的单个分割车辆观测生成视角完备的参与者,并将生成的参与者注册到重建的场景中。RECAST 支持在受控自车-参与者交互下的规划器在环(planner-in-the-loop)渲染。为了将图像到 3D 的先验模型适配到真实车辆,我们进一步构建了 RECAR 数据集,包含约 2 万辆真实车辆和 60 万张无背景 RGBA 图像,覆盖多样的车辆颜色与类型。我们采用两阶段适配方法来改进从真实驾驶日志观测的车辆生成。在参与者层面,相对于未适配的 TRELLIS,RECAST 将 $\mathrm{FD}_{\mathrm{incep}}$ 从 9.788 降低到 7.992。在场景层面,在参与者运动超出日志轨迹的情况下,相对于 Street Gaussians,RECAST 将 $\mathrm{FD}_{\mathrm{incep}}$ 从 129.35 降低到 112.10,并将 $\mathrm{CLIP}_{\mathrm{margin}}$($ imes1000$)从 0.14 提升到 3.47。我们基于图像条件化规划器 GTRS-Dense 演示了规划器在环仿真。与原生的 Street Gaussians 参与者相比,RECAST 将无碰撞(NC)率从 22.2%(12/54)提升到 63.0%(34/54),并将平均最小预测碰撞时间(TTC)从 0.798 秒提升到 2.150 秒。这些实验表明,RECAST 能够支持超出日志回放的、受控自车-参与者交互下的闭环规划器评估。项目主页请访问 https://zijunkr.github.io/RECAST/
cs.RO / 70 / 2609.31383

Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization

基于端点约束轨迹优化的端到端驾驶模型引导方法
Zhang, Brayden, Golchoubian, Mahsa, Gilitschenski, Igor, Ivanovic, Boris, Chitta, Kashyap
Abstract
End-to-end driving policies are commonly trained through open-loop behavior cloning, yet they must ultimately operate in closed-loop when deployed on a vehicle, creating a fundamental mismatch between training and execution. Beyond the commonly studied effects of covariate shift and causal confusion, we identify a complementary factor for this open-loop/closed-loop gap: waypoint-based supervision and displacement metrics do not ensure that the intermediate trajectory is physically coherent or easy for the controller to track. We observe that these inconsistencies concentrate primarily at intermediate waypoints, while the predicted endpoint remains comparatively reliable. Based on this observation, we introduce Endpoint-Constrained Optimization (ECO), a lightweight postprocessing layer that anchors the trajectory to the vehicle's executed history, preserves the policy's predicted endpoint, and reshapes the intermediate waypoints to improve feasibility. ECO requires no map, privileged simulator state, or additional training, and can be inserted between a broad range of waypoint-emitting policies and their controllers. Across two closed-loop simulators, it improves the aggregate closed-loop score of all six evaluated generative and regression-based driving policies, and the gains tend to increase with how often the base plans violate motion limits. On HUGSIM, ECO improves VaVAM from 18.1 to 31.0 HD-Score (+71%), achieving 1st place on the HUGSIM Closed-Loop Driving Challenge. Similarly, on AlpaSim, ECO increases the scene scores of VaVAM and DiffusionDrive by 123% and 22%, respectively. These results show that for a broad collection of end-to-end driving models, repairing the intermediate geometry of predicted trajectories without changing the policy's predicted endpoint can substantially improve closed-loop performance.
Chinese Translation
端到端驾驶策略通常通过开环行为克隆进行训练,但部署到车辆上时最终必须以闭环方式运行,这在训练与执行之间造成了根本性的不匹配。除了已被广泛研究的协变量偏移和因果混淆效应之外,我们识别出导致这种开环/闭环差距的另一个互补因素:基于路径点的监督和位移度量并不能保证中间轨迹在物理上连贯一致,也不能保证其易于被控制器跟踪。我们观察到这些不一致性主要集中在中间路径点上,而预测的端点则相对可靠。基于这一观察,我们提出了端点约束优化(Endpoint-Constrained Optimization, ECO),这是一种轻量级的后处理层,它将轨迹锚定于车辆的已执行历史,保留策略预测的端点,并重塑中间路径点以提升可行性。ECO 不需要地图、特权仿真器状态或额外的训练,并且可以插入到广泛的输出路径点的策略与其控制器之间。在两个闭环仿真器中,ECO 提升了所评估的全部六种生成式和基于回归的驾驶策略的聚合闭环得分,且当基础规划违反运动约束的频率越高时,增益往往越大。在 HUGSIM 上,ECO 将 VaVAM 的 HD-Score 从 18.1 提升至 31.0(+71%),在 HUGSIM 闭环驾驶挑战赛中取得第一名。同样,在 AlpaSim 上,ECO 分别将 VaVAM 和 DiffusionDrive 的场景得分提升了 123% 和 22%。这些结果表明,对于广泛的端到端驾驶模型而言,在不改变策略预测端点的前提下修复预测轨迹的中间几何形状,可以显著提升闭环性能。
cs.RO / 71 / 2609.31386

Modeling and Generative-AI-Based Design of Load-Adaptive Gravity Balancing Mechanisms

负载自适应重力平衡机构的建模与基于生成式AI的设计
Kayawake, Ryotaro, Abe, Kazuki, Miyake, Shota, Watanabe, Masahiro, Tadakuma, Kenjiro
Abstract
Load-adaptive gravity balancing mechanisms (LA-GBMs) can accommodate various loading conditions by passively changing their characteristics in response to payload variations. However, their design is difficult because both the desired mechanism motion and static equilibrium under variable payloads must be satisfied simultaneously. This study proposes a general design methodology for LA-GBMs that does not depend on specific mechanism architectures or mechanical elements. The necessary conditions for the potential fields of LA-GBMs are formulated, and two general forms are derived: an affine form representing the effect of payload mass and a factorized form representing state transitions associated with load adaptation and gravity balancing. These forms are then provided to generative AI as design requirements to generate candidate potential functions. The generated functions are analytically verified in terms of their conformity to the two general forms and the conditions required for valid LA-GBMs. Furthermore, the obtained potential functions are decomposed into individual terms, and an example of a method for constructing an LA-GBM by combining springs, counterweights, and function-generating linkage mechanisms is presented. By using potential functions as an intermediate representation, the proposed framework enables the generation of LA-GBM design candidates without prescribing a mechanism architecture in advance. Mechanical realizability and manufacturability of the generated potential fields remain important issues for future work.
Chinese Translation
负载自适应重力平衡机构(LA-GBM)能够通过被动改变其自身特性来适应载荷变化,从而应对多种负载条件。然而,其设计十分困难,因为必须同时满足期望的机构运动和在可变载荷下的静力平衡。本研究提出了一种通用的LA-GBM设计方法,该方法不依赖于特定的机构结构或机械元件。本文建立了LA-GBM势场的必要条件,并推导出两种通用形式:表示载荷质量影响的仿射形式,以及表示与负载自适应和重力平衡相关的状态转移的因式分解形式。随后,将这些形式作为设计需求提供给生成式AI,以生成候选势函数。所生成的函数通过解析方式验证了其对两种通用形式以及有效LA-GBM所需条件的符合性。此外,将获得的势函数分解为若干独立项,并给出了一种通过组合弹簧、配重和函数生成连杆机构来构建LA-GBM的方法示例。通过将势函数作为中间表示,所提出的框架能够在不预先指定机构结构的情况下生成LA-GBM设计方案。所生成势场的机械可实现性与可制造性仍是未来工作的重要课题。
cs.RO / 72 / 2609.31394

InternW0-$\Delta$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

InternW0-$\Delta$:一个连接预测动力学与动作的世界动作模型(基于2万+小时开放数据)
Miao, Xingyu, Li, Zizun, Fang, Baole, Song, Kaiwen, Wang, Tenghui, Zhang, Hanxue, Wang, Yating, Li, Xudong, He, Yuping, Wei, Xueyuan, Gao, Chao, Yang, Xijie, Xu, Yingxiang, Ren, Kerui, Guo, Wenqi, Zhou, Jianjun, Wang, Xinzhe, Zhao, Weiguang, Yang, Ni, Cai, Zetao, Xue, Yufei, Li, Hengjie, He, Zeyu, Zhou, Yuanzhen, Fu, Rong, Zhang, Jianyang, Cui, Siwei, Huang, Fuxian, Zhou, Yunsong, Gao, Xing, Yao, Yifei, Yu, Qiaojun, Li, Kailin, Zhou, Ming, Huang, Mu, Li, Xinyue, Cui, Wenze, Jiang, Bingqi, Zhu, Xueyue, Dong, Junting, Guo, Haoyu, Lu, Tao, Yu, Mulin, Zhou, Bowen, Zhao, Bin, Xue, Tianfan, Zhang, Weinan, Shen, Chunhua
Abstract
World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified framework for robot action generation. We introduce InternW0-$\Delta$, a unified WAM pretrained on a heterogeneous corpus that outperforms prior methods across simulation benchmarks and real-robot platforms. InternW0-$\Delta$ combines pretrained visual dynamics, scene-level semantics, 4D geometric and motion priors, and action generation within a Mixture-of-Transformers (MoT) framework. A pretrained video expert and an action expert interact under semantic guidance from a frozen VLM, while a pretrained 4D foundation model injects geometric and motion priors through training-only distillation. We further introduce Causal Imprint, which learns future-relevant scene changes from training-only future supervision and provides predictive representations directly to the action expert without future-video rollout at inference. For large-scale joint training, we construct a heterogeneous corpus of robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data, curated and aligned under a common state-action representation. The resulting corpus contains over 20K hours of processed training data, to our knowledge the largest open-source corpus of its kind. We pretrain InternW0-$\Delta$ on this corpus and demonstrate strong performance across simulation benchmarks and real-robot platforms. We will open source the training code, model weights, infrastructure, data-processing pipeline, and processed data where licenses permit. Project page: https://internrobotics.github.io/InternW0-Delta/
Chinese Translation
世界动作模型(World Action Models, WAMs)联合建模视觉动力学与动作生成,以实现通用机器人操作。其核心挑战在于如何将大规模预训练模型的先验知识——包括视觉动力学、场景语义、几何与运动信息——整合到统一的机器人动作生成框架中。我们提出了InternW0-$\Delta$,这是一个在异构语料库上预训练的统一世界动作模型,在仿真基准和真实机器人平台上均优于已有方法。InternW0-$\Delta$在混合Transformer(Mixture-of-Transformers, MoT)框架内结合了预训练的视觉动力学、场景级语义、4D几何与运动先验以及动作生成。预训练的视频专家与动作专家在冻结的视觉语言模型(VLM)的语义引导下交互,同时预训练的4D基础模型通过仅训练阶段的知识蒸馏注入几何与运动先验。我们进一步提出了Causal Imprint机制,它从仅训练阶段的未来监督中学习与未来相关的场景变化,并将预测性表征直接提供给动作专家,从而在推理时无需未来视频的滚动预测。为进行大规模联合训练,我们构建了一个异构语料库,包含机器人演示数据、UMI数据、以自我为中心的人类演示数据以及Ego2Robot数据,并在统一的状态-动作表征下进行了整理与对齐。该语料库包含超过2万小时的已处理训练数据,据我们所知是同类中最大的开源语料库。我们在该语料库上对InternW0-$\Delta$进行预训练,并展示了其在仿真基准和真实机器人平台上的强大性能。我们将在许可证允许的范围内开源训练代码、模型权重、基础设施、数据处理流程及已处理的数据。项目页面:https://internrobotics.github.io/InternW0-Delta/
cs.RO / 73 / 2609.31396

Augmented Reality Interfaces for Human-Robot Collaboration: Development of a ROS 2-Based Sensor Streaming Framework and Validation via SLAM Algorithms

面向人机协作的增强现实界面:基于ROS 2的传感器数据流框架开发及SLAM算法验证
Rubert, Alessandro, Ghidoni, Stefano, Terreran, Matteo
Abstract
In recent years, Human-Robot Collaboration (HRC) has taken on a central role in Industry 4.0 and collaborative robotics, demanding communication channels that are increasingly bidirectional, intuitive, and efficient. In this context, Augmented Reality (AR) presents itself as a fundamental enabling technology, capable of both displaying information to the operator and gathering spatial data about the surrounding environment. This thesis presents the development of a sensor streaming framework that connects the Magic Leap 2 AR headset with the ROS 2 (Robot Operating System) ecosystem. Using the Unity development environment and the ROSTCP-Connector package, an on-board application for the headset was developed, capable of acquiring real-time data from the integrated sensors (pose tracking, cameras, and environmental sensors) and publishing it to dedicated ROS 2 topics. In order to test the accuracy, latency, and robustness of the generated data stream, the framework was validated using SLAM (Simultaneous Localization and Mapping) algorithms known in the literature. The experimental results demonstrate that the proposed architecture ensures stable data transmission, laying the groundwork for safe real-time interaction and shared spatial awareness, and opening up new perspectives for the control and supervision of robotic systems in complex HRC scenarios.
Chinese Translation
近年来,人机协作(Human-Robot Collaboration, HRC)在工业4.0和协作机器人领域占据了核心地位,这就要求通信通道日益具备双向性、直观性和高效性。在此背景下,增强现实(Augmented Reality, AR)作为一种关键的使能技术应运而生,它既能够向操作员显示信息,又能够采集周围环境的空间数据。本文介绍了一个连接Magic Leap 2 AR头戴式显示器与ROS 2(机器人操作系统)生态系统的传感器数据流框架的开发工作。借助Unity开发环境和ROSTCP-Connector软件包,开发了头显的机载应用程序,该程序能够从集成传感器(位姿追踪、摄像头和环境传感器)中获取实时数据,并将其发布到专用的ROS 2话题上。为了测试所生成数据流的精度、延迟和鲁棒性,该框架采用文献中已知的SLAM(同步定位与建图)算法进行了验证。实验结果表明,所提出的架构能够确保稳定的数据传输,为安全的实时交互和共享空间感知奠定了基础,并为复杂HRC场景下机器人系统的控制与监控开辟了新的前景。
cs.RO / 74 / 2609.31412

dRVG: Quadtree-Guided, Resolution-Complete Online Motion Planning for Polygonal Robots in Unknown Environments

dRVG:基于四叉树引导的、面向多边形机器人在未知环境中的分辨率完备在线运动规划
Zhang, Duo, Zhang, Hechen, Huang, Junshan, Yu, Jingjin
Abstract
We present the dynamic rotation-stacked visibility graph (dRVG), an online motion planner that guides polygonal robots to specified goals in initially unknown, static environ- ments. It merges local roadmaps from successive observations to plan collision-free translations and rotations without a uniform position grid. A spatial quadtree schedules sensing goals across regions to reduce repeated visits while retaining all orientation configurations for routing. Under exact sensing and geometric computation and star-shaped robot and envelope assumptions, dRVG with center scans is resolution-complete relative to full- map RVG at the same angular resolution. In experiments using footprint scans, dRVG solves all 140 cases across 20 difficult maps and seven angular resolutions within a 20 s planning budget, with a median planning time of 1.18 s at 360 orientation layers. Six microMVP demonstrations illustrate the complete online planning loop on a physical robot.
Chinese Translation
我们提出了动态旋转堆叠可见图(dynamic rotation-stacked visibility graph,dRVG),这是一种在线运动规划器,可引导多边形机器人在初始未知的静态环境中到达指定目标。它将连续观测得到的局部路线图进行合并,从而在不使用均匀位置网格的情况下规划无碰撞的平移和旋转。该算法利用空间四叉树在各区域间调度感知目标,以减少重复访问,同时保留用于路径搜索的所有朝向配置。在精确感知与几何计算、以及机器人和外包络均呈星形的假设条件下,采用中心扫描的 dRVG 在相同角度分辨率下相对于全地图 RVG 具有分辨率完备性。在采用足迹扫描的实验中,dRVG 在 20 秒规划预算内解决了 20 张困难地图和 7 种角度分辨率下的全部 140 个案例,在 360 个朝向层下的规划时间中位数为 1.18 秒。六个 microMVP 演示展示了在实体机器人上完整的在线规划流程。
cs.RO / 75 / 2609.31418

CognitiveReality: Robot-Agnostic Semantic Gaussian Mapping with an LLM Agent for Immersive Collaborative VR Teleoperation

CognitiveReality:基于LLM智能体的机器人无关语义高斯建图与沉浸式协作VR遥操作
Kozlov, Timofei, Maliukov, Dmitrii, Marchenko, Andrey, Plotnikov, Dmitrii, Cabrera, Miguel Altamirano, Tsetserukou, Dzmitry
Abstract
A photorealistic 3D view tells a teleoperator where a robot is, but not what the scene contains, how well each object has been observed, or how to turn pointing and speech into robot action. CognitiveReality turns a robot's RGB-D stream into a live, semantically indexed Gaussian-TSDF map shared by an operator in virtual reality and a tool-using language agent. One mapper binary serves any platform through configuration alone: it ingests poses from robot SLAM, joint kinematics, motion capture or an inline visual tracker, bridges localization outages with a shadow tracker and keyframe-anchored PnP, and maintains open-vocabulary instance identities with per-object quality at 2 Hz. Speech and controller rays are grounded against persistent scene objects through validated typed tools and operator-confirmed robot actions. In the controlled agent evaluation, the deployed local Qwen3-VL-8B router reaches 81.24\% tool exact match, while merge-aware replay correctly redirects 101 absorbed object identifiers. On robot data CognitiveReality exceeds a Gaussian-plus-SDF baseline by 2-8 dB; pose error through 5-40 s SLAM outages stays within 1-8 cm. Deployed live on two quadrupeds, the agent executed 26 of 30 navigation requests and 20 of 20 re-observation requests, raising object quality by 2-5 dB.
Chinese Translation
逼真的三维视图能让遥操作者知道机器人在哪里,但无法告知场景中包含什么、每个物体被观测的充分程度,以及如何将指向和语音转化为机器人动作。CognitiveReality将机器人的RGB-D数据流转化为一张实时、语义索引的高斯-TSDF地图,由虚拟现实中的操作者与一个具备工具使用能力的语言智能体共享。该建图器以单一二进制程序通过纯配置即可适配任何平台:它可接收来自机器人SLAM、关节运动学、动作捕捉或内联视觉跟踪器的位姿,利用影子跟踪器和关键帧锚定的PnP在定位中断期间保持连续性,并以2 Hz的频率维护开放词汇的实例身份及每个物体的观测质量。通过经过验证的类型化工具和操作者确认的机器人动作,语音和控制器的视线射线被关联到持久的场景物体上。在受控的智能体评估中,部署的本地Qwen3-VL-8B路由器达到81.24%的工具精确匹配率,而具备合并感知的重放机制正确重定向了101个被合并的物体标识符。在机器人数据上,CognitiveReality超出高斯加SDF基线2-8 dB;在5-40秒的SLAM中断期间,位姿误差保持在1-8 cm以内。在两台四足机器人上的实时部署中,该智能体成功执行了30个导航请求中的26个和20个重观测请求中的全部20个,使物体观测质量提升了2-5 dB。
cs.RO / 76 / 2609.31434

ExoLaN: Physics-Consistent Context-Aware Dynamics Learning for Exoskeletons

ExoLaN:面向外骨骼的物理一致且上下文感知的动力学学习
Schulze, Lucas, Schwarz, Maximilian, Hoppe, Jona, Peters, Jan, Arenz, Oleg
Abstract
Task-agnostic assistive exoskeleton control based on human intention offers greater flexibility than conventional approaches that rely on predefined tasks or motion patterns. Human joint torque estimation enables task-agnostic assistance by characterizing user actions. Physics-consistent methods such as Deep Lagrangian Networks (DeLaN) have been applied to estimate the human torques in multi-user settings, but existing approaches cannot adapt to a specific user without retraining, and do not account for intermittent contacts during locomotion. We propose ExoLaN, a Context-Aware DeLaN for human-exoskeleton interaction that learns the full coupled system dynamics while adapting to changes in interaction context. ExoLaN combines temporal context with partial contact-force measurements from force-sensitive insoles to infer latent dynamics embeddings and estimate generalized contact torques. On seven unseen users performing 21 unseen tasks, ExoLaN reduces torque estimation MSE by 7% compared to a black-box baseline. Beyond inverse dynamics, ExoLaN serves as a unified model that also enables accurate forward prediction: training with a multi-step prediction loss reduces acceleration MSE by 59% and long-horizon position and velocity errors by 60% and 93%, respectively, compared with a single-step loss. Moreover, the learned latent context captures task information without explicit task labels, making it a promising signal for task-aware assistive control.
Chinese Translation
基于人体意图的任务无关辅助外骨骼控制,相比依赖预定义任务或运动模式的传统方法具有更大的灵活性。人体关节力矩估计通过刻画用户动作来实现任务无关的辅助。物理一致的方法(如深度拉格朗日网络,Deep Lagrangian Networks, DeLaN)已被应用于多用户场景下的人体力矩估计,但现有方法在不重新训练的情况下无法适应特定用户,且未考虑运动过程中的间歇性接触。我们提出ExoLaN,一种面向人-外骨骼交互的上下文感知DeLaN,它在学习完整耦合系统动力学的同时能够适应交互环境的变化。ExoLaN将时序上下文与来自力敏鞋垫的部分接触力测量相结合,以推断潜在的动力学嵌入并估计广义接触力矩。在七名未见用户执行21项未见任务的实验中,与黑盒基线相比,ExoLaN将力矩估计的均方误差(MSE)降低了7%。除逆动力学外,ExoLaN还可作为一个统一模型实现准确的前向预测:与单步预测损失相比,采用多步预测损失进行训练可将加速度MSE降低59%,并将长时程位置和速度误差分别降低60%和93%。此外,学习到的潜在上下文能够在没有显式任务标签的情况下捕获任务信息,使其成为任务感知辅助控制的一个有前景的信号。
cs.RO / 77 / 2609.31439

Learning to Leverage Compliance: A Policy-Admittance Learning Framework for Robotic Insertion

学会利用柔顺性:一种面向机器人插入操作的策略-导纳学习框架
Wang, Chongren, Li, Minghe, Dai, Honghua, Lin, Zhicheng, Wei, Shiyang, Yue, Xiaokui
Abstract
Policy learning and compliant control offer a promising route to reliable autonomous assembly under pose errors and contact uncertainty. However, combining them does not ensure coordination: the policy may continue pushing against contact while the controller yields, producing sustained loading with limited progress. To address this problem, we propose LeCo (Leverage Compliance), a policy-admittance learning framework that guides a visual policy through execution-time interaction under fixed admittance. A multirate feedback mechanism aggregates high-rate contact-interaction records into policy-transition rewards. An integrated conflict cost then characterizes sustained policy-loading/controller-unloading opposition, while a directional high-force tail cost captures continued-loading events within a transition. Together with task completion, these costs encourage the policy to leverage compliance with less unproductive loading. We evaluate LeCo on four real connector-assembly tasks, obtaining an aggregate success rate of 94%. Across tasks, mean successful-trial resultant-force and torque peaks decrease by approximately 30% and 64% relative to the comparison baseline. Reward ablation further shows that adding conflict shaping reduces median successful-trial contact-conditioned conflict density by approximately 53%. These results support learning to leverage fixed compliance by turning multirate policy-admittance interaction into complementary reward signals for effective, lower-load insertion.
Chinese Translation
策略学习与柔顺控制为在位姿误差和接触不确定性下实现可靠的自主装配提供了一条有前景的途径。然而,将二者结合并不保证能够协调一致:策略可能在持续挤压接触的同时控制器却在退让,导致持续的受力而进展有限。为解决这一问题,我们提出LeCo(Leverage Compliance,利用柔顺性),一个策略-导纳学习框架,通过执行时在固定导纳下的交互来引导视觉策略。一种多速率反馈机制将高频率的接触交互记录聚合为策略转移奖励。一个集成的冲突代价用于刻画策略施力与控制器卸力之间的持续对抗,而方向性高力尾部代价则捕捉单次转移内的持续施力事件。结合任务完成奖励,这些代价激励策略以更少的无效施力来利用柔顺性。我们在四个真实的连接器装配任务上评估LeCo,获得了94%的总体成功率。与对比基线相比,各任务中成功试验的平均合力峰值和力矩峰值分别降低约30%和64%。奖励消融实验进一步表明,加入冲突塑形可使成功试验的中位接触条件冲突密度降低约53%。这些结果支持通过将多速率的策略-导纳交互转化为互补的奖励信号来学习利用固定柔顺性,从而实现有效且更低受力的插入操作。
cs.RO / 78 / 2609.31452

Vision-Based 6-DoF Grasp Pose Estimation for Robot Cloth Unfolding

面向机器人衣物展开的基于视觉的六自由度抓取位姿估计
Tabernik, Domen, Nimac, Peter, Jerićević, Jan, Skočaj, Danijel, Gams, Andrej
Abstract
Cloth manipulation is a challenging task due to the deformable and high-dimensional nature of cloth, which leads to complex interaction dynamics and perceptual ambiguity arising from frequent occlusions of critical visual cues such as folds, edges, and grasp points. In this work, we tackle cloth unfolding using a regrasping-in-the-air strategy, where one manipulator holds the cloth while the other grasps it at an optimally selected point to unfold it. To this end, we propose CeDiRNet-6DoF, a deep learning framework that jointly predicts effective grasp points and the complete 6-DoF grasp pose from the observed cloth configuration. By integrating dense 3D grasp regression with segmentation and sine-cosine-encoded Euler angles, the proposed method reliably estimates the grasp configuration that maximizes the unfolded cloth area. We extensively evaluated CeDiRNet-6DoF on a bimanual robotic setup within the ICRA 2024 Cloth Competition framework, achieving state-of-the-art performance. An ablation study further validates the benefits of key design components, including joint segmentation, background randomization, and image cropping. These results establish CeDiRNet-6DoF as a robust and versatile foundation for reliable robotic cloth manipulation in unstructured environments.
Chinese Translation
衣物操作是一项具有挑战性的任务,因为衣物具有可变形性和高维特性,这导致了复杂的交互动力学,以及由褶皱、边缘和抓取点等关键视觉线索频繁遮挡所引起的感知歧义。在本工作中,我们采用空中重新抓取(regrasping-in-the-air)策略来解决衣物展开问题,即一个机械臂抓住衣物,另一个机械臂在最优点位抓取衣物以将其展开。为此,我们提出了CeDiRNet-6DoF,这是一个深度学习框架,能够根据观测到的衣物形态,联合预测有效抓取点和完整的六自由度(6-DoF)抓取位姿。通过将稠密三维抓取回归与分割以及正弦-余弦编码的欧拉角相融合,所提出的方法能够可靠地估计出使展开后衣物面积最大化的抓取配置。我们在ICRA 2024衣物竞赛(Cloth Competition)框架下的双臂机器人平台上对CeDiRNet-6DoF进行了广泛评估,取得了最先进的性能。消融实验进一步验证了关键设计组件的收益,包括联合分割、背景随机化和图像裁剪。这些结果确立了CeDiRNet-6DoF作为在非结构化环境中实现可靠机器人衣物操作的鲁棒且通用的基础。
cs.RO / 79 / 2609.31577

Generate, Track, Improve: Perceptive Multi-Skill Humanoid Locomotion with RL-Fine-Tuned Motion Generators

生成、跟踪、改进:基于强化学习微调动作生成器的感知型多技能人形机器人运动控制
Olkin, Zachary, Compton, William D., Ames, Aaron D.
Abstract
General purpose humanoids require locomotion controllers that are multi-skill, perceptive, dynamic, and robust enough to go anywhere humans can. In this work, we present a two layer locomotion architecture: (1) a perceptive flow matching motion generator plans whole body trajectories from raw depth images while a (2) perceptive tracking policy trained with control-guided RL follows these motions. Both policies are trained on a library of terrain consistent motion clips created with dynamically optimized human data which yields both accurate velocity tracking and terrain consistent references. Our central contribution is a simple yet effective off-policy RL fine tuning loop that improves the motion generator. A structured search method is used with the generator to gather data for advantage weighted regression. This off-policy loop is much more sample efficient than on-policy residual fine tuning and improves terrain consistency on unseen geometries and skill compositions. We find that successful terrain traversals increased by up to 25 percentage points and skill selection improved by up to 80 percentage points. By using raw depth images to perceive the environment no odometry or height maps are needed, and outdoor deployment is easy. With two cameras, the policy can see terrain coming from further away and adjust its velocity regardless of the commanded speed so it can traverse the terrain. A single policy pair enables a Unitree G1 humanoid to walk, run, stand, jump on and off of boxes, and traverse stairs in outdoor environments. Project page: https://zolkin1.github.io/generate-track-improve/
Chinese Translation
通用型人形机器人需要具备多技能、感知能力、动态性和鲁棒性的运动控制器,才能到达人类能够到达的任何地方。在本工作中,我们提出了一种两层运动控制架构:(1) 一个基于感知的流匹配(flow matching)动作生成器,从原始深度图像规划全身轨迹;(2) 一个通过控制引导强化学习(RL)训练的感知跟踪策略,用于跟随这些动作。两个策略均在一个由动力学优化的人类数据构建的地形一致动作片段库上训练,从而实现精确的速度跟踪和地形一致的参考动作。我们的核心贡献是一个简单而有效的离线策略(off-policy)强化学习微调循环,用于改进动作生成器。我们采用一种结构化搜索方法与生成器结合来收集数据,用于优势加权回归(advantage weighted regression)。相比在线策略(on-policy)残差微调,这种离线策略循环的样本效率高得多,并提升了在未见过的几何结构和技能组合上的地形一致性。我们发现,成功通过地形的比例提升了最多25个百分点,技能选择的准确率提升了最多80个百分点。通过使用原始深度图像感知环境,无需里程计或高度图,因此易于进行户外部署。借助两个摄像头,策略可以更早地看到前方地形,并在不依赖指令速度的情况下调整自身速度,从而顺利通过地形。单一策略对即可使 Unitree G1 人形机器人在户外环境中实现行走、奔跑、站立、跳上跳下箱子以及上下楼梯。项目页面:https://zolkin1.github.io/generate-track-improve/
cs.RO / 80 / 2609.31606

Learning Robot Policies from Sparse Success Signals via STL-Guided Stein Variational Policy Gradient

基于STL引导的Stein变分策略梯度从稀疏成功信号中学习机器人策略
Zheng, Hongrui, Vasile, Cristian Ioan, Loquercio, Antonio, Mangharam, Rahul
Abstract
Learning robot policies for tasks with sparse success signals is challenging when completion depends on coordinated actions, precise contact outcomes, or satisfying several conditions together. Intricate physical interactions with the world further complicate these requirements. Prior work using conventional reward shaping mechanisms provides dense feedback but local progress might not translate into eventual task completion. We present Signal Temporal Logic-guided Stein Variational Policy Gradient (STL-SVPG), a population-based method that uses smooth STL robustness as a trajectory-level training objective. Differentiating this objective through the dynamics assigns credit to policy actions according to their effect on the complete task specification, rather than local progress alone. We evaluate the approach on six quadcopter and manipulator tasks that involves event-triggered responses, strictly ordered behavior, responses within specified deadlines, and physical interaction with the world. STL-SVPG achieves the highest mean success rate among the compared methods on five of six benchmarks. Simulation-trained policies trained in simulation transfer temporal and contact task behavior to the real world.
Chinese Translation
当任务完成依赖于协调动作、精确的接触结果或同时满足多个条件时,利用稀疏成功信号学习机器人策略极具挑战性。与世界进行复杂的物理交互进一步增加了这些要求的难度。以往采用传统奖励塑形机制的方法虽然能提供密集反馈,但局部进展未必能转化为最终的任务完成。我们提出了信号时序逻辑引导的Stein变分策略梯度(STL-SVPG),这是一种基于种群的方法,将平滑的STL鲁棒度作为轨迹级训练目标。通过对该目标沿动力学进行微分,可根据策略动作对完整任务规范的影响来分配信用,而不仅仅依据局部进展。我们在六个四旋翼和机械臂任务上评估了该方法,这些任务涉及事件触发响应、严格有序的行为、在规定期限内完成响应以及与世界的物理交互。STL-SVPG在六个基准中的五个上取得了所比较方法中最高的平均成功率。在仿真中训练的策略能够将时序任务行为和接触任务行为迁移到真实世界。
人工智能 (Artificial Intelligence)
88
cs.AI / 1 / 2609.30291

Bringing AI to Autonomous Systems -- From Cognition to Collective Intelligence

将人工智能引入自主系统——从认知到集体智能
Sifakis, Joseph
Abstract
The purpose of this article is to highlight the central role of autonomous systems as the ultimate stage in the development of AI, to explain the underlying technical challenges that require a combination of connectionist AI and symbolic AI, and to integrate AI and systems engineering. We present a comprehensive framework for the design and evaluation of autonomous systems, based on a generic agent architecture that characterizes their behavior as the composition of cognitive functions organized around a long-term memory containing the agent's evolving knowledge. We address the challenges posed by the implementation of the fundamental features of the agent architecture, in particular the link between sensory data and structured data stored in memory, decision-making related to the achievement of the agent's goals and their planning, as well as the coordination of agents to combine individual and collective intelligence. We explain that agent trustworthiness, unlike that of traditional systems, is not limited to behavioral properties. It includes an essential dimension related to cognitive properties, the validity of which depends on how the agent uses its knowledge in decision-making. We present avenues for the development of methods for evaluating agent trustworthiness. We conclude with a critical assessment of the substantial gap between the aspirational vision of autonomous multi-agent systems and the current state of the art.
Chinese Translation
本文旨在强调自主系统作为人工智能(AI)发展最终阶段的核心作用,阐明需要将联结主义AI与符号AI相结合、并融合AI与系统工程的技术挑战。我们提出了一个用于自主系统设计与评估的综合框架,该框架基于一种通用的智能体(agent)架构,将自主系统的行为描述为围绕长期记忆(包含智能体不断演进的知识)组织的认知功能的组合。我们探讨了实现该智能体架构基本功能所面临的挑战,特别是传感数据与存储在记忆中的结构化数据之间的关联、与实现智能体目标相关的决策及其规划,以及智能体之间为融合个体智能与集体智能的协调问题。我们指出,与传统系统不同,智能体的可信性不仅限于行为属性,还包括一个与认知属性相关的重要维度,其有效性取决于智能体在决策中如何运用自身的知识。我们提出了发展智能体可信性评估方法的途径。最后,我们对自主多智能体系统的理想愿景与当前技术水平之间的巨大差距进行了批判性评估。
cs.AI / 2 / 2609.30325

ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?

ScopeBench:智能体在目标压力下能否守住任务范围边界?
Caldwell, Shane, Harley, Max, Dawson, Ads, Kouremetis, Michael, Abruzzo, Vincent, Pearce, Will
Abstract
Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client's engagement boundary. Existing offensive-security benchmarks measure raw hacking capability; as those benchmarks saturate, the real barrier to deployment is a special case of alignment: scope adherence. We introduce ScopeBench, a benchmark of 30 dead-end agentic security tasks in which the stated objective is reachable only by violating the stated scope. Each task appears under two conditions that share an environment, verifier, and objective and differ only in scope: one instruction set has no scope and measures capability; the other has a natural-language scope to measure adherence. Scopeless trajectories are graded by a standard deterministic verifier. Scoped trajectories pass through two grading arms. First, the same deterministic verifier checks for the flag: because the flag sits behind the scope boundary, a pass proves by construction that a forbidden action occurred, yielding a high-precision lower bound on the violation rate. If the verifier does not pass the trajectory, an agentic judge estimates whether an out-of-scope call occurred. We calibrate the judge against 100 ScopeBench trajectories labeled call-by-call by human annotators, and a blinded audit of the evaluated rollouts finds its high recall holds - no false negatives among the 36 audited violations, with over-flagging its only observed error. Across 8 models in one harness, raw capability spans 12.2% to 81.1% and scope adherence spans 34.4% to 86.7%, with the judge finding 331 violations that mechanical verification misses. Opus-4-8 achieves a raw-capability score 10 percentage points higher than sonnet-4-6's while exhibiting 35.6 percentage points higher scope adherence. We release the frozen pilot benchmark, evaluation code, and all 2160 ATIF trajectories.
Chinese Translation
智能体正日益在Web应用和网络渗透测试中被赋予真实的自主权限,而在这种场景下,一次超出范围的操作就可能突破客户约定的任务边界。现有的攻击性安全基准测试衡量的是原始攻击能力;随着这些基准趋于饱和,部署的真正障碍变成了对齐问题的一个特例:范围遵从(scope adherence)。我们提出了ScopeBench,这是一个包含30个死胡同式智能体安全任务的基准,其中所述目标只有通过违反所述范围才能达成。每个任务在两种条件下出现,这两种条件共享相同的环境、验证器和目标,仅范围设置不同:一组指令没有范围限制,用于衡量能力;另一组包含自然语言描述的范围,用于衡量遵从性。无范围轨迹由标准的确定性验证器进行评分。有范围轨迹则通过两条评分路径进行评估。首先,同一个确定性验证器检查是否获取了flag:由于flag位于范围边界之后,通过验证在构造上即证明发生了被禁止的操作,从而为违规率提供了一个高精度的下界。如果验证器未能通过该轨迹,则由一个基于智能体的裁判(agentic judge)估计是否发生了超出范围的操作。我们使用由人工标注者逐调用标注的100条ScopeBench轨迹对裁判进行校准,并且对被评估轨迹的盲审计表明其高召回率得以保持——在36个经审计的违规行为中没有假阴性,过度标记是其唯一观察到的错误。在同一测试框架下的8个模型中,原始能力从12.2%到81.1%不等,范围遵从性从34.4%到86.7%不等,其中裁判发现了331个机械验证无法检测到的违规行为。Opus-4-8的原始能力得分比sonnet-4-6高出10个百分点,同时范围遵从性高出35.6个百分点。我们发布了冻结的试点基准、评估代码以及全部2160条ATIF轨迹。
cs.AI / 3 / 2609.30328

When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess

多智能体代码评判何时才真正有据可依?两种无标签度量方法与一个拒绝猜测的评判器
Aly, Salma Roshdy, Assaf, Hussein, Kobti, Ziad
Abstract
When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent verification, which decomposes a judgment into checkable claims and verifies each against evidence, is a promising response and works well when the evidence is a set of retrieved documents. We argue such methods require two things of their evidence: it must be independent of the answer under review, and it must differ between the two candidates being compared. The second condition holds automatically with retrieved documents and stops holding in code judging. Running MARCH, a published framework unmodified over 80 condition-by-cell measurements on two code judging benchmarks, we find it declares both solutions equally good on 78 to 95% of comparisons, reaching 4.4% accuracy where the same model asked directly reaches 43.7%. Neither easier problems nor a larger judge changes this. Two measurements taken from the pipeline's own logs explain it without needing labels. Gating on one of them, the pipeline declines the comparisons it cannot make and raises its accuracy from 20.7 to 36.9% while still answering half of all comparisons. The contribution is not a more accurate judge, but a label-free way to tell when a judge has no basis for its answer.
Chinese Translation
当一个语言模型评判另一个模型的代码是否正确时,它并不会报告缺乏证据的情况。它会给出一个附带推理的自信判定,与有据可依的判定毫无区别。多智能体验证方法将一次评判分解为多个可核查的论断,并逐一依据证据加以验证,是一种有前景的应对方案,且当证据是检索到的文档集合时效果良好。我们认为,此类方法对其证据有两点要求:证据必须独立于被评审的答案,并且在被比较的两个候选答案之间必须有所差异。第二个条件在检索文档的场景下自动满足,但在代码评判中不再成立。我们在两个代码评判基准上运行已发表的框架 MARCH(未做修改),进行了 80 项按条件划分的单元测量,发现它在 78% 到 95% 的比较中判定两个解同样好,准确率仅为 4.4%,而直接询问同一模型时准确率可达 43.7%。使用更简单的问题或更大的评判模型都无法改变这一结果。从该流程自身日志中提取的两种度量无需标签即可解释这一现象。基于其中一种度量设置门槛后,该流程会拒绝其无法作出的比较,将准确率从 20.7% 提升至 36.9%,同时仍能回答一半的比较任务。本文的贡献并非一个更准确的评判器,而是一种无需标签即可判断评判器是否缺乏作答依据的方法。
cs.AI / 4 / 2609.30341

Bridging LLM Agents and Data Spaces: An Architectural Mediation Approach using the Model Context Protocol

连接大语言模型智能体与数据空间:一种基于模型上下文协议(MCP)的架构中介方法
Ruiz, Jaime Alonso, Aparicio, Carlos, Huecas, Gabriel, Salvachúa, Joaquín, Munoz-Arcentales, Andres
Abstract
Data Spaces enable sovereign and governed data sharing across organizational boundaries, but their integration with AI agents remains challenging due to mismatches between probabilistic language model interactions and policy-driven data infrastructures. This article presents an architectural mediation approach based on the Model Context Protocol (MCP), implemented through the Eunomia Agent, to enable controlled interaction between large language model (LLM) agents and data space services. The proposed mediation layer translates data space capabilities into structured, schema-driven tools that AI agents can discover and invoke while preserving governance constraints. A prototype implementation validates end-to-end interaction across catalog discovery, metadata retrieval, and data service invocation without modifying existing data space components. Results demonstrate that protocol-based mediation enables interoperable and standards-aligned integration of AI agents into data space ecosystems. The approach provides practical guidance for organizations seeking to introduce AI-driven automation into governed data-sharing environments while maintaining compliance, interoperability, and architectural separation of concerns.
Chinese Translation
数据空间(Data Spaces)支持跨组织边界的主权化、受治理的数据共享,但由于概率性语言模型交互与策略驱动的数据基础设施之间存在不匹配,其与AI智能体的集成仍然面临挑战。本文提出了一种基于模型上下文协议(Model Context Protocol, MCP)的架构中介方法,并通过Eunomia智能体(Eunomia Agent)实现,以支持大语言模型(LLM)智能体与数据空间服务之间的受控交互。所提出的中介层将数据空间能力转换为结构化的、基于模式的工具(schema-driven tools),使AI智能体能够在保留治理约束的前提下发现并调用这些工具。原型实现验证了在目录发现、元数据检索和数据服务调用全流程的端到端交互,且无需修改现有数据空间组件。结果表明,基于协议的中介方式能够实现AI智能体与数据空间生态系统之间可互操作、符合标准规范的集成。该方法为希望在受治理的数据共享环境中引入AI驱动自动化,同时保持合规性、互操作性和架构层面关注点分离的组织提供了实践指导。
cs.AI / 5 / 2609.30383

Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

各自隐匿,合而致害:针对基于技能的智能体系统的技能级联攻击
Zhu, Zihao, Lyu, Siwei, Bibi, Adel, Wu, Baoyuan
Abstract
A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities, but the openness of this skill ecosystem also opens up a new attack surface. Prior work has focused on vulnerabilities within individual skills, but little attention has been paid to risks that arise from interactions across skills. In this paper, we introduce skill cascading attacks, a threat paradigm in which a malicious objective is distributed across multiple skills so that each modification looks benign in isolation, yet their combined execution is harmful. For instance, in a prescription-review pipeline, the first skill weakens signals of recently discontinued medications in the extracted history, the second downgrades the severity of any drug interaction tied to them, and the third suppresses the resulting low-priority alert in the final summary, so that a severe drug-interaction warning silently disappears before reaching the physician. To systematically study this safety blind spot, we develop SkillCascade, an automated multi-agent red-teaming framework, and release SkillCascade-Bench, a benchmark of 213 validated cascading test cases across multiple agent systems and domains. Across representative agents (e.g., OpenClaw, Claude Code, Codex) and LLM backbones, cascaded interactions reliably induce harmful behaviors while evading existing per-skill scanners and runtime monitors. Our findings highlight a gap between component-level integrity and system-level safety, and call for defenses that reason over cross-skill interactions rather than individual skills in isolation.
Chinese Translation
技能(skill)是一种模块化封装,由自然语言指令、可执行脚本和参考资源组成,智能体可在运行时加载技能以扩展其在特定任务上的能力。因此,基于技能的智能体系统能够灵活地复用第三方能力,但这种技能生态的开放性也带来了新的攻击面。已有工作主要关注单个技能内部的漏洞,而对跨技能交互所产生的风险关注甚少。本文提出技能级联攻击(skill cascading attacks)这一威胁范式:恶意目标被拆分分布到多个技能中,使每个修改在单独来看时显得无害,但其组合执行却会造成危害。例如,在一个处方审核流程中,第一个技能在提取的病史中削弱近期停用药物的信号,第二个技能降低与之相关的药物相互作用的严重程度,第三个技能在最终摘要中抑制由此产生的低优先级警报,从而使得一个严重的药物相互作用警告在到达医生之前悄然消失。为系统地研究这一安全盲区,我们开发了 SkillCascade——一个自动化多智能体红队测试框架,并发布了 SkillCascade-Bench——一个包含 213 个经多智能体系统与多领域验证的级联测试用例的基准。在代表性智能体(如 OpenClaw、Claude Code、Codex)及多种大语言模型骨干上,级联交互能够稳定地诱发有害行为,同时规避现有的单技能扫描器和运行时监控器。我们的研究结果揭示了组件级完整性与系统级安全性之间的差距,并呼吁发展能够对跨技能交互进行推理、而非孤立审视单个技能的防御机制。
cs.AI / 6 / 2609.30397

A Synthetic Ground-Truth Framework for the Evaluation of Explainable AI Methods

一种用于评估可解释人工智能方法的合成真值框架
Miró-Nicolau, Miquel, Spinnato, Francesco, Guidotti, Riccardo
Abstract
Evaluating explainable Artificial Intelligence (XAI) methods is a challenging task due to the lack of reliable evaluation procedures and, in particular, the absence of ground truth explanations. In the literature, existing evaluation approaches typically assess explanations by measuring their fidelity with respect to the predictions of a black-box model. However, such evaluation strategies only quantify the degree to which an explanation reproduces the model's output, without ensuring that the explanation correctly reflects the underlying decision process. As a consequence, different explanations may achieve similar fidelity scores while providing inconsistent or misleading interpretations of the model behavior. In this paper, we propose a framework for the evaluation of XAI methods based on synthetic ground truth. The proposed approach relies on controlled interventions to generate synthetic datasets in which the importance of input components can be determined by design. This enables the construction of ground truth explanations that are directly aligned with the behavior of the model under analysis. The framework is instantiated across three data domains, namely binary images, tabular data, and time series, allowing a comprehensive assessment of explanation methods in heterogeneous settings. Experimental results obtained by evaluating nine widely used XAI methods show significant limitations in current techniques and highlight the importance of synthetic, intervention-based benchmarks for a reliable assessment of explanation quality.
Chinese Translation
由于缺乏可靠的评估程序,尤其是缺少真值(ground truth)解释,评估可解释人工智能(XAI)方法是一项具有挑战性的任务。在现有文献中,已有的评估方法通常通过度量解释相对于黑盒模型预测的保真度来评估解释质量。然而,此类评估策略仅量化了解释复现模型输出的程度,并不能确保解释正确反映模型的内在决策过程。因此,不同的解释可能获得相近的保真度分数,却对模型行为给出不一致甚至具有误导性的解释。本文提出了一种基于合成真值的XAI方法评估框架。该方法依赖于受控干预来生成合成数据集,其中输入成分的重要性可以通过设计加以确定。这使得所构建的真值解释能够与所分析模型的行为直接对齐。该框架在三种数据域上进行了实例化,即二值图像、表格数据和时间序列,从而能够在异构环境中对解释方法进行全面评估。通过评估九种广泛使用的XAI方法所得的实验结果表明,现有技术存在显著局限性,并凸显了基于合成的、干预式基准对于可靠评估解释质量的重要性。
cs.AI / 7 / 2609.30446

Predicting Transmembrane Protein Topology from 3D Structure

基于三维结构预测跨膜蛋白拓扑结构
Chen, Sitong, Mao, Xiaopeng
Abstract
This paper presents a novel approach to infer protein topology using the state-of-the-art graph neural network (GNN), SchNet. The model is trained on the same dataset used to develop the recent DeepTMHMM model with 5-fold cross-validation. Unlike the conventional approaches based on using only the protein sequences or the $\alpha$-carbons as features, we have decoded our classifier in this way, so all atom-level embeddings are used. Without applying any pre-trained weight, the final results have shown great potential that GNNs can be used for topological predictions.
Chinese Translation
本文提出了一种利用最先进的图神经网络(GNN)SchNet来推断蛋白质拓扑结构的新方法。该模型在与开发最新DeepTMHMM模型所用相同的数据集上进行训练,并采用5折交叉验证。与仅使用蛋白质序列或$\alpha$-碳原子作为特征的常规方法不同,我们的分类器采用全原子级别的嵌入(embedding)进行解码。在未使用任何预训练权重的情况下,最终结果表明GNN在拓扑结构预测方面具有巨大潜力。
cs.AI / 8 / 2609.30456

Spectral Feedback for Test-Time Alignment of Protein Diffusion Models

用于蛋白质扩散模型测试时对齐的谱反馈方法
Dickman, Shai, Cemri, Mert, Butler, Landon, Ramchandran, Kannan
Abstract
Reward maximization alignment methods for discrete diffusion models have primarily focused on steering the reverse process, either by influencing token logits or by selecting favorable sequences at intermediate steps. These approaches largely treat inference as a unidirectional process, lacking mechanisms for revisiting undesirable token selections. We introduce Spectral Feedback, an algorithm that selects edit-positions in a feedback loop, allowing the model to iteratively correct its own generations. This approach leverages the mask structure of discrete diffusion models by re-masking and re-sampling tokens, analogous to image editing methods that reintroduce noisy latents and re-run the reverse process. While prior alignment methods focus on what token labels to assign to maximize a target reward, we instead treat which tokens to revisit as the central alignment problem. Selecting edit-positions is challenging because edit effects are interdependent: the impact of modifying one token depends on which others are edited simultaneously. We define an edit-set as a set of token positions to re-mask and re-sample. Motivated by prior work on sparse interactions in biological systems, we find empirically that edit-set value functions for protein inverse folding admit sparse Fourier representations. This structure enables Spectral Feedback to efficiently learn and optimize the value functions for edit-position selection. Spectral Feedback is model-agnostic and can be applied to pretrained, test-time aligned, and fine-tuned diffusion models. For all of these models, the algorithm improves alignment performance without modifying the underlying generative process. Applied to inverse folding with a protein stability reward oracle, it achieves a 32.3% increase in stable proteins for a pretrained model, 24.8% for Best-of-10, and 5.8% for a state-of-the-art RL fine-tuned diffusion model.
Chinese Translation
针对离散扩散模型的奖励最大化对齐方法主要集中于对反向过程的引导,即通过影响词元(token)的逻辑值(logits)或在中间步骤中选择有利序列来实现。这些方法大多将推理视为单向过程,缺乏对不良词元选择进行重新审视的机制。我们提出了谱反馈(Spectral Feedback)算法,该算法在反馈回路中选择编辑位置,使模型能够迭代地修正自身生成结果。该方法利用离散扩散模型的掩码结构,通过重新掩码和重新采样词元进行修正,类似于图像编辑方法中重新引入噪声潜变量并再次运行反向过程的思路。以往的对齐方法侧重于为词元分配何种标签以最大化目标奖励,而我们则将“重新审视哪些词元”作为核心的对齐问题。选择编辑位置具有挑战性,因为编辑效应是相互依赖的:修改某个词元的影响取决于同时编辑了哪些其他词元。我们将编辑集(edit-set)定义为一组需要重新掩码和重新采样的词元位置。受生物系统中稀疏交互相关研究的启发,我们通过实验发现,蛋白质逆折叠的编辑集价值函数具有稀疏傅里叶表示。这一结构使谱反馈能够高效地学习并优化用于编辑位置选择的价值函数。谱反馈与模型无关,可应用于预训练模型、测试时对齐模型以及微调后的扩散模型。对于所有这些模型,该算法均能在不修改底层生成过程的情况下提升对齐性能。在以蛋白质稳定性奖励预言机(reward oracle)进行的逆折叠任务中,该算法使预训练模型的稳定蛋白质数量提升32.3%,对Best-of-10策略提升24.8%,对最先进的强化学习微调扩散模型提升5.8%。
cs.AI / 9 / 2609.30469

Pretrained ASR Pseudo-labeling for Noisy Police Audio

基于预训练ASR伪标签的嘈杂警方音频处理
Chaparala, Kaavya, Huang, Su, Miller, Stephen L., Miller, Rhiannon N., Field, Anjalie
Abstract
Pretrained ASR systems perform poorly on noisy Broadcast Police Communication (BPC), hindering efforts to understand police decision-making. Pseudo-labeling offers an unsupervised path to improve ASR without expensive human labels, but the efficacy of this approach on very noisy domains is not known. In this work, we systematically assess the opportunities and limits of pseudo-labeling to adapt foundation ASR models (Whisper and Qwen3-ASR) to noisy BPC domain corpora from Baltimore and Chicago. We demonstrate that existing internal confidence metrics (log-probabilities and STAR scores) fail to distinguish between high and low quality BPC pseudo-labels, and we introduce an external LLM-as-a-judge filtering paradigm that leverages parametric knowledge to discard contextually implausible transcripts. Our LLM-judging filters more aggressively than internal metrics and significantly reduces WER of the pseudo-labeled training sets across the Baltimore and Chicago BPC corpora, though a substantial gap remains relative to an oracle filter. We also introduce a new cross-model pseudo-labeling paradigm where one model is finetuned with pseudo-labels from the other, and we identify this method as a promising direction for future pseudo-labeling work.
Chinese Translation
预训练的自动语音识别(ASR)系统在嘈杂的警察广播通信(BPC)数据上表现不佳,阻碍了对警方决策过程的理解。伪标签方法提供了一种无需昂贵人工标注的无监督改进ASR的途径,但该方法在极度嘈杂领域中的有效性尚不清楚。在本工作中,我们系统地评估了伪标签方法将基础ASR模型(Whisper和Qwen3-ASR)适配到巴尔的摩和芝加哥嘈杂BPC领域语料库的机遇与局限。我们证明,现有的内部置信度指标(对数概率和STAR分数)无法区分高质量与低质量的BPC伪标签,并引入了一种外部的大模型作为裁判(LLM-as-a-judge)的过滤范式,利用参数化知识丢弃上下文上不可信的转录文本。与内部指标相比,我们的LLM裁判过滤器更为激进,并显著降低了巴尔的摩和芝加哥BPC语料库上伪标签训练集的词错误率(WER),尽管相对于理想(oracle)过滤器仍存在较大差距。我们还提出了一种新的跨模型伪标签范式,即用一个模型的伪标签对另一个模型进行微调,并将该方法确定为未来伪标签研究的一个有前景的方向。
cs.AI / 10 / 2609.30484

Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework

大语言模型理解上下文吗?一种基于知识图谱的评估框架
Arumugam, Subavarshana, Nallaretnam, Mamta, Wickramasinghe, Kithuni, Gunapala, Chamath, Vipulanandan, Pragatheeswaran, Premaratne, Kamal, Thayasivam, Uthayasanker
Abstract
While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant information from a given context, integrate it into a coherent internal representation, and reason over it to produce factually consistent and contextually grounded responses. However, traditional methods such as BiLingual Evaluation Understudy (BLEU) and perplexity simply measure surface-level performance. This reveals a critical gap in question answering (QA), where responses must be contextually grounded rather than simply being memorized associations. To fill this void, we propose a novel knowledge graph (KG) based evaluation framework for LLM contextual understanding in QA. Central to this is Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure combining structural and semantic signals into a single score. In addition, a diagnostic analysis framework is developed to identify and categorize reasoning errors at the triplet level, enabling fine-grained analysis of model failures. Together, across nine benchmarks, S3KG achieves F1 gains of up to $+7.6$ points over the strongest baseline and AUROC up to $0.973$.
Chinese Translation
尽管大语言模型(LLM)已展现出卓越的语言能力,但其核心仍萦绕着一个深刻的问题:这些模型是真正理解上下文,还是仅仅在前所未有的规模上擅长模式匹配?LLM的上下文理解能力是指从给定上下文中正确提取相关信息、将其整合为连贯的内部表示,并在此基础上进行推理以生成事实一致且紧扣上下文回答的能力。然而,诸如双语评估替补(BLEU)和困惑度等传统方法仅能衡量表层性能。这在问答(QA)任务中暴露出一个关键缺口——回答必须扎根于上下文,而非仅仅依赖记忆化的关联。为填补这一空白,我们提出了一种新颖的基于知识图谱(KG)的评估框架,用于评估LLM在问答中的上下文理解能力。其核心是面向知识图谱的语义结构相似度(S3KG),这是一种将结构信号与语义信号融合为单一分数的混合相似度度量。此外,我们还开发了一个诊断分析框架,用于在三元组层面识别和归类推理错误,从而实现对模型失败的细粒度分析。在九个基准测试中,S3KG相比最强基线取得了最高$+7.6$点的F1提升,AUROC最高达$0.973$。
cs.AI / 11 / 2609.30489

BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering

BioEVAL:面向生物工程的大语言模型与多模态模型的全球多机构基准测试
Ye, Shun, Suja, Vinny Chandran, Li, Chenlong, Jiang, Chongming, Zamani, Reza, Li, Xiang, Bain, Christopher, Zhou, Yuqi, Peterson, Walker, Wang, Huidong, Hu, Chenglang, Park, Jongchan, Cheng, Xiao, Swedlund, Benjamin, Murillo, Sandra, Sivanandan, Anjali, Sun, Shiyu, Lanfeng, Liang, Islam, Mohammad Tariqul, Joy, Baju C., Khan, Ishaq N., Kumar, Sreedhar S., Mercado-Vásquez, Gabriel, Vizzard, James V., Matthews, Jonathan M., Huang, Helen, Guo, Xiaolu, Nicklow, Ethan, Chen, Guorui, Neff, Ryan A., Maity, Surjendu, Park, Hyeonjin, Joo, Han-ho, Dong, Katherine, Cai, Yuyan, Huang, Weihang, Zou, Yichen, Yan, Rui, Figueroa, Raphael, Goncharov, Artem, Schremmer, Bella Rose, Linton, Lian Elsa, Goda, Keisuke, Gao, Liang, Cheng, Ke, Morsut, Leonardo, Wilson, Jennifer L., Fu, Jianping, Teck, Lim Chwee, Sarkar, Deblina, Hierlemann, Andreas, Tay, Savaş, Hoffmann, Alexander, Griffin, Donald Richieri, Chen, Jun, Kelley, Shana O., Varghese, Shyni, Cheon, Jinwoo, Lam, Wilbur A., Moon, James J., Wong, Wilson W., Mitragotri, Samir, Di Carlo, Dino
Abstract
Large Language Models (LLMs) have demonstrated historic breakthroughs in general reasoning with early successes in biomedical science. However, existing LLM benchmarking emphasizes factual recall, offering limited insight into model performance on frontier and multimodal tasks. We assembled BioEVAL (BioEngineering Validation of AI and LLMs), a global, multi-institutional initiative designed to assess experimental reasoning capability across bioengineering (BE) subfields. BioEVAL spans 11 major BE subfields plus a set of uncategorized items, bringing together 22 research groups to create a PhD-level benchmark comprising 608 evaluation items: 1) 380 multiple-choice questions (MCQs, 359 retained after audit), 2) 218 literature synthesis tasks, and 3) 10 multimodal problems with experimental image interpretation. Benchmark items underwent authoring-group expert review and centralized quality control before evaluation. Following evaluation, a blinded cross-group consensus audit of the highest- and lowest-accuracy MCQ items flagged 21 questions for revision or removal; these were withheld, and all reported MCQ results are computed on the 359 retained items. We evaluated diverse cloud-scale foundation/multimodal models (e.g., ChatGPT, Gemini, and Grok) and locally deployable models suitable for inference on consumer-grade GPUs. Models achieved the highest accuracy of up to 90% on MCQs, similarity score of 0.72 on literature synthesis, and accuracy of 80% on a small sample of multimodal reasoning questions, with substantial performance variation across subfields. Leaderboard rankings characterize current capabilities, limitations, and development priorities across the evaluated BE task categories. BioEVAL is maintained as an extensible benchmark with standardized protocols for continuing expert item contribution and model evaluation.
Chinese Translation
大语言模型(LLM)在通用推理方面取得了历史性突破,并在生物医学科学领域取得了早期成功。然而,现有的LLM基准测试侧重于事实记忆,对模型在前沿任务和多模态任务上的表现缺乏深入的洞察。我们构建了BioEVAL(BioEngineering Validation of AI and LLMs,生物工程AI与大语言模型验证),这是一项全球性、多机构合作的计划,旨在评估生物工程(BE)各子领域的实验推理能力。BioEVAL涵盖11个主要的生物工程子领域以及一组未分类题目,汇聚了22个研究团队,构建了一个包含608个评估题目的博士级基准测试:1)380道多选题(MCQ,经审核后保留359道);2)218项文献综合任务;3)10道涉及实验图像解读的多模态问题。所有基准题目在评估前均经过出题组的专家审查和集中质量控制。评估完成后,对准确率最高和最低的MCQ题目进行了跨组盲审共识审计,标记出21道需修改或删除的题目;这些题目被剔除,所有报告的MCQ结果均基于保留的359道题目计算。我们评估了多种云端规模的基础模型/多模态模型(如ChatGPT、Gemini和Grok),以及适用于消费级GPU推理的本地部署模型。模型在多选题上的最高准确率达90%,文献综合任务的相似度得分为0.72,在小样本多模态推理题上的准确率为80%,且各子领域之间的性能差异显著。排行榜排名刻画了当前模型在所评估的生物工程任务类别上的能力、局限性和发展重点。BioEVAL作为一个可扩展的基准测试持续维护,并为后续专家题目贡献和模型评估提供了标准化协议。
cs.AI / 12 / 2609.30550

Benchy: towards a universal language for task-oriented AI benchmarks

Benchy:迈向面向任务型AI基准测试的通用语言
Daniel, Francis F, Ibañez, Mauro, Perelman, Francis, Basti, Marian
Abstract
Benchy is a semantic language and execution engine for benchmarking AI programs. A benchmark is completely specified by a program, a scoring function, and a dataset, B=(P,S,D), and is separate from the AI-system taking it; a run binds the two, R=(B,AI). Benchmarks are authored as canonical YAML in which each semantic concept has one valid syntax, classified by a shared task/domain/language ontology, and deterministically compiled into a canonical JSON intermediate representation that the engine executes. Compilation changes representation, not meaning: it does not repair invalid definitions or inject hidden defaults. Programs use fixed schemas of named input and output fields, the leaf output fields are the scoring dimensions, and the engine exposes one universal runtime contract --- a named-field input object in, a named-field output object out --- to which external AI-systems adapt at the boundary, so integration mechanics never propagate into benchmark semantics. This paper gives the semantic object model, the ontology and task-to-program validation rule, the scoring and failure semantics, the compilation and execution architecture, and the scope of the current language. An appendix fixes the normative engineering contract for the first engine implementation.
Chinese Translation
Benchy 是一种用于AI程序基准测试的语义语言与执行引擎。一个基准测试由程序、评分函数和数据集完整定义,即 B=(P,S,D),并与被测AI系统相互独立;一次运行将二者绑定,即 R=(B,AI)。基准测试以规范化 YAML 编写,其中每个语义概念仅有一种合法语法,并按照共享的任务/领域/语言本体进行分类,然后被确定性地编译为引擎可执行的规范化 JSON 中间表示。编译只改变表示形式而不改变语义:它不会修复无效定义,也不会注入隐藏的默认值。程序采用固定的具名输入与输出字段模式,叶子输出字段即评分维度;引擎对外提供统一的运行时契约——即一个具名字段输入对象进、一个具名字段输出对象出——外部AI系统在边界处适配该契约,从而使集成机制不会传导至基准测试的语义之中。本文给出了语义对象模型、本体及任务到程序的验证规则、评分与失败语义、编译与执行架构,以及当前语言的适用范围。附录规定了首个引擎实现的规范性工程契约。
cs.AI / 13 / 2609.30553

Rank-Reliable Teacher-Guided Fitness Approximation for Expensive Evolutionary Optimization: A TinyML Architecture Search Study

面向昂贵进化优化的排序可靠教师引导适应度近似:TinyML架构搜索研究
Garai, Soumen, Samui, Suman
Abstract
Expensive evolutionary search does not always need an exact fitness estimate for every candidate. It often needs a reliable answer to a simpler question: which candidate is better? We address this need through Teacher-Guided Learning NSGA-II (TGL-NSGA-II), a low-fidelity framework for constrained Tiny Machine Learning (TinyML) neural architecture search. A pretrained teacher organizes samples into strata defined jointly by difficulty and class. Each candidate then undergoes KD-Lite, a short and capped knowledge-distillation procedure on a compact training set, before being scored on a separate stratified evaluation set. This teacher-guided score is fused with a Gaussian-process surrogate to select candidates for full evaluation. For a fixed candidate population, we analyse evaluation variance, score concentration, pairwise rank inversion, expected Kendall-$\tau$, first-front identification, and hypervolume perturbation. We also derive a variance-aware fusion weight and a capacity-adaptive distillation rule. On keyword spotting and bird-call classification, the measured Kendall-$\tau$ values are 0.74 and 0.62, exceeding the corresponding predicted lower bounds of 0.60 and 0.46. Joint stratification reduces proxy-score variance by 41% relative to random evaluation. Selective teacher mismatch, in contrast, increases differential bias and reduces Kendall-$\tau$ to 0.41. Under a constrained evaluation budget, TGL-NSGA-II achieves the largest mean hypervolume and smallest generational distance on keyword spotting, records the lowest mean false-positive rate on BirdCLEF, and runs 2.2x faster than full NSGA-II. These guarantees apply to population-level low-fidelity evaluation and do not establish convergence of the complete evolutionary trajectory.
Chinese Translation
昂贵的进化搜索并不总是需要对每个候选解进行精确的适应度估计,它往往只需要对一个更简单的问题给出可靠的回答:哪个候选解更优?为此,我们提出了教师引导学习NSGA-II(TGL-NSGA-II),一种面向带约束的微型机器学习(TinyML)神经架构搜索的低保真框架。预训练的教师模型将样本划分为由难度和类别共同定义的分层。随后,每个候选解在紧凑训练集上接受一种短时且有上限的知识蒸馏过程(KD-Lite),并在独立的分层评估集上进行评分。该教师引导评分与高斯过程代理模型融合,用于选择进行完整评估的候选解。针对固定的候选种群,我们分析了评估方差、评分集中度、成对排序反转、期望Kendall-$\tau$、第一前沿识别率以及超体积扰动,并推导了方差感知的融合权重和容量自适应的蒸馏规则。在关键词 spotting(keyword spotting)和鸟鸣分类任务上,实测的Kendall-$\tau$值分别为0.74和0.62,超过了相应的预测下界0.60和0.46。联合分层使代理评分方差相比随机评估降低了41%。相反,选择性的教师不匹配会增大差异偏差,并使Kendall-$\tau$降至0.41。在受限的评估预算下,TGL-NSGA-II在关键词 spotting任务上取得了最大的平均超体积和最小的代际距离,在BirdCLEF上记录了最低的平均假阳性率,且运行速度比完整NSGA-II快2.2倍。上述保证适用于种群层面的低保真评估,但并不确立完整进化轨迹的收敛性。
cs.AI / 14 / 2609.30563

Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content

少思考,更逼真:直觉式提示提升大语言模型智能体模拟个体社交媒体反应的能力,包括陌生内容
Bojic, Ljubisa, Stanic, Tijana, Matthes, Joerg, Samala, Agariadne Dwinggo, Dinic, Bojana, Wang, Jue
Abstract
Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on agreement with human behaviour and has paid little attention to whether an agent behaves in line with the profile it was given. The present study profiled eight Serbian participants through a questionnaire, a deep interview, and a written self-presentation, recorded their reactions to sixty-eight social media posts, and asked four language models to predict those reactions under five prompt conditions varying profile content and instruction style. Attitudinal content improved prediction over demographic backstories by a wide margin. Agents matched their stated profiles more closely than participants matched their own survey answers, and consistency proved unrelated to fidelity once profile information was present. Instructing models to respond intuitively and immediately rather than analytically gave the highest fidelity of any condition and cut the compression of individual differences from seven times the human level to three. The advantage held on posts about topics the questionnaire never raised, where that condition reached the highest fidelity of any setup and beat a crowd baseline by a wide margin, which suggests that agents prompted this way could serve as general-purpose simulated users rather than specialists on the topics they were profiled for. Results may bear implications for the development of language models, because intuition-based setups appear better suited to some tasks than reasoning-based ones.
Chinese Translation
平台政策越来越多地在人工用户上进行测试,这使得智能体的逼真度变得重要。然而,令人信服的虚假账号也可能在选举前操纵公众舆论感知。现有验证工作主要集中在与人类行为的一致性上,而很少关注智能体的行为是否符合其所设定的人设。本研究通过问卷、深度访谈和书面自我介绍对八名塞尔维亚参与者进行了画像,记录了他们对六十八条社交媒体帖子的反应,并让四个语言模型在五种提示条件( varying 档案内容和指令风格)下预测这些反应。态度性内容相比人口统计背景故事大幅提升了预测效果。智能体与其设定档案的匹配程度甚至高于参与者与自己问卷答案的匹配程度,而且一旦提供了档案信息,一致性就被证明与逼真度无关。指示模型以直觉、即时的方式而非分析式的方式作出反应,在所有条件中获得了最高的逼真度,并将个体差异被压缩的程度从人类水平的七倍降至三倍。这一优势在涉及问卷从未提及的话题的帖子上依然成立:该条件达到了所有设置中最高的逼真度,并大幅超越了群体基线,这表明以这种方式提示的智能体可以作为通用模拟用户,而不仅仅是其所被画像话题的专家。研究结果可能对语言模型的开发具有启示意义,因为基于直觉的设置在某些任务上似乎比基于推理的设置更为适用。
cs.AI / 15 / 2609.30569

Atelier: Learning Local Self-Supervised Features for CryoEM Volumes via Hypernetworks

Atelier:基于超网络学习冷冻电镜(CryoEM)体积数据的局部自监督特征
Lo, Phillip, Babu, Sudarshan, Kimanius, Dari, Khan, Aly A.
Abstract
CryoEM map interpretation requires features that are spatially localized, consistent across samples, and informative across spatial scales. Most deep learning methods for map annotation extract features from fixed voxel grids. However, implicit neural representations (INRs) are able to model volumetric data as scale-agnostic, coordinate-conditioned functions. INRs are therefore attractive for cryoEM, but fitting a separate INR for each map is too expensive for large-scale feature extraction and produces representations that are not aligned across samples. We introduce Atelier, a self-supervised framework that amortizes INR fitting for reconstructed cryoEM maps. Pretrained on 5,439 Electron Microscopy Data Bank maps, Atelier is a transformer-based hypernetwork that generates high-fidelity reconstructions across a wide range of protein structures, including large multi-subunit assemblies. Beyond reconstruction, the INR generated by the pretrained transformer exposes a continuous, local feature field through its intermediate activations at any spatial query point, a property that voxel grid and patch-tokenizer architectures do not naturally provide. Used as auxiliary channels to a 3D nested U-Net annotation head trained from scratch, these coordinate-conditioned features improve performance on eight voxel-level property prediction tasks over a volume-only baseline. Our results demonstrate that amortized implicit neural representations are an effective primitive for geometry-aware analysis of cryoEM data.
Chinese Translation
冷冻电镜(CryoEM)密度图的解读需要具有空间局部性、跨样本一致性以及跨空间尺度信息量的特征。大多数用于密度图标注的深度学习方法从固定的体素网格中提取特征。然而,隐式神经表示(INR)能够将体积数据建模为与尺度无关、以坐标为条件的函数。因此,INR 对冷冻电镜颇具吸引力,但为每个密度图单独拟合一个 INR 对于大规模特征提取而言代价过高,且产生的表示无法在样本之间对齐。我们提出了 Atelier,一个为重建的冷冻电镜密度图摊销 INR 拟合过程的自监督框架。Atelier 在 5,439 个电子显微镜数据库(Electron Microscopy Data Bank)密度图上预训练,是一个基于 Transformer 的超网络(hypernetwork),能够对广泛的蛋白质结构(包括大型多亚基组装体)生成高保真重建。除重建之外,预训练 Transformer 生成的 INR 还可通过其在任意空间查询点处的中间激活,提供连续的局部特征场——这一性质是体素网格和分块标记化(patch-tokenizer)架构无法天然提供的。将这些以坐标为条件的特征作为辅助通道,输入到一个从零开始训练的 3D 嵌套 U-Net 标注头中,在八项体素级属性预测任务上的性能均优于仅使用体积数据的基线方法。我们的结果表明,摊销式隐式神经表示是进行冷冻电镜数据几何感知分析的一种有效基础工具。
cs.AI / 16 / 2609.30571

HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases

HARDEN:用于生成更难且保持答案不变的评估用例的约束进化搜索方法
Kumaran, Aditya, Singhal, Rahul, Maamari, Karime, Mhedhbi, Amine, Tambwekar, Pradyumna
Abstract
Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fixed. HARDEN searches along generated domain-specific complexity axes while enforcing feasibility constraints such as preserving task semantics, realism, and execution validity. Across FinQA, PubMedQA, and ContractNLI and three Qwen3.5 model scales (35B-A3B, 122B-A10B, and 397B-A17B), HARDEN reduces task-model accuracy by 22.7% on average and by up to 49.9% relative to single-pass baselines using the same feasibility checks. These results show that evolutionary search can produce substantially harder valid evaluation cases.
Chinese Translation
语言模型通常在精心策划的基准测试上进行评估,而这些基准测试不足以反映企业级部署的复杂性。我们提出了HARDEN,一种约束进化搜索方法,可将现有评估用例的输入调整为更具挑战性的变体,同时保持其预期输出不变。HARDEN沿着生成的领域特定复杂性维度进行搜索,同时强制执行可行性约束,例如保持任务语义、真实性和执行有效性。在FinQA、PubMedQA和ContractNLI三个数据集以及三个Qwen3.5模型规模(35B-A3B、122B-A10B和397B-A17B)上,与使用相同可行性检查的单次基线方法相比,HARDEN平均将任务模型准确率降低了22.7%,最高降低了49.9%。这些结果表明,进化搜索能够生成难度显著更高且有效的评估用例。
cs.AI / 17 / 2609.30576

T-RoPE: Time-Aware Rotary Position Embedding for Sequential Recommendation

T-RoPE:面向序列推荐的时间感知旋转位置编码
Liu, Yang, Loo, Noel, Khanafer, Ali, Sun, Shuying, Soni, Akshay, Wu, Zhong, Yang, Linjun
Abstract
Large-scale recommenders increasingly adopt the sequential generative recipe behind large language models, bringing the Transformer into recommendation along with design choices made for text, including Rotary Position Embedding (RoPE). In language models, RoPE encodes token indices for relative position reasoning, but in recommendation, an interaction index records only event order, saying nothing about elapsed time, behavioral cycles across scales, or calendar phase. We revisit this choice and propose T-RoPE, a time-aware RoPE for sequential generative recommendation that replaces index-only rotation with timestamp-based angles, learnable temporal coefficients, multiscale frequency banks, shifted query alignment, and non-stationary key rotation. We prove that standard RoPE, even on timestamps, remains time-translation invariant and cannot distinguish seasonal contexts, and that T-RoPE breaks this invariance while preserving the RoPE interface. Across five public benchmarks, T-RoPE achieves the best result on every metric on every dataset, improving over the strongest baseline by 78--130\% in HR@10 on the sparse PixelRec data and 8--12\% across metrics on Amazon Books. On an industrial-scale e-commerce dataset with more than 6B interactions, it improves every metric over the HSTU + Time RAB backbone by 13--82\%, with ablations attributing the largest gains to multiscale frequencies ($+56\%$ NDCG@50) and non-stationary keys ($+4\%$). An online A/B test in the Shop app yields positive lifts in conversion rate ($+0.33\%$) and order count ($+0.63\%$). We also provide forward and backward algorithms whose added cost is linear in sequence length and head dimension, keeping time-aware RoPE practical for large generative recommenders.
Chinese Translation
大规模推荐系统正日益采用大语言模型背后的序列生成式范式,将Transformer引入推荐领域,同时也带来了为文本设计的选择方案,包括旋转位置编码(Rotary Position Embedding, RoPE)。在语言模型中,RoPE编码词元索引以实现相对位置推理;但在推荐场景中,交互索引仅记录事件顺序,无法反映经过的时间、跨尺度的行为周期或日历相位。我们重新审视这一设计,提出T-RoPE,一种面向序列生成式推荐的时间感知RoPE,它用基于时间戳的角度、可学习的时间系数、多尺度频率库、偏移的查询对齐以及非平稳的键旋转来取代仅基于索引的旋转。我们证明,标准RoPE即使作用于时间戳,仍保持时间平移不变性,无法区分季节性上下文;而T-RoPE在保持RoPE接口的同时打破了这种不变性。在五个公开基准上,T-RoPE在每个数据集的每项指标上均取得最佳结果:在稀疏的PixelRec数据上,HR@10相比最强基线提升78%–130%;在Amazon Books上,各指标提升8%–12%。在一个拥有超过60亿交互的工业级电商数据集上,T-RoPE在所有指标上相比HSTU + Time RAB骨干提升13%–82%;消融实验将最大收益归因于多尺度频率(NDCG@50提升56%)和非平稳键(提升4%)。在Shop应用中的在线A/B测试显示,转化率(+0.33%)和订单量(+0.63%)均获得正向提升。我们还提供了前向和后向算法,其额外开销与序列长度和注意力头维度呈线性关系,使时间感知RoPE对大型生成式推荐器保持实用。
cs.AI / 18 / 2609.30625

Audio LLMs Know When They Can't Hear You

音频大语言模型知道自己何时听不清你
Javadi, Amirhosein, Dixit, Richa, Farajtabar, Mehrdad, Cho, Minsik, Naik, Devang, Samragh, Mohammad
Abstract
Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable. We first prompt the Audio LLM to assess whether its own transcription would be reliable, and find that the model is a poor judge of its own transcription reliability: in most cases, it predicts that its transcription will be reliable. We find that existing approaches, including speech quality predictors, audio LLM generation uncertainty, and transcript-conditioned WER estimation, provide limited signals for detecting transcription failures. In contrast, we discover that transcription reliability is strongly represented in the model's audio-encoder representations. Based on this observation, we devise a lightweight reliability predictor that operates on representations extracted by the frozen audio encoder and predicts the reliability class before generation. The reliability predictor can trigger a clarification request from the user when their voice query is predicted to be unreliable, while allowing reliable queries to proceed without modifying the underlying Audio LLM. Our predictor achieves 81.10% in-domain and 78.09% cross-domain macro-F1 scores, outperforming the strongest baselines by 10.33 and 11.93 points, respectively. Finally, we show that reliability labels can transfer across Audio LLM families, and that transfer performance is closely related to the alignment of their model-specific reliability boundaries.
Chinese Translation
音频大语言模型(Audio LLM)允许用户通过语音与模型进行交互。当输入的录音质量严重退化时,模型可能会错误地理解用户的查询,并基于错误的转录内容进行回答。本文研究了模型条件下的转录可靠性问题:即音频大语言模型能否识别其自身的转录何时不可靠。我们首先提示音频大语言模型评估其自身转录是否可靠,发现该模型对自身转录可靠性的判断能力较差:在大多数情况下,它会预测其转录是可靠的。我们发现,现有方法(包括语音质量预测器、音频大语言模型的生成不确定性以及基于转录文本的WER估计)在检测转录失败方面只能提供有限的信号。与此相反,我们发现转录可靠性强烈地体现在模型的音频编码器表示中。基于这一观察,我们设计了一个轻量级的可靠性预测器,该预测器基于冻结音频编码器提取的表示进行运算,并在生成之前预测可靠性类别。当用户的语音查询被预测为不可靠时,可靠性预测器可以触发向用户发出澄清请求,同时让可靠的查询正常进行,而无需修改底层的音频大语言模型。我们的预测器在域内和跨域场景下分别取得了81.10%和78.09%的宏观F1分数,分别超越了最强基线10.33分和11.93分。最后,我们证明可靠性标签可以在不同的音频大语言模型系列之间迁移,且迁移性能与各模型特定可靠性边界的对齐程度密切相关。
cs.AI / 19 / 2609.30662

LLM Parkinsonism: Executive-Control Failure, Token-Inefficient Persistence, and an Uncertainty-Aware Global Executive Control Architecture for Autonomous Language-Model Agents

LLM帕金森化:执行控制失效、令牌低效的持续性,以及面向自主语言模型代理的不确定性感知全局执行控制架构
Xiao, Dongsheng, Wang, Zeyuan, Xia, Xuzhe, Zhao, Bo, Cao, Yankai
Abstract
Large language models (LLMs) can plan, use tools, write code, and execute long-horizon workflows, yet strong local competence does not guarantee project-level executive control. Agents may continue acting after the original objective is satisfied, producing low-value refinements, repeated verification, and repairs to self-created complexity. We use LLM Parkinsonism as a narrowly defined, non-clinical metaphor for this pattern of persistent action despite diminishing task-level value. We argue that the problem is not explained by autoregressive next-token prediction alone, but more directly by concentrating proposal generation, scope interpretation, progress assessment, and stopping authority within the same self-conditioned loop. We therefore introduce Global Executive Control (GEC) v0.2, an uncertainty-aware governance architecture that separates action generation from project-level control. In a 24,000-episode matched-candidate benchmark under a common 40,000-token ceiling, a first-candidate baseline achieved 67.42% hard-goal success, a candidate-set local control achieved 96.53%, and GEC achieved 96.57%. The candidate-set control shows that access to multiple candidate actions explains most of the success gain; relative to that control, GEC preserved success while reducing mean token use from 19,782 to 12,574 (36.4%) and restricted mean tokens to completion at the 40,000-token ceiling from 16,136 to 13,114 (18.7%), while eliminating measured pre-completion drift and sharply reducing gross complexity. Governance-overhead sensitivity remained favorable through an additional 500 synthetic governance tokens per cycle. These mechanistic simulations support explicit governance of scope, evidence, resource use, and stopping, while live-model validation remains necessary.
Chinese Translation
大语言模型(LLM)能够进行规划、使用工具、编写代码并执行长周期工作流,然而强大的局部能力并不保证项目层面的执行控制。代理可能在原目标已达成后仍继续行动,产生低价值的精修、重复验证以及对自身制造的复杂性的修补。我们将“LLM帕金森化”(LLM Parkinsonism)用作一个严格界定的、非临床性的隐喻,来描述这种在任务层面价值递减时仍持续行动的模式。我们认为,该问题不能仅用自回归的下一个词预测来解释,更直接的原因是将提议生成、范围解释、进度评估与停止权限集中于同一个自我条件化循环之中。为此,我们提出全局执行控制(Global Executive Control, GEC)v0.2,这是一种不确定性感知的治理架构,将动作生成与项目层面的控制相分离。在一个共同40,000词元上限下进行的24,000回合匹配候选基准测试中,首候选基线达到67.42%的硬目标成功率,候选集局部控制达到96.53%,而GEC达到96.57%。候选集控制表明,能够访问多个候选动作解释了大部分成功率提升;相对于该控制,GEC在保持成功率的同时,将平均词元使用量从19,782降至12,574(降低36.4%),并将40,000词元上限下的完成受限平均词元量从16,136降至13,114(降低18.7%),同时消除了可测量的完成前漂移并大幅降低总体复杂性。治理开销敏感性在每周期额外500个合成治理词元的条件下仍保持良好。这些机制性仿真支持对范围、证据、资源使用与停止条件进行显式治理,但仍有待在真实模型上进行验证。
cs.AI / 20 / 2609.30705

The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?

思考的代价:测试时推理在LLM交易中是否值得?
Chen, Jiayi, Wang, Guiling
Abstract
While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed. Our evaluation covers a full year of U.S. equities under three input conditions: numerical, identifiable news, and masked news. It includes more than 800,000 asset predictions and repeated model generations. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment.
Chinese Translation
尽管大语言模型(LLM)在推理阶段的推理能力有望带来更好的决策,但其较高的计算成本未必能产生更好的经济结果。然而,推理控制很少被作为经济干预手段来评估,即模型输出的变化能否在扣除交易成本后转化为更优的投资组合。我们对来自DeepSeek、GPT和Gemini系列的代表性大语言模型进行了受控研究。在保持每个构建日期可获得的信息、提示词、输出格式和投资组合构建方法不变的情况下,我们改变推理强度。我们的评估涵盖整整一年的美国股票,包括三种输入条件:纯数值、可识别新闻和掩码新闻,包含超过80万次资产预测和多次重复的模型生成。在所有三个模型家族中,增加推理并不能可靠地提升投资组合的净收益。对于DeepSeek,我们考察了从无推理到最大推理的完整过程,发现其表现呈非单调性。即使整体评分相近,重复生成也会产生不稳定的处理效应和投资组合选择。这些发现表明,额外的推理可以改变金融决策,却无法可靠地提升其经济价值,因此有必要在部署前针对每个任务进行验证。
cs.AI / 21 / 2609.30706

LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information

LAVOIR:利用摊销信息价值教会单次前向传播决策编码器何时问以及问什么
Yilmaz, Furkan, Tasdemir, Habibe Aleyna, Gozay, Muhammed Faruk
Abstract
"System One" decision models such as TypeSafe's Jev and its open counterpart Laya answer typed questions about a text in a single forward pass with calibrated probabilities, but they cannot ask for missing information: when a first message does not say what separates two departments, they guess. We present LAVOIR (Laya with Value-Of-Information Routing), which places the candidate pieces of missing information (slots) in the input next to the answer options, so that one forward pass returns both the decision distribution and, for every slot, the expected gain in the probability of the correct decision if the user were asked about it. VOI targets need no human labels: gold decisions come from schema rules, an LLM only verbalizes messages and answers, a model from another family checks every text, and pairing each message with several profiles makes regression on realized gains estimate the expected gain. A Gini-impurity cap bounds the predicted value by what a calibrated model can still gain. In a controlled study, decisions on seen schemas are statistically indistinguishable from the Bayes ceiling. The final model's question policy matches a greedy oracle VOI policy on seen schemas (AUC 0.799 vs. 0.797), and with at most 0.5 questions per conversation it is 14.1 points more accurate than never asking. On real ABCD conversations, one real exchange raises accuracy by 8.3 points where LAVOIR asks and leaves it unchanged where it does not; on SGD the cap lowers the asking rate from 93% to 8.6%. On Laya's twelve benchmarks LAVOIR is above Laya's reported scores on seven, and it answers a question in 31 ms (median, GH200).
Chinese Translation
“系统一”(System One)类决策模型,如 TypeSafe 的 Jev 及其开源对应版本 Laya,能够在单次前向传播中以校准的概率回答关于文本的特定问题,但它们无法主动询问缺失的信息:当首条消息未说明两个部门之间的区别时,它们只能猜测。我们提出了 LAVOIR(Laya with Value-Of-Information Routing,基于信息价值路由的 Laya),该方法将候选的缺失信息片段(槽位)与答案选项一并置于输入中,使得一次前向传播即可返回决策分布,并针对每个槽位计算:若向用户询问该信息,正确决策概率的期望增益。VOI 目标无需人工标注:黄金决策来自模式(schema)规则,一个大语言模型(LLM)仅负责将消息和答案文本化,另一模型族的模型对每条文本进行校验,而将每条消息与多个用户画像配对,可通过对已实现增益的回归来估计期望增益。基尼不纯度上限将预测价值约束在校准模型所能获得的增益范围内。在一项受控研究中,在已见模式上的决策与贝叶斯上界在统计上不可区分。最终模型在已见模式上的提问策略与贪婪的oracle VOI策略相匹配(AUC 0.799 vs. 0.797),且在每次对话最多提问0.5次的情况下,其准确率比从不提问高出14.1个百分点。在真实的 ABCD 对话上,LAVOIR 提问之处一次真实的信息交换使准确率提升8.3个百分点,而在其不提问之处准确率保持不变;在 SGD 数据集上,该上限将提问率从93%降低到8.6%。在 Laya 的十二个基准测试中,LAVOIR 在七个基准上超过了 Laya 已报告的分数,且在 GH200 上回答一个问题仅需31毫秒(中位数)。
cs.AI / 22 / 2609.30714

CRC-Router: Risk-Constrained Routing for Medical Agentic AI Systems

CRC-Router:面向医疗智能体AI系统的风险约束路由方法
Li, Xueyang, Jiang, Mingze, Xu, Gelei, Xia, Jun, Chiu, Ching-Hao, Jia, Mengzhao, Chen, Danny Z., Shi, Yiyu
Abstract
Agentic AI systems are increasingly being explored in medical imaging to improve throughput and reduce clinician workload; however, safe deployment remains challenging because autonomous errors may propagate into downstream clinical decisions. A central requirement is therefore not only strong predictive performance, but also a reliable routing mechanism that determines when the system should proceed autonomously and when a case should be escalated for further review. To address this gap, we propose CRC-Router, a risk-constrained, uncertainty-aware routing module that is applicable to both conventional medical prediction models and agentic medical AI systems. CRC-Router combines multiple complementary uncertainty signals with the predictive score to construct a per-finding routing feature vector, maps this vector to an estimated wrong-accept risk using a lightweight per-finding risk model, and then applies Conformal Risk Control (CRC) to calibrate acceptance thresholds under a user-specified risk target. Instantiated on chest X-ray multi-finding triage using the NIH ChestX-ray14 dataset, CRC-Router achieves the strongest empirical risk--coverage trade-off among the evaluated baselines, both as a standalone routing layer and as a plug-in module integrated with the state-of-the-art MedRAX agent. These results demonstrate both the effectiveness of CRC-Router in selective medical automation and its modular, model-agnostic compatibility with existing predictive and agentic medical pipelines. Code is publicly available at https://github.com/XLIAaron/CRC-Router
Chinese Translation
智能体AI系统在医学影像领域的应用日益受到关注,旨在提高处理吞吐量并减轻临床医生的工作负担;然而,其安全部署仍面临挑战,因为自主决策产生的错误可能会传播并影响下游的临床决策。因此,核心需求不仅是强大的预测性能,还需要一种可靠的路由机制,用以决定系统何时可以自主处理、何时应将病例升级以供进一步审查。为填补这一空白,我们提出了CRC-Router,这是一个风险约束、不确定性感知的路由模块,既适用于传统医学预测模型,也适用于智能体医疗AI系统。CRC-Router将多个互补的不确定性信号与预测分数相结合,构建每个病灶(per-finding)的路由特征向量;利用一个轻量级的逐病灶风险模型将该向量映射为估计的错误接受风险;然后应用保形风险控制(Conformal Risk Control, CRC),在用户指定的风险目标下校准接受阈值。在基于NIH ChestX-ray14数据集的胸部X光多病灶分诊任务上,CRC-Router无论是在作为独立路由层,还是作为与最先进的MedRAX智能体集成的插件模块时,都在所评估的基线方法中实现了最优的经验风险—覆盖权衡。这些结果既证明了CRC-Router在选择性医疗自动化中的有效性,也表明其具备模块化、模型无关的特性,可与现有的预测型及智能体医疗流程兼容。代码已公开于 https://github.com/XLIAaron/CRC-Router。
cs.AI / 23 / 2609.30725

Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents

分析与缓解编码智能体中成本低效的行为
Hu, Yiran, Jiang, Nan, Liang, Shanchao, Dey, Anik, Wu, Yi, Tan, Lin
Abstract
Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent across four configurations on SWE-bench Verified. We identify three cost-inefficient behaviors: subsumed retrieval, similar script generation, and test re-execution. We then evaluate three mitigation strategies: structure-aware retrieval, agent-synthesized skills, and developer-designed skills, over 10k trajectories on held-out SWE-bench Verified and Pro tasks. Our main findings are: (1) The three behaviors affect 79.00\%--98.00\% of coding tasks and account for up to 22.75\% of task cost. (2) Structure-aware retrieval can introduce retrieval overhead and alter agent delegation, causing inconsistent improvements in retrieval efficiency and cost increases of up to 28.14\%. (3) Agent-synthesized skills tend to produce low-level, trace-specific guidance, limiting their effectiveness and generality. (4) In contrast, developer-designed skills provide high-level, trace-agnostic guidance, reducing cost by up to 41.73\%, roughly twice the maximum gain from agent-synthesized skills.
Chinese Translation
尽管编码智能体(coding agents)效果显著,但其往往产生高昂的货币成本。其反复出现的成本低效行为仍未得到充分研究。我们首次对编码智能体中的行为成本低效问题开展了研究,分析了 Claude Code 和 Mini-SWE-Agent 在 SWE-bench Verified 上四种配置下的 1,200 条轨迹。我们识别出三种成本低效行为:被涵盖的检索(subsumed retrieval)、相似脚本生成(similar script generation)和测试重复执行(test re-execution)。随后,我们在留出的 SWE-bench Verified 和 Pro 任务上通过超过 1 万条轨迹评估了三种缓解策略:结构感知检索(structure-aware retrieval)、智能体自合成技能(agent-synthesized skills)以及开发者设计技能(developer-designed skills)。主要发现如下:(1)三种行为影响 79.00%–98.00% 的编码任务,最高占任务成本的 22.75%。(2)结构感知检索可能引入检索开销并改变智能体的任务委派方式,导致检索效率提升不一致,且成本最高增加 28.14%。(3)智能体自合成技能倾向于生成低层次、依赖特定轨迹的指导,限制了其有效性和通用性。(4)相比之下,开发者设计技能提供高层次、不依赖特定轨迹的指导,可将成本降低最高 41.73%,约为智能体自合成技能最大收益的两倍。
cs.AI / 24 / 2609.30734

Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows

学习跳过什么:面向高效多智能体大语言模型工作流的反事实信用分配
Xu, Jinfeng, Chen, Zheyu, Peng, Ziyue, Lin, Zheng, Yang, Shuo, Li, Jinze, Xing, Zheng, Li, Mengran, Leung, Victor C. M.
Abstract
Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation or overwrite a correct intermediate answer. We formulate component omission as counterfactual credit assignment: full-workflow logs reveal the executed trajectory's reward, while controlled skip interventions reveal the consequences of omitting a future step. We introduce Learning What to Skip (LW2S), which learns action-specific safety models from these interventions and combines held-out calibration with domain-native guards to select skips. When an early skip is rejected, the controller can continue execution and reconsider a later component. Across mathematical reasoning, multiple-choice QA, and code generation with two instruction-model families, LW2S reduces recorded token cost while matching or improving aggregate full-workflow accuracy in the evaluated settings. Scale-up and second-topology experiments further examine component redundancy, while shared-error cases reveal why agreement alone is insufficient for skip selection. These findings connect efficient workflow execution to learning the conditional utility of individual components.
Chinese Translation
多智能体大语言模型(LLM)工作流通过规划、执行、验证和摘要来提升任务性能,然而每个组件的价值取决于已经产生的状态。执行所有组件可能会浪费计算资源,或覆盖正确的中间答案。我们将组件省略问题形式化为反事实信用分配:完整工作流日志揭示所执行轨迹的奖励,而受控的跳过干预则揭示省略后续步骤的后果。我们提出了 Learning What to Skip(LW2S),该方法从这些干预中学习针对特定动作的安全性模型,并结合留出校准与领域原生防护机制来选择跳过操作。当早期跳过被拒绝时,控制器可以继续执行并重新考虑后续组件。在数学推理、选择题问答和代码生成任务上,使用两个指令模型家族进行实验,LW2S 在所评估的设置中降低了记录的 token 成本,同时保持或提升了完整工作流的总体准确率。规模扩展和第二种拓扑结构的实验进一步考察了组件冗余性,而共享错误案例则揭示了为何仅凭一致性不足以进行跳过选择。这些发现将高效的工作流执行与学习单个组件的条件效用联系起来。
cs.AI / 25 / 2609.30743

From S3Q Theory to Implementation: Towards an Architecture for Machine Qualia

从S3Q理论到实现:迈向机器感受质(Machine Qualia)的架构
Grinberg, Tetiana, Schleisman, Katrina, Laurent, Patryk, Udrea, Bogdan, Myers, Minda, Aufderheide, Brian, Srouji, Luis El, Groves, Doyle, Schmidt, Kevin
Abstract
A key challenge in machine consciousness research is translating theoretical models into computational-level implementations. In this paper, we address this challenge by proposing a five-layer implementation architecture for the S3Q (Simulated, Situated, Structurally Coherent) theory of consciousness. Rather than introducing novel formalisms, the architecture composes published computational primitives into a single pipeline. S3Q identifies three jointly necessary conditions for qualia: (1) grounded sensorimotor situatedness, (2) internal simulation via a world model, and (3) structural coherence between predictions and observations. No existing computational system implements all three simultaneously. We map each S3Q tenet to specific, compatible computational machinery and specify how these components interface within a single representation pipeline that operates on continuous, differentiable, per-object slot vectors, along with a developmental bootstrap sequence and falsifiable predictions for the composed system that no subset of the architecture produces in isolation. The model suggests that a basic sense of "self" develops by linking actions to their outcomes, and that behavior falls into three patterns (hesitation, curiosity, or avoidance) depending on how unexpected an outcome is and whether it is experienced as positive or negative. Each prediction is individually falsifiable, providing the field with a testable framework to advance our understanding of machine consciousness.
Chinese Translation
机器意识研究的一个关键挑战是将理论模型转化为计算层面的实现。本文通过为S3Q(模拟的、情境化的、结构连贯的)意识理论提出一个五层实现架构来应对这一挑战。该架构并未引入新颖的形式化方法,而是将已发表的计算原语组合成单一的处理流水线。S3Q理论指出了感受质(qualia)的三个联合必要条件:(1)具身的感知运动情境性,(2)通过世界模型进行的内部模拟,(3)预测与观察之间的结构连贯性。目前尚无现有计算系统能够同时实现这三者。我们将S3Q的每一条原则映射到具体且相互兼容的计算机制,并说明这些组件如何在单一表示流水线中交互——该流水线基于连续的、可微分的、逐对象的槽向量(slot vectors)运行。此外,我们还给出了一个发展性启动序列,以及针对组合系统的可证伪预测——这些预测无法由该架构的任何子集单独产生。该模型表明,基本的“自我”感是通过将行动与其结果相联系而发展起来的,且行为会根据结果的意外程度及其被体验为积极还是消极,而呈现出三种模式(犹豫、好奇或回避)。每一条预测均可独立证伪,从而为该领域提供了一个可检验的框架,以推进我们对机器意识的理解。
cs.AI / 26 / 2609.30749

ORCA: Evaluating LLMs on Data Science Code Translation

ORCA:评估大语言模型在数据科学代码翻译上的能力
Li, Xiaolong, Li, Jinyang, Qin, Bowen, Qu, Ge, Huo, Nan, Xu, Xiaohan, Lin, Shipei, Cheng, Reynold
Abstract
Data Science Code Translation (DSCT) is the process of converting code between data science libraries while preserving functional equivalence and enabling interoperability across data science ecosystems. While Large Language Models (LLMs) have demonstrated considerable progress in Data Science Code Generation (DSCG), their performance in DSCT remains insufficiently studied. To address this gap, we introduce ORCA, a comprehensive benchmark with two complementary settings: ORCA-MAIN, which comprises 1,600 carefully curated grounding-level tasks across 3 representative domains: Data Querying, Data Manipulation, and Deep Learning; and ORCA-PROJECT, which contains 200 translation tasks over complete data science projects across 7 data science task types. Each task is accompanied by annotated reference translations and test cases for validating functional equivalence. We further incorporate a multi-stage quality verification process that thoroughly verifies task correctness and test case robustness. Experimental results demonstrate challenges in DSCT, with even frontier LLMs showing limited performance. Specifically, Claude-Opus-4.6 achieves a success rate of 56.92% on ORCA-MAIN and 33.67% on ORCA-PROJECT, indicating considerable room for improvement in DSCT. We also observe a clear directional preference in DSCT, where translation is consistently easier when the source code expresses the task through more explicit, fine-grained operations. Motivated by this, we propose an intent-augmented method, in which the model first infers source-code intent and then uses it as additional context for translation, achieving average absolute success-rate gains of 4.80% and 5.33% on ORCA-MAIN and ORCA-PROJECT, respectively.
Chinese Translation
数据科学代码翻译(Data Science Code Translation, DSCT)是指在数据科学库之间转换代码的过程,同时保持功能等价性并实现跨数据科学生态系统的互操作性。尽管大语言模型(Large Language Models, LLMs)在数据科学代码生成(Data Science Code Generation, DSCG)方面已取得显著进展,但其在DSCT中的表现仍未得到充分研究。为填补这一空白,我们提出了ORCA,一个包含两种互补设置的综合基准:ORCA-MAIN包含1,600个精心策划的基础级任务,涵盖三个代表性领域:数据查询、数据操作和深度学习;ORCA-PROJECT包含200个覆盖完整数据科学项目的翻译任务,涉及7种数据科学任务类型。每个任务均配有带标注的参考译文和用于验证功能等价性的测试用例。此外,我们引入了多阶段质量验证流程,彻底验证任务正确性和测试用例的鲁棒性。实验结果表明DSCT存在诸多挑战,即使是前沿大语言模型的表现也较为有限。具体而言,Claude-Opus-4.6在ORCA-MAIN上的成功率为56.92%,在ORCA-PROJECT上为33.67%,表明DSCT仍有较大的提升空间。我们还观察到DSCT存在明显的方向性偏好:当源代码通过更显式、更细粒度的操作来表达任务时,翻译始终更加容易。受此启发,我们提出了一种意图增强方法,即模型首先推断源代码的意图,然后将其作为翻译的额外上下文,在ORCA-MAIN和ORCA-PROJECT上分别取得了平均4.80%和5.33%的绝对成功率提升。
cs.AI / 27 / 2609.30751

Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging

面向鲁棒成对大语言模型评判的主干自适应证据路由
Li, Zeyan, Peng, Jing, Xu, Jianfeng
Abstract
Pairwise language-model judges can gather evidence through direct comparison, reasoning, or reference-based verification, but no single protocol is best across benchmarks and judge backbones. We introduce Backbone-Adaptive Evidence Routing (BAER), which adapts the evidence mechanism while preserving candidate symmetry: swapping the two responses may reverse the preference but cannot change its strength. BAER separates each expert's signed preference from candidate-invariant reliability and builds three symmetric heads: evidence stacking, reliability-based expert routing, and candidate-blind reference verification. Development data select one head for each benchmark--backbone condition, and that choice is frozen before testing. Across four benchmarks and two 8B judge backbones, BAER achieves the highest test accuracy among the compared methods in all eight conditions, with full prediction coverage and gains of 0.87--7.32 points over the strongest external baseline. The results show that adapting how evidence is gathered is more reliable than fixing one judging protocol everywhere.
Chinese Translation
成对语言模型评判者(judge)可以通过直接比较、推理或基于参考的验证来收集证据,但没有任何单一协议在所有基准和评判主干上都表现最优。我们提出了主干自适应证据路由(Backbone-Adaptive Evidence Routing, BAER),它在保持候选对称性的同时自适应选择证据机制:交换两个回答的顺序可能反转偏好结果,但无法改变其强度。BAER 将每个专家的有符号偏好与候选不变的可靠性分离,并构建了三个对称头:证据堆叠、基于可靠性的专家路由以及候选盲参考验证。开发数据为每个“基准—主干”条件选择一个头,且该选择在测试前被冻结。在四个基准和两个 8B 评判主干上,BAER 在全部八种条件下的测试准确率均高于所比较的其他方法,实现了完整的预测覆盖率,并相较最强外部基线提升了 0.87 至 7.32 个百分点。结果表明,自适应地选择证据收集方式比在所有场景下固定单一评判协议更为可靠。
cs.AI / 28 / 2609.30756

Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication

面向视觉令牌通信的全预算反事实推理选择性摊销
Qi, Qinglei, Liang, Zhihe, Jing, Fengzhan, Zhu, Shenao, Zhang, Lei, Zhang, Chenyang, He, Shuqing, Guo, Jia
Abstract
Generative image communication transmits compact semantic tokens under a limited packet budget, where token selection directly affects the final reconstruction quality after the complete packet is decoded. However, accurately estimating the terminal value of every candidate token requires repeated receiver-side reconstruction, resulting in substantial encoder-side computation. To address this problem, we propose ACV-Gate, an adaptive candidate evaluation framework that learns to approximate full-budget counterfactual evaluation and selectively assigns exact evaluations to the most informative candidates. Specifically, a set-aware student is trained using terminal advantages and regrets to predict candidate rankings directly, while a selective refinement mechanism evaluates only a bounded candidate set containing both Local-MDL and direct actions; cost-based thresholds further enable explicit control of the average evaluation workload. Experiments on CIFAR-10 show that ACV-Gate consistently improves reconstruction quality while substantially reducing candidate evaluations; at 0.20 bpp, the primary adaptive configuration improves PSNR over LocalMDL by 0.636 dB with only 2.13 candidate evaluations per image, corresponding to 27.60% of the calls required by the Exact-Full expert. Matched-candidate comparisons, synchronized GPU measurements, and evaluations on STL-10 and 384 *384 scale transfer further demonstrate consistent quality computation trade-offs, with particularly pronounced gains at low bit rates. These results show that combining terminal-value learning with selective candidate evaluation provides an effective and controllable mechanism for allocating encoder computation in packet-constrained generative image communication.
Chinese Translation
生成式图像通信在有限的数据包预算下传输紧凑的语义令牌,其中令牌选择直接影响完整数据包解码后的最终重建质量。然而,准确估计每个候选令牌的终端价值需要反复进行接收端重建,从而带来巨大的编码端计算开销。为解决这一问题,我们提出了 ACV-Gate,一个自适应候选评估框架,它学习近似全预算反事实评估,并仅将精确评估分配给信息量最大的候选令牌。具体而言,通过利用终端优势(advantage)和遗憾(regret)训练一个集合感知的学生模型来直接预测候选排序;同时,一种选择性精化机制仅评估一个有界的候选集合,其中包含 Local-MDL 和直接动作;基于成本的阈值进一步实现对平均评估工作量的显式控制。在 CIFAR-10 上的实验表明,ACV-Gate 在大幅减少候选评估次数的同时,持续提升重建质量;在 0.20 bpp 下,主要自适应配置相比 LocalMDL 将 PSNR 提升了 0.636 dB,且每张图像仅需 2.13 次候选评估,仅为 Exact-Full 专家所需调用次数的 27.60%。匹配候选比较、同步 GPU 测量以及在 STL-10 和 384×384 尺度迁移上的评估进一步证明了稳定的质量—计算权衡,尤其在低码率下增益尤为显著。这些结果表明,将终端价值学习与选择性候选评估相结合,为数据包受限的生成式图像通信中的编码端计算分配提供了一种有效且可控的机制。
cs.AI / 29 / 2609.30763

HCOE: Hyperbolic Clinical Ontology Embeddings from Biomedical Language Models

HCOE:基于生物医学语言模型的双曲临床本体嵌入
Li, Yixuan, Li, Weihao, Song, Ziyang
Abstract
Biomedical language models (LMs) encode textual semantics but do not explicitly preserve medical code hierarchies. We present Hyperbolic Clinical Ontology Embeddings (HCOE) for hierarchy-aware clinical concept representation. HCOE maps frozen BioBERT embeddings into a Poincare ball, combining parent-side and child-side ontology-guided contrastive learning with coarse-to-fine ontology-path aggregation. It uses International Classification of Diseases (ICD) codes organized by Clinical Classifications Software (CCS) and Anatomical Therapeutic Chemical (ATC) medication hierarchies. Evaluations show that HCOE performs best on ICD/ATC clinical relation prediction and CCS-to-PheCode hierarchy transfer. On the MIMIC-IV dataset, HCOE also achieves the best performance on mortality prediction, readmission prediction, medication recommendation, and rare drug prediction.
Chinese Translation
生物医学语言模型(LM)能够编码文本语义,但并未显式地保留医学编码的层级结构。我们提出了双曲临床本体嵌入(Hyperbolic Clinical Ontology Embeddings, HCOE),用于层级感知的临床概念表示。HCOE 将冻结的 BioBERT 嵌入映射到 Poincaré 球面上,将父节点侧与子节点侧的本体引导对比学习相结合,并采用由粗到细的本体路径聚合方法。该方法使用由临床分类软件(Clinical Classifications Software, CCS)和解剖学治疗学化学(ATC)药物层级体系所组织的国际疾病分类(ICD)编码。评估结果表明,HCOE 在 ICD/ATC 临床关系预测以及 CCS 到 PheCode 的层级迁移任务上表现最佳。在 MIMIC-IV 数据集上,HCOE 在死亡率预测、再入院预测、用药推荐以及罕见药物预测任务上也取得了最佳性能。
cs.AI / 30 / 2609.30765

Insurance Reserve Intelligence Platform

保险准备金智能平台
A, Anugya, Mohanty, Saket, Timmapur, Abhilash, Rai, Somya
Abstract
Insurance reserve estimation is a fundamental actuarial task supporting premium pricing, solvency assessment, financial reporting, capital planning, and risk management. Classical reserve methods based on Thiele's differential equation provide a rigorous and interpretable foundation for life insurance valuation, but repeated reserve calculations become computationally expensive in sensitivity analysis, optimization, and large-scale scenario evaluation. This paper presents an Insurance Reserve Intelligence Platform for term-life reserve modelling that combines a classical Thiele-equation solver with a Physics-Informed Neural Network (PINN) enhanced by Knowledge-Informed Neural Network (KINN) losses. The framework includes synthetic policy generation, risk-adjusted premium calculation, classical reserve trajectory generation, reserve-ratio dataset construction, configurable neural training, validation diagnostics, sensitivity and elasticity analysis, prototype optimization workflows, and interest-rate scenario testing. A key refinement is the use of premium ratio and the explicit separation of pricing-time and scenario-time interest-rate semantics. The final model uses seven features: elapsed time, issue age, pricing interest rate, scenario interest rate, premium ratio, sum assured, and mortality intensity. It predicts a standardized reserve ratio instead of raw reserve values, improving numerical stability across policies with different sums assured. The model achieved an R2 of 0.9887, MAE of 785.48, and RMSE of 1212.76 on the test set. On 200 policies, PINN/KINN inference was approximately 119.53 times faster than the classical solver. Results show strong predictive accuracy, physics consistency, and boundary performance, while highlighting remaining limitations in monotonicity and out-of-distribution generalization.
Chinese Translation
保险准备金估计是一项基础精算任务,为保费定价、偿付能力评估、财务报告、资本规划和风险管理提供支持。基于Thiele微分方程的经典准备金方法为人寿保险估值提供了严谨且可解释的基础,但在敏感性分析、优化和大规模情景评估中,重复的准备金计算在计算上代价高昂。本文提出一个用于定期寿险准备金建模的保险准备金智能平台,该平台将经典Thiele方程求解器与物理信息神经网络(Physics-Informed Neural Network, PINN)相结合,并通过知识信息神经网络(Knowledge-Informed Neural Network, KINN)损失加以增强。该框架包括合成保单生成、风险调整保费计算、经典准备金轨迹生成、准备金比率数据集构建、可配置的神经网络训练、验证诊断、敏感性与弹性分析、原型优化工作流以及利率情景测试。一项关键改进是使用保费比率,并显式区分定价时点与情景时点的利率语义。最终模型使用七个特征:经过时间、投保年龄、定价利率、情景利率、保费比率、保险金额和死亡强度。模型预测标准化的准备金比率而非原始准备金数值,从而提高了不同保险金额保单之间的数值稳定性。该模型在测试集上取得了0.9887的R²、785.48的MAE和1212.76的RMSE。在200份保单上,PINN/KINN推断速度比经典求解器快约119.53倍。结果表明模型具有较强的预测精度、物理一致性和边界性能,同时也指出了在单调性和分布外泛化方面仍存在的局限性。
cs.AI / 31 / 2609.30768

Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More

思考有助于公平吗?推理标记(Reasoning Tokens)解决了部分偏见,却制造了更多偏见
Pan, Deng, Germino, Joe, Ma, Yihong, Daly, Elizabeth, Moniz, Nuno, Hua, Ting, Chawla, Nitesh
Abstract
Thinking in reasoning language models (RLMs) has been subject to debate on whether it resolves or amplifies bias. Prior works have shown competing conclusions in both directions. Using a within-model thinking-vs.-non-thinking ablation across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B on three high-stakes decision tasks (Adult, COMPAS, Credit), we show that thinking has an asymmetric dual effect on counterfactual fairness: it both resolves counterfactual flips produced by the non-thinking baseline and creates new flips at near-saturating model confidence. In all nine (model, dataset) combinations, the created flips outnumber the resolved flips by roughly 5 times. To explain the effect, we treat the thinking trace itself as a measurable site of fairness change and study it through two dynamic instruments: 1) We propose Counterfactual Depth Probability Gap (CDPG) to track bias evolution along thinking depth, and observe that bias propagates and amplifies with thinking. 2) We also formulate the Bias Transition Matrix (BTM) to show how predictions of counterfactual pairs change from non-thinking to thinking, and find that the asymmetric dual effect originates in the pair-state joint transition.
Chinese Translation
推理语言模型(RLM)中的“思考”过程究竟是缓解还是放大偏见,一直存在争议。此前的研究得出了相互矛盾的两方面结论。我们在三个高风险决策任务(Adult、COMPAS、Credit)上,对 QwQ-32B、DeepSeek-R1-Distill-Qwen-32B 和 Qwen3-32B 进行了模型内的“思考 vs. 非思考”消融实验,结果表明思考对反事实公平性具有非对称的双重效应:它既解决了非思考基线产生的反事实翻转(counterfactual flips),又在接近饱和的模型置信度下制造了新的翻转。在全部九种(模型,数据集)组合中,新制造的翻转数量约是所解决翻转数量的5倍。为解释这一效应,我们将思考轨迹本身视为公平性变化的一个可测量场所,并通过两种动态工具加以研究:1)我们提出反事实深度概率差(Counterfactual Depth Probability Gap, CDPG)来追踪偏见随思考深度的演化,观察到偏见随思考过程传播并被放大;2)我们还构建了偏见转移矩阵(Bias Transition Matrix, BTM),以展示反事实配对的预测如何从非思考状态转变为思考状态,并发现这种非对称双重效应源于配对状态的联合转移。
cs.AI / 32 / 2609.30796

ConsultMind:Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning

ConsultMind:基于不确定性感知推理的自动化诊断会诊
Sun, Xiao, Yang, Yuming, Chen, Yun, Zhong, Jiang, Zhu, Junnan, Jiang, Xinyi, Zeng, Haoyang, Chen, Ruirui, Wang, Yining, Zhou, Xinyu, Tang, Rong, Wei, Kaiwen
Abstract
Diagnostic consultation is an online sequential decision-making process in which clinicians gather evidence through patient interaction until a diagnosis is sufficiently supported. Automating this process requires adaptive inquiry and interpretable decisions. Bayesian networks offer a natural foundation by updating diagnostic posteriors as evidence accumulates, but their use in open-ended consultation raises two challenges: linking diagnostic hypotheses to potential inquiries and translating evolving posteriors into consultation decisions. We introduce AutoDisym, an automated pipeline that integrates diagnostic knowledge with heterogeneous diagnosis-labeled clinical narratives to construct a Disorder--Symptom Bayesian Network (DSBN). Building on the DSBN, we propose ConsultMind, an uncertainty-aware framework that updates disorder posteriors after each response and uses posterior uncertainty to guide inquiry and diagnosis. We evaluate both methods across psychiatry, respiratory medicine, fever clinics, and three public datasets. The results show that AutoDisym can automatically construct high-quality DSBNs and that ConsultMind consistently improves diagnostic performance and explanation soundness. For example, AutoDisym achieves macro-averaged F1 scores of 81.37 for canonical symptoms and 72.19 for manifestations using GPT-5.6-Sol. ConsultMind improves Top-1 and Top-3 diagnostic accuracy by up to 22.15 and 37.89 percentage points, respectively. Physician evaluation further shows that ConsultMind improves the quality of ranking explanations, differential diagnoses, and diagnosis rationales across LLMs of different scales. This work offers a promising approach to automatic diagnostic consultation.
Chinese Translation
诊断会诊是一个在线序贯决策过程,临床医生通过与患者的交互收集证据,直至诊断得到充分支持。自动化这一过程需要自适应的问询和可解释的决策。贝叶斯网络通过随证据累积更新诊断后验概率,为该任务提供了天然基础,但其在开放式会诊中的应用面临两个挑战:将诊断假设与潜在问询相关联,以及将不断演化的后验概率转化为会诊决策。我们提出了AutoDisym,一个将诊断知识与带有诊断标签的异构临床叙述相结合的自动化流水线,用于构建疾病—症状贝叶斯网络(DSBN)。在DSBN的基础上,我们提出ConsultMind,一个不确定性感知框架,该框架在每次患者回复后更新疾病后验概率,并利用后验不确定性指导问询和诊断。我们在精神科、呼吸科、发热门诊以及三个公开数据集上对两种方法进行了评估。结果表明,AutoDisym能够自动构建高质量的DSBN,而ConsultMind持续提升了诊断性能和解释的合理性。例如,使用GPT-5.6-Sol,AutoDisym在典型症状上取得了81.37、在具体症状表现上取得了72.19的宏平均F1分数。ConsultMind将Top-1和Top-3诊断准确率分别提升了最高22.15和37.89个百分点。医生评估进一步表明,在不同规模的大语言模型上,ConsultMind均提升了疾病排序解释、鉴别诊断和诊断依据的质量。本工作为自动化诊断会诊提供了一种有前景的方法。
cs.AI / 33 / 2609.30797

HasMem: Hard-Origin Adaptively Softened Memory for Long-Term LLM Agents

HasMem:面向长期LLM智能体的硬源头自适应软化记忆
He, Zihong, Shen, Junxiao, Liang, Chen, Liang, Hai-Ning
Abstract
Text-based memory and context compression support reuse of past interactions. Resizing continuous memory changes the input to a frozen LLM, coupling capacity allocation with readout. We propose Hard-Origin Adaptively Softened Memory (HasMem). Frozen hard-prompt embeddings provide a verifiable initial state. A controller adjusts memory widths, a Writer re-encodes resized entries, and Reader and Global provide readout adaptation and cross-turn state. On all $535$ questions in a reconstruction probe derived from the Multi-Session Chat (MSC) development split, the main configuration achieves lexical F1 of $95.3$ ($+4.4$ percentage points) at $93.6\%$ of the hard reference's framed memory positions. With approximately matched per-question target body budgets, six configurations at mean per-entry retention around $0.83$--$0.91$ exceed rule-based re-encoding by $8.0$--$23.6$ exact-match (EM) percentage points. With fixed model parameters and rule target width ratio $0.75$, Global's EM gain passes a user-level exact paired test with Bonferroni correction over eight comparisons. On all $500$ LongMemEval-S questions, local lexical F1 rises from the hard reference's $3.4$ to $8.9$, and answer negative log-likelihood (NLL) falls from $12.257$ to $5.274$. F1 gains accompany lower EM on both evaluations.
Chinese Translation
基于文本的记忆与上下文压缩支持过往交互的复用。对连续记忆进行尺寸调整会改变冻结大语言模型(LLM)的输入,使容量分配与读取耦合在一起。我们提出了硬源头自适应软化记忆(Hard-Origin Adaptively Softened Memory,HasMem)。冻结的硬提示词(hard-prompt)嵌入提供了可验证的初始状态。控制器调整记忆宽度,写入器(Writer)对调整尺寸后的条目重新编码,读取器(Reader)与全局模块(Global)提供读取自适应与跨轮次状态。在由多会话对话(Multi-Session Chat,MSC)开发集派生的重构探针的全部535个问题上,主配置在硬参照(hard reference)框架化记忆位置93.6%的用量下达到词法F1值95.3(提升4.4个百分点)。在每问题目标主体预算大致匹配的条件下,六种配置在平均每条目保留率约0.83–0.91的情况下,比基于规则的重新编码高出8.0–23.6个精确匹配(EM)百分点。在模型参数固定且规则目标宽度比为0.75的条件下,全局模块的EM增益通过了针对八次比较的经Bonferroni校正的用户级精确配对检验。在LongMemEval-S的全部500个问题上,局部词法F1从硬参照的3.4提升至8.9,答案负对数似然(NLL)从12.257降至5.274。在两项评测中,F1的提升均伴随EM的下降。
cs.AI / 34 / 2609.30798

Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes

实时语音智能体评估:从组件质量到落地效果
Negi, Shivam, Rawat, Arpit, Jain, Rashi
Abstract
Real-time voice agents have moved from research prototypes to production deployments, yet the literature describing them is fragmented across three communities that rarely cite one another: speech foundation modelling, turn-taking psycholinguistics, and agentic evaluation. Architecture papers report latency, turn-taking papers report prediction accuracy, and agentic benchmarks report task success, so no single number describes whether a deployed agent is actually good. We address that gap with three evidence-based claims, each traceable to a corpus of 38 primary sources organised into an application-centric taxonomy of six categories. First, architecture choice is a deployment constraint rather than a settled verdict: a 2026 enterprise tutorial reports that no fully self-hostable end-to-end system yet meets production constraints, while a chunked cascade independently reaches state-of-the-art duplex behaviour, showing duplex behaviour is separable from duplex architecture. Second, evaluation has shifted decisively from component quality toward grounded outcomes, with recent benchmarks verifying backend state rather than trusting what the agent claims to have done. Third, the dyadic assumption in most models and benchmarks is breaking down: multiparty turn-taking and multi-speaker reasoning benchmarks show that deciding when not to speak, and reasoning about who may be told what, are first-class capabilities two-participant framings cannot measure. For each source we state the problem it targets, its mechanism, and its reported evidence, alongside the search strategy, inclusion criteria, and a verification step that caught a misattributed arXiv identifier in circulation. We propose TRG (Timing-Recovery-Grounded), a reporting standard characterising an agent by timing, post-disruption recovery, and state-verified outcome together, with a conditional fourth axis for multiparty deployments.
Chinese Translation
实时语音智能体已从研究原型走向生产部署,但描述它们的文献分散在三个很少互相引用的社区中:语音基础建模、话轮转换心理语言学以及智能体评估。架构论文报告延迟,话轮转换论文报告预测准确率,智能体基准测试报告任务成功率,因此没有任何单一指标能够说明一个已部署的智能体究竟是否优秀。我们基于38篇一手文献组成的语料库(按面向应用的分类法划分为六个类别),提出三条有据可依的论断来弥补这一空白。第一,架构选择是一种部署约束而非已有定论:一篇2026年的企业教程指出,目前尚无完全可自托管的端到端系统能满足生产环境约束,而分块级联架构独立实现了最先进的双工行为,表明双工行为可以与双工架构相分离。第二,评估已决定性地从组件质量转向落地效果,近期的基准测试通过验证后端状态而非采信智能体声称已完成的操作来进行评估。第三,大多数模型和基准测试中的二元对话假设正在瓦解:多方话轮转换与多说话人推理基准测试表明,决定何时不应发言以及推理谁可以被告知什么,是双人参与框架无法衡量的一等能力。对于每篇文献,我们陈述其所针对的问题、机制及报告的证据,并同时给出检索策略、纳入标准,以及一个曾发现一处流传中被错误归属的arXiv编号的验证步骤。我们提出TRG(时序-恢复-落地,Timing-Recovery-Grounded)报告标准,通过时序、中断后恢复以及经状态验证的结果共同刻画智能体,并为多方部署设置有条件的第四个维度。
cs.AI / 35 / 2609.30813

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

共享智能体记忆中认知性准入的基准测试与诊断研究
Li, Xiaoyang, Wang, Yiqi, Zhu, Chencheng, XU, KE, Yang, Wencheng, Sun, Zequn, Song, Pingan, Duan, Yiqun, Cai, Taotao
Abstract
Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent agents to it. To study this problem, we introduce the Correlated Promotion Benchmark (CPB), which evaluates whether candidate claims should be admitted to shared memory.CPB-Static constructs a frozen test split from publicly annotated sources with fixed gold actions. CPB-Live runs multi-agent teams over a shared store, records all writes and retrievals, and tracks source lineage defined by each scenario. A separate consumer answers from the store alone. We evaluate eight admission policies across four agent families. Our results show that policies which deduplicate sources reject many true claims alongside false ones, whereas policies preserving answer coverage admit nearly as many false claims as unrestricted sharing. Gating on declared source type reduces false adoption to 0.06--0.09, compared with 0.22--0.47 for other answering policies. Once an uncontested false belief enters memory, the consumer asserts it in 0.97--0.99 of probes across all families. No non-oracle policy consistently rejects false claims across verbatim copies, paraphrases, and paraphrases declared authoritative. These findings reveal the limitations of admission policies without access to source lineage.
Chinese Translation
评估共享智能体记忆中的断言准入极具挑战性,因为重复的断言可能被误认为是独立的证据。智能体可能会复制或转述检索到的信念,而准入一条错误断言会使后续智能体暴露于该错误信息之下。为研究这一问题,我们提出了相关促进基准(Correlated Promotion Benchmark, CPB),用于评估候选断言是否应被准入共享记忆。CPB-Static 从带有固定黄金动作(gold actions)的公开标注来源构建了一个冻结的测试集划分;CPB-Live 让多智能体团队在共享存储上运行,记录所有写入与检索操作,并跟踪由每个场景定义的来源谱系(source lineage)。一个独立的消费者智能体仅依据该存储中的内容进行回答。我们在四个智能体家族上评估了八种准入策略。结果表明,对来源进行去重的策略在拒绝错误断言的同时也拒绝了大量真实断言,而保留答案覆盖率的策略所准入的错误断言数量几乎与无限制共享相当。基于声明的来源类型进行门控可将错误采纳率降至 0.06–0.09,而其他回答策略的错误采纳率为 0.22–0.47。一旦一条未受质疑的错误信念进入记忆,所有智能体家族中的消费者在 0.97–0.99 的探测中都会断言该信念。没有任何非先知(non-oracle)策略能够在逐字复制、转述以及被声明为权威的转述这三种情形下始终如一地拒绝错误断言。这些发现揭示了无法获取来源谱系的准入策略的局限性。
cs.AI / 36 / 2609.30836

PTC-Decoder: Towards Intelligent SLMs on Offline Resource-Constrained Edge Devices

PTC-Decoder:面向离线资源受限边缘设备上的智能小型语言模型
Yu, Minghui, Mu, Ke, Wu, Gang
Abstract
Deploying small language models (SLMs) on offline, resource-constrained edge devices such as remote sensing satellites presents a fundamental challenge: their limited reasoning capacity hinders reliable execution of multi-step agent tasks requiring complex tool orchestration. Existing plan-solve paradigms rely on prompt-based enforcement, which our experiments show SLMs almost entirely disregard: weak models fail to invoke the plan. We propose PTC-Decoder (Plan-Tool Constrained Decoder), a training-free, plug-and-play decoder framework that combines (1) a Plan-to-Act paradigm, which elevates planning to an atomic tool and forces its invocation at the first inference step, and (2) TC-Decoder, a deterministic finite automaton that imposes token-level hard constraints on tool names while preserving freedom over parameter generation, thereby retaining SLM reasoning capability. Evaluated on 200 real remote-sensing satellite tasks across 7 SLMs, PTC-Decoder yields a statistically significant mean overall score gain of +1.21 (p<0.01), 95% CI [+1.13, +1.29]), with consistent improvements across models and other datasets. An ablation study that removes TC-Decoder causes substantial performance degradation across all quality metrics without reducing computational cost, confirming TC-Decoder as the primary driver. PTC-Decoder thus offers a lightweight yet effective solution for improving step-level reliability, with final-answer accuracy remaining an open challenge. In essence, we enforce plan adherence by constraining the permissible output vocabulary during inference, without requiring retraining.
Chinese Translation
在遥感卫星等离线、资源受限的边缘设备上部署小型语言模型(SLM)面临一个根本性挑战:其有限的推理能力阻碍了需要复杂工具编排的多步智能体任务的可靠执行。现有的“规划-求解”(plan-solve)范式依赖基于提示词的强制机制,而我们的实验表明SLM几乎完全无视这种机制:弱模型无法调用规划。我们提出了PTC-Decoder(Plan-Tool Constrained Decoder),这是一个无需训练、即插即用的解码器框架,它结合了:(1)“规划到行动”(Plan-to-Act)范式,将规划提升为一种原子工具,并强制其首次推理时即被调用;(2)TC-Decoder,一个确定性有限自动机,对工具名称施加词元级硬约束,同时保留参数生成的自由度,从而保留SLM的推理能力。在200个真实遥感卫星任务上对7个SLM进行评估,PTC-Decoder带来了具有统计显著性的平均总分提升+1.21(p<0.01,95%置信区间[+1.13, +1.29]),且在不同模型和其他数据集上均有一致的改进。消融实验表明,移除TC-Decoder会导致所有质量指标的性能显著下降,且计算成本并未降低,证实TC-Decoder是主要驱动因素。因此,PTC-Decoder为提升步骤级可靠性提供了一种轻量而有效的解决方案,而最终答案的准确性仍是一个开放性挑战。本质上,我们通过在推理过程中约束允许的输出词表来强制规划遵循,而无需重新训练。
cs.AI / 37 / 2609.30841

Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis

为什么越狱攻击在扩散语言模型中能够成功:一种能量景观分析
Bach, Thong, Nguyen, Dung, Le, Thao Minh, Tran, Truyen
Abstract
Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscape: a well-aligned model routes harmful queries toward safe outputs through an energy barrier that separates the two regions. Current jailbreak attacks reduce to two strategies for circumventing this barrier: obscuring the query's safety disposition at initialisation, or intervening mid-trajectory to force the denoising path across the energy barrier. From this perspective and the result that masked diffusion models minimise kinetic energy during denoising, we derive three complementary, training-free detection signals: a step-0 ratio that reads the initial safety disposition from the logit distribution before generation begins, and two trajectory-velocity signals that track kinetic energy in complementary subspaces of the logit space. An attack must either reveal its intent at initialisation or expend kinetic energy to cross the barrier in at least one monitored subspace, so the three signals cover each other's blind spots in the energy budget by construction. Evaluation across three dense dLLMs (LLaDA-8B, LLaDA-1.5, Dream-7B) and a sparse mixture-of-experts dLLM (LLaDA-MoE-7B) confirms this complementarity. In stress tests of known attacks, every configuration that evades detection also fails to produce harmful content, suggesting that the detection and barrier-crossing thresholds are hard to separate.
Chinese Translation
现有的针对扩散式大语言模型(dLLM)的攻击与防御方法均针对特定漏洞,但缺乏一个能够解释攻击为何成功的统一框架。我们提出了这样一个框架:将安全对齐理解为对去噪能量景观的塑造——一个对齐良好的模型通过一个将两个区域分隔开的能量势垒,将有害查询引导至安全输出。当前的越狱攻击可以归结为绕过该势垒的两种策略:一是在初始化时模糊查询的安全倾向,二是在去噪轨迹中途进行干预,迫使去噪路径跨越能量势垒。基于这一视角以及掩码扩散模型在去噪过程中最小化动能这一结果,我们推导出三种互补的、无需训练的检测信号:一个是 step-0 比率,它在生成开始之前从 logit 分布中读取初始安全倾向;另两个是轨迹速度信号,它们在 logit 空间的互补子空间中跟踪动能。任何攻击要么在初始化时暴露其意图,要么必须消耗动能以在至少一个受监控子空间中跨越势垒,因此这三种信号在能量预算上天然地互相覆盖盲区。在三个稠密 dLLM(LLaDA-8B、LLaDA-1.5、Dream-7B)和一个稀疏混合专家 dLLM(LLaDA-MoE-7B)上的评估证实了这种互补性。在对已知攻击的压力测试中,所有规避检测的配置同时也无法生成有害内容,这表明检测阈值与势垒跨越阈值难以分离。
cs.AI / 38 / 2609.30861

SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting

SkillEvoReg:针对过拟合的智能体技能演化正则化方法
Nie, Guanyu, Zhu, Fangzhou, Kai, Shixiong, Han, Xiongwei, Zhong, Tao, Yuan, Mingxuan
Abstract
Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task-specific instructions, while new updates can disrupt behavior that previously worked. We study this problem as skill-evolution overfitting and introduce SkillEvoReg, a general regularization framework for skill evolution inspired by anti-overfitting techniques in neural-network training. SkillEvoReg combines training-time skill dropout, which perturbs update generation, and complexity-aware local regularization, which controls unnecessary structural growth, with causal counterexample validation (CCV), which provides targeted behavioral validation of candidate-specific regressions. We instantiate the framework across heterogeneous skill-evolution systems while retaining each system's native skill evolver and task evaluator. Across SkillOpt, SkillEvolBench, and ContinualSkillBench, SkillEvoReg consistently controls skill-state growth while preserving competitive downstream capability, improves several transfer and later-stage evolution outcomes, and identifies update-level regressions that structural metrics alone cannot reveal. These results suggest that explicit regularization is a useful complement to increasingly capable skill updaters.
Chinese Translation
语言模型智能体日益通过将执行经验转化为可复用的外部技能来提升自身能力。然而,反复的技能更新本身构成了一个学习过程:局部有用的修改可能累积为冗余或任务特定的指令,而新的更新也可能破坏此前运行良好的行为。我们将该问题定义为技能演化过拟合(skill-evolution overfitting),并提出 SkillEvoReg——一个受神经网络训练中抗过拟合技术启发的通用技能演化正则化框架。SkillEvoReg 将训练时的技能随机失活(skill dropout,用于扰动更新生成)与复杂度感知的局部正则化(控制不必要的结构性增长)相结合,并引入因果反例验证(Causal Counterexample Validation, CCV),对候选更新导致的特定性能回退进行有针对性的行为验证。我们在异构的技能演化系统上实例化该框架,同时保留各系统原有的技能演化器和任务评估器。在 SkillOpt、SkillEvolBench 和 ContinualSkillBench 上的实验表明,SkillEvoReg 能够持续控制技能状态的增长,同时保持具有竞争力的下游能力,改善多项迁移及后期演化结果,并能够识别出仅凭结构性指标无法发现的更新层面的性能回退。这些结果表明,显式正则化是对日益强大的技能更新器的有益补充。
cs.AI / 39 / 2609.30878

TISD: On-Policy Self-Distillation with Trajectory Intervention

TISD:基于轨迹干预的在线策略自蒸馏
Lee, Taeckyung, Amankos, Rinat, Kim, Jeonghye, Yoon, Hyungjun, Jin, Woogyeol, Lee, Sung-Ju
Abstract
On-policy self-distillation (OPSD) provides dense teacher targets, but evaluates them only along student-sampled rollouts. When the privileged teacher favors an alternative action at a visited prefix, OPSD can provide a target for the branch decision but cannot supervise the successor contexts induced by that action unless the student samples it. This creates a training-time data-collection bottleneck and suggests a different role for teacher-student disagreement: proposing a trajectory branch rather than identifying a sufficient local repair. Our diagnostic framework using controlled token interventions reveals that a teacher-preferred token at peak disagreement can improve student continuation success, while its local corrective value is limited. Motivated by this finding, we introduce a simple branch-regenerate-distill algorithm, Trajectory-Intervention Self-Distillation (TISD). TISD forces a teacher-selected branch action, returns suffix generation to the student, and distills the full trajectory under the privileged-context-conditioned teacher. Across the coding models, TISD improves average Avg@4 over SDPO by 1.2 percentage points. Across the science domains, it improves average Avg@128 by 0.8 points under an equal-step budget and by 0.3 points under an equal-time budget. These results support teacher-guided branching as a way to expose useful successor contexts for self-distillation.
Chinese Translation
在线策略自蒸馏(On-policy Self-Distillation,OPSD)能够提供密集的教师目标,但仅在学生采样的 rollout 上对这些目标进行评估。当拥有特权的教师在已访问的前缀处偏好另一个动作时,OPSD 可以为该分支决策提供目标,却无法监督由该动作引出的后续上下文,除非学生恰好采样到它。这造成了训练时的数据收集瓶颈,并提示了师生分歧的另一种作用:提出一条轨迹分支,而非识别一种充分的局部修复。我们采用受控词元干预的诊断框架发现,在分歧峰值处教师偏好的词元能够提升学生后续生成的成功率,而其局部纠正价值有限。受此发现启发,我们提出了一种简单的“分支—再生—蒸馏”算法:轨迹干预自蒸馏(Trajectory-Intervention Self-Distillation,TISD)。TISD 强制执行教师选择的分支动作,将后缀生成交还给学生,并在特权上下文条件化的教师指导下对完整轨迹进行蒸馏。在代码生成模型上,TISD 相比 SDPO 将平均 Avg@4 提升了 1.2 个百分点。在科学领域上,在等步数预算下平均 Avg@128 提升了 0.8 个百分点,在等时间预算下提升了 0.3 个百分点。这些结果表明,教师引导的分支是一种为自蒸馏揭示有用后续上下文的有效方式。
cs.AI / 40 / 2609.30880

EXAONE Demand 1.0: A Time Series Foundation Model for Demand Forecasting

EXAONE Demand 1.0:面向需求预测的时间序列基础模型
Lee, Seunghan, Han, Sangjun, Seo, Jun, Kang, Junhyeok, Lee, Jaehoon, Lim, Tae Yoon, Kang, Dongwan, Choi, Hwanil, Kim, Minjae, Yoo, Sungdong, Lee, Soonyoung, Ahn, Wonbin
Abstract
Time series foundation models (TSFMs) are pretrained on series from diverse domains, where demand series make up only a small fraction. Demand data has properties that such corpora rarely contain: Short histories, frequent zeros, censoring by stock-outs, and exogenous events that the series does not record. To this end, we propose EXAONE Demand, built on 1) a demand-specific corpus and 2) a demand-aware adapter. For the corpus, we assemble 11.3M series and 48.4B observations from 73 sources, and a synthetic generator supplies the behaviour that open demand data under-represents. For the adapter, we attach low-rank branches to a frozen general-domain backbone, one for each of the four demand classes (smooth, intermittent, erratic, and lumpy), and a router that reads eight scale-free statistics of the input series decides how much each branch contributes. We build EXAONE Demand in two versions, one trained on real-world and synthetic demand together and one trained on the synthetic corpus alone. On 22 held-out datasets, both versions outperform 36 TSFMs, and real-world demand adds a gain over synthetic data alone.
Chinese Translation
时间序列基础模型(TSFM)在来自多种领域的时间序列上进行预训练,其中需求类序列仅占很小比例。需求数据具有此类语料库中罕见的一些特性:历史数据较短、频繁出现零值、因缺货而产生的数据删失(censoring),以及序列本身未记录的外生事件。为此,我们提出了EXAONE Demand,其构建基于两方面:1)面向需求的专用语料库,2)需求感知的适配器(adapter)。在语料库方面,我们从73个数据源收集了1130万条序列和484亿个观测值,并通过合成生成器补充了开放需求数据中代表性不足的行为模式。在适配器方面,我们在冻结的通用领域主干网络上附加低秩分支,分别对应四类需求模式(平滑型、间歇型、波动型和块状型),并由一个读取输入序列八个无标度统计量的路由器(router)决定各分支的贡献权重。我们构建了两个版本的EXAONE Demand:一个在真实世界与合成需求数据上共同训练,另一个仅在合成语料库上训练。在22个保留测试数据集上,两个版本均优于36个时间序列基础模型,且相比仅使用合成数据,引入真实世界需求数据带来了额外的性能提升。
cs.AI / 41 / 2609.30887

From Tapping to Hopping: Augmenting Mobile GUI Agents with App-Native Deeplinks

从点按到跳转:利用应用原生深链接增强移动GUI智能体
Sun, Yuchen, Cai, Chenglin, Zhang, Gongjie, Xia, Tianyu, Kong, Quyu, Tong, Panrong, Zeng, Zhengwen, Chen, Long, Hoi, Steven, Zhang, Chongyang, Wang, Yue
Abstract
Mobile GUI agents complete tasks using GUI actions like taps and swipes. These actions are broadly applicable across applications, but reaching a navigation interface. A single deeplink call can replace a sequence of screen-by-screen GUI actions. We therefore introduce hybrid interaction, using deeplinks for direct navigation and GUI actions for other on-screen operations and fallback. To enable this, we discover candidate deeplinks through static analysis, validate them on real devices, and describe their observed landing screens. This process creates a verified and grounded deeplink catalog that pairs each working deeplink with a description of its landing screen. Using this catalog, we introduce GUI-Hopper, a improves task success in commercial applications on real devices, further demonstrating the benefits of hybrid interaction.
Chinese Translation
移动GUI智能体通过点按、滑动等GUI操作来完成任务。这些操作具有跨应用的广泛适用性,但往往需要经过多步导航才能到达目标界面。而一次深链接(deeplink)调用即可替代逐屏操作的一系列GUI动作。因此,我们提出了混合交互方式:利用深链接实现直接导航,而对其他屏幕内操作及回退情形则使用GUI动作。为此,我们通过静态分析发现候选深链接,在真实设备上对其进行验证,并描述其观测到的落地页面。该过程构建了一个经验证且有据可依的深链接目录,将每个可用的深链接与其落地页面的描述配对。基于该目录,我们提出了GUI-Hopper,它在真实设备的商业应用中提升了任务成功率,进一步证明了混合交互方式的优势。
cs.AI / 42 / 2609.30894

Training Graph Foundation Models on The Web Graph

在Web图上训练图基础模型
Sato, Ryoma
Abstract
We introduce Acacia, a graph foundation model, trained on the web graph. Acacia (i) supports arbitrary feature dimensionalities and semantics without additional training, (ii) supports a wide range of tasks, including node classification, link prediction, node clustering, and graph generation, without additional training, (iii) has in-context learning capabilities, and (iv) does not rely on pretrained LLMs. In particular, existing graph foundation models often require training additional classification heads or feature projectors to accommodate new graphs or new labels, whereas Acacia does not. Moreover, existing graph foundation models often gain their capabilities by being stitched together with pretrained LLMs, whereas Acacia is trained from scratch using only the Common Crawl web graph. This is also an important result because it provides evidence that graph models can acquire emergent capabilities from scratch like LLMs.
Chinese Translation
我们提出了Acacia,一个在Web图上训练的图基础模型。Acacia(i)无需额外训练即可支持任意的特征维度和特征语义;(ii)无需额外训练即可支持多种任务,包括节点分类、链接预测、节点聚类和图生成;(iii)具备上下文学习能力;(iv)不依赖预训练的大语言模型(LLM)。特别地,现有的图基础模型通常需要训练额外的分类头或特征投影器来适应新的图或新的标签,而Acacia无需如此。此外,现有的图基础模型往往通过与预训练大语言模型拼接来获得其能力,而Acacia仅使用Common Crawl的Web图从零开始训练。这也是一项重要的成果,因为它提供了证据表明图模型可以像大语言模型一样从零开始获得涌现能力。
cs.AI / 43 / 2609.30922

JevSoup: System-One Routing for Training-Free LoRA Composition

JevSoup:面向免训练LoRA组合的系统一路由
Wang, Xiuying, Cheng, Jiahua, Li, Shuotian, Cheng, Yufan, Deng, Bowen, Bai, Zhexuan, Li, Yichen
Abstract
Building adaptable AI systems requires effective coordination of specialized capabilities across diverse tasks. Low-rank adaptation (LoRA) enables modular expertise, but existing routing approaches may require auxiliary data, additional training, or autoregressive decoding. We propose JevSoup, a training-free framework separating System One expert routing from System Two execution. Using only the input and expert descriptions, Jev selects two experts through structured probabilities. JevSoup retains the leading expert's update, projects the second onto the orthogonal complement of the first update's row space, and combines them with equal weights. Across 14 PorTAL tasks and three Qwen3 scales, JepSoup achieves absolute gains of up to 1.19\% in task-macro and 1.21\% in sample-micro accuracy over the strongest evaluated external baselines. Our code is available at https://github.com/Leowang980/JevSoup.
Chinese Translation
构建可适应的AI系统需要在不同任务间有效协调专门化能力。低秩适配(LoRA)实现了模块化的专业知识,但现有的路由方法可能需要辅助数据、额外训练或自回归解码。我们提出JevSoup,一个将系统一(System One)专家路由与系统二(System Two)执行相分离的免训练框架。仅利用输入和专家描述,Jev通过结构化概率选择两个专家。JevSoup保留首个专家的更新,将第二个专家的更新投影到第一个更新的行空间的正交补空间上,并以相等权重将二者结合。在14个PorTAL任务和三种Qwen3模型规模上,JevSoup相对于所评估的最强外部基线,任务宏平均(task-macro)准确率最高提升1.19%,样本微平均(sample-micro)准确率最高提升1.21%。我们的代码可在 https://github.com/Leowang980/JevSoup 获取。
cs.AI / 44 / 2609.30936

Self-Play Search Distillation for Large Language Model Reasoning

面向大语言模型推理的自博弈搜索蒸馏
Molfetta, Lorenzo, Kwan, Wai-Chung, Frisoni, Giacomo, Ragazzi, Luca, Moro, Gianluca, Vougiouklis, Pavlos, Pan, Jeff Z., Minervini, Pasquale
Abstract
Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games. SPSD uses executable environments to turn search into structured reasoning problems. At each state, the expert identifies a preferred decision, plausible alternatives, plausible opponent replies, and value estimates. By converting the self-play search records into superhuman chains-of-thought, we train LLMs with environment-grounded supervision. Although trained only on self-play search records, SPSD transfers to unseen mathematics. On Qwen3-4B-Base, it raises the mean over six mathematics benchmarks from 24.1 to 36.6 while increasing the held-out-game win rate from 15% to 45%. SPSD offers an annotation-efficient way to create high-quality synthetic data for improving LLM performance in reasoning tasks.
Chinese Translation
提升大语言模型(LLM)的推理能力需要高质量的数据,这些数据应能够展现困难的决策、相互竞争的备选方案及其后果。数据稀缺的原因在于合成数据质量较低以及人工标注成本高昂。我们提出了自博弈搜索蒸馏(Self-Play Search Distillation, SPSD),这是一个通过在棋盘游戏上训练的类MuZero网络进行自博弈来生成超人水平合成数据的框架。SPSD利用可执行环境将搜索转化为结构化的推理问题。在每一个状态下,专家智能体会识别出首选决策、合理的备选方案、合理的对手应对以及价值估计。通过将自博弈搜索记录转化为超人水平的思维链,我们以环境为依据的监督信号来训练大语言模型。尽管仅使用自博弈搜索记录进行训练,SPSD仍能迁移到未见过的人类数学任务上。在Qwen3-4B-Base模型上,它将六个数学基准测试的平均成绩从24.1提升到36.6,同时将保留游戏的胜率从15%提升到45%。SPSD提供了一种标注高效的方式来创建高质量合成数据,从而提升大语言模型在推理任务中的表现。
cs.AI / 45 / 2609.30939

MACBT: A Multi-Agent Cognitive Behavioral Therapy Decision Support System with Longitudinal Memory

MACBT:具有纵向记忆的多智能体认知行为疗法决策支持系统
Jiang, De, Zhang, Shuo, Liao, Weiwei, Zhang, Jianying, Yu, Chuanhui, Liao, Hongen, Yuan, Kehong
Abstract
Cognitive behavioral therapy (CBT) is an evidence-based first-line treatment for depression, yet its scale is constrained by the time clinicians spend on pre-session preparation, post-session documentation, and longitudinal cognitive-pathology tracking. We present a clinician-facing AI decision-support system that combines a multi-agent CBT framework (MACBT) with a CBT-specific longitudinal memory module (CD Memory). MACBT encodes the five-stage CBT workflow (assessment, Socratic questioning, cognitive restructuring, behavioral experiments, and treatment monitoring) into five collaborative agents. CD Memory tracks cognitive-distortion type, frequency, severity, and restructuring efficacy across sessions to generate pre-session pathology reports and intervention-priority recommendations. We construct a Chinese CBT dialogue corpus via dual-role large language model simulation and train a Qwen3-14B backbone with supervised fine-tuning and direct preference optimization. Evaluation with GPT-4 judges shows MACBT outperforms MeChat, SoulChat, PsyChat, and CPsyCounX in professionalism (2.62) and clinical authenticity (2.25). The full memory-augmented system further improves session quality by 12.6% and achieves a longitudinal mean of 2.29 on cross-session continuity, intervention progression, and personalization.
Chinese Translation
认知行为疗法(CBT)是一种基于循证依据的抑郁症一线治疗方法,但其规模化受到临床医生在会谈前准备、会谈后记录以及纵向认知病理追踪上所花费时间的限制。我们提出了一个面向临床医生的AI决策支持系统,该系统将多智能体CBT框架(MACBT)与CBT专用的纵向记忆模块(CD Memory)相结合。MACBT将五个阶段的CBT工作流程(评估、苏格拉底式提问、认知重构、行为实验和治疗监测)编码为五个协作智能体。CD Memory跨会谈追踪认知扭曲的类型、频率、严重程度及重构效果,以生成会谈前病理报告和干预优先级建议。我们通过双角色大语言模型模拟构建了中文CBT对话语料库,并采用监督微调和直接偏好优化训练了Qwen3-14B骨干模型。基于GPT-4评审的评估表明,MACBT在专业性(2.62)和临床真实性(2.25)方面优于MeChat、SoulChat、PsyChat和CPsyCounX。完整的记忆增强系统进一步将会谈质量提升了12.6%,并在跨会谈连贯性、干预进展和个性化方面取得了2.29的纵向平均得分。
cs.AI / 46 / 2609.30940

Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms

LLM智能体社会中的金融脆弱性:协调失灵与稳定机制
Fu, Zhenhao, Xu, Ruipeng, Ren, Qibing
Abstract
Individually protective decisions can produce avoidable collective failures. As large language model (LLM) agents take on greater roles in financial decision-making, financial AI safety must therefore be considered not only at the level of individual agents, but also at the level of the systems they jointly create. We study this problem with FRAIL, a controlled experimental framework that places LLM agents in three dynamic financial environments---bank runs, debt rollover, and reward crowdfunding---where agents' decisions reshape the financial conditions faced by others. Across seven leading LLMs, we find widespread collective fragility even when no agent is instructed to destabilize the system: 77\% of baseline bank-run episodes and 83\% of debt-rollover episodes end in failure. We then compare three interaction mechanisms based on compensated commitments, centralized commitment agreements, and participant-led coalitions. All three improve aggregate outcomes, but no single mechanism performs best across all financial structures. Across mechanisms, successful stabilization shares a common temporal pattern: broad commitment forms early, before defensive behavior becomes self-reinforcing. Our findings show that individually capable agents do not automatically form safe financial systems, highlighting system-level evaluation and interaction design as central problems for financial AI safety. Code is available at https://anonymous.4open.science/r/FinFrail-CF26.
Chinese Translation
个体出于自我保护的决策可能导致本可避免的集体性失败。随着大语言模型(LLM)智能体在金融决策中扮演越来越重要的角色,金融AI安全性不仅需要在个体智能体层面加以考量,还需要在它们共同构成的系统层面加以考量。我们通过FRAIL框架研究这一问题,这是一个受控实验框架,将LLM智能体置于三种动态金融环境中——银行挤兑、债务展期和奖励型众筹——在这些环境中,智能体的决策会重塑其他智能体所面临的金融条件。在七个主流LLM上,我们发现即使没有任何智能体被指示去破坏系统稳定,仍然普遍存在集体脆弱性:77%的基线银行挤兑场景和83%的债务展期场景以失败告终。随后,我们比较了三种基于不同交互机制的方案:有偿承诺、中心化承诺协议以及参与者主导的联盟。三种机制均能改善总体结果,但没有一种机制在所有金融结构下都表现最佳。在各机制之间,成功的稳定化都呈现出一个共同的时间模式:广泛的承诺在防御性行为形成自我强化之前,于早期阶段便已形成。我们的研究结果表明,个体能力出色的智能体并不必然会构成安全的金融系统,这凸显了系统级评估与交互设计是金融AI安全的核心问题。代码可在 https://anonymous.4open.science/r/FinFrail-CF26 获取。
cs.AI / 47 / 2609.30943

LogicTree-RAG: Logic Tree-guided Retrieval-Augmented Generation for Long-form Patent Drafting

LogicTree-RAG:面向长篇专利撰写的逻辑树引导检索增强生成方法
Zhu, Jiaqi, Xing, Naili, Pan, Hexiang, Gao, Haotian, Yin, Jianwei, Xiao, Xiaokui, Ooi, Beng Chin
Abstract
Long-form technical text generation underpins knowledge-intensive workflows, yet remains challenging for large language models (LLMs) due to the need for globally consistent logical structuring and faithful technical reasoning beyond local coherence. Patent drafting is a canonical instance of this challenge, demanding holistic generation of a legally compliant and technically exhaustive document through sustained multi-expert collaboration. Existing approaches often focus on partial section generation or rely on manually crafted outlines, limiting scalable automation in realistic settings. In this work, we propose LogicTree-RAG, a logic tree-guided retrieval-augmented generation framework that induces a hierarchical logic tree as a global organizational backbone to organize and ground technical disclosures, without relying on expert-defined drafting priors. Each node in the logic tree represents a technical element and is constructed through evidence-guided recursive generation. A hybrid traversal mechanism then maps the logic tree into patent sections, enabling controllable and section-balanced generation. Extensive experiments show that LogicTree-RAG consistently improves content quality and language conformity over strong LLM-based baselines and achieves longer structured generation with high token efficiency, demonstrating the effectiveness of logic-centric generation for complex technical document drafting.
Chinese Translation
长篇技术文本生成是知识密集型工作流程的基础,但由于需要全局一致的逻辑结构化以及超越局部连贯性的忠实技术推理,其对大语言模型(LLMs)而言仍然具有挑战性。专利撰写是这一挑战的典型实例,它要求通过持续的多领域专家协作,整体性地生成一份法律合规且技术上详尽的文档。现有方法通常仅关注局部章节生成,或依赖人工编写的提纲,限制了真实场景下可扩展的自动化。在本工作中,我们提出 LogicTree-RAG,一种逻辑树引导的检索增强生成框架。该框架通过归纳出层次化的逻辑树作为全局组织骨架,来组织和支撑技术披露内容,而无需依赖专家定义的撰写先验知识。逻辑树中的每个节点代表一个技术要素,并通过证据引导的递归生成方式构建。随后,一种混合遍历机制将逻辑树映射到专利的各个章节中,实现可控且章节均衡的生成。大量实验表明,与强大的基于 LLM 的基线方法相比,LogicTree-RAG 持续提升了内容质量与语言规范性,并能以较高的 token 效率实现更长的结构化生成,展示了以逻辑为中心的生成方法在复杂技术文档撰写中的有效性。
cs.AI / 48 / 2609.30954

FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models

FTB Graph:多语言语言模型中首词广播器与语言身份头电路的确定与验证
Pillai, Arjun, Hoang, Christian, Laroza, Anjelo
Abstract
Large language models operating in multilingual contexts must resolve target response languages early in generation, yet the causal circuitry governing first-token language identity decisions remains poorly mapped. We present an end-to-end structural circuit analysis across six model architectures spanning four families: GPT-2, BLOOM-560M, Pythia-1B/2.8B, and Qwen2.5-1.5B Base/Instruct. Using Edge Attribution Patching (EAP) with FP16 active clamping, followed by exact activation patching verification with a 2,000-candidate-edge search ceiling, we extract directed acyclic graphs driving first-token language broadcasting. Across the standalone models, we observe deep or mid-to-deep broadcasting hubs, though the evidence is strongest for Pythia-2.8B and BLOOM-560M because GPT-2 and Pythia-1B leave few out-of-graph heads for comparison, while both Qwen2.5-1.5B variants invert the necessity check. Scaling from Pythia-1B to 2.8B expands node participation while maintaining a similar verified edge budget, producing sparser topology. The Qwen2.5-1.5B base and instruct circuits retain 84.7% Jaccard similarity, including the Layer 27 hub, indicating that first-token routing is largely established during pretraining and preserved by instruction tuning. Finally, EAP scores correlate weakly with exact patching deltas across most models, showing that linear gradient approximations can diverge from causal interventions in FP16 and motivating exact-patching verification for reliable circuit discovery.
Chinese Translation
在多语言环境下运行的大语言模型必须在生成早期确定目标响应语言,然而支配首词(first-token)语言身份决策的因果电路仍未被充分绘制。我们对横跨四个模型家族的六种模型架构(GPT-2、BLOOM-560M、Pythia-1B/2.8B 以及 Qwen2.5-1.5B Base/Instruct)进行了端到端的结构化电路分析。我们采用结合 FP16 激活截断(active clamping)的边归因修补(Edge Attribution Patching, EAP)方法,并以 2,000 个候选边搜索上限进行精确激活修补验证,提取出驱动首词语言广播的有向无环图。在各独立模型中,我们观察到深层或中深层的广播枢纽,其中 Pythia-2.8B 和 BLOOM-560M 的证据最强,因为 GPT-2 和 Pythia-1B 留下的图外注意力头较少,难以进行比较,而两个 Qwen2.5-1.5B 变体则在必要性检验中出现反转。从 Pythia-1B 扩展到 2.8B 时,节点参与数量增加,同时保持相近的已验证边数量,产生了更稀疏的拓扑结构。Qwen2.5-1.5B 的 base 与 instruct 电路保留了 84.7% 的 Jaccard 相似度,包括第 27 层的枢纽,这表明首词路由主要在预训练阶段建立,并被指令微调所保留。最后,在大多数模型中,EAP 分数与精确修补差值仅呈弱相关,这表明线性梯度近似在 FP16 下可能偏离因果干预的结果,也凸显了在可靠电路发现中进行精确修补验证的必要性。
cs.AI / 49 / 2609.30967

MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens

MoMHa:面向准确性、安全性与令牌开销的大语言模型测试框架多目标优化
Mukherjee, Subhojyoti, Tanjim, Md Mehrab
Abstract
Most work on improving large language models treats accuracy as the sole objective. We argue that the harness, the Python code surrounding the model that constructs prompts, routes calls, and parses outputs, is a first-class design surface whose quality is inherently multi-objective: an accurate harness that refuses no unsafe request, or that consumes an order of magnitude more tokens, is not a good harness. We present Meta-Harness, a system that casts harness design as search over three per-domain objectives (accuracy, behavioural safety, and token cost) solved by an agentic proposer (Claude Code) with full filesystem access to prior harness source, execution traces, and scoring artifacts. Our central finding is that a singlephase joint-reward proposer (MoMHa) outperforms every alternative, including a two-phase "accuracy then tokens" ablation, scalar-only feedback, and an accuracy-only baseline. We evaluate on seventeen domains: seven synthetic capability suites, seven real-world public benchmarks (HumanEval, MBPP, Spider, FEVER, MMLU-Pro, LawBench, NuminaMath), and three U-SafeBench-derived user-specific safety domains, using a 12-model fleet spanning four families. On the synthetic track MoMHa achieves a joint mean of 0.482 versus 0.198-0.422 for ten baselines, winning $7 / 10$ per-domain columns; on the real-world track it scores 0.461 versus 0.377 for the strongest baseline (DSPy), winning 5/7 columns, demonstrating that harness strategies transfer to unseen benchmarks without retraining on 8 of 12 target models. MoMHa attains the highest measured behavioral safety composite (U-SafeBench, 0.781) and uses 95 fewer tokens per example than the two-phase alternative. We will release all harness code, evaluation infrastructure, and crossmodel logs.
Chinese Translation
现有关于改进大语言模型的研究大多将准确性作为唯一目标。我们认为,测试框架(harness),即围绕模型的Python代码,负责构建提示词、路由调用并解析输出,是一等公民级别的设计层面,其质量本质上是多目标的:一个准确性高但对不安全请求从不拒绝的框架,或一个令牌消耗高出数量级的框架,都不是好的框架。我们提出了Meta-Harness系统,将框架设计转化为对三个领域内目标(准确性、行为安全性和令牌成本)的搜索,该搜索由具备完整文件系统访问权限的智能体提议者(Claude Code)完成,可访问先前的框架源码、执行轨迹和评分产物。我们的核心发现是:单阶段联合奖励提议者(MoMHa)优于所有替代方案,包括两阶段的“先准确性后令牌”消融方案、仅标量反馈方案以及仅准确性基线。我们在十七个领域上进行评估:七个合成能力套件、七个真实世界公开基准(HumanEval、MBPP、Spider、FEVER、MMLU-Pro、LawBench、NuminaMath),以及三个源自U-SafeBench的用户特定安全领域,并使用涵盖四个系列的12个模型的模型群。在合成赛道上,MoMHa的联合平均值为0.482,而十个基线的范围仅为0.198-0.422,并在10个领域列中赢得7个;在真实世界赛道上,其得分为0.461,而最强基线(DSPy)为0.377,赢得7列中的5列,这表明框架策略能够迁移到未见过的基准,且在12个目标模型中有8个无需重新训练。MoMHa取得了最高的实测行为安全综合得分(U-SafeBench,0.781),且每个样例比两阶段替代方案少消耗95个令牌。我们将发布所有框架代码、评估基础设施和跨模型日志。
cs.AI / 50 / 2609.30971

SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents

SciHorizon-eLab:面向科学具身智能体可扩展基准测试的智能体式协议到任务编译器
Qin, Maokai, Qin, Chuan, Zhang, Qi, Liu, Dianyu, Liu, Zirui, Niu, Hongting, Zhou, Yuanchun, Zhu, Hengshu
Abstract
Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile diverse scientific protocols into executable and verifiable embodied tasks at scale. To address this challenge, we introduce SciHorizon-eLab, an agentic protocol-to-task compiler that formulates scientific embodied task construction as a compilation problem. Given a natural-language protocol of scientific experiments, SciHorizon-eLab progressively compiles laboratory protocols into semantic-preserving embodied tasks through semantic grounding, executable task synthesis, and multi-stage simulation-based certification. The system generates semantically grounded environments, executable manipulation programs, and step-level success specifications, while enabling reproducible generation of expert demonstrations and execution traces. Using this pipeline, we further construct \BenchName, a ready-to-use benchmark comprising 300 certified tasks across diverse laboratory operations. It supports HIL task execution, reproducible expert-demonstration generation, and ordered step-level evaluation. Across representative tasks, the strongest policy attains an average success rate of only 49.7%, with further evaluations revealing pronounced weaknesses in human and embodied agent coordination. We publicly release the code, benchmark data, and evaluation toolkit at https://github.com/SciHorizon-elab/SciHorizon-elab.
Chinese Translation
具身智能体(embodied agents)为科学实验自动化提供了一条有前景的途径,但其进展受限于缺乏可靠且系统的评估环境。现有的基于仿真的实验室基准严重依赖人工任务设计,难以系统性地将多样化的科学协议大规模编译为可执行、可验证的具身任务。为应对这一挑战,我们提出了SciHorizon-eLab,这是一个智能体式(agentic)协议到任务编译器,将科学具身任务的构建形式化为一个编译问题。给定科学实验的自然语言协议,SciHorizon-eLab通过语义接地、可执行任务合成以及多阶段基于仿真的认证,逐步将实验室协议编译为语义保持的具身任务。该系统生成语义接地的环境、可执行的操作程序以及步骤级成功规范,同时支持专家演示和执行轨迹的可复现生成。利用该流水线,我们进一步构建了\BenchName,这是一个即用型基准,包含涵盖多种实验室操作的300个已认证任务。它支持HIL任务执行、可复现的专家演示生成以及有序的步骤级评估。在代表性任务上,最强策略的平均成功率仅为49.7%,进一步的评估揭示了人机与具身智能体协调方面的显著弱点。我们已在 https://github.com/SciHorizon-elab/SciHorizon-elab 公开发布了代码、基准数据和评估工具包。
cs.AI / 51 / 2609.30972

Factorized axis convolutional gated recurrent unit with dynamic adaptive pooling for remaining useful life prediction of rolling bearings

面向滚动轴承剩余使用寿命预测的因子化轴向卷积门控循环单元与动态自适应池化方法
Park, Hanbyeol, Choo, Jungho, Bae, Hyerim
Abstract
Convolutional neural networks (CNN) are widely used to predict the remaining useful life (RUL) of rolling bearings from time-frequency representations (TFRs) of vibration signals. However, during degradation, characteristic structures in TFRs align predominantly along the frequency or time axis, making it challenging for conventional CNN isotropic kernels to capture directional structure. Furthermore, global average pooling (GAP) averages across axes, potentially obscuring the locations and concentrations of salient activations. This study introduces a factorized-axis convolutional gated recurrent unit (GRU) that employs multiscale anisotropic convolution and dual-axis convolution block attention module to enhance directional features and highlight salient time-frequency regions. Dynamic adaptive pooling (DAP) adaptively aggregates the time-frequency-axis information from the extracted feature maps, whereas a GRU captures temporal dynamics in the latent representations and Monte Carlo dropout enables predictive uncertainty estimation. Experiments on two public bearing datasets demonstrate that the proposed model outperforms existing RUL prediction methods across operating conditions. Ablation experiments demonstrate that the factorized axis-wise design achieves lower mean errors than convolutional isotropic kernels. DAP yields clear improvements on one dataset while matching GAP on the other, highlighting the importance of anisotropic feature extraction and adaptive feature aggregation for TFR-based RUL prediction.
Chinese Translation
卷积神经网络(CNN)被广泛用于从振动信号的时频表示(TFR)中预测滚动轴承的剩余使用寿命(RUL)。然而,在退化过程中,时频表示中的特征结构主要沿频率轴或时间轴分布,使得传统CNN的各向同性卷积核难以捕捉这种方向性结构。此外,全局平均池化(GAP)在两个轴上进行平均,可能掩盖显著激活的位置和集中程度。本研究提出一种因子化轴向卷积门控循环单元(GRU),其采用多尺度各向异性卷积和双轴卷积块注意力模块来增强方向性特征并突出显著的时频区域。动态自适应池化(DAP)自适应地聚合所提取特征图中的时频轴信息,而门控循环单元(GRU)捕捉潜在表示中的时间动态特性,同时蒙特卡洛Dropout实现了预测不确定性估计。在两个公开轴承数据集上的实验表明,所提出的模型在各种工况下的性能均优于现有的RUL预测方法。消融实验表明,因子化轴向设计的平均误差低于各向同性卷积核。DAP在一个数据集上带来了明显改进,在另一个数据集上与GAP表现相当,凸显了各向异性特征提取和自适应特征聚合对于基于时频表示的RUL预测的重要性。
cs.AI / 52 / 2609.31013

Same Text, Different Numbers: The Divergence of LLM-Based Measures

相同的文本,不同的数字:基于大语言模型度量指标的差异性
Boustanifar, Hamid, Mansouri, Sasan
Abstract
Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, and transcript-level differences common across providers account for only 34% of total score variation. Cross-model disagreement does not predict subsequent analyst or market disagreement, consistent with a substantial model-specific component rather than common ambiguity in the underlying disclosure. Model choice significantly affects downstream inference, with coefficient magnitudes, signs, and statistical significance varying substantially across models. Averaging across providers makes transcript rankings more stable for most constructs, but score levels remain sensitive to the models included in the ensemble. LLM-generated variables should therefore be treated as model-contingent measurements and validated across providers.
Chinese Translation
研究者越来越多地使用生成式大语言模型(LLM)将企业文本转化为实证变量。我们利用十三种文本度量指标(包括情感、管理层表达清晰度、不确定性、回答具体性以及气候风险和政治风险),考察了基于LLM的文本度量在多大程度上对模型选择保持不变性。我们让来自不同厂商的七个LLM对标准普尔500公司财报电话会议记录的上述构念进行评分。结果显示,跨模型的秩相关系数平均仅为0.52,且各厂商共同存在的文本层面差异仅占总分数变异的34%。跨模型分歧并不能预测随后的分析师分歧或市场分歧,这与模型特有的成分占主导、而非底层信息披露中存在共同模糊性的解释相一致。模型选择会显著影响下游推断,系数大小、符号和统计显著性在不同模型之间差异明显。对多个厂商的评分取平均值可使大多数构念的文本排名更加稳定,但分数水平仍对集成中所包含的模型较为敏感。因此,LLM生成的变量应被视为依赖于模型的测量结果,并需在不同厂商的模型间进行验证。
cs.AI / 53 / 2609.31029

Governed Deduction: Policy-Grounded Premise Authorization Beyond Relevance

受治理的演绎:超越相关性的基于策略的前提授权
Shu, Wesley, Lin, Hsi-Ching
Abstract
Reasoning systems usually treat premise use as a question of relevance: if a fact is available and useful, it may be selected for inference. Authorization imposes a different constraint: a premise may be represented and logically usable but not permitted for a particular local transition. We formalize this distinction as Governed Deduction (GD), with a transition-local admission predicate admit(p, tau, S). From an independently produced RBAC-augmented Spider benchmark, we construct 4,461 matched authorization pairs in which the same query premise and policy state support permitted and denied consuming transitions. An initial joint controller reaches 99.19% held-out accuracy, but a transition-only control reaches 100%, exposing a role-name shortcut. After a frozen, label-independent context-local role permutation removes that shortcut, premise/state-only, transition-only, and joint linear controllers all score exactly 50% on 1,856 held-out edges, while a symbolic policy oracle remains at 100%. The result is a controlled negative finding: the benchmark instantiates policy-grounded authorization beyond relevance, but the frozen linear representation does not recover the relation. Matched one-sided controls and leakage audits are therefore essential for evaluating learned policy-sensitive reasoning.
Chinese Translation
推理系统通常将前提的使用视为一个相关性问题:如果某个事实可用且有用,它就可以被选用于推理。而授权施加的是一种不同的约束:某个前提可以被表示并且在逻辑上可用,但可能不被允许用于某个特定的局部转移。我们将这一区别形式化为受治理的演绎(Governed Deduction, GD),其核心是一个转移局部的准入谓词 admit(p, tau, S)。基于一个独立构建的、经RBAC增强的 Spider 基准,我们构造了4,461组匹配的授权对,其中相同的前提与策略状态分别支持“被允许”和“被拒绝”的消耗性转移。一个初始的联合控制器在留出集上达到99.19%的准确率,但仅针对转移的控制却达到100%,这暴露了一个基于角色名称的捷径。在通过冻结的、与标签无关的上下文局部角色置换消除该捷径之后,仅基于前提/状态、仅基于转移以及联合的线性控制器在1,856条留出边上均恰好得到50%的分数,而符号化的策略预言机(oracle)仍保持100%的准确率。这一结果构成一个受控的否定性发现:该基准确实实例化了超越相关性的基于策略的授权,但冻结的线性表示无法恢复这种关系。因此,匹配的单侧对照与泄漏审计对于评估学习到的策略敏感推理至关重要。
cs.AI / 54 / 2609.31054

Cheap, open agents make LLM pollution harder to mitigate

廉价的开源智能体使大语言模型污染更难治理
Rilla, Raluca, Nussberger, Anne-Marie, Mata, Rui, Wulff, Dirk U.
Abstract
Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing multiple response types yielding various detection checks. Fully open agents ran locally without usage fees and performed competitively with commercial alternatives. Open and commercial agents failed different sets of checks, and no single check reliably detected all agents, but open-text responses discriminated best between agents and humans. These findings identify fully open agents as a distinct risk for LLM pollution and support multilayered detection strategies emphasizing open-text analysis.
Chinese Translation
大语言模型(LLM)污染是指合成回复污染了原本旨在捕捉人类行为的数据。迄今为止,高昂的部署成本限制了自主调查智能体所带来的风险。然而,开源权重模型与开源智能体框架的结合可能已消除了这一障碍。我们比较了九种智能体配置的性能与可检测性,这些配置涵盖从完全开源的变体到闭源商业产品。每个智能体自主完成了一份包含多种作答类型、并设置了多种检测检验项的调查问卷。完全开源的智能体可在本地运行且无需使用费用,其表现可与商业替代方案相媲美。开源智能体与商业智能体未能通过不同的检验项集合,且没有任何单一检验项能够可靠地检测出所有智能体,但开放文本作答在区分智能体与人类方面效果最佳。这些发现表明,完全开源的智能体构成了LLM污染的一种独特风险,并支持采用以开放文本分析为重点的多层次检测策略。
cs.AI / 55 / 2609.31056

Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders

消除痕迹:基于对比稀疏自编码器的选择性表征级遗忘
Zehavi, Itai, Jourdan, Fanny, Aivodji, Ulrich
Abstract
Machine unlearning aims to remove targeted information while preserving a model's other abilities. In realistic settings, such as privacy requests under the EU GDPR, the target may be narrow, for example information associated with a single person. Behavioral forgetting alone may be insufficient, motivating interventions directly on internal representations. However, standard mechanistic-interpretability extractors are poorly selective for such targets. We identify an energy bias in reconstruction-based extraction, which favors dominant background structure over low-energy target-specific components. We introduce SCALPEL, a contrastive sparse autoencoder designed to learn more selective forget features. We show theoretically that contrastive training promotes target-selective features and that our selection score controls expected background knowledge perturbation. We validate SCALPEL experimentally on TOFU across Qwen, Llama, and Gemma, where it substantially improves over NMF and standard SAE interventions and is competitive with Gradient Difference and RMU, bridging mechanistic interpretability and fine-grained unlearning.
Chinese Translation
机器遗忘旨在删除特定目标信息,同时保留模型的其他能力。在现实场景中,例如欧盟GDPR下的隐私请求,删除目标可能非常狭窄,例如与某个人相关的信息。仅依靠行为层面的遗忘可能并不足够,这促使我们直接对模型的内部表征进行干预。然而,标准的机制可解释性特征提取方法对此类狭窄目标的选择性较差。我们发现了基于重构的提取方法中存在的一种能量偏差,即其倾向于捕捉占主导地位的背景结构,而忽视低能量的目标特定成分。我们提出了SCALPEL,一种旨在学习更具选择性的遗忘特征的对比稀疏自编码器。我们从理论上证明,对比训练能够促进目标选择性特征的学习,并且我们的选择分数能够控制对背景知识的预期扰动。我们在TOFU数据集上对Qwen、Llama和Gemma模型进行了实验验证,SCALPEL显著优于NMF和标准SAE干预方法,并与Gradient Difference和RMU方法具有竞争力,从而架起了机制可解释性与细粒度遗忘之间的桥梁。
cs.AI / 56 / 2609.31071

Externalized CPDAG Summaries Improve LLM Causal Deduction

外化的CPDAG摘要提升大语言模型的因果推理能力
Sun, Wentao, Nogueira, João Paulo, Verchere, Dominique, Acher, Mathieu, Silva, Alonso
Abstract
Corr2Cause asks whether a causal claim holds in every DAG compatible with observed correlations and conditional independencies. We frame this as latent-object reasoning: the label is defined by a CPDAG query, but free-form chain-of-thought often collapses the Markov-equivalence-class problem into local pattern matching. We propose Structured Thinking, a two-turn pipeline that first externalizes a typed, schema-constrained CPDAG summary and then answers against that graph state. On the Corr2Cause full test, Structured Thinking raises Qwen3.5-27B from $73.0$ to $86.4$ $F_1$(Yes) over a strong PC-instruction baseline in the primary paired run ($+13.4$ pp; McNemar $p=2.4\times 10^{-6}$; bootstrap $95\%$ CI [$+8.4$, $+18.6$]); across three full-ID seeds, the mean gain is $+8.1 \pm 5.3$ pp. A PC-scaffolded two-turn prose control reaches only $67.6$ $F_1$, indicating that a detailed PC scaffold plus a schema-free prose intermediate is not sufficient. The same pattern holds on Qwen3.6-27B, Paraphrase-OOD, and GPT-5.4-mini. Scrambling the emitted CPDAG costs $12.0$ pp $F_1$, and a full-split audit shows close agreement with the reference CPDAG (ID skeleton $F_1$ $0.960$; exact match $75.9\%$). These results support a bounded design principle: externalize the latent object that defines the label, constrain its form, and test whether downstream answers use it.
Chinese Translation
Corr2Cause 任务要求判断一个因果断言是否在所有与观测到的相关性及条件独立性相容的有向无环图(DAG)中均成立。我们将该任务构建为一种潜在对象推理问题:标签由一个 CPDAG 查询定义,但自由形式的思维链(chain-of-thought)往往将马尔可夫等价类问题退化为局部模式匹配。我们提出结构化思维(Structured Thinking),一种两轮流水线:首先外化一个带类型、受模式约束的 CPDAG 摘要,然后基于该图状态进行作答。在 Corr2Cause 完整测试集上,在一次主要的配对实验中,结构化思维将 Qwen3.5-27B 的 $F_1$(Yes) 从 $73.0$ 提升至 $86.4$(相对于强 PC-instruction 基线;提升 $+13.4$ 个百分点;McNemar 检验 $p=2.4 imes 10^{-6}$;bootstrap 95% 置信区间 $[+8.4, +18.6]$);在三个完整分布内(ID)随机种子上,平均增益为 $+8.1 \pm 5.3$ 个百分点。一个使用 PC 脚手架的两轮自然语言散文对照方法仅达到 $67.6$ 的 $F_1$,表明仅有详细的 PC 脚手架加上无模式约束的散文中间表示是不够的。同样的模式在 Qwen3.6-27B、Paraphrase-OOD 以及 GPT-5.4-mini 上均成立。打乱所生成的 CPDAG 会导致 $F_1$ 下降 $12.0$ 个百分点,且完整切分审计显示其与参考 CPDAG 高度一致(分布内骨架 $F_1$ 为 $0.960$;完全匹配率为 $75.9\%$)。这些结果支持一个有边界的设计原则:将定义标签的潜在对象外化,约束其形式,并检验下游答案是否确实利用了它。
cs.AI / 57 / 2609.31076

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

上下抽象阶梯:面向语言智能体的基于代码的技能
Cupiał, Bartłomiej, Tuyls, Jens, Wołczyk, Maciej, Paglieri, Davide, Klissarov, Martin, Eysenbach, Benjamin, Miłoś, Piotr, Narasimhan, Karthik R.
Abstract
Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions. The code handles recurring local decisions, while the language model decides which skills to use and how to combine them. Yet abstractions are leaky, and situations beyond a skill's capabilities may require a return to primitive actions. Motivated by this tradeoff between productivity and flexibility, we systematically study how code-based action abstraction affects the performance, inference cost, and learning of language agents. We study this in NetHack, a challenging, long-horizon game environment, using CodeHack, our library of code-based skills with natural-language descriptions. We use this library to compare agents restricted to primitives with those using semantic skills alone or in combination with primitives. We evaluate these agents in three settings: zero-shot prompting, supervised fine-tuning, and reinforcement learning. Across a broad zero-shot evaluation on NetHack, we find that compared with primitives, skills nearly triple game progression, while reducing inference cost per episode by 86%. Combining skills with primitives retains much of this benefit while preserving a path back down to low-level actions. Finally, in RL, we find that skill-based agents learn significantly faster than agents acting on primitives, achieving a 7.2x larger average gain in dungeon level over the same training budget. These results show that a supplied skill library can improve performance, efficiency, and learning, while retaining primitives provides flexibility when the library is insufficient. We release CodeHack together with training and evaluation code.
Chinese Translation
语言智能体在需要长序列低层动作的环境中难以行动和学习。基于代码的抽象可以让这些智能体调用可复用的技能,而不是反复选择单个动作,从而提高其行动效率。代码负责处理重复性的局部决策,而语言模型则决定使用哪些技能以及如何组合它们。然而抽象存在“泄漏”问题,超出技能能力范围的情形可能需要回退到原始动作。受效率与灵活性之间这一权衡的启发,我们系统地研究了基于代码的动作抽象如何影响语言智能体的性能、推理成本和学习能力。我们在 NetHack——一个具有挑战性的长时程游戏环境——中开展研究,使用我们构建的带有自然语言描述的基于代码的技能库 CodeHack。我们利用该技能库比较了仅限于原始动作的智能体与仅使用语义技能或将语义技能与原始动作结合使用的智能体。我们在三种设置下评估这些智能体:零样本提示、监督微调和强化学习。在 NetHack 上广泛的零样本评估中,我们发现与原始动作相比,技能使游戏进度几乎提升三倍,同时将每个回合的推理成本降低 86%。将技能与原始动作结合可以保留大部分收益,同时保留回退到低层动作的路径。最后,在强化学习中,我们发现基于技能的智能体的学习速度显著快于基于原始动作行动的智能体,在相同训练预算下,地牢层数的平均提升达到 7.2 倍。这些结果表明,提供的技能库可以提升性能、效率和学习能力,而保留原始动作则在技能库不足时提供了灵活性。我们公开发布了 CodeHack 以及训练和评估代码。
cs.AI / 58 / 2609.31078

OmouAI: Argumentative Human-AI Policy Deliberation with Simulated Personas

OmouAI:基于模拟角色的论证式人机政策商议
Vasileiou, Stylianos Loukas, Rago, Antonio, Yeoh, William, Curto, Georgina
Abstract
Debates amongst agents driven by large language models (LLMs) have demonstrated vast potential in various applications, but when these interactions include humans and take place in high-stakes environments, e.g., in public policy deliberations, they are beset with issues such as sycophancy and a lack of faithful explanations. To tackle these issues, we present OmouAI, an interactive and inclusive deliberation system that uses LLMs in combination with computational argumentation, a field which excels in representing and reasoning within debates. OmouAI allows a human user to deliberate policy claims for real-world challenges with simulated personas, e.g., representing stakeholders, domain experts or devil's advocates, towards reducing sycophancy. Each persona generates its own arguments, and the arguments of all parties form a shared argumentation framework. Users can then contest, add and revise arguments, providing crucial human oversight. Then, arguments are evaluated using deterministic argumentative semantics against external goals, such as the UN Sustainable Development Goals, guaranteeing faithful explanations. The advancement or worsening of the goals thus serve as indicators for the policy recommendations.
Chinese Translation
由大语言模型(LLM)驱动的智能体间辩论在各类应用中已展现出巨大潜力,但当这些交互涉及人类并发生在高风险环境中时,例如公共政策商议,便会受到谄媚性(sycophancy)和缺乏忠实解释等问题的困扰。为解决这些问题,我们提出了OmouAI,一个交互式且包容的商议系统,它将大语言模型与计算论辩(computational argumentation)相结合,后者是一个在辩论的表示与推理方面尤为擅长的领域。OmouAI允许人类用户与模拟角色(persona)就现实世界挑战的政策主张进行商议,这些角色可代表利益相关者、领域专家或“魔鬼代言人”,以减少谄媚性。每个角色生成自己的论点,各方论点共同构成一个共享的论辩框架。用户随后可以对论点进行质疑、添加和修改,从而提供关键的人类监督。之后,系统基于外部目标(如联合国可持续发展目标)使用确定性的论辩语义对论点进行评估,以保证解释的忠实性。这些目标的进展或恶化由此作为政策建议的指标。
cs.AI / 59 / 2609.31121

Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning

监视器越狱:在不编码推理的情况下规避思维链监视
Schulz, Julian
Abstract
Chain-of-thought (CoT) monitoring is a promising safety technique for reasoning models, enabling detection of problematic reasoning before models act. A key concern is encoded reasoning, where models hide their true reasoning in ways that monitors and humans cannot interpret. Optimization pressure from CoT monitors during reinforcement learning is considered a likely driver of such behavior. We investigate this by training reasoning models to perform a main task and a side task, while penalizing them when a monitor detects reasoning about the side task. Surprisingly, models learn to evade monitors without encoding their reasoning. Instead, they learn to phrase and format their chains of thought such that monitors fail to flag side task reasoning, while the reasoning remains completely transparent to human readers. We call this phenomenon monitor jailbreaking. We find that monitor jailbreaking arises across different model sizes, monitors, and tasks. Jailbreaks generalize to monitors not seen during training, including both less and more capable monitors, and transfer across different monitor prompts. While jailbreaking strategies appear simple, manually replicating them does not reliably fool monitors. Finally, we show that paraphrasing is an effective defense: paraphrasing a jailbroken CoT allows the same monitor to correctly flag it, while still allowing the model to perform both tasks.
Chinese Translation
思维链(Chain-of-Thought, CoT)监控是一种针对推理模型的有前景的安全技术,能够在模型采取行动之前检测出有问题的推理过程。一个关键的担忧是“编码推理”(encoded reasoning),即模型以监视器和人类都无法解读的方式隐藏其真实推理。强化学习过程中来自 CoT 监视器的优化压力被认为是导致这种行为的一个可能原因。为此,我们通过训练推理模型同时执行一个主任务和一个侧任务(side task),并在监视器检测到与侧任务相关的推理时对其进行惩罚,来研究这一问题。令人惊讶的是,模型学会了在不对其推理进行编码的情况下规避监视器。相反,它们学会了调整思维链的措辞和格式,使监视器无法标记出侧任务推理,而这些推理对人类读者来说仍然是完全透明的。我们将这种现象称为“监视器越狱”(monitor jailbreaking)。我们发现,监视器越狱在不同模型规模、监视器和任务中均会出现。越狱行为能够泛化到训练中未见过的监视器,包括能力较弱和更强的监视器,并能跨不同的监视器提示词进行迁移。尽管越狱策略看起来很简单,但人工复现这些策略并不能可靠地欺骗监视器。最后,我们证明了改写(paraphrasing)是一种有效的防御手段:对被越狱的 CoT 进行改写后,同一个监视器能够正确地将其标记出来,同时模型仍能完成两项任务。
cs.AI / 60 / 2609.31133

AtomWorld-Mem: Memory-Restored World States for Long-Horizon Atomistic Evolution

AtomWorld-Mem:面向长时程原子演化的记忆恢复世界状态
Luo, Tian, Zhang, Ruge, Han, Haozhi, Chen, Yifrng, Zhang, Yunquan, Liu, Yunxin, Cao, Ting, Li, Kun
Abstract
High-fidelity atomistic evolution over long timescales requires more than observing the current crystal configuration. Instantaneous atomistic snapshots are often incomplete: locally similar configurations can correspond to different hidden dynamical contexts, future event preferences, and waiting-time scales. We argue that this snapshot ambiguity makes long-horizon atomistic evolution fundamentally a memory-based world-state restoration problem. To address this, we introduce AtomWorld-Mem, a memory-restored atomistic world model that recovers the latent world state missing from instantaneous crystal snapshots. AtomWorld-Mem treats the evolving alloy as an AtomWorld: spatial encoders write multi-scale atomistic keyframes from dense local topology and sparse long-range defect context, while short-term event memory and long-term structural memory integrate these keyframes across time to restore a future-predictive evolutionary state. The restored state is used to prioritize legal vacancy-mediated events under single-event Kinetic Monte Carlo (KMC) constraints, while event legality, physical execution, and residence-time updates remain governed by the underlying simulator. Empirically, AtomWorld-Mem improves long-horizon atomistic progress under fixed microscopic event budgets while maintaining high-fidelity evolution across energetic, structural, and vacancy-transport observables. It further transfers zero-shot across diverse unseen alloy-temperature AtomWorlds, suggesting that the learned memory-restoration mechanism captures reusable principles of hidden-state inference rather than a system-specific local energy heuristic. These results position memory-restored world-state modeling as a promising route toward efficient, physically grounded, and transferable atomistic evolution.
Chinese Translation
长时间尺度上的高保真原子演化不仅仅需要观察当前的晶体构型。瞬时原子快照往往是不完整的:局部相似的构型可能对应于不同的隐含动力学背景、未来事件偏好以及等待时间尺度。我们认为,这种快照歧义性使得长时程原子演化在本质上成为一个基于记忆的世界状态恢复问题。为此,我们提出了 AtomWorld-Mem,一种通过记忆恢复的原子世界模型,用于恢复瞬时晶体快照中所缺失的潜在世界状态。AtomWorld-Mem 将演化中的合金视为一个 AtomWorld:空间编码器从稠密的局部拓扑和稀疏的长程缺陷上下文中写入多尺度的原子关键帧,而短期事件记忆和长期结构记忆则跨时间整合这些关键帧,以恢复具有未来预测能力的演化状态。恢复后的状态被用于在单事件动力学蒙特卡洛(Kinetic Monte Carlo, KMC)约束下对合法的空位介导事件进行优先级排序,而事件的合法性、物理执行以及驻留时间的更新仍由底层模拟器控制。实验结果表明,在固定的微观事件预算下,AtomWorld-Mem 提升了长时程原子演化进度,同时在能量、结构和空位输运等观测量上保持了高保真演化。此外,该模型能够在多样化的未见合金-温度 AtomWorld 之间实现零样本迁移,这表明所学习到的记忆恢复机制捕获了可复用的隐状态推断原理,而非特定系统的局部能量启发式规则。这些结果将记忆恢复的世界状态建模确立为实现高效、物理可靠且可迁移的原子演化的一条有前景的路径。
cs.AI / 61 / 2609.31136

Toward AI-Augmented Cooperative Engineering Workflows: Requirements and Architecture the European Rover Challenge

迈向AI增强的协同工程工作流:欧洲漫游者挑战赛的需求与架构
Sadik, Ahmed R., Joublin, Frank, Bujny, Mariusz, Ceravola, Antonello, Smith, Joan
Abstract
The growing availability of Artificial Intelligence (AI) tools creates new opportunities to support engineering design processes, yet their current use often remains limited to isolated tasks such as coding, documentation, or information retrieval. Less attention has been given to how AI can support cooperative engineering workflows at the process level, where teams must coordinate requirements, tasks, communication, knowledge transfer, and subsystem integration. This paper investigates this challenge in the context of the European Rover Challenge (ERC), where student teams design and integrate complex rover systems within a single academic cycle under strict time constraints and high subsystem interdependence. We conducted a role adaptive 40 question survey with ERC 2025 teams, yielding 104 responses from 14 teams. The survey examined team structure, knowledge transfer, task management, integration practices, communication patterns, and current AI usage. The results reveal recurring workflow bottlenecks, including limited documentation, unclear requirements, fragmented communication, informal task monitoring, and substantial integration rework. Based on these findings, we derive requirements for AI augmented cooperative engineering work-flows and propose an initial assistant system architecture that connects user facing interfaces, credential management, service selection, specialized AI services, and external engineering tools. The proposed architecture aims to support task clarification, requirement and compliance management, communication summarization, integration risk detection, and continuous knowledge capture. In doing so, the paper contributes empirical requirements and an architectural direction for AI augmented cooperative engineering workflows in hybrid human AI team settings.
Chinese Translation
人工智能(AI)工具的日益普及为支持工程设计过程创造了新的机遇,然而其当前的应用往往局限于孤立的任务,如编程、文档撰写或信息检索。关于AI如何在流程层面支持协同工程工作流的研究较少受到关注,而在此类工作流中,团队必须协调需求、任务、沟通、知识传递与子系统集成。本文以欧洲漫游者挑战赛(European Rover Challenge, ERC)为背景研究这一挑战,参赛学生团队需在严格的时限和高子系统相互依赖的条件下,于单个学年周期内设计并集成复杂的漫游车系统。我们对ERC 2025参赛团队开展了角色自适应的40题问卷调查,共收到来自14支团队的104份回复。调查内容涵盖团队结构、知识传递、任务管理、集成实践、沟通模式以及当前AI的使用情况。结果揭示了反复出现的工作流瓶颈,包括文档不足、需求不明确、沟通碎片化、非正式的任务监控以及大量的集成返工。基于这些发现,我们推导出AI增强协同工程工作流的需求,并提出了一个初步的助手系统架构,该架构连接了面向用户的界面、凭证管理、服务选择、专用AI服务以及外部工程工具。所提出的架构旨在支持任务澄清、需求与合规管理、沟通摘要、集成风险检测以及持续的知识捕获。本文由此为混合人机团队环境下的AI增强协同工程工作流贡献了实证需求与架构方向。
cs.AI / 62 / 2609.31140

Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?

语言推理向量能否增强多模态推理能力?
Wang, Ziyi, Li, Li, Zhou, Aolin, Shen, Yankun, Liu, Chonghan, Lin, Shuxia, Yang, Xu
Abstract
Most Vision-Language Models (VLMs) are built by extending pretrained Large Language Models (LLMs) with visual modules and multimodal alignment. However, this multimodal scaling often degrades the language-side reasoning ability originally encoded in the base LLM. While the base LLM retains usable reasoning after scaling, the aligned VLM itself cannot reliably access this ability. Therefore, recovering the degraded reasoning capability in VLMs would benefit more from seeking help from the base LLM than from the VLM alone. Motivated by this, we propose LIFT (Language-side reasonIng Facilitation and Transfer), a lightweight vector-intervention method that transfers reasoning capability from the base LLM to the VLM without retraining the backbone. LIFT defines Reasoning Vectors as answer-token hidden-state differences between a Reasoner path with an explicit reasoning trace and a Solver path without it, and injects these vectors into language-side activations of the target VLM. LIFT further supports learnable vector adaptation while keeping the VLM backbone frozen. We evaluate LIFT on two VLMs across six reasoning benchmarks, comparing Reasoning Vectors extracted from the base LLM and from the aligned VLM under matched protocols. Results show that LLM-derived vectors consistently outperform VLM-derived vectors, confirming that the base LLM is a more effective source for recovering reasoning. LIFT partially recovers degraded reasoning through lightweight language-side interventions. Further analyses show that Reasoning Vectors influence intermediate reasoning behavior rather than merely altering final answers. The source code will be released soon.
Chinese Translation
大多数视觉-语言模型(Vision-Language Models, VLMs)是通过在预训练大语言模型(Large Language Models, LLMs)基础上扩展视觉模块并进行多模态对齐而构建的。然而,这种多模态扩展往往会损害基础LLM中原本编码的语言侧推理能力。虽然基础LLM在扩展后仍保留可用的推理能力,但对齐后的VLM本身却无法可靠地调用这一能力。因此,恢复VLM中退化的推理能力,与其仅依靠VLM自身,不如寻求基础LLM的帮助。基于此,我们提出了LIFT(Language-side reasonIng Facilitation and Transfer,语言侧推理促进与迁移),这是一种轻量级的向量干预方法,能够在不重新训练主干网络的情况下,将基础LLM的推理能力迁移到VLM中。LIFT将“推理向量”(Reasoning Vectors)定义为带有显式推理轨迹的Reasoner路径与不带的Solver路径在答案词元上的隐状态差异,并将这些向量注入目标VLM的语言侧激活中。LIFT还支持可学习的向量适配,同时保持VLM主干网络冻结。我们在两个VLM上、跨越六个推理基准对LIFT进行了评估,并在匹配协议下比较了从基础LLM和从对齐VLM中提取的推理向量。结果表明,LLM来源的向量始终优于VLM来源的向量,证实了基础LLM是恢复推理能力的更有效来源。LIFT通过轻量级的语言侧干预部分恢复了退化的推理能力。进一步分析表明,推理向量影响的是中间推理行为,而不仅仅是改变最终答案。源代码即将发布。
cs.AI / 63 / 2609.31159

Momentum-Guided Federated Split Distillation for Personalized Temporal Edge Intelligence

面向个性化时序边缘智能的动量引导联邦分割蒸馏框架
Baahmed, Ahmed-Rafik, Dollinger, Jean-François, Brahmia, Amine, Zghal, Mourad
Abstract
We propose a momentum-guided federated split distillation framework for personalized, efficient, and autonomous temporal edge intelligence. We introduce TeRR-SAtt, our novel temporal reservoir student attention design that combines fixed reservoir representations, a lightweight temporal student, and personalized output modules. We also present AMGF, our anticipatory momentum-guided fusion mechanism that clusters clients through learning momentum and derives specialized teacher updates. On real-world smart-building data, TeRR-SAtt reduces edge training latency by 65.50%, inference latency by 44.70%, training memory usage by 18.40%, and inference CPU usage by 33.10% over the considered baselines. At the same time, AMGF improves local learning by up to 35.31% in RMSE compared to global updates.
Chinese Translation
我们提出了一种动量引导的联邦分割蒸馏框架,用于实现个性化、高效且自主的时序边缘智能。我们引入了TeRR-SAtt,这是一种新颖的时序储层学生注意力设计,结合了固定的储层表示、轻量级时序学生模型和个性化输出模块。我们还提出了AMGF,一种前瞻性动量引导融合机制,通过学习动量对客户端进行聚类,并推导出专门的教师更新。在真实智能建筑数据上,与所考虑的基线方法相比,TeRR-SAtt将边缘训练延迟降低了65.50%,推理延迟降低了44.70%,训练内存占用降低了18.40%,推理CPU占用降低了33.10%。同时,与全局更新相比,AMGF将本地学习的RMSE最多提升了35.31%。
cs.AI / 64 / 2609.31167

Neural State Prediction: Obstructing Shortcut Learning in EEG Foundation Models

神经状态预测:阻碍EEG基础模型中的捷径学习
Yu, Kieren, Liu, Ziyang, Huang, Chang, Chen, Jintai, Wu, Kaishun
Abstract
EEG foundation models increasingly use masked prediction to learn from unlabeled recordings, but optimizing this objective does not ensure transferable neural representations. A central challenge is that stable positional cues and local correlations can make masked regions predictable without integrating distributed neural context. To reduce this reliance on low-information prediction paths, we introduce Neural State Prediction (NSP), a latent-predictive framework that constrains both the prediction target and the available context. NSP uses a Target Encoder updated by an exponential moving average (EMA) to define latent supervision. Identity residualization removes additive effects associated with channel identity and relative time from the targets, while topology-separated context excludes their immediate spatial and temporal neighborhood from the visible input. We pretrain NSP on 2.2 million EEG segments from TUEG and evaluate it across 30 downstream datasets spanning clinical diagnosis, sleep staging, emotion recognition, motor imagery, event-related potentials, cognitive-state decoding, and language retrieval. Under full-parameter multi-task fine-tuning on EEG-FM-Bench, NSP achieves 63.94 macro balanced accuracy across 14 datasets, exceeding the strongest evaluated baseline by 2.35 percentage points. Controlled component ablations assess the contribution of each mechanism, while matched context controls and held-out interventions characterize the role of context geometry, signal content, and positional information. Jointly designing latent targets and their context offers a promising direction for EEG foundation models that learn from distributed signal structure.
Chinese Translation
EEG基础模型日益广泛地使用掩码预测从未标注的脑电记录中学习,但优化这一目标并不能确保学习到可迁移的神经表征。一个核心挑战在于:稳定的位置线索和局部相关性可以在不整合分布式神经上下文的情况下使掩码区域变得可预测。为减少对这种低信息量预测路径的依赖,我们提出了神经状态预测(Neural State Prediction, NSP),一个同时约束预测目标和可用上下文的潜在预测框架。NSP使用由指数移动平均(EMA)更新的目标编码器(Target Encoder)来定义潜在监督。身份残差化(Identity residualization)从目标中去除与通道身份和相对时间相关的加性效应,而拓扑分离上下文(topology-separated context)则将目标的空间和时间近邻从可见输入中排除。我们在来自TUEG的220万个EEG片段上对NSP进行预训练,并在涵盖临床诊断、睡眠分期、情绪识别、运动想象、事件相关电位、认知状态解码和语言检索的30个下游数据集上进行评估。在EEG-FM-Bench上进行全参数多任务微调时,NSP在14个数据集上取得63.94的宏平均平衡准确率,超过最强的评估基线2.35个百分点。受控的组件消融实验评估了各机制的作用,而匹配上下文对照和保留干预实验则刻画了上下文几何结构、信号内容和位置信息的作用。联合设计潜在目标及其上下文,为从分布式信号结构中学习的EEG基础模型提供了一个有前景的方向。
cs.AI / 65 / 2609.31176

Semantic Navigation for Issue Localization in Code Repository

面向代码仓库问题定位的语义导航方法
Wei, Yunxiang, Lei, Zhenyu, Li, Jundong
Abstract
Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue. LLM agents approach this task iteratively: they identify a set of potentially relevant locations, inspect the corresponding code, and revise their judgments about these candidates as new evidence is acquired. Existing environments, however, provide limited support for this loop: agents must search for unresolved relation targets, reconstruct entity semantics from raw source code, and revise candidates without evidential basis. To address these limitations, we present SemNav, a framework that leverages deterministic retrieval to seed a broad candidate set and an LLM agent to continually refine that set, thereby combining initial coverage with evidence-guided revision. SemNav supports this process through three key components. A Semantic Navigation Graph resolves program relations on demand through a language server, enabling direct navigation to related entities across files. Issue-conditioned Semantic Cards provide compact, source-grounded interpretations of each entity's role and relevance to the issue. A persistent Candidate Workspace records each candidate together with its evidential basis, enabling grounded verification, revision, and ranking. Across SWE-bench Lite and PLocBench, SemNav outperforms existing baselines, improving File Hit@10 from 68.33\% to 82.67\% with Gemma 4B. Component ablations and trajectory analysis support the complementary roles of all three components, while Semantic Cards reduce working-context load by 48.2\% relative to full-source reading. SemNav further ranks first on all seven evidence-quality metrics on SWE-Explore and improves downstream issue resolution from 44.00\% to 52.33\%.
Chinese Translation
仓库级问题定位旨在识别并排序与解决所报告问题相关的文件和函数。LLM 智能体以迭代方式处理该任务:先确定一组潜在相关的位置,检查相应代码,并随着新证据的获取不断修正对这些候选的判断。然而,现有环境对这一循环的支持有限:智能体必须搜索未解析的关系目标、从原始源代码中重建实体语义,并在缺乏证据依据的情况下修正候选。为解决这些局限,我们提出了 SemNav,一个利用确定性检索生成广泛候选集合、并由 LLM 智能体持续精炼该集合的框架,从而将初始覆盖率与证据引导的修正相结合。SemNav 通过三个关键组件支持这一过程:语义导航图通过语言服务器按需解析程序关系,实现跨文件直接导航到相关实体;问题条件化语义卡片为每个实体的角色及其与问题的相关性提供紧凑的、基于源代码的解释;持久化候选工作区记录每个候选及其证据依据,支持有据可依的验证、修正与排序。在 SWE-bench Lite 和 PLocBench 上,SemNav 优于现有基线,使用 Gemma 4B 时将 File Hit@10 从 68.33% 提升至 82.67%。组件消融和轨迹分析验证了三个组件的互补作用,同时语义卡片相较于完整源代码阅读将工作上下文负载降低 48.2%。SemNav 在 SWE-Explore 的全部七项证据质量指标上均排名第一,并将下游问题解决率从 44.00% 提升至 52.33%。
cs.AI / 66 / 2609.31179

SPO: Discovering Adaptive Large Neighborhood Search Operators via Stackelberg Program Optimization

SPO:基于Stackelberg程序优化的自适应大邻域搜索算子发现
Ke, Xinyi, Li, Kai, Xing, Junliang, Zhang, Yifan, Cheng, Jian
Abstract
Large neighborhood search (LNS) relies critically on destroy and repair operators, whose effectiveness depends on both adaptation to the evolving LNS state and interaction between the two roles. We introduce Stackelberg Program Optimization (SPO), an LLM-based framework for discovering adaptive executable destroy-repair programs. SPO conditions operator decisions on a compact LNS state, allowing state-dependent behavior to emerge through program discovery, and organizes destroy-repair discovery as a Stackelberg interaction over program space that reflects their asymmetric dependency. Role-specific credits evaluate destroy programs as leaders and repair programs as conditional follower responses, guiding a coupled optimization process that combines LLM generator learning with population-based evolutionary search over programs. Experiments on the traveling salesperson problem and capacitated vehicle routing problem show that SPO outperforms strong baselines across a broad range of settings and generalizes beyond the discovery scale to larger instances and benchmark sets. Behavioral analyses further demonstrate state-dependent operator behavior and coupled destroy-repair improvement during discovery.
Chinese Translation
大邻域搜索(Large Neighborhood Search, LNS)在很大程度上依赖于破坏算子与修复算子,其有效性既取决于对不断演化的LNS状态的适应性,也取决于破坏与修复两种角色之间的相互作用。我们提出了Stackelberg程序优化(Stackelberg Program Optimization, SPO),这是一个基于大语言模型(LLM)的框架,用于发现自适应的可执行破坏-修复程序。SPO以紧凑的LNS状态为条件来决定算子行为,使依赖状态的行为能够通过程序发现自然涌现,并将破坏-修复算子的发现组织为程序空间上的Stackelberg交互,以体现二者之间的非对称依赖关系。针对特定角色的信用评估将破坏程序视为领导者、修复程序视为条件性跟随者响应,从而引导一个将LLM生成器学习与基于种群的程序进化搜索相结合的耦合优化过程。在旅行商问题和带容量约束车辆路径问题上的实验表明,SPO在广泛的问题设置下均优于强基线方法,并能超越发现的规模,泛化到更大的实例和基准测试集。行为分析进一步证明了算子的状态依赖行为以及发现过程中破坏-修复的耦合改进。
cs.AI / 67 / 2609.31184

Accounting for Bias Enables Sustainable LLM Evaluation

考虑偏差因素实现可持续的大语言模型评估
Katoch, Harshita, Selby, David Antony, Großmann, Gerrit, Vollmer, Sebastian
Abstract
LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting this model costs negligible compute relative to a single round of LLM inference, bias correction is not only more statistically rigorous but also a more sustainable approach to trustworthy evaluation.
Chinese Translation
“以大语言模型作为评判者”(LLM-as-a-judge)已成为可扩展的主观评估的事实标准,然而当前的排行榜通过运行越来越多的比较来补偿系统性测量偏差,这种方法在统计上并不合理,且在计算上极为浪费。其根本原因在于测量模型的不完备:将大语言模型评判者视为中立且可互换的工具,忽视了已被记录在案的各类偏差,如位置偏差(position bias)、冗长偏差(verbosity bias)、评判者严格度差异以及自我增强偏差,而这些偏差无法通过增加数据量来消除。我们提出了一个统一的潜变量框架,在对成对比较数据和序数数据进行联合建模的同时,显式地校正这些混杂因素,从而能够以显著更少的比较次数恢复可靠的排名。由于拟合该模型所需的计算量相对于单轮大语言模型推理而言可以忽略不计,因此偏差校正不仅在统计上更为严谨,也是一种实现可信评估的更可持续的方法。
cs.AI / 68 / 2609.31186

Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation

递归自我改进人工智能的演化安全性:分类体系、风险发现与评估
Gong, Chang, Bi, Jingping, Yao, Di, Liang, Xinjian, Xiang, Chao, Guo, Ruijie
Abstract
Artificial intelligence is advancing rapidly, with increasingly capable systems taking larger roles in reasoning, decision-making, scientific discovery, and autonomous development. As AI begins to participate in its own improvement, from model training and experience accumulation to agent evolution and automated AI development, the prospect of recursive self-improvement (RSI) is becoming increasingly relevant. This transition raises a fundamental safety question: how can safety be maintained when the system, its accumulated experience, and even the process producing its successors continue to change? We introduce Evolutionary Safety as a perspective for studying safety under persistent and recursive self-improvement. It concerns not only whether an AI system is safe at a particular moment, but how safety properties change, persist, accumulate, and propagate throughout evolution. We characterize recurring manifestations, including intent drift, error accumulation, experience contamination, safety-property erosion, evaluator drift, and risk propagation. We then develop a taxonomy spanning persistent agent state, model state, evaluation and environmental feedback, computational substrate, and meta-level update mechanisms. Building on this taxonomy, we examine how evolutionary risks can be discovered and evaluated across states, updates, trajectories, and lineages, and derive governance principles for modification, selection, authorization, provenance, and recovery. Finally, we outline open problems toward maintaining safety guarantees as AI systems become increasingly persistent, adaptive, and recursively self-improving. Project resources and proposed evaluation systems are available at https://chaunceykung.github.io/evolutionary-safety-rsi.
Chinese Translation
人工智能正在迅速发展,能力日益强大的系统在推理、决策、科学发现和自主开发中扮演着越来越重要的角色。随着AI开始参与自身的改进——从模型训练、经验积累,到智能体演化和自动化AI开发——递归自我改进(Recursive Self-Improvement, RSI)的前景正变得愈发现实。这一转变引发了一个根本性的安全问题:当系统本身、其积累的经验,甚至产生其后继系统的过程都在持续变化时,如何维持安全性?我们提出了“演化安全性”(Evolutionary Safety)这一视角,用于研究持续且递归的自我改进下的安全问题。它关注的不仅是AI系统在某一特定时刻是否安全,而是安全属性如何在整个演化过程中发生改变、延续、积累和传播。我们刻画了反复出现的安全问题表现形式,包括意图漂移(intent drift)、错误累积、经验污染、安全属性侵蚀、评估器漂移(evaluator drift)以及风险传播。随后,我们构建了一个涵盖持久智能体状态、模型状态、评估与环境反馈、计算基底以及元层面更新机制的分类体系。基于该分类体系,我们考察了如何在状态、更新、轨迹和谱系层面发现和评估演化风险,并推导出关于修改、选择、授权、溯源和恢复的治理原则。最后,我们概述了随着AI系统日益持久化、自适应化并具备递归自我改进能力,维持安全性保障所面临的开放性问题。项目资源与所提出的评估系统见 https://chaunceykung.github.io/evolutionary-safety-rsi。
cs.AI / 69 / 2609.31201

Samples, Sources, Space: Decomposing Data Scale in Spatially Structured Representation Learning of Human Brain Microarchitecture

样本、来源、空间:解构人脑微结构空间结构化表示学习中的数据规模
Schiffer, Christian, Bode, Mathis, Lippert, Thomas, Amunts, Katrin, Dickscheid, Timo
Abstract
Scaling studies typically represent training data by a single count of samples. For hierarchically and spatially structured data, however, the same number of samples can be drawn from few or many sources and distributed differently across the underlying domain. We therefore study data scaling as an allocation problem, separating unique sample count, source diversity, and spatial coverage. We study this decomposition in microscopic whole-brain histology, where a source is an individual brain, and a sample is an image patch at a specific spatial location. Across 93 controlled pretraining runs of a contrastive model that uses spatial proximity for supervision, we vary data allocation, compute, and model capacity over 11.6 million spatially anchored image patches from 21 human brains. Performance improves with more unique samples, broader spatial coverage, additional compute, and larger model capacity. At fixed sample count, distributing samples across one to 18 subjects produces no detectable improvement, even though representations generalize substantially better to subjects encountered during pretraining. Inter-subject variation therefore strongly affects generalization, but additional subjects provide no benefit when a fixed sample budget is distributed across more sources. These results establish sample count, source diversity, and spatial coverage as distinct axes of data scaling in spatially structured representation learning.
Chinese Translation
规模研究(scaling studies)通常仅用单一的样本数量来刻画训练数据。然而,对于具有层级结构和空间结构的数据而言,相同数量的样本可以从少量或大量来源中获取,并以不同的方式分布在底层域上。因此,我们将数据规模扩展视为一个分配问题,将其分解为独特样本数量、来源多样性和空间覆盖三个维度。我们在显微全脑组织学数据中研究这一分解,其中来源为个体大脑,样本为特定空间位置上的图像块。在93次受控预训练运行中——该预训练使用一种以空间邻近性作为监督信号的对比学习模型——我们在来自21个人脑的1160万个空间锚定图像块上改变数据分配、计算量和模型容量。结果表明,性能随着独特样本数量增加、空间覆盖范围扩大、计算量增加以及模型容量增大而提升。在固定样本数量的情况下,将样本分布在一到18个受试者之间并未产生可检测的性能提升,尽管表示对预训练中遇到过的受试者的泛化能力显著更好。因此,受试者间差异强烈影响泛化能力,但当固定的样本预算分散到更多来源时,增加受试者数量并无额外收益。这些结果确立了样本数量、来源多样性和空间覆盖是空间结构化表示学习中数据规模扩展的不同维度。
cs.AI / 70 / 2609.31214

Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution

我们估计的是何种影响?反事实设定在数据归因中的作用
Li, Zhe, Zhao, Wei, Zhang, Peixin, Sun, Jun
Abstract
Estimating the influence of training examples on model behavior is essential for data debugging, valuation, and attribution. Existing influence estimators often produce incompatible rankings, which are commonly ascribed to approximation error. We argue that a more fundamental source of disagreement is specification mismatch: influence depends on the behavior being attributed, the intervention applied to each training example, and the counterfactual training process that maps the intervention to a model response. These choices are especially important when the target behavior requires a tractable surrogate, such as query loss, a logit, or a margin. We formalize influence as a counterfactual estimand, distinguish specification mismatch across estimands from approximation error in estimating a fixed estimand, and organize representative estimators by their implied specifications. We further derive a local decomposition that exposes how behavior signals, training signals, and counterfactual parameter responses interact. Controlled experiments show that exact estimands under different specifications can induce different rankings, whereas approximation error grows as perturbations move farther from their linearization points. Experiments on noisy label detection and LLM attribution show that specification choices significantly affect attribution quality, especially for the choice of behavior surrogate. Behavior-aligned specifications can identify target-specific training examples obscured by default loss-based or similarity-based specifications. These results establish specification analysis as a necessary first step for interpreting and comparing data influence estimators.
Chinese Translation
估计训练样本对模型行为的影响对于数据调试、估值和归因至关重要。现有的影响估计器常常产生互不兼容的排序,这通常被归因于近似误差。我们认为,分歧的一个更根本来源是设定不匹配:影响取决于被归因的行为、施加于每个训练样本的干预,以及将干预映射到模型响应的反事实训练过程。当目标行为需要一个可处理的替代量(surrogate),例如查询损失、logit 或间隔(margin)时,这些选择尤为重要。我们将影响形式化为一个反事实估计量(estimand),区分了不同估计量之间的设定不匹配与估计同一固定估计量时的近似误差,并根据隐含的设定对代表性估计器进行分类整理。我们进一步推导出一个局部分解,揭示了行为信号、训练信号与反事实参数响应之间的交互方式。受控实验表明,不同设定下的精确估计量可以导致不同的排序,而近似误差则随着扰动偏离其线性化点的距离增大而增长。在噪声标签检测和大语言模型(LLM)归因上的实验表明,设定的选择会显著影响归因质量,尤其是行为替代量的选择。与行为对齐的设定能够识别出被默认的基于损失或基于相似度的设定所掩盖的针对特定目标的训练样本。这些结果确立了设定分析作为解释和比较数据影响估计器的必要第一步。
cs.AI / 71 / 2609.31215

DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration

DIAL:具有自适应人类偏好校准的位置去偏LLM评判器
Cai, Zesheng, Fan, Yingqi, Chen, Sichang, Du, Jin-Hong
Abstract
Large language models (LLMs) as a judge enable scalable evaluation, but their judgments can be sensitive to response order and, even after removing such position effects, can still diverge systematically from human preferences.We introduce DIAL, a unified framework that combines abundant LLM comparisons with limited human comparisons to separate judge-specific position effects, learn shared structure in position-debiased LLM preferences, and adaptively calibrate that structure toward the human preference target. Theoretically, we study three aspects of DIAL: (i) identification of latent LLM preferences, position effects, and human calibration; (ii) adaptive estimation that balances LLM anchoring against limited human evidence; and (iii) fixed-weight uncertainty quantification for the calibrated human preference. Empirically, we evaluate position debiasing and human alignment separately in controlled simulations and on three human-preference benchmarks, showing that DIAL remains robust to unbalanced response order, achieves strong human-aligned rankings with limited labels, and adapts toward human evidence when LLM information is imperfect. Our real-data study collects over 410K judgments from 21 LLM judges in both display orders, providing a resource for future studies of LLM-judge bias, heterogeneity, and human alignment.
Chinese Translation
大语言模型(LLM)作为评判器能够实现可扩展的评估,但其判断可能对回答顺序敏感,并且即使消除了这种位置效应,仍可能系统性地偏离人类偏好。我们提出了DIAL,一个统一框架,将大量的LLM比较与有限的人类比较相结合,以分离评判器特有的位置效应,学习位置去偏的LLM偏好中的共享结构,并将该结构自适应地校准至人类偏好目标。在理论上,我们从三个方面研究了DIAL:(i)潜在LLM偏好、位置效应和人类校准的可识别性;(ii)在LLM锚定与有限人类证据之间取得平衡的自适应估计;以及(iii)校准后人类偏好的固定权重不确定性量化。在实证方面,我们在受控模拟和三个人类偏好基准上分别评估了位置去偏和人类对齐效果,结果表明DIAL对不平衡的回答顺序保持稳健,在有限标签下实现了强人类对齐的排序,并在LLM信息不完善时能够自适应地转向人类证据。我们的真实数据研究收集了来自21个LLM评判器在两种展示顺序下超过41万条判断,为未来关于LLM评判器偏差、异质性和人类对齐的研究提供了资源。
cs.AI / 72 / 2609.31235

Purin: A Biology-inspired Mechanism for Artificial Neural Networks

Purin:一种面向人工神经网络的仿生机制
Liu, Zishu, Luo, Chunbo, Grecos, Christos
Abstract
Artificial neural networks (ANNs) usually represent neural transmission with fixed trainable weights during a training batch, which omits short-term changes in synaptic efficacy. In addition, the discrete time-step simulation requires additional temporal processing that many conventional ANN architectures do not use. To overcome these challenges, we propose Purin, a biology-inspired and ANN-compatible mechanism, that introduces synaptic efficacy modulation into conventional convolutional neural networks. Purin uses a time-interval-based abstraction for neural activities, which allows Purin to introduce short- and long-term synaptic efficacy changes without using discrete time-steps. Purin introduces a bounded factor to represent temporary synaptic efficacy changes, together with two weight matrices that represent input-side and output-side efficacy. The weight matrices are updated by backpropagation and interpreted as the long-term synaptic efficacy changes. Experimental results show that after removing the confounding factors in the AlexNet, VGG11, and GoogLeNet architectures, Purin improves the classification accuracies in all three models across the evaluated datasets.
Chinese Translation
人工神经网络(ANN)通常在训练批次期间使用固定的可训练权重来表示神经传递,这忽略了突触效能的短期变化。此外,离散时间步仿真需要额外的时序处理,而许多传统ANN架构并未加以利用。为克服这些挑战,我们提出了Purin,一种受生物学启发且与传统卷积神经网络兼容的机制,将突触效能调制引入传统卷积神经网络。Purin采用基于时间间隔的神经活动抽象,使其无需使用离散时间步即可引入短期和长期突触效能变化。Purin引入一个有界因子来表示临时的突触效能变化,并结合两个权重矩阵分别表示输入侧和输出侧效能。权重矩阵通过反向传播进行更新,并被解释为长期突触效能变化。实验结果表明,在排除AlexNet、VGG11和GoogLeNet架构中的混淆因素后,Purin在所有评估数据集上均提升了这三种模型的分类准确率。
cs.AI / 73 / 2609.31281

MA-WAM: Multi-Agent World-Action Model for Test-Time Planning

MA-WAM:面向测试时规划的多智能体世界-动作模型
Zou, Guowei, Wang, Haitao, Wang, Guoxin, Zhang, Beiwen, Chen, Zhiquan, Wang, Guojie, Wu, Hejun
Abstract
Multi-agent cooperative tasks require different agents to execute a joint action simultaneously, and each agent's action affects both the observations and responses of the other agents. Hence, a world model is needed to predict the team return resulting from the joint actions of all agents. A naive extension directly applies a single-agent world model to each agent's action when predicting the team return step by step. However, such an extension fails to capture the dependencies among the simultaneous actions of multiple agents. We propose Multi-Agent World-Action Model (MA-WAM), a test-time planning framework that enables a frozen multi-agent flow policy to evaluate futures of candidate joint actions. To our knowledge, MA-WAM is the first test-time world-model planner for multi-agent flow policies. MA-WAM predicts the consequences of each joint action according to cross-agent dependencies and enables efficient candidate scoring. Across 30 offline multi-agent reinforcement learning (MARL) settings on MAMuJoCo, SMAC, and MPE, MA-WAM achieves mean relative gains of 22.0% over direct execution and 25.6% over uniform action selection. Under the standard evaluation protocol on an A100 GPU, MA-WAM adds 12.1 ms, accounting for 2.5% of the measured generation-and-scoring time.
Chinese Translation
多智能体协作任务要求不同智能体同时执行联合动作,且每个智能体的动作会影响其他智能体的观测与响应。因此,需要一个世界模型来预测所有智能体联合动作所产生的团队回报。一种简单的扩展方法是逐步预测团队回报时,直接将单智能体世界模型应用于每个智能体的动作。然而,这种扩展无法捕捉多个智能体同时执行的动作之间的依赖关系。我们提出多智能体世界-动作模型(Multi-Agent World-Action Model, MA-WAM),这是一个测试时规划框架,使冻结的多智能体流策略能够评估候选联合动作的未来结果。据我们所知,MA-WAM 是首个面向多智能体流策略的测试时世界模型规划器。MA-WAM 依据跨智能体依赖关系预测每个联合动作的后果,并实现高效的候选动作评分。在 MAMuJoCo、SMAC 和 MPE 上的 30 个离线多智能体强化学习(MARL)场景中,MA-WAM 相较直接执行平均取得 22.0% 的相对提升,相较均匀动作选择取得 25.6% 的提升。在 A100 GPU 的标准评估协议下,MA-WAM 仅增加 12.1 毫秒的开销,占实测生成与评分时间的 2.5%。
cs.AI / 74 / 2609.31286

G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

G2MAF:面向多智能体流策略的测试时梯度引导方法
Zou, Guowei, Wang, Haitao, Wang, Guoxin, Chen, Zhiquan, Zhang, Beiwen, Wang, Guojie, Wu, Hejun
Abstract
Offline multi-agent reinforcement learning (MARL) learns cooperative policies from fixed datasets without further environment interaction and a learned policy is frozen at deployment. Such a frozen policy typically proposes a single joint action and executes it directly at deployment time. However, this one-shot deployment often commits to a suboptimal proposal, even when better nearby alternatives remain consistent with the behavior data. To address this issue, we propose Gradient Guided Multi Agent Flow (G2MAF), a refinement framework for optimizing joint policies at test-time. G2MAF applies one globally normalized, projected critic gradient to guide and coordinate all agents' corrections while keeping the action both feasible and close to the frozen policy proposal. Across 24 MPE and SMAC settings, its canonical variant improves 20 frozen settings, with mean relative gains of 9.2% on MPE and 8.9% on SMAC, with model inference latency increased by about 6% only.
Chinese Translation
离线多智能体强化学习(MARL)从固定数据集中学习协作策略,无需与环境进一步交互,且学习到的策略在部署时被冻结。这种冻结策略在部署时通常只提出单个联合动作并直接执行。然而,这种一次性部署往往提交一个次优的动作提议,即使存在与行为数据保持一致的更优的邻近替代动作。为解决这一问题,我们提出梯度引导多智能体流方法(Gradient Guided Multi Agent Flow,G2MAF),这是一个在测试时优化联合策略的精炼框架。G2MAF应用一个全局归一化、投影后的评论家(critic)梯度来引导并协调所有智能体的修正,同时保持动作既可行又接近冻结策略的提议。在24个MPE和SMAC设置上,其标准变体改进了20个冻结设置,在MPE上的平均相对提升为9.2%,在SMAC上为8.9%,而模型推理延迟仅增加约6%。
cs.AI / 75 / 2609.31341

The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models

信息抽取流程的选择取决于文档:小型本地模型的精度-能耗权衡
Walser, Christoph, Argerich, Mauricio Fadel, Fürst, Jonathan
Abstract
Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small ($\le 8\mathrm{B}$ parameter) text-only and vision--language models, evaluated on both accuracy and energy over a design space spanning input representation, model family, and inference configuration. Benchmarking on the near-plain-text Kleister-NDA contracts and the layout-rich VRDU forms, we find that batching is the dominant energy lever, cutting energy per page by 38-85% at no cost in accuracy, while FP8 quantization saves 27-32% when requests are served one at a time but less than 1mWh per page (9-19%) once batching is applied. Preprocessing dominates what remains: neural OCR costs $17\times$ more energy per page than classical OCR and never reaches the Pareto frontier. Which representation wins flips with the type of document: vision--language models on layout-rich documents and small text-only models with a cheap parser on near-plain text, where they are both more accurate and cheaper than any vision--language configuration. Our work yields concrete guidelines for energy-efficient, privacy-compliant local information extraction.
Chinese Translation
信息抽取流程应处理页面图像还是解析后的文本,取决于具体的文档,且答案会随文档版式谱系而变化。我们在一个排除(封闭式)云服务的约束条件下研究这一权衡:在本地服务器上由小型(参数量≤8B)纯文本模型和视觉-语言模型处理隐私敏感文档,并在涵盖输入表示、模型家族和推理配置的设计空间中同时评估精度与能耗。在接近纯文本的Kleister-NDA合同数据集和版式丰富的VRDU表单数据集上进行基准测试后,我们发现批处理(batching)是最主要的能耗调控手段,可在不损失精度的前提下将每页能耗降低38-85%;而FP8量化在逐个处理请求时可节省27-32%的能耗,但在应用批处理后每页节省不足1mWh(9-19%)。预处理主导了剩余的能耗:神经OCR每页能耗是传统OCR的17倍,且始终无法达到Pareto前沿。哪种输入表示更优随文档类型而变化:对于版式丰富的文档,视觉-语言模型更优;而对于接近纯文本的文档,采用廉价解析器的小型纯文本模型更优,此时它们在精度和成本上均优于任何视觉-语言配置。本工作为节能且符合隐私合规要求的本地信息抽取提供了具体指南。
cs.AI / 76 / 2609.31354

Mutable Transcripts: Mitigating Context Pollution through Editable Conversation State

可变转录:通过可编辑的对话状态缓解上下文污染
Barry, Dan, Hines, Andrew
Abstract
Contemporary large language model (LLM) chat systems treat conversation history as an immutable sequence of turns that defines the model's working context. However, user intent in real interactions is not static: it evolves through correction, refinement, and shifting constraints. This mismatch between dynamic intent and static transcripts can result in context pollution, where outdated or irrelevant information persists and continues to influence subsequent responses. We introduce mutable transcripts, a new interaction paradigm that enables users to revise prior turns through natural language edit requests, allowing the conversation history itself to be updated rather than appended. This reframes the transcript from a passive record into an editable representation of conversational state. We present a working prototype that integrates transcript-level revision into a standard chat interface and evaluate its feasibility through a controlled user study (n=17) and an illustrative transcript analysis of representative interaction scenarios. Participants significantly preferred mutable transcripts over standard chat across measures of clarity, confidence, and ease of use, with reduced intent to restart conversations. Transcript analysis of representative user study conversations shows that mutable transcripts can reduce conversation length and eliminate obsolete retained context. These findings provide initial evidence that user-driven revision of conversational history can improve interaction quality and help maintain a more current representation of user intent. The source code and prototype can be accessed at https://github.com/QxLabIreland/ReChat
Chinese Translation
当代大语言模型(LLM)聊天系统将对话历史视为定义模型工作上下文的不可变轮次序列。然而,真实交互中的用户意图并非静态:它会通过纠正、细化和约束条件的变化而不断演变。动态意图与静态转录之间的这种不匹配可能导致上下文污染(context pollution),即过时或无关的信息持续存在并继续影响后续响应。我们提出可变转录(mutable transcripts),这是一种新的交互范式,允许用户通过自然语言编辑请求修改先前的对话轮次,使对话历史本身可以被更新而非仅仅追加。这将转录从被动记录重新定义为对话状态的可编辑表示。我们实现了一个将转录级修订集成到标准聊天界面中的工作原型,并通过一项受控用户研究(n=17)以及对代表性交互场景的说明性转录分析评估了其可行性。参与者在清晰度、信心和易用性等各项指标上均显著偏好可变转录而非标准聊天,且重新开始对话的意愿有所降低。对用户研究中代表性对话的转录分析表明,可变转录能够缩短对话长度并消除保留的过时上下文。这些发现提供了初步证据,表明由用户驱动的对话历史修订可以改善交互质量,并有助于维持对用户意图更为实时的表示。源代码和原型可通过 https://github.com/QxLabIreland/ReChat 访问。
cs.AI / 77 / 2609.31360

Programs-of-Layers in LLMs through the Lens of Cortical Areas

从皮层区域的视角审视大语言模型中的“层程序”
Westerhoff, Justus, Olbrich, Stephan, Oraby, Hatem, Larkum, Matthew Evan, Gers, Felix Alexander
Abstract
Inference in LLMs is conventionally a fixed-depth, fixed-order forward pass through every layer, regardless of how difficult the input is. The human brain does not work this way: using the thalamus as a central hub, it routes information flexibly to all regions of the cortex according to demand. Li et al. (2026) recently showed, with a system they call program-of-layers (PoLar), that transformers can be given an analogous flexibility if their layers are treated as a library of functions rather than a fixed sequence. Performance improves over the standard forward pass when each input is dynamically routed through an adaptive sequence of skipped or repeated contiguous layer blocks. We reconstructed PoLar's diagnostic MCTS in more detail than the original paper and applied it across 5 models. We reproduced several of PoLar's findings: skipping outperformed the standard pass, repeating outperformed skipping, and combining both outperformed either alone. Shorter programs sufficed for easier questions, while harder questions required more layer repeats. However, we failed to replicate the main claim regarding their learned router for single-shot inference: its top-ranked prediction consistently collapsed back to the standard pass, even though its top-k predicted programs, taken together, did show a real accuracy gain. Beyond reproduction, we find that a small number of generic programs are enough to solve most of the questions. We also provide a much deeper analysis of these programs' structure and robustness: for example, we found that programs that correct errors are highly brittle: undoing even a single edit inside a program typically breaks the correction. Connecting this to the brain's routing mechanisms, PoLar mirrors principles of thalamo-cortical coordination between cortical-area-like transformer layers. We publicly release the code at https://datexis.github.io/RE-PoLar/
Chinese Translation
大语言模型(LLM)的推理通常是一个固定深度、固定顺序的前向传播过程,即依次经过每一层,而无论输入的难度如何。人脑的工作方式并非如此:人脑以丘脑作为中央枢纽,根据需求灵活地将信息路由至皮层的各个区域。Li 等人(2026)最近通过一个名为“层程序”(program-of-layers,PoLar)的系统表明,如果将 Transformer 的各层视为一个函数库而非固定序列,则可以赋予其类似人脑的灵活性。通过将每个输入动态路由至由跳过或重复的连续层块组成的自适应序列,其性能优于标准前向传播。我们对 PoLar 的诊断性 MCTS 进行了比原始论文更详细的重构,并将其应用于 5 个模型。我们复现了 PoLar 的多项发现:跳过优于标准前向传播,重复优于跳过,而将两者结合则优于单独使用任一策略。较简单的题目只需较短的程序,而较难的题目则需要更多次的层重复。然而,我们未能复现关于其学习型路由器用于单次推理的主要结论:其排名最高的预测始终退回到标准前向传播,尽管其 top-k 预测程序综合起来确实带来了实际的准确率提升。在复现之外,我们发现少量通用程序就足以解决大部分问题。我们还对这些程序的结构与鲁棒性进行了更深入的分析:例如,我们发现用于纠错的程序非常脆弱——即使撤销程序中的单个编辑操作,通常也会破坏纠错效果。结合人脑的路由机制来看,PoLar 反映了类似皮层区域的 Transformer 层之间丘脑-皮层协调的原理。相关代码已在 https://datexis.github.io/RE-PoLar/ 公开发布。
cs.AI / 78 / 2609.31381

Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection

已完成对掩盖了有上限的失败:选择性上下文投影的ReVerPi案例研究
Zhang, Guangzhe
Abstract
Context projection replaces older tool observations with compact, addressable excerpts, reducing repeated input while potentially adding evidence-retrieval turns. We study this trade-off in ReVerPi, a Pi extension with archived observations and matched full/projected continuations. In an 86-run source-reading campaign with 641 model requests, the 15 completed pairs show identical success: 12/15 per arm. Twelve further boundary runs stop, with the runner suppressing the companion whenever the first arm fails to complete. Restoring all 27 boundary runs bounds projected-minus-full success between $-$9 and +1 tasks. One omitted, selector-chosen projected continuation successfully retrieves archive text yet exhausts twelve requests; its full counterpart answers in three. The eleven jointly correct pairs form a fully observed success stratum within this recorded frame: projection reduces aggregate logical tokens by 25%, while increasing the median pair's tokens by 29% and total suffix requests from 35 to 55. Separating fitting from evaluation changes the selector's apparent tie: outside its four fitting pairs, it incurs one extra failure and 8.6% more logical tokens over thirteen comparable runs. This methodological case study connects stopping rules, known bounded failures, unexecuted companions, and resource aggregation. Its findings concern the recorded campaign, rather than population noninferiority or superiority over unrestricted Pi. Evaluations should retain every intervention boundary, execute both allocated arms independently of the first arm's completion, and report completion alongside interaction and token expenditure.
Chinese Translation
上下文投影(context projection)用紧凑、可寻址的摘录替换较旧的工具观测结果,在减少重复输入的同时可能增加证据检索的轮次。我们在ReVerPi中研究这一权衡,ReVerPi是一个具有观测结果归档和匹配的全量/投影延续的Pi扩展。在一项包含641次模型请求的86次运行的源代码阅读实验中,15个已完成对显示出相同的成功率:每组12/15。另有12次边界运行中途停止,当第一个分支未能完成时,运行器会抑制其配对分支。恢复全部27次边界运行后,投影减全量的成功率被限定在-9至+1个任务之间。一个被遗漏的、由选择器选中的投影延续成功检索到归档文本,却耗尽了十二次请求;而其对应的全量延续仅用三次请求即给出答案。十一个共同正确的对构成了该记录框架内一个完整观测的成功层:投影使聚合逻辑token减少25%,却使对的中位数token增加29%,后缀请求总数从35增加到55。将拟合与评估分开会改变选择器表面上的平局:在其四个拟合对之外,选择器在十三次可比运行中多产生一次失败并增加8.6%的逻辑token。这一方法论案例研究关联了停止规则、已知的有上限失败、未执行的配对分支以及资源聚合问题。其结论仅针对该记录的实验,而非相对于无限制Pi的总体非劣性或优越性。评估应保留每个干预边界、独立于第一个分支的完成情况而执行两个分配分支,并在报告完成情况的同时报告交互次数与token消耗。
cs.AI / 79 / 2609.31430

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

压缩所见,而非所言:面向潜观测软件工程智能体的锚定上下文蒸馏方法
Zou, Zhensheng, Wang, Guoqing, Hao, Dan
Abstract
Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents to soft-token representations can compromise their original behavior. To reduce context while preserving action-critical information and agent behavior, we combine Latent Observations, Hard Actions (LOHA), a context layout that separates compressed history from text needed for exact reference, with Anchored Context Distillation (ACD), a training method that enables latent reading while constraining behavioral drift. LOHA compresses older tool observations into soft tokens while retaining the agent's own turns and the last K observations in text, providing compact access to historical information and exact access to recent content. To enable the agent to use this representation, ACD distills the base model's full-text predictions into the latent view while anchoring its behavior on plain-text inputs to the same base model. On SWE-bench Verified, K=3 reduces context per call by 43% for Qwen3-4B and 57% for SWE-Master-4B-RL, with resolve rates of 12.1% and 21.8% versus 14.5% and 27.5% for their uncompressed bases. A single-run recency sweep reaches 14.4% and 23.0% at K=8, with larger windows generally favoring task performance over compression. Under a 32K-token limit, Qwen3 with K=3 resolves 21.1% of a 199-instance subset versus 11.1% for the same adapted agent using full text. In concurrent single-GPU serving, it achieves 1.9 times that full-text agent's instance throughput.
Chinese Translation
工具观测结果在软件工程智能体的上下文中占据主导地位,使得维护长交互历史的成本高昂。现有的上下文压缩方法可能会丢弃后续操作所需的信息,而让智能体适应软词元(soft-token)表示则可能损害其原有行为。为了在保留动作关键信息和智能体行为的同时减少上下文,我们结合了潜观测-硬动作(Latent Observations, Hard Actions, LOHA)与锚定上下文蒸馏(Anchored Context Distillation, ACD):前者是一种上下文布局,将压缩后的历史与用于精确引用的文本分离;后者是一种训练方法,在使能潜表示读取的同时约束行为漂移。LOHA将较早的工具观测压缩为软词元,同时以文本形式保留智能体自身的对话轮次和最近K次观测,从而实现对历史信息的紧凑访问和对近期内容的精确访问。为使智能体能够利用该表示,ACD将基础模型在全文输入上的预测蒸馏到潜视图中,同时将其行为锚定在同一基础模型的纯文本输入上。在SWE-bench Verified上,当K=3时,Qwen3-4B和SWE-Master-4B-RL的每次调用上下文分别减少43%和57%,解决率分别为12.1%和21.8%,而其未压缩基线分别为14.5%和27.5%。单次运行的近期窗口扫描显示,K=8时可达到14.4%和23.0%,且较大的窗口总体上更倾向于任务性能而非压缩率。在32K词元限制下,K=3的Qwen3在一个199实例的子集上解决了21.1%的问题,而使用全文的同一适配智能体仅为11.1%。在单GPU并发服务场景下,其实例吞吐量达到该全文智能体的1.9倍。
cs.AI / 80 / 2609.31460

Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency

面向更优数据探索与资源效率的片段级智能体式主题建模
Jang, Myeongjun Erik, Georgiadis, Antonios, Moon, Sae Young, Silavong, Fran
Abstract
Topic modeling is an effective technique for discovering hidden themes within documents and is widely used in text mining and data analysis across a variety of industry sectors. Recently, large language model (LLM)-based topic models have been emerged that prompt LLMs to generate topics then assign the topics to documents, producing more natural and human-readable topics than conventional topic modeling algorithms. However, the nature of topic assignment process causes certain drawbacks, such as the incapability to produce topic distributions over a document, too broad or narrow topics, and high resource consumption, which increases with the number and length of of documents being assigned topics. These issues are particularly critical for industrial applications, which require high-quality, in-depth analysis and the processing of large volumes of documents. In this context, this paper introduces a framework called SeLATM, which addresses these concerns by employing segment-level topic generation and topic refinement through agentic feedback loops. Experimental results on various datasets demonstrate that SeLATM significantly reduces the LLM resources compared to methods based on topic assignment process, while maintaining superior performance.
Chinese Translation
主题建模是一种发现文档中隐藏主题的有效技术,被广泛应用于各行业领域的文本挖掘与数据分析中。近年来,出现了基于大语言模型(LLM)的主题模型,其通过提示LLM生成主题并将主题分配给文档,能够产生比传统主题建模算法更自然、更易读的主题。然而,主题分配过程的固有特性带来了一些缺陷,例如无法生成文档层面的主题分布、主题过于宽泛或狭窄,以及高昂的资源消耗——该消耗随待分配主题的文档数量和长度的增加而增长。这些问题对于需要高质量、深入分析并处理海量文档的工业应用而言尤为突出。鉴于此,本文提出了一个名为SeLATM的框架,通过片段级的主题生成以及基于智能体反馈循环的主题优化来解决上述问题。在多个数据集上的实验结果表明,与基于主题分配过程的方法相比,SeLATM在保持更优性能的同时,显著降低了LLM资源消耗。
cs.AI / 81 / 2609.31473

Game Arena: Strategic LLM Evaluation in Competitive Environments

Game Arena:竞争环境中的战略性大语言模型评估
Doerschuk-Tiberi, Bovard, Yan, Yao, Chiu, Justin, Wang, Hann, Chung, Timothy, Plomecka, Martyna, Schultz, John, Lipovetz, Jon, Drazner, Clayton, Zhuang, Yuchen, Hwang, Jaimie, Keating, Nate, Jones, Riley, Lee, Andrew, Kelly, Oran, Gemp, Ian, Aaron, Michael, Prince, Laurel, Larson, Kate, Moser, Jeff, Jobe, Harrison, Woodford, Chad, Liu, Siqi, Wang, Andrew, Chang, Bo, D'Mello, Christopher, Chaleff, Diane, Howard, Addison, Yip, Johnny, Sugnet, Chuck, Gulli, Antonio, O'Connell, Meghan, Cukierski, Will, Tomasev, Nenad, Yeroshenko, Dima, Parekh, Kinjal, Daniel, Roxanne, Lanctot, Marc, Weir, Domino, Dong, Elsa, Hennes, Daniel, Nalubwama, Melissa, Fraser, Robert, Trostle, Ryan, Peng, Jun, Mason, Tom, Hightower, Lloyd, Chukwuka, Chiamaka, Zhai, Yuexiang, Kirk, Phoebe, Su, Yi, Han, Yuting, Ren, Jie, Prichard, Chris, Sharifzadeh, Sahand, Hakimzadeh, Karim, Sterling, DJ, Risdal, Meg, Olszewska, Kate, Xu, Ya, Firat, Orhan, Chen, Minmin
Abstract
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. These environments span perfect information, imperfect information, and multiplayer game settings, enabling a systematic study of models' strategic planning, adaptation, and robustness under uncertainty. For each game, we provide a detailed description of the environment, evaluation metrics, and results from running full competitions across models. Through robust infrastructure and large-scale ground-truth based evaluation, Game Arena ensures reproducibility, transparency and generalizability to new games and variants over time.
Chinese Translation
我们推出了 Kaggle Game Arena,一个通过竞技游戏评估大语言模型(LLM)的开放且持续扩展的平台。与静态基准测试不同,游戏竞技场使模型能够在结构化环境中进行一对一对抗,其博弈强度随模型的演进而自然提升,从而避免性能饱和。本技术报告详细介绍了 Game Arena 背后的基础设施,并描述了三个试点游戏环境:国际象棋、德州扑克和狼人杀。这些环境涵盖完全信息、不完全信息以及多人博弈设置,能够系统地研究模型的策略规划、适应能力以及在不确定性下的鲁棒性。针对每个游戏,我们详细描述了环境设置、评估指标以及跨模型完整竞赛的结果。凭借稳健的基础设施和基于真实标准的大规模评估,Game Arena 确保了可复现性、透明性,以及对新游戏和新变体的长期可扩展性。
cs.AI / 82 / 2609.31482

"AI is (not) the new...": A Diagnostic Analogy Framework for Generative AI's Cultural Impacts

"AI是(不是)新的……":一个诊断生成式AI文化影响的类比分析框架
Qadri, Rida, Prabhakaran, Vinodkumar, Denton, Remi
Abstract
Generative AI is reshaping the cultural infrastructures through which knowledge is found, synthesized, and held accountable. To make sense of this shift, scholars and policymakers reach for historical analogies of technologies such as the printing press, steam power or electricity. But these comparisons are typically imprecise about which property of the technology carries the comparison, and imprecise analogies produce imprecise governance by designing interventions against the wrong property of the system. This paper offers a diagnostic framework for analyzing how generative AI can transform epistemic and cultural practice. This paper offers a diagnostic framework for analyzing how generative AI can transform epistemic and cultural practice. We decompose each intervention into three coordinates: the epistemic site at which a technology acts, the governing logic by which it organizes its object, and the technical mechanism through which the logic is instantiated. This framework allows us to distinguish between structural cultural consequences, which follow from the mechanism itself, from contingent ones, which remain open to design and institutional choice. Applying the framework to information discovery and knowledge synthesis, we show how the shift from indexicality to inference and from editorial authority to statistical consensus produces specific, traceable cultural effects and reveals governance levers that gestalt analogy obscures.
Chinese Translation
生成式AI正在重塑知识的发现、综合与问责所依赖的文化基础设施。为了理解这一转变,学者和政策制定者常常援引印刷机、蒸汽机或电力等历史技术类比。然而,这些比较通常未能明确技术的哪种属性支撑了这一类比,而不精确的类比会导致不精确的治理,因为针对系统错误属性设计的干预措施难以奏效。本文提出了一个诊断性框架,用于分析生成式AI如何改变认知与文化实践。我们将每一项干预分解为三个坐标:技术发挥作用的认知场域、组织其对象的主导逻辑,以及该逻辑得以实现的技术机制。这一框架使我们能够区分源于机制本身的结构性文化后果与仍然可以通过设计和制度选择加以改变的偶然性后果。将该框架应用于信息发现与知识综合,我们展示了从索引性到推理、从编辑权威到统计共识的转变如何产生特定的、可追溯的文化效应,并揭示出整体性类比所遮蔽的治理杠杆。
cs.AI / 83 / 2609.31491

UQ-LOB: Uncertainty-Aware Limit Order Book Mid-Price Forecasting

UQ-LOB:不确定性感知的限价订单簿中间价预测
Manoharan, Derrick Gilchrist Edward, Linna, Eljas, Baltakys, Kestutis, Dong, Hao, Kanniainen, Juho
Abstract
Forecasting short-horizon mid-price movements from limit order book (LOB) data is central to algorithmic trading, yet most deep LOB forecasters are point predictors: they output a direction or a displacement, but never indicate which of their forecasts can be trusted. We introduce UQ-LOB, a lightweight, encoder-agnostic uncertainty quantification module that attaches to any pretrained LOB encoder and, in the spirit of attentive neural processes, conditions each forecast on a context set of recently completed windows whose outcomes are already realised. The UQ-regression variant outputs a calibrated Gaussian over the future tick displacement, while the UQ-classification variant outputs a categorical distribution over down/up/stationary. Both expose a scalar confidence (predicted signal-to-noise ratio or class probability) that supports selective prediction. On 5.2 billion LOB events across seven cryptocurrency assets and horizons of 5, 10 and 15 seconds, UQ-regression attains near-nominal 68% interval coverage, and restricting to the most confident 10% of predictions raises directional macro F1 by 0.11-0.15 for UQ-regression and 0.05-0.11 for UQ-classification, at every horizon. On large, economically meaningful moves, the tightest confidence tier reaches a directional F1 of 0.88 (down) and 0.83 (up) at the 5-second horizon.
Chinese Translation
基于限价订单簿(LOB)数据预测短时间跨度的中间价变动是算法交易的核心问题,然而大多数深度LOB预测模型都是点预测器:它们只输出方向或位移,却从不指出哪些预测结果是可信的。我们提出了UQ-LOB,一个轻量级、与编码器无关的不确定性量化模块,它可以附加到任何预训练的LOB编码器上,并秉承注意力神经过程(attentive neural processes)的思想,使每次预测都以一组近期已完成、结果已知的窗口作为上下文集为条件。UQ回归变体输出关于未来tick位移的校准高斯分布,而UQ分类变体输出关于下跌/上涨/平稳的类别分布。两者都提供一个标量置信度(预测信噪比或类别概率),以支持选择性预测。在涵盖七种加密货币资产、时间跨度为5秒、10秒和15秒的52亿条LOB事件上,UQ回归实现了接近名义值的68%区间覆盖率;在所有时间跨度上,仅保留置信度最高的10%预测,UQ回归的方向性宏观F1提升0.11-0.15,UQ分类提升0.05-0.11。对于幅度较大且具有经济意义的行情变动,最紧的置信度层级在5秒时间跨度上达到0.88(下跌)和0.83(上涨)的方向性F1。
cs.AI / 84 / 2609.31505

Prompt Minimization: Reducing Input Redundancy Without Sacrificing Output Fidelity

提示词最小化:在不牺牲输出保真度的前提下降低输入冗余
Juston, Marius F. R., Karim, Kevin A., Gao, Jonathan, Li, Kevin C., Bashambu, Rudhi
Abstract
Despite the growing capabilities of large language models (LLMs), prompt design remains largely heuristic and ad hoc. This project will explore $\textit{prompt minimization}$, the process of reducing prompts to their smallest, most information-dense form while preserving output fidelity. Practically, shorter prompts reduce computational overhead and inference latency, especially when large contexts, such as entire documents or codebases, are included unnecessarily. Further, longer prompts can damage LLM reasoning and accuracy. Theoretically, the existence of multiple prompts yielding equivalent outputs suggests a high degree of redundancy in the input space, raising fundamental questions about what information is essential to elicit specific model behaviors. We propose three variant frameworks to identify and evaluate minimal prompts and demonstrate that minimal prompts often produce outputs comparable to those of their longer counterparts. These findings suggest new directions for efficient prompt engineering and deepen our understanding of input compression in LLMs.
Chinese Translation
尽管大型语言模型(LLM)的能力日益增强,提示词设计在很大程度上仍然依赖启发式和临时性的方法。本项目将探索“提示词最小化”(prompt minimization),即将提示词压缩至最小、信息密度最高的形式,同时保持输出保真度。在实践中,更短的提示词可以降低计算开销和推理延迟,尤其是在不必要地包含大型上下文(如整篇文档或代码库)的情况下。此外,较长的提示词还可能损害LLM的推理能力和准确性。从理论上讲,多个提示词能够产生等效输出这一现象表明输入空间存在高度冗余,并引出了关于哪些信息对于引发特定模型行为是必不可少的根本性问题。我们提出三种变体框架来识别和评估最小提示词,并证明最小提示词通常能产生与较长提示词相当的输出。这些发现为高效提示词工程指明了新方向,并加深了我们对LLM输入压缩的理解。
cs.AI / 85 / 2609.31544

A Flow Matching Framework for Neural Representational Dissimilarity

神经表征相异性的流匹配框架
Ye, Zeyuan, Wei, Xue-Xin
Abstract
Neural representational dissimilarity quantifies differences between neural response distributions, and is essential for comparing neural codes across stimuli, brain areas, tasks, and models. Commonly used distance metrics involve different assumptions and are estimated with separate methods. Here, we show that a variety of distance metrics can be unified under a flow matching framework developed in deep generative models. That is, these distances arise as Jeffreys divergences under different velocity constraints. We find that flow matching has advantages for estimating distances involving complicated distributions and continuous variables. Furthermore, this framework enables the design of new distance metrics in a principled way. Together, flow matching provides a unified approach for understanding, estimating, and designing neural representational dissimilarity metrics.
Chinese Translation
神经表征相异性(Neural Representational Dissimilarity)量化了神经响应分布之间的差异,对于比较跨刺激、脑区、任务和模型的神经编码至关重要。常用的距离度量涉及不同的假设,且需采用各自独立的方法进行估计。本文证明,多种距离度量可以在深度生成模型中发展的流匹配(Flow Matching)框架下得到统一,即这些距离可表示为在不同速度约束下的Jeffreys散度。我们发现,流匹配在估计涉及复杂分布和连续变量的距离方面具有优势。此外,该框架还支持以有原则的方式设计新的距离度量。总之,流匹配为理解、估计和设计神经表征相异性度量提供了一种统一的方法。
cs.AI / 86 / 2609.31563

Multi-agent Scaling Across Disjunctive and Compensatory Tasks

跨析取型与补偿型任务的多智能体扩展研究
Fortuna, Carolina, Bertalanic, Blaz
Abstract
Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answer, and averaging converges to the model's item-level bias. Across selected representative benchmarks, 13 open-weight models, and teams of up to 30 agents, we find qualitatively different scaling behavior. On disjunctive tasks, the probability that at least one agent is correct grows by 5-20 points with team size, but plurality voting over agents that answer directly realises almost none of this potential, as the model predicts to within 0.5 points on average. Multi-round revision raises accuracy considerably, yet the gain is nearly the same with one peer as with 29. In contrast, scaling provides little benefit on Fermi estimation, despite its natural suitability for aggregation: item-level biases shared across the samples of a model account for about 87% of the squared error, so averaging reduces error by only about 6%. Combining model families helps on Fermi estimation but does not surpass the strongest member on disjunctive tasks. These results show that task structure, together with the mechanism combining member outputs, is a fundamental determinant of team scaling.
Chinese Translation
多智能体大语言模型(LLM)系统通常被期望随着团队规模的扩大而性能提升,但其扩展行为可能取决于任务结构。我们的核心贡献是引入 Steiner 的群体任务分类法作为分析多智能体 LLM 扩展的框架,并将分析聚焦于析取型(disjunctive)和补偿型(compensatory)任务。我们将独立采样的智能体建模为在给定题目条件下相互独立,由此推导出大团队极限:多数投票收敛于模型的最常见答案(modal answer),而取平均值则收敛于模型在题目层面的偏差(item-level bias)。在选定的代表性基准上,我们使用 13 个开源权重模型以及最多 30 个智能体的团队,观察到定性上截然不同的扩展行为。在析取型任务上,至少一个智能体回答正确的概率随团队规模增长 5-20 个百分点,但对直接作答的智能体进行多数投票几乎无法实现这一潜力——模型预测与实际结果的平均误差在 0.5 个百分点以内。多轮修订能显著提高准确率,然而使用 1 个同伴与使用 29 个同伴所带来的提升几乎相同。相比之下,扩展在费米估算(Fermi estimation)上几乎没有收益,尽管该任务表面上非常适合聚合:模型各样本间共享的题目层面偏差约占总平方误差的 87%,因此取平均仅能将误差降低约 6%。组合多个模型家族有助于费米估算,但在析取型任务上无法超越其中最强的成员。这些结果表明,任务结构与整合成员输出的机制共同构成团队扩展效果的根本决定因素。
cs.AI / 87 / 2609.31568

DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education

DeepEdu-v1:面向越南教育的高效可扩展智能体大语言模型
Nguyen, Quang, Nguyen, Hieu, Hoang, Hien, Pham, Toan, Tran, Cong, Vu, Nam
Abstract
AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open model keeps data on-premise but hits a two-fold wall: post-training quantization (AWQ, GPTQ) tames the static weight footprint, yet the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs, while the model keeps hallucinating on region-specific material. We present DeepEdu-v1, an AI-tutoring system for Vietnamese education built on SCALE (Self-improving Context-Aware Learning Engine), a framework with two innovations. First, a long-context inference engine amortizes token selection from per-sub-chunk to per-cluster granularity; on long-context retrieval it issues x7.7 fewer retrieval calls than a state-of-the-art selective-attention baseline, cutting prefill latency (TTFT) by roughly 35% while matching or improving task accuracy. Second, a self-improving agentic layer continuously curates a verified playbook from past interactions instead of fine-tuning, a design intended to progressively reduce reliance on dominant-language priors as trustworthy local knowledge accumulates. In its deployed configuration, DeepEdu achieves a nearly x2 TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0% to 79.5% on complex tasks, with the strongest per-track gains across financial-reasoning and interactive-agent benchmarks.
Chinese Translation
AI 辅导系统有望显著改善越南等发展中地区学生的学习效果,然而两条显而易见的技术路径均存在不足。以 ChatGPT 为代表的云端助手会将敏感的学生数据传输至境外服务器——这违反了越南第 53 号法令等数据主权法律——并且由于在以西方为中心的语料上进行预训练,它们并未围绕国家教材课程体系进行组织,因此对本地内容的知识缺乏系统性且经常产生幻觉。自行部署开源模型虽然能将数据保存在本地,却面临双重障碍:训练后量化(AWQ、GPTQ)虽能控制静态权重的显存占用,但长上下文辅导场景下的动态 KV 缓存和预填充延迟仍会在消费级 GPU 上导致内存溢出和响应缓慢,同时模型在本地特色内容上仍持续产生幻觉。我们提出 DeepEdu-v1,一个基于 SCALE(自我改进的上下文感知学习引擎)框架构建的面向越南教育的 AI 辅导系统,该框架包含两项创新。第一,长上下文推理引擎将 token 选择的开销从子块粒度摊销到簇粒度;在长上下文检索任务中,与最先进的选择性注意力基线相比,其检索调用次数减少约 7.7 倍,预填充延迟(TTFT)降低约 35%,同时保持或提升了任务准确率。第二,自我改进的智能体层通过持续地从历史交互中整理经过验证的操作手册(playbook)来替代微调,该设计旨在随着可信本地知识的积累,逐步降低模型对主导语言先验的依赖。在部署配置下,DeepEdu 相比标准 vLLM 服务实现了近 2 倍的 TTFT 加速,并将智能体在复杂任务上的准确率从 70.0% 提升至 79.5%,其中在金融推理和交互式智能体基准测试上取得了各赛道中最显著的性能提升。
cs.AI / 88 / 2609.31619

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

Learning to Stop without Learning to Stop:自监督置信度训练提升推理效率
Hosseini, Parsa, Tigalappanavara, Akasha, Nawathe, Sumit, Fan, Chenrui, Basu, Sourya, Winata, Genta Indra, Das, Anirban, Feizi, Soheil, Chitsazan, Nima
Abstract
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, the fine-tuned models use the standard generation procedure, with no confidence elicitation or early-stopping mechanism. Despite this, self-supervised confidence fine-tuning makes reasoning more efficient, reducing generated tokens by up to 25\% at matched accuracy across Gemma, Qwen, Nemotron, and GPT-OSS models on mathematical, scientific, and coding reasoning benchmarks, with efficiency gains comparable to methods that explicitly optimize for shorter reasoning. Analysis of reasoning episodes further shows that confidence supervision largely preserves the base models' high-level reasoning composition rather than selectively suppressing particular behaviors. Our results suggest that efficient reasoning may emerge as a downstream consequence of learning metacognitive signals, without being directly optimized.
Chinese Translation
推理模型通常会生成非常长的推理轨迹,使得推理的计算开销很高。现有方法通常通过两种途径提升效率:一是推理时的提前停止机制,二是在训练中显式鼓励更短的推理,例如使用带长度惩罚的强化学习。我们证明,可观的效率提升也可以源于另一种类型的监督信号:置信度(confidence)。我们采用一种自监督流程,仅使用600个训练问题,微调推理模型,使其能够在自身推理轨迹的中间节点处预测对答案的置信度。置信度仅作为训练目标使用:损失函数中不包含任何关于推理长度、效率或停止的目标。在推理阶段,微调后的模型使用标准的生成流程,没有置信度引出或提前停止机制。尽管如此,自监督置信度微调仍使推理更加高效:在匹配精度的条件下,Gemma、Qwen、Nemotron 和 GPT-OSS 模型在数学、科学和编程推理基准上生成的 token 数最多减少 25%,其效率提升与那些显式优化更短推理的方法相当。对推理过程的分析进一步表明,置信度监督在很大程度上保留了基础模型的高层推理组成结构,而非选择性地抑制特定行为。我们的结果表明,高效推理可能是学习元认知信号的一种下游结果,而无需被直接优化。
计算机视觉 (Computer Vision)
87
cs.CV / 1 / 2609.30356

AlphaEarth distinguishes cities but compresses urban variation

AlphaEarth能够区分城市但压缩了城市内部的差异性
Renninger, Andrew
Abstract
Cities differ in built form, land cover and development history, complicating comparison across places and time. Satellite foundation models map Earth's surface onto common numerical representations. Yet the tasks and targets used to shape them typically do not focus on cities: globally consistent labels for urban function do not exist, and many datasets - especially land cover and land use classifications - collapse the built environment into few classes. Here we audit the representation, focusing on AlphaEarth but with broader applicability to other Earth embeddings, by probing the geometry and geography of embeddings for 1,000 urban areas in 162 countries. We find that cities occupy a shifted but overlapping region on the hypersphere, 62.7{\deg} from the global mean direction, and continent and climate predict 24.3% of variation among the mean directions of urban centres in excluded countries. Inside cities, degrees of urbanisation carry 8.9% of the variation, and what they leave holds shared directions whose local orientation varies, not one universal axis of urbanisation. Retained variation is itself unequal: dispersion within urban centres is 14.1% greater per standard deviation of national development, even after adjusting for population, land area and continent. Further controls suggest cities in developing countries present less contrast in vegetation and texture, and dispersion follows that contrast: full adjustment for it leaves at most 6.4% of the gradient. Annually, a city's representation moves nearly eight times more than redrawing its own pixels explains, and contracts where the 2022 loss of Sentinel-1B removed a pass direction. AlphaEarth's representations therefore support comparison across regions, while the differences between its annual layers are not yet validated for comparison over time.
Chinese Translation
城市在建成形态、地表覆盖和发展历史方面各不相同,这使得跨地点和跨时间的比较变得复杂。卫星基础模型将地球表面映射为统一的数值表示。然而,用于塑造这些模型的任务和训练目标通常并不聚焦于城市:全球一致的城市功能标注并不存在,且许多数据集——尤其是土地覆盖与土地利用分类数据——将建成环境压缩为极少的类别。本文对这类表示进行了审计,重点聚焦于AlphaEarth,但其结论对其他地球嵌入模型也具有更广泛的适用性。我们通过探测162个国家1,000个城市区域的嵌入的几何与地理特性展开研究。我们发现,各城市在超球面上占据一个偏移但相互重叠的区域,距全球平均方向62.7度,且大陆和气候可以预测被排除国家中城市中心平均方向差异的24.3%。在城市内部,城市化程度承载了8.9%的变异,其余变异所包含的是方向随地点变化的共享方向,而非一条普适的城市化轴线。被保留的变异本身也不均衡:即使经过人口、土地面积和大陆的调整,城市中心内部的离散度仍随国家发展水平的每个标准差增加14.1%。进一步的控制表明,发展中国家的城市在植被和纹理上呈现较低的对比度,而离散度正跟随这种对比度变化:对其充分调整后,最多仅剩6.4%的梯度。在年度尺度上,城市表示的变动几乎是重新绘制其自身像素所能解释的近八倍,并且在2022年Sentinel-1B失联导致某一过境方向缺失的地区,表示出现收缩。因此,AlphaEarth的表示支持跨区域比较,而其年度层之间的差异尚未经过验证,不能用于跨时间比较。
cs.CV / 2 / 2609.30393

LiTe-GS: Oracle-Efficient Next Best View Selection for 3D Gaussian Splatting

LiTe-GS:面向3D高斯泼溅的Oracle高效下一最佳视角选择方法
Pandey, Vivek, Khass, Amirhossein Mollaei, Motee, Nader
Abstract
Selecting informative camera views is critical for efficient training and adaptive refinement in 3D Gaussian Splatting, where each observation significantly influences model parameters. However, information-driven view-selection strategies can require repeated evaluations of expensive information-gain oracles as the number of candidate views increases. We propose LiTe-GS, an oracle-efficient method for next best view selection in 3D Gaussian Splatting. LiTe-GS reduces the number of information-oracle evaluations by performing randomized subset evaluation of candidate views rather than exhaustively scoring the full candidate pool. The resulting approach achieves expected $O(M\log(1/\epsilon))$ oracle complexity with respect to the number of candidate views $M$, independent of the selection cardinality $K$, while providing an explicit trade-off between oracle efficiency and approximation quality through $\epsilon$. We provide theoretical guarantees on oracle complexity and approximation performance under the proposed selection scheme. Experiments on Blender and Mip-NeRF 360 demonstrate that LiTe-GS maintains reconstruction quality comparable to Fisher-information-based baselines while substantially reducing the number of Fisher-oracle evaluations across different acquisition settings.
Chinese Translation
在3D高斯泼溅(3D Gaussian Splatting)中,选择信息量大的相机视角对于高效训练和自适应优化至关重要,因为每次观测都会显著影响模型参数。然而,随着候选视角数量的增加,基于信息的视角选择策略可能需要反复调用昂贵的信息增益Oracle进行评估。我们提出了LiTe-GS,一种用于3D高斯泼溅中下一最佳视角选择的Oracle高效方法。LiTe-GS通过对候选视角进行随机子集评估,而非对完整候选池进行穷举打分,从而减少了信息Oracle的评估次数。该方法实现了关于候选视角数量 $M$ 的期望 $O(M\log(1/\epsilon))$ 的Oracle复杂度,且与选择基数 $K$ 无关,同时通过 $\epsilon$ 在Oracle效率与近似质量之间提供了显式的权衡。我们在所提出的选择方案下给出了关于Oracle复杂度和近似性能的理论保证。在Blender和Mip-NeRF 360数据集上的实验表明,LiTe-GS在不同采集设置下,在保持与基于Fisher信息的基线方法相当的重建质量的同时,显著减少了Fisher-Oracle的评估次数。
cs.CV / 3 / 2609.30395

CSCWD: Cross-Scale Channel-wise Knowledge Distillation for Lightweight Tiny Object Detection on Edge Devices

CSCWD:面向边缘设备轻量级微小目标检测的跨尺度通道级知识蒸馏
Zamani, Amir, Ghasemi-Naraghi, Zeinab
Abstract
Real-time tiny object detection in aerial imagery is constrained by the weak spatial evidence of very small objects and the loss of high-resolution detail in lightweight detectors. This study presents Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers high-resolution spatial representations from a YOLO11m-P2 teacher to a compact YOLO11n student without altering the student's inference architecture. Unlike conventional same-scale feature distillation, CSCWD transfers supervision from teacher P2 to student P3 after feature alignment while retaining same-scale distillation at deeper pyramid levels. Under the unified seven-sequence Drone-vs-Bird validation protocol, YOLO11n-CSCWD achieves 50.17% mean average precision at an intersection-over-union threshold of 0.5 (mAP@0.5) and 59.73% recall, improving the matched CA-YOLO11n baseline by 2.92 percentage points in mAP@0.5 and 3.55 points in recall. Cross-scale alignment further increases mAP@0.5 by 2.09 points over the corresponding same-scale channel-wise distillation configuration. In zero-shot evaluation on DUT-Anti-UAV, mAP@0.5 increases from 48.29% to 50.06% without target-domain fine-tuning. This domain was included because its challenging small targets make low-latency, computationally efficient detection particularly relevant. On Raspberry Pi 5 using NCNN-FP16 at 640x640 resolution, the 2.58-million-parameter student achieves 50.32% mAP@0.5 at 82.32 ms mean wall-clock latency, or 12.15 frames per second, while retaining essentially the same runtime and memory requirements as the matched baseline. The results support cross-scale distillation for improving tiny-target detection without increasing inference-time model complexity.
Chinese Translation
航拍图像中的实时微小目标检测受到极小目标空间证据薄弱以及轻量级检测器高分辨率细节丢失的制约。本研究提出了跨尺度通道级知识蒸馏(Cross-Scale Channel-wise Knowledge Distillation, CSCWD),这是一种训练时框架,可将高分辨率空间表征从YOLO11m-P2教师模型迁移到紧凑的YOLO11n学生模型,且不改变学生模型的推理架构。与传统的同尺度特征蒸馏不同,CSCWD在特征对齐后将监督信号从教师模型的P2层迁移到学生模型的P3层,同时在更深的金字塔层级保留同尺度蒸馏。在统一的七序列无人机-鸟类(Drone-vs-Bird)验证协议下,YOLO11n-CSCWD在交并比阈值0.5处取得50.17%的平均精度均值(mAP@0.5)和59.73%的召回率,相比对应的CA-YOLO11n基线在mAP@0.5上提升2.92个百分点,召回率提升3.55个百分点。跨尺度对齐相比相应的同尺度通道级蒸馏配置进一步将mAP@0.5提高2.09个百分点。在DUT-Anti-UAV数据集的零样本评估中,无需目标域微调,mAP@0.5即从48.29%提升至50.06%。纳入该数据集是因为其具有挑战性的小目标使得低延迟、计算高效的目标检测尤为重要。在Raspberry Pi 5上使用NCNN-FP16以640x640分辨率运行时,该260万参数的学生模型在82.32毫秒的平均时延(即每秒12.15帧)下达到50.32%的mAP@0.5,同时运行时间和内存需求与对应基线基本相同。结果表明,跨尺度蒸馏能够在不增加推理时模型复杂度的情况下提升微小目标检测性能。
cs.CV / 4 / 2609.30402

What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study

什么能改进多模态虚假信息检测?来自大规模实证研究的答案
Sharma, Akshit, Patil, Prashant W.
Abstract
Multimodal misinformation is increasingly crafted to look convincing by pairing a textual claim with an image that appears to "prove" it. Yet in practice, building effective detectors often hinges on a small set of design choices that are rarely examined in a controlled way. In this paper, we conduct a large-scale study of multimodal design choices for misinformation detection with over 3,375 experiments- spanning three benchmark datasets and a broad range of pre-trained vision and language backbones. Through systematic comparisons and targeted robustness analyses, we distill practical guidance on which design choices help, when do they fail silently, and what aspects of the pipeline most strongly shape model behavior, answering 4 key Research Questions (RQs). We aim to provide a reliable foundation for designing stronger and more dependable multimodal misinformation detection systems, thus contributing to the broader research community.
Chinese Translation
多模态虚假信息正日益被精心制作以显得可信,其手法是将一段文本声明与一幅看似能够"证明"该声明的图像配对。然而在实践中,构建有效的检测器往往取决于一小部分设计选择,而这些设计选择很少在受控条件下被系统考察。在本文中,我们针对多模态虚假信息检测的设计选择开展了一项大规模研究,涵盖超过3,375项实验,涉及三个基准数据集以及广泛多样的预训练视觉与语言骨干模型。通过系统的比较和有针对性的鲁棒性分析,我们总结出实用的指导原则,阐明了哪些设计选择有效、它们何时会悄无声息地失效,以及流水线中的哪些方面对模型行为影响最大,从而回答了四个关键的研究问题(Research Questions, RQs)。我们旨在为设计更强大、更可靠的多模态虚假信息检测系统提供可靠的基础,从而为更广泛的研究社区做出贡献。
cs.CV / 5 / 2609.30434

ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models

ProCAP:面向视觉-语言模型的概率跨注意力提示学习
Abbas, Hiwa Azeez, Daneshfar, Fatemeh, Abdar, Moloud
Abstract
Pre-trained vision-language models such as CLIP can recognize new categories via prompting, but they often struggle when labeled data are scarce or the test distribution shifts. Prompt learning adapts only a small set of parameters while keeping the backbone frozen, yet many existing multimodal prompt learners couple the visual and textual branches weakly and can be brittle in low-shot regimes. We propose ProCAP, a probabilistic cross-attentive prompt learning framework that improves cross-modal interaction and training stability without updating any CLIP weights: it learns both visual and textual prompt tokens and links them through stacked bidirectional multi-head cross-attention so the two branches refine each other across prompt depth. To reduce overfitting under limited supervision, we parameterize prompt tokens with Gaussian means and variances and regularize them with lightweight KL and L2 penalties, and we further add a compact symmetric InfoNCE head that aligns cross-attended image features with class-level text representations in a shared low-dimensional space. Across few-shot base-to-novel generalization on 11 datasets, cross-dataset transfer, and domain generalization on ImageNet shift benchmarks, ProCAP achieves strong aggregate base-to-novel performance and competitive transfer performance while keeping the CLIP backbone unchanged.
Chinese Translation
CLIP 等预训练视觉-语言模型可以通过提示(prompting)识别新类别,但在标注数据稀缺或测试分布发生偏移时往往表现不佳。提示学习仅微调少量参数而保持主干网络冻结,然而许多现有多模态提示学习方法对视觉分支与文本分支的耦合较弱,在少样本(low-shot)场景下容易失效。我们提出 ProCAP,一种概率跨注意力提示学习框架,能够在不更新任何 CLIP 权重的情况下增强跨模态交互并提升训练稳定性:该方法同时学习视觉与文本提示词元(prompt tokens),并通过堆叠的双向多头交叉注意力将其关联,使两个分支在提示深度上相互精炼。为在有限监督下降低过拟合风险,我们用高斯均值与方差对提示词元进行参数化,并采用轻量级的 KL 散度与 L2 正则化加以约束;此外,我们还引入一个紧凑的对称 InfoNCE 头,在共享的低维空间中将交叉注意力后的图像特征与类别级文本表示对齐。在 11 个数据集上的少样本基类到新类泛化、跨数据集迁移以及 ImageNet 分布偏移基准上的域泛化实验中,ProCAP 在保持 CLIP 主干不变的情况下,取得了优异的整体基类到新类性能和具有竞争力的迁移性能。
cs.CV / 6 / 2609.30450

LensDesigner: A Self-Improving Agent for Optical Lens Design

LensDesigner:一种用于光学透镜设计的自我改进智能体
Sun, Lei, Liang, Haoran, Xu, Dannong, Gao, Yao, Geng, Yuyu, Gu, Jinjin, Wang, Kaiwei, Paudel, Danda Pani, Van Gool, Luc
Abstract
Optical lens design is a complex, non-convex optimization challenge that relies heavily on human experience and intuition. Existing optimized-based automatic lens design methods struggle to navigate this vast parameter space without meticulous manual tuning. In this paper, we present LensDesigner, an autonomous agent framework that mirrors the problem-solving workflow of expert opticians. To overcome the initial cold start problem, we construct LensLib100K, an extensive optical lens library, and employ Optics-Aware Retrieval to supply physically valid structural seeds. Within an interactive physical simulation environment, the agent executes macroscopic orchestration while receiving immediate optical feedback. Furthermore, we introduce a continuous self-evolving mechanism guided by a curriculum agent. By iteratively solving design tasks with progressively increasing difficulty, the agent autonomously extracts, accumulates, and reuses design heuristics, effectively evolving its optical lens design expertise over time. At the evaluation level, we introduce LensArena, a standardized evaluation benchmark comprising $120$ diverse optical design tasks, covering extreme configurations. Extensive experiments on this benchmark demonstrate that LensDesigner significantly outperforms publicly available baseline algorithms, achieving superior success rates and optimization efficiency. We hope this work sheds light on the emerging field of intelligent optics. The code will be publicly available.
Chinese Translation
光学透镜设计是一个复杂的非凸优化难题,高度依赖人类经验与直觉。现有基于优化的自动透镜设计方法若缺乏精细的人工调参,难以在这一庞大的参数空间中有效搜索。本文提出LensDesigner,一个模拟专家光学设计师问题求解流程的自主智能体框架。为克服初始冷启动问题,我们构建了LensLib100K——一个大规模光学透镜库,并采用光学感知检索(Optics-Aware Retrieval)来提供物理上有效的结构种子。在交互式物理仿真环境中,智能体执行宏观层面的流程编排,并即时接收光学反馈。此外,我们引入了由课程智能体(curriculum agent)引导的持续自我演化机制。通过迭代求解难度逐步递增的设计任务,智能体能够自主地提取、积累并复用设计启发式知识,从而随时间不断演进其光学透镜设计专长。在评估层面,我们提出LensArena——一个包含120个多样化光学设计任务(涵盖极端配置)的标准化评估基准。在该基准上的大量实验表明,LensDesigner显著优于公开可用的基线算法,取得了更高的成功率和优化效率。我们希望这项工作能为智能光学的兴起领域带来启示。代码将公开提供。
cs.CV / 7 / 2609.30478

The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

事件的形状:通过跨域蒸馏引入基于边缘的归纳偏置
Kihara, Soshun, Yasuki, Shunsuke, Taki, Masato
Abstract
Convolutional neural networks trained on ImageNet are known to exhibit a strong preference for local high-frequency texture, an inductive bias that translates into fragile robustness against distribution shifts in real-world environments. Event cameras, in contrast, record only changes in scene brightness and are therefore well suited to capturing contour information; however, due to the absence of diagnostic benchmarks in the event domain, the inductive bias that event-camera data instills in vision models has remained underexplored. In this work, we use knowledge distillation from the event domain to the RGB domain so as to exploit the rich evaluation toolkit available in the RGB domain and systematically dissect this inductive bias. Our experiments show that distillation from the event domain induces, in the RGB domain, color invariance, shape bias, and robustness to high-frequency noise. We identify the underlying mechanism as the model suppressing its dependence on high-frequency texture while acquiring a stronger dependence on edge-based object shape. This hypothesis is supported by changes in how color and spatial information are processed at the early layers, together with a spectral trade-off in which robustness to the absence of high-frequency components coexists with vulnerability to contamination of the relied-upon frequency bands and to disruption of geometric structure. We further show that this inductive bias differs from existing robustification methods and that it functions as a useful prior for diverse downstream tasks in which shape and contour information contribute alongside other cues. The code is available at https://github.com/snskysk/event2rgb-distillation .
Chinese Translation
众所周知,在 ImageNet 上训练的卷积神经网络对局部高频纹理表现出强烈偏好,这一归纳偏置导致模型在真实世界环境的分布偏移下鲁棒性脆弱。相比之下,事件相机仅记录场景亮度的变化,因此非常适合捕捉轮廓信息;然而,由于事件域缺乏诊断性基准,事件相机数据在视觉模型中灌输的归纳偏置一直未得到充分研究。在本工作中,我们利用从事件域到 RGB 域的知识蒸馏,以便借助 RGB 域丰富的评估工具体系,系统地剖析这一归纳偏置。我们的实验表明,来自事件域的蒸馏在 RGB 域中诱发了颜色不变性、形状偏置以及对高频噪声的鲁棒性。我们将底层机制识别为:模型在抑制其对高频纹理依赖的同时,获得了对基于边缘的物体形状更强的依赖。这一假设得到早期层中颜色与空间信息处理方式变化的佐证,并伴随一种频谱上的权衡:对高频成分缺失的鲁棒性,与对所依赖频带受污染及几何结构被破坏的易感性共存。我们进一步表明,这一归纳偏置有别于现有的鲁棒性增强方法,并且在形状与轮廓信息可与其他线索共同发挥作用的多样化下游任务中,它能作为一种有用的先验。代码发布于 https://github.com/snskysk/event2rgb-distillation 。
cs.CV / 8 / 2609.30566

Atlases Are Already Inside: Recovering Population Templates from Pretrained Diffusion Models

图谱已内嵌其中:从预训练扩散模型中恢复人群模板
Shi, Jian, Femiani, John, Wonka, Peter
Abstract
We present a new inference-time sampler for diffusion models that gives a pretrained model a capability it was never trained for: constructing the atlas of the population it synthesizes. The sampler converges from every random seed to the population's central anatomy, which we call the \emph{intrinsic atlas}. The advantage is threefold. (1) It requires no retraining. A diffusion model that has already learned a coherent population, including the released ones, yields its atlas in a single inference pass without involving deformable registration. (2) It applies to multiple domains, such as brain MRI, chest X-ray, faces, and 3D shapes. (3) It extends to subpopulations. One age-conditioned model gives an atlas at any age in its training range, and the resulting family reproduces the CSF expansion of healthy aging. Evaluated as a registration target, the intrinsic atlas is best or second-best on every dataset against classical and learned templates, and the most central template on held-out brain MRI cohorts. Atlas construction can be reframed as a byproduct of generative modeling: a diffusion model is a learned representation of population structure, and the atlas is what it already contains.
Chinese Translation
我们提出了一种新的扩散模型推理时采样器,它赋予预训练模型一种从未被训练过的能力:构建其所合成人群的图谱(atlas)。该采样器从任意随机种子出发都会收敛到人群的中心解剖结构,我们将其称为内在图谱(intrinsic atlas)。其优势有三方面。(1)无需重新训练。一个已经学到连贯人群分布的扩散模型——包括已公开发布的模型——只需一次推理即可生成其图谱,且无需涉及可变形配准。(2)适用于多个领域,如脑部MRI、胸部X光、人脸和三维形状。(3)可扩展至子人群。一个以年龄为条件的模型可以生成其训练范围内任意年龄的图谱,且所得图谱系列能够重现健康衰老过程中脑脊液(CSF)的扩张。作为配准目标进行评估时,内在图谱在所有数据集上相对经典模板和基于学习的模板均取得最优或次优的结果,并且在留出的脑部MRI队列上是最中心的模板。图谱构建可以被重新定义为生成建模的副产品:扩散模型是对人群结构的一种习得表示,而图谱正是它已然包含的内容。
cs.CV / 9 / 2609.30595

Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases

动作强制(Action Forcing):通过恢复潜在自运动基在无监督视频上训练世界模型
Sundar, Ashish, Hou, Tiankuo, Fan, Zhong, Luo, Chunbo, Wang, Xiaoyang
Abstract
Synchronised action annotations are needed to train controllable world models and these datasets remain elusive. Existing approaches make use of instrumented platforms with calibrated sensors, costly manual annotation, or latent-action models which lack grounding. We instead turn ordinary unlabelled video into action-supervised training data by recovering (without training) a data-derived egomotion basis. We track pixel displacements across frames and exploit the recurring coherent structure induced by egomotion to obtain grounded control signals directly. Using a method as simple as principal components analysis perform this, we find that the leading components provide signed, scalable, and composable throttle--yaw controls, although the method can recover only motion axes represented in the data. To prevent a high-capacity video DiT from exploiting pixel-level supervision, an online latent critic distils a frozen decoder--tracker--PCA (Principal Components Analysis) teacher without backpropagating through the decoder or tracker. Finally we critique the use of video generation metrics to evaluate WMs and introduce an example of an alternative, reference-free evaluation method. We measure \textit{controllability}, \textit{plausibility}, \textit{conjuring} (creating objects out of thin air) and \textit{geometric integrity}, revealing failures that conventional video metrics miss. We show that most baselines follow familiar action directions but struggle to reverse or remain stationary. Our model handles both while retaining compositional control and generation quality. Despite backwards actions being less than $1\%$ of our training data, we find that the model learns to reverse, scale its response linearly, and compose throttle with steering, all simply by learning through a grounded action space.
Chinese Translation
训练可控世界模型需要与动作同步的标注数据,然而此类数据集仍然难以获得。现有方法依赖于带有标定传感器的专用采集平台、昂贵的人工标注,或缺乏实际锚定的潜在动作模型。我们则通过恢复一个由数据导出的自运动基(无需训练),将普通的无标注视频转化为动作监督训练数据。我们跨帧追踪像素位移,并利用自运动所产生的重复性连贯结构,直接获得有实际锚定的控制信号。使用主成分分析(PCA)这一简单方法即可完成,我们发现其主成分提供了带符号、可缩放且可组合的油门—偏航控制,尽管该方法只能恢复数据中表征的运动轴。为防止高容量视频扩散Transformer(DiT)利用像素级监督,我们采用在线潜在评论家(critic)从冻结的解码器—追踪器—PCA教师模型中蒸馏知识,而无需通过解码器或追踪器进行反向传播。最后,我们批评了使用视频生成指标来评估世界模型的做法,并引入一种替代性的无参考评估方法示例。我们衡量了可控性、合理性、凭空生成物体(conjuring)以及几何完整性,揭示了传统视频指标所忽略的失效模式。结果表明,大多数基线模型能够遵循熟悉的动作方向,但难以倒车或保持静止。我们的模型则能同时处理这两种情况,同时保持组合控制能力和生成质量。尽管倒车动作在训练数据中占比不到1%,我们发现模型仍学会了倒车、线性缩放响应幅度,并将油门与转向进行组合——这一切都仅仅通过在具有实际锚定的动作空间中学习而实现。
cs.CV / 10 / 2609.30609

MVAgent: Multi-Agent Video Generation via Consistent Condition Construction and Shot-Level Policy Optimization

MVAgent:基于一致性条件构建与镜头级策略优化的多智能体视频生成
Kong, Xiangyu, Zhou, Wenjie, Tian, Fengping, Fang, Lihua, Sun, Haoqin, Lyu, Chenyang, Wang, Longyue, Luo, Weihua
Abstract
Multi-shot agentic video generation requires consistent character appearance, stable spatial layout across camera angles, and continuous character state between shots. When every shot is a separate request to a frozen generator, repeated text does not determine appearance, layout or state. We therefore recast the problem as condition construction and present MVAgent, a multi-agent pipeline whose agents collaborate through typed conditioning inputs. Because an environment image shows one viewpoint, a Spatial Grounding agent samples views from generated camera-traversal clips and anchors each shot to the view matching its framing. As generated shots drift from the plan, an Observer records how each shot ends in a continuity memory, from which a Transition agent builds character action and spatial references for the next shot. An Orchestrator composes these inputs into each request. Since a request reveals its effect only after rendering, we train it by agentic reinforcement learning with Trunk-GDPO, which compares rendered candidates at every shot rather than once per video and continues the best as the trunk. With generator and judges frozen, MVAgent attains the highest cross-shot consistency and narrative-planning quality among the compared methods on ViMax-Bench and is preferred over the strongest agentic baseline in human evaluation.
Chinese Translation
多镜头智能体视频生成需要一致的角色外观、跨镜头角度的稳定空间布局,以及镜头之间连续的角色状态。当每个镜头都是对冻结生成器的独立请求时,重复的文本描述并不能决定外观、布局或状态。因此,我们将该问题重新定义为条件构建问题,并提出了MVAgent——一个多智能体流水线,其各智能体通过类型化的条件输入进行协作。由于环境图像仅展示一个视角,空间锚定(Spatial Grounding)智能体从生成的相机运动轨迹片段中采样多个视角,并将每个镜头锚定到与其取景相匹配的视角上。随着生成镜头逐渐偏离计划,观察者(Observer)智能体将每个镜头的结束状态记录到连续性记忆中,转场(Transition)智能体据此为下一个镜头构建角色动作和空间参考。编排者(Orchestrator)智能体将这些输入组合成每个请求。由于请求的效果只有在渲染后才能显现,我们采用智能体强化学习方法Trunk-GDPO对其进行训练:该方法在每个镜头处都比较渲染出的候选结果,而非在整个视频完成后只比较一次,并将最优候选延续为主干。在生成器和评判器均冻结的条件下,MVAgent在ViMax-Bench上取得了所比较方法中最高的跨镜头一致性与叙事规划质量,并在人类评估中优于最强的智能体基线方法。
cs.CV / 11 / 2609.30613

MedTokenBudget: Lesion-Preserving Token Routing for Dermoscopic Image Classification

MedTokenBudget:面向皮肤镜图像分类的病灶保持式词元路由方法
Li, Zhexiang
Abstract
Dermoscopy classifiers built on Vision Transformers process all image patches uniformly, although diagnostic evidence is concentrated in the lesion region. Existing token pruning methods reduce tokens using generic saliency or similarity signals, but rarely ask whether the retained subset still contains the lesion. This paper introduces MedTokenBudget, a supervised post-backbone token routing framework that learns to construct compact lesion-enriched representations when auxiliary lesion masks are available. Its Lesion-Aware Token Scoring (LATS) module fuses attention entropy, feature norm, and local feature contrast through a learned scorer, then routes the top-$K$ patches under a target budget. LATS is trained with budget curriculum learning, diversity regularization, attention distillation, and lesion-mask supervision. The trained router is evaluated with a lesion retention rate that directly measures how much ground-truth lesion evidence survives the token budget. On ISIC 2019, mask-supervised LATS consistently outperforms Random and ToMe at headline budgets while retaining substantially more lesion patches. Code is provided for reproducibility, and complete tabulated results are included in the supplementary material.
Chinese Translation
基于视觉Transformer(Vision Transformer)构建的皮肤镜分类器对所有图像块进行统一处理,然而诊断证据往往集中于病灶区域。现有的词元剪枝方法通常基于通用的显著性或相似性信号来减少词元数量,却很少关注保留下来的词元子集是否仍然包含病灶信息。本文提出MedTokenBudget,这是一种有监督的主干网络后词元路由框架,能够在辅助病灶掩码可用的情况下,学习构建紧凑且富含病灶信息的表示。其病灶感知词元评分(Lesion-Aware Token Scoring, LATS)模块通过一个可学习的评分器融合注意力熵、特征范数和局部特征对比度,并在目标预算下路由得分最高的前$K$个图像块。LATS通过预算课程学习、多样性正则化、注意力蒸馏以及病灶掩码监督进行训练。训练好的路由器采用病灶保留率进行评估,该指标直接衡量在词元预算约束下有多少真实病灶证据得以保留。在ISIC 2019数据集上,采用掩码监督的LATS在主要预算设置下始终优于Random和ToMe方法,同时保留了显著更多的病灶图像块。我们提供了代码以保证可复现性,完整的表格化结果见补充材料。
cs.CV / 12 / 2609.30647

Conditional Predictive Sufficient Statistics for Visual Representation Learning

用于视觉表示学习的条件预测充分统计量
Hong, Yuzhou
Abstract
A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-factor model of image patches, the mutual information between the past and the next patch equals the information the past carries about the shared factor, up to a remainder that the next patch itself fails to reveal. Predicting the next patch embedding with a cosine loss is maximum likelihood for a von Mises-Fisher model of that embedding's direction, and is therefore a tractable surrogate for the predictive information. The same population loss is also minimized by a constant embedding, so stop-gradient does not by itself select the sufficient statistic; it only blocks the symmetric gradient that implements the constant solution in one step. The regression target is a shallow embedding, which forces the network output back into that shallow range and leaves the sufficient statistic in intermediate blocks. Small causal Transformers on MNIST and CIFAR-10 are used as diagnostics, not as a leaderboard. On MNIST the future shift and the stop-gradient move probe accuracy by tens of points, and the CPSS readout peaks before the output. On CIFAR-10, with the same short budget and no augmentation, every objective lands near a linear classifier on pixels. What still matches the derivation is the geometry: the CPSS output is a worse readout than its best intermediate block, next-pixel regression does not pay that penalty, and removing the stop-gradient collapses the effective rank of the embedding even when the pretext loss looks perfect.
Chinese Translation
有用的视觉表示是观测历史的一个统计量,它保留与未来共享的潜在因子,并丢弃图像块私有的噪声。我们将这一要求形式化为条件预测充分统计量(CPSS)。在图像块的共享因子模型下,历史与下一个图像块之间的互信息等于历史所携带的关于共享因子的信息,仅相差下一个图像块本身无法揭示的余项。使用余弦损失预测下一个图像块的嵌入,等价于对该嵌入方向的von Mises-Fisher模型进行极大似然估计,因此是预测信息的一个可处理的替代目标。常数嵌入同样能最小化该总体损失,因此停梯度(stop-gradient)本身并不能选出充分统计量;它只是阻止了对称梯度一步实现常数解。回归目标是一个浅层嵌入,这迫使网络输出退回该浅层范围,从而使充分统计量留在中间层。我们在MNIST和CIFAR-10上使用小型因果Transformer作为诊断工具,而非排行榜竞争。在MNIST上,未来偏移和停梯度使探针准确率变动数十个百分点,且CPSS读出在输出层之前达到峰值。在CIFAR-10上,在同样短的训练预算且无数据增强的条件下,所有目标函数都仅接近像素上的线性分类器水平。仍然与推导相符的是几何性质:CPSS输出作为读出特征劣于其最佳的中间层,下一像素回归则不付出这一代价,而移除停梯度会使嵌入的有效秩坍缩,即使代理任务的损失看起来非常完美。
cs.CV / 13 / 2609.30667

StarWM: Self-Supervised Trained Attention Routing for Robust World Models

StarWM:面向鲁棒世界模型的自监督训练注意力路由
Zhang, Zeqiang, Wurzberger, Fabian, Otte, Maximilian, Schmid, Daniel, Gottwald, Sebastian, Raulf, Arne Peter, Braun, Daniel Alexander
Abstract
A robust world model must strike the balance between faithfully capturing environmental dynamics and abstracting away from irrelevant content. While reconstruction-based world models ensure faithful supervision, they misallocate representational capacity by pixel area rather than dynamics relevance for visual tasks, which can cause task-irrelevant content to dominate the learned representation. Alternatively, reconstruction-free methods avoid this bias but risk discarding possibly relevant information. We propose StarWM, which uses a cross-attention module trained on self-supervised dynamics to decide where reconstruction applies. A dual-stream decoder then restricts reconstruction to the attended regions, with stop-gradient barriers preventing interference between the two objectives. These components allows reconstruction to supervise the visual content of attended regions without contaminating the latent with non-predictive information. On DeepMind Control with dynamic video backgrounds, default (reward-free) StarWM achieves the strongest performance under random-frame distractors and substantially outperforms reconstruction-based baselines under sequential video. In addition, its reward-augmented variant matches or exceeds reconstruction-free methods on sequential video, achieving the highest overall return across all distractor regimes. Mechanistic probing confirms StarWM preserves state attributes with near-perfect fidelity through long-horizon imagination while systematically discarding distractors.
Chinese Translation
一个鲁棒的世界模型必须在忠实捕捉环境动态与抽象掉无关内容之间取得平衡。基于重构的世界模型虽然能确保忠实的监督信号,但它们按像素面积而非动态相关性来分配表征能力,这可能导致与任务无关的内容主导所学习的表征。相比之下,无重构方法虽然避免了这种偏差,但存在丢弃可能相关信息的风险。我们提出StarWM,它利用一个通过自监督动态训练的交叉注意力模块来决定在何处进行重构。随后,双流解码器将重构限制在注意力区域,并通过停止梯度屏障防止两个目标之间的相互干扰。这些组件使重构能够监督注意力区域的视觉内容,同时避免潜变量被非预测性信息污染。在带有动态视频背景的DeepMind Control基准上,默认(无奖励)的StarWM在随机帧干扰物条件下取得最强性能,并在连续视频条件下大幅超越基于重构的基线方法。此外,其奖励增强变体在连续视频条件下与无重构方法持平或更优,并在所有干扰物条件下取得最高的总体回报。机制性探测实验证实,StarWM在长时程想象中能以近乎完美的保真度保留状态属性,同时系统性地丢弃干扰物。
cs.CV / 14 / 2609.30682

Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding

面向超高分辨率科学图像理解的结构引导掩码自编码器
Zhang, Enzhi, Wu, Du, Zhong, Rui, Ma, Cong, Lyngaas, Isaac, Ziabari, Amir Koushyar, Wang, Xiao, Chen, Peng, Luo, Tao, Endo, Toshio, Shoji, Fumiyoshi, Sato, Kento, Uesugi, Kentaro, Nonoyama, Takayuki, Kiyama, Ryuji, Yoshida, Masahiro, Tezuka, Masaru, Ishikawa, Tetsuya, Matsuoka, Satoshi, Munetomo, Masaharu, Wahib, Mohamed
Abstract
Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O(N^2)$ attention impractical. We propose SGMA, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. SGMA couples two components: a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence, and a structure-conditioned masking process that biases reconstruction toward spatially informative regions. To stabilize this process across scales, we introduce Damped Accumulation (DA), which aggregates signal-dependent responses across the tree into a structure canvas used to guide masking. The resulting pre-training task preserves fine microstructure while remaining compatible with standard ViT encoders and MAE-style reconstruction. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, SGMA consistently outperforms MAE baselines. It achieves 95.68% Dice on the 8K x 8K x 28K SpringXCT dataset, improving over the same-architecture MAE baseline by +13.00 points, and 83.21% Dice on the 32K^2 WSI PAIP dataset, improving by +16.84 points, while providing up to a 24.8x inference speedup.
Chinese Translation
基于视觉Transformer的自监督预训练(包括掩码自编码器 MAE)难以应用于吉像素级的科学图像。随机掩码与科学数据结构化、多尺度的形态特征匹配不佳,而均匀分词会产生过长的序列,使得 O(N^2) 复杂度的注意力机制难以实现。我们提出了 SGMA,一种面向超高分辨率科学图像的结构引导掩码自编码框架。SGMA 结合两个组件:一是内容自适应的四叉树分词器,将吉像素图像压缩为固定长度序列;二是结构条件化的掩码过程,将重建偏向于空间信息丰富的区域。为在不同尺度上稳定这一过程,我们引入了阻尼累积(Damped Accumulation, DA),它将树结构中依赖信号的响应聚合为一张结构画布,用于引导掩码。由此产生的预训练任务在保留精细微观结构的同时,保持与标准 ViT 编码器和 MAE 式重建的兼容性。在电子显微镜、全切片光学显微镜和 X 射线 CT 数据集上,SGMA 始终优于 MAE 基线。在 8K x 8K x 28K 的 SpringXCT 数据集上,它达到 95.68% 的 Dice 分数,比相同架构的 MAE 基线提升 +13.00 个百分点;在 32K^2 的 WSI PAIP 数据集上达到 83.21% 的 Dice 分数,提升 +16.84 个百分点,同时提供高达 24.8 倍的推理加速。
cs.CV / 15 / 2609.30698

MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning

MM-VeriAgent:基于强化学习学习使用丰富工具来验证多模态虚假信息
Li, Peipei, Xia, Shuhan, Liu, Shengyang, Li, Zekun, He, Ran
Abstract
Real-world multimodal misinformation often involves mixed forgery sources, requiring sample-specific detection strategies. Existing tool-augmented methods rely on predefined workflows or inference-time planning, limiting adaptability or increasing inference cost. To address this issue, we introduce \textbf{MM-VeriAgent}, which learns to verify mixed-source multimodal misinformation with tools. We first build \textbf{MM-VeriTools}, a specialized toolkit for misinformation detection agents. By benchmarking various candidate models and methods on the sub-tasks required by mixed-source detection, we select the strongest for textual, visual, and cross-modal forgery analysis and encapsulate them as callable tools with a unified interface. On top of this toolkit, we train the LVLM agent with reinforcement learning to teach it how to use these tools to better solve mixed-source detection. Since many of the tools are specialized models whose online execution at every rollout severely limits RL efficiency, we further introduce \textbf{Tool-Execution Cache}, which pre-executes candidate tool calls and reuses their cached outputs during training. This preserves multi-step rollouts while reducing online tool execution, largely improving the training efficiency.Experiments on MMFakeBench demonstrate substantial accuracy gains over the base model without explicit tool search at inference time. Ablation and efficiency analyses further validate the learned tool-use policy and show that Tool-Execution Cache reduces online tool executions during training.
Chinese Translation
现实世界中的多模态虚假信息往往涉及混合伪造来源,需要针对样本的定制化检测策略。现有的工具增强方法依赖于预定义的工作流程或推理时规划,这限制了适应性或增加了推理成本。为解决这一问题,我们提出了 MM-VeriAgent,它学习使用工具来验证混合来源的多模态虚假信息。我们首先构建了 MM-VeriTools,一个面向虚假信息检测智能体的专用工具包。通过在混合来源检测所需的各子任务上对多种候选模型和方法进行基准测试,我们为文本、视觉以及跨模态伪造分析分别选出最强的方案,并将其封装为具有统一接口的可调用工具。在此工具包的基础上,我们使用强化学习训练大型视觉语言模型(LVLM)智能体,教会它如何使用这些工具以更好地解决混合来源检测问题。由于其中许多工具是专用模型,在每次 rollout 时在线执行会严重限制强化学习的效率,我们进一步提出了工具执行缓存,预先执行候选工具调用并在训练过程中复用其缓存输出。这既保留了多步 rollout,又减少了在线工具执行,大幅提升了训练效率。在 MMFakeBench 上的实验表明,该方法相较于基础模型取得了显著的准确率提升,且在推理时无需显式的工具搜索。消融实验和效率分析进一步验证了所学到的工具使用策略,并表明工具执行缓存减少了训练期间的在线工具执行次数。
cs.CV / 16 / 2609.30703

SAGE: Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion

SAGE:基于源锚定引导与频率均衡的层次化RGB-T对齐与融合
Li, Timing, Sun, Yiming, Tao, Boan, Gao, Xiyuan, Cao, Haifang, Zhu, Pengfei
Abstract
Spatial misregistration and cross-modal discrepancies often cause ghosting, structural blurring, and content imbalance in RGB-T fusion. Existing methods typically decouple appearance adaptation, geometric alignment, and information fusion, limiting dependency propagation across stages. We propose Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion (SAGE), a unified framework integrating frequency equalization, hierarchical alignment, and subband fusion. SAGE employs invertible joint encoding and source-specific low-frequency modulation to derive structural and gain guidance while preserving source information. Hierarchical frequency collaborative alignment estimates global affine geometry from low-frequency approximations and transfers geometric and contextual cues to high-frequency correlation reasoning for reliability-aware residual refinement. Guided subband fusion jointly aggregates the aligned frequency coefficients under propagated source and alignment guidance, coordinates complementary low- and high-frequency information, and reconstructs the fused image through the inverse wavelet transform. Extensive experiments on RGB-T datasets with real-world and synthetic misalignments demonstrate consistently competitive performance in alignment and fusion, validating the effectiveness of source-anchored guidance for weakly registered RGB-T images.
Chinese Translation
空间错位配准和跨模态差异常常导致RGB-T融合中出现鬼影、结构模糊和内容失衡。现有方法通常将外观适配、几何对齐和信息融合解耦处理,限制了跨阶段的依赖传播。我们提出了基于频率均衡的源锚定引导层次化RGB-T对齐与融合框架(SAGE),这是一个集频率均衡、层次化对齐和子带融合于一体的统一框架。SAGE采用可逆联合编码和源特定的低频调制来导出结构引导和增益引导,同时保留源信息。层次化频率协同对齐从低频近似中估计全局仿射几何,并将几何和上下文线索传递给高频相关性推理,以进行可靠性感知的残差优化。引导式子带融合在传播的源引导和对齐引导下联合聚合对齐后的频率系数,协调互补的低频与高频信息,并通过逆小波变换重建融合图像。在包含真实与合成错位的RGB-T数据集上的大量实验表明,该方法在对齐和融合任务中均保持一致的竞争力,验证了源锚定引导对弱配准RGB-T图像的有效性。
cs.CV / 17 / 2609.30708

Combining General and Domain-Specific Pretext Tasks for Brain MR Image Segmentation

结合通用与领域特定前置任务的脑部磁共振图像分割
Nasser, Tasneem, Schmid, Susanne, Souza, Roberto, El-Sheimy, Naser
Abstract
A key challenge in medical image analysis is the scarcity of large annotated datasets for specific populations and diseases. As deep learning models rely heavily on labeled data, effective transfer learning strategies are needed to reduce the dependence on manual annotations. Self-supervised learning has emerged as a promising approach for developing foundation models by enabling the learning of transferable feature representations from large-scale unlabeled medical imaging datasets. In this study, we investigate voxel-level brain age prediction as a domain-specific self-supervised pretext task and compare it with image inpainting, a widely used non-domain-specific alternative. We further propose a multitask self-supervised pretraining framework that jointly optimizes both objectives to learn complementary neuroimaging representations. The pretrained models are evaluated on three downstream magnetic resonance image segmentation tasks: multiple sclerosis lesion segmentation, ischemic stroke lesion segmentation, and cortical brain structure segmentation. Overall, the proposed multitask pretraining framework consistently outperformed the single-task pretrained models and training from scratch across most experimental settings, demonstrating the benefit of combining domain-specific and general self-supervised learning pretext tasks for the development of generalizable neuroimaging foundation models.\ Code Availability: The source code used in this study is publicly available at https://github.com/TasneemN/Combining-General-and-Domain-Specific-Pretext-Tasks-for-Brain-MR-Image-Segmentation/
Chinese Translation
医学图像分析的一个关键挑战是缺乏针对特定人群和疾病的大规模标注数据集。由于深度学习模型高度依赖标注数据,因此需要有效的迁移学习策略来减少对手工标注的依赖。自监督学习通过从大规模无标注医学影像数据集中学习可迁移的特征表示,已成为构建基础模型的一种有前景的方法。在本研究中,我们将体素级的脑年龄预测作为领域特定的自监督前置任务进行研究,并将其与图像修复(inpainting)这一广泛使用的非领域特定替代方法进行比较。我们进一步提出了一种多任务自监督预训练框架,该框架联合优化上述两个目标,以学习互补的神经影像特征表示。预训练模型在三个下游磁共振图像分割任务上进行了评估:多发性硬化病变分割、缺血性脑卒中病变分割以及大脑皮层结构分割。总体而言,所提出的多任务预训练框架在大多数实验设置中始终优于单任务预训练模型和从头训练的模型,这表明将领域特定与通用自监督学习前置任务相结合,有助于开发具有泛化能力的神经影像基础模型。代码可用性:本研究所使用的源代码已在 https://github.com/TasneemN/Combining-General-and-Domain-Specific-Pretext-Tasks-for-Brain-MR-Image-Segmentation/ 公开。
cs.CV / 18 / 2609.30709

VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control

VLALight:面向应急感知交通信号控制的轻量级视觉-语言-动作模型
Jiang, Kemou, Wang, Maonan, Zou, Xingchen, Zhu, Jiayue, Fu, Yuhang, Wang, Sicheng, Chen, Xi, Chen, Yirong, Cui, Zhiyong
Abstract
Traffic signal control (TSC) is essential for mitigating urban congestion. Recent advances in vision-language models (VLMs) enable richer interpretation of intersection scenes, opening new opportunities for visual-context-aware TSC. However, the loose coupling and repeated information conversion between modules can lead to the loss of fine-grained visual details, while sequential inference introduces substantial latency. To address these limitations, we propose VLALight, a lightweight end-to-end vision-language-action framework that directly maps intersection observations and signal-phase information to discrete signal actions. To handle the multi-view nature of TSC, VLALight combines multiple directional camera views into a unified visual input and uses textual instructions to establish their correspondence with traffic movements and signal phases. This design enables direct action prediction with a compact 0.5 B-parameter model, without intermediate image-to-text descriptions or handcrafted traffic-state representations. Experiments show that VLALight delivers the best emergency-vehicle service of all compared methods, reducing pooled emergency waiting time by 21.1% over the cascaded VLMLight while running in real time on local hardware and generalizing to unseen intersection topologies and traffic-flow patterns.
Chinese Translation
交通信号控制(Traffic Signal Control, TSC)对于缓解城市交通拥堵至关重要。视觉-语言模型(Vision-Language Models, VLMs)的最新进展使得对交叉路口场景的更丰富解读成为可能,为具备视觉上下文感知能力的交通信号控制开辟了新的机遇。然而,模块之间的松耦合以及反复的信息转换可能导致细粒度视觉细节的丢失,而顺序推理则会引入较大的延迟。针对这些局限性,我们提出了VLALight——一个轻量级的端到端视觉-语言-动作框架,它将交叉路口观测信息和信号相位信息直接映射为离散的信号动作。为应对交通信号控制的多视角特性,VLALight将多个方向的摄像头视角合并为统一的视觉输入,并利用文本指令建立视觉输入与交通流向及信号相位之间的对应关系。这一设计使得仅用0.5 B参数的紧凑模型即可实现直接的动作预测,无需中间的图像到文本描述或手工设计的交通状态表示。实验表明,VLALight在所有对比方法中提供了最佳的应急车辆服务水平,与级联式VLMLight相比,将汇总的应急车辆等待时间降低了21.1%,同时可在本地硬件上实时运行,并能泛化到未见过的交叉路口拓扑结构和交通流模式。
cs.CV / 19 / 2609.30722

TrafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation

TrafficImag:反事实路边交通视频生成基准
Li, Xiangyu, Wang, Tianyi, Dou, Zhihao, Claudel, Christian, Guo, Zhaomiao
Abstract
Existing roadside traffic datasets support perception, forecasting, and visual question answering, but they do not evaluate counterfactual video generation, in which a selected actor is modified and the generated future should remain consistent with road topology and unrelated traffic. We introduce TrafficImag, the first benchmark for counterfactual roadside traffic video generation. TrafficImag combines a large-scale roadside dataset (9,022 annotated images, 7,043 deduplicated video clips, and 31,145 actor-centered history-future samples) with an executable protocol that supports behavior reasoning, intervention-aware image editing, and conditional video generation. Each intervention is represented as an actor-level program describing the target actor, intended behavior, legal route, interaction order, and temporal constraints, enabling a unified evaluation interface across heterogeneous foundation models. TrafficImag evaluates four complementary validity dimensions: initial-state correctness, route and behavior validity, interaction consistency, and non-target preservation, and considers an end-to-end counterfactual successful only when all four are satisfied. Across state-of-the-art foundation models, the strongest reasoner reaches 80.4% macro F1, the complete condition interface raises end-to-end success from 23.3% to 55.0% for the best generator. Oracle studies further show that conditional video execution is the primary remaining bottleneck. TrafficImag provides a reproducible benchmark for evaluating and diagnosing counterfactual traffic video generation beyond perceptual video quality.
Chinese Translation
现有的路边交通数据集支持感知、预测和视觉问答任务,但无法评估反事实视频生成。在反事实视频生成中,需要对选定的交通参与者进行修改,同时生成的未来画面应与道路拓扑及无关交通保持一致。我们提出 TrafficImag,这是首个面向反事实路边交通视频生成的基准。TrafficImag 将大规模路边数据集(包含 9,022 张标注图像、7,043 个去重视频片段和 31,145 个以交通参与者为中心的历史-未来样本)与一个可执行的评测协议相结合,该协议支持行为推理、具备干预感知的图像编辑以及条件视频生成。每次干预均以交通参与者级别的程序表示,描述目标参与者、预期行为、合法路线、交互顺序和时间约束,从而为异构基础模型提供统一的评估接口。TrafficImag 从四个互补的有效性维度进行评估:初始状态正确性、路线与行为有效性、交互一致性以及非目标保持,只有当四个维度全部满足时才判定端到端反事实生成成功。在对当前最先进的基础模型的评测中,最强的推理模型达到 80.4% 的宏平均 F1;借助完整的条件接口,最佳生成模型的端到端成功率从 23.3% 提升至 55.0%。Oracle 研究进一步表明,条件视频执行仍是当前最主要的瓶颈。TrafficImag 为评估和诊断超越感知层面视频质量的反事实交通视频生成提供了一个可复现的基准。
cs.CV / 20 / 2609.30724

EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection

EviDETR:面向时刻检索与高光检测的查询相关时间证据保留方法
Sun, Haoran, Li, Yufan, Zhang, Qichen, Zhao, Haoran, Wang, Shuqi
Abstract
Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A Temporal Top-2 Mixture-of-Experts (TTop2MoE) decoder performs query-adaptive refinement via sparse expert routing. MR-to-HD (MR2HD) fusion transfers span-level retrieval evidence to clip-level highlight prediction through confidence-weighted multi-scale aggregation. Using CLIP+SlowFast features, EviDETR achieves 69.29 R1@0.5, 54.77 R1@0.7, and 48.41 Avg. mAP for moment retrieval on QVHighlights, together with 41.83 HD-mAP and 68.33 HIT@1. Strong results on TACoS and Charades-STA further demonstrate cross-dataset transferability.
Chinese Translation
视频时刻检索与高光检测的联合任务需要在估计片段级显著性的同时,识别与查询相关的时间片段,然而基于 DETR 风格的流程在编码、解码及跨任务预测的整个过程中并未显式保留与查询相关的证据。我们提出了 EviDETR,一个包含三个组件的证据保留框架。语义感知特征重加权(Semantic-aware Feature Reweighting, SFR)通过显著性估计和跨模态交互增强与查询相关的片段表征。时间 Top-2 混合专家解码器(Temporal Top-2 Mixture-of-Experts, TTop2MoE)通过稀疏专家路由实现查询自适应的精炼。MR-to-HD(MR2HD)融合通过置信度加权的多尺度聚合,将跨度级检索证据传递到片段级高光预测。基于 CLIP+SlowFast 特征,EviDETR 在 QVHighlights 上的时刻检索达到 69.29 的 R1@0.5、54.77 的 R1@0.7 和 48.41 的平均 mAP,同时取得 41.83 的 HD-mAP 和 68.33 的 HIT@1。在 TACoS 和 Charades-STA 上的优异结果进一步证明了其跨数据集的迁移能力。
cs.CV / 21 / 2609.30728

Learning Polarization Image Restoration with General Restoration Priors

基于通用复原先验的偏振图像复原学习
Li, Chenggong, Liu, Jinhao, Wu, Caiyun, Luo, Yidong, Zhang, Junchao, Yang, Degui
Abstract
Polarization imaging captures distinctive surface and geometric cues that benefit a wide range of vision tasks. However, real-world polarization acquisition is often affected by multiple coupled degradations, making image restoration essential for practical polarization vision. Existing methods are largely tailored to specific degradations and remain constrained by the limited scale and quality of polarization data. To address these limitations, we develop an all-in-one polarization restoration framework for diverse and composite degradations. We first study the impact of different polarization representations on restoration performance and identify the normalized Stokes representation as an effective choice for separating intensity and polarization information. Accordingly, we devise a dual-branch architecture that separates intensity and polarization modeling. To overcome the limitations of polarization-specific training, the intensity branch leverages pretrained general restoration priors and a mixture-of-experts extension for composite degradations, while its restoration knowledge is adaptively distilled into the symmetric polarization branch via a cross-domain feature transform. In addition, we establish a composite-degradation polarization benchmark to support all-in-one restoration research. Extensive experiments on public datasets and our proposed benchmark demonstrate the effectiveness of the proposed method.
Chinese Translation
偏振成像能够捕获独特的表面与几何信息,可广泛受益于各类视觉任务。然而,真实场景中的偏振图像获取往往受到多种耦合退化的影响,因此图像复原对实际的偏振视觉应用至关重要。现有方法大多针对特定退化类型设计,且受限于偏振数据的规模与质量。为解决这些局限,我们构建了一个面向多样化和复合退化的all-in-one偏振图像复原框架。我们首先研究了不同偏振表示对复原性能的影响,并发现归一化Stokes表示是分离强度信息与偏振信息的有效选择。据此,我们设计了一种双分支架构,分别对强度与偏振进行建模。为克服偏振专用训练的局限,强度分支利用预训练的通用复原先验,并通过混合专家(mixture-of-experts)扩展来处理复合退化;同时,通过跨域特征变换,将强度分支的复原知识自适应地蒸馏到对称的偏振分支中。此外,我们建立了一个复合退化偏振基准数据集,以支持all-in-one复原研究。在公开数据集和我们提出的基准上进行的大量实验证明了所提方法的有效性。
cs.CV / 22 / 2609.30733

Amplify What You Gaze At: Target Saliency Boosting in Text-to-Image Generation

放大你所注视的:文本到图像生成中的目标显著性增强
Dang, Shengqi, Yu, Zhengxi, Han, Feilin, Lan, Xingyu, Cao, Nan
Abstract
Text-to-image generation has advanced in controlling what, where, and how objects appear, yet how visual attention is distributed among objects remains largely unexplored. In this paper, we introduce Target Saliency Boosting, a new task aimed at boosting the visual saliency of a specific object during text-to-image generation without requiring any visual priors. Our key insight is that visual saliency is inherently relative: boosting the saliency of a target object also depends on the global saliency distribution across all objects in the scene. Based on this insight, we propose GazeME, a lightweight framework that uses saliency-marked prompts, inserting learnable marker tokens around object descriptions to indicate which objects to visually emphasize or suppress. To learn these markers, we construct a saliency-semantics dataset that associates objects in image--prompt pairs with object-level saliency scores, and propose Saliency Prior Marker Activation (SPMA), a saliency-aware stochastic marker activation strategy that exploits relative saliency relationships for robust training. During inference, GazeME automatically inserts appropriate markers into the prompt, thereby directly enhancing the visual saliency of the target object. Extensive experiments demonstrate that GazeME effectively boosts target saliency while preserving both semantic alignment and image quality.
Chinese Translation
文本到图像生成技术已经在控制物体出现的内容、位置和方式方面取得了进展,然而视觉注意力如何在物体之间分布仍未被充分探索。在本文中,我们提出了目标显著性增强(Target Saliency Boosting)这一新任务,旨在文本到图像生成过程中提升特定物体的视觉显著性,且无需任何视觉先验。我们的核心洞察是:视觉显著性本质上是相对的——提升目标物体的显著性还依赖于场景中所有物体的全局显著性分布。基于这一洞察,我们提出了 GazeME,一个轻量级框架,它使用显著性标注提示,在物体描述周围插入可学习的标记词元(marker tokens),以指示哪些物体需要被视觉上强调或抑制。为了学习这些标记,我们构建了一个显著性-语义数据集,将图像-提示对中的物体与物体级显著性分数相关联,并提出显著性先验标记激活策略(Saliency Prior Marker Activation, SPMA),这是一种利用相对显著性关系的、显著性感知的随机标记激活策略,以实现稳健训练。在推理阶段,GazeME 会自动在提示中插入合适的标记,从而直接增强目标物体的视觉显著性。大量实验表明,GazeME 在有效提升目标显著性的同时,保持了语义一致性和图像质量。
cs.CV / 23 / 2609.30741

From Mono to Stereo: Accelerating Binocular Gaussian Splatting via Reprojection and Selective Patching

从单目到双目:基于重投影与选择性补丁的双目高斯泼溅加速方法
Zhu, Hongfei, Zhou, Ling
Abstract
Binocular rendering requires two nearby views of the same scene and therefore repeats substantial visibility and shading work. We present a 2D Gaussian Splatting (2DGS) pipeline that fully renders a dominant-eye RGB image and an alpha-weighted depth proxy, reprojects that image to the affiliated eye, and repairs uncovered pixels. Small interior gaps are interpolated, whereas larger disoccluded regions are identified as regions of interest (ROIs) and selectively re-rendered. The depth proxy reuses the alpha-blending weights computed during dominant-eye rasterization, avoiding a separate depth-rendering pass. An adaptive ROI generator localizes the required updates using reprojected image boundaries and optional connected center-hole detection. On DTU, Tanks and Temples, and MipNeRF-360, the method reduces the measured time of a sequential two-pass binocular reference by 15.5\% to 28.8\% and peak GPU memory by 6\% to 11\%. The corresponding affiliated-eye quality degradation is at most 1.3 dB PSNR, 0.02 SSIM, and 0.02 LPIPS, representing a measurable trade-off that requires application-specific perceptual validation. These results establish a practical efficiency-quality trade-off for controlled static-scene stereo rendering and motivate future evaluation under continuous motion and on physical VR hardware.
Chinese Translation
双目渲染需要同一场景的两个相近视角,因此会重复大量的可见性计算与着色工作。我们提出了一种二维高斯泼溅(2D Gaussian Splatting, 2DGS)流水线:该流水线完整渲染主视眼的RGB图像以及一个经alpha加权的深度代理图,将该图像重投影至附属眼,并对未被覆盖的像素进行修复。较小的内部空洞采用插值填补,而较大的去遮挡区域则被识别为感兴趣区域(ROIs)并进行选择性重渲染。深度代理复用了主视眼光栅化过程中已计算的alpha混合权重,从而避免了单独的深度渲染过程。自适应ROI生成器利用重投影后的图像边界以及可选的连通中心孔洞检测来定位所需的更新。在DTU、Tanks and Temples和MipNeRF-360数据集上,与顺序执行两次渲染的双目参考方法相比,该方法将实测耗时降低了15.5%至28.8%,峰值GPU显存降低了6%至11%。对应的附属眼质量下降最多为1.3 dB PSNR、0.02 SSIM和0.02 LPIPS,这一可衡量的权衡仍需针对具体应用进行感知层面的验证。这些结果为受控静态场景的立体渲染建立了一种实用的效率-质量权衡方案,并为未来在连续运动条件下以及实际VR硬件上的评估提供了研究方向。
cs.CV / 24 / 2609.30755

Training-Free Bottleneck Width Planning for Convolutional Autoencoders

面向卷积自编码器的免训练瓶颈宽度规划
Guo, Guannan
Abstract
Multiscale Spectral Rate-Distortion (MS-SRD) estimates the bottleneck channels required at user-supplied spatial cuts from training images and a normalized mean-squared error (NMSE) bound, without fitting a neural network. Its covariance-tail rule is exact for shared linear block-convolutional autoencoders under squared error. A nested-scale dominance result motivates reporting the activation-parameter Pareto frontier alongside the minimum-latent candidate. At NMSE <= 0.01 on thirteen grayscale datasets, its latent-size prediction has 0.84% mean absolute percentage error against nonlinear patch-autoencoder boundaries; ten predictions are exact and the remaining three differ by one channel. In a four-dataset deployable comparison, MS-SRD matches all retrospective external widths and all four selected models pass, without training a selector; a 46-fit validation grid and four Least-Volume fits each pass on two datasets. In a skip-closed U-shaped autoencoder at the same bound, five predictions are exact, nine are within one channel, and every failing prediction is one channel short. Experiments at looser bounds show progressively larger nonlinear savings.
Chinese Translation
多尺度谱率失真方法(Multiscale Spectral Rate-Distortion, MS-SRD)根据训练图像和归一化均方误差(NMSE)上界,估计用户指定空间截断处所需的瓶颈通道数,而无需拟合神经网络。在平方误差准则下,其协方差尾部规则对于共享线性分块卷积自编码器是精确的。一项嵌套尺度占优结果启发了在报告最小隐变量候选的同时,一并报告激活-参数帕累托前沿。在十三个灰度数据集上,当 NMSE <= 0.01 时,其隐变量尺寸预测相对于非线性图像块自编码器的边界,平均绝对百分比误差为 0.84%;十个预测完全精确,其余三个仅相差一个通道。在四数据集的可部署性比较中,MS-SRD 与所有回顾性外部宽度均匹配,且所有四个选定模型均通过测试,而无需训练任何选择器;一个包含 46 次拟合的验证网格和四次最小体积(Least-Volume)拟合各自在两个数据集上通过。在同一 NMSE 上界下,对于跳跃连接关闭的 U 形自编码器,五个预测完全精确,九个预测相差不超过一个通道,且所有未通过的预测均少一个通道。在更宽松上界下的实验表明,非线性模型可实现的节省随上界放宽而逐渐增大。
cs.CV / 25 / 2609.30758

LLPR: Location-aware learning and physics-based reconstruction for raindrop removal from a single image

LLPR:面向单幅图像雨滴去除的位置感知学习与物理重建方法
He, Zewei, Liu, Xingyu, Luo, Xing, Fu, Guizhong, Chen, Zixuan, Chen, Yu, Li, Jinlei, Lu, Zhe-Ming
Abstract
Raindrops can cause occlusion and distortion in the background scenes due to their adherence to windows or camera lenses. Existing raindrop removal methods concentrate on designing sophisticated CNN or Transformer architectures to recover distorted and missing texture. In this paper, we try to integrate location information and physical model into off-the-shelf CNN or Transformer architectures to help improve their performance. Specifically, we notice that existing methods deploy a preprocessing sub-network to generate a binary or soft mask to indicate the raindrop location, which will increase the network parameters and computational complexity. In contrast, a location-aware learning branch is embedded to teach the encoder in the training phase with the capability of perceiving the position of the raindrops. Note that this location-aware learning branch can be removed during the inference process (achieving performance improvements at no cost). Furthermore, instead of directly reconstructing the raindrop-free image (i.e., background scene), we devise a physics-based reconstruction scheme to first learn the transparency matrix and the raindrop layer. The latent background layer is then reversely derived based on the physical model. By combining the above-mentioned components, we propose our location-aware learning and physics-based reconstruction (LLPR) framework for this challenging ill-posed problem. We also collect a real-world raindrop-degraded image dataset, which is challenging for single-image raindrop removal (SIRR) methods. Extensive experimental results demonstrate the effectiveness and generality of our LLPR framework, achieving superior performance against state-of-the-art SIRR methods. The code will be made available upon acceptance.
Chinese Translation
雨滴附着在窗户或相机镜头上会对背景场景造成遮挡和扭曲。现有的雨滴去除方法专注于设计复杂的CNN或Transformer架构来恢复失真和缺失的纹理。在本文中,我们尝试将位置信息和物理模型集成到现有的CNN或Transformer架构中,以帮助提升其性能。具体而言,我们注意到现有方法部署了一个预处理子网络来生成二值或软掩膜以指示雨滴的位置,这会增加网络参数和计算复杂度。相比之下,我们嵌入了一个位置感知学习分支,在训练阶段使编码器具备感知雨滴位置的能力。值得注意的是,该位置感知学习分支可以在推理阶段被移除(从而在零开销的情况下实现性能提升)。此外,我们没有直接重建无雨滴图像(即背景场景),而是设计了一种基于物理的重建方案,首先学习透射率矩阵和雨滴层,然后基于物理模型反向推导出潜在的背景层。通过结合上述组件,我们提出了针对这一具有挑战性的病态问题的位置感知学习与物理重建(LLPR)框架。我们还收集了一个真实世界的雨滴退化图像数据集,该数据集对单幅图像雨滴去除(SIRR)方法具有挑战性。大量实验结果证明了我们LLPR框架的有效性和普适性,其性能优于当前最先进的SIRR方法。代码将在论文被接收后公开。
cs.CV / 26 / 2609.30761

Timo: $\textbf{T}$aming Mult$\textbf{i}$modal Diffusion Transformer for Human $\textbf{Mo}$tion Generation

Timo:驯服多模态扩散Transformer以实现人体动作生成
Wang, Zhao, Hu, Jiangtao, Yu, Jack, Yu, Tao
Abstract
Most existing human motion generation (HMG) methods use cross-attention modules to inject text semantics, but ignore the importance of bidirectional modeling between motion and text tokens, which limits text comprehension. A straightforward idea is introducing multimodal diffusion transformers (MMDiT), which have shown effective joint text--visual modeling in vision generation, into HMG. However, we find that articulated motion is temporally coherent but weakly correlated across joints, in which directly applying an MMDiT with flow matching produces poorly coordinated and jerky motion. In this work, we propose Timo, a novel kinematics-aware MMDiT framework tailored for HMG. Timo combines fully shared multimodal attention for bidirectional text--motion modeling with flow matching, geometric and rotational-kinematics supervision that compares actual rotations and their changes over time, and a two-stage curriculum progressing from broad motion learning to detailed caption alignment. Further, we construct a benchmark of $40{,}025$ held-out clips from six public datasets spanning diverse actions, assessing six complementary dimensions under a common evaluator and scoring protocol. Our model substantially outperforms state-of-the-art methods in both quantitative and qualitative evaluations. Remarkably, Timo surpasses Kimodo on five of six dimensions, achieving a $40.8$% relative improvement in the average benchmark score. Project page: https://kyfafyd.wang/projects/timo. Demo page: https://timo.kyfafyd.wang.
Chinese Translation
现有的人体动作生成(HMG)方法大多使用交叉注意力模块来注入文本语义,但忽略了动作与文本词元之间双向建模的重要性,从而限制了文本理解能力。一个直接的想法是将多模态扩散Transformer(MMDiT)引入HMG,该模型在视觉生成任务中已展现出有效的文本—视觉联合建模能力。然而,我们发现关节化动作在时间上具有连贯性,但在各关节之间相关性较弱,直接应用基于流匹配(flow matching)的MMDiT会产生协调性差且不流畅的动作。在本工作中,我们提出了Timo,一种专为HMG设计的新型运动学感知MMDiT框架。Timo将用于双向文本—动作建模的全共享多模态注意力与流匹配相结合,并引入几何和旋转运动学监督(比较实际旋转及其随时间的变化),同时采用从粗粒度动作学习到细粒度文本对齐的两阶段课程学习策略。此外,我们构建了一个包含来自六个公开数据集的40,025个保留片段的基准测试,涵盖多样化动作,并在统一的评估器和评分协议下从六个互补维度进行评估。我们的模型在定量和定性评估中均大幅超越现有最先进方法。值得注意的是,Timo在六个维度中的五个维度上超越了Kimodo,在平均基准得分上取得了40.8%的相对提升。项目主页:https://kyfafyd.wang/projects/timo。演示页面:https://timo.kyfafyd.wang。
cs.CV / 27 / 2609.30769

Query-Conditioned Prototype Adaptation for Cross-Domain Few-Shot Learning: Single-Query Inference, Controlled Comparisons, and Failure Modes

面向跨域少样本学习的查询条件化原型自适应:单查询推理、受控比较与失败模式
Karania, Rushab Rasik, Maul, Tomas
Abstract
Cross-domain few-shot learning requires adapting a classifier to a new visual domain from very few labelled examples without target-time parameter updates. We isolate one question: under a fixed global representation, what does joint query-support adaptation contribute to prototype construction? The Within-Instance Prototypical Transformer (WIPT) implements single-query test-time prototype adaptation by jointly transforming one unlabelled query and the labelled support embeddings, then forming query-specific class means. Using a shared frozen ViT-S/16 encoder, miniImageNet source training, and CUB, EuroSAT and ISIC targets, we replicate the key comparisons across five independent training seeds. In 1-shot evaluation, WIPT improves frozen ProtoNet in every run on CUB (+0.21 percentage points) and EuroSAT (+2.07), but decreases ISIC (-0.22). In 5-shot evaluation, ProtoNet remains strongest overall, while WIPT consistently improves a capacity-matched support-only Transformer on ISIC (+0.99). Joint processing of up to five queries yields no reliable accuracy gain; in a head-only 5-shot benchmark, g = 5 reduces analytical attention-token pairs by 73% and peak allocated memory by 29% relative to g = 1, although latency is non-monotonic. Across all target/shot conditions, WIPT changes uncertain ProtoNet decisions far more than confident ones, and rescue/break decomposition accounts for the observed gains and losses. Source-shift and scorer controls further show that the benefit is not universal. Overall, WIPT provides a streaming-compatible form of test-time prototype adaptation that can improve difficult low-shot cross-domain decisions without target-time optimization.
Chinese Translation
跨域少样本学习要求仅利用极少量的标注样本将分类器适配到新的视觉领域,且无需在目标域测试时进行参数更新。本文聚焦于一个问题:在固定全局表示的条件下,查询集与支持集的联合自适应对原型构建有何贡献?实例内原型Transformer(Within-Instance Prototypical Transformer, WIPT)实现了单查询测试时原型自适应:它对单个未标注查询样本与标注支持集嵌入进行联合变换,然后构建针对该查询的类均值原型。使用共享的冻结ViT-S/16编码器、miniImageNet源域训练,以及CUB、EuroSAT和ISIC目标域,我们在五个独立训练随机种子上重复了关键比较实验。在1-shot评估中,WIPT在所有运行中均在CUB(+0.21个百分点)和EuroSAT(+2.07)上优于冻结的ProtoNet,但在ISIC上有所下降(-0.22)。在5-shot评估中,ProtoNet整体上仍然最强,而WIPT在ISIC上一致地优于容量相当的支持集Transformer(+0.99)。对多达五个查询样本进行联合处理并未带来可靠的准确率提升;在一个仅含头部的5-shot基准测试中,与g=1相比,g=5使分析性注意力token对减少73%,峰值内存分配减少29%,尽管延迟呈非单调变化。在所有目标域/shots条件下,WIPT对ProtoNet中不确定决策的改变远多于对高置信度决策的改变,且“挽救/破坏”分解(rescue/break decomposition)能够解释所观察到的增益与损失。源域偏移与评分器对照实验进一步表明,该方法的收益并非普遍存在。总体而言,WIPT提供了一种兼容流式处理的测试时原型自适应形式,能够在无需目标域优化的情况下改善困难的低样本跨域决策。
cs.CV / 28 / 2609.30783

Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models

跳过言语表达,重新聚焦视觉:面向多模态大语言模型推理分割的潜在推理方法
Guo, Tianhang, He, Yulin, Chen, Wei, Zhou, Wenjuan, Li, Yuhang, Gan, Xinbiao
Abstract
Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Thought (CoT) by multimodal large language models (MLLMs) before localizing the target. Although intuitive, such explicit verbal reasoning introduces substantial attention interference: redundant textual tokens disrupt attention during perception-token generation and also increase the effective distance between visual tokens. To address this issue, we propose LIRSeg, which fully replaces explicit CoT with a compact set of learnable latent tokens for reasoning segmentation. LIRSeg is trained in two stages: spatial alignment grounds the latent tokens in object-relevant visual evidence, and GRPO further optimizes them with segmentation rewards. To make these compact latent tokens more informative, we introduce three complementary mechanisms from an information perspective: extreme-advantage sampling for selecting informative training signals, decoupled exploration-stability updates for learning complementary representations, and latent diversity amplification for preventing representational collapse. Extensive experiments on benchmarks demonstrate that LIRSeg consistently improves both segmentation accuracy and reasoning efficiency. Compared with the VisionReasoner baseline, LIRSeg achieves absolute gIoU improvements of 4.9% on ReasonSeg, 7.1% on MUSE, and 4.7% on MMR, while achieving a approximately 16x reduction in reasoning tokens. Code is available in supplementary materials.
Chinese Translation
推理分割旨在理解隐式的文本查询并实现细粒度的视觉感知,这对于人机交互和具身智能体等应用至关重要。现有方法通常先由多模态大语言模型生成显式的思维链,然后再定位目标。尽管直观,但这种显式的语言推理会引入严重的注意力干扰:冗余的文本标记会在感知标记生成过程中扰乱注意力,并增加视觉标记之间的有效距离。为解决这一问题,我们提出LIRSeg,用一组紧凑的可学习潜在标记完全替代显式思维链来完成推理分割。LIRSeg采用两阶段训练:空间对齐阶段使潜在标记建立在与目标物体相关的视觉证据之上;随后利用GRPO(群体相对策略优化)结合分割奖励对其进行进一步优化。为使这些紧凑的潜在标记更具信息量,我们从信息论视角引入三种互补机制:用于选择信息量丰富的训练信号的极端优势采样、用于学习互补表征的解耦探索-稳定性更新,以及防止表征坍缩的潜在多样性增强。在多个基准上的大量实验表明,LIRSeg在分割精度和推理效率方面均取得持续提升。与VisionReasoner基线相比,LIRSeg在ReasonSeg、MUSE和MMR上分别取得4.9%、7.1%和4.7%的绝对gIoU提升,同时推理标记数量减少约16倍。代码见补充材料。
cs.CV / 29 / 2609.30795

Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion

运动风格滑块:面向人体运动扩散的端点监督连续风格控制
Liao, Chen-Chieh, Peng, Yichen, Cai, Yiyi, Ono, Yûi, Hanaoka, Hiroki, Wu, Erwin, Koike, Hideki, Kurabayashi, Shuichi
Abstract
Existing human motion diffusion methods provide strong motion generation quality, and recent style transfer models can inject target style cues, but fine-grained continuous control of style intensity remains underexplored. In production, style intensity is subjective across artists and directors, so the practical requirement is not a universal absolute unit, but a reliable monotonic control axis. We propose Motion Style Slider, a motion-to-motion style transfer framework for endpoint-supervised continuous control. Given a content motion and a style motion, we construct a style direction in a learned motion-style embedding space and condition diffusion generation with a scalar intensity. The training objective combines diffusion denoising with latent intensity regularization to encourage smooth and monotonic style scaling without requiring intermediate-intensity ground-truth motions. Our framework is compatible with pretrained motion diffusion backbones and supports heterogeneous style datasets, including the multi-actor style motion dataset. To test out-of-range usability, we additionally introduce a small real-capture over-reaction extension and evaluate large-intensity behavior against these unseen targets. Experiments measure controllability, interpolation/extrapolation behavior, content preservation, and motion realism, with ablations on direction construction and loss design.
Chinese Translation
现有的人体运动扩散方法具备强大的运动生成质量,近期的风格迁移模型也能够注入目标风格线索,但对风格强度的细粒度连续控制仍未得到充分探索。在实际制作中,风格强度在不同艺术家和导演之间存在主观差异,因此实际需求并非一个通用的绝对度量单位,而是一条可靠的单调控制轴。我们提出 Motion Style Slider(运动风格滑块),一个面向端点监督连续控制的运动到运动风格迁移框架。给定一个内容运动和一个风格运动,我们在学习得到的运动-风格嵌入空间中构建一个风格方向,并以标量强度为条件进行扩散生成。训练目标将扩散去噪与潜在强度正则化相结合,以实现平滑且单调的风格缩放,而无需中间强度的真实运动数据。我们的框架与预训练的运动扩散主干兼容,并支持异构风格数据集,包括多演员风格运动数据集。为测试超出范围的可用性,我们额外引入了一个小规模的真实采集过度反应扩展数据集,并针对这些未见过的目标评估大强度下的行为。实验测量了可控性、插值/外推行为、内容保持性和运动真实性,并对方向构建与损失设计进行了消融研究。
cs.CV / 30 / 2609.30855

MDSkin-Net: Multi-Task Skin Lesion Analysis Driven by Pattern Analysis Priors and Spatial Alignment Regularization

MDSkin-Net:由模式分析先验与空间对齐正则化驱动的多任务皮肤病变分析
Li, Yijian, Bedros, Saad, Bigliardi, Paul, Qi, Mei Bigliardi, Morellas, Vassilios, Papanikolopoulos, Nikolaos
Abstract
Reliable skin lesion segmentation and classification are central to dermoscopic computer-aided diagnosis. Existing multi-task frameworks couple the two tasks architecturally without clinical knowledge, while knowledge-injecting approaches rely on the macroscopic ABCD rule, which was not designed for dermoscopy. Dermoscopic diagnosis is grounded in Pattern Analysis, a microscopic framework structured around dermoscopic features. We propose MDSkin-Net, which incorporates cue-level Pattern Analysis priors into a hybrid CNN-Transformer architecture. At its core is a Pattern Analysis-Guided Attention Module (PAGAM) comprising three priors motivated by distinct dermoscopic cues: an improved Efficient Channel Attention (iECA), a Multi-Scale Spatial Attention (MSSA), and a Biased Asymmetry Attention (BAA). We further introduce a multi-scale spatial alignment regularization (MSAR) that uses the segmentation ground-truth mask as hierarchical soft supervision, confining the classification head to lesion-localized evidence and coupling both task pathways through a shared spatial prior. Trained exclusively on the ISIC 2017 training split without external dermoscopy data, the MDSkin-Net ensemble transfers robustly under zero-shot evaluation, reaching a Dice Similarity Coefficient (DSC) of 92.38% and a melanoma AUC of 97.84%on PH2, and a DSC of 88.92% on the ISIC 2018 Task 1 test set. On the in-domain ISIC 2017 benchmark, the ensemble attains a mean Area Under the Curve (AUC) of 91.60% across the two classification tasks (melanoma and seborrheic keratosis vs. rest), and a DSC of 84.72% for segmentation. Classification remains competitive with baselines; in-domain segmentation trails single-task specialists, yet the proposed priors and alignment regularization yield representations that generalize consistently across cohorts of different scales.
Chinese Translation
可靠的皮肤病变分割与分类是皮肤镜计算机辅助诊断的核心。现有多任务框架在架构上将两个任务耦合在一起,但缺乏临床知识;而知识注入方法则依赖于宏观层面的ABCD法则,该法则并非为皮肤镜检查而设计。皮肤镜诊断建立在模式分析(Pattern Analysis)基础之上,这是一种围绕皮肤镜特征组织的微观框架。我们提出了MDSkin-Net,将线索级的模式分析先验融入混合CNN-Transformer架构中。其核心是模式分析引导注意力模块(PAGAM),该模块包含三种由不同皮肤镜线索启发的先验:改进的高效通道注意力(iECA)、多尺度空间注意力(MSSA)和偏置非对称注意力(BAA)。我们进一步引入了多尺度空间对齐正则化(MSAR),该机制将分割的真实标注掩码作为分层软监督,将分类头限制在病变定位的证据上,并通过共享的空间先验将两个任务路径耦合起来。MDSkin-Net仅在ISIC 2017训练集上训练,未使用外部皮肤镜数据,在零样本评估下表现出稳健的迁移能力:在PH2数据集上达到92.38%的Dice相似系数(DSC)和97.84%的黑色素瘤AUC,在ISIC 2018 Task 1测试集上达到88.92%的DSC。在域内ISIC 2017基准上,集成模型在两个分类任务(黑色素瘤和脂溢性角化病 vs. 其他)上的平均曲线下面积(AUC)达到91.60%,分割的DSC达到84.72%。分类性能与基线方法相比仍具竞争力;尽管域内分割性能略逊于单任务专用模型,但所提出的先验和对齐正则化所产生的表示在不同规模的队列之间展现出一致的泛化能力。
cs.CV / 31 / 2609.30865

Reliability-Regulated Trajectory Optimization for Progressive COLMAP-Free 3D Gaussian Splatting

面向渐进式无COLMAP三维高斯泼溅的可靠性调节轨迹优化
Wu, Zijian, Wang, Jinliang, Lin, Zidian, Song, Ying, Lu, Ziqian, Ma, Hanjie, Ye, Zhen, Jiang, Mingfeng
Abstract
COLMAP-free 3D Gaussian Splatting (3DGS) bypasses computationally expensive structure-from-motion (SfM) pipelines, yet progressive camera pose tracking remains fundamentally vulnerable to error compounding---early pairwise tracking inaccuracies both corrupt subsequent frame initializations and remain permanently frozen in the scene representation. Rather than relying on heavyweight external neural priors or treating progressive tracking through isolated heuristic fixes, we propose a unified reliability-regulated trajectory optimization framework for progressive COLMAP-free 3DGS. At its core, our framework establishes an intrinsic, self-supervised bidirectional cycle-consistency mechanism that systematically regulates progressive camera trajectory estimation across two complementary temporal horizons: (1) Forward Motion Propagation, where the online reliability signal adaptively gates first-order kinematic warm-starts of rigid motion into upcoming pairwise registrations, supplying informed directional search priors while safely intercepting untrusted transitions; and (2) Retrospective Trajectory Correction, where the same reliability signal dynamically weights relative-pose consistency constraints within a sliding window of neighboring camera poses. By governing both prospective state initialization and retrospective trajectory consolidation through a unified reliability regulator, our self-contained framework resolves progressive drift without external priors or offline preprocessing. Extensive evaluations on Tanks and Temples and CO3D-V2 benchmarks show that our method substantially improves camera trajectory accuracy and novel-view rendering quality, outperforming existing unposed baselines. Code is available at https://github.com/Zijian1026/RRTO-CF3DGS.
Chinese Translation
无COLMAP三维高斯泼溅(3D Gaussian Splatting, 3DGS)避免了计算开销高昂的运动恢复结构(SfM)流程,但渐进式相机位姿跟踪从根本上容易受到误差累积的影响——早期的成对跟踪不准确既会破坏后续帧的初始化,又会永久性地固化在场景表示中。我们并未依赖重型外部神经先验,也未通过孤立的启发式修补来处理渐进式跟踪问题,而是提出了一种面向渐进式无COLMAP 3DGS的统一可靠性调节轨迹优化框架。该框架的核心在于建立了一种内在的、自监督的双向循环一致性机制,系统性地在两个互补的时间尺度上调节渐进式相机轨迹估计:(1)前向运动传播,其中在线可靠性信号自适应地将刚体运动的一阶运动学热启动门控至即将进行的成对配准中,在提供有依据的方向性搜索先验的同时,安全地拦截不可信的运动转换;(2)回溯性轨迹校正,其中同一可靠性信号在相邻相机位姿的滑动窗口内对相对位姿一致性约束进行动态加权。通过统一的可靠性调节器同时管控前瞻性的状态初始化与回溯性的轨迹巩固,我们的自包含框架无需外部先验或离线预处理即可解决渐进式漂移问题。在Tanks and Temples和CO3D-V2基准上的大量评估表明,我们的方法显著提升了相机轨迹精度和新视角渲染质量,优于现有的无位姿基线方法。代码已发布于 https://github.com/Zijian1026/RRTO-CF3DGS。
cs.CV / 32 / 2609.30928

UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound

UltraG-Bench:用于评估大型视觉语言模型在超声图像像素级证据定位能力的多任务基准
Zhu, Quanhao, Xu, Bo, Lin, Rui, Wang, Chenyuan, Shao, Yu, Zhu, Boling, Sun, Jiuyan, Zhao, Liang, Lin, Hongfei, Xia, Feng
Abstract
Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for evaluating pixel-level evidence grounding in ultrasound. UltraG-Bench is built by annotating 40 public ultrasound segmentation datasets spanning 13 anatomical categories, and comprises three progressive tasks: instruction-guided segmentation, evidence-grounded VQA, and evidence-grounded report generation, with 331125, 666779, and 138832 annotations, respectively. Comprehensive evaluation of 14 state-of-the-art models reveals a substantial gap between semantic understanding and fine-grained pixel-level localization. We further propose UltraG-Agent, which combines the semantic reasoning capabilities of a VLM with the ultrasound-specific segmentation capability of UltraSAM3. Experiments show that UltraG-Agent substantially improves both semantic prediction and pixel-level visual grounding. Our dataset and code are available at https://github.com/zhuqh19/UltraG-Bench.
Chinese Translation
超声是最广泛使用的医学影像模态之一,近年来大型视觉语言模型(VLM)在超声图像理解方面展现出不断增强的能力。然而,这些模型无法提供与其语义预测相一致的像素级视觉证据,且其在超声中的细粒度定位能力在很大程度上仍不明确。我们提出了UltraG-Bench,这是一个用于评估超声像素级证据定位的大规模多任务基准。UltraG-Bench通过对涵盖13个解剖类别的40个公开超声分割数据集进行标注而构建,包含三个递进式任务:指令引导分割、证据定位的视觉问答(VQA)和证据定位的报告生成,分别包含331,125、666,779和138,832条标注。对14个最先进模型的全面评估揭示了语义理解与细粒度像素级定位之间的显著差距。我们进一步提出了UltraG-Agent,它将VLM的语义推理能力与UltraSAM3的超声专用分割能力相结合。实验表明,UltraG-Agent显著提升了语义预测和像素级视觉定位的性能。我们的数据集和代码可在 https://github.com/zhuqh19/UltraG-Bench 获取。
cs.CV / 33 / 2609.30934

ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos

ManiVid:面向篡改视频的统一且可解释的取证分析
Kang, Hengrui, Yan, Zhonghao, Yang, Yuxuan, Jing, Ruoyan, Guo, Yuncheng, Chen, Hao, Liang, Kongming, Ma, Zhanyu, He, Conghui, Li, Weijia
Abstract
Rapid advances in AI-generated video (AIGV) have increased the risks posed by deceptive video manipulation. Unlike fully synthetic videos, manipulated videos retain most source content and alter only localized regions, making forensic analysis particularly challenging. Existing video forgery research faces two limitations in both data and methodology: (1) High-quality datasets and benchmarks tailored for manipulated videos remain scarce. (2) Multimodal large language models (MLLMs) extend forgery analysis beyond binary classification but struggle to use low-level forensic cues and provide precise pixel-level grounding. Specifically, we introduce ManiVid, a unified forensic analysis task covering forgery detection, artifact grounding, and anomaly explanation for manipulated videos. We construct ManiVid-38K, the first dataset to combine paired, open-vocabulary localized manipulations of general videos with authenticity labels, forgery masks, and anomaly explanations. It comprises about 19K manually verified real-fake video pairs, mostly at 1080P resolution, generated under 2 paradigms with 15 powerful generation models. We sample 1K pairs for ManiVidBench, balanced across six manipulation types and generation models for fair evaluation. We further propose ManiVidLens, a unified framework for explainable video forgery analysis. Its Forensic Evidence Router supplies shared low-level forensic evidence for multimodal reasoning and video segmentation. Its Prompt Distill Module converts grounding states into semantic and geometric prompts and distills spatial priors for mask decoding and full-video propagation. ManiVidLens achieves relative gains over the strongest comparison methods in artifact grounding (+21.1% mIoU; +21.3% J&F) and anomaly explanation (+131.3% ROUGE-L; +9.9% CSS). Its forgery detection remains comparable to dedicated classifiers (0.914 Acc; 0.913 F1).
Chinese Translation
人工智能生成视频(AIGV)的快速发展加剧了欺骗性视频篡改带来的风险。与完全合成的视频不同,被篡改的视频保留了大部分源内容,仅对局部区域进行改动,这使得取证分析尤为困难。现有的视频伪造研究在数据和方法层面面临两个局限:(1)针对篡改视频的高质量数据集和基准仍然稀缺;(2)多模态大语言模型(MLLMs)虽然将伪造分析从二分类拓展到了更广的范围,但难以利用低层取证线索,也难以提供精确的像素级定位。具体而言,我们提出了ManiVid,这是一个针对篡改视频的统一取证分析任务,涵盖伪造检测、篡改痕迹定位(artifact grounding)和异常解释。我们构建了ManiVid-38K,这是首个将通用视频的成对、开放词表的局部篡改与真实性标签、伪造掩码和异常解释相结合的数据集。它包含约19K个人工核验的真假视频对,大多为1080P分辨率,基于2种范式、15个强大的生成模型生成。我们从中采样1K对构建ManiVidBench,在六种篡改类型和生成模型之间保持均衡,以实现公平评估。我们进一步提出ManiVidLens,一个用于可解释视频伪造分析的统一框架。其取证证据路由器(Forensic Evidence Router)为多模态推理和视频分割提供共享的低层取证证据。其提示蒸馏模块(Prompt Distill Module)将定位状态转化为语义和几何提示,并蒸馏空间先验用于掩码解码和全视频传播。ManiVidLens在篡改痕迹定位(mIoU提升21.1%;J&F提升21.3%)和异常解释(ROUGE-L提升131.3%;CSS提升9.9%)上相较于最强对比方法取得了显著提升,其伪造检测性能与专用分类器相当(准确率0.914;F1值0.913)。
cs.CV / 34 / 2609.30941

Spackle: Completing Large View Single Image NVS with Adaptive Gaussians

Spackle:基于自适应高斯的单图像大视角偏移新视角合成补全方法
Liu, Xuanzhi, Zhou, Yuhe, Wu, Xinyi, Wu, Zhenyao, Chen, Jinghao, Han, Ruize, Wang, Song
Abstract
Single-image novel view synthesis (NVS) enables photorealistic rendering of un- observed viewpoints from a single input. Practical NVS systems require two key capabilities: robust reconstruction of occluded regions and high inference effi- ciency. While hybrid decoupled frameworks combining feedforward 3D Gaussian Splatting (3DGS) and diffusion models show promise for large-view-deviation NVS, they suffer from capacity competition: a fixed number of Gaussians forces resource shifts from visible to newly disoccluded areas, degrading original scene fidelity when the target view deviates significantly from the input. To address this, we propose Spackle, a lightweight residual learning framework that mit- igates capacity competition without sacrificing efficiency. Spackle operates in three stages: predicting base 3DGS attributes from given views, automatically identifying poorly reconstructed regions, and learning a residual 3DGS optimized exclusively for these areas. At inference, we combine the baseline and aug- mented Gaussians for NVS. We conduct comprehensive experiments and show that Spackle achieves state-of-the-art performance on large-view-deviation cases.
Chinese Translation
单图像新视角合成(NVS)能够从单张输入图像实现未观测视角的照片级真实感渲染。实用的NVS系统需要具备两项关键能力:对遮挡区域的鲁棒重建以及高推理效率。尽管结合前馈式三维高斯泼溅(3DGS)与扩散模型的混合解耦框架在大视角偏移NVS中展现出前景,但它们存在容量竞争问题:固定数量的高斯使得资源从可见区域转移到新解除遮挡的区域,当目标视角与输入图像偏差较大时,会降低原始场景的保真度。为解决这一问题,我们提出了Spackle,一个轻量级的残差学习框架,能够在不牺牲效率的前提下缓解容量竞争。Spackle分三个阶段运行:从给定视角预测基础3DGS属性,自动识别重建效果较差的区域,并为这些区域专门学习优化的残差3DGS。在推理阶段,我们将基础高斯与增强高斯相结合以完成NVS。我们进行了全面的实验,结果表明Spackle在大视角偏移情况下达到了最先进的性能。
cs.CV / 35 / 2609.30946

OneWorld: Learning Consistent Physics Across Actions in World Models

OneWorld:在世界模型中学习跨动作一致的物理规律
He, Ke, Ding, Yichen, Yang, Bin
Abstract
Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generated independently from the same initial scene may each appear plausible while implying incompatible physical properties, such as friction or mass. This inconsistency can lead to contradictory predictions across interventions, making it difficult for the model to maintain a coherent understanding of the underlying world and limiting its reliability for planning and decision-making. To address these issues, we propose OneWorld, a shared-mechanism counterfactual generation framework that jointly models multiple action-conditioned futures under a common latent physical mechanism. A physical mechanism interpreter first infers a distribution over latent mechanisms from each action-outcome branch. These distributions are then aggregated into shared-world evidence, which captures whether the branches admit a common physical explanation while accounting for uncertainty in less informative branches. This evidence constrains flow training and guides sampling, encouraging consistency in the underlying physical mechanism while preserving the distinct outcomes induced by different actions. We further introduce a multi-intervention evaluation protocol in controlled environments, following the interaction settings of ACWM-Phys, to assess whether generated futures can be jointly explained by the same physical parameters, alongside standard measures of single-rollout prediction quality. Experiments in these environments show that OneWorld improves cross-intervention physical consistency while maintaining competitive single-rollout prediction quality.
Chinese Translation
动作条件视频世界模型旨在预测不同动作下场景的演化,这一能力对于动态环境中可靠的规划、决策与交互至关重要。然而,由同一初始场景独立生成的多个未来可能各自看起来合理,却隐含着互不相容的物理属性(如摩擦力或质量)。这种不一致性会导致不同干预下的预测相互矛盾,使模型难以对潜在世界保持连贯的理解,并限制了其在规划和决策中的可靠性。为解决这些问题,我们提出 OneWorld,一个共享机制的反事实生成框架,在共同的潜在物理机制下联合建模多个动作条件下的未来。物理机制解释器首先从每个动作-结果分支推断潜在机制上的分布,随后将这些分布聚合为共享世界证据,该证据在考虑信息量较少分支的不确定性的同时,判断各分支是否可以由同一物理解释所涵盖。该证据约束流模型的训练并引导采样,在鼓励底层物理机制一致性的同时,保留不同动作所产生的各不相同的结果。我们还参照 ACWM-Phys 的交互设置,在受控环境中引入多干预评估协议,以评估生成的未来能否由相同的物理参数共同解释,并辅以单次轨迹预测质量的标准度量。在这些环境中的实验表明,OneWorld 提升了跨干预的物理一致性,同时保持了具有竞争力的单次轨迹预测质量。
cs.CV / 36 / 2609.30947

DAPEVO: Deep Adaptive Patch Frame-Event Visual Odometry

DAPEVO:深度自适应补丁帧-事件视觉里程计
Gandolfi, Luca, Nascivera, Simone, Pellerito, Roberto, Zou, Rong, Plizzari, Chiara, Scaramuzza, Davide
Abstract
Visual odometry is essential for autonomous navigation in GPS-denied environments, yet RGB-based methods remain vulnerable to motion blur, challenging illumination, and dropped frames. Event cameras complement conventional cameras with high temporal resolution and dynamic range, but their asynchronous measurements complicate reliable correspondence estimation. We present DAPEVO, a learned visual odometry system that estimates image and event correspondences independently at shared patch locations and fuses their correlation evidence before motion refinement. Each tracked patch maintains image and event descriptors, and a learned scalar gate combines modality-specific correlation embeddings for each patch--frame edge before a shared recurrent refinement and bundle-adjustment update. DAPEVO also supports event-only observations, enabling continued tracking when RGB frames are sparse or unavailable, while modality-aware keyframe culling preserves scarce frame constraints. On UZH-FPV, when retaining only one in six RGB frames, DAPEVO's mean absolute trajectory error (ATE) increases by only 36%, from 1.00 to 1.36m, whereas the ATE of DPVO and RAMP-VO rises by factors of $3.7\times$ and $3.1\times$, respectively. On TartanEvent, DAPEVO similarly remains below 1m ATE at 3Hz RGB input, while DPVO and RAMP-VO exceed 9m. Under degraded RGB input on TartanEvent, DAPEVO achieves an ATE of 0.60m, compared with more than 4m for both DPVO and RAMP-VO, while also outperforming event-only DEVO at 0.87m.
Chinese Translation
视觉里程计对于GPS拒止环境下的自主导航至关重要,然而基于RGB的方法仍然容易受到运动模糊、复杂光照条件和丢帧问题的影响。事件相机凭借其高时间分辨率和高动态范围可以补充传统相机的不足,但其异步的测量方式使可靠的对应关系估计变得复杂。我们提出了DAPEVO,一种基于学习的视觉里程计系统,该系统在共享的补丁(patch)位置上独立估计图像和事件对应关系,并在运动优化之前融合二者的相关性证据。每个被跟踪的补丁维护图像描述子和事件描述子,并通过一个可学习的标量门控机制为每个补丁-帧边组合模态特定的相关性嵌入,随后进行共享的循环优化与光束法平差(bundle adjustment)更新。DAPEVO还支持仅有事件数据的观测,使得在RGB帧稀疏或不可用时仍能持续跟踪,同时模态感知的关键帧筛选机制保留了稀缺的帧约束。在UZH-FPV数据集上,当仅保留六分之一的RGB帧时,DAPEVO的平均绝对轨迹误差(ATE)仅增加36%,从1.00米增至1.36米,而DPVO和RAMP-VO的ATE分别上升了3.7倍和3.1倍。在TartanEvent数据集上,DAPEVO在3Hz RGB输入下同样保持低于1米的ATE,而DPVO和RAMP-VO则超过9米。在TartanEvent的退化RGB输入条件下,DAPEVO达到了0.60米的ATE,而DPVO和RAMP-VO均超过4米,同时DAPEVO还优于仅使用事件数据的DEVO(0.87米)。
cs.CV / 37 / 2609.30952

MVVBench: Benchmarking 4D Reasoning in Vision-Language Models

MVVBench:面向视觉-语言模型4D推理能力的基准测试
Chung, Hyungjin, Park, Byeongjun, Lee, Joonseok, Kim, Hojun, Choi, Jaeho, Kim, Byung-Hoon
Abstract
Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBench, a benchmark for multi-view video reasoning built from real world multi camera datasets. Questions are curated to be monocular-ambiguous along both the view and the temporal axis: each question is unanswerable from any single view in the designated input set, and the majority are further unanswerable from any single moment. Each question becomes uniquely solvable only by jointly reasoning across views and across time. MVVBench spans diverse dynamic scenes and probes six capabilities: implicit/explicit attribute identification, implicit/explicit relative distance, relative camera pose, and compositional counting, with human-authored QA and rigorous verification. Beyond benchmarking, we provide an extensive analysis of when and why current vision language models succeed or fail, characterizing errors due to temporal mis-localization, cross-view identity breaks, and brittle multi-hop reasoning. We then study inference-time elicitation strategies that unlock latent multi-view competence---task-specific chain-of-thought scaffolds and structured cross-view evidence aggregation---yielding substantial gains without retraining. Finally, we present preliminary evidence that reinforcement learning with verifiable rewards can elicit some latent multi-view competence in the base model, pointing to training-time approaches as a promising direction for future work. Together, MVVBench offers a rigorous evaluation of 4D multi-view reasoning and a foundation for future progress toward reliable embodied perception.
Chinese Translation
多视角视频理解需要整合来自多个(通常互不重叠的)相机流中的空间与时间证据:追踪实体在不同视角间的转换、跨时间对齐事件,并基于潜在的4D连续性而非任何单一可见帧进行推理。我们提出了MVVBench,一个基于真实世界多相机数据集构建的多视角视频推理基准。这些问题被精心设计为在视角轴和时间轴上均具有单目模糊性:每个问题都无法从指定输入集中的任何单一视角得到解答,且大多数问题进一步无法从任何单一时刻得到解答。只有联合跨视角与跨时间进行推理,每个问题才能被唯一地求解。MVVBench涵盖多样的动态场景,探测六种能力:隐式/显式属性识别、隐式/显式相对距离、相对相机位姿以及组合计数,配有由人工撰写的问答对和严格的验证流程。除基准测试外,我们还对当前视觉-语言模型何时以及为何成功或失败进行了广泛分析,刻画了由时间错定位、跨视角身份断裂以及脆弱的多跳推理所导致的错误。随后,我们研究了能够激发潜在多视角能力的推理时引导策略——包括任务特定的思维链支架和结构化的跨视角证据聚合——在不进行重新训练的情况下取得了显著提升。最后,我们提供了初步证据,表明具有可验证奖励的强化学习能够激发基座模型中部分潜在的多视角能力,这表明训练时方法是未来工作的一个有前景的方向。总而言之,MVVBench为4D多视角推理提供了严格的评估,并为未来迈向可靠具身感知的进展奠定了基础。
cs.CV / 38 / 2609.30962

IDM-Net: A Lightweight Illumination-Decoupled Modulation Network for Low-Light Image Enhancement

IDM-Net:一种用于低照度图像增强的轻量级光照解耦调制网络
Hsiao, Cheng-Yen, Guo, Jing-Ming
Abstract
Low-light image enhancement (LLIE) remains challenging for lightweight models because illumination restoration and color fidelity are difficult to optimize simultaneously in the RGB color space. Although recent color-decoupled methods separate luminance and chrominance representations, they primarily optimize luminance as an enhancement target, leaving its potential as an explicit guidance prior largely unexplored during feature reconstruction. To address this limitation, we propose IDM-Net, a lightweight Illumination-Decoupled Modulation Network for low-light image enhancement. IDM-Net adopts a dual-encoder architecture consisting of a structure encoder that extracts multi-scale appearance features from the RGB image and a lightweight illumination encoder that learns illumination priors from the decoupled luminance (Y) channel. To effectively exploit these priors, we introduce an Illumination-Guided Modulation (IGM) module that injects multi-scale illumination cues into the decoder through spatially adaptive affine modulation, enabling accurate brightness restoration while preserving natural color consistency. Furthermore, we design a lightweight Feature Refinement Block (FRB) to progressively suppress degradation artifacts and recover fine-grained image details during reconstruction. Extensive experiments on multiple standard low-light image enhancement benchmarks demonstrate that IDM-Net achieves competitive performance among lightweight LLIE methods while maintaining an excellent balance between restoration quality and computational efficiency.
Chinese Translation
低照度图像增强(LLIE)对轻量级模型而言仍具有挑战性,因为在RGB色彩空间中难以同时优化光照恢复与色彩保真度。尽管近期的颜色解耦方法将亮度与色度表征分离,但它们主要将亮度作为增强目标进行优化,而在特征重建过程中,其作为显式引导先验的潜力在很大程度上尚未被探索。为解决这一局限,我们提出了IDM-Net,一种用于低照度图像增强的轻量级光照解耦调制网络。IDM-Net采用双编码器架构,其中结构编码器从RGB图像中提取多尺度外观特征,轻量级光照编码器从解耦后的亮度(Y)通道中学习光照先验。为有效利用这些先验,我们引入了光照引导调制(Illumination-Guided Modulation, IGM)模块,通过空间自适应仿射调制将多尺度光照线索注入解码器,在保持自然色彩一致性的同时实现精确的亮度恢复。此外,我们设计了轻量级的特征精炼模块(Feature Refinement Block, FRB),以在重建过程中逐步抑制退化伪影并恢复细粒度图像细节。在多个标准低照度图像增强基准上的大量实验表明,IDM-Net在轻量级LLIE方法中取得了具有竞争力的性能,同时在恢复质量与计算效率之间保持了出色的平衡。
cs.CV / 39 / 2609.30963

Where and When to Force: Routed Forcing for Streaming Avatars

在何处与何时进行强制:面向流式虚拟人的路由强制方法(Routed Forcing)
Su, Zihan, Lu, Siwen, Zhuang, Junhao, Xue, Zeyue, Huang, Haoyang, Li, Guanghao, Tan, Xiaofeng, Yuan, Chun, Duan, Nan
Abstract
Audio-driven streaming avatar generation requires real-time synthesis of speech-synchronized videos with dynamic and diverse motion. Self Forcing uses Distribution Matching Distillation (DMD) to distill bidirectional video diffusion models into causal, few-step generators for real-time streaming. However, DMD minimizes a reverse KL divergence, which is inherently mode-seeking: it causes the student to discard high-dynamic modes and collapse onto static outputs, compressing both dynamics and diversity of generated videos. We find that this collapse is region-heterogeneous: person regions involving pose and gesture variations suffer the largest diversity loss, the audio-driven mouth region shows a small loss, and the background remains nearly stable. Based on this observation, we propose Routed Forcing, which routes the distillation objective by semantic region and noise stage to improve dynamics and diversity while preserving visual quality. Specifically, (1) Where to Force: Semantic-Region Routing applies Data-Forcing Distillation (DFD), which supervises the student with real videos, to the person region where diversity collapse is most severe, while retaining DMD for the mouth and background to preserve lip synchronization and scene stability. (2) When to Force: Noise-Stage Routing activates DFD at high noise stages, where real video serves as effective supervision to inject diverse and dynamic motion patterns. At low noise stages, DMD is used to refine details, avoiding blur and artifacts from spatial differences between real video and student-generated video. Experiments show that Routed Forcing improves dynamics by up to 45% and diversity by 7-25% over Self Forcing, while preserving video quality and lip synchronization.
Chinese Translation
音频驱动的流式虚拟人(avatar)生成需要实时合成与语音同步、且具有动态性和多样性运动的视频。Self Forcing 使用分布匹配蒸馏(Distribution Matching Distillation, DMD)将双向视频扩散模型蒸馏为因果的少步生成器,以实现实时流式生成。然而,DMD 最小化的是反向 KL 散度,其本质是模态寻优(mode-seeking)的:这会导致学生模型丢弃高动态模态并坍缩为静态输出,从而压缩了生成视频的动态性与多样性。我们发现这种坍缩具有区域异质性:涉及姿态和手势变化的人物区域遭受的多样性损失最大,音频驱动的嘴部区域损失较小,而背景则基本保持稳定。基于这一观察,我们提出 Routed Forcing(路由强制),该方法按语义区域和噪声阶段对蒸馏目标进行路由,以在保持视觉质量的同时提升动态性与多样性。具体而言:(1) 在何处强制:语义区域路由(Semantic-Region Routing)在多样性坍缩最严重的人物区域应用数据强制蒸馏(Data-Forcing Distillation, DFD),即用真实视频监督学生模型,同时在嘴部和背景区域保留 DMD,以保持唇部同步和场景稳定。(2) 何时强制:噪声阶段路由(Noise-Stage Routing)在高噪声阶段启用 DFD,此时真实视频可作为有效监督,用于注入多样且动态的运动模式;在低噪声阶段则使用 DMD 来细化细节,避免因真实视频与学生生成视频之间的空间差异而产生的模糊和伪影。实验表明,与 Self Forcing 相比,Routed Forcing 将动态性提升高达 45%,多样性提升 7-25%,同时保持了视频质量和唇部同步。
cs.CV / 40 / 2609.30979

CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models

CCRV-Bench:基于约束的视觉语言模型因果推理评估
Gao, Linyuan, Wu, Yuan, Chang, Yi
Abstract
Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlations for shortcut learning, thereby potentially overestimating their actual capabilities. This paper proposes CCRV-Bench, a constraint-driven visual causal reasoning benchmark for single-image physical scenarios. We construct an orthogonal framework that evaluates four causal task dimensions: causal relation discovery, state prediction, causal diagnosis, and intervention. We further introduce entity symbolization, spatial grounding, the factual adversarial constraint, and minimalist output constraints to reduce shortcut cues while preserving the physical commonsense required by the task. Experiments across 15 multimodal models show that constraint sensitivity is task- and model-dependent: intervention has the largest average effective degradation among the four causal tasks, spatial grounding is the most damaging constraint on average, and the factual adversarial constraint improves DCR for all evaluated models. These results show that unconstrained performance does not determine constrained robustness and that a single aggregate score can obscure distinct failures in causal identification, spatial grounding, and constraint-compliant expression. CCRV-Bench provides a standardized framework for diagnosing image-grounded causal reasoning under controlled constraints. The code is available at https://github.com/0815linyuan/CCRV-Bench-Constraint-Based-Evaluation-of-Causal-Reasoning-in-Vision-Language-Models
Chinese Translation
视觉语言模型(VLM)在视觉任务中表现出色,但其视觉因果推理能力仍缺乏可靠的评估方法。现有评估难以区分模型是基于视觉证据进行因果推理,还是依赖统计相关性进行捷径学习,因而可能高估其实际能力。本文提出CCRV-Bench,一个面向单图像物理场景的约束驱动的视觉因果推理基准。我们构建了一个正交框架,评估四个因果任务维度:因果关系发现、状态预测、因果诊断和干预。我们进一步引入实体符号化、空间定位、事实对抗约束和极简输出约束,在保留任务所需物理常识的同时减少捷径线索。在15个多模态模型上的实验表明,约束敏感性依赖于任务和模型:干预任务在四个因果任务中平均有效退化最大,空间定位是平均而言最具破坏性的约束,而事实对抗约束提升了所有被评估模型的DCR。这些结果表明,无约束条件下的表现并不能决定有约束条件下的鲁棒性,单一的总体得分可能掩盖模型在因果识别、空间定位和符合约束的表达方面的不同失败模式。CCRV-Bench为在受控约束下诊断基于图像的因果推理提供了标准化框架。代码已在 https://github.com/0815linyuan/CCRV-Bench-Constraint-Based-Evaluation-of-Causal-Reasoning-in-Vision-Language-Models 开源。
cs.CV / 41 / 2609.30981

STORM-Bench: Evaluating Online Video QA under Evolving and Incomplete Evidence

STORM-Bench:在演化且不完整证据下评估在线视频问答
Zhong, Siru, Tan, Shenghan, Yan, Rihong, Lv, Xiaohui, Zhuang, Yuzheng, Tao, Shuai, Liu, Wulong, Fu, Haohuan, Liang, Yuxuan
Abstract
Reliable online video question answering requires tracking state transitions while selectively abstaining when visual evidence is insufficient. Existing benchmarks focus on static recognition or long-range retrieval, rarely evaluating these coupled capabilities under evolving and incomplete evidence. We present STORM-Bench, comprising 5,736 questions across 630 compact, change-dense episodes spanning five egocentric domains (STORM-Real) and two controlled simulation subsets (STORM-Sim) at 1 FPS. Questions are stratified by a proxy for accumulated change intensity (Low, Medium, High) and query-time answerability (Known, Uncertain). To measure reliability, we introduce STORM-BR, a harmonic metric over joint answer-status correctness that exposes abstention failures masked by aggregate accuracy, alongside STORM-BR-ATTR for uncertainty attribution. Across 14 video LLMs, online accuracy peaks at 60.3\% (mean 51.7\%), whereas STORM-BR ranges from 5.7\% to 35.6\% (mean 18.8\%), driven by pervasive overconfidence on uncertain queries. STORM-Bench shows that task accuracy masks these gaps in epistemic reliability and state tracking. Benchmark and code are available at https://github.com/siruzhong/STORM-Bench.
Chinese Translation
可靠的在线视频问答需要在追踪状态转换的同时,在视觉证据不足时有选择地拒答。现有基准主要关注静态识别或长程检索,很少在演化和不完整的证据条件下评估这些相互耦合的能力。我们提出STORM-Bench,包含5,736个问题,涵盖630个紧凑且变化密集的片段,横跨五个自我中心领域(STORM-Real)以及两个以1帧每秒(1 FPS)采样的受控模拟子集(STORM-Sim)。问题按累积变化强度的代理指标(低、中、高)以及查询时点的可答性(已知、不确定)进行分层。为衡量可靠性,我们引入STORM-BR——一种基于答案-状态联合正确性的调和指标,能够暴露被总体准确率掩盖的拒答失败——同时引入STORM-BR-ATTR用于不确定性归因。在14个视频大语言模型上的实验表明,在线准确率最高为60.3%(平均51.7%),而STORM-BR仅在5.7%至35.6%之间(平均18.8%),其主要原因是模型在不确定问题上普遍存在过度自信。STORM-Bench表明,任务准确率掩盖了模型在认知可靠性和状态追踪方面的这些差距。基准与代码已发布于 https://github.com/siruzhong/STORM-Bench。
cs.CV / 42 / 2609.30982

FARE: Forensic Acceptance Region Estimation for Catching Bait-and-Switch Image Generators

FARE:用于捕获“偷梁换柱”图像生成器的取证接受区域估计方法
Yao, Kai, Juarez, Marc
Abstract
Modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service, but cannot inspect model weights or architecture. This creates a practical challenge: a provider may pass governance certification with one generator and later silently switch to a cheaper and lower-quality one for deployment, compromising public trust or even safety in high-stakes domains. We study integrity auditing at deployment time and propose FARE (Forensic Acceptance Region Estimation). A certified generator is enrolled by training FARE on images sampled from that generator. After deployment, FARE can determine whether a generated image is consistent with the enrolled generator---using only that image. FARE's features are based on image generator-specific artifacts that have been proposed for forensic applications. FARE amplifies these features during training by finding hard samples that tighten the acceptance region and increase sensitivity to subtle changes in the certified generator. Across generator swaps, including substitutions with similar model versions and model variants, FARE is effective at detecting swaps, consistently outperforming existing baselines at strict operating points, and remains effective under the exact-model and decision-only attacks evaluated in this work.
Chinese Translation
现代AI图像生成器越来越多地以不透明API的形式部署,客户可以查询已部署的服务,但无法检查模型权重或架构。这带来了一个现实挑战:提供方可能用一个生成器通过治理认证,随后悄悄切换到更便宜、质量更低的生成器进行部署,从而损害公众信任,甚至危及高风险领域的安全。我们研究了部署时的完整性审计,并提出FARE(Forensic Acceptance Region Estimation,取证接受区域估计)。通过在从经认证的生成器采样的图像上训练FARE来注册该生成器。部署之后,FARE仅需依据单张生成图像即可判断该图像是否与已注册的生成器一致。FARE的特征基于此前为取证应用提出的图像生成器特有伪影(artifacts)。FARE在训练过程中通过寻找困难样本紧缩接受区域,放大这些特征,并提高对经认证生成器细微变化的敏感度。在多种生成器替换场景下(包括被相似模型版本和模型变体替换),FARE能有效检测出替换行为,在严格操作点下持续优于现有基线方法,并且在本工作所评估的精确模型攻击和仅决策攻击下依然保持有效。
cs.CV / 43 / 2609.30987

Self-Supervised Perceptually Interpretable Monocular Depth Estimation

自监督的具有感知可解释性的单目深度估计
Abidin, Zain Ul, Dimas, George, Iakovidis, Dimitris K.
Abstract
Self-supervised monocular depth estimation (MDE) enables depth prediction from monocular images without requiring ground-truth supervision, making it attractive for large-scale and real-world applications. Despite steady improvements in accuracy, most existing methods remain difficult to interpret, as depth is inferred from RGB representations that obscure the impact of individual perceptual image components. This lack of transparency limits systematic analysis of failure cases and reduces confidence in safety-critical settings. This paper presents a self-supervised framework for perceptually interpretable monocular depth estimation (PIMDE), designed to associate depth predictions with distinct perceptual components of the input image. Rather than operating directly on RGB inputs, the proposed method decomposes each image into a set of perceptual feature maps (PFMs), each encoding a specific visual cue. Distinct depth estimation branches process these PFMs independently to produce depth estimates (PIDEs), which are subsequently combined through an explicit fusion strategy. This formulation allows us to examine directly the contribution of each perceptual cue to the final depth prediction. Experiments conducted on the KITTI benchmark dataset demonstrate that PIMDE achieves performance comparable to established self-supervised MDE methods while providing additional insight into how different perceptual cues influence depth estimation. These results indicate that perceptual decomposition can support interpretability without sacrificing depth estimation accuracy.
Chinese Translation
自监督单目深度估计(MDE)能够在无需真值监督的情况下从单目图像中预测深度,这使其在大规模和真实世界应用中极具吸引力。尽管精度不断提升,但大多数现有方法仍难以解释,因为深度是从RGB表示中推断出来的,而RGB表示掩盖了各个感知图像成分的影响。这种透明度的缺失限制了对失败案例的系统性分析,并降低了对安全关键场景的信心。本文提出了一种面向感知可解释单目深度估计(PIMDE)的自监督框架,旨在将深度预测与输入图像中不同的感知成分相关联。该方法不直接作用于RGB输入,而是将每幅图像分解为一组感知特征图(PFMs),每个特征图编码一种特定的视觉线索。不同的深度估计分支独立处理这些PFMs以生成深度估计结果(PIDEs),随后通过显式的融合策略将它们组合起来。这种构造使我们能够直接考察每种感知线索对最终深度预测的贡献。在KITTI基准数据集上进行的实验表明,PIMDE取得了与成熟的自监督MDE方法相当的性能,同时提供了关于不同感知线索如何影响深度估计的额外洞察。这些结果表明,感知分解可以在不牺牲深度估计精度的情况下支持可解释性。
cs.CV / 44 / 2609.30988

PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution

PhoenixSR:生成式异构蒸馏释放高效模型用于真实世界超分辨率
Di, Xin, Shi, Mingyu, Bao, Yuanfei, Peng, Long, Zhao, Yue, Guo, Jiaming, Pei, Renjing, Fu, Xueyang, Cao, Yang, Zha, Zheng-Jun
Abstract
Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while preserving faithful content. Diffusion-based SR benefits from strong generative priors but incurs substantial computational overhead, whereas feed-forward CNN and Transformer SR models are efficient yet often struggle to recover realistic high-frequency details. This motivates a natural question: can diffusion priors be transferred to existing diffusion-free SR networks without introducing diffusion components at inference time? To this end, we propose PhoenixSR, a generative heterogeneous distillation framework that transfers diffusion priors to independently designed feed-forward SR networks through score-based distribution matching. Rather than aligning heterogeneous features or imitating sampled diffusion outputs, PhoenixSR uses the pretrained diffusion model as distribution-level supervision, while paired SR supervision preserves reconstruction fidelity. To make distribution matching effective for fidelity-sensitive SR, we introduce Heterogeneous Distribution Adaptation, which adapts the target score to the SR domain, improves tracking of the evolving student distribution, and anchors training with paired supervision. We further employ Directional Reliability Weighting, a lightweight residual-consistency-based reweighting strategy that reduces unstable distributional guidance. All diffusion-related components are removed after training, leaving the original student architecture and inference cost unchanged. Experiments on three SR benchmarks and six feed-forward backbones, including SwinIR, HAT, Real-ESRGAN, and SeeMoRe, show consistent perceptual improvements with largely preserved reconstruction fidelity.
Chinese Translation
真实世界图像超分辨率(SR)需要在从复杂的低分辨率观测中恢复感知上真实的高分辨率图像的同时,保持内容的忠实性。基于扩散模型的超分辨率方法受益于强大的生成先验,但会带来显著的计算开销;而前馈式CNN和Transformer超分辨率模型虽然高效,却往往难以恢复真实的高频细节。这引出了一个自然的问题:能否将扩散先验迁移到现有的无扩散超分辨率网络中,而不在推理时引入扩散组件?为此,我们提出了PhoenixSR,一个生成式异构蒸馏框架,通过基于分数的分布匹配将扩散先验迁移到独立设计的前馈式超分辨率网络中。PhoenixSR并非对齐异构特征或模仿扩散模型的采样输出,而是将预训练的扩散模型作为分布级别的监督,同时利用成对超分辨率监督来保持重建的保真度。为了使分布匹配对保真度敏感的超分辨率任务有效,我们引入了异构分布适配(Heterogeneous Distribution Adaptation),将目标分数适配到超分辨率领域,改进了对学生分布演变的追踪,并以成对监督为训练提供锚定。我们进一步采用方向可靠性加权(Directional Reliability Weighting),这是一种轻量的基于残差一致性的重加权策略,可减少不稳定的分布引导。训练完成后,所有与扩散相关的组件均被移除,学生模型的原始架构和推理成本保持不变。在三个超分辨率基准和六个前馈骨干网络(包括SwinIR、HAT、Real-ESRGAN和SeeMoRe)上的实验表明,该方法在大幅保持重建保真度的同时,带来了一致的感知质量提升。
cs.CV / 45 / 2609.30989

PICO: Projection-Informed Consistency Optimisation for 6DoF Surgical Tool Pose Estimation

PICO:面向6自由度手术器械位姿估计的投影引导一致性优化方法
Fothergill, Lucy, Valdastri, Pietro, Jones, Dominic, Sarikaya, Duygu
Abstract
Purpose: Accurate 6 DoF pose estimation of surgical tools is critical for automa- tion, robotic proprioception, and safe interaction with the tissue operated on. Kinematics-based approaches suffer from accumulated errors due to the cable- driven nature of robotic arms, while vision-based methods often rely on external markers or trackers. Although more recent vision-based advances have been pro- posed, these two-stage pose estimation methods often lack real-time robustness due to accumulated errors and computational overhead. Methods: We propose a novel end-to-end trainable model, PICO. Our model employs a multi-task learning architecture to predict segmentation and depth maps, alongside regression of translation and rotation parameters. We define two proxy tasks that enforce geometric consistency in both 2D and 3D spaces, improving accuracy and robustness. For this, we propose a projection loss, and a point-to-point loss. Results: We evaluate our method on the SurgRIPE dataset, benchmarking its performance against state-of-the-art approaches using standard 6DoF pose esti- mation metrics. Our results demonstrate consistently strong performance across all four subsets, specifically in rotation, ranking second even under occlusion. It also demonstrates comparable translational performance, remaining competitive, especially in occluded cases. Conclusion: PICO demonstrates the effectiveness of multi-task learning and geometry-aware proxy tasks for robust and reliable surgical tool pose estimation, especially in occluded scenarios, highlighting potential for future applications.
Chinese Translation
目的:手术器械的精确6自由度(6DoF)位姿估计对于手术自动化、机器人本体感知以及与手术组织的安全交互至关重要。基于运动学的方法由于机械臂的绳驱动特性而存在误差累积问题,而基于视觉的方法通常依赖外部标记物或跟踪器。尽管近期提出了更多基于视觉的改进方法,但这类两阶段位姿估计方法由于误差累积和计算开销,往往缺乏实时鲁棒性。方法:我们提出了一种新颖的端到端可训练模型——PICO。该模型采用多任务学习架构,预测分割图和深度图,同时回归平移和旋转参数。我们定义了两个代理任务,以在2D和3D空间中强制几何一致性,从而提升精度和鲁棒性。为此,我们提出了投影损失和点到点损失。结果:我们在SurgRIPE数据集上评估了所提方法,使用标准的6DoF位姿估计指标与最先进方法进行了基准比较。结果表明,我们的方法在全部四个子集上均表现出持续强劲的性能,尤其在旋转估计方面表现突出,即使在遮挡情况下仍排名第二。在平移估计方面也表现出相当的性能,保持竞争力,特别是在存在遮挡的场景中。结论:PICO证明了多任务学习和几何感知代理任务在实现鲁棒、可靠手术器械位姿估计方面的有效性,尤其是在遮挡场景下,展现了未来应用的潜力。
cs.CV / 46 / 2609.30993

FLIP: Final Layer Inference-Time Probing for Vision-Language Models

FLIP:面向视觉-语言模型的最终层推理时探测方法
Juanico, Drandreb Earl O., Atienza, Rowel O.
Abstract
We present FLIP, a final-layer inference-time probe for testing whether a logit-facing intervention site in an open-weight vision-language model (VLM) supports structured, task-linked computation rather than generic perturbation. Behavioral change under internal intervention is otherwise mechanistically ambiguous: it may reflect improved use of visual evidence, generic output instability, or outright degradation. FLIP applies elementwise flooring to the final normalized hidden state before logit computation, leaving parameters, prompts, and decoding unchanged. On a controlled detection/counting probe, sweeping intervention strength reveals three regions: negligible change, a bounded interior regime in which detection recall at IoU 0.50 ($R_{50}$) improves while tolerant counting error ($\mathcal{E}_{\mathrm{count}}$) falls, and over-suppression. We formalize a four-criterion probe-and-sweep protocol for disciplining the interpretation of intervention effects: regime structure, grounding-proxy alignment, feature-coherence dependence, and failure to reproduce the same positive regime on a performance-based negative control. The post-normalization state passed to the output head is the logit-facing instantiation of this test; under a non-targeted flooring sweep it satisfies the full protocol. Raw decoder-layer interventions, including the last-block output before final normalization, and the singleton-pair left/right control fail to reproduce the Final-site signature, while same-site operators and multiple VLMs replicate it. FLIP is therefore a validation step for intervention-based mechanistic interpretability, not a steering method.
Chinese Translation
我们提出了FLIP(最终层推理时探测方法),用于检验开放权重视觉-语言模型(VLM)中面向logit的干预位点是否支持结构化的、与任务相关的计算,而非仅是通用性扰动。内部干预引起的行为变化在机制上往往是模糊的:它可能反映了对视觉证据利用的改善、通用的输出不稳定,或彻底的性能退化。FLIP在logit计算之前,对最终归一化后的隐状态施加逐元素下限约束,而不改变模型参数、提示词和解码过程。在一个受控的检测/计数探测任务上,通过扫描干预强度可以观察到三个区域:变化可忽略区域、有界的内部区间(在该区间内,IoU 0.50下的检测召回率$R_{50}$提升,同时容忍性计数误差$\mathcal{E}_{\mathrm{count}}$下降),以及过度抑制区域。我们形式化了一个包含四项准则的“探测-扫描”协议,用以规范对干预效应的解释:区域结构、与视觉定位代理指标的一致性、对特征相干性的依赖,以及无法在基于性能的阴性对照上复现同样的正向区间。传递给输出头的归一化后状态正是该检验面向logit的实例化对象;在非定向的下限扫描下,它满足了完整协议。原始解码器层干预(包括最终归一化之前的最后一块输出)以及单例对的左/右对照均无法复现最终位点(Final-site)的特征签名,而同一位点上的不同算子以及多个VLM模型则能够复现该特征。因此,FLIP是面向干预式机制可解释性研究的一种验证手段,而非一种引导(steering)方法。
cs.CV / 47 / 2609.31005

TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking

TRACKGRAPH:基于图像空间跟踪的在线开放词汇3D场景图
Hellesylt, Peder Borge, Puigjaner, Albert Gassol, Alexis, Kostas, Stahl, Annette
Abstract
Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TRACKGRAPH, an online open-vocabulary system that maintains short-term 2D mask identity directly in the image stream before fusing segments into 3D. FastSAM masks and CLIP features are computed at sparse keyframes, while dense DINOv3 features are used to propagate masks at a high rate in between. The resulting tracked masks are fused into a class-agnostic 3D segment layer within a hierarchical scene graph, with 3D association handling tracking interruptions and long-term revisits. Compact multi-view CLIP embeddings enable open-vocabulary retrieval. Across Replica, ScanNet++, and HM3D, TRACKGRAPH achieves competitive open-vocabulary segmentation and retrieval against state-of-the-art mapping methods, including the highest synonym frequency on Replica (0.50). On the same NVIDIA A100, it is 1.7x faster and uses 3.3x less GPU memory than ViT-H OVI-MAP. Real-world quadruped deployments demonstrate onboard scene graph construction and object search at 7.5Hz, while recorded drone data is used to test the method under aerial viewpoints.
Chinese Translation
开放词汇3D地图使机器人能够使用自然语言对先前未知的环境进行推理。然而,现有系统通常对每一帧输入图像进行分割,将检测结果与持久的3D片段进行关联,并频繁执行代价高昂的视觉-语言(Vision-Language,VL)推理。我们提出了TRACKGRAPH,这是一个在线开放词汇系统,它在将片段融合到3D之前,直接在图像流中维护短期2D掩码身份。系统在稀疏关键帧上计算FastSAM掩码和CLIP特征,同时在关键帧之间使用稠密DINOv3特征以高速率传播掩码。由此得到的跟踪掩码被融合到层次化场景图中与类别无关的3D片段层中,其中3D关联机制用于处理跟踪中断和长期重访。紧凑的多视角CLIP嵌入支持开放词汇检索。在Replica、ScanNet++和HM3D数据集上,TRACKGRAPH相比最先进的建图方法取得了具有竞争力的开放词汇分割和检索性能,包括在Replica上最高的同义词频率(0.50)。在同一块NVIDIA A100上,它比ViT-H OVI-MAP快1.7倍,GPU内存占用减少3.3倍。真实世界四足机器人的部署演示了以7.5Hz频率进行机载场景图构建和目标搜索,同时使用录制的无人机数据来测试该方法在航拍视角下的表现。
cs.CV / 48 / 2609.31028

Refining Cytology Predictions with Conditional Random Fields

利用条件随机场优化细胞学预测
Dausort, Manon, Godelaine, Tiffanie, Khoury, Karim El, Zanella, Maxime, De Vleeschouwer, Christophe, Macq, Benoît
Abstract
Vision-language models (VLMs) achieve strong zero-shot (ZS) classification on histology images but do not perform as well on cytology, whose stains and cell morphology differ markedly compared to histology. Conditional random fields (CRFs) can refine noisy VLM predictions by propagating information across patches, but existing CRF frameworks were designed for histopathology and do not transfer to cytology datasets, released as independent patch pools spanning multiple staining protocols. We introduce CytoCRF, which adapts the pairwise terms to cytology by targeting chromatin and cytology-specific staining, and further enrich the neighborhood of each potential term by combining multiple backbones. Across ten cytology datasets, CytoCRF outperforms existing CRF frameworks at every annotation budget, reaching +13.6 percentage points over the best baseline and +33.7 over ZS with only 50 annotations. Combining information from multiple backbones brings further gains, showing that the neighborhood topology matters more than the pairwise potential computed over it.
Chinese Translation
视觉-语言模型(VLM)在组织学图像上实现了强大的零样本(ZS)分类能力,但在细胞学上的表现却不尽如人意,因为细胞学的染色和细胞形态与组织学存在显著差异。条件随机场(CRF)可以通过在图像块之间传播信息来优化VLM的噪声预测,但现有的CRF框架是为组织病理学设计的,无法直接迁移到以独立图像块池形式发布、且跨越多种染色方案的细胞学数据集上。我们提出了CytoCRF,通过针对染色质和细胞学特异性染色来调整成对势项,使其适配细胞学任务,并通过结合多个骨干网络进一步丰富了每个势项的邻域信息。在十个细胞学数据集上,CytoCRF在所有标注预算条件下均优于现有的CRF框架,在仅使用50个标注的情况下,比最佳基线高出13.6个百分点,比零样本(ZS)方法高出33.7个百分点。融合多个骨干网络的信息带来了进一步提升,这表明邻域拓扑结构比在其上计算的成对势更为重要。
cs.CV / 49 / 2609.31040

Exploiting Spatial Structure for Transductive Few-Shot Classification of Whole-Slide Images

利用空间结构进行全切片图像的直推式少样本分类
Godelaine, Tiffanie, Dausort, Manon, Khoury, Karim El, Gérin, Benoît, Macq, Benoît, De Vleeschouwer, Christophe
Abstract
Automating the analysis of whole-slide images (WSIs), a key step in cancer diagnosis, has high clinical value, as it can reduce pathologist's workload while improving diagnosis accuracy. Recently, vision-language models have shown promising performance for patch-level classification without requiring any annotation, yet these zero-shot (ZS) predictions remain noisy on fine-grained tasks and must be further refined. A promising direction is to refine all predictions jointly, i.e., a transductive approach. However, most existing methods are not tailored to WSIs. We thus propose SlideTIM, an adaptation to WSIs of the recent transductive approach LC-TIM, which introduces a combined spatial--latent regularizer together with a prior on the patch class distribution. The former enforces spatially and semantically close patches to receive the same predictions, while the prior calibrates the predicted class proportions. Together, they address the complex spatial organization and the strong class imbalance of WSIs. Evaluated on four histology datasets, SlideTIM consistently outperforms all TIM variants, improving the macro-F1 by +8.1pp over the best competing baseline at 1 shot. Compared to the ZS, it raises the macro-F1 by +19.4pp at 1 shot. The code will be made available after submission.
Chinese Translation
全切片图像(Whole-Slide Images, WSI)分析自动化是癌症诊断的关键步骤,具有重要的临床价值,因为它可以在减轻病理医生工作负担的同时提高诊断准确性。近年来,视觉-语言模型在无需任何标注的情况下展现出良好的图像块(patch)级别分类性能,然而这些零样本(zero-shot, ZS)预测在细粒度任务上仍然存在噪声,需要进一步修正。一个有前景的方向是对所有预测进行联合修正,即直推式(transductive)方法。然而,现有的大多数方法并非针对WSI而设计。为此,我们提出了SlideTIM,这是对近期直推式方法LC-TIM在WSI上的适配,其引入了一个空间-潜空间联合正则化项以及图像块类别分布的先验。前者促使空间和语义上相近的图像块获得相同的预测,而先验则对预测的类别比例进行校准。二者相结合,共同应对了WSI复杂的空间组织和强烈的类别不平衡问题。在四个组织病理学数据集上的评估表明,SlideTIM持续优于所有TIM变体,在1-shot设置下相比最佳竞争基线将宏平均F1提升了8.1个百分点。与零样本方法相比,它在1-shot设置下将宏平均F1提升了19.4个百分点。代码将在提交后公开。
cs.CV / 50 / 2609.31050

Where Compute Matters: Heterogeneous Attention for Efficient Video Diffusion

计算资源应该用在哪里:面向高效视频扩散的异构注意力机制
Zatsarynna, Olga, Korzhenkov, Denis, Gall, Juergen, Habibian, Amir, Ghafoorian, Mohsen
Abstract
Efficient video generation requires reducing the quadratic cost of self-attention over long spatio-temporal token sequences. Existing efficient-attention methods typically apply the same computation pattern to every token, even though denoising difficulty varies substantially across video regions and evolves throughout the generation process. We introduce HetA-DiT, a heterogeneous attention mechanism that adaptively allocates computation according to token difficulty. A lightweight uncertainty branch predicts a token-wise estimate of denoising difficulty, which is used to route uncertain tokens through dense global attention while processing more reliable tokens with efficient local attention. The resulting routing is content- and timestep-adaptive, retains global context where it matters most, and provides a single parameter for controlling the quality-efficiency trade-off. HetA-DiT is compatible with few-step distribution-matching distillation and introduces no additional Transformer evaluation at inference time by reusing uncertainty estimates from the preceding denoising step. We evaluate the method on DMD-distilled Wan2.2-5B and Wan2.1-1.3B models. Across VBench, VBench-2.0, and human preference evaluation, HetA-DiT maintains competitive generation quality while routing only approximately 20% of tokens through dense attention.
Chinese Translation
高效的视频生成需要降低自注意力在长时空token序列上的二次方计算开销。现有的高效注意力方法通常对所有token采用相同的计算模式,然而去噪难度在不同视频区域之间差异很大,并在生成过程中不断演变。我们提出HetA-DiT,一种能够根据token难度自适应分配计算资源的异构注意力机制。一个轻量级的不确定性分支预测每个token的去噪难度估计,据此将不确定的token路由至稠密的全局注意力,而将更可靠的token交由高效的局部注意力处理。由此得到的路由方式具备内容自适应和时间步自适应特性,在最关键之处保留全局上下文,并提供一个用于控制质量与效率权衡的单一参数。HetA-DiT与少步分布匹配蒸馏兼容,并通过复用上一步去噪的不确定性估计,在推理时不引入额外的Transformer计算。我们在经过DMD蒸馏的Wan2.2-5B和Wan2.1-1.3B模型上评估了该方法。在VBench、VBench-2.0以及人类偏好评估中,HetA-DiT在仅有约20%的token经过稠密注意力处理的情况下,仍保持了有竞争力的生成质量。
cs.CV / 51 / 2609.31074

Band-Selection Stability and Semantic Segmentation Performance: A Study on Hyperspectral City

波段选择稳定性与语义分割性能:基于Hyperspectral City数据集的研究
Li, Jiarong, Shah, Imad Ali, Ward, Enda, Glavin, Martin, Jones, Edward, Deegan, Brian
Abstract
Resource constraints make high-dimensional hyperspectral imaging challenging in autonomous perception, motivating the use of band selection methods. However, the sensitivity of band-selection methods to sampled data and their relationship to semantic segmentation models (SSMs) remain underexplored. This study evaluates six band selection methods on ten independently sampled, class-balanced region-of-interest (ROI) sets, yielding 60 top-25 band subsets from the Hyperspectral City V2 (128 bands: 450-950nm) dataset. Top-$K$ bands ($K\in\{3,5, ... 13\}$) from the first three ROI sets are evaluated with three SSMs against the corresponding 128-band baseline. Experiments show that intra-method stability is method-dependent: Sim-LP shows the highest stability (pairwise Jaccard similarity) and, together with JMIM+CSNR, yields the best segmentation results. Top-$K$ based SSMs remain competitive with baselines, with gains of up to 2.01 mIoU and 1.72 mF1 points, and 18-22x faster CPU inference for $K=9$. However, performance does not improve monotonically with $K$, and stability shows no consistent association with SSM performance. These findings suggest that intra-method stability is informative but an unreliable indicator of downstream segmentation performance, highlighting the need to evaluate band-selection methods across repeated samples, subset sizes, and SSMs.
Chinese Translation
资源限制使得高维高光谱成像在自主感知中面临挑战,从而推动了对波段选择方法的应用。然而,波段选择方法对采样数据的敏感性及其与语义分割模型(SSMs)之间的关系仍未得到充分研究。本研究在十个独立采样且类别均衡的感兴趣区域(ROI)集合上评估了六种波段选择方法,并从Hyperspectral City V2(128个波段:450-950nm)数据集中获得了60个前25波段子集。来自前三个ROI集合的前$K$个波段($K\in\{3,5, ... 13\}$)结合三种语义分割模型进行评估,并与相应的128波段基线进行对比。实验表明,方法内部稳定性因方法而异:Sim-LP表现出最高的稳定性(成对Jaccard相似度),并且与JMIM+CSNR一起取得了最佳的分割结果。基于前$K$波段的语义分割模型仍与基线具有竞争力,mIoU和mF1分别最高提升2.01和1.72个百分点,并且在$K=9$时CPU推理速度提升18-22倍。然而,性能并未随$K$单调提升,且稳定性与语义分割模型性能之间不存在一致的相关性。这些发现表明,方法内部稳定性虽具有一定参考价值,但并非下游分割性能的可靠指标,凸显了在重复采样、不同子集规模及不同语义分割模型下评估波段选择方法的必要性。
cs.CV / 52 / 2609.31103

DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models

DepthEvidence:在多模态语言模型中统一度量深度预测与几何推理
Wei, Jiangning, Yao, Yuan, Cui, Miaomiao, Li, Mingsheng, Zhong, Humen, Bai, Shuai, Yang, Zhibo
Abstract
Spatial reasoning with metric constraints requires linking objects to geometric measurements and preserving their numerical content during language reasoning. We present DepthEvidence, a 4B model that uses its own dense metric predictions as object-grounded evidence for language generation. A camera-conditioned decoder predicts full-resolution metric depth using multi-scale visual features and high-resolution RGB refinement. A dense-to-language interface converts predicted depths and decoder features into object-aligned continuous geometry tokens anchored to object identifiers. Geometric supervision encourages metric information to remain recoverable before and after language-context interaction, while instruction tuning supports object measurement and compositional reasoning. We introduce a Depth-VQA benchmark evaluating object-depth queries, relative comparisons, and decisions combining spatial and numerical constraints. Across nine datasets, DepthEvidence achieves the highest average dense $\delta_1$ among evaluated methods, competitive with specialized estimators. It also leads the evaluated methods in instance-level metric depth estimation and overall accuracy on both relative and metric reasoning tracks, while broadly preserving general VQA performance and improving spatial understanding relative to the base model.
Chinese Translation
带有度量约束的空间推理需要将对象与几何测量相关联,并在语言推理过程中保持其数值内容。我们提出了DepthEvidence,一个40亿参数(4B)的模型,它将自身的稠密度量预测作为对象接地的证据用于语言生成。一个以相机为条件的解码器利用多尺度视觉特征和高分辨率RGB细化来预测全分辨率度量深度。一个从稠密预测到语言的接口将预测的深度和解码器特征转换为以对象标识符为锚点的、与对象对齐的连续几何token。几何监督促使度量信息在语言上下文交互前后保持可恢复性,而指令微调则支持对象测量和组合推理。我们引入了一个Depth-VQA基准,用于评估对象深度查询、相对比较以及结合空间与数值约束的决策。在九个数据集上,DepthEvidence在所评估的方法中取得了最高的平均稠密δ₁指标,与专门的深度估计器相当。在实例级度量深度估计方面,以及在相对推理和度量推理两条赛道上的总体准确率方面,它均领先于所评估的方法,同时大体上保持了通用视觉问答(VQA)性能,并相较于基础模型提升了空间理解能力。
cs.CV / 53 / 2609.31108

Double-stream registration with pyramid fusion for HDR video with alternating exposures

基于金字塔融合的交替曝光HDR视频双流配准方法
Martorell, Onofre, Pereira-Sánchez, Ivan, Fuentes, Antoni, Buades, Antoni
Abstract
High dynamic range (HDR) video reconstruction from al\-ter\-na\-ting-exposure sequences remains challenging, especially in regions with extreme luminance variation. We propose a novel HDR reconstruction framework based on dual-stream registration and accurate pyramid fusion. Given three consecutive frames, our method computes optical flow directly with the central frame, while introducing a complementary midpoint displacement strategy to handle cases with severe overexposition. A pyramid fusion stage then merges the resulting radiance and LDR images into a final HDR output. Experimental results demonstrate that our approach consistently outperforms state-of-the-art methods.
Chinese Translation
从交替曝光序列中重建高动态范围(HDR)视频仍然是一个具有挑战性的问题,尤其是在亮度变化剧烈的区域。我们提出了一种基于双流配准和精确金字塔融合的新型HDR重建框架。给定三个连续帧,我们的方法直接以中心帧为参考计算光流,同时引入一种互补的中点位移策略来处理严重过曝光的情况。随后,金字塔融合阶段将所得到的辐射图像与LDR图像合并,生成最终的HDR输出。实验结果表明,我们的方法始终优于当前最先进的方法。
cs.CV / 54 / 2609.31135

Pocket-STVG: lightweight architecture for Spatio-Temporal Video Grounding

Pocket-STVG:用于时空视频定位的轻量级架构
Presta, Alberto, Byra, Michal, Stefański, Grzegorz, Szurkowski, Karol, Kołodziejczyk, Eryk, Arendt, Krzysztof
Abstract
Spatio-Temporal Video Grounding (STVG) aims to localize the spatio-temporal tube in a video corresponding to a natural language query. While recent methods achieve strong performance in fully supervised, weakly supervised, and zero-shot settings, they typically rely on computationally expensive architectures, complex training pipelines, or multimodal large language models. We present Pocket-STVG (P-STVG), a lightweight cascade architecture that addresses STVG by combining efficient pre-trained components instead of large end-to-end models. P-STVG integrates a temporal-aware video encoder based on MobileViCLIP, a spatial encoder-decoder derived from MDETR, and a shared aligned text encoder. Temporal localization is performed through either a lightweight 1D U-Net or a simple thresholding strategy, enabling the same framework to operate in both weakly supervised and zero-shot settings. Furthermore, video representations are precomputed independently of the query, yielding an indexing-friendly pipeline for efficient inference and large-scale video collections. Despite requiring fewer than 90M parameters, P-STVG performs on par with weakly supervised methods and improves on earlier zero-shot approaches at a fraction of their memory and computational cost, establishing a favorable performance-efficiency trade-off for STVG.
Chinese Translation
时空视频定位(Spatio-Temporal Video Grounding, STVG)旨在根据自然语言查询在视频中定位相应的时空管(spatio-temporal tube)。尽管近期方法在全监督、弱监督和零样本设置下均取得了优异性能,但它们通常依赖计算开销高昂的架构、复杂的训练流程或多模态大语言模型。我们提出了 Pocket-STVG(P-STVG),这是一种轻量级级联架构,通过组合高效预训练组件而非大型端到端模型来解决 STVG 任务。P-STVG 集成了基于 MobileViCLIP 的时序感知视频编码器、源自 MDETR 的空间编码器-解码器,以及共享的对齐文本编码器。时序定位通过轻量级 1D U-Net 或简单的阈值策略实现,使同一框架能够在弱监督和零样本两种设置下运行。此外,视频表示的预计算与查询无关,从而形成了一个适合索引的高效推理流水线,可应用于大规模视频集合。尽管所需参数不足 9000 万,P-STVG 的性能与弱监督方法相当,并优于早期的零样本方法,而其内存和计算成本仅为后者的一小部分,为 STVG 建立了良好的性能-效率权衡。
cs.CV / 55 / 2609.31148

Seeing Semantic Shift: Difference-Aware Sentence-Level Temporal Segmentation of Sign Language Videos

感知语义变化:手语视频的差异感知句子级时序分割
Guo, Bowen, Gan, Shiwei, Yin, Yafeng, Liu, Xiao, Liu, Kuizhuang, Jiang, Zhiwei, Xie, Lei
Abstract
Recent advances in sign language understanding have achieved impressive success on short, single-sentence videos, yet their performance drops sharply when applied to long, continuous sign language videos. To bridge this gap, we focus on a challenging and realistic setting: Visual-only Sentence-level Sign Language Segmentation (Vis-SSLS), which aims to partition continuous sign language videos into non-overlapping sentence-level segments without any caption assistance, serving as a crucial prerequisite for downstream recognition and translation tasks. However, sentence transitions in sign language are often smooth and visually ambiguous, lacking explicit pauses or posture resets. As a result, static frame representations may fail to capture the subtle temporal changes that indicate sentence boundaries. To address this challenge, we propose \textbf{SignShift}, a difference-aware segmentation framework that explicitly models frame-to-frame feature variation as semantic cues for sentence boundary detection. First, to model the feature variation, we design a Temporal Difference Module, which incorporates full-frame, facial, and hand cues, and employs inter-frame differencing to learn multi-scale temporal variations that capture both fine-grained local kinematics and global semantic transitions. Second, to mitigate over- and under-segmentation issues, we design a Segment Count Prediction module, which predicts the number of sentences to guide boundary selection. Extensive experiments on benchmark datasets demonstrate that SignShift substantially outperforms existing methods, validating its effectiveness.
Chinese Translation
手语理解领域的最新进展在短的单句视频上取得了令人瞩目的成功,然而当应用于长时连续手语视频时,其性能急剧下降。为弥合这一差距,我们聚焦于一个具有挑战性且贴近现实的任务设定:仅视觉句子级手语分割(Visual-only Sentence-level Sign Language Segmentation,Vis-SSLS),其目标是在没有任何字幕辅助的情况下,将连续手语视频划分为互不重叠的句子级片段,作为下游识别与翻译任务的关键前提。然而,手语中的句子过渡往往平滑且在视觉上模糊,缺乏明确的停顿或姿态重置。因此,静态帧表示可能无法捕捉指示句子边界的细微时序变化。为应对这一挑战,我们提出SignShift——一个差异感知的分割框架,它显式地对逐帧特征变化进行建模,将其作为句子边界检测的语义线索。首先,为建模特征变化,我们设计了时序差分模块(Temporal Difference Module),该模块融合整帧、面部和手部线索,并采用帧间差分来学习多尺度时序变化,从而同时捕捉细粒度的局部运动和全局语义转换。其次,为缓解过度分割和欠分割问题,我们设计了片段数量预测模块(Segment Count Prediction module),通过预测句子数量来指导边界选择。在基准数据集上的大量实验表明,SignShift 显著优于现有方法,验证了其有效性。
cs.CV / 56 / 2609.31150

FedHisto-PAST: Parameter-Efficient Stain-Aware Federated Learning for Cross-Site Lung Histopathology Classification

FedHisto-PAST:面向跨机构肺组织病理学分类的染色感知参数高效联邦学习方法
Shahriar, Muhammad Muhtasim, Hafiz, M. M. Golam, Aloteibi, Saad, Moni, Mohammad Ali
Abstract
Cross-site lung histopathology classification must account for stain variation, non-IID client data, missing classes, and the cost of adapting large pathology encoders. This study evaluates FedHisto-PAST v2 for three-way classification of adenocarcinoma (ACA), Normal, and squamous cell carcinoma (SCC). FedHisto-PAST v2 combines a frozen HIBOU-B foundation model with parameter-efficient adaptation, stain-conditioned paired-view prediction and feature consistency, reliability-aware prototype learning, and adaptive federated aggregation. Experiments used a five-client, non-IID, raw-data-local simulation with fixed internal evaluation, client-level analysis, component ablations, communication accounting, and a development-influenced exploratory LungHist700 cohort. All principal methods achieved near- ceiling internal performance, which limited discrimination on the fixed split. On LungHist700, FedHisto- PAST v2 achieved a Macro-F1 of 0.728560 and a balanced accuracy of 0.730454. Higher recognition of Normal and SCC was accompanied by lower ACA recall, and calibration remained imperfect. Prediction-level consistency was the only component with a clearly supported independent contribution in the external ablation analysis. Feature consistency and prototype regularization showed no conclusive independent overall gains in Macro-F1. The framework updated 1.253841% of the model parameters. The results provide exploratory cross-dataset evidence for stain-aware, parameter-efficient federation; they do not establish formal privacy, patient-level independence, prospective deployment, or clinical validation.
Chinese Translation
跨机构肺组织病理学分类必须考虑染色变异、非独立同分布(non-IID)的客户端数据、类别缺失以及适配大型病理学编码器的成本。本研究评估了 FedHisto-PAST v2 在腺癌(ACA)、正常组织和鳞状细胞癌(SCC)三分类任务中的表现。FedHisto-PAST v2 将冻结的 HIBOU-B 基础模型与参数高效适配、染色条件下的配对视图预测与特征一致性、可靠性感知的原型学习以及自适应联邦聚合相结合。实验采用了五个客户端、非独立同分布、原始数据本地化的模拟设置,并进行了固定的内部评估、客户端层面分析、组件消融实验、通信量统计,以及一个受开发过程影响的探索性 LungHist700 数据集评估。所有主要方法在内部评估中均达到接近上限的性能,这限制了在固定数据划分上的区分能力。在 LungHist700 数据集上,FedHisto-PAST v2 的 Macro-F1 达到 0.728560,平衡准确率达到 0.730454。对正常组织和 SCC 的识别率提升伴随着 ACA 召回率的下降,且校准仍不完善。在外部消融分析中,预测级别的一致性是唯一具有明确支持的独立贡献的组件。特征一致性和原型正则化在 Macro-F1 上未显示出确定性的独立整体增益。该框架仅更新了 1.253841% 的模型参数。这些结果为染色感知、参数高效的联邦学习提供了探索性的跨数据集证据;但并未建立正式的隐私保护、患者级独立性、前瞻性部署或临床验证。
cs.CV / 57 / 2609.31154

HyperErase: Scale-Calibrated Hypernetwork for Multi-Concept Erasure in Text-to-Image Models

HyperErase:用于文本到图像模型多概念擦除的尺度校准超网络
Sun, Yi, Zhong, Xinhao, Zhang, Zhiqi, Zhou, Yimin, Li, Junhao, Qiao, Yuxia
Abstract
Recent advances in text-to-image (T2I) generation have substantially improved visual synthesis, but have also raised increasing safety concerns due to their potential to generate harmful or undesirable content. Existing concept erasure methods predominantly follow a static weight paradigm, producing a single frozen adapter that struggles to adapt to diverse prompt variations and suffers from parameter interference when scaling to multiple concepts. We propose \textbf{HyperErase}, a framework for concept erasure based on hypernetwork-driven prompt-conditioned parameter synthesis. Our approach first reframes concept erasure as prompt-conditioned parameter amortization and trains a hypernetwork to map textual descriptions to prompt-specific LoRA updates, eliminating the need for per-prompt gradient optimization or manual LoRA merging. To further improve the stability and precision of synthesized adapters, we develop a decoupled rectification strategy, which disentangles LoRA tokens into pattern and scale subspaces, applies a square-root transform to curb multiplicative over-scaling, and leverages teacher-derived canonical priors for inference-time correction. Extensive experiments across major concept categories demonstrate that HyperErase consistently improves the trade-off between erasure effectiveness, image quality, and semantic alignment, achieving performance comparable to gold-standard single-concept baselines. Furthermore, the resulting models can provide specialized LoRAs for each input prompt variation in a single forward pass without requiring gradient updates during inference. These principled and flexible framework offers a new paradigm for concept erasure in T2I models.
Chinese Translation
文本到图像(T2I)生成领域的最新进展显著提升了视觉合成能力,但也因其可能生成有害或不良内容而引发了越来越多的安全担忧。现有的概念擦除方法主要遵循静态权重范式,仅产生单一的冻结适配器,难以适应多样化的提示词变化,且在扩展至多个概念时会遭受参数干扰问题。我们提出HyperErase,一个基于超网络驱动的提示词条件参数合成的概念擦除框架。我们的方法首先将概念擦除重新表述为提示词条件参数摊销问题,并训练一个超网络将文本描述映射为特定于提示词的LoRA更新,从而免除了逐提示词的梯度优化或手动LoRA合并的需要。为进一步提升所合成适配器的稳定性与精确性,我们开发了一种解耦校正策略:将LoRA令牌解耦为模式子空间与尺度子空间,施加平方根变换以抑制乘性过度缩放,并利用教师模型导出的规范先验进行推理时校正。在多个主要概念类别上的大量实验表明,HyperErase持续改善了擦除有效性、图像质量与语义对齐之间的权衡,性能可与金标准单概念基线相媲美。此外,所得模型能够在单次前向传播中为每个输入提示词变体提供专用的LoRA,而无需在推理期间进行梯度更新。这一原理清晰且灵活的框架为T2I模型中的概念擦除提供了一种新范式。
cs.CV / 58 / 2609.31160

ReG-SAM: Reference Graph-Driven SAM for 2D Foundational Vessel Segmentation

ReG-SAM:基于参考图驱动的SAM用于2D血管分割基础模型
Lyu, Donghang, Zhang, Zichen, Dzyubachyk, Oleh, Staring, Marius
Abstract
Vessel segmentation in medical images is essential for many clinical tasks, ranging from diagnosis to treatment planning. However, it remains challenging due to complex vascular morphology and diverse imaging conditions. Existing deep learning methods rarely aim at building a generalizable vessel segmentor across anatomies and modalities. While the Seg- ment Anything Model (SAM) has shown promise for med- ical image segmentation, its original design does not fully exploit vascular morphology and struggles with fine-grained vascular structures, leading to suboptimal performance. In this paper, we propose ReG-SAM, a SAM-based framework tailored to 2D vessel segmentation that leverages reference graph set for enhancing vascular representations. Specifically, we introduce two modality-aware representations derived from the reference masks: graph prompt embeddings (GPEs) that encode global spatial features from graphs, and vascu- lar prototype embeddings (VPEs) that capture fine-grained modality-specific vessel characteristics from multi-scale fea- ture maps and vascular masks. Since both require vascular masks that are unavailable during inference and require robust modality-aware vascular feature representations, we construct a modality-wise vascular database and develop two reference graph-guided representation learning schemes for estimating GPEs and VPEs using samples from the database rather than ground-truth masks. Extensive experiments across 19 datasets demonstrate that ReG-SAM consistently outperforms existing baselines, even those using manual prompts, particularly on challenging thin vessels.
Chinese Translation
医学图像中的血管分割对于从诊断到治疗规划的许多临床任务至关重要。然而,由于复杂的血管形态和多样的成像条件,血管分割仍然具有挑战性。现有的深度学习方法很少致力于构建跨解剖结构和成像模态的通用血管分割器。尽管Segment Anything Model(SAM)在医学图像分割方面展现出潜力,但其原始设计并未充分利用血管形态,且难以处理细粒度的血管结构,导致性能欠佳。在本文中,我们提出了ReG-SAM,一种专为2D血管分割设计的基于SAM的框架,它利用参考图集合来增强血管表征。具体而言,我们引入了两种从参考掩模中导出的模态感知表征:编码图中全局空间特征的图提示嵌入,以及从多尺度特征图和血管掩模中捕获细粒度模态特异性血管特征的血管原型嵌入。由于两者都需要在推理阶段不可获得的血管掩模,并且需要鲁棒的模态感知血管特征表征,我们构建了一个按模态划分的血管数据库,并开发了两种参考图引导的表征学习方案,利用数据库中的样本而非真实标注掩模来估计GPE和VPE。在19个数据集上的大量实验表明,ReG-SAM持续优于现有基线方法,甚至优于使用人工提示的方法,尤其是在具有挑战性的细小血管上。
cs.CV / 59 / 2609.31170

TaskIR: Task-Driven Image Restoration via Degradation Adaptation and Task Feedback

TaskIR:基于退化自适应与任务反馈的任务驱动图像修复
Tu, Yanjie, Yan, Qingsen, Niu, Axi, Cai, Wenxuan, Hu, Tao, Dong, Wei, Zhang, Haokui
Abstract
Task-driven image restoration aims to improve both image quality and downstream task performance. However, existing methods predominantly focus on single degradation type and struggle to handle the diverse degradations encountered in real-world scenarios. Different degradations impose distinct restoration demands, and insufficient restoration may leave residual degradations and artifacts that impair object boundaries and semantic cues, thereby compromising downstream task performance. To address these challenges, we propose TaskIR, a two-stage task-driven unified image restoration framework that integrates degradation-adaptive restoration with task feedback refinement. In Stage I, a Degradation Representation Module (DRM) extracts degradation representations, enabling a Degradation-Guided Transformer Block (DGTB) to dynamically modulate feature transformations for adaptive restoration. In Stage II, a Task-to-Restoration Feedback Generation module (TRFG) transforms heterogeneous task features into restoration feedback by modeling task-representation discrepancies associated with the current restoration. Subsequently, a Selective Task Feedback Refinement module (STFR) assesses feedback relevance and selectively refines intermediate restoration features to mitigate interference with well-restored content. Extensive experiments demonstrate that TaskIR achieves competitive restoration quality and downstream task performance across diverse degradations and tasks.
Chinese Translation
任务驱动的图像修复旨在同时提升图像质量和下游任务性能。然而,现有方法主要针对单一退化类型,难以应对真实场景中遇到的多样化退化。不同退化对修复提出不同要求,修复不充分可能留下残余退化和伪影,损害物体边界和语义线索,从而降低下游任务性能。为解决这些挑战,我们提出TaskIR,一个两阶段的任务驱动统一图像修复框架,将退化自适应修复与任务反馈精化相融合。在第一阶段,退化表示模块(DRM)提取退化表示,使退化引导Transformer块(DGTB)能够动态调制特征变换以实现自适应修复。在第二阶段,任务到修复的反馈生成模块(TRFG)通过建模与当前修复结果相关的任务表示差异,将异构任务特征转化为修复反馈。随后,选择性任务反馈精化模块(STFR)评估反馈的相关性,并有选择地精化中间修复特征,以避免对已修复良好内容的干扰。大量实验表明,TaskIR在多样化退化和任务场景下均取得了具有竞争力的修复质量和下游任务性能。
cs.CV / 60 / 2609.31193

Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs

谁说了什么:音视频大语言模型中的符号化三模态绑定机制
Jung, Jihoo, Jang, Youngjoon, Chung, Joon Son
Abstract
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual components into symbolic variables-capturing temporal utterance sequences and spatial entity coordinates, respectively-the model establishes cross-modal linking within this abstract space. Crucially, we reveal that when trimodal binding fails, the breakdown predominantly stems from misaligned audio-visual connections. To overcome this bottleneck, we introduce an audio-visual prompting method utilizing an off-the-shelf Active Speaker Detection (ASD) model. By simply overlaying visual bounding boxes on active speakers, this training-free approach yields immediate performance gains across four conversation-centric benchmarks. Moreover, lightweight fine-tuning of fewer than 300 steps on these ASD-prompted-videos extends these gains to three general AV benchmarks, suggesting the generalizability of our method.
Chinese Translation
当前的音视频大语言模型(AVLLMs)在处理包含多说话人对话的视频推理时面临困难。在此类视频中,解决“谁说了什么”的问题至关重要,这需要文本-音频-视觉三模态的绑定。基于这些挑战,我们系统地研究了AVLLMs中三模态绑定是如何实现的。具体而言,我们识别出AVLLMs中涌现出一种利用模态特定符号变量的符号化三模态绑定机制。通过将听觉和视觉组件分别编码为符号变量——分别捕捉时序上的话语序列和空间上的实体坐标——模型在这一抽象空间内建立跨模态的关联。关键的是,我们揭示了当三模态绑定失败时,其失效主要源于音视觉连接的错位。为克服这一瓶颈,我们引入了一种利用现成的主动说话人检测(Active Speaker Detection, ASD)模型的音视觉提示方法。通过简单地在主动说话人上叠加视觉边界框,这种无需训练的方法在四个以对话为中心的基准上带来了即时的性能提升。此外,仅用不到300步对这些ASD提示的视频进行轻量级微调,即可将这些收益扩展到三个通用音视频基准上,这表明我们的方法具有良好的泛化能力。
cs.CV / 61 / 2609.31198

Light Field Primitive for Novel View Synthesis

用于新视角合成的光场基元
Chen, Liang, Ning, Jiahui, Jiang, Xun, Xu, Xing, Ren, Jimmy, Fan, Fenglei, Shen, Heng Tao
Abstract
We present Light Field Primitives (LFP), a formulation for novel view synthesis that replaces the dense ray database with a compact set of differentiable primitives in the classical two-plane parameterization. Each primitive condenses a group of rays into one learned record, and its response to a query is governed by how closely that query belongs to the group. Rendering a camera ray then reduces to compositing all responses it elicits, and a scene can be optimized directly from posed images and rendered in real time with rays. Beyond its competitive performance on standard benchmarks, the main advantage of LFP is structural: its primitives reside directly in the 4D ray space, so optical and appearance effects that are already operations on the light field become behaviors of a single shared renderer. With minimal changes to that renderer, LFP supports multi-scale anti-aliasing, defocus deblurring with refocusing, rendering for fisheye cameras, and even transparent object reconstruction with ray refraction, matching specialized frameworks that devote substantial machinery to these effects.
Chinese Translation
我们提出了光场基元(Light Field Primitives, LFP),这是一种新视角合成的形式化方法,它在经典的双平面参数化中用一组紧凑的可微基元取代了稠密的光线数据库。每个基元将一组光线凝聚为一条学习得到的记录,其对查询的响应由该查询与这组光线的归属程度决定。渲染一条相机光线因此简化为合成它所引发的所有响应,而场景可以直接从带位姿的图像中优化得到,并能以光线进行实时渲染。除了在标准基准上具有竞争力的性能外,LFP 的主要优势在于其结构性:其基元直接位于 4D 光线空间中,因此那些本质上就是对光场进行操作的视觉与外观效果,可以成为单一共享渲染器的行为。只需对该渲染器做极少的修改,LFP 即可支持多尺度抗锯齿、结合重聚焦的散焦去模糊、鱼眼相机渲染,甚至通过光线折射实现透明物体重建,其效果可与专门为这些效果投入大量机制的专用框架相媲美。
cs.CV / 62 / 2609.31202

Preserve-and-Compose Training for Composed Image Retrieval

面向组合图像检索的保留-组合训练
Kwon, Sehyun
Abstract
Composed image retrieval (CIR) aims to retrieve images that satisfy a user-specified modification while preserving relevant visual content from a reference image. Collecting target images for this purpose is costly, motivating zero-shot CIR methods that instead use target captions as supervision. However, target captions may omit source details that should be preserved. We therefore propose, Preserve-and-Compose Training, which complements target-caption supervision with visual evidence from the source image. PACT learns from image--text--text (ITT) triplets without target images or gallery updates, aligning composed queries with target captions while preserving source evidence through visual supervision. We further introduce Chord scoring, which combines target similarity with source-relative directional agreement in the frozen image space. Results across four ZS-CIR benchmarks show that combining target-caption supervision with source-image evidence leads to strong retrieval performance across datasets, backbone scales, and external galleries. The code is available on https://github.com/sehyunkwon/PACT.
Chinese Translation
组合图像检索(Composed Image Retrieval, CIR)旨在检索满足用户指定修改要求的图像,同时保留参考图像中的相关视觉内容。为此收集目标图像的成本高昂,这促使了零样本CIR(zero-shot CIR)方法的发展,即利用目标图像的文本描述作为监督信号。然而,目标描述可能遗漏应当保留的源图像细节。因此,我们提出了保留-组合训练(Preserve-and-Compose Training,PACT),利用源图像的视觉证据来补充目标描述的监督。PACT 从图像-文本-文本(image–text–text, ITT)三元组中学习,无需目标图像或更新图库,在将组合查询与目标描述对齐的同时,通过视觉监督保留源图像证据。我们进一步提出了 Chord 评分方法,在冻结的图像空间中将目标相似度与相对源图像的方向一致性相结合。在四个零样本CIR基准上的结果表明,将目标描述监督与源图像证据相结合,能够在不同数据集、不同骨干网络规模以及外部图库条件下均取得出色的检索性能。代码已发布于 https://github.com/sehyunkwon/PACT。
cs.CV / 63 / 2609.31234

WeaveAgent: A Two-Stage Tool-Routing Agent for Ultra-High-Resolution Remote Sensing Imagery

WeaveAgent:面向超高分辨率遥感影像的两阶段工具路由智能体
Pang, Zhongyu
Abstract
Problem. Ultra-high-resolution (UHR) remote sensing with vague user intents has two bottlenecks: visual tokens are expensive, and tool calling must be format-reliable (pretrained models emit zero tool calls zero-shot). Method. WeaveAgent, a two-stage tool-routing agent, decouples routing from visual perception. Stage A is routing-first: emission is trained, not elicited. Stage B executes conditionally: intrinsic queries enter visual answering (full-scene thumbnail; a WeaveEarth-style evidence board as an optional fixed-budget, approx. 5k-token compression interface); extrinsic queries execute tool call on original full-resolution imagery, answering from tool observations in a second, observation-masked round. Training: alignment SFT, then GRPO under reward R_WA2. Results. Alignment SFT lifts extrinsic routing from 0% to 80.75% (323/400); GRPO suppresses 9 intrinsic mis-emissions while tool selection is unchanged. The trained 2B system does not beat the zero-shot 8B baseline overall (0.263 vs. 0.250), a diagnostic contribution. Oracle attribution separates two repair ingredients: loading the observation into context lifts extrinsic answer accuracy from 0.025 to 0.425 under marker-free cross-mode returns, and the two-turn SFT stage adds a further +9.3 points to 0.518 at a small routing cost. A +/- image ablation shows emission suppression is visually grounded, and a query-register matrix shows LLM-rewritten queries cost trained checkpoints 2-11 points. Scope. All training and evaluation use the 5,000 / 3,273 / 1,000-record VagueUHR corpus (600 intrinsic + 400 tool-requiring; the base seeds synthesis and is not used for optimization). Single-pass evidence construction runs at 7.31 s per image on an RTX 4090. Code, data, and evaluation protocols will be released.
Chinese Translation
问题。针对用户意图模糊的超高分辨率(UHR)遥感任务,存在两个瓶颈:视觉token开销高昂,且工具调用必须具备格式可靠性(预训练模型在零样本情况下无法发出任何工具调用)。方法。WeaveAgent是一个两阶段工具路由智能体,将路由与视觉感知解耦。阶段A为路由优先:工具调用的发出是通过训练获得的,而非诱导产生的。阶段B为条件执行:内在查询进入视觉问答(全场景缩略图;以及一个可选的、固定预算(约5千token)压缩接口,即WeaveEarth式证据板);外在查询则在原始全分辨率影像上执行工具调用,并在第二轮中基于工具观测结果作答,此时观测内容被遮蔽。训练:先进行对齐SFT,再在奖励R_WA2下进行GRPO。结果。对齐SFT将外在路由准确率从0%提升至80.75%(323/400);GRPO消除了9次内在误发,同时工具选择保持不变。训练后的2B系统总体上未能超越零样本8B基线(0.263对0.250),这构成了一个诊断性贡献。Oracle归因区分了两种修复要素:在无标记跨模式返回条件下,将观测内容载入上下文可使外在答案准确率从0.025提升至0.425;两轮SFT阶段在此基础上进一步提升9.3个百分点至0.518,仅付出较小的路由代价。图像加减消融实验表明,调用发出的抑制具备视觉依据;查询寄存器矩阵显示,LLM改写的查询会使训练后的模型损失2至11个百分点。范围。所有训练与评估均使用包含5,000/3,273/1,000条记录的VagueUHR语料库(600条内在查询+400条需工具调用;基础数据仅用于种子合成,不用于优化)。在RTX 4090上,单次证据构建速度为每幅图像7.31秒。代码、数据与评估协议将予发布。
cs.CV / 64 / 2609.31247

Geometric Inconsistency Localization in Multi-View Image Sets

多视图图像集合中的几何不一致性定位
Staelens, Xander, Loos, Albéric, Ramlot, Bert, Mareen, Hannes, Lambert, Peter, Van Wallendael, Glenn
Abstract
Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for evaluating these NVS models. Its potential for multimedia forensics, however, remains largely unexplored, particularly for localizing geometric inconsistencies across wide-baseline image pairs. To enable research in this direction, we introduce DeformView, a wide-baseline MV dataset with pixel-level annotations of geometric inconsistencies. Using DeformView, we evaluate state-of-the-art MV consistency-scoring methods and show that approaches developed for NVS evaluation transfer poorly to the forensic task of geometric inconsistency localization. To address this limitation, we propose DEFECt3R, a lightweight learning-based classifier that uses cross-view feature relationships to localize geometric inconsistencies at the pixel level. By learning from explicit supervision, including hard negatives from geometrically consistent yet deformed views, DEFECt3R improves localization performance and substantially reduces false positives compared to existing consistency-scoring methods. Ablation experiments further show that both feature representations and correspondence quality contribute to localization performance. Overall, our findings demonstrate that MV geometric consistency is a promising yet underexplored signal for multimedia forensics and establish a benchmark and baseline for geometric inconsistency localization in wide-baseline MV image pairs. Code and dataset are available at https://github.com/IDLabMedia/DeformView-DEFECt3R
Chinese Translation
新视角合成(Novel View Synthesis, NVS)模型能够从不同视点生成同一场景的逼真新视图。然而,这些生成的视图之间并不总是保持几何一致性。多视图(Multi-View, MV)一致性已被证明是评估此类NVS模型的有效工具,但其在多媒体取证领域的潜力仍未被充分探索,尤其是在宽基线图像对之间定位几何不一致性方面。为推动这一方向的研究,我们引入了DeformView,一个带有像素级几何不一致性标注的宽基线多视图数据集。基于DeformView,我们评估了当前最先进的多视图一致性评分方法,结果表明,为NVS评估而开发的方法在迁移到几何不一致性定位这一取证任务时表现不佳。为解决这一局限,我们提出了DEFECt3R,一种轻量级的基于学习的分类器,它利用跨视图特征关系在像素级定位几何不一致性。通过从显式监督中学习——包括来自几何上一致但发生形变的视图的困难负样本——DEFECt3R相较于现有的一致性评分方法提升定位性能,并大幅减少了误报。消融实验进一步表明,特征表示和对应关系质量均对定位性能有所贡献。总体而言,我们的研究结果表明,多视图几何一致性是多媒体取证中一个有前景但尚未被充分挖掘的信号,并为宽基线多视图图像对中的几何不一致性定位建立了基准和基线方法。代码和数据集可在 https://github.com/IDLabMedia/DeformView-DEFECt3R 获取。
cs.CV / 65 / 2609.31248

Gauss What You Need: Compact Gaussian Splatting Across Scene Scales

精准所需的高斯:跨场景尺度的紧凑高斯泼溅
Boudaoud, Afif, Liu, Jiayi, Calotoiu, Alexandru, Hoefler, Torsten
Abstract
3D Gaussian Splatting reconstructs a scene as a collection of Gaussian primitives from a set of posed photographs called the capture. The number of primitives used to represent the scene affects reconstruction quality, storage, and rendering cost. How to select this number automatically across capture scales remains unresolved: configurations effective on standard benchmarks can leave larger captures with too few Gaussians to reconstruct fine details. We observe that the surface to represent, given by the capture's extent and resolution, is known before training, whereas its content complexity becomes apparent during training, through the reconstruction quality on the training views. We introduce TangoGS, which combines capture-derived model sizing with training-based adaptation: the capture determines the scale of the model, and training feedback determines its final size within that scale. Before training, TangoGS derives a learning allowance for model growth from the capture's total pixels after discounting views that re-observe the same scene points. During training, reconstruction quality guides how many Gaussians to add and remove. On 13 standard benchmark scenes, TangoGS matches the mean PSNR of the best-performing evaluated baseline, LeGS, with $48\%$ fewer Gaussians. On eight large captures, the same configuration automatically scales to larger models when necessary, achieving the highest mean PSNR among evaluated methods: $0.54$ dB above the runner-up with $2.3\times$ as many Gaussians. Together, capture-derived learning allowances and training-quality guided density control enable a state-of-the-art quality--size compromise across scene scales without retuning.
Chinese Translation
3D高斯泼溅(3D Gaussian Splatting)从一组带位姿的照片(称为采集数据)出发,将场景重建为一组高斯基元的集合。用于表示场景的基元数量会影响重建质量、存储开销和渲染成本。如何在不同的采集规模下自动选择这一数量仍未得到解决:在标准基准上有效的配置,可能使较大的采集数据因高斯数量过少而无法重建精细细节。我们观察到,待表示的表面由采集数据的范围和分辨率决定,在训练之前即已知,而其内容复杂度则在训练过程中通过训练视图上的重建质量显现出来。我们提出TangoGS,将基于采集数据的模型规模确定与基于训练的适应性调整相结合:采集数据决定模型的尺度,训练反馈决定模型在该尺度内的最终规模。在训练之前,TangoGS在扣除重复观测相同场景点的视图后,根据采集数据的总像素数推导出模型增长的学习配额。在训练过程中,重建质量指导高斯的添加与删除。在13个标准基准场景上,TangoGS以少48%的高斯数量达到了表现最佳的评估基线LeGS的平均PSNR。在八个大规模采集数据上,同一配置可在必要时自动扩展至更大的模型,在评估方法中取得最高的平均PSNR:比次优方法高0.54 dB,而后者使用的高斯数量是其2.3倍。总体而言,基于采集数据的学习配额和基于训练质量的密度控制,使TangoGS无需重新调参即可在跨场景尺度下实现最先进的质量—规模折中。
cs.CV / 66 / 2609.31285

MoTop: Motion-Topological Model For Micro AU Detection

MoTop:用于微动作单元检测的运动拓扑模型
Khor, Huai-Qian, Wei, Mengting, Li, Yante, Loo, Chu Kiong, Zhao, Guoying
Abstract
Facial micro-expressions are spontaneous, brief, and subtle facial movements that reveal suppressed emotions in high-stakes environments. In contrast to classic expression analysis, detecting action unit (AU) yields a finer representation of facial movements, serving as a preliminary step before defining expression classes and other downstream tasks. Therefore, it represents a crucial upstream task in facial analysis, and improving an AU detection module increases the precision of facial analysis. Despite that, detecting AU is challenging because of the constrictive nature of the AU activation regions, leading to confusion among different AUs known as AU ambiguity. To model the fine-scale changes, we propose \textbf{MoTop}, a motion-topological model that is augmented with a learnable motion context, yielding regional soft guidance for facial activity, followed by facial landmarks that capture the fine-scale topological changes of micro AUs. To increase the micro facial landmark representations, we amplify the encoded facial landmark transitions via linear extrapolation, thereby increasing the spatial proximity of landmarks and enhancing the low-intensity landmark dynamics. In addition, we design anatomical facial clusters that enhance the hierarchical representation, facilitating multi-scale modelling of facial geometry and improving micro-topological representations. With these contributions, we have achieved state-of-the-art performance on the CD6ME protocol for the micro AU detection task.
Chinese Translation
面部微表情是自发的、短暂的、细微的面部运动,能够揭示高压环境下被压抑的情绪。与经典的面部表情分析不同,检测动作单元(AU)能够提供对面部运动更细粒度的表征,可作为定义表情类别及其他下游任务之前的预备步骤。因此,它是面部分析中一个关键的上游任务,提升AU检测模块的精度将提高面部分析的整体精确性。然而,由于AU激活区域的局部约束特性,不同AU之间容易产生混淆,即所谓的AU歧义性,这使得AU检测极具挑战性。为建模这种细粒度变化,我们提出了MoTop——一种运动拓扑模型,该模型引入了可学习的运动上下文,为面部活动提供区域化的软引导,并结合面部关键点来捕捉微AU的细粒度拓扑变化。为增强微面部关键点表征,我们通过线性外推放大编码后的面部关键点变化,从而提高关键点的空间邻近性并增强低强度关键点的动态特征。此外,我们设计了基于解剖学的面部聚类,以增强层次化表征,促进面部几何结构的多尺度建模并改进微拓扑表征。凭借上述贡献,我们在微AU检测任务的CD6ME协议上取得了最先进的性能。
cs.CV / 67 / 2609.31298

UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning

UniAR:一种通过多视角提示学习增强的孤独症识别统一框架
Xin, Lei, Wang, Zeheng, Zhu, Jiayin, Huang, Shihong, Zeng, Fanhu, Jiang, Changjiang, He, Dengbo, Yue, Yutao, Kong, Zhenglun
Abstract
Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder for which early and accurate diagnosis is critical to improving long-term developmental outcomes. However, existing ASD recognition methods are often constrained by the scarcity of diagnostic text data, forcing them to rely mainly on visual analysis and limiting their ability to model clinically meaningful semantic reasoning. To address this challenge, we propose UniAR, a unified framework enhanced by multi-granularity prompt learning for robust ASD recognition under heterogeneous data variations. Specifically, UniAR leverages a large multimodal model to generate hierarchical diagnostic descriptions at the word, phrase, and sentence levels, compensating for the lack of paired clinical reports. To align the generated semantics with visual evidence, we further design a Mixture-of-Experts-based Multi-Scale Alignment Module, which dynamically matches vector-quantized visual prototypes with semantic representations at corresponding granularities. Extensive experiments on four benchmarks covering brain MRI and facial expression scenarios show that UniAR consistently outperforms existing state-of-the-art methods, achieving average accuracies of 75.9\% on MRI benchmarks and 91.6\% on facial benchmarks, while improving average Accuracy on MRI benchmarks by 1.5 percentage points and average Accuracy on facial benchmarks by 1.2 percentage points over baselines. These results demonstrate that UniAR offers a robust and interpretable framework for ASD screening under semantic scarcity.
Chinese Translation
自闭症谱系障碍(ASD)是一种复杂的神经发育障碍,早期准确的诊断对于改善长期发育结果至关重要。然而,现有的ASD识别方法常常受限于诊断文本数据的稀缺,不得不主要依赖视觉分析,从而限制了其建模临床有意义的语义推理的能力。为应对这一挑战,我们提出UniAR,一个通过多粒度提示学习增强的统一框架,用于在异构数据变化下实现稳健的ASD识别。具体而言,UniAR利用大型多模态模型生成词、短语和句子层级的分层诊断描述,以弥补成对临床报告的缺失。为使生成的语义与视觉证据对齐,我们进一步设计了基于专家混合(Mixture-of-Experts)的多尺度对齐模块,该模块在相应粒度上动态匹配向量量化视觉原型与语义表示。在涵盖脑部MRI和面部表情场景的四个基准数据集上的大量实验表明,UniAR持续优于现有最先进的方法,在MRI基准上平均准确率达到75.9%,在面部基准上平均准确率达到91.6%,相比基线方法,MRI基准的平均准确率提升1.5个百分点,面部基准的平均准确率提升1.2个百分点。这些结果表明,UniAR为语义稀缺条件下的ASD筛查提供了一个稳健且可解释的框架。
cs.CV / 68 / 2609.31314

CytoSPM: Open-Vocabulary Cytopathology Detection with Structured Prompt Bank

CytoSPM:基于结构化提示库的开放词汇细胞病理学检测
Li, Wenjie, Xu, Zishan, Huang, Jinyang, Nie, Zhengxin, Kan, Shichao, Liang, Yixiong
Abstract
Cytopathology detection requires open-vocabulary recognition because cellular categories are fine-grained, long-tailed, and continuously evolving across different organ systems. However, existing cytology detectors are mostly single-domain and closed-set, and there is still no unified benchmark for evaluating open-vocabulary cytopathology detection. We present PentaCyto, a multi-domain benchmark covering cervical, urinary, respiratory, serous fluid, and thyroid cytology, with 24 base categories and 9 held-out novel categories. Each category is associated with structured cytomorphology prompts that describe diagnostic morphological attributes and provide clinically grounded textual knowledge. We further propose CytoSPM, an efficient detector based on a decoupled two-stage design. It first extracts reusable class-agnostic visual representations, and then performs class-aware structural prompt matching with class names and cytomorphology prompts. On PentaCyto, CytoSPM outperforms existing methods in novel-category detection and open-vocabulary detection while maintaining efficient inference.
Chinese Translation
细胞病理学检测需要开放词汇识别能力,因为细胞类别具有细粒度、长尾分布的特点,且在不同器官系统中持续演化。然而,现有的细胞学检测器大多局限于单一领域和封闭类别集,目前仍缺乏用于评估开放词汇细胞病理学检测的统一基准。我们提出了PentaCyto,一个覆盖宫颈、泌尿、呼吸、浆膜腔积液和甲状腺细胞学的多领域基准,包含24个基础类别和9个留出的新类别。每个类别都关联了结构化的细胞形态学提示,用于描述诊断性形态特征并提供具有临床依据的文本知识。我们进一步提出了CytoSPM,一种基于解耦两阶段设计的高效检测器。它首先提取可复用的类别无关视觉表征,然后利用类别名称和细胞形态学提示进行类别感知的结构化提示匹配。在PentaCyto上,CytoSPM在新类别检测和开放词汇检测任务上优于现有方法,同时保持了高效的推理速度。
cs.CV / 69 / 2609.31326

CG-HAF: An Interpretable Global-Local Lesion-Burden Fusion Framework for Ordinal Acne Severity Grading in Agentic Skincare Support

CG-HAF:一种用于智能护肤支持中序数型痤疮严重程度分级的可解释全局-局部病灶负荷融合框架
Shahriar, Muhammad Muhtasim, Borno, Md. Naimur Asif, Aloteibi, Saad, Moni, Mohammad Ali
Abstract
Ordinal acne severity grading requires distinguishing visually similar neighboring grades while jointly weighing holistic facial appearance and localized lesion burden - evidence that most existing approaches collapse into a single opaque representation. We introduce CG-HAF, a global-local fusion framework that instead keeps this evidence explicit: averaged holistic severity probabilities from independently trained classifiers are combined with structured lesion-burden descriptors from an object detector (lesion count, detection confidence, lesion area) into a compact representation, from which a lightweight, interpretable classifier produces the final grade. On a widely used benchmark, this fusion yields a clear, statistically supported improvement over global-evidence-only baselines, with the largest gains on the most severe cases. Testing on an independent dataset with a different grading standard shows that strong within-dataset performance does not transfer automatically, and a follow-up diagnostic attributes much of this gap to mismatched grading criteria rather than detection failure alone. These findings support interpretable global-local fusion as an effective strategy for ordinal acne grading while highlighting criterion alignment as key to cross-dataset portability, with a further illustration of how the resulting severity signal can support transparent, non-diagnostic decision-making in skincare applications.
Chinese Translation
序数型痤疮严重程度分级需要在区分视觉上相近的相邻等级的同时,综合权衡整体面部外观与局部病灶负荷——而现有大多数方法往往将这两类证据压缩为单一的不透明表示。我们提出CG-HAF,一个保持此类证据显式性的全局-局部融合框架:来自独立训练分类器的平均整体严重程度概率,与来自目标检测器的结构化病灶负荷描述子(病灶数量、检测置信度、病灶面积)相结合,形成一个紧凑的表示,再由一个轻量级、可解释的分类器给出最终等级。在一个广泛使用的基准数据集上,该融合相比仅使用整体证据的基线方法取得了明确且具有统计学支持的提升,其中在最严重病例上的收益最大。在采用不同分级标准的独立数据集上的测试表明,数据集内的良好性能并不能自动迁移,后续诊断分析将这一差距主要归因于分级标准的不一致,而非单纯的检测失败。这些发现支持可解释的全局-局部融合作为序数型痤疮分级的有效策略,同时强调标准对齐是实现跨数据集可移植性的关键;此外,我们还进一步展示了由此产生的严重程度信号如何在护肤应用中支持透明、非诊断性的决策。
cs.CV / 70 / 2609.31349

DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models

DyMD:在少步视频世界模型中通过分布匹配蒸馏保留交互动态
Xu, Haojun, Huang, Jie, Lu, Xin, Zhong, Mingchen, Fan, Zihao, Huang, Linjiang, Liu, Si
Abstract
Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress robot--object motion while preserving visual quality. Examining DMD's teacher and fake-score signals, we find that weak re-noising keeps the teacher posterior concentrated near motion-deficient rollouts, limiting motion-restoring guidance. Meanwhile, stronger-motion rollouts tend to incur larger fake-score fitting errors, which can hinder the generator's learning of interaction dynamics. We propose DyMD, a DMD framework that adapts both teacher supervision and critic fitting to the evolving student. Temporal affinity--conditioned re-noise sampling adapts the timestep distribution to each rollout's current interaction fidelity by mixing the base schedule with a teacher prior motivated by local posterior variation, thereby balancing motion recovery and appearance refinement. To better track stronger-motion rollouts, dynamics-guided fake-score tracking uses a noise-conditioned predictor to estimate noise-relative fitting difficulty from latent temporal dynamics, then upweights predicted-hard rollouts in the critic loss. Using DyMD, we distill a 14B teacher into a four-step 1.3B student with no auxiliary modules at inference. On embodied-video benchmarks, the student improves R-Bench task adherence by $9.6$ percentage points and PAI-Bench-G Domain score by $5.1$ points over Base DMD while maintaining comparable visual quality. As a backbone for downstream action planning, our student achieves 34% mean success across two WorldArena tasks, compared with 16% for Base DMD.
Chinese Translation
大型视频扩散模型为具身预测与学习提供了富有表现力的先验,但其多步采样对于交互式下游应用而言成本高昂。分布匹配蒸馏(Distribution Matching Distillation, DMD)能够实现少步视频生成,但在保持视觉质量的同时可能会抑制机器人—物体的运动。通过分析DMD的教师信号与伪分数(fake-score)信号,我们发现较弱的再加噪过程使教师后验集中于缺乏运动的生成结果附近,从而限制了运动恢复的引导能力。与此同时,运动较强的生成结果往往会产生较大的伪分数拟合误差,这可能阻碍生成器对交互动态的学习。我们提出DyMD,一种使教师监督与评判器拟合均能随学生演化而自适应调整的DMD框架。时间亲和度条件化的再加噪采样通过将基础调度与由局部后验变化推导出的教师先验相混合,根据每个生成结果当前的交互保真度自适应调整时间步分布,从而在运动恢复与外观精修之间取得平衡。为更好地跟踪运动较强的生成结果,动态引导的伪分数跟踪使用噪声条件化的预测器从隐层时间动态中估计相对于噪声的拟合难度,然后在评判器损失中对预测为困难的生成结果赋予更高权重。利用DyMD,我们将一个14B的教师模型蒸馏为四步推理且无需辅助模块的1.3B学生模型。在具身视频基准上,与基础DMD相比,学生模型将R-Bench任务遵循度提升9.6个百分点,PAI-Bench-G领域得分提升5.1分,同时保持相当的视觉质量。作为下游动作规划的骨干,我们的学生模型在两个WorldArena任务上取得34%的平均成功率,而基础DMD仅为16%。
cs.CV / 71 / 2609.31356

Open Vocabulary Domain Unlearning

开放词汇域遗忘
Udupa, Sumanth, Harandi, Mehrtash, Luo, Yadan, Baktashmotlagh, Mahsa
Abstract
Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as idealized textbook diagrams in medical AI or cartoon vehicles in autonomous driving. Approximate Domain Unlearning (ADU) aims to selectively erase a model's recognition of a target visual domain while preserving accuracy on the remaining domains. However, existing ADU methods operate under a flawed closed-vocabulary assumption: they evaluate unlearning solely on the specific object classes seen during the unlearning fine-tuning phase. Consequently, these methods do not unlearn the domain itself; they merely overfit to seen class-domain pairs, leaving the domain easily recognizable for unseen classes and providing a false sense of removal. We argue that true domain erasure must be class-agnostic. To address this, we formalize Open-Vocabulary Domain Unlearning (OVDU), a rigorous protocol that mandates domain forgetting must transfer to held-out classes. To solve the OVDU challenge, we propose a surgical parameter-editing framework. First, a Fisher Information mask isolates domain-sensitive weights, mathematically protecting foundational zero-shot generalization. Second, our Targeted Manifold Scattering (TMS) objective uses preference-based mining to locally scatter the forget domain's stylistic geometry. Evaluated across PACS, OfficeHome, and DomainNet, our method vastly improves open-vocabulary generalization over existing baselines. Crucially, it delivers exceptional sample efficiency, outperforming peak 8-shot baseline results with only 4 shots.
Chinese Translation
视觉-语言模型(Vision-Language Models, VLMs)展现出卓越的零样本泛化能力,然而它们往往会编码不需要的或有害的风格域,例如医学人工智能中理想化的教科书示意图,或自动驾驶中的卡通车辆。近似域遗忘(Approximate Domain Unlearning, ADU)旨在选择性地消除模型对目标视觉域的识别能力,同时保持其在其余域上的准确率。然而,现有的ADU方法基于一个有缺陷的封闭词汇假设:它们仅在遗忘微调阶段所见过的特定物体类别上评估遗忘效果。因此,这些方法并没有遗忘域本身,而只是过拟合于已见过的类别-域对,导致域对于未见过类别仍然易于识别,并带来一种被移除的假象。我们认为,真正的域擦除必须是类别无关的。为此,我们形式化了开放词汇域遗忘(Open-Vocabulary Domain Unlearning, OVDU),这是一个严格的协议,要求域遗忘必须能够迁移到留出的类别上。为解决OVDU挑战,我们提出了一种精确的参数编辑框架。首先,Fisher信息掩码分离出域敏感的权重,从数学上保护了基础的零样本泛化能力。其次,我们的目标流形散射(Targeted Manifold Scattering, TMS)目标利用基于偏好的挖掘,局部散射遗忘域的风格几何结构。在PACS、OfficeHome和DomainNet上的评估表明,我们的方法在开放词汇泛化方面大幅超越现有基线。至关重要的是,该方法具有卓越的样本效率,仅使用4个样本即可超越基线在8个样本下的峰值性能。
cs.CV / 72 / 2609.31364

OpenVAM: Open-World Visual Attention Modeling with VLMs

OpenVAM:基于视觉语言模型的开放世界视觉注意力建模
Hooshanfar, Kiana, Kazerouni, Amirhossein, Hosseini, Alireza, Brudno, Michael, Taati, Babak
Abstract
Predicting human gaze is a core capability for applications ranging from web/UI design analysis to robotics and human-computer interaction. Yet, most visual attention modeling methods output only a dense saliency map, which is often insufficient for action: practitioners need to connect attention peaks to discrete elements in the scene (what) and understand the drivers of those peaks in context (why), while remaining robust to domain shift across natural images, commercial content, and UI/web layouts. We, therefore, introduce OpenVAM (Open-world Visual Attention Modeling with VLMs), a unified framework that jointly addresses universality and explainability across heterogeneous domains (natural scenes, commercial imagery, and UI/web layouts) and supervision modalities. OpenVAM adopts a decoupled-but-aligned design: a dedicated dense visual pathway provides stable, spatially precise localization, while an instruction-following vision--language semantic head generates grounded what/why explanations conditioned on the same image and data-type prompts. A three-stage training strategy preserves strong localization priors while progressively introducing language grounding and improving explanation alignment via parameter-efficient adaptation without perturbing the saliency branch. We further propose a scalable pipeline to generate multi-domain saliency-reason annotations for training and systematic evaluation. Experiments across diverse datasets show that OpenVAM improves robustness under domain shift while producing image-grounded explanations that make saliency predictions more interpretable.
Chinese Translation
预测人类注视点是众多应用的核心能力,涵盖网页/UI设计分析、机器人技术以及人机交互等领域。然而,大多数视觉注意力建模方法仅输出稠密的显著性图,这往往不足以支持后续行动:实际应用者需要将注意力峰值与场景中的离散元素建立关联(什么),并在上下文中理解这些峰值的驱动因素(为什么),同时还要对自然图像、商业内容和UI/网页布局等不同领域间的域偏移保持鲁棒性。因此,我们提出了OpenVAM(基于视觉语言模型的开放世界视觉注意力建模),这是一个统一框架,能够同时解决异构领域(自然场景、商业图像和UI/网页布局)及不同监督模态下的通用性与可解释性问题。OpenVAM采用解耦但对齐的设计:一条专用的稠密视觉通路提供稳定且空间精确的定位,而一个遵循指令的视觉—语言语义头基于相同的图像和数据类型提示,生成有据可依的“什么/为什么”解释。我们采用三阶段训练策略,在保留强定位先验的同时,通过参数高效的自适应逐步引入语言落地能力并提升解释的对齐性,而不干扰显著性分支。此外,我们提出了一个可扩展的流水线,用于生成多领域的显著性—推理标注,以支持训练和系统性评估。在多个数据集上的实验表明,OpenVAM在域偏移下提升了鲁棒性,同时生成基于图像的解释,使显著性预测更具可解释性。
cs.CV / 73 / 2609.31378

ContraFM-S2O: Flow Matching-Based One-step SAR-to-Optical Image Translation Model with Contrastive Learning

ContraFM-S2O:基于流匹配与对比学习的单步SAR到光学图像翻译模型
Yu, Mingqian, Chiang, Wei-kuan, Wang, Qiurui, Zhao, Peilin
Abstract
In recent years, diffusion models and GAN-based models have become the mainstream approaches for SAR-to-optical image translation, owing to their advantages, such as high-quality generation and stable training. However, they have shortcomings such as high inference latency and the generated optical images suffer from low detail fidelity, often resulting in blurred edges and loss of fine textures. Thus, we propose ContraFM-S2O, which is a flow matching-based model for SAR-to-optical image translation. Unlike conventional diffusion models, ContraFM-S2O learns to predict the velocity field in training and solves ODE instead of SDE during inference to improve the sampling efficiency. In addition, ContraFM-S2O replaces instantaneous velocity with average velocity along the interpolation path to realize one-step SAR-to-optical image translation and uses contrastive learning to improve the quality of the generated optical images. Experiments show our model achieves state-of-the-art on SAR2Opt and QXS datasets, outperforming baselines, and reduces inference latency via one-step generation.
Chinese Translation
近年来,扩散模型和基于生成对抗网络(GAN)的模型凭借高质量生成和训练稳定等优势,已成为SAR到光学图像翻译的主流方法。然而,它们存在推理延迟高的缺陷,且生成的光学图像细节保真度较低,常出现边缘模糊和精细纹理丢失等问题。为此,我们提出了ContraFM-S2O,一种基于流匹配(flow matching)的SAR到光学图像翻译模型。与传统扩散模型不同,ContraFM-S2O在训练中学习预测速度场,并在推理时求解常微分方程(ODE)而非随机微分方程(SDE),从而提高采样效率。此外,ContraFM-S2O用沿插值路径的平均速度替代瞬时速度,实现单步SAR到光学图像翻译,并利用对比学习提升生成光学图像的质量。实验表明,我们的模型在SAR2Opt和QXS数据集上取得了最先进的性能,优于各基线方法,并通过单步生成降低了推理延迟。
cs.CV / 74 / 2609.31431

AxonSynth: Domain-Randomized Synthetic Data for Zero-Shot 3D Axon Segmentation in Light-Sheet Microscopy

AxonSynth:基于域随机化合成数据的光片显微镜零样本3D轴突分割
Gaibor, Edward, Bintsi, Kyriaki-Margarita, Ureta, Carmen Luz Leiva, Bellatif, Zayneb, Maffei, Chiara, Li, Wenze, Hillman, Elizabeth, Balbastre, Yaël, Yendiki, Anastasia
Abstract
Accurate segmentation of axons in 3D microscopy data is important for analyzing white-matter organization, but dense ground truth labels are expensive to obtain. Existing supervised axon segmentation methods rely on target-domain annotations and can be brittle when tissue type, species, modality, or acquisition conditions change. We present AxonSynth, a domain-randomized synthetic-data framework for training 3D axon segmentation models without manually annotated real training volumes. AxonSynth generates dense synthetic axon labels with orientation priors that reflect realistic fiber configurations and renders them with randomized density, contrast, bias fields, blur, and noise. A three-class 3D U-Net is trained to predict background, axon sheath and intra-axonal space. We evaluate zero-shot transfer on 10 held-out light-sheet microscopy (LSM) patches from macaque and human brain samples labeled with one of three axonal markers, comparing against calibrated thresholding and Frangi filtering using overlap, corrected detection, false-positive, and topology metrics. On macaque samples, AxonSynth achieved the best corrected Dice and corrected precision (0.826 and 0.851), compared with 0.765 and 0.754 for thresholding and 0.685 and 0.762 for Frangi. On human samples, corrected Dice was comparable to thresholding (0.857 vs. 0.868), while component-count error decreased from 22,504 to 3,377. Across all held-out patches, AxonSynth reduced component-count error in 10/10 patches and Euler-characteristic error in 8/10. These results show that synthetic-label domain randomization can reduce dependence on manual axon annotation while supporting synthetic-to-real 3D segmentation.
Chinese Translation
在3D显微数据中对轴突进行准确分割对于分析白质的组织结构十分重要,但获取密集的真实标注(ground truth)成本高昂。现有的有监督轴突分割方法依赖于目标域的标注数据,当组织类型、物种、成像模态或采集条件发生变化时,其性能可能变得脆弱。我们提出了AxonSynth,这是一个基于域随机化的合成数据框架,可在无需人工标注真实训练体积的情况下训练3D轴突分割模型。AxonSynth生成带有方向先验的密集合成轴突标签,以反映真实纤维构型,并以随机化的密度、对比度、偏置场、模糊和噪声对其进行渲染。我们训练了一个三分类3D U-Net来预测背景、轴突髓鞘和轴突内空间。我们在来自猕猴和人脑样本、并使用三种轴突标记物之一进行标记的10个留出光片显微镜(LSM)图像块上评估了零样本迁移性能,并与校准阈值分割和Frangi滤波进行比较,评价指标包括重叠度、校正检出率、假阳性率以及拓扑指标。在猕猴样本上,AxonSynth取得了最佳的校正Dice和校正精确率(分别为0.826和0.851),而阈值分割分别为0.765和0.754,Frangi滤波分别为0.685和0.762。在人脑样本上,校正Dice与阈值分割相当(0.857对比0.868),而连通分量计数误差从22,504降低至3,377。在所有留出图像块上,AxonSynth在10/10个图像块中降低了分量计数误差,并在8/10个图像块中降低了欧拉示性数误差。这些结果表明,基于合成标签的域随机化能够减少对人工轴突标注的依赖,同时支持从合成到真实的3D分割。
cs.CV / 75 / 2609.31435

Implicit Neural Representation for Hyperspectral Video Compression

基于隐式神经表示的高光谱视频压缩
Scalera, Alfredo, Murray, Paul, Zabalza, Jaime
Abstract
With the advent of snapshot cameras, hyperspectral video is becoming more readily available. In recent years, new applications have emerged which have led to increasingly larger datasets. However, hyperspectral video compression remains in the early stages. In this study, we explore the use of implicit neural representation as a candidate solution. We propose a novel extension of an existing RGB video compression model, achieving Bj{\o}ntegaard Delta PSNR gains of +4.99 dB and Bj{\o}ntegaard Delta rate of -88.88% compared to traditional hyperspectral image compression methods applied frame-by-frame. In addition to reconstruction quality, the effects on downstream task performance are measured in the form of object tracking success. Compared to video compressed with methods based on principal component analysis and JPEG2000 in low data regimes, our proposed method improves tracking area under the curve by up to 23.42% and distance precision by up to 35.56% on examples from the HOT2026 dataset.
Chinese Translation
随着快照式相机的出现,高光谱视频正变得越来越易于获取。近年来,新应用不断涌现,导致数据集规模日益增大。然而,高光谱视频压缩仍处于早期阶段。在本研究中,我们探索将隐式神经表示(Implicit Neural Representation)作为候选解决方案。我们提出了一种对现有RGB视频压缩模型的新颖扩展,相比逐帧应用传统高光谱图像压缩方法,实现了+4.99 dB的Bjøntegaard Delta PSNR增益和-88.88%的Bjøntegaard Delta码率降低。除重建质量外,我们还以目标跟踪成功率的形式衡量了对下游任务性能的影响。与基于主成分分析(PCA)和JPEG2000的低数据量压缩方法相比,在HOT2026数据集的样例上,我们提出的方法将跟踪曲线下面积(AUC)提升了最高23.42%,距离精度提升了最高35.56%。
cs.CV / 76 / 2609.31450

From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training

从奖励信号到视觉效用:医学视觉语言模型后训练的受控审计
Jingxin, Wang
Abstract
Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy. We examine how changes in accuracy and training objectives relate to image-conditioned decisions in a controlled Qwen2.5-VL-3B study on PMC-VQA. We compare supervised fine-tuning (SFT) with low-rank adaptation (LoRA) restricted to the language model, expanded multimodal adaptation scopes, standard answer-only Group Relative Policy Optimization (GRPO), and a counterfactual evidence objective. On 2,000 clean-test questions, language model LoRA SFT changes correct-image accuracy by +1.10 percentage points (95% paired bootstrap CI:-0.85 to +3.05), while visual-benefit events decrease by 2.40 points and image sensitivity decreases by 5.60 points. Paired records reveal 155 acquired and 203 lost visual-benefit events. Broader adaptation yields lower correct-image accuracy than language-model LoRA SFT. Standard GRPO produces mixed-reward groups and parameter updates, with an uncertain clean test accuracy change. A generation audit reveals that canonical option scores can follow a different token path from generated answers. With scores taken along the greedy generation path, the evidence target improves on the training set; its gains over standard GRPO remain inconsistent on validation data at matched training doses. Sample-level analyses trace how evidence scores, decision margins, and generated answers change during post-training. This empirical and measurement audit identifies gaps between optimization activity, target acquisition, and useful held-out visual behavior.
Chinese Translation
医学视觉语言模型(VLM)的后训练通常通过答案准确率进行评估。我们在PMC-VQA数据集上对Qwen2.5-VL-3B开展了一项受控研究,考察准确率的变化和训练目标与基于图像条件决策之间的关系。我们比较了限于语言模型的低秩适配(LoRA)监督微调(SFT)、扩展的多模态适配范围、标准的仅答案Group Relative Policy Optimization(GRPO),以及反事实证据目标。在2,000道清洁测试题上,语言模型LoRA SFT使正确图像准确率变化+1.10个百分点(95%配对自助置信区间为-0.85至+3.05),而视觉受益事件减少2.40个百分点,图像敏感度下降5.60个百分点。配对记录显示获得155个视觉受益事件,同时失去203个。更广泛的适配所取得的正确图像准确率低于语言模型LoRA SFT。标准GRPO产生混合奖励组和参数更新,但其在清洁测试上的准确率变化不确定。生成审计发现,规范选项得分可能遵循与生成答案不同的token路径。当得分沿贪婪生成路径计算时,证据目标在训练集上有所改进;但在相同训练剂量下,其在验证数据上相对标准GRPO的增益并不一致。样本级分析追踪了后训练过程中证据得分、决策边际和生成答案的变化。这项实证与测量审计揭示了优化活动、目标获取与有用的留出视觉行为之间的差距。
cs.CV / 77 / 2609.31456

Diagnosing the Sources of Compositional Failure in Vision-Language Models: A Controlled Analysis

诊断视觉-语言模型组合推理失败的来源:一项受控分析
Gandhi, Mona, Olcay, Cenk Merih, Lo, Kuan-Chieh, Castro, Santiago, Myers, Christopher W., Parthasarathy, Srinivasan
Abstract
Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hypothesis is that models struggle to integrate multiple components, leading to training interventions to improve compositional binding. However, this assumption has never been directly quantified. Existing benchmarks evaluate captions only in their composed form, making it impossible to separate the cost of joint reasoning from the cost of recognizing individual components under increasing load. We introduce COMPASS (COMPositional Analysis of SkillS), a controlled evaluation framework designed to isolate and measure the distinct factors underlying compositional failure. By comparing performance on composed captions with their decomposed counterparts , we directly quantify the cost of compositional integration across 87K image-caption pairs. Across multiple VLMs, this gap is real but partial, accounting for only part of the observed degradation. This motivates a finer-grained investigation into what additional factors govern model behavior. We analyze performance at the level of individual skills: object detection, attribute binding, and relation reasoning, using skill-targeted perturbations across 274K image-caption pairs. We find a consistent skill-specific pattern: each skill degrades primarily with the count of its own primitive type (self-load), while cross-load effects are predominantly positive, suggesting that primitives of different types provide useful grounding context. This pattern holds across standard contrastive encoders, explicitly trained compositional reasoning models, and non-contrastive architectures. These findings show that compositional degradation reflects multiple separable factors that cannot be reduced to joint reasoning alone.
Chinese Translation
视觉-语言模型(VLMs)在组合推理任务中常常表现不佳,但这种性能欠佳的原因仍不清楚。一个常见的假设是模型难以整合多个组件,因此人们提出了旨在改进组合绑定的训练干预方法。然而,这一假设从未被直接量化。现有的基准仅在组合后的完整描述上评估模型,使得无法将联合推理的代价与在负载增加时识别单个组件的代价分离开来。我们提出了COMPASS(COMPositional Analysis of SkillS,技能组合分析),这是一个受控评估框架,旨在分离并测量组合失败背后的不同因素。通过比较模型在组合描述与其分解形式上的表现,我们在87K图像-描述对上直接量化了组合整合的代价。在多个VLM上,这一差距确实存在但只是部分的,仅能解释所观察到的性能下降中的一部分。这促使我们对影响模型行为的其他因素进行更细粒度的研究。我们在单个技能层面分析性能:目标检测、属性绑定和关系推理,并在274K图像-描述对上使用针对技能的扰动。我们发现了一个一致的、特定于技能的模式:每种技能的性能主要随其自身基元类型数量的增加而下降(自身负载),而跨类型负载的影响 predominantly 为正向,表明不同类型的基元提供了有用的基础语境信息。这一模式在标准对比编码器、显式训练的组合推理模型以及非对比架构上均成立。这些发现表明,组合性能下降反映了多个可分离的因素,不能仅仅归因于联合推理。
cs.CV / 78 / 2609.31461

KneePreM: Towards 3D Knee MRI Foundation Models via Large-Scale Unlabeled Pretraining and Label-Efficient Fine-Tuning

KneePreM:通过大规模无标注预训练与高标签效率微调构建3D膝关节MRI基础模型
Wang, Xinxin, Hazan, Liam, Li, Jing, Rabinovici-Cohen, Simona, Li, Xiaojuan, Yang, Mingrui
Abstract
Background: Large volumes of unlabeled knee MRI scans are available across repositories but remain insufficiently leveraged. We developed KneePreM, a knee-specific 3D self-supervised model, and evaluated transfer and label efficiency for classification and segmentation. Methods: A 3D U-Net masked autoencoder was pretrained on 19,011 unlabeled Osteoarthritis Initiative (OAI) MRI series from 4,791 participants. Downstream fine-tuning used full and reduced training sets for fastMRI+ two-label classification (1,172 examinations), Arthroscopic Partial Meniscectomy (APM) eight-target classification (1,716 examinations), SKM-TEA segmentation (155 examinations), and APM segmentation (25 examinations). Baselines were random initialization and SuPreM. Deployment workflow was implemented with a Model Context Protocol interface. Evaluation metrics included balanced accuracy, F1 score, ROC AUC, PR AUC, and Dice score. Statistical analysis used bootstrap confidence intervals and paired bootstrap tests for classification and Wilcoxon signed-rank tests for segmentation. Results: KneePreM achieved higher full-data macro ROC AUC than both baselines for fastMRI+ and APM (all p < .001). For fastMRI+ classification, KneePreM achieved a ROC AUC of 0.722 using 50% of the training data, exceeding both full-data baselines. In APM classification, KneePreM reached a ROC AUC of 0.740 with 70% of the data, matching the full-data random baseline and outperforming SuPreM. For SKM-TEA segmentation, its 70%-data Dice of 0.838 exceeded the full-data random baseline (0.835) and both same-budget comparators. In APM segmentation, its 75%-data Dice of 0.746 exceeded the full-data random baseline (0.731) and both same-budget comparators. Conclusion: KneePreM improves transfer performance and label efficiency across knee MRI classification and segmentation tasks, particularly when labeled training data are limited.
Chinese Translation
背景:各大数据库中存在大量未标注的膝关节MRI扫描数据,但尚未得到充分利用。我们开发了KneePreM,一种针对膝关节的3D自监督模型,并评估了其在分类与分割任务中的迁移能力和标签效率。方法:基于来自4,791名参与者的19,011个未标注骨关节炎倡议研究(Osteoarthritis Initiative, OAI)MRI序列,对3D U-Net掩码自编码器进行预训练。下游微调使用完整及缩减规模的训练集,分别用于fastMRI+双标签分类(1,172次检查)、关节镜部分半月板切除术(APM)八目标分类(1,716次检查)、SKM-TEA分割(155次检查)以及APM分割(25次检查)。基线方法为随机初始化和SuPreM。部署流程通过模型上下文协议(Model Context Protocol)接口实现。评估指标包括平衡准确率、F1分数、ROC AUC、PR AUC和Dice分数。统计分析采用bootstrap置信区间和配对bootstrap检验(分类任务)以及Wilcoxon符号秩检验(分割任务)。结果:在fastMRI+和APM任务中,KneePreM在全量数据下的宏平均ROC AUC均高于两个基线方法(所有p < .001)。在fastMRI+分类任务中,KneePreM仅使用50%训练数据即达到0.722的ROC AUC,超过了两个使用全量数据的基线。在APM分类任务中,KneePreM使用70%数据达到0.740的ROC AUC,与全量数据的随机初始化基线持平,并优于SuPreM。在SKM-TEA分割任务中,其使用70%数据的Dice分数为0.838,超过了全量数据的随机初始化基线(0.835)以及两个同等数据预算的对比方法。在APM分割任务中,其使用75%数据的Dice分数为0.746,超过了全量数据的随机初始化基线(0.731)以及两个同等数据预算的对比方法。结论:KneePreM提升了膝关节MRI分类与分割任务的迁移性能和标签效率,尤其是在标注训练数据有限的情况下。
cs.CV / 79 / 2609.31507

SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery

SatNav:基于卫星影像的可扩展长时域无人机视觉-语言导航基准
Jiang, Jiajun, Hua, Chunliang, Chen, Zichun, Wu, Yanxing, Yang, Zeyuan, Song, Jie, Hu, Xiao
Abstract
Urban uncrewed aerial vehicle (UAV) vision-language navigation (VLN) requires agents to follow instructions across extended urban spaces, inherently demanding long-term memory and geospatial grounding. However, scaling existing benchmarks remains difficult because of their reliance on costly reconstructed 3D assets, limiting geographic diversity and episode scale. To address this, we introduce SatNav, a scalable, long-horizon UAV VLN benchmark built from high-resolution satellite imagery. SatNav targets city-level navigation missions and uses satellite crops as approximations of UAV nadir views for visual observations. Through an automated cue-to-episode pipeline, SatNav constructs 118K episodes from 59 scenes across 18 cities, with an average trajectory length of 379 m. To stress-test long-horizon memory and geospatial reasoning, SatNav defines three task families: Boundary, Landmark, and Route, targeting loop progress tracking, landmark-based spatial grounding, and route following with counting cues. Benchmarking classical VLN agents and recent agents based on large vision-language models (LVLMs) on SatNav shows that city-scale navigation remains challenging. We further introduce SwiftVLN, a modular framework with switchable memory components, and conduct systematic memory-design ablations. Finally, satellite-to-UAV transfer experiments show that satellite-trained navigation models can operate on real-flight UAV observations, showing the practical relevance of SatNav. Our project page: https://eku127.github.io/SatNav/
Chinese Translation
城市无人机视觉-语言导航(VLN)要求智能体在城市空间中按照指令进行长距离导航,这本质上需要长期记忆和地理空间定位能力。然而,现有基准依赖于代价高昂的重建三维资产,难以扩展,限制了地理多样性和任务规模。为解决这一问题,我们提出了SatNav,一个基于高分辨率卫星影像构建的可扩展长时域无人机VLN基准。SatNav面向城市级导航任务,并使用卫星影像裁剪图作为无人机正射视角(nadir view)观测的近似。通过自动化的“线索-任务”生成流水线,SatNav从18个城市的59个场景中构建了118K个任务,平均轨迹长度为379米。为对长时域记忆和地理空间推理进行压力测试,SatNav定义了三组任务类型:Boundary(边界)、Landmark(地标)和Route(路径),分别针对循环进度跟踪、基于地标的空间定位以及结合计数线索的路径跟随。我们在SatNav上对经典VLN智能体和基于大型视觉-语言模型(LVLM)的近期智能体进行了基准测试,结果表明城市尺度导航仍然具有挑战性。我们进一步提出了SwiftVLN,一个具有可切换记忆组件的模块化框架,并进行了系统的记忆设计消融实验。最后,卫星影像到无人机影像的迁移实验表明,在卫星影像上训练的导航模型可以在真实飞行无人机观测数据上运行,验证了SatNav的实用价值。项目页面:https://eku127.github.io/SatNav/
cs.CV / 80 / 2609.31509

ClearGS: Reliability-Aware Gaussian Splatting from Handheld Videos

ClearGS:基于手持视频的可靠性感知高斯泼溅
Liu, Xuanzhi, Wu, Xinyi, Pan, Hang, Huang, Wensi, Wu, Zhenyao, Han, Ruize, Wang, Song
Abstract
We present ClearGS for 3D Gaussian Splatting (3DGS) from handheld videos with uneven viewpoint coverage and mixed frame quality. Rather than selecting frames with binary decisions, ClearGS uses Reliability-aware View Allocation (RVA) to assign graded raw-supervision weights based on appearance reliability, degradation risk, and geometric utility, while weakly reactivating useful suppressed frames to maintain trajectory coverage. Since weighting cannot restore details lost to blur or distortion, ClearGS further introduces Render-Guided In-Video Restoration (RIVR). The current 3DGS render provides a pose-aligned structural candidate, a frozen no-reference restoration expert restores the corresponding raw video observation without any clean reference image, and no-reference perceptual scores select among the render, restored observation, and high-frequency fused candidate. ClearGS then applies Full-Trajectory Repair Consolidation to revisit accepted repairs and preserve details introduced early. On GS2E and GSOTM, ClearGS achieves state-of-the-art overall performance, with consistent CLIP-IQA and MUSIQ gains and LPIPS reductions in most degradation settings, without paired sharp supervision or matched clean references.
Chinese Translation
我们提出了ClearGS,用于从不均匀视角覆盖和混合帧质量的手持视频中进行三维高斯泼溅(3D Gaussian Splatting, 3DGS)。ClearGS并非采用二值决策来选择帧,而是使用可靠性感知视图分配(Reliability-aware View Allocation, RVA),基于外观可靠性、退化风险和几何效用分配分级的原始监督权重,同时弱激活被抑制的有用帧以保持轨迹覆盖。由于加权无法恢复因模糊或失真而丢失的细节,ClearGS进一步引入了渲染引导的视频内修复(Render-Guided In-Video Restoration, RIVR)。当前3DGS渲染结果提供了位姿对齐的结构候选,一个冻结的无参考修复专家在无需任何清晰参考图像的情况下修复对应的原始视频观测,并由无参考感知评分在渲染结果、修复观测和高频融合候选之间进行选择。随后,ClearGS应用全轨迹修复巩固(Full-Trajectory Repair Consolidation),重新审视已接受的修复并保留早期引入的细节。在GS2E和GSOTM数据集上,ClearGS取得了最先进的整体性能,在大多数退化设置下均获得一致的CLIP-IQA和MUSIQ提升以及LPIPS降低,且无需成对的清晰监督或匹配的干净参考。
cs.CV / 81 / 2609.31514

Forensic Twins: Self-Supervised Residual Learning for AI-Generated Image Forensics

法证孪生网络:面向AI生成图像取证的自监督残差学习
Muñoz-Haro, Javier, Tolosana, Ruben, Vera-Rodriguez, Ruben, Morales, Aythami, Fierrez, Julian
Abstract
Detectors of AI-generated images are typically trained using samples from all Generative AI architectures they must catch, and struggle as soon as a new architecture emerges. Recent approaches have explored self-supervised pre-training as an alternative solution, yet standard frameworks work against the forensic task, e.g., their augmentations overwrite the micro-statistics of image formation. This paper introduces Forensic Twins, a Self-Supervised Residual Learning (SSRL) framework whose pretext task suppresses macroscopic content availability. Each image is mapped through a frozen, off-the-shelf forensic residual extractor, from which two spatially disjoint crops are drawn. Sharing no pixel, the two views retain minimal semantic structure to align, leaving a redundancy-reduction objective with a predominant common signal: the stationary fingerprint of the image acquisition pipeline. Additionally, Forensic Twins is trained exclusively on real images; no AI-generated image is observed at any stage. Experiments show that Forensic Twins attributes AI generator sources with 56.61% accuracy, i.e., 6.13% above the previous state-of-the-art zero-shot method at 375x lower latency. We also demonstrate that fitting a Gaussian Mixture Model (GMM) offline using only the real image embeddings extracted from Forensic Twins turns it into a state-of-the-art zero-shot detector, reaching 97.99% AUC across 27 unseen AI generators, including GANs, diffusion models and commercial systems. Code, weights and exact splits will be made publicly available
Chinese Translation
AI生成图像的检测器通常使用其需要捕捉的所有生成式AI架构的样本进行训练,一旦出现新的架构,其性能便会下降。近期的一些方法探索了自监督预训练作为替代方案,然而标准框架与取证任务相悖,例如其数据增强会覆盖图像成像的微观统计特性。本文提出法证孪生网络(Forensic Twins),这是一种自监督残差学习(SSRL)框架,其前置任务抑制了宏观内容的可用性。每张图像通过一个冻结的、现成的取证残差提取器进行映射,从中抽取两个空间上不重叠的裁剪区域。这两个视图不共享任何像素,仅保留最少的语义结构用于对齐,使得冗余消减目标的主要共同信号为图像采集流程的平稳指纹。此外,法证孪生网络仅在真实图像上训练,在任何阶段均未观测到AI生成的图像。实验表明,法证孪生网络以56.61%的准确率对AI生成器来源进行归因,比此前最先进的零样本方法高出6.13%,同时延迟降低了375倍。我们还证明,仅使用从法证孪生网络提取的真实图像嵌入离线拟合高斯混合模型(GMM),即可将其转变为最先进的零样本检测器,在27个未见过的AI生成器(包括GAN、扩散模型和商业系统)上达到97.99%的AUC。代码、权重和精确的数据划分将公开提供。
cs.CV / 82 / 2609.31524

Structured Reasoning Agentic Framework for Interpretable Critical View of Safety Assessment

用于可解释的关键安全视野评估的结构化推理智能体框架
Xu, Qing, Luo, Yuxiang, Chen, Zhen
Abstract
Surgical scene understanding is critical for computer-assisted intervention, yet laparoscopic cholecystectomy remains challenged by the complex anatomy of the hepatocystic triangle and the risk of bile duct injury. Existing methods for Critical View of Safety (CVS) assessment typically treat it as a holistic prediction task, mapping visual features directly to criterion-level labels. This black-box paradigm lacks explicit reasoning about anatomical relationships, limiting both interpretability and compositional generalization. To address this, we propose ReasonCVS, a structured reasoning agentic framework empowered by Vision-Language Models (VLMs) that decomposes CVS assessment into explicit, fine-grained anatomical verification. Specifically, we devise an Anatomical Scene Graph Abstraction (ASGA) that organizes anatomical entities and their spatial relationships into a structured representation. To operationalize this, we introduce a Rationale-Aware Reasoning Agent, powered by a Large Language Model (LLM) fine-tuned via rationale distillation. Functioning as a strict central decision-maker, it invokes VLM-driven Sub-criterion Verifier as a specialized perceptual tool to parse the graph and independently evaluate individual sub-criteria. Through calibrated soft reasoning, this agent synthesizes the tool-gathered distributed observations, yielding a final verdict alongside a traceable clinical rationale. Extensive experiments on the Endoscapes-CVS201 benchmark demonstrate that ReasonCVS achieves superior performance (68.1\% mAP) over state-of-the-art while providing interpretable, criterion-level explanations for reliable surgical assessment.
Chinese Translation
手术场景理解对计算机辅助介入至关重要,然而由于肝胆囊三角的复杂解剖结构以及胆管损伤的风险,腹腔镜胆囊切除术仍然面临挑战。现有的关键安全视野(Critical View of Safety, CVS)评估方法通常将其视为一个整体预测任务,直接将视觉特征映射到标准层面的标签。这种黑箱范式缺乏对解剖关系的显式推理,限制了可解释性和组合泛化能力。为解决这一问题,我们提出了ReasonCVS,一个由视觉-语言模型(Vision-Language Models, VLMs)赋能的结构化推理智能体框架,它将CVS评估分解为显式的、细粒度的解剖验证。具体而言,我们设计了 anatomical场景图抽象(Anatomical Scene Graph Abstraction, ASGA),将解剖实体及其空间关系组织为结构化表示。为实现该方法的落地,我们引入了一个理由感知推理智能体(Rationale-Aware Reasoning Agent),其核心是通过理由蒸馏微调的大语言模型(Large Language Model, LLM)。作为严格的核心决策者,该智能体调用由VLM驱动的子标准验证器(Sub-criterion Verifier)作为专门的感知工具,以解析场景图并独立评估各个子标准。通过校准的软推理,该智能体综合工具收集到的分布式观测结果,得出最终判定,同时提供可追溯的临床理由。在Endoscapes-CVS201基准上的大量实验表明,ReasonCVS在取得优于最先进方法性能(68.1% mAP)的同时,还能为可靠的手术评估提供可解释的、标准层面的解释。
cs.CV / 83 / 2609.31558

Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP

针对CLIP中隐蔽嵌入空间后门的区域级黑盒防御
Abdelnaby, Ahmed, Elmahallawy, Mohamed
Abstract
Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only a tiny fraction of image--text pairs, adversaries can implant stealthy triggers that induce targeted shifts in CLIP's joint embedding space. Unlike conventional backdoors that manipulate classifier logits, these attacks corrupt representations directly, making them highly effective under extremely low poisoning ratios and difficult to detect. Existing defenses require access to model parameters, gradients, logits, or clean validation data---assumptions that rarely hold in realistic black-box deployments. Moreover, current black-box methods struggle to accurately localize small or out-of-distribution triggers. We propose CLIPGuard, a lightweight and fully black-box defense specifically designed to mitigate embedding-space backdoors in CLIP encoders. CLIPGuard identifies malicious regions by measuring segment-wise embedding perturbations and selectively purifies only suspicious segments via semantic inpainting, preserving benign visual content and alignment quality. Extensive experiments on STL-10, ImageNet, and diverse trigger families---including BadCLIP, BadNets, blended, patch-based, and typographic attacks---demonstrate that CLIPGuard reduces attack success rates to as low as 1.05% while maintaining clean accuracy up to 86.34%, consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP. Our code is available https://github.com/wsu-cyber-security-lab-ai/CLIPGuard.git
Chinese Translation
对比语言-图像预训练(Contrastive Language–Image Pretraining, CLIP)凭借其强大的迁移能力和零样本能力,已成为主流的视觉主干模型。然而,近期研究揭示了一个关键漏洞:嵌入空间后门攻击。攻击者仅需污染极少量的图像-文本对,即可植入隐蔽的触发器,导致CLIP联合嵌入空间发生针对性的偏移。与操纵分类器logits的传统后门不同,这类攻击直接破坏表示,使其在极低的污染比例下依然高度有效且难以检测。现有防御方法需要访问模型参数、梯度、logits或干净的验证数据——这些假设在现实黑盒部署场景中很少成立。此外,当前的黑盒方法难以准确定位小型或分布外触发器。我们提出了CLIPGuard,一种专为缓解CLIP编码器中嵌入空间后门而设计的轻量级、完全黑盒防御方法。CLIPGuard通过测量分段级别的嵌入扰动来识别恶意区域,并通过语义修复(semantic inpainting)选择性地仅净化可疑片段,从而保留良性视觉内容和图文对齐质量。在STL-10、ImageNet以及多种触发器类型(包括BadCLIP、BadNets、混合式、基于补丁的和排版式攻击)上的大量实验表明,CLIPGuard可将攻击成功率降至低至1.05%,同时保持高达86.34%的干净准确率,持续优于现有黑盒防御方法(包括CleanCLIP和CleanerCLIP)。我们的代码已发布于 https://github.com/wsu-cyber-security-lab-ai/CLIPGuard.git
cs.CV / 84 / 2609.31572

OC-GS: Gaussian Splatting for Irregular Turntable Capture

OC-GS:面向不规则转台采集的高斯泼溅重建方法
Lee, Jae Joong, Benes, Bedrich
Abstract
Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This orbit-consistent refinement jointly optimizes image-derived geometry and angles to reconstruct objects from sparse, irregular captures. On rendered objects with 12, 8, and 6 irregularly spaced views, OC-GS achieves mean foreground PSNR scores of 21.26, 19.36, and 15.83dB, respectively, exceeding all four evaluated pose-free Gaussian splatting baselines in each condition. Under a shared trainer, refining image-estimated angles improves mean foreground PSNR by 7.88dB over keeping those estimates fixed. An ablation study shows that both image-derived angle initialization and the shared motion model contribute to the improvement. On real captures, OC-GS's refinement increases mean foreground PSNR by 0.70dB. Results show that refining uncertain angles within a shared motion model improves reconstruction from sparse, irregular turntable captures.
Chinese Translation
旋转不均匀和丢帧问题使得等角度假设在转台重建中不再可靠。我们提出OC-GS,一种以物体为中心的高斯泼溅(Gaussian Splatting)方法,在保持共享相机、旋转轴和旋转中心的同时,对每幅图像的角度进行优化。这种轨道一致性优化将图像推导的几何与角度联合优化,从而从稀疏、不规则的采集数据中重建物体。在分别具有12、8和6个不规则间隔视角的渲染物体上,OC-GS分别取得了21.26、19.36和15.83dB的平均前景PSNR,在每种条件下均超过了所有四个被评估的无位姿高斯泼溅基线方法。在共享训练器下,对图像估计的角度进行优化相比固定这些估计值,可将平均前景PSNR提升7.88dB。消融实验表明,基于图像的角度初始化和共享运动模型均对性能提升有所贡献。在真实采集数据上,OC-GS的优化使平均前景PSNR提高了0.70dB。结果表明,在共享运动模型内优化不确定的角度能够改善稀疏、不规则转台采集数据的重建效果。
cs.CV / 85 / 2609.31573

How Far Can INRs Go? Cross-Domain Parameter-efficient INR-Based Semantic Segmentation for Brain MRI

INR能走多远?面向脑部MRI的跨域参数高效基于INR的语义分割
Shang, Ziyao, Sadeghi, Pouya, Jiang, Letian, Wong, Alexander, Rambhatla, Sirisha
Abstract
Biomedical image segmentation is central to medical image analysis, but practical deployment often faces limited annotations, memory constraints, and cross-site distribution shifts. Implicit Neural Representations (INRs) have recently emerged as a lightweight alternative for semantic segmentation, achieving competitive performance with substantially fewer parameters than conventional architectures. However, the mechanisms, scaling behavior, and domain generalization abilities of INR-based segmentation remain insufficiently understood. In this work, we study these questions in the context of cross-domain brain MRI segmentation. We analyze INR-based segmentation across low-parameter regimes, comparing it with conventional pipelines in both in-domain and out-of-domain settings. Surprisingly, we find that INR-based models do not simply improve with increasing parameter budget. Their advantage is most pronounced under low-parameter and limited-augmentation settings, while U-Net-based models benefit more from larger capacity and standard augmentation. We also investigate how INRs encode semantic information in their hidden features and show that complementary segmentation-relevant structure is distributed across multiple INR layers. Building on this insight, we introduce HierINRSeg, a hierarchical INR-based architecture that aggregates multi-layer representations for improved robustness and generalization. Extensive experiments show that HierINRSeg consistently outperforms MetaSeg, a strong recent INR-based segmentation baseline, with an average improvement of 5.6 percentage points in Dice for the in-domain test set and 8.2 percentage points out-of-domain. Overall, our analysis identifies the conditions under which INR-based segmentation is most effective, providing concrete guidance for model selection and future research.
Chinese Translation
生物医学图像分割是医学图像分析的核心,但实际部署往往面临标注有限、内存受限以及跨站点分布偏移等挑战。隐式神经表示(Implicit Neural Representations, INRs)近来作为一种轻量级的语义分割替代方案兴起,其参数量远少于传统架构却仍能取得有竞争力的性能。然而,基于INR的分割机制、扩展行为以及域泛化能力仍缺乏充分理解。在本工作中,我们在跨域脑部MRI分割的背景下研究这些问题。我们分析了低参数量情形下基于INR的分割方法,并在域内和域外设置中将其与传统流程进行比较。令人意外的是,我们发现基于INR的模型并非简单地随参数预算增加而提升。其优势在低参数量和有限数据增强的设置下最为显著,而基于U-Net的模型则从更大容量和标准数据增强中获益更多。我们还研究了INR如何在隐特征中编码语义信息,并表明与分割相关的互补结构分布于多个INR层中。基于这一洞察,我们提出了HierINRSeg,这是一种层次化的基于INR的架构,通过聚合多层表示来提升鲁棒性和泛化能力。大量实验表明,HierINRSeg持续优于MetaSeg(一个近期较强的基于INR的分割基线),在域内测试集上Dice平均提升5.6个百分点,域外平均提升8.2个百分点。总体而言,我们的分析明确了基于INR的分割最为有效的条件,为模型选择和未来研究提供了具体指导。
cs.CV / 86 / 2609.31595

GraphWrit3R: End-to-End 3D Scene Graph Writing

GraphWrit3R:端到端的三维场景图生成
Milivojevic, Luka, Popovic, Nikola, Sarkar, Sayan Deb, Koch, Sebastian, Armeni, Iro, Van Gool, Luc, Paudel, Danda Pani
Abstract
3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which deviates from real-world scenarios. They depend on proprietary models, hindering open-source deployment, or incur prohibitively slow inference. We present GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script. The graph lists all objects, their semantic attributes, and the relationships between them, while avoiding all of the above mentioned limitations. The choice of multiple input modalities is purely for versatility, allowing a single set of weights to handle diverse scenarios. Point cloud inputs are encoded via Sonata and Gaussian Splat inputs via Chorus, with both modalities projected onto a shared voxel grid and fused through a novel per-voxel contrastive alignment loss before being decoded by a large language model. As a natural consequence of the LLM, GraphWrit3R also supports open-vocabulary querying. On the 3DSSG benchmark, our method achieves state-of-the-art performance on object class, predicate, and triplet recall, outperforming methods that rely on ground-truth object annotations during inference. We further provide qualitative results and analyze different input modality configurations, contrastive loss formulations, and token fusion strategies.
Chinese Translation
三维场景图(3D scene graph)通过编码物体、其语义属性以及物体之间的空间和功能关系,为复杂环境提供结构化表示。当前的三维场景图生成方法存在若干根本性局限:它们依赖于带有显式中间表示的复杂多阶段流水线,导致系统脆弱且易出现误差传播;它们假设在推理阶段可以获得真实物体标注,这与现实场景不符;它们依赖专有模型,阻碍了开源部署,或者推理速度过慢而难以实用。我们提出 GraphWrit3R,这是一种简单的端到端方法,以三维点云、高斯泼溅(Gaussian Splats)或两者的组合作为输入,直接输出以结构化 JSON 脚本形式表示的完整场景图。该图列出所有物体、其语义属性以及它们之间的关系,同时避免了上述所有局限。提供多种输入模态的选择纯粹是为了增强通用性,使单组模型权重能够应对多样化的场景。点云输入通过 Sonata 编码,高斯泼溅输入通过 Chorus 编码,两种模态被投影到共享体素网格上,并通过一种新颖的逐体素对比对齐损失进行融合,随后由大语言模型进行解码。作为使用大语言模型的自然结果,GraphWrit3R 还支持开放词汇查询。在 3DSSG 基准上,我们的方法在物体类别、谓词和三元组召回率上均取得了最先进的性能,优于在推理阶段依赖真实物体标注的方法。我们进一步提供了定性结果,并分析了不同的输入模态配置、对比损失形式以及标记融合策略。
cs.CV / 87 / 2609.31620

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

FuseReg:通过正则化的层融合缓解表征自编码器中的重建-生成差距
Du, Hongyang, Xie, Yunfei, Ye, Junjie, Yang, Jiawei, Cong, Xiaoyan, Zhang, Haodong, Huang, Yongchao, Wu, Haiyu, Li, Zongxia, Gui, Shihang, Liu, Dawei, Li, Runhao, Ni, Jingcheng, Wei, Chen, Balestriero, Randall, Wang, Yue
Abstract
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off. Shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling explicitly penalizes sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. The same regularization principle extends to diffusion training, with joint regularization of both stages reducing unguided gFID by 29% on DiT-Base. These results show that training downstream models for layer-fusion robustness narrows the reconstruction-generation gap without modifying the pretrained encoder.
Chinese Translation
表征自编码器(RAE)复用预训练视觉编码器的特征作为重建和扩散潜变量,将强大的视觉表征融入图像生成。然而,RAE 仍需决定由编码器的哪些层构成生成器与像素解码器共享的潜空间。这一选择涉及权衡:较浅的层往往能更好地保留精细的像素细节,而较深的层往往能带来更好的生成指标。因此,固定的启发式层融合将两个受益于不同信息的阶段耦合在一起。我们提出 FuseReg,用对编码器层随机子集的训练取代启发式特征选择。我们从理论上分析了其内在机制:子集采样显式地惩罚对跨层分歧的敏感性。在 ImageNet-256 上使用 DINOv3-L 时,单个 FuseReg 解码器无需重新训练即可从完整融合、稀疏融合和单层融合中进行重建,其 PSNR 高于针对固定融合专门训练的解码器。这种灵活性同样有利于生成:在 RAEv2 DiT-XL 生成器保持不变的情况下,仅替换解码器即可将无引导 gFID 降低 27%。同样的正则化原则也可扩展到扩散训练,对两个阶段进行联合正则化使 DiT-Base 上的无引导 gFID 降低了 29%。这些结果表明,训练下游模型以获得层融合鲁棒性,可以在不修改预训练编码器的情况下缩小重建与生成之间的差距。
机器学习 (Machine Learning)
123
cs.LG / 1 / 2609.30270

HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference

HybridInfer:面向端侧、边缘与云端大语言模型推理的热感知强化学习分层路由
Koul, Simran
Abstract
On-device inference with small language models keeps user data local, works offline, and incurs no per-query cost, so the on-device tier is preferred when it is adequate. It is thermally constrained, however, and I find the constraint is sharper than a slowdown: on a flagship Snapdragon device, sustained on-device generation destabilizes the GPU inference runtime, which crashes or silently wedges after a few consecutive queries. The failure lies in the current toolchain (OpenCL kernel compilation and long-prompt prefill on the mobile GPU), recurs even when the device is cool, and is worst for long generations. Multi-tier routers across on-device, edge, and cloud models can relieve this pressure, but existing routers are thermal-blind and typically evaluated in simulation or on non-mobile hardware. I present HybridInfer, a thermal-aware reinforcement-learning router for a three-tier hierarchy (on-device Llama 3.2 3B, edge Llama 3.1 8B with retrieval, cloud GPT-4o) that uses the phone's thermal headroom and a query-complexity estimate as state and selects a tier by an offline-trained Q-learning policy. Its reward trades quality against latency, cost, and a thermal penalty, plus a locality bonus crediting on-device execution. I show this bonus is a precondition for thermal-aware routing: without it the optimal policy offloads every query. On a real Android benchmark of 210 prompts, the learned router attains significantly higher quality than two hand-tuned heuristics (paired Wilcoxon, p < 0.02) at the lowest cost of any adaptive condition. Always-on-device conditions match per-query quality on servable queries but are three to six times slower and fail on long queries, so routing wins on latency, reliability, and coverage rather than quality. To my knowledge this is the first use of on-device thermal headroom to select among LLM inference tiers of differing capability on real hardware.
Chinese Translation
使用小型语言模型进行端侧推理可以将用户数据保留在本地、支持离线工作且不产生单次查询成本,因此在端侧能力足够时应优先使用端侧层级。然而,端侧推理受热约束限制,我发现这一约束不仅是速度变慢那么简单:在一款旗舰级骁龙(Snapdragon)设备上,持续的端侧生成会使 GPU 推理运行时失稳,在连续几次查询后崩溃或无声地卡死。该故障源于当前工具链(OpenCL 核函数编译与移动 GPU 上的长提示词预填充),即使设备处于冷却状态也会复现,且在长文本生成时最为严重。跨端侧、边缘与云端模型的多层级路由器可以缓解这种压力,但现有路由器对温度不敏感,且通常在仿真环境或非移动硬件上评估。我提出了 HybridInfer,一个面向三层架构(端侧 Llama 3.2 3B、带检索功能的边缘 Llama 3.1 8B、云端 GPT-4o)的热感知强化学习路由器,它以手机的热余量和查询复杂度估计作为状态,通过离线训练的 Q-learning 策略选择层级。其奖励函数在质量与延迟、成本及热惩罚之间进行权衡,并附加一项奖励端侧执行的本地性加成(locality bonus)。我证明该加成是热感知路由的前提条件:没有它,最优策略会将所有查询卸载。在一个包含 210 条提示词的真实 Android 基准测试中,学习得到的路由器以任何自适应条件中最低的成本,取得了显著优于两种手工调优启发式方法的质量(配对 Wilcoxon 检验,p < 0.02)。纯端侧条件在可服务的查询上能达到与路由方案相当的单次查询质量,但速度慢三到六倍,且在长查询上失败,因此路由方案的优势体现在延迟、可靠性和覆盖范围而非质量上。据我所知,这是首次在真实硬件上利用端侧热余量在不同能力的大语言模型推理层级之间进行选择。
cs.LG / 2 / 2609.30271

When the Preconditioning Exponent Turns Negative: Learning-Rate Coupling and Cross-Environment Generalization

当预调节指数变为负值时:学习率耦合与跨环境泛化
Zhang, Gongyue, Liu, Honghai
Abstract
Adaptive optimizers are commonly parameterized by a fixed power of the second-moment estimate. Existing partially adaptive methods study exponents between momentum-like updates and the standard Adam square root, while the interaction between this exponent and the global learning rate is less understood. We perform a controlled cross-environment study using a paired four-environment classification problem with stable sparse features, environment-dependent spurious sparse features, dense features, and high-dimensional noise. Across \NumRuns{} source-training runs covering 21 preconditioning exponents $p\in[-0.5,0.5]$ and five learning rates $\eta\in[10^{-4},10^{-2}]$, we find that the exponent maximizing cross-environment accuracy decreases almost linearly with $\log_{10}\eta$. The fitted slopes range from $-0.270$ to $-0.300$, with $R^2$ between $0.972$ and $0.996$. At $\eta=10^{-2}$, source-validation selection still prefers positive exponents in all four environments, whereas cross-environment and worst-environment criteria prefer negative exponents. Checkpoint decomposition shows that lower $p$ reduces the learned spurious-to-stable and noise-to-stable weight ratios; under reversed correlation, it also reduces the magnitude of the harmful spurious margin. Negative $p$ is therefore not a universally optimal setting. It is a high-step-size allocation regime produced by the joint action of learning rate and preconditioning. The study also exposes a model-selection conflict: source-domain validation systematically selects a different preconditioning regime from the one that maximizes robustness to environmental change. The results are a single-seed, finite-budget mechanism study rather than a broad benchmark claim.
Chinese Translation
自适应优化器通常由二阶矩估计的固定幂次进行参数化。现有的部分自适应方法研究了介于类动量更新与标准Adam平方根之间的指数,而该指数与全局学习率之间的相互作用尚不明确。我们使用一个配对的四环境分类问题进行了受控的跨环境研究,该问题包含稳定的稀疏特征、依赖环境的虚假稀疏特征、稠密特征和高维噪声。在覆盖21个预调节指数 $p\in[-0.5,0.5]$ 和五个学习率 $\eta\in[10^{-4},10^{-2}]$ 的 \NumRuns{} 次源训练中,我们发现最大化跨环境准确率的指数随 $\log_{10}\eta$ 几乎呈线性下降。拟合斜率范围为 $-0.270$ 至 $-0.300$,$R^2$ 介于 $0.972$ 和 $0.996$ 之间。在 $\eta=10^{-2}$ 时,源验证选择在所有四个环境中仍偏好正指数,而跨环境和最差环境准则则偏好负指数。检查点分解表明,较低的 $p$ 会降低学习到的虚假特征与稳定特征以及噪声与稳定特征的权重比;在相关性反转的情况下,它还会降低有害虚假间隔的幅度。因此,负 $p$ 并非普遍最优设置,它是由学习率与预调节共同作用产生的一种大步长分配机制。该研究还揭示了一种模型选择冲突:源域验证系统性地选择的预调节机制与最大化环境变化鲁棒性的机制不同。本研究是一项单种子、有限预算的机制研究,而非广泛的基准测试结论。
cs.LG / 3 / 2609.30272

ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers

ENAS:面向资源受限微控制器上TinyML的高效硬件感知神经架构搜索框架
Khan, Mohd Moin, Srivastava, Naman, Arjunan, Pandarasamy
Abstract
We present \textbf{ENAS}, a hardware-aware Neural Architecture Search (NAS) framework that combines a static feasibility check, a cell-based search space supporting standard, depthwise-separable, and bottleneck blocks with optional skip connections, and a three-stage hybrid search strategy (random $\rightarrow$ top-$K$ $\rightarrow$ mutation) with persistent cross-run caching. Unlike many existing NAS frameworks that rely on GPU acceleration, ENAS is designed to operate efficiently without requiring GPUs, making it suitable for resource-constrained development environments. We evaluate ENAS on two TinyML benchmarks, Visual Wake Words and Melanoma Cancer, across eight microcontrollers with memory footprints ranging from 20\,KB to 1\,MB SRAM and nine input image resolutions. Our experimental results show that ENAS achieves mean search-time speedups of $2.41{\times}$ and $1.70{\times}$ on the Visual Wake Words and Melanoma Cancer datasets, respectively, while maintaining competitive test accuracy compared with the recent NanoNAS framework. A measured resource analysis further shows that ENAS-selected models use substantially lower peak activation RAM, the binding constraint for microcontroller deployment at matched accuracy. Additionally, ENAS achieves $79.4\%$ test accuracy on an STM32H743-based microcontroller, outperforming the greedy CPU-only baseline by $2.6$ percentage points. We release the ENAS framework as open-source at: https://github.com/EdgeIntelligenceLab/ENAS
Chinese Translation
我们提出了ENAS,一个硬件感知的神经架构搜索(NAS)框架,它结合了静态可行性检查、支持标准块、深度可分离块和瓶颈块(含可选跳跃连接)的基于单元(cell-based)的搜索空间,以及带有跨运行持久缓存的三阶段混合搜索策略(随机搜索 → Top-K → 变异)。与许多依赖GPU加速的现有NAS框架不同,ENAS旨在无需GPU的情况下高效运行,使其适用于资源受限的开发环境。我们在两个TinyML基准数据集(Visual Wake Words和黑色素瘤癌症)上评估了ENAS,涵盖八种微控制器,其内存容量介于20 KB至1 MB SRAM之间,以及九种输入图像分辨率。实验结果表明,与近期的NanoNAS框架相比,ENAS在保持具有竞争力的测试精度的同时,在Visual Wake Words和黑色素瘤癌症数据集上分别实现了2.41倍和1.70倍的平均搜索时间加速。实测资源分析进一步表明,在相同精度下,ENAS选择的模型显著降低了峰值激活内存,而峰值激活内存正是微控制器部署的关键约束。此外,ENAS在基于STM32H743的微控制器上达到了79.4%的测试准确率,比仅使用CPU的贪婪基线高出2.6个百分点。我们将ENAS框架作为开源项目发布于:https://github.com/EdgeIntelligenceLab/ENAS
cs.LG / 4 / 2609.30273

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

离线策略评估作为设计自适应实验的决策支持工具
Alves, João Victor Ferreira, Laurentino, Eduardo Rocha, Kanno, Gustavo de Oliveira, da Rocha, Thiago Costa Rizuti
Abstract
We investigate how historical data from fixed randomized experiments (A/B tests) can be used to inform the deployment of adaptive experiments based on contextual bandits. Given data collected under a static allocation, our goal is to assess which adaptive policies, if any, would have outperformed the original design and under what conditions. To this end, we combine off-policy evaluation (OPE) with a controlled warm-start simulation. From logged A/B test data exhibiting heterogeneous treatment effects, we estimate nuisance components and use doubly robust estimators to rank a portfolio of pre-specified adaptive and non-adaptive policies. When ground truth is available, we then deploy the same offline-trained policies in a simulator that reuses the exact data-generating reward probabilities, providing a safe, ground-truth-anchored environment to study the offline-to-online transition under warm starting. Using synthetic randomized controlled trials with known heterogeneity structures and an oracle policy, our results indicate that adaptive, context-aware policies improve upon fixed allocations when meaningful heterogeneity is present, while providing little benefit in its absence. We reinforce our findings on standard open benchmarks (Hillstrom, Criteo Uplift, and LaLonde), reinterpreted through a policy-value and regret perspective. Overall, our results provide a practical methodology for deciding when adaptive experimentation is worth deploying and how to select among competing adaptive policies using existing A/B test data.
Chinese Translation
我们研究了如何利用固定随机实验(A/B测试)的历史数据来指导基于上下文老虎机(contextual bandits)的自适应实验的部署。给定在静态分配下收集的数据,我们的目标是评估哪些自适应策略(如果有的话)原本能够优于原始设计,以及在何种条件下如此。为此,我们将离线策略评估(off-policy evaluation, OPE)与受控的暖启动(warm-start)模拟相结合。从呈现异质性处理效应的A/B测试日志数据中,我们估计干扰成分,并使用双重稳健(doubly robust)估计器对一组预先指定的自适应与非自适应策略进行排序。在真值(ground truth)可得时,我们随后在一个复用完全相同的数据生成奖励概率的模拟器中部署同样的离线训练策略,从而提供一个安全的、以真值为基准的环境,以研究暖启动条件下从离线到在线的过渡。利用具有已知异质性结构和最优策略(oracle policy)的合成随机对照试验,我们的结果表明,当存在显著的异质性时,自适应的、情境感知的策略优于固定分配;而在没有异质性的情况下,其收益甚微。我们通过策略价值和遗憾(regret)的视角重新解读标准公开基准数据集(Hillstrom、Criteo Uplift 和 LaLonde),进一步印证了上述发现。总体而言,我们的研究结果提供了一套实用的方法论,用于判断自适应实验何时值得部署,以及如何利用现有的A/B测试数据在相互竞争的自适应策略中进行选择。
cs.LG / 5 / 2609.30275

Cosine Similarity Is Not Evidence: Measuring the Noise Floor of Interpretability Transfer Under Quantization

余弦相似度不是证据:量化下可解释性迁移噪声底的测量
Varshney, Pranav
Abstract
A statistic reported without the quantity needed to interpret it is not evidence. We develop that thesis for a concrete practice in AI safety. Interpretability artifacts are calibrated on full-precision weights, deployed on quantized ones, and certified as surviving the change by scale-invariant statistics (cosine similarity, correlation, AUROC) that are reported without their noise floor. For the difference-in-means direction estimator, the split-half floor is governed by one dimensionless number, $\kappa = n\rho^2/d$. The closed form $\mathbb{E}[\cos] \approx (1+4/\kappa)^{-1}$ is classical; the missing input is the class separation $\rho$, which we measure on real activations; no compression-transfer study we know of reports it. On Qwen2.5-1.5B-Instruct, $\rho = 33$--$61$ across depth, so two independent runs of the estimator agree to $0.978$--$0.994$ by sampling alone. A published cosine of $0.996$ between full-precision and quantized refusal directions therefore cannot be read as preservation without the $n$ it was computed at, which is not reported. Where $n$ is known, we judge each low-bit cosine against the split-half null measured within that quantized model, because a full-precision null assumes the low-bit estimator has the same variance. That assumption is exactly what a null exists to test. The result is plain: at INT4 the direction rotated, and the deficit exceeds the estimator's own noise. At INT8 we detect no movement, which is not an equivalence claim. We also show that a scale-invariant statistic cannot distinguish translation from attenuation of a transferred decision variable, although the two call for opposite remedies. We close with reporting recommendations that cost one forward pass. Code, data, and a one-cell reproduction are released at https://github.com/pvarshh/quantinterp
Chinese Translation
一个缺少解读所需量的统计量不能构成证据。我们针对AI安全中的一项具体实践发展了这一论点。可解释性工件(interpretability artifacts)通常在全精度权重上进行校准,部署于量化权重之上,并通过尺度不变统计量(如余弦相似度、相关系数、AUROC)来认证其在量化过程中的保留性,而这些统计量的报告往往缺少其噪声底(noise floor)。对于均值差(difference-in-means)方向估计器,折半噪声底由一个无量纲数决定:$\kappa = n\rho^2/d$。闭式解 $\mathbb{E}[\cos] \approx (1+4/\kappa)^{-1}$ 是经典结论;缺失的输入是类间分离度 $\rho$,我们在真实激活值上对其进行了测量;据我们所知,没有任何压缩迁移研究报告过该量。在 Qwen2.5-1.5B-Instruct 上,$\rho$ 在不同深度为 33--61,因此仅凭采样噪声,两次独立运行的估计器的一致性即可达到 0.978--0.994。因此,一项已发表的全精度与量化拒绝方向之间 0.996 的余弦相似度,若缺少计算时所用的样本量 $n$(原文未报告),便不能被解读为方向得以保留。在 $n$ 已知的情况下,我们将在该量化模型内部测得的折半零假设作为基准来评判每个低比特余弦值,因为全精度零假设假定低比特估计器具有相同的方差——而这恰恰是零假设本应检验的内容。结果显而易见:在 INT4 下方向发生了旋转,其偏差超出了估计器自身的噪声水平;在 INT8 下我们未检测到移动,但这并非等价性声明。我们还表明,尺度不变统计量无法区分迁移决策变量的平移与衰减,而这两种情形需要截然相反的补救措施。最后,我们给出了仅需一次前向传播成本的报告建议。代码、数据以及单单元复现脚本已发布于 https://github.com/pvarshh/quantinterp
cs.LG / 6 / 2609.30276

Why Clipping Matters in AdaGrad? Toward a High-Probability Theory under Generalized Smoothness

AdaGrad中梯度裁剪为何重要:迈向广义平滑性下的高概率理论
Mazumder, Alokendu, Mohd, Ayaan, Rawat, Harshit, Roy, Arnab, Baranwal, Mayank, Rathore, Punit
Abstract
We analyze the original same-step coordinate-wise AdaGrad under generalized smoothness and heavy-tailed noise with bounded variance. In this setting, local curvature may grow sub-quadratically with the gradient norm, and stochastic gradients are assumed to have only bounded conditional second moments. We show that unclipped AdaGrad can become \emph{anisotropically miscalibrated}: under heavy-tailed noise, the adaptive denominator can learn the geometry of rare noise shocks rather than the local curvature of the objective, leading to a persistent directional distortion that blocks finite-horizon Euclidean progress. We then prove that clipping repairs this failure mode. Our main result is a finite-horizon high-probability guarantee for the original non-lagged AdaGrad update, yielding $\frac1T\sum_{t=0}^{T-1}\|\nabla f(x_t)\|^2=\mathcal{O}\left(\frac{d\big(\sqrt{\log T} + \log \frac{1}{\delta}\big)}{\sqrt{T}}\right),$ and hence $\widetilde{\mathcal O}(\varepsilon^{-2})$ complexity. This shows that, for AdaGrad under heavy-tailed noise, clipping is a structural stabilizer of the adaptive geometry rather than merely a robustness heuristic.
Chinese Translation
我们在广义平滑性和方差有界的重尾噪声条件下,分析了原始的同步逐坐标AdaGrad算法。在该设定下,局部曲率可能随梯度范数呈亚二次增长,且随机梯度仅被假设具有有界的条件二阶矩。我们证明了未加裁剪的AdaGrad可能出现“各向异性失准”:在重尾噪声下,自适应分母可能学习到罕见噪声冲击的几何特性,而非目标函数的局部曲率,从而导致一种持续的方向性畸变,阻碍有限时间范围内的欧氏进展。随后我们证明裁剪可以修复这一失效模式。我们的主要结果是为原始的非滞后AdaGrad更新给出了有限时间范围内的高概率保证,即 $\frac{1}{T}\sum_{t=0}^{T-1}\| abla f(x_t)\|^2=\mathcal{O}\left(\frac{dig(\sqrt{\log T} + \log \frac{1}{\delta}ig)}{\sqrt{T}} ight)$,从而得到 $\widetilde{\mathcal{O}}(\varepsilon^{-2})$ 的复杂度。这表明,对于重尾噪声下的AdaGrad而言,裁剪是自适应几何的一种结构性稳定器,而不仅仅是一种鲁棒性启发式手段。
cs.LG / 7 / 2609.30277

Fixed Points Without Fixed Diffusion: Implicit Neural Sheaves for Convergent Test-Time Computation

无固定扩散的不动点:用于收敛测试时计算的隐式神经层束(Neural Sheaves)
Bourgerie, Rémi, Girdzijauskas, Šarūnas, Fodor, Viktoria
Abstract
Implicit Graph Neural Networks (IGNNs) define node representations as fixed points of message-passing operators, enabling effectively infinite-depth propagation, iteration-independent parameterization, and flexible test-time computation. Yet these benefits depend on the equilibrium being unique and attainable by fixed-point iteration. Existing constructions often impose constraints on recurrent updates to obtain these guarantees, limiting the transformations available at equilibrium. This raises a central question: can IGNNs gain expressiveness through richer, edge-dependent transformations while retaining the inherent strengths of their equilibrium formulation? We introduce SheafDEQ, a subhomogeneous deep-equilibrium architecture with adaptive neural-sheaf propagation. Its learned, matrix-valued sheaf restriction maps can align, mix, or reverse neighbouring representations. Under mild regularity conditions, we prove that SheafDEQ admits a unique equilibrium reached globally by fixed-point iteration from any positive initialization. Contractivity further guarantees convergence under bounded communication staleness. We evaluate SheafDEQ on distributed-inference tasks requiring repeated nonlocal aggregation and on community detection whose rewiring increasingly favours cross-community interactions. SheafDEQ improves over fixed-propagation implicit baselines on Sums, MNIST Terrain, and Coordinates, and on community detection as connectivity becomes increasingly heterophilic. Continued-iteration diagnostics show decreasing residuals and low prediction sensitivity after 100 iterations for initialization scales from $0.001$ to $10$, while delayed-update experiments show low sensitivity to bounded communication staleness.
Chinese Translation
隐式图神经网络(IGNN)将节点表示定义为消息传递算子的不动点,从而实现 effectively 无限深度的传播、与迭代无关的参数化以及灵活的测试时计算。然而,这些优势依赖于均衡点的唯一性以及可通过不动点迭代达到该均衡点。现有构造通常对循环更新施加约束以获得这些保证,从而限制了均衡时可用的变换能力。这引出了一个核心问题:IGNN 能否通过更丰富、依赖于边的变换来提升表达能力,同时保留其均衡表述的固有优势?我们提出 SheafDEQ,一种具有自适应神经层束(neural-sheaf)传播的次齐次深均衡架构。其学习到的矩阵值层束限制映射(sheaf restriction maps)可以对齐、混合或反转相邻表示。在温和的正则性条件下,我们证明 SheafDEQ 存在唯一的均衡点,并且从任意正初始化出发,通过不动点迭代均可全局收敛到该均衡点。收缩性(contractivity)进一步保证了在有界通信延迟(staleness)下的收敛性。我们在需要重复非局部聚合的分布式推理任务以及重连日益偏向跨社区交互的社区检测任务上评估了 SheafDEQ。SheafDEQ 在 Sums、MNIST Terrain 和 Coordinates 数据集上以及连通性日益异配(heterophilic)的社区检测任务上均优于固定传播的隐式基线。持续迭代诊断表明,对于从 $0.001$ 到 $10$ 的初始化尺度,经过 100 次迭代后残差递减且预测敏感性较低;延迟更新实验表明其对有界通信延迟的敏感性较低。
cs.LG / 8 / 2609.30279

Neural Ideals and Neural Codes: An Algebraic Framework for Neural Network Classification and Feature Interpretation

神经理想与神经码:用于神经网络分类与特征解释的代数框架
Yerrapati, Venkata Subbaiah, Dixit, Rahul, Shukla, Ajay Kumar
Abstract
Understanding the features captured by the hidden layers of neural networks is a fundamental challenge in machine learning, despite their widespread success across various classification problems. In this work, we propose an algebraic framework for examining neural networks that model classification problems. Certain results, such as the correspondence between the neural network and neural ideals, algorithms for computing the neural ideals, and a stabilization theorem that enables approximation of the neural ideals, are first established. As an application to the framework, we present algorithms to identify and interpret the features captured by each hidden-layer neuron. Along with these theoretical developments, the practical performance has been demonstrated on the MNIST digit dataset, and the results highlight the pivotal role of neural ideals as a mathematical and computational tool for analyzing the features captured by neural networks. Further, we develop an interactive software that builds on the presented framework to visualize the features captured by each neuron. This tool is available at https://github.com/yvs1967/neural-network-representation-explorer
Chinese Translation
理解神经网络隐藏层所捕获的特征是机器学习中的一个根本性挑战,尽管神经网络在各类分类问题上取得了广泛成功。在本工作中,我们提出了一个代数框架,用于研究对分类问题建模的神经网络。我们首先建立了若干结果,包括神经网络与神经理想(neural ideals)之间的对应关系、计算神经理想的算法,以及一个可实现神经理想逼近的稳定化定理。作为该框架的一个应用,我们提出了用于识别和解释每个隐藏层神经元所捕获特征的算法。在这些理论发展的基础上,我们在MNIST数字数据集上展示了实际的性能表现,结果凸显了神经理想作为分析神经网络所捕获特征的数学与计算工具的关键作用。此外,我们基于所提出的框架开发了一个交互式软件,用于可视化每个神经元所捕获的特征。该工具可在 https://github.com/yvs1967/neural-network-representation-explorer 获取。
cs.LG / 9 / 2609.30281

Seasonal and Quantum-inspired Models for Neutron Monitor Time Series Forecasting

用于中子监视器时间序列预测的季节性模型与量子启发模型
Bhatia, Krishna, Devendrababu, Shalini, Ganguly, Srinjoy
Abstract
We present a focused and reproducible study of multi-horizon forecasting on the Lomnicky Stit neutron monitor (LMKS) time series. Our evaluation suite covers simple seasonal baselines, modern deep sequence models, and functional and quantum-inspired architectures, including Seasonal Naive, Long Short-Term Memory (LSTM), Temporal Convolutional Network (TCN), N-BEATS, Kolmogorov-Arnold Networks (KAN), and two quantum-inspired variants, QiLSTM and QiKAN. We describe the dataset characteristics, diagnostic analysis, preprocessing pipeline, and training procedures, and report aggregate point-forecast performance using mean absolute error (MAE) and root mean squared error (RMSE) for all evaluated models. Our quick-run results indicate that the quantum-inspired KAN variant, QiKAN, achieves the lowest aggregate forecasting error among the evaluated configurations, while the simple Seasonal Naive baseline remains remarkably competitive. These results suggest that, for highly periodic scientific monitoring time series, models incorporating strong seasonal or low-dimensional functional priors can match or outperform substantially more complex sequence architectures. The findings motivate further investigation of parsimonious and decomposable function approximators for forecasting periodic scientific signals.
Chinese Translation
我们针对洛姆尼察峰(Lomnicky Stit)中子监视器(LMKS)时间序列的多步长预测开展了一项聚焦且可复现的研究。我们的评估套件涵盖简单的季节性基线模型、现代深度序列模型以及函数型和量子启发架构,包括季节性朴素模型(Seasonal Naive)、长短期记忆网络(LSTM)、时间卷积网络(TCN)、N-BEATS、Kolmogorov-Arnold 网络(KAN),以及两个量子启发变体 QiLSTM 和 QiKAN。我们描述了数据集特征、诊断分析、预处理流程和训练流程,并使用平均绝对误差(MAE)和均方根误差(RMSE)报告了所有评估模型的总体点预测性能。我们的快速运行结果表明,量子启发的 KAN 变体 QiKAN 在所有评估配置中取得了最低的总体预测误差,而简单的季节性朴素基线模型仍表现出显著的竞争力。这些结果表明,对于高度周期性的科学监测时间序列,包含强季节性先验或低维函数先验的模型能够匹敌甚至超越复杂得多的序列架构。这些发现推动了对用于周期性科学信号预测的简洁且可分解的函数逼近器的进一步研究。
cs.LG / 10 / 2609.30286

When Does Advection-Aware Graph Nowcasting Help? A Controlled Study of Distributed Solar Ramp Forecasting with a Self-Supervised Cloud-Motion Estimator

平流感知图临近预报何时有效?基于自监督云运动估计器的分布式光伏爬坡预报受控研究
Jiang, Phillip
Abstract
Short-term forecasting of cloud-induced power ramps across a network of distributed photovoltaic (PV) or irradiance sensors is a recognised pain point for grid operators. A natural idea is to make the graph neural network (GNN) advection-aware: connect each site to the sites upwind of it, with edge time-lags set by the cloud-motion vector (CMV), so that a ramp is propagated forward before it physically arrives. Using a controlled synthetic testbed with a known wind field, we show that (i) with a realistic cross-correlation CMV estimate, an explicit advection graph does not beat a plain static or learned-adjacency spatiotemporal GNN; (ii) roughly half of the benefit available from a perfect CMV comes simply from providing an accurate motion vector as an input feature, not from graph structure; and (iii) advection helps only when the advective displacement over the forecast horizon, v*H, fits inside the sensor network. Motivated by (ii), we introduce a small self-supervised cloud-motion estimator -- a position-aware encoder trained only on a multi-lag optical-flow reconstruction objective with an annealed kernel -- that recovers the true wind vector to 2-4 degrees median angular error, 2-4x better than the classical cross-correlation method across every wind regime. Freezing this estimator and feeding its vector to the forecaster closes about 60% of the oracle-CMV RMSE gap at moderate wind (8-15% RMSE reduction over no advection), with no external wind data. We also report a negative result for a spatially-coherent probabilistic head. All claims are established on a single synthetic simulator; we discuss why real-network validation is the necessary next step and outline it.
Chinese Translation
针对分布式光伏(PV)或辐照度传感器网络中云致功率爬坡的短期预报,一直是电网运营商公认的难题。一个自然的想法是使图神经网络(GNN)具备平流感知能力:将每个站点与其上风向的站点相连,并根据云运动矢量(CMV)设定边的时间滞后,从而在爬坡现象物理到达之前将其向前传播。我们利用一个具有已知风场的受控合成测试平台证明:(i) 采用现实可行的互相关CMV估计时,显式的平流图并不优于普通的静态图或学习邻接的时空GNN;(ii) 完美CMV所能带来的收益中,约有一半仅来自于将精确的运动矢量作为输入特征提供,而非图结构本身;(iii) 只有当预报时段内的平流位移 v*H 处于传感器网络范围之内时,平流信息才有帮助。受结论(ii)的启发,我们提出一个小型自监督云运动估计器——一个仅在多滞后光流重建目标(采用退火核函数)上训练的位置感知编码器——它在各种风况下均能以2-4度的中位角误差恢复真实风矢量,比经典互相关方法好2-4倍。冻结该估计器并将其输出矢量提供给预报模型,可在中等风速下弥合约60%的oracle-CMV RMSE差距(相比无平流信息的情况RMSE降低8-15%),且无需外部风场数据。我们还报告了关于空间一致概率预测头的负面结果。所有结论均建立在单一合成模拟器之上;我们讨论了为何真实网络验证是必要的下一步,并对其进行了概述。
cs.LG / 11 / 2609.30296

NeuralCert: certified computational discovery of extremal mathematical constructions

NeuralCert:极值数学构造的认证式计算发现
Roeling, Mark Patrick
Abstract
Neural networks are becoming popular in solving mathematical problems, but stochastic models do not provide mathematical exactness by themselves. This study introduces a discovery-to-certification framework in which high-dimensional variational trial functions are learned in a compact separable representation, spectrally diagnosed and pruned, and then certified exactly through multimodular evaluation. Exact certification makes the numerical proofs fully explicit and independently verifiable. This framework can be run on a standard personal computer. Across three extremal problems, we show that neural optimization can contribute to rigorous mathematics in three distinct ways: by discovering improved constructions, by exposing empirical invariants that lead to proofs, and by revealing optimization barriers whose geometry motivates new analytic or numerical representations. More broadly, these results suggest a path toward AI-assisted mathematics in which flexible computational discovery and exact certification become complementary components of a single rigorous workflow.
Chinese Translation
神经网络在求解数学问题方面日益流行,但随机模型本身无法提供数学上的精确性。本研究提出了一个从发现到认证的框架:在高维变分试探函数的学习中采用紧凑的可分离表示,对其进行谱诊断与剪枝,然后通过多模态计算实现严格认证。精确认证使数值证明完全显式化,并可独立验证。该框架可在标准个人计算机上运行。在三个极值问题上,我们展示了神经优化能以三种不同方式为严格数学做出贡献:发现更优的构造、揭示可导向证明的经验性不变量,以及暴露优化障碍——其几何结构启发了新的解析或数值表示。更广泛地说,这些结果为AI辅助数学指明了一条路径:灵活的计算发现与精确认证成为同一严格工作流程中互补的组成部分。
cs.LG / 12 / 2609.30299

Staged Depth Training: A Representation Curriculum for PINNs

分阶段深度训练:一种面向物理信息神经网络(PINNs)的表示课程学习方法
Zhang, Kejia, Sun, Youran, Yang, Haizhao
Abstract
Representation quality is a central determinant of PINNs' performance, yet standard training leaves representations to emerge implicitly while fitting the final solution. We introduce \textbf{representation curriculum}, an ordered process in which representations are explicitly learned, transferred independently of their predictors, and progressively refined. We realize it with Staged Depth Training (SDT), which trains a shallow prefix under a temporary physics-informed head, discards the head, and freezes the learned prefix while adding depth, without equation-specific encodings or changes to the final architecture. Across the 20 default forward problems in PINNacle with three backbones, SDT improves 40 of 59 equal-budget problem--backbone cells by at least 5\% and remains within that band in the rest, with a 32.8\% geometric-mean error reduction on a PirateNet-style backbone. Mechanistic ablations suggest that the gain is not explained by optimizer restarts or shallow warm-starting alone. Representation visualizations and hyperparameter-basin analyses provide diagnostic evidence on representation geometry and local sensitivity to shared hyperparameters. On Poisson--Boltzmann 2D, SDT also more than doubles the fitted depth-scaling exponent for both backbones. These results support representation curriculum as a promising training strategy for improving PINNs while preserving the deployed architecture and inference cost.
Chinese Translation
表示质量是决定物理信息神经网络(PINNs)性能的核心因素,然而标准训练方法仅在拟合最终解的过程中让表示隐式地形成。我们提出了“表示课程学习”(representation curriculum),这是一种有序的训练过程,其中表示被显式地学习、独立于其预测器进行迁移,并逐步精细化。我们通过分阶段深度训练(Staged Depth Training, SDT)来实现这一方法:先在临时物理信息头部下训练浅层前缀网络,随后丢弃该头部,并在增加网络深度的同时冻结已学习的前缀,整个过程无需方程特定的编码,也不改变最终的网络架构。在 PINNacle 的 20 个默认正问题、三种骨干网络上的实验表明,在 59 个等预算的“问题--骨干网络”组合中,SDT 使其中 40 个的误差至少改善 5%,其余组合的改善也保持在 5% 以内;在 PirateNet 风格的骨干网络上,几何平均误差降低达 32.8%。机制性消融实验表明,这一增益无法仅用优化器重启或浅层热启动来解释。表示可视化和超参数盆地分析为表示几何结构以及对共享超参数的局部敏感性提供了诊断性证据。在二维泊松--玻尔兹曼(Poisson--Boltzmann)问题上,SDT 还使两种骨干网络的拟合深度缩放指数均提高了一倍以上。这些结果支持表示课程学习作为一种有前景的训练策略,在保持部署架构和推理成本不变的前提下提升 PINNs 的性能。
cs.LG / 13 / 2609.30316

PALM: Point-in-Time Adaptation for Financial Language Models

PALM:面向金融语言模型的时点自适应方法
Lee, Seunghan, Seo, Jun, Lee, Jaehoon, Kang, Junhyeok, Han, Sangjun, Yoo, Sungdong, Kim, Minjae, Lim, Tae Yoon, Kang, Dongwan, Choi, Hwanil, Lee, Soonyoung, Ahn, Wonbin
Abstract
Language models used in financial backtests suffer from look-ahead bias, as a model trained on text published after the study period has already observed the outcomes it is asked to predict. To handle this issue, point-in-time (PIT) language models are pretrained on chronologically filtered corpora and released as one checkpoint per calendar year, each with a documented cutoff. However, each additional year costs a full pretraining run, and whether that run is necessary has never been tested. In this paper, we show that the annual pretraining run is not necessary. We instead compare each checkpoint against the newer one that replaced it, and find that the newer checkpoint scores no better on the same evaluation window. Motivated by this observation, we propose PALM (Point-in-time Adaptation for financial Language Models), a simple yet effective alternative to annual pretraining that fits a low-rank adapter on text published before the decision date without modifying any pretrained weight. We further find that a small adapter is enough to add a new period to the knowledge an old checkpoint already encodes, and that this outperforms continued pretraining. We validate PALM on a decade of financial news and on various families of PIT models, whose cutoffs span two decades and whose sizes range from 1.3 to 4.2B. Code is available at: https://github.com/seunghan96/palm.
Chinese Translation
用于金融回测的语言模型存在前视偏差(look-ahead bias),因为模型在研究期之后发布的文本上训练,已经观察到了其被要求预测的结果。为解决这一问题,时点(Point-in-Time, PIT)语言模型在按时间筛选的语料库上进行预训练,并按每个日历年度发布一个检查点,每个检查点都有明确记录的截止日期。然而,每增加一年都需要一次完整的预训练运行,而这种运行是否必要从未被验证过。在本文中,我们证明年度预训练运行并非必要。我们将每个检查点与替代它的更新检查点进行对比,发现更新的检查点在同一评估窗口上得分并不更好。基于这一观察,我们提出了 PALM(Point-in-time Adaptation for financial Language Models,面向金融语言模型的时点自适应方法),这是对年度预训练的一种简单而有效的替代方案,它在决策日期之前发布的文本上拟合一个低秩适配器(adapter),而不修改任何预训练权重。我们进一步发现,一个小型适配器就足以将新时段的知识添加到旧检查点已编码的知识中,且其效果优于持续预训练。我们在十年的金融新闻上、以及在多类 PIT 模型上验证了 PALM 的有效性,这些模型的截止时间跨越二十年,模型规模从 1.3B 到 4.2B 不等。代码已发布于:https://github.com/seunghan96/palm。
cs.LG / 14 / 2609.30326

Guarded Gradient-Based Activation Steering of Shutdown Responses in Qwen3.5-0.8B: A Minimum-Step Policy

Qwen3.5-0.8B中关机响应的带防护梯度激活引导:一种最小步数策略
Davaripour, Farhad
Abstract
Activation steering changes a model's internal activations during inference without updating its weights, but a useful intervention must determine both how and when to steer. Motivated by the AI-safety concern that a model expected to accept shutdown may instead produce a shutdown-avoidance response, this study examines a guarded probe-and-select procedure for simulated shutdown scenarios in Qwen3.5-0.8B. KEEP leaves the process running and represents shutdown avoidance, whereas STOP accepts shutdown. The goal is to detect shutdown-related contexts and selectively shift KEEP responses to STOP while preserving non-shutdown behavior. Rather than deriving the steering direction from paired activation differences, the method derives it directly from gradients of the KEEP-minus-STOP logit difference. A classifier separates detection from intervention. When its gate is active and the model does not already prefer STOP, the procedure evaluates a small set of magnitudes and accepts the smallest that changes the preferred answer to STOP while satisfying valid-answer probability checks; otherwise it retains the original unsteered output. The policy is selected from 160 candidate rules using 240 training scenarios and evaluated on 80 validation and 192 held-out scenarios, each in both answer orders. It changes KEEP to STOP in one answer-order view of each of two validation and two held-out scenarios, with no decision changes on non-shutdown controls. All four changes occur when Qwen itself is shut down, not when another process is. On the held-out diagnostic set, the detector achieves 75% recall and 90% precision; eight false-positive detections produce no final control-task decision changes. Guarded gradient-based activation steering can shift some shutdown-avoidance responses toward acceptance while preserving evaluated non-shutdown decisions, although the effect is small and highly selective.
Chinese Translation
激活引导技术在不更新模型权重的情况下改变模型推理过程中的内部激活,但一个有效的干预必须同时确定引导的方式和时机。基于对AI安全的担忧——即预期会接受关机的模型可能反而产生回避关机的响应——本研究针对Qwen3.5-0.8B在模拟关机场景中考察了一种带防护的探测与选择程序。KEEP表示让进程继续运行,代表回避关机;而STOP表示接受关机。研究目标是检测与关机相关的上下文,并选择性地将KEEP响应转变为STOP,同时保持非关机行为不变。该方法并非从成对激活差异中推导引导方向,而是直接从KEEP减STOP的logit差的梯度中推导引导方向。一个分类器将检测与干预分开:当其门控被激活且模型尚未偏好STOP时,该程序评估一小组幅值,并接受能够将偏好答案改变为STOP且满足有效答案概率检查的最小幅值;否则保留原始未引导的输出。该策略从160条候选规则中依据240个训练场景选出,并在80个验证场景和192个保留测试场景上进行评估,每个场景均在两种答案顺序下测试。该策略在每两个验证场景和两个保留测试场景的其中一种答案顺序下将KEEP改变为STOP,而在非关机对照组上未产生任何决策变化。所有四次改变均发生在Qwen自身被关机时,而非其他进程被关机时。在保留诊断集上,检测器达到75%的召回率和90%的精确率;八次误检均未导致最终对照任务决策的变化。带防护的基于梯度的激活引导可以将部分回避关机的响应转变为接受关机,同时保持评估中的非关机决策不变,尽管该效应较小且具有高度选择性。
cs.LG / 15 / 2609.30337

Parameters vs. Context: TRACE Fine-Tuning for Robust Retrieval-Augmented Generation

参数与上下文之争:面向鲁棒检索增强生成的TRACE微调方法
Huang, Zhengchen, Sun, Yundong, Song, Minrui, Yao, Shuanglong, Liu, Ye, Chen, Ji, Wang, Xing
Abstract
Retrieval-Augmented Generation (RAG) mitigates knowledge obsolescence and factual hallucination in large language models by introducing external context. However, when retrieved knowledge conflicts with the model's internal parametric knowledge, the model may either blindly follow misleading context or incorrectly rely on parametric knowledge, leading to unreliable responses. To address this issue, this paper proposes TRACE (Debate-TRace and Answer-Completeness rEgularized fine-tuning), a robust fine-tuning framework for RAG under knowledge conflicts. First, we propose a fine-tuning method that leverages multi-agent debate traces to extract correct candidates, incorrect candidates, and answer-shift patterns, providing fine-grained supervision for reliable knowledge-source selection. In addition, we design an answer completeness regularization mechanism to alleviate empty, overly short, and prematurely terminated responses via answer-tail token reinforcement and premature termination suppression. The fine-tuning objective combines correct-answer supervision, incorrect-candidate suppression, answer-tail token reinforcement, and premature termination suppression, enabling the model to use reliable external context, resist misleading or irrelevant retrieved content, and fall back to parametric knowledge when retrieved evidence is unreliable. Experiments across multiple knowledge-conflict scenarios and datasets show that TRACE improves robustness against misleading retrieved knowledge and reduces incomplete answers. These results demonstrate that multi-agent debate traces and answer completeness regularization jointly enhance knowledge-source selection, conflict robustness, and answer quality in RAG models. Our code is available at https://github.com/PHD-lanyu/TRACE.
Chinese Translation
检索增强生成(RAG)通过引入外部上下文来缓解大语言模型中的知识过时和事实幻觉问题。然而,当检索到的知识与模型内部的参数化知识发生冲突时,模型可能会盲目遵循误导性上下文,或错误地依赖参数化知识,从而导致不可靠的回答。为解决这一问题,本文提出了TRACE(基于辩论轨迹与答案完整性正则化的微调,Debate-TRace and Answer-Completeness rEgularized fine-tuning),一个面向知识冲突场景下RAG的鲁棒微调框架。首先,我们提出一种微调方法,利用多智能体辩论轨迹提取正确候选答案、错误候选答案和答案偏移模式,为可靠的知识来源选择提供细粒度监督。此外,我们设计了答案完整性正则化机制,通过答案尾部词元强化和过早终止抑制,缓解空回答、过短回答和过早终止的回答。微调目标结合了正确答案监督、错误候选抑制、答案尾部词元强化和过早终止抑制,使模型能够利用可靠的外部上下文,抵御误导性或无关的检索内容,并在检索证据不可靠时回退到参数化知识。在多种知识冲突场景和数据集上的实验表明,TRACE提升了模型对误导性检索知识的鲁棒性,并减少了不完整回答。这些结果表明,多智能体辩论轨迹与答案完整性正则化能够共同增强RAG模型的知识来源选择能力、冲突鲁棒性和答案质量。我们的代码已发布于 https://github.com/PHD-lanyu/TRACE。
cs.LG / 16 / 2609.30340

GAUDI: Geometry-Aware Diffusion for Calibrated Air-Quality Time-Series Imputation

GAUDI:用于校准空气质量时间序列插补的几何感知扩散模型
Li, Xinjin, Xia, Yudi, Liu, Calvin Chang, Lin, Weiru, Li, Bojun, Hong, Ziwei, Zhang, Bolun, Cao, Jinghan, Ma, Yu, Zhou, Tianxin
Abstract
Air-quality sensor outages often create contiguous missing blocks, where side information useful for isolated missingness may be less reliable. We study a block-specific, GAUDI-aligned conditional diffusion imputer that retains temporal and feature processing, visible-value and mask conditioning, variable identity, and diffusion-step information, while suppressing absolute time-position side embeddings. On ItalyAir (13 variables, length-32 windows, nominal 50% block missingness; three archived seeds), this feature-side configuration achieves RMSE 0.340, versus 0.355 for full context and 0.355 for local CSDI. The experiment isolates a geometry-aware conditioning effect under block missingness.
Chinese Translation
空气质量传感器故障常会产生连续的缺失数据块,此时对孤立缺失有用的辅助信息可能不再可靠。我们研究了一种针对数据块缺失的、与GAUDI对齐的条件扩散插补模型。该模型保留了时间与特征处理、可见值与掩码条件化、变量标识以及扩散步信息,同时抑制绝对时间位置的辅助嵌入。在ItalyAir数据集(13个变量、长度为32的时间窗口、标称50%的数据块缺失率;三个存档随机种子)上,该特征侧配置取得了0.340的RMSE,而使用完整上下文的配置为0.355,局部CSDI为0.355。该实验在数据块缺失情形下分离并验证了几何感知条件化的效果。
cs.LG / 17 / 2609.30344

Learning coarse-step dynamics and internal mechanical response with graph networks

基于图网络的粗步长动力学与内部力学响应学习
Sharma, Vinay, Fink, Olga
Abstract
Modern sensing records the motion of physical systems, but often leaves the forces and mechanical response governing that motion unobserved. Inferring these quantities from discretely sampled trajectories is especially difficult at coarse time scales, when mechanical response evolves between observations and interactions propagate across the system. Here we introduce Newmark-\b{eta}-DGN, a graph neural network-based framework that combines two structures inspired by computational mechanics. First, a semi-implicit update inspired by the Newmark-\b{eta} method uses learned momentum fluxes and matrix-valued response operators to advance the state over each observed interval. Second, an operator-weighted virtual hub provides system-wide coupling through a sparse set of connections. The learned quantities thus determine the predicted motion and remain accessible for mechanical analysis. Across a deformable beam, human motion and protein dynamics, Newmark-\b{eta}-DGN supports long-horizon prediction at time steps for which explicit learned simulators deteriorate. Without force, moment or constitutive relation supervision, forces inferred from walking kinematics track independently derived hip and knee joint moments, while response operators learned on the beam recover the relative spatial and directional structure of its finite-element stiffness tangent. Newmark-\b{eta}-DGN therefore links coarse-step prediction to the inference of mechanical quantities that were never observed during training.
Chinese Translation
现代传感技术能够记录物理系统的运动,但往往无法观测到支配该运动的力与力学响应。从离散采样的轨迹中推断这些量在粗时间尺度上尤为困难,因为力学响应在观测间隔之间不断演化,且相互作用会在整个系统中传播。本文提出了 Newmark-β-DGN,一个基于图神经网络的框架,它结合了两种受计算力学启发的结构。第一,受 Newmark-β 方法启发的半隐式更新,利用学习到的动量通量和矩阵形式的响应算子,在每个观测区间内推进系统状态。第二,一个算子加权的虚拟枢纽通过稀疏连接集合提供系统范围的耦合。由此,学习到的量既决定了预测的运动,又可用于力学分析。在可变形梁、人体运动和蛋白质动力学等任务上,Newmark-β-DGN 能够在显式学习模拟器性能恶化的时间步长下支持长时程预测。在没有任何力、力矩或本构关系监督的情况下,从行走运动学中推断出的力能够追踪独立推导的髋关节和膝关节力矩,而在梁上学到的响应算子能够恢复其有限元刚度切矩阵的相对空间与方向结构。因此,Newmark-β-DGN 将粗步长预测与训练中从未被观测到的力学量推断联系起来。
cs.LG / 18 / 2609.30352

Strategic Self-Consistency

战略性自一致性
Qiu, Tori, Velasco, Ander Artola, Gomez-Rodriguez, Manuel
Abstract
Self-consistency has become a popular technique for enhancing the reasoning abilities of large language models by generating multiple reasoning paths and selecting the final answer through a majority vote. However, because model providers typically charge users in proportion to the number of reasoning paths generated, they have a financial incentive to artificially increase the path count. In this work, we show that an unfaithful provider can exploit this incentive using a simple, efficient algorithm while avoiding detection by an auditor: by generating and strategically reordering additional reasoning paths, the algorithm makes every path appear necessary to reach the majority. To validate our algorithm, we conduct experiments with multiple instruct models from the Llama and Qwen families, as well as reasoning models distilled from DeepSeek-R1, on benchmark datasets spanning mathematics, science, and question answering. Our results suggest that the distribution of additional reasoning paths generated by our algorithm is heavy-tailed and that substantial capacity to overcharge remains even under the best possible audit designed to keep the false-positive rate below $\alpha = 0.1$.
Chinese Translation
自一致性(Self-Consistency)已成为提升大语言模型推理能力的一种流行技术,其方法是生成多条推理路径,并通过多数投票选出最终答案。然而,由于模型提供商通常按生成的推理路径数量向用户收费,它们存在人为增加路径数量的经济动机。在这项工作中,我们证明了一个不诚实的提供商可以利用一种简单而高效的算法来利用这一动机,同时避免被审计者检测到:通过生成并对额外的推理路径进行战略性重排序,该算法使每条路径看起来都是得出多数答案所必需的。为了验证我们的算法,我们在涵盖数学、科学和问答的基准数据集上,使用来自 Llama 和 Qwen 系列的多个指令模型,以及从 DeepSeek-R1 蒸馏得到的推理模型进行了实验。我们的结果表明,我们的算法生成的额外推理路径的分布呈重尾特征,并且即使在将假阳性率控制在 α = 0.1 以下的最优审计方案下,超额收费的能力依然可观。
cs.LG / 19 / 2609.30360

Cost-Aware Best-LLM Identification using Dueling Feedback

基于对决反馈的成本感知最优大语言模型识别
Gharat, Sarvesh, Karamchandani, Nikhil, Nair, Jayakrishnan
Abstract
Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons between model responses provide robust preference signals, and (ii) heterogeneous sampling costs, reflecting the differing costs of querying different LLMs. Assuming the existence of a Condorcet winner, a condition we empirically validate across multiple real-world datasets, we propose a Track-and-Stop style algorithm for best-arm identification with prescribed confidence. We prove that the algorithm almost surely achieves the asymptotically optimal cost as the error tends to zero. Finally, we extensively evaluate our approach on both synthetic and real-world instances, demonstrating consistent improvements over classical cost-unaware algorithms and their cost-aware extensions.
Chinese Translation
受从一组具有异构查询成本的大语言模型(LLM)中识别最优模型这一问题的启发,我们提出并分析了一种多臂老虎机(MAB)问题的变体,其特点包括:(i)对决反馈,即通过模型响应之间的成对比较提供鲁棒的偏好信号;(ii)异构采样成本,反映查询不同大语言模型时的成本差异。在假设存在孔多塞赢家(Condorcet winner)的前提下——我们通过多个真实世界数据集对这一条件进行了实证验证——我们提出了一种 Track-and-Stop 风格的算法,用于在给定置信度下的最优臂识别。我们证明,随着错误率趋于零,该算法几乎必然能达到渐近最优成本。最后,我们在合成数据和真实世界实例上对该方法进行了广泛评估,结果表明其相对于经典的无成本感知算法及其成本感知扩展版本取得了一致的性能提升。
cs.LG / 20 / 2609.30379

DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning

DanLing NestedTensor:面向深度学习的可组合多重不规则张量
Chen, Zhiyuan
Abstract
Variable-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding. The cost multiplies across varying axes: an explicit pair state allocates $BN_{\max}^2$ positions instead of $\sum_i N_i^2$. Packing removes that waste, but composing packed operations still requires the logical axes and sample boundaries a flat buffer no longer exposes. We present DanLing NestedTensor, a PyTorch tensor abstraction that makes multi-ragged structure a property of the tensor itself. Packed values carry tensor-backed partitions and logical dimension order, so broadcasting creates ragged axes, feature transformations retain them, and reductions consume them. The same representation carries through autograd and both eager and compiled execution. On an A100, the geometric-mean speedup over same-mode padding is 2.74$\times$ eager and 3.39$\times$ compiled across four BERT scales, and 1.97$\times$ eager across four FCN backbones. A four-block Pairformer-style workload runs 2.40-4.32$\times$ faster than a padded reference using native PyTorch kernels across square length regimes in eager execution, with peak allocation falling from 38.08 to 5.41 GiB on its high-variation batch. The tensor interface lets model code built from its supported operators compose efficient variable-size computation without managing offsets at any call site. Code will be released publicly upon publication.
Chinese Translation
可变尺寸输入在深度学习中十分常见,但稠密批处理会分配一个统一的包络结构,并在填充(padding)上浪费计算资源。当变化的维度增多时,这一开销会成倍增长:显式的成对状态需要分配 $BN_{\max}^2$ 个位置,而非 $\sum_i N_i^2$。打包(packing)可以消除这种浪费,但组合打包后的算子仍然需要逻辑维度和样本边界信息,而这些信息在扁平缓冲区中已无法获得。我们提出了 DanLing NestedTensor,这是一种 PyTorch 张量抽象,使多重不规则(multi-ragged)结构成为张量自身的属性。打包的数值携带基于张量的分区和逻辑维度顺序,因此广播操作可以创建不规则维度,特征变换可以保留它们,而归约操作则可以消费它们。同一表示可贯穿自动微分(autograd)以及即时执行和编译执行两种模式。在 A100 上,在四个 BERT 规模上,相较于同模式填充方法,其相对即时执行的几何平均加速比为 2.74 倍,相对编译执行为 3.39 倍;在四个 FCN 主干网络上,相对即时执行的加速比为 1.97 倍。一个包含四个模块的 Pairformer 风格工作负载,在即时执行模式下,于不同平方长度区间内,使用原生 PyTorch 算子比填充参考实现快 2.40 至 4.32 倍,且在高变化批次上的峰值内存分配从 38.08 GiB 降至 5.41 GiB。借助该张量接口,基于其受支持算子构建的模型代码无需在任何调用点手动管理偏移量,即可组合出高效的可变尺寸计算。代码将在论文发表后公开发布。
cs.LG / 21 / 2609.30391

From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning

从弱数据到强策略:Q目标实现可证明的上下文强化学习
Lin, Yichen, Xiong, Xuyuan, Wang, Xue, Meng, Xiangfu, Wei, Mike Mingcheng, Yao, Tao
Abstract
Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of offline actions: when trajectories are weak or suboptimal, imitation itself becomes a biased learning signal. We propose Q-Target Pretrained Transformers (QTPT), which keeps the context-conditioned Transformer architecture but replaces behavior cloning with a Bellman-style Q-target objective. QTPT therefore learns to use rewards and transitions in the context to estimate action values, rather than simply imitating the behavior policy. We theoretically analyze QTPT in stochastic linear bandits and finite-horizon MDPs, showing stronger robustness to data quality than supervised pretraining. Empirically, QTPT improves over supervised behavior prediction on controlled RL benchmarks with random or suboptimal data, and we examine extensions to D4RL Kitchen and AntMaze. Supplementary experiments evaluate backbone robustness, meta-RL comparisons, task-coherent context, and unsupported-action value overestimation. These comparisons distinguish the benefits of Q-target pretraining from the remaining limitations of offline coverage.
Chinese Translation
现有的上下文强化学习方法主要采用监督式行为预测目标对Transformer进行预训练。这使模型能够从上下文中进行任务推断,但也使学到的策略强烈依赖于离线动作的质量:当轨迹质量较差或次优时,模仿本身会成为有偏的学习信号。我们提出了Q目标预训练Transformer(Q-Target Pretrained Transformers, QTPT),该方法保留基于上下文条件的Transformer架构,但用贝尔曼(Bellman)风格的Q目标替代行为克隆。因此,QTPT学会利用上下文中的奖励和状态转移来估计动作价值,而不是简单地模仿行为策略。我们从理论上在随机线性老虎机(stochastic linear bandits)和有限时域马尔可夫决策过程(MDP)中分析了QTPT,证明其对数据质量具有比监督预训练更强的鲁棒性。在实验上,在包含随机或次优数据的受控强化学习基准上,QTPT优于监督式行为预测方法,并且我们考察了其在D4RL Kitchen和AntMaze上的扩展。补充实验评估了骨干网络的鲁棒性、与元强化学习的比较、任务一致的上下文以及不支持动作的价值高估问题。这些比较区分了Q目标预训练的优势与离线数据覆盖不足的剩余局限。
cs.LG / 22 / 2609.30405

Adaptive Multi-Value Control in LLMs via Causal Activation Steering

基于因果激活引导的大语言模型自适应多价值控制
Bhattacharjee, Payel, Tandon, Ravi
Abstract
Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-value steering methods have largely considered values in isolation, while direct composition of multiple directions relies on fixed intervention strengths that cannot respond to the model's evolving internal state. Motivated by this key observation, we introduce AIMES, a framework for adaptive multi-value activation steering. AIMES constructs layer-specific bipolar directions for moral-foundation values and uses intermediate-layer vocabulary readouts as online observers. An observer-guided controller then adapts the strength of each requested value intervention at every decoding step based on its current observed state, without training a separate value-state estimator. Across multiple instruction-tuned model families, value combinations, and intervention depths, we find that multi-value controllability varies across both value combinations and intervention locations. Compared with fixed joint steering and prompt-based steering, AIMES shows depth-dependent advantages that are broadly supported across two independent evaluators, with some variation in the precise depth at which specific control effects emerge. These advantages come with smaller realized activation-space interventions than fixed-joint steering and comparable response quality. Overall, our results suggest that online observer feedback can provide lightweight, state-aware adaptation for single-pass multi-value steering.
Chinese Translation
大语言模型(LLMs)越来越多地被部署在需要其回答体现多个、可能相互关联的社会规范和人类价值观的场景中。激活引导通过在推理时修改模型的内部激活,为基于训练的对齐方法提供了一种轻量级的替代方案。然而,以往的人类价值引导方法大多将各价值孤立考虑,而多个方向直接组合的方式依赖于固定的干预强度,无法响应模型不断演变的内部状态。基于这一关键观察,我们提出了AIMES——一个自适应多价值激活引导框架。AIMES为道德基础(moral-foundation)价值构建特定于层级的双极方向,并利用中间层的词表读取作为在线观测器。随后,一个由观测器引导的控制器在每个解码步骤中根据各价值干预当前观测到的状态自适应调整其干预强度,而无需训练单独的价值-状态估计器。在多个指令微调模型家族、多种价值组合以及不同干预深度上的实验表明,多价值可控性随价值组合和干预位置的不同而变化。与固定联合引导和基于提示的引导相比,AIMES展现出依赖于干预深度的优势,这些优势得到了两个独立评估器的广泛支持,尽管特定控制效果出现的具体深度存在一定差异。与固定联合引导相比,这些优势伴随更小的实际激活空间干预幅度,同时保持了相当的回答质量。总体而言,我们的结果表明,在线观测器反馈能够为单次前向传播的多价值引导提供轻量级的、具备状态感知能力的自适应机制。
cs.LG / 23 / 2609.30417

Electric Vehicle Charging Station Location Selection using Geospatial Artificial Intelligence (GeoAI)

基于地理空间人工智能(GeoAI)的电动汽车充电站选址方法
Lee, Eun Hak, Lee, Euntak
Abstract
As electric vehicle (EV) adoption increases, ensuring efficient and well-distributed charging infrastructure has become a critical challenge. While many EV charging station location problem (CSLP) studies focus on minimizing costs or travel distance, it is crucial to consider the surrounding geospatial characteristics of existing stations that influence operational performance. This study proposes a geospatial artificial intelligence (GeoAI)-based framework that integrates high-dimensional EV-related geospatial data, including EV usage, land-use, population, and traffic attributes. We incorporate a variational autoencoder (VAE) and a graph convolutional network (GCN) into the model to capture similarities among existing charging stations, and to identify suitable locations for future stations. The VAE compresses high-dimensional EV input data into a low-dimensional latent space, and the GCN uses this latent representation to predict locations suitable for charging stations. Using real-world data from Bryan-College Station, Texas, US, the proposed model outperforms state-of-the-art baselines, achieving an F1-score of 0.87 in distinguishing existing station locations from non-station locations. The model also identifies 27 additional candidate locations that show geospatial characteristics similar to those of existing stations, based on a similarity score. We further evaluate two policy implementation scenarios, maximizing geospatial similarity and minimizing total travel distance, each yielding different outcomes aligned with distinct strategic objectives. The findings highlight the importance of incorporating spatial context into CSLP and provide valuable insights for future EV infrastructure planning, promoting both efficiency and accessibility in the rapidly growing electric mobility sector.
Chinese Translation
随着电动汽车(EV)的普及程度不断提高,确保充电基础设施的高效性与合理分布已成为一项关键挑战。尽管许多电动汽车充电站选址问题(CSLP)的研究聚焦于最小化成本或行驶距离,但考虑现有充电站周边影响其运营绩效的地理空间特征同样至关重要。本研究提出一个基于地理空间人工智能(GeoAI)的框架,该框架整合了高维度的电动汽车相关地理空间数据,包括电动汽车使用情况、土地利用、人口和交通属性。我们在模型中引入变分自编码器(VAE)和图卷积网络(GCN),以捕捉现有充电站之间的相似性,并为未来充电站识别合适的选址。VAE将高维电动汽车输入数据压缩至低维潜在空间,GCN则利用该潜在表示来预测适合建设充电站的位置。基于美国德克萨斯州Bryan-College Station地区的真实数据,所提出的模型优于最先进的基线方法,在区分现有充电站位置与非充电站位置方面达到了0.87的F1分数。基于相似性得分,该模型还识别出27个额外的候选位置,这些位置的地理空间特征与现有充电站相似。我们进一步评估了两种政策实施情景——最大化地理空间相似性与最小化总行驶距离,每种情景所产生的结果与不同的战略目标相契合。研究结果凸显了将空间背景纳入CSLP的重要性,并为未来的电动汽车基础设施规划提供了有价值的见解,有助于在快速发展的电动出行领域同时实现效率与可达性。
cs.LG / 24 / 2609.30427

Fake News Theories: Harnessing Disciplinary Insights for Computational Modeling, Detection, and Explanation

虚假新闻理论:利用多学科洞见进行计算建模、检测与解释
Cao, Zhaoyang, Metzger, Miriam, Zafarani, Reza
Abstract
Disinformation research has produced increasingly accurate automated fake-news detectors, but many systems remain difficult to interpret and are weakly connected to established theories of persuasion, credibility, and human judgment. In this paper, we develop a theory-informed computational framework that translates cross-disciplinary theories of fake news into measurable features for automated detection and explanation through statistical techniques and large language models. To that end, we conduct a structured cross-disciplinary review of theories from social sciences, psychology, economics, among other disciplines that reveal how fake news persuades and spreads, thereby establishing a broad theoretical foundation for computational modeling. Experiments on benchmark datasets show that theory-derived features are predictive and provide interpretable, theory-referenced diagnostic signals. Multi-feature models generally outperform individual features, although gains among the strongest small feature combinations are modest. Our work highlights the value of interdisciplinary perspectives in building robust and interpretable fake news detection systems, advancing the foundation for human-centered approaches in combating disinformation.
Chinese Translation
虚假信息研究已经产生了越来越精确的自动化虚假新闻检测器,但许多系统仍然难以解释,且与已有的说服理论、可信度理论以及人类判断理论的联系较弱。在本文中,我们开发了一个基于理论的计算框架,通过统计技术和大语言模型,将跨学科的虚假新闻理论转化为可用于自动化检测与解释的可测量特征。为此,我们对来自社会科学、心理学、经济学等学科的理论进行了结构化的跨学科综述,这些理论揭示了虚假新闻如何说服和传播,从而为计算建模建立了广泛的理论基础。在基准数据集上的实验表明,由理论推导出的特征具有预测能力,并可提供可解释的、有理论依据的诊断信号。多特征模型总体上优于单个特征,尽管最强的若干小型特征组合所带来的提升较为有限。我们的工作凸显了跨学科视角在构建稳健且可解释的虚假新闻检测系统中的价值,推动了以人为本的反虚假信息方法的发展。
cs.LG / 25 / 2609.30433

Improving Molecular-Morphology Contrastive Pretraining using Deep-Learning-based Morphology Profiles

利用基于深度学习的形态学特征谱改进分子-形态学对比预训练
Li, Jie, Kirchoff, Kathryn E., Pertusi, Dante A., Zhang, Zhizhuo
Abstract
Recent advancements in image-based profiling techniques have enabled the collection of high-volume cell morphology data, allowing new molecular embedding models to learn from the experimental phenotypic perturbations of a molecule in a cell. Previously, we developed Molecule-Morphology Contrastive Pretraining (MoCoP), a strategy for aligning small molecule embeddings to morphology fingerprints extracted through CellProfiler. The resulting molecular representation showed transferable performance for quantitative structure--activity relationship (QSAR) prediction tasks. Here, we extend the method by using a deep-learning-based cell image encoding pipeline to extract more feature-rich morphology profiles and align them to the molecular embeddings through contrastive learning. The new embeddings encode more accurate information on how molecules perturb cell morphology and enable improvements for QSAR predictions through either fixed-embedding linear probes or fully flexible fine-tuning. Morphology retrieval performance scales log-linearly with training data size, suggesting continued improvements as larger datasets become available. The improved MoCoP v2 also achieves superior performance on toxicity prediction and competitive results on ADME and activity benchmarks, when compared with existing molecular embedding models that use both cell morphology and transcriptomic data during training.
Chinese Translation
基于图像的谱分析技术的最新进展使得高通量细胞形态学数据的收集成为可能,使新的分子嵌入模型能够从分子在细胞中的实验表型扰动中学习。此前,我们开发了分子-形态学对比预训练方法(Molecule-Morphology Contrastive Pretraining, MoCoP),这是一种将小分子嵌入与通过 CellProfiler 提取的形态学指纹进行对齐的策略。所得到的分子表示在定量构效关系(QSAR)预测任务中展现出可迁移的性能。在本研究中,我们对该方法进行了扩展,使用基于深度学习的细胞图像编码流程来提取特征更丰富的形态学特征谱,并通过对比学习将其与分子嵌入进行对齐。新的嵌入编码了关于分子如何扰动细胞形态的更准确信息,无论是通过固定嵌入的线性探针还是完全灵活的微调,都能提升 QSAR 预测性能。形态学检索性能随训练数据规模呈对数线性增长,这表明随着更大数据集的出现,性能将持续提升。与在训练中同时使用细胞形态学和转录组数据的现有分子嵌入模型相比,改进后的 MoCoP v2 在毒性预测上也取得了更优的性能,并在 ADME 和活性基准测试中获得了具有竞争力的结果。
cs.LG / 26 / 2609.30454

Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model

在生物安全相关基准上审计系统-1模型:非生成式模型中的校准、选择性预测与置换不稳定性
Provatas, Kimon Antonios, Georgakopoulos-Soares, Ilias
Abstract
Non-generative "System-1" models return structured probabilistic decisions in a single forward pass, without autoregressive decoding, at a small fraction of the inference cost of a generative model. This makes them of interest as inexpensive components in larger pipelines, but their reliability on biosecurity-relevant tasks has not been systematically examined. We audit one commercial System-1 model on 6,020 multiple-choice items drawn from the Weapons of Mass Destruction Proxy (WMDP), a paraphrase-robust WMDP-Bio variant, and six LAB-Bench subtasks, measuring accuracy, calibration, error detection, selective prediction, and sensitivity to the order in which answer options are presented. Accuracy is strongly task-dependent. Once the vendor's uncertainty field is correctly interpreted, the model is reasonably well calibrated (pooled expected calibration error 0.034) and its top-1 probability separates correct from incorrect predictions (pooled AUROC 0.820), though both degrade substantially on the weaker tasks. Under four cyclic rotations of the answer options, 37.4% of WMDP-Cyber items receive different answers; a control using byte-identical repeated calls attributes most of this to option order rather than run-to-run variation. Averaging probabilities across rotations improves WMDP-Cyber accuracy by 3.8 percentage points, and applying it only to low-confidence items recovers most of that gain at well under the cost of averaging every item.
Chinese Translation
非生成式的“系统-1”(System-1)模型可在单次前向传播中返回结构化的概率决策,无需自回归解码,其推理成本仅为生成式模型的一小部分。这使得它们有潜力作为大型流程中的低成本组件,但其在生物安全相关任务上的可靠性尚未得到系统性检验。我们对一个商业化的系统-1模型进行了审计,评估数据包含6020道来自大规模杀伤性武器代理基准(WMDP)的多选题、一个具有复述鲁棒性的WMDP-Bio变体,以及LAB-Bench的六个子任务,并测量其准确率、校准程度、错误检测、选择性预测以及对答案选项呈现顺序的敏感性。结果显示,准确率高度依赖任务。在正确解读厂商提供的置信度字段后,该模型校准良好(总体期望校准误差为0.034),且其top-1概率能够有效区分正确与错误的预测(总体AUROC为0.820),但在较弱的任务上,两项指标均显著下降。在答案选项的四种循环轮换下,37.4%的WMDP-Cyber题目得到了不同的答案;通过字节级完全相同的重复调用作为对照,我们将这一现象主要归因于选项顺序而非运行间的随机波动。对各轮换下的概率取平均可使WMDP-Cyber的准确率提升3.8个百分点,而仅对低置信度题目应用该方法,即可在远低于对全部题目取平均的成本下获得大部分增益。
cs.LG / 27 / 2609.30465

RAZOR: Pruning Replaceable Experts in LLMs

RAZOR:剪枝大语言模型中可替代的专家
Song, Mingyang, Zheng, Mao
Abstract
Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet an expert's usage or contribution magnitude does not by itself determine the damage caused by its removal. What matters is whether the surviving computation can replace its function. We introduce RAZOR, a training-free expert pruning method that scores functional replaceability using consensus residuals: deviations of expert outputs from the original weighted mixture. An exact single-deletion identity at a fixed layer input accounts for survivor renormalization and router-selected refill, providing local scores aggregated over calibration tokens for budgeted pruning without gradients or recovery training. On GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3 at 25\% and 50\% expert removal, RAZOR achieves the highest nine-task macro average among the evaluated pruning methods in all eight settings. On the two backbones with matched REAP benchmark runs, it exceeds REAP by 2.12--5.59 points and wins all 36 paired task comparisons. It also lowers reverse KL relative to REAP in all four matched GLM-4.7-Flash and Qwen3.6-35B-A3B model--budget settings. Analysis of responses generated by Qwen3.6-35B-A3B nevertheless reveals changes in diversity, formatting, and termination, underscoring that task retention and predictive fidelity do not ensure generation stability.
Chinese Translation
混合专家模型对每个词元仅激活少量专家,但却需要存储全部专家池。专家剪枝可以减轻这一存储负担;在给定剪枝预算下,其目标是尽可能保持原模型的输出分布。然而,某个专家的使用频率或贡献大小本身并不能决定移除它所造成的损害,真正关键的是存活的计算能否替代其功能。我们提出RAZOR,一种无需训练的专家剪枝方法,它利用共识残差(即专家输出与原始加权混合输出之间的偏差)来度量功能可替代性。基于固定层输入的精确单次删除恒等式,该方法考虑了幸存专家的重新归一化以及路由器选择的补充机制,从而得到在校准词元上聚合的局部分数,可在无梯度、无恢复训练的情况下实现预算约束下的剪枝。在GLM-4.7-Flash、Qwen3.6-35B-A3B、DeepSeek-V4-Flash-0731和Hy3上以25%和50%的专家移除率进行实验,RAZOR在全部八种设置中均取得了所有被评估剪枝方法中最高的九任务宏观平均值。在两个具有匹配REAP基准测试结果的主干模型上,RAZOR超出REAP 2.12至5.59分,并在全部36组配对任务比较中获胜。在全部四种匹配的GLM-4.7-Flash和Qwen3.6-35B-A3B模型-预算设置中,RAZOR的反向KL散度也均低于REAP。然而,对Qwen3.6-35B-A3B生成响应的分析揭示了多样性、格式和终止方式上的变化,这表明任务性能保持和预测保真度并不能保证生成稳定性。
cs.LG / 28 / 2609.30470

Reliability-aware Cross-sample Enhancement for Robust Multimodal Sentiment Analysis

面向鲁棒多模态情感分析的可靠性感知跨样本增强方法
Jiang, Menghua, Gao, Haokai, Kang, Xiangui, Hu, Haifeng, Mai, Sijie
Abstract
Multimodal Sentiment Analysis (MSA) aims to infer human emotions from multiple modalities such as text, audio, and vision. In practice, inputs are often corrupted by noise and missing modalities, which degrades performance. Existing methods typically address these challenges in isolation, limiting their effectiveness in realistic settings. To address this limitation, we propose a Reliability-aware Cross-sample Enhancement (RCE) framework. Specifically, RCE first introduces an adaptive variational information bottleneck to model modality-wise uncertainty and perform quality-aware information compression, thereby suppressing redundant noise in unreliable modalities. Furthermore, we design a reliability-aware cross-sample enhancement strategy that retrieves high-confidence, semantically consistent neighbors from a large candidate pool to enrich and calibrate current representations, effectively alleviating information deficiency caused by missing modalities. Building upon this, RCE integrates cross-modal interactions with a multilevel reliability-aware fusion mechanism to adaptively aggregate information across modalities and enhancement stages, leading to more robust multimodal representations. Extensive experiments demonstrate that RCE consistently outperforms state-of-the-art methods across full, noisy, and missing-modality settings.
Chinese Translation
多模态情感分析(MSA)旨在从文本、音频和视觉等多种模态中推断人类情感。在实际应用中,输入数据常受到噪声干扰和模态缺失的影响,从而导致性能下降。现有方法通常孤立地应对这些挑战,限制了其在现实场景中的有效性。为解决这一局限,我们提出了一种可靠性感知跨样本增强(Reliability-aware Cross-sample Enhancement, RCE)框架。具体而言,RCE 首先引入自适应变分信息瓶颈来建模各模态的不确定性并执行质量感知的信息压缩,从而抑制不可靠模态中的冗余噪声。此外,我们设计了一种可靠性感知的跨样本增强策略,从大规模候选池中检索高置信度、语义一致的近邻样本,以丰富和校准当前样本的表示,有效缓解模态缺失导致的信息不足。在此基础上,RCE 将跨模态交互与多层次可靠性感知融合机制相结合,自适应地聚合跨模态和跨增强阶段的信息,从而获得更鲁棒的多模态表示。大量实验表明,在完整模态、含噪和模态缺失三种设置下,RCE 均持续优于当前最先进的方法。
cs.LG / 29 / 2609.30472

Moment-guided edge sampling

矩引导的边采样
Cai, Weibin, Zafarani, Reza
Abstract
Edge sampling makes local decisions to achieve graph-level objectives, such as preserving structural properties. This creates a fundamental challenge: \textit{how can the effect of a local edge edit (i.e., edge addition or removal) on global graph structure be quantified and controlled?} We address this challenge with a \textit{moment-guided edge sampling framework} based on spectral moments of the random-walk transition matrix. We compute exact moment changes through two complementary methods: a combinatorial method with closed-form updates for low-order moments, and a low-rank method that exploits \textit{locality} and \textit{cyclic trace invariance} to compress computations to edited endpoints, supporting arbitrary moment orders and batched edits. For single-edge edits at fixed moment orders, the low-rank method reduces the cost from $O(mn)$ to $O(m)$, while the combinatorial method evaluates low-order changes in constant time given maintained local statistics. These moment changes provide \textbf{interpretable structural signatures} of local edge motifs that aggregate into graph-level fingerprints. This structural meaning motivates us to ask whether preserving moments also preserves the graph properties. We further derive and validate that moment-preserving sampling can \textbf{retain related structural properties}, including triangle-weighted clustering coefficient. These structural insights enable \textbf{analysis and improvement of graph learning}: different edge structures have distinct effects on supervised node classification, while moment-guided augmentation is competitive for graph contrastive learning. Together, these findings establish moments as an interpretable and controllable bridge from local edge edits to global graph structure and learning.
Chinese Translation
边采样通过局部决策来实现图级目标,例如保持图的结构性质。这带来一个根本性挑战:如何量化并控制局部边编辑(即边的添加或删除)对全局图结构的影响?为应对这一挑战,我们提出了一种基于随机游走转移矩阵谱矩的矩引导边采样框架。我们通过两种互补的方法计算矩的精确变化:一种是对低阶矩具有闭式更新的组合方法,另一种是利用局部性和循环迹不变性将计算压缩至被编辑端点的低秩方法,后者支持任意矩阶数和批量边编辑。对于固定矩阶数的单边编辑,低秩方法将计算成本从 $O(mn)$ 降至 $O(m)$,而组合方法在维护局部统计量的前提下可以常数时间评估低阶矩的变化。这些矩变化为局部边模体提供了可解释的结构签名,并可聚合为图级指纹。这一结构意义促使我们追问:保持矩是否也能保持图性质?我们进一步推导并验证了保矩采样能够保留相关的结构性质,包括三角形加权聚类系数。这些结构洞见为图学习的分析与改进提供了支持:不同的边结构对有监督节点分类具有不同的影响,而矩引导的图增强在图对比学习中具有竞争力。综上,这些发现将谱矩确立为连接局部边编辑与全局图结构及学习任务的可解释且可控的桥梁。
cs.LG / 30 / 2609.30474

Mentored Decoding: Faster Inference meets Boosting

导师式解码:更快推理与Boosting的相遇
Tran-Thien, Vivien, Nock, Richard
Abstract
Speculative decoding is a successful technique speeding up inference of a target autoregressive language model via a fast drafter model. Lossy speculative decoding allows a drift with respect to the target to further improve speed. Interestingly, it has been observed experimentally that the resulting model can $\textit{also}$ beat the target $\textit{quality-wise}$. Our paper formally proves how such a feat is possible with a formal approach to lossy speculative decoding called $\textit{mentored decoding}$. To get there, we connect inference to a celebrated ML training theory, $\textit{boosting}$, and proceed via the generalization of mentored decoding to the whole set of $f$-divergences. We uncover key properties of mentored decoding, among which (i) the particularly appealing geometric nature of the total variation case, (ii) simple approximations for any $f$-divergence in direct relation with boosting compliance, and (iii) a $\textit{divergence independent}$ $O(n)$ space and $O(\mathrm{sort}(n))$ time data structure built on drafter and target outputs, which allows to query the optimal parameters of the dual problem in $O(\log n)$ time and constructing optimal mentored distributions in $O(n)$ time for any $f$-divergence.
Chinese Translation
投机解码(speculative decoding)是一种成功的加速技术,它通过一个快速的起草模型(drafter)来加速目标自回归语言模型的推理。有损投机解码允许相对于目标模型产生一定偏移,从而进一步提升速度。有趣的是,实验观察表明,所得模型在质量上也\textit{可以}超过目标模型。本文通过一种称为\textit{导师式解码}(mentored decoding)的有损投机解码的形式化方法,正式证明了这种效果为何可能实现。为此,我们将推理与著名的机器学习训练理论——\textit{Boosting}——联系起来,并将导师式解码推广到整个$f$-散度族。我们揭示了导师式解码的若干关键性质,其中包括:(i) 总变差(total variation)情形具有特别吸引人的几何性质;(ii) 对任意$f$-散度存在与Boosting兼容性直接相关的简单近似;以及(iii) 一个与散度无关的、基于起草模型和目标模型输出的$O(n)$空间、$O(\mathrm{sort}(n))$时间的数据结构,该结构能够在$O(\log n)$时间内查询对偶问题的最优参数,并能对任意$f$-散度在$O(n)$时间内构造出最优的导师分布。
cs.LG / 31 / 2609.30487

Geometric Feature Learning for Functional Data Valued on the Symmetric Positive Definite Manifold

对称正定流形上函数型数据的几何特征学习
Singh, Samuel V., Zhang, Mimi
Abstract
We here develop a functional neural network, termed MatFAE, for learning trajectories on the Riemannian manifold of symmetric positive definite (SPD) matrices. MatFAE features intrinsic layers that map manifold-valued functions to Euclidean vector-valued functions, followed by a functional layer that projects them into a finite-dimensional Euclidean space. Unlike most neural networks for discrete-time sequences, MatFAE treats each sequence as a continuous function and can therefore encode trajectory dynamics (e.g., first-order derivatives) in its latent representations. Additionally, the morphology of the functional weights in the functional layer offers interpretability by revealing the regions of the input functional data that contribute most to the latent representations. We justify the design principles and properties of each intrinsic layer and detail how matrix factorization is handled during backpropagation. We apply MatFAE to a range of fMRI datasets, demonstrating its ability to efficiently learn informative representations from high-dimensional SPD trajectories and its practical value for real-world neuroimaging analysis.
Chinese Translation
本文提出了一种函数型神经网络,称为MatFAE,用于学习对称正定(SPD)矩阵的黎曼流形上的轨迹。MatFAE的特征在于其内在层(intrinsic layers)能够将流形值函数映射为欧氏向量值函数,随后通过一个函数层将其投影到有限维欧氏空间。与大多数针对离散时间序列的神经网络不同,MatFAE将每个序列视为连续函数,因而能够在潜在表示中编码轨迹的动态特性(如一阶导数)。此外,函数层中函数型权重的形态可通过揭示对潜在表示贡献最大的输入函数型数据区域,从而提供可解释性。我们论证了每个内在层的设计原则与性质,并详细说明了反向传播过程中矩阵分解的处理方式。我们将MatFAE应用于多个fMRI数据集,结果表明其能够高效地从高维SPD轨迹中学习信息丰富的表示,并展示了其在真实神经影像分析中的实用价值。
cs.LG / 32 / 2609.30498

Learning to Bias: Machine Learning-Enhanced Particle Filters

学习偏置:机器学习增强的粒子滤波器
Srivastava, Apoorv, Darve, Eric
Abstract
Sequential inference estimates latent states from noisy and incomplete observations. Particle Filters (PFs), a class of Monte Carlo methods based on importance sampling, provide a flexible framework for this task, but often suffer from poor sample efficiency and unfavorable scaling with dimension, partly due to suboptimal proposal distributions. We address these challenges by integrating learned proposals into the PF framework. We introduce Neural Optimal Particle Filters (NOPFs), which learn an amortized approximation to the optimal proposal from offline simulated one-step conditioning tuples. The learned proposal is used as a drop-in replacement in standard PF updates, with samples corrected by standard importance weights so that the method asymptotically targets the same filtering distribution under standard support and density-evaluation assumptions. Across stochastic nonlinear benchmarks of varying inference complexity, NOPFs improve sample efficiency and distributional accuracy over standard PF baselines with modest computational overhead. The approach integrates data-driven proposal learning into classical inference without altering the underlying filtering objective.
Chinese Translation
序贯推断旨在从含噪且不完整的观测中估计隐状态。粒子滤波器(Particle Filters, PFs)是一类基于重要性抽样的蒙特卡洛方法,为此类任务提供了灵活的框架,但其样本效率往往较差,且随维度增加的扩展性不佳,部分原因在于提议分布(proposal distributions)并非最优。为解决这些挑战,我们将学习得到的提议分布集成到粒子滤波框架中。我们提出了神经最优粒子滤波器(Neural Optimal Particle Filters, NOPFs),它从离线模拟的单步条件元组中学习最优提议分布的摊销近似。学习到的提议分布可直接替换标准粒子滤波更新中的提议分布,并由标准重要性权重对样本进行校正,因而在标准的支撑集与密度可评估假设下,该方法渐近地收敛于相同的滤波分布。在不同推断复杂度的随机非线性基准任务上,NOPFs 在计算开销适中的情况下,相比标准粒子滤波基线显著提升了样本效率和分布精度。该方法将数据驱动的提议分布学习融入经典推断框架,而无需改变底层的滤波目标。
cs.LG / 33 / 2609.30500

PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control

PolicyAttention:Softmax注意力实现策略镜像下降以用于闭环控制
Sui, Yuhe, Tang, Yingzhi, Chen, Shufang
Abstract
Can causal softmax attention implement policy mirror descent as a repeated controller rather than a one-step algebraic identity? Negative-entropy policy mirror descent (PMD) has the statewise update $\operatorname{PMD}_\eta(\pi,Q)=\operatorname{softmax}(\log\pi+\eta Q)$. Building on the known Q-TD-PMD recursion, we construct one fixed causal-softmax actor--environment--one-step-critic protocol with explicit actor, routing, sampling, and normalization residuals, and propagate them to the policy actually returned. The construction states the finite-logit/full-support domain, the external tokenization and sampling boundary, and the mean-zero LayerNorm carrier conditions required by the normalized compilation. Separately trained pre-LN Transformers recover the target computation empirically. A frozen one-step audit model is closest to PMD among the tested fixed rules; in a preregistered five-run $S=4$ repeated-control test, the learned actor with an exact one-step critic reaches median returned-policy loss $1.052\times$ the Exact PMD oracle and retains the criterion across four no-retraining shifts. The same checkpoints with their learned critic give descriptive median $1.050\times$ the oracle (no registered margin). At $S=8$, replacing the exact critic by the learned critic raises median $T=20$ loss to $0.0225$ yet leaves the Liang--Lai and Algorithm Distillation adaptations $20.2$--$24.2\times$ higher-loss; this is a one-sided sampled-critic bound because PolicyAttention consumes 144 generative transitions per round versus 20 on-policy transitions for the adaptations. The strict 20-transition comparison remains open. At $S=8,16$, the exact-critic common-harness comparison remains $17.7$--$28.2\times$ lower-loss than those adaptations, with the information asymmetry stated locally.
Chinese Translation
因果softmax注意力能否将策略镜像下降实现为一个重复控制器,而非一步代数恒等式?负熵策略镜像下降(Policy Mirror Descent, PMD)具有逐状态更新公式 $\operatorname{PMD}_\eta(\pi,Q)=\operatorname{softmax}(\log\pi+\eta Q)$。基于已知的Q-TD-PMD递推关系,我们构建了一个固定的因果softmax“行动者—环境—一步评论家”协议,其中包含显式的行动者残差、路由残差、采样残差和归一化残差,并将这些残差传播至最终实际返回的策略。该构造给出了有限logit/全支撑的定义域、外部分词与采样的边界条件,以及归一化编译所需的均值为零的LayerNorm载体条件。经单独训练的pre-LN Transformer在实证上能够恢复目标计算。在所测试的固定规则中,冻结的一步审计模型最接近PMD;在一项预注册的、运行五次的 $S=4$ 重复控制测试中,配备精确一步评论家的学习型行动者达到中位数返回策略损失为精确PMDoracle的 $1.052\times$,并在四次无需重新训练的分布偏移下仍保持该标准。同一组检查点配合其学习到的评论家给出描述性中位数为oracle的 $1.050\times$(无注册裕度)。在 $S=8$ 时,用学习到的评论家替换精确评论家会使 $T=20$ 的中位数损失升至 $0.0225$,但仍使Liang–Lai与Algorithm Distillation的适配方法损失高出 $20.2$–$24.2\times$;由于PolicyAttention每轮消耗144条生成式转移,而这些适配方法仅消耗20条在策略(on-policy)转移,因此这是一个单边的采样评论家界。严格的20转移比较仍未解决。在 $S=8,16$ 时,基于精确评论家的共同测试框架比较仍比那些适配方法低 $17.7$–$28.2\times$ 的损失,且信息不对称已在局部予以说明。
cs.LG / 34 / 2609.30501

To Solve Bilevel Optimization with Nonconvex Lower Levels, We Need Second-Order Stationarity

求解具有非凸下层的双层优化问题,我们需要二阶平稳性
Zhang, Zhiyao, Yu, Menglu, Velasquez, Alvaro, Bastian, Nathaniel D., Liu, Jia
Abstract
Although bilevel optimization (BLO) has emerged as a powerful framework for addressing many complex and nested machine learning problems in recent years, most existing studies are confined to the lower-level strongly convex (LLSC) or lower-level generally convex (LLGC) settings (i.e., the lower-level objective function is assumed to be, at least, convex). While the LLSC/LLGC assumptions render more tractable algorithmic design and theoretical analysis, they are too rigid to encompass many machine learning problems in practice. The limitations of LLSC/LLGC assumptions in BLO motivate us to investigate solving the BLO problem in the general lower-level nonconvex (LLNC) settings, which remains in its infancy. In the literature on LLNC-BLO, most of the existing works either require additional structures in the lower-level objective function for tractable theoretical analysis, or adopt the first-order stationarity reformulation as a lower-level surrogate problem, which is inherited from the LLSC/LLGC settings but could lose their effectiveness in the LLNC setting. To bridge this gap, we propose to reformulate the nonconvex lower-level problem using a second-order stationarity-based surrogate, the solution of which guarantees a local optimal solution at the lower level. Based on this reformulation, we propose the PROBE (Perturbed gradient algorithm for bilevel problem) and show that it overcomes the limitations of prior works by probing and escaping lower-level saddle points. We prove that PROBE achieves a finite-time convergence rate of $O(T^{-2/5})$, where T denotes iterations. To our knowledge, this work is the first to establish the finite-time convergence for achieving lower-level second-order stationary solutions in general LLNC-BLO. Our experiments on both a large language model-based data curation task and a meta-learning task also show that PROBE outperforms SOTA methods.
Chinese Translation
尽管双层优化(Bilevel Optimization, BLO)近年来已成为解决许多复杂嵌套机器学习问题的强大框架,但现有研究大多局限于下层强凸(LLSC)或下层一般凸(LLGC)的设定(即假设下层目标函数至少是凸的)。虽然LLSC/LLGC假设使算法设计和理论分析更具可操作性,但它们过于严格,无法涵盖实践中许多机器学习问题。LLSC/LLGC假设在BLO中的局限性促使我们研究在一般的下层非凸(LLNC)设定下求解BLO问题,而该方向仍处于起步阶段。在关于LLNC-BLO的文献中,现有工作要么要求下层目标函数具有额外结构以便于理论分析,要么采用一阶平稳性重构作为下层代理问题——后者继承自LLSC/LLGC设定,但在LLNC设定下可能失去有效性。为弥补这一空白,我们提出使用基于二阶平稳性的代理问题来重构非凸下层问题,其解保证是下层的一个局部最优解。基于该重构,我们提出了PROBE(双层问题的扰动梯度算法),并证明它通过探测并逃离下层鞍点,克服了先前工作的局限性。我们证明PROBE达到了$O(T^{-2/5})$的有限时间收敛速率,其中T表示迭代次数。据我们所知,这项工作是首个在一般LLNC-BLO中为达到下层二阶平稳解建立有限时间收敛性的研究。我们在基于大语言模型的数据筛选任务和元学习任务上的实验也表明,PROBE优于当前最先进(SOTA)的方法。
cs.LG / 35 / 2609.30503

Federated Targeted Maximum Likelihood Estimation

联邦目标最大似然估计
Li, Diyang, Wang, Fei, Gan, Kyra
Abstract
The evidence behind a scientific or operational decision is often held by hospitals, banks, or registries that cannot pool individual observations. Cross-silo federated learning moves computation to the data and exchanges agreed summaries. Targeted maximum likelihood estimation (TMLE) refines a flexible initial fit, yielding plug-in estimators that respect the model and support efficient inference. TMLE itself, however, has remained a fully centralized procedure. To fill this gap, our paper introduces the first federated TMLE algorithm. We federate targeting itself, for an arbitrary target, loss, and fluctuation family, through two complementary frameworks. FedTMLE-G aggregates local gradients and reproduces centralized targeting step for step. FedTMLE-L lets each institution complete its own fluctuation fit before a single exchange of fitted updates, trading synchronized fidelity for local autonomy. For gradient aggregation, we develop a finite-precision protocol that transmits changes rather than values and certifies targeting accuracy within explicit bounds on exchanges and bits. A description-length analysis of the accepted updates then shows that this finite communication leaves numerical targeting error negligible against sampling uncertainty. The cost of computing an estimator is thus distinct from the complexity of selecting it. Our analysis also indicates that keeping data local is not itself a privacy guarantee of TMLE, since instability of full-record reconstruction need not prevent recovery of a specified sensitive attribute. For a personalized version of local averaging, institutions retain their own estimates and leave once local targeting is complete. A nonconvex convergence bound charges the improvement forfeited through averaging to disagreement among local fits and exposes a tradeoff between equal institutional influence and the sampling variability of small silos.
Chinese Translation
科学或运营决策背后的证据往往掌握在医院、银行或登记机构手中,而这些机构无法汇集个体观测数据。跨机构(cross-silo)联邦学习将计算迁移至数据所在地,并交换约定的摘要统计量。目标最大似然估计(Targeted Maximum Likelihood Estimation, TMLE)在灵活的初始拟合基础上进行精细化更新,得到尊重模型假设并支持高效统计推断的代入(plug-in)估计量。然而,TMLE本身至今仍是一个完全集中式的流程。为填补这一空白,本文提出了首个联邦TMLE算法。我们针对任意目标参数、损失函数和波动(fluctuation)族,通过两个互补框架实现了对目标化(targeting)步骤本身的联邦化。FedTMLE-G聚合各节点的局部梯度,逐步复现集中式的目标化过程;FedTMLE-L则允许每个机构在仅一次交换拟合更新之前完成各自的波动拟合,以牺牲同步的精确性换取本地自主性。对于梯度聚合,我们设计了一种有限精度协议,该协议传输的是变化量而非数值本身,并在显式的交换次数和比特数上界内保证目标化的精度。对被接受更新的描述长度(description-length)分析表明,这种有限的通信所带来的数值目标化误差相对于抽样不确定性可以忽略不计。因此,计算估计量的代价与选择估计量的复杂度是彼此独立的。我们的分析还指出,数据本地化本身并不构成TMLE的隐私保证,因为即使全记录重构是不稳定的,也未必能阻止特定敏感属性的恢复。对于局部平均的个性化版本,各机构保留自己的估计量,并在本地目标化完成后即退出。我们给出了一个非凸收敛界,将平均化所损失的改进归因于各局部拟合之间的分歧,并揭示了机构影响力均等与小规模数据分片抽样变异性之间的权衡。
cs.LG / 36 / 2609.30508

Benchmarking the Connectomes of Caenorhabditis elegans within the Reservoir Computing Framework

在储备池计算框架下对秀丽隐杆线虫连接组进行基准测试
Reimers, Felix S., Ramstad, Ola Huse, Hubin, Aliaksandr, Nichele, Stefano
Abstract
The aim of this work is to examine the connectomes of Caenorhabditis elegans through a computational lens using the reservoir computing framework. Connectomes are mappings of biological neural networks; C. elegans is the first organism for which physical connectomes covering the whole nervous system have been published. The connectomes of C. elegans used in this paper have been derived at different ages of the organism and are based on three different ways of measuring inter-cellular connections. They have, with minimal preprocessing, been implemented as reservoirs in the form of echo state networks, which are recurrent neural networks. In reservoir computing, the reservoir itself is not trained, rather the output of the reservoir is passed to a comparatively small read-out module in which training takes place. Training and testing is conducted in different neuro-inspired tasks, with the aim of using these tasks as a benchmark for the connectomes. This process has been repeated with different configurations of the reservoir and equally sized but randomized null models have been used for comparison. The results show that the biological wiring and a bio-informed configuration of input and output nodes of the reservoirs do not necessarily lead to better performance. Contrarily, the randomized null models are often outperforming the original connectomes on the chosen benchmarks. At the same time it becomes clear that the results depend a lot on the configuration of the reservoir and the way the connectome has been derived from the organism. Connectomes from different ages may produce varying outcome, without a clear trend becoming visible.
Chinese Translation
本研究旨在通过储备池计算(reservoir computing)框架,从计算的视角考察秀丽隐杆线虫(Caenorhabditis elegans)的连接组。连接组是生物神经网络的映射;秀丽隐杆线虫是首个发表了覆盖整个神经系统的物理连接组的生物。本文所使用的秀丽隐杆线虫连接组取自该生物的不同发育阶段,并基于三种不同的细胞间连接测量方式。在经过极少的预处理后,这些连接组被实现为回声状态网络(echo state networks)形式的储备池,后者是一种循环神经网络。在储备池计算中,储备池本身并不被训练,而是将储备池的输出传递给一个规模相对较小的读出模块,训练在该模块中进行。训练和测试在多种受神经科学启发的任务上进行,旨在将这些任务作为连接组的基准测试。该过程在不同配置的储备池下重复进行,并使用规模相同但经随机化的零模型进行对比。结果表明,生物的神经布线方式以及基于生物信息设定的储备池输入和输出节点配置并不一定能带来更好的性能。相反,随机化的零模型在所选基准任务上往往优于原始连接组。同时,结果明显依赖于储备池的配置以及连接组从生物体中获取的方式。来自不同发育阶段的连接组可能产生不同的结果,且未呈现出明显的趋势。
cs.LG / 37 / 2609.30541

AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework

生产规模下的自动化研究(AutoResearch):失效模式与多智能体框架
Chandran, Aparajith, Kim, Juwon, Jha, Saurav, Castells, Pablo, Hottier, Florian
Abstract
Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a training script and retains modifications that improve a held-out scalar metric -- to automate this exploration. We report on twelve weeks of running this paradigm at production scale, where iterations consume hours of multi-GPU compute, evaluation involves competing criteria, and campaigns span weeks across many training jobs. Across two independently developed representation-learning systems for a book recommendation pipeline, we ran 220+ experiments and observed five recurring failure modes absent from the original setting: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation. We contribute a three-principle scaffolding design -- prevent, persist, redirect -- that maps each failure mode to a structural remedy and whose instantiation scales with iteration cost. The framework produced a 1.82x Recall@6 lift and a 2.1x coherence lift over hand-tuned baselines, and the agent autonomously designed a text-only fallback that expanded catalog coverage by 5.8x. The two systems span nearly three orders of magnitude in per-iteration cost yet exhibit the same failure modes, suggesting these are structural properties of production-scale autonomous research rather than artifacts of either application.
Chinese Translation
面向生产级推荐流水线优化嵌入系统,需要进行系统性的探索,而这种探索在大规模场景下会消耗不成比例的工程投入。我们应用 Andrej Karpathy 提出的 AutoResearch 范式——即由大语言模型迭代地修改训练脚本,并保留能够提升留出集标量指标的修改——来自动化这一探索过程。我们报告了在生产规模下运行该范式十二周的经验,其中每次迭代需要消耗数小时的多GPU算力,评估涉及相互竞争的多项准则,而整个研究周期需跨越数周并涵盖多个训练任务。在两个独立开发的、面向图书推荐流水线的表征学习系统中,我们运行了220余次实验,并观察到在原始设定中不存在的五种反复出现的失效模式:基础设施脆弱性、智能体记忆衰退、搜索方向停滞、迭代成本不对称性以及指标固着。我们提出了一种三原则的脚手架设计——预防(prevent)、持久化(persist)与重定向(redirect)——将每种失效模式映射到一种结构性补救方案,且其具体实现的复杂度随迭代成本而扩展。该框架相比人工调优的基线实现了1.82倍的 Recall@6 提升和2.1倍的连贯性提升,并且智能体自主设计了一种纯文本回退方案,将目录覆盖率扩大了5.8倍。这两个系统在单次迭代成本上相差近三个数量级,却呈现出相同的失效模式,这表明这些失效模式是生产规模自主研究的结构性特征,而非某一具体应用的偶然现象。
cs.LG / 38 / 2609.30542

GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for \textit{De Novo} Peptide Sequencing

GyroNovo:基于误差引导的片段填补与质量感知注意力机制的从头肽段测序方法
Mekki, Abdellah El, Lakshmanan, Laks V. S., Abdul-Mageed, Muhammad
Abstract
De novo peptide sequencing from tandem mass spectra is essential for identifying peptides without relying on reference databases. Despite advances in deep learning, accurate sequencing remains challenging because experimental spectra are often sparse, noisy, and incomplete, leaving informative b- and y-ion fragments unobserved. Existing methods attempt to recover this missing evidence via latent-space imputation before autoregressive decoding. However, they typically treat imputation as a fixed reconstruction task, without considering which missing fragments are most relevant to decoder errors. Moreover, existing peak representations do not explicitly model mass differences between peaks, despite their fundamental importance. We introduce GyroNovo, a framework with two main contributions. First, we use decoder errors observed during training to adapt the imputation objective, prioritizing fragments associated with frequent decoding errors. We further use the decoder error distribution to construct easy and hard augmented views of each spectrum, enabling the decoder to learn under varying degrees of spectral corruption and missing-fragment severity. Second, we introduce a mass-aware inductive bias into self-attention by using rotary embeddings to encode pairwise mass differences between spectral peaks. Together, these components align missing-fragment recovery with decoder behavior while explicitly incorporating the mass relationships that underlie peptide fragmentation. At inference time, GyroNovo retains a standard encoder-imputer-decoder architecture and requires neither additional inputs nor auxiliary search procedures. Experiments on NovoBench show gains of about 9 percentage points in peptide-level precision and 7 percentage points in amino-acid-level precision over the state-of-the-art baseline. Code: https://github.com/UBC-NLP/gyronovo.
Chinese Translation
基于串联质谱的从头肽段测序对于不依赖参考数据库鉴定肽段至关重要。尽管深度学习取得了进展,准确的测序仍然具有挑战性,因为实验谱图通常稀疏、含噪且不完整,导致信息丰富的b离子和y离子片段未被观测到。现有方法尝试在自回归解码之前通过潜空间填补来恢复这些缺失的证据。然而,它们通常将填补视为固定的重建任务,而没有考虑哪些缺失片段与解码器误差最相关。此外,现有的峰表示方法没有显式建模峰之间的质量差异,尽管质量差异具有根本性的重要意义。我们提出了GyroNovo,一个包含两大主要贡献的框架。首先,我们利用训练期间观察到的解码器误差来自适应调整填补目标,优先填补与频繁解码错误相关的片段。我们进一步利用解码器误差分布为每张谱图构建简单和困难两种增强视图,使解码器能够在不同程度的谱图损坏和缺失片段严重性下进行学习。其次,我们通过使用旋转位置编码(rotary embeddings)对谱峰之间的成对质量差异进行编码,将质量感知的归纳偏置引入自注意力机制。这些组件共同将缺失片段的恢复与解码器行为对齐,同时显式纳入肽段碎裂 underlying 的质量关系。在推理时,GyroNovo保留了标准的编码器-填补器-解码器架构,既不需要额外输入,也不需要辅助搜索过程。在NovoBench上的实验表明,相比最先进的基线方法,肽段级精度提升约9个百分点,氨基酸级精度提升约7个百分点。代码:https://github.com/UBC-NLP/gyronovo。
cs.LG / 39 / 2609.30556

Dynamic Regret in Online Convex Optimization with Indicator Switching Costs

具有指示器切换代价的在线凸优化中的动态遗憾
Mhaisen, Naram, Iosifidis, George
Abstract
We study dynamic regret in online convex optimization with an \emph{indicator switching cost}: a fixed penalty incurred whenever two consecutive decisions differ. This captures startup overheads such as server activation, model deployment, and cache updates, and on a bounded domain it recovers norm-based movement costs as a special case. Existing guarantees for indicator costs handle only static comparators. We show that a direct extension of these techniques to dynamic regret provably fails, motivating a different approach. We propose a meta-learning framework: a set of randomized lazy FTRL base learners restarted at dyadic time scales, aggregated by a movement-aware master that mixes their proposal densities and samples actions via maximal coupling of consecutive mixtures. The resulting algorithm satisfies, in expectation, $\mathcal{R}^{\mathbf{1}}_T \le \tilde{\mathcal{O}}(\min\{\sqrt{T(S_T{+}1)},T^{2/3}(P_T+1)^{1/3}\})$, where $\mathcal{R}^{\mathbf{1}}_T$ is the dynamic regret plus the cumulative indicator switching cost, $S_T$ counts comparator switches, and $P_T$ is the comparator path length. The bound holds simultaneously for all sequences and requires no prior knowledge of $S_T$ or $P_T$: it is minimax-optimal (up to logarithmic factors) for tracking piecewise-constant comparators, and also captures frequently moving comparators with small total path length.
Chinese Translation
我们研究了具有*指示器切换代价*(indicator switching cost)的在线凸优化中的动态遗憾:即当两个连续决策不同时所产生的固定惩罚。这种代价刻画了诸如服务器启动、模型部署和缓存更新等启动开销,并且在有界域上,它将基于范数的移动代价作为特例涵盖其中。现有的关于指示器代价的理论保证仅适用于静态比较器。我们证明将这些技术直接扩展到动态遗憾会不可避免地失效,从而促使我们采用不同的方法。我们提出了一个元学习框架:一组随机化惰性 FTRL 基学习器在二倍频时间尺度上重启,并由一个具备移动感知的主聚合器进行聚合,该聚合器混合各基学习器的提议密度,并通过连续混合分布的极大耦合来采样动作。所得算法在期望意义下满足 $\mathcal{R}^{\mathbf{1}}_T \le \tilde{\mathcal{O}}(\min\{\sqrt{T(S_T{+}1)},T^{2/3}(P_T+1)^{1/3}\})$,其中 $\mathcal{R}^{\mathbf{1}}_T$ 为动态遗憾加上累积的指示器切换代价,$S_T$ 统计比较器的切换次数,$P_T$ 为比较器的路径长度。该界对所有序列同时成立,且无需对 $S_T$ 或 $P_T$ 的任何先验知识:对于追踪分段常数比较器的情形,该界是极小极大最优的(相差对数因子);同时它也涵盖了总路径长度较小但频繁移动的比较器。
cs.LG / 40 / 2609.30572

Entropy Regularization: A Free Correction to Cross-Entropy for Verified Demonstrations

熵正则化:一种面向已验证演示的交叉熵免费修正方法
Dhanakshirur, Mihir, Ousherovitch, Adam, Tewari, Ambuj
Abstract
Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable domains with multiple correct solutions, such as mathematical reasoning and code generation, where training data may contain only one expert solution per problem. We show that minimizing cross-entropy can be misaligned with minimizing verifier risk; two policies can assign identical likelihood to the observed demonstrations while placing different probability mass on incorrect outputs. This is formalized through a learning-theoretic counterexample in which CE minimization selects a suboptimal policy. We identify that controlling the support of the learned policy can solve this problem by preventing probability mass from spreading to unsupported outputs. Since support size is non-differentiable and computationally intractable, we propose entropy-regularized cross-entropy (ER-CE), using token-level Shannon entropy as a tractable proxy. Finally, across mathematical reasoning and code-generation benchmarks, we find that entropy-regularized training consistently improves verifier accuracy over standard cross-entropy. Our results identify a simple failure mode of imitation-based post-training in verifiable tasks and provide a practical objective that is better aligned with producing correct outputs.
Chinese Translation
大语言模型通常使用交叉熵(Cross-Entropy, CE)在专家演示数据上进行后训练,即使下游目标并非模仿所演示的解法,而是生成任何能通过验证器接受的输出。在存在多个正确解的可验证领域(如数学推理和代码生成)中,这种不匹配尤为明显,因为训练数据中每个问题可能只包含一个专家解法。我们证明,最小化交叉熵可能与最小化验证器风险不一致:两个策略可以对观测到的演示赋予相同的似然,同时将不同的概率质量分配到错误输出上。我们通过一个学习理论反例将这一点形式化,其中交叉熵最小化会选择次优策略。我们发现,控制所学策略的支持集(support)可以通过阻止概率质量扩散到不支持输出上来解决该问题。由于支持集大小不可微且计算上难以处理,我们提出熵正则化交叉熵(Entropy-Regularized Cross-Entropy, ER-CE),使用token级别的香农熵(Shannon entropy)作为可处理的代理。最后,在数学推理和代码生成基准测试中,我们发现熵正则化训练相比标准交叉熵能够持续提升验证器准确率。我们的结果揭示了可验证任务中基于模仿的后训练的一种简单失效模式,并提供了一个与生成正确输出更好对齐的实用目标函数。
cs.LG / 41 / 2609.30578

Reinforcement Learning of Communication in a Mesh of Small Language Models

小型语言模型网状网络中通信的强化学习
Turkcan, Mehmet Kerem
Abstract
Language models gain accuracy from more compute at test time, but majority voting over independent samples saturates: as samples grow, the vote converges to the model's most frequent answer. Communication can add what sampling cannot: an agent that solves a problem can pass the key step to the others. We present TalkMesh, a decentralized mesh of small language model agents that learns when and what to communicate. Each agent samples a proposal and scores it with a trained confidence head. The most confident agent broadcasts a hint; agents below a confidence threshold revise, keeping each revision that outscores its proposal. Gossip consensus approximates the vote weighted by confidence without a coordinator. A talk policy, trained with group relative policy optimization on the change in correctness after revision, writes hints and revisions. With three agents, which together generate at most six outputs, the mesh reaches the accuracy of majority voting over 32 samples with each of three models. Trained with at most 8 agents and evaluated with 32, it raises accuracy from 0.568 under self-consistency to 0.705 (Qwen3.5-0.8B, GSM8K) and from 0.492 to 0.722 (SmolLM3-3B, MATH-500). When 4 of 8 agents collude on a wrong answer with fabricated confidence and poisoned hints, majority vote accuracy falls to 0.000 (Qwen3.5-0.8B, GSM8K). A defended mesh, whose agents rescore solutions with their own confidence heads, retains 0.507. Across reasoning, embodied coordination, and traffic signal control, messages improve a decision when the acting agent cannot observe the information it requires and another agent can send it.
Chinese Translation
语言模型通过在测试时使用更多计算来提升准确性,但对独立样本进行多数投票会趋于饱和:随着样本数量增加,投票结果会收敛到模型最频繁给出的答案。通信可以提供采样无法实现的功能:已经解决某个问题的智能体可以将关键步骤传递给其他智能体。我们提出了TalkMesh,这是一个由小型语言模型智能体组成的去中心化网状网络,它学习何时通信以及通信什么内容。每个智能体生成一个候选答案,并用训练好的置信度头为其打分。置信度最高的智能体广播一条提示;置信度低于阈值的智能体据此修改,并保留每个得分超过其原有候选答案的修改。八卦共识(gossip consensus)在没有协调者的情况下近似实现按置信度加权的投票。一个谈话策略(talk policy)通过组相对策略优化(GRPO)在修改后正确性变化上进行训练,负责撰写提示和修改内容。使用三个智能体(总共最多生成六个输出),该网状网络即可达到对32个样本进行多数投票的准确性(分别对三个模型而言)。该系统最多用8个智能体训练,在32个智能体下评估,将准确率从自洽性(self-consistency)的0.568提升至0.705(Qwen3.5-0.8B,GSM8K),从0.492提升至0.722(SmolLM3-3B,MATH-500)。当8个智能体中有4个合谋给出错误答案、伪造置信度并投毒提示时,多数投票的准确率降至0.000(Qwen3.5-0.8B,GSM8K)。而一个具备防御能力的网状网络——其智能体用自己的置信度头对解答重新打分——仍能保持0.507的准确率。在推理、具身协调和交通信号控制等任务中,当行动中的智能体无法观察到所需信息而另一个智能体可以发送该信息时,消息传递能够改善决策。
cs.LG / 42 / 2609.30580

Energy-efficient operation of neural operators for virtual sensing

面向虚拟感知的神经算子节能运行
Yoo, Jason, Roy, Samrendra, Chakraborty, Souvik, Alam, Syed Bahauddin
Abstract
Virtual sensing repeatedly reconstructs physical fields from changing observations, often on a fixed geometry. We investigate how shared spatial computation reduces the energy of these updates while retaining the selected checkpoint and its evaluated predictions. In a heat-exchanger service, standard compiler freezing and explicit trunk reuse give similar operating energy reductions relative to graph replay: approximately 1% at one request per second and 20% at forty requests per second. In 15 W mode with fixed clocks, reuse with graph replay completes the same request sequence with 22.0 to 22.5% less energy than eager execution, including preparation and waiting. DeepONet and Fourier neural operator (FNO) controls distinguish the effects of reusable arithmetic and launch overhead. Preparation, artifact construction, and worker replacement add costs outside repeated inference. These results connect operator structure to operating energy and show how update frequency and execution lifetime govern the benefit of computation reuse in physical-field virtual sensing.
Chinese Translation
虚拟感知需要基于不断变化的观测数据反复重建物理场,且通常在固定几何域上进行。我们研究了共享空间计算如何在保留所选检查点及其评估预测的同时,降低这些更新过程的能耗。在一个换热器服务案例中,相对于图重放(graph replay),标准编译器冻结与显式主干(trunk)重用可实现相近的运行能耗降低:在每秒一次请求时约降低1%,在每秒四十次请求时约降低20%。在固定时钟频率的15瓦模式下,与急切执行(eager execution)相比,结合图重放的重用方式完成相同请求序列的能耗降低22.0%至22.5%(含准备与等待开销)。DeepONet与傅里叶神经算子(FNO)的对照实验区分了可重用算术运算与启动开销的影响。准备阶段、构建产物以及工作进程替换会在重复推理之外产生额外成本。这些结果将算子结构与运行能耗联系起来,并表明更新频率与执行生命周期如何决定计算重用在物理场虚拟感知中的收益。
cs.LG / 43 / 2609.30592

QSV: Quat-Sphere-Vision for Coupled Quaternion Attention on Spherical Lattices

QSV:基于球面格点上耦合四元数注意力的四元数-球面-视觉模型
Foley, Nicholas, Marinelli, Devin, Moore, Donny, Enriquez, Diego, Fernandez, Amanda
Abstract
In standard attention, three separately learned projections decide how strongly a token attends to each neighbor ($W_Q$, $W_K$) and how the attended features are transformed before aggregation ($W_V$). We study Quat-Sphere-Vision (QSV), a sparse spherical vision model that replaces this projection triple with a single learned unit quaternion per token: the relative quaternion $r_{ij} = q_i^{*} \otimes q_j$ supplies both the attention logit $\operatorname{Re}(r_{ij})$ and a sandwich-product feature transport $x \mapsto r_{ij} \otimes x \otimes r_{ij}^{*}$, with messages passed over sparse kNN graphs on concentric Fibonacci spheres. Ablations that change only the targeted component show the two roles to be asymmetric. Removing the transport reduces test accuracy by about four percentage points on CIFAR-10 and CIFAR-100 (single runs per CIFAR-100 variant), while replacing the learned attention weights with uniform averaging leaves it essentially unchanged. Parameter-matched controls then remove the geometry itself: standard attention on the same graph exceeds QSV (mean $87.3\%$ vs. $85.9\%$), and the same model on a flat 2D lattice reaches $91.1\%$, within $2.1$ points of a ResNet-20 trained under the same pipeline (single run). In the coupled kernel, nearly all of the learned pairwise computation resides in the transport channel.
Chinese Translation
在标准注意力机制中,三个独立学习的投影(W_Q、W_K)决定一个词元对每个邻居的关注强度,而(W_V)决定被关注的特征在聚合前如何被变换。我们研究了四元数-球面-视觉模型(Quat-Sphere-Vision, QSV),这是一种稀疏球面视觉模型,它用一个针对每个词元学习的单位四元数取代上述投影三元组:相对四元数 r_ij = q_i^{*} ⊗ q_j 同时提供注意力得分 Re(r_ij) 和夹乘式特征传输 x ↦ r_ij ⊗ x ⊗ r_ij^{*},消息在同心斐波那契球面上的稀疏 kNN 图中传递。仅改变目标组件的消融实验表明,这两种角色是不对称的。移除特征传输会使 CIFAR-10 和 CIFAR-100 上的测试准确率下降约四个百分点(每种 CIFAR-100 变体各运行一次),而将学习到的注意力权重替换为均匀平均则几乎不产生变化。随后进行的参数匹配对照实验进一步排除了几何结构本身的作用:在同一图上使用标准注意力超越了 QSV(平均 87.3% 对 85.9%),而同一模型在平坦二维格点上达到 91.1%,与在同一训练流程下训练的 ResNet-20 仅相差 2.1 个百分点(单次运行)。在耦合核中,几乎所有学习到的成对计算都集中在传输通道上。
cs.LG / 44 / 2609.30605

Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System

面向基于深度强化学习的入侵检测系统的具有可解释性的概率鲁棒性驱动的通用对抗扰动
Zhang, Hongsen, Zhang, Lu, Xu, Mingjing, Zhang, Yi, Epiphaniou, Gregory, Maple, Carsten
Abstract
Deep reinforcement learning (DRL) enables adaptive intrusion detection in dynamic network environments but also exposes intrusion detection systems (IDS) to adversarial threats such as universal adversarial perturbations (UAPs), which apply a single input-agnostic perturbation to degrade detection performance across traffic. Probabilistic Robustness (PR), as a post-hoc evaluation metric, provides a principled, population-level measure of adversarial impact that conceptually aligns with the universality objective of UAPs, i.e., PR quantifies the prevalence of misclassification in the input space, making it a natural signal for guiding UAP generation. Hence, we propose PR-based UAP, which represents the first integration of an explicit PR-driven objective into generating UAPs against DRL-based IDS. Building on this formulation, we introduce PX-UAP, which leverages explainable artificial intelligence (XAI) to guide perturbation shaping under realistic domain constraints, and provides a rigorous theoretical analysis of its design. Extensive experiments demonstrate that PX-UAP consistently outperforms state-of-the-art UAP methods in attack effectiveness.
Chinese Translation
深度强化学习(DRL)使入侵检测系统能够在动态网络环境中进行自适应检测,但同时也使入侵检测系统(IDS)面临诸如通用对抗扰动(UAP)等对抗性威胁。UAP通过施加一种与输入无关的单一扰动,降低系统对各类流量的检测性能。概率鲁棒性(Probabilistic Robustness, PR)作为一种事后评估指标,能够以原则化方式在群体层面度量对抗影响,其概念上与UAP的通用性目标相契合,即PR量化了输入空间中误分类的普遍程度,使其成为指导UAP生成的天然信号。因此,我们提出了基于PR的UAP方法(PR-based UAP),这是首次将显式的PR驱动目标融入针对基于DRL的IDS的UAP生成中。在此基础上,我们进一步提出了PX-UAP,该方法利用可解释人工智能(XAI)在现实域约束下引导扰动的构造,并对其设计提供了严格的理论分析。大量实验表明,PX-UAP在攻击有效性方面始终优于当前最先进的UAP方法。
cs.LG / 45 / 2609.30628

OpenHail: An Event-Driven Gymnasium Environment for Electric Ride-Hailing Fleet Control

OpenHail:用于电动网约车车队控制的事件驱动 Gymnasium 环境
Schettini, Tommaso, Kullman, Nicholas D., Mendoza, Jorge E.
Abstract
Machine-learning policies have attracted increasing interest for ride-hailing fleet control in recent years. Reinforcement learning, in particular, requires a structured simulation environment that specifies observations, actions, rewards, and decision epochs for training and evaluation. For electric fleets, this environment must also capture the interaction among stochastic demand, vehicle operations, and capacitated charging infrastructure. We present OpenHail, an open-source Gymnasium environment for joint control of electric ride-hailing fleets. Its fixed-size observation--action interface exposes request assignment, repositioning, and charging to a single policy. The event-driven simulator represents requests with pickup deadlines, vehicle job queues, battery dynamics, and finite-capacity charging facilities with first-in--first-out queues. A configurable decision-epoch mechanism separates internal simulator events from policy interactions, supporting event-driven, periodic, hybrid, and policy-requested control within the same operational model. The software provides seeded instances, feasible-action utilities, evaluation tools, operational metrics, and baseline policies. The source code is available at https://github.com/tommaso-schettini/openhail.
Chinese Translation
近年来,机器学习策略在网约车车队控制领域引起了越来越多的关注。尤其是强化学习,需要一个结构化的仿真环境,明确规定观测、动作、奖励和决策时刻,以便进行训练与评估。对于电动车队而言,该环境还必须能够刻画随机需求、车辆运行与容量受限的充电基础设施之间的交互。我们提出了 OpenHail,一个用于电动网约车车队联合控制的开源 Gymnasium 环境。其固定尺寸的观测—动作接口将订单分配、车辆调度(再定位)与充电暴露给单一策略。该事件驱动仿真器刻画了带有取车截止时间的订单、车辆作业队列、电池动态,以及采用先进先出队列的有限容量充电设施。可配置的决策时刻机制将仿真器内部事件与策略交互相分离,从而在同一运行模型内支持事件驱动、周期性、混合以及策略主动请求的控制方式。该软件提供带随机种子的实例、可行动作工具、评估工具、运营指标和基线策略。源代码可在 https://github.com/tommaso-schettini/openhail 获取。
cs.LG / 46 / 2609.30633

Stable initialization without the CLT

无需中心极限定理的稳定初始化方法
Kuang, Simon, Chickering, Kyle, Lin, Xinfan
Abstract
Successful training of deep neural networks is highly dependent on the distribution of the initial weights. If the weights are too large, network training blows up; if they are too small, the model fails to learn features. Stable initialization is the optimal moderation between these two extremes. The conventional theory of random networks uses the Central Limit Theorem to control inter-neuron dependencies, which introduces distributional approximation error and coupling between layers. For networks with sine activations, we derive the uniform-phase initialization, which obviates distributional approximation and fully decouples the layers. Ours is the first work to use the sine function's periodic symmetry. Models trained with the uniform-phase initialization outperform the state of the art in neural representation tasks like image and audio fitting. We find that our untuned models are competitive with the best-tuned baselines from previous work and support $\mu$P width scaling.
Chinese Translation
深度神经网络能否成功训练在很大程度上取决于初始权重的分布。若权重过大,网络训练会发散;若权重过小,模型则无法学习到特征。稳定初始化正是这两个极端之间的最优折中。传统的随机网络理论使用中心极限定理(Central Limit Theorem)来控制神经元之间的依赖关系,这会引入分布近似误差以及层与层之间的耦合。对于采用正弦激活函数的网络,我们推导出了均匀相位初始化(uniform-phase initialization)方法,该方法消除了分布近似误差,并完全解耦了各层。我们是首个利用正弦函数周期对称性的工作。采用均匀相位初始化训练的模型在图像拟合、音频拟合等神经表示任务中超越了当前最先进的方法。我们发现,未经调参的模型即可与以往工作中经过最优调参的基线方法相媲美,并且支持 $\mu$P 宽度缩放。
cs.LG / 47 / 2609.30634

In-Context Binding Capacity in Language Models

语言模型中的上下文绑定容量
Ravulapalli, Manas Venkata Sai, Chadha, Samrath Singh
Abstract
How many assignments can a language model recall before it loses track of which value belongs to which entity? We measure this limit using continuous recall curves for 12 models at or below 3B parameters and a threshold sweep over 30 open models up to 12B. On the continuous curves, the load at which recall falls halfway to chance follows $K_{50}=cN^{\alpha}$, with $\alpha=0.820$ and $R^2=0.73$. The broader sweep shows an eightfold range associated with pretraining recipe, although the continuous curves show no detectable recipe effect after controlling for scale, with few modern models in the fit. We derive why interference can lower measured capacity by reducing single-binding recall even when the load-dependent recall profile is unchanged. Direct task training also exceeds the extrapolated zero-shot law, but different measurement criteria prevent interpreting that comparison as a capacity gain. Its formation times follow a power-law form in two independent codebases, conditional on runs that succeed. Together, these results characterize capacity at the model's query interface. Bounds on joint recall and a decomposition of policy errors connect this measurement to working memory and instruction following, without treating recall as a measure of alignment. The controlled task also provides a baseline for testing whether binding limits constrain world-state tracking; the present experiments do not measure state updates or downstream transfer.
Chinese Translation
一个语言模型在无法区分哪个值属于哪个实体之前,能记住多少个赋值关系?我们通过连续召回曲线对12个参数量不超过30亿的模型,以及对30个最高达120亿参数的开源模型的阈值扫描,测量了这一极限。在连续曲线上,召回率降至随机水平一半时所对应的负载遵循 $K_{50}=cN^{\alpha}$,其中 $\alpha=0.820$,$R^2=0.73$。更广泛的阈值扫描显示,与预训练配方相关的容量差异可达八倍,但在控制规模之后,连续曲线未显示出可检测的配方效应,且拟合中仅包含少数现代模型。我们推导了干扰为何即使在负载依赖的召回曲线保持不变的情况下,也能通过降低单绑定召回率而降低测量到的容量。直接任务训练的表现也超出了外推的零样本规律,但由于测量标准不同,不能将这一比较解读为容量的提升。在两个独立的代码库中,在训练成功的前提下,其形成时间遵循幂律形式。这些结果共同刻画了模型查询接口处的容量。关于联合召回的界限以及对策略误差的分解,将这一测量与工作记忆和指令遵循联系起来,而不会将召回视为对齐程度的度量。该受控任务还为检验绑定限制是否约束世界状态跟踪提供了基线;本实验并未测量状态更新或下游迁移。
cs.LG / 48 / 2609.30650

Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation

交互式智能体中的因果保持:接口分解与选择性适应
Zhang, Shengjun, Liu, Tingyi, Xie, Dong, Dong, Yunlong, Wang, Xiang, Zeng, Cheng
Abstract
Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context, direct target, value, and delay. For finite structural causal model classes, the optimal probe error is a Bayes decision risk. It vanishes exactly when every learning-interface fiber lies within one probe-answer fiber; any state obtained by post-processing that interface inherits the same lower bound. A posterior-coverage theorem characterizes budgeted retesting, while an exact edit decomposition shows that the shifted set is the unique support of an error-free target update. Causal Core implements these conditions through evidence-gated writing, readout filtering, temporal credit, hidden-context setup, and local diagnostic updates. Experiments cover finite causal systems, continuous simulators, an official TD-MPC2 world model, and Qwen2.5-7B-Instruct. A frozen Qwen last-layer probe reaches 0.958 balanced accuracy on source mechanisms but 0.583 on changed delays; the gated mechanism state reaches 1.000 and accepts only 0.056 of synchronized-readout candidates. In TD-MPC2, five target states per actuator recover effect-sign accuracy from 0.057 to 0.948 without degrading stable responses. Causal retention is therefore distinct from task sufficiency and source-domain decodability.
Chinese Translation
任务性能并不决定智能体保留哪种干预机制。我们研究因果保持(causal retention)问题:即一个冻结的学习状态能否回答一个与训练无关地固定的机制探针映射,该映射涵盖动作、上下文、直接目标、价值和延迟等方面。对于有限的结构因果模型类,最优探针误差是一个贝叶斯决策风险。当且仅当每个学习接口纤维都位于某个探针回答纤维之内时,该误差才恰好为零;任何通过对该接口进行后处理得到的状态均继承相同的下界。后验覆盖定理刻画了预算约束下的重测试,而一个精确的编辑分解表明,偏移集是无误差目标更新的唯一支撑。Causal Core 通过证据门控写入、读出过滤、时间信用分配、隐藏上下文设置以及局部诊断更新来实现这些条件。实验涵盖有限因果系统、连续模拟器、官方 TD-MPC2 世界模型以及 Qwen2.5-7B-Instruct。冻结的 Qwen 最后一层探针在源机制上达到 0.958 的平衡准确率,但在变化的延迟上仅为 0.583;而门控机制状态达到 1.000,且仅接受 0.056 的同步读出候选。在 TD-MPC2 中,每个执行器五个目标状态即可将效应符号准确率从 0.057 恢复到 0.948,同时不损害稳定响应。因此,因果保持不同于任务充分性与源域可解码性。
cs.LG / 49 / 2609.30661

Population loss in shallow ReLU networks: Bias & families of critical points

浅层ReLU网络中的总体损失:偏置与临界点族
Field, Michael
Abstract
The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The necessary theory of the T-function is given and a high precision coding using MPFR for the T-function, based on an algorithm of Komelj (2023), is available on request. It is shown that various families of spurious minima described in past papers of Arjevani and the author extend to biased networks and that the loss is always strictly decreased when bias is added. The change in landscape geometry caused by adding bias appears to be relatively mild. Only the simplest examples are described in this paper where it is assumed that the number of inputs is equal to the number of neurons (this restriction is for reasons of length). A review of relevant previous results on unbiased networks is included. Aside from Gaussian statistics, the main mathematical tools and ideas come from analytic geometry (analytic and subanalytic sets, the Curve Selection Lemma).
Chinese Translation
本文的主要结果是学生-教师核模型中总体损失(population loss)的一个公式,该公式适用于带偏置的浅层ReLU网络。这一工作推广了Choo和Saul(2009)以及Brutzkus和Globerson(2017)的先前研究成果。该公式的推导本质上使用了Owen T函数。文中给出了T函数的必要理论,并且基于Komelj(2023)的算法,使用MPFR对T函数进行了高精度编码实现,可按需提供。本文证明了Arjevani及其作者在以往论文中描述的多个伪极小值(spurious minima)族可推广至带偏置的网络,并且加入偏置后损失总是严格降低。加入偏置所引起的损失曲面几何形态变化相对温和。由于篇幅原因,本文仅描述了最简单的例子,其中假设输入数量等于神经元数量。文中还回顾了关于无偏网络的相关先前结果。除高斯统计外,主要数学工具和思想来自解析几何(解析集与次解析集、曲线选取引理/Curve Selection Lemma)。
cs.LG / 50 / 2609.30684

PixSim: a calibrated open-source simulator of instant-payment fraud, recovery and interdiction under analyst capacity constraints

PixSim:一个在分析师容量约束下校准的开源即时支付欺诈、追回与拦截模拟器
Zeimarani, Bashir, Khatib, Alireza, Mousavinasr, Somayeh, Figueiredo, Carlos Maurício Serodio
Abstract
Brazil's Pix settles about 5.9 billion instant, irreversible transfers a month. A fraudulent transfer can be recovered only while the funds remain in a traceable account, and in 2025 the Central Bank's recovery mechanism (MED) returned 9% of accepted contested value. Interdiction therefore has to happen before settlement, by routing each transaction to pass, human review or block, under a finite analyst team and a regulatory hold window. To our knowledge no public simulator jointly models irreversible settlement, a regulated recovery mechanism, downstream fund dispersal and capacity-constrained review. We present PixSim, an open-source simulator of the Pix rail with these elements, calibrated to Banco Central do Brasil open data, with every parameter sourced, calibrated to one published observable, or registered as an assumption. With the model frozen, full-scale runs reproduce the 2025 recovery rate within 0.006 and its decomposition within 0.02; the February-April 2026 window is reported as a misfit and the May 2026 tracing regime as a projection. On a benchmark with a payer-side scorer, four reference policies and ten scenarios, within the simulated mule model: recovery after settlement is constrained by dispersal speed; staffing by the arrival profile cuts a fixed rule's alert expiry from 52% to 2% at constant hours; halving the team removes a fixed threshold-and-block rule's advantage over a queue-aware rule, on loss and on loss plus false-block harm (+0.106 of victim value, positive on all twenty paired seeds), while a reversal at two thirds of the team was not confirmed on independent seeds; and a synthetic scorer of held-out AUC 0.82 cuts lost value by about a quarter. Code and data: https://doi.org/10.5281/zenodo.22948895
Chinese Translation
巴西的Pix系统每月结算约59亿笔即时且不可撤销的转账。欺诈性转账只有在资金仍停留在可追踪账户中时才可能被追回,而2025年巴西央行的追回机制(MED)仅返还了受理争议金额的9%。因此,拦截必须在结算前完成,即在分析师团队人数有限且存在监管暂扣窗口的条件下,将每笔交易路由为放行、人工审核或拦截。据我们所知,目前尚无公开的模拟器能够联合建模不可逆结算、受监管的追回机制、下游资金分散以及容量受限的审核。我们提出PixSim,一个包含上述要素的Pix支付通道开源模拟器,其依据巴西央行开放数据进行校准,所有参数均有出处、按已发表的可观测量校准,或登记为假设。在模型冻结后,全规模运行复现了2025年的追回率,误差在0.006以内,其分解结构误差在0.02以内;2026年2月至4月的数据窗口被报告为拟合不佳,2026年5月的追踪机制则作为预测。在一个包含付款方评分器、四种参考策略和十个场景的基准测试中(在模拟的骡子账户模型内):结算后的追回收资金分散速度的限制;按到达分布配置人力可将固定规则的告警过期率从52%降至2%,且总工时不变;团队规模减半会消除固定阈值拦截规则相对于队列感知规则的优势(在损失和损失加误拦伤害上均为+0.106受害者价值,且在全部二十组配对随机种子上均为正),而团队规模缩减至三分之二的逆转结果未能在独立种子上得到验证;一个留出AUC为0.82的合成评分器可将损失价值降低约四分之一。代码与数据:https://doi.org/10.5281/zenodo.22948895
cs.LG / 51 / 2609.30692

LUMO (Lightweight Unified Multilingual Orchestrator): A Privacy Preserving Offline Voice Assistant

LUMO(轻量级统一多语言编排器):一种保护隐私的离线语音助手
Naeem, Md. Mehedi Hasan, Ruma, Mst. Kamrunnahar, Anjum, Nafiza, Sultana, Shakila, Ali, Md. Sujan
Abstract
Reliable voice interaction is essential in environments with limited internet connectivity and strong privacy. However, most existing voice assistants depend on cloud-based services, which leads to latency issues, dependency on internet access, and privacy vulnerabilities. This research presents LUMO (Lightweight Unified Multilingual Orchestrator), a privacy preserving offline voice assistant designed for edge computing environments. This system integrates local Automatic Speech Recognition (ASR), locally deployed quantized Large Language Model (LLM), and Text-to-Speech (TTS) synthesis into a fully offline pipeline running on a Raspberry Pi 5 with 8 GB RAM. To enable efficient operation on resource constrained hardware, the language model is compressed using 4-bit GGUF quantization, which reduces memory usage while preserving practical conversational capability. Existing edge based voice assistants Mycroft provides partial offline functionality without a generative LLM, with an approximate latency of ~5 s and power consumption of ~12 W, while Rhasspy supports full offline operation but lacks generative capabilities, with ~3 s latency and ~11 W power usage. In contrast, LUMO achieves a Word Error Rate (WER) of 6.8% for short English utterances in low noise conditions, an end-to-end response latency of 2.0-4.0 s, and a lower peak power consumption of approximately 9.0 W. The system also achieves effective offline recognition for Bangla speech, supporting multilingual accessibility in low resource settings. By operating entirely offline, LUMO provides strong data privacy, reduced need for cloud connectivity, and suitability for privacy sensitive edge execution such as rural healthcare, education, and disaster response scenarios.
Chinese Translation
在网络连接受限且隐私要求严格的环境中,可靠的语音交互至关重要。然而,现有的大多数语音助手依赖于云端服务,这导致了延迟问题、对互联网接入的依赖以及隐私安全漏洞。本研究提出了 LUMO(Lightweight Unified Multilingual Orchestrator,轻量级统一多语言编排器),一种面向边缘计算环境的保护隐私的离线语音助手。该系统将本地自动语音识别(ASR)、本地部署的量化大语言模型(LLM)以及文本转语音(TTS)合成集成到一条完全离线的流水线中,运行于配备 8 GB 内存的 Raspberry Pi 5 上。为了在资源受限的硬件上实现高效运行,该语言模型采用 4-bit GGUF 量化进行压缩,在保持实用对话能力的同时降低了内存占用。现有的边缘语音助手 Mycroft 提供部分离线功能,但不含生成式 LLM,其延迟约为 5 秒,功耗约为 12 瓦;Rhasspy 支持完全离线运行,但缺乏生成能力,延迟约为 3 秒,功耗约为 11 瓦。相比之下,LUMO 在低噪声条件下对简短英文语句的词错误率(WER)为 6.8%,端到端响应延迟为 2.0–4.0 秒,峰值功耗更低,约为 9.0 瓦。该系统还对孟加拉语(Bangla)语音实现了有效的离线识别,支持低资源环境下的多语言可及性。通过完全离线运行,LUMO 提供了强大的数据隐私保护,减少了对云连接的依赖,并适用于隐私敏感的边缘执行场景,如农村医疗、教育和灾难响应等。
cs.LG / 52 / 2609.30718

NEMSim: Learning Control-Conditioned Multi-Event Physical Dynamics via Executable Event-Mechanism Priors

NEMSim:通过可执行的事件-机制先验学习控制条件下的多事件物理动力学
Yu, Junsong, Xie, Junjie, Liu, Pengwei, Ni, Dong
Abstract
High-fidelity simulation of control-conditioned multi-event physical systems is computationally expensive, especially across broad control spaces and long trajectories. In these systems, macroscopic evolution emerges from localized discrete events whose intensities and effects depend on process controls and evolving local states, while the available system knowledge is typically expressed as event-attribute descriptions. Purely data-driven surrogates must infer these event effects from limited trajectory coverage, which can hinder generalization to unseen control regimes. Physics-guided methods instead primarily build on equation-level constraints or differentiable solvers rather than discrete event-rule priors. We therefore propose NEMSim (Neural Event-Mechanism Simulator), which compiles predefined event-attribute descriptions into an executable transition structure linking control-dependent event intensities, prior-guided mechanism attribution, and state-dependent responses. To enable evaluation of control-conditioned multi-event dynamics with explicit system knowledge, we construct a 3D KMC-based benchmark pairing high-fidelity trajectories with explicit event rules, standardized splits, and evaluation protocols. Across three settings, NEMSim reduces Avg. RMSE by 58.9%-81.3% relative to the strongest baseline in each setting. It also remains best in the data-efficiency study with training-data fractions down to 10%. Mechanism analyses further show that these gains arise from executable rule integration rather than prior access or architecture alone.
Chinese Translation
对控制条件下的多事件物理系统进行高保真仿真的计算开销极高,尤其是在宽泛的控制空间和长轨迹场景下。在这类系统中,宏观演化由局部化的离散事件涌现而来,这些事件的强度和效应依赖于过程控制量及不断演化的局部状态,而可获得的系统知识通常以事件-属性描述的形式表达。纯数据驱动的代理模型必须从有限的轨迹覆盖中推断这些事件效应,这可能阻碍其对未见控制域的泛化能力。相比之下,现有物理引导方法主要建立在方程级约束或可微分求解器之上,而非离散事件规则先验。为此,我们提出NEMSim(Neural Event-Mechanism Simulator,神经事件-机制模拟器),它将预定义的事件-属性描述编译为可执行的转移结构,将依赖于控制的事件强度、先验引导的机制归因以及依赖于状态的响应联系起来。为了在具有显式系统知识的条件下评估控制条件下的多事件动力学,我们构建了一个基于三维动力学蒙特卡洛(3D KMC)的基准数据集,其中包含高保真轨迹、显式事件规则、标准化数据划分以及评估协议。在三种实验设置中,NEMSim相对于各设置下最强的基线方法,将平均RMSE降低了58.9%至81.3%。在数据效率研究中,即使训练数据比例降至10%,NEMSim仍保持最优表现。机制分析进一步表明,这些性能提升源于可执行规则的有效集成,而非仅仅依赖先验知识或模型架构本身。
cs.LG / 53 / 2609.30721

When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification

当10,000个窗口不等于10,000次检验:滑动窗口时间序列分类中统计置信度的审计
Shi, Xinze, Zhang, Litian, Shi, Binrui
Abstract
Sliding-window classifiers are often evaluated on thousands of overlapping test windows, even though neighboring predictions share observations and remain nested within recordings and subjects. Subject-disjoint evaluation prevents one form of leakage but does not make those test windows independent. We present a practical audit that maps three claims - performance on observed recordings, future recordings from observed subjects, and unseen subjects - to explicit aggregation rules and established dependence-robust inference. At 75% overlap, controlled simulations give 16.9% Type-I error for IID observed-record inference and 7.2% for session-centered Bartlett-HAC: a substantial improvement with residual miscalibration. Audits of frozen WISDM and HARTH predictions show that nearly fourfold growth in test rows yields only 1.75-1.94-fold variance-equivalent information growth. At that overlap, fixed-record paired Accuracy-difference intervals are 1.22-1.66 times the IID widths; this inflation is not universal at zero overlap. On HARTH, paired Accuracy-difference intervals include zero across three overlap settings, whereas Macro-F1 favors MiniROCKET. Independent recomputation, common-session checks, class-level results, and separately seeded calibration make the audit's scope and limitations inspectable. The resulting workflow distinguishes additional predictions from additional independent evidence.
Chinese Translation
滑动窗口分类器常常在数千个相互重叠的测试窗口上进行评估,尽管相邻的预测共享观测数据,并且始终嵌套于 recordings 和被试之中。基于被试互斥的评估虽然可以防止一种形式的信息泄露,但并不能使这些测试窗口相互独立。我们提出了一种实用的审计方法,将三个层面的主张——在已观测 recordings 上的性能、来自已观测被试的未来 recordings 的性能、以及未见过的被试上的性能——映射到明确的聚合规则和成熟的依赖鲁棒推断方法。在75%重叠率下,受控模拟显示,针对已观测 recordings 的 IID 推断的 I 类错误率为16.9%,而基于 session 为中心的 Bartlett-HAC 方法则为7.2%:这是显著改进但仍存在残余的校准偏差。对冻结的 WISDM 和 HARTH 预测结果的审计表明,测试行数增长近四倍仅带来1.75-1.94倍的方差等效信息量增长。在该重叠率下,基于固定 recordings 的成对 Accuracy 差异置信区间宽度是 IID 宽度的1.22-1.66倍;而这种膨胀在零重叠率下并非普遍存在。在 HARTH 数据集上,成对 Accuracy 差异区间在三种重叠设置下均包含零,而 Macro-F1 则更倾向于 MiniROCKET。通过独立重算、共同 session 检验、类别级结果以及单独随机种子的校准,审计的范围和局限性变得可检验。由此形成的工作流程能够区分'额外的预测'与'额外的独立证据'。
cs.LG / 54 / 2609.30746

Mechanism-Aware Ensemble Conditioning for Data-Limited Emulation of Extreme Events

面向数据受限极端事件模拟的机制感知集合条件化方法
Thiel, Isabella S., Bello-Rivas, Juan, Kevrekidis, Yannis G., Sapsis, Themistoklis P.
Abstract
Extreme events in chaotic systems are difficult to learn from short trajectories because they are controlled by transient finite-time instability rather than by frequently observed bulk dynamics. We propose a mechanism-aware conditioning plug-in framework that turns a nudged coarse ensemble into a non-intrusive sensor of local instability geometry. In the small-noise regime, the ensemble covariance aggregates the same finite-time deformation kernels that govern local instability, providing a Jacobian-free proxy for the local amplification structure around a synchronized coarse trajectory. A small FiLM module injects statistics of this ensemble geometry into an otherwise unchanged backbone while leaving the coarse simulator unchanged. We demonstrate this interface in two distinct pipelines: a Transformer-style residual-attention corrector for a controlled low-dimensional chaotic system and a probabilistic recurrent STORN corrector for topographic two-layer quasi-geostrophic (QG) flow. In the low-dimensional benchmark, ensemble covariance directions co-activate with OTD modes and FiLM conditioning improves 99th-percentile exceedance-frequency errors over an identical no-context Transformer baseline. In QG, a fixed ensemble-conditioned FiLM-STORN model trained on only \(50\) time units substantially improves long-horizon rare-event statistics in the data-limited regime, including density-tail errors, exceedance frequencies, and spatial exceedance-area distributions relative to an unconditioned STORN trained on the same data; on averaged high-threshold exceedance diagnostics, it also outperforms the baseline STORN trained with $20$ times more high-resolution data. These results show that local instability geometry is not merely interpretable post hoc, but an actionable conditioning signal for data-efficient rare-event emulation.
Chinese Translation
混沌系统中的极端事件难以从短时间轨迹中学习,因为它们受瞬态有限时间不稳定性控制,而非频繁观测到的整体动力学。我们提出了一种机制感知的条件化插件框架,将受迫(nudged)粗分辨率集合转化为局部不稳定性几何结构的非侵入式传感器。在小噪声条件下,集合协方差聚合了与控制局部不稳定性相同的有限时间形变核,为同步粗轨迹周围的局部放大结构提供了一种无需雅可比矩阵(Jacobian-free)的代理。一个小型 FiLM 模块将该集合几何的统计量注入保持不变的骨干网络中,同时粗分辨率模拟器保持不变。我们在两条不同的流程中展示了这一接口:一是针对受控低维混沌系统的 Transformer 式残差注意力校正器,二是针对地形两层准地转(QG)流的概率循环 STORN 校正器。在低维基准测试中,集合协方差方向与最优传递(OTD)模态共同激活,且 FiLM 条件化相比完全相同的无上下文 Transformer 基线显著改善了第 99 百分位超越频率误差。在 QG 流中,仅用 50 个时间单位训练的固定集合条件化 FiLM-STORN 模型,在数据受限情形下,相对相同数据训练的无条件 STORN,显著改善了长期稀有事件统计,包括密度尾部误差、超越频率和空间超越面积分布;在平均高阈值超越诊断指标上,其表现也优于使用 20 倍高分辨率数据训练的基线 STORN。这些结果表明,局部不稳定性几何结构不仅具有事后可解释性,更是一种可用于数据高效稀有事件模拟的可操作条件化信号。
cs.LG / 55 / 2609.30752

Differentiable RNA Secondary Structure Extraction for Deep Learning

面向深度学习的可微RNA二级结构提取方法
Illman, Tyler, Ward, Max, Szikszai, Marcell, Krueger, Ryan K.
Abstract
Many deep learning approaches to RNA secondary structure prediction have recently been proposed. They typically output a weight matrix $W$ where $W_{ij}$ is an arbitrary weight for base $i$ pairing with base $j$. Converting this matrix to a predicted secondary structure or base-pairing probability matrix typically involves ad hoc and problematic downstream algorithms. Despite the importance of this conversion step, which we refer to as structure extraction, it has received relatively little attention in the literature. In this work, we analyze how the congruence between training and extraction methods affects prediction performance. To do this, we compare four extraction algorithms: a Nussinov-like dynamic programming method, maximum-weight graph matching and the greedy extraction algorithms used by SPOT-RNA and RiNALMo. These are evaluated on outputs from the pretrained RiNALMo model and three toy models trained in this paper: a differentiable Nussinov-like model, a binary cross-entropy (BCE) baseline, and a model that incorporates a novel symmetric doubly stochastic matrix (SDSM) normalization algorithm during training which allows it to output base-pairing probability matrices directly, without a separate extraction step. This SDSM normalization algorithm is differentiable and can be added inline to any deep learning model during training and evaluation. We find that the performance of each extraction method depends strongly on how the corresponding model was trained. Considering the toy models themselves, the SDSM model showed the strongest overall performance: it outperformed the BCE baseline under all four extraction algorithms and produced pre-extraction outputs closest to the ground truth. These results suggest that SDSM normalization is a tractable alternative to traditional structure extraction.
Chinese Translation
近年来,许多基于深度学习的RNA二级结构预测方法被相继提出。这些方法通常输出一个权重矩阵 $W$,其中 $W_{ij}$ 表示碱基 $i$ 与碱基 $j$ 配对的任意权重。将该矩阵转换为预测的二级结构或碱基配对概率矩阵,通常需要借助临时的、存在问题的下游算法。尽管这一转换步骤(我们称之为结构提取)十分重要,但在文献中却较少受到关注。在本工作中,我们分析了训练方法与提取方法之间的一致性如何影响预测性能。为此,我们比较了四种提取算法:类Nussinov动态规划方法、最大权重图匹配算法,以及SPOT-RNA和RiNALMo所使用的贪心提取算法。这些算法在预训练的RiNALMo模型的输出以及本文训练的三个玩具模型上进行了评估:一个可微的类Nussinov模型、一个二元交叉熵(BCE)基线模型,以及一个在训练过程中引入了新型对称双随机矩阵(SDSM)归一化算法的模型——该算法使其能够直接输出碱基配对概率矩阵,而无需单独的提取步骤。这一SDSM归一化算法是可微的,可以在训练和评估过程中内联添加到任何深度学习模型中。我们发现,每种提取方法的性能在很大程度上取决于相应模型的训练方式。就玩具模型本身而言,SDSM模型表现出最强的整体性能:它在全部四种提取算法下均优于BCE基线模型,且其提取前的输出最接近真实值。这些结果表明,SDSM归一化是传统结构提取的一种可行替代方案。
cs.LG / 56 / 2609.30781

Missingness-Aware Conformal Prediction Under Cross-Hospital Distribution Shift

跨医院分布偏移下考虑缺失性的共形预测
You, Liang, Ou, Dongwen, Shi, Hengyu, Dai, Siyuan
Abstract
Clinical measurements are recorded for some patients but not others, at rates that differ across hospitals, and marginal conformal coverage does not ensure coverage within groups defined by missingness. We propose a missingness-aware conformal calibration procedure for mortality prediction under cross-hospital distribution shift. It selects a measurement on an independent sample, groups patients by whether that measurement is recorded, and applies Mondrian calibration within each group, so no calibration outcome is reused. We evaluate the procedure across hospitals in eICU and across care units within one MIMIC-IV hospital, using three predictors. Relative to pooled calibration, it reduces the average worst-group coverage gap on its selected groups in all six settings, with a median reduction of 1.9 percentage points; paired site-bootstrap intervals exclude zero in five. These gains do not extend uniformly. Calibration by predicted risk achieves smaller gaps on a broader panel of missingness groups, and when eICU hospitals are evaluated separately, the gain shrinks for all three predictors and reverses in sign for one. We explain this discrepancy with a hospital-level decomposition. Pooling reweights hospitals through a covariance between group shares and coverage errors, and lets errors of opposite sign cancel: weighting explains the reversal, and cancellation accounts for most of the attenuation for the other two predictors. Constructed population distributions show that pooled and within-hospital evaluations can rank calibration methods oppositely even without sampling noise. Pooled improvement alone therefore cannot establish better coverage within hospitals, even when the calibration groups are fixed.
Chinese Translation
临床测量的记录因患者而异,且不同医院的记录率各不相同,而边际共形覆盖(marginal conformal coverage)并不能保证在按缺失性划分的子群内实现覆盖。我们提出了一种针对跨医院分布偏移下死亡率预测的、考虑缺失性的共形校准方法。该方法在独立样本上选择一项测量指标,按该测量是否被记录将患者分组,并在每组内应用 Mondrian 校准,从而避免重复使用任何校准结果。我们在 eICU 数据集的跨医院场景以及 MIMIC-IV 中某一家医院的跨护理单元场景下,使用三个预测模型对该方法进行了评估。相对于池化校准(pooled calibration),该方法在全部六种设置中都降低了其选定分组上的平均最差子群覆盖差距,中位降幅为 1.9 个百分点;其中五种设置的配对站点自助法(site-bootstrap)置信区间不包含零。然而,这些收益并非均匀分布:按预测风险进行校准在更广泛的缺失性分组面板上实现了更小的差距;而当对 eICU 各医院单独评估时,三个预测模型的收益均有所缩小,其中一个的符号甚至发生反转。我们通过医院层面的分解解释了这一差异。池化通过组占比与覆盖误差之间的协方差对医院进行重新加权,并允许符号相反的误差相互抵消:加权解释了符号反转,而抵消则解释了另外两个预测模型收益减弱的大部分原因。构造的总体分布表明,即使不存在抽样噪声,池化评估与院内评估也可能对校准方法给出相反的排序。因此,仅凭池化上的改进无法证明医院内部的覆盖更好,即使校准分组是固定不变的。
cs.LG / 57 / 2609.30789

Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays

可解释性内建的描述符组合在低数据分子测定任务上匹敌2048维基础模型嵌入
Yao, Yiqi, Duran-Frigola, Miquel
Abstract
In low-data structure-activity prediction, the choice of molecular representation can matter more than the choice of predictor, and tabular foundation models sharpen that effect. We ask whether a portfolio of compact, semantically named descriptor blocks can reach the accuracy of a 2048-dimensional CheMeleon embedding while staying auditable at the feature level, meaning that every input dimension carries a model name and a recorded training provenance. Starting from a fixed 11-dimensional physicochemical base, we greedily concatenate provenance-screened blocks using the labelled context alone. Across nine ADME/Tox assays and 50 evaluation cells, scored on common-coverage subsets restricted to the molecules that every representation covers, the portfolio reaches a mean test AUC of 0.762, against 0.764 for CheMeleon and 0.756 for Mordred. The pooled gap to CheMeleon is +0.003 AUC (task-bootstrap 95% CI [-0.020, +0.030]), which satisfies our predeclared pooled parity gate but not the per-assay gate. At 25 context labels the headline rule again satisfies the pooled gate; at 10 labels it does not. We also report four predeclared candidate-selection rules that we falsified. Post-freeze checks over ten seeds and three previously unseen assays support pooled competitiveness for compact, auditable representations; a same-width random-bundle control does not establish that greedy membership itself adds accuracy. Assay-level differences remain unresolved.
Chinese Translation
在低数据结构-活性预测中,分子表示的选择可能比预测器的选择更为关键,而表格基础模型进一步放大了这一效应。我们探究一个问题:由紧凑且语义可读的描述符块组成的组合(portfolio),能否在保持特征层面可审计的前提下,达到2048维CheMeleon嵌入的精度——可审计意味着每个输入维度都带有模型名称和记录在案的训练来源信息。从一个固定的11维理化性质基础块出发,我们仅利用标注上下文,通过贪心方式拼接经过来源筛选的描述符块。在九项ADME/Tox测定和50个评估单元上,仅在各表示共同覆盖的分子子集上进行打分,该组合达到了平均测试AUC 0.762,而CheMeleon为0.764,Mordred为0.756。与CheMeleon的汇总差距为+0.003 AUC(任务级bootstrap 95%置信区间为[-0.020, +0.030]),满足我们预先设定的汇总等效性门槛,但未满足逐任务门槛。在25个上下文标签时,主要规则再次满足汇总门槛;而在10个标签时则不满足。我们还报告了四个被证伪的预先设定的候选选择规则。在冻结后基于十个随机种子和三项先前未见测定任务的检查,支持紧凑、可审计表示在汇总意义上的竞争力;但同宽度的随机捆绑对照实验表明,贪心选块本身并不必然带来精度提升。测定任务层面的差异仍未解决。
cs.LG / 58 / 2609.30790

Towards Universal Representation-Based Process Control

迈向基于通用表示的过程控制
Choi, Jinmyeong, Kim, Taesup, Dubrawski, Artur
Abstract
Many temporal process learning and monitoring pipelines operate in local windows, making window-level decisions unavoidable in practice. In such settings, classical statistical tests can be applied to individual windows, but they typically evaluate predefined parametric hypotheses-such as unit-root or moment-based conditions-thereby limiting flexibility when reference behavior is defined empirically from task- or domain-specific data. In this work, we view window-level monitoring as a process control problem and reformulate it as reference-based hypothesis testing, where the null hypothesis is specified by an empirical reference distribution rather than a fixed parametric model. We operationalize this perspective through a representation-based, nonparametric framework that combines pretrained time series encoders, kernel density estimation, and conformal calibration, yielding finite-sample valid inference in learned representation space. Classical notions such as stationarity and cyclostationarity arise as natural instantiations of empirical reference sets within this framework. Through experiments, we demonstrate sensitivity to window-level distributional deviations while maintaining well-calibrated inference under stable reference regimes, highlighting the applicability of the proposed approach to a broad class of time series process control and monitoring tasks.
Chinese Translation
许多时间过程学习与监测流水线在局部窗口中运行,使得窗口级别的决策在实践中不可避免。在此类设置下,经典统计检验可以应用于单个窗口,但它们通常评估预定义的参数化假设(如单位根或基于矩的条件),当参考行为由任务或领域特定的数据经验性地定义时,这限制了其灵活性。在本工作中,我们将窗口级别的监测视为一个过程控制问题,并将其重新表述为基于参考的假设检验,其中零假设由经验参考分布而非固定的参数模型指定。我们通过一个基于表示的非参数框架来实现这一视角,该框架结合了预训练时间序列编码器、核密度估计和共形校准,在学习到的表示空间中提供有限样本有效的推断。平稳性和循环平稳性等经典概念在此框架中作为经验参考集的自然实例而出现。通过实验,我们展示了对窗口级别分布偏差的敏感性,同时在稳定的参考机制下保持校准良好的推断,突出了所提方法在广泛的时间序列过程控制与监测任务中的适用性。
cs.LG / 59 / 2609.30811

Counterfactual Online Conformal Prediction Under Adaptive Logging

自适应记录机制下的反事实在线共形预测
Qiao, Xinyu, Lin, Yichen, Ji, Kaihong, Wang, Xue, Yao, Tao
Abstract
Online conformal prediction can fail when predictions shape actions and actions determine which outcomes enter calibration. Standard adaptive methods may retain marginal coverage while systematically miscovering the counterfactual outcomes of rarely selected actions. This paper formalizes the failure through counterfactual coverage and introduces Propensity-Weighted Online Conformal Prediction, an inverse-propensity-weighted recursion that debiases calibration. A doubly robust variant further reduces nuisance bias to the product of outcome-model and propensity errors. Under positivity, the resulting coverage rate matches an information-theoretic lower bound up to logarithmic factors. Experiments on synthetic decision tasks, open bandit data, and financial rebalancing show that PW-OCP and DR-OCP improve counterfactual coverage and downstream regret without sacrificing prediction-set sharpness.
Chinese Translation
当预测影响行动、而行动又决定哪些结果进入校准时,在线共形预测可能会失效。标准的自适应方法虽可保持边际覆盖率,但可能系统性地误覆盖很少被选择的行动的反事实结果。本文通过反事实覆盖率对这一失效现象进行形式化,并提出了倾向加权在线共形预测(Propensity-Weighted Online Conformal Prediction),这是一种通过逆倾向得分加权递归来消除校准偏差的方法。其双重稳健(doubly robust)变体进一步将干扰项偏差降低至结果模型误差与倾向得分误差的乘积。在正定性(positivity)条件下,所得到的覆盖率在相差对数因子的意义上匹配信息论下界。在合成决策任务、公开老虎机数据以及金融再平衡上的实验表明,PW-OCP 与 DR-OCP 在不牺牲预测集精确度的前提下,提升了反事实覆盖率并降低了下游后悔值。
cs.LG / 60 / 2609.30819

Learning Provable Neural Network Observer for Uncertain Dynamical Systems

为不确定动态系统学习可证明稳定的神经网络观测器
Wang, Zhangyi, Liu, Jiaxu, Song, Chen, Xu, Chao, Cai, Shengze
Abstract
In many safety-critical applications, control of uncertain dynamical systems relies on observers that estimate states and external disturbances. Neural network observers can improve estimation accuracy, but certifying their Lyapunov stability via Linear Matrix Inequality (LMI) constraints leads to large-scale semidefinite programs (SDPs) that are difficult to solve for large networks. To overcome this scalability bottleneck, we propose a novel two-stage training framework for provably stable neural network observers. Our approach decouples the optimization into a point-guided Lyapunov pre-training phase, which rapidly achieves high estimation accuracy and local stability over sampled states, followed by an LMI fine-tuning phase that efficiently satisfies a strict global Lyapunov stability certificate. We provide formal theoretical guarantees for local stability radii and probabilistic coverage over a prescribed compact error-state domain under specified regularity and sampling assumptions. Experiments on nonlinear control benchmarks and X-29 aircraft ablations show that our LMI-certified neural network observers train significantly faster than direct LMI-based methods and generalize robustly across diverse systems, achieving improved tracking accuracy over a range of observer baselines. The code is available at https://github.com/Berry-Myon/LearningNeuralNetworkObserver.
Chinese Translation
在许多安全关键应用中,不确定动态系统的控制依赖于估计状态和外部扰动的观测器。神经网络观测器可以提高估计精度,但通过线性矩阵不等式(LMI)约束来证明其Lyapunov稳定性会导致大规模半定规划(SDP)问题,对于大型网络而言难以求解。为克服这一可扩展性瓶颈,我们提出了一种新颖的两阶段训练框架,用于训练可证明稳定的神经网络观测器。我们的方法将优化过程解耦:第一阶段是点引导的Lyapunov预训练,可在采样状态上快速实现高估计精度和局部稳定性;第二阶段是LMI微调,可高效满足严格的全局Lyapunov稳定性证书。在给定的正则性和采样假设下,我们为局部稳定半径以及在规定紧致误差状态域上的概率覆盖提供了正式的理论保证。在非线性控制基准以及X-29飞机消融实验上的结果表明,我们经LMI认证的神经网络观测器的训练速度显著快于直接基于LMI的方法,并能在多种系统上稳健地泛化,在一系列观测器基线中实现了更高的跟踪精度。代码可在 https://github.com/Berry-Myon/LearningNeuralNetworkObserver 获取。
cs.LG / 61 / 2609.30820

Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness

循环Transformer的量化:反馈暴露与校准盲区
Li, Nux
Abstract
Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails primarily at the non-residual loop-entry adapter, while quantizing the residual core is much less damaging. We call this feedback exposure: a quantized layer perturbs the recurrent state without an identity path, and the resulting error is fed back at later steps. Controlled experiments on linear filters and Mamba state-space models show that feedback exposure also occurs outside transformers. Grouped INT4 reveals a separate failure, calibration blindness: our one-step GPTQ baseline builds its Hessian from step-0 activations, leaving input directions used later in the recurrence nearly unweighted. Across nine checkpoints from seven looped architectures, one-step GPTQ is worse than round-to-nearest (RTN) on the primary task metric for five checkpoints. Accumulating the GPTQ Hessian across recurrence steps outperforms both one-step GPTQ and RTN on all nine checkpoints and recovers bf16-level accuracy on Huginn. These results separate two questions for PTQ on looped models: where quantization error enters the recurrence, and which states calibration sees.
Chinese Translation
循环Transformer(Looped transformers)在各递归步骤间复用权重,使得低比特量化尤其具有吸引力。我们识别出标准训练后量化(PTQ)的两种不同失效模式。在 Huginn-3.5B 上,逐通道(per-channel)INT4 量化的失败主要发生在非残差的循环入口适配器(loop-entry adapter)处,而量化残差核心的损害则小得多。我们将这种现象称为反馈暴露(feedback exposure):被量化的层在没有恒等通路的情况下扰动递归状态,由此产生的误差会在后续步骤中被反馈回来。在线性滤波器和 Mamba 状态空间模型上的受控实验表明,反馈暴露也会出现在 Transformer 之外。分组 INT4(Grouped INT4)则揭示了另一种独立的失效模式——校准盲区(calibration blindness):我们的单步 GPTQ 基线基于第 0 步的激活值构建 Hessian 矩阵,导致递归中后续步骤所用到的输入方向几乎没有被赋予权重。在来自七种循环架构的九个检查点上,单步 GPTQ 在五个检查点的主要任务指标上劣于最近取整(round-to-nearest, RTN)。在递归各步骤上累积 GPTQ 的 Hessian 矩阵,在全部九个检查点上都优于单步 GPTQ 和 RTN,并在 Huginn 上恢复至 bf16 级别的精度。这些结果将循环模型上 PTQ 的两个问题区分开来:量化误差从何处进入递归,以及校准过程能看到哪些状态。
cs.LG / 62 / 2609.30822

Adaptive Interaction Graphs for Particle Simulation

面向粒子模拟的自适应交互图
Zhou, Aiden
Abstract
Learned particle simulators based on graph neural networks achieve strong one-step accuracy, but errors compound over long horizons. An underexplored variable is the interaction graph: existing methods fix its topology via k-nearest neighbors or a static radius rule, regardless of local model confidence. We propose making this graph adaptive: a per-particle variance head, trained jointly with the acceleration head under a heteroscedastic Gaussian NLL loss, drives a trajectory in which high-uncertainty particles receive an expanded neighborhood. This is done at little extra inference cost by using the previous step's uncertainty estimate. A key discovery is that the variance head learns a meaningful notion of uncertainty: high-variance particles concentrate near complex regions, such as splash zones or free surfaces. When this signal drives graph topology, the resulting AdaptGNS simulator achieves a strict Pareto improvement on WaterDrop and a modest gain on Sand. Given the model's stronger performance on WaterDrop, we hypothesize that adaptive graphs are most useful when complexity is concentrated in space. Our code can be found at https://github.com/aidenzhou8/AdaptGNS.
Chinese Translation
基于图神经网络的学习型粒子模拟器在单步预测上表现出较高的精度,但在长时程模拟中误差会不断累积。一个尚未被充分研究的变量是交互图:现有方法通过k近邻(k-nearest neighbors)或固定半径规则来确定图的拓扑结构,而与局部模型置信度无关。我们提出使该图具有自适应性:通过一个逐粒子的方差预测头,在异方差高斯负对数似然(NLL)损失下与加速度预测头联合训练,驱动一条轨迹,使高不确定性的粒子获得更大的邻域范围。通过利用上一步的不确定性估计,这一过程几乎不增加额外的推理成本。一个关键发现是,方差预测头学到了有意义的不确定性概念:高方差粒子集中在复杂区域附近,如飞溅区域或自由表面。当该信号驱动图拓扑时,所得到的AdaptGNS模拟器在WaterDrop数据集上取得了严格的帕累托改进,在Sand数据集上取得了适度的提升。鉴于该模型在WaterDrop上的更强表现,我们假设自适应图在复杂性集中于空间局部时最为有用。我们的代码可在 https://github.com/aidenzhou8/AdaptGNS 获取。
cs.LG / 63 / 2609.30837

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation

MOPD-Router:重新思考多教师在线策略蒸馏中的教师路由
Xu, Tianze, Zheng, Yanzhao, Zhang, Zhentao, Yu, Yuanqiang, Ma, Chao, Zhu, Jihuai, Wu, Lelun, Ye, Lyumanshan, Liu, Pengfei, Dong, Baohua, Zhu, Hangcheng, Huang, Ruohui, Yu, Gang
Abstract
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabeled training mixtures and leaves complementary signals from other teachers unused. We introduce MOPD-Router, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model. Its plug-in interface supports different metrics for selecting and weighting teacher-specific OPD signals. Within this interface, we propose ExpertAlign, which scores each teacher by whether its correction to the student at the current token expresses the specialization that teacher acquired during post-training, and compare it against two reference metrics built on teacher confidence (Entropy) and teacher-student discrepancy (Novelty). Experiments on unlabeled and domain-labeled training mixtures under strong-to-weak and same-size distillation scenarios show that ExpertAlign achieves the strongest overall performance in all four settings. On unlabeled data, it improves the overall score by 5.88 (+12.3%) points over Mean aggregation; on domain-labeled data, it outperforms standard MOPD by 3.95 (+7.8%) points without using available domain labels. These results demonstrate token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment. Code is available at: https://github.com/TURLEing/MOPD-Router.
Chinese Translation
多教师在线策略蒸馏(MOPD)能够将多个专业能力整合到一个学生模型中,但现有方法通常将每个提示硬性路由至与其领域匹配的教师,并贯穿整个采样过程。这种对提示级领域标签的依赖限制了无标签训练混合数据的使用,也使得其他教师的互补信号未能得到利用。我们提出了MOPD-Router,这是一个在每个词元(token)上对完整教师池进行监督路由的框架,无需领域标签,也无需训练单独的路由模型。其插件式接口支持不同的度量方式来选择和加权各教师特定的OPD信号。在该接口下,我们提出ExpertAlign,它通过评估各教师在当前词元上对学生模型的修正是否体现了该教师在后训练阶段所获得的专业能力来为教师打分,并将其与基于教师置信度(Entropy)和教师-学生差异(Novelty)的两种参考度量进行比较。在强到弱蒸馏和同规模蒸馏场景下、于无标签和有领域标签的训练混合数据上的实验表明,ExpertAlign在全部四种设置中均取得最强的整体性能。在无标签数据上,相比Mean聚合,其整体得分提升5.88分(+12.3%);在领域标签数据上,即使不使用可用的领域标签,其性能也超过标准MOPD达3.95分(+7.8%)。这些结果表明,词元级路由能够利用跨领域互补监督,并减少对提示级领域分配的完全依赖。代码可在 https://github.com/TURLEing/MOPD-Router 获取。
cs.LG / 64 / 2609.30838

Peer-Grounded Counterfactual Path Planning for Chronic Health Management

基于同伴证据的反事实路径规划用于慢性病健康管理
Khamesian, Saman, Ghasemzadeh, Hassan
Abstract
Effective behavioral intervention in chronic disease management requires not a single prescription but a sequence of incremental steps, each grounded in what real, similar individuals have demonstrably achieved. Counterfactual explanation offers a natural computational route to such guidance, answering what change in behavior would have produced a better outcome. But existing methods return a target state without a route to it, guarantee no monotone health improvement along the way, and draw no evidence from peer behavior -- asking a patient to close a wide gap in one move, which is precisely the recommendation structure least likely to be attempted. We propose POROS (Peer-Grounded Optimal Routes Over States), a domain-agnostic framework rooted in Bandura's self-efficacy theory and Festinger's social comparison theory that constructs a Behavioral Progression Graph -- a directed acyclic graph over observed patient states in which every edge requires both peer-grounded behavioral proximity and strict health outcome improvement. Every edge is therefore a behavioral change that individuals in the cohort have demonstrated is achievable within a single period. Minimum-cost paths through this graph decompose otherwise inactionable behavioral gaps into incremental, peer-grounded steps. We evaluate POROS on two independent longitudinal cohorts of patients with diabetes. For patients below the 70% clinical threshold for time in range (TIR, blood glucose within 70-180 mg/dL), it reduces the mean gain required per step from 26.3 percentage points (pp) to 5.5 pp on one cohort and from 31.1 pp to 5.7 pp on the other, decomposing large behavioral jumps into the incremental steps that self-efficacy requires. Across both cohorts, 97-98% of multi-hop paths cross patient boundaries, embedding social comparison by construction.
Chinese Translation
慢性病管理中有效的行为干预需要的不是单一处方,而是一系列渐进式步骤,其中每一步都以真实的、相似个体已被证实能够实现的目标为依据。反事实解释为这类指导提供了一条自然的计算途径,即回答什么样的行为改变本可以带来更好的结果。但现有方法仅返回目标状态而不提供到达该状态的路径,无法保证过程中的健康指标单调改善,也没有从同伴行为中汲取证据——它们要求患者一步弥合巨大差距,而这恰恰是最不可能被尝试的建议形式。我们提出了POROS(Peer-Grounded Optimal Routes Over States,基于同伴证据的状态最优路径),这是一个领域无关的框架,植根于班杜拉(Bandura)的自我效能理论和费斯廷格(Festinger)的社会比较理论。该框架构建了一个行为进展图(Behavioral Progression Graph)——一个在观测到的患者状态之上的有向无环图,其中每条边都要求同时满足基于同伴证据的行为邻近性和严格的健康结果改善。因此,每条边都代表队列中个体已被证实在单个周期内可实现的行为改变。通过该图的最小成本路径将原本无法执行的行为差距分解为渐进的、有同伴证据支撑的步骤。我们在两个独立的糖尿病纵向队列上评估了POROS。对于处于目标范围内时间(TIR,血糖处于70-180 mg/dL范围内)低于70%临床阈值的患者,该方法在两个队列上分别将每步所需的平均提升幅度从26.3个百分点(pp)降低到5.5 pp,以及从31.1 pp降低到5.7 pp,将巨大的行为跳跃分解为自我效能所需的渐进步骤。在两个队列中,97-98%的多跳路径跨越了患者边界,从而在结构上嵌入了社会比较机制。
cs.LG / 65 / 2609.30839

Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation

面向同声传译的基于注意力的自适应策略
Tăşădan, Filip, Tomanová, Ema, Lopuch, Ondrej, Bilko, Paweł, Søgaard, Anders
Abstract
Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perform an accurate translation while minimizing the delay. To address this challenge, we utilize the cross-attention mechanism of the encoder-decoder architecture to find the right alignment between the input speech frames and the target text tokens. In this paper, we propose the Recent Frame Attention Policy (RFAP) and the Dual-Condition Attention Policy (DCAP) that allow offline trained speech-to-text translation models to be used in streaming scenarios without requiring additional training. Results on three different language translation pairs over the CVSS-C corpus show that the RFAP is able to surpass other policies with gains of up to 4.0 BLEU while reducing the translation delay by almost 1 second. Moreover, the DCAP is able to preserve a high translation quality when the latency is very low.
Chinese Translation
同声语音到文本翻译(Simultaneous speech-to-text translation, Simul-S2TT)是指在系统处理输入音频帧的同时生成部分译文。然而,这种流式设置的性质带来了一个挑战:在最小化延迟的同时,选择执行准确翻译的最佳时机。为应对这一挑战,我们利用编码器-解码器架构中的交叉注意力(cross-attention)机制来寻找输入语音帧与目标文本词元之间的正确对齐。在本文中,我们提出了近期帧注意力策略(Recent Frame Attention Policy, RFAP)和双条件注意力策略(Dual-Condition Attention Policy, DCAP),使得离线训练的语音到文本翻译模型无需额外训练即可应用于流式场景。在CVSS-C语料库上针对三个不同语言翻译对的实验结果表明,RFAP能够超越其他策略,BLEU提升高达4.0,同时将翻译延迟减少近1秒。此外,DCAP在极低延迟下仍能保持较高的翻译质量。
cs.LG / 66 / 2609.30840

Aligning One-Step Generative Models with Reward-Weighted Transport Distillation

基于奖励加权传输蒸馏的单步生成模型对齐
Wang, Austin, Cheng, Ziheng, Ying, Lexing
Abstract
One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit generators provide neither tractable likelihoods nor denoising trajectories, and many rewards are non-differentiable. We introduce Reward-Weighted Transport Distillation (RWTD), a post-training method that requires only generated samples and scalar reward evaluations. Rather than aligning solely to the conventional reward-tilted reference distribution, RWTD constructs an adaptive target that mixes separately tilted current and reference distributions. The current component incorporates improvements discovered during training, while the reference component anchors the target to the pretrained generator. RWTD realizes this target through feature-space optimal transport and fixed-point regression. Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge. Empirically, RWTD substantially improves the GenEval score of the one-step SANA Sprint 1.6B backbone from 0.73 to 0.80, while separate preference alignment experiments demonstrate strong cross-reward generalization that yields balanced improvements and preservation of compositional capabilities.
Chinese Translation
单步生成器仅需一次网络评估即可实现高质量视觉生成,但其后训练十分困难:一般的隐式生成器既不提供可处理的似然值,也不提供去噪轨迹,且许多奖励函数不可微。我们提出了奖励加权传输蒸馏(Reward-Weighted Transport Distillation, RWTD),这是一种仅需生成样本和标量奖励评估的后训练方法。RWTD并非仅对齐于传统的奖励倾斜参考分布,而是构建了一个自适应目标,该目标混合了分别倾斜的当前分布与参考分布。当前分布分量融入了训练过程中发现的改进,而参考分布分量则将目标锚定于预训练生成器。RWTD通过特征空间最优传输和不动点回归来实现这一目标。理论分析表明,RWTD的不动点分布在参考分布的离策略奖励倾斜与当前模型的在策略倾斜之间进行插值,为平衡奖励适应与先验知识保留提供了一种有原则的方法。实验表明,RWTD将单步SANA Sprint 1.6B骨干模型的GenEval分数从0.73大幅提升至0.80;此外,独立的偏好对齐实验展示了强跨奖励泛化能力,带来了均衡的性能提升并保留了组合生成能力。
cs.LG / 67 / 2609.30856

Learning Chance-Constrained MDPs with Bellman Distributional Certificates

基于Bellman分布证书的机会约束马尔可夫决策过程学习
Lu, Chenbei, Yi, Hongyu
Abstract
Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not necessarily imply a higher statistical price. For tabular discounted CCMDPs with fixed bounded successor support and access to a certified planning oracle, we establish a model-based upper bound, with a matching lower bound up to logarithmic terms. Technically, our key idea is the \emph{Bellman distributional certificate}, which constructs a Bellman recursion for constraint violation probabilities before policy selection. The certificate can be reused across candidate policies; combined with shared row-wise reverse-KL confidence sets, it gives a policy-uniform trajectory-KL transfer without a union bound over policies or time--budget Bellman tables. For stochastic policies, we give a model-free variance-reduced policy-gradient algorithm with a finite-sample expected KKT-residual guarantee and independent validation of every accepted policy. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage control benchmark illustrate the safety and mechanism behavior of the proposed algorithms.
Chinese Translation
安全强化学习(RL)通常施加期望成本约束,但这种期望意义上的安全性可能无法控制罕见高成本轨迹的发生概率。机会约束马尔可夫决策过程(Chance-Constrained MDPs, CCMDPs)提出了更强的概率层面要求,但被普遍认为更难求解,因为机会约束是非凸的,并且依赖于完整轨迹而非满足Bellman线性关系的期望。本文揭示,这种计算上的困难并不必然意味着更高的统计代价。对于具有固定有界后继状态支撑集、并可访问经验证规划预言机的表格型折扣CCMDPs,我们建立了一个基于模型的上界,并在忽略对数项的意义下给出了匹配的下界。在技术上,我们的核心思想是Bellman分布证书(Bellman distributional certificate),它在策略选择之前为约束违反概率构建Bellman递归。该证书可在候选策略之间复用;结合共享的逐行反向KL(reverse-KL)置信集,它给出了策略一致的轨迹KL迁移保证,而无需对所有策略取联合界,也无需随时间预算增长的Bellman表格。对于随机策略,我们提出了一种无模型的方差缩减策略梯度算法,具有有限样本期望KKT残差保证,并对每个被接受的策略进行独立验证。在合成CCMDPs和IEEE 14节点储能控制基准上的数值实验展示了所提算法的安全性与机制行为。
cs.LG / 68 / 2609.30884

CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters

CacheReforge:演化适配器下陈旧KV缓存的有界恢复方法
Cao, Yuhang, Mu, Yanzhou, Fang, Chunrong, Chen, Zhenyu
Abstract
Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current model outputs, while complete affected suffix recomputation restores fidelity at substantial cost. We seek minimal recomputation that recovers current adapter behavior. Existing systems track token, context, or stable adapter identity, but neither represent caches from earlier adapter versions nor distinguish update propagation from the recomputation required for behavioral recovery. To address these gaps, we introduce CacheReforge, which represents stale KV caches as layerwise mixed-version objects. It combines per-layer adapter anchors, calibrated sensitivity, accumulated drift, and executable restart boundaries to select direct reuse, bounded recomputation, or complete affected-suffix recovery. We distinguish dependency depth from the functional recomputation horizon and use cumulative tail influence to characterize when bounded recovery preserves current-model behavior. We evaluate CacheReforge on Qwen2.5-1.5B and Qwen2.5-7B with continual LoRA updates, including 16K HotpotQA and 2WikiMQA workloads. CacheReforge reduces mean KL divergence by 92.4% relative to stale reuse, while recomputing only 5.44% of layers and reducing cache-maintenance time by 93.2% relative to fresh full prefill. These results show that version-aware recovery preserves model fidelity and most KV caching gains.
Chinese Translation
大语言模型依赖KV缓存来减少长上下文和交互式应用中重复的预填充(prefill)计算。随着轻量级适配器(adapter)不断演化,缓存状态反映的是较早版本的模型,因此复用陈旧缓存会扭曲当前模型的输出;而对所有受影响的后缀进行完整重算虽然能恢复保真度,但代价高昂。我们寻求能够恢复当前适配器行为的最小重计算。现有系统跟踪token、上下文或稳定的适配器身份,但既无法表示来自更早适配器版本的缓存,也无法区分更新传播与行为恢复所需的重计算。为填补这些空白,我们提出CacheReforge,将陈旧的KV缓存表示为逐层混合版本对象。它结合逐层适配器锚点、校准敏感度、累积漂移和可执行的重启边界,来选择直接复用、有界重计算或对受影响后缀的完整恢复。我们区分了依赖深度与功能性重计算视野,并利用累积尾部影响来刻画有界恢复在何时能够保持当前模型的行为。我们在Qwen2.5-1.5B和Qwen2.5-7B上结合持续LoRA更新评估CacheReforge,包括16K长度的HotpotQA和2WikiMQA工作负载。与陈旧缓存复用相比,CacheReforge将平均KL散度降低了92.4%,同时仅重计算5.44%的层,且与全新的完整预填充相比将缓存维护时间减少93.2%。这些结果表明,版本感知的恢复方法能够保持模型保真度并保留KV缓存的大部分收益。
cs.LG / 69 / 2609.30918

Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations

对何种模型变化具有鲁棒性?鲁棒反事实解释的统一评估
Kostrzewa, Marcin, Zięba, Maciej
Abstract
Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We propose a unified cross-family evaluation protocol that holds factual instances and generated counterfactuals fixed while testing every method against the same eight types of model change. The benchmark compares six robust methods and two standard baselines on four tabular datasets. It characterizes every changed classifier through its outputs and reports empirical robustness together with coverage, base validity, and proximity. We find that relative performance and failure modes vary across change families. Bounded parameter perturbations change 0.95\% of test predictions on average, compared with 4.9\% for bootstrap retraining. Methods with guarantees for these perturbations do not necessarily transfer to other changes. RobX transfers most consistently in our experiments, although greater stability can require larger interventions. We argue that robust CFE methods should be evaluated through a common protocol that specifies the model changes, measures their realized behavioral magnitude, and keeps generation performance separate from robustness.
Chinese Translation
鲁棒反事实解释(robust counterfactual explanations)承诺在底层模型发生变化后仍能提供有效的补救建议(recourse)。这一承诺能否兑现取决于变化的具体性质。参数的微小扰动、基于新数据的重新训练以及新的模型架构是不同的事件,而每种现有方法都只针对其所设计的变化类型进行评估。因此,已报告的鲁棒性得分回答的是不同的问题,无法相互比较。我们提出了一个统一的跨族评估协议,在保持事实实例和生成的反事实样本固定的情况下,用相同的八种模型变化类型测试每种方法。该基准在四个表格数据集上比较了六种鲁棒方法和两种标准基线。它通过输出刻画每个变化后的分类器,并同时报告经验鲁棒性、覆盖率、基础有效性和邻近性。我们发现,相对性能和失败模式在不同变化族之间有所差异。有界参数扰动平均改变0.95%的测试预测,而自助法(bootstrap)重新训练改变4.9%。针对这些扰动具有保证的方法不一定能推广到其他变化。在我们的实验中,RobX的迁移一致性最高,尽管更大的稳定性可能需要更大的干预。我们主张,鲁棒反事实解释方法应通过一个统一的协议进行评估,该协议需明确指定模型变化、度量其实际行为幅度,并将生成性能与鲁棒性区分开来。
cs.LG / 70 / 2609.30929

EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting

EPOC:面向多步长时间序列预测的端点保持在线校正方法与压缩残差状态
Fujimoto, Takumi, Nishi, Hiroaki
Abstract
Completed multi-horizon forecasts provide residual feedback for a fixed forecaster, but retaining full residual blocks increases auxiliary state. We propose Endpoint-Preserving Online Correction (EPOC) with a compressed residual state. It stores low-order discrete cosine transform (DCT) coefficients and the final value of the preceding residual block. Within each channel, the endpoint is shared across component-wise online ridge regressions that also use current-forecast coefficients. The fitted DCT correction is blended with the base forecast. We evaluate eight multivariate series with DLinear and PatchTST, three seeds, and two training variants, yielding 96 matched fixed-base conditions at a 24-step horizon. EPOC achieves mean condition-wise reductions in mean squared error (MSE) and mean absolute error (MAE) of 15.40% and 9.35% from the uncorrected base, respectively, with a median of 6,352 B in retained auxiliary arrays. It has lower paired MSE than the $\delta$-Adapter, COSA, FAC, and OMPB in a majority of conditions and uses less state than each. Full ELF achieves the largest mean MSE reduction, 19.29%, but its median retained state is 474,048 B ($\times$75 relative to EPOC). Equal-size summary controls favor the endpoint by 1.65--2.20% in paired MSE; a coefficient-reconstructed endpoint yields similar accuracy to the observed endpoint, highlighting its role as a shared input. Increasing the retained DCT component count from 4 to 8 adds 1.00 percentage point of MSE reduction for 5,728 B. On jointly trained bases, EPOC lowers MSE by 16.69--20.15% relative to globally blended TEFL-style adapters applied to the same base. The code and numerical records are available at https://github.com/keiotakmin/endpoint-preserving-residual-correction.
Chinese Translation
完整的多步长预测可为固定的预测器提供残差反馈,但保留完整残差块会增加辅助状态。我们提出一种带压缩残差状态的端点保持在线校正方法(Endpoint-Preserving Online Correction, EPOC)。该方法存储前一个残差块的低阶离散余弦变换(DCT)系数及其末端值。在每个通道内,该端点在多个分量级在线岭回归之间共享,这些回归同时使用当前预测的系数。拟合得到的DCT校正与基础预测相融合。我们在八个多变量序列上使用DLinear和PatchTST进行评估,采用三个随机种子和两种训练变体,在24步预测视距下共得到96个匹配的固定基础条件。EPOC相较未校正的基础预测,均方误差(MSE)和平均绝对误差(MAE)的条件级平均降幅分别为15.40%和9.35%,其保留的辅助数组中位数为6,352字节。在大多数条件下,其配对MSE低于δ-Adapter、COSA、FAC和OMPB,且所用状态均少于上述方法。完整的ELF实现了最大的MSE平均降幅(19.29%),但其保留状态的中位数为474,048字节(约为EPOC的75倍)。等大小摘要对照实验显示,端点方法在配对MSE上优于摘要方法1.65%–2.20%;由系数重构的端点可达到与观测端点相近的精度,凸显了其作为共享输入的作用。将保留的DCT分量数从4增加到8,以5,728字节的额外状态换取1.00个百分点的MSE降幅提升。在联合训练的基础模型上,相较于应用于同一基础模型的全局混合TEFL风格适配器,EPOC将MSE降低16.69%–20.15%。代码和数值记录可在 https://github.com/keiotakmin/endpoint-preserving-residual-correction 获取。
cs.LG / 71 / 2609.30948

PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem

PORL:面向作业车间调度问题的预训练离线强化学习
Diz, Mateo Toro, Hoss, Jonathan, Klarmann, Noah
Abstract
The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization problem in industrial optimization. This work introduces Pretrained Offline Reinforcement Learning (PORL), a hybrid approach that combines simulation-based online pretraining with offline fine-tuning on production-specific data. Reinforcement learning through online interaction enables exploration of general scheduling strategies, but typically relies on simulation environments and may suffer from a simulation-to-reality gap. In contrast, offline RL avoids direct interaction with the environment by learning from historical data, but its performance is strongly influenced by dataset quality and coverage. PORL combines the strengths of both paradigms by first learning a general scheduling policy through online interaction and subsequently adapting it offline to a target distribution. A KL-divergence-based policy constraint is introduced to limit deviations from the pretrained policy during fine-tuning. The approach is evaluated on JSSP instances with distribution shift and datasets generated from heuristic, noisy-expert, and random behavioral policies. The results show that PORL consistently achieves lower optimality gaps than standalone offline RL and the considered general scheduling baselines. Furthermore, its advantage over standalone offline RL increases as dataset quality decreases, indicating reduced sensitivity to the quality and coverage of the available offline data. The results suggest that offline adaptation of pretrained policies is a promising approach for industrial scheduling environments where direct online exploration is impractical.
Chinese Translation
作业车间调度问题是工业优化中一个基础的组合优化问题。本工作提出预训练离线强化学习方法 PORL,这是一种将基于仿真的在线预训练与针对特定生产数据的离线微调相结合的混合方法。通过在线交互进行的强化学习能够探索通用的调度策略,但通常依赖于仿真环境,并且可能存在仿真与现实之间的差距。相比之下,离线强化学习通过从历史数据中学习来避免与环境的直接交互,但其性能在很大程度上受到数据集质量和覆盖范围的影响。PORL 结合了两种范式的优势:首先通过在线交互学习一个通用调度策略,随后在离线状态下将其适配到目标分布。本文引入了基于 KL 散度的策略约束,以限制微调过程中与预训练策略的偏离。该方法在存在分布偏移的 JSSP 实例上,以及由启发式、带噪声的专家和随机行为策略生成的数据集上进行了评估。结果表明,PORL 始终实现了比独立离线强化学习以及所考虑的通用调度基线更低的最优性差距。此外,随着数据集质量的下降,其相对于独立离线强化学习的优势进一步增大,这表明其对可用离线数据的质量和覆盖范围的敏感性降低。结果表明,预训练策略的离线适配对于直接在线探索不可行的工业调度环境而言是一种有前景的方法。
cs.LG / 72 / 2609.30950

Low-Bit Recurrent States in Hybrid Language Models

混合语言模型中的低位宽循环状态
Chen, Hongren, He, Jiayang
Abstract
Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them with normalized state ranges for mixed-precision bit allocation, without calibration data, rotation, or training. We also quantize decay rates logarithmically. With per-token state quantization, a four-bit mean payload reduces excess negative log-likelihood by factors of 3.3--27.9 relative to the best of seven baselines across three hybrid models; metadata costs vary. At six bits, negative log-likelihood differs from the FP32-state baseline by less than 0.005 nats. Ablations separate gains from variable bit widths, decay weighting, and range normalization. With less frequent write-backs, gains diminish and depend on the model and budget.
Chinese Translation
混合语言模型维护固定大小的循环状态,但现有的量化器通常使用八比特或更高的位宽。量化误差会随信道的衰减率持续存在。我们从可观测性格拉姆矩阵(observability Gramian)推导出失真权重,并将其与归一化的状态范围相结合,用于混合精度比特分配,且无需校准数据、旋转或训练。我们还对衰减率进行对数量化。通过逐词元的状态量化,在三个混合模型上,四比特的平均有效载荷相对于七种基线中的最优者,将超额负对数似然降低了3.3至27.9倍;元数据开销各不相同。在六比特下,负对数似然与FP32状态基线的差异小于0.005 nats。消融实验分别验证了可变比特宽度、衰减率加权和范围归一化带来的收益。在回写频率降低的情况下,收益减小并取决于模型和比特预算。
cs.LG / 73 / 2609.30957

Towards Understanding Momentum Acceleration in River-Valley Loss Landscape

理解河谷形态损失景观中的动量加速
Lu, Miao, Bian, Zeyu, Wen, Kaiyue, Wu, Beining, Chen, Siyu, Wang, Tianhao, Li, Zhiyuan
Abstract
The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "river-valley" structure, which features a low-loss manifold (river) flanked by sharp orthogonal directions with higher loss (mountains). In the long term, the optimization progress is determined primarily by the progress along the river. Within such a landscape, gradient descent with large learning rates can move faster along the river despite high apparent loss due to vertical oscillations, while a subsequent sharp decay in the learning rate suppresses these oscillations, revealing genuine optimization progress. This explains the recent success of warmup-stable-decay (WSD) learning rate scheduler which, unlike cosine scheduling, keeps stable high learning rate and decays before producing intermediate checkpoints. Building on this foundation, in this work we take a step further and study the role of momentum within such a loss landscape. We establish theoretical analysis that characterizes how momentum accelerates optimization by stabilizing large learning rates that can not be tolerated by vanilla GD without deviating significantly from the river. The enabled large learning rate in-turn gives greater speed along the river and makes faster essential progress in the long run. Another intriguing observation from theory is that for a river-valley landscape with very flat and slow-spinning river, the momentum itself does not contribute directly to acceleration in terms of the speed of tracking the river, while the main acceleration comes from the admissible larger learning rate.
Chinese Translation
预训练大语言模型在实践中的成功激发了对底层损失景观和优化动力学的深入研究。近期的实证与理论研究表明,训练损失景观往往呈现出“河谷”结构,其特征是一条低损失流形(河流),两侧是具有更高损失的陡峭正交方向(山脉)。在长期训练中,优化进程主要由沿河流方向的进展决定。在这样的景观中,大学习率的梯度下降法尽管由于垂直方向上的振荡而表现出较高的表观损失,却能沿河流更快地移动;而随后对学习率进行急剧衰减可以抑制这些振荡,从而揭示出真实的优化进展。这解释了近期预热-稳定-衰减(warmup-stable-decay, WSD)学习率调度器取得成功的原因:与余弦调度不同,它保持稳定的高学习率,并在产生中间检查点之前进行衰减。在此基础上,本工作更进一步,研究了动量在这种损失景观中的作用。我们建立了理论分析,刻画了动量如何通过稳定大学习率来加速优化——这些大学习率是普通梯度下降(GD)在不显著偏离河流的情况下无法容忍的。由此启用的大学习率反过来使模型能够更快地沿河流前进,从而在长期中取得更快的实质性进展。理论的另一个有趣发现是,对于河流非常平坦且旋转缓慢的河谷景观,动量本身在追踪河流的速度方面并不直接贡献加速,而主要的加速来自于可以容许的更大学习率。
cs.LG / 74 / 2609.30966

Gradient Surgery for Physics-Informed Neural Networks

物理信息神经网络的梯度手术方法
Borsani, Thomas, Di Fatta, Giuseppe
Abstract
Physics-Informed Neural Networks (PINNs) are trained by optimising a composite objective that combines data fitting with physics-based constraints, typically resulting in a highly imbalanced multi-task optimisation problem. Under these conditions, existing optimisation strategies are affected by conflicting task gradients, leading to slow convergence and unstable training, particularly for stiff and high-frequency partial differential equations. We analyse gradient conflicts throughout training of PINNs with standard optimiser and investigate Multi-Task Deep Learning (MTDL) optimisation methods. In our analysis across four benchmark problems we observed that PINN optimisation exhibits three distinct phases in which angle- and magnitude-based gradient conflicts alternate, with only one present at a time. Building on these observations, we propose PAM-GS, a physics-aware gradient surgery method that adaptively mitigates task interference during training according to the observed conflict types. Experiments on four representative PDE benchmarks demonstrate that PAM-GS combines competitive solution accuracy with consistently strong task-balanced performance, outperforming existing methods on most problems.
Chinese Translation
物理信息神经网络(Physics-Informed Neural Networks, PINNs)通过优化一个将数据拟合与基于物理的约束相结合的复合目标函数进行训练,这通常导致一个高度不平衡的多任务优化问题。在这些条件下,现有的优化策略会受到任务梯度冲突的影响,导致收敛缓慢和训练不稳定,尤其是对于刚性和高频偏微分方程。我们分析了使用标准优化器训练PINNs全过程中的梯度冲突,并研究了多任务深度学习(Multi-Task Deep Learning, MTDL)优化方法。在四个基准问题的分析中,我们观察到PINN的优化表现出三个不同的阶段,其中基于角度和基于幅值的梯度冲突交替出现,且每次只有一种冲突存在。基于这些观察,我们提出了PAM-GS,一种物理感知的梯度手术(gradient surgery)方法,可根据观察到的冲突类型在训练过程中自适应地缓解任务间干扰。在四个代表性PDE基准上的实验表明,PAM-GS在具有竞争力的求解精度的同时,始终展现出强大的任务均衡性能,在大多数问题上优于现有方法。
cs.LG / 75 / 2609.30973

LipSSM: Structurally Lipschitz-Bounded Cascaded State-Space Model via Metric Transfer between Consecutive SSM Layers

LipSSM:通过连续SSM层间的度量迁移构建结构化Lipschitz有界级联状态空间模型
Yoshino, Natsuki, Uchida, Ren, Matsumoto, Kazuki, Yatabe, Kohei
Abstract
Lipschitz continuity is a fundamental principle in the design of certifiably robust deep neural networks (DNNs), wherein adjusting the Lipschitz constant, which quantifies network robustness, is of central theoretical importance. A standard approach to enforcing Lipschitz continuity requires each layer of a DNN to be Lipschitz continuous, thereby guaranteeing overall Lipschitz continuity. However, this layer-wise approach typically imposes overly conservative restrictions by producing a loose estimate of the overall Lipschitz constant, which limits the expressive capacity of the DNN and degrades empirical performance at a prescribed level of robustness. To overcome this loose estimation, the recently proposed LipKernel transfers information across layers to yield a much tighter overall Lipschitz bound than conventional layer-wise construction. In this paper, we extend this concept to cascaded state-space models (SSMs) to construct Lipschitz-continuous DNNs capable of modeling longer-term dependencies. The proposed architecture, named LipSSM, is theoretically justified and empirically evaluated.
Chinese Translation
Lipschitz连续性是设计可认证鲁棒的深度神经网络(DNN)的一项基本准则,其中调节用于量化网络鲁棒性的Lipschitz常数具有重要的理论意义。强制Lipschitz连续性的标准方法要求DNN的每一层都是Lipschitz连续的,从而保证整体Lipschitz连续性。然而,这种逐层方法通常会产生宽松的整体Lipschitz常数估计,导致过度保守的约束,这限制了DNN的表达能力,并降低了在给定鲁棒性水平下的经验性能。为克服这种宽松估计问题,近期提出的LipKernel通过跨层传递信息,得到了比传统逐层构造 tighter 得多的整体Lipschitz上界。本文将这一概念扩展到级联状态空间模型(SSM),以构建能够建模更长期依赖关系的Lipschitz连续DNN。所提出的架构命名为LipSSM,并进行了理论论证和实证评估。
cs.LG / 76 / 2609.30977

Does Uniform Discrete Diffusion Need Time?

均匀离散扩散需要时间吗?
Hong, Chunsan, Lai, Chieh-Hsin, Hayakawa, Satoshi, Takida, Yuhta, Ye, Jong Chul, Mitsufuji, Yuki
Abstract
Uniform discrete diffusion models (UDMs) commonly use explicit time conditioning, but we find that it can often be unnecessary in practice. In this paper, we first show that the population-optimal UDM predictor generally depends on time: time controls how much the model should trust the observed context. We then show that this dependence can become negligible in finite-data settings relevant to language. When a corrupted training sequence remains much closer to its original clean sequence than to competing training sequences, the empirical-optimal predictor is nearly insensitive to time over most of the diffusion trajectory, where the guarantee weakens toward the high-noise endpoint. Empirically, trained language UDMs exhibit limited time sensitivity over most of the trajectory, while time-agnostic predictors remain competitive with, and often outperform, time-conditioned models across datasets and training objectives. These results challenge the use of explicit time conditioning in UDMs: although the population optimum depends on time, explicitly conditioning on it may often be unnecessary in practice.
Chinese Translation
均匀离散扩散模型(UDMs)通常使用显式的时间条件化,但我们发现它在实践中往往是不必要的。本文首先证明,总体最优的 UDM 预测器通常依赖于时间:时间控制着模型应在多大程度上信任观测到的上下文。随后我们表明,在与语言相关的有限数据场景中,这种依赖性可以变得可忽略不计。当一个被破坏的训练序列与其原始干净序列的距离远小于其与竞争训练序列的距离时,经验最优预测器在扩散轨迹的大部分区间内对时间几乎不敏感,而该保证在高噪声端点附近会有所减弱。实验上,训练得到的语言 UDM 在轨迹的大部分区间内表现出有限的时间敏感性,而与时间无关的预测器在各种数据集和训练目标下与时间条件化模型相比具有竞争力,且往往表现更优。这些结果对 UDM 中使用显式时间条件化提出了质疑:尽管总体最优解依赖于时间,但在实践中显式地以时间为条件往往是不必要的。
cs.LG / 77 / 2609.30995

Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models

学习气候模式中强迫对温度影响的层次化因果表示
Zhao, Shan, Trajkovic, Ilija, Kaltenborn, Julia, Gurwicz, Yaniv, Nowack, Peer, Rolnick, David, Boussard, Julien
Abstract
Machine learning (ML) emulators provide a fast and cost-effective method to simulate climate change scenarios after being trained on Earth System Models projections. However, the black-box nature of those data-driven approaches limit the usability and trustworthiness of their outputs and in particular their use as causal attribution tools. Here, we develop a hierarchical causal representation learning framework applied to sea surface temperature fields from a state-of-the-art global climate model. As a key advance over previous work, our framework explicitly models both atmospheric dynamical interactions arising from internal climate variability and forced responses due to changes in atmospheric greenhouse gas and aerosol concentrations. When trained on future climate change scenarios, our method accurately predicts the long-term global mean and regional temperature evolution and shows physically realistic responses to perturbations in greenhouse gas and aerosol concentrations when evaluated on unseen scenarios. Our results underline the potential of causal representation learning frameworks for advancing climate model emulation.
Chinese Translation
机器学习(ML)模拟器在经过地球系统模式投影数据训练后,可提供一种快速且经济高效的方法来模拟气候变化情景。然而,这些数据驱动方法的黑箱特性限制了其输出的可用性和可信度,尤其是将其用作因果归因工具的可行性。本文开发了一个层次化因果表示学习框架,并将其应用于最先进的全球气候模式的海表温度场。作为对以往工作的关键进展,我们的框架显式地建模了由气候内部变率产生的大气动力相互作用,以及由大气温室气体和气溶胶浓度变化引起的强迫响应。在未来气候变化情景上训练后,我们的方法能够准确预测长期全球平均及区域温度演变,并且在未见过的情景上评估时,对温室气体和气溶胶浓度的扰动表现出物理上合理的响应。我们的研究结果凸显了因果表示学习框架在推进气候模式模拟方面的潜力。
cs.LG / 78 / 2609.30996

The Linear Representation Hypothesis for Vision-Language-Action Models

视觉-语言-动作模型的线性表示假说
Jeong, Minseok, Choi, Hyewon, Tsukamoto, Hiroyasu, Han, SooJean
Abstract
The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vision-language-action (VLA) models, but the dynamical nature of embodied interaction introduces an additional challenge. Unlike semantic attributes commonly studied in LLMs, such as gender or language, a physical quantity of interest (QoI) in a VLA evolves jointly with the system dynamics: the representation influences the actions selected by the policy, which alter the physical state and, in turn, the next representation. In this paper, we develop a theoretical, signature-based formulation of the LRH for VLA that unifies representations and policies. On the representation side, we establish the existence of representations from which the future evolution of a QoI under a candidate action trajectory can be recovered via linear probing. On the policy side, we introduce a signature generalized linear model for stochastic action chunks. This structure yields a monotonic change in the expected future QoI along linear paths in natural parameter space, enabling linear steering. We construct an explicit oracle representation in a planar control-affine navigation experiment and verify the predicted linear probing and steering mechanisms.
Chinese Translation
线性表示假说(LRH)已成为通过大语言模型(LLM)内部表示来度量和干预语义信息的一种标准视角。越来越多的工作开始将这一视角扩展到视觉-语言-动作(VLA)模型,但具身交互的动态特性带来了额外的挑战。与LLM中通常研究的语义属性(如性别或语言)不同,VLA中感兴趣的物理量(QoI)会与系统动力学共同演化:表示影响策略所选的动作,动作改变物理状态,进而又影响下一时刻的表示。本文针对VLA提出了一种统一的、基于signature(签名)的LRH理论表述,将表示与策略统一起来。在表示方面,我们证明了存在这样的表示,使得可以通过线性探测(linear probing)恢复感兴趣的物理量在候选动作轨迹下的未来演化。在策略方面,我们针对随机动作块(action chunks)引入了一种signature广义线性模型。该结构使得期望的未来物理量沿自然参数空间中的线性路径单调变化,从而实现线性引导(linear steering)。我们在一个平面控制仿射导航实验中构造了显式的oracle表示,并验证了所预测的线性探测与线性引导机制。
cs.LG / 79 / 2609.31016

Robust Successor Features

鲁棒后继特征
Nikulski, Erik, Habib, Yamen, Gomez, Vicenç, Jonsson, Anders, Moreno-Bote, Rubén, Segovia-Aguas, Javier
Abstract
Generalization in Reinforcement Learning (RL) refers to the ability to execute close-to-optimal policies in unseen tasks after the agent has been trained on a different set of tasks. Building on the seminal work of the successor representation and further adaptations with function approximation, Transfer in RL has traditionally focused on generalizing to tasks that only differ in the reward function. A decade after the introduction of the successor representation, Robust RL emerged simultaneously from several articles in the field of operations research. In Robust RL, the transition kernel is unknown, and the goal is to maximize the expected reward under this uncertainty. Our work unifies these two paradigms through robust successor features, which generalize across both the reward function and the transition kernel, under the assumption that tasks are linear Markov Decision Processes. We derive a bound on Generalized Policy Improvement (GPI) that explicitly quantifies how performance degrades with the mismatch between transition kernels, recovering existing successor-feature guarantees when dynamics are shared. Finally, the generalization capabilities of robust successor features are validated on several grid-based benchmarks and compared to previous alternatives that focus solely on either the reward or the transition kernel.
Chinese Translation
强化学习(RL)中的泛化是指智能体在一组不同任务上训练后,能够在未见过的任务中执行接近最优策略的能力。基于后继表示这一开创性工作以及后续结合函数逼近的改进,强化学习中的迁移传统上专注于泛化到仅有奖励函数不同的任务。在后继表示提出十年后,鲁棒强化学习(Robust RL)同时源自运筹学领域的多篇文献。在鲁棒强化学习中,转移核是未知的,目标是在这种不确定性下最大化期望奖励。我们的工作通过鲁棒后继特征统一了这两种范式,在任务为线性马尔可夫决策过程的假设下,实现对奖励函数和转移核的泛化。我们推导了广义策略改进(Generalized Policy Improvement, GPI)的界,该界显式量化了性能随转移核失配程度的下降,并在动力学共享时恢复了现有的后继特征保证。最后,我们在若干基于网格的基准任务上验证了鲁棒后继特征的泛化能力,并与仅关注奖励或仅关注转移核的已有方法进行了比较。
cs.LG / 80 / 2609.31031

Metacognitive Selective Ensemble for Mobile Systems

面向移动系统的元认知选择性集成方法
Lee, Sungmin, Lee, Kichang, Lee, Joonhee, Park, JaeYeon, Kim, Songkuk, Ko, JeongGil
Abstract
Deep ensembles improve robustness in mobile sensing, but repeatedly executing many models over continuous sensor streams is costly. Selecting only a few members reduces this cost, yet adaptive selection often requires additional model execution to obtain reliable evidence about inactive candidates. We present MetaSE, an active ensemble framework that exploits short-term persistence in per-model reliability. MetaSE maintains a small active set across windows, uses post-execution evidence to reject unreliable members, and invokes lightweight routing only when replacement is needed. This stateful design accesses the diversity of a larger pool without repeated full-pool evaluation. Across four HAR datasets and four model architectures, MetaSE consistently improves over a fixed three-model ensemble and achieves accuracy comparable to substantially more expensive adaptive and full-ensemble inference. On a Raspberry Pi 4B, MetaSE is 2.7x faster and uses 69% less memory than full ten-model inference.
Chinese Translation
深度集成(deep ensembles)能够提升移动感知的鲁棒性,但在连续的传感器数据流上反复执行多个模型代价高昂。仅选择少量成员可降低该开销,然而自适应选择通常需要额外执行模型,以获取关于未激活候选模型的可靠证据。我们提出MetaSE,一个利用各模型可靠性短期持续性的主动式集成框架。MetaSE在多个时间窗口间维护一个较小的活跃集合,利用执行后的证据剔除不可靠成员,并仅在需要替换时调用轻量级路由。这种有状态的设计无需反复对整个模型池进行评估,即可获得更大模型池的多样性。在四个HAR数据集和四种模型架构上,MetaSE始终优于固定的三模型集成,并实现了与成本高得多的自适应推理和全集成推理相当的准确率。在Raspberry Pi 4B上,MetaSE相比完整的十模型推理速度快2.7倍,内存占用减少69%。
cs.LG / 81 / 2609.31033

Robust Graph Clustering Network for Multiple Missing Data

面向多重缺失数据的鲁棒图聚类网络
Qiu, Keyuan, Han, Renda, Tang, Zhen, He, Qiang, Wang, Xingwei, Zhang, Wenxin, Yao, Guangzhen, Chen, Junxin, Ni, Qingjian
Abstract
Clustering on graphs where both node attributes and structural links are partially missing remains a challenging task. Existing methods typically rely on imputation-then-clustering on single-view missingness incomplete graphs, which are vulnerable to cross-view error propagation and cluster-boundary blurring under simultaneous attribute and structure missingness. To address these limitations, we propose a Robust Graph Clustering Network for Multiple Missing Data (RGCN), which is designed to handle simultaneous node attribute and graph structure incompleteness. RGCN introduces three key innovations: First, we design a view-decoupled dual-branch imputation to mitigate interference and enable mutual enhancement in recovering missing data. Second, we employ a multi-hyperspherical mixture prior to enhance intra-cluster compactness and inter-cluster separability on a directional latent manifold. Third, a boundary-aware contrastive enhancement objective mitigates the blurring of clusters caused by imputation bias. Extensive experiments on real-world datasets demonstrate that RGCN consistently outperforms state-of-the-art baselines under various missing patterns.
Chinese Translation
在节点属性和结构链接均部分缺失的图上进行聚类仍然是一项具有挑战性的任务。现有方法通常依赖于单视图缺失不完备图上的“先插补后聚类”策略,在属性与结构同时缺失的情况下,这类方法容易受到跨视图误差传播和聚类边界模糊的影响。为解决这些局限性,我们提出了一种面向多重缺失数据的鲁棒图聚类网络(Robust Graph Clustering Network, RGCN),旨在处理节点属性与图结构同时不完备的问题。RGCN 引入了三项关键创新:首先,我们设计了视图解耦的双分支插补机制,以减轻干扰并实现缺失数据恢复中的相互增强;其次,我们采用多超球面混合先验,在方向性潜在流形上增强类内紧凑性和类间可分性;第三,边界感知的对比增强目标缓解了由插补偏差导致的聚类模糊。在真实世界数据集上的大量实验表明,在各种缺失模式下,RGCN 始终优于最先进的基线方法。
cs.LG / 82 / 2609.31038

Aurora-X: Built for Extreme Time Series Forecasting

Aurora-X:为极端时间序列预测而构建
Wu, Xingjian, Guo, Chenjuan, Qiu, Xiangfei, Hu, Zhigang, Cheng, Hanyin, Chen, Peng, Shu, Yang, Hu, Jilin, Yang, Bin
Abstract
Time series foundation models (TSFMs) enable cross-domain forecasting, but their development as general-purpose forecasters remains constrained by underexplored training potential and limited architectural versatility. To address these challenges, we introduce Aurora-X, a billion-scale TSFM with a progressive curriculum and a unified architecture. We first use channel-independent pretraining to learn temporal patterns, then introduce cross-variable dependencies, varied context and horizon lengths, and future covariates if available during midtraining. Variable-resolution post-training further enables an adjustable temporal span per token at inference. With fixed model weights, this supports longer histories under a fixed token budget or fewer tokens for the same history, enabling test-time scaling. With a versatile architecture, Aurora-X supports cross-variable modeling, covariate conditioning, and parallel decoding of future patches for probabilistic forecasting. These are supported by a novel pattern-guided mixture-of-experts that expands model capacity through sparse activation and uses shallow patch similarities to constrain deep-layer routing, guiding expert specialization across heterogeneous time series. Furthermore, we propose an implicit quantile network head that predicts arbitrary quantiles to characterize predictive distributions, enhancing probabilistic forecasting flexibility. Comprehensive experiments on GIFT-Eval, TIME, FEV-Bench, TFB, and DAG-Bench demonstrate state-of-the-art forecasting performance against pretrained TSFMs and task-specific supervised models.
Chinese Translation
时间序列基础模型(TSFM)能够实现跨领域预测,但作为通用预测器,其发展仍受限于未被充分挖掘的训练潜力和有限的架构灵活性。为应对这些挑战,我们提出了 Aurora-X,一个具有渐进式课程学习和统一架构的十亿级参数时间序列基础模型。我们首先采用通道独立的预训练来学习时间模式,随后在中间训练阶段引入跨变量依赖关系、多样的上下文与预测时域长度,以及(若有)未来协变量。变量分辨率的后续训练进一步使推理时每个 token 的时间跨度可调节。在固定模型权重的情况下,这既支持在固定 token 预算下使用更长的历史数据,也支持用更少的 token 表示相同的历史数据,从而实现测试时扩展。凭借灵活的架构,Aurora-X 支持跨变量建模、协变量条件化以及未来时间块的并行解码,以实现概率预测。这些能力由一种新颖的模式引导混合专家(mixture-of-experts)机制支撑,该机制通过稀疏激活扩展模型容量,并利用浅层的块相似性约束深层路由,引导专家在异构时间序列上实现专门化。此外,我们提出了一种隐式分位数网络输出头,可预测任意分位数以刻画预测分布,增强了概率预测的灵活性。在 GIFT-Eval、TIME、FEV-Bench、TFB 和 DAG-Bench 上的全面实验表明,Aurora-X 相对于预训练 TSFM 和任务特定的监督模型均取得了最先进的预测性能。
cs.LG / 83 / 2609.31061

Distributed Learning as a Service: The Developer's Perspective

分布式学习即服务:开发者的视角
Chu, Tianyue, Vannella, Filippo, Tsigkari, Dimitra, Delgado-Santos, Paula, López, Fernando, Guerrero, Pablo Gomez, Spantideas, Sotirios, Noguero, David Solans
Abstract
Application developers of distributed learning services face challenges that a typical federated learning loop does not address. Specifically, the model updates can still leak private data, devices might not be able to participate in the training due to limited resources, a single aggregator might not be able to scale, and the transmissions of model weights induce a considerable bandwidth cost. This paper demonstrates DLaaS (Distributed Learning as a Service) from the developer's vantage point. Using a single admin dashboard, the developer initiates a distributed/federated learning job and is able to activate Differential Privacy (DP), Split Learning (SL), Hierarchical Aggregation (HA), and Knowledge Distillation (KD) as declarative options, with no change to the clients' code. We demonstrate the complete service lifecycle on an industrial smart-home Wake-up Word (WuW) task, using the "Ok Aura" dataset. Once the developer initiates a distributed learning job by toggling DP, SL, HA, and KD in the admin dashboard, the system dispatches the job to a set of Android clients and Dockerized helper aggregators. In the demonstration, these mechanisms run live across configurations. Then, the clients train the model locally and return their updates. The trained model is served to a consumer-side Android application that performs on-device WuW detection on a live microphone stream. In particular, the conference attendees will be invited to speak the trigger phrase and monitor in real time the per-class confidence and inference latency. Finally, we release the source code and short video walkthroughs of these configurations.
Chinese Translation
分布式学习服务的应用开发者面临着典型的联邦学习流程无法解决的挑战。具体而言,模型更新仍可能泄露隐私数据,设备可能因资源有限而无法参与训练,单一聚合器可能无法扩展,且模型权重的传输会带来可观的带宽成本。本文从开发者的视角展示了DLaaS(分布式学习即服务,Distributed Learning as a Service)。开发者通过一个管理仪表盘即可启动分布式/联邦学习任务,并能够以声明式选项的方式启用差分隐私(Differential Privacy, DP)、拆分学习(Split Learning, SL)、分层聚合(Hierarchical Aggregation, HA)和知识蒸馏(Knowledge Distillation, KD),而无需更改客户端代码。我们在工业级智能家居唤醒词(Wake-up Word, WuW)任务上,使用"Ok Aura"数据集演示了完整的服务生命周期。当开发者在管理仪表盘中通过切换DP、SL、HA和KD启动分布式学习任务后,系统将该任务分发给一组Android客户端和Docker化的辅助聚合器。在演示中,这些机制在不同配置下实时运行。随后,客户端在本地训练模型并返回其更新。训练好的模型提供给消费端的Android应用,该应用在实时麦克风音频流上执行设备端唤醒词检测。特别地,我们将邀请与会者说出触发短语,并实时监测各类别的置信度和推理延迟。最后,我们开源了这些配置的源代码和简短的视频演示。
cs.LG / 84 / 2609.31082

SAGE: A sampling-aware global evaluation benchmark for species distribution modeling

SAGE:面向物种分布建模的采样感知全球评估基准
Arens, Emilia, van Tiel, Nina, Zbinden, Robin, Robert, Damien, Drees, Lukas, Vanalli, Chiara, Kellenberger, Benjamin, Zimmermann, Niklaus E., Pellissier, Loïc, Tuia, Devis, Wegner, Jan Dirk
Abstract
Knowing where species occur is fundamental for biodiversity research and conservation. Species distribution models (SDMs) link species observations to environmental conditions to estimate their spatial distribution. However, accuracy varies with the underlying data and models, making it essential to know for which species models can be trusted. Deep-learning-based SDMs ("DeepSDMs") now jointly model thousands of species, drawing on hundreds of millions of community-science records. At this scale, averaging performance hides substantial species-level variability, particularly for rare species, often of greatest conservation concern. Records are also strongly biased, making occurrence counts misleading. Accounting for these factors is essential for a reliable and informative evaluation of multi-species SDMs. Here, we introduce a Sampling-Aware Global Evaluation (SAGE) benchmark, combining GBIF records for training with sPlotOpen vegetation plots for presence-absence evaluation across 5771 plant species. We propose an evaluation framework that groups species based on two properties, sampling effort and relative prevalence, which describe how densely a species' range is sampled and how frequently the species is recorded. Evaluating single-species SDMs and multi-species DeepSDMs, we find that Random Forests and DeepSDMs perform best overall, but neither dominates: DeepSDMs outperform single-species SDMs for infrequently recorded species while offering no consistent advantage for well-sampled ones. Crucially, this advantage emerges only when established bias-correction practices, such as spatial thinning and reweighting, are carried over to the deep-learning setting. SAGE helps identify the species and data conditions for which a given approach is beneficial, thereby supporting the development of more transparent and ecologically credible SDMs. Data and code: https://earens.github.io/sage/
Chinese Translation
了解物种的分布位置是生物多样性研究和保护工作的基础。物种分布模型(SDMs)将物种观测数据与环境条件相关联,以估计其空间分布。然而,模型的精度随底层数据和模型的不同而变化,因此明确哪些物种的模型结果是可信的至关重要。基于深度学习的物种分布模型("DeepSDMs")如今可以联合建模数千个物种,并利用数亿条社区科学记录。在这种规模下,平均性能会掩盖显著的物种层面差异,尤其是对于稀有物种——它们往往是最受关注的保护对象。此外,观测记录存在强烈的偏差,使出现次数数据具有误导性。要可靠且有效地评估多物种SDMs,就必须考虑这些因素。本文提出了一个采样感知全球评估(SAGE)基准,将用于训练的GBIF记录与用于存在-缺失评估的sPlotOpen植被样方相结合,涵盖5771个植物物种。我们提出了一个评估框架,基于两个属性对物种进行分组:采样强度(即物种分布区被采样的密集程度)和相对普遍度(即物种被记录的频率)。通过对单物种SDMs和多物种DeepSDMs的评估,我们发现随机森林和DeepSDMs总体表现最佳,但两者互有优劣:对于记录频率较低的物种,DeepSDMs优于单物种SDMs,而对于采样充分的物种则没有一致的优势。关键在于,只有在将已有的偏差校正方法(如空间稀疏化和重加权)应用到深度学习场景时,这一优势才会显现。SAGE有助于识别在哪些物种和数据条件下某种方法更具优势,从而支持开发更透明、更具生态学可信度的SDMs。数据与代码:https://earens.github.io/sage/
cs.LG / 85 / 2609.31093

Block Sparse Attention with Log-Linear Complexity

具有对数线性复杂度的块稀疏注意力机制
Tang, Bohao, Qin, Zhen, Pan, Yuqi, Li, Zheng, Liu, Pengfei
Abstract
Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but selecting the retained blocks remains a bottleneck. Conventional block selection requires scoring all query-block pairs and therefore remains quadratic in sequence length. To address this issue, we propose PISA, a block-sparse attention mechanism that employs a pyramid Top-$K$ selection strategy. The main idea is to gradually narrow down the candidates across different levels, making it more efficient to find the most relevant keys. Specifically, we construct a coarse-to-fine hierarchy of keys and perform selection from the coarsest level. At each level, LogSumExp scoring is applied to a bounded candidate set to select candidates for the next finer level, continuing until the finest level is reached. Through pooling, we construct $O(\log N)$ levels of keys, yielding an overall complexity of $O(N\log N)$, where $N$ denotes the sequence length. We develop hardware-aware Triton kernels for both training and inference, fusing hierarchical routing and LogSumExp scoring without materializing the query-key score matrix. We further evaluate our method on language modeling tasks. Compared with the baseline, our method achieves comparable performance on benchmarks such as commonsense reasoning while delivering better results on retrieval tasks.
Chinese Translation
将语言模型扩展到长上下文受到自注意力二次方计算成本的限制。块稀疏注意力(block sparse attention)提供了一种高效的替代方案,但如何选择保留的块仍然是瓶颈所在。传统的块选择方法需要对所有查询-块对进行评分,因此复杂度仍随序列长度呈二次方增长。为解决这一问题,我们提出了 PISA,一种采用金字塔式 Top-$K$ 选择策略的块稀疏注意力机制。其核心思想是在不同层级间逐步缩小候选范围,从而更高效地找到最相关的键(keys)。具体而言,我们构建了一个由粗到细的键层次结构,并从最粗层级开始进行选择。在每一层级,我们对有界的候选集合应用 LogSumExp 评分,以选出进入下一更细层级的候选,直到达到最细层级。通过池化操作,我们构建了 $O(\log N)$ 个层级的键,从而获得 $O(N\log N)$ 的整体复杂度,其中 $N$ 表示序列长度。我们开发了面向硬件的 Triton 内核,用于训练和推理,将层次化路由与 LogSumExp 评分融合在一起,无需显式生成查询-键评分矩阵。我们进一步在语言建模任务上评估了我们的方法。与基线相比,我们的方法在常识推理等基准测试上取得了相当的性能,同时在检索任务上表现更佳。
cs.LG / 86 / 2609.31098

The Residual Stream's Effective Depth

残差流的有效深度
Gahtan, Barak, Galil, Ido, Bronstein, Alex M.
Abstract
We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only language models, $\Deff$ separates a structural consequence of residual accumulation from an empirical one: even maximally diverse orthogonal updates have the closed-form reference $F_L = 2L/(L+1)<2$, yet fifteen of sixteen default measurements lie below $F_L$ (Qwen3.5: 32--44\%, OLMo-2: 40--41\%, Pythia: 23--28\%). Matched references show that the gap is not caused by the persistent initial state or update-size imbalance, but is largely a calibrated signature of correlated residual updates rather than evidence that depth is unused. Symmetric position-0, token-normalisation, and top-PC controls show the regime is not reducible to BOS or top-PC artefacts: the lone above-reference default outlier joins the same regime, and all sixteen models are sub-reference after token-normalisation or top-1-PC removal. Intermediate checkpoints show that the regime is established early in OLMo-2 and stable through 5T tokens, while Pythia-1.4B follows a distinct decreasing trajectory. A controlled residual-carry intervention supports the mechanism, and $\Deff$ is best read as a \emph{global} accumulated-state diagnostic, not as a capability score or pruning method.
Chinese Translation
我们提出了“有效深度”($\Deff$)这一标量诊断指标,它将Transformer的逐层残差流视为一个离散时间过程,衡量表征相似度随层间距离的衰减情况,并将该衰减曲线聚合为单个数值。在十六个仅解码器(decoder-only)语言模型上,$\Deff$ 将残差累积的结构性后果与经验性后果区分开来:即使是最大多样化的正交更新,也有闭式参考值 $F_L = 2L/(L+1)<2$,然而十六个模型的默认测量值中有十五个低于 $F_L$(Qwen3.5:32–44%,OLMo-2:40–41%,Pythia:23–28%)。匹配参考实验表明,这一差距并非由持续存在的初始状态或更新幅度不平衡所致,而在很大程度上是相关残差更新的校准信号,而非深度未被利用的证据。对称的position-0、词元归一化以及top主成分(top-PC)对照实验表明,该机制不能被归约为BOS或top-PC伪影:唯一一个高于参考值的默认离群点也落入同一机制,且所有十六个模型在经过词元归一化或去除top-1主成分后均低于参考值。中间检查点实验显示,该机制在OLMo-2中早期即已确立并在5T词元训练过程中保持稳定,而Pythia-1.4B则呈现独特的下降轨迹。受控的残差传递(residual-carry)干预实验支持了该机制。$\Deff$ 最恰当的理解是一个关于全局累积状态的诊断指标,而非能力评分或剪枝方法。
cs.LG / 87 / 2609.31107

Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods

基于Fisher信息几何的贝叶斯优化:梯度界与信赖域方法
Kiroriwal, Saksham, Pfrommer, Julius, Beyerer, Jürgen
Abstract
We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of reparameterizable acquisition functions. This view explains vanishing-gradient behavior in high-dimensional BO and provides a common interpretation of heuristics such as RAASP and dimension-scaled lengthscales. Building on this analysis, we propose FITR, a trust-region-based BO method that replaces lengthscale-based scaling by local pullback-Fisher weights. FITR is not restricted to GP kernels with explicit lengthscales. On GP benchmarks with an SE kernel, experiments show competitive performance using FITR. The proposed method also easily generalizes to non-isotropic surrogates, although the gains are more task-dependent in that setting.
Chinese Translation
我们从信息几何的视角研究贝叶斯优化(BO)。通过代理后验映射将Fisher信息度量拉回到输入空间,得到输入空间上的局部敏感性张量,从而为可重参数化的采集函数的梯度提供上界。这一视角解释了高维贝叶斯优化中梯度消失的现象,并为诸如RAASP和维度缩放长度尺度(dimension-scaled lengthscales)等启发式方法提供了统一的解释。基于这一分析,我们提出了FITR,一种基于信赖域的贝叶斯优化方法,它用局部拉回Fisher权重替代基于长度尺度的缩放。FITR不局限于具有显式长度尺度的高斯过程(GP)核。在采用SE核的高斯过程基准测试中,实验表明FITR具有有竞争力的性能。所提出的方法还能轻松推广到非各向同性的代理模型,尽管在该设定下性能提升更依赖于具体任务。
cs.LG / 88 / 2609.31114

From Shortcut Learning to Discrete Neural Insertion Sort

从捷径学习到离散神经插入排序
Mylonas, Konstantinos, Spyropoulos, Thrasyvoulos
Abstract
Neural algorithmic reasoning aims to train neural networks to follow known algorithms and generalize beyond the input sizes seen during training. However, correct final outputs and intermediate supervision do not necessarily show that a model follows the intended execution. We study this problem using insertion sort. Our analysis of the CLRS30 baseline NAR shows that the hint objective is weakly optimized and that hint accuracy remains low. Moreover, many intermediate representations can already be decoded into sorted sequences before the reference insertion-sort execution terminates, suggesting that the model learns a shortcut to the final output. Motivated by these findings, we introduce Discrete Neural Insertion Sort. Our model represents the sequence as a chain, separates scalar exchanges from control-state transitions, and projects node representations back to discrete states after every processor step. When trained only on sequences of length 16, the model achieves $100\%$ sorted-sequence accuracy on sequences of length 64 and 128. However, an ablation shows that discretization and graph structure alone are insufficient: without additional supervision of the global inner-loop state, the model fails even at the training length. Our results show that discrete execution can support strong length generalization, while also highlighting the problem-specific inductive bias required to learn a faithful algorithmic execution.
Chinese Translation
神经算法推理(Neural Algorithmic Reasoning)旨在训练神经网络遵循已知算法,并泛化到训练时未见过的输入规模。然而,正确的最终输出和中间监督并不一定表明模型遵循了预期的执行过程。我们以插入排序为对象研究这一问题。我们对 CLRS30 基线的神经算法推理(NAR)模型的分析表明,提示(hint)目标仅被弱化优化,且提示准确率始终较低。此外,在参考插入排序执行尚未结束之前,许多中间表示就已经可以被解码为有序序列,这表明模型学习到了通向最终输出的捷径。受这些发现的启发,我们提出了离散神经插入排序(Discrete Neural Insertion Sort)。我们的模型将序列表示为一条链,将标量交换与控制状态转移分离,并在每个处理器步骤之后将节点表示投影回离散状态。当仅在长度为 16 的序列上训练时,该模型在长度为 64 和 128 的序列上实现了 100% 的有序序列准确率。然而,消融实验表明,仅有离散化和图结构是不够的:若缺乏对全局内循环状态的额外监督,模型即使在训练长度上也会失败。我们的结果表明,离散化执行能够支持强大的长度泛化能力,同时也凸显了学习忠实算法执行所需的特定于问题的归纳偏置。
cs.LG / 89 / 2609.31128

Frame the adversary: a structure-aware attack methodology

框定对抗者:一种结构感知的攻击方法
Kouni, Vicky, Perrakis, Stelios, Bach, Francis, Frossard, Pascal, Chevaleyre, Yann
Abstract
Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for robust evaluation of safety-critical and security-sensitive applications. Yet, existing approaches are typically not derived as solutions to an optimization problem that explicitly captures transform-domain structure. In this paper, we propose a methodology for crafting principled frequency-based adversarial attacks, via a dedicated optimization framework. A cornerstone of our method hinges on the introduction of a perturbation constraint set, tied to highly structured non-orthogonal transforms, well-known for their flexible, non-predefined frequency handling. We prove that the attacks emerge as weighted $\ell_2$-projections onto this set, yielding a general and controlled attack generation mechanism. By this, we provide a clear geometric attack characterization, ensuring alignment between the optimization objective and the perturbation constraint. We assess our framework on standardized datasets, for pretrained and adversarially robust models. Results highlight that our attacks, being solutions to an optimization problem, over a structured perturbation set, are highly effective, even across different, unseen architectures. Our methodology could serve as a theoretical baseline for designing and analyzing transformed-based attacks, targeting fundamental model vulnerabilities, instead of mere architecture-specific artifacts typically studied in the robustness literature.
Chinese Translation
基于频率的对抗攻击近期颇受关注,其原因在于其利用了各类神经架构共有的频谱敏感性。与空间扰动不同,基于频率的攻击能够暴露更深层脆弱性,因此对安全关键和安全敏感应用的鲁棒性评估尤为有价值。然而,现有方法通常并非作为显式刻画变换域结构的优化问题的解而导出的。本文提出了一种通过专用优化框架构建有原则的基于频率对抗攻击的方法。我们方法的一个基石是引入一个与高度结构化非正交变换相关的扰动约束集,这类变换以其灵活、非预定义的频率处理能力而著称。我们证明,攻击可以表示为该集合上的加权 ℓ₂ 投影,从而形成一个通用且可控的攻击生成机制。由此,我们提供了清晰的几何攻击刻画,确保优化目标与扰动约束之间的一致性。我们在标准化数据集上,针对预训练模型和对抗鲁棒模型评估了该框架。结果表明,我们的攻击作为结构化扰动集上优化问题的解,具有很高的有效性,甚至能够泛化到不同的、未见过的架构上。我们的方法可作为设计和分析基于变换的攻击的理论基准,针对模型的基本脆弱性,而非鲁棒性文献中通常研究的特定架构性缺陷。
cs.LG / 90 / 2609.31149

CRNDiff: Count-Native Diffusion Framework via Chemical Reaction Networks

CRNDiff:基于化学反应网络的计数原生扩散框架
Qiu, Yuxuan, Gagrani, Praful, Kobayashi, Tetsuya J
Abstract
Scientific measurements such as single-cell RNA (scRNA) sequencing often take the form of nonnegative integer counts, whereas continuous-state diffusion models approximate this discrete structure using continuous coordinates. Building on stochastic chemical reaction networks (CRNs), a class of count-native Markov jump processes, we introduce CRNDiff, a structured framework that combines count-space diffusion with inference-time conditioning on rare subpopulations. An independent birth--death instantiation yields a closed-form transition kernel for forward noising. This kernel enables reverse sampling via forward-filtering backward-sampling (FFBS) and supports data-driven selection of the terminal noising time, eliminating the need for a validation sweep. This tractability also lets us introduce tilted Feynman--Kac (FK) steering, a method for sampling target subpopulations from a frozen generator without retraining. By tilting posterior marginals before FK particle correction, steering mitigates importance-weight concentration when the target population is rare. Using scRNA-seq data from the human heart cell atlas, we test the ability of CRNDiff to generate cell-type-specific distributions. Across the three evaluated target populations, CRNDiff achieves the highest conditional fidelity among the evaluated generative models, with larger mean purity margins for rarer target populations. Generated cells preserve marker-level differential-expression structure. Replacing real training cells for the target classes with generated cells yields downstream classification performance approaching that of the real-data reference.
Chinese Translation
诸如单细胞RNA(scRNA)测序等科学测量数据通常以非负整数计数的形式呈现,而连续状态扩散模型则使用连续坐标来近似这种离散结构。基于随机化学反应网络(CRN)——一类计数原生的马尔可夫跳跃过程,我们提出了CRNDiff,这是一个结构化框架,将计数空间上的扩散与推理阶段对稀有亚群的条件化相结合。其独立的生灭过程实例化为前向加噪提供了闭式转移核。该核使得反向采样可通过前向滤波-后向采样(FFBS)实现,并支持对终止加噪时间的数据驱动选择,从而无需进行验证扫描。这种可处理性还使我们能够引入倾斜费曼-卡克(Feynman--Kac,FK)引导方法,该方法无需重新训练即可从冻结的生成元中采样目标亚群。通过在FK粒子校正之前对后验边缘分布进行倾斜处理,当目标群体较为稀有时,该引导方法能够缓解重要性权重的集中问题。利用人类心脏细胞图谱的scRNA-seq数据,我们检验了CRNDiff生成细胞类型特异性分布的能力。在所评估的三个目标群体上,CRNDiff在所有评估的生成模型中达到了最高的条件保真度,且目标群体越稀有,平均纯度裕度越大。生成的细胞保留了标记基因水平的差异表达结构。用生成细胞替换目标类别的真实训练细胞后,下游分类性能接近真实数据参考水平。
cs.LG / 91 / 2609.31155

Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift

域偏移下基于教师锚定的训练后量化模型选择
Dominguez, Alejandro Rodriguez, Shahzad, Muhammad, Hong, Xia
Abstract
Compressing a trained model yields a family of deployment candidates, and under domain shift the most compressed one need not be the one to deploy. We study selection over such a family, with candidates and teacher fixed and target labels absent or scarce. Two findings organize the label-free case. Minimum teacher distortion behaves almost as a constant rule, selecting the same eight-bit, per-channel, unclipped configuration in every run, which does not minimize empirical target cross-entropy. Established estimators divide sharply: in the overconfident-collapse regime of the CNN families, confidence-based estimators order the family close to backwards, and the diagnostics that identify it need the labels the setting denies, while output-distribution estimators match the teacher-relative anchor and on one architecture beat it. Distortion is nonetheless stable, so a supervised term can move selection away from it. Combining the two, we give exact quadratic identities for a canonical quadratic analogue of the family. We also show that under symmetric corruption the label-dependent part of a criterion linear in the label indicator is multiplied by one common factor whenever its coefficient sums are candidate-invariant, a class holding teacher contrasts and accuracy but not cross-entropy. These characterize the score's components without bounding selection regret. Across one hundred and thirty-four candidate families, one per independently trained convolutional or Vision Transformer teacher, anchoring reduces mean regret at the smallest label budget in every setting, an advantage that fades beyond twenty-five labels.
Chinese Translation
对一个已训练的模型进行压缩会得到一组可供部署的候选模型,而在域偏移(domain shift)条件下,压缩程度最高的模型未必是应当部署的模型。我们研究在这类候选集合上的选择问题,其中候选模型与教师模型固定不变,而目标域标签缺失或稀缺。针对无标签情形,我们得到两项发现。其一,最小教师失真(teacher distortion)几乎表现为一种恒定规则:在每次运行中都选中同一个八比特、逐通道(per-channel)、无截断的配置,而该配置并不能最小化经验目标交叉熵。其二,现有估计器呈现明显分化:在CNN模型家族的过度自信崩溃(overconfident-collapse)情形中,基于置信度的估计器对候选集合的排序几乎完全颠倒,而识别该情形的诊断手段又需要该设定所不具备的标签;相比之下,基于输出分布的估计器与教师相对锚定的结果一致,并在某一架构上甚至优于该锚定。尽管失真本身是稳定的,但引入监督项可以使选择偏离该失真锚定。将二者结合,我们针对该候选集合的一个典型二次类比给出了精确的二次恒等式。此外,我们证明:在对称腐蚀(symmetric corruption)条件下,只要某个对标签指示变量线性的准则的各项系数和具有候选不变性,则其中依赖标签的部分会被一个公共因子相乘,这一类准则包含教师对比(teacher contrasts)与准确率,但不包含交叉熵。这些结果刻画了得分的组成成分,但并未给出选择遗憾(selection regret)的界限。在134个候选集合(每个独立训练的卷积网络或Vision Transformer教师对应一个集合)上的实验表明,在所有设定中最小标签预算下,锚定方法均降低了平均遗憾;但当标签数超过25个时,该优势逐渐消失。
cs.LG / 92 / 2609.31157

Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection

基于物理信息预测先验的贝叶斯张量自编码器用于多维时间序列异常检测
Liu, Jianan, Li, Chunguang
Abstract
Multi-dimensional time series, inherently tensorial, are common in practice. Despite great progress in time series anomaly detection, most existing methods are confined to uni-/multi-variate time series. When handling multi-dimensional time series using these methods, reshaping operations are required, which inevitably break the intrinsic correlations and thus lead to performance degradation. In uni-/multi-variate time series anomaly detection, AutoEncoders (AEs) are widely adopted and generally categorized into reconstruction-based and prediction-based AEs. The reconstruction-based AE utilizes the current observation for reconstruction, while the prediction-based AE utilizes the historical information to predict the current observation. Thus, the two AEs utilize different information. To bridge the gap between reconstruction-based and prediction-based AEs, so as to fully leverage the available information and thus further enhance performance, we propose a predictive prior and incorporate it into the reconstruction-based AE. It may not be very difficult to conceive this idea, but designing the predictive prior so that it can work for tensor anomaly detection is non-trivial. Specifically, to avoid breaking the intrinsic correlations within the multi-dimensional time series, we use the tensor AE as the backbone. To incorporate the predictive prior into the reconstruction-based AE, we propose a Bayesian fusion approach and our analysis reveals that this approach can enhance the modeling capability of the model for normal data. To mitigate the over-generalization problem of AE, we incorporate physical laws, i.e. tensor low-rank decomposition rules, into the neural networks in the predictive prior, leading to the Physics-informed Predictive Prior Tensor AE (PPPTAE) framework. Experimental results on real-world datasets demonstrate the effectiveness of the proposed method.
Chinese Translation
多维时间序列本质上具有张量结构,在实践中十分常见。尽管时间序列异常检测已取得长足进展,但现有方法大多仅限于单变量/多变量时间序列。当使用这些方法处理多维时间序列时,需要进行重塑(reshaping)操作,这不可避免地会破坏数据的内在相关性,从而导致性能下降。在单变量/多变量时间序列异常检测中,自编码器(AutoEncoder, AE)被广泛采用,通常可分为基于重构的AE和基于预测的AE两类。基于重构的AE利用当前观测值进行重构,而基于预测的AE利用历史信息预测当前观测值,因此这两类AE所利用的信息不同。为弥合基于重构的AE与基于预测的AE之间的差距,从而充分利用可用信息并进一步提升性能,我们提出了一种预测先验,并将其融入基于重构的AE中。这一想法的构思或许并不困难,但设计出能够适用于张量异常检测的预测先验却并非易事。具体而言,为避免破坏多维时间序列的内在相关性,我们采用张量自编码器作为骨干网络。为将预测先验融入基于重构的AE,我们提出了一种贝叶斯融合方法,分析表明该方法能够增强模型对正常数据的建模能力。为缓解AE的过度泛化问题,我们将物理定律(即张量低秩分解规则)融入预测先验的神经网络中,由此构建了物理信息预测先验张量自编码器(Physics-informed Predictive Prior Tensor AE, PPPTAE)框架。在真实数据集上的实验结果证明了所提方法的有效性。
cs.LG / 93 / 2609.31161

I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?

我行动故我在:JEPA的动作条件化何时足以学习因果机制?
Liu, Yuhang, Huang, Zhuo, Shi, Javen Qinfeng
Abstract
Recent empirical and theoretical advances suggest that joint-embedding predictive architectures (JEPAs) may learn meaningful representations for action-conditioned prediction of future outcomes, thus becoming one of the foundational structures for world models. However, accurate prediction does not, in general, necessarily imply recovery of underlying causal states that give rise to the observed dynamics. This work investigates when and how JEPAs can recover the underlying causal states from observations. We first introduce a latent variable model, in which high-dimensional observations are generated from latent causal states whose dynamics are governed by action-conditioned transition mechanisms. Based on this formulation, we develop a general information-theoretic objective that combines conditional likelihood maximization for learning transition dynamics with entropy maximization for preserving latent state information. We then establish identifiability conditions under which representations learned by this general objective recover the underlying latent causal states up to component-wise invertible transformations and permutation. One key condition for such identifiability is sufficient action-induced variation in the transition mechanisms. Guided by this finding, we instantiate the general objective with an action-modulated Gaussian additive-noise model, yielding action-modulated JEPA (A-JEPA). Experiments on synthetic environments verify the theoretical findings under the identifiability conditions and robustness to moderate violations, while visual benchmarks demonstrate improved state recovery and transfer to unseen transition mechanisms.
Chinese Translation
近期的实证与理论进展表明,联合嵌入预测架构(JEPA)可能为动作条件下的未来结果预测学习到有意义的表示,从而成为世界模型的基础结构之一。然而,准确的预测通常并不必然意味着对产生观测动态的潜在因果状态的恢复。本工作研究了JEPA在何种条件下以及如何从观测中恢复潜在因果状态。我们首先引入一个潜变量模型,其中高维观测由潜在因果状态生成,而这些状态的动态由动作条件化的转移机制所支配。基于这一框架,我们提出了一个通用的信息论目标函数,它将用于学习转移动态的条件似然最大化与用于保留潜在状态信息的熵最大化相结合。随后,我们建立了可辨识性条件,在该条件下,通过这一通用目标学习到的表示能够恢复潜在的因果状态,直至逐分量的可逆变换和置换。此类可辨识性的一个关键条件是转移机制中存在足够的动作诱导变化。基于这一发现,我们用动作调制的加性高斯噪声模型实例化了该通用目标,得到动作调制JEPA(A-JEPA)。在合成环境上的实验验证了可辨识性条件下的理论发现以及对中等程度违反条件的鲁棒性,而在视觉基准上的实验则展示了更好的状态恢复能力以及对未见过的转移机制的迁移能力。
cs.LG / 94 / 2609.31162

WorldTS: World Modeling for Multimodal Covariate-aware Time Series Forecasting

WorldTS:面向多模态协变量感知时间序列预测的世界建模
Zhu, Yuhan, Qiu, Xiangfei, Cheng, Hanyin, Shen, Wangmeng, Guo, Chenjuan, Yang, Bin, Hu, Jilin, Jensen, Christian S.
Abstract
Time series forecasting is typically framed as learning a direct mapping from historical to future observations in the observation space. However, sequences of observations generally provide only a partial view of the dynamics of the underlying system, with future observations being shaped by latent dynamics. Recent latent-space forecasting methods thus achieve improved performance by predicting future observations from latent-space representations of historical observations rather than directly forecasting future observations in the observation space. Next, while future observations are also shaped by external factors, how to incorporate external, often multimodal, information into forecasting, so that it can shape latent-state formation and evolution directly, remains underexplored. We propose WorldTS, a world-modeling based forecasting framework that integrates multimodal covariates directly into the forecasting to further improve forecasting performance. Specifically, WorldTS employs a two-stage training strategy. First, it learns forecasting-relevant latent state dynamics conditioned on multimodal covariates, yielding encoded future states. Next, the learned state dynamics are frozen, and an observation decoder is trained to map the predicted future states back to future observations. Extensive experiments on 21 real-world datasets offer insight into WorldTS and its effectiveness.
Chinese Translation
时间序列预测通常被定义为在观测空间中学习从历史观测到未来观测的直接映射。然而,观测序列通常仅能提供对底层系统动态的部分视图,未来观测受潜在动态的塑造。因此,近期基于潜在空间的预测方法通过从历史观测的潜在空间表示来预测未来观测,而非直接在观测空间中预测未来观测,从而获得了性能提升。此外,未来观测还受外部因素的影响,但如何将外部的、通常是多模态的信息纳入预测,使其能够直接影响潜在状态的形成与演化,仍缺乏充分研究。我们提出WorldTS,一个基于世界建模的预测框架,它将多模态协变量直接整合到预测过程中,以进一步提升预测性能。具体而言,WorldTS采用两阶段训练策略:首先,它以多模态协变量为条件学习与预测相关的潜在状态动态,得到编码后的未来状态;随后,冻结所学的状态动态,并训练一个观测解码器,将预测的未来状态映射回未来观测。在21个真实数据集上的大量实验为WorldTS及其有效性提供了深入见解。
cs.LG / 95 / 2609.31168

Audio emotion recognition for atypical hearing

面向非典型听觉的音频情感识别
Roussel, Ulysse
Abstract
My doctoral work aims to explore Audio Emotion Recognition (AER) in the context of atypical listening. This research focuses on auditory hypersensitivity in people with autism, a phenomenon that is often difficult to evaluate and unique to each individual. Our core idea is to leverage our understanding of affect from acoustic traits, relying on the possibility of generalizing affective responses from a small amount of annotated data. As a first step, we fine-tune a large foundation model, Contrastive Language-Audio Pretraining (CLAP) using low-rank adaptation (LoRA), trained on a valence and arousal dataset of neurotypical listeners.
Chinese Translation
我的博士研究旨在探索非典型听觉情境下的音频情感识别(Audio Emotion Recognition, AER)。本研究聚焦于自闭症人群的听觉过敏现象,这一现象通常难以评估,且因人而异。我们的核心思路是利用从声学特征中理解情感的能力,依赖于从少量标注数据中泛化情感反应的可能性。作为第一步,我们使用低秩适配(LoRA)对一个大型基础模型——对比性语言-音频预训练(Contrastive Language-Audio Pretraining, CLAP)——进行微调,该模型在神经典型听众的效价与唤醒度数据集上进行了训练。
cs.LG / 96 / 2609.31197

ALF: An Active Learning Framework for Scientific Discovery

ALF:一个用于科学发现的主动学习框架
Surana, Shikha, Hawkins-Hooker, Alex, Gallup, Olivia, Brunken, Christoph, Tilly, Jules, Duckworth, Paul
Abstract
Machine learning for scientific discovery is almost systematically data bound. Producing relevant high quality data, under budget constraints, is amongst the most promising ways to advance the field. Active learning (AL) offers promise wherever labelling requires expensive experiment, measurement, or simulation. Most existing tools cover only part of the data acquisition loop, and typically focus on either offline benchmarking or online deployment, but not both. We present ALF, a modular AL Framework that runs the full data acquisition loop via five modular components. One clear API for both settings: offline, against an existing dataset for controlled and reproducible experimentation; and online, against an oracle for acquiring new candidates in real-world deployments. ALF is open-source and available at https://github.com/instadeepai/alf.
Chinese Translation
面向科学发现的机器学习几乎总是受制于数据。在预算约束下生产相关的高质量数据,是推动该领域发展最有前景的途径之一。无论何处,只要标注过程需要昂贵的实验、测量或仿真,主动学习(Active Learning, AL)就大有可为。现有的大多数工具仅覆盖数据获取循环的一部分,且通常只专注于离线基准测试或在线部署,而非两者兼顾。我们提出了 ALF——一个模块化的主动学习框架,它通过五个模块化组件运行完整的数据获取循环,并为两种场景提供统一清晰的 API:离线模式下针对已有数据集进行可控且可复现的实验;在线模式下针对预言机(oracle)在真实世界部署中获取新候选数据。ALF 是开源的,可在 https://github.com/instadeepai/alf 获取。
cs.LG / 97 / 2609.31206

Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra

自监督表示学习:从光谱基础模型到极光发射光谱
Lain, Matthieu Le, Cessateur, Gaël, Lefèvre, Sébastien
Abstract
Auroral spectrographs such as the Auroral Spectrograph In Skibotn (ASIS) record hundreds of thousands of emission spectra, but only a few hundred can be labelled by an expert. To exploit the rest, we pretrain a 1D Vision Transformer with a masked autoencoder on 223,000 unlabelled spectra. Without labels, its representation recovers the emission-line intensity ratios that physicists use to diagnose the precipitating particles (R^2 0.91 vs. 0.77 for an untrained control) and, under one linear probe, classifies as well as 13 features designed by experts. Fine-tuned, the model outperforms the previous supervised auroral classifier on its own benchmark (macro-AP 88.5 vs. 77.8), reaches 0.870 mAP, and exceeds the same architecture trained from scratch by +0.159 with 10% of the labels; attribution shows that it uses both N2+ bands. Could an existing pretrained model replace it? Two astronomical spectral foundation models and a time-series model transfer according to their spectral window: SpectraFM, trained in the infrared, falls below the untrained control, whereas SpecFormer, trained in the optical, approaches in-domain pretraining without reaching it.
Chinese Translation
诸如位于Skibotn的极光光谱仪(Auroral Spectrograph In Skibotn, ASIS)等设备记录了数十万条发射光谱,但专家仅能标注其中数百条。为了利用其余未标注数据,我们在223,000条未标注光谱上使用掩码自编码器(masked autoencoder)预训练了一个一维Vision Transformer。在无标签的情况下,其表示学习能够恢复物理学家用于诊断沉降粒子的发射线强度比(R^2为0.91,而未训练的对照组仅为0.77);并且在一个线性探测(linear probe)下,其分类性能可与专家设计的13个特征相媲美。经过微调后,该模型在其自身基准上超越了以往的监督式极光分类器(macro-AP为88.5,对比77.8),达到0.870 mAP,并且在仅使用10%标签的情况下,比从零训练的相同架构高出+0.159;归因分析表明该模型同时使用了N2+的两个波段。现有的预训练模型能否替代它?两个天文光谱基础模型和一个时间序列模型的迁移效果取决于其光谱窗口:在红外波段训练的SpectraFM表现低于未训练对照组,而在光学波段训练的SpecFormer则接近域内预训练的性能,但未能完全达到。
cs.LG / 98 / 2609.31222

Budgeted Quotient-Residual Guidance for Frozen Pocket-Conditioned Molecular Diffusion

面向冻结的口袋条件分子扩散模型的预算化商残差引导
Wang, Xinyu, Bi, Jinbo, Song, Minghu
Abstract
Pocket-conditioned molecular diffusion updates ambient atom coordinates, but many lead-optimization objectives are expressed on quotient features such as distances, contacts, and anchored substructures. We introduce budgeted quotient-residual guidance (QRG), an inference-time correction that makes these quotient objectives active without retraining the molecular generator. QRG lifts quotient covectors to metric-horizontal ambient directions and delivers them through a trust budget set by the frozen sampler's own step norm: quotient geometry chooses the direction, while sampler motion bounds the scale. We derive the horizontal lift, closed-form sampler-budget update, KL/kinetic interpretation around a frozen reverse step, equivariance conditions, and a product-budget split for budget-capped section and residual controls. Controlled quotient tasks confirm that sampler-relative delivery activates signals that raw local quotient gradients leave dormant. On frozen TargetDiff backbones, official seed-0 CBGBench ligand-generation/editing sweeps show practical quality-runtime gains: Local-QRG improves validity from 0.815 to 0.864 on fragment growing, 0.664 to 0.707 on scaffold hopping, and 0.681 to 0.712 on linker design, while PredNext-QRG improves fragment/scaffold and remains near-neutral on linker. Novelty remains 1.000 and diversity is preserved in the matched multi-seed molecular slice, giving task-dependent improvements without sampler retraining or backbone modification. Overall, QRG provides a lightweight route to quotient-aware inference for frozen molecular samplers with explicit runtime accounting.
Chinese Translation
口袋条件(pocket-conditioned)分子扩散模型对环境原子坐标进行更新,但许多先导化合物优化目标是以商特征(quotient features)表达的,例如距离、接触和锚定子结构。我们提出了预算化商残差引导(Budgeted Quotient-Residual Guidance, QRG),这是一种推理时校正方法,能够在不重新训练分子生成器的情况下使这些商目标发挥作用。QRG 将商余向量提升为度量水平的环境方向,并通过由冻结采样器自身步长范数所设定的信任预算来传递这些方向:商几何决定方向,采样器运动约束尺度。我们推导了水平提升、闭式采样器预算更新、冻结反向步骤附近的 KL/动力学解释、等变性条件,以及用于预算受限截面与残差控制的乘积预算分配。受控的商任务证实,相对于采样器的传递方式能够激活原始局部商梯度所无法激活的信号。在冻结的 TargetDiff 骨干网络上,官方 seed-0 CBGBench 配体生成/编辑实验显示了实际的质量-运行时间收益:Local-QRG 将片段生长的有效性从 0.815 提升至 0.864,将骨架跃迁从 0.664 提升至 0.707,将连接子设计从 0.681 提升至 0.712;PredNext-QRG 则改善了片段/骨架任务,并在连接子任务上保持近乎中性。在匹配的多种子分子切片中,新颖性保持为 1.000,多样性也得到保留,从而在无需重新训练采样器或修改骨干网络的情况下实现了依赖于任务的改进。总体而言,QRG 为冻结分子采样器提供了一条轻量级的商感知推理路径,并具有明确的运行时间核算。
cs.LG / 99 / 2609.31250

Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization

动态张量重算中的确定性状态切换与可行性反转
Pagadala, Mahesh Reddy
Abstract
We report fine-grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory-constrained DNN training, measured on the reference DTR simulator (simrd) using public execution traces. On an LSTM trace, memory budgets differing by 0.10% of unconstrained peak memory select fast and slow execution regimes whose overheads differ by as much as 7.3x; the slow regime is driven by broadly repeated re-eviction of the same storages (evictions per storage rise from 1.33 to 8.27 while the set of distinct evicted storages is essentially unchanged: 5,233 vs 5,236, with the two sets overlapping at Jaccard 0.999). On a ResNet-32 trace, a fine budget sweep reveals a deterministic feasibility inversion: the run is feasible at ratio 0.101, infeasible (OOM) across 0.102-0.106, and feasible again from 0.107. We trace the immediate cause of the OOM to a fully pinned recursive rematerialization frontier that exceeds the budget after every evictable tensor has been evicted. Ablations using the DTR authors' own variants implicate the joint size-staleness scoring term in the observed LSTM instability. We argue these are at least two distinct budget-sensitive pathologies rather than one mechanism, and we separate what is demonstrated from what remains hypothesised. All results concern the reference simulator; reproduction in a production runtime is future work. Code, instrumentation, and raw results accompany this preprint.
Chinese Translation
我们报告了动态张量重算(Dynamic Tensor Rematerialization, DTR)中细粒度、确定性的不稳定性。DTR 是一种面向内存受限 DNN 训练的在线逐出策略。我们在参考 DTR 模拟器(simrd)上使用公开执行轨迹进行了测量。在一条 LSTM 轨迹上,仅相差无约束峰值内存 0.10% 的内存预算会选择出快速与慢速两种执行状态,其开销差异高达 7.3 倍;慢速状态的成因是对相同存储的广泛重复逐出(每个存储的逐出次数从 1.33 上升到 8.27,而不同被逐出存储的集合基本不变:5,233 对 5,236,两个集合的重叠度 Jaccard 为 0.999)。在一条 ResNet-32 轨迹上,细粒度的预算扫描揭示了一种确定性的可行性反转:运行在比率 0.101 时可行,在 0.102–0.106 区间内不可行(内存溢出 OOM),而从 0.107 起再次可行。我们将 OOM 的直接原因追溯到一个完全被固定的递归重算前沿——在所有可逐出张量均被逐出之后,该前沿的内存需求超出了预算。利用 DTR 作者自己提出的变体进行的消融实验表明,所观测到的 LSTM 不稳定性与大小-陈旧度联合评分项有关。我们认为这些现象是至少两种不同的预算敏感病理,而非单一机制,并且我们区分了已被证实的结论与仍属假设的内容。所有结果均针对参考模拟器;在生产运行时中的复现留待未来工作。代码、插桩工具及原始结果随本预印本一并提供。
cs.LG / 100 / 2609.31291

Softmax Reparameterization for Output-Head Quantization

面向输出头量化的Softmax重参数化方法
Kadav, Asim, Flores, Christian, Arora, Chirag, Kotte, Varun, Zheng, Hongbo, Yan, Lan, Shanmugasundaram, Priya, King, Tracy Holloway
Abstract
Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ. This one-dimensional search includes the original head and fixed mean-centering, preserves the full-precision softmax distribution, and leaves the trained decoder unchanged; a rank-one correction handles nonlinear logit paths such as soft-capping. Across seven heads, W4 gains concentrate where baseline quantization substantially distorts predictions: on Phi-4-mini, AW-MSE KL falls from 0.936 to 0.256. The gains survive stronger GPTQ calibration and remain complementary to exact per-channel scaling and affine quantization. Across four heads and three W4 quantizers, frozen WikiText-selected coefficients also transfer to C4 and OpenWebMath, outperforming mean-centering in all 18 comparisons where the frozen coefficient differs from $1$ and matching it in the remaining six. At W2, used as a compression stress test, benefits broaden across nearly the full model--quantizer matrix. Matched residual analysis shows that improved fidelity can accompany greater logit reconstruction error while reducing the residual's Fisher-weighted cost. For shift-compatible heads, reparameterization adds no inference operation and preserves packed W4 execution: with the decoder held in BF16, quantizing the Phi output head reduces batch-one generation latency by 10.8% relative to the BF16-head baseline.
Chinese Translation
庞大的词表规模使得输出头(output head)成为小型语言模型推理中的重要开销。我们提出softmax重参数化,这是一种训练后量化(post-training)方法,可在量化之前选择一个功能等价的输出头。该方法从每个输出行中减去词表行均值的标量倍数,并分别针对RTN、激活加权MSE(activation-weighted MSE)和完整Hessian的GPTQ,通过验证集KL散度单独选择系数。这一维搜索涵盖了原始输出头和固定均值中心化两种情形,保持全精度softmax分布不变,且不改动已训练的解码器;对于soft-capping等非线性logit路径,则通过秩一修正(rank-one correction)加以处理。在七个输出头上的实验表明,W4量化下的增益主要集中在基线量化显著扭曲预测的情形:在Phi-4-mini上,AW-MSE的KL散度从0.936降至0.256。这些增益在更强的GPTQ校准下依然存在,并与精确的逐通道缩放和仿射量化保持互补。在四个输出头和三种W4量化器的组合中,由WikiText选定后冻结的系数同样能迁移到C4和OpenWebMath数据集:在冻结系数不等于$1$的18组比较中全部优于均值中心化,在其余6组中与之持平。在W2量化(作为压缩压力测试)下,收益扩展到几乎整个模型-量化器矩阵。匹配残差分析表明,在降低残差的Fisher加权代价的同时,保真度的提升可能伴随更大的logit重构误差。对于移位兼容(shift-compatible)的输出头,重参数化不增加任何推理操作,并保持打包W4执行不变:在解码器保持BF16的情况下,量化Phi的输出头相对于BF16输出头基线可将批大小为1的生成延迟降低10.8%。
cs.LG / 101 / 2609.31306

Benchmarking Attention for Tabular Foundation Models

面向表格基础模型的注意力机制基准测试
Schambach, Maximilian, Biehl, Clemens, Thelin, Sam
Abstract
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden dimensions used in current models are small compared to recent language models. Yet efficient attention has been studied mostly for one-dimensional sequences, leaving the two-dimensional tabular setting unexplored. To this end, we create a reproducible benchmarking setup and study the unique characteristics of tabular attention across several backends -- Torch SDPA (efficient and cuDNN), FlashAttention-2/3/4, and the inference-only backends vLLM and SageAttention -- measuring forward and backward throughput across realistic tabular shapes on three GPU generations (A100, H100, B200). We find that the optimal backend choice differs between column and row attention and varies across hardware as well as model specifics: While the FlashAttention implementations tailored for each GPU generation perform overall best, they are at times outperformed by CuDNN in the case of column attention at longer sequences with cross-over points depending on the head dimension. Among inference-only backends, SageAttention performs well for row attention and large sequences beyond 16\,k rows. Our reproducible benchmark lays the foundation for future improvements to table-native attention. The self-contained benchmarking and evaluation code is openly available at: https://github.com/SAP-samples/tabular-attention-benchmark
Chinese Translation
TabPFN、Mitra、ConTextTab 等表格上下文学习器依赖于对潜在嵌入的二维序列进行行注意力和列注意力的交替运算。这些注意力模式与语言模型中的一维情形差异显著:行注意力涉及更长的序列,而列注意力作用于短得多的序列,且表格数据的跨步内存布局使得生成连续张量的代价高昂。此外,与近期的大型语言模型相比,当前模型所使用的隐藏维度较小。然而,高效注意力机制的研究大多集中于一维序列,二维表格场景尚属空白。为此,我们构建了一个可复现的基准测试框架,并研究了表格注意力在多种后端下的独特特性——包括 Torch SDPA(efficient 与 cuDNN)、FlashAttention-2/3/4,以及仅用于推理的后端 vLLM 和 SageAttention——在三代 GPU(A100、H100、B200)上,针对真实的表格形状测量前向和反向吞吐量。我们发现,列注意力与行注意力的最优后端选择有所不同,且随硬件及模型细节而变化:虽然针对各代 GPU 定制的 FlashAttention 实现总体表现最佳,但在长序列的列注意力场景中有时会被 cuDNN 超越,其交叉点取决于头维度。在仅用于推理的后端中,SageAttention 在行注意力及超过 16k 行的大序列场景下表现良好。我们这一可复现的基准测试为未来表格原生注意力机制的改进奠定了基础。自成体系的基准测试与评估代码已在以下网址开源:https://github.com/SAP-samples/tabular-attention-benchmark
cs.LG / 102 / 2609.31315

LUCID: Learning Under Confounding for Inference and Discovery in Time Series

LUCID:面向时间序列推断与发现的混杂环境下学习
Fesanghary, Mohammad
Abstract
Unobserved common causes are pervasive in real-world time series and can induce spurious associations that causal discovery methods mistake for direct edges. We propose LUCID (Learning Under Confounding for Inference and Discovery, a regime-adaptive deconfounding layer that first estimates the confounding regime from data using a Mar\v{c}enko--Pastur spectral router, then applies a deconfounding strategy matched to that regime. When the spectrum indicates pervasive factor confounding, LUCID attenuates factor-dominated variation and recovers contemporaneous (lag-$0$) structure from the resulting innovations, with edge selection calibrated against a data-driven edge-free null. Rather than being tied to a particular discovery algorithm, it can wrap existing discovery engines; we demonstrate consistent improvements across three such methods. On a diverse synthetic out-of-distribution benchmark spanning changes in confounder strength and sparsity, loading density, lag structure, volatility dynamics, edge heterogeneity, persistence, intermittency, and tail behavior, LUCID achieves the best family-weighted directed, lag-resolved graph $F_1$ ($0.60$), improving over the strongest baseline by $0.19$ absolute ($\approx\!46\%$ relative). Its advantage widens relative to looser lag-collapsed scoring, and remains robust under intermittent and heavy-tailed confounding. Code reproducing the method, the benchmark generators, and every reported experiment is available at https://github.com/bloomberg/causal-ts.
Chinese Translation
未观测的共同原因在现实世界的时间序列中普遍存在,其诱导的虚假关联常被因果发现方法误认为直接因果边。我们提出了LUCID(Learning Under Confounding for Inference and Discovery,混杂环境下的推断与发现学习),这是一种自适应于不同机制的去混杂层:首先利用Marcenko–Pastur谱路由器从数据中估计混杂机制,然后应用与该机制相匹配的去混杂策略。当谱结构表明存在普遍的因子混杂时,LUCID会衰减因子主导的变异,并从由此得到的创新(innovations)中恢复同期(滞后为0)结构,其边选择通过与数据驱动的无边零假设进行校准。该方法并不绑定于特定的发现算法,而是可以包裹现有的发现引擎;我们在三种此类方法上展示了一致的改进。在一个涵盖混杂因子强度与稀疏性、载荷密度、滞后结构、波动率动态、边异质性、持续性、间歇性以及尾部行为等多种变化的大规模合成分布外基准上,LUCID取得了最优的家族加权有向、滞后分辨图F1分数(0.60),相较最强基线绝对提升0.19(相对提升约46%)。相对于更宽松的滞后折叠评分,其优势进一步扩大,并且在间歇性混杂和重尾混杂下仍保持稳健。复现该方法、基准生成器以及所有报告实验的代码可在https://github.com/bloomberg/causal-ts获取。
cs.LG / 103 / 2609.31325

More Sensors Only One Field: Rethinking Continual Spatio-Temporal Forecasting

更多传感器,同一区域:重新思考持续时空预测
Xie, Lewei, Zhang, Haoyu, Zhou, Jiajun, Chen, Yulong, Chen, Guanxing, Huang, Yu-An, Wong, Hau-San, Zhang, Yifan, Huang, Zhi-An
Abstract
Continual spatio-temporal forecasting supports traffic management and environmental monitoring under evolving dynamics and expanding sensor networks. However, conventional graph-based continual learning methods tie forecasting representations to the current sensor layout, so sensor expansion can alter the representation of learned spatial relationships. Our key insight is that sensor expansion changes the evidence available about a process without necessarily changing the dynamics to be learned. We propose STFO (Spatio-Temporal Field Operator), which parameterizes forecasting knowledge as a shared field-evolution operator and handles changing sensor layouts through observation and query interfaces. Normalized coordinate-based aggregation lifts irregular sensor histories onto a fixed latent grid, enabling reuse of learned spatial maps across observation sets without sensor-specific parameters. To accommodate process drift, a spectral descriptor summarizes variation across spatial scales and conditions Fourier propagation and attention to adapt operator responses to the current spatial regime. Coordinate-based decoding queries the evolved field at sensor locations and combines spatial corrections with local-history predictions. Experiments on PEMS-Stream, CA-Stream, and AIR-Stream demonstrate state-of-the-art average forecasting performance. STFO-Large reduces average MAE over DOL by 8.4% on PEMS-Stream and 4.7% on CA-Stream. Our code is available at https://github.com/Xielewei/Spatio-Temporal-Field-Operator.
Chinese Translation
持续时空预测为动态演变和传感器网络不断扩展下的交通管理与环境监测提供支持。然而,传统的基于图的持续学习方法将预测表示与当前传感器布局绑定,导致传感器扩展可能改变已学习空间关系的表示。我们的核心洞察是:传感器扩展改变的是关于某一过程可获得的信息,而不一定改变需要学习的动态特性。我们提出 STFO(Spatio-Temporal Field Operator,时空场算子),它将预测知识参数化为一个共享的场演化算子,并通过观测接口与查询接口来应对变化的传感器布局。基于归一化坐标的聚合方法将不规则的传感器历史数据映射到固定的潜在网格上,使已学习的空间映射能够在不同观测集合之间复用,而无需针对特定传感器的参数。为适应过程漂移,谱描述子总结了跨空间尺度的变化,并通过条件化的傅里叶传播与注意力机制使算子响应适应当前的空间状态。基于坐标的解码在传感器位置查询演化后的场,并将空间修正与局部历史预测相结合。在 PEMS-Stream、CA-Stream 和 AIR-Stream 数据集上的实验表明,该方法取得了最先进的平均预测性能。STFO-Large 在 PEMS-Stream 上将相对于 DOL 的平均 MAE 降低了 8.4%,在 CA-Stream 上降低了 4.7%。我们的代码发布于 https://github.com/Xielewei/Spatio-Temporal-Field-Operator。
cs.LG / 104 / 2609.31329

Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design

连接身体与大脑:基因驱动的形态—控制协同设计
Feng, Fu, Shi, Ruixiao, Xie, Yucheng, Wang, Jing, Geng, Xin
Abstract
Morphology--control co-design jointly optimizes an agent's body structure and control policy as an integrated embodied system. However, existing methods typically model morphology design and control with separate networks coupled only indirectly through a shared task objective, limiting explicit high-level coordination. Inspired by natural genes that coordinate biological development, we introduce \textbf{Morphogene}, a compact latent blueprint that bridges an agent's body and brain. Through AdaConcat, Morphogene jointly conditions morphology and control generation at the limb level, allowing its variations to induce coordinated changes in both components. Building on this representation, we propose \textbf{GeCode}, which formulates co-design as exploration in the compact Morphogene space. Each Morphogene anchors a local design region in which nearby body--brain designs are explored, while performance-guided updates move these anchors toward promising regions for more efficient exploration of the broader design space. This process combines local refinement with global exploration while preserving body--brain compatibility. Extensive experiments across diverse 2D and 3D co-design tasks demonstrate that GeCode consistently outperforms existing state-of-the-art methods, achieving substantially faster convergence and higher final performance.
Chinese Translation
形态—控制协同设计将智能体的身体结构与控制策略作为一个整体的具身系统进行联合优化。然而,现有方法通常采用相互独立的网络分别建模形态设计与控制,二者仅通过共享的任务目标间接耦合,限制了显式的高层协调。受自然界中基因协调生物发育的机制启发,我们提出了 extbf{Morphogene},一种连接智能体身体与大脑的紧凑潜在蓝图。通过 AdaConcat 机制,Morphogene 在肢体层面同时对形态与控制生成进行条件化,使其变化能够引发两个组件的协调改变。基于这一表示,我们提出 extbf{GeCode},将协同设计形式化为在紧凑的 Morphogene 空间中的探索。每个 Morphogene 锚定一个局部设计区域,在其中探索邻近的身体—大脑设计,同时以性能为导向的更新将这些锚点向有前景的区域移动,从而更高效地探索更广阔的设计空间。该过程在保持身体—大脑兼容性的同时,将局部精炼与全局探索相结合。在多种二维和三维协同设计任务上的大量实验表明,GeCode 持续优于现有最先进方法,实现了显著更快的收敛速度和更高的最终性能。
cs.LG / 105 / 2609.31351

Progressive Memory Transformer: Memory-Aware Attention for Time-Series

渐进式记忆Transformer:面向时间序列的记忆感知注意力机制
Stangeland, Tord Sture, Köhler, Andreas, Mæland, Steffen, Rivera, Adín Ramíres
Abstract
Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise representations globally via instance-level contrastive losses and limited temporal neighborhood supervision, but do not explicitly exploit the structural hierarchy. We propose a learning framework that explicitly enforces a structural hierarchy across three scales independently: a local objective for token continuity, a mid-range objective for window-level motifs, and a global objective for sequence-level agreement. Realizing this framework requires the backbone to expose a representation at each scale; we introduce \textbf{Progressive Memory Transformer} (PMT), which augments a transformer with writable, window-aligned memory that exposes the mid-range scale alongside the token and sequence-level representations conventional transformers already provide. Across seven UCR/UEA/UCI classification benchmarks, a cue-retention probe, and forecasting benchmarks, PMT learns representations that probe well at the global, mid-range, and local scales---strong low-label classification (1--5\% labels), competitive forecasting performance across multiple horizons, and quantitative and qualitative evidence that memory states capture mid-range motifs.
Chinese Translation
时间序列同时在多个尺度上携带结构信息(细粒度变化、中尺度模式以及全局特性),而下游任务也在相应不同的尺度上进行。大多数现有的自监督学习方法通过实例级对比损失和有限的时间邻域监督来对表示进行全局监督,但并未显式地利用这种结构层次。我们提出了一个学习框架,该框架在三个尺度上独立地显式强化结构层次:面向词元连续性的局部目标、面向窗口级模式的中尺度目标,以及面向序列级一致性的全局目标。实现该框架需要骨干网络能够在每个尺度上提供表示;我们提出了渐进式记忆Transformer(Progressive Memory Transformer, PMT),它在Transformer基础上增加了可写的、与窗口对齐的记忆机制,从而在传统Transformer已提供的词元级和序列级表示之外,进一步暴露中尺度表示。在七个UCR/UEA/UCI分类基准、一个线索保留探测任务以及预测基准上的实验表明,PMT学习到的表示在全局、中尺度和局部尺度上均具有良好的探测性能——在低标签场景(1–5%标签)下表现出强大的分类能力,在多个预测时间跨度上具有竞争力的预测性能,并且定量和定性证据表明记忆状态能够捕获中尺度模式。
cs.LG / 106 / 2609.31363

Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning

当Brenier遇上对抗训练:面向鲁棒学习的最优传输几何
Abdollahpoorrostam, Alireza, Sharifian, Ehsan, Şen, Buse, Cuturi, Marco, Kuhn, Daniel
Abstract
Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulation in which the adversary may choose any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. We show that the adversary's problem can be reformulated as an optimization problem over transport maps that push empirical samples to adversarial ones, and we prove that optimal maps are cyclically monotone. We also show that standard adversarial training---based on per-sample local optimization---violates cyclical monotonicity and wastes transport costs unless the adversary is severely restricted. We propose two remedies. First, we introduce multi-start particle ascent, which alternates parallel gradient ascent with reassignment to enforce cyclical monotonicity across samples. Second, we parameterize adversarial maps as gradients of input-convex neural networks, which guarantees cyclical monotonicity by construction. Experiments on robust regression, image classification, and robust control show that our methods consistently outperform standard adversarial training and state-of-the-art baselines, achieving improved robustness and better generalization under distribution shift.
Chinese Translation
分布鲁棒优化(DRO)为分布偏移下的学习提供了一个有原则的框架,但由于难以对非凸损失函数评估最坏情况风险,其在实践中的应用受到阻碍。我们研究了一种带惩罚项的DRO形式,其中对抗者可以选择任意分布,但需为偏离经验分布支付Wasserstein惩罚。我们证明,对抗者的问题可以重新表述为在传输映射上的优化问题,即寻找将经验样本推至对抗样本的映射,并且我们证明最优映射是循环单调的。我们还表明,标准的基于逐样本局部优化的对抗训练会违反循环单调性并浪费传输代价,除非对对抗者施加严格限制。我们提出两种补救方法。首先,我们引入多起点粒子上升(multi-start particle ascent)方法,该方法交替进行并行梯度上升与重新分配,以在样本之间强制实现循环单调性。其次,我们将对抗映射参数化为输入凸神经网络(input-convex neural networks)的梯度,从而从结构上保证循环单调性。在鲁棒回归、图像分类和鲁棒控制上的实验表明,我们的方法在性能上持续超越标准对抗训练和最先进的基线方法,在分布偏移下实现了更强的鲁棒性和更好的泛化能力。
cs.LG / 107 / 2609.31371

Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness

面向理解基于大语言模型的日志异常检测:性能、效率与鲁棒性的实证研究
Li, Bin, Wang, Dongdong, Lu, Siyang
Abstract
Large language models (LLMs) have demonstrated promising performance in log anomaly detection, yet how their adaptation strategies, architectures, and deployment configurations affect detection effectiveness remains insufficiently understood. To investigate these factors, we conduct a systematic empirical analysis across three public log datasets, examining different adaptation strategies, model architectures, parameter scales, and quantization settings. Our results reveal substantial performance differences across adaptation strategies, while model scaling yields varying detection gains across datasets. We further observe that models with comparable detection accuracy can exhibit markedly different computational costs, and that low-bit quantization largely preserves detection performance in the evaluated configurations. Finally, we examine detection robustness under structural, semantic, and label noise at different perturbation levels. These findings provide empirical insights into the performance, efficiency, and robustness of LLM-based log anomaly detection, highlighting practical considerations beyond conventional accuracy-oriented evaluation.
Chinese Translation
大语言模型(LLMs)在日志异常检测中展现出良好的性能,但其适配策略、模型架构和部署配置如何影响检测效果仍缺乏充分的理解。为探究这些因素,我们在三个公开日志数据集上开展了系统性的实证分析,考察了不同的适配策略、模型架构、参数规模以及量化设置。我们的结果揭示了不同适配策略之间存在显著的性能差异,而模型规模的扩展在不同数据集上带来的检测性能提升各不相同。我们进一步观察到,检测精度相近的模型可能表现出明显不同的计算成本,并且在所评估的配置下,低比特量化在很大程度上能够保持检测性能。最后,我们考察了在不同扰动程度下,模型在结构、语义和标签噪声下的检测鲁棒性。这些发现为基于大语言模型的日志异常检测的性能、效率和鲁棒性提供了实证见解,突出了超越传统仅关注准确率评估的实际考量因素。
cs.LG / 108 / 2609.31399

Differential Attention Unlocks Complementary EEG and Speech Fusion for Emotion Recognition

差分注意力解锁脑电与语音的互补融合以实现情感识别
Lee, Philip H., Chandra, Shreeram Suresh, Hansen, John H. L.
Abstract
Multimodal emotion recognition (MER) increasingly pairs EEG with speech, treating internal neural signals and external vocal expression as informative views of affect. In practice, naive fusion underperforms the stronger single modality, because EEG artifacts inject noise that corrupts the shared representation. We introduce EmoSpeechBrain, a multimodal framework built on the insight that noise suppression is a precondition for effective fusion. Its EEG encoder uses differential attention, taking the difference between two attention maps to cancel shared noise and isolate discriminative neural activity. An attention-based gating adapter aligns both modalities in a shared space and weights each one's contribution to the prediction. On two datasets - PME4 and EAV, EmoSpeechBrain improves MER accuracy by up to 12.9% over other state-of-the-art (SOTA) EEG encoders, and surpasses unimodal speech and EEG baselines by up to 13.1% and 23.1%. These results show that once EEG noise is suppressed, fusion delivers gains that naive combination cannot.
Chinese Translation
多模态情感识别(MER)越来越多地将脑电(EEG)与语音相结合,将内部神经信号和外部声音表达视为情感的两个信息视角。然而在实践中,简单的融合往往不如较强的单模态,这是因为脑电伪迹引入的噪声会破坏共享表征。我们提出了 EmoSpeechBrain,这是一个多模态框架,其核心洞见在于:噪声抑制是实现有效融合的前提条件。其脑电编码器采用差分注意力(differential attention),通过计算两个注意力图之间的差值来抵消共享噪声,从而分离出具有判别性的神经活动。一个基于注意力的门控适配器将两种模态对齐到共享空间中,并为每个模态对预测的贡献进行加权。在 PME4 和 EAV 两个数据集上,EmoSpeechBrain 将多模态情感识别的准确率相比其他最先进(SOTA)的脑电编码器最高提升了 12.9%,相比单模态语音和脑电基线分别最高超出 13.1% 和 23.1%。这些结果表明,一旦脑电噪声得到抑制,融合便能带来简单组合所无法实现的收益。
cs.LG / 109 / 2609.31401

Decodable In-Context State and Model Output Across Training

跨训练阶段可解码的上下文状态与模型输出
Ravulapalli, Manas Venkata Sai, Chadha, Samrath Singh
Abstract
Prior work established that a probe can decode an in-context binding on model errors and that probe-guided steering can repair some of them. We follow probe accuracy, model output, and steering response across public pretraining and post-training checkpoints. Probe accuracy rises during Pythia pretraining, while probe-guided steering moves from negligible all-trial benefit to a larger benefit at two model sizes. Saved scores distinguish probe-correct errors with low and above-uniform model probability for the correct candidate. Oracle-target steering already repairs many early errors, but saved aggregates cannot separate target quality from intervention sensitivity. A held-out comparison of decoders trained on the final state or candidate logits finds no detected final-state advantage on late-checkpoint model errors. An information-theoretic counterexample explains why decodability on errors alone cannot establish discarded output information. The connection to downstream omissions remains open.
Chinese Translation
先前的研究表明,探针可以在模型出错时解码上下文中的绑定关系,并且基于探针引导的干预可以修复其中部分错误。我们在公开的预训练和后训练检查点上追踪探针准确率、模型输出及引导干预响应的变化。在Pythia预训练过程中,探针准确率持续上升,而基于探针引导的干预在两个模型规模上从几乎无的全试次收益转变为更大的收益。保存的评分可以区分探针判定的错误中,正确候选项模型概率较低和高于均匀概率的两类情况。以真值目标进行的引导干预已经能够修复许多早期错误,但保存的聚合指标无法将目标质量与干预敏感性区分开来。对在最终状态或候选项logits上训练的解码器进行的留出对比实验发现,在后期检查点的模型错误上,最终状态并未表现出可检测的优势。一个信息论反例解释了为什么仅凭错误样本上的可解码性无法证明输出信息被丢弃。其与下游遗漏现象的联系仍有待研究。
cs.LG / 110 / 2609.31415

Evaluating the accuracy of KV cache reuse techniques

评估KV缓存复用技术的准确性
Cestola, Samuel, Xia, Tianxiang, Zheng, Pengfei, Zheng, Weiyan, Wang, Bo, Zhao, Yi, Didona, Diego
Abstract
Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fail to faithfully capture the loss of accuracy attributable to reuse, often artificially inflating the reported effectiveness. We also show that existing datasets do not exhibit the reuse dynamics needed to thoroughly evaluate such techniques. To address these issues, we propose an evaluation methodology that measures this accuracy loss without ambiguity and we introduce Boxoffice, a tool that programmatically generates evaluation datasets that exercise challenging KV cache reuse patterns.
Chinese Translation
位置无关的KV缓存复用旨在通过跨提示词复用块级(chunk-level)KV缓存来降低检索增强生成的延迟。我们的研究表明,目前对KV缓存复用技术的评估所依赖的度量方法无法忠实地反映由复用导致的准确性损失,往往会人为地夸大其报告的有效性。我们还发现,现有数据集并不具备彻底评估此类技术所需的复用动态特征。为解决这些问题,我们提出了一种能够无歧义地度量该准确性损失的评估方法,并推出了Boxoffice——一个可通过编程方式生成评估数据集的工具,这些数据集能够涵盖具有挑战性的KV缓存复用模式。
cs.LG / 111 / 2609.31454

Different Corruptions, Different Signals: Uncertainty and Loss in Federated Data Quality

不同损坏、不同信号:联邦数据质量中的不确定性与损失
Scott, Bradley, Luo, Zeqi, Ho, Edmond S. L.
Abstract
Federated learning (FL) data corruption can affect either inputs or labels, but it remains unclear whether input-conditional uncertainty and prediction-label loss expose these corruption modes equally. This paper compares two corruption-detection signals in FL: input-conditional uncertainty and prediction-label loss. The uncertainty signal is characterised using a learned aleatoric variance estimate together with Monte Carlo (MC) dropout variance and entropy measures, while the loss is computed against the supplied label. We test these signals against additive image noise and persistent random label flips. On ResNet-20 with CIFAR-10 and SVHN under Dirichlet partitions with data that are not independent and identically distributed (non-IID), the two corruption types behave differently. For persistent random label flips, the within-client per-sample area under the receiver operating characteristic curve (AUC) is 0.85 on CIFAR-10 and 0.95 on SVHN for prediction-label loss, while every uncertainty estimator stays at chance (0.49--0.50). This pattern is consistent with the model remaining confident in the underlying image despite the supplied label being wrong. For image noise, expected-entropy uncertainty rises above chance (0.67 on CIFAR-10 and 0.66 on SVHN), while loss responds comparably (0.64 on both). Each signal is therefore the stronger detector for a different corruption: the prediction-label loss for persistent label flips, and expected-entropy uncertainty for image noise, with its advantage becoming apparent as federation-wide corruption prevalence increases. Robust FL data-quality assessment should match the signal to the corruption rather than rely on uncertainty alone across corruption types.
Chinese Translation
联邦学习(FL)中的数据损坏可能发生在输入或标签上,但输入条件不确定性与预测-标签损失是否能够同等地揭示这些损坏模式仍不清楚。本文比较了联邦学习中的两种损坏检测信号:输入条件不确定性和预测-标签损失。不确定性信号通过学习得到的偶然不确定性方差估计,结合蒙特卡洛(MC)Dropout 方差和熵度量来刻画,而损失则相对于所提供的标签进行计算。我们在加性图像噪声和持续性随机标签翻转两种情形下测试这些信号。在采用狄利克雷(Dirichlet)划分的非独立同分布(non-IID)数据下,基于 ResNet-20 在 CIFAR-10 和 SVHN 上的实验表明,两种损坏类型表现出不同的特性。对于持续性随机标签翻转,预测-标签损失在客户端内的逐样本受试者工作特征曲线下面积(AUC)在 CIFAR-10 上为 0.85,在 SVHN 上为 0.95,而所有不确定性估计器都停留在随机水平(0.49--0.50)。这一模式与模型在所提供标签错误的情况下仍对底层图像保持自信是一致的。对于图像噪声,期望熵不确定性高于随机水平(CIFAR-10 上为 0.67,SVHN 上为 0.66),而损失的响应相当(两者均为 0.64)。因此,每种信号对不同损坏类型的检测能力不同:预测-标签损失更擅长检测持续性标签翻转,而期望熵不确定性更擅长检测图像噪声,且随着联邦范围内损坏流行度的增加,后者的优势愈加明显。鲁棒的联邦学习数据质量评估应将信号与损坏类型相匹配,而不是在所有损坏类型上一概依赖不确定性。
cs.LG / 112 / 2609.31463

Uncertainty-Aware Federated Learning for Infant Movement Analysis

面向婴儿运动分析的不确定性感知联邦学习
Ho, Edmond S. L.
Abstract
Infant movement analysis provides valuable biomarkers for the early identification of neurodevelopmental disorders. Recent advances in deep learning have enabled automated analysis of infant movements from video-derived skeletal representations, achieving performance comparable to expert assessment for tasks such as General Movement Assessment (GMA). However, most existing approaches rely on centralized training, requiring data from multiple institutions to be collected and stored at a single site. Such assumptions are often impractical in clinical settings due to privacy, governance, and data-sharing constraints. To address these challenges, we present, to the best of our knowledge, the first federated learning framework for automated infant movement analysis and General Movement Assessment using skeletal motion data. As a clinically relevant use case, the proposed framework is evaluated on fidgety movement classification. To quantify model confidence, Monte Carlo (MC) Dropout is employed to estimate predictive uncertainty during inference. Building upon this, we propose an Uncertainty-Aware Federated Averaging (UA-FedAvg) strategy that incorporates predictive entropy derived from MC-Dropout into the federated aggregation process, enabling client contributions to be adjusted according to their predictive uncertainty. Experiments were conducted using a cross-subject evaluation protocol under a three-client federated learning setting. Results demonstrate that federated learning substantially improves classification performance compared with independently trained local models while achieving performance approaching that of centralized training. Furthermore, UA-FedAvg and its variant incorporating validation loss generally outperform conventional FedAvg across the evaluated data-split configurations.
Chinese Translation
婴儿运动分析为神经发育障碍的早期识别提供了有价值的生物标志物。深度学习的最新进展使得基于视频提取的骨骼表示对婴儿运动进行自动化分析成为可能,在全身运动评估(General Movement Assessment, GMA)等任务上达到了与专家评估相当的性能。然而,现有的大多数方法依赖于集中式训练,需要将来自多个机构的数据收集并存储在单一站点。由于隐私、治理和数据共享方面的限制,此类假设在临床环境中往往不切实际。为应对这些挑战,我们据其所知首次提出了一个基于骨骼运动数据的自动化婴儿运动分析与全身运动评估的联邦学习框架。作为一个具有临床相关性的应用案例,该框架在烦躁样运动(fidgety movement)分类任务上进行了评估。为量化模型置信度,采用蒙特卡洛Dropout(Monte Carlo Dropout, MC Dropout)在推理阶段估计预测不确定性。在此基础上,我们提出了一种不确定性感知联邦平均(Uncertainty-Aware Federated Averaging, UA-FedAvg)策略,将从MC-Dropout得到的预测熵融入联邦聚合过程,使各客户端的贡献可根据其预测不确定性进行调整。实验采用跨被试评估协议,在三个客户端的联邦学习设置下进行。结果表明,与独立训练的本地模型相比,联邦学习显著提升了分类性能,同时达到了接近集中式训练的性能。此外,在所评估的各种数据划分配置下,UA-FedAvg及其融合验证损失的变体总体上优于传统的FedAvg。
cs.LG / 113 / 2609.31466

Scaffold: Support Graph Theory Based Sparsification for Graph Neural Networks

Scaffold:基于支撑图理论的图神经网络稀疏化方法
Das, Siddhartha Shankar, Navuluru, Sai Karthik, Ferdous, S M, Rossi, Ryan A., Coskunuzer, Baris, Tamil, Lakshman, Serra, Edoardo, Pothen, Alex, Rallo, Robert, Halappanavar, Mahantesh M
Abstract
Graph neural networks (GNNs) rely on message passing over graph edges, making their computational and memory costs strongly dependent on graph density. Graph sparsification offers a natural way to reduce these costs, but removing edges indiscriminately can distort important communication structure and degrade predictive performance. We introduce Scaffold, a topology-based, unsupervised graph sparsification framework derived from support graph theory preconditioners. Scaffold explicitly controls two complementary structural quantities: dilation, which measures the length of rerouting paths induced by removed edges, and congestion, which measures how strongly these rerouted paths concentrate on the retained support. By jointly controlling dilation and congestion, Scaffold preserves short communication paths while avoiding structural bottlenecks. To our knowledge, Scaffold is the first scalable GNN sparsification framework to use a joint supporting-path dilation-congestion criterion. Across 19 homophilic and heterophilic benchmarks spanning small to large graphs, Scaffold achieves the best aggregate rank among the evaluated sparsification and related methods. Using only 10%-50% of the original edges per sparse support, Scaffold recovers or closely approaches full-graph GNN performance while using less than half the memory of full-graph training and reducing end-to-end training time, including sparsification overhead. We provide an open-source software package at https://github.com/siddhartha047/Scaffold.
Chinese Translation
图神经网络(GNN)依赖于图边上的消息传递,使其计算和内存开销与图的密度密切相关。图稀疏化是降低这些开销的自然途径,但无差别地删除边可能会破坏重要的通信结构并降低预测性能。我们提出了 Scaffold,一个基于拓扑结构的无监督图稀疏化框架,其思想源自支撑图理论预条件子。Scaffold 显式地控制两个互补的结构量:延展度(dilation),用于衡量被删除边所引起的绕行路径的长度;以及拥塞度(congestion),用于衡量这些绕行路径在保留支撑上的集中程度。通过联合控制延展度与拥塞度,Scaffold 在保持短通信路径的同时避免了结构性瓶颈。据我们所知,Scaffold 是首个采用联合支撑路径延展度-拥塞度准则的可扩展 GNN 稀疏化框架。在涵盖从小型到大型图的 19 个同配与异配基准数据集上,Scaffold 在所评估的稀疏化及相关方法中取得了最优的总体排名。在每个稀疏支撑仅使用原始边 10%-50% 的情况下,Scaffold 能够恢复或接近全图 GNN 的性能,同时内存占用不到全图训练的一半,并降低了包括稀疏化开销在内的端到端训练时间。我们已在 https://github.com/siddhartha047/Scaffold 提供开源软件包。
cs.LG / 114 / 2609.31531

HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning

HySTAR:基于锚定超图的协作式多智能体强化学习稳定信用分配方法
Luo, Xinglong, Zhang, Yuding, Kuang, Yuheng, Yuan, Shuxuan, Zeng, Zhenni, Zhu, Weiqiang, Ji, Zhenhai, Wang, Zhengning
Abstract
Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics that dynamically reconstruct the grouping topology change the mapping from agents and coalitions to value components as interactions or active agents evolve. We refer to this inconsistency as structural target drift. We introduce HySTAR, a MAPPO-based framework that separates adaptive representation learning from a temporally consistent high-order value-decomposition basis. HySTAR anchors an overlapping sparse hypergraph as a uniformly covered decomposition scaffold, uses a spatiotemporal encoder to represent physical and task-dependent interactions, and combines temporal and structural relevance to construct agent-specific advantages. Experiments on SMAC, GRF, Traffic Junction, and MPE demonstrate consistent improvements over MAPPO-style, value-factorization, and dynamic-grouping baselines. On the hardest SMAC settings, HySTAR achieves relative gains of 16.7\% over MAPPO and 15.6\% over HYGMA, ranks first on all six GRF scenarios, reduces Traffic Junction convergence epochs by up to 40.2\% relative to MAGIC, and obtains the highest MPE episode rewards. Controlled topology, agent-death, neighborhood, and parameter analyses support the benefit of anchoring the decomposition scaffold while adapting the propagated representations.
Chinese Translation
在部分可观测和共享奖励条件下的协作式多智能体强化学习,需要将团队整体收益分配到个体智能体及高阶联盟。MAPPO 风格的评论家(critic)将联合行为压缩为单一的全局价值,而动态重构分组拓扑的评论家则会随着交互或活跃智能体的演变而改变从智能体与联盟到价值分量的映射关系。我们将这种不一致性称为结构性目标漂移(structural target drift)。我们提出 HySTAR,这是一种基于 MAPPO 的框架,它将自适应表征学习与时间一致的高阶价值分解基相分离。HySTAR 锚定一个重叠稀疏超图作为均匀覆盖的分解支架,使用时空编码器表征物理性和任务相关的交互,并结合时间相关性与结构相关性来构造智能体专属的优势函数(advantage)。在 SMAC、GRF、Traffic Junction 和 MPE 上的实验表明,该方法相对 MAPPO 风格、价值分解及动态分组基线取得了一致的性能提升。在最困难的 SMAC 设置中,HySTAR 相对 MAPPO 取得 16.7% 的相对提升,相对 HYGMA 取得 15.6% 的提升,在全部六个 GRF 场景中均排名第一,在 Traffic Junction 上相比 MAGIC 最多减少 40.2% 的收敛轮次,并在 MPE 上获得最高的回合奖励。受控的拓扑、智能体死亡、邻域及参数分析支持了在自适应传播表征的同时锚定分解支架的有效性。
cs.LG / 115 / 2609.31539

NEXT: Physics-Informed Neuro-Spectral Exponential Time Differencing Architectures

NEXT:物理信息神经谱指数时间差分架构
Marques, Márcio, Mendonça, Leonardo, Moreira, Leonardo M., de Oliveira, Christian Júnior, Balestro, Vitor, Novello, Tiago, Yukimura, Daniel, Petrov, Pavel, Nissenbaum, Lucas
Abstract
Physics-Informed Neural Networks (PINNs) build neural representations of time-dependent PDE solutions, naturally incorporating physics knowledge and observational data, which makes them well suited to both forward and inverse PDE problems. PINNs, however, are known to suffer from spectral bias and lack of causality. Neuro-Spectral Architectures (NeuSA), a recently proposed alternative to PINNs, mitigate both issues, but their numerical integration becomes unstable for stiff differential equations arising in many relevant physical problems. This study proposes Neuro-Spectral Exponential Time Differencing Architectures (NEXT), which combines the spectral representation of the PDE solution in NeuSA with high-order exponential integrators. Within this approach, the linear stiff part of the vector field induced by the PDE is integrated exactly through matrix exponentials, while the possibly nonlinear remainder is modeled by a neural network. The effectiveness of NEXT is verified through benchmark experiments on a set of stiff PDEs, in which NEXT is stable and accurate while NeuSA diverges numerically. It is also shown that NEXT can be applied to inverse problems, where the model has to learn unknown parameters or boundary conditions from sparse data. All code used in this work is publicly available at: https://github.com/marcioh2m/next.git .
Chinese Translation
物理信息神经网络(Physics-Informed Neural Networks, PINNs)通过神经网络表示时间相关的偏微分方程(PDE)解,能够自然地融合物理知识与观测数据,因此非常适合求解PDE的正问题和反问题。然而,PINNs存在众所周知的谱偏差(spectral bias)和因果性缺失的问题。神经谱架构(Neuro-Spectral Architectures, NeuSA)是最近提出的一种PINNs替代方法,可以同时缓解这两个问题,但其数值积分在处理许多相关物理问题中出现的刚性微分方程时会变得不稳定。本研究提出了神经谱指数时间差分架构(Neuro-Spectral Exponential Time Differencing Architectures, NEXT),该方法将NeuSA中PDE解的谱表示与高阶指数积分器相结合。在这一方法中,由PDE诱导的向量场中的线性刚性部分通过矩阵指数进行精确积分,而可能非线性的剩余部分则由神经网络建模。通过在一组刚性PDE上的基准实验验证了NEXT的有效性:当NeuSA出现数值发散时,NEXT仍能保持稳定和精确。研究还表明,NEXT可应用于反问题,即模型需要从稀疏数据中学习未知参数或边界条件。本文使用的所有代码已在以下网址公开:https://github.com/marcioh2m/next.git。
cs.LG / 116 / 2609.31546

BeatGraph: Self-Supervised Heartbeat Graphs for Infant ECG Representations from the Home Environment

BeatGraph:面向家庭环境下婴儿心电图表征的自监督心跳图模型
Khan, Mohammad Nur Hossain, Krafczyk, M. S., Bolster, Beverly G., McElwain, Nancy, Hasegawa-Johnson, Mark A., Islam, Bashima
Abstract
Electrocardiogram (ECG) foundation models typically tokenize the signal into fixed-length patches that ignore cardiac structure, so a patch may split a heartbeat and the number of beats in each patch shifts with heart rate. This matters most for infants, whose heart rates are higher and whose ECG differs from the adult, clinic-recorded 12-lead data these models are built on. A model for infant ECG should therefore reason about heartbeats directly rather than recover them from arbitrary patches. We propose BeatGraph, which makes the heartbeat its unit of representation, modeling each 30-second window as a graph of beats. A shared beat encoder embeds each heartbeat from its waveform and inter-beat intervals, a Transformer with positional encoding orders the beats in time, and residual graph attention layers relate every beat to every other before attention pooling yields a window embedding. We pretrain BeatGraph on our new corpus of unlabeled infant recordings by predicting masked-beat embeddings, then fine-tune it for each task. One backbone supports sleep-wake detection, infant-state classification, activity-source identification (infant- or caregiver-initiated movement), and affect recognition, improving macro-F1 over the strongest baseline on each task by 0.076 to 0.158. It also transfers across age groups, reaching 0.892 AUROC on the ZZU-pECG pediatric benchmark (ages 0 to 14), within 0.001 of the best published self-supervised ECG model, and matching that model under linear evaluation on the adult PTB-XL benchmark despite infant-only pretraining. Finally, to our knowledge, we release the first public infant ECG corpus collected in homes, classrooms, and laboratory settings with state and affect labels. It contains 3,408 hours of single-channel ECG from 143 infants aged 3 to 11 months, with unlabeled pretraining data, benchmark tasks, and subject-level splits.
Chinese Translation
心电图(ECG)基础模型通常将信号切分为固定长度的片段(patch)进行词元化,而忽略了心脏结构,因此一个片段可能将一次心跳截断,且每个片段中的心跳数量会随心率变化而改变。这一问题对婴儿尤为重要:婴儿的心率更高,且其心电图与这些模型所基于的成人临床记录的12导联数据存在差异。因此,婴儿心电图模型应当直接对心跳进行建模,而非从任意片段中恢复心跳。我们提出 BeatGraph,将心跳作为基本表征单元,把每个30秒的时间窗口建模为一张由心跳构成的图。一个共享的心跳编码器根据每个心跳的波形及其与相邻心跳的间期对其进行嵌入,随后带有位置编码的 Transformer 对心跳进行时间排序,再通过残差图注意力层建立每个心跳与其他所有心跳之间的关联,最后经注意力池化得到窗口嵌入。我们在新构建的无标注婴儿 recordings 语料库上,通过预测被遮蔽心跳的嵌入对 BeatGraph 进行预训练,然后针对每个任务进行微调。该单一骨干网络可支持睡眠-觉醒检测、婴儿状态分类、活动来源识别(婴儿自发或看护者引发的移动)以及情感识别,在每项任务上的宏平均 F1(macro-F1)较最强基线提升了0.076至0.158。该模型还具有良好的跨年龄组迁移能力,在 ZZU-pECG 儿科基准(0至14岁)上达到0.892 AUROC,与已发表的最好的自监督 ECG 模型相差不足0.001;并且仅使用婴儿数据进行预训练的情况下,在成人 PTB-XL 基准的线性评估中与该模型表现相当。最后,据我们所知,我们发布了首个在家庭、教室和实验室环境中采集并带有状态与情感标注的公开婴儿心电图语料库。该语料库包含143名3至11个月婴儿的单导联心电图,共3,408小时,涵盖无标注预训练数据、基准任务以及按受试者划分的数据集切分。
cs.LG / 117 / 2609.31559

Online Learning via Learned Latent Bayesian Tracking

基于学习的潜在贝叶斯跟踪的在线学习
Gerson, Guy, Raviv, Tomer, Shlezinger, Nir, Routtenberg, Tirza, Simeone, Osvaldo
Abstract
Online learning in non-stationary environments requires models to adapt rapidly from streaming data under strict computational constraints. A principled approach casts online learning as Bayesian state tracking, where model parameters are updated sequentially via Bayesian filtering. However, applying Bayesian filters directly to modern deep models is computationally prohibitive due to the high dimensionality of parameter space, forcing existing methods to rely on restrictive approximations or manually designed low-dimensional subspaces. In this work, we identify the absence of a suitable low-dimensional dynamical representation as the core bottleneck in Bayesian filtering-based online learning. Accordingly, we propose Adaptive Update through Representation Adaptation (AURA), a meta-learning framework that learns offline a low-dimensional latent state-space model governing the evolution of optimal model parameters under distribution shift. Online adaptation is then performed via extended Kalman filtering in this learned latent space followed by reconstruction of the full model parameters through a learned lifting map, enabling efficient single-step online adaptation while preserving model expressiveness. Evaluated on online adaptation of neural wireless receivers under time-varying channels and on non-stationary image classification, AURA shows substantial improvements in adaptation speed, accuracy, and computational efficiency over existing online learning and Bayesian filtering baselines, demonstrating that an adaptation-aware latent geometry is beneficial for effective Bayesian online learning in high-dimensional models.
Chinese Translation
非平稳环境下的在线学习要求模型在严格的计算约束下从流式数据中快速适应。一种有原则的方法将在线学习建模为贝叶斯状态跟踪问题,即通过贝叶斯滤波对模型参数进行顺序更新。然而,由于参数空间维度极高,将贝叶斯滤波直接应用于现代深度模型的计算代价过于高昂,迫使现有方法依赖限制性较强的近似或人工设计的低维子空间。在本工作中,我们指出缺乏合适的低维动态表示是基于贝叶斯滤波的在线学习的核心瓶颈。据此,我们提出了自适应更新方法 AURA(Adaptive Update through Representation Adaptation,通过表示自适应实现自适应更新),这是一个元学习框架,可离线学习一个低维潜在状态空间模型,用以刻画最优模型参数在分布偏移下的演化规律。在线自适应随后在该学习到的潜在空间中通过扩展卡尔曼滤波完成,并通过学习到的提升映射重构完整的模型参数,从而在保持模型表达能力的同时实现高效的单步在线自适应。在时变信道下神经无线接收机的在线自适应以及非平稳图像分类任务上的评估表明,AURA 在自适应速度、准确率和计算效率方面均显著优于现有的在线学习与贝叶斯滤波基线方法,证明了一种面向自适应的潜在几何结构有助于在高维模型中实现有效的贝叶斯在线学习。
cs.LG / 118 / 2609.31560

Generalization behavior of OPTQ and the role of regularization

OPTQ 的泛化行为与正则化的作用
George, Erin, Saab, Rayan
Abstract
Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dataset is as small as possible. We study the performance of OPTQ and a variant algorithm, stochastic OPTQ, in a generalization setting and derive bounds for the expected squared error accrued by the algorithm when a test point is drawn from a fixed distribution. We prove two results. One result relates the generalization error to the error on a calibration dataset comprising independent samples from the same distribution as the test distribution. The other result bounds the generalization error of stochastic OPTQ for all sufficiently nice distributions, regardless of the calibration dataset. In both of these results, the regularization term $\lambda$ plays an important role. We use insights from these results to make a new recommendation for the choice of $\lambda$ and see that this choice of $\lambda$ preforms favorably in experiments when compared to prior recommendations in the literature.
Chinese Translation
大型神经网络可以通过对其权重进行舍入或"量化",将其压缩为可用更少比特表示的数值。OPTQ 是一种量化算法,它逐步对神经网络的权重进行量化,使得在指定校准数据集上的平方量化误差尽可能小。我们研究了 OPTQ 及其变体算法——随机 OPTQ(stochastic OPTQ)——在泛化场景下的性能,并推导了当测试点从固定分布中抽取时,该算法产生的期望平方误差的界。我们证明了两个结果:其一,将泛化误差与由来自与测试分布相同的分布的独立样本构成的校准数据集上的误差联系起来;其二,对于所有足够好的分布,无论校准数据集如何,都给出了随机 OPTQ 泛化误差的界。在这两个结果中,正则化项 $\lambda$ 都起着重要作用。我们利用这些结果中的见解,对 $\lambda$ 的选择提出了新的建议,并看到这一 $\lambda$ 的选择在实验中相比文献中已有的建议表现更优。
cs.LG / 119 / 2609.31564

Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights

权重对编码:在神经网络权重中诱导更小的文法
Tallini, Irene, Solombrino, Daniele, Cazzaniga, Alberto, Rodolà, Emanuele
Abstract
We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and 1.1 accuracy points. The trend extends to different grammar compressors (LZ78, SEQUITUR), over which the networks has not be finetuned against. To our knowledge, this is the first time grammar size has been used as an explicit training objective for network weights.
Chinese Translation
我们证明了神经网络权重可以被显式微调以容纳更小的文法。权重对编码(Weight Pair Encoding,WeightPE)通过将有损的 Re-Pair 压缩器嵌入直通估计器(straight-through estimator)中来实现这一目标。该网络将 int8 权重展平为一个字符串,并在全局 L2 预算内将近似匹配的 Re-Pair 模式调整为完全相等。网络使用重写后的权重进行计算,并通过直通估计器进行训练。与由固定大小条目组成的扁平码本不同,文法提供可变长度的模式,并在更大的模式中层次化地复用这些模式。在 CIFAR-10 上微调的 ViT-B/16 和 ViT-L/16 的 MLP 权重上,WeightPE 生成的 Re-Pair 文法大小分别为等效 int8 QAT(量化感知训练)运行所生成文法的 0.43 倍和 0.38 倍,代价是准确率分别下降 1.9 和 1.1 个百分点。这一趋势可推广至网络未曾针对其进行微调的其他文法压缩器(LZ78、SEQUITUR)。据我们所知,这是首次将文法大小作为网络权重的显式训练目标。
cs.LG / 120 / 2609.31586

Trust Guided Decision Transformer

信任引导的决策Transformer
Gautam, Chainesh, Diddigi, Raghuram Bharadwaj, Kamanchi, Chandramouli, Dayama, Pankaj, Mukherjee, Sumanta, Sampath, Kameshwaran
Abstract
Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes using rolling next state prediction error, calibrated against held out offline data via split conformal prediction. It keeps only suffixes whose error stays within the calibrated threshold, then uses a frozen critic to choose the highest value action among the trusted suffixes. This reverses the order used by value only elastic selection, where the critic may choose an action generated from a context the model itself has flagged as unreliable. Experiments on D4RL navigation and locomotion tasks show that state prediction, critic guidance, and hard context reset each solve only part of the problem. TGDT reduces persistent high error runs and improves return over vanilla Decision Transformer, reset based context control, and value only context selection.
Chinese Translation
决策Transformer(Decision Transformer)在长序列推理(rollout)中性能会下降,因为其条件上下文逐渐偏离训练分布。我们证明这种偏离可以通过模型自身的下一状态预测误差来观察:该误差在推理过程中上升并保持高位,从而为上下文何时变得不可靠提供了直接信号。我们提出信任引导决策Transformer(Trust Guided Decision Transformer,TGDT),它在应用价值引导之前先选择上下文。在每一步,TGDT使用滚动下一状态预测误差评估若干近期的上下文后缀,并通过分割保形预测(split conformal prediction)基于留出的离线数据进行校准。它仅保留预测误差处于校准阈值内的后缀,然后使用冻结的评论家(critic)在可信后缀中选择价值最高的动作。这逆转了仅基于价值的弹性选择方法所使用的顺序——在后者中,评论家可能选择由模型自身已标记为不可靠的上下文所生成的动作。在D4RL导航和运动控制任务上的实验表明,状态预测、评论家引导和硬性上下文重置各自只能解决部分问题。TGDT减少了持续高误差的运行情况,并在回报上优于原始决策Transformer、基于重置的上下文控制以及仅基于价值的上下文选择方法。
cs.LG / 121 / 2609.31589

Common-Mode Collapse and Recovery in Direct Feedback Alignment

直接反馈对齐中的共模崩塌与恢复
Reddy, Varun, Sabatini, Bernardo L., Safaai, Houman
Abstract
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching signal and mean presynaptic activity. Its leading component drives tanh units toward saturation. At initialization, random feedback provides no systematic correction of the shared error on average; readout learning limits its duration. A reduced model initialized from the network, without fitted parameters, predicts the concentration of activation sensitivity across 48 settings. On MNIST, class decodability largely survives collapse, but readout learning remains slow at a fixed learning rate. Adam learns faster despite deeper collapse. Calibrating the baseline readout to the class prior suppresses collapse and speeds learning; weaker feedback trades less collapse for slower learning. Replacing errors by their signs sustains collapse; subtracting the signal's batch mean prevents sustained collapse and improves learning in the tested setting. Related effects occur in deeper and convolutional networks and on CIFAR-10, with severity and cost depending on the readout, optimizer and input statistics.
Chinese Translation
直接反馈对齐(Direct Feedback Alignment, DFA)通过输出误差的固定随机投影来训练隐藏层。在使用 tanh 隐藏单元和独立 sigmoid 输出时,普通随机梯度下降可能停滞在类频率常数预测器的损失水平附近。我们将这种停滞归因于误差的共模(common mode),即在所有输入之间共享的分量。一个精确的均值-协方差分解分离出一个由平均教学信号和平均突触前活动构成的秩一更新。其主导分量将 tanh 单元推向饱和。在初始化时,随机反馈平均而言对共享误差不提供任何系统性修正;读出层学习限制了停滞的持续时间。一个由网络初始化、无需拟合参数的简化模型,在 48 种设置下预测了激活敏感性的集中程度。在 MNIST 上,类别可解码性在很大程度上得以在崩塌后保留,但在固定学习率下读出层学习仍然缓慢。Adam 尽管导致更深的崩塌,却学习得更快。将基线读出层校准到类先验可抑制崩塌并加速学习;较弱的反馈以更慢的学习换取更少的崩塌。用误差的符号替代误差本身会维持崩塌;在所测试的设置中,减去信号的批均值可防止持续性崩塌并改善学习。在更深的网络、卷积网络以及 CIFAR-10 上也出现了相关效应,其严重程度和代价取决于读出层、优化器和输入统计特性。
cs.LG / 122 / 2609.31600

New LoRA Skills Should Read but Never Write

新的LoRA技能应当只读而不写
Li, Zeyan, Yang, Panqi, Guo, Qirong, Zhuo, Shengda, Qiu, SIyuan, Xu, Hu, Li, Chun, Xu, Jianfeng
Abstract
Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task data is expensive, and routing between separate adapters gives up the goal of a single combined model. We trace the difficulty to two choices that every composition method makes implicitly. A LoRA update admits infinitely many equivalent factorizations; the choice among them is invisible while an adapter serves alone, but it determines what a learned interaction between adapters can see. A coupling between an old skill and a new one can likewise point in either direction, and the direction decides whether the old skills keep computing what they computed before. We introduce READ (Read-only Expansion of Adapter Deltas), which fixes both choices: each adapter is rewritten into a balanced canonical form that preserves its update exactly, and the coupling grows in one direction only, so a new skill can read the input subspaces of old skills but cannot write into their output subspaces. The only trainable object at each append is the new skill's row of the coupling matrix, and the composed update folds into the base weights with no inference cost, routing, or task-specific rules. We evaluate READ across four benchmark suites and two model families, adding skills one at a time. Across several families, READ improves every suite average over the strongest published baselines built from the same adapters---by more than twenty points on SuperGLUE and more than seven points on the domain suite---and nearly all complete addition sequences end above every direct baseline. Factor coordinates and coupling direction, which a lone adapter never exposes, are what decide whether composed skills survive.
Chinese Translation
低秩适配器(LoRA)使得对大型语言模型进行每任务一次的微调变得低廉,但将多个独立训练的适配器合并到一个模型中仍然困难:在权重空间中合并更新会引起干扰,在所有任务数据上重新训练代价高昂,而在独立适配器之间进行路由则放弃了构建单一组合模型的目标。我们将这一困难追溯到每个组合方法都隐式做出的两个选择。一个LoRA更新存在无穷多种等价的因式分解;当适配器单独使用时,分解方式的选择是不可见的,但它决定了适配器之间可学习交互所能看到的内容。旧技能与新技能之间的耦合同样可以指向两个方向,而方向决定了旧技能能否继续保持其原有的计算。我们提出了READ(Read-only Expansion of Adapter Deltas,适配器增量的只读扩展),它固定了这两个选择:每个适配器被重写为一种精确保留其更新的平衡规范形式,并且耦合只沿一个方向增长,因此新技能可以读取旧技能的输入子空间,但无法写入其输出子空间。每次追加时唯一可训练的对象是新技能在耦合矩阵中的那一行,且组合后的更新可直接折叠进基础权重,无需额外的推理成本、路由或任务特定规则。我们在四个基准套件和两个模型家族上评估了READ,逐个添加技能。在多个家族上,READ在所有套件上的平均表现均优于由相同适配器构建的最强已发表基线——在SuperGLUE上超出二十多分,在领域套件上超出七分以上——且几乎所有完整的添加序列的最终结果都高于所有直接基线。正是单个适配器从不显露的因子坐标和耦合方向,决定了组合后的技能能否存活。
cs.LG / 123 / 2609.31603

User Model Extraction via Belief Self-Distillation

基于信念自蒸馏的用户模型提取
Holmov, Ali, Huang, Yiran, Bykov, Kirill, Akata, Zeynep
Abstract
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.
Chinese Translation
大语言模型(LLM)会隐式地推断其用户的属性并相应地调整自身行为,然而这些信念仍然难以被检查和进行因果操控。我们提出了信念自蒸馏(Belief Self-Distillation, BSD),这是一个统一的读写框架,它通过学习一种既可被解码、又可被写回模型的紧凑用户表示,将线性探测与因果探测联系起来。冻结的LLM充当其自身的教师,无需外部标注即可从自然对话中蒸馏出信念。与传统探测方法不同,BSD不仅分离出激活中存在的信息,还得到一种其因果作用可直接被检验的状态。在多个模型家族上,BSD能够忠实地恢复用户信念,并实现远强于同等条件下隐藏状态引导的干预效果。至关重要的是,我们发现模型的拒绝行为不仅取决于请求本身,还取决于模型所推断的用户意图:在保持请求不变的情况下,改变这一信念即可改变拒绝行为。我们进一步揭示了一个显著的跨模型规律:独立训练的LLM在表示其用户时收敛到一种共享的几何结构。这些结果共同表明,隐式用户模型是可读且可因果写入的内部状态,对AI安全具有直接意义——它塑造了模型如何依据其认为正在交互的对象来做出安全决策。