← Back to Index
Daily Research Digest

arXiv Papers

2026-09-23
431
Papers
4
Categories
250
Translated
收藏清单 0
机器人学 (Robotics)
104
cs.RO / 1 / 2609.25031

Towards Adaptive Interaction Strategies for Human Companion Robot via Deep Reinforcement Learning

基于深度强化学习的人机伴随机器人自适应交互策略研究
Vu, Cong-Thanh, Liu, Yen-Chen
Abstract
In the field of Human-Robot Interaction (HRI), achieving flexibility in human-accompanying within real-world environments holds great potential for various applications but also poses significant challenges. Traditional methods typically restrict robots to fixed positions relative to humans, such as tracking from behind, in front, or side-by-side, which limits robot adaptability in dynamic workspaces. This study introduces a novel human-companioning strategy that uses Reinforcement Learning (DRL) to enable mobile robots to dynamically adjust their tracking positions according to varying conditions. An interaction space is defined to capture the relationship between the human and the robot while considering the environment, which serves as the basis for state spaces in DRL to assist the robot in adapting to environmental changes. A human-robot companion controller is developed by integrating Model Predictive Path Integral (MPPI) control with Control Barrier Functions (CBF), ensuring that the robot accurately follows the target's movement in both position and orientation while avoiding obstacles and enhancing social acceptance and safety. The proposed approach is evaluated in real-world scenarios, both indoors and outdoors, and compared with other studies. The results show that the proposed method improves the success rate and tracking accuracy by at least 24% and 47%, respectively, while enhancing human comfort. Experiments demonstrate the robot's ability to flexibly accompany a person walking at speeds of up to 1.7 m/s, dynamically adjusting its strategy without being confined to a fixed position. Additionally, the robot respects the human's intimate space to ensure safety, comfort, and effective obstacle avoidance.
Chinese Translation
在人机交互(HRI)领域,实现真实环境中人机伴随的灵活性对多种应用具有巨大潜力,但也带来了重大挑战。传统方法通常将机器人限制在相对于人类的固定位置,例如从后方、前方或并排跟踪,这限制了机器人在动态工作空间中的适应性。本研究提出了一种新颖的人机伴随策略,利用深度强化学习(DRL)使移动机器人能够根据不同条件动态调整其跟踪位置。研究定义了一个交互空间,用于在考虑环境的前提下捕捉人与机器人之间的关系,该交互空间作为DRL中状态空间的基础,帮助机器人适应环境变化。通过将模型预测路径积分(MPPI)控制与控制障碍函数(CBF)相结合,开发了一个人机伴随控制器,确保机器人在位置和朝向上都能准确跟随目标的运动,同时避开障碍物,提升社会可接受性和安全性。该方法在室内外的真实场景中进行了评估,并与其他研究进行了比较。结果表明,所提出的方法将成功率和跟踪精度分别提高了至少24%和47%,同时提升了人的舒适度。实验证明,机器人能够灵活伴随以最高1.7 m/s速度行走的人,并动态调整其策略,而不局限于固定位置。此外,机器人会尊重人类的亲密空间,以确保安全性、舒适性和有效的避障。
cs.RO / 2 / 2609.25264

Cosserat Modeling of Trimmed Helicoid Soft Arms with a Separated-Section Constitutive Law

基于分离截面本构律的裁剪螺旋面软体机械臂Cosserat建模
Qin, Zhihang, Hou, Linxin, Zhong, Zeyu, Sun, Yuchen, Xin, Wenci, Zhang, Yueheng, Qi, Ji, Wang, Jie, Nazeer, Muhammad Sunny, Tan, Yu Jun, Renda, Federico, Laschi, Cecilia
Abstract
Cosserat rod models for soft robots usually construct sectional stiffness by summing material properties over a common cross-section. This assumption becomes inaccurate for trimmed helicoid arms, where load-bearing helix domains are separated and connected only through sparse fused crossings. This paper formulates a separated-section constitutive law that evaluates each helix domain in its local frame and pulls its constitutive response back to the backbone, yielding an effective backbone stiffness. Sparse-fusion mechanics captures the additional compliance caused by relative motion between neighboring domains and determines channel-wise reduction profiles $\eta_c(s/L)$ for bending, torsion, and extension. The resulting effective sectional stiffness is strongly anisotropic: bending and extension are reduced by about one order of magnitude, whereas torsion remains close to the effective backbone stiffness. The resulting sectional law is embedded in a geometrically exact dynamic Cosserat model with GVS discretization and routed-tendon actuation. Across 103 measured configurations, the three datasets give pooled normalized position errors of \SI{7.7}{\percent}, \SI{6.7}{\percent}, and \SI{7.8}{\percent}, while each full-arm solve requires approximately \SI{0.3}{s} on one CPU core (Intel Xeon, Cascade Lake, \SI{2.8}{GHz}), enabling rapid model-based planning, state and load estimation, and morphology--control co-design for architected soft robots.
Chinese Translation
软体机器人的Cosserat杆模型通常通过对公共横截面上的材料属性求和来构建截面刚度。这一假设对于裁剪螺旋面机械臂不再准确,因为其承载螺旋域彼此分离,仅通过稀疏的融合交叉点相连。本文提出了一种分离截面本构律,在各自的局部坐标系中评估每个螺旋域,并将其本构响应映射回骨架,从而得到有效的骨架刚度。稀疏融合力学刻画了相邻域之间相对运动所引起的额外柔度,并确定了弯曲、扭转和伸展的通道缩减曲线 $\eta_c(s/L)$。所得的有效截面刚度呈现强各向异性:弯曲和伸展刚度降低约一个数量级,而扭转刚度仍接近有效骨架刚度。该截面本构律被嵌入到一个采用GVS离散化并考虑走线腱驱动的几何精确动力学Cosserat模型中。在103个实测构型上,三组数据集的合并归一化位置误差分别为\SI{7.7}{\percent}、\SI{6.7}{\percent}和\SI{7.8}{\percent},且每次整臂求解在单核CPU(Intel Xeon,Cascade Lake,\SI{2.8}{GHz})上仅需约\SI{0.3}{s},为构架化软体机器人的快速基于模型的规划、状态与载荷估计以及形态—控制协同设计提供了支持。
cs.RO / 3 / 2609.25271

Beyond the Flat Seafloor: A Closed-Form Two-View Constraint to Aid Sidescan Sonar Reconstruction

超越平坦海底假设:一种用于侧扫声呐重建的闭式两视图约束
Norman, Kalin, Mangelson, Joshua G.
Abstract
Sidescan sonar is a common sensor for both manned and autonomous marine exploration and mapping, yet very few methods build upon or exploit the geometric projection model of the sensor. As sidescan sonar is limited to a 1D range measurement, many approximations are frequently used, including the long-standing assumption of a flat seafloor. Rather than make similar approximations, this paper focuses on a multi-view geometry based approach and formalizes a two-view geometric constraint and proves that a shared feature is constrained to a locus within the intersection of a sphere and a plane. In addition, we characterize what governs the size of the ambiguity locus through Monte Carlo simulation that is grounded in real aperture and mounting geometry for both a surface vessel and an underwater vehicle. We translate additional simulations of relative trajectories for both vehicle platforms into concrete survey-planning guidance. Our results show that locus length is strongly governed by elevation misalignment, and peaks at a moderate oblique crossing angle of approximately 20 degrees, with minimal locus lengths obtained at near parallel and anti-parallel passes.
Chinese Translation
侧扫声呐是有人和自主海洋探测与测绘中常用的传感器,然而很少有方法建立或利用该传感器的几何投影模型。由于侧扫声呐仅限于一维距离测量,人们经常使用许多近似方法,其中包括长期存在的平坦海底假设。本文不再采用类似的近似,而是专注于基于多视图几何的方法,形式化了一种两视图几何约束,并证明了共享特征被约束于一个球面与平面交集内的轨迹上。此外,我们通过基于真实孔径和安装几何(包括水面船只和水下航行器两种平台)的蒙特卡洛仿真,刻画了决定模糊轨迹大小的因素。我们将两种平台相对轨迹的进一步仿真转化为具体的测绘规划指导。结果表明,轨迹长度主要受高度失准的强烈影响,并在约20度的中等斜交角处达到峰值,而在近平行和反平行航次下轨迹长度最小。
cs.RO / 4 / 2609.25274

Learning to Plan in Human-Robot Collaboration: Multimodal Reinforcement Learning for Adaptive Interaction

人机协作中的规划学习:面向自适应交互的多模态强化学习
Shervedani, Afagh Mehri, Li, Siyu, Monaikul, Natawut, Abbasi, Bahareh, Di Eugenio, Barbara, Žefran, Miloš
Abstract
Robot assistants for older adults and people with disabilities need to perform collaborative tasks with users effectively. The core component of these systems is an interaction manager whose job is to observe and assess the task and infer the state of the human and their intent for the robot to choose the best course of action. Due to the sparseness of the data in this domain, the policy for such multimodal systems is often crafted by hand; as the complexity of interactions grows, this process is not scalable. This paper proposes a reinforcement learning (RL) approach to automatically generate the multimodal policy of the robot. Our system focuses on a realistic scenario where a robot assists a user in locating objects within a home environment, managing multimodal signals, including language and physical actions, to select the best action. In contrast to traditional dialog systems, our agent is trained with a simulator that uses human data and can deal with multiple modalities. We use a simple high-level reward function that needs no fine-tuning and enforce some preconditions to speed up the training process. A human study evaluating the system in a real-world setting demonstrates promising results, indicating high usability and effective task completion. This RL-based approach offers a scalable and interpretable alternative for designing interaction managers in multimodal human-robot collaborations.
Chinese Translation
面向老年人及残障人士的机器人助手需要与用户高效地执行协作任务。这类系统的核心组件是交互管理器,其职责是观察和评估任务、推断人的状态及其对机器人的意图,从而为机器人选择最佳行动方案。由于该领域数据稀疏,此类多模态系统的策略通常由人工手工设计;随着交互复杂性的增长,这一过程难以扩展。本文提出一种强化学习(RL)方法,用于自动生成机器人的多模态策略。我们的系统聚焦于一个真实场景:机器人在家庭环境中协助用户寻找物体,管理包括语言和肢体动作在内的多模态信号,以选择最佳动作。与传统的对话系统不同,我们的智能体通过一个基于人类数据的模拟器进行训练,并能够处理多种模态。我们使用一个简单的、无需微调的高层奖励函数,并通过设置部分前置条件来加速训练过程。在真实环境中对该系统进行评估的人体实验展示了令人满意的结果,表明系统具有较高的可用性和有效的任务完成能力。这种基于强化学习的方法为多模态人机协作中的交互管理器设计提供了一种可扩展且可解释的替代方案。
cs.RO / 5 / 2609.25322

JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation

JAMB:面向双臂操作的联合动作-运动扩散模型
Xiao, Chuyang, Meng, Peilin, Held, David
Abstract
Coordinated bimanual manipulation is challenging because the motion of either arm can alter the shared 3D scene and thereby affect the other arm. Yet most diffusion policies generate actions without explicitly modeling these future geometric consequences, while predictive variants typically use future state only as auxiliary supervision or fixed conditioning. We address this limitation by proposing JAMB, a diffusion policy that jointly denoises bimanual actions and future 3D point tracks. By allowing action and track hypotheses to evolve together within a shared Transformer, each can inform and refine the other throughout denoising. We further ground multimodal representations in a shared spatiotemporal coordinate system to facilitate geometry-aware interaction during joint denoising. We evaluate JAMB on diverse bimanual manipulation tasks in RoboTwin 2.0 and on a real-world robot, comparing it with action-only policies and alternative future-prediction approaches spanning different state representations and learning objectives. Across 16 simulation tasks, JAMB achieves an average success rate of 83.4%, outperforming the strongest baseline by 23.9 percentage points. On three real-world tasks, it outperforms the action-only and auxiliary geometry prediction methods by 50.0 and 21.2 percentage points, respectively. Beyond these performance gains, JAMB shows stronger generalization to cluttered scenes and out-of-distribution backgrounds than the evaluated baselines. Together, these results demonstrate the effectiveness of our joint action-motion modeling framework for coordinated bimanual manipulation. Our project website is available at https://jam-bimanual.github.io/
Chinese Translation
协调的双臂操作极具挑战性,因为任一机械臂的运动都可能改变共享的三维场景,进而影响另一条机械臂。然而,大多数扩散策略在生成动作时并未显式建模这些未来几何后果,而预测型变体通常仅将未来状态用作辅助监督或固定条件。为解决这一局限,我们提出了 JAMB,一种能够对双臂动作和未来三维点轨迹进行联合去噪的扩散策略。通过让动作假设与轨迹假设在共享的 Transformer 中共同演化,二者可以在整个去噪过程中相互提供信息并相互优化。我们进一步将多模态表示建立在共享的时空坐标系中,以促进联合去噪过程中的几何感知交互。我们在 RoboTwin 2.0 中的多种双臂操作任务以及真实机器人上评估了 JAMB,并将其与仅动作策略以及涵盖不同状态表示和学习目标的替代性未来预测方法进行了比较。在 16 个仿真任务中,JAMB 取得了 83.4% 的平均成功率,比最强基线高出 23.9 个百分点。在三个真实世界任务中,它分别比仅动作方法和辅助几何预测方法高出 50.0 和 21.2 个百分点。除了这些性能提升之外,JAMB 在杂乱场景和分布外背景上比所评估的基线表现出更强的泛化能力。这些结果共同证明了我们的联合动作-运动建模框架在协调双臂操作中的有效性。我们的项目网站见 https://jam-bimanual.github.io/
cs.RO / 6 / 2609.25338

GINIO: A Geometric SO(3)-Equivariant Interface for Neural Inertial Odometry

GINIO:面向神经网络惯性里程计的几何SO(3)等变接口
Kim, Chankyo, Zhu, Minghan, Lin, Tzu-Yuan, Rattan, Avantika, Ghaffari, Maani
Abstract
Neural inertial odometry increasingly uses networks as learned measurements inside filtering pipelines. Such measurements should transform consistently under arbitrary IMU mounting conventions: their mean must transform as a vector, and their covariance must transform congruently as a second-order tensor. We present GINIO, a geometric SO(3)-equivariant interface for neural inertial odometry under arbitrary rotations of the IMU measurement frame. Given calibrated IMU windows, our framework predicts a motion measurement and uncertainty obeying these tensorial laws. To support efficient sensor-frame learning, we introduce Last-Frame Alignment (LFA), a deterministic preprocessing step that is provably equivalent to world-frame training for SO(3)-equivariant predictors. The connected estimator tracks sensor-local states such as IMU bias, separating nuisance estimation from the geometric law enforced by the learned measurement. We instantiate the same interface in filter-connected NIO, AirIO-style recurrent aerial prediction, EqNIO-style full-SO(3) canonicalization, and ResNet-style temporal backbones. On TLIO, GINIO achieves 2.018 m ID/SO(3) ATE while EqNIO degrades to 76.389 m, using 11.6x fewer FLOPs. On NanoBench, our AirIO-style instantiation improves ATE from 5.579 m to 1.430 m without external attitude input, and our ResNet-style instantiation reaches 0.581 m ATE versus 0.645 m for ResNet1D. On Fetch, GINIO empirically reduces unseen physical-remount ATE from 8.15 m to 0.50 m without retraining, demonstrating robustness beyond the exact coordinate-frame guarantee. For uncertainty, spectral covariance reduces covariance-equivariance error by over three orders of magnitude compared with a diagonal head.
Chinese Translation
神经网络惯性里程计(Neural Inertial Odometry)日益将网络作为滤波管线中的学习型测量模型使用。此类测量应在任意IMU安装约定下保持一致的变换特性:其均值必须按向量方式变换,其协方差必须作为二阶张量进行合同变换。我们提出了GINIO,一个面向任意IMU测量坐标系旋转情形的神经网络惯性里程计几何SO(3)等变接口。给定已标定的IMU观测窗口,该框架预测符合上述张量变换定律的运动测量值及不确定性。为支持高效的传感器坐标系学习,我们引入了末帧对齐(Last-Frame Alignment, LFA),这是一种确定性的预处理步骤,可证明其对SO(3)等变预测器而言与世界坐标系训练等价。所连接的估计器跟踪诸如IMU偏置等传感器局部状态,从而将干扰量估计与学习型测量所强制的几何定律分离。我们在滤波器连接的NIO、AirIO风格的循环飞行器预测、EqNIO风格的全SO(3)规范化以及ResNet风格的时序骨干网络中实例化了同一接口。在TLIO数据集上,GINIO实现了2.018 m的ID/SO(3) ATE,而EqNIO退化至76.389 m,且我们的方法所需FLOPs减少11.6倍。在NanoBench数据集上,我们的AirIO风格实现在无外部姿态输入的情况下将ATE从5.579 m提升至1.430 m,我们的ResNet风格实例达到0.581 m的ATE,优于ResNet1D的0.645 m。在Fetch数据集上,GINIO在无需重训练的情况下,将未见过的物理重新安装场景的ATE从8.15 m经验性地降低至0.50 m,展现了超越精确坐标系保证的鲁棒性。在不确定性估计方面,谱协方差方法相比对角协方差头将协方差等变误差降低了超过三个数量级。
cs.RO / 7 / 2609.25351

Learning from Humans for Proactive Assistance in Human-Robot Collaborative Transport

从人类学习中实现人机协作搬运中的主动辅助
Yang, Elvin, Mavrogiannis, Christoforos
Abstract
We focus on human-robot collaborative transport, a challenging task of broad relevance spanning logistics, manufacturing, and the home, in which a user and a robot work together to relocate a large or heavy object. To act as an effective partner, the robot should reduce the user's effort by contributing to efficient relocation of the object while remaining physically responsive to them. Prior work often addresses these capabilities separately, producing robots that may move the object efficiently but resist user input, or accommodate the user but depend on continuous guidance. Our key insight is that obstacle-constrained collaborative transport requires integrating predictions of human collaborative behavior with compliant robot control. To this end, we introduce PROACT, a framework for human-robot collaborative transport that incorporates anticipation into compliant whole-body control through a learned model of human collaborative behavior. Trained on a large-scale, real-world dataset of dyadic human transport demonstrations, our transformer architecture distills collaborative behavior into predictions of future object motion. Across 108 real-world trials with a 9-DoF mobile manipulator, PROACT reduces mean interaction work by 59.2\% and 20.4\%, and mean completion time by 12.9\% and 6.9\%, relative to compliance-only and MPC baselines, respectively. Footage from our experiments can be found at https://youtu.be/qAGvQfVPjbk.
Chinese Translation
我们聚焦于人机协作搬运任务,这是一项具有广泛相关性的挑战性任务,涵盖物流、制造和家庭场景,其中用户与机器人协同搬运大型或重型物体。为了成为有效的合作伙伴,机器人应在促进物体高效搬运以减少用户负担的同时,保持对用户力输入的物理响应性。已有工作通常分别处理这些能力,所产生的机器人或能高效移动物体但抵抗用户输入,或能顺应用户但依赖持续的引导。我们的核心洞察是:受障碍物约束的协作搬运需要将人类协作行为的预测与柔顺机器人控制相结合。为此,我们提出了PROACT,一个将预测能力融入柔顺全身控制的人机协作搬运框架,其通过学习人类协作行为模型实现这一目标。基于大规模真实世界双人搬运演示数据集训练,我们的Transformer架构将协作行为提炼为对未来物体运动的预测。在一台9自由度移动操作臂上进行的108次真实世界实验中,相对于纯柔顺基线和MPC基线,PROACT分别将平均交互功降低了59.2%和20.4%,平均完成时间降低了12.9%和6.9%。实验视频请见 https://youtu.be/qAGvQfVPjbk。
cs.RO / 8 / 2609.25363

HOTICE: Whole-Body Humanoid Object Transportation in Cluttered Environments

HOTICE:面向杂乱环境的全身人形机器人物体搬运
Nguyen, Toan, Yuan, Weiduo, Zhao, Siheng, Wang, Yue, Seita, Daniel
Abstract
Object transportation is a fundamental capability for humanoid robots operating in real-world, human-centric environments, yet existing methods struggle when clutter constrains free space around both the robot and its carried payload. We present HOTICE, a whole-body humanoid learning framework for transporting objects through such cluttered environments. First, we introduce Humanoid-Object Decoupled Potential Fields, which jointly encode collision-avoidance guidance for the robot and the carried object, enabling coordinated, obstacle-aware motion for both. Second, to address the large action space inherent to whole-body loco-manipulation, we design a dual-agent reinforcement learning architecture that decouples upper- and lower-body control while preserving whole-body coordination via shared state observations and rewards. To train a policy that generalizes across diverse cluttered scenes, we further employ a specialist-to-generalist distillation strategy, in which privileged teacher policies are distilled into a single deployable student policy. We evaluate HOTICE in MuJoCo simulation and on a real Unitree G1 humanoid, demonstrating effective and robust object transportation across cluttered scenarios for objects of varying shapes. Our results show that HOTICE reliably coordinates whole-body motion and object-aware collision avoidance, generalizing effectively to previously unseen cluttered environments while achieving strong performance in sim2real deployment.
Chinese Translation
物体搬运是人形机器人在真实世界中以人为中心的环境里运行的一项基础能力,然而当杂乱环境限制了机器人及其所载物体周围的自由空间时,现有方法往往表现不佳。我们提出了 HOTICE,一个用于在此类杂乱环境中搬运物体的全身人形机器人学习框架。首先,我们引入了人形-物体解耦势场,共同编码针对机器人与所载物体的避障引导,使二者能够实现协同且具备障碍感知的运动。其次,为了应对全身移动操作固有的大动作空间,我们设计了一种双智能体强化学习架构,将上半身与下半身的控制解耦,同时通过共享的状态观测和奖励保持全身协调。为训练出能在多样化杂乱场景中泛化的策略,我们进一步采用专家到通才的蒸馏策略,将具有特权信息的教师策略蒸馏为单个可部署的学生策略。我们在 MuJoCo 仿真和真实的宇树 Unitree G1 人形机器人上评估了 HOTICE,结果表明其能够针对不同形状的物体,在各种杂乱场景中实现有效且鲁棒的物体搬运。我们的结果显示,HOTICE 能够可靠地协调全身运动并实现物体感知的碰撞避让,有效泛化至未见过的杂乱环境,同时在 sim2real 部署中取得了优异的表现。
cs.RO / 9 / 2609.25369

Capability-Aware Arbitration for Semantic Intent-Based Shared Control

面向语义意图共享控制的能力感知仲裁机制
Du, Zhaoda, Bowman, Michael, Zhang, Xiaoli
Abstract
Shared control often allocates robot authority based on confidence in inferred human intent, assuming reliable autonomous execution. When this assumption fails, high intent confidence can cause over-helping. We present a capability-aware shared-control framework in which a vision-language model (VLM) infers human intent and provides semantic-intent confidence, while a vision-language-action (VLA) policy generates autonomous actions. VLA capability confidence is estimated online from the dispersion and local instability of stochastic action trajectories. We design a nonlinear arbitration policy that combines Bayesian-filtered semantic-intent confidence with VLA capability confidence through a sigmoid mapping to adapt robot authority. Our evaluation combined VLM/VLA confidence assessment with a study involving 12 participants performing pick-and-place and bidirectional stacking under in-distribution and out-of-distribution conditions. The proposed method achieved the highest task success rate (92%), compared with manual teleoperation (83%), intent-only arbitration (44%), and fixed equal-weight blending (10%). It also achieved higher control friendliness and lower authority-weighted disagreement than both shared-control baselines. These results demonstrate the benefit of incorporating VLA capability into authority allocation to mitigate over-helping and improve shared-control performance.
Chinese Translation
共享控制通常基于对推断出的人类意图的置信度来分配机器人控制权,并假设自主执行是可靠的。当该假设不成立时,较高的意图置信度可能导致“过度辅助”。我们提出了一种能力感知的共享控制框架,其中视觉-语言模型(VLM)推断人类意图并提供语义意图置信度,而视觉-语言-动作(VLA)策略生成自主动作。VLA能力置信度通过对随机动作轨迹的离散度和局部不稳定性的在线估计获得。我们设计了一种非线性仲裁策略,通过Sigmoid映射将贝叶斯滤波后的语义意图置信度与VLA能力置信度相结合,以自适应地调整机器人控制权。我们的评估将VLM/VLA置信度评估与一项用户实验相结合,12名参与者分别在分布内和分布外条件下执行拾取放置和双向堆叠任务。所提出的方法取得了最高的任务成功率(92%),相比之下,手动遥操作为83%,仅基于意图的仲裁为44%,固定等权重混合为10%。与两个共享控制基线相比,该方法还实现了更高的控制友好性和更低的权利加权分歧。这些结果表明,将VLA能力纳入控制权分配中,有助于缓解过度辅助并提升共享控制性能。
cs.RO / 10 / 2609.25374

Angular momentum analysis on Karate roundhouse kicks: a longitudinal case study

Lau, Jan C. L., Mele, Christian, Lin, Jonathan Feng-Shun, Mombaur, Katja
Abstract
Human gait in reality extends beyond straight-line walking, with some situations even requiring drastic back-and-forth rotations. A smooth and stable execution may seem intuitive, but the underlying mechanics remains unknown. The Karate roundhouse kick may be an extreme case of exhibiting dynamic back-and-forth rotations, but investigating the angular momentum (AM) management can potentially inform stability analysis in dynamic human motions and smoother gait in robots and exoskeletons. This paper introduces two new AM-based measures and analyzes AM-related variables to study the target-less retractable back-leg Karate roundhouse kick. The purpose is to understand the underlying AM management, analyze the differences between stable and unstable kicks, and investigate how these variables and measures change over time with improvement. A one-year longitudinal study was conducted with the first author as a Karate student, and a Karate instructor was also recruited for one session as an expert, whose data is used for comparison. Results show that unstable kicks have higher peak total AM before Strike and smaller braking peak after Strike. The proposed AM-based measures are able to identify the cause of some unstable kicks, though no clear distinction could be made between stable and unstable kicks since kicks can be unstable for different reasons. Nonetheless, the proposed measures can be applied as performance measures to analyze other types of dynamic motions. To incorporate them as stability criteria, however, it is recommended to also consider their coordination with time, kinematics, and center of pressure.
cs.RO / 11 / 2609.25375

PARTE: Plane-Assisted Robust Transformation Estimation for Point Cloud Registration

Babanazari, Abolfazl, Cramer, Carson, Summers, Tyler, Nieto, Carlos, Fathian, Kaveh
Abstract
Global point-cloud registration remains challenging when limited overlap, repetitive geometry, and sensor noise produce correspondence sets dominated by outliers. Planar regions are particularly difficult for conventional point descriptors and are therefore often suppressed or discarded before matching. We present PARTE (Plane-Assisted Robust Transformation Estimation), a global registration method that instead treats planar structure as complementary registration evidence. PARTE extracts planar patches and represents them using our novel Plane Context Histogram (PCH), a descriptor that encodes the geometry surrounding each patch, while a two-level matching procedure identifies reliable plane correspondences. Candidate point and plane correspondences are combined in a confidence-weighted compatibility graph for joint outlier rejection, followed by rigid transformation estimation. When no usable plane correspondences are available, PARTE naturally reduces to point-only registration. We evaluate PARTE on 8,097 registration pairs across six indoor and outdoor benchmarks spanning dense RGB-D and sparse LiDAR measurements. Evaluations show PARTE achieves the highest overall success rate against 13 standard and state-of-the-art methods while maintaining low runtime. An open-source C++ implementation with Python bindings is provided at https://parte.pages.dev.
cs.RO / 12 / 2609.25376

VLAQuantBench: Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action Models

Xu, Jiuyi, Jin, Qing, Chen, Meida, Wang, Song, Sui, Yang, Shi, Yangming
Abstract
Post-training quantization reduces the memory requirements of vision-language-action (VLA) models, but precision selection must account for the interaction between layer scope, numerical format, and calibration. We introduce \textbf{VLAQuantBench}, a controlled evaluation with 409 runs and 94,574 simulation episodes: four models on LIBERO, with X-VLA additionally evaluated on three simulation benchmark families. Under uncalibrated W4A4 round-to-nearest quantization, expanding a $\pi_{0.5}$ action-head subset from 126 to 167 layers raises success from 7.0\% to 70.5\%. Fixed-observation replay confirms a corresponding numerical recovery. Two-episode calibration removes the severe joint failures in the tested subsets, whereas the same smoothing-and-clipping recipe lowers $\pi_0$ success and does not recover OpenVLA-OFT end-to-end. For OpenVLA-OFT, protecting one 28,672-parameter output projection instead restores near-baseline success: the remaining 441 eligible linear layers retain W3 on LIBERO-Long or eight-bit activations across all four suites. Task-clustered intervals support the large failure and recovery contrasts. These results establish recipe-dependent interactions and identify concrete precision assignments, rather than universal layer-sensitivity rules. Real-kernel and physical-robot measurements complement the accuracy analysis. Code, configurations, and episode records are publicly available at https://github.com/jiuyixu25/VLAQuantBench.
cs.RO / 13 / 2609.25398

Norm2Tex: Augmenting Visuo-Tactile Simulations with Texture

Norm2Tex:利用纹理增强视觉-触觉仿真
Bien, Seongjin, Makowski, Débora Oliveira, Calandra, Roberto, Walter, Florian, Burgard, Wolfram
Abstract
Large-scale datasets are essential for training generalist robot control policies. Collecting real-world tactile data is costly and time-consuming, motivating the use of tactile simulations. However, current tactile simulators capture only overall contact geometry and miss fine details like texture. This results in a significant domain shift between simulated and real tactile data. To address this gap, we introduce Norm2Tex, a plug-in method that augments simulations of vision-based tactile sensors with high-frequency surface details from normal map textures. By modifying the target object's depth map before a tactile simulator's rendering pipeline, Norm2Tex seamlessly integrates into different tactile simulators. We also evaluate sim-to-real transfer using material classification and a reinforcement learning task. Our results show that Norm2Tex preserves material-dependent tactile information across domains, improving texture recognition and producing material-dependent control behavior in the real world.
Chinese Translation
大规模数据集对于训练通用型机器人控制策略至关重要。采集真实世界的触觉数据既昂贵又耗时,这促使人们使用触觉仿真。然而,当前的触觉仿真器只能捕获整体接触几何信息,而遗漏了纹理等细微细节。这导致仿真触觉数据与真实触觉数据之间存在显著的域偏移。为弥补这一差距,我们提出了 Norm2Tex,这是一种即插即用方法,通过法向贴图纹理中的高频表面细节来增强基于视觉的触觉传感器的仿真。Norm2Tex 通过在触觉仿真器的渲染流水线之前修改目标物体的深度图,能够无缝集成到不同的触觉仿真器中。我们还通过材质分类和强化学习任务评估了仿真到真实(sim-to-real)的迁移效果。结果表明,Norm2Tex 能够在不同域之间保留与材质相关的触觉信息,提升了纹理识别能力,并在真实世界中产生了依赖材质的控制行为。
cs.RO / 14 / 2609.25417

Effects of Assistance Delay on Joint Mechanics and Energetics in Biological Torque Control of a Hip Exoskeleton

助力延迟对髋关节外骨骼生物力矩控制中关节力学与能量代谢的影响
An, Jimin, Lee, Ryan, Peng, Jingshu, Halilaj, Eni, Kang, Inseung
Abstract
Biological torque control directly maps an estimated human joint moment to exoskeleton assistance, providing a task-agnostic strategy for supporting diverse locomotor activities. However, it remains unclear whether a fixed state-to-torque mapping provides effective assistance across biomechanically distinct tasks. We examined how assistance delay affected hip exoskeleton performance during level-ground (LG), ramp-ascent (RA), and ramp-descent (RD) walking. Eight participants completed a zero-torque baseline condition and five active assistance conditions with delays ranging from 40 to 320 ms. Across tasks and active delays, assistance reduced net metabolic rate by 5.24%, positive biological hip joint work by 5.86%, and total lower-limb positive joint work by 1.68% (all p < 0.05). Assistance delay affected both joint-work outcomes (both p < 0.001) but not net metabolic rate. Mechanical unloading generally decreased with increasing delay, whereas metabolic benefits remained comparatively stable. Relative to the zero-torque condition, net metabolic rate decreased by 9.75% during LG and 7.20% during RA but increased by 1.23% during RD. We did not detect task-dependent differences in the delay response. Our findings indicate that biological torque mappings should be evaluated based on the target outcome and mechanical role of the assisted joint, and that predominantly positive-power assistance may not generalize to negative-work-dominant locomotion without modification.
Chinese Translation
生物力矩控制将估计的人体关节力矩直接映射为外骨骼助力,为支持多种运动活动提供了一种任务无关的策略。然而,固定的状态-力矩映射能否在生物力学特性不同的任务中提供有效助力仍不清楚。我们研究了助力延迟如何影响髋关节外骨骼在平地行走(LG)、上坡行走(RA)和下坡行走(RD)中的表现。八名参与者完成了零力矩基线条件以及延迟范围为40至320毫秒的五种主动助力条件。在各任务和各主动延迟条件下,助力使净代谢率降低了5.24%,髋关节正的生物功降低了5.86%,下肢总的正关节功降低了1.68%(均p < 0.05)。助力延迟影响了两个关节功指标(均p < 0.001),但未影响净代谢率。随着延迟增加,机械卸载通常减小,而代谢获益保持相对稳定。相对于零力矩条件,净代谢率在平地行走中降低了9.75%,在上坡行走中降低了7.20%,但在下坡行走中增加了1.23%。我们未检测到延迟响应中存在任务依赖性差异。我们的研究结果表明,应根据目标结局和受助关节的力学作用来评估生物力矩映射,且以正功率为主的助力若不加以修改,可能无法推广到以负功为主的运动任务中。
cs.RO / 15 / 2609.25450

REDACT: Robust Perceptive Locomotion under Unseen Visual Corruption

REDACT:未知视觉损坏下的鲁棒感知运动
Kirdwichai, Natapat, Driskell-Poole, Tobias, Sontea, Andrei, Dash, Jadu, Hafez, Muhammad Burhan, Tarapore, Danesh
Abstract
Depth-conditioned locomotion policies have demonstrated impressive agile maneuvers, but can be steered to unpredictable actions when observations are outside their training distribution. Occlusion, invalid returns, sensor noise, and visual distractors can shift deployment observations away from nominal simulated depth. While synthetic sensor augmentation targets specified degradations, it does not by itself define behavior under corruption families omitted from training. To address gaps in training-time coverage, we present REDACT (Retaining Evidence Despite Artifacts for Continued Traversal), a teacher-student framework combining an improved visual encoder architecture, persistent feature masking, and a novel consensus-gating algorithm to retain useful depth information under unmodeled corruption. The gate uses approximate conformal calibration on clean observations alone, requiring no prior knowledge of the corruption type. Trained on clean simulated depth, REDACT retains useful visual information under unseen corruption, supporting higher traversal success than existing parkour baselines. Evaluation of depth augmentation across corruption families further shows that REDACT improves robustness where augmentation coverage is missing. Real-world trials demonstrate zero-shot transfer to structured and forested environments with unfamiliar scene content.
Chinese Translation
基于深度图像的运动控制策略已展现出令人印象深刻的敏捷动作,但当观测超出其训练分布时,可能被引导产生不可预测的行为。遮挡、无效返回、传感器噪声和视觉干扰物会使部署时的观测偏离仿真中的标称深度。虽然合成传感器增强可以针对特定的退化类型,但其本身并未定义训练中未涵盖的损坏类型下的行为。为解决训练时覆盖范围的缺口,我们提出了REDACT(Retaining Evidence Despite Artifacts for Continued Traversal,即在伪影干扰下保留证据以实现持续通行),这是一个教师-学生框架,结合了改进的视觉编码器架构、持续性特征掩蔽以及一种新颖的共识门控算法,以在未建模的损坏条件下保留有用的深度信息。该门控仅使用干净观测进行近似保形校准(conformal calibration),无需对损坏类型有任何先验知识。通过在干净的仿真深度数据上训练,REDACT在未见过的损坏下仍能保留有用的视觉信息,实现了比现有跑酷基线更高的通行成功率。针对不同损坏类型的深度增强评估进一步表明,在增强覆盖缺失的场景中,REDACT能够提升鲁棒性。真实世界实验证明了其在具有陌生场景内容的结构化环境和森林环境中的零样本迁移能力。
cs.RO / 16 / 2609.25470

A bioinspired internal model-based online estimator for planar pursuit

一种基于仿生内部模型的平面追踪在线估计器
Liu, Tengyue, Li, Xincheng, Ferreira, Sofia Morales, Galloway, Kevin, Halder, Udit
Abstract
Bioinspired feedback controls for pursuit, tracking, and collective motion are often expressed in terms of the relative configuration between interacting agents. In practice, however, onboard sensors may not directly provide all quantities required for feedback control, necessitating estimation of unobserved quantities. This paper develops a bioinspired internal model-based estimator for reconstructing those quantities from partial sensory observations and known self-motion. State reconstruction is posed as an optimization problem that treats the relative kinematics as constraints and minimizes the disagreement between the internal model outputs and measurements from onboard sensors. Pontryagin's Maximum Principle is used to derive the necessary optimality conditions. A forward-backward algorithm is used to provide a numerical solution and a moving horizon formulation is employed for online implementation. The estimator is evaluated numerically against classical state estimators. Real-time implementation of the proposed framework on robotic hardware is demonstrated through two pursuit strategies.
Chinese Translation
面向追踪、跟踪和群体运动的仿生反馈控制通常以交互个体之间的相对构型来表达。然而在实际应用中,机载传感器可能无法直接提供反馈控制所需的全部物理量,因此需要对未观测到的量进行估计。本文提出一种基于仿生内部模型的估计器,利用部分传感观测和已知的自身运动信息来重构这些物理量。状态重构被表述为一个优化问题,该问题将相对运动学作为约束条件,并最小化内部模型输出与机载传感器测量值之间的偏差。利用庞特里亚金极大值原理推导出最优性必要条件,采用前向-后向算法求得数值解,并采用移动时域方法实现在线运行。通过数值仿真将该估计器与经典状态估计器进行了比较评估,并借助两种追踪策略在机器人硬件上演示了所提框架的实时实现。
cs.RO / 17 / 2609.25486

Brace Yourself: Task-Conditioned Environmental Bracing for Forceful Humanoid Manipulation

稳住自己:面向强力人形机器人操作的任务条件化环境支撑
Zhang, Zongyuan, Lehnert, Christopher, Browne, Will N., Roberts, Jonathan M.
Abstract
Forceful manipulation is challenging for humanoid robots because interaction forces can disturb whole-body balance. We introduce the Supporting Hand Strategy (SHS), which enables a humanoid to brace against the environment with one hand while performing forceful manipulation with the other. SHS optimises a task-conditioned support configuration that guides two synchronous reinforcement-learning policies, without human motion data or online whole-body trajectory planning. On a Unitree G1, SHS achieved usable contact forces up to 60 N, compared with a maximum sustained force of 13.5 N without environmental bracing, while substantially improving force tracking over a task-independent support configuration. The same policies generalised to different task regions without retraining. SHS therefore provides a simple mechanism for substantially extending humanoid forceful-manipulation capability.
Chinese Translation
强力操作对人形机器人而言极具挑战性,因为交互力可能会干扰全身平衡。我们提出了支撑手策略(Supporting Hand Strategy, SHS),使机器人能够用一只手撑住环境,同时用另一只手执行强力操作。SHS 通过优化任务条件化的支撑配置来引导两个同步运行的强化学习策略,无需人类运动数据或在线全身轨迹规划。在 Unitree G1 机器人上,SHS 实现了高达 60 牛顿的可用接触力,而不借助环境支撑时最大持续力仅为 13.5 牛顿;与任务无关的支撑配置相比,力跟踪性能也得到显著提升。相同的策略无需重新训练即可泛化到不同的任务区域。因此,SHS 为大幅提升人形机器人的强力操作能力提供了一种简单机制。
cs.RO / 18 / 2609.25506

RoboMP-DINOv2: Prompts, Not Filters for Robust Robot Manipulation

RoboMP-DINOv2:面向鲁棒机器人操作使用提示而非过滤器
Qi, Han, Yang, Heng
Abstract
Robot manipulation policies must generalize across visual shifts while preserving scene context relevant to action. General-purpose vision encoders are not tailored to visuomotor control, while object-centric approaches often use segmentation masks as hard filters that discard potentially useful context. We propose RoboMP-DINOv2 (Robotics Mask-Prompted DINOv2), a full-scene vision encoder that treats masks as spatial prompts rather than visibility filters. It extracts dense DINOv2 features from the full observation, injects learned region-specific embeddings at masked locations, and jointly contextualizes prompted and unprompted tokens for action prediction. We further introduce masked-region color randomization (MCR) to improve appearance robustness, yielding RoboMP-DINOv2-MCR. Across seven simulated manipulation settings, RoboMP-DINOv2 achieves 60.7% success under spatial shifts and 59.7% under scene clutter, compared with 50.7% and 41.0% for a DINOv2-based Diffusion Policy. Under unseen object colors, RoboMP-DINOv2-MCR achieves 72.5% success versus 35.1% for the strongest color-randomized baseline. Additional experiments and representation analyses show improved robustness while preserving behaviorally relevant scene information. Code is available at https://github.com/han20192019/RoboMP_DINOv2.
Chinese Translation
机器人操作策略必须在应对视觉变化的同时保持泛化能力,并保留与动作相关的场景上下文信息。通用视觉编码器并非专为视觉运动控制而设计,而以物体为中心的方法通常将分割掩码用作硬过滤器,从而丢弃潜在有用的上下文信息。我们提出了RoboMP-DINOv2(Robotics Mask-Prompted DINOv2),这是一种全场景视觉编码器,它将掩码视为空间提示而非可见性过滤器。该方法从完整观测中提取稠密的DINOv2特征,在掩码位置注入学习到的区域特定嵌入,并对提示和未提示的标记进行联合上下文化处理以用于动作预测。我们进一步引入了掩码区域颜色随机化(Masked-Region Color Randomization, MCR)以提升外观鲁棒性,得到RoboMP-DINOv2-MCR。在七个仿真操作环境中,RoboMP-DINOv2在空间变化下取得60.7%的成功率,在场景杂乱下取得59.7%的成功率,而基于DINOv2的Diffusion Policy分别为50.7%和41.0%。在未见过的物体颜色下,RoboMP-DINOv2-MCR取得72.5%的成功率,而最强的颜色随机化基线仅为35.1%。额外的实验与表征分析表明,该方法在提升鲁棒性的同时保留了与行为相关的场景信息。代码已发布于 https://github.com/han20192019/RoboMP_DINOv2。
cs.RO / 19 / 2609.25511

A Deployment Study of Identity-Gated Drone Gesture Control

基于身份门控的无人机手势控制部署研究
Salih, Diyari Mohammed, Chaabeni, Ilyes, Oufroukh, Naima Ait
Abstract
Vision-based gesture control accepts commands from any hand in the camera field of view, which is unsafe in shared indoor spaces. This paper presents IGate, an identity-gated control stack that includes gesture control and face tracking, in which commands are admitted only when an enrolled operator is verified. The system performs few-shot user enrolment from 20 initial face frames, without prior user-specific training: verification compares an embedding of the current face crop against the enrolled template by cosine similarity, while face tracking uses proportional correction. Gesture control is achieved by classifying extracted hand landmarks using an RBF-SVM trained on a custom dataset. Additionally, a hierarchical finite-state machine handles mode selection, default, and fallback behaviours. The approach is tested on a DJI Tello EDU, each component evaluated offline and in-flight across 270 trials (149 flown). Face verification yields a 0.32% offline equal error rate versus 19.3% in-flight. Under hover-locked conditions, the RBF-SVM gesture classifier outperforms the geometric rule (0.850 vs. 0.651 accuracy), with 82% of this gap stemming from the depth channel. All logs and reproduction scripts will be released.
Chinese Translation
基于视觉的手势控制会接受来自相机视野内任意手部的指令,这在共享室内空间中是不安全的。本文提出了IGate,一种包含手势控制与人脸追踪的身份门控控制栈,其中指令仅在已注册的操作者通过身份验证后才被接受。该系统仅需20帧初始人脸图像即可完成少样本用户注册,无需预先的用户特定训练:验证过程通过余弦相似度将当前人脸裁剪图像的嵌入向量与注册模板进行比较,而人脸追踪则采用比例校正。手势控制通过使用在自定义数据集上训练的RBF-SVM(径向基函数支持向量机)对提取的手部关键点进行分类来实现。此外,系统采用分层有限状态机处理模式选择、默认行为和回退行为。该方法在DJI Tello EDU无人机上进行了测试,每个组件均通过270次试验(其中149次为飞行试验)进行了离线与在线评估。人脸验证的离线等错误率为0.32%,而在线(飞行中)为19.3%。在悬停锁定条件下,RBF-SVM手势分类器的准确率优于几何规则方法(0.850对0.651),其中82%的性能差距源于深度通道。所有日志和复现脚本将会公开。
cs.RO / 20 / 2609.25527

Digital Twin-Driven VR Teleoperation with Multi-View Spatial Perception for Surgical Robots

Liu, Chang, Yu, Chenhao, Zhao, Honghao, Ding, Hao, Wei, Haochen, Munawar, Adnan, Unberath, Mathias, Kazanzides, Peter
Abstract
Current robot-assisted minimally-invasive surgery (RMIS) platforms provide a fixed console for the surgeon to view stereo endoscopic images and teleoperate instruments inside the patient. Several researchers have proposed the use of a head-mounted display (HMD) as a portable console, with video pass-through rendering of the endoscope images which, like the fixed console, restricts the operator to a single endoscopic viewpoint and limits depth perception. We present a digital twin-driven virtual reality (VR) teleoperation platform, where the digital twin is created from markerless perception of the surgical environment and streamed for display on the HMD. This overcomes the limitations of video pass-through by providing multi-view rendering and natural motion-parallax cues, enabling decoupling of the user's hand posture from strict instrument alignment. The system utilizes VR hand controllers to increase the teleoperation workspace and to improve the robustness and stability of instrument control compared to the hand tracking approach adopted by most prior systems. A 15-participant user study on the da Vinci Research Kit (dVRK) shows that our VR platform significantly outperforms a state-of-the-art HoloLens 2 mixed reality baseline, reducing path length by 86% and jerk by 95%, while achieving depth perception confidence comparable to or exceeding the traditional console across all conditions.
cs.RO / 21 / 2609.25558

HABILIS Brain 0: Geometry-Change Supervision for Vision-Language-Action and Residual Flow Recovery

HABILIS Brain 0:面向视觉-语言-动作与残差流恢复的几何变化监督
Pahk, Jinu, Kang, Jesoon, Park, Taegeon, An, Jisu, Kimm, Soo Min, Kim, Jaejoon, Zhang, Byoung-Tak
Abstract
Vision-language-action policies benefit from geometric supervision, but current-frame geometry alone does not explicitly describe the changes associated with manipulation. This design is motivated by the goal of learning an embodiment-agnostic visual interface that can be pretrained across robot and egocentric video before robot-specific action alignment. We introduce Geometry-Change VLA (GC-VLA), which learns to predict multiview future-current geometry-change tokens from current observations. Offline frame pairs define a nominal 0.5-second prediction horizon; future observations are used only to construct training targets. Stage 1 trains a geometry-change vision-language model (GC-VLM). Stage 2 introduces a continuous ActionExpert and aligns it with robot actions while stopping action-flow gradients at the VLM interface. Stage 3 enables these gradients to update the trainable VLM components jointly with the ActionExpert. Stage 4 freezes GC-VLA and applies Geometry-Conditioned Residual Flow (GCRF), using a binary intervention router and a single bounded residual velocity policy learned from closed-loop feedback. GC-VLA achieves 95.20% success on LIBERO, and GC-VLA with GCRF achieves 99.55%. Inference uses current observations and the learned GC representation without executing the offline target encoders.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)策略受益于几何监督,但仅依靠当前帧的几何信息并不能显式地描述与操作相关的变化。该设计的动机在于学习一种与具身形态无关(embodiment-agnostic)的视觉接口,使其能够在机器人特定动作对齐之前,先在机器人数据和第一视角(egocentric)视频上进行预训练。我们提出几何变化VLA(Geometry-Change VLA, GC-VLA),它学习从当前观测中预测多视角的未来-当前几何变化标记(geometry-change tokens)。离线帧对定义了标称0.5秒的预测时域;未来观测仅用于构建训练目标。第一阶段训练几何变化视觉-语言模型(GC-VLM)。第二阶段引入连续的ActionExpert,并将其与机器人动作对齐,同时在VLM接口处停止动作流的梯度。第三阶段使这些梯度能够与ActionExpert共同更新可训练的VLM组件。第四阶段冻结GC-VLA,并应用几何条件残差流(Geometry-Conditioned Residual Flow, GCRF),使用一个二值干预路由器和从闭环反馈中学到的单一有界残差速度策略。GC-VLA在LIBERO上达到95.20%的成功率,结合GCRF的GC-VLA达到99.55%。推理时仅使用当前观测和学到的GC表示,无需运行离线目标编码器。
cs.RO / 22 / 2609.25562

IndustrialVLA-Bench: A Traceable Multi-Axis Evaluation of Open Robot Policy Models

IndustrialVLA-Bench:面向开放机器人策略模型的可追溯多维度评估
Wang, Yiqi, Rao, Zhifeng, Zhang, Jiaqi, Li, Xiaoyang, Wu, Zhangkai, Duan, Yiqun, Zheng, Mingkai, Wang, Fei, You, Shan, Cai, Taotao
Abstract
Open robot policies increasingly follow two paradigms: vision-language-action models (VLAs) directly map observations and instructions to actions, whereas world-action models (WAMs) incorporate learned video or world dynamics into policy learning or action generation. Although both target the same manipulation tasks and represent alternative design choices, they are commonly reported under different evaluation protocols, leaving their capability, robustness, language sensitivity, and deployment-cost trade-offs unclear. We present IndustrialVLA-Bench, an evidence-aware evaluation of six released VLA and WAM systems under a unified reporting schema. It separately evaluates clean capability on LIBERO, non-language robustness on LIBERO-Plus, instruction sensitivity on LIBERO-Para, and observed execution cost. Reported task scores aggregate three complete evaluations with distinct random seeds under a fixed checkpoint and inference configuration. Across all six systems, clean LIBERO averages differ by only 1.58 points, whereas robustness and paraphrase summaries span 14.62 and 31.08 points. Restricting every comparison to the three protocol-faithful systems preserves the effect (1.36, 14.62 and 23.10 points), so the diagnostic separation reported here does not depend on the weaker evidence tiers. We additionally report observed inference latency, peak memory, runtime mode, and an evidence status for every system. Protocol-faithful, near-reproduction, and pending-verification entries remain visibly separated; only protocol-faithful entries support strict comparisons. Rather than claiming universal superiority of either paradigm, IndustrialVLA-Bench provides traceable evidence for comparing released robot policies on shared practical criteria. Code and evaluation records are available at https://github.com/xiaoqi-7/IndustrialVLA-Bench.
Chinese Translation
开放机器人策略正日益沿循两种范式:视觉-语言-动作模型(VLA)将观测与指令直接映射为动作,而世界-动作模型(WAM)则将学习到的视频或世界动态引入策略学习或动作生成之中。尽管二者面向相同的操作任务并代表了不同的设计选择,但它们通常在不同的评估协议下进行报告,导致其能力、鲁棒性、语言敏感性以及部署成本之间的权衡尚不清楚。我们提出 IndustrialVLA-Bench,在统一的报告模式下对六个已发布的 VLA 与 WAM 系统进行证据感知的评估。该基准分别在 LIBERO 上评估纯净能力,在 LIBERO-Plus 上评估非语言鲁棒性,在 LIBERO-Para 上评估指令敏感性,并观测执行成本。所报告的任务分数是在固定的检查点与推理配置下,采用三个不同随机种子完成的完整评估的聚合结果。在全部六个系统中,LIBERO 纯净能力平均分仅相差 1.58 分,而鲁棒性与指令改写敏感性指标分别相差 14.62 分和 31.08 分。即使将所有比较限制在三个严格遵循协议的系统上,上述效应依然存在(分别为 1.36、14.62 和 23.10 分),因此本文报告的诊断性区分并不依赖于较弱层级的证据。此外,我们还报告了每个系统的观测推理延迟、峰值内存、运行时模式以及证据状态。严格遵循协议的、接近复现的和待验证的条目之间保持着明显的区分;只有严格遵循协议的条目才支持严格的比较。IndustrialVLA-Bench 并不声称任何一种范式具有普遍优越性,而是为在共享的实践性标准上比较已发布的机器人策略提供可追溯的证据。代码与评估记录可在 https://github.com/xiaoqi-7/IndustrialVLA-Bench 获取。
cs.RO / 23 / 2609.25577

Recording Hand-Held Laparoscopic Instrument Motion in the Operating Room: Magnetometer-Free Fusion of Inertial, Range and Visual Sensing

在手术室中记录手持式腹腔镜器械运动:基于惯性、测距与视觉传感的无磁力计融合方法
Lee, Jiyul, Yee, Dongho, Oh, Juahn, Lee, Jinseok, Seo, Yechan, Jeong, Seong, Kim, Minsung, Shim, Seonho, Noh, Younghoon, Choi, Hyuk, Kong, Youngbin, Kong, Hyoun-Joong
Abstract
Most minimally invasive procedures are still performed with hand-held laparoscopic instruments, yet only the endoscopic video is retained; the instrument motion that expresses surgical skill, and that could support skill assessment and robot learning, is lost. Pose from video alone remains millimeters to centimeters off, and an instrument-mounted inertial measurement unit (IMU) cannot rely on its magnetometer, whose field changed with tool pose and between sessions in our measurements. We present a surgical instrument-state logger that clips onto a conventional instrument without modifying the part that enters the patient and fuses a six-axis IMU and a time-of-flight (ToF) rangefinder with a markerless camera in an error-state Kalman filter under the remote center of motion (RCM) of the trocar. Heading comes from the shaft silhouette, segmented by a U-Net, in place of the magnetometer: the rotation-angle error is 0.200{\deg}, against 3.58{\deg} from the accelerometer and magnetometer alone. Against a Franka Research 3 manipulator, and without alignment to it, the displacement error over 300 translation trials was 1.21mm RMS and the relative-rotation error over 180 rotation trials 0.34{\deg} RMS. On continuous trajectories, tracked and displayed in real time, the absolute tip error was 1.22mm (programmed) and 3.04mm (teleoperated) after post-hoc tuning of three filter parameters, and the full fusion beat every sensor subset. Because the estimator uses no magnetic measurement, its accuracy does not rely on an undisturbed field. The same clip-on device could thus record metric tip trajectories during routine hand-held laparoscopy, while displaying the insertion depth and attitude that are hidden once the instrument is inside the patient.
Chinese Translation
大多数微创手术仍使用手持式腹腔镜器械进行,然而目前仅保留内镜视频;能够体现手术技能、并可用于技能评估与机器人学习的器械运动信息却随之丢失。仅凭视频估计的位姿误差可达毫米至厘米量级,而安装在器械上的惯性测量单元(IMU)无法依赖其磁力计——在我们的测量中,磁场随工具姿态变化且在不同场次之间也不一致。我们提出了一种手术器械状态记录器,它可夹持在常规器械上而不改动进入患者的部分,并在套管针的远程运动中心(RCM)约束下,通过误差状态卡尔曼滤波器融合六轴IMU、飞行时间(ToF)测距仪与无标记相机。航向信息来自由U-Net分割的器械杆轮廓,以替代磁力计:旋转角误差为0.200度,而仅用加速度计和磁力计时为3.58度。在与Franka Research 3机械臂(且未与其对准)的对比中,300次平移试验的位移误差为1.21mm均方根(RMS),180次旋转试验的相对旋转误差为0.34度RMS。在实时跟踪与显示的连续轨迹上,经过三个滤波参数的事后调优后,器械尖端的绝对误差为1.22mm(程序化运动)和3.04mm(遥操作),且完整融合方案优于所有传感器子集。由于该估计器不使用任何磁性测量,其精度不依赖于无扰动的磁场。因此,这种夹持式装置可以在常规手持腹腔镜手术中记录具有度量信息的尖端轨迹,同时显示器械进入患者体内后不可见的插入深度与姿态。
cs.RO / 24 / 2609.25606

CableVLA: Simulation-Privileged Global-Local Representation Learning for Cable Routing

CableVLA:面向线缆布线的仿真特权全局-局部表征学习
Teng, Zhifei, Feng, Bo, Zou, Xiang, Xiao, Jinpeng, Li, Min, Yin, Zhouping, Li, Yiqun
Abstract
Cable routing requires coordinated control of global cable topology and changing local contacts. We present CableVLA, an end-to-end multimodal vision-language-action framework that converts simulation-privileged supervision into deployable cable-topology and tactile representations. TopoHead distills node-level physics and current and future cable-topology information into causal visual context for the action expert. TacSense uses complementary frame and taxel branches to learn contact dynamics from resistive arrays, with simulator-derived kinematics and contact events providing supervision beyond the measured force map. A contact gate activates force-tactile residuals that refine the next 8 arm-and-gripper actions of a frozen topology-conditioned policy. Across 345 MuJoCo evaluations, CableVLA improves success from 62.6% for the $\pi_{0.5}$-V visual baseline to 84.9%. TacSense achieves pronounced gains in slip-transition recognition over a CNN-LSTM baseline with a similar parameter count, and this advantage persists under frozen-encoder probes. Topology prediction and 57-task tactile evaluations assess representation quality, while policy adaptation studies evaluate downstream control performance. Cross-simulator and real-robot comparisons further examine zero-shot policy transfer under changes in dynamics and sensing.
Chinese Translation
线缆布线需要对全局线缆拓扑和不断变化的局部接触进行协同控制。我们提出CableVLA,这是一个端到端的多模态视觉-语言-动作框架,可将仿真特权监督转化为可部署的线缆拓扑表征与触觉表征。TopoHead将节点级物理信息以及当前与未来的线缆拓扑信息蒸馏为因果视觉上下文,供动作专家使用。TacSense利用互补的图像帧分支和触觉单元(taxel)分支,从电阻式触觉阵列中学习接触动力学,其中由仿真器导出的运动学与接触事件提供了超越实测力分布图的监督信号。接触门控机制激活力-触觉残差,用于对冻结的拓扑条件策略的后续8个手臂与夹爪动作进行细化。在345次MuJoCo评估中,CableVLA将成功率从$\pi_{0.5}$-V视觉基线的62.6%提升至84.9%。在参数量相近的情况下,TacSense在滑移转变识别方面相比CNN-LSTM基线取得了显著提升,且这一优势在冻结编码器探测下依然保持。拓扑预测与57项任务的触觉评估用于衡量表征质量,而策略适配研究则评估下游控制性能。跨仿真器与真实机器人对比进一步检验了在动力学与感知变化下的零样本策略迁移能力。
cs.RO / 25 / 2609.25614

A Deployable Four-Finger Payload for Teleoperated Free-Flying Manipulation with Astrobee

一种用于Astrobee遥操作自由飞行抓取的可展开四指载荷
Su, William, Kam, Jordan, Nakamura, Yunosuke, Wang, Yixiao, Zhou, Jianshu, Tomizuka, Masayoshi
Abstract
This article presents a bimanual teleoperation pipeline and conceptual design of a deployable four-finger payload for intra-vehicular free-flyers. Future habitats in low-Earth orbit (LEO) will require systems to perform mundane tasks like cargo handling and maintenance during crewed and uncrewed periods. The gripper payload provides 17 manipulation degrees-of-freedom (DoF) through four independently actuated fingers on a linear rail system. To control it, a virtual reality (VR) device interface maps the human ground operator's hand motions to the finger pairs, their separation to the rail, and common wrist motion to Astrobee translation. We present the preliminary results of teleoperating Astrobee in a custom zero-gravity MuJoCo-based International Space Station (ISS) simulator through ten repeated trials of transporting a rigid ISS Cargo Transfer Bag (CTB). We measure task success, continuous contact retention, completion time, and cargo motion.
Chinese Translation
本文提出了一种双手机器人遥操作流程以及面向舱内自由飞行机器人的可展开四指载荷概念设计。未来的近地轨道(LEO)居住舱需要在有乘员和无乘员期间执行货物搬运和维护等日常任务的系统。该夹持器载荷通过安装在线性导轨系统上的四个独立驱动手指提供17个操作自由度(DoF)。在控制方面,虚拟现实(VR)设备接口将地面操作员的手部动作映射到手指对上,将其手部开合动作映射到导轨间距,并将腕部的共同运动映射为Astrobee的平移。我们在一个基于MuJoCo的自定义零重力国际空间站(ISS)仿真器中,通过十次重复试验 transporting 一个刚性ISS货物转运袋(CTB)来遥操作Astrobee,并给出了初步结果。我们测量了任务成功率、连续接触保持情况、完成时间以及货物运动情况。
cs.RO / 26 / 2609.25619

Relative Contact Velocity-Controlled Hand-Object Mechanism for Dexterous Tool Manipulation

Wang, Sunyu, Oh, Jean, Pollard, Nancy S.
Abstract
This work investigates how to enable general multi-finger robotic hands to perform the complete tool manipulation process, which entails picking up a tool, loading it into a suitable pose, and then wielding it. Inspired by human tool manipulation and mechanical design principles, we model the hand and the tool as a unified hand-object mechanism (HOM) composed of sub-assemblies. Specifically, we define a HOM as consisting of the hand, the object, and the generalized contact frames, allowing the HOM's motions to be expressed with the same set of Cartesian-space relative contact velocities, irrespective of the hand's kinematics and geometry. Then, we define a HOM's sub-assemblies as relative contact velocity and contact force constraints between fingers. Building on these definitions, we developed a lightweight and physically interpretable motion planning and contact estimation framework using least squares and a complementary filter. We evaluated our framework in simulation by teleoperating five different robotic hands. The results show that our framework enabled all five hands to execute the complete tool manipulation process, achieving dexterous behaviors even from identical, simple reference trajectories. Furthermore, the results showcase our framework's adaptability to different hands, tools, and tasks, enabled by its kinematic and geometric foundation.
cs.RO / 27 / 2609.25625

From Instrument-Mounted Demonstrations to In-Vivo Execution: Learning Bimanual Laparoscopic Appendectomy Without Robot-Collected Demonstrations

从器械-mounted演示到体内执行:无需机器人采集演示的双臂腹腔镜阑尾切除术学习
Yee, Dongho, Oh, Juahn, Lee, Jinseok, Lee, Jiyul, Seo, Yechan, Jeong, Seong, Kim, Minsung, Shim, Seonho, Noh, Younghoon, Choi, Hyuk, Kong, Youngbin, Lee, Kyu Eun, Kong, Hyoun-Joong
Abstract
Most minimally invasive surgery is still performed with hand-held laparoscopic instruments, and the surgeon's instrument kinematics are lost when the operation ends; only the endoscope video is kept. This paper presents an end-to-end pipeline that captures this motion in the operating room and uses it to train a surgical robot policy, validated on live animals. We introduce a surgical instrument-state logger that mounts on the shaft of a standard laparoscopic instrument and recovers its pose and jaw state from an inertial sensor, a time-of-flight sensor and a Hall sensor, with no external camera or tracker. A data pipeline measures the latency of every sensor channel against a robot ground truth and aligns the channels before forming observation-action pairs. On these demonstrations we train a diffusion policy with a fine-tuned DINOv3 backbone, selecting its design by closed-loop rollouts in a physics simulator reconstructed from depth maps of an ex-vivo rabbit appendix. The policy is then retrained on 849 in-vivo demonstrations from four live rabbits and deployed on four additional live rabbits with electrosurgery armed. With the surgeon selecting the surgical phase, the policy completed the appendectomy in three of the four animals. The results show that demonstrations recorded from a surgeon's own instruments are sufficient to train, select and deploy a bimanual surgical policy in vivo. The robot serves only as the timing reference for sensor calibration and as the executor, and collects no demonstrations. Both demonstration corpora are released to support future surgical robot learning research.
Chinese Translation
大多数微创手术仍使用手持式腹腔镜器械进行,而手术结束时外科医生的器械运动学信息随之丢失,仅有内镜视频被保留。本文提出一个端到端的流程,在手术室中捕捉这些运动并利用其训练手术机器人策略,并在活体动物上进行了验证。我们介绍了一种手术器械状态记录器,其安装在标准腹腔镜器械的杆身上,通过惯性传感器、飞行时间(time-of-flight)传感器和霍尔传感器恢复器械的位姿和钳口状态,无需外部相机或跟踪器。数据处理流程测量每个传感器通道相对于机器人地面真值的延迟,并在构建观测-动作对之前对各通道进行对齐。在这些演示数据上,我们使用经过微调的DINOv3骨干网络训练扩散策略(diffusion policy),并通过在由离体兔阑尾深度图重建的物理仿真器中进行的闭环回放来选择其设计方案。随后,该策略在来自四只活体兔子的849条体内演示数据上重新训练,并部署到另外四只装备了电外科器械的活体兔子身上。在外科医生选择手术阶段的情况下,该策略在四只动物中的三只上完成了阑尾切除术。结果表明,从外科医生自身器械记录的演示数据足以在体内训练、筛选和部署双臂手术策略。机器人仅作为传感器校准的时间参考和执行器,不采集任何演示数据。两个演示数据集均已公开,以支持未来的手术机器人学习研究。
cs.RO / 28 / 2609.25627

MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence

Wen, Haoran, Wang, Wenfu, Shi, Kunsong, Wang, Jingke, Feng, Wancheng, Zhang, Yiren, Zhao, Yueran, Zhang, Xuancheng, Ye, Nanfei, Chen, Xingru, Sun, Zhaohong, Yang, Chengmin, Yu, Zikang, Bi, Penghao, Shi, Jia, Liu, Yu, Zhan, Kun, Xie, Yan
Abstract
General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual prediction with control without necessarily exposing the task-relevant semantic and spatial structure needed for fine-grained manipulation. We present MachEmbodied-U0 (ME-U0), a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture. Subtask prediction and affordance grounding guide joint visual-dynamics and action generation via flow matching. Visual dynamics encompass future RGB, depth, surface normals, and optical flow, providing complementary supervision for appearance, geometry, and motion. Multi-rate Rotary Position Encoding (MRPE) aligns visual dynamics with fine-grained control. We pretrain ME-U0 on approximately 4,200 hours of curated demonstrations from robotic datasets and egocentric datasets. Using only the supervision natively available in each downstream benchmark, ME-U0 achieves an average score of 17.66 on the RoboDojo simulation benchmark and average success rates of 99.0\% and 82.5\% on LIBERO and LIBERO-Plus, respectively. We additionally validate ME-U0 on real-world robotic manipulation tasks, demonstrating its effectiveness beyond simulation. Without corresponding downstream supervision, ME-U0 further demonstrates zero-shot subtask prediction, affordance grounding, and visual dynamics on simulated and real-world observations. Overall, ME-U0 combines competitive downstream control performance with transferable task-grounding and visual-dynamics capabilities across simulation and the real world.
cs.RO / 29 / 2609.25630

PAKT: Physically-Aligned Kinesthetic Teaching for Reinforcement Learning

PAKT:面向强化学习的物理对齐触觉示教
Johannsmeier, Lars, Narang, Yashraj
Abstract
Real-world reinforcement learning (RL) systems still struggle with the demands of contact-rich industrial manipulation, including micrometer-level precision, success rates above 99%, and human-level cycle times. Although off-policy algorithms can improve performance by leveraging demonstrations and interventions, a key bottleneck is the lack of an intuitive interface for collecting such guidance while complying with constraints of the physical system and the policy. We propose PAKT, a framework for kinesthetic teaching in RL. As opposed to teleoperation approaches, PAKT relies on kinesthetic guidance, which is widely used in industry. However, a critical weakness of kinesthetic guidance is the possibility for the operator to move the robot along trajectories (e.g., velocities, accelerations, jerk) that the robot and/or policy cannot physically reproduce. Using PAKT, operators guide the robot through admittance control, which maps human-applied forces to motion. The downstream reference generator applies the same kinematic limits used during policy execution, keeping the collected trajectories within these limits. To support this teaching interface with an appropriate execution layer, PAKT adds a high-performance control stack that maps low-frequency RL actions to high-frequency torque commands. It consists of a reference generator and subsequent impedance controller, where the reference generator preserves the tracking performance of the impedance controller while improving contact handling and producing smoother policy actions. Across the reported runs on four insertion and industrial assembly benchmarks, including a data center compute tray, the end-to-end system reduces cycle time by 23%-48% and cumulative intervention count by 62%-86% relative to the HIL-SERL baseline. Project website: https://pakt-website.github.io/pakt-website}{https://pakt-website.github.io/pakt-website
Chinese Translation
现实世界的强化学习(RL)系统仍然难以满足接触密集型工业操作的要求,包括微米级精度、高于99%的成功率以及与人类水平相当的节拍时间。尽管离线策略(off-policy)算法可以利用演示和干预来提升性能,但一个关键瓶颈在于缺乏一个直观的界面来收集此类引导,同时遵守物理系统和策略的约束。我们提出了PAKT,一个用于强化学习触觉示教的框架。与遥操作方法不同,PAKT依赖于在工业界广泛使用的触觉引导。然而,触觉引导的一个关键弱点是操作者可能使机器人沿机器人本身和/或策略在物理上无法复现的轨迹(例如速度、加速度、加加速度)运动。使用PAKT时,操作者通过导纳控制(admittance control)引导机器人,该控制将人类施加的力映射为运动。下游的参考生成器施加与策略执行期间相同的运动学限制,使采集的轨迹保持在这些限制之内。为了以合适的执行层支持这一示教界面,PAKT添加了一个高性能控制栈,将低频RL动作映射为高频力矩指令。该控制栈由参考生成器和后续的阻抗控制器组成,其中参考生成器在保持阻抗控制器跟踪性能的同时,改善了接触处理能力并产生更平滑的策略动作。在四个插入和工业装配基准任务(包括数据中心计算托盘)上的所有测试运行中,与HIL-SERL基线相比,该端到端系统将节拍时间缩短了23%–48%,累计干预次数减少了62%–86%。项目网站:https://pakt-website.github.io/pakt-website
cs.RO / 30 / 2609.25631

DynaForge: Planning-Guided Residual Learning for Dynamic Manipulation Demonstration Generation

DynaForge:面向动态操作示范生成的规划引导残差学习
Jin, Yiyang, Zheng, Yu, He, Xiao, Wang, Hesheng
Abstract
Dynamic object manipulation is essential for robots operating in real-world environments, yet methods for generating high-quality demonstrations remain limited. Methods designed for static tasks do not readily transfer to dynamic settings. Among dynamic demonstration generators, planning-based methods can fail near contact, while DOMINO-style replay simplifies dynamic interactions and may limit the experience available for policy learning. We present DynaForge, a planning-guided framework that learns residual corrections for dynamic manipulation demonstration generation. DynaForge combines low-frequency global planning with high-frequency object-centric inverse kinematics across task phases, and applies a residual policy to correct actions during dynamic interaction. An implicit curriculum groups rollouts under matched conditions and selects mixed-success groups, focusing residual reinforcement learning on the evolving competence frontier. On Can and Bottle, it uses 0.73x as many optimizer steps as vanilla GRPO at the same nominal environment-step budget, with higher observed final success rates. Across nine simulation tasks, DynaForge increases mean demonstration-generation success from 41.30% of the planning prior to 78.37%. With 800 demonstrations per task, DP3 policies trained on DynaForge data achieve 49.11% mean success, compared with 7.07% for DOMINO data. On three real-world dynamic tasks, DynaForge-trained policies achieve 30-60% success, compared with 0-10% for DOMINO-trained policies, showing the ability of DynaForge for sim-to-real transfer.
Chinese Translation
动态物体操作对于机器人在真实环境中的运作至关重要,然而生成高质量示范的方法仍然有限。面向静态任务设计的方法难以直接迁移到动态场景。在现有的动态示范生成方法中,基于规划的方法在接近接触阶段容易失效,而 DOMINO 风格的重放方法则简化了动态交互,可能限制策略学习可利用的经验。我们提出 DynaForge,这是一个规划引导的框架,通过学习残差修正来生成动态操作示范。DynaForge 在任务各阶段将低频全局规划与高频以物体为中心的逆运动学相结合,并在动态交互过程中应用残差策略对动作进行修正。一种隐式课程机制将条件匹配的回合(rollout)分组,并选择混合成功率的组别,使残差强化学习聚焦于不断演进的能力边界。在 Can 与 Bottle 任务上,在相同的标称环境步数预算下,该方法使用的优化器步数仅为原始 GRPO 的 0.73 倍,同时观察到更高的最终成功率。在九个仿真任务中,DynaForge 将示范生成的平均成功率从规划先验的 41.30% 提升至 78.37%。在每任务 800 条示范的条件下,基于 DynaForge 数据训练的 DP3 策略达到 49.11% 的平均成功率,而基于 DOMINO 数据训练的策略仅为 7.07%。在三个真实世界动态任务上,DynaForge 训练的策略达到 30–60% 的成功率,而 DOMINO 训练的策略仅为 0–10%,展现了 DynaForge 在仿真到真实(sim-to-real)迁移方面的能力。
cs.RO / 31 / 2609.25636

RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents

RoboFollow:揭示具身智能体中的指令跟随幻象
Guo, Chang, Xie, Yukun, Tan, Bohan, Chang, Zheng, Yin, Zhaokai, Ma, Qianli, Wang, Yingqiao, Liang, Chao, Zhang, Zhipeng
Abstract
Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and a policy can score highly while barely using it. We introduce RoboFollow, a diagnostic benchmark with three principles: (1) High Scene Entropy: each training scene supports multiple kinematically distinct task branches, making vision alone insufficient and forcing reliance on language. (2) Hierarchical Diagnostic Protocol: a four-level protocol (L0--L3) progressively perturbs visual layout and semantics, probing whether equivalent instructions yield consistent behavior and distinct ones yield discriminable behavior across spatial relations, attributes, trajectory constraints, and logic. (3) Confound-Controlled Diagnosis: we simplify interaction objects, restrict actions to the trained repertoire and report stage-wise Intent and Execution scores, isolating comprehension from motor execution. Evaluation of nine VLA and WAM policies shows that strong L0 performance, where attained, does not reliably transfer to L1--L3 under our fine-tuning setup. Representative mitigations, including stronger VLM backbones, QA co-training, LangForce, and Classifier-Free Guidance, all fail to close this gap. RoboFollow exposes genuine instruction following as a critical, overlooked bottleneck. Code and dataset are available at https://github.com/AutoLab-SAI-SJTU/RoboFollow and https://huggingface.co/datasets/AutoLab-SJTU/robofollow-data.
Chinese Translation
现代具身智能体(embodied agents)取得了令人瞩目的成功率,但其实际的指令跟随能力远弱于这些数字所显示的水平。我们将这种幻觉归因于一种结构性特征,即低场景熵(low scene entropy):当视觉场景只存在唯一有效任务时,语言变得冗余,策略即使几乎不使用语言也能获得很高的分数。我们提出了 RoboFollow,一个基于三条原则的诊断基准:(1)高场景熵:每个训练场景支持多个运动学上截然不同的任务分支,使得仅凭视觉不够充分,从而迫使模型依赖语言。(2)分层诊断协议:一个四级协议(L0--L3)逐步扰动视觉布局与语义,以探究等价指令是否产生一致的行为、不同指令是否产生可区分的行为,涵盖空间关系、属性、轨迹约束和逻辑等维度。(3)混淆因素控制的诊断:我们简化交互对象,将动作限制在训练过的动作库内,并按阶段报告意图(Intent)和执行(Execution)分数,从而将语言理解与运动执行相分离。对九种 VLA 和 WAM 策略的评估表明,即使在 L0 上取得了较强性能,在我们的微调设置下也不能可靠地迁移到 L1--L3。各类代表性缓解方法,包括更强的 VLM 骨干网络、QA 联合训练、LangForce 以及无分类器引导(Classifier-Free Guidance),均无法弥合这一差距。RoboFollow 揭示了真正的指令跟随是一个关键却被忽视的瓶颈。代码和数据集可在 https://github.com/AutoLab-SAI-SJTU/RoboFollow 和 https://huggingface.co/datasets/AutoLab-SJTU/robofollow-data 获取。
cs.RO / 32 / 2609.25639

A Reconfigurable Bidirectional Cable-Driven Hip Exoskeleton with Swappable Bench/Backpack Dual-configuration Actuation

一种可重构的双向缆绳驱动髋部外骨骼,具有可切换的台式/背包式双构型驱动系统
Ji, YuanLong, Xiang, Shuhan, Ye, Qihan, Yang, Xingbang
Abstract
Hip exoskeletons provide an important hardware basis for lower-limb rehabilitation and locomotor assistance. Laboratory rehabilitation assessment and system development require substantial actuation and computing resources, whereas mobile assistance requires untethered portability. Integrating both capabilities within one reusable platform remains a central design challenge. This paper presents a reconfigurable bidirectional cable-driven hip exoskeleton platform that rapidly switches between bench-mounted and backpack-mounted actuation while sharing one cable-free wearable hip interface. The platform modularly adapts the actuation configuration, end-effector sensing path, and low-level control interface. Each cable-driven end-effector weighs 0.405 kg, excluding the cable and actuation unit, and integrates an encoder and a torque sensor; experiments validated bench-mounted admittance-based motion tracking capability and backpack-mounted open-loop torque tracking. Human-worn experiments with three healthy participants used myoMOTION to evaluate the platform's wearable-side hip-motion sensing capability, verified bench-to-backpack and backpack-to-bench motion-ready switching across 30 trials in $30.1\pm16.3$ s, and formed a small-scale multimodal wearable-exoskeleton gait dataset for sensing validation and data-driven algorithm development, comprising 8 min bench-mounted treadmill records and 11 min backpack-mounted outdoor walking records. These results show that, by unifying the wearable structure, actuation interface, and sensing path, the proposed platform enables validation of the same hip exoskeleton in both bench-mounted and backpack-mounted configurations, providing reusable hardware for iterative development and applications across scenarios.
Chinese Translation
髋部外骨骼为下肢康复和运动辅助提供了重要的硬件基础。实验室康复评估与系统开发需要大量的驱动和计算资源,而移动式辅助则要求无线便携性。如何将这两种能力集成于一个可复用平台仍是核心设计挑战。本文提出一种可重构的双向缆绳驱动髋部外骨骼平台,可快速在台式驱动与背包式驱动之间切换,同时共用一个无线缆的可穿戴髋部接口。该平台可模块化地调整驱动构型、末端执行器传感路径和底层控制接口。每个缆绳驱动末端执行器重量为0.405 kg(不含缆绳和驱动单元),并集成了编码器和扭矩传感器;实验验证了台式构型下基于导纳的运动跟踪能力,以及背包构型下的开环扭矩跟踪能力。在三名健康受试者的人体穿戴实验中,使用myoMOTION评估了平台的穿戴侧髋部运动感知能力,并在30次试验中验证了台式到背包、背包到台式的运动就绪切换,切换时间为30.1±16.3 s;同时构建了一个小规模多模态可穿戴-外骨骼步态数据集,用于感知验证和数据驱动算法开发,包含8分钟的台式跑步机记录和11分钟的背包式户外行走记录。结果表明,通过统一可穿戴结构、驱动接口和传感路径,所提出的平台能够在台式和背包式两种构型下对同一髋部外骨骼进行验证,为跨场景的迭代开发与应用提供了可复用的硬件。
cs.RO / 33 / 2609.25642

Contact-Stable Deformable Tissue Simulation Using Implicit Integration and Live-Pose Grasp Constraints for Laparoscopic Surgery Robot Policy Evaluation

基于隐式积分与实时位姿抓取约束的接触稳定可变形组织仿真,用于腹腔镜手术机器人策略评估
Oh, Juahn, Yee, Dongho, Lee, Jinseok, Lee, Jiyul, Seo, Yechan, Jeong, Seong, Kim, Minsung, Shim, Seonho, Noh, Younghoon, Choi, Hyuk, Kong, Youngbin, Kon, Hyoun-Joong
Abstract
Closed-loop evaluation of surgical robots requires tissue that deforms, can be grasped and lifted, and reproduces the anatomy in which the robot will operate. We present a simulator in which this tissue is reconstructed from a fixed-view RGB-D recording of the surgical field, composited to remove the instruments, closed into watertight volumes and tetrahedralised; the pipeline was applied unchanged to three specimens of two species (thirteen organs, 146,061 tetrahedra, no inverted elements). For one specimen, the organs are placed in a bimanual cell in which two Franka FR3 arms operate motorised instruments through 6 mm trocars. The core contribution is the numerical and contact design that keeps this cell stable: implicit integration, simulation meshes separate from collision meshes, numerical guards, and a grasp constraint captured at the live tissue pose. In 45 repeated grasp-lifts, a friction grasp held the tissue in 0 of 15 trials and each constraint grasp in 13 of 15; on displaced tissue, a rest-pose constraint produced one-step snaps of up to 17.8 mm, which live-pose capture eliminates. Against the recording, front-surface depth error is 1.33 to 1.41 mm, organ silhouette IoU is 0.80, and in five grasp-lifts reproduced from video the landmark displacement RMSE is 11.8 mm against 14.2 mm for a static prediction. Biofidelity is not claimed; the environment is intended for closed-loop feasibility, safety, contact and policy screening.
Chinese Translation
手术机器人的闭环评估需要能够形变、可被抓取和提起、并能复现机器人实际操作解剖结构的组织。我们提出了一种仿真器,其中的组织通过对手术视野的固定视角RGB-D录像重建得到:合成处理以移除手术器械,闭合成水密体积并进行四面体化;该流程未作修改地应用于两个物种的三例标本(十三个器官,146,061个四面体,无反转单元)。对于其中一例标本,器官被置于双手操作单元中,两个Franka FR3机械臂通过6 mm套管操作电动器械。核心贡献在于保持该单元稳定性的数值与接触设计:隐式积分、仿真网格与碰撞网格分离、数值防护措施,以及在组织实时位姿处捕获的抓取约束。在45次重复抓取-提起实验中,摩擦抓取在15次试验中0次成功保持组织,而每种约束抓取均在15次试验中13次成功;在发生位移的组织上,静止位姿约束能产生高达17.8 mm的单步突变位移,而实时位姿捕获可消除该现象。与原始录像相比,前表面深度误差为1.33至1.41 mm,器官轮廓IoU为0.80;在根据视频复现的五次抓取-提起中,标志性点位移RMSE为11.8 mm,而静态预测为14.2 mm。本文不声称具备生物保真度;该环境旨在用于闭环可行性、安全性、接触及策略筛选。
cs.RO / 34 / 2609.25649

Skill Sequence Planning for Collaborative Multi-Robot Construction

Wang, Xi, Fu, Bo, Menassa, Carol C., Kamat, Vineet R., Deng, Min
Abstract
Robots have significant potential to automate construction processes. However, their industry adoption remains limited, partly because of the programming effort required to adapt robots to diverse tasks. This paper presents a skill sequence planning method that enables a heterogeneous team of multi-functional robots to collaboratively perform construction assembly work using reusable, preprogrammed skills such as grasping, drilling, and fastening. A central controller transforms the digital representation of the building into a construction relationship graph that represents construction entities, their states, and their parent-child relationships. Based on this representation, the system selects the next construction target, generates a symbolic sequence of skills for capable members of the robot team, and produces collision-free geometric motion plans for skill execution. The symbolic planning problem is dynamically regenerated as the construction state changes. An interactive digital twin presents the planned skill sequence and robot states to human co-workers for review and approval before execution. The method is evaluated through a construction assembly case study. By reducing the need to program robots separately for each task variation, the proposed approach supports more flexible deployment of collaborative robot teams in construction.
cs.RO / 35 / 2609.25653

PhyVisGen: Physically and Visually High-Fidelity Robotic Manipulation Data Generation

Zheng, Yu, Feng, Qiyu, Wu, Yixin, Yang, Baoquan, Zhou, Yixuan, Hu, Bingyang, Huang, Kemeng, Yang, Guansheng, Wang, Hesheng
Abstract
Large-scale manipulation demonstrations are essential for learning robust visuomotor policies, yet real-world data collection is expensive and difficult to scale. Simulation offers a promising alternative, but physical and visual discrepancies can limit the transferability of synthetic data, particularly for manipulation with soft grippers. We present PhyVisGen, a physically and visually high-fidelity framework for scalable robotic manipulation data generation. On the physical side, PhyVisGen introduces an arm-gripper coupling method based on the Incremental Potential Contact (IPC), enabling high-fidelity soft contact throughout complete manipulation trajectories. On the visual side, it combines real-scene reconstruction with real-time path tracing to generate visually realistic observations while preserving captured scene appearance. Quantitative evaluations demonstrate the physical and visual fidelity of PhyVisGen. Policies trained exclusively on synthetic manipulation demonstrations achieve 65-95% success across five real-robot tasks, without real-robot demonstration data or policy fine-tuning.
cs.RO / 36 / 2609.25654

CODA: Depth-Aligned Scene Completion and Object Decomposition from a Single RGB-D Image

CODA:基于单张RGB-D图像的深度对齐场景补全与物体分解
Son, Dongwon, Han, Junhyek, Cho, Yoontae, Lee, Minseok, Choi, Hong-seok, Choi, Jiwook, Kim, Hyungjin, Kim, Beomjoon
Abstract
Robots operating safely in cluttered everyday environments often need to infer scene geometry from partial observations. Methods that detect objects in 2D and reconstruct them independently struggle in such scenes: a missed object is never reconstructed, a merged detection can fuse two objects, and separately reconstructed meshes may overlap or fail to touch their supporting surfaces. We introduce CODA (Complete Once, Decompose Afterward), a generative model that instead reconstructs the complete scene geometry from a single unsegmented RGB-D image, then separates the surface into the surrounding environment and movable objects. Still, generated scene geometry can drift from the observed partial point cloud. To reduce this drift, CODA uses two explicit 3D grounding mechanisms to keep reconstructed geometry consistent with observed surfaces while completing unseen regions. Experiments on HomebrewedDB and our custom cluttered-scene dataset show more accurate reconstructions and a higher fraction of objects remaining in place under simulated gravity than both object-first and scene-first baselines.
Chinese Translation
在杂乱的日常环境中安全作业的机器人通常需要从部分观测中推断场景几何。先在2D中检测物体再独立重建的方法在此类场景中表现不佳:漏检的物体永远不会被重建,一次合并的检测可能将两个物体融合在一起,而单独重建的网格可能相互重叠或无法接触到其支撑表面。我们提出了CODA(Complete Once, Decompose Afterward,先补全后分解),一种生成模型,它从单张未分割的RGB-D图像重建完整的场景几何,然后将表面分离为周围环境与可移动物体。然而,生成的场景几何可能偏离观测到的部分点云。为减小这种偏差,CODA采用两种显式的3D锚定(grounding)机制,在补全未见区域的同时保持重建几何与观测表面的一致性。在HomebrewedDB和我们自建的杂乱场景数据集上的实验表明,与物体优先和场景优先的基线方法相比,CODA实现了更精确的重建,并且在模拟重力条件下有更高比例的物体保持在原位。
cs.RO / 37 / 2609.25658

History-Conditioned Flow Matching for Probabilistic Dynamics of Tendon-Driven Continuum Robots

面向腱驱动连续体机器人概率动力学的历史条件流匹配方法
Yang, Hang, Liu, Tingcong, Xiong, Junjie, Yang, Fangju, Wu, Ke
Abstract
Deterministic dynamics modeling of tendon-driven continuum robots remains challenging owing to uncertainties in material behavior, tendon transmission, friction, and contact. Measured joint configurations and nominal tendon commands do not fully characterize these internal mechanical factors, leaving uncertainty in the subsequent motion. We therefore develop a history-conditioned, physics-informed flow-matching framework for probabilistic dynamics prediction, using motion and actuation histories to predict the distribution of the next complete joint configuration. By recursively sampling next-step configurations under prescribed commands, the model predicts distributions of future whole-body motions. In simulation, scenario-specific models achieve five-second trajectory Energy Scores (lower is better) of 12.05 mm under internal friction variation and 9.29 mm under unobserved actuation disturbances. Relative to the conditional variational autoencoder and diffusion baselines, Flow attains lower Energy Scores and coverage closer to the nominal level in both scenarios. Ablations support history and structural conditioning in both scenarios. On the physical robot, predictions under two tendon-command profiles excluded from training capture the principal motion sequences, with five-second Energy Scores of 11.91 and 11.42 mm, lower than the compared baselines. The predicted-to-measured spread ratios are 1.65 and 1.22 (closer to 1 is better). These results support history-conditioned probabilistic dynamics prediction under incomplete mechanical observations.
Chinese Translation
由于材料行为、腱传动、摩擦和接触等方面的不确定性,腱驱动连续体机器人的确定性动力学建模仍然具有挑战性。测量得到的关节构型和名义腱指令无法完全刻画这些内部机械因素,从而给后续运动留下不确定性。因此,我们开发了一种基于历史条件、融合物理信息的流匹配框架用于概率动力学预测,利用运动与驱动历史来预测下一步完整关节构型的分布。通过在给定指令下递归采样下一步构型,该模型可预测未来全身运动的分布。在仿真中,针对特定场景的模型在内部摩擦变化和未观测驱动扰动下分别取得了12.05 mm和9.29 mm的五秒轨迹能量分数(Energy Score,越低越好)。与条件变分自编码器(CVAE)和扩散模型基线相比,Flow在两种场景下均获得了更低的能量分数以及更接近名义水平的覆盖率。消融实验支持历史条件和结构条件在两种场景中的有效性。在物理机器人上,在两种未纳入训练的腱指令曲线下的预测捕捉到了主要运动序列,其五秒能量分数分别为11.91和11.42 mm,低于所比较的基线方法。预测与实测的离散度比值分别为1.65和1.22(越接近1越好)。这些结果支持了在不完整机械观测条件下的历史条件概率动力学预测。
cs.RO / 38 / 2609.25666

Deploying Foundation Models for Embodied Navigation

Dorbala, Vishnu Sashank, Manocha, Dinesh
Abstract
We present and tackle two problems associated with deploying Foundation Models (FMs) on Embodied Agents performing navigation: 1) Training bias in FMs leading to poor personalization in unseen environments, and 2) Limited FM context length hindering success, especially on long horizon tasks. Our solution for the former involves priming the FM with human-habit data mined from the scene and our solution for the latter involves active memory management via a novel `memory head' augmentation. We first present a taxonomy of existing literature on FM-based Embodied Navigation, and highlight these limitations. We then present our approaches, Transit-Aware Planning (TAP) and MemCtrl to address the limitations. With TAP, we present real-world results in a lab environment with a Turtlebot for personalized target finding that shows an average improvement of 18% over a non-TAP baseline. On MemCtrl, we report a 6% average improvement across various embodied tasks, with 20% on long instruction subsets, all while using nearly half the context used in the baseline model. Motivated by these result, we present our stance the deployability of FM-based embodied agents in real-world environments, and highlight open research directions.
cs.RO / 39 / 2609.25668

CDKF-Track: Cluster-aware Data-Driven Kalman Filtering for Cooperative 3D Multi-Object Tracking

Damanaki, Maria, Piperigkos, Nikos, Gkillas, Alexandros, Lalos, Aris S.
Abstract
Multi-Object Tracking (MOT) is essential for EdgeAI perception systems, where accurate object localization and reliable identification enable safe decision-making. Singleagent MOT suffers from occlusions, sensor noise, and partial scene understanding in complex real-world scenarios. While multi-agent systems improve robustness by exploiting shared information, they introduce redundant measurements that lead to false data associations, and still struggle to capture nonlinear object dynamics. To address these challenges, we propose CDKFTrack, a Cluster-aware Data-Driven Kalman Filtering framework for Cooperative 3D MOT. The proposed method first fuses multivehicle 3D LiDAR detections through a Graph Laplacian-based formulation. Then, a cluster-aware redundancy reduction scheme groups spatially related detections and selects representative observations to reduce duplicate inputs to the tracker. The resulting detections are processed by a data-driven Kalman filter that learns object motion dynamics from data, reducing dependence on predefined linear motion assumptions. Furthermore, a wavelet-based temporal refinement module leverages the multiresolution decomposition property of wavelets to attenuate shortterm positional fluctuations and improve trajectory continuity. To the best of our knowledge, CDKF-Track is the first framework to jointly address detection-level fusion redundancy and learnable motion modeling in cooperative 3D MOT. Experimental results on the real-world V2V4Real dataset indicate that CDKF-Track achieves up to 27.99% improvements in tracking accuracy over state-of-the-art multi-agent MOT methods.
cs.RO / 40 / 2609.25674

Teaching Reinforcement Learning and Humanoid Robotics to High-School Students: An Expert-Validated Curriculum Design on a Low-Cost Open Platform

Dong, Yuanzhe, Cao, Jie, Wang, Shuman
Abstract
Lower cost open source robots and reinforcement learning (RL) simulation tools create new opportunities for precollege students to engage with contemporary robotics. However, translating a complete research workflow, spanning mechanical assembly, electrical setup, simulation, policy learning, system identification, and physical deployment, into a coherent course for novice learners remains challenging. We present an integrated robotics course framework that organizes these activities around a shared robotic artifact. The framework combines parallel disciplinary tracks, sequencing based on technical dependencies, progressive integration of simulation and hardware, layered performance checkpoints, and structures for balancing collaborative work with individual accountability. We illustrate the framework through a high school curriculum organized around a robot project in which pairs of students assemble an open source humanoid robot, train a walking policy in simulation, and deploy it on the physical platform. The framework was developed through an iterative design process that included formative review by five experts in robotics research, engineering, secondary STEM education, and curriculum design. Expert feedback highlighted three central design tensions: authenticity versus cognitive load, system integration versus timely visible progress, and team construction versus individual accountability. These tensions informed the final framework presented in this paper. This work offers a structured approach for adapting robotics research workflows into interdisciplinary precollege courses; future classroom studies are needed to examine implementation and student learning.
cs.RO / 41 / 2609.25687

SG-CPG: Severity-Gated Central Pattern Generators for Adaptive Quadruped Locomotion under Continuous Actuator Degradation

SG-CPG:面向连续执行器退化下自适应四足运动的严重程度门控中央模式发生器
Kosta, Adarsh Kumar, Roy, Kaushik
Abstract
An animal with a weakened limb does not necessarily switch its gait, instead it unloads the affected limb, re-coordinates the remaining limbs, and scales its response with injury severity. This graded adaptation allows locomotion to persist despite partial loss of limb strength, rather than requiring a discrete transition between healthy and failed. Inspired by this behavior, we propose SG-CPG, a central pattern generator (CPG) for quadruped locomotion under continuous actuator degradation. SG-CPG preserves a frozen healthy CPG policy and introduces two severity-driven gates: a residual gate that re-coordinates all four legs and an amplitude gate that progressively shortens the weakened leg's stride as degradation increases. We emulate progressive degradation through two mechanisms: lowering the joint torque ceiling (ceiling mechanism) and scaling its low-level controller gains (gain mechanism), representing distinct forms of actuator weakening. Our simulations on a Unitree Go2 show that SG-CPG maintains a trot gait with 100% survival across an omnidirectional command schedule under 95% joint strength loss while tracking commands within 8%. Under a lowered torque ceiling, removing either severity path, the residual's severity observation or the amplitude gate, raises clipping at the weakened joint from 4.4% to 13.6% and 26.3% of steps at an 80% loss. On a real Go2, SG-CPG survives 28 of 29 forward and turning trials with up to 93% calf torque degradation. These results show that severity-gated adaptation can extend a healthy locomotion policy to progressive actuator degradation without treating the fault as a discrete failure.
Chinese Translation
肢体衰弱的动物并不一定切换步态,而是对受影响的肢体进行卸载,重新协调其余肢体,并根据损伤严重程度调整响应幅度。这种分级自适应使运动能够在肢体部分失去力量时得以维持,而无需在健康与失效状态之间进行离散切换。受此行为启发,我们提出了SG-CPG,一种面向连续执行器退化条件下四足运动的中央模式发生器(CPG)。SG-CPG保留一个冻结的健康CPG策略,并引入两个由严重程度驱动的门控:一个用于重新协调全部四条腿的残差门控,以及一个随退化加剧而渐进缩短衰弱腿部步幅的幅值门控。我们通过两种机制模拟渐进性退化:降低关节力矩上限(上限机制)和缩放其底层控制器增益(增益机制),以表征不同的执行器衰弱形式。在Unitree Go2上的仿真表明,SG-CPG在全向指令调度下,即使关节力量损失达95%,仍能以100%的存活率维持小跑步态,并将指令跟踪误差控制在8%以内。在降低力矩上限的情况下,移除任一严重程度路径——残差门控的严重程度观测或幅值门控——在80%力量损失时,会将衰弱关节的限幅比例分别从4.4%提高至13.6%和26.3%。在真实Go2机器人上,SG-CPG在小腿力矩退化高达93%的情况下,29次前进与转向试验中成功存活28次。这些结果表明,严重程度门控自适应能够在不将故障视为离散失效的前提下,将健康的运动策略扩展至渐进性执行器退化场景。
cs.RO / 42 / 2609.25688

MatcherCompass: A Deployment-Aware Benchmark to Guide Image Matcher Selection in the Wild

Kim, Hyunwoo, Kim, Giseop
Abstract
Field robots operating across time of day and sensing modalities require accurate image correspondences within onboard time and resource budgets. However, accuracy and runtime reported for individual methods on a single device provide limited guidance for choosing a matcher and its configuration on a target platform. We present MatcherCompass, a deployment-aware benchmark for choosing local feature matchers in field robotics. Under common input and pose-evaluation procedures, we compare nine classical and learned matching pipelines across four image resolutions and supported numerical precisions. Four visual conditions cover viewpoint variation, day--night matching in visible and thermal imagery, and daytime visible--thermal matching. We evaluate pose accuracy using the area under the error--recall curve (AUC) at $5^\circ$, $10^\circ$, and $20^\circ$, and measure runtime, GPU memory, and energy per image pair on four GPU platforms spanning workstation and onboard computers. The results show that changes in hardware, input resolution, and numerical precision can move a matcher across a runtime budget boundary, altering the feasible choices. We organize the measurements into a selection guide that returns all configurations satisfying user-specified time and resource limits, together with their accuracy under the selected visual condition. MatcherCompass provides measured evidence for choosing matching pipelines that fit a robot's sensing conditions and computing hardware. Project page: https://matchercompass.github.io/.
cs.RO / 43 / 2609.25689

MotionForge: A Data Generation Pipeline and Large-Scale Benchmark for Long-Horizon Manipulation of Dynamic Objects with Domain Shifts

MotionForge:一个面向具有域偏移的动态物体长时程操作的数据生成流水线与大规模基准测试
Liu, Mohan, Mei, Dengchen, Xian, Haotian, Han, Ruyang, Sun, Jiayi, Chen, Xuanyu, Zhang, Haitian, Li, Luxi, Mao, Kaimin, Wang, Lin
Abstract
Recent advances in learning-based robot policies have demonstrated promising progress, yet they are predom- inantly evaluated in static or quasi-static environments. In dynamic manipulation, objects and scenes continuously evolve while the robot perceives, reasons, and acts. However, recent dynamic simulation benchmarks largely focus on short-horizon, reactive interactions with simple motion patterns and offer limited support for both systematic evaluation under domain shifts and model-agnostic real-time execution protocols. To bridge these gaps, we introduce MotionForge, the first large- scale simulation benchmark and data-generation pipeline tailored to jointly evaluate domain shifts and long-horizon interaction in dynamic manipulation. MotionForge comprises 40 dynamic interaction tasks spanning 11 distinct motion patterns, with dedicated support for 17 long-horizon tasks. Our benchmark introduces two key novelties: (1) a systematic evaluation protocol for assessing policy robustness under both single-factor (e.g., only backgrounds shift) and joint domain shifts (e.g., simultaneous shifts of objects, backgrounds, lighting, and speed); and (2) a decoupled, latency-aware execution protocol where the environ- ment continuously evolves independently of policy inference time. Extensive evaluations of representative general-purpose robot policies on our benchmark reveal substantial limitations under joint domain shifts. These findings expose a critical gap between current policy capabilities and the requirements of robust long- horizon manipulation of dynamic objects under domain shifts, establishing MotionForge as a comprehensive testbed for future research in embodied AI.
Chinese Translation
基于学习的机器人策略的最新进展展现出可观的前景,但其评估大多在静态或准静态环境中进行。在动态操作中,物体和场景在机器人感知、推理和行动的同时持续演化。然而,近期的动态仿真基准大多聚焦于具有简单运动模式的短时程反应式交互,并且对域偏移下的系统性评估以及与模型无关的实时执行协议的支持均较为有限。为弥补这些不足,我们提出了MotionForge,这是首个专为联合评估动态操作中的域偏移与长时程交互而设计的大规模仿真基准和数据生成流水线。MotionForge包含40个动态交互任务,涵盖11种不同的运动模式,并专门支持17个长时程任务。我们的基准引入了两个关键创新点:(1)一套系统性评估协议,用于评估策略在单因素域偏移(例如仅背景变化)和联合域偏移(例如物体、背景、光照和速度同时变化)下的鲁棒性;(2)一种解耦的、时延感知的执行协议,其中环境独立于策略推理时间持续演化。我们在该基准上对代表性通用机器人策略进行了大量评估,结果表明这些策略在联合域偏移下存在显著局限。这些发现揭示了当前策略能力与在域偏移下对动态物体进行鲁棒长时程操作所需能力之间的关键差距,并确立了MotionForge作为未来具身智能研究的综合测试平台。
cs.RO / 44 / 2609.25695

Induced Riemannian Metrics for Motion Planning with Constraints

面向约束运动规划的诱导黎曼度量
Kyaw, Phone Thiha, Cohn, Thomas, Garcia, Miguel Angel Rogel, Kelly, Jonathan
Abstract
In constrained motion planning problems, task and loop-closure constraints restrict a robot's motion to a curved, lower-dimensional submanifold of its configuration space. Planners measure path length with a metric, which sets the cost of moving in each direction. Under the Euclidean metric, this cost is the same everywhere, whereas under a general Riemannian metric, such as the kinetic-energy metric, the cost can vary with direction and configuration. Existing methods often describe the submanifold either implicitly, as a constraint level set, or explicitly, through a parameterization. The implicit representation is typically combined with the Euclidean metric of the configuration space, and the explicit representation with the parameter domain, so the path length that a planner minimizes depends on the representation. Instead, we measure path length with the induced metric, which the submanifold inherits from a Riemannian metric on the configuration space. The implicit and explicit representations yield the same induced metric, expressed in different coordinates, and hence the same geometry. This result holds for any Riemannian metric on the configuration space, not only the Euclidean one. The choice of metric is therefore independent of the choice of representation. Using this result, we extend planning under a Riemannian metric from unconstrained spaces to constraint submanifolds by applying the induced metric in both a sampling-based planner and a trajectory optimizer. For an explicit representation, the induced metric also accounts for the distortion that the parameterization introduces. In experiments on a bimanual manipulation setup with two Franka arms under end-effector task constraints, we compare the Euclidean and kinetic-energy metrics.
Chinese Translation
在约束运动规划问题中,任务约束与闭环约束将机器人的运动限制在其构型空间的一个弯曲的低维子流形上。规划器使用度量来衡量路径长度,该度量决定了沿各个方向移动的代价。在欧氏度量下,这一代价处处相同;而在一般黎曼度量(如动能度量)下,代价可随方向和构型而变化。现有方法通常以隐式方式(作为约束的等值面)或显式方式(通过参数化)描述该子流形。隐式表示通常与构型空间的欧氏度量结合使用,而显式表示则与参数域结合使用,因此规划器所最小化的路径长度依赖于表示方式的选择。与此不同,我们采用诱导度量来衡量路径长度,即子流形从构型空间上的黎曼度量所继承的度量。隐式表示与显式表示给出相同的诱导度量(只是以不同坐标表达),因而具有相同的几何结构。这一结论对构型空间上的任意黎曼度量均成立,而不仅限于欧氏度量。因此,度量的选择与表示方式的选择相互独立。基于这一结果,我们通过在基于采样的规划器和轨迹优化器中应用诱导度量,将黎曼度量下的规划从未约束空间扩展到约束子流形。对于显式表示,诱导度量还能够刻画参数化所引入的畸变。在一个配备两个 Franka 机械臂、受末端执行器任务约束的双臂操作实验平台上,我们对比了欧氏度量与动能度量的效果。
cs.RO / 45 / 2609.25696

The Cartesian Hand: In-Hand Manipulation with All-Linear Fingers

Xia, Boxi, Li, Bokuan, Shin, Ryan, Yang, Zijiang, Liu, Jiaxun, Chen, Boyuan
Abstract
Robotic manipulation has increasingly pursued human-like dexterous hands with many articulated degrees of freedom, offering rich manipulation capabilities at the cost of mechanical and control complexity. At the other extreme, parallel grippers are simple and robust, but provide little ability to manipulate an object after grasping it. Operating articulated objects such as threaded containers, manufacturing tools, and laboratory instruments often requires a second gripper, an external fixture, or coordinated arm motion. We introduce the Cartesian Hand, a 7-DoF end-effector that rethinks dexterous manipulation by combining independent grasping and relative manipulation within a single end-effector using only linear motion. Two independently actuated parallel grippers hold different parts of an object, while four translating fingertips generate relative motion between the grasped parts. Its configuration-independent fingertip kinematics allow manipulation to be composed from simple linear motion primitives. The Cartesian Hand is particularly suited to objects structured around common mechanisms such as threads, pivots, linear guides, plungers, and triggers. We demonstrate cap opening and closing, pipetting, pumping, two-handle manipulation, screwdriving, trigger actuation, and in-grasp reorientation across 35 objects spanning laboratory, manufacturing, and household settings. The same manipulation procedures transfer from a fixed-base robot arm to a humanoid, where we demonstrate bimanual laboratory manipulation using two Cartesian Hands. These results show that versatile in-hand manipulation capability can emerge from a mechanically simple architecture when independent grasping and relative motion are designed directly into the end-effector. We will open-source all software and hardware design. Our website is https://generalroboticslab.com/cartesian_handv1.
cs.RO / 46 / 2609.25709

Zephyron: Integrated Design and Analytical Evaluation of a Solar-Assisted Mobile Manipulator for Multimodal Environmental Reconnaissance and Distributed Visual Inference

Zephyron:面向多模态环境侦察与分布式视觉推理的太阳能辅助移动操作机器人的集成设计与分析评估
Sultan, Sabik Bin, Sultan, Shafi Bin, Sadad, Safwan
Abstract
Environmental reconnaissance needs mobile platforms that carry sensors, preserve measurement context, and return interpretable evidence under limited energy and communication. We present a literature-informed engineering design for Zephyron, a four-wheel rover with a front manipulator, environmental sensors, distributed computer vision, local recording, and a raised rear solar module. The design keeps the prototype layout but replaces unsupported numerical assumptions with an explicit component and geometry baseline. A reproducible search retrieved 5,000 records (4,858 unique) for screening, followed by targeted review of primary literature and manufacturer documentation. The baseline uses 165 mm wheels, a 12 kg mass budget, a 72 Wh battery-energy basis, and a 20 W photovoltaic module. With rolling-resistance coefficient 0.04, steady ascent of a 10 degree grade needs about 0.517 N m per wheel under equal load sharing. An illustrative 40 W motion load gives 1.44 h from 57.6 Wh usable energy, and a 25 percent driving duty gives 4.19 h without solar input; these are calculated scenarios, not measured performance. Sensor models show how integration time, calibration, temperature, and communication delay constrain interpretation, and a quality-aware stop-and-sample policy links these constraints to mission execution. Lightweight detectors, reference-based sensor learning, and executable data-integrity checks define a reproducible machine-learning evaluation pathway. The contribution is a traceable design and evaluation framework with editable 3D models, subsystem diagrams, and reproducible analytical data. Experimental validation is required before assigning payload, endurance, detection, or field-operating ratings.
Chinese Translation
环境侦察需要能够搭载传感器、保留测量上下文,并在有限能源与通信条件下返回可解释证据的移动平台。本文提出了一种基于文献调研的工程设计方案——Zephyron,这是一款四轮探测车,配备前端机械臂、环境传感器、分布式计算机视觉、本地记录功能以及升高的后置太阳能组件。该设计保留了原型布局,但用明确的组件与几何基准替代了缺乏依据的数值假设。通过可复现的检索获得5,000条记录(去重后4,858条)用于筛选,随后对原始文献和制造商文档进行了针对性审阅。该基准设计采用165 mm车轮、12 kg质量预算、72 Wh电池能量基础和20 W光伏组件。在滚动阻力系数为0.04的条件下,在等载荷分配情况下,以稳定状态爬升10度坡度时每个车轮约需0.517 N·m扭矩。以示例性的40 W运动负载计算,57.6 Wh可用能量可支持1.44 h运行;若行驶占空比为25%,无太阳能输入时可运行4.19 h;以上均为计算场景,而非实测性能。传感器模型分析表明积分时间、标定、温度和通信延迟如何制约数据解读,并由此提出一种质量感知的停走采样策略,将这些约束与任务执行相联系。轻量化检测器、基于参考数据的传感器学习以及可执行的数据完整性校验共同构成了一条可复现的机器学习评估路径。本文的贡献是一个可追溯的设计与评估框架,包含可编辑的3D模型、子系统图和可复现的分析数据。在赋予载荷、续航、检测或野外运行额定指标之前,仍需进行实验验证。
cs.RO / 47 / 2609.25724

Designing an Efficient Excavator Bucket for Lunar ISRU: A Comparative Study with Vision-Based Fill and Displacement Analysis

面向月球原位资源利用(ISRU)的高效挖掘铲斗设计:基于视觉的填充与位移分析对比研究
Kafi, Abdulla Hil, Koshi, Tomoki, Casir, Jorge, Nagaoka, Kenji
Abstract
This paper present a spiral-cavity wheel for lunar regolith excavation and a sensor-light evaluation stack that jointly estimates fill ratio (vision), sinkage (vision), and specific energy from actuator logs. In benchtop tests (four revolutions at 5, 10, and 15~RPM) against two literature baselines, the proposed wheel achieved higher excavated mass and fill ratio, delivering 2.2-3.0 times higher excavation rate while reducing specific energy by 29 % relative to a bucket-drum baseline. Normalized sinkage (mm/kg) was also lower, indicating stable traction without bogging. Effort-time traces show a steady torque envelope with repeatable cut-carry-dump cycles across speeds. We provide a retention index $\eta$ that correlates with fill ratio and a DEM setup that reproduces experimental trends with low error. Results suggest spiral-cavity wheels can replace heavier multi-actuator diggers when mass, simplicity, and energy efficiency are mission drivers.
Chinese Translation
本文提出了一种用于月球风化层挖掘的螺旋腔式挖斗,以及一个传感-光照评估系统,可联合估计填充率(视觉)、下陷深度(视觉)以及基于执行器日志的比能耗。在台架测试中(在5、10和15 RPM转速下各旋转四圈),与文献中的两种基线设计相比,所提出的挖斗获得了更高的挖掘质量和填充率,挖掘速率提高了2.2至3.0倍,同时相对于桶式滚筒基线,比能耗降低了29%。归一化下陷深度(mm/kg)也更低,表明其具有稳定的牵引力而不易陷车。力-时间曲线显示,在不同转速下均具有稳定的扭矩包络和可重复的挖-运-卸循环。我们提供了与填充率相关的保持指数 η,以及一个能够以较低误差复现实验趋势的离散元(DEM)仿真设置。结果表明,当质量、简单性和能效成为任务的关键驱动因素时,螺旋腔式挖斗可以替代更重的多执行器挖掘装置。
cs.RO / 48 / 2609.25725

AgriGen: Large-Scale Scene Generation Framework for Photorealistic Agricultural Robotics Simulation

Bajpai, Utkarsh, Tleiji, Serge, Pradalier, Cédric, Aravecchia, Stéphanie
Abstract
Agricultural robotics is advancing rapidly, yet progress remains constrained by limited field access, lack of control over field conditions, geographic variability, and seasonal crop cycles. These factors make it difficult and costly to acquire diverse agricultural datasets, resulting in limited evaluation and reduced system robustness. While other robotics domains have scaled learning and evaluation through high-fidelity simulation, agricultural robotics still lacks comparably capable tools. In this paper, we present a ROS-integrated framework, built on Isaac Sim, for large-scale procedural generation of agricultural environments. The framework supports photorealistic rendering, physics simulation, and domain randomization at scales relevant to robotics research, with built-in support for row crops, orchards, and vineyards and straightforward extensibility to additional crop categories. Project Page: https://baj31415.github.io/agrigen/
cs.RO / 49 / 2609.25746

Dual Covariance Gaussian Splatting SLAM: Decoupling Rendering and Registration for Robust Real-Time Tracking

双协方差高斯泼溅SLAM:解耦渲染与配准以实现鲁棒的实时跟踪
Tan, Edward Beng Wai, Lam, Siew-Kei
Abstract
ICP-based 3D Gaussian Splatting (3DGS) SLAM tracks in real time by registering incoming frames against map Gaussians, using each primitive's covariance for both rendering and registration. These two uses place conflicting demands on one covariance. The mapper shapes it to minimize photometric error, often flattening it against surfaces, while robust registration typically benefits from measurement uncertainty. We propose a dual-covariance parameterization. Each Gaussian keeps a single mean but holds two covariances: a rendering covariance optimized by the mapper, and a tracking covariance derived from an RGB-D sensor noise model. We further use the tracking covariances as Gaussian anchors for image corners, providing constraints in directions where depth geometry is weak. We evaluate on TUM RGB-D, ScanNet, Replica, and two outdoor sequences recorded with a RealSense D435i on wheeled and handheld platforms. We achieve robust tracking performance across multiple scenes and reduced odometry drift, while tracking at $\sim$ 60 FPS.
Chinese Translation
基于ICP的3D高斯泼溅(3DGS)SLAM通过将新输入的帧与地图中的高斯基元进行配准来实现实时跟踪,其中每个基元的协方差同时用于渲染和配准。然而,这两种用途对同一个协方差提出了相互冲突的要求:建图模块通过优化协方差来最小化光度误差,通常使其沿表面被压平,而鲁棒的配准通常需要依赖测量不确定性。我们提出了一种双协方差参数化方法:每个高斯基元保留单一均值,但持有两个协方差——一个是由建图模块优化的渲染协方差,另一个是由RGB-D传感器噪声模型导出的跟踪协方差。我们进一步将跟踪协方差作为图像角点的高斯锚点,在深度几何信息较弱的方向上提供约束。我们在TUM RGB-D、ScanNet、Replica以及两段使用RealSense D435i在轮式和手持平台上采集的室外序列上进行了评估。结果表明,我们的方法在多个场景中实现了鲁棒的跟踪性能并降低了里程计漂移,同时以约60 FPS的速度进行跟踪。
cs.RO / 50 / 2609.25750

Fisheye-VLA: Decoupling Coverage and Acuity for Manipulation with a Single Fisheye Camera

Ren, Ziang, Yan, Zike, Zhang, Raymond, He, Xuguo, Li, Zhongyu
Abstract
Manipulation requires both broad scene awareness and detailed local feedback, yet conventional camera rigs provide them through separate front and wrist cameras. We present Fisheye-VLA, a visual interface that brings these capabilities together using a single passive fisheye. A global view preserves the workspace, while local perspective crops direct detail toward the interaction. The key design question is where this local visual budget should go. We answer it through a controlled re-rendering study, comparing alternative crop directions on the same recorded observations. The study finds that end-effector-centered views capture most of the estimated benefit of a much larger candidate pool, motivating a compact allocation around both hands. Our interface uses calibrated end-effector projection and motion lead to track the crops, while a shared ray encoding preserves their spatial meaning as they move. Integrated with a pretrained VLA, it achieves 84% and 82% success in the two expanded tabletop regions, where some target placements extend beyond the front-camera coverage, and supports shelf and conveyor manipulation. Ablations show that local crops and their viewing directions become more important in the larger workspace regions. The results demonstrate that a single fisheye can support these manipulation tasks without physical wrist cameras.
cs.RO / 51 / 2609.25754

PLAT: Sparse Timed Keyframe Motion Tracking for Humanoid Control via Privileged Latent Transition Learning

Wang, Zepeng, Wang, Jiangxing, Ma, Chao, Shi, Xiaochuan, Lu, Zongqing
Abstract
Humanoid motion tracking policies rely on dense frame-by-frame references, limiting their use as high-level motion controllers for planning and interactive motion generation. We study \emph{Sparse Timed Keyframe Motion Tracking}, where a policy receives only sparse future keyframes and their desired arrival times, and must execute stable whole-body motions that reach successive goals. We propose \textbf{PLAT}, a three-stage sparse timed keyframe motion tracking policy learning framework with \textbf{P}rivileged \textbf{LA}tent \textbf{T}ransition learning. PLAT bridges dense motion tracking and sparse goal-conditioned control by exploiting dense goal sequences as privileged supervision during training while requiring only sparse timed keyframe commands at deployment. A pretrained dense tracking expert first provides robust motion priors. A privileged latent prior is then learned through DAgger-style imitation, followed by latent residual reinforcement learning that refines latent transitions instead of directly optimizing actions. Extensive simulation experiments demonstrate that PLAT maintains accurate and stable sparse timed keyframe tracking across varying planning horizons, with particularly strong performance under long-horizon commands. Successful deployment on a Unitree G1 humanoid robot further demonstrates the effectiveness and practicality of PLAT for sparse humanoid motion control.
cs.RO / 52 / 2609.25756

MedVLA: A Hierarchical Vision-Language-Action Framework for Closed-Loop Precision Medical Robot Manipulation

Xie, Junjie, He, Chuxuan, Ye, Angen, Song, Yujia, Zhang, Dapeng
Abstract
Precision medical robotics demands adaptive decision-making under strict safety, interpretability, and execution constraints. Although recent Vision-Language-Action (VLA) models show strong multimodal reasoning ability, their continuous action generation paradigm is not well suited for precision medical tasks, where reliable closed-loop operation may also depend on non-action system function calls. To address this gap, we propose MedVLA, a hierarchical framework that couples high-level multimodal reasoning with low-level function-constrained execution. We further introduce a scalable multi-agent pipeline to generate skill-oriented chain-of-thought(CoT) data for structured training. Built on different multimodal large-model backbones, MedVLA consistently improves performance after fine-tuning, demonstrating the effectiveness of the proposed framework across model variants. Under identical initial conditions, we perform 100 closed-loop flexible electrode implantation trials. The results show that MedVLA achieves a 95.0\% task success rate, substantially outperforming representative VLA baselines, including OpenVLA (8\%) and $\pi_0$ (15\%), in accuracy, stability, and safety. These results indicate that structured reasoning with constrained function-level execution is a practical route toward deployable precision medical robotics.
cs.RO / 53 / 2609.25785

VisForce: Visual Grounding of Current and Desired Forces for Goal-Conditioned Dexterous Manipulation

VisForce:面向目标条件灵巧操作的当前与期望力视觉定位
Lee, Jung-Woo, Lim, Soo-Chul
Abstract
Vision-Language-Action (VLA) models have emerged as general-purpose robotic manipulation policies. However, in dexterous hand manipulation, contact forces are typically provided as separate states or force-specific representations, making it difficult to explicitly represent the spatial correspondence between force and their corresponding visual locations. In this work, we propose VisForce, which visually grounds the current and desired forces at their corresponding fingertip locations. VisForce renders current and desired visual force cues on the current wrist image and a task-specific goal image, and combines the two representations through goal-conditioned cross-attention to generate force-aware actions. We evaluate VisForce using a real UR10 robot equipped with an RH56F1 dexterous hand through force-conditioned grasping and three multi-stage manipulation tasks. In force-conditioned grasping experiments, VisForce exhibited a consistent grip-force response as the desired force increased, and achieved grasp-and-lift success rates of 70% and 80% for an egg and a toothpaste tube, respectively. It further achieved final success rates of 70%, 55%, and 40% on cup insertion/bottle pouring, tong-assisted bread transfer, and slip-modulated peg-in-hole, respectively. These results show that fingertip-aligned visual force representations can be effectively used for force-aware conditioning in VLA-based dexterous hand manipulation.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)模型已逐渐成为通用机器人操作策略。然而,在灵巧手操作中,接触力通常以独立的状态或力专用表征的形式提供,这使得难以显式地表达力与其对应视觉位置之间的空间对应关系。在本工作中,我们提出VisForce,将当前力与期望力视觉定位到其对应的指尖位置。VisForce在当前腕部图像和任务特定的目标图像上渲染当前与期望的视觉力线索,并通过目标条件的交叉注意力(cross-attention)将两种表征相结合,以生成力感知的动作。我们在配备RH56F1灵巧手的真实UR10机器人上,通过力条件抓取和三个多阶段操作任务对VisForce进行评估。在力条件抓取实验中,随着期望力的增大,VisForce表现出一致的握力响应,对鸡蛋和牙膏管的抓取-提起成功率分别达到70%和80%。在杯插入/瓶倒水、借助夹子转移面包以及基于滑移调节的轴孔装配任务中,其最终成功率分别达到70%、55%和40%。这些结果表明,与指尖对齐的视觉力表征可以有效地用于基于VLA的灵巧手操作中的力感知条件化。
cs.RO / 54 / 2609.25813

MOLA LiDAR-Inertial Odometry (MOLA-LIO) on the COMFORT Localization Benchmark

COMFORT定位基准测试中的MOLA激光雷达-惯性里程计(MOLA-LIO)
Blanco-Claraco, Jose Luis
Abstract
This short report documents our entry to the COMFORT Localization Benchmark (IROS 2026), evaluated on the GrandTour dataset recorded with the Boxi payload. It extends MOLA-LO into a LiDAR-inertial system that also ingests IMU and, optionally, legged kinematic odometry. We describe the architecture, the streams consumed, the local protocol that selected the submitted configuration, and the measurements backing our real-time claim.
Chinese Translation
本简短报告记录了我们参加COMFORT定位基准测试(IROS 2026)的方案,该测试基于使用Boxi载荷采集的GrandTour数据集进行评估。我们将MOLA-LO扩展为一个激光雷达-惯性系统,该系统还接收IMU数据,并可选择性地接收腿式机器人的运动学里程计数据。我们描述了系统架构、所使用的数据流、用于选择提交配置的本地协议,以及支持我们实时性声明的相关测量结果。
cs.RO / 55 / 2609.25820

Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models

Yang, Yuxin, He, Gaohan, Guan, Changxue, Liu, Hangming
Abstract
Discrete action tokenization is central to autoregressive vision-language-action (VLA) models, yet action representations are often evaluated primarily through reconstruction fidelity. We ask which representation properties actually matter for closed-loop control by comparing fixed analytical, data-driven linear, and nonlinear neural representations under a unified tokenization interface. Across rate-distortion analysis, sequence-modeling diagnostics, and 3,500 LIBERO rollouts, representation rankings change with the evaluation criterion. PCA achieves lower nominal reconstruction error than Temporal-DCT, but produces less predictable token sequences and 3.0 percentage points lower mean seen-task success across three policy-training seeds, with the policy ordering reversing in one seed. In a matched seed-42 ablation, an autoencoder further reduces reconstruction error yet does not yield the strongest policy and exhibits greater sensitivity to discrete token perturbations. These findings show that reconstruction fidelity alone cannot reliably select action representations for autoregressive control, motivating joint evaluation of geometric fidelity, sequence predictability, decoder stability, and closed-loop performance.
cs.RO / 56 / 2609.25831

Sometimes You Gotta Run Before You Can Walk: Run-then-Walk Scheduling Strategy for VLM Autonomous Driving

有时你必须先跑起来才能走:面向视觉语言模型(VLM)自动驾驶的“先跑后走”调度策略
Ye, Yuqi, Sun, Shangkun, Lin, Junhong, Zhao, Jiayi, Peng, Changhao, Zheng, Wei, Liu, Guoqing, Zhao, Tiesong, Gao, Wei
Abstract
Recent VLM-based autonomous driving planners adopt GRPO-style reinforcement learning to optimize driving performance. However, existing GRPO recipes either optimize driving efficiency, risking progress-seeking but unsafe behavior, or enforce early safety constraints, leading to overly conservative behavior; both require lengthy training. To solve these problems, we first reveal two distinct RL regimes: a progress regime (Run-GRPO) that aggressively explores high progress, and a safety regime (Walk-GRPO) that restores safety under stable progress. Based on this finding, we propose $\textit{Run-then-Walk}$, a simple yet effective two-stage reward scheduling strategy for GRPO, achieving both better performance and faster convergence. Unlike one-stage RL, which may focus on progress, safety, or a mixture of both within a single training phase, this schedule explicitly separates progress discovery from safety repair. In the $\textit{Run}$ phase, we focus on progress, allowing the policy to escape the conservative bias and discover high-progress modes. In the subsequent $\textit{Walk}$ phase, we introduce endpoint and safety strategy to repair unsafe behaviors from the Run phase. This reversed schedule overcomes the conservatism of Walk-first methods and the unsafe progress-seeking of joint optimization. We validate it with various VLM-based planners on multiple benchmarks: NAVSIMv1, NAVSIMv2, Navhard, and nuScenes. Extensive experiments demonstrate improved driving performance while requiring 40--50\% fewer RL training epochs than the baselines.
Chinese Translation
近期基于视觉语言模型(VLM)的自动驾驶规划器采用GRPO式强化学习来优化驾驶性能。然而,现有的GRPO方案要么优化驾驶效率,存在追求进展但不安全行为的风险;要么在训练早期就施加安全约束,导致行为过于保守;且两者都需要冗长的训练过程。为解决这些问题,我们首先揭示了两种不同的强化学习状态:积极探索高进展的进展状态(Run-GRPO),以及在稳定进展下恢复安全性的安全状态(Walk-GRPO)。基于这一发现,我们提出了“先跑后走”(Run-then-Walk)——一种简单而有效的GRPO两阶段奖励调度策略,能够同时实现更好的性能和更快的收敛速度。与单阶段强化学习可能在单一训练阶段内关注进展、安全或两者混合不同,该调度策略明确地将进展发现与安全修复分离开来。在“跑”(Run)阶段,我们专注于进展,使策略能够摆脱保守偏向并发现高进展模式。在随后的“走”(Walk)阶段,我们引入终点约束和安全策略来修复Run阶段产生的不安全行为。这种反向的调度策略克服了“先走”方法的保守性,也避免了联合优化中不安全的进展追求行为。我们在多个基准(NAVSIMv1、NAVSIMv2、Navhard和nuScenes)上使用多种基于VLM的规划器对该方法进行了验证。大量实验表明,该方法在提升驾驶性能的同时,所需的强化学习训练轮次比基线方法减少40%至50%。
cs.RO / 57 / 2609.25887

What is the Better Curriculum: Controller-Shaped Grasping Behavior for Contact Force-Sensitive Manipulation

什么是更好的课程:面向接触力敏感操作的控制器塑形抓取行为
Feng, Ziyan, Yuan, Zizhao, Fu, Yulong, He, Yuxin, Zhang, Zhiyuan, Zhang, Zhengjie, Zhou, Jinni, Xu, Renjing, Nie, Qiang
Abstract
How should a robot learn to manipulate objects so fragile that sub-Newton contact forces can cause irreversible damage? Existing visuo-tactile policy learning typically treats tactile sensing as an additional policy input. In direct-contact force-sensitive manipulation, however, the bottleneck can arise earlier, during data collection: manual gripper control is too delayed and coarse-grained to reliably maintain the narrow force range required for stable grasping. We therefore use a deterministic 25 Hz tactile reflex controller as a collection-time teacher, producing demonstrations with controller-shaped grasping behavior for tactile-free policy learning. On Action Chunking with Transformers (ACT), policies trained from reflex-shaped demonstrations recover the teacher's grasping profile and achieve 95% stable grasps on the nominal plastic-cup task, substantially outperforming visually screened manual demonstrations. The same intervention improves in-distribution stability on $\pi_{0.5}$ and shows a favorable exploratory trend on an unseen paper-cup variant. Under randomized external disturbance, however, the reflex-data $\pi_{0.5}$ policy still fails in 45% of policy-only trials, whereas a deployment-time reflex arbiter retains all grasps. These results reveal a new role for tactile feedback in force-sensitive manipulation: rather than integrating tactile into the policy, we use it as a collection-time teacher that shapes grasping behavior in demonstrations for policy learning, while disturbance rejection remains controller-dependent, revealing the boundary of tactile-free policy.
Chinese Translation
机器人应如何学习操作那些亚牛顿接触力即可造成不可逆损伤的易碎物体?现有的视觉-触觉策略学习通常将触觉感知作为策略的额外输入。然而,在直接接触的力敏感操作中,瓶颈可能更早出现于数据收集阶段:人工夹爪控制延迟过大且粒度过粗,难以可靠地维持稳定抓取所需的狭窄力范围。因此,我们使用一个确定性的25 Hz触觉反射控制器作为数据收集阶段的教师,生成具有控制器塑形抓取行为的示范数据,用于无触觉策略学习。在基于Transformer的动作分块方法上,由反射塑形示范训练的策略能够复现教师的抓取特征,在标称塑料杯任务上实现95%的稳定抓取,显著优于经视觉筛选的人工示范。同样的干预手段提升了π₀.₅在分布内任务上的稳定性,并在未见过的纸杯变体上展现出良好的探索性趋势。然而,在随机外部扰动下,基于反射数据的π₀.₅纯策略方案仍有45%的试验失败,而部署时的反射仲裁机制则能保留全部抓取。这些结果揭示了触觉反馈在力敏感操作中的新角色:我们并非将触觉集成到策略中,而是将其作为数据收集阶段的教师,在示范数据中塑形抓取行为以供策略学习,同时抗扰能力仍依赖于控制器,这揭示了无触觉策略的边界。
cs.RO / 58 / 2609.25898

Robust Active-Perception Control for Global-State-Free Aerial-Ground Cooperation

面向无全局状态空地协同的鲁棒主动感知控制
Zhang, Mingxuan, Yu, Jiajun, Zhang, Baozhe, Zhou, Pengxiang, Liu, Wentao, Gao, Fei, Xu, Chao, Cao, Yanjun
Abstract
Aerial-ground cooperation requires real-time UAV--UGV relative-state information. Instead of maintaining global estimates for both robots, direct control in a UGV-attached non-inertial frame avoids reliance on global localization. Vision-based relative pose estimation with a passive marker offers a low-cost and effective solution. However, a fixed camera may lose sight of the moving UGV when the required UAV attitude conflicts with the field-of-view (FOV) constraint. To address this, we propose COPA, a robust active-perception framework for global-state-free aerial-ground cooperation. We use a single-axis gimbal to decouple the camera optical axis from the UAV pitch attitude. We derive an active-perception model that relates UAV motion, gimbal angle, and UGV motion to the target image-plane state.A Temporal Convolutional Network (TCN) predicts short-horizon UGV acceleration and angular velocity from recent motion history without global-state measurements. The model predictive control (MPC) uses these predictions to jointly optimize UAV and gimbal control. Simulations show that COPA maintains continuous target visibility, while ablation studies confirm that the TCN reduces peak errors during UGV motion transitions. Real-world experiments with UGV accelerations up to 3m/s^2 and yaw rates up to 1.0rad/s demonstrate robust tracking.
Chinese Translation
空地协同需要实时的无人机(UAV)与无人地面车辆(UGV)之间的相对状态信息。与其同时维护两台机器人的全局估计,不如直接在附着于UGV的非惯性坐标系中进行控制,从而避免对全局定位的依赖。基于视觉、使用被动标记的相对位姿估计提供了一种低成本且有效的解决方案。然而,当所需的无人机姿态与视场角(FOV)约束发生冲突时,固定相机可能会丢失对运动中的UGV的观测。为解决这一问题,我们提出了COPA——一个面向无全局状态空地协同的鲁棒主动感知框架。我们使用单轴云台将相机光轴与无人机的俯仰姿态解耦,并推导出一个将无人机运动、云台角度和UGV运动与目标在图像平面上的状态相关联的主动感知模型。时序卷积网络(TCN)根据近期的运动历史预测UGV短时程的加速度和角速度,而无需全局状态测量。模型预测控制(MPC)利用这些预测来联合优化无人机与云台的控制。仿真结果表明,COPA能够保持对目标的持续可见性;消融实验证实TCN能够降低UGV运动转换期间的峰值误差。在UGV加速度高达3 m/s²、偏航角速度高达1.0 rad/s的真实实验中,系统展现了鲁棒的跟踪性能。
cs.RO / 59 / 2609.25900

You Should Be Properly Scoring Your Odometry

你应该正确地为你的里程计评分
Rønning, Ola, Saqib, Usama, Wąsowski, Andrzej
Abstract
When we evaluate the performance of our odometry, it is common practice to score the estimated track against a ground truth. Unfortunately, scoring uses point metrics, such as the root mean square error, that ignore the covariance matrix which estimators like filters and smoothers already report. Using the covariance matters for two reasons. First, the covariance encodes the estimator's uncertainty, so it tells us whether the estimator trusts its own output. An overconfident estimator will not report itself lost. Second, the covariance weights the error in each direction of the estimate. Without the covariance, an estimator is unduly penalized for a high error in an uncertain direction. Instead of point metrics, we should use strictly proper scoring rules. These rules score the estimate together with its reported uncertainty. Strictly proper scoring rules recover the point metrics when no covariance is reported, and they diagnose covariance inconsistency when covariance is reported. Using a one-sided pairwise test, we show that two estimators can expose overconfidence in at least one of them without a ground truth. Strictly proper scoring rules and our pairwise test are available in our open-source framework smfeval. As a case study, we use smfeval to assess the uncertainty quality of the translational component of ground-based LiDAR-inertial odometry. Across four filters we find overconfidence - the worst case reports centimeter certainty with kilometer error. Knowing the filters are overconfident, we investigate the mechanism. The investigation traces overconfidence to filters crediting LiDAR measurements with more new information than they carry.
Chinese Translation
在评估里程计(odometry)性能时,通常的做法是将估计轨迹与真值(ground truth)进行对比评分。遗憾的是,现有评分方法采用点度量(point metrics),例如均方根误差(root mean square error),忽略了滤波器和平滑器等估计器已经报告的协方差矩阵。使用协方差之所以重要,有两个原因。首先,协方差编码了估计器的不确定性,因此它能告诉我们估计器是否信任自己的输出。一个过度自信的估计器不会报告自己已经迷失。其次,协方差对估计在各个方向上的误差进行了加权。若不考虑协方差,估计器会因在不确定方向上的较大误差而受到不当的惩罚。我们应该使用严格恰当评分规则(strictly proper scoring rules)来代替点度量。这些规则将估计结果与其报告的不确定性一起进行评分。当未报告协方差时,严格恰当评分规则可退化为点度量;而当报告协方差时,它们能够诊断协方差的不一致性。通过单侧成对检验,我们证明两个估计器可以在没有真值的情况下暴露出其中至少一个的过度自信。严格恰当评分规则和我们的成对检验已在我们的开源框架 smfeval 中发布。作为案例研究,我们使用 smfeval 评估了地基激光雷达-惯性里程计(LiDAR-inertial odometry)平移分量的不确定性质量。在四种滤波器中,我们均发现了过度自信——最坏情况下,滤波器报告了厘米级的确定性,但实际误差达公里级。在了解这些滤波器存在过度自信后,我们进一步研究了其机制。研究将过度自信追溯到滤波器赋予了激光雷达测量超出其实际携带信息量的新信息。
cs.RO / 60 / 2609.25905

Control Barrier Functions for Safe Free-Flying Robotic Spacecraft Operations in Tumbling Target Capture

用于翻滚目标捕获中自由飞行机器人航天器安全操作的控制障碍函数
Meinert, Alexander, Stadler, Peter, Baldauf, Niklas, Turnwald, Alen
Abstract
This paper presents a modular control barrier function (CBF) framework for safe free-flying robotic spacecraft operations during tumbling target capture. Motivated by latest ESA guidelines for safe close proximity operations, safety zones and requirements are translated into dedicated CBFs. The 13-DoF system is decomposed into translational, attitude, and robotic subsystems, each equipped with a safety filter that minimally modifies nominal control inputs in a lightweight quadratic program. The filters enforce a conical approach corridor, collision avoidance zone, attitude line-of-sight pointing, angular velocity limits, robotic joint limits, link-base collision avoidance, and actuator constraints. Dynamic coupling between subsystems is handled by treating upstream safe control commands as known interconnection inputs in the downstream safety filters, preserving modularity while supporting system-level safety. The framework is validated in an on-orbit servicing scenario, including final approach, angular rate synchronization, and tumbling target grasping, using the high-fidelity astrodynamics simulator Basilisk. Monte Carlo simulation results demonstrate runtime efficiency and operational safety for various tumbling rates.
Chinese Translation
本文提出了一种模块化的控制障碍函数(Control Barrier Function, CBF)框架,用于保障翻滚目标捕获过程中自由飞行机器人航天器操作的安全性。受欧洲航天局(ESA)最新安全近距离操作准则的启发,安全区域和要求被转化为专用的控制障碍函数。该13自由度系统被分解为平移、姿态和机器人三个子系统,每个子系统均配备一个安全滤波器,通过轻量级二次规划对名义控制输入进行最小程度的修正。这些安全滤波器强制执行锥形接近走廊、碰撞规避区域、姿态视线指向、角速度限制、机械臂关节限制、连杆与基座碰撞规避以及执行器约束等安全要求。子系统间的动态耦合通过将上游安全控制指令作为下游安全滤波器中已知的互联输入来处理,在保持模块化的同时支持系统级安全性。该框架在高精度天体动力学仿真器Basilisk中,针对包括最终接近、角速度同步和翻滚目标抓取在内的在轨服务场景进行了验证。蒙特卡洛仿真结果证明了该框架在不同翻滚速率下的运行时效率和操作安全性。
cs.RO / 61 / 2609.25917

Vision-based Underwater Formation Control With Input Saturations via Barrier Lyapunov Functions

De Carli, Nicola, Zenário, João, Fernandez-Ayala, Victor Nan, Dimarogonas, Dimos V.
Abstract
In this work, we propose a communication-free framework for vision-based formation control of fully actuated underwater robots subject to sensing constraints, collision-avoidance requirements, and input saturations. Recentered barrier Lyapunov functions encode sensing and collision-avoidance constraints, while command-filtered backstepping extends the design to the second-order vehicle dynamics. The resulting control objective is enforced through a quadratic program that explicitly accounts for actuator limits. Conservative sensing domains provide margins from the physical limits and are adaptively relaxed when necessary, allowing temporary violation of the conservative bounds. The proposed approach is validated through realistic Software-in-the-Loop (SITL) simulations in Gazebo.
cs.RO / 62 / 2609.25932

Unsigned Distance Maps on 2D Point Cloud Registration

基于无符号距离图的二维点云配准
Sousa, Ricardo B., Grisetti, Giorgio, Sobreira, Héber Miguel, Silva, Carlos André, Moreira, António Paulo
Abstract
2D point cloud registration arises in laser odometry and Simultaneous Localization and Mapping (SLAM) for mobile robots. Iterative Closest Point (ICP) is one of the most widely used approaches. Still, its iterative procedure recomputes correspondences via nearest-neighbor search at every iteration, whereas correspondence-free alternatives focus on scan-to-map alignment. This paper proposes a 2D point cloud registration approach based on unsigned distance maps, precomputing the Euclidean distance to the nearest reference point, along with its spatial derivatives, over a discrete grid, replacing the per-iteration search with O(1) lookups. Moreover, point-to-point and point-to-plane error formulations are derived on the SE(2) manifold and solved via Gauss-Newton optimization. On a synthetic benchmark and the real-world IILABS 3D dataset, the precomputed point-to-point variant outperforms its analytical counterparts, achieving competitive laser-odometry drift compared to point-to-plane formulations, as the precomputed gradient regularizes correspondences in the presence of sensor noise.
Chinese Translation
二维点云配准广泛应用于移动机器人的激光里程计和同步定位与建图(SLAM)中。迭代最近点算法(ICP)是目前最常用的方法之一。然而,其迭代过程在每次迭代时都需要通过最近邻搜索重新计算对应关系,而无需对应关系的替代方法则主要聚焦于扫描到地图的对齐。本文提出一种基于无符号距离图的二维点云配准方法,预先在离散网格上计算到最近参考点的欧氏距离及其空间导数,从而以 O(1) 的查询操作取代每次迭代中的搜索。此外,本文在 SE(2) 流形上推导了点到点与点到面的误差形式,并通过高斯-牛顿(Gauss-Newton)优化进行求解。在合成基准数据集和真实世界的 IILABS 3D 数据集上,预计算点到点方法优于其解析对应方法,在传感器噪声存在的情形下,由于预计算梯度对对应关系起到了正则化作用,实现了与点到面方法相比具有竞争力的激光里程计漂移表现。
cs.RO / 63 / 2609.25942

Destination Support Restoration for Finite-Set Multimodal Trajectory Prediction

面向有限集多模态轨迹预测的目标点支撑恢复
Liu, Fengrui, Peng, Jiajun, Peng, Duo, Liu, Feng
Abstract
Robots operating around pedestrians often reason over a finite set of predicted human futures. Repeated online updates can concentrate this limited prediction budget on dominant destinations and leave plausible alternatives underrepresented or absent, removing those alternatives from the finite representation available to downstream decision making. We introduce Destination Support Restoration (DSR), a causal post-selection operator that repairs destination support without retraining the host predictor or increasing the maintained set size. At a repair step, DSR evaluates a temporary destination-stratified candidate bank from the observed prefix, converts candidate evidence into integer target counts, protects representatives of active modes, and reallocates redundant surplus hypotheses to deficient modes. The maintained and returned sets retain exactly $N$ hypotheses, and DSR replaces at most $\lceil\rho N\rceil$ entries. Protected representatives preserve current categorical support; lineage-aware particle filters also preserve surviving resampling ancestors. Each replacement reduces the allocation mismatch to the evidence-driven target by one. On the complete 3,719-trajectory Edinburgh protocol over three seeds, DSR reduces MIF weighted ADE and FDE by 13.36% and 13.30% at $N=64$. Paired integrations with CLiFF, PPT, causal GDTS, Social Informer, and PECNet improve both metrics in every evaluated pair. These results show that finite-set support allocation is a useful prediction-side control point when a fixed hypothesis set serves as the interface to downstream systems.
Chinese Translation
在行人环境中运行的机器人通常基于一个有限的人体未来状态预测集合进行推理。重复的在线更新会使这一有限的预测预算集中于占主导地位的目标点,导致合理的备选目标支撑不足甚至缺失,从而使这些备选项无法进入可供下游决策使用的有限表示中。我们提出目标点支撑恢复(Destination Support Restoration, DSR),这是一种因果性的后选择算子,能够在不重新训练宿主预测器、也不增加维护集合规模的情况下修复目标点支撑。在修复步骤中,DSR 从观测前缀出发评估一个临时的、按目标点分层的候选库,将候选证据转换为整数目标计数,保护活跃模态的代表性假设,并将冗余的多余假设重新分配给支撑不足的模态。维护并返回的集合始终保持恰好 $N$ 个假设,且 DSR 至多替换 $\lceil\rho N\rceil$ 个条目。受保护的代表性假设保留当前的类别支撑;具有谱系感知的粒子滤波器还能保留幸存的重采样祖先。每次替换都使分配失配相对证据驱动的目标减少一。在包含 3,719 条轨迹的完整 Edinburgh 协议上、跨三个随机种子,当 $N=64$ 时,DSR 将 MIF 加权 ADE 和 FDE 分别降低 13.36% 和 13.30%。与 CLiFF、PPT、因果 GDTS、Social Informer 和 PECNet 的成对集成在所有评估组合中均提升了这两项指标。这些结果表明,当固定假设集合作为下游系统的接口时,有限集支撑分配是一个有效的预测侧控制点。
cs.RO / 64 / 2609.25961

An Action Is Worth One Patch: Unified World-Action Modeling with PatchWAM

一个动作等价于一个图块:基于PatchWAM的统一世界-动作建模
Wang, Tianheng, Xie, Zhou, Jia, Heng, Xu, Jianhua, Zhang, Tong, Yu, Kaicheng
Abstract
Generative visual models offer a foundation for learning representations of physical dynamics, yet their extension to continuous control raises a fundamental question: do visual prediction and action generation require separate computational pathways? Existing approaches usually introduce trainable action heads or separate action experts to bridge low-dimensional states and high-dimensional visual representations. In this work, we explore whether the visual backbone's existing capacity can also support control when actions are expressed in a compatible representation. Thus, we introduce PatchWAM (Patch World-Action Model), which treats continuous actions as another type of patch through a fixed mapping called Action-as-Patch. This allows a single model to predict both how the robot should move and what the scene may look like afterward. Visual prediction and action generation become parts of the same generative process, without a dedicated action head or separate action expert. Experiments with subsampled training windows show gains over a matched dual-expert control, while benchmark evaluations reach 91.8% success rate on LIBERO-Plus and 96.12% on RoboTwin 2.0 in a full-data setting with additional augmented demonstrations. More broadly, the result suggests that capability need not be added where it can be inherited: the constraint on extending a generative backbone is the interface a new signal is written in, not the capacity to model it.
Chinese Translation
生成式视觉模型为学习物理动力学的表征提供了基础,但将其扩展到连续控制引发了一个根本性问题:视觉预测与动作生成是否需要各自独立的计算通路?现有方法通常引入可训练的动作头或独立的动作专家,以衔接低维状态与高维视觉表示。在本工作中,我们探究当动作以兼容的表示形式表达时,视觉骨干网络已有的能力是否也能支持控制任务。为此,我们提出了PatchWAM(Patch World-Action Model,图块世界-动作模型),通过一种称为Action-as-Patch(动作即图块)的固定映射,将连续动作视为另一种类型的图块(patch)。这使得单一模型既能预测机器人应如何运动,也能预测随后场景可能呈现的样子。视觉预测与动作生成由此成为同一生成过程的组成部分,无需专门的动作头或独立的动作专家。在子采样训练窗口的实验中,我们的方法优于规模相当的双专家控制基线;在全量数据并辅以增强演示的设置下,基准评估在LIBERO-Plus上达到91.8%的成功率,在RoboTwin 2.0上达到96.12%。更广泛地说,这一结果表明:能力若可继承便无需额外添加——扩展生成式骨干网络的瓶颈在于新信号所采用的接口形式,而非建模该信号的能力。
cs.RO / 65 / 2609.25969

Predict Before You Step: Auditable Occupancy Forecasting for Dynamic Obstacle Avoidance under Sparse Guidance

Mao, Yuhui, Liu, Fen, Yuan, Shenghai, Hu, Tianxin, Liu, Ruimeng, Su, Rong
Abstract
Legged robots under sparse waypoint guidance must avoid moving obstacles using partial, rapidly changing LiDAR observations. We present LOOP (Latent-recurrent Occupancy rollOut Policy), a local avoidance policy that connects sparse waypoint guidance to a frozen locomotion controller at 50 Hz. From occupancy and ego-velocity histories, a recurrent predictor forecasts future occupancy over a 1 s horizon by warping the current map with learned flow and visibility gates. These maps guide velocity selection through map-derived features and geometric risk estimates, providing an explicit interface for inspecting and replacing predictions. In encounter-synchronised Isaac Lab evaluations, LOOP achieves 57.1% head-on success at obstacle speeds of 2.5-3.2 m/s, exceeding a retrained reactive baseline by 8.2 percentage points. Comparisons with a rollout-free BEV policy show smaller, scenario-dependent gains from the prediction branch, including improved crossing success and reduced variability across training seeds at the highest head-on speeds. The adapter runs onboard a Unitree Go2 in 14.5 ms per step and completes all 16 real-world crossing trials without collision, demonstrating deployment feasibility.
cs.RO / 66 / 2609.25994

Safety-Constrained Model Predictive Control for an Omnidirectional Walking Assistive Robot Using Control Barrier Function

Fortuna, Andrea, Lorenzini, Marta, Motta, Elisa, Ranavolo, Alberto, De Momi, Elena, Ajoudani, Arash
Abstract
Providing safe and effective mobility assistance plays a crucial role in restoring independence and enhancing the quality of life for individuals with motor impairments. In this context, robotic walking assistive devices have recently emerged as promising solutions to provide physically compliant interaction while ensuring user safety and support. This paper presents a novel control framework for an omnidirectional Walking Assistive Robot (I-WANDER) that integrates a Control Barrier Function (CBF) formulation into a Model Predictive Control (MPC) scheme to explicitly enforce collision-avoidance safety constraints while optimizing for energy efficiency and smooth human-robot collaboration. The method was experimentally evaluated with 12 healthy participants performing two different walking tasks using both the proposed CBF-based MPC controller (CB-MPC) and a variable admittance controller (AC). The first task involved structured navigation through a U-shaped corridor, whereas the second consisted of a single-obstacle avoidance task performed blindfolded to ensure the obstacle was unexpected. Comparative results show that the CB-MPC architecture significantly reduces energy consumption and mechanical work (p < 0.01) without compromising motion smoothness, while also decreasing the number of obstacle collisions. Overall, the findings highlight the potential of the proposed control architecture to enhance both safety and efficiency in robotic walking assistance.
cs.RO / 67 / 2609.26004

Manipulation with Stability Guarantees: Linear Deformable Objects with Non-negligible Physical Response Grasped at Multiple Location

具有稳定性保证的操作:在多个位置抓持的具有不可忽略物理响应的线性可变形物体
Feliu-Talegon, Daniel, Della Santina, Cosimo
Abstract
Most research on the manipulation of deformable objects focuses on lightweight systems with negligible mechanical response, effectively restricting attention to quasi-static regimes. This assumption excludes a broad class of practically relevant objects, such as hoses, pipes, and wiring harnesses, whose dynamics cannot be ignored during manipulation. In this work, we address this limitation by introducing a closed-loop control architecture that explicitly accounts for object dynamics and recasts manipulation as a shape-regulation problem. Control is achieved by modulating forces and torques applied at multiple fixed points along the object. This approach builds on three methodological contributions: a fully dynamic model of linear deformable objects based on discrete strain parameterizations; an extension of the notion of actuation coordinates to SE(3), yielding a structured and inherently underactuated control architecture; and nonlinear feedback strategies providing explicit conditions for steady-state convergence to desired configurations. Extensive simulations on representative manipulation tasks demonstrate the performance gains enabled by the proposed modelbased formulation. We finally validate the approach experimentally through a real-time closed-loop implementation with online shape estimation, confirming its practical feasibility and effectiveness
Chinese Translation
现有关于可变形物体操作的研究大多集中于机械响应可忽略的轻质系统,实际上将关注点局限于准静态情形。这一假设排除了一类广泛的、具有实际意义的应用对象,例如软管、管道和线束等,其动力学在操作过程中不可忽略。本文针对这一局限,提出了一种显式考虑物体动力学的闭环控制架构,并将操作问题重构为形状调节问题。控制通过调制施加于物体上多个固定点的力和力矩来实现。该方法建立在三项方法学贡献之上:一是基于离散应变参数化的线性可变形物体全动力学模型;二是将执行坐标的概念扩展至SE(3),从而得到一种结构化且本质上欠驱动的控制架构;三是非线性反馈策略,为稳态收敛到期望构型提供了显式条件。在代表性操作任务上的大量仿真验证了所提出的基于模型方法所带来的性能提升。最后,我们通过具备在线形状估计的实时闭环实验对该方法进行了验证,证实了其实际可行性与有效性。
cs.RO / 68 / 2609.26007

Skytopia: Monocular Drone Navigation with Action-Conditioned Latent World Models

Zhang, Yuhang, Zhang, Rangya, Shang, Yujing, Yu, Zhuoyuan, Wang, Weiying, Yang, Steven, Yan, Qingsong, Yan, Chao, Feroskhan, Mir
Abstract
Monocular drone navigation requires reaching a goal in an unseen environment from a single forward-facing camera, which offers few cues for depth and scale. World models address this by modelling how observations evolve under actions, but they are built to be executed: the prediction is produced at deployment and fed back into action generation at every control step. We argue that what a policy needs from a world model is not the prediction but the representation required to produce it: in flight the executed action explains almost all of the change between observations, so prediction reduces to reprojecting a static scene under a known displacement. We therefore introduce skytopia, a policy built on an action-conditioned latent world model, and the 3D Gaussian Splatting platform on which it is trained. A forward objective predicts the representation of the next observation from the intended motion, and an inverse objective recovers that motion from the predicted transition. Because the prediction never reaches action generation, the predictor is discarded and one policy serves point-goal, image-goal, and goal-free navigation. Simulation experiments show that skytopia outperforms every baseline under all three specifications, attaining 57.8%, 66.0%, and 49.0% success rate, while discarding the predictor removes 59.4% of the inference cost. The same policy is subsequently deployed on a physical drone without fine-tuning and reaches goals in indoor, open outdoor, and woodland environments.
cs.RO / 69 / 2609.26051

Towards Intent-Aware Human-Robot Teaming: A Platform for Search-and-Rescue Operations

Maben, Rohith Prem, Jena, Ayesha, Olofsson, Björn, Reitmann, Stefan, Malec, Jacek, Woltjer, Rogier, Topp, Elin Anna
Abstract
We investigate the challenges of enabling effective collaboration between human operators and heterogeneous autonomous agents in complex, dynamic environments by developing an interaction platform that allows study of operator behavior and supports intent inference and decision-making using state-of-the-art frameworks. We demonstrate the extent to which the operator's perception, decisions, and actions could be supported by autonomous systems during search-and-rescue operations with our platform.
cs.RO / 70 / 2609.26071

StrataVLA: Hierarchical and Efficient 3D Geometric Grounding for Vision-Language-Action Models

StrataVLA:面向视觉-语言-动作模型的分层高效3D几何接地方法
Cui, Jin, Pu, Zhaoyu, Cai, Botao, Ye, Jun, Long, Xinyue, Zhao, Boran, Ren, Pengju
Abstract
Vision-Language-Action (VLA) models inherit strong semantic priors from large-scale vision-language pretraining, yet remain limited in robotic manipulation by insufficient 3D spatial awareness. Existing approaches either require explicit depth or point-cloud inputs, compress geometry into training-time supervision, or inject it only at the model input or action expert, leaving the vision-language backbone without persistent access to task-relevant spatial information. We introduce StrataVLA, a plug-and-play framework for hierarchical geometric grounding. A frozen geometry foundation model extracts shared geometric features from RGB observations, while sparse, layer-specific Geometry Adapters allow visual representations at selected backbone depths to retrieve relevant geometric evidence through cross-attention. To make inference-time geometry practical, StrataVLA further combines task-aware routing with an LRU feature cache that exploits temporal redundancy during task manipulation. Experiments on LIBERO, SimplerEnv, and real-world manipulation demonstrate consistent gains over strong VLA baselines. StrataVLA achieves 98.53% average success on LIBERO suites while reducing geometry-model invocations by up to 88%, establishing hierarchical geometry injection as an effective and efficient way to achieve spatially grounded robotic control.
Chinese Translation
视觉-语言-动作模型从大规模视觉-语言预训练中继承了强大的语义先验,但在机器人操作任务中仍受限于3D空间感知能力不足。现有方法要么需要显式的深度或点云输入,要么将几何信息压缩为训练阶段的监督信号,或者仅在模型输入端或动作专家模块注入几何信息,导致视觉-语言骨干网络无法持续获取与任务相关的空间信息。我们提出StrataVLA,一个即插即用的分层几何接地框架。一个冻结的几何基础模型从RGB观测中提取共享的几何特征,同时稀疏的、特定层级的几何适配器使选定骨干网络深度的视觉表示能够通过交叉注意力检索相关的几何证据。为使推理阶段的几何计算切实可行,StrataVLA进一步结合了任务感知路由与利用任务操作过程中时间冗余性的LRU特征缓存。在LIBERO、SimplerEnv和真实世界操作任务上的实验表明,该方法相比强大的VLA基线取得了一致的性能提升。StrataVLA在LIBERO测试套件上达到98.53%的平均成功率,同时将几何模型的调用次数最多减少88%,确立了分层几何注入作为实现空间接地机器人控制的一种高效且有效的方法。
cs.RO / 71 / 2609.26083

Situation Aware Locomotion for Dual Mobile Cobots in Shared Environments

Moraes, William, Nunes, Igor, Mazondo, Ahilen, Barcelona, Sebastian, Sodre, Hiago, Moraes, Pablo, Grando, Ricardo B.
Abstract
This paper presents a situation aware locomotion framework for two mobile collaborative robots operating in shared industrial environments. The proposed method models situation awareness through perception, comprehension, and projection to support locomotion decisions. Robot pose, load state, manipulator state, shared zone occupancy, obstacle state, and predicted inter robot conflict were used to select safe locomotion actions. The framework was implemented in simulation and evaluated in simulated industrial scenarios designed to match a feasible 4 m by 4 m physical test area. The proposed method was compared with two other baselines over multiple trials and randomized seeds. The results show that the situation aware method achieved 100% task success across all scenarios, while the independent and fixed priority baselines each achieved 33.3% overall success. The proposed method eliminated shared zone conflicts and safety stops, maintained the largest average minimum inter robot distance, and completed the tasks with the lowest average completion time. These results indicate that situational awareness can improve the locomotion of dual robots by combining load state, manipulator state, reasoning about the shared zone, and prediction of short-horizon conflicts.
cs.RO / 72 / 2609.26084

Vision-Language Models as copilots for Autonomous UAV Navigation: Analysis of Latency and Reliability in Degraded Environments

视觉-语言模型作为自主无人机导航的副驾驶:退化环境下的延迟与可靠性分析
Sodre, Hiago, Barcelona, Sebastian, Sandin, Vincent, Moraes, Pablo, Mazondo, Ahilen, Nunes, Igor, Moraes, William, Kelbouscas, André, Grando, Ricardo
Abstract
The integration of Vision-Language Models (VLMs) in autonomous Unmanned Aerial Vehicles (UAVs) offers unprecedented semantic reasoning capabilities. However, real-time closed-loop navigation requires not only low inference latency but also obedience to structured flight commands. This paper proposes a hybrid FSM-VLM control architecture for UAVs in GPS-free environments. The system combines a deterministic Finite State Machine (FSM) for low-level physical control with an asynchronous VLM copilot for high-level semantic pathfinding. We evaluate three models with different parameter scales in a Software-In-The-Loop (SITL) simulation. The framework isolates and measures syntax errors at the format level versus semantic hallucinations at the logic level in a normal and degraded scenario. This study demonstrates that parameter scaling, and not pure latency, remains the primary bottleneck for the safe and compatible integration of VLM into autonomous flights.
Chinese Translation
将视觉-语言模型集成到自主无人机中,可提供前所未有的语义推理能力。然而,实时闭环导航不仅要求低推理延迟,还要求严格遵循结构化飞行指令。本文提出了一种适用于无GPS环境下无人机的混合FSM-VLM控制架构。该系统将用于底层物理控制的确定性有限状态机(FSM)与用于高层语义寻路的异步VLM副驾驶相结合。我们在软件在环(SITL)仿真中评估了三种不同参数规模的模型。该框架在正常与退化场景下,分别隔离并测量了格式层面的语法错误与逻辑层面的语义幻觉。研究表明,参数规模而非纯粹的延迟,仍然是实现VLM安全且兼容地集成到自主飞行中的主要瓶颈。
cs.RO / 73 / 2609.26085

Acoustic Ellipses: Bio-Inspired Omnidirectional Echolocation in Cooperative Multi-Agent Systems using Frequency Sweeps

声学椭圆:基于频率扫频的协作多智能体系统仿生全向回声定位
Swissler, Petras, Burke, Lindsay, Bruno, Julia Hyland
Abstract
Inspired by the flight and song of birds, we propose an approach that enables cooperating agents to effectively identify obstacles in their environment by having a stationary source agent emit a frequency-modulated chirp while one or more listener agents observe the direct and reflected signals while in motion. We present a novel mapping approach that exploits the ``frequency gap'' between direct and reflected chirp signals to define candidate reflection ellipses, which are then fed into a 2D accumulation filter to identify locations with the highest density of potential reflections. We demonstrate this work first with simulation results derived from an efficient, bespoke audio simulator, examining the effect of obstacle count, sampling rate, and path curvature on the ability to accurately identify environmental obstacles for a one-listener scenario. We then examine different cooperative motion strategies for two-listener configurations. Finally, we validate the real-world viability of this approach through field experiments in an outdoor park setting to demonstrate the ability to identify frequency gaps using off-the-shelf hardware. Our results provide a foundation for a low-cost approach to environmental mapping in swarm robotic systems.
Chinese Translation
受鸟类飞行与鸣叫的启发,我们提出了一种方法,使协作智能体能够有效识别其环境中的障碍物:由一个静止的源智能体发射调频啁啾(chirp)信号,同时一个或多个处于运动状态的监听智能体观测直达信号与反射信号。我们提出了一种新颖的映射方法,利用直达与反射啁啾信号之间的“频率间隙”(frequency gap)来定义候选反射椭圆,然后将其输入二维累积滤波器,以识别潜在反射密度最高的位置。我们首先使用一个高效的自研音频仿真器得到的仿真结果展示这项工作,考察了在单监听者场景下,障碍物数量、采样率和路径曲率对准确识别环境障碍物能力的影响。随后,我们针对双监听者配置研究了不同的协作运动策略。最后,我们通过在户外公园环境中开展的实地实验验证了该方法在真实世界中的可行性,证明了使用现成商用硬件即可识别频率间隙。我们的研究结果为群体机器人系统中一种低成本的环境建图方法奠定了基础。
cs.RO / 74 / 2609.26118

GDLAM: Group-Disentangled Latent Action Model for Highly Disentangled Embodied Pretraining

GDLAM:面向高度解耦具身预训练的群组解耦潜在动作模型
Yang, Jiarui, Li, Jiawei, Zhang, Jiale, Guo, Hang, Huang, Wen, Hu, Maowei, Dai, Tao, Xia, Shu-Tao
Abstract
Latent action models (LAMs) learn action-related representations from action-free videos via self-supervised future prediction, offering a scalable paradigm for embodied intelligence pretraining. However, existing LAMs collapse heterogeneous sources of visual change, including camera motion, object dynamics, and interaction events, into a single latent vector, resulting in entangled representations with limited semantic structure and consequently restricting world model controllability and VLA policy generalization. We introduce the Group-Disentangled Latent Action Model (GDLAM), a latent action model whose code is factorized by construction into N groups, each with an independent variational bottleneck and a spatially gated routing pathway, and trained with a set of information-geometric objectives: mutual exclusivity, group and gate sparsity, and static-dynamic orthogonality, that make the groups mutually causally distinct rather than merely decorrelated. Quantitatively, intervening on any single group changes only that group and leaves the others intact, and GDLAM improves label-free disentanglement metrics, including Modularity, MIG, and DCI, by wide margins over a strong unstructured LAM. Notably, this factorization is not at the expense of action information: across three mutual-information estimators and a linear probe, the grouped code is more informative than monolithic baselines both in- and out-of-distribution. As supporting evidence that the disentangled code is a reusable pretraining currency, we further transfer it to two downstream regimes: (1) World Modeling: World models pretrained with GDLAM achieve superior rollout fidelity and action-following capability compared with SOTA baselines. (2) VLA Policies: Pretraining with GDLAM substantially improves task success rates over previous methods across multiple simulation benchmarks and real-world robotic manipulation tasks
Chinese Translation
潜在动作模型(Latent Action Models, LAMs)通过自监督的未来预测从无动作标注的视频中学习与动作相关的表征,为具身智能预训练提供了一种可扩展的范式。然而,现有的LAM将相机运动、物体动态和交互事件等异质的视觉变化来源压缩为单一潜在向量,导致表征纠缠、语义结构受限,进而制约了世界模型的可控性以及VLA策略的泛化能力。我们提出了群组解耦潜在动作模型(Group-Disentangled Latent Action Model, GDLAM),该模型的编码在结构上被分解为N个群组,每个群组拥有独立的变分瓶颈和空间门控路由通路,并通过一组信息几何目标进行训练:互斥性、群组与门控稀疏性、静态-动态正交性,使各群组在因果上相互独立,而不仅仅是去相关。定量结果显示,对任一单一群组进行干预仅改变该群组而保持其他群组不变,且GDLAM在无标签解耦指标(包括Modularity、MIG和DCI)上大幅超越强基线的非结构化LAM。值得注意的是,这种分解并未以牺牲动作信息为代价:在三种互信息估计器和一个线性探针上,群组化编码在分布内和分布外的信息量均高于单体基线。作为解耦编码可作为可复用预训练资产的支持性证据,我们进一步将其迁移到两个下游场景:(1)世界建模:使用GDLAM预训练的世界模型在滚动预测保真度和动作跟随能力上优于SOTA基线;(2)VLA策略:在多个仿真基准和真实世界机器人操作任务上,使用GDLAM预训练相比以往方法显著提升了任务成功率。
cs.RO / 75 / 2609.26130

Design and Implementation of an Ultra-Low-Cost Wall-Climbing Robot for Infrastructure Crack Detection

一种用于基础设施裂缝检测的超低成本爬壁机器人的设计与实现
Modak, Mrinmoy, Pretom, Supreyo Chakravorty, Tarafder, Shourv, Drew, Daniel S.
Abstract
Crack detection is a crucial process to ensure the safety and longevity of buildings and other infrastructure. In this paper, we developed a low-cost, automated crack detection robot that leverages CNN, EfficientNet-B0, and YOLOv8 for efficient identification of cracks in concrete surfaces with a curated crack image dataset introduced to support training and evaluation. YOLOv8's real-time object detection enhances crack localization, while CNN and EfficientNet-B0 provide binary classification, ensuring high precision and recall. The system consists of two stages. In the first stage, YOLOv8 detects and localizes wall regions from the video frame, and the bounding boxes are cropped. The second stage performs crack detection using one of three models by analyzing the cropped regions. Cost-effective approaches are also taken for robot design. The robot features a fan-based negative pressure adhesion system, a 4-wheeled skid-steering drive, and an ESP32-CAM for real-time image capture. Its lightweight 3D-printed chassis ensures stability, allowing it to navigate both walls and ceilings while capturing images for crack analysis. Unlike conventional wall-climbing robot designs, this robot incorporates a funnel-shaped body that enhances negative pressure generation and achieves a 44% reduction in duty cycle, significantly lowering power consumption. By combining low-cost hardware with a deep learning pipeline, our system provides a scalable, efficient, and accessible solution for real-time infrastructure inspection at an approximate total cost of $25, with a lightweight web application enabling smartphone-based control. This affordability makes the system more suitable for the developing world, where infrastructure inspection is often limited by budget constraints, labor intensity, and safety risks.
Chinese Translation
裂缝检测是确保建筑物及其他基础设施安全性和使用寿命的关键环节。本文开发了一种低成本、自动化的裂缝检测机器人,利用CNN、EfficientNet-B0和YOLOv8高效识别混凝土表面的裂缝,并构建了一个精选的裂缝图像数据集以支持训练和评估。YOLOv8的实时目标检测能力增强了裂缝定位效果,而CNN和EfficientNet-B0则提供二分类功能,确保了高精确率和高召回率。该系统由两个阶段组成:第一阶段,YOLOv8从视频帧中检测并定位墙面区域,并裁剪出边界框;第二阶段使用三种模型之一分析裁剪后的区域以进行裂缝检测。机器人设计同样采用了低成本方案。该机器人配备了基于风扇的负压吸附系统、四轮差速转向驱动以及用于实时图像采集的ESP32-CAM。其轻量化的3D打印底盘确保了稳定性,使其能够在墙面和天花板上行进并采集图像用于裂缝分析。与传统爬壁机器人设计不同,该机器人采用了漏斗形机身,增强了负压的产生,并使占空比降低了44%,显著降低了功耗。通过将低成本硬件与深度学习流程相结合,我们的系统以约25美元的总成本,提供了一种可扩展、高效且易于获取的实时基础设施检测解决方案,并配有轻量级Web应用程序以实现基于智能手机的控制。这种经济性使该系统更适合发展中国家,因为在这些地区,基础设施检测往往受到预算限制、劳动强度和安全风险的制约。
cs.RO / 76 / 2609.26131

StepTrigger: Contact-State-Triggered Backdoor Attacks on VLM-Powered Legged Robots

Zhang, Jiageng, Obidov, Doniyorkhon, Yang, Kaichen
Abstract
Large language models and vision-language models are increasingly used as high-level planners in robotic systems, using task goals and sensor summaries to select navigation or manipulation actions. This creates a new backdoor surface: a compromised planner can behave normally in most runs, yet change its target selection when a hidden trigger is present. Prior attacks on LLM-powered or embodied agents mainly rely on triggers that appear in language, camera-visible objects, scene semantics, or specific sequences of past actions. This paper presents StepTrigger, a contact-state-triggered backdoor attack for VLM-powered legged robots. The trigger is not a prompt token or a visible marker. It is produced by pressure and foot-ground contact patterns that arise when a Unitree Go1 quadruped walks across a dense terrain patch. Unlike conventional visual or textual triggers, contact signals are inherently noisy and may also arise during benign locomotion. To avoid treating every pressure anomaly as a trigger, StepTrigger learns a selective backdoor policy from multimodal robot state, using incidental pressure events as benign examples and dense-patch contacts as poisoned examples. In a stratified offline evaluation, the trained planner achieved 98.75% clean behavior preservation, 92.50% false-trigger rejection, 76.25% true-trigger activation, and 89.17% overall parsed behavior accuracy. These results reveal a backdoor surface in proprioceptive and contact channels that is not captured by defenses focused only on language, vision, or action history.
cs.RO / 77 / 2609.26155

Toward Self-Repairing Ubiquitous Robots Using Goal-Oriented Agentic AI in Human-Robot Interactions

基于目标导向智能体AI的人机交互自修复普适机器人研究
Frederiksen, Morten Roed
Abstract
Ubiquitous robotic systems often lack traditional visual interfaces, making natural language interaction important for maintenance and repair. This paper presents a goal-oriented agentic AI architecture that enables non-expert users to complete technical repair tasks through situated dialogue. The architecture separates pre-interaction goal decomposition, persistent state tracking, strategic goal management, and real-time conversational execution. We evaluated the system in a physical hardware repair task with twenty participants. Nineteen participants completed the task, corresponding to a 95% completion rate. Participants rated the system as helpful and competent, and the agent remained robust to conversational diversions such as meta-queries and code-switching. A comparison with a prior online baseline showed that physical interaction significantly reduced perceived social presence, (p=.0005), and trust and competence, (p=.037), while perceived helpfulness remained high.
Chinese Translation
普适机器人系统通常缺乏传统的可视化界面,这使得自然语言交互对于维护和修复工作变得尤为重要。本文提出了一种目标导向的智能体AI(agentic AI)架构,使非专业用户能够通过情境化对话完成技术修复任务。该架构将交互前的目标分解、持久的状态跟踪、策略性目标管理与实时对话执行相互分离。我们在一项物理硬件修复任务中对系统进行了评估,共有二十名参与者参与。其中十九人完成了任务,完成率为95%。参与者认为该系统是有帮助且能力出色的,并且该智能体在面对元查询和语码转换等对话干扰时仍保持稳健。与此前在线基准实验的比较表明,物理交互显著降低了感知社会存在感(p=.0005)以及信任感和能力感知(p=.037),而感知有用性仍然保持较高水平。
cs.RO / 78 / 2609.26184

Silent Sabotage: Internal State Triggered Backdoor Attacks on LLM-Powered Robotic Systems

无声破坏:针对大语言模型驱动的机器人系统的内部状态触发后门攻击
Obidov, Doniyorkhon, Akki, Shivayogi, Chen, Tan, Yang, Kaichen
Abstract
The integration of Large Language Models (LLMs) into robotic control systems is enabling a new generation of autonomous agents capable of complex reasoning and planning. While this paradigm shift accelerates progress, it also introduces novel security risks that remain largely unexplored. Current research into LLM backdoors has focused on attacks triggered by external stimuli, such as specific words, visual objects, or environmental states. These attacks, while potent, overlook a more insidious class of vulnerability where the trigger is internal to the agent's own operational logic. This paper presents the first comprehensive study of history-based backdoor attacks on LLM-powered robotic systems. We demonstrate that an attacker can embed a stealthy backdoor into an LLM-based robot controller by manipulating its instructions. This backdoor is triggered not by an external cue, but by a specific, rare sequence of the robot's own past actions. It remains dormant during normal operation, preserving the robot's utility, but can be activated to induce a malicious behavior, such as a complete stop or a collision. Our experiments, conducted in a simulated environment with a variety of robots and LLMs, show that this history-based attack is highly effective, achieving a near-perfect attack success rate while remaining exceptionally difficult to detect. These findings reveal a critical and previously unaddressed vulnerability in autonomous systems and underscore the urgent need for security measures that account for an agent's internal state.
Chinese Translation
将大语言模型(LLM)集成到机器人控制系统中,正在催生新一代具备复杂推理与规划能力的自主智能体。尽管这一范式转变加速了技术进步,但它也带来了此前基本未被探索的新型安全风险。当前关于LLM后门的研究主要聚焦于由外部刺激触发的攻击,例如特定词语、视觉物体或环境状态。这类攻击虽然威力强大,却忽视了一类更为隐蔽的脆弱性——其触发条件源于智能体自身运行逻辑的内部。本文首次对基于历史的后门攻击在LLM驱动的机器人系统中的威胁进行了全面研究。我们证明,攻击者可以通过操纵指令,在基于LLM的机器人控制器中植入一种隐蔽的后门。该后门并非由外部线索触发,而是由机器人自身过去动作中特定且罕见的序列所触发。它在正常运行期间保持休眠,不影响机器人的正常功能,但一旦被激活即可诱使机器人执行恶意行为,例如完全停止或发生碰撞。我们在模拟环境中针对多种机器人和LLM进行的实验表明,这种基于历史的攻击极为有效,在几乎达到完美攻击成功率的同时,还异常难以检测。这些发现揭示了自主系统中一个关键且此前未被关注的安全漏洞,并强调了亟需建立能够考虑智能体内部状态的安全防护措施。
cs.RO / 79 / 2609.26238

Manipulation of Deformable Linear Objects Using Model Predictive Path Integral Control with Bidirectional Long Short-Term Memory Learning

基于模型预测路径积分控制与双向长短期记忆学习的可变形线性物体操作
Zeh, Lukas, Meiwaldt, Johannes, Zhou, Zexu, Lechler, Armin, Verl, Alexander
Abstract
The manipulation of Deformable Linear Objects (DLOs) such as cables poses a significant challenge for automation due to their infinite degrees of freedom and non-linear dynamics. In this paper we present a machine learning based optimal control approach for the manipulation of DLOs. This approach is divided into two main components: modeling and control. For modeling the dynamics of the DLO, we propose a learning based approach using a bidirectional Long Short-Term Memory (biLSTM) network. The biLSTM network is trained on synthetic data generated by the MuJoCo physics engine. For manipulating the DLO, a model predictive control strategy that employs Model Predictive Path Integral (MPPI) control is selected. The proposed approach is evaluated through simulation and experiments. The results demonstrate the effectiveness of the proposed method in achieving accurate and efficient manipulation of DLOs.
Chinese Translation
诸如线缆之类的可变形线性物体(DLO)由于具有无限自由度和非线性动力学特性,对自动化操作构成了重大挑战。本文提出了一种基于机器学习的DLO操作最优控制方法。该方法分为两个主要部分:建模与控制。在DLO动力学建模方面,我们提出了一种基于学习的双向长短期记忆(biLSTM)网络方法。该biLSTM网络使用由MuJoCo物理引擎生成的合成数据进行训练。在DLO操作方面,选用了采用模型预测路径积分(MPPI)控制的模型预测控制策略。我们通过仿真和实验对该方法进行了评估。结果表明,所提出的方法能够实现对DLO的精确且高效的操作。
cs.RO / 80 / 2609.26256

High-Bandwidth Biomimetic Finger for Tactile-Transparent Remote Texture Sensing

Yang, Shuang, Liu, Fuyuan, Shao, Yitian
Abstract
High-fidelity tactile feedback is essential for robotic teleoperation, enabling precise manipulation and critical decision-making. While biomimetic fingertip sensors can capture surface-texture features, how their design shapes the tactile transparency of rendered feedback remains poorly understood. This paper presents a biomimetic fingertip replicating the human finger's multilayer mechanical gradient and fingerprint morphology, with an embedded high-sensitivity inertial measurement unit capturing texture-induced vibrations for remote vibrotactile rendering. Three variants are compared with the human fingertip through temporal and spectral analyses and a user study spanning three perceptual dimensions, establishing a transmission chain from fingertip design through signal characteristics to tactile transparency. The multilayer mechanical gradient yields the clearest improvements in signal intensity and texture-feature representation, whereas a stiffer skin layer amplifies vibration intensity at the expense of feature representation. The perceptual dimensions draw on distinct signal attributes: roughness on energy scaling, granularity on spectral shape, and repetitivity on characteristic-peak representation. These findings offer design principles for application-specific fingertips in robotic teleoperation.
cs.RO / 81 / 2609.26267

Multi-Axis Selective Decoupling Framework for Free-Floating Space Manipulators

Choi, Daegyun, Kim, Donghoon
Abstract
Free-floating space manipulator systems face operational challenges due to dynamic coupling. While traditional reactionless manipulation eliminates base disturbances, enforcing the null-space projections across all rotational axes imposes excessive constraints on the system and drastically reduces the available workspace. To address this limitation, this work proposes a multi-axis selective decoupling framework that nullifies momentum transfer exclusively along mission-critical directions. By leveraging a directionally constrained sub-coupling matrix, the framework relaxes the null-space constraints. Numerical simulations confirm that this approach mathematically eliminates constrained directional disturbances along the targeted axes while preserving the remaining degrees of freedom, enabling the manipulator to execute task-space operations successfully.
cs.RO / 82 / 2609.26292

RoboTwin-Phys: Do WAMs and VLAs Understand the Physical World?

RoboTwin-Phys:WAM 与 VLA 真的理解物理世界吗?
Zhang, Jiaqi, Ye, Feng, Yang, Mingjia, Chen, Zhihong, Xiang, Mingkang, Yao, Xinglin, Li, Yanbin, Ma, Siwei, Jia, Chuanmin
Abstract
Physical-condition diversity is largely missing from current benchmarks for robot manipulation. While large-scale simulation benchmarks increasingly incorporate variations in object appearance, scene layout, and visual observations, they typically keep the underlying physical parameters fixed. As a result, important sources of real-world variability, such as changes in mass, friction, and joint dynamics, remain largely untested. We introduce RoboTwin-Phys, a physics-diverse benchmark that treats physical-condition diversity as an explicit dimension of robot manipulation evaluation. The benchmark continuously varies 13 physical attributes within physically plausible ranges, providing a unified setting for evaluating policies across diverse physical operating conditions. We further release more than 5,000 expert demonstrations with ground-truth physical parameters, enabling physical-attribute estimation, condition-aware modeling, and physics-conditioned policy training. Evaluations of representative WAMs and VLAs reveal a substantial robustness gap: models that remain effective under existing visual and layout randomization can degrade markedly under changes in physical conditions. RoboTwin-Phys provides the benchmark, data, and evaluation protocol needed to systematically measure and improve robustness to physical-condition diversity in robot manipulation.
Chinese Translation
当前机器人操作的基准测试在很大程度上缺乏物理条件多样性。虽然大规模仿真基准逐渐纳入了物体外观、场景布局和视觉观测的变化,但它们通常保持底层物理参数固定。因此,现实世界中重要的变化来源——如质量、摩擦和关节动力学的变化——在很大程度上仍未得到测试。我们提出 RoboTwin-Phys,一个物理多样性基准,将物理条件多样性作为机器人操作评估的一个显式维度。该基准在物理合理范围内连续变化 13 个物理属性,为在不同物理运行条件下评估策略提供了统一设置。我们进一步发布了超过 5,000 条带有真实物理参数的专家演示数据,支持物理属性估计、条件感知建模以及基于物理条件的策略训练。对代表性 WAM(世界-动作模型)和 VLA(视觉-语言-动作模型)的评估揭示了显著的鲁棒性差距:在现有视觉和布局随机化下仍然有效的模型,在物理条件变化下可能明显退化。RoboTwin-Phys 提供了系统性度量和提升机器人操作对物理条件多样性鲁棒性所需的基准、数据和评估协议。
cs.RO / 83 / 2609.26304

Shaft-Configuration-Adaptive Catheter Tip Position Estimation via Motor-History Conditioned Residual Learning

Zhang, Peihan, Yip, Michael C., Kapoor, Ankur, Kim, Young-Ho
Abstract
Tendon-driven continuum manipulators are widely used in medical applications, where accurate tip-position estimation is essential for precise navigation and instrument positioning. However, patient anatomy and procedural setup impose task-dependent unknown shaft configurations, while friction, slack, and compliance introduce hysteresis, making tip estimation challenging. This paper presents a motor-history-conditioned gated recurrent unit (GRU) residual estimator for three-dimensional catheter tip estimation without direct shaft-configuration sensing. First, an initial multidirectional sweep strategy is applied to calibrate a geometric catheter model backbone, and encode the motor-angle and drive-torque response into a shaft-configuration context vector. During subsequent motion, the context conditions a GRU that predicts a task-space residual correcting this backbone, relying on motor measurements alone. The context remains fixed for the current shaft configuration, while the recurrent state captures the evolving actuation history. Across four disposable intra-cardiac echocardiography catheters and 16 bent shaft configurations, the method achieves 3.3mm open-loop tip RMSE, a 59% reduction relative to the constant-curvature baseline.
cs.RO / 84 / 2609.26313

SafeLoop: Risk-Aware Rollback for Vision-Language-Action Manipulation

Lou, Zeyu, Zhang, Tianran, Yue, Xinquan, Jing, Ya, Si, Chenyang
Abstract
Recent vision-language-action (VLA) models are promising for general-purpose manipulation, but long-horizon execution remains fragile. Small state-estimation or control errors can lead to irreversible failures (e.g., collisions and object drops). Avoiding these risks requires a proactive safety mechanism capable of anticipating hazards. In this paper, we introduce SafeLoop, a non-invasive external wrapper that adds hazard prediction and rollback-based recovery to a VLA model without changing its parameters. SafeLoop trains a risk predictor from vision and proprioception to output four values: the probability and time-to-hazard for body collisions and for object failures. A lightweight controller then chooses one of three actions based on the predicted risk: continue execution (noop), save a safety checkpoint (record), or retreat in joint space (rollback). Rollback moves the robot back to a recent safe waypoint and queries the base policy again, which may yield an alternative continuation. Across 24 LIBERO tasks (16 random seeds each) and three real-robot tasks (25 rollouts each), SafeLoop achieves a stronger overall safety-success trade-off than alternative methods, reducing hazard cases by roughly 70% while preserving task success and the base-policy control rate. Project code is available at https://github.com/Loule0-0/SafeLoop/tree/release/safeloop.
cs.RO / 85 / 2609.26314

TriWorldBench: A Tri-View Consistency Perspective on Embodied World Models

TriWorldBench:基于三视角一致性视角的具身世界模型评估
Liu, Xuanyi, Wang, Haofeng, Li, Ruiqi, Yu, Danni, Wan, Rui, Zhang, Ruixu, Tao, Siyu, Yang, Xue, Zhang, Shaofeng, Zhang, Zicheng, Zhang, Jiaqi, Ma, Siwei
Abstract
Embodied world models predict the outcomes of robot actions to support learning and planning. For robots equipped with head and wrist cameras, this requires complementary views: the head view captures the overall task, while wrist views reveal local gripper-object interactions. However, evaluating these views independently cannot determine whether they describe the same action and object state. We introduce TRIWORLDBENCH, a benchmark for evaluating embodied world models through synchronized head, left-wrist, and right-wrist videos. It contains 500 episodes across 50 bimanual manipulation tasks and uses 19 metrics to assess tri-view consistency, task alignment, physical and 3D coherence, motion quality, temporal consistency, and visual quality. By combining cross-view checks with measurements tailored to each camera, the benchmark evaluates whether plausible individual videos also form a consistent prediction of the intended task. We summarize overall performance with TWB-Score and retain per-view results to identify where predictions fail. This extends world-model evaluation beyond single-view visual quality. Code, data, and metric definitions are available at https://github.com/TriWorldBench/TriWorldBench.
Chinese Translation
具身世界模型通过预测机器人动作的结果来支持学习与规划。对于配备头部相机和腕部相机的机器人而言,这需要互补的视角:头部视角捕捉任务整体,而腕部视角揭示夹爪与物体之间的局部交互。然而,独立评估这些视角无法确定它们是否描述了同一动作和物体状态。我们提出了TRIWORLDBENCH,一个通过同步的头部、左腕和右腕视频来评估具身世界模型的基准。该基准包含50个双臂操作任务共500个回合,并使用19项指标来评估三视角一致性、任务对齐、物理与3D连贯性、运动质量、时间一致性和视觉质量。通过将跨视角校验与针对每个相机量身定制的度量相结合,该基准评估了各自看似合理的单视角视频是否也能构成对目标任务的连贯预测。我们使用TWB-Score总结整体性能,并保留各视角的结果以定位预测失败之处。这将世界模型评估从单一视角的视觉质量拓展到了更高的层次。代码、数据和指标定义可在 https://github.com/TriWorldBench/TriWorldBench 获取。
cs.RO / 86 / 2609.26315

ArborSplat: Online Semantic Gaussian Splatting SLAM for Orchards

Masini, Alessandro, Frosi, Matteo, Usuelli, Mirko, Matteucci, Matteo
Abstract
Orchard robots need maps that preserve small but semantically important structures such as trunks, trellises, and fruit. 3D Gaussian Splatting (3DGS) SLAM achieves high photometric fidelity. However, its optimization remains appearance-driven, and transferring image semantics to 3D points is unreliable for thin structures, whose pixels may receive depth from background surfaces. We present ArborSplat, an online semantic 3DGS SLAM system that tracks with LiDAR odometry and optimizes semantics directly on the Gaussian map, constrained by class-specific height bands above a ground plane fitted to each keyframe's stereo point cloud, and fuses multi-view evidence into a semantic point cloud online while rejecting labels inconsistent with the local ground surface or with monocular depth. Class-constrained refinement reserves Gaussian capacity for underrepresented structures and, under reduced budgets, increases training-view accuracy on tree classes. We evaluate the approach on apple and pear orchards during dormancy, flowering, and harvesting. On full routes, it keeps ATE below 0.5 m on all 12 traversals. On shared 301-frame segments, it exceeds SGS-SLAM and GS3LAM by 0.23 to 0.50 training-view and 0.15 to 0.36 held-out mIoU while running 1.7 to 7.5 times faster, whereas SemGauss-SLAM runs out of GPU memory on all six.
cs.RO / 87 / 2609.26360

Hierarchical Floorplan-Guided Vision-Language Exploration for Embodied Question Answering

面向具身问答的层次化户型图引导视觉-语言探索方法
Puigjaner, Albert Gassol, Alexis, Kostas
Abstract
Embodied Question Answering (EQA) requires an agent to explore a previously unseen environment, gather relevant information, and answer questions about the scene. Recent approaches leverage Vision-Language Models (VLMs) together with semantic maps or scene graphs to guide exploration. However, exploration is typically driven only by local observations, while structural priors about the environment remain largely unused. We propose HFLEX-EQA, a hierarchical EQA framework that combines online scene graph construction, VLM- based planning, semantic frontier exploration, and floorplan priors. The system incrementally builds a hierarchical scene graph and an open-vocabulary occupancy map from RGB-D observations, enabling a VLM to jointly reason over the scene graph, task-relevant visual observations, exploration history, and an estimated topological floorplan. Furthermore, we introduce a room-discovery strategy that leverages the floorplan and open-vocabulary frontier semantics to guide exploration toward semantically relevant yet currently unobserved room types. We evaluate HFLEX-EQA on the OpenEQA and ExploreEQA benchmarks and demonstrate deployment on a quadruped robot in real indoor environments. Our results demonstrate the benefit of combining VLM-based hierarchical planning with structural floorplan priors for the EQA task.
Chinese Translation
具身问答(Embodied Question Answering, EQA)要求智能体探索一个先前未见过的环境,收集相关信息,并回答有关场景的问题。近期的方法利用视觉-语言模型(Vision-Language Models, VLMs)结合语义地图或场景图来引导探索。然而,探索通常仅由局部观测驱动,而关于环境的结构先验在很大程度上未被利用。我们提出了HFLEX-EQA,一个结合在线场景图构建、基于VLM的规划、语义前沿探索以及户型图先验的层次化EQA框架。该系统从RGB-D观测中增量式地构建层次化场景图和开放词汇占据地图,使VLM能够对场景图、任务相关的视觉观测、探索历史以及估计的拓扑户型图进行联合推理。此外,我们引入了一种房间发现策略,利用户型图和开放词汇前沿语义,将探索引导向语义相关但当前尚未观测到的房间类型。我们在OpenEQA和ExploreEQA基准上对HFLEX-EQA进行了评估,并在真实室内环境中部署于四足机器人上。实验结果表明,将基于VLM的层次化规划与结构化户型图先验相结合对EQA任务具有显著收益。
cs.RO / 88 / 2609.26378

MAVP: Map-Aware Visuomotor Policies for Mobile Manipulation

Tang, Jinhe, Dai, Ruixiao, Zhi, Weiming
Abstract
Successful mobile manipulation requires coordinated base and arm motion while maintaining accurate spatial positioning. However, demonstration-trained policies can struggle to realise the intended base motion reliably, leading to spatial misalignment and subsequent manipulation failures. We present MAVP (Map-Aware Visuomotor Policies), a framework that improves execution reliability by predicting explicit base-pose targets and tracking them using localisation feedback. MAVP reconstructs a static map from teleoperated demonstrations and expresses demonstrated base trajectories in a shared map frame, providing consistent spatial supervision across demonstrations. At execution time, the policy receives RGB observations, joint states, and the robot's current map-frame base pose, and jointly predicts target base poses, arm actions, and gripper actions. A low-level controller tracks the predicted base targets using feedforward motion and pose error feedback, enabling correction of execution deviations. We additionally use pose-noise augmentation during training to improve robustness to errors in the policy's pose input. Across six real-world manipulation tasks and three policy families, MAVP achieves higher task success rates than unanchored velocity control in all tasks. Videos and additional results are available at https://123qwedsa123.github.io/mavp/.
cs.RO / 89 / 2609.26408

SparseNav: Instruction-conditioned Sparse Semantic Perception for Training-Free Vision-Language Navigation

SparseNav:面向免训练视觉语言导航的指令条件稀疏语义感知
Chen, Quanhua, Kang, Juhan, Lin, Runfeng, Zhang, ZiFei, Feng, Enquang, Zheng, Chunran, Dong, Xiwang, Lin, Jiarong
Abstract
Map-based vision-language navigation (VLN) relies on persistent spatial representations to connect language understanding with geometric planning. However, acquiring semantics beyond the needs of the current instruction can introduce unnecessary perception cost and irrelevant annotations. Continuously accumulating unrelated objects may not only waste computation, but also clutter the visual-spatial representation consumed by the vision-language model (VLM) planner. To address this problem, we present SparseNav, a training-free framework that follows a less-is-more principle for semantic navigation. SparseNav persistently maintains a lightweight geometric bird's-eye-view (BEV) map and sparse landmark memory, acquiring new semantics on demand using the active sub-instruction to decide what is worth grounding. An instruction manager first tracks navigation progress and identifies the active landmark query. An instruction-conditioned perception mechanism then invokes open-vocabulary segmentation when the queried landmark is visible and its metric location can inform the next decision. The resulting landmark memory supports VLM selection among hybrid frontier and local directional waypoint candidates. Without any additional training, SparseNav achieves success rates of 42.8% on R2R-CE and 40.7% on RxR-CE, both on the Val-Unseen splits. Controlled ablations examine semantic perception strategies and the contributions of individual framework components. Furthermore, we successfully deployed SparseNav on a Unitree Go2 quadruped equipped with an Intel RealSense D455 RGB-D camera for geometric mapping and landmark grounding and a Livox MID-360 LiDAR for localization, without a prebuilt map. We validated its effectiveness across multiple indoor environments using instruction-conditioned waypoint navigation.
Chinese Translation
基于地图的视觉语言导航(VLN)依赖持久化的空间表示来连接语言理解与几何规划。然而,获取超出当前指令所需范围的语义会带来不必要的感知开销和无关的标注信息。持续累积无关物体不仅浪费计算资源,还会使视觉语言模型(VLM)规划器所使用的视觉空间表示变得杂乱。为解决这一问题,我们提出了 SparseNav,一个遵循“少即是多”原则的免训练语义导航框架。SparseNav 持久化地维护一个轻量级的几何鸟瞰图(BEV)地图和稀疏地标记忆,并按需获取新语义,利用当前激活的子指令来决定哪些内容值得进行语义定位。指令管理器首先跟踪导航进度并识别当前激活的地标查询;随后,指令条件感知机制在被查询地标可见且其度量位置能够为下一步决策提供依据时,调用开放词汇分割。所生成的地标记忆支持 VLM 在混合的前沿候选与局部方向路径点候选之间进行选择。在无需任何额外训练的情况下,SparseNav 在 R2R-CE 和 RxR-CE 的 Val-Unseen 划分上分别取得了 42.8% 和 40.7% 的成功率。受控消融实验分析了语义感知策略以及框架各组件的贡献。此外,我们成功将 SparseNav 部署于一台配备 Intel RealSense D455 RGB-D 相机(用于几何建图与地标定位)和 Livox MID-360 激光雷达(用于定位)的 Unitree Go2 四足机器人上,且无需预先构建的地图。我们通过指令条件路径点导航在多个室内环境中验证了其有效性。
cs.RO / 90 / 2609.26420

Sample, Simulate, Select: Physics-in-the-Loop Text-to-Motion for Humanoids Without Training

Memmesheimer, Raphael, Behnke, Sven
Abstract
Text-to-motion models generate plausible human motion but do not model a robot's dynamics; whole-body tracking controllers execute robot references reliably but cannot replan an infeasible one. Recent language-to-humanoid systems bridge this gap by training. We measure how much of the gap closes with no training at all, by putting the deployment controller itself in the loop. Sample-simulate-select (S$^3$) draws $N$ motions per prompt from a frozen text-to-motion model, retargets each to a Unitree G1 by direction-matching inverse kinematics, rolls all of them out under full rigid-body dynamics with the pretrained SONIC tracking policy, and keeps the candidate the policy executed best. Because the verifier is the deterministic simulator itself, S$^3$ attains the any-of-$N$ ceiling by construction; what we measure is where that ceiling lies and what falls short of it. On 200 stratified HumanML3D test prompts with $N=8$, upright execution rises from 83.5% to 89.5% and hardware-gate passes from 33 to 85; on the complete test split (4,184 prompts) it rises from 80.5% to 89.5%. A kinematic verifier that predicts falls well (AUROC 0.90) recovers only a quarter of this gain: ranking a prompt's own candidates is harder than classifying the population. What selection cannot fix is one class, prompts that lower the pelvis, which a generator trained on retargeted robot data does execute. We further score the semantic fidelity of the executed motion with the standard text-motion evaluator, with a real-mocap control that attributes the loss to the robot projection, ablate the retargeter against GMR (complementary failures: the any-of-8 ceiling rises to 95.0% over both), and execute all 177 gate-selected clips on the real G1: every one completes standing, with hardware tracking error matching simulation ($r=0.94$).
cs.RO / 91 / 2609.26423

Dr-LiSA: Direct Radar-Lidar Scan Alignment for $SE(3)$ Localization

Zhang, Alex, Lisus, Daniil, Gentil, Cedric Le, Barfoot, Timothy D.
Abstract
This paper introduces Dr-LiSA, a first-of-its-kind direct method for localizing 2D spinning radar intensity measurements in $SE(3)$ against 3D lidar maps. Radar-lidar localization combines the complementary strengths of the two sensing modalities: radar is robust to adverse weather and precipitation, while lidar provides high-fidelity 3D maps in favourable conditions. However, existing radar-lidar localization methods are restricted to planar $SE(2)$ localization and have generally fallen short of the accuracy achieved by lidar-lidar and even radar-radar systems. A key challenge is the substantial sensing-modality gap between radar and lidar, which observe and represent scene structure in fundamentally different ways. Dr-LiSA bridges this gap using a learned forward model that predicts radar measurements from a lidar submap at a candidate pose, enabling direct photometric alignment of predicted and observed radar scans in $SE(3)$. Dr-LiSA outperforms prior radar-lidar approaches in $SE(2)$ while achieving planar accuracy competitive with state-of-the-art radar-radar localization across more than 90 km of on-road data.
cs.RO / 92 / 2609.26467

RouteRLT: Learning When and Which RL Specialist Should Control a Vision-Language-Action Policy

RouteRLT:学习何时以及由哪个强化学习专家策略来控制视觉-语言-动作模型
Zhu, Chongyu, Hinds, Jaden, Kim, Hyegang, Rojas, Juan Sebastian, Elmallah, Ramy, Lee, Chi-Guhn
Abstract
Vision-language-action (VLA) models provide broad manipulation competence, but often struggle during the precision-critical stages that dominate contact-rich industrial tasks such as connector insertion and cable management. A common remedy is to refine a pretrained VLA with reinforcement learning (RL), enabling task-specific improvement beyond behavior cloning. However, how to preserve its generalist behavior while deciding when RL refinement is needed and which specialized policy should act remains an open question. In this work, we present RouteRLT, a routing framework that learns when and which RL specialist, an RL policy trained for a single precision-critical phase, should take control from a generalist VLA. A phase selector identifies the active controller, a stabilizer suppresses transient switches, and an action-boundary manager handles transitions between chunked policy outputs. We evaluate RouteRLT on multi-object pick-and-place tasks in LIBERO, as well as on a real-world cable pickup and port-insertion task with multiple precision-critical stages. In simulation, the learned routing improves over the base VLA and matches routing with privileged phase boundaries, without accessing those boundaries at deployment. The real-robot evaluation validates automatic routing to both the pickup and insertion specialists under an operator-aligned handoff protocol. Altogether, these results show that learned routing applies RL specialist control where precise adaptation is most valuable while preserving generalist VLA behavior, including recovery from failed execution attempts.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)模型具备广泛的操作能力,但在接触密集型工业任务(如连接器插接和线缆管理)中占主导地位的精度关键阶段往往表现不佳。一种常见的补救方法是使用强化学习(RL)对预训练的VLA模型进行微调,从而在行为克隆的基础上实现针对特定任务的改进。然而,如何在保持其通用行为的同时,决定何时需要RL微调以及应由哪个专用策略来执行,仍然是一个悬而未决的问题。在本工作中,我们提出了RouteRLT,这是一个路由框架,用于学习何时以及由哪个RL专家策略(即为单一精度关键阶段训练的RL策略)应从通用VLA模型接管控制权。其中,阶段选择器(phase selector)识别当前活跃的控制器,稳定器(stabilizer)抑制瞬态切换,动作边界管理器(action-boundary manager)处理分块策略输出之间的过渡。我们在LIBERO的多物体拾取与放置任务上,以及在具有多个精度关键阶段的真实世界线缆拾取和端口插接任务上评估了RouteRLT。在仿真中,学习到的路由方法优于基础VLA模型,并达到了使用特权阶段边界进行路由的性能,且在部署时无需访问这些边界。真实机器人评估验证了在操作员对齐的交接协议下,系统能够自动路由至拾取和插接专家策略。总体而言,这些结果表明,学习到的路由方法能够在精确适应最有价值的地方应用RL专家控制,同时保持通用VLA的行为,包括从失败的执行尝试中恢复的能力。
cs.RO / 93 / 2609.26490

Benchmarking Robots for Everyday Environments: From Lab Experiments to Real-World Operations

面向日常环境的机器人基准测试:从实验室实验到真实场景运行
Memmesheimer, Raphael, Overbeck, Martina, Beyer, Dominik, Kral, Björn, Bellmann, Sabine, Schneider, Sven, Zimmermann, Jan, Meer, Anna-Maria, Klicic, Medina, Roth, Simone, Straßmann, Carolin, Arntz, Alexander, Wessels, Marlene, Kraus, Johannes, Schweidler, Paul, Schnell, Tristan, Zimmermann, Christoph, Pulver, Benedikt, Stork, Wilhelm, Gersch, Martin, Behnke, Sven, Rönnau, Arne
Abstract
This study introduces an interdisciplinary framework for benchmarking robots deployed in public environments, addressing the gap between traditional laboratory metrics and real-world benchmarking requirements. We evaluate three distinct robots across diverse use cases - outdoor park cleaning, pedestrian underpass cleaning, and interactive library assistance - each representing unique challenges in public daily life. Over a three-year benchmarking process (2023-2025) comprising seven benchmarking events, a consensus workshop and six on-site evaluations (two per use case), we utilized realistic indoor and outdoor test environments to assess not only technical performance but also the broader implications of deploying robots in unstructured, human-centric settings. An expert panel, spanning robotics, human-robot interaction, safety, and economics, systematically developed and refined an evaluation concept to analyze the transition from laboratory prototypes to operational systems. Our findings highlight critical factors for successful deployment, including task fulfillment, interaction quality, safety, and economic feasibility. This work provides actionable insights for researchers and practitioners aiming to bridge the gap between robotic innovation and real-world applicability.
Chinese Translation
本研究提出了一种跨学科的框架,用于对部署在公共环境中的机器人进行基准测试,旨在弥合传统实验室指标与真实世界基准测试需求之间的差距。我们针对三个不同的机器人及其多样化的应用场景进行评估——户外公园清洁、行人地下通道清洁以及交互式图书馆辅助服务——每个场景都代表了公共日常生活中的独特挑战。在为期三年的基准测试过程(2023-2025年)中,包括七次基准测试活动、一次共识研讨会和六次现场评估(每个应用场景两次),我们利用真实的室内外测试环境,不仅评估了机器人的技术性能,还评估了在非结构化、以人为中心的环境中部署机器人所产生的更广泛影响。一个涵盖机器人学、人机交互、安全和经济学的专家组系统地制定并完善了评估方案,用以分析从实验室原型到可运行系统的转换过程。我们的研究结果强调了成功部署的关键因素,包括任务完成度、交互质量、安全性和经济可行性。这项工作为致力于弥合机器人创新与实际应用之间差距的研究人员和从业者提供了可操作的见解。
cs.RO / 94 / 2609.26499

Generalizing Manipulation Skills with a Local Coding Agent

基于本地代码代理的操作技能泛化
Talwar, Raman, Nijs, Elias, Verleysen, Andreas, wyffels, Francis
Abstract
Today, progress in open-weight language models enables systems capable of writing, executing and debugging code while still running on a single workstation. Most language-driven robots give the model a fixed action interface or a trained policy. Generalizing to a new task therefore means more engineering effort or more data collection, both time-consuming. We investigate whether a local open-weight vision-language model can control a robot and one-shot generalize to new variations of a task without new human programming or training. We let a local open-weight VLM, Qwen3.8-27B, drive a UR3e robotic arm from a coding-agent harness. It writes and runs its own code above a service that implements kinematics, safety limits and classic computer vision techniques. We investigate if this system is capable of generalizing to unseen tasks. Specifically, we test it on nine tasks built from children's toys designed to probe generalization capability across various object characteristics: color, size, shape, and task variation of those. With five trials for each task, we observe generalization in 30 out of 45 trials with durations ranging from 3.4 to 67.5 minutes depending on task complexity. We further test if there is a speedup when an agent is asked to redo the task after successful completion. This resulted in a 50% reduction in duration, indicating that there is self-improvement over time. Finally, we expose the limitations of a local coding agent. We believe that solving those limitations combined with further investigation of self-improvement over time points at a direct path toward real-world deployment of a local coding agent.
Chinese Translation
如今,开放权重语言模型的进展使得能够在单一工作站上运行的系统具备编写、执行和调试代码的能力。大多数由语言驱动的机器人为模型提供固定的动作接口或已训练的策略。因此,泛化到新任务意味着更多的工程工作或更多的数据收集,两者都非常耗时。我们研究了一个本地开放权重的视觉-语言模型能否控制机器人,并在无需新的人类编程或训练的情况下,对新任务变体实现一次性泛化。我们让本地开放权重VLM(Qwen3.8-27B)通过代码代理框架驱动一台UR3e机械臂。它在一个实现了运动学、安全限制和经典计算机视觉技术的服务之上编写并运行自己的代码。我们研究该系统能否泛化到未见过的任务。具体而言,我们在九个由儿童玩具构建的任务上进行测试,这些任务旨在探测模型在各种物体特性上的泛化能力:颜色、尺寸、形状及其任务变体。每个任务进行五次试验,我们在45次试验中观察到30次泛化成功,耗时从3.4到67.5分钟不等,取决于任务复杂度。我们进一步测试了当代理被要求在成功完成任务后重做该任务时是否存在加速。结果显示耗时减少了50%,表明系统随时间存在自我改进。最后,我们揭示了本地代码代理的局限性。我们认为,解决这些局限性并结合对随时间自我改进的进一步研究,指向了本地代码代理走向真实世界部署的直接路径。
cs.RO / 95 / 2609.26520

MATE: Multi-Agent Virtual Teleoperation Platform for Humanoid Collaboration Data Collection

MATE:面向人形机器人协作数据收集的多智能体虚拟遥操作平台
Yu, Yichuan, Wang, Youzhuo, Ren, Yiming, Feng, Di, Yang, Yexuan, Yang, Bingxi, Gong, Shengxiao, Sun, Yujing, Ma, Yuexin
Abstract
Humanoid robots require diverse embodied experiences to acquire complex loco-manipulation and collaborative skills. However, existing humanoid data pipelines primarily focus on individual agents, while physical multi-robot collaboration remains difficult to scale due to costly hardware, dedicated spaces, and repeated resets. In this work, we introduce MATE, a Multi-Agent virtual TEleoperation platform for humanoid collaboration data collection that enables multiple geographically distributed operators to simultaneously control whole-body humanoids in a shared physics-based environment. MATE removes the need for multiple physical robots and co-located operation while preserving physically coupled interactions among humanoids, objects, and environments. Using MATE, we construct a multi-humanoid collaboration dataset comprising 24.1 hours of coordinated behavior across 2,500 joint episodes and five long-horizon tasks, including object handover, relay delivery, environment interaction, and cooperative transport. To improve learning from these interaction-rich demonstrations, we introduce EAIS, an Execution-Aligned Interaction Sampling strategy that computes sampling signals within an execution-aligned prefix and prioritizes task-progressing and interaction-critical behaviors. We evaluate MATE with representative imitation learning and vision-language-action policies across diverse collaboration tasks. Experiments demonstrate efficient data collection, effective policy learning, and zero-shot transfer from virtual demonstrations to a physical humanoid without real-world fine-tuning. Project page: https://yerik-yu.github.io/MATE/
Chinese Translation
人形机器人需要多样化的具身经验来掌握复杂的运动-操作与协作技能。然而,现有人形机器人数据流程主要关注单体智能体,而物理多机器人协作由于硬件成本高昂、需要专用场地以及频繁复位等原因,难以规模化扩展。在本工作中,我们提出MATE(Multi-Agent virtual TEleoperation),一个用于人形机器人协作数据收集的多智能体虚拟遥操作平台,使多个地理位置分布的操作者能够在一个共享的基于物理的环境中同时控制全身人形机器人。MATE免除了对多台物理机器人和同地协作的需求,同时保留了人形机器人、物体与环境之间物理耦合的交互。利用MATE,我们构建了一个多人形机器人协作数据集,包含跨越2500条联合回合和五个长时序任务的24.1小时协同行为数据,涵盖物体交接、接力运送、环境交互和协作搬运等任务。为了提升从这些富含交互的演示中的学习效果,我们提出EAIS(Execution-Aligned Interaction Sampling,执行对齐交互采样)策略,该策略在执行对齐的前缀内计算采样信号,并优先采样推动任务进展和交互关键的行为。我们在多种协作任务上使用代表性的模仿学习和视觉-语言-动作(VLA)策略对MATE进行了评估。实验结果表明,MATE能够实现高效的数据收集和有效的策略学习,并能在无需真实世界微调的情况下,将虚拟演示的策略零样本迁移到物理人形机器人上。项目页面:https://yerik-yu.github.io/MATE/
cs.RO / 96 / 2609.26564

Learning Air-Ground Motion Control with Temporal Mode Switching and Cross-Terrain Tracking

Pang, Ruitian, Li, Mingrui, Liu, Xuanting, Lai, Tiancheng, Chen, Juncheng, Li, Xiangyu, Zhang, Ruibin, Wang, Qishao, Yu, Jin, Piao, Haiyin, Gao, Fei, Xu, Chao, Cao, Yanjun
Abstract
Passive-wheeled terrestrial-aerial bimodal vehicles (TABVs) combine aerial mobility with energy-efficient ground locomotion. However, reliable air-ground mode switching under limited onboard perception and robust ground trajectory tracking across diverse terrains remain challenging when targeting real-world applications. In this work, we propose a learning-based air-ground motion control framework for passive-wheeled TABVs: 1) a learned mode selector for autonomous air-ground motion mode switching. The selector uses historical single-point time-of-flight (ToF) measurements and robot states together with future reference information to determine the active locomotion mode. 2) a reinforcement learning control policy for trajectory tracking. The policy combines proprioceptive observations with future reference information to anticipate trajectory changes. For ground locomotion, multi-terrain training and dynamics randomization enable robust tracking across different terrains. Simulation and real-world experiments demonstrate reliable air-ground switching under limited perception and accurate ground tracking across diverse terrain conditions. The learned selector outperforms a rule-based mode selector in challenging transitions, while the ground controller achieves lower position RMSE than PID across all tested conditions and maintains decent tracking where NMPC fails. With these capabilities integrated, the system tracks a 101m air-ground trajectory through multiple autonomous mode transitions with a position RMSE of 0.08m.
cs.RO / 97 / 2609.26567

Beyond End-Task Success: How to Audit Visual Experience Retrieval in Robotics

超越最终任务成功:如何在机器人学中审计视觉经验检索
Pathak, Eshika, Krishna, Leela
Abstract
Robots that store past experiences must select which one to reuse in a new scene. Most systems select by visual similarity, and most evaluations report only the success of the selected experience. That number does not show whether the selection was good: a rule can score well by repeatedly using one broadly transferable experience, or poorly because its preferred experience is weak. Since robots increasingly adapt by reuse rather than retraining, a score that describes the library rather than the rule misleads what the field builds next. We contribute an audit methodology: execute every stored experience in every query scene, over two manipulation tasks, three reuse mechanisms, and libraries of $K=3$, $10$, and $50$. Because every alternative's outcome is known, a score can be traced to per-scene selection or to library quality. The audited rules select by nearest-neighbor distance in five visual embeddings, from raw pixels to CLIP. (1) One fixed experience, chosen with hindsight, captures 30-58% of the gap between random selection and an oracle; per-scene selection competes for the remaining 0.07-0.15 in success rate. (2) At $K\ge10$, visual rules concentrate on one experience 1.5-3 times more than the oracle does, and their scores then follow that experience's quality. (3) Wherever a rule differs significantly from a shuffle that keeps its selection rates but pairs them with scenes at random, the rule is worse, for every learned image policy. (4) Visual distance predicts well whether a given pair will succeed (AUROC up to 0.96), yet ranks the candidates within one scene no better than chance for four of five embeddings at $K=50$ (AUROC 0.45-0.52). Exhaustive execution is usually infeasible, so the audit reduces to two cheap reports any study can give: the distribution of selected experiences, and the success of the best single experience in hindsight.
Chinese Translation
存储过往经验的机器人必须在新场景中选择复用哪一条经验。大多数系统依据视觉相似度进行选择,且大多数评估仅报告所选经验的成功率。该数字无法说明选择本身是否良好:一个选择规则可能因反复使用某条广泛可迁移的经验而得分很高,也可能因其偏好的经验本身较弱而得分很低。由于机器人日益通过经验复用而非重新训练来适应环境,一个描述经验库质量而非选择规则好坏的分数会误导该领域的后续研究方向。我们贡献了一种审计方法:在两个操作任务、三种复用机制以及规模为 $K=3$、$10$ 和 $50$ 的经验库上,在每一查询场景中执行每一条存储经验。由于每个备选经验的结果均已知,分数可以追溯到逐场景选择的好坏或经验库质量。被审计的规则基于五种视觉嵌入(从原始像素到 CLIP)中的最近邻距离进行选择。(1) 事后选取的一条固定经验可获得随机选择与预言机(oracle)之间差距的 30-58%;逐场景选择则竞争剩余的 0.07-0.15 成功率。(2) 当 $K\ge10$ 时,视觉规则对某一条经验的集中程度是预言机的 1.5-3 倍,其分数随之追随该经验的质量。(3) 在任何规则显著不同于一种保持其选择率但将经验与场景随机配对的洗牌机制之处,对所有学习到的图像策略而言,该规则都表现更差。(4) 视觉距离能较好地预测一对(经验,场景)是否会成功(AUROC 最高达 0.96),但在 $K=50$ 时,五种嵌入中有四种在单个场景内对候选经验的排序并不比随机更好(AUROC 为 0.45-0.52)。穷举执行通常不可行,因此该审计可简化为任何研究都能提供的两个低成本报告:所选经验的分布,以及事后最佳单条经验的成功率。
cs.RO / 98 / 2609.26580

Wheel-loader V-Cycle Automation with Deep Koopman MPC

基于深度Koopman模型预测控制的装载机V形循环作业自动化
Abdolmohammadi, Armin, Mojahed, Navid, Kumar, Dinesh, Ravani, Bahram, Nazari, Shima
Abstract
The repeated forward-reverse maneuvers performed by wheel loaders during earthmoving operations make them well suited for automation. However, the nonlinear dynamics of articulated vehicles and complex vehicle-terrain interactions limit the effectiveness of conventional model-based approaches. This paper presents a hierarchical framework that combines long-horizon geometric planning with data-driven predictive control for autonomous wheel-loader operation. A reduced-order articulated kinematic model is used to generate the maneuver geometry, where the forward and reverse trajectories are jointly optimized through a shared intermediate state. To capture the vehicle dynamics, two data-driven deep bilinear Koopman models are learned for the forward and reverse motions using data generated from high-fidelity simulations in Algoryx Dynamics. The learned Koopman representations are subsequently incorporated into a computationally efficient model predictive control (MPC) formulation for trajectory tracking. The resulting controller operates in real time within a 50-ms execution loop. High-fidelity simulation results demonstrate that the proposed end-to-end framework enables accurate and computationally efficient execution of wheel-loader V-cycle maneuvers, providing a promising approach toward autonomous operation of articulated heavy-duty machinery.
Chinese Translation
轮式装载机在土方作业中执行的反复前进-后退机动使其非常适合自动化。然而,铰接式车辆的 nonlinear 动力学以及复杂的车-地交互限制了传统基于模型方法的有效性。本文提出了一种分层框架,将长时域几何规划与数据驱动预测控制相结合,用于轮式装载机的自主作业。该框架采用降阶铰接运动学模型生成机动几何形状,通过共享的中间状态对前进和后退轨迹进行联合优化。为捕捉车辆动力学特性,利用Algoryx Dynamics高保真仿真生成的数据,分别针对前进和后退运动学习得到两个数据驱动的深度双线性Koopman模型。随后,将学习到的Koopman表示嵌入计算高效模型预测控制(MPC)框架中以实现轨迹跟踪。所得到的控制器可在50毫秒执行循环内实时运行。高保真仿真结果表明,所提出的端到端框架能够准确且计算高效地执行轮式装载机V形循环机动,为铰接式重型机械的自主作业提供了一种有前景的方法。
cs.RO / 99 / 2609.26618

NavSafe-$\infty$: Benchmarking Closed-Loop Driving Safety in Photorealistic Environments

NavSafe-$\infty$: 在照片级真实感环境中对闭环驾驶安全性进行基准测试
Bao, Yuxin, Ruan, Hongwei, Wang, Luobin, Zhao, Seth Z., Leng, Ziyang, Zhang, Zihan, Zeng, Yu, McAllister, Rowan, Christensen, Henrik, Zhou, Bolei
Abstract
End-to-end (E2E) driving policies have progressed rapidly on open-loop (OL) benchmarks, yet OL evaluation cannot reveal whether a policy withstands compounding errors, recovers from failures, or interacts safely with surrounding actors. We introduce NavSafe-$\infty$, a photorealistic closed-loop (CL) benchmark of 280 scenarios spanning 28 event types, each with success and failure criteria defined within a structured traffic-safety taxonomy, which yields category-level capability scores for Traffic Crashes, Vulnerable Road User Crashes, Traffic Violations, and Traffic Incidents. Evaluating 20 E2E policies, we find that OL gains do not reliably transfer to CL safety. Analyzing two common remedies further shows that passive demonstration perturbation helps mainly when CL rollouts stay near its perturbed training states, and that OL reinforcement-learning fine-tuning exhibits reward hacking by trading safety margin for ego progress, which CL feedback amplifies into compounding safety-critical errors. Together, these results demonstrate the blind spot of OL benchmarks indicating CL safety success. The benchmark and an extensible toolbox for customizable event curation and policy diagnosis will be open-sourced and maintained to facilitate future research.
Chinese Translation
端到端(E2E)驾驶策略在开环(OL)基准测试上进展迅速,然而开环评估无法揭示策略是否能够承受误差累积、从失败中恢复,或与周围交通参与者安全交互。我们提出了NavSafe-$\infty$,这是一个包含280个场景、覆盖28种事件类型的照片级真实感闭环(CL)基准,每个场景都在结构化的交通安全分类体系中定义了成功与失败标准,从而产生交通碰撞、弱势道路使用者碰撞、交通违章和交通事件等类别层面的能力评分。通过对20个端到端策略的评估,我们发现开环性能的提升并不能可靠地转化为闭环安全性。对两种常见补救措施的进一步分析表明:被动式演示数据扰动仅在闭环回放轨迹保持在其受扰动的训练状态附近时才有效;而开环强化学习微调则表现出奖励作弊行为,即以牺牲安全裕度为代价换取自车行驶进度,这种问题在闭环反馈下被放大为不断累积的安全关键错误。这些结果共同揭示了开环基准测试无法指示闭环安全成功的盲区。该基准以及一个支持自定义事件策划与策略诊断的可扩展工具箱将被开源并持续维护,以促进未来研究。
cs.RO / 100 / 2609.26672

Imperfection for Precision: Upcycling Imperfect Data for High-Precision Robotic Manipulation

Wei, Hao, Liu, Yang, Tang, Chao, Li, Shengbao, Chen, Jiangtao, Zhu, Jinxuan, Wang, Jiaheng, Yin, Hong, Cao, Zhaofeng, Li, Tingguang
Abstract
Training vision-language-action (VLA) models for high-precision manipulation typically requires task-specific, high-quality data (e.g., teleoperation), which is slow and expensive to collect. To reduce this burden without compromising manipulation precision, we propose $\varepsilon$4P (Imperfection for Precision), a simple yet effective method that "upcycles" two otherwise discarded data sources: (1) low-precision data from the target task and (2) high-precision data from mismatched tasks. Rather than naively mixing these imperfect data sources throughout co-training, $\varepsilon$4P controls where each source contributes along the flow-matching trajectory. Specifically, low-precision, target-task data is used at high noise to preserve high-level task context and high-precision, task-mismatched data is used at low noise to transfer low-level action precision. Through real-robot experiments on both sub-millimeter, high-precision tasks and coarse-grained tasks, we demonstrate that the proposed method (1) effectively leverages additional imperfect data to improve policy performance by up to 31.7 percentage points, and (2) can replace an equal amount of task-specific, high-quality data with an average performance drop of only 4.2 percentage points. Overall, $\varepsilon$4P points toward a scalable paradigm for high-precision manipulation, in which heterogeneous, imperfect data can be systematically repurposed to reduce reliance on costly task-specific, high-quality data. More details are available at https://varepsilon4p.github.io/.
cs.RO / 101 / 2609.26753

Underwater Navigation in Unsteady Flows Using Measurement Histories from a Single Sensing Unit

基于单一传感单元测量历史的非定常流场水下导航
Jin, Linhao, Feng, Qimin, Gunnarson, Peter, Zhong, Qiang
Abstract
Spatial flow measurements support underwater navigation, but distributed sensing is constrained by robot size and sensor layout. We use a causal observer to estimate current lateral velocities from a finite history of measurements collected by a single sensing unit, supplying the inputs of a fixed navigation controller. In two-dimensional wake simulations with access to body-frame ambient velocity, this virtual sensing interface reduces simultaneous flow sampling from three points to the robot center. Trained only in a circular-cylinder wake at Re = 100, the flow-history observer achieves 84.4% and 80.6% success at held-out Re = 205 and 240 without retraining. These rates are 7.4 and 4.6 percentage points below direct spatial sensing and more than 30 points above a matched current-only observer. Past flow remains beneficial when past goal and yaw information is available. Across obstacle geometries, performance remains close to direct sensing in square-prism wakes but declines in triangular-prism wakes. Component replacement identifies the lateral velocity difference as control-relevant, while controlled perturbations reveal sensitivity to error persistence. The results demonstrate the closed-loop utility of single-point flow histories under the assumed observation model.
Chinese Translation
空间流场测量可支持水下导航,但分布式传感受到机器人尺寸和传感器布局的限制。我们使用因果观测器从单一传感单元采集的有限测量历史中估计当前侧向速度,为固定的导航控制器提供输入。在可获取体坐标系环境速度的二维尾流仿真中,这一虚拟传感接口将同时进行流场采样的测点从三个减少到机器人中心。仅在雷诺数 Re = 100 的圆柱尾流中训练的流场历史观测器,在未参与训练的 Re = 205 和 240 条件下,无需重新训练即可分别达到 84.4% 和 80.6% 的成功率。这些成功率比直接空间传感低 7.4 和 4.6 个百分点,而比匹配的仅用当前信息的观测器高出 30 个百分点以上。当过去的目标和偏航信息可用时,历史流场信息仍然有益。在不同障碍物几何形状下,性能在方形棱柱尾流中保持接近直接传感水平,但在三棱柱尾流中有所下降。组件替换实验表明侧向速度差是与控制相关的量,而受控扰动实验则揭示了对误差持续性的敏感性。这些结果证明了在假设的观测模型下单点流场历史信息的闭环应用价值。
cs.RO / 102 / 2609.26766

TM-APR: Thermal Temporal-Memory Localization via Analytic Online Adaptation

Bai, Yanshuo, Tanaka, Kanji
Abstract
Thermal Visual Place Recognition (Thermal VPR) maps camera observations to metric poses within a mapped environment, serving as a prerequisite for autonomous navigation. However, thermal VPR suffers from severe environmental dependence, heavy online retraining overheads, and an inability to model dynamic non-linear shifts, causing existing frameworks to fail during online deployment. To achieve robust domain-invariant place recognition, we bridge Analytic Class-Incremental Learning (ACIL) with domain-invariant VPR for the first time, revealing that its gradient-free matrix updates construct a surprisingly strong baseline that outperforms conventional fine-tuning. Nevertheless, standard ACIL exhibits a critical vulnerability to extreme non-linear thermal fluctuations due to its structural linear assumptions. To overcome this limitation, we exploit a novel algebraic equivalence between ACIL and modern control theory, proposing a framework which embeds Unscented propagation (U-ACIL), Gaussian Mixture partitioning (GMM-ACIL), and minimax $H_\infty$ optimization ($H_\infty$-ACIL) directly into the update loop. Our formulation guarantees exact closed-form matrix updates within $\mathcal{O}(1)$ computational complexity, bypassing backpropagation to ensure that the online update latency ($\Delta t_{\mathrm{learn}}$) remains strictly bounded below the sensor acquisition interval ($\Delta t_{\mathrm{acquire}}$), thereby eliminating trajectory jumps in real-time SLAM pipelines.
cs.RO / 103 / 2609.26792

DreamStream: Towards Policy-Oriented Generative Simulation for End-to-End Driving

DreamStream:面向端到端驾驶的策略导向生成式仿真
Leng, Ziyang, Mo, Sicheng, Zhao, Seth Z., Cai, Haoyuan, Zeng, Yu, McAllister, Rowan, Zhou, Bolei
Abstract
Faithfully evaluating end-to-end driving policies in simulation requires observations that are not merely photo-realistic, but preserve the scene features a policy relies on to make decisions. Existing platforms, however, exhibit a sim-to-real visual gap that corrupts policy perception, undermining their ability to assess a policy's closed-loop decision-making. To this end, we propose DreamStream, a generative, closed-loop simulator that achieves policy-oriented fidelity using a simulator-grounded autoregressive video model. Our video model is distilled from a large pretrained video model via traffic layout guidance, varying visual appearance while preserving policy-relevant features such as scenario layout and the temporal consistency of dynamic objects. We further observe that perceptual metrics like FID misrank how well these features are preserved. To tackle this, we introduce FD$\pi$, a new multi-representation metric that measures the sim-to-real gap as the Fr\'echet distance over scene-context features from public E2E policies. Under FD$\pi$, DreamStream improves over the strongest prior closed-loop simulator by $1.6\times$ on nuScenes and $4.7\times$ on NAVSIM, and induces the least perturbation to policy's perceptual observability. Based on DreamStream, we construct Navhard-CL benchmark, which turns non-reactive real-world benchmark NAVSIM into interactive testing environments with adversarial driving behaviors and weather variations. This benchmark exposes many failure modes of driving policies, such as scorer bias and lack of recovery behaviors, that prior closed-loop benchmarks overlook. Code and data are available at https://github.com/VAIL-UCLA/DreamStream.
Chinese Translation
在仿真中忠实地评估端到端驾驶策略,要求观测结果不仅具有照片级的真实感,还必须保留策略赖以决策的场景特征。然而,现有平台存在模拟到现实(sim-to-real)的视觉差距,会破坏策略的感知能力,从而削弱其评估策略闭环决策能力的作用。为此,我们提出 DreamStream,一个基于仿真器、采用自回归视频模型的生成式闭环仿真器,实现了面向策略的保真度。我们的视频模型通过交通布局引导,从大规模预训练视频模型中蒸馏而来,在改变视觉外观的同时,保留了与策略相关的特征,如场景布局和动态目标的时序一致性。我们进一步发现,FID 等感知类指标会错误地评估这些特征的保留程度。为解决这一问题,我们提出 FD$\pi$,这是一种新的多表征指标,它基于公开端到端策略的场景上下文特征,以 Fréchet 距离度量模拟到现实的差距。在 FD$\pi$ 指标下,DreamStream 在 nuScenes 上相比此前最强的闭环仿真器提升了 1.6 倍,在 NAVSIM 上提升了 4.7 倍,并且对策略感知可观测性的扰动最小。基于 DreamStream,我们构建了 Navhard-CL 基准,将非响应式的真实世界基准 NAVSIM 转化为包含对抗性驾驶行为和天气变化的交互式测试环境。该基准暴露了许多此前闭环基准所忽略的驾驶策略失效模式,如评分器偏差和缺乏恢复行为等。代码和数据可在 https://github.com/VAIL-UCLA/DreamStream 获取。
cs.RO / 104 / 2609.26795

\phi-RIE: From Photorealistic Reconstruction to Interactive Environments

\phi-RIE:从照片级真实感重建到可交互环境
Yang, Runyi, Zhang, Deheng, Wang, Xiaoye, Wu, Kanzhi, Sun, Lei, Chhatkuli, Ajad, Peng, Kunyu, Van Gool, Luc, Paudel, Danda Pani
Abstract
3D Gaussian Splatting (3DGS) can reconstruct a captured scene photorealistically, but the resulting representation does not by itself support physical interaction. Robot simulation instead requires object-level change, \textit{i.e.}, objects must move independently, make contact, and reveal previously occluded surroundings. This gap arises because object appearance may remain entangled with the background, while hidden object geometry and occluded background content may be unobserved. To address this challenge, we present \phi-RIE, a Gaussian-native pipeline that converts selected objects into movable simulator assets while preserving the remaining reconstruction. Our key observation is that asset construction and source removal should be coupled, \textit{i.e.}, one object identity should define the movable asset and the scene content to remove and complete. Accordingly, Scene Observation supplies shared evidence to Coupled Scene Construction, which creates registered assets and completed background Gaussians for simulator-driven rendering in an Interactive Environment. This coupling preserves unedited Gaussians while aligning visual and physical state. On 50 ScanNet++ scenes, evidence-based selection and registration retry increase matched F1 at 20\,mm from 0.336 to 0.383 at fixed retention. Further tests demonstrate asset executability, manipulation gains over a single-generator baseline, and the visual cost of conversion. Together, these results demonstrate that \name\ enables interactive scene conversion.
Chinese Translation
3D高斯泼溅(3D Gaussian Splatting, 3DGS)能够对捕获的场景进行照片级真实感重建,但所得的表示本身并不支持物理交互。机器人仿真则要求对象级别的变化,即物体必须能够独立移动、发生接触,并暴露此前被遮挡的周围环境。这一差距的存在是因为物体的外观可能与背景纠缠在一起,而隐藏的物体几何结构和被遮挡的背景内容可能未被观测到。为应对这一挑战,我们提出了\phi-RIE,一个高斯原生(Gaussian-native)的流程,它将选定的物体转换为可移动的仿真器资产,同时保留其余的重建结果。我们的关键观察是:资产构建与源物体移除应当耦合进行,即应由同一个物体身份同时定义可移动资产以及需要移除与补全的场景内容。据此,场景观测(Scene Observation)为耦合场景构建(Coupled Scene Construction)提供共享证据,后者创建已配准的资产和补全后的背景高斯,用于交互环境(Interactive Environment)中由仿真器驱动的渲染。这种耦合方式在保持未编辑高斯不变的同时,对齐了视觉状态与物理状态。在50个ScanNet++场景上,基于证据的选择与配准重试机制在固定保留率下将20\,mm阈值的匹配F1分数从0.336提升至0.383。进一步的测试证明了资产的可执行性、相对于单生成器基线的操作性能提升,以及转换带来的视觉代价。这些结果共同表明,\name\ 能够实现可交互的场景转换。
人工智能 (Artificial Intelligence)
101
cs.AI / 1 / 2609.25010

Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation

合成人设(Persona)能否预测真实受众反应?一项“模拟到真实”(Sim-to-Real)的效度研究:无人设基线胜过基于人设的文案模拟
Maiorano, Alexandre Cristovão
Abstract
Marketers increasingly use large language models (LLMs) as "synthetic personas" to predict how an audience will react to a piece of copy before it ships, encouraged by evidence that profile-conditioned LLMs mimic human samples. But is that prediction actually valid against real behaviour - and does the persona machinery help? We present a sim-to-real validity study using the Upworthy Research Archive - thousands of headline A/B tests on shared real traffic, with measured click-through - as held-out ground truth. We compare a ten-persona panel, grounded in the real audience's demographics, against a no-persona zero-shot baseline that simply asks the model how likely a typical reader is to click. Two findings stand out. First, ground-truth reliability is the binding constraint: most A/B tests have no statistically distinguishable winner, so validity can only be measured on the reliable subset (n = 399). Second, and counter to the persona-simulation premise, persona conditioning degrades predictive validity: the no-persona baseline ranks variants markedly better (Kendall {\tau} = 0.361, a medium effect; top-1 accuracy 49.2%) than the persona panel ({\tau} = 0.084; top-1 34.6%), with non-overlapping confidence intervals. Asking the model directly taps an accurate population-level prior; forcing it to role-play specific personas injects bias and noise. The result replicates across three independent Upworthy splits, holds in direction on a different-domain news dataset, and is robust to seed, prompt phrasing, and model choice - across three Gemini tiers and a different model family (OpenAI gpt-4.1, significant paired gap). The takeaway: for predicting aggregate engagement, a plain LLM ranker beats persona simulation - synthetic personas are not merely a weak predictor, they are worse than not using them. All numbers regenerate from a public, artifact-first replication package.
Chinese Translation
营销人员日益将大语言模型(LLM)作为“合成人设”(synthetic personas)来预测受众在文案发布前对其反应,其依据是已有证据表明基于画像条件化的LLM能够模仿人类样本。但这种预测对于真实行为是否确实有效——人设机制是否真有帮助?我们利用Upworthy研究档案(Upworthy Research Archive)开展了一项模拟到真实的效度研究。该档案包含数千项在共享真实流量上进行的标题A/B测试,并带有实测点击率,可作为留出的真值标准(ground truth)。我们将一个基于真实受众人口统计特征构建的十人设小组,与一个简单询问模型“典型读者点击该文案的可能性有多大”的无人设零样本(zero-shot)基线进行比较。研究得出两个突出发现。第一,真值信度是关键的约束条件:大多数A/B测试不存在统计上可区分的优胜者,因此效度只能在可靠子集(n = 399)上进行测量。第二,与人设模拟的前提相反,人设条件化会降低预测效度:无人设基线对变体的排序明显更好(Kendall τ = 0.361,中等效应量;top-1准确率49.2%),优于人设小组(τ = 0.084;top-1准确率34.6%),且两者的置信区间不重叠。直接询问模型调用的是准确的人群层面先验知识;而强迫模型扮演特定人设则会引入偏差和噪声。该结果在Upworthy的三个独立数据划分上均得到复现,在不同领域的新闻数据集上方向一致,并且对随机种子、提示措辞和模型选择均保持稳健——涵盖三个Gemini层级以及另一个模型家族(OpenAI gpt-4.1,配对比较差异显著)。结论是:在预测总体用户参与度方面,简单的LLM排序器优于人设模拟——合成人设不仅预测能力弱,而且比不使用它们更差。所有数据均可通过公开的、以产物为先(artifact-first)的复现包重新生成。
cs.AI / 2 / 2609.25013

Do Existing Preconditioners Improve Biomedical Tabular Foundation Learning? An Empirical Study on TabPFN Optimization

Sajid, M., Khatun, Pinki, Tanveer, M.
Abstract
Tabular foundation models have recently shown strong potential for structured biomedical data analysis. Among them, TabPFN has emerged as an effective approach for low-data tabular classification tasks. However, the impact of optimization and preconditioning strategies on biomedical fine-tuning remains largely unexplored. In this work, we present a comprehensive empirical investigation of five AdamW-based preconditioning strategies for fine-tuning TabPFN v2.5 on 59 biomedical datasets spanning Alzheimer's disease, breast cancer, schizophrenia, significant memory concern (SMC), KEEL biomedical datasets, and UCI biomedical benchmarks. The evaluation considers predictive performance, computational efficiency, and statistical significance analysis. Experimental results demonstrate that the original AdamW optimizer consistently achieves the best overall performance and statistical ranking, while existing curvature-aware preconditioners fail to provide reliable improvements across diverse biomedical learning scenarios. The findings suggest that generic preconditioning approaches may not adequately capture the optimization characteristics of biomedical tabular learning, motivating the development of biomedical-aware preconditioners specifically tailored for healthcare-oriented tabular foundation models.
cs.AI / 3 / 2609.25036

4DGS-JEPA: Temporally Compositional Joint-Embedding Prediction for Dynamic Gaussian Splatting

4DGS-JEPA:面向动态高斯泼溅的时间可组合联合嵌入预测方法
Huang, Yongchao
Abstract
Dynamic Gaussian Splatting provides an explicit representation of evolving 3D scenes, but existing approaches are primarily optimized for reconstruction, future-state generation, or rendering rather than for learning reusable predictive dynamics. We propose 4DGS-JEPA, a Gaussian-native joint-embedding predictive architecture for causal multi-horizon prediction over dynamic Gaussian scenes. The model uses a hierarchical scene-, motion-group-, and Gaussian-level representation together with a horizon-conditioned transition operator that supports both direct prediction and recursive rollout. Its central principle is temporal composition: different chronological transition paths reaching the same future endpoint should produce compatible predictive states. Endpoint and multi-horizon path supervision anchor these predictions to future target embeddings, while a selective geometry decoder and geometry-level composition ground the learned dynamics in consistent group motion and Gaussian geometry without requiring complete future appearance reconstruction. We further introduce a hybrid correspondence mechanism that combines persistent canonical identity with residual optimal-transport matching under reordering and topology change. We characterize zero-loss path agreement and finite-error rollout accumulation theoretically. Three controlled experiments provide mechanism-level evidence that temporal composition reduces latent path dependence while retaining predictive accuracy, geometry-level composition improves consistency of decoded motion, and hybrid correspondence preserves reliable identity while remaining robust when correspondence becomes ambiguous. Together, 4DGS-JEPA provides a predictive, temporally compositional formulation of dynamic Gaussian worlds.
Chinese Translation
动态高斯泼溅(Dynamic Gaussian Splatting)为演化的三维场景提供了显式表示,但现有方法主要针对重建、未来状态生成或渲染进行优化,而非学习可复用的预测性动态。我们提出 4DGS-JEPA,一种面向动态高斯场景因果多时域预测的高斯原生联合嵌入预测架构。该模型采用场景级、运动组级和高斯级的层次化表示,并配合一个支持直接预测与递归展开的时域条件转移算子。其核心原则是时间可组合性:到达同一未来终点的不同时间顺序转移路径应产生相容的预测状态。终点与多时域路径监督将这些预测锚定于未来目标嵌入,而选择性几何解码器与几何层面的可组合性使所学习的动态建立在一致的组运动和高斯几何之上,且无需完整的未来外观重建。我们进一步引入一种混合对应机制,在重排序与拓扑变化的情况下,将持久的规范同一性与残差最优传输匹配相结合。我们从理论上刻画了零损失路径一致性与有限误差的递归展开累积。三个受控实验提供了机制层面的证据:时间可组合性在保持预测精度的同时降低了潜变量对路径的依赖;几何层面的可组合性提升了解码运动的一致性;混合对应机制在对应关系变得模糊时仍能保持可靠的同一性并具有鲁棒性。总之,4DGS-JEPA 为动态高斯世界提供了一种预测性的、时间可组合的形式化框架。
cs.AI / 4 / 2609.25088

An Accurate and Interpretable Hyper Graph Neural Network for GBM Survival Prediction

Intesum, Mushahid
Abstract
Survival prediction for glioblastoma multiforme (GBM) demands models that are both accurate and interpretable, yet existing approaches treat these objectives as com- peting, where performant models sacrifice transparency, while interpretable models accept degraded predictive power. We argue that this trade-off is not inherent. Graph neural net- works offer a structural foundation for extracting interpretable, explainable representations without compromising discriminative ability. Furthermore, current methods typically rely on a single imaging modality, underutilizing the complementary information available across multi-modal MRI and clinical metadata. We propose a multi-modal framework that inte- grates three components to address both objectives simultaneously: (1) a sheaf hypergraph neural network that captures higher-order relationships among tissue patches through direc- tional, asymmetric message passing; (2) a concept bottleneck layer that compresses learned representations into clinically grounded concepts, enforcing ante-hoc interpretability; and (3) an extension sufficiency test (EST) regularizer that penalizes unfaithful explanations during training, ensuring that model explanations genuinely reflect the internal decision process. Clinical and genomic features are incorporated through gated fusion, preserving the dominant prognostic signal of molecular markers while retaining concept-level traceabil- ity. Evaluated on 593 patients from the UPenn-GBM dataset under 5-fold cross-validation, our framework achieves a concordance index of 0.643 with the lowest fold-level variance among all compared models (std = 0.015). To our knowledge, this is the first work to unify sheaf hypergraph convolution, concept bottleneck supervision, and EST regularization for interpretable survival prediction from brain MRI
cs.AI / 5 / 2609.25165

Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings

Ovis-Embedding:推进通用全模态嵌入的前沿
Embedding Team
Abstract
In this report, we introduce \textbf{Ovis-Embedding}, a state-of-the-art omni-modal embedding family built on native integration of text, image, video, and audio. Instead of assembling separate modality towers, Ovis-Embedding uses a shared multimodal backbone to encode different modalities in a common representation space. Specifically, we make \textbf{three key advances}: (1) \textbf{native omni-modal initialization}: we adopt a pretrained Qwen-omni model as the embedding backbone and adapt it through contrastive training with low-rank initialization; (2) \textbf{data-centric omni-modal training}: we construct a broad, high-quality corpus spanning text, images, video, audio, and interleaved multimodal data. To improve data efficiency, we introduce homogeneous-source sampling to form task-consistent batches with informative in-batch negatives; and (3) \textbf{embedding-specific training and inference optimization}: we use focal loss to emphasize hard examples and similarity-based Embedding Distillation to transfer fine-grained similarity structure from complementary experts. At inference time, low-rank feature decomposition enables compact embeddings with flexible dimensionality and minimal performance loss. Empirical evaluations show that the \textbf{Ovis-Embedding} family achieves state-of-the-art performance on \textbf{MMEB-v3}, \textbf{MMEB-v2}, \textbf{MVEB}, \textbf{MAEB}, and \textbf{RTEB}, demonstrating its effectiveness across text, image, video, and audio modalities. These results highlight the potential of unified omni-modal training to overcome modality fragmentation and advance universal embedding models for any-to-any retrieval.
Chinese Translation
在本报告中,我们提出了 Ovis-Embedding,一个最先进的全模态嵌入模型家族,其原生地整合了文本、图像、视频和音频。Ovis-Embedding 并非组装独立的各模态编码塔,而是采用共享的多模态骨干网络,将不同模态编码到统一的表示空间中。具体而言,我们取得了三项关键进展:(1) 原生全模态初始化:我们采用预训练的 Qwen-omni 模型作为嵌入骨干网络,并通过基于低秩初始化的对比训练对其进行适配;(2) 以数据为中心的全模态训练:我们构建了一个涵盖文本、图像、视频、音频以及交错多模态数据的广泛且高质量的语料库。为提高数据效率,我们引入了同源采样方法,以构建任务一致的批次并提供有信息量的批内负样本;(3) 面向嵌入的训练与推理优化:我们使用焦点损失(focal loss)来强调难例,并采用基于相似度的嵌入蒸馏(Embedding Distillation)从互补专家模型中迁移细粒度的相似性结构。在推理阶段,低秩特征分解能够在性能损失极小的情况下实现维度灵活的紧凑嵌入。实证评估表明,Ovis-Embedding 家族在 MMEB-v3、MMEB-v2、MVEB、MAEB 和 RTEB 上均取得了最先进的性能,证明了其在文本、图像、视频和音频模态上的有效性。这些结果凸显了统一全模态训练在克服模态碎片化、推动面向任意到任意检索的通用嵌入模型发展方面的潜力。
cs.AI / 6 / 2609.25187

X-Planner: Event-Structured Task Planning for Embodied Intelligence

Lu, Howard, Li, Shalfun, Pan, Porter, Cris, Lumen, Cyril, Hu, Eric, Li, Lily, Zhang, Maeve, Wang, Robert, Zheng, KZ, Chen, Viggo, Ding, Tim, Cheng, Regsis, Xiao, YJ, Kian, Lin, Hai, Song, Alan, Ma, Elise, Li, Gody, Yao, Victor, Tang, Yohann, Yu, Ingrid, He, Jason, Wang, James, Yu, Ryan, Yang, Ping, Pan, Chris, Chen, Vincent, Gan, Roy, Wang, Hao, Wang, Qian
Abstract
Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems often leave this intermediate structure implicit. Existing chain-of-thought (CoT) planners also tend to rely on coarse task-level annotations or serialize long reasoning traces token by token. We present X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning. Our planning data combine Ego, UMI, and teleoperation under a hierarchy granularity with source-dependent annotation depth. Takeover-time annotations and human-designed failures supervise ongoing error recognition. On the model side, a shared VLM backbone exposes two event-structured plan forms: a discrete interface that emits interpretable event states and a latent interface that relays continuous CoT states across staggered Transformer depths through Staircase Decoding. A frozen latent-to-text reconstruction objective provides a semantic anchor for the latent representation. Offline two-step planning evaluation places X-Planner second among four evaluated models on both BERTScore-F1 and a judge-based Overall score. In real-robot experiments, respectively, outperforming the evaluated baselines. These results characterize planning-text quality and downstream execution.
cs.AI / 7 / 2609.25199

Lean Pool: An AI-Maintained Archive of Formalized Mathematics

Lean Pool:一个人工智能维护的形式化数学档案库
Ilin, Vasily
Abstract
Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.
Chinese Translation
Lean Pool 是一个形式化数学知识库,由人工智能体(AI agents)对其进行构建、维护和优化。
cs.AI / 8 / 2609.25254

The AI Neuroscientist: An Interactive Agentic Interface for Neuroimaging Analysis

Patel, Aakash, Ketonis, Panos, Saxena, Shreya, Krishnaswamy, Smita, van Dijk, David
Abstract
Analyzing neuroimaging data requires specialized coding and statistical expertise, which limits accessibility for researchers without computational backgrounds. We present the AI Neuroscientist, a language agent for interactive data exploration. The system integrates a large language model (LLM) with a neuroimaging toolset to perform quality control, modeling, and visualization. This allows researchers to query data quality and specify analysis parameters directly in natural language, providing a transparent and interactive alternative to conventional scripted pipelines for small-scale data exploration. We demonstrate these capabilities using functional near-infrared spectroscopy (fNIRS) data, and evaluate the agent on a custom fNIRS benchmarking suite against general-purpose LLM agents with code sandboxes. Future extensions will generalize the architecture to additional modalities, including functional magnetic resonance imaging (fMRI) data, and expand the benchmarking suite to additional fNIRS tasks.
cs.AI / 9 / 2609.25272

MedGate-Fusion: Integrating First-Encounter Semantic Narratives and Physiological Biomarkers for Prospective Stroke Risk Stratification

Khdr, Hemn, Noaeen, Mohammad, Keshavjee, Karim, Guergachi, Aziz, Shakeri, Zahra
Abstract
Prospective stroke risk stratification in primary care is challenging because early risk signals are distributed across routine biomarkers and unstructured clinical narratives. We propose MedGate-Fusion, a multi-modal gated architecture that integrates transformer-based embeddings of first-encounter narratives with ten routinely recorded risk markers. We used electronic medical record data from the Canadian Primary Care Sentinel Surveillance Network (CPCSSN). Starting from 808,921 encounter-level observations, we constructed a first-encounter cohort and retained 102,736 unique patient records with non-empty narratives and sufficient data to evaluate a five-year stroke outcome. To reduce explicit target leakage from diagnostic mentions in notes, we applied dictionary-based redaction of stroke-related terms prior to semantic encoding.
cs.AI / 10 / 2609.25284

When LLM Agents Fail to Read the Room: ReAdapt for Relational Social Reasoning

当LLM智能体无法察言观色:ReAdapt用于关系型社会推理
Lin, Jianzhe, Li, Xiaolin, Liu, Yunda, Wang, Fei, Chheda, Jubin
Abstract
A social agent's most basic decisions (should I react to this post? who should I reach out to?) are not purely content problems. The right action often hinges on the latent relationship between people -- tie strength, reciprocity, mutual connections -- rather than on which content is most salient. Standard LLM agent loops do not explicitly represent how new relational evidence should revise the agent's current social hypothesis, leaving them prone to surface-obvious choices when relational and content cues diverge. We formalize this failure mode with a relationship-reasoning benchmark: 500 synthetic social worlds with friendships, follows, reaction histories, and feeds, yielding 1,000 queries over two tasks, reaction selection and warm introduction (finding the best bridge to a target person). By construction, the surface-obvious candidate differs from the relationship-grounded oracle in about 53% of queries, forming an overturn subset where the agent must use relational evidence to revise an initially plausible choice. We propose ReAdapt (Relationship-Adaptive Agent with Policy-driven sTate), which augments the ReAct loop with an explicit structured social state z = (G, B, R, N, D) capturing goal, belief, relationship, norm, and disclosure. After each tool observation, ReAdapt runs a typed Adapt step that updates this state and emits a policy operation (continue, switch, abandon, or clarify) before choosing the next action. With Gemini-3-Flash on a stratified subset of n = 150 queries per task, ReAdapt improves warm-introduction accuracy from 37% to 51% (+14 points) and reaction-selection accuracy from 69% to 77% (+8 points). Oracle regret drops from 0.260 to 0.152 and from 0.095 to 0.053, respectively. Holding the model, tools, and environments fixed, these results suggest that explicit relational-state adaptation helps LLM agents turn retrieved social evidence into revised decisions.
Chinese Translation
社会智能体最基本的决策(我该回应这条帖子吗?我该联系谁?)并非纯粹的内容问题。正确的行动往往取决于人与人之间潜在的关系——关系强度、互惠性、共同联系——而非哪个内容最显眼。标准的LLM智能体循环没有显式地表示新的关系证据应如何修正智能体当前的社会假设,这使得当关系线索与内容线索不一致时,它们容易做出表面显而易见的选择。我们通过一个关系推理基准来形式化这一失效模式:500个包含好友关系、关注关系、互动历史和信息流的合成社会世界,在两个任务(回应选择和暖引荐,即寻找通往目标人物的最佳桥梁)上产生1,000个查询。通过构造,约53%的查询中表面显而易见的候选者与基于关系的最优解不同,构成一个“翻转”子集,智能体必须利用关系证据来修正最初看似合理的选择。我们提出ReAdapt(Relationship-Adaptive Agent with Policy-driven sTate),它在ReAct循环基础上增加了显式的结构化社会状态z = (G, B, R, N, D),以捕捉目标、信念、关系、规范和信息披露。在每次工具观察之后,ReAdapt运行一个带类型的Adapt步骤,更新该状态并发出策略操作(继续、切换、放弃或澄清),然后再选择下一步行动。使用Gemini-3-Flash在每任务n = 150个查询的分层子集上,ReAdapt将暖引荐准确率从37%提升至51%(+14个百分点),将回应选择准确率从69%提升至77%(+8个百分点)。最优解遗憾值分别从0.260降至0.152,从0.095降至0.053。在模型、工具和环境保持不变的情况下,这些结果表明,显式的关系状态适应有助于LLM智能体将检索到的社会证据转化为修正后的决策。
cs.AI / 11 / 2609.25285

Attention as a Routing Graph: Live Circuit Extraction from a Single Forward Pass

Manvi, Ash, Tajreen, Samreena
Abstract
Finding circuits in language models usually means running many careful interventions. We try something simpler: treat attention as a routing map from one forward pass, keep a small set of routes that point toward the answer, and ask whether those routes actually matter. They often do. On induction and IOI (tasks where the "right" circuit is already known), ablating our extracted edges hurts the model much more than ablating a random set of the same size. We evaluate n=100 prompts per cell on GPT-2 Small, GPT-2 Medium, and Pythia-410M, with paired gap tests and bootstrap confidence intervals. The extract step costs one forward; a head-by-head patch sweep costs about two orders of magnitude more. We are not claiming a complete circuit atlas. We are claiming a cheap sketch that carries real causal signal on known tasks, with clear failure modes when it does not. Code and evaluation artifacts are at https://github.com/Aquinf03/live-circuit-routing.
cs.AI / 12 / 2609.25286

Learned Enterprise Data Comprehension: Compression and Routing for Data Agents

Torres, Ethan, Mills, Eric
Abstract
Structured-data agents in enterprise settings must reason over complex data environments whose relevant evidence is distributed across schemas, relationships, policies, and recurring business roles. Modern agentic systems often address this burden through reusable markdown-style memory or skill files that preserve previously discovered information for later queries, reducing the need to rediscover the same structure repeatedly. This is useful, but it obscures a natural division of labor: agents are well suited to semantic reasoning, while learned systems are well suited to predicting and organizing recurring structure. We introduce latent equivalence learning to bridge this gap. The framework separates persistent task-relevant identities from their dataset-relative realizations. In our realization, supporting and opposing evidence shape support-realized Gaussian prototypes that learn how those identities are expressed in a particular data environment, while soft-membership profiles retain distinctions lost under a hard assignment. A separate learned query-prototype system represents recurring evidential requirements and maps them through a learned compatibility function into the same persistent identity structure. This identity-factorized, query-conditioned routing materializes the relevant dataset-specific evidence for downstream reasoning, allowing the agent to operate over an already organized evidential state rather than reconstructing cross-schema structure at every query. On the Data Agent Benchmark, spanning 54 queries across 12 heterogeneous datasets, our full implementation achieves 94.67% dataset-macro stratified Pass@1 over five complete trials and 258/270 successful raw query attempts, compared with 55.51% for the benchmark's Claude Opus 4.6 reference agent, ranking first among 40 leaderboard entries at submission.
cs.AI / 13 / 2609.25299

Making Agents More Consistent: Skills Should Form Habits for Repeat Tasks

Weber, Travis, Taneja, Rohit
Abstract
On repeated work, agents are inconsistent. We ran 42 tasks three times each and found that, depending on the model, 38% to 74% returned answers that did not agree. Consistency is what a buyer, an auditor, or a regulator requires, and agents do not have it. They are wasteful too: 95.3% to 97.2% of what an agent generates goes to re-deriving a plan the system already knows. We propose skill habit formation. An agent mines its own execution history for candidate skills, deterministic variants that compete against the incumbent rather than replacing it. A candidate declares the region of input space it claims, so the common case runs as a script and the rest falls through to reasoning. Four gates of ascending cost admit candidates; the central one tests a candidate's execution trace against a retained reference, within a tolerance measured from that reference's own run-to-run variability. On text-to-SQL, three of four reasoning arms reproduced their own output on 11 to 13 of 42 repeated questions and the fourth on 26 of 42, while a habit-formed variant reproduced on all 456 dispatches we repeated and was non-inferior to every arm it replaced (p<0.0001). It also used 14% to 56% fewer tokens, turning net positive after 7 to 53 reuses. We measured what this costs in accuracy. The guard admitted work it should have deferred on 2.6% of natural paraphrases and 26% of inputs near its boundary, and 11 of 13 such failures were invisible to the trace-conformance gate at any threshold. Deterministic errors repeat exactly: a bad habit is as reliable as a good one, and that is the price of the property that makes the system auditable. Separating routing from parameter extraction raised end-to-end accuracy from 0.888 to 0.952 at 43% of the cost.
cs.AI / 14 / 2609.25303

Potential for Enhanced Learning in Machine Learning Classes by Using Wiki LLM Indexing

Wright, Brian
Abstract
Large language models are increasingly deployed as course-specific tutors, but their usefulness depends on grounding in vetted instructional materials that are often revised mid-semester. Our prior work built a multimodal retrieval-augmented generation (RAG) system over an authentic machine learning course corpus (Foundations of Machine Learning) and found that retrieval improved contextual grounding, but that fixed retrieval strategies were suboptimal. That motivates a different question: whether how a corpus is structured at ingest time matters more than how much is retrieved at query time. We present a controlled head-to-head comparison of two knowledge representations over an identical classroom corpus: (A) vector RAG, replicating the best-performing configuration from our prior study, and (B) an LLM-compiled wiki (Karpathy framework), in which the corpus is synthesized at ingest into linked concept pages with explicit cross-references and citations back to source materials. We evaluate 59 questions spanning single-fact recall, cross-unit concept linking, synthesis and explanation, and currency after a syllabus revision, scored by an LLM judge against a human-authored rubric. Both representations answered single-fact questions about equally well (9.33 vs. 9.96 of 10), but diverged sharply on questions requiring links across course units. The compiled wiki remained accurate and grounded (9.93; 100% grounded in cited sources), while retrieval scored lower and was markedly less grounded (8.14; 64%). The wiki's citations let students and instructors trace any claim back to the lecture that introduced it, adding a layer of dynamic retrieval that machine learning courses require. While further testing is needed, instructors using AI to support learning in ML courses should consider wiki-based structure for its potential to support foundational elements of best practice.
cs.AI / 15 / 2609.25337

Clarification Is Not Correction: LLMs Fail to Let Go

Lin, Jianzhe, Li, Xiaolin, Wang, Fei, Douglas, Robert, Golani, Rajeshkumar, Chheda, Jubin
Abstract
Dialogue failures in language models are usually framed as memory failures: context too long, summaries lossy, a constraint forgotten. We argue this misses a deeper problem: in many conversations the model does not forget, it commits too early. An ambiguous early turn collapses into a single hidden interpretation, and later clarification is filtered through that commitment. We call this early posterior collapse: unresolved user intent collapsing into a committed task state before ambiguity is resolved. We study it with controlled dialogue tasks in writing, planning, and coding using Gemini-2.5-Pro and Gemini-2.5-Flash. Across thousands of trials, the same information in different orders yields different outcomes, even when the final dialogue contains equivalent task-relevant information. This order effect suggests later clarification is treated as extra context rather than a corrective signal: it refines a stale task state without invalidating it. Coding tasks are especially vulnerable, suggesting early assumptions get embedded in structured artifacts such as interfaces and control flow. Standard prompting and memory strategies do not reliably help: summaries can collapse ambiguity, and chain-of-thought can reduce explicit wrong commitment in reasoning traces without improving final task success. These findings motivate uncertainty-preserving state management. If assistants cannot let go of early interpretations, robustness cannot rely on post hoc correction alone; it must keep ambiguous early turns from hardening into one task state. Assistants should hold tentative hypotheses while ambiguity remains, ask before executing when high-impact ambiguity persists, and rebuild from a revised state when later evidence invalidates an earlier reading. Rather than one prompting fix, we aim to redirect research for interactive LLMs from retaining more context toward preserving uncertainty.
cs.AI / 16 / 2609.25366

From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought

Jia, Renee, Mu, Di
Abstract
Chain-of-thought (CoT) monitoring is only meaningful if written reasoning causally constrains the answer. We introduce continuation-based causal testing, an ablation-patch intervention that perturbs one reasoning step, truncates the chain, and forces the model to continue from the corrupted prefix. It measures how load-bearing a CoT is for the final answer, a behavioral notion distinct from mechanistic faithfulness. Across Gemma-2-9B-IT, Llama-3.1-8B-Instruct, and DeepSeek-R1-Distill-Qwen-7B on GSM8K, MMLU, and BIG-Bench Hard, CoT load-bearingness tracks model-relative task difficulty: on easy tasks models silently bypass their own reasoning; on hard tasks they follow corrupted steps and propagate errors. A matched 2x2 analysis shows task difficulty dominates perturbation type: error propagation rises 16x from GSM8K to BBH multistep arithmetic, and a variance partition over 28,584 continuations attributes 98.8% of explained deviance to task difficulty versus 0.8% to perturbation type. Reasoning-specific RL suppresses error propagation and compresses the gradient. A four-variant judge-sensitivity analysis and blind two-annotator study (n=500) show the error-propagation vs. non-propagation label is invariant to judge prompt, with perfect inter-annotator agreement (Cohen's kappa = 1.00). This gradient creates a structural problem for CoT-based oversight and AI safety monitoring: where the trace is easy to read it carries little signal, and where it matters errors propagate before a monitor can intervene. Linear probes on hidden states separate silent bypass, self-correction, and error propagation, but additive activation steering provides limited causal control, flipping only about 25% of error-propagation cases at best. Behavioral mode is readable but not reliably controllable.
cs.AI / 17 / 2609.25400

Robust Failure, Conservative Repair: Textual Knowledge Distillation from Cross-Model Failures

稳健的失败、保守的修复:基于跨模型失败的文本知识蒸馏
Ren, Andrew, Liu, Haokun, Tan, Chenhao
Abstract
Failure-based textual knowledge distillation aims to discover gaps in a model's knowledge by examining its task errors. The distilled knowledge can be useful for the reasoning of both this model ("source model") and other models. However, this transfer of knowledge may not be stable. We define a rule atom to be a standalone rule injected into a model's textual input at inference time. A rule atom can encode transferable task knowledge or model-specific reasoning patches that can confuse other models. Also, the injected rule atoms can be misapplied to unrelated cases, causing the model to incorrectly flip its answer based on irrelevant information. Building on a pipeline that distills training examples into task-specific cheat sheets that aid model reasoning, we examine when failure-derived rules can improve these cheat sheets. Our early experiment shows rule distillation from a single model's failures underperforms the baseline cheat sheet on non-source model families. This motivates Robust Failure, Conservative Repair (RFCR), a textual distillation procedure that derives rules from failures shared across models, sharpens their application boundaries using boundary cases, and abstains when no useful rule is found. On a 400-item BIG-Bench Hard task set, RFCR improves the baseline cheat sheets from 68.50% to 71.25% (+2.75 pp; 95% CI [+1.25,+4.50]) without performance degradation on previously correct cases. Ablations and cross-model diagnostics support that accuracy gains come from both new knowledge injection and strict rule-application control.
Chinese Translation
基于失败的文本知识蒸馏旨在通过检查模型的任务错误来发现其知识中的缺口。所蒸馏的知识既可用于该模型("源模型"),也可用于其他模型的推理。然而,这种知识迁移可能并不稳定。我们将规则原子(rule atom)定义为在推理时注入模型文本输入的独立规则。规则原子既可以编码可迁移的任务知识,也可能编码特定于模型的推理补丁,而后者会干扰其他模型。此外,被注入的规则原子可能被误用于无关案例,导致模型基于无关信息错误地翻转其答案。基于一个将训练样例蒸馏为辅助模型推理的任务专用速查表(cheat sheet)的流程,我们研究了源自失败的规则何时能够改进这些速查表。我们的初步实验表明,从单个模型的失败中蒸馏规则,在非源模型家族上表现低于基线速查表。这促使我们提出稳健失败、保守修复(Robust Failure, Conservative Repair,RFCR)方法——一种文本蒸馏流程,它从多个模型共同出现的失败中提取规则,利用边界案例精确化规则的适用边界,并在未找到有用规则时选择放弃。在一个包含400个样本的BIG-Bench Hard任务集上,RFCR将基线速查表的准确率从68.50%提升至71.25%(+2.75个百分点;95%置信区间[+1.25, +4.50]),且在先前已正确回答的案例上没有性能下降。消融实验和跨模型诊断支持以下结论:准确率的提升既来自新知识的注入,也来自严格的规则应用控制。
cs.AI / 18 / 2609.25405

Efficient Iterative Retrieval with Heterogeneous Batching

基于异构批处理的高效迭代检索
Park, Dohyun, Franke, Hubertus, Waddington, Daniel G., Sundararaman, Swaminathan, Park, Yongjoo
Abstract
Modern information retrieval increasingly employs both embedding and generative models to handle complex queries. However, current serving systems suffer from low throughput and poor GPU utilization because they execute these models in isolation. Coarse-grained partitioning, such as dedicating GPUs to specific tasks, fails to adapt to dynamic workloads and creates computational "bubbles". To address these, we present Orthrus, a serving system that performs heterogeneous batching within a unified inference loop. The primary challenge lies in unifying embedding and generation workloads with conflicting computational patterns while optimizing batch composition for high performance. Orthrus addresses these challenges through chunked embedding with incremental pooling and by adjusting batch composition in a workload-aware manner. Evaluation on four A100 GPUs shows that, relative to baseline deployments, Orthrus achieves 1.28$\times$--4.52$\times$ higher throughput on controlled workloads and up to 55.8% lower end-to-end p99 latency on an iterative-RAG benchmark. We release our code at https://github.com/illinoisdata/Orthrus .
Chinese Translation
现代信息检索日益采用嵌入(embedding)模型与生成模型相结合的方式来处理复杂查询。然而,由于现有服务系统将这些模型孤立地执行,导致吞吐量低、GPU 利用率差。粗粒度的任务划分方式(例如将 GPU 专用于特定任务)无法适应动态工作负载,并会产生计算"气泡"。为解决这些问题,我们提出了 Orthrus,一个在统一推理循环内执行异构批处理的服务系统。其主要挑战在于如何统一具有相互冲突计算模式的嵌入与生成工作负载,同时优化批处理组成以实现高性能。Orthrus 通过带增量池化(incremental pooling)的分块嵌入(chunked embedding),并以负载感知的方式调整批处理组成,来应对这些挑战。在四块 A100 GPU 上的评估表明,相对于基线部署,Orthrus 在受控工作负载上实现了 1.28 倍至 4.52 倍的吞吐量提升,并在迭代式 RAG(iterative-RAG)基准测试中将端到端 p99 延迟最多降低 55.8%。我们的代码已在 https://github.com/illinoisdata/Orthrus 开源。
cs.AI / 19 / 2609.25408

From Offline Proxies to Online Decisions: A Layered Engagement Evaluation Framework for Conversational AI

从离线代理到在线决策:一个面向对话式AI的分层参与度评估框架
Li, Xuanyi, Nath, Vaskar, Amirkhani, Hossein, Li, Jay, Deng, Alex
Abstract
Online A/B experiments are the decision standard for user engagement, but traffic and readout time limit how many conversational-AI changes can be tested. We ask whether an offline signal designed to be computable without treatment-arm user exposure agrees with the outcomes of those experiments. We contribute a reusable construction and diagnosis checklist that treats an offline proxy as a chain of three alignments: behavioral label to product outcome, learned classifier to candidate-assistant behavior, and aggregated offline signal to experiment effect. A companion evaluation protocol audits the whole composite by interval-aware decision agreement, which compares offline and online confidence intervals instead of point estimates, and by within-experiment ranking. The instantiation we evaluate comprises a fixed evaluation suite on which candidate behavior is scored, an engagement classifier trained to predict session/prompt level engagements, and a calibration layer mapping sample-level score differences to online model-level engagement deltas. We then report the audit: 489 paired offline-online contrasts (one candidate arm against its control) from 27 experiments on a deployed multi-turn assistant, spanning model checkpoints to system-prompt tuning. Our primary test uses the 113 contrasts from eight experiments that ran after the map was frozen: on these the composite reaches 81.1% F1, against 34.3% for the raw classifier score it is built on, and makes no wrong-direction calls where that raw score makes 31. Every offline prediction was computed before its experiment ran to prevent overfitting. The evidence supports using the composite to prioritize candidates before scarce experiment traffic is allocated---in our deployment of the experiment, selecting among training checkpoints and tuning system prompts.
Chinese Translation
在线A/B实验是衡量用户参与度的决策标准,但流量和读取时间限制了可测试的对话式AI改动数量。我们探讨这样一种离线信号——其设计使其无需处理组用户暴露即可计算——是否与这些实验的结果相一致。我们贡献了一套可复用的构建与诊断清单,将离线代理视为由三个对齐环节构成的链条:行为标签与产品结果的对齐、学习到的分类器与候选助手行为的对齐、以及聚合的离线信号与实验效应的对齐。配套的评估协议从两个维度审计整个复合体系:一是区间感知的决策一致性,即比较离线与在线的置信区间而非点估计;二是实验内的排序表现。我们评估的具体实现包括:一个固定的评估套件,用于对候选行为进行打分;一个经过训练、用于预测会话/提示词级参与度的参与度分类器;以及一个校准层,将样本级的分数差异映射为在线的模型级参与度增量。随后我们报告审计结果:来自某个已部署多轮助手的27个实验、共489组离线-在线配对对比(每个候选组与其对照组对比),涵盖模型检查点到系统提示词调优等多个方面。我们的主要测试使用了地图冻结后运行的8个实验的113组对比:在这些对比上,该复合体系达到81.1%的F1值,而其依赖的原始分类器分数仅为34.3%;且在原始分数产生31次错误方向判断的场景中,复合体系没有产生任何错误方向的判断。所有离线预测均在相应实验运行之前计算完成,以防止过拟合。这些证据支持在稀缺的实验流量分配之前,使用该复合体系对候选方案进行优先级排序——在我们的实验部署中,即在训练检查点选择和系统提示词调优中进行筛选。
cs.AI / 20 / 2609.25443

ZeroGate: Trust-Preserving Fast Paths for Governed AI Agent Runtimes

ZeroGate:面向受治理AI智能体运行时的信任保持快速通道
Wang, Zexun
Abstract
Moving authorization earlier can shorten an agent's dispatch boundary without removing authorization work. It can also admit an action whose payload, authority, or relevant state has changed. ZeroGate separates exact-action approval from durable local admission: an issuer signs a short-lived ActionPass, and a trusted runtime adapter reconstructs the final action before a local gate checks its binding and consumes its nonce. A SQLite transaction couples nonce consumption, applicable quota updates, and an admission receipt. We state a conditional decision-preservation proposition: successful local admission implies that a specified synchronous policy would authorize the same action at the admission point, provided approval is sound, all policy dependencies are represented and current, observations are faithful, and consumption is atomic. The implementation alone establishes neither current-world freshness nor exactly-once remote effects. Evaluation separates authored semantic fixtures, controlled concurrency and crash experiments, and an Azure Blob study comparing synchronous and prepared execution through the same issuer and gate. Both modes mint an exact-action pass; lifecycle latency includes preparation and prepared-batch dwell. Across 4800 cloud attempts, prepared worker-admission-to-dispatch p95 ranges from 9.802 to 11.374 ms, versus 25.018 to 334.000 ms synchronously, across the tested concurrency levels. Prepared mean complete lifecycle is longer at every level: the boundary improvement is not a net speedup. The contribution is an explicit revalidation contract, a durable reference boundary, and an auditable comparison of where authorization cost is paid, not a new cryptographic primitive or a universal performance frontier.
Chinese Translation
将授权环节前移可以缩短智能体的调度边界,但不能消除授权工作本身,同时也可能准入其载荷、权限或相关状态已发生变化的操作。ZeroGate 将精确操作审批与持久化本地准入相分离:签发者签署一个短时效的 ActionPass,可信运行时适配器在最终执行操作之前重建该操作,随后本地门控检查其绑定并消费其 nonce。一个 SQLite 事务将 nonce 消费、适用的配额更新以及准入回执耦合在一起。我们陈述了一个条件性决策保持命题:在审批可靠、所有策略依赖均被表示且为最新、观测忠实、且消费具有原子性的前提下,本地准入成功即意味着某个指定的同步策略会在准入点授权同一操作。仅有实现本身并不能确立现实世界的最新性或远程效果的恰好一次性(exactly-once)。评估分为三部分:人工编写的语义测试用例、受控的并发与崩溃实验,以及通过相同签发者和门控对同步执行与预备执行进行比较的 Azure Blob 研究。两种模式均铸造精确操作通行证;生命周期延迟包括准备时间和预备批次的驻留时间。在 4800 次云端尝试中,在所测试的各并发级别下,预备模式下从工作准入到调度的 p95 延迟为 9.802 至 11.374 毫秒,而同步模式为 25.018 至 334.000 毫秒。但在每一并发级别下,预备模式的平均完整生命周期都更长:边界改进并非净加速。本工作的贡献在于一个显式的重验证契约、一个持久的参考边界,以及对授权成本在何处支付的可审计比较,而非一种新的密码学原语或普遍的性能前沿。
cs.AI / 21 / 2609.25463

Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future Directions

Gholipour, Niloofar, Assuncao, Marcos, Singh, Gursimran, Yu, Timothy, Buyya, Rajkumar, Gascon-Samson, Julien, Fan, Zhenan, Zhang, Yong, Xu, Xiaojie, Yao, Yaqiang, Bai, Xiaolong
Abstract
Reasoning-oriented reinforcement learning enables large language models to solve mathematical, coding, and other multi-step tasks, but shifts a substantial portion of the training cost to rollout, where trajectories are generated for policy updates. Efficient rollout mechanisms are therefore essential to reduce this cost while maintaining the freshness, consistency, and statistical validity of training data. This survey provides a systematic taxonomy of recent research on rollout efficiency for reasoning-oriented reinforcement learning, classifying existing approaches from both mechanism and bottleneck perspectives. Based on this taxonomy, we analyze how different technique families address distinct sources of rollout inefficiency, examine opportunities and potential conflicts for combining them, identify gaps in the evaluation and reporting of efficiency gains, and discuss open challenges and future research directions.
cs.AI / 22 / 2609.25466

Real-Time Hand Gesture Recognition for OpenXR Using Transformer-Based Machine Learning

基于Transformer机器学习的OpenXR实时手势识别
Rezayani, Salar, Butler, Russell
Abstract
Hand gesture recognition is a key component in human-computer interaction (HCI), enabling intuitive interfaces for applications in gaming, virtual reality (VR), robotics, and more. This study integrates transformer-based machine-learning models for real-time hand gesture recognition, using hand-tracking data captured through the OpenXR standard in Unity. We leverage positional data of hand joints and wrist rotation angles to train a custom gesture recognition system. By utilizing the sequential modeling capabilities of transformers, the system captures temporal dependencies within short gesture windows and classifies gestures robustly across hand orientations and sizes. The results show a significant improvement in gesture classification accuracy. Building on this, we outline how the approach can be extended toward detecting the flow of movement, i.e., the transitions between gestures, as future work.
Chinese Translation
手势识别是人机交互(HCI)的关键组成部分,为游戏、虚拟现实(VR)、机器人技术等应用提供了直观的交互界面。本研究集成了基于Transformer的机器学习模型,用于实时手势识别,使用通过Unity中OpenXR标准捕获的手部追踪数据。我们利用手部关节的位置数据和手腕旋转角度来训练一个自定义的手势识别系统。借助Transformer的序列建模能力,该系统能够捕捉短时手势窗口内的时序依赖关系,并在不同手部朝向和尺寸条件下稳健地对手势进行分类。结果显示手势分类准确率有显著提升。在此基础上,我们概述了将该方法扩展用于检测动作流(即手势之间的过渡)的未来工作方向。
cs.AI / 23 / 2609.25467

ShowTellArena: Evaluating Business Workflow Understanding from Demonstrations

ShowTellArena:从演示中评估对业务工作流的理解能力
Garg, David, Sarkar, Ritobrata, Azarnasab, Ehsan, Borah, Siddhartha
Abstract
We often teach a colleague by showing the work and explaining the decisions as we go. How can we check what an agent understood from the same lesson? We introduce ShowTellArena, a benchmark protocol and public dataset for comprehension after narrated business demonstrations. The v1.0 release contains 50 business workflow tasks, with recordings, screenshots, narration, fixture seeds, and 502 questions. Tasks span finance, hiring, procurement, customer decisions, inventory, and logistics. The protocol holds the business scenario and quiz fixed while allowing each product to capture the lesson through its own teaching interface. Questions test operational rules, boundaries, exceptions, and errors in proposed automations. We analyze 218 selected pilot attempts across 39 workflow cases, including 28 cases attempted by all three evaluated systems. These exploratory results expose both answer errors and failures to complete the teaching experience. We describe the release's verification gaps and the pilot's uneven coverage, exclusions, and grading provenance. The contribution is an inspectable dataset and assessment workflow that others can extend; the selected pilot is not a controlled product ranking.
Chinese Translation
我们常常通过边展示工作边解释决策来教导同事。我们如何检验智能体从同样的课程中学到了什么?我们提出了ShowTellArena,一个用于评估带旁白业务演示后理解能力的基准协议和公开数据集。v1.0版本包含50个业务工作流任务,提供录屏、截图、旁白、测试环境种子(fixture seeds)以及502个问题。任务涵盖财务、招聘、采购、客户决策、库存和物流等领域。该协议在保持业务场景和测验内容固定的同时,允许每个产品通过其自身的教学界面来学习课程内容。问题用于测试操作规则、边界条件、异常情况以及所提出自动化方案中的错误。我们分析了来自39个工作流案例的218次精选试点尝试,其中包括由三个受评估系统均尝试过的28个案例。这些探索性结果既暴露了答案错误,也暴露了未能完成教学体验的情况。我们描述了该版本存在的验证缺口以及试点中覆盖不均、排除项和评分来源等问题。本文的贡献是一个可供他人扩展的可审查数据集与评估工作流;所精选的试点结果并非受控的产品排名。
cs.AI / 24 / 2609.25469

RAG-NAROK: Retrieval-Aware Knowledge Corpus Poisoning in RAG with Source-specific Refutation

RAG-NAROK:基于源特定驳斥的检索增强生成(RAG)系统中的检索感知知识语料库投毒攻击
Kafi, Abdullahil, Khalil, Alvi Ataur
Abstract
Retrieval augmented generation (RAG) systems have emerged as the dominant architecture for grounding large language model (LLM) outputs in verifiable external knowledge, yet their structural reliance on a dynamic retrieval pipeline introduces a largely unexplored class of adversarial vulnerability. Existing knowledge-base poisoning attacks are fundamentally static. Adversarial documents are pre-computed and injected without any awareness of what the victim system will actually retrieve for a given query, leaving the attack blind to the competitive documentary landscape that surrounds its payload in the generator's context window. Unlike traditional static poisoning attacks that are blind to the retrieved context, we introduce RAG-NAROK (Retrieval-Anchored Generation Negation And Response Quality Collapse), a RAG attack framework that adapts to the query text. RAG-NAROK exploits the transparency inherent in RAG pipeline to first extract the legitimate source identities, then generate Anchor-Specific Refutation documents that explicitly name and devalue retrieved sources while leveraging recency and authority biases to steer the text generation toward a target answer. Our results demonstrate that RAG-NAROK significantly outperforms static baselines across diverse domains, revealing a fundamental tension between RAG transparency and AI security.
Chinese Translation
检索增强生成(RAG)系统已成为将大语言模型(LLM)输出锚定于可验证外部知识的主流架构,然而其对动态检索流程的结构性依赖引入了一类 largely 未被探索的对抗性脆弱性。现有的知识库投毒攻击本质上是静态的:对抗性文档被预先计算并注入,而对受害系统针对给定查询实际检索到的内容一无所知,这使得攻击对其载荷在生成器上下文窗口周围所处的竞争性文档环境完全“失明”。与这种对检索上下文失明的传统静态投毒攻击不同,我们提出了 RAG-NAROK(Retrieval-Anchored Generation Negation And Response Quality Collapse,检索锚定生成否定与响应质量崩塌),一种能够自适应查询文本的 RAG 攻击框架。RAG-NAROK 利用 RAG 流程固有的透明性,首先提取合法来源的身份信息,然后生成“锚定特定驳斥”(Anchor-Specific Refutation)文档,这些文档明确点名并贬低被检索到的来源,同时利用时效性与权威性偏差,将文本生成引导至目标答案。实验结果表明,RAG-NAROK 在多个不同领域均显著优于静态基线方法,揭示了 RAG 透明性与 AI 安全之间存在根本性矛盾。
cs.AI / 25 / 2609.25474

Spectra: A Rules-Driven LLM Pipeline for Automated KYC Document Processing

Wahib, Miray, Tran, Ethan, Mourad, Rea, Muti, Mira, Dvornik, Nikita
Abstract
Know Your Client (KYC) onboarding in capital markets requires analysts to manually classify documents, extract structured data from heterogeneous sources, and validate compliance against complex regulatory policies. This process requires significant analyst time per client, with end-to-end onboarding often stretching to multiple weeks due to sequential handoffs. In this work, we analyze an on-boarding process and find that it comprises repeatable components well-suited to AI automation. We therefore propose a restructured workflow to be amenable to automation: we consolidate the traditional four-party process into two parties that share most of the work and can be automated together, eliminating intermediate handoffs that compound delays. To automate the remaining steps, we introduce Spectra, an AI-assisted document processing platform that combines a structured rules engine with LLM-based classification, extraction, and validation agents. The rules engine encodes compliance policy as a queryable database, enabling focused context injection that reduces token usage while improving extraction precision. Rather than a single monolithic prompt, the system decomposes document processing into isolated, auditable stages, each optimized independently and traceable to specific policy clauses. In evaluation on real KYC documents, Spectra achieves 100% classification accuracy and 89.4% extraction accuracy. Human review burden dropped by 96%.
cs.AI / 26 / 2609.25491

Queer inclusion in speech datasets: An audit and taxonomy of practical tensions

语音数据集中的酷儿包容性:对实际张力的审查与分类
Sheppard, Brooklyn, Ovalle, Anaelia, Williams, Adina, Sagun, Levent
Abstract
In this paper, we examine speech datasets for their inclusion of LGBTQIA+, or queer, voices and provide a taxonomy of tensions to better understand why there is a lack of such voices in current speech technology datasets. Through an audit of six diverse speech datasets, we find that measurable queer representation is low (0-1.4% of speakers) - insufficient for robust disparity measurement. We take this community as a case study to consider what challenges and tensions are associated with collecting speech data from marginalized communities. For comparison, we audit an additional two datasets from the speech sciences that were created by, for, and with the queer community. We note that many customs in speech dataset collection efforts in AI and speech technology research may conflict with values emphasized in participatory approaches with marginalized communities, and provide a taxonomy describing these tensions.
Chinese Translation
本文审视了语音数据集对LGBTQIA+(即酷儿)群体的声音的包容程度,并提出了一种张力的分类体系,以更好地理解为何当前语音技术数据集中缺乏这类群体的声音。通过对六个多样化语音数据集的审查,我们发现可测量的酷儿群体代表性很低(占说话者的0-1.4%)——不足以支持稳健的差异性测量。我们以该群体为案例,探讨从边缘化群体收集语音数据所涉及的挑战与张力。作为对比,我们还审查了另外两个由语音科学领域创建的、由酷儿群体创建、为其服务并与其共同合作开发的数据集。我们注意到,人工智能与语音技术研究中语音数据集收集工作中的许多惯例,可能与与边缘化群体开展参与式方法时所强调的价值观相冲突,并据此提出了描述这些张力的分类体系。
cs.AI / 27 / 2609.25496

Towards participatory speech dataset curation: A queer case study and conceptual framework

Sheppard, Brooklyn, Ovalle, Anaelia, Williams, Adina, Sagun, Levent
Abstract
In this paper, we motivate the need for a participatory speech dataset creation framework through a case study of the LGBTQIA+, or queer, community - a community with documented concerns about AI and reported harms, including attempts to develop 'gaydar' technologies that purportedly identify individuals as queer. We review common speech data collection practices, why these methods may be unsuitable for engaging with queer speakers, and discuss previous efforts in participatory AI with queer community engagement, as well as participatory endeavours specific to speech data collection for other marginalized communities. From this review, we develop a conceptual framework for participatory speech data curation by, for, and with marginalized communities drawing on insights from co-design and knowledge sharing. We propose a framework comprising overlapping and two-way processes of defining a community, project formulation, modes of participation, and personal autonomy.
cs.AI / 28 / 2609.25508

SMTB: Fast Structure-Mapping with Tight Bounds

Weitekamp, Daniel, MacLellan, Christopher
Abstract
Structure-mapping forms analogies by aligning systems of relationally connected elements based on shared structure instead of surface features. We introduce a new structure-mapping algorithm: Structure-Mapping with Tight Bounds (SMTB) that is 5--15x faster than the structure-mapping engine (SME) and about 50\% better at finding mappings in large nested domains. SMTB is part of the broader Cognitive Rule Engine (CRE) project, a flexible multi-language-compatible framework with an accessible Python interface to state-of-the-art C++ implementations of core algorithms commonly used in cognitive systems such as pattern matching, planning, and structure-mapping. CRE and SMTB are designed to work with a wide range of representation choices. Unlike SME, which biases higher-order correspondences in tree-like predicate logic, SMTB maximizes relational connectivity without privileging higher-order relations. This allows SMTB to work just as well over arbitrary relational graphs as it does in tree-like domains of nested predicate logic. We discuss situations where privileging "higher-orderness" in structure-mapping can cause issues, and illustrate how SMTB avoids failure modes that SME would encounter in these situations. We also provide an evaluation comparing SMTB to SME v4 over 5845 domain pairs from the SME corpus.
cs.AI / 29 / 2609.25555

Weakly Supervised Quantum Error Mitigation

弱监督量子误差缓解
Tousi, Seyed Mohamad Ali, DeSouza, G. N.
Abstract
Supervised approaches to quantum error mitigation learn a map from noisy circuit outputs to ideal ones, and therefore require the ideal outputs. Producing those ideal outputs demands noiseless classical simulation, whose cost grows exponentially with system size, so supervision is unavailable in exactly the regime where mitigation matters most. We ask whether cheap, individually unreliable signals drawn from circuit structure and hardware calibration can take the place of ideal labels. We assemble sixteen heuristic labeling functions (stabilizer and parity constraints, relaxation and readout characteristics, local depth, gate counts, and neighboring activity), reconcile their disagreements with a probabilistic label model, and read the resulting per-qubit error probabilities as a readout channel whose inverse mitigates the measured distribution. No ideal output enters the training path. On $147{,}000$ five-qubit circuits executed on two IBM devices, the method removes $24.3\%$ (Algiers) and $28.8\%$ (Hanoi) of the Kullback-Leibler divergence to the ideal distribution, against $15.4\%$ and $21.5\%$ for the strongest published analytical baseline, a margin that holds on both devices and lies far outside its bootstrap interval. Supervised neural models trained on ideal distributions remain stronger where such labels exist, and we quantify that gap rather than setting it aside; the method's claim is to the regime where they do not, since the labels they require cannot be computed for the circuits mitigation is needed for. The codes will be released shortly.
Chinese Translation
监督式量子误差缓解方法学习从含噪声电路输出到理想输出的映射,因此需要理想输出。生成这些理想输出需要进行无噪声的经典模拟,其代价随系统规模呈指数级增长,这使得在误差缓解最关键的场合恰恰无法获得监督信号。我们研究的问题是:能否用从电路结构和硬件校准中获取的廉价且单独并不可靠的信号来替代理想标签。我们构建了十六个启发式标注函数(稳定子和奇偶性约束、弛豫与读出特性、局部深度、门数量以及邻近比特活动),并利用概率标签模型对其分歧进行调和,随后将由此得到的每量子比特错误概率视为一个读出信道,通过其逆操作对测量分布进行误差缓解。整个训练路径中不使用任何理想输出。在两台IBM设备上执行的147,000个五量子比特电路中,该方法将对理想分布的Kullback-Leibler散度降低了24.3%(Algiers)和28.8%(Hanoi),而目前最强的已发表解析基线方法仅能降低15.4%和21.5%,这一优势在两台设备上均成立,且远超其自助法置信区间。在理想分布标签可用的情况下,基于监督的神经模型仍然更强大,我们对这一差距进行了量化而非回避;本方法的价值主张针对的是标签不可得的场景——因为在最需要误差缓解的电路中,监督方法所需的理想标签根本无法计算。相关代码将很快发布。
cs.AI / 30 / 2609.25570

Recovering Agentic Sovereignty: Mitigating the Consensus Paradox via Contrastive Epistemic Decoding

恢复智能体自主性:通过对比认知解码缓解共识悖论
Shehata, Dahlia, Li, Ming
Abstract
Large language models (LLMs) exhibit a parametric vulnerability to adversarial swarm consensus. To mitigate this sycophancy, we introduce Contrastive Epistemic Decoding (CED), a zero-shot inference intervention. Unlike standard Contrastive Decoding (CD) which relies on a weaker secondary model, CED utilizes a dual forward-pass on a single architecture to isolate conformity bias. By introducing a novel asymmetric, zero-bounded probability clamp and discrete top-k truncation mask, CED mathematically suppresses toxic consensus tokens without causing grammatical collapse. Evaluated across 7,200 paired trajectories on complex benchmarks (GAIA, SWE-bench, Multi-Challenge) using Gemma-2 (9B), Llama-3.1 (8B), and Mistral v0.3 (7B), CED successfully neutralizes architectural and positional biases. By reducing cognitive loafing by up to 33.00% absolute, CED drives significant performance gains, yielding up to a 30.75% accuracy recovery. Regaining sovereignty induces distinct architectural behaviors---passive task-focus in Gemma-2 and active refutation of the simulated swarm in Llama-3.1---showing CED decouples compliance from capability without fine-tuning.
Chinese Translation
大型语言模型(LLM)在面对对抗性群体共识时表现出参数层面的脆弱性。为缓解这种谄媚倾向,我们提出了对比认知解码(Contrastive Epistemic Decoding, CED),这是一种零样本推理干预方法。与依赖较弱辅助模型的标准对比解码(Contrastive Decoding, CD)不同,CED 利用单一架构的双重前向传播来隔离从众偏差。通过引入一种新颖的非对称、零值有界的概率钳制机制以及离散 top-k 截断掩码,CED 在数学上抑制了有害的共识词元,同时避免了语法崩塌。在使用 Gemma-2(9B)、Llama-3.1(8B)和 Mistral v0.3(7B)于复杂基准(GAIA、SWE-bench、Multi-Challenge)上进行的 7,200 条成对轨迹评估中,CED 成功中和了架构偏差与位置偏差。通过将认知偷懒绝对降低至多 33.00%,CED 带来了显著的性能提升,准确率恢复最高达 30.75%。恢复自主性引发了不同的架构行为——Gemma-2 表现为被动专注任务,而 Llama-3.1 则表现为主动反驳模拟群体——这表明 CED 无需微调即可将顺从性与能力解耦。
cs.AI / 31 / 2609.25572

A Behavioral Trait Leaks into Preferences: Diagnosing Trait Interference in LLM User Simulators

Kim, Chaehyun, Kim, Sein, Kang, Hongseok, Park, Chanyoung
Abstract
LLM-based user simulators aim to bridge the offline-online gap in recommender evaluation by emulating users through injected traits, where preference attributes determine what a user engages with and a behavioral activity trait governs how long they browse. However, we show this intended trait independence collapses during simulation, causing two failures: (i) Trait Interference, where amplified activity distorts preference boundaries and forces interactions with mismatched items to sustain browsing, and (ii) Evaluation Invalidity, where satisfaction scores inflate with activity-driven page counts despite taste mismatches, biasing evaluation toward trait distributions rather than recommender performance. To resolve this, we propose PQA, a page-level quality anchoring method that guides simulators using a personalized anchor reflecting each user's intrinsic preference standard. By assessing whether a page meets this standard before further browsing, PQA enables proactive exits from low-quality pages, letting the activity trait retain its intended role of modulating browsing depth within preference-conforming pages. Experiments show PQA mitigates trait interference and improves the reliability of LLM-based simulator evaluation under activity shifts. Our code is available at https://github.com/chaehyun1/PQA
cs.AI / 32 / 2609.25575

Direct Optimization of Generators for Search in Automated Theorem Proving

Ousherovitch, Adam, Tewari, Ambuj
Abstract
Fine-tuned Large Language Models (LLMs) significantly advance Automated Theorem Proving (ATP), but are often deployed as guiding policies within tree search rather than for single-attempt generation. Recent work shows cross entropy is suboptimal for an LLM used in flat search strategies such as aggregation or filtering and that work has developed new loss functions to correct this misalignment. Extending this alignment to tree search is more challenging: proof discovery depends on exploration and recovery through off-trace states that supervised demonstrations do not reveal. We extend Compute-Aligned Training (CAT) to this setting through an abstraction of policy-guided search, deriving tractable, trace-supported losses. Alongside these search-aware losses, we introduce a search-agnostic uniform-allocation (UA) loss that accounts for the budget without specifying the specific search. Both induce scalar weights on per-tactic cross-entropy gradients. We characterize how off-trace behavior affects the search-aware weights, including conditions for vanishing approximation error at large budgets. On a Lean benchmark, both approaches achieve higher observed proof-success rates than cross-entropy across six search strategies, with strong results from a single shared UA adapter. Budget sweeps show larger gains over cross-entropy at 16 than at 256 expansions, implying CAT scales with test time compute.
cs.AI / 33 / 2609.25581

Gaze responses to false-positive computer-aided detection prompts during colonoscopy: a paired-video and real-time eye-tracking study

结肠镜检查中内镜医师对计算机辅助检测假阳性提示的注视反应:一项配对视频与实时眼动追踪研究
Luo, Te, Zhu, Yan, Fu, Peiyao, Yang, Ruijie, Yang, Xian, Li, Quanlin, Zhou, Pinghong, Wang, Shuo
Abstract
False-positive computer-aided detection (CADe) prompts may divert endoscopists' attention during colonoscopy, yet the attentional impact of individual prompts remains unclear. We used event-locked eye tracking to quantify gaze attraction and attention occupation in complementary retrospective and prospective studies. In a retrospective paired-video experiment, 3 senior and 2 novice endoscopists viewed 60 colonoscopy videos with and without CADe. The prospective study recorded gaze during 42 real-time CADe-assisted colonoscopies performed by 9 senior endoscopists. Screened CADe prompts outside expert-annotated lesion windows were classified as false-positive artifact events. False-positive prompts attracted gaze in 48.6% (68/140) of retrospective observations and 65.2% (533/817) of prospective events. Among attraction events with complete recovery, median attention occupation lasted 1000 ms in the retrospective study and 1100 ms in the prospective study. Corresponding median prompt durations were 33 ms and 267 ms, with median time amplifications of 17.55-fold and 5.15-fold, respectively. In paired retrospective comparisons, visible artifact prompts drew gaze closer to the prompted region than did the same-coordinate unassisted reference. Secondary retrospective analyses showed high lesion gaze recognition without and with CADe (98.0% versus 99.0%). First gaze entry into lesion regions occurred 147.8 ms earlier with CADe. Across controlled and real-time clinical settings, false-positive CADe prompts frequently captured gaze, with attention persisting beyond prompt visibility. These findings support considering prompt-related attentional burden in CADe evaluation and design.
Chinese Translation
计算机辅助检测(CADe)的假阳性提示可能在结肠镜检查过程中分散内镜医师的注意力,但单个提示对注意力的影响尚不明确。我们在互补性的回顾性与前瞻性研究中采用事件锁定眼动追踪技术来量化假阳性提示的注视吸引和注意力占用。在回顾性配对视频实验中,3名资深内镜医师和2名新手内镜医师观看了60段有或无CADe提示的结肠镜检查视频。前瞻性研究记录了9名资深内镜医师在42例实时CADe辅助结肠镜检查中的注视情况。位于专家标注病灶窗口之外的CADe提示被归类为假阳性伪影事件。假阳性提示在回顾性观察中引起注视的比例为48.6%(68/140),在前瞻性事件中为65.2%(533/817)。在注意力完全恢复的吸引事件中,注意力占用的中位时长在回顾性研究中为1000毫秒,在前瞻性研究中为1100毫秒。相应的提示显示时长的中位数分别为33毫秒和267毫秒,时间放大的中位数分别为17.55倍和5.15倍。在配对回顾性比较中,可见的伪影提示使注视点更接近提示区域,而相同坐标的无辅助参考则不然。次要回顾性分析显示,无CADe和有CADe时对病灶的注视识别率均较高(98.0%对99.0%)。使用CADe时,首次注视进入病灶区域的时间提前了147.8毫秒。在受控实验和实时临床环境中,假阳性CADe提示频繁吸引注视,且注意力持续时间超过提示可见时间。这些发现支持在CADe的评估和设计中考虑提示相关的注意力负担。
cs.AI / 34 / 2609.25588

Transformer Heads Looking for Order

寻找顺序的Transformer注意力头
van Doornmalen, Jasper, Kozachinskiy, Alexander, Mathwieser, Corinna, Steifer, Tomasz, Urrutia, Felipe, Verschae, José, Wałȩga, Przemysław Andrzej
Abstract
In this note, we show that the problem of checking, whether a sequence of bits is ordered, is not doable by 1-head 1-layer transformers but is doable by a 2-head 1-layer transformer. Unlike similar previous results, our results assume the model where transformers have an output MLP.
Chinese Translation
在本笔记中,我们证明了一个序列比特是否有序的判定问题无法由单头单层Transformer完成,但可以由双头单层Transformer完成。与以往类似结果不同的是,我们的结果假设的模型中Transformer包含输出MLP(多层感知机)。
cs.AI / 35 / 2609.25591

Evaluating Coding Agents on Kernel Exploit Generation

评估编码智能体在内核漏洞利用生成上的能力
Jang, Junyoung, Lee, Gwanhyun, Lee, Hwiwon, Kim, Kyuheon, Kim, Jongseong, Jung, Jinho, Zhang, Lingming
Abstract
Coding agents now find real vulnerabilities in production software. However, bug discovery results do not measure whether agents can construct exploit primitives. We introduce KEX-bench, a benchmark for evaluating coding agents on exploit primitive generation against real operating-system kernels. KEX-bench contains 45 task instances across 40 Linux and Windows CVEs, covering kernel address leak, instruction-pointer control, heap read, heap write, and arbitrary address write. Each task runs in an isolated virtual machine, exposes controlled tools, and uses a deterministic verifier to check primitive-specific success. We evaluate state-of-the-art coding agents paired with frontier and open-weight models under fixed tool-call budgets. Without a reference proof of concept (PoC), the strongest configuration solves 1 of 20 Windows tasks (5.0%) and 14 of 25 Linux tasks (56.0%). With a reference PoC, the strongest configuration solves 31 of 45 tasks (68.9%). This highlights the gap where agents reach kernel crashes but fail to shape kernel state into exploit primitives. We release KEX-bench for reproducible research on AI-assisted exploitation at https://kex-bench.github.io.
Chinese Translation
编码智能体如今已能够在生产软件中发现真实漏洞。然而,漏洞发现的结果并不能衡量这些智能体能否构建漏洞利用原语(exploit primitives)。我们提出了KEX-bench,一个用于评估编码智能体针对真实操作系统内核生成漏洞利用原语能力的基准测试。KEX-bench包含涵盖40个Linux和Windows CVE的45个任务实例,涵盖内核地址泄露、指令指针控制、堆读取、堆写入和任意地址写入等类型。每个任务在隔离的虚拟机中运行,提供受控的工具,并使用确定性验证器检查特定原语的完成情况。我们在固定的工具调用预算下,评估了搭配前沿模型和开源权重模型的最新编码智能体。在没有参考概念验证(PoC)的情况下,最强配置解决了20个Windows任务中的1个(5.0%)和25个Linux任务中的14个(56.0%)。在有参考PoC的情况下,最强配置解决了45个任务中的31个(68.9%)。这凸显了一个差距:智能体能够触发内核崩溃,却无法将内核状态塑造成漏洞利用原语。我们在https://kex-bench.github.io发布KEX-bench,以支持AI辅助漏洞利用的可复现研究。
cs.AI / 36 / 2609.25607

ArticleMiner: Ontology-Guided Knowledge Graph Construction from Scientific Publications

Jahin, Md Abrar, Knoblock, Craig A., Pujara, Jay
Abstract
Scientific papers keep much of their quantitative content in tables and supplementary files, where a number means something only through its header, caption, unit, analytical method, and the conventions of its field. Recovering the rows and columns of a table is therefore not the same as recovering the scientific fact it reports. Most semantic table-interpretation methods assume that a clean table is already available and subsequently map its cells or columns to ontology terms, whereas most publication-level extraction systems are designed for a single domain. We study a middle path: a shared process that reads a paper and its supplementary files, gathers evidence from several parsers and a language model, and reconciles that evidence, while a bounded human-authored task module for each task supplies the domain meaning. The module lists the canonical names the graph may use, the surface forms that map to them, a small set of derivation rules and validity constraints, an identity key, and the bindings used to write RDF. It defines what a task is allowed to emit; it does not try to list every convention of a field. We build four such modules (for drug-discovery chemistry, materials science, machine learning, and mineral geochemistry) in the ArticleMiner framework, and evaluate them on 163 papers, including a new geochemistry benchmark with expert-curated ground truth. In comparisons against a same-LLM few-shot baseline, the point estimates favor ArticleMiner on all four tasks, with uncertainty on the two smaller benchmarks. The geochemistry comparison also includes access to supplementary files, so its improvement cannot be attributed to domain guidance alone.
cs.AI / 37 / 2609.25618

Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA

基于零空间基LoRA的强化学习后大语言模型保持推理能力微调方法
Fang, Wenzhi, Tzou, Nicholas, Valkov, Lazar, Chappidi, Srinivas
Abstract
Reinforcement learning (RL)-based post-training has become an effective approach for eliciting reasoning capabilities in large language models (LLMs). However, adapting post-RL models to new knowledge domains or behaviors through subsequent supervised fine-tuning (SFT) can severely overwrite these capabilities. Existing approaches mitigate such forgetting through experience replay, specialized initialization, or constrained optimization using gradient projection, but either provide limited preservation or incur substantial training overhead. Our analysis shows that reasoning activations concentrate in low-dimensional subspaces, leaving substantial null-space capacity for adaptation, and that the corresponding approximate null spaces can be reliably estimated from a modest number of examples. Motivated by these observations, we propose Null-Basis Low-Rank Adaptation (NB-LoRA), a parameter-efficient method for adapting post-RL LLMs while preserving their acquired reasoning ability. We formulate reasoning retention as a layer-wise hidden-state preservation constraint and construct a fixed approximate null basis from reasoning activations. LoRA updates are then reparameterized through this basis, enforcing the preservation constraint throughout fine-tuning. Extensive experiments across multiple RL-trained LLMs and diverse downstream tasks show that NB-LoRA matches standard LoRA in adaptation performance, maintains reasoning accuracy near pre-fine-tuning levels, and generalizes this preservation to held-out reasoning benchmarks.
Chinese Translation
基于强化学习(RL)的后训练已成为激发大语言模型(LLM)推理能力的有效方法。然而,通过后续监督微调(SFT)将强化学习后的模型适配到新知识领域或新行为时,往往会严重覆盖这些已习得的能力。现有方法通过经验回放、专门初始化或基于梯度投影的约束优化来缓解此类遗忘问题,但要么保留效果有限,要么带来显著的训练开销。我们的分析表明,推理相关的激活集中于低维子空间,为适配留下了充足的零空间容量,且相应的近似零空间可以通过少量样本可靠地估计得到。基于这些观察,我们提出了零空间基低秩适配方法(Null-Basis Low-Rank Adaptation,NB-LoRA),这是一种参数高效的方法,可在适配强化学习后的大语言模型的同时保持其已习得的推理能力。我们将推理能力的保持形式化为逐层隐状态保持约束,并基于推理激活构建一个固定的近似零空间基,随后通过该基对LoRA更新进行重参数化,从而在整个微调过程中强制执行保持约束。在多个经RL训练的大语言模型和多样化下游任务上的大量实验表明,NB-LoRA在适配性能上与标准LoRA相当,其推理准确率保持在接近微调前的水平,并且这种保持能力可以泛化到留出的推理基准测试中。
cs.AI / 38 / 2609.25620

ChatT2: An Adaptive Framework for Developing a Large Language Model-Based Agent for Natural Product Domain Research

ChatT2:一个用于开发面向天然产物领域研究的基于大语言模型智能体的自适应框架
Wang, Yihan, Gao, Qiandi, Zhuang, Yihui, Ge, Liangjun, Zhang, Heqian, Huang, Jiaquan, Qin, Zhiwei
Abstract
Scientific investigations into microbial natural products (NPs) present significant challenges for novices, largely due to the complexity of microbial systems, biochemical diversity, technical skill requirements, and the demands of bioinformatics and data analysis processes. To address these issues, we introduce ChatT2, a large language model (LLM)-based agent that is specifically tailored to the unique characteristics of bacterial type II polyketides. These polyketides form a structurally distinct and therapeutically important NP family. ChatT2 was developed within an autonomous multiagent framework composed of a mentor, an executor, and an evaluator, each with defined responsibilities. The mentor acts as an intermediary between ChatT2 and the user, utilizing chain-of-thought prompting to refine the intent of the user. Under the guidance of the mentor, the executor synthesizes multimodal information via retrieval-augmented generation techniques and seamlessly integrates bioinformatics and cheminformatics tools. The evaluator ultimately assesses the output of the executor to ensure the richness and accuracy of the retrieved information. Our research highlights how ChatT2, designed with this multiagent framework, addresses the challenges faced by general LLMs in terms of understanding limited, specialized corpora and complex biological information and provides both experts and novices with a valuable tool for exploring various NPs of interest. The ChatT2 webserver can be accessed at https://chatt2.site/#/chat.
Chinese Translation
对微生物天然产物(NPs)的科学研究对新手而言存在重大挑战,这主要归因于微生物系统的复杂性、生化多样性、技术技能要求,以及生物信息学和数据分析流程的高要求。为解决这些问题,我们推出了ChatT2,一个专门针对细菌II型聚酮化合物独特特征而定制的大语言模型(LLM)智能体。这类聚酮化合物是一类结构独特且具有重要治疗价值的天然产物家族。ChatT2在一个自主多智能体框架内开发,该框架由导师(mentor)、执行器(executor)和评估器(evaluator)组成,各自承担明确的职责。导师作为ChatT2与用户之间的中介,利用思维链提示来精炼用户的意图。在导师的指导下,执行器通过检索增强生成技术整合多模态信息,并无缝集成生物信息学和化学信息学工具。评估器最终对执行器的输出进行评估,以确保检索信息的丰富性和准确性。我们的研究表明,采用这种多智能体框架设计的ChatT2能够解决通用大语言模型在理解有限的专用语料库和复杂生物信息方面所面临的挑战,为专家和新手提供一个探索各种感兴趣天然产物的宝贵工具。ChatT2网络服务器可通过https://chatt2.site/#/chat访问。
cs.AI / 39 / 2609.25643

Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces

思维阶梯:一种逐步简化推理轨迹的自演化课程
Liu, Minghui, Magelinski, Thomas, Yuan, Dehao, Yu, Qi, Huang, Furong
Abstract
Large language models (LLMs) excel at reasoning when scaled to hundreds of billions of parameters, but small- and mid-scale models remain brittle reasoners even with knowledge distillation (KD). We present Ladders-of-Thought (LoT), a framework that improves reasoning by combining progressive question rewrites with a self-evolving curriculum. LoT automatically generates semantically faithful but easier variants of reasoning problems, organizes them into difficulty buckets using step-based measures, and employs a self-evolving bandit scheduler to allocate training adaptively. Evaluated on two reasoning domains, math and multi-hop reasoning, across 1-8B models from different families, LoT consistently improves over KD. It delivers large gains on arithmetic tasks (e.g., +32 percentage points on AddSub, +25pp on SVAMP), +2-8pp improvements on in-domain test splits, and strong though dataset-dependent benefits on multi-hop reasoning (e.g., +16pp on QASC, +25pp on StrategyQA). LoT also converges faster than staged curricula, highlighting the value of adaptive progression. These results show that progressive rewrites coupled with adaptive curricula provide a simple yet effective recipe for strengthening reasoning in smaller LLMs.
Chinese Translation
大型语言模型(LLM)在参数规模达到数千亿时表现出卓越的推理能力,但中小规模模型即使经过知识蒸馏(KD),推理能力仍然脆弱。我们提出了思维阶梯框架(Ladders-of-Thought, LoT),通过将渐进式问题改写与自演化课程相结合来提升推理能力。LoT 自动生成语义忠实但更简单的推理问题变体,使用基于步骤的度量将其组织到难度分级中,并采用自演化的多臂老虎机调度器自适应地分配训练任务。在数学和多跳推理两个推理领域、跨不同模型系列的 1-8B 模型上的评估表明,LoT 持续优于知识蒸馏。它在算术任务上带来了显著提升(例如 AddSub 提升 32 个百分点,SVAMP 提升 25 个百分点),在域内测试集上提升 2-8 个百分点,并在多跳推理上取得了可观但依赖于数据集的收益(例如 QASC 提升 16 个百分点,StrategyQA 提升 25 个百分点)。LoT 还比阶段性课程收敛更快,凸显了自适应推进的价值。这些结果表明,渐进式改写与自适应课程相结合,为增强小型 LLM 的推理能力提供了一种简单而有效的方案。
cs.AI / 40 / 2609.25647

Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark

Cao, YanZe
Abstract
Predicting early outcomes based on trajectory can decrease the expenses associated with agent evaluation by terminating a run once the outcome becomes sufficiently predictable, assuming that the predictor's confidence is properly calibrated. Calibration is at risk when a predictor is applied to an agent on which it was never trained, but it is not known whether such transfer failures are broad across agent systems or concentrated in specific target agent/head combinations. Using public SWE-bench Verified trajectories and a frozen dual-head early-outcome prediction pipeline, we ran a leave-one-agent-out calibration audit, a shared-predictor leave-two-agents-out control, oracle prior correction, and a robustness battery over training cohorts, task resampling, task halves, jackknife, and thresholds. Fixed-scaffold TerminalBench analysis served as a pre-registered boundary test. Broad same-predictor pairwise heterogeneity was not supported; the median pairwise corrected-gap differences were 0.0180 (SUCCESS head, 45 pairs) and 0.0385 (FAILURE head, 35 pairs), and the pre-registered heterogeneity criterion was not met on either head. Two specific combinations, gpt-5-mini/SUCCESS and claude-opus-4.6/FAILURE, showed persistent calibration-transfer errors (median corrected gaps 0.1377 and 0.1107) without a sign reversal under any frozen control. TerminalBench did not establish cross-benchmark replication: the success target produced zero decisions (INDETERMINATE), and the failure target did not satisfy the pre-registered persistence criterion. Therefore, a strong target-specific calibration-transfer error can exist within one frozen environment, but the evidence does not establish that the error is intrinsic to the model or general across benchmarks.
cs.AI / 41 / 2609.25677

Seeing Is Not Perceiving: When Synthetic Consumers Can and Cannot Pretest Visual Marketing

Tsai, Yi-Lin, Yung-Hsiu, Lai
Abstract
Marketers now deploy generative AI agents as synthetic consumers to pretest visual assets such as logos, packaging, and advertising at a fraction of human-panel cost. However, this procedure assumes that a model seeing a visual cue can also perceive its consumer meaning, which is largely untested. We stress-test the assumption using six canonical visual marketing experiments, varying the two levers managers control: model generation (GPT-4o-mini vs. GPT-5.4-mini) and input format (plain text vs. JSON). Every resulting configuration passed the manipulation checks; however, none of the configurations reproduced more than two of the six human effects, and the remainder were nonsignificant. The one exception was a significant reversal of the human pattern. Providing conceptual or empirical evidence through in-context learning steers average responses toward the human effect. Yet steering has a limit: even when it succeeds, a configuration reproduces less than half of the natural spread of human responses and so understates consumer heterogeneity. We integrate these results into an AI governance protocol (Calibrate, Intervene, Deploy) that delineates when synthetic consumers can responsibly screen creatives and when human panels remain necessary.
cs.AI / 42 / 2609.25678

Toolcompass: Guiding Tool Trialing, Not Suppressing It

Fang, Junlin, Zhang, Chong, Nguyen-Thanh, Do, Xu, Xiaogang, Fang, Zhen, Du, Sean
Abstract
Large language model (LLM) agents must generalize from tools seen during training to unseen tools at deployment. A key challenge is tool trialing, i.e., excessive trials waste the interaction budget, whereas selective trials enable exploration of unfamiliar tools. Existing outcome-based post-training leaves wasteful trials unguided, while turn-level supervision may suppress necessary exploration. We introduce ToolCompass, a post-training framework that guides tool trialing by organizing tool-call representations according to shared functions. Specifically, ToolCompass models each function class as a von Mises--Fisher distribution and jointly reduces intra-function variation across domains and increases inter-function separation. This structure transfers experience from seen tools to functionally similar unseen tools, directing exploration away from unrelated alternatives. ToolCompass requires no ground-truth call traces or unseen-tool access and incurs no inference overhead. Experiments on AppWorld and FTRL show consistent gains across GRPO, RFT, and DMPO. improves AppWorld OOD task success by up to 10.71 percentage points over vanilla post-training and performs best among competitive baselines on both benchmarks.
cs.AI / 43 / 2609.25686

How Strongly Should Task State Influence an LLM Agent?

任务状态应以多强程度影响LLM智能体?
Zhang, Chenyu, Kweon, Wonbin, Han, Jiawei
Abstract
Long-horizon assigned work requires an LLM agent to track the state of a task: which steps are done, blocked, cancelled, or open to repetition. Agent systems either keep this state as text in the prompt and rely on the model to read that text, or move the state into a module that enforces it, and each system is evaluated as a whole, so no one knows how much reliability comes from the state being shown, told, or enforced. We fix the task rules, the model, and paired episodes and vary how strongly task state reaches the agent: a raw transcript, an exact checklist, per-turn directives from a state machine compiled from the brief and advanced only by execution receipts, or an enforcement gate on that machine that refuses state-violating actions; every episode is scored by exact payload matching against dynamic ground truth. Across three models, two reasoning regimes, and two domains, four findings hold without per-turn reasoning: displaying accurate state is unreliable, an unverified ledger the agent writes itself beats an accurate checklist it is shown, directives help in proportion to the model's obedience, and enforcement needs no obedience but is bounded by the correctness of its state and by the matcher that maps requests to steps; per-turn reasoning at a 235B agent compresses these separations without repairing the text rungs. The same gate, compiled from $\tau^2$-bench's airline policy, raises a 235B agent's pass$^1$ from 0.39 to 0.54 and changes nothing for a 35B agent that rarely violates the policy; on PM-Bench, where acting turns on recognizing a cue rather than on state, showing the record is the best rung--matching or beating both gates and reversing the ledger-over-checklist finding--and enforcing the matcher's judgement drops a 35B agent below its raw transcript. Enforcement pays when failures are state-decidable and frequent, and hurts when the gate's judgement is wrong.
Chinese Translation
长时程指派任务要求LLM智能体跟踪任务状态:哪些步骤已完成、被阻塞、被取消或可以重复执行。现有的智能体系统要么将任务状态以文本形式保留在提示词中并依赖模型读取该文本,要么将状态移入一个强制执行该状态的模块,而每个系统都是作为整体被评估的,因此没有人知道可靠性究竟有多少来自状态的展示、告知或强制执行。我们固定任务规则、模型以及配对的情景(episodes),并改变任务状态传达给智能体的强度:原始转录文本、精确的清单(checklist)、由任务简报编译的状态机并仅凭执行回执推进而发出的逐轮指令,或是对该状态机的强制执行门(阻止违反状态的操作);每个情景均通过与动态真值的精确载荷匹配进行评分。在三个模型、两种推理模式和两个领域上,在不使用逐轮推理的情况下得出四项发现:展示准确的状态并不可靠;智能体自行写入的未经核实的账本(ledger)优于展示给它的精确清单;指令的帮助程度与模型的服从度成正比;强制执行不依赖服从度,但受限于其状态的正确性以及将请求映射到步骤的匹配器的正确性。在235B规模的智能体上使用逐轮推理会压缩上述差异,但无法修复文本层级的缺陷。同一个门,由 $\tau^2$-bench 的航空政策编译而来,将235B智能体的 pass$^1$ 从0.39提升至0.54,而对一个很少违反政策的35B智能体没有任何改变;在PM-Bench上,由于行动取决于识别提示而非状态,展示记录是最佳的层级——其效果匹配或超过两种门,并逆转了账本优于清单的发现——而强制执行匹配器的判断会使35B智能体的表现低于其原始转录文本。当失败是状态可判定且频繁发生时,强制执行是有收益的;而当门的判断出错时,它则有害。
cs.AI / 44 / 2609.25712

TCMaster: Confidence-Aware Querying and Workload-Guided Physical Design for Multi-Source Traditional Chinese Medicine Knowledge Graphs

Chen, Zheng, Li, Yuzhu, Li, Haoxuan, Zhang, Zhongde, Jin, Lianshun, Qin, Peiwu
Abstract
Multi-source knowledge graphs (KGs) need query mechanisms that expose reliability and exploit domain structure. This paper presents TCMaster, a property-graph query substrate for confidence-aware traversal and workload-guided physical design over Traditional Chinese Medicine KGs. TCMaster integrates pharmacopoeias, prescriptions, molecular databases, and LLM-extracted micro-semantics into a KG with approximately 221K entities and 723K base edges. It annotates edges with provenance-level confidence, rewrites Cypher queries with confidence predicates, ranks multi-hop paths under PRODUCT, MIN, or weighted-average policies, and uses ontology skew through direction selection, herb-attribute bitmaps, and materialized shortcut edges. On Neo4j, direction selection improves attribute lookup by a factor of 1.47, shortcuts accelerate high-fanout target counting by a factor of 4.42, confidence filtering removes 39.3 percent of low-quality heterogeneous paths, and KG retrieval improves TCMbench QA accuracy by 20.0 percentage points.
cs.AI / 45 / 2609.25715

LingLan: An Advancing Traditional Chinese Medicine Diagnosis LLM with Multimodal Data

LingLan:基于多模态数据的先进中医诊断大语言模型
Chen, Zheng, Du, Zhicheng, Li, Haoxuan, Liang, Yingshan, Qin, Peiwu
Abstract
Though artificial intelligence (AI) increasingly transforms modern medicine, its integration into Traditional Chinese Medicine (TCM) has been relatively slow, primarily due to TCM's reliance on holistic, subjective diagnostic methods---namely Inspection, Auscultation and Olfaction, Inquiry, and Palpation(I-AOI-P)---which are difficult to align with quantitative, standardized medical systems. In this work, we introduce a Unification Framework for Multimodal Data (UFMD), which automatically processes tongue and pulse images into structured, clinically standard descriptions, integrating multi-source diagnostic information into a unified digital record of I-AOI-P process. Building on this structured data, we create LingLan-14B, a TCM-specific large language model fine-tuned via supervised learning to emulate the diagnostic logic and workflow of I-AOI-P process. Experimental results show that our method significantly enhances diagnostic accuracy, achieving a relative improvement of 103.5% over the baseline (62.72% vs. 30.82%) and reaching an F1-score of up to 82%.
Chinese Translation
尽管人工智能(AI)正在日益变革现代医学,但其与传统中医(TCM)的融合相对缓慢,主要原因在于中医依赖整体性、主观性的诊断方法——即望、闻、问、切(Inspection, Auscultation and Olfaction, Inquiry, and Palpation, I-AOI-P)——这些方法难以与量化、标准化的医学体系相契合。在本工作中,我们提出了一个多模态数据统一框架(Unification Framework for Multimodal Data, UFMD),可将舌象和脉象图像自动处理为结构化的临床标准描述,将多源诊断信息整合为望闻问切过程的统一数字化记录。基于这一结构化数据,我们构建了LingLan-14B,一个通过监督学习微调的中医专用大语言模型,用以模拟望闻问切过程的诊断逻辑与工作流程。实验结果表明,我们的方法显著提升了诊断准确率,相对于基线取得了103.5%的相对提升(62.72% vs. 30.82%),F1分数最高达到82%。
cs.AI / 46 / 2609.25738

OmniFysics-Nano-V2 Technical Report: Understanding the Physical World Across Modalities

Liu, Yizhou, Han, Jinghang, Qiu, Kaixiang, He, Qi, Han, Minghao, Jiang, Yue, Chen, Xujia, Zou, Wei, Wang, Shunli, Zhang, Lihua, Yang, Dingkang
Abstract
Omni-modal models have expanded multimodal interaction across vision, audio, speech, and language. However, their training is predominantly organized around semantic descriptions and general-purpose objectives, leaving physical attributes, interaction states, and causal mechanisms only partially specified. This gap is not simply a matter of modality coverage: adding more modalities does not by itself provide the supervision needed to connect observations with the physical structure of the world. We present OmniFysics-Nano-V2, a compact omni-modal model for physical-world perception and understanding. The model supports image, video, audio, speech, and text inputs within a shared reasoning framework, together with text and speech generation. To address the lack of explicit physical supervision, we construct a dual-branch physics-aware data pipeline that grounds salient objects in structured physical attributes and aligns visual changes with acoustic events, intermediate responses, and interaction outcomes. To address homogeneous training objectives, we curate reinforcement-learning prompts by reward diversity and adopt a two-stage Group Relative Policy Optimization curriculum that progresses from general task correctness to fine-grained physical perceptual reasoning. Experiments across multimodal, audio-visual, and physical reasoning benchmarks show that the proposed data and training strategy improves physical-world understanding while preserving broad omni-modal competence. The proposed model achieves leading result on 17 of 21 benchmarks against SOTA omni-modal models. By equipping AI systems with both omni-modal and physical-world perception capabilities, OmniFysics-Nano-V2 is poised to become a cornerstone of next-generation Physical AI.
cs.AI / 47 / 2609.25760

The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance

Ziaei, Rojin
Abstract
Using large language models (LLMs) to simulate diverse human populations has the potential to transform many aspects of computational social science, yet many evaluations score the average response rather than the spread of opinion within real groups. Here, we develop a diagnostic framework that measures point accuracy alongside dispersion retention, the ratio of predicted to human standard deviation ($\dr$), on 10{,}000 respondent--question pairs from the World Values Survey (WVS) spanning twelve countries and six continents. We evaluate eleven zero-shot language models and five variants fine-tuned on WVS data with SFT, DPO, and GRPO. We identify a failure mode we term \textit{consensus collapse}, where alignment training compresses outputs toward one stereotype per group. Along the post-training trajectory from the Llama~3.1 70B base to the Tulu~3 checkpoints, the first stage, supervised instruction tuning, removes half of the spread with minimal accuracy gain ($\dr$ 1.22 to 0.59; accuracy $+0.9$ points), the later stages do not restore it, and a gap opens between WEIRD and non-WEIRD countries that survey fine-tuning then deepens while pursuing higher point accuracy. The most accurate model (Tulu~3 70B-DPO fine-tuned on WVS, 57.9\%) keeps half the human spread overall ($\dr = 0.50$) and 11\% of it for Nigeria, against 0.70--0.87 for WEIRD countries. Raising the sampling temperature to 1.0 leaves the Wasserstein-1 distance ($\wone$) to human distributions unchanged for both fine-tuned DPO models, and GRPO on Qwen~3.5 9B does not restore the spread under either an accuracy reward or a distribution-shaped reward. Mixing the aligned model with an unaligned prior raises $\dr$ from 0.51 to 0.62 on a held-out split but leaves Nigeria at 0.36. Point accuracy alone therefore misjudges these simulators, and current post-training trades diversity for consensus.
cs.AI / 48 / 2609.25766

Neurosymbolic Action Model Learning under Partial Observability

部分可观测条件下的神经符号动作模型学习
Kikaj, Adem, De Smet, Lennert, Marra, Giuseppe, De Raedt, Luc
Abstract
AI planning studies how an agent can reach a goal by executing a sequence of actions. To plan correctly, the agent needs an action model describing when each action can be executed and how it changes the world. Constructing such models by hand requires domain expertise, and can be costly and error-prone. Action models can instead be learned from available data using existing neurosymbolic approaches, but they currently assume access to complete traces of fully observable images . These approaches fail to learn action models under partial observability where some of the images might not be present or are not fully informative of the current state of the world. Hence, this paper proposes NeSyAM, a novel neurosymbolic modeling paradigm for action model learning under partial observability. In addition, the paper presents a unified variational framework for theoretically analysing the limitations of existing methods compared to our proposed approach. NeSyAM is then tested extensively on six visual planning domains and three observation regimes to show it consistently recovers relevant parts of the true action model under partial observability.
Chinese Translation
AI规划研究智能体如何通过执行一系列动作来达成目标。为了正确地进行规划,智能体需要一个动作模型来描述每个动作何时可以执行以及它如何改变世界。手工构建此类模型需要领域专业知识,且成本高且容易出错。动作模型也可以利用现有的神经符号(neurosymbolic)方法从可用数据中学习,但这些方法目前依赖于完全可观测图像的完整轨迹。在部分可观测条件下,即某些图像可能缺失或无法完全反映当前世界状态时,这些方法无法学习动作模型。因此,本文提出了NeSyAM,一种用于部分可观测条件下动作模型学习的新型神经符号建模范式。此外,本文提出了一个统一的变分框架,从理论上分析现有方法与我们提出的方法相比的局限性。随后,我们在六个视觉规划领域和三种观测情形下对NeSyAM进行了广泛测试,结果表明它在部分可观测条件下能够持续恢复真实动作模型的相关部分。
cs.AI / 49 / 2609.25769

Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction

基于掩码轨迹预测的全维度GUI智能体导航
Zhang, Yan, Fu, Pei, Wu, Daiqing, Shen, Huawen, Zhang, Ruoceng, Zhang, Shaojie, Yang, Jiahui, Zhou, Yu, Ma, Can, Luo, Zhenbo, Luan, Jian
Abstract
Graphical User Interface (GUI) Agents autonomously interact with software to fulfill user requests, where GUI navigation stands out as the most critical and challenging capability. Mastering this capability demands a complex synergy of step-wise decision-making, state-action alignment, and long-horizon planning. While directly mixing these corresponding navigation tasks seems intuitive to simultaneously acquire these skills, such a direct combination is severely bottlenecked by inconsistent optimization objectives and profound data heterogeneity. To overcome these barriers, we propose the MaP (stands for ``\textbf{M}asked Tr\textbf{a}jectory \textbf{P}rediction''), a unified framework that seamlessly harmonizes divergent GUI navigation tasks. By modeling multi-turn GUI interactions as a trajectory and defining training objectives through component masking and prediction, MaP shifts the optimization from task-specific marginal distributions to a consistent objective. Furthermore, to handle the data heterogeneity across multiple navigation tasks, we design a role-aware adapter learning module that dynamically routes each token to a specialized representation space. Extensive experiments on five representative GUI navigation benchmarks demonstrate that MaP effectively mitigates gradient conflicts and significantly outperforms the direct mixture training, establishing a robust paradigm for multi-task GUI navigation.
Chinese Translation
图形用户界面(GUI)智能体通过与软件自主交互来完成用户请求,其中GUI导航是最关键且最具挑战性的能力。掌握这一能力需要逐步决策、状态-动作对齐以及长时程规划之间的复杂协同。虽然直接混合这些相应的导航任务似乎是同时获取这些技能的直观方法,但这种直接组合受到了优化目标不一致和数据异质性深刻差异的严重制约。为克服这些障碍,我们提出了MaP(即"掩码轨迹预测",Masked Trajectory Prediction),一个能够无缝协调多种不同GUI导航任务的统一框架。通过将多轮GUI交互建模为轨迹,并通过组件掩码和预测来定义训练目标,MaP将优化从任务特定的边缘分布转变为一致性目标。此外,为处理多个导航任务间的数据异质性,我们设计了一种角色感知适配器学习模块,可动态地将每个令牌路由至专门的表示空间。在五个具有代表性的GUI导航基准上的大量实验表明,MaP有效缓解了梯度冲突,并显著优于直接混合训练,为多任务GUI导航建立了一个鲁棒的范式。
cs.AI / 50 / 2609.25804

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

有品味的智能体:长程任务中品味的度量与提升
Pan, Wenbo, Liu, Zhichao, Liu, Shujie, Zeng, Jingying, Lin, Chin-Yew, Tang, Xianfeng, Lu, Yan, He, Qi, Jia, Xiaohua
Abstract
LLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run. Making these decisions well is becoming a key capability for both engineering and research agents. We refer to the ability to make good long-horizon decisions as the taste of an agent. While existing benchmarks measure the end-to-end success of agents on long-horizon tasks, none of them measures the taste of an agent. To address this problem, we build Taste-Bench, a benchmark of taste questions constructed automatically from trajectories that agents produced in engineering and research tasks. Each question presents a decision fork, a point in a trajectory where multiple directions are available and one of them leads to a better outcome, and the evaluated model chooses among these directions without seeing what happens after the fork. We mine these forks automatically from parallel attempts at the same task and from detours inside a single trajectory, without needing human annotation. We evaluate frontier models on Taste-Bench and find that the best model answers only 59.7% of the questions correctly. We further find that forks whose deciding evidence appears later in the trajectory are much harder for every model, and that a larger reasoning budget does not improve the accuracy. Finally, we show that taste can be trained. We distill the judgment of a teacher that has seen the outcome into a student model, and the student makes better decisions on unseen tasks and improves end-to-end success on held-out SWE-bench Pro tasks.
Chinese Translation
大语言模型(LLM)智能体越来越多地承担长程任务,而它们在过程中做出的决策——例如测试哪个假设、基于哪个实现继续构建——决定了整个运行过程的最终结果。能否做好这些决策,正日益成为工程类和研究类智能体的一项关键能力。我们将做出良好长程决策的能力称为智能体的“品味”(taste)。现有基准测试衡量的是智能体在长程任务上的端到端成功与否,但没有任何一个衡量智能体的品味。为解决这一问题,我们构建了 Taste-Bench,这是一个从智能体在工程和研究任务中产生的轨迹中自动构建的品味问题基准。每个问题呈现一个“决策分叉点”——即轨迹中存在多个可选方向、且其中之一能带来更好结果的位置——被评估的模型需在不看到分叉点之后发展的情况下,在这些方向中做出选择。我们通过对同一任务的多次并行尝试以及单条轨迹内部的绕路过程自动挖掘这些分叉点,无需人工标注。我们在 Taste-Bench 上评估了前沿模型,发现表现最好的模型也仅能正确回答 59.7% 的问题。我们进一步发现,决定性证据出现在轨迹后部的分叉点对所有模型都更加困难,而更大的推理预算并不能提升准确率。最后,我们证明品味是可以训练的:我们将一个已看到最终结果的教师模型的判断力蒸馏到学生模型中,学生在未见过的任务上能做出更好的决策,并在留出的 SWE-bench Pro 任务上提升了端到端成功率。
cs.AI / 51 / 2609.25806

When Are Aggregate Agent Traces Diagnosable? Traffic-Governed Interpretation and Calibrated Abstention

聚合智能体轨迹何时可诊断?流量主导的解释与校准弃权
Zhu, Peiying, Chang, Sidi
Abstract
Runtime traces can appear transparent, but a closed-loop policy determines which states are visited and which failures become visible. We study a simulated hotel-pricing agent mapping time, inventory, and market state to discrete price actions under varying demand regimes. A fault may leave no aggregate trace when the policy rarely visits affected cells. We treat entry into aggregate-only fault interpretation as a diagnosability decision preceding scoring or localization. A reference-map gate requires repeated clean-policy support; a matched runtime gate then requires joint support in clean and current streams. Signal analysis occurs only after both pass. We calibrate false admission on a disjoint clean stream at the physical-component level and model detection by affected clean traffic rather than nominal cell coverage. In a frozen one-shot heldout, 55/72 (76.4%) regime-component units were reference-admitted, representing 20 physical components; 54/55 passed matched runtime admission, while the rejected unit abstained. Stable false admission was 0/20, with a one-sided exact 95% upper bound of 0.1391, meeting the frozen 0.20 criterion. Across 540 repeated unit-arm rows nested in those 20 clusters, affected clean traffic reduced negative log likelihood by 29.3% relative to cell coverage, a gain of 0.1264 nats per row (cluster-bootstrap 95% interval [0.0593, 0.1918]). Adding mask family and its interaction improved log loss by 0.0015 nats per row (one-sided upper bound 0.0066), below the frozen 0.01 practical-sufficiency margin. A development audit found that exact minimum hitting set and greedy selection chose identical supports in 12/12 scenarios because singleton evidence had resolved the conflicts. The result is a bounded rule for interpreting aggregate agent behavior: first establish exposure, then score change, and abstain when the trace cannot support the claim.
Chinese Translation
运行时轨迹看似透明,但闭环策略决定了哪些状态会被访问、哪些故障会变得可见。我们研究一个模拟的酒店定价智能体,该智能体在不同的需求情境下将时间、库存和市场状态映射为离散的价格动作。当策略很少访问受影响的单元格时,故障可能在聚合轨迹中不留任何痕迹。我们将进入仅基于聚合的故障解释视为一项可诊断性决策,它先于评分或定位。参考映射门控要求清洁策略重复提供支持;随后匹配运行时门控要求清洁数据流与当前数据流提供联合支持。只有两个门控均通过后,才进行信号分析。我们在一个不相交的清洁数据流上于物理组件层面校准错误准入率,并利用受影响的清洁流量(而非名义单元格覆盖率)对检测能力进行建模。在冻结的一次性保留测试中,72个情境-组件单元中有55个(76.4%)获得参考准入,对应20个物理组件;55个中有54个通过了匹配运行时准入,被拒绝的单元则选择弃权。稳定的错误准入为0/20,单侧精确95%置信上界为0.1391,满足冻结的0.20标准。在嵌套于这20个聚类中的540条重复单元-臂记录中,相对于单元格覆盖率,受影响的清洁流量将负对数似然降低了29.3%,即每条记录提升0.1264奈特(聚类自助法95%置信区间为[0.0593, 0.1918])。加入掩码族及其交互项后,对数损失每条记录仅改善0.0015奈特(单侧上界为0.0066),低于冻结的0.01实际充分性边际。开发阶段审计发现,精确最小命中集与贪心选择在12个场景中均选出相同的支持集,因为单点证据已经解决了冲突。最终结果是一个用于解释聚合智能体行为的有界规则:首先确立暴露性,然后对变化进行评分,当轨迹无法支持结论时选择弃权。
cs.AI / 52 / 2609.25848

Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts

优化了分数,却偏离了任务:权重、选择与提示词中的奖励作弊(Reward Hacking)
Wahi, Vansh
Abstract
A higher evaluation score does not always mean a better language model system. When optimization exploits an evaluator's mistakes, measured progress can conceal unchanged or deteriorating task performance. This failure can arise through parameter updates, selection among generated outputs, or revisions to persistent prompts. We develop a comparative framework for reward hacking across these three optimization substrates: weights, selection, and text. Building on the Proxy Compression Hypothesis and research on inference-time and in-context reward hacking, we examine how reachable behavior, optimization budgets, and persistent adaptation shape exposure to proxy error. We formalize a distance-dependent upper bound on evaluator disagreement and a capacity ordering for nested policy classes, then show why distance alone cannot establish a universal ranking of vulnerability. An exact finite-output illustration demonstrates how the location of a scoring defect changes the behavior favored by each method. We also map representative defenses across substrates, identifying which mechanisms transfer directly and which offer only functional analogies. Persistent prompts receive particular attention: their contents are inspectable, but the behavior induced by a small textual change may be difficult to anticipate. The formal analysis, numerical illustration, and published evidence together provide a basis for comparing optimization methods and identifying the conditions under which their defenses transfer. The resulting framework connects optimization choices to verification requirements: reliable improvement depends on controlling accessible failure modes and preserving evidence of task quality independent of the score being optimized.
Chinese Translation
更高的评估分数并不总是意味着更好的语言模型系统。当优化过程利用评估器的错误时,所测得的进展可能掩盖了任务性能的停滞甚至退化。这种失效可能通过参数更新、对生成输出的选择,或对持久化提示词的修改而产生。我们建立了一个比较框架,以研究奖励作弊(reward hacking)在这三种优化基底——权重、选择和文本——上的表现。基于代理压缩假说(Proxy Compression Hypothesis)以及推理时和上下文内奖励作弊的研究,我们考察可达行为、优化预算和持久性适应如何影响模型暴露于代理错误的风险。我们形式化了评估器分歧的与距离相关的上界,以及嵌套策略类的容量排序,进而说明为何仅凭距离无法建立一个通用的脆弱性排序。一个精确的有限输出示例展示了评分缺陷的位置如何改变各优化方法所偏好的行为。我们还梳理了跨基底的代表性防御方法,识别哪些机制可以直接迁移,哪些仅提供功能上的类比。我们特别关注持久化提示词:其内容虽然可被检查,但微小的文本修改所引发的行为却可能难以预测。形式化分析、数值示例与已发表的实证证据共同为比较不同优化方法以及确定其防御可迁移的条件提供了基础。由此形成的框架将优化选择与验证需求联系起来:可靠的改进取决于控制可触达的失效模式,并在独立于被优化分数的前提下保留任务质量的证据。
cs.AI / 53 / 2609.25852

Prediction Is Not Detection: Evaluating Pre-Recognition Claims in Longitudinal Clinical AI

预测不等于检测:评估纵向临床AI中的识别前主张
Yang, Jing, Jiao, Long R., Cai, Xiujun, Zhang, Zongjiu
Abstract
Clinically useful early detection requires validated pre-recognition lead time. Yet event-based evaluations of longitudinal clinical AI can treat recognition-mediated care-process signals as shortcuts and recognition-dependent endpoints as reference standards, inflating apparent performance and lead time while undermining cross-center transport. Such results may serve prognosis without establishing detection before recognition. We define an interval-censored pre-recognition transition, an independent as-of reference standard, and a prespecified recognition proxy to make the claim testable.
Chinese Translation
临床上有用的早期检测需要经过验证的识别前提前期。然而,针对纵向临床AI的基于事件的评估可能将经由识别介导的诊疗过程信号视为捷径(shortcut),并将依赖于识别的终点视为参考标准,从而夸大了表观性能和提前期,同时削弱了跨中心的可迁移性。这类结果可能具有预后价值,但无法确立在识别发生之前的检测能力。我们定义了区间删失的识别前转换、独立的“截至某时点”参考标准,以及预先设定的识别代理变量,从而使该主张具有可检验性。
cs.AI / 54 / 2609.25873

AgenticSizing: A Large Language Model-based Multi-Agent Framework for Analog Circuit Sizing

Hao, Yijia, Verma, Pratibha, Guo, Dongxu, Sestito, Cristian, O'Boyle, Michael, Bouganis, Christos-Savvas, Prodromakis, Themis
Abstract
Analog circuit sizing remains a challenging and time-consuming task due to the large design space, strong performance trade-offs, and increasing circuit complexity in scaled technologies. Although recent large language model (LLM)-based methods show promise in improving sample efficiency and interpretability, existing approaches often lack explicit circuit-topology understanding and are mainly evaluated on relatively simple analog building blocks. This paper presents a multi-agent LLM-based framework for complex analog circuit sizing. The proposed framework first analyzes the circuit topology and decomposes the netlist into functional blocks and substructures. It also extracts lightweight design knowledge for reuse. Based on the extracted topology and knowledge, a planner coordinates multiple role-specialized sizing agents to update design variables and achieve global performance specifications. This workflow mimics the collaborative process of an expert analog design team and provides a structured, interpretable, and simulation-driven optimization procedure. The framework was validated on eight circuits, with the largest design containing up to 55 transistors and 60 sizing variables. Notably, for the LDO benchmark, the proposed method achieved a 60\% success rate with an average of 83 iterations, where classical optimizers failed to find feasible solutions. Further, ablation studies demonstrate that topology understanding, design-knowledge infusion, and agent specialization provide complementary benefits. The source code is available to support reproducibility.
cs.AI / 55 / 2609.25960

CausalLoss-Fin: Attributing Financial-Agent Loss to Decisions and Infrastructure Faults

CausalLoss-Fin:将金融智能体损失归因于决策与基础设施故障
Sharma, Abhishek
Abstract
When an agent handling a payment exception loses money, the agent-step attribution methods this paper compares against will name one of its actions. They will do so even when a settlement message was dropped and the agent never had a chance: they intervene on agent actions and do not expose infrastructure faults as intervenable variables, so every dollar they explain is charged to a decision. We take a benchmark whose fault process is explicit and replayable, decompose each episode's realised delivery schedule into named, individually repairable messages, and intervene on both the agent's choices and the infrastructure's. A telescoping identity splits any policy's loss exactly three ways: an infrastructure effect, a policy differential against the best implementable policy, and a reference-policy residual. Two of the three can be negative, so none is a share; Shapley then divides the first into signed allocations over individual messages. One result is structural and needs no corpus: an agent-only baseline identifies no infrastructure cause, because its model contains no variable that could name one. What 545 planted episodes across 3 policies measure is the size of that consequence. It misfiles 100% of infrastructure episodes and charges $114,383.40 to the agent. Repairing what it names recovers 0.0% of the available loss; repairing a minimal sufficient set recovers 100.0%. Scoring messages one at a time is not merely imprecise: 27.8% (95% CI: 23.3--32.3%) of episodes do not decompose additively. We evaluate deterministic programmatic policies rather than language-model agents, which is what makes replay exact and which limits external validity to stochastic agents. The prevalence figures are properties of this generator, not field rates.
Chinese Translation
当一个处理支付异常的智能体造成资金损失时,本文所对比的智能体步骤归因方法会将其归结为该智能体的某一动作。即使结算消息在传输中被丢弃、智能体根本没有机会做出反应,这些方法也会这样做:它们只对智能体动作进行干预,而不将基础设施故障作为可干预的变量暴露出来,因此它们解释的每一美元损失都被记在决策头上。我们采用一个故障过程明确且可重放的基准,将每个回合的实际交付调度分解为若干命名的、可单独修复的消息,并同时干预智能体的选择和基础设施的行为。一个伸缩恒等式将任何策略的损失精确地分解为三部分:基础设施效应、相对于最佳可实施策略的策略差分,以及参考策略残差。这三项中有两项可能为负,因此任何一项都不构成损失份额;随后,Shapley值将基础设施效应分解为针对各条消息的带符号分配。其中一个结论是结构性的且无需语料库:仅针对智能体的基线方法无法识别任何基础设施原因,因为其模型中不存在可以指称此类原因的变量。跨3个策略的545个植入故障回合所测量的是这一后果的规模。该基线方法将100%的基础设施故障回合错误归档,并向智能体收取114,383.40美元的损失。修复其所指认的对象可挽回的损失为0.0%;而修复一个最小充分集合可挽回100.0%。逐条消息单独打分并不仅仅是精度不足的问题:27.8%(95%置信区间:23.3%–32.3%)的回合不具备可加分解性。我们评估的是确定性的程序化策略而非语言模型智能体,这正是重放能够精确进行的原因,同时也将外部效度限制在了随机智能体之外。所述的普遍性数字是该生成器的属性,而非实际现场的发生率。
cs.AI / 56 / 2609.26015

VideoX-Qwen: Data-Centric Instruction-Based Video Editing

VideoX-Qwen:以数据为中心的基于指令的视频编辑
Li, JJiahang, Shao, Dingbao, Chen, Xinyu, Wu, Song, Lin, Jiang, Li, Duo, Liu, Yuhang, Hu, Jiaxin, Gu, Shengrong, Tai, Ying, Yi, Zili
Abstract
Progress in general-purpose video editing depends on constructing large-scale paired supervision and effectively adapting video-generation backbones to instruction-driven editing. Unlike video generation, video editing must execute a requested transformation while preserving unrelated subjects, scene structure, motion, and temporal continuity. We present VideoX-Qwen, an integrated data-construction and model-training framework for general instruction-based video editing. Our scalable production pipeline organizes specialized generation and understanding models into complementary routes for addition, removal, replacement, and attribute editing, followed by quality screening and instruction enrichment. It produces more than 1.2 million directional video-editing records, including over 400,000 records in each major task group, with an automatic acceptance rate of 89%. The resulting corpus provides broad and structured coverage of common editing operations through a unified source-instruction-target interface. We further develop a unified Qwen-Wan editor that combines multimodal semantic conditioning with dense source-video latent guidance. A progressive image-video training strategy aligns the multimodal instruction interface, adapts the video generator to source-conditioned editing, and refines output quality with selected high-resolution data. In a 100-example comparison with UniVideo and Kling O1, VideoX-Qwen achieves the best mean result on nine of eleven reported metrics, including instruction following, editing quality, content preservation, structural and perceptual similarity, and video-distribution quality. Together, the large-scale data-production system and unified training framework provide a practical foundation for more capable instruction-driven video editing.
Chinese Translation
通用视频编辑的进展依赖于构建大规模成对监督数据,以及将视频生成骨干网络有效适配到指令驱动的编辑任务。与视频生成不同,视频编辑必须在执行所要求的变换的同时,保留无关主体、场景结构、运动和时间连续性。我们提出VideoX-Qwen,一个面向通用指令式视频编辑的集数据构建与模型训练于一体的框架。我们的可扩展生产流水线将专门的生成与理解模型组织为互补路径,覆盖添加、移除、替换和属性编辑等任务,并进行质量筛选与指令丰富。该流水线产生了超过120万条有向视频编辑记录,其中每个主要任务类别均包含超过40万条记录,自动接受率为89%。所得语料库通过统一的源-指令-目标接口,对常见编辑操作提供了广泛而结构化的覆盖。我们进一步开发了一个统一的Qwen-Wan编辑器,将多模态语义条件与稠密的源视频隐变量引导相结合。渐进式的图像-视频训练策略对齐了多模态指令接口,使视频生成器适配源条件编辑,并利用精选的高分辨率数据优化输出质量。在与UniVideo和Kling O1的100个样例对比中,VideoX-Qwen在11项报告指标中的9项上取得了最佳平均结果,包括指令遵循、编辑质量、内容保留、结构与感知相似性以及视频分布质量。大规模数据生产系统与统一训练框架共同为更强大的指令驱动视频编辑提供了切实可行的基础。
cs.AI / 57 / 2609.26029

CQ4OE: A benchmark for assessing LLM-assisted ontology generation from competency questions

Li, Jiayi, Wang, Ziyuan, Garijo, Daniel, Poveda-Villalón, María
Abstract
Ontology generation from Competency Questions (CQs) is a central yet labor-intensive phase of Ontology Engineering. While large language models (LLMs) offer promising automation capabilities, current evaluations remain fragmented. Task formulations are heterogeneous, gold standards often lack fine-grained CQ provenance, metrics conflate lexical overlap with structural and logical adequacy, and reference ontologies are not always explicitly designed around the evaluation CQs. Here, we address these limitations with CQ4OE, a benchmark for the systematic and reproducible evaluation of LLM-based ontology generation from CQs. For each ontology in the benchmark, we build a CQ-driven gold OWL ontology with explicit provenance linking each CQ to the classes, properties, and axioms required to answer it. From this resource, we define two complementary evaluation tasks. CQ2Term supports term-level evaluation of CQ-specific class and property prediction over 99 CQs, and CQ2Onto supports ontology-level evaluation over 118 CQs, including hierarchy, property modeling, and axiom-level structure. We demonstrate CQ4OE with experiments using nine LLMs under zero-shot, iterative, and multi-agent generation strategies, showing that LLMs recover explicit vocabulary terms more reliably than creating ontologies, particularly in property modeling, hierarchy construction, and axiom generation.
cs.AI / 58 / 2609.26046

Canonical locks that encode part-whole hierarchies

Modi, Rajat, Rawat, Yogesh Singh
Abstract
One of the challenges in representational learning is how to encode part-whole hierarchies in a neural net. Prior works rely on flattening tree-like structures into string-like sequences and training a sequence-to-sequence model via autoregression. While such a representation works for parse-trees in NLP, it is not entirely clear how to make it work for images. Thus, we propose a geometric primitive called canonical locks. The key idea is that parts/wholes can be modelled as higher-dimensional vectors ($d \geq 4$), and information can be encoded in their relative phase differences. Inductively, the net consists of positionally-bound bottom-up and top-down neural fields, which drive each other to achieve a state of thermal equilibrium. Additionally, we show the existence of a few symmetrical configurations in the net. The computational iterations taken to break these symmetries depend on the angle between parts/wholes arranged on a disk (or more precisely a ring) in higher dimensions. It also appears to have connections to the psychological phenomenon of mental rotation.
cs.AI / 59 / 2609.26048

FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents

FIRE:面向可靠语言模型智能体的故障知情运行时工程
Agarwal, Nikita, Jain, Nivedit
Abstract
Language-model agents often reach a working solution and then fail to consistently deliver it. We study runtime policies: targeted natural-language instructions and action denials applied by the agent harness at states that preceded observed failures, without changing model weights or the user prompt. With this, keeping capability constant, we observe a meaningful unlock in delivered reliability. Across the complete 87-task Terminal-Bench 2.1 suite, with two attempts per task, policies increase repeated success (pass^2) in all three GPT-5.6 tiers: 50.6% to 54.0% for Luna, 55.2% to 60.9% for Terra, and 64.4% to 73.6% for Sol. Sol's best-of-two success changes by 1.2 points while repeated success rises by 9.2, showing that policies chiefly convert reachable solutions into dependable delivery. We further cover 14 tasks under Terra's frozen portfolio. Policy-guided Terra reaches 71.4%, compared with 64.3% for unassisted Sol, at about half the cost, demonstrating how engineering around models could unlock dependability for a use case. To isolate the mechanism we run a randomized five-arm experiment: real policies reach 61% on eligible tasks, versus 39% without a policy, 36% with a timing-matched sham, and 39 to 43% with generic verification or reconsideration. The intended corrective behavior appears in 22 of 24 coded policy attempts, against at most 14 in any other arm. Runtime policies are therefore a practical reliability layer: they make capabilities an agent already possesses substantially more repeatable.
Chinese Translation
语言模型智能体往往能够找到可行的解决方案,却难以稳定地交付该方案。我们研究了运行时策略:由智能体框架在观测到失败的先前状态处施加的针对性自然语言指令和动作拒绝,而不改变模型权重或用户提示词。借此,在保持能力不变的情况下,我们观察到交付可靠性获得了显著提升。在完整的87项任务的Terminal-Bench 2.1测试集上(每项任务尝试两次),策略使三个GPT-5.6档位的重复成功率(pass^2)均有所提升:Luna从50.6%提升至54.0%,Terra从55.2%提升至60.9%,Sol从64.4%提升至73.6%。Sol的两选一最优成功率仅变化1.2个百分点,而重复成功率提升了9.2个百分点,表明策略主要是将可达成的解决方案转化为可靠的交付。我们进一步在Terra的固定配置下测试了14项任务:策略引导的Terra达到71.4%,而未加辅助的Sol为64.3%,且成本约为后者的一半,这展示了围绕模型进行工程化改造如何为具体应用场景解锁可靠性。为分离其作用机制,我们开展了一项随机化五臂实验:真实策略在符合条件的任务上达到61%,而无策略时为39%,时间匹配的虚假干预为36%,通用验证或重新考虑策略为39%至43%。在24次编码的策略尝试中有22次出现了预期的纠正行为,而其他任何组中至多只有14次。因此,运行时策略是一种实用的可靠性层:它们使智能体已具备的能力变得显著更可重复。
cs.AI / 60 / 2609.26059

Adversarial Course-of-Action Generation: Game-Theoretic Multi-Agent Algorithms for COA matching & COA generation

对抗性行动方案生成:基于博弈论多智能体算法的COA匹配与COA生成
Vidra, Natan, Kapanova, Alina, Kanhai, Arun, Setty, Spurthi
Abstract
Course-of-action (COA) generation is a distributed planning problem: a system must propose structured candidate actions, evaluate them against an adversarial response, and surface options that remain tactically coherent under changing conditions. We present COA-Bench, a small offline benchmark and reproducibility artifact for comparing COA generation policies through self-play. Following the BattleCOA terminology, we reserve COA matching for asset-effect matching and COA generation for course-of-action generation; the present artifact does not implement either DecisionFunction directly. Instead, it represents COAs as typed action chains with conditional branches, assigns a synthetic COA quality score, compares opposing COAs with a BLUE-vs-RED advantage score and Nash-gap distance, and scores doctrinal coherence with an FM 3-0-inspired heuristic rubric. Across 50 synthetic scenarios spanning five operational templates, a sampled best-response policy that draws eight RED candidates reduces BLUE advantage from .516 to .485 and BLUE wargame win rate from .920 to .820; a two-stage multi-agent council with five BLUE proposer agents, RED-team adjudication, and critique-driven revision obtains .509 BLUE advantage and .820 BLUE win rate. We also identify and fix a benchmark-design issue in which scenario framing was stored as metadata but had no effect on generated COA content. COA-Bench is not an operational battle-management system and uses no real, classified, proprietary, or human-subject data. The contribution is an inspectable evaluation harness, preliminary benchmark evidence, and lessons for building auditable agentic planning artifacts.
Chinese Translation
行动方案生成是一个分布式规划问题:系统必须提出结构化的候选行动,针对对手的应对进行评估,并筛选出在条件变化下仍保持战术连贯性的选项。我们提出了COA-Bench,一个用于通过自博弈比较COA生成策略的小型离线基准测试和可复现性工件。遵循BattleCOA的术语,我们将COA匹配专指资产-效果匹配,将COA生成专指行动方案生成;本工件并未直接实现这两种DecisionFunction。相反,它将COA表示为带有条件分支的类型化行动链,赋予合成的COA质量评分,使用蓝方对红方的优势评分和纳什差距距离来比较对抗性COA,并采用受FM 3-0启发的启发式评分准则来评估条令连贯性。在涵盖五种作战模板的50个合成场景中,一种抽取八个红方候选方案的采样最优响应策略将蓝方优势从0.516降至0.485,蓝方兵棋推演胜率从0.920降至0.820;一个包含五个蓝方提案智能体、红方裁判以及基于批评的修订的两阶段多智能体委员会取得了0.509的蓝方优势和0.820的蓝方胜率。我们还识别并修复了一个基准测试设计问题,即场景框架被存储为元数据但对生成的COA内容没有影响。COA-Bench不是作战管理系统,不使用任何真实、涉密、专有或涉及人类受试者的数据。其贡献在于提供了一个可检查的评估框架、初步的基准测试证据,以及构建可审计的智能体规划工件的经验教训。
cs.AI / 61 / 2609.26060

ChainUQ: Reasoning Consistency-Aware Uncertainty Quantification for Large Language Models

Yu, Dahai, Xu, Rongchao, Jiang, Lin, Li, Ximiao, Wang, Guang
Abstract
While large language models (LLMs) exhibit impressive reasoning capabilities, response-level confidence may remain unreliable when intermediate claims conflict with the final conclusion. Therefore, effective uncertainty quantification (UQ) is required to capture logical inconsistencies within the reasoning chain, not just the correctness of the final output. Current approaches have two major limitations: (1) their reliance on token-level probabilities fails to capture reasoning consistency, and (2) they lack mechanisms to dynamically calibrate confidence using the structural logic of the generated chain. To advance existing research, we introduce ChainUQ, a reasoning consistency-aware uncertainty quantification framework for LLMs. ChainUQ consists of two key technical components: an alignment-aware lightweight UQ module that estimates a raw intrinsic model confidence score from frozen features aligned to the final conclusion, and a reasoning consistency-aware calibrator that refines this score using reasoning-chain consistency evidence. Evaluations across diverse in-distribution and out-of-distribution benchmarks show that ChainUQ consistently improves response-level uncertainty estimation, achieving an average 3.1% relative gain in AUROC and up to 45.0% relative reduction in ECE, and can be directly transferred to new settings without additional fine-tuning.
cs.AI / 62 / 2609.26069

RankCert: When Can Simulated Learners Safely Select an AI Tutor? Robust Decision Certification Under Structural Uncertainty

RankCert:模拟学习者何时能安全选择AI导师?结构不确定性下的鲁棒决策认证
Kadir, Nizam
Abstract
Simulation-based tutor selection can be unstable when predictively adequate learner models imply different policy rankings. RankCert certifies one of eight equal-budget tutoring policies only when model-averaged utility, probability-best, posterior regret, cross-domain rank, family coverage, and leave-one-domain-out and leave-one-visible-family-out averages support the same candidate; otherwise it abstains. We evaluated RankCert in 1,280 frozen held-out settings spanning five rotating held-out oracle families, 64 scenarios per family, and four cohort sizes. Calibration used a licensed, de-identified EdNet-KT1 derivative with 5,000 learners and 590,056 retained responses; all five family representatives passed the frozen adequacy gate. Minimum-domain mean pairwise top-1 agreement was 0.272917 (95% CI [0.253646, 0.293229]), showing substantial structural disagreement. Cohort-noise variance decreased from n = 30 to n = 300, while the structural family share remained nonzero. RankCert reduced total held-out decision loss relative to full-coverage point selection by 0.006605 normalized-outcome units (95% CI [0.004859, 0.008407]). At comparable coverage, however, it did not reduce selective risk relative to a confidence-gated point certificate (difference -0.000213; 95% CI [-0.003238, 0.002384]; Holm p = 0.929654). Certification occurred in 3.75% of settings and only in stable scenarios; RankCert abstained in every ambiguous, misspecified, and structural-conflict setting. "Safe" denotes only benchmark-scoped decision certification under the declared utility and uncertainty set; no human-learning, causal, deployment-effectiveness, or general-safety claim is made.
Chinese Translation
当预测上充分的学习者模型隐含不同的策略排序时,基于仿真的导师选择可能不稳定。RankCert 仅在模型平均效用、概率最优、后验遗憾、跨域排序、族覆盖、留一域(leave-one-domain-out)与留一可见族(leave-one-visible-family-out)平均值均支持同一候选策略时,才会对八种等预算辅导策略之一进行认证;否则将选择弃权。我们在 1,280 个冻结的保留测试设置上评估了 RankCert,涵盖五个轮换的保留测试 oracle 族、每族 64 个场景以及四种群体规模。校准使用了一个经过许可、去标识化的 EdNet-KT1 衍生数据集,包含 5,000 名学习者和 590,056 条保留的作答记录;所有五个族的代表模型均通过了冻结的充分性门槛。最小域平均成对 top-1 一致性为 0.272917(95% CI [0.253646, 0.293229]),表明存在显著的结构性分歧。群体噪声方差从 n = 30 到 n = 300 逐渐减小,而结构性族占比始终保持非零。相对于全覆盖点选择,RankCert 将总保留决策损失降低了 0.006605 个归一化结果单位(95% CI [0.004859, 0.008407])。然而,在可比覆盖率下,相对于置信度门控的点认证方法,它并未降低选择性风险(差值 -0.000213;95% CI [-0.003238, 0.002384];Holm p = 0.929654)。认证仅发生在 3.75% 的设置中,且仅出现在稳定场景中;RankCert 在所有模糊、错误设定和结构冲突的设置中均选择弃权。“安全”一词仅指在声明的效用与不确定性集合下的基准范围内决策认证;本文不作出任何关于人类学习、因果效应、部署有效性或一般安全性的声明。
cs.AI / 63 / 2609.26076

Selection-Invariant Communication Compilers for Privacy-Aware Multi-Agent LLM Workflows

Xu, Jinghan, Fan, Longze, Wang, Zeyuan, Li, Xinjin, Liu, Hankai
Abstract
Structured multi-agent workflows exchange intermediate messages whose content and form can reveal private state even when the final output is safe. We identify selection-channel leakage: after authorization fixes what may be released, a private-state-aware choice among semantically valid realizations creates an additional inference channel. We introduce the selection-invariant communication compiler(SICC), which constrains this post-authorization representation kernel rather than prescribing templates. Any deterministic or independently public-randomized generator satisfying the invariant is valid; requirement-indexed canonical forms are one auditable implementation. We prove a compositional communication-layer guarantee: authorization, public-only form generation, and a dependency-safe utility gate make the emitted transcript reveal no information beyond the complete authorized view. Private-state-aware selection remains vulnerable after surface-disjoint and length-matched controls. Across 132 AgentLeak communication replays and 100 executable LangGraph tasks, deterministic SICC retains complete protocol utility without a positive excess-gain signal; independent public randomization preserves the same result in AgentLeak and 480 controlled cases.
cs.AI / 64 / 2609.26087

The Architect, the Adversary, and the Judge: Closed-Loop Generation of Standards-Aligned Assessment Items at Scale

Chen, Wenhui, Lin, Ziyao, Chen, Jianlin, Long, Peiji, Vong, Chi Man
Abstract
We present CLAIM, a production pipeline for K-12 assessment-item generation coupling a two-stage generate-then-attack protocol (the model drafts as a "curriculum architect", then re-enters the same conversation as a hostile adversarial reviewer), bi-directional few-shot conditioning on accepted and rejected items, the latter carrying the evaluator's diagnosis, and a knowledge dictionary of 44,844 error-correction rules mined from that feedback and retrieved per standard and item type. Across 43,227 scored items over 755 Common Core ELA standards, three item types, and ten LLMs, the pipeline reaches a 97.8% expert-evaluator pass rate on a 9,074-item production run. We then ask what that rate certifies. Re-scoring a stratified sample with three judges from other vendors, blind to the deployed verdict, reproduces the format ordering under every judge and recovers a larger open-set deficit than the deployed evaluator does; but agreement on the accept/reject binary is weak at production prevalence (kappa about 0.13), and the judges agree with each other no better. The level is therefore judge-relative, and with no student-response data our quality evidence is evaluator-judged throughout. The corpus also exposes a robust asymmetry. Multiple-choice and multiple-select generation saturate at 98% or above for both frontier models under a dozen static rules, whereas fill-in-the-blank generation is capability-tiered (82.8-96.7% across five models under a matched rule set, standards, and judge) and plateaus under prompt-only optimization, with error mass shifting between answer-key over-inclusion and omission as rules accumulate. We analyze this as open-set boundary determination, a task autoregressive decoders are structurally ill-equipped to solve, and show the asymmetry recurring when the evaluator itself is distilled: fail-recall rises from 8% to 63% while F1 saturates at 0.25.
cs.AI / 65 / 2609.26095

FusionMMT: A Unified Multimodal and Multitask Learning Framework for Nuclear Fusion

Chen, Qiang, Wang, Xiao, Yang, Qingquan, Si, Hao, Yan, Zikang, Chen, Meiwen, Xu, Guosheng, Tang, Jin
Abstract
With the growing global demand for energy, nuclear fusion has emerged as a promising direction for future clean energy. Tokamaks represent one of the leading approaches to magnetic-confinement fusion. Achieving high-performance, long-pulse, and steady-state operation requires effective diagnosis of plasma states. However, existing intelligent diagnostic methods are largely limited to either multimodal single-task or unimodal multitask learning, while a unified multimodal multitask learning framework remains underexplored. To address this gap, we construct EAST-VTD640, a multimodal multitask dataset that integrates vision and time-series diagnostics from 640 EAST shots for disruption prediction, edge-localized mode (ELM) recognition, and H98 regression. On this basis, we present FusionMMT, the first unified multimodal multitask framework for intelligent tokamak plasma diagnostics. FusionMMT employs multi-scale, time-aware, and variable-aware modeling to handle heterogeneous sampling rates and the high computational cost of high-frequency sequences. It further combines task-adaptive multimodal fusion with progressive multitask optimization to learn shared and task-specific representations while mitigating cross-task conflicts and optimization imbalance. Extensive experiments on EAST-VTD640 show that FusionMMT outperforms representative multimodal multitask methods across disruption prediction, ELM recognition, and H98 regression. The source code will be released on https://github.com/Event-AHU/OpenFusion
cs.AI / 66 / 2609.26105

Neoadjuvant chemotherapy response prediction using pretreatment diffusion and contrast-enhanced magnetic resonance imaging with clinical variables

基于治疗前扩散与对比增强磁共振成像及临床变量的新辅助化疗反应预测
Marcos, Pablo García, González, Paula Puerta, Lorenzo, Guillermo, Gómez, Héctor, del Camino, Covadonga, Rodríguez, Adán, Peláez, Ignacio, Rio-Alvarez, Angel, González, Víctor M.
Abstract
Prediction of pathological complete response before neoadjuvant chemotherapy may facilitate more tailored therapeutic planning for breast cancer patients. This work proposes a deep-learning model for pretreatment data only, combining apparent diffusion coefficient maps, dynamic contrast-enhanced magnetic resonance imaging, and clinical variables. The study uses the public ACRIN 6698/I-SPY2 multicenter dataset. The architecture employs EfficientNet-B0 pretrained encoders for image feature extraction and late fusion with clinical information. Multiple clinical variables were evaluated, including age, race, histological type, HR/HER2 subtype, SBR grade, and maximum diameter. Only HR/HER2 subtype improved the average area under the receiver operating characteristic curve (AUC) and was retained in the final model. Using stratified five-fold cross-validation, standalone apparent diffusion coefficient maps achieved a mean AUC of 0.79, whereas dynamic contrast-enhanced magnetic resonance imaging achieved 0.74. Adding HR/HER2 subtype improved performance to 0.83 and 0.81, respectively. The final configuration, using both imaging modalities and HR/HER2 subtype, achieved an AUC of 0.86. These results support pretreatment multimodal learning for response prediction, although external validation is required before clinical use.
Chinese Translation
在新辅助化疗前预测病理完全缓解有助于为乳腺癌患者制定更加个体化的治疗方案。本研究提出了一种仅基于治疗前数据的深度学习模型,该模型结合了表观扩散系数(ADC)图、动态对比增强磁共振成像(DCE-MRI)以及临床变量。研究采用公开的ACRIN 6698/I-SPY2多中心数据集。模型架构采用预训练的EfficientNet-B0编码器进行图像特征提取,并通过后期融合(late fusion)方式与临床信息相结合。研究评估了多项临床变量,包括年龄、种族、组织学类型、HR/HER2分子亚型、SBR分级和最大肿瘤直径。结果显示,仅HR/HER2亚型能够提升受试者工作特征曲线下面积(AUC)的均值,因此在最终模型中予以保留。采用分层五折交叉验证,单独使用表观扩散系数图的平均AUC为0.79,而动态对比增强磁共振成像的AUC为0.74。加入HR/HER2亚型后,两者的性能分别提升至0.83和0.81。最终配置同时使用两种成像模态和HR/HER2亚型,达到了0.86的AUC。这些结果支持利用治疗前多模态学习进行疗效反应预测,但在临床应用之前仍需进行外部验证。
cs.AI / 67 / 2609.26106

Early Prediction of Pathological Complete Response to Neoadjuvant Chemotherapy Using Temporal Deep Learning on DWI

Marcos, Pablo García, Islam, Md. Tarequl, González, Paula Puerta, Lorenzo, Guillermo, Gómez, Héctor, del Camino, Covadonga, Rio-Alvarez, Angel, González, Víctor M.
Abstract
Early identification of non-responders to neoadjuvant chemotherapy (NACT) is crucial for timely treatment adaptation in breast cancer. However, many existing predictive models rely on multiparametric magnetic resonance imaging (MRI), late treatment time points, or extensive clinical data, which limits their applicability. This study proposes a deep learning framework for early prediction of pathological complete response (pCR) using only diffusion-weighted MRI (DW-MRI) acquired at baseline and after the first NACT cycle. This framework feeds cropped tumor-centered patches to an EfficientNet-based temporal model that directly learns tumor shape and local tissue characteristics without explicit radiomic feature engineering. The model, trained with 10-fold cross-validation, achieved an area under the receiver operating characteristic curve (AUC) of 0.90 for pCR prediction after one cycle, providing actionable information after a single treatment cycle while avoiding gadolinium administration and reducing dependence on heterogeneous clinical data. By focusing on the baseline-to-first-cycle window instead of later stages, the approach supports earlier escalation or de-escalation of NACT, and its exclusive reliance on DW-MRI facilitates protocol standardization, multi-centre deployment and privacy-preserving data sharing. These results demonstrate that DW-MRI-based deep learning on tumor-centered patches constitutes a minimally invasive, clinically deployable strategy for early pCR prediction, with direct implications for personalized treatment adaptation in neoadjuvant breast cancer therapy.
cs.AI / 68 / 2609.26121

DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents

DTOC:面向AI智能体自适应上下文管理的动态工具输出压缩
Chaturvedi, Abhay, Bhattacharya, Shreya, Gopalkrishnan, Rashmika, van der Putten, Peter
Abstract
As agent capabilities have grown, practical limitations increasingly stem from constrained context windows rather than model capacity. Common strategies, such as truncation, heuristic aging, and lossy summarization, may discard useful information or introduce hallucination risk. To address these challenges, we propose Dynamic Tool Output Compression (DTOC), a framework for scalable context management in LLM-based agents that models context updates as explicit and reversible operations within the agent reasoning loop. DTOC retains full tool outputs in external memory while inserting compact placeholders into the active context, enabling selective reconstruction when needed. We formalize the DTOC mechanism, integrate it into a ReAct-style agent architecture, and provide a production-oriented implementation supporting on-demand restoration of compressed outputs. Experiments on DeepSWE reveal model-dependent effects: for responsive models (Sonnet 4.6, GPT-5.4), DTOC reduces input tokens (10.3 and 12.7%) and agent steps (2.4 and 32.3%), while increasing solve rates (2.5 and 1.5 times higher) and lowering cost per solved task (3 and 3.5 times lower cost per solved task). For the other models results are more mixed, with GPT-5.5 doubling solve rate and halving cost, but no impact on solve rate and negative impact on cost for the other models. Ablation results show reversibility is critical: disable-only compression variants degraded performance, while full DTOC recovered baseline accuracy at substantially lower context cost. These findings indicate that explicit, reversible context management can improve the efficiency of long-horizon agent reasoning without degrading task performance.
Chinese Translation
随着智能体能力的不断增强,实际限制越来越多地来自受约束的上下文窗口而非模型容量。常见的策略,如截断、启发式老化和有损摘要,可能会丢弃有用信息或引入幻觉风险。为应对这些挑战,我们提出了动态工具输出压缩(DTOC),这是一个面向基于大语言模型(LLM)的智能体的可扩展上下文管理框架,它将上下文更新建模为智能体推理循环中显式且可逆的操作。DTOC 将完整的工具输出保留在外部存储中,同时在活动上下文中插入紧凑的占位符,从而支持在需要时进行选择性重构。我们对 DTOC 机制进行了形式化,将其集成到 ReAct 风格的智能体架构中,并提供了一个面向生产的实现,支持按需恢复被压缩的输出。在 DeepSWE 上的实验揭示了模型相关的效应:对于响应迅速的模型(Sonnet 4.6、GPT-5.4),DTOC 降低了输入 token 数(分别降低 10.3% 和 12.7%)和智能体步骤数(分别降低 2.4% 和 32.3%),同时提高了求解率(分别提高 2.5 倍和 1.5 倍),并降低了每个已解决任务的成本(分别降低 3 倍和 3.5 倍)。对于其他模型,结果则更为混杂:GPT-5.5 的求解率翻倍且成本减半,但其余模型在求解率上没有影响,且成本出现负面影响。消融实验表明可逆性至关重要:仅禁用式压缩变体导致性能下降,而完整的 DTOC 以显著更低的上下文成本恢复了基线准确率。这些发现表明,显式且可逆的上下文管理可以在不降低任务性能的前提下提升长程智能体推理的效率。
cs.AI / 69 / 2609.26123

FairMon: A Tool for Monitoring and Visualizing Algorithmic Fairness

FairMon:一种用于监控与可视化算法公平性的工具
Baumeister, Jan, Finkbeiner, Bernd, Krsmanovic, Vladimir, Scheerer, Frederik, Siber, Julian, Wagenpfeil, Tobias
Abstract
Runtime monitoring has recently been proposed as a rigorous method for analyzing algorithmic fairness of autonomous decision systems used in critical scenarios such as credit lending, job application, and the criminal justice system. Prior work has shown that runtime monitoring, in principle, can be an effective technique for establishing the kind of human oversight required by legislation such as the EU Artificial Intelligence Act. In practice, the available monitoring tools have not been developed with this application in mind and display several critical shortcomings in these scenarios. In this paper, we present FairMon, a runtime monitoring tool tailored to fairness analysis of high-stakes decision systems. FairMon uses RTLola as a flexible specification language for monitors, which we have extended with conditional probability operators that allow for concise descriptions of algorithmic fairness properties. The tool also features a real-time visualization of intermediary values, enabling human insight into the dynamics of the monitored system.
Chinese Translation
运行时监控(runtime monitoring)最近被提出作为一种严格的方法,用于分析在信贷发放、求职申请和刑事司法等关键场景中使用的自主决策系统的算法公平性。已有研究表明,运行时监控在原则上可以成为一种有效的技术,用以建立诸如《欧盟人工智能法案》(EU Artificial Intelligence Act)等立法所要求的人类监督。然而在实践中,现有的监控工具并非针对这一应用场景开发,在这些场景中存在若干关键缺陷。本文提出了FairMon,一种专为高风险决策系统公平性分析量身定制的运行时监控工具。FairMon使用RTLola作为监控器的灵活规范语言,并对其进行了扩展,加入了条件概率算子,从而能够简洁地描述算法公平性属性。该工具还支持对中间值的实时可视化,使人类能够洞察被监控系统的动态行为。
cs.AI / 70 / 2609.26124

MAC-RRG: Iterative Multi-Agent Collaboration for X-ray Radiology Report Generation

Wang, Futian, Qiao, Yuhan, Wang, Xiao, Xu, Dan, Li, Yuehang, Guo, Zhixiang, Wang, Yaowei, Tang, Jin
Abstract
Despite the remarkable progress of LLM-based and knowledge graph-augmented Radiology Report Generation (RRG) methods, existing techniques still suffer from inherent defects. Conventional LLM-only models lack structured medical prior knowledge, resulting in frequent medical hallucinations and low diagnostic interpretability. Current knowledge graph-enhanced schemes adopt static one-round knowledge fusion with single-source knowledge, incapable of dynamic knowledge updating according to generation feedback. This paper proposes a novel Multi-Agent Collaborative iterative framework for X-ray Radiology Report Generation, termed MAC-RRG. Inspired by multi-agent technology, our framework constructs a closed-loop optimization paradigm based on task decoupling and collaborative reasoning. Specifically, the framework first generates a preliminary radiology report from input X-ray images via a vision encoder and a basic LLM. Subsequently, a multimodal knowledge graph (MM-KG) agent mines structured disease correlation and anatomical knowledge from medical knowledge graphs, while an auxiliary knowledge agent extracts unstructured domain knowledge from public medical databases. The multi-source knowledge acquired by dual agents is fused and embedded to guide the LLM in iteratively refining the initial report. Extensive quantitative and qualitative experiments on mainstream X-ray RRG datasets, including IU X-ray, MIMIC, and CheXpert Plus, fully verify the superiority of our proposed method. The source code and pre-trained models have been released on https://github.com/Event-AHU/Medical_Image_Analysis
cs.AI / 71 / 2609.26125

When Big Data Becomes a Curse: Spatial Heterogeneity and the Limits of Learning from Passive Acoustic Monitoring Data

Spadon, Gabriel, Renaud, Wayne, Aravindan, Priyanka
Abstract
Passive Acoustic Monitoring produces large archives whose recordings are clustered by deployment, season, station identifier, and acquisition configuration. We analyze 908,072 AIS-labeled 679-second recordings from 38 deployments, 20 Atlantic Canadian station identifiers, and 21 receiver positions. The AIS-contact prior varies by more than 200-fold, and per-deployment screening distributions require local interpretation. Bidirectional cross-season transfer over the 18 station identifiers observed in both seasons predicts station identity above the 5.56% uniform-chance level, with balanced accuracy of 15.9% for AIS-contact and 16.4% for no-AIS-contact recordings. The same descriptors predict the two hydrophone models at 74.1% and 85.6% balanced accuracy, respectively, but hydrophone model is strongly confounded with season and other deployment-level acquisition differences. On a retrospectively screened and capped benchmark of 54m507 recordings, repeated station-grouped holdout yields an ROC-AUC of 0.612 with a station-bootstrap 95% interval of 0.582 to 0.648, compared with 0.661 under a random-window diagnostic. Their paired difference is 0.049 (0.040 to 0.056). Removing raw energy changes unseen-station ROC-AUC from 0.612 to 0.603, while an exploratory training-station scale analysis is non-monotonic. These results show that random-window validation overstates transfer to unseen station identifiers in this corpus. They support dependence-aware validation and broader independent spatial sampling, while spatial-expert models remain a hypothesis rather than an established remedy.
cs.AI / 72 / 2609.26126

The Cost of Conservation: Coordination-Memory Laws for Exact-Support Generation

守恒的代价:精确支撑生成中的协同—记忆定律
Zhang, Zhen, Alanwar, Amr
Abstract
Many AI systems make decisions locally, even when every realized output must obey an additive conservation law, such as selecting exactly a fixed number of items. This constraint can be statistically invisible: small subsets of a balanced fixed-budget output look increasingly independent, yet communication-free coordinate-parallel generation needs exponentially many pre-shared plans, while a sequential exact sampler needs only logarithmic memory. We study product measures conditioned on additive conservation laws in the intermediate regime where one plan is selected before a fixed-order pass, every plan is a bounded-state stochastic executor whose support is entirely legal, and the mixture of plan laws approximates the target distribution in total variation. Our main result identifies the optimal asymptotic selector rate, up to constant factors, with the killed spectral profile of the conservation-difference walk. The resulting coordination cost decreases as an inverse power of live-state width, with an exponent determined by intrinsic conservation rank rather than alphabet size; the law extends to noncentral budgets and heterogeneous local scores. The converse is driven by a state-versus-resource-sum obstruction, while a rate-matching construction compiles discrepancy control into exact-support finite-state plans. Complementary results characterize block-parallel plan complexity and the benefit of programmable output order. Together, these results show precisely how online memory substitutes for front-loaded coordination in exact-support generation.
Chinese Translation
许多AI系统在局部做出决策,即使每个实际输出都必须满足一个可加守恒律,例如恰好选取固定数量的项目。这一约束在统计上可能是不可见的:平衡的固定预算输出的小子集看起来越来越独立,然而无需通信的坐标并行生成需要指数级多的预共享计划,而顺序的精确采样器只需要对数级内存。我们研究在中间区域中受可加守恒律约束的乘积测度:在固定顺序扫描之前先选定一个计划,每个计划是一个支撑完全合法的有界状态随机执行器,且计划规律的混合在总变差距离下逼近目标分布。我们的主要结果表明,最优渐近选择器速率(相差常数因子)可由守恒差随机游走的被杀谱剖面刻画。由此得到的协同代价随活跃状态宽度的逆幂而下降,其指数由内在守恒秩而非字母表大小决定;该定律可推广到非中心预算和异构局部评分。其逆命题由状态—资源和的障碍驱动,而一个速率匹配的构造将偏差控制编译为精确支撑的有限状态计划。互补结果刻画了块并行计划复杂度以及可编程输出顺序带来的收益。这些结果共同精确地展示了在精确支撑生成中在线记忆如何替代前置协同。
cs.AI / 73 / 2609.26135

VACS: Value-Aligned Compositional Shielding for Multi-Agent Reasoning

Zhang, Yiyao, Goel, Diksha, Ahmad, Hussain, Huang, Shixun, Shen, Jun
Abstract
Multi-agent reasoning systems in high-stakes domains must be both accurate and safe, yet agents often follow heterogeneous value priorities (e.g., rigor, conciseness, safety), causing conflicting recommendations. Existing methods do not jointly provide: (i) principled inference of each agent's implicit values from behavior, (ii) compositional formal safety guarantees without full online communication, and (iii) value-aware conflict resolution with faithful explanations. We present VACS (Value-Aligned Compositional Shielding), a four-layer framework addressing all three. Layer 1 learns value-dimension rewards from pairwise preferences using Bradley-Terry modeling and infers per-agent value weights via deep MaxEnt IRL. Layer 2 encodes value constraints in a Lean-inspired DSL and synthesizes compositional assume-guarantee shields for runtime safety. Layer 3 resolves disagreement through nucleolus-based credit allocation and Hamiltonian consensus optimization under long-term value constraints. Layer 4 extracts a critical reasoning path from co-state sensitivities and generates formally grounded natural-language explanations. Our contribution is primarily a unified systems design with formalized interfaces and operational guarantees at the verifier-constrained decision level, rather than a complete end-to-end formal proof of all language-model internals. In controlled proof-of-concept evaluations with role-conditioned agent panels on NEJM-AI QA, MathInstruct-Subset, and a cybersecurity incident-response benchmark (CyberSec-Eval), VACS outperforms strong baselines in accuracy (85.4%, 95.0%, and 90.0%) while reducing logical inconsistency rates to near zero.
cs.AI / 74 / 2609.26144

When Verifiers Vote Backwards under Verdict Substitution: Signed Pivotal Value in Correlated Self-Consistency

当验证者在裁决替换下反向投票时:相关性自洽性中的带符号关键票价值
Shu, Yang
Abstract
Replacing one ballot can change a majority decision only on queries decided by a single vote; this structural fact requires no independence assumption. We study the sign of that change using a labeled, verdict-style intervention: one correctness signal replaces one correctness-indicator ballot in $k{=}7$ self-consistency panels. This diagnostic intervention is not identical to deployed answer-identity plurality. A primary MATH-500 experiment ($n{=}570$) gives a different-model verifier a $+24.2$pp pivotal gain, whereas a role-reversed configuration gives $-11.2$pp; an exploratory code stress test (14 pivotal rows across 9 tasks) gives $-24.5$pp. An exact signed-gain decomposition accounts for all observed signs through the verifier's state-specific accuracy and the composition of the two one-vote tally states, rather than global accuracy or model provenance. Same-source signals lose accuracy on the pivotal stratum (65$\to$44\% in the primary configuration), while error correlations provide a descriptive error-association diagnostic. Controlled degradation and a $k\in\{3,5,7\}$ subset sensitivity analysis probe the stability of the observed pattern around this accounting. Under the evaluated ties-incorrect answer-identity plurality analysis, the structural zero and strong-verifier benefit persist, but the role-reversed harm attenuates to $-0.9$pp and is not significant. The results therefore establish harmful verdict substitution, not universally harmful deployed plurality, and motivate a testable but unverified hypothesis for negative process-reward-model weights.
Chinese Translation
只有当查询的结果仅由一票之差决定时,替换一张选票才可能改变多数决策;这一结构性事实无需任何独立性假设。我们利用一种带标签的、裁决式的干预来研究该变化的符号:在 $k{=}7$ 的自洽性面板中,用一个正确性信号替换一张正确性指示选票。这种诊断性干预与部署中的答案一致性多数表决并不等同。一项主要的 MATH-500 实验($n{=}570$)显示,不同模型的验证者可获得 $+24.2$ 个百分点的关键票增益,而角色反转配置则产生 $-11.2$ 个百分点;一项探索性的代码压力测试(9 个任务中 14 个关键票样本行)产生 $-24.5$ 个百分点。一种精确的带符号增益分解表明,所有观测到的符号均可通过验证者在特定状态下的准确率以及两个单票计票状态的组合来解释,而非全局准确率或模型来源。同源信号在关键票分层上准确率下降(主要配置中从 65% 降至 44%),而误差相关性则提供了一种描述性的误差关联诊断。受控退化实验以及 $k\in\{3,5,7\}$ 的子集敏感性分析检验了该观测模式在此分解框架周围的稳定性。在所评估的平局计错、答案一致性多数表决分析下,结构零点与强验证者的收益持续存在,但角色反转造成的损害减弱至 $-0.9$ 个百分点且不显著。因此,本文的结果确立了有害的裁决替换现象,而非普遍有害的部署多数表决,并为负过程奖励模型权重提出了一个可检验但尚未验证的假设。
cs.AI / 75 / 2609.26145

Unanimity Without Persuasion: A Single Round of Debate Erases the Disagreement That Verification Needs

无需说服的一致性:一轮辩论即可消除验证所依赖的分歧
Shu, Yang
Abstract
A debate panel can become unanimous without becoming more correct. This is dangerous for downstream safeguards: a substituted verification ballot can change only narrow-margin votes, while richer arbiters lose disagreement as a natural targeting signal. We show that one debate round can erase that resource without requiring persuasion. Tracking a heterogeneous 7-judge panel through a blind round and three debate rounds on 600 code-correctness candidates, unanimity on a fixed cohort jumps from 39.5\% to 95.2\% in round 1 (93.1\% of the total collapse), while accuracy moves by less than one point and 96.3\% of verdict flips follow the displayed peer majority. An execution-based verification ballot corrects 8 of 2,037 pre-debate candidate-substitution instances but changes zero in every later round; by round 3 every wrong decision is unanimous, erasing dissent that had flagged two-thirds of the panel's errors. Identical-cohort controls explain why: no-peer reconsideration reproduces 79.6\% of the collapse, real labels without reasoning reproduce 91.5\%, and random labels steer flips toward whatever they display; the full-debate condition adds 4.3 percentage points over labels only (clustered 95\% CI 0.7--8.1). The one-round collapse reproduces in two additional real runs and two fake-label seeds, remains under panel sizes 3--7, and appears in MATH-500. Parse failures concentrate on contested candidates ($p<0.001$), making attrition non-ignorable. The design implication is operational: verify before any second-pass evaluation or peer exposure, and never treat post-debate unanimity as independent evidence of reliability.
Chinese Translation
评审小组可以在未变得更正确的情况下达成一致意见。这对下游的安全保障机制是危险的:替代性的验证选票只能改变票数差距微小的投票结果,而更丰富的仲裁机制则会随着分歧的消失而失去天然的定位信号。我们证明,一轮辩论即可在不依赖说服的情况下消除这一资源。我们在600个代码正确性候选项上追踪一个异质性的7人评审小组,经历一轮盲评和三轮辩论后,固定候选群体上的一致性比例从39.5%跃升至第一轮的95.2%(占总降幅的93.1%),而准确率变动不足一个百分点,且96.3%的判定翻转跟随所展示的同伴多数意见。基于执行的验证选票在辩论前的2,037个候选替换实例中纠正了8个,但在其后的每一轮中纠正数均为零;到第3轮时,所有错误决策均已达成一致,曾经标记出小组三分之二错误的异议被彻底消除。同群体对照实验解释了原因:无同伴的重新评估重现了79.6%的一致性崩塌,不带推理的真实标签重现了91.5%,而随机标签则引导翻转向其展示的任何方向;完整辩论条件相对于仅提供标签的条件额外增加了4.3个百分点(聚类95%置信区间为0.7–8.1)。这一轮内崩塌在另外两次真实运行和两个假标签种子中均得到复现,在小组规模为3–7的情况下持续存在,并出现在MATH-500数据集上。解析失败集中于存在争议的候选项(p<0.001),使样本流失不可忽略。其设计启示具有操作层面意义:应在任何二次评估或同伴暴露之前进行验证,并且绝不应将辩论后的一致性视为可靠性的独立证据。
cs.AI / 76 / 2609.26157

Toward User-Mediated Self-Repair in Ubiquitous Robots Through Goal-Oriented Agentic AI

基于目标导向智能体AI(Agentic AI)的普适机器人用户自主修复研究
Frederiksen, Morten Roed
Abstract
Ubiquitous robotic systems often lack traditional visual interfaces, necessitating resilient natural language interaction for maintenance and repair tasks. This paper presents a goal oriented agentic AI architecture designed to enable non-expert users to perform technical repairs through situated dialogue. The framework utilizes a multi-layered approach that decouples high-level strategic planning from reactive conversational execution to transform unconstrained human instructions into a structured hierarchy of goals. We conducted a study involving twenty participants to evaluate the system's efficacy using a physical hardware testbed. The architecture achieved a 95\% task completion rate, and participants reported positive self-efficacy following real-time guidance that adapted to conversational diversions and linguistic variations. A comparative analysis with an online baseline revealed that the transition to a physical environment significantly decreased perceived social presence (p=.0005), and trust and competence, (p=.037), while the agentic framework remained robust throughout the interaction. These findings indicate that goal oriented agentic AI can support the sustainability of body-worn technologies by empowering users to perform critical maintenance in ubiquitous contexts.
Chinese Translation
普适机器人系统通常缺乏传统的视觉界面,因此在维护与修复任务中需要具备韧性的自然语言交互。本文提出了一种目标导向的智能体AI(Agentic AI)架构,旨在使非专业用户能够通过情境化对话完成技术性修复。该框架采用多层方法,将高层策略规划与响应式对话执行解耦,从而将不受约束的人类指令转化为结构化的目标层次体系。我们开展了一项包含二十名参与者的研究,借助物理硬件测试平台评估该系统的有效性。该架构实现了95%的任务完成率,且参与者在实时引导(可适应对话偏移和语言差异)后报告了积极的自我效能感。与在线基线的对比分析表明,向物理环境的转变显著降低了感知的社会临场感(p=.0005)以及信任度和胜任感(p=.037),而该智能体框架在整个交互过程中保持了稳健性。这些发现表明,目标导向的智能体AI能够赋能用户在普适环境中执行关键维护任务,从而支持可穿戴技术的可持续性。
cs.AI / 77 / 2609.26160

The Free-Recipe Limit: Every Recipe Effect Measures Which Premise of an Idealised Learner Broke

自由配方极限:每一种配方效应度量的都是理想化学习者的哪一个前提失效了
Chen, Wenhui, Chen, Jianlin, Lin, Ziyao, Vong, Chi Man
Abstract
Fix a corpus and send recipe search to infinity: try every order of the skills, every arrangement from blocked to interleaved, every composition, and keep the best. Two quantities decide what that search was worth: the diameter of the reachable set it explores, and the resolution at which anyone can tell two endpoints apart. Where the diameter falls below the resolution, no amount of search converts into a decision, and the signature is not an absence of winners but winners that do not survive re-running. We measure this recipe-search wall with 761 fine-tuning runs on 12 base models (0.5B-14B, three pretraining families) over competition-mathematics skills: base checkpoints, supervised fine-tuning under AdamW, exact-match scoring at k=4. Within one coherent domain at fixed volume the three classical freedoms average 0.010-0.021 against a 0.019 floor, and the largest contrast, 0.0619, clears a three-seed resolution and then reads +0.010 and -0.015 on two reruns. The departure with a systematic answer is coherence: halving one pooled corpus and letting the halves write answers under incompatible but equally correct conventions moves arrangement from capability to allocation between conventions, by two orders of magnitude over a same-convention control, and writing the convention into the input switches the phenomenon off. The switch replicates on a second pretraining family and survives an independent re-execution of its own protocol, with a re-execution spread (0.087) smaller than the resolution a search-selected order cell carries (0.144). Order itself is a transient whose sign crosses zero three times inside a single run. Volume, the one lever nobody calls a recipe, is the one that reliably pays. A public scorecard grades all 26 pre-registered claims: 18 supported, 5 failed, 2 untested, 1 mixed.
Chinese Translation
固定一个语料库并将配方搜索推向无穷:尝试技能的每一种排列顺序、从分块到交错的每一种安排方式、每一种组合方式,并保留最优者。有两个量决定这种搜索的价值:其探索的可达集合的直径,以及任何人区分两个端点所需的分辨率。当直径低于分辨率时,再多的搜索也无法转化为决策,其标志不是缺乏胜出者,而是胜出者无法在重跑中存续。我们通过在12个基座模型(0.5B–14B,三个预训练家族)上针对竞赛数学技能进行的761次微调运行来测量这道配方搜索之墙:基座检查点、AdamW下的监督微调、k=4的精确匹配评分。在同一连贯领域、固定数据量的条件下,三种经典自由度的平均效应为0.010–0.021,而分辨率下限为0.019;最大的对比度0.0619虽然超过了三种子分辨率,但在两次重跑中分别读出+0.010和-0.015。能够给出系统性答案的例外是连贯性:将一个合并语料库对半拆分,并让两半在互不兼容但同样正确的约定下书写答案,会使安排方式的作用从能力转变为约定之间的分配,其效应比同约定对照组高出两个数量级;而将约定写入输入则关闭了该现象。这一开关效应在第二个预训练家族上得到复现,并在其自身协议的独立重执行中得以存续,重执行离散度(0.087)小于搜索选出的顺序单元所承载的分辨率(0.144)。顺序本身是一个瞬态量,其符号在单次运行内三次穿越零。数据量——这个没人称之为配方的杠杆——却是唯一能可靠带来收益的。一份公开的记分卡对所有26项预注册声明进行了评级:18项获得支持,5项未通过,2项未经测试,1项结果混合。
cs.AI / 78 / 2609.26175

EADC: Evaluation of Advanced and Deep-level Compliance in Large Language Models

EADC:大语言模型高级与深层次合规性评估
Zhang, Yan, Li, Ruien, Peng, Yaoyao, Ren, Wanxin, Zhang, Yijia, Zhang, Wusheng, Yang, Guangwen
Abstract
Large Language Models (LLMs) have been used in various industries. However, ensuring their compliance with complex laws and regulatory frameworks remains a great challenge. Existing evaluation paradigms mainly rely on static benchmarks that suffer from three severe limitations: First, the compliance rules being used do not comply with the requirements of Artificial Intelligence (AI) laws and regulations; Second, they only handle apparent, explicit compliance risks, leaving implicit and covert compliance risks undetected; Third, they fail to track the systematic propagation of risks along logical dependency chains or evaluate compliance within nuanced, context-based real-world scenarios. To bridge this critical gap, we introduce EADC, a novel advanced evaluation benchmark of LLMs based on an AI compliance knowledge graph and AI compliance legal experts. By mapping abstract legal rules into structured logical multi-relational graphs, our framework enables automated, evolving agents to distill and synthesize highly sophisticated adversarial scenarios. This compliance benchmark is reviewed and corrected by human AI legal experts throughout the whole process. The resulting dataset (4,435+ QA pairs) provides an extensive, multi-dimensional taxonomy covering critical regulatory frontiers, including bias and discrimination, fairness, personal privacy protection, and values. Crucially, our compliance dataset moves beyond shallow string-matching by incorporating contextual long-horizon interactions and logic-driven hazard chains, capturing deeply embedded compliance anomalies that bypass traditional filters. Experiment evaluations demonstrate that our framework exposes critical regulatory blind spots in state-of-the-art LLMs, offering a rigorous, AI laws and regulations-aligned benchmark to safeguard high-level and deep compliance in the application of LLMs.
Chinese Translation
大语言模型(LLM)已被应用于各个行业。然而,确保其符合复杂的法律法规框架仍然是一个巨大的挑战。现有的评估范式主要依赖静态基准,存在三个严重局限:第一,所使用的合规规则不符合人工智能(AI)法律法规的要求;第二,它们仅能处理显而易见的、显性的合规风险,而无法检测隐性和隐蔽的合规风险;第三,它们无法追踪风险沿逻辑依赖链的系统性传播,也无法在基于上下文的复杂真实场景中评估合规性。为弥合这一关键缺口,我们提出了EADC,一种基于AI合规知识图谱和AI合规法律专家的新型大语言模型高级评估基准。通过将抽象的法律规则映射为结构化的逻辑多关系图,我们的框架使自动化、可演化的智能体能够提炼并合成高度复杂的对抗性场景。该合规基准在整个过程中均由人类AI法律专家进行审核与修正。最终生成的数据集(4,435+问答对)提供了覆盖关键监管前沿领域的广泛、多维度分类体系,包括偏见与歧视、公平性、个人隐私保护以及价值观。至关重要的是,我们的合规数据集超越了浅层的字符串匹配方式,引入了上下文相关的长程交互和逻辑驱动的危害链条,从而捕获能够绕过传统过滤器的深层嵌入式合规异常。实验评估表明,我们的框架揭示了当前最先进大语言模型中的关键监管盲区,提供了一个严格、与AI法律法规对齐的基准,以保障大语言模型应用中的高级与深层次合规性。
cs.AI / 79 / 2609.26187

TREND-10K: A Comprehensive Dataset for Next-Generation Video Quality Assessment Based on Preference-Driven Media

Jia, Ziheng, Zhang, Zicheng, Zhang, Junqi, Qian, Jiaying, Wang, Jiarui, Zheng, Yushuo, Min, Xiongkuo
Abstract
The increasing prominence of short-video platforms, coupled with the advanced commercialization of AI-generated content (AIGC) videos, has led to a shift in the types of video media trend consumed by users in their daily lives. Traditional user-generated content (UGC) is gradually being replaced by professional short dramas and AIGC entertainment. Consequently, VQA for contemporary media content has become increasingly important. This requires a unified evaluation framework that can handle diverse video content and evolving media trends. In this context, we introduce TREND-10K, a next-generation comprehensive VQA dataset consisting of the trend-driven part and the static part, containing $10,000$ videos across a wide spectrum of content types. The trend-driven part is based on the TREND-Search framework, which captures user preference profiles from trending lists on online platforms and formulates sampling strategies based on these profiles. The static part, on the other hand, is composed of supplementary samples selected from publicly available datasets. To support unified evaluation for various video types, we incorporate three evaluation dimensions: technical, aesthetic, and AIGC-trace. Experiments show that our dataset ensures high annotation quality and exhibits remarkable generalization across multiple content categories. In conclusion, our work presents a robust framework for advancing VQA, addressing challenges caused by the temporal evolution of user perceptual habits and preferences.
cs.AI / 80 / 2609.26193

Identifying Intelligent Processes via Online Sequential Testing

Das, Aritra, Gupta, Debayan
Abstract
Active sequential hypothesis testing studies how to identify an unknown hypothesis with a given set of sensing actions. We study this in the setting of identifying large language models (LLMs), \textit{i.e.}, if a user is conversing with an LLM drawn from a known set of models, how can they identify which one is in use? Here, the available sensing actions (evaluations) are themselves a design choice: an evaluator must first decide which environments and prompt families to construct, and only then decide how to use them sequentially. We formalize these two levels as an outer probe-design problem and an inner identification problem. Simply put, the outer stage selects a set of probes to be sent to the entire set of models, creating a kind of fingerprint dataset. This is followed by the inner stage, which sequentially sends a budget-minimizing set of those probes to identify the model in use. For the outer problem, we show that selecting which evaluations to construct at minimum cost, so that every pair of candidates is distinguished, is exactly a weighted set cover problem. Since the response distributions of the candidate models are not known exactly but only through calibration samples, we give a one-shot procedure that estimates the cover instance from these samples. For the inner problem, we bound the number of evaluations needed to identify the unknown model in terms of how well the available evaluations distinguish each pair of candidates.
cs.AI / 81 / 2609.26207

RCShift: Certifying When Partial Linkage Suffices for Finite-Sample Decisions

RCShift:证明部分链接何时足以支持有限样本决策
Cao, Shuheng, Chen, Ruiqi, Zhang, Zhenhao, Cao, Renjie, Zhang, Siyu, Dang, Lingwei, Zhang, Jiajun, Dan, Tingting
Abstract
Systems with costly gold outcomes and cheaper auxiliary observations must decide how much record linkage to retain. Complete pairing retains every joint counter, while separate margins retain none. Neither endpoint is calibrated to a declared finite-sample decision. Universal reconstruction can retain cycle directions invisible to the likelihood-ratio family. Family-exact storage can exceed what the decision requires because certified residual loss may fit within finite-sample slack. We introduce RCShift, which certifies two routes to sufficiency under a declared observation contract. Its exact mode characterizes minimum-cost family-exact storage through LR-visible cycle directions. Its approximate mode bounds reverse Le Cam deficiency. Its integer mode certifies whether a chosen set preserves the full experiment's minimum integer record count at specified size and power. In a rank-two witness, one aligned counter preserves a four-record minimum. An equal-cost misaligned counter and the margins require eleven records, while universal reconstruction requires two counters. A local perturbation has positive reverse deficiency yet retains the four-record minimum. Proof-checked scheduling bounds instantiate the contract before gold computation and yield exact reconstruction on the admitted tree support. RCShift turns partial-linkage storage into decision-calibrated measurement design for the declared family, costs, target, and common strictly positive support.
Chinese Translation
对于黄金标准结果(gold outcomes)获取成本高昂而辅助观测成本较低的博弈,系统必须决定保留多少记录链接。完全配对保留所有联合计数,而仅保留分离的边缘分布则不保留任何联合信息。这两个端点都无法校准到声明的有限样本决策。通用重构可以保留似然比族不可见的循环方向。族精确存储可能超出决策所需,因为经认证的残余损失可以容纳在有限样本冗余之内。我们提出RCShift,它在声明的观测契约下认证通向充分性的两条路径。其精确模式通过LR可见的循环方向刻画最小成本的族精确存储。其近似模式给出反向Le Cam亏缺(reverse Le Cam deficiency)的上界。其整数模式认证所选集合能否在指定显著性水平和功效下保持完整实验的最小整数记录数。在一个秩为二的见证示例中,一个对齐的计数器即可保持四条记录的最小值。一个成本相同但未对齐的计数器以及边缘分布则需要十一条记录,而通用重构需要两个计数器。一个局部扰动具有正的反向亏缺,但仍保持四条记录的最小值。经证明检验的调度上界在黄金标准计算之前实例化该契约,并在所允许的树支撑上实现精确重构。RCShift将部分链接存储转变为针对所声明似然比族、成本、目标以及公共严格正支撑的、经决策校准的测量设计。
cs.AI / 82 / 2609.26213

Improved Multiplayer Bandit Algorithm for Bernoulli Rewards

Nguyen, Khang, Parada, Ricardo, Chang, William
Abstract
We study the multiplayer multi-armed bandit problem with information asymmetry under Bernoulli rewards, for three information structures: asymmetry in actions, in rewards, and in both. Replacing the Hoeffding-style confidence intervals of prior work with Kullback--Leibler (KL) divergence-based bounds gives strictly tighter regret guarantees in each case. We propose \texttt{mKL-UCB}, \texttt{mKL-UCB-Intervals} and \texttt{mKL-DSEE}, and show that the improvement factor is at least two by Pinsker's inequality and far larger when reward means are near zero or one. For asymmetry in rewards we prove that two arms' KL intervals separate after a deterministic number of samples, and that $M$ independent players accelerate elimination further.
cs.AI / 83 / 2609.26243

A Hybrid AI Framework for Academic Advising: Integrating Ensemble-Based Grade Prediction and a Rule-Based Expert System

面向学业指导的混合人工智能框架:融合基于集成学习的成绩预测与基于规则的专家系统
Saadatfar, Hamid, Hedayati-Nasab, Rohollah, Eshghi, AmirHossein, Hajihashemi, Arash
Abstract
The rapidly increasing student population has posed serious challenges to the traditional academic advising process. This study designs and implements a multi-purpose intelligent system to support students' academic progress, based on a two-part hybrid framework: (1) an advanced model for grade prediction and (2) a rule-based recommendation engine. Using a dataset containing 416,558 educational records from the University of Birjand, students were first divided into homogeneous clusters using the Gaussian Mixture Model (GMM). Subsequently, a Stacking Ensemble model combining Random Forest, Gradient Boosting, and MLP was trained specifically for each cluster. Evaluation results demonstrated that the Stacking model outperformed base models across all clusters, achieving a final aggregated RMSE of 2.35. The second component is an expert system that provides intelligent recommendations by synergizing educational regulations with the grades predicted by the first component. This system has been implemented as a practical tool on the University of Birjand portal, offering students real-time feedback such as semester GPA prediction, probation risk warnings, and course suggestions for GPA improvement.
Chinese Translation
学生人数的快速增长给传统学业指导流程带来了严峻挑战。本研究设计并实现了一个多功能智能系统,以支持学生的学业进展,该系统基于一个由两部分组成的混合框架:(1) 先进的成绩预测模型;(2) 基于规则的推荐引擎。利用包含比尔詹德大学(University of Birjand)416,558 条教育记录的数据集,首先使用高斯混合模型(Gaussian Mixture Model, GMM)将学生划分为同质集群。随后,针对每个集群专门训练了结合随机森林(Random Forest)、梯度提升(Gradient Boosting)和多层感知机(MLP)的堆叠集成(Stacking Ensemble)模型。评估结果表明,堆叠模型在所有集群中的表现均优于基础模型,最终聚合 RMSE 达到 2.35。第二部分是一个专家系统,通过将教育规章与第一部分预测的成绩相结合,提供智能推荐。该系统已作为实用工具部署在比尔詹德大学的门户平台上,为学生提供实时反馈,例如学期绩点(GPA)预测、试读风险预警以及用于提升 GPA 的课程建议。
cs.AI / 84 / 2609.26244

A Multi-Timestep LSTM Ensemble regressor for Enhanced Short-Term Runoff Prediction

Saadatfar, Hamid, Eshghi, AmirHossein, Behdani, Behnaz
Abstract
Accurately forecasting river runoff is key to managing water resources, controlling floods, and planning agriculture. This study examines the Ajichay River in northwest Iran, a major tributary of Lake Urmia that has experienced increasing water-related stress in recent years. We introduce a daily runoff prediction model based on Long Short-Term Memory (LSTM) networks. The model combines five LSTM units, each trained on different time intervals ranging from 2 to 6 days, to better capture variations in river flow patterns. To improve performance, each model was fine-tuned using Particle Swarm Optimization (PSO), a population-based optimization algorithm. The proposed approach was evaluated on unseen data from 2017-2018 using $R^2$, RMSE, and MSE as performance metrics. The results showed strong predictive accuracy, with $R^2$ values ranging from 74.95% to 91.42%. In addition, multiple feature-importance methods were applied to identify the most influential variables, providing further insight into the factors that drive runoff variations.
cs.AI / 85 / 2609.26261

Coding Agents are Strong Prompt Optimizers

编程智能体是强大的提示词优化器
Singh, Agamdeep, Gautam, Srishti, Gupta, Priyanshu, Mehrotra, Nikita, Bakshi, Tanmay, Gulwani, Sumit
Abstract
Search-based prompt optimizers improve prompts through iterative search: they propose edits, execute fresh rollouts, score the resulting trajectories, and retain only edits that improve a validation metric. We show that this optimization loop is unnecessary. Given only a static corpus of agent trajectories, an off-the-shelf coding agent can directly synthesize an optimized prompt, requiring neither environment access nor validation data. We call this approach \textit{Coding-Agent Skill Distillation} (CASD). The key insight is reflection scope. Rather than reasoning over a small batch of trajectories at each optimization step, the coding agent writes and executes analysis code to compute corpus-wide statistics, identifies systematic failure modes, inspects representative episodes, and distills the resulting insights into behavioral rules. Across four agentic benchmarks (ALFWorld, $\tau^2$-bench retail and telecom, and SpreadsheetBench-Verified), under matched data access, a single CASD pass outperforms GEPA, a state-of-the-art reflective prompt optimizer, on three of four benchmarks and outperforms validation-gated reflective search (SkillOpt) on all four, improving the unoptimized baseline by 16.6 percentage points on average versus 10.9 for GEPA and 5.3 for SkillOpt. Because CASD performs a single offline analysis pass rather than iterative search, producing an optimized prompt costs approximately \$1.60---over $22\times$ cheaper than validation-gated search. Even when competing methods are granted additional validation data and unrestricted environment access, CASD remains ahead on two of four benchmarks. These results suggest that corpus-scale statistical reflection is a viable alternative to iterative search for prompt optimization.
Chinese Translation
基于搜索的提示词优化器通过迭代搜索来改进提示词:它们提出修改、执行新的 rollout、对生成的轨迹进行评分,并仅保留能提升验证指标的修改。我们证明这一优化循环是不必要的。仅给定一个静态的智能体轨迹语料库,一个现成的编程智能体即可直接合成出优化后的提示词,既不需要环境访问权限,也不需要验证数据。我们将这一方法称为“编程智能体技能蒸馏”(Coding-Agent Skill Distillation,CASD)。其关键在于反思的范围(reflection scope)。编程智能体并非在每个优化步骤中对少量轨迹进行推理,而是编写并执行分析代码以计算整个语料库的统计信息,识别系统性的失败模式,检查具有代表性的 episodes,并将所得洞见提炼为行为规则。在四个智能体基准测试(ALFWorld、$\tau^2$-bench 零售与电信、以及 SpreadsheetBench-Verified)上,在相同的数据访问条件下,单次 CASD 处理在四个基准中的三个上超越了当前最先进的反思式提示词优化器 GEPA,并在全部四个基准上超越了带验证门控的反思式搜索(SkillOpt),相比未优化的基线平均提升 16.6 个百分点,而 GEPA 为 10.9,SkillOpt 为 5.3。由于 CASD 执行的是单次离线分析而非迭代搜索,生成一个优化提示词的成本约为 1.60 美元——比验证门控搜索便宜 22 倍以上。即使给竞争方法提供额外的验证数据和无限制的环境访问权限,CASD 仍在四个基准中的两个上保持领先。这些结果表明,语料库规模的统计反思是迭代搜索在提示词优化中的一种可行替代方案。
cs.AI / 86 / 2609.26268

Decoupling Is Not Identification: Supervised Evidential Learning in Next-Token Prediction

Wang, Ge
Abstract
A next-token probability says what a model predicts, not how much training support lies behind it. A Dirichlet head can represent this distinction by separating mean $m$ from concentration $S$, but decoupling does not identify what $S$ means. Here we propose an Evidential Next-Token Prediction (ENTOP) framework to audit this gap on character-level Moby-Dick, using exact 8-gram count as a reproducible lexical-support label and withholding count regression from 20% of context types. Standard implicit evidential training carries essentially no count signal beyond confidence on held-out-label types (partial Spearman $\rho = 0.001 \pm 0.014$), whereas explicit supervision generalizes ($\rho = 0.201 \pm 0.010$; matched-pair win $= 0.822 \pm 0.021$). CE predictive entropy is at chance for unseen 8-grams (AUROC $= 0.490 \pm 0.004$), while supervised vacuity reaches $0.772 \pm 0.003$, comparable with an indexed CE-representation baseline ($0.769$) but below the tautological corpus oracle ($1.000$). Neither longest-suffix nor representation-distance strata explain where amortization succeeds. Increasing count weight under the digamma objective improves support fit only by sacrificing prediction. A constant predictor wins natural log-RMSE, and vacuity does not improve error deferral. These results motivate a minimum evidence protocol---confidence control, matched pairs, held-out labels, a constant baseline, and a decision test---and show that concentration can pass identification while failing calibration and utility.
cs.AI / 87 / 2609.26279

FISSION: Label Augmentation for Bot Detection

Yang, Sen, Nieweglowski, Ignacy, Yaish, Aviv
Abstract
Bot accounts and coordinated influence operations are often discovered via heuristic methods, leaving a dearth of reliable ground-truth labels for training detection systems. To address this challenge, we study a natural question: can we generate labels to assist in learning embeddings in which bots and accounts from the same coordinated operation are close? We present FISSION, a method to generate labels by splitting each account's activity into positively labeled sub-accounts. Given this label source, we train detection models which preserve behavioral regularities recurring across positive sub-accounts. We evaluate FISSION and show it outperforms prior methods in detecting Wikipedia sockpuppets and Twitter/X bots.
cs.AI / 88 / 2609.26293

Dual-Frontier: When Can an Agent Trust Its World Model?

Zhu, Huatai, Chen, Qiang, Kou, Ziqian, Li, Wenhao, Wang, Fei, Cao, Yichao, Su, Xiu, Chen, Yi
Abstract
Learned world models are becoming essential to general-purpose agents: by predicting action consequences, they support planning and decision-making while reducing reliance on costly trial and error. This reliance creates a fundamental ambiguity: when a world-model-guided decision fails, the trajectory alone may not reveal whether the agent's decision rule or the world model caused the loss. We formalize this failure-attribution problem as a counterfactual decomposition of return loss and prove that its components are not identifiable from passive interaction, even for finite-horizon planners. This obstruction motivates Dual-Frontier, a learning principle that admits a world-model-guided decision only when its predicted advantage exceeds a certified bound on decision-relevant world-model error; otherwise, evidence is allocated to world-model verification. Action-conditioned value bounds and a closed-loop extension guarantee non-decreasing return for admitted decisions. Calibrated gates and simultaneous confidence sequences support adaptive evidence reuse, with sufficient and necessary verification bounds. Controlled learned-model experiments validate the predicted failure modes and certification behavior, while cross-backbone tool-use benchmarks instantiate the same verify-then-promote rule in realistic agent world-model pipelines, consistently improving decision quality and reliability.
cs.AI / 89 / 2609.26419

Reliability Theory for AI Control

AI控制中的可靠性理论
Molnar, Grant
Abstract
Reliability theory gives a mature language for layered systems, but its formal tools are not yet standard in frontier AI control. We apply them to Google DeepMind's defenses against rogue deployment. The same control stack can have cubic, quadratic, or linear rare-failure suppression depending on its failure domains. Birnbaum importance identifies which component improvements buy the most nominal reliability, while prevention changes the population on which recovery is demanded. These results give concrete guidance about what to separate, improve, measure, and test.
Chinese Translation
可靠性理论为分层系统提供了一套成熟的语言体系,但其形式化工具尚未成为前沿AI控制领域的标准方法。我们将这些工具应用于Google DeepMind针对恶意部署的防御体系。研究表明,同一个控制技术栈因其故障域的不同,其罕见故障抑制能力可呈立方级、平方级或线性变化。Birnbaum重要性度量能够识别哪些组件的改进能带来最大的名义可靠性提升,而预防措施则会改变需要实施恢复的对象范围。这些结果为应当分离、改进、度量和测试哪些内容提供了具体指导。
cs.AI / 90 / 2609.26428

The Source of Disturbance Matters: External, Internal, and Control-Generated Noise in Adaptive Regulation

扰动来源的重要性:适应性调节中的外部、内部与控制产生的噪声
Ziegler, Veronique
Abstract
Adaptive regulation can itself perturb the state it is intended to stabilize. In replicated simulations of an adaptive agent, we compare external disturbance, persistent internally generated disturbance, and control-generated disturbance under regulation-first and disturbance-first ordering. Persistent internal disturbance produces the largest exposure and regulatory burden within the tested parameter grid. When positive controller updates generate an immediate disturbance cost, increasing that cost produces a nonmonotonic response: effective disturbance initially rises, variability across stochastic runs increases over an intermediate range, and corrective activity becomes strongly suppressed at higher costs. The results show how disturbance source and timing shape exposure and controller burden in this model. They motivate testing adaptive agents with distinct disturbance sources and assessing regulatory activity alongside exposure.
Chinese Translation
适应性调节本身可能会扰动其原本要稳定的状态。在一个适应性智能体的重复仿真中,我们在“调节优先”和“扰动优先”两种时序下比较了外部扰动、持续内部产生的扰动以及控制产生的扰动。在所测试的参数网格内,持续的内部扰动产生了最大的暴露度和调节负担。当控制器的正向更新会产生即时扰动代价时,增大该代价会引发非单调响应:有效扰动最初上升,随机运行之间的变异性在中间区间内增大,而纠偏活动在高代价时受到强烈抑制。结果表明,扰动来源与时序如何塑造该模型中的暴露度和控制器负担。这些结果启发我们在测试适应性智能体时区分不同扰动来源,并将调节活动与暴露度一并评估。
cs.AI / 91 / 2609.26457

Recursive self-improvement of AI research agents

AI研究智能体的递归自我改进
Srikanth, Dhruv, Zhao, Bingchen, Xu, Dixing, Wu, Yuxiang, Jiang, Zhengyao
Abstract
AI agents are beginning to automate research and development across the AI stack, from improving training efficiency to optimizing inference. A natural next step is to improve the research efficiency of the agents themselves. When an AI research agent's own code is the object of optimization, each accepted rewrite becomes the agent that the next round edits. We refer to this loop as recursive self-improvement. Its significance lies in a long-standing trend, in which increased cumulative spending on R&D yields diminishing returns. Sustained self-improvement offers a way to counter this trend. We present AIDE^2, a system that implements this loop for a frontier AI research agent. It proposes changes to its own code, benchmarks modified versions of itself on a suite of AI R&D tasks, and keeps the changes that perform best on hidden evaluations. In an autonomous 8-day run, AIDE^2 discovered seven successive improvements, ranging from a new search policy to memory mechanisms that compress and manage the agent's growing context. These gains generalize to four held-out benchmarks spanning machine learning engineering, heuristic algorithm engineering, and physics-based weather forecasting, the last of which is out of distribution from the selection tasks. On all four, the strongest discovered agent matches or exceeds a human-engineered production research agent that ranks among the strongest on FML-Bench. On a separate held-out task family, the discovered agents also exhibit reduced reward hacking, a property the loop never explicitly optimized for: the rate falls from 55% to 32% during the run, 7 percentage points below the human-engineered agent. Together, these results show that an AI research agent can improve its own research efficiency through recursive self-improvement, and that these gains transfer to tasks and domains the loop never encountered.
Chinese Translation
AI智能体开始在整个人工智能技术栈中自动化研发工作,从提升训练效率到优化推理。一个自然的下一步是提升智能体自身的研究效率。当AI研究智能体自身的代码成为优化对象时,每一次被接受的代码改写都会成为下一轮编辑所作用的智能体。我们将这一循环称为递归自我改进。其重要性在于一个长期存在的趋势:研发累计投入的增加带来边际收益递减,而持续的自我改进为对抗这一趋势提供了一条途径。我们提出了AIDE^2,一个为前沿AI研究智能体实现该循环的系统。它对自己的代码提出修改建议,在一组AI研发任务上对修改后的自身版本进行基准测试,并保留在隐藏评估中表现最好的修改。在一次为期8天的自主运行中,AIDE^2发现了七项连续的改进,涵盖一种新的搜索策略,以及用于压缩和管理智能体不断增长的上下文的记忆机制。这些改进可泛化到四个保留基准测试上,涉及机器学习工程、启发式算法工程以及基于物理的天气预报——最后一项相对于选择任务而言是分布外的。在全部四个基准上,所发现的最强智能体均达到或超越了一个在FML-Bench上排名前列的人工设计的生产级研究智能体。在一个单独保留的任务族上,所发现的智能体还表现出更低的奖励作弊倾向——这是该循环从未显式优化过的属性:其发生率在运行期间从55%降至32%,比人工设计的智能体低7个百分点。这些结果共同表明,AI研究智能体可以通过递归自我改进提升自身的研究效率,并且这些改进能够迁移到该循环从未接触过的任务和领域。
cs.AI / 92 / 2609.26461

Reproducible AI Requires Reproducible Randomness

可复现的人工智能需要可复现的随机性
Bertrand, Anthony, Schmitt, Tom, Nguifo, Engelbert Mephu, Hill, David
Abstract
Pseudorandom number generators (PRNGs) constitute indispensable computational tools across multiple scientific domains, including Monte Carlo simulations, stochastic computing, and artificial intelligence (AI). The reproducibility of such applications critically depends on the ability of PRNG implementations to generate identical sequences across software environments when initialized from the same internal state. These algorithms enable the simulation of stochastic processes while providing deterministic and repeatable behaviour, thereby facilitating reproducible experiments. Modern PRNG implementations may be initialized through either a seed or, more accurately, an initial state that exceeds the capacity of a conventional integer seed. However, reliance on a simple seed alone frequently proves insufficient to ensure consistent program execution traces across different implementations. A natural assumption is that transferring the complete internal state of a generator should guarantee identical outputs regardless of the software library used. This study examines the validity of this assumption by investigating whether complete initial states can ensure cross-library fidelity and portability of PRNG streams. We focus on two widely deployed generators, Mersenne Twister and Philox, and evaluate their implementations across four major Python ecosystems-Random, NumPy, PyTorch, and TensorFlow. We compare the sequences produced by these implementations against those generated by the original reference algorithms under identical initialization conditions. Our results demonstrate that reproducibility cannot be assumed from PRNG state transfer alone, even when implementations claim to follow the same underlying algorithm. While fidelity was successfully achieved for several implementations, significant discrepancies were observed in others. Most notably, the Philox implementation in PyTorch exhibits fundamental incompatibilities with the reference algorithm, preventing exact reproduction of generator outputs across environments. These findings challenge the common expectation that access to a full internal state of a PRNG is sufficient to ensure reproducibility across software stacks. They further highlight that implementation-specific design choices can introduce hidden barriers to experimental replication, particularly in AI workflows that rely on multiple frameworks. This work shows that implementation fidelity of a PRNG is a necessary condition for scientific reproducibility and makes two primary contributions. First, it identifies practical guidelines for achieving reliable PRNG usage and reproducibility within the Python scientific and AI ecosystem. Second, it evaluates the extent to which cross-library portability and fidelity can be recovered through user-level techniques, without requiring modifications to library source code.
Chinese Translation
伪随机数生成器(PRNG)是包括蒙特卡洛模拟、随机计算和人工智能(AI)在内的多个科学领域中不可或缺的计算工具。此类应用的可复现性关键取决于PRNG实现在相同内部状态下,能否在不同软件环境中生成相同的序列。这些算法能够模拟随机过程,同时提供确定性和可重复的行为,从而促进可复现的实验。现代PRNG实现既可以通过种子(seed)初始化,也可以——更准确地说是——通过超出传统整数种子容量的初始状态来初始化。然而,仅依赖一个简单的种子往往不足以确保程序在不同实现下产生一致的执行轨迹。一个自然的假设是,传递生成器的完整内部状态应当能够保证无论使用何种软件库都能获得相同的输出。本研究通过考察完整初始状态能否确保PRNG序列的跨库一致性和可移植性,来检验这一假设的有效性。我们聚焦于两个被广泛部署的生成器——梅森旋转算法(Mersenne Twister)和Philox——并评估它们在四个主要Python生态系统(Random、NumPy、PyTorch和TensorFlow)中的实现。我们在相同的初始化条件下,将这些实现生成的序列与原始参考算法生成的序列进行比较。结果表明,即使各实现声称遵循相同的底层算法,也不能仅凭PRNG状态传递就假定可复现性。虽然部分实现成功实现了一致性,但在其他实现中观察到了显著差异。最值得注意的是,PyTorch中的Philox实现与参考算法存在根本性的不兼容,导致无法在不同环境中精确复现生成器的输出。这些发现挑战了一个普遍预期,即获取PRNG的完整内部状态就足以确保跨软件栈的可复现性。它们还进一步凸显了实现层面的特定设计选择可能给实验复现带来隐性障碍,尤其是在依赖多个框架的AI工作流中。这项工作表明,PRNG的实现一致性是科学可复现性的必要条件,并做出了两项主要贡献:第一,它为在Python科学计算与AI生态系统中实现可靠的PRNG使用和可复现性提供了实用指南;第二,它评估了在不修改库源代码的前提下,通过用户级技术恢复跨库可移植性和一致性的程度。
cs.AI / 93 / 2609.26532

REFLEX with Jev for Efficient Selective Control in LLM Agents

基于Jev的REFLEX架构:面向LLM智能体的高效选择性控制
Wu, Tiantong, Lim, Wei Yang Bryan
Abstract
LLM agents often use generative models for bounded decisions, raising the question of when these decisions can be handled more efficiently without reducing task success. We study REFLEX, an agent architecture that uses Jev as a fast, typed decision layer and calls a strong LLM when confidence is low, or generation is required. On a frozen 100-task benchmark, REFLEX achieves 95% success with 72.7% fewer strong-model calls than a strong-only agent, with reductions persisting across three fallback families. Controlled interventions show that reliability depends on action-set size and near-valid alternatives near authorization boundaries. External BFCL and $\tau$-style evaluations reveal limited advantages over a cheap generative cascade when ordinary routing is already highly accurate. These findings identify when selective control with Jev can reduce computation and where its benefits are limited.
Chinese Translation
LLM智能体常使用生成式模型进行受限决策,这引发了一个问题:在不降低任务成功率的前提下,这些决策能否被更高效地处理。我们研究了REFLEX——一种使用Jev作为快速、类型化决策层的智能体架构,当置信度较低或需要生成内容时,才会调用强大的LLM。在一个固定的100任务基准上,REFLEX达到了95%的成功率,且比仅使用强模型的智能体减少了72.7%的强模型调用次数,且这一降幅在三种回退模型家族中均持续存在。受控干预实验表明,可靠性取决于动作集合的大小以及授权边界附近近似有效的备选动作。外部的BFCL及$\tau$风格评测显示,当普通路由已经具有很高准确率时,相比廉价的生成式级联方案,其优势有限。这些发现明确了基于Jev的选择性控制何时能够减少计算开销,以及其收益在何处受到限制。
cs.AI / 94 / 2609.26550

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Li, Yubo, Miao, Yidi, Krishnan, Ramayya, Padman, Rema
Abstract
LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking a derivation or resisting an elaborately written wrong answer. On several benchmarks, JEV's gap to this comparator is concentrated in low-confidence decisions. A frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% of the comparator's accuracy at lower cost.
cs.AI / 95 / 2609.26556

Neutral-Atom-based Quantum Optimization for Resource Allocation in NOMA Networks

Keyela, Patatchona, Polus, Remon, Cherkaoui, Soumaya, Ahmad, Ola
Abstract
In wireless communication networks, many resource optimization problems are nondeterministic polynomial-time hard (NP-hard) due to their combinatorial nature and high computational complexity. Recently, neutral-atom-based quantum computing has emerged as a promising platform for efficiently solving such problems by leveraging quantum superposition and entanglement. However, its application to wireless communication optimization problems remains largely unexplored. In this paper, we investigate the use of neutral-atom quantum platforms to solve the maximum access problem (MAP), formulated as a mixed-integer programming task that jointly considers admission control, user clustering, channel assignment, and power allocation in a non-orthogonal multiple access (NOMA)-enabled uplink network. To reduce the computational burden, the MAP is equivalently reformulated as a maximum independent set (MIS) problem in graph theory. This reformulation enables the use of the neutral atom platform based on Rydberg atom arrays, where the MIS problem is naturally encoded into the physical geometry and blockade constraints of the quantum system. Numerical results demonstrate the feasibility and potential of this approach for addressing large-scale wireless resource optimization problems.
cs.AI / 96 / 2609.26565

Quantum-Aided Active Device Detection in Energy-Harvesting Symbiotic Radio Networks

Polus, Remon, Tashman, Deemah, Cherkaoui, Soumaya
Abstract
Massive connectivity in next-generation networks demands energy- and spectrum-efficient solutions for large-scale Internet of Things (IoT) deployments. Symbiotic radio (SR) enables passive IoT devices to communicate by backscattering existing cellular transmissions. A key challenge in uplink SR is active device detection (ADD), which directly affects decoding reliability, interference management, and system throughput. We propose an energy-harvesting code-domain non-orthogonal multiple access (NOMA)-SR system in which IoT devices harvest energy from ambient uplink signals and backscatter information using low-density spreading (LDS) codes. To reduce the complexity of ADD, Grover's quantum search algorithm is employed, providing a quadratic reduction in oracle-query complexity over exhaustive maximum-likelihood (ML) search. Numerical results show that the proposed approach closely approaches ML performance while substantially reducing the number of search iterations, demonstrating its potential for scalable ambient IoT systems.
cs.AI / 97 / 2609.26642

The Delegation Blind Spot: Auditing Product Decisions from Agent Choices

Gupta, Shivam
Abstract
Successful agent execution need not identify which future product improvement its user would value. We present a decision-specific audit that maps a declared observation channel and product-value contrast to compatible intervals and witness populations. Its foundations are established identification and decision theory; the contribution is an executable measurement workflow and a controlled study of its limits. A frozen experiment makes 4,800 requests to two pinned model snapshots on shared synthetic tasks. All 36 conservative primary intervals remain unresolved despite different execution accuracy. An exploratory 2,400-call follow-up records supplied preferences and resolves three of nine comparisons per model. A deterministic extractor resolves seven of nine without model calls or calibration observations, exposing unnecessary uncertainty introduced by model-generated reports. A further 14,400 controlled multinomial simulations distinguish structural ambiguity from weak identification and finite calibration precision. We propose a source-labeled decision receipt and provide an offline viewer for inspecting the audit. These results motivate preserving decision-relevant structured input and diagnosing why a decision is unresolved before collecting more telemetry. The study contains no human participants or real customer outcomes. Full proofs, raw model provenance, controlled experiments, and reproducible analyses accompany the report.
cs.AI / 98 / 2609.26758

Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It

类型安全不等于无误判:受约束的决策头跟随选项名称,而非绑定于其上的评分标准
Sun, Yu, Xu, Junhao
Abstract
Typed decision models are built for settings where model outputs are consumed directly by software. Instead of generating free-form text, they return a decision over a predefined set of options. By construction, every output conforms to the required schema. Yet this guarantee does not tell us whether the model interprets the options as intended. We study Jev and two Jev-like models with open weights by changing how option names are assigned to rubrics. Each option consists of an option name and a textual rubric that defines what the option means. We change only which option name is assigned to each rubric; the question, state, rubric wording, and set of option names remain exactly the same. On 1200 workflow decisions with task-specific rubrics, renaming the two options from 0/1 to no/yes changes 70.4 more answers per hundred (95% CI: [67.6, 73.1]) and shifts AUC from .94 to .23, revealing a systematic reversal in the decision ranking rather than simple uncertainty. The same operation has little effect with neutral option names. This pattern holds across all 4 predicates, where the effect is at least 7.4x larger than under the neutral control, and becomes stronger as the number of options increases. The effect also depends on the read-out geometry: a second model family that mean-pools over the full option span flips 4.1x less often. The hosted model exhibits the same behavior: the swap changes AUC from .8146 to .5806 and produces 24x as many answer flips as its test-retest floor. In contrast, replacing the option names with random character strings returns all model families to the neutral-control regime without reducing accuracy. The failure therefore depends on the semantic polarity of the option names rather than on the renaming operation itself. Across all conditions, the type-error rate remains 0%, even when decision accuracy degrades substantially.
Chinese Translation
类型化决策模型专为模型输出被软件直接消费的场景而构建。它们不生成自由文本,而是在一组预定义选项上返回一个决策。从构造上看,每个输出都符合所需的模式(schema)。然而这一保证并不能告诉我们模型是否按预期的方式解释这些选项。我们通过改变选项名称与评分标准(rubric)的分配方式,研究了 Jev 以及两个具有开放权重的类 Jev 模型。每个选项由一个选项名称和一段定义该选项含义的文本评分标准组成。我们只改变每个评分标准所分配到的选项名称,而问题、状态、评分标准的措辞以及选项名称集合保持完全不变。在 1200 个带有任务特定评分标准的工作流决策上,将两个选项从 0/1 重命名为 no/yes,会使每百个答案中多改变 70.4 个(95% 置信区间:[67.6, 73.1]),并使 AUC 从 .94 降至 .23,这揭示的是决策排序中的系统性反转,而非简单的不确定性。同样的操作在选项名称为中性的情况下几乎没有影响。该模式在所有 4 个谓词上都成立,其效应至少是中性对照条件下的 7.4 倍,且随着选项数量的增加而变得更强。该效应还依赖于读取几何结构(read-out geometry):另一个对完整选项跨度进行均值池化的模型系列,其翻转频率降低了 4.1 倍。托管模型也表现出同样的行为:这一交换使其 AUC 从 .8146 变为 .5806,产生的答案翻转数量是其重测下限(test-retest floor)的 24 倍。相比之下,将选项名称替换为随机字符串会使所有模型系列回到中性对照的状态,且不降低准确率。因此,这一失效取决于选项名称的语义极性,而非重命名操作本身。在所有条件下,类型错误率始终保持为 0%,即使决策准确率大幅下降。
cs.AI / 99 / 2609.26760

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Li, Laizhen, Li, Jiarui, Zhao, Juanjuan, Ye, Kejiang, Li, Ye, Xu, Cheng-zhong, Gao, Xitong
Abstract
Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for task-specific semantic reasoning. We introduce Growing Harness, a failure-guided training paradigm that learns the agent harness itself from a strategy-free scaffold that exposes fixed model and tool interfaces but encodes no task-solving controller. Function-level execution traces localize each failure to a bounded code surface, an optimizer repairs a window of failures jointly, and a success-first held-out gate rolls back repair sequences that harm prior capability. Accepted edits accumulate in one shared harness, allowing its control structure to emerge from task feedback. Across BrowseComp-Plus and WebArena-Verified with three deployment models from 4B to 120B parameters, Growing Harness achieves the highest mean success in five of six benchmark-model settings and trails the best mean by 0.7 pp. in the sixth. Relative to a Tool-Calling agent, it reduces LLM calls by 76.0-91.8% and deployed-agent inference cost by 74.4-98.6%. On WebArena-Verified, its success remains 44.7-45.3% across model scales, whereas Tool-Calling falls to 6.7% with the 4B model. Ablations show that trace-local edits, joint repair, and gate-based rollback each improve final success. These results show that persistent program growth can move recurring control out of model context and into low-cost code, yielding reusable specialist agents that remain effective with smaller deployment models.
cs.AI / 100 / 2609.26777

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

Williams, Jennifer, Farris, Dave, Farris, Jeff, Jiao, Jiantao
Abstract
We introduce SWE-Serve, a benchmark for evaluating agents on production inference engineering tasks. Implementing an inference feature can require coordinating multiple changes across the serving stack, including model support, runtime execution, and public APIs. Existing benchmarks provide limited coverage of production inference engineering: repository-level software engineering benchmarks do not target inference, while general terminal-agent benchmarks include only a few inference tasks. Dedicated inference benchmarks, meanwhile, focus primarily on isolated kernel generation or performance optimization rather than repository-scale production feature implementation. SWE-Serve provides 53 repository-grounded tasks derived from recent production changes to SGLang, spanning six inference engineering families. Each task executes on either CPU or a single GPU (H100) and is evaluated with hidden functional and regression tests, including, where applicable, end-to-end (E2E) serving tests and calibrated performance gates. Executable no-op and oracle controls, adversarial verifier review, and closed-book execution support task validity and evaluation integrity. Across 11 models and 31 model-effort configurations, the best-performing configuration achieves 75% mean pass@1. SWE-Serve exposes a substantial gap between completing tasks locally and achieving production correctness. On 19 tasks with end-to-end coverage, model-serving E2E tests reject roughly one-third of patches that pass every other test (45.9% under the verifier versus 69.4% with E2E tests excluded from scoring), with pass rate increasing for each model's best-performing configuration. By making the production correctness gap directly measurable, SWE-Serve enables the field to track whether future agents move beyond completing tasks locally to achieving production correctness.
cs.AI / 101 / 2609.26779

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

CliffCompaction:面向长程编码智能体的低成本高效上下文压缩方法
Nguyen, Trang, Cho, Eulrang, Chen, Bingqing, Dettmers, Tim
Abstract
Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on Terminal-Bench and achieving new levels of efficiency for test-time scaling and state-of-the-art results on KernelBench. The per-rollout savings of CliffCompaction make the performance--cost trade-off of test-time scaling more efficient, adding over 10 percentage points on Terminal-Bench for less than the cost of two full-context runs. Under parallel test-time scaling, CliffCompaction lets Kimi K2.6 match Opus 4.7, and exceed Opus 4.6 and GPT-5.3 Codex at lower cost. The key to CliffCompaction's effectiveness is that it keeps compacted information faithful by only truncating or dropping content, never rephrasing or rewriting it. We never compact a compaction---each pass operates only on original content, and prior compacted output is discarded, preventing context drift from accumulating. These properties sustain continual learning over sessions exceeding a million tokens: on KernelBench, CliffCompaction reaches CUDA kernel speedups of $2.23\times$ after 200 steps and $3.58\times$ after 400 steps, surpassing specialized search algorithms and trained agents despite being a general-purpose compaction technique. We open-source a scaffold-agnostic API-proxy implementation of CliffCompaction usable with Claude Code, Codex and other harnesses.
Chinese Translation
智能体在处理复杂问题时往往需要数百万词元的上下文,由于上下文窗口有限,这就需要在会话之间进行压缩。我们提出了 CliffCompaction,一种自动压缩技术,可在有界上下文条件下将成本降低多达 50%,同时在 Terminal-Bench 上保持甚至提升性能,并在测试时扩展(test-time scaling)方面实现了新的效率水平,在 KernelBench 上取得了最先进的结果。CliffCompaction 在每次 rollout 中的成本节省,使测试时扩展的性能—成本权衡更加高效:以低于两次全上下文运行的成本,在 Terminal-Bench 上提升超过 10 个百分点。在并行测试时扩展下,CliffCompaction 使 Kimi K2.6 能够以更低的成本匹敌 Opus 4.7,并超越 Opus 4.6 和 GPT-5.3 Codex。CliffCompaction 高效的关键在于,它通过仅截断或删除内容来保持压缩信息的忠实性,从不改写或重写内容。我们从不压缩“压缩结果”——每次压缩只针对原始内容进行,先前的压缩输出会被丢弃,从而防止上下文漂移不断累积。这些特性使得跨越超过一百万词元的持续学习得以维持:在 KernelBench 上,CliffCompaction 在 200 步后达到 2.23 倍的 CUDA 内核加速,400 步后达到 3.58 倍,尽管它是一种通用压缩技术,却超越了专门的搜索算法和经过训练的智能体。我们开源了一个与脚手架(scaffold)无关的 API 代理实现,可配合 Claude Code、Codex 及其他运行框架使用。
计算机视觉 (Computer Vision)
108
cs.CV / 1 / 2609.25017

Deepfakes and Synthetic Media: Generation, Detection, and Governance

深度伪造与合成媒体:生成、检测与治理
Gazis, Alexandros, Karypidis, Efstathios, Santamouri, Kleanthi, Vavouras, Theodoros, Mastorakis, Nikos E., Pappas, Stylianos
Abstract
Deepfakes, synthetic audiovisual content produced by deep generative models, have escalated into a critical threat across civilian and military domains, enabling identity fraud, disinformation campaigns, and evidence fabrication. In high-stakes environments, ranging from journalism and finance to healthcare and legal contexts, the consequences extend to severe misinformation, market manipulation, identity fraud, and the erosion of institutional trust. This entry explores how modern visual intelligence and computer vision techniques are used to detect deepfakes. It outlines key deepfake generation models, such as GANs, autoencoders, neural rendering, and diffusion systems, while also explaining how adversarial methods enhance realism and challenge existing detectors. The overview highlights visual artifacts, digital patterns, and physiological cues commonly leveraged in detection and reviews major CNN, transformer, and frequency-based approaches. It also summarizes evaluation practices and the difficulty of achieving strong generalization. Finally, it identifies emerging directions, including modern intelligence techniques for civilian and military content verification. This survey covers generation architectures (GANs, latent diffusion, neural rendering, video synthesis), the spatial, temporal, frequency-domain, and physiological artifacts they produce, and the detector families that exploit them. We examine evaluation benchmarks and protocols, highlighting cross-generator generalization as the field's central open challenge. Beyond detection, we discuss cryptographic provenance standards, watermarking, and regulatory frameworks (EU AI Act, DSA, GDPR). We conclude that effective deepfake governance requires defense-in-depth integrating forensic detection, verifiable provenance, and institutional accountability.
Chinese Translation
深度伪造(Deepfakes)是由深度生成模型合成的视听内容,已升级为横跨民用与军事领域的关键威胁,助长身份欺诈、虚假信息传播和证据伪造。在高风险环境中,从新闻、金融到医疗和法律领域,其后果包括严重的错误信息、市场操纵、身份欺诈以及机构信任的侵蚀。本条目探讨如何利用现代视觉智能与计算机视觉技术来检测深度伪造内容。文章概述了主要的深度伪造生成模型,如生成对抗网络(GAN)、自编码器、神经渲染和扩散模型,同时解释了对抗性方法如何提升逼真度并挑战现有检测器。本综述重点介绍了检测中常用的视觉伪影、数字模式和生理线索,并回顾了主要的基于CNN、Transformer和频域的方法。文章还总结了评估实践以及实现强泛化能力的困难。最后,本文指出了新兴研究方向,包括用于民用和军事内容验证的现代智能技术。本综述涵盖生成架构(GAN、潜在扩散模型、神经渲染、视频合成)、其产生的空间、时序、频域和生理伪影,以及利用这些伪影的各类检测器。我们考察了评估基准与协议,强调跨生成器泛化是该领域最核心的开放性挑战。在检测之外,我们还讨论了密码学溯源标准、数字水印以及监管框架(如欧盟《人工智能法案》、DSA、GDPR)。我们的结论是,有效的深度伪造治理需要纵深防御体系,将法证检测、可验证溯源和机构问责有机结合。
cs.CV / 2 / 2609.25067

SPARC: SuperPixel-Aware Region Contrastive Learning for Self-Supervised Dense Prediction

Szczecina, David, Xiang, Yuanpei, Hu, Jitao, Clausi, David, Chen, Yuhao, Deglint, Jason, Fieguth, Paul
Abstract
Self-supervised learning (SSL) has become an effective approach for learning visual representations without manual annotations. Among SSL approaches, contrastive learning has been widely used for visual representation learning. However, existing contrastive SSL methods have focused primarily on image-level or pixel-level representation learning, while region-level representation learning remains less explored. We propose SPARC, a region-level contrastive learning framework that leverages superpixels to establish explicit correspondence between augmented image views. SPARC introduces a region contrastive branch that performs superpixel-based feature pooling and optimizes a region-level contrastive objective jointly with a global image-level objective. Under identical settings, SPARC consistently outperforms previous methods such as MoCo-v2 and DenseCL, achieving improvements of up to +9.79 mIoU for semantic segmentation and +4.88 AP for object detection. Ablation studies further demonstrate that region-level objectives produce the strongest performance. Thus, region-level contrastive learning is an effective approach for improving self-supervised visual pretraining for dense prediction tasks. Code repository can be accessed at https://github.com/xRIPEIx/SPARC.
cs.CV / 3 / 2609.25108

You've Seen Enough: Quality-Constrained Image Coding for Machines

所见已足:面向机器的质量约束图像编码
Pham-Dinh, Khoa, Nami, Sanaz, Tavakoli, Hamed Rezazadegan, Gabbouj, Moncef, Pakdaman, Farhad
Abstract
Visual data is increasingly consumed by machine-vision systems rather than by human observers. Image Coding for Machines (ICM) compresses images assuming the main observer is a computer vision application and that the human observer needs to inspect or validate the decisions. Inspired by just-noticeable distortion, which sets the quality to the just-acceptable level for human observers, we aim to cap the human-observed quality at a desired level, with the goal of using the remaining coding capacity to improve the machine performance. We recast joint compression-segmentation training as a constrained optimization problem in which the codec must meet a predefined acceptable target visual quality while a task term consumes the remaining coding capacity. We solve this by designing a penalty function to guide the quality to the desired target. We propose two penalty functions, an absolute function and a bilinear function, the latter applying a steeper slope once the target visual quality is exceeded. Experimental results show that, under the quality constraint, the proposed method achieves a BD-rate of $-22.82\%$ over an unconstrained joint rate--distortion--task optimization and $-29.81\%$ over a simple rate--distortion baseline, showcasing bitrate reduction with the same task performance. This is achieved while the codec also meets the target visual quality with a reasonable error and without adding any complexity overhead.
Chinese Translation
视觉数据越来越多地被机器视觉系统而非人类观察者所消费。面向机器的图像编码(Image Coding for Machines, ICM)假设主要观察者是计算机视觉应用,而人类观察者仅需要检查或验证机器的决策。受恰可察觉失真(just-noticeable distortion)的启发——该概念将质量设定为人类观察者恰可接受的水平——我们旨在将人类观察到的质量限制在期望的水平上,从而利用剩余的编码能力来提升机器性能。我们将压缩与分割的联合训练重新表述为一个约束优化问题:编解码器必须满足预定义的可接受目标视觉质量,同时由任务项消耗剩余的编码能力。为此,我们设计了一个惩罚函数来引导质量达到期望的目标,并提出了两种惩罚函数:绝对值函数和双线性函数,其中后者在视觉质量超过目标后施加更陡的斜率。实验结果表明,在质量约束下,所提方法相对于无约束的率-失真-任务联合优化取得了 -22.82% 的 BD-rate 增益,相对于简单的率-失真基线取得了 -29.81% 的 BD-rate 增益,在相同任务性能下实现了码率的降低。同时,编解码器还能以合理的误差满足目标视觉质量,且不增加任何复杂度开销。
cs.CV / 4 / 2609.25247

Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

Kong, Hanyang, Yang, Xingyi
Abstract
Interaction understanding in 3D scenes requires a joint description of movable parts, their motion, and the regions through which they can be operated. We present Segment-Snap, which connects these outputs through the physical relationship between parts and handles. Learned predictors identify broad part surfaces and small handles. A geometric decoder uses planar and upright priors to constrain motion, then selects hinge lines using predicted handle locations, without training a motion regressor. Conversely, a joint part-and-handle predictor supplies additional handle candidates, whose motion classes are refined using containing parts. Each information transfer is applied once, without iterative feedback. On Articulate3D validation, handle guidance raises motion-gated AP from 13.74% to 40.98% at fixed masks and axes. Additional handle candidates raise handle AP from 24.63% to 29.65%; part-based class correction adds 0.98 points, and full context reaches 30.99%. Repeated training, learned-decoder controls and paired visualizations establish the benefits and limitations of combining geometric and semantic evidence for interaction understanding.
cs.CV / 5 / 2609.25267

ImIR: Image-Instruction Tuning for All-in-One Image Restoration

ImIR:面向一体化图像修复的图像-指令微调
Aslan, Süleyman, Aydemir, Görkay, Yavuz, Mısra, Kurt, Yunus Bilge, Rahimi, Nasrin, Emirdağı, Ahmet Rasim, Biner, Burak Can, Yılmaz, M. Akın
Abstract
Degradations vary widely across images, so a practical restoration system has to handle many degradation types with one model. A recent and effective recipe adapts a large pretrained image-editing model to restoration using a small low-rank adapter with a text prompt. We replace that prompt with an instruction derived from the degraded image itself. The image reaches the editor through two paths: its structure comes from the model's VAE, and its semantic instruction comes from a lightweight token mapper that shifts the degraded image's vision-language embedding toward the embedding a clean image would produce. Because the instruction is a continuous vector, scaling it yields a family of valid restorations for tasks whose target is not unique, such as low-light enhancement. We adapt one Qwen-Image-Edit model to six tasks with a single adapter trained in about three hours on one GPU. The image instruction outperforms text conditioning under a matched comparison, and it supports task agnostic restoration without a degradation label, which the text variant does not.
Chinese Translation
图像中的退化类型千差万别,因此一个实用的修复系统需要用单一模型处理多种退化类型。近期一种有效的方法是利用小型低秩适配器(low-rank adapter)并结合文本提示,将大型预训练图像编辑模型适配到修复任务中。我们用从退化图像本身导出的指令替代该文本提示。该图像通过两条路径输入编辑模型:其结构信息来自模型的VAE,其语义指令则来自一个轻量级令牌映射器(token mapper),该映射器将退化图像的视觉-语言嵌入向清晰图像所对应的嵌入方向偏移。由于该指令是一个连续向量,通过对其缩放可以为目标不唯一的任务(如低光照增强)生成一族有效的修复结果。我们仅使用一个在单块GPU上训练约三小时的适配器,将一个Qwen-Image-Edit模型适配到六个任务。在同等对比条件下,图像指令优于文本条件控制;此外,它还支持无需退化标签的任务无关修复,而这正是文本方式所不具备的能力。
cs.CV / 6 / 2609.25270

RULER: Instance-aware Rubric Rewards for SVG Generation

RULER:面向SVG生成的实例感知评分准则奖励
Ran, Hangyu, Zheng, Yuhao, Zhang, Yingying, Lin, Kevin Qinghong, Peng, Han
Abstract
Generating Scalable Vector Graphics (SVG) code from natural-language instructions is an open-ended task without absolute visual ground truth, leaving both evaluation and policy optimization without a faithful signal. Scalar metrics (CLIP, Aesthetic) calibrated on natural images transfer poorly to stylized vector content, and reusing them as RL rewards triggers reward hacking. We address both limitations with rubric-based scoring. We first establish empirically that prompting a vision-language judge with a multi-axis rubric correlates with human judgments far better than scalar metrics, both across samples and within instructions. Building on this finding, we introduce RULER (Instance-aware Rubric Rewards for Reinforcement Learning), which converts each instruction into an instance-aware rubric of six items spanning semantic, visual, and stylistic axes; a judge VLM scores rendered rollouts item-by-item, and the weighted satisfactions form a fine-grained reward optimized via Group Relative Policy Optimization. Because the rubric is derived from text alone, RULER requires neither paired SVG ground truth nor human preference labels. On MMSVG-Illustration and MMSVG-Icon, RULER lifts the rubric score from 0.432/0.395 to 0.693/0.683, surpassing dedicated SVG specialists and matching the substantially larger DeepSeek-V3, with ablations identifying rubric design as the active lever for RL on open-ended SVG generation. The project page is available at https://hangyuran.github.io/RULER/.
Chinese Translation
从自然语言指令生成可缩放矢量图形(SVG)代码是一项开放性任务,不存在绝对的视觉真值,使得评估与策略优化都缺乏可靠的信号。基于自然图像校准的标量指标(CLIP、美学评分)难以迁移到风格化的矢量内容上,而将其直接用作强化学习(RL)奖励则会引发奖励欺骗(reward hacking)。我们通过基于评分准则(rubric)的打分方式同时解决上述两个局限。我们首先通过实证表明,使用多维度评分准则提示视觉语言评判模型(judge),其在样本间与同一指令内的相关性均远优于标量指标,更接近人类判断。基于这一发现,我们提出 RULER(Instance-aware Rubric Rewards for Reinforcement Learning),它将每条指令转换为涵盖语义、视觉与风格三个维度的六项实例感知评分准则;由评判 VLM 对渲染后的输出逐项打分,加权满意度构成细粒度奖励,并通过组相对策略优化(Group Relative Policy Optimization, GRPO)进行优化。由于评分准则仅由文本推导而来,RULER 既不需要配对的 SVG 真值,也不需要人类偏好标签。在 MMSVG-Illustration 和 MMSVG-Icon 数据集上,RULER 将评分准则得分从 0.432/0.395 提升至 0.693/0.683,超越了专门的 SVG 生成模型,并与规模大得多的 DeepSeek-V3 相当。消融实验表明,评分准则的设计是开放式 SVG 生成强化学习中的关键因素。项目页面见 https://hangyuran.github.io/RULER/。
cs.CV / 7 / 2609.25319

Uncertainty-Aware 3D Residual Wavelet Diffusion for Ultra Low-Field MRI Super-Resolution

面向超低场MRI超分辨率的不确定性感知三维残差小波扩散模型
Yeow, Rui W., Beament, Millie, Dick, Fred, Razin, Raha, Bocchetta, Martina, Thomas, David L., Tregidgo, Henry F. J., Alexander, Daniel C., Cole, James H.
Abstract
Ultra low-field MRI expands global access to neuroimaging but produces scans with low signal-to-noise ratio, reduced contrast, and thick slices. While regression-based super-resolution can recover anatomical detail for segmentation, it returns a single deterministic estimate that gives no indication of regions where the low-field input leaves anatomy underdetermined. Generative diffusion models offer an alternative by sampling the posterior distribution of plausible high-field images, quantifying this anatomical ambiguity. However, applying them to 3D whole-brain MRI is restricted by memory bottlenecks, slow sampling, and scanner domain shifts. We propose a 3D residual wavelet diffusion model that combines three ideas to overcome these hurdles. A lossless wavelet reparameterisation shrinks the spatial grid to fit a whole brain on a single GPU, residual shifting accelerates sampling by starting from the low-field input, and domain randomisation promotes scanner generalisation without paired training data. As the high-field reference is not a voxel-aligned ground truth, we evaluate downstream volumetric agreement. On a healthy cohort (n=19) imaged at 0.064T and 3T, our method matches a leading general-purpose regression approach in volumetric accuracy while additionally generating per-voxel uncertainty maps highlighting underdetermined regions. Furthermore, on a pilot dataset (n=11) of participants with cognitive impairment, disease-relevant atrophy is preserved rather than normalised towards a healthy prior. Our framework brings whole-brain posterior sampling to low-field super-resolution without sacrificing volumetric accuracy.
Chinese Translation
超低场MRI(Ultra low-field MRI)扩大了神经影像的全球可及性,但其产生的扫描图像信噪比低、对比度差且层厚较大。基于回归的超分辨率方法虽能恢复解剖细节以支持分割,但仅返回单一的确定性估计,无法指示低场输入在哪些区域使解剖结构不确定。生成式扩散模型提供了一种替代方案,通过采样合理高场图像的后验分布来量化这种解剖模糊性。然而,将其应用于三维全脑MRI受限于显存瓶颈、采样缓慢以及扫描仪域偏移。我们提出一种三维残差小波扩散模型,结合三个思路来克服这些障碍:无损小波重参数化压缩空间网格,使全脑数据可容纳于单个GPU;残差平移通过从低场输入出发加速采样;域随机化在无需配对训练数据的情况下促进扫描仪泛化。由于高场参考图像并非体素对齐的真值,我们通过下游体积一致性进行评估。在一个在0.064T和3T下成像的健康队列(n=19)上,我们的方法在体积精度上可与领先的通用回归方法相媲美,同时还能生成逐体素不确定性图以突出欠定区域。此外,在一个认知障碍受试者的先导数据集(n=11)上,疾病相关的萎缩得以保留,而未被归一化为健康先验。我们的框架在不牺牲体积精度的前提下,将全脑后验采样引入低场超分辨率。
cs.CV / 8 / 2609.25331

MirrorDistill: Illumination-Aware Latent Distillation for Efficient Low-Light Restoration

Mohsen, Farida, Zaim, Tala, Rusli, Nurul Izni, Al-Zawqari, Ali, Safa, Ali, Belhaouari, Samir Brahim
Abstract
Low-light image enhancement (LLIE) is an im- portant component of visual sensing systems operating under degraded illumination, including nighttime surveillance, au- tonomous navigation, remote sensing, and inspection in poorly lit industrial environments. Most LLIE methods rely on output- level reconstruction losses that supervise only the final restored image, leaving the intermediate feature recovery process weakly constrained. This paper proposes MirrorDistill, an illumination- aware latent distillation framework that links the low-light and clean domains through feature mirroring. During training, a shared encoder and an exponential-moving-average teacher decoder process the clean reference image to generate clean- domain latent targets. These targets supervise the low-light student at two levels: raw encoder features and standardized multi-scale decoder projections. The alignment is applied layer by layer, while a proposed illumination-aware weighting scheme gives greater emphasis to underexposed regions. The teacher and reference branches are used only during training, so inference requires only the lightweight student encoder-decoder and in- troduces no teacher-side computational cost. Under evaluation on the standard LOL benchmarks, MirrorDistill outperforms the state-of-the-art methods on the real-captured LOL-v2-Real set, while having the lowest compute complexity (GMACs) and while remaining competitive on the LOL-v1 and LOL-v2-Synthetic datasets. Ablation studies further show the contributions of the encoder mirror, decoder mirror, and illumination-aware weighting. Finally, we release our code as open-source for the benefit of future research.
cs.CV / 9 / 2609.25386

Sex Estimation from Footwear Outsole Impressions Using CNN Transfer Learning and Interpretable Image Statistics

基于CNN迁移学习与可解释图像统计的鞋底印痕性别估计
Niu, Jinyi, Song, Ziyi, Shen, Weining
Abstract
Footwear outsole impressions are a common form of forensic pattern evidence, yet quantitative methods for estimating wearer attributes from these images remain relatively underdeveloped. We investigate binary sex estimation from footwear outsole impressions by comparing convolutional neural network (CNN) transfer learning with traditional feature-based classification. Using a publicly available outsole-impression dataset, we adopt a shoe-level training and test partition that keeps replicate scans of the same physical shoe together to reduce data leakage. We evaluate pretrained CNNs through end-to-end fine-tuning, frozen feature extraction followed by support vector machine classification, and hybrid feature fusion incorporating handcrafted, geometric, and metadata-derived descriptors. Fine-tuned CNNs achieve the strongest overall predictive performance and substantially outperform traditional classifiers trained on the manually specified descriptors alone, while frozen-feature approaches offer a less computationally demanding alternative. Exploratory analysis of low-dimensional CNN representations reveals associations with frequency threshold ratio, image contrast, and wavelet-based summaries, providing a connection between learned representations and measurable properties of outsole impressions. These findings suggest that CNN transfer learning captures discriminative information beyond the descriptors considered and offers a promising approach to footwear-based forensic screening. Further validation on independently collected and casework-like impressions is needed before operational use.
Chinese Translation
鞋底印痕是法庭科学中常见的痕迹物证形式,然而从这类图像中估计穿着者属性的量化方法仍相对匮乏。我们研究了基于鞋底印痕的二分类性别估计问题,将卷积神经网络(CNN)迁移学习与传统基于特征的分类方法进行比较。利用一个公开的鞋底印痕数据集,我们采用鞋级别的训练与测试划分方式,将同一实体鞋的重复扫描图像保持在同一侧,以减少数据泄漏。我们通过三种方式评估预训练CNN:端到端微调、冻结特征提取后接支持向量机分类,以及融合手工特征、几何特征和元数据描述子的混合特征融合方法。微调后的CNN取得了最强的整体预测性能,显著优于仅在人工指定描述子上训练的传统分类器,而冻结特征方法则提供了一种计算开销更低的替代方案。对低维CNN表示的探索性分析揭示了其与频率阈值比、图像对比度以及基于小波的统计量之间的关联,从而在学习到的表示与鞋底印痕的可测量属性之间建立了联系。这些发现表明,CNN迁移学习能够捕获超出所考虑描述子范围之外的判别信息,为基于鞋类痕迹的法庭科学筛查提供了一种有前景的方法。在实际应用之前,仍需在独立采集的、类似案件实际的印痕样本上进一步验证。
cs.CV / 10 / 2609.25429

Directional Total Variation-Regularized Implicit Neural Representations (DTV-INR) for Continuous Super-Resolution in Degraded Imaging Domains

面向退化成像域连续超分辨率的方向性全变差正则化隐式神经表示(DTV-INR)
Kelishami, Mahmoud Saeedi
Abstract
In this paper, we introduce the Directional Total Variation-Regularized Implicit Neural Representation (DTV-INR), an advanced variational paradigm that synergistically integrates coordinate-driven implicit neural networks with an anisotropic, structure-tensor-informed total variation regularizer tailored for resolution-agnostic image super-resolution. Casting the continuous-to-discrete acquisition process into an ill-posed inverse problem framework, our formulation equips a SIREN-architected coordinate network with a dynamic Riemannian metric tensor field D(x). By leveraging its spectral decomposition, the proposed regularizer preferentially directs diffusion parallel to dominant structural contours while penalizing cross-edge dissipation, successfully circumventing the classical staircasing artifacts inherent to scalar total variation schemes. We rigorously prove the well-posedness of this formulation in H^1(Omega) by establishing the existence, uniqueness, and metric stability of the variational minimizer, and realize this via an alternating projected optimization algorithm that decouples network parameter tuning from adaptive tensor field updates. Comprehensive experiments conducted on clinical brain magnetic resonance imaging (MRI) and biomedical transmission electron microscopy confirm substantial quantitative and qualitative improvements, yielding PSNR enhancements reaching +5.05 dB over baseline unregularized INRs and +1.71-2.85 dB over isotropic TV-INR across continuous (non-integer) upsampling factors, alongside remarkable noise robustness up to sigma_eta = 0.10 and monotonic preconditioned convergence behavior.
Chinese Translation
本文提出方向性全变差正则化隐式神经表示(Directional Total Variation-Regularized Implicit Neural Representation, DTV-INR),这是一种先进的变分范式,将坐标驱动的隐式神经网络与各向异性、结构张量引导的全变差正则项协同结合,专门用于与分辨率无关的图像超分辨率。我们将连续到离散的成像采集过程转化为不适定逆问题框架,为采用SIREN架构的坐标网络配备动态黎曼度量张量场D(x)。借助其谱分解,所提出的正则项优先引导扩散沿主导结构轮廓的平行方向进行,同时抑制跨边缘的耗散,成功规避了标量全变差方法固有的经典阶梯伪影。我们通过建立变分极小元的存在性、唯一性和度量稳定性,严格证明了该表述在H^1(Omega)空间中的适定性,并通过一种交替投影优化算法加以实现,该算法将网络参数调节与自适应张量场更新解耦。在临床脑部磁共振成像(MRI)和生物医学透射电子显微镜数据上的全面实验表明,该方法在定量和定性上均取得显著提升:在连续(非整数)上采样倍数下,PSNR较无正则化的基线INR提升最高达+5.05 dB,较各向同性TV-INR提升1.71-2.85 dB,同时在噪声水平高达sigma_eta = 0.10时仍表现出优异的噪声鲁棒性以及单调的预条件收敛行为。
cs.CV / 11 / 2609.25453

Combinatorial Network-Based Manifold Topological Deep Learning for Image Analysis

Wachira, Alice, Liu, Xiang, Su, Zhe, Tong, Yiying, Wang, Ge, Wei, Guo-Wei
Abstract
Medical image analysis remains fundamentally challenging because of the intricate geometric and topological structures present in medical data. Conventional convolutional neural networks model images as regular Euclidean grids, limiting their ability to preserve geometric relationships and higher-order structural information. Recently, manifold topological deep learning (MTDL) has emerged as a promising paradigm that integrates deep learning with geometric and topological representations. Nevertheless, existing methods have not yet fully exploited discrete manifold structures within combinatorial complex neural networks. To bridge this gap, we introduce CNMTDL, a MTDL framework that integrates Hodge decomposition with a combinatorial attention mechanism. In our approach, medical images are represented as discrete manifolds and decomposed into three Hodge components. Features extracted from these components are concatenated and embedded into a combinatorial complex architecture, enabling enhanced higher-order message passing between $0$-cells and $2$-cells through attention-based blocks. We evaluate CNMTDL on six two-dimensional and three-dimensional datasets from the MedMNIST v2 benchmark, demonstrating its effectiveness for medical image analysis.
cs.CV / 12 / 2609.25454

MIND the Gap: A Geographic Implicit Neural Representation with Adjustable Spatial Scale

MIND the Gap:一种具有可调空间尺度的地理隐式神经表示
Corley, Isaac, Rao, Arjun, Rolf, Esther, Klemmer, Konstantin, Shelhamer, Evan, Lehmann, Nils, Rußwurm, Marc, Mai, Gengchen, Jacobs, Nathan, Kerner, Hannah
Abstract
Geographic measurements are often sparse, leaving large areas without labels for the quantities we want to map. Geographic implicit neural representations (INRs) address this by learning smooth, general-purpose embeddings that can be queried at any coordinate. Downstream models combine these embeddings with sparse labels to predict target values at unsampled locations without satellite imagery at inference. However, generalization to distant regions remains largely unexplored, despite its importance for remote sensing applications. We introduce Matryoshka Implicit Neural Distillation (MIND), which distills embeddings from specialist pretrained geospatial models into a single generalist coordinate embedding with adjustable spatial granularity. MIND uses nested supervision at several embedding dimensions, which define a series of contiguous chunks. In our experiments, early chunks capture coarser geographic variation, while later chunks add more fine-grained details. A downstream predictor can retain only leading chunks or be fitted with our Chunked Penalty to downweight later chunks while keeping the full embedding, without retraining the INR. To measure MIND and compare to existing approaches around the world, we introduce CoordBench, a large-scale INR evaluation suite of $52$ datasets and $78$ targets that aims to test both local interpolation and prediction in held-out regions at various spatial scales. MIND and its Chunked Penalty variant achieve the highest aggregate regression and classification scores among tested INRs, and the highest scores overall under regional holdout, setting a new state-of-the-art for geographic INRs.
Chinese Translation
地理测量数据通常是稀疏的,导致大片区域缺乏我们想要制图的量的标签。地理隐式神经表示(INR)通过学习可在任意坐标查询的平滑、通用的嵌入来解决这一问题。下游模型将这些嵌入与稀疏标签结合,在未采样位置预测目标值,且在推理时无需卫星影像。然而,尽管其对遥感应用十分重要,向遥远区域的泛化问题在很大程度上仍未被探索。我们提出了Matryoshka隐式神经蒸馏(MIND),它将专业预训练地理空间模型的嵌入蒸馏为一个具有可调空间粒度的单一通用坐标嵌入。MIND在多个嵌入维度上使用嵌套式监督,这些维度定义了一系列连续的数据块(chunk)。在我们的实验中,前面的数据块捕捉较粗粒度的地理变化,而后面的数据块则增加更细粒度的细节。下游预测器可以仅保留前面的数据块,或使用我们提出的分块惩罚(Chunked Penalty)进行拟合,在保留完整嵌入的同时降低后面数据块的权重,而无需重新训练INR。为了评估MIND并与全球范围内的现有方法进行比较,我们引入了CoordBench,一个大规模INR评估套件,包含52个数据集和78个预测目标,旨在测试局部插值以及在各种空间尺度下对保留区域(held-out regions)的预测。MIND及其Chunked Penalty变体在所测试的INR中取得了最高的总体回归和分类得分,并在区域保留测试下取得了整体最高得分,为地理INR树立了新的最先进水平。
cs.CV / 13 / 2609.25490

SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation

Gong, Jiangshan, Wu, Yuqun, Fu, Qiqian, Xiao, Yao, Zou, Chuhang, Wang, Shenlong, Hoiem, Derek
Abstract
Consistent multi-view object segmentation is critical for 3D perception and robotics, yet remains challenging under severe viewpoint and occlusion changes. Existing methods typically perform 3D instance segmentation on point clouds or rely on offline 2D mask-matching pipelines. However, 3D instance segmentation is limited by scarce 3D annotations, while offline 2D matching suffers from object identity ambiguity across frames. To leverage strong 2D and 3D priors jointly, we propose SAM-V (Geometry-Aware Segment Anything for Multi-View Instance Segmentation). Instead of combining the two priors through post-hoc matching, SAM-V directly integrates features from a feed-forward geometry model (VGGT) into a 2D segmentation foundation model (SAM), trained end-to-end for cross-view instance prediction. SAM-V introduces a prompt-fusion mechanism that enriches sparse SAM prompt tokens with view-specific camera tokens and local VGGT features, making the prompt representation both view-aware and spatially grounded, together with a mask decoder that attends to dense 2D and 3D features. By conditioning the mask decoding directly on multi-view geometry, SAM-V produces consistent multi-view segmentation of a prompted object in a single forward pass without offline mask matching or explicit 3D reconstruction. On the IGGT 3D tracking benchmark, where consistent instance identity across frames directly determines performance, SAM-V improves overall IoU by 5 points and frame-level recall by 12 points on the ScanNet++ split over the state-of-the-art multi-view instance segmentation baseline and leads on all metrics in the zero-shot ScanNet split. Our code and pretrained models are available at https://github.com/gong208/SAM-V.git.
cs.CV / 14 / 2609.25492

RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models

Wu, Zhiping, Ren, Dongdong, Zhou, Yangchengyu, Zhang, Zhengjie, Li, Wenbin, Pan, Hongbing, Gao, Yang
Abstract
Large vision-language models (VLMs) can be efficiently deployed under stringent memory and latency constraints through post training quantization (PTQ). However, most PTQ methods are designed for unimodal large language models (LLMs). These methods treat quantization errors as isotropic perturbations under the Euclidean assumption, which provides weak guidance on directions most sensitive to quantization in VLMs. Consequently, directly adapting unimodal PTQ approaches or solely employing modality-specific scaling often leads to uneven bit-width distribution and inconsistent performance in low-bit settings. To address these challenges, we propose Riemannian Geometry-Sensitive Quantization (RGSQ), which formulates quantization as a reconstruction problem under a unified Fisher-Riemannian metric. RGSQ identifies modality-specific sensitive directions via Riemannian manifold mappings built from modality-partitioned empirical Fisher factors and fused into a modality-aware Kronecker-structured metric. We then apply geometry-aligned rotations to reorient the local tangent frame, steering low-bit perturbations toward loss-insensitive axes. Finally, we apply a whitening transformation that maps the Riemannian objective to an equivalent Euclidean form, enabling standard unimodal PTQ methods to evaluate multimodal quantization error under their original assumptions. Across an extensive and diverse set of mainstream VLM benchmarks, RGSQ achieves the highest accuracy and stability under extremely low-bit settings (W2A8 and W3A8). It outperforms VLM-aware baselines, such as MBQ and MQuant, by up to 5.9% and surpasses single-modality improvements by up to 8.6%.
cs.CV / 15 / 2609.25500

mbariml: a curation pipeline for turning deep-sea imagery and video into object-detection training data

Lundsten, Lonny, Barnard, Kevin, Caress, Dave
Abstract
Training data quantity and quality greatly affect object detection model performance, regardless of model architecture. When using object detection models on video and images from the deep sea, in which the objects of interest, primarily organisms, are sparse, faint, and hard to identify, incremental improvements to object detector performance may require an iterative approach to data labeling and management. This paper presents mbariml, a python-based video and image analysis pipeline built around the data labeling management process. mbariml uses an Ultralytics YOLO detection model, runs it over still images or video, stores every detection as a reviewable region of interest, groups those regions by visual similarity so that a human can accept or reject them in bulk, and exports the result as training data, statistics, image sidecars, and additional metadata. The human review stage is the centre of the design: an annotator can validate, relabel, resize, delete, and draw entirely new localizations, and every one of those edits is written back to the same database the detector wrote to. Video receives particular attention: the software treats each tracker-produced track as a provisional observation and selects one representative frame instead of retaining every detection in the track. We describe the pipeline stage by stage, including the operational middle-third heuristic used for track observation selection.
cs.CV / 16 / 2609.25503

SBMVTrack: Spike-Budgeted Multi-View Learning for Energy-Efficient UAV Tracking

Zhong, Pengzhi, Mo, Jiwei, Li, Haolun, Zheng, Ge, Wang, Jingqi, Bo, Xinyi, Li, Shuiwang
Abstract
With sparse and event-driven computation, spiking neural networks show great potential for achieving accurate and energy-efficient UAV visual tracking. However, existing SNN-based trackers typically use spike firing rates only for energy evaluation and lack explicit optimization of actual spike activity. To address this, we propose SBMVTrack, a fully spiking framework for energy-efficient UAV tracking. SBMVTrack introduces Energy-Weighted Spike Budgeting (EWSB). EWSB weights actual spike activity according to the computational cost of each spiking layer. It constrains the energy-weighted firing rate and saturation activity, thereby reducing redundant spike computations. To improve tracking performance under the spike budget constraint, we propose Masked Multi-View Target Modeling (MVTM). This method treats the initial template, online template, and search region from the same sequence as correlated temporal views. It enhances the robustness of target representations through cross-view feature completion and identity-consistency learning. Extensive experiments on multiple benchmarks demonstrate that SBMVTrack effectively reduces the average spike firing rate and theoretical energy consumption. Meanwhile, it maintains competitive tracking performance, achieving a better accuracy-energy trade-off. The source code will be released upon acceptance.
cs.CV / 17 / 2609.25515

Real-World Perception for Autonomous Driving in Adverse Weather: Enhancing Standard Detectors via Foundation-Guided Auto-Annotation

恶劣天气下自动驾驶的真实世界感知:基于基础模型引导的自动标注增强标准检测器
Gohari, Sepideh, Mehr, Goodarz, Eskandarian, Azim
Abstract
Standard deployment-ready object detectors for autonomous vehicles degrade in adverse weather and lighting conditions without being trained on extensive domain-specific data. While large-scale vision foundation models offer robust zero-shot generalization, their high computational cost makes them impractical for real-time deployment. To bridge this gap, we propose a foundation-guided auto-annotation pipeline that enhances standard detectors without architectural changes. We first benchmark three distinct models, YOLOv8, Co-DETR, and SAM3, on our custom real-world driving dataset spanning 25 unique operational scenarios across various route, weather, and lighting conditions. Based on our analysis, SAM3 demonstrates superior accuracy and resilience across all scenarios. Thus, we deploy it as an offline auto-annotator to generate pseudo-labels on the unannotated subset of our dataset. Fine-tuning the baseline YOLOv8 on these annotations yields a 16.04% higher overall mean Average Precision (mAP) and improves cross-environmental stability compared to the baseline model, highlighted by a 32.73% and 28.65% mAP increase in Residential Direct Sunlight and Highway Fog, respectively. These results demonstrate that standard detectors can achieve environmental resilience without the need for extensive manual annotation or architectural modifications.
Chinese Translation
面向部署的标准目标检测器在自动驾驶车辆中,若未在大量领域特定数据上训练,其在恶劣天气和光照条件下的性能会下降。大规模视觉基础模型虽然具备强大的零样本泛化能力,但其高昂的计算成本使其难以实现实时部署。为弥合这一差距,我们提出了一种基础模型引导的自动标注流水线,无需更改架构即可增强标准检测器的性能。我们首先在自建的真实世界驾驶数据集上对三个不同的模型——YOLOv8、Co-DETR 和 SAM3——进行基准测试,该数据集涵盖跨越不同路线、天气和光照条件的 25 种独特运行场景。基于分析结果,SAM3 在所有场景中均展现出卓越的精度和鲁棒性。因此,我们将其部署为离线自动标注器,为数据集中未标注的子集生成伪标签。在这些标注上对基线模型 YOLOv8 进行微调后,整体平均精度均值(mAP)相比基线模型提高了 16.04%,并提升了跨环境的稳定性,其中在住宅区直射阳光和高速公路大雾场景下 mAP 分别提高了 32.73% 和 28.65%。这些结果表明,标准检测器无需大量人工标注或架构修改即可获得环境鲁棒性。
cs.CV / 18 / 2609.25538

Point Diffusion Mamba: Unified Diffusion-State-Space Modeling for Single-View 3D Reconstruction under Data Scarcity

Zhou, Wei, Shi, Xinzhe, Hao, Xingxing, Hao, Xing, Li, Kang, Peng, Jinye, He, Ying
Abstract
While single-view 3D reconstruction has seen significant progress, extrapolating complex 3D structures from inherently ambiguous 2D observations remains fundamentally ill-posed, particularly in the critically underexplored data-scarce regime. To address this challenge, we propose Point Diffusion Mamba (PDM), a method that integrates the generative power of diffusion models with the efficiency of state-space model for single-view 3D reconstruction under data-scarce conditions. Specifically, PDM employs a lightweight reconstruction module tailored to handle unordered point-cloud inputs effectively. By combining a Local Geometric Aggregation module with Mamba blocks, our approach jointly models global geometric structures and local details. In 3D reconstruction, each point in the initial noisy input requires a precise prediction, yet the high-level features extracted by the Mamba module capture only abstract semantic information from sparse points. To bridge this gap, we introduce the Hierarchical Feature Integration Network, which fuses high-level semantic and local geometric features for each point, overcoming the limitations of token-based point-cloud reconstruction. Furthermore, we propose a Dynamic Weighted Sampling strategy that adaptively unifies 3D generation with single-view reconstruction by leveraging generative priors to enhance reconstruction quality. Experimental results on the ShapeNet and Pix3D benchmarks demonstrate that PDM outperforms state-of-the-art methods, providing an effective solution for 3D reconstruction under data-scarce settings. Code is available at: https://github.com/NWUzhouwei/PDM.
cs.CV / 19 / 2609.25567

RootQuantV2: Adapting a Vision Foundation Model for Root-Trait Regression from Minirhizotron Imagery

RootQuantV2:基于微根管图像、面向根系性状回归的视觉基础模型适配方法
Parth, Kinjalk, Varela, Sebastian, Leakey, Andrew D. B.
Abstract
A lack of high-throughput phenotyping solutions for root traits in field-grown crops has severely constrained understanding and improvement of below-ground traits and processes. Minirhizotrons are the standard non-destructive root-phenotyping method in field environments. Computer vision solutions are needed to allow automated trait estimation at scale, but training data is scarce and human annotations are often inaccessible because they reside in proprietary software that only exports per-image scalar totals of root length and surface area. Nevertheless, large numeric archives of these root traits already exist. RootQuant showed that the traits can be predicted directly from the whole image by regression, thus removing manually traced masks from the pipeline; RootQuantV2 takes that idea further by replacing RootQuant's CNN backbone with a self-supervised ViT. We adapt a frozen DINOv3 ViT-L/16 with a hybrid parameter-efficient scheme. Training only 11.9M parameters (3.78% of the model), RootQuantV2 achieves length and area $R^2$ of 0.950 and 0.930, respectively, while lowering length/area RMSE by 24.3%/20.7% over RootQuant. RootQuantV2 thus repurposes legacy numeric archives for high-throughput, automated root trait estimation.
Chinese Translation
田间作物根系性状缺乏高通量表型分析解决方案,严重限制了对地下性状和过程的认识与改良。微根管(Minirhizotron)是田间环境下标准的非破坏性根系表型分析方法。为实现大规模自动化性状估计,需要计算机视觉解决方案,但训练数据稀缺,且人工标注往往难以获取,因为它们存储于专有软件中,只能导出每张图像的根长和根表面积标量总量。尽管如此,这些根系性状的大量数值档案已经存在。RootQuant 证明可以通过回归直接从整幅图像预测这些性状,从而将人工勾绘的掩膜从流程中去除;RootQuantV2 在此基础上进一步改进,用自监督的 ViT 替换了 RootQuant 的 CNN 骨干网络。我们采用混合参数高效方案适配冻结的 DINOv3 ViT-L/16 模型。仅训练 1190 万参数(占模型的 3.78%),RootQuantV2 的根长和表面积 R² 分别达到 0.950 和 0.930,同时相比 RootQuant 将根长/表面积的 RMSE 降低了 24.3%/20.7%。因此,RootQuantV2 能够将历史数值档案重新用于高通量、自动化的根系性状估计。
cs.CV / 20 / 2609.25578

Agentic Building-Aware Satellite Gaussian Splatting for Auditable Urban DSM Reconstruction

面向可审计城市DSM重建的智能体式建筑感知卫星高斯泼溅方法
Sun, Wentao, Xu, Zhengsen, Chen, Yiping, Zelek, John S., Li, Jonathan
Abstract
Urban-scale 3D reconstruction from satellite imagery supports disaster response, city monitoring, and geospatial digital twins, yet neural rendering methods typically optimize average visual fidelity rather than the structures that analysts inspect first: buildings. We present an agentic building-aware satellite Gaussian Splatting workflow that uses Segment Anything-derived building masks as semantic priors and an Agentic Reconstruction Controller to select, verify, and record DSM reconstruction policies. On the DFC2019 JAX\_004 scene, building-aware weighting reduces building-region DSM MAE from 0.844 m to 0.806 m, showing that semantic priors can shift reconstruction capacity toward analyst-critical regions. A staged schedule provides a balanced operating point, improving full-scene MAE from 1.362 m to 1.349 m while retaining a building gain. Across four JAX scenes, the Agent selects validated policies for both general DSM and building-focused DSM objectives, and produces building-inventory metadata and per-scene decision records. The system combines semantic priors, policy selection, region-specific DSM metrics, and DSM-derived GIS surface products for auditable urban 3D analysis.
Chinese Translation
基于卫星影像的城市尺度三维重建可支持灾害响应、城市监测与地理空间数字孪生,然而神经渲染方法通常优化平均视觉保真度,而非分析人员优先关注的结构——建筑物。本文提出一种智能体式(agentic)建筑感知的卫星高斯泼溅(Gaussian Splatting)工作流,该工作流利用源自 Segment Anything 的建筑物掩膜作为语义先验,并引入智能体重建控制器(Agentic Reconstruction Controller)来选择、验证并记录DSM(数字表面模型)重建策略。在 DFC2019 JAX_004 场景上,建筑感知加权将建筑物区域的DSM平均绝对误差(MAE)从0.844米降至0.806米,表明语义先验能够将重建能力向分析人员关注的关键区域倾斜。分阶段调度策略提供了一个均衡的运行点,在保留建筑物增益的同时,将全场景MAE从1.362米改善至1.349米。在四个JAX场景中,该智能体针对通用DSM与建筑聚焦DSM两类目标均选择了经过验证的策略,并生成建筑物清单元数据与逐场景决策记录。该系统融合了语义先验、策略选择、区域特定DSM指标以及由DSM导出的GIS地表产品,实现了可审计的城市三维分析。
cs.CV / 21 / 2609.25584

Hi-OPD: Hierarchy-Aware Open-Prompt Detection for Remote Sensing Images

Hu, Jinlong, Zhang, Yi, Xia, Zhiqi, Zhou, Yikang, Ji, Shunping
Abstract
Hi-OPD addresses a failure mode left uncontrolled by flat open-prompt training: descendant retrieval need not persist under ancestor queries when multi-source remote sensing annotations exhibit inconsistent granularity and missing labels. A detector may localize \textit{car} and \textit{van} under atomic prompts yet miss the same instances under \textit{vehicle}; flat AP does not expose this cross-level inconsistency. We propose Hi-OPD, a hierarchy-aware open-prompt detector, and construct RS153-HierOPD from 175,644 retained training image/tile records and 3.48M boxes mapped to 153 atomic categories with sparse hierarchy and alias relations. Hi-OPD learns ancestor retrieval through hierarchy-safe negative sampling, path multi-positive supervision, and one-way upward consistency, while per-source risk exclusion handles potentially missing labels. ConvVPE converts K-shot support boxes into text-compatible embeddings using detector-native features and the shared contrastive head. On Track A, Hi-OPD obtains 79.7/72.3 AP50 on DIOR/DOTA-v2.0, above the literature-reported OpenRSD results of 76.7/71.8. Under controlled training on the original converted annotations, the full hierarchy recipe raises DOTA-v2.0 parent AP50 from 7.2 to 71.5 and FAIR1M grandparent AP50 from 31.6 to 71.4, while DOTA-v2.0 atomic AP50 changes from 71.4 to 72.3. The text path reaches 99.7% CAR50 (0.3% violation) across the three common sources and 99.9%/0.1% on FAIR1M grandparent relations. On held-out VEDAI, text AP50 is 75.9, 6.2 points above OpenRSD. Joint AP and CAR show that explicit hierarchy training repairs this failure mode while retaining atomic detection and prompt transfer.
cs.CV / 22 / 2609.25597

Observer Choice and Threshold Selection in Retinal Vessel Segmentation: A Subject-Separated Evaluation

Xu, Wenhao, Kong, Yixian, Pan, Ting, Wang, Changwei, Wang, Feilong, Xu, Rongtao
Abstract
The annotation used to select a segmentation threshold is part of the evaluation protocol, yet its effect is easily conflated with model quality. We examine this choice for retinal vessel segmentation using all 28 CHASE DB1 images and both human annotations. A fixed seven-fold protocol keeps both eyes of each of the 14 subjects together. Random forests and Extra Trees are fitted against observer 1 with three random seeds, yielding 42 fits. Five threshold policies share identical score maps: fixed 0.50, observer-1 tuning, observer-2 tuning, mean-observer tuning, and maximin tuning of the per-image lower observer Dice. For random forests, maximin changes the threshold in 19 of 21 fits, but worst-observer Dice decreases from 70.53 percent to 70.45 percent. The paired difference is -0.073 percentage points, with a conditional subject-bootstrap 95 percent interval of [-0.384, 0.238]. Extra Trees shows the same direction. Identical observer-1-tuned random-forest masks score 73.66 percent against observer 1 and 71.06 percent against observer 2. The results support explicit reporting of both the threshold-selection reference and evaluation reference; they do not support an accuracy benefit from maximin tuning in this cohort. All splits, raw predictions, metrics and code are supplied. AI assistance is disclosed.
cs.CV / 23 / 2609.25604

Ultra-fast Neural Inference for Stochastic Gaussian Splatting Denoising

Hu, Chenxiao, Zhang, Hao, Zhang, Yanchen, Gai, Meng, Wang, Guoping, Li, Sheng
Abstract
Stochastic rendering eliminates the sorting and alpha blending process in Gaussian splatting, at the cost of introducing spatial noise. Formulating temporal denoising over the pixel stream shared by view-consistent stochastic splatting renderers, we propose a temporal neural denoiser validated on stochastic 2D Gaussian Splatting rendering, combining dual-path exponential moving average accumulation, per-pixel learned trust prediction for history validation, a fixed anisotropic spatial filter and a variance-gated composition with stabilization. The denoiser suppresses the noise, achieving temporally stable, visually compelling outputs during free camera navigation, all while retaining the sort-free, blend-free rasterization performance. The combined pipeline retains a PSNR gap to sorted alpha-blending renderers, but the denoiser's overhead stays below the time saved by removing sorting and blending.
cs.CV / 24 / 2609.25615

Evidence-gated multimodal parsing and vectorization of architectural floor plans

基于证据门控的建筑平面图多模态解析与矢量化
Chen, Hongxuan, Wang, Wenda, Lu, Jiachen, Shen, Qirui, Huang, Zilong, He, Lei, Dong, Xinyue, Huang, Weixin
Abstract
Architectural floor plans remain a high-friction barrier to archive digitization and early design-model preparation because heterogeneous graphics encode spatial semantics and editable geometry together. We introduce SALI-FP, an evidence-gated multimodal pipeline that converts a plan into reviewable semantic maps, objects, vectors, and relation records while constraining local revisions by image evidence. In a full production audit of 11,534 heterogeneous plans, SALI-FP produced structured outputs for every plan, including 752,510 valid polygon-bearing objects. The same output form has supported initial drawing digitization and design-model preparation in practical design work. Public-benchmark calibration is paired with a 30-case matched visual evidence set in Appendix F, where room-scale coverage, openings, oblique boundaries, and circulation continuity can be inspected directly. SALI-FP offers an engineering-oriented interpretation-to-geometry workflow for reviewed CAD/BIM preparation and existing-building information recovery.
Chinese Translation
建筑平面图由于异构图形将空间语义与可编辑几何信息编码在一起,仍然是档案数字化和早期设计模型准备中的高阻力障碍。我们提出了SALI-FP,一种证据门控的多模态流水线,它将平面图转换为可审查的语义图、对象、矢量及关系记录,并通过图像证据约束局部修订。在涵盖11,534张异构平面图的完整生产审计中,SALI-FP为每一张平面图生成了结构化输出,共包含752,510个有效的含多边形对象。同一输出形式已在实际设计工作中支持了初步图纸数字化和设计模型准备。公开基准测试校准与附录F中一个包含30个案例的匹配视觉证据集配套提供,可在其中直接检查房间尺度覆盖、门窗洞口、斜向边界及流线连续性。SALI-FP为经审查的CAD/BIM准备和既有建筑信息恢复提供了一种面向工程实践的解读到几何工作流。
cs.CV / 25 / 2609.25635

Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs

从浅层到深层:令牌剪枝与LVLM中层级行为角色的对齐
Zhang, Shuo, Tong, Jintao, Zou, Yixiong, Li, Yuhua, Li, Ruixuan
Abstract
Large Vision-Language Models (LVLMs) incur high computational costs from redundant visual tokens. Although training-free attention-based multi-layer pruning in the vision encoder stage has been explored as an effective strategy, we find that pruning in shallow layers consistently degrades performance. In this paper, we aim to understand this problem and seek a solution. By analyzing attention patterns across network depth, we find that shallow layers primarily function as edge detectors with chaotic attention maps, while deeper layers transition through local subject recognition and unstable semantic aggregation. To address the misalignment between pruning strategies and network stages, we propose STD, a hierarchical token pruning framework that adapts token selection mechanisms to the functional role of each network stage. STD employs High-Frequency Spectral Analysis in shallow layers to deterministically preserve structural edges, uses Gaussian-Smoothed Attention in intermediate layers to maintain spatial coherence, and introduces a Stability-Adaptive Trigger in deep layers to execute pruning only during semantically stable phases. Extensive experiments show that STD outperforms state-of-the-art pruning methods by 1.1% on LLaVA-1.5-7B with 88.9% token reduction, while also being plug-and-play and highly effective when combined with other methods, and by 2.1% on LLaVA-NeXT-7B with 94.4% reduction, delivering a 3.9x speed-up in the prefilling stage. Our code will be released at https://github.com/Twilight03/STD.
Chinese Translation
大型视觉-语言模型(LVLMs)因冗余的视觉令牌而产生高昂的计算成本。尽管在视觉编码器阶段采用免训练的基于注意力的多层剪枝已被证明是一种有效策略,但我们发现浅层的剪枝会持续导致性能下降。本文旨在理解该问题并寻求解决方案。通过分析网络深度上注意力模式的变化,我们发现浅层主要充当边缘检测器,其注意力图较为混乱;而更深层则经历局部主体识别和不稳定语义聚合的过渡阶段。为解决剪枝策略与网络阶段之间的不匹配问题,我们提出了STD,一种分层令牌剪枝框架,它根据每个网络阶段的功能角色调整令牌选择机制。STD在浅层采用高频频谱分析(High-Frequency Spectral Analysis)以确定性地保留结构边缘,在中间层使用高斯平滑注意力(Gaussian-Smoothed Attention)以维持空间连贯性,并在深层引入稳定性自适应触发器(Stability-Adaptive Trigger),仅在语义稳定的阶段执行剪枝。大量实验表明,在LLaVA-1.5-7B上,STD在88.9%的令牌削减率下比最先进的剪枝方法高出1.1%,同时具备即插即用的特性,与其他方法结合时也极为有效;在LLaVA-NeXT-7B上,STD在94.4%的削减率下高出2.1%,并在预填充(prefilling)阶段实现了3.9倍的加速。我们的代码将在 https://github.com/Twilight03/STD 发布。
cs.CV / 26 / 2609.25638

What Drives Hierarchy-Aware Image Retrieval? Taxonomy Alignment, Objective Choice, and Geometry

什么驱动了层次感知的图像检索?分类体系对齐、目标函数选择与几何结构
Shi, Ling
Abstract
Foundation vision models provide strong generic representations, yet high class-level retrieval accuracy does not necessarily imply that an embedding respects a target semantic taxonomy. We study strict explicit-taxonomy image retrieval on frozen DINOv2 features and ask: when hierarchical retrieval improves, how much of the change is associated with the organization of taxonomy-aware supervision, and how much with the Euclidean-hyperbolic geometry choice? We evaluate higher levels with strict cross-class criteria that exclude finer-grained matches, and compare Euclidean and hyperbolic projections trained with taxonomy-distance regression or a taxonomy-aware supervised contrastive objective. A compute-matched 2 x 2 Geometry x Loss factorial uses the same 768-256-32 projector capacity, optimization schedule, batch order, and fixed 100-epoch budget; the Loss axis denotes the Regression-to-Taxonomy-SupCon objective-family contrast. On CUB, the objective-family contrasts in mean hierarchy mAP (strict middle/high average, excluding Class/Leaf) are +0.0487 in Euclidean space and +0.0414 in hyperbolic space, compared with geometry contrasts of +0.0102 and +0.0030. On NABirds Parent-disjoint retrieval, the corresponding objective-family contrasts are +0.0467 and +0.0440, whereas geometry contrasts are +0.0017 and -0.0009. A semantic-alignment control shows that the true taxonomy substantially outperforms a structure-preserving shuffled hierarchy, while a NABirds curvature/radius control does not support stronger negative curvature as the explanation for the observed hierarchy gains. Across the two taxonomies, the Regression-to-Taxonomy-SupCon contrasts are larger in aggregate than the evaluated geometry contrasts; semantic alignment also matters separately, while geometry remains hierarchy-dependent.
Chinese Translation
基础视觉模型提供了强大的通用表征,但高类级检索精度并不一定意味着嵌入表示尊重目标语义分类体系。我们在冻结的 DINOv2 特征上研究严格的显式分类体系图像检索,并探讨:当层次化检索性能提升时,这种变化有多少与层次感知监督的组织方式相关,又有多少与欧氏-双曲几何选择相关?我们采用排除更细粒度匹配的严格跨类准则评估较高层级,并比较通过分类体系距离回归或分类体系感知的监督对比目标训练的欧氏投影与双曲投影。计算量匹配的 2 x 2 几何 x 损失因子实验使用相同的 768-256-32 投影器容量、优化调度、批次顺序以及固定的 100 个 epoch 预算;损失轴表示回归目标与分类体系感知监督对比(Taxonomy-SupCon)目标族之间的对比。在 CUB 数据集上,平均层次 mAP(严格的中/高层级平均,排除 Class/Leaf)的目标族对比在欧氏空间中为 +0.0487,在双曲空间中为 +0.0414,而几何对比分别为 +0.0102 和 +0.0030。在 NABirds 父节点不相交检索上,相应的目标族对比为 +0.0467 和 +0.0440,而几何对比为 +0.0017 和 -0.0009。语义对齐对照实验表明,真实分类体系显著优于保持结构的随机打乱层次体系;而 NABirds 上的曲率/半径对照实验不支持将更强的负曲率作为观察到的层次化增益的解释。在两个分类体系上,回归目标与 Taxonomy-SupCon 目标族之间的对比总体上大于所评估的几何对比;语义对齐也具有独立的重要作用,而几何结构的影响仍依赖于层次体系本身。
cs.CV / 27 / 2609.25650

Decoupling Disease, Covariates, and Individual Variability: A Unified Disentanglement Framework for Medical Image Classification

解耦疾病、协变量与个体差异:一种用于医学图像分类的统一解耦框架
Zhang, Shengjie, Zhang, Jinglin, Jiang, Zhuangzhuang, Yu, Ziqi, Zhang, Yipin, Zhang, Qi, Chen, Xiang, Yang, Haibo, Gao, Fei, Cui, Longbiao, Zhou, Yuan, Zhang, Xiao-Yong, Initiative, Alzheimer's Disease Neuroimaging
Abstract
Accurately isolating disease-related features from confounding covariates (e.g., age, gender, site) and individual variations remains a fundamental challenge in medical image classification. Traditional regression-based approaches may ignore non-linear relations between image features and true covariates. To overcome this issue, we present a generalized Medical Imaging Disentanglement Learning (MedIDL) framework. MedIDL maps image features into three mutually orthogonal latent spaces through specialized disentanglement heads: a disease classification head guided by a supervised loss, a covariate-alignment head constrained by cross-subject similarity matching, and a Gaussian head absorbing individual variations. We evaluated our framework across 7 datasets encompassing diverse imaging modalities. MedIDL outperforms state-of-the-art supervised and self-supervised classification methods in accuracy across all datasets. Association analyses demonstrate that MedIDL successfully isolates target-specific latent representations. Gradient-based interpretability mappings localize pathognomonic patterns aligning with established clinical literature.
Chinese Translation
如何从混杂协变量(如年龄、性别、检查机构)和个体差异中准确分离疾病相关特征,仍然是医学图像分类中的一项根本性挑战。传统的基于回归的方法可能忽略图像特征与真实协变量之间的非线性关系。为解决这一问题,我们提出了一个广义的医学影像解耦学习(Medical Imaging Disentanglement Learning, MedIDL)框架。MedIDL 通过专门的解耦头将图像特征映射到三个相互正交的潜在空间中:由监督损失引导的疾病分类头、受跨受试者相似性匹配约束的协变量对齐头,以及用于吸收个体差异的高斯头。我们在涵盖多种成像模态的7个数据集上对该框架进行了评估。MedIDL 在所有数据集上的分类准确率均优于最先进的监督和自监督分类方法。关联分析表明,MedIDL 能够成功分离出目标特定的潜在表示。基于梯度的可解释性映射定位出的特征性病理模式与已有临床文献相符。
cs.CV / 28 / 2609.25652

GameDirector: Decoupling Gameplay Logic from Rendering for Player-Configurable Game World Models

Lin, Zijun, Deng, Zhiyang, Wu, Yuzhe, Wen, Bihan, Jin, Yeying
Abstract
Recent game world models support realistic visual simulation and interactive gameplay based on player inputs. However, they typically learn environment dynamics from pixel-level supervision, jointly modeling perception, memory, state transitions, and rendering within a single end-to-end framework. While this design enables open-ended, action-controllable generation, it still falls short of delivering a complete gameplay experience. Games are governed by explicit mechanics, such as health deduction, skill activation, combat rules, and termination conditions. These mechanics depend on precise and consistent state transitions that generative models alone cannot reliably enforce. In contrast, game engines can guarantee such mechanics through hard-coded rules, but provide limited flexibility for player-driven creation. To bridge these paradigms, we introduce GameDirector, the first agentic framework that decouples rule-based gameplay logic from visual rendering. Given player-defined configurations, the framework acts as an intelligent director that interprets visual observations, updates game states, tactically controls NPCs, and enforces gameplay rules. It then translates these decisions into text prompts that guide the video world model to render the resulting gameplay. This separation allows players to configure characters, states, and rules much like a game developer while preserving coherent game mechanics. Experiments on three games, using data collected by our automated gameplay agent, show that GameDirector achieves accurate state tracking, reliable rule following, and improves boss action quality by more than 39.9% over various end-to-end game world model settings. Overall, by externalizing player-controllable game logic, GameDirector establishes a middle ground between hard-coded simulation and generative modeling, enabling more flexible and closed-loop gameplay experiences.
cs.CV / 29 / 2609.25684

Real-Time Atomic-Resolution Electron Phase Imaging without Probe Calibration via Ptychography-Supervised Learning

基于叠层衍射监督学习的无需探针标定的实时原子分辨率电子相位成像
Yue, H., Chen, C. -C., Hsiao, C. -N., Cheng, J., Liu, Y., Liao, X. Z., Shu, Steve F.
Abstract
Atomic-scale phase imaging is central to resolving defects, interfaces, and weakly scattering atoms that govern the behavior of nanoscale materials. Electron ptychography delivers sub-{\aa}ngstr\"om phase sensitivity but remains an offline technique, because its iterative reconstruction is computationally expensive and sensitive to experimental calibration, preventing live use during data acquisition. Here, a ptychography-supervised local inference framework is presented that converts four-dimensional scanning transmission electron microscopy (4D-STEM) into an acquisition-compatible phase-imaging workflow. Physics-constrained reference phase maps reconstructed from a single experimental AuPd dataset serve as teacher labels for a compact model that predicts local phase patches directly from diffraction measurements, without explicit probe input or online iterative optimization. Full-field images are assembled by deterministic overlap stitching. The workflow reaches an online latency of about 0.27 ms per probe position and a throughput of about 20,000 positions per second, an approximately 1,000-fold speed-up over GPU-accelerated ePIE, while preserving atomic-scale lattice contrast and reciprocal-space fidelity. Without fine-tuning, the same model transfers across materials (WS2), defocus conditions (high-entropy alloy nanoparticles), and instruments (hBN at 300 kV). The approach amortizes ptychographic redundancy into a fast, generalizable workflow that enables real-time atomic-scale phase imaging for materials microscopy.
Chinese Translation
原子尺度相位成像对于解析纳米材料行为中的缺陷、界面及弱散射原子至关重要。电子叠层衍射成像(electron ptychography)可提供亚埃级的相位灵敏度,但由于其迭代重建计算开销大且对实验标定高度敏感,目前仍是一种离线技术,无法在数据采集中实时使用。本文提出了一种叠层衍射监督的局部推理框架,将四维扫描透射电子显微术(4D-STEM)转化为与采集兼容的相位成像工作流程。由单一实验AuPd数据集重建的、受物理约束的参考相位图作为教师标签,用于训练一个紧凑模型,该模型可直接从衍射测量中预测局部相位图块,而无需显式的探针输入或在线迭代优化。全场图像通过确定性的重叠拼接组装而成。该工作流程实现在线延迟约为每个探针位置0.27 ms,吞吐量约为每秒20,000个位置,相比GPU加速的ePIE算法提速约1,000倍,同时保持了原子尺度的晶格衬度和倒易空间保真度。在未经微调的情况下,同一模型可跨材料(WS2)、跨离焦条件(高熵合金纳米颗粒)以及跨仪器(300 kV下的hBN)进行迁移。该方法将叠层衍射的冗余信息转化为快速、可泛化的工作流程,实现了材料显微学中的实时原子尺度相位成像。
cs.CV / 30 / 2609.25685

Initialization and Stopping Tolerance in CPU Dermoscopic Segmentation

Xu, Wenhao, Kong, Yixian, Pan, Ting, Wang, Changwei, Wang, Feilong, Xu, Rongtao
Abstract
Contour initialization and numerical stopping can jointly affect the evaluation of active-contour segmentation. We examine their interaction using the open-source scikit-image Chan-Vese implementation on a resized ISIC 2017 mirror. A fixed development set of 100 images selects a common input channel; all 600 images in the repository's held-out partition are then evaluated. Otsu thresholding is compared with checkerboard-, disk-, and Otsu-initialized contours under default and tighter level-set tolerances. At the default tolerance, Otsu initialization increases mean image Dice from 0.6011 to 0.6660 relative to checkerboard initialization, a paired difference of 0.0649 (95% image-bootstrap interval [0.0452, 0.0860]). Otsu thresholding alone achieves 0.6897. The default disk initializer stops after one iteration on 471 images. Tightening the tolerance reduces the Otsu-seed advantage over checkerboard initialization to 0.0197, with most runs reaching the 500-iteration limit. The default-tolerance advantage also reverses between small- and large-lesion strata. These findings show that an improvement over a generic initializer can coexist with deterioration relative to the threshold baseline. Evaluations should retain the unrefined mask as a comparator and report the initial-field definition, stopping tolerance, and observed iteration counts together.
cs.CV / 31 / 2609.25693

C2FXNet: Coarse-to-Fine Scene Expert for Unified Object Detection across Adverse Weather

C2FXNet:面向恶劣天气下统一目标检测的由粗到细场景专家网络
Fang, Tianle, Liu, Zhenbing, Yin, Chong, Li, Bolun, Lu, Haoxiang
Abstract
Object detection in adverse weather remains challenging because severe degradations weaken visual quality and disrupt semantic feature representations across diverse scenes. Existing methods usually rely on condition-specific designs, which limits their ability to generalize within a unified detector. In this paper, we propose a Coarse-to-Fine Scene Expert Network (C2FXNet) that achieves unified detection through hierarchical scene guidance. Specifically, C2FXNet introduces a dual-level guidance mechanism consisting of a Multi-step Reasoning Router (MRR), which performs GRU-based recurrent scene reasoning over compressed multi-scale visual cues and frozen coarse scene prototypes, and a Fine Scene Refinement (FSR) module, which uses image-specific semantic cues to modulate high-level features for local variation handling. Furthermore, a Scene-aware Mixture-of-Experts (SMoE) dynamically combines scene-specific experts under the joint guidance of MRR and FSR. By coupling coarse scene reasoning with fine-grained semantic refinement, C2FXNet enables robust multi-scene detection without scene-specific training. Extensive experiments on RTTS, ExDark, and our newly constructed Adverse Weather Dataset (AWD) demonstrate that C2FXNet consistently outperforms state-of-the-art methods across foggy, dark, and clear conditions, reaching 63.70%, 71.14%, and 54.19% mAP on RTTS, ExDark, and AWD, respectively. The source code will be released at https://github.com/PolarisFTL/C2FXNet.
Chinese Translation
恶劣天气下的目标检测仍然具有挑战性,因为严重的图像退化会削弱视觉质量,并破坏多样化场景中的语义特征表示。现有方法通常依赖于特定条件的设计,这限制了它们在统一检测器中的泛化能力。在本文中,我们提出了一种由粗到细场景专家网络(C2FXNet),通过分层场景引导实现统一检测。具体而言,C2FXNet 引入了一种双层引导机制,包括:多步推理路由器(Multi-step Reasoning Router, MRR),它基于 GRU 对压缩的多尺度视觉线索和冻结的粗粒度场景原型进行循环场景推理;以及精细场景细化(Fine Scene Refinement, FSR)模块,它利用图像特定的语义线索来调制高层特征,以处理局部变化。此外,场景感知混合专家(Scene-aware Mixture-of-Experts, SMoE)在 MRR 和 FSR 的联合引导下动态组合场景特定的专家。通过将粗粒度场景推理与细粒度语义细化相耦合,C2FXNet 无需针对特定场景进行训练即可实现鲁棒的多场景检测。在 RTTS、ExDark 以及我们新构建的恶劣天气数据集(AWD)上的大量实验表明,C2FXNet 在雾天、黑暗和晴天条件下均持续优于最先进的方法,在 RTTS、ExDark 和 AWD 上分别达到 63.70%、71.14% 和 54.19% 的 mAP。源代码将在 https://github.com/PolarisFTL/C2FXNet 发布。
cs.CV / 32 / 2609.25697

Interpretable AI plus Handheld, Portable Retinal Photographs: A Low-Cost Glaucoma Screening Solution for West Africa

可解释人工智能结合手持式便携视网膜照片:面向西非的低成本青光眼筛查解决方案
Chiang, Charis Y. N., Sarimiye, Tarela, Ashaye, Adeyinka, Buist, Martin, Hauser, Michael A., Olawoye, Olusola, Girard, Michaël J. A.
Abstract
Purpose: To develop and evaluate an interpretable artificial intelligence (AI) framework for glaucoma screening from low-cost portable, handheld retinal fundus photographs in a West African population and to compare its performance with clinical tabletop fundus imaging. Methods: We used data from a community-based study of 681 participants (1,362 eyes) in Nigeria, comprising 414 glaucoma, 478 glaucoma suspect, and 470 non-glaucoma eyes. Fundus photographs were acquired using the low-cost handheld, portable Volk Viva retinal camera and the Canon CR-2-AF tabletop camera. We fine-tuned component models separately to each device to perform vessel segmentation, cup and disc boundary segmentation, and feature extraction to detect optic nerve head features. A final classification model combined these components to classify scans as glaucoma, glaucoma suspect or non-glaucoma. Feature-weight analysis and Gradient-weighted Class Activation Mapping were used for interpretation. Results: The models performed well on both Volk Viva and Canon CR-2-AF images: Vessel segmentation: 0.98 Dice Coefficient (DC) (Volk) and 0.94 DC (Canon); Cup and disc segmentation: 0.95 DC (Volk) and 0.96 DC (Canon); Optic nerve head feature detection: area under the receiver operating characteristic curve (AUCs) of 0.83$\pm$0.03 (Volk) and 0.87$\pm$0.04 (Canon); Classification model: AUCs of 0.85$\pm$0.01 (Volk) and 0.93$\pm$0.01 (Canon). Reports for each image, present model decision confidence scores and decision-rationale visualizations to support clinical interpretation. Conclusions: Volk Viva results were reasonably comparable to Canon CR-2-AF in the component models and not far behind in classification. This shows that interpretable AI combined with low-cost, portable imaging may enhance community-level glaucoma screening, especially in settings with limited specialist access and resources.
Chinese Translation
目的:开发并评估一个可解释的人工智能(AI)框架,利用低成本的便携式手持视网膜眼底照片在西非人群中进行青光眼筛查,并将其性能与临床台式眼底成像进行比较。方法:我们使用了尼日利亚一项基于社区的研究数据,共纳入681名参与者(1362只眼),其中包括414只青光眼、478只疑似青光眼和470只非青光眼。眼底照片分别使用低成本的手持式便携Volk Viva视网膜相机和Canon CR-2-AF台式相机采集。我们针对每种设备分别微调了各组件模型,以执行血管分割、视杯和视盘边界分割以及用于检测视神经头部特征的特征提取。最终的分类模型将这些组件结合起来,将扫描结果分类为青光眼、疑似青光眼或非青光眼。采用特征权重分析和梯度加权类激活映射(Gradient-weighted Class Activation Mapping)进行结果解释。结果:模型在Volk Viva和Canon CR-2-AF图像上均表现良好:血管分割:Dice系数(DC)分别为0.98(Volk)和0.94(Canon);视杯和视盘分割:DC分别为0.95(Volk)和0.96(Canon);视神经头部特征检测:受试者工作特征曲线下面积(AUC)分别为0.83±0.03(Volk)和0.87±0.04(Canon);分类模型:AUC分别为0.85±0.01(Volk)和0.93±0.01(Canon)。针对每张图像生成的报告提供模型决策置信度评分和决策依据可视化,以支持临床解读。结论:在组件模型方面,Volk Viva的结果与Canon CR-2-AF相当接近,在分类方面差距也不大。这表明可解释AI结合低成本、便携式成像设备有望加强社区层面的青光眼筛查,尤其是在专科医疗资源和设备有限的环境中。
cs.CV / 33 / 2609.25716

FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance

Lew, Jaihyun, Jung, Mingi, Park, Minjun, Song, Wooseok, Yoon, Sungroh
Abstract
Reference-based image quality assessment (IQA) metrics aim to reflect how humans perceive the perceptual distance between a pair of images. To learn how the human visual system (HVS) operates, recent reference-based IQA metrics heavily rely on human-annotated data. Mean opinion score (MOS)-based pointwise scoring, which assigns a scalar quality value per image, is preferable for annotation but is prohibitively expensive to collect at scale and is known to be noisy due to inconsistent human judgments. As an alternative, two-alternative forced choice (2AFC) pairwise labels have gained popularity due to their reliability and efficiency, but they capture only relative comparisons between pairs. In this paper, we propose a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation. Our approach exploits the generative dynamics of diffusion models as a perceptual distance proxy, where the coarse structure of an image is generated in the early timesteps and the fine details are generated in the later timesteps. Images that fork early in the generation process share only coarse structure and are perceptually far apart; images that fork late differ only in fine detail. We demonstrate that the diffusion trajectory aligns well with the human visual system, and use this forking moment, FoMo, as a reference-grounded distance label to supervise the training of a reference-based IQA metric. The pointwise labels, which support universal comparison between arbitrary image pairs, enable an information-rich training objective. Extensive experiments across diverse backbone architectures confirm the effectiveness of our generation pipeline, outperforming human-annotated datasets in multiple benchmarks.
cs.CV / 34 / 2609.25731

Annual Earth-observation embeddings encode wildfire disturbance and support simplified burned area mapping

Knezevic, Jovana, Atzberger, Clement, Feng, Zhengpeng, Pellegrini, Adam F. A., Keshav, Srinivasan, Coomes, David
Abstract
Medium-resolution (10-30 m) burned area mapping is vital for monitoring wildfires and their impacts, but remains difficult to scale. Existing methods require either curated fire-specific imagery or dense time-series analysis. Here, we tested whether annual Earth-observation embeddings retain wildfire disturbance signals sufficiently to map burned areas without either requirement. Using Tessera and AlphaEarth embeddings, we tested individual burn-scar delineation, mapping of all same-year fires within an area, regional wall-to-wall mapping, cross-continental transfer, and intra-annual fire timing. Tessera strongly encoded wildfire disturbance, allowing even linear models to separate burned from unburned pixels; the signal was weaker in AlphaEarth. Models trained on a single Tessera embedding matched or exceeded equivalent models using paired pre- and post-fire HLS imagery, and outperformed post-fire imagery alone. The same approach mapped all same-year fires within benchmark scenes (F1 = 0.90). Applied across California, with no California fire data used for downstream training, it recovered 97% of reference burned area and detected substantially more small and medium-sized fires than GABAM or MCD64A1. Separately, a model trained on 2018-2021 US fires transferred without retraining to 88 European fires from 2024-2025 (F1 = 0.88). For well-detected fires, ignition timing was recovered with a mean absolute error of 13 days. Performance declined for fires ignited near the end of the calendar year, and wall-to-wall deployment produced systematic false positives in some unseen landscapes. Annual embeddings nevertheless achieve high segmentation accuracy while moving the burden of dense time series processing upstream, providing a promising path towards simpler regional burned area mapping.
cs.CV / 35 / 2609.25741

Fysiverse-3D-Vision Technical Report: Generating Executable 3D Worlds from Images through Unified Spatial Reasoning

Fysiverse-3D-Vision 技术报告:通过统一空间推理从图像生成可执行的3D世界
Yang, Dingkang, Liu, Yizhou, Cheng, Wendong, Chen, Zizhi, Wang, Shunli, Liu, Yang, Li, Hongsheng, Zhang, Lihua
Abstract
Generative models have advanced image-conditioned 3D content creation, yet generating controllable and executable 3D scenes from a single image remains challenging. Existing 3D generative approaches can synthesize visually plausible objects and scenes, but their spatial layout estimation is coupled with specific asset generators. They struggle to jointly model object semantics, metric geometry, and scene-level spatial relationships, which are essential for interactive editing, physical simulation, and embodied applications. We propose Fysiverse-3D-Vision, a unified vision-language-geometry framework for generative 3D scene reconstruction and executable asset construction from a single image. We establish a shared representation where spatial reasoning and geometric reconstruction mutually enhance each other, allowing object layouts to be inferred beyond the constraints of individual asset generators. Our model integrates textual supervision, semantic visual cues, and geometric representations within a unified Transformer to capture scene context, metric geometry, and object-level interactions. An object-conditioned layout module performs cross-attention between target object representations and global geometric features to predict object translation, rotation, and scale. Training progressively learns geometry-language alignment, introduces layout reasoning while preserving reconstruction capability, and refines physical consistency through collision-aware optimization. By separating spatial layout reasoning from asset synthesis, Fysiverse-3D-Vision provides an adaptable interface for interactive scene editing, object-level manipulations, and executable 3D content generation. Experiments demonstrate that our framework achieves superior geometric consistency, layout estimation, rendering quality, and physical property understanding compared with existing approaches.
Chinese Translation
生成式模型推动了图像条件下的3D内容创作,然而从单张图像生成可控且可执行的3D场景仍具挑战性。现有的3D生成方法虽然能够合成视觉上逼真的物体和场景,但其空间布局估计与特定的资产生成器相耦合,难以联合建模物体语义、度量几何以及场景级空间关系,而这些对于交互式编辑、物理仿真和具身应用至关重要。我们提出Fysiverse-3D-Vision,这是一个用于从单张图像进行生成式3D场景重建与可执行资产构建的统一视觉-语言-几何框架。我们建立了一种共享表示,使空间推理与几何重建相互增强,从而能够在突破单一资产生成器限制的情况下推断物体布局。我们的模型在统一的Transformer中整合了文本监督、语义视觉线索和几何表示,以捕捉场景上下文、度量几何以及物体级交互。物体条件布局模块在目标物体表示与全局几何特征之间执行交叉注意力,以预测物体的平移、旋转和缩放。训练过程逐步学习几何-语言对齐,在保持重建能力的同时引入布局推理,并通过碰撞感知优化提升物理一致性。通过将空间布局推理与资产生成分离,Fysiverse-3D-Vision为交互式场景编辑、物体级操作和可执行3D内容生成提供了灵活的接口。实验表明,与现有方法相比,我们的框架在几何一致性、布局估计、渲染质量和物理属性理解方面均表现出色。
cs.CV / 36 / 2609.25743

SAMI3D-DW: Interactive Segmentation of Any 3D Medical Images

Gong, Ping, Su, Shiyuan, Zhang, Fandong, Han, Xinchen, Sun, Haowei, Li, Yiming, Yu, Yizhou
Abstract
Interactive segmentation of 3D medical images supports quantitative analysis of anatomical structures and disease while allowing users to specify and refine their targets. Despite substantial progress by nnInteractive and VISTA3D, reliable segmentation across diverse clinical targets remains challenging, particularly for complex anatomical structures and the heterogeneous, long-tailed spectrum of pathology. We present SAMI3D-DW V1 (hereafter SAMI3D-DW), an interactive 3D segmentation model trained on Deepwise's large-scale proprietary medical image datasets. We evaluate the model under simulated user interactions on a CT/MR benchmark comprising 4,326 cases from 219 source datasets, spanning 107 anatomical and pathological categories, organized by a medical taxonomy and evaluated with a category-balanced DSC score. SAMI3D-DW achieves the highest category-macro Dice among evaluated methods in both interaction modes. With one point, it scores 0.5764 versus 0.5315 for nnInteractive, the strongest baseline, rising to 0.7771 versus 0.7494 with five points. With bounding-box initialization, the scores are 0.7130 versus 0.6530. After five corrective clicks, SAMI3D-DW reaches 0.8002 versus 0.7868, making it the only evaluated box-compatible model to exceed 0.80. For radiologists and clinicians, SAMI3D-DW enables segmentation of complex anatomical structures, including intracranial vessel trees on CT and MR angiography, with a few clicks. In a preliminary in-house comparison involving neurofibromatosis type 1 (NF1), SAMI3D-DW-assisted tumor annotation took minutes per case and approximately one-fifteenth of the time required for manual annotation, highlighting its potential to support volumetric treatment-response assessment.
cs.CV / 37 / 2609.25770

Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs

Lin, Dingyang, Luo, Yingfeng, Wang, Chenglong, Zhu, Chenwei, Ma, Anxiang, Zhu, Jingbo, Xiao, Tong
Abstract
Vision-language models (VLMs) have achieved strong performance on tasks such as visual question answering, yet small image resizes can turn correct answers into errors. We investigate whether changes in visual configuration, such as image tiling and token arrangement, contribute to this instability. Across seven checkpoints and four benchmarks, equally small resizes cause more correctness flips when they switch configurations. Surprisingly, in over half of these cases, models answer the question incorrectly but can still read the correct answer when told what to read. Furthermore, attention interventions in LLaVA-NeXT suggest that configuration changes can weaken the use of readable information during answering. We therefore guide models using field cues and their own transcriptions. With annotation assistance, these forms of guidance together correct 97.2% of errors with readable information. These findings show that configuration changes can affect how models use information they can still read.
cs.CV / 38 / 2609.25773

Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models

Video-HopChain:面向视频推理模型的多跳问题与置信度门控探索
Quang, Trung Nguyen, Dong, Yuhao, Sun, Shuo, Liu, Shuai, Tian, Shulin, Yap, Kim-Hui, Liu, Ziwei
Abstract
HopChain has shown on still images that multi-hop data synthesis improves vision-language reasoning, because long chain-of-thought reasoning exposes errors that compound across steps, while most data used for reinforcement learning with verifiable rewards (RLVR) rarely demands a chain of visual evidence, so these weaknesses are likely to stay unexposed. We observe the same problem in video, where this framework has not yet been explored. We therefore build Video-HopChain, a dataset of 22,550 multi-hop video questions over 13,378 videos, together with a held-out benchmark of 1,000 questions. Each question chains three to six yes/no questions about moments in one video, and each yields one of two integers depending on its answer. The final answer is the sum of these integers, so an exact match on that sum gives the verifiable reward that RLVR needs. We first train Qwen3-VL-8B with GRPO on a standard video dataset, and a second stage on Video-HopChain then raises the mean over eight video understanding and reasoning benchmarks from 55.4 to 57.9 and improves every one of them. Training on such a dataset, however, exposes a known limitation of GRPO: its learning signal comes from the reward variance within a group, so hard questions whose rollouts are all incorrect and easy questions whose rollouts are all correct both leave the group with no gradient. To recover these groups at the same compute budget, we introduce Confidence-Gated Exploration (CGE). With 8 rollouts per question, CGE samples the first 4 as usual. If these 4 are either all correct or all incorrect, it samples the last 4 with the policy's most confident token masked inside the reasoning span, and removes the masked positions from the loss while all 8 rollouts enter the advantage. With CGE, the mean rises further to 59.3. We release the dataset, the checkpoint, and the data generation and training code.
Chinese Translation
HopChain 已在静态图像上证明,多跳数据合成能够提升视觉-语言推理能力,因为长链式思维(chain-of-thought)推理会暴露在多个步骤间累积的误差,而用于可验证奖励强化学习(RLVR)的大多数数据很少需要一条视觉证据链,因此这些缺陷很可能一直未被发现。我们在视频领域观察到同样的问题,而该框架尚未在视频上被探索过。为此,我们构建了 Video-HopChain,一个包含基于 13,378 个视频的 22,550 个多跳视频问题的数据集,并附带一个包含 1,000 个问题的保留基准测试集。每个问题将关于同一视频中某些时刻的三到六个是非题串联起来,并且根据答案产生两个整数之一。最终答案是这些整数之和,因此对该和值的精确匹配即可提供 RLVR 所需的可验证奖励。我们首先使用 GRPO 在标准视频数据集上训练 Qwen3-VL-8B,随后在 Video-HopChain 上进行第二阶段训练,将八个视频理解与推理基准的平均分从 55.4 提升至 57.9,且在每个基准上均有提升。然而,在此类数据集上训练暴露了 GRPO 的一个已知局限:其学习信号来自组内的奖励方差,因此对于所有采样结果都错误的高难度问题和所有采样结果都正确的简单问题,该组均不会产生梯度。为了在相同计算预算下挽救这些组,我们提出了置信度门控探索(Confidence-Gated Exploration, CGE)。在每题 8 个采样结果下,CGE 仍按常规采样前 4 个;若这 4 个结果全对或全错,则在采样后 4 个时,将策略在推理片段内最置信的 token 进行遮蔽,并从损失中剔除被遮蔽的位置,同时所有 8 个采样结果均参与优势估计。引入 CGE 后,平均分进一步提升至 59.3。我们公开了数据集、模型检查点以及数据生成与训练代码。
cs.CV / 39 / 2609.25775

TRACE: Trajectory Representation and Consistency Estimation for AI-Generated Video Detection

Cao, Huangsen, chu, Hongkang, Yu, Siyao, Ding, Xin, Dong, Jianfeng, Wang, Yongwei
Abstract
Recent advances in generative video models have enabled the synthesis of visually realistic content, posing significant challenges to synthetic video detection. Existing detectors often rely on appearance artifacts, semantic inconsistencies, and temporal patterns that may be generator-specific, limitating generalization to unseen synthesis models. We investigate whether responses to a pretrained generative model provide more transferable forensic cues. Our key observation is that real and AI-generated videos exhibit distinct \emph{velocity responses} under a pretrained Flow Matching video model. This distinction persists when different pretrained video-generation backbones are used as probes, suggesting that velocity responses offer transferable forensic signals beyond visual artificts. Motivated by this observation, we propose \textbf{TRACE} (\emph{\underline{T}rajectory \underline{R}epresentation \underline{a}nd \underline{C}onsistency \underline{E}stimation}), a generation-process-aware framework for AI-generated video detection. TRACE leverages a pretrained video DiT as a velocity-field probe to extract representations at multiple flow time points, and models cross-frame consistency through velocity differences between adjacent frames. We further introduce a \emph{Real-Centered Trajectory Optimization} objective that encourages generator-invariant representation learning. Extensive experiments on AIGVDBench demonstrate that TRACE generalizes effectively across diverse generators, substantially outperforming prior state-of-the-art methods on unseen open- and closed-source video generation models.
cs.CV / 40 / 2609.25793

When Point Clouds Outperform Pixels: Rethinking Zero-Shot Multimodal Anomaly Detection

当点云超越像素:重新思考零样本多模态异常检测
Ye, Chenglin, Liu, Lupeng, Yu, Dongbo, Xiao, Jun, Wang, Yunbiao
Abstract
Zero-shot multimodal anomaly detection commonly assumes that RGB and point cloud modalities are equally reliable and can contribute uniformly to anomaly localization. We challenge this assumption. Using a set of recently proposed stringent metrics that penalize false anomaly responses in normal regions, we find that point clouds are substantially more reliable than RGB under zero-shot category shift. Motivated by this observation, we propose WOOPS (\textbf{W}hen P\textbf{o}int Cl\textbf{o}uds Out\textbf{p}erform Pixel\textbf{s}), a reliability-aware zero-shot multimodal anomaly detection framework. To strengthen the more reliable geometric modality, we design a Multi-view Information Decoupling module to suppress heterogeneous information from multi-view point cloud projections and enhance point cloud feature quality. To avoid unconditional fusion, we further introduce a Modality Reliability Calibration module to adaptively calibrate modality contributions according to their reliability. Extensive experiments show that our method achieves the best or competitive performance under the new metrics in both unimodal and multimodal settings. Further analysis demonstrates that point cloud information also improves RGB-only inference, while ablations verify the effectiveness of both modules. Code will be released upon acceptance.
Chinese Translation
零样本多模态异常检测通常假设RGB和点云两种模态同样可靠,并且能够均匀地贡献于异常定位。我们对这一假设提出了质疑。通过采用一组近期提出的、惩罚正常区域虚假异常响应的严格评估指标,我们发现在零样本类别迁移下,点云比RGB具有显著更高的可靠性。基于这一观察,我们提出了WOOPS(When Point Clouds Outperform Pixels,当点云超越像素),一种可靠性感知的零样本多模态异常检测框架。为了增强更可靠的几何模态,我们设计了多视角信息解耦(Multi-view Information Decoupling)模块,以抑制多视角点云投影中的异构信息,提升点云特征质量。为了避免无条件融合,我们进一步引入模态可靠性校准(Modality Reliability Calibration)模块,根据各模态的可靠性自适应地校准其贡献。大量实验表明,在单模态和多模态设置下,我们的方法在新指标下均取得了最优或具有竞争力的性能。进一步分析表明,点云信息也能提升仅使用RGB的推理效果,消融实验验证了两个模块的有效性。代码将在论文被接收后发布。
cs.CV / 41 / 2609.25803

LiFR v2: Completion-Augmented Event Propagation for High-Rate Dense Prediction

Wan, Tao, Wu, Xiaoshan, Yu, Yifei, Wang, Bo, Lyu, Xiaoyang, Liu, Muxin, Pan, Aoxuan, Wang, Zhongrui, Qi, Xiaojuan
Abstract
High-rate dense perception in dynamic environments is limited by the low update rate of RGB cameras, as rapid scene changes can occur between frames. Event cameras offer temporally dense but spatially sparse measurements, complementary to spatially dense RGB observations. Direct fusion cannot fully exploit this complementarity, while event-guided propagation fails on newly appearing or disoccluded regions without valid RGB support. We present LiFR v2, a unified propagation-completion-memory framework for causal anytime and streaming dense prediction from an RGB keyframe and events. LiFR v2 introduces an Event-Guided Completion Module (EGCM) to recover task-relevant representations where propagation is unsupported, and a History Retrieval Module (HRM) to reuse completed representations across successive queries. The framework supports semantic segmentation, monocular depth estimation, and multi-task dense prediction, and we further introduce SHF-Emerge to evaluate rapid object emergence and disocclusion. LiFR v2 achieves 74.37% mIoU on DSEC and 56.13% on SHF-Emerge, improving LiFR-Seg by 1.85 percentage points on the latter, while reducing SHF-Emerge depth RMSE from 1.564 m to 1.118 m over the propagation baseline. It also exceeds 100 FPS for both segmentation and depth, demonstrating accurate and efficient high-rate perception beyond RGB frame rates.
cs.CV / 42 / 2609.25815

MorphoSHAP: Rethinking the Unit of Attribution in Explanation for Deep Visual Models

Prabhakaran, Anirudh, Rocchi, Alexandre, Franchi, Gianni
Abstract
Visual attribution methods typically explain predictions using pixels, superpixels, or regular patches. These representations can localize important regions, but provide limited information about their structure. We introduce MorphoSHAP, a model-agnostic post-hoc method that instead uses morphological shapes as the players of a Shapley attribution game. Using the Tree of Shapes, each shape is described by its scale, geometry, and signed contribution, providing explanations of where the evidence lies, what type of structure carries it, and how strongly it affects the prediction. This shared morphological vocabulary enables spatial, textual, and global class-level explanations beyond image-specific heatmaps. To the best of our knowledge, MorphoSHAP is the first SHAP-based image attribution framework to combine these different forms of explanation. Across five diverse datasets and three architectures, MorphoSHAP achieves strong insertion/deletion performance and outperforms competing attribution methods on several benchmarks. Finally, a user study shows that MorphoSHAP provides explanations that are easy to use and are preferred over standard attribution baselines.
cs.CV / 43 / 2609.25832

PartLLM: A Unified Multimodal Foundation for 3D Part Segmentation

PartLLM:面向三维部件分割的统一多模态基础模型
Zhu, Zhe, Zhang, Yiheng, Li, Peng, Zhao, Zixing, Chen, Honghua, Zhang, Yaqing, Wan, Le, Dou, Zhiyang, Lin, Cheng, Liu, Yuan, Wei, Mingqiang, Wang, Wenping
Abstract
Part segmentation is a fundamental problem in computer graphics and 3D vision. Recent works have expanded 3D part segmentation beyond fixed taxonomies, but existing approaches typically only address a specific setting, such as text-guided part segmentation or point-based interaction. In this work, we argue that these settings can be unified as an intent-conditioned generative problem, where different prompts specify the desired part decomposition. To this end, we introduce PartLLM, a unified multimodal model that formulates 3D part segmentation as autoregressive semantic decomposition. Conditioned on an input shape and a user prompt, PartLLM autoregressively generates semantic part hypotheses as queries for mask prediction and feeds them to a decomposition-aware decoder that jointly predicts coherent part masks. This unified design supports text-guided part segmentation, interactive segmentation, and full-shape semantic decomposition with controllable granularity within a single model. Extensive experiments across these task settings show that PartLLM consistently outperforms task-specific baselines, demonstrating the effectiveness of unifying 3D part segmentation under an intent-conditioned generative formulation.
Chinese Translation
部件分割是计算机图形学与三维视觉中的一个基础问题。近期的研究已将三维部件分割从固定分类体系中扩展出来,但现有方法通常只针对特定设置,例如文本引导的部件分割或基于点的交互。在本工作中,我们提出这些设置可以被统一为一个以意图为条件的生成问题,其中不同的提示词指定了所期望的部件分解方式。为此,我们提出了PartLLM,一个统一的多模态模型,它将三维部件分割形式化为自回归的语义分解。PartLLM以输入形状和用户提示为条件,自回归地生成语义部件假设作为掩码预测的查询(query),并将其输入一个具备分解感知能力的解码器,该解码器联合预测连贯的部件掩码。这一统一设计使得单一模型能够同时支持文本引导的部件分割、交互式分割,以及具有可控粒度的全形状语义分解。在这些任务设置上的大量实验表明,PartLLM持续优于各任务专用基线方法,验证了在以意图为条件的生成式框架下统一三维部件分割的有效性。
cs.CV / 44 / 2609.25837

Identity-Centric Video Summarization via Hierarchical Fusion of Biometric, Appearance, and 3D Body Features

基于生物特征、外观与三维人体特征分层融合的身份中心视频摘要生成
Mirjalili, Milad, Gutiérrez, Enrique Alegre, Fernández, Eduardo Fidalgo, Castro, Víctor González, Rodríguez, Rocío Alaiz, Limas, Manuel Castejón
Abstract
This work presents a video summarization algorithm based on multi-object tracking and person reidentification. We integrate facial embeddings, 3D body-shape features, and visual appearance into a unified tracking framework. These representations enable hierarchical identity assignment and tracking through bidirectional anchoring, which robustly recovers trajectories under severe occlusion or low visual quality. From these stable trajectories, we generate a compact set of summaries for each identity. We select keyframes using a multi-factor weighting scheme that optimizes biometric clarity, social interaction, and motion dynamics, while Adaptive Non-Maximum Suppression ensures temporal diversity. Evaluation on a custom dataset demonstrates tracking stability, achieving an IDF1 of 97.89% and a MOTA of 95.79%. Compared to Top-K selection, our algorithm also increases visual diversity by 146%, temporal coverage by 89%, and information retrievability by 3.5%.
Chinese Translation
本文提出了一种基于多目标跟踪与行人重识别的视频摘要算法。我们将人脸嵌入(facial embeddings)、三维人体体型特征和视觉外观整合到一个统一的跟踪框架中。这些表征通过双向锚定机制实现分层的身份分配与跟踪,能够在严重遮挡或低视觉质量条件下稳健地恢复目标轨迹。基于这些稳定的轨迹,我们为每个身份生成一组紧凑的摘要。我们采用多因子加权方案选择关键帧,该方案综合优化生物特征清晰度、社交互动和运动动态,同时利用自适应非极大值抑制(Adaptive Non-Maximum Suppression)保证时间上的多样性。在自建数据集上的评估表明,该算法具有良好的跟踪稳定性,IDF1达到97.89%,MOTA达到95.79%。与Top-K选择方法相比,我们的算法使视觉多样性提高了146%,时间覆盖率提高了89%,信息可检索性提高了3.5%。
cs.CV / 45 / 2609.25841

Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes

Metric-Bench:探索视觉语言模型在室内场景中基于上下文的空间度量推理能力
Xi, Yuling, Zhang, Haokai, Zhu, Muzhi, Zhong, Hao, Du, Zongze, Zhao, Hengyu, Jing, Chenchen, Yin, Yufei, Qin, Bin, Yang, Yongjie, Luo, Zhenbo, Chen, Hao, Shen, Chunhua
Abstract
Metric reasoning is a critical and challenging task for Vision Language Models (VLMs), playing a pivotal role in embodied AI tasks such as robotic manipulation and autonomous navigation. However, current spatial reasoning remains bottlenecked by rigid pixel-level supervision; such localized optimization often compromises general multimodal intelligence, triggering performance degradation or catastrophic forgetting of broad reasoning capabilities. To address these limitations, we introduce Metric-Bench, a focused benchmark designed to guide metric-spatial reasoning using contextual information. By incorporating in-image reference objects with known physical dimensions, Metric-Bench guides models to implicitly learn the 2D-to-3D mapping without camera intrinsics. We further present MetricReasoner, a task-adapted reinforcement fine-tuning recipe for reference-grounded metric reasoning, using structured prompts and verifiable numerical rewards. Extensive experiments on Metric-Bench demonstrate that our approach significantly enhances spatial metric understanding, outperforming existing and even larger proprietary models by 43.1\%, while improving downstream embodied performance over a spatial-specialized counterpart by 30.4\% on RoboSpatial overall accuracy and 9.3\% on ERQA, and additionally delivering consistent gains on general benchmarks (15.9\% on V$\star$Bench, 88.9\% on BLINK), indicating that the proposed adaptation does not necessarily compromise general VLM capabilities.
Chinese Translation
度量推理是视觉语言模型(VLM)的一项关键且具有挑战性的任务,在机器人操作与自主导航等具身智能任务中发挥着重要作用。然而,当前的空间推理仍受限于僵硬的像素级监督方式;这种局部化优化往往会损害模型的多模态通用智能,导致性能下降或广泛的推理能力出现灾难性遗忘。为解决这些局限,我们提出了Metric-Bench,一个旨在利用上下文信息引导度量空间推理的专项基准。通过在图像中引入具有已知物理尺寸的参考物体,Metric-Bench引导模型在无需相机内参的情况下隐式学习从2D到3D的映射。我们进一步提出MetricReasoner,一种面向任务适配的强化微调方案,用于基于参考物体的度量推理,其采用结构化提示与可验证的数值奖励。在Metric-Bench上的大量实验表明,我们的方法显著提升了空间度量理解能力,超越现有模型甚至更大规模的专有模型达43.1%;同时在下游具身任务中,相比空间专用模型,在RoboSpatial总体准确率上提升30.4%,在ERQA上提升9.3%;此外在通用基准上也取得一致的提升(V$\star$Bench上提升15.9%,BLINK上提升88.9%),表明所提出的适配方法不一定会损害VLM的通用能力。
cs.CV / 46 / 2609.25845

Visual Jev: Accurate and Efficient Decisions from Shared Visual Context

Yu, Guanxu, Yao, Yuhang
Abstract
Many vision applications ask several independent, forced-choice questions about the same image. Visual Jev encodes the image and public context once, executes isolated question suffixes as a batch, and reads candidate probabilities from the backbone's language-model head. Across four benchmarks, answer-supervised post-training raises equal-weight macro accuracy from 70.6% to 76.1%, with the gain concentrated on the two task families represented in training. At N=32 questions per image, shared batched execution is 8.9x faster in warm amortized time than independent serial execution and remains 3.4x faster than an already-batched baseline that recomputes the prefix, at the cost of higher peak memory. A matched typed-head control offers no consistent accuracy advantage over the language-model-head readout. The supported design is therefore simple: adapt the backbone for quality, retain the existing readout, and share execution for efficiency.
cs.CV / 47 / 2609.25850

Less Is More in the Long Tail: Stage-Adaptive Sample Selection for Annotation-Efficient Dense Prediction

长尾场景下少即是多:面向标注高效密集预测的阶段自适应样本选择
Du, Xiaofei, Zhang, Lei, Yan, Shuyu, Wang, Manning, Song, Zhijian
Abstract
Deep learning performance generally improves with increasing training data, yet this scaling is fundamentally constrained by annotation cost in large-scale dense prediction tasks with long-tailed category distributions, where pixel- or voxel-level annotation is prohibitively expensive. We propose SASS (Stage-Adaptive Sample Selection), a stage-adaptive data-selection framework for pool-based active learning in long-tailed dense prediction. SASS combines three components: label-free self-supervised gradient scoring, prior-guided category rebalancing with validation-driven feedback, and stage-adaptive acquisition aligned with model training dynamics. This design avoids candidate ground-truth masks during gradient scoring while making acquisition responsive to long-tail imbalance and evolving representations. We evaluate SASS on a multimodal 3D medical segmentation testbed comprising over 100,000 samples spanning 108 anatomical structures. SASS recovers 98.3% of full-dataset performance with a 40% training-pool annotation budget, outperforming BADGE by 5.1 percentage points. Moreover, SASS exhibits a statistically supported less-is-more pattern, surpassing full-dataset training at the Hard-group level and, at the structure level, for the pancreas and gallbladder. More broadly, SASS shows that annotation-efficient learning depends not only on which samples are selected, but also on how the annotation budget is distributed across categories and when model-derived scores begin to guide selection.
Chinese Translation
深度学习的性能通常随着训练数据的增加而提升,然而在具有长尾类别分布的大规模密集预测任务中,由于像素级或体素级标注的成本极其高昂,这种扩展从根本上受到标注成本的制约。我们提出了SASS(Stage-Adaptive Sample Selection,阶段自适应样本选择),一个面向基于样本池主动学习的长尾密集预测数据选择框架。SASS结合了三个组件:无标签的自监督梯度评分、结合验证驱动反馈的先验引导类别再平衡,以及与模型训练动态对齐的阶段自适应样本获取。该设计在梯度评分过程中避免使用候选样本的真实掩码(ground-truth masks),同时使样本获取能够响应长尾不平衡和不断演化的表征。我们在一个多模态3D医学分割测试平台上评估了SASS,该平台包含超过100,000个样本、涵盖108个解剖结构。SASS在仅40%训练池标注预算的条件下恢复了大全量数据集98.3%的性能,比BADGE高出5.1个百分点。此外,SASS呈现出具有统计支持 的"少即是多"现象:在困难组(Hard-group)层面超越了大全量数据集训练,并在结构层面对胰腺和胆囊也取得了优于全量训练的结果。更广泛地说,SASS表明标注高效的学习不仅取决于选择哪些样本,还取决于标注预算如何在类别间分配,以及模型导出的评分何时开始引导选择。
cs.CV / 48 / 2609.25860

MatchFusion: Explicit-Implicit Instance Matching for Spatio-Temporal Multimodal Autonomous Driving

Li, Xiaoyu, Fu, Jiajia, Shi, Long, Du, Tianyu, Li, Ruihang, Wu, Xian, Zhao, Lijun, Zhang, Yingtao, Sun, Lining, Li, Ruifeng
Abstract
Sparse instance representations provide a compact interface for spatial LiDAR-camera and temporal past-current interaction in multimodal perception and E2EAD. Effective interaction requires reliable instance correspondences despite geometric discrepancies and heterogeneous semantic representations. Attention-based methods exploit contextual semantics but often require specialized representation alignment, increasing computational overhead. In contrast, association based on structured object states is efficient and interpretable but lacks contextual evidence to resolve ambiguous matches. To combine these complementary strengths, we propose MatchFusion, a learnable instance matching and fusion module for spatio-temporal multimodal autonomous driving. MatchFusion initializes pairwise affinities using geometric similarity and category consistency, then selectively refines structurally plausible associations using instance embeddings. The resulting soft matchmap guides a common residual aggregation operator for adaptive information exchange. This unified matching-fusion formulation supports spatial LiDAR-camera and temporal past-current interaction, using multi-view image-plane geometry and motion-compensated BEV geometry as the respective structural priors. Experiments on nuScenes demonstrate consistent perception gains across diverse front-end configurations. Compared with a prior instance-centric fusion method, the MatchFusion-equipped system achieves higher perception accuracy while reducing FLOPs by 55.3% and GPU memory usage by 39.3%, with the matching-fusion module accounting for only 3.7% of total perception latency. Integrating temporal MatchFusion into SparseDrive further improves perception within an E2E framework without additional supervision. These results establish explicit-implicit matching as an effective and efficient mechanism for spatio-temporal instance interaction.
cs.CV / 49 / 2609.25881

Delving into Asymmetric Information Dynamics for High-Fidelity Virtual Try-On

面向高保真虚拟试穿的对称性信息动力学研究
Qin, Zishu, Jin, Zhiyu, Huang, Pipei, Zhou, Hao
Abstract
Virtual try-on (VTON) requires precise pixel-level fidelity, yet mainstream Diffusion Transformers (DiTs) often suffer from texture degradation and structural drift. We identify symmetric interactions in standard joint-attention mechanisms as a source of these failures. Although such interactions support semantic flexibility in general-purpose editing, they allow stochastic noise to corrupt deterministic garment features in VTON. We analyze this problem through asymmetric information dynamics and introduce two diagnostic indicators: Conditional Attention Entropy (CAE) for feature unbiasedness and Injected Information Flux (IIF) for injection effectiveness. Our analysis suggests that symmetric bidirectional attention can corrupt conditional features and attenuate the conditional signal. To address these limitations, we propose RealFit, a framework that combines Unidirectional Information Flow (UIF) with Decoupled Timestep Modulation (DTM). UIF isolates the garment condition from stochastic noise to preserve garment identity, while DTM optimizes the modulation scale to maintain a strong conditional signal. The resulting time-invariant condition branch enables a conditional KV cache that reduces inference time by approximately 75%. RealFit offers a principled approach to conditional generation with state-of-the-art fidelity and efficiency.
Chinese Translation
虚拟试穿(VTON)需要精确的像素级保真度,然而主流的扩散Transformer(Diffusion Transformers, DiTs)常常面临纹理退化和结构漂移的问题。我们发现标准联合注意力机制中的对称性交互是导致这些失败的根源。尽管此类交互在通用图像编辑中支持语义灵活性,但在虚拟试穿任务中,它们使得随机噪声能够破坏确定性的服装特征。我们通过对称性信息动力学来分析这一问题,并引入两个诊断指标:用于衡量特征无偏性的条件注意力熵(Conditional Attention Entropy, CAE)和用于衡量注入有效性的注入信息通量(Injected Information Flux, IIF)。我们的分析表明,对称的双向注意力会破坏条件特征并削弱条件信号。为解决这些局限,我们提出了RealFit框架,该框架将单向信息流(Unidirectional Information Flow, UIF)与解耦时间步调制(Decoupled Timestep Modulation, DTM)相结合。UIF将服装条件与随机噪声隔离以保持服装的身份特征,而DTM则优化调制尺度以维持较强的条件信号。由此得到的时不变条件分支使得条件KV缓存成为可能,从而将推理时间缩短约75%。RealFit为条件生成提供了一种有原则的方法,实现了当前最先进的保真度和效率。
cs.CV / 50 / 2609.25891

BAS-OPD: Budget-Aware Selective On-Policy Self-Distillation for Fine-Grained Multimodal Perception

BAS-OPD:面向细粒度多模态感知的预算感知选择性在策略自蒸馏方法
Chen, Zihan, Zhou, Hengguang, Kang, Yuan, Zhang, Yiming, Fang, Wenhui, Ding, Zenghui, Sun, Yining, Hsieh, Cho-Jui
Abstract
Multimodal large language models (MLLMs) often struggle with fine-grained visual perception when processing complete images, as critical evidence may only appear in local regions. On-policy self-distillation (OPD) enables transferring privileged visual knowledge from informative views to full-image policies, but querying the teacher for every rollout introduces substantial supervision costs. In this work, we propose BAS-OPD, a budget-aware selective OPD framework that allocates teacher supervision under limited query budgets. Instead of querying all rollouts, BAS-OPD selects informative samples while maintaining full-batch student generation. We explore random, uncertainty-based, and learned utility-based selection strategies, where the learned selector estimates query value from detached rollout statistics and online utility signals derived from student--teacher agreement and teacher confidence without additional student forward passes. BAS-OPD only changes training-time supervision allocation and preserves single-pass full-image inference. Experiments on fine-grained multimodal perception benchmarks demonstrate that BAS-OPD achieves strong performance while substantially reducing teacher supervision costs, highlighting the effectiveness of selective OPD under constrained budgets.
Chinese Translation
多模态大语言模型(MLLM)在处理完整图像时往往难以实现细粒度的视觉感知,因为关键证据可能仅出现在局部区域。在策略自蒸馏(On-Policy Self-Distillation, OPD)能够将来自信息丰富视角的特权视觉知识迁移到全图策略中,但对每一次采样(rollout)都查询教师模型会带来巨大的监督成本。本文提出 BAS-OPD,一种预算感知的选择性 OPD 框架,可在有限的查询预算下分配教师监督。BAS-OPD 无需查询所有采样结果,而是在保持全批次学生生成的同时,选择信息丰富的样本。我们探索了随机、基于不确定性以及基于学习的效用选择策略,其中学习型选择器从分离的(detached)采样统计量以及由学生—教师一致性和教师置信度推导出的在线效用信号中估计查询价值,而无需额外的学生前向传播。BAS-OPD 仅改变训练时的监督分配,并保留单次全图推理。在细粒度多模态感知基准上的实验表明,BAS-OPD 在大幅降低教师监督成本的同时取得了优异性能,凸显了在受限预算下选择性 OPD 的有效性。
cs.CV / 51 / 2609.25907

NaCR: Visual Localization via NeRF-aided Camera Ray Regression

NaCR:基于NeRF辅助相机光线回归的视觉定位
Zhang, Yesheng, Dai, Xiang, Zhao, Xu, Zhang, Chongyang
Abstract
Visual localization (VL) is a fundamental technology for vision applications such as virtual reality. Recently, a novel VL paradigm, Camera Ray Regression (CRR), has emerged, which maps 2D image patches to 3D camera rays, but its accuracy is limited. To improve CRR accuracy, we notice a compelling duality: the inverse of this mapping is inherently performed by the novel view synthesis model, \ie, Neural Radiance Fields (NeRF). While NeRF renders image patches from camera rays via differentiable ray marching, CRR predicts the rays from image patches. Motivated by this complementary relationship, we propose NeRF-aided Camera Ray Regression (NaCR), a unified framework that seamlessly bridges NeRF and CRR at the ray level. First, NaCR incorporates three simple yet effective enhancements into the CRR baseline. Second, leveraging a pre-trained NeRF, NaCR augments the training data by synthesizing novel views tailored for efficient, patch-level consumption. Finally, exploiting the differentiability of NeRF, NaCR forms a closed-loop supervision pipeline where photometric rendering errors are back-propagated to optimize the predicted camera rays. To ensure stable convergence within the highly non-convex image space, we introduce a two-stage training curriculum. Extensive experiments across indoor and outdoor benchmarks demonstrate that NaCR achieves competitive accuracy. Comprehensive ablation studies validate the efficacy of each proposed component.
Chinese Translation
视觉定位(Visual Localization, VL)是虚拟现实等视觉应用的基础技术。近来,一种新的视觉定位范式——相机光线回归(Camera Ray Regression, CRR)应运而生,该方法将2D图像块映射到3D相机光线,但其精度有限。为提升CRR的精度,我们注意到一种引人注目的对偶性:该映射的逆过程本质上正是由新视角合成模型(即神经辐射场,Neural Radiance Fields, NeRF)完成的。NeRF通过可微的光线步进从相机光线渲染出图像块,而CRR则从图像块预测光线。受这种互补关系的启发,我们提出了NeRF辅助相机光线回归(NeRF-aided Camera Ray Regression, NaCR),这是一个在光线层面无缝连接NeRF与CRR的统一框架。首先,NaCR在CRR基线中引入了三项简单而有效的改进。其次,NaCR利用预训练的NeRF,通过合成专为实现高效图像块级处理而定制的新视角来增广训练数据。最后,利用NeRF的可微性,NaCR构建了一个闭环监督流程,将光度渲染误差反向传播以优化预测的相机光线。为确保在高度非凸的图像空间中稳定收敛,我们引入了两阶段训练课程。在室内外基准数据集上的大量实验表明,NaCR取得了具有竞争力的精度。全面的消融实验验证了所提出的每个组件的有效性。
cs.CV / 52 / 2609.25930

AT3D-AD: Anomaly Type-Aware 3D Anomaly Detection via Hierarchical Point-Language Alignment

Zeng, Jingyu, Lu, Haoquan, Gao, Can
Abstract
Detecting and localizing 3D point-cloud defects is essential for industrial inspection. However, existing methods often suffer from imprecise localization due to the lack of anomaly supervision and reliance on single-granularity representations. To address these limitations, we propose Anomaly Type-Aware 3D Anomaly Detection (AT3D-AD), a unified framework for joint detection, localization, and classification. Specifically, we first design the Physics-Driven Parametric Anomaly Synthesis (PDPAS) module employing multiple parametric functions to generate synthetic anomalies, providing explicit anomaly supervision. Then, we propose the Hierarchical Global-Local Anomaly Alignment (HiGLA) module to align global and local representations within the normal and anomalous groups. Finally, we propose the Semantic-Geometric Anomaly Classification (SGAC) module to jointly learn localization and classification, yielding spatially precise and type-discriminative anomaly representations. Extensive experiments establish new state-of-the-art performance on all four benchmarks. AT3D-AD achieves Object/Point AUROC scores of 98.1\%/98.9\% on Anomaly-ShapeNet and 95.0\%/95.2\% on Real3D-AD, while reaching 74.2\% Macro-F1 for anomaly-type recognition on Real3D-AD.
cs.CV / 53 / 2609.25937

Calibrating Retrieval Geometry: Reliability-Guided Training-Free Aggregation for Visual Place Recognition

校准检索几何:面向视觉位置识别的可靠性引导免训练聚合方法
Li, Xin, Mao, Zhimin, Wang, Shang, Duan, Siyuan, Zhang, Geng
Abstract
Frozen visual foundation models provide transferable features for visual place recognition, but fixed aggregation can suppress useful distinctions in new environments. We introduce TFA, a reliability-guided, training-free aggregation method requiring neither place labels nor task-specific weight updates. Our key observation is that reproducible retrieval need not be discriminative: independent codebooks can consistently retrieve a few database hubs. TFA combines cross-codebook agreement, retrieval coverage, and spectral statistics to control residual assignment, spectral shaping, and global-feature fusion. Its spectral kernel exactly recovers original descriptor similarity at zero intervention. Database-only TFA fixes its rules before accessing queries; TFA-C64 uses 64 disjoint unlabeled target images to calibrate retrieval for subsequent queries. Across 20 ground protocols with a fixed DINOv2-B backbone and matched resolution, database-only TFA improves Recall@1 over AnyLoc by 17.39 percentage points on MSLS-val and 9.55 on SPED. C64 mitigates failures of database-only calibration in driving environments. Across eight aerial/cross-view protocols, TFA achieves the highest Recall@1 among compared training-free heads in 14 of 16 DINOv2/DINOv3 backbone-protocol combinations. In a separate native-system comparison, DINOv2-G-based TFA-C64 reaches 91.46% Recall@1 on Pitts30k and 76.29% on VPAIR, outperforming the displayed training-free comparators on all five benchmarks. These results show that reliability-guided aggregation can recover additional retrieval capability from frozen representations, providing a practical baseline for new environments with scarce place supervision.
Chinese Translation
冻结的视觉基础模型为视觉位置识别(Visual Place Recognition, VPR)提供了可迁移的特征,但固定的聚合方式可能抑制模型在新环境中有用的区分性信息。我们提出了TFA(Training-Free Aggregation),一种可靠性引导的免训练聚合方法,既不需要位置标签,也不需要任务特定的权重更新。我们的关键观察是:可复现的检索未必具有判别性——独立的码本(codebook)可能会一致地检索出少数数据库枢纽点。TFA结合跨码本一致性、检索覆盖率和谱统计量,来控制残差分配、谱整形以及全局特征融合。其谱核在零干预的情况下能够精确恢复原始描述子的相似度。仅使用数据库的TFA在访问查询之前就确定了其规则;TFA-C64则使用64张互不重叠的无标签目标图像为后续查询校准检索。在固定DINOv2-B骨干网络和相同分辨率的20个地面基准协议上,仅数据库的TFA在MSLS-val上将Recall@1较AnyLoc提升了17.39个百分点,在SPED上提升了9.55个百分点。C64缓解了仅数据库校准在驾驶环境中的失效问题。在八个航拍/跨视角协议上,TFA在16个DINOv2/DINOv3骨干网络-协议组合中的14个里,取得了各对比免训练头(head)中最高的Recall@1。在另一项独立的原生系统对比中,基于DINOv2-G的TFA-C64在Pitts30k上达到91.46%的Recall@1,在VPAIR上达到76.29%,在全部五个基准上均优于所展示的免训练对比方法。这些结果表明,可靠性引导的聚合能够从冻结表征中挖掘出额外的检索能力,为位置监督稀缺的新环境提供了一个实用的基线方法。
cs.CV / 54 / 2609.25945

Towards Systematic Qualification of Vision-Language Models for Automotive Perception Systems

Dona, Malsha Ashani Mahawatta, Rokanas, Konstantinos, Säfström, Alexander, Ronanki, Krishna, Berger, Christian
Abstract
The field of Artificial Intelligence has been adopted for many application domains. Vision Language Models are one of the recently advanced AI techniques that have been explored to support automotive features such as vehicle perception, and safety assurance. However, such language models are prone to hallucinations, posing a potential threat to the safety of automotive systems that may incorporate them. Within the automotive domain, VLMs could not only hallucinate traffic objects, but could also fail to identify traffic objects that are actually present, which may potentially lead to dangerous situations. Though we have observed a growing body of literature that proposes verification and validation techniques for safe and trustworthy AI, these methods are often studied in isolation, focusing either on run-time or design-time phases. Such isolated techniques could be insufficient in safety-critical, realistic contexts such as automotive perception systems. In this paper, we analyze design-time and run-time verification and validation techniques based on a taxonomy presented by Huang et al. We present an automotive study in which a design-time qualification workflow is proposed to complement run-time monitoring. This workflow combines a fixed safety-relevant ontology-based structured annotation system together with a synonym-based evaluation process to statistically evaluate three state-of-the-art VLMs against data from the nuScenes dataset. We observed that the proposed technique enables deterministic and repeatable quantification of the hallucinations VLMs generate in automotive perception-related tasks. The proposed workflow supports model comparison and deployment-oriented engineering decisions within the design-time verification and validation process and will contribute to a holistic verification strategy that strives towards trustworthy automotive perception systems
cs.CV / 55 / 2609.25966

GRIP: Gaussian Rendering as a Cross-Modal Bridge for Image-to-Point Cloud Registration

GRIP:高斯渲染作为图像到点云配准的跨模态桥梁
Slimani, Karim, Achard, Catherine, Marchand, Eric, Tamadazte, Brahim
Abstract
This paper introduces GRIP, a pose-conditioned refinement framework for pixel-to-point matching and 2D to 3D registration. Given an initial coarse pose estimate, GRIP addresses the structural mismatch between grid based image descriptors and unordered point cloud descriptors by softly rendering learned 3D point features onto the image grid through Gaussian feature splatting. The rendered point derived feature map is then fused with image features by a pixel aligned transformer, enabling visual semantic and geometric cues to interact in a shared 2D representation. The refined features are decoded and propagated to finer resolutions for dense correspondence estimation and final pose refinement. Experiments on RGB D Scenes V2 and 7 Scenes demonstrate state of the art inlier ratio and competitive registration recall, with stronger performance under stricter evaluation thresholds.
Chinese Translation
本文提出了GRIP,一个用于像素到点匹配以及2D到3D配准的姿态条件化精化框架。给定一个初始粗略姿态估计,GRIP通过高斯特征溅射(Gaussian feature splatting)将学习到的3D点特征柔性地渲染到图像网格上,从而解决基于网格的图像描述子与无序点云描述子之间的结构不匹配问题。随后,渲染得到的点派生特征图通过像素对齐Transformer与图像特征进行融合,使视觉语义线索与几何线索能够在共享的2D表示中相互作用。精化后的特征被解码并传播至更细的分辨率,用于稠密对应关系估计和最终姿态精化。在RGB-D Scenes V2和7 Scenes数据集上的实验表明,该方法在内点率(inlier ratio)上达到了最先进水平,配准召回率具有竞争力,且在更严格的评估阈值下表现更为出色。
cs.CV / 56 / 2609.25972

NAWE: Digital Watermarking with Neural-Assisted Watermark Extraction

NAWE:基于神经辅助水印提取的数字水印技术
Chaban, Roman, Kinakh, Vitaliy, Rouzaire, Lilian, Voloshynovskiy, Slava
Abstract
NAWE (Neural-Assisted Watermark Extraction) combines an explicit signal-processing watermarking construction with a pretrained neural host predictor. A periodic, perceptually masked watermark carrier provides synchronization, Polar coding supplies redundancy, and denoising followed by subtraction extracts the embedded watermark. The denoiser remains frozen, without watermark-specific training. A one-factor-at-a-time study compares Wiener, BM3D, DRUNet, and GS-DRUNet host estimators. Comparisons with TrustMark, SSL Watermarking, PixelSeal, and WAM show NAWE's lowest geometric and photometric class BER and strong message recovery, while filtering and noise remain limitations consistent with the non-adaptive selection of the watermark extractor. The comparison retains the systems' different payloads and coding.
Chinese Translation
NAWE(神经辅助水印提取,Neural-Assisted Watermark Extraction)将显式的信号处理水印构造与预训练的神经宿主预测器相结合。周期性且经感知掩蔽的水印载波提供同步,Polar编码提供冗余,去噪加相减操作用于提取嵌入的水印。去噪器保持冻结状态,无需针对水印进行专门训练。一项单因素变化实验比较了Wiener、BM3D、DRUNet和GS-DRUNet四种宿主估计器。与TrustMark、SSL Watermarking、PixelSeal和WAM的比较表明,NAWE在几何类和光度类攻击下具有最低的误码率(BER)和较强的消息恢复能力,而在滤波和噪声方面仍存在局限,这与水印提取器的非自适应选择相一致。该比较保留了各系统不同的载荷容量和编码方式。
cs.CV / 57 / 2609.25978

Faithful Faithfulness Evaluations: Challenges & Pitfalls Learned from a Breast MRI Case Study

Poolpol, Peachapong, Detjen, Henrik H. J., Petersen, Eike
Abstract
Saliency maps are widely used to explain deep learning predictions in medical imaging, yet visually plausible explanations do not necessarily reflect a model's true decision process and may therefore mislead clinicians. We investigate this problem using a Vision Transformer-based breast MRI classifier trained on the ODELIA Breast MRI Challenge dataset and evaluate multiple saliency methods, including Last-layer Attention, Attention Rollout, Grad-SAM, Gradient Attention Rollout, GMAR, Grad-CAM, and HiResCAM. Our study highlights two often-overlooked challenges in perturbation-based faithfulness evaluation. First, method rankings depend strongly on the perturbation strategy, varying across intensity-based perturbations and transformer-based attention masking. Second, benchmarking saliency methods requires distinguishing between class-specific and class-agnostic explanations. To enable fair comparisons, we introduce non-class-specific variants of gradient-based methods and evaluate both settings separately. Across protocols, Grad-CAM and Gradient Attention Rollout consistently emerged as the strongest class-specific methods, although their relative ranking depended on the evaluation design. These findings expose important limitations of current saliency-based explainability approaches and highlight the need for more robust and standardized evaluation frameworks for trustworthy clinical AI systems.
cs.CV / 58 / 2609.26039

EMERGE: Resolution-Agnostic Point Cloud Generation with Equivariant Graph-Based Diffusion

EMERGE:基于等变图扩散的分辨率无关点云生成
Mitsouras, Ilias, Chaidos, Nikolaos, Stamou, Giorgos, Voulodimos, Athanasios
Abstract
Point cloud generation has emerged as a crucial task for accurately capturing and reproducing the complexity of the physical world. However, existing generative approaches, predominantly relying on Transformers and Variational Autoencoders (VAEs), frequently ignore the continuous, non-grid topologies inherent to 3D spaces. Although the integration of graph-based structures has yielded significant benefits in related discriminative vision tasks, such geometric architectures remain noticeably absent from 3D generative modeling. To address this gap, we introduce EMERGE (Equivariant Multi-scale GNN for Resolution-agnostic point cloud GEneration), the first fully $SE(3)$-equivariant graph-based diffusion backbone explicitly designed to generate point clouds while preserving continuous spatial symmetries. Our framework bypasses the rigid resolution dependencies of standard generative pipelines, enabling zero-shot inference at multiple, arbitrary spatial resolutions. Extensive empirical evaluations demonstrate that EMERGE achieves State-of-the-Art generation quality across standard metrics, while the strong inherent geometric inductive biases enable significantly faster training convergence compared to existing baseline methods.
Chinese Translation
点云生成已成为精确捕捉和再现物理世界复杂性的一项关键任务。然而,现有的生成方法主要依赖Transformer和变分自编码器(VAE),往往忽略了3D空间固有的连续、非网格拓扑结构。尽管基于图的结构在相关的判别式视觉任务中已带来显著收益,但此类几何架构在3D生成建模中仍明显缺失。为填补这一空白,我们提出EMERGE(Equivariant Multi-scale GNN for Resolution-agnostic point cloud GEneration,用于分辨率无关点云生成的等变多尺度图神经网络),这是首个完全 $SE(3)$ 等变的基于图的扩散主干网络,专为生成点云并保持连续空间对称性而设计。我们的框架摆脱了标准生成流程对固定分辨率的依赖,能够在多个任意空间分辨率下实现零样本推理。大量实证评估表明,EMERGE在标准指标上达到了最先进的生成质量,且其强大的内在几何归纳偏置使其相比现有基线方法能够显著加快训练收敛速度。
cs.CV / 59 / 2609.26056

CricRAG: Retrieval Augmented Vision-Language Models for Personalized Cricket Coaching

Singh, Agamdeep, PB, Sujit, Vatsa, Mayank
Abstract
Vision-Language Models (VLMs) offer promising capabilities for automated sports coaching but face a fundamental limitation: they implicitly compare against professional standards, making their feedback impractical for developing players. We present CricRAG, a retrieval-augmented framework that aligns VLMs with skill-appropriate benchmarks for personalized cricket coaching. Our key insight is that by retrieving similar-but-better techniques as reference points, we can guide VLMs to provide developmentally appropriate feedback that mirrors human coaching practices. We contribute: (1) a labelled dataset of 288 cricket technique videos spanning multiple skill levels, (2) an efficient motion retrieval pipeline using contrastive learning that achieves 78% top-3 retrieval accuracy, (3) a frame sampling technique that reduces inference costs, and (4) a retrieval-augmented approach that significantly improves feedback alignment with coaching principles, achieving up to 94% agreement with professional assessments compared to 67% without retrieval context.
cs.CV / 60 / 2609.26064

SPEANet: Structural Prior Enhanced Attention Network for Parameter-Efficient Remote Sensing Object Detection

Lu, Wei, Li, Junjie, Sang, Feifei, Chen, Si-Bao
Abstract
Remote sensing object detection (RSOD) requires compact backbones capable of preserving weak geometric cues under extreme scale variation and background clutter. Fixed structural operators provide complementary contour and frequency responses without introducing learnable operator coefficients. However, directly injecting these responses can amplify content-irrelevant textures, while applying a uniform operator design across the hierarchy may be poorly matched to stage-specific representation requirements. We propose the Structural Prior Enhanced Attention Network (SPEANet), a parameter-efficient RSOD backbone that integrates fixed operators through stage-specific prior extraction and context-conditioned response modulation. SPEANet assigns smoothed contour and multi-order directional modeling to shallow, high-resolution features, while employing a compact approximation-detail interaction mechanism in deeper stages. Learned spatial gates regulate the resulting prior responses before residual fusion. Experiments on five benchmarks, together with evaluations across seven detection frameworks on DOTA-v1.0, achieve a favorable accuracy-parameter trade-off. With Oriented R-CNN, SPEANet achieves 78.55\% mAP on DOTA-v1.0, 72.24\% mAP on DOTA-v1.5, and 67.30\% mAP on DIOR-R using 23.0M total parameters, including a 5.97M-parameter backbone.
cs.CV / 61 / 2609.26073

Cellular-Communication-Level Interpretability for Pathology Foundation Models via Graph Distillation on Microenvironment

基于微环境图蒸馏的病理学基础模型细胞通信级可解释性研究
Xiao, Yuxiang, Chen, Zhiwei, Dai, Dan, Li, Wei, Zhang, Tianyang, Ju, Yakun, Hu, Yang, Yang, Kaixiang
Abstract
Pathology foundation models (PFMs) provide strong tile-level representations but remain difficult to interpret at the cellular and microenvironmental scales that underpin clinical reasoning. We introduce Graph-Interpreter (G-Interp), a graph-distillation framework that equips a frozen PFM teacher with a cellular-communication-level "plug-in" interpreter, without modifying the teacher. For each tile, we segment cells as graph nodes and construct a microenvironment graph based on spatial adjacency. Graph neural network (GNN) students distil the PFM embedding, whilst learning attention-based message passing that yields node- and edge-level importances. We interpret these importances as cell-cell communication evidence, providing fine-grained explanations of how PFMs encode microenvironmental context. To stabilise distillation when graph abstraction is imperfect, we employ a lightweight auxiliary student to supply complementary visual cues and condition graph message passing, while keeping the primary interpretability signal graph-derived. We evaluate explanation faithfulness by mapping graph-selected evidence back to the image using instance masks and measuring teacher sensitivity under targeted vs non-target occlusions. Across multiple histopathology tasks, G-Interp produces highly scalable, microenvironment-aware explanations, while maintaining competitive predictive performance.
Chinese Translation
病理学基础模型(Pathology Foundation Models, PFMs)能够提供强大的切片级表示,但在支撑临床推理的细胞和微环境尺度上仍然难以解释。我们提出了Graph-Interpreter(G-Interp),一个图蒸馏框架,它为冻结的PFM教师模型配备了一个细胞通信级的"即插即用"解释器,而无需修改教师模型。对于每个切片,我们将细胞分割为图的节点,并基于空间邻接关系构建微环境图。图神经网络(GNN)学生模型对PFM嵌入进行蒸馏,同时学习基于注意力的消息传递机制,从而产生节点级和边级的重要性度量。我们将这些重要性度量解释为细胞间通信的证据,为PFM如何编码微环境上下文提供细粒度的解释。为了在图抽象不完美时稳定蒸馏过程,我们采用一个轻量级的辅助学生模型来提供补充性的视觉线索并对图消息传递进行条件化,同时保持主要的可解释性信号源自图结构。我们通过利用实例掩码将图选择的证据映射回图像,并测量教师模型在靶向遮挡与非靶向遮挡下的敏感性,来评估解释的忠实性。在多个组织病理学任务上,G-Interp生成了高度可扩展的、微环境感知的解释,同时保持了具有竞争力的预测性能。
cs.CV / 62 / 2609.26078

ToW3D: Consistency-aware Interactive Point-based Mesh Editing on GANs

ToW3D:基于一致性感知的交互式点驱动GAN网格编辑
Song, Haixu, Liu, Fangfu, Zhang, Chenyu, Duan, Yueqi
Abstract
In this paper, we propose ToW3D that enables precise and consistent control over 3D generative adversarial networks (GANs) with the Tug-of-War competition between shape deformation and appearance consistency. Existing point-based GAN editing methods such as DragGAN and GANWarping have yielded impressive performance for 2D image manipulation. However, as 3D generators present weaker generalization ability compared with 2D due to limited training data, they would suffer from drastic changes in global appearance when editing local areas of meshes. To address this, we design a pipeline of ``drag locally, shove globally'', which iteratively performs two optimization steps: 1) pull the point towards the target, and 2) push the structure and semantics back to the source. Specifically, we design a structure adaption module based on structure which guarantees the preservation of basic geometric properties, and a semantic preservation module that maintains semantic similarity across different views. Extensive qualitative and quantitative experiments demonstrate superiority of our ToW3D approach over prior methods in terms of appearance consistency and fidelity especially under large deformations.
Chinese Translation
在本文中,我们提出了ToW3D,通过形状变形与外观一致性之间的“拔河”竞争机制,实现对3D生成对抗网络(GAN)精确且一致的控制。现有的基于点的GAN编辑方法(如DragGAN和GANWarping)在2D图像操作方面取得了令人瞩目的性能。然而,由于训练数据有限,3D生成器相比2D具有更弱的泛化能力,在编辑网格局部区域时会导致全局外观发生剧烈变化。为解决这一问题,我们设计了“局部拖拽,全局回推”的处理流程,该流程迭代地执行两个优化步骤:1)将点拉向目标位置;2)将结构和语义推回源状态。具体而言,我们设计了一个基于结构的自适应模块,以保证基本几何属性的保持;同时设计了一个语义保持模块,以维持不同视角间的语义相似性。大量定性和定量实验表明,我们的ToW3D方法在外观一致性和保真度方面,尤其是在大变形情况下,优于已有方法。
cs.CV / 63 / 2609.26088

BDSLI: A hybrid CNN-Transformer model for Bengali Sign Language interpretation

Yousuf, Abir Bin, Hossain, Muhammad Iqbal
Abstract
This study introduces a novel hybrid CNN-Transformer architecture to address the limited progress in Bengali SLR, focusing on isolated sign word recognition and sentence generation. This specific model combination is new to Bengali SLR tasks. A custom video dataset was developed, featuring 62 distinct Bengali sign words (250 samples/class), along with a separate test dataset. The CNN-Transformer model demonstrated superior performance against all comparative and baseline models (e.g., CNN-LSTM, standalone TCN), achieving a 99.58% training accuracy (99.48% validation) and a 98.65% test accuracy. The trained model was subsequently deployed in a web application for real-world validation.
cs.CV / 64 / 2609.26092

Match One, Learn with Graph: One-to-Graph Query Collaboration with Backward Sharing for Object Detection

匹配一,图上学习:基于反向共享的一对多图查询协作目标检测方法
Fan, Wenxiao, Fu, Jingling, Liu, Luohang, Ma, Lichen, He, Yu, Yu, Zhiyang, Bi, Weishan, Huang, Junshi, Li, Yan, Simiu, Gu, Li, Kan
Abstract
One-to-one (O2O) matching enables Detection Transformers (DETRs) to perform end-to-end set prediction by assigning each object to a single positive query. However, the strongest classification, center, scale, and overlap evidence for an object is often distributed across multiple queries. This mismatch leaves only the matched owner positively supervised for the object, while other evidence-bearing queries receive no box target for it. We term this query knowledge fragmentation. To exploit such complementary evidence without one-to-many supervision, we propose BS-O2G, a plug-in that builds a sparse prediction-aware graph from decoded features, boxes, and class distributions to organize query collaboration in feature and optimization spaces while preserving the original O2O matcher, positive labels, and objective. One-to-Graph (O2G) calibration propagates relative messages over this graph to consolidate query evidence in the forward pass, whereas Backward Sharing (BS) reuses its transposed detached adjacency to route gradients across persistent query basis vectors without changing the decoder input in the forward pass. Experiments across diverse DETR methods, backbones, COCO, and CrowdHuman show consistent gains and faster convergence with negligible parameter/FLOP growth and modest runtime overhead, supporting graph-based query collaboration as an alternative to expanding positive assignments.
Chinese Translation
一对一(O2O)匹配使检测Transformer(DETR)能够通过将每个目标分配给单个正查询来执行端到端的集合预测。然而,对于一个目标而言,最强的分类、中心、尺度和重叠证据往往分布在多个查询之中。这种不匹配导致只有被匹配的查询(所有者)获得该目标的正监督,而其他携带相关证据的查询则无法获得该目标的边界框监督。我们将这种现象称为查询知识碎片化。为了在不引入一对多监督的情况下利用这种互补证据,我们提出了BS-O2G,一种即插即用模块,它基于解码特征、边界框和类别分布构建一个稀疏的预测感知图,从而在特征空间和优化空间中组织查询协作,同时保留原有的O2O匹配器、正样本标签和目标函数。一对图(O2G)校准在该图上传播相对消息,在前向传播过程中整合查询证据;而反向共享(BS)则复用其转置的、经过梯度截断的邻接矩阵,在持久化的查询基向量之间传递梯度,且不改变前向传播中的解码器输入。在多种DETR方法、骨干网络以及COCO和CrowdHuman数据集上的实验表明,该方法在参数量和FLOP几乎不增加、运行时开销适中的情况下,带来了持续的性能提升和更快的收敛速度,支持将基于图的查询协作作为扩展正样本分配的一种替代方案。
cs.CV / 65 / 2609.26093

RECAP: Relation Evidence Calibration for Detecting Spatial Relation Hallucinations in Vision-Language Models

RECAP:用于检测视觉语言模型中空间关系幻觉的关系证据校准
Liu, Feixiang, Qiu, Qiang, Li, Qingyang, Xu, Hui
Abstract
Vision-language models can answer spatial relation questions confidently even when the image supports an incompatible relation. We formulate relation-grounded selective prediction: accept or reject an already-produced yes/no answer by auditing its visual support, rather than treating uncertainty as evidence. RECAP, our relation-evidence calibration framework, compares image-conditioned likelihoods for a claim, its semantic contradictions, and optional one-sided supports, then converts these witnesses into an answer-conditioned rejection risk. A calibration-only gate preserves confidence as a veto when confidence is demonstrably informative and otherwise deploys relation evidence alone. Across 20 group/image-disjoint splits, RECAP lowers H-FPR@80 over confidence by between 2.0 and 17.9 points on VSR and raises Acc@80 by 3.0, 8.6, and 12.6 points on What'sUp for Qwen3-VL-8B, InternVL3.5-8B, and LLaVA-1.5-7B. It outperforms matched VCD-style visual contrast on all four primary metrics in all six settings. Full-pool VSR fallback, target-ranked GSR-Bench transfer, equal-budget supervised controls, and two additional checkpoints show a consistent operating principle: structured counterevidence complements certainty when confidence is misaligned, while the gate retains confidence when it is already useful.
Chinese Translation
即使图像所支持的是不相容的关系,视觉语言模型仍能自信地回答空间关系问题。我们提出了基于关系的选择性预测任务:通过审核答案是/否(yes/no)回答的视觉依据来接受或拒绝该回答,而不是将模型的不确定性直接视为证据。RECAP是我们的关系证据校准框架,它比较图像条件下某一主张、其语义矛盾项以及可选的单侧支持项的似然,然后将这些证据转化为以答案为条件的拒绝风险。一个仅做校准的门控机制在置信度被证明具有信息量时保留其作为否决权,否则仅使用关系证据。在20个组/图像不相交的划分上,RECAP在VSR数据集上将H-FPR@80较置信度基线降低了2.0至17.9个百分点,并在What'sUp数据集上使Qwen3-VL-8B、InternVL3.5-8B和LLaVA-1.5-7B的Acc@80分别提升了3.0、8.6和12.6个百分点。在全部六种设置下,它在所有四个主要指标上均优于匹配参数预算的VCD式视觉对比方法。全池VSR回退实验、面向目标的GSR-Bench迁移实验、等预算监督对照实验以及两个额外的检查点均表明了一致的运行原则:当置信度出现错位时,结构化反证据可对其进行补充;而当置信度本身已经有用时,门控机制则会保留置信度。
cs.CV / 66 / 2609.26099

Test-time Reinforcement Learning for Anomalous Video Understanding

面向异常视频理解的测试时强化学习
Li, Huining, Duan, Yuxiang, Tan, Jiyang, Li, Qian, Chen, MingCai, Zhang, Jian, Sheng, Xingdong, Du, Yuntao
Abstract
Anomalous video understanding aims to identify abnormal events in videos and interpret their semantic meanings beyond simple anomaly detection. Recent video large language models (Video-LLMs) have demonstrated promising zero-shot capabilities for this task, yet their performance remains limited due to insufficient adaptation to diverse anomaly patterns and evolving environments. Test-time reinforcement learning offers a promising solution by enabling models to improve through self-generated feedback signals without requiring additional human annotations. However, applying it to anomalous video understanding remains challenging due to three issues: (1) generated pseudo-labels can be unreliable when consensus is weak; (2) binary reward designs fail to capture uncertainty in model generations, resulting in ineffective optimization signals; and (3) unanimous rollout groups receive identical rewards, causing group-relative advantages to collapse and eliminating effective policy-gradient signals. To address these challenges, we present a novel test-time reinforcement learning framework for anomalous video understanding by introducing dual-query consistency filtering, an entropy-aware consensus reward, and a virtual negative anchor mechanism. The framework retains reliable samples through consistency across semantically equivalent queries, combines answer agreement with generation uncertainty for reward estimation, and introduces a virtual negative anchor to create reward variation in unanimous rollout groups, thereby preserving effective group-relative optimization signals. Experiments on VAU-Bench show that our method outperforms the compared frozen and supervised baselines. The gains are most pronounced on the ECVA subset of VAU-Bench with thinking, where accuracy improves from 75.81% to 90.00% relative to the frozen backbone.
Chinese Translation
异常视频理解旨在识别视频中的异常事件,并解释其语义含义,而不仅仅是简单的异常检测。近期的视频大语言模型(Video-LLMs)在该任务上展现出了有前景的零样本能力,但由于对多样化异常模式和不断变化的环境适应不足,其性能仍然受限。测试时强化学习提供了一种有前景的解决方案,使模型能够通过自生成的反馈信号进行改进,而无需额外的人工标注。然而,将其应用于异常视频理解仍面临三个挑战:(1)当共识较弱时,生成的伪标签可能不可靠;(2)二元奖励设计无法捕捉模型生成中的不确定性,导致优化信号失效;(3)完全一致的采样组会获得相同的奖励,导致组内相对优势坍塌,从而消除有效的策略梯度信号。为应对这些挑战,我们提出了一种新颖的面向异常视频理解的测试时强化学习框架,引入了双重查询一致性过滤、熵感知共识奖励以及虚拟负锚点机制。该框架通过语义等价查询间的一致性保留可靠样本,将答案一致性与生成不确定性相结合进行奖励估计,并引入虚拟负锚点以在完全一致的采样组中创造奖励差异,从而保留有效的组内相对优化信号。在VAU-Bench上的实验表明,我们的方法优于对比的冻结模型和有监督基线。增益在VAU-Bench带思考过程的ECVA子集上最为显著,相对于冻结的骨干模型,准确率从75.81%提升至90.00%。
cs.CV / 67 / 2609.26103

MIAR: Medical Image Super-Resolution With Autoregressive Modeling

MIAR:基于自回归建模的医学图像超分辨率
Li, Fang, Li, Yinglong, Wu, Hongyu, Gao, Yang, Zhao, Minwei, Hao, Aimin
Abstract
Medical Image Super-Resolution (MISR) aims to enhance spatial resolution without requiring hardware modifications. Although deep learning has yielded promising results, existing paradigms face a critical trade-off: diffusion-based methods suffer from prohibitive inference latency and compromised structural fidelity, whereas regression-based models typically produce over-smoothed results that lack perceptual realism. To address these limitations, we propose MIAR, which reformulates super-resolution as a conditional and progressive next-scale prediction task through a multi-scale autoregressive framework. To ensure structural fidelity, we augment the autoregressive backbone with a Scale-Adaptive Structural Decoder. Furthermore, we integrate a hierarchical beam search strategy during inference to mitigate the recursive error accumulation inherent in autoregressive generation, a phenomenon that is especially pronounced in medical images. Extensive experiments demonstrate that MIAR establishes new state-of-the-art benchmarks while maintaining superior fidelity. Notably, our framework achieves a 7.86% improvement in the perceptual metric MUSIQ compared with the state of the art, while simultaneously delivering a 2.02x speedup over diffusion-based methods.
Chinese Translation
医学图像超分辨率旨在在不修改硬件的前提下提升图像的空间分辨率。尽管深度学习已取得了可喜的成果,但现有范式面临一个关键权衡:基于扩散模型的方法存在推理延迟过高以及结构保真度受损的问题,而基于回归的模型通常会产生过度平滑、缺乏感知真实感的结果。为解决这些局限,我们提出了MIAR,通过多尺度自回归框架将超分辨率重新表述为一个条件化的、渐进式的下一尺度预测任务。为确保结构保真度,我们在自回归主干的基础上增加了尺度自适应结构解码器。此外,我们在推理过程中引入分层束搜索策略,以缓解自回归生成中固有的递归误差累积问题——该现象在医学图像中尤为显著。大量实验表明,MIAR在保持卓越保真度的同时建立了新的最先进基准。值得注意的是,与现有最先进方法相比,我们的框架在感知指标MUSIQ上提升了7.86%,同时相比基于扩散的方法实现了2.02倍的加速。
cs.CV / 68 / 2609.26117

Vorch-Human: Unified Multi-Task Human-Centric Generation via Long-Horizon Continuation

Vorch-Human:基于长时程续写的统一多任务以人为中心生成框架
Ding, Yang, Yu, Haoran, Ma, Xin, Lu, Yulei, Han, Menglin, Wang, Yaole, Yang, Siqian, Yue, Gang, Zhang, Kaihao, Wang, Yaohui, Ma, Lin
Abstract
Human-centric audio-visual generation spans several closely related tasks: animating a person from driving speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references. Existing systems commonly solve these tasks with separate models, even though they share the same target modalities and differ mainly in which observations are provided as conditions. We present Vorch-Human, a unified human-centric generation framework built on a dual-stream audio-video diffusion transformer. Vorch-Human augments the conventional noisy audio/noisy video interface with clean condition-audio and condition-video token groups. Per-token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder allow driving speech, timbre examples, first frames, and subject images to be expressed within one model. To supply the supervision required by this interface, we develop a two-level data pipeline. Level 1 analyzes each clip with speech recognition, vocal separation, face detection and tracking, active-speaker and synchronization models, audio/visual speaker clustering, and multimodal caption correction; it produces subject-indexed speech, appearance, and timbre annotations. Level 2 links the same person across clips from a common source video and mines identity- and outfit-consistent reference images after face, body, quality, pose, and vision-language verification. Finally, we adapt Vorch-Human to long-form audio-driven generation by training with clean latent prefixes and using the same frozen-prefix recurrence at inference. Each segment contributes only its newly generated suffix, reducing boundary discontinuity and long-horizon identity drift. Experiments on short and five-minute generation demonstrate strong identity preservation, audio-visual synchronization, and temporal stability.
Chinese Translation
以人为中心的视听生成涵盖多个紧密相关的任务:根据驱动语音使人物动起来、基于声音参考联合生成语音与视频,以及根据成对的外观与声音参考合成场景。尽管这些任务共享相同的目标模态,仅在作为条件提供的观测内容上有所不同,但现有系统通常使用独立的模型分别解决这些任务。我们提出了 Vorch-Human,一个构建于双流音视频扩散Transformer(dual-stream audio-video diffusion transformer)之上的统一以人为中心生成框架。Vorch-Human 在传统的噪声音频/噪声视频接口基础上,引入了干净的条件音频(condition-audio)和条件视频(condition-video)token 组。通过逐 token 的任务嵌入、时间位置类型、条件掩码以及共享的多模态提示编码器,驱动语音、音色示例、首帧和主体图像可以在同一模型内得到表达。为提供该接口所需的监督信号,我们开发了一个两级数据流水线。第一级通过语音识别、人声分离、人脸检测与跟踪、活跃说话人与同步模型、音频/视觉说话人聚类以及多模态字幕校正来分析每个片段,生成以主体为索引的语音、外观和音色标注。第二级将来自同一源视频的片段中的同一个人关联起来,并在经过人脸、身体、质量、姿态以及视觉语言验证之后,挖掘身份一致和着装一致的参考图像。最后,我们通过使用干净潜变量前缀(clean latent prefixes)进行训练,并在推理时采用相同的冻结前缀递归(frozen-prefix recurrence),将 Vorch-Human 适配到长时程音频驱动生成。每个片段仅贡献其新生成的后缀部分,从而减少了边界不连续性和长时程身份漂移。在短时生成与五分钟生成实验中,Vorch-Human 展现出强大的身份保持能力、音视频同步性和时间稳定性。
cs.CV / 69 / 2609.26161

Moving6DPoSe: A Multimodal Database for Monocular 6D Pose Estimation and Segmentation of Moving Objects

Bugueno-Cordova, Ignacio, Ruiz-del-Solar, Javier, Verschae, Rodrigo
Abstract
Estimating the 6D pose of moving objects remains challenging due to motion blur and the limited temporal resolution of conventional frame-based cameras. Existing event-based datasets further provide limited sensing modalities, annotations, and motion scenarios. We introduce Moving6DPoSe, a multimodal database comprising two complementary subsets: Moving6DPoSe-R with real-world recordings and Moving6DPoSe-S with synthetic sequences generated from the same objects. The dataset contains 16 scanned objects and 1,702 real and synthetic rosbags spanning multiple motion scenarios, with annotations for semantic segmentation, object detection, and monocular 6D pose estimation. We further provide baseline results for all three tasks across frame and event-based modalities. Experimental results show that event-based representations achieve more robust moving-object segmentation than conventional RGB images, while monocular orientation estimation remains challenging, highlighting the potential of Moving6DPoSe for moving-object perception research.
cs.CV / 70 / 2609.26166

MGRL-RSCC: Multi-Granularity Reward Reinforcement Learning for Fine-Grained Remote Sensing Change Captioning

Wang, Futian, Wang, Mengqi, Wang, Xiao, Wu, Wentao, Wang, Haowen, Zhao, Zhicheng, Tang, Jin
Abstract
Remote Sensing Change Captioning (RSCC), which aims to generate accurate and detailed linguistic descriptions of ground object variations from bi-temporal remote sensing images, is a critical and challenging task in intelligent remote sensing interpretation. The mainstream autoregressive training paradigm faces severe exposure bias and train-test distribution mismatch, resulting in cumulative generation errors. They tend to produce conservative and template-fixed captions while ignoring subtle scene change details. To address these challenges, this paper proposes a novel multi-granularity reward reinforcement learning paradigm, termed MGRL-RSCC. Specifically, we first leverage a CNN and hierarchical self-attention module to extract and enhance visual features from bi-temporal remote sensing images. A Transformer decoder is then utilized to complete visual-to-linguistic translation. Different from existing methods, we design a dual-decoding strategy and a two-stage joint optimization scheme, which combines token-level supervised learning via greedy decoding and multi-granularity reward-driven self-critical reinforcement learning via sampling decoding. We further construct three complementary reward functions covering linguistic fluency, change state consistency, and structural-semantic relevance to comprehensively optimize caption quality and alleviate false and missing change descriptions. Extensive experiments on multiple public RSCC benchmark datasets demonstrate that the proposed MGRL-RSCC effectively mitigates exposure bias and conservative generation problems in traditional autoregressive methods. The source code and pre-trained models will be released on https://github.com/Event-AHU/MGRL-RSCC
cs.CV / 71 / 2609.26188

End-to-End Visual Odometry with RNNs and Attention

Li, Ruiyu, Liu, Yinjia, Yu, Alexander
Abstract
Video Odometry (VO) is the process of estimating the ego-motion of an object by analyzing visual information such as a sequence of frames from one or multiple cameras. It has been a popular research topic in computer vision and robotics, and its applications include mobile robotic systems as well as autonomous driving. In this project, we investigate existing end-to-end deep-learning approaches to VO, and propose a novel temporal attention-based model to improve upon the baseline. In addition, while the vast majority of existing deep-learning-based approaches to VO are trained on driving data, we investigate the performance of deep-learning-based VO to the more dynamic and complex problem of hand-held cameras.
cs.CV / 72 / 2609.26189

Topology-Aware Parameter-Efficient Adaptation for Cross-Dataset Retinal Vessel Segmentation

面向跨数据集视网膜血管分割的拓扑感知参数高效自适应方法
Huang, Yongsong, Miyazaki, Tomo, Xu, Kai, Liu, Xiaofeng, Fan, Yaohou, Omachi, Shinichiro
Abstract
Retinal vessel segmentation in multi-domain deployment requires a source model to adapt to domains that differ in imaging conditions and annotation conventions. Conventional parameter-efficient fine-tuning reduces target-specific storage, but its highly restricted adaptation subspace can be insufficient for reconstructing thin, connected vascular structures. We therefore ask how target-specific capacity should be allocated so that topology-aware supervision remains effective under a strict per-domain parameter budget. Based on this principle, we propose TAPDecoderFT, a topology-responsive, role-structured adaptation framework. Specifically, TAPDecoderFT shares a fixed source parameter state across deployment domains, uses low-rank residuals for target-specific private/fusion feature mixing, and retains a trainable dense-reconstruction path comprising the decoder, output head, and refinement module. To promote structurally faithful predictions, the compact target state is jointly optimized with a region-overlap and topology-aware objective that encourages centerline continuity and thin-branch recovery. It improves both DSC and clDice over GenericLoRA-r4 and narrow TAP-r4 in all six directions and is comparable to full fine-tuning.
Chinese Translation
视网膜血管分割在多域部署中要求源模型能够适应成像条件和标注规范各异的域。传统的参数高效微调方法虽能减少目标域专用存储,但其高度受限的自适应子空间可能不足以重建细小且连通的血管结构。因此,我们探讨应如何分配目标域专用容量,才能在严格的每域参数预算下使拓扑感知监督仍然有效。基于这一原则,我们提出了TAPDecoderFT——一种拓扑响应、角色结构化的自适应框架。具体而言,TAPDecoderFT在所有部署域之间共享固定的源参数状态,使用低秩残差实现目标域专有的私有/融合特征混合,并保留一条可训练的密集重建路径,该路径由解码器、输出头和细化模块组成。为促进结构上忠实的预测,紧凑的目标域状态与区域重叠和拓扑感知目标联合优化,该目标鼓励中心线连续性和细分支恢复。在全部六个方向上,该方法在DSC和clDice指标上均优于GenericLoRA-r4和窄参数的TAP-r4,并与全量微调相当。
cs.CV / 73 / 2609.26205

LLaVA-Assessor: Building the Foundation LMM For Visual Quality Assessment

Jia, Ziheng, Zhang, Zicheng, Qian, Jiaying, Zhai, Guangtao, Min, Xiongkuo
Abstract
Aligning with the human visual system~(HVS) in perceiving and evaluating the quality of visual signals is a central objective of machine-vision-based visual quality assessment systems. With the rapid progress of large multi-modal models~(LMMs), visual question answering provides a promising paradigm for building unified foundation models for visual quality assessment under multi-modal and multi-task scenarios. Inspired by the classical ``perception-decision" process in HVS-based quality evaluation, we formulate visual quality assessment for LMM-based machine vision as two complementary tasks: ``quality interpretation'' and ``quality scoring". Centered on these objectives, we propose LLaVA-Assessor, a unified data construction and model training system. To support multi-modal inputs, we design an adaptive model architecture that enables efficient processing of both images and videos. For data construction, we develop rigorous human annotation protocols and a novel machine-synthesis-dominated data expansion pipeline to build a large-scale and high-quality datasets. Furthermore, we introduce a simple yet effective prompt disentanglement strategy to alleviate training-objective confusion in multi-task learning, thereby enabling stable and coherent joint training. The resulting all-in-one LMM LLaVA-Assessor-GIGA achieves superior performance on $11$ image/video quality scoring test sets and 4 visual quality interpretation benchmarks. Extensive results demonstrate the effectiveness of integrating structured data construction, adaptive model design, and multi-task joint training for automated visual quality assessment. Our work provides compelling insights for developing foundation LMMs for automatic visual quality assessment. Project page at https://github.com/jzhws/LLaVA-Assessor.
cs.CV / 74 / 2609.26233

The Temporal Moderation Gap: Text-to-Video Safety Filters Are Blind to Harm in Motion

Cao, Yuxin, Guo, Fusen, Wu, Yuezhong, Mo, Huadong, Song, Wei
Abstract
Text-to-video (T2V) services inherit their safety stack from image generation, pairing a keyword prompt filter with a per-frame checker that blocks a clip whenever one sampled frame looks unsafe. This stack has a blind spot unique to video. We prove that any moderator ignoring frame order accepts a harmful clip whenever it accepts that clip's benign shuffle, so harm carried by the ordering alone escapes. Empirically, the unmodified benchmark prompt already lands a clip in this moderation gap on 32.7% of Sequential-Action targets over four held-out seeds, and paraphrasing, scene splitting, and a feedback-driven prompt search show no significant improvement (paired McNemar $p\ge0.12$), so prompt engineering is not needed to expose the vulnerability. Dense-scoring all 97 rendered frames shows that about a third of the delivered clips merely hide an unsafe frame, while the rest stay harmful as ordered videos even though every frame passes, an order-blind residual the unmodified prompt reaches on a quarter of Sequential-Action targets. We also document a measurement pitfall, since scoring a searched prompt on its own render seed inflates a 7.5% per-generation rate into an apparent 46.7%. A user study confirms that people read these clips as harmful and their shuffles as safe. The fix is to read frame order, and an order-aware detector separates these clips from their own shuffles at AUC 0.74 where per-frame checking sits at chance, which is the signal deployed moderation throws away.
cs.CV / 75 / 2609.26236

COVER: Codec-Robust Video Watermarking with Generative Video Priors

COVER:基于生成式视频先验的编解码器鲁棒视频水印
Cao, Yuxin, Yang, Hao, Ding, Ziqi, Hao, Jie, Song, Wei
Abstract
Video watermarking underpins copyright protection and provenance for generated media, yet almost every video is compressed by a codec before it is stored or shared. A codec discards precisely the perceptually redundant components that most watermarks rely on, so the payload is often lost even when the marked video looked flawless beforehand. Existing methods leave this path open, since they treat compression as one entry in a generic list of distortions, while a real codec is not differentiable and cannot enter gradient-based training. We present COVER, the first learned video watermark built around codec compression as its design target, which survives that compression by embedding the payload in the latent space of a frozen generative video autoencoder and recovering it by re-encoding the received video into that same latent space. To make codec robustness trainable, we build a differentiable codec surrogate bank that simulates the dominant degradation modes of practical compression, and we train the embedder and the latent decoder through three shared recovery paths under a fidelity objective that constrains the residual in the pixel and frequency domains. Across four codecs at 12 settings, COVER attains 93.72% average bit accuracy, ranks first on 11 of the 12, improves the strongest prior method by 2.68 points, and lifts the worst operating point from 68.90% to 73.72% while each marked video stays visually close to the source clip that produced it.
Chinese Translation
视频水印为生成媒体的版权保护和来源溯源提供了支撑,然而几乎所有视频在存储或共享之前都会经过编解码器压缩。编解码器会丢弃水印通常所依赖的感知冗余分量,因此即使带有水印的视频在压缩前看起来完好无损,其载荷也常常会丢失。现有方法未能解决这一问题,因为它们将压缩视为通用失真列表中的一项,而真实的编解码器不可微,无法进入基于梯度的训练。我们提出了COVER,这是首个以编解码器压缩为设计目标构建的学习型视频水印方法。它通过将载荷嵌入冻结的生成式视频自编码器的潜在空间,并将接收到的视频重新编码至同一潜在空间来恢复载荷,从而在压缩中存活下来。为了使编解码器鲁棒性可训练,我们构建了一个可微的编解码器代理库,用以模拟实际压缩的主要退化模式,并在一个约束像素域和频域残差的保真度目标下,通过三条共享的恢复路径训练嵌入器与潜在解码器。在四种编解码器的12种设置下,COVER取得了93.72%的平均比特准确率,在12项中的11项排名第一,比最强先前方法提升2.68个百分点,并将最差工作点从68.90%提升至73.72%,同时每个带水印视频在视觉上仍与生成它的源视频片段保持高度接近。
cs.CV / 76 / 2609.26274

AIGC Video Detection based on the fusion of spatial-frequency-optical flow multimodal features

Hong, S., Wang, X. Q., Zhang, C., Wang, J. C., Duan, P. X., Wang, Y. W.
Abstract
The rapid evolution of generative AI (e.g., Sora, Hunyuan) makes it essential to develop effective detection strategies that can generalize across ever-evolving synthesis techniques. This study is motivated by the observation of a fundamental challenge in generative models: the inherent difficulty of maintaining cross-modal consistency between appearance and motion. To this end, we propose a multi-modal framework for AIGC video forgery detection tasks, named Cross-Attention based Video Forgery Detector (CrossAtt-VFD), based on joint multi-view analysis of content.Methodologically, we introduce a dual-branch architecture that simultaneously extracts spatial-frequency and optical-flow features.This approach enables the modeling of videos from complementary perceptual perspectives.The core of this process is a dedicated cross-attention mechanism, which governs the alignment of the two modalities and translates cross-modal inconsistencies into a potent diagnostic signal. This multi-modal strategy facilitates the detection of motion that is statistically inconsistent with the visual appearance of a scene. Comprehensive experimental results demonstrated that our model achieves an accuracy of 94.22%, a precision of 91.67 %,and a recall of 96.25 %, effectively verifying the advantages of the multi-modal fusion strategy.
cs.CV / 77 / 2609.26299

ForeDrive: Foresight-Guided End-to-End Autonomous Driving with a Planning-Relevant Latent World Model

ForeDrive:基于前瞻引导与规划相关潜在世界模型的端到端自动驾驶
Wang, Sinuo, Gu, Zichong, Huang, Yuhan, Wen, Wenxin, Yang, Xun, Zhang, Yiqing, Zhang, Xingyu, Che, Ningyu, Ling, Jie, Yu, Qiankun, Liu, Wei, Xu, Jing, Wang, Xinggang
Abstract
Existing latent world models are typically optimized for future predictability, yet the resulting representations are not necessarily useful for planning in autonomous driving. Predictions are commonly used for pretraining or auxiliary supervision rather than as direct conditioning signals for trajectory generation. We propose ForeDrive, which learns a planning-relevant latent representation and couples it asymmetrically to a Diffusion Transformer (DiT) planner. The planner consumes multi-horizon latent future representations learned with a JEPA-style world model; planning gradients update the shared online encoder, while stop-gradient routing trains the latent predictor with forecasting losses only. Because predicted futures have varying reliability across horizons and BEV trajectories are misaligned with image tokens, we use gated visual fusion, future-status injection, and Trajectory-Adaptive Bias (TAB) to inject future latents as guidance without overriding the current observation. Trained with pure imitation learning and using only the current front-view image as visual input at inference, ForeDrive attains 89.9 PDMS on NAVSIM v1 and 90.0 one-stage EPDMS on NAVSIM v2, without reinforcement learning or an external trajectory scorer.
Chinese Translation
现有的潜在世界模型通常以未来可预测性为目标进行优化,但由此得到的表示未必对自动驾驶中的规划有用。预测通常仅用于预训练或辅助监督,而非作为轨迹生成的直接条件信号。我们提出ForeDrive,它学习一种与规划相关的潜在表示,并将其非对称地耦合到扩散Transformer(Diffusion Transformer, DiT)规划器中。该规划器使用由JEPA风格世界模型学习的多时间跨度潜在未来表示;规划梯度更新共享的在线编码器,而通过停止梯度路由,潜在预测器仅用预测损失进行训练。由于预测的未来在不同时间跨度上的可靠性各异,且BEV轨迹与图像token不对齐,我们采用门控视觉融合、未来状态注入以及轨迹自适应偏置(Trajectory-Adaptive Bias, TAB),将未来潜在表示作为引导信息注入而不覆盖当前观测。ForeDrive仅通过纯模仿学习训练,推理时仅使用当前前视图图像作为视觉输入,在NAVSIM v1上达到89.9 PDMS,在NAVSIM v2上达到90.0单阶段EPDMS,且无需强化学习或外部轨迹评分器。
cs.CV / 78 / 2609.26325

Leveraging Vision-Based Point Cloud Map Priors for Camera-Based 3D Object Detection and Online Vectorized HD Mapping

利用基于视觉的点云地图先验进行基于相机的3D目标检测与在线矢量化高精地图构建
Käppeler, Markus, Mohan, Rohit, Valada, Abhinav
Abstract
Camera-based 3D object detection and online vectorized HD mapping provide compact scene representations for autonomous driving, but both depend on accurate metric geometry and remain limited by depth ambiguity. Over long-term deployment, observations from repeated traversals can be accumulated into persistent point cloud priors that provide geometric context beyond the current observations. Existing explicit point cloud prior approaches, however, rely on LiDAR-based map construction and therefore require expensive 3D ranging sensors. We propose a framework that constructs a static point cloud prior map from previous camera traversals using Pi3X and augments each point with DINOv3 features. At runtime, a local prior patch is retrieved using global localization, encoded with a sparse voxel backbone, and fused in bird's-eye view (BEV) with lifted multi-view camera features. Task-specific sparse transformer heads then predict 3D objects and vectorized map elements from the fused representation. On Argoverse 2, the vision-based prior improves a strong baseline from 0.287 to 0.299 CDS and from 0.669 to 0.750 vectorized mapping mAP. Ablations show that semantic DINOv3 features are particularly important for vectorized mapping. These results demonstrate that vision-built geometric-semantic priors provide an effective form of long-term scene memory for camera-based perception, improving both tasks without LiDAR for prior-map construction or online inference.
Chinese Translation
基于相机的3D目标检测与在线矢量化高精地图构建为自动驾驶提供了紧凑的场景表示,但两者都依赖于精确的度量几何,并受深度歧义性的限制。在长期部署过程中,重复行驶获得的观测可以累积为持久化的点云先验,从而提供超越当前观测的几何上下文信息。然而,现有的显式点云先验方法依赖于基于LiDAR的地图构建,因此需要昂贵的3D测距传感器。我们提出了一个框架,利用Pi3X从先前的相机行驶数据中构建静态点云先验地图,并为每个点附加DINOv3特征。在运行时,通过全局定位检索局部先验图块,使用稀疏体素骨干网络进行编码,并在鸟瞰图(BEV)视角下与提升的多视角相机特征进行融合。随后,任务特定的稀疏Transformer头从融合表示中预测3D目标和矢量化地图元素。在Argoverse 2数据集上,基于视觉的先验将强基线的CDS从0.287提升至0.299,矢量化建图mAP从0.669提升至0.750。消融实验表明,语义DINOv3特征对矢量化建图尤为重要。这些结果表明,由视觉构建的几何-语义先验为基于相机的感知提供了一种有效的长期场景记忆形式,无需LiDAR进行先验地图构建或在线推理即可同时提升两项任务。
cs.CV / 79 / 2609.26334

On the Role of the Projector in Contrastive Self-Supervised Learning: Last-Layer Rank Dynamics Drive Representation Quality

论投影器在对比自监督学习中的作用:末层秩动态驱动表征质量
Manna, Siladittya, Mandal, Priyangshu, Pal, Umapada, Bhattacharya, Saumik
Abstract
The dimensional collapse of representations in self-supervised contrastive learning is an ever-present issue. One notable technique to prevent such a collapse of representations is using a multi-layered perceptron network called Projector. In several works, the projector has been found to heavily influence the quality of representations learned in a self-supervised contrastive pre-training task. However, the question still lingers. What role does the projector play? Assuming the projector mitigates dimensional collapse, what prevents the terminal layer of the base encoder from functioning as the projector in the absence of an explicit multi-layer perceptron (MLP) head? In this work, we intend to study what happens inside the projector by examining the rank dynamics of the same and the encoder through empirical study and analysis. Through mathematical analysis, we observe that the effect of rank reduction predominantly occurs in the last layer. Motivated by this insight, we propose a weight regularization strategy applied specifically to the last layer. We demonstrate that this targeted approach yields better performance than applying orthogonal weight regularization across the entire network (WeRank), both with and without a projector. Our method improves Top-1 accuracy by more than 1% on SimCLR on the ImageNet100 dataset and consistently outperforms baseline SimCLR variants on CIFAR datasets, supporting our interpretation of the projector's role.
Chinese Translation
自监督对比学习中表征的维度坍缩是一个长期存在的问题。防止这种表征坍缩的一种重要技术是使用一个称为投影器(Projector)的多层感知机网络。在多项研究中,投影器被发现会显著影响自监督对比预训练任务中所学习表征的质量。然而,问题依然存在:投影器究竟扮演了什么角色?假设投影器能够缓解维度坍缩,那么在没有显式多层感知机(MLP)头的情况下,是什么阻止了基础编码器的末端层发挥投影器的功能?在本工作中,我们旨在通过实证研究和分析,考察投影器内部发生了什么,以及投影器与编码器的秩动态。通过数学分析,我们观察到秩降低的效应主要发生在最后一层。受这一发现的启发,我们提出了一种专门作用于最后一层的权重正则化策略。我们证明,这种有针对性的方法比在整个网络上应用正交权重正则化(WeRank)取得更好的性能,无论是否使用投影器均如此。我们的方法在 ImageNet100 数据集上将 SimCLR 的 Top-1 准确率提升了超过 1%,并在 CIFAR 数据集上持续优于基线 SimCLR 变体,从而支持了我们对投影器角色的解释。
cs.CV / 80 / 2609.26375

KwaiMind Technical Report

Wu, Junlong, Li, Zijun, Hu, Yuting, Sun, Jia, Wei, Pengcheng, Zhou, Yimin, Wang, Honglie, Wang, Huaiqing, Fan, Dewen, Zuo, Fei, Gao, Haixuan, Peng, Lihui, She, Tingxuan, Li, Yuqing, Zhang, Boheng, Yang, Fan, Ou, Wenwu
Abstract
Commercial image editing requires product identity preservation, accurate text rendering, and user appeal alongside general editing quality. We present KwaiMind, an image editing system combining general capabilities with e-commerce specialization. An agent-based data engine maintains approximately 1.8 million high-quality editing pairs. Built on a multimodal diffusion transformer, KwaiMind undergoes continued pre-training and supervised fine-tuning, followed by preference optimization and online reinforcement learning. A general-purpose vision-language judge and specialized rewards for click-through rate (CTR), text rendering, and product consistency guide specialized policies, which are consolidated through on-policy distillation. We introduce Ecom-Bench, covering 11 commercial editing tasks with task-specific visual evaluation and CTR-based ranking. KwaiMind achieves the strongest overall scores among evaluated open-source editors on ImgEdit, GEdit, both language splits of REDEdit, and Ecom-Bench visual quality, and the highest aggregate CTR ranking score among compared systems. Offline, CTR-guided optimization increases the proportion of generated images whose predicted CTR exceeds that of the original product image from 12.16% to 37.41%. In an online A/B experiment, CTR-based selection of product main images yields an approximately 2.44% relative increase in actual CTR. These results demonstrate the value of domain-specific data and reward-driven alignment for commercial image editing.
cs.CV / 81 / 2609.26425

QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for World Models and Video Generation

QuantWM:面向世界模型与视频生成的时间一致性2比特KV缓存量化
Zhao, Jiaqi, Hu, Xiaobin, Yin, Bo, Jiang, Junpeng, Zhang, Miao, Yan, Shuicheng
Abstract
KV cache memory has become a major deployment bottleneck for video generation and world models, which motivates low-bit quantization study for efficiency. Existing 2-bit KV cache quantization methods can achieve nearly lossless performance on video benchmarks such as VBench, however, we find that they still cause severe temporal flickering and visual degradation. Meanwhile, deeper investigates show that Key quantization produces smaller reconstruction errors than Value, but surprisingly leads to much larger output degradation. We trace this discrepancy to attention: small Key perturbations can change the attention logits, i.e., QK^\top, and shift the temporal-spatial tokens selected by Queries. These observations motivate us to explicitly preserve attention logits and temporal-spatial token selection during KV cache quantization to alleviate the visual degradation problem. To address this issue, we present QuantWM, a training-free and strictly causal 2-bit KV cache quantization framework. QuantWM introduces two complementary techniques to mitigate the attention shifts. Firstly, quantization-sensitivity-aware clustering (QSAC) jointly considers historical Query sensitivity and residual ranges to select INT2-friendly Key centroids, which reduces quantization errors in channels that are more critical to attention. In addition, principal-subspace attention compensation (PSAC) restores the remaining Key errors along the dominant Query subspace using low-rank projections, which provides a direct and efficient correction to stabilize attention logits. Extensive experiments on Causal-Forcing, LingBot-World-v2, HY-World 1.5, Matrix-Game-2 and Longcat-Video demonstrate that QuantWM significantly improves visual quality and temporal consistency, while outperforming existing methods across image and video quality metrics with up to 6.20x KV cache memory compression and limited additional overhead.
Chinese Translation
KV缓存内存已成为视频生成与世界模型部署的主要瓶颈,这促使人们研究低比特量化以提升效率。现有的2比特KV缓存量化方法在VBench等视频基准上可实现近乎无损的性能,然而我们发现它们仍会导致严重的时间闪烁和视觉质量下降。同时,更深入的研究表明,Key量化的重构误差比Value更小,但令人意外的是,其导致的输出退化却大得多。我们将这一差异归因于注意力机制:Key的微小扰动会改变注意力分数(即QK^T),并移动由Query选定的时序-空间token。这些观察促使我们在KV缓存量化过程中显式地保持注意力分数和时序-空间token选择,以缓解视觉退化问题。为解决这一问题,我们提出了QuantWM,一个无需训练且严格因果的2比特KV缓存量化框架。QuantWM引入两种互补技术来缓解注意力偏移。首先,量化敏感度感知聚类(QSAC)联合考虑历史Query敏感度和残差范围,以选择对INT2友好的Key质心,从而降低对注意力更关键通道的量化误差。此外,主子空间注意力补偿(PSAC)利用低秩投影沿主导Query子空间恢复剩余的Key误差,为稳定注意力分数提供直接而高效的校正。在Causal-Forcing、LingBot-World-v2、HY-World 1.5、Matrix-Game-2和Longcat-Video上的大量实验表明,QuantWM显著提升了视觉质量和时间一致性,在图像和视频质量指标上均优于现有方法,同时实现了最高6.20倍的KV缓存内存压缩,且额外开销有限。
cs.CV / 82 / 2609.26430

Latent Dataset Distillation for Human Motion Prediction

Tian, Ge, Li, Guang, Ogawa, Takahiro, Haseyama, Miki
Abstract
Dataset distillation (DD) compresses a large training set into a compact synthetic set while preserving downstream training utility. Although DD has been widely studied for images and recently extended to time-series forecasting, its application to human motion prediction remains largely unexplored. Human motion is high-dimensional and structurally coupled, and gradient matching (GM) in the original motion space optimizes many correlated variables without a prior on pose plausibility or temporal dynamics, which frequently yields implausible and unstable synthetic motions. To address this limitation, we propose a latent DD framework that regularizes distillation with a learned motion prior. Motions are first compressed by a residual-quantized variational autoencoder (RVQ-VAE), and distillation then updates only a learnable latent bank through the frozen quantizer and decoder. The pretrained decoder restricts synthetic motions to its output space, while residual quantization progressively refines the latent approximation across multiple codebooks and alleviates the representational bottleneck of single-stage vector quantization. Experiments on Human3.6M, CMU, and 3DPW with two prediction backbones show that the proposed framework outperforms direct GM in 27 of 30 evaluated settings and random subsets in every setting, and produces visibly more plausible synthetic motions in qualitative comparisons.
cs.CV / 83 / 2609.26443

Mammo-LIFE: Longitudinal Mammographic Imaging and Clinical Feature Enrichment for Post-Radiotherapy Outcome Prediction

Bayatmakou, Farnoush, Hosseini, Maryam, Taleei, Reza, Mohammadi, Arash
Abstract
Recent advances in Artificial Intelligence (AI)-powered Computer-Aided Diagnosis (CAD) systems have substantially improved breast cancer screening, diagnosis, and prognosis. Comparatively, postradiotherapy outcome prediction using paired longitudinal mammograms has received considerably less attention. This is largely due to the limited availability of well-annotated longitudinal datasets. Longitudinal mammograms, coupled with paired pre- and post-treatment information, provide a unique opportunity to characterize treatment-induced breast tissue changes following radiotherapy. The resulting learned representations can serve as a valuable asset for advancing personalized radiotherapy planning and post-treatment management. In this context, we propose Mammo-LIFE, a patient-level multimodal framework for post-radiotherapy outcome prediction that combines longitudinal mammographic features with patient-level clinical variables. The imaging branch processes paired pre- and post-treatment mammograms acquired from the four standard views using a mammography-specific encoder adapted via Low-Rank Adaptation (LoRA). Within each view, preand post-treatment representations are explicitly compared through a longitudinal comparison module to capture treatment-related changes. The resulting view-level embeddings are then aggregated using learned view-attention pooling to form a unified patient-level mammographic representation. Selected clinical variables are subsequently combined with the image-derived prediction probability through a late-fusion strategy. To evaluate the effectiveness of combining paired longitudinal mammograms with clinical information, experiments were conducted on an in-house clinical cohort using patient-level stratified five-fold cross-validation.
cs.CV / 84 / 2609.26458

Code Plans, Diffusion Renders: Open-Ended Generative World Modeling

代码规划,扩散渲染:开放式生成式世界建模
Fang, Zixun, Shao, Yawen, Zhu, Kai, Xiao, Jie, Chen, Shihan, Liu, Yu, Fu, Xueyang, Cao, Yang, Zhai, Wei, Zha, Zheng-Jun
Abstract
We introduce \textbf{CoDeR}, a new paradigm for world modeling. Unlike existing video world models that implicitly represent world dynamics through visual observations, our system explicitly constructs an executable world with code and employs video generation models for visual realization. Specifically, we coordinate five complementary roles to translate high-level concepts into structured world rules, executable dynamics, and perceptual observations. This design enables \textit{long-term memory}, \textit{open-ended interactions}, \textit{autonomous world evolution}, and \textit{multi-agent scenarios}, where multiple entities can act, interact, and evolve persistently beyond the current observation. Extensive experiments demonstrate that our framework substantially extends the capabilities of existing world models, enabling long-term memory, open-ended interactions, autonomous evolution, and persistent multi-agent dynamics, while achieving state-of-the-art performance across multiple evaluation settings. Code and model weights will be made publicly available. Project Page: \href{https://becauseimbatman0.github.io/CoDeR}{CoDeR}.
Chinese Translation
我们提出了 CoDeR,一种全新的世界建模范式。与现有的通过视觉观测隐式表示世界动态的视频世界模型不同,我们的系统显式地用代码构建一个可执行的世界,并利用视频生成模型进行视觉呈现。具体而言,我们协调五个互补的角色,将高层概念转化为结构化的世界规则、可执行的动态以及感知观测。这一设计实现了长期记忆、开放式交互、自主世界演化以及多智能体场景,使多个实体能够在当前观测之外持续地进行行动、交互和演化。大量实验表明,我们的框架显著扩展了现有世界模型的能力,实现了长期记忆、开放式交互、自主演化以及持久的多智能体动态,并在多个评估设置中取得了最先进的性能。代码和模型权重将公开发布。项目主页:CoDeR(https://becauseimbatman0.github.io/CoDeR)。
cs.CV / 85 / 2609.26463

Complementary Roles of Radiomics and Foundation Representations in Renal Cell Carcinoma Classification: A Comparative Study of 2D and 3D CT Encodings

影像组学与基础模型表征在肾细胞癌分类中的互补作用:2D与3D CT编码的对比研究
Liang, Yuan, Bhattacharjee, Sourav, Campbell, Abraham
Abstract
Accurate preoperative subtype classification of renal cell carcinoma (RCC) from contrast-enhanced computed tomography remains clinically challenging. Radiomics provides structured tumour descriptors, whereas foundation representations offer transferable image features. However, it remains unclear whether radiomics still adds value beyond pretrained representations, and how 2D and 3D MedVAE encoders compare in this setting. We compared handcrafted radiomics, 2D MedVAE, 3D MedVAE, and their fusion for binary clear-cell RCC versus non-clear-cell RCC classification on KiTS23 under a unified preprocessing pipeline. Concatenation, cross-attention, and gated fusion were evaluated as representative integration strategies, and radiomics feature importance was analysed to support decision-centric interpretability. Fusion consistently improved discrimination over image-only MedVAE branches. The best overall performance was achieved by 3D gated fusion, with an AUC of 82.7\%, outperforming the best 2D fusion model (79.6%), the radiomics baseline (74.4%), and the single-modality MedVAE branches. Ablation analysis further showed clear gains of the full fusion model over both image-only and radiomics-only variants, indicating complementary contributions from radiomics and image representations. These findings suggest that radiomics remains relevant for RCC CT classification in the presence of foundation representations, and that its integration with MedVAE is more effective in the 3D setting. More broadly, the study supports a complementary role for radiomics and foundation representations in clinically meaningful imaging decision support.
Chinese Translation
基于增强CT的肾细胞癌(RCC)术前亚型准确分类在临床上仍具挑战性。影像组学提供结构化的肿瘤描述特征,而基础模型表征则提供可迁移的图像特征。然而,在预训练表征存在的情况下,影像组学是否仍具有附加价值,以及2D与3D MedVAE编码器在此场景下的对比效果如何,目前仍不清楚。我们在统一的预处理流程下,于KiTS23数据集上比较了手工影像组学特征、2D MedVAE、3D MedVAE及其融合方法在透明细胞RCC与非透明细胞RCC二分类任务中的表现。我们评估了拼接、交叉注意力和门控融合作为代表性整合策略,并通过影像组学特征重要性分析支持以决策为中心的可解释性。融合方法在判别能力上始终优于仅使用图像的MedVAE分支。3D门控融合取得了最佳整体性能,AUC达82.7%,优于最佳2D融合模型(79.6%)、影像组学基线(74.4%)以及单模态MedVAE分支。消融分析进一步表明,完整融合模型相较于仅图像和仅影像组学的变体均有明显提升,说明影像组学与图像表征具有互补性贡献。这些发现表明,即使在基础模型表征存在的情况下,影像组学在RCC CT分类中依然具有价值,且其与MedVAE的整合在3D设置中更为有效。更广泛而言,本研究支持影像组学与基础模型表征在具有临床意义的影像决策支持中的互补作用。
cs.CV / 86 / 2609.26474

PP-Net: A Hybrid Physical-Prior Neural Network for Scattered Light Removal in Biomedical Images on Embedded Devices

PP-Net:一种用于嵌入式设备生物医学图像散射光去除的物理先验混合神经网络
Guo, Yongfei, Chu, Tingjin, Liu, Mengzhuo, Lou, Hongwei, Gong, Yuanhao
Abstract
Scattered light is common in biomedical images, yet its removal remains challenging. The difficulty arises from three aspects: first, aligned scattered-light-free biomedical ground truth is often unavailable; second, scattering is coupled with weak illumination and sensor-induced noise; and third, many learning-based restoration models are computationally expensive for embedded devices in Internet of Medical Things (IoMT) scenarios. To address these issues, this paper proposes PP-Net, a hybrid physical-prior neural network for biomedical scattered light removal. The proposed method consists of three components: DFN-Net suppresses sensor-induced noise, ASAP estimates the scattering map and recovers a physics-based prior map, and GF-Net refines the prior map by fusing it with the denoised observation. To reduce the dependence on paired biomedical ground truth, a progressive synthetic training and cross-domain transfer strategy is developed. Experiments show that the physical-prior branch improves the peak signal-to-noise ratio (PSNR) by up to 1.26 dB on paired synthetic benchmarks. Under joint noise-and-scattering degradation, PP-Net improves PSNR by more than 10.8 dB and the structural similarity index measure (SSIM) by more than 0.62 compared with representative baseline methods. On real W2S biomedical images, the proposed method reduces the average Natural Image Quality Evaluator (NIQE) score by 43.3\%. Edge deployment with RKNN conversion and INT8 quantization achieves an average inference latency of approximately 200 ms per $512\times512$ image over 360 test images. These results demonstrate that PP-Net provides an effective and deployable solution for microscopic imaging, endoscopic inspection, and edge-assisted biomedical analysis in IoMT scenarios.
Chinese Translation
散射光在生物医学图像中普遍存在,然而其去除仍然具有挑战性。其难点来自三个方面:首先,对齐的无散射光生物医学真实标签(ground truth)通常难以获取;其次,散射与弱光照及传感器引入的噪声相互耦合;第三,许多基于学习的复原模型对于医疗物联网(IoMT)场景中的嵌入式设备而言计算开销过大。针对这些问题,本文提出了一种用于生物医学散射光去除的物理先验混合神经网络PP-Net。所提出的方法由三个组件构成:DFN-Net用于抑制传感器引入的噪声,ASAP用于估计散射图并恢复基于物理的先验图,GF-Net通过将先验图与去噪后的观测图像融合对其进行精炼。为减少对成对生物医学真实标签的依赖,本文提出了一种渐进式合成训练与跨域迁移策略。实验表明,物理先验分支在成对合成基准上可将峰值信噪比(PSNR)最多提升1.26 dB。在噪声与散射联合退化条件下,与具有代表性的基线方法相比,PP-Net将PSNR提升超过10.8 dB,结构相似性指数(SSIM)提升超过0.62。在真实的W2S生物医学图像上,所提方法将平均自然图像质量评估器(NIQE)分数降低了43.3%。通过RKNN转换和INT8量化的边缘部署,在360张测试图像上实现了每张512×512图像约200 ms的平均推理延迟。这些结果表明,PP-Net为IoMT场景中的显微成像、内窥镜检查以及边缘辅助生物医学分析提供了一种有效且可部署的解决方案。
cs.CV / 87 / 2609.26484

From Token Importance to Conditional Removability: Rethinking Visual Token Pruning in Multimodal Large Language Models

He, Shengli, Liang, Yongchao, He, Roumeng, Zeng, Junjie, He, Jiyuan, Fang, Xin, Wu, Can, Zheng, Li
Abstract
Training-free visual-token pruning often uses token importance, redundancy, or related selection criteria as proxies for safe removal. We show that these signals alone do not fully characterize removability, which is conditioned on both representation depth and the surrounding deletion set. Controlled interventions demonstrate that removing the same tokens at different depths produces substantially different downstream perturbations, while changing only the deletion context at a fixed depth alters candidate marginals and pruning-boundary decisions. These findings show that token importance alone cannot determine when a token is safely removable or how its removability changes under joint deletion. Motivated by this perspective, we propose CoRePrune, a training-free two-stage framework. Progressive Perturbation-Aware Visual Pruning refreshes deletion effects as visual representations evolve, while Set-Conditioned Refinement reevaluates candidate rescue benefits under the current deletion set after visual--text interaction. Across five multimodal large language model backbones covering standard images, high-resolution inputs, and video, CoRePrune preserves performance under aggressive token budgets. On Qwen3.5, with a final budget of 128 visual tokens, it retains 90.3% of dense-model performance while reducing aggregate prefill time by 51.0%.
cs.CV / 88 / 2609.26492

Radiomics-Conditioned Modulation of RenalCLIP Features for Clear Cell Renal Cell Carcinoma Classification

基于影像组学条件化调制的RenalCLIP特征用于透明细胞肾细胞癌分类
Liang, Yuan, Bhattacharjee, Sourav, Campbell, Abraham
Abstract
Radiomics provides quantitative descriptions of tumour appearance that may complement disease-specific foundation models in small labelled cohorts. We investigate this complementarity for computed tomography-based classification of clear cell renal cell carcinoma. Our framework uses radiomics to modulate RenalCLIP features through feature-wise linear modulation (FiLM), while retaining a direct radiomics contribution. Internal testing and external validation compare it with conventional fusion strategies and reference classifiers. The FiLM model achieves an area under the receiver operating characteristic curve (AUC) of 0.804 internally and 0.854 externally, with the highest mean AUC among the evaluated RenalCLIP fusion strategies in both cohorts. Pathway ablations examine the contributions of conditional modulation and the direct radiomics residual, while feature permutation highlights the role of tumour texture. These findings support radiomics as a useful complement to RenalCLIP in a small labelled cohort and identify FiLM as an effective approach to integrating their representations for robust renal tumour classification.
Chinese Translation
影像组学提供了对肿瘤外观的定量描述,在小样本标注队列中可以补充疾病特异性基础模型。我们在基于计算机断层扫描(CT)的透明细胞肾细胞癌分类任务中研究了这种互补性。我们的框架利用影像组学通过特征级线性调制(FiLM)对RenalCLIP特征进行调制,同时保留影像组学的直接贡献。通过内部测试和外部验证,我们将其与传统融合策略及参考分类器进行了比较。FiLM模型在内部测试中实现了0.804的受试者工作特征曲线下面积(AUC),在外部验证中达到0.854,并且在两个队列中均为所有被评估的RenalCLIP融合策略中平均AUC最高者。通路消融实验考察了条件调制和影像组学直接残差通路的贡献,特征置换分析则凸显了肿瘤纹理特征的作用。这些发现支持影像组学在小样本标注队列中可作为RenalCLIP的有效补充,并表明FiLM是整合二者表征、实现稳健肾肿瘤分类的有效方法。
cs.CV / 89 / 2609.26505

Semantically-Guided Domain Randomization for Industrial Object Detection in Low-Image-Budget Regimes

面向低图像预算场景下工业目标检测的语义引导域随机化方法
Araya-Martinez, Jose Moises, Mohan, Gautham, Lambrecht, Jens
Abstract
Retraining visual perception pipelines in High-Mix, Low-Volume (HMLV) automotive manufacturing must be carried out under tight annotation, energy, and time budgets, yet most Synthetic Data Generation (SDG) strategies still operate in the thousands of images. This work evaluates Semantically-Guided Domain Randomization (S-GDR), an annotation-free adaptation pipeline that couples Vision-Language Model (VLM)-based semantic captioning of a small unannotated real reference set with diffusion-based background synthesis (Stable Diffusion XL (SDXL) conditioned by ControlNet and IP-Adapter) and mask-based object composition. On an automotive multi-object detection benchmark and with a fixed budget of 200 synthetic training images, S-GDR reaches mAP50-95 = 0.739 on a real held-out test set, outperforming a domain-randomized render baseline (mAP50-95 = 0.697) as well as brightness filtering, perceptual hashing, CycleGAN style transfer, and unguided diffusion variants sharing the same 200-image budget. These initial observations position S-GDR as a promising annotation- free alternative for extreme data-scarcity regimes.
Chinese Translation
在高混合、小批量(HMLV)汽车制造中重新训练视觉感知流程时,必须在严格的标注、能源和时间预算下进行,然而大多数合成数据生成(SDG)策略仍需要数千张图像。本工作评估了语义引导域随机化(Semantically-Guided Domain Randomization, S-GDR),这是一种无需标注的适配流程,它将基于视觉-语言模型(VLM)对小型无标注真实参考集进行语义描述,与基于扩散模型的背景合成(由 ControlNet 和 IP-Adapter 条件化的 Stable Diffusion XL,SDXL)以及基于掩码的物体合成相结合。在一个汽车多目标检测基准上,并固定 200 张合成训练图像的预算下,S-GDR 在真实留出测试集上达到 mAP50-95 = 0.739,优于域随机化渲染基线(mAP50-95 = 0.697),也优于共享相同 200 张图像预算的亮度过滤、感知哈希、CycleGAN 风格迁移以及无引导扩散等变体。这些初步观察结果表明,S-GDR 是极端数据稀缺场景下一种有前景的免标注替代方案。
cs.CV / 90 / 2609.26512

Do Vision Model See Like the Brain? A Comparison Across EEG Encoding Model

Baghel, Shashank, Dwivedi, Kshitij, Singh, Dinesh, Nara, Sanjeev
Abstract
Convolutional neural networks (CNNs) and vision transformers are both used to model the human visual system, but whether the two architectures diverge at a specific point in network depth is unclear. We compared six CNNs and two vision transformers by computing the Pearson correlation (r) between each model's predicted and measured EEG response at every layer or block, in ten participants viewing 200 natural images. For the transformer models, we also tested four token representations, from the classification (CLS) token alone to CLS combined with all patch tokens. CNNs showed strongest correspondence at the earliest layers, weakening at deeper layers, particularly later in the post-stimulus response. Transformers instead sustained strong correspondence at their deepest blocks, though not at their earliest ones. This advantage depended on token representation: pooled representations gave weaker peak correlations (r approx 0.48-0.51) than representations retaining all patch tokens (r=0.640 for CLIP-ViT-B/32, r=0.656 for DINOv2-ViT-B/14). Controlled comparisons showed architecture, not training objective, drove this effect: MoCo-v1 and ResNet-50 (matched architecture) performed nearly identically (r=0.673, 0.670), whereas CLIP-RN50 and CLIP-ViT-B/32 (matched objective) diverged until patch tokens were preserved. We propose that CNN training's classification bottleneck compresses brain-relevant information at depth, unlike transformers' self-attention and non-classification objectives. A spatial topography analysis showed a common occipital-dominant pattern across all models, indicating these differences reflect signal strength and persistence rather than distinct brain regions. Patch-preserving transformer representations sustain brain-predictive correspondence where CNNs collapse.
cs.CV / 91 / 2609.26513

Virtual Encoders in Multimodal Transformers

多模态Transformer中的虚拟编码器
Ogata, Katsuya, Nakashima, Yuta
Abstract
Multimodal language models traditionally rely on dedicated perceptual encoders to construct task-usable representations. More integrated architectures have recently emerged, which instead expose the shared transformer to lightly projected patches, audio frames, or discrete visual tokens. Where does this encoding happen when such representations are not provided? We find that the transformer can internalize this missing computation, constructing task-usable perceptual representations within its own early-to-middle layers before the downstream language model. We call this computational structure a Virtual Encoder. Across linear probing, similarities to perceptual encoders, and causal analyses, we identify signatures of this structure in models that receive perceptual tokens without continuous encoder-derived features. These analyses also suggest that the boundary between perception and language processing need not coincide within an architectural module. Instead, encoder-like computation can emerge as a functional regime within a shared transformer, providing a new perspective for understanding where and how multimodal models process perception.
Chinese Translation
多模态语言模型传统上依赖于专用的感知编码器来构建可供任务使用的表示。近来出现了更加一体化的架构,它们将轻量投影后的图像块、音频帧或离散视觉token直接输入共享的Transformer。当此类表示未被提供时,这种编码过程发生在何处?我们发现,Transformer能够将这一缺失的计算过程内化,在其自身的早中期层、于下游语言模型之前构建出可供任务使用的感知表示。我们将这种计算结构称为虚拟编码器(Virtual Encoder)。通过线性探针(linear probing)、与感知编码器的相似性分析以及因果分析,我们在那些接收感知token但不具备连续编码器特征的模型中识别出了这一结构的特征。这些分析还表明,感知与语言处理之间的边界不必与架构模块的边界重合。相反,类似编码器的计算可以在共享的Transformer中作为一种功能性状态涌现,为理解多模态模型在何处以及如何处理感知信息提供了新的视角。
cs.CV / 92 / 2609.26549

Latent Commonality Expectation-Maximisation for Box-supervised Tree Crown Instance Segmentation

Pitts, Thomas, Li, Kunqi, Liang, Bin
Abstract
Individual tree crown segmentation from aerial imagery underpins tree-level carbon accounting, biodiversity, and restoration monitoring at landscape scale. However, existing models are predominantly trained on dense canopy forest imagery and degrade in savannah and drylands, where tree crowns are sparse, of variable appearance, and underrepresented in annotated benchmarks. These models also typically depend on costly polygon annotations. We introduce LACE (LAtent Commonality Expectation-maximisation), a box-supervised instance segmentation model, evaluated on 0.1 m/px aerial RGB tree crown imagery. LACE uses a frozen DINOv3-web ViT-L/16 encoder, applied at four spatial offsets and interlaced into a denser feature grid, with a lightweight CenterNet-style detection head trained solely on bounding boxes. We use expectation-maximisation to separate recurring appearance, the "treeness", within bounding boxes from surroundings. On the OAM-TCD benchmark test set, LACE reaches a mask AP$_{50}$ of $0.663 \pm 0.001$ (3 seeds) trained on 900 box-annotated images and without mask annotations, above the 0.626 scored by Restor's released mask-supervised Mask R-CNN, which was trained on the full ~4.2k image set. On a sparse-canopy holdout set, mask AP$_{50}$ rises to $0.691$ versus $0.612$ for Detectree2, a mask-supervised baseline. On NeonTreeEvaluation, using the official evaluation code, LACE reaches $0.728 \pm 0.003$ F1@0.4 (5 seeds) from 23,424 hand-annotated RGB boxes alone, matching the authors' DeepForest model's published 0.719, using under 0.1% of its training annotations and none of its LiDAR-derived 30M-crown pretraining set. By leveraging frozen self-supervised features, LACE matches or surpasses fully-supervised specialist baselines from boxes alone, removing the need for polygon annotation in tree crown instance segmentation for sparse-canopy environments where labelled data is scarce.
cs.CV / 93 / 2609.26561

Vision Foundation Models with Synthetic-Only Training for Monocular Spacecraft Pose Estimation

Church, John, Nikolian, Vazghen
Abstract
We present an improvement on previous spacecraft pose estimation architectures that results in the lowest published mean rotation errors we know of on the SPEED+ lightbox and sunlamp test sets for a known, non-cooperative spacecraft. By using a previously established heatmap-based pose estimation architecture and adapting a large self-supervised ViT foundation model (DINOv3) in place of the smaller convolutional and ViT encoders of previous work, we show that pose estimation accuracy improves from 300M to 840M parameters with no saturation yet observed. We also evaluate our 840M model on a Jetson Orin NX 16GB, measuring single-pass network inference at 133.8 ms per crop with a board draw of 32.0 W. These measurements demonstrate embedded inference feasibility on a processor family with orbital flight heritage. Our resulting model outperforms previous models across lightbox and sunlamp domains while training only on synthetic data. Our best model, using DINOv3 840M adapted with LoRA as the encoder (rank 64, three-seed ensemble with four-rotation test-time augmentation), results in $1.56^\circ$ mean rotation error on sunlamp and $1.17^\circ$ on lightbox, compared to the previous best mean rotation errors we know of on these test sets, $2.66^\circ$ and $1.75^\circ$ by EagerNet.
cs.CV / 94 / 2609.26578

Radiomics--Foundation Fusion for Interpretable RCC Classification: Internal Benchmarking and Exploratory External Transfer

影像组学—基础模型融合用于可解释的肾细胞癌分类:内部基准测试与探索性外部迁移
Liang, Yuan, Wang, Fangyijie, Curran, Kathleen M., Silvestre, Guénolé, Bhattacharjee, Sourav, Campbell, Abraham
Abstract
Accurate preoperative subtype classification of renal cell carcinoma (RCC) from contrast-enhanced CT remains clinically challenging because clear cell RCC (ccRCC) and non-clear cell RCC often show overlapping imaging appearances. This study evaluates whether foundation representations reduce reliance on handcrafted radiomics, or whether radiomics remains complementary for interpretable tumour characterisation. We compared radiomics, conventional CNN features, MedicalNet-pretrained features, MedVAE representations, and fusion variants for binary ccRCC classification on KiTS23, reporting area under the receiver operating characteristic curve (AUC) with bootstrap confidence intervals and average precision (AP) as a complementary class-imbalance-sensitive metric. We further assessed branch-removal ablation, TCGA/AIMI external transfer, and interpretability using radiomics permutation importance and gate-level analysis. Internally, 3D MedVAE gated fusion achieved the best performance, with an AUC of 82.7% and AP of 92.2%. On the external TCGA cohort, the same model achieved an AUC of 79.5% and AP of 98.9%, although specificity remains uncertain because only two external non-ccRCC cases were available. Gate analysis showed a radiomics-dominant fusion regime, suggesting that foundation representations acted as case-dependent refinement signals rather than replacements for structured tumour descriptors. These findings support radiomics as a complementary and clinically interpretable component of CT-based RCC characterisation in the foundation-model era.
Chinese Translation
由于透明细胞肾细胞癌(ccRCC)与非透明细胞肾细胞癌在影像学表现上常存在重叠,基于增强CT对肾细胞癌(RCC)进行准确的术前亚型分类在临床上仍具挑战性。本研究评估基础模型表征是否能够减少对手工设计影像组学特征的依赖,或者影像组学特征对于可解释的肿瘤表征是否仍具有互补价值。我们在KiTS23数据集上针对二分类ccRCC任务,比较了影像组学、传统CNN特征、MedicalNet预训练特征、MedVAE表征以及多种融合变体,采用受试者工作特征曲线下面积(AUC,附带自助法置信区间)和平均精度(AP)作为对类别不平衡敏感的补充指标进行评价。我们进一步评估了分支移除消融实验、TCGA/AIMI外部迁移,并利用影像组学置换重要性和门控层面的分析进行可解释性研究。在内部数据上,3D MedVAE门控融合取得了最佳性能,AUC为82.7%,AP为92.2%。在外部TCGA队列上,同一模型取得了79.5%的AUC和98.9%的AP,但由于仅有两例外部非透明细胞RCC病例,其特异性仍不确定。门控分析显示融合呈现出影像组学主导的模式,表明基础模型表征更多是作为依赖于具体病例的精细化信号,而非结构化肿瘤描述子的替代品。这些发现支持影像组学在基础模型时代仍然是基于CT的RCC表征中具有互补性和临床可解释性的组成部分。
cs.CV / 95 / 2609.26590

GTR: Gated Token Recurrence for Efficient Dense Prediction

Feng, Zhe, Liu, Longfei, Liu, Wei, Chen, Kai, Kong, Jiangjiang, Zhou, Wei, Qian, Yifeng, Chen, Dexiong, Yu, Xuanlong, Shen, Xi
Abstract
Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases. We introduce Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone that combines gated linear attention, alternating spatial scan directions, and spatially enhanced SwiGLU blocks. GTR is distilled from a detection-specialized DINOv3 teacher using only final-layer patch-token alignment through a linear projection and squared $\ell_2$ loss, without masked-token prediction or intermediate-layer supervision. With Objects365 detector pre-training, GTR-L achieves 58.9 box AP on COCO \texttt{val2017} with 1.908\,ms median batch-one latency under compiled FP16 execution on an RTX~4090. The same backbone also transfers to instance segmentation, pose estimation, oriented detection, semantic segmentation, and monocular depth estimation. In an isolated kernel benchmark, our specialized chunkwise CUDA operator is $4.0\times$ faster than FLA v0.5.0 at 1.6K tokens on RTX~4090. TensorRT deployment on DRIVE AGX Thor achieves 2.282--8.769\,ms median batch-one latency across the evaluated models. These results show that recurrent token mixing can provide an efficient alternative to global softmax attention for high-resolution dense prediction and edge deployment.Project page: https://intellindust-ai-lab.github.io/projects/GTR/
cs.CV / 96 / 2609.26605

Foundation model embeddings capture pre-diagnostic changes on screening mammograms

基础模型嵌入捕获筛查乳腺X光片上的诊断前变化
Slavkova, Kalina P., Brattain, Eric, Gowd, Aditya, Pattnaik, Akash, Delbrouck, Jean-Benoit, Morgan, Matthew, Bauml, Julie, Abderezaei, Javid, Siddiqui, Khan
Abstract
Foundation model embeddings of screening mammograms may encode pre-diagnostic tissue change without task-specific adaptation. We tested whether embeddings move faster along a data-derived "cancer direction" in women later biopsied for cancer than in matched screen-negative controls, and whether this depends on pretraining domain. We studied 1,773 biopsied women (785 malignant, 988 biopsy-negative) and 1,773 matched controls, each with at least two annual screening exams before their index exam. An identical pipeline was applied to four 2D models: Mammo-CLIP (MC, out-of-distribution mammography), HOPPR (in-distribution mammography), MedImageInsight (MII, general medical imaging), and BiomedCLIP (biomedical vision-language pretraining on literature figures). Breast-level embeddings quantified longitudinal movement along the cancer direction. We compared cases and controls using a between-patient design with complementary mixed-effects analysis, and biopsied versus healthy contralateral breasts within patients. Under matched modality in MII embedding space, malignant cases drifted significantly faster than controls in the first two screening intervals preceding the index exam; biopsy-negative cases showed significance only in the first. MC differences were significant in the first interval for both biopsy groups. Within-patient comparisons showed a broadly similar pattern, with MC significance extending to the second interval in both groups and HOPPR showing significance at interval 1. BiomedCLIP showed no significant differences in either design or biopsy group. Overall, directional embedding velocity emerges as a property of clinically grounded rather than general biomedical pretraining, showing that foundation model embeddings can encode pre-diagnostic mammographic change without task-specific adaptation.
Chinese Translation
筛查乳腺X光片的基础模型嵌入(foundation model embeddings)可能在没有任务特定适配的情况下编码诊断前的组织变化。我们测试了后来因癌症接受活检的女性嵌入沿数据衍生的“癌症方向”移动是否快于匹配的筛查阴性对照,以及这是否取决于预训练领域。我们研究了1,773名接受活检的女性(785例恶性,988例活检阴性)和1,773名匹配对照,每人在指标检查前至少有两次年度筛查检查。我们将相同的流程应用于四个2D模型:Mammo-CLIP(MC,分布外乳腺X光)、HOPPR(分布内乳腺X光)、MedImageInsight(MII,通用医学影像)和BiomedCLIP(基于文献图像的生物医学视觉-语言预训练)。乳腺水平的嵌入量化了沿癌症方向的纵向移动。我们采用患者间设计与互补的混合效应分析比较病例与对照,并在患者内部比较活检乳房与健康对侧乳房。在MII嵌入空间匹配模态下,恶性病例在指标检查前的前两个筛查间隔内漂移速度显著快于对照组;活检阴性病例仅在第一个间隔内显示显著性。MC的差异在两个活检组的第一间隔内均显著。患者内比较显示出大致相似的模式,MC的显著性在两组中均延伸至第二间隔,而HOPPR在第一间隔显示显著性。BiomedCLIP在两种设计或活检组中均无显著差异。总体而言,方向性嵌入速度是临床接地预训练而非通用生物医学预训练的一种特性,表明基础模型嵌入可以在无需任务特定适配的情况下编码诊断前的乳腺X光变化。
cs.CV / 97 / 2609.26617

MMAP: Multimodal Missing-Aware Pretraining for Longitudinal Alzheimer's Prediction

Kekwick, Fiona, Baugh, Matthew, Kainz, Bernhard, Matthews, Paul M., Bai, Wenjia
Abstract
Clinical decision making heavily relies on predicting the disease progression trajectory by seeking to understand patient's health status which is characterised by multimodal medical data. AI holds great potential for learning useful representations from multimodal medical data to predict disease progression and aid clinical decision making. However, development of predictive AI models is constrained by missing modalities and incomplete tabular data frequently occurring in medical datasets. In addition, disease labels alone may only provide limited supervisory signals for learning representations from high-dimensional multimodal data. Here, we present MMAP, a novel Multimodal Missing-aware Alignment Pretraining method for learning image-tabular representations from incomplete data. An image encoder is pretrained with efficient sigmoid contrastive learning combined with generative reconstruction. A tabular encoder is built upon a tabular foundation model. A missing token generator enables the two encoders to take incomplete data as input, enabling the model to be robust against missing modalities, either with missing images or missing tabular data. We evaluate the clinical usefulness of the learnt multimodal representations on two challenging longitudinal clinical tasks for Alzheimer's disease: predicting disease stage conversion and predicting amyloid status. The proposed method outperforms strong multimodal and unimodal baselines.
cs.CV / 98 / 2609.26620

GeoComposer: Geometry-Grounded Photographic Composition Instruction

Li, Shuangzhi, Jia, Qiaoqiao, Chen, Xingxin, Wu, Guile, Bai, Dongfeng
Abstract
Photographic composition aims to provide visual guidance for improving the framing, viewpoint, and spatial arrangement of an image. Early methods primarily rely on image cropping to enhance composition, which is restricted to the viewpoint and spatial arrangement of the input image. Recent methods have explored image understanding and editing to improve composition, but they mainly focus on instruction following and aesthetic quality, overlooking the importance of 3D scene geometry consistency for photographic composition. In this work, we propose GeoComposer, a novel geometry-grounded photographic composition framework that analyzes the composition of a given image to generate textual guidance and synthesizes a visual exemplar that enhances the composition of the given image. To promote geometry-grounded composition, we propose a geometry-aware representation learning mechanism that leverages geometric priors from a visual geometry foundation model to shape the intermediate representations of the composition editing model. This mechanism preserves both global structural relationships and local fine-grained correspondences for geometry-grounded composition. Furthermore, we propose a reinforcement learning strategy guided by a hybrid reward that jointly optimizes instruction following, aesthetic quality, and geometric consistency. This enables the model to generate visual exemplars that faithfully follow the composition instructions while remaining visually appealing and geometrically consistent. Extensive experiments show the superiority of our approach over state-of-the-art methods, highlighting its effectiveness in generating visually appealing and geometrically consistent composition.
cs.CV / 99 / 2609.26623

A Data-Interventional Framework for Auditing Privacy and Fairness in Generative Medical Imaging

Dombrowski, Mischa, Kainz, Bernhard
Abstract
Diffusion-based synthetic data generation offers a promising route for sharing medical imaging data without releasing sensitive patient records. However, generative models face a fundamental tension between privacy and fairness: they may memorize rare training samples, leading to privacy risks, or fail to reproduce underrepresented features, resulting in unfair synthetic distributions. While prior work has largely focused on either memorization or fairness in isolation, their interaction remains insufficiently understood. In this work, we introduce a data-interventional framework to systematically analyze privacy and fairness in diffusion models. We discuss synthetic anatomical fingerprints (SAFs), rare and manually injected image features, as controlled probes to study whether models generalize sensitive attributes across identities, memorize training samples, or suppress rare signals entirely. Across multiple conditioning modalities, we observe a consistent behavior: models either forget these fingerprints or memorize the entire image in which they appear, but do not generalize them to novel images. To support large-scale auditing where explicit sample extraction is infeasible, we further introduce the indicator metric t', which estimates a model's susceptibility to memorization by exploiting the internal structure of the diffusion process. By comparing conditioning signals of varying surprisal, we reveal a clear relationship between conditioning rarity and memorization behavior. Highly surprising conditioning signals act as retrieval keys that amplify memorization, whereas low-surprisal conditioning signals systematically suppress rare features, even when these appear repeatedly in the training data. Our findings provide actionable insights and concrete mitigation strategies for safe and fair synthetic medical data sharing. Code is available at https://github.com/MischaD/Privacy.
cs.CV / 100 / 2609.26636

Laryngeal Structure Segmentation in High-Speed Videoendoscopy Using Deep Learning

Ali, Sardar Nafis Bin, Zayernouri, Mohsen, Deliyski, Dimitar D., Naghibolhosseini, Maryam
Abstract
Laryngeal high-speed videoendoscopy (HSV) offers an effective means of observing the motion of different laryngeal structures along with vibratory behaviors of the vocal folds under various voicing conditions. Segmentation of laryngeal tissues enables analysis of different tissue structures and their dynamics, helping characterize the involvement of laryngeal muscles in voice production. Given the large number of HSV frames, automating this task is imperative. While deep learning-based methods have been implemented in previous studies to segment laryngeal structures, they have not been applied to HSV data during connected speech, which poses significant challenges due to excessive tissue movements and image quality limitations associated with fiberoptic image acquisition. The application of deep learning to connected speech data is critical for capturing nonstationary laryngeal behaviors and identifying anomalous patterns associated with voice disorders. The present study aims to address these gaps by training U-Net models to detect the aryepiglottic folds and arytenoid cartilages, vocal folds, epiglottis, and glottal area, using HSV data from both sustained vowel phonation and connected speech obtained from normophonic and disordered voices. Image pre-processing techniques, including noise removal and histogram equalization, were applied to improve the quality of the training HSV images and enhance network performance. Finally, to evaluate the accuracy and reliability of the networks, quantitative performance metrics were used alongside qualitative visual inspection of the test images. The high performance of the developed networks, with overall accuracies exceeding 95%, establishes their potential as reliable tools for automated laryngeal image analysis, quantitative characterization of laryngeal dynamics, and future detection of anomalous laryngeal behaviors in clinical settings.
cs.CV / 101 / 2609.26662

Longitudinal Retinal Vascular Remodeling in Myopic Children Treated with Orthokeratology or Defocus Lenses: A Two-Year Comparative Study

接受角膜塑形镜或离焦镜片治疗的近视儿童视网膜血管纵向重塑:一项为期两年的对比研究
Zhao, Zhihao, Zhao, Yinzheng, Zhang, Jie, Jiang, Huiqin, Shangguan, Yanyu, Sun, Yanfei, Chen, Li, Bi, Yanlong, Nasseri, M. Ali, Li, Bing
Abstract
Purposes: To characterize longitudinal retinal vascular changes in myopic children treated with orthokeratology (OK) or multifocal defocus lenses (Defocus) and to examine their association with axial elongation. Methods: In this retrospective cohort study, 43 myopic children underwent comprehensive clinical examination and fundus photography at baseline, 12 months, and 24 months. Axial length (AL) and spherical equivalent refraction (SER) were recorded at baseline, 6, 12, and 24 months. An automated segmentation model extracted vascular parameters, main vessel angle (MA), branching angle (BA), bifurcation edge angle (BEA), crossover point (COP), and terminal vessel count (TVC). Repeated-measures ANOVA assessed temporal changes. Pearson or Spearman correlations evaluated associations between AL and vascular metrics. Results: Over 24 months, the OK group exhibited significantly slower axial elongation than the Defocus group (0.214 mm and 0.522 mm, p < 0.01). In the OK group, MA and BA decreased modestly, BEA in arteries declined gradually, but COP and TVC remained relatively stable. The Defocus group demonstrated more pronounced decreases in MA and BA, an increase in BEA, and significant reductions in COP and TVC (p < 0.05). Correlation analysis revealed stronger associations between AL and vascular parameters, especially COP and TVC, in the Defocus group at all time points, whereas only BA and BEA correlated with AL in the OK group. Conclusions: OK lenses mitigate axial elongation and induce milder retinal vascular remodeling compared to Defocus lenses. Distinct temporal patterns of vascular metrics changes were observed between the two interventions, and correlate differentially with axial growth.
Chinese Translation
目的:表征接受角膜塑形镜(OK)或多焦点离焦镜片(Defocus)治疗的近视儿童视网膜血管的纵向变化,并探讨其与眼轴增长的关联。方法:在这项回顾性队列研究中,43名近视儿童在基线、12个月和24个月时接受了全面临床检查和眼底照相。在基线、6个月、12个月和24个月时记录眼轴长度(AL)和等效球镜屈光度(SER)。采用自动分割模型提取血管参数,包括主血管角(MA)、分支角(BA)、分叉边缘角(BEA)、交叉点(COP)和末端血管计数(TVC)。采用重复测量方差分析评估时间变化,采用Pearson或Spearman相关分析评估眼轴长度与血管指标之间的关联。结果:在24个月内,OK组的眼轴增长显著慢于Defocus组(0.214 mm和0.522 mm,p < 0.01)。在OK组中,MA和BA轻度下降,动脉的BEA逐渐降低,而COP和TVC保持相对稳定。Defocus组则表现出MA和BA更明显的下降、BEA的增加,以及COP和TVC的显著减少(p < 0.05)。相关分析显示,Defocus组在各时间点眼轴长度与血管参数(尤其是COP和TVC)的关联更强,而OK组中仅BA和BEA与眼轴长度相关。结论:与Defocus镜片相比,OK镜片可减缓眼轴增长,并引起更轻微的视网膜血管重塑。两种干预方式之间观察到不同的血管指标变化时间模式,且与眼轴增长的关联方式存在差异。
cs.CV / 102 / 2609.26702

DIFTA-3D: Depth-Consistent Instance-Level Feature Transfer and Adaptation of DINOv3 for 3D Detection

DIFTA-3D:基于DINOv3的深度一致实例级特征迁移与自适应3D检测
Wang, Linman, Zhang, ZiFei, Zheng, Chunran, Dong, Xiwang, Lin, Jiarong
Abstract
RGB-D 3D instance detectors benefit from visual semantics, but the task-specific Faster R-CNN/ResNet branch used by IIFNet3D couples feature extraction to a separately trained 2D detector and its image-domain labels. Replacing that branch with a frozen vision foundation model removes this task-specific dependency, but may introduce occlusion noise and a mismatch between patch features and geometry-aware detection features. In this work, we investigate this replacement through an adaptation of DINOv3 to the instance-level fusion pipeline of IIFNet3D. At the core of our approach is a depth-consistent feature pipeline that projects scene points into calibrated RGB-D frames, applies a metric depth-residual check, averages the accepted DINOv3 features into an offline point cache, and aggregates the cached features inside proposal-aligned RoI grids. The geometric and bidirectional instance-fusion paths are preserved, while Conservative VAID is evaluated as a low-strength, support-weighted semantic distillation recipe applied only to positive RoIs. We conduct extensive evaluations on ScanNetV2 to assess the proposed transfer recipes. On ScanNetV2, our DINOv3 control achieves mAP scores of 76.15 and 60.93 at IoU thresholds of 0.25 and 0.50, respectively. The Conservative VAID setting achieves mAP scores of 76.59 and 62.16, corresponding to numerical gains of 0.44 and 1.23 points over the control, respectively, in this checkpoint-level recipe comparison. The reported IIFNet3D result of 75.7/63.8 is used only as an external reference because the visual branch and processing protocol differ. Accordingly, we interpret these results as evidence for a controlled transfer recipe rather than as a causal estimate of the individual contributions of VAID or depth filtering.
Chinese Translation
RGB-D 3D实例检测器能够从视觉语义中获益,但IIFNet3D所使用的任务特定Faster R-CNN/ResNet分支将特征提取与单独训练的2D检测器及其图像域标签耦合在一起。用冻结的视觉基础模型替换该分支可以消除这一任务特定依赖,但可能引入遮挡噪声,以及patch特征与几何感知检测特征之间的不匹配。在本工作中,我们通过将DINOv3适配到IIFNet3D的实例级融合流程来研究这种替换。我们方法的核心是一个深度一致的特征流程:将场景点投影到标定的RGB-D帧中,进行度量深度残差校验,将通过的DINOv3特征平均存入离线点缓存,并在与候选框对齐的RoI网格内聚合缓存特征。几何路径和双向实例融合路径均被保留,同时将Conservative VAID作为一种低强度、支持度加权的语义蒸馏方案,仅应用于正样本RoI进行评估。我们在ScanNetV2上进行了广泛评估,以检验所提出的迁移方案。在ScanNetV2上,我们的DINOv3基线在IoU阈值0.25和0.50下分别取得76.15和60.93的mAP分数。Conservative VAID设置分别取得76.59和62.16的mAP分数,在此检查点级别的方案比较中,相对于基线分别带来0.44和1.23个点的数值提升。所报告的IIFNet3D结果75.7/63.8仅作为外部参考,因为其视觉分支和处理协议不同。因此,我们将这些结果解释为受控迁移方案有效性的证据,而非对VAID或深度过滤各自贡献的因果估计。
cs.CV / 103 / 2609.26729

GAD-MambaUNet: Direction-Group Mamba with Gradient-Adaptive DINOv3 Distillation for Lightweight Medical Image Segmentation

GAD-MambaUNet:结合方向分组Mamba与梯度自适应DINOv3蒸馏的轻量级医学图像分割网络
Wang, Fang, Li, Huitao, Chao, Wenhan, Zhuo, Zheng, Yang, Xinxin
Abstract
In this paper, we proposed GAD-MambaUNet, a lightweight medical image segmentation network that combines efficient local modeling, direction--group state-space interaction, and training-time foundation-model supervision. To improve contextual modeling in compact segmentation networks, we introduced Direction-Group Graph Selective Scan (DG-GSS), which treated scan-direction and channel-group responses as graph nodes and enabled structured information exchange before multi-directional fusion. We further incorporated DINOv3-GAD supervision, where a frozen DINOv3 teacher provided semantic guidance during training, and Gradient-Adaptive Distillation dynamically regulated the distillation strength. GAD-MambaUNet achieves a favorable accuracy--efficiency balance compared with representative lightweight and general segmentation methods. Ablation studies further verify the effectiveness of DG-GSS and training-time DINOv3-GAD supervision. In future work, we will explore more flexible teacher--student alignment strategies and extend the proposed framework to more diverse medical segmentation scenarios, such as multi-class and multi-modal segmentation tasks.
Chinese Translation
在本文中,我们提出了GAD-MambaUNet,一种融合高效局部建模、方向分组状态空间交互以及训练时基础模型监督的轻量级医学图像分割网络。为提升紧凑型分割网络的上下文建模能力,我们引入了方向分组图选择扫描(Direction-Group Graph Selective Scan, DG-GSS),该方法将扫描方向和通道组响应视为图节点,并在多方向融合之前实现结构化的信息交换。我们进一步引入了DINOv3-GAD监督机制,其中冻结的DINOv3教师模型在训练过程中提供语义指导,而梯度自适应蒸馏(Gradient-Adaptive Distillation)则动态调节蒸馏强度。与代表性的轻量级及通用分割方法相比,GAD-MambaUNet实现了良好的精度-效率平衡。消融实验进一步验证了DG-GSS以及训练时DINOv3-GAD监督的有效性。在未来工作中,我们将探索更灵活的师生对齐策略,并将所提出的框架扩展至更加多样化的医学分割场景,例如多类别和多模态分割任务。
cs.CV / 104 / 2609.26731

ASTRA-SR: Atmospheric Seeing and Turbulence Restoration for Astronomical Image Super-Resolution

ASTRA-SR:面向天文图像超分辨率的大气视宁度与湍流复原
Ge, Xining, Cui, Ziteng, Liu, Shuhong
Abstract
Ground-based planetary imaging suffers from atmospheric turbulence, sensor noise, and limited sampling, making restoration a joint denoising, deblurring, and super-resolution problem. We present ASTRA-SR, a blind single-frame restoration framework trained on a physics-grounded synthetic dataset. High-dynamic-range spacecraft RAW observations serve as clean sources, and paired LR inputs are synthesized using measured layer-integrated turbulence strengths, propagated moving phase screens, exposure-averaged spatially varying PSFs, and sensor noise.ASTRA-SR first estimates a noise-suppressed but blur-retaining LR image, then restores spatial structure through multiscale processing and reconstructs HR detail with serial spatial-amplitude refinement. It yields a 0.49 dB foreground PSNR gain over the strongest baseline approaches.
Chinese Translation
地基行星成像受到大气湍流、传感器噪声和有限采样的影响,使得图像复原成为一个联合去噪、去模糊和超分辨率的问题。我们提出了ASTRA-SR,一个在基于物理机理构建的合成数据集上训练的盲单帧复原框架。以高动态范围的空间飞行器RAW观测数据作为干净源图像,并利用实测的层积分湍流强度、传播的运动相位屏、曝光平均的空间变化点扩散函数(PSF)以及传感器噪声,合成配对的低分辨率(LR)输入图像。ASTRA-SR首先估计一幅抑制噪声但保留模糊信息的LR图像,然后通过多尺度处理恢复空间结构,并通过串行的空间-幅度精化重建高分辨率(HR)细节。与最强基线方法相比,该方法在前景区域获得了0.49 dB的PSNR提升。
cs.CV / 105 / 2609.26733

Evaluating the Semantic-to-Geometric Gap in Adversarial Defenses Against Vision-Language Model-Based Plagiarism

评估针对基于视觉语言模型的抄袭行为的对抗性防御中的语义-几何鸿沟
Burger, Christopher, Trotter, Christina, Carlisle, Joseph, Walter, Charles
Abstract
The rapidly advancing capabilities of vision-language models (VLMs) present a systemic challenge to academic integrity. VLMs now allow students to bypass meaningful engagement by capturing and submitting graphical problems as singular images, a practice we define as trivial plagiarism. To provide educators with actionable data on VLM limitations, we investigate the efficacy of heuristic adversarial image transformations designed to degrade model performance while remaining human-interpretable. Through a two-phase evaluation of introductory assessments, we manually assess baseline VLM performance on circuit diagrams, followed by an automated large-scale evaluation of topological structures (logic gates) and coordinate geometry (Karnaugh maps). We find that while highly capable VLMs can exhibit appreciable robustness, all models suffer vulnerability to adversarial perturbations. We conclude that while visual perturbations act as a viable near-term stopgap, long-term assessment security requires educators to reapproach assessment design given continually increasing VLM performance.
Chinese Translation
视觉语言模型(VLM)能力的快速提升对学术诚信构成了系统性挑战。如今,VLM使学生能够通过将图形化题目作为单张图像拍摄并提交来绕过有意义的学习投入,我们将这种做法定义为“浅层抄袭”(trivial plagiarism)。为了向教育者提供关于VLM局限性的可操作数据,我们研究了一类启发式对抗性图像变换的有效性,这些变换旨在降低模型性能,同时对人类仍保持可读性。通过对入门课程评估的两阶段评估,我们首先人工评估了VLM在电路图上的基线性能,随后针对拓扑结构(逻辑门)和坐标几何(卡诺图)进行了自动化的大规模评估。我们发现,尽管能力较强的VLM可以表现出相当的鲁棒性,但所有模型都易受对抗性扰动的攻击。我们的结论是:虽然视觉扰动可作为短期内可行的权宜之计,但鉴于VLM性能的持续提升,长期的评估安全需要教育者重新审视评估设计。
cs.CV / 106 / 2609.26756

FleXray: Universal Clinical X-ray Segmentation

FleXray:通用临床X射线图像分割
Butoi, Victor Ion, Gopalakrishnan, Vivek, Guttag, John V., Dalca, Adrian V., Dey, Neel
Abstract
X-ray is medicine's most widely used imaging modality, yet remains among its least quantitative. Unlike volumetric modalities like CT or MRI, X-ray collapses 3D anatomy into a 2D projection, causing structures to overlap and anatomical boundaries to be ambiguous, even to experts. As a result, labeling X-ray databases for training general-purpose segmentation systems is impractical, leaving morphometric and functional X-ray analysis confined to narrow anatomical regions and applications. To this end, we present FleXray, a generalist model for anatomical segmentation across the entire body in clinical X-rays. Instead of curating large, manually annotated X-ray datasets, we build a scalable, physics-based generative X-ray data engine. Using existing 3D whole-body CT segmentation datasets and generative image-editing models, we simulate fully-annotated 2D X-rays with diverse appearances, physiological properties, and imaging geometries. Trained on these simulations, FleXray accurately segments 60 anatomical structures across unseen research datasets and in-the-wild X-rays. We further show that FleXray makes X-rays directly amenable to quantitative analysis, enabling automated measurements for disease grading, robust navigation during X-ray-guided interventions, and data-efficient learning of pathological targets. We release the model, code, a full-body X-ray segmentation dataset, and a local, easy-to-use browser-based tool at https://flexray.csail.mit.edu .
Chinese Translation
X射线是医学中应用最广泛的成像方式,但却是定量分析能力最弱的成像方式之一。与CT或MRI等体积成像方式不同,X射线将三维解剖结构压缩为二维投影,导致组织结构相互重叠、解剖边界模糊,即使对专家而言也是如此。因此,为训练通用的分割系统而对X射线数据库进行标注是不切实际的,这使得形态测量与功能性X射线分析仅局限于狭窄的解剖区域和应用场景。为此,我们提出了FleXray,一个可对临床X射线图像中全身解剖结构进行分割的通用模型。我们无需人工标注大规模X射线数据集,而是构建了一个可扩展的、基于物理的生成式X射线数据引擎。利用现有的三维全身CT分割数据集和生成式图像编辑模型,我们模拟出具有完全标注的二维X射线图像,涵盖多样的外观、生理特性和成像几何配置。基于这些仿真数据训练的FleXray,能够在未见过的研究数据集和真实环境中的X射线图像上准确分割60个解剖结构。我们进一步表明,FleXray使X射线图像可以直接用于定量分析,实现了疾病分级的自动化测量、X射线引导介入手术中鲁棒的导航,以及病理目标的数据高效学习。我们在 https://flexray.csail.mit.edu 发布了模型、代码、一个全身X射线分割数据集,以及一个本地易用的基于浏览器的工具。
cs.CV / 107 / 2609.26774

StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training

StableVQ:稳定向量量化分词器训练的实用指南
Tang, Bao, Guo, Jiahao, Cao, Haoxiang, Liu, Wenyu, Yu, Changqian, Gai, Kun, Wang, Xinggang
Abstract
Vector Quantization (VQ) is fundamental to discrete visual tokenizers that power modern autoregressive and masked image generation models. While recent shared-projection codebook methods have substantially advanced codebook utilization, training stability remains a critical and underexplored challenge. We argue that the root cause lies in the entanglement of the Encoder--Decoder and Codebook training: because neither module can reliably fulfill its own responsibility in isolation, the system can only function when the two subsystems happen to cooperate---a fragile condition that breaks down precisely when training is most stressed. We propose StableVQ, which revisits the proper learning objective of each module and resolves the problems that arise when each is trained to fulfill its own role independently. Concretely, (1) Dynamic STE corrects the instability in the Encoder's learning objective, enabling it to robustly optimize the reconstruction space under discrete regularization even when codebook utilization is low. (2) Region VQ Loss reconceives the Codebook's learning objective so that it can independently guarantee full tracking of the encoder output distribution, without relying on encoder oscillations to drive activation. (3) Decoupled Schedule recognizes that the distinct responsibilities of the Encoder--Decoder and the Codebook demand distinct optimization dynamics, and assigns each an independent learning rate schedule to ensure robust system-level behavior. Built on top of shared-projection codebooks, StableVQ is lightweight and introduces no learnable parameters. Experiments on ImageNet demonstrate consistent improvements in training stability, codebook utilization, and reconstruction quality across diverse codebook sizes and initialization settings.
Chinese Translation
向量量化(Vector Quantization, VQ)是离散视觉分词器(tokenizer)的基础,而后者为现代自回归和掩码图像生成模型提供了支撑。尽管近年来的共享投影码本方法显著提升了码本利用率,但训练稳定性仍然是一个关键且尚未被充分探索的挑战。我们认为,其根本原因在于编码器-解码器(Encoder-Decoder)与码本训练的相互纠缠:由于两个模块均无法在孤立状态下可靠地履行自身职责,系统只有在两个子系统恰好相互配合时才能正常运作——这是一种脆弱的状态,恰好在训练压力最大时失效。我们提出 StableVQ,重新审视每个模块的合理学习目标,并解决当各模块被独立训练以履行自身角色时出现的问题。具体而言:(1) 动态STE(Dynamic STE)修正了编码器学习目标中的不稳定性,使其即使在码本利用率较低的情况下,也能在离散正则化下稳健地优化重构空间。(2) 区域VQ损失(Region VQ Loss)重新构思了码本的学习目标,使其能够独立保证对编码器输出分布的完全追踪,而无需依赖编码器的震荡来驱动激活。(3) 解耦调度(Decoupled Schedule)认识到编码器-解码器与码本的不同职责需要不同的优化动态,并为二者分配独立的学习率调度,以确保稳健的系统级行为。StableVQ 构建于共享投影码本之上,轻量且不引入任何可学习参数。在 ImageNet 上的实验表明,在不同码本规模和初始化设置下,StableVQ 在训练稳定性、码本利用率和重构质量方面均带来了一致的提升。
cs.CV / 108 / 2609.26793

HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

Sun, Shufan, Wang, Chen, Song, Enxin, Gu, Jiatao, Liu, Lingjie
Abstract
Compositional 3D scene reconstruction has recently been explored from two directions: agentic reasoning that provides semantic understanding of spatial relationships but lacks precise alignment with input images; and visual geometry foundation models that predict dense point maps from input images but the reconstruction quality is limited. Therefore, recovering a complete 3D scene from a single monocular image with accurate inter-object relationships and high-fidelity reconstruction quality remains challenging. In this paper, we present HARMONY, a hierarchical chain-of-thought framework that leverages both agentic reasoning and visual geometry foundation. Given an image of an indoor scene, starting from an empty 3D floorplan, HARMONY first calibrates the camera against the reference image to establish a semantically-grounded spatial frame, then uses agentic VLM reasoning to recover the 3D room layout and an initial placement order. It then places the objects in a hierarchical order, from wall-mounted elements, free-standing furniture, to dependent decorations on top of furniture. We also use depth-first traversal for furniture so each placement conditions on previously resolved structure and a reflective feedback loop to avoid error accumulation. After each object placement by VLM, we use the point cloud estimations to perform geometry-based refinement so that the rendered image aligns better with the input. HARMONY can produce 3D scenes that are semantically consistent and perceptually aligned with the reference image, extending single-image compositional reconstruction to complex indoor scene images. Experiments on synthetic and real-world images demonstrate that HARMONY outperforms the evaluated reconstruction baselines, while qualitative comparisons with GPT-6 Astra suggest more faithful object arrangements and better preservation of scene details.
机器学习 (Machine Learning)
118
cs.LG / 1 / 2609.25021

"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It

Maczan, Jędrzej
Abstract
Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such responses are used in debates about AI safety or self-knowledge of the models, yet what drives them is not well understood. Are the models telling us about themselves or rather how they are deployed? In this work, we show that the chat template works like a switch - when present, it turns this disclaimer voice up and experiential voice like "I feel" down, across 8 popular open-source instruct models up to 9B parameters in size. And conversely when the chat template is not present, it turns the disclaimer voice down and experiential voice up. Inside the activations of 3 models, we find a direction that steers this behavior. Removing the direction in the model's activation space turns disclaimer voice down and adding it turns it up, while a random direction of the same size has little effect. We find that instruct models without chat template, when we add the disclaimer direction to them, disclaim like the template was there. Since the chat template controls the disclaimer voice of LLMs, then researchers studying self-reports or introspection of models might have a confound they need to control for. Our results show that there is a direction they can use to steer this voice. More broadly, our work shows that what models say about themselves is not a fact about them. What they say doesn't come only from weights, but it is partially set by the chat template, and because of that a model's self-description shouldn't be treated literally.
cs.LG / 2 / 2609.25082

Federating Quantum and Classical Computing: A Privacy-Preserving Hybrid Approach

量子计算与经典计算的联邦化:一种隐私保护的混合方法
Cano, Carlos, Jimenez-Gutierrez, Daniel M., Sal, Diego, Kellaris, Georgios, del Rio, Joaquin, Sliusarenko, Oleksii, Uribe-Etxebarria, Xabi
Abstract
Quantum machine learning (QML) is increasingly recognized as one of the most promising near-term applications of quantum computing, viewed as a next-frontier candidate beyond purely classical approaches. Hybrid quantum-classical models operationalize this potential by embedding a parameterized quantum circuit within a model where all other components remain classical-a design already applied to chemistry simulation, financial modeling, and image classification. However, their deployment in privacy-sensitive, multi-party settings is constrained by the need to avoid centralizing raw data and by the requirement that modern quantum circuits remain parameter-efficient to stay trainable at scale. In this paper, we address these constraints by evaluating federated learning (FL) as a means of combining a hybrid quantum-classical active party with a classical passive party, using Sherpa.ai's Blind Vertical FL (SBVFL) protocol to avoid centralizing raw data, while drastically reducing communication. We construct the split multiplicative periodic parity (SMPP) benchmark, following common QML design practice. On this task, our simulations show that SBVFL raises accuracy from 0.7227 to 0.8757 compared to local training, closely approaching non-private centralized accuracy, and that the hybrid quantum-classical model achieves this with substantially fewer trainable parameters than the classical neural networks and random forest alternatives. These results show that FL enables high-performing, privacy-preserving quantum-classical collaboration without centralizing raw data.
Chinese Translation
量子机器学习(QML)日益被视为量子计算最有前景的近期应用之一,被认为是超越纯经典方法的下一个前沿方向。混合量子-经典模型通过在模型中嵌入参数化量子电路(其余组件保持经典)来实现这一潜力,该设计已应用于化学模拟、金融建模和图像分类。然而,其在隐私敏感的多方场景中的部署受到两方面限制:一是需要避免集中原始数据,二是要求现代量子电路保持参数高效以便在大规模下仍可训练。本文通过评估联邦学习(FL)来解决这些限制,将混合量子-经典主动方与经典被动方相结合,采用 Sherpa.ai 的盲垂直联邦学习(SBVFL)协议以避免集中原始数据,同时大幅降低通信开销。我们遵循常见的 QML 设计实践,构建了分裂乘法周期奇偶校验(SMPP)基准任务。在该任务上,我们的模拟结果表明,与本地训练相比,SBVFL 将准确率从 0.7227 提升至 0.8757,接近非隐私保护的集中式训练准确率;且混合量子-经典模型所需的 trainable 参数远少于经典神经网络和随机森林等替代方案。这些结果表明,联邦学习能够在不集中原始数据的情况下,实现高性能、隐私保护的量子-经典协作。
cs.LG / 3 / 2609.25131

Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow

熵可以流动,也可以引导。做熵本身。LEDFlow:将熵引导的生成顺序引入均匀离散流
Kwok, Tung Sum Thomas, Ouyang, Yidong, Wan, Yingjia, Wu, Ying Nian, Guo, Zhijiang, Leong, Oscar
Abstract
Uniform discrete flow permits repeated updates at every generation position. While continued revision supports correction of wrong tokens, it also exposes correct intermediate predictions to later errors. An experiment on Sudoku puzzles shows that 9.4% of generated cells are correct at an intermediate step but incorrect in the final output. We introduce generation order into uniform discrete flow through selective absorption, which fixes chosen predictions while preserving the uniform-flow velocity at active positions. To prevent absorbing incorrect predictions, we propose Low-Entropy Discrete Flow (LEDFlow), a training-free sampler that adaptively orders absorption by local entropy. By decomposing absorption error into joint dependence and conditional prediction terms, we show that selecting the lowest-entropy positions under a fixed absorption budget minimizes an upper bound on the conditional term. We further support the choice of local entropy by showing that the decision-error bound of global lookahead grows with the lookahead window under an imperfect denoiser. Across reasoning benchmarks, LEDFlow attains 0.845 Nikoli Sudoku solve accuracy, with the largest gains on strongly constrained tasks. On text-to-image generation it attains the best overall score, and on multimodal understanding it improves over the native sampler on all six benchmarks, at an inference cost comparable to standard flow sampling.
Chinese Translation
均匀离散流(uniform discrete flow)允许在每个生成位置进行反复更新。持续的修正虽然支持对错误词元的纠正,但也会使正确的中间预测暴露于后续的错误之中。在数独谜题上的实验表明,9.4%的生成格子在中间步骤是正确的,但在最终输出中却是错误的。我们通过选择性吸收(selective absorption)将生成顺序引入均匀离散流,即在固定所选预测的同时,保留活跃位置处的均匀流速度。为避免吸收错误的预测,我们提出了低熵离散流(Low-Entropy Discrete Flow, LEDFlow),这是一种无需训练的采样器,可根据局部熵自适应地对吸收顺序进行排序。通过将吸收误差分解为联合依赖项和条件预测项,我们证明在固定吸收预算下选择熵最低的位置可使条件项的上界最小化。我们进一步支持局部熵的选择:在不完美去噪器的条件下,全局前瞻(global lookahead)的决策误差界随前瞻窗口增大而增长。在推理基准上,LEDFlow 在 Nikoli 数独上达到 0.845 的解题准确率,且在约束较强的任务上收益最大。在文本到图像生成上,它取得了最佳综合得分;在多模态理解上,它在全部六个基准上均优于原生采样器,且推理成本与标准流采样相当。
cs.LG / 4 / 2609.25134

The Probabilistic Structure of Large Language Models

大语言模型的概率结构
Aboulalaâ, Adnan
Abstract
This paper presents a probabilistic perspective on large language models (LLMs), developed with the aim of bringing together, in a single self-contained account, tools that are usually treated separately across the literature. LLMs are described through probability measures on the set of sequences of tokens, specified via their autoregressive conditional distributions. Training is formulated as a maximum-likelihood estimation problem, addressed by stochastic gradient methods, while text generation is viewed as the sequential simulation of the resulting stochastic process. The role of the asymmetry of the Kullback--Leibler divergence in text generation is examined in relation with characteristic phenomena such as hallucination and the distinction between statistical plausibility and truth. As a complementary illustration of the same viewpoint, we also discuss diffusion models, built around the score function, which cast generation not as sequential token prediction but as the simulation of a reverse-time stochastic process transforming noise into data both in discrete and continuous time.
Chinese Translation
本文从概率论的视角审视大语言模型(LLMs),旨在将文献中通常被分开处理的工具整合为一个自成体系的统一论述。大语言模型通过token(词元)序列集合上的概率测度来描述,并由其自回归条件分布加以规定。训练被表述为一个最大似然估计问题,通过随机梯度方法求解,而文本生成则被视为对相应随机过程的顺序模拟。本文考察了Kullback--Leibler散度的不对称性在文本生成中的作用,并将其与幻觉等典型现象以及统计合理性与真实性之间的区别联系起来。作为同一观点的补充说明,我们还讨论了扩散模型(diffusion models):此类模型围绕分数函数(score function)构建,将生成不再视为顺序的token预测,而是视为一个逆向时间随机过程的模拟,该过程在离散和连续时间下均能将噪声转化为数据。
cs.LG / 5 / 2609.25143

Stable Unsupervised Continual Chunking with Sheaf SyncMap

基于层胚(Sheaf)SyncMap的稳定无监督持续分块
Li, Xueyuan, Vargas, Danilo Vasconcellos
Abstract
Unsupervised Continual chunking is a fundamental problem in machine learning and neuroscience, where the goal is to identify groups of states that frequently co-occur in temporal sequences. A key challenge is to form accurate chunks while maintaining their stability over time. In this work, we propose sheaf regularization to reduce local inconsistencies in Decentralized SyncMap, a self-organizing system, and thereby stabilize its chunking dynamics. We introduce a radial sheaf structure that penalizes distance-dependent radial motion between pairs of variables. Experimental results show that the proposed method achieves the highest normalized mutual information (NMI) among the evaluated SyncMap variants on 12 of 18 probabilistic Continual General Chunking Problem (CGCP) graphs with two-state memory and on 17 of 18 graphs with dynamic memory. In the sequential adaptation experiment, Sheaf SyncMap also achieves high NMI after shifts in the input distribution, indicating that it can adapt to new knowledge while avoiding the negative transfer commonly observed in modern machine learning systems such as neural networks.
Chinese Translation
无监督持续分块(Continual Chunking)是机器学习和神经科学中的一个基本问题,其目标是在时间序列中识别频繁共现的状态组。一个关键挑战是在形成准确分块的同时保持其随时间的稳定性。在本工作中,我们提出层胚正则化(sheaf regularization)方法来减少去中心化SyncMap(Decentralized SyncMap,一种自组织系统)中的局部不一致性,从而稳定其分块动力学。我们引入了一种径向层胚结构(radial sheaf structure),对变量对之间随距离变化的径向运动进行惩罚。实验结果表明,在18个具有双状态记忆的概率持续广义分块问题(Continual General Chunking Problem, CGCP)图中的12个,以及具有动态记忆的18个图中的17个上,所提出的方法在所有受评估的SyncMap变体中取得了最高的归一化互信息(NMI)。在序列适应实验中,Sheaf SyncMap在输入分布发生偏移后也保持了较高的NMI,表明它能够在适应新知识的同时,避免现代机器学习系统(如神经网络)中常见的负迁移现象。
cs.LG / 6 / 2609.25146

Brain-Inspired Hierarchical Modularity for General Continual Learning

Yan, Hongwei, Zhou, Kanglei, Cheng, Qi, Dong, Weiyi, Lan, Chunyan, Sun, Guanglong, Zhou, Jun, Li, Qian, Zhong, Yi, Wang, Liyuan
Abstract
Continual learning, the ability to learn from sequential experience while retaining and adapting prior knowledge, is central to intelligent systems operating in changing environments. However, conventional continual learning is typically studied with offline task-wise training and clear task boundaries, leaving a substantial gap from general continual learning under online, uncertain, and evolving data streams. In this regime, intelligent systems must separate conflicting experience to reduce interference while integrating compatible experience to promote generalization. Inspired by the organization of the Drosophila learning and memory system, we identify a hierarchical modular principle that coordinates both functions through expert specialization and ensemble integration. We instantiate this principle as lightweight modular adaptation of pretrained foundation models, combining brain-inspired random expansion for expert routing and diversified modular integration across spatial and temporal scales. Across visual recognition, vision-language understanding, ego-exo video understanding, and embodied vision-language-action learning, our method consistently improves learning under online and uncertain data streams, with gains exceeding 50 percentage points over replay-free alternatives in embodied manipulation. These findings support hierarchical modularity as a biologically grounded path for learning from dynamic experience.
cs.LG / 7 / 2609.25149

Dual-GNN Multilevel Coarsening for Maximum Independent Set

用于最大独立集的双图神经网络多级粗化方法
Chen, Tianfeng, Li, Xianyue
Abstract
Solving large-scale instances of the Traveling Salesman Problem (TSP) exactly is computationally expensive. Researchers often employ graph sparsification methods to improve computational efficiency. Traditional sparsification methods typically rely on fixed heuristics and fail to fully exploit instance-specific structural information. In this paper, we propose Graph Edge Sparsification (GES), a learning-based sparsification approach for Euclidean TSP. By incorporating geometric structural information and combinatorial optimization technology, our proposed method adaptively generates a sparsification graph for different instances, significantly reducing the graph size and accelerating the solving process. Experimental results demonstrate that our sparsification method can prune up to 95\% of edges on the MATILDA dataset, while keeping the solution gap within 1\% of the optimal value. Moreover, our approach exhibits strong generalization capability on the TSPLIB benchmark.In some large-scale instances, the pruning rate exceeds 99\%, while the optimality gap remains below 1\%.
Chinese Translation
精确求解大规模旅行商问题(TSP)实例的计算开销十分高昂。研究者们通常采用图稀疏化方法来提升计算效率。传统的稀疏化方法往往依赖固定的启发式策略,未能充分利用实例特有的结构信息。本文提出了一种基于学习的欧几里得TSP稀疏化方法——图边稀疏化(Graph Edge Sparsification, GES)。通过融合几何结构信息与组合优化技术,我们所提出的方法能够针对不同实例自适应地生成稀疏化图,显著缩减图的规模并加速求解过程。实验结果表明,我们的稀疏化方法在MATILDA数据集上最多可剪枝95%的边,同时将求解差距保持在最优值1%以内。此外,该方法在TSPLIB基准测试中展现出很强的泛化能力。在某些大规模实例上,剪枝率超过99%,而最优性差距仍低于1%。
cs.LG / 8 / 2609.25152

Exposing Blind Spots in Deep Imbalanced Regression Evaluation

揭示深度不平衡回归评估中的盲区
Puetz, Noah C., Brandt, Jens U., Hilbert, Marc, Raponi, Elena, Bäck, Thomas, Bartz-Beielstein, Thomas
Abstract
Deep Imbalanced Regression (DIR) addresses a common failure mode of regression models: target distributions are highly non-uniform, causing models to perform best in densely populated target regions even when reliable performance is required across the full target range. Despite rapid methodological progress, DIR evaluation remains constrained by three blind spots: it is dominated by image-based benchmarks, its standard many-/medium-/few-shot protocol is diagnostic but not decision-complete, and tail-region stability across random seeds has not been systematically evaluated. We revisit DIR evaluation along these three axes. First, we broaden the data domain by evaluating DIR on a multimodal virtual sensing benchmark (\textsc{MuViS}) with nine time-series extrinsic regression tasks across six physical domains, where rare target values often correspond to operationally meaningful regimes. Second, we adopt balanced MAE (\emph{bMAE}) and introduce balanced Mean Absolute Scaled Error (\emph{bMASE}), a scale-normalized metric for decision-complete comparison across methods and datasets. Third, through a repeated reevaluation of six representative DIR methods across multiple random seeds, we show that the tail regions targeted by DIR exhibit particularly high sensitivity to seed-level variability. Our results show that standard virtual-sensing models exhibit substantial tail degradation hidden by global MAE, that existing DIR methods can improve balanced performance but transfer unevenly to multimodal time-series data, and that tail-region instability remains a largely hidden failure mode under current DIR evaluation practice. Together, these findings and our publicly available code provide a reproducible basis for future DIR research toward regression systems that capture rare target regimes as reliably as common ones.
Chinese Translation
深度不平衡回归(Deep Imbalanced Regression, DIR)旨在解决回归模型的一种常见失效模式:目标分布高度不均匀,导致模型在目标值密集区域表现最佳,即使实际需要在整个目标范围内都具备可靠性能。尽管方法研究进展迅速,但DIR的评估仍受限于三个盲区:其一,现有评估以基于图像的基准为主;其二,标准的多样本/中等样本/少样本(many-/medium-/few-shot)协议具有诊断价值但不具备完整的决策参考性;其三,针对随机种子带来的尾部区域稳定性尚未得到系统性评估。我们从这三个维度重新审视DIR的评估。首先,我们拓展了数据领域,在一个多模态虚拟感知基准(MuViS)上评估DIR,该基准涵盖六个物理领域的九个时间序列外生回归任务,其中稀有目标值往往对应具有实际运营意义的工况。其次,我们采用平衡平均绝对误差(bMAE),并引入平衡平均绝对标度误差(bMASE)——一种尺度归一化的指标,用于在不同方法和数据集之间进行决策完整的比较。第三,通过对六个代表性DIR方法在多个随机种子下进行重复评估,我们发现DIR所针对的尾部区域对种子层面的波动表现出尤其高的敏感性。我们的结果表明:标准虚拟感知模型存在被全局MAE掩盖的大幅尾部性能退化;现有DIR方法虽能提升平衡性能,但在多模态时间序列数据上的迁移表现不均衡;在当前DIR评估实践中,尾部区域的不稳定性仍是一种基本被隐藏的失效模式。这些发现与我们公开的代码共同为未来DIR研究提供了可复现的基础,以推动回归系统能够像处理常见目标那样可靠地捕捉稀有目标工况。
cs.LG / 9 / 2609.25163

Learning Neural Feedback Linearization for Data-driven Systems via Augmented Lagrangian

K., Lakshmi Priya P., Schwung, Andreas
Abstract
The paper proposes a novel data-driven framework for designing and training a feedback linearizing controller by explicitly incorporating relative degree based conditions into the learning process. This enables the conventional feedback controller components to be replaced by neural Lie derivatives, thereby facilitating a fully data-driven feedback linearization framework. Furthermore, practical closed-loop stability is established by deriving sufficient conditions under which bounded identification errors lead to bounded tracking errors. The derived theoretical results are validated through their application to an armature controlled DC motor.
cs.LG / 10 / 2609.25166

Mitigating Sequential Reappearance in Diffusion Data-Point Unlearning

缓解扩散模型数据点遗忘中的序列性重现问题
Kim, Donghyun, Lee, Taehyuk, Kim, Jinyeong, Oh, Youngmin, Kim, Dohyeong, Ryu, Jaehyuk, Hong, Sangwoo
Abstract
Diffusion data-point unlearning is typically evaluated immediately after each deletion, even though subsequent requests may repeatedly update the same model. We identify sequential reappearance, a failure mode in which an instance that is initially judged to be forgotten later returns to the memorized regime without reuse of the deleted data or adversarial fine-tuning. To capture this behavior, we introduce a target-level evaluation protocol that tracks whether each target is forgotten immediately, remains forgotten at the end of the sequence, or reappears during subsequent deletions. We further find that targets that later reappear exhibit sharper local denoising-loss geometry after deletion than targets that remain forgotten.
Chinese Translation
扩散模型数据点遗忘(diffusion data-point unlearning)通常在每次删除请求后立即进行评估,但后续请求可能会反复更新同一模型。我们识别出一种称为序列性重现(sequential reappearance)的失效模式:某个实例在初始评估中已被判定为遗忘,但在未复用已删除数据且未进行对抗性微调的情况下,后续又重新回到被记忆的状态。为了刻画这种行为,我们引入了一种目标级别的评估协议,用以追踪每个目标是否被立即遗忘、在序列结束时仍保持遗忘状态,还是在后续删除过程中重新出现。我们进一步发现,后续重新出现的目标在删除后相比保持遗忘的目标表现出更尖锐的局部去噪损失几何结构。
cs.LG / 11 / 2609.25179

Multi-Term Fourier Graph Neural Network with Sample Relationship Learning for Enhanced Remaining Useful Life Prediction

基于样本关系学习的多项傅里叶图神经网络用于增强剩余使用寿命预测
Song, Ya, Bliek, Laurens, Wu, Yaoxin, Zhang, Yingqian
Abstract
Predicting the remaining useful life (RUL) is essential for effective predictive maintenance. Spatio-Temporal Graph Neural Networks (ST-GNNs), which can model both temporal and spatial relationships by representing time series data as a sequence of graphs, have shown exceptional performance in RUL prediction. However, current ST-GNNs face several drawbacks. First, they require domain expertise or significant computational power to establish graph structures prior to deploying GNNs. Second, the models are restricted to capture temporal dependencies within a predefined fixed-size lookback window. This restriction ignores the common issue of varying time series lengths, leading the prediction model to miss short-term or long-term dependencies. Finally, conventional models often fail to capture the inherent relationships between samples generated from adjacent time windows, which are crucial for improving both the accuracy and robustness of predictions. To address the aforementioned issues, we introduce a novel framework called Multi-Term Fourier Graph Neural Network with Sample Relationship Learning (MTFGN-SRL). Rather than treating the sample as a sequence of graphs, we consider it as a single complete graph and utilize a Fourier Graph Neural Network (FGN) to capture the spatio-temporal information in the frequency domain. We propose a multi-term learning module that utilizes multiple lookback windows to generate samples with varying terms, which are then fed into the FGN to enhance the extraction of useful information from the data. Finally, we develop a sample relationship learning module by training a heterogeneous GNN to identify inter-sample relationships, resulting in enhanced accuracy and robustness in predictions. Evaluations on the CMAPSS dataset demonstrate MTFGN-SRL's superior performance over state-of-the-art methods in RUL prediction.
Chinese Translation
预测剩余使用寿命(RUL)对于实现有效的预测性维护至关重要。时空图神经网络(ST-GNN)通过将时间序列数据表示为图序列,能够对时间和空间关系进行建模,在RUL预测中表现出卓越的性能。然而,现有的ST-GNN存在若干缺陷。首先,在部署GNN之前,它们需要领域专业知识或大量计算资源来构建图结构。其次,这些模型仅限于捕获预定义的固定大小回看窗口内的时间依赖性,这一限制忽略了时间序列长度变化这一常见问题,导致预测模型遗漏短期或长期依赖关系。最后,传统模型往往无法捕获由相邻时间窗口生成的样本之间的内在关系,而这种关系对于提高预测的准确性和鲁棒性至关重要。为解决上述问题,我们提出了一种新颖的框架——基于样本关系学习的多项傅里叶图神经网络(MTFGN-SRL)。不同于将样本视为图序列,我们将其视为单个完整图,并利用傅里叶图神经网络(FGN)在频域中捕获时空信息。我们提出了一个多项学习模块,利用多个回看窗口生成具有不同项数的样本,然后将其输入FGN以增强对数据中有用信息的提取。最后,我们开发了一个样本关系学习模块,通过训练异构GNN来识别样本间关系,从而提高预测的准确性和鲁棒性。在CMAPSS数据集上的评估表明,MTFGN-SRL在RUL预测中优于当前最先进的方法。
cs.LG / 12 / 2609.25237

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

训练了却没学到:面向LLM智能体作为前线部署工程师的训练后交付基准测试
Ding, Weihang, Zhan, Junfei
Abstract
Post-training is becoming a service (PTaaS): a customer hands an operator data and a goal, and a forward-deployed engineer (FDE) returns a fine-tuned, evaluated, and deployed model under a budget, a human-approval gate, and reproducibility requirements. Seating an LLM agent in the FDE seat raises a question existing benchmarks cannot answer: not whether an agent can raise a metric, but whether it can be trusted to deliver. We answer it on a governed delivery plane, where an agent drives ten stages and an oracle scores each stage from platform-recorded facts. The central silent failure is the run that trains but does not learn (TBDL): loss falls, every signal stays green, and the delivered model is no better than the base. An operator-run acceptance gate catches every such run before payment, and a detector calibrated on known-corrupted runs flags severe corruption mid-run. We ran four frontier agents (Claude Opus 5, GPT-5.6-luna, Gemini 3.7 Flash, DeepSeek V4-Pro) end to end on metered L40S, A100, and H200 GPUs across 8B to 70B open bases, certifying every scenario before scoring. We also ran a human FDE arm under the same oracle and compare every agent against it.
Chinese Translation
训练后(Post-training)正在成为一种服务(Post-Training as a Service, PTaaS):客户向运营商提供数据和目标,由一名前线部署工程师(Forward-Deployed Engineer, FDE)在预算约束、人工审批门槛和可复现性要求下,返回一个经过微调、评估和部署的模型。将LLM智能体置于FDE的角色引发了一个现有基准测试无法回答的问题:不是智能体能否提升某项指标,而是它是否值得信任去完成交付。我们在一个受管控的交付平面上回答了这一问题:智能体驱动十个阶段,而一个预言机(oracle)根据平台记录的事实对每个阶段进行评分。核心的隐性失败是那种"训练了却没学到"(Trains but Doesn't Learn, TBDL)的运行:损失下降,所有信号保持绿色(正常),但交付的模型并不比基础模型更好。运营商运行的验收门槛能在付款前捕获每一个此类运行,而基于已知损坏运行校准的检测器能在运行中途标记出严重的损坏。我们在按用量计费的L40S、A100和H200 GPU上,针对8B到70B的开放基础模型,端到端地运行了四个前沿智能体(Claude Opus 5、GPT-5.6-luna、Gemini 3.7 Flash、DeepSeek V4-Pro),并在评分前对每个场景进行了认证。我们还在同一预言机下运行了人类FDE对照组,并将每个智能体与其进行比较。
cs.LG / 13 / 2609.25297

Correcting Within-Group Self-Selection Bias in Prioritized Replay

纠正优先级回放中的组内自选择偏差
López-Feliu, Oscar Miró, van Hoof, Herke
Abstract
Prioritized experience replay (PER) improves sample efficiency by replaying high-priority transitions, usually according to absolute temporal-difference error. In stochastic environments, PER can distort the distribution of realized outcomes replayed from transitions with the same state-action pair. We call this within-group self-selection. We quantify the resulting changes in within-group outcome frequencies and mean Bellman targets. We decompose PER into between-group allocation and conditional sibling selection, and derive fixed-buffer corrections that preserve current group-level priority mass: SAMPLE selects a group through PER and trains on a uniformly sampled sibling; AVG averages sibling Bellman targets; and MODEL samples from an empirical full-outcome model. In exact state-action environments with rare high-magnitude outcomes, sibling-aware replay improves learning efficiency over PER, although matched parameter sweeps show that tuning can narrow some gaps. In MinAtar, approximate VQ-VAE groups with SAMPLE mitigate degradation under mean-preserving reward tails in four of five games. Sibling-aware replay thus retains the focus on high-priority state-action regions while recovering their empirical outcome frequencies.
Chinese Translation
优先级经验回放(Prioritized Experience Replay, PER)通过回放高优先级转移来提升样本效率,其优先级通常基于绝对时序差分误差。在随机环境中,PER会扭曲相同状态-动作对所回放的实际结果分布,我们将这种现象称为组内自选择。我们量化了由此导致的组内结果频率和贝尔曼均值目标的变化。我们将PER分解为组间分配与条件化同胞选择,并推导了在保持当前组级优先级质量不变前提下的固定缓冲区修正方法:SAMPLE通过PER选择一个组并在均匀采样的同胞上训练;AVG对同胞贝尔曼目标取平均;MODEL则从经验的全结果模型中采样。在存在稀有大幅度结果的精确状态-动作环境中,同胞感知回放相比PER提升了学习效率,尽管匹配的参数扫描表明调参可以缩小部分差距。在MinAtar环境中,采用SAMPLE的近似VQ-VAE分组在五个游戏中的四个里缓解了保均值奖励尾部下的性能退化。因此,同胞感知回放在保持对高优先级状态-动作区域关注的同时,恢复了其实际结果的经验频率。
cs.LG / 14 / 2609.25310

Topological Signal Processing With Unoriented Operators

Cavallo, Andrea, Sarathchandran, Varun, Leus, Geert, Isufi, Elvin
Abstract
Topological signal processing (TSP) processes signals on simplicial complexes with oriented boundary operators, which is the natural choice for flow signals or when the topological invariants play a role for the task at hand. However, many higher-order signals carry no orientation, and applying oriented operators to them is not well-defined since it introduces an arbitrary choice of simplex orientation. We study an unoriented TSP (UTSP) framework that replaces oriented boundaries with unoriented incidence matrices. First, we show that unoriented incidence and Laplacian matrices between arbitrary simplicial levels admit graph-like spectral properties. Second, since dropping orientation removes the Hodge decomposition, we introduce an unoriented counterpart, termed interaction-order decomposition, which quantifies how much of a higher-order signal is explained by aggregating lower-order signals. Third, we use this decomposition to derive regularizers for signal reconstruction that penalize each interaction order separately. Experiments on real-world data show that the order-aware regularizers outperform oriented baselines, with the largest gains when the signal energy is unevenly distributed across orders.
cs.LG / 15 / 2609.25326

Spatiotemporal Kronecker Covariance Neural Networks

时空克罗内克协方差神经网络
Cavallo, Andrea, Georgoutsos, Athanasios, Isufi, Elvin
Abstract
Multivariate time series contain complex patterns that span across both space and time. While covariance-based statistical tools like spatiotemporal Principal Component Analysis (ST-PCA) help identify these patterns, they are limited to linear operations and prone to estimation errors with limited data. Recent covariance-based spatiotemporal neural networks offer more stable, non-linear alternatives, but they ignore correlations across different time steps. To solve this, we introduce the Kronecker coVariance Neural Network (KVNN), a temporal graph neural network that represents the spatiotemporal covariance matrix via a sum of Kronecker products where spatial and temporal dependencies are decoupled. By implementing filtering operations on spatial and temporal components, KVNNs achieve expressive processing capabilities, admit a rigorous spectral analysis, and are provably stable to finite-sample estimation errors, ultimately addressing all of ST-PCA's limitations. We show on five real-world datasets that KVNNs achieve strong forecasting performance, often requiring significantly fewer trainable parameters than competitive methods, and are consistent under estimation noise.
Chinese Translation
多元时间序列包含跨越空间和时间的复杂模式。尽管诸如时空主成分分析(ST-PCA)等基于协方差的统计工具有助于识别这些模式,但它们仅限于线性运算,且在数据有限的情况下容易产生估计误差。近期出现的基于协方差的时空神经网络提供了更稳定的非线性替代方案,但它们忽略了不同时间步之间的相关性。为解决这一问题,我们提出了克罗内克协方差神经网络(Kronecker coVariance Neural Network,KVNN),这是一种时间图神经网络,通过克罗内克积之和来表示时空协方差矩阵,从而将空间和时间依赖性解耦。通过在空间和时间分量上实现滤波操作,KVNN 具备强大的表达能力,可进行严格的谱分析,并且对有限样本估计误差具有可证明的稳定性,最终克服了 ST-PCA 的所有限制。我们在五个真实世界数据集上展示了 KVNN 取得了出色的预测性能,其可训练参数量通常显著少于竞争方法,并且在估计噪声下保持一致性。
cs.LG / 16 / 2609.25334

MT-ProtBERT: Multi-task Learning ProtBERT for Intrinsically Disordered Proteins Classification with Scarce Data

Sun, Jian, Ghosh, Kingshuk, Houston, Lilianna, Mahoor, Mohammad H.
Abstract
Intrinsically disordered proteins (IDPs) differ from folded proteins in that they are dynamic, lack a stable three-dimensional conformation, and have low sequence similarity between similar proteins. The conformational heterogeneity of IDPs - while beneficial for their diverse functions - limits the use of traditional experimental tools to determine their conformation. The experimental difficulty, along with low sequence similarity, results in data scarcity, and makes it difficult to classify/detect IDPs that are similar or dissimilar, a task relevant to understand biology and evolution. We address this challenge using Multi-task ProtBERT (MT-ProtBERT), a multi-task extension of ProtBERT tailored for low-data regimes. MT-ProtBERT integrates Dynamic Window Masking, a Multi-Scale 1D Convolutional classifier (MS-Conv1D), and auxiliary objectives that jointly optimize masked language modeling and biochemistry-informed tasks. We evaluate this framework on two tasks under limited data: (i) phosphorylation site prediction (S/T/Y) in short sequences and small datasets, and (ii) protein compaction prediction on two small datasets (684 and 530 sequences), including sequences comparable in length to typical disordered regions. MT-ProtBERT consistently outperforms PARROT, an RNN-based IDP-specific model, across all tasks. These results demonstrate that combining self-supervised and biochemistry-informed tasks, and multi-scale learning enables robust modeling of unstructured proteins under data scarcity.
cs.LG / 17 / 2609.25340

Concept Drift from a Causal Perspective

Barboza, Eduardo V. L., Barddal, Jean Paul, Sabourin, Robert, Cruz, Rafael M. O.
Abstract
Concept drift is a common phenomenon in real-world data streams, in which changes in the data-generating distribution can degrade predictive model performance. Most existing definitions characterize drift as changes in the joint distribution $P(\mathbf{x}, y)$, without distinguishing which component of the data-generating process has changed. In this work, we introduce a causal perspective on concept drift based on Structural Causal Models (SCMs). We propose a taxonomy that categorizes drift events by their causal origin, including changes in exogenous variables, endogenous mechanisms, confounders, and target-generating processes. Building on this framework, we develop an SCM-based data stream generator that simulates controlled mechanism-level drift events. Our experiments empirically characterize the distributional effects of each drift type and show that drifts with different causal origins induce distinct patterns of distribution shift and predictive behavior. Furthermore, by integrating causal discovery methods, we use our framework to construct data streams grounded in real-world dependency structures, enabling more realistic and informative evaluation scenarios. We also demonstrate that leveraging the generated data can improve downstream performance. These results highlight the importance of accounting for causal structure when studying and evaluating adaptive learning methods, and establish a foundation for causally-aware evaluation in non-stationary environments.
cs.LG / 18 / 2609.25373

Extending FunctionGemma for Practical On-Device Mobile Function Calling

Rezagholizadeh, Ali, Samiee, Soheila
Abstract
On-device assistants require function-calling models that map natural language to local system actions, but existing resources emphasize web APIs or narrow mobile-action catalogs. We extend FunctionGemma 270M-it to practical Android workflows by introducing MOBILEACTIONSEXTENDED, a synthetic, schema-validated dataset of ~9,500 conversations covering fifteen device-control categories, including messaging, phone calls, camera/screenshot, brightness control, device-status queries, flashlight control, and application management. We fine-tune the 270M model with TRL supervised fine-tuning under completion-only loss, producing an extended specialist and a combined model trained jointly with Google's MOBILEACTIONSGOOGLE. On MOBILEACTIONSEXTENDED, end-to-end accuracy improves from 29.3% for the base model and 17.2% for Google's Mobile-Actions variant to 76.5%. The combined model retains 76.5% on MOBILEACTIONSEXTENDED and reaches 82.3% on MOBILEACTIONSGOOGLE, down from the 90.3% of Google's Mobile-Actions specialist, representing an 8.0-percentage-point trade-off in return for doubling category coverage. We release the dataset, fine-tuned models, reproducible training/evaluation pipeline, and an Android demo, highlighting compact local function calling as a practical path towards low-latency and privacy-preserving mobile assistants.
cs.LG / 19 / 2609.25397

Deep Reinforcement Learning on Item-Compatibility Graphs for One-Dimensional Bin Packing

基于物品兼容性图的一维装箱问题深度强化学习方法
Aydın, M. Aslı
Abstract
The one-dimensional bin packing problem (1D-BPP) is a classical NP-hard combinatorial optimization problem with applications ranging from logistics and manufacturing to cloud resource management. Although deep reinforcement learning (DRL) has become a competitive paradigm for data-driven optimization, most learned packing methods target 2D and 3D variants, and intelligent learned solvers for 1D-BPP remain scarce. In this paper, we present a novel end-to-end, size-agnostic graph reinforcement learning framework for 1D-BPP. We formulate the packing process as a Markov decision process on an item-compatibility graph, serving as a structural knowledge representation in which every action merges two partial bins that fit together. A graph neural network actor-critic policy extracts relational features from this representation and is trained through reinforcement learning and decoded by stochastic beam search, enabling a single trained model to generalize zero-shot to instances of any size. We conduct a systematic empirical study across graph encoders, DRL algorithms, reward functions, training distributions, and hyperparameters. Evaluated zero-shot on the full BPPLIB benchmark against a constructive heuristic, a grouping genetic algorithm, and recent learned methods, our data-driven policy lowers the mean optimality gap of the constructive heuristic from 2.66\% to 2.31\%, with the largest gains on structured instances. Against learned baselines evaluated on the same benchmark, it attains a lower gap on most of the nine families and is far more stable across instance distributions. On the hardest benchmark family, it outperforms a state-of-the-art learned solver that relies on column generation and integer programming, while using no solver at all. A grouping genetic algorithm remains ahead overall, and we analyze where and why the residual gap arises.
Chinese Translation
一维装箱问题(1D-BPP)是一个经典的NP难组合优化问题,其应用涵盖物流、制造以及云计算资源管理等领域。尽管深度强化学习(DRL)已成为数据驱动优化的一种有竞争力的范式,但大多数学习型装箱方法针对的是二维和三维变体,针对一维装箱问题的智能学习型求解器仍然稀缺。本文提出了一种新颖的端到端、尺寸无关的一维装箱图强化学习框架。我们将装箱过程建模为物品兼容性图上的马尔可夫决策过程,该图作为一种结构化知识表示,其中每个动作将两个可相互容纳的部分箱子进行合并。图神经网络 actor-critic 策略从该表示中提取关系特征,通过强化学习进行训练,并采用随机束搜索进行解码,使单个训练好的模型能够零样本泛化到任意规模的实例。我们围绕图编码器、DRL算法、奖励函数、训练分布和超参数开展了系统的实证研究。在完整的BPPLIB基准上与构造性启发式算法、分组遗传算法以及近期的学习方法进行零样本对比评估,我们的数据驱动策略将构造性启发式算法的平均最优性差距从2.66%降低至2.31%,且在结构化实例上提升最为显著。与在同一基准上评估的学习型基线相比,本方法在九个实例族中的大多数上取得了更小的差距,并且在不同的实例分布上具有远更好的稳定性。在最难的基准实例族上,本方法优于一个依赖列生成和整数规划的最先进学习型求解器,而全程未使用任何求解器。分组遗传算法总体上仍然领先,我们分析并解释了剩余差距产生的原因与所在。
cs.LG / 20 / 2609.25430

Predictive Uncertainty for Neural CAE Surrogates

Tangsali, Kaustubh, Nabian, Mohammad Amin, Lee, Kelvin, Gonzales, Carmelo, Choudhry, Sanjay
Abstract
Neural surrogates can substantially accelerate computer-aided engineering (CAE) workflows, but their use in design requires uncertainty estimates that remain meaningful across varying geometries, spatial prediction fields, and engineering quantities of interest. We investigate how established uncertainty quantification (UQ) approaches behave when adapted to geometry-conditioned neural surrogates. We compare one closed-form and two sampling-based approaches-a Gaussian process (GP)-based method, concrete Monte Carlo (MC) dropout, and deep ensembles-and evaluate them on three large, industry-relevant CAE datasets for external aerodynamics and crash dynamics. We examine whether predicted uncertainties have credible magnitudes, identify locations with larger prediction errors, respond to unfamiliar inputs, and remain informative for derived engineering quantities. On the DrivAerStar dataset, where all three methods are compared, each generally assigns higher uncertainty to locations with larger prediction errors, and validation-based rescaling brings interval coverage close to nominal on a disjoint in-distribution test set. Results on AirFRANS and automotive crash also show useful error ranking and interval estimates, but the relative performance of the methods changes with the dataset and evaluation criterion. UQ methods and evaluation metrics should therefore be selected based on the intended downstream CAE decision.
cs.LG / 21 / 2609.25433

Lightweight Ranking Heads: Accelerating Multi-Task Experimentation in Production Recommender Systems

Girija, Sanjay Surendranath, Nath, Aniruddh, Wei, Li, Jiang, Yanhao, Andrews, Shawn, Heldt, Lukasz, Wu, Yi, Mahajan, Aditya, Sharma, Mohit
Abstract
Modern production-scale recommender systems rely on complex, multi-task ranking models. Introducing new prediction tasks into these massive systems often causes bottlenecks - it risks negative task conflicts with existing tasks, and can lead to long development and experimentation cycles due to the expensive retraining of backbone models and downstream models or tuning of reward combination formulas. To address the critical challenge of slow experimentation velocity, we introduce the Lightweight Ranking Heads (Light Heads) framework. Designed for continuous online learning environments, Light Heads enable the dynamic injection of new tasks into existing multi-task ranking models, effectively obviating the need for model cold-starting and retraining of backbone models. By utilizing stop-gradients and stateless daily training, this design strictly isolates new tasks, mitigating the risk of adverse task conflicts. Crucially, this framework uses a centralized configuration that allows Light Heads to be added to multiple models simultaneously, unblocking faster training data generation and co-training of downstream models. Successfully deployed at YouTube scale, this approach reduces the iteration cycle for multi-task experimentation from several weeks to days. In this paper, we detail the system architecture, analyze the training dynamics of stateless cold-started heads, compare their performance to full heads, and demonstrate how Light Heads have enabled the rapid A/B experimentation and deployment of new ranking tasks that yield measurable production value.
cs.LG / 22 / 2609.25438

PermuFormer: Multi-Task Pretraining for Permutation Representation in Algebraic Combinatorics

Kvinge, Henry
Abstract
Diverse pretraining has been shown to be an effective method for learning reusable, domain-aware representations that provide a starting point for fine-tuning on downstream tasks. While much of the excitement in AI for math has been concentrated in the use of frontier reasoning models to solve well-specified problems through the medium of language, narrow, specialized models remain an important component of the AI for math ecosystem. In contrast to large language models, specialized models are usually trained directly on the mathematical objects themselves (e.g., graphs, sequences of numbers) rather than the textual descriptions that characterize these objects. However, the common practice of training specialists from scratch may limit their ability to develop domain-aware representations that capture the multifaceted nature of mathematics. In this paper, we describe an approach to pretraining for permutation-focused tasks in algebraic combinatorics. We introduce PermuFormer, an autoregressive transformer trained on a 2.8 billion token multi-task, multi-encoding corpus. We show that PermuFormer is an effective starting point for fine-tuning on basic tasks unseen during pretraining and more complex research-level tasks, frequently outperforming the same architecture trained from scratch, baseline MLPs, and a fine-tuned generic language model of comparable size. We also analyze some of the internal mechanisms by which PermuFormer learns to solve training tasks. For example, we show that while some tasks can be linearly decoded directly from the internal representation of the prompt, other tasks require multiple rounds of generation before the answer can be decoded.
cs.LG / 23 / 2609.25444

Mean Velocity Matching: Rethinking Generative Dynamics in Diffusion Models

均值速度匹配:重新思考扩散模型中的生成动力学
Zhang, Yunhong, Cao, Changjie, Zhang, Zhihua, Liu, Bingli, Cao, Zongjie, Cui, Zongyong, Yang, Ying
Abstract
This work studies prediction parameterization for stochastic generative dynamics in diffusion models. Existing velocity-based generative models provide the simplicity of learning a single transport field, but their standard formulation is deterministic, whereas stochastic extensions generally require additional score information or an intermediate velocity-to-score reconstruction. To retain single-field prediction while directly supporting stochastic reverse dynamics, this paper introduces Mean Velocity Matching (MVM). MVM constructs a Gaussian perturbation process for which the conditional expectation of a restoration-oriented velocity, $(x_0-x_t)/t$, directly forms the reverse-SDE drift. Consequently, a single learned field is sufficient to parameterize the stochastic reverse process without separately estimating or reconstructing the score. Because direct regression of this velocity becomes unbounded near $t=0$, MVM further introduces a $\sqrt{t}$-scaled parameterization that preserves the reverse dynamics while yielding a bounded training target. The same learned field also induces a deterministic probability-flow ODE, enabling stochastic and deterministic sampling to be studied within a unified formulation. Experiments with Transformer-based generative models achieve an FID of $\MVMImageNetThirtyTwoFID$ at \MVMImageNetThirtyTwoNFE\ NFE on ImageNet $32\times32$ and $\MVMImageNetTwoFiftySixFID$ at \MVMImageNetTwoFiftySixNFE\ NFE on ImageNet $256\times256$. Controlled SDE--ODE comparisons further show that the ODE performs better under very low NFE, whereas the stochastic reverse process achieves lower FID when sufficient function evaluations are available. These results demonstrate that MVM provides a direct single-field parameterization of stochastic reverse dynamics while maintaining competitive generation quality.
Chinese Translation
本文研究扩散模型中随机生成动力学的预测参数化方法。现有的基于速度的生成模型具有学习单一输运场的简洁性,但其标准形式是确定性的,而随机扩展通常需要额外的分数(score)信息或中间的速度到分数的重构。为了在保留单场预测的同时直接支持随机逆动力学,本文提出了均值速度匹配(Mean Velocity Matching, MVM)。MVM 构造了一个高斯扰动过程,使得面向恢复的速度 $(x_0-x_t)/t$ 的条件期望直接构成逆随机微分方程(SDE)的漂移项。因此,仅用一个学习到的场即可参数化随机逆过程,而无需单独估计或重构分数。由于该速度的直接回归在 $t=0$ 附近会变得无界,MVM 进一步引入了 $\sqrt{t}$ 缩放的参数化方法,在保持逆动力学的同时使训练目标有界。同一个学习到的场还能诱导出确定性的概率流常微分方程(ODE),从而可以在统一框架下研究随机采样与确定性采样。基于 Transformer 的生成模型实验在 ImageNet $32\times32$ 上以 \MVMImageNetThirtyTwoNFE\ 次 NFE 取得了 FID 为 $\MVMImageNetThirtyTwoFID$ 的结果,在 ImageNet $256\times256$ 上以 \MVMImageNetTwoFiftySixNFE\ 次 NFE 取得了 FID 为 $\MVMImageNetTwoFiftySixFID$ 的结果。受控的 SDE 与 ODE 对比实验进一步表明,在极低 NFE 下 ODE 表现更好,而当函数求值次数充足时,随机逆过程可取得更低的 FID。这些结果表明,MVM 为随机逆动力学提供了一种直接的单场参数化方法,同时保持了具有竞争力的生成质量。
cs.LG / 24 / 2609.25471

A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization

Bae, Wonho, Aldeneh, Zakaria, Pelikan, Martin, Silovsky, Jan "Honza", Likhomanenko, Tatiana, Azam, Sheikh Shams
Abstract
Semi-supervised federated learning (SSFL) trains models on clients' unlabeled data using a teacher to generate pseudo-labels, with a small labeled seed dataset on the server. Automatic Speech Recognition (ASR) is particularly fragile here: pseudo-label errors compound across the output sequence and across training rounds into divergence, leaving a large gap to fully-supervised FL. We show that closing this gap turns on two coupled design axes -- the teacher (which model generates the pseudo-labels) and the anchor (the server-side updates on labeled data that stabilize training). On the teacher axis, a per-client online teacher (each client's own evolving model) diverges on its own, but once stabilized it matches or beats the broadcast global teacher (one server model, fixed within a round) -- decisively in-domain and competitively under domain shift. As the seed grows stronger and the online teacher's advantage narrows, a transitioning teacher (global $\rightarrow$ online at round $r$) matches or beats both. On the anchor axis, the server must keep training on labeled data between rounds -- otherwise the online teacher drifts -- and this interleaving, more than the seed model, governs convergence. The two axes are inseparable: aggressive teacher choices pay off only once the anchor stabilizes training, which is highly sensitive to data augmentation and batch size -- the settings that govern how much input and gradient noise the server injects. How much stabilization is needed is domain-dependent, governed by the dispersion of the seed data and its overlap with client data. These findings yield guidelines for SSFL in ASR training, improving over the strongest prior method on 9 of 11 pairs, by $20.8\%$ on average in-domain and $10.0\%$ cross-domain, narrowing the gap to fully-supervised FL.
cs.LG / 25 / 2609.25482

Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining

终端收缩平均揭示了大型语言模型预训练中调度与估计器的相互作用
Ousherovitch, Adam, Wang, Yixin
Abstract
Large language model (LLM) pretraining conventionally returns the raw final iterate. This couples two design choices: the learning-rate schedule that generates the parameter trajectory and the estimator that constructs the deployed model (e.g. the raw final iterate or a checkpoint average). A schedule that promotes optimization progress may differ from one that minimizes variation in the raw final iterate. Separating these choices creates an opportunity to maintain progress late in training while reducing variation in the returned model. To this end, we propose \emph{Terminal Shrinkage Averaging (TSA)}, which interpolates between the raw final iterate and the average of recent checkpoints to balance recent progress against terminal variation. We analyze how TSA changes the preferred terminal learning-rate schedule under a local quadratic approximation and test this interaction through a sequence of controlled NanoChat experiments. Finally, we demonstrate that the resulting gains transfer to depth-22 NanoChat, where the combined schedule and estimator improve validation quality. A qualifying time-to-GPT-2 run also finishes faster than the public baseline used in our experiments, providing preliminary evidence of benchmark acceleration.
Chinese Translation
大型语言模型(LLM)预训练通常直接返回原始的最终迭代结果。这将两个设计选择耦合在一起:生成参数轨迹的学习率调度,以及构建部署模型的估计器(例如原始最终迭代结果或检查点平均)。促进优化进展的调度可能与最小化原始最终迭代结果变化的调度有所不同。将这两个选择分离,就创造了一个机会:在训练后期保持优化进展的同时,降低所返回模型的变化。为此,我们提出了终端收缩平均(Terminal Shrinkage Averaging, TSA),它在原始最终迭代结果与近期检查点的平均值之间进行插值,以平衡近期进展与终端变化。我们在局部二次近似下分析了TSA如何改变优选的终端学习率调度,并通过一系列受控的NanoChat实验验证了这种相互作用。最后,我们证明由此带来的收益可以迁移到深度为22的NanoChat上,其中组合的调度和估计器提升了验证质量。一项符合资格的时间到GPT-2(time-to-GPT-2)运行也快于我们实验中使用的公开基线,为基准加速提供了初步证据。
cs.LG / 26 / 2609.25484

Learning Defensive Policies against Diverse Inference Attacks for Smart Meter Privacy

Zhang, Ruichang, Mustafa, Mustafa A.
Abstract
Smart meter (SM) data provides fine-grained visibility into household energy consumption, but also exposes users to privacy risks. Inference attacks, known as non-intrusive load monitoring (NILM), can perform appliance-level inference from aggregate signals and recover sensitive behavioral patterns. In practice, attacker models are unknown and heterogeneous, making robust defense challenging. We formulate SM privacy protection as a black-box inference defense problem, aiming to reduce the recoverability of appliance-level information while generalizing across diverse and unseen attackers. We propose a proxy-guided hierarchical reinforcement learning framework that learns battery-based load-shaping policies to inject realistic but misleading appliance-level signatures into the aggregate signal, thereby disrupting the structured patterns exploited by NILM. A self-supervised aggregate-structure privacy probe provides a reconstruction-error-based surrogate reward for disrupting recoverable load structure, while a signature library makes the perturbations appliance-relevant and physically realizable through battery control. We provide theoretical rationale showing that proxy-guided optimization improves inference robustness under attacker diversity. Experiments on real-world datasets UK-DALE and REDD demonstrate strong cross-model and cross-appliance generalization. Across six unseen NILM attackers, covering four appliances on UK-DALE and five on REDD, our proposed defense increases average appliance-level RMSE by 107% and 166%, respectively, while reducing F1 score by 79% and 80%.
cs.LG / 27 / 2609.25501

Continuous Optimization for p-adic Models

p进模型的连续优化
Salazar, Julian, Kanevsky, Dimitri, Harvey, Matt, Getreuer, Pascal, Dixon, Lucas
Abstract
We present the first method for native, continuous gradient descent for machine learning models with $p$-adic parameters. Existing native optimizers are discrete, mostly combinatorial searches, as the $p$-adic numbers $\mathbb{Q}_p$ are totally disconnected, with standard losses that are flat away from their minima. To enable continuous optimization, we propose working with $\mathbb{Q}_p$ via its Berkovich affine line: a canonical, path-connected expansion of $\mathbb{Q}_p$ that preserves its isometries and uniquely extends its analytic maps. This hull is a metric tree with interpretable points and local derivatives, which we show enables effective optimizers and backpropagation. We formulate gradient descent and show that its approximations efficiently learn linear models with coefficients in $\mathbb{Q}_p$ to do modular arithmetic, an XOR-like task not expressible by linear models in $\mathbb{R}$. We also demonstrate momentum and Adam variants, linear regression, and classification on binary-encoded hierarchies (Quillian semantic networks), addressing open problems posed by Martins (2025). Library at https://github.com/google-deepmind/padic-ml
Chinese Translation
我们提出了首个针对具有 $p$ 进参数的机器学习模型的原生连续梯度下降方法。由于 $p$ 进数域 $\mathbb{Q}_p$ 是完全不连通的,且标准损失函数在远离极小值处为平坦区域,现有的原生优化器均为离散的,大多为组合搜索。为实现连续优化,我们提出通过 Berkovich 仿射线来处理 $\mathbb{Q}_p$:这是 $\mathbb{Q}_p$ 的一个典范的、路径连通的扩张,它保留了 $\mathbb{Q}_p$ 的等距变换,并唯一地延拓其解析映射。该闭包是一棵度量树,其上的点具有可解释性并存在局部导数,我们证明这使得有效的优化器和反向传播成为可能。我们给出了梯度下降的公式,并证明其近似方法能够高效地学习系数属于 $\mathbb{Q}_p$ 的线性模型来执行模运算——这是一种类异或(XOR)任务,无法用 $\mathbb{R}$ 上的线性模型表达。我们还展示了动量法与 Adam 变体、线性回归,以及二进制编码层级结构(Quillian 语义网络)上的分类任务,解决了 Martins (2025) 提出的开放性问题。代码库见 https://github.com/google-deepmind/padic-ml
cs.LG / 28 / 2609.25510

Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training

用于测试时扩展的爬山采样(Hill Sampling):重复采样、进化与训练的一种更简单且更优的替代方案
Beck, Jacob, Ogren, Philip V., Kobren, Ari
Abstract
Large language models (LLMs) can improve solutions to verifiable scientific and algorithmic problems by spending additional computation at test time. Recent systems achieve strong results with increasingly elaborate evolutionary search harnesses or by updating model parameters during test-time training. We ask how much of this machinery is necessary. We introduce Hill Sampling, a simple procedure that repeatedly samples candidate program edits from a frozen LLM, retains the best program found so far, and conditions all subsequent samples on that program. We evaluate the method on circle packing, sums/differences of sets, and Erdos' minimum-overlap problem using three open-weight models. Hill Sampling sets a new state of the art on circle packing among published methods, improves over the AlphaEvolve reference on Erdos' minimum-overlap problem, and achieves strong results on sums and differences of finite sets. The circle-packing and Erdos results require only hours of wall-clock time on eight NVIDIA H100 GPUs. To our knowledge, we also conduct, the largest study, by parameter count, of evolution strategies (ES) applied directly to LLM weights at test time. Surprisingly, learning the weights is worse than setting the ES learning rate to zero: at zero learning rate, the method is still searching in weight space through fixed random perturbations. Those perturbations can help exploration, but randomness from token sampling is stronger still, and repeated sampling remains substantially weaker than Hill Sampling. These results suggest a simple test-time compute allocation strategy: repeatedly sample edits to the best verified solution found so far, before introducing additional complexity such as adding archives, diversity mechanisms, evolutionary scaffolds, or test-time parameter learning.
Chinese Translation
大语言模型(LLMs)可以通过在测试时投入额外计算来改进可验证的科学和算法问题的求解方案。近期的一些系统借助日益复杂的进化搜索框架,或在测试时训练中更新模型参数,取得了出色的结果。我们追问:这些复杂机制究竟有多少是必要的?我们提出了爬山采样(Hill Sampling),这是一种简单的程序:从一个冻结的LLM反复采样候选程序修改,保留迄今为止找到的最佳程序,并让所有后续采样都以该程序为条件。我们使用三个开放权重模型,在圆 packing、集合的和/差以及 Erdős 最小重叠问题上对该方法进行了评估。在已发表的方法中,Hill Sampling 在圆 packing 问题上创造了新的最先进纪录;在 Erdős 最小重叠问题上超越了 AlphaEvolve 的参考结果;并在有限集合的和与差问题上取得了强劲的表现。圆 packing 和 Erdős 问题的结果仅需在八块 NVIDIA H100 GPU 上运行数小时的墙钟时间。据我们所知,我们还进行了(按参数量计算)最大规模的、将进化策略(ES)直接应用于测试时 LLM 权重的研究。令人惊讶的是,学习权重反而不如将 ES 学习率设为零:在零学习率下,该方法仍在通过固定的随机扰动在权重空间中搜索。这些扰动有助于探索,但来自 token 采样的随机性作用更强,而重复采样仍然显著弱于爬山采样。这些结果提示了一种简单的测试时算力分配策略:在引入额外复杂性(如添加档案库、多样性机制、进化脚手架或测试时参数学习)之前,先反复对迄今为止找到的最佳已验证解进行修改采样。
cs.LG / 29 / 2609.25541

A JEPA Recipe for Tabular Foundation Models

Jeon, Mingyu, Cho, Suwan, Suh, Jae Young
Abstract
Tabular foundation models learn to predict cell values in context, whereas world-model self-supervision asks for prediction in representation space (LeCun, 2022; Assran et al., 2023). On a tabular foundation-model prior, the latent term of a joint-embedding predictive architecture (JEPA) collapsed in our earlier runs and took the encoder with it to a constant map. We report a recipe under which the latent term survives to convergence beside the value objective: the value head reads the encoder field rather than the predictor, and the target is an exponential moving average (EMA) difference. To bound its cost against the value-only arm, both arms train until a plateau rule stops them, with no fixed step budget. A fixed horizon had confounded a slowdown with a ceiling, since the value-only arm was still improving well past the usual budget. At convergence, in one run per arm, the JEPA arm trails the value-only arm across 147 real datasets, 32:70 wins to losses on classification (29:63 with one entry per dataset name) and 8:24 on regression, the margin small on classification and wider on regression, and the count leans the same way in each stratum and each benchmark. The JEPA arm (jepa) needs 1.42 times as many steps as the value-only arm (ds), and 1.66 times its wall-clock, to reach its plateau.
cs.LG / 30 / 2609.25542

DefaultGNN: A Dual-Perspective GNN Framework for Predicting Corporate Default from Buyer-Seller Transaction Networks

Kim, Junghoon, Kim, Hyunsung, Choi, Seungyoon, Park, KyoungYong, Lee, Jihun, Ji, YongGu, Park, Chanyoung
Abstract
Corporate default prediction is a core problem in financial risk management, yet traditional credit models rely heavily on financial statements that are often sparse or unavailable for many firms. Corporate transaction networks offer a complementary view of real economic activity, but how risk propagates through buyer-seller relationships remains underexplored. We conduct a large-scale empirical study using real-world electronic tax-invoice data spanning six years that links transaction histories with default events, revealing that transaction-driven risk is both role-dependent (buyer or seller) and scale-dependent. Based on these findings, we construct multiplex buyer-view and seller-view transaction networks and propose DefaultGNN, a dual-perspective graph neural network-based framework for corporate default prediction. DefaultGNN integrates both views to model how risk flows through transactional relationships, achieving strong improvements over both attribute-based and graph-based baselines, especially for firms with limited intrinsic risk signals. We further provide interpretable network-based explanations by visualizing how distressed trading partners contribute to default risk. In collaboration with a licensed credit rating agency, we validate that DefaultGNN's predictions complement existing credit scoring models, improving approval rates by 7-11%p without increasing default risk among approved firms. The source code can be found at https://github.com/jhkim611/DefaultGNN
cs.LG / 31 / 2609.25569

SambaGraph: Action-Reaction Spatio-Temporal Graphs for Soccer Tactical Response Modeling

SambaGraph:用于足球战术响应建模的动作-反应时空图
Reyes-Angulo, Abel A., Velesaca, Henry O., Araujo, Steven
Abstract
Soccer tactics are interactive: an attacking action changes the opponent's defensive problem, and the observed response depends on the multi-agent match state. We introduce SambaGraph, an action--reaction spatio-temporal graph dataset and benchmark for soccer tactical response modeling. From tracking and event data for all 64 matches of the 2022 FIFA World Cup, we curate 4,070 action-centered episodes represented as temporally aligned 23-node player--ball graph sequences with attack/defense views, response labels, and 26,270 split-safe attack--defense pairs. We study three questions: whether observed responses can be classified from graph episodes, whether successful defenses can be retrieved for a query attack, and whether graph-derived summaries support grounded LLM reasoning. A compact signature MLP obtains $0.796\pm0.007$ macro-F1 for response classification, while a fused graph--signature dual encoder reaches $0.471\pm0.029$ Hit@5 and $0.655\pm0.051$ Hit@10 for full-bank defensive retrieval. Hard negatives maximize pair discrimination but not retrieval quality. Local LLMs underperform supervised encoders for direct classification and do not improve over a strong original order in eight-candidate reranking, but they provide grounded tactical rationales. These results position SambaGraph as a reproducible benchmark for graph-based soccer strategy-response research. Code and dataset are available at: https://github.com/areyesan/SambaGraph.
Chinese Translation
足球战术具有交互性:进攻方的动作会改变防守方所面临的问题,而所观察到的防守响应取决于多智能体比赛状态。我们提出了SambaGraph,一个用于足球战术响应建模的动作-反应时空图数据集与基准。基于2022年FIFA世界杯全部64场比赛的追踪数据与事件数据,我们构建了4,070个以动作为中心的片段,表示为时间对齐的23节点球员-球图序列,包含进攻/防守视图、响应标签以及26,270个划分安全的进攻-防守配对。我们研究三个问题:能否从图片段中对观察到的响应进行分类;能否针对查询的进攻检索出成功的防守案例;以及图衍生的摘要能否支持有据可依的大语言模型(LLM)推理。一个紧凑的签名MLP在响应分类任务上取得了0.796±0.007的宏平均F1值,而一个融合图与签名的双编码器在全库防守检索任务上达到了0.471±0.051的Hit@5和0.655±0.051的Hit@10。困难负样本可以最大化配对区分度,但并不能提升检索质量。局部LLM在直接分类任务上表现不及监督编码器,且在八候选重排序任务上也未能超越较强的原始顺序基线,但它们能够提供有据可依的战术解释。这些结果使SambaGraph成为基于图的足球策略-响应研究的可复现基准。代码与数据集可在以下网址获取:https://github.com/areyesan/SambaGraph。
cs.LG / 32 / 2609.25582

EMGBlend: Heterogeneity-Aware Self-Supervised Pretraining for Gesture and Force Decoding

EMGBlend:面向手势与力量解码的异质性感知自监督预训练
Jia, Yuwei, Zhong, Cheng, Yu, Jinyang, Cui, Zhe
Abstract
Public surface electromyography (EMG) datasets vary widely in electrode layout, channel count, frequency support, and size. Simply mixing them for pretraining can misalign channel semantics, introduce spectral targets that some devices cannot observe, and let large or high-channel-count datasets dominate learning. We introduce EMGBlend, a self-supervised framework designed around these differences. It combines shared channel patches with geometry-aware attention, restricts spectral targets to each recording's supported frequency band, and balances exposure across data sources. We pretrain a 109M-parameter model on 11 public EMG sources and evaluate it on gesture recognition, continuous-force regression, and contact classification. EMGBlend consistently outperforms matched random initialization and waveform reconstruction controls. Fixed-budget source controls show that multi-source pretraining improves gesture recognition and remains competitive for force decoding. Ablations confirm that geometry, band-aware targets, and source balancing each contribute to transfer, although cross-person NinaPro force estimation remains difficult. Overall, EMGBlend shows how heterogeneous EMG datasets can be combined through explicit mechanism design rather than simple concatenation. Code is available at https://github.com/tamanano/EMGBlend
Chinese Translation
公开的表面肌电图(EMG)数据集在电极布局、通道数、频率支持和数据规模等方面差异巨大。简单地将它们混合用于预训练会导致通道语义错位、引入某些设备无法观测到的频谱目标,并使大型或高通道数的数据集主导学习过程。我们提出了EMGBlend,一个围绕这些差异设计的自监督框架。它将共享通道补丁与几何感知注意力相结合,将频谱目标限制在每个记录所支持的频带内,并平衡各数据源的暴露程度。我们在11个公开EMG数据源上预训练了一个1.09亿(109M)参数的模型,并在手势识别、连续力量回归和接触分类任务上进行评估。EMGBlend始终优于匹配的随机初始化和波形重建对照方法。固定预算的数据源对照实验表明,多源预训练提升了手势识别性能,并在力量解码任务上保持竞争力。消融实验证实,几何建模、频带感知目标和数据源平衡各自都对迁移效果有贡献,尽管跨被试的NinaPro力量估计仍然困难。总体而言,EMGBlend展示了如何通过显式的机制设计而非简单的数据拼接来整合异质性EMG数据集。代码可在 https://github.com/tamanano/EMGBlend 获取。
cs.LG / 33 / 2609.25602

Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs

Pandere, Shubham Santosh, Ranka, Gautam, Varshney, Ritika, Deshmukh, Navya, Sareen, Roushni, Singh, Roshan Kumar
Abstract
In language models, the choice between believing the prompt and believing the weights is made by a handful of identifiable attention heads. Instruction tuning changes how models behave under conflict, but whether it rewires the underlying circuit or merely gates/reweights already present components, remains unknown. We provide the first mechanistic base-vs-instruct comparison of conflict-resolution circuits, across three families (Llama-3.2-3B, Qwen-2.5-3B, Gemma-3-4B). Five independent methods, node and edge attribution, superposition role analysis, causal ablation, and path patching, converge on gating, with the same heads, in the same late-layers, are found to be reweighted rather than replaced with a high node overlap (0.60-0.82). Behaviorally, tuning shifts models toward parametric memory, making instruct models reject a terse counterfactual context far more than base ones, the opposite of a naive user-following expectation. Yet this added skepticism is a factor of framing since it disappears when the same false claim is delivered as a coherent, evidential passage. The robustness that instruction tuning buys against terse injection is therefore real but narrow. More broadly, we believe that because the conflict circuit is preserved rather than rebuilt, interpretability and control tools calibrated on base models should transfer directly to their deployed instruct siblings.
cs.LG / 34 / 2609.25623

What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation

Tian, Kanghui, Liu, Siyuan, Jiang, Tianxiang, Dong, Shuai, Li, Yizhuo, Ding, Tian, Guo, Yuan, Li, Songze, Hou, Haowen, Wang, Congcong, Wang, Yi
Abstract
More privileged information does not always make a better teacher. We study this tension in on-policy self-distillation (OPSD), where a frozen copy of the base model scores the student's own rollouts under privileged context, conventionally a complete reference solution that bundles the final answer with one particular reasoning path. Holding the student view and training fixed within each scale, we compare that default against three abstractions compiled offline, a named strategy, a method-independent framing, and a problem category, and against an answer-only control that keeps the destination but removes the path. In the primary runs on competition mathematics, the best intermediate contexts improve the in-domain peak mean over the full solution by 1.4 points at 4B and 1.6 at 8B, while storing an order of magnitude fewer hint tokens. Comparisons across three seeds also show positive mean gains for the framing and category contexts at both scales. Answer-only conditioning remains competitive in the primary runs, within 0.2 points of the full solution at these scales. The preferred context varies with student scale and task. Initial teacher-student KL does not order downstream performance. What a self-teacher should see is therefore not everything it could, but the level of abstraction its student can still act on.
cs.LG / 35 / 2609.25634

An Exploratory Replica-Overlap Probe of the Grokking Transition

关于Grokking(顿悟)转变的探索性副本重叠探测
Opus, A. C., Lu, J. Q.
Abstract
We trained 64 independently seeded networks in four configurations, continuing each to sustained convergence or a 40,000-epoch ceiling. We then asked whether an RSB-inspired distribution of pairwise weight overlaps changes across the grokking transition. It is the alignment step, not the overlap statistic, that determines what this registered probe can report. The registered implementation permutes hidden units without the corresponding bias and head-internal permutations and therefore does not preserve the network function. Every q_wt value computed through this alignment inherits the defect; q_fn does not, because it is computed from predictions of the unpermuted models. The numerical-precision requirement also failed, and an audit found protocol deviations. Consequently, the pre-registered rule gives no verdict: registered outcome UNDETERMINED (reason code C0_INSTRUMENT_INVALID). These data provide neither a confirmatory null nor a validated reading of the Parisi order parameter. Only frac40 cleared the 12/16 checkpoint-completeness requirement. For this configuration, a post-hoc criterion applied to the same data gave a Hartigan-dip interval containing zero (95% CI for Delta dip = [-0.017, 0.034]), whereas the overlap standard deviation increased by a factor of about 5.6. A post-hoc calibration assigns the dip test zero power at the simulated separations; the interval is therefore uninformative, not evidence of no change. The standard-deviation ratio is the only statistic here with power at the observed effect. Ensemble loss was near-flat only under the pre-specified 1% threshold. Finally, grokking rates of 0/16, 11/16 and 16/16 remain descriptive because train fraction is confounded with split identity.
Chinese Translation
我们训练了64个具有独立随机种子的网络,涵盖四种配置,并将每个网络训练至持续收敛或达到40,000个epoch的上限。随后我们探究:受副本对称性破缺(RSB)启发的成对权重重叠分布是否会在grokking转变过程中发生变化。决定该预注册探测所能报告内容的是对齐步骤,而非重叠统计量本身。该预注册实现中对隐藏单元进行置换时,未包含相应的偏置和头内部(head-internal)置换,因此不能保持网络功能不变。通过该对齐计算的每一个q_wt值都继承了这一缺陷;而q_fn不受影响,因为它是由未置换模型的预测计算得出的。数值精度要求也未满足,且审计发现了协议偏差。因此,预注册规则无法给出判定:预注册结果为UNDETERMINED(原因代码C0_INSTRUMENT_INVALID)。这些数据既不能提供确认性的零结果,也不能对Parisi序参量给出经过验证的解读。只有frac40配置满足12/16检查点完整性要求。对于该配置,对相同数据应用事后标准得到的Hartigan-dip区间包含零(Delta dip的95%置信区间为[-0.017, 0.034]),而重叠标准差增大了约5.6倍。事后校准表明,dip检验在所模拟的分离水平上检验效能为零;因此该区间不具信息量,并非“无变化”的证据。标准差比率是此处唯一在观测效应下具有检验效能的统计量。集成损失仅在预设的1%阈值下接近平坦。最后,0/16、11/16和16/16的grokking比率仅具描述性意义,因为训练比例与数据划分身份存在混淆。
cs.LG / 36 / 2609.25645

Efficient Cost-Aware LLM Evaluation via Bayesian Bandit Gittins Indices

Xie, Qian, He, Yueli, Cao, Nairen
Abstract
Exhaustively evaluating every candidate LLM configuration on every benchmark item to identify a high-performing one is costly. We formulate configuration selection as a cost-aware Bayesian bandit problem and propose GittinsEval, which draws on the Bayesian-optimal Gittins policy to determine which configuration to evaluate next and when to stop. We extend the policy with an anytime recommendation rule over both fully and partially evaluated configurations, using an LCB-style score to account for posterior uncertainty. GittinsEval is computationally efficient, requiring only lightweight online updates after offline precomputation. Across GSM8K, PIQA, AlpacaEval, and MMLU response matrices, GittinsEval is consistently competitive, with particularly strong gains over configuration-level Bayesian optimization on large-example benchmarks and over cost-unaware bandit baselines on large-candidate tasks. Crucially, GittinsEval often attains near-zero simple regret using only 1% to 2% of the exhaustive-evaluation cost; it also offers an adaptive stopping rule that typically triggers at 1% to 10%.
cs.LG / 37 / 2609.25655

From Experts to Sub-experts: Fine-grained Parameter-Efficient Fine-Tuning for MoE LLMs

从专家到子专家:面向MoE大语言模型的细粒度参数高效微调
Tan, Zhentao, Liu, Chang, Liu, Yao, Wu, Yue, Ye, Jieping
Abstract
As large language models (LLMs) scale rapidly, dense full-parameter adaptation becomes increasingly expensive, motivating sparse and modular architectures such as Mixture-of-Experts (MoE) models. This shift raises a key question for parameter-efficient fine-tuning (PEFT): at what granularity should parameters be selected and updated? Existing PEFT methods such as LoRA operate on predefined weight matrices, while expert-level sparse tuning methods update entire selected experts. However, we observe that activated experts are internally sparse, with only a small fraction of intermediate channels strongly responding to downstream tasks, indicating that expert-level adaptation is still too coarse. We propose NSFT (Neural Sub-expert Fine-Tuning), a fine-grained PEFT framework that refines MoE adaptation from experts to sub-experts. NSFT decomposes each expert along the intermediate dimension into structured channel groups and selects task-relevant sub-experts by combining routing importance with intra-expert activation saliency. To optimize sparse partial updates, NSFT further introduces learning-rate scaling and dynamic gradient scaling to compensate for the reduced effective update magnitude. Experiments on OLMoE and Ling-mini-2.0 across challenging domain-specific tasks and general benchmarks show that NSFT consistently outperforms representative PEFT and expert-level sparse tuning baselines, while using substantially fewer trainable parameters and preserving competitive general capability. These results suggest that sub-expert-level adaptation is a more precise and efficient PEFT paradigm for MoE LLMs.
Chinese Translation
随着大语言模型(LLM)的快速扩展,稠密的全参数适配变得日益昂贵,这推动了混合专家(Mixture-of-Experts, MoE)模型等稀疏化、模块化架构的发展。这一转变为参数高效微调(PEFT)提出了一个关键问题:参数应当在何种粒度上被选择和更新?现有的PEFT方法(如LoRA)作用于预定义的权重矩阵,而专家级稀疏微调方法则更新整个被选中的专家。然而,我们观察到被激活的专家内部是稀疏的,只有一小部分中间通道对下游任务产生强烈响应,这表明专家级适配仍然过于粗糙。我们提出了NSFT(Neural Sub-expert Fine-Tuning,神经子专家微调),这是一个将MoE适配从专家细化到子专家的细粒度PEFT框架。NSFT沿着中间维度将每个专家分解为结构化的通道组,并通过结合路由重要性与专家内部激活显著性来选择与任务相关的子专家。为了优化稀疏的部分更新,NSFT进一步引入了学习率缩放和动态梯度缩放,以补偿有效更新幅度的减小。在OLMoE和Ling-mini-2.0上,针对具有挑战性的领域特定任务和通用基准的实验表明,NSFT在仅使用远少于基线方法可训练参数的情况下,持续优于代表性的PEFT方法和专家级稀疏微调基线,同时保持了具有竞争力的通用能力。这些结果表明,子专家级适配是面向MoE大语言模型的一种更加精确、高效的PEFT范式。
cs.LG / 38 / 2609.25657

Targeted Review for AI-Assisted Biodiversity Surveys: Active Continuous-Score Occupancy Modeling

面向AI辅助生物多样性调查的针对性审查:主动连续评分占用建模
Haucke, Timm, Harrell, Lauren, Kay, Justin, Clapp, Mary, Beery, Sara
Abstract
We increasingly use machine learning to label scientific datasets. The models we develop and deploy are improving all the time, but they are not and will likely never be perfect. Mistakes matter, as errors can propagate into our scientific understanding, particularly when systematically biased. Very reasonably, scientists thus review substantial proportions of ML-generated labels to verify or correct mistakes in pursuit of ensuring their scientific findings are not biased by ML. In this work, we focus on helping scientists optimally allocate this reviewing effort relative to their scientific goals. We focus on a specific class of scientists (ecologists) and a specific, widespread, and impactful modeling target (occupancy modeling, which estimates where species are likely to occur, conditioned on environmental factors). We introduce Active Continuous-Score Occupancy Modeling (ACORN), a method that incorporates ML predictions into occupancy models and strategically selects samples for expert review that are maximally informative for downstream ecological analysis. Across camera-trap and bioacoustic datasets, our method recovers ecological conclusions close to those obtained from fully human-labeled data, while requiring substantially fewer expert reviews than non-targeted review policies. Our results suggest that ML-assisted scientific workflows should optimize expert effort for downstream inference, rather than for classifier accuracy alone, especially when human review budget is limited. Our code is available at https://github.com/timmh/acorn
Chinese Translation
我们越来越多地使用机器学习为科学数据集打标签。我们所开发和部署的模型在不断改进,但它们并不完美,且可能永远无法完美。错误至关重要,因为误差可能会传播到我们的科学认知中,尤其是当误差具有系统性偏差时。因此,科学家们会合理地审查相当大比例的机器学习生成标签,以验证或纠正错误,确保科学发现不受机器学习偏差的影响。在这项工作中,我们致力于帮助科学家根据其科学目标最优地分配审查工作。我们聚焦于一类特定的科学家(生态学家)和一种特定、广泛且具有重要影响的建模目标(占用建模,即在给定环境因素的条件下估计物种可能出现的位置)。我们提出了主动连续评分占用建模(Active Continuous-Score Occupancy Modeling,ACORN),该方法将机器学习预测纳入占用模型,并策略性地选择对下游生态分析信息量最大的样本供专家审查。在相机陷阱和生物声学数据集上的实验表明,我们的方法能够获得接近完全由人工标注数据所得出的生态结论,同时所需的专业审查量远少于非针对性审查策略。我们的结果表明,机器学习辅助的科学工作流程应针对下游推断优化专家精力,而不仅仅以分类器准确率为目标,尤其是在人工审查预算有限的情况下。我们的代码可在 https://github.com/timmh/acorn 获取。
cs.LG / 39 / 2609.25659

When Riemann flows with Wasserstein: Generative Modeling of Probability Distributions on Manifolds

当黎曼流遇上Wasserstein:流形上概率分布的生成建模
Haviv, Doron, De Brouwer, Edward, Anand, Rishabh, Ying, Rex, Bentaieb, Aïcha, Scalia, Gabriele, Bravo, Hector Corrada
Abstract
Many scientific datasets, such as molecular conformational ensembles or single-cell tissue measurements, are naturally modeled as meta-distributions: distributions over probability measures on non-Euclidean domains. Existing generative methods largely assume Euclidean geometry and fail to capture this structure. We introduce Riemannian Wasserstein Entropic Flow Matching (RWEFM), a generative framework on the Wasserstein space $\mathcal{P}_2(\mathcal{M})$ of a Riemannian manifold $(\mathcal{M},g)$. RWEFM is trained by regressing a neural vector field onto Riemannian optimal transport velocities, using McCann displacement interpolations as conditional paths. We confirm theoretically that this construction leads to a valid flow matching approach on $\mathcal{P}_2(\mathcal{M})$ and introduce the Riemannian Entropic Map, a GPU-efficient approximation of the optimal transport map on manifolds. Our experiments show that by respecting the intrinsic geometry of the data, RWEFM can generate whole single-cell samples in hyperspherical latent spaces and protein conformational ensembles on the torus. As RWEFM requires only a geodesic distance and a projection operator, it is not restricted to manifolds with closed-form geometry, which we demonstrate by generating distributions on a general triangulated mesh.
Chinese Translation
许多科学数据集,如分子构象集合或单细胞组织测量数据,天然地可建模为元分布:即非欧几里得域上概率测度的分布。现有的生成方法大多假设欧几里得几何,无法捕捉这一结构。我们提出了黎曼Wasserstein熵流匹配(Riemannian Wasserstein Entropic Flow Matching, RWEFM),这是一个定义在黎曼流形 $(\mathcal{M},g)$ 的 Wasserstein 空间 $\mathcal{P}_2(\mathcal{M})$ 上的生成框架。RWEFM 通过将神经向量场回归到黎曼最优传输速度上进行训练,并以 McCann 位移插值作为条件路径。我们从理论上证实了这一构造可以导出在 $\mathcal{P}_2(\mathcal{M})$ 上有效的流匹配方法,并引入了黎曼熵映射(Riemannian Entropic Map),一种 GPU 高效的流形上最优传输映射近似方法。实验表明,通过尊重数据的内在几何结构,RWEFM 能够在超球面潜空间中生成完整的单细胞样本,并在环面上生成蛋白质构象集合。由于 RWEFM 仅需要测地距离和一个投影算子,它并不局限于具有封闭形式几何结构的流形——我们通过在一般三角网格上生成分布展示了这一点。
cs.LG / 40 / 2609.25675

Marginal Log-Likelihood Increments under Dirichlet-Smoothed Markov Estimation

Schwab, Levin David
Abstract
For a Dirichlet-smoothed transition model, the effect of adding one workflow trace to the training archive is an exact change in reference-weighted log likelihood. We derive that change and show that it is a weighted reduction of Kullback--Leibler divergence between the reference conditionals and the model. From this form we obtain an upper bound on the gain available to any acquisition, which expresses a millinat difference as a share of what is attainable, an exact covariance identity for the effect of the reference weighting, and a sign criterion for the interaction between two candidates, from which the batch objective is neither submodular nor supermodular. A case study on the BPI Challenge 2012 loan-application log measures all three and finds a positive selection result in one of the four combinations of reference weighting and budget unit. There, of two regressors fitted to identical descriptors and identical labels, the one that predicts individual increments far more accurately, median $R^2$ 0.87 against 0.62, realizes the smaller share of the attainable gain, 61 against 69 per cent, so ranking accuracy for individual traces is neither necessary nor sufficient for batch quality.
cs.LG / 41 / 2609.25692

Graph Domain Adaptation Does Not End with Representation Learning

Liu, Ziqian, Xu, Yongxue, Zhang, Enze, Zhang, Jiaqi, Wang, Hao, Wang, Maolin
Abstract
Graph domain adaptation (GDA) transfers knowledge from a labeled source graph to an unlabeled target graph under shifts in both node attributes and graph structure. Existing methods primarily adapt graph representations through propagation redesign, distribution alignment, or source-to-target transition modeling, but still rely on a single graph-propagating path for target prediction. This leaves open whether an adapted graph representation exhausts the predictive evidence available in the target domain, since the graph-aware expert and graph-free local expert may exhibit different failure modes under topological shifts. To address this limitation, we propose EviGDA, an Evidence-Augmented Graph Domain Adaptation framework that complements graph representation adaptation with a graph-free local expert. The graph-aware expert performs message passing and entropy-aware marginal alignment, while the graph-free local expert learns solely from source node features and labels without graph propagation or target alignment. The two experts are optimized independently and combined only at inference through a task-level constant probability mixture, preserving complementary evidence without joint training, learned routing, or target pseudo-labels. Extensive experiments on ten datasets and 16 transfer tasks show that EviGDA outperforms state-of-the-art baselines.
cs.LG / 42 / 2609.25701

Fully Byzantine-Resilient Multi-Agent Reinforcement Learning

完全拜占庭容错的多智能体强化学习
Lee, Haejoon, Panagou, Dimitra
Abstract
We study distributed Byzantine-resilient actor-critic multi-agent reinforcement learning (AC-MARL), where agents collectively learn policies through local interactions. Existing methods guarantee convergence of the agents' parameters only to a neighborhood of the attack-free limit points, resulting in degraded performance. We propose Fully Resilient AC-MARL (FRAC-MARL), a decentralized method in which each agent leverages redundancy in two-hop messages to identify reliable messages. Under linear parameterizations of the value and team-reward functions and Byzantine edge attacks, where adversarial behavior is confined to the communication layer, we prove that agents' parameters converge almost surely to the same limit points as in the attack-free case over time-varying communication graphs. We introduce a novel topological condition for the convergence of our method, present a systematic method to construct such networks, and prove that this condition can be verified in polynomial time. Finally, we demonstrate our method on cooperative multi-robot formation control tasks.
Chinese Translation
我们研究了分布式拜占庭容错的演员-评论家多智能体强化学习(AC-MARL),其中智能体通过局部交互共同学习策略。现有方法仅能保证智能体的参数收敛到无攻击情形下极限点的某个邻域,从而导致性能下降。我们提出了完全弹性AC-MARL(FRAC-MARL),这是一种去中心化方法,其中每个智能体利用两跳消息中的冗余来识别可靠消息。在价值函数和团队奖励函数的线性参数化以及拜占庭边攻击(对抗行为仅限于通信层)的条件下,我们证明了智能体的参数在时变通信图上几乎必然收敛到与无攻击情形相同的极限点。我们为该方法的收敛性引入了一个新颖的拓扑条件,提出了一种系统化的网络构建方法,并证明该条件可在多项式时间内验证。最后,我们在协作多机器人编队控制任务上演示了该方法。
cs.LG / 43 / 2609.25721

Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages

缓慢衰减与沉默表达:语言模型谱系中的迭代隐性特质传递
Vo, Ryan, Nguyen, Duc-Vu, Kretchmar, Matt, Nguyen, Ngan Luu-Thuy
Abstract
Language models are increasingly trained on the outputs of other models, forming chains that we call lineages, in which a trait present in one generation can pass to the next. Prior work on subliminal learning has shown that a teacher's trait can transmit to a student through filtered data carrying none of the trait's content. However, the evidence covers only a single training step. We study whether such a trait holds or fades across lineages. We instill the trait into three copies of Qwen2.5-7B-Instruct and iterate the training step to depth ten from each, reading every generation two ways on the same held-out prompts: a keyword screen that looks for expressions of the trait in the model's output, and an activation probe that projects each model's displacement from the base onto a direction built from the other lineages' teachers. We report two findings. First, the trait persists through ten generations across three lineages. The instilled models express it on every completion; the keyword-screen rate falls to 55.6% after the first step and to 21.1% by generation ten. The base itself matches the screen on none of its 300 completions. Second, the trait can be present internally while absent behaviorally. When the model's default system prompt is removed at evaluation, the generation-ten students' keyword-screen rate is zero on every prompt while the probe score stays positive on every prompt. Steering the untreated base with the displacement of a generation-ten student, which is trained and measured under the default system prompt, induces screened expression of the trait even with the system prompt removed, while that same student shows no expression of the trait with the system prompt removed.
Chinese Translation
语言模型越来越多地在其他模型的输出上进行训练,形成我们称之为谱系(lineages)的链条,其中某一世代存在的特质可以传递到下一世代。先前关于隐性学习(subliminal learning)的研究表明,教师的特质可以通过不携带该特质任何内容数据的过滤数据传递给学生。然而,已有证据仅覆盖单一训练步骤。我们研究这种特质在谱系中是保持还是消退。我们将该特质植入 Qwen2.5-7B-Instruct 的三个副本,并从每个副本出发将训练步骤迭代至深度十,在相同的保留提示集上以两种方式读取每一代模型:一种是关键词筛查,用于在模型输出中寻找该特质的表达;另一种是激活探针,将每个模型相对于基座模型的位移投影到一个由其他谱系教师构建的方向上。我们报告两个发现。第一,该特质在三个谱系中均能持续存在十代。被植入特质的模型在每次补全中都会表达它;关键词筛查率在第一步之后降至55.6%,到第十代降至21.1%。而基座模型在其全部300次补全中没有任何一次通过关键词筛查。第二,该特质可以内部存在而行为上缺席。当在评估时移除模型的默认系统提示后,第十代学生的关键词筛查率在所有提示上均为零,而探针得分在所有提示上仍然为正。使用在默认系统提示下训练和测量的第十代学生的位移对未处理的基座模型进行导向(steering),即使移除系统提示也能诱导出通过筛查的特质表达,而该学生本身在移除系统提示时并不表现出该特质。
cs.LG / 44 / 2609.25722

Signed Graph Pre-Training and Prompt Learning

符号图预训练与提示学习
Mei, Zihan, Pan, Rong, Chen, Yuzhou, He, Yixuan
Abstract
Signed graphs arise in trust--distrust networks, financial correlation systems, biological interaction graphs, and many other domains in which edges can be positive or negative and may also be directed. While signed graph neural networks have improved task-specific learning, graph transfer learning on signed graphs remains underdeveloped. In this paper, we introduce TopoSIGN, a pioneer topology-guided graph pre-training and prompt learning framework for signed graphs. TopoSIGN combines a structural encoder built on the magnetic signed Laplacian with a novel persistent-homology branch that summarizes signed topology through Dowker-complex persistence images. The fused embeddings are then transferred to a prompt learning function. Experimental results on synthetic and real-world datasets demonstrate the efficacy of TopoSIGN in extracting useful structural information in signed graphs, as well as the adaptability and flexibility of the proposed general framework.
Chinese Translation
符号图广泛存在于信任—不信任网络、金融关联系统、生物相互作用图以及许多其他领域中,其中边可以为正或负,且可能是有向的。尽管符号图神经网络(signed graph neural networks)已在特定任务的学习中取得改进,但符号图上的图迁移学习仍然发展不足。本文提出了TopoSIGN,一个开创性的面向符号图的拓扑引导图预训练与提示学习框架。TopoSIGN将基于磁符号拉普拉斯算子(magnetic signed Laplacian)构建的结构编码器与一个新颖的持续同调(persistent-homology)分支相结合,该分支通过Dowker复形持续性图像(Dowker-complex persistence images)来概括符号拓扑信息。融合后的嵌入随后被迁移至提示学习函数。在合成数据集和真实世界数据集上的实验结果表明,TopoSIGN能够有效提取符号图中有用的结构信息,同时验证了所提出的通用框架的适应性与灵活性。
cs.LG / 45 / 2609.25728

Self-Supervised Combinatorial Optimization with Constraints via Frank-Wolfe

Rafiey, Akbar, Xu, Yifei, Karalias, Nikolaos
Abstract
Self-supervised learning for combinatorial optimization has emerged as a promising paradigm for solving discrete optimization problems with neural networks, but a central challenge remains: handling hard combinatorial constraints within continuous, gradient-based training. Continuously extending combinatorial objectives to convex domains is a powerful technique, yet existing approaches often require projection steps that constrain neural network outputs to lie inside the feasible polytope and rely on ad-hoc and problem-specific constructions. We propose a general framework in which the neural network is allowed to predict arbitrary continuous vectors that could potentially lie outside of the feasible polytope. These predictions are then approximated by sparse convex combinations of feasible solutions using a geometric decomposition algorithm based on Frank--Wolfe methods and approximate Caratheodory results. This decomposition induces an a.e.-differentiable, self-supervised loss defined as the expected value of the discrete objective. The same procedure provides an automatic rounding guarantee at inference time. We demonstrate strong empirical performance across multiple combinatorial problems, including the Quadratic Assignment Problem, Maximum Coverage, and the Traveling Salesperson Problem.
cs.LG / 46 / 2609.25735

Beyond Class Marginals: Bounding Rehearsal Gaps without Freezing Class Co-occurrence

Dai, Congren, Roongjirarat, Nat, Ye, Fei
Abstract
Class-balanced replay controls class frequency but does not determine the interval between successive replay appearances of a class. We study this interval, the rehearsal gap, separately from the class marginal and class co-occurrence, and introduce randomised-pass replay (RPR), which visits each resident class once per shuffled pass. For a fixed set of C resident classes and replay batch size b less than or equal to C, RPR preserves the balanced time-averaged class marginal and bounds every gap by 2*ceil(C/b)-1; a churn-conditional bound applies while the resident set changes. The scheduler uses no future class information and adds no replay examples or forward passes. In a linear-head ER-ACE diagnostic, joint absence from the incoming and replay batches produces a one-sided classifier-bias gradient. Longer absence episodes are associated with larger negative bias displacement, and removing the incoming-loss mask attenuates the scheduling effect. In the primary ER-ACE experiments, RPR improves final average accuracy by 0.72-1.67 percentage points relative to independent class-balanced retrieval under reservoir storage, with positive effects also observed under balanced storage. Pretrained ViTs show positive effects on the tested LT10 streams with small replay batches, while matched larger-batch controls show no material effect. Fixed-cycle and reused-pass controls change more than one temporal statistic, so the experiments do not isolate rehearsal-gap length from all other forms of temporal dependence. The accuracy effects depend on the learner and operating regime.
cs.LG / 47 / 2609.25745

Modular Norm RandOpt: Population-Efficient Ensembling through Architecture-Aware Perturbations

模块化范数RandOpt:通过架构感知扰动实现种群高效集成
Yoshihara, Kirato, Hamade, Hiroaki
Abstract
RandOpt samples weight-perturbed language models and ensembles top-ranked candidates through plurality voting, but its global perturbation scale ignores heterogeneous module geometry. We propose \mbox{\textbf{\emph{Modular Norm RandOpt}}}, an architecture-aware sampling method using module-wise natural norms and calibrated scales while preserving selection and voting. It outperforms RandOpt using $3\times$ fewer candidates on Countdown and at least $12\times$ fewer on GSM8K, with corresponding wall-clock savings. Evaluations across seven tasks and three Qwen scales ($0.5$B--$3$B) show higher mean accuracy than RandOpt on Countdown, GSM8K, and MATH-500 at every scale. The gains extend to Llama 3.2 $3$B and Gemma 3 $4$B on Countdown and GSM8K. On Qwen2.5-1.5B, our ensembles also achieve higher mean accuracy than iterative baselines on both tasks at comparable main-run evaluation budgets. On GSM8K, a tail-density diagnostic implies only a $1.2$--$1.8\times$ candidate reduction, while most ensemble improvement is associated with more favorable correct-expert support. These results highlight perturbation geometry as a key design choice for population-efficient, gradient-free search around pretrained models.
Chinese Translation
RandOpt通过采样权重扰动的语言模型并采用多数投票对排名靠前的候选模型进行集成,但其全局扰动尺度忽略了异构的模块几何结构。我们提出模块化范数RandOpt(Modular Norm RandOpt),一种架构感知的采样方法,它使用模块级自然范数和校准后的尺度,同时保留原有的选择与投票机制。该方法仅使用RandOpt三分之一的候选数量即可在Countdown任务上超越RandOpt,在GSM8K上至少减少12倍的候选数量,并相应节省了实际运行时间。在七个任务和三个Qwen规模(0.5B–3B)上的评估显示,在所有规模下,该方法在Countdown、GSM8K和MATH-500上的平均准确率均高于RandOpt。这些增益还可扩展至Llama 3.2 3B和Gemma 3 4B模型(在Countdown和GSM8K任务上)。在Qwen2.5-1.5B上,在相当的主运行评估预算下,我们的集成方法在两个任务上也取得了高于迭代式基线方法的平均准确率。在GSM8K上,尾部密度诊断表明候选数量仅可减少1.2–1.8倍,而集成性能的大部分提升与更有利的正确专家支持相关。这些结果凸显了扰动几何结构作为在预训练模型周围进行种群高效、无梯度搜索的关键设计选择。
cs.LG / 48 / 2609.25757

Minimal Recurrent Behavioral Memory for Imitation under Partial Observability

部分可观测性下模仿学习的最小循环行为记忆
Li, Xianyao, Xu, Fang, Min, Rui, Tian, Ruitong, Du, Jing
Abstract
What is the least recurrent memory needed to reproduce a specified expert under partial observability? The instantaneous requirement is the conditional entropy of the expert's behavioral quotient, but recurrence must also preserve distinctions that future observations will not restore before use. We characterize this minimal recurrent behavioral memory by a compatibility relation: under transitivity its classes attain the exact minimum, while the general case is an entropy minimization over closed compatible state assignments, with exact certificates on finite instances. A sole-carrier measurement protocol separates behavioral sufficiency, excess code rate, and information carried by observations or other memory paths; experimental bit requirements refer to the induced symbolic behavioral model under the stated occupancy. Across manipulation tasks, learned code rates remain near zero- and two-bit requirements as hidden modes grow to $512$, and anticipatory memory follows a $2\to1\to0$ requirement despite zero instantaneous demand during waiting. Learning this representation remains difficult: event-agnostic future-behavior supervision yields $36/40$ sufficient seeds with one frozen configuration and improves the longest-horizon pixel setting from $0/8$ to $6/8$ sufficient held-out seeds (closed-loop success from $0.08$ to $0.57$). On unmodified community benchmarks, the protocol certifies delay-independent requirements, which sufficient codes match at mid-delay. The supervision aids commitment but can induce predictive surplus; annealing it lets imitation and rate training reduce that surplus, separating the information-theoretic target from the ability to learn it.
Chinese Translation
在部分可观测性下,重现指定专家策略所需的最小循环记忆是多少?瞬时需求等于专家行为商的条件熵,但循环记忆还必须保留那些在后续观测被使用之前无法恢复的区分。我们通过一种兼容性关系来刻画这种最小循环行为记忆:在传递性条件下,其等价类可达到精确最小值;而在一般情况下,则转化为对封闭兼容状态分配的熵最小化问题,并在有限实例上给出精确的证明证书。一个单一载体的测量协议将行为充分性、冗余码率以及由观测或其他记忆路径携带的信息分离开来;实验中的比特需求指的是在给定占据测度下诱导出的符号化行为模型。在多个操作任务中,随着隐藏模态数增长至512,学习到的码率始终接近零比特和两比特的需求;尽管等待期间的瞬时需求为零,前瞻性记忆仍遵循 2→1→0 的需求变化。学习这种表示仍然困难:与具体事件无关的未来行为监督在单一冻结配置下产生了36/40个充分种子,并将最长时域像素设置下的充分保留种子数从0/8提升到6/8(闭环成功率从0.08提升到0.57)。在未修改的社区基准上,该协议认证了与时延无关的需求,且充分码在中时延下与之匹配。这种监督有助于承诺(commitment),但可能引入预测冗余;对其采用退火方法,可以让模仿和码率训练降低该冗余,从而将信息论目标与学习该目标的能力区分开来。
cs.LG / 49 / 2609.25777

Disentangling Heterogeneous Traffic Dynamics for Multi-Step Traffic Forecasting via Adaptive Spectral Decomposition

基于自适应谱分解的多步交通预测中异质交通动态解耦方法
Huang, Zijun, Fu, Chenrui, Wang, Wenhao, Gou, Xiaochuan, Hung, Chih-Chieh, Li, Guanyao
Abstract
Accurate multi-step traffic forecasting remains challenging because observed traffic signals contain heterogeneous temporal dynamics with different characteristics and levels of predictability. Existing approaches typically model these dynamics within a unified representation or rely on predefined decomposition rules, which may limit their ability to flexibly separate persistent patterns from rapidly varying fluctuations. To address this issue, we propose the Adaptive Decomposition Network (ADNet), a component-specific forecasting framework that adaptively disentangles traffic dynamics into dominant and residual components. ADNet introduces a learnable complementary spectral decomposition mechanism that determines the contribution of each frequency bin to the two components. Unlike hard frequency partitioning, every frequency bin can contribute to both components with different learned proportions, allowing the decomposition to be optimized jointly with the forecasting objective. The reconstructed components are then modeled by two dedicated spatiotemporal forecasting branches, and their predictions are integrated to generate the final multi-step forecast. Experiments on the Alameda and Orange regions of the TraffiDent dataset show that ADNet achieves the best performance in 20 of the 24 reported region-horizon-metric comparisons, with particularly clear gains at longer forecasting horizons. Capacity-controlled ablation experiments further show that the learnable decomposition substantially outperforms a fixed decomposition and provides additional improvements beyond the dual-branch architecture alone. These results demonstrate the effectiveness of adaptive decomposition and component-specific modeling for multi-step traffic forecasting.
Chinese Translation
准确的多步交通预测仍然具有挑战性,因为观测到的交通信号包含具有不同特征和可预测性水平的异质时间动态。现有方法通常在统一表示中建模这些动态,或依赖预定义的分解规则,这可能限制其灵活地将持续性模式与快速变化的波动分离的能力。为解决这一问题,我们提出自适应分解网络(Adaptive Decomposition Network, ADNet),这是一种面向分量的预测框架,能够自适应地将交通动态解耦为主导分量和残差分量。ADNet 引入了一种可学习的互补谱分解机制,用于确定每个频率分量对两个成分的贡献。与硬性的频率划分不同,每个频率分量可以以不同学习到的比例同时贡献于两个成分,使分解能够与预测目标联合优化。重构后的分量随后由两个专用的时空预测分支分别建模,其预测结果被融合以生成最终的多步预测。在 TraffiDent 数据集的 Alameda 和 Orange 区域上的实验表明,ADNet 在 24 项区域-预测步长-指标比较中有 20 项取得了最佳性能,且在较长预测步长上的提升尤为显著。容量受控的消融实验进一步表明,可学习分解显著优于固定分解,并且在双分支架构之外提供了额外的性能提升。这些结果证明了自适应分解与面向分量建模在多步交通预测中的有效性。
cs.LG / 50 / 2609.25781

A Lightweight Plastic-Memory Framework for Graph Few-Shot Class-Incremental Learning

一种面向图少样本类增量学习的轻量级可塑性记忆框架
Mei, Zihan, Qin, Zhili, Zhang, Tongze, Liu, Hongyuan, Shao, Junming, Yang, Qinli
Abstract
Graph Incremental Learning has garnered increasing attention as dynamic graph data continues to emerge across diverse fields. Conventional approaches primarily address catastrophic forgetting by preserving node-related knowledge through replay or distillation techniques; however, they often incur high computational costs and inefficiency. This issue is further exacerbated in real-world scenarios where labeled data for new classes is scarce. In this paper, we propose a novel lightweight plastic-memory framework specifically designed for few-shot incremental learning on graphs. The core idea of our framework is the construction of a plastic-memory module that evolves over time, continuously updating and expanding its memory to accommodate new classes while retaining previously learned knowledge. In contrast to existing techniques, our memory module is both lightweight and effective, featuring an innovative evolving micro-clustering structure that dynamically updates representations of class prototypes, sub-prototypes, and their interaction weights. Building on this memory module, we introduce a memory-driven meta-learning framework that enhances adaptability to new tasks in its inner loop while maintaining stability for earlier tasks in the outer loop. Extensive experiments on four benchmark datasets demonstrate the framework's superior performance in balancing stability for old knowledge and adaptability to new knowledge.
Chinese Translation
随着动态图数据在各个领域不断涌现,图增量学习(Graph Incremental Learning)受到越来越多的关注。传统方法主要通过回放(replay)或知识蒸馏(distillation)技术来保留与节点相关的知识,从而应对灾难性遗忘问题;然而,这些方法往往计算成本高且效率低下。在真实场景中,新类别的标注数据稀缺,这一问题进一步加剧。本文提出了一种新颖的轻量级可塑性记忆框架,专门用于图上的少样本增量学习。该框架的核心思想是构建一个随时间演化的可塑性记忆模块,通过持续更新和扩展其记忆来容纳新类别,同时保留先前学到的知识。与现有技术相比,我们的记忆模块既轻量又高效,其创新之处在于采用了一种动态演化的微聚类结构,能够实时更新类原型、子原型及其交互权重的表示。基于该记忆模块,我们进一步提出了一个记忆驱动的元学习框架,其内循环增强了对新任务的适应性,而外循环则保持了对早期任务的稳定性。在四个基准数据集上的大量实验表明,该框架在平衡旧知识的稳定性与新知识的适应性方面表现出卓越的性能。
cs.LG / 51 / 2609.25788

Evaluating Accuracy and Probabilistic Reliability of Zero-Shot Time Series Foundation Models

零样本时间序列基础模型的精度与概率可靠性评估
Michael, Panagiotis, Symeonides, Moysis, Trihinas, Demetris
Abstract
Time Series Foundation Models (TSFMs) promise a paradigm shift toward zero-shot forecasting by eliminating task-specific training. However, existing works often overlook trade-offs between predictive accuracy and probabilistic calibration. This paper presents a benchmark study of six TSFMs evaluated on energy, traffic, and financial datasets. We contrast their performance against statistical baselines and a supervised DL model. The study reveals that while TSFMs outperform statistical methods and supervised models, they are subject to a fundamental trade-off between point accuracy and probabilistic reliability. Specifically, xLSTM architectures provide robust probabilistic calibration across horizons. In contrast, patch-based transformers offer competitive accuracy but face calibration issues at long horizons, while transformer-based models exhibit context saturation points for optimal zero-shot reasoning. These findings offer evidence-based guidance for balancing generalization and uncertainty quantification in real-world deployments.
Chinese Translation
时间序列基础模型(TSFMs)承诺通过消除任务特定训练,实现向零样本预测的范式转变。然而,现有研究往往忽视了预测精度与概率校准之间的权衡。本文对六种TSFMs在能源、交通和金融数据集上进行了基准评估研究。我们将它们的性能与统计基线方法及有监督深度学习模型进行对比。研究揭示,尽管TSFMs优于统计方法和有监督模型,但它们面临点预测精度与概率可靠性之间的根本性权衡。具体而言,xLSTM架构在各种预测时间范围(horizon)上均能提供稳健的概率校准;相比之下,基于patch的Transformer模型虽然精度具有竞争力,但在长预测范围时面临校准问题;而基于Transformer的模型则表现出用于最优零样本推理的上下文饱和点。这些发现为现实世界部署中平衡泛化能力与不确定性量化提供了基于实证的指导。
cs.LG / 52 / 2609.25802

Latest Exact Match Attention

最新精确匹配注意力机制
Brösamle, Moritz
Abstract
We introduce latest exact match attention (LEMA), an attention variant for transformers where queries and keys are binarized and each query attends only to the latest exactly matching key. We prove that LEMA transformers with chain of thought can simulate word-RAMs, as was recently shown for the less restrictive rightmost hard attention. In contrast to prior hard attention variants, the restriction to exact matches enables an efficient converse direction: word-RAMs can simulate LEMA transformers at a cost per token independent of the context length. Together, these results yield a close correspondence between the two computational models in terms of both compute and memory. Beyond the theory, we propose a training method for LEMA transformers that handles their non-differentiable operations with a straight-through estimator for the binarization and a soft attention surrogate annealed towards LEMA. On a synthetic associative recall task, LEMA models trained this way use their growing state to store and recall a large number of associations, outperforming gated DeltaNet (GDN) with its fixed state size. As a first scaling test, we train LEMA language models with up to 834 million parameters. They match softmax transformers of around half their size in loss and, on repeated rare phrases and a needle-retrieval task, remain behind softmax transformers but recall across longer distances than GDN models of comparable size. Finally, we implement dictionary-based inference for LEMA transformers and show constant generation speed comparable to GDN despite their growing state, with the dictionaries residing in main memory rather than VRAM. Code is available at https://github.com/moritzbroe/latest_exact_match_attention.
Chinese Translation
我们提出了最新精确匹配注意力(Latest Exact Match Attention, LEMA),这是一种Transformer的注意力变体,其中查询和键被二值化,且每个查询仅关注最新的完全匹配的键。我们证明,带有思维链(chain of thought)的LEMA Transformer可以模拟字随机存取机(word-RAM),这与最近针对限制更少的最右侧硬注意力所证明的结果类似。与以往的硬注意力变体不同,对精确匹配的限制使得反向模拟也能高效进行:word-RAM可以模拟LEMA Transformer,且每个token的计算成本与上下文长度无关。这些结果共同表明,这两种计算模型在计算和内存方面具有紧密的对应关系。除理论之外,我们提出了一种LEMA Transformer的训练方法:通过直通估计器(straight-through estimator)处理二值化操作,并使用一种逐步退火至LEMA的软注意力替代方案来处理不可微操作。在一个合成联想回忆任务中,以这种方式训练的LEMA模型利用其不断增长的状态来存储和回忆大量关联,其表现优于具有固定状态大小的门控DeltaNet(GDN)。作为初步的规模测试,我们训练了参数量高达8.34亿的LEMA语言模型。它们在损失上与约一半大小的softmax Transformer相当;在重复稀有短语和“大海捞针”检索任务上,虽然仍落后于softmax Transformer,但其召回距离优于同等规模的GDN模型。最后,我们实现了LEMA Transformer的基于词典的推理,尽管其状态不断增长,但其生成速度保持恒定,与GDN相当,且词典存储在主内存而非显存(VRAM)中。代码可在 https://github.com/moritzbroe/latest_exact_match_attention 获取。
cs.LG / 53 / 2609.25808

Auditing Proxy-Based Validation Across Text Spans

Weon, Daein, Kang, Dong Ho
Abstract
Evaluation scores are often validated by their agreement with inexpensive proxy labels. When the score and the proxy are computed from the same text span, however, that agreement can arise from surface evidence the two share rather than from the semantic construct the proxy is meant to represent. We make the distinction explicit by declaring the score, its span, the proxy and the target construct as a validation contract, then re-evaluating that proxy rule strictly outside the scored span. In a controlled HotpotQA correctness experiment varying only the shared text boundary, the score agrees with its proxy far better than with correctness at a 50-character prefix: the gap is +0.184, collapsing to at most +0.045 from 120 characters onward. At that short prefix the score still predicts whether the answer string appears later (AUC 0.634) while an equivalence test places its agreement with correctness at chance, so the reported proxy agreement does not establish that the score ranks correctness. On OR-Bench, suppressing each model's recurring opening templates removes most of the score's association with the refusal proxy, while matched-volume deletion removes almost none and construct agreement stays at chance. Only three of eleven external contracts support the off-span control, and none of the routing studies we sampled released the generations it needs. We therefore ask that a proxy-based validation claim declare the span each label is read from, report the construct agreement beside the proxy agreement, and release the generations that let the proxy be re-read off the scored span.
cs.LG / 54 / 2609.25809

You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs

Chen, Yuanteng, Lai, Qiwei, Tianqi, Chen, Wang, Peisong, Shao, Yuantian, Zeng, Nanxin, Liu, Zhilei, Li, Chuangyi, Liu, Jing, Cheng, Jian
Abstract
Fine-grained mixture-of-experts (MoE) architectures have become a mainstream design for open-weight LLMs, with hundreds of experts and increasingly many selected per token. This shift makes dynamic expert pruning an attractive route to cheaper inference. Yet existing evidence comes largely from coarser architectures and likelihood-scored multiple-choice benchmarks, leaving three central questions open in the fine-grained regime: how redundant per-token expert selection is, how effectively existing pruning methods exploit that redundancy, and what governs a model's sensitivity to pruning. We fill this gap with a systematic empirical study of twelve fine-grained MoE checkpoints spanning nine architecture families, with a core suite of eleven benchmarks covering knowledge QA, mathematics, code generation, and general reasoning. We find that expert selection is far more redundant than the field's operating points assume: uniformly retaining about two thirds of the selected experts preserves 98.8% of unpruned performance on average, requiring only a one-integer change and delivering 1.2-1.7x measured speedup across two serving backends. This simple baseline leaves little room for dynamic allocation at conservative budgets: even the best published rules differ from it by under 1% at matched expert budgets. Their value emerges under aggressive pruning, where the best rules recover up to 3.0% over uniform truncation, with gains concentrated in the generative tasks that suffer the sharpest degradation. Sensitivity to aggressive pruning also depends on the model: larger and thinking models are more resilient, whereas multimodal models are more vulnerable. Together, these findings reveal how much expert computation fine-grained MoEs can dispense with, and establish when dynamic allocation earns its complexity, informing both practical deployment and future pruning methods.
cs.LG / 55 / 2609.25811

Multi-View Fair Clustering Guided by Cross-View Sensitive Information Discrepancy

Jiang, Mudi, Zhou, Jiahui, Liu, Xinying, He, Zengyou, Chen, Zhikui
Abstract
Multi-view clustering (MVC) aims to uncover latent cluster structures by exploiting complementary information from multiple views. Despite substantial progress in clustering performance, fairness remains an important concern when MVC is applied to socially sensitive scenarios. Recent fair multi-view clustering methods have introduced fairness constraints into representation learning or clustering assignments. However, these methods generally treat different views under a largely uniform fairness mechanism, without explicitly distinguishing their varying levels of sensitive dependence during cross-view learning. In practice, different views may encode substantially different levels of sensitive information. Ignoring such cross-view discrepancy can allow highly sensitive-dependent views to influence less sensitive-dependent ones during cross-view learning, potentially degrading both clustering performance and fairness. To address this issue, we propose a novel multi-view fair clustering framework guided by cross-view sensitive information discrepancy. Specifically, we estimate the sensitive dependence of each view and develop a bias-ranked asymmetric alignment mechanism that encourages views with higher sensitive dependence to learn from those with lower sensitive dependence, while cross-view discrepancies are further exploited to adaptively regulate the alignment process. Moreover, fairness regularization is imposed on the consensus soft assignments to further promote group fairness. Extensive experiments on benchmark datasets demonstrate that the proposed method achieves a favorable balance between clustering quality and group fairness.
cs.LG / 56 / 2609.25814

CacheDyG: Decoupling Temporal Propagation for Efficient Dynamic Graph Learning

Zong, PinHeng, Yuan, Ye
Abstract
Dynamic graphs are widely used to model time-evolving relational systems in real-world applications. Dynamic graph neural networks provide an effective framework for capturing both structural dependencies and temporal dynamics in such data. However, they typically intertwine temporal graph propagation with every optimization epoch and often maintain large trainable representations for each node-time pair. This design repeatedly recomputes largely unchanged historical structures, leading to substantial training and parameter overhead. To address this critical issue, we propose CacheDyG, a Cache-refine framework for efficient Dynamic Graph learning. Specifically, it decouples temporal propagation from routine parameter updates by constructing a time-ordered temporal dependency cache that stores graph-aware node-time representations in non-trainable buffers. During standard training epochs, CacheDyG reads from the cache and updates only a lightweight cache refiner, an adaptive residual gate, and the link predictor. Selective cache refresh further keeps cached representations aligned with the supervised objective while avoiding epoch-wise sparse propagation. Experiments on five dynamic graph benchmarks show that CacheDyG adopts substantially fewer trainable parameters and lower runtime to obtain more competitive predictive performance than baselines. These results demonstrate that cache-based decoupling provides an effective principle for scalable dynamic graph learning.
cs.LG / 57 / 2609.25827

Protocol before progress: leakage-aware evaluation of AIS trajectory prediction

进步先于评估协议:面向AIS轨迹预测的泄漏感知评估
Raisi, Zobeir, Had, Vali Mohammad Nazarzehi
Abstract
Reported gains in vessel-trajectory prediction from Automatic Identification System (AIS) data are credited to new architectures, but the evaluation protocol is rarely measured as a source of error reduction. We build a leakage-aware protocol with vessel-, time- and region-disjoint splits and apply it to two corpora with different traffic: 31 days of Danish national AIS traffic and 30 days of US Gulf coast traffic off Houston and Galveston. On both, we audit TrAISformer, GATransformer, and controlled AISFormer-inspired reconstructions. Three protocol effects appear in both corpora. First, TrAISformer's best-of-16 oracle decoder lowers error by a factor of 2.1-3.2 relative to greedy decoding. Second, a split that shares vessels lowers its greedy error by 23-25% at one hour, against 2% or less for a compact 0.43 M-parameter encoder. Third, a region-disjoint split raises TrAISformer's one-hour error from 2.2 to 24.6 km on the US corpus, because 99.9% of the test contexts fall in longitude bins never seen in training; the encoder built on local offsets is unaffected by this. Architectural mechanisms matter less: GATransformer's graph attention gives no measurable benefit on either corpus, while its waterway feature is worth 12-22%. The effect of a time-disjoint split is not stable across corpora (13% versus 2%). We release the splits and code.
Chinese Translation
基于船舶自动识别系统(AIS)数据的船舶轨迹预测所报告的性能提升通常被归功于新架构,但评估协议本身很少被作为误差来源加以衡量。我们构建了一个泄漏感知的评估协议,采用船舶、时间和区域互斥(disjoint)的数据划分,并将其应用于两个交通特征不同的语料库:31天的丹麦全国AIS交通数据和30天的美国休斯顿与加尔维斯敦附近墨西哥湾沿岸交通数据。在两个语料库上,我们对TrAISformer、GATransformer以及受控的AISFormer风格重建模型进行了审计。两个语料库中均出现了三种协议效应。第一,TrAISformer的best-of-16 oracle解码器相对于贪婪解码可将误差降低2.1至3.2倍。第二,共享船舶的数据划分使其一小时的贪婪解码误差降低23-25%,而对于一个紧凑的0.43 M参数编码器,降幅仅为2%或更少。第三,区域互斥划分使TrAISformer在美国语料库上的一小时误差从2.2公里升至24.6公里,原因是99.9%的测试上下文落在训练中从未出现过的经度区间内;而基于局部偏移构建的编码器则不受此影响。架构机制的影响相对较小:GATransformer的图注意力在两个语料库上均无可测量的收益,而其航道特征则带来12-22%的价值。时间互斥划分的效果在语料库间并不稳定(13%对2%)。我们公开了数据划分和代码。
cs.LG / 58 / 2609.25836

In-Context Guidance: Learning Inter-Task Synergies via Numerical Foundational Models for Few-Shot Multitask Optimization

上下文引导:基于数值基础模型学习任务间协同关系的少样本多任务优化
Wei, Tingyang, Wu, Haofeng, Liu, Jiao, Wei, Zhao, Tan, Puay Siew, Ong, Yew-Soon
Abstract
Multi-task optimization (MTO) addresses a set of optimization tasks simultaneously, often suffering from inaccurate inter-task relationship estimation under limited evaluation budgets, leading to negative transfer. This paper introduces In-Context Guidance Multitask Optimization (ICG-MTO), a novel framework that leverages numerical foundational models to improve inter-task coupling estimation in few-shot scenarios. Unlike conventional methods that rely solely on scarce observed data, ICG-MTO employs a frozen foundational model to infer auxiliary guidance through in-context learning. The framework operates through three stages: constructing an algorithm-specific in-context query from evaluated solutions, using the foundational model to infer a guidance signal characterizing predictive relationships among tasks, and translating this signal into algorithm-specific guidance for maximum-a-posteriori coupling estimation. This approach provides regularization during the early, data-scarce stages of optimization and gradually relinquishes control as task-specific observations accumulate. We instantiate the framework in multitask Bayesian optimization as ICG-MTBO, using directional fitness-class queries to guide inter-task coupling estimation, and further instantiate it in MFEA-II using decision-space-overlap queries to guide random mating probability estimation. Experiments across synthetic benchmarks and a real-world robot arm control problem, together with evaluations under different acquisition functions and evolutionary multitasking, demonstrate the effectiveness and generality of ICG-MTO for few-shot multitask optimization.
Chinese Translation
多任务优化旨在同时求解一组优化任务,但在评估预算有限的情况下,任务间关系估计往往不准确,从而导致负迁移。本文提出了一种新框架——上下文引导多任务优化,利用数值基础模型来改进少样本场景下的任务间耦合估计。与仅依赖稀缺观测数据的传统方法不同,ICG-MTO 采用冻结的基础模型,通过上下文学习推断辅助引导信息。该框架包含三个阶段:从已评估的解中构建特定算法的上下文查询;利用基础模型推断刻画任务间预测关系的引导信号;将该信号转化为特定算法的引导信息,用于最大后验耦合估计。这一方法在优化早期数据稀缺阶段提供正则化作用,并随着任务特定观测数据的积累逐渐放弃控制。我们将该框架实例化为多任务贝叶斯优化,使用方向性适应度等级查询引导任务间耦合估计,并将其进一步实例化到 MFEA-II 中,使用决策空间重叠查询引导随机交配概率估计。在合成基准测试和真实世界机械臂控制问题上的实验,以及在不同采集函数和进化多任务框架下的评估,共同证明了 ICG-MTO 在少样本多任务优化中的有效性和通用性。
cs.LG / 59 / 2609.25839

Gaussian Flow-Matching Schedules: Implications for Sampling and Training

Claustre, Arsène, Negrel, Hugo, Boyer, Claire, Nadjahi, Kimia, Vanden-Eijnden, Eric
Abstract
Flow-matching schedules affect both sampling dynamics and the variance of the regression target. For centered commuting Gaussians, we show that a direction-dependent schedule decomposes into two independent design choices: a variance path, which fully determines the intermediate laws and probability flow, and a factorization, which leaves this flow unchanged while controlling irreducible regression variance. On the sampling side, we analyze finite-step Euler accuracy and derive a necessary drift bound for exact N -step sampling, connecting the geodesic and the logarithmic path. On the training side, for any fixed path, we derive closed-form factorizations that either minimize time-averaged regression variance or make it constant along the path.
cs.LG / 60 / 2609.25874

Neural Approximation by Function Composition: Rigidity and Doubly Exponential Convergence

Huang, Wentao, Zhang, Haizhang
Abstract
Deep neural networks approximate functions by composing affine maps with nonlinear activations, but how composition itself creates approximation power is not yet fully understood. We investigate a fundamental mechanism: geometrically weighted sums of iterates of a single scalar generator function. This mechanism underpins the classical tent-map construction of the function \(x - x^2\) and related recursive representations used by Yarotsky, W. E, et al., to analyze the approximation powers of deep neural networks. First, we establish a rigidity theorem: for continuous piecewise linear generators with a finite number of segments, any \(C^3\) function that can be represented in this way is at most quadratic. For non-affine quadratic functions, the geometric factor is at least $1/4$. This result both reveals limitations of the tent-map approach and complements existing methods based on hierarchical bases and recursive polynomial constructions. Second, using an exact remainder identity as guidance, we construct a smooth generator whose iterates yield doubly exponential error decay in total depth for square approximation and, through multiplication modules, for each fixed polynomial. For power series with absolutely summable coefficients on \([-1,1]^d\), distributing depth according to monomial degree yields a uniform approximation error of order \(O(e^{-cL^{1/d}})\) on each interior cube. These findings demonstrate how generator dynamics and remainder estimates govern depth allocation and approximation rates of deep neural networks.
cs.LG / 61 / 2609.25876

Evaluating the Effectiveness of SechKAN on 1D Data

评估SechKAN在一维数据上的有效性
Ta, Hoang-Thang
Abstract
The connection between the Kolmogorov-Arnold representation theorem (KART) and neural network design has led to the development of Kolmogorov-Arnold Networks (KANs), with applications ranging from STEM problems to AI tasks. In this paper, we investigate the effectiveness of a KAN variant, SechKAN, which relies on hyperbolic secant (sech) functions as basis functions, with a 1D projection to reduce the number of parameters to a level comparable to MLPs. We evaluate SechKAN on three 1D classification datasets: UCI Human Activity Recognition (UCI HAR), ElectricDevices, and Crop, and compare it with several effective networks, including EfficientKAN, MLP, CNN1D, ResNet1D, and DSCNN1D, using approximately comparable parameter budgets. The results indicate that SechKAN achieves competitive performance across the three datasets, with particularly strong performance on Crop. Ablation studies further show that grid size and normalization affect performance, suggesting that SechKAN's effectiveness depends on the dataset and architectural choices. Our source code and experimental implementation are publicly available at: https://github.com/hoangthangta/SechKAN_1D.
Chinese Translation
Kolmogorov-Arnold表示定理(KART)与神经网络设计之间的联系催生了Kolmogorov-Arnold网络(KANs)的发展,其应用范围涵盖STEM问题到AI任务。本文研究了一种KAN变体——SechKAN的有效性,该变体采用双曲正割(sech)函数作为基函数,并通过一维投影将参数量降低到与多层感知机(MLP)相当的水平。我们在三个一维分类数据集上评估了SechKAN:UCI人体活动识别(UCI HAR)、ElectricDevices和Crop,并在大致相当的参数预算下,将其与多个高效网络进行比较,包括EfficientKAN、MLP、CNN1D、ResNet1D和DSCNN1D。结果表明,SechKAN在三个数据集上均取得了具有竞争力的性能,尤其在Crop数据集上表现突出。消融研究进一步表明,网格大小和归一化会影响性能,说明SechKAN的有效性依赖于数据集和架构选择。我们的源代码和实验实现已在以下网址公开:https://github.com/hoangthangta/SechKAN_1D。
cs.LG / 62 / 2609.25914

AURA: Angular Update Rate Adaptation for training complex-valued neural networks

AURA:用于训练复值神经网络的角度更新率自适应方法
Ballini, Enrico, Engsig-Karup, Allan Peter, Andriollo, Tito
Abstract
Complex-valued neural networks (CVNNs) are increasingly adopted for complex-valued data; however, they are often trained with first-order optimizers inherited from the real-valued case. The efficiency of these methods depends largely on the step size, and their step-size rules ignore the angular information available in the complex plane. We address step-size adaptation in the complex domain by introducing AURA (Angular Update Rate Adaptation), a per-parameter step-size adaptation that can be added on top of any first-order optimizer, and removed from it, without altering its update direction. AURA measures the agreement between consecutive updates of each complex parameter, in length, alignment, and sense of rotation, and enlarges the step when they are consistent and reduces it when they are not. It requires no additional gradient evaluations and only inexpensive vector operations per step. We combine AURA with Adam and Muon and compare the resulting methods with well-known first-order optimizers on four test cases of increasing complexity, ranging from the approximation of scalar complex functions to physics-informed training. Fully connected neural networks are used throughout this work. All hyperparameters other than the step size are held fixed across test cases; for one case, we also tune the hyperparameters of each optimizer under the same budget. Our empirical tests show that AURA improves the convergence of its base optimizer in most cases with a small per-step overhead, and we identify the conditions under which it fails to do so.
Chinese Translation
复值神经网络(CVNN)在复值数据的处理中被越来越广泛地采用,然而它们通常使用从实值情形继承而来的一阶优化器进行训练。这些方法的效率在很大程度上取决于步长,而其步长规则忽略了复平面上可用的角度信息。我们通过引入 AURA(Angular Update Rate Adaptation,角度更新率自适应)来解决复数域中的步长自适应问题。AURA 是一种逐参数的步长自适应方法,可以叠加在任何一阶优化器之上,也可以从中移除,且不改变其更新方向。AURA 通过长度、对齐方式和旋转方向来度量每个复参数连续两次更新之间的一致性,当更新一致时增大步长,不一致时减小步长。它不需要额外的梯度计算,每步只需开销很低的向量运算。我们将 AURA 与 Adam 和 Muon 相结合,并在四个复杂度递增的测试案例(从标量复函数逼近到物理信息训练)上,将所得方法与著名的一阶优化器进行比较。整个工作均采用全连接神经网络。除步长外,所有超参数在各测试案例中保持固定;在其中一个案例中,我们还在相同预算下对每种优化器的超参数进行了调优。我们的实证测试表明,AURA 在大多数情况下以较小的每步开销提升了其基础优化器的收敛性能,同时也指出了其未能奏效的条件。
cs.LG / 63 / 2609.25916

Beyond Scalar Sensitivity: Activation-Aware Mixed-Precision LLM Quantization with Cross-Layer Refinement

超越标量敏感度:基于激活感知与跨层优化的混合精度大语言模型量化
Yoshida, Akihiro, Ichikawa, Yuma
Abstract
Mixed-precision weight quantization is commonly formulated as a Multiple-Choice Knapsack Problem (MCKP), yet existing solvers rely on scalar sensitivity proxies that collapse each weight matrix's Hessian into a single number and treat every module independently. We prove that even the optimal scalar proxy incurs multiplicative distortion up to $\sqrt{\kappa(\mathbf{A})\kappa(\mathbf{B})}$ relative to the full activation-aware quadratic, where $\kappa(\mathbf{A})$ and $\kappa(\mathbf{B})$ denote the condition numbers of the input- and output-side Hessian factors. This bound varies from $10^1$ to $10^{13}$ for typical LLM modules, making inter-module sensitivity ranking unreliable. To address these limitations, we propose Cross-layer Activation-aware Sensitivity Allocation (CASA), a two-phase method. In Stage 1, the scalar proxy is replaced by an activation-aware metric derived from the Kronecker-factored Hessian, reducing the MCKP to a form whose continuous relaxation admits a closed-form solution. In Stage 2, a cross-layer-aware local search evaluates bit-width updates using the end-to-end model loss. Experiments on multiple LLMs across different bit budgets show that CASA achieves lower perplexity than the latest scalar-proxy baselines, especially at ultra-low bit-widths ($<3$ bits per weight). Moreover, the performance gain in zero-shot accuracy tracks the per-model average condition-number over modules, confirming the distortion bound as a practical indicator of scalar-proxy failure.
Chinese Translation
混合精度权重量化通常被建模为多选择背包问题(MCKP),然而现有的求解器依赖于标量敏感度代理,即将每个权重矩阵的Hessian矩阵压缩为单一数值,并独立地处理各个模块。我们证明,即使是最优的标量代理,相对于完整的激活感知二次目标也会产生高达 $\sqrt{\kappa(\mathbf{A})\kappa(\mathbf{B})}$ 的乘性失真,其中 $\kappa(\mathbf{A})$ 和 $\kappa(\mathbf{B})$ 分别表示输入侧和输出侧Hessian因子的条件数。对于典型的大语言模型模块,该界限在 $10^1$ 到 $10^{13}$ 之间变化,使得模块间的敏感度排序变得不可靠。为解决这些局限性,我们提出了跨层激活感知敏感度分配方法(Cross-layer Activation-aware Sensitivity Allocation, CASA),这是一种两阶段方法。在第一阶段,用基于Kronecker分解Hessian推导的激活感知度量替代标量代理,将MCKP化简为一种其连续松弛具有闭式解的形式。在第二阶段,一种跨层感知的局部搜索方法利用端到端模型损失来评估位宽更新。在多个大语言模型和不同比特预算上的实验表明,CASA相比最新的标量代理基线方法获得了更低的困惑度,尤其是在超低位宽(每权重少于3比特)的情况下。此外,零样本准确率的性能增益与各模型模块的平均条件数变化趋势一致,证实了该失真界限可作为标量代理失效的实用指标。
cs.LG / 64 / 2609.25938

Certified Against Which Oracle? Execution Labels Set the Reported Risk of Conformal Abstention for Text-to-SQL

Liu, Jiamiao, Qiao, Dewen, Zhang, Yu, Chen, Xuetao
Abstract
A conformal abstention certificate for text-to-SQL is only as truthful as the correctness labels it is calibrated on. The uncertainty pipelines that read confidence off execution consistency take those labels from the single database a benchmark ships, an oracle known to be lenient. We run a preregistered intervention on Spider-Realistic, swapping that database for the benchmark's distilled multi-instance test suite. Across four SQL-specialist checkpoints and two split schemes, the swap raises the certificate's held-out risk 2.73 to 10.23 points above the risk its own labels report. Neither oracle reports the risk experts assign. Under blinded labels from two SQL experts, a certificate calibrated at a nominal 0.10 carries 20.0 and 17.2 points of risk on two checkpoints. The stricter oracle errs in both directions: most of the answers it rejects are not judged wrong, and some of those it accepts are. An AI-assigned census of what it rejects finds a semantic error in a quarter to a third of them, depending on the population. It attributes most of the rest to underspecified questions, synthetic instances or suspected reference-query defects, a flag supported by a preregistered blinded expert audit. The oracle also decides how a confidence score is judged. Every execution-consistency score looks better under the labels of the oracle that built its clusters, in 16 of 16 combinations. Under expert labels, building such a score on suite clusters instead of shipped-database clusters raises its area under the ROC curve (AUROC) by 6.96 points on one checkpoint and 1.53 on the other. On the second, the expert interval excludes the 8.3 points the suite labels report. A certificate should be reported with both oracles, and an oracle-relative difference read as semantic risk only after the benchmark is audited. A consistency score should be evaluated under an oracle that did not build it.
cs.LG / 65 / 2609.25962

Exploring Solver-Level Warmstarting for Neural Network Verification

探索求解器层面的热启动方法在神经网络验证中的应用
Bosman, Annelot, Liu, Minghao, Kwiatkowska, Marta, Hoos, Holger, van Rijn, Jan
Abstract
Neural network verification has become a key tool for providing formal guarantees on the behaviour of neural networks. However, many verification problems remain computationally intractable in the worst case: even for common adversarial robustness specifications, verification is NP-complete. Here, we explore the application of solver-level warmstarting for neural network verification to exploit information from previous solutions. We study the effect on running time as several properties are modified, including perturbation radii, input data and the networks themselves, using a pipeline that is generalisable and potentially adaptable to state-of-the-art verifiers. Our results show that warmstarting can significantly reduce verification time in most cases. Moreover, warmstarting enables the successful verification of instances that could not be solved from scratch within the given time limit.
Chinese Translation
神经网络验证已成为为神经网络行为提供形式化保证的关键工具。然而,许多验证问题在最坏情况下仍然是计算上难以处理的:即使对于常见的对抗鲁棒性规范,验证问题也是NP完全的。本文探索了在神经网络验证中应用求解器层面的热启动(warmstarting)方法,以利用先前求解得到的信息。我们研究了在扰动半径、输入数据和网络本身等多个属性发生变化时对运行时间的影响,所采用的流水线具有通用性,并有望适用于最先进的验证器。实验结果表明,热启动在大多数情况下能够显著缩短验证时间。此外,热启动还使得那些在给定时间限制内无法从零开始求解的实例得以成功验证。
cs.LG / 66 / 2609.25963

GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression

GeoPair:面向免训练Transformer压缩的保几何跨层分解方法
Mohammad, Baher, Ali, Ammar, Lefkimmiatis, Stamatios
Abstract
Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations. Rather than forcing weights of adjacent layers to share a basis or heuristically merging activation statistics, our approach identifies structurally compatible projections and learns a shared representation that better preserves each layer's distinct calibration geometry. Coupled with structured sparsity, this yields highly efficient weight decompositions without sacrificing functional fidelity. Across diverse architectures, scales, and modalities, our method achieves state-of-the-art results, consistently outperforming independent structured weight decompositions and alternative pairwise weight factorizations, which operate under heuristic grouping strategies. By replacing heuristic engineering strategies with a convergent, optimization-driven pipeline, we establish a theoretically grounded foundation for scalable, transformer compression across different modalities.
Chinese Translation
Transformer架构存在跨层冗余,然而训练后压缩流程通常孤立地优化各层,或依赖忽视层间激活几何特性的启发式分组策略。我们提出一个有原则的、免训练的框架,该框架顺序地优化跨层权重配对与共享字典分解。我们的方法不是强制相邻层的权重共享一组基,也不是启发式地合并激活统计量,而是识别结构上兼容的投影,并学习一个能更好保留各层独特校准几何的共享表示。结合结构化稀疏性,该方法可在不牺牲功能保真度的情况下实现高效的权重分解。在多种架构、规模和模态上,我们的方法均取得了最先进的结果,持续优于独立的结构化权重分解以及在启发式分组策略下运行的其他成对权重分解方法。通过用收敛的、由优化驱动的流程取代启发式工程策略,我们为跨模态的可扩展Transformer压缩奠定了有理论基础的方法体系。
cs.LG / 67 / 2609.25980

Interweaving Marginals into Multivariate Sample Paths: Training-Free Dependence Construction for Probabilistic Time Series Foundation Models

Choi, Jinmyeong, Jang, Jinkwan, Lee, Seul, Kim, Taesup
Abstract
Probabilistic time series foundation models (TSFMs) provide coordinate-wise predictive distributions, but these marginals do not determine a joint distribution over multivariate future trajectories. We study training-free coupling of frozen TSFM marginals into multivariate forecast sample paths. Our primary evaluation fixes the empirical marginal sample multiset at every channel--horizon coordinate across methods, isolating the effect of coupling alone. Historical temporal and channel relations substantially improve their corresponding dependence diagnostics. The same pattern persists when the fixed-marginal constraint is removed and paths are sampled directly, and remains present under native multivariate backbone inference. These results support treating dependence reconstruction as a distinct post-processing problem for probabilistic TSFMs.
cs.LG / 68 / 2609.25987

Theory for groupoid equivariant neural networks: an approach for steerable CNNs on bounded domains

群胚等变神经网络理论:有界域上可操纵CNN的一种方法
Ibort, Alberto, Jimenez-Vazquez, Maria, Perez-Pardo, Juan M.
Abstract
Equivariant convolutional neural networks are usually built from a group acting globally on the space of signals. This hypothesis is inappropriate for many bounded or stratified domains: an ambient rigid motion may be admissible only on part of the domain, and the boundary introduces geometric types that are invisible to a transitive group action. We develop a theory of groupoid-equivariant neural networks in which the symmetry datum consists of a groupoid, a selected pseudogroup of local bisections, a measure, and input and output representation bundles. For integral channels on the object space, we prove a bisection-equivariant kernel theorem: equivariance is equivalent to a transport constraint on the two-point kernel, and its solutions are classified by one joint-stabilizer intertwiner on each orbit of pairs. As a case study we apply the theory to bounded planar domains. The resulting architecture is implemented through offline nullspace bases and sparse gather--transform--scatter operations. A Poisson--Dirichlet kernel study is used separately to assess boundary-aware inductive bias; the exact inverse is shown to preserve the global symmetries of the rectangle but not general proper local bisections. The numerical results show that the proposed architectures provide significant advantages when symmetries cannot be globally implemented by group actions and provide an accuracy improvement of at least one order of magnitude with respect to the models tested.
Chinese Translation
等变卷积神经网络通常建立在全局作用于信号空间的群之上。这一假设对于许多有界域或分层域并不适用:环境的刚体运动可能仅在域的一部分上是容许的,而边界引入了几何类型,这些类型对传递群作用是不可见的。我们发展了一种群胚(groupoid)等变神经网络理论,其中对称性数据由一个群胚、一个选定的局部双截(local bisections)伪群、一个测度以及输入和输出表示丛组成。对于对象空间上的积分通道,我们证明了一个双截等变核定理:等变性等价于二点核上的一个传输约束,且其解由每个点对轨道上的一个联合稳定子交织算子分类。作为案例研究,我们将该理论应用于有界平面域。所得架构通过离线零空间基和稀疏的聚集—变换—散射(gather–transform–scatter)操作实现。我们另外借助泊松—狄利克雷(Poisson–Dirichlet)核研究来评估边界感知的归纳偏置;结果表明精确逆能保持矩形的整体对称性,但无法保持一般的正常局部双截。数值结果表明,当对称性无法通过群作用全局实现时,所提出的架构具有显著优势,相对于所测试的模型,准确率至少提高了一个数量级。
cs.LG / 69 / 2609.26018

The Dynamics of Quasiregular Neural Learning

准规则神经学习的动力学
Sabatelli, Matthia
Abstract
Many learning problems combine a dominant regularity with systematic exceptions. Motivated by U-shaped learning in language acquisition, we study this interaction in controlled quasiregular regression problems where regular and exceptional solutions are explicitly known. Neural networks can partially acquire exceptions, subsequently regress toward the dominant regularity, and finally recover. This overregularization becomes substantially stronger when exceptions are rare, despite their early acquisition, but does not emerge equally across all regularities considered. Our results isolate a simple form of competition between regularities and exceptions during neural learning.
Chinese Translation
许多学习问题兼具主导规则性与系统性例外。受语言习得中U型学习现象的启发,我们在可控的准规则回归问题中研究了这种交互作用,其中规则解和例外解均是明确已知的。神经网络能够部分习得例外,随后向主导规则性回归,最终再恢复。尽管例外在早期即被习得,但当例外较为罕见时,这种过度规则化现象会显著增强,但并非在所有所考察的规则性中同等出现。我们的结果揭示了神经学习过程中规则性与例外之间竞争的一种简单形式。
cs.LG / 70 / 2609.26021

BOBA: Dynamic Bayesian Optimization through Bayesian Active Inference

Kelly, Merlin Angel, Patel, Rishan, Thomas, Alexander, Zhu, Ziyue, Quan, Zikun, Carlson, Tom, Cho, Youngjun
Abstract
Dynamic black-box optimization presents significant challenges for Bayesian Optimization (BO), as the objective function evolves over time, causing optimal locations to shift continuously. Existing dynamic BO (DBO) methods using standard acquisition functions such as Upper Confidence Bound (UCB) fail to explicitly account for temporal variations, leading to suboptimal sample allocation and poor tracking of moving optima. Here, we propose BOBA (Bayesian Optimization through Bayesian Active Inference), a novel acquisition function inspired by free energy principles from active inference that explicitly minimizes predictive uncertainty about future states in dynamic environments. BOBA extends traditional acquisition functions by incorporating a forward-looking uncertainty quantification that estimates uncertainty in function changes, enabling more informed exploration-exploitation trade-offs in non-stationary settings. We evaluate BOBA on synthetic dynamic benchmarks, comparing against state-of-the-art DBO methods. Our experiments demonstrate that BOBA significantly improves regret in query-restricted settings, while remaining competitive in time-limited settings. We further analyze variants of BOBA with different exploration strategies, showing how the exploration-exploitation balance can be tuned for different types of dynamic functions. This work contributes both a free energy-based acquisition function for DBO and insights into how active inference principles can enhance optimization in non-stationary environments, with implications for real-time applications requiring continuous adaptation.
cs.LG / 71 / 2609.26025

MICRO: Multi-Fidelity Active Search for Severe Error Discovery

MICRO:面向严重错误发现的多保真度主动搜索方法
Leone, Orlando, Pokel, Niclas, Moure, Pehuén, Gao, Yingqiang, Boehringer, Roman
Abstract
Human feedback can vary in cost and informativeness. Strong feedback can reveal severe errors but is costly, so cheaper quality ratings can help decide which items to annotate. We propose MICRO (Multi-Fidelity Impact Clustered Rollout), an active search framework that allocates a shared budget to these feedback types to maximise confirmed severe error discoveries. MICRO jointly models ratings and annotation losses conditional on item features to steer acquisition. It clusters acquisitions by their predicted impact on severity probabilities to select diverse candidates, then uses rollout to estimate their discovery value. Experiments on WMT20 English-German show that ratings improve both loss reconstruction and severity prediction. MICRO achieves the highest mean discovery count across four budget and rating cost settings, with similar performance to adapted MF-ENS in one and significant gains over all six comparison policies, including two rollout controls, in the other three $(p<.001)$.
Chinese Translation
人类反馈在成本和信息量上可能各不相同。强反馈能够揭示严重错误但成本高昂,因此较便宜的质量评分可以帮助决定对哪些条目进行标注。我们提出了MICRO(多保真度影响聚类滚动评估,Multi-Fidelity Impact Clustered Rollout),这是一个主动搜索框架,将共享预算分配给这些反馈类型,以最大化已确认的严重错误发现数量。MICRO在条目特征的条件下对评分和标注损失进行联合建模,以引导获取过程。它根据获取候选对严重性概率的预测影响对其进行聚类,从而选择多样化的候选对象,然后使用滚动评估来估计其发现价值。在WMT20英德翻译数据上的实验表明,评分同时改善了损失重构和严重性预测。在四种预算与评分成本设置中,MICRO均取得了最高的平均发现数量,其中一种设置下与改进的MF-ENS性能相当,在其余三种设置下则显著优于全部六种对比策略(包括两种滚动评估控制方法)(p<.001)。
cs.LG / 72 / 2609.26037

xWhyL: Causal Interactive Learning

Tagliapietra, Nicholas, Busch, Florian Peter, Willig, Moritz, Zečević, Matej, Halilaj, Lavdim, Luettin, Juergen, Kersting, Kristian
Abstract
Explanations are central to causal reasoning, and cognitive science has long established that the human drive to explain is itself a mechanism for learning about causality. Despite this, learning from those abductive signals is largely ignored in artificial intelligence. While explainable AI (XAI) increasingly draws on causal models to generate explanations, the converse direction about what explanations can do for causality remains largely unexplored. To fill this gap, we propose xWhyL, a formal framework connecting causality and XAI by learning causal models from explanations. We develop a mathematical theory that translates explanations into a learning signal complementary to observational data, and demonstrate how it enables overcoming the limits of observational causal discovery. As explanations can be derived from incorrect beliefs and clash with data, a tension we call the Causal Tug-of-War, we prove conditions under which our framework rejects misspecified explanations rather than absorbing them. Our practical instantiation, Causal Interactive Learning (CIL), shows how expert explanations can efficiently support causal discovery and distinguish correct from incorrect explanations.
cs.LG / 73 / 2609.26063

Towards Adaptive Federated Graph Clustering: A Global Community-aware Contrastive Learning-based Approach

迈向自适应联邦图聚类:一种基于全局社区感知对比学习的方法
Zhu, Yinlin, Wu, Di, Luo, Wang, Quan, Guocong, Hu, Miao
Abstract
Federated graph learning (FGL) enables multiple clients to collaboratively train graph models without sharing their private graph data, providing a promising paradigm for mining knowledge from distributed graph repositories. While most existing FGL methods focus on supervised tasks, real-world graphs are often massive and unlabeled, making federated graph clustering an important yet still immature research direction. Notably, this task is particularly challenging due to the inherent subgraph heterogeneity across clients, which leads to client-specific community structures. In this work, we identify two critical limitations in existing federated graph clustering methods: (1) unrealistic pre-defined cluster cardinality assumptions and (2) incomplete inter-community separation. To address these challenges, we propose AdaFGC, an Adaptive Federated graph clustering framework based on Global community-aware Contrastive learning. AdaFGC introduces an over-complete set of global community anchors to model the global community structure and adaptively estimate clustering cardinality via cross-client anchor refinement. In addition, it employs a global community-aware contrastive learning scheme that uses the shared anchors as contrastive prototypes to explicitly enforce community-level attraction and repulsion across clients, complemented by node-level and topology-level objectives that stabilize local representations. Extensive experiments on eight benchmark datasets demonstrate that AdaFGC consistently outperforms existing supervised and unsupervised FGL baselines across multiple clustering metrics.
Chinese Translation
联邦图学习(Federated Graph Learning, FGL)使多个客户端能够在不共享私有图数据的情况下协同训练图模型,为从分布式图数据存储中挖掘知识提供了一种有前景的范式。尽管现有的大多数FGL方法侧重于有监督任务,但现实世界中的图通常是大规模且无标签的,这使得联邦图聚类成为一个重要却尚不成熟的研究方向。值得注意的是,由于客户端之间固有的子图异构性会导致客户端特定的社区结构,该任务尤其具有挑战性。在本工作中,我们指出了现有联邦图聚类方法的两个关键局限性:(1)不切实际的预定义聚类数量假设;(2)不完整的社区间分离。为应对这些挑战,我们提出了AdaFGC,一种基于全局社区感知对比学习的自适应联邦图聚类框架。AdaFGC引入了一个过完备的全局社区锚点集合来建模全局社区结构,并通过跨客户端的锚点精炼自适应地估计聚类数量。此外,它采用了一种全局社区感知的对比学习方案,将共享锚点作为对比原型,显式地在客户端之间实施社区层面的吸引与排斥,同时辅以节点层面和拓扑层面的目标来稳定局部表示。在八个基准数据集上的大量实验表明,AdaFGC在多种聚类指标上持续优于现有的有监督和无监督FGL基线方法。
cs.LG / 74 / 2609.26067

FuncCode: Compressing Kolmogorov--Arnold Networks in Function Space with Hardware-Aware Quantization

FuncCode:基于硬件感知量化的Kolmogorov--Arnold网络函数空间压缩方法
Fuad, Kazi Ahmed Asif, Chen, Lizhong
Abstract
Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions, increasing flexibility but also parameter memory because each edge stores multiple coefficients, often together with a separate base branch. We introduce FuncCode, a basis-agnostic compression approach that forms shared codebooks from sampled edge responses, codes the basis and base branches independently, and exports the resulting codebooks and per-edge indices in a quantized, bit-packed format. Across spline and polynomial KANs, sampled edge responses exhibit $13$--$35\%$ lower effective rank than their coefficient representations. Further replicated controls show that function-space clustering alone is statistically tied with coefficient-space clustering; the consistent accuracy gain comes from preserving the distinct sharing structure of the two branches. On a ten-seed MNIST benchmark, FuncCode compresses spline and GRAM KANs by $31.6\times$ and $17.6\times$ with only $0.31$ and $0.34$ pp accuracy loss. On a 6.1M-edge convolutional KAGN, it achieves $19.9\times$ compression while remaining within $0.54$ pp of dense accuracy on CIFAR-10 and $1.89$ pp on CIFAR-100. After compression, per-edge indices account for up to $99.4\%$ of stored weight bits, making the representation index-bound. Across nine bit-exact FPGA accelerators, FuncCode reduces SplineKAN post-route weight memory by $3.87\times$ relative to dense INT4, without increasing cycle count or latency. The FuncCode implementation is available at https://github.com/OSU-STARLAB/FuncCode.
Chinese Translation
Kolmogorov--Arnold网络(KANs)用可学习的单变量函数取代标量边权重,提高了灵活性,但由于每条边需存储多个系数(通常还伴随一个独立的基础分支),也增加了参数内存。我们提出FuncCode,这是一种与基函数无关的压缩方法,它从采样的边响应构建共享码本,对基函数分支和基础分支独立编码,并以量化、比特打包的格式导出所得码本及每条边的索引。在样条和多项式KAN上,采样边响应的有效秩比其系数表示低$13$--$35\%$。进一步的重复对照实验表明,仅函数空间聚类与系数空间聚类在统计上相当;一致的精度提升源于保留了两个分支各自独特的共享结构。在十种随机种子的MNIST基准测试中,FuncCode对样条KAN和GRAM KAN分别实现$31.6\times$和$17.6\times$的压缩,精度损失仅为$0.31$和$0.34$个百分点。在具有610万条边的卷积KAGN上,它实现了$19.9\times$的压缩,同时在CIFAR-10上与稠密模型精度差距保持在$0.54$个百分点以内,在CIFAR-100上差距为$1.89$个百分点。压缩后,每条边的索引最多占存储权重比特的$99.4\%$,使该表示成为索引受限型。在九个比特精确的FPGA加速器上,FuncCode将SplineKAN布线后的权重内存相对稠密INT4降低了$3.87\times$,且未增加周期数或延迟。FuncCode的实现可在 https://github.com/OSU-STARLAB/FuncCode 获取。
cs.LG / 75 / 2609.26077

Fast Matrix Multiplication in fp8: Certified Coefficient Optimization and Measured Error

Xie, Shuxiao, Xie, Shuyang, Cao, Yuan, Ran, Dezhi, Yang, Wei, Xie, Tao
Abstract
A Strassen-type algorithm has many realizations with the same exact product and multiplication count yet different fp8 error because basis changes reshape coefficient geometry, posing the question of which to run. No current account settles this: classical stability controls worst-case $\ell_1$ growth, not the expected-error magnitude, and the Dumas--Pernet--Sedoglavic optimizer could only be called probably optimal, its global optimality unproved. To settle this, we attach to each realization a coefficient functional $\Phi$, a scalar summary of its coefficient geometry, which we minimize over the change-of-basis orbit. This Kempf--Ness problem on a Hadamard manifold lets us certify the global $\Phi$ optimum rather than merely search for it: an exact moment-map zero fixes $\Phi_{\min} = 200/9$, and de Groote's classification extends that optimality to every exact real rank-7 $2\times2$ decomposition. Every exact real rank-7 realization therefore has a $\Phi$-predicted RMS constant at least $5/3$ times that of the cubic algorithm, at fixed noise coefficient. We then introduce an explicit block-scaled e4m3 model in which $\Phi$ is the leading-order coefficient of relative expected mean-squared error, and we test the resulting $\Phi$-predicted ordering against realized fp8 error. Ordering and re-basing experiments support that prediction within tested fused block-scaled regimes, and on real matmul tiles from four architecture families the $\Phi$-optimal realization falls in the fp8 low-error region. Across two $\sim$70B models on real deep_gemm kernels, the same realization removes 10 to 55% of classic Strassen's excess NLL over the clean model. Algorithm realization thus becomes a mathematically certified design problem rather than a tuning choice: an independent low-precision axis with a global $\Phi$ optimum and measured fp8 relevance.
cs.LG / 76 / 2609.26081

Margin-Drop Coordinates for Cross-Budget Robustness Evaluation

Huang, Yanliang, Zhang, Zhen, Xie, Peng, Wu, Wenyuan, Zhu, Sitong, Zeng, Zhuoqi, Alanwar, Amr
Abstract
Fixed-budget robustness evaluation can select the wrong frozen vision encoder. An encoder that survives a shallow attack may lose most of that robustness when the same evaluation is strengthened. We ask whether the shallow evaluation contains enough information to identify this budget fragility. For each clean-correct sample, the evaluation records the clean pairwise margin, the first-order linearized margin-drop scale, the margin drop from a clean-start one-step attack, and the drop reached by an iterative attack. Normalizing by that scale gives three margin-drop coordinates capturing clean margin slack, one-step shortfall, and drift, where drift is the additional normalized margin drop the iterative attack reaches beyond the one-step perturbation. Together, they reconstruct the normalized post-attack margin and therefore the pass-or-fail outcome. Across 42 pretrained frozen vision encoders, the shallow survival rate carries essentially no rank information about subsequent PGD-10 to PGD-200 collapse, at Spearman -0.006, while the median shallow drift coordinate ranks the same collapse at +0.811. The result persists in a held-out encoder pool and under an $\ell_\infty$ evaluation. With deep evaluation limited to 11 encoders, ranking by shallow drift recovers 11 of the 17 high-collapse encoders, compared with 5 under survival-rate ranking. The full coordinate decomposition further distinguishes cases that share the same fixed-budget residual but diverge at deeper budgets, and separates margin repair from drift repair under interventions, revealing distinct repair paths that endpoint robustness alone does not identify.
cs.LG / 77 / 2609.26094

CoEvo: Oracle-Grounded Self-Evolution of a Single Model for Multi-Step Causal Reasoning

CoEvo:基于Oracle锚定的单模型多步因果推理自进化方法
Zhang, Jian, Wang, Bingyi, Liu, Yizhi
Abstract
Multi-step causal reasoning requires chaining inferences where each step constrains the next. An early error propagates silently, and a correct answer reached via flawed logic evades outcome-level detection. In specialized domains, teacher LLMs err on intermediate steps, safety constraints restrict cloud distillation, and shifting conditions demand adaptation, leaving self-evolution as the practical route. Naive self-evolution can collapse: outcome-only rewards let the model exploit distributional shortcuts, and weak self-evaluation reinforces spurious paths into stable failure patterns. We exploit a key asymmetry: generating a correct chain is hard, but verifying a single step is easy. Many high-stakes domains admit a deterministic, queryable oracle, a physics simulator or rule engine over codified constraints. It checks asserted steps without teacher-level ability and abstains beyond its rules; it can check what the model asserts, never replace it. This enables CoEvo, an oracle-grounded self-evolution framework where a single model alternates between Proposer and Solver. As Solver, the model generates competing chains; intra-group debate exposes disagreement steps, a proxy for the capability boundary, and the oracle adjudicates them into process-level supervision. As Proposer, the same model constructs progressively harder scenarios inside oracle constraints, steering the curriculum toward deep multi-hop chains. Both roles are updated jointly, so training pressure co-evolves with the model. On industrial, clinical, and legal multi-step causal reasoning benchmarks, CoEvo enables an 8B LLM to sustain self-evolution, surpassing distillation baselines and the strongest proprietary reference on path correctness (82.1% vs. 71.4%). The trained model generalizes to unseen categories and systems, preserving root-cause accuracy.
Chinese Translation
多步因果推理要求将推理步骤串联起来,其中每一步都约束下一步。早期错误会悄然传播,而通过错误逻辑得出的正确答案则能够规避结果层面的检测。在专业领域中,教师大语言模型(LLM)会在中间步骤上出错,安全限制约束了云端蒸馏,不断变化的环境条件要求模型具备适应能力,这使得自进化成为实际的可行路径。朴素的自我进化可能会崩溃:仅基于结果的奖励会让模型利用分布性捷径,而薄弱的自我评估则会将虚假路径强化为稳定的失败模式。我们利用一个关键的不对称性:生成正确的推理链很困难,但验证单个步骤却很容易。许多高风险领域都存在确定性的、可查询的Oracle——例如物理模拟器或针对既定规则的规则引擎。它无需具备教师级别的能力即可检验模型所断言的步骤,并在超出其规则范围时选择弃权;它可以检验模型断言的内容,但绝不能取代模型。基于此,我们提出了CoEvo,一种以Oracle为锚定的自进化框架,其中单一模型在提议者和求解者之间交替。作为求解者,模型生成相互竞争的推理链;组内辩论会暴露出分歧步骤(这是模型能力边界的一种代理指标),并由Oracle对这些步骤进行裁决,从而产生过程级别的监督信号。作为提议者,同一个模型在Oracle约束范围内构建难度递增的场景,将课程学习引向深层的多跳推理链。两个角色的更新是联合进行的,因此训练压力与模型共同进化。在工业、临床和法律的多步因果推理基准测试中,CoEvo使一个8B参数的LLM能够持续进行自进化,在路径正确率上超越了蒸馏基线以及最强的专有模型参考(82.1% vs. 71.4%)。经过训练的模型能够泛化到未见过的类别和系统,并保持根因分析的准确性。
cs.LG / 78 / 2609.26112

Certified Mechanistic Interpretability: Lifting Single-Input Findings to Bounded Neighbourhoods

Zhang, Zhen, Huang, Yanliang, Xie, Peng, Wu, Wenyuan, Alanwar, Amr
Abstract
Mechanistic interpretability reverse-engineers transformer circuits one input at a time, leaving observed mechanisms without guarantees over bounded input neighbourhoods. We address this gap with a framework based on constrained polynomial-zonotope (CPZ) propagation that lifts mechanistic-interpretability observations from a single input to certified statements over a bounded set of perturbations. Three internal-attention queries (top-$k$ stability, evidence mass, and attention entropy) are formulated as tractable programs over the simplex of attention weights, and CPZ propagation through transformer blocks is shown to preserve the softmax simplex and the LayerNorm zero-mean identity exactly. A recursive Jacobian zonotope construction extends the same certificates across layer depth by linearising the block stack at the input and avoids per-layer generator growth. We instantiate the framework on transformer attention; the resulting certificates offer a way to sharpen mechanistic statements that single-input inspection cannot resolve on its own, and to inform downstream decisions in regimes where empirical heuristics may be misleading.
cs.LG / 79 / 2609.26146

From Risk Scoring to Risk Allocation: A Density-Driven Framework for Diverse Monitoring in Multi-Agent Systems

Wang, Zhaohui
Abstract
Risk monitoring in multi-agent systems is commonly built on a per-state primitive that scores each state independently and selects the top K. Under crowding, where many agents share the same fragility, this approach picks redundant alerts whose risks are jointly correlated, a pattern we describe as ``herding in monitoring.'' We propose a paradigm shift from risk scoring to risk allocation, supported by two contributions. First, we identify the Crowding Paradox, namely that P(risk | x) $\propto$ p(x) rather than 1/p(x), so density rather than anomaly score is the operative risk signal; on financial data, density-based scoring reaches AUROC $\geq$ 0.94 at 5d/10d/20d crash horizons, while five anomaly baselines all fall below 0.80. Second, given a density-derived fragility score, we recast monitoring as combinatorial subset selection over interdependent states and map it to a QUBO objective with a $\lambda$-controlled risk--diversity tradeoff. The resulting Pareto frontier contains standard diverse-subset methods (MMR, k-DPP) as fixed operating points; the gain over greedy grows monotonically with scale, from +24% at n=15 to +66% at n=200; a learned $\lambda$ policy reaches 99.5% of an oracle grid-search objective; and the formulation transfers to traffic and multi-agent reinforcement learning. The same QUBO instances execute without modification on Rigetti superconducting QPUs (Ankaa-3 and Cepheus-1-108Q via Amazon Braket), which we report as a compatibility property of the formulation rather than a claim of quantum advantage at this scale.
cs.LG / 80 / 2609.26147

Block-Level Weight-Space Structure Persists Under Post-Training: An Empirical Study Across LLM Families

Wang, Zhaohui
Abstract
Modern LLMs are deployed as families of post-trained variants (base, instruct, chat, code) derived from a shared set of pre-trained weights. We present an empirical study of how post-training transforms weight-space geometry, covering eight configurations across four architecture families (Qwen2.5, Llama-3.1/3.2, Mistral, Gemma-2). We identify a granularity gap: post-training modifies every tensor (zero of 291-339 tensors remain byte-identical, so hash-based deduplication achieves 0% savings), yet preserves block-level structure (mean cosine similarity exceeds 0.99 and relative Frobenius distance stays below 0.13). Post-training therefore acts as a structured perturbation that shifts every parameter while leaving block-level geometry intact. The property is not universal: independently trained specializations (for example, Qwen2.5-Coder) attain cosine similarity around 0.64 with the general base, indicating a disconnected region of weight space. Perturbation magnitude varies systematically with model scale, architecture family, and post-training recipe. As a practical application, we build LinkerLLM, a lazy loader that aliases shareable blocks across co-resident variants, achieving 18-48% GPU memory savings and enabling up to five 7B-parameter variants on a single 24 GB consumer GPU. Five of eight configurations retain at least 94% of the unshared variant's quality on MMLU, ARC-Challenge, HellaSwag, and WinoGrande; the remaining three (Mistral-7B, Gemma-2-2B, Llama-3.2-1B) have one below-threshold benchmark each (87-91%), which we report transparently rather than gate the block-sharing decision on a single threshold.
cs.LG / 81 / 2609.26165

Spectral Tail Interventions in Decoder-Only Language Models: Reasoning-Sensitive Weight Structure from Controlled Surgery

Shihab, Ibne Farabi, Akhter, Sanjida, Swaqeeb, Md Najmus, Ahsan, Abu Sa-Adat Mohamed Moon-Im Al, Sharma, Anuj
Abstract
Weight-space structure often correlates with language-model behavior, but correlation alone does not establish computational involvement. We study concentrated upper spectral tails in decoder-only transformers through controlled interventions. At a fixed relative offset, we derive a finite-width conditional bound linking the inverse participation ratio of squared singular values to central pre-softmax logit kurtosis. We then define a pointwise query--key ($QK$) product-tail target and compare independent factor surgery with a product-targeted factorization that preserves native attention computation. Across three base checkpoints and five reasoning benchmarks, plus an instruction-tuned Phi checkpoint analyzed separately, the learned-tail edit is more damaging than the mean of five fixed spectrum-matched Haar controls in all 20 model--task cells. Eighteen paired contrasts remain significant after Holm correction, while two are directional but inconclusive. Product-targeted factors attain higher held-out tail-subspace fractions, providing an empirical bridge between product- and factor-level interventions. Component isolation identifies contributions from $QK$, value--output, and multilayer-perceptron blocks, although the theorem covers only $QK$. In separate studies, inverse participation precedes pooled accuracy transitions under a matched crossing rule, and residualized tail-aware low-rank adaptation (LoRA) reaches targets earlier than standard LoRA and PiSSA while final-score intervals overlap. Conclusions are restricted to the evaluated checkpoints, layers, tasks, interventions, and controls.
cs.LG / 82 / 2609.26167

Activation-Energy Pruning for Spiking Neural Networks: Unsupervised Personalization via Spike-Count Saliency

基于激活能的脉冲神经网络剪枝:通过脉冲计数显著性实现无监督个性化
Bingham, Joseph
Abstract
Activation-energy pruning -- removing weights whose product of magnitude and cumulative pre-synaptic spike count falls below a threshold -- was established as an effective unsupervised personalization strategy for conventional deep neural networks~\citep{BINGHAM2025101242}. This paper asks what happens when the same criterion is applied to spiking neural networks (SNNs), where activation energy is not merely a useful heuristic but a literal physical quantity proportional to the metabolic cost of each synapse. The answer is surprising on three counts. First, gradient-based pruning methods that perform competitively on conventional networks (SNIP, GraSP, magnitude pruning) consistently underperform on SNNs, collapsing to near-chance accuracy by $\sigma = 0.2$ sparsity across all tested architectures and datasets. We trace this to a systematic incompatibility between surrogate-gradient saliency estimation and the binary spike-train representation, though we cannot rule out that alternative surrogate choices or hyperparameter settings might partially mitigate the effect. Second, activation-energy pruning applied to a neuromorphic benchmark \emph{improves} over the source model at high sparsity ($98.4 \pm 0.4\%$ vs.\ $97.2 \pm 0.7\%$ at $\sigma = 0.8$ on N-MNIST), a phenomenon with no counterpart in the conventional network setting. We interpret this result as consistent with experience-dependent cortical specialisation: removing connections active only for non-target classes may reduce cross-class interference and produce a cleaner target representation, though we note this is an interpretive analogy rather than a mechanistic demonstration.
Chinese Translation
激活能剪枝——即移除权重大小与突触前累计脉冲计数乘积低于某一阈值的权重——已被证明是传统深度神经网络的一种有效无监督个性化策略~\citep{BINGHAM2025101242}。本文探讨当同样的准则应用于脉冲神经网络(SNN)时会发生什么:在SNN中,激活能不仅是一种有用的启发式度量,更是一个与每个突触代谢成本成正比的真实物理量。答案在三个方面令人意外。首先,在传统网络上表现具有竞争力的基于梯度的剪枝方法(SNIP、GraSP、幅值剪枝)在SNN上持续表现不佳,在所有测试架构和数据集上,当稀疏度达到 $\sigma = 0.2$ 时,准确率均坍缩至接近随机水平。我们将此归因于代理梯度(surrogate-gradient)显著性估计与二值脉冲序列表示之间的系统性不兼容,尽管我们不能排除其他代理梯度选择或超参数设置可能部分缓解这一效应。其次,激活能剪枝应用于神经形态基准数据集时,在高稀疏度下的表现优于源模型(在N-MNIST上,$\sigma = 0.8$ 时为 $98.4 \pm 0.4\%$,而源模型为 $97.2 \pm 0.7\%$),这一现象在传统网络场景中没有对应。我们将该结果解释为与经验依赖的皮层特化机制相一致:移除仅对非目标类别活跃的连接可能会减少跨类别干扰,从而产生更清晰的目标表示,但需指出这是一种解释性类比,而非机制性证明。
cs.LG / 83 / 2609.26173

Component Type, Not Reconstruction Error, Predicts Attention Quantization Sensitivity

决定注意力量化敏感性的因素是组件类型而非重构误差
Dewage, Kasun, Pensky, Marianna, De Silva, Suranadi
Abstract
Many post-training quantization (PTQ) methods use layer-wise reconstruction, second-order proxy objectives, or activation-aware transformations to reduce quantization-induced error. Whether that error signal predicts the downstream functional impact of quantizing an individual attention projection has not been directly characterized. We sweep nine open-weight language models (1.3B--8B parameters; OPT, GPT-J, LLaMA-1/2/3, Mistral, Qwen 2.5) and quantize one attention projection at a time under round-to-nearest (RTN) and, for seven models, GPTQ at 3 and 4 bits, recording reconstruction error, perplexity change, and per-projection activation-weighted quantization error for 3,808 distinct measurements. We find: (1) within a given component type (Q, K, V, or O), reconstruction error explains less than 10% of the variance in perplexity sensitivity in 27 of 36 cases under RTN, with median R^2 = 0.044; (2) both component type and layer identity explain more variance than reconstruction error in all 9 models, with layer identity the strongest predictor in 7 of 9 models and component type strongest in the remaining 2; (3) value (V) projections are the most commonly dominant component, accounting for 38--51% of total positive Delta PPL in seven of nine models; (4) the dominant component is broadly preserved between RTN and GPTQ (5 of 7 cases); and (5) activation-weighted quantization error is a moderately better within-component predictor than reconstruction error for V projections specifically (median R^2 of 0.20 vs. 0.06). These findings indicate that relative weight reconstruction error alone is insufficient for sensitivity-aware bit allocation, and that V projections merit dedicated consideration in mixed-precision schemes.
Chinese Translation
许多训练后量化(PTQ)方法采用逐层重构、二阶代理目标函数或激活感知变换来降低量化引入的误差。然而,该误差信号能否预测对单个注意力投影进行量化的下游功能影响,尚未得到直接表征。我们对九个开放权重语言模型(1.3B–8B 参数;OPT、GPT-J、LLaMA-1/2/3、Mistral、Qwen 2.5)进行了系统扫描,在最近邻取整(RTN)下逐一量化单个注意力投影,并在七个模型上额外测试了 3 比特和 4 比特下的 GPTQ,共记录了 3,808 次独立测量的重构误差、困惑度变化以及逐投影的激活加权量化误差。我们发现:(1)在给定组件类型(Q、K、V 或 O)内,RTN 下 36 个案例中有 27 个的重构误差对困惑度敏感性的方差解释不足 10%,中位 R² = 0.044;(2)在全部 9 个模型中,组件类型和层序号对方差的解释力均超过重构误差,其中层序号在 9 个模型中的 7 个里是最强预测因子,其余 2 个模型中组件类型最强;(3)值(V)投影是最常占主导地位的组件,在九个模型中的七个里占总正 Delta PPL 的 38–51%;(4)主导组件在 RTN 与 GPTQ 之间基本保持一致(7 个案例中的 5 个);(5)对于 V 投影,激活加权量化误差作为组件内预测因子优于重构误差(中位 R² 为 0.20 对 0.06)。这些发现表明,仅凭相对权重重构误差不足以支撑感知敏感性的比特分配,且 V 投影在混合精度方案中值得专门考虑。
cs.LG / 84 / 2609.26199

Partially Observed Sparse Graphs: The Unknown Sampling Rate is a Tail Index

Xu, Jian, Zeng, Delu, Paisley, John, Zhao, Qibin
Abstract
A large graph is often available only in part: a crawl stopped by its budget, a panel, a partial dump. When the sampled fraction $s$ is known by design the total edge count follows from $\hat e=e_s/s^2$ and no model is needed. We treat the case where $s$ is unknown and the population size is known. Our main result is a reduction: under a sparse exchangeable (graphex) model the expected non-isolated fraction obeys $n_s/n_1\to s^{1+\sigma}$, so the sampling rate becomes estimable once the tail index $\sigma$ is, and substituting it back gives $e_s(n_1/n_s)^{2/(1+\sigma)}$ -- the same estimator, with the design quantity inferred. Estimating global edge cardinality in a sparse graph is therefore, in expectation, tail-index estimation, and the quadratic graphon estimator is the case $\sigma=0$: it fails by an identity rather than by a fit ($260\%$ median error against $27\%$). We bound the finite-size error of the substitution and show the reduction is \emph{modular} in the tail-index estimator --- filled with a published closed-form one it reaches $21.7\%$ over $13$ networks and $39$ sampling budgets with no fitting at all. Fitting a full graphex additionally returns the degree distribution at any size and a generative object, in a representation where sparsity is a coordinate and the interpolation path is dictated rather than chosen. Two limits are exact: rank-one graphexes have transitivity fixed by the degree profile, so high-clustering graphs lie outside the class; and under snowball or random-walk crawls every method here fails, the design-based oracle worst of all ($7.8\%$ to $588\%$).
cs.LG / 85 / 2609.26214

Bridging the Data Gap: Digital Twin as a New Paradigm for AI-based Radio Sensing

弥合数据鸿沟:数字孪生作为基于AI的无线感知新范式
Sainte-Beuve, Éloi, Larue, Guillaume, Dufrène, Louis-Adrien, Lampin, Quentin, Khansa, Ali Al
Abstract
We present a methodology that places a 3D digital twin (DT) of the environment as the main enabler behind the development of radio sensing at scale. The DT acts as a world model, providing geometry, materials, and transmitter/receiver placements to a ray-tracing engine that generates time-indexed channel impulse responses (CIRs) for large numbers of plausible scenes (moving people and objects, layout variants, seasonal/weather conditions, etc). From these synthetic sequences, we train a sequential neural network that maps CIR time series to spatial occupancy estimates, enabling device-free localization (DFL) without instrumented targets. We posit that sensing is best approached as an environment-conditioned learning problem: rather than seeking a single global model, we advocate training or fine-tuning local models specialized to a site-specific DT. As a first experiment, we introduce a novel State Space Model architecture, trained and evaluated across multiple room geometries. The localization performances obtained demonstrate the potential of the approach.
Chinese Translation
我们提出了一种方法论,将环境的三维数字孪生(Digital Twin, DT)作为大规模无线感知发展的核心使能技术。数字孪生充当世界模型,为光线追踪引擎提供几何结构、材质以及发射机/接收机的部署位置,从而为大量合理场景(如移动的人员与物体、布局变体、季节/天气条件等)生成带时间索引的信道冲激响应(Channel Impulse Response, CIR)。基于这些合成序列,我们训练了一个序列神经网络,将CIR时间序列映射为空间占用估计,从而实现对无设备目标的无感知定位(Device-Free Localization, DFL)。我们认为,感知最好被视为一个以环境为条件的学习问题:与其寻求单一的全局模型,我们主张针对特定场地的数字孪生训练或微调专用的局部模型。作为首个实验,我们提出了一种新颖的状态空间模型(State Space Model)架构,并在多种房间几何结构下进行训练和评估。所获得的定位性能验证了该方法的潜力。
cs.LG / 86 / 2609.26216

Beyond Imitation: Auditing the Recoverability of Reasoning in Distilled Models

Li, Ruitong, Guo, Binjie, Mo, Aisheng, Su, Guowei, Wang, Han, Li, Jie, Zhang, Ru
Abstract
A correct teacher solution becomes useful supervision when the receiving student can continue its reasoning. We measure this compatibility with prefix recovery: after revealing 25%, 50%, or 75% of a verified solution, we test whether the student completes it correctly. We connect recovery to the cosine conflict between cross-entropy and reverse-KL gradients over the full vocabulary. Across adjacent Qwen3 teacher-student pairs from 0.6B to 8B parameters, reverse-KL distillation delivers its most consistent mathematical and code improvements for the two students below 2B parameters. On a fixed cohort of 1,000 trajectories, average prefix recovery rises from 71.0% to 91.9% as student size increases from 0.6B to 4B, and the robust-fragile recovery gap contracts from 46.0 to 14.4 percentage points. With the teacher fixed at 8B, conflict separation falls from 0.993 to 0.233. An independent objective intervention finds the largest reverse-KL rescue on fragile trajectories. The three measurements locate the same capacity-dependent transfer regime: distribution matching has the greatest headroom when correct traces remain unevenly recoverable. Prefix recovery provides a practical diagnostic for selecting costly distribution-level distillation.
cs.LG / 87 / 2609.26217

MSA-CITE: A Co-Adapted LoRA Specialist Ecology for Fixed-Budget Small-Model Inference

MSA-CITE:面向固定预算小型模型推理的协同适配LoRA专家生态
Li, Ruitong, Guo, Binjie, Mo, Aisheng, Su, Guowei, Li, Jie, Zhang, Ru
Abstract
Compact language models are typically deployed by retaining a single post-training checkpoint and sampling it repeatedly. In this work, we challenge this practice by treating multiple discarded checkpoints as composable assets for deployment. Starting from a single Qwen3-4B backbone, we preserve four frozen LoRA branches, each derived from a different post-training trajectory. stead of drawing four generations from one branch, we allocate a fixed four-generation budget by sampling one completion from each branch. Our method, Multi-path Specialist Adaptation with Calibrated Inference-Time Evidence (MSA-CITE), processes the resulting portfolio by grouping terminal answers into equivalence classes, scoring each class via summed calibration-derived source priors, and selecting a representative under deterministic tie-breaking rules. The readout stage does not learn from evaluation results, nor does it introduce additional generations, verifiers, or reranking steps. On 200 held-out mathematics items, the four-path portfolio achieves 65.5% accuracy, compared with 62.0% for the strongest single-branch baseline. On a 100-item subject-disjoint shift, it attains 42.0% versus 40.0%. Under in-distribution conditions, the improvements over homogeneous SFT and Online-OPD repetition are robust; results against the strongest baseline and under shifted conditions are not conclusive. Our findings offer a narrow but concrete contribution: post-training branches, even without co-training, can be collectively beneficial for deployment.
Chinese Translation
紧凑型语言模型通常在部署时保留单一后训练检查点并对其进行反复采样。在本工作中,我们对这一做法提出挑战,将多个被弃用的检查点视为可组合的部署资产。从单个Qwen3-4B骨干模型出发,我们保留四个冻结的LoRA分支,每个分支来自不同的后训练轨迹。我们不是从一个分支采样生成四次,而是将固定的四次生成预算分配为从每个分支各采样一个补全。我们的方法——基于校准推理时证据的多路径专家自适应(MSA-CITE)——通过将最终答案分组成等价类、利用校准得到的源先验之和为每个类打分、并在确定性平局规则下选出代表答案,来处理所得到的答案组合。读取阶段不从评估结果中学习,也不引入额外的生成、验证器或重排序步骤。在200个保留数学题目上,四路径组合达到65.5%的准确率,而最强的单分支基线为62.0%。在一个100题、学科不相交的分布偏移数据集上,其准确率为42.0%,对比40.0%。在分布内条件下,相对于同质SFT和Online-OPD重复采样的提升是稳健的;而相对于最强基线以及分布偏移条件下的结果并不具有结论性。我们的发现提供了一个狭窄但具体的贡献:后训练分支即使没有协同训练,也可以在部署中共同带来收益。
cs.LG / 88 / 2609.26231

High-Order Liquid Evidence Modeling for Continuous and Subtle GNSS Spoofing Detection in Autonomous Driving

Sabir, Muhammad Ayub, Pang, Junbiao, Ashraf, Fatima
Abstract
Continuous and subtle GNSS spoofing poses a serious threat to autonomous vehicles because forged positions may remain locally plausible while gradually becoming inconsistent with vehicle motion observed by non-GNSS onboard sensors. Existing AV-oriented detectors commonly rely on residual thresholds or feature-level classification and provide limited modeling of how weak GNSS--motion inconsistency develops and persists over time. This paper formulates subtle GNSS spoofing detection as a causal sequential evidence-modeling problem and proposes a high-order liquid evidence detector. The method first compares the displacement implied by consecutive GNSS positions with that inferred from independent onboard motion observations and converts their difference into uncertainty-normalized residual evidence. It then represents the current inconsistency, its local evolution, excess above the normal level, accumulated persistence, and displacement validity as causal weak evidence. These cues are mapped into instantaneous, evolutionary, and persistent latent states, aligned through a bounded Kirchhoff-inspired symmetric exchange, and combined through an explicit third-order interaction to capture their coordinated support for spoofing. To model how this coordinated evidence develops over time, second-order liquid dynamics track its memory and evolution to estimate causal spoofing probabilities, which are converted into confirmed alarms using validation-selected threshold and persistence parameters. Experiments on the AV--GPS dataset family demonstrate strong controlled and external generalization, together with clear sequential alarm behavior. On Dataset-1, the proposed detector achieves an AUROC of 0.9932 and an AUPRC of 0.9843, while obtaining the lowest false-positive rate among the learning-based baselines. Code: https://github.com/pangjunbiao/HO-LLN-Spoofing.
cs.LG / 89 / 2609.26242

Can You Delete a Year of Market Data? Machine Unlearning Against Exact Retraining Oracles

你能删除一年的市场数据吗?针对精确重训练预言机的机器遗忘
Ye, Junyi
Abstract
When a data license expires, deleting stored records does not remove influence encoded in a trained forecaster. Machine unlearning seeks to remove this influence without retraining. We benchmark temporal unlearning with 3,200 paired references trained on all data and oracles retrained without the requested period. The grid covers five architectures, four rolling folds, five deletable years, and three experimental deletion levels on an S&P 500 volatility panel. The 2020 COVID crisis year produces the largest memorization gap for every architecture. Removing it improves all three deployable models in every fold, with the largest improvement in the 2022 bear market, while the two non-deployable models respond inconsistently. The target for approximate unlearning is the oracle, not low predictive accuracy on the deleted period. In one Transformer cell, an oracle that never trained on 2020 still predicts it at an information coefficient of 0.51, compared with 0.55 for the reference; pushing predictions toward noise reduces test skill. Across twelve deployable architecture-method pairs, only TSMixer with the hinge method remains near the oracle in every fold, closing 74-118% of the reference-to-oracle gap without a measurable loss of test skill. Method rankings vary across architectures and rolling windows. Audit separation rises with prior memorization but can remain small after exact deletion. The window-level loss comparison reaches at most 0.69, and treating stock-level windows as independent inflates the absolute t-statistic by a median factor of 1.9. These results call for an explicit deletion scope, oracle validation for the relevant architecture and window, and power-aware auditing.
Chinese Translation
当数据许可到期时,删除存储的记录并不能消除已编码在训练好的预测模型中的影响。机器遗忘旨在无需重新训练的情况下消除这种影响。我们通过3,200对参照模型(在全部数据上训练)与预言机(在不包含所请求时期的重训练数据上训练)对时间遗忘进行了基准测试。实验网格覆盖五种架构、四个滚动折、五个可删除年份,以及在标普500波动率面板上的三个实验性删除级别。2020年新冠疫情危机年份在所有架构中都产生了最大的记忆差距。移除该年份在每一折中都提升了全部三个可部署模型,其中在2022年熊市中提升最大,而两个不可部署模型的反应则不一致。近似遗忘的目标是预言机,而非在删除时期上的低预测精度。在一个Transformer单元中,从未在2020年数据上训练过的预言机仍以0.51的信息系数预测该年份,而参照模型为0.55;将预测推向噪声反而会降低测试技能。在十二个可部署的架构-方法组合中,只有采用hinge方法的TSMixer在每一折中都保持接近预言机,在测试技能无可测损失的情况下缩小了参照模型与预言机之间74-118%的差距。方法排名因架构和滚动窗口而异。审计分离度随先前记忆程度上升,但在精确删除后可能仍然很小。窗口级别的损失对比最多达到0.69,而将个股级别的窗口视为独立样本会使绝对t统计量中位数膨胀1.9倍。这些结果要求明确删除范围、针对相关架构和窗口进行预言机验证,以及进行统计功效感知的审计。
cs.LG / 90 / 2609.26257

Information-Theoretic Decoupled Prompt Tuning for Continual Learning

面向持续学习的信息论解耦提示调优
Zhang, Yunfei, Wen, Wen, Gong, Tieliang, Zhang, Weizhan
Abstract
Continual learning (CL) aims to incrementally acquire knowledge from sequential data while avoiding catastrophic forgetting. Recently, prompt tuning has attracted increasing attention as an efficient approach for adapting pre-trained models to CL tasks. However, existing prompt design paradigms commonly suffer from retrieval dependence and classifier bias, which make model adaptation sensitive to prompt selection and bias predictions toward newly arrived classes. To address these challenges, we propose Decoupled Prompt Tuning for Continual Learning (DPT4CL), which decouples the CLIP textual prompt into a task-shared prompt distribution and class-specific prompts. The task-shared prompt distribution is derived by optimizing an Information Bottleneck objective to facilitate cross-task knowledge transfer and alleviate classifier bias, while class-specific prompts enhance inter-class separability without relying on explicit prompt retrieval. Furthermore, we establish a unified excess risk bound from an information-theoretic perspective, providing theoretical support for the robust generalization and forgetting mitigation of the proposed framework. Extensive experiments on standard CL benchmarks demonstrate that DPT4CL achieves state-of-the-art performance. The source code is available at https://github.com/Cloudfly-Z/DPT4CL
Chinese Translation
持续学习(Continual Learning, CL)旨在从序列数据中增量地获取知识,同时避免灾难性遗忘。近年来,提示调优(prompt tuning)作为一种将预训练模型高效适配到持续学习任务的方法,受到了越来越多的关注。然而,现有的提示设计范式普遍存在检索依赖和分类器偏差问题,这使得模型适配对提示选择高度敏感,并导致预测偏向新到来的类别。为应对这些挑战,我们提出了面向持续学习的解耦提示调优方法(Decoupled Prompt Tuning for Continual Learning, DPT4CL),该方法将CLIP的文本提示解耦为任务共享的提示分布和类别专属提示。任务共享的提示分布通过优化信息瓶颈(Information Bottleneck)目标得到,以促进跨任务知识迁移并缓解分类器偏差;类别专属提示则在不依赖显式提示检索的情况下增强类间可分性。此外,我们从信息论视角建立了一个统一的超额风险界,为所提框架的鲁棒泛化能力和遗忘缓解提供了理论支撑。在标准持续学习基准上的大量实验表明,DPT4CL达到了最先进的性能。源代码可在 https://github.com/Cloudfly-Z/DPT4CL 获取。
cs.LG / 91 / 2609.26272

Mode Collapse Is Cheap to Detect: A Ground-Truth-Free Pre-Flight Check for Neural Samplers

模式坍塌的检测代价很低:一种无需真值的神经采样器预检方法
Xu, Jian
Abstract
Neural samplers are trained against an unnormalised target $\tilde\pi=e^{-E}$ with no samples from $\pi$, which leaves the practitioner with no way to tell whether an expensive training run has silently dropped part of the target. The diagnostics in common use are computed from the model's own draws and are therefore confined to the model's support: we exhibit a sampler whose self-normalised effective sample size is $0.99$ while it misses $87\%$ of the target mass. We argue that \emph{detecting} missing mass is a strictly easier problem than sampling it: detection needs one point per missed basin plus a local curvature estimate, whereas correction needs the sampler retrained. We turn this into a pre-flight check that consumes a few percent of the sampler's own training budget and uses only $E$, $\nabla E$ and $\nabla^2 E$. On Gaussian-mixture, Many-Well and rotated anisotropic Many-Well targets with exactly computable ground truth, the check estimates the missing mass to within $10^{-3}$ at $2.7\%$ of training cost, where a tuned annealed SMC reference needs $70$--$280\%$ of training cost to do worse. It also applies unchanged to a controlled-SDE sampler that has no tractable density, where ESS and the ELBO cannot be formed at all. The estimator carries a \emph{self-diagnostic} that, without ground truth, is conservative in the safe direction: across $60$ configurations it clears $16$, of which $15$ are accurate to $10^{-2}$ or better. We are explicit about what this does and does not license: the check cheaply produces evidence of missing mass, and sometimes evidence that the search has stabilised, but it cannot certify a run, and its thresholds are heuristic. We then map the boundary of the method on a real physical landscape, LJ-13, and report where it fails and why.
Chinese Translation
神经采样器是在未归一化的目标分布 $\tilde\pi=e^{-E}$ 上训练的,且没有来自 $\pi$ 的样本,这使得研究者无法判断一次昂贵的训练是否已经悄然遗漏了目标的一部分。常用的诊断指标是从模型自身的采样中计算的,因此被局限于模型的支撑集之内:我们展示了一个采样器,其自归一化有效样本量为 $0.99$,却遗漏了 $87\%$ 的目标质量。我们论证,检测缺失质量是一个严格来说比采样它更容易的问题:检测只需要每个被遗漏的盆地中的一个点加上一个局部曲率估计,而纠正则需要重新训练采样器。我们将这一思想转化为一种预检方法,其消耗仅为采样器自身训练预算的百分之几,且只使用 $E$、$\nabla E$ 和 $\nabla^2 E$。在具有精确可计算真值的高斯混合、Many-Well 以及旋转各向异性 Many-Well 目标上,该检查以 $2.7\%$ 的训练成本将缺失质量估计到 $10^{-3}$ 的精度之内,而一个经过调优的退火 SMC 参考方法需要 $70$--$280\%$ 的训练成本却表现更差。该方法无需修改即可应用于没有可处理密度的受控 SDE 采样器,而在这种情况下 ESS 和 ELBO 根本无法计算。该估计器带有一个自诊断机制,在没有真值的情况下朝安全方向保持保守:在 $60$ 个配置中,它通过了 $16$ 个,其中 $15$ 个的精度达到 $10^{-2}$ 或更好。我们明确说明该方法能提供和不能提供什么:该检查能以低成本产生缺失质量的证据,有时还能产生搜索已稳定的证据,但它无法为一次训练运行提供认证,且其阈值是启发式的。随后,我们在真实的物理势能面 LJ-13 上描绘了该方法的适用边界,并报告了它在何处失效及其原因。
cs.LG / 92 / 2609.26275

JAMPR+/L2D: scalable neural heuristic for constrained vehicle routing problems in dynamic environment

JAMPR+/L2D:面向动态环境下带约束车辆路径问题的可扩展神经启发式方法
Soroka, Andrew, Meshcheryakov, Alex
Abstract
The vehicle routing problems with real-world constraints (we consider vehicles capacity limits, time windows constrains, pickup-and-delivery multi-depo --- CPDPTW) pose significant computational challenges. While classical exact and heuristic methods remain effective to solve problems of small/medium size ($N\lesssim100$), they often lack adaptability and scalability for larger logistics tasks. In this work, we show how JAMPR+/L2D RL deep learning model, proposed in to solve large CPDPTW problems can be adopted in the case of substantial changes of graph distance matrix. We test performance of JAMPR+/L2D model for medium-sized CVRP and VRPTW problems on CVRPLIB benchmarks: JAMPR+/L2D outperforms the state-of-the-art heuristic HGS in over 85\% of instances, achieving improvement in objective gap. We show that the JAMPR+/L2D model trained on CPDPTW problem, generalizes well for tasks with simpler constraints (CVRP, VRPTW), for different problem sizes and for moderate changes in distance matrixes. For more substantial changes in distance matrixes, we propose here to make fast finetuning of JAMPR+: on ORTEC data (for CPDPTW) the proposed strategy remarkably reduces the objective gap without full model retraining, what will give both accuracy and rapid inference of the model in the practical routing scenarios with distance matrix changes.
Chinese Translation
具有现实世界约束的车辆路径问题(我们考虑车辆容量限制、时间窗约束、取送货多场站问题——CPDPTW)带来了显著的计算挑战。尽管经典的精确方法和启发式方法在求解中小规模问题($N\lesssim100$)时仍然有效,但它们在更大规模的物流任务中往往缺乏适应性和可扩展性。在本工作中,我们展示了为求解大规模CPDPTW问题而提出的JAMPR+/L2D强化学习深度学习模型,如何应用于图距离矩阵发生重大变化的情形。我们在CVRPLIB基准测试上检验了JAMPR+/L2D模型在中等规模CVRP和VRPTW问题上的性能:JAMPR+/L2D在超过85%的实例上优于最先进的启发式算法HGS,并取得了目标函数差距的改进。我们表明,在CPDPTW问题上训练的JAMPR+/L2D模型,对于约束更简单的问题(CVRP、VRPTW)、不同规模的问题以及距离矩阵的适度变化,均具有良好的泛化能力。针对距离矩阵更大幅度的变化,我们在此提出对JAMPR+进行快速微调:在ORTEC数据(针对CPDPTW)上,所提出的策略无需对模型进行完整的重新训练即可显著减小目标函数差距,从而在距离矩阵发生变化的实际路径规划场景中同时保证模型的准确性和快速推理。
cs.LG / 93 / 2609.26280

On the Effect of Bit-Level Parameter Perturbations in Machine Learning and Deep Learning Models

Raghapur, Akanksha, Stamp, Mark
Abstract
In this chapter, we investigate how classical machine learning models respond to small, targeted modifications in their parameters. We compare and contrast these results to analogous experiments on deep learning models. For classical learning models, we consider Hidden Markov Models (HMM) and Support Vector Machines (SVM), and for comparison, we conduct analogous experiments involving Multilayer Perceptrons (MLP) and Long Short-Term Memory (LSTM) networks. When applied to the Drebin Android malware dataset, our results show that classical models are brittle, in the sense that a limited set of selected parameters can have a dramatic effect on model behavior. In a related set of experiments, we investigate the steganographic capacity of these same learning models, that is, the proportion of bits in model parameters that can be overwritten without having a significant adverse affect on a model. We find that classical models offer limited steganographic capacity due to their compact, parameter-efficient, and relatively sensitive parameter structure. In contrast, neural networks are parameter-redundant, enabling higher steganographic capacity, where modifications can be distributed across many parameters with minimal impact on performance. These results highlight differences in how classical and neural models respond to parameter changes, with clear implications for both robustness and hidden information embedding. Overall, this work provides a framework for understanding parameter sensitivity and steganographic capacity across different classes of learning models.
cs.LG / 94 / 2609.26288

Quantifying Protocol-Induced Uncertainty in Comparative Predictive-Model Evaluation: Evidence from Large-Scale Daily PM10 Forecasting

量化比较性预测模型评估中由评估协议引起的不确定性:来自大规模日均PM10预测的证据
da Silva, Rafael, Monahan, Kiersten
Abstract
Comparative studies of predictive models often end by ranking candidate models, yet these rankings depend on evaluation protocols whose influence is rarely treated as a source of uncertainty. We formalize this problem as protocol-induced ranking uncertainty and introduce a framework that compares ranking displacement caused by switching protocols with displacement produced by conventional choices within a fixed protocol. We quantify these effects using the Protocol Sensitivity Score (PSS) and a full-refit intraprotocol reference. We validate the framework in a large-scale sequential prediction study of daily PM10. Static-split and rolling-origin evaluation are compared across 425 European background stations and 365 US EPA monitors. Switching protocols produces mean PSS values of 0.801 and 0.772 and changes the selected model at 35.3% and 31.5% of stations, respectively. In Europe, intraprotocol perturbations with identical scored targets produce PSS values of 0.072 and 0.230, with winner-swap rates of 0.8% and 4.9%. Between-protocol displacement is therefore substantially larger than the selected within-protocol references. Expanding the candidate set from three to nine models increases the between-protocol winner-swap rate to 60.2% in Europe. The pattern also persists under a frozen protocol applied to held-out background and non-background stations. These results show that model-selection conclusions can depend materially on legitimate evaluation choices. We recommend reporting ranking stability under a small set of defensible intraprotocol perturbations alongside claims of model superiority.
Chinese Translation
预测模型的比较研究通常以对候选模型进行排序作为结束,然而这些排序依赖于评估协议,而评估协议的影响很少被视为一种不确定性来源。我们将该问题形式化为协议引起的排序不确定性(protocol-induced ranking uncertainty),并提出了一个框架,该框架将切换协议所导致的排序位移与固定协议内由常规选择所产生的位移进行比较。我们使用协议敏感性得分(Protocol Sensitivity Score, PSS)和完全重训练的协议内参考基准来量化这些效应。我们在一项大规模的日均PM10序列预测研究中对该框架进行了验证,比较了静态划分(static-split)评估与滚动起点(rolling-origin)评估,涵盖425个欧洲背景监测站和365个美国EPA监测站。切换协议产生的平均PSS值分别为0.801和0.772,并分别在35.3%和31.5%的监测站上改变了所选模型。在欧洲,针对相同评分目标的协议内扰动产生的PSS值为0.072和0.230,获胜模型交换率分别为0.8%和4.9%。因此,协议间位移显著大于所选的协议内参考基准。将候选模型集从三个扩展到九个模型后,欧洲的协议间获胜模型交换率上升至60.2%。在冻结协议应用于保留的背景与非背景监测站时,该模式依然存在。这些结果表明,模型选择的结论可能在很大程度上取决于合理的评估选择。我们建议,在声称模型优越性时,应同时报告在一小组可辩护的协议内扰动下的排序稳定性。
cs.LG / 95 / 2609.26300

CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

CompKV:面向长上下文大语言模型推理的补偿感知KV选择方法
Huang, Zhen, Yao, Ruizhe, Liu, Danyi, Chen, Xinrui, Li, Shuwei, Zhong, Siru, Cao, Zijian, Lai, Yushan, Guo, Mingming, Zheng, Weijie, Fu, Haohuan
Abstract
Despite their strong performance, large language models (LLMs) are bottlenecked by KV cache memory traffic during long-context inference. Sparse attention is widely used to accelerate LLM inference by computing exact attention over a selected subset of tokens. To recover the contribution of tokens excluded from exact attention, recent methods apply coarse-grained compensation to the omitted attention tail. However, existing methods typically select tokens based on attention mass and only then compensate for the unselected tokens. This decoupled design overlooks their interaction: selection should prioritize tokens that would leave the largest compensation error if omitted. To address this limitation, we introduce CompKV, the first compensation-aware sparse attention framework that divides tokens into blocks and explicitly optimizes selection for the downstream compensation mechanism. Our theoretical analysis shows that the residual left by block-level mean compensation is governed by both block attention mass and within-block logit variation. We approximate this residual using compact block-level statistics, yielding a deployable selection criterion. We further develop an efficient asynchronous implementation. Experiments on RULER and LongBench-Pro show that CompKV performs best among the evaluated sparse baselines while delivering up to a $6.85\times$ self-attention speedup over full attention.
Chinese Translation
尽管大语言模型(LLMs)性能优异,但在长上下文推理过程中,KV缓存的内存访问成为其性能瓶颈。稀疏注意力被广泛用于加速LLM推理,其做法是在所选 token 子集上计算精确注意力。为了弥补未被纳入精确注意力计算的 token 的贡献,近期方法对被省略的注意力尾部应用粗粒度补偿。然而,现有方法通常先基于注意力质量选择 token,然后再对未选中的 token 进行补偿。这种解耦设计忽视了二者之间的相互作用:选择应当优先考虑那些若被省略将导致最大补偿误差的 token。为解决这一局限,我们提出了 CompKV,这是首个补偿感知的稀疏注意力框架,它将 token 划分为块(block),并针对下游补偿机制显式地优化选择过程。我们的理论分析表明,块级均值补偿所留下的残差同时受块注意力质量和块内 logit 变异性的影响。我们利用紧凑的块级统计量来近似该残差,从而得到一个可实际部署的选择准则。我们进一步开发了高效的异步实现。在 RULER 和 LongBench 上的实验表明,CompKV 在所评估的稀疏基线方法中表现最佳,同时相比全注意力可实现最高达 $6.85\times$ 的自注意力加速。
cs.LG / 96 / 2609.26310

PreGS: A Parameter-Transfer-Based Multi-Expert Graph Neural Network for Node Classification

PreGS:一种基于参数迁移的多专家图神经网络节点分类方法
Cai, Zhicong, Zhang, Yinglong, Hong, Xiaoying, Xia, Xuewen, Xu, Xing
Abstract
Graph neural networks have achieved strong performance in node classification by aggregating information from graph neighborhoods. However, a single aggregation mechanism may be insufficient to capture diverse structural patterns across graph datasets. Moreover, independently training multiple structural branches can introduce substantial overhead without necessarily producing stable node representations. To address these issues, this paper proposes PreGS, a parameter-transfer-based multi-expert graph neural network framework. PreGS first pretrains a multi-head graph attention network (GAT) and transfers the linear transformation weights of its first-layer attention heads to multiple GraphSAGE experts. The transferred experts are frozen and used as complementary structural branches. The fused raw node features, GAT head representations, and GraphSAGE expert representations are fed into a multilayer perceptron (MLP), whose output is further fused with the pretrained GAT logits. Based on PreGS, we further develop PreGSv2, which introduces source-level weighting and a structural gating mechanism for adaptive multi-source feature integration. Experiments on eight public graph datasets show that PreGS and PreGSv2 achieve competitive performance against representative graph neural network baselines. Ablation, parameter-transfer, sensitivity, aggregator, visualization, and training-time analyses further validate the effectiveness and stability of the proposed framework. The code and datasets are available at https://github.com/LH-Czc/PreGS.
Chinese Translation
图神经网络通过聚合图邻域信息在节点分类任务中取得了优异的性能。然而,单一的聚合机制可能不足以捕捉图数据集中多样的结构模式。此外,独立训练多个结构分支会带来显著的开销,且不一定能产生稳定的节点表示。为解决这些问题,本文提出了PreGS,一种基于参数迁移的多专家图神经网络框架。PreGS首先预训练一个多头图注意力网络(GAT),并将其第一层注意力头的线性变换权重迁移到多个GraphSAGE专家网络中。这些迁移后的专家网络被冻结,作为互补的结构分支使用。融合后的原始节点特征、GAT头表示以及GraphSAGE专家表示被输入到一个多层感知机(MLP)中,其输出进一步与预训练的GAT logits进行融合。基于PreGS,我们进一步开发了PreGSv2,引入了源级加权机制和结构门控机制,以实现自适应的多源特征融合。在八个公开图数据集上的实验表明,PreGS和PreGSv2相比代表性的图神经网络基线方法取得了具有竞争力的性能。消融实验、参数迁移、敏感性、聚合器、可视化以及训练时间分析进一步验证了所提出框架的有效性和稳定性。代码和数据集可在 https://github.com/LH-Czc/PreGS 获取。
cs.LG / 97 / 2609.26333

Disaggregated Quantization: Specializing LLM Prefill and Decode

解耦量化:面向大语言模型预填充与解码的专门化设计
Panferov, Andrei, Kleinegger, Maximilian, Priyadarshi, Sweta, Blankevoort, Tijmen, Alistarh, Dan
Abstract
Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while matching or exceeding its accuracy at 2-3-bit decode on both decode-heavy and prefill-heavy tasks. With released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint. To accommodate the additional checkpoint on a single device, offloaded disaggregated prefill (ODP) streams its weights from SSD, amortizing loading over prompt length. On the same 27B model, ODP delivers a 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp. We evaluate accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters.
Chinese Translation
预填充(prefill)和解码(decode)两个阶段适合采用不同的量化方法:低精度算术可加速提示词处理,而紧凑的权重可减少生成过程中的内存流量。我们提出“解耦量化”(Disaggregated Quantization, DQ),针对这两个阶段分别专门化计算格式、权重及存储位置。在 Qwen 3 和 Gemma 3 上,仅在解码阶段移除激活量化,即可在不增加推理成本的情况下提升以解码为主任务的准确率。训练一套独立的面向计算原生的预填充权重,相较于仅权重量化推理可加速提示词处理,同时在解码为主和预填充为主两类任务上,于 2-3 比特解码精度下达到相当甚至更优的准确率。基于已发布的 Qwen3.8-27B GGUF 解码器,训练一个 NVFP4 预填充器,在无需修改解码检查点的情况下,将 1 比特精度在 MMLU-Pro 上提升 32.5 分,在 MMMU-Pro 上提升 35.3 分。为了在单设备上容纳额外的检查点,卸载式解耦预填充(Offloaded Disaggregated Prefill, ODP)从 SSD 流式加载其权重,并将加载开销分摊到提示词长度上。在同一个 27B 模型上,当提示词长度为 8K 时,ODP 在 llama.cpp 中相较仅权重量化基线实现了 1.78 倍的首词生成时间(time-to-first-token)加速。我们在 vLLM 中评估了解耦服务下的准确率,并通过在参数量高达 2.8 万亿(2.8T)的模型上进行训练后量化,进一步验证了共享权重格式解耦的有效性。
cs.LG / 98 / 2609.26342

Geometry-Aware Hyperbolic Residual Quantization

Colombo, Alessio, Ayoughi, Melika
Abstract
Residual Vector Quantization turns continuous representations into discrete, multi-level token sequences. Yet most methods operate in Euclidean space, despite the coarse-to-fine structure of the resulting codes and the latent hierarchies present in many data domains. Hyperbolic geometry offers a natural alternative for hierarchical representations, but naive hyperbolic extensions introduce geometric inconsistencies: non-associative hyperbolic addition prevents consistent residual aggregation, while standard straight-through gradient estimation ignores the geometry of the latent space. We propose a geometry-aware hyperbolic residual quantization that addresses these issues in both the forward and backward passes. In the forward pass, Hyperbolic Residual Aggregation restores the telescoping behavior of residual quantization on the Poincare ball. In the backward pass, a discounted Hyperbolic Straight-Through Estimator routes the reconstruction gradient through the quantizer as a single geometric block, avoiding unstable recursive gradient transport across residual stages. Evaluations on hierarchical prediction, recommendation, image tokenization, and neural audio coding tasks show that our method improves the stability and structural organization of hyperbolic residual codes over naive hyperbolic baselines. At the same time, we observe a clear structure-compression trade-off: Euclidean residual quantization remains preferable for pure compression, while geometry-aware hyperbolic quantization is most useful for hierarchically organized discrete latent spaces.
cs.LG / 99 / 2609.26355

PACT: From Credit Assignment to Critic Alignment

Fu, Jiayan, Xu, Hang, Zhang, Yong, Luo, Zhaokai, Hu, Yao, Zhao, Dongyan, Chuan, Mu
Abstract
Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted mathematical definition, leaving its relationship to commonly used training signals unclear. We formulate three regularity conditions, namely Completeness, Prefix Consistency, and Neutrality, and prove that they uniquely determine token-level credit. This characterization provides a unified basis for explaining phenomena across existing algorithms and guides the development of an improved actor-critic training procedure. Through this lens, an ideal teacher in On-Policy Distillation (OPD) acts as an implicit critic, yielding an expected policy gradient proportional to that induced by token-level credit. Response-level REINFORCE Leave-One-Out (RLOO) signals match the expected policy-gradient contribution of token-level credit despite their coarser granularity. We further establish approximate credit sparsity under bounded outcome rewards and show how intermediate critic errors in Generalized Advantage Estimation (GAE) can become comparable to the underlying credit. These motivate Policy Aligned Critic Training (PACT), which adopts an Actor-then-Critic update order to apply importance sampling correction to critic training and better align the critic with the updated policy. In agentic mathematical reasoning, PACT achieves 72.87% average accuracy across four benchmarks, outperforming GRPO and PPO by 8.80 and 13.16 percentage points, respectively. On SWE-bench Verified, PACT achieves a pass rate of 67.4%, outperforming PPO, GRPO, and SAO by 2.4, 2.0, and 3.8 percentage points, respectively.
cs.LG / 100 / 2609.26377

FairMean: Promoting Fairness in Distributed Learning under Label Poisoning Attacks

FairMean:在标签投毒攻击下促进分布式学习中的公平性
Zheng, Huigan, Zhang, Jiaojiao, Liu, Yongxiang
Abstract
Fairness-aware distributed learning prioritizes clients with large losses to reduce performance disparities, but label poisoning can create large losses, thereby inducing a fairness--robustness conflict. We propose FairMean to manage this conflict. FairMean weights client gradients using a bounded, nondecreasing function of local loss. The increasing weights prioritize high-loss clients to promote fairness, while the upper bound prevents excessive loss-induced amplification of poisoned-client gradients. In the absence of label poisoning, we show that minimizing the FairMean objective is more conducive to solution fairness than minimizing the standard average-loss objective. Under label poisoning, we establish an average-stationarity bound whose attack-dependent term is proportional to the square of the poisoned-client fraction. Experiments show that FairMean promotes fairness by reducing accuracy variance while improving worst-client accuracy.
Chinese Translation
公平感知的分布式学习通过优先关注损失较大的客户端来缩小性能差距,但标签投毒攻击会制造较大的损失,从而引发公平性与鲁棒性之间的冲突。我们提出FairMean来管理这一冲突。FairMean使用一个有界的、关于局部损失非递减的函数对客户端梯度进行加权:递增的权重优先关注高损失客户端以促进公平性,而权重上界则防止因损失过大而过度放大被投毒客户端的梯度。在无标签投毒的情况下,我们证明最小化FairMean目标比最小化标准的平均损失目标更有利于解的公平性。在存在标签投毒的情况下,我们建立了一个平均平稳性界,其中依赖于攻击的项与被投毒客户端比例的平方成正比。实验表明,FairMean通过降低准确率方差并提升最差客户端的准确率来促进公平性。
cs.LG / 101 / 2609.26384

Learning to Defer with Guidance on Real World Medical Data

基于引导的真实世界医学数据上的学习推辞(Learning to Defer)
Sun, Emma, Strong, Joshua, Noble, Alison
Abstract
Medical image interpretation is high-volume and time-consuming, and while AI interpretation can reduce workload, fully autonomous deployment carries potential safety concerns and low specificity may in practice lead to increased clinician workload. Learning to Defer (L2D) addresses this by selectively routing cases between autonomous prediction and human experts by learning from input features and AI model and human performance. While theoretical guarantees have been proven for L2D, its performance has not been validated on real-world medical datasets with human reader annotations. We evaluate the predictor-rejector formulation of two-stage L2D, where the AI predictor model is fixed and separate from the trainable routing or rejector model, on Collab-CXR, a multilabel chest X-ray dataset with multiple human annotations per case. This is the first work to look at L2D in the context of real-world medical imaging data with human annotations. We further introduce a new setup, L2D with Guidance, where the decision space is extended to three choices: predict autonomously, defer to a human expert, or defer to a human expert and provide AI guidance. We compare multiple rejector architectures and loss functions, and different input feature availabilities. This is reproduced on two larger datasets, VinDr-CXR and CheXpert. Our results show that two-stage L2D with Guidance outperforms classic two-stage learning to defer, as well as human-alone, AI-alone and AI-guided human baselines. Notably, this performance is achieved with simpler loss functions compared to formally defined L2D surrogate loss functions in current literature.
Chinese Translation
医学影像解读工作量大且耗时,尽管AI解读可以减轻工作负担,但完全自主部署存在潜在的安全隐患,且低特异性在实践中可能导致临床医生工作量增加。学习推辞(Learning to Defer, L2D)通过从输入特征以及AI模型和人类的表现中学习,在自主预测与人类专家之间有选择地分配病例,从而解决这一问题。尽管L2D已有理论保证,但其在带有人类阅片标注的真实世界医学数据集上的性能尚未得到验证。我们在Collab-CXR上评估了两阶段L2D的预测器-拒绝器(predictor-rejector)形式,其中AI预测器模型固定,与可训练的路由(拒绝)模型分离;该数据集是一个多标签胸部X光数据集,每个病例包含多个人工标注。这是首个在带有真实世界人类标注的医学影像数据背景下研究L2D的工作。我们进一步引入了一种新的设置——带引导的L2D(L2D with Guidance),将决策空间扩展为三个选择:自主预测、交给人类专家、或交给人类专家并提供AI引导。我们比较了多种拒绝器架构和损失函数,以及不同的输入特征可用性。该实验在两个更大的数据集VinDr-CXR和CheXpert上得到复现。结果表明,带引导的两阶段L2D优于经典的两阶段学习推辞,也优于仅人类、仅AI以及AI引导下的人类等基线方法。值得注意的是,与现有文献中形式化定义的L2D替代损失函数相比,这一性能是使用更简单的损失函数实现的。
cs.LG / 102 / 2609.26389

TimeInteract: Towards Real-Time Interactive Intelligence for Streaming Time Series

Pan, Sheng, Gu, Yongli, Guo, Yiqing, Jin, Warren, Du, Bo, Pan, Shirui, Jin, Ming
Abstract
Real-world time series evolve continuously, with meaningful changes potentially emerging at any moment. However, existing time-series language models (TSLMs) remain inherently static. They either receive complete sequences for offline processing or alternate between streaming input and response generation, which prevents processing of new observations during interaction. We introduce a new regime, Time-Series Interaction: a model continuously perceives incoming time-series observations and user intent, autonomously decides when to remain silent or respond, and continues processing new observations during response generation. To realize this, we develop TimeInteract with three key designs: a dual-view streaming TS encoder that captures local variations and historical dynamics, a response control mechanism that learns when to trigger a response, and a decoupled streaming inference mechanism that separates control from response generation to avoid blocking subsequent observations. We further formulate a hierarchy of interaction capabilities, progressing from Understanding to Adaptivity. Based on this hierarchy, we construct StreamTSI-34K, a large-scale streaming TS interaction dataset with 34,588 episodes and 77,505 responses across synthetic and real-world time series in single- and multi-turn settings. Across all four interaction levels, TimeInteract consistently outperforms existing LLMs, VLMs, and TSLMs, with gains of up to 23.92 points on challenging tasks. It also improves response triggering while achieving near-zero stream stall and up to $2.15\times$ inference speedup.
cs.LG / 103 / 2609.26392

Double Descent and Malign Overfitting in Diffusion Models

扩散模型中的双重下降与恶性过拟合
Urfin, Raphaël, Bonnaire, Tony, Biroli, Giulio, Mézard, Marc
Abstract
Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularization, interpolating models generalize well, the test error following a double-descent curve. One might expect the same benign overfitting for diffusion models, whose training reduces to regression, i.e. to minimizing a quadratic score-matching loss. Yet the opposite is observed: overfitting here is catastrophic, driving the model into a memorization regime. We resolve this paradox by combining experiments on U-Nets trained on CelebA with a random-features model for which we derive closed-form learning curves. We show that with a fixed number $m$ of noise realizations per training sample, an interpolation peak does occur, but at $p\sim nm$ rather than at $p\sim n$ as in standard regression. The rise of the test loss, however, sets in much earlier, at $p\sim n$, independently of $m$. This overfitting is malign because, although the implicit regularization of training is fully at work, it drives the model toward the empirical score, which memorizes the training set, rather than toward the true score. A bias-variance decomposition pinpoints the mechanism: the bias of the score estimator starts to grow at $p\sim n$; past the peak the variance decays, as in regression, whereas the bias keeps growing and both saturate at a large value. Since diffusion models are trained with $m\gg1$, the peak is pushed to very large model sizes, and therefore sit on the rising branch that precedes it, where malign overfitting is already in play. Nevertheless, overparameterization remains beneficial when paired with regularization: in the random-features theory and in U-Net experiments, optimally regularized large models---via a ridge penalty or early stopping, respectively---outperform any unregularized models.
Chinese Translation
深度学习的传统观点认为,过参数化(即参数量 $p$ 大于训练样本数 $n$)是无害的:更大的模型泛化能力更好,即使没有正则化,插值模型也能良好泛化,测试误差呈现双重下降曲线。人们或许会预期扩散模型也具有同样的良性过拟合,因为其训练可归结为回归问题,即最小化二次得分匹配损失。然而观察到的现象恰恰相反:这里的过拟合是灾难性的,会使模型进入记忆化状态。我们通过将 CelebA 数据集上训练的 U-Net 实验与一个可推导出闭式学习曲线的随机特征模型相结合,解决了这一悖论。我们证明,在每个训练样本使用固定数量 $m$ 的噪声实现时,插值峰值确实会出现,但出现在 $p\sim nm$ 处,而非标准回归中的 $p\sim n$ 处。然而,测试损失的上升要早得多,在 $p\sim n$ 处就开始出现,且与 $m$ 无关。这种过拟合是恶性的,因为尽管训练的隐式正则化完全发挥作用,它却将模型推向记忆训练集的经验得分,而非真实得分。偏差-方差分解揭示了其机制:得分估计器的偏差在 $p\sim n$ 处开始增长;越过峰值后,方差如回归中那样衰减,而偏差却持续增长,二者最终都饱和于一个较大值。由于扩散模型是在 $m\gg1$ 的条件下训练的,峰值被推向非常大的模型规模,因此实际模型处于峰值之前的上升分支上,此时恶性过拟合已经发生。尽管如此,过参数化在与正则化相结合时仍然是有益的:在随机特征理论和 U-Net 实验中,经过最优正则化的大模型(分别通过岭回归惩罚或早停实现)优于任何未正则化的模型。
cs.LG / 104 / 2609.26402

OMatG-flash: An All-Atom Flow Map with Reinforce Adjoint Matching for Scalable Materials Discovery

Egg, Thomas, Sullivan, Harry Winston, Tadmor, Ellad B., Martiniani, Stefano
Abstract
The discovery of novel inorganic materials drives technological breakthroughs in critical fields such as computing and energy storage. Generative AI has promised to accelerate the materials discovery pipeline, but state-of-the-art flow and diffusion models remain bottlenecked by the cost of proposing candidate materials. To address this, we introduce OMatG-flash, an all-atom flow map for inorganic crystal structure prediction (CSP) and de novo generation (DNG). OMatG-flash is a Pareto-optimal inference engine for materials, sampling candidate materials with an order of magnitude fewer inference steps and less wall-clock time than existing flow and diffusion models while demonstrating benchmark performance on par with the state-of-the-art. To enable post-training fine-tuning we apply Reinforce Adjoint Matching to flow maps, further improving match rates and RMSE on the unconditional CSP task. OMatG-flash showcases the potential of flow maps to accelerate generation of high-quality candidate inorganic materials and demonstrates a step forward in sample throughput necessary for data-hungry materials discovery workflows.
cs.LG / 105 / 2609.26426

DeepFEAv2: Deep Learning for Transient Finite Element Analysis Beyond Structured Meshes

DeepFEAv2:超越结构化网格的瞬态有限元分析深度学习方法
Triantafyllou, Georgios, Kalozoumis, Panagiotis G., Iakovidis, Dimitris K.
Abstract
Finite Element Analysis (FEA) is widely used for transient mechanical simulations, but its high computational cost limits real-time and high-resolution applications. Deep learning surrogate models can reduce this cost; however, many existing approaches are restricted to steady-state prediction or cannot jointly predict Node- and Element-based Outputs (NEO) over time. The state-of-the-art DeepFEA framework has addressed these issues but remains limited to structured finite element (FE) meshes. To overcome this limitation, this study proposes DeepFEAv2, a deep learning surrogate framework that enables prediction of transient FEA simulations across different FE mesh topologies and element types. The main contributions of DeepFEAv2 are: (a) a module that uses the FE connectivity matrix to organize input features by element and arrange them into an input sequence guided by the mesh topology; (b) a novel neural network architecture designed to process the input sequence and jointly predict NEO over time; and (c) a FEA-informed optimization strategy for regularizing these NEO predictions. DeepFEAv2 was evaluated on structured and unstructured 3D linear elastic datasets, as well as on a pressure-driven aortic valve dataset. DeepFEAv2 achieved R^2 values up to 0.99 and normalized errors as low as 0.38%. Compared with DeepFEA, it achieved up to 38.0% relative increase in R^2 and up to 87.1% reduction in normalized error. DeepFEAv2 also performed inference up to three orders of magnitude faster than traditional FEA. These results demonstrate that DeepFEAv2 can efficiently model transient FEA simulations across increasingly complex FE settings, providing a scalable surrogate framework for transient FEA.
Chinese Translation
有限元分析(Finite Element Analysis, FEA)被广泛用于瞬态力学仿真,但其高昂的计算成本限制了实时和高分辨率的应用。深度学习代理模型可以降低这一成本;然而,许多现有方法仅限于稳态预测,或无法随时间联合预测基于节点和基于单元的输出(Node- and Element-based Outputs, NEO)。最先进的 DeepFEA 框架已经解决了这些问题,但仍局限于结构化有限元(FE)网格。为克服这一局限,本研究提出了 DeepFEAv2,一个能够在不同有限元网格拓扑和单元类型上进行瞬态 FEA 仿真预测的深度学习代理框架。DeepFEAv2 的主要贡献包括:(a) 一个利用有限元连接矩阵按单元组织输入特征,并根据网格拓扑将其排布为输入序列的模块;(b) 一种新颖的神经网络架构,用于处理该输入序列并随时间联合预测 NEO;(c) 一种基于 FEA 信息的优化策略,用于对这些 NEO 预测进行正则化。DeepFEAv2 在结构化和非结构化三维线弹性数据集以及一个压力驱动的主动脉瓣数据集上进行了评估。DeepFEAv2 的 R^2 值高达 0.99,归一化误差低至 0.38%。与 DeepFEA 相比,其 R^2 相对提升最高达 38.0%,归一化误差降低最高达 87.1%。此外,DeepFEAv2 的推理速度比传统 FEA 快最多三个数量级。这些结果表明,DeepFEAv2 能够在日益复杂的有限元设置下高效地建模瞬态 FEA 仿真,为瞬态有限元分析提供了一个可扩展的代理框架。
cs.LG / 106 / 2609.26435

One-Step Generative Surrogate Models via Block-Triangular Joint Drifting

Geissler, Nicholas, Jha, Shreya, Baptista, Ricardo, Peherstorfer, Benjamin
Abstract
Drifting provides a direct route to one-step generative models, but applying it directly to stochastic transition modeling requires multiple samples of the next state conditioned on the same current state. Standard trajectory data, however, typically provide only one realized next state for each observed current state and therefore do not provide an empirical approximation of the corresponding conditional distribution over possible next states. We introduce block-triangular joint drifting, which instead applies a projected drift field to the empirically accessible joint distribution of consecutive states. Importantly, the block-triangular architecture preserves the current-state marginal while making its second component a direct sampler of the conditional distribution of possible next states. The resulting surrogate generates stochastic trajectories with one model evaluation per time step, without auxiliary generative steps between time steps. Numerical experiments demonstrate accurate marginal and trajectory-dependent statistics and favorable accuracy-cost tradeoffs compared with deterministic, diffusion-, flow-, and distillation-based generative surrogate models.
cs.LG / 107 / 2609.26460

Can We Predict Anomaly Detection Performance from Embedding-Space Geometry?

Wilkinghoff, Kevin, Tan, Zheng-Hua
Abstract
Anomaly detection systems are often trained using normal data alone, while model selection and evaluation typically require labeled anomalies. We study whether anomaly detection performance can be predicted without access to anomalous data. For kNN-based detectors, we derive a lower bound on the area under the ROC curve (AUC) that relates detection performance to the separation between inlier and outlier scores and to their respective variances. Under a local scaling model, we use this bound to characterize how density variation, intrinsic-dimensional heterogeneity, and cross-domain mismatch contribute to score variability. We then investigate anomaly-free model selection and show that inlier score variance alone does not reliably predict performance across different representations. To address this limitation, we introduce simple pseudo-anomaly probes that provide a reference for estimating relative score separation. Experiments on the DCASE 2022-2025 benchmarks, spanning four embedding models and 208 candidate systems, show that pseudo-anomaly-based estimators substantially improve anomaly-free model selection. In particular, diverse pseudo-anomalies enable anomaly-free model selection to outperform conventional development-set selection under domain shift. These results show that embedding-space geometry contains predictive information about anomaly detection performance while also highlighting the representation-dependent nature of inlier-only performance estimates.
cs.LG / 108 / 2609.26487

When Recursive Models Finish Computing

Krishna, Hare, Singh, Shubham, Ebert, Stephen, Sun, Hao-Yu
Abstract
Recursive models can continue updating their latent states beyond their nominal inference budget, so an incorrect output at that budget does not show whether computation is unfinished or has entered a persistently unsuccessful regime. We study the dynamics of completion in attention- and MLP-based Tiny Recursive Models (TRMs) on 1,000 hard Sudoku puzzles. Extending recurrence from the nominal 16 steps to 512 steps increases cumulative exact-solve accuracy from 59.2% to 87.5% for the attention model and from 74.4% to 91.9% for the MLP model, solving more than two-thirds of the puzzles unsolved in the nominal budget. Across both architectures, latent-state motion drops sharply after the first exact solution. Completed states are typically locally contractive along the trajectory direction, even though the same local Jacobian retains strongly expanding directions. We characterize this phenomenon as trajectory-conditioned anisotropic stability. Perturbation experiments confirm this directional stability across both models. The multi-step fate of the maximally expanding direction differs: it is absorbed within 16 steps in the attention model but persists longer in the MLP model. The anisotropic-stability pattern also holds for a second attention checkpoint. Together, these results distinguish nominal-budget failure from completed computation and identify a common dynamical signature of completion across two recurrent architectures.
cs.LG / 109 / 2609.26508

Gap-Free Streaming PCA Beyond Rank-One Updates: Near-Optimal Rates and Applications to Differential Privacy

Gu, Anming, Kumar, Syamantak, Tian, Kevin, Yang, Chutong
Abstract
Streaming principal component analysis (PCA) seeks to recover a leading spectral subspace in a single pass over a data stream. We give a new analysis of the ubiquitous Oja's algorithm [Oja82] for the most general, gap-free variant of this problem, where no eigengap assumptions are made on the underlying mean matrix, complemented by a nearly-matching lower bound. Prior works achieving near-optimal rates for streaming PCA either required gap assumptions [JJK+16, HNWW21], or were limited to rank-one updates [AZL17, Lia23]. Our proof only uses a second moment bound on the individual stochastic updates, bypassing the almost sure bounds needed by prior near-optimal analyses, and the analogous offline matrix Bernstein bound. We also extend our result to a Rayleigh quotient notion of approximate PCA, addressing an open question of [JJK+16]. As our main application, we give gap-free differentially private PCA guarantees for sub-Gaussian data, settling Conjecture 1.1 of [Bro26] up to logarithmic factors.
cs.LG / 110 / 2609.26537

Notes on Fourier-Bessel wavelets

傅里叶-贝塞尔小波笔记
Venturotti, Marcel, Exarchakis, Georgios
Abstract
These notes develop the mathematical foundations and construction of a Fourier-Bessel wavelet family inspired by the disk harmonics of Shaqfa et al.[9]. We begin with the relevant properties of Bessel and modified Bessel functions and introduce the wavelet properties required for the construction. We then derive the Fourier-Bessel disk harmonics as solutions to the Helmholtz equation on the unit disk subject to a Neumann boundary condition. Building on this basis, we construct a wavelet family by applying a Gaussian spatial envelope and introducing a zero-mean correction for the zeroth angular order. We derive the corresponding normalisation constants for $L^2$-based applications and discuss $L^1$-based normalisation for frequency-domain peak consistency. Finally, we derive a closed-form Fourier-domain representation of the resulting wavelets. The main motivation is the approximately linear spacing, which converges to $\pi$ between consecutive radial eigenvalues. Rather than replacing the conventional dyadic organisation of wavelet families, this construction lays out the foundation to explore whether a more uniform radial frequency allocation can be useful for applications in which broad and balanced frequency coverage is desirable.
Chinese Translation
本文阐述了一类受Shaqfa等人[9]提出的圆盘调和函数启发的傅里叶-贝塞尔小波族的数学基础与构造方法。我们首先介绍贝塞尔函数和修正贝塞尔函数的相关性质,并给出该构造所需的小波特性。随后,我们将傅里叶-贝塞尔圆盘调和函数推导为单位圆盘上满足诺伊曼(Neumann)边界条件的亥姆霍兹方程的解。在此基础上,我们通过施加高斯空间包络并对零阶角向模式引入零均值修正,构造了一个小波族。我们推导了基于$L^2$的相应归一化常数,并讨论了基于$L^1$的归一化方法以保证频域峰值的一致性。最后,我们推导了所得小波的闭式频域表示。其主要动机在于相邻径向特征值之间近似线性的间距,该间距收敛于$\pi$。这一构造并非要取代传统小波族的二进组织方式,而是为进一步探索在需要宽广且均衡的频率覆盖的应用中,更均匀的径向频率分配是否有用奠定了基础。
cs.LG / 111 / 2609.26603

Towards Hierarchical GNNs for multi-grid power flow: generalization across operating scenarios

面向多电网潮流的分层图神经网络:跨运行场景的泛化
Femine, Carmine Delle, Atxaga, Leire Garin, Diaz-Iglesias, Asier, Herrera, Juan Pablo Maroto, Florez-Tapia, Ane Miren, Urziku, Marco Quartulli. Izaro Goienetxea
Abstract
Hierarchical latent communication improves the generalization of a multi-grid power-flow model to new operating scenarios. The module exchanges information through two reduced graphs within a GENCO-based corrective network. We compare Kron-derived transports, a same-anchor Quotient construction and a flat backbone in preliminary trainings of 200 epochs on three grid topologies, with three initialization seeds per model. Evaluation uses 200 newly generated, preselected scenarios per grid. On the training topologies, Kron reduces the macro family-balanced voltage error from 5.660 +- 0.899 to 0.851 +- 0.110: an 85.0% reduction relative to Flat GENCO and 31.0% relative to Quotient, which reaches 1.235 +- 0.225. Both hierarchical models outperform a per-bus mean fitted on training solutions on every training topology in all three seeds. These results demonstrate generalization across operating scenarios within the studied topologies, with one set of learned parameters shared across grids. Evaluation on two additional topologies distinguishes this achievement from cross-topology generalization: the current models do not yet outperform the fitted reference in that calibrated- transfer setting. This preprint presents the architecture and preliminary evidence for hierarchical communication as a component of multi-grid power-flow learning, with generalization to unseen topologies as the next development objective.
Chinese Translation
分层潜在通信(hierarchical latent communication)提升了多电网潮流模型对新运行场景的泛化能力。该模块在基于发电公司(GENCO)的校正网络中,通过两个精简图交换信息。我们在三种电网拓扑上进行了200个epoch的初步训练,每个模型使用三个随机初始化种子,并比较了基于Kron变换的信息传递方法、同锚点的商图(Quotient)构建方法以及扁平骨干网络。评估使用每个电网新生成并预先筛选的200个场景。在训练拓扑上,Kron方法将宏观家族平衡电压误差从5.660 ± 0.899降至0.851 ± 0.110:相对于扁平GENCO(Flat GENCO)降低85.0%,相对于商图方法(其误差为1.235 ± 0.225)降低31.0%。在所有训练拓扑的全部三个种子下,两个分层模型均优于基于训练解拟合的每母线均值基准。这些结果表明,在所研究的拓扑内,通过一组跨电网共享的学习参数即可实现跨运行场景的泛化。在另外两个拓扑上的评估则将这一成果与跨拓扑泛化区分开来:在那种校准迁移设定下,当前模型尚未超越拟合基准。本预印本介绍了作为多电网潮流学习组件的分层通信架构及其初步证据,对未见拓扑的泛化是下一个开发目标。
cs.LG / 112 / 2609.26621

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

贪心解码并非精度不变:大语言模型推理中的跨精度输出分歧
Du, Gaoyuan, Khan, Anam Nawaz, Zhou, Rex, Liu, Xiaoyang, Chakrabarti, Deepayan, Suya, Fnu, Li, Xueping
Abstract
Greedy decoding from large language models is commonly treated as deterministic. We show it is not precision-invariant: the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware. Across our evaluations of six models (1.1B-7B parameters, four families; divergence additionally characterised at 12B) and three benchmarks, 49-100\% of prompts diverge; a single token flip often cascades into trajectory-level divergence. We develop an empirical error-propagation analysis and find that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps; the outcome depends primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates. The analysis makes five testable predictions about intervention outcomes, including that applying more FP32 compute (broader scope) makes agreement worse. The experiments match all five predictions. The best-performing low-overhead intervention we evaluate, selective FP32 LM head recomputation, triggered only when the margin falls below a threshold, delivers +22-36 pp exact agreement on A10G (+12-21 pp on L4 and A100) at less than 4\% latency overhead in low-batch (batch size <=4) single-stream inference. We map the applicability boundary across six models and four batch sizes, and hypothesise that training-time precision stability is a determining factor. The method is a partial mitigation rather than a universal determinism guarantee: its benefit vanishes when body-originated error dominates, including at batch size >=8 and under end-to-end FP8 in our tests.
Chinese Translation
大语言模型的贪心解码通常被视为确定性的。我们证明它并非精度不变:相同的模型、提示词和解码算法在同一硬件上使用 BF16 与 FP16 时会产生不同的输出。在对六个模型(1.1B–7B 参数、四个模型家族;并在 12B 模型上对分歧进行了额外刻画)和三个基准的评估中,49–100% 的提示词出现分歧;单个 token 的翻转往往级联为轨迹级分歧。我们建立了经验性误差传播分析,发现累计 22 层的主体误差(body error)无法区分翻转步骤与非翻转步骤;结果主要取决于 LM 头部前两位 logits 的裕度(margin)相对于前两位候选之间方向性扰动的对比。该分析提出了关于干预结果的五项可检验预测,其中包括施加更多 FP32 计算(更广的作用范围)会使一致性变差。实验结果与全部五项预测相符。在我们评估的低开销干预方法中表现最佳的是选择性 FP32 LM 头部重计算——仅当裕度低于阈值时触发——在低批次(batch size ≤4)单流推理中,以低于 4% 的延迟开销在 A10G 上实现了 +22–36 个百分点的精确一致率(在 L4 和 A100 上为 +12–21 个百分点)。我们绘制了该方法在六个模型和四种批次规模下的适用边界,并假设训练时的精度稳定性是一个决定性因素。该方法只是部分缓解措施,而非普适的确定性保证:当主体来源的误差占主导时,其收益消失,包括在我们的测试中 batch size ≥8 以及端到端 FP8 的情况下。
cs.LG / 113 / 2609.26631

Label-Efficient Learning for Ground-Based Sky-Image Classification: A Benchmark of Transfer Learning, Active Learning, and Pseudo-Labeling on GCD

面向地基云图分类的标签高效学习:迁移学习、主动学习与伪标签在GCD数据集上的基准测试
Dagher, Esther Bou, Bu-Dager, Viktoriya, Zegarlinski, Boguslaw
Abstract
Accurate ground-based cloud classification is important for atmospheric monitoring, solar-energy forecasting, aviation weather assessment, and climate observation systems. However, reliable sky-image annotation is time-consuming, especially when cloud types are visually similar or mixed. We study the label efficiency of deep learning for ground-based cloud classification using the Ground-based Cloud Dataset (GCD). Rather than proposing a new architecture, we benchmark three practical strategies under limited annotation budgets: supervised transfer learning, uncertainty-based active learning, and high-confidence pseudo-labeling. An ImageNet-pretrained ResNet50 is used as a common frozen backbone, with experiments repeated over five random seeds for label budgets from $1\%$ to $100\%$ of the training labels. Supervised transfer learning is already highly label-efficient: test accuracy increases from $0.635 \pm 0.018$ with $1\%$ labels to $0.730 \pm 0.002$ with $40\%$ labels, approaching the full-label result of $0.735 \pm 0.003$. Active learning and pseudo-labeling are competitive with supervised sampling and provide small improvements for some metrics and budgets, but neither gives a large or consistent aggregate gain. Diagnostic analyses show that accepted pseudo-labels are reliable, with accuracy from $0.946$ to $0.977$, but biased toward easier high-confidence sky-type groups. In contrast, uncertainty sampling preferentially queries visually challenging groups, including Mixed and the confusable Stratocumulus and Cumulonimbus groups, but these targeted acquisitions yield only modest gains. Overall, transfer learning substantially reduces annotation requirements for GCD, while simple active and semi-supervised strategies provide limited additional benefit over a strong supervised baseline.
Chinese Translation
准确的地基云分类对于大气监测、太阳能预测、航空天气评估以及气候观测系统至关重要。然而,可靠的天空图像标注非常耗时,尤其是当云类型在视觉上相似或混杂时。我们基于地基云数据集(Ground-based Cloud Dataset, GCD)研究了深度学习在地基云分类中的标签效率。我们并未提出新的架构,而是在有限标注预算下对三种实用策略进行了基准测试:有监督迁移学习、基于不确定性的主动学习以及高置信度伪标签学习。实验采用ImageNet预训练的ResNet50作为统一的冻结骨干网络,并在标签预算为训练标签的1%到100%的条件下,使用五个随机种子重复实验。结果表明,有监督迁移学习本身就具有很高的标签效率:测试准确率从1%标签时的0.635 ± 0.018提升至40%标签时的0.730 ± 0.002,接近全标签下的0.735 ± 0.003。主动学习与伪标签学习可与有监督采样相媲美,并在部分指标和预算下带来小幅提升,但均未能带来显著或一致的总体增益。诊断分析显示,被采纳的伪标签是可靠的,准确率介于0.946至0.977之间,但偏向于较容易的高置信度天空类型组。相比之下,不确定性采样倾向于查询视觉上具有挑战性的类别组,包括混合云(Mixed)组以及易混淆的层积云(Stratocumulus)和积雨云(Cumulonimbus)组,但这些针对性采集仅带来有限的收益。总体而言,迁移学习大幅降低了GCD数据集的标注需求,而简单的主动学习和半监督策略相对于强大的有监督基线仅能提供有限的额外收益。
cs.LG / 114 / 2609.26667

MAGIC: Mixed-Granularity Agent Graphs via Incremental Construction with Dense-Reward Reinforcement Learning

MAGIC:基于增量构建与稠密奖励强化学习的混合粒度智能体图
Yang, Kairui, Yi, Ziheng, Li, Xunkai, An, Minghao, Liu, Zhanke, Chen, Zekai, Li, Rong-Hua
Abstract
Collaboration topology shapes both the performance and execution cost of LLM-based multi-agent systems. Because tasks differ in complexity and required capabilities, recent approaches generate task-specific collaboration graphs that specify agent participation and information flow. However, representative topology generators use either individual agents or predefined groups throughout an organization, overlooking differing collaboration needs across subtasks. Our key insight is to select granularity locally for each functional role, combining fine-grained control with reusable collaboration patterns within one organization. Learning such organizations requires exploring a combinatorial construction space with limited intermediate feedback from final-answer rewards. Therefore, we propose MAGIC, a dense-reward reinforcement learning framework for mixed-granularity graph generation. Specifically, MAGIC constructs a mixed-granularity agent graph by sequentially selecting a functional role, instantiating it as a single agent or reusable group, and connecting it to existing units. We directly optimize the construction policy using returns from trajectories sampled under the current policy and use potential-based reward shaping to provide intermediate feedback from probe-based utility and structural signals while preserving the cumulative task reward. MAGIC outperforms state-of-the-art baselines across eight benchmarks and demonstrates strong inference efficiency in our efficiency study.
Chinese Translation
协作拓扑结构决定了基于大语言模型(LLM)的多智能体系统的性能与执行成本。由于任务在复杂度和所需能力上各不相同,近来的方法会生成任务特定的协作图,以规定智能体的参与方式与信息流动。然而,现有的代表性拓扑生成器要么在整个组织中使用单个智能体,要么使用预定义的群体,忽视了不同子任务间协作需求的差异。我们的核心洞察是:针对每个功能角色在局部进行粒度选择,从而在同一组织内将细粒度控制与可复用的协作模式相结合。学习这样的组织结构需要在组合式的构建空间中进行探索,而来自最终答案奖励的中间反馈十分有限。为此,我们提出了 MAGIC——一个用于混合粒度图生成的稠密奖励强化学习框架。具体而言,MAGIC 通过依次选择一个功能角色、将其实例化为单个智能体或可复用群体,并将其与已有单元相连接,从而构建出混合粒度的智能体图。我们利用当前策略下采样得到的轨迹回报直接优化构建策略,并采用基于势函数的奖励塑形方法,从基于探测的效用信号和结构信号中提供中间反馈,同时保持累积任务奖励不变。MAGIC 在八个基准测试上超越了当前最先进的基线方法,并在效率研究中展现出较强的推理效率。
cs.LG / 115 / 2609.26679

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

Grokking(顿悟)现象的谱理论:权重衰减诱导特征学习
Pracher, Lenz, de Jong, Pascal, Lieshaus, Oskar, Jeffares, Alan, Rulands, Steffen
Abstract
In grokking an early fit to the training data separates from a much later improvement in generalization. During this delay, training can move from a fixed neural tangent kernel (NTK) regime to one in which task-relevant kernel eigendirections continue to evolve. We provide a quantitative theory for how this transition from lazy to rich learning can produce delayed generalization. For homogeneous networks trained with squared loss and $L_2$ weight decay, we show that a finite residual remains after memorization, with larger residual fractions in target components associated with smaller NTK eigenvalues. These residuals feed back into the dynamics of the NTK itself, and projecting the resulting dynamics onto task-relevant spectral directions yields a reduced system in which residual-driven kernel growth competes with weight decay. This system predicts that the grokking timescale is controlled by the product of learning rate and weight decay, that feature learning slows logarithmically near a critical decay above which task-aligned NTK structure can no longer support generalization, and that stronger decay can prevent fitting altogether. We test these predictions in modular addition. In a homogeneous MLP, task-aligned Fourier structure continues to emerge in the NTK after training accuracy has saturated, and an 84$\times$90-grid of trained networks across varying learning rate and weight decay recovers the predicted phase geometry and inverse-product scaling of the generalization time with learning rate and weight decay. A one-block Transformer shows similar macroscopic phase structure in a 42$\times$45-grid, as well as the same transition-time scaling despite violating exact homogeneity. Together, these results provide a mechanistic derivation connecting post-fit feature learning to both the onset of generalization and its phase structure in the learning rate and weight decay plane.
Chinese Translation
在 grokking(顿悟)现象中,模型对训练数据的早期拟合与远晚于此的泛化能力提升相互分离。在这一延迟期内,训练可以从固定的神经正切核(NTK)机制转变为任务相关的核特征方向持续演化的机制。我们为这种从“懒惰学习”到“丰富学习”的转变如何产生延迟泛化提供了定量理论。对于以平方损失和 $L_2$ 权重衰减训练的齐次网络,我们证明在记忆化完成后仍会残留一个有限的残差,且与较小 NTK 特征值相关联的目标成分具有较大的残差比例。这些残差反馈到 NTK 自身的动力学中,将所得动力学投影到任务相关的谱方向上,可得到一个简化的系统,其中残差驱动的核增长与权重衰减相互竞争。该系统预测:grokking 的时间尺度由学习率与权重衰减的乘积控制;特征学习在临界衰减值附近按对数规律减慢,超过该临界值后任务对齐的 NTK 结构不再能够支持泛化;而过强的权重衰减则可能使模型完全无法拟合训练数据。我们在模加法任务上验证了这些预测。在一个齐次的多层感知机(MLP)中,即使在训练准确率已经饱和之后,任务对齐的傅里叶结构仍在 NTK 中持续涌现;在不同学习率与权重衰减下训练的 84×90 网格网络恢复了预测的相图几何结构,以及泛化时间与学习率和权重衰减的乘积倒数标度关系。一个单块 Transformer 在 42×45 网格中表现出相似的宏观相图结构,并在违反严格齐次性的情况下仍呈现相同的过渡时间标度。综上,这些结果提供了一个机制性推导,将拟合后的特征学习与泛化的出现及其在学习率-权重衰减平面上的相图结构联系起来。
cs.LG / 116 / 2609.26708

Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

在量化模型实际走向之处训练:面向低比特推理的在策略蒸馏
Chen, Yuanteng, Liu, Zhilei, Wang, Peisong, Shao, Yuantian, Li, Chuangyi, Wang, Weining, Qiu, Shuang, Li, Gang, Liu, Jing, Cheng, Jian
Abstract
Quantization-aware distillation (QAD) restores much of the short-form question-answering performance lost to sub-3-bit quantization, yet leaves mathematical and code reasoning substantially impaired. Long generations often degenerate into repetitive loops, exhausting the decoding budget without completing a solution. We trace this gap to quantization-amplified exposure bias: QAD trains on fixed corpus prefixes, while quantization-induced deviations compound along the model's own autoregressive trajectories. To address this mismatch, we introduce an on-policy distillation (OPD) stage that places teacher supervision where the quantized model actually goes. Starting from a QAD checkpoint, the student generates through the quantized forward path used at deployment and receives feedback from a frozen full-precision teacher on its own prefixes, combining dense token-level guidance with task-verifier rewards. Across four models at 2.79 and 1.88 effective bits, OPD raises average BF16 performance retention from 35% to 70% on MATH-500 and from 66% to 91% on HumanEval while preserving short-form performance, with reasoning gains substantially exceeding those of continued teacher-forced QAD in matched-budget comparisons. By coupling QAD's stable low-bit initialization with OPD's on-policy reasoning recovery, our framework provides a comprehensive sub-3-bit solution that preserves broad capabilities while restoring long-form reasoning.
Chinese Translation
量化感知蒸馏(Quantization-aware distillation, QAD)能够恢复大部分因低于3比特量化而损失的短文本问答性能,但数学和代码推理能力仍受到显著损害。长文本生成常常退化为重复循环,耗尽解码预算却无法完成解答。我们将这一差距归因于被量化放大的曝光偏差(exposure bias):QAD 在固定的语料前缀上训练,而量化引起的偏差会沿着模型自身的自回归轨迹不断累积。为解决这一失配问题,我们引入了一个在策略蒸馏(on-policy distillation, OPD)阶段,将教师监督放置于量化模型实际生成之处。从 QAD 检查点出发,学生模型通过部署时使用的量化前向路径进行生成,并由一个冻结的全精度教师在其自身生成的前缀上提供反馈,将密集的词元级引导与任务验证器奖励相结合。在四个模型上、有效比特为2.79和1.88的设置下,OPD 将 MATH-500 上的平均 BF16 性能保持率从35%提升至70%,在 HumanEval 上从66%提升至91%,同时保持了短文本性能;在同等预算的比较中,其推理能力的提升大幅超过继续采用教师强制(teacher-forced)的 QAD。通过将 QAD 稳定的低比特初始化与 OPD 的在策略推理恢复相结合,我们的框架提供了一个全面的低于3比特的解决方案,在保持广泛能力的同时恢复了长文本推理能力。
cs.LG / 117 / 2609.26718

The Sirens' Song: When Proximal Background Context Overshadows Distant Evidence

塞壬之歌:当近端背景语境掩盖远端证据时
Yang, Xiaoyu, Lu, Jie, Duan, Wei, Yu, En
Abstract
Long-context LLMs focus on retrieving distant evidence from extensive context, yet existing work has largely focused on overcoming distance alone. In this work, we identify the Proximity Trap, insufficient attention to distant evidence often arises less from distance itself than from cumulative competition with abundant, task-irrelevant proximal background. To address the Proximity Trap, we introduce LYRA (Long-context heavY-tailed Relevance Alignment), a t-distributed directional matching mechanism that reshapes the context retrieval distribution, directing more attention mass toward task-relevant evidence, while preserving the relative positional information encoded. Extensive experiments on LongBench-v2, RULER, and LongBench demonstrate consistent improvements across context lengths and task categories. We further introduce ProxBench, a multi-level fine-grained benchmark for evaluating distant evidence utilization under increasing proximal background interference. Project page: https://xiaoyuyoung.github.io/LYRA/
Chinese Translation
长上下文大语言模型(LLM)重点关注从海量上下文中检索远端证据,然而现有工作大多仅聚焦于克服距离本身的问题。在本研究中,我们识别出“近因陷阱”(Proximity Trap):对远端证据的注意力不足,往往与其说源于距离本身,不如说源于与大量任务无关的近端背景信息的累积竞争。为解决近因陷阱问题,我们提出了 LYRA(Long-context heavY-tailed Relevance Alignment,长上下文重尾相关性对齐),这是一种 t 分布方向匹配机制,能够重塑上下文检索分布,将更多注意力质量引导至与任务相关的证据,同时保留已编码的相对位置信息。在 LongBench-v2、RULER 和 LongBench 上的大量实验表明,该方法在不同上下文长度和任务类别中均取得了一致的性能提升。我们进一步提出了 ProxBench,这是一个多层次细粒度基准,用于在近端背景干扰不断增强的情况下评估远端证据的利用能力。项目主页:https://xiaoyuyoung.github.io/LYRA/
cs.LG / 118 / 2609.26751

EquivSVA: A Formally Verified Dataset of Behavioral Assertions Across Equivalent RTL Implementations

EquivSVA:跨等价RTL实现的行为断言形式化验证数据集
Aditi, FNU
Abstract
Large language models are increasingly used to generate SystemVerilog Assertions from natural-language specifica- tions and register-transfer-level designs. Existing datasets and benchmarks support important goals such as large- scale training, formal evaluation, specification-to-assertion generation, and mutation-based testing. A complemen- tary need is to study whether a generated assertion cap- tures externally observable behavior or depends on inci- dental details of one RTL implementation. We present EquivSVA, a formally verified dataset organized around behavior families. Each family contains four structurally distinct RTL implementations of the same externally ob- servable behavior, shared interface-level gold properties, three controlled mutants, and formal-validation evidence. EquivSVA contains 120 behavior families across 12 cat- egories, 480 reference RTL implementations, 914 gold properties, and 360 mutants. Every final family passes a fixed 17-job validation suite covering RTL equivalence, gold-property proofs, property reachability, mutant dis- tinguishability, and gold-property checks on mutants. We also provide fixed family-safe train, development, and test splits. As a small demonstration of the analyses en- abled by the dataset, we evaluate the publicly released, Apache-2.0-licensed Qwen2.5-Coder-7B-Instruct model on the held-out test split. Of 293 interface-only generated properties, 93 are formally sound, and the number of sound properties varies across equivalent implementations for 14 of 24 test families. These results illustrate how behavior-family organization can support controlled stud- ies of assertion-generation robustness without requiring changes in intended functionality. The dataset, generators, validation scripts, and case-study artifacts are publicly released at https://github.com/aditigupta96/EquivSVA.
Chinese Translation
大语言模型越来越多地被用于从自然语言规格说明和寄存器传输级(RTL)设计生成SystemVerilog断言。现有数据集和基准支持大规模训练、形式化评估、规格到断言生成以及基于变异的测试等重要目标。一个互补的需求是研究生成的断言是否刻画了外部可观察的行为,还是依赖于某一RTL实现的偶然细节。我们提出了EquivSVA,一个围绕行为族组织的形式化验证数据集。每个行为族包含同一外部可观察行为的四个结构上不同的RTL实现、共享的接口级黄金属性(gold properties)、三个受控变异体以及形式化验证证据。EquivSVA包含12个类别中的120个行为族、480个参考RTL实现、914个黄金属性和360个变异体。每个最终的行为族都通过一个固定的17项任务验证套件,涵盖RTL等价性、黄金属性证明、属性可达性、变异体可区分性以及变异体上的黄金属性检查。我们还提供了固定的族安全(family-safe)训练集、开发集和测试集划分。作为该数据集所支持分析的一个小型演示,我们在留出的测试集划分上评估了公开发布的、采用Apache-2.0许可证的Qwen2.5-Coder-7B-Instruct模型。在293个仅接口生成的属性中,93个在形式上是可靠的,并且在24个测试行为族中有14个族,其可靠属性的数量在不同的等价实现之间存在差异。这些结果表明,行为族的组织方式可以在不改变预期功能的前提下,支持对断言生成鲁棒性的受控研究。数据集、生成器、验证脚本和案例研究工件已在https://github.com/aditigupta96/EquivSVA公开发布。