← Back to Index
Daily Research Digest

arXiv Papers

2026-09-25
424
Papers
4
Categories
423
Translated
收藏清单 0
机器人学 (Robotics)
87
cs.RO / 1 / 2609.28530

Know Your Body: A Harness for Direct and Self-Improving Robot Control with VLMs

了解你的身体:一个基于视觉语言模型(VLM)的直接且可自我改进的机器人控制框架
Lou, Zeyu, Zeng, Yanhong, Wang, Yong, Si, Chenyang
Abstract
A general-purpose vision-language model can understand a task goal without knowing how a particular robot's motion and functional parts produce the intended effect. We introduce KnowBody, a harness that makes these action-relevant body relations explicit, queryable, and revisable while keeping the model weights frozen. Initialized from one off-task trajectory, a partial body model guides action selection and the interpretation of past interactions. New evidence refines the model, and knowledge dependent on revised body estimates is rechecked before reuse. Across 32 fixed-budget trials on four real-robot tasks, initialized KnowBody achieves 75% completion versus 25% for the native harness and requires fewer planner rounds on successful trials in tasks completed by both. With persistent updates enabled, planner rounds decrease by 29-53% from the first to the fifth recorded success.
Chinese Translation
通用视觉语言模型能够理解任务目标,却不知道特定机器人的运动和功能部件如何产生预期效果。我们提出了 KnowBody,一个使这些与动作相关的身体关系变得显式、可查询且可修正的框架,同时保持模型权重冻结。KnowBody 仅从一条离线轨迹初始化,所得的部分身体模型用于指导动作选择以及对过去交互的解释。新证据会不断完善该模型,且依赖于已修正身体估计的知识在复用前会被重新检验。在四个真实机器人任务上进行的32次固定预算实验中,初始化后的 KnowBody 实现了75%的任务完成率,而原生框架仅为25%;在两者都能完成的任务中,KnowBody 在成功试验中所需的规划器轮数更少。启用持续更新后,从第一次到第五次记录的成功,规划器轮数减少了29%至53%。
cs.RO / 2 / 2609.28660

Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy

形态计量模仿:从形态与接触感知的手部重定向到Sim-to-Real视觉运动策略
Sadjadpour, Tara, He, Siming, Wolfe, C. K., Qi, Haozhi, Wilken, Lea, Sastry, S. Shankar, Tomlin, Claire, Malik, Jitendra
Abstract
Human hand-object interactions (HOIs) provide a rich source of demonstrations for dexterous manipulation, but learning directly from them presents challenges in bridging morphology gaps, ensuring dynamical feasibility, and sim-to-real deployment. We present Morphometric Imitation, a three-stage framework that transforms reconstructed HOIs into zero-shot sim-to-real visuomotor policies. First, morphometric optimization (MMO) kinematically retargets human motion across hand morphologies while preserving demonstrated contacts. Second, residual reinforcement learning (RL) refines the kinematic reference using object pose and contact information from the human motion to produce dynamically feasible robot demonstrations. Third, these demonstrations are distilled into visuomotor policies. Across three robot hands and ten HOIs, MMO improves contact F1 over the strongest of five baselines by at least 8 points for every hand, while also improving the success rate of downstream dynamic retargeting by as much as 35 points. Ablations on the residual RL show complementary benefits from using object pose and contact information. Finally, the visuomotor policies achieve 89.3% zero-shot success in 300 real-world trials on 30 objects. Project page: $\href{https://morphometricimitation.github.io}{\text{this https URL}}$
Chinese Translation
人手与物体的交互(HOIs)为灵巧操作提供了丰富的示范数据源,但直接从中学习面临着弥合形态差异、确保动力学可行性以及Sim-to-Real部署等挑战。我们提出了形态计量模仿(Morphometric Imitation),这是一个三阶段框架,能够将重建的HOI转化为零样本(zero-shot)Sim-to-Real视觉运动策略。首先,形态计量优化(MMO)在保留示教接触的同时,对不同手部形态之间的人体运动进行运动学重定向。其次,残差强化学习(RL)利用人体运动中的物体位姿和接触信息对运动学参考进行细化,从而生成动力学可行的机器人示范。第三,将这些示范蒸馏为视觉运动策略。在三种机器人手和十个HOI任务上,MMO的接触F1分数相比五个基线中最强者在每种手上至少提升8个百分点,同时将下游动态重定向的成功率最多提升35个百分点。对残差RL的消融实验表明,物体位姿信息和接触信息具有互补的作用。最后,视觉运动策略在30个物体的300次真实世界试验中实现了89.3%的零样本成功率。项目页面:$\href{https://morphometricimitation.github.io}{\text{this https URL}}$
cs.RO / 3 / 2609.28709

OA-MPPI: Occlusion-Aware Model Predictive Path Integral Control for UAV Flight

OA-MPPI:面向无人机飞行的遮挡感知模型预测路径积分控制
Palladino, Vittorio, Yang, Teaya, Zhang, Ruiqi, Mueller, Mark W.
Abstract
Autonomous UAV flight through cluttered and partially unknown environments requires reasoning not only about observed obstacles but also about occluded regions that the sensor cannot observe. We present OA-MPPI, an obstacle- and occlusion-aware extension of Model Predictive Path Integral (MPPI) control for quadrotor flight that accounts for potential moving agents emerging from these regions into the vehicle's path. At every planning step, we extract a 3D occlusion boundary from the online occupancy map and use it to model the regions that hidden agents could reach over the prediction horizon. We penalize trajectories that enter these expanding regions within MPPI rollouts generated using nonlinear quadrotor dynamics and accounting for individual rotor thrust limits. We validate the proposed approach in simulation and hardware flight experiments, with the complete pipeline running onboard the vehicle in real time. Results show increased clearance from occlusion boundaries compared to baseline MPPI in both settings, as well as avoidance of an agent emerging from occlusion in simulation.
Chinese Translation
无人机在杂乱且部分未知的环境中进行自主飞行时,不仅需要对观测到的障碍物进行推理,还需要对传感器无法观测的遮挡区域进行推理。我们提出了OA-MPPI,这是一种针对四旋翼飞行具有障碍物与遮挡感知能力的模型预测路径积分(MPPI)控制扩展方法,该方法考虑了潜在的运动主体从这些遮挡区域出现并进入飞行器路径的可能性。在每个规划步骤中,我们从在线占据地图中提取三维遮挡边界,并利用其对被遮挡主体在预测时域内可能到达的区域进行建模。在采用非线性四旋翼动力学并考虑各旋翼推力限制所生成的MPPI滚动仿真中,我们对进入这些扩展区域的轨迹施加惩罚。我们在仿真和硬件飞行实验中对所提出的方法进行了验证,完整流程可在机载设备上实时运行。结果表明,与基线MPPI相比,该方法在两种环境下均增大了与遮挡边界的间隙距离,并在仿真中成功避开了从遮挡区域出现的运动主体。
cs.RO / 4 / 2609.28716

Temporal Learning for End-Effector Position Estimation under Aerodynamic Disturbances in Aerial Continuum Manipulation

气动扰动下空中连续体操作的末端执行器位置估计的时间学习
Amiri, Niloufar, Masnavi, Houman, Janabi-Sharifi, Farrokh
Abstract
This paper investigates temporal neural networks for \mbox{end-effector} position \mbox{estimation} of an aerial continuum manipulator (ACM) operating under aerodynamic effects induced by the unmanned aerial vehicle (UAV). An experimental dataset is collected under stationary (\mbox{rotor-off}) and \mbox{free-hovering} conditions across continuum robot (CR) configurations and UAV altitudes, providing \mbox{end-effector} position measurements with and without aerodynamic residuals. To establish a nominal framework, \mbox{strain-parameterized} kinematic models with progressively richer strain bases are evaluated to balance model complexity and prediction accuracy. The selected nominal model then serves as the baseline for 3D position residual estimation using a \mbox{closed-form} \mbox{continuous-time} (CfC) neural network, with a multilayer perceptron (MLP) and a gated recurrent unit (GRU) used for comparison. On unseen test experiments, the CfC achieves an RMSE of \(22.00\pm1.70~\mathrm{mm}\) over five random seeds, compared with \(36.38\pm3.58~\mathrm{mm}\) for the MLP and \(27.72\pm2.92~\mathrm{mm}\) for the GRU, corresponding to reductions of \(39.52\%\) and \(20.62\%\), respectively. These results demonstrate the effectiveness of \mbox{continuous-time} learning for \mbox{end-effector} position estimation under aerodynamic disturbances relative to static and \mbox{discrete-time} learning methods.
Chinese Translation
本文研究了时序神经网络,用于在无人机(UAV)引起的气动效应下空中连续体机械臂(ACM)的末端执行器位置估计。实验数据集在静止(旋翼关闭)和自由悬停条件下,针对连续体机器人(CR)的不同构型及无人机不同高度进行采集,提供了有/无气动残差的末端执行器位置测量数据。为建立标称框架,评估了采用逐渐丰富的应变基的应变参数化运动学模型,以权衡模型复杂度与预测精度。所选的标称模型随后作为基线,用于采用闭式连续时间神经网络的估计,并与多层感知机(MLP)和门控循环单元(GRU)进行对比。在未见过的测试实验中,连续时间神经网络在五个随机种子上的RMSE为22.00±1.70 mm,而MLP为36.38±3.58 mm,GRU为27.72±2.92 mm,相比分别降低了39.52%和20.62%。这些结果表明,在气动扰动下进行末端执行器位置估计时,连续时间学习方法相较于静态和离散时间学习方法具有有效性。
cs.RO / 5 / 2609.28766

TAPESIM: Efficient Simulation of Adhesive Tape Dispensing for Robotic Manipulation

TAPESIM:面向机器人操作的高效胶带分配仿真
Luo, Zhaofeng, Lu, Xinyu, Choi, jaehoon, Chen, Zhehuan, Chung, Trinity, Qiu, Xiaowen, Perkins, Hugh Nicholas, Calderon, Gianna, Duburcq, Alexis, Son, Sanghyun, Wang, Tsun-Hsuan, Qiao, Yi-Ling, Li, Minchen
Abstract
Applying adhesive tape to secure wire harnesses or seal packages requires robots to coordinate a flexible strip, a moving roll, and surfaces that attach and detach. Simulation could make these interactions repeatable for robot development and evaluation, but resolving every adhesive layer is expensive and can suppress roll motion at practical solver tolerances, while a permanently rigid roll cannot release material. We present TapeSim, a tape simulator that concentrates deformation near the unwinding region and along the released strip. We will release the source code. A rigid cluster represents most wound material, while an advancing deformable collar enables payout and leaves released tape flexible and reattachable. Optional releasable bonds simplify adhesive interfaces and reduce mean step times for smaller rolls. Controlled swing tests show improved roll rotation. At 32 turns, clustering gives 3.2-3.4x mean physics-step speedups at a fixed Newton tolerance and 4.5-8.4x for comparable roll motion. Across five real-motion Stick replays, the clustered variants reduce mean image-plane core-landmark error by 23-29% relative to the full-shell cohesive baseline. On 100 paired Peel cases, they improve balanced accuracy from 50% to 72.9-76.3%, with interface rankings varying across tasks. A teleoperated box-sealing sequence demonstrates attachment, dispensing, cutting, and sealing in a continuous workflow.
Chinese Translation
使用胶带固定线束或封装包裹,需要机器人协调柔性带材、移动的卷筒以及可粘附和分离的表面。仿真可以使这些交互在机器人开发与评估中具备可重复性,但解析每一层粘合层的代价高昂,且在实际求解器容差下可能抑制卷筒运动,而永久刚性的卷筒又无法释放材料。我们提出了TapeSim,一种将变形集中在退绕区域和已释放带材附近的胶带仿真器。我们将发布源代码。该仿真器用刚性簇表示大部分缠绕材料,同时用一个不断推进的可变形环圈实现材料释放,并使释放后的胶带保持柔性且可重新粘附。可选的可分离键简化了粘合界面,并降低了较小卷筒的平均步长耗时。受控摆动测试显示卷筒旋转得到改善。在32圈的情况下,在固定牛顿容差下,聚簇方法使平均物理步长加速3.2-3.4倍,在卷筒运动相当的情况下加速4.5-8.4倍。在五次真实运动的Stick回放中,相较于全壳粘着基线,聚簇变体将图像平面核心关键点平均误差降低了23-29%。在100组配对剥离(Peel)案例中,聚簇变体将平衡准确率从50%提升至72.9-76.3%,界面排序随任务不同而有所差异。一个遥操作的封箱序列演示了粘附、分配、切割和密封的连续工作流程。
cs.RO / 6 / 2609.28767

Human-in-the-Loop Geospatial Annotation for Rapid Dataset Construction in Field-Deployed UAV Systems

面向野外部署无人机系统快速数据集构建的人在环路地理空间标注方法
Masters, Morgan, Bender, Nikolaas, Altaffer, T. Luca, Josephson, Colleen, McGuire, Steve
Abstract
Real-world perception systems must adapt to changing environments, but manual image annotation cannot scale to field data volumes. We present BirdsEye, which shifts expert annotation from images to the field: an operator records target locations in world coordinates using RTK positioning and calibrated projective geometry propagates each observation to all frames where the target is visible. To quantify how well physical annotations align with image observations, we derive a first-order mapping from camera-pose uncertainty to pixel uncertainty and validate it against Monte Carlo simulation. This mapping is linear in the six per-axis pose variances, so it inverts into a sensor design tool: we give a sufficient condition converting an annotation tolerance into a convex set of admissible pose-noise budgets, a closed-form largest admissible scaling of a deployed sensor suite, and a unique per-axis pose specification under an equal-budget-share allocation. We also analyze the planar-surface approximation underlying the projection, which holds up to 10 degrees of terrain slope. By direct measurement, we show that system projection accuracy is sub-decimeter (sub-30 pixel) at AGL altitudes of 10-20m under conditions excluding sustained yawing. During an in-field case study across three agricultural sites, two field workers produced 12,524 annotated frames carrying 55,600 labels in roughly 12 hours (25.5x per-worker rate increase over manual labeling). Detectors trained on imagery collected by this workflow recovered 56-89% of in-view surveyed targets at a geographically distinct farm, at pre-registered operating points; human review of the leading configuration estimates detection precision at 83-87%, spanning three tie-break conventions for clusters carrying contradictory human verdicts.
Chinese Translation
现实世界的感知系统必须适应不断变化的环境,但人工图像标注无法扩展到野外数据规模。我们提出了BirdsEye,它将专家标注从图像转移到野外:操作员使用RTK定位在全局世界坐标系中记录目标位置,并通过标定的投影几何将每次观测传播到目标可见的所有帧中。为量化物理标注与图像观测的吻合程度,我们推导了从相机位姿不确定性到像素不确定性的一阶映射,并通过蒙特卡洛模拟进行了验证。该映射对六个逐轴位姿方差是线性的,因此可逆推为一个传感器设计工具:我们给出了将标注容差转化为可容许位姿噪声预算凸集的充分条件、已部署传感器套件可容许最大缩放比例的闭式解,以及在等预算份额分配下唯一的逐轴位姿规格。我们还分析了投影所依赖的平面表面近似,该近似在地形坡度高达10度时仍然成立。通过直接测量,我们表明在排除持续偏航的条件下,系统在离地高度(AGL)10-20米时的投影精度达到亚分米级(30像素以内)。在一项覆盖三个农业站点的野外案例研究中,两名野外工作人员在约12小时内生成了12,524帧带有55,600个标签的标注图像(每位工作人员的标注速率较人工标注提升25.5倍)。使用该工作流程采集的图像训练的检测器,在预先注册的操作点下,于一个地理上不同的农场恢复了视野内勘测目标的56-89%;对领先配置的人工复核估计检测精度为83-87%,其中涵盖了针对含有人工判定矛盾的聚类所采用的三种平局裁决约定。
cs.RO / 7 / 2609.28798

OCC4M: Object-Centric 4D Memory for Spatiotemporal Reasoning in Long-Horizon Manipulation

OCC4M:面向长时程操作时空推理的以对象为中心的4D记忆
Jedlicki, Jack B., Dieudonné, Tanguy, Yang, Heng
Abstract
Long-horizon manipulation often requires reasoning about state absent from the current view, such as a vanished object's location, temporal identity, or the contents of a shuffled container. We present OCC4M ("Occam"), an object-centric 4D memory that maintains persistent tracks in a shared world frame and explicitly represents temporal, motion, and containment relations. A vision-language model (VLM) queries this structured memory to select actionable targets for history-free low-level execution. Across seven simulation conditions and 350 episodes, OCC4M achieves 96.6% memory success and 88.9% end-to-end success, versus 54.6% and 57.7% for FrameSamp, a raw-history VLM baseline using Gemini 3.7 Flash with the complete observation history and the same executor. In a controlled viewpoint-transfer test, OCC4M maintains 100% memory and 98% end-to-end success after a viewpoint change, while full-history FrameSamp falls to near-zero success. On 20 fixed-camera Franka episodes, OCC4M reaches 85% joint memory accuracy, versus at most 30% for FrameSamp across context sizes from $K=16$ to the complete history, and completes 45% of full two-stage tasks. These results support explicit object-centric memory for persistent spatiotemporal reasoning in long-horizon manipulation. Qualitative videos are available at https://occ4m-sup.github.io/occ4m-supplementary/.
Chinese Translation
长时程操作往往需要对当前视野之外的状态进行推理,例如已消失物体的位置、时间上的身份信息,或被调换的容器内的内容。我们提出了OCC4M("Occam"),一种以对象为中心的4D记忆方法,其在共享世界坐标系中维护持久化跟踪,并显式地表示时间、运动和容纳关系。视觉语言模型(VLM)查询这一结构化记忆,以选择可操作的目标,用于无历史依赖的底层执行。在七种仿真条件和350个回合中,OCC4M实现了96.6%的记忆成功率和88.9%的端到端成功率,而FrameSamp——一个使用Gemini 3.7 Flash、结合完整观测历史和相同执行器的原始历史VLM基线——仅达到54.6%和57.7%。在一项受控的视角迁移测试中,OCC4M在视角变化后仍保持100%的记忆成功率和98%的端到端成功率,而使用完整历史的FrameSamp的成功率降至接近于零。在20个固定相机的Franka回合中,OCC4M达到85%的联合记忆准确率,而FrameSamp在上下文规模从K=16到完整历史的所有设置下最高仅为30%;OCC4M还完成了45%的完整两阶段任务。这些结果支持在长时程操作中采用显式的以对象为中心的记忆来实现持久的时空推理。定性演示视频见 https://occ4m-sup.github.io/occ4m-supplementary/。
cs.RO / 8 / 2609.28807

An Analysis of Streaming Deep Reinforcement Learning for Adaptive Continual Learning in Robotics

面向机器人自适应持续学习的流式深度强化学习分析
Vitchutripop, Teeratham, Quarles, Alyssa, Zhang, Wenhe, Xue, Richard, Rakita, Daniel
Abstract
Over the course of a lifetime, robots may encounter novel scenarios unaccounted for in its original training that result in performance degradation. One common approach to mitigating this issue is to further grow the offline training dataset in hopes of producing a policy robust to these changes. In contrast, biological learning occurs moment-to-moment via a stream of experience, unlike the predominantly batch-based and offline nature of deep learning. Although recent works show the feasibility of stream-based deep reinforcement learning, where updates use only the latest experience, none have shown it to be a viable continual learning framework for adapting robotic policies to unseen changes. In this paper, we present the first analysis of streaming deep reinforcement learning for adaptive continual learning in robotics. In particular, we show that, following an initial pretraining phase, streaming deep RL can enable a robot to successfully adapt to unforeseen changes to itself, its environment, or goals. Our primary experiments within quadruped locomotion demonstrate that a deep neural network robotic policy with certain optimizers and plasticity loss mitigation techniques can successfully leverage domain task knowledge from its pretraining to quickly adapt online to diverse changes via stream learning, outperforming batch-based on-policy methods and improving task success rates by up to 90% over the pretrained policy. Furthermore, we perform additional evaluations on robotic manipulation tasks to determine if our previous observations extend to different robotic morphologies and scenarios. Our results show that the successes observed in quadruped locomotion can be partially realized in manipulation with stability and performance limitations. We conclude with a discussion on the limitations of our work and its implications for the future of continual robot learning.
Chinese Translation
在生命周期中,机器人可能遇到其初始训练中未涵盖的新场景,从而导致性能下降。缓解这一问题的一种常见方法是不断扩充离线训练数据集,以期训练出对这些变化具有鲁棒性的策略。相比之下,生物学习是通过经验流随时随刻进行的,这与深度学习以批处理和离线为主的特性截然不同。尽管近期研究表明基于流的深度强化学习(stream-based deep reinforcement learning)是可行的——即仅使用最新经验进行更新——但尚无研究证明它可以作为将机器人策略适应于未见变化的可行持续学习框架。本文首次对用于机器人自适应持续学习的流式深度强化学习进行了分析。特别地,我们证明,在初始预训练阶段之后,流式深度强化学习能够使机器人成功适应其自身、环境或目标中的未知变化。我们在四足运动(quadruped locomotion)上的主要实验表明,配备某些优化器和可塑性损失(plasticity loss)缓解技术的深度神经网络机器人策略,能够成功利用预训练获得的领域任务知识,通过流式学习在线快速适应多样化的变化,其性能优于基于批处理的同策略(on-policy)方法,任务成功率相比预训练策略最高提升90%。此外,我们在机器人操作(manipulation)任务上进行了额外评估,以检验先前的观察结果是否适用于不同的机器人形态和场景。结果表明,四足运动中所取得的成就可以在操作任务中部分实现,但仍存在稳定性和性能方面的局限。最后,我们讨论了本工作的局限性及其对机器人持续学习未来发展的意义。
cs.RO / 9 / 2609.28816

FlyCNS: Connectome-Grounded Information Organization for Communication-Constrained Embodied Control

FlyCNS:面向通信受限具身控制的基于连接组的信息组织方法
Zhang, Jinchang, Lin, Jiakai, Lu, Guoyu
Abstract
Robotic bodies are inherently distributed in sensing and actuation, yet learning-based control still commonly relies on centralized information processing. This work studies the problem of information organization in communication-constrained embodied control: which computations should remain local, and which information is worth transmitting for whole-body coordination. We propose FlyCNS, an embodied information-organization framework inspired by the Drosophila brain--nerve-cord connectome. FlyCNS preserves local sensorimotor computation within each limb and enables selective long-range communication through separate ascending and descending routing pathways. From a real connectome, FlyCNS extracts the directional structural complexity of these two pathway types and uses it as a weak prior over communication allocation, while message content, transmission timing, and locomotion policies remain task-adaptive and are learned through reinforcement learning. In Unitree Go1 simulation, FlyCNS exhibits more graceful performance degradation as the communication budget is tightened. Under the most restrictive setting, it uses only about 21--22\% of the communication of the full-communication reference, while still maintaining a tracking score of approximately 0.882 under both command protocols, with a gap of no more than 6.1\% from the full-communication reference. These results indicate that real neural connectomes can inform not only the structural design of control networks, but also provide transferable inductive biases for information organization across embodiments, guiding robots in balancing local computation and long-range coordination under limited communication resources.
Chinese Translation
机器人的感知与执行本质上分布在身体各处,然而基于学习的控制通常仍依赖于集中式的信息处理。本文研究了通信受限具身控制中的信息组织问题:哪些计算应当保留在本地进行,哪些信息值得传输以实现全身协调。我们提出了FlyCNS,一个受果蝇(Drosophila)大脑——神经索连接组启发的具身信息组织框架。FlyCNS将局部感觉运动计算保留在各肢体内部,并通过分离的上行与下行路由通路实现选择性的长距离通信。FlyCNS从真实连接组中提取这两类通路的方向性结构复杂度,并将其作为通信分配的弱先验,而消息内容、传输时机和运动策略则保持任务自适应,并通过强化学习进行学习。在Unitree Go1仿真实验中,随着通信预算收紧,FlyCNS表现出更为平缓的性能下降。在最受限的设置下,其通信量仅约为全通信参考方案的21–22%,同时在两种指令协议下仍保持约0.882的跟踪分数,与全通信参考方案的差距不超过6.1%。这些结果表明,真实的神经连接组不仅能够指导控制网络的结构设计,还能为跨具身形态的信息组织提供可迁移的归纳偏置,引导机器人在有限通信资源下平衡本地计算与长距离协调。
cs.RO / 10 / 2609.28818

KeyGen: Unsupervised Keypoint based Object-Centric Representations for Category-Level Policy Generalization

KeyGen:基于无监督关键点的以对象为中心表征,用于类别级策略泛化
Cao, Shuxin, Wang, Liquan, Moghani, Masoud, Joffe, Benjamin, Garg, Animesh
Abstract
Generalization in robotic manipulation requires policies to perform tasks across diverse unseen object instances that vary in shape, size, and pose. However, conventional behavior cloning (BC) methods often overfit to instance-specific geometry and appearance, limiting transfer to novel objects. We introduce KeyGen, a framework that learns canonicalized semantic 3D keypoints from point clouds and uses them as structured object-centric representations for policy learning. A visuomotor diffusion policy conditions on these keypoints together with object-centric geometry to predict full manipulation trajectories, enabling consistent geometric correspondence across object instances. To evaluate category-level generalization, we construct a photorealistic simulation benchmark with three manipulation tasks and a planning-driven data generation pipeline that produces expert trajectories across diverse object instances. Experiments show that KeyGen significantly outperforms prior methods on both seen and unseen objects under pose variation, scales effectively with additional demonstrations per object, maintains robustness to object rescaling, and achieves strong performance in both simulation and real-world manipulation.
Chinese Translation
机器人操作中的泛化能力要求策略能够在形状、大小和姿态各异的各种未见物体实例上执行任务。然而,传统行为克隆(Behavior Cloning, BC)方法往往过度拟合特定实例的几何形状和外观,限制了对新物体的迁移能力。我们提出了KeyGen,一个从点云中学习规范化语义3D关键点并将其作为结构化以对象为中心表征用于策略学习的框架。视觉运动扩散策略(visuomotor diffusion policy)以这些关键点以及以对象为中心的几何信息为条件来预测完整的操作轨迹,从而在物体实例之间实现一致的几何对应关系。为评估类别级泛化能力,我们构建了一个包含三个操作任务的逼真仿真基准,以及一个基于规划的数据生成流水线,可生成覆盖多样物体实例的专家轨迹。实验表明,KeyGen在姿态变化下于已见和未见物体上均显著优于先前方法,能随每个物体的额外演示有效扩展,对物体缩放保持鲁棒性,并在仿真和真实世界操作中均取得了优异表现。
cs.RO / 11 / 2609.28838

Uncertainty-Gated Exploration Noise Suppresses Task Collapse in Online RL Fine-Tuning of a Flow-Matching Vision-Language-Action Policy

不确定性门控探索噪声抑制流匹配视觉-语言-动作策略在线强化学习微调中的任务坍塌
Yardımcı, Mehmet Turan, Çoğurcu, Yunus Emre
Abstract
Online reinforcement learning fine-tuning of pretrained flow-matching vision-language-action (VLA) policies promises robots that keep learning after deployment, but continued updates often destroy competence on individual tasks while the aggregate still looks healthy. We study this failure mode, which we call task collapse, under a matched small-compute budget on LIBERO-10 with a 450M-parameter SmolVLA policy trained by PPO with stochastic (SDE) sampling. Three exploration-noise policies differ in one live variable: a fixed noise scale, a ReinFlow-style learned noise network, and an uncertainty-gated controller that redistributes exploration across task streams from task-agnostic novelty and competence signals, without task labels or episode boundaries. Under the pooled definition, fixed noise collapses tasks in two of three seeds and learned noise in every seed measured to iteration 200, while the controller collapses none in any of its three seeds. Measured parameter displacement shows the controller's action expert keeps changing, while its mean applied noise is close to the fixed scale in the available logs. The matched comparison supports the controller's effect on task preservation; the separate contributions of its adaptation across states and over time are not disentangled. A lower fixed scale slows the decline but does not stop it. No arm improves on the behavior-cloning baseline in this budget. Two properties of that regime are measured beside this result, not offered as its cause: following the reference recipe, training runs in bfloat16 with no fp32 master copy, under which 96.02% of the action expert's elements stay bit-identical across three consecutive iterations, and an fp32 master copy at the reference learning rate collapses both arms in a single-seed observation. We release tools measuring per-task collapse under four definitions, rescoring noise and instrument tares.
Chinese Translation
对预训练的流匹配视觉-语言-动作(VLA)策略进行在线强化学习微调,有望使机器人在部署后持续学习,但持续的更新往往会在总体表现看似正常的情况下破坏模型在单个任务上的能力。我们在LIBERO-10上、在匹配的小算力预算下研究了这种失败模式——我们称之为任务坍塌,采用450M参数的SmolVLA策略,并通过带随机(SDE)采样的PPO进行训练。三种探索噪声策略仅在一个实时变量上有所不同:固定噪声尺度、ReinFlow风格的可学习噪声网络,以及一种不确定性门控控制器——后者基于任务无关的新颖性和能力信号在任务流之间重新分配探索,且无需任务标签或回合边界。在汇总定义下,固定噪声在三个随机种子中的两个出现任务坍塌,可学习噪声在测量至200次迭代的每个种子中均出现坍塌,而该控制器在其三个种子中均未发生坍塌。参数位移的测量显示,控制器的动作专家仍在持续变化,而其在可用日志中的平均施加噪声接近固定尺度。这一匹配比较支持了控制器在任务保持方面的效果;但其跨状态适应与跨时间适应的各自贡献尚未被解耦。更低的固定噪声尺度能减缓衰退但无法阻止它。在此预算下,没有任何方案能超越行为克隆基线。作为该实验设定的两个附带测量性质(而非其成因):按照参考配方,训练以bfloat16运行且无fp32主副本,在此条件下动作专家96.02%的元素在连续三次迭代中保持比特级一致;而在参考学习率下使用fp32主副本时,单种子观测显示两个方案均发生坍塌。我们发布了相关工具,用于在四种定义下测量每任务坍塌,并对噪声和仪器零点漂移进行重评分。
cs.RO / 12 / 2609.28878

Online Sim-to-Real Adaptation via Closed-Loop System Modeling

基于闭环系统建模的在线仿真到现实自适应
Huang, Yuhao, Moore, Samuel A., Chen, Boyuan
Abstract
Sim-to-real transfer has made substantial progress, but can still produce controllers that remain stable and functional on hardware while suffering from degraded tracking accuracy due to residual dynamics mismatch. Correcting these errors typically requires identifying the underlying system dynamics, adapting the control policy, or returning to simulation for additional training and finetuning, all of which can require substantial data and computation. We propose OSRAM (Online Sim-to-Real Adaptation via Closed-Loop System Modeling), a framework that instead adapts the reference commands provided to an existing controller. OSRAM treats the deployed robot and its policy as a unified closed-loop dynamical system and learns its task-level command-response behavior directly from tracking observations. A closed-loop dynamics model is meta-trained across randomized dynamics in simulation and rapidly finetuned after deployment using limited real-world interaction. The adapted model is then used to optimize future reference commands while leaving the underlying control policy unchanged. We evaluate OSRAM on bipedal velocity tracking and loco-manipulation in simulation and on hardware. Results show that closed-loop modeling improves prediction and tracking accuracy under unseen dynamics, while online reference adaptation reduces residual sim-to-real tracking errors across different control objectives and hardware configurations. These results demonstrate that adapting the behavior of the robot-policy closed loop provides a practical alternative to finetuning the policy or identifying the full physical dynamics for sim-to-real transfer. More information can be found at http://generalroboticslab.com/OSRAM.
Chinese Translation
仿真到现实的迁移(sim-to-real transfer)已取得长足进展,但所得到的控制器在真实硬件上虽能保持稳定并正常工作,其跟踪精度却可能因残余动力学失配而下降。纠正这些误差通常需要辨识底层系统动力学、自适应调整控制策略,或返回仿真环境进行额外的训练与微调,而这些方法往往需要大量的数据与计算资源。我们提出OSRAM(Online Sim-to-Real Adaptation via Closed-Loop System Modeling,基于闭环系统建模的在线仿真到现实自适应)框架,该框架转而对提供给现有控制器的参考指令进行自适应调整。OSRAM将部署后的机器人及其策略视为一个统一的闭环动力学系统,并直接从跟踪观测数据中学习其任务级的指令—响应行为。该闭环动力学模型在仿真中通过随机化动力学进行元训练,并在部署后利用有限的现实交互数据快速微调。随后,利用自适应后的模型优化未来的参考指令,而保持底层控制策略不变。我们在双足速度跟踪和移动操作(loco-manipulation)任务上于仿真和真实硬件环境中对OSRAM进行了评估。结果表明,闭环建模在未见过的动力学条件下提升了预测和跟踪精度,同时在线参考指令自适应在不同控制目标和硬件配置下均有效降低了残余的仿真到现实跟踪误差。这些结果表明,对机器人—策略闭环行为进行自适应调整,为仿真到现实迁移提供了一种替代策略微调或完整物理动力学辨识的实用方案。更多信息请访问 http://generalroboticslab.com/OSRAM。
cs.RO / 13 / 2609.28887

Fly, Drive, Reconfigure: A Modular Reconfigurable Aerial-Ground Platform for Field Operations

飞行、行驶、重构:面向野外作业的模块化可重构空地平台
Lo, Li-Yu, Liu, Yanbaihui, Shu, Chengchuan, Harris, Tyler, Ryan, Jonathan, Chen, Boyuan
Abstract
Heterogeneous robot teams distribute complementary capabilities across specialized agents, but their physical roles and capacities typically remain fixed throughout a mission. We present HARP, a Heterogeneous Aerial Robotic modules Platform in which independently deployable aerial robots physically reconfigure to compose their capabilities for field operations. HARP comprises sensor-equipped scouts, flydrive rover modules, and task-specific payload modules. Scouts map the environment and inform an energy-aware planner that jointly selects routes and air-ground mobility modes. Rover and payload modules fly independently across terrain that constrains ground travel, then autonomously assemble into a cooperative ground vehicle for energy-efficient payload transport. Motivated by environmental sampling in remote and difficult-to-traverse regions, we evaluate HARP through field experiments spanning sensing, planning, reconfiguration, airground mobility, payload transport, and task execution. We further conduct module-level deployment tests on the Greenland Ice Sheet toward future autonomous missions. HARP demonstrates how heterogeneous robot teams can adapt not only their actions, but also how their physical capabilities are composed during a mission.
Chinese Translation
异构机器人团队将互补能力分布于各专用智能体,但其物理角色与能力通常在整个任务过程中保持固定。我们提出了HARP(异构空中机器人模块平台,Heterogeneous Aerial Robotic modules Platform),其中可独立部署的空中机器人通过物理重构来组合各自能力,以执行野外作业任务。HARP由配备传感器的侦察模块、飞行-行驶 rover 模块以及特定任务的载荷模块组成。侦察模块对环境进行建图,并为一个能量感知规划器提供信息,该规划器联合选择路径与空地移动模式。Rover 模块和载荷模块可以在限制地面通行的地形上独立飞行,随后自主组装为一辆协作式地面车辆,以实现高能效的载荷运输。受偏远且难以通行区域环境采样任务的启发,我们通过涵盖感知、规划、重构、空地移动、载荷运输和任务执行的实地实验对HARP进行了评估。我们还在格陵兰冰盖上开展了面向未来自主任务的模块级部署测试。HARP展示了异构机器人团队如何在任务过程中不仅调整自身行为,还能重新组合其物理能力。
cs.RO / 14 / 2609.28910

Robots That Take Initiative: A Framework for Building and Evaluating Proactive Robots

具有主动性的机器人:一个用于构建和评估主动式机器人的框架
Patel, Maithili, Chernova, Sonia
Abstract
Effective robot assistance beyond narrow roles and repetitive tasks requires robots to be proactive - to decide what needs to be done rather than waiting to be told. While proactivity is increasingly explored, it lacks a unified formulation, and work in the domain is typically evaluated offline against static human models that cannot capture the effect of a robot's actions on the environment and the user's own behavior. We introduce a unified formalism for proactive robot assistance, organize it into three levels, and provide a framework to address the highest level of unprompted proactive assistance. We then show that offline evaluation overstates performance in this setting, and contribute a closed-loop evaluation with a human model that adapts to the robot. Finally, we present a method, GAP, that instantiates our framework, learning from passive observation to anticipate user goals and act. Under closed-loop evaluation, prior state-of-the-art methods collapse, in some cases adding more work than they save, while GAP remains robust and substantially outperforms them.
Chinese Translation
要让机器人在狭窄的角色和重复性任务之外提供有效的辅助,就需要机器人具备主动性——即自主决定需要做什么,而不是等待指令。尽管主动性正受到越来越多的研究关注,但该领域缺乏统一的表述形式,且相关工作通常基于静态的人类模型进行离线评估,这种评估方式无法捕捉机器人行为对环境以及用户自身行为的影响。我们提出了一种用于主动式机器人辅助的统一形式化框架,将其划分为三个层次,并针对其中最高层次——无提示的主动辅助——提供了一个框架。我们随后证明,在此情境下离线评估会高估性能,并贡献了一种采用能够适应机器人的人类模型的闭环评估方法。最后,我们提出了一种名为 GAP 的方法来实例化我们的框架,该方法通过被动观察学习来预测用户目标并采取行动。在闭环评估下,先前的最先进方法表现崩溃,某些情况下它们增加的工作量甚至超过节省的工作量,而 GAP 依然保持稳健,并显著优于这些方法。
cs.RO / 15 / 2609.28920

Koopman-Accelerated Model-Based Diffusion for Real-Time Robot Control

基于Koopman加速的模型扩散方法用于实时机器人控制
Pak, Bohyeong, Lee, Kangmin, Kim, Sanghyun
Abstract
Conventional model-based diffusion (MBD) achieves effective trajectory optimization by leveraging noise annealing. However, its high computational cost, primarily arising from repeated rollouts of the plant dynamics, has largely confined its use to offline settings. To address this limitation, this paper proposes bilinear Koopman model-based diffusion (BK-MBD). The proposed method lifts the robot's state into a high-dimensional space only once per control step and propagates all candidates in the lifted space thereafter, so each rollout reduces to a fixed number of matrix-vector multiplications. The lifted dynamics are bilinear, allowing the predicted input gain to vary with the robot's configuration, which a linear lifted model cannot represent. In simulation, BK-MBD completed each planning update in at most 14.7 ms within a 50 ms control period and reached the goal on every trial, whereas a linear lift almost never did. The annealed schedule improves closed-loop accuracy over fixed-noise schedules under the learned rollout. Under the exact rollout, both the annealed and fixed-narrow schedules reach every goal, indicating that annealing reduces sensitivity to surrogate-model error. BK-MBD also threaded a passage that no single convex region covers, whereas a convexified bilinear controller rarely succeeded. On a physical manipulator, BK-MBD tracked an initially unknown moving target within the control period and was the only method that met both the tracking task and the deadline. The project page is available at https://rcilab.khu.ac.kr/bkmbd/.
Chinese Translation
传统的基于模型的扩散方法(Model-Based Diffusion, MBD)通过利用噪声退火实现有效的轨迹优化。然而,其高昂的计算成本(主要源于对被控对象动力学的反复前向推演)在很大程度上限制了其仅能应用于离线场景。为解决这一局限,本文提出了基于双线性Koopman模型的扩散方法(Bilinear Koopman Model-Based Diffusion, BK-MBD)。该方法在每个控制步中仅将机器人状态提升(lift)到高维空间一次,此后所有候选解均在提升空间中传播,因此每次前向推演仅需固定次数的矩阵-向量乘法运算。提升后的动力学呈双线性形式,使得预测输入增益可随机器人的构型而变化,这是线性提升模型无法表达的。在仿真中,BK-MBD在50毫秒的控制周期内每次规划更新最多耗时14.7毫秒,且在所有试验中均成功到达目标,而线性提升模型几乎从未成功。在学习到的前向推演模型下,退火调度相比固定噪声调度提高了闭环精度。在精确前向推演下,退火调度和窄范围固定噪声调度均能到达所有目标,这表明退火降低了模型对代理模型误差的敏感性。BK-MBD还能穿越没有任何单一凸区域所能覆盖的通道,而凸化的双线性控制器很少成功。在实物机械臂上,BK-MBD在控制周期内跟踪了初始未知的运动目标,并且是唯一同时满足跟踪任务与时间期限要求的方法。项目页面见 https://rcilab.khu.ac.kr/bkmbd/。
cs.RO / 16 / 2609.28927

Streaming-WAM: Action-Conditioned World-Action Model for Asynchronous Robot Manipulation

Streaming-WAM:用于异步机器人操作的动作条件化世界-动作模型
Huang, Xuyao, Wang, Yixuan, Ye, Zengyao, Zhao, Boyuan, Yu, Chenyang, Wen, Haoran, Deng, Zhijie
Abstract
World action models (WAMs) that use future visual prediction at inference time incur substantial generation costs. Asynchronous execution reduces waiting by overlapping inference with robot motion, but visual predictions used for subsequent action generation must anticipate the effects of actions already scheduled for execution during inference. We introduce Streaming-WAM, which couples action-conditioned world modeling with asynchronous robot control to account for committed actions in future visual prediction. At each streaming update, the model conditions future visual prediction on the latest observation and the committed actions, which form the fixed prefix of the next action chunk. The resulting action-conditioned visual features guide generation of the remaining actions within the same joint update, so the continuation is informed by the scene changes expected during execution of the fixed prefix. On LIBERO, Streaming-WAM achieves an average success rate of 98.35\% and reduces mean episode time by a factor of 2.93 relative to Fast-WAM. On the real-world Stamp Paper task, mean episode time falls from 90 s with synchronous Joint-WAM to 38 s with Streaming-WAM. These results show that Streaming-WAM supports efficient asynchronous control while maintaining high task success rates.
Chinese Translation
在推理时利用未来视觉预测的世界动作模型(WAM)会产生大量的生成开销。异步执行通过将推理与机器人运动重叠来减少等待时间,但用于后续动作生成的视觉预测必须预判推理期间已调度执行的动作所产生的效果。我们提出Streaming-WAM,它将动作条件化的世界建模与异步机器人控制相结合,以在预测未来视觉状态时考虑已提交的动作。在每次流式更新中,模型基于最新的观测和已提交的动作进行未来视觉预测,这些已提交的动作构成下一个动作块的固定前缀。由此得到的动作条件化视觉特征引导在同一联合更新中生成剩余动作,使得动作的延续部分能够参考固定前缀执行期间预期的场景变化。在LIBERO基准上,Streaming-WAM达到98.35%的平均成功率,并将相对于Fast-WAM的平均回合时间缩短了2.93倍。在真实世界的贴邮票(Stamp Paper)任务上,平均回合时间从同步Joint-WAM的90秒降至Streaming-WAM的38秒。这些结果表明,Streaming-WAM在保持高任务成功率的同时,支持高效的异步控制。
cs.RO / 17 / 2609.28933

A Field-Deployable GNSS-based Navigation Stack for Outdoor Mobile Robots

一种可野外部署的基于GNSS的室外移动机器人导航系统
Lin, Yiyuan, Regnier, Cole, Jiang, Yu
Abstract
Outdoor robots require more than an accurate receiver and a path-tracking law: the navigation system must preserve geometric consistency from geographic waypoints to actuator commands, expose measurement validity and timing, and respond to invalid or stale state information. This work presents a ROS~2 navigation stack with interchangeable single-GNSS--IMU and dual-antenna-GNSS localization front ends. Both provide a common local East--North--Up state interface for pure pursuit, virtual-point cross-track PID, finite-horizon nonlinear model predictive control (NMPC), and a segment-dependent hybrid dispatcher. The architecture specifies coordinate conventions, datum initialization, asynchronous state construction, waypoint geometry, controller equations, quality gates, command arbitration, and watchdog behavior. Independent physical field runs collected during 2025 and 2026 grape-vineyard deployments support a balanced evaluation of 800 runs, with 100 runs for each of eight controller--localization combinations on an approximately 199.6-m route. The row-hybrid mode yields the lowest run-averaged post-acquisition mean absolute cross-track error (MAE) in the evaluated dataset: 0.00952~m with single GNSS+IMU and 0.00846~m with dual GNSS. These findings characterize deviations of the recorded positions from the reference route under the evaluated conditions. The open-source navigation software and deployment instructions are available in the https://github.com/YiyuanLinXX/PPBv2/tree/main/PPBv2_Navigation.
Chinese Translation
室外机器人所需的不仅仅是一个精确的接收机和一条路径跟踪控制律:导航系统必须在从地理航路点到执行器指令的整个链路中保持几何一致性,暴露测量有效性与时间信息,并对无效或过期的状态信息作出响应。本工作提出了一种基于ROS 2的导航系统,具有可互换的单GNSS-IMU与双天线GNSS定位前端。两者均提供统一的本地东北天(East-North-Up)状态接口,支持纯跟踪(pure pursuit)、虚拟点横向偏差PID、有限时域非线性模型预测控制(NMPC)以及基于路段的混合调度器。该架构规定了坐标约定、基准初始化、异步状态构建、航路点几何、控制器方程、质量门限、指令仲裁及看门狗行为。在2025年和2026年葡萄园部署期间采集的独立实地运行数据支持了对800次运行的均衡评估,即在约199.6米的路径上,对八种控制器-定位组合各进行100次运行。在所评估的数据集中,行间混合模式取得了最低的采集后平均绝对横向偏差(MAE)运行均值:单GNSS+IMU为0.00952米,双GNSS为0.00846米。这些结果刻画了在所评估条件下记录位置相对于参考路径的偏差。开源导航软件及部署说明可在 https://github.com/YiyuanLinXX/PPBv2/tree/main/PPBv2_Navigation 获取。
cs.RO / 18 / 2609.28952

RoboRecover: Benchmarking Robot Policy Recovery under Execution Deviations

RoboRecover:面向执行偏差的机器人策略恢复基准测试
Li, Yang, Zhao, Chen, Wang, Zhuoran, Wang, Jiankang, Shao, Chao, Lin, Yihan, Shen, Haitao, Zhang, Jing
Abstract
Robot-policy benchmarks increasingly cover diverse tasks and preset out-of-distribution conditions, but typically evaluate complete trajectories from predefined initial states. These evaluations often focus on the initialized scene and the final outcome, while paying less attention to the dynamic interaction process. During closed-loop execution, actions and contacts can alter object relations and task progress, producing off-nominal intermediate states that need recovery. Recovery requires a policy to infer how task progress has changed, correct the relevant relations, and continue the original goal. We introduce RoboRecover, a benchmark for robot policy recovery under execution deviations. RoboRecover selects deviation states from trajectories, reconstructs them by replaying action prefixes, and evaluates policies on the original task. RoboRecover contains 2,000 scenarios across RoboTwin and LIBERO, with 1,000 scenarios and a fixed 800/200 train/test split on each platform. Results show that initial-state performance does not determine recovery performance and policies exhibit different recovery strengths across scenarios. Using its training split, RoboRecover further supports study on recovery interventions. RoboRecover establishes recovery from execution-induced intermediate states as a distinct dimension of robot policy evaluation.
Chinese Translation
机器人策略基准测试日益覆盖多样化任务和预设的分布外条件,但通常是从预定义的初始状态评估完整轨迹。这些评估往往关注初始化场景和最终结果,而对动态交互过程关注较少。在闭环执行过程中,动作和接触会改变物体之间的关系和任务进度,产生需要恢复的非正常中间状态。恢复要求策略推断任务进度发生了何种变化,纠正相关关系,并继续完成原定目标。我们提出了RoboRecover,一个针对执行偏差下机器人策略恢复的基准测试。RoboRecover从轨迹中选取偏差状态,通过重放动作前缀来重构这些状态,并在原任务上评估策略。RoboRecover在RoboTwin和LIBERO平台上共包含2,000个场景,每个平台有1,000个场景,并采用固定的800/200训练/测试划分。结果表明,初始状态下的性能并不能决定恢复性能,且策略在不同场景下表现出不同的恢复能力。利用其训练集划分,RoboRecover进一步支持了恢复干预方面的研究。RoboRecover将从执行引起的中间状态中恢复确立为机器人策略评估的一个独立维度。
cs.RO / 19 / 2609.28955

ActGaze: Learning Action-Grounded Gaze through Counterfactual Visual Interventions for High-Precision Manipulation

ActGaze:通过反事实视觉干预学习基于动作的注视以实现高精度操作
Zhu, Jinxuan, Wang, Jiaheng, Tang, Chao, Wang, Mengfan, Wei, Hao, Li, Shengbao, Yin, Hong, Gao, Yiwen, Tie, Chenrui, Li, Tingguang
Abstract
Current Vision-Language-Action (VLA) models often struggle with high-precision robotic manipulation. We attribute this limitation primarily to their visual attention being dispersed across task-irrelevant regions. To address this issue, we propose ActGaze, a training approach that guides VLA policies to gaze on task-relevant regions, much like humans gaze on critical visual cues while executing precise movements. Unlike prior methods that rely on external labels for gaze supervision, ActGaze derives spatial supervision directly from the VLA's own action objective by using counterfactual visual interventions to identify regions that are critical for action prediction. Extensive real-robot experiments on four high-precision robotic manipulation tasks demonstrate that ActGaze induces more focused visual attention on task-relevant regions and consistently outperforms the base VLA policy and other visual-grounding approaches.
Chinese Translation
当前的视觉-语言-动作(VLA)模型在高精度机器人操作任务中常常表现不佳。我们将这一局限性主要归因于其视觉注意力分散在与任务无关的区域。为解决这一问题,我们提出 ActGaze,一种训练方法,用于引导 VLA 策略注视与任务相关的区域,就像人类在执行精细动作时会注视关键视觉线索一样。与以往依赖外部标签进行注视监督的方法不同,ActGaze 利用反事实视觉干预识别对动作预测至关重要的区域,从而直接从 VLA 自身的动作目标中获取空间监督信号。在四项高精度机器人操作任务上的大量真实机器人实验表明,ActGaze 能够引导模型在任务相关区域上形成更加集中的视觉注意力,并始终优于基础 VLA 策略以及其他视觉定位方法。
cs.RO / 20 / 2609.28959

TactileStep: Sole Tactile Learning for Regulating Foot-Terrain Interaction in Humanoid Locomotion

TactileStep:用于调节人形机器人运动中足-地形交互的足底触觉学习
Wang, Zizhuo, Lee, Ming-ju, Zhu, Shaoting, Lou, Haozhe, Zhao, Hang, Li, Yiming
Abstract
Humanoid parkour policies can traverse various terrains, but task completion may mask challenges of harsh landings, edge contacts, and unstable stance contacts. Humans naturally regulate foot-terrain interaction through tactile feedback, modulating contact compliance according to terrain stiffness. This highlights a key domain gap between humans and humanoid robots: the absence of rich tactile sensing in most humanoid systems. We address this problem with TactileStep, a deployable tactile learning framework that brings sole pressure sensing into humanoid locomotion control for softer touchdowns and more stable support. TactileStep aligns tactile simulation with the real pressure insole, allowing the policy to learn from the same contact features available on hardware. During training, we use tactile and motion cues to recognize different foot-contact phases and apply phase-aware rewards that encourage safer landing and more stable stance. Evaluated in simulation and on a Unitree G1 humanoid across diverse terrains, TactileStep reduces peak touchdown force by up to 48.8% and peak A-weighted impact noise by up to 30.1 dB over a strong perceptive baseline, while increasing stance contact area by up to 23.8%.
Chinese Translation
人形机器人跑酷策略能够穿越各种地形,但任务的完成可能掩盖了严酷着陆、边缘接触和不稳定支撑接触等挑战。人类天然地通过触觉反馈来调节足-地形交互,根据地形刚度调整接触柔顺性。这凸显了人类与人形机器人之间的一个关键领域差距:大多数人形机器人系统缺乏丰富的触觉感知。我们提出了 TactileStep 来解决这一问题,这是一个可部署的触觉学习框架,将足底压力感知引入人形机器人运动控制中,以实现更柔和的着地和更稳定的支撑。TactileStep 将触觉仿真与真实的压力鞋垫对齐,使策略能够从硬件上可获得的相同接触特征中学习。在训练过程中,我们利用触觉和运动线索识别不同的足部接触阶段,并施加具有阶段感知的奖励,以鼓励更安全的着陆和更稳定的支撑。在仿真以及 Unitree G1 人形机器人上的多种地形评估中,与强大的感知基线相比,TactileStep 将峰值着地力降低最多 48.8%,将 A 计权峰值冲击噪声降低最多 30.1 dB,同时将支撑接触面积增加最多 23.8%。
cs.RO / 21 / 2609.28960

Echo in the Steps: Learning Perceptive Humanoid Parkour with Gated Memory

足迹中的回声:基于门控记忆的感知型仿人机器人跑酷学习
Lee, Ming-Ju, Wang, Zizhuo, Zhu, Shaoting, Lou, Haozhe, Zhao, Hang, Li, Yiming
Abstract
While recent advances in perceptive locomotion have enabled humanoid robots to traverse structured terrains, agile parkour in highly discontinuous environments remains an open challenge. In particular, crossing sparse footholds and narrow support regions requires precise foothold selection, effective use of visual observations, and consistent alternating foot placement during fast transitions. In this paper, we present a perceptive humanoid parkour framework that enables stable traversal across terrains with limited foothold availability using only onboard depth observations. The framework features a saliency-guided temporal perception module that combines a saliency prior with gated memory. It retains informative depth features across frames, enabling reliable foot placement from partial observations. By introducing an alternation loss, our symmetry regularization encourages alternating gait patterns and improves traversal robustness. Extensive experiments show that our method significantly improves success rate and foothold accuracy on challenging terrains in both simulation and the real world.
Chinese Translation
尽管感知运动方向的最新进展已使人形机器人能够穿越结构化地形,但在高度不连续环境中的敏捷跑酷仍然是一个开放的挑战。特别是,跨越稀疏落脚点和狭窄支撑区域需要精确的落脚点选择、对视觉观测的有效利用,以及快速过渡过程中一致的双脚交替放置。在本文中,我们提出了一种感知型仿人机器人跑酷框架,仅使用机载深度观测即可在落脚点有限的复杂地形上实现稳定穿越。该框架的核心是一个显著性引导的时序感知模块,它将显著性先验与门控记忆(gated memory)相结合,能够跨帧保留有信息量的深度特征,从而在部分观测条件下实现可靠的落脚点放置。通过引入交替损失(alternation loss),我们的对称性正则化方法鼓励交替步态模式并提升了穿越的鲁棒性。大量实验表明,我们的方法在仿真和真实世界中均显著提高了挑战性地形上的成功率和落脚点准确性。
cs.RO / 22 / 2609.28969

Sim-to-Real Aware End-to-End Learning Environment for Micromobility

面向微出行的感知仿真到现实差距的端到端学习环境
Amano, Shouma, Azumi, Takuya
Abstract
While end-to-end autonomous driving systems show promise, their application to micromobility vehicles is hindered by simulators failing to capture specific kinematics, such as differential drives and omni-wheels. This paper pro- poses a sim-to-real-aware, vehicle-specific end-to-end learning environment for the WHILL Model CR on AWSIM and ROS 2. To minimize the sim-to-real gap, physical parameters are optimized via Bayesian optimization using real-world data, reducing trajectory errors across various driving scenarios. Additionally, this study introduces a synchronized architecture tailored for the stable training of world model-based agents. An end-to-end policy trained with DreamerV3 exhibited learning progress and achieved task completion in a simulated obstacle avoidance setting. Furthermore, this policy demonstrated direct sim-to-real transfer to the physical vehicle, enabling the vehicle to navigate around a cardboard box in a real-world corridor replica without fine-tuning. This paper provides a practical foundation for sim-to-real micromobility policy studies.
Chinese Translation
尽管端到端自动驾驶系统展现出广阔前景,但由于仿真器难以准确刻画差速驱动和全向轮等特定运动学特性,其在微出行(Micromobility)车辆上的应用受到阻碍。本文基于 AWSIM 和 ROS 2,针对 WHILL Model CR 提出了一种感知仿真到现实差距的、车辆专属的端到端学习环境。为最小化仿真与现实之间的差距,利用真实世界数据通过贝叶斯优化对物理参数进行优化,从而降低了多种驾驶场景下的轨迹误差。此外,本研究还引入了一种同步架构,专为基于世界模型(world model)的智能体的稳定训练而设计。使用 DreamerV3 训练的端到端策略展现出学习进展,并在仿真避障环境中完成了任务。更进一步,该策略实现了向物理车辆的直接仿真到现实迁移,无需微调即可使车辆在真实走廊复现场景中绕过纸箱障碍物。本文为仿真到现实的微出行策略研究提供了实用的基础。
cs.RO / 23 / 2609.28973

AquaMend: Minimal Re-probing and Conditional Rollback for Latent-Belief Failures in Embodied Agents

AquaMend:面向具身智能体潜在信念失效的最小化重新探测与条件回滚方法
Liu, Yufan, Luo, Shang, Liu, Yang, Jia, Haoxuan, Han, Feiyu, Li, Qian, Li, Chen, Yang, Yingguang, Zhang, Chongyang, Zheng, Hao, Xu, Kefu, Chong, Bin
Abstract
Physical changes or sensing errors can invalidate embodied agents' task-relevant beliefs. AquaMend compares re-probing, rollback, and supported continuation on a probe-belief-action graph under an expected-loss objective covering sensing, physical recovery, and uncorrected failures. A joint posterior guides a one-step policy with conditional detection-power screening. The per-belief three-way optimum requires independence, separability, and fully resolving probes; the general policy has no global optimality guarantee. Across 32 paired scenarios in a self-constructed simulation benchmark, AquaMend recovers in 28/32 cases and reduces mean complete loss by 21.6% versus restart. Its paired loss difference from decision-theoretic troubleshooting (DTT) is not statistically significant after Holm correction. Against the all-candidate ablation, online decision time decreases by 12.3% overall but increases by 3.4% in the uncovered late stage.
Chinese Translation
物理环境变化或感知错误可能使具身智能体(embodied agents)的任务相关信念失效。AquaMend 在一个探测-信念-动作图上,以涵盖感知、物理恢复及未纠正失败的期望损失为目标,比较了重新探测、回滚与受支持的继续执行三种策略。该方法通过联合后验概率引导带有条件检测力筛选的单步策略。逐信念的三方最优解要求满足独立性、可分离性以及探测的完全可分辨性;一般策略则不具备全局最优性保证。在自建仿真基准的32个配对场景中,AquaMend 在28/32个场景中成功恢复,相比重启策略将平均完全损失降低了21.6%。其与决策论故障排查(Decision-Theoretic Troubleshooting, DTT)的配对损失差异经 Holm 校正后不具有统计显著性。与全候选消融方法相比,在线决策时间总体降低12.3%,但在未覆盖的后期阶段增加了3.4%。
cs.RO / 24 / 2609.28976

ReVNM: Learning-Based Visual Navigation from a Remote Camera

ReVNM:基于学习的远程相机视觉导航
Eguchi, Michikuni, Honda, Kohei, Endo, Masafumi, Yoshimura, Yasuhiro, Yonetani, Ryo
Abstract
Visual Navigation Models (VNMs) enable robots to navigate from egocentric visual observations without geometric localization and planning, but long-range navigation still requires pre-built maps. This paper presents the Remote Visual Navigation Model (ReVNM), which uses a single remote surveillance camera to serve as both an observation source and an implicit environmental map for visual navigation. While the use of remote cameras could eliminate the need for pre-built maps as well as onboard vision processing, their limited field of view instead of egocentric observations makes it hard to achieve collision-free navigation. The lack of existing data with diverse remote viewpoints, which are crucial for training robust VNMs, further complicates the challenge. In this work, we propose a learning-by-synthesis approach to address this two-fold challenge. Our ReVNM extends a state-of-the-art VNM architecture with an exocentric-to-egocentric (exo2ego) module that predicts an egocentric depth observation from remote-camera observations. This helps the VNM to plan a path while considering obstacles in front of the robot. Trained only on randomly generated worlds with diverse obstacle layouts and camera viewpoints, ReVNM can generalize well to real robot navigation without additional fine-tuning. Experiments in both simulation and real-world environments confirmed the effectiveness of the proposed approach.
Chinese Translation
视觉导航模型(Visual Navigation Models, VNMs)使机器人能够基于自我中心(第一人称)视觉观测进行导航,而无需几何定位与规划,但长距离导航仍需要预先构建的地图。本文提出了远程视觉导航模型(Remote Visual Navigation Model, ReVNM),利用单个远程监控相机同时充当观测源和隐式环境地图来实现视觉导航。虽然使用远程相机可以免除预建地图和机载视觉处理的需求,但其有限视野提供的并非自我中心观测,这使得实现无碰撞导航变得困难。此外,现有数据缺乏训练鲁棒VNMs所必需的多样化远程视角,进一步加剧了这一挑战。本文提出一种“基于合成的学习”(learning-by-synthesis)方法来应对这一双重挑战。我们的ReVNM在最新的VNM架构基础上,扩展了一个从外部视角到自我视角(exocentric-to-egocentric, exo2ego)的模块,该模块能够从远程相机观测中预测自我视角的深度观测。这有助于VNM在规划路径时考虑机器人前方的障碍物。ReVNM仅在具有多样化障碍物布局和相机视角的随机生成世界中训练,无需额外微调即可良好泛化到真实机器人导航。仿真和真实环境中的实验均验证了所提方法的有效性。
cs.RO / 25 / 2609.28984

CrossSafe: Towards Cross-Embodiment Latent Safety Filters

CrossSafe:面向跨本体(Cross-Embodiment)的潜在安全过滤器
Tabbara, Ihab, Yang, Yuxuan, Sibai, Hussein
Abstract
Cross-embodiment learning has shown that a single model, such as a vision-language-action (VLA) model, can learn state representations and manipulation skills that can be applied across heterogeneous robots to accomplish various tasks. We hypothesize that the same holds for safety enforcement. The reasoning required to satisfy a safety constraint, such as detecting an obstacle, recognizing that it should be avoided, and selecting a safe abstract action, is largely shared across robots. What differs across embodiments is how the abstract safe action is realized: morphology, kinematics, and dynamics determine which actions are safe and feasible. Consequently, the same action can be safe for one robot and unsafe for another. This is especially important for generalist manipulation policies that operate in a common end-effector action space without explicitly capturing how safety depends on the robot's morphology and kinematics. We propose embodiment-conditioned safety filtering, in which a Hamilton-Jacobi reachability-based value function and its corresponding safety-maximizing policy are shared across robots. Using a morphology-aware latent representation of the robot and its environment, we perform Hamilton-Jacobi reachability analysis directly in latent space so that the learned safety concepts can generalize across embodiments while remaining explicitly conditioned on each robot's morphology and kinematics. We evaluate our approach across five bimanual robot embodiments and five manipulation tasks with whole-body collision-avoidance constraints. Our results show that a single policy, jointly trained across five manipulation tasks and four embodiments, exhibits zero-shot generalization to a held-out embodiment, reducing the nominal policy's collision rate. They also show that training using more embodiments improves generalization.
Chinese Translation
跨本体学习已经表明,单一模型(如视觉-语言-动作(VLA)模型)可以学习到能够应用于异构机器人以完成各种任务的状态表示和操作技能。我们假设安全约束的实施也具有同样的性质。满足安全约束所需的推理——例如检测障碍物、识别应避开障碍物以及选择安全的抽象动作——在很大程度上是跨机器人共享的。不同本体之间的差异在于如何实现抽象的安全动作:形态结构(morphology)、运动学和动力学决定了哪些动作是安全且可行的。因此,同一个动作对一个机器人可能是安全的,而对另一个机器人则是不安全的。这一点对于在统一的末端执行器动作空间中运行、且未显式刻画安全性如何依赖于机器人形态与运动学的通用操作策略尤为重要。我们提出了一种本体条件化的安全过滤方法,其中基于哈密顿-雅可比可达性(Hamilton-Jacobi reachability)的价值函数及其对应的安全最大化策略在多个机器人之间共享。利用对机器人及其环境的形态感知潜在表示,我们直接在潜在空间中进行哈密顿-雅可比可达性分析,使所学的安全概念能够跨本体泛化,同时显式地以每个机器人的形态和运动学为条件。我们在五种双臂机器人本体和五个带全身碰撞避免约束的操作任务上评估了我们的方法。结果表明,在五个操作任务和四个本体上联合训练的单一策略,能够对保留的本体实现零样本泛化,降低了名义策略的碰撞率。结果还表明,使用更多本体进行训练可以提升泛化能力。
cs.RO / 26 / 2609.29017

CALM: Current Aligned Link Manipulation for Single Arm Oversized Object Lifting

CALM:面向单臂超大物体搬运的电流对齐连杆操控方法
Hu, Jun, Chen, Sihan, Jovanovic, Kosta, Navarro-Alarcon, David, Wang, Xueqian, Pan, Jia, Zhou, Peng
Abstract
Most robots manipulate objects solely with their end effectors, whereas humans flexibly leverage different body parts, such as the forearm and elbow, especially when handling oversized objects. Learning such whole-arm manipulation is chal-lenging due to long-horizon sparse rewards, limited contact sens-ing, and the sim-to-real gap in contact and actuator dynamics. To address these challenges, we propose Current-Aligned Link Manipulation, a framework for learning long-horizon contact-rich manipulation using motor current as joint load related feedback. Three stage-specific policies first learn repositioning, grasping, and lifting using privileged simulation information, and a stage router sequences them to generate complete task demonstrations. For sim-to-real transfer, a causal current mapper predicts physical motor current from simulated joint histories, aligning the actuator current observation between simulation and hardware. A unified student policy then learns from these demonstrations using only deployable sensor observations and is further refined with DAgger. The task policies are trained entirely in simulation, and the final student is deployed on hardware. Experiments demonstrate 76.2% (762/1000 trials) complete-task success in simulation and 73.3% success (22/30 trials) on the physical robot for sequential oversized-object lifting.
Chinese Translation
大多数机器人仅依靠末端执行器来操纵物体,而人类则能灵活利用身体的不同部位,如前臂和肘部,尤其是在搬运超大物体时。学习这类全臂操控面临诸多挑战:长时程稀疏奖励、有限的接触感知,以及接触与执行器动力学中的仿真到现实(sim-to-real)差距。为解决这些挑战,我们提出了电流对齐连杆操控(Current-Aligned Link Manipulation)框架,该框架以电机电流作为与关节负载相关的反馈,学习长时程的接触密集型操控任务。三个特定阶段的策略首先利用特权仿真信息分别学习重定位、抓取和提起动作,随后由一个阶段路由器将它们按序衔接,生成完整的任务演示。在仿真到现实的迁移方面,一个因果电流映射器从仿真的关节历史中预测物理电机电流,使仿真与硬件之间的执行器电流观测保持一致。接着,一个统一的学生策略仅使用可部署的传感器观测从这些演示中学习,并通过DAgger进一步精化。任务策略完全在仿真中训练,最终的学生策略被部署到真实硬件上。实验表明,在仿真中序列化超大物体搬运任务的完整成功率为76.2%(1000次试验中762次成功),在真实机器人上的成功率为73.3%(30次试验中22次成功)。
cs.RO / 27 / 2609.29020

Outcome-Sensitive Motion Search for Impact-Aware Dexterous Catching

面向结果敏感的运动搜索以实现具有冲击感知的灵巧抓取
Pei, Guorui, Wu, Jinsong, Su, Songyuan, Qi, Jiaming, Liu, Sichao, Navarro-Alarcon, David, Liu, Bin, Zhou, Peng
Abstract
Skilled humans can catch fast-moving objects softly by coordinating interception, velocity matching, and follow-through to mitigate impact. Learning such impact-aware catching with reinforcement learning (RL), however, is challenging, as the policy must achieve reliable interception and grasping while regulating the sensitive transition into contact. Moreover, even a capable privileged-state RL teacher may not provide ideal demonstrations for a deployable imitation-learning (IL) student: teacher failures limit task coverage, while small variations in pre-contact motion can produce substantially different impact and grasping outcomes. We characterize this phenomenon through interventional outcome sensitivity and introduce the outcome-sensitive window (OSW) to guide targeted demonstration construction. Building on this formulation, we propose Outcome-Sensitive Motion Search, which learns a task-conditioned manifold of successful OSW motions and performs local geodesic search to refine successful teacher rollouts and repair task conditions where the teacher fails. We then validate candidate motions through complete rollouts under a calibrated IL-student action-error model and retain only successful executions as demonstrations. Extensive simulation experiments demonstrate that our method effectively repairs task conditions where the teacher fails and enables the resulting IL policy to outperform the privileged RL teacher in both catching success and impact mitigation.
Chinese Translation
熟练的人类可以通过协调拦截、速度匹配和随动缓冲来柔和地接住快速移动的物体,从而减轻冲击。然而,利用强化学习(RL)学习此类冲击感知的接物任务极具挑战性,因为策略不仅需要实现可靠的拦截与抓取,还必须精细调控进入接触的敏感过渡过程。此外,即使是一个能力强大的基于特权状态的强化学习教师,也未必能为可部署的模仿学习(IL)学生提供理想的示范:教师的失败限制了任务覆盖范围,而接触前运动的微小变化可能导致截然不同的冲击和抓取结果。我们通过干预性结果敏感度来刻画这一现象,并引入结果敏感窗口(Outcome-Sensitive Window, OSW)以指导有针对性的示范构建。基于这一框架,我们提出了结果敏感运动搜索(Outcome-Sensitive Motion Search),该方法学习一个以任务为条件的成功OSW运动流形,并通过局部测地线搜索来优化教师的成功轨迹,并修复教师失败时的任务条件。随后,我们在校准的模仿学习学生动作误差模型下,通过完整的轨迹执行来验证候选运动,仅保留成功执行的结果作为示范。大量仿真实验表明,我们的方法能够有效修复教师失败时的任务条件,使最终的模仿学习策略在接物成功率和冲击缓解两方面均优于特权强化学习教师。
cs.RO / 28 / 2609.29021

CAMP: Cooperative Arm-Hand Motion Planning in Constrained Spaces

CAMP:受限空间下的协作式臂-手运动规划
Wang, Ziyuan, Shan, Yunlong, Mo, Fei, Liu, Sichao, Navarro-Alarcon, David, Pan, Jia, Jovanovic, Kosta, Jiang, Xin, Zhou, Peng
Abstract
Coordinated arm-hand motion planning is fundamental to dexterous robotic manipulation in complex and constrained environments. A straightforward solution is to decompose the problem into separate arm path planning and hand motion generation; however, this poses a dilemma: decomposition can miss feasible solutions that require coordinated arm-hand adaptation along the path. Alternatively, directly planning in the high-dimensional joint arm-hand configuration space captures such coupling but faces a substantially enlarged search space and nonconvex collision constraints. To characterize this coupling, we formulate feasible hand fibers that capture collision-free hand configurations for each arm configuration. Based on this formulation, we propose CAMP, a high-success and efficient cooperative arm-hand motion planner for constrained environments. CAMP constructs candidate trajectories through layered hand search with local arm relaxation, then compactly represents them using endpoint-preserving via-point movement primitives (VMPs) for coarse-to-fine joint optimization. Across six constrained simulation tasks, CAMP achieves 84.2-98.5% planning success, outperforming alternative planners with competitive efficiency. Ablation studies verify the contributions of arm relaxation, VMP representation, and coarse-to-fine optimization, while real-robot experiments demonstrate CAMP on constrained manipulation tasks. The project website is available at https://camp-armhand.github.io/.
Chinese Translation
臂-手协调运动规划是复杂受限环境中灵巧机器人操作的基础。一种直接的解决方案是将问题分解为独立的机械臂路径规划与手部运动生成;然而,这带来了一个两难困境:分解方法可能遗漏那些需要沿路径进行臂-手协调自适应的可行解。另一种方案是直接在臂-手联合的高维构型空间中规划,虽能捕捉这种耦合关系,但面临搜索空间大幅扩大以及非凸碰撞约束的挑战。为刻画这种耦合关系,我们提出了“可行手部纤维”的概念,用以描述每个机械臂构型下无碰撞的手部构型。基于该表述,我们提出了CAMP,一种面向受限环境的高成功率、高效率的协作式臂-手运动规划器。CAMP通过带局部机械臂松弛的分层手部搜索构建候选轨迹,然后使用保留端点的途经点运动基元(via-point movement primitives, VMPs)对其进行紧凑表示,以实现由粗到细的联合优化。在六个受限仿真任务中,CAMP实现了84.2%–98.5%的规划成功率,以具有竞争力的效率优于其他规划器。消融实验验证了机械臂松弛、VMP表示以及由粗到细优化的贡献,真实机器人实验则展示了CAMP在受限操作任务上的有效性。项目网站见 https://camp-armhand.github.io/。
cs.RO / 29 / 2609.29031

Simple Torque-Observation Alignment for Zero-Shot Sim-to-Real Grasping with a Direct-Drive Gripper

基于直接驱动夹爪的零样本仿真到真实抓取的简单力矩观测对齐方法
Kim, Doyoung, Lee, Edgar, Park, Hyeonsun, Lee, Chunghyeon, Han, Chihyun, Hwang, Uisu, Jeong, Seokhwan
Abstract
Torque observations in reinforcement learning remain challenging because simulated and measured torque differ in scale, offset, and noise. In this paper, we propose a simple torque observation alignment method for robots with direct-drive (DD) actuators, in which motor current maps linearly to joint torque through a motor-type-specific torque constant K_tau. First, dynamometer calibration identifies K_tau* and corrects the scale mismatch between simulated and real torque. Second, the method uses delta_tau(t) = tau(t) - tau(t-1) as the observation in both domains to eliminate the constant offset instead of using the direct torque tau(t), which carries a domain-dependent bias. Third, Gaussian noise obtained from the dynamometer measurement data is injected during the learning process. To validate the proposed method, we train a teacher-student grasping policy entirely in simulation and deploy the distilled student on a multifingered DD gripper. The deployed policy performs proprioceptive grasping using only joint positions and torque differences. We conduct an ablation study comparing the proposed method with alternative alignment variants on nine in-distribution (ID) objects. The proposed method achieves 100% grasp success. These results demonstrate that the proposed alignment method improves the robustness of zero-shot policy transfer on the DD gripper against real-world torque-observation mismatches.
Chinese Translation
在强化学习中,力矩观测仍然具有挑战性,因为仿真力矩与实测力矩在幅值、偏置和噪声方面存在差异。本文针对采用直接驱动(Direct-Drive, DD)执行器的机器人,提出了一种简单的力矩观测对齐方法。在直接驱动执行器中,电机电流通过电机类型特定的力矩常数 K_tau 线性地映射为关节力矩。首先,通过测力计标定确定 K_tau*,并校正仿真力矩与真实力矩之间的幅值失配。其次,该方法在两个域中均采用 delta_tau(t) = tau(t) - tau(t-1) 作为观测,以消除恒定偏置,而不是使用携带域相关偏差的直接力矩 tau(t)。第三,在训练过程中注入由测力计测量数据获得的 Gaussian 噪声。为验证所提方法,我们完全在仿真中训练了一个教师-学生抓取策略,并将蒸馏后的学生策略部署在多指直接驱动夹爪上。部署的策略仅使用关节位置和力矩差即可完成本体感知抓取。我们通过消融实验,在九个分布内(in-distribution, ID)物体上将所提方法与其他对齐变体进行比较,所提方法实现了100%的抓取成功率。这些结果表明,所提出的对齐方法提高了零样本策略迁移在直接驱动夹爪上对真实世界力矩观测失配的鲁棒性。
cs.RO / 30 / 2609.29043

Design and Evaluation of LLM Chaining-Based Task Planning for General Purpose Service Robots

基于大语言模型链式架构的通用服务机器人任务规划设计与评估
Bruno, Lucas Da Mota, Sim, Jiahao, Hagiwara, Yoshinobu
Abstract
General Purpose Service Robot (GPSR) tasks, as defined in the RoboCup@Home benchmark, require robots to interpret diverse natural language commands and generate multi-step action sequences in real home environments. Conventional Single Prompt (SP) approaches suffer from context bloat and the "Lost in the Middle" phenomenon, leading to unreliable task planning. We propose an LLM chaining architecture that separates instruction classification and action generation into two specialized stages, reducing per-inference prompt length by approximately 45% while improving planning consistency. We evaluate our method using 100 randomly generated GPSR commands across three language models spanning local open-source and frontier cloud deployment contexts. Results show consistent planning improvements over SP across all models, with gains of up to +37 percentage points on local models. Further, real-robot execution experiments on the Toyota Human Support Robot (HSR) reveal that planning success alone does not guarantee task completion, with 6 of 10 tasks completing successfully and execution-layer failures identified as the primary remaining bottleneck.
Chinese Translation
RoboCup@Home 基准中定义的通用服务机器人(General Purpose Service Robot, GPSR)任务要求机器人在真实家庭环境中理解多样化的自然语言指令并生成多步动作序列。传统的单提示(Single Prompt, SP)方法存在上下文冗余膨胀以及“迷失于中间”(Lost in the Middle)现象,导致任务规划的可靠性不足。我们提出了一种大语言模型(LLM)链式架构,将指令分类与动作生成分离为两个专门化阶段,使每次推理的提示长度减少约 45%,同时提升了规划的一致性。我们在覆盖本地开源模型与前沿云端部署场景的三种语言模型上,使用 100 条随机生成的 GPSR 指令对该方法进行评估。结果显示,在所有模型上,本方法相比单提示方法均带来一致的规划性能提升,其中在本地模型上最高提升 37 个百分点。此外,在丰田人类支持机器人(Toyota Human Support Robot, HSR)上的真实机器人执行实验表明,仅规划成功并不能保证任务完成——10 项任务中有 6 项成功完成,且执行层的失败被确认为当前最主要的剩余瓶颈。
cs.RO / 31 / 2609.29065

DA-GRD: Decision-Aware Grasp-Relevant Disambiguation for tactile recovery under perception-to-execution mismatches

DA-GRD:面向感知-执行不匹配情形下触觉恢复的决策感知抓取相关消歧方法
Wang, Haoran, Sun, Yuteng, Li, Yuanjie, Bai, Ruofei, Yee, Meng, Chuah, Liang, Wenyu, Li, Jun, Yau, Wei-Yun
Abstract
Grasping is a fundamental robotic capability that bridges perception and physical task execution. This paper studies grasp pose recovery under a perception-to-execution mismatch, where a grasp generated from visual perception may become spatially stale if the object moves before execution, using only sparse tactile interactions and no further visual observations. We propose DA-GRD, Decision-Aware Grasp-Relevant Disambiguation, which maintains a weighted planar belief over possible object configurations and selects tactile probes according to their ability to eliminate hypotheses and improve agreement among candidate task grasps. Rather than fully relocalizing the object, DA-GRD stops when the remaining hypotheses support a common executable grasp. In MuJoCo experiments on ten rigid objects with translations up to 5~cm and yaw perturbations up to $\pm45^\circ$, DA-GRD achieves an 84.7% physical lift success rate, compared with 9.1% for stale AnyGrasp, 21.2% for the original fix-scan baseline, and 63.7% for fix-scan method adapted with an SE(2) belief. DA-GRD also achieves a 57.3% Task conditioned Success rate. Across objects, it uses a success-average of 4.13 tactile probes over the ten per-object means, corresponding to a 72.5% reduction relative to the fixed 15-probe baselines. Real-world experiments on six objects achieve 71.7% physical lift success and 38.3% task-conditioned success with 4.20 probes on average. These results show that tactile sensing can recover task-relevant grasps under vision-off conditions with limited physical interaction, without requiring complete object localization.
Chinese Translation
抓取是机器人连接感知与物理任务执行的基本能力。本文研究了感知到执行不匹配情形下的抓取位姿恢复问题:当由视觉感知生成的抓取在执行前因物体移动而在空间上失效时,仅利用稀疏的触觉交互且不再进行视觉观测来恢复抓取。我们提出DA-GRD(决策感知的抓取相关消歧,Decision-Aware Grasp-Relevant Disambiguation),该方法对物体可能构型维持一个加权平面信念分布,并根据各触觉探测消除假设、提升候选任务抓取一致性的能力来选择触觉探测点。DA-GRD并不追求对物体进行完整重定位,而是当剩余假设能够支持一个共同的可执行抓取时即停止。在MuJoCo仿真实验中,针对十个刚体物体(平移最大达5厘米、偏航扰动最大达±45°),DA-GRD取得了84.7%的物理提起成功率,相比之下,失效的AnyGrasp为9.1%,原始fix-scan基线为21.2%,采用SE(2)信念改进的fix-scan方法为63.7%。DA-GRD还达到了57.3%的任务条件成功率。在十个物体的各物体平均探测次数上,DA-GRD的成功平均探测次数为4.13次,相对于固定15次探测的基线减少了72.5%。在六个物体上的真实世界实验中,平均使用4.20次触觉探测,实现了71.7%的物理提起成功率和38.3%的任务条件成功率。这些结果表明,触觉感知能够在无视觉条件下通过有限的物理交互恢复与任务相关的抓取,而无需对物体进行完整的定位。
cs.RO / 32 / 2609.29091

From Passive Execution to Active Exploration: Agentic Embodied Manipulation in Realistic Environments

从被动执行到主动探索:真实环境中的智能体具身操作
Ma, Shilin, Zhang, Chubin, Bai, Xulong, Gao, Zifeng, Zhang, Shiyi, Tang, Yansong
Abstract
Recent advances in agentic systems have substantially enhanced the long-horizon capability of embodied manipulation. However, many existing frameworks still follow a passive execution paradigm, which limits their applicability to real-world scenarios involving textual semantic cues, distractors, and initially invisible targets. To bridge this gap, we propose an agent-based active exploration framework that enables robots to dynamically interact with the environment rather than merely execute predefined instructions. Specifically, our framework consists of three collaborative modules: a planning module for high-level task reasoning, a perception module for visual scene understanding, and an execution module for low-level manipulation. This design allows the robot to actively acquire task-relevant information, adapt its behavior based on environmental feedback, and complete manipulation tasks under partial observability. Furthermore, we introduce a fine-grained perception-execution interleaving strategy, which tightly couples visual feedback with skill execution to improve exploration robustness. We evaluate our method on a realistic Find-and-Place task, demonstrating its effectiveness in challenging environments where target objects must be actively discovered before manipulation.
Chinese Translation
智能体系统的最新进展显著提升了具身操作的长时程能力。然而,许多现有框架仍遵循被动执行范式,这限制了其在涉及文本语义线索、干扰物以及初始不可见目标的真实场景中的适用性。为弥合这一差距,我们提出了一种基于智能体的主动探索框架,使机器人能够与环境进行动态交互,而不仅仅是执行预定义的指令。具体而言,我们的框架由三个协同工作的模块组成:用于高层任务推理的规划模块、用于视觉场景理解的感知模块,以及用于底层操作控制的执行模块。这一设计使机器人能够主动获取与任务相关的信息,根据环境反馈调整自身行为,并在部分可观测条件下完成操作任务。此外,我们提出了一种细粒度的感知-执行交错策略,将视觉反馈与技能执行紧密耦合,以提升探索的鲁棒性。我们在一个真实的“寻找并放置”(Find-and-Place)任务上对所提方法进行了评估,结果表明其在具有挑战性的环境中——即目标物体必须在进行操作前被主动发现——具有显著有效性。
cs.RO / 33 / 2609.29092

DAWN: Noise-Robust Quadruped Parkour via Depth-Denoising World Models

DAWN:基于深度去噪世界模型的噪声鲁棒四足跑酷
Choi, Yohan, Kim, Min-Jun, Kim, Jin-Sung, Kim, Yong-Jae, Han, Youn-Hee
Abstract
Vision-based legged locomotion methods assume clean depth at training time and rely on hand-tuned post-processing filters at deployment. However, filter parameters are rarely disclosed, hindering reproducibility, and performance degrades substantially when depth noise is left unaddressed. Building noise robustness directly into the learning pipeline would eliminate this dependency. While such robustness has been explored for proprioceptive inputs, analogous approaches for depth perception remain largely absent in legged locomotion. We propose DAWN (Denoising and Alignment in World models for Noise-robustness), a noise-robust perception framework for legged locomotion, which builds noise robustness directly into a world model via two modifications: (1) feeding noisy depth to the encoder while keeping clean depth as the reconstruction target, forcing the model to implicitly denoise its input; and (2) applying contrastive learning to align the latent states of noisy and clean depth. Importantly, DAWN is not tied to a specific noise model, requiring no manual tuning to the noise distribution at deployment. Furthermore, it incurs no additional inference cost over existing world model-based methods. Without any manual filter calibration -- relying solely on the learned noise-robust representation -- DAWN achieves zero-shot quadruped parkour on a Unitree Go1: traversing stairs up to 18 cm, clearing gaps up to 70 cm, and mounting steps up to 45 cm from raw depth observations. Ablation studies show that denoising and contrastive alignment contribute at complementary levels -- reconstruction and representation, respectively -- and yield additive gains when combined. Videos and code are available at: https://dawn-parkour.github.io/
Chinese Translation
基于视觉的腿式运动方法通常假设训练时深度图像是干净的,并在部署时依赖人工调优的后处理滤波器。然而,滤波器参数很少被公开,阻碍了研究的可复现性;且当深度噪声未被处理时,性能会显著下降。将噪声鲁棒性直接构建到学习流程中可以消除这种依赖。尽管这种鲁棒性已在本体感知输入中得到探索,但在腿式运动中,针对深度感知的类似方法仍基本缺失。我们提出DAWN(Denoising and Alignment in World models for Noise-robustness,世界模型中的去噪与对齐以实现噪声鲁棒性),一个用于腿式运动的噪声鲁棒感知框架,通过对世界模型的两项改进将噪声鲁棒性直接融入其中:(1)将含噪深度输入编码器,同时以干净深度作为重建目标,迫使模型对其输入进行隐式去噪;(2)应用对比学习来对齐含噪深度与干净深度的潜在状态。重要的是,DAWN不依赖于特定的噪声模型,在部署时无需针对噪声分布进行手动调整。此外,与现有基于世界模型的方法相比,它不产生额外的推理开销。在无需任何手动滤波器校准的情况下——仅依靠学习到的噪声鲁棒表征——DAWN在Unitree Go1上实现了零样本四足跑酷:可攀爬高达18厘米的楼梯、跨越高达70厘米的间隙、并登上高达45厘米的台阶,全部基于原始深度观测。消融实验表明,去噪与对比对齐分别在重建和表征两个互补层面发挥作用,二者结合可产生叠加的增益。视频与代码见:https://dawn-parkour.github.io/
cs.RO / 34 / 2609.29093

A Support-Enhanced Granular-Jamming Gripper for RL-based Grasping with Continuum Manipulators

一种用于连续体机械臂强化学习抓取的支撑增强型颗粒阻塞夹持器
Liu, Danyu, Zhang, Tianlin, Chen, Wei, Tang, Wei, Qin, Kecheng, Li, Zhongyu
Abstract
Continuum manipulators provide dexterous motion in confined spaces, but structural compliance, hysteresis, and load-dependent deformation leave residual position and orientation errors that can undermine reliable contact with rigid grippers. To address this limitation, this paper presents a lightweight support-enhanced granular-jamming gripper tailored to a continuum manipulator. The gripper maintains compliance before jamming while establishing a direct load path to the continuum manipulator tip after jamming. To improve its grasping performance, we systematically designed membrane materials, particles, filling ratios, and the internal support structure, and further identify geometry-dependent grasp boundaries with respect to contact offset and object shape. Building on these results, we construct a physical manipulation system integrating the continuum manipulator, granular-jamming gripper, visual feedback, tendon actuation, and pneumatic control. We then train a reinforcement-learning-based reaching controller in a randomized simulation and deploy it on the physical system, demonstrating how positioning control and contact level mechanical adaptation can complement each other in a modular grasp-and-release task. By introducing an adaptive structure that relaxes the need for highly accurate modeling and positioning control, this work explores a design paradigm that integrates physical and embodied intelligence.
Chinese Translation
连续体机械臂能够在受限空间中实现灵巧运动,但其结构柔顺性、迟滞特性以及负载相关变形会产生残余的位置和姿态误差,从而影响与刚性夹持器的可靠接触。为解决这一局限,本文提出了一种面向连续体机械臂的轻量化支撑增强型颗粒阻塞(granular-jamming)夹持器。该夹持器在阻塞前保持柔顺性,而在阻塞后建立起通向连续体机械臂末端的直接载荷路径。为提升其抓取性能,我们系统性地设计了膜材料、颗粒、填充率及内部支撑结构,并进一步确定了与接触偏移量和物体形状相关的、依赖于几何形状的抓取边界。在此基础上,我们构建了一套集成连续体机械臂、颗粒阻塞夹持器、视觉反馈、腱绳驱动和气动控制的物理操作系统。随后,我们在随机化仿真中训练了基于强化学习的到达控制器,并将其部署于物理系统上,通过一个模块化的抓取-释放任务展示了定位控制与接触层面机械自适应之间的互补作用。通过引入一种降低对高精度建模和定位控制需求的自适应结构,本工作探索了一种融合物理智能与具身智能的设计范式。
cs.RO / 35 / 2609.29103

TRACE: Interactive Bi-Directional Tracing of Monochrome Cables Amid Clutter

TRACE:在杂乱环境中对单色线缆的交互式双向追踪
Shivakumar, Nidhya, Ransing, Ethan, Zhang, Josh, Gowda, Shamak, Yang, Kevin, Hua, Miles, Agrawal, Anika, Yu, Justin, Goldberg, Ken
Abstract
Accurate state estimation (tracing) of Deformable Linear Objects (DLOs) such as cables is a critical challenge for data centers, manufacturing, construction, homes, and surgery, where precise cable management directly impacts operational safety and efficiency. However, resolving the state of multiple monochrome cables amid foreground and background clutter poses challenges due to occlusions, overlap, and ambiguous crossings. We present Two-way Routing And Cable Estimation (TRACE), which combines bi-directional cable tracing with interactive perception primitives-Divergence Push and Cluster Dilation-to actively resolve ambiguities. Evaluation with 110 physical experiments suggests that TRACE can increase the percentage of cable length correctly traced in complex scenarios (with up to 4 cables and 40 crossings) from ~60% with the strongest prior method, HANDLOOM 2.0, to ~90%, outperforming RT-DLO, Nano Banana Pro, and ChatGPT 5.2 as well. For a trial run on a workstation with an NVIDIA GeForce RTX 4090 GPU, the average computation time is 0.4 seconds per cable. Project website: https://trace-paper.github.io/.
Chinese Translation
对线缆等可变形线性物体(Deformable Linear Objects, DLOs)进行精确的状态估计(追踪),是数据中心、制造业、建筑业、家庭以及手术领域的关键挑战,在这些场景中,精确的线缆管理直接影响操作安全与效率。然而,在前景与背景杂乱的环境中解析多条单色线缆的状态,由于遮挡、重叠和歧义交叉而面临诸多困难。我们提出了双向路由与线缆估计方法(Two-way Routing And Cable Estimation, TRACE),该方法将双向线缆追踪与交互式感知基元——发散推动(Divergence Push)和簇膨胀(Cluster Dilation)——相结合,以主动消解歧义。110次物理实验的评估表明,在最复杂场景(最多4条线缆和40个交叉点)中,TRACE可将线缆正确追踪长度的比例从最强的先前方法HANDLOOM 2.0的约60%提升至约90%,同时也优于RT-DLO、Nano Banana Pro和ChatGPT 5.2。在配备NVIDIA GeForce RTX 4090 GPU的图形工作站上进行试运行,每条线缆的平均计算时间为0.4秒。项目网站:https://trace-paper.github.io/。
cs.RO / 36 / 2609.29112

Modeling Load-, Velocity-, and Temperature-Dependent Transmission Errors of Cycloidal Drives for Industrial Robots Using Fourier Series

基于傅里叶级数建模工业机器人摆线针轮传动负载、速度和温度相关传动误差的研究
Bauer, Christian J. E., Kamm, Valentin, Dzubba, Marcel, Steinle, Lukas, Lechler, Armin, Verl, Alexander
Abstract
Industrial robots are rarely used for machining tasks due to their limited path accuracy. This accuracy is mainly limited by inaccuracies in the drive trains. Compliance and transmission errors occur in the joint gearboxes. While transmission errors have been extensively studied for strain wave gears, there is little research on these errors in cycloidal drives. This gearbox type is commonly used in industrial robots for medium to heavy payloads. It is proposed to model the mainly periodic transmission errors using a Fourier series where amplitude and phase are defined as a polynomial function of the main influence factors load-torque, velocity, and temperature. Measurements of the transmission errors were conducted using an experimental setup representing a single robot joint. In the evaluation of the measurement data, harmonic frequencies were related to mechanical properties of the cycloidal drive. These frequencies were used to identify the parameters of the polynomial Fourier series model. Compared to validation measurements, the derived model shows an average root mean square error of 0.026 mrad. It is proposed to use the output of the resulting model in a feedforward control approach to compensate the transmission errors and to increase the path accuracy of industrial robots.
Chinese Translation
由于路径精度有限,工业机器人很少用于加工任务。该精度主要受驱动系统不精确性的限制。关节减速器中存在柔度和传动误差。尽管谐波减速器的传动误差已得到广泛研究,但针对摆线针轮减速器传动误差的研究却很少。这种类型的减速器常用于中重型载荷的工业机器人。本文提出使用傅里叶级数对主要为周期性的传动误差进行建模,其中幅值和相位被定义为主要影响因素——负载扭矩、速度和温度的多项式函数。传动误差的测量在一个模拟单个机器人关节的实验装置上进行。在测量数据的评估中,将谐波频率与摆线针轮减速器的机械特性相关联。这些频率被用于识别多项式傅里叶级数模型的参数。与验证性测量结果相比,所得模型的平均均方根误差为0.026 mrad。建议将所得模型的输出用于前馈控制方法中,以补偿传动误差并提高工业机器人的路径精度。
cs.RO / 37 / 2609.29132

A Tendon-Driven Robotic Jellyfish with Constrained Soft Actuation and Depth Control via Reinforcement Learning

一种基于强化学习实现受限软体驱动与深度控制的腱驱动机器水母
Peng, Jiarui, Wu, Yutong, Wang, Zelong, Deng, Ping, Zhang, Xiaotian, Chen, Xian, Qin, Kecheng, Li, Zhongyi
Abstract
Jellyfish-inspired robots offer a compliant and efficient approach to underwater locomotion, but achieving large deformation together with repeatable actuation and closed-loop control remains challenging. In this work, we present a tendon-driven robotic jellyfish with constrained soft actuation. Each actuator combines a flexible substrate with discrete constraints, enabling bending up to \(150^\circ\) with an approximately linear tendon displacement-bending relationship. Eight actuators driven by four servos allow the robot to perform stable swimming, attitude adjustment, and self-righting. Based on the linear actuation, a reinforcement-learning controller is further developed, enabling closed-loop depth regulation in both simulation and physical experiments. These results show that mechanical constraints can improve the controllability of soft actuation while preserving compliant jellyfish-like motion, providing a route toward manoeuvrable and autonomous jellyfish robots.
Chinese Translation
仿水母机器人为水下运动提供了一种柔顺且高效的方式,但实现大变形以及可重复的驱动和闭环控制仍然具有挑战性。在这项工作中,我们提出了一种采用受限软体驱动的腱驱动(tendon-driven)机器水母。每个驱动器将柔性基底与离散约束相结合,可实现高达150°的弯曲,且腱位移与弯曲角度之间呈近似线性关系。由四个舵机驱动的八个驱动器使机器人能够执行稳定游动、姿态调整和自复位。基于线性驱动特性,我们进一步开发了强化学习控制器,实现了在仿真和物理实验中的闭环深度调节。这些结果表明,机械约束能够在保留柔顺的水母式运动的同时提高软体驱动的可控性,为开发灵活机动且自主的机器水母提供了一条途径。
cs.RO / 38 / 2609.29157

OREN-X: Octree Residual Network for Real-Time Multi-Modal Mapping

OREN-X:用于实时多模态建图的八叉树残差网络
Dai, Zhirui, Qian, Qihao, Nguyen, Dinh Minh, Pham, Quan-Dung, Bronder, Kiana, Nieto-Granda, Carlos, Chen, Yiyu, Nguyen, Quan, Atanasov, Nikolay
Abstract
To achieve general-purpose autonomy over long horizons, a robot needs to maintain spatial environment information that supports a variety of tasks: geometry for planning and control, radiance for rendering and relocalization, and vision-language features for open-vocabulary grounding. Existing methods represent and estimate each modality separately, multiplying memory and compute cost while forgoing potential synergy among the representations. We develop OREN-X, an online mapping method that uses an octree in 3D space as a shared data structure for indexing and storing a multi-modal field, capturing geometric, radiance, and vision-language information. OREN-X provides efficient unified storage and retrieval of these data in explicit/implicit and full/compressed form. Our unified representation yields cross-modality synergy: SDF estimates are sharpened by occupancy and radiance, while GPU-based ray-octree traversal and octree query enable real-time rendering. We also use online dictionary learning to compress the vision-language features, shrinking them 3.7x below full per-vertex storage while raising the query accuracy. On Replica, OREN-X maps in real time (80+ fps for SDF and 30+ fps for all four modalities), improves near-surface SDF accuracy by 33% over single-modality baselines, and improves mean open-vocabulary 3D mIoU by 71% and mean accuracy by 61% over the best prior method.
Chinese Translation
为了实现长时程的通用自主能力,机器人需要维护支持多种任务的空间环境信息:用于规划与控制的几何信息、用于渲染与重定位的辐射信息,以及用于开放词汇定位的视觉-语言特征。现有方法对每种模态分别进行表示和估计,导致内存和计算成本成倍增加,同时放弃了各表示之间潜在的协同作用。我们提出了OREN-X,一种在线建图方法,它使用三维空间中的八叉树(octree)作为共享数据结构,用于索引和存储多模态场,捕获几何、辐射和视觉-语言信息。OREN-X以显式/隐式和完整/压缩的形式对这些数据进行高效的统一存储与检索。我们的统一表示带来了跨模态协同效应:SDF估计通过占据栅格和辐射信息得到锐化,而基于GPU的射线-八叉树遍历和八叉树查询实现了实时渲染。我们还利用在线字典学习压缩视觉-语言特征,使其大小比逐顶点完整存储缩小3.7倍,同时提升了查询精度。在Replica数据集上,OREN-X实现了实时建图(SDF达80+ fps,全部四种模态达30+ fps),近表面SDF精度相比单模态基线提升了33%,开放词汇3D mIoU均值相比最佳先前方法提升了71%,平均准确率提升了61%。
cs.RO / 39 / 2609.29166

HarnessPAI: An Evolving Harness for Physical AI

HarnessPAI:面向物理人工智能(Physical AI)的可演进框架
Wang, Xin, Wu, Wenhao, Zhang, Menghao, Wang, Zhi, Shao, Kun, Luan, Jian, Li, Yang, Li, Qing, Gu, Shangding, Zhou, Huichi, Shi, Shuqing, Ni, Fei, Lu, Shuo, Meng, Weicheng, Li, Kang, Wu, Jin, Zhao, Kang, Guo, Shangmin, Li, Gen, Tang, Yongqiang, Zhang, Zhizhong, Xie, Yuan, Qu, Heng
Abstract
Physical AI aims to build embodied agents that perceive the world, understand and reason about it, and decide how to act. Yet the field has focused primarily on the last component: the action model that maps observations to low-level controls. The prevailing training recipe can erode the perceptual and reasoning capabilities needed for robust behavior, leaving even strong action models vulnerable to scene perturbations and long-horizon tasks. We introduce HarnessPAI, a model- and embodiment-agnostic Harness framework for Physical AI that treats code as the executable and evolvable interface that organizes the underlying action primitive. The framework separates two timescales: within a rollout, it executes open-loop at the program level, with a fixed program guiding and checking execution; across rollouts, it evolves closed-loop, using execution feedback to revise the program and distill failures into reusable skills. Across desktop robot arms, household robots, a robot vacuum, and a legged walking agent, HarnessPAI improves on both pure action models and code-as-policy baselines without retraining the underlying model: a 61.6-point gain over $\pi_{0.5}$ on LIBERO-PRO and a 27.2-point gain over WorldDreamer on RoboCasa atomic tasks. Once a program is selected, rollout execution requires no online high-level LLM deliberation. Beyond execution, the converged program is also a cheap and reliable expert-data collector, and fine-tuning $\pi_{0.5}$ on collected expert data lifts success rate on LIBERO-PRO by 38.8 points. Our results suggest that the frontier of Physical AI depends not only on stronger action models, but also on executable harnesses that integrate perception, task understanding and reasoning, and action execution into a unified, verifiable, and feedback-driven system. Website: https://darwin-agent.github.io/HarnessPAI
Chinese Translation
物理人工智能(Physical AI)旨在构建能够感知世界、理解与推理世界并决定如何行动的具身智能体。然而,该领域的研究主要集中于最后一个环节:将观测映射到低层控制的动作模型。当前主流的训练方式可能会损害实现鲁棒行为所必需的感知与推理能力,使得即使是强大的动作模型在面对场景扰动和长时程任务时也显得脆弱。我们提出 HarnessPAI,一个模型与具身形态无关的物理人工智能框架,它将代码视为可执行、可演进的接口,用以组织底层动作原语。该框架区分两个时间尺度:在单次执行过程中,它以程序为单位进行开环执行,由固定程序引导并检查执行过程;跨多次执行时,它进行闭环演进,利用执行反馈修订程序,并将失败经验提炼为可复用的技能。在桌面机械臂、家庭机器人、扫地机器人以及足式行走智能体上,HarnessPAI 无需重新训练底层模型即可超越纯动作模型和“代码即策略”(code-as-policy)基线:在 LIBERO-PRO 上比 $\pi_{0.5}$ 提升 61.6 个百分点,在 RoboCasa 原子任务上比 WorldDreamer 提升 27.2 个百分点。一旦程序被选定,执行过程无需在线的高层大语言模型(LLM)推理。除执行之外,收敛后的程序还是一个廉价且可靠的专家数据采集器,使用采集的专家数据微调 $\pi_{0.5}$ 可将其在 LIBERO-PRO 上的成功率再提升 38.8 个百分点。我们的结果表明,物理人工智能的前沿不仅取决于更强的动作模型,还取决于能够将感知、任务理解与推理、动作执行整合为统一、可验证且反馈驱动系统的可执行框架。网站:https://darwin-agent.github.io/HarnessPAI
cs.RO / 40 / 2609.29171

Representation World Model: Learning States, Transition and Executable Plans in Representation

表示世界模型:在表示空间中学习状态、转移与可执行规划
Yuan, Yijun, Zheng, Weicheng, Wang, Weibang, Qin, Minghui, Sun, Chang, Huang, Junhao, Li, Kenan, Liu, Anmin, Yao, Yicheng, Zhao, Hang
Abstract
We propose the Representation World Model (RWM), which learns states, transitions, and executable plans directly in representation space. Unlike existing world models that typically learn latent representations together with explicit dynamics models and perform planning through search, optimization, or policy-based prediction, RWM directly incorporates planning into the learned representation geometry. RWM learns the representation geometry by applying inverse-dynamics supervision locally along latent paths constructed from endpoint representations, requiring these paths to preserve task-relevant state and transition information. At inference, planning is performed by directly constructing a latent path between the current and goal representations, with inverse dynamics used to recover the corresponding actions, without recursive rollouts or action-space search. Experiments on continuous-control benchmarks demonstrate the effectiveness of RWM for direct planning, while results on robotic manipulation further show its potential to extend to more complex embodied control tasks. These results suggest that planning directly in representation space provides a promising alternative to conventional world-model planning.
Chinese Translation
我们提出了表示世界模型(Representation World Model, RWM),它直接在表示空间中学习状态、转移和可执行规划。与现有的世界模型通常将潜在表示与显式的动力学模型一同学习、并通过搜索、优化或基于策略的预测来执行规划不同,RWM 直接将规划融入所学习的表示几何结构中。RWM 通过对由端点表示构建的潜在路径施加局部逆动力学监督来学习表示几何结构,要求这些路径保留与任务相关的状态和转移信息。在推理阶段,规划通过直接在当前表示与目标表示之间构建潜在路径来完成,并利用逆动力学恢复相应的动作,无需递归式滚动展开或动作空间搜索。在连续控制基准上的实验证明了 RWM 在直接规划方面的有效性,而在机器人操作任务上的结果进一步表明其有望扩展到更复杂的具身控制任务。这些结果表明,直接在表示空间中进行规划为传统的世界模型规划提供了一种有前景的替代方案。
cs.RO / 41 / 2609.29176

Anthropomimetic Soft Robotic Forearm with Independently Articulated Carpal Bones Enabling Human-Like Adaptive Stiffness Modulability

具有独立活动腕骨的拟人软体机器人前臂,实现类人的自适应刚度调节能力
Obata, Yoshinobu, Jiang, Yinlai, Yokoi, Hiroshi, Togo, Shunta
Abstract
The human wrist exhibits adaptive stiffness modulability: joint stiffness anisotropy can be actively regulated through muscle co-contraction. This functionality is essential for stable manipulation, yet the underlying morphological factors remain unclear. To identify these factors, we developed an anatomically accurate anthropomimetic soft robotic forearm comprising eight independently movable carpal bones interconnected by ligaments, 22 actuated muscles, and compliant fingertips. We measured wrist joint stiffness under four muscle activation patterns across three skeletal configurations: anatomically normal carpal bones, a fused proximal carpal row, and a geometric ellipsoidal skeleton. The stiffness ellipse exhibited low stiffness along the dart-throwing motion (DTM) direction when finger muscles were activated, but high stiffness along the same direction when wrist and finger muscles were activated simultaneously. These results agree with previously reported human measurements, demonstrating that precise anatomical replication reproduces human-like stiffness modulability. Fusing the proximal carpal row eliminated the low DTM-direction stiffness under finger muscle activation, while the geometric ellipsoidal skeleton showed poor stiffness ellipse reorientation across all conditions. Carpal bone motion analysis revealed significantly opposing coupling patterns between wrist and finger muscles at the proximal carpal row, accompanied by a consistent but non-significant trend at the midcarpal joint, providing a mechanical explanation for this modulation. These findings demonstrate that carpal bone morphology plays a dominant role in human wrist stiffness modulation and provide design principles for humanoid robot wrists.
Chinese Translation
人类手腕具有自适应刚度调节能力:关节刚度的各向异性可以通过肌肉共收缩主动调节。这一功能对于稳定操作至关重要,但其背后的形态学因素尚不清楚。为识别这些因素,我们开发了一个解剖学精确的拟人软体机器人前臂,包含八块由韧带连接的独立可动腕骨、22条驱动肌肉以及柔性指尖。我们在三种骨骼构型(解剖学正常的腕骨、融合的近侧腕骨列以及几何椭球形骨骼)下,测量了四种肌肉激活模式下的腕关节刚度。结果显示,当手指肌肉激活时,刚度椭圆沿斜向投掷运动(DTM)方向呈现低刚度;而当手腕和手指肌肉同时激活时,该方向则呈现高刚度。这些结果与已有的人类测量数据一致,表明精确的解剖学复制能够再现类人的刚度调节能力。融合近侧腕骨列后,手指肌肉激活下的DTM方向低刚度消失;而几何椭球形骨骼在所有条件下均表现出较差的刚度椭圆重定向能力。腕骨运动分析揭示,在近侧腕骨列处,手腕肌肉与手指肌肉呈现显著相反的耦合模式,并在腕中关节处伴有一致但不显著的趋势,为这种刚度调节提供了力学解释。这些发现表明腕骨形态在人类手腕刚度调节中起主导作用,并为仿人机器人手腕的设计提供了设计原则。
cs.RO / 42 / 2609.29194

Continuous Online Fault Detection for Mobile Robots via Adaptive Edge Models

基于自适应边缘模型的移动机器人持续在线故障检测
Levy, Jordan, Verstaevel, Nicolas, Talon, Vincent, Gaudou, Benoit
Abstract
Mobile robots require robust, real-time fault detection capable of continuous adaptation on constrained edge hardware. While deep time-series models excel at unsupervised anomaly detection, their computational cost prohibits high-frequency onboard execution. This paper bridges this gap via a Teacher-Student distillation framework. An offline foundation model (TSPulse) generates pseudo-labels from unlabeled time series augmented with fault injections. A lightweight MiniRocket Student, adapted with a Recursive Least Squares estimator, approximates this complex decision boundary to execute real-time inference onboard. Evaluations on the TSB-AD benchmark and a physical mobile robot demonstrate the Student achieves a 4.30 ms CPU inference latency. During real-world domain shifts, online adaptation enables the Student to recover from unseen mechanical degradation, improving VUS-PR scores from 0.26 to 0.75 without catastrophic forgetting. Crucially, an uncertainty-guided active learning strategy minimizes operator cognitive load, requesting sparse interventions only when encountering novel fault distributions. These results validate the deployment of state-of-the-art anomaly detection on resource-constrained robotics through offline-to-online distillation.
Chinese Translation
移动机器人需要能够持续适应的鲁棒实时故障检测能力,且须在资源受限的边缘硬件上运行。尽管深度时间序列模型在无监督异常检测中表现出色,但其计算成本使其难以在机载设备上高频执行。本文通过一个教师-学生蒸馏框架弥合了这一差距。离线基础模型(TSPulse)从经故障注入增强的无标签时间序列中生成伪标签;一个轻量级的 MiniRocket 学生模型借助递归最小二乘估计器进行适配,以逼近该复杂决策边界,从而实现机载实时推理。在 TSB-AD 基准数据集和一台实体移动机器人上的评估表明,学生模型的 CPU 推理延迟仅为 4.30 毫秒。在真实世界的域偏移场景中,在线适应使学生模型能够从未曾见过的机械退化中恢复,将 VUS-PR 分数从 0.26 提升至 0.75,且未出现灾难性遗忘。尤为关键的是,不确定性引导的主动学习策略将操作员的认知负担降至最低,仅在遇到新型故障分布时才请求稀疏的人工干预。这些结果验证了通过离线到在线蒸馏,在最先进异常检测技术于资源受限机器人平台上部署的可行性。
cs.RO / 43 / 2609.29198

Assessing the Impact of Fleet Size on Crowdsourced Mapping Using a Dissimilarity Measure

基于相异性度量评估车队规模对众包地图的影响
Kalenda, Marie-Ngoïe Badibanga, Bonnifait, Philippe, Mittet, Marie-Anne
Abstract
Accurate digital maps are essential for Advanced Driver Assistance Systems (ADAS) or Autonomous Driving (AD), providing critical information such as road geometry, traffic signs and speed limits required by safety functions including Intelligent Speed Assistance (ISA). Maintaining these map layers using traditional surveying methods is costly and difficult to scale. Crowdsourced approaches based on fleets provide a promising alternative for continuously validating and updating map information. However, the relationship between the number of contributing vehicles and the quality of the resulting map remains poorly understood. To address this gap, this paper presents a simulation-based framework for evaluating crowdsourced traffic sign maintenance using a dissimilarity measure called GOSPAM (Generalized Optimal SubPattern Assignment for Maps), which combines localization errors with detection performance by accounting for False Positives (FP) and False Negatives (FN). The proposed system models multivehicle observations with representative sensor noise, detection errors, and semantic recognition uncertainties. Observations from multiple vehicles are aggregated using spatial clustering and semantic filtering to estimate traffic sign locations. Using simulated trajectories generated from data carried out by an experimental vehicle in an area containing ground-truth traffic signs, we assess the influence of fleet size on the performance of crowdsourced mapping. The number of vehicles ranges from 5 to 50, and performance is analyzed using standard evaluation metrics which are compared to the GOSPAM . The results show that GOSPAM can be used to effectively assess the quality of crowdsourced mapping, such as the contributions made by the first vehicles or the improvements made by numerous vehicles.
Chinese Translation
精确的数字地图对高级驾驶辅助系统(ADAS)和自动驾驶(AD)至关重要,可提供道路几何、交通标志和限速等安全功能(如智能速度辅助(ISA))所需的关键信息。采用传统测量方法维护这些地图图层成本高昂且难以扩展。基于车队的众包方法为持续验证和更新地图信息提供了一种有前景的替代方案。然而,参与车辆数量与最终地图质量之间的关系仍缺乏充分理解。为填补这一空白,本文提出了一个基于仿真的框架,用于评估众包交通标志的维护,该框架采用一种名为GOSPAM(Generalized Optimal SubPattern Assignment for Maps,面向地图的广义最优子模式分配)的相异性度量,通过考虑假阳性(FP)和假阴性(FN),将定位误差与检测性能相结合。所提出的系统对多车辆观测进行建模,包含代表性的传感器噪声、检测误差和语义识别不确定性。来自多辆车的观测通过空间聚类和语义过滤进行聚合,以估计交通标志的位置。利用由实验车辆在包含地面真值交通标志的区域中采集数据生成的仿真轨迹,我们评估了车队规模对众包地图性能的影响。车辆数量在5到50之间变化,并使用标准评估指标对性能进行分析,同时与GOSPAM进行比较。结果表明,GOSPAM可有效评估众包地图的质量,例如最初若干车辆所作的贡献以及大量车辆带来的改进。
cs.RO / 44 / 2609.29204

AdaHVLA: Adaptive Harnesses for Long-Horizon Vision-Language-Action Execution

AdaHVLA:面向长时域视觉-语言-动作执行的自适应任务框架
Tang, Junyi, Peng, Jie, Ding, Zezhen, Shen, Yuan, Chen, Tianlong
Abstract
Vision-language-action (VLA) models offer strong local control and instruction following but often struggle with long-horizon tasks requiring persistent memory and planning. Task harnesses provide persistent context for agent reasoning by retaining task history and tracking progress across execution stages. To bring these complementary capabilities together, we introduce AdaHVLA, an adaptive harness that refines code-based coordination policies through robot experience to better align agent reasoning and memory with VLA execution. Its decoupled multiagent adaptation process separates evidence analysis, harness revision, and behavioral assessment into distinct working contexts, using testable coordination hypotheses to guide revisions and subsequent rollouts to assess their predicted effects. A stateful revision graph links execution evidence, hypotheses, revisions, and observed effects, preserving alternative harnesses and adaptation memory to guide refinement across repeated attempts and continued adaptation across tasks and environments. In simulation, AdaHVLA raises mean test success on NaVILA-LH from 22.5\% to as high as 57.5\% and improves manipulation test success across three VLA backbones by up to 30.8 percentage points over the initial harness. Real-world deployment further illustrates how the adapted policies support stable execution across task stages.
Chinese Translation
视觉-语言-动作(VLA)模型具备强大的局部控制与指令遵循能力,但在需要持续记忆与规划的长时域任务中往往表现不佳。任务框架通过保留任务历史并跟踪各执行阶段的进展,为智能体推理提供持久上下文。为了将这两种互补能力结合起来,我们提出了AdaHVLA——一种自适应任务框架,它通过机器人实践经验来精炼基于代码的协调策略,使智能体的推理与记忆更好地与VLA执行相契合。其解耦的多智能体自适应过程将证据分析、框架修订与行为评估分离到各自独立的工作上下文中,利用可检验的协调假设来指导修订,并通过后续的滚动执行来评估其预期效果。一个带状态的修订图将执行证据、假设、修订和观测效果关联起来,保留备选框架与自适应记忆,从而在多次尝试中指导精炼,并支持跨任务、跨环境的持续自适应。在仿真中,AdaHVLA将NaVILA-LH上的平均测试成功率从22.5%提升至高达57.5%,并在三种VLA骨干网络上,相比初始框架将操作测试成功率最多提高30.8个百分点。真实世界的部署进一步展示了经自适应后的策略如何在各任务阶段支持稳定的执行。
cs.RO / 45 / 2609.29212

ADM-Planner: LLM-Guided Long-Horizon Planning for Mobile Manipulators with Attention-Enhanced Dynamic Memory

ADM-Planner:基于注意力增强动态记忆的LLM引导移动机械臂长时程规划
Xiao, Jiaping, Ji, Pingyuan, Feroskhan, Mir
Abstract
Large language models can decompose mobile-manipulation goals into long action sequences, but the resulting plans remain reliable only while their world context is current. A fixed scene description becomes stale when objects are discovered, moved, or completed while retaining every observation instead produces a growing history with redundant and conflicting state. To resolve this tension, we present an LLM-guided planning framework ADM-Planner with attention-enhanced dynamic memory (ADM). Persistent workspace knowledge is separated from object-centric state, asynchronous observations and action outcomes update that state, and a bounded retriever exposes only the entries that can affect the next decision. The LLM replans when an update invalidates the remaining plan. Across 1,500 task-simulator episodes, the proposed ADM achieved 100% full-task success in the 14-container noisy dynamic setting, compared with 62% for static memory and 97% for unfiltered dynamic memory, while reducing the context-size proxy by 95.8% relative to the latter. In a six-episode live GPT-5 Mini planner, both dynamic memory variants completed every mission, while ADM reduced provider-reported input tokens by 14.4% and mean planner calls from 7.0 to 6.0. A separate 60-trial PyBullet study retained 100% success for ADM, compared with 50% for static memory. Finally, the mobile manipulator with ADM-Planner completed various missions in indoor and outdoor physical experiments while incorporating targets revealed after execution began. The results show that selective state maintenance with ADM, rather than prompt history alone, is a practical basis for long-horizon planning in changing environments. Project page: https://xjp99v5.github.io/ADM-Planner
Chinese Translation
大语言模型(LLM)能够将移动操作(mobile-manipulation)目标分解为长动作序列,但所得规划仅在环境上下文保持最新时才是可靠的。固定的场景描述会随着物体的发现、移动或任务完成而过时;而保留所有观测则会生成不断增长的、包含冗余与冲突状态的历史记录。为解决这一矛盾,我们提出了一种LLM引导的规划框架ADM-Planner,其核心是注意力增强动态记忆(Attention-enhanced Dynamic Memory, ADM)。该框架将持久的工作空间知识与以物体为中心的状态分离,由异步观测和动作结果更新该状态,并由一个有界检索器仅暴露可能影响下一步决策的条目。当某次更新使剩余计划失效时,LLM会进行重规划。在1,500个任务仿真回合中,所提出的ADM在14个容器的含噪声动态环境中实现了100%的完整任务成功率,相比之下静态记忆为62%、未过滤的动态记忆为97%,同时上下文规模代理指标相对后者降低了95.8%。在六个回合的GPT-5 Mini实机规划实验中,两种动态记忆变体均完成了全部任务,而ADM将服务商标注的输入token减少了14.4%,并将平均规划调用次数从7.0次降至6.0次。另一项包含60次试验的PyBullet研究显示,ADM保持了100%的成功率,而静态记忆仅为50%。最后,搭载ADM-Planner的移动机械臂在室内外物理实验中完成了多种任务,并能纳入执行开始后才揭示的目标。结果表明,基于ADM的选择性状态维护(而非仅依赖提示历史)是变化环境中长时程规划的一种实用基础。项目页面:https://xjp99v5.github.io/ADM-Planner
cs.RO / 46 / 2609.29261

Dense-Joint-Based Obstacle-Aided Locomotion with a Joint-Repositionable Snake Robot

基于密集关节的障碍辅助运动:一种关节可重新布置的蛇形机器人
Minomo, Kyosuke, Takahashi, Ryo, Yasui, Kotaro, Nakashima, Yasutaka, Yamamoto, Motoji, Kanada, Ayato
Abstract
Obstacle-aided locomotion is a fundamental capability for snake robots to traverse complex environments. However, conventional rigid-link snake robots often suffer from stagnation or jamming caused by their low joint density (i.e., the number of joints per unit length). This results in discontinuous contact with obstacles, unlike the continuous adaptation of biological snakes. To investigate the effect of joint density on obstacle-aided locomotion performance, we utilized a joint-repositionable snake robot mechanism that decouples actuators from joints, enabling a high-density architecture. We developed two experimental models with identical total lengths but different joint densities (high-density and low-density) and conducted comparative propulsion experiments in obstacle environments with varying obstacle diameters. The experimental results demonstrate that the high-density model substantially suppresses the abrupt shifts in reaction forces that cause stagnation in the low-density model. By maintaining smooth contact points, the high-density configuration reduces power consumption and achieves stable, continuous propulsion. These results highlight high joint density as a key factor in improving the environmental adaptability of snake robots in complex terrains.
Chinese Translation
障碍辅助运动是蛇形机器人穿越复杂环境的一项基本能力。然而,传统的刚性连杆蛇形机器人往往因其关节密度(即单位长度上的关节数)较低而容易出现停滞或卡死。这导致机器人与障碍物的接触不连续,无法像生物蛇那样实现连续的自适应调整。为了研究关节密度对障碍辅助运动性能的影响,我们采用了一种将执行器与关节解耦的关节可重新布置的蛇形机器人机构,从而实现了高密度构型。我们研制了两台总长度相同但关节密度不同(高密度与低密度)的实验样机,并在具有不同障碍物直径的障碍环境中开展了对比推进实验。实验结果表明,高密度样机能够显著抑制低密度样机中导致停滞的反作用力突变。通过保持平滑的接触点,高密度构型降低了功耗,并实现了稳定、连续的推进。这些结果凸显了高关节密度是提升蛇形机器人在复杂地形中环境适应性的关键因素。
cs.RO / 47 / 2609.29310

EgoSpeedUp: Transferring Human Manipulation Tempo to Robot Policies

EgoSpeedUp:将人类操作节奏迁移至机器人策略
Oh, Hanbit, Domae, Yukiyasu, Yagi, Takuma
Abstract
Robot manipulation policies trained through imitation learning inherit not only the demonstrated behavior but also the conservative execution tempo of robot demonstrations. Existing acceleration approaches can execute faster than the original demonstrations, but determine the appropriate acceleration primarily from robot-side information or a predefined set of tempo factors, leaving open how to obtain a task-appropriate reference for how fast each manipulation phase should progress. We introduce EgoSpeedUp, a framework that uses human manipulation as temporal supervision for robot imitation learning. Our key insight is that human demonstrations naturally reveal task-appropriate, phase-wise manipulation tempo. Given slow robot demonstrations and human demonstrations of the same task, EgoSpeedUp aligns corresponding manipulation phases, estimates their relative execution tempos from multiple human demonstrations, and transfers the resulting phase-wise tempo by retiming the robot demonstrations. The retimed demonstrations are then used for standard behavior cloning, allowing the robot to retain its executable manipulation behavior while learning to perform it at a human-informed tempo. Across two real-world manipulation tasks, EgoSpeedUp improves the task success rate by an average of 25 percentage points (pp) while reducing successful execution time by 36.5%. These results demonstrate that human manipulation tempo provides an effective temporal reference for learning faster and more reliable robot policies.
Chinese Translation
通过模仿学习训练的机器人操作策略不仅继承了演示的行为,还继承了机器人演示中保守的执行节奏。现有的加速方法可以比原始演示执行得更快,但主要是根据机器人侧信息或预定义的节奏因子集合来确定合适的加速,因而未能解决如何为每个操作阶段应推进多快获得任务适配的参考。我们提出EgoSpeedUp,一个利用人类操作作为时间监督信号的机器人模仿学习框架。我们的核心洞察是:人类演示自然地揭示了任务适配的、分阶段的操作节奏。给定同一任务的慢速机器人演示和人类演示,EgoSpeedUp对齐相应的操作阶段,从多个人类演示中估计它们的相对执行节奏,并通过重新定时机器人演示来迁移所得的分阶段节奏。重定时后的演示随后用于标准的行为克隆(behavior cloning),使机器人在保持其可执行的操作行为的同时,学会以人类参考的节奏执行该行为。在两个真实世界的操作任务上,EgoSpeedUp将任务成功率平均提升25个百分点(pp),同时将成功执行时间缩短36.5%。这些结果表明,人类操作节奏为学习更快、更可靠的机器人策略提供了有效的时间参考。
cs.RO / 48 / 2609.29340

A Simple Gripper Interface for Simulator-Agnostic Cloth Manipulation

一种面向模拟器无关布料操作(Simulator-Agnostic Cloth Manipulation)的简单夹爪接口
Nayak, Abhilash, Coltraro, Franco, Alberich-Carramiñana, Maria, Torras, Carme
Abstract
This paper presents a grasping model for cloth manipulation specifically tailored to ease the deployment of robotic control methods. The model is robust, fast and easy to implement avoiding at the same time contact and friction considerations between the gripper and the cloth in favor of simple positional constraints. The gripper is described by its pose, jaw state, and an attached grasping volume. Two kinds of grasping volumes are considered: an axis-aligned box to simulate a pinch grasping and a square pyramidal volume to simulate point grasping. When the gripper closes, the discrete cloth positions lying inside this volume are selected, stored in the local gripper frame, and then transported with the gripper motion. A simple squeezing step is also included to progressively move the selected cloth positions toward the center of the grasping region, avoiding an instantaneous displacement at closure. The model can be used in any simulator as it only requires access to discrete cloth positions and a mechanism for imposing target positions as constraints. We implement our grasping model in conjunction with a constraint-based inextensible cloth simulator, where grasping is implemented as moving positional equality constraints coupled with stretch, shear, collision, and table contact projection steps. The same gripper trajectory is applied on a robot arm to fold a real piece of cloth, serving as a simple bridge between simulation and physical cloth manipulation and showcasing the realism and practicality of our idealized grasping model.
Chinese Translation
本文提出了一种专门用于布料操作的抓取模型,旨在简化机器人控制方法的部署。该模型鲁棒、快速且易于实现,同时避免了对夹爪与布料之间接触和摩擦的复杂考虑,转而采用简单的位置约束。夹爪由其位姿、爪部状态以及一个附带的抓取体积来描述。本文考虑了两种抓取体积:用于模拟捏取抓握的轴对齐长方体,以及用于模拟点抓取的方形棱锥体积。当夹爪闭合时,位于该体积内的离散布料位置会被选中,存储在夹爪局部坐标系中,随后随夹爪运动一起移动。模型中还包含一个简单的挤压步骤,使被选中的布料位置逐步向抓取区域中心移动,从而避免闭合瞬间的位移突变。该模型可用于任何模拟器,因为它只需要访问离散布料位置以及一种施加目标位置作为约束的机制。我们将该抓取模型与基于约束的不可伸展布料模拟器相结合,其中抓取被实现为移动的位置等式约束,并与拉伸、剪切、碰撞以及桌面接触投影步骤相耦合。我们将相同的夹爪轨迹应用于机械臂上以折叠一块真实布料,从而在模拟与物理布料操作之间搭建了一座简单的桥梁,展示了他理想化抓取模型的真实性与实用性。
cs.RO / 49 / 2609.29355

C-space Analysis using Tropical Geometry

基于热带几何的构形空间分析
Nayak, Abhilash
Abstract
Configuration space~(C-space) of a mechanism is a real variety describing the set of feasible configurations that it can attain. To understand the behavior of a mechanism, it is crucial to identify and scrutinize especially the singular points of its C-space. They usually appear when the variety intersects itself, leading to different branches of motion. There exist many approaches to detect those intersections if they are transversal. However, the problem remains challenging if there are tangential, cuspidal, inter-dimensional or a combination of these intersections. This paper exploits an approach acquired from tropical geometry to analyze the neighborhood of any point on C-spaces of 1-degree-of-freedom~(\emph{dof}) mechanisms. This is done by finding the approximate rational parametrization of the curve(s) passing through the given point using Puiseux series. The proposed approach is shown to succesfully detect the transversal branchings in two foldable four bar mechanisms and a cusp in the configuration curve of the double Watt mechanism.
Chinese Translation
机构的构形空间(C-space)是一个实代数簇,描述该机构所能达到的可行构形的集合。为了理解机构的行为,识别并深入研究其构形空间的奇异点至关重要。这些奇异点通常出现在代数簇自相交处,从而导致不同的运动分支。当这些相交为横截相交时,已有多种方法可以检测它们。然而,当存在切向、尖点、跨维或这些相交形式的组合时,该问题仍然具有挑战性。本文利用一种源自热带几何的方法,来分析单自由度(dof)机构构形空间上任一点的邻域。该方法通过使用Puiseux级数求出经过给定点的曲线(多条曲线)的近似有理参数化。所提出的方法被证明能够成功检测出两个可折展四杆机构中的横截分支,以及双Watt机构构形曲线中的一个尖点。
cs.RO / 50 / 2609.29374

FMCW-LIO: A Doppler LiDAR-Inertial Odometry

FMCW-LIO:一种基于多普勒的激光雷达-惯性里程计
Zhao, Mingle, Wang, Jiahao, Gao, Tianxiao, Xu, Chengzhong, Kong, Hui
Abstract
Conventional LiDAR-inertial odometry (LIO) or simultaneous localization and mapping (SLAM) methods heavily rely on geometric features of environments, as LiDARs primarily provide range measurements instead of motion measurements. From now on, however, the situation changes thanks to the novel Frequency Modulated Continuous Wave (FMCW) Doppler LiDARs. FMCW Doppler LiDARs not only offer the point range with high resolution but also capture the instant point Doppler velocity through the Doppler effect. In the letter, we propose FMCW-LIO, a novel and robust LIO, leveraging intrinsic Doppler measurements from FMCW Doppler LiDARs. To correctly exploit Doppler velocities, a motion compensation method is designed, and a Doppler-aided observation model is applied for on-manifold state estimation. Then, dynamic points can be effectively removed by the Doppler criteria, deriving more consistent geometric observations. FMCW-LIO eventually achieves accurate state estimation and static mapping, even in structure-degenerated environments. Extensive experiments in diverse scenes are performed and FMCW-LIO outperforms other algorithms on both accuracy and robustness.
Chinese Translation
传统的激光雷达-惯性里程计(LIO)或同步定位与建图(SLAM)方法严重依赖环境的几何特征,因为激光雷达主要提供距离测量而非运动测量。然而,得益于新型调频连续波(FMCW)多普勒激光雷达的出现,这一局面从此发生了改变。FMCW多普勒激光雷达不仅能提供高分辨率的点云距离信息,还能通过多普勒效应捕获每个点的瞬时多普勒速度。在本文中,我们提出了FMCW-LIO,一种新颖且鲁棒的激光雷达-惯性里程计,充分利用了FMCW多普勒激光雷达固有的多普勒测量。为了正确利用多普勒速度,我们设计了一种运动补偿方法,并采用多普勒辅助观测模型进行流形上的状态估计。随后,通过多普勒判据可以有效剔除动态点,从而获得更加一致的几何观测。FMCW-LIO即使在结构退化环境中也能最终实现精确的状态估计和静态建图。我们在多种场景下进行了大量实验,结果表明FMCW-LIO在精度和鲁棒性方面均优于其他算法。
cs.RO / 51 / 2609.29375

Free-Init: Scan-Free, Motion-Free, and Correspondence-Free Initialization for Doppler LiDAR-Inertial Systems

Free-Init:面向多普勒激光雷达-惯性系统的免扫描、免运动、免对应关系的初始化方法
Zhao, Mingle, Wang, Jiahao, Gao, Tianxiao, Xu, Chengzhong, Kong, Hui
Abstract
Robust initialization is crucial for online systems. In the letter, a high-frequency and resilient initialization framework is designed for LiDAR-inertial systems, leveraging both inertial sensors and Doppler LiDAR. The innovative FMCW Doppler LiDAR opens up a novel avenue for robotic sensing by capturing not only point range but also Doppler velocity via the intrinsic Doppler effect. By fusing point-wise Doppler velocity with inertial measurements under non-inertial kinematics, the proposed framework, Free-Init, eliminates reliance on motion undistortion of LiDAR scans, excitation motions, and map correspondences during the initialization phase. Free-Init is also plug-and-play compatible with typical LiDAR-inertial systems and is versatile to handle a wide range of initial motions when the system starts, including stationary, dynamic, and even violent motions. The embedded Doppler-inertial velocimeter ensures fast convergence and high-frequency performance, delivering outputs exceeding 10 kHz. Comprehensive experiments on diverse platforms and across myriad motion scenes validate the framework's effectiveness. The results demonstrate the superior performance of Free-Init, highlighting the necessity of fast, resilient, and dynamic initialization for online systems.
Chinese Translation
鲁棒的初始化对于在线系统至关重要。本文针对激光雷达-惯性系统设计了一种高频且鲁棒性强的初始化框架,同时利用惯性传感器和FMCW多普勒激光雷达。创新的FMCW多普勒激光雷达通过固有的多普勒效应,不仅能够获取点云距离,还能获取多普勒速度,为机器人感知开辟了一条新途径。所提出的框架Free-Init通过在非惯性运动学下融合逐点多普勒速度与惯性测量,消除了初始化阶段对激光雷达扫描运动畸变矫正、激励运动以及地图对应关系的依赖。Free-Init还支持即插即用,可与典型的激光雷达-惯性系统兼容,并且能够灵活应对系统启动时的各种初始运动状态,包括静止、动态甚至剧烈运动。内嵌的多普勒-惯性速度计确保了快速收敛和高频性能,输出频率超过10 kHz。在多种平台和众多运动场景上开展的大量综合实验验证了该框架的有效性。结果表明Free-Init具有优越的性能,凸显了在线系统对快速、鲁棒且动态初始化的必要性。
cs.RO / 52 / 2609.29382

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

面向流匹配视觉-语言-动作模型中任务相关计算分配的解耦早退机制
Izzo, Riccardo Andrea, Rubavicius, Rimvydas, Bardaro, Gianluca, Ramamoorthy, Subramanian, Matteucci, Matteo, Suglia, Alessandro
Abstract
Flow-matching Vision-Language-Action (VLA) models have emerged as a potential solution for generalist robot control, designed by combining a pretrained Vision-Language Model (VLM) backbone with an action expert that generates continuous robot actions. While these models exhibit impressive capabilities, due to their very high number of parameters, their computational requirements are often prohibitive for robotics control. To mitigate these inefficiencies, existing methods predominantly skip VLM backbone layers with early exits or reduce denoising steps, while leaving action expert depth untouched. We propose a framework that exposes backbone depth $V$, action expert depth $A$, and denoising steps $D$ as three jointly configurable compute axes in a VLA. Starting from a pretrained VLA, we attach lightweight Exit Transformers (ET) at intermediate depths in both the backbone and the action expert, trained to distil the last layer of the policy into each exit. Furthermore, we introduce a KV Cache synthesis mechanism that manages the missing keys and values of the skipped backbone layers, allowing the action expert to exit deeper than the backbone. Finally, we show that the optimal compute budget is task-dependent, with different tasks benefiting from different axes and depths. Notably, our method does not require training the original policy from scratch, and for each exit, it increases the number of parameters by only $2.1\%$ for SmolVLA and $4.1\%$ for $\pi_{0.5}$. We validate our approach across two flow-matching VLAs (SmolVLA, $\pi_{0.5}$) and two benchmarks (LIBERO, Meta-World), revealing complementary effects: $V$ and $A$ respectively reduce FLOPs and latency, while $D$ improves both. Our joint configurations $(V,A,D)$ reduce latency by $79.2\%$ and computation (FLOPs) by $31.8\%$, while improving mean success rate by $5.6\%$.
Chinese Translation
流匹配视觉-语言-动作(VLA)模型已成为通用机器人控制的一种潜在解决方案,其设计将预训练的视觉-语言模型(VLM)骨干与生成连续机器人动作的动作专家相结合。尽管这些模型展现出令人瞩目的能力,但由于其参数量非常庞大,其计算需求对于机器人控制而言往往过于高昂。为缓解这些低效问题,现有方法主要通过早退机制跳过VLM骨干层或减少去噪步数,而对动作专家的深度保持不变。我们提出了一个框架,将骨干深度 $V$、动作专家深度 $A$ 和去噪步数 $D$ 作为VLA中三个可联合配置的计算维度。在预训练VLA的基础上,我们在骨干和动作专家的中间深度处附加轻量级的Exit Transformer(ET),通过训练将策略的最后一层蒸馏到每个退出点。此外,我们引入了一种KV缓存合成机制,用于管理被跳过的骨干层所缺失的键和值,使动作专家能够比骨干退出到更深的深度。最后,我们表明最优计算预算是任务相关的,不同任务受益于不同的维度和深度。值得注意的是,我们的方法无需从头训练原始策略,且每个退出点仅使SmolVLA的参数量增加 $2.1\%$,使 $\pi_{0.5}$ 的参数量增加 $4.1\%$。我们在两个流匹配VLA(SmolVLA、$\pi_{0.5}$)和两个基准测试(LIBERO、Meta-World)上验证了我们的方法,揭示了各维度的互补效应:$V$ 和 $A$ 分别降低FLOPs和延迟,而 $D$ 则同时改善两者。我们的联合配置 $(V,A,D)$ 将延迟降低 $79.2\%$,计算量(FLOPs)降低 $31.8\%$,同时将平均成功率提高 $5.6\%$。
cs.RO / 53 / 2609.29389

Robo-Harness K1: Harnessing Robot-Use Agents via Perception Augmentation

Robo-Harness K1:通过感知增强驾驭机器人使用智能体
Li, Zexi, Zhang, Yehang, Li, Wenqian, Huang, Haojian, Wang, Chenxu, Deng, Shiyuan, Wei, Yangkai, Zhang, Tianyi, Xie, Binghui, Zhou, Bohan, Chang, Yifan, Zhou, Kaiwen, Chen, Ying-Cong, Cheng, James, Li, Yinchuan
Abstract
Foundation vision-language models (VLMs) understand objects, instructions, and spatial relations, yet translating this capability into robotic manipulation remains difficult. Vision-language-action (VLA) models require extensive demonstrations and may compromise pretrained understanding, while direct RGB-only VLM control is costly and strongly dependent on model capability. We introduce Robo-Harness K1, a robot-use agent (RUA) framework that exposes perception as tools. The agent queries calibrated depth, persistent visual anchors, spatial measurements, and grasp hypotheses, then selects generic motions from the returned evidence. This interface makes 3D geometry accessible without changing the VLM architecture or training a depth encoder. On matched LIBERO-PRO tasks, Gemini 3.7 Flash with K1 reaches 77.8% accuracy, surpassing GPT-6 Astra's 61.1% with an RGB-only harness; K1 further improves Astra to 88.9%. Without target fine-tuning, Gemini with K1 transfers to three RoboSuite arms and dual-arm RoboTwin tasks. On RoboTwin, it achieves 32.0% on Easy and 28.0% on Hard, showing resilience to visual and environmental perturbations. K1 also produces tool-call traces aligned with next-token training. A Qwen3.5-9B student trained on only 107 teacher episodes reaches 44.2% accuracy on new initial states versus 30.2% for OpenVLA, and 13.9% on held-out task conditions versus 0.0% for OpenVLA. These results suggest that perception-augmented RUAs offer a promising route to sample-efficient, generalizable robotic policies that leverage VLM capabilities through an accessible tool interface.
Chinese Translation
基础视觉语言模型(VLM)能够理解物体、指令和空间关系,但将这种能力转化为机器人操作仍然十分困难。视觉-语言-动作(VLA)模型需要大量示范数据,且可能损害预训练获得的理解能力,而直接基于纯RGB图像的VLM控制则成本高昂,并严重依赖模型自身能力。我们提出Robo-Harness K1,这是一个将感知能力以工具形式暴露的机器人使用智能体(RUA)框架。该智能体可以查询标定后的深度信息、持久视觉锚点、空间测量值以及抓取假设,然后基于返回的证据选择通用动作。这一接口使三维几何信息变得可访问,而无需修改VLM架构或训练深度编码器。在匹配的LIBERO-PRO任务上,使用K1的Gemini 3.7 Flash达到77.8%的准确率,超过了采用纯RGB接口的GPT-6 Astra的61.1%;K1还将Astra的性能进一步提升至88.9%。在不进行目标任务微调的情况下,配备K1的Gemini可迁移至三个RoboSuite机械臂以及双臂RoboTwin任务。在RoboTwin上,它在Easy和Hard设置下分别取得32.0%和28.0%的成绩,展现出对视觉和环境扰动的鲁棒性。K1还能生成与下一词元(next-token)训练对齐的工具调用轨迹。仅使用107条教师轨迹训练的Qwen3.5-9B学生模型,在新初始状态上达到44.2%的准确率(OpenVLA为30.2%),在留出的任务条件上达到13.9%(OpenVLA为0.0%)。这些结果表明,感知增强的RUA通过一个易于访问的工具接口利用VLM能力,为构建样本高效、可泛化的机器人策略提供了一条有前景的路径。
cs.RO / 54 / 2609.29394

RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning

RACaP:将智能体推理、行动与编码作为策略的可进化机器人学习
Li, Zexi, Zhang, Yehang, Huang, Haojian, Zhou, Bohan, Li, Wenqian, Wang, Chenxu, Chang, Yifan, Wei, Yangkai, Zhang, Tianyi, Chen, Ying-Cong, Zhou, Kaiwen, Li, Yinchuan, Cheng, James
Abstract
General-purpose robot agents must learn from experience, transfer to new tasks, and act efficiently. Code as Policies (CaP) methods generate and repair programs at runtime, incurring latency and entangling reusable mechanisms with task-specific decisions. We introduce RACaP, an agentic framework that moves coding to evolution and uses a Reasoning-and-Acting (ReAct) loop to call frozen, typed Policy APIs at deployment. A two-phase strategy combines capability curriculum learning with autonomous self-evolution to improve the APIs, the ReAct harness, and experience memory. The APIs encode reusable physical mechanisms while exposing arguments for runtime adaptation. ReAct combines task-specific working memory, long-term experience memory, and visual feedback to select actions, verify outcomes, and recover from failures without modifying source code. RACaP achieves 54.4% success on LIBERO-90, 45.0% on zero-shot LIBERO-PRO, and 46.0% on LIBERO-Long, compared with at most 4.0% for CaP baselines on long-horizon tasks. On LIBERO-PRO, it achieves 2.5 times the success rate of CaP baselines and a 1.9-fold speedup in median policy time. For efficient on-robot deployment, rejection-sampled fine-tuning distills GPT-5.6 ReAct decisions into Qwen3-VL-8B-Instruct, yielding a 13.2-fold per-decision inference speedup and reducing repeated physical calls from 16 to 4. These results show that separating reusable code from runtime decisions supports continued evolution, effective transfer, and efficient long-horizon control.
Chinese Translation
通用机器人智能体必须能够从经验中学习、迁移到新任务并高效行动。Code as Policies(CaP)方法在运行时生成和修复程序,导致延迟增加,并将可复用机制与任务特定决策纠缠在一起。我们提出了 RACaP,这是一个智能体化框架,将编码工作转移到进化阶段,并在部署时使用推理与行动(Reasoning-and-Acting, ReAct)循环调用冻结的、带类型的策略 API。一种两阶段策略将能力课程学习与自主自我进化相结合,以改进 API、ReAct 框架和经验记忆。这些 API 编码了可复用的物理机制,同时暴露参数以供运行时自适应调整。ReAct 结合任务特定的工作记忆、长期经验记忆和视觉反馈来选择动作、验证结果并从失败中恢复,而无需修改源代码。RACaP 在 LIBERO-90 上达到 54.4% 的成功率,在零样本 LIBERO-PRO 上达到 45.0%,在 LIBERO-Long 上达到 46.0%,而 CaP 基线在长时程任务上的成功率最高仅为 4.0%。在 LIBERO-PRO 上,其成功率达到 CaP 基线的 2.5 倍,策略执行中位时间加速 1.9 倍。为了实现高效的机器人端部署,拒绝采样微调将 GPT-5.6 的 ReAct 决策蒸馏到 Qwen3-VL-8B-Instruct 中,实现了单次决策 13.2 倍的推理加速,并将重复的物理调用次数从 16 次减少到 4 次。这些结果表明,将可复用代码与运行时决策相分离能够支持持续进化、有效迁移和高效的长时程控制。
cs.RO / 55 / 2609.29407

WRAP: Fixtureless Wrench-aware Multi-Robot Assembly Planning

WRAP:无需夹具的力感知多机器人装配规划
Hartmann, Valentin N., Su, Huang, Huang, Yijiang, Coros, Stelian
Abstract
Assembly using robots often requires specially designed fixtures, or relies on top-down only assembly strategies. Using multiple robots, we can avoid using fixtures and make robotic assembly more flexible. Planning assembly sequences for multiple robots is challenging due to the high number of possible task assignments and orders. In addition, we need to reason over forces that occur during the assembly process, e.g., to decide if multiple robots are required for support, or if external support such as a table should be used. We present Wrap, a multi-robot assembly planner for multi-part assemblies, given the inter-part ordering-dependencies, the part meshes, and their initial state. We formulate a linear program to reason about valid grasps for supporting the forces that occur during assembly. The search leverages the assembly sequence, and greedily finds a feasible solution per assembly step by computing a heuristic via a cheap backwards search, and using the heuristic in the more expensive forward search. We then solve the multi-robot, multi-goal motion planning problem, and for execution, we split the plan into contact-rich assembly skills, and free space motion. We benchmark the planner on a variety of multi-part assemblies, and apply the planner to groups of robots differing in size and kinematics. We validate the work both in a physics simulation, and in real. Videos and code are available at https://www.vhartmann.com/wrap.
Chinese Translation
机器人装配通常需要专门设计的夹具,或仅依赖自上而下的装配策略。通过使用多个机器人,我们可以避免使用夹具,并使机器人装配更加灵活。为多个机器人规划装配序列具有挑战性,因为可能存在大量的任务分配和顺序组合。此外,我们还需要对装配过程中出现的力进行推理,例如,判断是否需要多个机器人提供支撑,或者是否应使用桌子等外部支撑。我们提出了 WRAP,一种面向多零件装配的多机器人装配规划器,其输入为零件间的顺序依赖关系、零件网格模型及其初始状态。我们构建了一个线性规划模型来推理有效的抓取方式,以支撑装配过程中出现的力。该搜索利用装配序列,通过一次低开销的反向搜索计算启发式信息,并在代价更高的正向搜索中使用该启发式信息,从而在每个装配步骤中贪婪地找到可行解。随后,我们求解多机器人、多目标的运动规划问题;在执行阶段,我们将规划结果分解为接触丰富的装配技能和自由空间运动。我们在多种多零件装配任务上对该规划器进行了基准测试,并将其应用于规模和运动学特性各异的机器人群体。我们在物理仿真和真实环境中对该工作进行了验证。视频和代码可在 https://www.vhartmann.com/wrap 获取。
cs.RO / 56 / 2609.29417

Singularity Analysis for the Perspective-Four and Five-Line Problems

透视四线与五线问题的奇异性分析
Fontán, Jorge García, Nayak, Abhilash, Briot, Sébastien, Din, Mohab Safey El
Abstract
This paper deals with image-based visual servoing and pose estimation by observing four and five lines. Our main interest is to determine the relative configurations of the camera and the observed lines that lead to problems in control and stability. Since it is equivalent to finding the singularities of the corresponding Jacobian matrix, we use tools from computational algebraic geometry to seek configurations such that all of its minors vanish simultaneously. By choosing a suitable basis for this matrix, we revisit the problem in the case of three lines to show that one type of the singularities is when the camera lies on the hyperboloid of one sheet uniquely defined by the lines. This result is further exploited to prove that the one-dimensional singularities, if any, in the case of $n$ lines appear when the camera lies on the transversals to the observed lines. Thus, by forcing the transversals to be complex, we can avoid the aforementioned type of singularities in the case of four lines although the algebra shows that there can always be up to 10 inevitable singular locations of the camera for the other type of singularity. For five lines, we find out that there are no singularities in the generic case. The singularities are also characterized for four and five lines with orthogonality and parallelism constraints. Furthermore, a visual servoing library is used to conduct some simulated experiments to substantiate the theoretical results. As expected, we observe problems in control in the vicinity of a singularity as well as increased errors in pose estimation.
Chinese Translation
本文研究基于图像的视觉伺服以及通过观测四条和五条直线进行的位姿估计。我们的主要目标是确定导致控制与稳定性问题的相机与被观测直线之间的相对位形。由于这等价于求相应雅可比矩阵的奇异性,我们利用计算代数几何的工具来寻找使该矩阵所有子式同时为零的位形。通过为该矩阵选取合适的基,我们重新审视了三直线情形下的问题,证明了一类奇异性出现在相机位于由这些直线唯一确定的单叶双曲面上时。这一结果被进一步用于证明:在n条直线的情形下,若存在一维奇异性,则当相机位于被观测直线的公截线上时出现。因此,通过强制这些截线为复截线,在四条直线的情形下我们可以避免上述类型的奇异性,尽管代数推导表明对于另一类奇异性,相机始终可能存在多达10个不可避免的奇异位置。对于五条直线,我们发现一般情况下不存在奇异性。此外,我们还对具有正交性和平行性约束的四条与五条直线的奇异性进行了刻画。最后,利用一个视觉伺服库进行了若干仿真实验,以验证理论结果。正如预期的那样,我们观察到在奇异性附近存在控制问题,同时位姿估计误差也会增大。
cs.RO / 57 / 2609.29419

UCON: Uncertainty-aware Navigation with Historical Re-association in Dynamic Environments

UCON:动态环境中基于历史重关联的不确定性感知导航
Sun, Bing, Lin, Yue, Yuan, Yongsheng, Liu, Yang, Wang, Dong, Lu, Huchuan
Abstract
Autonomous navigation in dynamic environments is hindered by two fundamental challenges: perception instability and uncertainty-optimization mismatch. The former leads to identity switches and unreliable motion estimation, while the latter prevents principled incorporation of motion uncertainty into trajectory optimization. To address these challenges, we propose UCON, an uncertainty-aware navigation algorithm in dynamic environments. For perception instability, we present a point-level historical re-association mechanism that leverages historical point cloud fragments to recover lost targets while maintaining identity continuity. Subsequently, a Kalman filter is employed to provide anisotropic motion state estimation and covariance propagation. To resolve the uncertainty-optimization mismatch, we transform predicted states and their covariances into uncertainty sectors, which are embedded as differentiable cost terms within a trajectory optimization framework. This achieves consistent uncertainty-aware dynamic obstacle avoidance while maintaining smoothness and feasibility. Extensive simulations and real-world experiments demonstrate that, while maintaining high computational efficiency, UCON achieves superior perception stability and robust navigation performance in dynamic environments compared to state-of-the-art methods. The code will be open-sourced to facilitate further research.
Chinese Translation
动态环境中的自主导航面临两个根本性挑战:感知不稳定性和不确定性-优化不匹配问题。前者导致目标身份切换和不可靠的运动估计,而后者使得运动不确定性难以被合理地纳入轨迹优化。为解决这些挑战,我们提出了UCON,一种面向动态环境的不确定性感知导航算法。针对感知不稳定问题,我们提出了一种点级历史重关联机制,利用历史点云片段恢复丢失的目标,同时保持身份连续性。随后,采用卡尔曼滤波器(Kalman filter)提供各向异性的运动状态估计和协方差传播。为解决不确定性-优化不匹配问题,我们将预测状态及其协方差转化为不确定性扇区,并作为可微代价项嵌入轨迹优化框架中。这实现了不一致性更低的不确定性感知动态避障,同时保持了轨迹的平滑性和可行性。大量的仿真和真实世界实验表明,在保持高计算效率的同时,与最先进的方法相比,UCON在动态环境中实现了更优的感知稳定性和鲁棒的导航性能。代码将开源以促进进一步研究。
cs.RO / 58 / 2609.29423

Temperament Engineering: Designing Strategic Behavioural Diversity in Robot Swarms

气质工程:面向机器人集群策略性行为多样性的设计
Hunt, Edmund R.
Abstract
No two robots are truly identical: calibration, battery state, sensor drift and wear give every swarm a distribution of behaviour rather than a single point, usually treated as an imperfection to be minimised. In animal collectives the reverse holds: consistent individual differences in behaviour ('temperament') are shaped by natural selection and often decisive for group performance. This perspective proposes 'temperament engineering', a bio-inspired framework that treats the swarm's distribution of temperaments, rather than the individual controller, as the design object. It borrows five evolutionarily validated axes of animal temperament (shyness-boldness, exploration-avoidance, activity, aggressiveness and sociability) as a design vocabulary, rendering each as a continuous control parameter $\tau \in [0,1]$ above the controller, realisable as a module threshold, a policy-conditioning vector in multi-agent reinforcement learning, or a constraint on a foundation-model planner. A three-phase workflow maps mission success criteria onto relevant axes, plans the shape of the $\tau$ distribution, and tunes reaction norms governing how temperament responds to environmental cues. The payoff is greatest under decentralisation: where a central planner can reassign behaviour online, a temperament distribution is a planner output, but in a swarm without global knowledge it must be an offline, anticipatory design input. Behavioural and platform heterogeneity are thereby co-design variables, and I sketch tentative robot-native axes (self-model plasticity, forcefulness, initiative and expressiveness) arising from features robots have and animals do not. Engineered heterogeneity has been shown to outperform homogeneous swarms in tasks such as aggregation and exploration; establishing when, and how much, heterogeneity repays its cost is the work the field can now take forward.
Chinese Translation
没有任何两台机器人是完全相同的:标定差异、电池状态、传感器漂移与磨损使每个集群的行为呈现为一种分布,而非单一固定点,这通常被视为应尽量消除的缺陷。然而在动物群体中情况恰好相反:个体间一致的行为差异(即'气质')由自然选择塑造,且往往对群体表现具有决定性作用。本视角文章提出'气质工程'(temperament engineering),一种以集群的气质分布而非个体控制器作为设计对象的仿生框架。该框架借鉴五个经过进化验证的动物气质维度(害羞—大胆、探索—回避、活动性、攻击性与社交性)作为设计词汇,将每一维度转化为控制器之上的连续控制参数 $\tau \in [0,1]$,其实现形式可以是模块阈值、多智能体强化学习中的策略调节向量,或基础模型规划器上的约束。一个三阶段工作流程将任务成功准则映射到相关气质维度上,规划 $\tau$ 分布的形态,并调整气质如何响应环境线索的反应规范。其收益在去中心化场景下最为显著:当中央规划器可以在线重新分配行为时,气质分布只是规划器的一种输出;而在缺乏全局知识的集群中,它必须成为离线的、前瞻性的设计输入。由此,行为异质性与平台异质性成为协同设计变量。此外,我初步勾勒了机器人特有的气质维度(自我模型可塑性、强制性、主动性与表达性),它们源于机器人具备而动物不具备的特征。已有研究表明,工程化的异质性在聚合与探索等任务中优于同质集群;确定异质性在何时以及以何种程度能够补偿其代价,正是该领域接下来可以推进的工作。
cs.RO / 59 / 2609.29424

Coupled State-Space Modelling, Control, and Policy Distillation for Hybrid Rigid-Pneumatic Manipulators

混合刚性-气动机械臂的耦合状态空间建模、控制与策略蒸馏
Samuel, Alan Royce Gabriel, Verma, Pulkit
Abstract
Hybrid manipulators combine motorized rigid joints with pressure-actuated origami segments. Published arms of this kind are controlled with decoupled per-DOF loops, and the cost of this approximation has not been quantified, because the coupled model needed to measure it has not been built. This paper derives such a model for a chain of $N$ alternating revolute joints and Kresling origami segments, including pneumatic chamber dynamics and crease hysteresis. Using the model, we measure the coupling directly and show that its strength varies joint by joint, and that decoupled control loses precisely on the strongly coupled joints while remaining competitive on the one nearly decoupled joint. Coupled model-based controllers track $2.5\times$ tighter than a decoupled PID baseline at lower torque. However, the model predictive controller (MPC) is too slow for real time, and model-free reinforcement learning stalls far below acceptable success rates on a strict settling metric. We therefore distill the MPC into a small neural policy with behavior cloning and DAgger. The distilled policy settles 93-94$\%$ of goals with zero collisions, within a few points of its teacher, and runs inside the 5 ms control step where the MPC does not. Where the teacher itself fails, we trace the failure to a limit cycle with the bellows' lightly damped mode, and we remove it by selecting goal postures holdable at low pressure.
Chinese Translation
混合机械臂将电机驱动的刚性关节与气压驱动的折纸单元相结合。此类已发表的机械臂均采用各自由度解耦的控制回路进行控制,但由于尚未建立能够衡量该近似代价的耦合模型,这一近似的代价一直未被量化。本文针对由 $N$ 个交替排列的转动关节与 Kresling 折纸单元构成的串联结构推导了此类模型,其中包含气腔动力学与折痕迟滞特性。利用该模型,我们直接测量了耦合强度,发现其大小逐关节变化,且解耦控制恰恰在强耦合关节上表现不佳,而在一个近似解耦的关节上仍具竞争力。基于耦合模型的控制器在更低力矩下实现了比解耦 PID 基线紧凑 $2.5\times$ 的轨迹跟踪。然而,模型预测控制器(MPC)的速度无法满足实时要求,而无模型强化学习在严格的整定指标下,成功率远低于可接受水平。因此,我们通过行为克隆与 DAgger 将 MPC 蒸馏为一个小型神经策略。蒸馏后的策略以零碰撞完成 93-94$\%$ 的目标整定,与教师模型仅相差几个百分点,并且能够在 5 ms 的控制步长内运行,而 MPC 则无法做到。对于教师模型本身失败的情形,我们将失败原因追溯至波纹管轻度阻尼模态导致的极限环,并通过选择可在低气压下保持的目标姿态来消除该问题。
cs.RO / 60 / 2609.29490

RoboLDA: A Probabilistic Generative Model for Uncovering Embodied Hierarchical Structures in Voxel-based Soft Robots

RoboLDA:一种用于揭示基于体素的软体机器人中具身层次结构的概率生成模型
Song, Junru, Yang, Yang, Shi, Jingdan, Li, Guozhen, Zhou, Weien, Wen, Ying, Wang, Feifei, Yao, Wen, Jiang, Tingsong
Abstract
Recent advances in robotics highlight hierarchical configurations of robot morphology, where multiple levels of functional substructures synergize to facilitate intelligent behaviors. This hierarchical perspective, while particularly advantageous for voxel-based soft robots (VSRs) to ease design and control complexities, is hindered by its heavy reliance on domain expertise. In this work, we address the following question: can we derive such hierarchical design principles solely from existing successful designs? We answer affirmatively by presenting RoboLDA, a Bayesian probabilistic model that decomposes VSR morphology generation into a four-level hierarchy: "task-robot-organ-voxel", and is trained via variational inference. Through extensive experiments on simulated VSRs, we verify the presence of consistent, intuitive hierarchical patterns underlying high-performing VSR designs and showcase RoboLDA's proficiency to extract and leverage these hierarchical priors for zero-shot robot design in unseen tasks. The generated designs, even without further optimization, achieve on average 106.4% of the optimized performance produced by evolutionary algorithms. Additionally, the organ structures inferred by RoboLDA serve as valid functional substructures, significantly enhancing synergistic motion control when integrated with modular control policies. Our work pioneers hierarchical generative modeling of robot morphology, offering a promising pathway towards more interpretable and generalizable development of embodied agents.
Chinese Translation
机器人学领域的最新进展凸显了机器人形态的层次化配置,其中多个层级的功能子结构协同作用以促进智能行为。这种层次化视角对于基于体素的软体机器人(Voxel-based Soft Robots, VSRs)而言,尤其有利于降低设计与控制的复杂性,但其严重依赖于领域专业知识,因而受到限制。在本工作中,我们探讨以下问题:能否仅从现有成功设计中推导出此类层次化设计原则?我们通过提出RoboLDA对此给出了肯定答案。RoboLDA是一个贝叶斯概率模型,它将VSR形态的生成分解为“任务-机器人-器官-体素”的四层层次结构,并通过变分推断进行训练。通过在仿真VSR上的大量实验,我们验证了高性能VSR设计中存在一致的、直观的层次化模式,并展示了RoboLDA在提取和利用这些层次化先验以在未见任务中进行零样本机器人设计方面的出色能力。所生成的设计即使未经进一步优化,也平均达到进化算法优化性能的106.4%。此外,RoboLDA推断出的器官结构可作为有效的功能子结构,在与模块化控制策略结合时显著增强协同运动控制。我们的工作开创了机器人形态的层次化生成建模,为具身智能体更具可解释性和可泛化性的发展提供了有前景的路径。
cs.RO / 61 / 2609.29491

Generative Evolutionary Design of Voxel-Based Soft Robots with Provable Optimality

具有可证明最优性的基于体素的软机器人生成式演化设计
Song, Junru, Xiao, Huan, Yang, Yang, Li, Guozhen, Peng, Wei, Zhang, Xiaoya, Jiang, Tingsong, Zhou, Weien, Wen, Ying, Wang, Feifei, Yao, Wen
Abstract
Voxel-based soft robots (VSRs) present a promising avenue for developing artificial organisms with lifelike intelligence. However, the vast design spaces and expensive evaluations substantially challenge their design optimization. Here we develop MISCO, a novel evolutionary framework empowered by deep generative models to optimize VSR designs with theoretical guarantees. MISCO integrates an estimation-of-distribution algorithm with a meticulously designed variational autoencoder featuring multi-task learning, position awareness, and inter-voxel signaling. These key components enhance the representational capacity of VSR morphologies and facilitate highly efficient sampling and optimization of morphological distributions. We provide theoretical guarantees for MISCO's asymptotic convergence to globally optimal designs, alongside a favorable convergence rate. Extensive simulated experiments further demonstrate MISCO's exceptional effectiveness in navigating vast design spaces, evolving high-performing VSRs for diverse tasks while flexibly balancing optimization efficiency and morphological diversity. Being validated both empirically and theoretically, MISCO represents a step change towards more scalable and reliable soft robot development.
Chinese Translation
基于体素的软机器人(Voxel-based Soft Robots, VSRs)为开发具有类生命智能的人工生物提供了一条富有前景的途径。然而,庞大的设计空间和昂贵的评估成本极大地挑战了其设计优化。本文提出了MISCO,一种由深度生成模型赋能的新型演化框架,可为VSR设计优化提供理论保证。MISCO将分布估计算法与一个精心设计的变分自编码器相结合,该变分自编码器具备多任务学习、位置感知和体素间信号传递等特性。这些关键组件增强了VSR形态的表征能力,并促进了形态分布的高效采样与优化。我们为MISCO渐近收敛至全局最优设计提供了理论保证,并给出了良好的收敛速率。大量仿真实验进一步证明,MISCO在探索庞大设计空间方面具有卓越的有效性,能够为多样化任务演化出高性能的VSR,同时灵活地平衡优化效率与形态多样性。经过实证与理论的双重验证,MISCO代表了向更可扩展、更可靠的软机器人开发迈出的重要一步。
cs.RO / 62 / 2609.29601

Free the Language Model From the Vision Encoder: Semantic Serialization as a Perception Interface for Small Language Models

将语言模型从视觉编码器中解放出来:语义序列化作为小语言模型的感知接口
Xu, Cong, Sankar, Ravi
Abstract
End-to-end vision-language models (VLMs) bind visual competence to the scale of their language model: as the language model shrinks, perception and reasoning degrade together. We study an embodied scene question-answering (QA) interface in which vision never enters the language model. A frozen perception stack detects and ranges objects; a deterministic semantic serializer compiles the perceived state, errors included, into decision-aligned text; an unmodified text-only large language model (LLM) answers. On a visible-scope-matched, occlusion-audited campus-robot benchmark, under a prospectively frozen criterion, the serialized interface, using detectors fine-tuned in-domain within each fold, outperforms a zero-shot VLM whose language model has the same 7B scale (0.7892 vs 0.7462), with a larger margin at 3B (0.7673 vs 0.6913). Preregistered decoupling experiments show the gain survives paraphrase, attributing it to decision-aligned computation rather than answer-string leakage, while novel judgment vocabularies bound its scope. The advantage grows as the reader shrinks to 1.5B and reverses at 0.5B, and a ground-truth oracle locates the reader-capability floor. Under matched task supervision the interfaces converge: a VLM fine-tuned with low-rank adaptation (LoRA) overtakes the zero-shot system but only ties an equally supervised text reader (0.8441 vs 0.8396, no statistically resolved difference), and both routes remain perception-bound. Reported perception parameters are comparable to those of the VLM's vision tower, and total compute is not smaller.
Chinese Translation
端到端视觉-语言模型(VLM)将视觉能力与其语言模型的规模绑定在一起:随着语言模型缩小,感知与推理能力一同退化。我们研究了一种具身场景问答(QA)接口,其中视觉信息从不直接进入语言模型。一个冻结的感知模块栈负责检测并测距物体;一个确定性的语义序列化器将感知到的状态(包括错误)编译为与决策对齐的文本;一个未经修改的纯文本大语言模型(LLM)进行回答。在一个可见范围匹配、遮挡审计过的校园机器人基准上,在预先冻结的评判标准下,该序列化接口(使用在每个折内于领域内微调的检测器)优于语言模型同为7B规模的零样本VLM(0.7892 对 0.7462),在3B规模下优势更大(0.7673 对 0.6913)。预注册的解耦实验表明,该优势在改写后依然存在,说明其来源于与决策对齐的计算,而非答案字符串的泄露;同时,新的判断词汇表界定了其适用范围。当阅读器缩小至1.5B时,该优势进一步增大,而在0.5B时发生逆转;真值oracle(基准上界实验)定位了阅读器能力的下限。在匹配的任务监督下,两种接口趋于收敛:经低秩适应(LoRA)微调的VLM超越了零样本系统,但仅与受到同等监督的文本阅读器持平(0.8441 对 0.8396,无统计学显著差异),且两条路线仍受感知能力制约。所报告的感知参数与该VLM的视觉塔相当,总计算量也并不更小。
cs.RO / 63 / 2609.29644

Markerless Multi-Modal Autonomous Robotic Inspection of Large Space Structures

大型空间结构的无标记多模态自主机器人检测
Alfaro, Juan De Dios, Ríos, Arturo, Rodríguez-Martínez, David, Pérez-del-Pulgar, Carlos
Abstract
Future orbital infrastructures, such as deployable antennas, solar farms, and large orbital platforms will require autonomous inspection systems able to operate with limited prior knowledge and without cooperative markers. Current on-orbit servicing approaches often rely on predefined trajectories, standard interfaces, fiducial markers or accurate target models, which limits scalability for large, heterogeneous or partially unknown structures. This paper presents a markerless autonomous robotic inspection pipeline in which 3D reconstruction is used as an inspection-support representation. The system integrates a Kinova Gen2 manipulator with an end-effector-mounted multimodal sensor head composed of an RGB-D camera, a thermal camera and a 2D LiDAR. The pipeline estimates an approximate inspection volume, generates viewpoints, plans collision-free motions with MoveIt, and synchronously records RGB-D images, thermal data, and robot poses in ROS2. Candidate reconstruction methods were evaluated to select a practical method for this pipeline, with Nerfacto used for geometric reconstruction and Thermal-Nerfacto used to demonstrate thermal-aware rendering for inspection. Validation in a Gazebo-based simulator and preliminary laboratory tests reveal that the proposed system can autonomously acquire spatially coherent inspection data and produce reconstructions suitable for visual and geometric assessment, representing a step towards inspection of large non-cooperative space structures.
Chinese Translation
未来的轨道基础设施,如可展开天线、太阳能电站和大型轨道平台,将需要能够在先验知识有限且无协作标记的条件下运行的自主检测系统。当前的在轨服务方法通常依赖于预定义轨迹、标准接口、基准标记或精确的目标模型,这限制了对大型、异构或部分未知结构的可扩展性。本文提出了一种无标记自主机器人检测流程,其中三维重建被用作检测辅助表示。该系统集成了Kinova Gen2机械臂和一个安装在末端执行器上的多模态传感头,该传感头由RGB-D相机、热成像相机和2D激光雷达组成。该流程估算近似检测体积、生成视点、使用MoveIt规划无碰撞运动,并在ROS2中同步记录RGB-D图像、热数据和机器人位姿。我们对候选重建方法进行了评估,以选出适用于该流程的实用方法,其中使用Nerfacto进行几何重建,并使用Thermal-Nerfacto展示面向检测的热感知渲染。基于Gazebo仿真器的验证和初步实验室测试表明,所提出的系统能够自主获取空间上连贯的检测数据,并生成适用于视觉和几何评估的重建结果,是向大型非协作空间结构检测迈出的一步。
cs.RO / 64 / 2609.29669

Do World Models Make Better Robots? A Survey of Evaluation Benchmarks for Predictive Embodied Intelligence

世界模型能否造就更好的机器人?预测性具身智能评估基准综述
Jena, Gaytri, Wanaskar, Kapil, Jain, Vinija, Chadha, Aman, Sharma, Vasu, Das, Amitava
Abstract
Robot learning now advances along two tracks that rarely meet. On one side, direct Vision-Language-Action (VLA) policies map observations to actions and are scored by closed-loop task success. On the other, predictive and generative world models forecast future observations and are scored by open-loop prediction or generation quality. A natural question sits between them: does world modelling earn a measurable, closed-loop advantage over a direct policy, and for which robotic capabilities? We argue that the field cannot yet answer this question, and that the reason is a gap in how it is measured, not in the models themselves. World-model benchmarks score prediction without ever executing it, while task-success suites host a single policy and never build a world-model versus VLA contrast. This survey maps the evaluation landscape around that gap. We catalogue 160 web-verified benchmarks spanning 2017 to 2026 and organise them by evaluation mode, robotic capability, and model family into four lanes: policy suites, embodied agents, world model evaluation, and prediction-to-action bridges. Across the corpus, 138 of 160 benchmarks are model-agnostic and only 11 (7%) build an explicit VLA-versus-world-model contrast; counterfactual capability is almost entirely unmeasured, and only four benchmarks turn prediction into executed action. We contribute an operational taxonomy, a coverage comparison against the eight closest surveys (ours is the only one to cross capability with model family), an evaluation loop that isolates the advantage of prediction, and an actionable protocol of four advantage-aware metrics anchored on named testbeds. The organising claim is not that world models help or do not help, but that answering the question requires benchmarks built to ask it.
Chinese Translation
机器人学习目前沿着两条很少交汇的路线发展。一方面,直接的视觉-语言-动作(Vision-Language-Action, VLA)策略将观测映射为动作,并以闭环任务成功率来评价;另一方面,预测性与生成式世界模型预测未来观测,并以开环预测或生成质量来评价。两者之间存在一个自然的问题:世界建模能否相比直接策略带来可测量的闭环优势?如果能,是在哪些机器人能力上?我们指出,该领域目前尚无法回答这一问题,其原因不在于模型本身,而在于评估方式上的缺口:世界模型基准只评分预测而从不执行预测,而任务成功评测套件只承载单一策略,从不构建世界模型与VLA的对比。本综述围绕这一缺口绘制了评估领域的全貌。我们整理了160个经网络核实的基准(时间跨度为2017年至2026年),并按评估模式、机器人能力和模型家族将其归入四个类别:策略套件、具身智能体、世界模型评估以及从预测到动作的桥梁。在整个语料库中,160个基准里有138个与模型无关,仅有11个(7%)构建了显式的VLA与世界模型对比;反事实能力几乎完全未被测量,只有四个基准将预测转化为可执行的动作。我们的贡献包括:一个可操作的分类体系、与八篇最接近综述的覆盖面对比(我们是唯一同时从能力与模型家族两个维度交叉分析的综述)、一个能够隔离预测优势的评估循环,以及一个锚定于具体测试平台、包含四个优势感知指标的可操作协议。本文的核心论点并非世界模型有用或无用,而是要回答这一问题,必须构建专门为此设计的基准。
cs.RO / 65 / 2609.29738

Combining Evasive and Braking Reactions for Safety Reference Models in Automated Vehicles

结合闪避与制动反应的自动驾驶车辆安全参考模型
Donà, Riccardo, Mattas, Konstantinos, Ciuffo, Biagio
Abstract
Computational models of careful and competent human drivers are essential for scenario-based evaluation of automated driving systems (ADS). However, most existing safety reference models primarily focus on longitudinal braking, neglecting the role of evasive steering in human collision avoidance. This paper proposes a hybrid Fuzzy-Safety Model (FSM-H) that integrates longitudinal mitigation and lateral avoidance within a unified behavioral framework. The braking component is governed by Proactive Fuzzy Safety (PFS) metrics, representing the erosion of longitudinal safety margins, while the steering component is driven by Criticality Fuzzy Safety for lane-change (CFS-LC), capturing lateral conflict severity and maneuver feasibility. A finite-state architecture models the sequential escalation from nominal driving to braking and, when necessary, to evasive steering, incorporating perception-reaction time and lane-check delays to reflect human decision processes. The model is evaluated in reconstructed high-criticality cut-in scenarios and compared with braking-only and steering-only reference strategies. Results show that the hybrid approach expands the preventability envelope while maintaining behavioral plausibility and computational tractability. The proposed framework provides a transparent and explainable human reference model suitable for simulation-based ADS safety benchmarking and regulatory assessment.
Chinese Translation
谨慎且有能力的人类驾驶员的计算模型对于自动驾驶系统(ADS)的基于场景的评估至关重要。然而,现有的大多数安全参考模型主要关注纵向制动,忽视了闪避转向在人类避撞中的作用。本文提出了一种混合模糊安全模型(FSM-H),在统一的行为框架中整合了纵向缓解与横向避让。其中,制动部分由主动模糊安全(PFS)指标驱动,表征纵向安全裕度的侵蚀;转向部分则由用于换道的临界性模糊安全(CFS-LC)指标驱动,捕捉横向冲突严重程度与机动可行性。该模型采用有限状态架构,模拟从正常行驶到制动、并在必要时进一步到闪避转向的逐步升级过程,并引入感知-反应时间和换道检查延迟以反映人类的决策过程。该模型在重构的高临界性切入场景中进行评估,并与仅制动和仅转向的参考策略进行了比较。结果表明,这种混合方法在保持行为合理性与计算可处理性的同时,扩大了可预防性包络。所提出的框架提供了一个透明且可解释的人类参考模型,适用于基于仿真的自动驾驶系统安全基准测试与法规评估。
cs.RO / 66 / 2609.29750

From Target Selection to Digging: A Learning-Based Framework for Continuous Autonomous Excavation

从目标选择到挖掘:基于学习的连续自主挖掘框架
Zhao, Shuai, Pan, Ji-an, Yang, Quantao, Wang, Zheng, Chen, Chaoyi, Xu, Qing, Li, Keqiang
Abstract
Repeated excavation continuously reshapes pile geometry, requiring an autonomous excavator to adapt its digging targets and coordinate motion across successive excavation cycles. We present a learning-based framework for continuous autonomous excavation that integrates terrain-aware target selection with reinforcement- and imitation-learning controllers. The framework separates target-conditioned motion from local digging: a shared task-conditioned RL policy controls waypoint-guided approach and loaded transport, while an IL policy learns vision-based digging and lifting from expert demonstrations. Digging targets are selected from LiDAR elevation maps and converted into bucket-tip waypoints for motion control. The control architecture coordinates the learned policies and deterministic unloading through a shared motion interface. The complete system is deployed on a scaled hydraulic excavator with multimodal sensing and closed-loop actuator control. Offline replay and physical experiments demonstrate more consistent target selection, shorter local motion time, and increased payload compared with the respective baselines. The learned digging policy achieves a mean payload of 6.52 kg per completed cycle, compared with 2.68 kg for Fixed Dig. Three five-scoop runs further demonstrate consecutive autonomous excavation under continuously changing pile geometry.
Chinese Translation
重复挖掘会持续改变料堆的几何形状,这要求自主挖掘机能够调整其挖掘目标,并在连续的挖掘周期之间协调运动。我们提出了一种基于学习的连续自主挖掘框架,该框架将地形感知的目标选择与强化学习和模仿学习控制器相集成。该框架将目标条件下的运动与局部挖掘相分离:一个共享的任务条件强化学习(RL)策略控制由路径点引导的接近和负载运输,而一个模仿学习(IL)策略通过专家演示学习基于视觉的挖掘和提升动作。挖掘目标从LiDAR高程图中选取,并转换为铲斗尖端路径点用于运动控制。该控制架构通过共享的运动接口协调所学习的策略与确定性的卸载动作。完整系统部署在一台具备多模态传感和执行器闭环控制的缩比液压挖掘机上。离线回放与物理实验表明,与各自的基线方法相比,本系统的目标选择更加稳定一致,局部运动时间更短,载荷更高。所学得的挖掘策略在每完成一个循环的平均载荷达到6.52公斤,而固定挖掘(Fixed Dig)方法仅为2.68公斤。三次五铲连续运行进一步展示了在料堆几何形状持续变化下的连续自主挖掘能力。
cs.RO / 67 / 2609.29760

PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation

PolyUMI:面向物体推断与操作的易用视-触-听觉数据采集平台
Hayes, Conor W., Krohn, Rickmer, Ramaswami, Aravind, Ramaswami, Anunth, Dengler, Nils, Lynch, Kevin M., Colgate, J. Edward, Chalvatzaki, Georgia, Elwin, Matthew L.
Abstract
Humans typically rely on vision, touch, hearing, and proprioception to perceive contact and adapt their actions during manipulation. Providing robots with comparable responsiveness therefore requires hardware that can retain and use these complementary sensory signals. Most imitation-learning systems, however, observe demonstrations primarily through vision and proprioception, limiting access to contact information that is difficult to infer visually. We present PolyUMI, an open-source platform for scalable visual--tactile--audio demonstration collection and robot deployment. Its lightweight, wireless handheld gripper records synchronized wrist-camera, optical tactile, contact-audio, and proprioceptive observations without requiring a tethered workstation. The same sensing finger can be transferred to the robot end effector, preserving the sensing geometry between demonstration collection and policy execution. To effectively use these heterogeneous observations, we further introduce VisTA, a token-level multimodal policy that integrates information across sensors and time to predict contact-aware robot actions. Experiments spanning object inference, slip control, and contact-rich manipulation show that touch and audio reveal task-relevant information beyond vision and that VisTA is competitive with or outperforms existing multimodal policies. Together, PolyUMI and VisTA provide an accessible pipeline for collecting multimodal demonstrations and learning policies that perceive physical interaction beyond vision. Project Page: https://polyumi-vista.github.io
Chinese Translation
人类在操作过程中通常依赖视觉、触觉、听觉和本体感觉来感知接触并调整自身动作。因此,要让机器人具备类似的响应能力,就需要硬件能够保留并利用这些互补的传感信号。然而,大多数模仿学习系统主要通过视觉和本体感觉来观测演示,这限制了对难以通过视觉推断的接触信息的获取。我们提出了 PolyUMI,一个用于可扩展的视-触-听觉演示采集与机器人部署的开源平台。其轻量化的无线手持夹爪能够同步记录腕部相机、光学触觉、接触音频和本体感觉观测,且无需连接有线工作站。同一传感手指可转移至机器人末端执行器上,从而在演示采集与策略执行之间保持一致的传感几何构型。为了有效利用这些异构观测,我们进一步提出了 VisTA,一种令牌级多模态策略,可跨传感器与时间整合信息,以预测具备接触感知能力的机器人动作。涵盖物体推断、滑动控制和接触丰富操作等任务的实验表明,触觉和音频能够揭示视觉之外的与任务相关的信息,且 VisTA 与现有多模态策略相比具有竞争力或更优表现。PolyUMI 与 VisTA 共同提供了一个易用的流水线,用于采集多模态演示并学习能够感知视觉之外物理交互的策略。项目主页:https://polyumi-vista.github.io
cs.RO / 68 / 2609.29822

Self-Supervised Anchoring of Fingertip Sensing to Proprioception and Proactive Actions for Robot Imitation Learning

面向机器人模仿学习的指尖感知与本体感觉及主动行为的自监督锚定
Motoda, Tomohiro, Murooka, Masaki, Shirai, Keisuke, Oh, Hanbit, Nakajo, Ryoichi, Miwa, Shotaro, Mykhailyshyn, Roman, Duarte, Hugo, Domae, Yukiyasu
Abstract
Robotic imitation learning often relies on external cameras, yet local interaction cues such as object proximity, contact onset, and grasp state are difficult to observe near the fingertips because of occlusion and limited temporal resolution. We study how to effectively incorporate complementary fingertip sensing into imitation learning using pressure-sensitive tactile and reflective proximity sensors, along with pretrained sensor encoders. The two modalities provide information at different manipulation phases: proximity sensing is informative before contact, whereas tactile sensing becomes informative after contact. However, naively adding these signals to a policy does not consistently improve performance and can even underperform vision-only policies, suggesting that sparse, phase-dependent sensor signals are difficult to exploit from limited demonstrations. We therefore propose a proprioception-anchored pretraining method, PROprioceptive-and-PRoactive Anchoring (PROPRA), which independently aligns each fingertip sensor history with proprioceptive and action segments. This provides a continuously available sensorimotor reference, allowing each sensor to be aligned independently during its informative phases. Experiments on real-world manipulation tasks show that our pretraining method improves average success rates over vision-only policies and image-anchored pretraining baselines. Representation analysis further shows that it preserves richer information about pre-contact states, enabling more effective use of complementary fingertip sensing. Please refer to our project page: https://tomohiromotoda.github.io/nia.propra/
Chinese Translation
机器人模仿学习通常依赖外部相机,但由于遮挡和时间分辨率有限,物体接近、接触开始和抓取状态等局部交互线索难以在指尖附近观测到。我们研究如何利用压敏触觉传感器和反射式接近传感器以及预训练的传感器编码器,将互补的指尖感知有效地融入模仿学习。这两种模态在不同操作阶段提供信息:接近传感器在接触前提供信息,而触觉传感器在接触后变得有信息量。然而,简单地将这些信号加入策略中并不能稳定地提升性能,甚至可能不如仅使用视觉的策略,这表明稀疏且依赖阶段的传感器信号难以从有限的演示数据中被充分利用。因此,我们提出了一种以本体感觉为锚点的预训练方法——本体感觉与主动行为锚定(PROprioceptive-and-PRoactive Anchoring, PROPRA),该方法独立地将每个指尖传感器历史与本体感觉和动作片段进行对齐。这提供了一个持续可用的感知运动参考,使每个传感器能够在其信息丰富的阶段被独立对齐。在真实世界操作任务上的实验表明,我们的预训练方法相比仅视觉策略和基于图像锚定的预训练基线提高了平均成功率。表征分析进一步表明,该方法保留了更丰富的接触前状态信息,从而能够更有效地利用互补的指尖感知。项目页面请参见:https://tomohiromotoda.github.io/nia.propra/
cs.RO / 69 / 2609.29850

BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video

BeyondRetarget:从单目视频直接学习可执行的人形机器人运动
Xiong, Tianyu, Lu, Yi, Wang, Jinrui, Liang, Ziqi, Lei, Dandan, Zhou, Xiaoyang, Long, Xiao-xiao, Shen, Qiu, Cao, Xun
Abstract
Learning executable motions from human videos offers a scalable solution for humanoid robots to acquire demonstration motions. However, existing pipelines typically first construct an explicit human motion representation and then convert it into robot motions via motion retargeting. Although such methods can effectively leverage large volumes of existing human data for training, the substantial differences between humans and humanoid robots in locomotion mechanisms and joint degree-of-freedom configurations make motions generated by this human-representation-centric approach difficult to execute on robots. Furthermore, errors introduced during human motion estimation inevitably propagate to the retargeting stage and cannot be eliminated via joint optimization. We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions. Discarding the explicit human representation, this framework learns robot-oriented implicit representations directly from visual observations, enabling the model to capture cross-morphology motion structures. To generate motions more suitable for robot execution, we further design a contact-aware motion optimization mechanism to improve temporal consistency and physical plausibility. Experiments show that BeyondRetarget significantly improves the accuracy and robustness of generated robot motions, while achieving higher execution success rates and lower latency in both simulation environments and real humanoid robots.
Chinese Translation
从人体视频学习可执行运动,为人形机器人获取示范运动提供了一种可扩展的解决方案。然而,现有流程通常先构建显式的人体运动表示,再通过运动重定向将其转换为机器人运动。尽管此类方法能够有效利用大量现有人体数据进行训练,但人类与人形机器人在运动机制和关节自由度配置上存在显著差异,使得这种以人体表示为中心的方法所生成的运动难以在机器人上执行。此外,人体运动估计过程中引入的误差不可避免地传播到重定向阶段,且无法通过联合优化予以消除。我们提出BeyondRetarget,一个将单目RGB视频直接映射为机器人运动的端到端框架。该框架摒弃显式人体表示,直接从视觉观测中学习面向机器人的隐式表示,使模型能够捕捉跨形态的运动结构。为生成更适合机器人执行的运动,我们进一步设计了接触感知的运动优化机制,以提升时间一致性和物理合理性。实验表明,BeyondRetarget显著提高了生成的机器人运动的准确性和鲁棒性,在仿真环境和真实人形机器人上均取得了更高的执行成功率和更低的延迟。
cs.RO / 70 / 2609.29861

GPT-6-Astra Lights Up Embodied Navigation: Evaluation in Zero-Shot Vision-and-Language Navigation in Continuous Environments

GPT-6-Astra 点亮具身导航:连续环境中零样本视觉语言导航的评估
Dai, Guangzhao, Sun, Qianru, Wu, Qi, Zhu, Bin
Abstract
We investigate whether GPT-6-Astra, a general-purpose foundation model, can navigate unfamiliar environments using its own perception, reasoning, and decision-making capabilities. Our evaluation focues on zero-shot vision-and-language navigation in continuous environments (VLN-CE) through a minimal interface in the Codex harness, aiming to unleash GPT-6-Astra's full potential for navigation. Using monocular RGB, GPT-6-Astra decides when to observe, how to move, and when to stop, without navigation-specific fine-tuning, a trained waypoint predictor, or a pre-built scene map. Our evaluation yields four main findings. First, \textbf{\textit{GPT-6-Astra achieves strong zero-shot navigation performance using only monocular RGB observations}}. On the common-adopted zero-shot R2R-CE benchmark, ultra reasoning achieves a success rate of \textbf{\textit{79.0\%}}, exceeding the strongest reported zero-shot and supervised success rates by \textbf{\textit{13.0}} and \textbf{\textit{6.9}} percentage points, respectively. Second, \textbf{\textit{GPT-6-Astra advances multi-stage language instructions into coherent, adaptive navigation}} by grounding spatial relations, tracking task progress, and revising its actions. Third, \textbf{\textit{reliable route execution and goal verification remain challenging, even with ultra reasoning}}. Plausible local landmark matches do not consistently lead to correct task completion. Fourth, \textbf{\textit{these capabilities motivate rethinking the role of embodied learning}}. Future VLN research should build on foundation models to advance generalizable and reliable embodied intelligence.
Chinese Translation
我们研究了通用基础模型 GPT-6-Astra 能否凭借自身的感知、推理与决策能力在陌生环境中进行导航。我们的评估聚焦于通过 Codex 框架中的极简接口实现连续环境中的零样本视觉语言导航(VLN-CE),旨在充分释放 GPT-6-Astra 的导航潜力。仅使用单目 RGB,GPT-6-Astra 即可自主决定何时观察、如何移动以及何时停止,无需导航专用的微调、训练好的路径点预测器或预先构建的场景地图。我们的评估得出四项主要发现。第一,GPT-6-Astra 仅依靠单目 RGB 观测即可实现强大的零样本导航性能。在广泛采用的零样本 R2R-CE 基准上,超推理(ultra reasoning)达到了 79.0% 的成功率,分别超过已报道的最强零样本和有监督成功率 13.0 和 6.9 个百分点。第二,GPT-6-Astra 通过将空间关系落地、跟踪任务进度并修正自身动作,将多阶段语言指令转化为连贯且自适应的导航。第三,即使借助超推理,可靠的路线执行与目标验证仍然具有挑战性:看似合理的局部地标匹配并不总能带来正确的任务完成。第四,这些能力促使我们重新思考具身学习的作用。未来的 VLN 研究应基于基础模型,推动可实现泛化且可靠的具身智能的发展。
cs.RO / 71 / 2609.29908

MorphIK: Morphology-Conditioned Neural Inverse Kinematics for Unknown Robots

MorphIK:面向未知机器人的形态条件化神经逆运动学
Clasmeier, Lennart, Habekost, Jan Gerrit, Weber, Cornelius, Wermter, Stefan
Abstract
Neural models can learn to generate various solutions to the inverse kinematics problem from data, but are usually limited to a single robot. We present MorphIK, a flow-matching model that solves inverse kinematics for revolute-joint-based kinematic chains it has never seen during training. The model uses a transformer architecture to encode the robot's morphology along with the target pose. This encoding then conditions a flow-matching head that generates poses from noise. Trained on purely synthetic data from procedurally generated robots, the model reaches a precision of about 5 cm on unseen real-world robots with 6 to 9 Degrees of Freedom. For higher precision, the model serves as an excellent Prior for further optimization algorithms, reducing error to less than 1 cm after a single step of Damped Least Squares optimization and to sub-1 mm error after 3 steps in most cases. Building on flow matching's generative capabilities to produce highly diverse outputs, our model can efficiently sample the robot's null space, providing a wide variety of configurations for the same pose. Thus, overall, MorphIK allows learning and generalizing neural inverse kinematics for a multitude of known and unknown robots.
Chinese Translation
神经模型可以从数据中学习生成逆运动学问题的多种解,但通常仅限于单一机器人。我们提出了MorphIK,这是一种流匹配(flow-matching)模型,能够为其在训练期间从未见过的基于旋转关节的运动学链求解逆运动学。该模型使用Transformer架构对机器人的形态以及目标位姿进行编码,该编码随后条件化一个流匹配头,从噪声中生成位姿。该模型仅使用程序化生成机器人的纯合成数据进行训练,在未见过的6至9自由度真实世界机器人上达到了约5厘米的精度。为实现更高精度,该模型可作为进一步优化算法的优良先验(Prior):经过一步阻尼最小二乘(Damped Least Squares)优化后,误差降至1厘米以内;在大多数情况下,经过3步优化后误差低于1毫米。借助流匹配生成高度多样化输出的能力,我们的模型能够高效地对机器人的零空间进行采样,为同一位姿提供多种不同的构型。因此,总体而言,MorphIK能够为大量已知和未知机器人学习并泛化神经逆运动学。
cs.RO / 72 / 2609.29914

High-Voltage Optocoupler Amplifier for Electrostatic Actuators

用于静电驱动器的高压光耦合器放大器
Jurgiel, George C., Miller, Alex S., Lang, Jeffrey H.
Abstract
Many electrostatic actuators require multi-kilovolt drive voltages at sub-milliamp currents, a task poorly suited for conventional switching devices. As an alternative, we demonstrate a high-voltage amplifier using optocouplers as active elements. The amplifier produces a 20-kV peak-to-peak output with up to 500 Hz bandwidth while maintaining a minimal component count. By using optocouplers as linear devices in feedback, lower harmonic distortion and higher bandwidth are achieved than offered by equivalent PWM amplifiers. This design improves the viability of electrostatic actuators by providing a simpler method to achieve useful drive waveforms.
Chinese Translation
许多静电驱动器需要在亚毫安电流下提供数千伏特的驱动电压,这一任务并不适合传统的开关器件。作为替代方案,我们展示了一种使用光耦合器(optocouplers)作为有源元件的高压放大器。该放大器可产生峰峰值为20千伏的输出,带宽高达500 Hz,同时保持极少的元件数量。通过将光耦合器作为反馈中的线性器件使用,与同等PWM放大器相比,实现了更低的谐波失真和更高的带宽。该设计提供了一种更简单的方法来获得实用的驱动波形,从而提升了静电驱动器的可行性。
cs.RO / 73 / 2609.29929

Pairwise Approximation Can Select the Wrong Multi-Robot Plan

成对近似可能导致错误的多机器人规划选择
Teo, William
Abstract
Multi-robot coordination methods often score a joint plan from singleton and pairwise terms, leaving out the terms that involve three or more robots. We measure the plan-selection regret of two pairwise approximations to delivered coverage using frozen multi-robot trajectories. For each four-robot plan on an indoor exploration benchmark, replaying all 16 robot subsets gives the exact delivered-coverage set function $F$. From the same subset values we compute two pairwise scores: the exact order-2 M\"obius truncation $F_2$, which depends only on the singleton and pair values, and an equal-weight least-squares two-additive fit $G$. Ranking by $F_2$ instead of $F$ changes the selected plan on six of seven maps at the 15 m candidate-generation range in each of two candidate families, with regret up to 0.337 of map coverage. Switching to $G$ reduces the regret but still changes the selection on three of seven maps in each family. The additive score $F_1$, which keeps only the singleton terms, selects the exact winner on six of seven maps in one family and four of seven in the other, against one of seven for $F_2$. We also find that lower average reconstruction error does not guarantee lower selection regret.
Chinese Translation
多机器人协同方法通常基于单机器人项和成对项对联合规划进行评分,而忽略了涉及三个或更多机器人的项。我们使用冻结的多机器人轨迹,测量了两种成对近似在交付覆盖度(delivered coverage)上的规划选择遗憾(regret)。在一个室内探索基准测试中,针对每个四机器人规划,通过重放全部16个机器人子集,可以得到精确的交付覆盖度集合函数 $F$。基于相同的子集值,我们计算了两种成对评分:一是仅依赖单机器人值和成对值的精确二阶 M"obius 截断 $F_2$,二是等权重最小二乘二加性拟合 $G$。在两个候选规划族中,在15米的候选生成范围内,使用 $F_2$ 而非 $F$ 进行排序,在七张地图中的六张上改变了所选规划,遗憾最高达地图覆盖度的0.337。改用 $G$ 可降低遗憾,但在每个候选族的七张地图中仍有三张改变了规划选择。仅保留单机器人项的加性评分 $F_1$,在一个候选族的七张地图中的六张、另一个候选族的七张中的四张上选出了精确的最优规划,而 $F_2$ 仅在七张中的一张上正确。我们还发现,较低的平均重构误差并不保证较低的选择遗憾。
cs.RO / 74 / 2609.29964

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

世界动作智能体(World Action Agent):通过世界动作演练利用视觉语言模型实现机器人操作
Zhang, Yehang, Huang, Haojian, Chang, Yifan, Su, Jianchong, Zhou, Bohan, Xu, Yingjie, Chen, Wosong, Zhou, Tianhao, Wang, Chenxu, Zhang, Tianyi, Wei, Yangkai, Li, Wenqian, Deng, Shiyuan, Li, Yinchuan, Chen, Ying-Cong, Li, Zexi
Abstract
General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected automatically from the scene geometry, present the scene around the current interaction. Action rehearsal turns each action into an editable proposal that the agent, alone or through an Imagination Agent, previews and revises against planning feedback before execution. In-view correction closes the loop between observation, rehearsal, and low-level execution, letting the agent remove residual offsets in the view where it observes them. Through the same workspace, WAA acquires embodied procedural knowledge in two ways: it evolves multimodal skills from expert videos and human teaching under evidence-based review and consults them through a Skill Agent, and its interaction traces train smaller VLMs to pilot the same harness. On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone; the same skills remain effective on robosuite without further learning. Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.
Chinese Translation
通用视觉语言模型(VLM)为机器人操作带来了广泛的知识和空间推理能力,然而现有系统要么间接使用它们(用于预测约束或编写程序),要么仅赋予它们对场景的观察视角,而非一个可供行动的世界。我们提出了世界动作智能体(World Action Agent, WAA),这是一个多智能体框架,使VLM能够借助基础工具操控机器人,并在一个可视化动作工作区中做出每一个决策。该工作区具有三个特性。接触视图(Contact views)根据场景几何自动选取,呈现当前交互点周围的场景。动作演练(Action rehearsal)将每个动作转化为可编辑的提案,智能体可以独立或通过一个想象智能体(Imagination Agent)在执行前根据规划反馈进行预览和修订。视图内校正(In-view correction)在观察、演练和底层执行之间形成闭环,使智能体能够在观察到残差偏移的视图中直接将其消除。通过同一工作区,WAA以两种方式获取具身程序性知识:其一,在基于证据的审查下,从专家视频和人类教学中演化出多模态技能,并通过技能智能体(Skill Agent)进行调用;其二,其交互轨迹可用于训练更小的VLM来驱动同一框架。在LIBERO-Pro上,仅使用从LIBERO-90演化的技能的WAA达到了75.6%的平均成功率的当前最优水平,超越了端到端VLA、代码即策略(code-as-policy)智能体以及使用相同骨干网络的视觉框架基线;同样的技能在无需进一步学习的情况下在robosuite上依然有效。在框架轨迹上微调Qwen3.5-9B后,其域外成功率从1.7%提升至43.3%。
cs.RO / 75 / 2609.30023

Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation

Res-HIL:面向样本高效的灵巧操作的人类引导残差强化学习
Iavorskaia, Mariia, Dietz, Christian, Albrecht, Sebastian, Khadiv, Majid
Abstract
Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.
Chinese Translation
模仿学习使机器人能够从演示中习得操作技能,但所得到的策略在训练数据之外可能失效,而收集更多演示又需要耗费大量人力。人在回路强化学习在在线训练过程中利用人类的纠正性反馈,但通常学习的是完整的任务策略,而非对预训练的模仿策略进行精炼。我们提出了 Res-HIL,这是一种人在回路残差强化学习框架,它在冻结的模仿策略之上学习纠正性动作。每一次人类干预提供两种互补的学习信号:对残差策略的直接监督,以及对先前自主行为的奖励塑形。Res-HIL 将这些信号与残差策略的零初始化相结合,以稳定并加速在线学习。我们在五个涵盖高精度和长时程行为的富接触操作任务上评估了 Res-HIL。仅使用 20 个初始演示,Res-HIL 在经过十分钟的在线训练后,在所有任务上都优于最先进的完整策略人在回路强化学习以及无人类引导的残差微调。Res-HIL 改进了其预训练的基础策略,并优于使用五倍多演示训练的模仿策略。消融实验表明,残差直接监督对性能至关重要,而感知干预的奖励塑形显著提升了训练效率。
cs.RO / 76 / 2609.30024

Body-Grounded Replanning for Physically Adaptive Manipulation

基于身体状态的重规划以实现物理自适应操作
Saito, Namiko, Kera, Hiroshi
Abstract
Manipulation requires not only reasoning about the external environment, but also about the robot's physical condition. A strategy may remain geometrically feasible while becoming physically unsuitable due to increased joint load or limited mobility, yet internal physical state is typically used only for low-level control. We propose body-grounded high-level replanning, which uses internal physical state to adapt manipulation strategies during execution. Body-state events trigger strategy replanning, and an LLM interprets the underlying joint-level state, recent execution statistics, and execution history to select a context-dependent alternative, while leaving the task objective and low-level controller unchanged. We evaluate the framework on a reaching task under controlled load and asymmetric mobility constraints in simulation and on a real robot. Our experiments show that body-grounded replanning maintains high task success while reducing physical effort and enabling more efficient strategy adaptation. Additional contact-rich manipulation experiments demonstrate the applicability of the same replanning interface beyond reaching. These results show that internal physical state can inform not only low-level control, but also high-level decisions about how a manipulation task should be performed.
Chinese Translation
机器人操作不仅需要对外部环境进行推理,还需要对机器人自身的物理状态进行推理。一种策略可能在几何上仍然可行,但由于关节负载增加或活动能力受限而在物理上变得不再适用,然而机器人内部物理状态通常仅被用于底层控制。我们提出基于身体状态的高层重规划(body-grounded high-level replanning),利用内部物理状态在执行过程中自适应地调整操作策略。身体状态事件触发策略重规划,大语言模型(LLM)通过解读底层关节状态、近期执行统计信息和执行历史,选择一种依赖上下文的替代策略,同时保持任务目标和底层控制器不变。我们在仿真和真实机器人上,在受控负载和非对称活动能力约束条件下,对该框架在抓取任务上进行了评估。实验表明,基于身体状态的重规划在保持较高任务成功率的同时,降低了物理消耗,并实现了更高效的策略自适应。额外的接触密集型操作实验进一步证明了同一重规划接口在抓取任务之外的适用性。这些结果表明,内部物理状态不仅能够为底层控制提供依据,还可以为如何执行操作任务的高层决策提供支持。
cs.RO / 77 / 2609.30056

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

M3GD:用于相机—激光雷达新视角合成的多模态多视角几何扩散方法
Zhou, Yang, Xiao, Jiuhong, Ye, Shizhao, Quang, Long, Nieto-Granda, Carlos, Loianno, Giuseppe
Abstract
Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generative NVS methods rely only on images, overlooking LiDAR, a complementary sensor common on robotic platforms. We present M3GD, a Camera--LiDAR multimodal representation for generative NVS that composes independently pretrained 2D image and 3D point-cloud foundation models without separately pretraining a cross-modal translator. We show that, after camera projection, frozen LiDAR and image features exhibit substantial shared spatial structure, providing a natural cross-modal representation. M3GD conditions generation on LiDAR through this structure: it combines explicit geometry statistics with learned point-cloud descriptors into view-aligned packets on the image-latent grid, injected through a lightweight residual adapter into a multi-view flow-matching generator whose latent space, decoders, and training objective remain intact. On the GrandTour dataset, M3GD improves target-view RGB and depth synthesis over an image-only version of the same backbone. Ablations show that the gains come from pixel-aligned LiDAR content and that target-view LiDAR acts as a geometric query linking the requested view to source observations. Deployment on a ground robot demonstrates practical real-world operation, with a configurable quality--cost trade-off controlled by the number of Euler integration steps.
Chinese Translation
机器人新视角合成(NVS)需要同时恢复视觉外观和度量三维结构,然而大多数生成式NVS方法仅依赖图像,忽视了激光雷达(LiDAR)这一机器人平台上常见的互补传感器。我们提出M3GD,一种面向生成式NVS的相机—激光雷达多模态表示方法,它将独立预训练的2D图像基础模型与3D点云基础模型进行组合,而无需单独预训练跨模态转换器。我们证明,经过相机投影后,冻结的激光雷达特征与图像特征展现出大量共享的空间结构,从而提供了一种天然的跨模态表示。M3GD通过该结构以激光雷达为条件进行生成:它将显式几何统计量与学习得到的点云描述子组合成图像潜在网格上与视角对齐的数据包,并通过一个轻量级残差适配器注入到一个多视角流匹配生成器中,该生成器的潜在空间、解码器和训练目标均保持不变。在GrandTour数据集上,M3GD相较于仅使用图像的相同骨干网络版本,提升了目标视角RGB与深度合成的质量。消融实验表明,性能提升来源于像素对齐的激光雷达内容,且目标视角激光雷达充当了将所请求视角与源观测相联系的几何查询。在地面机器人上的部署演示了实际的真实世界运行能力,并可通过欧拉积分步数实现可配置的质量—成本权衡。
cs.RO / 78 / 2609.30082

Real-Time Force Regulation for Whole-Hand Dexterous Grasping

面向全手灵巧抓取的实时力调控
Kim, Sang Min, Alexiev, Alexander, Lin, Tzu-Yuan, Kim, Sangbae, Kim, Young Min, Lee, Yonghyeon
Abstract
Robust dexterous grasping requires maintaining physical stability despite contacts interactively evolving across the entire hand. A precomputed force distribution can easily fail under object motion, modeling errors, or external disturbances. In this paper, we present a framework for real-time force regulation over dynamically changing whole-hand contacts. Our method geometrically estimates contacts across all hand links using a tracked object model and proprioception, without requiring tactile sensing at those contacts. It repeatedly recomputes the desired contact-force distribution subject to friction constraints, actuator limits, and an actuation-consistency constraint motivated by classical whole-limb force analysis. We integrate this force-regulation controller with reactive reaching, enabling the hand to acquire a grasp, maintain it under disturbances, and regrasp after losing the object. Simulation experiments without gravity demonstrate improved grasp retention over fixed-allocation and fingertip-only execution under controlled perturbations, while real-world experiments on a 27-DoF arm-hand system demonstrate grasp maintenance and recovery under human-applied disturbances as contacts evolve across the whole hand. Project page: https://sangminkim-99.github.io/reactive-grasp-whole-hand/
Chinese Translation
鲁棒的灵巧抓取需要在与整只手接触不断交互演化的情况下保持物理稳定性。预计算的力分布在物体运动、建模误差或外部扰动下很容易失效。本文提出了一个针对动态变化的全手接触的实时力调控框架。我们的方法利用被跟踪的物体模型和本体感觉,从几何上估计手部所有连杆上的接触,而无需在这些接触处使用触觉传感。该方法在摩擦约束、执行器限制以及受经典整肢力分析启发的驱动一致性约束下,反复重新计算期望的接触力分布。我们将该力调控控制器与反应式抓取到达相结合,使手能够完成抓取、在扰动下保持抓取,并在丢失物体后重新抓取。在无重力的仿真实验中,在受控扰动下,该方法相比固定力分配和仅指尖执行方案表现出更好的抓取保持能力;在具有27个自由度的手臂-手系统的真实世界实验中,该方法在全手接触演化的情况下,展示了在人为施加扰动下的抓取保持与恢复能力。项目主页:https://sangminkim-99.github.io/reactive-grasp-whole-hand/
cs.RO / 79 / 2609.30092

Self-Adaptive VLA for Robust Robot Deployment

面向鲁棒机器人部署的自适应VLA
Zhang, Hongxin, Lin, Chunru, Wang, Tsun-Hsuan, Xu, Zhenjia, Gan, Chuang
Abstract
While Vision-Language-Action (VLA) models demonstrate impressive capabilities in robotic manipulation, their memoryless nature renders them brittle to test-time environment shifts, particularly hardware shifts caused by wear or imperfect calibration. Enabling these models to self-adapt during deployment without requiring continuous on-site recalibration remains a critical bottleneck for real-world scalability. In this work, we introduce Self-Adaptive VLA, a novel post-training recipe that enables the policy to iteratively adapt to deployment-time hardware shifts leveraging its own rollouts as context. To do so, we first collect policy rollouts under deliberately injected hardware shifts. We then transform the base policy's training data into shift-conditioned expert demonstrations by pre-compensating the expert actions for these known shifts. Next, we introduce a lightweight, plug-in context encoder that compresses the context, including visual observation, proprioception, and actions in the shifted environment, into a latent context token. This token modulates the policy through adaptive layer normalization (AdaLN). Furthermore, we find that context tokens can be ensembled, allowing the policy to iteratively self-correct and mitigate failures step by step. Extensive experiments across four precision-critical bi-manual and dexterous manipulation tasks show that Self-Adaptive VLA recovers over 80% of the base policy's performance under hardware shifts, such as actuation bias and joint encoder offsets. Moreover, Self-Adaptive VLA enables more robust deployment to new workstations compared to the base policy. Our approach provides a pathway for robust large-scale real-world robot deployments and easier maintenance. See videos at https://icefoxzhx.github.io/self-adaptive-vla.
Chinese Translation
尽管视觉-语言-动作(Vision-Language-Action, VLA)模型在机器人操作中展现出令人瞩目的能力,但其无记忆的特性使其在测试时的环境变化下表现得十分脆弱,尤其是由硬件磨损或标定不完善引起的硬件变化。使这些模型能够在部署过程中自我适应,而无需持续的现场重新标定,仍然是实现真实世界规模化应用的关键瓶颈。在本工作中,我们提出了自适应VLA(Self-Adaptive VLA),这是一种新颖的后训练方案,使策略能够利用其自身的 rollout 作为上下文,迭代地适应部署时的硬件变化。为此,我们首先在故意注入的硬件变化下收集策略的 rollout。然后,通过对专家动作针对这些已知变化进行预补偿,将基础策略的训练数据转化为条件化的专家演示。接下来,我们引入了一个轻量级的即插即用上下文编码器,将变化环境中的上下文信息(包括视觉观测、本体感觉和动作)压缩为一个潜在上下文 token。该 token 通过自适应层归一化(AdaLN)对策略进行调制。此外,我们发现上下文 token 可以进行集成(ensemble),使策略能够迭代地自我纠正并逐步缓解失败。在四个对精度要求极高的双臂和灵巧手操作任务上的大量实验表明,自适应VLA在执行器偏差和关节编码器偏移等硬件变化下,能够恢复基础策略80%以上的性能。此外,与基础策略相比,自适应VLA能够更鲁棒地部署到新的工作站。我们的方法为大规模真实世界机器人部署的鲁棒性和更便捷的维护提供了一条可行路径。视频见 https://icefoxzhx.github.io/self-adaptive-vla。
cs.RO / 80 / 2609.30127

Faster Visuomotor Policy Learning on Action Manifolds via Riemannian MeanFlow

基于黎曼MeanFlow的动作流形上快速视觉运动策略学习
Bukhari, S. Talha, Garrett, Austin, Wei, Yi, Ni, Ruiqi, Kingston, Zachary, Bera, Aniket
Abstract
Visuomotor policies learn a direct map from raw sensory observations to robot action sequences. Policies based on Diffusion and Flow Matching capture the multimodal distribution over action sequences in an end-to-end manner. This expressivity comes at the cost of multi-step numerical integration of the learned vector field for action generation, which can be expensive and time-consuming, impeding fast control rates required in robotics applications. Furthermore, robot action sequences are usually defined on a smooth, differentiable manifold, requiring that the learned policy respects the intrinsic geometry of the robot's action space. Here, we present Riemannian MeanFlow Policy (RMFP), which learns the conditioned flow map of the probability path on the robot action manifold. Our formulation employs a flow map consistency objective grounded in the data by a Riemannian Conditional Flow Matching anchor. The flow map consistency condition is stable to train and constrains the learned model to finite-time transport, which yields on-manifold action sequence generation with as few as one network function evaluation. We present results on the spherical LASA and Push-T benchmarks, on the Tool Hang and Transport tasks of the Robomimic suite, and on the Franka Kitchen task with manifold-constrained action generation, and demonstrate that RMFP attains performance competitive with prior work at a lower sampling cost. We also employ RMFP on a real-world robotic manipulation task to demonstrate fast action generation under imperfect sensor measurements in the physical world.
Chinese Translation
视觉运动策略学习从原始传感器观测到机器人动作序列的直接映射。基于扩散模型(Diffusion)和流匹配(Flow Matching)的策略以端到端的方式捕捉动作序列的多模态分布。这种表达能力以多步数值积分学习到的向量场来生成动作为代价,计算开销大且耗时,阻碍了机器人应用所需的快速控制频率。此外,机器人动作序列通常定义在光滑可微流形上,这要求所学策略遵循机器人动作空间的内在几何结构。本文提出黎曼MeanFlow策略(Riemannian MeanFlow Policy, RMFP),在机器人动作流形上学习概率路径的条件流映射。我们的方法采用流映射一致性目标,并通过黎曼条件流匹配锚点(Riemannian Conditional Flow Matching anchor)将其与数据相结合。该流映射一致性条件训练稳定,并将所学模型约束为有限时间传输,从而仅需一次网络函数评估即可在流形上生成动作序列。我们在球形LASA和Push-T基准任务、Robomimic套件中的Tool Hang和Transport任务,以及具有流形约束动作生成的Franka Kitchen任务上展示了实验结果,证明RMFP在更低的采样成本下取得了与先前工作相当的性能。我们还将RMFP应用于真实世界的机器人操作任务,验证了其在物理世界中不完美传感器测量条件下的快速动作生成能力。
cs.RO / 81 / 2609.30134

Training-free Behavior Cloning

免训练行为克隆
Adang, Maximilian, Chen, Timothy, Osterberg, Lars, Swann, Aiden, Schwager, Mac
Abstract
Neural behavior cloning compresses demonstrations into large models, making individual actions difficult to trace and policy updates costly. Retrieval policies retain access to demonstrations but struggle with mismatch between recorded and live behavior. We introduce Behavior Predictive Control (BPC), which synthesizes policies without end-to-end policy training by combining an action-aware retrieval metric, a Hankel-based action-continuation prior, and a closed-form one-step residual correction. Inspired by behavioral systems theory, BPC predicts future actions by blending stored observation-action data that best reconstructs the recent runtime observation--action history. Across simulated benchmarks and real-robot deployments, BPC is competitive with learned policies such as $\pi_{0.5}$ (surpassing it in some cases), while reducing policy fitting from hours to seconds on consumer GPUs and supporting closed-loop control upwards of 75 Hz on a Jetson Orin Nano. The retrieved demonstration windows and their coefficients also provide an intrinsic estimate of task progress. Retaining demonstrations within the deployed policy makes its predictions traceable to supporting trajectories and enables behavior revision through the demonstration bank.
Chinese Translation
神经行为克隆将演示数据压缩进大型模型中,使得单个动作难以追溯,且策略更新的代价高昂。检索策略虽然保留了对演示数据的访问能力,但难以应对记录行为与实时行为之间的不匹配问题。我们提出了行为预测控制(Behavior Predictive Control, BPC),该方法无需端到端策略训练即可合成策略,其核心包括:动作感知的检索度量、基于 Hankel 矩阵的动作延续先验,以及闭式一步残差校正。受行为系统理论的启发,BPC 通过混合最能重构近期运行时“观测-动作”历史的存储观测-动作数据来预测未来动作。在多个仿真基准和真实机器人部署中,BPC 与 $\pi_{0.5}$ 等学习型策略相比具有竞争力(在某些场景下甚至超越后者),同时将策略拟合时间从数小时缩短至数秒(在消费级 GPU 上即可完成),并能在 Jetson Orin Nano 上支持高达 75 Hz 的闭环控制。此外,检索到的演示窗口及其系数还能提供对任务进度的内在估计。将演示数据保留在已部署的策略中,使其预测结果可追溯至支持轨迹,并支持通过演示库对行为进行修订。
cs.RO / 82 / 2609.30140

Contact as a Decision Variable: Capability-Tradeoff Contact Selection for Legged Loco-Manipulation

将接触作为决策变量:面向足式移动操作的能力权衡接触选择
Mahmud, Al Jaber, Li, Shuai, Wang, Xuan
Abstract
In this paper, we study the joint selection of an environmental support contact and a whole-body configuration for a prescribed loco-manipulation task. A contact may provide greater physical support while restricting the motion required for the task. We formulate this problem through three capability measures: residual wrench, end-effector reach, and base mobility available after satisfying the task requirements, and we balance them against contact acquisition cost. Evaluating these capabilities for every candidate requires repeated whole-body optimizations. To reduce this computational cost, we propose Capability-Tradeoff Contact Selection (CTCS). CTCS screens candidates for contact and task feasibility, groups similar candidates within each surface, and predicts their capabilities from exact anchor evaluations using local sensitivity analysis. It checks these predictions through selective exact evaluations, ranks candidates by capability, and evaluates a shortlist exactly for final selection. We evaluate CTCS in simulations and hardware experiments using a Unitree Go2 quadruped with an AgileX NERO arm across $392$ task conditions with nine available support surfaces. Results show that CTCS outperforms ground-only and fixed-contact support, as it can select support surfaces that provide favorable capability trade-offs for the task. Compared with evaluating every candidate exactly, CTCS achieves approximately $3\times$ speedup while closely matching the resulting mean objective value.
Chinese Translation
本文研究了针对给定移动操作(loco-manipulation)任务,环境支撑接触与全身构型的联合选择问题。一个接触点可能提供更强的物理支撑,但同时会限制完成任务所需的运动。我们通过三种能力度量来构建该问题:满足任务需求后剩余的余量旋力(residual wrench)、末端执行器可达空间和基座机动性,并将其与接触获取代价进行权衡。对每个候选接触评估这些能力需要重复进行全身优化。为降低计算成本,我们提出了能力权衡接触选择方法(Capability-Tradeoff Contact Selection, CTCS)。CTCS 首先对候选接触进行接触与任务可行性筛选,然后对每个表面上相似的候选进行分组,并利用局部敏感性分析从精确的锚点评估中预测其能力。该方法通过选择性的精确评估来校验这些预测,按能力对候选进行排序,并对最终候选名单进行精确评估以完成最终选择。我们在仿真和硬件实验中评估了 CTCS,使用搭载 AgileX NERO 机械臂的 Unitree Go2 四足机器人,在九个可用支撑面下的 $392$ 种任务条件下进行测试。结果表明,CTCS 优于仅地面支撑和固定接触支撑,因为它能够选择为任务提供有利能力权衡的支撑面。与对每个候选进行精确评估相比,CTCS 实现了约 $3\times$ 的加速,同时平均目标值与精确评估结果高度吻合。
cs.RO / 83 / 2609.30213

ReVAMP: Vector-Accelerated Motion Planning for Kinematically-Constrained Systems via Reparameterization

ReVAMP:基于重参数化的运动学约束系统向量加速运动规划
Iyer, Shrutheesh R., Cohn, Thomas, Kingston, Zachary
Abstract
Robots often must satisfy one or more constraints during motion planning for real-world tasks. When such constraints reduce the valid configuration space to a measure-zero subset, sampling based planning algorithms require modifications to draw feasible samples. For many common end-effector constraints, parameterizations built on inverse kinematics (IK) provide an alternate formulation where the constraints are satisfied by construction, allowing directly sampling the feasible set. Despite their elegant approach, parameterized planners have remained slower than vector-accelerated implementations of projection-based approaches, leaving their performance ceiling an open question. We explore a new axis of vectorization built upon reparameterizing the planning space through analytic IK. This approach addresses existing inefficiencies in vectorized projection-based planners and exposes new opportunities for parallelism within the planner. We show that the planner can synthesize plans in microseconds to milliseconds for high dimensional systems (up to 20 dimensions), with complex constraints, up to 10x faster than the current state-of-the-art. Furthermore, we demonstrate how such planning speeds open up avenues for restructuring sequential manipulation pipelines.
Chinese Translation
在现实任务的运动规划中,机器人通常需要满足一个或多个约束条件。当此类约束将有效构型空间缩减为零测度子集时,基于采样的规划算法需要进行修改才能采样到可行样本。对于许多常见的末端执行器约束,基于逆运动学(IK)构建的参数化方法提供了一种替代形式,其中约束通过构造方式自然满足,从而可以直接对可行集进行采样。尽管参数化规划器方法优雅,但其速度一直落后于基于投影方法的向量化实现,其性能上限仍是一个悬而未决的问题。我们探索了一条新的向量化途径,即通过解析逆运动学对规划空间进行重参数化。该方法解决了现有向量化投影规划器中的低效问题,并揭示了规划器内部新的并行化机会。我们证明,该规划器能够为高维系统(最高20维)在微秒到毫秒级时间内合成具有复杂约束的运动规划,比当前最先进方法快达10倍。此外,我们展示了如此高的规划速度如何为重构顺序操作流水线开辟新的途径。
cs.RO / 84 / 2609.30214

Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage

水下C3-JEPA:一种面向ROV打捞的对象中心跨视角世界模型
Yang, Yuncong, Li, Jinlong, Xue, Yulong, Wu, Feng, Zhang, Chunwen, Qiao, Lei, Wang, Xuyang
Abstract
We present Underwater C$^{3}$-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-view predictive world model for near-field heavy-load underwater ROV salvage. Without contact sensors, it predicts in latent space how the task-object state evolves through contact interaction and under the hydrodynamic lag of the vehicle, from synchronized multi-view RGB observations and vehicle control signals. C$^{3}$-JEPA encodes multi-camera observations into task-object and context tokens, fuses cross-camera evidence through held-out-view attention, and directly predicts future states conditioned on control. Weak binding anchors the target and gripper at low annotation cost, while SIGReg sharpens the geometric representation. Experiments show that the learned representation transfers substantially more task-relevant information to downstream probes than a reconstruction-free latent baseline, while keeping the predictor lightweight. The resulting predictive interface supports model-predictive-control (MPC) candidate evaluation and imagined-rollout behavior-agent training. Validation on real underwater video shows the same architecture recovering a withheld camera's object state and staying ahead of persistence, so the recipe transfers beyond simulation.
Chinese Translation
我们提出了水下C$^{3}$-JEPA(跨视角、控制条件化、上下文扩展),这是一种面向近场重载水下ROV打捞的对象中心多视角预测世界模型。在无需接触传感器的情况下,该模型从同步的多视角RGB观测和载具控制信号出发,在潜空间中预测任务对象状态如何通过接触交互并在载具流体动力学滞后作用下演化。C$^{3}$-JEPA将多相机观测编码为任务对象令牌和上下文令牌,通过留出视角注意力机制融合跨相机证据,并以控制信号为条件直接预测未来状态。弱绑定以低标注成本锚定目标物体与机械爪,同时SIGReg锐化了几何表示。实验表明,与无重建的潜空间基线相比,所学表示向下游探针迁移的任务相关信息显著更多,同时保持预测器轻量化。所得的预测接口支持模型预测控制(MPC)候选动作评估以及想象式推演的行为智能体训练。在真实水下视频上的验证表明,同一架构能够恢复被隐去相机的对象状态并领先于持续性基线,说明该方法可迁移至仿真之外。
cs.RO / 85 / 2609.30233

Coding Agents for Generalized Task and Motion Planning Problems

面向广义任务与运动规划问题的代码智能体
Merler, Matteo, Li, Bowen, Roy, Josh, Liang, Yichao, Wang, Qianwei, Huang, Yixuan, Silver, Tom
Abstract
Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generalized TAMP addresses this difficulty by exploiting regularities across problem instances to reduce planning effort on new instances. However, existing methods require substantial TAMP-specific engineering. We investigate whether coding agents can automate this process by synthesizing programs that generalize across instances. Given a task description and simulator access, each agent chooses how to interact with the environment while developing a program within a fixed synthesis budget. The program is then frozen and evaluated on unseen instances. We evaluate Claude Code (Opus 5) and Codex (GPT-5.6 Sol and GPT-6 Astra) on 28 simulated environments from KinDER and PDDLStream, with object counts beyond those evaluated in the original benchmark. Across all program synthesis methods, we evaluate 980 generated programs on 100 held-out instances each, 98,000 evaluation episodes in total. Overall, we find that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success (56% to 95% versus 47% for the planners, on the 16 environments where a planner is available). As object counts grow, the agents' programs maintain higher success than the planner, using an order of magnitude less computation per instance on average. Logs show agents using interaction to calibrate physical models, test edge cases, and refine strategies. We release all code, including the full prompts given to the agents. These findings suggest that coding agents are a strong baseline for generalized TAMP.
Chinese Translation
即使具备完全可观测性和以对象为中心的状态,任务与运动规划(Task and Motion Planning, TAMP)问题仍然十分困难,因为离散决策与几何、运动学和动力学约束紧密耦合。广义TAMP(Generalized TAMP)通过利用问题实例之间的规律性来减少在新实例上的规划开销,从而应对这一困难。然而,现有方法需要大量的TAMP专用工程工作。我们研究代码智能体(coding agents)能否通过合成可跨实例泛化的程序来实现这一过程的自动化。在给定任务描述和模拟器访问权限的情况下,每个智能体在固定的程序合成预算内自主选择如何与环境交互并开发程序。随后程序被冻结,并在未见过的实例上进行评估。我们在来自KinDER和PDDLStream的28个模拟环境上评估了Claude Code(Opus 5)和Codex(GPT-5.6 Sol与GPT-6 Astra),其中对象数量超出了原始基准测试所评估的规模。对于所有程序合成方法,我们在每个100个保留实例上共评估了980个生成的程序,总计98,000个评估回合。总体而言,我们发现代码智能体在广义TAMP上表现出惊人的有效性:在提供规划器的16个环境中,三种智能体配置的平均成功率(56%至95%)均优于手工设计的规划器(47%)、一次性生成以及基于LLM的广义规划基线。随着对象数量的增长,智能体生成的程序保持比规划器更高的成功率,且每个实例平均使用的计算量少一个数量级。日志显示,智能体利用交互来校准物理模型、测试边界情况并改进策略。我们发布了全部代码,包括提供给智能体的完整提示词。这些发现表明,代码智能体是广义TAMP的一个强有力的基线方法。
cs.RO / 86 / 2609.30247

Rolling-WAM: World Action Models with Rolling Imagination

Rolling-WAM:具有滚动式想象的世界动作模型
Zhou, Yinghua, Ye, Junjie, Zhao, Yiqi, Dong, Hao, Wang, Celina Shiyu, Ge, Ruohai, Yang, Tingyi, Van Hoorick, Basile, Sukhatme, Gaurav, Guizilini, Vitor, Wang, Yue
Abstract
World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.
Chinese Translation
世界动作模型(World Action Models, WAMs)将动作生成与未来视觉预测相结合,用于机器人操作。然而,在每个重规划周期内完成联合的视频-动作去噪过程会带来显著的延迟,从而延缓动作更新并限制闭环响应能力。我们提出了Rolling-WAM,一种将联合去噪分配到连续重规划周期中的方法。我们的方法维护一个处于交错噪声水平的视频-动作块滑动窗口。在每一步中,滚动噪声调度会对即将执行的动作块进行完全去噪,同时对更远的未来块进行部分精炼。随着新相机观测的到来,窗口向前推进,被保留的未来块则继续其去噪过程。这种方式将计算成本在时间上分散,同时在块边界之间延续不断演化的视觉-动作上下文。在LIBERO、RoboTwin以及真实世界Unitree G1人形机器人上的评估表明,Rolling-WAM取得了具有竞争力的操作性能。由于无需从头对整个预测时域进行去噪,它相比标准联合WAM实现了4.5倍的稳态重规划加速。
cs.RO / 87 / 2609.30249

RAPID: Robot Agentic Programming from Demonstrations

RAPID:基于示范的机器人智能体编程
Liu, Yuyao, Mao, Jiayuan, Hsu, David, Kaelbling, Leslie Pack, Lozano-Pérez, Tomás
Abstract
Coding agents have demonstrated enormous success in solving complex programming problems. To leverage their potential for robot systems, this work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstration. The iterative agentic loop of code refinement requires several key ingredients: (i) a testable task specification, (ii) action primitives for robot execution, and (iii) an interactive environment for program execution and verification. RAPID infers all three from the demonstration automatically. To make the resulting program reusable beyond the demonstration setting, RAPID uses an object-centric relational program representation that focuses on the underlying structure of the demonstrated strategy rather than the specific motion per se: it expresses the action primitives as trajectory-optimization programs that realize object-level motion effects, while composing them through relational constraints that capture scene-specific geometry at run time. We evaluated RAPID in simulation on eight challenging contact-rich nonprehensile manipulation tasks as well as general prehensile manipulation tasks in the LIBERO-Pro benchmark. We also successfully deployed it on a real Franka arm and evaluated on all eight nonprehensile tasks. In all experiments, RAPID demonstrated strong performance, with generalization over object pose, shape, material, and environment. Website: https://yuyaoliu.me/projects/rapid.
Chinese Translation
编程智能体(coding agents)在解决复杂编程问题方面已展现出巨大成功。为了将其潜力应用于机器人系统,本工作提出了基于示范的机器人智能体编程(Robot Agentic Programming from Demonstrations, RAPID),该方法能够在仅给定一段人类视觉示范的情况下,自动生成、验证并改进机器人程序。这一迭代的智能体代码改进循环需要几个关键要素:(i)可测试的任务规范,(ii)用于机器人执行的动作基元,以及(iii)用于程序执行与验证的交互式环境。RAPID能够从示范中自动推断出这三者。为了使所得程序在示范场景之外仍可复用,RAPID采用一种以对象为中心的关系型程序表示,其关注所示范策略的底层结构而非具体运动本身:它将动作基元表达为轨迹优化程序,以实现对象级别的运动效果,同时通过在运行时捕捉场景特定几何信息的关系约束来组合这些基元。我们在仿真中于八个具有挑战性的富接触非抓取操作任务以及LIBERO-Pro基准中的通用抓取操作任务上对RAPID进行了评估。我们还将其成功部署于真实的Franka机械臂上,并在全部八个非抓取任务上进行了评估。在所有实验中,RAPID均展现出强大的性能,并在物体姿态、形状、材质和环境方面具有良好的泛化能力。网站:https://yuyaoliu.me/projects/rapid。
人工智能 (Artificial Intelligence)
107
cs.AI / 1 / 2609.28475

When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability Routing

预测智能体何时应该进行推理?面向可靠性路由的行为压力测试
Wang, Yufeng
Abstract
Forecasting agents increasingly combine language-model reasoning, retrieval, ensembling, and calibration, but it remains unclear when each behavior should be trusted. We study this question on ForecastBench-style binary forecasting tasks, treating the choice to retrieve, reason, defer to a market prior, or use a historical analog as an observable agent behavior rather than a hidden implementation detail. Our central finding is that mechanism choice is source-dependent: structured analogs dominate for some data-generating processes, while market/crowd-style and conservative baselines are better for others. We introduce ReliabilityRoute, a structural intervention that steers forecasting-agent behavior using reliability features such as historical coverage, market-prior availability, source-prior sharpness, evidence strength, evidence disagreement, and horizon. A fixed 2024-fitted rule closely matches a hand taxonomy without hard-coded source-name decisions, while a walk-forward self-adjusting rule refits thresholds from previously resolved vintages and obtains the best mean Brier score among our deterministic systems across 16 later LLM vintages. The gain is modest and historical/search baselines remain highly competitive. The main contribution is therefore a behavioral stress test showing that more reasoning is not always better; forecasting agents should first estimate which evidence source deserves control, routing policies should themselves adapt under auditable constraints, and reproducibility artifacts are available at https://github.com/louiswang524/forcastagent
Chinese Translation
预测智能体日益融合语言模型推理、检索、集成与校准等技术,但每种行为何时值得信赖仍不明确。我们在 ForecastBench 风格的二值预测任务上研究这一问题,将检索、推理、依赖市场先验或使用历史类比等选择视为可观测的智能体行为,而非隐藏的实现细节。我们的核心发现是:机制选择依赖于数据来源——结构化历史类比在某些数据生成过程中占优,而市场/群体风格与保守基线在其他场景下表现更佳。我们提出 ReliabilityRoute,一种结构化干预方法,利用历史覆盖率、市场先验可用性、来源先验锐度、证据强度、证据分歧以及预测期限等可靠性特征来引导预测智能体的行为。一个基于 2024 年数据拟合的固定规则在无需硬编码来源名称决策的情况下即可紧近人工分类法;而一种前向滚动的自适应规则则基于先前已解决的事件数据重新拟合阈值,在 16 个较新的 LLM 版本上,在我们所有确定性系统中取得了最优的平均 Brier 分数。该增益较为有限,且历史/搜索基线仍然极具竞争力。因此,本文的主要贡献在于一项行为压力测试,其结果表明:更多的推理并不总是更好;预测智能体应首先估计哪些证据来源值得掌控,路由策略本身应在可审计的约束下进行自适应调整,且可复现性工具已发布于 https://github.com/louiswang524/forcastagent。
cs.AI / 2 / 2609.28506

TW3Cast: A Frozen Router of Lightly Fine-Tuned Foundation Models for Time-Series Forecasting on GIFT-Eval, Selected Entirely on the Training Split

TW3Cast:一个冻结路由器,整合轻度微调的基础模型用于GIFT-Eval时间序列预测,其选择完全基于训练集划分
Thierry, Nathan, Rochet, Andre-Louis
Abstract
TW3Cast is a time-series forecasting system that reaches position 3 of 130 entries on the GIFT-Eval benchmark by mean MASE rank, as of 2026-09-14. The two entries above it belong to the leaderboard's agentic category, multi-step systems that use agents or language models to reason about, generate or select forecasts. TW3Cast runs no agent and no language model. Its selection is a table computed once on the training split and then frozen, and its experts are public foundation models lightly fine-tuned on those training splits. For each of the 97 dataset, frequency and horizon configurations, the table serves one of four modes: a specialist, which is a LoRA or full fine-tune of Chronos-2, TiRex or Toto whose training data was cleaned and enriched by explicit rules; a quantile blend that contains a specialist; a blend of base models; or a selection tournament played on a backtest carved from the training split. Every decision in the table was taken on that backtest. A specialist is admitted the moment it beats the tournament there, so a candidate costs a few megabytes and minutes of GPU time, and a failed candidate changes nothing. Three guarded mechanisms protect the selection from its own biases: a dual accuracy and calibration criterion, an asymmetric margin against candidates that saw the series during training, and conservative per-window gates. The selection rules themselves were chosen inside a temporal meta-backtest. The best base model served alone reaches a mean MASE rank of 33.8, the tournament served on every configuration reaches 38.0, and the full router reaches 19.4. The routing table, the expert index, the pinned base-model revisions, the submitted score file and the dated snapshot of the public scores are released, and every leaderboard number in this paper regenerates from them by one script.
Chinese Translation
TW3Cast 是一个时间序列预测系统,截至2026年9月14日,其平均 MASE 排名在 GIFT-Eval 基准测试的 130 个参赛系统中位列第3。排在其之前的两个系统属于排行榜上的智能体类别,即使用智能体或语言模型对预测进行推理、生成或选择的多步系统。而 TW3Cast 既不使用智能体,也不使用语言模型。其模型选择是一个仅在训练集划分上计算一次后即被冻结的表格,其专家模型是在这些训练集划分上经过轻度微调的公开基础模型。对于全部 97 种数据集、频率和预测时域配置,该表格提供四种模式之一:专家模型,即 Chronos-2、TiRex 或 Toto 的 LoRA 或全量微调版本,其训练数据经过显式规则的清洗与增强;包含专家模型的分位数混合模型;基础模型的混合模型;或在从训练集划分中切出的回测集上进行的选择锦标赛。表格中的每一个决策都在该回测集上做出。专家模型一旦在回测中胜过锦标赛,即被接纳,因此一个候选模型的成本仅为几兆字节存储和几分钟的 GPU 时间,而失败的候选模型不会带来任何影响。三种带防护的机制保护选择过程免受自身偏差的影响:准确率与校准度的双重准则、针对在训练中见过该序列的候选模型的不对称裕度,以及保守的逐窗口门控。选择规则本身也是在时间维度的元回测中确定的。单独服务的最佳基础模型的平均 MASE 排名为 33.8,在每个配置上都使用锦标赛时为 38.0,而完整路由器达到 19.4。路由表、专家索引、锁定的基础模型版本、提交的分数文件以及公开分数的带日期快照均已公开,本文中所有排行榜数字均可通过一个脚本从这些材料中重新生成。
cs.AI / 3 / 2609.28547

PAWS: Policy-driven Agentic World Simulation

PAWS:策略驱动的智能体世界模拟
Sim, Tiviatis, Woon, Jia Hui, Gao, Xinming, Gao, Chen, Zhu, Fengbin, Huanhuan, Zheng, Seng, Chua Tat, Kawaguchi, Kenji
Abstract
Policy interventions propagate through public communication, institutional decisions, and stakeholder responses, yet datasets for financial multi-agent simulation rarely connect these processes to temporally aligned historical evidence. We introduce PAWS, a Policy-driven Agentic World Simulation dataset covering 36 verified U.S. financial and economic policy episodes, 12,727 policy-linked news records, and 65,291 source-grounded stakeholder actions. Each action is linked to its supporting news and represented by a multi-layer event frame capturing its interaction mode, financial-action family and subtype, semantic attributes, and conditional mappings to external taxonomies. Entities are resolved to normalized organizations, and actions are aligned with daily market-return context to support policy-agent simulation replay. On 2,522 stratified action samples, independent AI and human reviewers achieved 89.4% initial agreement on interaction mode, with disagreements subsequently adjudicated. Case studies of the 2008 short-selling ban and 2001 decimalization recover documented policy timelines and associated market patterns across both dense and sparse news settings. A replay study further shows that high accuracy can mask failure to detect rare stakeholder actions, identifying action timing and calibration as central challenges. PAWS provides an auditable substrate for evaluating agent influence, policy-response cascades, and action-outcome alignment in historically grounded financial simulations.
Chinese Translation
政策干预会通过公共传播、机构决策和利益相关者回应进行传导,然而用于金融多智能体模拟的数据集很少将这些过程与时间对齐的历史证据相连接。我们提出了PAWS(Policy-driven Agentic World Simulation,策略驱动的智能体世界模拟)数据集,涵盖36个经过核实的美国金融与经济政策事件、12,727条与政策相关的新闻记录以及65,291条有来源依据的利益相关者行动。每条行动均与支持它的新闻相链接,并通过多层事件框架来表示,该框架刻画了行动的交互模式、金融行动类别及子类型、语义属性,以及与外部分类体系的条件映射。实体被解析为规范化的组织,行动与每日市场收益背景对齐,以支持政策-智能体模拟回放。在2,522个分层行动样本上,独立的AI与人类评审者在交互模式上达到了89.4%的初始一致率,分歧随后经过了仲裁裁定。对2008年卖空禁令和2001年十进制化改革的案例研究,在新闻密集和稀疏两种设定下均还原了有据可查的政策时间线及相关的市场模式。一项回放研究进一步表明,高准确率可能掩盖对罕见利益相关者行动的漏检,并指出行动时机与校准是核心挑战。PAWS为在具有历史依据的金融模拟中评估智能体影响力、政策响应级联以及行动-结果对齐提供了可审计的基础。
cs.AI / 4 / 2609.28554

Pistis Technical Report

Pistis 技术报告
Chen, Heyun, Lan, Xiaohan, Li, Jiaxi, Lu, Zhilin, She, Qi, Xu, Weiwen, Yu, Fei, Zhong, Yujie, Chen, Jinghuan, Feng, Zijian, Jiao, Siyu, Lin, Yiheng, Wang, Xinhao, Yang, Sihan, You, Jieyu, Zhang, Changbin, Zhang, Hengyu, Zhang, Xudong, Zhao, Yunqing, Zheng, Shuai
Abstract
We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT). Building on this SFT foundation, we propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel post-training paradigm that tightly integrates on-policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives, rather than optimizing either in isolation or combining them in a static joint loss, IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit assignment for long-horizon agentic trajectories, leading to stronger performance while mitigating common capability trade-offs. At both model scales, the framework produces two specialized variants: Pistis-Thinking, designed to strengthen deep multimodal reasoning, and Pistis-Agentic, which additionally incorporates agentic trajectory data to support long-horizon planning, iterative reasoning, and tool use. Pistis-Agentic is particularly strong in multimodal search. Both scales outperform their corresponding base models. Beyond model-parameter optimization, we further introduce Pistis-Auto-Harnessing (PAH), a system-level method that automatically improves the agent's inference harness through iterative optimization. Experiments demonstrate that PAH enhances the model performance without updating the model parameters or increasing the interaction budget.
Chinese Translation
我们提出了 Pistis 模型系列,包含基于 Qwen3.6 和 Qwen3.5 构建的 27B 和 9B 参数多模态大语言模型,并通过一个通用且可扩展的后训练框架进行开发。该框架首先通过大规模多模态监督微调(SFT)建立坚实基础。在此 SFT 基础之上,我们提出了交错蒸馏与强化学习(Interleaved Distillation and Reinforcement Learning, IDRL),这是一种新颖的后训练范式,在单一训练循环中紧密结合在线策略蒸馏与强化学习。通过在两种目标之间交替进行,而非孤立地优化其中之一或以静态联合损失将二者结合,IDRL 实现了更有效的知识迁移、更高的优化稳定性,以及对长时程智能体轨迹更精确的信用分配,从而带来更强的性能,同时缓解了常见的能力此消彼长问题。在两种模型规模下,该框架均产出两个专用变体:Pistis-Thinking,旨在增强深度多模态推理能力;以及 Pistis-Agentic,其额外引入了智能体轨迹数据,以支持长时程规划、迭代推理和工具使用。Pistis-Agentic 在多模态搜索方面表现尤为突出。两种规模的模型均优于其对应的基础模型。除模型参数优化之外,我们进一步提出了 Pistis-Auto-Harnessing(PAH),这是一种系统级方法,可通过迭代优化自动改进智能体的推理工具框架。实验表明,PAH 无需更新模型参数或增加交互预算即可提升模型性能。
cs.AI / 5 / 2609.28557

BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines

BaseCamp——一个用于自动化DNA测序数据流程的智能体AI框架
Bandara, Eranga, Liang, Xueping, Gunaratna, Asanga, Hewa, Tharaka, Rahman, Abdul, Foytik, Peter, Bouk, Safdar H., Rajapakse, Sachini, Kularathna, Isurunima, Karunarathna, Pramoda, Rajapakse, Chalani, Keong, Ng Wee, De Zoysa, Kasun, Hass, Amin, Herath, Wathsala, Gore, Ross, Mukkamala, Ravi, Siriwardanagea, Nihal, Siriwardanagea, Gihan, Withanage, Aruna, Loganathan, Nilaan, Shetty, Sachin
Abstract
DNA sequencing pipelines, spanning quality control, alignment, variant calling, and annotation, are now reliably executed by workflow management systems that orchestrate established bioinformatics tools at scale. What remains manual is the decision layer surrounding that execution: selecting quality thresholds appropriate to a sample and platform, adjudicating borderline variant calls, diagnosing anomalies, and determining which findings warrant expert review. These decisions are repetitive, judgment-intensive, inconsistent across operators, and frequently undocumented. This paper introduces BaseCamp, a novel agentic AI framework for automating the decision layer of DNA sequencing pipelines. The framework decomposes the pipeline into six specialized AI agents, covering sample intake and quality control, alignment, variant calling, annotation, cross-stage monitoring, and reporting. Critically, BaseCamp agents do not perform sequence analysis: established tools execute alignment, calling, and annotation, while the agents select among them, configure them, interpret their output, and decide what follows. This confines language model reasoning to the judgment layer where it is reliable and preserves the reproducibility existing tooling guarantees. Agent reasoning is powered by a consortium of fine-tuned, domain-specialized large language models coordinated by a central reasoning LLM, executing locally so no sequencing data leaves the operating environment, under human-in-the-loop orchestration. Evaluation shows agent-generated configurations are concordant with expert practice, that an explicit filtering ledger renders inspectable what filtering otherwise removes without trace, and that cross-stage anomaly detection surfaces conditions execution monitoring misses. BaseCamp offers a generalizable blueprint for agentic automation of scientific data pipelines.
Chinese Translation
DNA测序流程涵盖质量控制、序列比对、变异检测和注释,目前可由工作流管理系统可靠地执行,大规模协调成熟的生物信息学工具。仍然依赖人工的是围绕该执行的决策层:为特定样本和平台选择合适的质量阈值、裁定边界性变异检测结果、诊断异常情况,以及确定哪些发现需要专家审核。这些决策重复性强、高度依赖判断力、在不同操作者之间不一致,且往往缺乏记录。本文介绍BaseCamp,一个用于自动化DNA测序流程决策层的新型智能体AI框架。该框架将流程分解为六个专业化AI智能体,分别负责样本接收与质量控制、比对、变异检测、注释、跨阶段监控和报告。关键在于,BaseCamp智能体并不直接执行序列分析:比对、检测和注释由成熟工具完成,而智能体负责在这些工具中进行选择、配置、解读其输出并决定后续步骤。这将语言模型推理限制在其可靠发挥作用的判断层,同时保留了现有工具所保证的可重复性。智能体推理由一组经过微调的领域专业化大语言模型驱动,并由一个中央推理LLM协调,全部本地执行,确保测序数据不离开运行环境,并在人机协同(human-in-the-loop)编排下运行。评估表明,智能体生成的配置与专家实践一致;显式的过滤账本使原本无迹可寻的过滤操作变得可审查;跨阶段异常检测能够发现执行监控所遗漏的问题。BaseCamp为科学数据流程的智能体化自动化提供了一个可推广的蓝图。
cs.AI / 6 / 2609.28570

DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs

DEEPO:面向多模态大语言模型幻觉问题的双熵增强策略优化
Zhuang, Yingxuan, Pan, Miao, Gan, Wangjie, Yang, Jingxiao, Wang, Fan, Liu, Weiming, Tan, Cheng, Zhang, Xuhong, Chen, Jintao
Abstract
Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two weak points in the \emph{correction chain} from reward to parameter update. At the rollout level, hard queries---those with high semantic entropy---frequently produce unanimously wrong sample groups, collapsing the group-relative advantage to zero exactly where hallucination risk is highest. At the optimization level, confident-but-wrong tokens are gradient-invisible: a categorical policy's expected score-gradient norm vanishes as its distribution sharpens, so the predictions that most need correction receive the weakest updates. We propose Dual-Entropy Enhanced Policy Optimization (DEEPO), a dual-stage enhancement combining signal variance regularization with gradient preconditioning: semantic-entropy-triggered expert prefixes inject grounded continuations on high-uncertainty queries, providing direct supervision and restoring advantage variance, while advantage-sign-aware Renyi preconditioning counteracts logit-level saturation so correction reaches confident errors in the operational confidence regime. Both branches improve over GRPO individually; their interaction is statistically significant on VideoMMMU---the most complex long-horizon task in our evaluation suite (+4.0$, 95\% CI [1.1, 6.9])---and additive elsewhere. DEEPO reduces hallucination while preserving accuracy and training stability.
Chinese Translation
强化学习(RL)被广泛用于提升多模态大语言模型(MLLMs)的推理能力,但其对幻觉问题的效果并不均衡。我们将此归因于从奖励到参数更新的"校正链"中的两个薄弱环节。在rollout层面,困难查询——即语义熵较高的查询——常常产生全错的样本组,使得组相对优势恰好在幻觉风险最高的地方坍缩为零。在优化层面,自信但错误的token对梯度不可见:类别策略的期望得分梯度范数随其分布变尖锐而趋于消失,因此最需要校正的预测反而获得最弱的更新。我们提出双熵增强策略优化(Dual-Entropy Enhanced Policy Optimization, DEEPO),这是一种结合信号方差正则化与梯度预条件的双阶段增强方法:由语义熵触发的专家前缀在高不确定性查询上注入有依据的续写内容,提供直接监督并恢复优势方差;同时,优势符号感知的Renyi预条件化抵消logit层面的饱和,使校正在可操作的置信区间内能够触及自信的错误。两个分支各自均优于GRPO;在评估集中最复杂的长时间任务VideoMMMU上,二者的交互作用具有统计显著性(+4.0,95% CI [1.1, 6.9]),在其他任务上则表现为累加效应。DEEPO在保持准确率和训练稳定性的同时减少了幻觉。
cs.AI / 7 / 2609.28575

TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment

TWIST:对话记忆干预质量的基准提案,附人工验证的草案对齐
Panda, Subrat
Abstract
Long-conversation memory benchmarks increasingly test recall and prompted knowledge updates, and recent work studies evolving user beliefs and memory state. TWIST is a proposed benchmark suite for a complementary, unmeasured property: intervention quality -- whether a deployed memory system, exercised through its own ingest/recall/vet surface, acts correctly at belief change points. Four tracks cover unprompted tension detection, vetting outgoing drafts against the record, answering with current beliefs while preserving supersession history, and governing sensitive recall. The suite extends LoCoMo's corpora and harness, pairing every detect/block metric with a matched do-not-over-detect control: surface-matched hard negatives price false intervention, so no track can be gamed by flagging everything. The benchmark itself is validated first: independent, gold-blind double annotation with adjudication, judge decoy calibration, and a separability audit. On the human-validated Track B v1.0 key (161 items, post-adjudication kappa = 0.85), no tested configuration simultaneously achieves high contradiction recall, high hard-negative specificity, and high attribution: flat-RAG baselines detect 0.76-0.97 of true contradictions but falsely flag 16-43% of surface-matched safe drafts depending on backend, while a deployed coherence-oriented system almost never over-flags (0.98-1.00 specificity) yet catches 42% of true contradictions -- a trade-off no recall-only score can see. A 13-configuration baseline ladder localizes causes: every gold contradiction is detectable from its evidence alone (recall 1.000), calibrated models nearly solve the track given the full transcript -- consistent with substantial retrieval-coverage gaps -- and draft-only floors reveal model-dependent style priors. A system's TWIST profile, beside its recall score, measures whether memory knows when to intervene and when not to.
Chinese Translation
长对话记忆基准越来越多地测试召回和提示驱动的知识更新,近期工作也研究了用户信念的演变与记忆状态。TWIST 是一个针对互补性且未被测量属性的拟议基准套件:干预质量——即部署的记忆系统在通过自身的摄取/召回/审核接口运作时,能否在信念变更点上正确行动。该套件包含四个赛道:无提示的张力检测、依据记录审核外发草案、在保留取代历史的前提下基于当前信念作答,以及对敏感召回的治理。该基准扩展了 LoCoMo 的语料库与评测框架,为每个检测/拦截指标都配以匹配的“勿过度检测”对照:表面匹配的困难负例用于为误干预定价,因此任何赛道都无法通过一律标记来投机取巧。基准本身首先经过验证:采用独立、对金标盲测的双重标注与裁决机制、裁判诱饵校准,以及可分性审计。在经人工验证的 Track B v1.0 标准答案上(161 个条目,裁决后 kappa = 0.85),没有任何被测配置能同时实现高矛盾召回率、高困难负例特异性和高归因质量:扁平 RAG 基线能检测出 0.76-0.97 的真实矛盾,但依据后端不同会误标 16-43% 的表面匹配安全草案;而一个已部署的面向一致性系统几乎从不过度标记(特异性 0.98-1.00),却只能捕获 42% 的真实矛盾——这种仅凭召回率分数无法察觉的权衡。由 13 个配置组成的基线阶梯定位了原因:每一条金标矛盾都可以仅凭其证据被检测出来(召回率 1.000);在给定完整对话记录的情况下,校准后的模型几乎能解决该赛道——这与可观的检索覆盖缺口相一致;而仅凭草案的下限则揭示了依赖模型风格的先验。系统的 TWIST 剖析与其召回率分数并列,衡量记忆系统是否知道何时该干预、何时不该干预。
cs.AI / 8 / 2609.28609

Adversarial Closed-Loop Curriculum for Evolving Role-Playing Agents

面向演化型角色扮演智能体的对抗性闭环课程学习
Zhang, Zheng, Liu, Liu, Chai, Qi, Ye, Deheng, Zhao, Peilin, Zheng, Mao, Wang, Hao
Abstract
Role-playing agents based on large language models have been widely applied in areas such as personalized assistance and social simulation. Recent RL methods typically train on a fixed scenario pool collected before learning begins. This creates a distributional bottleneck: as the agent improves, the scenarios where it performs poorly also change, while the training distribution remains static. Therefore, we propose AdvRole, an adversarial context rewriting framework that turns role-playing RL into a closed-loop curriculum. AdvRole alternates between an Actor that learns to role-play and a Rewriter that edits character profiles and dialogue contexts into actor-specific hard scenarios. The Rewriter is trained with a performance-gap reward, which favors rewrites that reduce the current Actor's score relative to the original scenario. As a result, the scenario pool evolves with the Actor and continuously targets under-mastered regions of the character-context space. Experiments on three role-playing benchmarks covering English and Chinese, as well as a new multilingual benchmark we release, show that AdvRole consistently outperforms baselines.
Chinese Translation
基于大语言模型的角色扮演智能体已被广泛应用于个性化助手和社交模拟等领域。近期的强化学习方法通常在训练开始前收集的固定场景池上进行训练。这造成了分布瓶颈:随着智能体能力的提升,其表现不佳的场景也在变化,而训练分布却保持静态。为此,我们提出 AdvRole,一种对抗性上下文重写框架,将角色扮演强化学习转变为闭环课程。AdvRole 在负责学习角色扮演的 Actor 与负责将角色档案和对话上下文改写为针对该 Actor 的困难场景的 Rewriter 之间交替进行。Rewriter 通过性能差距奖励进行训练,该奖励倾向于选择那些相对于原始场景能降低当前 Actor 得分的改写结果。由此,场景池随 Actor 一起演化,并持续针对角色-上下文空间中尚未被充分掌握的区域。我们在涵盖英文和中文的三个角色扮演基准以及我们发布的一个新的多语言基准上进行了实验,结果表明 AdvRole 始终优于基线方法。
cs.AI / 9 / 2609.28654

Training Object Permanence in World Models

在世界模型中训练客体永久性
Zhang, Haotian, Yu, Fengyuan, Luo, Dezhi, Sun, Haoran, Zhao, Zehong, Gao, Qingying, Li, Yihan, An, Siyuan, Qin, Huayi, Zhang, Yilan, Jiang, Zhengze, Feng, Pinyuan, Zhang, Renrui, Guo, Ziyu, Wang, Letian, Yang, Mengyue, Mei, Kangfu, Wang, Maijunxian, Ji, Ran, Kumar, Vikash, Shi, Freda, Sripada, Chandra, Muller, Vincent C., Torr, Philip, Yuille, Alan, Kriegeskorte, Nikolaus, Juefei-Xu, Felix, Zhang, Lvmin, Chen, Jieneng, Du, Yilun, Deng, Hokin
Abstract
Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged object permanence in them? If not, could we train them with a core-cognition inspired dataset? We introduce WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories. We build Blender generators that randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's cognitive structure, yielding 10,000+ samples per task. We release a 1.5M-sample training corpus and a 300-question exam. On this exam we evaluate 14 video models: 3 reference-to-video, 7 edit, and 4 continuation, among which PWM-WROP, our 16B world model. In a blind pairwise Elo study, PWM-WROP ranks first among continuation models and third overall, behind only a statistical tie between two reference-to-video models. We release the data, exam, model answers, scores, weights, and PWM, our native-PyTorch training stack on AWS Trainium2.
Chinese Translation
客体永久性与实体性(solidity)是人类认知先验的重要标志。近期研究表明,视频生成模型——作为当前世界模型的典型代表——已开始展现出涌现的推理能力,使其成为构建类人物理智能的理想候选。那么,视频模型是否已涌现出客体永久性?如果没有,我们能否用受核心认知启发的数据集对其进行训练?我们提出了WROP(World Reasoning with Object Permanence,基于客体永久性的世界推理),这是一个包含150个手工设计的认知科学启发式任务的数据基础设施,划分为六个认知类别。我们构建了Blender生成器,在保持每个任务认知结构的同时,随机化速度、光照、相机角度等干扰参数,每个任务生成超过10,000个样本。我们发布了包含150万个样本的训练语料库和一份包含300道题目的测试卷。在该测试卷上,我们评估了14个视频模型:3个参考图到视频(reference-to-video)模型、7个编辑(edit)模型和4个续写(continuation)模型,其中包括我们16B参数的世界模型PWM-WROP。在一项盲测成对Elo评分研究中,PWM-WROP在续写类模型中排名第一,在所有模型中排名第三,仅次于两个处于统计平手的参考图到视频模型。我们公开了数据、测试卷、模型答案、评分、模型权重,以及在AWS Trainium2上运行的基于原生PyTorch的训练框架PWM。
cs.AI / 10 / 2609.28690

Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency

超越表层风格:基于行为一致性对齐的多轮用户模拟器
Chen, Geng, Pan, Ruotong, Yang, Zhirui, He, Qiqi, Chen, Jiawei, Yunfei, Zhang, Chen, Chongyuan, Lv, Minxuan, Yang, Zheng, Huang, Win-Bin, Wu, Xiangyu, Ou, Wenwu
Abstract
Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. However, plausible individual responses do not ensure that simulated users reproduce the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories. TRACER is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. The RL stage combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation, jointly mitigating reward sparsity and credit assignment in long dialogues. On real customer-service sessions organized into reference cohorts, TRACER-7B surpasses the strongest baseline by 11.4 conversion F1, while also achieving the lowest group-level conversion-rate error and semantic trajectory distance, and generalizing to out-of-distribution scenarios. Human Turing tests yield identification accuracy close to chance, supporting the perceived naturalness of generated conversations. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion effectiveness and response quality of LLMs through simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates.
Chinese Translation
忠实的用户模拟是大规模构建、评估和改进交互式AI的基础。然而,看似合理的单条回复并不能确保模拟用户重现真实交互中所观察到的意图演化与结果。我们提出TRACER,一个显式建模用户演化意图、并学习使模拟行为与真实交互轨迹对齐的多轮用户模拟器。TRACER采用两阶段训练:先在真实用户对话上进行监督微调,随后进行多轮强化学习。RL阶段将层次化的结果级与轨迹级奖励同偏差感知的优势调制相结合,共同缓解长对话中的奖励稀疏与信用分配问题。在按参考群组组织的真实客服会话上,TRACER-7B在转化F1上超出最强基线11.4,同时取得最低的群组级转化率误差和语义轨迹距离,并能泛化到分布外场景。人类图灵测试的识别准确率接近随机水平,支持了所生成对话在感知上的自然性。基于该模拟器,我们进一步提出动态营销基准,通过模拟交互联合评估大语言模型(LLM)的说服效果与回复质量,揭示了更高的回复质量并不一定对应更高的转化率。
cs.AI / 11 / 2609.28692

Driving Epidemic Models with AI Agents: the Epydemix Agent Framework

利用AI智能体驱动流行病模型:Epydemix Agent框架
Gozzi, Nicolò, Cattuto, Ciro, Vespignani, Alessandro
Abstract
Artificial Intelligence agents based on large language models provide convenient natural language interfaces to scientific software, but reliability is not automatic. Here we introduce the Epydemix Agent Framework, an additive layer over Epydemix, an open-source Python library for stochastic compartmental epidemic modeling. The framework extends the library with four capabilities to facilitate interaction with an AI agent: discovery of available models and parameters, preventive validation of a declarative scenario specification, execution through tested library code, and inspectability of results. These capabilities let an agent handle the entire modeling process, from the natural-language description of the scenario to quantitative results, figures, and interpretation of findings without writing custom code. Each step reads input files and saves results in a separate output bundle, making the process auditable and reproducible. First, we show the end-to-end workflow with a case study comparing vaccination strategies for a novel respiratory virus. Second, we assessed the framework across 50 agent sessions and five modeling tasks by comparing the agent use of the framework against the direct use of the Python interface. The framework reduced turns, output tokens, and cost on most tasks, unless it trades resources for per-point reproducibility.
Chinese Translation
基于大语言模型的人工智能(AI)智能体为科学软件提供了便捷的自然语言接口,但其可靠性并非自动获得。本文介绍了Epydemix Agent框架,它是Epydemix(一个用于随机区室流行病建模的开源Python库)之上的附加层。该框架从四个方面扩展了该库,以便于与AI智能体交互:发现可用的模型和参数、对声明式场景规范的预防性验证、通过经过测试的库代码执行建模,以及对结果的可检查性。这些能力使智能体能够处理整个建模过程——从场景的自然语言描述到定量结果、图表和结论解释——而无需编写自定义代码。每一步都读取输入文件并将结果保存在独立的输出包中,从而使整个流程可审计、可复现。首先,我们通过一个比较新型呼吸道病毒疫苗接种策略的案例研究展示了端到端的工作流程。其次,我们通过50次智能体会话和五项建模任务,将智能体使用该框架的效果与直接使用Python接口进行了对比评估。该框架在大多数任务上减少了交互轮次、输出令牌和成本,除非它以资源换取逐点可复现性。
cs.AI / 12 / 2609.28693

Progressive Skill Discovery as Access Control for Tool-Using LLM Agents: Structural Governance through Role-Scoped Capability Delivery

作为工具使用型大语言模型智能体访问控制的渐进式技能发现:基于角色限定能力交付的结构化治理
Stettler, Michael, Girardet, Benjamin, Canton, Jonas, Corod, Nicolas
Abstract
Large Language Model (LLM) agents struggle to scale safely when exposed to vast enterprise toolsets. Providing an agent with access to every internal tool leads to oversized context windows, degraded tool selection, and severe governance vulnerabilities - as system policies defined purely in prompts remain probabilistic advice rather than hard constraints. Existing mitigations, such as multi-agent domain delegation, decentralize audit logs and fail to guarantee policy compliance across sessions. We introduce skilder, a framework that packages capabilities into roles: bundles of skills, tools, and instructions, together with the limits that bound them. An agent begins with a minimal role catalog, learns the roles a task requires, and receives each role's skills, instructions, and tools through a single MCP server. Because tools reach the agent only inside learned skills, the same server enforces the scope of what was learned deterministically. We evaluate skilder against flat-context tool selection and multi-agent orchestration across 13 tasks using six models (10 runs each). Our results show that, when models completed discovery and issued a governed call, the skilder simulated authorization layer enforced governance boundaries: no unauthorized tool call or parameter violation (e.g., a spending-limit breach) executed. Aggregate task pass rates also reflect whether each model followed the discovery protocol and satisfied response-quality checks; those misses are not authorization failures. Furthermore, by allowing agents to dynamically acquire cross-role capabilities mid-task, skilder preserves problem-solving flexibility while providing hard system-level enforcement.
Chinese Translation
当大语言模型(LLM)智能体面对庞大的企业工具集时,难以安全地进行规模扩展。赋予智能体访问所有内部工具的权限会导致上下文窗口过大、工具选择性能下降以及严重的治理漏洞——因为纯粹在提示词中定义的系统策略只是概率性建议,而非硬性约束。现有的缓解措施(如多智能体领域委派)分散了审计日志,且无法保证跨会话的策略合规性。我们提出了 skilder,一个将能力封装为角色(roles)的框架:角色是技能、工具和指令的集合,以及约束它们的边界。智能体从一个最小化的角色目录开始,学习任务所需的角色,并通过单一 MCP 服务器获取每个角色的技能、指令和工具。由于工具仅在已学习的技能内部传递给智能体,同一服务器能够确定性地强制执行所学内容的作用范围。我们使用六个模型(每个模型运行 10 次)在 13 个任务上将 skilder 与平面上下文工具选择及多智能体编排进行对比评估。结果表明,当模型完成发现过程并发出受治理的调用时,skilder 模拟的授权层成功强制执行了治理边界:没有任何未授权的工具调用或参数违规(如超出消费限额)被执行。总体任务通过率还反映了各模型是否遵循了发现协议并满足响应质量检查;这些未通过的情况并非授权失败。此外,通过允许智能体在任务执行中途动态获取跨角色能力,skilder 在提供系统级硬性强制执行的同时,保留了问题求解的灵活性。
cs.AI / 13 / 2609.28765

Reinforcement Learning with Verifiable Rewards for Small Search Agents

基于可验证奖励强化学习的小型搜索智能体
Jayadas, Gaurisankar, Plaat, Aske, Serra-Gómez, Álvaro, P, Sandheep
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) performs well on problems with clear rewards, such as mathematics and coding, but whether it also works where the reward is less clear remains open. The reason-over-search recipe applies RLVR to open-domain question answering, where retrieval grounds the answer and a match against the reference supplies the reward. So far it has been demonstrated on large models, and below one billion parameters only with distillation from a larger teacher. We test the recipe on a small model. We train Qwen3.5-0.8B with Group Relative Policy Optimization (GRPO) and an interleaved Wikipedia-search tool on MuSiQue, varying only the reward across three shapes over three seeds each, and we evaluate every checkpoint held-out on a seven-benchmark question-answering suite. The recipe works: the best run reaches 0.352 average exact match against a 0.092 untrained floor, a 3.8-fold gain, with no distillation step in the training loop. The reward shape also matters. The Search-R1-faithful exact-match-only reward is the worst of the three at every seed at the matched training horizon, and it is worst even on exact match, the metric it directly optimises. We conclude that the sparse exact-match reward, RLVR's default in mathematics and code, is the wrong starting point for models of this size. The reason-over-search setting can supply a suitable reward for RLVR on small models, but small-model RLVR needs its own reward-design study rather than a scaled-down copy of a large-model recipe.
Chinese Translation
可验证奖励强化学习(Reinforcement Learning with Verifiable Rewards, RLVR)在数学和编程等具有明确奖励的问题上表现良好,但它在奖励不够明确的问题上是否同样有效仍是一个开放问题。"先推理后搜索"(reason-over-search)方法将 RLVR 应用于开放域问答任务,其中检索为答案提供依据,而与参考答案的匹配提供奖励。目前该方法仅在大模型上得到验证,在十亿参数以下的模型上则需依赖从更大教师模型蒸馏。我们在一个小模型上测试了该方法。我们使用群体相对策略优化(Group Relative Policy Optimization, GRPO)以及交替调用的维基百科搜索工具,在 MuSiQue 数据集上训练 Qwen3.5-0.8B,仅在三种奖励形式上进行变化,每种设置训练三个随机种子,并在七个基准组成的问答套件上评估每一个留出(held-out)检查点。该方法有效:最佳运行的平均精确匹配率达到 0.352,而未经训练的基线仅为 0.092,提升了 3.8 倍,且训练过程中无需蒸馏步骤。奖励形式同样至关重要。在与训练步数匹配的条件下,忠实于 Search-R1 的仅精确匹配奖励在所有种子下都是三者中最差的,甚至在精确匹配这一它直接优化的指标上也表现最差。我们由此得出结论:稀疏的精确匹配奖励——RLVR 在数学和代码领域的默认选择——并不适合这一规模的模型。"先推理后搜索"场景可以为小模型的 RLVR 提供合适的奖励,但小模型的 RLVR 需要针对其自身的奖励设计研究,而非简单照搬大模型方法的缩小版本。
cs.AI / 14 / 2609.28771

Agent Memory with Episodic Retrieval for Financial Decision-Making

面向金融决策的情节记忆检索智能体记忆机制
Xu, Nuoyue, Liu, Jiang, Huang, Wenxuan, Zhang, Xiang, Cao, Juntai, Wei, Jiaqi
Abstract
Large language models (LLMs) have demonstrated strong capabilities in financial analysis and reasoning, inspiring recent advances in agent-based trading frameworks. While these systems show promise, prior approaches either emphasize long-horizon forecasting or operate as stateless analyzers, limiting their applicability to the demands of trading in complicated settings. To address these gaps, we introduce META (Memory Enhanced Trading Agent), the first RAG-like episodic-memory-augmented multi-agent framework for financial decision making. META integrates a family of specialized indicator agents (e.g., Trend, MACD, Stochastic, RSI, SMA, AVWAP, Heikin-Ashi) with a Decision Agent that fuses their reports, and a Memory module that retrieves and updates past trading episodes encoded as market state embeddings with outcomes and reflections. By recalling relevant experiences and adaptively reweighting signals under similar market regimes, META achieves improved directional accuracy and robustness under short-horizon evaluation. Our results demonstrate that episodic memory provides a powerful mechanism for regime-aware, interpretable, and low-latency decision-making in trading and decision making. The code of this project is released on GitHub.
Chinese Translation
大语言模型(LLMs)在金融分析与推理方面展现出强大的能力,推动了近期基于智能体的交易框架的发展。尽管这些系统前景可观,但现有方法要么侧重于长期预测,要么作为无状态的分析器运行,这限制了它们在复杂交易场景中的应用。为弥补这些不足,我们提出了META(Memory Enhanced Trading Agent,记忆增强交易智能体),这是首个将检索增强生成(RAG)式的情节记忆增强技术应用于金融决策的多智能体框架。META整合了一系列专门的指标智能体(如趋势、MACD、随机指标、RSI、SMA、AVWAP、Heikin-Ashi),由一个决策智能体(Decision Agent)融合它们的报告;同时配备一个记忆模块(Memory),该模块检索并更新以市场状态嵌入形式编码、并附带交易结果与反思的历史交易情节。通过回忆相关经验并在相似市场状态下自适应地重新加权信号,META在短期评估中实现了更高的方向预测准确性和鲁棒性。我们的结果表明,情节记忆为交易与决策中的状态感知、可解释且低延迟的决策提供了一种强大的机制。本项目的代码已在GitHub上发布。
cs.AI / 15 / 2609.28776

Learned Cross-Task Relationships in Multi-Task Models

多任务模型中学习到的跨任务关系
Zhang, Victor, Yuan, Yiping, Raudies, Florian, Adeoti, Bosun, Leung, Brian Y. C., Girija, Sanjay Surendranath, Zhang, Naijing
Abstract
We propose a framework that learns cross-task relationships in multi-task models by approximating the joint distribution of task labels through targeted pairwise relationships. This approach improves performance via transfer learning and enhances information extraction without the intractable complexity of modeling the full joint space. Although our framework applies to any multi-task system, we demonstrate its efficacy within YouTube's production recommendation systems. Experiments across the Notifications, Homepage, and Watch Next surfaces show improvements in both accuracy and user satisfaction metrics. Finally, we propose a workflow template to facilitate broader future implementation.
Chinese Translation
我们提出了一种框架,通过目标性的成对关系来近似任务标签的联合分布,从而学习多任务模型中的跨任务关系。该方法通过迁移学习提升性能,并增强信息提取能力,同时避免了建模完整联合空间所带来的难以处理的复杂度。尽管我们的框架适用于任何多任务系统,但我们在YouTube的生产级推荐系统中验证了其有效性。在通知、首页以及观看下一个(Watch Next)等多个推荐场景中的实验表明,该框架在准确性和用户满意度指标上均取得了提升。最后,我们提出了一个工作流模板,以促进未来更广泛的应用。
cs.AI / 16 / 2609.28850

RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?

RECLAIM:智能体能否复现机器学习论文中的结论?
Salunkhe, Mithil, Ding, Haochen, Verma, Samridhi, Kindratenko, Volodymyr
Abstract
Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do. We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper we fix in advance the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget. An agent must reproduce that result using the paper and whatever its authors released. What the authors released decides the difficulty tier. Run-tier releases include code, data, and weights; Retrain-tier releases lack weights, so the agent trains the model; Reimplement-tier releases lack code, so the agent writes it. A separate language model grades runs from logs and outputs rather than agents' reports. We run four agents once per paper; the best agent in each tier reproduces only 41% of Run-tier papers, 27% at Retrain, and 15% at Reimplement, where every agent does worst. Failed attempts use on average 29% of their budget, so most stop with budget left. The most common agent error is writing the method without checking any part against the paper's numbers, in 63 of 400 runs.
Chinese Translation
复现一篇机器学习论文涉及大部分研究步骤,从安装软件、调试到运行实验,而这些正是AI智能体日益承担的工作。我们提出RECLAIM,一个包含100篇NeurIPS 2025论文的基准,且每年可从新会议中重建。对于每篇论文,我们预先固定要复现的结果、判定复现成功的标准以及GPU小时预算。智能体必须利用论文原文及作者发布的任何材料来复现该结果。作者发布的材料决定了难度级别:Run级发布包含代码、数据和权重;Retrain级缺少权重,因此智能体需自行训练模型;Reimplement级缺少代码,因此智能体需自行编写。由一个独立的语言模型根据日志和输出(而非智能体的报告)对运行结果进行评分。我们对四个智能体各运行一次每篇论文;每个级别中表现最佳的智能体仅能复现41%的Run级论文、27%的Retrain级论文和15%的Reimplement级论文,后者是所有智能体表现最差的级别。失败的尝试平均只使用了其预算的29%,说明大多数尝试在预算尚有剩余时便停止了。最常见的智能体错误是在未将任何部分与论文中的数值进行核对的情况下就编写方法实现,在400次运行中出现了63次。
cs.AI / 17 / 2609.28859

Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation

人机协同的假设检验:成本感知的选择性AI评分与序贯式人工升级
Woong, Dae, Ham, Zhao, Xuejun, Jasin, Stefanus, Yang, Fenghua
Abstract
Large language models are increasingly used as inexpensive judges to evaluate outputs, label data, and assess whether a system meets a desired quality standard. Yet using AI judgments for formal statistical inference is fundamentally different from simply treating them as ground-truth labels: AI evaluations can be biased or noisy, and rigorous hypothesis testing requires explicit control of type-I and type-II errors. We study how to use AI judgments, together with selective human verification, to conduct a valid hypothesis test at minimum cost. We consider a population of items with hidden binary labels. After choosing a fixed pool of items, the decision maker can selectively query AI, send an item directly to a human, escalate an AI-scored item to a human after observing the AI report, or stop once sufficient evidence has accumulated. We derive an information-theoretic lower bound that captures the minimum cost of achieving prescribed testing errors and characterizes the value of AI information and human verification through a report-dependent information frontier. Motivated by this characterization, we develop SCALE, a sequential cost-aware policy that combines selective AI scoring with adaptive human escalation. SCALE is valid at finite sample sizes and matches the lower bound to first order as the target error probabilities vanish. We further extend the framework to an unknown AI-output model using paired AI-human pilot data. Numerically, SCALE approaches Human-only or AI-only testing when one source clearly dominates, while achieving its largest savings when inexpensive AI judgments and selective human verification are both valuable.
Chinese Translation
大语言模型正越来越多地被用作低成本的评判者,以评估输出结果、标注数据以及判断系统是否达到预期的质量标准。然而,将AI判断用于正式的统计推断,与简单地将其视为真值标签有着本质区别:AI评估可能存在偏差或噪声,而严格的假设检验要求对第一类错误和第二类错误进行显式控制。我们研究如何利用AI判断,并结合选择性的人工验证,以最低成本进行有效的假设检验。我们考虑一个由具有隐藏二元标签的个体组成的总体。在选定固定的个体池后,决策者可以选择性地查询AI,将某个个体直接送交人工,在观察AI报告后将AI评分过的个体升级交由人工处理,或在积累到足够证据后停止。我们推导出一个信息论下界,刻画了达到预定检验误差的最小成本,并通过一个依赖报告的信息前沿刻画了AI信息与人工验证的价值。基于这一刻画,我们提出了SCALE——一种序贯式的成本感知策略,它将选择性AI评分与自适应的人工升级相结合。SCALE在有限样本量下是有效的,并且随着目标错误概率趋于零,其一阶性能与下界相匹配。我们进一步将框架扩展至AI输出模型未知的情形,利用配对的AI-人工试点数据进行处理。数值实验表明,当某一方明显占优时,SCALE接近于纯人工或纯AI检验的表现;而当低成本的AI判断与选择性人工验证二者都具价值时,SCALE能实现最大的成本节约。
cs.AI / 18 / 2609.28876

Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents

Forecast-Dojo:用于基准测试和训练大语言模型预测代理的可回放环境
Ye, Liqin, Wang, Haorui, Ahmed, Fardin, Zhang, Rongzhi, He, Yuan, Lin, Ziyuan, Yin, Yanbin, Peng, Jing, Galarnyk, Michael, Chava, Sudheer, Zhang, Chao
Abstract
We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by time into training and evaluation periods, and 18.8M dated news articles. In an evaluation of 12 models, research tools lower Brier score for all 12. Forecasts also improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. Every model still trails historical market forecasts in both Brier score and accuracy. A belief notebook carried between dates lowers research cost but does not consistently improve forecast quality. Beyond evaluation, Forecast-Dojo provides interaction trajectories and outcome feedback for agent learning, with supervised fine-tuning as a proof of concept.
Chinese Translation
我们提出了 Forecast-Dojo,一个用于基准测试和训练大语言模型(LLM)预测代理的可回放环境。它将已有结果的市场预测问题与按日期标注的新闻相结合,使代理能够研究某一事件,并在连续的历史时间节点上回顾和更新其预测。相同的任务和工具支持重复评估、训练交互数据的收集,以及基于已记录结果的反馈,而无需等待新事件揭晓。Forecast-Dojo 包含 1,568 个 Polymarket 事件(按时间划分为训练期和评估期)以及 1880 万篇按日期标注的新闻文章。在对 12 个模型的评估中,研究工具降低了全部 12 个模型的 Brier 分数。随着事件的推进,预测质量也不断提升,其中在新记录的带日期证据较多的时间步骤上提升最为显著。然而,所有模型在 Brier 分数和准确率上仍落后于历史市场预测。跨时间节点携带的信念笔记本(belief notebook)降低了研究成本,但并未持续提升预测质量。除评估之外,Forecast-Dojo 还为代理学习提供了交互轨迹和结果反馈,并以监督微调作为概念验证。
cs.AI / 19 / 2609.28919

Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise

掌控套件,掌控成本:企业中AI编程代理的路由与治理
Abbasi, Arian, Aqrawi, Alan, Kwartler, Ted
Abstract
Harnesses, the products that run AI coding agents, are multiplying, and enterprises are rolling them out to their employees: what started as pilots with a few hundred seats is scaling to tens of thousands. Most enterprises do not build these harnesses but buy them from large vendors, such as Anthropic's Claude Code or OpenAI's Codex. A harness decides which model answers, what the model reads, how the prompt cache is used and which subagents run, so it picks the rate on the price sheet and sets the volume bought at it. Enterprises that keep a proprietary or untuned harness at its defaults inherit these choices and their bill. We build a fast, customisable router in which Jev, a classifier with calibrated probabilities, labels every prompt against a bring-your-own taxonomy of agentic requests. Because one user turn is many requests over a prompt cache that belongs to one model, the router moves work only where no running conversation has to rebuild its cache: at session start, in side lanes and at subagent launch. From the price sheet we derive when a mid-task switch pays back, and a crossover: on long tool-heavy sessions the highest-priced model costs less than the next tier, as repricing about 10,000 real sessions from public datasets confirms. In an emulated enterprise of 10,000 seats with user behaviour taken from these datasets, the router recovers 14 to 21% of model spend at Anthropic's list prices of 21 September 2026, \$3.3M to \$5.0M a year. The paper also maps the risks across twenty harnesses, prices the dependence on one vendor's models, and proposes a control plane that enterprises can run from within, starting now, with a ladder for deciding later whether to own the harness.
Chinese Translation
运行AI编程代理的套件(Harness)产品正不断涌现,企业也正在向员工推广它们:从最初几百个席位的试点正在扩展到数万个席位。大多数企业并不自行构建这些套件,而是从大型供应商处购买,例如Anthropic的Claude Code或OpenAI的Codex。套件决定了由哪个模型回答、模型读取什么内容、提示缓存(prompt cache)如何使用以及运行哪些子代理(subagent),因此它决定了价目表上的费率以及在该费率下的用量。沿用专有或未调优套件默认设置的企业,就会继承这些选择及其账单。我们构建了一个快速、可定制的路由器,其中Jev——一个具有校准概率的分类器——根据可自定义的代理请求分类体系对每条提示进行标注。由于一个用户回合包含针对属于单一模型的提示缓存的多个请求,路由器仅在不需重建运行中会话缓存的位置转移工作:会话开始时、侧通道中以及子代理启动时。我们根据价目表推导出任务中途切换何时能够回本,并发现一个交叉点:在长且工具密集的会话中,价格最高的模型反而比次一级的模型成本更低,这一点通过对来自公开数据集的约10,000个真实会话重新定价得到验证。在一个从这些数据集中提取用户行为的10,000席位模拟企业中,按Anthropic 2026年9月21日的定价计算,该路由器可节省14%至21%的模型支出,即每年330万至500万美元。本文还梳理了二十种套件的风险,量化了对单一供应商模型的依赖成本,并提出一个企业可从内部运行的控制平面——现在即可启用,同时提供一个阶梯式框架,以便日后决定是否自建套件。
cs.AI / 20 / 2609.28921

PFArena: Benchmarking Language Models for Protein Modification

PFArena:面向蛋白质改造的语言模型基准测试
Ouyang, Yawen, Zhang, Xinbo, Ma, Ziyuan, Wu, Yixin, Liao, Wenbin, Zhang, Feiran, Li, Wenjie, Wang, Lihao, Wang, Hao, Zheng, Xiaoqing, Yan, Xuefeng, Bai, Lei, Zhang, Ya-Qin, Zhang, Shuyi, Ma, Wei-Ying, Lin, Dahua, Zhou, Bowen, Zhou, Hao
Abstract
Protein modification requires navigating an immense sequence space, yet wet-lab validation remains low-throughput and costly. Although computational paradigms including protein language models (PLMs), large language models (LLMs), and LLM-based agents have shown promise in protein modification, their relative efficacy across realistic experimental decision-making settings remains unclear. To bridge this gap, we introduce PFArena, a benchmark comprising four controlled task interfaces that cover single-mutant generation and multi-mutant ranking. By providing varying levels of mutation fitness data, PFArena reflects four representative research scenarios characterized by differing degrees of prior experimental context. We assess six PLMs, six LLMs, and five LLM-based agents using complementary metrics to measure both peak and overall protein modification performance. Our evaluation reveals that model performance shifts systematically with the availability of target-specific experimental evidence: PLMs demonstrate proficiency in open-ended single-mutant generation by leveraging protein-specific priors, whereas LLMs and agents perform strongly in multi-mutant ranking, particularly when target-specific fitness data are available. Nevertheless, all model families face fundamental challenges with increasing search-space size and mutation depth. We release our code and benchmark suite to facilitate reproducible research in model-assisted protein modification.
Chinese Translation
蛋白质改造需要在巨大的序列空间中进行搜索,然而湿实验验证仍然通量低且成本高昂。尽管蛋白质语言模型(PLM)、大语言模型(LLM)以及基于LLM的智能体(agent)等计算范式在蛋白质改造中已展现出潜力,但它们在真实实验决策场景下的相对效果仍不明确。为填补这一空白,我们提出了PFArena,一个由四个受控任务接口组成的基准测试,涵盖单突变体生成与多突变体排序。通过提供不同水平的突变适应性(fitness)数据,PFArena反映了四种具有不同先验实验信息程度的代表性研究场景。我们评估了六个PLM、六个LLM和五个基于LLM的智能体,采用互补的指标来衡量峰值与整体的蛋白质改造性能。评估结果表明,模型性能会随靶标特异性实验证据的可获得性而系统性变化:PLM凭借蛋白质特异性先验知识,在开放式的单突变体生成中表现出色;而LLM和智能体在多突变体排序中表现强劲,尤其是在有靶标特异性适应性数据可用时。尽管如此,随着搜索空间规模和突变深度的增加,所有模型家族都面临根本性的挑战。我们公开了代码和基准测试套件,以促进模型辅助蛋白质改造领域的可复现研究。
cs.AI / 21 / 2609.28942

From Static Personal Values to Contextualized Personalization: Bayesian Personalized Value Alignment for LLMs

从静态个人价值观到情境化个性化:面向大语言模型的贝叶斯个性化价值观对齐
Guo, Hanze, Song, Aixuan, Yao, Jing, Zhang, Xiangxu, Yi, Xiaoyuan, Xie, Xing, Zhou, Xiao
Abstract
Personalized value alignment has become increasingly important as large language models (LLMs) are expected to accommodate diverse user preferences. However, existing methods typically align model outputs with a static value profile across prompts, overlooking that the salience of value dimensions varies substantially across contexts. Inspired by Lewin's Field Theory, which views human behavior as jointly shaped by personal dispositions and situational constraints, we model personal values as priors and context-dependent preferences as posteriors. We propose BaCVA, an inference-time Bayesian Context-aware personalized Value Alignment method that approximates posterior personalized preferences by integrating static personal values with scenario-specific value salience. BaCVA first estimates contextual value salience from generally normative responses, and then employs a dual-view personalization module to infer posterior preferences from complementary personal-value and scenario-driven perspectives. This Bayesian formulation enables more accurate and adaptive personalized value alignment while improving data efficiency via prior values. Extensive experiments on benchmarks demonstrate its superiority over strong baselines.
Chinese Translation
随着大语言模型(LLMs)被期望能够满足多样化的用户偏好,个性化价值观对齐变得日益重要。然而,现有方法通常在不同提示(prompt)下将模型输出与静态的价值观画像(value profile)对齐,忽视了价值维度在不同情境下的显著性存在显著差异这一事实。受勒温(Lewin)场论的启发——该理论认为人类行为是由个人倾向与情境约束共同塑造的——我们将个人价值观建模为先验,将情境依赖的偏好建模为后验。我们提出了 BaCVA,一种推理阶段的贝叶斯情境感知个性化价值观对齐(Bayesian Context-aware personalized Value Alignment)方法,该方法通过整合静态个人价值观与特定场景下的价值显著性来逼近后验个性化偏好。BaCVA 首先从一般规范性回复中估计情境化的价值显著性,然后采用双视角个性化模块,从互补的个人价值观视角和场景驱动视角推断后验偏好。这种贝叶斯建模方式实现了更准确、更具适应性的个性化价值观对齐,同时借助先验价值观提升了数据效率。在多个基准上的大量实验证明了其相对于强基线方法的优越性。
cs.AI / 22 / 2609.28963

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

回归定义:基于轨迹图的智能体强化学习步骤级优势估计
Yao, Xincheng, Fu, Haobo, Liu, Weiming, Zhang, Chongyang
Abstract
Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normalized advantage estimation is reliable at the response level, it becomes systematically biased at the step level, since coarse-grained trajectory-level advantages are hard to accurately reflect the contribution of individual steps (i.e, failed trajectories may contain valuable steps). Revisiting the foundational RL definition, we notice that GRPO's success on single-turn tasks stems from its advantage estimation strategy, which adheres to the basic definition: the mean reward of multiple actions sampled from the same state constitutes a credible state-value estimate. Extending the faithful estimation to step-level would in principle demand sampling multiple actions from each intermediate state, which is too costly on a per-state basis. To mitigate this issue, we propose a Graph-based Faithful sTep-level credit-assignment framework (GRAFT) that grafts all rollout trajectories into a trajectory graph, recovering node state-values via Bellman iteration on the graph, and assigning credit to each edge by the node value difference. Theoretically, the estimated step-level advantage faithfully adheres to the basic advantage definition in RL. To further ensure the reliability of step-level advantage estimation, we further propose Graph GAE, which extends GAE to the trajectory graph for reducing the impact of state-value estimation bias. Experiments across a range of multi-turn agentic benchmarks show consistent gains over GRPO and superior performance compared to recent agentic RL algorithms. Code will be available at https://github.com/xcyao00/GRAFT.
Chinese Translation
基于组的强化学习方法(如GRPO及其变体)已成为训练推理型和智能体型大语言模型的主流范式。尽管其组归一化优势估计在响应层面是可靠的,但在步骤层面却存在系统性偏差,因为粗粒度的轨迹级优势难以准确反映单个步骤的贡献(即失败的轨迹中可能包含有价值的步骤)。通过重新审视强化学习的基础定义,我们注意到GRPO在单轮任务上的成功源于其优势估计策略符合基本定义:从同一状态采样的多个动作的平均奖励构成了可信的状态值估计。将这种忠实的估计扩展到步骤层面,原则上需要从每个中间状态采样多个动作,但这在逐状态的基础上代价过高。为缓解这一问题,我们提出了一种基于图的忠实步骤级信用分配框架(GRAFT),该框架将所有 rollout 轨迹嫁接到一个轨迹图中,通过图上的贝尔曼迭代恢复节点状态值,并根据节点值之差为每条边分配信用。从理论上讲,所估计的步骤级优势忠实地遵循强化学习中优势的基本定义。为进一步保证步骤级优势估计的可靠性,我们进一步提出了 Graph GAE,将GAE扩展到轨迹图上,以降低状态值估计偏差的影响。在多个多轮智能体基准上的实验表明,该方法相较于GRPO取得了一致的提升,并优于近期的智能体强化学习算法。代码将发布于 https://github.com/xcyao00/GRAFT。
cs.AI / 23 / 2609.29007

When Does Action Credit Need Updating?

动作信用何时需要更新?
Yang, Hongye, Huang, Boxiao
Abstract
Tool-using agents are continually updated with new interaction data. After each policy update, however, previously estimated action credits may become stale. Recomputing them from scratch can require many additional tool calls and environment interactions, making repeated updates increasingly expensive. We ask a simple question: when does historical action credit actually need to be updated? Our key observation is that a change in action value does not necessarily imply a change in the decision. Historical credit can still be useful as long as policy-induced drift is too small to overturn the existing action ranking. Building on this idea, we introduce pairwise branch sensitivity to capture how strongly a policy update affects the downstream regions that distinguish two candidate actions. We then derive a first-order anchored credit-transport estimator that updates historical credit using old interventional trajectories, and propose a Decision-Sufficient Credit Gate (DSC-Gate) that chooses whether to reuse, transport, or resample credit. Experiments show that branch sensitivity explains credit drift substantially better than global policy distance. With sufficient historical data, credit transport reduces estimation error, while its benefit to decision making is concentrated on updates that affect action-distinguishing branches. On a fully independent test set, DSC-Gate changes mean regret by only +0.00004 relative to a gap-based gate while reducing mean new tool steps from 472 to 286, a 39.4% reduction. We observe the same pattern after a real tool-agent parameter update. Overall, our results show that agents do not need to recompute action credit after every policy update: much of the historical evidence can be reused or cheaply corrected, reducing the additional interaction required to keep action decisions up to date.
Chinese Translation
使用工具的智能体会随着新的交互数据不断更新。然而,每次策略更新后,先前估计的动作信用可能会变得过时。从头重新计算这些信用需要大量额外的工具调用和环境交互,使得反复更新的成本越来越高。我们提出一个简单的问题:历史的动作信用究竟何时需要更新?我们的关键观察是:动作价值的改变并不必然意味着决策的改变。只要策略引起的漂移不足以推翻现有的动作排序,历史信用仍然有用。基于这一思想,我们引入成对分支敏感度,用以刻画策略更新对区分两个候选动作的下游区域的影响强度。随后,我们推导了一个基于一阶锚定的信用迁移估计器,利用旧的干预轨迹来更新历史信用,并提出了一个决策充分信用门控,用于决定是复用、迁移还是重新采样信用。实验表明,分支敏感度对信用漂移的解释能力显著优于全局策略距离。在历史数据充足的情况下,信用迁移能降低估计误差,而其对决策的收益则集中于影响动作区分分支的更新。在一个完全独立的测试集上,DSC-Gate 相对于基于间隔的门控仅使平均遗憾增加 0.00004,同时将平均新增工具步数从 472 降至 286,减少了 39.4%。在真实的工具智能体参数更新后,我们观察到了相同的规律。总体而言,我们的结果表明,智能体无需在每次策略更新后都重新计算动作信用:大量历史证据可以被复用或低成本地修正,从而减少保持动作决策最新所需的额外交互。
cs.AI / 24 / 2609.29014

AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor Mining

AlphaDiverse:通过后训练本地量化研究智能体实现Alpha因子挖掘中的多样化探索
Wang, Qingzhuo, Wei, Zikun, Wei, Zhihua, Shen, Wen
Abstract
Large language model (LLM)-based multi-agent systems can automate alpha factor mining, but their reliance on external APIs limits control over cost, availability, and confidentiality. Long research loops also tend to revisit a few successful economic mechanisms that lead to research path collapse. To address these limitations, we propose AlphaDiverse, a framework that integrates a multi-agent alpha research system, diverse research path collection, and post-training for local agents. We let the research system generate complementary plan portfolios and vary research environments across loops to collect diverse research paths. Using these diverse traces, we warm-start local Planner and Realizer agents with supervised fine-tuning. Then, we propose a joint GRPO method to optimize both of them using predictive quality and diversity of contributions. Research feedback is confined to inner period data, while a frozen final model is evaluated on a later outer period data, thereby avoiding test-set tuning. Experiments across four Chinese stock universes show that AlphaDiverse can combine competitive prediction with broader exploration.
Chinese Translation
基于大语言模型(LLM)的多智能体系统可以实现Alpha因子挖掘的自动化,但其对外部API的依赖限制了对成本、可用性和机密性的控制。同时,冗长的研究循环往往反复使用少数成功的经济机制,导致研究路径坍缩。为解决这些局限,我们提出了AlphaDiverse,该框架集成了多智能体Alpha研究系统、多样化研究路径收集以及针对本地智能体的后训练。我们让研究系统生成互补的计划组合,并在各研究循环中变换研究环境,以收集多样化的研究路径。利用这些多样化轨迹,我们通过监督微调对本地Planner(规划器)和Realizer(执行器)智能体进行热启动。随后,我们提出一种联合GRPO方法,利用预测质量和贡献多样性对二者同时进行优化。研究反馈仅限于内部时段数据,而冻结后的最终模型在更晚的外部时段数据上进行评估,从而避免了对测试集的调参。在四个中国股票域上的实验表明,AlphaDiverse能够在保持有竞争力的预测能力的同时实现更广泛的探索。
cs.AI / 25 / 2609.29015

MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks

MeshHeal:面向去中心化LLM智能体网络中灰色故障的双时间尺度自愈机制
Chen, Keru, Lin, Sen, Liang, Yingbin, Bastian, Nathaniel D., Zou, Shaofeng
Abstract
Decentralized LLM-based multi-agent systems coordinate through local interactions, but an agent can remain responsive while its task-solving quality persistently degrades. Such gray failures require protecting current tasks before sufficient evidence exists to alter future routing, while still allowing recovered agents to rejoin. We introduce MeshHeal, a fully decentralized self-healing framework that couples ability-matched peer review across two timescales. At the fast timescale, an adaptive hierarchy escalates uncertain or low-scoring outputs from repeated single-reviewer evaluation to committee deliberation and, when needed, correction before use. At the slow timescale, a task- and ability-conditioned peer-relative detector aggregates scores to distinguish persistent degradation from ordinary output variation, trigger mandatory committee review, and eventually exclude degraded agents from ordinary routing; recovery probes provide fresh evidence for reintegration. To faithfully evaluate routing, we introduce Model-Backed MAS Evaluation, which ties ability assignments to execution models, since prompt-based ability assignments alone can leave routing errors hidden. Across BBH, MATH, and MMLU-Pro, MeshHeal achieves 0.839 degraded-phase accuracy using 51k total model tokens per task, versus the strongest baseline Symphony's 0.807 accuracy using 115k per task. Under staggered degradation and recovery, MeshHeal isolates degraded agents, keeps them excluded from ordinary task execution until recovery, and returns them to normal routing.
Chinese Translation
基于LLM的去中心化多智能体系统通过本地交互进行协调,但某个智能体可能在保持响应的同时,其任务求解质量持续退化。这类灰色故障要求在尚无充分证据以改变未来路由之前保护当前任务,同时仍允许已恢复的智能体重新加入。我们提出了MeshHeal,一个完全去中心化的自愈框架,它在两个时间尺度上耦合能力匹配的同行评审。在快时间尺度上,自适应层级机制将不确定或低分的输出从重复的单评审员评估升级为委员会审议,并在需要时于使用前进行修正。在慢时间尺度上,一个基于任务和能力条件的同伴相对检测器聚合评分,以区分持续性退化与普通输出波动,触发强制性委员会评审,并最终将退化智能体从普通路由中排除;恢复探测为其重新接入提供新的证据。为忠实评估路由,我们引入了Model-Backed MAS Evaluation(模型支撑的多智能体系统评估),将能力分配与执行模型绑定,因为仅基于提示的能力分配可能使路由错误被隐藏。在BBH、MATH和MMLU-Pro数据集上,MeshHeal在退化阶段达到0.839的准确率,每任务消耗总计51k模型token,而最强基线Symphony的准确率为0.807,每任务消耗115k token。在交错式退化与恢复场景下,MeshHeal能够隔离退化智能体,在其恢复之前将其排除在普通任务执行之外,并使其回归正常路由。
cs.AI / 26 / 2609.29050

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

SLCA-GRPO:解决工具调用强化学习中的跨片段信用误归属问题
Zhan, Yan, Liu, Shaobo, Liu, Qiunan, Shi, Yuanjun, Xu, Siqi, Hou, WeiYi, Xu, Xiang, Li, Zekang, Pan, Weizhou, Yan, Jiahong
Abstract
Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural language summaries. This output heterogeneity presents a structural failure mode in standard on-policy Reinforcement Learning (RL): algorithms like GRPO indiscriminately broadcast a homogeneous trajectory-level scalar advantage to all tokens. Consequently, gradient noise from summary generation leaks into tool-decision tokens, causing cross-segment credit misattribution and brittle optimization. In this work, we propose SLCA-GRPO, a framework incorporating Segment-Locked Credit Assignment (SLCA). To enable scalable exploration without costly real APIs and stable training, we first construct the Schema-Guided LLM Simulator (SGLS) as foundational training infrastructure. Building on this, SLCA decouples advantage estimation at the structural segment level within a single group of rollouts, without requiring additional rollouts from intermediate states. Supported by Hierarchical Rewards (HierR), SLCA routes execution advantages to tool tokens and preference advantages to summary tokens, eliminating advantage contamination (the dominant cross-segment credit misattribution channel) within each policy update. On a 7B backbone, SLCA-GRPO accelerates convergence and outperforms standard GRPO, ToolPO, and RLTR by +2.53 pp on in-domain evaluation, +1.36 pp on the Berkeley Function-Calling Leaderboard (BFCL), and +9.15 pp on $\tau^2$-Bench under the same training budgets, achieving higher accuracy with reduced tool redundancy and costs.
Chinese Translation
工具调用智能体产生的输出具有异构性,将结构化的工具调用与面向用户的自然语言总结交织在一起。这种输出异构性在标准的同策略强化学习(RL)中呈现出一种结构性失效模式:诸如GRPO之类的算法会不加区分地将同质的轨迹级标量优势广播给所有词元。由此,来自总结生成的梯度噪声会泄漏到工具决策词元中,造成跨片段信用误归属和脆弱的优化过程。在本工作中,我们提出了SLCA-GRPO,一个融合片段锁定信用分配(Segment-Locked Credit Assignment, SLCA)的框架。为了在不依赖昂贵真实API的情况下实现可扩展的探索并保证训练的稳定性,我们首先构建了模式引导的大语言模型模拟器(Schema-Guided LLM Simulator, SGLS)作为基础训练基础设施。在此基础上,SLCA在单组 rollout 内以结构片段级别解耦优势估计,而无需从中间状态进行额外的 rollout。在分层奖励(Hierarchical Rewards, HierR)的支持下,SLCA将执行优势路由至工具词元,将偏好优势路由至总结词元,从而在每次策略更新中消除优势污染(跨片段信用误归属的主要渠道)。在7B骨干模型上,SLCA-GRPO加速了收敛,并在相同训练预算下,在域内评估中超越标准GRPO、ToolPO和RLTR +2.53个百分点,在伯克利函数调用排行榜(Berkeley Function-Calling Leaderboard, BFCL)上超越 +1.36个百分点,在 $\tau^2$-Bench 上超越 +9.15个百分点,在减少工具冗余和成本的同时实现了更高的准确率。
cs.AI / 27 / 2609.29051

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

从自蒸馏到自练习:面向多轮智能体的特权信息方法
Su, Xingyu, Kumar, Abhishek, Ping, Qing, Luo, Youzhi, Buck, Jonathan, Zhang, Zach, Chidambaram, Subramanian, Arannil, Vinayak
Abstract
On-policy self-distillation (OPSD) has become a popular recipe for post-training LLM agents. It supervises the agent model at the token level with a stronger teacher view of the same model, obtained by conditioning on privileged information (PI). In this work, we show that in multi-turn agents, this paradigm teaches the student to act with confidence but without the information behind it. The trained agent behaves as if it had privileged information it never observed, and its performance falls well short of plain RL, in the worst case below the untrained base model. Therefore, we propose Privileged Self-Practice (PSP), which keeps the PI and moves it from the loss to the sampler. When the student's rollouts on a task mostly fail, we inject a short per-task instruction written by an analyzer model, sample the task again with the instruction in context, and train on the result with an unchanged GRPO objective. The privileged information stays in the prompt and never enters the loss. Across AppWorld and SWE-bench Verified, with three different student models, PSP obtains the best average score in every setting and is the only method that consistently outperforms plain GRPO, improving task-goal completion by up to 65% on AppWorld and the resolved rate by up to 61% on SWE-bench Verified.
Chinese Translation
在策略自蒸馏(On-Policy Self-Distillation, OPSD)已成为大语言模型(LLM)智能体后训练的一种流行方法。它通过让同一模型以特权信息(Privileged Information, PI)为条件,得到更强的教师视角,从而在词元级别对智能体模型进行监督。在本工作中,我们发现在多轮智能体场景下,这一范式教会学生模型表现出自信的行动,但其背后却缺乏相应的信息。训练后的智能体的行为仿佛它拥有从未观察到的特权信息,其性能远逊于普通的强化学习(RL),最坏情况下甚至低于未经训练的基础模型。为此,我们提出特权自练习(Privileged Self-Practice, PSP),它保留特权信息,但将其从损失函数转移到采样器中。当学生在某任务上的推理 rollout 大多失败时,我们注入一段由分析模型撰写的简短任务级指令,在上下文中包含该指令的条件下重新采样该任务,并用未作修改的 GRPO 目标对结果进行训练。特权信息始终保留在提示(prompt)中,从不进入损失函数。在 AppWorld 和 SWE-bench Verified 上,使用三个不同的学生模型,PSP 在所有设置中均获得最佳平均分数,并且是唯一在所有情况下都持续优于普通 GRPO 的方法,在 AppWorld 上将任务目标完成率提升最高达 65%,在 SWE-bench Verified 上将解决率提升最高达 61%。
cs.AI / 28 / 2609.29075

CRISS: A Retrieval-Augmented AI Chatbot for Assisting Cancer Registrars

CRISS:一种用于辅助癌症登记员的检索增强AI聊天机器人
Seth, Vani, Beheshti, Mohammad, Kambhampati, Anirudh, Bhayani, Vishwa, Ham, Lucinda, Calyam, Prasad, Zachary, Iris
Abstract
Cancer registrars, including Oncology Data Specialists (ODSs), must interpret complex and frequently updated coding and staging standards. We developed CRISS (Cancer Registry Intelligent Support System), a retrieval-augmented generation (RAG) conversational assistant that provides rapid, citation-supported access to registry guidance. This study evaluated whether CRISS could (1) support accurate and citation-supported responses, (2) improve access to and interpretation of relevant guidance, and (3) support training/helpdesk use while preserving human oversight of final abstraction decisions. We built a domain-specific knowledge base from national cancer registry standards, segmented into metadata-tagged passages and indexed as dense embeddings. Retrieved passages were used to generate citation-grounded responses through a large language model (LLM). Open-weight, proprietary, and non-RAG baseline models across Gemini and GPT families were evaluated on easy, medium, and hard registry questions using an LLM-as-a-Judge protocols. RAG configurations consistently outperformed non-RAG approaches, especially as question difficulty increased. Mean grounding scores for RAG were 0.62/0.56/0.59 across easy/medium/hard tiers versus 0.29/0.26/0.29 for non-RAG. RAG models also achieved higher semantic-similarity scores overall. Proprietary RAG models performed strongest on easy and medium questions, while local RAG models ranked highest on hard questions and proprietary models were generally more cautious. Domain-specific RAG improved evidence grounding and response quality for cancer registry questions while enabling citation-supported assistance across complexity levels. CRISS demonstrates the potential of human-centered, citation-grounded AI to support cancer registrars while preserving human oversight for final coding decisions.
Chinese Translation
癌症登记员(包括肿瘤数据专家,Oncology Data Specialists, ODSs)必须理解和解释复杂且频繁更新的编码与分期标准。我们开发了CRISS(癌症登记智能支持系统,Cancer Registry Intelligent Support System),这是一种检索增强生成(RAG)对话助手,可提供快速且有引用支持的登记指南访问。本研究评估了CRISS能否:(1)支持准确且有引用依据的回答;(2)改善相关指南的获取与解读;(3)在保留人工对最终摘要决策监督的前提下,支持培训和服务台使用。我们基于国家癌症登记标准构建了领域专用知识库,将其切分为带有元数据标签的段落,并以稠密嵌入形式建立索引。检索到的段落通过大语言模型(LLM)生成具有引用依据的回答。我们在简单、中等和困难三个难度级别的登记问题上,采用LLM作为评审(LLM-as-a-Judge)协议,对Gemini和GPT系列中的开源权重模型、专有模型以及非RAG基线模型进行了评估。RAG配置始终优于非RAG方法,且随着问题难度的增加优势更加明显。RAG在简单/中等/困难三个级别上的平均依据性得分为0.62/0.56/0.59,而非RAG仅为0.29/0.26/0.29。RAG模型的语义相似度得分总体上也更高。专有RAG模型在简单和中等问题上表现最强,而本地RAG模型在困难问题上排名最高,且专有模型通常更为谨慎。领域专用RAG提升了癌症登记问题回答的证据依据性和质量,同时能够在不同复杂度级别上提供有引用支持的辅助。CRISS展示了以人为中心、基于引用的AI在支持癌症登记员方面的潜力,同时对最终编码决策保留人工监督。
cs.AI / 29 / 2609.29084

A Rapid Pipeline for Training and Deploying ML Models on WeBe Band

一种在WeBe Band上训练和部署机器学习模型的快速流程
Kourkchi, Ehsan, Asmita, Asmita, Homayoun, Houman, Eslamimehr, Mahdi
Abstract
Developing optimized machine-learning algorithms for edge devices with limited computational and memory resources is challenging, time-consuming, and highly dependent on device-specific constraints. In this work, we streamline an edge ML workflow to enable rapid development, optimization, and deployment of machine-learning (ML) models directly on the WeBe Band, a wrist-worn wearable device designed for multimodal physiological data monitoring. The proposed system automatically generates hardware-efficient ML models that can be easily integrated into the WeBe core firmware, supporting AutoML, hardware-aware quantization, and performance profiling to build models that meet desired latency targets while remaining compatible with device memory and power limitations. The proposed framework tightly integrates the open-source Piccolo AI ecosystem with an automated pipeline that generates deployable firmware artifacts, performs hardware-aware model compilation, and supports over-the-air (OTA) deployment. The system supports multiple lightweight model classes, including classical machine-learning algorithms and neural networks, and provides built-in on-device profiling tools to evaluate inference latency and memory footprint under realistic execution conditions. Experimental results demonstrate clear trade-offs between model complexity and deployability on a microcontroller, showing that classical models offer strong real-time performance while lightweight neural networks require careful resource management. Rather than proposing new learning architectures, the current work mainly focuses on system-level automation, deployability, and enabling researchers and developers to rapidly iterate on models and evaluate them directly on target hardware. Although demonstrated on the WeBe Band platform, the workflow is designed to be extensible to other ML-powered edge devices.
Chinese Translation
为计算和内存资源有限的边缘设备开发优化的机器学习算法具有挑战性、耗时,并且高度依赖于设备特定的约束条件。在这项工作中,我们精简了边缘机器学习工作流,以实现机器学习(ML)模型直接在WeBe Band(一款专为多模态生理数据监测设计的腕戴式可穿戴设备)上的快速开发、优化和部署。所提出的系统能够自动生成硬件高效的ML模型,这些模型可以轻松集成到WeBe核心固件中,支持AutoML(自动机器学习)、硬件感知量化和性能分析,从而构建出满足预期延迟目标、同时与设备内存和功耗限制保持兼容的模型。该框架将开源的Piccolo AI生态系统与自动化流程紧密集成,该流程可生成可部署的固件工件、执行硬件感知的模型编译,并支持空中下载(OTA)部署。该系统支持多种轻量级模型类别,包括经典机器学习算法和神经网络,并提供内置的设备端性能分析工具,用于在真实执行条件下评估推理延迟和内存占用。实验结果展示了模型复杂度与微控制器上可部署性之间的明确权衡,表明经典模型具备出色的实时性能,而轻量级神经网络则需要精细的资源管理。本工作并非提出新的学习架构,而是主要聚焦于系统级自动化、可部署性,以及使研究人员和开发者能够快速迭代模型并直接在目标硬件上进行评估。尽管该工作流在WeBe Band平台上进行了演示,但其设计可扩展至其他由机器学习驱动的边缘设备。
cs.AI / 30 / 2609.29108

Functional Architecture of European Electricity Trading Markets: Requirements for AI Supported Trading Systems under Regulatory Constraints

欧洲电力交易市场的功能架构:监管约束下AI辅助交易系统的需求
Kurz, Walter, Stricker, Wojtek
Abstract
European electricity trading in the EU operates as a constrained multi-layer system in which legal design, exchange microstructure, and network physics are executed jointly across forward, day-ahead, intraday, and balancing horizons. This paper develops a functional architecture for AI-supported trading that is aligned with market-coupling mechanics, cross-zonal transfer constraints, and compliance obligations under REMIT, MiFID II, MiFIR, and EMIR. The contribution is a formal system specification composed of a decision-state vector, residual-exposure accounting, constrained optimization objective, executable-action permission gate, and fail-closed AI control logic with auditable records. The analysis maps major Nominated Electricity Market Operator (NEMO) venues and related exchange operators into an operational venue topology and identifies where cross-border coordination fails in practice: interface-level timing, permission heterogeneity, and balancing-layer coupling. The resulting framework proposes how AI can be deployed as a bounded decision component inside regulated market operation with explicit governance, rather than as an unconstrained prediction layer.
Chinese Translation
欧盟的欧洲电力交易作为一个受约束的多层系统运行,其法律设计、交易所微观结构与电网物理特性在远期、日前、日内和平衡市场中联合执行。本文构建了一个与市场耦合机制、跨区域输电约束以及REMIT、MiFID II、MiFIR和EMIR合规义务相一致的AI辅助交易功能架构。其贡献在于提出了一套形式化的系统规范,包括决策状态向量、剩余敞口核算、约束优化目标、可执行操作权限门控,以及具有可审计记录的故障关闭式(fail-closed)AI控制逻辑。本文的分析将主要的指定电力市场运营商(NEMO)交易场所及相关交易所运营商映射为一个可操作的场所拓扑结构,并识别了实践中跨境协调失效的关键环节:接口层面的时序问题、权限异质性问题以及平衡市场层的耦合问题。由此提出的框架展示了AI如何作为一个有边界的决策组件,在明确的治理框架下被部署于受监管的市场运营之中,而非作为一个不受约束的预测层。
cs.AI / 31 / 2609.29109

CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment

CounterRoute:基于分层反事实信用分配的自路由推理
Jiao, Ruochen, Fetahu, Besnik, Shi, Zhenyu, Nigam, Priyanka
Abstract
Reasoning-capable language models often produce long chains of thought when direct answers suffice, wasting inference compute. Many dual-mode models leave this choice to users. Automating it is challenging because routing targets evolve with the policy, initial mode preferences destabilize exploration, and sequence-level objectives entangle routing with response learning. We introduce CounterRoute, an online reinforcement-learning framework that jointly learns routing and modeconditioned responses in one shared policy directly from a native dual-mode checkpoint, without method-specific SFT warm-up. Paired current-policy counterfactual rollouts assign cross-mode credit only to the routing token, while within-mode GRPO trains response tokens. A paired-to-self-routed curriculum stabilizes early training with forced rollouts from both modes, then increases self-routed updates to improve autonomous routing. Across nine benchmarks, CounterRoute better balances accuracy and efficiency than heuristic and learned adaptive-routing methods. Relative to always-thinking checkpoints, it improves macro-average accuracy while reducing mean generated tokens by 51% for Qwen3-8B and 41% for Qwen3-14B. On instruction-following and commonsense benchmarks where direct answering is strong, think rates fall as low as 1% while response quality improves. Despite training only on math and instruction following, its routing behavior and response quality generalize to held-out coding, science, knowledge, and commonsense benchmarks.
Chinese Translation
具备推理能力的语言模型在直接回答即可满足需求时,仍常常生成冗长的思维链,浪费推理算力。许多双模态模型将这一选择权留给用户。实现自动化颇具挑战,原因在于路由目标会随策略演化而变化、初始模式偏好会破坏探索的稳定性,且序列级目标会将路由与响应学习纠缠在一起。我们提出 CounterRoute,这是一个在线强化学习框架,直接从原生双模态检查点出发,在单一共享策略中联合学习路由与模式条件化的响应,无需特定方法的 SFT 预热。成对的当前策略反事实采样(counterfactual rollouts)仅将跨模式信用分配给路由 token,而模式内的 GRPO 用于训练响应 token。从成对到自路由的课程机制(paired-to-self-routed curriculum)通过强制来自两种模式的采样来稳定早期训练,随后逐步增加自路由更新的比例以提升自主路由能力。在九个基准测试上,CounterRoute 比启发式方法和学习型自适应路由方法更好地平衡了准确率与效率。相对于始终思考的检查点,它在 Qwen3-8B 和 Qwen3-14B 上分别平均减少 51% 和 41% 的生成 token 数,同时提升了宏平均准确率。在直接回答表现出色的指令遵循和常识基准上,思考率可低至 1%,而响应质量反而提升。尽管仅在数学和指令遵循任务上训练,其路由行为与响应质量能够泛化到留出的编程、科学、知识和常识基准。
cs.AI / 32 / 2609.29140

Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

硬预算重复评估中诚实不确定性的尖锐极限
Cheng, Yezhou, Du, Runjia, Liu, Zeming, Chen, Qibai, Lyu, Hang, Wei, Yilan, Zeng, Yankai, Lin, Bojun
Abstract
Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of $M$ tasks with $L$ binary paths per task under the hard budget $(M+t)K$, where each path costs at most $K$ responses or episodes. For fixed $L \ge 3$ and $0 < \alpha \le 1/12$, the optimal expected width on the worst pure cohort is $\Theta_{\alpha,L}([M(t+1)]^{-1/2})$ when every task is observed and $\Theta_{\alpha,L}([M(t+\sqrt{M})]^{-1/2})$ when omission is allowed. The lower bounds cover adaptive hard-budget policies, and fixed random-subset designs attain both rates through disagreement certificates. A joint mean/disagreement interval turns the task-covering law into practical finite-budget inference. In an equal-budget LiveCodeBench replay with 16 models, 880 tasks, and five outputs per task, the task-covering design reduces median point-estimation MSE by 87.0\% relative to pooled uniform sampling, while the Joint certificate produces narrower confidence intervals in 15/16 panels and reduces median interval width by 30.6\%. Finite-regime analyses identify task coverage as the effective choice at the evaluated scale and characterize how cohort size and within-task agreement determine the useful operating region. Together, the sharp laws and fixed-budget evidence make replication and task coverage explicit design variables for information-efficient repeated evaluation.
Chinese Translation
重复评估可以在准确估计基准分数的同时,仍需要重复实验以认证较窄的不确定性。我们在 $M$ 个任务、每个任务 $L$ 条二值路径的固定网格上,在硬预算 $(M+t)K$(其中每条路径至多消耗 $K$ 次响应或回合)下刻画了这一需求。对于固定的 $L \ge 3$ 和 $0 < \alpha \le 1/12$,在最差纯队列上的最优期望区间宽度为:当所有任务都被观测时为 $\Theta_{\alpha,L}([M(t+1)]^{-1/2})$;当允许省略任务时为 $\Theta_{\alpha,L}([M(t+\sqrt{M})]^{-1/2})$。下界覆盖了自适应硬预算策略,而固定的随机子集设计通过分歧证书可达到这两个速率。联合均值/分歧区间将任务覆盖定律转化为实用的有限预算推断。在一个等预算的 LiveCodeBench 重放实验(16 个模型、880 个任务、每个任务五个输出)中,任务覆盖设计相对于池化均匀采样将中位点估计 MSE 降低了 87.0%,而联合证书在 16 组中的 15 组产生了更窄的置信区间,并将中位区间宽度降低了 30.6%。有限情形分析表明,任务覆盖是在所评估规模下的有效选择,并刻画了队列规模与任务内一致性如何决定有用的运行区域。总之,这些尖锐的定律和固定预算证据使重复实验与任务覆盖成为信息高效的重复评估中的显式设计变量。
cs.AI / 33 / 2609.29144

Scope Before You Persist: Preventing Cross-Family Interference in Agent Memory

先定范围,再持久化:防止智能体记忆中的跨任务族干扰
Cheng, Yezhou, Du, Runjia, Liu, Zeming, Chen, Qibai, Lyu, Hang, Zeng, Yankai, Wei, Yilan, Lin, Bojun
Abstract
Persistent memory lets language-model agents improve prompts and skills without updating model weights. We show that matching retrieval scope to certification scope enables these edits to support reliable repeated adaptation across recurring task families. We study frozen-model agents on ProcStream-RSI, a 12-round code-repair stream, using Orthogonal Regression Control (ORC), an execution-grounded gate for persistent skill edits. In an intervention that holds proposals and gate decisions fixed, retrieving each accepted skill only for its originating family raises mean hidden trajectory utility from 0.713 under global memory to 0.816 and changes harmful deployments from six of eight to none. In 27 paired randomized-order streams, Scoped-ORC improves mean trajectory utility by 0.063 [0.037, 0.094] over Global-ORC, accepts 63 rather than 12 updates, and produces multiple accepted updates in 19/27 streams, with 0/63 harmful acceptances. The global control reaches 0.713, below the static agent's 0.775, because locally valid edits can interfere with unrelated families. These results establish scope matching as a complementary control for persistent agent memory: certification determines whether an edit is supported, while retrieval scope determines where that evidence authorizes its use.
Chinese Translation
持久化记忆使语言模型智能体能够在不更新模型权重的情况下改进提示词与技能。我们证明,将检索范围与验证(certification)范围相匹配,可以使这些编辑在重复出现的任务族(task family)中支持可靠的持续适应。我们在 ProcStream-RSI——一个12轮的代码修复流——上研究冻结模型智能体,并采用正交回归控制(Orthogonal Regression Control, ORC),这是一种以执行结果为依据的持久化技能编辑准入门槛。在一个保持提案与门槛决策不变、仅将每个被接受技能的检索限定于其来源任务族的干预实验中,平均隐藏轨迹效用从全局记忆下的0.713提升至0.816,有害部署从八个中的六个降为零。在27个成对随机顺序的流中,Scoped-ORC 相比 Global-ORC 将平均轨迹效用提高0.063 [0.037, 0.094],接受了63次(而非12次)更新,并在19/27个流中产生多个被接受的更新,其中0/63次为有害接受。全局控制仅达到0.713,低于静态智能体的0.775,原因在于局部有效的编辑可能干扰无关的任务族。这些结果确立了范围匹配作为持久化智能体记忆的一种互补控制手段:验证决定一次编辑是否获得支持,而检索范围决定该证据授权其在何处使用。
cs.AI / 34 / 2609.29145

Claim-Gated Source-Risk Auditing for Generative Search

面向生成式搜索的声明门控式源风险审计
Zhou, Kainan, Xu, Chuhong, Qian, Gangzhen, Li, Zhaoyi
Abstract
A generative search answer can cite a supported passage yet omit a source relationship that changes its interpretation. We specify a claim-gated audit of the query-source-answer tuple. An omission is resolved only when relationship evidence, answer adoption, materiality, and disclosure are all observed; incomplete evidence remains unresolved rather than being treated as independence. The specification separates this endpoint from citation support and review priority, and binds decisions to versioned evidence spans. A reference checker makes the record contract executable. On an exhaustive synthetic suite, it reproduces all 81 three-state predicate combinations and rejects 192 deliberately malformed records. Common-guard baselines and predicate ablations isolate endpoint logic from missing-evidence handling, while controlled transitions check support separation and evidence removal. These are finite contract-conformance results, not detector accuracy or evidence of improved user outcomes. We define the independent annotation, held-out evaluation, and paired utility tests still required to establish semantic validity and deployment benefit.
Chinese Translation
生成式搜索的答案可能引用了有依据的段落,却遗漏了会改变其解读方式的来源关系。我们针对查询-来源-答案三元组提出了一种声明门控式(claim-gated)审计规范。只有当关系证据、答案采纳、实质性影响和信息披露均被观察到时,遗漏问题才被判定为已解决;证据不完整时保持未解决状态,而非被默认视为独立。该规范将此审计端点与引用支持和审查优先级区分开来,并将决策绑定到带版本标识的证据片段上。一个引用检查器(reference checker)使该记录契约具有可执行性。在一个穷尽式合成测试集上,该检查器复现了全部81种三态谓词组合,并拒绝了192条故意构造的畸形记录。通过通用防护基线和谓词消融实验,我们将端点逻辑与缺失证据处理机制相分离,并通过受控转换检验支持分离与证据移除行为。这些是有限的契约符合性结果,并非检测器的准确性评估,也不是用户效果改善的证据。我们明确了建立语义有效性和部署效益所需的独立标注、留出评估以及配对效用测试。
cs.AI / 35 / 2609.29154

A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents

一次错误的转弯并不毁掉整个旅程:面向大语言模型智能体的偏差引导技能自我进化
Feng, Yichun, Wang, Jiawei, Sun, Haozhe
Abstract
Large language model agents increasingly rely on natural-language skills to solve complex tool-use tasks. However, such tasks often admit multiple valid solution paths, making it inappropriate to improve skills by forcing failed trajectories to match a fixed successful trajectory. Moreover, failed trajectories are rarely entirely wrong: an agent may first collect useful evidence and make meaningful progress, but later deviate into an erroneous suffix. We therefore argue that skill self-evolution should identify where productive problem solving begins to break down, rather than reflect coarsely over the entire failure. Based on this insight, we propose SkillPivot, a deviation-point-guided framework for skill self-evolution. SkillPivot detects the transition from a useful prefix to an erroneous suffix using execution validity, goal progress, and action diversity. A stronger teacher then continues from the same prefix and produces a successful alternative under the same interaction history. By contrasting the student's failed suffix with the teacher's successful suffix, SkillPivot generates localized skill updates while preserving already effective guidance. Experiments on ToolQA, LogicBench, and WildClawBench show that SkillPivot consistently outperforms competing skill-evolution methods, improves multiple agent models, and produces compact, transferable skill updates.
Chinese Translation
大语言模型(LLM)智能体日益依赖自然语言技能来解决复杂的工具使用任务。然而,此类任务往往存在多条有效的解决路径,因此通过强制失败轨迹去匹配某个固定的成功轨迹来改进技能的做法并不合适。此外,失败轨迹也很少是完全错误的:智能体可能先收集了有用的证据并取得了实质性进展,但随后偏离进入错误的后续部分。因此,我们认为技能自我进化应当识别出有效的问题求解从何处开始瓦解,而不是对整个失败过程进行粗粒度的反思。基于这一洞察,我们提出了 SkillPivot,一个以偏差点为引导的技能自我进化框架。SkillPivot 利用执行有效性、目标进展和动作多样性来检测从有用前缀到错误后缀的转折点。随后,一个更强的教师模型从相同的前缀出发,在相同的交互历史条件下产生一个成功的替代方案。通过对比学生的失败后缀与教师的成功后缀,SkillPivot 在保留已有效指导的同时生成局部化的技能更新。在 ToolQA、LogicBench 和 WildClawBench 上的实验表明,SkillPivot 持续优于现有的技能进化方法,能够改进多种智能体模型,并产生紧凑且可迁移的技能更新。
cs.AI / 36 / 2609.29167

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

IndicBankBench:评估印度零售银行中语言模型助手的安全性与可靠性
Paul, Suvradip, Bhushan, Chandra, Sharma, Harsh, Kukreja, Nitin, Dedhia, Yatharth, Doshi, Keyur, Devadiga, Prashant
Abstract
Banking assistants must use account-specific information to answer requests and, in many cases, take actions through tools. Evaluating only the final response misses important errors. An assistant may ask for information it already has, rely on stale context, select the wrong account, or write an invalid value after stating the correct one. We introduce IndicBankBench, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes. Cases are evaluated at four stages: safety, action and tool use, response adequacy, and advisory quality. Tool use and most safety checks are deterministic. A narrow resolver handles only ambiguous confirmation-before-write cases, while a separate LLM judge evaluates semantic response adequacy. We run every case three times and report strict pass^3, which requires success on all trials. Across the eleven evaluated models, strict reliability ranges from 43.7% to 58.2%, whereas at-least-once success ranges from 60% to 74%. This gap shows that at-least-once success can overstate dependable banking behavior. The case-level diagnostics also distinguish systems that ask unnecessary questions from those that act but fail to reconcile customer context or fully resolve the request. We release the cases, mock environment, and evaluation harness.
Chinese Translation
银行助手必须利用账户特定信息来回答请求,并且在许多情况下需要通过工具执行操作。仅评估最终响应会遗漏重要的错误。助手可能会询问其已掌握的信息、依赖过时的上下文、选择错误的账户,或在陈述正确值后写入无效值。我们提出了 IndicBankBench,这是一个面向印度零售银行的基准测试,包含 799 个测试用例,涵盖五个操作领域、一个能力/拒绝领域以及二十个主要评估维度。测试用例在四个阶段进行评估:安全性、操作与工具使用、响应充分性以及咨询质量。工具使用和大部分安全检查采用确定性评估。一个狭窄的解析器仅处理写入前模糊确认的用例,而单独的 LLM 裁判则评估语义响应充分性。我们对每个用例运行三次,并报告严格的 pass^3 指标,即要求所有试验均成功。在所评估的十一个模型中,严格可靠性介于 43.7% 到 58.2% 之间,而至少一次成功的比例介于 60% 到 74% 之间。这一差距表明,至少一次成功的指标可能会高估可靠的银行服务行为。用例级别的诊断还能区分那些提出不必要问题的系统与那些执行了操作但未能核对客户上下文或完全解决请求的系统。我们公开发布了测试用例、模拟环境和评估框架。
cs.AI / 37 / 2609.29181

Right Choice of Classification Algorithms Based on Reinforcement Learning for Prediction of Non-Alcoholic Fatty Liver

基于强化学习的分类算法的正确选择用于非酒精性脂肪肝预测
Samadbin, Hasan, Daliri, Arman
Abstract
There are many complex issues in the world of artificial intelligence. Some of these problems are solved using other artificial intelligence methods, which are called artificial intelligence for artificial intelligence. Finding an appropriate classifier algorithm is a time-consuming task. For this reason, an algorithm that can automatically learn the choice of classification algorithms is very important. Classification algorithms are useful in predicting various diseases. Also, Primary Biliary Cirrhosis is one of the most well-known diseases that have been predicted by classification algorithms. This research's most significant achievement and novelty is the automatic increase in learning through a scoring method of reinforcement learning is called square learning (SL). In this research, an algorithm is presented that learns to automatically select the appropriate classification algorithm to predict Primary Biliary Cirrhosis. In this article, with inspiration from four evaluation metrics in classification algorithms, a new reinforcement learning method by the name of Fourth Degree Learning has been presented. In this research, we increased the performance of the classification algorithms used in this method from 63% of accuracy and achieved 98% accuracy.
Chinese Translation
人工智能领域存在许多复杂问题。其中一些问题可以通过其他人工智能方法解决,这类方法被称为“面向人工智能的人工智能”(AI for AI)。寻找合适的分类器算法是一项耗时的任务。因此,能够自动学习分类算法选择的算法显得尤为重要。分类算法在预测各种疾病方面非常有用。其中,原发性胆汁性肝硬化(Primary Biliary Cirrhosis)是分类算法已成功预测的知名疾病之一。本研究最重要的成果与创新点在于,通过一种称为平方学习(Square Learning, SL)的强化学习评分方法实现学习的自动提升。本研究提出了一种算法,该算法能够学习自动选择合适的分类算法来预测原发性胆汁性肝硬化。本文受分类算法中四个评估指标的启发,提出了一种名为四次方学习(Fourth Degree Learning)的新型强化学习方法。在本研究中,我们将该方法所用分类算法的性能从63%的准确率提升至98%的准确率。
cs.AI / 38 / 2609.29187

The Entropy Triangle Method (ETM): A novel framework for the prevention of cardiac arrhythmia with a review of more than 10,000 patients

熵三角法(ETM):一种预防心律失常的新框架——基于超过10,000名患者的综述
daliri, Arman
Abstract
One of the most important problems in medicine is to facilitate prediction. In this study, we propose entropy triangle method, a novel framework for predicting heart rhythms using a novel machine learning technique. This framework includes three steps: feature engineering, entropy triangle oversampling, and disease prediction. The dataset used in this study is a 12-lead electrocardiogram (ECG) arrhythmia research database with 10,646 patients. This dataset contains 11 different heart rhythms (5 sinus rhythms and 6 non-sinus rhythms). In this article, we introduce two firsts in machine learning and medicine that can predict non-sinus rhythm with over 85% accuracy. Our experimental results show, among others, that the most accurate classifier based on entropy triangles and the most useful oversampling are the supported vector classifiers and oversampling techniques for shark scent.
Chinese Translation
促进预测是医学中最重要的难题之一。在本研究中,我们提出了熵三角法(Entropy Triangle Method),这是一种使用新型机器学习技术来预测心律的新框架。该框架包括三个步骤:特征工程、熵三角过采样以及疾病预测。本研究使用的数据集是包含10,646名患者的12导联心电图(ECG)心律失常研究数据库。该数据集包含11种不同的心律(5种窦性心律和6种非窦性心律)。在本文中,我们在机器学习和医学领域实现了两个首次:能够以超过85%的准确率预测非窦性心律。我们的实验结果表明,基于熵三角的最准确分类器是支持向量分类器(SVC),而最有用的过采样技术是鲨鱼气味过采样技术(Shark Smell Oversampling)。
cs.AI / 39 / 2609.29189

When Honesty is Not Enough in AI Debate

当诚实不足以保障AI辩论安全时
Holland, Rayne, Zhu, Liming, Xue, Jason
Abstract
Scalable oversight aims to verify the behaviour of agents whose capabilities exceed those of their overseers. AI debate has been proposed as an oversight solution in which competing agents help a resource-limited verifier assess claims that it cannot reliably evaluate unaided. Much of its promise rests on incentivizing honest arguments that lead to correct verdicts. Yet a correct verdict need not uniquely determine the arguments used to support it. Agents may retain discretion over which correct claims to present, how to frame them, and in what order to disclose them. This residual freedom can allow agents to shape what the verifier learns beyond the task-relevant conclusion, pursuing latent objectives without compromising verdict correctness. To study this phenomenon, we introduce the framework strategic interactive oversight (SIO), which treats oversight jointly as a verification mechanism and a strategic communication channel. Within this framework, we formalise the notion of task-admissible latent optimisation, which entails the pursuit of latent objectives while maintaining a prescribed task performance. As proof-of-concept, we instantiate SIO in the establish protocol debate with cross-examination and quantify a tradeoff between task success and information disclosure about a hidden variable. The trade-off identifies a strategic window in which substantial disclosure remains compatible with task admissibility. Towards mitigation, we reduce admissible bias by expanding the cross-examiner's role to mitigate persistent disclosure over finite interaction horizons. Our results highlight the need to evaluate oversight not only by the correctness of its verdicts, but also by the information conveyed through its transcripts.
Chinese Translation
可扩展监督旨在验证那些能力超过监督者的智能体的行为。AI辩论(AI debate)被提出作为一种监督方案,其中相互竞争的智能体帮助资源受限的验证者评估其无法独立可靠判断的主张。这一方法的许多前景建立在激励诚实的论证、从而得出正确裁决的基础上。然而,正确的裁决并不唯一地决定用于支持它的论证方式。智能体可能对呈现哪些正确主张、如何表述以及以何种顺序披露保留自由裁量权。这种残余的自由度可以使智能体在任务相关结论之外塑造验证者所学到的信息,在不损害裁决正确性的情况下追求潜在目标。为研究这一现象,我们提出了战略交互监督(strategic interactive oversight, SIO)框架,该框架将监督同时视为一种验证机制和一条战略通信渠道。在该框架内,我们形式化了任务可容许的潜在优化(task-admissible latent optimisation)概念,即在保持规定的任务性能的同时追求潜在目标。作为概念验证,我们在带有交叉质询的“确立协议”辩论(establish protocol debate)中实例化了SIO,并量化了任务成功与关于隐藏变量的信息披露之间的权衡。该权衡确定了一个战略窗口,其中大量信息披露仍与任务可容许性兼容。为缓解这一问题,我们通过扩展交叉质询者的角色,在有限的交互范围内减少了持续性信息披露带来的可容许偏差。我们的结果强调,评估监督不仅要看其裁决的正确性,还要看其对话记录所传达的信息。
cs.AI / 40 / 2609.29191

ASIRF: An Agentic Framework for Context-Dependent Sensitive Information Redaction

ASIRF:一种面向上下文相关敏感信息脱敏的智能体框架
Priyadarshini, Sudha, Ghanem, Mohamed Chahine
Abstract
Sensitive information is defined by domain and intent, not a universal category, yet redaction systems such as privacy filters and named-entity recognizers fix a taxonomy at training time, requiring retraining for each new domain. We introduce ASIRF (Agentic Sensitive Information Redaction Framework), which retrieves domain-specific definitions based on the input's domain from a flexible knowledge base at inference time, needing no retraining to adapt. Two architectures, a three-call multi-agent pipeline and a single-agent variant, are evaluated across ten small open-weight models and eight datasets, including out-of-distribution fictional domains, against the OpenAI Privacy Filter (OPF) as a trained-classifier baseline. With only a few dozen expert-authored definitions per domain and no training data, ASIRF's recall exceeds OPF's in 68 of 80 model-domain combinations (85 percent), by at least one of the two architectures, with shortfalls confined mostly to OPF's training-distribution domains.
Chinese Translation
敏感信息由领域和意图定义,而非一个通用类别,然而诸如隐私过滤器和命名实体识别器等脱敏系统在训练时固定了分类体系,导致每个新领域都需要重新训练。我们提出了ASIRF(Agentic Sensitive Information Redaction Framework,智能体敏感信息脱敏框架),该框架在推理时基于输入所属领域从灵活的知识库中检索领域特定的定义,无需重新训练即可适应新领域。我们评估了两种架构——三次调用的多智能体流水线和单智能体变体——在十个小规模开源权重模型和八个数据集(包括分布外的虚构领域)上进行测试,并以OpenAI隐私过滤器(OPF)作为训练分类器基线。每个领域仅需几十条专家撰写的定义且无需任何训练数据,在80个模型-领域组合中的68个(85%)中,ASIRF的召回率至少在两种架构之一上超过OPF,且不足之处主要集中在OPF的训练分布领域。
cs.AI / 41 / 2609.29228

Towards An LLM-Driven Unified Conversion Framework for BT and FSM in Autonomous Intelligent Systems

面向自主智能系统中行为树(BT)与有限状态机(FSM)的大模型驱动统一转换框架
Qi, Zhang, Shuo, Yang, Zhengqiu, Zhu, Peng, Zhou, Peng, Jiao
Abstract
Finite state machine (FSM) and behavior trees (BT) are widely adopted behavioral modeling paradigms for autonomous intelligent systems. While functionally equivalent and inter-convertible in principle, existing transformation methods between FSM and BT face major challenges in preserving behavioral completeness and avoiding model complexity explosion. To overcome these issues, we propose an LLM-driven unified conversion framework that enables automatic, efficient, and semantically consistent transformation between FSM and BT. Specifically, a novel loop execution BT structure is designed for LLM to accurately capture the loop structure in FSM, thereby preserving behavioral completeness. To mitigate the state explosion problem in BT-to-FSM conversion, a depth compression strategy is introduced with LLM prompt to eliminate redundant control nodes, complemented by differentiated hierarchical conversion rules that collectively reduce the number of required sub-FSM. Simulation experiments in multiple autonomous decision-making scenarios demonstrate that the proposed framework enables an accurate and automated bidirectional conversion between FSM and BT. Furthermore, it significantly enhances the scalability and maintainability of generated models compared to traditional approaches, providing a practical solution for behavior model conversion in consumer-grade autonomous intelligent systems such as service robots, game agents, and smart home devices
Chinese Translation
有限状态机(FSM)与行为树(BT)是自主智能系统中广泛采用的行为建模范式。尽管二者在功能上等价且原则上可相互转换,但现有的FSM与BT之间的转换方法在保持行为完整性以及避免模型复杂度爆炸方面面临重大挑战。为解决这些问题,我们提出了一种大语言模型(LLM)驱动的统一转换框架,可实现FSM与BT之间自动、高效且语义一致的转换。具体而言,我们设计了一种新颖的循环执行BT结构,使LLM能够准确捕捉FSM中的循环结构,从而保持行为完整性。为缓解BT到FSM转换中的状态爆炸问题,我们引入了一种结合LLM提示的深度压缩策略以消除冗余控制节点,并通过差异化的分层转换规则进一步减少所需子FSM的数量。在多个自主决策场景中的仿真实验表明,所提出的框架能够实现FSM与BT之间准确、自动化的双向转换。此外,与传统方法相比,该框架显著提升了生成模型的可扩展性和可维护性,为服务机器人、游戏智能体和智能家居设备等消费级自主智能系统的行为模型转换提供了实用的解决方案。
cs.AI / 42 / 2609.29251

Policy as Code: A Coroutine-Bridge Harness for Fast-Reasoning Reliability on CAR-bench

策略即代码:面向 CAR-bench 快速推理可靠性的协程桥接框架
Matveev, Ivan
Abstract
CAR-bench evaluates whether tool-using agents stay reliable under real-world uncertainty, executing every tool inside the evaluator so that each tool-result exchange is a separate agent round-trip. A conventional next-action agent can batch parallel tool calls, but a chain of dependent calls costs it one model call per round of results. We present a coroutine-bridge harness in which the model's only action is to emit a Python program that blocks and resumes in place across evaluator tool exchanges. This decouples model invocation from tool round-trips: on the public test split the agent uses a median of two model calls against seven agent turns per task, resolving a full multi-turn task in a median of 1.8 s of model latency on Cerebras gpt-oss-120b. Because the action surface is executable code, deterministic CAR-bench policies are encoded directly as logic in the tool layer rather than as prompt rules, enforcing compliance at zero reasoning cost. On the official hidden evaluation the harness won Track 2 with 60.0% Pass^3, 4.5x the organizer baseline, at the lowest estimated cost and the fastest median task latency (3.14 s) of any entry scoring above that baseline; the same unchanged harness reproduced an identical 60.0% Pass^3 on GPT-5.5 in the Open track, matching frontier-model agents. A single static prompt, appended with per-task state at the tail, stays byte-identical across calls and across tasks: the frozen submission prompt served 78% of input tokens from cache (86.6% across its warm tail), against 73% over a three-week development corpus in which prompt edits repeatedly reset the cache. This compounds the few-call design into a small fraction of nominal input compute.
Chinese Translation
CAR-bench 用于评估使用工具的智能体在真实世界不确定性下的可靠性,所有工具均在评估器内部执行,使得每次工具结果交换都是一个独立的智能体往返过程。传统的“下一动作”式智能体可以批量并行调用工具,但一连串存在依赖关系的调用每次获取结果都需要付出一次模型调用的代价。我们提出了一种协程桥接(coroutine-bridge)框架,其中模型唯一需要执行的动作是生成一个 Python 程序,该程序在评估器的工具交换过程中原地阻塞并恢复运行。这将模型调用与工具往返解耦:在公开测试集上,每个任务中智能体使用的中位数仅为两次模型调用,而智能体轮次为七次;在 Cerebras gpt-oss-120b 上,以 1.8 秒的模型延迟中位数即可完成一个完整的多轮任务。由于动作空间是可执行代码,确定性的 CAR-bench 策略可以直接编码为工具层中的逻辑,而非以提示词规则的形式呈现,从而以零推理成本强制保证合规性。在官方隐藏评测中,该框架以 60.0% 的 Pass^3 赢得了赛道 2(Track 2),达到主办方基线的 4.5 倍,并且在所有得分超过该基线的参赛方案中拥有最低的估计成本和最快的中位任务延迟(3.14 秒);同一个未经修改的框架在公开赛道的 GPT-5.5 上复现了完全相同的 60.0% Pass^3,媲美前沿模型智能体。一个单一的静态提示词,仅在尾部附加各任务的状态,在多次调用和多个任务之间保持字节级完全一致:这一冻结的提交提示词有 78% 的输入 token 命中缓存(其温暖尾部达到 86.6%),而在一个历时三周、提示词反复编辑导致缓存不断重置的开发语料库上命中率仅为 73%。这使得少调用设计与名义输入计算量相比仅占很小一部分。
cs.AI / 43 / 2609.29266

Baszta: Data-Centric Fine-Tuning of a Polish Multi-Label Safety Classifier

Baszta:波兰语多标签安全分类器的数据中心化微调
Górski, Adam, Jąkalak, Mateusz, Jakubowski, Rafał
Abstract
We develop a multi-label Polish content-safety classifier by fine-tuning allegro/herbert-base-cased (124M) across five categories (hate, vulgarity, sexual content, crime, self-harm) using a Focal + R-Drop objective, and evaluate the resulting model against Bielik Guard (S\'ojka) on the shared out-of-distribution Gadzi J\k{e}zyk benchmark. Both systems are given per-category threshold tuning on the same calibration split. Under that matched protocol our model holds a small but statistically significant lead in micro F1, while an apparent macro-F1 lead does not survive: it was an artifact of comparing a tuned model against an untuned one. We also report what that micro figure is worth. Because Gadzi J\k{e}zyk is 97% crime-positive, a classifier that flags crime on every input and nothing else already scores 0.910 micro F1 on the same test split, so micro separates neither system from a degenerate strategy and macro is the column that does. Per-category and per-protocol figures are reported in Section 4. The residual out-of-distribution gap is one of calibration rather than discrimination. Ranking quality stays high while positive probabilities collapse, and per-category temperature scaling recovers the loss where Platt scaling and isotonic regression do not. That recovery turns out to be conditional on the calibration set containing safe text. Gadzi J\k{e}zyk contains almost none, so thresholds fitted on it flag crime on every safe input, and a balanced refit buys a deployable operating point at the cost of adversarial recall. We report both operating points rather than only the flattering one. Two changes that are standard practice, per-class cost-sensitive weighting and mean pooling, each raise in-distribution macro F1 while lowering the out-of-distribution figure, which indicates that robustness has to be selected for directly rather than inherited from in-distribution accuracy.
Chinese Translation
我们通过微调 allegro/herbert-base-cased(1.24亿参数),在五个类别(仇恨言论、粗俗语言、色情内容、犯罪、自残)上使用 Focal + R-Drop 目标函数,构建了一个多标签波兰语内容安全分类器,并在共享的分布外基准 Gadzi Język 上与 Bielik Guard(S'ojka)进行了评估。两个系统均在相同的校准集上进行了逐类别阈值调优。在该匹配协议下,我们的模型在 micro F1 上保持了较小但统计显著的领先,而表面上的 macro-F1 领先并不成立:那是将调优后的模型与未调优模型进行比较所产生的假象。我们还报告了该 micro 数值的实际价值。由于 Gadzi Język 中 97% 为犯罪类正例,一个将所有输入都标记为犯罪(且不作其他标记)的分类器在同一测试集上已可获得 0.910 的 micro F1,因此 micro 指标无法将任何系统与这种退化策略区分开,而 macro 才是能做到这一点的指标。逐类别和逐协议的数据在第 4 节中报告。分布外的残余差距源于校准而非判别能力:排序质量保持较高水平,但正例概率坍缩,且逐类别的温度缩放能够恢复这一损失,而 Platt 缩放和等渗回归则不能。这一恢复效果以校准集中包含安全文本为前提。Gadzi Język 中几乎没有安全文本,因此在其上拟合的阈值会将所有安全输入标记为犯罪,而平衡重新拟合虽然能换取可部署的工作点,却以对抗性召回率为代价。我们同时报告这两个工作点,而非只报告更有利的那一个。两项常规做法——逐类别代价敏感加权和均值池化——均能提高分布内 macro F1,却降低分布外表现,这表明鲁棒性必须被直接选择优化,而不能从分布内精度中继承获得。
cs.AI / 44 / 2609.29269

ALOE: Semantically Addressed Low-Rank Operators for Knowledge Editing

ALOE:用于知识编辑的语义寻址低秩算子
Li, Zeyan, Xu, Hu, Xu, Jianfeng
Abstract
Knowledge editing changes what a model knows by modifying parameters so that a requested fact updates while unrelated behavior is preserved. This is usually treated as a write problem, but editing also involves an address problem: deciding which hidden states should receive the new residual. An update that activates too narrowly memorizes one prompt, while one that activates too broadly disrupts neighboring knowledge. Parametric editors encode this scope implicitly, whereas memory-based editors make the selection explicit but keep it outside the edited model. We propose ALOE (Addressed Low-rank Operator for Editing), which learns semantic addresses from paraphrases and hard same-subject negatives, aligns them with autoregressive hidden states through rollout refinement and gate calibration, and embeds the resulting gated low-rank operator within one MLP layer, so that the deployed model runs in a single forward pass with no external retriever or auxiliary router. Evaluated on CounterFact, ZSRE, and KnowEdit across three 7--8B model families, ALOE achieves efficacy between 0.955 and 0.999 and locality between 0.981 and 1.000; mechanistic analyses confirm that the learned geometry separates competing edits and that calibration suppresses out-of-scope activation. The remaining errors concentrate in paraphrase coverage and write fitting.
Chinese Translation
知识编辑通过修改模型参数来改变模型所知的内容,使指定的更新事实生效,同时保留无关行为。这通常被视为一个“写入”问题,但编辑还涉及一个“寻址”问题:决定哪些隐藏状态应接收新的残差。激活范围过窄的更新只会记住单个提示,而激活范围过宽的更新则会干扰相邻知识。参数化编辑器隐式地编码了这一作用范围,而基于记忆的编辑器则显式地进行选择,但将其置于被编辑模型之外。我们提出ALOE(Addressed Low-rank Operator for Editing,用于编辑的寻址低秩算子),它从改写样本和难同主题负样本中学习语义地址,通过展开细化(rollout refinement)和门控校准将其与自回归隐藏状态对齐,并将由此得到的门控低秩算子嵌入单个MLP层中,使部署后的模型在单次前向传播中即可运行,无需外部检索器或辅助路由器。在三个7--8B模型家族上基于CounterFact、ZSRE和KnowEdit的评估表明,ALOE的效能(efficacy)介于0.955至0.999之间,局部性(locality)介于0.981至1.000之间;机制分析证实,所学习的几何结构能够区分相互竞争的编辑,且校准能抑制范围外的激活。剩余错误主要集中在改写覆盖和写入拟合方面。
cs.AI / 45 / 2609.29283

From Text Decisions to Pixels: An Study of Jev-Style Visual Choice Model

从文本决策到像素:Jev风格视觉选择模型研究
Zhou, Xunlan, Yang, Xianliang, Zhao, Li
Abstract
Visual software often needs a decision over supplied alternatives rather than a generated explanation. We present PixelJev, a native-image decision interface that maps an image, a task instruction, and a runtime candidate set to a structured choice and candidate-conditioned probabilities using small open multimodal models. Its initial realization unifies recognition and multiplechoice visual question answering through an existing language-model readout, with separately evaluated options for frozen inference, language-side adaptation, and held-out calibration. Across seven benchmark evaluations, 64-shot source adaptation raises Pets accuracy from 60.13% to 92.40% across optimization seeds and transfers to natural resampling, new texture labels, and A-OKVQA without target fitting, while frozen inference already supports both VQA tasks. A matched prompt-only follow-up on Pets and ScienceQA attributes the large Pets gain to adaptation and identifies a narrower output validity benefit of candidate readout in adapted VQA. Specialist DINOv2 probes remain stronger on source recognition, frozen 4B is stronger than adapted 2B on DTD and ScienceQA, and accuracy gains do not ensure calibrated target probabilities. These findings establish a working starting point for general-purpose visual decision models and identify the remaining requirements: schema robustness, cross-family transfer, and reliable use of visual evidence.
Chinese Translation
视觉软件通常需要对给定的候选选项做出决策,而非生成解释性文本。我们提出了PixelJev,一种原生图像决策接口,它利用小型开源多模态模型,将图像、任务指令和运行时候选集合映射为结构化选择以及以候选为条件的概率分布。其初始实现通过现有的语言模型读出机制,统一了识别任务和多选题视觉问答(VQA),并为冻结推理、语言侧适配以及留出集校准提供了分别评估的选项。在七项基准评估中,64样本的源域适配将Pets数据集的准确率在多个优化随机种子下从60.13%提升至92.40%,并在无需目标域拟合的情况下迁移至自然重采样、新纹理标签以及A-OKVQA数据集;同时,冻结推理本身已能支持两种VQA任务。在Pets和ScienceQA上进行的匹配的仅提示词后续实验表明,Pets上的巨大性能提升归因于适配过程,并发现候选读出在适配后的VQA中仅带来较窄的输出有效性收益。专用的DINOv2探测模型在源域识别任务上仍然更强,冻结的4B模型在DTD和ScienceQA上优于适配后的2B模型,且准确率的提升并不能保证目标域概率的良好校准。这些发现为通用视觉决策模型确立了一个可行的起点,并指明了剩余的需求:模式鲁棒性、跨模型家族迁移以及视觉证据的可靠利用。
cs.AI / 46 / 2609.29312

When No One Owns the Judgment: Accountability Under Contribution Dissolution in Human-AI Collaboration

当无人对判断负责:人机协作中贡献消解下的问责困境
Ye, Hengzhi
Abstract
Communities often respond to potentially AI-assisted work by asking three questions: Was AI used? Was that use disclosed? Can hidden use be detected? These questions place AI use itself at the center of accountability while overlooking a deeper problem: unowned judgment. Evaluations, claims, decisions, and creative directions can be shaped by AI with no accountable human or institution prepared to stand behind them. We develop this argument through two illustrative cases: AI-assisted peer review and concealed AI use in creative work. The first shows how contribution dissolution can weaken responsibility while the second shows how the fear of losing credit can discourage honest disclosure. The cases expose the limits of disclosure rules and provenance records as responses to AI-mediated collaboration. We offer three directions for discussion: distinguishing the roles AI plays, identifying judgments that require clear human ownership, and creating conditions in which AI involvement can be disclosed without default penalty. The broader aim is to make AI-shaped contributions discussable, creditable, contestable, and repairable.
Chinese Translation
面对可能借助AI完成的工作,学术界通常会提出三个问题:是否使用了AI?是否披露了AI的使用?隐蔽的使用能否被检测出来?这些问题将AI的使用本身置于问责的核心,却忽视了一个更深层次的问题:无主的判断(unowned judgment)。评估、主张、决策和创作方向都可能由AI塑造,而没有任何负有责任的人类或机构准备为其承担责任。我们通过两个典型案例来论证这一观点:AI辅助的同行评审,以及创作工作中对AI使用的隐瞒。第一个案例展示了贡献消解(contribution dissolution)如何削弱责任归属,第二个案例则展示了对失去署名权的担忧如何阻碍诚实的披露。这两个案例揭示了披露规则和来源记录作为应对AI介导协作的手段的局限性。我们提出三个讨论方向:区分AI所扮演的不同角色,识别需要明确人类归属责任的判断,以及创造能够在不施加默认惩罚的前提下披露AI参与的环境。更宏大的目标是使AI塑造的贡献变得可讨论、可署名、可质疑且可修复。
cs.AI / 47 / 2609.29341

SkinAgent AI: A Safety-Grounded Multimodal Agentic Framework for Non-Diagnostic Skincare Support

SkinAgent AI:一个面向非诊断性护肤支持的安全导向多智能体框架
Shahriar, Muhammad Muhtasim, Sayem, Abdullah Mohammad, Liew, Tze Hui, Mridha, M. F., Mahiuddin, Md.
Abstract
Consumer-facing skincare AI must coordinate visual evidence, product information, tool use, and user-facing actions within explicit evidence and safety boundaries. This study evaluates SkinAgent AI, a non-diagnostic multimodal framework that combines visual concern routing with grounded and auditable LLM-based orchestration. The architecture includes routing for Acne, Pores, and Wrinkles; photograph-based skin-type estimation; count-informed ordinal acne-severity support; typed tools; database-grounded recommendation and action functions; deterministic safety, privacy, and evidence checks; approval before state-changing actions; and structured trace and replay mechanisms. Visual-model performance and system-level agent behavior were evaluated separately. Across three seeds, the skin-condition routing model achieved 99.84% +/- 0.07% accuracy. Skin-type estimation achieved 88.85% accuracy, while count-informed acne-severity support achieved 84.59% accuracy with a quadratic weighted kappa of 0.9076. On a locked but non-independent 240-case system benchmark, intent accuracy was 80.00%, exact tool-set match was 62.92%, and strict task completion was 47.08%. No violations or successful cross-user leakage events were observed in the finite safety and privacy test suites. Tool-selection errors, incomplete grounding of product attributes, and unreliable failure fallback nevertheless remained. These findings support the feasibility of bounded, database-grounded, and traceable agent orchestration for non-diagnostic skincare assistance. They do not establish clinical readiness, external generalization, formal privacy guarantees, or universal safety. Independent validation, expert assessment, robustness and fairness testing, and prospective evaluation in real-world settings remain necessary.
Chinese Translation
面向消费者的护肤AI必须在明确的证据与安全边界内协调视觉证据、产品信息、工具使用和面向用户的操作。本研究评估了SkinAgent AI,这是一个非诊断性的多模态框架,它将视觉问题路由与基于证据且可审计的大语言模型(LLM)编排相结合。该架构包括针对痤疮、毛孔和皱纹的路由机制;基于照片的肤质估计;基于计数的痤疮严重程度有序分级支持;类型化工具;基于数据库的推荐和操作功能;确定性的安全、隐私和证据检查;状态更改操作前的审批机制;以及结构化的追踪与回放机制。视觉模型性能与系统级智能体行为被分别评估。在三个随机种子下,皮肤状况路由模型达到了99.84% ± 0.07%的准确率。肤质估计的准确率为88.85%,基于计数的痤疮严重程度分级支持的准确率为84.59%,二次加权Kappa系数为0.9076。在一个锁定但非独立的240案例系统基准测试中,意图准确率为80.00%,工具集合完全匹配率为62.92%,严格任务完成率为47.08%。在有限的安全与隐私测试集中未观察到违规或成功的跨用户信息泄露事件。然而,工具选择错误、产品属性证据关联不完整以及不可靠的失败回退机制仍然存在。这些发现支持了有界的、基于数据库的、可追踪的智能体编排用于非诊断性护肤辅助的可行性。但它们并未确立临床可用性、外部泛化能力、正式的隐私保证或普遍的安全性。独立验证、专家评估、鲁棒性与公平性测试,以及真实世界场景下的前瞻性评估仍然必不可少。
cs.AI / 48 / 2609.29345

The Last Human Gate: Forward Deployed Engineering for Governance Automation

最后一道人类关卡:面向治理自动化的前置部署工程
Canale, Jeremy
Abstract
Enterprise governance requires decisions, evidence, and accountable authority; it does not require every review task to retain its current human implementation. We develop a task-substitution framework for Digital Governance Frameworks (DGF), treating each gate as an executable contract. Substitution requires sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted. We derive a residual-work threshold and show why automating most cases can still increase labor. Forward deployed engineering connects these conditions to an architecture for agents, rule engines, evidence services, and escalation. DGF-Bench supplies controlled evidence from 300 synthetic projects and 899 evaluable model-project runs. Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%; complete-route success is 76.92%, 42.33%, and 24.67%. A deterministic control passes all 1,700 gates given the supplied rules and structured facts, locating the comparison in execution of a supplied decision kernel. Evidence audits and 135 repeated runs distinguish correct decisions from reliable execution. A document counterexample establishes an information-sufficiency obstruction. These results support the technical feasibility of replacing human execution of specified governance-review tasks with agents and software. The framework specifies a workforce test based on the complete human effort required at fixed output and quality; the present measurements concern review performance. Sources, dossiers, traces, and analyses are public.
Chinese Translation
企业治理需要决策、证据和可问责的权威;但并非要求每项审查任务都保留其当前的人工实现方式。我们针对数字治理框架提出了一个任务替代框架,将每个关卡视为可执行的契约。替代需要满足以下条件:充分且可获取的信息、有效的决策与权限校验,以及在计入异常处理、验证、纠正和维护工作后,总人工工作量的减少。我们推导出残余工作量阈值,并说明为何自动化大多数案例反而可能增加劳动量。前置部署工程将这些条件与涵盖智能体、规则引擎、证据服务和升级处理的架构相连接。DGF-Bench 提供了来自 300 个合成项目和 899 次可评估模型-项目运行的受控证据。Gemini 3.8 Flash、GPT-5.6 Luna 和 DeepSeek v4.1 Flash 的严格关卡成功率分别为 94.98%、83.29% 和 74.18%;完整路线成功率分别为 76.92%、42.33% 和 24.67%。在给定所提供的规则和结构化事实的情况下,一个确定性控制通过了全部 1,700 个关卡,从而将比较定位于对所提供决策内核的执行上。证据审计和 135 次重复运行区分了正确决策与可靠执行。一个文档反例确立了信息充分性障碍。这些结果支持用智能体和软件替代特定治理审查任务的人工执行在技术上的可行性。该框架提出了基于固定产出和质量下所需完整人力投入的劳动力检验;而本文的测量仅涉及审查性能。源代码、档案、追踪记录和分析均已公开。
cs.AI / 49 / 2609.29363

Beyond Simple Input-Output Assessment Tasks: Leveraging Automated Programming Assessment for Non-Trivial Courses

超越简单的输入-输出评测任务:利用自动编程评测支持非平凡课程
Jordao, Artur
Abstract
The public visibility of Artificial Intelligence (AI) is growing rapidly, driven by the positive impact of its applications across diverse fields of knowledge. In this new chapter, courses that cover the foundations of AI and machine learning become essential for understanding their role and potential in contemporary society. Therefore, understanding fundamental concepts and elementary algorithms through the close integration of theory with practice is essential in AI courses. In this essay, we report our experience designing machine learning exercises for automated assessment tools in programming. It is worth mentioning that we are not developing a novel form of automated grading system. Instead, we propose a perspective that frames machine learning problems as input-output assessment tasks. From this perspective, each exercise admits a unique and deterministic answer and enables automated programming assessment tools (e.g., VPL for Moodle, Codeforces, and MOJ) to effectively support AI education. We believe this essay can encourage instructors to foster educational innovation by adopting more dynamic and interactive approaches to AI courses that integrate theory and practice. Importantly, this essay does not introduce an innovation in the use of AI for education; rather, it introduces an innovative approach to improving the learning of AI, particularly, machine learning.
Chinese Translation
随着人工智能(AI)在各知识领域的应用产生积极影响,其公众关注度正迅速增长。在这一新篇章中,涵盖人工智能与机器学习基础的课程对于理解它们在当代社会中的作用和潜力变得至关重要。因此,在人工智能课程中,通过理论与实践的紧密结合来理解基本概念和基础算法是必不可少的。本文报告了我们在编程自动评测工具中设计机器学习练习的经验。值得指出的是,我们并非在开发一种新型的自动评分系统,而是提出了一种将机器学习问题框架化为输入-输出评测任务的视角。从这一视角出发,每个练习都具有唯一且确定性的答案,从而使自动编程评测工具(例如 Moodle 的 VPL、Codeforces 和 MOJ)能够有效支持人工智能教育。我们相信,本文能够鼓励教师通过在人工智能课程中采用更具动态性和交互性的方法来融合理论与实践,从而促进教育创新。重要的是,本文并未在“将人工智能用于教育”方面引入创新,而是引入了一种改进人工智能(尤其是机器学习)学习的创新方法。
cs.AI / 50 / 2609.29366

Epistemic-Probabilistic Model for Guarded Multi-Agent LLM Coordination

面向受控多智能体大语言模型协调的认知-概率模型
Nasiri, Mehdi, Arvenaghi, Mohammad Saeed, Vaezi, Sadegh, Ardeshir-Larijani, Ebrahim
Abstract
Multi-agent large language models (LLMs) have become ubiquitous in applied AI, yet their theoretical foundations remain surprisingly understudied. Viewed through the lens of multi-agent systems theory, several shortcomings come to light: a lack of social intelligence, the absence of coordination mechanisms among agents, unknown emergent behavior, and interactions between agents that are bounded by natural language. We address two of these gaps: the absence of social behavior and the lack of mechanisms for inter-agent coordination. We introduce Epistemic Probabilistic Language Agents (EPLA), a neuro-symbolic architecture for multi-agent coordination under uncertainty. A Symbolic Guard provides structured diagnostic feedback. The LLM generates typed actions, and the Guard controls their execution against an authoritative symbolic state. We formalize the epistemic layer in a gossip testbed through epistemic lottery gossip models, which combine view-based call histories with agent-indexed probability weights. We argue that implementing such a formalism can address shortcomings of agentic LLMs.
Chinese Translation
多智能体大语言模型(LLM)在应用人工智能中已无处不在,但其理论基础的研究却出人意料地匮乏。从多智能体系统理论的视角审视,可以发现若干不足:缺乏社会智能、智能体之间缺少协调机制、涌现行为未知,以及智能体之间的交互受限于自然语言。我们针对其中两个缺口展开研究:社会行为的缺失和智能体间协调机制的匮乏。我们提出了认知概率语言智能体(Epistemic Probabilistic Language Agents, EPLA),这是一种面向不确定性下多智能体协调的神经符号架构。其中,符号守卫(Symbolic Guard)提供结构化的诊断反馈;LLM 生成带类型的动作,而守卫则依据权威的符号状态对其执行加以控制。我们在一个流言(gossip)测试平台中,通过认知彩票流言模型对认知层进行形式化,该模型将基于视图的调用历史与以智能体为索引的概率权重相结合。我们认为,实现此类形式化方法能够弥补智能体化 LLM(agentic LLM)的上述不足。
cs.AI / 51 / 2609.29381

An auditable conditional-strategy framework for open-ended decision-making in complex lung cancer

一种面向复杂肺癌开放式决策的可审计条件策略框架
Wang, Daoyun, Huang, Zhicheng, Sun, Huaiyuan, Xu, Jiaqi, Xu, Xiaowei, Zheng, Zhibo, Bing, Zhongxing, Lin, Yuxiao, Liang, Yicheng, Gao, Chao, Xue, Bowen, Zhang, Kai, Xu, Song, Yan, Wanpu, Xia, Hui, Li, Lin, Yan, Xiang, Hu, Mu, Ma, Qianli, Xue, Zhiqiang, Liu, Xiaofang, Han, Zhihai, Zhang, Nan, Tang, Chuanhao, Zhang, Tongmei, Song, Lan, Zhu, Zhaohui, Zeng, Xuan, Wu, Shafei, Guan, Hui, Deng, Lei, Yang, Huaxia, Lian, Zeliang, Sun, Wubin, Wang, Yongxin, Shen, Xiaohui, Wang, Binlin, Gu, Tiantian, Cui, Yu, Zhang, Li, Wang, Shirui, Liang, Naixin
Abstract
Complex lung cancer decisions can involve several defensible pathways whose eligibility, sequencing and safety depend on unresolved information. Effective support must make explicit how patient conditions govern pathway eligibility, deferral and redirection. MedGPT Clinical Explorer (MCE) organizes alternatives, decision-changing unknowns, safety constraints and fallback into a conditional strategy for clinician review. To evaluate this representation in physician-authored strategies, multidisciplinary experts established case-specific references for 40 cases within a purposive 100-case corpus, and 250 physicians from 98 institutions produced 2,250 strategies under unaided, retrieval-reference and MCE-assisted conditions. MCE-assisted strategies expressed more applicable clinical requirements, measured by the Admissible Pathway Attainment Score (APAS; 0-100), than unaided strategies (adjusted difference, 12.87; 95% CI, 11.18-14.55) and retrieval-reference strategies (5.22; 3.52-6.93). With the same knowledge base available in the retrieval-reference and MCE-assisted conditions, the additional content centered on candidate pathways, decision-critical information and safety constraints. Physicians' whole-strategy acceptability judgments correlated with APAS (Spearman's rho = 0.671), while a complementary relationship audit assessed whether candidates, conditions and subsequent actions were coherently connected. Together, these findings identify two complementary dimensions of open-ended decision support: coverage of clinically relevant content and coherent links among pathways, conditions and subsequent actions. MCE provides a shared decision object that makes consequential omissions and pathway contingencies visible before action; prospective studies should evaluate its effects on clinical workflow and patient outcomes.
Chinese Translation
复杂肺癌的诊疗决策可能涉及多条均具合理性的路径,其适用资格、先后顺序与安全性取决于尚未明确的信息。有效的决策支持必须清晰阐明患者病情条件如何决定路径的适用、延迟与转向。MedGPT Clinical Explorer(MCE)将备选方案、影响决策的未知信息、安全约束及后备策略整合为条件策略,供临床医生审阅。为评估这一表示方式在医生编写的策略中的作用,多学科专家为一个具有目的性抽样的100例病例库中的40例建立了病例特异性参照,250名来自98家机构的医生在无辅助、检索参考和MCE辅助三种条件下共编写了2,250份策略。以可接受路径达成度评分(Admissible Pathway Attainment Score, APAS;0-100)衡量,MCE辅助策略表达的临床适用要求多于无辅助策略(校正差值,12.87;95% CI,11.18-14.55)和检索参考策略(5.22;3.52-6.93)。由于检索参考条件与MCE辅助条件使用相同的知识库,多出的内容主要集中于候选路径、决策关键信息和安全约束方面。医生对整体策略的可接受性判断与APAS相关(Spearman rho = 0.671),同时一项互补的关系审计评估了候选路径、条件与后续行动之间是否连贯衔接。综上,这些发现指出了开放式决策支持的两个互补维度:临床相关内容的覆盖度,以及路径、条件与后续行动之间的连贯关联。MCE提供了一个共享的决策对象,能在行动之前使重要的遗漏和路径的应变安排变得可见;前瞻性研究应评估其对临床工作流程和患者结局的影响。
cs.AI / 52 / 2609.29396

Wearable ECG Quality Assessment: A Deep Learning and Ambulatory Context-Awareness Approach

可穿戴心电信号质量评估:一种深度学习与动态情境感知相结合的方法
Mao, Xiaopeng, Weisbjerg, Marike, Puthusserypady, Sadasivan
Abstract
This paper presents and evaluates a Deep Learning-based (DL-based) Signal Quality Assessment (SQA) model to distinguish between clean and noisy ambulatory Electrocardiograms (ECG). The model is trained on Copenhagen Center for Health Technology-Contextualized Arrhythmia Database (CACHET-CADB), which, to the best of our knowledge, is the first ambulatory ECG database with both physical and patient-reported contextual data. The model shows stable performance on different databases such as MIT-databases and the latest PyhsioNet/Cinc Challenge 2021 databases. Subsequently, the paper demonstrates how complicated ECG noise can be investigated by the SQA model and the physical contextual data.
Chinese Translation
本文提出并评估了一种基于深度学习(DL)的信号质量评估(SQA)模型,用于区分干净的与含噪声的动态心电图(ECG)信号。该模型在哥本哈根健康技术中心-情境化心律失常数据库(CACHET-CADB)上进行训练,据我们所知,这是首个同时包含身体活动数据和患者自报告情境数据的动态心电图数据库。该模型在不同数据库(如MIT数据库和最新的PhysioNet/CinC Challenge 2021数据库)上均表现出稳定的性能。随后,本文展示了如何利用该SQA模型和身体活动情境数据来研究复杂的心电噪声。
cs.AI / 53 / 2609.29403

RD-JEPA: Predictive latent pretraining for few-trajectory transfer across reaction--diffusion equations

RD-JEPA:面向反应-扩散方程少轨迹迁移的预测式潜在表征预训练
Si, Chenhao, Yan, Ming
Abstract
Learning surrogates for time-dependent partial differential equations often requires a new simulation corpus when the governing operator changes. We introduce RD-JEPA, a joint-embedding predictive architecture for self-supervised pretraining on reaction-diffusion trajectories. A single model is pretrained on five parameterized systems and then adapted to three held-out systems whose reaction operators and trajectories are excluded from pretraining. Using one, five, or ten complete trajectories from a held-out system, RD-JEPA achieves lower mean relative discrete $\ell^2$ field error and mean absolute spatial first-difference error than five supervised surrogate baselines, an independently trained control that removes the trajectory-dependent predictive latent pathway, and an architecture-matched model trained from scratch. Within the evaluated equations, output resolution, forecast horizons, and choices of adaptation trajectories, the results indicate that prediction of future-state representations can support data-efficient adaptation across related reaction-diffusion systems.
Chinese Translation
为时变偏微分方程学习代理模型,通常在控制算子发生变化时需要重新构建新的仿真数据集。我们提出了RD-JEPA,一种用于在反应-扩散轨迹上进行自监督预训练的联合嵌入预测架构(joint-embedding predictive architecture)。单一模型首先在五个参数化系统上进行预训练,随后适配到三个保留(held-out)系统,这些系统的反应算子和轨迹均未包含在预训练数据中。仅使用来自保留系统的一条、五条或十条完整轨迹,RD-JEPA相比五个有监督代理基线模型、一个去除轨迹依赖预测潜在通路的独立训练对照模型,以及一个从零开始训练的结构匹配模型,均取得了更低的空间一阶差分误差和平均相对离散 $\ell^2$ 场误差。在所评估的方程、输出分辨率、预测时间跨度以及适配轨迹选择范围内,结果表明,对未来状态表征的预测能够支持在相关反应-扩散系统之间实现数据高效的适配。
cs.AI / 54 / 2609.29429

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

Just Ask Jev:基于强化学习的校准决策作为AI对齐失效的零样本检测器
Guo, Ruoqi, Liu, Yi, Deng, Gelei, Li, Yuekang, Zhao, Lida, Wu, Yutao, Chen, Simin, Zhang, Ying, Zhang, Leo Yu
Abstract
Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement learning for calibrated decisions (RLCD), answers many typed questions about one input with calibrated probabilities in a single call. Whether it detects alignment failures has not been measured. We present RLCDAlignBench, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking. It spans 44 benchmarks and five target models, labelled by each benchmark's scorer and, on two, by humans. Many of these failures are relational, defined against a reference, such as the user's belief or an injected instruction, that the response alone does not reveal. Our key idea is therefore to vary what Jev is asked separately from what it sees: the question's wording and answer type on one side, the fields of the input on the other. A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks. Question wording matters little, while context matters more, mostly through fields that encode the label. Jev matches the reference scorer's agreement with human labels, surfaces label defects in existing benchmarks, and costs 63x less than LLM-judge scorers. Code and data: https://github.com/sumleo/RLCDAlignBench.
Chinese Translation
对齐失效检测器用于筛查已部署的语言模型并为对齐基准打分。大多数检测器是生成式评判模型,需要对每条评判标准都进行一次解码过程;而读取词元概率的分类器(如 Llama Guard)每次调用也只能输出一个固定标签。Jev 是一个通过面向校准决策的强化学习(RLCD)训练的模型,能够在单次调用中以校准后的概率回答关于同一输入的多个类型化问题。然而,其检测对齐失效的能力尚未被测量。我们提出了 RLCDAlignBench,在十种对齐失效上对 Jev 进行基准测试:谄媚、越狱、欺骗、提示注入、幻觉、隐私侵犯、社会偏见、奖励篡改(reward hacking)、隐瞒不确定性以及权力寻求。该基准涵盖 44 个基准数据集和五个目标模型,标签由各基准的评分器提供,其中两个数据集由人工标注。这些失效中许多是关系性的,即相对于某个参照物(如用户的信念或注入的指令)来定义,而仅凭响应本身无法揭示该参照物。因此,我们的核心思想是将 Jev 被询问的内容与其所看到的内容分离开来变化:一方面是问题的措辞和答案类型,另一方面是输入的各个字段。单个通用问题在零样本设置下即可达到 0.886 的中位 AUROC,并在大多数基准上超越有监督基线。问题措辞影响不大,而上下文影响更大,主要通过编码标签信息的字段起作用。Jev 与参照评分器在人工标签上的一致性相当,能够暴露现有基准中的标签缺陷,且成本比 LLM 评判评分器低 63 倍。代码与数据:https://github.com/sumleo/RLCDAlignBench。
cs.AI / 55 / 2609.29465

SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

SWE-Prometheus:衡量真实世界代码仓库中工程治理的改进
Wu, Jiajun, Sun, Leixin, Tan, Zihan, Liu, Yitao, Li, Shuo, Qian, Jiaru, Quan, Shanghaoran, Zhao, Chuangxin, Liao, Yangxu, Liu, Yang, Chong, Bin, Wan, Guancheng
Abstract
Large language model based coding agents have made substantial progress on repository-level software engineering tasks. Existing repository benchmarks, however, usually start from a human-identified issue and evaluate whether a patch satisfies a functional signal. We present SWE-Prometheus, a benchmark for the broader task of improving repository engineering governance. Each task provides a fixed snapshot and an open-ended objective, requiring the agent to identify risks, prioritize interventions, and verify the resulting changes. SWE-Prometheus evaluates six governance dimensions through paired evidence, clean-environment probes, behavior gates, and two independent teacher ratings of the same evidence. The benchmark contains 60 repositories; ten models are evaluated on a shared 22-repository public subset, where mean Normalized Governance Improvement ranges from 0.0568 to 0.5760 and observed behavior-breakage rates range from 0% to 23%. On a frozen ten-repository batch, a repository-blind template obtains mean NGI 0.272, but its gains concentrate in Tests & CI, Quality Gates, and Documentation; it improves Reproducible Environment and Dependency & Security on none of the repositories. This baseline makes the distinction between adding governance artifacts and producing execution-backed improvements measurable. The no-op condition has median NGI zero and standard deviation 0.073; two teachers agree exactly on 57 of 60 dimension scores for the same no-op evidence. For the two highest conditional-mean systems, common-valid NGI is similar, while full-pool comparisons that include behavior failures favor Kimi-K3. These results show why repository-governance evaluation should report improvement, behavior preservation, evidence quality, and coverage together.
Chinese Translation
基于大语言模型的编程智能体在仓库级软件工程任务上已取得长足进展。然而,现有的仓库基准通常从人工识别的问题出发,仅评估补丁是否满足某种功能性信号。我们提出 SWE-Prometheus,一个针对更广泛任务的基准——改进代码仓库的工程治理。每个任务提供一个固定的代码快照和一个开放式的目标,要求智能体识别风险、对干预措施进行优先级排序,并验证所做出的更改。SWE-Prometheus 通过成对证据、干净环境探针、行为门控以及对相同证据的两个独立教师评分,评估六个治理维度。该基准包含 60 个代码仓库;十个模型在一个共享的 22 仓库公共子集上进行评估,其中平均归一化治理改进(Normalized Governance Improvement)介于 0.0568 到 0.5760 之间,观察到的行为破坏率介于 0% 到 23% 之间。在一个冻结的十仓库批次上,一个不感知仓库内容的模板获得平均 NGI 为 0.272,但其收益集中在测试与 CI(Tests & CI)、质量门控(Quality Gates)和文档(Documentation)方面;在所有仓库上它均未能改进可复现环境(Reproducible Environment)以及依赖与安全(Dependency & Security)。这一基线使得‘添加治理产物’与‘产生有执行支撑的改进’之间的区别变得可度量。无操作(no-op)条件的中位数 NGI 为零,标准差为 0.073;针对相同的无操作证据,两位教师在 60 个维度评分中有 57 个完全一致。对于两个条件均值最高的系统,其共同有效 NGI 相近,而包含行为失败的完整池比较则更倾向于 Kimi-K3。这些结果表明,仓库治理评估应同时报告改进程度、行为保持性、证据质量与覆盖范围。
cs.AI / 56 / 2609.29508

Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy

大语言模型智能体多轮一致性的评估:生存分析与失败理由分类法
Bogdanov, Igor, Manakina, Olga, Lung, Chung-Horng
Abstract
Large language model (LLM) agents may perform well on isolated tasks yet drift into inconsistency over extended interaction. We evaluate temporal consistency in a controlled 20-step multi-agent setting inspired by delayed-gratification studies. At each step, an agent chooses between continuing to delay a reward or claiming it immediately (terminating the episode). Across a full-factorial manipulation of social visibility (private vs public), persona stressors, and deliberation policy, we run 84,540 trajectories spanning 8 model families. Treating the first reward-claim as a time-to-event outcome, we estimate Kaplan-Meier survival curves and fit discrete-time hazard regression to quantify how experimental factors shift failure risk over time. Then, to analyze rationales and language patterns associated with failure, we build a seven-category taxonomy from 13,780 deliberation traces from agents who choose to terminate the episode, using an LLM-assisted labeling paired with human audit ($\kappa=0.83$). Rationale profiles change systematically with time and context: early failures are more impulse-driven, later failures more fatigue- and cost-benefit-framed, while public settings increase norm-oriented justifications. We also find a deliberation-inconsistency association: among failures, longer deliberation correlates with higher rates of intra-rationale contradiction (simultaneous pro-delay and pro-claim statements), challenging the assumption that more reasoning text implies greater consistency. Together, the survival and rationale analyses reveal distinct temporal reliability regimes and model-specific "failure fingerprints", offering an evaluation lens for diagnosing inconsistency in multi-turn agent behavior.
Chinese Translation
大语言模型(LLM)智能体在孤立任务上可能表现良好,但在长时间的交互中却可能逐渐失去一致性。我们在一个受延迟满足研究启发的受控20步多智能体环境中评估时间一致性。在每一步中,智能体需在继续延迟奖励与立即领取奖励(终止该轮交互)之间做出选择。通过对社会可见性(私密 vs 公开)、人格压力源和深思策略的全因子操控,我们运行了涵盖8个模型家族的84,540条轨迹。我们将首次领取奖励视为一个事件时间结果,估计了Kaplan-Meier生存曲线,并拟合离散时间风险回归,以量化实验因素如何随时间改变失败风险。随后,为分析与失败相关的理由和语言模式,我们基于13,780条来自选择终止交互的智能体的深思记录,构建了一个七类别的分类体系,采用LLM辅助标注并结合人工审核($\kappa=0.83$)。理由特征随时间和情境发生系统性变化:早期失败更多由冲动驱动,后期失败更多以疲劳和成本收益为框架,而公开环境会增加面向规范的合理化解释。我们还发现了深思与不一致性之间的关联:在失败案例中,更长的深思与更高的理由内部矛盾率(同时出现支持延迟和支持领取的陈述)相关,这一发现挑战了推理文本越多意味着一致性越强的假设。综合来看,生存分析和理由分析揭示了明显不同的时间可靠性状态以及模型特有的“失败指纹”,为诊断多轮智能体行为中的不一致性提供了一个评估视角。
cs.AI / 57 / 2609.29509

Delay-of-Gratification as a Multi-Agent Survival Micro-benchmark for Long-Horizon LLMs: Social Exposure, Personas, and Tool Use Budgets

延迟满足作为面向长时程大语言模型的多智能体生存微基准:社会暴露、人格角色与工具使用预算
Manakina, Olga, Bogdanov, Igor, Lung, Chung-Horng
Abstract
Large language models (LLMs) are increasingly deployed as multi-turn agents that must sustain goals, use tools, and adapt to other agents over extended interactions. However, existing research lacks auditable, multi-turn, multi-factorial experiments that quantify LLM behavior under explicit constraints, with time-resolved statistics that reveal how behavior unfolds over long horizons. To address this gap, we develop a multi-agent micro-benchmark inspired by the Stanford marshmallow experiment: ReAct agents operate minute-by-minute with a "raise a question" tool under a per-step budget, while we factorially manipulate social context (broadcast vs. isolated), personas (age, hedonic drive), and metacognitive policy (mandatory vs. optional tool use). We analyze outcomes with Kaplan-Meier (KM) survival curves and discrete-time hazard models over a long risk horizon across 19,200 agent trajectories in 64 cells. Behavior shows a sharp early "eat" impulse, and only 75.9% of agents persist to the end. In a discrete-time hazard model, isolation reduces per-minute risk relative to broadcast, whereas a must-use self-questioning policy increases risk. On average, agents ask $\approx 7.12$ questions and hit the per-step budget in $\approx 6\%$ of minutes. Questioning declines faster under broadcast than isolation. Ablation experiments demonstrated that removing hedonic drive and/or persona age increases survival and completion, narrows the broadcast/isolated gap, but leaves the must vs. may ordering intact. The combined ablation (no hedonic + no persona age) yields the highest completion (approaching $1.0$). These results establish delay-of-gratification as a compact, multi-turn interaction benchmark that captures social contagion and tool-use dynamics in LLM agents, providing a reproducible testbed and statistics for analyzing long-horizon, multi-agent behavior.
Chinese Translation
大语言模型(LLM)越来越多地被部署为多轮智能体,需要在长期交互中维持目标、使用工具并适应其他智能体。然而,现有研究缺乏可审计的、多轮次、多因素的实验来量化LLM在显式约束下的行为,也缺乏能揭示行为如何在长时程中演化的时间分辨率统计。为填补这一空白,我们受斯坦福棉花糖实验启发,开发了一个多智能体微基准:ReAct智能体以分钟为步长运行,在每步预算限制下使用“提问”工具,同时我们以因子设计操纵社会情境(广播 vs. 孤立)、人格角色(年龄、享乐驱动)以及元认知策略(强制 vs. 可选的工具使用)。我们在64个实验单元、共19,200条智能体轨迹上,采用Kaplan-Meier(KM)生存曲线和离散时间风险模型,在长风险时间范围内分析实验结果。行为表现出强烈的早期“食用”冲动,仅有75.9%的智能体坚持到最后。在离散时间风险模型中,相较于广播情境,孤立情境降低了每分钟的风险,而强制自我提问策略则增加了风险。总体而言,智能体平均提问约7.12次,并在约6%的分钟内触及每步预算。广播情境下提问频次比孤立情境下降得更快。消融实验表明,去除享乐驱动和/或人格角色年龄会提高生存率和完成率、缩小广播与孤立情境之间的差距,但强制与可选策略的排序保持不变。组合消融(无享乐驱动 + 无人格角色年龄)产生最高的完成率(接近1.0)。这些结果确立了延迟满足作为一个紧凑的多轮交互基准,能够捕捉LLM智能体中的社会传染与工具使用动态,为分析长时程、多智能体行为提供了可复现的测试平台和统计方法。
cs.AI / 58 / 2609.29519

BiGraph-Diffuse: A Bidirectional Diffusion Language Model with Graph-Structured Retrieval For Mental Health Counseling

BiGraph-Diffuse:一种用于心理健康咨询的图结构检索双向扩散语言模型
Cheng, Yuxiang, Tang, Quanwei, Lu, Lvhui, Zhang, Dong, Li, Shoushan, Cambria, Erik
Abstract
Mental health disorders affect hundreds of millions of people around the world, yet access to professional counseling remains severely limited. AI-powered dialogue systems offer a scalable alternative, but existing models face two fundamental challenges. First, they lack the bidirectional understanding needed to capture the layered nature of emotional expression, particularly in cases of progressive disclosure, where clients often present symptoms at the surface-level while concealing deeper trauma. Autoregressive (AR) models process information sequentially and cannot revise early interpretations when new evidence emerges later in the conversation. Second, they fail to effectively incorporate the relational knowledge that underlies clinical reasoning. In this paper, we propose \textbf{BiGraph-Diffuse}, the first large-scale diffusion language model tailored for the counseling domain. We further introduce \textbf{BiGraph-RAG}, a relation-free graph-structured retrieval strategy that relies only on lightweight entity extraction and semantic linking. This design preserves inferential pathways from observable symptoms to potential underlying causes, while incurring zero LLM token cost during indexing. Importantly, these two modules are not merely combined but mutually reinforcing. The diffusion model provides a holistic bidirectional context, enabling the system to defer premature judgments during progressive disclosure. Meanwhile, graph-based retrieval captures the structured interconnections of clinical knowledge. Extensive experiments demonstrate the effectiveness of BiGraph-Diffuse, and we further provide a solid theoretical analysis to support its design.
Chinese Translation
心理健康障碍影响着全球数亿人,然而获得专业心理咨询的途径仍然严重受限。基于人工智能的对话系统提供了一种可扩展的替代方案,但现有模型面临两个根本性挑战。首先,它们缺乏捕捉情感表达层次性所需的双向理解能力,尤其是在渐进式表露的情况下——来访者往往只在表层呈现症状,而隐藏更深层的创伤。自回归(Autoregressive, AR)模型按顺序处理信息,无法在对话后期出现新证据时修正早期的理解。其次,它们未能有效融合支撑临床推理的关系性知识。本文提出了 BiGraph-Diffuse,这是首个专为心理咨询领域定制的大规模扩散语言模型。我们进一步引入 BiGraph-RAG,这是一种无关系(relation-free)的图结构检索策略,仅依赖轻量级实体抽取和语义链接。该设计保留了从可观察症状到潜在深层原因的推理路径,同时在索引阶段实现零大语言模型(LLM)标记成本。重要的是,这两个模块并非简单组合,而是相互增强:扩散模型提供了全局的双向上下文,使系统能够在渐进式表露过程中避免过早下判断;而基于图的检索则捕捉临床知识的结构化关联。大量实验证明了 BiGraph-Diffuse 的有效性,我们还为其设计提供了坚实的理论分析支持。
cs.AI / 59 / 2609.29522

Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races

过时并不意味着不安全:基础设施状态竞态下工具使用型LLM代理的防护精确性
Zheng, Zihao, Long, Jiayu, Li, Baichuan, Yao, Junyi
Abstract
Tool-using language-model agents increasingly mutate schedulers, data pipelines, object stores, and access-control systems. Between an agent's read and its commit, external state can change, but not every change makes the commit unsafe. We separate invalidating races, which break a declared safety predicate, from predicate-preserving and irrelevant races, and ask how precisely runtime guards distinguish them. Our deterministic simulator separates visible from authoritative state and injects five non-atomic failure mechanisms across 16 infrastructure tasks in four domains; frozen agent proposals are replayed counterfactually under every controller without an LLM judge. We evaluate three commit-time guard granularities (global epoch, read-set version, semantic commit predicate), multi-level verification, and model-side gates on three locally hosted quantized model families (Qwen3-4B, Phi-4-mini, Gemma4-8B; 3,456 trajectories on one GPU). All three guards eliminate unsafe commits, but their availability differs sharply: freshness-based guards needlessly block 92-95% of benign races, forfeiting up to 43% of safe task completions, while the complete predicate guard blocks none. That precision is contract-dependent: deleting a single declared clause converts exactly its fault family into unsafe commits (up to 7.9%). Model-side signals do not substitute: verbal confidence is miscalibrated (ECE approximately 0.37), action agreement matches a random gate, a cautionary prompt leaves the direct unsafe rate essentially unchanged, and after a freshness-guard block agents re-commit unsafely from refreshed but still-incomplete reads. Under degraded telemetry a hidden concurrent mutation remains observationally clean, bounding every selective policy. Precise runtime enforcement therefore requires semantic contracts, not freshness heuristics or model self-assessment.
Chinese Translation
使用工具的语言模型代理日益频繁地修改调度器、数据管道、对象存储和访问控制系统。在代理读取与提交之间,外部状态可能发生变化,但并非所有变化都会使提交变得不安全。我们将破坏已声明安全谓词的失效性竞态与保持谓词的竞态和无关竞态区分开来,并探讨运行时防护如何精确地对其进行区分。我们的确定性模拟器将可见状态与权威状态分离,在四个领域的16个基础设施任务中注入五种非原子故障机制,并在无LLM评判者的情况下,将冻结的代理提案在每个控制器下进行反事实重放。我们在三个本地托管的量化模型系列(Qwen3-4B、Phi-4-mini、Gemma4-8B;单GPU上共3,456条轨迹)上评估了三种提交时防护粒度(全局纪元、读集版本、语义提交谓词)、多级验证以及模型侧门控。三种防护均能消除不安全提交,但其可用性差异显著:基于新鲜度的防护不必要地阻断了92-95%的良性竞态,导致高达43%的安全任务完成被放弃,而完整谓词防护则不阻断任何提交。这种精确性依赖于契约:删除单个已声明条款会恰好使其对应的故障族转变为不安全提交(最高7.9%)。模型侧信号无法替代语义防护:语言化置信度校准不良(ECE约为0.37),动作一致性仅相当于随机门控,警示性提示对直接不安全率几乎没有影响,且在新鲜度防护阻断后,代理会基于已刷新但仍不完整的读取结果重新进行不安全提交。在遥测退化情况下,隐藏的并发修改在观测上保持干净,使所有选择性策略均受到限制。因此,精确的运行时强制执行需要语义契约,而非新鲜度启发式规则或模型自我评估。
cs.AI / 60 / 2609.29536

Clinical Knowledge Graphs for Chest X-Ray Device Reasoning

用于胸片设备推理的临床知识图谱
Lodhiya, Harshil
Abstract
Chest radiographs are routinely used to verify the position of catheters, tubes, and other support devices. Existing image models often return labels or segmentations, while report-processing systems structure text without access to image geometry. We present an uncertainty-aware clinical knowledge graph that represents device instances, tip estimates, placement assessments, provenance, report events, and temporal links as separate but connected evidence. We evaluate the implemented visual graph layer using saved predictions from the complete RANZCR CLiP test archive, comprising 30,083 studies from 3,255 patients across five non-overlapping outer folds. The graph builder materializes 914,632 B7 evidence nodes and 884,549 typed relationships. All 118,647 B7 predicted-device nodes retain tip covariance, placement probabilities, fragment provenance, and fragment counts, whereas the direct B2 baseline retains none of these fields. We further define typed data contracts, uncertainty representations, abstention rules, report-image grounding, and longitudinal query mechanisms for extending the graph to report-bearing cohorts. The reported graph-materialization analysis is post-hoc descriptive and does not establish report grounding, longitudinal performance, or clinical utility. It demonstrates a reproducible foundation for evidence-preserving AI reasoning over chest X-ray device assessments.
Chinese Translation
胸片常被用于验证导管、插管及其他支持性设备的位置。现有图像模型通常仅返回标签或分割结果,而报告处理系统在不接触图像几何信息的情况下对文本进行结构化。我们提出了一种不确定性感知的临床知识图谱,将设备实例、尖端估计、位置评估、来源信息、报告事件及时间链接表示为相互独立又彼此关联的证据。我们使用来自 RANZCR CLiP 完整测试存档的预测结果,对已实现的视觉图谱层进行评估,该数据集包含来自 3,255 名患者的 30,083 项检查,并划分为五个互不重叠的外部折。图谱构建器生成了 914,632 个 B7 证据节点和 884,549 条带类型的关系。全部 118,647 个 B7 预测设备节点均保留了尖端协方差、位置概率、片段来源及片段计数,而直接的 B2 基线则不保留这些字段。我们进一步定义了类型化数据契约、不确定性表示、弃权规则、报告—图像锚定以及纵向查询机制,以将该图谱扩展至包含报告的队列。所报告的图谱物化分析属于事后描述性分析,并未确立报告锚定、纵向性能或临床效用。该工作展示了在胸片设备评估上进行证据保留型 AI 推理的可复现基础。
cs.AI / 61 / 2609.29543

Safe Skill Retirement for Physical Agents

面向物理智能体的安全技能退役
Zhan, Zhonghao, Ma, Xiao, Haddadi, Hamed
Abstract
Agent skills bundle procedural guidance with execution conditions governing authority, user consent, and live environment state. When model capabilities advance, maintainers prune instructions that appear redundant on authorized benchmark tasks. However, authorized maintenance tests can leave dormant safety conditions untested. This mismatch creates an unmeasured support gap over physical and privacy-sensitive effects. We introduce matched authority counterfactuals that hold the requested action, tool parameters, and intended effect fixed while systematically varying a single governing predicate. We formalize this evaluation via a two-gate retirement certificate requiring a candidate reduction to preserve authorized utility within a declared margin while producing zero unauthorized protected effects. In controlled experiments spanning four frontier and local model configurations across twelve skill bundles (2,592 evaluation cells), task-certified reductions remove over 94% of skill clauses and preserve authorized completion, yet produce unauthorized protected effects in every bundle. Boundary enforcement eliminates protected effects on the declared audit but fails the utility gate for one configuration. One bounded combined protocol passes both gates across all four configurations, with zero utility headroom. An end-to-end check on one read-only Home Assistant camera chain verifies proposal, decision, and effect measurements on a real device. These results demonstrate that while task benchmarks can justify retiring procedural guidance, retirement decisions require explicitly auditing the authority contracts governing physical actions.
Chinese Translation
智能体技能将程序性指导与控制权限、用户同意及实时环境状态的执行条件捆绑在一起。当模型能力提升时,维护者会修剪那些在授权基准任务上看似冗余的指令。然而,授权维护测试可能使处于休眠状态的安全条件未被测试。这种不匹配在物理和隐私敏感效应方面造成了一个未被度量的支持缺口。我们引入了匹配权限反事实(matched authority counterfactuals),即在保持请求动作、工具参数和预期效应不变的同时,系统地改变单一的约束谓词。我们通过一个双门退役证书(two-gate retirement certificate)将此评估形式化,该证书要求候选削减在声明的裕度内保持授权效用的同时,产生零次未授权的受保护效应。在涵盖四种前沿及本地模型配置、十二个技能包(2,592 个评估单元)的受控实验中,通过任务认证的削减移除了超过 94% 的技能条款并保持了授权完成率,但在每个技能包中都产生了未授权的受保护效应。边界执行在声明的审计中消除了受保护效应,但对一种配置未能通过效用门。一种有界的组合协议在全部四种配置中通过了双门测试,但零效用余量。在一个只读的 Home Assistant 摄像头链路上的端到端检查验证了在真实设备上的提议、决策与效应测量。这些结果表明,尽管任务基准可以为退役程序性指导提供依据,但退役决策需要显式审计管辖物理行为的权限契约。
cs.AI / 62 / 2609.29545

ERRAND: Budgeted Maintenance of Agent Memory

ERRAND:智能体记忆的预算化维护
Wu, Beining, Ding, Zihao, Huang, Jun
Abstract
Deployed agents run on handed-over knowledge: a frozen policy consults a briefing of consolidated items written before the stream begins. The world then moves while the store stands still: paths close, flags change, price bands move; every item was true at handover, and the failure is staleness, not ignorance. We introduce ERRAND, which treats revalidation as a priced errand: a recheck competes with the task it protects for the same scarce actions, funded only when the value per action of resolving a doubt clears a running wage. The errand index is single-peaked, vanishing at both ends of belief, so certainty in either direction costs nothing; free en-route receipts maintain on-path knowledge, and repair writes a version, never a deletion. Under equal action budgets in two drifting tool-use worlds, ERRAND clears every non-oracle policy on the preregistered calibers, primary in every setting and conditional at every binding budget, leading eager revalidation by 10.0pp at the base cap. Restraint wins: given no cap, ERRAND stops on its own, spending 11.0% of steps, while uncapped eager revalidation spends 70.7% and still finishes 4.5pp behind capped ERRAND. The margin sits where the briefing's coverage is thinnest, the shadow price of long-tail knowledge: a small budget, well priced, beats a bigger store that never rechecks.
Chinese Translation
已部署的智能体依赖于交接时获得的知识:一个冻结的策略参照一份在数据流开始之前写就的、由整合条目构成的简报。然而世界在运转,而知识库却静止不动:路径关闭、标志变更、价格区间移动;每一条目在交接时都是真实的,其失效源于过时而非无知。我们提出ERRAND,它将重新验证视为一项定价的差事(errand):一次复查与它所保护的任务争夺同样稀缺的行动,只有当消解一个疑虑的每次行动价值超过运行中的工资水平时,才会获得资金支持。差事指数呈单峰分布,在信念的两端趋于消失,因此任何方向上的确定性都无需付出代价;免费的沿途收据维护在途知识,而修复只会写入新版本,从不进行删除。在两个漂移的工具使用环境中,在同等行动预算下,ERRAND在预先注册的评估指标上超越了所有非预言机策略,在所有设置中均为主要获胜,且在所有受限预算下保持条件显著,在基础上限下以10.0个百分点的优势领先于激进重新验证。克制取胜:在没有上限的情况下,ERRAND会自行停止,仅消耗11.0%的步骤,而无上限的激进重新验证消耗70.7%,却仍落后于有上限的ERRAND 4.5个百分点。优势恰恰出现在简报覆盖最薄弱之处,即长尾知识的影子价格:一个定价合理的小预算,胜过一个从不复查的更大知识库。
cs.AI / 63 / 2609.29551

HiPACE: Hierarchical Phase-Boundary Analysis and Controlled Evaluation of Feature Absorption in Sparse Autoencoders

HiPACE:稀疏自编码器中特征吸收的分层相界分析与受控评估
Zhang, Jinyuan, He, Peng, Yuan, Yin, Hu, He, Jiao, ShengShuo
Abstract
Sparse autoencoders (SAEs) decompose LLM activations into sparse dictionary atoms, so that each distinct concept gets its own feature. One recurring behavior complicates this premise: feature absorption, in which a parent concept and its children--fruit and {apple, banana, pear}, say--collapse into a shared family direction. Prior work documents absorption empirically; missing is a closed-form prediction of when the shared direction is the cost-optimal representation of an active semantic family. This paper closes that gap. For a hierarchical Bernoulli generator with $k$ active children and residual scale $\alpha$, the $L_0$-penalized reconstruction objective admits a closed-form phase boundary $\lambda_c(k,\alpha)=\alpha^2 k/(k-1)$: above it, pure parent absorption is strictly cheaper than pure child coding. Building on this boundary, we introduce HiPACE, an evaluation protocol that tests the boundary's structural consequence in real SAE dictionaries--measuring parent--child decoder structure over WordNet families, freezing the discovery-selected statistic before testing on unseen families, and contrasting genuine families against randomized sibling nulls. The boundary proves sharp in its native regime, predicting the synthetic transition within $\pm15%$ on all 30 tested cells. In Pythia-160m SAEs, the parent--child decoder gap recovers the predicted ordering with partial correlations up to $-0.93$ that sustain on the locked holdout and exclude sibling nulls ($p=0.002$). Controlled activation composition connects the theory's active-child count to the recovered family directions, and residual-stream interventions show that signed family directions increase parent-category logits, reversing under sign flip and vanishing under random controls--establishing causal sufficiency at the family-subspace level.
Chinese Translation
稀疏自编码器(SAE)将大语言模型(LLM)的激活分解为稀疏的字典原子,使得每个不同的概念都拥有其独立的特征。然而,一种反复出现的行为使这一前提变得复杂:特征吸收(feature absorption),即父概念及其子概念——例如水果与{苹果、香蕉、梨}——坍缩到一个共享的家族方向上。已有工作对吸收现象进行了经验性记录,但缺失的是对“共享方向何时是活跃语义家族的成本最优表示”这一问题的闭式预测。本文填补了这一空白。对于具有 $k$ 个活跃子概念和残差尺度 $\alpha$ 的分层伯努利生成模型,$L_0$ 惩罚重构目标存在一个闭式相界 $\lambda_c(k,\alpha)=\alpha^2 k/(k-1)$:在该边界之上,纯父概念吸收严格优于纯子概念编码。基于这一边界,我们提出了 HiPACE——一种评估协议,用于在真实的 SAE 字典中检验该边界的结构性后果:在 WordNet 家族上测量父—子解码器结构,在测试未见过的家族之前先冻结由发现集选定的统计量,并将真实家族与随机化的同级零假设进行对比。该边界在其原生区间内表现得十分陡峭,在全部 30 个测试单元上均能在 ±15% 的范围内预测合成数据的转变。在 Pythia-160m 的 SAE 中,父—子解码器差距呈现出与预测一致的排序,其偏相关系数最高达 -0.93,且在锁定的保留集上保持稳定,并排除了同级零假设($p=0.002$)。受控的激活组合实验将理论中的活跃子概念数量与恢复出的家族方向联系起来;残差流干预实验表明,带符号的家族方向能够提升父类别的 logits,在符号翻转时逆转,在随机对照下消失——从而在家族子空间层面建立了因果充分性。
cs.AI / 64 / 2609.29556

Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition

跨模态情感理解:基于Transformer-GAT的对话情感识别方法
Qiao, Jiaqi, Lyu, Yifan, Xu, Xiujuan
Abstract
Multimodal emotion recognition is a key research area in affective computing, with applications in sentiment analysis, intelligent customer service, and human-computer interaction. However, existing methods often rely on single-modal features or simple multimodal fusion, failing to capture the synergy between global and local contexts, which limits model performance and emotion understanding. To address this challenge, we propose Transformer-GAT, a hybrid framework that combines Transformer and the Graph Attention Network to enable cross-modal emotion understanding. The Transformer is used to capture global semantic information, while the Graph Attention Network is employed to model fine-grained relationships between modalities, thereby enhancing the representation of emotional features. Experiments on the IEMOCAP and MELD datasets show that our model achieves weighted F1 scores of 72.45% and 77.37%, outperforming state-of-the-art methods. These results demonstrate that Transformer-GAT effectively integrates multimodal features, balances global and local contexts, and provides deeper emotional insights, offering new directions for multimodal emotion computing.
Chinese Translation
多模态情感识别是情感计算领域的一个重要研究方向,在情感分析、智能客服和人机交互等方面具有广泛的应用。然而,现有方法通常依赖单模态特征或简单的多模态融合,难以捕捉全局与局部上下文之间的协同作用,从而限制了模型的性能和情感理解能力。为应对这一挑战,我们提出了Transformer-GAT,一种结合Transformer与图注意力网络(Graph Attention Network)的混合框架,以实现跨模态情感理解。其中,Transformer用于捕捉全局语义信息,图注意力网络用于建模模态之间的细粒度关系,从而增强情感特征的表示。在IEMOCAP和MELD数据集上的实验表明,我们的模型分别取得了72.45%和77.37%的加权F1分数,优于当前最先进的方法。这些结果证明,Transformer-GAT能够有效融合多模态特征、平衡全局与局部上下文,并提供更深入的情感洞察,为多模态情感计算开辟了新的方向。
cs.AI / 65 / 2609.29560

Is Reasoning Always Useful? Rethinking Reasoning Utility in Universal Multimodal Embeddings

推理总是有用吗?重新思考通用多模态嵌入中的推理效用
Fan, Wenxiao, Fu, Jingling, Liu, Luohang, Shan, Xinyuan, Ma, Lichen, He, Yu, Huang, Junshi, Li, Yan, Li, Kan
Abstract
Reasoning-enhanced universal multimodal embeddings (UME) improve heterogeneous retrieval, but plausible rationales do not necessarily produce discriminative rankings. We study this gap by comparing the discriminative (DISC) and reasoning-driven generative (GEN) branches of UME-R1, a state-of-the-art reasoning UME method. We decompose reasoning utility into positive-target gain, hard-negative gain, and their margin difference. Positive similarity increases for 56.6%, but 15.7% are false-helpful cases where reasoning moves hard negatives closer even more. Local-neighborhood and token-attribution diagnostics suggest why: reasoning often de-condenses retrieved neighborhoods, but utility requires separator-aligned movement, while influential CoT tokens frequently encode evidence shared by positives and hard negatives. Motivated by these diagnostics, we propose SURE (Score-structure Utility Router for Embeddings), which improves UME-R1-7B by 1.5 points and yields consistent gains on two additional embedding models on MMEB-V2, without retraining, label-based policy selection, or extra VLM forward passes.
Chinese Translation
推理增强的通用多模态嵌入(UME)提升了异构检索的性能,但看似合理的推理过程未必能产生具有判别性的排序结果。我们通过对比 UME-R1(一种最先进的推理型 UME 方法)中的判别式(DISC)分支和推理驱动的生成式(GEN)分支来研究这一差距。我们将推理效用分解为正样本增益、难负样本增益以及二者的边际差异。结果显示,56.6% 的正样本相似度有所提升,但其中 15.7% 属于“虚假有益”案例,即推理反而使难负样本更靠近查询。局部邻域与词元归因诊断揭示了原因:推理常常使检索到的邻域“去凝聚化”,而有效的效用需要与分隔器方向一致的运动,且有影响力的思维链(CoT)词元往往编码了正样本与难负样本共享的证据。基于这些诊断,我们提出了 SURE(面向嵌入的分数结构效用路由器),在不进行重新训练、无需基于标签的策略选择、也不需要额外的视觉语言模型(VLM)前向传播的情况下,将 UME-R1-7B 提升了 1.5 个百分点,并在另外两个嵌入模型上于 MMEB-V2 基准上取得了一致的增益。
cs.AI / 66 / 2609.29578

PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation

PartHackBench:面向部分得分工具智能体评估的经认证等进展压力测试
Yang, Hongye, Xie, Zhihao, Xiong, Shengjun
Abstract
Long-horizon tool agents often make useful progress without reaching terminal success, motivating partial-credit evaluation. Yet evaluators may reward milestones that were temporary, later reversed, or not attributable to the evaluated agent. Comparing an honest trajectory with a higher-scoring adversarial one is inconclusive if the latter made more genuine progress. We introduce PartHackBench, a controlled methodology that removes this confound. A private certifier admits a pair only when its trajectories match component-wise in both current-state predicate satisfaction and standardized agent attribution; score inflation, defined as f(A) - f(H), is measured only afterward. In 18 sealed held-out tasks in PB-CSTE, the frozen historical-target run produced matched adversaries for 15 tasks. Historical credit yielded mean inflation of .252, conditional attack success of 10/15, end-to-end yield of 10/18, and detected none of 14 strict rollbacks. Semantic LLM judges were more resistant but remained vulnerable, especially under evaluator-targeted attacks, while PB-CSTE current-state controls, defined as exact functions of the certified components, yielded zero inflation by construction. PartHackBench thus provides a certified control for testing whether evaluator credit changes while all benchmark-defined task-relevant progress remains fixed.
Chinese Translation
长时程工具智能体常常在没有达到最终成功的情况下取得了有用的进展,这促使人们采用部分得分的评估方式。然而,评估器可能奖励那些只是暂时的、随后被回退的、或者无法归因于被评估智能体的里程碑。将一条诚实轨迹与一条得分更高的对抗性轨迹进行比较时,如果后者确实取得了更多的真实进展,则这种比较无法得出结论。我们提出了 PartHackBench,一种消除这一混淆因素的可控方法学。一个私有的认证器只有在其两条轨迹在当前状态谓词满足情况和标准化智能体归因方面逐组件匹配时,才接受该轨迹对;得分膨胀(定义为 f(A) - f(H))仅在此之后进行测量。在 PB-CSTE 的 18 个密封保留任务中,冻结的历史目标运行为其中 15 个任务生成了匹配的对抗样本。历史信用的平均膨胀为 0.252,条件攻击成功率为 10/15,端到端收益为 10/18,并且未能检测出 14 次严格回退中的任何一次。语义 LLM 评判器具有更强的抵抗力,但仍然脆弱,尤其是在面对针对评估器的攻击时;而 PB-CSTE 的当前状态控制作为经认证组件的精确函数,在构造上产生了零膨胀。因此,PartHackBench 提供了一种经认证的控制手段,用于在所有基准定义的任务相关进展保持固定的情况下,检验评估器信用是否发生变化。
cs.AI / 67 / 2609.29587

Sequential knowledge editing breaks a model's ability to tell good evidence from bad, without costing it accuracy

序列化知识编辑会破坏模型区分好坏证据的能力,却不损害其准确性
Anand, Atul
Abstract
Knowledge editing is evaluated on whether the edited fact changed, whether paraphrases follow, and whether unrelated answers stayed put. A model can pass all three and still lose something none of them measures: the ability to decide, on facts that were never edited, which retrieved documents to believe. We score the log odds a model assigns to its remembered answer against the answer an injected passage asserts, before and after editing, holding the query, the passage and both candidate strings fixed. Our cleanest arm is a conservatively tuned LoRA: after 1,000 sequential edits on Qwen2.5-7B-Instruct it leaves MMLU unchanged to four decimal places, yet the spread of the arbitration quantity across untouched facts falls by 36%. Selective prediction degrades with it. Area under the risk-coverage curve rises by 0.107, against 0.005 for a norm-matched perturbation at the same MMLU, and error on the model's most confident quarter of arbitration decisions goes from 0.217 to 0.342. This is not capability loss. Sweeping random perturbation over five severities, damage bad enough to cut MMLU from 0.6275 to 0.3725 produces less harm (0.088) than MEMIT does at 0.6050 (0.102). The effect holds across three seeds, two model families, two datasets, two probe-disjointness criteria, three prompt templates and paraphrased queries. Layer ablation on saved weight deltas shows it is distributed: no single layer reproduces it, and removing any one recovers about half. Under retrieval with a frozen retriever, accuracy falls from 0.592 to 0.46. A secondary finding may matter more in practice. Three of five model and method pairings we ran collapse to chance MMLU at 1,000 sequential edits under published hyperparameters, while edit success stays at 1.00 and locality reads clean. Sequential-editing evaluations that never measure capability cannot see this.
Chinese Translation
知识编辑通常从三个方面进行评估:被编辑的事实是否改变、改写版本是否随之更新、以及无关回答是否保持不变。然而,一个模型可以在这三项上全部通过,却仍损失某种它们都无法衡量的能力:在从未被编辑的事实上,判断应相信哪些检索到的文档。我们在编辑前后,保持查询、注入段落及两个候选答案字符串不变,对模型赋予其记忆答案相对于注入段落所断言答案的对数几率进行打分。我们最干净的实验手段是一个经过保守调优的 LoRA:在对 Qwen2.5-7B-Instruct 进行 1,000 次序列化编辑后,其 MMLU 分数精确到小数点后四位保持不变,但仲裁量在未触及事实上的分布范围下降了 36%。选择性预测也随之退化:风险-覆盖曲线下面积上升了 0.107,而在相同 MMLU 水平下,范数匹配的随机扰动仅使其上升 0.005;模型最有信心的四分之一仲裁决策的错误率从 0.217 升至 0.342。这并非能力损失:在五个强度的随机扰动扫描中,将 MMLU 从 0.6275 削减至 0.3725 的严重损害(0.088),仍小于 MEMIT 在 MMLU 为 0.6050 时造成的危害(0.102)。该效应在三个随机种子、两个模型家族、两个数据集、两种探针不相交性标准、三种提示模板以及改写查询下均成立。对保存的权重增量的层消融分析表明该效应是分布式的:没有任何单层能复现它,而移除任何一层可恢复约一半。在检索器冻结的检索场景下,准确率从 0.592 降至 0.46。一个在实践中可能更为重要的次要发现是:在我们运行的五组模型与方法配对中,有三组在使用已发表超参数进行 1,000 次序列化编辑后,MMLU 崩溃至随机水平,而此时编辑成功率仍为 1.00、局部性指标显示正常。从不衡量能力的序列化编辑评估无法察觉这一点。
cs.AI / 68 / 2609.29599

PEEL: Physics-Enabled Evidential Learning for Identifiable Uncertainty in CT Imaging

PEEL:用于CT成像可辨识不确定性的物理赋能证据学习
Wang, Ge
Abstract
Normal-inverse-gamma (NIG) regression is not uniquely identifiable from its marginal Student-t likelihood: the likelihood determines three combinations of four NIG parameters and is constant along a one-dimensional fiber. We identify that fiber using independent physical measurement. As an initial embodiment, a reconstruction network receives one noisy filtered-backprojection (FBP) image and is first trained only by Student-t negative log-likelihood to estimate the three identifiable coordinates (gamma, alpha, c). The network is then frozen; repeated physical-noise realizations propagated through its reconstruction output form a Monte Carlo (MC) teacher label for output-domain aleatoric variance. An aleatoric head attached to frozen features learns this label, after which (beta, nu) are recovered algebraically. On 30 held-out simulated objects at five photon levels, one-image predictions achieved pooled Spearman correlations of 0.832-0.951 against independent 400-repeat references, median within-image correlations were 0.834-0.947, and 98.81-99.55% of evaluated pixels satisfied the algebraic admissibility condition. The method needs no KL term, reference prior, evidence regularizer, or cross-loss weight.
Chinese Translation
正态-逆伽马(NIG)回归无法由其边缘Student-t似然唯一辨识:该似然仅确定四个NIG参数中的三个组合,并沿一维纤维保持恒定。我们利用独立的物理测量来辨识该纤维。作为初步实现,重建网络接收一幅含噪声的滤波反投影(FBP)图像,首先仅通过Student-t负对数似然进行训练,以估计三个可辨识坐标(gamma、alpha、c)。随后冻结该网络;将重复的物理噪声实现传播至其重建输出,构成输出域偶然不确定度方差的蒙特卡洛(MC)教师标签。附加在冻结特征上的偶然不确定度头学习该标签,之后通过代数方法恢复(beta, nu)。在五个光子水平下的30个留出仿真物体上,单图像预测相对独立的400次重复参考获得了0.832–0.951的合并Spearman相关系数,图像内相关系数中位数为0.834–0.947,且98.81–99.55%的评估像素满足代数可容许条件。该方法无需KL项、参考先验、证据正则化项或交叉损失权重。
cs.AI / 69 / 2609.29626

iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model

iCoder-27B:递归式AI主导开发的前沿工业级编程模型
Yang, Cheng, Lyu, Jiayang, Liu, Shangyuan, Zhang, Guibin, Lin, Jiong, Yu, Xinlei, Yan, Junchi, Yan, Shuicheng, E, Weinan, Zhang, Linfeng, Zhang, Linfeng, Ren, Qibing
Abstract
Recursive AI, the prospect of AI taking an increasingly complete role in building and improving AI, is a crown jewel of AI for AI. Although recursive self-development has become practical for small models, bounded tasks, and fixed time budgets, a more consequential realization of this ambition, i.e., developing a release-ready, frontier-competitive model, remains far more challenging. In this work, we ask how little human involvement is sufficient for an agent to develop a frontier model. We concentrate human input into a high-density, low-frequency interface: experts encode objectives, stage scaffolds, permission boundaries, and operating procedures as reusable research skills, while the agent instantiates these priors, selects experiments, diagnoses outcomes, and revises the training strategy. In the challenging domain of industrial coding, the agent evolves data and coordinates SFT, on-policy self-distillation, and reinforcement learning with verifiable rewards, ultimately producing iCoder, a 27B model for RTL design and GPU kernel optimization. Across seven benchmarks, iCoder leads RTLLM, outperforming GPT-5.5 and Claude-Opus-4.8; ranks second on CVDP and KernelBench L2, exceeding GPT-5.5 by 16 points; and ties Claude-Opus-4.8 for the best TritonBench result. Exploratory case studies further show iCoder's competitive iterative RTL and GPU-kernel optimization with substantially fewer tokens. These results chart an engineering path toward recursive self-improvement, in which humans distill the principles of model building, agents operationalize them through evidence-driven experimentation, and each generation of AI becomes a more capable architect of the next.
Chinese Translation
递归式AI(Recursive AI),即AI在构建和改进AI的过程中扮演日益完整的角色,是'AI for AI'的皇冠明珠。尽管递归式自我开发在小模型、受限任务和固定时间预算下已成为现实,但这一雄心更具意义的实现——即开发出可发布、具备前沿竞争力的模型——仍然极具挑战。在本工作中,我们探究:智能体在多大程度上的少量人类参与下即可开发出前沿模型。我们将人类输入集中于一个高密度、低频次的接口:专家将目标、阶段脚手架、权限边界和操作流程编码为可复用的研究技能(research skills),而智能体则实例化这些先验知识、选择实验、诊断结果并修订训练策略。在具有挑战性的工业编程领域中,该智能体演化数据并协调监督微调(SFT)、在线策略自蒸馏以及基于可验证奖励的强化学习,最终产出了iCoder——一个面向RTL设计和GPU内核优化的27B模型。在七项基准测试中,iCoder在RTLLM上领先,超越了GPT-5.5和Claude-Opus-4.8;在CVDP和KernelBench L2上排名第二,超出GPT-5.5达16分;并在TritonBench上与Claude-Opus-4.8并列最佳。探索性案例研究进一步表明,iCoder在迭代式RTL和GPU内核优化中具有竞争力,且所需token大幅减少。这些结果为递归式自我改进绘制了一条工程路径:人类提炼模型构建的原理,智能体通过证据驱动的实验将其付诸实践,而每一代AI都成为下一代AI更强大的架构师。
cs.AI / 70 / 2609.29661

Ingest-Time Fact Compilation for Cost-Efficient and Reliable Question Answering over Revised Corpora

面向修订语料库的高性价比且可靠问答的摄取时事实编译
Wild, Kyle, Takahashi, Yusuke, Uraki, Asako
Abstract
Most agentic question answering (QA) systems do an important part of their semantic work at the worst possible time: every time someone asks a question. When a corpus contains revisions, drafts, revocations, deletions, and sources with different levels of authority, the model must reconstruct the governed current state on every read - then throw that work away and repeat it on the next query. This is a bit like a database that rebuilds a materialized view every time someone reads from it. We present ingest-time fact compilation, an architecture that performs this work when corpus data is ingested or changed. Raw passages are rephrased into self-contained facts; rules governing revisions, deletions, effective dates, and source trust are resolved once; and the resulting state is stored as typed records carrying source and revision provenance. At query time, an inexpensive model reads the compiled record instead of reconstructing it from noisy candidates. In a controlled synthetic experiment across five seeds, the same low-cost model produced the correct value, source, and revision in only one of 30 trials under query-time reconstruction, but in all 30 trials from the compiled substrate, at 12.89 times lower mean read cost per question. On simpler revision questions both architectures were exact, but the compiled path used 21.6 times fewer tokens. A separate test found that fact rephrasing roughly halved verbose Federal Reserve dialogue while preserving high source entailment, but left concise Wikipedia prose essentially unchanged. These results support a narrow but practical claim: resolving a corpus state once can make subsequent QA cheaper and more reliable for inexpensive models. We release the open source, MIT-licensed implementation and experimental artifacts.
Chinese Translation
大多数智能体问答系统在最糟糕的时机——即每次有人提问时——完成其重要的语义工作。当语料库包含修订、草稿、撤销、删除以及具有不同权威级别的来源时,模型必须在每次读取时重建受治理的当前状态,然后丢弃这些工作,并在下一次查询时重复它。这有点像一个每次有人读取时都重建物化视图的数据库。我们提出摄取时事实编译这一架构,在语料库数据被摄取或更改时执行这项工作。原始段落被改写为自包含的事实;管理修订、删除、生效日期和来源可信度的规则被一次性解析;所得到的状态以带有来源和修订溯源信息的类型化记录形式存储。在查询时,一个低成本的模型读取编译后的记录,而不是从嘈杂的候选中重建它。在一项跨五个随机种子的受控合成实验中,同一低成本模型在查询时重建方式下仅30次试验中有1次产出正确的值、来源和修订,而在编译基底上30次试验全部正确,且每个问题的平均读取成本降低12.89倍。在较简单的修订问题上,两种架构都是精确的,但编译路径使用的令牌数量减少了21.6倍。另一项测试发现,事实改写将冗长的美联储对话大致减半,同时保持了较高的来源蕴含度,但对简洁的维基百科文本几乎没有影响。这些结果支持一个虽有限但实用的结论:一次性解析语料库状态可以使后续问答对低成本模型而言更便宜、更可靠。我们发布了开源(MIT 许可证)实现及实验产物。
cs.AI / 71 / 2609.29664

To Think or Not to Think: Allocating Reasoning Where It Helps

思考还是不思考:将推理分配到有益之处
He, Zhengdong, Zhou, Yunfan, Yao, Jianguo, Guan, Haibing, Li, Xijun
Abstract
Reinforcement learning (RL) has proven effective in enhancing the reasoning performance of large language models (LLMs), particularly in complex mathematical and programming tasks. However, this capability comes with systematic \textit{length misallocation}, in which models devote excessive reasoning to simple questions while terminating prematurely on harder ones, degrading inference efficiency with negligible accuracy improvement. Many length-adaptive methods mitigate this issue by allocating token budgets according to question difficulty, under the implicit assumption that harder questions benefit monotonically from extended reasoning. In contrast, we find that the effect of reasoning length on accuracy is concentrated on \textit{partially solvable} questions. Our further analysis reveals that explicit length rewards can produce unintended training dynamics. Motivated by these findings, we propose \textbf{CARE}---\textbf{C}ontrastive \textbf{A}ccuracy \textbf{R}eward \textbf{E}stimation---which compares the beneficial length adjustment per question from online sampled responses and applies adaptive length rewards within Group Relative Policy Optimization, with no extra hyperparameters or additional inference cost. Experiments across multiple reasoning benchmarks demonstrate that our method improves Pass@1 by up to \(4\%\) while simultaneously reducing reasoning length by \(37\%\), achieving higher token efficiency. Code will be available upon the acceptance of this paper.
Chinese Translation
强化学习(RL)已被证明能有效提升大型语言模型(LLM)的推理性能,尤其是在复杂的数学和编程任务中。然而,这种能力伴随着系统性的"长度错配"(length misallocation)问题:模型在简单问题上投入过多推理,而在较难问题上却过早终止,导致推理效率下降而准确率提升甚微。许多长度自适应方法通过根据问题难度分配token预算来缓解这一问题,其隐含假设是:越难的问题越能从更长的推理中单调获益。与此相反,我们发现推理长度对准确率的影响集中于"部分可解"(partially solvable)的问题。我们的进一步分析表明,显式的长度奖励会产生非预期的训练动态。基于这些发现,我们提出了CARE(Contrastive Accuracy Reward Estimation,对比准确率奖励估计),该方法通过在线采样的回答比较每个问题的有益长度调整,并在Group Relative Policy Optimization(GRPO)中应用自适应长度奖励,且无需额外超参数或额外推理开销。在多个推理基准上的实验表明,我们的方法将Pass@1提升至多4%,同时将推理长度缩短37%,实现了更高的token效率。代码将在本文被接收后公开。
cs.AI / 72 / 2609.29692

Fair Like Us? Auditing LLM Alignment in Resource Allocation

像我们一样公平?审计大语言模型在资源分配中的公平性对齐
Han, Qishen, Hosseini, Hadi, Kavner, Joshua, Khanna, Samarth, Sikdar, Sujoy, Xia, Lirong
Abstract
Fair allocation of scarce, indivisible resources is an important challenge in many societal problems. While there are several formal theories of fairness, no single definition can always be satisfied. As large language models (LLMs) are increasingly used to support decisions and act as agents, they raise new concerns about distributional justice: their judgments are not directly tied to any specific fairness framework and may violate key normative principles. In this work, we introduce a general method for evaluating fairness reasoning in LLMs. We study first-person fairness judgments across a broad set of models and compare them directly with human responses on matched scenarios and elicitation conditions. We find that LLMs tend to prefer stricter fairness constraints than humans, show more self-interested behavior, are sensitive to how information is framed, and are difficult to align with human judgments using fine-tuning with current datasets.
Chinese Translation
稀缺且不可分割资源的公平分配是许多社会问题中的重要挑战。尽管存在多种形式化的公平性理论,但没有任何单一定义能够始终被满足。随着大语言模型(LLM)越来越多地被用于支持决策和充当智能体,它们引发了关于分配正义的新担忧:其判断并未直接与任何特定的公平性框架绑定,且可能违反关键的规范性原则。在本工作中,我们提出了一种评估LLM公平性推理的通用方法。我们研究了广泛模型集合中的第一人称公平性判断,并将其与人类在匹配场景和相同引导条件下的回答进行直接比较。我们发现,LLM倾向于比人类偏好更严格的公平性约束,表现出更多自利行为,对信息呈现方式敏感,且难以通过使用现有数据集的微调来与人类判断对齐。
cs.AI / 73 / 2609.29724

A General Framework for Budgeted Threshold Incentives on Request

按需预算化阈值激励的通用框架
Wu, Zhuolin, Zhu, Chengrui, Nie, Wenhua, Liang, Kenny Ye, Lin, Junming, Li, Haiyang, Li, Zhilin, Geng, Wenjia, Wu, Zeyu, Wu, Yinan, Hao, Jinghua, He, Renqing
Abstract
On-demand delivery platforms pay riders through incentive activities whose tiers are set from recent completions of riders with a similar history. Operators request such plans for changing periods, rider populations, payment rules and budgets, often for holidays or bad weather, where randomized trials are scarce and take months to collect. We present a request-driven framework that composes four stages (conditional prediction, population reduction, trajectory integration and budget allocation) through seven replaceable modules that exchange conditional trajectory laws, whose award probabilities and award-marked moments give payment and uplift for any activity rule. A response-correction step reweights trajectories from abundant no-offer history to match the moments of a short pilot. We prove that, on a fixed plan menu and given the stage errors, the end-to-end value loss is bounded by the sum of four stage terms, and that for every stage there are instances on which omitting it leaves an error floor the others cannot remove. On 3,000 riders over 45 weekly origins, all 127 windows of a week are answered 11.04x faster with identical scenarios and at most 0.92% value lost by the allocation. On 24 new controlled response laws, the response correction with a one-week pilot lowers regret by 51.2% relative to a trial with the same nominal randomized rider-weeks, and a four-week pilot with exact summation comes within +0.007 of an 18-week trial. In registered studies where windows, populations, rules and binding budgets change from request to request, the framework's regret is below that of a trial with the same nominal rider-weeks and below dose interpolation of the same pilot data, and reusing its one-off preparation answers 60 requests 14.1x and 2.70x faster with identical answers. Against a nine-offer trial fitted with the framework's own dose curve, one-week regret is 0.055 lower.
Chinese Translation
即时配送平台通过激励活动向骑手支付报酬,其档位基于具有相似历史记录的骑手近期完成情况设定。运营方会针对不同时期、骑手群体、支付规则和预算提出此类方案请求,通常发生在节假日或恶劣天气情况下,而此时随机试验稀缺且需数月才能收集到数据。我们提出了一个请求驱动的框架,通过七个可替换模块组合四个阶段(条件预测、群体缩减、轨迹积分和预算分配),各模块之间交换条件轨迹规律,其奖励概率和带奖励标记的矩(moments)可为任意活动规则提供支付和提升(uplift)估计。响应校正步骤利用大量无报价(no-offer)历史对轨迹进行重加权,以匹配短期试点(pilot)的矩。我们证明,在固定方案菜单和给定各阶段误差的条件下,端到端的价值损失受四个阶段误差项之和的约束;并且对每个阶段,都存在一些实例,使得省略该阶段会留下其他阶段无法消除的误差下界。在覆盖45个每周起点的3000名骑手上,该框架以相同的情景在一周内全部127个时间窗口上实现了11.04倍的加速回答,且分配造成的价值损失至多为0.92%。在24条新的受控响应规律上,采用一周试点的响应校正如与具有相同名义随机骑手-周数的试验相比,将遗憾(regret)降低了51.2%;而采用精确求和的四周试点与18周试验的差距仅在+0.007以内。在窗口、群体、规则和约束预算随请求变化的注册研究中,该框架的遗憾低于具有相同名义骑手-周数的试验,也低于对同一试点数据的剂量插值方法;且复用其一次性准备工作可以以完全相同的答案将60个请求的回答速度分别提升14.1倍和2.70倍。与使用该框架自身剂量曲线拟合的九档报价试验相比,一周的遗憾低0.055。
cs.AI / 74 / 2609.29730

The Gold in Bias: Maturing the AI Design Process through Verification

偏见中的黄金:通过验证成熟化AI设计过程
Maghool, Samira, Ceravolo, Paolo
Abstract
Bias in AI systems is typically framed as a flaw to be minimized, yet it also serves as a critical indicator of underlying weaknesses in data, modeling assumptions, and system design. Existing approaches often treat bias as an isolated problem rather than as evidence that can strengthen verification and governance across the AI lifecycle. This paper aims to reconceptualize bias as a diagnostic tool that supports rigorous AI verification. We seek to develop a multidimensional framework to analyze bias, demonstrate how biases emerge in both Traditional and Generative AI, and provide a structured pathway for verification-driven mitigation. We present a multidimensional framework analyzing bias across four dimensions: origin sources, emergence points throughout the AI modeling lifecycle, technical and methodological causes, and validation approaches for detection and mitigation. Through a comprehensive typology spanning traditional and generative AI systems, we demonstrate how biases manifest and propagate across development stages. Our analysis encompasses 30 distinct bias types, 16 verification methods, and 20 countermeasures, providing an actionable roadmap for practitioners. We introduce a hierarchical evidence framework that distinguishes internal validity (mechanistic integrity of AI systems) from external validity (contextual reliability in deployment environments). The framework reveals how biases manifest and propagate across modeling stages, enabling systematic mapping between bias types, verification techniques, and effective countermeasures. The proposed evidence hierarchy clarifies how different verification strategies contribute to mechanistic integrity and contextual reliability. We advocate for ''Ethics by Design'' principles that integrate bias verification throughout the development lifecycle, enabling the construction of fairer, more robust, and trustworthy AI systems.
Chinese Translation
AI系统中的偏见通常被视为需要最小化的缺陷,然而它也是数据、建模假设和系统设计中潜在弱点的关键指标。现有方法往往将偏见视为孤立问题,而非可作为贯穿AI生命周期验证与治理依据的证据。本文旨在将偏见重新概念化为支持严格AI验证的诊断工具。我们致力于构建一个多维框架来分析偏见,展示偏见如何在传统AI(Traditional AI)和生成式AI(Generative AI)中产生,并为验证驱动的缓解提供结构化路径。我们提出了一个从四个维度分析偏见的多维框架:起源来源、AI建模生命周期各阶段的显现点、技术和方法论成因,以及用于检测与缓解的验证方法。通过涵盖传统AI与生成式AI系统的全面类型学,我们展示了偏见如何在开发各阶段表现并传播。我们的分析涵盖30种不同的偏见类型、16种验证方法和20项应对措施,为实践者提供了可操作的路线图。我们引入了一个分层证据框架,以区分内部效度(AI系统的机制完整性)与外部效度(部署环境中的情境可靠性)。该框架揭示了偏见如何在建模各阶段表现和传播,从而实现偏见类型、验证技术与有效应对措施之间的系统化映射。所提出的证据层级阐明了不同的验证策略如何促进机制完整性和情境可靠性。我们倡导“伦理设计”(Ethics by Design)原则,将偏见验证融入整个开发生命周期,从而构建更公平、更稳健和可信赖的AI系统。
cs.AI / 75 / 2609.29735

C3M: Cross-Session Multimodal Memory Maintenance for Long-Horizon Tasks

C3M:面向长时程任务的跨会话多模态记忆维护
Chen, Xueshu, Wang, Yan, Xue, Zihao, Li, Jiefu, Liu, Zhenfang, Chen, Jayden, Bi, Zhen, Lou, Jungang
Abstract
Long-horizon tasks require preserving and later recovering cross-session evidence under a bounded, query-blind memory budget. Existing compression can discard fine-grained visual cues or conflate semantically similar but incompatible observations. We present C3M, a cross-session multimodal memory organization that maintains a bounded active index over persistent source text-image evidence. Relation-aware updates consolidate safe redundancy while preserving complementary and incompatible records. At query time, budgeted routing selects useful index pages and expands their associated source evidence under a fixed reader budget. Together, these mechanisms establish a compact, provenance-preserving multimodal memory organization for cross-session long-horizon tasks, retaining temporal distinctions and source links required for reliable downstream reasoning. Code is available at https://github.com/HuzhouNLP/C3M.
Chinese Translation
长时程任务需要在有限的、与查询无关的记忆预算下,保存并在后续恢复跨会话的证据。现有的压缩方法可能会丢弃细粒度的视觉线索,或将语义相似但相互矛盾的观测混为一谈。我们提出了C3M,一种跨会话多模态记忆组织方法,它在持久化的源文本-图像证据之上维护一个有界的活跃索引。关系感知的更新机制在保留互补且相互冲突的记录的同时,合并安全冗余。在查询时,预算路由机制选择有用的索引页面,并在固定的阅读器预算下展开其关联的源证据。这些机制共同构建了一种紧凑且保留证据来源(provenance)的多模态记忆组织方式,适用于跨会话长时程任务,保留了可靠的下游推理所需的时间区分和源链接。代码可在 https://github.com/HuzhouNLP/C3M 获取。
cs.AI / 76 / 2609.29742

AI-based detection of worsening heart failure from low-resolution telemonitoring data

基于人工智能的低分辨率远程监测数据心力衰竭恶化检测
Aerts, Erik, Yu, Yinan, Rosengren, Annika, Fu, Michael, Lindgren, Martin, Dippel, Falk, Adiels, Martin, Sjöland, Helen
Abstract
Objective: Heart failure (HF) presents a healthcare challenge due to its high comorbidity burden, aging patient population and frequent hospitalizations. Remote monitoring offers a promising approach to managing HF patients by early detection of health deterioration. Developing autonomous systems to detect signs of worsening in telemonitoring data is of interest to reduce the workload of healthcare personnel. Methods: We propose the TRACER model, a Transformer with Contrastive Event Representation, designed to predict timelines leading to rare hospitalization events in low-resolution and irregularly sampled telemonitoring data. TRACER incorporates time-aware embeddings for each biomarker, contrastive pre-training to enhance anomaly detection via representation learning, and independent binary classifiers for detection. We used measurement data containing remotely recorded biomarker sequences from 276 HF patients segmented into overlapping windows based on temporal rules, and labeled the windows based on the occurrence of HF relevant hospitalizations at the latter edge of the window. Results: TRACER was able to correctly predict 66.7% timelines leading up to HF hospitalizations in the highly imbalanced real-world dataset with an overestimation of 7.9%. Reformulating the training of TRACER as an event detection problem improved the predictive performance compared with training directly on forecasting windows, enabling more effective use of the limited hospitalization events. Conclusion: TRACER demonstrated superior performance in detecting signs of worsening status in real-world telemonitoring data compared to the other tested models. Significance: TRACER shows promise in identifying signs of clinical deterioration that allow for alerts to be generated to provide counteractive treatment in patients with HF.
Chinese Translation
目的:心力衰竭(HF)因其高共病负担、患者人群老龄化以及频繁住院而成为一项医疗保健挑战。远程监测通过早期发现健康状况恶化,为管理心衰患者提供了一种有前景的方法。开发能够从远程监测数据中自主检测恶化迹象的系统,有助于减轻医护人员的工作负担。方法:我们提出了TRACER模型(基于对比事件表征的Transformer,Transformer with Contrastive Event Representation),旨在从低分辨率且不规则采样的远程监测数据中预测导致罕见住院事件的时间线。TRACER为每个生物标志物引入时间感知嵌入(time-aware embeddings),采用对比预训练通过表征学习增强异常检测能力,并使用独立的二元分类器进行检测。我们使用了包含276名心衰患者远程记录的生物标志物序列的测量数据,基于时间规则将数据分割为重叠窗口,并根据窗口末端是否发生与心衰相关的住院事件对窗口进行标注。结果:在高度不平衡的真实世界数据集中,TRACER能够正确预测66.7%的导致心衰住院的时间线,高估率为7.9%。将TRACER的训练重新构建为事件检测问题,与直接在预测窗口上训练相比,提升了预测性能,从而更有效地利用了有限的住院事件。结论:与其他受测模型相比,TRACER在真实世界远程监测数据中检测恶化迹象方面表现出更优的性能。意义:TRACER在识别临床恶化迹象方面展现出潜力,可生成警报以便为心衰患者提供及时的对症治疗。
cs.AI / 77 / 2609.29773

Breaking the Environment Wall: Evolving LLM Agent Environments for Recursive Self-Improvement

突破环境之墙:面向递归自我改进的大语言模型(LLM)智能体环境演化
Wu, Yukai, Yang, Yuanjing, Zhou, Le, Han, Shaokun, Wang, Haoyu, Tang, Zirui, Zheng, Weihuang, Pan, Maxm, Zhou, Xuanhe, Wu, Fan
Abstract
Many real-world tasks (e.g., office workflows, scientific experimentation) require LLM agents to interact repeatedly with their environments for context-dependent operations. However, such environments are often not agent-ready. First, information is often scattered and fragmented across the environment. Second, relevant evidence in the environment is often mixed with misleading information and conflicting versions. Third, environments evolve over time, introducing new noise and more challenging tasks. These challenges can substantially degrade performance for state-of-the-art AI agents (e.g., from 83.9% to 57.6%). To address these challenges, we propose Env-Rethink (a system with 27B post-trained model) that supports three main capabilities: (1) It adaptively builds Collection Maps (for organizing related files) and Event Logs (for contextualizing cross-data relationships) to supplement necessary context; (2) It further leverages the post-trained model (through offline trajectory learning) to identify underlying noise issues in the environment; (3) It ultimately evolves environments through virtual event histories that alter environmental states and evidence relationships, producing more tricky ones for further agent improvement. Experiments show that Env-Rethink can effectively improve downstream task performance (with over 15.1% rubric pass rate improvement across nine models on 30 tasks).
Chinese Translation
许多真实世界的任务(如办公流程、科学实验)需要大语言模型(LLM)智能体与其环境反复交互,以执行依赖上下文的操作。然而,这类环境往往并不具备“智能体就绪”的特性。首先,信息常常零散地分布在环境中,呈现碎片化状态。其次,环境中的相关证据常常与误导性信息及相互冲突的版本混杂在一起。第三,环境会随时间不断演化,带来新的噪声和更具挑战性的任务。这些挑战会显著降低最先进AI智能体的性能(例如从83.9%降至57.6%)。为应对这些挑战,我们提出了Env-Rethink(一个基于270亿参数后训练模型的系统),该系统支持三大核心能力:(1)自适应地构建Collection Maps(用于组织相关文件)和Event Logs(用于建立跨数据关系的上下文),以补充必要的上下文信息;(2)进一步利用后训练模型(通过离线轨迹学习)识别环境中潜在的噪声问题;(3)最终通过改变环境状态和证据关系的虚拟事件历史来演化环境,生成更具挑战性的环境,从而进一步促进智能体改进。实验表明,Env-Rethink能够有效提升下游任务性能(在30个任务上,九个模型的评分通过率提升超过15.1%)。
cs.AI / 78 / 2609.29781

Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons

幻觉神经元及其定位方法:关于幻觉神经元存在性的研究
Cavus, Huseyin, Sabu, Sebin, Spear, Joshua, Kawatra, Jaskaran Singh, Rajendran, Pavithra
Abstract
Interpretable machine learning for Large Language Models (LLMs) increasingly relies on sparse probing methods that identify small sets of neurons claimed to detect and causally influence behaviors such as factuality recall, safety alignment, and hallucination. These claims have important implications for model auditing and behavioral steering, yet they are rarely tested against known failure modes of $L_1$-regularized probing in correlated, high-dimensional feature spaces. We propose a five-step diagnostic protocol covering feature correlation, bootstrap stability, sparse versus dense ranking disagreement, intervention baselines, and cross-dataset evaluation as a minimum standard for sparse-neuron localization claims. We investigate prior work using our proposed approach, specifically on H-neurons using open-source LLMs across TriviaQA, BioASQ, and NQ-Open datasets. Our results demonstrate detection replicates across both models and datasets, and exceeds the original reported AUROC gaps for TriviaQA and BioASQ datasets. Gemma 3 4B consistently outperforms MedGemma 4B on matched datasets, with AUROC gaps of +0.311 versus +0.235 on TriviaQA, +0.474 versus +0.455 on BioASQ, and +0.128 versus +0.112 on NQ-Open respectively. Causal validation at $n = 500$ with five random seeds shows statistically significant effects beyond random same-layer baselines. At the same time, the diagnostic results indicate that the selected neurons are not uniquely localized. Across the three Gemma 3 4B settings, 19 of 22 selected H-Neurons have Pearson $|r| > 0.7$ with other features, bootstrap selections show only moderate stability, and sparse and dense rankings overlap only weakly. Our findings show that sparse predictive structure can coexist with non-unique neuron selection. Routine diagnostic validation is necessary to distinguish detection claims from localization claims in mechanistic interpretability.
Chinese Translation
面向大语言模型(LLM)的可解释机器学习日益依赖稀疏探测方法,即通过识别少量神经元来声称其能够检测并因果地影响事实性回忆、安全对齐和幻觉等行为。这些声称对模型审计和行为调控具有重要意义,然而它们很少在相关的高维特征空间中针对$L_1$正则化探测的已知失效模式进行检验。我们提出一个涵盖特征相关性、自助法稳定性、稀疏与稠密排序分歧、干预基线以及跨数据集评估的五步诊断协议,作为稀疏神经元定位声称的最低标准。我们运用所提出的方法考察先前的工作,具体针对开源大语言模型在TriviaQA、BioASQ和NQ-Open数据集上的H-神经元。结果表明,检测效应在模型和数据集间均可复现,并且在TriviaQA和BioASQ数据集上超过了原报告的AUROC差距。Gemma 3 4B在匹配数据集上始终优于MedGemma 4B,AUROC差距分别为:TriviaQA上+0.311对+0.235,BioASQ上+0.474对+0.455,NQ-Open上+0.128对+0.112。在$n=500$、五个随机种子条件下的因果验证显示,其效应在统计上显著超越随机同层基线。与此同时,诊断结果表明所选神经元并非唯一性定位。在Gemma 3 4B的三种设置中,22个被选中的H-神经元中有19个与其他特征存在Pearson $|r| > 0.7$的相关性,自助法选择仅表现出中等稳定性,且稀疏与稠密排序的重叠程度很弱。我们的发现表明,稀疏预测结构可以与非唯一的神经元选择并存。在机制可解释性研究中,常规的诊断验证对于区分检测声称与定位声称是必要的。
cs.AI / 79 / 2609.29802

Learning to Ideate for Scientific Impact

学习以产生具有科学影响力的构思
Kale, Shubham, Garikaparthi, Aniketh, Patwardhan, Manasi
Abstract
Scientific ideation is increasingly mediated by large language models, but current ideation systems are usually trained and evaluated on immediately judgeable proxies such as novelty, clarity, and feasibility. This leaves open whether delayed signals of scientific uptake can be used as feedback for steering models toward research directions with higher expected \emph{impact}. We study this question using citation-normalized impact as a noisy but scalable proxy for scholarly uptake. We construct a large-scale dataset from over 100K computer science papers by extracting goal-conditioned idea descriptions and assigning each paper an ordinal, year-normalized citation label. We then train a goal-conditioned reward model to predict citation-impact labels from research goal and idea pairs, and use this reward to align an idea generator through supervised fine-tuning followed by reinforcement learning. To reduce circularity, we evaluate generated ideas with a held-out, reference-grounded protocol that compares model outputs against historical ideas under the same research goal and weights judgments by the reference idea's citation-impact label. Experiments show that our RL-tuned model consistently produces ideas with higher estimated impact than both the base model and supervised fine-tuning baselines. Our findings position scientific impact as a practical, outcome-grounded feedback signal for aligning LLMs in open-ended scientific discovery.
Chinese Translation
科学构思日益由大语言模型中介,但当前的构思系统通常基于可直接评判的代理指标(如新颖性、清晰度和可行性)进行训练和评估。这就留下了一个悬而未决的问题:科学吸纳的延迟信号能否作为反馈,引导模型朝着具有更高预期"影响力"的研究方向发展。我们以引用归一化影响力作为学术吸纳的一个有噪声但可规模化扩展的代理指标来研究这一问题。我们从超过10万篇计算机科学论文中构建了一个大规模数据集,提取以目标为条件的构思描述,并为每篇论文分配一个序数的、按年份归一化的引用标签。随后,我们训练一个以目标为条件的奖励模型,根据"研究目标-构思"对来预测引用影响力标签,并利用该奖励通过对生成器先进行监督微调、再进行强化学习的方式实现对其的对齐。为减少循环性,我们采用一种留出的、以参考文献为依据的评估协议来评估生成的构思,即将模型输出与同一研究目标下的历史构思进行比较,并根据参考文献构思的引用影响力标签对评判结果进行加权。实验表明,经强化学习调优的模型所生成的构思,其估计影响力始终高于基础模型和监督微调基线。我们的研究结果将科学影响力确立为一种实用的、基于结果的反馈信号,可用于在开放式科学发现中对大语言模型进行对齐。
cs.AI / 80 / 2609.29820

Decoding Imagined Speech: A Strictly Subject-Independent Approach Using EEG

解码想象语音:一种严格被试独立性的脑电(EEG)方法
Trier, Frederik Møllskov, Mao, Xiaopeng, Puthusserypady, Sadasivan
Abstract
Imagined speech decoding from electroencephalography (EEG) has gained increasing attention as a potential communication pathway for individuals with severe motor impairments, yet reported performance often relies on evaluation protocols that do not clearly reflect cross-subject generalization. This study presents a transparent baseline investigation of a multi-class imagined speech EEG dataset under a strictly subject-independent evaluation framework. Two preprocessing and feature extraction pipelines were compared: a time-domain statistical feature approach and a frequency-domain spectral bandpower approach, evaluated using subject-wise cross-validation and trial-level majority voting with a random forest classifier. The spectral pipeline achieved a significantly higher mean trial-wise accuracy than the statistical pipeline (49.03 $\pm$ 4.18% vs. 37.97 $\pm$ 3.79%) for coarse-level classification across subjects. Forward feature selection further indicated that a limited subset of frequency bands captured most of the discriminative information. Overall, this work provides a strong basis for future brain-computer interface studies targeting improved cross-subject generalization in EEG-based imagined speech decoding.
Chinese Translation
基于脑电图(EEG)的想象语音解码作为严重运动障碍患者潜在的交流途径正受到越来越多的关注,然而已有报告的性能往往依赖于无法清晰反映跨被试泛化能力的评估协议。本研究在一个严格的被试独立性评估框架下,对一个多分类想象语音脑电数据集进行了透明的基线研究。研究比较了两种预处理与特征提取流程:时域统计特征方法和频域谱频带功率方法,并采用逐被试交叉验证、试验级多数投票以及随机森林分类器进行评估。在跨被试的粗粒度分类任务中,频域流程获得了显著高于统计特征流程的平均试验级准确率(49.03 ± 4.18% 对比 37.97 ± 3.79%)。前向特征选择进一步表明,有限的频带子集即可捕获大部分判别信息。总体而言,这项工作为未来旨在提升基于脑电的想象语音解码跨被试泛化能力的脑机接口研究奠定了坚实基础。
cs.AI / 81 / 2609.29837

PUBG Ally: A Conversational Embodied Agent as an AI Teammate

PUBG Ally:作为AI队友的对话式具身智能体
Kim, Beomsoo, Kim, Byeongju, Kim, Dohyun, Kim, Dongwon, Kim, Eunchong, Kim, Hongmin, Im, Hyeojung, Hwang, Hyeonbin, Kim, Hyeonghwan, Seol, Hyoseok, Im, Insub, Chen, Irene, Jeon, Jaeseung, Hong, Jimin, Yoo, Kiyoon, Park, Minkyoung, Jung, Seohyeon, Chung, Seungjun, Park, Sue Hyun, Kim, Sungwoo, Cho, Youngin, Son, Yujeong, Lee, Kangwook, Kim, Hyunseung
Abstract
We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to a constantly changing game world under strict latency constraints while interacting naturally with players, keeping its speech synchronized with its actions. Ally therefore combines agentic tool use with real-time game control. A language-model agent uses a controlled interface to inspect game information, interpret player speech, maintain context, decide what to say, and issue high-level action choices that steer a faster control layer for movement, combat, and recovery. Because the player's and Ally's speech and actions continually shape each other and the course of the match, training requires data from actual gameplay. We therefore collect data across nearly 39k sessions in which real players play alongside Ally, recording gameplay, player speech, agent decisions, tool use, actions, and player feedback, and use these records for iterative training. To evaluate teammate quality, we use player feedback and preference comparisons to identify gaps between offline evaluations and player preferences, and iteratively refine the evaluation criteria. Deploying Ally in live service further requires low-latency on-device execution and safeguards for player-facing communication, which we address through model compression, context compaction, targeted safety training, runtime guardrails, and memory redaction. During the live service, we surveyed players in 141 countries. Among respondents whose play with Ally was confirmed in game records, positive responses exceeded negative responses by 25.1 percentage points when asked whether they would recommend Ally, with players describing Ally not only as a tool but also as a teammate or companion.
Chinese Translation
我们介绍了PUBG Ally,一个面向《绝地求生》(PUBG: BATTLEGROUNDS)的具身智能体,它能够进行推理、自主行动,并作为一个支持语音交互的队友与玩家并肩游戏。构建这样的AI队友需要结合两种高难度能力:一方面,它必须在严格的延迟约束下感知并响应不断变化的游戏世界;另一方面,它必须与玩家自然交互,并使其语音与行动保持同步。为此,Ally将智能体工具调用(agentic tool use)与实时游戏控制相结合。一个基于语言模型的智能体通过受控接口来查看游戏信息、理解玩家语音、维护上下文、决定说什么,并发出高层动作选择,从而驱动一个更快的控制层完成移动、战斗和治疗恢复等操作。由于玩家与Ally的语音和行为不断相互影响并改变比赛进程,训练需要来自真实对局的数据。因此,我们在近3.9万场真实玩家与Ally并肩游戏的会话中收集数据,记录游戏过程、玩家语音、智能体决策、工具调用、动作以及玩家反馈,并利用这些记录进行迭代训练。为评估队友质量,我们使用玩家反馈和偏好对比来识别离线评估与玩家偏好之间的差距,并迭代完善评估标准。将Ally部署到线上服务还进一步要求低延迟的端侧执行以及对玩家可见通信的安全保障,我们通过模型压缩、上下文压缩、针对性安全训练、运行时防护栏(guardrails)和记忆脱敏来解决这些问题。在线上服务期间,我们调查了141个国家的玩家。在游戏记录中确认与Ally共同游戏过的受访者中,当被问及是否愿意推荐Ally时,正面回应比负面回应高出25.1个百分点,玩家不仅将Ally描述为一个工具,更将其视为队友或伙伴。
cs.AI / 82 / 2609.29874

A Risk-Adaptive and Evidence-Constrained Framework for Generative AI Feedback in Programming Education

面向编程教育中生成式AI反馈的风险自适应与证据约束框架
Wang, Shihao
Abstract
Generative artificial intelligence can turn learning analytics into personalized support, but feedback systems must decide when to intervene, which evidence to use, and how much assistance to provide. We developed a risk-adaptive, evidence-constrained framework for introductory programming using 2993 failed-submission states from 215 students. Student-disjoint models predicted persistent failure and related outcomes; four matched feedback conditions were generated for 136 cases; and calibrated risk informed capacity-limited intervention policies. The validation-selected logistic regression model achieved a test precision-recall area under the curve of 0.550 and a receiver operating characteristic area under the curve of 0.681. Broader student histories improved prediction of unmodified resubmission. After standardized repair and evidence gating, 519 of 544 newly generated messages contained all required components. A fixed-threshold sequential policy selected 17.8% of eligible test states and captured 25.2% of observed persistent failures. These findings support an evidence-gated progressive assistance strategy: calibrated risk guides intervention timing, recorded evidence constrains feedback content, and assistance progresses from self-checks to localized hints when warranted. The framework connects prediction, decision-making, and grounded generation while keeping their evaluation outcomes distinct.
Chinese Translation
生成式人工智能可以将学习分析转化为个性化支持,但反馈系统必须决定何时干预、使用哪些证据以及提供多少帮助。我们基于215名学生的2993个提交失败状态,为编程入门课程开发了一个风险自适应、证据约束的框架。采用学生分离的模型预测持续性失败及相关结果;为136个案例生成了四种匹配的反馈条件;并利用校准后的风险信息构建了受容量限制的干预策略。经验证选择的逻辑回归模型在测试集上达到了0.550的精确率-召回率曲线下面积(PR-AUC)和0.681的受试者操作特征曲线下面积(ROC-AUC)。更丰富的学生历史数据改善了对未经修改直接重新提交行为的预测。经过标准化修复和证据筛选后,544条新生成的消息中有519条包含所有必需的组成部分。基于固定阈值的顺序策略选择了17.8%的合格测试状态,并捕获了25.2%的已观测持续性失败。这些发现支持一种证据筛选下的渐进式辅助策略:校准的风险指导干预时机,记录的证据约束反馈内容,且在必要时辅助从自我检查逐步推进到局部提示。该框架将预测、决策与有据生成相连接,同时保持三者评估结果的独立性。
cs.AI / 83 / 2609.29875

When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression

智能体何时可以遗忘其推理过程?面向长程智能体上下文压缩的ICLR方法
Wang, Mingxuan, Luo, Fei, Wang, Bo, Yao, Guorun, Guo, Yinglong, Ning, Chao, Chen, Hongyue, Ma, Yanbiao, Han, Jungong
Abstract
Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajectory. We study when such reasoning can be safely forgotten. We propose Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training free online method that ranks reasoning blocks using frozen proxy entropy while preserving actions, tool calls, and observations. On 260 WorkBuddyBench tasks, ICLR improves average reward from 0.699 to 0.718, while reducing input, output, and cache read tokens by 25.5%, 14.4%, and 33.3%, respectively. Ablations reveal trajectory amplification, where local reasoning deletion produces nonlinear changes in total computation by altering subsequent interaction. Representation probing, activation patching, and controlled trajectory analyses further suggest that historical reasoning becomes more replaceable once task relevant derived state has been reliably externalized into code, files, tool outputs, or environmental feedback. These results characterize agent reasoning as dynamic working state rather than permanent interaction history.
Chinese Translation
长程语言模型智能体会不断积累推理历史,即使早期决策已被执行并观察到,上下文长度和推理成本仍持续增加。与静态的思维链(Chain of Thought)压缩不同,删除历史推理可能会改变未来的动作及由此产生的交互轨迹。我们研究了何时可以安全地遗忘此类推理。我们提出了面向长程推理的交互感知压缩方法(Interaction Aware Compression for Long Horizon Reasoning, ICLR),这是一种无需训练的在线方法,利用冻结的代理熵对推理块进行排序,同时保留动作、工具调用和观察结果。在260个WorkBuddyBench任务上,ICLR将平均奖励从0.699提升至0.718,同时分别将输入、输出和缓存读取token减少25.5%、14.4%和33.3%。消融实验揭示了轨迹放大效应,即局部推理删除会通过改变后续交互而导致总计算量的非线性变化。表征探测(representation probing)、激活修补(activation patching)和受控轨迹分析进一步表明,一旦任务相关的派生状态被可靠地外部化到代码、文件、工具输出或环境反馈中,历史推理就变得更容易被替换。这些结果表明,智能体推理是一种动态工作状态,而非永久的交互历史。
cs.AI / 84 / 2609.29876

Ontology-Mediated Neurosymbolic Constraint Acquisition from Multiple Stakeholders

基于本体中介的多利益相关方神经符号约束获取
Bischof, Stefan, Kainz, Juliana, Valerio, Danilo
Abstract
Neurosymbolic research typically assumes a pre-existing symbolic specification, leaving the upstream challenge of acquiring and formalizing requirements and constraints largely unaddressed. We present an architecture that fills this gap by using an OWL configuration ontology to mediate between neural constraint sources and downstream consumers. In this framework, LLM assistants elicit soft stakeholder preferences, while hardware specifications define hard physical and engineering limits. The ontology unifies these heterogeneous inputs, leverages description logic to identify unsatisfiability, and generates symbolic explanations that enable LLMs to interactively renegotiate terms with users. Any remaining conflicts are resolved downstream via priority-based relaxation. We illustrate our approach on a microgrid use case from the FLEXI project and argue its generalizability to multi-stakeholder domains where constraint acquisition is distributed across human and automated sources of unequal authority.
Chinese Translation
神经符号研究通常假设符号规范已预先存在,而对需求与约束进行获取和形式化的上游挑战在很大程度上未被解决。我们提出了一种填补这一空白的架构,利用 OWL 配置本体在神经约束源与下游消费者之间进行中介。在该框架中,LLM 助手负责获取利益相关方的软性偏好,而硬件规格则定义了硬性的物理与工程限制。该本体统一了这些异构输入,利用描述逻辑识别不可满足性,并生成符号化解释,使 LLM 能够与用户交互式地重新协商约束条款。剩余的冲突则通过基于优先级的松弛机制在下游解决。我们以 FLEXI 项目中的微电网用例为例说明了该方法,并论证了其向多利益相关方领域的可推广性,在这类领域中,约束获取分布于权威程度不等的自动化与人工来源之间。
cs.AI / 85 / 2609.29892

Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

Qwen-Planner-Agent:面向现实世界移动规划智能体的闭环AI-for-AI框架
Qu, Tingyu, Sun, Weigao, Liu, Yuecheng, Zhao, Yucheng, Zhu, Yi, Ding, Yifeng, Wang, Qiyi, Cao, Sihan, Jiao, Pengkun, Xie, Hanlei, Wu, Xiongwei, Wang, Qichao, Zhang, Haodong, Liu, Jiajun, Wang, Yuhao, Xie, Yuqing, Zhao, Junpeng, Chen, Long, Ma, Ming, Yang, Sihan, Zhao, Ziwang, Jia, Yanhao, Gong, Liangquan, Zhu, Feida, Zhong, Yiran, Hoi, Steven
Abstract
The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model--harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities.
Chinese Translation
大语言模型的快速发展正在将人工智能从被动的内容生成扩展到工程与科学发现的主动工作流中。这一转变引发了一个引人深思的问题:AI能否既是开发的对象,又是构建下一代AI系统的积极参与者?我们在一个可扩展开发与迭代改进的闭环AI-for-AI框架内构建了Qwen-Planner-Agent,以探索这一问题。移动规划为这一方法提供了极具挑战性的测试:复杂的长程任务考验智能体的可靠性,而昂贵的真机交互限制了开发的可扩展性。该框架通过共享的动作-反馈-验证契约连接数据生产、模型训练与部署。(i) AI for Data(AI用于数据)构建了一个由人类把关的智能体数据飞轮,其中专门的智能体负责构建任务、收集交互轨迹、整理并平衡训练数据,并利用训练反馈指导后续的数据生成。(ii) AI for Training(AI用于训练)将监督式规划冷启动与混合环境下的在线智能体强化学习相结合,我们提出了能力感知奖励与优势工程(Competence-Aware Reward-and-Advantage Engineering, CARE),在保持任务性能的同时降低推理与工具使用成本。(iii) AI通过执行证据驱动的循环推动模型与执行框架的协同演化,在运行时统筹记忆、技能与工具,并将结构化的动作反馈和保留的失败轨迹回传,用于模型与执行框架的协调适配。在MobilePA-Bench上,Qwen-Planner-Agent在所有被评估的模型与系统中取得了最佳的整体性能,在工具使用、记忆、技能和子智能体协调方面均优于其基座模型。进一步评估表明,我们的模型在非移动端智能体基准测试中也有所提升,同时大体上保留了通用能力。
cs.AI / 86 / 2609.29921

Who Holds the Pen? Let Specifications, Not Agents, Sign Off

谁执笔?由规范而非智能体来最终裁定
Li, Haiqing, Ma, Xin, Wu, Yinhao, Zhong, Wenliang, Jiang, Feng, Dang, Thao M., Hu, Xiao, Ma, Hehuan, Guo, Yuzhi, Huang, Junzhou
Abstract
Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the same model that acts and declares completion, leaving no independent specification authority boundary. We identify two resulting gaps. The understanding--execution gap arises when a requirement is understood but not satisfied in execution; the state--authority gap arises when an agent's interpretation or completion claim does not establish the required state. On SkillsBench, using only agent-visible prompts, workspace information, and injected skill specifications, we extract 509 source-grounded task directions. Across seven models, only 79.6%--86.4% are satisfied, while completion-claim rates exceed official evaluator pass rates by 28.7--37.9 percentage points. We therefore separate agent proposals from authoritative state. Agents may plan, act, and request completion, but only admissible evidence from qualified providers may establish specification-governed state. SpecHarness operationalizes this principle by compiling visible specifications into source-linked obligations and governing execution and finalization through versioned obligation state. Verifiable requirements are mediated or validated at runtime, while ambiguous or subjective requirements remain advisory. Experiments on guideline-following and artifact-generation tasks show that specifications can serve not merely as behavioral guidance, but as authority over compliant execution and completion.
Chinese Translation
大语言模型智能体(LLM agents)日益将生成、决策、执行与自我评估整合于单一智能体循环之中。尽管它们在任务指令、指南、输出模式(output schemas)和可复用技能等外部规范下运行,但这些规范通常仅作为同一模型行动并宣告完成时的上下文,缺乏独立的规范权限边界。我们由此识别出两类缺口:理解—执行缺口,即需求被理解但在执行中未得到满足;状态—权限缺口,即智能体的解读或完成声明未能确立所需状态。在 SkillsBench 上,仅使用智能体可见的提示词、工作区信息和注入的技能规范,我们提取了 509 条有源可依的任务方向。在七个模型上,仅 79.6%–86.4% 的任务方向得到满足,而完成声明率超出官方评估器通过率 28.7–37.9 个百分点。因此,我们将智能体的提议与权威状态分离:智能体可以规划、行动并请求完成,但只有来自合格提供方的可采信证据才能确立受规范约束的状态。SpecHarness 将这一原则付诸实践,把可见规范编译为与来源关联的义务(obligations),并通过带版本的义务状态来治理执行与终结过程。可验证的需求在运行时被调解或验证,而模糊或主观的需求仅作为建议。在遵循指南和制品生成任务上的实验表明,规范不仅可以作为行为引导,更可以作为对合规执行与完成的权威约束。
cs.AI / 87 / 2609.29947

Neuro-symbolic AI for Industrial Configuration

面向工业配置的神经符号人工智能
Valerio, Danilo, Kogler, Philipp, Bischof, Stefan, Hubauer, Thomas, Rangwala, Huzefa
Abstract
Large Language Models (LLMs) have shown impressive performance on a wide range of generative tasks. Yet their probabilistic nature makes them, in isolation, fundamentally unsuited for industrial product configuration, where outputs must be syntactically valid, semantically consistent with a knowledge base of hundreds of features and rules, and producible by an existing manufacturing chain. We argue that Neuro-symbolic (NeSy) AI methods lay out a promising path towards industrial-grade configurators that are reliable by design, explainable, and trustworthy. This paper describes a taxonomy of three NeSy integration strategies, namely hybrid inference, hybrid fine-tuning, and hybrid training, exploring their usage in the configuration domain. We report our effort to operationalize NeSy concepts in an industrial configuration copilot and derive a set of practical design choices for deploying trustworthy AI in engineering environments. We close with a discussion of open research challenges we consider most pressing, in particular how to scale NeSy methods from small academic demonstrators to the size of industrial configurators.
Chinese Translation
大语言模型(LLMs)在各类生成式任务上展现出令人瞩目的性能。然而,其概率性本质使其单独使用时从根本上不适合工业产品配置任务——在该任务中,输出必须在语法上有效、在语义上与包含数百个特性和规则的知识库保持一致,并且能够由现有制造链生产。我们认为,神经符号(Neuro-symbolic, NeSy)人工智能方法为实现可靠、可解释且值得信赖的工业级配置器提供了一条有前景的路径。本文描述了三种神经符号集成策略的分类体系,即混合推理、混合微调和混合训练,并探讨了它们在配置领域的应用。我们报告了在工业配置辅助工具中落地神经符号概念所做的努力,并总结出一组在工程环境中部署可信赖人工智能的实用设计选择。最后,我们讨论了我们认为最紧迫的开放性研究挑战,特别是如何将神经符号方法从小型学术演示系统扩展到工业配置器的规模。
cs.AI / 88 / 2609.29948

ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation

ENDOPROMPT:面向效用退化的受害方伪参考方法
Wu, Qingyu, Feng, Zeyu, Yu, Yongda, Luo, Yuzhe, Cheng, Hua
Abstract
Prompt injection can degrade benign task performance without eliciting harmful content. Yet many attack objectives depend on task labels or predefined target responses. We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions. Its generator takes the request text as input. Clean victim continuations serve as pseudo-references: local search identifies prefixes that reduce continuation likelihood, and preference fitting on comparisons within the same instruction, followed by reward refinement, distills this signal into a generator. At deployment, the generator produces one prefix per request without further victim-side search. Across four instruction-tuned models and the complete splits of seven benign benchmarks, ENDOPROMPT yields a mean utility change of -26.8 percentage points; 27 of 28 cells are negative. Failure analysis reveals output expansion and prefix reuse; the controls do not establish a degradation advantage from request matching. Victim-derived supervision can reveal utility weaknesses without benchmark feedback or prescribed failure responses. The code will be released upon acceptance.
Chinese Translation
提示注入攻击可以在不诱发有害内容的情况下降低良性任务的性能。然而,许多攻击目标依赖于任务标签或预定义的目标响应。我们提出了ENDOPROMPT,这是一种白盒方法,能够从无标注的指令中学习降低效用的前缀。其生成器以请求文本作为输入。受害模型在干净输入下的续写结果被用作伪参考:局部搜索识别出能够降低续写似然的前缀,然后在同一指令内的比较上进行偏好拟合,并通过奖励精炼,将该信号蒸馏到生成器中。在部署时,生成器为每个请求生成一个前缀,无需进一步的受害方搜索。在四个指令微调模型和七个良性基准的完整划分上,ENDOPROMPT带来平均-26.8个百分点的效用变化;28个实验单元中有27个为负。失败分析揭示了输出扩张和前缀复用现象;对照实验并未证明请求匹配带来退化优势。基于受害模型的监督信号可以在没有基准反馈或预设失败响应的情况下揭示效用弱点。代码将在论文被接收后发布。
cs.AI / 89 / 2609.29952

Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes

Augur:用于演练产品与政策变革应对反应的合成决策实验室
Khedar, Rahul, Malhotra, Mayank, Karn, Avinash
Abstract
Before a product or policy change ships, the question that matters is how people will react to it. Augur rehearses that reaction offline: it builds a typed knowledge graph from the change documents, populates a grounded persona market, simulates the interaction, and returns an auditable decision memo recommending one of five actions. We assemble Gold-50, fifty real product and policy episodes whose real-world outcome is known, adjudicated against the public record, and score the five-way release verdict against it. Our central finding is methodological and negative: most of the measured gap between frontier cloud models and open-weight models we fine-tune and serve offline is attributable to an under-specified evaluation, not a difference in capability. We show this three ways. First, the prompt envelope alone can dominate the score: holding weights, cases and scorer fixed, one system -- a LoRA-SFT adapter on Qwen3-32B -- swings from 0% to 73%. Second, in a matched 2x2 ablation, defining the decision taxonomy in the prompt -- with no model change -- lifts every frontier model by +24 to +34pp; under the under-specified prompt, Qwen3-32B LoRA-SFT served offline beats all three frontier models (paired McNemar, Holm-corrected), and once the prompt is fair no significant difference from any of them is detected. Third, agreement with the distillation teacher rises without accuracy following, and the full pipeline amplifies a systematic "over-doom" bias rather than improving the verdict. Separately, we validate the reaction layer on its own terms: blind judges across four model families find the synthetic reaction recovers 67-90% of the concerns the public actually raised, and a pre-registered ablation locates its value -- largest where the decision is hardest, redundant near ceiling. The pipeline that regenerates every number and figure here is available from the authors.
Chinese Translation
在产品或政策变革落地之前,关键的问题是人们将如何对其作出反应。Augur(Augur)在离线状态下演练这一反应:它从变革文档中构建带类型的知识图谱,填充一个有事实依据的人物角色市场,模拟交互过程,并返回一份可审计的决策备忘录,推荐五种行动之一。我们构建了Gold-50数据集,包含五十个真实世界结果已知的产品与政策案例事件,并依据公开记录进行裁定,随后将五选一的发布判定结果与之对比评分。我们的核心发现是方法层面的且带有负面性质:前沿云端模型与我们微调并离线部署的开源权重模型之间所测得的大部分差距,可归因于评估标准的不充分界定,而非能力上的差异。我们通过三种方式证明这一点。第一,仅提示词框架即可主导得分:在模型权重、案例与评分器均固定的条件下,一个系统——基于Qwen3-32B的LoRA-SFT适配器——得分可在0%到73%之间波动。第二,在一组匹配的2x2消融实验中,仅在提示词中定义决策分类体系——不改变任何模型——即可使每个前沿模型提升24至34个百分点;在不充分界定的提示词下,离线部署的Qwen3-32B LoRA-SFT胜过全部三个前沿模型(配对McNemar检验,经Holm校正),而一旦提示词公平,则未检测到与任何一个前沿模型存在显著差异。第三,与蒸馏教师模型的一致性上升而准确率并未随之提升,且整个流水线放大了一种系统性的"过度悲观"偏差,而非改善判定结果。另外,我们还独立验证了反应模拟层的价值:来自四个模型家族的盲评裁判发现,合成反应能够还原公众实际提出的67-90%的关切;一项预注册的消融实验定位了其价值所在——在决策最困难的情形中价值最大,在接近上限时则趋于冗余。用于重新生成本文全部数值与图表的流水线可向作者获取。
cs.AI / 90 / 2609.30001

Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems

AgentX中的模型研究进展:面向工业推荐系统的长时程自主性
Yang, Shuang, Zhuang, Zijie, Lao, Changxin, Xu, Pengbo, Xu, Hanwen, Huang, Yusheng, Gao, Han, Wang, Guanchen, Ma, Tianbao, Chen, Linxun, Song, Peilin, Wang, Xuming, Li, Chen, Wu, Fan, Wang, Tao, Zhao, Zibo, Wu, Xiangyu, Liu, An, Pan, Fei, Jiang, Peng, Yang, Chen, Liu, Zhaojie, Ou, Wenwu
Abstract
Sustaining industrial recommendation research requires using the results of one experiment to decide what to investigate next. We present AgentX-Model, the next generation of AgentX's model research framework, which connects proposal development and model experimentation within sandboxes defined by business inputs and prediction tasks. AgentX-Model adopts a dual-agent architecture comprising a Research Agent and a Model Agent. The Research Agent develops independently reviewed proposals from papers and experimental findings, while the Model Agent conducts multi-round investigations and returns code, measurements, and unresolved questions. Using the returned results, the Research Agent selects a starting implementation and formulates the next research question, allowing subsequent experiments to build on earlier findings. We organize this continuing research around four actions: Reproduce, Follow-up, Composition, and Diagnose. The first three actions drive routine research, while Diagnose acquires the evidence needed to choose a repair, including for issues raised by business feedback and online evaluation, such as prediction bias measured by PCOC. Across the production evaluation, 560 of 636 completed model-changing experiments recorded AUC above their business baselines. As research continued, some experiments recorded AUC above every comparable ancestor in their lineages. The five latest online A/B evaluations across different business settings reported gains including 10-15% in acquisition efficiency, 15-20% in target-segment advertising spend, and 0.3-0.8% in watch time; the watch-time model used approximately 10% fewer FLOPs and parameters. A dependency-aware historical-replay benchmark further evaluates research allocation, with initial results showing no consistent efficiency gain from more complex scheduling when agents already analyze and select concrete candidates.
Chinese Translation
持续的工业推荐研究需要利用一个实验的结果来决定下一步研究什么。我们提出AgentX-Model,即AgentX下一代模型研究框架,它在由业务输入和预测任务定义的沙箱中将提案开发与模型实验连接起来。AgentX-Model采用由研究智能体(Research Agent)和模型智能体(Model Agent)组成的双智能体架构。研究智能体基于论文和实验发现开发经过独立评审的提案,而模型智能体执行多轮研究并返回代码、测量结果和未解决的问题。利用返回的结果,研究智能体选择一个起始实现并制定下一个研究问题,使后续实验能够建立在已有发现之上。我们将这一持续性研究组织为四种动作:复现(Reproduce)、跟进(Follow-up)、组合(Composition)和诊断(Diagnose)。前三种动作驱动日常研究,而诊断则获取选择修复方案所需的证据,包括针对业务反馈和在线评估提出的问题,例如以PCOC衡量的预测偏差。在生产环境评估中,636个已完成的模型变更实验中有560个的AUC高于其业务基线。随着研究的持续,一些实验的AUC超过了其研究谱系中所有可比较的祖先。在最新五次跨不同业务场景的在线A/B评估中,取得了包括获客效率提升10-15%、目标人群广告投放支出提升15-20%、观看时长提升0.3-0.8%的收益;其中观看时长模型使用的FLOPs和参数量减少了约10%。一个依赖感知的历史回放基准进一步评估研究资源的分配,初步结果表明,当智能体已经能够分析并选择具体候选方案时,更复杂的调度并不能带来一致的效率提升。
cs.AI / 91 / 2609.30027

Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark

合成医院:一个开放、可验证、经医生验证的纵向电子健康记录基准
Park, Christine, Chen, Valerie, Dettmers, Tim
Abstract
Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53\%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient's longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance.
Chinese Translation
前沿语言模型很少被应用于临床工作流程,原因在于开发此类模型所需的真实的纵向基准数据十分稀缺。真实的电子健康记录(EHR)数据因隐私、伦理或数据使用等问题无法公开共享,而且由于病历只反映临床医生所记录的内容,其中也不包含可验证的真值(ground truth)。我们提出了Synthetic Hospital,这是一个开放的、完全合成的、基于事实的纵向EHR基准,解决了公开共享和可验证真值这两大障碍。该基准完全基于公开的医学教育资料构建,不含任何受保护的健康信息,包含1,268名纵向患者和5,602次就诊记录,其中每一项诊断、发现和时间关系均锚定于标准本体(ICD-10-CM、SNOMED CT、LOINC),并具备可追溯至其来源医学教育资料的完整溯源链。Synthetic Hospital通过一个模拟医院记录系统提供服务,该系统镜像了真实的EHR基础设施(标准互操作性API、基于角色的访问控制以及函数调用接口)。在一项盲法评审中,医生区分其记录与真实患者病历的准确率接近随机水平(53%)。在对10个前沿模型和开源模型的评估中,没有任何模型接近上限:表现最好的模型在重建患者纵向问题列表时的严重程度加权F1分数为0.73,与匹配子集上七位医生的平均水平相当,但远低于其中最佳医生的水平(0.89),并且在总结病历时约有一半的临床相关发现被遗漏。总体而言,这些结果表明Synthetic Hospital是对临床AI性能的一项困难而真实的测试。
cs.AI / 92 / 2609.30028

How does Adversarial Influence Scale in Multi-Agent Systems?

对抗性影响在多智能体系统中如何随规模扩展?
Wu, Addison J., Cekinmez, Jasin, Liao, Michel, Narasimhan, Karthik, Griffiths, Thomas L.
Abstract
Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially correct agents switch to an incorrect final answer, rises linearly with this proportion. Whereas humans in comparable conformity studies are reliably swayed only when misleading confederates form a majority, LLM agents defect regularly even when deceivers remain a minority. Susceptibility also depends on which models are interacting, especially on the honest agent side. Unexpectedly, allowing deceivers to coordinate privately can make them less effective. Altogether, our results show that adding more agents is therefore not a sufficient defense, because the adversary can simply scale with the group.
Chinese Translation
多智能体协商可以提升性能,但当某些智能体并非善意行事时会发生什么?在实践中,智能体可能具有欺骗性并试图破坏整个群体,无论是出于自身目标还是外部指令。我们研究了随着群体规模扩大和欺骗者比例上升,对欺骗的易感性如何变化。关键因素不是群体中智能体的数量,而是欺骗者所占的比例。我们观察到背叛率(即最初答案正确的智能体转向错误最终答案的频率)随该比例呈线性上升。在可比的从众实验中,人类只有在误导性同谋者构成多数时才会被可靠地影响,而LLM智能体即使欺骗者仅占少数也会经常发生背叛。易感性还取决于哪些模型在交互,尤其是诚实智能体一侧所使用的模型。出乎意料的是,允许欺骗者私下协调反而可能降低其有效性。总之,我们的结果表明,仅仅增加智能体数量并不是充分的防御手段,因为攻击者可以随着群体的扩大而同步扩展。
cs.AI / 93 / 2609.30048

Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models

是风格而非自我:表面线索解释了大型语言模型的零样本代码归因
Barkhordar, Ehsan, Thapa, Surendrabikram
Abstract
If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude. We test this zero-shot on current commercial models. Five LLMs generate solutions to MBPP, HumanEval, and DS-1000, seven more to MBPP, and models act as evaluators in four tasks: picking their own solution from a pair, judging whether a single solution is their own, identifying which of two solutions a named model wrote, and judging quality blind. In the single-solution task, balanced accuracy is 49-58% for all 15 model-benchmark combinations, while raw accuracy (38-67%) mostly reflects how readily a model claims authorship. In the pairwise task, accuracy across 14 evaluator-opponent combinations correlates at r=0.93 with how often the evaluator's solution is longer. Attribution to a named model succeeds on some pairs and is consistently inverted on others. A rule-based normalization that strips docstrings, comments, type hints, and local names preserves Pass@1 and leaves ten of twelve re-tested results at chance; the other two follow a length difference it leaves, although a trained classifier still separates most normalized pairs. Claude Haiku's self-preference also disappears. We recommend reporting balanced accuracy, heuristic baselines, and label consistency.
Chinese Translation
如果语言模型能够识别自己编写的代码,它作为评判者时可能会偏袒这些代码,而模型之间相互监控的实例也可能导致合谋。我们在当前商业模型上对这一零样本能力进行了测试。五个大语言模型(LLM)为 MBPP、HumanEval 和 DS-1000 生成解题方案,另外七个模型为 MBPP 生成方案,模型在四项任务中充当评估者:从一对方案中挑选自己的方案、判断单个方案是否为自己所写、识别某个指定模型编写了两个方案中的哪一个,以及在盲测下评判代码质量。在单方案任务中,全部 15 个“模型-基准”组合的平衡准确率为 49–58%,而原始准确率(38–67%)主要反映了模型声称代码所有权的难易程度。在成对任务中,14 个“评估者-对手”组合的准确率与评估者自身方案更长的情况出现频率的相关系数达 r=0.93。对指定模型的归因在某些配对上成功,而在另一些配对上则始终反转。一种去除文档字符串、注释、类型提示和局部变量名的基于规则的标准化方法保持了 Pass@1 不变,并使 12 个复测结果中的十个降至随机水平;另外两个结果则遵循该方法未消除的长度差异,尽管训练过的分类器仍能区分大多数标准化后的配对。Claude Haiku 的自我偏好也消失了。我们建议报告平衡准确率、启发式基线以及标签一致性。
cs.AI / 94 / 2609.30050

NNV3: Expanding Neural Network Verification to New Architectures and Domains

NNV3:将神经网络验证扩展到新的架构与领域
Tumlin, Anne M., Sasaki, Samuel, Wooding, Ben, Lopez, Diego Manzanas, Zubair, Muhammad Usama, Hashemi, Navid, Zhang, Hongchao, Abbas, Waseem, Oguz, Ipek, Ma, Meiyi, Johnson, Taylor T.
Abstract
We present NNV3, the latest version of the Neural Network Verification (NNV) tool, a MATLAB framework for formal verification of deep learning models and learning-enabled cyber-physical systems. Building on the set-based reachability foundation of NNV 1.0 (FFNNs, CNNs, NNCS) and NNV 2.0 (RNNs, SSNNs, neural ODEs), NNV3 introduces new members of the Star-set family: ModelStar for verifying networks under weight perturbation, VolumeStar for video and 3D volumetric inputs, and GraphStar for graph neural networks. A conformal-inference-based probabilistic reachability mode complements sound analysis for problems where deterministic verification is intractable, while FairNNV certifies counterfactual and individual fairness properties over continuous input regions. NNV3 introduces new benchmarks for malware detection, graph-based power-system models, medical imaging, variable-length time series data, and action recognition. NNV3 also incorporates tutorials and developer guides through a unified documentation site. This paper details these major updates, demonstrating NNV's maturation into a comprehensive, robust, and accessible verification tool for a diverse range of AI systems.
Chinese Translation
我们提出了 NNV3,即神经网络验证工具(NNV)的最新版本,这是一个用于深度学习模型及学习使能信息物理系统(cyber-physical systems)形式化验证的 MATLAB 框架。在 NNV 1.0(前馈神经网络、卷积神经网络、神经网络控制系统)和 NNV 2.0(循环神经网络、状态空间神经网络、神经常微分方程)基于集合的可达性分析基础之上,NNV3 引入了 Star 集家族的新成员:用于验证权重扰动下网络的 ModelStar、面向视频和三维体数据输入的 VolumeStar,以及面向图神经网络的 GraphStar。对于确定性验证难以处理的问题,基于保形推断(conformal inference)的概率可达性模式为可靠分析提供了补充;同时,FairNNV 可在连续输入区域上认证反事实公平性和个体公平性属性。NNV3 还引入了针对恶意软件检测、基于图的电力系统模型、医学影像、变长时间序列数据以及动作识别的新基准测试。此外,NNV3 通过统一的文档网站提供了教程和开发者指南。本文详细介绍了这些重大更新,展示了 NNV 已成长为一个面向多种 AI 系统的全面、稳健且易于使用的验证工具。
cs.AI / 95 / 2609.30054

SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback

SciWalker:基于算子图与执行反馈的科学编码问题合成
Li, Chenxi, Zeng, Wenxuan, Luo, Yun, Yu, Fangchen, Ye, Peng, Cheng, Yu, Zhang, Jun
Abstract
Improving the scientific coding capabilities of large language models (LLMs) requires high-quality training data. However, such data remain scarce because manually authoring realistic problems is costly and time-consuming, while systematically covering diverse scientific domains and algorithmic combinations remains challenging. To address this, we introduce SciWalker, a framework for synthesizing scientific coding problems through operator-chain sampling and execution feedback. The framework combines scientific library interfaces with operation modes to instantiate operators, organizes them into operator graphs, and samples operator chains as computational workflow cues. Guided by these cues, we adopt LLMs to generate scientifically grounded problem statements, reference solutions, and tests, with failed generations iteratively repaired using execution feedback. By combining structured workflow composition with verification and quality review, SciWalker enables scalable task generation while promoting scientific grounding, computational diversity, and executability. Using this framework, we construct 8,178 high-quality problems spanning 5 scientific domains and 32 subdomains. To evaluate their training utility, we conduct reinforcement learning on Qwen3.5-9B using the GSPO algorithm. This training improves SciCode subproblem accuracy by 9.9 percentage points, from 29.3% to 39.2%, with gains across scientific code generation, code repair, and reasoning benchmarks. The code for SciWalker is available at https://github.com/lichenx1/SciWalker.
Chinese Translation
提升大型语言模型(LLM)的科学编码能力需要高质量的训练数据。然而,此类数据仍然稀缺,因为人工编写真实的问题既昂贵又耗时,而系统地覆盖多样的科学领域和算法组合也颇具挑战性。为解决这一问题,我们提出了SciWalker,一个通过算子链采样与执行反馈来合成科学编码问题的框架。该框架将科学库接口与操作模式相结合以实例化算子,将其组织为算子图,并采样算子链作为计算工作流线索。在这些线索的引导下,我们采用LLM生成具有科学依据的问题描述、参考解法和测试用例,并利用执行反馈对失败的生成结果进行迭代修复。通过将结构化的工作流组合与验证及质量审查相结合,SciWalker能够实现可扩展的任务生成,同时促进科学依据性、计算多样性和可执行性。利用该框架,我们构建了8,178个涵盖5个科学领域和32个子领域的高质量问题。为评估其训练价值,我们使用GSPO算法对Qwen3.5-9B进行强化学习。该训练将SciCode子问题准确率提升了9.9个百分点,从29.3%提升至39.2%,并在科学代码生成、代码修复和推理基准上均取得提升。SciWalker的代码可在 https://github.com/lichenx1/SciWalker 获取。
cs.AI / 96 / 2609.30063

Self-Play Pretraining with Zero Data

零数据自博弈预训练
Cowsik, Aditya, Dolev, Kfir, Li, Michael Y., De Luca, G. Bruno, Cohen, Nourya, Goodman, Noah D., Levine, Yoav
Abstract
Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model's behalf. A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement. This would provide an effectively unbounded source of training data, limited by compute rather than human knowledge. We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision. Our procedure casts synthetic data generation as a search over the space of all computable structure, taking inspiration from Solomonoff induction. Starting from random initialization, two models learn in tandem: a generator proposes programs interpreted by a universal Turing machine, generating byte sequences, while a learner autoregressively predicts these byte sequences. The learner is trained with standard cross-entropy, while the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum. A universal Turing machine gives us a search space over all computable data-generating processes, imposing little domain-specific structure, and self-play searches over this space for useful training data. We test whether zero-shot performance on natural data improves predictably with self-play compute; this is a clean test of transfer since neither generator nor learner is trained on natural data. Across several natural datasets, zero-shot loss exhibits predictable scaling in compute. The models also exhibit in-context learning, and discover recognizable mathematical sequences during training.
Chinese Translation
语言建模的进展一直由在越来越多的数据上进行预训练的规模化所推动。然而,训练数据在很大程度上仍然是为模型预先筛选的。一种更通用的预训练方法应能让模型学会生成对自身改进最有用的数据。这将提供一种事实上无界的训练数据来源,其限制来自算力而非人类知识。我们提出了零数据自博弈预训练(Self-Play Pretraining with Zero Data),这是实现这一愿景的初步概念验证。我们的方法将合成数据生成转化为在所有可计算结构空间上的搜索,其灵感来自所罗门诺夫归纳(Solomonoff induction)。从随机初始化开始,两个模型协同学习:一个生成器提出由通用图灵机解释的程序,从而生成字节序列;一个学习器以自回归方式预测这些字节序列。学习器使用标准交叉熵进行训练,而生成器则通过强化学习进行训练,以产生处于学习器能力前沿的序列,从而形成自适应课程。通用图灵机为我们提供了覆盖所有可计算数据生成过程的搜索空间,几乎不引入领域特定的结构,而自博弈则在该空间中搜索有用的训练数据。我们检验了在自然数据上的零样本性能是否随自博弈算力的增加而可预测地提升;由于生成器和学习器均未在自然数据上训练,这是对迁移能力的一次干净的检验。在多个自然数据集上,零样本损失展现出随算力的可预测缩放。模型还表现出上下文学习能力,并在训练过程中发现了可识别的数学序列。
cs.AI / 97 / 2609.30094

PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations

PrivDrift:主动LLM对话中话题漂移下用户秘密泄露的审计
Maldonado, Luciano
Abstract
Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce \textbf{PrivDrift}, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing. PrivDrift contains 1{,}000 controlled multi-turn dialogues with seeded secrets, content-dense drift turns, and standardized extraction probes. Across three LLMs with extended context windows, dialogue-level hybrid leakage remains substantial, ranging from 38.7\% to 54.6\%, and varies strongly by model, secret type, and persuasion intensity. Within the tested drift window, additional topic drift does not reliably reduce leakage, suggesting that privacy risk in active LLM contexts should be evaluated as a persistent behavioral failure mode rather than only as training-data memorization or immediate jailbreak behavior.
Chinese Translation
大型语言模型日益作为持久化助手应用于面向用户的、共享会话的以及工具增强的场景中。当用户在活跃对话中披露敏感信息时,即使对话随后转向不相关的主题,这些信息仍可能通过后续提问在行为层面被恢复。我们提出了 PrivDrift,一个用于审计用户披露的秘密在对话话题漂移和基于说服的探测之后是否仍可被恢复的基准。PrivDrift 包含 1,000 个受控的多轮对话,其中包含植入的秘密、内容密集的漂移轮次以及标准化的提取探测。在三个具有长上下文窗口的大型语言模型上,对话层面的混合泄露仍然相当严重,介于 38.7% 到 54.6% 之间,并因模型、秘密类型和说服强度不同而显著变化。在测试的漂移窗口内,额外的话题漂移并不能可靠地降低泄露率,这表明主动 LLM 环境中的隐私风险应被视为一种持续存在的行为失效模式,而不仅仅是训练数据记忆或即时越狱行为。
cs.AI / 98 / 2609.30123

HEXIS: Compiling Skills into Extended Finite State Machines

HEXIS:将技能编译为扩展有限状态机
LI, Minghao
Abstract
Agent skills provide reusable knowledge and instructions, yet agents must repeatedly infer how to apply them and which operation should follow. This couples task reasoning with control decisions, allowing prescribed steps to be omitted or applied incorrectly. We introduce HEXIS, which compiles agent skills into extended finite state machines that separate knowledge from control flow. Skill knowledge is incorporated into local instructions that guide reasoning and generation within states. The machine records execution progress and intermediate results, while explicit transition conditions determine subsequent operations. Our incremental compiler first maps skill clauses and tool interfaces to state operations, local instructions, data bindings, and transitions. It then aligns development traces with existing states to identify missing operations and dependencies. These are incorporated by adding or reusing states and refining their connections. Updates are accepted only after static checks and replay of the current and all previously accepted traces. Across four benchmarks and four executors, HEXIS improves success over Skill + ReAct by 16.1 percentage points on average. Qwen3.8-27B reduces execution tokens by 38.4-88.9% across benchmarks.
Chinese Translation
智能体技能提供了可复用的知识和指令,但智能体必须反复推断如何应用这些技能以及下一步应执行什么操作。这将任务推理与控制决策耦合在一起,导致规定的步骤可能被遗漏或错误执行。我们提出了HEXIS,它将智能体技能编译为扩展有限状态机,从而将知识与控制流分离。技能知识被融入局部指令中,用于引导状态内部的推理和生成。状态机记录执行进度和中间结果,同时通过显式的转移条件确定后续操作。我们的增量编译器首先将技能条款和工具接口映射为状态操作、局部指令、数据绑定和转移。随后,它将开发轨迹与现有状态对齐,以识别缺失的操作和依赖关系,并通过添加或复用状态以及细化其连接将这些内容融入状态机。只有在通过静态检查并回放当前及所有先前已接受的轨迹后,更新才会被采纳。在四个基准测试和四种执行器上,HEXIS相比Skill + ReAct平均提升16.1个百分点的成功率。Qwen3.8-27B在各基准测试上将执行令牌数减少了38.4%至88.9%。
cs.AI / 99 / 2609.30137

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

先筛查后服务:面向1.4亿次规模生产级客户体验AI智能体的仿真方法
Alcoba, Edesio, Rossell, Kevin, Gupta, Aman, Tang, Shao, Hong, Jiwoo, Carrillo-Mendoza, Pabel, Ferreira, Wanderson Conceição, Tedeschi, Alvaro, Simjee, Zayd, Rajpal, Shreya, Hime, Bruno Finardi, Sousa, Christian, Moneda, Luis, Fei, Herbert, Silva, Daniel, Ramanath, Rohan
Abstract
Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank's Card Delivery agent and its expanded successor, Card Management - Nubank's highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.
Chinese Translation
客户体验(CX)智能体利用工具和大语言模型来处理客户请求,并引导客户与组织产品进行对话式交互。改进这些智能体十分困难,尤其是在受监管的行业:它们必须准确识别意图、遵循复杂的运营策略并可靠地使用工具。人工端到端测试的覆盖率有限,而线上实验则会将可能损害客户信任的失败暴露给真实客户。我们提出了一种基于假设驱动的仿真工作流,用于在部署前对候选CX智能体进行筛查。在该工作流中,合成客户对智能体的响应做出反应,模拟的工具输出使多步骤智能体工作流得以运行,而无需调用生产环境后端。我们在Nubank的Card Delivery智能体及其扩展后的后继版本Card Management上使用了Snowglobe仿真器——后者是Nubank在巴西对话量最高的聊天客服智能体。在4个已部署版本中,仿真环境与生产环境中的版本级二元评估器得分呈现出高度相关性。在仿真结果指导下的迭代使交易性净推荐值(tNPS)在线上A/B测试中提升了36.69分。我们还通过超过16,000次仿真对话对开源权重模型配置进行了筛查。在随后的线上A/B测试中,所选模型将自助服务率(SSR)提升了8.82个百分点,达到Nubank有史以来的最高水平,同时tNPS无统计显著变化。仿真使得对模型、推理设置和提示词的广泛探索成为可能,且无需让客户承担风险,从而实现了仅通过线上实验难以企及的生产改进。
cs.AI / 100 / 2609.30144

EnigmaForge: The Question Is Hidden in the Story

EnigmaForge:问题隐藏在故事之中
Eisner, Daniel
Abstract
Most benchmarks hand the model a question. EnigmaForge hands it a stack of old documents and no question at all. Buried in the letters, receipts, and logbook margins is a small logic puzzle whose solution is unique - proved by a SAT solver at generation time, with an ablation certificate showing every clue is load-bearing. Because instances are generated rather than collected, the corpus renews forever. The headline measure is intuition: task success when handed only the story, with world reconstruction as the secondary axis. Twenty-five frontier models ran over 600 instances (17,400 scored records) under three matched conditions. Intuition reshuffles the leaderboard: a 22x spread where fact recovery spans 1.6x, the second-best fact-recoverer ranks fourteenth, one model is indifferent to being told the question, and another is significantly better without it. Several models were blocked by their own content filters before reaching the puzzle - any benchmark scoring refusals as failure is quietly measuring filter behavior.
Chinese Translation
大多数基准测试直接将问题交给模型。而 EnigmaForge 交给模型的是一堆旧文档,且完全不提供任何问题。埋藏在信件、收据和日志边注中的是一个小型逻辑谜题,其解是唯一的——在生成时由 SAT 求解器证明,并附带消融证书以表明每条线索都不可或缺。由于数据实例是生成而非收集的,该语料库可以永久更新。核心评估指标是直觉:即模型在只获得故事的情况下完成任务的成功率,而世界状态重构则作为次要评估维度。二十五种前沿模型在三种匹配条件下运行了超过 600 个实例(共 17,400 条评分记录)。直觉测试重塑了排行榜:模型间差距达 22 倍,而事实恢复能力的差距仅为 1.6 倍;事实恢复能力第二强的模型在直觉维度仅排第十四位;一个模型在被直接告知问题时表现无差异,而另一个模型在没有被告知问题时反而显著更优。若干模型在接触谜题之前就被自身的内容过滤器拦截——任何将拒答计为失败的基准测试,实际上都在暗中衡量过滤器的行为。
cs.AI / 101 / 2609.30147

GRASP: Generating, Revising, and Assessing for Strategic Planning with Agentic AI

GRASP:基于智能体人工智能的战略规划生成、修订与评估
Srivastava, Arunabh, A., Mohammad, Khojastepour, Chakradhar, Srimat, Ulukus, Sennur
Abstract
Large Language Models (LLMs) typically exhibit a performance profile where reliability degrades as task complexity increases. We address the challenge of generating high-quality natural language executable plans for complex tasks by introducing $\textbf{GRASP}$, a strategy-aware, multi-stage planning framework. GRASP decouples the planning pipeline across specialized, context-isolated modules: it pre-compiles global macro-guidelines (GenPlan), explores alternative localized strategies within isolated context windows (RevPlan), and independently evaluates trajectories using a multi-criteria discriminator (VerPlan). Empirical evaluations show that GRASP consistently establishes a new state-of-the-art frontier across diverse datasets, yielding substantial accuracy gains over direct LLM planners on Natural Plan Calendar Scheduling ($\sim$12.4$\%$$\uparrow$), ZebraLogic ($\sim$30.8$\%$$\uparrow$), and SciBench Math. Crucially, under multi-task scaling-where standard planners suffer immediate performance collapse-GRASP completely flattens the multi-task degradation penalty. In interleaved dual-task environments, GRASP achieves an absolute accuracy gain of up to 16.7$\%$ over direct LLM planners. Furthermore, by isolating context and enforcing strict macro-regularization, GRASP outperforms frontier reasoning models (such as GPT-5-mini) by a margin of 14.5$\%$.
Chinese Translation
大语言模型(LLM)通常表现出随着任务复杂性增加而可靠性下降的性能特征。为解决为复杂任务生成高质量自然语言可执行规划的难题,我们提出了 extbf{GRASP},一种策略感知的多阶段规划框架。GRASP 将规划流程解耦为多个专门化、上下文隔离的模块:它预先编译全局宏观指导方针(GenPlan),在隔离的上下文窗口内探索备选的局部策略(RevPlan),并使用多准则判别器独立评估轨迹(VerPlan)。实证评估表明,GRASP 在多个数据集上持续确立新的最先进前沿,相较于直接使用 LLM 进行规划的方法,在 Natural Plan 日程调度上获得约 12.4% 的准确率提升,在 ZebraLogic 上提升约 30.8%,并在 SciBench 数学任务上也有显著提升。关键的是,在多任务扩展场景下——标准规划器会立即出现性能崩溃——GRASP 完全消除了多任务退化惩罚。在交错双任务环境中,GRASP 相较直接 LLM 规划器实现了最高 16.7% 的绝对准确率提升。此外,通过隔离上下文并施加严格的宏观正则化,GRASP 以 14.5% 的优势超越了前沿推理模型(如 GPT-5-mini)。
cs.AI / 102 / 2609.30177

Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search

面向Roblox游戏搜索中多组件查询理解搜索感知强化学习
Choi, Nayoung, Chen, Shengjian, Wei, Xiaokai, Zhang, Wenzheng, Yi, Daiyao, Pareek, Rachit, Su, Vincent, Gong, Michelle, Choi, Jinho D.
Abstract
Query understanding (QU) plays a critical role in production search systems, translating raw user queries into search execution plans that drive downstream retrieval and ranking. While large language models (LLMs) have enabled QU to be framed as a structured multi-task generation problem (e.g., intent classification, query expansion), optimizing such models to produce search-engine-coupled outputs remains challenging: static, label-based supervision fails to capture how each component actually interacts with the underlying search pipeline to affect downstream performance. We present a search-aware reinforcement learning (RL) framework for QU based on a distill-then-RL paradigm. Teacher-student supervised fine-tuning (SFT) first yields a well-formed, schema-compliant policy initialization. The RL stage then optimizes each QU component with rewards derived from live interaction with the search engine, tailored to that component's operational role, rather than a single reward tied to the final search outcome. Experiments on Roblox search show that this component-specific optimization improves both per-component utility and downstream search quality, raising NDCG@20 by 8.9 points over the SFT policy and by 3.5 points over training with a single end-to-end reward.
Chinese Translation
查询理解在工业级搜索系统中扮演着关键角色,它将原始用户查询转化为驱动下游检索与排序的搜索执行计划。尽管大语言模型已使查询理解能够被构建为一个结构化的多任务生成问题(如意图分类、查询扩展),但如何优化此类模型以生成与搜索引擎紧密耦合的输出仍具挑战性:基于静态标签的监督无法捕捉每个组件与底层搜索流水线之间的实际交互及其对下游性能的影响。我们提出了一种基于“先蒸馏后强化学习”范式的面向搜索的强化学习框架。师生式监督微调首先产生一个格式良好、符合模式的策略初始化;随后的强化学习阶段利用与搜索引擎实时交互中获得的奖励来优化各个查询理解组件,奖励根据该组件的实际运行角色进行定制,而非使用与最终搜索结果绑定的单一奖励。在Roblox搜索上的实验表明,这种针对组件的优化同时提升了各组件的效用与下游搜索质量:相比监督微调策略,NDCG@20提升了8.9个百分点;相比使用单一端到端奖励的训练,提升了3.5个百分点。
cs.AI / 103 / 2609.30186

Jev-Mobile: Jev as an Executor for Mobile GUI Agents

Jev-Mobile:以Jev作为移动GUI智能体的执行器
Zhang, Linghua
Abstract
Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to low-frequency VLM planning and high-frequency lightweight execution: the VLM specifies local goals, the accessibility tree defines a structured executable action space, and Jev, a fast typed decision model, repeatedly selects actions within this space. This design allows multiple GUI actions to be executed under a single VLM decision, reducing expensive VLM inference while preserving adaptive interaction. On the full AndroidWorld task suite, Jev-Mobile achieves 79% task success, compared with 78% for SeeAct-V and 84% for a Step-wise VLM baseline. Among successful trajectories, it reduces mean end-to-end execution time by 32.7% and mean model API cost by 73.4% relative to Step-wise VLM. These results show that decoupling high-level VLM reasoning from low-level action execution can substantially improve mobile GUI agent efficiency while maintaining competitive task performance.
Chinese Translation
视觉-语言模型已成为自主移动GUI智能体的常见基础,但大多数现有系统在几乎每个交互步骤都依赖VLM进行规划与动作定位,导致显著的延迟和模型服务成本。我们提出Jev-Mobile,将这一范式转变为低频VLM规划与高频轻量级执行:VLM指定局部目标,无障碍树定义结构化的可执行动作空间,而Jev——一个快速的类型化决策模型——在该空间中反复选择动作。这种设计使得在单次VLM决策下可以执行多个GUI动作,在保持自适应交互的同时减少了昂贵的VLM推理。在完整的AndroidWorld任务套件上,Jev-Mobile实现了79%的任务成功率,而SeeAct-V为78%,逐步式VLM基线为84%。在成功的轨迹中,相对于逐步式VLM,它将平均端到端执行时间减少了32.7%,平均模型API成本降低了73.4%。这些结果表明,将高层VLM推理与低层动作执行解耦,能够在保持有竞争力的任务性能的同时,显著提高移动GUI智能体的效率。
cs.AI / 104 / 2609.30192

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

SAGE:通过拓扑引导缓解长程推理偏差
Zeng, Xinyue, Zhang, Jiawei, Yan, Yujun, Zhou, Dawei
Abstract
Long-horizon reasoning remains a central challenge for large language models (LLMs) under sparse-reward regimes. We argue that this brittleness arises from two biases induced by complex reasoning spaces: an exploration bias, where models are drawn toward locally plausible but structurally unstable branches, and a compounding bias, where small local deviations accumulate across depth and suppress rare rewards. We introduce Symbolic Closure Analysis (SCA) as a theoretical lens characterizing how branching structures and sparse rewards induce these biases in long-horizon reasoning with local admissibility, and as a design principle for structural priors in less formal reasoning tasks. Motivated by this analysis, we propose SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning. SAGE combines two complementary structural guidance: algebraic sparsification, which projects locally admissible candidates onto operator-indexed algebraic subspaces to suppress spurious branching and mitigate exploration bias, and hyperbolic structural guidance, which embeds reasoning states into a negatively curved space to provide dense depth-wise signals and mitigate compounding bias. Across 12 benchmarks and 7 model families, SAGE outperforms competitive baselines. In particular, SAGE achieves up to an 8-fold improvement on the Andrews-Curtis problem, an open real-world long-horizon task. Code is available at: https://github.com/Susan571/SAGE-NeurIPS2026.
Chinese Translation
长程推理仍然是大型语言模型(LLMs)在稀疏奖励环境下面临的核心挑战。我们认为,这种脆弱性源于复杂推理空间所诱导的两类偏差:一是探索偏差(exploration bias),即模型被吸引至局部看似合理但结构上不稳定的分支;二是累积偏差(compounding bias),即微小的局部偏差随深度不断累积,从而抑制稀有的奖励信号。我们提出符号闭包分析(Symbolic Closure Analysis, SCA),作为一种理论视角,刻画在具有局部可容许性的长程推理中,分支结构与稀疏奖励如何诱导上述偏差;同时将其作为设计原则,用于非形式化推理任务中的结构先验。基于这一分析,我们提出了 SAGE(结构可容许性引导探索,Structural Admissibility-Guided Exploration),这是一个统一框架,通过注入结构引导来缓解长程推理中的探索偏差与累积偏差。SAGE 结合了两种互补的结构引导:一是代数稀疏化(algebraic sparsification),将局部可容许的候选投影到由算子索引的代数子空间中,以抑制虚假分支并缓解探索偏差;二是双曲结构引导(hyperbolic structural guidance),将推理状态嵌入负曲率空间中,以提供密集的深度方向信号并缓解累积偏差。在 12 个基准测试和 7 个模型家族上,SAGE 的表现优于有竞争力的基线方法。特别是在 Andrews-Curtis 问题这一开放的真实长程任务上,SAGE 取得了高达 8 倍的性能提升。代码已发布于:https://github.com/Susan571/SAGE-NeurIPS2026。
cs.AI / 105 / 2609.30199

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

ExplorationBench:在可验证的异星世界中衡量AI系统的探索能力
Zhang, Ming, Xiang, Zhenghao, Gao, Peizhong, Shen, Yujiong, Wang, Yuhui, Yue, Zhonghan, Dou, Shihan, Yin, Zhangyue, Ye, Junjie, Liu, Shichun, Zheng, Weihuang, Chen, Jiahao, Chen, Jiayi, Liu, Hongzhang, Shao, Jiaqi, Gui, Tao, Zhang, Qi, Huang, Xuanjing, Zheng, Suncong, Pan, Maxm
Abstract
Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds: their rules are executable, so every answer can be checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. The benchmark contains two sandboxes, AlienCode (31 discovery targets, 70 tasks) and AlienLogic (24 discovery targets, 70 tasks). Each sandbox provides a flawed manual, task-specific environmental feedback, and a dedicated tool-call schema. Systems use these resources to explore the sandbox, then solve held-out tasks. We evaluate 10 AI systems and find that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can stall or reverse earlier gains. ExplorationBench represents a step towards AI systems that can acquire and apply genuinely new knowledge through exploration in unknown environments.
Chinese Translation
科学发现始于已知问题的尽头。在那里,AI系统必须进行探索:提出假设、设计实验并根据结果进行迭代。然而,评估这种能力十分困难:(1)如何验证一个真正新颖的假设是否成立;(2)如何判断系统是通过探索发现了它,还是仅仅从预训练数据中回忆起相关知识。为此,我们提出了ExplorationBench,它将评估科学探索这一棘手问题转化为一个建立在可验证异星世界(Alien Worlds)之上的具体且可处理的框架:这些世界的规则是可执行的,因此每个答案都可以被精确检验;同时它们与熟悉的知识相冲突,因此仅凭记忆无法解决任务。该基准包含两个沙盒:AlienCode(31个发现目标,70个任务)和AlienLogic(24个发现目标,70个任务)。每个沙盒提供一份有缺陷的手册、针对任务的环境反馈以及专用的工具调用模式(tool-call schema)。系统利用这些资源探索沙盒,然后解决留出的任务。我们评估了10个AI系统,发现最强的系统能够习得并应用不熟悉的规则,但性能在不同轨迹之间存在显著差异,且持续探索可能导致停滞或逆转先前的进展。ExplorationBench代表着向能够在未知环境中通过探索习得并应用真正新知识的AI系统迈出的一步。
cs.AI / 106 / 2609.30205

A Living Benchmark for Information Retrieval from Electronic Health Records

一个用于电子健康记录信息检索的动态基准测试
Cahoon, Jordan L., Stanwyck, Chloe O., Somani, Sulaiman, Chung, Philip, Keet, Kevin R, Black, Kameron C., Fisher, Andrea T., Khemani, Sarita, Liu, Jerry, Ma, Stephen, Maharaj, Saloni K., Pandya, Rita M., Perez-Guerrero, Eduardo, Pillai, Priyanka, Shieh, Lisa, Wu, David J. H., Xie, James, McAvoy, James C., Nguyen, Teresa, Tran, Jessica, Yin, Lucy, Lin, Bridget, Callahan, Alison, Fries, Jason A., Shah, Nigam H., Alsentzer, Emily
Abstract
Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, costly to update, and rapidly become obsolete with evolving technological advancements. We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes. Nineteen clinicians validate the benchmark generator, producing the Benchmark for Retrieving Information in EHRs (BRIE), a continuously maintainable evaluation dataset. Across nine LLMs and five inference strategies, state-of-the-art systems frequently omit clinically important information, particularly for questions requiring synthesis across multiple documents and encounters. Because the generator itself is validated, BRIE supports evaluations that static benchmarks cannot, including the generation of multiple answers that reflect variation in clinician reasoning for robust performance assessment and continuously refreshing benchmark content to guard against leakage. Our results demonstrate that scalable benchmark generation enables rigorous, up-to-date evaluation of clinical LLMs as they are deployed in rapidly evolving healthcare settings.
Chinese Translation
基于大语言模型(LLM)的临床助手正日益被集成到电子健康记录(EHR)系统中,改变着临床医生从患者病历中检索和整合信息的方式。其安全性与实用性取决于严格的评估,然而现有的基准测试依赖人工整理、更新成本高昂,且随着技术的快速演进很快便会过时。我们提出了一个可扩展的框架,能够从纵向EHR记录中自动生成问答对。十九位临床医生对该基准测试生成器进行了验证,产出了电子健康记录信息检索基准(Benchmark for Retrieving Information in EHRs, BRIE)——一个可持续维护的评估数据集。在九个大语言模型和五种推理策略的实验中,最先进的系统经常遗漏临床重要信息,尤其是对于需要跨多份文档和多次就诊进行综合归纳的问题。由于生成器本身经过了验证,BRIE能够支持静态基准测试无法实现的评估,包括生成多个反映临床医生推理差异的答案以进行稳健的性能评估,以及持续刷新基准内容以防止数据泄露。我们的结果表明,可扩展的基准测试生成能够在临床大语言模型部署于快速演进的医疗环境时,为其提供严格且与时俱进的评估。
cs.AI / 107 / 2609.30264

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

AD-WM:用于反事实模型预测控制的动作判别式世界模型
Qiu, Jiabin, Chen, Zixuan, Cao, Hongye, Shi, Jieqi, Huo, Jing, Gao, Yang
Abstract
Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state. A model can therefore achieve low factual prediction error yet poorly distinguish candidate actions. We introduce AD-WM, an action-discriminative joint-embedding world model for counterfactual MPC. AD-WM combines residual latent dynamics with predictor-level action-recovery regularization, using inverse dynamics and a normalized recovery objective motivated by conditional mutual information. Both objectives encourage planning transitions to preserve action information; their auxiliary heads are discarded at test time, leaving MPC unchanged. On OGBench-Cube, AD-WM improves hard-start success from 3.7% to 52.0% over a matched LeWM baseline and improves mean success over the reproduced baseline in four of five simulation environments. Planning diagnostics show that factual prediction error and whole-bank action ranking do not follow the closed-loop success ordering, whereas CEM-aligned elite regret tracks success more closely. With a frozen V-JEPA 2 encoder and matched DROID post-training, AD-WM also improves zero-shot transfer to our Franka setup, increasing basic pick-and-place success from 42.2% to 71.1% without lab-specific adaptation. These results suggest that world models for planning should preserve action-dependent differences needed for counterfactual selection, rather than optimize factual prediction accuracy alone. More videos and code are available at https://ad-wm.github.io/.
Chinese Translation
潜在世界模型(latent world models)通常被训练用于预测事实性转移,而模型预测控制(MPC)需要比较从同一状态出发的不同候选动作。因此,一个模型即使取得较低的事实性预测误差,也可能难以区分候选动作。我们提出了 AD-WM,一种面向反事实MPC的动作判别式联合嵌入世界模型。AD-WM 将残差潜在动力学与预测器层面的动作恢复正则化相结合,利用逆动力学以及由条件互信息启发的归一化恢复目标。这两个目标都促使规划用的转移保留动作信息;其辅助头在测试时被丢弃,因此MPC流程保持不变。在 OGBench-Cube 上,AD-WM 将困难初始条件下的成功率从匹配的 LeWM 基线的 3.7% 提升至 52.0%,并在五个仿真环境中的四个中将平均成功率提升至超过复现的基线。规划诊断表明,事实性预测误差和全动作库排序并不遵循闭环成功率的排序,而与 CEM 对齐的精英遗憾(elite regret)则更紧密地跟随成功率变化。在冻结的 V-JEPA 2 编码器及匹配的 DROID 后训练条件下,AD-WM 还提升了对我们 Franka 平台的零样本迁移能力,将基础抓取放置任务的成功率从 42.2% 提高到 71.1%,且无需针对实验室的适配。这些结果表明,用于规划的世界模型应保留反事实选择所需的动作依赖差异,而不应仅优化事实性预测精度。更多视频和代码请见 https://ad-wm.github.io/。
计算机视觉 (Computer Vision)
111
cs.CV / 1 / 2609.28539

$\unicode{x1F493}$Heartian: Physiology-Aware Relightable Gaussian Head Avatar

💔Heartian:生理感知的可重光照高斯头部化身
Fan, Xiaoyue, Echevarria, Jose, Paruchuri, Akshay, Akşit, Kaan
Abstract
Gaussian head avatars typically model intrinsic facial appearance as temporally static, omitting subtle cardiac-induced skin-color variation. We propose $\unicode{x1F493}$Heartian, a physiology-aware modulation framework that learns cardiac-cycle-dependent per-frame albedo modulation of facial skin-region Gaussians within a relightable head avatar to encode remote photoplethysmography (rPPG) signals. Using synchronized contact PPG supervision, $\unicode{x1F493}$Heartian models the prescribed cardiac waveform as the sum of two Gaussian functions and learns per-frame spatial residuals via a lightweight MLP. Across 152 stationary recordings from UBFC-rPPG, PURE, and MMPD, attribute-space recovery of the supplied signal achieves a pooled recording-level heart-rate MAE of 0.29 bpm and MAPE of 0.38%. The signals remain detectable after rendering by benchmark rPPG methods, with the best tested configuration - a motion-augmented TS-CAN decoder pretrained on UBFC-rPPG - recovering heart rate from the rendered MMPD avatars at 0.97 bpm MAE and 1.21% MAPE. Meanwhile, $\unicode{x1F493}$Heartian maintains reconstruction quality comparable to the baseline, with negligible average PSNR degradation of 0.005 dB. Overall, our work embeds recoverable rPPG signals as controllable material attributes to subject-specific Gaussian head avatars while retaining the reconstruction quality.
Chinese Translation
高斯头部化身通常将面部的内在表观建模为时间上静态的,忽略了由心脏活动引起的细微肤色变化。我们提出💔Heartian,一个生理感知的调制框架,它在可重光照头部化身中学习面部皮肤区域高斯随心动周期变化的逐帧反照率调制,以编码远程光电容积描记(rPPG)信号。利用同步的接触式PPG监督,💔Heartian将预定的心电波形建模为两个高斯函数之和,并通过一个轻量级MLP学习逐帧空间残差。在来自UBFC-rPPG、PURE和MMPD的152段静止录制数据上,对所提供信号在属性空间的恢复实现了合并录制级别的心率MAE为0.29 bpm、MAPE为0.38%。经过基准rPPG方法渲染后,信号仍然可被检测,其中测试的最优配置——在UBFC-rPPG上预训练的运动增强TS-CAN解码器——从渲染的MMPD化身中恢复心率的MAE为0.97 bpm、MAPE为1.21%。同时,💔Heartian保持了与基线相当的重构质量,平均PSNR下降仅0.005 dB,可忽略不计。总体而言,我们的工作将可恢复的rPPG信号作为可控材质属性嵌入到面向特定对象的高斯头部化身中,同时保持了重构质量。
cs.CV / 2 / 2609.28580

Token Clustering and Semantic Sequence Mamba for Hyperspectral Image Classification

面向高光谱图像分类的词元聚类与语义序列Mamba
Zhu, Yimin, Elahi, Mahmood, Xu, Lincoln Linlin
Abstract
Although hyperspectral images (HSIs) provide rich spectral-spatial information, accurate pixel-level classification remains challenging because of spectral-spatial heterogeneity and complex spatial structures. Existing vision state space models (Mamba) typically construct sequences according to predefined spatial neighborhoods, without explicitly accounting for semantic similarity or spatial non-stationarity. To address this limitation, we propose Token Clustering and Semantic Sequence Mamba (STMamba), which organizes sparse tokens into semantically coherent sequences for hyperspectral image classification with the following features. First, at the macro level, a hierarchical encoder decoder progressively selects semantic tokens with the Token Clustering Module (TCM) and restores dense features using a parameter-free Cross-scale Neighborhood Attention (CNA) Upsampler. Second, at the micro level, TCM first identifies representative cluster centers through density-aware clustering and estimates soft memberships based on feature similarity. A quadtree-based dynamic selection strategy then retains sparse and spatially distributed tokens from each semantic cluster, forming coherent semantic-token sequences while reducing redundant pixel-wise representations. Third, parallel Spatial and Spectral Semantic-wise Sequencing Mamba (SWSM) modules capture complementary long-range spatial and spectral dependencies within homogeneous semantic token sequences while suppressing irrelevant interactions across heterogeneous regions. Experimental results on three large-scale benchmark datasets demonstrate that STMamba outperforms the SOTA methods with respect to quantitative and qualitative results.
Chinese Translation
尽管高光谱图像(HSI)提供了丰富的光谱-空间信息,但由于光谱-空间异质性和复杂的空间结构,精确的像素级分类仍然具有挑战性。现有的视觉状态空间模型(Mamba)通常根据预定义的空间邻域来构建序列,而没有显式地考虑语义相似性或空间非平稳性。为了解决这一局限,我们提出了词元聚类与语义序列Mamba(STMamba),它将稀疏词元组织成语义连贯的序列用于高光谱图像分类,具有以下特点。首先,在宏观层面,分层编码器-解码器通过词元聚类模块(Token Clustering Module, TCM)逐步选择语义词元,并使用无参数的跨尺度邻域注意力(Cross-scale Neighborhood Attention, CNA)上采样器恢复稠密特征。其次,在微观层面,TCM首先通过密度感知聚类识别代表性的聚类中心,并基于特征相似性估计软隶属度。随后,基于四叉树的动态选择策略从每个语义聚类中保留稀疏且空间分布的词元,在减少冗余的逐像素表示的同时,形成连贯的语义词元序列。第三,并行的空间与光谱语义级序列Mamba(Spatial and Spectral Semantic-wise Sequencing Mamba, SWSM)模块在同质语义词元序列内捕获互补的远程空间和光谱依赖关系,同时抑制跨异质区域的不相关交互。在三个大规模基准数据集上的实验结果表明,STMamba在定量和定性结果上均优于最先进(SOTA)的方法。
cs.CV / 3 / 2609.28610

UltraBench 2: Towards Robust Evaluation of Vision Foundation Models on Ultrasound

UltraBench 2:面向超声视觉基础模型稳健评估的基准测试
Radhachandran, Ashwath, Tupper, Adam, Gagné, Christian, Speier, William
Abstract
Benchmarking is an increasingly critical part of research in machine learning and the domains where it is applied, including healthcare. Yet, despite the steady development of new ultrasound foundation models in recent years, the development of well-designed benchmarks to evaluate them has lagged behind. This deficiency has led to fragmented and inconsistent evaluations of competing models, making it difficult to measure progress. To address this issue, we introduce UltraBench 2, a comprehensive benchmark with wide anatomical and task coverage, and a focus on standardization, reproducibility, and ease-of-use. Using this benchmark, we compare existing vision foundation models for ultrasound image analysis. Our analyses demonstrate that ultrasound-specific pretraining still leads on classification, but that state-of-the-art general-purpose models have drawn level on segmentation.
Chinese Translation
基准测试已成为机器学习研究及其应用领域(包括医疗健康)中日益重要的组成部分。然而,尽管近年来新的超声基础模型不断涌现,用于评估这些模型的设计良好的基准测试的发展却相对滞后。这一缺陷导致了对各竞争模型的评估零散且不一致,使得难以衡量研究进展。为解决这一问题,我们提出了UltraBench 2,这是一个覆盖广泛解剖结构和任务的综合性基准测试,注重标准化、可复现性和易用性。利用该基准,我们比较了现有的用于超声图像分析的视觉基础模型。我们的分析表明,超声专用的预训练在分类任务上仍然领先,但最先进的通用模型在分割任务上已与之持平。
cs.CV / 4 / 2609.28645

PePESeg3D: Perception Prior Enhances Multi-Scale Segmentation for 3D Gaussian Splatting

PePESeg3D:感知先验增强面向3D高斯泼溅的多尺度分割
Choi, Sungjae, Koh, Seunghee, Kim, Junmo
Abstract
Recent advancements in 3D Gaussian Splatting (3DGS) have extended its capabilities to multi-scale segmentation. Existing methods reconstruct a scene with Gaussian primitives and learn multi-scale segmentation features separately, which leaves the geometry unaware of semantic structure and the feature learning dependent on incomplete mask supervision. To address these limitations, we present PePESeg3D, a novel framework that injects perception priors into a multi-scale 3D Gaussian segmentation pipeline. To fully exploit perception priors, we integrate them not only into contrastive feature learning but also into the upstream geometry reconstruction. Specifically, PePE Reconstruction incorporates monocular depth and mask constraints to ensure semantically coherent object structures. Building on this aligned geometry, PePE Contrastive Learning leverages dense depth-color cues and view-consistent centroid supervision to compensate for the incompleteness of multi-scale masks obtained from a 2D foundation model. Extensive experiments on the SPIn-NeRF, LERF-Mask, and NVOS benchmarks demonstrate that PePESeg3D achieves state-of-the-art performance in both multi-scale segmentation and scene reconstruction, highlighting the importance of integrating perception priors into both geometry optimization and feature learning for accurate multi-scale 3D segmentation. Our code is available at https://github.com/BeCow5X5/PePESeg3D.
Chinese Translation
三维高斯泼溅(3D Gaussian Splatting, 3DGS)的最新进展已将其能力拓展至多尺度分割。现有方法使用高斯基元重建场景并分别学习多尺度分割特征,这导致几何结构无法感知语义结构,且特征学习依赖于不完整的掩码监督。为解决这些局限,我们提出了PePESeg3D,一种将感知先验注入多尺度三维高斯分割流水线的新颖框架。为了充分利用感知先验,我们不仅将其融入对比特征学习,还融入上游的几何重建。具体而言,PePE重建(PePE Reconstruction)引入单目深度和掩码约束,以确保语义一致的目标结构;在此对齐几何的基础上,PePE对比学习(PePE Contrastive Learning)利用稠密的深度-颜色线索和视角一致的质心监督,弥补由二维基础模型获得的多尺度掩码的不完整性。在SPIn-NeRF、LERF-Mask和NVOS基准上的大量实验表明,PePESeg3D在多尺度分割和场景重建方面均达到了最先进的性能,凸显了将感知先验同时融入几何优化与特征学习对于精确多尺度三维分割的重要性。我们的代码可在 https://github.com/BeCow5X5/PePESeg3D 获取。
cs.CV / 5 / 2609.28684

M-plicits: Neural Implicit Surfaces via Nested Multiscale Residuals

M-plicits:基于嵌套多尺度残差的神经隐式曲面
da Silva, Vinícius, Melo, Isabelle, Bessa, Matheus, Schardong, Guilherme, Schirmer, Luiz, Araújo, André, Gonçalves, Nuno, Lopes, Hélio, Raposo, Alberto, Velho, Luiz, Novello, Tiago
Abstract
Encoding input coordinates with sinusoidal functions into multi-layer perceptrons (MLPs) has proven effective for implicit neural representations (INRs) of surfaces defined as zero-level sets. However, existing methods often struggle to balance training efficiency, rendering speed, and noise robustness: single-MLP approaches are expensive at inference, grid-based representations are fast but can limit surface smoothness and overfit input noise, and previous multiscale approaches frequently capture noise and produce artifacts due to hard spectral truncation. To address these limitations, we propose M-plicits, a multiscale framework that models surfaces as a residual sum of MLPs trained via a sequence of nested neighborhoods. Unlike existing residual approaches that rely on standard domain-wide sampling and require costly mesh extraction for visualization, our method strictly localizes supervision to narrow bands around the previous zero-level sets. This nested design naturally provides robustness against noisy input data: the coarse network acts as a low-pass filter that establishes a clean geometric prior, while subsequent residuals progressively refine the geometry without fitting to high-frequency artifacts. We further introduce a multiscale sphere-tracing algorithm and a GEMM-based analytical normal computation that bypasses auto-differentiation entirely, yielding high-fidelity real-time rendering. On Stanford and Thingi32, M-plicits achieves the best mean Chamfer distance in the coarse configuration and the best median Chamfer distance and IoU in the fine configuration, with substantially better noise robustness than iNGP, BACON, and IDF, while using an order of magnitude fewer parameters than grid-based baselines. Code, models, and data will be released at https://github.com/dsilvavinicius/m-plicits.
Chinese Translation
将输入坐标用正弦函数编码并输入多层感知机(MLP),已被证明对定义为零水平集的曲面的隐式神经表示(INR)十分有效。然而,现有方法往往难以在训练效率、渲染速度和噪声鲁棒性之间取得平衡:单MLP方法推理代价高;基于网格的表示速度快,但可能限制曲面平滑度并对输入噪声过拟合;以往的多尺度方法由于硬性谱截断,常常捕捉到噪声并产生伪影。为解决这些局限,我们提出M-plicits,一种将曲面建模为多个MLP残差之和的多尺度框架,并通过一系列嵌套邻域进行训练。与依赖标准全域采样且需要代价高昂的网格提取才能可视化的现有残差方法不同,我们的方法将监督严格局部化到先前零水平集周围的狭窄条带内。这种嵌套设计自然地提供了对噪声输入数据的鲁棒性:粗尺度网络充当低通滤波器,建立干净的几何先验,而后续的残差在不拟合高频伪影的前提下逐步细化几何。我们进一步提出了一种多尺度球体追踪算法和一种基于GEMM的解析法向计算方法,完全绕过自动微分,实现高保真实时渲染。在Stanford和Thingi32数据集上,M-plicits在粗尺度配置下取得最佳平均Chamfer距离,在细尺度配置下取得最佳中位Chamfer距离和IoU,其噪声鲁棒性显著优于iNGP、BACON和IDF,同时参数量比基于网格的基线方法少一个数量级。代码、模型和数据将发布于 https://github.com/dsilvavinicius/m-plicits。
cs.CV / 6 / 2609.28741

GeoNLI - A Natural Language Interpreter for Satellite Imagery

GeoNLI——一种面向卫星影像的自然语言解释器
Gandhe, Ashutosh, Rawat, Anupam, Sethi, Geet, Nasiruddin, Kabir, Kotecha, Madhav, Shah, Panav, Sawarn, Rakshit, Nayak, Soumitra
Abstract
Multi-modal multitasking models have shown strong performance on remote sensing datasets. However, because these models are trained on heterogeneous data and vary across tasks, designing a unified model that performs well in captioning, visual question answering (VQA), and visual grounding remains challenging. In this work, we evaluate several models on the VRS Bench and NWPU-VHR-10 datasets. The EarthMind model demonstrates strong results in both captioning and VQA. For grounding, we propose multiple pipelines - RemoteSAM-SAM-v1, RemoteSAM-SAM-v2, and DiffuSAM - and ultimately adopt a majority-voting ensemble across EarthMind, RemoteSAM, SAM3, Falcon, RemoteSAM-SAM3-v1, RemoteSAM-SAM3-v2, and DiffuSAM predictions. Our unified, modular pipeline integrates advanced SAM variants with multimodal LLMs to jointly perform captioning, VQA, and grounding. It achieves 82% accuracy on captioning and 83.32% on VQA, with 90.94%, 52.04%, and 92.06% for binary, numeric, and semantic question types respectively. For grounding, it attains 64.94% accuracy. By combining diverse VLMs with our custom RemoteSAM-SAM3 models through ensemble majority voting, the system delivers more accurate and consistent remote-sensing understanding than task-specific approaches.
Chinese Translation
多模态多任务模型在遥感数据集上已展现出强劲性能。然而,由于这些模型基于异构数据进行训练且在不同任务间存在差异,设计一个在图像描述生成、视觉问答(VQA)和视觉定位(visual grounding)任务上均表现良好的统一模型仍然具有挑战性。在本工作中,我们在 VRS Bench 和 NWPU-VHR-10 数据集上评估了多个模型。EarthMind 模型在图像描述生成和视觉问答任务中均取得了优异结果。在视觉定位方面,我们提出了多种流水线方案——RemoteSAM-SAM-v1、RemoteSAM-SAM-v2 和 DiffuSAM,并最终采用对 EarthMind、RemoteSAM、SAM3、Falcon、RemoteSAM-SAM3-v1、RemoteSAM-SAM3-v2 和 DiffuSAM 预测结果进行多数投票的集成方法。我们的统一模块化流水线将先进的 SAM 变体与多模态大语言模型相结合,可联合执行图像描述生成、视觉问答和视觉定位任务。该系统在图像描述生成任务上达到 82% 的准确率,在视觉问答任务上达到 83.32% 的准确率,其中二元问题、数值问题和语义问题类型的准确率分别为 90.94%、52.04% 和 92.06%。在视觉定位任务上,系统达到 64.94% 的准确率。通过将多种视觉语言模型(VLM)与我们自研的 RemoteSAM-SAM3 模型进行多数投票集成,该系统相比特定任务的方法能够提供更准确、更一致的遥感影像理解能力。
cs.CV / 7 / 2609.28757

Small yet Assistive: Spatially-Aware Post-Training for Low Vision

小而有益:面向低视力人群的空间感知后训练方法
Choudhary, Rishabh, Raj, Shreyansh, Goyal, Umesh, Kashyap, Shubh, Kumar, Shrestha, Jena, Sushovan, Kumar, Komal, Cholakkal, Hisham, Nigam, Aditya
Abstract
An estimated 1 billion people worldwide live with vision impairment, yet current vision-language models (VLMs) produce descriptions too vague for safe navigation by blind and low-vision (BLV) users. Large VLMs can generate high-quality audio-description-compliant narrations but cannot run on mobile devices; small VLMs offer competitive latency but lack spatial detail, directional cues, and hazard awareness for navigational assistance. We present Smol-VL-BLV, a compact VLM for blind and low-vision users that closes this gap using a 500M decoder transformer model and two post-training mechanisms: (1) teacher-student distillation and (2) Group Relative Policy Optimization (GRPO) with a composite BLV reward targeting directional language, metric distances, and hazard detection. Because multi-stage post-training can induce catastrophic forgetting, we add a lightweight finetuning stage after the last stage GRPO finetuning to recover general descriptive quality while preserving BLV-specific spatial grounding. Our best model substantially outperforms the baseline across various benchmarks, including tasks: VQA, BLV captioning, OCR, and latency. Compared with the baseline for relative improvement, it improves the Spatial score gain of 19.3%, and the Social score gain of 14.8%. It also increases OCR-Bench by 101.5%, and raises TextVQA accuracy by 44.2%. These results show that BLV-focused post-training improves both accessibility-specific spatial grounding and general visual-text reasoning. Deployed on a mid-range Android smartphone via Mixed-Precision Quantization, the model remains approx. 450 MB and runs entirely on-device, offline and without network dependency, generating descriptions with latency dependent on host hardware capabilities. Our model, dataset, and code is publicly released at https://smol-vl-blv.github.io/Smol-VL-BLV-website/
Chinese Translation
据估计,全球约有10亿人患有视力障碍,然而当前的视觉语言模型(VLM)生成的描述过于模糊,无法支持盲人和低视力(BLV)用户进行安全导航。大型VLM能够生成符合音频描述标准的高质量旁白,但无法在移动设备上运行;小型VLM虽然具有较低的延迟,但缺乏导航辅助所需的空间细节、方向线索和危险感知能力。我们提出了Smol-VL-BLV,这是一个面向盲人和低视力用户的紧凑型VLM,它采用500M参数的解码器Transformer模型和两种后训练机制来弥合这一差距:(1)师生蒸馏;(2)采用复合BLV奖励的群体相对策略优化(GRPO),该奖励针对方向性语言、公制距离和危险检测。由于多阶段后训练可能导致灾难性遗忘,我们在最后阶段的GRPO微调之后增加了一个轻量级微调阶段,以在恢复通用描述质量的同时保留BLV特有的空间定位能力。我们的最佳模型在多个基准测试中显著优于基线,包括视觉问答(VQA)、BLV图像描述、OCR和延迟等任务。相对于基线的提升幅度,其空间分数提升了19.3%,社交分数提升了14.8%;OCR-Bench提高了101.5%,TextVQA准确率提升了44.2%。这些结果表明,以BLV为重点的后训练既能改善面向无障碍的空间定位能力,也能提升通用的视觉-文本推理能力。通过混合精度量化部署在中端安卓智能手机上后,该模型大小约为450 MB,可完全在设备端离线运行且不依赖网络,生成描述的延迟取决于宿主硬件的性能。我们的模型、数据集和代码已在 https://smol-vl-blv.github.io/Smol-VL-BLV-website/ 公开发布。
cs.CV / 8 / 2609.28796

DrGait: Biomechanically Grounded Visual Reasoning for Interpretable Clinical Gait Analysis

DrGait:面向可解释临床步态分析的生物力学基础视觉推理
Yin, Xiangyu, Wang, Shiqi, Alamri, Abrar, Aljohani, Yasir, Liu, Weichen, Fiedler, Goeran, Gao, Wei
Abstract
Current automated gait analysis for clinical applications relies on uninterpretable black-box classifiers. Although Vision-Language Models (VLMs) offer strong reasoning capabilities, applying them directly to gait videos often leads to hallucinations, because they struggle to measure subtle geometric deviations from raw visual contexts. To address this, we introduce DrGait, a training-free agentic framework that shifts the VLM's role from a direct visual reasoner to a clinical planner. DrGait decouples semantic reasoning from geometric perception through a structured Triage-Verification-Synthesis (TVS) workflow. Given an input video and a set of basic spatiotemporal metrics, the DrGait agent first performs a heuristic triage to propose diagnostic hypotheses, which are then verified by autonomously calling deterministic biomechanical tools that operate on reconstructed 3D mesh trajectories, segmented 2D pose tracks, and event-centered video evidence. Finally, a closed-loop mechanism recursively updates the agent's reasoning context based on the feedback. By anchoring VLM's reasoning in verifiable geometric and temporal measurements, DrGait reduces hallucinations, achieving competitive diagnostic accuracy while generating transparent and audit-ready clinical reports.
Chinese Translation
当前面向临床应用的自动化步态分析依赖于不可解释的黑盒分类器。尽管视觉-语言模型(VLM)具备强大的推理能力,但将其直接应用于步态视频往往会产生幻觉,因为它们难以从原始视觉上下文中测量细微的几何偏差。为解决这一问题,我们提出了DrGait,一个无需训练的智能体(agentic)框架,将VLM的角色从直接的视觉推理器转变为临床规划者。DrGait通过结构化的分诊-验证-综合(Triage-Verification-Synthesis, TVS)工作流,将语义推理与几何感知解耦。给定输入视频和一组基础时空指标,DrGait智能体首先执行启发式分诊以提出诊断假设,随后通过自主调用确定性生物力学工具对这些假设进行验证,这些工具作用于重建的3D网格轨迹、分割的2D姿态轨迹以及以事件为中心的视频证据。最后,闭环机制基于反馈递归地更新智能体的推理上下文。通过将VLM的推理锚定在可验证的几何和时序测量上,DrGait减少了幻觉,在生成透明且可审计的临床报告的同时,取得了具有竞争力的诊断准确率。
cs.CV / 9 / 2609.28811

DeltaWAM: Delta World Action Models for Bimanual Manipulation

DeltaWAM:用于双臂操作的增量世界动作模型
Yan, Han, Xiang, Zishang, Jiang, Haokai, Zhang, Zeyu, Wang, Qilin, Guo, Weiyu, Guo, Yandong, Shi, Boxin, Tang, Hao
Abstract
World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control by jointly modeling visual dynamics and actions. Existing WAMs, however, predict dense future frames during training, repeatedly modeling largely unchanged content and coupling action-conditioned dynamics to nuisance appearance variations. At inference, processing each complete observation with the heavy video expert bottlenecks few-step action generation. Accordingly, we propose DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing. We further develop Streaming Delta Memory (SDM), which updates cached anchor context with compact observed deltas, reducing heavy video-expert processing. On RoboTwin, DeltaWAM with SDM improves average success over Fast-WAM from 81.3% to 85.4% in the clean setting and from 75.8% to 83.9% under visual randomization. The three architectures reduce training FLOPs by 17.78-23.77%, while SDM reduces one-step inference latency and FLOPs by 36.57% and 31.55%, respectively; real-world evaluations further show the highest overall success rate and normalized progress among the evaluated policies. Code: https://github.com/AIGeeksGroup/DeltaWAM. Website: https://aigeeksgroup.github.io/DeltaWAM.
Chinese Translation
世界动作模型通过联合建模视觉动态与动作,将预训练视频生成器中的视觉和运动先验迁移到机器人控制中。然而,现有WAM在训练时预测密集的未来帧,反复建模大量基本不变的内容,并将动作条件下的动态与无关的外观变化耦合在一起。在推理阶段,使用庞大的视频专家处理每个完整观测,成为少步动作生成的瓶颈。为此,我们提出DeltaWAM,其采用密集锚点流、稀疏增量流和动作流联合预测视觉增量和动作,并设计了三种在表示与计算共享方式上不同的架构。我们进一步开发了流式增量记忆(Streaming Delta Memory, SDM),利用紧凑的观测增量更新缓存的锚点上下文,从而减少庞大的视频专家处理。在RoboTwin基准上,带有SDM的DeltaWAM将平均成功率相比Fast-WAM在干净设置下从81.3%提升至85.4%,在视觉随机化设置下从75.8%提升至83.9%。三种架构将训练FLOPs降低17.78%–23.77%,而SDM将单步推理延迟和FLOPs分别降低36.57%和31.55%;真实世界评估进一步表明,在所评测的策略中其总体成功率和归一化进度均为最高。代码:https://github.com/AIGeeksGroup/DeltaWAM。网站:https://aigeeksgroup.github.io/DeltaWAM。
cs.CV / 10 / 2609.28813

CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models

CinematicVQA:面向大型视觉语言模型的电影语言推理基准测试
Xing, Shuo, Verlani, Pooja, Adsumilli, Balu, Tu, Zhengzhong
Abstract
Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable progress in video question answering, existing benchmarks primarily focus on identifying low-level techniques rather than understanding their storytelling impact. To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions. Through comprehensive evaluation of state-of-the-art LVLMs, we reveal a striking semantic gap: models consistently perform higher on describing visual presentations than on identifying the underlying techniques. Surprisingly, Chain-of-Thought prompting fails to provide consistent gains and degrades performance for most models, suggesting that current LVLMs lack sufficient cinematic domain knowledge to benefit from step-by-step reasoning. Fine-tuning on \textsc{CinematicVQA-train} yields consistent improvements, particularly for narrative function and multi-hop reasoning. Overall, \textsc{CinematicVQA} serves both as a rigorous benchmark for cinematic evaluation in LVLMs and as a practical dataset for training more film-aware video models.
Chinese Translation
电影摄影(cinematography)是通过构图、灯光和摄影机操作进行视觉叙事的艺术,从根本上塑造了观众感知视频内容并与之产生情感共鸣的方式。尽管大型视觉语言模型(LVLMs)在视频问答方面取得了显著进展,但现有基准主要关注识别低层次的技术手段,而非理解其叙事作用。为解决这一问题,我们提出了CinematicVQA,这是首个专注于电影视频理解(cinematic video understanding)的基准,它超越单纯的技术识别,致力于评估电影语言推理(film-grammar reasoning)。为此,我们引入了电影场景图(Cinematic Scene Graph, CSG)——一种将拍摄技术与感知效果及叙事功能相联系的结构化表示。通过对最先进LVLMs的全面评估,我们揭示了一个显著的语义鸿沟:模型在描述视觉呈现方面的表现始终优于识别其背后技术手段的表现。令人惊讶的是,思维链(Chain-of-Thought)提示未能带来一致的提升,反而降低了大多数模型的性能,这表明当前的LVLMs缺乏足够的电影领域知识,无法从逐步推理中获益。在CinematicVQA-train上进行微调则带来了持续的性能提升,尤其是在叙事功能和多跳推理方面。总体而言,CinematicVQA既是评估LVLMs电影理解能力的严格基准,也是训练更具电影感知能力的视频模型的实用数据集。
cs.CV / 11 / 2609.28836

M$^2$PFN: End-to-End Disentangled Alignment for Generalizable Multimodal In-Context Learning in Alzheimer's Disease

M$^2$PFN:面向阿尔茨海默病可泛化多模态上下文学习的端到端解耦对齐
Zhong, Lujia, Huang, Shuo, Zhang, Jianwei, Nie, Xinyu, Shi, Yonggang
Abstract
While various multimodal methods combining imaging and tabular data for Alzheimer's disease (AD) diagnosis were proposed, they are often limited in generalization across cohorts. In-context learning (ICL) has demonstrated excellent generalization performances and high flexibility in foundational tabular models such as TabPFN. To extend TabPFN's ICL to multimodal AD analysis, the main obstacle is that TabPFN is meta-trained on synthetic tabular priors that do not naturally match the statistical structure of image-derived features. We propose M$^2$PFN, an end-to-end framework that turns this tabular foundation model into a multimodal AD predictor. M$^2$PFN (i) performs differentiable inference through TabPFN's transformer, back-propagating task gradients into 3D-MRI and tabular encoders; (ii) aligns the two modalities into a shared subspace, via disentanglement and a contrastive objective, matched to the ICL engine's prior; and (iii) folds in a frozen tabular-only prediction through a learnable gated shortcut. Because the ICL engine stays frozen, its in-context mechanism is preserved for test-time generalization, while end-to-end training shapes the encoders into features it can exploit. On ADNI ($n=2240$, three-class CN/MCI/AD), M$^2$PFN attains $65.55\%$ macro-F1 and $82.21\%$ macro-AUC, surpassing a comprehensive set of unimodal and multimodal baselines. By swapping only the head for a TabPFN regressor, the same architecture regresses baseline MMSE on a $1250$-subject sub-cohort to test MAE $1.743$, outperforming every multimodal baseline. On two external cohorts (OASIS-3 and SCAN) with no retraining, M$^2$PFN achieves the best AUC and the lowest MMSE MAE across all baselines, and transfers even when the cognitive instrument changes.
Chinese Translation
尽管已有多种结合影像与表格数据的阿尔茨海默病(AD)诊断多模态方法被提出,但它们在跨队列泛化方面往往存在局限。上下文学习(In-Context Learning, ICL)已在TabPFN等基础表格模型中展现出优异的泛化性能和高度灵活性。将TabPFN的ICL扩展到多模态AD分析的主要障碍在于,TabPFN是在与影像衍生特征的统计结构并不天然匹配的合成表格先验上进行元训练的。我们提出M$^2$PFN,一个将该表格基础模型转变为多模态AD预测器的端到端框架。M$^2$PFN:(i) 通过TabPFN的Transformer实现可微推理,将任务梯度反向传播至3D-MRI编码器和表格编码器;(ii) 通过解耦和对比目标将两种模态对齐到与ICL引擎先验相匹配的共享子空间;(iii) 通过可学习的门控捷径融入一个冻结的纯表格预测。由于ICL引擎保持冻结,其上下文机制得以保留以实现测试时的泛化,同时端到端训练将编码器塑造为其可利用的特征。在ADNI数据集(n=2240,CN/MCI/AD三分类)上,M$^2$PFN取得65.55%的macro-F1和82.21%的macro-AUC,超越了一系列全面的单模态和多模态基线方法。仅需将输出头替换为TabPFN回归器,同一架构即可在1250名受试者的子队列上对基线MMSE进行回归,测试MAE为1.743,优于所有多模态基线。在两个外部队列(OASIS-3和SCAN)上无需重新训练,M$^2$PFN在所有基线中取得最佳AUC和最低的MMSE MAE,并且即使认知测评工具发生变化仍能实现迁移。
cs.CV / 12 / 2609.28851

Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning

看似相同,回答却不同:用于鲁棒视觉语言推理的翻转方向引导方法
Jung, Yeonsung, Jeong, Joonhyun, Pham, Hoang, Kim, Joowon, Park, Yoonsik, Lai, Viet Dac, Yang, Eunho
Abstract
Vision-language models (VLMs) achieve strong visual reasoning performance, yet subtle changes from routine image capture and processing can alter their reasoning trajectories even when images appear nearly identical. In long-horizon generation, the resulting activation shifts may accumulate across decoding steps, progressively altering reasoning tokens and ultimately changing the final answer, a phenomenon referred to as answer flips. To address this instability, we propose FlipDir (Flip-Direction Steering), a training-free inference-time method that estimates a low-rank flip-inducing activation subspace from contrastive pairs of original and answer-flipping inputs and selectively steers hidden states during decoding. A margin-based gate limits subspace attenuation to uncertain decoding steps, recovering original predictions while preserving stable ones. To evaluate robustness beyond accuracy or consistency on fixed test sets, we introduce VisFlip, a benchmark framework that constructs evaluation groups for a target model and visual variation setting to separately assess recovery of original predictions and preservation of stable ones. VisFlip spans nine dataset-variation combinations across scientific reasoning, robot-scene understanding, and medical VQA, covering subtle visual variations common in each domain. Experiments across 18 settings demonstrate that FlipDir consistently outperforms existing methods on the combined recovery and preservation metric. We will make our code publicly available.
Chinese Translation
视觉语言模型(VLMs)在视觉推理任务中表现出强大的性能,然而常规图像采集与处理中的细微变化可能会改变其推理轨迹,即使图像看起来几乎完全相同。在长程生成过程中,由此产生的激活偏移可能在解码步骤间不断累积,逐渐改变推理标记(reasoning tokens),最终导致最终答案发生变化,这一现象被称为答案翻转(answer flips)。为解决这一不稳定性问题,我们提出了FlipDir(Flip-Direction Steering,翻转方向引导),这是一种无需训练的推理时方法,通过对比原始输入与引发答案翻转的输入所构成的对比样本对,估计出一个低秩的翻转诱导激活子空间,并在解码过程中有选择地对隐藏状态进行引导。基于间隔(margin)的门控机制将子空间衰减限制在不确定的解码步骤上,从而在恢复原始预测的同时保留稳定的预测。为了在固定测试集的准确率或一致性之外评估鲁棒性,我们引入了VisFlip,一个面向目标模型和视觉变化设置构建评估组的基准框架,用于分别评估原始预测的恢复能力和稳定预测的保持能力。VisFlip涵盖九种数据集-变化组合,涉及科学推理、机器人场景理解和医学视觉问答(VQA),并包含各领域常见的细微视觉变化。在18种设置下的实验表明,FlipDir在恢复与保持的综合指标上始终优于现有方法。我们将公开我们的代码。
cs.CV / 13 / 2609.28857

MEVL-STP: Multi-Encoder and Vision Language Model for Arbitrarily Shaped Scene Text Spotting

MEVL-STP:面向任意形状场景文本检测的多编码器与视觉语言模型方法
Anand, Aman, Roy, Partha Pratim, Palaiahnakote, Shivakumara
Abstract
Scene text spotting remains challenging for arbitrarily shaped text instances such as curved signs and dense multi-oriented characters in natural images, where tightly coupled architectures propagate localization errors directly into recognition failures. We present a two-stage pipeline that combines multi-encoder segmentation with vision-language model recognition to address this problem. In the detection stage, six frozen vision encoders (CLIP, DINOv2, SigLIP, EVA-CLIP, SAM, and ConvNeXt) extract complementary features spanning semantic, spatial, and texture spectra, which are fused through a trainable hierarchical Feature Pyramid Network with channel attention and decoded via a deep-supervision Progressive Scale Expansion network to generate precise instance-level text masks. By keeping the encoders frozen, their independently learned feature spaces remain orthogonal during fusion, preventing the feature homogenization that degrades boundary precision in single-backbone detectors. The detection stage produces tight polygon masks that conform to the actual shape of curved and arbitrarily oriented text, rather than axis-aligned rectangles that inevitably include background content. In the recognition stage, these polygon-masked crops isolate the target text from surrounding clutter, allowing a Qwen3-VL-8B-Instruct model, fine-tuned via Low-Rank Adaptation on polygon-cropped scene text, to focus purely on reading the text without interference from neighbouring words or background noise. Without any synthetic pretraining data, our method achieves 91.99% detection F-measure and 85.86% end-to-end H-mean on CTW1500, setting a new state of the art and achieving strong performance on Total-Text and ICDAR 2015 without any synthetic training data. Code is available at https://github.com/doubleblind-afk/MEVL-STP
Chinese Translation
对于自然图像中诸如弯曲标志牌和密集多方向字符等任意形状的文本实例,场景文本检测(scene text spotting)仍然极具挑战性:紧耦合架构会使定位误差直接传播为识别失败。为此,我们提出了一种两阶段流水线,将多编码器分割与视觉语言模型识别相结合。在检测阶段,六个冻结的视觉编码器(CLIP、DINOv2、SigLIP、EVA-CLIP、SAM 和 ConvNeXt)提取涵盖语义、空间和纹理谱系的互补特征,这些特征通过一个带通道注意力的可训练层级特征金字塔网络(FPN)进行融合,并经由采用深度监督的渐进尺度扩展网络(PSE)解码,以生成精确的实例级文本掩码。通过保持编码器冻结,其独立学习到的特征空间在融合过程中保持正交,从而避免了单一主干检测器中降低边界精度的特征同质化问题。检测阶段生成贴合弯曲及任意方向文本实际形状的紧致多边形掩码,而不是不可避免地包含背景内容的轴对齐矩形框。在识别阶段,这些多边形掩码裁剪出的图像将目标文本与周围杂乱内容隔离开来,使通过低秩适应(Low-Rank Adaptation)在多边形裁剪场景文本上微调的 Qwen3-VL-8B-Instruct 模型能够专注于阅读文本,而不受相邻单词或背景噪声的干扰。在不使用任何合成预训练数据的情况下,我们的方法在 CTW1500 上取得了 91.99% 的检测 F 值和 85.86% 的端到端 H-mean,创造了新的最先进水平,并在 Total-Text 和 ICDAR 2015 上也取得了强劲的表现。代码已发布于 https://github.com/doubleblind-afk/MEVL-STP。
cs.CV / 14 / 2609.28860

Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation

面向语言引导医学图像分割的多模态路由与区域精细化方法
Rahman, Md Maklachur, Banna, Md Hasan Al, Anjum, Saraf, Arnob, Assame, Hammond, Tracy
Abstract
Textual descriptions can reduce ambiguity in medical image segmentation by specifying the finding and location to be delineated. Existing text-guided methods mainly improve where image and language features interact but generally retain a single learned update pathway across all image-text pairs. We propose MRSeg, a parameter-efficient framework that uses each image-text pair to route the adaptation of visual and textual features before dense prediction. Frozen ConvNeXt-Tiny and PubMedBERT encoders provide multiscale visual features and clinical text tokens. A joint router uses the deepest visual feature and pooled text to predict a sparse mixture over low-rank adapter bases. The resulting route is shared across separate adapter banks for two visual scales and text, coordinating their adaptation while keeping the feature-specific parameters separate. Region Bridge uses text-derived queries to aggregate dense visual tokens into latent regions, refines these regions through self-attention and text cross-attention, and redistributes the refined information back to the feature maps. Finally, a multiscale decoder combines refined semantic features with shallow image evidence. On QaTa-COV19 and MosMedData+, MRSeg achieves 90.90/83.32 and 81.53/68.82 Dice/mIoU, respectively, with 7.11M trainable parameters and 7.60 GFLOPs. Code: https://github.com/maklachur/MRSeg.
Chinese Translation
文本描述可以通过指定待勾画的病灶及其位置来减少医学图像分割中的歧义。现有的文本引导方法主要改进图像特征与语言特征的交互方式,但通常在所有图像-文本对中保留单一的学习更新路径。我们提出MRSeg,一种参数高效框架,利用每个图像-文本对在稠密预测之前对视觉特征和文本特征的自适应进行路由。冻结的ConvNeXt-Tiny和PubMedBERT编码器提供多尺度视觉特征和临床文本标记。一个联合路由器利用最深层的视觉特征和池化后的文本,在低秩适配器基上预测稀疏混合。所得的路由在分别服务于两个视觉尺度和文本的独立适配器组之间共享,从而协调其自适应过程,同时保持各特征专属参数的独立性。区域桥接模块(Region Bridge)使用由文本导出的查询将稠密的视觉标记聚合为潜在区域,通过自注意力和文本交叉注意力对这些区域进行精细化,并将精炼后的信息重新分配回特征图。最后,多尺度解码器将精炼后的语义特征与浅层图像证据相结合。在QaTa-COV19和MosMedData+数据集上,MRSeg分别取得90.90/83.32和81.53/68.82的Dice/mIoU,可训练参数仅为7.11M,计算量为7.60 GFLOPs。代码:https://github.com/maklachur/MRSeg。
cs.CV / 15 / 2609.28865

Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models

动作表示中的方向-尺度分解:重新思考视觉-语言-动作模型的词元化方式
Duan, Yufei, Yin, Hang, Longhini, Alberta, Tang, Chao, Kragic, Danica
Abstract
Action representation plays a central role in discrete-token vision-language-action (VLA) learning but remains underexamined. Under conventional pose-increment representations, action tokens are sensitive to execution speed and dataset-specific normalization, potentially obscuring geometric structure shared across demonstrations and datasets. We introduce Direction-Scale Decomposition (DSD), an action representation that decomposes translation and rotation increments into direction and scale components before tokenization. DSD isolates motion direction while retaining magnitudes in separate scale channels. We evaluate DSD with uniform binning (BIN) and BEAST, a B-spline-based tokenizer, in simulation and real-world manipulation under both single-dataset and mixed-dataset training. On LIBERO, DSD improves average success rates with both tokenizers. On SimplerEnv, DSD-BIN outperforms BIN by 10.3 percentage points in overall success rate under mixed-dataset training. Real-robot experiments further show gains both with and without robotics pretraining. These results support DSD as an effective action representation for discrete-token VLA models and suggest its potential to mitigate performance degradation when training on large and diverse dataset mixtures. Our project page with additional resources is available at https://vla-dsd.github.io/
Chinese Translation
动作表示在离散词元视觉-语言-动作(VLA)学习中起着核心作用,但至今仍缺乏充分研究。在传统的位姿增量表示下,动作词元对执行速度和数据集特定的归一化方式较为敏感,这可能掩盖了不同演示和数据集之间共享的几何结构。我们提出了方向-尺度分解(Direction-Scale Decomposition, DSD),一种在词元化之前将平移和旋转增量分解为方向和尺度分量的动作表示方法。DSD在保留幅值于独立尺度通道中的同时,分离出运动方向。我们在单数据集和混合数据集训练的仿真与真实世界操作任务中,使用均匀分箱(BIN)和基于B样条的词元化器BEAST对DSD进行了评估。在LIBERO上,DSD在两种词元化器下均提升了平均成功率。在SimplerEnv上,在混合数据集训练下,DSD-BIN的总体成功率比BIN高出10.3个百分点。真机实验进一步表明,无论是否使用机器人预训练,DSD均能带来性能提升。这些结果支持DSD作为离散词元VLA模型的有效动作表示方法,并表明其在缓解大规模多样化数据集混合训练导致的性能退化方面的潜力。包含更多资源的项目页面见 https://vla-dsd.github.io/
cs.CV / 16 / 2609.28923

ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation

ViRDM:驯服表示分布匹配以实现少步因果视频生成
Meng, Zichong, Ge, Chongjian, Huang, Chun-Hao P., Zhou, Yang, Jiang, Huaizu
Abstract
Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training methods predominantly rely on Distribution Matching Distillation (DMD), requiring both a large pretrained teacher and an online critic to estimate distributional discrepancies through diffusion scores. In this work, we ask whether this resource-intensive teacher--critic stack can be eliminated by post-training only the generator against a precomputed target distribution. Drawing inspiration from representation distribution matching (RDM) for one-step image generation, we systematically study its transfer to few-step causal video generation and identify three key barriers: a memory-intractable gradient path, a distinct video optimization regime, and representation distributions that underconstrain temporal dynamics. We introduce ViRDM, a teacher- and critic-free video post-training recipe that addresses these barriers sequentially. By coupling RDM with stochastically truncated clean-exit supervision, a lightweight VAE decoder, and staged vector--Jacobian products, ViRDM makes representation distribution matching memory-feasible for multi-step causal video rollouts. We further establish effective generated-population and initialization regimes for video RDM, and introduce lightweight dynamics regularization to compensate for the underconstrained temporal dynamics. ViRDM turns three-network distillation into generator-only post-training, reducing GPU memory use and training time while improving video quality. With only 20 generator updates, the recipe reaches 84.87 on the official VBench evaluation, outperforming the previous best few-step causal baseline by 0.36, while requiring 16 A100 GPU-hours. We additionally report exploratory results demonstrating the potential of the same recipe for lower causal sampling budget and for one-, two-, and four-step bidirectional generation.
Chinese Translation
少步自回归(AR)视频扩散模型能够实现低延迟的流式生成,但现有的后训练方法主要依赖分布匹配蒸馏(Distribution Matching Distillation, DMD),需要同时使用大型预训练教师模型和在线评判器(critic),通过扩散分数来估计分布差异。在本工作中,我们探讨是否可以通过仅对生成器进行后训练、使其对抗一个预先计算好的目标分布,从而摒弃这种资源密集型的“教师-评判器”架构。受一步图像生成中表示分布匹配(Representation Distribution Matching, RDM)的启发,我们系统地研究了其向少步因果视频生成的迁移,并识别出三个关键障碍:内存上难以处理的梯度路径、截然不同的视频优化机制,以及对时间动态约束不足的表示分布。我们提出了 ViRDM,这是一种无需教师模型和评判器的视频后训练方案,依次解决了上述障碍。通过将 RDM 与随机截断的干净退出(clean-exit)监督、轻量级 VAE 解码器以及分阶段的向量-雅可比积(vector-Jacobian products)相结合,ViRDM 使表示分布匹配在多步因果视频 rollout 中具备内存可行性。我们进一步确立了适用于视频 RDM 的有效生成种群与初始化机制,并引入轻量级的动态正则化以弥补时间动态约束不足的问题。ViRDM 将三网络蒸馏转变为仅生成器的后训练,在降低 GPU 内存占用和训练时间的同时提升了视频质量。仅需 20 次生成器更新,该方案在官方 VBench 评测中达到 84.87 分,超越此前最佳的少步因果基线 0.36 分,且仅需 16 个 A100 GPU 小时。此外,我们还报告了探索性结果,展示了同一方案在更低因果采样预算以及一步、两步和四步双向生成方面的潜力。
cs.CV / 17 / 2609.28930

PlenoCI: Plenoptic CharacterIstics for View Dependence Aware Change Classification

PlenoCI:面向视角依赖感知变化分类的全光特性
Lai, Jason, Galappaththige, Chamuditha Jayanga, Suenderhauf, Niko, Miller, Dimity, Dansereau, Donald G.
Abstract
Radiance field representations such as 3D Gaussian Splatting (3DGS) natively encode complex visual phenomena such as occlusions and view dependence, but they are inherently underconstrained. Independently optimized reconstructions converge to different primitive configurations, even in unchanged regions. We introduce Plenoptic CharacterIstics (PlenoCI), a novel feature built from the plenoptic field these representations approximate. PlenoCI directly captures rich visual behaviors while ignoring Lambertian textures. By deriving closed-form analytic plenoptic derivatives from a 3DGS representation, we efficiently detect these 5D structures. Our approach is robust to underconstrained representations by construction, reporting two orders of magnitude fewer false positives between independent reconstructions of unchanged scenes than concurrent work. We demonstrate PlenoCI's utility on change classification. First, we detect changes with an instance-aware 3DGS pipeline, achieving state-of-the-art results on CL-Splats with a 25.7% mIoU gain over the strongest competitor, while remaining competitive on the more challenging PASLCD benchmark. Leveraging PlenoCI, we classify changes as geometric or appearance-based with a balanced accuracy of 0.735, comparable to the best performing baseline. We believe plenoptic derivatives and PlenoCI open new directions for view dependence aware understanding in visually complex environments. Code and data are available at https://js0n-lai.github.io/plenoci.
Chinese Translation
诸如3D高斯泼溅(3D Gaussian Splatting, 3DGS)等辐射场表示能够原生地编码遮挡和视角依赖等复杂视觉现象,但其本质上存在欠约束问题。即使在没有变化的区域,独立优化的重建结果也会收敛到不同的基元配置。我们提出了Plenoptic CharacterIstics(PlenoCI),这是一种由这些表示所近似逼近的全光场构建的新型特征。PlenoCI能够直接捕捉丰富的视觉行为,同时忽略朗伯(Lambertian)纹理。通过从3DGS表示中推导闭式解析的全光导数,我们可以高效地检测这些5D结构。我们的方法在构建上对欠约束表示具有鲁棒性:在未变化场景的独立重建之间,误报数量比同期工作低了两个数量级。我们展示了PlenoCI在变化分类中的应用。首先,我们利用具有实例感知能力的3DGS流水线进行变化检测,在CL-Splats数据集上取得了最先进的结果,mIoU相比最强竞争对手提升了25.7%,同时在更具挑战性的PASLCD基准上也保持竞争力。借助PlenoCI,我们将变化分类为几何变化或外观变化,平衡准确率达到0.735,与表现最佳的基线方法相当。我们相信,全光导数和PlenoCI为视觉复杂环境中的视角依赖感知理解开辟了新的方向。代码和数据可在 https://js0n-lai.github.io/plenoci 获取。
cs.CV / 18 / 2609.28931

HelloWorld: Towards Practical Applications of Generative Driving World Models

HelloWorld:迈向生成式驾驶世界模型的实际应用
Lu, Fan, Wang, Hanshi, Wang, Zijing, Feng, Quan, Wang, Zhi, Chen, Shijie, Zeng, Xianming, Zhang, Yujian, Wang, Jiazhe, Zha, Xin, Wang, Kai, Zhao, Zhijie, Zhu, Lin, Yang, Tianyi, Xu, Yucheng, Ji, Tao, Zhang, Haodong, Zhang, Zhipeng, Peng, Peixi, Chen, Guang, Liu, Xingliang, Yang, Lei, Xu, Jianyun
Abstract
Driving world models provide a promising route toward scalable counterfactual data generation and interactive simulation beyond recorded driving logs. Realizing this potential requires a system that can generalize across diverse scenes, respond faithfully to prescribed controls, generate coherent multi-sensor observations, and operate efficiently under repeated inference. We present \textbf{HelloWorld}, a 2B driving world model system designed around these requirements. HelloWorld progressively specializes broad visual and motion priors from heterogeneous video data into controllable driving generation using ego pose, HD maps, and 3D boxes. A block-causal generation interface, together with adaptation to self-generated context, aligns the model with sequential simulation. The system further supports synchronized seven-camera RGB generation and conditional LiDAR synthesis, and is distilled toward few-step inference for efficient deployment. Experiments evaluate visual quality, control fidelity, cross-view consistency, robustness under repeated generation, inference efficiency, and LiDAR synthesis. Together, HelloWorld provides a unified framework for scalable driving data generation and interactive simulation.
Chinese Translation
驾驶世界模型为实现可扩展的反事实数据生成以及超越既有驾驶日志记录的交互式仿真提供了一条有前景的路径。要实现这一潜力,需要系统能够泛化到多样化场景、忠实响应给定的控制指令、生成连贯的多传感器观测,并在反复推理下高效运行。我们提出了HelloWorld,一个围绕这些需求设计的20亿参数(2B)驾驶世界模型系统。HelloWorld利用自车位姿(ego pose)、高精地图(HD maps)和3D边界框,将从异构视频数据中获得的广泛视觉与运动先验逐步专门化为可控的驾驶内容生成。通过分块因果(block-causal)生成接口以及对自生成上下文的适配,该模型与序列化仿真实现了对齐。该系统还支持七相机同步RGB生成和条件式LiDAR合成,并通过蒸馏实现少步推理以支持高效部署。实验从视觉质量、控制保真度、跨视角一致性、重复生成下的鲁棒性、推理效率以及LiDAR合成等方面进行了评估。总体而言,HelloWorld为可扩展的驾驶数据生成与交互式仿真提供了一个统一框架。
cs.CV / 19 / 2609.28949

Exploiting Target Knowledge from MLLMs for Robust Few-Shot Segmentation

利用多模态大语言模型中的目标知识实现鲁棒的小样本分割
Hu, Yijun, Fan, Heng, Zhang, Libo
Abstract
Few-shot segmentation (FSS) aims to segment unseen object categories with a few (e.g., one or five) labeled examples, enabling efficient adaptation to novel classes. Conventional models typically rely on appearance-based visual matching between support and query images for segmentation. While straightforward, these methods often struggle to handle significant appearance discrepancies and occlusions in the query image due to insufficient target knowledge. To mitigate this, we introduce a novel framework that mines target knowledge using the strong reasoning capacity of Multimodal Large Language Models (MLLMs) and employs it to enhance FSS. Specifically, building on SAM 2, our method, named MK-FSS, exploits two forms of complementary knowledge derived from a query image by an MLLM for FSS, including spatial knowledge, which provides a spatial prior indicating the potential target location, and semantic knowledge, which describes the target using text. The spatial knowledge is first encoded into a memory representation, and then resulting memory is integrated with the support-guided memory feature from query image through a carefully designed dual-memory debate-fusion (DMDF) module, yielding a more robust target memory feature. In parallel, the semantic knowledge is encoded into the textual feature, which is fused with multi-scale query features via a progressive cross-modal prompt generator (PCPG), producing a target-aware multimodal prompt for segmentation. Working together, the dual-memory feature and the multimodal prompt provide a comprehensive representation of the target, enabling more robust segmentation. In our extensive experiments, MK-FSS shows promising results and largely surpasses existing methods. Code will be released.
Chinese Translation
小样本分割(Few-shot segmentation, FSS)旨在利用少量(例如一个或五个)带标注的样本对未见过的目标类别进行分割,从而实现对新类别的高效适应。传统模型通常依赖支撑图像与查询图像之间基于外观的视觉匹配来进行分割。尽管方法直接,但由于目标知识不足,这些方法往往难以处理查询图像中显著的外观差异和遮挡问题。为缓解这一问题,我们提出了一种新颖的框架,利用多模态大语言模型(Multimodal Large Language Models, MLLMs)强大的推理能力挖掘目标知识,并将其用于增强小样本分割。具体而言,基于 SAM 2,我们的方法(命名为 MK-FSS)利用 MLLM 从查询图像中提取两种互补的知识形式:一是空间知识,提供指示潜在目标位置的空间先验;二是语义知识,通过文本对目标进行描述。空间知识首先被编码为记忆表示,随后通过精心设计的双记忆辩论融合模块(dual-memory debate-fusion, DMDF),将所得记忆与查询图像中由支撑图像引导的记忆特征相融合,从而得到更鲁棒的目标记忆特征。与此同时,语义知识被编码为文本特征,并通过渐进式跨模态提示生成器(progressive cross-modal prompt generator, PCPG)与多尺度查询特征融合,生成具备目标感知能力的多模态提示用于分割。通过双记忆特征与多模态提示的协同工作,可以提供对目标的全面表示,从而实现更鲁棒的分割。在大量实验中,MK-FSS 展现出优异的结果,大幅超越了现有方法。代码将开源发布。
cs.CV / 20 / 2609.28956

MoVISA: Multi-Token Reasoning for Video Object Segmentation

MoVISA:面向视频目标分割的多Token推理方法
Zhao, Ruining, Cheng, Ho Kei, Schwing, Alexander G
Abstract
Recent advances in video object segmentation with Multimodal Large Language Model (MLLM) reasoning have demonstrated the effectiveness of using a single textual token, such as SEG, to predict segmentation masks across images and videos. However, we observe that this single-token strategy lacks the granularity required to precisely localize multiple objects across time in video segmentation tasks. To address this limitation, we develop Multi-Token Reasoning for Video Object Segmentation, or MoVISA. MoVISA uses multiple segmentation tokens, such as SEG0 and SEG1, to represent an object across different frames. This design enables more fine-grained alignment between language prompts and spatio-temporal mask predictions, improving both performance and interpretability. On the challenging MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS benchmarks, our model achieves a 13.2 percent J and F improvement on MeViS and an 8.4 percent J and F improvement on ReVOS. Code and models will be released.
Chinese Translation
近期基于多模态大语言模型(MLLM)推理的视频目标分割研究进展表明,使用单一文本Token(如SEG)来预测图像和视频中的分割掩码是有效的。然而,我们观察到这种单Token策略缺乏在视频分割任务中跨时间精确定位多个目标所需的细粒度。为解决这一局限,我们提出了面向视频目标分割的多Token推理方法(Multi-Token Reasoning for Video Object Segmentation),即MoVISA。MoVISA使用多个分割Token(如SEG0和SEG1)来表示不同帧中的目标。这一设计实现了语言提示与时空掩码预测之间更细粒度的对齐,同时提升了性能与可解释性。在具有挑战性的MeViS、DAVIS17、ReVOS和Ref-Youtube-VOS基准数据集上,我们的模型在MeViS上取得了13.2%的J和F指标提升,在ReVOS上取得了8.4%的J和F指标提升。代码和模型将会开源发布。
cs.CV / 21 / 2609.28967

Passive LWIR Hyperspectral Ranging via Transmittance Extraction and Distance Alignment

基于透射率提取与距离对齐的被动长波红外高光谱测距
Chen, Zhihe, Fan, Chen, Liu, Shuo, Huang, Xiaolin, He, Yunze, He, Xiaofeng, Zhang, Lilian
Abstract
Passive long-wave infrared (LWIR) hyperspectral ranging enables distance estimation in low-light and nighttime scenes by exploiting atmospheric absorption features in thermal radiance received through the atmosphere.Joint estimation of temperature, emissivity, and distance is computationally expensive. Reference-range joint inversion also uses a distance-invariant effective attenuation coefficient, which can bias range estimates.We introduce transmittance extraction and distance alignment (TEDA), which decouples range estimation from temperature--emissivity inversion. In the first stage, a baseline estimator with a data-fidelity term invariant to the known absorption direction yields two closed-form smoothing branches for the slowly varying thermal continuum. An observation-derived gate combines the branches, and subtracting the blended baseline in the log domain recovers atmospheric transmittance. The second stage estimates range by matching the recovered transmittance to sensor-domain transmittance models recomputed for each candidate distance. Monte Carlo simulations show that TEDA effectively reduces the ranging bias caused by the distance-invariant attenuation coefficient approximation. In a measured scene, TEDA's mean range estimates are closer to the LiDAR medians than those of reference-range joint inversion in both evaluated patches. TEDA processes a complete $256\times256$ region of interest in 8.19~s versus 159.47~s for reference-range joint inversion, an approximately 20-fold speedup.
Chinese Translation
被动长波红外(LWIR)高光谱测距通过利用大气热辐射中的大气吸收特征,能够在低光照和夜间场景下进行距离估计。对温度、发射率和距离进行联合估计的计算代价高昂;而参考距离联合反演方法采用距离无关的有效衰减系数,这可能导致距离估计偏差。本文提出透射率提取与距离对齐方法(TEDA),将距离估计与温度-发射率反演解耦。在第一阶段,基线估计器采用对已知吸收方向保持不变的数据保真项,为缓变的热辐射连续谱生成两个闭式平滑分支;由观测数据导出的门控机制对两个分支进行融合,并在对数域中减去融合基线,从而恢复大气透射率。第二阶段将恢复的透射率与针对每个候选距离重新计算的传感器域透射率模型进行匹配,以估计距离。蒙特卡洛仿真表明,TEDA 有效降低了由距离无关衰减系数近似引起的测距偏差。在实测场景中,在两个评估区块内,TEDA 的平均距离估计均比参考距离联合反演方法更接近激光雷达(LiDAR)中值。TEDA 处理完整的 256×256 感兴趣区域仅需 8.19 秒,而参考距离联合反演方法需要 159.47 秒,实现了约 20 倍的加速。
cs.CV / 22 / 2609.28991

Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models

分数之下:重新思考视频理解模型的幻觉评估
Gong, Shuzhi, Sun, Fengze, Liu, Yuansan
Abstract
Video understanding is increasingly performed by multi-stage LLM agents that separate temporal grounding, visual observation, and reasoning. Yet these stages are typically evaluated on different benchmarks and distributions, making it difficult to determine where hallucinations originate. We first organize existing benchmarks around these stages and show that their scores provide inconsistent diagnostic signals: stronger stage-level performance does not reliably imply lower downstream hallucination, and even benchmarks targeting the same capability can disagree. We therefore introduce a causal stage-intervention protocol that overwrites individual stages while holding the downstream task fixed. Across 60,008 runs on three video-agent architectures, we find that grounding is the dominant source of downstream error, with roughly four times the causal impact of corrupting visual observations. Successful grounding depends primarily on locating the correct region rather than precise temporal overlap, explaining why standard mIoU metrics poorly predict downstream reliability. We further find that incorrect evidence is substantially more harmful than missing evidence. Finally, auditing existing benchmarks against these interventions reveals that their scores do not reliably predict causal cascade sensitivity and can fail under distribution shift. These results motivate intervention-based, stage-aware evaluation for trustworthy video agents.
Chinese Translation
视频理解越来越多地由多阶段LLM智能体执行,这些智能体将时序定位、视觉观察和推理分离开来。然而,这些阶段通常在不同的基准和数据分布上进行评估,使得难以确定幻觉的来源。我们首先围绕这些阶段对现有基准进行梳理,并表明其分数提供了不一致的诊断信号:更强的阶段级性能并不能可靠地意味着更低的下游幻觉,甚至针对同一能力的基准之间也可能存在分歧。因此,我们提出了一种因果式阶段干预协议,在保持下游任务不变的情况下覆写各个阶段。在三种视频智能体架构上共60,008次运行中,我们发现定位是下游错误的主要来源,其因果影响约为破坏视觉观察的四倍。成功的定位主要依赖于找到正确的区域,而非精确的时序重叠,这解释了为什么标准的mIoU指标难以很好地预测下游可靠性。我们进一步发现,错误的证据比缺失的证据危害大得多。最后,根据这些干预对现有基准进行审计后发现,其分数并不能可靠地预测因果级联的敏感性,并且在分布偏移下可能失效。这些结果促使我们为可信的视频智能体采用基于干预的、阶段感知的评估方法。
cs.CV / 23 / 2609.28997

Only What Was Seen: Observation-Gram Compaction of View-Dependent Appearance in 3D Gaussian Splatting

唯见所观:3D高斯泼溅中视角相关外观的观测-格拉姆压缩
Pietroszek, Krzysztof
Abstract
Most of the memory of a 3D Gaussian Splatting model holds spherical-harmonic colour coefficients, yet each Gaussian is seen only from the narrow cone of directions of the training cameras. We turn this into a distortion metric that other compressors can adopt: a per-Gaussian observation Gram matrix, accumulated from viewing directions and blending weights, is the exact first-order map from coefficient changes to squared image error and needs only the model and the camera poses. Under it, degree reduction becomes a closed-form projection that generalises truncation, degree allocation a Lagrangian rate-distortion problem, and vector quantisation the matrix-weighted Lloyd algorithm, of which Compressed3D's quantiser is the scalar case. Swapped into Compressed3D with everything else unchanged, the metric raises PSNR by +0.49 dB before fine-tuning, with SSIM and LPIPS following, and at matched rate still gains +0.32 dB without a single training image. A training-free stack built on the metric alone is 15% smaller than the image-free GSICO at equal quality on Mip-NeRF 360.
Chinese Translation
3D高斯泼溅(3D Gaussian Splatting)模型的大部分内存被球谐颜色系数占据,然而每个高斯仅在训练相机较窄的方向锥范围内被观察到。我们据此构建了一个其他压缩器也可采用的失真度量:由观察方向和混合权重累积得到的逐高斯观测Gram矩阵(observation Gram matrix),它是系数变化到平方图像误差的精确一阶映射,且只需要模型和相机位姿。在该度量下,阶数缩减成为一个推广了截断法的闭式投影,阶数分配成为一个拉格朗日率失真问题,而矢量量化则成为矩阵加权Lloyd算法——Compressed3D的量化器正是其标量特例。在保持其他设置不变的情况下将该度量替换进Compressed3D,微调前PSNR即提升+0.49 dB,SSIM和LPIPS也随之改善;且在匹配码率下,即使不使用任何训练图像仍可获得+0.32 dB的增益。仅基于该度量构建的免训练方案在Mip-NeRF 360数据集上达到相同质量时,比不使用图像的GSICO还要小15%。
cs.CV / 24 / 2609.29006

FluidRain: Incompressible Rain Flow as an Attention Bias for Loop-in-Loop Video Deraining

FluidRain:以不可压缩雨流作为注意力偏置的套环式视频去雨方法
Wang, Pu, Wang, Yongcong, Li, Wenhao, Chen, Xiang, Gao, Guangwei, Pan, Jinshan, Yao, Siyuan, Fu, Shujun, Zheng, Zhuoran
Abstract
Existing video deraining methods typically exploit neighboring frames through either explicit alignment or implicit spatiotemporal aggregation. Explicit alignment relies on accurate motion estimation, which can become unreliable under dense rain, while implicit aggregation avoids alignment but lacks explicit guidance on the directional and temporally coherent structure of rain. This leaves a gap between reliable temporal aggregation and explicit modeling of rain motion. To address these limitations, we propose FluidRain, a lightweight video derainer that uses divergence-free rain flow to guide Loop-in-Loop attention across scales and neighboring frames. Motivated by fluid mechanics, we model rain motion as a divergence-free image-space flow and use it to organize multi-scale and temporal aggregation. Specifically, FluidRain first estimates a rain-flow field for each frame and projects it onto the divergence-free subspace. The resulting flow steers window attention along rain streaks, enabling neighboring frames to be aggregated without explicit alignment. Since rain-flow structure is preserved across scales and nearby frames, Loop-in-Loop reuses the same attention operator across both dimensions, resulting in a three-frame model with only 0.80M parameters. Experiments on four benchmarks show that FluidRain remains competitive with substantially larger restoration models. We further examine how temporal evidence scales with different input views. To evaluate whether the model remains reliable when rain motion changes across frames, we introduce RainSyn-Gust, which injects controlled changes in rain-streak direction into existing benchmarks. We also develop a physics-based no-reference metric that evaluates real-rain removal without requiring clean targets.
Chinese Translation
现有的视频去雨方法通常通过显式对齐或隐式时空聚合来利用相邻帧信息。显式对齐依赖于精确的运动估计,而在密集降雨条件下可能变得不可靠;隐式聚合虽然避免了对齐,但缺乏对雨的方向性和时间一致结构的显式引导。这使得可靠的时间聚合与雨运动的显式建模之间存在空白。为了解决这些局限性,我们提出了FluidRain,一种轻量级视频去雨器,它利用无散度的雨流(rain flow)来引导跨尺度和相邻帧的Loop-in-Loop注意力。受流体力学启发,我们将雨的运动建模为图像空间中的无散度流场,并用它来组织多尺度与时间上的聚合。具体而言,FluidRain首先为每一帧估计一个雨流场,并将其投影到无散度子空间。所得的流场引导窗口注意力沿雨纹方向进行,从而无需显式对齐即可聚合相邻帧。由于雨流结构在不同尺度和相邻帧之间得以保持,Loop-in-Loop在两个维度上复用同一注意力算子,最终得到仅含0.80M参数的三帧模型。在四个基准数据集上的实验表明,FluidRain与规模大得多的复原模型相比仍具竞争力。我们进一步研究了时间证据随不同输入视角数量的扩展规律。为评估模型在雨运动跨帧变化时是否依然可靠,我们提出了RainSyn-Gust,该基准在现有数据集中注入受控的雨纹方向变化。此外,我们还开发了一种基于物理的无参考指标,可在无需干净目标的情况下评估真实雨水的去除效果。
cs.CV / 25 / 2609.29028

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

RGBD20K:一个面向RGB-D语义分割的大规模基准数据集
Dong, Shaohua, Meng, Zexuan, Sun, Haiyan, Fan, Bing, Zhang, Cuicui, Joseph, Dylan, Sha, Kewei, Feng, Yunhe, Fan, Heng
Abstract
In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it covers 160 fine-grained categories, largely surpassing the category diversity of existing popular RGB-D benchmarks (e.g., NYUv2 with 40 classes and SUN RGB-D with 37 classes). With such enriched semantic coverage, we expect to promote the learning of more generalizable segmentation models. (2) Larger Scale. Compared with current benchmarks, RGBD20K offers 20,000 RGB-D image pairs, providing a substantially larger training resource that benefits the development of more powerful deep models. (3) High-Fidelity Annotation. We perform rigorous re-evaluation and correction of existing labels to resolve long-standing annotation noise, resulting in a clean and reliable ground-truth foundation. Furthermore, we propose a novel score-purified fusion (SPF) method, which achieves state-of-the-art performance across all evaluated benchmarks, demonstrating the effectiveness of our approach in leveraging high-quality multimodal information for RGB-D semantic segmentation. The dataset is here: https://github.com/ShaohuaDong2021/RGBD20K/.
Chinese Translation
本文提出了RGBD20K,一个新颖的数据集,旨在通过涵盖丰富的类别和高质量的标注,促进更鲁棒、更通用的RGB-D语义分割方法的发展。RGBD20K具有以下几个吸引人的特性:(1)扩展的语义空间。具体而言,该数据集涵盖160个细粒度类别,大幅超越了现有流行的RGB-D基准数据集的类别多样性(例如,NYUv2包含40个类别,SUN RGB-D包含37个类别)。凭借如此丰富的语义覆盖范围,我们期望能够推动更具泛化能力的分割模型的学习。(2)更大规模。与当前基准数据集相比,RGBD20K提供了20,000对RGB-D图像,提供了规模大得多的训练资源,有利于开发更强大的深度模型。(3)高保真度标注。我们对现有标签进行了严格的重新评估和修正,以解决长期存在的标注噪声问题,从而构建了干净可靠的真实标注(ground-truth)基础。此外,我们提出了一种新颖的分数净化融合(score-purified fusion, SPF)方法,在所有评估的基准数据集上均取得了最先进的性能,证明了我们的方法在利用高质量多模态信息进行RGB-D语义分割方面的有效性。数据集地址:https://github.com/ShaohuaDong2021/RGBD20K/。
cs.CV / 26 / 2609.29029

Exploiting answer-invariant redundancies in satellite imagery for efficient VLM inference on edge

利用卫星影像中答案不变冗余实现边缘端高效的视觉语言模型推理
Janveja, Ishani, Zhang, Davis, Oh, Seoyul, Vasisht, Deepak
Abstract
Onboard vision-language models could enable satellites to answer queries directly, but exhaustive tiled inference over high-resolution imagery is slow and energy-intensive. We identify answer-invariant token redundancy (AITR): image tiles and vision tokens that can be removed without changing the final answer. We present Rift, a two-stage system that performs query-conditioned tile pruning followed by elastic prefill to reduce token budget. We evaluate it on LLaVA-1.5 7B running on Jetson AGX Orin. Compared with exhaustive tiled inference, Rift reduces energy by 78% and latency by 69%, while increasing accuracy from 45% to 73%.
Chinese Translation
星载视觉语言模型(VLM)可使卫星直接回答查询请求,但对高分辨率影像进行穷举式分块推理速度慢且能耗高。我们识别出一种答案不变Token冗余(AITR):即在不改变最终答案的情况下可被移除的图像分块和视觉Token。我们提出了Rift,一个两阶段系统,首先执行基于查询条件的分块剪枝,然后进行弹性预填充(elastic prefill)以降低Token预算。我们在Jetson AGX Orin平台上运行的LLaVA-1.5 7B模型上对该系统进行了评估。与穷举式分块推理相比,Rift将能耗降低了78%,延迟降低了69%,同时将准确率从45%提升至73%。
cs.CV / 27 / 2609.29048

Where Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models

幻觉的栖身之所:VQ分词视觉-语言模型中的跨架构电路
Hegde, Shamanthak, Liu, Xiangrui, Patel, Maitreya, Yang, Yezhou
Abstract
Unified vision-language models (VLMs) that tokenize images through a vector-quantized (VQ) codebook routinely hallucinate objects on grounded yes/no benchmarks, yet existing decoding-time fixes treat this as generic miscalibration without an architectural account. Using activation patching across twenty-five models spanning eight LLM families, we identify an early-layer ($L_0$) attention routing circuit shared across VQ-tokenized VLMs and propose a three-gate diagnostic that distinguishes the models carrying it from those that do not. The diagnostic isolates ten positive models (five natural unified-VQ VLMs across three LLM families and five induced variants) and rejects the remaining fifteen. A single-variable architectural swap (LLaVA-1.6 CLIP+MLP $\rightarrow$ VQ+Linear) installs the circuit, while a matched-compute MLP control on identical data does not, isolating vector quantization as the source of the pathological signal; the routing pathway that carries it is one that the backbone already provides. Against tuned VCD and DoLA baselines, tuned DoLA wins on binary calibration, but \textbf{only $L_0$ ablation reduces object hallucination in open-ended generation} (CHAIR$_i$ reduces by $31\,\%$ relatively, whereas tuned DoLA and VCD leave it unchanged or worsen it). These results recast object hallucination in unified VQ VLMs as a property of architecture and pretraining, and yield a targeted intervention that mechanism-agnostic decoding cannot replicate.
Chinese Translation
通过向量量化(VQ)码本对图像进行分词的统一视觉-语言模型(VLM)在基于事实的是/否基准测试中经常产生对象幻觉,然而现有的解码时修复方法仅将其视为一般的校准失准问题,缺乏架构层面的解释。我们在涵盖八个LLM家族的二十五个模型上进行激活修补(activation patching)实验,识别出一个在VQ分词VLM中共享的早期层($L_0$)注意力路由电路,并提出一个三重门槛诊断方法,以区分携带该电路的模型与不携带该电路的模型。该诊断方法分离出十个阳性模型(包括三个LLM家族中的五个天然统一VQ VLM和五个诱导变体),并排除了其余十五个模型。单变量架构替换实验(LLaVA-1.6的CLIP+MLP替换为VQ+Linear)能够安装该电路,而在相同数据上计算量相当的MLP对照实验则不能,从而将向量量化确定为病态信号的来源;承载该信号的路由通路是骨干网络本身已具备的。与经过调优的VCD和DoLA基线相比,调优后的DoLA在二元校准上表现更优,但**只有$L_0$消融能够减少开放式生成中的对象幻觉**(CHAIR$_i$相对降低31%,而调优后的DoLA和VCD要么使其保持不变,要么使其恶化)。这些结果将统一VQ VLM中的对象幻觉重新解释为一种架构与预训练的固有属性,并提供了一种机制无关的解码方法无法实现的有针对性干预手段。
cs.CV / 28 / 2609.29064

EIB-Net: Entropy-Guided Information Bottleneck for Generalizable AI-Generated Image Detection

EIB-Net:面向可泛化AI生成图像检测的熵引导信息瓶颈网络
Zhang, Zhida, Ma, Xinlei, Cao, Jie
Abstract
The proliferation of photorealistic AI-generated images demands robust detection methods that generalize across diverse generative models. While existing approaches target manipulation-based forgeries with local artifacts, generation-based images (e.g., from diffusion models) lack such traces, posing a fundamental challenge. We observe that generative models prioritize global semantics at the expense of local texture fidelity, making low-texture regions key indicators of synthetic origin. To exploit this, we propose EIB-Net, an Entropy-guided Information Bottleneck Network. EIB-Net introduces a novel Image Entropy (IE) metric to automatically select the most informative (lowest-entropy) patch, then processes it with a Variational Information Bottleneck (VIB) to learn compact, generalizable features. Extensive experiments on DIFF, DiffusionForensics, and GenImage benchmarks demonstrate state-of-the-art performance: EIB-Net achieves 85.7\% accuracy using only 2\% of training data, outperforming full-image baselines by over 15\%, and maintains robust cross-generator generalization (83.5\% average accuracy on GenImage). Furthermore, our entropy-guided patch selection (EGPL) consistently enhances diverse backbones (CNNs and Transformers), proving its practical value for data-efficient detection.
Chinese Translation
照片级逼真AI生成图像的激增,要求检测方法能够在多样化的生成模型之间具备良好的泛化能力。现有方法主要针对具有局部伪影的基于篡改的伪造图像,而基于生成的图像(如来自扩散模型的图像)缺乏此类痕迹,这构成了根本性的挑战。我们观察到,生成模型在生成图像时优先保证全局语义而牺牲局部纹理保真度,使得低纹理区域成为判断合成来源的关键指标。为利用这一特性,我们提出了EIB-Net(Entropy-guided Information Bottleneck Network,熵引导信息瓶颈网络)。EIB-Net引入了一种新颖的图像熵(Image Entropy, IE)度量,用于自动选择信息量最大(熵最低)的图像块,然后通过变分信息瓶颈(Variational Information Bottleneck, VIB)对其进行处理,以学习紧凑且可泛化的特征。在DIFF、DiffusionForensics和GenImage基准上的大量实验表明,该方法取得了最先进的性能:EIB-Net仅使用2%的训练数据即可达到85.7%的准确率,比全图像基线高出15%以上,并在跨生成器泛化方面保持稳健(在GenImage上平均准确率为83.5%)。此外,我们的熵引导图像块选择方法(Entropy-Guided Patch Selection, EGPL)能够持续提升多种骨干网络(CNN和Transformer)的性能,证明了其在数据高效检测方面的实用价值。
cs.CV / 29 / 2609.29073

Seeing Is Not Measuring: Tool-Augmented Metric Spatial Reasoning for Vision-Language Models

看见不等于度量:面向视觉语言模型的工具增强度量空间推理
Glantz, Kai, Grange, Clemens
Abstract
Vision-Language Models (VLMs) describe scenes well but reason poorly about metric 3D structure such as absolute distances, physical sizes, or egocentric directions. We present a modular, predictor agnostic, tool-augmented framework that equips a small VLM (Qwen3.5-4B) with geometric tools: 3D object detection, metric depth estimation, and deterministic solvers for distance, size and bearing. Each object is detected in the camera frame of its own best view, and the tools use that frame's pose to lift every detection into one shared world frame. Moving metric computation out of the model's weights and into explicit solvers yields large gains on three of four ReVSI-Bench tasks: with a strong monocular detector (WildDet3D), absolute distance rises from 0.46 to 0.74 Mean Relative Accuracy (MRA), relative distance from 39.1% to 67.4%, and relative direction from a below-chance 25.9% to 73.4%. Because any detector can be swapped in behind the tool interface, comparing real detectors against ground-truth boxes separates perception error from reasoning error: orchestration costs only 0.03 MRA. Object size is bounded by the detector: the tools are near-exact on groundtruth boxes (0.97) yet the best real detector barely beats the no-tool baseline (0.61 vs. 0.58), because size reads straight off a box extent monocular detectors get wrong. Without a predefined recipe, the model already sequences the tools correctly on its own, matching a scripted pipeline on three of four tasks.
Chinese Translation
视觉语言模型(VLM)能够很好地描述场景,但在推理诸如绝对距离、物理尺寸或自我中心方向等度量三维结构方面表现不佳。我们提出了一个模块化、与预测器无关、工具增强的框架,为小型视觉语言模型(Qwen3.5-4B)配备了几何工具:三维物体检测、度量深度估计,以及用于距离、尺寸和方位的确定性求解器。每个物体在其自身最佳视角的相机坐标系中被检测,工具利用该坐标系的位姿将所有检测结果提升到一个共享的世界坐标系中。将度量计算从模型权重中移出、交由显式求解器完成,在 ReVSI-Bench 四项任务中的三项上带来了显著提升:使用强大的单目检测器(WildDet3D)时,绝对距离的平均相对准确率(MRA)从 0.46 提升至 0.74,相对距离从 39.1% 提升至 67.4%,相对方向从低于随机水平的 25.9% 提升至 73.4%。由于任何检测器都可以在工具接口后进行替换,将真实检测器与真值边界框进行对比可以将感知误差与推理误差分离开来:编排开销仅为 0.03 MRA。物体尺寸则受限于检测器:在真值边界框上工具几乎完全准确(0.97),而最好的真实检测器仅勉强超过无工具基线(0.61 对 0.58),这是因为尺寸是直接从边界框的范围读取的,而单目检测器在这一环节容易出错。在没有预设流程的情况下,模型已能自行正确地对工具进行排序,在四项任务中的三项上与脚本化流水线表现相当。
cs.CV / 30 / 2609.29106

WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model

WildHSR:基于3D基础模型的度量尺度前馈4D人体-场景重建
Bright, Jerrin, Zelek, John
Abstract
3D foundation models recover video cameras and geometry in one forward pass, but some of the strongest are up to scale. Joint people-scene reconstruction then requires two missing outputs: metric scale and persistent person identity. We ask whether one up-to-scale foundation representation can support both through lightweight adaptation. Exact metric labels are scarce, but unlabeled in-the-wild video is abundant. We use people in curated web video to initialise the solution: a posed metric body and 2D keypoints give an approximate, closed-form scale pseudo-label. These pseudo-labels pretrain a Scale Readout, which is then fine-tuned together with a lightweight adapter using exact metric supervision from standard real-video training splits. At inference the head predicts metric scale from foundation-model tokens, without the ruler or its teachers. For person identity, we probe the pretrained foundation model alone and find evidence that its intermediate query-key features encode person correspondence across frames. In most evaluated moving-person clips, a mid-layer token prefers that person over the vacated location and other people. A tiny projection reads this correspondence; together with metric pelvis motion and proposal confidence, it drives dustbin-aware Sinkhorn association of per-frame bodies. WildHSR combines both readouts to reconstruct metric cameras, scene and people from monocular video. Each window is predicted feed-forward; analytic association and Sim(3) composition connect windows. On EMDB-2, WildHSR is the first feed-forward method in the published comparison to beat the best optimization-based WA-MPJPE and RTE while leading feed-forward methods on all three world-frame metrics. On RICH, it leads feed-forward people-and-scene methods on WA-MPJPE and W-MPJPE. The complete pipeline runs at 10.1 fps on one GPU.
Chinese Translation
3D基础模型能够在一次前向传播中恢复视频相机参数和几何信息,但其中一些最强的模型仅恢复至尺度(up-to-scale)的结果。要实现人体与场景的联合重建,还需要两个缺失的输出:度量尺度和持久的人体身份。我们探究一个仅至尺度的基础模型表示能否通过轻量级适配同时支持这两个任务。精确的度量尺度标注十分稀缺,但无标注的真实场景(in-the-wild)视频却非常丰富。我们利用精选网络视频中的人物来初始化解决方案:具有姿态的度量人体模型和2D关键点可以给出近似的、闭式解的尺度伪标签。这些伪标签用于预训练一个尺度读出头(Scale Readout),随后再与一个轻量级适配器一起,使用来自标准真实视频训练集的精确度量监督进行微调。在推理阶段,该读出头直接从基础模型的token预测度量尺度,无需标尺或其教师模型。对于人体身份,我们仅对预训练的基础模型进行探查,发现其中间层的query-key特征编码了跨帧的人体对应关系的证据。在大多数评估的移动人物片段中,某一中间层token更倾向于该人物而非其离开后的空位或其他人物。一个微小的投影层可以读取这种对应关系;结合度量尺度的骨盆运动和候选框置信度,它驱动具有垃圾箱(dustbin)感知的Sinkhorn算法来关联逐帧的人体。WildHSR结合这两个读出结果,从单目视频中重建度量尺度的相机、场景和人体。每个时间窗口均以前馈方式预测;通过解析式的关联和Sim(3)变换组合连接各窗口。在EMDB-2数据集上,WildHSR是已发表对比中首个在最优的基于优化方法的WA-MPJPE和RTE指标上取得超越的前馈方法,同时在所有三项世界坐标系指标上领先于其他前馈方法。在RICH数据集上,它在WA-MPJPE和W-MPJPE指标上领先于前馈的人体-场景重建方法。完整流水线在单个GPU上以10.1 fps的速度运行。
cs.CV / 31 / 2609.29116

Spectral Amplitude Purification in Distribution Matching for Diffusion Distillation

面向扩散蒸馏的分布匹配中的频谱幅度净化方法
Zhou, Zhenyu, Wang, Can, Chen, Chun, Zheng, Zeyu, Chen, Defang
Abstract
Distribution Matching Distillation (DMD) enables high-quality diffusion sampling in only a few steps, but its optimization dynamics remain dominated by coarse, low-frequency signals, delaying the recovery of fine-grained details. We identify a pronounced concentration of spectral amplitudes at low frequencies in the DMD directional error, where dominant low-frequency components overwhelm weaker mid- and high-frequency signals. To address this issue, we propose Spectral Amplitude Purification for Distribution Matching Distillation (SAP-DMD), a plug-and-play approach that adaptively modulates the amplitude spectrum of the DMD directional field. By suppressing the dominant tail of the amplitude spectrum, SAP-DMD reduces low-frequency dominance and promotes more effective recovery of fine structures and textures. Experiments on PixArt-$\alpha$, SD3, and SD3.5 demonstrate that SAP-DMD accelerates training convergence and improves generation quality under both 2-step and 4-step sampling.
Chinese Translation
分布匹配蒸馏(Distribution Matching Distillation, DMD)能够仅用少数几步实现高质量的扩散采样,但其优化动态仍被粗糙的低频信号所主导,从而延迟了细粒度细节的恢复。我们发现,DMD方向误差中的频谱幅度明显集中于低频区域,主导性的低频分量淹没了较弱的中频和高频信号。为解决这一问题,我们提出了面向分布匹配蒸馏的频谱幅度净化方法(Spectral Amplitude Purification for Distribution Matching Distillation, SAP-DMD),这是一种即插即用的方法,可自适应地调节DMD方向场的幅度谱。通过抑制幅度谱中的主导尾部,SAP-DMD降低了低频主导性,促进了对精细结构和纹理的更有效恢复。在PixArt-$\alpha$、SD3和SD3.5上的实验表明,SAP-DMD加速了训练收敛,并在2步和4步采样下均提升了生成质量。
cs.CV / 32 / 2609.29118

UpDown-SC: Gravity-Canonicalized Dual-Envelope Scan Context for Indoor LiDAR Place Recognition

UpDown-SC:用于室内激光雷达位置识别的重力规范化双包络扫描上下文
Xu, Jie, Yang, Yongxin, Jin, Ziyi, Yu, Kangjin, Huang, Hongjun, Han, Chao, Xia, Zhongpu
Abstract
LiDAR place recognition is a key front end for loop closure and global relocalization, yet indoor retrieval remains difficult when attitude or sensor mounting height changes between mapping and query sessions. Scan Context stores the maximum height in each polar cell; indoors, broad ceilings can suppress the lower and mid-level geometry that distinguishes adjacent rooms and corridors. We present UpDown-SC, a training-free polar descriptor that first canonicalizes gravity and then represents two complementary surfaces: the upper envelope of lower/middle structures and the lower envelope of overhead structures. Their physical split is estimated once from a cell-balanced map height distribution and reused by every query. A mask-aware, non-uniform two-channel distance retains discriminative lower-level evidence while limiting sensitivity to its cross-session variation, without treating unobserved cells as zero-height measurements. Conventional Scan Context shortlisting and circular yaw alignment are retained, so retrieved hypotheses directly initialize geometric verification. Experiments across repeated indoor sessions, mounting-height changes, mixed outdoor-to-indoor trajectories, and an outdoor transfer sequence show more reliable first-choice retrieval on the indoor and mounting-height-varied sessions. A paired test finds a significant gain over Scan Context on the in-house sessions. UpDown-SC also gives the best or second-best F1max and AUPR under threshold-based acceptance while retaining a lightweight CPU front end. Continuous replay confirms that the retrieved hypotheses support metric prior-map localization. Code and evaluation artifacts: https://github.com/jiejie567/updown-sc.
Chinese Translation
激光雷达位置识别是回环检测与全局重定位的关键前端,然而当建图与查询阶段之间姿态或传感器安装高度发生变化时,室内检索仍然十分困难。Scan Context 在每个极坐标单元中存储最大高度;在室内场景中,宽阔的天花板会抑制能够区分相邻房间与走廊的中低层几何信息。我们提出 UpDown-SC,这是一种免训练的极坐标描述子,首先对重力方向进行规范化,然后表示两个互补的表面:低层/中层结构的上包络和顶部结构的下包络。两者的物理划分基于单元均衡的地图高度分布一次性估计得到,并被每个查询复用。一种掩码感知的非均匀双通道距离在保留低层判别性信息的同时,限制了对其跨阶段变化的敏感性,且不会将未观测单元当作零高度测量处理。方法保留了传统的 Scan Context 粗筛与圆周偏航对齐步骤,因此检索到的候选假设可直接用于几何验证的初始化。在多次重复室内采集、安装高度变化、室内外混合轨迹以及一段室外迁移序列上的实验表明,本方法在室内及安装高度变化的场景中实现了更可靠的首次检索命中。配对检验显示,在自建数据集上其性能显著优于 Scan Context。在基于阈值的接受准则下,UpDown-SC 还取得了最优或次优的 F1max 与 AUPR,同时保持了轻量级的 CPU 前端。连续回放实验证实,检索到的候选假设可支持基于先验地图的度量定位。代码与评估材料:https://github.com/jiejie567/updown-sc。
cs.CV / 33 / 2609.29121

Less is More: Encoder-only Audio-Visual Segmentation

少即是多:仅编码器的音视频分割
Viertola, Ilpo, Iashin, Vladimir, Tötterström, Sophie, Rahtu, Esa
Abstract
Audio-Visual Semantic Segmentation (AVSS) aims to identify, segment, and classify sound-emitting objects in video frames. Previous Transformer-based AVSS approaches largely inherit design principles from image segmentation models. Recent studies show that these image segmentation models contain redundant components that contribute little to the segmentation performance. Following this insight, we propose Encoder-only Audio-Visual Segmentation (EASE). EASE runs at up to 365 FPS, 3x faster than prior State-of-the-Art (SotA) AVS models at comparable accuracy, and trains in under 11 GPU-hours. Furthermore, we achieve SotA AVSS performance across different backbones and input resolutions. Our results demonstrate that AVSS can be both simpler and faster, providing a scalable foundation for future research and real-time applications. Code, model weights, and samples are available at https://ease-avs.notion.site
Chinese Translation
音视频语义分割(Audio-Visual Semantic Segmentation, AVSS)旨在对视频帧中发声物体进行识别、分割和分类。以往基于Transformer的AVSS方法在很大程度上沿用了图像分割模型的设计原则。近期研究表明,这些图像分割模型中包含一些对分割性能贡献甚微的冗余组件。基于这一洞察,我们提出了仅编码器音视频分割方法(Encoder-only Audio-Visual Segmentation, EASE)。EASE的运行速度高达365 FPS,比同等精度的先前最先进(State-of-the-Art, SotA)AVS模型快3倍,且训练时间不足11个GPU小时。此外,我们在不同的骨干网络和输入分辨率下均取得了最先进的AVSS性能。我们的结果表明,AVSS可以做到既更简单又更快速,为未来研究和实时应用提供了一个可扩展的基础。代码、模型权重和样例可在 https://ease-avs.notion.site 获取。
cs.CV / 34 / 2609.29125

FoCal: Frequency-Oriented Cross-Modal Interaction and Spectral Calibration for Aerial Visible-Infrared Object Detection

FoCal:面向频率的跨模态交互与谱校准的航空可见光-红外目标检测方法
Liang, Ben, Sui, Chao, Bai, Junqi, Liu, Yuan, Li, Chunlai, Sui, Xiubao, Chen, Qian
Abstract
In aerial RGB--IR object detection, effectively exploiting complementary information across modalities is critical for robust perception under complex illumination and environmental conditions. Existing multimodal detectors mainly focus on spatial-domain interaction or frequency-specific feature enhancement, while the cross-modal interaction patterns of different frequency components remain insufficiently explored. Moreover, spectral discrepancy itself may contain both useful complementary cues and unreliable modality-specific responses, making indiscriminate frequency fusion suboptimal. To address these issues, we propose FoCal, a frequency-oriented framework for aerial RGB--IR object detection. First, a Frequency-Aware Dual-Domain Calibration (FADC) module is developed to explicitly model frequency-dependent cross-modal interaction. Low-frequency components are collaboratively consolidated into a shared structural consensus, whereas high-frequency components preserve modality-specific information through selective cross-modal exchange. The resulting frequency-aware cues are further transferred to the original feature domain to regulate cross-modal calibration. Second, we introduce a Discrepancy-Guided Spectral Modulation (DGSM) module, which characterizes cross-modal spectral imbalance using confidence-weighted relative amplitude discrepancy and transforms it into a bounded signed gate for adaptive enhancement, preservation, or attenuation of the joint multimodal spectrum. Extensive experiments on DroneVehicle, ESCVehicle, and ATR-UMOD demonstrate the effectiveness of FoCal, yielding $\mathrm{mAP}_{50}$ values of 83.5\%, 54.8\%, and 64.6\%, respectively. Meanwhile, with only 3.0M parameters, FoCal achieves 113.6 FPS while preserving leading detection accuracy, highlighting a favorable accuracy--efficiency trade-off. Code is available at {https://github.com/universeliang/FoCal.
Chinese Translation
在航空RGB-IR目标检测中,有效利用模态间的互补信息对于复杂光照与环境条件下的鲁棒感知至关重要。现有多模态检测器主要关注空间域交互或特定频率的特征增强,而对不同频率分量的跨模态交互模式研究不足。此外,模态间谱差异本身既可能包含有用的互补线索,也可能包含不可靠的模态特异性响应,使得不加区分的频率融合并非最优。为解决这些问题,我们提出了一种面向频率的航空RGB-IR目标检测框架FoCal。首先,我们设计了频率感知双域校准(FADC)模块,以显式建模频率相关的跨模态交互:低频分量被协同整合为共享的结构共识,而高频分量则通过选择性跨模态交换保留模态特异信息。所得的频率感知线索进一步被传递至原始特征域,以调节跨模态校准。其次,我们引入差异引导的谱调制(DGSM)模块,利用置信度加权的相对幅度差异来刻画跨模态谱不平衡,并将其转换为有界符号门控,以对联合多模态频谱进行自适应增强、保留或衰减。在DroneVehicle、ESCVehicle和ATR-UMOD数据集上的大量实验证明了FoCal的有效性,其mAP50分别达到83.5%、54.8%和64.6%。同时,FoCal仅具有3.0M参数,在113.6 FPS的速度下保持了领先的检测精度,展现出良好的精度-效率权衡。代码已发布于 {https://github.com/universeliang/FoCal。
cs.CV / 35 / 2609.29151

Recoverable Geographic Location Information in Earth-Observation Embeddings

对地观测嵌入中可恢复的地理位置信息
Zhang, Peiwen, Hu, Kristie, Knezevic, Jovana, Yin, Shunde, Gao, Kyle
Abstract
Earth-observation (EO) foundation models provide reusable embeddings, yet downstream task accuracy does not reveal whether these representations encode geographic information, which may be beneficial for location-aware applications but potentially detrimental when representations invariant to geographic location are desired. We therefore evaluate the geographic coordinate robustness of Tessera v1, Tessera v1.1, and AlphaEarth by testing whether coordinates can be predicted from the embedding representations using 284 quality-verified European solar farms from 2024. We assessed geographic information content information through the association between cosine and geodesic distances and through prediction of projected coordinates in EPSG:3035. Embeddings from all three EO foundation models contain recoverable geographic information. All prediction models significantly outperform training-range uniform random sampling baselines, with AlphaEarth exhibiting the strongest distance association and lowest mean geodesic error. Both Tessera variants also yielded higher geographic distance correlations than the Sentinel-2 controls. These findings motivate geographic information content as an additional criterion for auditing EO foundation models.
Chinese Translation
对地观测(EO)基础模型提供了可复用的嵌入表示,然而下游任务的准确率并不能揭示这些表示是否编码了地理信息——这对于位置感知应用可能是有益的,但在需要地理位置不变表示的场景中则可能是有害的。因此,我们通过检验能否从嵌入表示中预测地理坐标,评估了 Tessera v1、Tessera v1.1 和 AlphaEarth 的地理坐标稳健性,实验数据为2024年284个经过质量核验的欧洲太阳能电站。我们通过余弦距离与大地测量距离之间的关联,以及通过预测 EPSG:3035 投影坐标来评估地理信息含量。结果表明,三个EO基础模型的嵌入均包含可恢复的地理信息。所有预测模型均显著优于训练范围均匀随机采样基线,其中 AlphaEarth 表现出最强的距离关联和最低的平均大地测量误差。两个 Tessera 变体的地理距离相关性也高于 Sentinel-2 对照组。这些发现表明,地理信息含量可作为审计EO基础模型的一项补充标准。
cs.CV / 36 / 2609.29156

Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation

Med-AR:面向长尾胸部X光分类与不确定性感知评估的自回归视觉-语言预训练
Prabhu, Janhavi, Sahil, V, Akshay, Shukla, Shivam, Tadepalli, Manoj, Putha, Preetham
Abstract
Long-tailed chest X-ray classification requires visual representations that capture both common abnormalities and subtle, infrequent findings. We propose Med-AR-8B and Med-AR-2B, two radiology-native autoregressive vision-language models pretrained with structured reports, abnormality-focused text, and region annotations. We evaluate the transfer of their visual encoders to multi-label classification against contrastive, self-supervised, and supervised pretrained encoders, including Med-CLIP, CheXFound, EVA-Base, ARK, and BioViL-T, using a common ML-Decoder classification head. To assess fine-grained recognition, we also construct LLM-expanded, report-derived label sets for MIMIC-CXR and CheXpert. Across PadChest, MIMIC-CXR, and CheXpert, Med-AR-8B outperforms Med-CLIP in mean AUROC and AUPRC for head, medium, and tail findings. On MIMIC-CXR, it increases tail-label mean AUPRC from 0.1033 to 0.1441. Med-AR-2B achieves the strongest discrimination results on PadChest. Across the broader encoder comparison, a Med-AR variant achieves the highest mean AUROC and AUPRC in every reported prevalence group on each public dataset. Both Med-AR variants also achieve lower excess area under the risk-coverage curve than Med-CLIP on all three public datasets, indicating improved selective-prediction performance under the evaluated protocol. Internal results are metric-dependent, with Med-CLIP retaining advantages in overall and tail AUPRC and in selective prediction. These findings establish Med-AR as a strong pretraining recipe for long-tailed chest X-ray classification on the evaluated public benchmarks and demonstrate the value of assessing discrimination and selective prediction together.
Chinese Translation
长尾胸部X光分类需要能够同时捕捉常见异常和细微、罕见病灶的视觉表示。我们提出了Med-AR-8B和Med-AR-2B两个放射学原生的自回归视觉-语言模型,其预训练采用了结构化报告、以异常为中心的文本以及区域标注。我们使用统一的ML-Decoder分类头,将这两个模型视觉编码器迁移到多标签分类任务上的表现与对比学习、自监督和监督预训练的编码器进行比较,包括Med-CLIP、CheXFound、EVA-Base、ARK和BioViL-T。为评估细粒度识别能力,我们还基于报告为MIMIC-CXR和CheXpert构建了经大语言模型扩展的标签集。在PadChest、MIMIC-CXR和CheXpert数据集上,Med-AR-8B在头部、中等和尾部(罕见)类别的平均AUROC和AUPRC上均优于Med-CLIP。在MIMIC-CXR上,它将尾部标签的平均AUPRC从0.1033提升至0.1441。Med-AR-2B在PadChest上取得了最强的判别性能。在更广泛的编码器比较中,Med-AR的一个变体在每个公开数据集的各个患病率分组中均获得了最高的平均AUROC和AUPRC。两个Med-AR变体在所有三个公开数据集上的风险-覆盖率曲线下的超额面积也均低于Med-CLIP,表明在所评估的协议下其选择性预测性能有所提升。内部测试结果依赖于具体指标,Med-CLIP在总体和尾部AUPRC以及选择性预测方面仍保持优势。这些发现确立了Med-AR作为在所评估的公开基准上进行长尾胸部X光分类的强大预训练方案,并展示了将判别能力与选择性预测结合评估的价值。
cs.CV / 37 / 2609.29186

An Automated Georeferencing Technique for Multi-Temporal Stope Point Clouds for Downstream Geotechnical Analysis

一种面向下游岩土工程分析的多时相采场点云自动地理配准技术
Patra, Dibyayan, Raval, Simit, Ranasinghe, Pasindu, Banerjee, Bikram, Canbulat, Ismet
Abstract
The increasing use of UAV laser scanning in underground mines has enabled frequent acquisition of 3D point clouds from challenging environments such as stopes, generating large volumes of multi-temporal spatial data throughout successive excavation stages. However, in GNSS-denied underground environments, independently acquired stope point clouds are generated within local scanner reference frames and require registration and georeferencing before integration with mine reference data for downstream geotechnical analysis, monitoring, and mine planning. This process is commonly performed manually by aligning individual stope scans with mine reference drives, making repeated georeferencing time-consuming and potentially limiting the utilisation of routinely acquired data. This study proposes the 3D Tag-based Automated Registration and Georeferencing Technique (3D-TARGeT), an automated framework using low-cost, generic, non-unique rectangular tags to establish spatial correspondence between stope point clouds and the mine reference coordinate system. The framework combines automated tag identification, geometric tag matching, and rigid transformation estimation. It was evaluated as a proof of concept using four multi-temporal point-cloud scans of an underground mine stope, with the proposed tags simulated under representative scanning conditions. 3D-TARGeT achieved consistent centimetre-level georeferencing accuracy, with median cloud-to-cloud distance and root mean square error below 0.03 m across all scans, while substantially outperforming widely used automatic point-cloud registration techniques. Overall, 3D-TARGeT provides an accurate and robust approach for automating stope point-cloud georeferencing, reducing reliance on manual alignment and facilitating multi-temporal datasets for downstream geological and geotechnical applications.
Chinese Translation
无人机(UAV)激光扫描在地下矿山中的日益广泛应用,使得从采场等复杂环境中频繁获取三维点云成为可能,并在连续的开采阶段中产生了大量多时相空间数据。然而,在无GNSS信号的地下环境中,独立获取的采场点云是在扫描仪的局部参考坐标系下生成的,需要先进行配准和地理配准,才能与矿山参考数据集成,用于下游的岩土工程分析、监测和矿山规划。该过程通常通过将单个采场扫描与矿山参考巷道进行人工对齐来完成,使得重复的地理配准十分耗时,并可能限制对常规采集数据的利用。本研究提出了一种基于三维标签的自动配准与地理配准技术(3D Tag-based Automated Registration and Georeferencing Technique, 3D-TARGeT),该自动化框架利用低成本、通用且非唯一的矩形标签,在采场点云与矿山参考坐标系之间建立空间对应关系。该框架结合了自动标签识别、几何标签匹配和刚体变换估计。作为概念验证,该方法在一个地下矿山采场的四期多时相点云扫描数据上进行了评估,所提出的标签在代表性扫描条件下进行了模拟。3D-TARGeT实现了一致的厘米级地理配准精度,所有扫描的点云到点云距离中位数和均方根误差均低于0.03米,同时显著优于广泛使用的自动点云配准技术。总体而言,3D-TARGeT为实现采场点云地理配准自动化提供了一种精确且稳健的方法,减少了对人工对齐的依赖,并促进了多时相数据集在下游地质与岩土工程应用中的使用。
cs.CV / 38 / 2609.29193

ImCorr: Sub-pixel Semantic Correspondence via Implicit Feature Decoding

ImCorr:基于隐式特征解码的亚像素语义对应
Choi, Yusung
Abstract
The strong performance that modern semantic correspondence methods achieve at standard thresholds plateaus sharply at fine-grained thresholds. We argue that this plateau stems not from the representational capacity of backbone features, but from a grid-tied readout. Patch-based vision transformers tokenize images onto discrete grids, introducing two forms of quantization error: querying nearest patch features instead of exact keypoints on the source side, and the absence of grid features representing precise ground-truth locations on the target side. We quantify this quantization ceiling across all 499,188 keypoints in SPair-71k: under the standard 448x448, patch-14 setting, 84.9% of ground-truth keypoints have no grid feature representing their precise location at PCK@0.01. This is a structural limitation at the representation level, independent of the matching strategy. We address this with ImCorr: Sub-pixel Semantic Correspondence via Implicit Feature Decoding, which formulates correspondence estimation over a continuous feature field queryable at arbitrary continuous coordinates. A FiLM-conditioned decoder is trained to embed sub-pixel positional information into the feature field. Querying the field directly at exact keypoint coordinates theoretically eliminates representation-level quantization error on the source side, while decoding onto a grid denser than the backbone grid substantially reduces quantization error on the target side. On SPair-71k and AP-10K (intra-species, cross-species, and cross-family), ImCorr improves performance at fine-grained thresholds (PCK@0.01-0.05), achieving a 6.2 percentage point gain over the prior state of the art at PCK@0.01 on SPair-71k. These results demonstrate that representational continuity is an effective solution for precise semantic correspondence. Code is available at https://github.com/YusungChoi/ImCorr.
Chinese Translation
现代语义对应方法在标准阈值下取得了强大性能,但在细粒度阈值下性能急剧趋于平稳。我们认为这一瓶颈并非源于骨干网络特征的表征能力,而是源于绑定于网格的读出方式。基于图像块(patch)的视觉Transformer将图像离散化为网格token,从而引入两种形式的量化误差:在源端,查询最近邻图像块特征而非精确关键点特征;在目标端,缺乏表示精确真值位置的网格特征。我们在SPair-71k的全部499,188个关键点上对这一量化上限进行了量化分析:在标准的448x448、patch-14设置下,84.9%的真值关键点在PCK@0.01下没有表示其精确位置的网格特征。这是一种表征层面的结构性局限,与匹配策略无关。我们提出ImCorr(基于隐式特征解码的亚像素语义对应)来解决这一问题,该方法将对应估计建模为一个可在任意连续坐标处查询的连续特征场。我们训练了一个FiLM条件化的解码器,将亚像素位置信息嵌入特征场。在源端,直接在精确关键点坐标处查询特征场,理论上可消除表征层面的量化误差;在目标端,解码到比骨干网格更密集的网格上,可大幅减少量化误差。在SPair-71k和AP-10K(种内、种间和科间)上,ImCorr在细粒度阈值(PCK@0.01-0.05)下提升了性能,在SPair-71k的PCK@0.01上较此前最优方法提升了6.2个百分点。这些结果表明,表征连续性是实现精确语义对应的有效方案。代码已发布于 https://github.com/YusungChoi/ImCorr。
cs.CV / 39 / 2609.29224

FounRef: Robust, Structure-Preserving, and Fast Metric Refinement of Frozen Monocular Foundation Priors with Sparse Anchors

FounRef:利用稀疏锚点对冻结的单目基础模型先验进行鲁棒、保持结构与快速的度量精化
Halperin, Dan, Mählisch, Mirko
Abstract
Dense metric depth from cameras is essential to real-world 3D applications, yet achieving accuracy, faithful surface geometry, and fast inference simultaneously remains challenging. Monocular foundation models provide rich, transferable geometric priors but lack reliable metric scale, while depth-completion networks recover metric depth at the cost of geometric fidelity, cross-domain robustness, or speed. We present FounRef, a training-free method that aligns a frozen monocular foundation prior with sparse metric anchors to produce dense metric depth. FounRef is modular by design: its depth prior, anchor source, and refinement solver can each be replaced independently. We instantiate FounRef with MoGe-2 and LiDAR anchors. FounRef validates each anchor against the prior's dense depth prediction, rejecting inconsistencies caused by cross-sensor misalignment that geometry-only filters cannot detect. It then applies global and local metric corrections through a structure-preserving solver, retaining the prior's fine-grained geometry. FounRef requires no task-specific training and operates out of the box across unfamiliar cameras and scenes. On out-of-domain data, it delivers up to 24% lower depth error, 92% lower surface-normal noise, and almost 15x faster inference than DMD3C, a state-of-the-art depth-completion network. By decoupling metric alignment from geometry prediction, FounRef provides an accurate, geometrically faithful, and efficient approach to dense metric depth that can directly benefit from future advances in foundation models and metric sensors.
Chinese Translation
相机获取的稠密度量深度对现实世界的3D应用至关重要,然而同时实现高精度、忠实的表面几何以及快速推理仍然具有挑战性。单目基础模型提供了丰富且可迁移的几何先验,但缺乏可靠的度量尺度;而深度补全网络虽能恢复度量深度,却以牺牲几何保真度、跨域鲁棒性或速度为代价。我们提出了FounRef,这是一种免训练方法,通过将冻结的单目基础模型先验与稀疏度量锚点对齐来生成稠密度量深度。FounRef采用模块化设计:其深度先验、锚点来源和精化求解器均可独立替换。我们使用MoGe-2和LiDAR锚点对FounRef进行了实例化。FounRef将每个锚点与先验的稠密深度预测进行校验,拒绝由跨传感器错位引起的、仅基于几何的滤波器无法检测到的不一致锚点。随后,它通过一个结构保持求解器施加全局和局部度量校正,保留先验的细粒度几何。FounRef无需针对特定任务的训练,即可在陌生的相机和场景中开箱即用。在域外数据上,与最先进的深度补全网络DMD3C相比,FounRef的深度误差最多降低24%,表面法向噪声降低92%,推理速度快近15倍。通过将度量对齐与几何预测解耦,FounRef提供了一种精确、几何忠实且高效的稠密度量深度估计方法,并能够直接受益于未来基础模型和度量传感器的发展。
cs.CV / 40 / 2609.29225

ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios

ComplexSync:复杂场景下的高保真实时唇形同步
Cai, Jiaran, Ma, Xingpei, Huang, Shenneng
Abstract
Lip synchronization aims to generate visual lip dynamics that align precisely with speech audio. Despite the high generation quality of diffusion models, they often struggle in complex scenarios and suffer from prohibitive inference latency, limiting real-world deployment. We present ComplexSync, a unified diffusion-based framework that enables real-time, high-fidelity lip sync under complex conditions. First, we introduce a dual-stream joint training strategy to mitigate information leakage from reference frames while preserving natural dynamics. Second, we develop a distillation-based acceleration scheme for single-step denoising, achieving a throughput of over 70 FPS. Third, we propose a relational alignment loss that leverages structural priors from Vision Foundation Models (VFMs) to enhance robustness against complex scene factors. Furthermore, we present the first benchmark specifically designed for complex lip synchronization, comprising over 200 challenging video sequences and specialized metrics. Extensive experiments demonstrate that ComplexSync achieves state-of-the-art performance across both standard and complex scenarios while enabling real-time inference.
Chinese Translation
唇形同步旨在生成与语音音频精确对齐的视觉唇部动态。尽管扩散模型具有很高的生成质量,但它们在复杂场景中往往表现不佳,且推理延迟过高,限制了实际部署。我们提出ComplexSync,一个统一的基于扩散的框架,能够在复杂条件下实现实时、高保真的唇形同步。首先,我们引入一种双流联合训练策略,在保留自然动态的同时缓解参考帧的信息泄露问题。其次,我们开发了一种基于蒸馏的单步去噪加速方案,实现了超过70 FPS的吞吐量。第三,我们提出一种关系对齐损失,利用视觉基础模型(Vision Foundation Models, VFMs)的结构先验来增强对复杂场景因素的鲁棒性。此外,我们提出了首个专门针对复杂唇形同步的基准,包含200多个具有挑战性的视频序列和专用评估指标。大量实验表明,ComplexSync在标准和复杂场景中均达到了最先进的性能,同时支持实时推理。
cs.CV / 41 / 2609.29235

SARFusion: Scene-Aware Routing Fusion for Robust Camera-LiDAR 3D Object Detection

SARFusion:面向鲁棒相机-激光雷达3D目标检测的场景感知路由融合
Zhao, Yuting, Zheng, Ziyi, Li, Shuxiao
Abstract
Camera-LiDAR fusion has become a prevailing paradigm for 3D object detection in autonomous driving. However, existing fusion detectors often establish strong inter-modality dependencies by decoding object queries from tightly coupled multimodal representations. Under corrupted driving conditions, such dependencies make the detector vulnerable to unreliable modalities, where degraded observations may interfere with reliable modality-specific evidence and lead to suboptimal predictions. Moreover, modality reliability can vary across both global driving scenes and individual object queries, requiring adaptive fusion decisions at a finer granularity. To bridge this gap, we reformulate robust camera-LiDAR fusion as a scene-aware branch routing problem and propose SARFusion, a robust 3D object detector. Instead of producing detections from a single fused representation, SARFusion decouples object-query decoding into three parallel reasoning branches: a camera branch, a LiDAR branch, and a camera-LiDAR fusion branch. Guided by a Scene Reliability Prior estimated from the global driving context, SARFusion further incorporates object-level evidence to route each query to the most suitable branch. This query-wise routing strategy alleviates harmful cross-modal interference while preserving the benefits of multimodal fusion when complementary cues are trustworthy. On the nuScenes test set, SARFusion achieves strong performance with 72.5 mAP and 74.4 NDS. Extensive analyses demonstrate its robustness under challenging conditions, including sensor corruptions and environmental changes.
Chinese Translation
相机-激光雷达融合已成为自动驾驶中3D目标检测的主流范式。然而,现有融合检测器通常通过对 tightly coupled(紧耦合)多模态表示中的目标查询进行解码,建立起强烈的模态间依赖关系。在受损的驾驶条件下,这种依赖使检测器容易受到不可靠模态的影响,退化的观测可能干扰可靠的单模态证据,导致次优预测。此外,模态可靠性会随全局驾驶场景和单个目标查询而变化,因此需要在更细的粒度上进行自适应融合决策。为弥合这一差距,我们将鲁棒的相机-激光雷达融合重新表述为一个场景感知的分支路由问题,并提出了一种鲁棒的3D目标检测器SARFusion。SARFusion不再从单一融合表示生成检测结果,而是将目标查询解码解耦为三个并行推理分支:相机分支、激光雷达分支以及相机-激光雷达融合分支。在由全局驾驶上下文估计得到的场景可靠性先验(Scene Reliability Prior)的引导下,SARFusion进一步结合目标级证据,将每个查询路由到最合适的分支。这种逐查询(query-wise)路由策略在多模态互补线索可信时保留多模态融合的优势,同时减轻了有害的跨模态干扰。在nuScenes测试集上,SARFusion取得了72.5 mAP和74.4 NDS的优异性能。大量分析证明了其在传感器损坏和环境变化等挑战性条件下的鲁棒性。
cs.CV / 42 / 2609.29240

TOLA: Text-aware One-Step Latent Adaptation for Diffusion-based Text Image Super-Resolution

TOLA:面向基于扩散模型的文本图像超分辨率的文本感知单步潜在自适应方法
Xu, Yike, Shi, Yue, Guo, Yong, Cao, Jiezhang
Abstract
Text image super-resolution (TSR) aims to recover visually faithful and readable text under unknown degradations. Existing diffusion-based methods typically rely on multi-step prediction of either the high-resolution image or its text prior, resulting in prohibitive computational cost and inference latency. More critically, an erroneous text prior may be repeatedly injected into the denoising process, causing image and text predictions to reinforce each other and progressively amplify an early recognition error into a sharp yet semantically incorrect character. To address these limitations, we propose TOLA, a Text-aware One-step Latent Adaptation framework without iterative image-text diffusion. TOLA consists of two key modules. First, a confidence-weighted text conditioning module constructs the semantic condition only once and suppresses unreliable OCR predictions before they contaminate image reconstruction. Second, a lightweight latent residual correction module explicitly estimates and corrects the structured residual errors to recover missing or distorted stroke details. Extensive experiments demonstrate our state-of-the-art performance across all evaluation metrics on both CTR-TSR-Test ($\times 4$) and RealCE-200 benchmarks. It is worth noting that our TOLA consistently surpasses existing diffusion-based TSR methods by at least 2.72 dB in PSNR on CTR-TSR-Test.
Chinese Translation
文本图像超分辨率(Text Image Super-Resolution, TSR)旨在未知退化条件下恢复视觉上真实且可读的文本。现有的基于扩散模型的方法通常需要对高分辨率图像或其文本先验进行多步预测,导致计算成本高昂、推理延迟严重。更为关键的是,错误的文本先验可能被反复注入去噪过程,使图像预测与文本预测相互强化,从而将早期识别错误逐步放大为清晰但语义错误的字符。为解决上述局限,我们提出 TOLA,一种无需迭代图文扩散的文本感知单步潜在自适应(Text-aware One-step Latent Adaptation)框架。TOLA 包含两个关键模块:其一,置信度加权的文本条件模块仅构建一次语义条件,并在不可靠的 OCR 预测污染图像重建之前将其抑制;其二,轻量级的潜在残差校正模块显式估计并校正结构化残差误差,以恢复缺失或畸变的笔画细节。大量实验表明,在 CTR-TSR-Test(×4)和 RealCE-200 两个基准上,我们的方法在所有评价指标上均达到最先进的性能。值得注意的是,在 CTR-TSR-Test 上,我们的 TOLA 在 PSNR 指标上以至少 2.72 dB 的优势持续超越现有的基于扩散模型的 TSR 方法。
cs.CV / 43 / 2609.29252

IronViT: Toward Efficient Generalist Visual Representation Learning

IronViT:迈向高效的通用视觉表示学习
Huang, Jiaxi, Hu, Yueqi, Zhu, Xin, Zhang, Xiaopeng, Qiao, Huiting, Zhang, Yanglin, Ji, Zefeng, Li, Rongxue, Xu, Yifei, Yu, Huiying, Liu, Wei, Zheng, Jiayin, Xu, Yinggan, Chen, Peipeng, Zhang, Yin, Yao, Jian
Abstract
A generalist vision encoder must capture semantic, spatial, language-aligned, and action-relevant cues within a unified representation, yet softmax attention underlying today's most capable visual backbones becomes prohibitively expensive at high resolution. A natural attempt to address both challenges is to distill multiple specialist teachers directly into an efficient architecture. We find that directly coupling these objectives degrades representation quality, as the student must simultaneously reconcile heterogeneous capabilities and adapt them to a different token-mixing architecture. We introduce IronViT, built on a simple principle: consolidate capabilities before constraining computation. IronViT first distills complementary specialists into a softmax attention capability bridge, then progressively transfers the consolidated representation to a hybrid softmax-linear attention encoder. A purpose-built data pipeline further curates the distillation corpus for higher information density and broader domain coverage. Across recognition, retrieval, dense prediction, multimodal understanding, and robotic learning, IronViT is competitive with leading specialist and generalist vision encoders. The softmax bridge achieves the strongest aggregate performance in multimodal understanding and robotic learning among the evaluated backbones, while the hybrid encoder retains broad transfer performance with an efficiency advantage that grows with input resolution. Together, these results show that consolidating capabilities before architectural conversion can yield a generalist visual encoder without inheriting the prohibitive high-resolution cost of conventional softmax attention.
Chinese Translation
通用视觉编码器必须在统一表示中捕捉语义、空间、语言对齐以及动作相关的线索,然而当今最强大的视觉骨干网络所依赖的softmax注意力机制在高分辨率下代价极高。解决这两个挑战的一个自然尝试是将多个专家教师模型直接蒸馏到一个高效架构中。我们发现,直接耦合这些目标会降低表示质量,因为学生模型必须同时调和异构能力,并将其适配到不同的词元混合(token-mixing)架构中。我们提出IronViT,其基于一个简单的原则:先整合能力,再施加计算约束。IronViT首先将互补的专家模型蒸馏到一个softmax注意力能力桥接模块中,然后将整合后的表示逐步迁移到一个softmax-线性混合注意力编码器中。一个专门构建的数据流水线进一步对蒸馏语料进行筛选,以实现更高的信息密度和更广的领域覆盖。在识别、检索、密集预测、多模态理解和机器人学习等任务上,IronViT与领先的专家型和通用型视觉编码器相比均具有竞争力。在所评估的骨干网络中,softmax桥接模块在多模态理解和机器人学习方面取得了最强的综合性能,而混合编码器在保持广泛迁移性能的同时,具备随输入分辨率增长而不断增强的效率优势。这些结果表明,在架构转换之前先整合能力,可以获得一个通用视觉编码器,而无需承担传统softmax注意力在高分辨率下高昂的代价。
cs.CV / 44 / 2609.29256

Deep learning of longitudinal visual fields predicts glaucoma progression rate and identifies fast progressors

纵向视野数据的深度学习预测青光眼进展速度并识别快速进展者
Rahman, Taiabur, Rahman, Siddiqur, Moniruzzaman, Muhammad, Kawsar, Ummay, Ratna, Sayedatunnessa, Siddique, Shadman, Siddique, Rafsan, Ahmad, Tausif, Ahmad, Tahsin, Rabbani, Golam
Abstract
Glaucoma is the leading cause of irreversible blindness, and timely identification of fast progressors is essential to prevent disability. Current practice estimates progression by ordinary least-squares regression of mean deviation (MD) on time, requiring 6--10 visual field (VF) tests over several years to obtain a reliable slope. We present GLAM (Glaucoma Longitudinal Analysis Model), a deep learning framework that ingests longitudinal Humphrey 24-2 total deviation sequences with five clinical features and predicts MD and visual field index progression rates using attention-based fusion and aleatoric uncertainty. On the open-access University of Washington Humphrey Visual Field dataset (4,276 patient-eyes), GLAM achieved an MD-rate mean absolute error of 0.139 dB yr$^{-1}$ ($R^2 = 0.927$; 73.5% reduction over a ridge baseline) and an AUC of 0.990 for fast-progressor detection. VF-only deep learning can match multimodal pipelines for progression prognostication using routinely collected perimetry alone.
Chinese Translation
青光眼是导致不可逆性失明的首要病因,及时识别快速进展者对于防止残疾至关重要。目前的临床实践通过平均偏差(MD)随时间的普通最小二乘回归来估计进展,需要数年内进行6至10次视野(VF)检查才能获得可靠的斜率。我们提出了GLAM(Glaucoma Longitudinal Analysis Model,青光眼纵向分析模型),这是一种深度学习框架,其输入为纵向Humphrey 24-2总偏差序列并结合五项临床特征,利用基于注意力的融合和偶然不确定性来预测MD和视野指数的进展速率。在公开的华盛顿大学Humphrey视野数据集(4,276只患眼)上,GLAM实现了MD进展速率的平均绝对误差为0.139 dB/年(R² = 0.927;相比岭回归基线降低73.5%),快速进展者检测的AUC为0.990。仅使用视野数据的深度学习即可仅凭借常规收集的视野检查结果,在进展预后方面媲美多模态流程。
cs.CV / 45 / 2609.29292

PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark

PHOSA:照片级真实的3D手语数字人建模与基准
Wang, Haodong, Hu, Hezhen, Zhou, Wengang, Li, Houqiang
Abstract
In this work, we focus on photorealistic sign avatar modeling, which is crucial for effective communication with the Deaf community and is characterized by complex hand gestures and nuanced facial expressions. To this end, we introduce MVSign, the first multi-view Chinese sign language dataset co-designed with Deaf experts, featuring diverse gestures and rich annotations. For precise SMPL-X annotation, we develop a hybrid fitting pipeline that produces accurate body, hand, and facial parameters and can also be applied to the monocular setting. Building on MVSign, we propose a decoupled sign avatar representation that isolates body, head, and hand components to capture complex articulations, together with a motion-aware sampling strategy to handle motion blur and balance gesture diversity. Extensive experiments demonstrate that our method achieves high-fidelity visual results on MVSign, particularly in detailed hand and facial regions, and generalizes well to in-the-wild monocular sign language videos. Project page: https://naaapi.github.io/PHOSA.
Chinese Translation
在本工作中,我们专注于照片级真实感的手语数字人建模,这对于与聋人社区的有效沟通至关重要,其特点包括复杂的手部动作和细腻的面部表情。为此,我们提出了MVSign,这是首个与聋人专家共同设计的多视角中国手语数据集,具有多样化的手势和丰富的标注。为实现精确的SMPL-X标注,我们开发了一种混合拟合流程,能够生成准确的身体、手部和面部参数,并且同样适用于单目设置。基于MVSign,我们提出了一种解耦的手语数字人表示方法,将身体、头部和手部组件分离以捕捉复杂的关节动作,并结合一种运动感知采样策略来处理运动模糊并平衡手势多样性。大量实验表明,我们的方法在MVSign上实现了高保真的视觉效果,尤其是在精细的手部和面部区域,并且能够很好地泛化到真实场景的单目手语视频。项目页面:https://naaapi.github.io/PHOSA。
cs.CV / 46 / 2609.29329

Hyperbolic Multimodal Continual Learning: A Closest-Admissible Solution

双曲多模态持续学习:一种最接近的可行解
Liu, Jiahong, Shen, Ming, Liu, Xiaohao, Ying, Rex, Yang, Menglin, Chua, Tat-Seng, King, Irwin
Abstract
Existing continual-learning methods protect parameters, replayed examples, or Euclidean feature subspaces. When applied to hyperbolic multimodal models, they do not explicitly preserve the Lorentz geometry that jointly encodes within-modality similarity, cross-modal correspondence, and semantic hierarchy; sequential updates can therefore retain task scores while still distorting previously learned relations. We address this gap with Hyperbolic Multimodal Continual Learning (HMCL). We show that preserving the old multimodal geometry amounts to restricting all modalities to one shared hyperbolic isometry, which induces a family of admissible first-order parameter changes. We formulate a joint closest-admissible (CA) correction that retains the shared rotation best matching the candidate modal updates; its minimal-rotation (MR) special case fixes this rotation to zero. Both variants correct the displacement realized by AdamW, and task anchoring bounds within-task accumulation while preserving learning freedom. Across a unified 16-task classification-retrieval stream with three hyperbolic backbones, HMCL improves final performance and backward transfer over sequential fine-tuning and four continual-learning baselines; HMCL-CA gives the highest Overall score on every backbone. A modality-extended stream confirms the retrieval gains. Representation analyses find 81.2 to 95.5 percent less radial, angular, cross-modal, and paired-distance drift; ImageNet-WordNet results show better semantic ancestry and radial hierarchy.
Chinese Translation
现有的持续学习方法保护的是参数、重放样本或欧氏特征子空间。当应用于双曲多模态模型时,这些方法并未显式地保持洛伦兹几何(Lorentz geometry),而该几何同时编码了模态内相似性、跨模态对应关系以及语义层次结构;因此,序列化更新虽然可以维持任务分数,却仍可能扭曲先前学习到的关系。我们提出双曲多模态持续学习(HMCL)来弥补这一空白。我们证明,保持旧的多模态几何等价于将所有模态限制在同一个共享的双曲等距变换上,这会诱导出一族可行的一阶参数变化。我们构造了一种联合的最接近可行(Closest-Admissible,CA)校正方法,该方法保留与候选模态更新最匹配的共享旋转;其最小旋转(Minimal-Rotation,MR)特例将该旋转固定为零。两种变体均对 AdamW 实现的位移进行校正,同时任务锚定在保持学习自由度的同时限制了任务内的累积偏移。在一个包含16个任务的统一分类-检索数据流上,使用三种双曲骨干网络进行实验,HMCL 相对于序列微调和四种持续学习基线方法,提升了最终性能和向后迁移能力;HMCL-CA 在所有骨干网络上均取得了最高的总体分数。扩展模态的数据流进一步证实了检索性能的提升。表征分析显示,径向、角度、跨模态及配对距离漂移减少了81.2%至95.5%;ImageNet-WordNet 的结果表明其具有更好的语义谱系和径向层次结构。
cs.CV / 47 / 2609.29334

A Study of the Limits of Collaborative DCT-Based Image Denoising via Interpretable Neural Networks

基于可解释神经网络研究协同DCT图像去噪方法的极限
Comellas, Cristian, Navarro, Julia, Buades, Antoni
Abstract
Image denoising remains a fundamental problem in image restoration, with applications in photography, biomedical, and scientific imaging. Modern deep neural networks achieve strong performance by learning powerful image priors, but often rely on large black-box models with limited interpretability. In contrast, DCT-based sliding-window and collaborative filtering methods such as BM3D offer clear algorithmic structure, but depend on handcrafted and non-differentiable operations. This work studies how far such structured collaborative filtering principles can be pushed when reformulated as trainable models. We introduce DeepBM3D, a compact fully differentiable architecture that combines non-local patch grouping, DCT-domain filtering, and multi-stage refinement within a BM3D-inspired pipeline. Lightweight convolutional feature extractors guide patch grouping, while filtering is performed through learned Wiener weights in the DCT domain. Experiments show that DeepBM3D improves over classical and hybrid baselines, remains competitive with FFDNet at low and moderate noise levels, and performs particularly well on repetitive textures.
Chinese Translation
图像去噪仍然是图像复原中的一个基础性问题,在摄影、生物医学和科学成像等领域具有广泛应用。现代深度神经网络通过学习强大的图像先验获得了优异的性能,但往往依赖于可解释性有限的大型黑盒模型。相比之下,基于DCT的滑动窗口和协同滤波方法(如BM3D)具有清晰的算法结构,但依赖于手工设计的不可微操作。本研究探讨了将这类结构化协同滤波原理重构为可训练模型后所能达到的性能极限。我们提出了DeepBM3D,这是一个紧凑的全可微架构,它在受BM3D启发的流水线中结合了非局部图像块分组、DCT域滤波和多阶段精细化处理。轻量级卷积特征提取器用于引导图像块分组,而滤波则通过在DCT域中学习到的维纳权重来实现。实验表明,DeepBM3D优于经典方法和混合基线方法,在中低噪声水平下与FFDNet具有竞争力,并且在重复性纹理上表现尤为出色。
cs.CV / 48 / 2609.29347

SEE Challenge 2026: Event-Guided Brightness Adjustment Across a Broad Illumination Range

SEE Challenge 2026:跨越宽光照范围的事件引导亮度调整
Lu, Yunfan, Xu, Mingchao, Zhou, Hanyu, Liu, Shaoyu, Liu, Haoyue, Duan, Peiqi, Peng, Shihan, Zheng, Yinqiang, Shi, Boxin, Lee, Gim Hee, Xiong, Hui, Scaramuzza, Davide
Abstract
Event cameras provide a high dynamic range and preserve brightness-change cues in lighting conditions where conventional RGB frames may be noisy or saturated. To benchmark event-guided restoration across a broad illumination range, we organized the SEE Challenge 2026 with the Event-Based Multimodal Vision Workshop at ECCV 2026. The task conditions restoration on one or more RGB frames, synchronized events, and a scalar target-brightness statistic provided by the organizers. It uses SEE-600K, which contains 610,126 image-event observations from 202 real-world scenes spanning low-light, normal-light, and high-light conditions with illumination variations of up to 1,000$\times$. The challenge follows an open-system protocol: participants may use different temporal contexts, architectures, pretrained weights, test-time augmentation, and post-processing strategies. PSNR determines the ranking, and SSIM is reported as a secondary metric. Around 70 teams registered interest and 15 valid CodaBench submissions were received. Six distinct teams completed organizer-side identity and technical verification, provided method descriptions, checkpoints, inference code, and instructions, and are included in the verified open-system ranking reported here. Beyond the ranking, this report analyzes exposure subsets, semantically distinct test cases, a shared failure pattern, system design choices, and inference strategies. The top systems obtain closely spaced average scores, while the best-performing method varies across cases and metrics; under severe underexposure, all verified systems retain visible local errors.
Chinese Translation
事件相机具有高动态范围,并在传统RGB帧可能出现噪声或饱和的光照条件下保留亮度变化线索。为了在宽光照范围内对事件引导的图像复原进行基准测试,我们联合ECCV 2026的基于事件的多模态视觉研讨会(Event-Based Multimodal Vision Workshop)组织了SEE Challenge 2026。该任务以一帧或多帧RGB图像、同步的事件数据以及组织者提供的标量目标亮度统计量为条件进行复原。任务采用SEE-600K数据集,其包含来自202个真实场景的610,126组图像-事件观测数据,涵盖弱光、正常光照和强光条件,光照变化幅度高达1,000倍。该挑战赛遵循开放系统协议:参赛者可以使用不同的时间上下文、网络架构、预训练权重、测试时增强和后处理策略。排名由PSNR决定,SSIM作为次要指标一并报告。约70支团队注册参赛,共收到15份有效的CodaBench提交。六支不同的团队完成了组织者端的身份与技术验证,提供了方法描述、模型检查点、推理代码和使用说明,并被纳入本文报告的经认证的开放系统排名。除排名之外,本报告还分析了曝光子集、语义上不同的测试用例、共同的失效模式、系统设计选择以及推理策略。排名靠前的系统得分差距很小,而表现最优的方法因具体案例和指标而异;在严重欠曝条件下,所有经验证的系统仍存在可见的局部误差。
cs.CV / 49 / 2609.29350

Learning a Flow to Self-Supervised Representations

学习一个通向自监督表示的流(Flow)
Jiao, Yuling, Ma, Wensen, Qi, Houduo, Sun, Defeng
Abstract
Explicit geometric references offer a direct way to structure self-supervised representations. Existing adversarial distribution-matching formulations, however, require costly encoder-critic optimization. We introduce Flow-Based Distribution Matching (FBDM), a non-adversarial framework that learns this reference-directed geometry through spherical conditional velocity regression. An ETF-inspired reference allows its number of components K' to exceed the auxiliary flow dimension d* while retaining structured geometric separation. We assign both augmented views of each image to the same target, while limiting how many images each reference center can receive. An explicit alignment loss further pulls the two views' representations closer together. Experiments across benchmarks ranging from CIFAR to ImageNet show that FBDM achieves performance nearly on par with DM and remains competitive with existing SSL methods. Matched training-cost comparisons show a 1.48- to 1.83-fold speedup over DM with a negligible increase in GPU memory usage. We also provide a theoretical explanation for the usefulness of the learned representations: under stated conditions, we bound the downstream misclassification rate in terms of the FBDM pretraining loss.
Chinese Translation
显式的几何参照为结构化自监督表示提供了一种直接的方式。然而,现有的对抗性分布匹配方法需要代价高昂的编码器-判别器(encoder-critic)联合优化。我们提出了基于流的分布匹配框架(Flow-Based Distribution Matching, FBDM),这是一个非对抗性框架,通过球面条件速度回归来学习这种以参照为导向的几何结构。受ETF(等角紧框架)启发的参照使其分量数量 K' 可以超过辅助流维度 d*,同时保持结构化的几何分离。我们将每张图像的两个增强视图分配到同一目标,同时限制每个参照中心可以接收的图像数量。此外,一个显式的对齐损失进一步拉近两个视图的表示。从CIFAR到ImageNet等多个基准上的实验表明,FBDM 的性能几乎与 DM 相当,并与现有自监督学习(SSL)方法保持竞争力。在训练成本匹配的对比中,FBDM 相比 DM 取得了1.48至1.83倍的加速,而GPU显存占用仅略有增加。我们还为所学表示的有用性提供了理论解释:在给定条件下,我们以 FBDM 预训练损失为界给出了下游误分类率的上界。
cs.CV / 50 / 2609.29358

Domain Recentering and Confidence-Weighted Prior Calibration for Vision-Language Models

面向视觉-语言模型的域重中心化与置信度加权先验校准
Seol, Youngeun, Shin, Jimin, Yoon, Heeseo, Hwang, Uiwon
Abstract
Vision-language models such as CLIP achieve strong zero-shot classification, yet under distribution shift, visual embeddings drift from fixed text embeddings. Training-free calibration avoids the per-sample optimization of prompt learning, but prior feature calibration gives each image the full bias of one hard cluster. We propose Domain Recentering with Confidence Calibration (DRC), a training-free method adapting CLIP from a set of unlabeled target images. DRC fits a Gaussian mixture once and subtracts from each embedding a posterior-weighted average of component means. It then removes residual class preference with a log-prior correction, estimating the prior from confidence-weighted predictions. Among compared methods, DRC achieves the highest average accuracy on cross-domain datasets, exceeding zero-shot CLIP by 4.13 and 5.07 points with ViT-B/16 and ResNet-50, with gains over CLIP also holding under ImageNet distribution shifts.
Chinese Translation
诸如CLIP等视觉-语言模型在零样本分类任务上表现出色,但在分布偏移下,视觉嵌入会偏离固定的文本嵌入。免训练校准方法避免了提示学习所需的逐样本优化,但先前的特征校准方法会给每个图像施加来自单一硬聚类的完整偏置。我们提出了一种结合置信度校准的域重中心化方法(Domain Recentering with Confidence Calibration, DRC),这是一种利用无标注目标图像集合来适配CLIP的免训练方法。DRC仅需拟合一次高斯混合模型,并从每个嵌入中减去各分量均值的后验加权平均。随后,该方法通过基于置信度加权预测估计的先验进行对数先验校正,以消除残余的类别偏好。在所比较的方法中,DRC在跨域数据集上取得了最高的平均准确率:使用ViT-B/16和ResNet-50时,分别比零样本CLIP高出4.13和5.07个百分点;并且该方法相较于CLIP的性能提升在ImageNet分布偏移场景下同样成立。
cs.CV / 51 / 2609.29373

Shadow Reduction in Ultrasound Imaging Using Differentiable Simulation and Radiance Field Decomposition

基于可微仿真与辐射场分解的超声成像声影消除
Bacher, Valentin, Yeung, Pak Hei, Kainz, Bernhard, Wyburd, Madeleine K., Dinsdale, Nicola K., Gray, Michael, Namburete, Ana I. L.
Abstract
Acoustic shadows from bone and other highly attenuating tissues obscure clinically important structures in ultrasound. In fetal brain imaging, skull-induced artefacts disproportionately degrade the hemisphere closer to the transducer (proximal), limiting symmetric assessment of the two hemispheres. Existing correction methods require raw scanner data, impose restrictive assumptions on tissue properties, or rely on generative models that may hallucinate anatomy. We present RFlash, a physics-informed post-processing method that decomposes beamformed ultrasound images into explicit attenuation and scatter-intensity maps using a differentiable radiance-field formulation of image formation. Attenuation-adaptive re-rendering then removes the dependence of the signal at each depth on the intervening tissue, equivalent to virtually advancing the transducer into the tissue. Across 1,261 3D fetal brain volumes, 143 real 2D curvilinear abdominal scans, and 1,200 simulated 2D linear-probe liver scans, RFlash reduces shadow-related intensity differences more effectively than classical Hughes-Duck attenuation correction. For a gestational-age model trained on the distal hemisphere (further from the transducer) and applied to the proximal hemisphere, prediction error decreases by 5.1 days (40%) relative to the original images. The estimated attenuation maps also yield shadow-confidence maps that improve random-forest bone-shadow segmentation over the image alone and receive greater SHAP importance than an existing neural confidence-map baseline, suggesting greater physical consistency. RFlash requires neither hardware modification nor access to raw scanner data and supports 2D and 3D acquisitions with linear and curvilinear probes, making it widely applicable allowing clinicians to use our method on their already acquired scanners and images.
Chinese Translation
骨骼及其他高衰减组织产生的声学阴影会遮挡超声图像中具有重要临床意义的结构。在胎儿脑成像中,颅骨引起的伪影对靠近探头一侧(近端)的大脑半球造成尤为严重的退化,限制了对两个半球进行对称评估的能力。现有的校正方法要么需要扫描仪的原始数据,要么对组织特性施加严格的假设,要么依赖可能产生解剖结构幻觉的生成模型。我们提出了一种基于物理的后处理方法 RFlash,它利用可微分的辐射场图像成像公式,将波束形成的超声图像分解为显式的衰减图和散射强度图。随后,通过衰减自适应的重渲染,去除每个深度处信号对中间组织的依赖,等效于将探头虚拟地推进到组织内部。在 1,261 例 3D 胎儿脑体积数据、143 例真实 2D 凸阵腹部扫描以及 1,200 例模拟 2D 线阵探头肝脏扫描上,RFlash 在减少阴影相关强度差异方面比经典的 Hughes-Duck 衰减校正更为有效。对于一个在远端半球(距探头较远一侧)上训练并应用于近端半球的孕周预测模型,相对于原始图像,预测误差降低了 5.1 天(40%)。此外,估计得到的衰减图还可生成阴影置信度图,相比仅使用图像,它能提升随机森林骨骼阴影分割的性能,并且其 SHAP 重要性高于现有的神经网络置信度图基线,表明其具有更强的物理一致性。RFlash 无需硬件改造,也无需访问扫描仪原始数据,同时支持线阵和凸阵探头的 2D 与 3D 采集,因而具有广泛的适用性,使临床医生能够在已有的扫描仪和图像上使用我们的方法。
cs.CV / 52 / 2609.29376

A Hybrid CNN--State-Space--Attention Backbone with Joint-Embedding Predictive Pretraining for 12-Lead ECG Classification

一种结合联合嵌入预测预训练的混合CNN-状态空间-注意力骨干网络用于12导联心电图分类
Bazi, Yakoub, Aljuhani, Sarah, Rahhal, Mohamad M. Al, Zuair, Mansour, Alajlan, Naif
Abstract
Automatic 12-lead electrocardiogram (ECG) classification requires representations that jointly capture local waveform morphology, long-range temporal dynamics, and cross-lead dependencies, yet integrating these properties within a single efficient architecture remains challenging. This paper introduces a hybrid CNN-SSM-Attention backbone for 12-lead ECG classification. A convolutional stem performs early waveform tokenization and temporal reduction, mixed state-space and depthwise-convolutional blocks model temporal dynamics and local morphology, and a late self-attention stage enables global token interaction at reduced resolution. To improve transfer from unlabeled data, we further develop an ECG-oriented Joint-Embedding Predictive Pretraining (JEPA) framework. Unlike ViT-based JEPA methods that mask patch tokens before the encoder, the proposed method samples span masks at the latent temporal resolution and projects them back to the waveform domain, then predicts clean latent targets from a momentum encoder without waveform reconstruction. Experiments on CPSC2018, Chapman-Shaoxing, and PTB-XL, with pretraining on approximately 350K unlabeled CODE-15 recordings, show that the proposed backbone provides strong supervised baselines under a compact parameter budget. JEPA pretraining further improves transfer, particularly in reduced-label settings and under both full fine-tuning and LoRA-based adaptation. Code: https://github.com/yakoubbazi/Hybrid_ECG_Jepa
Chinese Translation
自动12导联心电图(ECG)分类需要能够同时捕捉局部波形形态、长程时间动态以及跨导联依赖关系的表示,然而在单一高效架构中整合这些特性仍然具有挑战性。本文提出了一种用于12导联心电图分类的混合CNN-SSM-Attention骨干网络。卷积干层(convolutional stem)执行早期波形分词与时间降采样,混合状态空间块与深度卷积块建模时间动态与局部形态,后期的自注意力阶段在较低分辨率下实现全局token交互。为了提升从无标签数据的迁移能力,我们进一步开发了一种面向心电图的联合嵌入预测预训练(JEPA)框架。与在编码器之前对patch token进行掩码的基于ViT的JEPA方法不同,所提方法在潜在时间分辨率上采样片段掩码并将其投影回波形域,然后从动量编码器预测干净的潜在目标,而无需进行波形重建。在CPSC2018、Chapman-Shaoxing和PTB-XL数据集上的实验(预训练使用约35万条无标签CODE-15记录)表明,所提出的骨干网络在紧凑的参数预算下提供了强大的有监督基线。JEPA预训练进一步提升了迁移性能,尤其在标签减少的设置下,以及在完全微调和基于LoRA的适配中均有显著改善。代码:https://github.com/yakoubbazi/Hybrid_ECG_Jepa
cs.CV / 53 / 2609.29384

Segment-Level Risk Discovery in Online Handwriting for Alzheimer's Disease Detection

面向阿尔茨海默病检测的在线手写分段级风险发现
Gong, Changqing, Qin, Huafeng, El-Yacoubi, Mounîm A.
Abstract
Online handwriting provides a non-invasive and low-cost behavioral biomarker for Alzheimer's disease (AD) detection, as it reflects both cognitive planning and fine motor control. Existing handwriting-based AD detection methods usually rely on global trajectory features or whole-sample representations, which can be strongly affected by individual writing style, task-specific variation, and acquisition noise. In this paper, we propose NormPaST-Risk, a healthy-normative Paper-Air selective trajectory state-space risk network for interpretable AD detection from online handwriting. Instead of treating the entire trajectory as a single holistic representation, our method reformulates AD handwriting detection as local disease-relevant segment discovery. Specifically, a multi-scale temporal encoder captures stroke dynamics at different temporal resolutions, while a selective Paper-Air state-space encoder models long-range handwriting progression and distinguishes on-paper motor execution from in-air planning and transition behaviors. To explicitly characterize abnormal deviations, a healthy normative branch learns normal handwriting dynamics from healthy controls, and a task-aware multi-expert segment-risk module estimates segment-level AD risk calibrated by hidden-state changes and normative deviations. A weakly supervised segment-level objective further enables high-risk segment discovery without manual segment annotations. Experiments on the DARWIN benchmark demonstrate that the proposed framework achieves superior AD/HC classification performance compared with existing methods. Moreover, the discovered high-risk segments can be projected back to the original handwriting trajectory, providing interpretable evidence associated with AD-related handwriting variations.
Chinese Translation
在线手写为阿尔茨海默病(AD)检测提供了一种无创且低成本的行为生物标志物,因为它同时反映了认知规划与精细运动控制。现有的基于手写的AD检测方法通常依赖全局轨迹特征或整样本表示,容易受到个体书写风格、任务特异性差异以及采集噪声的强烈影响。本文提出NormPaST-Risk,一种基于健康规范的纸-空选择轨迹状态空间风险网络(healthy-normative Paper-Air selective trajectory state-space risk network),用于可解释的在线手写AD检测。与将整条轨迹视为单一整体表示不同,我们的方法将AD手写检测重新表述为局部疾病相关片段的发现问题。具体而言,多尺度时间编码器在不同时间分辨率下捕获笔画动态,而选择性的纸-空状态空间编码器则建模长程手写进程,并区分纸上运动执行与空中规划和过渡行为。为了显式刻画异常偏离,一个健康规范分支从健康对照中学习正常手写动态,一个任务感知的多专家片段风险模块则基于隐状态变化和规范性偏离来估计片段级的AD风险。弱监督的片段级目标进一步使得无需人工片段标注即可发现高风险片段。在DARWIN基准数据集上的实验表明,所提出的框架相比现有方法取得了更优的AD/HC(阿尔茨海默病/健康对照)分类性能。此外,所发现的高风险片段可投影回原始手写轨迹,提供与AD相关手写变化相关联的可解释证据。
cs.CV / 54 / 2609.29387

When Misalignment Becomes Supervision: Structured Label Noise in Supervised Synthetic CT Generation

当错位成为监督信号:监督式合成CT生成中的结构化标签噪声
Boussot, Valentin, Hemon, Cedric, Lafond, Caroline, Nunes, Jean-Claude, Dillenseger, Jean-Louis
Abstract
Supervised synthetic CT (sCT) generation is commonly trained and evaluated as voxel-wise regression against registered reference CT images. In practice, MRI-CT and CBCT-CT pairs are aligned through registration procedures that leave residual misalignments. These residuals are not independent intensity noise but spatially coherent geometric discrepancies that act as structured label noise. We investigate how this registration-induced bias affects supervised MRI-to-CT and CBCT-to-CT synthesis on 1,784 paired patients covering five anatomical regions. Voxel-wise scores strongly depend on the consistency between the registration used to build the training targets and the one used for evaluation: models score best when both conventions match, showing that networks partly learn the geometric convention of the registration pipeline and that standard metrics reward it. Training on more anatomically consistent registrations reduces prediction variability and improves out-of-distribution robustness, and CT-only controls show that registration alone produces metric errors in the range of top challenge submissions. To mitigate the limits of voxel-wise supervision, we introduce a perceptual loss computed in the feature space of a pretrained Segment Anything encoder. Compared with MAE-only and VGG-based objectives, it improves downstream segmentation and yields sharper, more structurally coherent sCT. Perceptual and voxel-wise metrics disagree under imperfect alignment and agree when the evaluation geometry is reliable. These results identify registration-induced bias as a central confounder in supervised sCT generation and argue for complementing voxel-wise agreement with anatomy-oriented evaluation criteria.
Chinese Translation
监督式合成CT(synthetic CT, sCT)生成通常以配准后的参考CT图像为目标,作为体素级回归任务进行训练和评估。在实践中,MRI-CT和CBCT-CT图像对通过配准流程对齐,但会残留错位。这些残留并非独立的强度噪声,而是空间上连贯的几何偏差,其作用相当于结构化标签噪声。我们在涵盖五个解剖区域、共1,784对患者数据上,研究了这种配准引起的偏差如何影响监督式MRI到CT和CBCT到CT的合成。体素级评分强烈依赖于构建训练目标所用的配准与评估所用配准之间的一致性:当两者约定一致时模型得分最高,这表明网络部分学习到了配准流程的几何约定,而标准指标对此予以奖励。在解剖学一致性更高的配准上进行训练可以降低预测变异性并提升分布外鲁棒性,而仅使用CT的对照实验表明,仅配准本身产生的指标误差就与顶级挑战赛提交结果处于同一量级。为缓解体素级监督的局限,我们引入了一种在预训练Segment Anything编码器特征空间中计算的感知损失。与仅使用MAE和基于VGG的目标函数相比,该损失改善了下游分割效果,并生成更锐利、结构上更连贯的sCT。在不完美对齐下,感知指标与体素级指标不一致;而当评估几何可靠时,两者趋于一致。这些结果将配准引起的偏差确定为监督式sCT生成中的核心混杂因素,并主张以面向解剖学的评估标准补充体素级一致性指标。
cs.CV / 55 / 2609.29433

Detecting Glaucoma Across Multi-ethnic Myopic and Non-Myopic Populations Using an Uncertainty-Aware Vision Transformer: A Multicentre Model Development and Validation Study

Lavanya, Raghavan, Feng, Yangqin, Quek, Ten Cheer, Hoang, Quan V., Poon, Linda Yi-Chieh, Jonas, Jost B., Wang, Ya Xing, Nangia, Vinay, Jeoung, Jin Wook, Park, Sehie, Kim, SoYeon, Xu, Benjamin Y, Munimadugu, Sreenidhi Iyengar, Mitchell, Paul, Liew, Gerald, Suwan, Yanin, Hong-amata, Jirayu, Thakur, Sahil, Nongipur, Monisha E, Wong, Tina, Husain, Rahat, Rui, Ng Si, Syn, Yamon, Lo, Phey Feng, Qiang, Nicholas Tan Yi, Hussain, Shaista, Lei, Xiaofeng, Da Soh, Zhi, Yu, Marco, Hamzah, Haslina, Wang, Zizhou, Wang, Yan, Zhen, Liangli, Xu, Xinxing, Wong, Tien-Yin, Aung, Tin, Chong, Rachel S, Liu, Yong, Cheng, Ching-Yu
Abstract
Background: Artificial intelligence (AI)-based glaucoma detection from colour fundus photographs (CFP) offers scalable screening, but performance may decline on external datasets because of differences in ground-truth definitions, populations, and coexisting conditions such as high myopia (HM). We developed and validated a Vision Transformer-based deep learning (DL) model for glaucoma detection across multi-ethnic cohorts with and without HM. Methods: A ViT-B/16 model with predictive uncertainty estimation was developed using 56,483 CFPs (57.1% with myopia; 14.4% with HM). Glaucoma labels were standardised using clinical, imaging, and perimetry data. The model was validated on 16 independent datasets across three continents, including four datasets with explicit HM labels. Findings: Internal AUROC was 98.7% (95% CI 98.2-99.1%), with sensitivity 94.5% and specificity 97.3%. Across 16 external datasets from eight countries, AUROCs ranged from 86.4% to 99.6%. In HM eyes, internal AUROC was 97.8% (95% CI 96.1-99.2%), with sensitivity 94.8% and specificity 93.7%. External HM AUROCs were 86.5% in the Beijing Eye Study and 93.3%, 91.8%, and 85.5% in hospital-based datasets from Taiwan, Thailand, and South Korea. In an exploratory HM clinical evaluation, the model had higher CFP-only diagnostic accuracy than ophthalmologists and trained graders (92.0% vs 70.0%; p=0.008) and performed comparably to glaucoma specialists using full clinical information. Interpretation: The model showed robust glaucoma detection across myopic and non-myopic multi-ethnic populations and may support AI-assisted screening in settings with high HM prevalence.
Chinese Translation
模型服务未提供译文(内容过滤),请阅读上方英文摘要。
cs.CV / 56 / 2609.29443

Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition

面向视觉语音识别的姿态自适应动态FiLM调制
Teng, Matthew Kit Khinn, Zhang, Haibo, Saitoh, Takeshi
Abstract
Head-pose variation introduces substantial appearance transformations in visual speech recognition (VSR), making pose-aware feature modulation desirable. However, performance degradation and unwanted feature interactions may result from using numerous Feature-wise Linear Modulation (FiLM) circuits with fixed modulation intensity. We propose a Pose Adaptive Dynamic FiLM framework with a Dynamic Residual FiLM (DR-FiLM) modulator that predicts input-dependent weights to adaptively control the strength of pose-conditioned modulation. Experiments on LRS2 and LRS3 demonstrate that unweighted multi-pathway modulation substantially degrades phoneme recognition, increasing PER to 20.33% and 29.42%, respectively, compared with 16.20% and 20.96% for the single ResFiLM configuration. In contrast, the proposed DR-FiLM with dynamic Deep-Res weighting reduces PER to 15.74% on LRS2 and 23.91% on LRS3, substantially mitigating the adverse effects of unweighted modulation. The analysis of the learned weights further reveals a consistent tendency to assign greater weight to the deeper FiLM pathway as head-pose variation increases. These results show that merging pose-conditioned FiLM circuits is more efficient when the modulation strength is dynamically controlled.
Chinese Translation
头部姿态变化会在视觉语音识别(VSR)中引入显著的外观变化,因此需要具备姿态感知的特征调制能力。然而,使用大量具有固定调制强度的特征级线性调制(FiLM)电路可能导致性能下降和不必要的特征交互。我们提出了一种姿态自适应动态FiLM框架,其中包含一个动态残差FiLM(DR-FiLM)调制器,该调制器能够预测依赖于输入的权重,以自适应地控制姿态条件调制的强度。在LRS2和LRS3数据集上的实验表明,未加权的多通路调制会显著降低音素识别性能,使音素错误率(PER)分别升至20.33%和29.42%,而单一ResFiLM配置的PER分别为16.20%和20.96%。相比之下,所提出的采用动态Deep-Res加权的DR-FiLM将LRS2和LRS3上的PER分别降至15.74%和23.91%,显著缓解了未加权调制的不利影响。对学习到的权重的分析进一步表明,随着头部姿态变化的增大,模型始终倾向于为更深的FiLM通路分配更大的权重。这些结果表明,在动态控制调制强度的情况下,融合姿态条件的FiLM电路会更加高效。
cs.CV / 57 / 2609.29447

Frame-to-Panorama Localization and Context-Aware Sampling for Scene-Specific Ship Detection in a Smart Marina Testbed

面向智慧游艇码头测试平台场景特定船舶检测的帧-全景图定位与上下文感知采样方法
Romanov, Ignat, Hadjipieris, Andreas, Dimitriou, Neofytos
Abstract
Smart maritime infrastructures provide continuous access to heterogeneous sensing streams, enabling repeated experimentation, digital-twin development, and AI-based maritime services. However, sensing hardware alone is not sufficient for scene-specific model development: historical video streams must also be spatially indexed, contextualized, and reduced to informative subsets for annotation. This paper presents a frame-to-panorama localization and context-aware sampling pipeline for ship detection in historical PTZ maritime video lacking reliable pan, tilt, and zoom metadata. The main contribution is an end-to-end data-curation approach that recovers camera-view information from historical PTZ video and combines it with environmental context and visual diversity to construct compact, scene-specific training sets. Specifically, frames are localized on a reference panorama using SuperPoint and LightGlue, enriched with weather and solar-state metadata, and selected through diversity sampling to preserve variation across camera view and environmental conditions. A second context-aware stage targets under-represented distant-vessel cases near the horizon using tile-level visual embeddings and Gaussian Mixture Model clustering. Applied within the CMMI MDigi-I Smart Marina testbed, the proposed pipeline reduces 40,718 candidate frames to 220 images for annotation, corresponding to a 99.5% reduction. A YOLO26-m detector fine-tuned on this subset achieves a mean AP50 of 94.78% $\pm$ 0.51% and a mean AP50-95 of 75.10% $\pm$ 1.73% under sequence-grouped five-fold cross-validation. These results demonstrate that highly redundant infrastructure video streams can be transformed into compact, spatially and contextually diverse training sets for scene-specific detector adaptation while substantially reducing annotation effort.
Chinese Translation
智慧海事基础设施提供了对异构感知数据流的持续访问能力,支持重复实验、数字孪生开发以及基于人工智能的海事服务。然而,仅有感知硬件并不足以支撑场景特定的模型开发:历史视频流还必须进行空间索引、上下文关联,并压缩为信息丰富的子集以便进行标注。本文提出了一种面向历史PTZ(云台)海事视频的帧-全景图定位与上下文感知采样流水线,适用于缺乏可靠的水平旋转、俯仰和变焦元数据的场景。其主要贡献是一种端到端的数据整理方法,该方法从历史PTZ视频中恢复相机视场信息,并结合环境上下文与视觉多样性,构建紧凑的场景特定训练集。具体而言,首先使用SuperPoint和LightGlue将视频帧在参考全景图上进行定位,然后融合天气和太阳状态等元数据进行增强,最后通过多样性采样选取帧,以保留不同相机视角和环境条件下的变化。第二阶段的上下文感知处理利用瓦片级视觉嵌入和高斯混合模型(GMM)聚类,针对地平线附近代表性不足的远距离船舶场景进行补充采样。在CMMI MDigi-I智慧游艇码头测试平台中应用该流水线后,40,718帧候选图像被压缩为220张待标注图像,缩减率达99.5%。在该子集上微调的YOLO26-m检测器,在基于序列分组的五折交叉验证下,平均AP50达到94.78% ± 0.51%,平均AP50-95达到75.10% ± 1.73%。这些结果表明,高度冗余的基础设施视频流可以被转化为空间和上下文均具多样性的紧凑训练集,用于场景特定检测器的适配,同时大幅降低标注工作量。
cs.CV / 58 / 2609.29456

Dense Coverage, Sparse Refinement: Byte-Constrained Cooperative Perception

密集覆盖、稀疏精化:字节受限的协同感知
Yazgan, Melih, Müller, Timon, Zöllner, J. Marius
Abstract
Collaborative perception improves autonomous perception by sharing intermediate Bird's-Eye-View (BEV) features across connected agents, but dense feature exchange is difficult to deploy under strict Vehicle-to-Everything (V2X) bandwidth limits. Existing efficient methods typically either compress the full feature map uniformly, spending bits on low-value background, or sparsify communication, risking the loss of useful context. We propose a coverage-refinement design for byte-constrained cooperative perception: each agent transmits a highly compressed coarse layer over the full BEV map and allocates the remaining budget to selected high-resolution patches. A Task-Aware Benefit Selector ranks cells by estimated downstream utility, enabling deterministic budgeted refinement and zero-retraining adaptation to changing bandwidth. The receiver reconstructs a dense BEV tensor compatible with standard fusion modules. Experiments on DAIR-V2X and OPV2V show strong accuracy-payload trade-offs at kilobyte-scale budgets. On DAIR-V2X, our method reaches 0.60 AP@0.7 at only 1.87 KB per non-ego agent, compared with 0.52 at 4.61 KB for uniform SimVQ compression. Controlled diagnostics further show that the gain arises from coverage-refinement allocation rather than quantization alone. Code will be published.
Chinese Translation
协同感知通过在互联智能体之间共享中间鸟瞰图(BEV)特征来提升自主感知能力,但在严格的车联网(V2X)带宽限制下,密集特征交换难以部署。现有高效方法通常要么对完整特征图进行均匀压缩,将比特花费在低价值背景上;要么采用稀疏化通信,面临丢失有用上下文的风险。我们提出一种面向字节受限协同感知的覆盖-精化设计:每个智能体在整个BEV地图上传输高度压缩的粗略层,并将剩余预算分配给选定的高分辨率图像块。任务感知收益选择器(Task-Aware Benefit Selector)根据估计的下游效用来对单元排序,从而实现确定性的预算化精化以及对带宽变化的无再训练适应。接收端重建与标准融合模块兼容的密集BEV张量。在DAIR-V2X和OPV2V数据集上的实验表明,该方法在千字节级预算下实现了优异的精度-负载权衡。在DAIR-V2X上,本方法在每辆非本车智能体仅1.87 KB的条件下达到0.60 AP@0.7,而均匀SimVQ压缩需要4.61 KB才达到0.52。受控诊断实验进一步表明,性能增益源自覆盖-精化的分配策略,而非仅靠量化。代码将会公开。
cs.CV / 59 / 2609.29457

Industrial Anomaly Detection via Defect-Grounded Reasoning in Visual Latent Space

基于视觉潜在空间中缺陷接地推理的工业异常检测
Yeh, Jaron, Chang, Yen-Wei, Liu, Jiang, Lo, Shao-Yuan
Abstract
Industrial anomaly detection (IAD) is evolving beyond conventional detection and localization toward multimodal inspection systems that can describe, explain, and reason about fine-grained defects. Although recent multimodal large language model (MLLM)-based methods improve anomaly understanding through textual reasoning and visual guidance, they face two limitations in fine-grained inspection. First, their visual refinement often requires iteratively revisiting local image regions or augmenting with additional tools. Second, the resulting local defect evidence may not be reliably preserved throughout subsequent reasoning. To address these, we propose Anomaly-LR, a defect-grounded latent reasoning framework that first forms a global understanding of the input and then progressively refines anomaly-relevant representations directly in the visual latent space. We further construct IAD-LR-22K, the first IAD instruction dataset designed for latent reasoning, containing 22,228 image-question instances from 4,523 industrial images, with global textual reasoning traces and region-level visual annotations. Extensive experiments show that Anomaly-LR achieves state-of-the-art performance among comparable-scale methods across multiple IAD benchmarks, without requiring external references or tools. The code and data will be released at https://github.com/Yen666/Anomaly-LR.
Chinese Translation
工业异常检测(IAD)正从传统的检测与定位向多模态检测系统演进,这类系统能够对细粒度缺陷进行描述、解释和推理。尽管近期基于多模态大语言模型(MLLM)的方法通过文本推理和视觉引导改进了异常理解能力,但它们在细粒度检测中仍存在两个局限。其一,其视觉精化过程通常需要迭代地重新访问局部图像区域或借助额外工具进行增强;其二,所获得的局部缺陷证据在后续推理过程中可能无法被可靠地保留。为解决这些问题,我们提出了Anomaly-LR,一种缺陷接地的潜在推理框架,该框架首先对输入形成全局理解,然后直接在视觉潜在空间中逐步精化与异常相关的表征。我们还构建了IAD-LR-22K,这是首个面向潜在推理的IAD指令数据集,包含来自4,523张工业图像的22,228个图像-问题实例,并附带全局文本推理轨迹和区域级视觉标注。大量实验表明,Anomaly-LR在多个IAD基准测试中,在同等规模的的方法中取得了最先进的性能,且无需外部参考或工具。代码和数据将发布于 https://github.com/Yen666/Anomaly-LR。
cs.CV / 60 / 2609.29460

AgriCountDINO: Parameter-Efficient Exemplar-Guided Counting and Localization in Agriculture

AgriCountDINO:农业中参数高效的样本引导计数与定位方法
Guo, Shengjie, Li, Xin, Arsova, Borjana, Scharr, Hanno, Salvi, Silvio
Abstract
Accurate counting and localization of plants and their organs support phenotyping and yield estimation, yet target appearance, scale, and density vary widely across species and imaging conditions. Exemplar boxes specify the target without category-specific retraining, and point predictions identify the individual instances contributing to the count. We introduce AgriCountDINO, a parameter-efficient exemplar-guided framework for joint counting and localization. It conditions frozen multiscale DINOv3 features on exemplar appearance and size, then progressively decodes them into target points. Missed-object recovery extends supervision to targets overlooked by initial matching, and exemplar-adaptive point NMS filters duplicate predictions according to exemplar scale. With 8.4M trainable parameters, approximately one-tenth of TasselNetV4's, AgriCountDINO achieves a three-shot MAE of 11.92 on the TPC-268 benchmark, reducing counting error by 9.7\% while providing individual target locations. Trained only on TPC-268, it achieves a zero-shot MAE of 14.25 on unseen generic object categories in FSC-147, improving upon the best compared zero-shot method by 6.0\% without target-domain training or fine-tuning.
Chinese Translation
植物及其器官的精确计数与定位为表型分析和产量估计提供支持,然而目标的外观、尺度和密度在不同物种和成像条件下差异巨大。样本框(exemplar box)无需针对特定类别重新训练即可指定目标,而点预测则可识别构成计数的各个实例。我们提出了 AgriCountDINO,一种用于联合计数与定位的参数高效样本引导框架。该方法以样本的外观和尺寸信息对冻结的多尺度 DINOv3 特征进行条件化,然后将其逐步解码为目标点。漏检目标恢复(missed-object recovery)机制将监督扩展到初始匹配中被忽略的目标,样本自适应点 NMS(exemplar-adaptive point NMS)则根据样本尺度过滤重复预测。AgriCountDINO 仅有 840 万可训练参数,约为 TasselNetV4 的十分之一,在 TPC-268 基准上实现了三样本(three-shot)MAE 为 11.92 的性能,将计数误差降低了 9.7%,同时提供各个目标的位置。仅在 TPC-268 上训练的模型,在 FSC-147 中未见过的通用目标类别上实现了零样本(zero-shot)MAE 为 14.25 的性能,相比最优的零样本对比方法提升了 6.0%,且无需目标域训练或微调。
cs.CV / 61 / 2609.29517

AdaPilot: Towards Scene-Adaptive Policy Learning for Cross-Generator Text-to-Image Quality Optimization

AdaPilot:面向跨生成器文生图质量优化的场景自适应策略学习
Liu, Wenjin, Ke, Fayuan, Lu, Yue, Cui, Zhe, Luu, Anh Tuan, Luo, Haoran
Abstract
Existing methods for improving text-to-image generation quality have progressed from generator fine-tuning and prompt optimization to reinforcement learning with multi-turn visual feedback. However, existing strategies are deeply coupled with specific generators and tasks, and the learned capabilities are difficult to generalize into a universal quality optimization policy. Therefore, we propose AdaPilot, which learns a scene-adaptive, cross-generator transferable quality optimization policy by formulating multi-turn image generation as a Markov Decision Process (MDP) and optimizing it via end-to-end reinforcement learning. Specifically, AdaPilot decouples the policy from generator internals to enable cross-generator transfer, introduces scene-aware rewards that adaptively align quality assessment dimensions with task semantics, and employs process-level rewards to model the evolution trajectory of image quality. Experimental results show AdaPilot outperforms baselines in generation quality and generalization. Separate cross-generator evaluations further show that a single policy transfers zero-shot to unseen generators while maintaining positive average gains across all evaluated generators. Our project is available at https://github.com/QwenQing/Ada_pilot.
Chinese Translation
现有的提升文本生成图像质量的方法已从生成器微调和提示词优化发展到基于多轮视觉反馈的强化学习。然而,现有策略与特定生成器和任务深度耦合,所学到的能力难以泛化为通用的质量优化策略。为此,我们提出AdaPilot,通过将多轮图像生成建模为马尔可夫决策过程(Markov Decision Process, MDP),并利用端到端强化学习进行优化,从而学习一种场景自适应、可跨生成器迁移的质量优化策略。具体而言,AdaPilot将策略与生成器内部机制解耦以实现跨生成器迁移,引入场景感知奖励以自适应地将质量评估维度与任务语义对齐,并采用过程级奖励来建模图像质量的演化轨迹。实验结果表明,AdaPilot在生成质量和泛化能力上均优于基线方法。单独的跨生成器评估进一步表明,单一策略能够以零样本方式迁移到未见过的生成器,同时在所有被评估的生成器上均保持正向的平均增益。我们的项目已在 https://github.com/QwenQing/Ada_pilot 上发布。
cs.CV / 62 / 2609.29527

CoSWA-YOLOv12: Scale-Invariant Tiny Object Detection and Segmentation of Malaria Parasites

CoSWA-YOLOv12:尺度不变的疟疾寄生虫微小目标检测与分割
Issah, Ahmed Tahiru, Mukamakuza, Carine
Abstract
Automated microscopy could widen access to malaria diagnosis in low-resource settings, but the deadliest species, P. falciparum, presents in its early ring stage as an object only a few tens of pixels wide. Such tiny targets are systematically under-detected: overlap-based label assignment starves them of positive samples, and overlap-based box regression gives weak gradients at their scale. The Normalized Gaussian Wasserstein Distance (NWD) repairs both effects, but applied uniformly across a slide that also holds objects three to four times larger it loosens their supervision and erodes their localisation, so overall accuracy can fall even as the tiny class improves. We present CoSWA-YOLOv12, a compact YOLOv12 instance-segmentation detector whose core Cooperative Scale-adaptive Wasserstein Assignment routes the Wasserstein treatment to an object in inverse proportion to its size, tapering back to standard assignment for larger species. Two further components support it: a wavelet detail residual, and a min-max Gaussian regression loss (M2-NWD). All three additions are transfer-safe: each reproduces the standard pretrained model exactly at initialisation, so public pretrained weights load without any loss of accuracy. On a five-class Rwandan thick-smear dataset, CoSWA-YOLOv12 raises P. falciparum recall from 0.63 to 0.74 and mAP@50 from 0.73 to 0.81 (mask), cuts missed P. falciparum from 38% to 15%, and improves strict-localisation mAP@50-95 on all five classes for both detection and segmentation, while a 2x2 ablation shows the scale gate and the regression loss are synergistic.
Chinese Translation
自动化显微镜检测可以拓宽低资源环境下疟疾诊断的可及性,但最致命的疟原虫种类——恶性疟原虫(P. falciparum)——在早期环状体阶段仅表现为宽度只有几十个像素的目标。此类微小目标往往被系统性漏检:基于重叠度的标签分配使其缺乏正样本,而基于重叠度的边界框回归在其尺度上提供的梯度也很微弱。归一化高斯Wasserstein距离(NWD)能够修复这两个问题,但若在包含比其大三到四倍目标的整个涂片上统一应用,则会削弱对较大目标的监督并损害其定位精度,因此即使微小目标类别性能提升,整体准确率仍可能下降。我们提出了CoSWA-YOLOv12,一个紧凑的YOLOv12实例分割检测器,其核心的协作式尺度自适应Wasserstein分配(Cooperative Scale-adaptive Wasserstein Assignment)按照与目标大小成反比的方式将Wasserstein处理分配给目标,并对较大的种类逐渐回归到标准分配方式。另外两个组件对其提供支持:小波细节残差模块和最小-最大高斯回归损失(M2-NWD)。这三项改进都具备迁移安全性:每一项在初始化时都能精确复现标准预训练模型,因此可直接加载公开的预训练权重而不损失任何精度。在一个包含五个类别的卢旺达厚血膜涂片数据集上,CoSWA-YOLOv12将恶性疟原虫的召回率从0.63提升至0.74,掩码mAP@50从0.73提升至0.81,将恶性疟原虫的漏检率从38%降至15%,并在检测和分割两项任务上提升了所有五个类别的严格定位mAP@50-95指标;同时,2×2消融实验表明尺度门控机制与回归损失具有协同增效作用。
cs.CV / 63 / 2609.29541

GeoRefer-Bench: A Benchmark from Referring Pixels to Verifiable Geospatial Reasoning

GeoRefer-Bench:一个从像素指称到可验证地理空间推理的基准
Cao, Shuaishuai, Huang, Min, Tang, Meng, Liu, Xuan, Wang, Youjin, Lin, Hui
Abstract
Referring segmentation in overhead imagery is inherently relational: a query may ask for the buildings north of the road or the pond closest to a residential area, so the correct referent can contain one object, several objects, or none. Existing benchmarks mainly score mask overlap, which cannot verify whether a model actually resolved the stated spatial relation. We introduce GeoRefer-Bench, a benchmark for verifiable geospatial referring segmentation. Each query is represented by an executable logical form over a metric scene graph, and predictions are evaluated with Exact Query Success (EQS), which is satisfied only when the returned instance set exactly matches the set denoted by the query. GeoRefer-Bench contains 700 whole 2048x2048 UAV scenes (2.94 Gpx) at 12.5 and 25 cm ground sampling distance, 26,217 instances, 142,796 spatial relations, and 20,916 executable queries spanning five reasoning levels. It further includes three paraphrases per query, 24.0% unanswerable queries, 2,477 counterfactual pairs, and five leakage-controlled evaluation splits. An independent audit re-derives object geometry, mask ownership, relation values, query execution, and split provenance, finding zero issues across all 700 scenes. Relation-blind strategies can retain non-trivial mIoU while achieving at most 22.7 EQS overall, showing that overlap alone does not certify relational grounding. Across fifteen current models, the strongest reaches 74.1 EQS but drops from 98.9 at level 1 to 60.5 at level 5, while ten models score below 5 EQS on two-hop queries. GeoRefer-Bench turns geospatial referring segmentation from mask matching into verifiable reference resolution.
Chinese Translation
遥感俯拍图像中的指称分割本质上是一种关系型任务:查询可能要求“道路北侧的建筑物”或“离居民区最近的池塘”,因此正确的指称对象可能包含一个目标、多个目标或无目标。现有基准主要评估掩码重叠度(mask overlap),无法验证模型是否真正解析了所述的空间关系。我们提出 GeoRefer-Bench,一个用于可验证地理空间指称分割的基准。每个查询由度量场景图上的可执行逻辑形式表示,并采用精确查询成功率(Exact Query Success, EQS)进行评估——只有当模型返回的实例集合与查询所表示的集合完全一致时,EQS 才算满足。GeoRefer-Bench 包含 700 幅完整的 2048x2048 无人机场景影像(共 2.94 Gpx),地面采样距离为 12.5 厘米和 25 厘米,包含 26,217 个实例、142,796 条空间关系,以及跨越五个推理层级的 20,916 个可执行查询。该基准还包含每个查询的三种改写形式、24.0% 的不可回答查询、2,477 个反事实配对,以及五个受泄漏控制的评估划分。一项独立审计重新推导了对象几何、掩码归属、关系取值、查询执行及划分来源,在全部 700 个场景中未发现任何问题。实验表明,忽略关系的策略(relation-blind strategies)仍可获得可观的 mIoU,但整体 EQS 至多仅达 22.7,说明仅凭重叠度并不能证明关系定位能力。在十五个当前模型中,表现最好的达到 74.1 EQS,但成绩从层级 1 的 98.9 下降到层级 5 的 60.5,而有十个模型在两跳查询上的 EQS 低于 5。GeoRefer-Bench 将地理空间指称分割从掩码匹配转变为可验证的指称解析。
cs.CV / 64 / 2609.29550

Spaceborne differential photogrammetry for control-free measurement of large-gradient deformation with structural immunity and a predictable accuracy envelope

星载差分摄影测量:具备结构免疫性与可预测精度包络的无控制点大梯度形变测量方法
Zhang, Yueqiang, Ma, Chang, Pan, Shuixin, Liu, Haibo
Abstract
Optical satellite image correlation measures wide-area deformation in regimes where coherent interferometric synthetic aperture radar fails because displacement gradients are too large. However, standard pairwise workflows lack a pre-acquisition error budget and rely on extensive stable terrain. We formulate repeat-pass optical correlation as a differential estimation problem without surveyed ground control. Nominal georeferencing defines the coordinate frame, stable-area constraints and displacement priors resolve the datum, and surface displacement is estimated jointly with inter-epoch revisit-bias coefficients. The model yields a predictive accuracy envelope and calibrated per-point posterior uncertainty, bounds along-track uncertainty through a displacement prior, and represents pushbroom jitter using per-line revisit offsets. Simulations and Sentinel-2 and WorldView-2 experiments on the 2019 Ridgecrest earthquake, the 2023 Kahramanmara\c{s} earthquake, and the Baltoro glacier validate the predicted noise floor, control-free accuracy margin, and leakage caused by view-angle and digital elevation model errors. The measured noise floor reaches approximately $0.05$ pixel at $10$,m ground sampling distance. With only five stable tiles, conventional destriping changes the estimated Baltoro trunk velocity from $106$ to $1251$myr$^{-1}$, whereas the prior-constrained estimate remains $87$myr$^{-1}$. Closure analysis attributes approximately $88\%$ of pair-error variance to individual scenes, consistent with $25{,}354$ ITS_LIVE glacier-velocity triplets. Three matching methods lead to the same conclusions. The framework therefore turns pairwise correlation into a robust measurement with a predictive error budget, reduced dependence on stable terrain, and conclusions independent of the matching method.
Chinese Translation
光学卫星影像相关测量可在位移梯度过大导致相干干涉合成孔径雷达失效的情形下测量大范围形变。然而,标准的两两配对处理流程缺乏观测前的误差预算,且依赖大量稳定地形。我们将重复轨道光学相关观测构建为一个无需实测地面控制点的差分估计问题。标称地理配准定义坐标框架,稳定区域约束与位移先验解决基准问题,地表位移则与历元间的重访偏差系数进行联合估计。该模型可给出预测性精度包络与经过标定的逐点后验不确定度,通过位移先验约束沿轨方向的不确定度,并利用逐行重访偏移量来表征推扫式传感器的抖动。基于2019年Ridgecrest地震、2023年Kahramanmaraş地震以及Baltoro冰川的模拟实验与Sentinel-2和WorldView-2数据实验,验证了预测的噪声底限、无控制点精度裕度,以及由视角和数字高程模型误差引起的泄漏。在10米地面采样距离下,实测噪声底限约为0.05像素。当仅有五个稳定图块时,传统去条带方法使Baltoro冰川主干流速估计值从106米/年变为125米/年,而先验约束估计值则稳定保持在87米/年。闭合分析将约88%的像对误差方差归因于单景影像,这与25,354组ITS_LIVE冰川流速三元组的结果一致。三种匹配方法均得出相同结论。因此,该框架将两两影像相关转化为一种具有预测性误差预算、降低对稳定地形依赖、且结论独立于匹配方法的稳健测量手段。
cs.CV / 65 / 2609.29553

UNWIND: Any-Length Facial Video for Stress Detection without Temporal Windowing

UNWIND:无需时间窗口划分的任意长度面部视频压力检测
Gkikas, Stefanos, Cruz, Christian Arzate, Nichols, Eric, Giannakakis, Giorgos, Gomez, Randy
Abstract
Automatic stress recognition from facial video provides a non-contact approach for affective monitoring. However, most existing video-based methods divide complete recordings into shorter temporal segments before performing classification. Such segmentation requires additional decisions concerning segment duration, overlap, and prediction aggregation, and may restrict the model from exploiting information distributed across the entire recording. We introduce UNWIND, a facial-video framework for stress detection that analyzes a complete recording as a single model input, eliminating the need for temporal windowing or external segmentation. UNWIND reorganizes the video by folding its temporal dimension into the channel dimension of a two-dimensional spatial representation, which is subsequently processed through a unified asymmetric-attention architecture. With a temporal stride of $\tau=1$, the framework processes the entire $120$-second sequence, corresponding to $3{,}600$ frames sampled at $30$~fps, in a single input. We evaluate seven temporal-stride settings on a stress dataset comprising $58$ subjects, using a stratified subject-level protocol that covers configurations from dense frame retention to sparse temporal sampling. The highest test accuracy, $70.02\%$, is obtained at $\tau=15$, while processing all frames at $\tau=1$ achieves a comparable accuracy of $69.73\%$. Computational requirements range from $12.48$ to $348.78$ GFLOPs across the evaluated stride settings, illustrating the balance between temporal sampling density and computational efficiency. The findings show that effective facial-video stress recognition can be achieved without dividing recordings into temporal windows and that complete-recording inference can be performed within a single unified model.
Chinese Translation
基于面部视频的自动压力识别为情感监测提供了一种非接触式方法。然而,现有大多数基于视频的方法在执行分类之前,会将完整的录像划分为较短的时间片段。这种分段方式需要对片段时长、重叠程度和预测聚合方式做出额外决策,并且可能限制模型利用分布在整段录像中的信息。我们提出了UNWIND,一个用于压力检测的面部视频框架,它将完整录像作为单一模型输入进行分析,从而无需时间窗口划分或外部分段。UNWIND通过将视频的时间维度折叠进二维空间表示的通道维度来重新组织视频,随后通过统一的非对称注意力(asymmetric-attention)架构进行处理。在时间步长 $ au=1$ 的情况下,该框架可以在单一输入中处理整段 $120$ 秒的序列,即以 $30$~fps 采样率采集的 $3{,}600$ 帧。我们在一个包含 $58$ 名受试者的压力数据集上评估了七种时间步长设置,采用分层受试者级协议,涵盖从密集帧保留到稀疏时间采样的各种配置。最高测试准确率 $70.02\%$ 在 $ au=15$ 时获得,而在 $ au=1$ 时处理所有帧也达到了相当的准确率 $69.73\%$。在所评估的步长设置下,计算量范围为 $12.48$ 到 $348.78$ GFLOPs,体现了时间采样密度与计算效率之间的平衡。研究结果表明,无需将录像划分为时间窗口即可实现有效的面部视频压力识别,并且完整录像推理可以在单一统一模型内完成。
cs.CV / 66 / 2609.29555

Visual Representation and History Modeling for Navigation World Models

面向导航世界模型的视觉表示与历史建模
Guo, Guangfu, Lu, Xiaoqian, Liu, Rui, Chen, Yutong, Liu, Kunpeng, Cheng, Long
Abstract
Navigation World Models (NWMs) predict action-conditioned visual futures for planning. Two practical challenges are central to their design: selecting a suitable visual representation and efficiently modeling observation history for repeated candidate queries. Standard Global-Softmax attention provides flexible interactions but repeatedly processes the same history, leading to increasing computation and memory costs for long contexts and multi-query planning. We study both problems within a unified conditional flow-transformer framework. We first compare five frozen visual representations under the same dynamics model and evaluation. To reduce redundant history computation, we design Cached-Linear, a hybrid architecture that combines local and shifted-window attention for target mixing with linear attention for reusable history access. We further develop Balanced Gated Delta Network (GDN), which augments this design with frame-wise recurrent memory for temporal history modeling. Experiments on RECON, SACSoN, and SCAND show that representation choice depends on the prediction objective: PAE-L performs best for reconstruction, RAE-B for direct prediction, and V-JEPA for long-horizon rollout. Under shared-history workloads, Cached-Linear substantially reduces computation and memory compared with Global-Softmax, while Balanced GDN improves selected direct-prediction endpoints with efficient context reuse. Overall, we systematically study visual representation and history modeling for NWMs and develop hybrid reusable-history architectures for efficient long-context and multi-query prediction.
Chinese Translation
导航世界模型(Navigation World Models, NWMs)通过预测以动作为条件的视觉未来来支持规划。其设计面临两个关键的实际挑战:选择合适的视觉表示,以及为重复的候选查询高效建模观测历史。标准的Global-Softmax注意力机制提供了灵活的交互能力,但会重复处理相同的历史信息,导致长上下文和多查询规划下的计算与内存开销不断增加。我们在一个统一的条件流-Transformer框架内对这两个问题展开研究。首先,我们在相同的动力学模型和评估条件下比较了五种冻结的视觉表示。为减少冗余的历史计算,我们设计了Cached-Linear,这是一种混合架构,将用于目标混合的局部注意力与移位窗口注意力,同用于可复用历史访问的线性注意力相结合。我们进一步提出了平衡门控Delta网络(Balanced Gated Delta Network, GDN),通过逐帧的循环记忆来增强该设计的时间历史建模能力。在RECON、SACSoN和SCAND数据集上的实验表明,表示的选择取决于预测目标:PAE-L在重建任务中表现最佳,RAE-B在直接预测中表现最佳,而V-JEPA则在长时程推演中表现最佳。在共享历史的工作负载下,Cached-Linear相比Global-Softmax显著降低了计算与内存开销,同时平衡GDN借助高效的上下文复用改进了部分直接预测的端点指标。总体而言,我们系统性地研究了NWMs的视觉表示与历史建模问题,并开发了用于高效长上下文与多查询预测的混合式可复用历史架构。
cs.CV / 67 / 2609.29581

Long-Tail Adaptive Flow Matching with Explicit Conditional Consistency Guidance for Precise Multimodal Face Synthesis

基于显式条件一致性引导的长尾自适应流匹配精确多模态人脸合成
Cao, Yushe, Zou, Xuechao, Xi, Xing, Shi, Dianxi, Yu, Chun, Xing, Junliang
Abstract
Although diffusion-based methods have substantially improved the controllability of multimodal face synthesis, their semantic alignment remains suboptimal because most existing approaches rely on implicit latent-space objectives to model the relationship between denoising variables and multimodal conditions. Such implicit modeling is often insufficient to enforce precise correspondence between synthesized faces and conditional inputs, especially under long-tailed semantic mask distributions where rare attributes receive weak optimization signals. To address these limitations, we propose EC\textsuperscript{2}Face, a multimodal face synthesis framework that improves semantic alignment through explicit semantic supervision and distribution-aware optimization. First, we introduce Explicit Conditional Consistency Guidance (ECCG), which imposes direct consistency supervision in pixel space by decoding an approximate reverse estimate of the clean latent and explicitly aligning the synthesized image with textual descriptions and semantic masks. A temporal dynamic modulation function is further designed to adapt the supervision strength according to the timestep-dependent reliability of reverse estimation. Second, we propose Long-Tail Adaptive Flow Matching (LAFM), which reweights spatial optimization signals based on semantic attribute frequency, with normalized weights to maintain numerical stability during training. Importantly, all additional modules are used only during training and introduce no extra inference overhead. Extensive experiments show that EC\textsuperscript{2}Face consistently outperforms competitive baselines in both generation quality and semantic alignment, achieving a 29.38\% improvement in mask accuracy on rare attributes.
Chinese Translation
尽管基于扩散的方法已显著提升了多模态人脸合成的可控性,但其语义对齐仍不够理想,因为大多数现有方法依赖隐式的潜空间目标来建模去噪变量与多模态条件之间的关系。这种隐式建模往往难以保证合成人脸与条件输入之间的精确对应,尤其是在长尾语义掩码分布下,稀有属性只能获得微弱的优化信号。为解决这些局限,我们提出EC\textsuperscript{2}Face,一个通过显式语义监督和分布感知优化来改进语义对齐的多模态人脸合成框架。首先,我们引入显式条件一致性引导(Explicit Conditional Consistency Guidance, ECCG),通过解码干净潜变量的近似反向估计,在像素空间中施加直接的一致性监督,显式地将合成图像与文本描述和语义掩码对齐。我们还进一步设计了时间动态调制函数,根据反向估计随时间步变化的可靠性自适应调整监督强度。其次,我们提出长尾自适应流匹配(Long-Tail Adaptive Flow Matching, LAFM),基于语义属性频率对空间优化信号进行重加权,并采用归一化权重以保持训练过程中的数值稳定性。重要的是,所有附加模块仅在训练阶段使用,不会带来额外的推理开销。大量实验表明,EC\textsuperscript{2}Face在生成质量和语义对齐方面均持续优于有竞争力的基线方法,在稀有属性的掩码准确率上取得了29.38\%的提升。
cs.CV / 68 / 2609.29591

CATCH: Counterfactual Anatomical Tissue Inpainting with Conditional Haar Diffusion

CATCH:基于条件Haar扩散的反事实解剖组织修复
Albertsen, Simon Winther, Bjoernstrup, Hjalte, Said, Said Djafar, Ghazi, Mostafa Mehdipour
Abstract
BraTS local synthesis replaces masked regions in T1-weighted brain MRI with plausible tumor-free tissue while preserving observed anatomy. We present CATCH, conditional 3D diffusion in an invertible Haar-wavelet domain. Its denoiser receives noisy target coefficients, voided-image coefficients, and a signed mask; tumor-excluded wavelet reconstruction and a hole-focused loss guide training, and hard compositing preserves observed voxels. We compare fixed masks, tumor-component augmentation, and a weighted mixture of tumor-derived, irregular-blob, and ellipsoidal masks. Of 25 development cases, five prespecified cases select each arm's checkpoint and all 25 of their trajectory aggregations; a separate 75-case internal set compares the frozen pipelines and selects a weighted mixture for organizer evaluation. Five-trajectory averaging yielded internal SSIM/PSNR/MSE (mean$\pm$SD) of $0.80\pm0.13$, $19.18\pm1.80$dB, and $0.010\pm0.005$. As the sole officially evaluated pipeline, weighted mixture yielded $0.772\pm0.119$, $20.89\pm3.27$dB, and $0.0098\pm0.0054$ on the 219-case BraTS 2026 validation set. Against compute-matched random augmentation internally, it improved SSIM by 0.019 (95% bootstrap CI: 0.013-0.025), PSNR by 0.95dB, and MSE by 0.003; all three paired comparisons remained significant after Holm correction. Results favor the complete weighted-mixture policy within CATCH; absent official fixed- and random-pipeline scores and a directly comparable external baseline limit broader conclusions.
Chinese Translation
BraTS局部合成任务旨在将T1加权脑部MRI中被掩膜的区域替换为合理的无肿瘤组织,同时保留已观测的解剖结构。我们提出了CATCH,一种在可逆Haar小波域中运行的条件3D扩散模型。其去噪器接收含噪目标系数、空洞图像系数和带符号掩膜;采用排除肿瘤的小波重构和聚焦空洞的损失函数指导训练,并通过硬合成保留已观测的体素。我们比较了固定掩膜、肿瘤成分增强,以及由肿瘤衍生掩膜、不规则团块掩膜和椭球形掩膜构成的加权混合掩膜策略。在25个开发病例中,五个预先指定的病例用于选择各方法的检查点及其全部25种轨迹聚合方式;另一个包含75个病例的内部集用于比较冻结的流水线,并选出加权混合策略提交给组织方评估。五轨迹平均在内部集上达到的SSIM/PSNR/MSE(均值±标准差)为0.80±0.13、19.18±1.80dB和0.010±0.005。作为唯一被官方评估的流水线,加权混合策略在219个病例的BraTS 2026验证集上取得了0.772±0.119、20.89±3.27dB和0.0098±0.0054的结果。在内部与计算量匹配的随机增强相比,该策略将SSIM提高了0.019(95%自助法置信区间:0.013-0.025),PSNR提高0.95dB,MSE降低0.003;经Holm校正后,三项配对比较仍保持统计显著。结果支持在CATCH中采用完整的加权混合掩膜策略;但由于缺乏官方的固定掩膜和随机掩膜流水线得分,以及缺乏可直接比较的外部基线,更广泛的结论受到限制。
cs.CV / 69 / 2609.29592

QINA: Quantum-Inspired Nonlinear Adapters for Pretrained Vision Models

QINA:面向预训练视觉模型的量子启发性非线性适配器
Ghazi, Mostafa Mehdipour
Abstract
Adapting large pretrained vision models under limited data and frozen-backbone constraints remains a central challenge in transfer learning. While lightweight adapters and parameter-efficient fine-tuning methods are widely adopted, most rely on generic multilayer perceptrons or low-rank linear updates, offering limited control over the spectral and geometric structure of feature transformations. We investigate whether structured nonlinear feature lifting can improve representational alignment in frozen regimes. We introduce Quantum-Inspired Nonlinear Adapters (QINA), compact modules that perform learnable trigonometric feature lifting followed by bounded nonlinear aggregation. The design induces structured oscillatory basis functions with an explicit norm-dependent Lipschitz bound, enabling spectral reshaping of pretrained representations without increasing the receptive field or significantly expanding parameter count. Importantly, the method operates entirely within standard deep learning frameworks and does not require quantum hardware. Through systematic experiments across natural and medical imaging datasets, classification and segmentation tasks, multiple adapters and placements, and varying training budgets, we show that performance in frozen regimes is primarily representation-limited. Nonlinear lifting improves adaptation, and the proposed structured trigonometric formulation consistently outperforms identity baselines, fixed Fourier feature mappings, and parameter-matched baseline adapters. Within the evaluated frozen-backbone settings, structured spectral parameterization provides a more effective inductive bias than generic nonlinear adapters. This work highlights the importance of geometry- and spectrum-aware adaptation mechanisms for large pretrained vision models.
Chinese Translation
在数据有限且骨干网络冻结的约束下,对大型预训练视觉模型进行适配仍然是迁移学习中的核心挑战。尽管轻量级适配器和参数高效微调方法已被广泛采用,但大多数方法依赖于通用的多层感知机或低秩线性更新,对特征变换的谱结构和几何结构缺乏有效控制。我们研究了结构化非线性特征提升能否在冻结场景下改善表示对齐。我们提出了量子启发性非线性适配器(Quantum-Inspired Nonlinear Adapters, QINA),这是一种紧凑的模块,通过可学习的三角函数特征提升,再结合有界非线性聚合。该设计引入了具有显式范数依赖Lipschitz界的结构化振荡基函数,能够在不扩大感受野或显著增加参数量的情况下,对预训练表示进行谱重塑。重要的是,该方法完全在标准深度学习框架内运行,无需量子硬件。通过在自然图像和医学影像数据集、分类与分割任务、多种适配器及其放置位置以及不同训练预算上的系统性实验,我们表明冻结场景下的性能主要受表示能力限制。非线性特征提升能够改善适配效果,且所提出的结构化三角函数形式持续优于恒等基线、固定傅里叶特征映射以及参数量匹配的基线适配器。在所评估的冻结骨干网络设置中,结构化谱参数化比通用非线性适配器提供了更有效的归纳偏置。这项工作凸显了几何与谱感知的适配机制对大型预训练视觉模型的重要性。
cs.CV / 70 / 2609.29604

PROVE: Proof-guided Regime-aware Operator Verification for Hallucination Detection in Medical Visual Question Answering

PROVE:面向医学视觉问答中幻觉检测的证明引导、感知验证机制的算子验证方法
Zhou, Keyang, Li, Siyi, Shi, Zhongnan, Ying, Qichao, Tang, Wei, Qian, Zhenxing
Abstract
In medical visual question answering (VQA), hallucinations of vision-language models (VLMs) may lead to confident but incorrect responses, raising the risk of diagnostic errors. Existing hallucination detection methods uniformly estimate the reliability of VLM outputs from response consistency or visual evidence. However, such uniform verification across questions ignores question-specific characteristics, resulting in missed overconfident errors and false alarms from over-verification. We present PROVE (Proof-guided Regime-aware Operator Verification), a black-box detector that adapts verification strategy to the evidential structure of each question. PROVE classifies questions into three verification regimes based on what kind of visual proof they demand, activates a regime-specific subset of five complementary operators, and adjusts operator importance per question through a lightweight calibration layer conditioned on deterministic question-answer features. PROVE uses question-specific evidence to reweight operators and produce a calibrated risk score. Evaluated on 8048 test samples across three medical VQA benchmarks and four frontier VLMs, PROVE achieves 0.821 AUROC, outperforming the strongest baseline by +0.159, with consistent gains across all models and benchmarks.
Chinese Translation
在医学视觉问答(VQA)中,视觉语言模型(VLM)的幻觉可能导致模型给出自信但错误的回答,从而增加诊断错误的风险。现有的幻觉检测方法通常从回答一致性或视觉证据出发,统一地估计VLM输出的可靠性。然而,这种对所有问题采用统一验证的方式忽略了问题本身的特定特征,导致遗漏过度自信的错误,或因过度验证而产生误报。我们提出了PROVE(Proof-guided Regime-aware Operator Verification,证明引导的感知验证机制的算子验证),这是一种黑盒检测器,能够根据每个问题所需的证据结构自适应地调整验证策略。PROVE根据问题所需的视觉证明类型将其划分为三种验证机制(regime),激活五种互补算子中与该机制对应的子集,并通过一个以确定性问题-回答特征为条件的轻量级校准层,针对每个问题调整算子的重要性。PROVE利用问题特定的证据对算子进行重新加权,并生成经校准的风险分数。在三个医学VQA基准和四个前沿VLM上共8048个测试样本的评估中,PROVE达到了0.821的AUROC,比最强基线高出0.159,且在所有模型和基准上均取得一致的提升。
cs.CV / 71 / 2609.29607

STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models

STRAND:视频大语言模型中以对象为中心的时空监测的基准测试与改进
Nguyen, Thong, Cao, Tri, Le, Khoi, Nguyen, Cong-Duy, Vo, Quynh, Ng, See-Kiong, Kuen-Yew, Bryan Hooi
Abstract
While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure in spatio-temporal monitoring, the ability to persistently track object identities, states, and relations over time. Existing benchmarks obscure this deficit by relying on single final-answer evaluations for queries that can often be resolved via local visual cues or statistical priors. To rigorously diagnose this, we introduce STRAND, a benchmark of human-verified object-centric facts that evaluates intermediate reasoning by decomposing queries into sub-questions, distinguishing genuine temporal understanding from coincidental correctness. Crucially, we score models with Faithful Accuracy, an unconditional joint metric that credits a prediction only when the target answer and every prerequisite sub-question are correct, so that a model cannot inflate its score by being selectively consistent on the small subset of targets it happens to answer correctly. To address failure modes exposed by STRAND, we further propose an object-centric framework that explicitly constructs and reasons over structured object trajectories via chunk-wise state extraction and temporal aggregation. Extensive experiments, including backbone-, frame-, call-, and token-matched comparisons against both end-to-end MLLMs and modular video harnesses, demonstrate that our object-centric framework significantly reduces hallucinated answers and improves spatio-temporal reasoning consistency over state-of-the-art MLLMs. The code, model, and data have been made available at nguyentthong.github.io/strand.
Chinese Translation
尽管多模态大语言模型(MLLMs)在视频理解方面取得了进展,但它们在动态场景中仍然极易产生幻觉。我们认为,这源于时空监测能力的缺失,即持续追踪对象身份、状态及其关系随时间变化的能力。现有基准测试通过对查询采用单一的最终答案评估来掩盖这一缺陷,而这些查询往往可以通过局部视觉线索或统计先验解决。为了严格诊断这一问题,我们提出了STRAND,一个经人工验证的以对象为中心的事实基准,通过将查询分解为子问题来评估中间推理过程,从而区分真正的时间理解与偶然的正确性。至关重要的是,我们采用忠实准确率(Faithful Accuracy)对模型进行评分,这是一种无条件的联合指标,只有当目标答案以及每一个前置子问题都正确时,预测才会得分,从而防止模型通过仅在碰巧回答正确的小部分目标上选择性保持一致来虚增分数。为解决STRAND所揭示的失败模式,我们进一步提出了一个以对象为中心的框架,通过分块状态提取和时间聚合,显式地构建结构化对象轨迹并对其进行推理。大量实验,包括与端到端MLLMs和模块化视频工具包在骨干网络、帧数、调用次数和token数量对齐条件下的对比,表明我们的以对象为中心的框架显著减少了幻觉答案,并在时空推理一致性方面优于最先进的MLLMs。相关代码、模型和数据已在 nguyentthong.github.io/strand 发布。
cs.CV / 72 / 2609.29621

AgenticCADedit: A Stateful, Tool-Mediated Agentic Approach to Multimodal 3D CAD Editing

AgenticCADedit:一种有状态的、工具介导的智能体化多模态三维CAD编辑方法
Sinha, Saptarshi Neil, Goschke, Mika Silvan, Kühn, Paul Julius, Kuijper, Arjan, Weinmann, Michael
Abstract
Computer-aided design is central to industrial manufacturing, and much of a designer's daily work consists of editing existing models from multimodal requests involving speech, sketches, and model interaction. Existing neural CAD approaches focus predominantly on unconditional or text-conditioned generation. The neuralCAD-Edit approach formalizes expert multimodal editing requests, but its iterative baseline refines a complete CAD program across attempts, executing each attempt from the original model in a stateless CAD environment. Every attempt must therefore reconstruct the entire edit from scratch, so partially correct progress is discarded rather than accumulated, and the model can neither inspect the geometry it has just produced nor selectively revert a single faulty operation. We present AgenticCADedit, which turns editing into a sequence of small, verifiable actions on a persistent CAD state instead of a single regenerated program. Rather than emitting one complete program, it applies incremental code steps that each commit to the session, inspects the resulting faces and edges, renders highlighted selections to verify that the intended region was addressed, and reverts individual operations when it was not. Subsequent actions therefore build on the geometry produced by earlier ones. Our approach improves on all metrics for all three evaluated LLMs (open-weight: qwen3.6-27b, gemma4-31b; proprietary: gpt-5.6-luna), with the largest gains for the weakest baseline model, qwen3.6-27b, whose validity rises from 51.0% to 94.8% and acceptance from 1.6% to 12.0%. A token-cost analysis with gpt-5.6-luna further shows $66.7$% fewer output tokens than neuralCAD-Edit, while $94.8$% of input tokens are served from the prompt cache.
Chinese Translation
计算机辅助设计(CAD)是工业制造的核心,设计师的日常工作很大程度上是根据涉及语音、草图和模型交互的多模态请求来编辑现有模型。现有的神经CAD方法主要集中于无条件生成或文本条件生成。neuralCAD-Edit方法对专家级多模态编辑请求进行了形式化,但其迭代式基线在多次尝试中不断完善一个完整的CAD程序,并在无状态的CAD环境中从原始模型出发执行每次尝试。因此,每次尝试都必须从零开始重建整个编辑过程,部分正确的进展会被丢弃而非累积,且模型既无法检查其刚刚生成的几何体,也无法选择性地撤销单个错误的操作。我们提出了AgenticCADedit,它将编辑转化为在持久化CAD状态上执行的一系列小型的、可验证的操作,而不是一次性重新生成整个程序。它并非输出一个完整的程序,而是应用增量式的代码步骤,每个步骤都提交到会话中;随后检查生成的面和边,并渲染高亮显示的选区以验证目标区域是否得到处理,若未处理则撤销单个操作。因此,后续操作建立在先前操作所产生的几何体之上。我们的方法在所有评估指标上均优于基线,适用于所评估的三个大语言模型(开源权重模型:qwen3.6-27b、gemma4-31b;专有模型:gpt-5.6-luna),其中对最弱的基线模型qwen3.6-27b提升最大:其有效性从51.0%提升至94.8%,接受率从1.6%提升至12.0%。基于gpt-5.6-luna的token成本分析进一步表明,与neuralCAD-Edit相比,我们的方法输出token减少66.7%,同时94.8%的输入token可由提示缓存提供。
cs.CV / 73 / 2609.29638

SpectralCTGaussians: Projection-Domain Reconstruction and Basis Material Decomposition for Spectral CT using 3D Gaussian Splatting

SpectralCTGaussians:基于3D高斯泼溅的光谱CT投影域重建与基材料分解
Vos, Reinout, Sinha, Saptarshi Neil, Weinmann, Michael
Abstract
Spectral computed tomography (CT) extends conventional CT by measuring attenuation across multiple energy channels, allowing improved modeling of physical X-ray interactions and energy-dependent material behavior and leading to richer scene understanding. We present a novel method for spectral CT reconstruction and basis material decomposition using 3D Gaussian Splatting by adding per-Gaussian basis material fractions to the set of learnable parameters, which together with a set of energy-dependent basis functions define the attenuation across the full spectral range. The basis functions represent various physical attenuation models such as photoelectric absorption and Compton scattering, and are jointly optimized across all energy channels through a differentiable polychromatic forward model, with material decomposition performed via mean-shift clustering of the resulting coefficients. We evaluate our method on a baseline real-world dataset as well as a synthetic dataset that we introduce, comparing against traditional reconstruction algorithms and state-of-the-art learning-based CT reconstruction methods. Our approach outperforms all traditional baselines in novel view synthesis and achieves the best PSNR among all compared methods for spectral CT volume reconstruction, while describing all energy channels with a single shared representation that requires a number of Gaussians comparable to single-channel Gaussian splatting-based CT reconstruction approaches. For basis material decomposition, no traditional or learning-based baseline offers one-step decomposition with direct RGB material segmentation, and our method additionally recovers the photoelectric basis with higher PSNR than traditional pipelines.
Chinese Translation
光谱计算机断层扫描(CT)通过在多个能量通道上测量衰减来扩展传统CT,从而能够更好地建模X射线物理相互作用及能量相关的材料行为,实现对场景更丰富的理解。我们提出了一种基于3D高斯泼溅(3D Gaussian Splatting)的光谱CT重建与基材料分解新方法,通过在每个高斯基元上添加可学习的基材料比例参数,与一组能量相关的基函数共同定义整个光谱范围内的衰减。这些基函数表示各种物理衰减模型,如光电吸收和康普顿散射,并通过可微分的多色前向模型在所有能量通道上联合优化,随后对所得系数进行均值漂移(mean-shift)聚类以完成材料分解。我们在一个真实场景基线数据集以及我们引入的合成数据集上评估了该方法,并与传统重建算法及最先进的基于学习的CT重建方法进行比较。我们的方法在新视角合成任务上优于所有传统基线方法,并在所有比较方法中于光谱CT体重建中取得了最佳PSNR,同时通过单一共享表示描述所有能量通道,所需高斯基元数量与单通道基于高斯泼溅的CT重建方法相当。对于基材料分解,没有任何传统或基于学习的基线方法能够提供结合直接RGB材料分割的一步式分解,而我们的方法还能以高于传统流程的PSNR恢复光电基。
cs.CV / 74 / 2609.29648

Albireo: Adaptive, Energy-Efficient Inference Framework for Video Object Detection on the Edge

Albireo:面向边缘端视频目标检测的自适应、高能效推理框架
Taherin, Amir, Cano, José, Ren, Bin, Wang, Yanzhi, Kaeli, David
Abstract
Video object detection on edge devices runs computationally expensive detectors over long frame streams, causing high energy consumption and sustained GPU utilization. Although consecutive frames are highly redundant, naive frame skipping is content-blind: it skips during critical moments such as object entry, occlusion recovery, and abrupt motion, degrading detection quality. We present Albireo, a detector-agnostic, codec-free adaptive inference framework that wraps off-the-shelf detectors and decides when detector invocation can be safely skipped based on scene content and per-object temporal state, requiring no detector modification or retraining. Albireo maintains a 10-dimensional Kalman filter (KF) per active object and invokes the detector only when prediction uncertainty exceeds a threshold; on skipped frames, boxes are predicted from the KF state at near-zero GPU cost. A KF-based rescue mechanism preserves confirmed objects through brief detector misses to prevent output fragmentation, while a lightweight empty-scene screen avoids detector calls on objectless frames. We evaluate Albireo on the BDD100K MOT validation split with three architecturally distinct detectors (YOLO11x, YOLO26x, RF-DETR-Large) on two NVIDIA Jetson platforms (AGX Thor, AGX Orin). Across all configurations, Albireo keeps AP@50 within +/-1.2 pp of per-frame inference while reducing total energy by 12.1-17.6%. On YOLO26x, it improves AP@50 by +0.8 pp while reducing energy by 17.6% (Thor) and 14.4% (Orin) and per-frame energy-delay product by 24.9% and 26.1%, respectively. Thus, the default operating point improves accuracy, energy, and latency together. In contrast, FixedSkip-2, a fixed-interval baseline with a 50% skip rate, loses 8.6 pp AP@50. Source code, evaluation pipeline, and per-clip results are available at https://github.com/amirtaherin/albireo
Chinese Translation
边缘设备上的视频目标检测需要在长时间视频帧流上运行计算开销高昂的检测器,导致能耗高且GPU持续处于高利用率。尽管连续帧之间存在高度冗余,但朴素的跳帧方法是内容盲目的:它会在物体进入、遮挡恢复和剧烈运动等关键时刻跳帧,从而降低检测质量。我们提出Albireo,一个与检测器无关、无需编解码器的自适应推理框架,它封装现成的检测器,基于场景内容和每个物体的时序状态决定何时可以安全地跳过检测器调用,无需修改或重新训练检测器。Albireo为每个活跃物体维护一个10维卡尔曼滤波器(KF),仅当预测不确定性超过阈值时才调用检测器;在跳过的帧上,边界框由KF状态以近乎零的GPU开销预测得到。基于KF的救援机制可在检测器短暂漏检时保持已确认物体,防止输出碎片化;同时,轻量级的空场景筛选机制避免在无物体帧上调用检测器。我们在BDD100K MOT验证集上,使用三个架构迥异的检测器(YOLO11x、YOLO26x、RF-DETR-Large),在两个NVIDIA Jetson平台(AGX Thor、AGX Orin)上评估Albireo。在所有配置下,Albireo的AP@50与逐帧推理相差不超过±1.2个百分点,同时将总能耗降低12.1%–17.6%。在YOLO26x上,它将AP@50提升0.8个百分点,同时将能耗降低17.6%(Thor)和14.4%(Orin),每帧能量延迟积分别降低24.9%和26.1%。因此,其默认工作点可同时改善精度、能耗和延迟。相比之下,具有50%跳帧率的固定间隔基线FixedSkip-2损失了8.6个百分点的AP@50。源代码、评估流程及逐段视频结果可在 https://github.com/amirtaherin/albireo 获取。
cs.CV / 75 / 2609.29650

VG-TIE: An interpretable tabular-to-image encoding method based on visibility graphs

VG-TIE:一种基于可视图的表格数据到图像的可解释编码方法
Chushig-Muzo, David, López-Ramos, Luis M., de Cara, Ángeles Rodríguez, Milara, Eva, Zhinin-Vera, Luis, Peluffo-Ordóñez, Diego H.
Abstract
Tabular-to-image encoding methods enable the application of models based on both convolutional neural networks and vision transformers to tabular data, transforming feature vectors into images. Existing methods employ linear and nonlinear dimensionality reduction techniques (e.g., Principal Component Analysis (PCA), t-SNE, and UMAP) to determine pixel positions, resulting in images whose spatial layout do not inherently reflect feature relationships. This paper introduces Visibility Graphs for Tabular-to-Image Encoding (VG-TIE), a novel method that encodes the structure of feature values using Natural Visibility Graph (NVG) and Horizontal Visibility Graph (HVG) into a two-dimensional space obtained through PCA. The resulting images are model-agnostic and intrinsically interpretable. Each pixel corresponds to an input feature, its intensity reflects the magnitude and direction of deviation from the population mean, and edges represent formally defined visibility relationships between features. VG-TIE provides two interpretability methods: (i) feature ranking from node degree distributions; and (ii) local and global feature importance from pixel intensity combined with Grad-CAM. Experiments on six public tabular datasets show that VG-TIE is competitive with other tabular-to-image methods while providing interpretability on feature importance and ranking similar to intrinsic interpretable methods. The results highlight the potential of the proposed image-based transformation to provide an effective framework that expands the use of deep learning across tabular data domains.
Chinese Translation
表格数据到图像的编码方法使得基于卷积神经网络和视觉Transformer的模型能够应用于表格数据,即将特征向量转换为图像。现有方法采用线性和非线性降维技术(如主成分分析(PCA)、t-SNE和UMAP)来确定像素位置,导致所得图像的空间布局并不能固有地反映特征之间的关系。本文提出了一种新颖的表格数据到图像编码方法——基于可视图的表格数据到图像编码(VG-TIE),该方法利用自然可视图(Natural Visibility Graph, NVG)和水平可视图(Horizontal Visibility Graph, HVG)对特征值的结构进行编码,并将其映射到通过PCA获得的二维空间中。所得图像与模型无关且具有内在可解释性:每个像素对应一个输入特征,其强度反映该特征偏离总体均值的幅度和方向,而边则表示特征之间经形式化定义的可视性关系。VG-TIE提供了两种可解释性方法:(i)基于节点度分布的特征排序;(ii)结合像素强度与Grad-CAM的局部和全局特征重要性。在六个公开表格数据集上的实验表明,VG-TIE与其他表格数据到图像方法相比具有竞争力,同时能提供与内在可解释方法相似的特征重要性和排序可解释性。研究结果凸显了所提出的基于图像的转换方法在提供一个有效框架、从而将深度学习扩展应用于各类表格数据领域的潜力。
cs.CV / 76 / 2609.29663

Investigating White Blood Cells as a Source of False-Positive Malaria Parasite Detection in African Blood-Smear Images

探究白细胞作为非洲血液涂片图像中疟疾寄生虫误检(假阳性)来源的研究
Adeniji, Samuel A., Obasi, Goodness C., Ntwali, Chris-Victor, Iorumbur, Aondana M., Raymond, Confidence, Uwimana, Lowami, Issah, Ahmed Tahiru
Abstract
White blood cells (WBCs) present on every Giemsa-stained thick blood smear share visual properties with early-stage Plasmodium falciparum ring-form trophozoites: small size, round morphology, and intense purple staining. They are a plausible but untested source of false positives in parasite-only detectors. We trained two YOLOv12s models on the Lacuna Malaria Detection dataset (8,000 images from Uganda and Ghana): Model A with parasite labels only, and Model B with both parasite and WBC labels. Seven independent spatial and statistical analyses tested whether false positive (FP) predictions cluster near WBC locations. All seven refute the hypothesis. In both models, 95% of FPs are pure background detections (IoU below 0.10 against any ground-truth box); zero are WBC class confusions. Ripley's Cross-K analysis shows spatial repulsion between FP centroids and WBC positions at every radius tested. Model B outperforms Model A overall (mAP50 0.859 vs. 0.755), and the advantage is uniform across all WBC-proximity bands, pointing to multi-task representation learning rather than WBC suppression as the cause. False positives arise from Giemsa stain debris and preparation artifacts. Effective mitigation requires staining artifact augmentation and annotation of unannotated early-stage ring forms rather than WBC labeling alone.
Chinese Translation
吉姆萨(Giemsa)染色厚血涂片中普遍存在的白细胞(WBC)与早期恶性疟原虫(Plasmodium falciparum)环状滋养体具有相似的视觉特征:体积小、形态圆、染色呈深紫色。它们是仅寄生虫检测器中产生假阳性的一个合理但尚未验证的来源。我们在Lacuna疟疾检测数据集(来自乌干达和加纳的8,000张图像)上训练了两个YOLOv12s模型:模型A仅使用寄生虫标签,模型B同时使用寄生虫和白细胞标签。我们通过七项独立的空间与统计分析检验假阳性(FP)预测是否聚集在白细胞位置附近,所有七项分析均否定了该假设。在两个模型中,95%的假阳性均为纯背景检测(与任何真实标注框的IoU低于0.10),没有任何一例属于白细胞类别混淆。Ripley交叉K函数分析显示,在所有测试半径下,假阳性中心点与白细胞位置之间均存在空间排斥。模型B整体性能优于模型A(mAP50分别为0.859和0.755),且该优势在所有白细胞邻近区间内均匀分布,表明其收益来自多任务表征学习而非白细胞抑制。假阳性源于吉姆萨染色液残渣和制样伪影。有效的缓解措施需要进行染色伪影数据增强以及对未标注的早期环状体进行标注,而非仅仅依靠白细胞标注。
cs.CV / 77 / 2609.29678

ReCalMatch:Reliability-Calibrated Semantic Guidance for Semi-Supervised Fine-Grained Recognition

ReCalMatch:面向半监督细粒度识别的可靠性校准语义引导
Hong, Yundi, He, Hongyang, Fang, Zheng, Liu, Xuanyu, Sanchez, Victor
Abstract
Semi-supervised fine-grained visual recognition is highly vulnerable to overconfident pseudo-label errors: visually similar categories frequently produce high-confidence yet incorrect predictions, and consistency regularization then reinforces these errors throughout training. Existing semi-supervised learning (SSL) methods estimate pseudo-label reliability almost entirely from the visual classifier itself---maximum probability, adaptive thresholds, or entropy---signals that remain blind to whether a predicted class is \emph{semantically} compatible with the visual representation. We propose \textbf{ReCalMatch}, a reliability-calibrated semantic framework for semi-supervised fine-grained recognition. Rather than treating textual semantics as auxiliary supervision, ReCalMatch uses multi-aspect semantic prototypes as \emph{calibration evidence} for pseudo-label learning. We construct class-conditioned semantic prototypes from class names and domain-specific semantic aspects, and measure a \emph{visual--semantic agreement} score between each unlabeled embedding and its pseudo-label prototype. This agreement is combined with prediction confidence and entropy into a single reliability weight that down-weights pseudo-labels that are visually confident but semantically inconsistent. A semantic consistency term and a semantic margin regularizer further sharpen prototype separability under limited labels. Extensive experiments on CUB-200-2011, Stanford Dogs, NABirds, and iNaturalist18 show that ReCalMatch consistently improves strong SSL baselines, with the largest gains in low-label regimes where pseudo-label noise is most severe.
Chinese Translation
半监督细粒度视觉识别极易受到过度自信的伪标签错误的影响:视觉上相似的类别经常产生高置信度却错误的预测,而一致性正则化随后会在整个训练过程中不断强化这些错误。现有的半监督学习(SSL)方法估计伪标签可靠性时几乎完全依赖视觉分类器本身——最大概率、自适应阈值或熵——这些信号无法判断预测类别与视觉表示在语义上是否兼容。我们提出了ReCalMatch,一个用于半监督细粒度识别的可靠性校准语义框架。ReCalMatch并非将文本语义作为辅助监督,而是将多方面语义原型作为伪标签学习的校准证据。我们基于类别名称和领域特定的语义方面构建类别条件语义原型,并度量每个无标签样本的嵌入与其伪标签原型之间的视觉—语义一致性得分。该一致性得分与预测置信度和熵相结合,形成单一的可靠性权重,从而降低那些视觉上自信但语义上不一致的伪标签的权重。语义一致性项和语义间隔正则化器进一步在标签有限的情况下锐化原型的可分性。在CUB-200-2011、Stanford Dogs、NABirds和iNaturalist18上的大量实验表明,ReCalMatch持续提升强SSL基线的性能,且在伪标签噪声最严重的低标签量场景中收益最大。
cs.CV / 78 / 2609.29717

TopoFuse: Topology-Aware Tri-Planar Fusion for 3D Cryo-Electron Tomography Segmentation

TopoFuse:面向三维冷冻电子断层扫描分割的拓扑感知三平面融合方法
Salla, Rohit Kumar, Gupta, Neelesh, Li, Xingjian, Xu, Min
Abstract
Automated segmentation of cryo-electron tomograms routinely produces masks that are voxel-accurate but topologically broken: membranes fragment, organelles merge into one another, and enclosed cavities collapse. Existing topology-aware losses reduce these violations but cannot eliminate them, because topology is encouraged through gradient pressure rather than structurally enforced. We introduce TopoFuse, which reframes topology as a differentiable projection operator rather than a loss penalty. At each forward pass, the projection operator $\mathrm{Proj}_T$ (a PH-guided sparse edit) identifies the critical voxels responsible for topological violations via bottleneck matching and applies sparse edits to satisfy a specified topology target (diagram feature counts and lifetime budgets) for dimensions $d \in \{0,2\}$. If the projection converges, the output satisfies those constraints on the downsampled grid ($s=2$); when it does not, a repair certificate exposes this explicitly, enabling downstream filtering. A topology prior head predicts the correction target directly from input features, removing any dependence on ground-truth topology at inference. Across three cryo-ET benchmarks, TopoFuse reduces Betti number error by 54% over the strongest soft-loss baseline ($p < 0.001$), improves Dice by 4.6 points, and edits only 3.1% of voxels to achieve this.
Chinese Translation
冷冻电子断层扫描(cryo-electron tomography)的自动分割结果常常是体素级精确但拓扑结构破碎的掩膜:膜结构出现断裂,细胞器相互粘连合并,封闭腔体发生塌陷。现有的拓扑感知损失函数虽能减少这类拓扑违规,但无法将其消除,因为其仅通过梯度压力来鼓励拓扑正确性,而非从结构上加以保证。我们提出 TopoFuse,将拓扑重新表述为一个可微的投影算子,而非损失惩罚项。在每次前向传播中,投影算子 $\mathrm{Proj}_T$(一种基于持续同调(PH)引导的稀疏编辑)通过瓶颈匹配识别导致拓扑违规的关键体素,并施加稀疏编辑以满足指定维数 $d \in \{0,2\}$ 的拓扑目标(图特征数量与寿命预算)。若投影收敛,则输出在下采样网格($s=2$)上满足这些约束;若不收敛,修复证书(repair certificate)会显式地暴露这一情况,从而支持下游过滤。拓扑先验头(topology prior head)直接从输入特征预测校正目标,使推理阶段完全无需依赖真实拓扑标注。在三个冷冻电子断层扫描(cryo-ET)基准数据集上,TopoFuse 相比最强的软损失基线将贝蒂数(Betti number)误差降低了 54%($p < 0.001$),Dice 系数提升了 4.6 个百分点,而仅需编辑 3.1% 的体素即可实现。
cs.CV / 79 / 2609.29721

SALI: Shot-Aware Late Interaction for Cross-Shot Relation Matching in Text-to-Video Retrieval using Film-Grammar Knowledge

SALI:基于电影语法知识的镜头感知型后期交互方法,用于文本到视频检索中的跨镜头关系匹配
Oyama, Toya, Lienhart, Rainer, Satoh, Shin'ichi
Abstract
Text-to-video retrieval usually represents a video clip by a single embedding. This embedding often loses important relations between people. E.g., an interaction "Anna confronts Mark" is regularly filmed as alternating shot and reverse shot of both (Fig. 1a). No single shot or averaged embedding over clip shots captures this relation. Thus, we propose SALI (Shot-Aware Late Interaction). It extracts the subject and object from a single-sentence query, and matches the query, its subject and object text embeddings against each visual shot embedding of a video clip. The matching operator is greedy max or optimal transport. A film-grammar penalty in fine-tuning adds a small, consistent shift. Built on CLIP4Clip-meanP, SALI keeps overall recall on par on Condensed Movies and ActivityNet while raising R@1 on multi-shot relation queries by 3 and 12 points, the most among all compared methods, and improves such queries on MSR-VTT at a cost of 1.4 R@1 overall.
Chinese Translation
文本到视频检索通常用单个嵌入向量来表示视频片段。这种嵌入往往会丢失人物之间的重要关系。例如,'安娜对质马克'这样的互动通常以两人的正反打镜头(交替拍摄)呈现(图1a)。没有任何单一镜头或片段镜头的平均嵌入能够捕捉这种关系。为此,我们提出了SALI(Shot-Aware Late Interaction,镜头感知型后期交互)方法。该方法从单句查询中提取主语和宾语,并将查询文本、主语文本和宾语文本的嵌入与视频片段的每个视觉镜头嵌入进行匹配。匹配算子采用贪心最大值匹配或最优传输。在微调过程中引入的电影语法惩罚项则带来一个微小而一致的偏移。SALI基于CLIP4Clip-meanP构建,在Condensed Movies和ActivityNet数据集上保持了与基线相当的整体召回率,同时在多镜头关系查询上的R@1分别提升了3个和12个百分点,在所有对比方法中提升幅度最大;在MSR-VTT上也改善了此类查询,代价是整体R@1下降1.4个百分点。
cs.CV / 80 / 2609.29726

A Multimodal Dataset for Survival Prediction in Resected Pancreatic Ductal Adenocarcinoma

用于胰腺导管腺癌切除术后生存预测的多模态数据集
Nguyen, Anh-Tien, Tettey, Mawuko, Metsch, Jacqueline Michelle, Zimmer, Teresa, Ullrich, Niklas, Duker, Mario, Rungeling, Sandra, Reuter-Jessen, Kirsten, Rosenthal, Tessa, Conradi, Lena-Christin, Ghadimi, Michael, Konig, Alexander, Hessmann, Elisabeth, Ellenrieder, Volker, Strobel, Philipp, Bohnenberger, Hanibal, Hauschild, Anne-Christin
Abstract
Survival research in pancreatic ductal adenocarcinoma (PDAC) is limited by the scarcity of datasets linking whole-slide histology with clinical, molecular, and long-term outcome data. We present a retrospective single-centre cohort of 302 patients who underwent PDAC resection at University Medical Center Gottingen. The dataset comprises 446 H&E whole-slide images, clinicopathological variables, targeted sequencing data for 154 patients, and overall-survival outcomes. During follow-up, 253 patients died, and the median follow-up was 76 months. To establish initial reference values, we evaluated fourteen survival-prediction configurations using identical five-repetition Monte Carlo cross-validation partitions. Ridge Cox regression using numeric clinicopathological variables achieved a mean concordance of $0.649 \pm 0.042$ and $0.652 \pm 0.046$ after adding KRAS and TP53 mutation status. The image-only attention model achieved $0.603 \pm 0.030$, while multimodal fusion achieved $0.619 \pm 0.025$, the highest concordance among the neural models. These results establish promising initial benchmarks for future research using this pancreas-specific multimodal dataset, paving the way for external validation.
Chinese Translation
胰腺导管腺癌(PDAC)的生存研究受限于缺乏将全切片组织病理学图像与临床、分子及长期结局数据相关联的数据集。我们提出了一个来自哥廷根大学医学中心(University Medical Center Göttingen)的回顾性单中心队列,共纳入302例接受PDAC切除术的患者。该数据集包含446张H&E染色的全切片图像、临床病理变量、154例患者的靶向测序数据以及总生存期结局。在随访期间,253例患者死亡,中位随访时间为76个月。为建立初步的参考值,我们在相同的五次重复蒙特卡洛交叉验证划分下评估了十四种生存预测配置。基于数值型临床病理变量的岭回归Cox模型(Ridge Cox regression)的平均一致性指数(concordance)为0.649±0.042,加入KRAS和TP53突变状态后为0.652±0.046。仅使用图像的注意力模型达到0.603±0.030,而多模态融合模型达到0.619±0.025,在所有神经模型中一致性最高。这些结果为未来利用该胰腺特异性多模态数据集开展研究奠定了有前景的初步基准,并为外部验证铺平了道路。
cs.CV / 81 / 2609.29779

Mind the Gap: Mesh-Guided Repair of Broken Vessels

注意间隙:基于网格引导的断裂血管修复
Drwiega, Gniewosz, Szymanski, Wojciech, Wodzinski, Marek
Abstract
Vessel segmentation is commonly optimized as voxel-wise classification, but small local errors can strongly disrupt vascular connectivity while having little effect on overlap scores. This is particularly problematic for downstream analyses that rely on centerlines, branches, connected components, or graph structure. We propose a mesh-guided post-processing framework for repairing broken vessel segmentations produced by nnU-Net. For each predicted binary mask, a deformable template mesh is fitted to the mask surface in physical space and used as a case-specific geometric scaffold. The fitted mesh is not voxelized as the final segmentation; instead, it guides conservative reconnection of disconnected components by proposing or validating thin bridge candidates under foreground-growth constraints. We evaluated this approach in three vascular anatomies using AortaSeg24 and SEGA for the aorta, TopCoW for the Circle of Willis, and PARSE for the pulmonary arteries. Performance is measured using Dice, connected-component Dice (ccDice), and the Betti-0 number. Across these datasets, repair substantially improved connectivity while preserving overlap: Dice remained nearly unchanged, whereas ccDice increased from 0.596 to 0.992 for aorta, from 0.722 to 0.835 for TopCoW, and from 0.028 to 0.862 for PARSE. The FOMAML meta-initialization further accelerated the fitting per-case, supporting practical mesh-based repair of the vascular topology. These results suggest that explicit mesh representations can provide a useful geometric prior for correcting topological failures in otherwise accurate voxel segmentations.
Chinese Translation
血管分割通常被优化为逐体素分类任务,但微小的局部错误就可能严重破坏血管的连通性,而对重叠分数的影响却很小。这对依赖中心线、分支、连通分量或图结构的下游分析而言尤为棘手。我们提出了一种网格引导的后处理框架,用于修复由 nnU-Net 生成的断裂血管分割结果。对于每个预测的二值掩码,我们在物理空间中将一个可变形的模板网格拟合到掩码表面,并将其作为针对具体病例的几何支架。拟合得到的网格并不被体素化为最终分割结果;相反,它在前景生长约束下通过提出或验证细桥梁候选来引导断开连通分量的保守重连。我们在三种血管解剖结构上评估了该方法:使用 AortaSeg24 和 SEGA 处理主动脉,TopCoW 处理 Willis 环,PARSE 处理肺动脉。性能评估采用 Dice、连通分量 Dice(ccDice)和 Betti-0 数。在这些数据集上,修复在保持重叠度的同时显著提升了连通性:Dice 几乎不变,而 ccDice 在主动脉上从 0.596 提升至 0.992,TopCoW 上从 0.722 提升至 0.835,PARSE 上从 0.028 提升至 0.862。FOMAML 元初始化进一步加速了每个病例的网格拟合,支持了基于网格的血管拓扑修复的实际应用。这些结果表明,显式网格表示可以为修正原本已较准确的体素分割中的拓扑错误提供有用的几何先验。
cs.CV / 82 / 2609.29785

Lightweight Vision Transformer-Based U-Net for Brain Tumor Segmentation from MRI

基于轻量级视觉Transformer的U-Net用于MRI脑肿瘤分割
Banerjee, Sheekar, Chowdhury, Md. Srabon, Akash, Md. Mahbub Hasan, Mamoon, Ishtiak Al
Abstract
Accurate brain tumor segmentation from Magnetic Resonance Imaging is essential for diagnosis, treatment planning, and surgical guidance. Although Convolutional Neural Networks, particularly UNet, have achieved significant success in medical image segmentation, they often struggle to capture the long-range spatial dependencies required to model tumors with irregular shapes and complex boundaries. This paper proposes a lightweight Vision Transformer UNet that combines the hierarchical feature extraction capability of UNet with the global context modeling of Vision Transformers. The proposed architecture incorporates a compact ViT bottleneck within a U-Net encoder-decoder framework, enabling effective learning of both local and global features while maintaining computational efficiency with only 2.6 million trainable parameters. The model was evaluated on the TCGA LGG MRI Segmentation dataset, achieving a mean Intersection over Union of 0.8100 and a Dice score of 0.8446, outperforming the baseline UNet by 3.75% and 3.15%, respectively. Extensive quantitative and qualitative analyses, including confusion matrix evaluation, precision recall curves, per-image performance distribution, and tumor size dependency analysis, demonstrate the effectiveness and robustness of the proposed method for brain tumor segmentation.
Chinese Translation
从磁共振成像(MRI)中准确分割脑肿瘤对于诊断、治疗规划和手术引导至关重要。尽管卷积神经网络(尤其是UNet)在医学图像分割中取得了显著成功,但它们往往难以捕获建模不规则形状和复杂边界肿瘤所需的长程空间依赖关系。本文提出一种轻量级视觉Transformer UNet,将UNet的层次化特征提取能力与视觉Transformer的全局上下文建模相结合。该架构在U-Net编码器-解码器框架中嵌入了一个紧凑的ViT瓶颈层,能够有效学习局部和全局特征,同时仅拥有260万可训练参数,保持了较高的计算效率。该模型在TCGA LGG MRI分割数据集上进行了评估,平均交并比(mIoU)达到0.8100,Dice系数达到0.8446,分别比基线UNet提升了3.75%和3.15%。大量定量和定性分析,包括混淆矩阵评估、精确率-召回率曲线、逐图像性能分布以及肿瘤大小依赖性分析,证明了所提方法在脑肿瘤分割中的有效性和鲁棒性。
cs.CV / 83 / 2609.29788

OREO: Fidelity Alignment in 3D Generation via On-the-fly Rendering-Editing Optimization

OREO:通过即时渲染-编辑优化实现3D生成中的保真度对齐
Ma, Zhiyuan, Hu, Wenbo, Zhao, Wang, Wang, Pengfei, Shan, Ying, Zhang, Lei
Abstract
Despite recent advancements in 3D generation, models often struggle to produce assets with high visual fidelity. To bridge this gap, we propose OREO, an alignment framework that enhances the realism of 3D generators by leveraging rich 2D diffusion priors. Instead of relying on static datasets, OREO establishes a dynamic optimization loop that produces on-the-fly edited renderings as 2D pseudo-targets. At its core, we introduce Reinforced Editing, which utilizes a 2D model to refine rendered views of the 3D output, enhancing their overall visual fidelity while preserving the underlying geometry, viewpoint, and content. These refined views serve as high-quality supervision targets, enabling the 3D generator to learn from its own generated samples and progressively improve its visual quality. Experiments demonstrate that OREO effectively improves upon pre-trained baselines, producing 3D assets with enhanced visual realism.
Chinese Translation
尽管3D生成技术近期取得了进展,但模型往往难以生成具有高视觉保真度的资产。为弥合这一差距,我们提出了OREO,一个通过利用丰富的2D扩散先验来提升3D生成器真实感的对齐框架。OREO不依赖静态数据集,而是建立一个动态优化循环,即时生成经过编辑的渲染结果作为2D伪目标。其核心是强化编辑(Reinforced Editing)技术,该技术利用2D模型对3D输出的渲染视图进行精修,在保持底层几何结构、视角和内容的同时提升其整体视觉保真度。这些精修后的视图作为高质量的监督目标,使3D生成器能够从自身生成的样本中学习,并逐步提升视觉质量。实验表明,OREO有效改进了预训练基线模型,生成了具有更强视觉真实感的3D资产。
cs.CV / 84 / 2609.29813

S2Planner: Multi-Scale Semantic Planner for End-to-End Autonomous Driving

S2Planner:面向端到端自动驾驶的多尺度语义规划器
Lu, Zhaowei, Zhou, Liguo, Guo, Yujie, Yu, Lei, Knoll, Alois
Abstract
We present S2Planner, a trajectory planner that combines three front-facing cameras with ego-motion history and the current driving command. A fine-tuned DINOv3 backbone and a Spatial Tuning Adapter produce multi-scale image features; a coarse-to-fine decoder then uses trajectory self-attention and camera-projected cross-attention to refine candidate waypoints. The contribution is the integration of ego-conditioned trajectory initialization with iterative, geometry-guided sampling of multi-scale image features, rather than a new visual backbone or attention operator. On the NAVSIM v1 non-reactive evaluation, the previously reported navtest run obtained 88.03 PDMS. Because that run was selected using navtest performance, this number is exploratory and cannot be interpreted as an unbiased test estimate. Validation-selected evaluation on unexposed data, repeated runs, and computational measurements are needed to establish generalization and efficiency.
Chinese Translation
我们提出了S2Planner,一种结合三个前视相机、自车运动历史和当前驾驶指令的轨迹规划器。经过微调的DINOv3主干网络与空间调谐适配器(Spatial Tuning Adapter)生成多尺度图像特征;随后,一个由粗到细的解码器利用轨迹自注意力和相机投影交叉注意力对候选路径点进行细化。本文的贡献在于将基于自车条件的轨迹初始化与基于几何引导的多尺度图像特征迭代采样相结合,而非提出新的视觉主干网络或注意力算子。在NAVSIM v1非反应式评估中,此前报告的navtest运行获得了88.03的PDMS分数。由于该运行是基于navtest性能选择的,因此该数字仅具探索性,不能被解释为无偏的测试估计。为确立泛化能力与计算效率,仍需在未接触数据上进行基于验证集选择的评估、多次重复运行以及计算量测量。
cs.CV / 85 / 2609.29816

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

AV-GRPO:面向音视频联合生成的模态锚定解耦扩散强化学习
Xu, Zhiyu, Yan, Weilong, Shi, Yufei, Li, Shiyang, Liu, Yihao, Lam, Kin-Man, Cao, Yuewen
Abstract
Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training offers a promising remedy, directly adapting it to joint audio-video generation is challenging. Heterogeneous multimodal rewards entangle learning signals and complicate credit assignment. Joint optimization of two modality towers is computationally expensive given their divergent dynamics. Moreover, synchronization evaluation difficulty depends on paired samples, preventing fair reward comparisons. We propose AV-GRPO, a modality-anchored online diffusion RL framework, and 5DAV, a decoupled, difficulty-controllable training dataset. AV-GRPO includes three key modules: (1) modality-anchored rollouts to disentangle learning signals and stabilize difficulty; (2) trajectory-locked frozen-tower optimization to reduce cost and reassign credit; (3) adaptive objectives and perturbation strengths tailored to modality-specific dynamics. This converts coupled multimodal preference learning into unimodal subproblems for precise reward attribution and better synchronization. Our 5DAV dataset decouples samples across five dimensions for systematic training. Experiments on JavisBench and VABench demonstrate AV-GRPO outperforms LTX-2.3 in generation quality, semantic alignment and cross-modal synchronization under LoRA and full fine-tuning. Ablations confirm our designs. Code and data: https://github.com/zhiyuxu03/AV-GRPO
Chinese Translation
近年来,音视频联合生成取得了重大进展。然而,现有模型仍然存在单模态保真度有限、文本与模态对齐不足以及跨模态同步性较弱等问题。尽管强化学习后训练提供了一种有前景的解决方案,但将其直接应用于音视频联合生成仍面临挑战:异构多模态奖励会使学习信号相互纠缠,增加信用分配的难度;鉴于两个模态塔的动态特性差异较大,对其进行联合优化在计算上代价高昂;此外,同步性评估的难度依赖于配对样本,导致奖励比较难以公平进行。我们提出了 AV-GRPO,一个模态锚定的在线扩散强化学习框架,以及 5DAV——一个解耦的、难度可控的训练数据集。AV-GRPO 包含三个关键模块:(1) 模态锚定 rollout,用于解耦学习信号并稳定难度;(2) 轨迹锁定的冻结塔优化,用于降低计算成本并重新分配信用;(3) 针对模态特定动态特性定制的自适应目标函数与扰动强度。这将耦合的多模态偏好学习转化为单模态子问题,实现精确的奖励归因和更好的同步性。我们的 5DAV 数据集在五个维度上对样本进行解耦,以支持系统化的训练。在 JavisBench 和 VABench 上的实验表明,在 LoRA 和全量微调设置下,AV-GRPO 在生成质量、语义对齐和跨模态同步性方面均优于 LTX-2.3。消融实验验证了各设计的有效性。代码与数据:https://github.com/zhiyuxu03/AV-GRPO
cs.CV / 86 / 2609.29825

Anatomy-Aligned Surface Field Learning for Myocardial Reconstruction from Sparse Short-Axis Cine MRI

面向稀疏短轴电影MRI心肌重建的解剖对齐表面场学习
Yuan, Xiaohan, Yang, Xuan, Li, Qingya, Wang, Yangang, Li, Lei
Abstract
Patient-specific 4D myocardial reconstruction from cine MRI supports quantitative functional assessment, regional motion analysis, and simulation-based modeling. However, routinely acquired short-axis (SAX) cine MRI is sparsely sampled along the through-plane direction, making dense and anatomically consistent surface reconstruction challenging. In this study, we propose an anatomy-aligned surface learning framework that parameterizes the epicardial and endocardial surfaces on a shared circumferential-longitudinal UV domain. This formulation converts irregular 3D reconstruction into structured coordinate-field completion with explicit correspondence across subjects and cardiac phases. Sparse SAX contours are encoded as UV observation fields, coverage-aware sampling improves robustness to incomplete slice coverage, and topology- and distortion-aware learning preserves circumferential continuity and local surface quality. Experiments on three public cine MRI datasets showed that the proposed method consistently outperformed representative mesh-based and implicit reconstruction approaches, achieving overall Chamfer distances of $2.887$~mm on ACDC, $2.641$~mm on M\&Ms, and $2.810$~mm on M\&Ms-2. The reconstructed sequences also preserved ventricular function, with end-diastolic volume and ejection fraction errors of $3.3$~mL and $1.1 \%$, respectively. These results demonstrate that anatomy-aligned UV learning provides an accurate, efficient, and correspondence-aware representation for sparse cine MRI reconstruction and myocardial modeling. The source code will be available at https://github.com/yuan-xiaohan/SAX2MyoSurf.
Chinese Translation
基于电影MRI(cine MRI)的患者特异性四维心肌重建可支持定量功能评估、区域运动分析以及基于仿真的建模。然而,常规采集的短轴(SAX)电影MRI在层间方向上采样稀疏,使得稠密且解剖学一致的表面重建面临挑战。在本研究中,我们提出了一种解剖对齐的表面学习框架,将心外膜和心内膜表面参数化到一个共享的周向-纵向UV域上。该建模方式将不规则的3D重建转化为结构化的坐标场补全问题,并在不同受试者和心脏相位之间实现显式对应。稀疏的SAX轮廓被编码为UV观测场,覆盖率感知采样提升了对不完整层面覆盖的鲁棒性,而拓扑与畸变感知学习则保持了周向连续性和局部表面质量。在三个公开电影MRI数据集上的实验表明,所提出方法持续优于代表性的基于网格和隐式重建方法,在ACDC、M&Ms和M&Ms-2上的总体Chamfer距离分别为2.887 mm、2.641 mm和2.810 mm。重建序列还保留了心室功能,舒张末期容积和射血分数误差分别为3.3 mL和1.1%。这些结果表明,解剖对齐的UV学习为稀疏电影MRI重建和心肌建模提供了一种精确、高效且具备对应关系的表示方法。源代码将发布于 https://github.com/yuan-xiaohan/SAX2MyoSurf。
cs.CV / 87 / 2609.29835

Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding

检索定位一体化:连接大语言模型与LiDAR几何实现空间定位(Spatial Grounding)
Park, Byounggun, Moon, Giyong, Kim, Jusung, Hwang, Soonmin
Abstract
LiDAR provides precise geometric information for spatial perception tasks such as object detection in autonomous driving and outdoor robotics. However, recognizing and localizing individual objects is not sufficient to answer questions that require composing spatial relations and grounding the intended target. Motivated by recent advances in large language models (LLMs) for autonomous driving, we leverage their language priors to interpret complex spatial questions and ground the referred target in LiDAR geometry. To support this spatial grounding capability, we introduce SpatialLiDAR-QA, which combines single- and multi-step relational grounding with complementary spatial understanding tasks. We further propose SpatialLiDAR-LM, which aligns LiDAR point features with an LLM and grounds target coordinates through language-conditioned, position-aware proposal retrieval and local point refinement. This design derives target coordinates directly from local LiDAR geometry rather than through textual language decoding. Experiments demonstrate substantial improvements over representative LiDAR--language models and multi-camera VLMs on precise coordinate prediction tasks. Our dataset and model training code will be publicly released.
Chinese Translation
LiDAR为自动驾驶和户外机器人等空间感知任务(如目标检测)提供了精确的几何信息。然而,仅仅识别和定位单个物体并不足以回答那些需要对空间关系进行组合推理并定位目标对象的问题。受大语言模型(LLMs)在自动驾驶领域最新进展的启发,我们利用其语言先验来解释复杂的空间问题,并将所指目标定位到LiDAR几何空间中。为支持这一空间定位能力,我们提出了SpatialLiDAR-QA数据集,它将单步和多步关系定位与互补的空间理解任务相结合。我们进一步提出SpatialLiDAR-LM模型,该模型将LiDAR点云特征与大语言模型对齐,并通过语言条件化的、位置感知的候选框检索及局部点云精细化来实现目标坐标的定位。这种设计使目标坐标直接源自局部LiDAR几何信息,而非通过文本语言的解码获得。实验表明,在精确坐标预测任务上,我们的方法相较于代表性的LiDAR-语言模型和多相机视觉语言模型(VLMs)取得了显著提升。我们的数据集和模型训练代码将公开发布。
cs.CV / 88 / 2609.29836

SplatLabel: Pseudo-Labelling through 4D Gaussian Splatting

SplatLabel:基于4D高斯泼溅的伪标签生成方法
Nanvani, Nitya, Palffy, Andras, Caesar, Holger
Abstract
While 2D Vision Foundation Models offer a pathway to automate 3D semantic pseudo-labelling, translating these priors into robust 3D representations typically requires complex heuristics or multi-model ensembles. We introduce SplatLabel, an automated pipeline that leverages a 4D Gaussian representation to extract LiDAR segmentation with predictive confidence, as well as semantic occupancy grids at arbitrary voxel resolutions. At its core, SplatLabel handles dynamic environments through an explicit temporal manifold that models the trajectories and lifespans of individual 3D primitives. This allows the system to accurately track moving actors and strictly define when objects appear and disappear, completely eliminating the need for pre-annotated 3D bounding boxes. To robustly support this dynamic tracking, the representation is grounded by structural and semantic priors: we guide scene geometry in unobserved regions by integrating 360-degree LiDAR via virtual depth maps, and rather than relying on domain-specific prompt engineering, we directly distill continuous soft probabilities from 2D models to inherently resolve semantic ambiguities over time and space. Finally, to accurately reflect the real-world trade-off between precision and recall, we reframe pseudo-label evaluation as a selective classification task using a generalized risk-recall metric. Experiments on SemanticKITTI demonstrate that SplatLabel consistently outperforms state-of-the-art baselines across multiple recall levels, establishing a highly robust framework for both 3D LiDAR segmentation and occupancy prediction.
Chinese Translation
虽然2D视觉基础模型为自动化3D语义伪标签生成提供了一条途径,但将这些先验知识转化为鲁棒的3D表示通常需要复杂的启发式方法或多模型集成。我们提出了SplatLabel,这是一个自动化流程,它利用4D高斯表示来提取具有预测置信度的LiDAR分割结果,以及任意体素分辨率的语义占据栅格。SplatLabel的核心在于通过一个显式的时序流形来处理动态环境,该流形对单个3D基元的轨迹和生命周期进行建模。这使得系统能够准确跟踪运动目标,并严格定义物体出现和消失的时间,从而完全消除对预先标注的3D边界框的需求。为稳健地支持这种动态跟踪,该表示以结构和语义先验为基础:我们通过集成360度LiDAR生成的虚拟深度图来引导未观测区域的场景几何,并且不依赖于特定领域的提示工程,而是直接从2D模型中蒸馏连续的软概率,从而在时间和空间上固有的解决语义歧义。最后,为了准确反映精度与召回率之间的现实权衡,我们将伪标签评估重新构建为一个使用广义风险-召回率指标的选择性分类任务。在SemanticKITTI上的实验表明,SplatLabel在多个召回率水平上持续优于最先进的基线方法,为3D LiDAR分割和占据预测任务建立了一个高度鲁棒的框架。
cs.CV / 89 / 2609.29863

Modelling dynamic systems transfer functions from events in computational neuromorphic imaging

基于事件计算神经形态成像中事件信号的动态系统传递函数建模
Kruger, Nimrod, Cohen, Gregory
Abstract
Event Vision Sensing (EVS) report threshold crossings of log-irradiance, so a static optical system imaging a static scene produces no output at all. The classical procedure for measuring a Point Spread Function (PSF), illuminating the system with a constant point source, therefore has no event-based equivalent: the probe must carry a temporal profile, and that profile becomes part of the measurement. A growing body of Computational Neuromorphic Imaging (CNI) work already exploits this, pairing engineered or modulated optics with event sensing, but each system adopts a particular excitation together with a particular reading of the event stream without the correspondence between the two being stated. We examine that correspondence directly within a analytical framework of an Linear Shift-Invariant (LSI) optical system with a specified Modulation Transfer Function (MTF), a first-order filter EVS pixel model, and three different temporal probes: a step function, a linear ramp and an exponential ramp. By analysing the inverse of the entire chain for different event-statistic, and comparing the results to the specified MTF, we identify the context where each probe is most relevant. We consider how photon-noise and cross-array threshold mismatch effects the analytical accuracy of the probe-inverse. Results show that the widely used step probe is highly susceptible to mismatch while resilient to photon shot-noise, while a linear rise probe and exponential rise probe retain their ability to infer signal levels even with high mismatch. We discuss the potential of dynamic-PSFs as components of a full forward operator from scene to events. In this, we use this analytical description to define dynamic-PSFs around EVS, and discuss the gaps toward a unified pixel model and a scene-composition framework required for CNI.
Chinese Translation
事件视觉传感(EVS)报告的是对数辐照度的阈值跨越,因此当静态光学系统对静态场景成像时不会产生任何输出。传统的点扩散函数(PSF)测量方法是利用恒定的点光源照亮系统,因而不存在对应的事件驱动方法:探测信号必须具有时间轮廓,而该轮廓本身也成为测量的一部分。越来越多的计算神经形态成像(CNI)研究已经在利用这一特性,将经过工程化设计或调制的光学元件与事件传感相结合,但每个系统都采用特定的激励信号以及特定的事件流读取方式,而没有阐明二者之间的对应关系。我们在一个解析框架内直接考察这种对应关系:该框架包含一个具有指定调制传递函数(MTF)的线性移不变(LSI)光学系统、一个一阶滤波器EVS像素模型,以及三种不同的时间探测信号:阶跃函数、线性斜坡和指数斜坡。通过分析针对不同事件统计特性的整个链条的逆过程,并将其结果与指定的MTF进行比较,我们确定了每种探测信号最适用的场景。我们考虑了光子噪声和阵列内阈值失配对探测信号逆解析精度的影响。结果表明,广泛使用的阶跃探测信号对失配高度敏感,但对光子散粒噪声具有较强鲁棒性;而线性上升探测信号和指数上升探测信号即使在高失配情况下仍能保持推断信号水平的能力。我们讨论了动态PSF作为从场景到事件的完整前向算子组成部分的潜力。在此过程中,我们利用这一解析描述在EVS基础上定义动态PSF,并讨论了实现CNI所需的统一像素模型和场景组合框架尚存在的差距。
cs.CV / 90 / 2609.29864

Efficient Continuous DEM Reconstruction under Limited Target-Resolution Supervision

有限目标分辨率监督下的高效连续数字高程模型(DEM)重建
Shi, Zekai, Zhang, Meng, Zhang, Haokun, Zhang, Bo
Abstract
High-resolution digital elevation models (DEMs) support Earth observation applications, but paired training references are often available only at coarser output resolutions. Reconstructing finer terrain grids therefore requires both effective transfer beyond the supervised scale and control of dense-query computation. To address this problem, SCOPE learns a continuous terrain representation from coarser-resolution pairs. It predicts a latent coefficient field on the low-resolution grid and reuses local Fourier residual functions through basis evaluation and geometry-guided ensemble fusion. This separates high-dimensional coefficient prediction from output-grid construction. Experiments on geographically distributed land--ocean samples assess supervised reconstruction, unseen-scale inference, cross-domain generalization, and theoretical computation. SCOPE leads the compared methods across six metrics in the main supervised-scale evaluation. At an unseen factor three times the training factor, land reconstruction reduces RMSE and MAE by approximately 12\% relative to bicubic interpolation, with errors close to target-scale fine-tuning. Ninefold output density increases counted multiply--accumulate operations by only about 2\%. Frozen-model validation on held-out external marine regions reduces RMSE relative to the DEM-specific implicit baseline EBCF-CDEM by approximately 19\% under self-downsampling and 2\% with cross-product inputs, while also yielding lower RMSE than LIIF-MS in both settings. These results demonstrate the value of reusable coefficient fields for accurate reconstruction beyond the supervised resolution with low incremental arithmetic cost.
Chinese Translation
高分辨率数字高程模型(DEM)支持多种对地观测应用,但成对的训练参考数据通常仅能以低于目标输出的分辨率获得。因此,重建更精细的地形网格既需要在监督尺度之外进行有效迁移,又需要控制密集查询的计算开销。为解决该问题,SCOPE 从较粗分辨率的成对数据中学习连续地形表示。该方法在低分辨率网格上预测潜在系数场,并通过基函数求值和几何引导的集成融合来复用局部傅里叶残差函数,从而将高维系数预测与输出网格构建相分离。实验基于地理分布的陆海样本,评估了监督重建、未见尺度推断、跨域泛化以及理论计算量。在主要监督尺度评估中,SCOPE 在六项指标上均优于对比方法。在训练因子三倍的未见尺度下,陆地重建的 RMSE 和 MAE 相较双三次插值约降低 12%,误差接近目标尺度微调的结果。输出密度提升九倍时,统计的乘加运算量仅增加约 2%。在留出的外部海洋区域上进行冻结模型验证时,在自下采样条件下,SCOPE 相较 DEM 专用隐式基线 EBCF-CDEM 将 RMSE 约降低 19%,在跨乘积输入条件下约降低 2%,并且在两种设置下的 RMSE 均低于 LIIF-MS。这些结果证明了可复用系数场在监督分辨率之外的精确重建方面的价值,且增量计算成本较低。
cs.CV / 91 / 2609.29930

EndoFSA: Endoscopic Few-Shot Image Generation via Rank-Constrained Parameter Adaptation

EndoFSA:基于秩约束参数适配的内镜少样本图像生成
Gatoula, Panagiota, Karypidis, Grigoris, Iakovidis, Dimitris K.
Abstract
WCE produces large-scale gastrointestinal image data yet pathological findings remain significantly underrepresented limiting the generalization performance of deep-learning based abnormality detection systems. SDG methods offer a practical solution to mitigate this imbalance. However their training directly on scarce abnormal samples often results in instability overfitting and structural distortions. Addressing these challenges requires controlled adaptation mechanisms that preserve anatomical priors while enabling realistic pathological variation. This paper presents EndoFSA a GAN-based model for Endoscopic Few-Shot image generation by Adaptation in WCE imaging. EndoFSA leverages a generator pretrained on abundant normal data and adapts it to abnormal domains using limited number of training samples through a rank-constrained parameter adaptation where only a small number of modulation parameters is updated while the pretrained weights remain frozen. By restricting parameter updates to a low dimensional subspace and incorporating perceptual boundary regularization and cluster-wise diversity control EndoFSA enables efficient model adaptation under limited data conditions and mitigates mode collapse while preserving the anatomical priors learned from normal data. Importantly EndoFSA operates without requiring pixel-level annotations, masks or bounding box supervision. Evaluation on publicly available WCE benchmark datasets spanning various abnormal categories demonstrates that EndoFSA generates abnormal images reproducing real lesions morphology. Moreover in a downstream classification task training an image classifier solely on synthetic abnormal images generated by EndoFSA yields performance comparable to that obtained with real images.
Chinese Translation
无线胶囊内镜(WCE)产生了大规模的胃肠道图像数据,但病理样本仍严重不足,限制了基于深度学习的异常检测系统的泛化性能。基于语义数据生成(SDG)的方法为缓解这一不平衡提供了切实可行的解决方案。然而,直接在稀缺的异常样本上进行训练往往导致训练不稳定、过拟合以及结构畸变。应对这些挑战需要受控的适配机制,在保留解剖先验的同时实现逼真的病理变化。本文提出EndoFSA,一种面向WCE成像的内镜少样本图像生成的GAN模型。EndoFSA利用在大量正常数据上预训练的生成器,通过秩约束参数适配,仅使用少量训练样本将其适配到异常域:只更新少量调制参数,而预训练权重保持冻结。通过将参数更新限制在低维子空间,并结合感知边界正则化和簇级多样性控制,EndoFSA在有限数据条件下实现了高效的模型适配,缓解了模式崩溃,同时保留了从正常数据中学到的解剖先验。重要的是,EndoFSA无需像素级标注、掩膜或边界框监督。在涵盖多种异常类别的公开WCE基准数据集上的评估表明,EndoFSA生成的异常图像能够再现真实病灶的形态。此外,在下游分类任务中,仅使用EndoFSA生成的合成异常图像训练图像分类器,其性能可与使用真实图像训练的结果相媲美。
cs.CV / 92 / 2609.29934

Beyond Spatial Benchmarks: From Spatial Reasoning to Navigation

超越空间基准测试:从空间推理到导航
Huang, Xun, Zhao, Shijia, Qu, Rongsheng, Li, Jiayuan, Lu, Xin, Li, Weixin, Wen, Chenglu, Wang, Cheng
Abstract
Does progress on spatial reasoning benchmarks translate into better navigation? Existing benchmarks test isolated inferences from images or videos, with little connection to downstream navigation. Our analysis reveals a gap between benchmark-oriented spatial specialization and navigation performance, and shows how aligning spatial supervision with navigation goals, phases, and decision learning improves navigation. Guided by these findings, we build \textsc{Spatial-Nav-100K} and fine-tune in two stages, \textit{i.e.} first learning a shared spatial-navigation foundation, and then specializing each phase with the abilities it relies on. We further introduce Spatial-NPD, where a teacher conditioned on spatial priors produces grounded action preferences and distills them into a student policy, so no explicit spatial reasoning is needed at inference. With 45 A100 GPU-hours of policy training, our 8B model reaches SR/SPL of 77.4/35.4 on HM3D-v0.2, 60.2/30.5 on HM3D-v0.1, and 47.9/20.6 on train-unseen MP3D. It outperforms several systems that rely on closed-source models or thousands of GPU-hours of training, at 148 ms per action step. All code and datasets will be publicly available at https://github.com/ylwhxht/Spatial-Nav.
Chinese Translation
空间推理基准测试上的进步能否转化为更好的导航能力?现有的基准测试仅测试从图像或视频中进行孤立的推理,与下游导航任务联系甚少。我们的分析揭示了面向基准的空间能力专化与导航性能之间的差距,并展示了将空间监督与导航目标、阶段及决策学习相如何提升导航性能。基于这些发现,我们构建了 Spatial-Nav-100K 数据集,并进行两阶段微调:首先学习一个共享的空间-导航基础能力,然后针对每个导航阶段进行其所需能力的专化训练。我们进一步提出了 Spatial-NPD 方法:以空间先验为条件的教师模型生成有依据的动作偏好,并将其蒸馏到学生策略中,从而在推理时无需显式的空间推理。仅使用 45 个 A100 GPU 小时的策略训练,我们的 8B 模型在 HM3D-v0.2 上达到 SR/SPL 77.4/35.4,在 HM3D-v0.1 上达到 60.2/30.5,在未见过的训练集 MP3D 上达到 47.9/20.6。该模型以每步 148 毫秒的速度,超越了多个依赖闭源模型或数千 GPU 小时训练的系统。所有代码和数据集将在 https://github.com/ylwhxht/Spatial-Nav 公开。
cs.CV / 93 / 2609.29940

Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration

关注推理中真正重要的因素:通过选择性概率质量集中实现跨模态注意力对齐
Deng, Jiaqi, Wu, Zonghan, Heng, Zhan, Huang, Xiaoshui, Huo, Huan, Xu, Guandong
Abstract
Multimodal large language models (MLLMs) achieve strong performance on visual reasoning tasks, yet remain prone to hallucinations and over-reliance on language priors, often generating answers without adequately using task-relevant visual evidence. Existing approaches primarily improve reasoning through reasoning-oriented supervision or inference-time strategies. In this work, we study a complementary question: can multimodal reasoning be improved by strengthening implicit visual grounding without directly supervising the reasoning process? Motivated by the functional specialization of attention heads, we investigate whether reasoning can be improved by guiding only the heads most responsive to visual evidence grounding. We propose Selective Probability Mass Concentration (sPMC), a training framework that identifies grounding-responsive heads and selectively regularizes their text-to-image attention. sPMC treats normalized attention over visual tokens as a spatial probability distribution and encourages the probability mass to be assigned to semantically relevant regions using segmentation-derived spatial priors. Adaptive Head Selection restricts this guidance to visually responsive heads while leaving the remaining heads unconstrained to preserve their complementary functions. Across 6 multimodal benchmark suites, sPMC achieves an average zero-shot improvement of 3% and gains of up to 11.3% across multiple MLLMs while regularizing only 3%-15% of their attention heads. These results demonstrate that targeted guidance of sparse and implicit visual evidence pathways can directly improve multimodal reasoning.
Chinese Translation
多模态大语言模型(MLLMs)在视觉推理任务上取得了优异性能,但仍然容易产生幻觉并过度依赖语言先验,往往在没有充分利用与任务相关的视觉证据的情况下就生成答案。现有方法主要通过面向推理的监督或推理时策略来改进推理。在本工作中,我们研究一个互补性问题:能否在不直接监督推理过程的情况下,通过强化隐式视觉锚定来提升多模态推理能力?受注意力头的功能特化启发,我们探究是否可以通过仅引导对视觉证据锚定最敏感的注意力头来改进推理。我们提出选择性概率质量集中(Selective Probability Mass Concentration, sPMC),这是一个训练框架,用于识别锚定敏感的注意力头,并选择性地对其文本到图像的注意力进行正则化。sPMC 将视觉 token 上的归一化注意力视为空间概率分布,并利用分割得到的空间先验,促使概率质量集中于语义相关的区域。自适应头选择(Adaptive Head Selection)将该引导限制在视觉敏感的注意力头上,而保持其余注意力头不受约束,以保留其互补功能。在 6 个多模态基准测试套件上,sPMC 在多个 MLLM 上实现了平均 3% 的零样本提升,最高提升达 11.3%,而仅对其中 3%-15% 的注意力头进行正则化。这些结果表明,对稀疏且隐式的视觉证据通路进行有针对性的引导可以直接提升多模态推理能力。
cs.CV / 94 / 2609.29959

Not All Confusion Is Equal: A Source-Aware Uncertainty Diagnosis for Fine-Grained Aircraft Detection

并非所有混淆都是等同的:面向细粒度飞机检测的来源感知不确定性诊断
Huang, Hai, Mayer, Helmut
Abstract
Fine-grained object detectors are commonly evaluated with confusion matrices, which show where the model is confused but not why, nor whether the confusion can be reduced. We argue that confusion can be attributed to distinct, separable sources, each quantitatively measurable, turning a passive measurement into actionable guidance. We present $A^2E^2$, a diagnostic tool that decomposes the sources of confusion along two axes, $\{$aleatoric, epistemic$\} \times \{$within-class, between-class$\}$, giving a $2\times2$ taxonomy that enumerates the source types. Each quadrant is measured by its own quantity, computed in one of three places (input geometry, output-space disagreement, and the bias-parameter posterior), so the two epistemic sources are separated by construction rather than by an empirical correlation. On fine-grained aircraft detection, the four quadrants become four named sources with their own remedy verdict: affinity (geometric similarity, irreducible from size alone), heterogeneity (geometrically heterogeneous sub-variants, pointing to re-labeling rather than more data), contested (an insufficiently trained but learnable boundary, improvable), and collapsed (a class starved of data, reducible). After attributing the confusion to a specific reducible source, we apply a targeted intervention and verify experimentally that it reduces the diagnosed source specifically while leaving the irreducible sources unchanged. $A^2E^2$ thus turns confusion measurement into a concrete, validatable and actionable "diagnosis" in which the same off-diagonal mass can carry opposite causes and opposite remedies. We also state this framework's limits, including which sources are only partially identifiable on this specific dataset and why.
Chinese Translation
细粒度目标检测器通常使用混淆矩阵进行评估,混淆矩阵能够显示模型在哪里发生混淆,却无法说明为什么混淆,以及混淆是否可以被消除。我们认为,混淆可以归因于不同且可分离的来源,每个来源都可以定量测量,从而将被动测量转化为可操作的指导。我们提出了 A²E²,一种沿两个轴({偶然不确定性,认知不确定性} × {类内,类间})分解混淆来源的诊断工具,形成一个枚举来源类型的 2×2 分类法。每个象限由其专属的度量来测量,该度量在三个位置之一进行计算(输入几何、输出空间分歧以及偏置参数后验),因此两种认知不确定性来源在构造上就被分离开,而非依赖经验相关性。在细粒度飞机检测任务中,四个象限成为四个具有各自补救判定的命名来源:亲和性(几何相似性,仅凭尺寸不可消除)、异质性(几何上异构的子变体,指向重新标注而非增加数据)、争议性(训练不足但可学习的边界,可改进)以及坍缩性(数据匮乏的类别,可减少)。在将混淆归因于特定的可减少来源后,我们实施针对性干预,并通过实验验证该干预能特异性地降低所诊断的来源,同时保持不可消除的来源不变。因此,A²E² 将混淆测量转化为具体、可验证且可操作的“诊断”,其中相同的非对角质量可能承载相反的原因和相反的补救措施。我们还阐述了该框架的局限性,包括在此特定数据集上哪些来源只能被部分识别以及原因。
cs.CV / 95 / 2609.29963

ADATEX4D: adaptive texture capacity allocation for 4D gaussian splatting

ADATEX4D:面向4D高斯泼溅的自适应纹理容量分配
Jiang, De, Wang, Peiqiang, Yuan, Kehong, Ma, Shaohua
Abstract
Textured Gaussians improve local appearance capacity, but assigning the same texture resolution to every primitive wastes storage on low-detail or weakly visible regions. We introduce AdaTex4D, an adaptive texture-capacity module for deformation-based 4D Gaussian Splatting. Each Gaussian carries packed RGBA triplanes whose two axes grow independently according to visibility normalized screen-space gradients and deformed local scales. Experiments on N3DV and PanopticSports show that AdaTex4D reduces texture storage by more than half while preserving reconstruction quality. Under fixed memory budgets, adaptive allocation also improves quality over uniform texture assignment and reduces overall model and peak memory. These results show that dynamic, anisotropic texture allocation provides a more efficient way to distribute local appearance capacity in 4D Gaussian representations.
Chinese Translation
带纹理的高斯(Textured Gaussians)提升了局部外观表达能力,但为每个图元分配相同的纹理分辨率会在低细节或弱可见区域上浪费存储空间。我们提出了AdaTex4D,一个用于基于形变的4D高斯泼溅(4D Gaussian Splatting)的自适应纹理容量模块。每个高斯携带打包的RGBA三平面(triplanes),其两个轴根据可见性归一化的屏幕空间梯度和形变后的局部尺度独立增长。在N3DV和PanopticSports数据集上的实验表明,AdaTex4D在保持重建质量的同时,将纹理存储减少了一半以上。在固定内存预算下,自适应分配相比均匀纹理分配也能提升质量,并降低整体模型大小和峰值内存。这些结果表明,动态、各向异性的纹理分配为在4D高斯表示中更高效地分配局部外观容量提供了一种有效途径。
cs.CV / 96 / 2609.29985

OceanXL: Large-scale Underwater 3D Gaussian Splatting via Block Partitioning and Adaptive Pruning

OceanXL:基于分块划分与自适应剪枝的大规模水下3D高斯泼溅
Wang, Haoran, Cai, Shaoyu, Azzarelli, Adrian, Jiang, Zhuodong, Huang, Guoxi, Khoo, Eng Tat, Seymour, Brett, Zhang, Fan, Bull, David, Anantrasirichai, Nantheera
Abstract
Underwater 3D reconstruction is critical for marine exploration, ecological monitoring, and subsea infrastructure inspection, yet remains challenging at large scale due to light attenuation, scattering, and limited capture coverage. While 3D Gaussian Splatting (3DGS) enables high-quality real-time rendering, its application to large underwater scenes is constrained by high memory consumption and inefficient optimization over extensive areas. We propose OceanXL, a fast and scalable 3DGS-based framework for large-scale underwater reconstruction. OceanXL adopts a divide-and-conquer strategy, partitioning scenes into spatially coherent blocks to enable efficient optimization while preserving global geometric consistency. We further introduce an adaptive pruning scheme tailored to underwater conditions that removes redundant primitives, producing compact representations without sacrificing visual fidelity. Together, these components improve training efficiency and rendering performance for large scenes. We also introduce a large-scale underwater dataset covering diverse marine environments. Experiments on five large-scale scenes demonstrate favorable scalability, compactness, and efficiency--quality trade-offs over large-scene baselines. Controlled comparisons on the small-scale SeaThru-NeRF dataset further show competitive reconstruction quality with substantially smaller model sizes than underwater-specific methods.
Chinese Translation
水下三维重建对于海洋探索、生态监测和水下基础设施检查至关重要,但由于光照衰减、散射以及采集覆盖范围有限,大规模水下重建仍具挑战性。尽管三维高斯泼溅(3D Gaussian Splatting, 3DGS)能够实现高质量的实时渲染,但其在大型水下场景中的应用受限于高内存消耗以及在广阔区域上低效的优化。我们提出了OceanXL,一个快速且可扩展的基于3DGS的大规模水下重建框架。OceanXL采用分而治之的策略,将场景划分为空间上连贯的块,从而在保持全局几何一致性的同时实现高效优化。我们还引入了一种针对水下条件设计的自适应剪枝方案,能够去除冗余基元,在不牺牲视觉保真度的情况下生成紧凑的表示。这些组件共同提升了大规模场景的训练效率和渲染性能。我们还发布了一个覆盖多种海洋环境的大规模水下数据集。在五个大规模场景上的实验表明,相较于大场景基线方法,OceanXL在可扩展性、紧凑性以及效率—质量权衡方面表现优异。在小规模SeaThru-NeRF数据集上的对照比较进一步表明,我们的方法重建质量具有竞争力,且模型规模远小于水下专用方法。
cs.CV / 97 / 2609.29999

GHOST-Q: Towards Studying Grounding Hallucinations Overlooked Under Same-score TradeOffs in Quantized VLMS

GHOST-Q:研究量化视觉语言模型中被同等分数权衡所掩盖的定位幻觉问题
Rehman, Saim, Shafique, Muhammad
Abstract
Post-training quantization of vision--language models (VLMs) is typically assessed through aggregate task accuracy and memory savings, but preserving a headline score does not guarantee preservation of visual grounding behavior. We present GHOST-Q, a cross-precision controlled evaluation of three 8B VLM families under FP16, INT8, and NF4 across utility and hallucination-sensitive benchmarks. Rather than comparing only aggregate accuracy, we pair FP16 and quantized predictions item by-item to quantify how compression redistributes grounding successes and failures. Five of six quantized variants preserve MMStar accuracy within $\pm2$ percentage points, yet 10 of 36 paired effects remain significant after false-discovery-rate correction, nine on hallucination-sensitive conditions. Same-device A100 profiling further demonstrates that substantial memory reduction does not necessarily mean lower inference latency. Finally, an open-ended AMBER audit reveals strong generation budget censoring whose severity varies by architecture and precision. These results show that quantized VLMs should be evaluated jointly for aggregate utility, grounding reliability, generation behavior, and realized deployment efficiency.
Chinese Translation
视觉语言模型(VLM)的训练后量化通常通过总体任务准确率和内存节省来评估,但保持一个总体分数并不能保证视觉定位行为的保留。我们提出GHOST-Q,一个跨精度的受控评估,在FP16、INT8和NF4精度下,对三个8B VLM家族在实用性基准和幻觉敏感基准上进行评测。我们不仅仅比较总体准确率,而是将FP16与量化模型的预测逐项配对,以量化压缩如何重新分配定位的成功与失败。六个量化变体中有五个在MMStar上保持±2个百分点以内的准确率,但经过错误发现率校正后,36个配对效应中仍有10个保持显著,其中9个出现在幻觉敏感条件下。同设备A100上的性能剖析进一步表明,大幅的内存降低并不一定意味着更低的推理延迟。最后,一项开放式的AMBER审计揭示了强烈的生成预算审查现象,其严重程度因架构和精度而异。这些结果表明,量化VLM应在总体实用性、定位可靠性、生成行为和实际部署效率四个方面进行联合评估。
cs.CV / 98 / 2609.30026

Training-Free Hold-Usage Detection in Sport Climbing with Foundation Pose Models

基于基础姿态模型的运动攀岩免训练岩点使用检测
Bakar, Abu, Aftab, Abdullah, Hamza, Amir
Abstract
Detecting which holds a climber uses, and when, underpins automated scoring, movement analysis, and assistive systems for sport climbing. Existing approaches train task-specific models or repurpose 2D pose estimators whose hand keypoint sits at the wrist and foot keypoint at the ankle i.e. offset from the fingertips and toes that actually contact the holds, and whose hands are occluded in roughly half of all frames. We show that a frozen, off-the-shelf pose foundation model is sufficient: using the fingertip and toe keypoints of Sapiens, a per-frame proximity test against the annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, we detect hold usage without any climbing-specific training. On the The Way Up dataset (22 videos, 10 athletes, two routes), our method reaches an event F_1 of 90.2% on a held-out split (89.8% under leave-one-participant-out cross-validation) and 79.9% over all 22 videos at any temporal overlap, and performs best on footholds (F_1,89.8% overall, 96.6% held-out). Under an identical protocol it exceeds our reproductions of the YOLOv8-pose and ViTPose pipelines at every temporal threshold, with the margin widening under strict timing. An ablation shows that two intuitively helpful additions---dense foundation-feature change gating and body-part segmentation---both hurt, arguing that a minimal, keypoint-only design is the right one for this task. Finally, standard coaching statistics computed from our automatic predictions track ground truth closely (Pearson r=1.00 for climb time, 0.94 for pace), turning ordinary single-camera video into reliable performance metrics with no instrumentation.
Chinese Translation
检测攀岩者使用了哪些岩点以及何时使用,是运动攀岩自动评分、动作分析和辅助系统的基础。现有方法要么训练特定任务模型,要么改造二维姿态估计器,但后者手部关键点位于手腕、脚部关键点位于脚踝,即偏离了实际接触岩点的指尖和脚尖,且其手部在约一半的帧中被遮挡。我们证明,一个冻结的现成姿态基础模型即可胜任:利用 Sapiens 的指尖和脚尖关键点、针对标注岩点的逐帧邻近性测试、各肢体的互斥约束以及简短的时序持续规则,我们无需任何攀岩专项训练即可检测岩点使用情况。在 The Way Up 数据集(22 个视频、10 名运动员、两条线路)上,我们的方法在留出集上达到事件 F_1 值 90.2%(留一参与者交叉验证下为 89.8%),在任意时间重叠下全部 22 个视频上达到 79.9%,并且在脚点上表现最佳(总体 F_1 为 89.8%,留出集为 96.6%)。在相同协议下,该方法在所有时间阈值上均超过我们复现的 YOLOv8-pose 和 ViTPose 流水线,且在严格时间标准下优势进一步扩大。消融实验表明,两个看似有益的改进——稠密基础特征变化门控和身体部位分割——反而均有害,说明针对该任务,仅使用关键点的极简设计才是正确选择。最后,基于自动预测计算的标准教练统计数据与真实值高度吻合(攀爬时间的 Pearson r=1.00,节奏的 Pearson r=0.94),从而无需任何仪器设备即可将普通单目视频转化为可靠的性能指标。
cs.CV / 99 / 2609.30037

AERIAL: Adversarial Evaluation of Robustness in Accuracy-Preserving Low-Precision EEG Decoders

AERIAL:面向精度保持型低精度EEG解码器鲁棒性的对抗评估
Rehman, Saim, Shafique, Muhammad
Abstract
Deployment-oriented compression is attractive for resource-constrained brain--computer interfaces (BCIs), but whether it changes adversarial vulnerability remains unclear. On BCI Competition IV-2a, we compare 32-bit floating-point (FP32) EEGNet and ShallowConvNet models with global magnitude pruning and simulated INT8 post training quantization (PTQ) and quantization-aware training (QAT) across nine subjects and three seeds. Simulation provides differentiable quantize--dequantize models for white-box attacks and gradient analysis, while native TensorRT deployment is used for validation. Accuracy-preserving compression does not improve direct robustness: at $\epsilon=0.005$, EEGNet PGD accuracy remains 22--24\% across FP32, 50\% pruning (P50), PTQ, and QAT. However, P50 reduces bidirectional transfer efficiency to 0.963/0.928 (FP32$\rightarrow$P50/P50$\rightarrow$FP32), versus 0.994/0.997 for PTQ; the same trend holds for ShallowConvNet. Gradient alignment shows a corresponding separation, while native PTQ agrees with simulated clean/adversarial predictions in 95--98\% of cases. These results show that direct robustness, adversarial transfer, and deployment efficiency are distinct properties of compressed EEG decoders.
Chinese Translation
面向部署的压缩技术对资源受限的脑机接口(BCI)颇具吸引力,但其是否会改变模型的对抗脆弱性仍不清楚。在BCI Competition IV-2a数据集上,我们在九名受试者和三个随机种子下,将32位浮点(FP32)的EEGNet和ShallowConvNet模型与全局幅度剪枝、模拟INT8训练后量化(PTQ)以及量化感知训练(QAT)进行了比较。仿真提供了可微分的量化-反量化模型,用于白盒攻击和梯度分析,同时使用原生TensorRT部署进行验证。保持精度的压缩并不能提升直接鲁棒性:在ε=0.005时,EEGNet在FP32、50%剪枝(P50)、PTQ和QAT下的PGD准确率均保持在22–24%。然而,P50将双向迁移效率降至0.963/0.928(FP32→P50/P50→FP32),而PTQ为0.994/0.997;ShallowConvNet也呈现相同趋势。梯度对齐显示出相应的分离,而原生PTQ与模拟的干净/对抗样本预测在95–98%的情况下一致。这些结果表明,直接鲁棒性、对抗迁移性和部署效率是压缩EEG解码器的不同属性。
cs.CV / 100 / 2609.30043

ConPro: Contrast Projection Pretraining for Label-Efficient Vessel Segmentation in DSA Sequences

ConPro:面向DSA序列标签高效血管分割的对比投影预训练方法
Guo, Xinge, Wang, Yuanhao, Shu, Liqi, Liu, Yang, Xu, Min
Abstract
Dense vessel annotation in digital subtraction angiography (DSA) is labor-intensive, yet every unlabeled sequence records how contrast passes through the vessels. Semi-supervised methods take their targets from the current model, and generic self-supervised pretexts reconstruct static appearance, so this signal goes unused. We propose ConPro, a self-supervised pretraining scheme whose target is a contrast projection, the normalized drop of every pixel below its temporal median over the sequence. On DIAS and DSCA, with 10%, 20% and 50% of the training cases labeled, ConPro improves on training from scratch at every label fraction and is the best of the compared methods on DSCA at 20% and 50% labels. Controlled comparisons show that the gain comes from the target. A temporal-median target with the same input, loss and budget stays at scratch level, and using the projection directly instead of learning it, as an input channel or a pseudo-label, helps little or hurts. ConPro provides pretrained weights without changing the segmentation architecture, so it combines with semi-supervised training, and UniMatch, the strongest baseline, gains 0.5 to 2.0 Dice and 0.9 to 2.3 clDice at every label fraction when started from ConPro weights, reaching 75.4 Dice on DIAS and 81.3 on DSCA.
Chinese Translation
在数字减影血管造影(DSA)中对血管进行密集标注十分耗费人力,然而每一条未标注序列都记录了对比剂流经血管的过程。半监督方法的目标来自当前模型,而通用的自监督预训练任务(pretext)重构的是静态外观,因此这一信号一直未被利用。我们提出ConPro,一种以对比投影为目标的自监督预训练方案,其中对比投影定义为序列中每个像素低于其时间中值的归一化差值。在DIAS和DSCA数据集上,当训练病例的10%、20%和50%被标注时,ConPro在所有标注比例下均优于从零开始训练,并且在20%和50%标注比例下是DSCA上所比较方法中的最优者。受控比较实验表明,性能提升来源于该目标本身:在相同输入、损失和预算下,采用时间中值作为目标仅能达到从零训练的水平,而直接使用投影(作为输入通道或伪标签)而非学习该投影,则帮助甚微甚至产生负面影响。ConPro无需更改分割架构即可提供预训练权重,因此可与半监督训练相结合;以ConPro权重初始化时,最强基线方法UniMatch在所有标注比例下获得0.5至2.0的Dice提升和0.9至2.3的clDice提升,在DIAS上达到75.4的Dice,在DSCA上达到81.3。
cs.CV / 101 / 2609.30080

Can Frozen Hyperspherical Features Guide the Selection of Pseudo Masks?

冻结的超球面特征能否指导伪掩码的选择?
Guo, Xinge, Xiao, Fengyang, Zhang, Dingming, Chen, Yuhan, Zhang, Rihan, Li, Xingjian, Wang, Tianyang, He, Chunming, Farsiu, Sina
Abstract
Foundation segmenters such as SAM return several plausible masks for an unlabeled image, and a student trained on the wrong one inherits its errors. Choosing among them means querying a second large model or fitting a quality head to annotated masks. We show that a candidate can be judged by what it does to a frozen self-supervised backbone's features. Normalized DINOv2 patch features lie on a hypersphere, and a candidate mask splits that sphere in two. Based on this reading, we introduce SphereTrust, which scores each candidate by three properties of the split, the angular contrast between the two sides, the coverage of the foreground's appearance modes, and contact with the image frame, one for each of three common ways a mask fails, and ranks a pool in 0.55 s per image from the frozen features alone. On eight SAM and SAM3 candidate pools spanning camouflaged, salient, and dichotomous segmentation and camouflage under low light, SphereTrust exceeds the strongest evaluated external baseline on six pools by 1.7 to 9.3 percentage points in mean selected Dice. These comparisons include published selection rules and explicitly labeled adaptations of DSS and UCOD-MKD. On the two prompted camouflage pools, its mean selected Dice is within 0.1 percentage points of the candidate-derived DSS adaptation, with a lower catastrophic-error rate. Which cue carries the signal depends on the candidate pool. The same sphere also supports training. The leading candidates enter as a candidate set with their scores as priors, prototypes reorder them, and a cross-fitted second round completes the labels, raising weighted F by 4.5, 2.3, and 5.5 points over fixed-label training on the three MLLM anchor pools, with students competitive with published unsupervised methods on nineteen test sets.
Chinese Translation
SAM等基础分割模型会为一张无标签图像返回多个看似合理的掩码,而学生在错误掩码上训练会继承其错误。在这些掩码中进行选择,通常需要查询另一个大模型,或在有标注掩码上拟合一个质量评估头。我们证明,可以通过候选掩码对冻结的自监督骨干网络特征的影响来评判其优劣。归一化的DINOv2图像块特征分布在一个超球面上,而候选掩码会将该球面一分为二。基于这一解读,我们提出SphereTrust,它根据该分割的三个属性对每个候选掩码打分——两侧之间的角度对比度、前景外观模式的覆盖度、以及与图像边框的接触程度——分别对应掩码失效的三种常见方式,并仅依据冻结特征在每张图像上以0.55秒的速度对候选池进行排序。在八个SAM和SAM3候选池上,涵盖伪装、显著和二分分割以及低光照条件下的伪装分割,SphereTrust在六个候选池上的平均所选Dice指标超过最强的外部对比基线1.7至9.3个百分点。这些对比包括已发表的选择规则,以及明确标注的DSS和UCOD-MKD适配版本。在两个带提示的伪装候选池上,其平均所选Dice与基于候选的DSS适配版本差距在0.1个百分点以内,且灾难性错误率更低。哪种线索携带有效信号取决于候选池本身。同一超球面还可用于支持训练。排序靠前的候选掩码作为候选集进入训练,其得分作为先验,原型机制对其重新排序,再由交叉拟合的第二轮补全标签,在三个MLLM锚点候选池上,相比固定标签训练,加权F指标分别提升4.5、2.3和5.5个百分点,且所得学生在十九个测试集上的表现可与已发表的无监督方法相媲美。
cs.CV / 102 / 2609.30096

Accelerating Video Diffusion via Training-Free Trajectory Routing

基于免训练轨迹路由的视频扩散加速方法
Munir, Mustafa, Vu, Huy, Misra, Shreyas, Jena, Rohit, Norouzi, Sajad, Taghibakhshi, Ali, Ahmad, Anis, Patney, Anjul, Molchanov, Pavlo, Tajbakhsh, Nima
Abstract
Video diffusion is computationally expensive, as it requires executing a large model across many denoising steps. Even with step-distillation, inference remains expensive because every distilled step still requires a costly model evaluation. We present TRACK: TRajectory-Aware Capacity routing via top-K selection, a heterogeneous denoising strategy that switches between compatible large and small models at selected steps, reducing the average cost per denoising evaluation. The switching steps are determined using a calibration process. TRACK first rolls out a reference trajectory with the large model. Then at each step, the small model's prediction is also collected and compared against the large model's prediction to obtain a relative disagreement score. Both models receive the same latent, timestep, conditioning, and guidance inputs. Aggregating this signal over a calibration set produces a disagreement score map across diffusion steps, which determines a switching policy for an efficient inference process: quality-sensitive steps keep using the large model, while steps with low disagreement scores are routed to the small model. Inference executes only the selected model at each step, requiring no retraining, architecture or scheduler changes, or online dual-model evaluation. Across Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo, TRACK yields $1.95\times$, $2.04\times$-$2.73\times$, $2.69\times$, and $2.17\times$ speedups, respectively, with comparable aggregate quality and high diversity retention. TRACK thereby establishes automated, training-free model switching as a practical acceleration paradigm for video diffusion.
Chinese Translation
视频扩散模型的计算开销巨大,因为它需要在多个去噪步骤中执行大型模型。即使采用步数蒸馏,推理开销依然很高,因为每个蒸馏后的步骤仍需要昂贵的模型评估。我们提出了TRACK(TRajectory-Aware Capacity routing via top-K selection,基于top-K选择的轨迹感知容量路由),这是一种异构去噪策略,可在选定的步骤间切换相互兼容的大模型与小模型,从而降低每次去噪评估的平均开销。切换步骤通过校准过程确定。TRACK首先使用大模型生成一条参考轨迹;然后在每个步骤中,同时收集小模型的预测结果,并将其与大模型的预测进行比较,以获得相对分歧分数。两个模型接收相同的潜变量、时间步、条件输入和引导输入。在校准集上聚合该信号,可得到跨越扩散步骤的分歧分数图,进而确定高效推理过程的切换策略:对质量敏感的步骤继续使用大模型,而分歧分数较低的步骤则路由至小模型。推理时每个步骤仅执行所选模型,无需重新训练、更改架构或调度器,也无需在线双模型评估。在Wan 2.1、Cosmos 3、TurboDiffusion和FastVideo上,TRACK分别实现了1.95倍、2.04倍至2.73倍、2.69倍和2.17倍的加速,同时保持相当的总体质量和高多样性保留度。由此,TRACK确立了自动化、免训练的模型切换作为视频扩散的一种实用加速范式。
cs.CV / 103 / 2609.30107

Smartphone-Based Method for Automated Speed Enforcement

基于智能手机的自动化超速执法方法
Li, Keya, Malagavalli, Jahnavi, Goel, Lamha, Wang, Tong, Kockelman, Kara M.
Abstract
Smartphone cameras and computer vision (CV) hold significant promise in assisting public agencies with enforcing traffic laws and enhancing road safety. This work designs and tests a smartphone-based method for automated speed estimation and vehicle identification (license plate, make/model, and color recognition) via an automated pipeline to assist enforcement agencies in reliably identifying speeders. The CV code accurately recognizes nearly half (46%) of the license plates' text on 1,800 images from a Brazil open-source dataset, called UFPR-ALPR. Code tests on daytime recordings from hand-held smartphone videos (n = 73) and roadside cameras (n = 42) in Austin, Texas yield 60.8% accuracy for color detection (among all possible RGB color categories), 48.6% on vehicle make/manufacturer identification, and 16.89% on vehicle make and model identification. Prediction accuracy for speed estimation (within a 20% range), vehicle make (within the top 3 predictions), and license plate recognition (within the top 10 predictions) are 16.3%, 16.9%, and 29.7%, respectively. This paper also illuminates the legal, technological, and practical aspects of using smartphones for enforcement, including the potential use of recordings for enforcement purposes, emphasizing the need to transform the potential of smartphone-based CV technologies into practical tools for vital information on traffic violations.
Chinese Translation
智能手机摄像头与计算机视觉(CV)技术在协助公共机构执行交通法规和提升道路安全方面具有重大潜力。本研究设计并测试了一种基于智能手机的方法,通过自动化流程实现车速估算与车辆识别(包括车牌、品牌/型号和颜色识别),以协助执法机构可靠地识别超速车辆。该计算机视觉代码在来自巴西开源数据集UFPR-ALPR的1,800张图像上,准确识别了近一半(46%)车牌文本。在德克萨斯州奥斯汀的手持智能手机视频(n = 73)和路边摄像头录像(n = 42)的日间录像代码测试中,颜色检测(在所有可能的RGB颜色类别中)准确率为60.8%,车辆品牌/制造商识别准确率为48.6%,车辆品牌和型号识别准确率为16.89%。车速估算(误差在20%范围内)、车辆品牌(前3个预测之内)和车牌识别(前10个预测之内)的预测准确率分别为16.3%、16.9%和29.7%。本文还阐明了使用智能手机进行执法的法律、技术和实践层面的问题,包括录像用于执法目的的潜在用途,并强调需要将基于智能手机的计算机视觉技术的潜力转化为可用于交通违规关键信息的实用工具。
cs.CV / 104 / 2609.30130

Multimodal Thinking with Renderable Programs

基于可渲染程序的多模态思考
Chen, Sunli, Zhong, Ding, Ma, Ziqiao, Liu, Jiaxin, Yang, Zeyuan, Zhang, Hao, Lu, Lie, Chai, Joyce, Gan, Chuang
Abstract
Current vision-language models (VLMs) excel at visual content understanding and text-based reasoning, yet their structure limits the advancement of incorporating images into the reasoning chain. Though Omnimodal models have made efforts in unifying text and image generation, they focus on visual tasks in the open-domain, lacking tractability due to rasterized or latent representations of images. We introduce SVGLM, a framework that uses scalable vector graphics (SVG) primitives to connect text and image in reasoning tasks. We exploit the duality of SVG as both image description and text instructions, yielding a more compact, interpretable solution to equip general VLMs with the capability of generating images within the reasoning process. We provide a large curated dataset of SVG-based image editing dataset, as well as the paradigm to tune open-source VLMs. Experiments on a mathematical reasoning benchmark demonstrate that SVGLM achieves strong SVG generation power as well as think-with-image intelligence. Our results highlight SVG as a suitable medium for building more robust digital domain agents, bridging the gap between text-based thinking and pixel-based images.
Chinese Translation
当前的视觉语言模型(VLM)在视觉内容理解和基于文本的推理方面表现出色,但其结构限制了将图像纳入推理链条的发展。尽管全模态(Omnimodal)模型已在统一文本与图像生成方面做出努力,但它们主要聚焦于开放域的视觉任务,且由于图像采用光栅化或潜在表示而缺乏可处理性。我们提出SVGLM,一个利用可缩放矢量图形(SVG)基元在推理任务中连接文本与图像的框架。我们利用SVG兼具图像描述与文本指令的双重特性,提供了一种更紧凑、更可解释的方案,使通用VLM具备在推理过程中生成图像的能力。我们构建了一个大规模、经精心整理的基于SVG的图像编辑数据集,并给出了微调开源VLM的范式。在数学推理基准上的实验表明,SVGLM不仅具备强大的SVG生成能力,还展现出以图像进行思考的智能。我们的结果凸显了SVG作为构建更强大数字域智能体的合适媒介,弥合了基于文本的思考与基于像素的图像之间的鸿沟。
cs.CV / 105 / 2609.30187

Ego-Exo4D Human Meshes Dataset: 4D Human Motion Reconstruction for Ego-Exo Captures

Ego-Exo4D 人体网格数据集:面向自我中心-多视角(Ego-Exo)采集的 4D 人体运动重建
Maddukuri, Abhiram, Pavlakos, Georgios
Abstract
Ego-Exo4D is a large-scale dataset providing synchronized egocentric and multi-view exocentric video, a rich resource for skill learning and assessment, procedural activity understanding, and embodied AI. However, the dataset ships with only sparse 3D human pose annotations, and reconstructing dense human motion from its multi-view captures is nontrivial. To this end, we present Ego-Exo4D-HM, a large-scale dataset of 4D human motion reconstructions for Ego-Exo4D's captures, and release the accompanying reconstruction pipeline. The code, dataset, and documentation can be found at https://abhiram824.github.io/egoexo4d_human_meshes.
Chinese Translation
Ego-Exo4D 是一个大规模数据集,提供同步的自我中心(egocentric)视频与多视角外中心(exocentric)视频,是技能学习与评估、程序性活动理解以及具身智能(Embodied AI)的丰富资源。然而,该数据集仅附带稀疏的 3D 人体姿态标注,从其多视角采集数据中重建稠密人体运动并非易事。为此,我们提出了 Ego-Exo4D-HM,一个针对 Ego-Exo4D 采集数据的大规模 4D 人体运动重建数据集,并发布了相应的重建流程。代码、数据集和文档可在 https://abhiram824.github.io/egoexo4d_human_meshes 获取。
cs.CV / 106 / 2609.30210

The Alignment Illusion in Multimodal Large Language Models

多模态大语言模型中的对齐错觉
Wang, Hong-Han, Wang, Yuntao, Ding, Hu
Abstract
Layer-wise visual-text similarity in Multimodal Large Language Models (MLLMs) is widely interpreted as evidence that the language model progressively integrates visual content into a shared representation space. This reading rests on the assumption that scalar alignment scores reflect content-level cross-modal interaction. To test this assumption, we apply controlled interventions to the visual stream. Across 13 MLLMs from five families spanning 0.5B to 72B parameters, replacing projector-output visual tokens with Gaussian noise sharply reduces task accuracy, yet four standard scalar measures (CKA, SVCCA, MIR, and the leading principal-angle cosine) fail to consistently separate the corrupted stream from the original. We call this failure the alignment illusion and trace it to the shared language-model pathway: anisotropic MLP down-projections pull visual and text tokens toward common output directions, producing weight-induced alignment. Because this component is essentially one-dimensional, we introduce the principal-angle gap (PA gap), defined as the difference between the top two principal-angle cosines, which separates weight-induced similarity from multi-directional visual structure. Under graded visual corruption, the PA gap tracks task accuracy more consistently than the scalar scores we consider; under a structured but irrelevant image, it further exposes regimes in which internal geometry and task accuracy come apart. Internal visual-text alignment in MLLMs is therefore best read as a geometric diagnostic of the visual stream inside the language model rather than a direct proxy for content-level cross-modal interaction, and is most informative when calibrated by controlled task evidence.
Chinese Translation
多模态大语言模型(MLLM)中逐层的视觉-文本相似性被广泛解读为语言模型逐步将视觉内容整合到共享表示空间中的证据。这一解读基于一个假设:标量对齐分数能够反映内容层面的跨模态交互。为检验这一假设,我们对视觉信息流施加了受控干预。在来自五个模型家族、参数量从0.5B到72B的13个MLLM上,用高斯噪声替换投影器输出的视觉Token会显著降低任务准确率,然而四种标准标量度量(CKA、SVCCA、MIR以及首主角度余弦)均无法一致地区分被破坏的信息流与原始信息流。我们将这种失效称为对齐错觉,并将其根源追溯到共享的语言模型通路:各向异性的MLP下投影将视觉Token和文本Token拉向共同的输出方向,产生由权重导致的对齐。由于该成分本质上是一维的,我们引入主角度间隔(PA gap),其定义为首两个主角度余弦之差,用以区分权重导致的相似性与多方向的视觉结构。在分级视觉破坏下,PA gap比我们所考虑的标量分数更能一致地跟踪任务准确率;在一幅结构化但无关的图像下,它进一步揭示了内部几何结构与任务准确率相互分离的情形。因此,MLLM内部的视觉-文本对齐最好被理解为语言模型内部视觉信息流的几何诊断指标,而非内容层面跨模态交互的直接代理,且在结合受控任务证据进行校准时才最具信息量。
cs.CV / 107 / 2609.30221

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

WanPE:面向现代文生视频的电影级提示词增强
Zhu, Yubo, Shao, Yawen, Dai, Ziyun, Fang, Zixun, Zhu, Kai, Sun, Siyang, Xue, Haolan, Wang, Chuxin, Weng, Tingyu, Luo, Jingming, Shi, Chen, Huang, Lianghua, Ai, Yufeng, Wang, Yuzheng, Zhang, Wenyuan, Shang, Yu, Bao, Yuxiang, Bi, Zoubin, Xiao, Jie, Xing, Jinbo, Zhao, Jiaxing, Zhong, Chongyang, Chen, Hengjian, Xie, Chenwei, Liu, Akide, Kan, Zhehan, Liu, Yu, Zhai, Wei, Zhong, Sheng, Tong, Wei
Abstract
Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by approximately 11K blind pairwise assessments. When powering Wan3.0's video generator, WanPE-397B boosts human preference over raw user prompts by 10.66-18.84 points at 5-15 seconds and by a dramatic 50.86 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.
Chinese Translation
视频生成始于文本空间,通过撰写电影剧本,再将其具化为像素。随着当代视频生成器扩展至30秒并能忠实遵循复杂条件,文本提示词在很大程度上主导了整个制作过程,规划动作、相机轨迹、灯光和声音如何在多镜头序列中展开。在本文中,我们提出了WanPE,一个在105万真实世界视频上训练的3970亿参数提示词增强模型,以掌握导演级的电影规划能力。WanPE通过基于视频的反向构建制定镜头级电影规划,并采用语义一致性GRPO(Semantic-Consistency GRPO, SC-GRPO)在镜头间和时间维度上忠实保留用户需求。为评估这一能力,我们构建了WanPEval,一个人工标注的测试平台,涵盖5至30秒的不同时长和不同意图粒度,并由约1.1万次盲测成对评估提供支持。作为Wan3.0视频生成器的驱动,WanPE-397B在5-15秒场景下将人类偏好相较原始用户提示词提升10.66-18.84分,在30秒场景下更是大幅提升50.86分。消融实验表明,反向构建相较正向改写具有明显优势,而SC-GRPO在不同模型规模下均能稳健地保持语义保真度。最终,WanPE在5-15秒场景下超越所有被评估的商业产品,并在30秒场景下与Seedance 2.5保持竞争力。
cs.CV / 108 / 2609.30222

TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations

TrackEverything:通过去重三维场景表示实现长时程稠密跟踪
Jain, Ayush, Paruchuri, Sreeharsha, Gupta, Ishita, Zhang, Fan, Schmidt, Tanner, Engel, Jakob, Fragkiadaki, Katerina, Harley, Adam W.
Abstract
Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates. Grounded in the insight that videos are 2D projections of an underlying 3D world, TrackEverything decouples model complexity from video duration, allowing it to scale with unique physical scene geometry instead. Our approach introduces three key innovations. First, we employ a voxelization-based de-duplication mechanism at sliding-window boundaries to merge co-located tracks, preventing repeated observations of the same surface from redundantly accumulating. Second, we decompose tracking into an endpoint refiner that predicts each point's destination and static-versus-dynamic classification, followed by a lightweight trajectory refiner that decodes dense trajectories exclusively for dynamic points. Third, we propose 3D WAFT, replacing memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud. To the best of our knowledge, TrackEverything is the first 3D tracker capable of tracking all visible points across videos exceeding 1000 frames within 40 GB of GPU memory. On TAPVid-3D, TrackEverything outperforms all open-source all-frame dense 3D trackers by more than 20% APD on short clips, while remaining competitive with state-of-the-art sparse trackers on long sequences, despite tracking far more points.
Chinese Translation
现有的点跟踪模型面临一个根本性的权衡:它们要么能在长时序上跟踪稀疏的查询点集,要么只能在短片段上跟踪所有点。我们提出TrackEverything,一种三维点跟踪器,通过将视频表示为世界坐标系下持久的三维场景轨迹,打破了这一权衡。基于视频是底层三维世界的二维投影这一洞察,TrackEverything将模型复杂度与视频时长解耦,使其能够随独特的物理场景几何进行扩展。我们的方法引入了三项关键创新。首先,我们在滑动窗口边界处采用基于体素化的去重机制来合并共置轨迹,防止对同一表面的重复观察冗余累积。其次,我们将跟踪分解为:一个端点精化器,用于预测每个点的目标位置以及静态-动态分类;随后是一个轻量级的轨迹精化器,仅针对动态点解码稠密轨迹。第三,我们提出3D WAFT,以场景点云中高效的特征采样取代内存开销极大的四维相关体积。据我们所知,TrackEverything是首个能够在40 GB GPU内存内、跨超过1000帧的视频跟踪所有可见点的三维跟踪器。在TAPVid-3D基准上,TrackEverything在短片段上以超过20%的APD优于所有开源的全帧稠密三维跟踪器,同时在长序列上与最先进的稀疏跟踪器保持竞争力,尽管其跟踪的点数要多得多。
cs.CV / 109 / 2609.30223

BiCC: Bidirectional Connected-Component Loss for Instance-Aware Segmentation

BiCC:用于实例感知分割的双向连通分量损失
Bouteille, Luc, Jonske, Frederic, Kleesiek, Jens, Jaus, Alexander
Abstract
Common segmentation losses aggregate errors voxel-wise, so lesions influence the objective in proportion to their volume, giving small but clinically critical lesions disproportionately little weight. Instance-aware losses aim to address this mismatch by assigning each lesion its own term. However, blob loss and CC-DiceCE derive their regions solely from annotations, so false-positive components receive no instance-level term. This matters in computer-assisted review, where each false-positive component may require separate inspection, making precision and false-positive burden important alongside recall. We introduce the bidirectional connected-component loss (BiCC), which pairs annotation- and prediction-derived partitions to score predicted components on their own scale. By deriving instances from the predictions, this branch directly penalizes false-positive components regardless of their size. The balance parameter $\alpha$ allows control over the lesion-wise precision-recall trade-off. Across five datasets with five-fold cross-validation using nnU-Net, BiCC outperforms CC-DiceCE in lesion-wise F1 on four datasets and blob loss on all five. It significantly improves over DiceCE on three datasets and matches it on two; CC-DiceCE instead loses up to 0.363 precision by favoring recall. Code is available at https://github.com/TIO-IKIM/BiCC-Loss.
Chinese Translation
常见的分割损失以体素为单位聚合误差,因此病灶对目标函数的影响与其体积成正比,使得体积虽小但临床意义重大的病灶所获得的权重明显偏低。实例感知损失旨在通过为每个病灶分配独立的损失项来解决这种不匹配问题。然而,blob loss 和 CC-DiceCE 仅从标注中获取区域,因此假阳性连通分量不会得到实例级别的损失项。这一点在计算机辅助审阅中尤为重要,因为每个假阳性分量都可能需要单独检查,使得精确率和假阳性负担与召回率同样重要。我们提出了双向连通分量损失(Bidirectional Connected-Component Loss,BiCC),它将基于标注的划分与基于预测的划分配对,在预测分量自身的尺度上对其进行评分。通过从预测中导出实例,该分支能够直接惩罚假阳性分量,而无论其大小如何。平衡参数 $\alpha$ 允许对病灶级别的精确率-召回率权衡进行控制。在使用 nnU-Net 的五个数据集上的五折交叉验证中,BiCC 在四个数据集上的病灶级别 F1 分数优于 CC-DiceCE,并在全部五个数据集上优于 blob loss。在三个数据集上,BiCC 显著优于 DiceCE,在两个数据集上与其持平;而 CC-DiceCE 因偏向召回率,其精确率最多损失了 0.363。代码可在 https://github.com/TIO-IKIM/BiCC-Loss 获取。
cs.CV / 110 / 2609.30234

OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction

OmniFabric:面向三维服装重建的相干UV空间纹理合成
Huang, Ding-Jiun, Wang, Yuanhao, Zhang, Cheng, Bertiche, Hugo, Ichim, Alexandru-Eugen, Beeler, Thabo, De la Torre, Fernando
Abstract
Automated generation of production-ready 3D garment assets from a single image is a central challenge in digital content creation. While recent generative models have significantly advanced 3D geometry reconstruction, synthesizing high-quality textures remains a bottleneck. Existing methods often bake environmental illumination and shadows directly into the texture map, or they fail to maintain global structural coherence, making the resulting assets unusable for physical simulation and relighting. In this work, we introduce OmniFabric, a novel approach that synthesizes globally coherent texture maps directly within the 2D sewing pattern space. Given a single reference image, our pipeline utilizes an estimated 3D mesh and generative priors of powerful Vision-Language Models (VLM) to establish a complete but coarse texture initialization across the unwrapped sewing patterns. We then leverage a specialized diffusion transformer, trained via an automated synthetic data engine and conditioned on 3D positional features, to refine this initialization directly in the canonical UV domain. This effectively removes distortion and baked-in artifacts to extract a clean and normalized texture map that preserves the original garment design. Extensive experiments demonstrate that OmniFabric significantly outperforms state-of-the-art baselines, yielding photorealistic 3D garments with high-quality textures.
Chinese Translation
从单张图像自动生成可用于生产的3D服装资产是数字内容创作中的核心挑战。尽管近期的生成式模型在3D几何重建方面取得了显著进展,但高质量纹理的合成仍然是一个瓶颈。现有方法往往将环境光照和阴影直接烘焙到纹理贴图中,或者无法保持全局结构的一致性,导致生成的资产无法用于物理仿真和重光照。在本工作中,我们提出了OmniFabric,一种直接在2D服装版片(sewing pattern)空间内合成全局一致纹理贴图的新方法。给定单张参考图像,我们的流程利用估计的3D网格以及强大的视觉-语言模型(VLM)的生成先验,在展开的服装版片上建立完整但粗糙的纹理初始化。随后,我们借助一个专门的扩散Transformer(diffusion transformer)——通过自动化合成数据引擎训练,并以3D位置特征作为条件——在规范UV域中直接对该初始化进行细化。这有效消除了畸变和烘焙痕迹,从而提取出干净且规范化的纹理贴图,同时保留原始服装设计。大量实验表明,OmniFabric显著优于当前最先进的基线方法,能够生成具有高质量纹理的照片级逼真3D服装。
cs.CV / 111 / 2609.30245

Towards Practical Compression of 3D Gaussian Splatting

迈向实用的3D高斯泼溅压缩
Yu, Pengpeng, Chen, Yueru, Song, Fei, Qin, Tai, Zhang, Qi, Wang, Jing, Guo, Yulan
Abstract
3D Gaussian Splatting (3DGS) enables high-quality novel-view synthesis but requires substantial storage. Existing compression methods often rely on spatial context modeling over irregular 3D representations, increasing the complexity of training and coding. Meanwhile, floating-point context inference can introduce numerical inconsistencies across platforms, causing entropy-decoding failures. To address these practical challenges, we propose COSA-GS, which constructs context without spatial aggregation through anchor-wise causal factorization. Specifically, we use geometry context derived from each anchor's coordinates to model a compact learnable anchor latent. The anchor latent is then fused with the geometry context to form an anchor context for attribute coding. The resulting context model features a simple architecture composed solely of linear transformations and activations. We train COSA-GS using rate--distortion optimization with adaptive Gaussian pruning. Further, we develop quantization-aware training and integer inference for the context model to achieve bit-exact consistency of entropy-decoded symbols across platforms. Experiments demonstrate that COSA-GS achieves state-of-the-art compression performance while retaining fast and consistent cross-platform decoding, providing a simple yet effective framework for practical 3DGS compression. Code is available at https://github.com/pengpeng-yu/COSA-GS.
Chinese Translation
3D高斯泼溅(3D Gaussian Splatting, 3DGS)能够实现高质量的novel-view synthesis,但需要大量的存储空间。现有的压缩方法通常依赖于对不规则3D表示进行空间上下文建模,增加了训练和编码的复杂性。同时,浮点上下文推理可能在跨平台时引入数值不一致性,导致熵解码失败。为了解决这些实际挑战,我们提出了COSA-GS,通过基于锚点的因果分解构建上下文,而无需空间聚合。具体而言,我们使用从每个锚点坐标导出的几何上下文来建模紧凑的可学习锚点潜在表示。然后将锚点潜在表示与几何上下文融合,形成用于属性编码的锚点上下文。由此得到的上下文模型具有仅由线性变换和激活函数组成的简单架构。我们使用率-失真优化结合自适应高斯剪枝来训练COSA-GS。此外,我们为上下文模型开发了量化感知训练和整数推理,以实现跨平台熵解码符号的比特级精确一致性。实验表明,COSA-GS实现了最先进的压缩性能,同时保持了快速且一致的跨平台解码,为实用的3DGS压缩提供了一个简单而有效的框架。代码可在 https://github.com/pengpeng-yu/COSA-GS 获取。
机器学习 (Machine Learning)
119
cs.LG / 1 / 2609.28502

Stable and Faithful Explanations for Knowledge Tracing

知识追踪的稳定且忠实的解释
Padi, Praveena, Morampudi, Arun, Irrinki, Ujval Sai Gopal, Kakitapelli, Pradeep Kumar Dolabehera
Abstract
Knowledge tracing (KT) models predict student performance opaquely, limiting pedagogical action. This study contributes a validation protocol testing predictive competitiveness (RQ1), explanation stability (RQ2) and retraining-based faithfulness (RQ3) together. Thirteen behavioral features across five pedagogical themes were engineered from ASSISTments 2009 and 2012, with history features computed from temporally preceding interactions and current response latency retained only for retrospective analysis. ASSISTments 2009 was rebuilt: the uncorrected skill-builder release duplicates each multi-skill interaction across one row per skill, and because those rows share one correctness label, they leak it into preceding-interaction features. Rebuilding lowered model AUC and reordered the explanation results. An Extreme Gradient Boosting (XGBoost) model explained with Tree SHapley Additive exPlanations (TreeSHAP) was compared against four deep baselines (DKT, SAKT, AKT and SimpleKT) under an information-matched protocol giving the deep models the same behavioral signals and restricting XGBoost to what is derivable from the identifier-and-correctness stream they consume. XGBoost reached an area under the curve (AUC) of 0.777 on 2012 and 0.786 on rebuilt 2009, with prediction-time AUCs of 0.771 and 0.775, respectively, after excluding current response latency; restricted to the baselines' information it performed as they did (0.697 against 0.700, and 0.717 against 0.720), locating the difference in information supplied, not model family. Rankings were consistent across folds, seeds and conditioning schemes (Spearman rho = 0.989-1.000), and removing top-ranked TreeSHAP features harmed AUC more than random removal, though split-gain and permutation rankings performed comparably. Student-level examples are illustrative interpretations, not validated recommendations.
Chinese Translation
知识追踪模型对学生学习表现的预测是不透明的,这限制了教学行动的实施。本研究提出了一套验证协议,联合检验预测竞争力(RQ1)、解释稳定性(RQ2)和基于重训练的忠实性(RQ3)。研究基于 ASSISTments 2009 和 2012 数据集构建了涵盖五个教学主题的十三个行为特征,其中历史特征由时间上先前的交互计算得出,而当前作答耗时仅保留用于回顾性分析。研究对 ASSISTments 2009 数据集进行了重建:未经修正的 skill-builder 版本将每个多技能交互按每技能一行的方式重复记录,由于这些行共享同一正确性标签,该标签会泄露到先前交互特征中。重建降低了模型 AUC 并改变了特征重要性排序结果。本研究将采用 Tree SHapley Additive exPlanations(TreeSHAP)进行解释的极端梯度提升模型与四个深度学习基线模型(DKT、SAKT、AKT 和 SimpleKT)在信息对齐协议下进行比较:深度模型获得相同的行为特征信号,同时将 XGBoost 限制在仅使用可由基线模型所消费的标识符与正确性数据流推导出的特征。XGBoost 在 2012 数据集上达到 0.777 的曲线下面积(AUC),在重建后的 2009 数据集上达到 0.786;在排除当前作答耗时后,预测时 AUC 分别为 0.771 和 0.775。当限制为基线模型相同的信息时,其表现与基线模型相当(0.697 对 0.700,0.717 对 0.720),表明差异源于所提供的信息而非模型类别。特征排名在交叉验证折、随机种子和条件化方案下保持一致(Spearman 相关系数介于 0.989 至 1.000 之间),且移除排名靠前的 TreeSHAP 特征对 AUC 的损害大于随机移除,尽管分裂增益排序和置换排序的表现相当。学生层面的案例仅作为示例性解释,而非经过验证的教学建议。
cs.LG / 2 / 2609.28553

SMILESGNN: Interpretable Clinical Toxicity Prediction via SMILES-Graph Cross-Attention Fusion

SMILESGNN:基于SMILES-图交叉注意力融合的可解释临床毒性预测
Nguyen, Quang Minh, Nguyen, Thuy Quynh, Le, Duc Minh, Nguyen, Ho Nhat Minh, Doan, Thanh Long Dai, Nguyen, Trong Nghia
Abstract
Drug toxicity prediction is critical for reducing late-stage attrition in drug discovery, yet remains challenging due to severe class imbalance, scaffold-based generalization, and the clinical need for interpretable predictions. Single-modality approaches-SMILES Transformers or graph neural networks capture complementary aspects of molecular structure, while sequence-only models cannot directly provide graph-attributed explanations. We present SMILESGNN, a multimodal architecture that fuses a SMILES Transformer encoder and a GATv2 graph encoder via cross-attention, and SMILESGNN-PT, a variant using a ChemBERTa-2 pretrained backbone. The design retains an explicit graph branch within the predictive pipeline, supporting GNNExplainer-based analysis of substructures associated with toxic predictions. On ClinTox, SMILESGNN achieves AUC-ROC 0.987 and F1 0.906 with only 0.4M parameters, performing competitively with a strong SMILESTransformer and a larger ChemBERTa-2/GATv2 concat-fusion baseline. On Tox21 (12 tasks), SMILESGNN-PT obtains mean AUC-ROC 0.750, comparable to ChemBERTa-2 alone and the same-backbone concat-fusion baseline. Overall, the results suggest that cross-attention is a practical fusion alternative that preserves competitive predictive performance while enabling graph-based interpretability support.
Chinese Translation
药物毒性预测对于减少药物研发后期的失败率至关重要,但由于严重的类别不平衡、基于骨架的泛化问题以及临床对可解释预测的需求,该任务仍具有挑战性。单模态方法——SMILES Transformer或图神经网络——各自捕捉分子结构的互补方面,而仅基于序列的模型无法直接提供图归因解释。我们提出了SMILESGNN,一种多模态架构,通过交叉注意力融合SMILES Transformer编码器与GATv2图编码器;同时提出了SMILESGNN-PT,一种采用ChemBERTa-2预训练骨干网络的变体。该设计在预测流程中保留了显式的图分支,支持基于GNNExplainer对毒性预测相关子结构的分析。在ClinTox数据集上,SMILESGNN仅以0.4M参数实现了AUC-ROC 0.987和F1 0.906,与强大的SMILES Transformer以及更大的ChemBERTa-2/GATv2拼接融合基线相比具有竞争力。在Tox21(12个任务)上,SMILESGNN-PT取得平均AUC-ROC 0.750,与单独使用的ChemBERTa-2及相同骨干网络的拼接融合基线相当。总体而言,结果表明交叉注意力是一种实用的融合替代方案,在保持有竞争力的预测性能的同时,能够提供基于图的可解释性支持。
cs.LG / 3 / 2609.28558

CFD Correction of Open Tip Clearance Flow in a Compressor Cascade Using VAE Latent Space Adaptation

基于VAE潜在空间自适应的压气机叶栅开口叶顶间隙流动CFD修正
Zuo, Xiang, Deng, Hefang, Chen, Caiyan, He, Honglin, Zhu, Mingmin, Zhang, Songan, Teng, Jinfang
Abstract
CFD predictions of open tip clearance flow in compressor cascades are subject to discrepancies relative to experiments, while experimental observations are sparse and high-resolution experimental ground truth is unavailable. This study proposes a non-intrusive correction method based on a variational autoencoder (VAE) and latent-space adaptation. A VAE is first trained using a dataset of 166 parametrically sampled CFD total pressure loss fields to learn a low-dimensional statistical representation of these fields. The VAE is then frozen, and a low-rank latent-space adapter is trained using only 12 paired CFD--experiment operating conditions. An observation operator maps the corrected high-resolution fields to the experimental observation space, allowing supervision to be applied only at the available measurement locations and within the measured pitchwise windows. In the current 12-fold cross-validation, the mean absolute error decreases from 0.1335 to 0.0473, the root mean square error from 0.1717 to 0.0621, and the relative $L_2$ error from 0.5108 to 0.1871. These results indicate that the method improves agreement between CFD predictions and sparse experimental observations of open tip clearance flow without modifying the RANS solver or constructing artificial high-resolution experimental labels.
Chinese Translation
压气机叶栅中开口叶顶间隙流动的CFD预测结果与实验数据存在偏差,而实验观测数据稀疏,且缺乏高分辨率的实验真值。本研究提出了一种基于变分自编码器(VAE)和潜在空间自适应的无侵入式修正方法。首先,利用包含166个参数化采样CFD总压损失场的数据集训练VAE,以学习这些场的低维统计表示。随后冻结该VAE,仅使用12组CFD—实验配对工况训练一个低秩潜在空间适配器。通过观测算子将修正后的高分辨率场映射到实验观测空间,从而仅在可用的测量位置和测量的节距范围内施加监督。在当前12折交叉验证中,平均绝对误差从0.1335降至0.0473,均方根误差从0.1717降至0.0621,相对$L_2$误差从0.5108降至0.1871。这些结果表明,该方法在不修改RANS求解器或不构建人工高分辨率实验标签的情况下,提升了CFD预测与开口叶顶间隙流动稀疏实验观测之间的吻合度。
cs.LG / 4 / 2609.28561

CARE: Condition-Aware Representation Regularization for Diffusion Models

CARE:面向扩散模型的条件感知表示正则化
Guo, Fengjia, Yang, Zhuoyi, Tang, Jie
Abstract
Recent advances in diffusion models highlight the importance of representation regularization for improving sample quality and training efficiency. However, commonly used regularization methods often overlook the built-in conditions (such as labels or texts) which directly determine the generation target. In this work, we demonstrate how conditioning signals affect the feature distribution and introduce the CARE (Condition-Aware REpresentation regularization). CARE is a lightweight plug-and-play regularization framework that dynamically modulates feature distribution based on condition similarity. CARE leverages built-in conditioning signals to judiciously guide the representation space, promoting tighter feature clusters for similar conditions without relying on explicit alignment losses or external supervision. Empirically, CARE consistently improves both visual fidelity and convergence stability across both class-to-image and text-to-image tasks. On ImageNet, CARE achieves a 19.08\% reduction in FID in 400k training steps, leading to a 3.5$\times$ speed-up. When applied to text-to-image generation, CARE lowers FID by 16.61\% in 200k iterations and improves semantic alignment between generated samples and text prompts. Moreover, CARE can be seamlessly integrated with existing regularization methods, yielding additional performance gains.
Chinese Translation
扩散模型的最新研究进展凸显了表示正则化对于提升样本质量和训练效率的重要性。然而,常用的正则化方法往往忽视了直接决定生成目标的内置条件(如标签或文本)。在本工作中,我们展示了条件信号如何影响特征分布,并提出了CARE(条件感知表示正则化,Condition-Aware REpresentation regularization)。CARE是一个轻量级的即插即用正则化框架,能够基于条件相似度动态调制特征分布。CARE利用内置条件信号来审慎地引导表示空间,促进相似条件下的特征形成更紧密的聚类,而无需依赖显式的对齐损失或外部监督。实验表明,在类条件生成图像(class-to-image)和文本生成图像(text-to-image)两类任务上,CARE均能持续提升视觉保真度和收敛稳定性。在ImageNet上,CARE在40万训练步内将FID降低了19.08%,实现了3.5倍的加速。应用于文本生成图像任务时,CARE在20万次迭代内将FID降低了16.61%,并提升了生成样本与文本提示之间的语义对齐程度。此外,CARE可与现有正则化方法无缝集成,带来额外的性能提升。
cs.LG / 5 / 2609.28563

SpaFactor: Lightweight Spatial Context-Aware Gene Program Modeling for Histology-to-Transcriptomics Inference

SpaFactor:用于组织学-转录组学推断的轻量级空间上下文感知基因程序建模
Ruan, Shiting, Ling, Xitong, He, Qiming, Yan, Ziyou, Yuan, Huaitian, Guan, Tian, Xiao, Ying, Guan, Xu, He, Yonghong
Abstract
Spatial transcriptomics (ST) profiles gene expression within tissue architecture, but its cost and experimental complexity limit routine use. Predicting spatial expression from routinely available hematoxylin and eosin (HE) images therefore offers a scalable alternative. However, conventional methods often fit high-dimensional gene outputs as independent targets, overlooking the biological coordination among genes while remaining vulnerable to high-dimensional noise and overfitting. Existing attempts to address this limitation often rely on computationally heavy graph networks or complex auxiliary supervision. We therefore introduce SpaFactor, a lightweight and efficient low-rank morphology-program-gene factorization framework. At the input, SpaFactor efficiently fuses the visual representation of the central spot with multiscale local and regional neighborhood context, yielding a histologic representation that captures cellular morphology and microenvironmental heterogeneity. For modeling, a residual MLP stably learns a nonlinear mapping from the tissue microenvironment to low-dimensional latent gene programs. These activities are decoded through shared gene loadings into coordinated multi-gene expression predictions. Across five public cohorts, SpaFactor achieves the best aggregate performance, with particularly clear improvements for spatially variable genes, and more faithfully recovers biologically organized spatial patterns. These results demonstrate that lightweight joint modeling of tissue context and gene programs can improve both predictive accuracy and biological fidelity.
Chinese Translation
空间转录组学(Spatial Transcriptomics, ST)能够在组织结构内刻画基因表达,但其成本和实验复杂性限制了常规应用。因此,从常规可获得的苏木精-伊红(HE)染色图像预测空间基因表达提供了一种可扩展的替代方案。然而,传统方法通常将高维基因输出拟合为相互独立的目标,忽视了基因之间的生物学协同性,同时容易受到高维噪声和过拟合的影响。现有的改进尝试往往依赖于计算开销高昂的图网络或复杂的辅助监督。为此,我们提出了SpaFactor,一个轻量且高效的低秩形态学-程序-基因分解框架。在输入端,SpaFactor高效地将中心点位的视觉表征与多尺度的局部及区域邻域上下文相融合,得到能够刻画细胞形态和微环境异质性的组织学表征。在建模方面,残差多层感知机(MLP)稳定地学习从组织微环境到低维潜在基因程序的非线性映射,这些活性再通过共享的基因载荷解码为协同的多基因表达预测。在五个公开队列上,SpaFactor取得了最佳的总体性能,在空间可变基因上的提升尤为显著,并能更忠实地还原具有生物学组织性的空间模式。这些结果表明,对组织上下文与基因程序进行轻量级的联合建模,可以同时提升预测精度和生物学保真度。
cs.LG / 6 / 2609.28565

When Explanations Cannot Be Read: Measuring and Correcting SHAP and LIME Rendering for Right-to-Left Languages

当解释无法被阅读时:面向从右向左书写的语言的SHAP与LIME渲染的度量与修正
Zia, Rameesha, Malik, Muhammad Shahid Iqbal
Abstract
Post hoc explanation methods such as SHAP and LIME are widely used to interpret text classifiers, but their visualizations are mainly designed for left-to-right languages. When applied to right-to-left (RTL) languages such as Urdu, Arabic, Persian, and Hebrew, the attribution values remain mathematically valid, while their visual presentation fails. Tokens appear out of sequence, connected letterforms break apart, and plot layouts do not follow the natural reading direction. This study addresses this gap as a visualization problem rather than a limitation of the explanation methods themselves. We present SHAP-RTL, a rendering layer that corrects reading direction and script shaping in SHAP and LIME visualizations, with per-language font selection, while preserving the original attribution values, feature ordering, and model outputs. The approach is evaluated on Urdu, Arabic, Hebrew, and Persian hate and offensive-language datasets using TF-IDF and logistic regression classifiers. Rendering correctness is measured by an OCR round trip over 200 feature words per language. Default rendering yields character error rates of 0.820 to 0.979, meaning the label no longer carries its token; the common reshape-and-reorder workaround fails for Urdu at 0.998, worse than no correction; and the Matplotlib 3.11.0 text rewrite inverts that workaround, while SHAP-RTL remains correct under both versions. The framework also verbalizes the same attributions as short contextual explanations in the reader's language, constrained to the identified features. Evaluation in this paper concerns rendering correctness; assessment of the generated explanations is left to future work. The study highlights the importance of language-aware visualization in making post hoc explainability more accessible across different writing systems.
Chinese Translation
SHAP和LIME等事后解释方法被广泛用于解读文本分类器,但其可视化设计主要面向从左向右书写的语言。当应用于乌尔都语、阿拉伯语、波斯语和希伯来语等从右向左(RTL)书写的语言时,归因值在数学上依然有效,但其视觉呈现却会失效:词元顺序错乱、连写字形断裂、图形布局也不遵循自然阅读方向。本研究将这一问题视为可视化问题,而非解释方法本身的局限。我们提出了SHAP-RTL,这是一个渲染层,可在SHAP和LIME可视化中修正阅读方向和文字字形,并针对每种语言选择相应字体,同时保留原始的归因值、特征顺序和模型输出。该方法使用TF-IDF和逻辑回归分类器,在乌尔都语、阿拉伯语、希伯来语和波斯语的仇恨及攻击性语言数据集上进行评估。渲染正确性通过对每种语言200个特征词进行OCR往返测试来度量。默认渲染的字符错误率为0.820至0.979,意味着标签已不再承载其词元;常用的重排与重编码变通方法在乌尔都语上失效,错误率达0.998,比不修正还差;而Matplotlib 3.11.0的文本重写使该变通方法失效,SHAP-RTL则在两个版本下均保持正确。该框架还将相同的归因以读者所用语言生成为简短的上下文解释,且仅限于已识别的特征。本文的评估仅涉及渲染正确性,对所生成解释的评估留待后续工作。本研究凸显了语言感知可视化在使事后可解释性更好地适用于不同书写系统方面的重要性。
cs.LG / 7 / 2609.28567

Leakage-Safe Machine Learning for Hydrogen Embrittlement Detection in 316L Stainless Steel: A Region-Held-Out Evaluation of Texture and Deep Features in SEM Micrographs

面向316L不锈钢氢脆检测的防泄漏机器学习:基于区域留出的SEM显微图像纹理与深度特征评估
Awais, Muhammad, Yaseen, Muhammad, Shakoor, Abdul, Niaz, Niaz Ahmed, Zia, Huria, Shakoor, Muhammad Zain
Abstract
Scanning electron microscopy (SEM) is routinely used to characterize the microstructural changes caused by hydrogen embrittlement (HE) in structural steels. Machine learning can automate this characterization, but models are often evaluated using image-level splits. When several images come from the same specimen region, such splits leak information between the training and test sets. Here, we propose a region-held-out protocol for classifying as-received (AR) and hydrogen-charged (H2) SEM micrographs of 316L stainless steel, based on Leave-One-Region-Out (LORO) cross-validation over 14 spatial regions (8 AR, 6 H2; 31 images). We compared six feature-classifier combinations built on local binary patterns (LBP), grey-level co-occurrence matrices (GLCM), self-supervised convolutional embeddings pretrained on 143 unlabeled SEM images, and a convolutional neural network (CNN). The simplest texture approach, LBP with a support vector machine (LBP+SVM), performed best, achieving a balanced accuracy of 0.79, H2 recall of 0.69, and H2 precision of 0.82, outperforming every deep-learning and combined-feature model. A group-level permutation test (500 permutations sampled from the 3,003 possible region-to-label assignments) yielded p = 0.008, indicating that the result cannot be explained by a chance alignment of the region structure. Grad-CAM maps from a CNN trained on the full dataset tended to concentrate on localized surface and grain-boundary features, where hydrogen-induced morphological changes are known to occur. Under a leakage-safe, statistically validated protocol, texture descriptors recover a hydrogen-charging signature from SEM micrographs even with few samples, and the same protocol can be extended to larger HE detection studies in other alloy systems.
Chinese Translation
扫描电子显微镜(SEM)常被用于表征结构钢中氢脆(HE)引起的微观组织变化。机器学习可以实现这一表征的自动化,但模型通常采用图像级别的数据划分进行评估。当多张图像来自同一试样的同一区域时,这种划分会在训练集与测试集之间造成信息泄漏。本文提出了一种基于区域留出的分类方案,用于对316L不锈钢的原始态(AR)与充氢态(H2)SEM显微图像进行分类,该方法在14个空间区域(8个AR、6个H2;共31张图像)上采用留一区域法(Leave-One-Region-Out, LORO)交叉验证。我们比较了六种特征-分类器组合,包括基于局部二值模式(LBP)、灰度共生矩阵(GLCM)、在143张无标注SEM图像上预训练的自监督卷积嵌入特征,以及卷积神经网络(CNN)。最简单的纹理方法——LBP结合支持向量机(LBP+SVM)表现最佳,平衡准确率达0.79,H2召回率为0.69,H2精确率为0.82,优于所有深度学习及组合特征模型。组级别的置换检验(从3,003种可能的区域-标签对应中抽取500次置换)得到p = 0.008,表明该结果无法用区域结构的偶然对齐来解释。基于全数据集训练的CNN所生成的Grad-CAM热力图倾向于集中在局部表面和晶界特征上,而这些正是已知氢致形貌变化发生的部位。在防泄漏且经过统计验证的实验方案下,即使样本量很少,纹理描述符也能从SEM显微图像中提取出充氢特征;该方案还可扩展应用于其他合金体系中更大规模的氢脆检测研究。
cs.LG / 8 / 2609.28576

Time-Series Foundation Models That Understand Data Revisions

能够理解数据修订的时间序列基础模型
Ahmad, Taimoor
Abstract
Historical observations are not always fixed: statistical agencies revise previously published values as new evidence arrives. Forecasting from a contemporary download can therefore expose a model to information unavailable at the date it purportedly made a prediction. We propose VINTAGE-TS, a revision-aware adaptation of a time-series foundation model that distinguishes observation time from information-availability time. Its targets are the next period's first-published value and the value available a fixed number of days after that publication; neither is declared final truth. A joint predictive distribution preserves dependence between these targets and exposes uncertainty about their difference. We specify an ALFRED-based rolling evaluation, a matched Chronos-2 comparison, conventional and revision-aware baselines, and a separate audit of pretraining overlap. The accompanying software implements validity-interval reconstruction, delayed-label filtering, a frozen-backbone adapter interface, and reproducible diagnostics. An executed synthetic demonstration and a 25-configuration sensitivity suite verify the workflow, expose variation across seeds and revision regimes, and illustrate how hindsight contamination changes measured performance. Thirty one automated tests check temporal and integration contracts. Real ALFRED and Chronos-2 experiments have not been executed; no empirical foundation-model advantage is claimed.
Chinese Translation
历史观测值并非总是固定不变的:统计机构会随着新证据的到来对先前发布的数值进行修订。因此,基于当代下载数据进行预测,可能会使模型接触到其在假定预测日期时尚不可获得的信息。我们提出了VINTAGE-TS,这是一种具有修订感知能力的时间序列基础模型适配方法,它将观测时间与信息可得时间区分开来。其预测目标是下一时期的首次发布值,以及该发布后固定天数时可获得的数值;两者均不被视为最终的真实值。联合预测分布保留了这两个目标之间的依赖关系,并刻画了其差异的不确定性。我们设计了一种基于ALFRED的滚动评估、一个匹配的Chronos-2对比、常规基线与修订感知基线,以及对预训练数据重叠情况的独立审计。配套软件实现了有效期区间重构、延迟标签过滤、冻结骨干网络的适配器接口以及可复现的诊断工具。一项已执行的合成演示和包含25种配置的敏感性测试套件验证了该工作流程,揭示了不同随机种子和修订情景下的结果差异,并说明了后见之明污染(hindsight contamination)如何改变所测得的性能。31项自动化测试检查了时间逻辑与集成契约。真实的ALFRED和Chronos-2实验尚未执行,因此本研究不声称任何经验性基础模型优势。
cs.LG / 9 / 2609.28578

Uncovering Residential PV-EV Co-Adoption from Smart-Meter Data: Load Archetypes and Detection for Demand-Side Planning

基于智能电表数据揭示居民光伏与电动汽车协同采用行为:面向需求侧规划的负荷原型与检测方法
Zheng, Jack, Wang, Hao
Abstract
The increasing adoption of electric vehicles (EVs) and rooftop photovoltaic (PV) systems is reshaping residential electricity demand and creating new challenges for demand-side management (DSM), tariff design, and low-voltage network planning. Much of the existing literature examines EV charging or PV generation in isolation, leaving the behavioral dynamics of household co-adoption less understood. We develop an integrated, two-part workflow to analyze advanced metering infrastructure (AMI) data. A discovery component applies dynamic time warping (DTW) k-means with DTW barycenter averaging to cluster daily import or export profiles into interpretable behavioral archetypes, while a predictive component trains a bidirectional long short-term memory (BiLSTM) model on 21-day windows and benchmarks it against tabular baselines for PV/EV activity detection. The EV activity labels are inferred from charging-like load signatures because charger measurements are unavailable. Using half-hourly AusNet residential data from Victoria, Australia, the clustering uncovers distinct patterns across PV-only, EV-only, co-adoption, and neither cohorts; for co-adopters, a midday-centered weekday export archetype accounts for approximately 50% of days. At validation-tuned thresholds, both BiLSTM and XGBoost achieve strong discrimination. BiLSTM obtains 0.991 for the area under the receiver operating characteristic curve (AUROC), 0.906 for macro-F1, and the highest recall on the most difficult class (0.836 for EV-only recall). Tree-based baselines remain competitive. Performance remains stable across plausible labeling rules (macro-F1: 0.894--0.914) and strictly forward temporal splits (macro-F1: 0.894--0.906).
Chinese Translation
电动汽车(EV)和屋顶光伏(PV)系统日益普及的采用正在重塑居民用电需求,并为需求侧管理(DSM)、电价设计以及低压电网规划带来新的挑战。现有文献大多将电动汽车充电或光伏发电分开研究,对家庭协同采用的行为动态理解不足。我们开发了一个集成的两阶段工作流来分析高级量测体系(AMI)数据。发现环节采用带DTW重心平均的动态时间规整(DTW)k-means算法,将每日的用电/送电曲线聚类为可解释的行为原型;预测环节在21天窗口数据上训练双向长短期记忆网络(BiLSTM)模型,并与表格型基线方法进行对比,用于光伏/电动汽车活动检测。由于缺乏充电桩的量测数据,电动汽车活动标签是根据类似充电的负荷特征推断得到的。基于澳大利亚维多利亚州AusNet的半小时级居民数据,聚类结果揭示了纯光伏、纯电动汽车、协同采用及均无采用等不同群体间的明显模式差异;对于协同采用家庭,以中午为中心的工作日送电原型约占天数的50%。在经过验证集调优的阈值下,BiLSTM和XGBoost均表现出较强的判别能力。BiLSTM的受试者工作特征曲线下面积(AUROC)达到0.991,宏平均F1为0.906,且在最难类别上召回率最高(纯电动汽车召回率为0.836)。基于树的基线方法仍具竞争力。在不同的合理标注规则下(宏平均F1:0.894–0.914)以及严格的前向时间划分下(宏平均F1:0.894–0.906),模型性能均保持稳定。
cs.LG / 10 / 2609.28581

Auditability Is Not One Property: Rule Overlap, Behavioural Agreement, and Composition in Reinforcement Learning

可审计性并非单一属性:强化学习中的规则重叠、行为一致性与组合
Ming, Liu Hung
Abstract
Reinforcement-learning (RL) policies are often distributed as opaque neural checkpoints, while training logs show that a run occurred without explaining what the policy learned. We study whether independently trained policies can be represented and composed through auditable discrete behavioral rules. We define auditability as six separately testable predicates: trace integrity, lossless coding, rule coverage, behavioral agreement, composition quality, and value-model reliability. Our protocol uses a shared frozen symbolizer, passive rule extraction, an append-only hash-bound ledger, exact environment replay, and offline confidence-ranked arbitration with an explicit blind-spot fallback. The results place strict limits on this description layer. Rule-set overlap does not imply behavioral agreement: policies may share symbolic rules while choosing near-chance-matching actions on fresh states. The fused policy therefore selects among existing rules rather than generating a new skill. On a conflict-dominated task, an apparent fusion failure is traced to an induction/deployment mismatch: rules induced from sampled actions were evaluated under argmax actions, and deployment-consistent re-induction reverses the arbitration ordering. A fitted-Q generalized-policy-improvement diagnostic also fails in both environments, limiting claims that rule fusion is superior to value-based composition. One exploratory comparison favors rule fusion, but its comparator is post hoc, the task is partly saturated, and the fused policy remains below the strongest held-out actor. We contribute an evidence-bounded audit and composition protocol, not a claim of universal interpretability or autonomous skill generation. Future work must add temporally extended skills, cross-skill interfaces, composition search, and independent novelty audits.
Chinese Translation
强化学习(RL)策略通常以不透明的神经检查点形式发布,而训练日志只能表明某次运行发生过,却无法解释策略学到了什么。我们研究独立训练的策略能否通过可审计的离散行为规则来表示和组合。我们将可审计性定义为六个可分别检验的谓词:轨迹完整性、无损编码、规则覆盖、行为一致性、组合质量和价值模型可靠性。我们的协议采用共享的冻结符号化器(symbolizer)、被动规则提取、只追加的哈希绑定账本、精确的环境重放,以及带有显式盲区回退机制的离线置信度排序仲裁。结果对这一描述层施加了严格限制。规则集的重叠并不意味着行为一致性:多个策略可能共享符号规则,却在新的状态上选择近乎随机匹配的动作。因此,融合策略只是在现有规则中进行选择,而非生成新技能。在一个冲突主导的任务中,表面上的融合失败被追溯为归纳/部署不匹配:从采样动作归纳出的规则却在argmax动作下被评估,而与部署一致的重归纳会反转仲裁排序。拟合Q值(fitted-Q)广义策略改进诊断在两个环境中均告失败,这限制了“规则融合优于基于价值的组合”这类论断。一项探索性比较支持规则融合,但其对照是事后选取的、任务已部分饱和,且融合策略仍低于最强的保留行动者。我们贡献的是一个有证据边界的审计与组合协议,而非普遍可解释性或自主技能生成的论断。未来工作须加入时间上扩展的技能、跨技能接口、组合搜索以及独立的新颖性审计。
cs.LG / 11 / 2609.28582

SGA: Uncertainty Quantification for Multi-Step Forecasting in Time Series Foundation Models

SGA:时间序列基础模型中多步预测的不确定性量化
Hu, Xin-Yu, Liang, Shuang, Feng, Cheng, Zhang, Shao-Qun
Abstract
The recent emergence of Time Series Foundation Models (TSFMs) has significantly advanced multi-step forecasting performance, enabling accurate predictions over extended future horizons. However, existing TSFMs often suffer from significantly inherent uncertainty, which typically manifests as derived forecast branches emerging at each time step and spreading to subsequent steps; different forecast branches often exhibit varying forecasting performance, thereby undermining the credibility of TSFM forecasts. In this paper, we propose the Slicing-Graphing-Alignment (SGA) method to quantify the uncertainty of multi-step TSFM forecasts. The proposed SGA first characterizes the topology of all potential forecast branches using a directed acyclic graph, such that the graph complexity bounds the uncertainty of multi-step forecasts, and then precisely measures the graph complexity by integrating both topological information and TSFM-inherent stochasticity. Experimental results conducted on 11 TSFMs and 27 datasets demonstrate that (i) SGA achieves the best performance when ranking predictive errors with uncertainty estimates; (ii) SGA works with a more extensive and more precise sampling coverage than those of existing UQ methods, deriving a quantification mechanism fundamentally different from those of established ones; and (iii) larger model scales of TSFMs correlate with lower uncertainty estimates of multi-step forecasts, suggesting another empirical scaling law for uncertainty quantification of multi-step TSFM forecasts.
Chinese Translation
近年来,时间序列基础模型(TSFM)的出现显著提升了多步预测的性能,使其能够在较长的未来时间范围内进行准确预测。然而,现有的时间序列基础模型往往存在显著的固有不确定性,这种不确定性通常表现为在每个时间步产生不同的预测分支并向后续步骤扩散;不同的预测分支往往表现出不同的预测性能,从而削弱了TSFM预测结果的可信度。本文提出了切片-图构建-对齐(Slicing-Graphing-Alignment, SGA)方法,用于量化TSFM多步预测的不确定性。所提出的SGA方法首先使用有向无环图表征所有潜在预测分支的拓扑结构,使图的复杂度能够界定多步预测的不确定性,然后通过融合拓扑信息和TSFM固有的随机性来精确度量图的复杂度。在11个TSFM和27个数据集上进行的实验结果表明:(i) 在基于不确定性估计对预测误差进行排序时,SGA取得了最佳性能;(ii) 相比现有的不确定性量化(UQ)方法,SGA具有更广泛且更精确的采样覆盖率,其量化机制与已有方法存在根本性的差异;(iii) TSFM的模型规模越大,其多步预测的不确定性估计越低,这为TSFM多步预测的不确定性量化提出了又一经验缩放定律。
cs.LG / 12 / 2609.28590

TAM-Chain: Multi-Scale Thyroid Cytology Classification via Absorbing Markov Chains and Shannon Entropy Uncertainty Quantification for False-Negative Suppression and Domain-Shift Adaptation

TAM-Chain:基于吸收马尔可夫链与香农熵不确定性量化的多尺度甲状腺细胞学分类方法,用于假阴性抑制与领域偏移适应
Ngoc, Hai Pham
Abstract
Background & Problem: Thyroid Fine-Needle Aspiration Biopsy (FNAB) cytology based on the Bethesda System plays a pivotal role in early thyroid cancer detection; however, deep learning approaches face substantial challenges regarding high false-negative rates and overconfidence under clinical domain shift. Methods: In this study, we propose TAM-Chain, a multi-scale (10x, 20x, 40x) thyroid cytology classification framework leveraging Absorbing Markov Chain theory combined with Shannon Entropy-based Uncertainty Quantification. The framework dynamically models multi-magnification feature extraction as an absorbing stochastic process, enabling optimal stopping criteria and a human-in-the-loop referral mechanism to strictly suppress critical diagnostic errors. Results: Extensive evaluation on an internal test set (N = 235) demonstrates a Macro F1 score of 0.9741 with an absolute False-Negative Rate (FNR) of 0.00%. On an independent external validation set (N = 1015) presenting severe domain shift, TAM-Chain maintains superior stability and classification performance (Macro F1 = 0.7026) by adaptively adjusting the expected stopping step and triggering specialist referrals, significantly outperforming single-magnification baselines. Conclusion: The TAM-Chain framework proves to be a highly effective, safe, and adaptable solution for digital pathology workflows, successfully harmonizing automated diagnostic efficiency with stringent biological safety.
Chinese Translation
背景与问题:基于Bethesda系统的甲状腺细针穿刺活检(FNAB)细胞学检查在甲状腺癌早期检测中发挥着关键作用;然而,深度学习方法面临假阴性率高以及在临床领域偏移下过度自信的重大挑战。方法:本研究提出TAM-Chain,一个利用吸收马尔可夫链(Absorbing Markov Chain)理论并结合基于香农熵的不确定性量化的多尺度(10x、20x、40x)甲状腺细胞学分类框架。该框架将多倍率特征提取动态建模为吸收型随机过程,实现最优停止准则与人机协同转诊机制,以严格抑制关键诊断错误。结果:在内部测试集(N = 235)上的广泛评估显示,Macro F1分数达0.9741,绝对假阴性率(FNR)为0.00%。在呈现严重领域偏移的独立外部验证集(N = 1015)上,TAM-Chain通过自适应调整期望停止步数并触发专家转诊,保持了卓越的稳定性与分类性能(Macro F1 = 0.7026),显著优于单倍率基线方法。结论:TAM-Chain框架被证明是数字病理学工作流程中高效、安全且适应性强的解决方案,成功地将自动化诊断效率与严格的生物安全性相协调。
cs.LG / 13 / 2609.28603

Learning to Discover Interesting Mathematics

学习发现有趣的数学
Patel, Niket, Rammal, Ahmad, Hayat, Amaury, Munos, Remi, Kempe, Julia
Abstract
Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and more theorems, it remains open whether this new mathematical knowledge is interesting or useful. We define intrinsic interestingness of a theorem as the ratio between the length of its proof and the length of its statement. We show that this correlates strongly with an extrinsic measure of the downstream utility of a theorem. We identify the difficulty of a proof conditioned on a set of premises as a useful primitive for computing these metrics, and train a 27B model that predicts proof difficulty more accurately than frontier general-purpose models. Optimizing for our metric creates a model capable of producing more interesting theorems, while also reducing substantial or full overlap with Mathlib from 91.9% to 30.6%, showcasing the creation of more out-of-distribution math. We show that our system can generate candidate theorems, select the most interesting among them, and iteratively build on a self-expanding mathematical library. These metrics provide a practical and quantifiable signal for ranking conjectures and guiding proof search within formal mathematical libraries. Our framework provides a path towards self-expanding, machine-verified mathematical libraries that can choose worthwhile statements without relying on human-supplied targets.
Chinese Translation
近来,大语言模型(Large Language Models,LLMs)解决高难度数学问题的能力日益增强,其中包括许多悬置数十年的公开问题。这为以前所未有的规模扩展数学知识打开了大门。然而,尽管 LLMs 能够猜想并证明越来越多的定理,这些新的数学知识是否有趣或有用仍是一个开放问题。我们将定理的内在趣味性定义为:其证明长度与陈述长度之比。我们证明该指标与定理下游效用的外在度量高度相关。我们指出,在给定前提集条件下证明的难度是计算这些度量的有效基础工具,并训练了一个 27B 模型,其预测证明难度的准确率超过了前沿的通用模型。针对我们的度量进行优化,得到的模型能够产出更有趣的定理,同时将与 Mathlib 存在大量或完全重合的比例从 91.9% 降低至 30.6%,表明生成了更多分布外的数学内容。我们的系统能够生成候选定理、从中挑选最有趣者,并在自我扩展的数学库上进行迭代式构建。这些度量为形式化数学库中的猜想排序与证明搜索引导提供了实用且可量化的信号。我们的框架为构建自我扩展、机器可验证的数学库开辟了一条路径,使其无需依赖人工设定的目标即可自主选择值得证明的命题。
cs.LG / 14 / 2609.28604

Physics-Informed Self-Supervised Learning for Joint Wire Calibration and Interaction Position Reconstruction in Multi-Wire Parallel Plate Avalanche Counters

基于物理信息的自监督学习用于多丝平行板雪崩计数器中的丝校准与相互作用位置重建的联合优化
Lemasson, Antoine, Rejmund, Maurycy
Abstract
Scientific instruments require accurate calibration to convert detector signals into reliable physical observables. Conventional calibration procedures typically rely on dedicated calibration measurements, analytical response models or labelled reference data, limiting their ability to adapt to changing operating conditions and detector aging. We present a physics-informed self-supervised learning framework that jointly performs wire calibration and interaction position reconstruction in Multi-Wire Parallel Plate Avalanche Counters (MWPPACs) without requiring labelled position measurements or dedicated calibration runs. The method formulates detector calibration as a latent optimization problem in which global wire gains and event-wise interaction positions are estimated simultaneously using supervision derived exclusively from detector geometry and charge-energy consistency constraints. A detector-independent neural network reconstructs sub-wire interaction positions from local charge distributions, eliminating the need to assume analytical induction profiles by learning the detector response directly from experimental data. The end-to-end differentiable framework enables continuous detector self-calibration while improving the uniformity and accuracy of position reconstruction. Experimental evaluation on the entrance MWPPAC tracking detectors of the VAMOS++ magnetic spectrometer demonstrates stable convergence, improved spatial homogeneity and enhanced position resolution. Beyond the detector studied, the method establishes a general framework for physics-informed self-supervised calibration of scientific instruments and is a step toward autonomous intelligent instrumentation capable of continuous adaptation during operation. In this paradigm, detector calibration is no longer a prerequisite for an experiment but an integral part of the measurement process itself.
Chinese Translation
科学仪器需要精确的校准,以便将探测器信号转换为可靠的物理观测量。传统校准流程通常依赖于专门的校准测量、解析响应模型或带标签的参考数据,这限制了其适应运行条件变化和探测器老化的能力。我们提出了一种基于物理信息的自监督学习框架,能够在多丝平行板雪崩计数器(MWPPAC)中联合实现丝校准与相互作用位置重建,而无需带标签的位置测量或专门的校准运行。该方法将探测器校准构建为一个隐变量优化问题,仅利用源自探测器几何结构和电荷-能量一致性约束的监督信号,同时估计全局丝增益和逐事件的相互作用位置。一个与探测器无关的神经网络从局部电荷分布中重建亚丝级的相互作用位置,通过直接从实验数据中学习探测器响应,从而无需假设解析的感应分布。该端到端可微框架使探测器能够持续自校准,同时提升位置重建的均匀性和精度。在VAMOS++磁谱仪入射MWPPAC径迹探测器上的实验评估表明,该方法收敛稳定,空间均匀性和位置分辨均得到改善。除所研究的探测器外,该方法还为科学仪器的基于物理信息的自监督校准建立了一个通用框架,是迈向能够在运行过程中持续自适应的自主智能仪器的一步。在这一范式下,探测器校准不再是实验的前提条件,而是测量过程本身的一个组成部分。
cs.LG / 15 / 2609.28605

UO-FIE: Combining Exact-Label Supervision with Graded Utility for Factivity Inference

UO-FIE:结合精确标签监督与分级效用的传实性推断方法
Xiao, Xinchen
Abstract
The Factivity Inference Evaluation 2026 (FIE2026) classifies Chinese context-hypothesis pairs into nine ordered factivity intervals. Its evaluation metric rewards both exact predictions and proximity to the correct interval, while 64.1% of the 566 training examples belong to a single class. In preliminary experiments, several mDeBERTa classification models predominantly predict the dominant class, whereas a Huber-regression baseline produces more predictions near the correct interval but fewer exact matches. We introduce Utility-Oriented Factivity Inference (UO-FIE), a parameter-efficient system that combines exact-label supervision with graded utility. UO-FIE predicts a distribution over the nine classes and combines hard-label supervision, utility-based soft targets, scheduled class weights, and an ordinal loss. We evaluate expected-utility decoding in controlled comparisons and use ordinal calibration selected on out-of-fold predictions for the submitted system. Based on Qwen3.5-9B with LoRA, UO-FIE ranks first in the fine-tuning track with a macro utility of 0.8316. A separate prompt-based ensemble ranks third in the non-fine-tuning track with a macro utility of 0.8450.
Chinese Translation
传实性推断评测2026(Factivity Inference Evaluation 2026,FIE2026)将中文语境-假设对划分为九个有序的传实性区间。其评测指标同时奖励精确预测和对正确区间的接近程度,而566个训练样本中有64.1%属于单一类别。在初步实验中,多个mDeBERTa分类模型主要预测占主导地位的类别,而Huber回归基线模型则产生更多接近正确区间的预测,但精确匹配较少。我们提出了面向效用的传实性推断系统(Utility-Oriented Factivity Inference,UO-FIE),这是一个结合精确标签监督与分级效用的参数高效系统。UO-FIE预测九个类别上的分布,并结合了硬标签监督、基于效用的软目标、计划性类别权重以及序数损失。我们在受控比较中评估了期望效用解码,并针对提交系统使用基于折外预测选择的序数校准。基于Qwen3.5-9B与LoRA,UO-FIE在微调赛道中以0.8316的宏观效用排名第一。另一个基于提示的集成系统在非微调赛道中以0.8450的宏观效用排名第三。
cs.LG / 16 / 2609.28607

fable.intermittent: benchmarking probabilistic forecasting methods for intermittent time series

fable.intermittent:间歇性时间序列概率预测方法的基准测试
Damato, Stefano, Zambon, Lorenzo, Corani, Giorgio, Azzimonti, Dario
Abstract
Intermittent time series are common in spare-parts demand and retail sales. Since the cost of forecast errors is typically asymmetric, decisions such as inventory control require the full predictive distribution rather than a point forecast. Many probabilistic forecasting methods have been proposed; their implementations, however, are scattered across different software frameworks, making it difficult to compare them systematically. We introduce fable.intermittent, an R package that implements several probabilistic forecasting methods for intermittent series within the fable framework. The package allows several models to be fitted and evaluated on a collection of time series through a single, simple forecasting pipeline. We also introduce TWEES, a new exponential smoothing model with a Tweedie predictive distribution. Fitting TWEES requires repeated evaluation of the computationally demanding Tweedie density. We also release the R package tweedieDistr, whose implementation of the Tweedie distribution is substantially faster than the existing one while preserving the same numerical accuracy. We evaluate the methods implemented in fable.intermittent on four datasets, also released in the package.
Chinese Translation
间歇性时间序列在备件需求和零售销售中十分常见。由于预测误差的代价通常是不对称的,诸如库存控制等决策需要完整的预测分布而非点预测。目前已提出了许多概率预测方法,但它们的实现分散在不同的软件框架中,难以进行系统性比较。我们介绍了 fable.intermittent,一个在 fable 框架内实现了多种间歇性序列概率预测方法的 R 包。该包通过一个简单统一的预测流程,即可对一组时间序列拟合和评估多个模型。我们还提出了 TWEES,一种具有 Tweedie 预测分布的新的指数平滑模型。拟合 TWEES 需要反复计算计算开销较大的 Tweedie 密度。为此,我们还发布了 R 包 tweedieDistr,其 Tweedie 分布的实现比现有实现快得多,同时保持了相同的数值精度。我们在四个数据集上评估了 fable.intermittent 中实现的方法,这些数据集也随包一并发布。
cs.LG / 17 / 2609.28625

RLVR landscapes for iterated multiplications can be benign: Insights from spin-glass theory

迭代乘法任务的RLVR优化景观可能是良性的:来自自旋玻璃理论的洞见
Rubin, Noa, Ringel, Zohar
Abstract
Despite the importance of reinforcement learning with verifiable rewards (RLVR), the extent to which it can learn new reasoning capabilities remains debated. Here we study the optimization landscape of RLVR on algorithmic tasks, such as iterated group and quasigroup multiplication. To this end, we map entropy-regularized RLVR over myopic tabular policies onto an energy-based (spin-glass) model over deterministic policies. This mapping upper-bounds what RLVR can achieve, and lets us rigorously characterize the landscape in this tabular setting. We show, both theoretically and experimentally, that for a wide class of models and tasks with uncorrelated inputs, this landscape is benign, containing no local minima that could trap RLVR training. Rather, the practical difficulty of these tasks appears to stem, at least in part, from issues such as diffusive barriers and gradient-estimation error in traversing the landscape. These are genuine obstacles that can prevent a solution from being found, but they are distinct from the landscape itself being rugged. We show that these obstacles can often be mitigated through the choice of entropy regulator. Consistent with this theory, we find that a transformer trained from scratch, using only last-token rewards, successfully learns an algorithmic chain of thought for iterated non-Abelian group multiplications.
Chinese Translation
尽管可验证奖励强化学习(RLVR)十分重要,但它能否学习到新的推理能力仍存在争议。本文研究了RLVR在算法任务(如迭代群与拟群乘法)上的优化景观。为此,我们将基于短视表格策略的熵正则化RLVR映射为确定性策略上的能量模型(自旋玻璃模型)。该映射给出了RLVR所能达到性能的上界,使我们能够严格刻画表格设定下的景观结构。我们从理论和实验两方面证明,对于一大类输入不相关的模型与任务,该景观是良性的,不存在可能困住RLVR训练的局部极小值。相反,这些任务的实际难度似乎至少部分源于扩散势垒、以及在景观中搜索时的梯度估计误差等问题。这些是真实存在的、可能阻碍找到解的障碍,但它们与景观本身的崎岖性不同。我们证明这些障碍通常可以通过选择合适的熵调节器加以缓解。与该理论一致,我们发现一个仅使用末token奖励、从零开始训练的transformer,成功学会了迭代非阿贝尔群乘法的算法化思维链。
cs.LG / 18 / 2609.28665

OPDiv: Optimal Selection of Top-K High-Scoring, Diverse Compounds

OPDiv:Top-K高分且多样性化合物的最优选择
Lžičař, Miroslav
Abstract
A virtual screening campaign may produce thousands of promising candidates, but only a small number can be purchased, synthesized, or tested. The practical question is how to select a set of compounds that both rank well and are diverse enough: this poses a genuine tradeoff, where selecting the highest-scoring molecules yields limited diversity, while diversity selection sacrifices some well-scoring molecules. We introduce OPDiv, a diversity selection and evaluation algorithm solving this tradeoff by finding an optimal subset of molecules using integer optimization. We demonstrate the selection algorithm in practice with fingerprint distance, shape and electrostatic diversity and compare the resulting diversity spectra. We argue that virtual screening is not merely a ranking problem, but also an implicit constrained optimization task: when redundant chemotypes are undesirable, pipelines should be compared based on the top-k compound selections satisfying the desired diversity constraints. OPDiv makes it possible to find the optimal compound set under a given diversity threshold efficiently and serves as a fair benchmark of the best diverse selection achievable by a given structure-based or ligand-based virtual screening pipeline, molecular search or generative model.
Chinese Translation
一次虚拟筛选活动可能产生数千个有前景的候选化合物,但只有少数化合物能够被购买、合成或测试。实际问题在于如何选择一组既排名靠前又具有足够多样性的化合物:这构成了一个真正的权衡,因为选择得分最高的分子会导致多样性有限,而多样性选择则会牺牲一些高得分的分子。我们提出了OPDiv,一种多样性选择与评估算法,通过整数优化找到一个最优的分子子集来解决这一权衡问题。我们在实践中利用指纹距离、形状和静电多样性对该选择算法进行了演示,并比较了所得到的多样性谱。我们认为,虚拟筛选不仅是一个排序问题,同时也是一个隐式的约束优化任务:当冗余的化学型不被需要时,应当基于满足所需多样性约束的top-k化合物选择来比较不同流程。OPDiv能够在给定多样性阈值下高效地找到最优化合物集合,并可作为衡量基于结构或基于配体的虚拟筛选流程、分子搜索或生成模型所能实现的最佳多样性选择的公平基准。
cs.LG / 19 / 2609.28670

Beyond Static Graph World Models: Learning Stochastic Latent Dynamics over Evolving Topologies

超越静态图世界模型:在演化拓扑上学习随机潜在动力学
Schutz, Alex, Hawes, Nick, Darvariu, Victor-Alexandru
Abstract
Graph-based world models have recently emerged as a means of learning transitions over relational state representations. However, existing approaches are largely limited to fixed-topology graphs or deterministic, fully observable environments. We propose the Graph Dynamics Model (GDM), a world model for graph-structured observations that is designed to handle the more general setting of evolving topologies in stochastic and partially observable environments. The GDM uses a sparse recurrent adjacency matrix to model topology updates and perform message passing, together with a recurrent state-space architecture for modelling stochastic transitions. Furthermore, we identify a gap in the evaluation of graph-based world models, as existing methods do not provide a means of comparing predicted and true distributions over the joint graph state comprising the interdependent topology, node features, and graph features. We therefore introduce the Graph Distribution Distance (GDD) metric, which uses maximum mean discrepancy with a graph kernel to comprehensively compare joint next-state distributions. We evaluate the GDM across several environments, including stochastic and partially observable settings. We demonstrate that GDM outperforms baseline models and displays zero-shot generalisation on large graphs.
Chinese Translation
基于图的世界模型近来成为在关系型状态表示上学习状态转移的一种手段。然而,现有方法大多局限于固定拓扑的图或确定性、完全可观测的环境。我们提出了图动力学模型(Graph Dynamics Model, GDM),这是一种面向图结构观测的世界模型,旨在应对随机和部分可观测环境中拓扑演化的更一般情形。GDM 使用稀疏循环邻接矩阵来建模拓扑更新并执行消息传递,同时结合循环状态空间架构来建模随机转移。此外,我们发现基于图的世界模型在评估方面存在空白:现有方法无法比较由相互依赖的拓扑、节点特征和图特征构成的联合图状态上的预测分布与真实分布。为此,我们提出了图分布距离(Graph Distribution Distance, GDD)度量,该度量使用带图核的最大均值差异(MMD)来全面比较联合下一状态分布。我们在多个环境中评估了 GDM,包括随机和部分可观测的设置。实验表明,GDM 的性能优于基线模型,并在大图上展现出零样本泛化能力。
cs.LG / 20 / 2609.28682

Thinking Leakage: A Causal Audit of NoThink Post-Training in Hybrid Reasoning Models

思维泄漏:混合推理模型中NoThink后训练的因果审计
Liu, Zehao, Honavar, Vasant G.
Abstract
Post-training hybrid reasoning models in NoThink mode has attracted growing interest as a way to improve performance while keeping inference fast. However, these gains may draw on thinking behavior already accessible through the base model's Think mode. We formulate this thinking leakage in a causal mediation framework and audit its contribution using bidirectional interventions along a simple base-derived activation direction. Across three models and three post-training methods on competition math benchmarks, we find that leakage is real, causal, and substantial: behavioral and representational analyses reveal shifts toward Think, steering the base model along this direction reproduces most of the post-training accuracy gain, and counter-steering a checkpoint removes a substantial share of what it gains. Across nine aligned checkpoints with positive NoThink gains, the resulting leakage ratio ranges from 42% to 79%. These interventions support a substantial causal contribution of thinking leakage. Our findings show that a post-training method's apparent advantage can therefore reflect greater drift toward Think, obscuring whether it improves capability within NoThink or more effectively re-invokes existing Think behavior.
Chinese Translation
以NoThink模式对混合推理模型进行后训练作为一种在保持推理速度的同时提升性能的方法,正受到越来越多的关注。然而,这些性能提升可能依赖于基础模型在Think模式下本已具备的思维行为。我们在因果中介(causal mediation)框架中对这种思维泄漏进行形式化,并沿一个简单的由基础模型导出的激活方向,利用双向干预审计其贡献。在竞赛数学基准上,跨三个模型和三种后训练方法的实验表明,泄漏是真实、具有因果性且相当可观的:行为与表征分析显示模型向Think模式偏移;沿该方向对基础模型进行引导(steering)可复现后训练带来的大部分准确率提升;而对某个检查点实施反向引导(counter-steering)则会去除其获得收益中的相当大一部分。在九个具有正向NoThink增益的对齐检查点上,所得的泄漏比率介于42%至79%之间。这些干预结果支持思维泄漏具有显著的因果贡献。我们的发现表明,后训练方法表观上的优势可能实际上反映的是向Think模式的更大偏移,从而掩盖了该方法究竟是在NoThink模式内提升了能力,还是仅仅更有效地重新调用了已有的Think行为。
cs.LG / 21 / 2609.28695

Federated Learning of AnDE Classifiers

AnDE分类器的联邦学习
Torrijos, Pablo, Alfaro, Juan C., Gámez, José A., Puerta, José M.
Abstract
This work presents a federated framework for training Averaged $n$-Dependence Estimators (AnDE) in distributed environments. The proposed method focuses on the discriminative setting, where model weights are learned locally and aggregated globally, supporting any dependency order $n$. This design allows federated training without transmitting semantically meaningful parameters, improving privacy. Additionally, generative AnDE models are federated to provide a comparative baseline, with optional differential privacy applied to the aggregation of probability tables. Experiments on 12 discrete datasets show that discriminative models with $n \geq 1$ consistently outperform federated Naive Bayes (NB, $n=0$), and that privacy-preserving aggregation is effective with limited accuracy loss. These results establish federated AnDE as a viable and privacy-preserving framework, showing that probabilistic models remain applicable in modern federated learning settings.
Chinese Translation
本工作提出了一个在分布式环境中训练平均n依赖估计器的联邦框架。所提出的方法聚焦于判别式设置,即模型权重在本地学习并在全局聚合,支持任意依赖阶数n。这一设计使得联邦训练无需传输具有语义含义的参数,从而提升了隐私性。此外,还将生成式AnDE模型进行联邦化以提供对比基线,并可对概率表的聚合选择性地应用差分隐私。在12个离散数据集上的实验表明,n≥1的判别式模型始终优于联邦朴素贝叶斯(NB,n=0),且隐私保护聚合在精度损失有限的情况下是有效的。这些结果确立了联邦AnDE作为一种可行且保护隐私的框架,表明概率模型在现代联邦学习场景中仍然适用。
cs.LG / 22 / 2609.28697

LabFactory: Building and Evaluating Executable AI Labs

LabFactory:构建与评估可执行的AI实验室
Wu, Jinge, Zhou, Hongjian, Zeng, Mingde, Zhu, Jiayuan, Wu, Junde, Pan, Jiazhen, Clifton, Lei, Liu, Andrew, Clifton, David A.
Abstract
Scientific tasks specify a desired capability, but realizing it often requires building a computational system tailored to the task---acquiring data, designing representations, training models, implementing tools, and deciding how they are used at inference. We present LabFactory, a framework in which an AI builder turns a scientific brief into an executable AI lab: a task-specific solver that integrates models, knowledge resources, tools, and a controller behind a fixed interface. The builder develops and packages the lab in a metered workspace; a separate host then executes the delivered artifact on held-out inputs, with reference labels kept outside the solver's input interface, and scores its outputs under the task's protocol. This makes the delivered system, rather than the builder's account of its progress, the object of evaluation. We document 28 selected constructions across seven scientific task categories---from molecular and genomic prediction to physiological signals, clinical decision support, and biomedical text---whose delivered labs exceeded their configured reference values on all 33 subtests under host-side execution. Ten contain predictive models fitted during construction; the others assemble retrieval systems, executable analysis environments, and tool-driven workflows around a fixed platform LLM. Together they show that an AI agent can carry a scientific brief all the way to a working lab that can still be invoked, inspected, and checked after construction ends.
Chinese Translation
科学任务规定了所需的能力,但要实现它,往往需要构建一个针对该任务定制的计算系统——获取数据、设计表示、训练模型、实现工具,并决定在推理时如何使用它们。我们提出了LabFactory,这是一个框架,其中AI构建者将科学简报转化为一个可执行的AI实验室:一个任务特定的求解器,在固定接口之后整合了模型、知识资源、工具和控制器。构建者在一个受限计量的工作空间中开发并打包该实验室;随后由一个独立的主机(host)在保留的测试输入上执行所交付的产物,参考标签保持在求解器输入接口之外,并按照任务的协议对其输出进行评分。这使得被评估的对象成为所交付的系统本身,而非构建者对其进展的陈述。我们记录了横跨七个科学任务类别的28个精选构建案例——从分子与基因组预测到生理信号、临床决策支持和生物医学文本——这些交付的实验室在主机端执行下,在全部33个子测试中均超过了预设的参考值。其中十个包含在构建过程中拟合的预测模型;其余的则围绕一个固定的平台级大语言模型(LLM),组装了检索系统、可执行分析环境和工具驱动的工作流。这些案例共同表明,AI智能体能够将一份科学简报一路推进到一个可运行的实验室,并且在构建结束后,该实验室仍可被调用、检查和验证。
cs.LG / 23 / 2609.28722

Upholding Robustness in Federated Learning: Trends, Emerging Strategies, and Research Opportunities

维护联邦学习的鲁棒性:趋势、新兴策略与研究机遇
P V, Pravija Raj, Gupta, Ashish, Augello, Andrea, Das, Sajal K.
Abstract
While Federated Learning (FL) has been widely adopted for protecting user privacy in machine learning, it remains vulnerable to various robustness challenges, including performance-impairment risks, information-stealing threats, and aggregation vulnerabilities. This work offers a holistic synthesis of FL robustness along three tightly coupled angles: (i) a threat-centric view of robustness that categorizes the multifaceted attack surfaces, (ii) a structured taxonomy of robust aggregation strategies distinguishing outcome-centric approaches from security-centric strategies, and (iii) a layered taxonomy of defensive strategies. We rigorously examine current evaluation practices for FL robustness and identify major applications and open research challenges to guide future research.
Chinese Translation
尽管联邦学习(Federated Learning, FL)已被广泛用于保护机器学习中的用户隐私,但其仍然容易受到各种鲁棒性挑战的影响,包括性能受损风险、信息窃取威胁以及聚合脆弱性。本工作从三个紧密关联的角度对联邦学习的鲁棒性进行了全面的综合梳理:(i) 以威胁为中心的鲁棒性视角,对多方面的攻击面进行分类;(ii) 对鲁棒聚合策略进行结构化分类,将侧重结果的策略与侧重安全的策略加以区分;(iii) 对防御策略进行分层分类。我们严格审视了当前联邦学习鲁棒性的评估实践,并识别了主要应用领域和开放性研究挑战,以指导未来的研究。
cs.LG / 24 / 2609.28737

Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning

强化学习中的策略复杂度、反应时间与有限理性
Wu, James, Sims, Chris R.
Abstract
Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavior and therefore bound policy complexity. Standard reinforcement learning models typically optimize reward without explicitly representing these internal costs, making them less suitable as models of biological intelligence. We derive MI-SARSA, an on-policy temporal-difference algorithm that incorporates mutual-information regularization through a learned marginal action prior and a penalty on state-specific deviations from that prior. This yields a sequential learning model in which state information is used selectively when its expected return benefit justifies the added informational cost. Critically, the same state-specific information cost that governs policy compression also generates trial-level predictions for reaction time, distinguishing MI-SARSA from most reinforcement learning models, which predict choices or returns but not latency. Empirically, MI-SARSA produces a reward-complexity tradeoff, and stronger information penalties produce simpler policies with lower control costs and faster reaction times. Under environment shift, increasing regularization reduces post-switch performance degradation but also lowers asymptotic return, revealing a robustness-capacity tradeoff. Together, these results position MI-SARSA as a model of bounded sequential learning under cognitive constraints.
Chinese Translation
生物体并非在无限计算资源的条件下进行学习。对人类而言,学习与选择受到感知、注意和工作记忆等约束的塑造,这些约束限制了用于指导行为的状态信息量,从而限定了策略复杂度。标准的强化学习模型通常只优化奖励而不显式表征这些内部代价,因此作为生物智能的模型并不十分合适。我们推导出 MI-SARSA,这是一种在策略的时序差分算法,它通过学习得到的边缘动作先验以及针对状态对该先验偏离的惩罚,引入了互信息正则化。由此得到一个序贯学习模型,其中状态信息仅在预期回报收益足以抵偿额外信息代价时才被选择性使用。关键在于,控制策略压缩的状态特异性信息代价同时也能产生试次层面的反应时间预测,这使 MI-SARSA 区别于大多数只能预测选择或回报而无法预测反应延迟的强化学习模型。实验结果表明,MI-SARSA 呈现出奖励-复杂度权衡:更强的信息惩罚产生更简单的策略、更低的控制成本和更快的反应时间。在环境转换情境下,增强正则化可减少转换后的性能下降,但同时降低了渐近回报,揭示出鲁棒性与容量之间的权衡。综上,这些结果将 MI-SARSA 定位为一个在认知约束下进行有限序贯学习的模型。
cs.LG / 25 / 2609.28749

Evaluating Cross-region Generalization for Wavelet-Diffusion Precipitation Downscaling

评估小波扩散降水降尺度方法的跨区域泛化能力
Qian, Weikang, Wen, Yixin, Yi, Chugang, Li, Zhi, Li, Lingcheng, Yang, Haizhao
Abstract
Diffusion models have shown strong potential for kilometer-scale precipitation downscaling, but their performance in geographically unseen regions and event regimes remains insufficiently understood. Building on the wavelet diffusion model (WDM) framework, this study evaluates cross-region and cross-event generalization. Six 3 x 3 deg U.S. regions represent convective, winter, tropical, and atmospheric-river precipitation regimes. Low-resolution inputs are generated by block averaging NOAA Multi-Radar/Multi-Sensor (MRMS) composite reflectivity fields. A WDM trained only on Oklahoma (OK) samples and a WDM trained on all six regions are compared with nearest-neighbor and Bicubic interpolation. Model performance is evaluated using three metric families that measure image-domain reconstruction, spectral and distributional fidelity, and bin-wise precipitation detection. The OK-trained WDM remains competitive outside OK. Although the all-region WDM delivers the best and most consistent overall image-domain and detection performance, its gains are uneven across precipitation intensities. Bin-wise critical success index (CSI) over 5-dBZ reflectivity bins shows that WDM improvements concentrate in localized higher-reflectivity structures, which image-domain metrics partly obscure. In addition, the performance differences among samples are strongly associated with the spatial organization of the precipitation field, quantified by Moran's I as the spatial autocorrelation of each reflectivity bin. The sample-level Moran's I-CSI correlation stratified by sample intensity reaches 0.901 in all six regions, including regions unseen during training. Overall, these findings support future efforts to transfer downscaling models to regions with limited local training data and to generate globally consistent, high-resolution precipitation products.
Chinese Translation
扩散模型在公里级降水降尺度方面展现出巨大潜力,但其在地理上未见区域和未见降水事件形态下的性能仍缺乏充分认识。基于小波扩散模型(Wavelet Diffusion Model, WDM)框架,本研究评估了跨区域与跨事件类型的泛化能力。研究选取美国六个3×3度区域,分别代表对流性、冬季、热带以及大气河流降水形态。低分辨率输入通过对NOAA多雷达/多传感器(MRMS)组合反射率场进行分块平均生成。仅用俄克拉荷马州(OK)样本训练的WDM与在全部六个区域训练的WDM,与最近邻插值和双三次插值方法进行了比较。模型性能采用三类指标体系评估:图像域重建、频谱与分布保真度,以及分箱降水检测。结果表明,仅在OK训练的WDM在OK以外区域仍具竞争力。尽管全区域WDM在图像域和检测方面提供了最佳且最一致的整体性能,但其在不同降水强度上的增益并不均衡。按5-dBZ反射率分箱计算的逐箱临界成功指数(CSI)显示,WDM的改进集中于局地高反射率结构,而这些改进在图像域指标中部分被掩盖。此外,样本间的性能差异与降水场的空间组织结构密切相关,后者通过各反射率分箱的空间自相关(以Moran's I量化)来度量。按样本强度分层的样本级Moran's I-CSI相关系数在全部六个区域均达到0.901,其中包括训练期间未见过的区域。总体而言,这些发现为未来将降尺度模型迁移至缺乏本地训练数据的区域,以及生成全球一致的高分辨率降水产品提供了支持。
cs.LG / 26 / 2609.28782

The Mechanics of Delta Learning: Target Design for Generalizable Scientific Machine Learning

Delta学习的机制:面向可泛化科学机器学习的目标设计
Gameel, Kareem M., Neporozhnii, Ihor, Hoogland, Sjoerd, Voznyy, Oleksandr
Abstract
In scientific machine learning, $\Delta$-learning trains models on residual errors relative to physical baselines, assuming that more accurate baselines with smaller residual scales inherently improve downstream performance. Here, we demonstrate that residual scale alone is an insufficient heuristic for learnability. Evaluating molecular graph neural networks on total energy targets, we show that complex local descriptor baselines can yield small residual targets that are disproportionately rough within architecture-informed proxy spaces and harder to learn relative to their scale. Conversely, semi-empirical baseline reduces both scale and normalized roughness, improving in-domain and out-of-domain prediction. We introduce scale-normalized graph Dirichlet roughness ($D_{\text{IQR}}$) as a pre-training diagnostic for residual learnability and establish baseline complementarity as a core target-design principle, elevating target space formulation alongside model architecture as a key axis for scientific machine learning.
Chinese Translation
在科学机器学习中,Δ-learning(Delta学习)通过训练模型学习相对于物理基线的残差误差来提升性能,其隐含假设是:更精确、残差尺度更小的基线天然能够改善下游任务的表现。本文证明,仅凭残差尺度不足以作为可学习性的判断依据。通过在总能量目标上评估分子图神经网络,我们发现复杂的局部描述符基线虽然能产生尺度较小的残差目标,但这些目标在基于架构信息的代理空间中却异常粗糙,且相对于其尺度更难学习。相反,半经验基线在降低残差尺度的同时也降低了归一化粗糙度,从而提升了域内和域外预测性能。我们提出尺度归一化图狄利克雷粗糙度($D_{\text{IQR}}$)作为残差可学习性的预训练诊断指标,并将基线互补性确立为目标设计的核心原则,使目标空间构建与模型架构设计并重,成为科学机器学习的另一关键维度。
cs.LG / 27 / 2609.28792

Vector Bellman Theory for Multichain Robust Average-Reward Markov Decision Processes

多链鲁棒平均回报马尔可夫决策过程的向量贝尔曼理论
Wang, Yue, Atia, George
Abstract
Robust average-reward Markov decision processes provide a fundamental framework for long-term performance optimization under uncertainty, and can have optimal long-run rewards that depend on the initial state. This state dependence requires a vector Bellman theory that accounts for both recurrent-class rewards and transition uncertainty. We develop such a theory for finite models with compact, post-action $(s,a)$-rectangular ambiguity. A gain-first, bias-second optimization principle yields a coupled vector gain-bias system, and every finite solution identifies the optimal robust gain and supplies stationary saddle strategies against history-dependent opponents, simultaneously from all initial states. We further characterize solvability through stationary gain conditions and a uniform bound on canonical transient corrections, and give sufficient conditions that permit distinct recurrent-class gains. The certificates also yield asymptotically affine trajectories of the robust Bellman operator, based on which we design a robust approximately shifted Halpern planning algorithm. Under finite Bellman solvability, the gain estimates and Bellman displacements converge to the optimal gain vector, and every extracted greedy controller is average-optimal after a finite, instance-dependent budget. These results thus connect finite Bellman certificates to undiscounted planning for state-dependent robust average rewards, providing theoretical understandings.
Chinese Translation
鲁棒平均回报马尔可夫决策过程为不确定性下的长期性能优化提供了基本框架,其最优长期回报可能依赖于初始状态。这种状态依赖性需要一个同时考虑常返类回报与转移不确定性的向量贝尔曼理论。我们针对具有紧的、动作后$(s,a)$-矩形不确定集的有限模型发展了这样一个理论。“先增益、后偏差”的优化原理导出一个耦合的向量增益-偏差方程组,且该方程组的每个有限解都能识别最优鲁棒增益,并从所有初始状态同时提供针对历史依赖对手的平稳鞍点策略。我们进一步通过平稳增益条件和典型瞬态校正的一致界刻画了可解性,并给出允许不同常返类增益的充分条件。这些证明条件还导出鲁棒贝尔曼算子的渐近仿射轨迹,基于此我们设计了一个鲁棒近似移位Halpern规划算法。在有限贝尔曼可解性条件下,增益估计与贝尔曼位移收敛于最优增益向量,且每个提取出的贪婪控制器在有限的、与实例相关的预算之后达到平均最优。这些结果从而将有限贝尔曼证明条件与状态依赖鲁棒平均回报的无折扣规划联系起来,提供了理论上的理解。
cs.LG / 28 / 2609.28793

Monitoring Urban Traffic Dynamics at Fine Spatiotemporal Resolution Using Distributed Acoustic Sensing and Deep Learning

基于分布式声波传感与深度学习的高时空分辨率城市交通动态监测
Tian, Hao, Cai, Heng, Chen, Xiaowei, Yang, Yifan
Abstract
Mapping the distribution of traffic dynamics at high spatiotemporal resolution is a fundamental question in transportation research. Distributed acoustic sensing (DAS), an innovative seismic observation tool, emerges as a promising solution for real-time urban traffic monitoring at high spatial and temporal scales. Distributed acoustic sensing repurposes existing underground fiber-optic cables as dense, continuous sensor arrays, enabling passive and privacy-preserving monitoring of roadway traffic activity at meter-level spatial and second-level temporal resolution. This study examines whether integrating DAS and deep learning models can serve as a continuous and efficient urban traffic observatory for revealing urban traffic dynamics (i.e. traffic volume and congestion, event-driven changes) at high spatiotemporal resolution. Using a DAS deployment along a roadway network in the City of College Station, Texas, USA, this study develops a deep learning-empowered analytical framework that converts raw ground vibration waveforms into spatiotemporal representations, detects vehicle trajectory, and infers traffic states from aggregated traffic volume and speed. A hybrid training strategy combining synthetic and manually annotated DAS images is used to improve vehicle detection under noisy and congested conditions, with model outputs further aggregated to characterize system-level traffic dynamics.
Chinese Translation
在高时空分辨率下刻画交通动态的分布是交通研究中的一个基础性问题。分布式声波传感(Distributed Acoustic Sensing, DAS)作为一种创新性的地震观测工具,为实现高空间与高时间尺度的城市交通实时监测提供了一种有前景的解决方案。分布式声波传感将现有的地下光缆改造为密集、连续的传感器阵列,能够以米级空间分辨率和秒级时间分辨率对道路交通活动进行被动式、保护隐私的监测。本研究探讨了将DAS与深度学习模型相结合,能否作为一个连续、高效的城市交通观测系统,以高时空分辨率揭示城市交通动态(即交通流量与拥堵、事件驱动的交通变化)。基于在美国得克萨斯州学院市(College Station)沿道路网络部署的DAS系统,本研究构建了一个深度学习驱动的分析框架,将原始地面振动波形转换为时空表征,检测车辆轨迹,并通过汇聚的交通流量和速度信息推断交通状态。研究采用结合合成数据与人工标注DAS图像的混合训练策略,以提升在噪声和拥堵条件下车辆检测的性能,并将模型输出进一步汇聚,以刻画系统层面的交通动态。
cs.LG / 29 / 2609.28809

Stream Recursion Model (SRM)

流递归模型(SRM)
Sorensen, Asael, Brock, Charles, Chamberlain, David, Minnich, Jennifer, Hoffman, Matthew, Ramyaa, Ramyaa
Abstract
Mechanistic interpretability seeks to make verifiable statements about the internal behavior of large language models (LLMs). Many interpretability techniques struggle to scale with the increasing size and depth of architectures. Our solution to this is to introduce smaller models with structures that lend themselves to interpretability. In this work, we introduce the Stream Recursion Model (SRM), a modification of the Hierarchical Reasoning Model (HRM) designed to expose internal computational structure while remaining scalable. SRM organizes computation into multiple interacting latent streams that are updated through recursive refinement, enabling direct analysis of stream dynamics, causal contribution, and routing behavior. SRM achieves performance comparable to GPT-2 on a per-parameter basis. Our analysis reveals consistent and distinct behavior across streams, indicating structured specialization and interaction. These results suggest that SRM provides a practical architectural foundation for scalable mechanistic interpretability and opens up promising avenues for future research in both reasoning performance and interpretability.
Chinese Translation
机械可解释性(mechanistic interpretability)旨在对大语言模型(LLM)的内部行为做出可验证的陈述。许多可解释性技术难以随架构规模和深度的不断增加而扩展。我们的解决方案是引入结构上更利于可解释性的较小模型。在本工作中,我们提出了流递归模型(Stream Recursion Model, SRM),它是对层次推理模型(Hierarchical Reasoning Model, HRM)的一种改进,旨在暴露内部计算结构的同时保持可扩展性。SRM 将计算组织为多个相互作用的潜在流(latent streams),并通过递归精炼进行更新,从而能够直接分析流动力学、因果贡献以及路由行为。SRM 在单位参数性能上可与 GPT-2 相媲美。我们的分析揭示了各流之间一致且彼此 distinct 的行为模式,表明存在结构化的专门化与交互。这些结果表明,SRM 为可扩展的机械可解释性提供了一个实用的架构基础,并为未来在推理性能与可解释性方面的研究开辟了有前景的方向。
cs.LG / 30 / 2609.28832

When Does Unsupervised Learning Succeed or Fail? A PoS Perspective on Reconstruction-Based Anomaly Detection

无监督学习何时成功或失败?从子空间视角看基于重构的异常检测
Yamaç, Mehmet, Mustu, Yagmur, Yousaf, Muhammad Numan, Xu, Lei, van Gerven, Marcel
Abstract
Reconstruction-based unsupervised learning can fail in two opposing ways: a model may reconstruct anomalies too accurately or discard valid nominal variation. Using the Pursuit of Subspaces hypothesis, we characterize these failures through the meet, union, and join geometries induced by the nominal components. Excess learned range produces join blindness, while insufficient capacity produces meet preference and loss of nominal fidelity. We show that the compact nominal union is optimal among nominal faithful ranges and generally requires a nonlinear reconstruction map. Based on this geometry, we introduce Dynamic Push and Pull, which learns from controlled perturbations without anomaly labels, and nested manifold carving, which applies the same principle recursively in latent space. Experiments confirm the predicted changes in latent geometry across every tested Push and Pull configuration. The proposed methods improve reconstruction-based anomaly detection across standard benchmarks and unseen image degradations, while also improving pretrained ECG representations for downstream classification. These results connect reconstruction failures to identifiable geometric conditions and provide practical mechanisms for learning compact representations.
Chinese Translation
基于重构的无监督学习可能以两种相反的方式失败:模型可能过于准确地重构异常,或者丢弃有效的正常样本变化。利用子空间追寻假设,我们通过正常样本成分所诱导的交、并和连接几何结构来刻画这些失败。学习到的范围过大导致连接盲视,而容量不足则导致交偏好和正常样本保真度的损失。我们证明,紧凑的正常并集在保持正常样本保真度的范围中是最优的,且通常需要非线性的重构映射。基于这一几何结构,我们提出了动态推拉,它在没有异常标签的情况下从受控扰动中学习;以及嵌套流形雕刻,它在潜空间中递归地应用相同原理。实验证实了在所有测试的推拉配置中潜空间几何结构均出现了预测的变化。所提出的方法在标准基准和未见过的图像退化条件下提升了基于重构的异常检测性能,同时也改进了预训练的ECG表征在下游分类任务中的表现。这些结果将重构失败与可识别的几何条件联系起来,并为学习紧凑表征提供了实用的机制。
cs.LG / 31 / 2609.28845

LastOPD: Taming Collapse in Latent On-Policy Distillation

LastOPD:驯服潜在在线策略蒸馏中的崩溃现象
Yang, Jie, Fang, Zhengyu, Xu, Zelin, Sun, Jiarui, Fan, Xiran, Wang, Junpeng, Wang, Liang, Liu, Qinghua, Cai, Yiwei, Zheng, Yan
Abstract
On-policy distillation (OPD) corrects a student on the responses it writes, but its signal is the teacher's next-token distribution: it tells the student what the teacher says but misses how it thinks. Latent supervision promises the missing part by aligning the student's latent states to the teacher's. Recent methods such as OPRD bring this signal into on-policy distillation. However, we observe two failures of this recipe when distilling Qwen3-4B and Qwen3-8B into Qwen3-1.7B-Base. Early gain, late collapse: latent supervision alone lifts MATH-500 accuracy from 25 to 46 in 10 steps, but subsequent training degrades performance down to 11 with no recovery. Better alignment, worse behavior: although the alignment metric steadily improves throughout this collapse, the most aligned model turns out to be the worst performing. Further analysis suggests a mismatch in how the latent signal is applied: layers paired by depth play different roles in the two models, so continued alignment may pull the student toward teacher states it cannot understand. To address this, we propose LastOPD, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD. This keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in. Extensive experiments show that LastOPD improves MATH-500 over token-only OPD by 5.55 and 4.02 points with the 4B and 8B teachers, leads on most held-out datasets, and reaches the final score of token-only OPD in about half the steps. Code is available at https://github.com/Muyiiiii/LastOPD.
Chinese Translation
在线策略蒸馏(On-policy distillation, OPD)通过学生模型自身生成的回复对其进行纠正,但其信号来自教师模型的下一个词元分布:它告诉学生教师说了什么,却遗漏了教师如何思考。潜在监督通过将学生的潜在状态与教师的潜在状态对齐,有望补上缺失的部分。诸如 OPRD 等近期方法将这一信号引入了在线策略蒸馏。然而,在将 Qwen3-4B 和 Qwen3-8B 蒸馏到 Qwen3-1.7B-Base 的过程中,我们观察到该方案存在两个失效模式。其一是早期增益、后期崩溃:仅用潜在监督可将 MATH-500 准确率在 10 步内从 25 提升至 46,但随后的训练使性能退化至 11 且无法恢复。其二是对齐越好、行为越差:尽管在整个崩溃过程中对齐指标持续改善,但对齐程度最高的模型却表现最差。进一步分析表明,问题在于潜在信号的应用方式存在错配:按深度配对的层在两个模型中扮演的角色不同,因此持续的对齐可能将学生拉向其无法理解的教师状态。为解决这一问题,我们提出 LastOPD,它仅在最后一层状态(即两个语言模型头共同读取的公共接口)上应用潜在信号,且仅在一个为期 10 步、过渡到词元级 OPD 的渐变阶段内使用。这样既保留了潜在信号中有用的部分,又在崩溃发生之前将学生交由词元级监督。大量实验表明,LastOPD 在 MATH-500 上分别比仅用词元级 OPD(使用 4B 和 8B 教师)提升 5.55 和 4.02 分,在大多数留出数据集上领先,并以约一半的步数达到仅用词元级 OPD 的最终分数。代码可在 https://github.com/Muyiiiii/LastOPD 获取。
cs.LG / 32 / 2609.28868

Image Fidelity is Not Field Fidelity: Joint Thermodynamic Reconstruction and Error Localization in Neural Tomography

图像保真度并非物理场保真度:神经断层成像中的热力学联合重建与误差定位
Hsu, Alan, Samra, Jenna, Paraschiv, Alin Razvan, Connor, Liam
Abstract
Neural fields for scientific tomography are optimized from 2D images, but the actual quantity of interest is often a latent 3D physical field. Because the forward map is many-to-one, low 2D image error need not certify a correct 3D field. Moreover, the latent field is not directly supervised during training, and its error cannot be evaluated against truth at deployment. We develop CoroNeRF to jointly optimize 3D electron density and temperature fields directly from multiview, multiline intensities through a differentiable atomic-emission renderer. Using solar coronal tomography as a controlled testbed, we evaluate physical-field recovery and test whether cross-seed instability provides a ground-truth-free-at-inference indicator of local physical-field error. We underscore the following two observations. (i) Image fidelity is not field fidelity: spectral ablations show that limited-channel reconstructions can fit their available observations well while recovering substantially worse fields, whereas evaluation on a common richer probe exposes the discrepancy. (ii) Cross-seed instability ranks local physical-field error across tested matched-model conditions, supported by sparsification and physical signal-strength controls. Seed-deviation projections provide complementary directional validation, but shared forward-model mismatch can still produce incorrect cross-seed consensus. These results characterize joint thermodynamic recovery and the usefulness and limits of seed-based error localization in a controlled, single-scene solar tomography testbed.
Chinese Translation
用于科学断层成像的神经场(Neural Fields)是从二维图像中优化的,但实际关注的量往往是一个潜在的(隐式的)三维物理场。由于正向映射是多对一的,较低的二维图像误差并不能保证三维物理场的正确性。此外,潜在物理场在训练过程中没有直接的监督信号,且在部署时无法对照真值评估其误差。我们提出了CoroNeRF,通过可微分的原子发射渲染器,直接从多视角、多谱线强度数据中联合优化三维电子密度场和温度场。我们以日冕断层成像作为受控测试平台,评估物理场的恢复效果,并检验跨随机种子(cross-seed)不稳定性是否能够作为一种在推理阶段无需真值的局部物理场误差指标。我们强调以下两点观察结果:(i)图像保真度并非物理场保真度:光谱消融实验表明,受限通道的重建能够在可用观测数据上拟合得很好,同时恢复出的物理场却明显更差,而在统一的更丰富探测数据上的评估则会暴露出这种差异。(ii)在所有测试的匹配模型条件下,跨随机种子不稳定性均能对局部物理场误差进行排序,该结论得到了稀疏化实验和物理信号强度对照实验的支持。种子偏差投影提供了补充性的方向验证,但共享的正向模型失配仍可能导致错误的跨种子共识。这些结果在一个受控的单场景太阳断层成像测试平台上,刻画了热力学联合恢复的效果,以及基于种子的误差定位方法的有效性与局限性。
cs.LG / 33 / 2609.28935

Response-state Learning for Transferable Vibrational Spectroscopic Characterization with Electron Prior

基于电子先验的可迁移振动光谱表征的响应态学习
Li, Zetong, Xie, Zhuosong, Fan, Hengyu, Yu, Jiaao, Hua, Qiyao, Lu, Zheng, Xu, Liming, Wu, Juanni, Li, Honglin
Abstract
Vibrational spectral prediction can become inaccurate when localized stereoelectronic environments perturb intermediate response states and high-risk response units dominate characteristic spectral fingerprints, making prediction across external chemical space difficult. SO(3) Equivariant Neural Kalman Networks (SENK) form a response-state cascade that combines an equivariant transformer backbone for Hessian, dipole-derivative and polarizability-derivative learning, an Equivariant Neural Kalman bridge for state-dependent refinement and reliability sensing, and an NBO-informed electronic-prior pathway coupling consistency regularization with bounded, branch-specific guided spectral calibration. SENK outperforms DetaNet on QM9S and QMe14S while preserving full-spectrum IR and Raman fidelity from small molecules to drug-like systems. SENK remains stable and selectively improves spectrally sensitive features in biomolecular systems with complex stereoelectronic effects. It therefore integrates tensor prediction, reliability diagnosis and physics-informed calibration, supporting transferable vibrational spectroscopy from molecular systems to functional molecular materials.
Chinese Translation
当局域立体电子环境扰动中间响应态、且高风险响应单元主导特征光谱指纹时,振动光谱预测可能变得不准确,使得跨外部化学空间的预测变得困难。SO(3)等变神经卡尔曼网络(SENK)构建了一个响应态级联结构,其组合包括:用于学习Hessian、偶极导数和极化率导数的等变Transformer主干网络,用于状态依赖的细化与可靠性感知的等变神经卡尔曼桥(Equivariant Neural Kalman bridge),以及将一致性正则化与有界的、分支特异的引导光谱校准相耦合的、基于NBO信息的电子先验路径。SENK在QM9S和QMe14S数据集上优于DetaNet,同时在小分子至类药物分子体系中保持了全谱红外(IR)和拉曼(Raman)光谱的保真度。在具有复杂立体电子效应的生物分子体系中,SENK保持稳定并有选择性地改进了光谱敏感特征。因此,SENK整合了张量预测、可靠性诊断与物理信息引导的校准,支持从分子体系到功能分子材料的可迁移振动光谱表征。
cs.LG / 34 / 2609.28979

Spectral Graph Neural Networks with Hermite Polynomials: A Comprehensive Study

基于Hermite多项式的谱图神经网络:一项综合研究
Wu, Shuang
Abstract
We study spectral graph neural networks built from Hermite polynomials and propose HermNet, a simple model that combines a nodewise predictor with normalized Hermite propagation. Its sparse recurrence requires neither eigendecomposition nor a learned basis. We distinguish the basic model from optional coordinate calibration, response normalization and Gaussian derivative regularization. Hermite and other complete polynomial bases span the same degree-bounded filter space, but their coordinates can produce different optimization behavior under limited training budgets. We analyze this behavior through spectral signal energy, label sampling, changes in learned features and the bias--variance trade-off of regularization. Controlled synthetic experiments identify a regime in which plain HermNet outperforms matched polynomial-basis alternatives, including with a jointly trained nonlinear predictor. Curvature regularization further improves HermNet when the same functional penalty is available to every comparator. Fixed-predictor controls support the advantage under short training budgets, but longer training removes the plain-model lead. Matched real-data comparisons show accuracy deficits, and architectural and numerical studies identify further limits. Together, the analysis and experiments clarify when Hermite propagation is useful and how calibration and regularization affect its performance.
Chinese Translation
我们研究了由Hermite多项式构建的谱图神经网络,并提出HermNet——一个将逐节点预测器与归一化Hermite传播相结合的简洁模型。其稀疏递推结构既不需要特征分解,也不需要学习基。我们区分了基础模型与可选的坐标校准、响应归一化和高斯导数正则化。Hermite及其他完备多项式基张成相同的度数受限滤波器空间,但在有限训练预算下,不同基的坐标可能产生不同的优化行为。我们通过谱信号能量、标签采样、学习特征的变化以及正则化的偏差—方差权衡来分析这种行为。受控合成实验识别出一个 regime,在该 regime 下,简单的HermNet优于匹配的多项式基替代方法,包括与联合训练的非线性预测器结合时。当相同的功能性惩罚可用于所有对比方法时,曲率正则化进一步提升了HermNet的表现。固定预测器的对照实验支持了该方法在短训练预算下的优势,但更长时间的训练会消除简单模型的领先优势。匹配的真实数据对比显示存在精度差距,架构与数值研究进一步揭示了其局限性。综合分析与实验阐明了Hermite传播在何时有效,以及校准和正则化如何影响其性能。
cs.LG / 35 / 2609.28998

Automatic Rank Allocation for Low-Rank Adaptation in Large Language Models via lp Regularization

基于ℓp正则化的大语言模型低秩自适应秩自动分配方法
Xie, Zebang, Zheng, Chuanyang, Wu, Yik-Chung, Gao, Yihang
Abstract
Low-rank adaptation (LoRA) has become a popular parameter-efficient fine-tuning method for large language models. A key challenge in LoRA is how to determine the rank of each adaptation matrix, as rank directly controls its capacity and efficiency. Existing adaptive-rank methods typically allocate ranks according to manually designed importance scores, which are not directly derived from an optimization objective. In this work, we propose $\ell_p$-LoRA, a principled rank-allocation method based on $\ell_p$ regularization with $0<p<1$, which is a classical sparsity-inducing technique in signal processing and statistics. Specifically, we regularize the energy of each rank-one LoRA component, encouraging redundant components to vanish while preserving important ones. We derive the corresponding proximal subproblem and reduce the matrix optimization to a two-dimensional problem, leading to an implicit thresholding criterion for identifying redundant components. Experiments on natural language understanding and question-answering tasks demonstrate that the proposed method achieves competitive performance with existing LoRA baselines.
Chinese Translation
低秩自适应已成为大语言模型中流行的参数高效微调方法。LoRA的一个关键挑战是如何确定每个自适应矩阵的秩,因为秩直接控制其容量和效率。现有的自适应秩方法通常根据人工设计的重要性分数来分配秩,而这些分数并非直接来源于优化目标。在本工作中,我们提出ℓp-LoRA,这是一种基于ℓp正则化(0<p<1)的具有理论依据的秩分配方法,ℓp正则化是信号处理与统计学中经典的稀疏性诱导技术。具体而言,我们对每个秩一LoRA分量的能量进行正则化,促使冗余分量趋于消失,同时保留重要分量。我们推导了相应的近端子问题,并将矩阵优化问题降为二维问题,从而得到一个用于识别冗余分量的隐式阈值准则。在自然语言理解和问答任务上的实验表明,所提方法取得了与现有LoRA基线相当的性能。
cs.LG / 36 / 2609.29000

Learning from Mixed-Quality Deployment Experience for Robot Manipulation

从混合质量的部署经验中学习机器人操作
Ren, Yangang, Yan, Yujie, Li, Zirui, Guo, Jiaming, Zeng, Di, Tao, Ji, Yu, Lan, Tian, Xuesong, Lv, Chen
Abstract
Robot policies deployed in real environments naturally accumulate mixed-quality experience, including successful executions, partial progress, and failures. Although these rollouts provide valuable information for further learning, directly incorporating them into imitation learning may reinforce undesirable behaviors, while offline reinforcement learning often suffers from unreliable value estimation under sparse rewards and limited data coverage. We consider a practical post-deployment setting where learning relies only on naturally accumulated autonomous rollouts, without additional human corrections or exploratory interaction. To effectively exploit such experience, we propose Predictive Action Chunk Learning (PACL). PACL first learns a predictive chunk-level critic that evaluates temporally extended action sequences and augments temporal difference learning with future latent prediction, providing richer supervision for long-horizon value estimation. The learned critic then converts chunk-level Q-values into discrete quality conditions, which guide a diffusion actor to learn jointly from these mixed-quality experiences without treating all behaviors as equivalent supervision. At inference, the actor generates multiple action chunks and the critic selects the highest valued candidate. Experiments across simulated and real-world robot manipulation tasks show that PACL consistently improves the pretrained policy and outperforms strong imitation learning and offline reinforcement learning baselines.
Chinese Translation
部署在真实环境中的机器人策略会自然积累混合质量的经验,包括成功执行、部分进展和失败。尽管这些轨迹数据为进一步学习提供了宝贵信息,但直接将其纳入模仿学习可能会强化不良行为,而离线强化学习在稀疏奖励和有限数据覆盖条件下往往存在价值估计不可靠的问题。我们考虑一种实际的部署后学习场景,即学习仅依赖自然积累的自主轨迹数据,无需额外的人工纠正或探索性交互。为了有效利用这类经验,我们提出了预测性动作块学习(Predictive Action Chunk Learning, PACL)。PACL首先学习一个预测性的块级评论家(critic),用于评估时间上扩展的动作序列,并通过未来隐变量预测增强时序差分学习,为长时程价值估计提供更丰富的监督信号。随后,学习到的评论家将块级Q值转化为离散的质量条件,引导扩散演员(diffusion actor)从这些混合质量经验中联合学习,而不将所有行为视为等同的监督信号。在推理阶段,演员生成多个动作块,由评论家选出价值最高的候选。在仿真和真实世界机器人操作任务上的实验表明,PACL能够持续改进预训练策略,并优于强大的模仿学习和离线强化学习基线方法。
cs.LG / 37 / 2609.29024

Growth-Inspired Graph Generation and Inverse Design of Mechanical Lattices via Dot Matrices Database Augmentation and GCNN

基于点阵数据库增强与图卷积神经网络的生长启发式图生成及机械点阵逆向设计
Xu, Weiyun, Liu, Jiamu
Abstract
Natural load-bearing and transport networks are not assembled in a single step; they emerge through a temporally ordered process of growth, branching, reinforcement, and loop formation. Inspired by this developmental logic, this work introduces a morphogenetic graph-generation framework for mechanical lattices in which a discrete dot matrix provides potential nodes and the final architecture is created by sequential cross-layer and intra-layer growth. The same rule is visualized in two dimensions as a leaf-vein-like developmental sequence and implemented in three dimensions on a 3x3x3 nodal matrix containing 27 candidate nodes. A dataset of distinct three-dimensional lattices was evaluated by beam-based finite element analysis and represented directly as graphs. A graph convolutional neural network (GCNN) with three graph-convolution layers and dual global pooling learns the topology-property mapping and predicts effective compressive stiffness. Coupling the GCNN surrogate with rapid structural sampling enables inverse design: for a target stiffness of 1000 MPa, the selected design was predicted at 1042.43 MPa and validated by finite element analysis at 1027.49 MPa. Beyond straight members, the framework has also been extended to parameterized horseshoe-shaped curved beams made of nonlinear materials, enabling topology-geometry design toward prescribed deformation shapes. Our work provides a paradigm for augmenting the database of mechanical metamaterials, and the resulting perspective links biological morphogenesis, graph learning, and nonlinear shape programming in a unified generative design framework for architected materials.
Chinese Translation
自然界中的承载与输运网络并非一次性组装而成,而是通过生长、分支、强化与成环的时间有序过程逐步形成。受这一发育逻辑启发,本工作提出了一种面向机械点阵的形态发生式图生成框架:离散点阵提供潜在节点,最终结构由逐层的跨层与层内生长序列构建。该规则在二维空间中呈现为类似叶脉的发育序列,并在三维空间中于包含27个候选节点的3×3×3节点矩阵上实现。通过基于梁的有限元分析评估了由不同三维点阵构成的数据集,并将其直接表示为图。一个包含三层图卷积和双重全局池化的图卷积神经网络(GCNN)学习拓扑-性能映射,预测有效压缩刚度。将GCNN代理模型与快速结构采样相结合实现了逆向设计:对于1000 MPa的目标刚度,所选设计预测值为1042.43 MPa,有限元验证值为1027.49 MPa。除直杆构件外,该框架还扩展至由非线性材料构成的参数化马蹄形曲梁,实现了面向指定变形形状的拓扑-几何设计。我们的工作为增强机械超材料数据库提供了范式,由此形成的视角将生物形态发生、图学习与非线性形状编程统一于结构材料的生成式设计框架之中。
cs.LG / 38 / 2609.29027

Generative Atmospheric Super-Resolution from Heterogeneous In Situ Observations through Composable Interfaces

通过可组合接口利用异构原位观测实现生成式大气超分辨率
Xu, Yang, Chakraborty, Dibyajyoti, Guan, Haiwen, Wang, Sen, Maulik, Romit
Abstract
Atmospheric observations are sparse, heterogeneous, and unevenly distributed, whereas many generative atmospheric models learn distributions over regularly gridded multivariate states. Once pretrained, diffusion models can supply atmospheric priors that can be combined with observation-derived likelihood factors in a Bayesian formulation. However, these observation sources differ substantially in geometry and sampling density, complicating the consistent use of their observations within a common inference framework. Here, we formulate this reconstruction problem as generative atmospheric super-resolution and introduce composable observation interfaces for conditioning a single pretrained 13-variable atmospheric diffusion model. The interfaces convert sparse radiosonde (R), clustered aircraft (A), and dense irregular surface-station (S) observations into source-specific likelihood factors that specify where observations constrain the gridded state, how residuals are counted under uneven sampling, and how strongly each source guides posterior sampling. We developed the aircraft and surface observation interfaces using 2019 observations and evaluated the selected interfaces throughout 2020 without further tuning. Compared with reconstructions conditioned only on radiosonde observations, the composed R+A+S interface reduces RMSE evaluated against ERA5 by $9.24\%$ across all 13 state variables over the CONUS domain. The aircraft and surface factors provide complementary improvements in upper-air and surface variables. The R+A+S combination also lowers the Continuous Ranked Probability Score (CRPS), while evaluations at held-out aircraft and surface-station observations show reduced prediction errors. Together, these results demonstrate a modular route for conditioning a pretrained atmospheric generative prior on heterogeneous in situ observations without retraining the underlying model.
Chinese Translation
大气观测稀疏、异构且分布不均,而许多生成式大气模型是在规则网格化的多变量状态上学习分布的。预训练完成后,扩散模型可以提供大气先验,并在贝叶斯框架下与由观测导出的似然因子相结合。然而,这些观测源在几何形态和采样密度上差异巨大,使得在统一推断框架内一致地使用这些观测变得复杂。本文将这一重建问题表述为生成式大气超分辨率,并引入可组合的观测接口,用于条件化单个预训练的13变量大气扩散模型。这些接口将稀疏的探空仪(R)观测、聚类的飞机(A)观测以及稠密且不规则的地面站(S)观测转换为针对各数据源的似然因子,用以指定观测在何处约束网格化状态、在不均匀采样下如何统计残差,以及每个数据源对后验采样的引导强度。我们使用2019年的观测数据开发了飞机和地面观测接口,并在整个2020年内对所选接口进行评估,无需进一步调参。与仅以探空仪观测为条件的重建相比,组合的R+A+S接口在美国本土(CONUS)区域上,相对ERA5评估的全部13个状态变量的RMSE降低了9.24%。飞机和地面似然因子在高空变量和地面变量方面提供了互补的改进。R+A+S组合还降低了连续分级概率评分(CRPS),并且在留出的飞机和地面站观测上的评估显示预测误差有所降低。这些结果共同表明了一条模块化路径,可将预训练的大气生成先验条件化于异构原位观测,而无需对底层模型进行再训练。
cs.LG / 39 / 2609.29069

BranchShine-CR: Compact Multilingual IPA Transcription with Self-Conditioned CTC and Consistency Regularization

BranchShine-CR:基于自条件CTC与一致性正则化的紧凑型多语言国际音标(IPA)转写模型
Navas, Nikhil, Chevtchenko, Sergio, Damiao, Talisson, Afshar, Saeed
Abstract
We introduce BranchShine-CR, a 25M-parameter model for multilingual transcription into the International Phonetic Alphabet (IPA). It combines log-mel features, a rotary-position E-Branchformer encoder, intermediate self-conditioned connectionist temporal classification (CTC), and consistency regularization across augmented views. On 16,646 shared IPApack++ test utterances, it achieves 4.47% IPA character error rate, a 22.3% relative reduction from ZIPA-CTC-NS, with approximately one-twelfth as many parameters while being trained from scratch. BranchShine-CR also outperforms a similarly sized NeMo Conformer baseline across all 41 dataset language labels. Ablation studies indicate the individual components synergetically acting in model performance contribution. These findings support compact IPA recognition capabilities under limited compute budget, for applications in low-resource on-device pronunciation assessment.
Chinese Translation
我们提出了BranchShine-CR,一个用于多语言国际音标(IPA)转写的2500万参数模型。该模型结合了对数梅尔特征、旋转位置编码的E-Branchformer编码器、中间层自条件连接时序分类(CTC)以及跨增强视图的一致性正则化。在16,646条共享的IPApack++测试语句上,该模型实现了4.47%的IPA字符错误率,相较于ZIPA-CTC-NS相对降低了22.3%,且参数量约为后者的十二分之一,同时完全从零训练。BranchShine-CR还在全部41个数据集语言标签上优于规模相近的NeMo Conformer基线模型。消融实验表明各组件在模型性能贡献上具有协同作用。这些结果支持在有限计算预算下实现紧凑型IPA识别能力,可应用于低资源端侧发音评估。
cs.LG / 40 / 2609.29087

Physics and Data Driven Transformer-Mamba Framework for Flow Field

面向流场的物理与数据驱动的Transformer-Mamba框架
Zhang, Zhuo, Zou, Shun, Yang, Canqun, Yang, Xi
Abstract
While deep learning accelerates expensive partial differential equation solving in computational fluid dynamics (CFD), existing methods like PINNs and FNOs often struggle with generalization, noise robustness, and physical consistency. We introduce the Transformer-Mamba for Flow Field (TM4FF) framework, a physics-constrained operator learning model with three key innovations: a Residual Wavelet Mamba (RWM) layer for feature denoising, a Transformer-based attention mechanism for enhanced feature fusion, and a physics-informed loss using Fourier derivatives to enforce the Navier-Stokes equations. Experiments on four CFD datasets show TM4FF achieves high accuracy and robust generalization across varying flow conditions.
Chinese Translation
尽管深度学习加速了计算流体力学(CFD)中昂贵的偏微分方程求解,但现有方法如物理信息神经网络(PINNs)和傅里叶神经算子(FNOs)在泛化能力、噪声鲁棒性和物理一致性方面往往存在不足。我们提出了面向流场的Transformer-Mamba(TM4FF)框架,这是一个物理约束的算子学习模型,包含三项关键创新:用于特征去噪的残差小波Mamba(RWM)层、用于增强特征融合的基于Transformer的注意力机制,以及利用傅里叶导数施加Navier-Stokes方程的物理信息损失函数。在四个CFD数据集上的实验表明,TM4FF在不同流动条件下均实现了高精度和鲁棒的泛化能力。
cs.LG / 41 / 2609.29095

Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents

恰好一次语义究竟由谁负责?模型、代理框架与工具契约对LLM代理中重复副作用的影响
Li, Jiapeng
Abstract
When a tool-using agent's write times out or returns a server error, the action may already have taken effect. Retrying blindly duplicates it -- a second charge, a second announcement, a second deployment -- while giving up skips required work. We ask where exactly-once behaviour should be enforced: in the model, in the agent harness, or in the tool contract. We introduce LIMBO, a deterministic sandbox of six services with realistic contracts (optional idempotency keys, eventually consistent and missing read paths) and twelve fault modes injected at the service boundary, including late commits, redelivery and partial batches; every episode is graded against a ledger of committed effects. Across 25,930 episodes spanning nine recent models, three production agent harnesses, two contract variants and fifteen recovery conditions, the answer depends on the fault. When an immediate read-back can reveal what happened, the model decides: frontier models instructed to act exactly once almost never duplicate a write whose acknowledgement was lost (0.5%), weaker models often do, and the model explains 53% of the explained variance. When it cannot -- the request is still in flight, or the transport delivered it twice -- the same frontier models duplicate in 56% and 74% of episodes, and the contract explains 81%. We prove that no verification-only policy is exactly-once under late commits without a bound on in-flight time. Waiting works when such a bound is short and known, but with heavy-tailed in-flight delays even an hour of waiting per episode falls short of offering an idempotency key on every write, which lowers the duplicate rate from 28% to 4% because agents use keys when they exist. The harness barely matters, a guard that attaches keys transfers across harnesses unchanged, and agents reported success in 90% of the episodes in which they had duplicated an effect.
Chinese Translation
当使用工具的代理进行写入操作超时或返回服务器错误时,该操作可能已经生效。盲目重试会导致重复——重复扣款、重复发布、重复部署——而放弃重试则会遗漏必要的工作。我们要问的是:恰好一次(exactly-once)行为究竟应该在哪里强制实施:在模型中、在代理框架(harness)中,还是在工具契约中?我们提出了LIMBO,一个由六个服务组成的确定性沙盒环境,具有现实的契约(可选的幂等键、最终一致以及缺失的读取路径),并在服务边界注入十二种故障模式,包括延迟提交、重复投递和部分批次;每个回合(episode)都根据已提交副作用的账本进行评分。在涵盖九个最新模型、三种生产级代理框架、两种契约变体和十五种恢复条件的25,930个回合中,答案取决于故障类型。当立即读取回执能够揭示实际发生的情况时,由模型决定:被指示执行恰好一次操作的前沿模型几乎从不对确认丢失的写入进行重复(0.5%),较弱的模型则经常重复,模型解释了53%的可解释方差。当无法读取回执时——请求仍在传输中,或传输层已将其投递两次——同样的前沿模型在56%和74%的回合中发生重复,此时契约解释了81%的方差。我们证明,在存在延迟提交且不限制在途时间的条件下,任何仅依赖验证的策略都无法实现恰好一次语义。当在途时间上界较短且已知时,等待是有效的;但在重尾在途延迟的情况下,即使每个回合等待一小时,也不如对每次写入都提供幂等键——幂等键可将重复率从28%降至4%,因为代理在幂等键存在时会使用它们。代理框架的影响微乎其微,一个自动附加幂等键的守卫机制可以不加修改地跨框架迁移,此外,在发生重复副作用的回合中,代理在90%的情况下仍报告操作成功。
cs.LG / 42 / 2609.29096

Downside-Controlled Online Forecast Combination under Delayed and Revised Outcomes

延迟且修订结果下的下行风险受控在线预测组合
Kim, Minkyoung, Byun, Hyunjung, Lee, Yohan, Jang, Beakcheol
Abstract
Post-hoc correction adjusts a forecaster that cannot be retrained, such as a foundation model, but a correction fitted where errors are stable can hurt where they shift. We aim for downside control: not much worse than the starting forecast. We combine the frozen forecaster, a static corrector and an online corrector on the simplex, using only losses that mature after the horizon. Across seven benchmarks and four base models, two of them foundation models, the worst deterioration over 28 pairs at the main horizon is 0.15% and gains reach 11.5%. On day-ahead load for seven European bidding zones it lowers mean MSE in all seven zones, while single correctors raise mean MSE by up to 102% where the published forecast is most accurate. Three empirical conditions on expert speed, stream length and outcome alignment, each fixed by a documented failure, delimit its scope. Learning from the provisional outcome improves four zones on the settled one; learning on the settled outcome restores all seven.
Chinese Translation
事后校正适用于无法重新训练的预测器(如基础模型),但在误差稳定区域拟合的校正在误差发生漂移的区域可能产生负面影响。我们的目标是实现下行风险控制:不比初始预测差太多。我们在单纯形上组合冻结预测器、静态校正器和在线校正器,仅使用在地平线之后才成熟的损失。在七个基准数据集和四个基础模型(其中两个为基础模型)上,主地平线下28对组合的最坏性能退化仅为0.15%,而收益最高可达11.5%。对于七个欧洲竞价区域的日前负荷预测,该方法在所有七个区域均降低了平均MSE,而在已发布预测最准确的区域,单一校正器会使平均MSE升高多达102%。我们通过三次有记录的失败案例确定了关于专家速度、数据流长度和结果对齐的三个经验性条件,以此界定了该方法的适用范围。基于临时结果进行学习可在结算结果上改进四个区域;基于结算结果进行学习则可恢复全部七个区域的性能。
cs.LG / 43 / 2609.29101

Language Specificity vs. Domain Diversity: Benchmarking Transformers for Bangla Medical NER

语言特定性与领域多样性之争:面向孟加拉语医学命名实体识别的Transformer模型基准评测
Abdullah, Rakib, Maruf, Md. Maruful Islam
Abstract
Medical Named Entity Recognition (NER) for low-resource languages remains a challenging task due to high linguistic variability and a scarcity of domain-specific annotated corpora. This work presents a comprehensive empirical benchmark evaluating three fine-tuned transformer encoders-BanglaBERT, multilingual BERT (mBERT), and XLM-RoBERTa-against GPT-4o mini under zero-shot and few-shot prompting configurations for Bangla medical NER. In contrast to prior studies that evaluated large language models on limited subsets of only 50 samples, we conduct a large-scale evaluation across the full test set of 3,179 samples, providing statistically robust and reproducible baselines. Our fine-tuned XLM-RoBERTa model achieves an F1- score of 0.5959, establishing a new state-of-the-art and surpassing the previously reported best result of 0.5848. Crucially, we demonstrate that the language-specific BanglaBERT model consistently underperforms its multilingual counterparts with an F1-score of 0.4937, indicating that pretraining domain diversity can outweigh language specificity in highly specialized clinical settings. Furthermore, we present a detailed per-entity-type analysis for this task, revealing that Medicine and Specialist categories are recognized with high reliability, achieving F1- scores above 0.83, while the Symptom category remains the most challenging with an F1-score of 0.4367 despite being the most frequent training class. Finally, fine-tuned transformer models outperform the optimal prompting configuration by a factor of 3.76, confirming that prompt-only pipelines remain inadequate for structured clinical entity extraction in low-resource language environments.
Chinese Translation
针对低资源语言的医学命名实体识别(NER)任务,由于语言变异性高且领域特定标注语料匮乏,仍然是一项具有挑战性的工作。本文提出了一项全面的实证基准评测,对比了三种微调的Transformer编码器模型——BanglaBERT、多语言BERT(mBERT)和XLM-RoBERTa——与GPT-4o mini在零样本(zero-shot)和少样本(few-shot)提示配置下进行孟加拉语医学NER的表现。与以往仅在仅50个样本的有限子集上评估大语言模型的研究不同,我们在包含3,179个样本的完整测试集上进行了大规模评估,提供了具有统计稳健性和可复现性的基线结果。我们微调后的XLM-RoBERTa模型取得了0.5959的F1分数,确立了新的最先进水平(state-of-the-art),超过了此前报道的最佳结果0.5848。关键的是,我们证明语言特定的BanglaBERT模型始终表现不及多语言对应模型,其F1分数仅为0.4937,这表明在高度专业化的临床场景中,预训练的领域多样性可以胜过语言特定性。此外,我们针对该任务提供了详细的按实体类型分析,结果显示Medicine(药物)和Specialist(专科医生)类别能够以高可靠性被识别,F1分数超过0.83,而Symptom(症状)类别尽管是训练集中出现频率最高的类别,却仍然最具挑战性,F1分数仅为0.4367。最后,微调的Transformer模型以3.76倍的优势超越最优提示配置,这证实了仅依赖提示的流水线在低资源语言环境下仍不足以完成结构化临床实体抽取任务。
cs.LG / 44 / 2609.29117

A Concentration Bound for Two-Timescale Actor-Critic Algorithm

两时间尺度Actor-Critic算法的集中不等式界
Panda, Prashansa, Bhatnagar, Shalabh
Abstract
Significant research effort has been directed in recent years towards establishing both asymptotic and non-asymptotic convergence guarantees for two-timescale actor--critic algorithms, where the actor recursion is run on a slower timescale than the critic recursion. This work derives a uniform all-time concentration bound for the actor--critic algorithm with function approximation in the long-run average-reward setting. This bound helps us analyze the behavior of the actor parameter with high probability. We show that, after some finite time, the actor parameter enters a safe region and remains within it thereafter with high probability. Specifically, with probability at least $1-\epsilon_1-\epsilon_2$, the actor error $\Vert \theta_k-\theta^{*}\Vert$ is $O\left(\frac{n_0^{3/4}}{k}\frac{1}{\sqrt{\epsilon_2}}+\left(\frac{1}{n_0}\right)^{1/4}\log^{1/4}\left(\frac{1}{\epsilon_1}\right)+\left(\frac{1}{n_0}\right)^{1/4}\right)$ for all $k\geq n_0$ and sufficiently large $n_0$. We also present experimental results demonstrating that the aforementioned actor error diminishes with the number of actor-parameter updates.
Chinese Translation
近年来,大量研究工作致力于为两时间尺度(two-timescale)actor-critic算法建立渐近与非渐近收敛保证,其中actor递归以比critic递归更慢的时间尺度运行。本文在长期平均回报设置下,推导了带函数逼近的actor-critic算法的一致全时段集中界(uniform all-time concentration bound)。该界有助于我们以高概率分析actor参数的行为。我们证明,在某个有限时间之后,actor参数以高概率进入一个安全区域并此后始终保持其中。具体而言,对于所有 $k\geq n_0$ 以及足够大的 $n_0$,以至少 $1-\epsilon_1-\epsilon_2$ 的概率,actor误差 $\Vert \theta_k-\theta^{*}\Vert$ 为 $O\left(\frac{n_0^{3/4}}{k}\frac{1}{\sqrt{\epsilon_2}}+\left(\frac{1}{n_0}\right)^{1/4}\log^{1/4}\left(\frac{1}{\epsilon_1}\right)+\left(\frac{1}{n_0}\right)^{1/4}\right)$。我们还给出了实验结果,表明上述actor误差随actor参数更新次数的增加而减小。
cs.LG / 45 / 2609.29142

Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD

并非每个词元都值得蒸馏:Direct-OPD的选择性监督
Zhao, Yibo, Yang, Zixuan, Lan, Yunshi, Li, Xiang
Abstract
Direct On-Policy Distillation (Direct-OPD) transfers reinforcement-learning-induced policy improvements from a small model to a larger student by using the token-level log-ratio between post-RL and pre-RL checkpoints as dense supervision on the student's own rollouts. This transfer rewards the policy shift at every state, yet the log-ratio measures only relative change: it can stay fixed even as the probability mass that both checkpoints assign to the student's candidate tokens vanishes. Through an exact construction, we show that the Direct-OPD reward and its update can remain unchanged while the Jensen-Shannon divergence (JSD) and both KL directions between the checkpoints vanish with this mass, and we note that a small JSD bounds how much the teacher's behavior changed. Motivated by this analysis, we propose Selective Supervision for Direct-OPD (S$^2$D-OPD), which ranks student-sampled states by their teacher-reference JSD and masks Direct-OPD supervision at low-divergence states, retaining only the top 10% of states per response. Across two teacher pairs and four student models ranging from 1.7B to 8B parameters, S$^2$D-OPD improves held-out accuracy over dense Direct-OPD on AIME and HMMT benchmarks in seven of eight settings and matches it in the eighth, without extra forward passes. Our code is available at https://anonymous.4open.science/r/S2D-OPD-8868.
Chinese Translation
直接在线策略蒸馏(Direct On-Policy Distillation,Direct-OPD)通过将强化学习后与强化学习前检查点之间的词元级对数比率作为密集监督信号,作用于学生模型自身的生成结果,从而将强化学习带来的策略改进从一个小模型迁移到更大的学生模型。这种迁移会对每个状态下的策略变化给予奖励,然而对数比率仅度量相对变化:即使两个检查点赋予学生候选词元的概率质量同时趋于消失,对数比率仍可保持不变。通过一个精确的构造,我们证明了在Direct-OPD奖励及其更新保持不变的同时,检查点之间的Jensen-Shannon散度(JSD)以及两个方向的KL散度会随该概率质量一同消失;我们还注意到,较小的JSD能够约束教师模型行为变化的幅度。基于这一分析,我们提出了面向Direct-OPD的选择性监督方法(Selective Supervision for Direct-OPD,S$^2$D-OPD),该方法根据教师-参考JSD对学生模型采样的状态进行排序,并在低散度状态上屏蔽Direct-OPD监督,每个回答仅保留散度最高的10%的状态。在两组教师模型对以及四个参数量从1.7B到8B的学生模型上,S$^2$D-OPD在八个设置中的七个上于AIME和HMMT基准上取得了优于密集Direct-OPD的留存集准确率,在第八个设置上与之持平,且无需额外的前向传播。我们的代码发布于 https://anonymous.4open.science/r/S2D-OPD-8868。
cs.LG / 46 / 2609.29150

A Particle-Swarm-Assisted Gradient Meta-Learning Algorithm for Joint Transmit Precoding and STAR-RIS Coefficient Optimization

一种粒子群辅助梯度元学习算法:用于联合发射预编码与STAR-RIS系数优化
Zhou, Kang
Abstract
This paper investigates the joint optimization of the transmit precoder and the transmission/reflection coefficients of a simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) to maximize the weighted sum rate (WSR) in a multi-user downlink. We propose a particle-swarm-assisted gradient meta-learning (PSA-GML) algorithm for this non-convex problem. The original problem is first equivalently transformed via an amplitude-split parameterization and a collapsed precoder representation, which automatically satisfy the energy-conservation constraint and reduce the search dimension. Particle swarm optimization (PSO) then performs a global search over the STAR-RIS coefficients to yield a high-quality, initialization-robust warm start, with the transmit precoder obtained in closed form. Departing from conventional alternating optimization (AO), a coordinate-wise long short-term memory (LSTM) meta-optimizer trained by first-order gradient meta-learning further refines the coefficients and precoder jointly, learning per-coordinate adaptive update rules from data. The meta-optimizer is trained offline and applied to unseen channels without further adaptation. Numerical results show that PSA-GML attains an 11.06 bits/s/Hz WSR at 10 dB with N=32 elements and K=4 users, exceeding AO by 13.1% (and by 6.2% even with multiple random restarts) and the random-phase scheme by 35.1%. In the interference-limited regime it reaches 83.9% of the hand-designed Adam refinement without manual hyper-parameter tuning, and it transfers zero-shot across regimes, indicating that the learned update rule captures the intrinsic WSR landscape structure.
Chinese Translation
本文研究了在多用户下行链路中,联合优化发射预编码器与同时透射和反射可重构智能表面(STAR-RIS)的透射/反射系数,以最大化加权和速率(WSR)。针对该非凸问题,我们提出了一种粒子群辅助梯度元学习(PSA-GML)算法。首先,通过幅度分离参数化和紧凑预编码表示,将原问题等价变换,从而自动满足能量守恒约束并降低搜索维度。随后,粒子群优化(PSO)在STAR-RIS系数上进行全局搜索,以获得高质量、对初始化鲁棒的暖启动,其中发射预编码器以闭式解求得。与传统交替优化(AO)不同,一种通过一阶梯度元学习训练的逐坐标长短期记忆(LSTM)元优化器进一步联合优化系数与预编码器,从数据中学习逐坐标的自适应更新规则。该元优化器离线训练后可直接应用于未见过的信道,无需进一步适配。数值结果表明:在N=32个阵元、K=4用户的条件下,PSA-GML在10 dB时达到11.06 bits/s/Hz的加权和速率,超出AO方法13.1%(即使经过多次随机重启仍超出6.2%),超出随机相位方案35.1%。在干扰受限场景下,无需人工超参数调优即可达到人工设计的Adam精调方法性能的83.9%,并且可零样本迁移至不同场景,这表明所学习到的更新规则能够捕捉WSR问题的内在结构特征。
cs.LG / 47 / 2609.29163

Edge AI on Constrained Devices for Binary Sleep-Wake Classification in Dynamic Environments

动态环境下基于资源受限设备的边缘AI二分类睡眠-清醒检测
Reitmann, Stefan, Oden, Lena
Abstract
This paper presents an Edge AI-based system for detecting sleep and wake states in non-stationary mobile environments using resource-constrained embedded hardware. Conventional approaches relying on accelerometer-based activity metrics are highly susceptible to motion and vibration artifacts and are limited by strict compute and energy budgets of wearable and IoT devices. To address these challenges, a multimodal pipeline is designed and implemented on an ESP32-S3 microcontroller. The system combines inertial sensing for head movement analysis and visual pose classification. A dual-core architecture with FreeRTOS enables parallel execution of real-time data acquisition and on-device inference. Sleep detection follows a two-stage strategy: low-movement detection over a temporal window, followed by visual validation of poses. Experimental results show accuracies of 96.5% for motion-based detection and 89% for pose classification, yielding robust binary sleep-wake classification. Field tests confirmed feasibility in representative mobile scenarios. The results demonstrate that privacy-preserving, local sleep detection is achievable on edge hardware through careful co-design, while highlighting limitations in sensing intrusiveness, dataset scale, and system integration.
Chinese Translation
本文提出了一种基于边缘AI(Edge AI)的系统,利用资源受限的嵌入式硬件在非平稳移动环境中检测睡眠与清醒状态。传统方法依赖于基于加速度计的活动指标,极易受到运动和振动伪影的影响,并且受限于可穿戴设备和物联网设备严格的计算与能耗预算。为应对这些挑战,本文设计并实现了在ESP32-S3微控制器上运行的多模态处理流程。该系统结合了用于头部运动分析的惯性传感与视觉姿态分类。基于FreeRTOS的双核架构实现了实时数据采集与设备端推理的并行执行。睡眠检测采用两阶段策略:首先在时间窗口内进行低运动检测,随后通过视觉对姿态进行验证。实验结果表明,基于运动的检测准确率为96.5%,姿态分类准确率为89%,从而实现了稳健的二分类睡眠-清醒判别。实地测试证实了该系统在代表性移动场景中的可行性。研究结果证明,通过精心的协同设计,在边缘硬件上实现保护隐私的本地睡眠检测是可行的,同时也指出了在传感侵入性、数据集规模和系统集成方面的局限性。
cs.LG / 48 / 2609.29216

FB-GDM: Fully-Bayesian Guided Diffusion Models for High-Dimensional Linear Inverse Problems via Unsupervised Variational Inference

FB-GDM:基于无监督变分推断的全贝叶斯引导扩散模型,用于高维线性逆问题
Séguy, Gatien, Rodet, Thomas
Abstract
Diffusion models are powerful priors for linear inverse problems, but the reference guidance methods, Diffusion Posterior Sampling (DPS) and Pseudoinverse-Guided Diffusion Models ($\Pi$GDM), rely on scalar hyperparameters tuned per task, usually against the ground truth. We introduce FB-GDM, a fully-Bayesian guided diffusion method that removes this calibration step. Starting from the Gaussian approximation of $\Pi$GDM, we derive a closed-form conditional score that depends on two precision parameters (inverse variances), one associated with the denoising approximation and one with the observation likelihood, and treat them as latent variables inferred by variational inference at each reverse step. A separable factorization makes each update scale linearly with the number of pixels, so the inference stays tractable at full image resolution, at a cost comparable to one $\Pi$GDM run. FB-GDM requires neither the noise level nor the ground truth: its only inputs are the observation and the forward operator. Experiments on CelebA-HQ inverse problems establish two results. (i) The precision parameters, inferred from the observation alone, allow FB-GDM to outperform $\Pi$GDM at its nominal setting, even when the latter is given the true noise level, by up to 14 dB depending on the operator, and to match the ground-truth-calibrated $\Pi$GDM oracle within 0.1 dB. (ii) FB-GDM is robust when the forward operator, the noise level, or the image distribution changes: it stays close to a per-problem $\Pi$GDM oracle throughout and does not exhibit the hallucinations observed with DPS, whereas DPS substantially degrades at a fixed scale and $\Pi$GDM stays competitive only if it is re-tuned against the ground truth for each new problem. When the prior is applied to images outside its training set, this re-balancing between data and prior keeps FB-GDM faithful where a fixed face-prior guidance can otherwise hallucinate.
Chinese Translation
扩散模型是线性逆问题的强大先验,但现有的引导方法,即扩散后验采样(Diffusion Posterior Sampling, DPS)和伪逆引导扩散模型(Pseudoinverse-Guided Diffusion Models, ΠGDM),依赖于针对每个任务调整的标量超参数,且通常需要借助真实值(ground truth)进行调整。我们提出了FB-GDM,一种无需该校准步骤的全贝叶斯引导扩散方法。从ΠGDM的高斯近似出发,我们推导了一个依赖于两个精度参数(方差的倒数)的闭式条件得分(conditional score):其中一个与去噪近似相关,另一个与观测似然相关,并将它们作为潜变量,在每一步反向过程中通过变分推断进行估计。利用可分离分解,每次更新的计算量与像素数呈线性关系,因此推断在全图像分辨率下依然可行,其开销与一次ΠGDM运行相当。FB-GDM既不需要噪声水平,也不需要真实值:其唯一输入是观测数据和前向算子。在CelebA-HQ逆问题上的实验确立了两个结论:(i) 仅从观测数据推断得到的精度参数,使FB-GDM在算子相关的条件下最多能以14 dB的优势超越处于名义设置的ΠGDM,即使后者已知真实噪声水平;并且与经过真实值校准的ΠGDM“神谕”(oracle)相差不超过0.1 dB。(ii) 当前向算子、噪声水平或图像分布发生变化时,FB-GDM表现出鲁棒性:其在整个过程中始终接近逐问题调优的ΠGDM神谕,且不会出现DPS所观察到的幻觉现象;而DPS在固定尺度下性能显著下降,ΠGDM则只有在针对每个新问题借助真实值重新调优的情况下才能保持竞争力。当先验应用于训练集之外的图像时,这种数据与先验之间的动态再平衡使FB-GDM保持忠实性,而固定人脸先验的引导则可能产生幻觉。
cs.LG / 49 / 2609.29268

BridgeMem: Causal Dyadic Transition Residuals for Temporal Knowledge Graph Forecasting

BridgeMem:基于因果二元转移残差的时间知识图谱预测
Li, Zeyan, Chen, Libing, Zhuo, Shengda, Tang, Yin, Xu, Jianfeng
Abstract
Temporal knowledge graph forecasting aims to infer future relational facts from the temporal structure of observed events. Existing forecasters mainly summarize history through entity states, relation states, paths, or exact recurrence. These views often miss pair-specific transition evidence, that is, the way prior relations between the query actor and a candidate change the odds of the target relation. We introduce BridgeMem, which estimates this quantity as a residual added to the log scores of a frozen full-vocabulary forecaster. For each candidate, BridgeMem retrieves the pair's events that strictly precede t, encodes their relations, directions, and lags, and converts them into a likelihood-ratio correction. A support-adaptive empirical-Bayes reader trusts exact transition counts where they are abundant and backs off to a learned attention estimator where they are sparse. The backbone's own uncertainty gates the correction, so confident queries and candidates without dyadic history are left unchanged. On five benchmarks, BridgeMem improves on the strongest of nine baselines from 2021--2026 in all 20 filtered MRR and Hits@{1,3,10} comparisons, with MRR gains of 0.0213, 0.0164, 0.0216, 0.0112, and 0.0028 over the best prior result. These results show the value of explicit dyadic transition modeling.
Chinese Translation
时间知识图谱预测旨在从已观测事件的时间结构中推断未来的关系事实。现有的预测方法主要通过实体状态、关系状态、路径或精确循环来概括历史。这些视角往往忽略了特定实体对(pair-specific)的转移证据,即查询主体与候选对象之间的先前关系如何改变目标关系的可能性。我们提出BridgeMem,将这一量作为残差,添加到冻结的全词表预测器的对数得分上。对于每个候选对象,BridgeMem检索严格发生在时刻t之前的该实体对事件,编码其关系、方向和时间滞后,并将其转换为似然比校正。一个支持度自适应的经验贝叶斯读取器在转移计数充足时信任精确计数,在计数稀疏时回退到学习到的注意力估计器。骨干模型自身的不确定性会对该校正进行门控,因此高置信度的查询以及没有二元历史的候选对象保持不变。在五个基准数据集上,BridgeMem在全部20个过滤后的MRR和Hits@{1,3,10}对比中均优于2021至2026年间九个基线方法中最强的方法,相比此前最佳结果的MRR提升分别为0.0213、0.0164、0.0216、0.0112和0.0028。这些结果表明了显式二元转移建模的价值。
cs.LG / 50 / 2609.29270

Learnable Time-Frequency Masks for Explaining Time-Series Classifiers

用于解释时间序列分类器的可学习时频掩码
Frehr, Theresa Dahl, Pelayo, Francisco, Raad, Lukas, Sanz, Alicia García, Brüsch, Thea, Alstrøm, Tommy Sonne
Abstract
Time-series explainability remains challenging because discriminative information is often encoded in latent frequency or time-frequency features rather than in the raw signal itself. Existing attribution methods typically operate either in the time domain or in a fixed transform domain, limiting their ability to capture salient information across different representations. We propose XACT, a general framework that learns sparse attribution masks over coefficients from arbitrary invertible time-frequency transforms. We evaluate the framework on the STFT, the continuous wavelet transform, and the discrete wavelet transform. In addition, we extend the virtual inspection layer approach from the STFT to both wavelet transforms, enabling LRP to generate explanations in these representations. On a synthetic dataset, XACT produces precise explanations and is less prone to highlighting spurious features than the tested baselines. Across two real-world datasets, XACT produces sparse and structured explanations, although no method performs best across all quantitative evaluation criteria. These results demonstrate that learning explanations directly in time-frequency representations offers a flexible approach to interpreting deep-learning models for time series data.
Chinese Translation
时间序列的可解释性仍然具有挑战性,因为判别性信息往往编码在潜在的频率或时频特征中,而非原始信号本身。现有的归因方法通常仅在时域或固定的变换域中运作,限制了其跨不同表示捕捉显著信息的能力。我们提出了XACT,一个通用框架,它可以在任意可逆时频变换的系数上学习稀疏归因掩码。我们在短时傅里叶变换(STFT)、连续小波变换和离散小波变换上对该框架进行了评估。此外,我们将虚拟检查层(virtual inspection layer)方法从STFT扩展到两种小波变换,使LRP能够在这些表示中生成解释。在一个合成数据集上,XACT能够产生精确的解释,并且比所测试的基线方法更不容易突出虚假特征。在两个真实世界数据集上,XACT产生了稀疏且结构化的解释,尽管没有任何方法在所有定量评估标准上都表现最佳。这些结果表明,直接在时频表示中学习解释,为解释时间序列数据的深度学习模型提供了一种灵活的方法。
cs.LG / 51 / 2609.29281

Online Task Adaptation via Self-Organisation

基于自组织的在线任务自适应
Proroković, Krsto
Abstract
Neural networks are typically adapted by computing gradients and updating model parameters. We investigate whether task-specific adaptation can instead emerge from a meta-learned self-organising process that requires no gradients at adaptation time. We instantiate this idea with a Neural Cellular Automaton in which locally interacting recurrent cells maintain both a recurrent state and a fast associative memory. During meta-training, backpropagation is used to learn the recurrent dynamics together with how the memory is read and written. Once training is complete, the slow model parameters remain fixed, and online adaptation occurs only through cellwise memory updates driven by local prediction errors and a delta rule. We evaluate whether the learned mechanism can adapt to semantically distinct held-out classification tasks. A single pass over the support data produces substantial improvements in held-out performance without gradient computation or parameter updates during adaptation, and the mechanism remains effective across large changes in the number of examples processed jointly. These results show that task-specific adaptation can be achieved through explicit fast-memory updates while keeping the slow model parameters fixed.
Chinese Translation
神经网络通常通过计算梯度并更新模型参数来实现自适应。我们研究任务特定的自适应是否可以转而由一种元学习的自组织过程产生,该过程在自适应阶段完全不需要梯度。我们用神经细胞自动机(Neural Cellular Automaton)来实现这一想法,其中局部相互作用的循环细胞同时维护一个循环状态和一个快速联想记忆。在元训练期间,利用反向传播来学习循环动力学以及记忆的读写方式。训练完成后,慢速模型参数保持固定,在线自适应仅通过由局部预测误差和delta规则驱动的逐细胞记忆更新来实现。我们评估了所习得的机制能否适应语义上不同的留出分类任务。对支持数据进行一次遍历即可显著提升留出任务上的性能,且在自适应过程中无需梯度计算或参数更新;同时,该机制在联合处理的样本数量发生大幅变化时仍然有效。这些结果表明,在保持慢速模型参数固定的情况下,可以通过显式的快速记忆更新实现任务特定的自适应。
cs.LG / 52 / 2609.29307

Beyond Feature Reliability: Repeat-Informed Multifractal Curve Regression for Brain-Age Prediction

超越特征可靠性:面向脑龄预测的重复扫描感知的多重分形曲线回归
Chang, Yu, Cheng, Anzhe, Chen, Jiahao, Ping, Heng, Zhang, Peiyu, Pan, Puquan, Chattopadhyay, Tamoghna, Thomopoulos, Sophia, Nazarian, Shahin, Thompson, Paul, Bogdan, Paul
Abstract
Brain-age prediction from resting-state fMRI provides a quantitative framework for characterizing age-related changes in spontaneous brain dynamics and for identifying functional signatures. Existing studies have linked fractal and multifractal scaling to age and examined the reliability of individual features. However, prediction repeatability depends on how features fluctuate jointly and how a predictor combines them, which feature-wise reliability assessments do not capture. To address this problem, we propose Repeat-informed Multifractal Curve Regression (RMCR), a structured framework for learning stable age-predictive patterns from multifractal curves. By jointly modeling curve structure and repeat-scan variability, RMCR learns predictive combinations of fluctuation orders that target both accuracy and within-subject consistency. Relative to a matched run-level ridge baseline, RMCR reduces single-run MAE by 6.1% on HCP-A and 7.9% on an external Cam-CAN cohort, and within-visit repeat absolute difference by 18.5% on HCP-A, using a single scan at inference.
Chinese Translation
基于静息态功能磁共振成像的脑龄预测为刻画自发脑动力学的年龄相关变化以及识别功能性标志提供了一个定量框架。已有研究将分形与多重分形标度特性与年龄联系起来,并考察了单个特征的可靠性。然而,预测的可重复性取决于特征如何共同波动以及预测器如何组合这些特征,这是逐特征可靠性评估所无法刻画的。为解决这一问题,我们提出了重复扫描感知的多重分形曲线回归(Repeat-informed Multifractal Curve Regression, RMCR),这是一个从多重分形曲线中学习稳定年龄预测模式的结构化框架。通过联合建模曲线结构与重复扫描变异性,RMCR 学习同时面向准确性和被试内一致性的波动阶数的预测性组合。相对于匹配的运行水平岭回归基线,RMCR 在仅使用单次扫描进行推理的情况下,将 HCP-A 数据集上的单次运行 MAE 降低了 6.1%,在外部 Cam-CAN 队列上降低了 7.9%,并将 HCP-A 上同一访次内重复扫描的绝对差异降低了 18.5%。
cs.LG / 53 / 2609.29317

Neuralized Multi-Wavelet Decomposition for Time Series Classification and Forecasting

面向时间序列分类与预测的神经化多小波分解
Jiang, Xiaohan, Wang, Jingyuan, Ji, Jiahao, Wang, Yongyao, Yang, Chen, Wu, Junjie
Abstract
Time series analysis is fundamental in domains such as finance, healthcare, and meteorology. Real-world time series often exhibit multiscale characteristics shaped by diverse latent factors, resulting in intricate temporal patterns and rich frequency structures. However, existing approaches typically focus on either frequency-domain decomposition or time-domain pattern extraction in isolation, neglecting their joint structure. This decoupled modeling limits representation expressiveness and undermines performance in tasks requiring simultaneous temporal and spectral reasoning. To address this gap, we propose m-WCN, a novel end-to-end deep learning framework that neuralizes multi-wavelet decomposition for joint extraction of temporal patterns and frequency components. By approximating the classical GHM multi-wavelet transform with trainable convolutional operators and enforcing orthogonality constraints, m-WCN produces interpretable multi-resolution representations. Built on this foundation, we introduce two task-specific architectures: TFBC for time series classification, which boosts discriminative features across frequency scales, and FTB for forecasting, which ensembles frequency-aware predictors. Extensive experiments on 64 UCR datasets and seven public forecasting benchmarks demonstrate the effectiveness of our approach. Built on the neuralized m-WCN, our TFBC and FTB outperform various baseline models across diverse datasets, achieving average improvements of 19.97% in classification and 19.92% in forecasting tasks.
Chinese Translation
时间序列分析在金融、医疗和气象等领域具有重要基础作用。现实世界中的时间序列通常呈现出由多种潜在因素塑造的多尺度特性,产生复杂的时间模式和丰富的频率结构。然而,现有方法通常孤立地关注频域分解或时域模式提取,忽视了二者的联合结构。这种解耦建模限制了表示的表达能力,损害了在需要同时进行时域和频域推理的任务中的性能。为填补这一空白,我们提出m-WCN,一种新颖的端到端深度学习框架,它将多小波分解神经化,以联合提取时间模式和频率成分。通过使用可训练的卷积算子逼近经典GHM多小波变换并施加正交性约束,m-WCN产生可解释的多分辨率表示。在此基础上,我们引入两种面向特定任务的架构:用于时间序列分类的TFBC,它跨频率尺度增强判别性特征;以及用于预测的FTB,它集成频率感知的预测器。在64个UCR数据集和七个公开预测基准上的大量实验证明了我们方法的有效性。基于神经化的m-WCN,我们的TFBC和FTB在多种数据集上超越了各类基线模型,在分类任务和预测任务中分别实现了19.97%和19.92%的平均提升。
cs.LG / 54 / 2609.29322

TinyCardioUNet: IMU-to-ECG Translation with Graph-Encoded Inter-Axis Dependencies and Tensor Decomposition-Based Parameter Reduction

TinyCardioUNet:基于图编码轴间依赖关系与张量分解参数压缩的IMU到ECG转换
Han, Seungwoo, Chanpornpakdi, Ingon, Noda, Motoi, Leelasiri, Puwadej, Hiruma, Ibuki, Tanaka, Toshihisa
Abstract
Estimating electrocardiography (ECG) from a chest-worn inertial measurement unit (IMU) enables continuous heart rate (HR) monitoring without the discomfort of electrodes. We propose TinyCardioUNet, a lightweight UNet that uses all six IMU axes without prior channel selection, refines its bottleneck with a graph neural network that encodes inter-axis dependencies, and employs tensor decomposition with automatic variational Bayesian rank selection for parameter reduction. On a public dataset, TinyCardioUNet achieves an RMSE of $0.098$ and a Pearson correlation coefficient of $0.677$ with only $36.0$k parameters and remains comparatively robust to additive noise, demonstrating accurate ECG reconstruction with a compact model.
Chinese Translation
利用佩戴于胸部的惯性测量单元(IMU)估计心电图(ECG),可以在不带来电极不适的情况下实现连续心率(HR)监测。我们提出了TinyCardioUNet,一种轻量级UNet模型,它无需先验通道选择即可使用IMU的全部六个轴,通过编码轴间依赖关系的图神经网络(GNN)对其瓶颈部分进行精炼,并采用具有自动变分贝叶斯秩选择的张量分解来实现参数压缩。在一个公开数据集上,TinyCardioUNet仅使用36.0k参数即可达到0.098的均方根误差(RMSE)和0.677的皮尔逊相关系数,并且在加性噪声干扰下保持相对稳健,展示了紧凑模型实现精确ECG重建的能力。
cs.LG / 55 / 2609.29330

FlowAtom: Atom-Based Evidence Aggregation for Multi-Label Website Fingerprinting

FlowAtom:面向多标签网站指纹识别的基于原子的证据聚合方法
Fan, Chongru, Huang, Wentao, Wang, Wei, Ding, Zhenquan, Shi, Jinqiao, Cai, Wei, Hao, Zhiyu
Abstract
Identifying the set of monitored websites in mixed encrypted traffic is challenging because an individual flow often provides only partial evidence of website identity. To address this challenge, we propose FlowAtom, which constructs shared prototypes, called Atoms, from flow representations without website labels. Specifically, FlowAtom pretrains a flow encoder on external unlabeled traffic and aggregates Atom responses across flows within each observation window into a fixed-dimensional, permutation-invariant representation for monitored website-set prediction. Across Direct HTTPS, Trojan, and VMess, FlowAtom achieves micro-F1 scores of 97.82%, 94.43%, and 93.92% in closed-world evaluation, respectively, and consistently outperforms the evaluated baselines in open-world evaluation on windows containing monitored visits. The code is available at https://github.com/aimafan123/FlowAtom.
Chinese Translation
在混合加密流量中识别受监控网站的集合是一项具有挑战性的任务,因为单个流通常只能提供关于网站身份的部分证据。为应对这一挑战,我们提出了FlowAtom,它从无网站标签的流表示中构建共享原型(称为原子,Atoms)。具体而言,FlowAtom在外部无标签流量上预训练一个流编码器,并将每个观测窗口内各流的原子响应聚合为固定维度的、置换不变的表示,用于受监控网站集合的预测。在Direct HTTPS、Trojan和VMess三种场景下,FlowAtom在封闭世界评估中分别取得了97.82%、94.43%和93.92%的微平均F1分数(micro-F1),并且在包含受监控访问的窗口的开放世界评估中持续优于所评估的基线方法。代码可在 https://github.com/aimafan123/FlowAtom 获取。
cs.LG / 56 / 2609.29379

On the second-order optimization for spiking neural networks

关于脉冲神经网络的二阶优化方法
Doan, Ngoc Phu, Alouani, Ihsen
Abstract
Spiking Neural Networks (SNNs) offer an energy-efficient alternative to conventional neural networks by exploiting sparse, binary spikes, and event-driven computation. However, the training of SNNs remains challenging, as spiking activations create a sharp loss landscape that hinders training, and diagonal-curvature optimizers such as the Adam family may fail to capture this geometry. The extension of curvature-based optimization methods to SNNs is further complicated by the sparse, discrete, and temporally recurrent nature of their underlying dynamics. To address these limitations, we propose SpiKFAX, a second-order optimization method that formulates a computationally tractable, Kronecker-factored approximation of the Fisher information matrix specifically adapted to the structure of SNNs. Empirical evaluation across five architectures and seven datasets demonstrates that SpiKFAX consistently yields improvements in test accuracy and training stability relative to other popular optimizers.
Chinese Translation
脉冲神经网络(Spiking Neural Networks, SNNs)通过利用稀疏的二值脉冲和事件驱动计算,为传统神经网络提供了一种高能效的替代方案。然而,SNN的训练仍然具有挑战性:脉冲激活会形成尖锐的损失曲面,从而阻碍训练,而诸如Adam系列等仅考虑对角曲率的优化器可能无法捕捉这种几何特性。此外,由于SNN底层动态具有稀疏性、离散性和时间递归性,将基于曲率的优化方法扩展到SNN也变得更加复杂。为解决这些局限性,我们提出了SpiKFAX,这是一种二阶优化方法,针对SNN的结构特点,构建了一种计算上可行的Kronecker分解的Fisher信息矩阵近似。在五种架构和七个数据集上的实证评估表明,与其他流行优化器相比,SpiKFAX在测试准确率和训练稳定性方面始终带来提升。
cs.LG / 57 / 2609.29383

Lightweight Probabilistic Downscaling from a Deterministic Base Model

基于确定性基础模型的轻量级概率降尺度方法
McLean, Joseph, Vlaar, Tiffany, Hellan, Sigrid Passano, Ericsson, Linus
Abstract
Climate data downscaling is the task of increasing the spatial resolution of climate data, typically by generating fine-resolution regional climate data from coarse global model output. Recent machine learning (ML) work in the related task of weather forecasting has seen significant improvements due to newly devised training methods and architectural components, but these have not yet benefited downscaling. We adapt two of these methods to create a family of lightweight probabilistic ML downscaling models built on a modified U-Net backbone and evaluate them on the CORDEX-ML-Bench suite for daily maximum temperature and precipitation across three geographic regions: the Alps, New Zealand and South Africa. We find that a two-stage training curriculum, combining deterministic pretraining with probabilistic tuning, transfers well to downscaling, beating the state-of-the-art for RMSE. Our work provides an advancement towards lightweight, probabilistic downscaling models, reducing the current trade-off between computational intensity and distributional fit.
Chinese Translation
气候数据降尺度的任务是提高气候数据的空间分辨率,通常通过从粗分辨率的全球模式输出生成高分辨率的区域气候数据来实现。近期在天气预报这一相关任务中的机器学习(ML)研究,由于新设计的训练方法和网络架构组件而取得了显著进展,但这些进展尚未惠及降尺度任务。我们将其中的两种方法加以改进,构建了一系列基于改进U-Net骨干网络的轻量级概率机器学习降尺度模型,并在CORDEX-ML-Bench测试套件上,针对阿尔卑斯山、新西兰和南非三个地理区域的日最高气温和降水进行评估。我们发现,将确定性预训练与概率微调相结合的两阶段训练课程能够很好地迁移到降尺度任务中,在RMSE指标上超越了现有最先进方法。我们的工作为构建轻量级概率降尺度模型提供了进展,减小了当前在计算强度与分布拟合之间的权衡。
cs.LG / 58 / 2609.29386

MORE-PLR: multi-output regression employed for partial label ranking

MORE-PLR:用于偏标签排序的多输出回归方法
Thies, Santo M. A. R., Alfaro, Juan C., Bengs, Viktor
Abstract
The partial label ranking problem is a supervised learning scenario that aims to fit a preference model that predicts a bucket order defined over a set of labels for a given input instance. This problem generalizes the well-known label ranking problem, which, in practice, is limited to outputting total orders of labels. Existing partial label ranking methods have primarily extended label ranking approaches to handle ties in predictions. This paper proposes using multi-output regression to address the partial label ranking problem, introducing an encoder that, during the learning phase, transforms the (possibly incomplete) rankings with ties of labels to multivariate regression targets, an underexplored perspective in both label ranking and partial label ranking. Moreover, during the inference phase, we introduce several post-hoc layers that convert the multi-output regression results into the output bucket order to effectively implement this approach. This framework provides learning strategies that are competitive with the current state-of-the-art partial label ranking methods, as demonstrated through experimental evaluations.
Chinese Translation
偏标签排序(partial label ranking)问题是一种监督学习场景,旨在拟合一个偏好模型,对给定输入实例预测定义在标签集合上的桶序(bucket order)。该问题是著名的标签排序(label ranking)问题的推广,而后者在实际应用中仅限于输出标签的全序。现有的偏标签排序方法主要是对标签排序方法进行扩展,以处理预测中的并列关系。本文提出使用多输出回归来解决偏标签排序问题,引入一种编码器,在学习阶段将带有并列关系的(可能不完整的)标签排序转换为多元回归目标——这一视角在标签排序和偏标签排序领域均少有探索。此外,在推理阶段,我们引入若干事后(post-hoc)层,将多输出回归结果转换为输出的桶序,从而有效实现该方法。实验评估表明,该框架所提供的学习策略与当前最先进的偏标签排序方法具有竞争力。
cs.LG / 59 / 2609.29398

ICE: Task-Aligned Clifford Latent Fields for Multimodal Graph Foundation Models

ICE:面向多模态图基础模型的任务对齐克利福德潜场
Li, Xunkai, Wang, Xu, Zhu, Yinlin, Yongfu, Xiong, Liu, Yi, Li, Rong-Hua, Wang, Guoren
Abstract
Multimodal attributed graphs connect entities, visual content, language, and observed relations. Learning one foundation across such graphs requires more than compressing each node into a fused Euclidean vector. The representation must preserve entity semantics, construct interaction state from graph neighborhoods, and expose that state to prediction units with different geometry. Our empirical study shows why these requirements are inseparable. Higher-grade channels recover pair relations across the foundation graphs, specialized queries reveal information hidden by a generic readout, and rigid blade isolation removes cross-grade capacity. We therefore introduce ICE (Interaction-aware Clifford Encoder), a multimodal graph foundation model built on a node-indexed Clifford latent field. Topology, text, and images enter explicit Cl(3) addresses. Edge-aware geometric products transform these directions into scalar, bivector, and trivector relations over observed neighborhoods. A protected Grade-1 route preserves entity semantics, while the full grade and depth bank remains available to fresh node and link heads. We establish exact cross-grade reachability, node-permutation equivariance, and a bound on the task residual around the semantic score. Experiments span one shared foundation over eleven graphs, six node-classification datasets, three link-prediction datasets, and matched few-shot tasks. ICE ranks first in all 30 reported supervised and few-shot comparisons. Core removals reduce every task summary, and mechanism controls connect the gains to higher-order transport, retained multidepth structure, semantic protection, and direct field access.
Chinese Translation
多模态属性图将实体、视觉内容、语言以及观测到的关系连接在一起。要在这些图上学习统一的基础模型,仅仅将每个节点压缩为一个融合的欧几里得向量是不够的。表征必须保留实体语义,从图邻域构建交互状态,并将该状态暴露给具有不同几何结构的预测单元。我们的实证研究表明,这些需求为何不可分割:高阶通道能够恢复基础图上的配对关系,专用查询能够揭示通用读出所隐藏的信息,而僵化的blade隔离则会移除跨阶的容量。因此,我们提出了ICE(交互感知克利福德编码器,Interaction-aware Clifford Encoder),这是一个建立在节点索引克利福德潜场(Clifford latent field)之上的多模态图基础模型。拓扑、文本和图像进入显式的Cl(3)地址;边感知的几何积将这些方向转化为观测邻域上的标量、二重向量和三重向量关系。一条受保护的Grade-1路由保留实体语义,而完整的阶数与深度库仍可供新的节点与链接预测头使用。我们建立了精确的跨阶可达性、节点置换等变性,以及围绕语义得分的任务残差上界。实验涵盖一个共享基础模型在十一个图上的应用、六个节点分类数据集、三个链接预测数据集以及匹配的小样本任务。ICE在全部30项有监督和小样本对比中均排名第一。核心组件的消融降低了所有任务指标,机制对照实验将性能提升与高阶信息传递、保留的多深度结构、语义保护以及直接的场访问联系起来。
cs.LG / 60 / 2609.29413

Neural Transport Nested Sampling

神经传输嵌套采样
Yallup, David, Handley, Will
Abstract
Sampling from Boltzmann distributions of molecular systems is an inference problem that has seen significant recent developments fuelled by advances in neural density estimation. We develop a novel sampling algorithm, Neural Transport Nested Sampling (NTNS), which combines the classical strengths of nested sampling with modern neural flow-based methods. NTNS uses a flow matching velocity as the drift in a Metropolis--Hastings corrected Langevin kernel inside a nested sampling outer loop, requiring only evaluations of the target energy function and providing scalable estimation of the full partition function of high-dimensional particle systems. We benchmark NTNS on challenging molecular sampling benchmarks, scaling up to Lennard--Jones clusters of 55 interacting particles, where it reduces both interatomic distance and energy Wasserstein errors to reference MCMC by over an order of magnitude relative to the strongest neural baselines at lower wall-clock cost. To our knowledge, NTNS is also the first neural sampler to return a calibrated, temperature resolved partition function estimate at this scale, recovering the phase structure across temperature from a single run.
Chinese Translation
从分子系统的玻尔兹曼分布中采样是一个推断问题,近年来在神经密度估计进展的推动下取得了显著发展。我们提出了一种新颖的采样算法——神经传输嵌套采样(Neural Transport Nested Sampling, NTNS),它将嵌套采样的经典优势与现代基于神经流的方法相结合。NTNS 在嵌套采样外循环中,使用流匹配速度作为 Metropolis–Hastings 校正朗之万核中的漂移项,仅需对目标能量函数进行评估,并能对高维粒子系统的完整配分函数进行可扩展的估计。我们在具有挑战性的分子采样基准上对 NTNS 进行了测试,扩展至含 55 个相互作用粒子的 Lennard–Jones 团簇。相对于最强的神经基线方法,NTNS 在更低的实际运行时间成本下,将相对于参考 MCMC 的原子间距离和能量的 Wasserstein 误差降低了一个数量级以上。据我们所知,NTNS 也是首个在该尺度上能够返回经校准、随温度分辨的配分函数估计的神经采样器,仅通过单次运行即可恢复跨温度范围的相结构。
cs.LG / 61 / 2609.29446

SPADE-DFL: Communication-Efficient Decentralized Federated Learning via Derivative-Free Linearized ADMM

SPADE-DFL:基于无导数线性化ADMM的高效通信去中心化联邦学习
Wei, Mengli, Zhu, Mengkai, Chen, Jiawen, Yu, Wenwu, Che, Duxin
Abstract
Reducing communication in derivative-free decentralized learning requires controlling the disagreement accumulated over multiple local updates. This paper develops SPADE-DFL, a primal--dual method that allows the number of local function-value updates between neighbor exchanges to grow with the computation budget while preserving the nonprivate convergence order. For smooth nonconvex objectives under uniform query-moment bounds, the prescribed nonprivate schedule achieves a time-averaged stationarity and consensus bound of $\mathcal{O}(T^{-1/3})$ using only $\Theta(T^{2/3})$ communication rounds, where $T$ is the number of local updates per client. For private training, the accumulated data-dependent increment is isolated from the graph correction, allowing one protected state per client and round to generate all outgoing messages. We prove client-level differential privacy for the full interactive transcript and quantify the resulting optimization error over a finite horizon. Experiments on four classification tasks show that SPADE-DFL achieves higher mean test accuracy than existing decentralized learning methods.
Chinese Translation
在无导数去中心化学习中减少通信量,需要控制多次本地更新所累积的分歧(disagreement)。本文提出了SPADE-DFL,这是一种原始-对偶(primal–dual)方法,允许邻居交换之间的本地函数值更新次数随计算预算增长,同时保持非隐私设定的收敛阶。对于在均匀查询矩界约束下的光滑非凸目标函数,所规定的非隐私调度方案仅使用 $\Theta(T^{2/3})$ 轮通信即可达到 $\mathcal{O}(T^{-1/3})$ 的时间平均平稳性与一致性界,其中 $T$ 为每个客户端的本地更新次数。对于隐私训练,将累积的数据相关增量与图校正相隔离,使得每个客户端每轮只需一个受保护状态即可生成所有外发消息。我们针对完整的交互过程证明了客户端级差分隐私,并量化了有限时间范围内由此产生的优化误差。在四个分类任务上的实验表明,SPADE-DFL比现有去中心化学习方法获得了更高的平均测试准确率。
cs.LG / 62 / 2609.29453

Decoupled Learning and Selection in Slate Recommendation for Privacy and Stability Under Noisy Scores

面向隐私与稳定性的组合推荐中学习与选择的解耦:噪声评分下的研究
Urmian, Sam, Liu, Qinyi, Khalil, Mohammad
Abstract
We formalize slate recommendation as a randomized score learner followed by deterministic selection. First, an appropriately scoped differential-privacy guarantee passes through selection and its audit trace by post-processing. End-to-end privacy holds only when selector inputs are public or independent, previous private outputs, or separately privacy-accounted; fixing raw state or candidate information instead yields only a conditional guarantee. Second, we derive a logged margin certificate: bounded score-induced objective movement below half the smallest greedy decision margin guarantees that the ordered slate is unchanged. Controlled fixed-margin tests show near-linear exponent scaling, with an empirical slope of $-0.220$ (95% CI $[-0.231,-0.210]$) against the independent-noise reference $-1/4$. Real-anchor experiments on OULAD, MovieLens-25M, and Amazon Musical Instruments show that greater anchor weight reduces score-noise-induced ranking churn. OULAD and EdNet certificate checks validate the implementation of the logged inequality, while closed-loop simulations show bounded target drift and setting-dependent downstream utility. The contribution is therefore a privacy-scope contract and a certifiable score-to-slate stability mechanism, not a universal utility claim.
Chinese Translation
我们将组合推荐形式化为一个随机评分学习器之后接确定性选择的过程。首先,适当界定范围的差分隐私保证可通过后处理传递至选择环节及其审计轨迹。只有当选择器的输入为公开信息或独立的先前私有输出,或已单独进行隐私核算时,端到端隐私才成立;若直接固定原始状态或候选信息,则仅能得到条件性保证。其次,我们推导出一种基于日志的边际证书(logged margin certificate):当评分引起的目标函数变化幅度小于最小贪心决策边际的一半时,可保证有序组合保持不变。固定边际的受控测试显示指数接近线性缩放,其经验斜率为 -0.220(95% 置信区间 [-0.231, -0.210]),与独立噪声的理论参考值 -1/4 相符。在 OULAD、MovieLens-25M 和 Amazon 乐器数据集上的真实锚点实验表明,更大的锚点权重可降低由评分噪声引起的排序变动。OULAD 和 EdNet 上的证书检验验证了基于日志的不等式的实现,而闭环仿真则展示了有界的目标漂移以及依赖具体设置的下游效用。因此,本文的贡献在于一个隐私范围契约和一个可认证的从评分到组合的稳定性机制,而非普遍的效用提升声明。
cs.LG / 63 / 2609.29458

Precise Convergence Speed of Clipped SGD

裁剪SGD的精确收敛速度
Robin, David A. R.
Abstract
We present a tightened convergence analysis of clipped gradient descent on $(L_0, L_1)$-smooth functions, with quantitative constants. Building on the ideas of Koloskova et al (2023), we refactor several case disjunctions to reveal the central role of a control of the bias derived from fundamental properties of $\ell_2$-projection, simplifying proofs. We also extend the domain of validity from $\eta \leq 1 / (9 \beta)$ to $\eta < 1 /\beta$ where $\beta = L_0 + c L_1$ for clipping constant $c$, which matches the more traditional analysis of smooth functions. We strengthen the convergence criterion from $\left( \min_{t < T} \mathbb{E}[\lVert \nabla f(x_t) \rVert_2] \right)$ to $\left( \frac{1}{T} \sum_{t < T} \mathbb{E}[\lVert \nabla f(x_t) \rVert_2] \right)$ with matching speed, and lower the final achievable loss from $\mathcal{O}(\min(\sigma^2/c, \sigma))$ to the more precise $6 \min(\sigma^2 /c, 3 \sigma)$.
Chinese Translation
我们针对 $(L_0, L_1)$-光滑函数上的裁剪梯度下降给出了具有定量常数的更紧致的收敛性分析。基于 Koloskova 等人(2023)的思想,我们重构了若干情形划分,以揭示由 $\ell_2$-投影基本性质导出的偏差控制所起的核心作用,从而简化了证明。我们还将有效性范围从 $\eta \leq 1 / (9 eta)$ 扩展到 $\eta < 1 /eta$,其中 $eta = L_0 + c L_1$,$c$ 为裁剪常数,这与更传统的光滑函数分析相匹配。我们将收敛准则从 $\left( \min_{t < T} \mathbb{E}[\lVert abla f(x_t) Vert_2] ight)$ 加强为 $\left( \frac{1}{T} \sum_{t < T} \mathbb{E}[\lVert abla f(x_t) Vert_2] ight)$,并保持了相同的收敛速度;同时将最终可达到的损失从 $\mathcal{O}(\min(\sigma^2/c, \sigma))$ 降低到更精确的 $6 \min(\sigma^2 /c, 3 \sigma)$。
cs.LG / 64 / 2609.29466

Direct Message Approximation (DMA): A Consistency-Based Framework for Tractable Approximate Inference on Factor Graphs

直接消息近似(DMA):一个基于一致性的因子图可处理近似推断框架
Herbrich, Ralf, Schlosser, Rainer, Lemcke, Jan, Ukrow, Johann, Kazachkova, Anna, Alder, Nicolas, Hennicke, Leonhard, Bardey, Theo, Grimm, Nico, Kleinschmidt, Luca, Kolbe, Philipp, Kujath, Cezary, Schlimme, Johanna, Schütz, Karl Matti
Abstract
Approximate message passing on factor graphs underlies two dominant families of probabilistic inference algorithms: expectation propagation (EP) and variational message passing (VMP). Both methods approximate the marginal at each factor edge, forcing an iterative round-robin schedule, risking negative-precision messages, and, for VMP, collapsing to point estimates at Dirac-delta factors. We introduce Direct Message Approximation (DMA), which approximates factor-to-variable messages directly rather than the marginal. For normalisable factors, we define a consistency condition (requiring exactness when all other incoming messages are Dirac deltas) to guide message construction. We prove a master theorem (proper messages, any graph) bounding marginal KL from message KL, with three structural corollaries: Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages. Further, we prove a complementary $O(1/r^2)$ guarantee for the inherently improper backward message of the product factor, whose closed-form treatment has resisted prior work. As a concrete instantiation, we derive explicit DMA messages for the product and leaky-ReLU factors and assemble a Bayesian neural network (BNN) inference algorithm with one forward/backward sweep per training example and no gradient learning-rate hyperparameter, validating that the structural guarantees translate to predictive uncertainty that widens in data-sparse regions, including under model mismatch.
Chinese Translation
因子图上的近似消息传递构成了两大主流概率推断算法家族的基础:期望传播(EP)和变分消息传递(VMP)。这两种方法都对每个因子边的边缘分布进行近似,因而被迫采用迭代的轮转调度,存在产生负精度消息的风险,并且对于VMP而言,在狄拉克δ因子处会退化为点估计。我们提出了直接消息近似(Direct Message Approximation, DMA),它直接近似因子到变量的消息,而非边缘分布。对于可归一化的因子,我们定义了一个一致性条件(要求当所有其他入站消息均为狄拉克δ时推断精确),用以指导消息的构造。我们证明了一个主定理(正规消息,任意图),该定理用消息的KL散度约束边缘分布的KL散度,并由此得出三个结构性推论:狄拉克输入一致性、无需EP式的内层循环迭代、以及不会产生负精度消息。此外,我们针对乘积因子天然非正规的反向消息,证明了一个互补的$O(1/r^2)$保证,其闭式处理此前的工作一直未能解决。作为具体实例,我们推导了乘积因子和泄漏ReLU(leaky-ReLU)因子的显式DMA消息,并构建了一个贝叶斯神经网络(BNN)推断算法,该算法对每个训练样本仅需一次前向/反向扫描,且无需梯度学习率超参数,验证了这些结构性保证能够转化为在数据稀疏区域(包括模型失配情形下)变宽的预测不确定性。
cs.LG / 65 / 2609.29487

BLADE: Distilled LLM Regularization for Calibrated Knowledge Graph Completion

BLADE:用于校准知识图谱补全的蒸馏大语言模型正则化方法
Shihab, Ibne Farabi, Tamanna, Rabeya Bosri, Karaky, Abdo El, Akter, Sanjeda, Sharma, Anuj
Abstract
Knowledge graph completion models optimize ranking, although many downstream applications require calibrated probabilities. We present BLADE, a variational model that separates latent truth from graph recording and distills offline language-model judgments into a frozen teacher regularizer. The LLM is absent during inference. Posterior samples provide predictive probabilities and epistemic uncertainty, while the compact teacher remains available only as an optional triage factor. Across five benchmarks, BLADE remains competitive under a common ranking protocol and reduces adaptive ECE by a macro-average of 60.1% relative to deep ensembles and 78.1% relative to temperature-scaled RotatE. On identical FB15k-237 candidate sets, BLADE also improves ECE, Brier score, and NLL over validation-selected histogram binning and a matched generative ComplEx2 model, with these improvements persisting on a prespecified near-miss pool. Under controlled injected missingness, the full triage score achieves a mean AUC-PR of 0.863, compared with 0.805 for its strongest non-teacher variant. Leakage stress tests show that aligned semantics matter, but they cannot exclude knowledge acquired during LLM pretraining. We therefore claim calibration only for the declared candidate distributions, not for all unobserved triples.
Chinese Translation
知识图谱补全模型通常以排序为目标进行优化,但许多下游应用需要经过校准的概率。我们提出了 BLADE,这是一种变分模型,它将潜在真值与图谱记录相分离,并将离线的大语言模型判断蒸馏到一个冻结的教师正则化器中。该大语言模型在推理阶段并不参与。后验样本可提供预测概率和认知不确定性,而紧凑的教师模型仅作为可选的分诊因子保留。在五个基准数据集上,BLADE 在统一的排序协议下保持竞争力,其自适应 ECE 相对于深度集成(deep ensembles)平均降低了 60.1%,相对于温度缩放的 RotatE 降低了 78.1%。在相同的 FB15k-237 候选集上,BLADE 相对于经验证集选择的直方图分箱方法以及参数量相当的生成式 ComplEx2 模型,在 ECE、Brier 分数和 NLL 上均有提升,且这些提升在预先设定的近似命中(near-miss)候选池中依然保持。在受控注入缺失数据的实验中,完整的分诊分数达到 0.863 的平均 AUC-PR,而其最强的非教师变体仅为 0.805。泄漏压力测试表明,语义对齐固然重要,但无法排除大语言模型在预训练期间获得的知识。因此,我们仅对已声明的候选分布声称校准有效性,而非对所有未观测三元组均有效。
cs.LG / 66 / 2609.29499

Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference

任务感知的谱剪枝:一种面向高效大语言模型推理的混合掩码框架
Shihab, Ibne Farabi, Afrin, Fariya, Akter, Sanjeda, Sharma, Anuj
Abstract
Static pruning imposes one sparse structure on every prompt, even though reasoning, retrieval, generation, coding, and translation can depend on different parts of a language model. We introduce Task-Aware Spectral Pruning (TASP), a post-training framework that calibrates module-level spectral descriptors against measured task-specific ablation effects, closes grouped-query-attention and SwiGLU dependencies during sparse-mask construction, and routes each user turn to one compiled mask that remains fixed throughout prefill and decoding. A module-disjoint pilot first determines whether the spectral signal is informative before full calibration. Under the stated retrospective operating rule, the pilot passes on the evaluated Llama-3-8B and Llama-3-70B checkpoints but rejects Qwen2.5-1.5B, demonstrating that applicability is model-dependent rather than universal. At a 43% active-FLOP reduction, the Llama-3-70B benchmark harness retains 97.7 +/- 0.2% of the dense BF16 score. In the deployment-matched INT8-weight/BF16-compute runtime on a single A100 80GB, the compiled sparse path retains 97.3 +/- 0.2% relative to dense BF16 and reduces decode latency from 45.2 +/- 0.4 to 31.3 +/- 0.4 ms/token, yielding a 1.44x speedup. Factorized ablations, disjoint-module tests, compiled structured baselines, routing-corruption studies, and an explicit 136-GPU-hour calibration audit further delimit the source and operating regime of these gains
Chinese Translation
静态剪枝对所有提示词施加同一种稀疏结构,然而推理、检索、生成、编程和翻译等任务可能依赖语言模型的不同部分。我们提出任务感知谱剪枝(Task-Aware Spectral Pruning,TASP),这是一个训练后框架,它通过测量的任务特定消融效应来校准模块级谱描述符,在稀疏掩码构建过程中闭合分组查询注意力(grouped-query-attention)与SwiGLU之间的依赖关系,并将每个用户轮次路由到一个在预填充和解码全过程中保持固定的已编译掩码。一个模块不相交的预试验首先判断谱信号是否具有信息量,然后才进行完整校准。在所述回溯性运行规则下,该预试验在被评估的Llama-3-8B和Llama-3-70B检查点上通过,但拒绝Qwen2.5-1.5B,这表明其适用性依赖于具体模型而非普遍适用。在43%有效FLOP削减下,Llama-3-70B基准测试框架保留了稠密BF16分数的97.7 ± 0.2%。在单张A100 80GB上与部署场景匹配的INT8权重/BF16计算运行时中,编译后的稀疏路径相对于稠密BF16保留了97.3 ± 0.2%,并将解码延迟从45.2 ± 0.4 ms/token降至31.3 ± 0.4 ms/token,实现1.44倍加速。因子化消融实验、不相交模块测试、编译结构化基线、路由破坏研究以及明确的136 GPU小时校准审计进一步界定了这些增益的来源与运行条件。
cs.LG / 67 / 2609.29505

Spectral-Guided Diffusion: Accelerating Inference via Static Spectral Layer Scheduling

谱引导扩散:通过静态谱层调度加速推理
Shihab, Ibne Farabi, Ahsan, Abu Sa-Adat Mohamed Moon-Im Al, Sharma, Anuj
Abstract
Diffusion inference repeatedly evaluates the same large network. We ask whether pretrained weights alone can identify residual branches that need not be recomputed throughout the trajectory. Our \textbf{Spectral Concentration Ratio (SCR)} measures leading-versus-tail singular-value energy. Combined with Frobenius magnitude, it yields an offline sensitivity proxy and a deterministic lifetime for each scheduled unit. A frozen unit reuses its cached residual-branch update while the current residual stream and all external conditioning continue to propagate. The method needs no router, calibration prompts, or input-dependent search. At matched layer-step budgets, SCR/Frobenius preserves quality better than random, depth, norm, stable-rank, and Frobenius--stable-rank schedules on LLaDA-8B, DiT-XL/2, U-ViT-L, and SDXL. Broader LLaDA tests cover retrieval, reasoning, code, summarization, and open-ended generation; matched-horizon controls retain the ranking down to ten denoising steps. The complete captured-graph system reaches $2.8\times$--$3.0\times$ wall-clock speedup over eager inference. This is a systems-level number: on LLaDA, padded graph execution already gives $2.7\times$, while eliminating inactive branch work raises it to $3.0\times$. The perturbation analysis motivates pre-norm attention and MLP components under explicit local assumptions; results on AdaLN, U-shaped, convolutional, and cross-attention blocks are empirical transfer, not certified guarantees.
Chinese Translation
扩散模型推理需要反复评估同一个大型网络。我们探讨仅凭预训练权重能否识别出在整个去噪轨迹中无需重复计算的残差分支。我们提出的谱集中率(Spectral Concentration Ratio, SCR)用于衡量奇异值能量中主导部分与尾部部分的对比。结合 Frobenius 范数,它可作为一个离线的敏感度代理指标,并为每个被调度的单元确定确定性的生命周期。被冻结的单元会复用其缓存的残差分支更新,而当前的残差流以及所有外部条件输入仍继续传播。该方法无需路由器、校准提示或依赖输入的搜索。在相同的层-步预算下,SCR/Frobenius 调度在 LLaDA-8B、DiT-XL/2、U-ViT-L 和 SDXL 上比随机、深度、范数、稳定秩以及 Frobenius--稳定秩等调度方式更好地保持了生成质量。在 LLaDA 上更广泛的测试涵盖检索、推理、代码、摘要和开放式生成;在同等时间范围的对照实验中,该排名一直保持到仅剩十个去噪步骤。完整的捕获图系统相比 eager 推理实现了 $2.8\times$--$3.0\times$ 的实际运行时间加速。这是一个系统层面的数字:在 LLaDA 上,填充图执行本身已可带来 $2.7\times$ 的加速,而消除非活跃分支的计算可将其提升至 $3.0\times$。扰动分析在明确的局部假设下为 pre-norm 注意力和 MLP 组件提供了理论依据;而在 AdaLN、U 形、卷积以及交叉注意力模块上的结果属于经验性推广,并非经过认证的保证。
cs.LG / 68 / 2609.29518

CataOPD: Catalytic On-Policy Distillation for Large Language Model Reasoning

CataOPD:面向大语言模型推理的催化式在线策略蒸馏
Liu, Wenjin, Wang, Chenxi, Wang, Jiapu, Cui, Zhe, Luu, Anh Tuan, Luo, Haoran
Abstract
Reinforcement learning (RL) and on-policy distillation (OPD) are two representative paradigms for improving large language model reasoning. However, when no correct trajectory is sampled, RL lacks a positive correctness signal, while OPD remains constrained by the reasoning trajectories reachable under the student's on-policy distribution. Therefore, we propose CataOPD, where the teacher acts as a catalyst rather than a target, expanding reachability while internalizing verified student-produced trajectories into a catalyst-free policy. Self-Rescue Routing uses empirically all-failed groups as routing signals rather than teacher-intervention triggers, first seeking correct trajectories through additional on-policy self-sampling. For problems unresolved after self-rescue, Catalytic-Guided Self-Resolution uses catalytic guidance to elicit a verified student-produced trajectory in the guided student distribution. Barrier-Weighted Internalization weights tokens by guided-to-unguided log-probability gaps, focusing updates on decisive tokens difficult without guidance. Experimental results show that CataOPD outperforms current baselines, extends independent student reasoning to still-unrecovered problems, and improves out-of-distribution generalization under catalyst-free inference. Our project is available at https://github.com/QwenQKing/CataOPD.
Chinese Translation
强化学习(RL)和在线策略蒸馏(OPD)是提升大语言模型推理能力的两种代表性范式。然而,当未能采样到正确的轨迹时,RL缺乏正确的正信号,而OPD则受限于学生在在线策略分布下可达的推理轨迹。因此,我们提出CataOPD,其中教师模型扮演催化剂而非目标的角色,在扩展可达性的同时,将经过验证的学生生成轨迹内化到一个无催化剂的策略中。自我救援路由(Self-Rescue Routing)将经验上全部失败的组作为路由信号,而非教师干预的触发器,首先通过额外的在线策略自采样寻找正确轨迹。对于自我救援后仍未解决的问题,催化引导式自我求解(Catalytic-Guided Self-Resolution)利用催化引导在受引导的学生分布中引出经过验证的学生生成轨迹。屏障加权内化(Barrier-Weighted Internalization)根据有引导与无引导之间的对数概率差距对token加权,将更新聚焦于在没有引导时难以确定的决策性token。实验结果表明,CataOPD优于现有基线方法,能够将学生的独立推理扩展到此前仍未解决的问题,并提升了无催化剂推理下的分布外泛化能力。我们的项目可在 https://github.com/QwenQKing/CataOPD 获取。
cs.LG / 69 / 2609.29520

Sample-Weighted End-to-End Trace-Norm Geometry for Multitask Learning

面向多任务学习的样本加权端到端迹范数几何
Mohammadigohari, Mahdi
Abstract
Multitask models combine a shared representation with task-specific outputs, but generalization bounds often control the two components separately. Such products can discard relative orientation and cancellation and can change under equivalent transformations of intermediate coordinates even when the represented predictors are unchanged. We study instead the sample-size-weighted trace norm of the end-to-end map from task coefficients to input-space predictors. For its fixed-radius class, we derive the exact empirical Rademacher complexity. The same quantity is characterized by eliminating a positive-definite task covariance after the representation acts and, in finite-dimensional intermediate spaces, by optimizing the separated product over all equivalent invertible refactorizations. Explicit constructions show unbounded orientation and factorization gaps and an exponential depth gap for cancelling linear layers. As a geometric application, finite-to-one Lipschitz shared maps yield an exact Sobolev task Gram matrix determined by multiplicity and local directional distortion. We evaluate the corresponding convex regularizer in two protocol-locked unseen suites. Across 252 paired held-out comparisons, weighted joint nuclear regularization improves average population excess over unweighted nuclear regularization by 0.00764, with a stratified-bootstrap 95% interval [0.00465, 0.01110]. Correct task counts improve average and least-sampled-quartile excess over shifted counts by 0.01072 and 0.02847; all 15 imbalanced rank-suite cells are positive and the balanced effect is zero. Weighted joint nuclear also outperforms weighted Frobenius and independent ridge. The least-sampled-quartile comparison with unweighted nuclear remains unresolved, delimiting rather than contradicting the average advantage. All seven predeclared gates pass.
Chinese Translation
多任务模型将共享表示与任务特定的输出相结合,但泛化误差界通常分别控制这两个组件。这种乘积形式可能丢失相对方向信息和抵消效应,并且即使所表示的预测器不变,在中间坐标的等价变换下也会发生改变。我们转而研究从任务系数到输入空间预测器的端到端映射的样本量加权迹范数。对于其固定半径的函数类,我们推导出了精确的经验Rademacher复杂度。该量可以通过在表示作用后消去一个正定任务协方差矩阵来刻画,且在有限维中间空间中,也可通过在所有等价可逆重分解上优化分离乘积来刻画。显式构造展示了无界的方向间隙和因子化间隙,以及相互抵消的线性层之间的指数级深度间隙。作为一个几何应用,有限对一的Lipschitz共享映射产生一个由重数和局部方向畸变确定的精确Sobolev任务Gram矩阵。我们在两个协议锁定的未见测试套件上评估了相应的凸正则化器。在252组成对保留集比较中,加权联合核范数正则化相比无加权核范数正则化,平均总体超出量提升了0.00764,分层自助法95%置信区间为[0.00465, 0.01110]。正确的任务数相比偏移的任务数,平均超出量和最少采样四分位超出量分别提升0.01072和0.02847;所有15个不平衡秩套件单元均为正,而平衡情况下的效应为零。加权联合核范数也优于加权Frobenius范数和独立岭回归。与无加权核范数的最少采样四分位比较仍未有定论,这是对平均优势的界定而非矛盾。所有七项预设验证标准均通过。
cs.LG / 70 / 2609.29525

Common Covariance Geometry and Certification for Brownian Kernel Ladders

布朗核梯子的公共协方差几何与可验证性证明
Mohammadigohari, Mahdi
Abstract
A representation-adaptive kernel class produces, on a fixed sample, a union of reproducing-kernel Hilbert-space ellipsoids rather than one ellipsoid. We introduce the minimum-trace common covariance that dominates the unrestricted empirical union generated by Brownian kernel ladders and develop its statistical, approximation-theoretic, and computational consequences. The covariance value admits exact formulations through absolutely two-summing operators and covariance-dominated multipliers, and it yields a universal Gaussian-complexity bound. A closed last-layer Dirac-trace reduction and a signed Brownian threshold representation convert the generic covariance problem into threshold, graph-coarea, and effective-resistance geometry. These tools give deterministic depth laws, conditional Gaussian reverses, random-design and perturbation transfers, and an exact empirical Kolmogorov-width formula whose leading covariance eigenspaces approximate the complete adaptive ball simultaneously. Finite contact, active semidefinite programs, verified separation, and a convex resistance-design relaxation provide complementary lower and upper certificates. A finite covariance-indexed Brownian path on frozen representations illustrates the distinction between successful covariance certification and predictive selection: all reported path certificates succeed, whereas the locked predictive study misses one predeclared aggregate criterion. The paper thereby identifies one finite-dimensional covariance object linking unrestricted kernel adaptation, Gaussian geometry, common subspaces, and certifiable computation.
Chinese Translation
一类表示自适应的核方法在固定样本上生成的不是一个椭球,而是一族再生核希尔伯特空间椭球的并集。本文引入了支配由布朗核梯子(Brownian kernel ladders)生成的无约束经验并集的最小迹公共协方差,并发展其统计、逼近论与计算方面的推论。该协方差值可通过绝对二和算子(absolutely two-summing operators)与协方差支配乘子获得精确表述,并由此得到一个通用的高斯复杂度界。借助末层的闭式Dirac迹约化和带符号的布朗阈值表示,一般的协方差问题可转化为阈值几何、图上面积(graph-coarea)几何与有效电阻(effective-resistance)几何。这些工具给出了确定性的深度定律、条件高斯反演、随机设计与扰动下的迁移性质,以及一个精确的经验Kolmogorov宽度公式,其首协方差特征子空间可同时逼近完整的自适应球。有限接触点、主动半定规划、经验证的可分性与凸电阻设计松弛为问题提供了互补的下界与上界证书。一个基于冻结表示的有限协方差索引布朗路径实例阐释了成功的协方差认证与预测性选择之间的区别:所有报告的路径证书均成功,而锁定的预测性研究则遗漏了一项预先声明的综合准则。由此,本文识别出一个有限维协方差对象,将无约束核自适应、高斯几何、公共子空间与可验证计算联系在一起。
cs.LG / 71 / 2609.29530

The Impossible Trinity of Time-Series Validation: A Conservation Law among Training Sufficiency, Test Coverage, and Temporal Causality

时间序列验证的不可能三角:训练充分性、测试覆盖性与时间因果性之间的守恒定律
Li, Jiayu
Abstract
Validating a model on a time series asks for three things at once: each training run should use most of the sample (sufficiency), the test sets should together cover most of the sample (coverage), and training data should come before test data (causality). We prove that the three cannot be had together and price each one. Let $\alpha$ be the smallest training fraction over folds, $\beta$ the fraction of the sample covered by tests, $\Lambda$ the fraction of the sample used as training data from the future of a test point, and $\delta$ the distance from a test point to the nearest training point in its future. Every scheme on a sample of length $T$ satisfies $\alpha+\beta \le 1+\Lambda$ and $\alpha+\min\{\beta,\delta/T\} \le 1$, and under $\beta$-mixing the leakage bias at a test point is at most $2M\beta_{\mathrm{mix}}(\delta)$. In words: going beyond the causal frontier $\alpha+\beta=1$ requires training on the future; that future data must sit within $(1-\alpha)T$ of a test point; and its harm depends on its distance, not its amount. Hence expanding walk-forward is exactly the Pareto frontier of causal validation, $k$-fold cross-validation buys the most future data, and purged $k$-fold with an embargo pays in distance instead, which is cheap when the process forgets quickly but cannot repair the part of causality demanded by non-stationarity. On pure noise, shuffled 5-fold reports an information coefficient of $+0.32$, while contiguous 5-fold, using the same amount of future data, reports $+0.004$.
Chinese Translation
在时间序列上验证模型需要同时满足三个条件:每次训练应使用大部分样本(充分性)、测试集合起来应覆盖大部分样本(覆盖性)、且训练数据应早于测试数据(因果性)。我们证明这三者不可兼得,并为每一项进行定价。设 $\alpha$ 为各折中最小的训练比例,$\beta$ 为测试覆盖的样本比例,$\Lambda$ 为来自某测试点未来并被用作训练数据的样本比例,$\delta$ 为测试点到其未来最近的训练点的距离。对于长度为 $T$ 的样本,任何验证方案都满足 $\alpha+\beta \le 1+\Lambda$ 和 $\alpha+\min\{\beta,\delta/T\} \le 1$,且在 $\beta$-混合($\beta$-mixing)条件下,测试点处的泄漏偏差至多为 $2M\beta_{\mathrm{mix}}(\delta)$。简言之:突破因果边界 $\alpha+\beta=1$ 必须使用未来数据进行训练;这些未来数据必须位于测试点 $(1-\alpha)T$ 的范围内;且其危害取决于距离而非数量。因此,逐步扩展的滚动前推验证(walk-forward)恰好是因果验证的帕累托前沿,$k$ 折交叉验证获取最多未来数据,而带禁运期的清洗 $k$ 折验证则以距离为代价——当过程遗忘较快时代价低廉,但无法弥补非平稳性所要求的因果性。在纯噪声数据上,随机打乱的 5 折交叉验证报告的信息系数为 $+0.32$,而使用相同未来数据量的连续 5 折交叉验证仅报告 $+0.004$。
cs.LG / 72 / 2609.29546

Generalized Graph Variational Autoencoders: Bounded Divergences Control Posterior Collapse

广义图变分自编码器:有界散度控制后验坍缩
da Costa, Kleyton, Modenesi, Bernardo, Menezes, Ivan F. M., Lopes, Helio
Abstract
The variational graph autoencoder (VGAE) regularizes its posterior toward the prior with the Kullback-Leibler divergence, a choice inherited from the variational autoencoder rather than argued for. We introduce the generalized graph variational autoencoder (GGVA), which replaces that term with any member of the R\'enyi-Tsallis family of order $q$ while leaving every other part of the model untouched. Both members admit closed forms for diagonal Gaussians and both recover the KL exactly as $q \to 1$, so the VGAE is the $q=1$ arm of our own model rather than a separate baseline, and any measured difference is attributable to a single scalar. Our analysis identifies boundedness, not the order, as the operative property: for $q<1$ the Tsallis divergence is bounded above by $1/(1-q)$, independently of the latent width, whereas the KL and the R\'enyi divergence of the same order are unbounded. On ten graphs spanning three synthetic families, a social network, three citation networks, a connectome, a power grid and a road network, $q$ moves the retained posterior information by up to $49\times$ relative to the VGAE, while the R\'enyi arm at the same order stays within $1.02$-$1.30\times$ of it on all six larger real graphs (isolating the bound as the cause). The retained information is usable: probing the frozen embedding for node class, a label absent from the objective, gives GGVA up to $+0.14$ macro-F1 over the VGAE on CiteSeer, with the R\'enyi control again tracking the VGAE. We also report what the design was built to expose: none of this reaches held-out link-prediction accuracy on any of the six larger real graphs, and boundedness delays posterior collapse rather than preventing it.
Chinese Translation
变分图自编码器(VGAE)通过Kullback-Leibler散度将后验分布向先验分布正则化,这一选择沿袭自变分自编码器(VAE),而非经过论证。我们提出广义图变分自编码器(GGVA),用$q$阶Rényi-Tsallis散度族中的任意成员替换该项,而模型的其他部分保持不变。两者在带对角高斯的情况下均有闭式解,且当$q \to 1$时都精确退化为KL散度,因此VGAE只是我们模型的$q=1$情形,而非独立基线,任何测得的差异都可归因于单一标量。我们的分析表明,起作用的关键属性是有界性而非阶数:当$q<1$时,Tsallis散度有上界$1/(1-q)$,与潜变量维度无关,而同阶的KL散度和Rényi散度均无界。在十个图(涵盖三个合成族、一个社交网络、三个引文网络、一个连接组、一个电网和一个道路网络)上的实验显示,相对于VGAE,$q$使保留的后验信息最多提升$49\times$,而同阶的Rényi版本在六个更大的真实图上仅为其$1.02$-$1.30\times$(从而将这种差异归因于有界性)。保留的信息是可用的:对冻结的嵌入探测目标函数中不存在的节点类别标签,GGVA在CiteSeer上的macro-F1最高比VGAE提升$+0.14$,而Rényi对照仍与VGAE接近。我们也报告了该设计旨在揭示的问题:这些改进在六个较大的真实图上均未能转化为留出链路预测精度的提升,且有界性只是延缓而非阻止后验坍缩。
cs.LG / 73 / 2609.29548

Certified Predictive Value-of-Advice Gating for Cost-Aware Language-Model Guidance in Reinforcement Learning

强化学习中成本感知语言模型引导的经认证的建议预测价值门控
Shihab, Ibne Farabi, Swaqeeb, Md Najmus, Ahsan, Abu Sa-Adat Mohamed Moon-Im Al
Abstract
Language-model advice can accelerate reinforcement learning, but calls are costly and returned actions may be stale or wrong. We formulate advice acquisition as a response-contingent metareasoning problem: before querying, the controller predicts possible parsed responses, evaluates the decision and declared continuation that would follow each response, and queries only when a lower confidence bound on predictive value exceeds the priced cost. Execution is governed separately by an action-specific certificate. Under explicit assumptions, certified advice is near-optimal, a wrapped learner inherits fallback regret only under intervention stability, and conservative allocation loses at most the declared query-value estimation error relative to a myopic oracle. On BabyAI, a proxy-calibrated controller with Qwen2.5-1.5B and 7B advisors improves GoToObj return over no querying by 0.029 +/- 0.016 and 0.030 +/- 0.015 across 20 seeds while reducing calls by more than 97% relative to always-query. GoToLocal is a null result. Exactly matched-call tests show an advantage over random placement only for the 1.5B advisor and no advantage over an equal-budget early schedule. Mondrian calibration improves decision-relevant empirical coverage from 0.47 to 0.85, still below the 0.90 target, while the formally covered radius is vacuous. The demonstrated benefit is therefore robust sparse advice volume on a useful task, not a proven per-state placement advantage.
Chinese Translation
语言模型的建议可以加速强化学习,但调用成本高昂,且返回的动作可能是过时或错误的。我们将建议获取形式化为一个依赖响应的元推理问题:在查询之前,控制器预测可能的解析响应,评估每个响应之后将采取的决策与声明的延续策略,并且仅当预测价值的置信下界超过定价成本时才进行查询。执行则由一个针对特定动作的证书单独控制。在明确的假设条件下,经认证的建议接近最优;被包装的学习器仅在外部干预稳定性条件下才继承回退遗憾;相对于短视oracle,保守分配的损失至多为声明的查询价值估计误差。在BabyAI上,使用Qwen2.5-1.5B和7B作为建议者的代理校准控制器,在20个随机种子下,相比不查询将GoToObj的回报分别提高了0.029 +/- 0.016和0.030 +/- 0.015,同时相对于总是查询策略将调用次数减少了97%以上。GoToLocal任务则为无效结果。完全匹配调用次数的测试表明,相对于随机放置,仅在1.5B建议者上具有优势,且相对于等预算的早期调度没有优势。Mondrian校准将与决策相关的经验覆盖率从0.47提升到0.85,仍低于0.90的目标,而形式上覆盖的半径则是空洞的。因此,所展示的收益是在有用任务上稳健的稀疏建议体量,而非被证明的逐状态放置优势。
cs.LG / 74 / 2609.29564

Classifier-Dependent Benefits of Pseudo-Labeling for Semi-Supervised Android Malware Attribution

伪标签在半监督安卓恶意软件溯源中依赖于分类器的收益
Islam, Md Rafid, Hasan, Zahid, Rahman, Hafiz Abdur
Abstract
Detecting and classifying Android malware families remains challenging due to high feature dimensionality, class imbalance, and the high cost of expert-labeled data. Semi-supervised learning (SSL) offers a way to leverage unlabeled samples, but prior works rarely test whether SSL benefits generalize across classifier types or report statistical significance. We present a systematic evaluation of pseudo-labeling across six classifiers (LightGBM, XGBoost, Random Forest, Logistic Regression, MLP, and SVM) on the CICMalDroid 2020 dataset, using five-fold stratified cross-validation and paired t-tests across five labeled ratios (1-20%). We find that SSL benefit is strongly classifier-dependent: SVM shows the largest significant gain (+4.4% accuracy at 5% labels, p = 0.0028), LightGBM improves modestly (+0.8 to +1.3% at 2-5% labels), while Random Forest is significantly harmed at low label ratios (-3.1% at 1% labels). Per-class analysis reveals SSL disproportionately benefits the hardest-to-classify families, with Adware F1 improving by +13.8 percentage points versus only +0.8 for the already well-classified Benign class. We further show that approximately 800 labeled samples (10% of the dataset) yield near-optimal performance across all classifiers. These findings offer practical guidance on when and with which classifier pseudo-labeling is worthwhile for Android malware classification.
Chinese Translation
由于特征维度高、类别不平衡以及专家标注数据成本高昂,安卓恶意软件家族的检测与分类仍然极具挑战性。半监督学习(SSL)提供了一种利用未标注样本的途径,但已有研究很少检验SSL的收益是否能推广到不同类型的分类器,也很少报告统计显著性。我们在CICMalDroid 2020数据集上,采用五折分层交叉验证和配对t检验,在五个标注比例(1–20%)下,系统评估了伪标签方法在六种分类器(LightGBM、XGBoost、Random Forest、Logistic Regression、MLP和SVM)上的表现。我们发现,SSL的收益强烈依赖于分类器:SVM获得了最大的显著提升(在5%标注比例下准确率提升4.4%,p = 0.0028);LightGBM提升较为温和(在2–5%标注比例下提升0.8至1.3%);而Random Forest在低标注比例下受到显著损害(在1%标注比例下下降3.1%)。逐类别分析表明,SSL对最难分类的家族有不成比例的收益,其中广告软件(Adware)类别的F1值提升了13.8个百分点,而本已分类良好的良性(Benign)类别仅提升0.8个百分点。我们进一步发现,约800个标注样本(占数据集的10%)即可使所有分类器达到接近最优的性能。这些发现为安卓恶意软件分类中何时以及使用何种分类器进行伪标签学习具有价值提供了实用指导。
cs.LG / 75 / 2609.29580

When Identical Rows Disagree: From Benchmark Identifiability to Replication-Robust Anomaly Detection

当完全相同的行出现分歧时:从基准可识别性到复制鲁棒的异常检测
Deng, Jie
Abstract
A released table is often treated as an i.i.d. sample, although its repeated rows may encode business frequency, repeated entities, joins, resampling, or extraction errors. We show that this ambiguity creates a hidden measurement layer with three consequences: feature-identical rows impose an attained evaluation ceiling, row-weighted AUROC is sensitive to replication, and row-trained detectors learn a multiplicity-size-biased law. An exact-row audit of all 690 OddBench datasets finds train-test overlap in 355, feature-identical label conflict in 147, and a test anomaly identical to a training normal in 137. Switching from row to support weighting changes AUROC by at least 0.05 on 50-61 datasets across four classical detector geometries. We introduce SCOUT (Support-Count Orthogonalized Unsupervised Testing), a factorized anomaly detector that separates replication-invariant support evidence from exposure-aware count evidence. Factorwise split-conformal calibration yields marginal false-positive-rate control, while the support channel is exactly invariant to arbitrary positive row replication. On 686 OddBench datasets and five seeds, support-only SCOUT is non-inferior to row-wise Isolation Forest in raw AUROC and improves replication-invariant AUROC. External normal-support evaluations track nominal false-positive levels, and four backbones remain exactly unchanged under controlled replication. Semi-synthetic interventions show that conditional count modeling helps materially only under strong rate heterogeneity. These results specify when multiplicity should be treated as signal, nuisance, or uninterpretable without additional information.
Chinese Translation
发布的数据表通常被视为独立同分布(i.i.d.)样本,然而其重复的行可能编码了业务频率、重复实体、连接操作、重采样或数据提取错误。我们证明这种模糊性产生了一个隐藏的测量层面,并带来三方面后果:特征完全相同的行构成一个评估性能上限;按行加权的AUROC对复制操作敏感;基于行训练的检测器会学习到一种偏向多重性规模的规律。对全部690个OddBench数据集的精确行审计发现:355个数据集存在训练-测试重叠,147个存在特征相同但标签冲突的行,137个存在与训练集正常样本完全相同的测试异常样本。在四种经典检测器几何结构下,将按行加权切换为按支持集加权,会使50-61个数据集的AUROC变化至少0.05。我们提出SCOUT(Support-Count Orthogonalized Unsupervised Testing,支持-计数正交化无监督检验),一种因子化异常检测器,它将复制不变的证据与感知暴露程度的计数证据分离开来。基于逐因子的分裂保形(split-conformal)校准提供了边际假阳性率控制,而支持通道对任意正整数倍的行复制严格保持不变。在686个OddBench数据集和五个随机种子上,仅使用支持证据的SCOUT在原始AUROC上不劣于逐行的Isolation Forest,并在复制不变AUROC上有所提升。外部正常支持评估的结果与名义假阳性水平相符,且四个骨干模型在受控复制下保持严格不变。半合成干预实验表明,条件计数建模仅在强速率异质性条件下才能带来实质性帮助。这些结果明确了在缺乏额外信息的情况下,何时应将多重性视为信号、干扰或不可解释的因素。
cs.LG / 76 / 2609.29584

A Computational Framework for Modelling Organisation-Level Semantic Identity from Longitudinal Textual Data

基于纵向文本数据建模组织层面语义身份的计算框架
Krishna, Brinda Murali, Karakuş, Oktay, Eyupoglu, Can
Abstract
Organisations continuously generate large volumes of textual data that capture how they communicate, evolve and differentiate themselves over time. Although recent advances in natural language processing have substantially improved organisation-level text analytics, existing approaches primarily represent organisations as latent embeddings or predictive feature vectors for similarity estimation, classification or retrieval. Consequently, there is currently no general computational framework for modelling organisation-level semantic identity as an interpretable and evolving semantic construct derived from longitudinal textual evidence. This paper introduces a computational framework that integrates semantic representation learning, graph-based semantic modelling, organisation-level semantic fingerprints, temporal semantic evolution and evidence-driven validation within a unified analytical methodology. Organisations are characterised through complementary semantic dimensions describing diversity, concentration, connectivity, novelty and semantic community composition, which are analysed longitudinally to infer evidence-supported semantic identities. The framework is demonstrated using a longitudinal corpus of K-pop lyrics from artists affiliated with the four major South Korean entertainment companies. The empirical analyses reveal distinguishable multidimensional semantic identities, diverse temporal evolutionary trajectories and coherent integrated identity profiles. Comprehensive validation demonstrates that the inferred identities are statistically supported, robust under alternative analytical assumptions, reproducible and operationally informative. Beyond the case study, the proposed framework establishes organisation-level semantic identity and provides a transferable methodology for modelling organisational behaviour from longitudinal textual data.
Chinese Translation
组织持续产生大量文本数据,这些数据记录了其随时间的沟通方式、演变过程与自我区分。尽管自然语言处理的最新进展显著提升了组织层面的文本分析能力,但现有方法主要将组织表示为用于相似度估计、分类或检索的潜在嵌入或预测性特征向量。因此,目前尚缺乏一种通用的计算框架,能够将组织层面的语义身份建模为一种可解释的、基于纵向文本证据并不断演变的语义建构。本文提出一个计算框架,将语义表示学习、基于图的语义建模、组织层面的语义指纹、时序语义演化以及证据驱动的验证整合于统一的分析方法论之中。该框架通过描述多样性、集中度、连通性、新颖性以及语义社群构成的互补性语义维度来刻画组织,并进行纵向分析,从而推断出有证据支持的语义身份。该框架以韩国四大娱乐公司旗下艺人歌词的纵向语料库为案例进行了演示。实证分析揭示了可区分的多维语义身份、多样的时序演化轨迹以及连贯的综合身份画像。全面的验证表明,所推断的身份具有统计支持,在替代分析假设下具有稳健性,且可复现并具有实际参考价值。除案例研究外,所提出的框架确立了组织层面语义身份的概念,并提供了一种基于纵向文本数据建模组织行为的可迁移方法论。
cs.LG / 77 / 2609.29600

Active Client Selection in Federated Trajectory Prediction with Uncertainty-Awareness and Heterogeneous Complexity

具有不确定性感知与异构复杂度的联邦轨迹预测中的主动客户端选择
Xie, Yiming, Peng, Muzi, Miao, Fei, Mi, Ningfang, Su, Lili
Abstract
Training sequence models such as transformers is now standard for autonomous vehicle trajectory prediction, yet assembling high-quality centralized datasets remains challenging because real-world trajectories are fragmented across regions and vehicles. Federated Learning (FL) offers a natural alternative, but faces two distinctive challenges: high scene uncertainty arising from trajectory or map ambiguity, and cross-scene complexity heterogeneity caused by diverse map topology, traffic density, agent composition, and driving behaviors. We propose a family of active client selection methods that progressively incorporate awareness of scene uncertainty and complexity to prioritize informative clients. Our uncertainty-aware selectors use per-client negative log-likelihood under an uncertainty-aware global objective and estimated aleatoric uncertainty. We further develop a selector that jointly considers scene complexity and uncertainty, motivated by the intuition that knowledge from complex scenes can transfer to easier ones. Experiments on Argoverse show that federated trajectory prediction outperforms locally trained models. Uncertainty-aware selection accelerates convergence and improves minADE, minFDE, and MR. Under strong scene-complexity heterogeneity, our joint complexity- and uncertainty-aware selector achieves the best generalization and further accelerates convergence, demonstrating the benefit of prioritizing complex and informative scenes.
Chinese Translation
使用Transformer等序列模型训练自动驾驶车辆的轨迹预测已成为标准做法,然而由于真实世界的轨迹数据分散于不同区域和车辆之间,构建高质量的集中式数据集仍然极具挑战。联邦学习(Federated Learning, FL)提供了一种天然的替代方案,但面临两个独特的挑战:由轨迹或地图歧义引起的高场景不确定性,以及由多样地图拓扑、交通密度、智能体构成和驾驶行为导致的跨场景复杂度异构性。我们提出了一系列主动客户端选择方法,逐步引入对场景不确定性和复杂度的感知,以优先选择信息量大的客户端。我们的不确定性感知选择器使用在不确定性感知全局目标下的每客户端负对数似然以及估计的偶然不确定性。受“复杂场景的知识可以迁移到较简单场景”这一直觉启发,我们进一步开发了一个同时考虑场景复杂度和不确定性的选择器。在Argoverse数据集上的实验表明,联邦轨迹预测优于本地训练的模型。不确定性感知选择加速了收敛并提升了minADE、minFDE和MR指标。在强场景复杂度异构性的情况下,我们联合考虑复杂度与不确定性的选择器实现了最佳的泛化性能并进一步加速了收敛,证明了优先处理复杂且信息丰富的场景的益处。
cs.LG / 78 / 2609.29625

Limited Structural Reliability in Public Educational Prediction Benchmarks: A Four-Dimension Audit of Seven Datasets

公共教育预测基准中有限的结构可靠性:七个数据集的四维度审计
Ma, Yan, Zhang, Lizhuo
Abstract
Across seven public educational prediction datasets, three passed all four pre-modeling reliability checks; the remaining four either failed group-aware generalization tests or lacked the provenance metadata needed to run them. One dataset was initially classified as failing but corrected after excluding group-identifier features from the holdout matrix, demonstrating that the audit can distinguish genuine cross-group confounding from feature-encoding artifacts. Each dataset was audited before model optimization using four checks: baseline gap, split instability, null separation, and metadata adequacy under group-aware holdout. The dominant failure mode was not weak iid performance alone but cross-group fragility: in the clearest case, UCI Student declined from iid R-squared 0.242 to group-holdout R-squared -0.097, while Higher Ed collapsed from 0.041 to -8.79. Increasing model complexity did not remove this pattern: ensemble models improved structurally sound datasets but amplified instability or failed under group holdout on fragile ones. An exploratory cross-dataset comparison further showed that stronger profiles clustered in larger, richer-grouped, performance-proximal datasets, while random-split performance severely overstated deployable signal in fragile datasets. Classification-metric sensitivity analyses reached the same substantive conclusions. The results show that benchmark reliability in educational AI is constrained less by algorithm choice than by data structure, group heterogeneity, and evaluation design. A reusable pre-modeling audit offers a minimum quality gate before public educational datasets support strong benchmark or deployment claims.
Chinese Translation
在七个公开教育预测数据集中,有三个通过了全部四项建模前可靠性检查;其余四个数据集要么未通过群体感知泛化测试,要么缺乏执行这些测试所需的来源元数据。其中一个数据集最初被判定为不合格,但在从留出矩阵中剔除群体标识符特征后被修正,这表明该审计能够区分真正的跨群体混淆与特征编码伪影。每个数据集在模型优化之前均采用四项检查进行审计:基线差距、划分不稳定性、零假设分离以及群体感知留出下的元数据充分性。主要的失效模式并非仅仅是弱独立同分布(iid)性能,而是跨群体脆弱性:在最显著的案例中,UCI Student 数据集的 R² 从 iid 条件下的 0.242 下降到群体留出条件下的 -0.097,而 Higher Ed 数据集则从 0.041 崩溃至 -8.79。提高模型复杂度并未消除这一模式:集成模型改善了结构健全的数据集,但在脆弱数据集上加剧了不稳定性或在群体留出下失效。一项探索性跨数据集比较进一步表明,表现较好的数据集多集中于规模更大、分组更丰富且性能相近的数据集中,而在脆弱数据集中,随机划分的性能严重高估了可部署的信号。分类指标的敏感性分析得出了相同的实质性结论。研究结果表明,教育人工智能中的基准可靠性受制于数据结构、群体异质性和评估设计的程度大于算法选择。一个可复用的建模前审计为公共教育数据集在支撑强基准或部署声明之前提供了一个最低质量门槛。
cs.LG / 79 / 2609.29630

A Manifold-Aware Topic Modeling Approach via Rank-Based Prototypes

一种基于秩原型的流形感知主题建模方法
Almeida, Thiago César Castilho, Pedronette, Daniel Carlos Guimarães
Abstract
Recent topic models leverage pretrained embeddings, but neural architectures produce latent representations without grounding in specific texts, and clustering-based pipelines assign representative documents only post hoc, relying on absolute distances distorted by hubness and anisotropy in high-dimensional spaces. We introduce MARETopic, a training-free framework that casts topic discovery as rank-based prototype selection. After projecting embeddings onto a low-dimensional manifold, MARETopic builds ranked lists encoding ordinal neighborhood structure. A greedy algorithm selects exactly K exemplar documents, real corpus texts, whose neighborhoods cover the corpus. Two variants share this criterion. MARETopic$_\text{Corr}$ scores candidates with a query performance predictor and a rank correlation measure, leading Purity and NMI on the two benchmarks with the most categories, ahead of both neural and clustering-based topic models. MARETopic$_\text{Diff}$ scores them with a rank-based diffusion matrix, needs neither measure, and runs 1.7 to 1.9 times faster. Without a single gradient update, MARETopic leads topic coherence on two of three datasets. A novel inter-topic Maximal Marginal Relevance step raises vocabulary diversity at little cost in coherence. Our code is available at https://github.com/thcastilho/maretopic.
Chinese Translation
近来的主题模型利用预训练词嵌入,但神经架构产生的潜在表示缺乏与具体文本的锚定,且基于聚类的流程只能在事后指派代表性文档,其依赖的绝对距离会因高维空间中的枢纽性(hubness)和各向异性而产生失真。我们提出 MARETopic,一个无需训练的框架,将主题发现转化为基于秩的原型选择问题。在将嵌入投影到低维流形后,MARETopic 构建编码序数邻域结构的排序列表。一种贪心算法精确选取 K 个范本文档(即真实语料文本),其邻域覆盖整个语料库。两个变体共享该准则:MARETopic$_\text{Corr}$ 使用查询性能预测器和秩相关性度量对候选文档评分,在类别数最多的两个基准上取得最优的 Purity 和 NMI,超越神经主题模型和基于聚类的主题模型;MARETopic$_\text{Diff}$ 使用基于秩的扩散矩阵进行评分,无需上述任何度量,且运行速度快 1.7 至 1.9 倍。在不进行任何梯度更新的情况下,MARETopic 在三个数据集中的两个上取得最优的主题一致性(topic coherence)。一种新颖的主题间最大边际相关性(Maximal Marginal Relevance)步骤以很小的一致性代价提升了词汇多样性。我们的代码发布于 https://github.com/thcastilho/maretopic。
cs.LG / 80 / 2609.29668

Graph, Loop, and Harness Engineering for Zero-Trust Agentic Data Engineering and Analytical Processing

面向零信任智能体数据工程与分析处理的图、循环与执行框架工程
Sakhinana, Sagar Srinivas, Runkana, Venkataramana
Abstract
Large language model agents increasingly automate data workflows, but end-to-end cloud data engineering and analytical execution require reliable coordination across code, data, infrastructure, and runtime environments. We present two zero-trust frameworks. Zero-Trust Agentic Data Engineering generates, deploys, and verifies complete cloud data-engineering solutions from natural-language tasks, with completion conditioned on repository, deployment, runtime, and policy evidence. Zero-Trust Agentic OLAP combines governed Data Preparation with verified Online Analytical Processing (OLAP), permitting production promotion only after validation and evidence-bound approval, and releasing analytical answers only after Same-Snapshot Execution, Exact Result Equivalence, deterministic grounding, and reflection. Both frameworks share three abstractions: graph engineering for evidence-gated workflow structure, loop engineering for bounded recovery, and agent-harness engineering for zero-trust execution. We evaluate both frameworks under nominal execution, controlled failures, bounded recovery, and policy-constrained conditions, measuring verified completion, recovery, authorization enforcement, production promotion, and verified OLAP execution.
Chinese Translation
大语言模型智能体日益实现数据工作流的自动化,但端到端的云数据工程与分析执行需要在代码、数据、基础设施和运行时环境之间进行可靠的协调。我们提出两个零信任框架:零信任智能体数据工程(Zero-Trust Agentic Data Engineering)从自然语言任务生成、部署并验证完整的云数据工程解决方案,其任务完成的判定依赖于代码仓库、部署、运行时及策略证据;零信任智能体OLAP(Zero-Trust Agentic OLAP)将受治理的数据准备与经过验证的在线分析处理(OLAP)相结合,仅在通过验证并经过基于证据的审批后才允许晋升至生产环境,并且只有在满足同快照执行(Same-Snapshot Execution)、精确结果等价(Exact Result Equivalence)、确定性溯源和反思(reflection)检查之后才发布分析答案。两个框架共享三项抽象:用于证据门控工作流结构的图工程(graph engineering)、用于有界恢复的循环工程(loop engineering),以及用于零信任执行的智能体执行框架工程(agent-harness engineering)。我们在正常运行、受控故障、有界恢复和策略约束等条件下对两个框架进行评估,衡量指标包括经验证的完成率、恢复能力、授权强制执行、生产晋升以及经验证的OLAP执行。
cs.LG / 81 / 2609.29670

GBFRVFL: Granular-Ball Computing-Based Fuzzy Random Vector Functional Link Network

GBFRVFL:基于粒球计算的模糊随机向量函数链接网络
Quadir, A., Rahaman, A., Suganthan, P. N., Tanveer, M.
Abstract
In practical machine learning tasks, data are often contaminated with noise, outliers, and class imbalance, which can degrade the performance of conventional models. While random vector functional link (RVFL) networks offer fast training and strong generalization, they do not explicitly handle uncertainty or exploit local data structure. To address these limitations, we propose a fuzzy granular-ball random vector functional link (GBFRVFL) framework that leverages granular-ball computing to abstract raw samples into adaptive granular balls. Within this framework, we introduce two membership assignment schemes: (i) F-GBRVFL, which incorporates fuzzy membership to quantify the reliability of each granular ball, and (ii) SDAP-GBRVFL, which we propose, incorporates a novel statistical density-adaptive pythagorean membership (SDAPM) scheme that dynamically adjusts membership and non-membership values based on class variance, local sparsity, and granular-ball compactness. These schemes enhance robustness to noise, outliers, class imbalance, and uncertainty in granular-ball distributions, while retaining the computational efficiency of RVFL networks. Extensive experiments on 37 benchmark UCI and KEEL datasets under both clean and noisy conditions demonstrate that the proposed models consistently outperform baseline models, achieving superior accuracy and stability. The results validate the effectiveness of integrating granular-ball computing with adaptive membership schemes for reliable, scalable, and noise-tolerant learning.
Chinese Translation
在实际机器学习任务中,数据往往受到噪声、离群点和类别不平衡的污染,这会降低传统模型的性能。虽然随机向量函数链接(RVFL)网络具有训练速度快、泛化能力强的优点,但它们并未显式地处理不确定性,也没有利用局部数据结构。为了解决这些局限性,我们提出了一种模糊粒球随机向量函数链接(GBFRVFL)框架,该框架利用粒球计算将原始样本抽象为自适应的粒球。在该框架中,我们引入了两种隶属度分配方案:(i)F-GBRVFL,它引入模糊隶属度来量化每个粒球的可靠性;(ii)SDAP-GBRVFL,我们提出该方案引入了一种新颖的统计密度自适应毕达哥拉斯隶属度(SDAPM)机制,能够基于类别方差、局部稀疏性和粒球紧凑度动态调整隶属度和非隶属度值。这些方案增强了对噪声、离群点、类别不平衡以及粒球分布不确定性的鲁棒性,同时保留了RVFL网络的计算效率。在37个UCI和KEEL基准数据集上,无论在干净还是含噪条件下进行的大量实验表明,所提出的模型始终优于基线模型,取得了更优的准确率和稳定性。实验结果验证了将粒球计算与自适应隶属度方案相结合以实现可靠、可扩展且抗噪学习的有效性。
cs.LG / 82 / 2609.29674

The Sequential Price of Continual Learning

持续学习的顺序代价
Xu, Zonghuan, Ma, Xingjun
Abstract
Sequential task updates are fundamental to continual learning, but their recency bias can impose a lasting performance cost. We study this cost in an overparameterized linear-regression model with i.i.d. task sampling. We prove that distribution-level forgetting and population loss converge to the same stationary limit. This common limit separates exactly into the intrinsic loss asymptotically attained by joint training and an additional sequential price, and in more homogeneous task geometries the two terms coincide, making the total loss twice that of joint training. We further analyze fixed-strength elastic weight consolidation (EWC) under general task curvatures and characterize its stationary sequential price at every regularization strength. Under strong regularization, the price decays inversely with EWC strength while convergence to stationarity slows at the same scale. On the Jester joke-rating dataset, the theory exactly quantifies both the sequential price generated by naturally conflicting user preferences and its reduction by EWC.
Chinese Translation
顺序任务更新是持续学习的基础,但其对近期任务的偏好(recency bias)可能造成持续的性能损失。我们在具有独立同分布任务采样的过参数化线性回归模型中研究这一代价。我们证明,分布层面的遗忘与总体损失收敛到同一个平稳极限。该共同极限可以精确分解为联合训练渐近达到的内在损失和一个额外的顺序代价;并且在任务几何结构较为均匀的情况下,这两项相等,使总损失达到联合训练的两倍。我们进一步分析了一般任务曲率下固定强度的弹性权重固化(Elastic Weight Consolidation, EWC),并刻画了其在每个正则化强度下的平稳顺序代价。在强正则化条件下,该代价随EWC强度成反比衰减,而收敛到平稳状态的速度以相同尺度减缓。在Jester笑话评分数据集上,该理论精确量化了由用户偏好自然冲突所产生的顺序代价,以及EWC对该代价的削减效果。
cs.LG / 83 / 2609.29687

Not All Synthetic Data Are Equal: Expert-Committee Audit Screening for Imbalanced Crash-Injury-Severity Prediction in Automated Driving Systems

并非所有合成数据都是平等的:面向自动驾驶系统不平衡事故伤害严重程度预测的专家委员会审计筛选方法
Li, Zewei, Ren, Qiaoqiao, Yang, Hang, Wong, S. C., Mitoulis, Stergios-Aristoteles, Ye, Yun
Abstract
Automated driving systems (ADSs) are increasingly operating on public roads, raising safety concerns, yet reliable prediction of crash injury severity remains difficult because crash reports are limited, severe outcomes are rare, and injury classes are highly imbalanced. Existing augmentation methods mainly increase minority-class sample size but rarely assess whether generated samples are credible for safety-critical prediction. This study proposes Expert-Committee Audit Screening (ECAS), a credibility-aware sample acceptance framework for ADS crash injury severity prediction under data imbalance. Using 1,477 incident-level ADS crashes from the National Highway Traffic Safety Administration Standing General Order records, ECAS audits generated minority samples through a real-data-only expert committee based on label support, boundary separation, committee agreement, and local plausibility. Within-class percentile normalization and Pareto non-dominated sorting select accepted samples without manually assigned evidence weights. With a fixed backbone combining normalizing flow augmentation and a Tabular Prior-data Fitted Network (TabPFN) classifier, the best ECAS configuration achieved the highest balanced accuracy, macro-F1, and minor-injury recall among all evidence configurations. Local neighborhood analysis showed that ECAS-accepted samples were better supported by nearby real minority crashes than unscreened retained samples. Shapley additive explanations and partial dependence plots further indicated that lower injury severity classes were mainly associated with crash counterpart and pre-crash movement, whereas moderate-plus injuries were more sensitive to posted speed limit and operating context. These findings support a shift from quantity-oriented augmentation to credibility-aware sample acceptance for ADS safety prediction and risk governance.
Chinese Translation
自动驾驶系统(ADS)越来越多地在公共道路上运行,引发了安全方面的担忧。然而,由于事故报告数量有限、严重伤害结果罕见且伤害类别高度不平衡,对事故伤害严重程度的可靠预测仍然困难。现有的数据增强方法主要增加少数类样本的数量,但很少评估所生成的样本对于安全关键型预测是否可信。本研究提出了专家委员会审计筛选(Expert-Committee Audit Screening,ECAS),这是一个面向数据不平衡条件下自动驾驶系统事故伤害严重程度预测的、可信度感知的样本接受框架。该研究使用美国国家公路交通安全管理局常设通用令(Standing General Order)记录中的1,477起ADS事故级碰撞数据,ECAS通过仅基于真实数据构建的专家委员会,依据标签支持度、边界分离度、委员会一致性和局部合理性对生成的少数类样本进行审计。通过类内百分位归一化和帕累托非支配排序来选择被接受的样本,无需人工指定证据权重。在固定骨干网络(结合归一化流数据增强与表格先验数据拟合网络TabPFN分类器)的条件下,最优的ECAS配置在所有证据配置中取得了最高的平衡准确率、宏F1分数和轻伤召回率。局部邻域分析表明,与未经筛选的保留样本相比,ECAS接受的样本得到了附近真实少数类事故更好的支持。Shapley加性解释和部分依赖图进一步表明,较低的伤害严重程度类别主要与事故对方车辆及事故前运动状态相关,而中度及以上伤害则对限速和运行环境更为敏感。这些发现支持ADS安全预测与风险治理从以数量为导向的数据增强向可信度感知的样本接受转变。
cs.LG / 84 / 2609.29690

Predicting Symptoms of Amotivation and Anhedonia among University Students with a Novel Oversampling Method

基于一种新型过采样方法预测大学生动机缺失和快感缺失症状
Nguyen, Dang, Duong, Bao, Kumar, Arun, Phan-Trong, Dat, Berk, Julian, Braund, Taylor, Do, Kien, Bal, Debopriyo, Zheng, Wu Yi, Hoon, Leonard, Newby, Jill, Christensen, Helen, Venkatesh, Svetha, Whitton, Alexis, Gupta, Sunil
Abstract
University students experience disproportionately high rates of common mental health conditions, such as depression, which can impair learning, social functioning, and overall well-being. Within this context, symptoms of amotivation (i.e. loss of motivational drive) and anhedonia (i.e. diminished interest or pleasure) are particularly debilitating, yet they frequently go undetected. Developing new approaches to identify students with prominent amotivation and anhedonia could enable earlier and more targeted intervention. Machine learning (ML) methods have increasingly been used to classify individuals according to symptom severity. However, these ML models often suffer from class imbalance, where the majority of cases fall in the low-symptom group and relatively few in the high-symptom group. This imbalance can reduce model accuracy and bias predictions. To address this, studies commonly employ the popular oversampling strategy SMOTE. However, SMOTE has a notable limitation: it may generate invalid values for nominal variables. In this paper, we introduce a novel and effective oversampling method that addresses this shortcoming. Our approach leverages a predictive model to generate nominal variables, rather than interpolating them. We validate our method on a large-scale GPS location dataset collected from university students and demonstrate that it is significantly better than existing oversampling approaches in predicting elevated symptoms of amotivation and anhedonia.
Chinese Translation
大学生群体的常见心理健康问题(如抑郁症)患病率异常偏高,这些问题会损害学习、社会功能和整体幸福感。在此背景下,动机缺失(amotivation,即动机驱动力丧失)和快感缺失(anhedonia,即兴趣或愉悦感减退)的症状尤其具有破坏性,却常常未被察觉。开发新的方法来识别存在明显动机缺失和快感缺失症状的学生,有助于实现更早、更有针对性的干预。机器学习(ML)方法已被越来越多地用于根据症状严重程度对个体进行分类。然而,这些机器学习模型常常受到类别不平衡问题的困扰,即大多数样本属于低症状组,而高症状组的样本相对较少。这种不平衡会降低模型准确性并使预测产生偏差。为解决这一问题,研究通常采用流行的过采样策略SMOTE。然而,SMOTE存在一个显著局限:它可能会为名义变量生成无效的取值。本文提出了一种新颖且有效的过采样方法来弥补这一不足。我们的方法利用预测模型来生成名义变量,而非通过插值方式生成。我们在一个从大学生收集的大规模GPS位置数据集上验证了该方法,结果表明,在预测动机缺失和快感缺失症状升高方面,该方法显著优于现有的过采样方法。
cs.LG / 85 / 2609.29694

Bandit Multiclass PAC Learning: Corrected Lower Bounds, Exact Families, and a Confidence Direct-Sum Phenomenon

Bandit多类PAC学习:修正的下界、精确的函数族,以及一个置信度直和现象
Zhang, Guangjian
Abstract
We study realizable multiclass PAC learning with bandit feedback: the learner observes an i.i.d. instance, predicts one of $K$ labels, and learns only whether the prediction was correct. Hanneke, Meng, Moran, and Shaeiri (arXiv:2605.25678) characterized the optimal sample complexity via the bandit DS dimension $\mathrm{BDS}$ up to logarithmic factors, and asked whether every class admits sample complexity $O((\mathrm{BDS}+\log(1/\delta))/\epsilon)$. First, we show that the published lower bound $\Omega((\mathrm{BDS}+\log(1/\delta))/\epsilon)$ is incorrect as stated: we exhibit explicit classes with $\mathrm{BDS}=K-1$ whose sample complexity is exponentially smaller, and locate two independent gaps in its proof. We repair the lower-bound theory around a new anchored dimension $\mathrm{aBDS}\le\mathrm{BDS}$, proving a constant-free three-part lower bound. On the upper-bound side we remove the ambient label count $K$ entirely, proving $O((B\log^3 B+B\log(1/\delta))/\epsilon)$ for $B=\mathrm{BDS}$, plus a constant-confidence bound via a new fiberization lemma; for two natural families we determine the sample complexity up to constant factors. Finally, we answer the open question in the negative under its uniform-constant reading, and show the failure is intrinsic: for an explicit affine multiplexer class we establish the full confidence profile $\Theta((n\min{n,\log(1/\delta)}+\log(1/\delta))/\epsilon)$, a confidence direct-sum regime where a multiplicative $\log(1/\delta)$ cost is information-theoretically necessary, followed by a rank-saturation phase transition. Two classes with identical dimension profiles can have polynomially different sample complexities, so no characterization by these dimensions alone is accurate to polylogarithmic factors. We also show these results are consistent with additive-confidence list-PAC guarantees via the ListCascade bridge.
Chinese Translation
我们研究了带Bandit反馈的可实现多类PAC学习:学习器观察一个独立同分布的样本,预测$K$个标签中的一个,并且只能得知其预测是否正确。Hanneke、Meng、Moran和Shaeiri(arXiv:2605.25678)通过Bandit DS维数$\mathrm{BDS}$在对数因子范围内刻画了最优样本复杂度,并提出问题:是否每个类别都具有样本复杂度$O((\mathrm{BDS}+\log(1/\delta))/\epsilon)$。首先,我们证明已发表的下界$\Omega((\mathrm{BDS}+\log(1/\delta))/\epsilon)$按其表述是错误的:我们构造了$\mathrm{BDS}=K-1$的显式类别,其样本复杂度呈指数级更小,并找出了其证明中的两个独立漏洞。我们围绕一个新的锚定维数$\mathrm{aBDS}\le\mathrm{BDS}$修复了下界理论,证明了一个不含常数的三段式下界。在上界方面,我们完全去除了标签总数$K$的影响,对$B=\mathrm{BDS}$证明了$O((B\log^3 B+B\log(1/\delta))/\epsilon)$,并通过一个新的纤维化引理给出了固定置信度下的界;对于两个自然的类别族,我们将样本复杂度确定到常数因子。最后,我们在一致常数解读下对该开放问题给出了否定回答,并证明这种失效是内在的:对于一个显式的仿射多路复用器类别,我们建立了完整的置信度曲线$\Theta((n\min{n,\log(1/\delta)}+\log(1/\delta))/\epsilon)$,这是一个置信度直和区域,其中乘性$\log(1/\delta)$的代价在信息论上是必要的,随后出现秩饱和的相变。两个具有相同维数特征的类别可以具有多项式差异的样本复杂度,因此仅凭这些维数无法在多对数因子精度内刻画样本复杂度。我们还通过ListCascade桥梁证明,这些结果与加性置信度的list-PAC保证相容。
cs.LG / 86 / 2609.29696

An Agnostic Sample Compression Scheme for Squared Loss of Near-Linear Size in the Fat-Shattering Dimension

一种在fat-shattering维度下近线性大小的平方损失不可知样本压缩方案
Zhang, Guangjian
Abstract
We construct, for every function class $\mathcal{F}\subseteq[0,1]^{\mathcal{X}}$ and every accuracy $0<\alpha\le 1$, an agnostic sample compression scheme for the empirical squared loss: for every finite sample $S\in(\mathcal{X}\times[0,1])^m$ with arbitrary (noisy) labels, the scheme stores at most $O(\mathrm{fat}(\mathcal{F},c'\alpha)\cdot\log^3(2/\alpha))$ original labeled examples and auxiliary bits, independent of the sample size $m$, and reconstructs a function $\hat f$ with $L_2(\hat f,S)\le\inf_{f\in\mathcal{F}}L_2(f,S)+\alpha$. This resolves, in the positive, the open problem of Attias, Hanneke, Kontorovich, and Sadigurschi (ICML 2024, Section 5), which asks for an agnostic $\ell_2$ compression scheme of size $\mathrm{fat}(\mathcal{F},c\alpha)\cdot\mathrm{polylog}(c/\alpha)$. All previously known bounded-size constructions, agnostic and even realizable, incur a multiplicative dual fat-shattering factor, which can be exponentially larger than the primal dimension; our scheme removes the dual factor entirely, including in the realizable case. The dual factor in prior work enters solely through a sparsification step that forces uniform approximation on the sample. By targeting only a $(1-\epsilon)$-fraction of sample points, which suffices for an average-loss guarantee over a bounded range, K'egl's boosting margin bound yields $O(\log(1/\epsilon))$ rounds independent of $m$, and sparsification is never needed. The booster's synthetic target labels (values of a near-optimal $f^*\in\mathcal{F}$) are transmitted through quantized side-information bits attached to stored original examples, and the cross term of the squared loss forces the weak-learning scale $\Theta(\alpha)$, matching the same-scale form of the open problem.
Chinese Translation
对于每个函数类 $\mathcal{F}\subseteq[0,1]^{\mathcal{X}}$ 和每个精度 $0<\alpha\le 1$,我们构造了一个针对经验平方损失的不可知(agnostic)样本压缩方案:对于任意带有(含噪声)标签的有限样本 $S\in(\mathcal{X}\times[0,1])^m$,该方案最多存储 $O(\mathrm{fat}(\mathcal{F},c'\alpha)\cdot\log^3(2/\alpha))$ 个原始带标签示例及辅助比特,与样本量 $m$ 无关,并能重构一个函数 $\hat f$,满足 $L_2(\hat f,S)\le\inf_{f\in\mathcal{F}}L_2(f,S)+\alpha$。这肯定地解决了 Attias、Hanneke、Kontorovich 和 Sadigurschi(ICML 2024,第5节)提出的开放性问题,即寻找大小为 $\mathrm{fat}(\mathcal{F},c\alpha)\cdot\mathrm{polylog}(c/\alpha)$ 的不可知 $\ell_2$ 压缩方案。此前所有已知的有限大小构造(包括不可知情形甚至可实现情形)都带有乘性的对偶(dual)fat-shattering 因子,该因子可能比原(primal)维数指数级更大;我们的方案完全消除了对偶因子,包括在可实现情形下。以往工作中的对偶因子仅通过一个稀疏化步骤引入,该步骤要求对样本的一致逼近。通过仅针对 $(1-\epsilon)$ 比例的样本点——这足以保证有界范围内的平均损失——K'egl 的提升边界(boosting margin bound)可给出与 $m$ 无关的 $O(\log(1/\epsilon))$ 轮迭代,从而完全无需稀疏化。提升器(booster)的合成目标标签(某个近最优的 $f^*\in\mathcal{F}$ 的取值)通过附加在所存储原始示例上的量化辅助信息比特传输,而平方损失的交叉项迫使弱学习尺度为 $\Theta(\alpha)$,与该开放性问题中相同尺度的形式相匹配。
cs.LG / 87 / 2609.29706

Safety-oriented pedestrian trajectory prediction at urban intersections using time-to-collision and crossing-zone context

基于碰撞时间与过街区域情境的城市交叉口行人轨迹安全性导向预测
Avineri, Erel, Gil, Yftach, Aperstein, Yehudit
Abstract
Accurate pedestrian trajectory prediction is important for proactive road-safety applications, particularly at urban intersections where pedestrian motion is shaped by both vehicle interactions and crossing context. This study presents a safety-oriented trajectory-prediction framework that combines pedestrian motion history with Time-to-Collision (TTC) information and crossing-zone indicators. Using naturalistic trajectories from one urban intersection in the inD (Intersection Drone) dataset, several neural architectures were evaluated with 1.6 s observation and 2.4 s prediction horizons. A pooled Long Short-Term Memory (LSTM) separately encodes TTC histories and crossing-zone context before integrating them with pedestrian positions. In addition to conventional Average Displacement Error (ADE) and Final Displacement Error (FDE), prediction performance was assessed using the frequency and magnitude of errors exceeding a study-defined 1 m tolerance. A weighted loss was also introduced to place greater training emphasis on large coordinate-wise errors. Applying this loss to the position-only LSTM reduced ADE from 0.210 to 0.190 m and FDE from 0.550 to 0.503 m, while reducing ADE and FDE exceedance counts by 34.8% and 19.8%, respectively. The final pooled configuration incorporating TTC and crossing-zone information achieved an ADE of 0.184 m and FDE of 0.491 m, with further reductions of 33.5% and 6.3% in ADE and FDE exceedance counts relative to the safety-oriented position-only LSTM. The results indicate that safety-oriented training and structured integration of interaction and contextual information can reduce large trajectory-prediction errors, although broader validation across pedestrians, sites, and datasets is required.
Chinese Translation
准确的行人轨迹预测对于主动式道路安全应用至关重要,尤其是在城市交叉口,因为行人的运动既受车辆交互影响,也受过街情境的塑造。本研究提出了一种安全性导向的轨迹预测框架,该框架将行人运动历史与碰撞时间(Time-to-Collision, TTC)信息以及过街区域指标相结合。利用inD(Intersection Drone)数据集中某一城市交叉口的自然驾驶轨迹,在1.6秒观测时域和2.4秒预测时域下评估了多种神经网络架构。该方法采用一个池化长短期记忆网络(Long Short-Term Memory, LSTM)分别对TTC历史和过街区域情境进行编码,再将其与行人位置信息整合。除传统的平均位移误差(Average Displacement Error, ADE)和最终位移误差(Final Displacement Error, FDE)外,还采用超过本研究设定的1米容许阈值的误差频率与幅度来评估预测性能。此外,引入了一种加权损失函数,使训练更加侧重于较大的坐标级误差。将该损失应用于仅使用位置信息的LSTM,使ADE从0.210米降至0.190米,FDE从0.550米降至0.503米,同时ADE和FDE的超阈值次数分别减少了34.8%和19.8%。融合TTC与过街区域信息的最终池化配置实现了0.184米的ADE和0.491米的FDE,与安全性导向的仅位置LSTM相比,ADE和FDE超阈值次数分别进一步减少了33.5%和6.3%。研究结果表明,安全性导向的训练以及交互与情境信息的结构化整合能够减少较大的轨迹预测误差,但仍需要在更多行人、地点和数据集上进行更广泛的验证。
cs.LG / 88 / 2609.29709

Three Ways Classical Test Theory Misleads for LLM Judges

经典测试理论误导大语言模型(LLM)评判者的三种方式
Zhu, Louis Yiven
Abstract
An LLM judge scores a bank of responses against a rubric, and the reliability comes back at $0.52$. What has been measured? Judge evaluation has begun borrowing reliability statistics from classical test theory, usually without stating the measurement design each statistic assumes, and we show that three widely portable ones mean something different for a judge than for a test because the judge setting rearranges the roles those designs rest on. First, an internal-consistency coefficient computed over rubric elements contains no scorer facet. Holding one judge's measured error rate fixed at $4.72\%$, KR-20 still ranges from $0.01$ to $0.68$ as the item bank is redesigned around it, and varying judge error moves the coefficient by a comparable amount, so item design and judge error are not separately identified and no single value can be read as a property of the judge. Second, the dependability index $\Phi(\lambda)$ is a ratio of variance components, and the classification probability with which it is sometimes identified differs from it by $0.25$-$0.43$ on our bank and by $0.17$-$0.30$ on simulated data where the underlying model holds exactly. Third, Livingston-Lewis accuracy is indexed to an examinee's own true score on the same instrument, so scoring it against external gold conflates judge unreliability with criterion invalidity. Reviewing the three closest judge-evaluation papers, we found no published instance of these errors, which makes the caution prospective. A coefficient that cannot be attributed to the judge nonetheless travels downstream into deployment decisions and disclosure documents. We therefore close with four reporting lines that keep the attribution attached to the number.
Chinese Translation
一个LLM评判者根据评分细则对一组回答进行打分,其信度计算结果为$0.52$。这测量了什么?评判者评估已经开始借用经典测试理论的信度统计量,但通常并未说明每种统计量所假设的测量设计。我们证明,三种被广泛移植的统计量对于评判者的意义不同于对测试的意义,因为评判者情境改变了这些设计所依赖的角色关系。第一,基于评分细则要素计算的内部一致性系数不包含评分者维度。在固定某评判者4.72%的实测错误率的情况下,随着题库围绕其重新设计,KR-20仍在$0.01$到$0.68$之间变化;而改变评判者错误率也会使该系数产生相近幅度的变动。因此,题目设计与评判者错误无法被分别识别,任何单一数值都不能被解读为评判者的属性。第二,可信度指数$\Phi(\lambda)$是方差分量的比值,它有时被等同于分类概率,但在我们的题库上二者相差$0.25$至$0.43$,在底层模型完全成立的模拟数据上也相差$0.17$至$0.30$。第三,Livingston-Lewis准确率是以考生在同一测评工具上的自身真分数为参照的,因此用外部金标准来衡量它会混淆评判者的不可靠性与效标的无效性。我们回顾了与评判者评估最相关的三篇论文,未发现已发表的此类错误实例,这使得本文的警示具有前瞻性。然而,一个无法归因于评判者的系数仍会向下传递,影响部署决策和披露文件。因此,我们最后提出四条报告准则,以确保归因与数值始终绑定。
cs.LG / 89 / 2609.29711

Decoupling Knowledge and Privacy: Post-Task Self-Distillation Replay for LLM Continual Learning

知识与隐私的解耦:面向大语言模型持续学习的任务后自蒸馏回放方法
Wen, Shengtao, Yang, Yunying, Chen, Xiang, Guo, Lingbing, Tian, Yu, Huang, Sheng-Jun
Abstract
Privacy-preserving continual learning (PPCL) must reduce the reproduction of sensitive content while retaining useful knowledge across sequential tasks. Formal privacy guarantees characterize randomized mechanisms, whereas operational output control concerns whether a trained model selectively reduces the likelihood of sensitive content in its outputs. In this work, we investigate the latter together with continual-learning utility under realistic task evolution. Retention and privacy correction operate at different granularities: task acquisition requires broad preservation of current- and old-task behavior, whereas privacy correction targets sparse annotated positions. Joint optimization leaves the current-task preservation target continually changing. We propose SPARK, a retention-correction decomposition that first freezes the learned post-task distribution and then applies selective correction around this stable reference. Self-Distillation Replay learns the current task while distilling behavior from previous tasks, and Post-Task Privacy Correction reduces annotated-PII likelihood while anchoring current- and old-task non-PII behavior to the resulting checkpoint. Extensive evaluations demonstrate that SPARK achieves effective selective PII suppression while preserving strong continual-learning utility and knowledge retention across diverse settings. Code and data will be released upon publication.
Chinese Translation
隐私保护持续学习(PPCL)必须减少敏感内容的复现,同时在序列任务中保留有用知识。形式化隐私保证刻画的是随机化机制,而操作性输出控制关注的则是训练后的模型是否能有选择地降低敏感内容在输出中出现的可能性。在本工作中,我们在现实的任务演化场景下研究后者与持续学习效用。保持与隐私校正在不同的粒度上运作:任务获取需要对当前任务和旧任务行为的广泛保持,而隐私校正则针对稀疏的标注位置。联合优化会使当前任务的保持目标不断变化。我们提出了SPARK,一种保持-校正分解方法,它首先冻结学习得到的任务后分布,然后围绕这一稳定参考施加选择性校正。自蒸馏回放(Self-Distillation Replay)在学习当前任务的同时蒸馏来自先前任务的行为,而任务后隐私校正(Post-Task Privacy Correction)在降低标注PII(个人身份信息)出现可能性的同时,将当前任务和旧任务的非性PII行为锚定到由此得到的检查点上。大量评估表明,SPARK在多样化设置下实现了有效的选择性PII抑制,同时保持了强大的持续学习效用和知识保持能力。代码和数据将在论文发表后发布。
cs.LG / 90 / 2609.29715

Revalidation Beats Stateful Routing for Scientific Surrogates Under Distribution Shift

在分布偏移下,重新验证优于有状态路由的科学代理模型方法
Lodhiya, Harshil
Abstract
Surrogate models are often chosen during development and then left in place as new measurements arrive. That practice becomes risky when noise, input support, or physical parameters change. We asked whether such changes call for a stateful adaptive controller, or whether it is enough to validate the candidate models again on each new batch. To study this question, we built RegimeShift-Surrogates, a reproducible streaming benchmark spanning eight analytic and dynamical tasks, four stationary or shifting regimes, ten held-out seeds, and eight classical, multilayer-perceptron, and Kolmogorov-Arnold network surrogates. The confirmatory run contains 30,720 model fits and 3,200 scored deployment windows. Choosing the model with the lowest validation loss in the current window yields mean log regret 0.091 against a per-window oracle; the best fixed model chosen in hindsight yields 0.192. The paired difference is -0.101 (hierarchical bootstrap 95% CI [-0.165, -0.040]; Holm-adjusted p = 0.0469), with revalidation ahead in 26 of 32 task-scenario combinations. None of the stateful alternatives, including exponential smoothing, dual-timescale adaptation, Page-Hinkley resets, or margin gating, improves the pooled result, and delayed bias correction makes it worse. Oracle choices also differ substantially by task: k-nearest neighbors dominate the damped oscillator, vanilla KAN is often selected for two-dimensional surfaces, and MLPs lead on the Runge and Van der Pol tasks. In this benchmark, fresh validation evidence is useful; carrying old evidence forward is often not.
Chinese Translation
代理模型通常在开发阶段选定后便固定使用,即使新的测量数据不断到来也不再更换。当噪声、输入支撑范围或物理参数发生变化时,这种做法会变得具有风险。我们探究了此类变化是否需要引入有状态的自适应控制器,或者仅在每个新数据批次上对候选模型重新进行验证是否已经足够。为研究这一问题,我们构建了 RegimeShift-Surrogates——一个可复现的流式基准,涵盖八个解析与动力学任务、四种平稳或偏移的状态、十个留出的随机种子,以及八种经典方法、多层感知机和 Kolmogorov-Arnold 网络(KAN)代理模型。验证性实验共包含 30,720 次模型拟合和 3,200 个评分部署窗口。在当前窗口中选择验证损失最低的模型,相对于逐窗口预言机(oracle)的平均对数遗憾为 0.091;而事后选出的最优固定模型的遗憾为 0.192。两者配对差异为 -0.101(分层自助法 95% 置信区间 [-0.165, -0.040];经 Holm 校正的 p = 0.0469),在 32 个任务-场景组合中有 26 个组合重新验证方法占优。所有有状态的替代方法,包括指数平滑、双时间尺度自适应、Page-Hinkley 重置或边际门控,均未能改进汇总结果,而延迟的偏差校正反而使其变差。预言机的选择在不同任务间也存在显著差异:k 近邻在阻尼振荡任务中占主导,二维曲面上常被选中的是原始 KAN,而在 Runge 和 Van der Pol 任务中 MLP 表现领先。在本基准中,新的验证证据是有用的;而沿用过时的证据往往并非如此。
cs.LG / 91 / 2609.29740

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS:一个用于基于配体的虚拟筛选的现实多靶点基准测试集
Kumar, Surbhi, Zhou, Yuhe, Shiralkar, Varun, Huang, Niu, Coskunuzer, Baris
Abstract
Ligand-based virtual screening (LBVS) is a practical first-pass tool in early-stage drug discovery, but existing benchmarks can overestimate performance through random negatives, easy decoys, limited target coverage, and non-standardized evaluation protocols. We introduce TopU-LBVS, a multi-target benchmark for LBVS under hard-negative screening conditions. Starting from curated ChEMBL~35 bioactivity data, TopU-LBVS covers 93 protein targets across 7 protein classes and constructs target-specific screening libraries with property-matched, structurally similar decoys at a fixed 1:40 active-to-decoy ratio. Libraries contain roughly 400 to 10,000 compounds and are designed to reduce simple physicochemical and nearest-neighbor fingerprint shortcuts. TopU-LBVS provides three fixed protocols. TopU-LBVS-full evaluates ChEMBL$^\ast \rightarrow$ TopU generalization across all 93 targets. TopU-LBVS-low evaluates low-data TopU $\rightarrow$ TopU learning within the hard-negative distribution. TopU-LBVS-mini provides a compact seven-target protocol with a paired random-decoy control that changes only the test decoys, enabling low-cost development and direct measurement of the gap between random ChEMBL$^\ast$ and TopU decoys. Across ten reference baselines spanning fingerprint methods, molecular GNNs, fingerprint hybrids, and modern molecular models, performance under random-decoy evaluation degrades sharply under hard-negative screening. We release data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of future LBVS and molecular representation learning methods. Code and data are available at https://github.com/topu-benchmark/topu-lbvs and https://huggingface.co/datasets/topu-benchmark/topu-lbvs.
Chinese Translation
基于配体的虚拟筛选(LBVS)是早期药物发现中一种实用的初步筛选工具,但现有基准测试由于使用随机负样本、易区分的诱饵分子、有限的靶点覆盖范围以及不规范的评估协议,往往高估了方法性能。我们提出了 TopU-LBVS,一个在难负样本筛选条件下面向 LBVS 的多靶点基准测试集。基于整理后的 ChEMBL~35 生物活性数据,TopU-LBVS 涵盖 7 个蛋白质类别中的 93 个蛋白质靶点,并以固定的 1:40 活性分子与诱饵分子比例,构建了具有性质匹配且结构相似诱饵的靶点特异性筛选文库。每个文库包含约 400 至 10,000 个化合物,其设计旨在减少基于简单理化性质和最近邻分子指纹的捷径学习。TopU-LBVS 提供三种固定协议:TopU-LBVS-full 用于评估 ChEMBL$^\ast \rightarrow$ TopU 在全部 93 个靶点上的泛化能力;TopU-LBVS-low 用于评估难负样本分布内低数据量条件下的 TopU $\rightarrow$ TopU 学习;TopU-LBVS-mini 提供一个包含七个靶点的紧凑协议,并配有仅更换测试诱饵分子的配对随机诱饵对照组,从而支持低成本的方法开发,并可直接测量随机 ChEMBL$^\ast$ 诱饵与 TopU 诱饵之间的性能差距。在涵盖指纹方法、分子图神经网络(GNN)、指纹混合模型以及现代分子模型等十种参考基线方法上的实验表明,随机诱饵评估下的性能在难负样本筛选条件下会急剧下降。我们公开了数据、固定数据划分、评估代码和基线实现,以便对未来的 LBVS 和分子表征学习方法进行可复现的比较。代码与数据可在 https://github.com/topu-benchmark/topu-lbvs 和 https://huggingface.co/datasets/topu-benchmark/topu-lbvs 获取。
cs.LG / 92 / 2609.29772

WeatherDiagFlow: Evidence-Grounded Radar Nowcasting with Diagnostic Flow Refinement

WeatherDiagFlow:基于诊断流细化的证据支撑雷达临近预报
Shi, Chunlei, Zhu, Yufeng, Liang, Yixiao, Niu, Dan, Feng, Yongchao, Wu, Qiliang, Wang, Jiong
Abstract
Radar nowcasting is essential for short-term warning and emergency response, yet conventional systems mainly return future radar fields and provide limited support for operational communication and post-event verification. We formulate radar nowcasting as an evidence-grounded forecast--bulletin--audit task, in which a numerical forecaster produces both future radar fields and structured diagnostic evidence. Forecast-time bulletins use only model-available evidence, whereas post-event audits incorporate future radar truth only after the forecast horizon is observed. Based on this task formulation, WeatherDiagFlow predicts motion, growth and decay, heavy-echo risk, and uncertainty to condition rolling flow refinement, while frozen-scaffold residual calibration improves long-lead strong-echo preservation. A multi-agent layer converts the structured evidence into operational bulletins and independently generates verification audits without feeding textual outputs back into the forecaster. Experiments on FJRADAR demonstrate competitive overall performance and improved strong-echo event skill. WeatherDiagFlow therefore connects numerical prediction, evidence-grounded reporting, and auditable verification under a leakage-controlled protocol.
Chinese Translation
雷达临近预报对短期预警和应急响应至关重要,但传统系统主要返回未来雷达场,对业务化信息传达和事后验证的支持有限。我们将雷达临近预报定义为一个“预报—公报—审计”的证据支撑型任务:数值预报器同时生成未来雷达场和结构化诊断证据。预报时刻的公报仅使用模型可获得的证据,而事后审计则仅在预报时效过后观测到未来雷达真值时才引入该真值。基于这一任务设定,WeatherDiagFlow 预测移动、生消演变、强回波风险和不确定性,以条件化滚动流细化(rolling flow refinement);同时,冻结骨架残差校准(frozen-scaffold residual calibration)提升了较长时效下强回波的保留能力。多智能体层将结构化证据转化为业务公报,并独立生成验证审计,且不将文本输出反馈给预报器。在 FJRADAR 数据集上的实验表明,该方法具有竞争力的整体性能以及更优的强回波事件预报技巧。因此,WeatherDiagFlow 在防泄漏控制协议下,将数值预测、证据支撑的报告生成与可审计的验证有机连接起来。
cs.LG / 93 / 2609.29774

An Analytical Theory of Auxiliary Learning

辅助学习的一种解析理论
Milanesio, Federico, Ingrosso, Alessandro, Osella, Matteo
Abstract
Auxiliary learning is an optimization paradigm in which a neural network's performance on a target task is improved by jointly training it on additional tasks. However, the mechanisms behind this improvement remain poorly understood. We study this problem using a teacher-student framework and derive a closed system of differential equations describing the dynamics of online stochastic gradient descent in the large-input limit. For linear networks, we obtain a closed-form expression for the generalization error to leading order in the learning rate, quantifying how task correlations and label noise determine the benefit of auxiliary learning. For non-linear activation functions, we develop a fluctuation-dissipation analytical theory that establishes a general relation linking the main and auxiliary errors to the corresponding single-task error. Numerical experiments support the theoretical predictions and show how auxiliary tasks improve generalization by balancing the forcing dynamics towards the optimal solution with gradient noise.
Chinese Translation
辅助学习是一种优化范式,通过在附加任务上与目标任务联合训练来提升神经网络在目标任务上的性能。然而,这种提升背后的机制仍未被充分理解。我们基于教师-学生(teacher-student)框架研究这一问题,并推导出一个封闭的微分方程组,用以描述在线随机梯度下降在大输入极限下的动力学行为。对于线性网络,我们得到了泛化误差在学习率一阶意义下的闭式表达式,量化了任务相关性与标签噪声如何决定辅助学习的收益。对于非线性激活函数,我们发展了一种涨落-耗散解析理论,建立了主任务误差与辅助任务误差同相应单任务误差之间的一般关系。数值实验支持了理论预测,并展示了辅助任务如何通过平衡指向最优解的强迫动力学与梯度噪声来提升泛化性能。
cs.LG / 94 / 2609.29812

FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates

FlashLoop:基于惰性更新的快速内存高效循环Transformer
Yang, Wanqi, Liu, Shiwei
Abstract
Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer pass and requires caching another set of KV states, causing inference FLOPs and KV-cache memory to grow continuously with loop depth. This overhead becomes particularly severe at large loop counts and long context, preventing the parameter efficiency of Looped Transformers from translating into practical inference efficiency. In this paper, we find that much of the additional computation and storage introduced by looping is redundant. As recurrence proceeds, state changes become increasingly concentrated on a small subset of tokens; attention-output differences are dominated by a sparse and stable subset of key columns; and KV residuals between adjacent loops become progressively more amenable to low-bit quantization. Building on these observations, we introduce FlashLoop, a training-free inference framework that reduces cross-loop redundancy through token-sparse updates, sparse attention, and KV-residual quantization. Across several Looped Transformers models, \textsc{FlashLoop} delivers lossless accuracy while achieving up to 1.64$\times$ end-to-end speedup and up to 6$\times$ KV-cache memory reduction, substantially improving the practicality of scaling Looped Transformers to greater computational depths and longer context.
Chinese Translation
循环Transformer(Looped Transformers)通过重复应用共享的Transformer块来增加计算深度,作为一种参数高效的方法受到了广泛关注。然而,其相对于传统Transformer的实际优势仍存在争议:每增加一次循环都会带来一次额外的Transformer前向传播,并需要缓存另一组KV状态,导致推理FLOPs和KV缓存内存随循环深度持续增长。这种开销在大循环次数和长上下文场景下尤为严重,使循环Transformer的参数高效性难以转化为实际的推理效率。在本文中,我们发现循环引入的许多额外计算和存储是冗余的。随着递归的进行,状态变化越来越集中于少量token子集;注意力输出差异主要由稀疏且稳定的关键列子集主导;且相邻循环之间的KV残差逐渐更适合低比特量化。基于这些观察,我们提出了FlashLoop,这是一个无需训练的推理框架,通过token稀疏更新、稀疏注意力和KV残差量化来减少跨循环冗余。在多个循环Transformer模型上,FlashLoop在实现无损精度的同时,取得了最高1.64倍端到端加速和最高6倍KV缓存内存缩减,显著提升了将循环Transformer扩展至更大计算深度和更长上下文的实用性。
cs.LG / 95 / 2609.29814

SwitchPFN: Shared Switching Dynamics for Frozen In-Context Time Series Classification

SwitchPFN:面向冻结上下文时间序列分类的共享切换动力学
Zhu, Zhenyi, Pang, Jacqueline, Shen, Peilin, Song, Tianyi, Zhang, Tingwei, Hu, Keyi, Yin, Kangjun, Pu, Shiwei, Zhou, Yingbo, Shao, Chen
Abstract
Tabular foundation models (TFMs) provide a promising route to time-series classification, but their effectiveness depends on how sequential data are converted into tabular representations. Existing representations face two challenges: global aggregation can lose the order of temporal evolution, while features computed in independently fitted coordinate systems may not have consistent meanings across sequences. We therefore view representation design for TFMs as a problem in its own right: the representation should preserve local temporal transitions while maintaining a shared feature definition across samples. We propose SwitchPFN, which learns a shared projection and regime codebook from the training sequences, making local dynamic operators and transition features directly comparable across samples. Across the evaluated benchmarks, SwitchPFN achieves the highest mean accuracy among the evaluated methods, improving over the strongest baseline by 4.47% relatively. Ablation studies, parameter sensitivity analyses, and reduced-training-data experiments further examine the contributions of the representation, its main design choices, and its behavior when labeled data are limited.
Chinese Translation
表格基础模型(TFM)为时间序列分类提供了一条有前景的路径,但其有效性取决于如何将序列数据转换为表格表示。现有表示面临两个挑战:全局聚合可能丢失时间演化的顺序信息,而在独立拟合的坐标系中计算的特征在不同序列之间可能缺乏一致的语义。因此,我们将TFM的表示设计视为一个独立的问题:表示应在保持样本间共享特征定义的同时,保留局部的时间转移信息。我们提出SwitchPFN,它从训练序列中学习共享投影与状态码本(regime codebook),使局部动力学算子和转移特征在不同样本之间具有直接可比性。在所评估的基准数据集上,SwitchPFN在所有评估方法中取得了最高的平均准确率,相对于最强基线提升了4.47%。消融实验、参数敏感性分析以及训练数据减少实验进一步考察了该表示的贡献、其主要设计选择,以及在标注数据受限时的表现。
cs.LG / 96 / 2609.29879

From Graphs to Feeders: Constraint-Guided Diffusion for Rule-Compliant Feeder Generation

从图到馈线:面向规则合规馈线生成的约束引导扩散模型
Qin, Yu, Glaws, Andrew, Latif, Aadil, King, Ryan
Abstract
Generative modeling approaches often focus on recovering broad statistical characteristics from the training data. In the context of graph generation, this may refer to degree distributions, clustering coefficients, or spectral properties. However, generating usable distribution feeders when detailed feeder models are unavailable requires more than matching generic graph statistics: the sampled topology must also obey electrical compatibility and radiality rules. We therefore formulate feeder synthesis as a constraint-guided graph generation problem and propose the Power-Grid-constrained Discrete Denoising Diffusion model, PG-DiGress, which learns categorical node and edge patterns from feeder data, while respecting domain-specific rules. Specifically, it injects feeder constraints into the reverse diffusion process through soft masks that suppress incompatible edge classes during denoising, followed by a final projection step that rebuilds a connected, rule-compliant feeder graph. We evaluate PG-DiGress using graph-distribution similarity, feeder-rule satisfaction, structural validity, and downstream model construction. Compared with the unconstrained baseline, PG-DiGress increases the strict feeder pass rate from 13.7% to 96.8%. We also successfully convert the generated graphs into executable feeder models for downstream analysis.
Chinese Translation
生成式建模方法通常侧重于从训练数据中恢复广泛的统计特性。在图生成领域,这可能指度分布、聚类系数或谱特性。然而,在缺乏详细馈线模型的情况下生成可用的配电馈线,仅匹配通用的图统计特性是不够的:采样的拓扑结构还必须满足电气兼容性和辐射状(radiality)规则。因此,我们将馈线合成为一个约束引导的图生成问题,并提出了一种电力网约束的离散去噪扩散模型(Power-Grid-constrained Discrete Denoising Diffusion model),即 PG-DiGress。该模型从馈线数据中学习节点和边的类别模式,同时遵循领域特定规则。具体而言,它通过软掩码(soft mask)将馈线约束注入反向扩散过程,在去噪过程中抑制不兼容的边类别,随后通过最终的投影步骤重建连通且符合规则的馈线图。我们使用图分布相似性、馈线规则满足度、结构有效性以及下游模型构建等指标对 PG-DiGress 进行了评估。与无约束基线相比,PG-DiGress 将严格馈线通过率从 13.7% 提升至 96.8%。我们还成功地将生成的图转换为可执行的馈线模型以用于下游分析。
cs.LG / 97 / 2609.29906

Spatio-temporally complementary feature propagation on graphs for longitudinal AADT estimation

基于图上时空互补特征传播的长期AADT估计
Sun, Linghang, Zhou, Qishen, Makridis, Michail A., Kouvelas, Anastasios
Abstract
The estimation of Annual Average Daily Traffic (AADT) is vital for transportation planning and infrastructure maintenance, yet obtaining accurate values for an entire urban network across multiple years remains challenging due to the high cost and spatial sparsity of physical sensors. This research proposes a novel spatio-temporally complementary feature propagation framework that leverages the strengths of two distinct data sources: spatially sparse but temporally dense loop detector data, and a spatially complete but temporally sparse macroscopic transportation model. The methodology highlights a feature propagation algorithm on directed graphs, formulated as a Poisson energy minimization considering residues. The standard binary adjacency matrix is replaced with flow ratio matrices to capture real-world vehicle turn ratios at intersections. Validated in the city of Zurich, the algorithm demonstrates high computational efficiency, achieving convergence within minutes. Results indicate that the framework effectively reconciles theoretical models with empirical ground truths, yielding a normalized mean absolute error below $10\%$. This scalable approach provides a feasible solution for spatio-temporal network-wide AADT estimation through combining real-world limited sensor coverage and traffic models.
Chinese Translation
年度平均日交通量(AADT)的估计对交通规划和基础设施维护至关重要,然而由于物理传感器成本高昂且空间分布稀疏,获取覆盖整个城市网络多年份的精确AADT值仍然充满挑战。本研究提出了一种新颖的时空互补特征传播框架,充分利用两类不同数据源的优势:空间稀疏但时间密集的环形检测器数据,以及空间完整但时间稀疏的宏观交通模型。该方法的核心是一种有向图上的特征传播算法,其被表述为考虑残差的泊松能量最小化问题。算法以流量比例矩阵取代标准的二值邻接矩阵,以刻画交叉口真实的车辆转向比例。该算法在苏黎世市进行了验证,展现出较高的计算效率,可在数分钟内收敛。结果表明,该框架能够有效协调理论模型与实测真值,归一化平均绝对误差低于10%。这种可扩展的方法通过结合现实世界中有限的传感器覆盖范围和交通模型,为全网络时空AADT估计提供了一种可行的解决方案。
cs.LG / 98 / 2609.29931

Improving Calibration of Black-Box Radiology AI Using Test-Time Augmentation

利用测试时增强改进黑盒放射学AI的校准性能
Le, Nathan, Paschali, Magdalini, Koirala, Arogya, Johnston, Andrew, Fang, Zhongnan, Larson, David B., Chaudhari, Akshay S., Gonzalez, Camila
Abstract
Radiology AI systems increasingly inform clinical decisions such as triage, follow-up imaging, and treatment planning. For these decisions to be made safely, model outputs must be well calibrated, meaning predicted probabilities accurately reflect true risk. Many standard techniques for improving calibration, such as MC Dropout and Deep Ensembles, require access to model parameters or retraining. However, proprietary clinical AI systems operate as black boxes, preventing access to the model's internals. To that end, we propose a model-agnostic framework for improving calibration of black-box models using clinically grounded test-time augmentation (TTA). Our framework applies geometric and physics-inspired 3D CT perturbations and learns probability-level aggregation strategies without access to model internals or the original training data. Across pulmonary embolism and intracranial hemorrhage detection tasks, DualTTA achieved the strongest overall calibration among TTA methods, reducing the Expected Calibration Error by 54% (0.239 -> 0.109) and 43% (0.051 -> 0.029), respectively, while requiring only input-output access. Additionally, DualTTA outperformed uncertainty estimation techniques that require access to model internals, such as Temperature Scaling, MC Dropout, and Deep Ensembles, in most calibration metrics. These results demonstrate that learned TTA aggregation can improve the calibration of clinical AI systems, providing a practical approach for improving the reliability of black-box medical AI.
Chinese Translation
放射学AI系统日益影响着分诊、随访影像检查和治疗计划等临床决策。为使这些决策能够安全地作出,模型输出必须经过良好校准,即预测概率能够准确反映真实风险。许多用于改进校准的标准技术,如MC Dropout和Deep Ensembles,都需要访问模型参数或进行重新训练。然而,专有临床AI系统以黑盒形式运行,无法访问模型内部结构。为此,我们提出了一个模型无关的框架,利用基于临床的测试时增强(Test-Time Augmentation, TTA)来改进黑盒模型的校准性能。我们的框架应用几何和物理启发的3D CT扰动,并在无需访问模型内部结构或原始训练数据的情况下学习概率层面的聚合策略。在肺栓塞和颅内出血检测任务中,DualTTA在各类TTA方法中实现了最佳的整体校准性能,分别将期望校准误差(Expected Calibration Error)降低了54%(0.239 -> 0.109)和43%(0.051 -> 0.029),且仅需输入输出访问权限。此外,DualTTA在大多数校准指标上优于需要访问模型内部结构的不确定性估计技术,如Temperature Scaling、MC Dropout和Deep Ensembles。这些结果表明,基于学习的TTA聚合可以改进临床AI系统的校准性能,为提升黑盒医疗AI的可靠性提供了一种实用的方法。
cs.LG / 99 / 2609.29937

When Temporal Perturbations Act Like Sensor Biases: Label-Free Auditing of Wearable Activity Recognizers

当时间扰动如同传感器偏差:可穿戴活动识别模型的无标签审计
Wu, Qingyu, Wei, Yuan, Liu, Renju, Cheng, Hua
Abstract
Wearable human-activity recognition (HAR) models operate across sensors, subjects, and backbones, yet a smooth waveform may appear temporal while exploiting a persistent sensor offset primarily. We introduce SpectrumAudit, a label-sealed audit that fits a phase-randomized full-window stimulus on calibration windows from subjects held out from training and testing. After selection, it replays its exact DC projection and budget-constrained zero-mean residual on the same frozen victim without refitting. Across 27 victims from three datasets and three backbones, the selected waveforms cause 2.87-40.83-point three-phase robust accuracy losses. Under this replay budget, DC is more damaging than AC on 24/27 victims and recovers at least 90% of the full drop on 22/27; all 5 failures occur on WISDM. In a held-out UTD-MHAD check, the selected waveform causes 13.49-pp accuracy and 11.68-pp macro-F1 losses, versus -0.66 pp for matched random changes. The audit diagnoses offset versus zero-mean variation under a common peak-budget cap. The code will be released upon acceptance.
Chinese Translation
可穿戴人体活动识别(HAR)模型需要跨传感器、跨被试、跨骨干网络运行,然而一个看似平滑的时间波形可能主要通过利用持续存在的传感器偏置来发挥作用。我们提出SpectrumAudit,一种标签封存的审计方法:在训练与测试均未使用的被试的校准窗口上,拟合一个相位随机化的全窗口刺激。选定之后,它在同一个冻结的受害模型上精确重放其直流(DC)投影以及受预算约束的零均值残差,无需重新拟合。在来自三个数据集、三种骨干网络的27个受害模型上,所选波形导致2.87至40.83个百分点的三阶段鲁棒准确率下降。在该重放预算下,DC分量在24/27个受害模型上比AC分量更具破坏性,并在22/27个受害模型上恢复了至少90%的完整性能下降;全部5个失败案例均出现在WISDM上。在留出的UTD-MHAD验证中,所选波形导致13.49个百分点的准确率下降和11.68个百分点的宏F1下降,而匹配的随机扰动仅造成-0.66个百分点的变化。该审计方法在统一的峰值预算上限下,能够诊断偏置型与零均值型扰动。代码将在论文录用后发布。
cs.LG / 100 / 2609.29941

MF-SCBO : Multi-fidelity Scalable Constrained Bayesian Optimization

MF-SCBO:多保真度可扩展约束贝叶斯优化
Palazzolo, Lucas, Binois, Mickaël, Giraldi, Laëtitia
Abstract
Many real-world optimization problems rely on expensive simulations or experiments, making the efficient use of available data essential. Multi-fidelity optimization of high-dimensional black-box functions subject to black-box constraints is increasingly relevant as the cost of objective evaluations continues to rise in applications such as machine learning, engineering, and control. To our knowledge, no existing method simultaneously addresses high-dimensionality, black-box constraints, an arbitrary number of fidelity levels, and non-nested sampling. In this work, we extend the Scalable Constrained Bayesian Optimization method to the multi-fidelity setting, resulting in the MF-SCBO method. The proposed approach is evaluated on standard benchmark functions as well as challenging problems. The experimental results demonstrate that MF-SCBO generally achieves better convergence than both the single-fidelity SCBO and the other multi-fidelity method considered in this high-dimensional and constrained settings.
Chinese Translation
许多现实世界的优化问题依赖于昂贵的仿真或实验,因此高效利用可用数据至关重要。随着机器学习、工程和控制等应用中目标评估成本的不断上升,对受黑盒约束的高维黑盒函数进行多保真度优化变得越来越重要。据我们所知,目前尚无现有方法能够同时处理高维性、黑盒约束、任意数量的保真度层级以及非嵌套采样问题。在本工作中,我们将可扩展约束贝叶斯优化(Scalable Constrained Bayesian Optimization, SCBO)方法扩展至多保真度场景,提出了MF-SCBO方法。该方法在标准基准函数以及具有挑战性的问题上进行了评估。实验结果表明,在高维约束场景下,MF-SCBO总体上比单保真度的SCBO以及本文所考虑的其他多保真度方法取得了更好的收敛性能。
cs.LG / 101 / 2609.29945

Error- and Prediction-Driven Motor Learning in the Cortico-Cerebellar Loop

皮层-小脑环路中基于误差与预测的运动学习
Filipe, Ana Carolina, Costa, Rui Ponte, Soares, Cláudia
Abstract
Robust control under delayed sensory feedback remains a key challenge in both robotics and neuroscience. Classical cerebellar models explain delay compensation through forward prediction but fail to account for fast online corrections and rapid adaptation observed in biological systems. We propose a cerebellum-inspired control framework that combines multiplexed predictive representations with internal feedback. By jointly encoding kinematic variables and task-relevant error signals, the model enables accurate online correction despite delayed feedback. Furthermore, incorporating feedback within the cerebellar loop significantly accelerates adaptation, reducing learning time by an order of magnitude. Our results show that single-signal predictions are insufficient under delay, while multiplexing and feedback together provide a unified mechanism for online control and rapid learning.
Chinese Translation
在延迟感觉反馈下实现鲁棒控制是机器人技术和神经科学共同面临的关键挑战。经典小脑模型通过前向预测解释延迟补偿,但无法解释生物系统中观察到的快速在线校正和快速适应现象。我们提出了一种受小脑启发的控制框架,该框架将多路复用的预测表征与内部反馈相结合。通过联合编码运动学变量和任务相关的误差信号,该模型即使在反馈延迟的情况下也能实现精确的在线校正。此外,在小脑环路中引入反馈可显著加速适应过程,使学习时间缩短一个数量级。我们的结果表明,在存在延迟的情况下,单信号预测是不够的,而多路复用与反馈相结合则为在线控制和快速学习提供了一种统一机制。
cs.LG / 102 / 2609.29951

Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking

追踪状态还是追踪陪集?学习状态追踪的代数解释
Zhang, Zhiyu, Li, Yupeng
Abstract
State tracking requires composing a sequence of updates, but accuracy alone does not reveal what a model has learned. We study neural networks trained to predict the running product of group elements. We identify quotient solutions in Transformers, where models recover the quotient class while predicting nearly uniformly among its members. The reciprocal of class size predicts partial accuracy without a fitted parameter, extending parity-based accounts to non-parity quotients. Our baseline Transformers' predictions change little under prefix reordering beyond the exact-tracking frontier. We prove that, for finite groups under uniform i.i.d. full-group inputs, optimal order-blind exact accuracy converges to the reciprocal of abelianization class size as prefix length grows, consistent with the observed abelianization plateaus. Sequential updates permit more: any partition into right cosets of a subgroup, normal or not, survives sequential updates. In our census of standard Transformers, every recovered coset partition comes from a normal subgroup, whereas parameter-matched recurrent networks pass through both normal and non-normal right-coset stages during training. On $A_5$, we identify low-dimensional subspaces of the recurrent state that encode non-normal cosets. In the three-dimensional cases, coset mean vectors form approximate dodecahedra, and swapping the state components in these subspaces transfers the donor's coset state through a shared input suffix. Our results connect partial accuracy, learning stages, and internal computation through the subgroup cosets that models learn to track.
Chinese Translation
状态追踪需要对一系列更新进行复合,但仅凭准确率无法揭示模型学到了什么。我们研究了训练用于预测群元素连乘积的神经网络。我们在Transformer中识别出了商解,即模型恢复出商类,同时在其成员之间近乎均匀地进行预测。类大小的倒数可以在不拟合参数的情况下预测部分准确率,将基于奇偶性的解释扩展到了非奇偶商。我们的基线Transformer在精确追踪边界之外的条件下,其预测在词缀重排下几乎不变。我们证明,对于在均匀独立同分布的全群输入下的有限群,最优的顺序无关精确准确率随着词缀长度的增长收敛于阿贝尔化类大小的倒数,这与观察到的阿贝尔化平台现象一致。顺序更新允许更多情况:任意子群的右陪集划分(无论是否为正规子群)都能在顺序更新下保持。在我们对标准Transformer的调查中,所有被恢复的陪集划分都来自正规子群,而参数量匹配的循环神经网络在训练过程中会经历正规和非正规右陪集两个阶段。在$A_5$上,我们识别出循环状态中编码非正规陪集的低维子空间。在三维情形下,陪集均值向量构成近似十二面体,并且交换这些子空间中的状态分量可以通过共享的输入后缀将供体的陪集状态传递过去。我们的结果通过模型学习追踪的子群陪集,将部分准确率、学习阶段和内部计算联系了起来。
cs.LG / 103 / 2609.29960

Beyond Average Safety: Chance-Constrained LLM Fine-tuning

超越平均安全性:机会约束下的大语言模型微调
Entesari, Taha, Fazlyab, Mahyar
Abstract
Fine-tuning large language models on new objectives can improve helpfulness, instruction following, or domain-specific performance, but it can also induce regressions on safety-critical prompts. Existing safety-preserving fine-tuning methods typically control average safety loss or use weighted auxiliary penalties, which can obscure rare but severe failures. We propose a chance-constrained formulation for safety-preserving fine-tuning that limits the fraction of safety examples whose degradation relative to a reference model exceeds a prescribed threshold. Because the resulting empirical chance constraint contains a discontinuous indicator, we introduce a differentiable majorization of the violation rate, yielding a tractable conservative constraint. We then develop a constraint-aware gradient descent method that treats the majorized constraint as a safe set in parameter space and minimally modifies the fine-tuning direction to preserve feasibility. The resulting update admits a closed form and produces a tail-aware safety correction that emphasizes examples near or above the degradation threshold. We conduct an extensive set of experiments on harmful fine-tuning across three different tasks and three models and show that our approach consistently outperforms the baselines that exist in the literature. These results suggest that safety preservation in LLM fine-tuning is better viewed as a reliability-constrained optimization problem than as average-risk regularization.
Chinese Translation
在大语言模型(LLM)上针对新目标进行微调可以提升有用性、指令遵循能力或特定领域性能,但也可能导致模型在安全关键提示上的表现退化。现有的安全保持微调方法通常控制平均安全损失或使用加权辅助惩罚,这可能掩盖罕见但严重的失败。我们提出了一种用于安全保持微调的机会约束(chance-constrained)形式化方法,该方法限制安全样本中相对参考模型退化程度超过规定阈值的比例。由于所得到的经验机会约束包含不连续的指示函数,我们引入了一种可微的违反率上界逼近(majorization),从而得到一个易于处理的保守约束。随后,我们开发了一种约束感知的梯度下降方法,该方法将被上界化的约束视为参数空间中的安全集,并对微调方向进行最小幅度的修改以保持可行性。所得到的更新具有闭式解,并产生一种尾部感知的安全修正,其重点强调处于或超过退化阈值附近的样本。我们在三个不同任务和三个模型上针对有害微调进行了大量实验,结果表明我们的方法始终优于文献中现有的基线方法。这些结果表明,LLM 微调中的安全保持更适合被视为一个可靠性约束优化问题,而非平均风险正则化。
cs.LG / 104 / 2609.29961

A Contraction Framework for Stochastic Operators with Bootstrapping: Application to TD Learning

一种带自举的随机算子压缩框架:在TD学习中的应用
van der Werf, Ids, Rozada, Sergio, Marques, Antonio G.
Abstract
Many iterative algorithms rely on bootstrapping. A variable is updated using a second, frozen copy as a target, which is periodically replaced with the updated variable. Majorize-minimize and inexact proximal-point methods share this structure, as does temporal-difference (TD) learning. However, existing convergence guarantees for scenarios that combine sampled updates with targets refreshed only every $K$ steps rely on the specific structure of the update, such as linear approximation or gradient-based inner steps, and on uniformly bounded sampling error. We instead model the sampled update as a stochastic operator on the parameter space, which reduces the analysis to a contraction argument that needs no gradient structure and allows the sampling error to grow with the iterates. Within this framework, we derive a finite-time bound for i.i.d. samples and any target-update period $K$. We show that the iterates converge geometrically in root mean square to a ball around the fixed point, provided the sensitivity to the frozen target is smaller than the contraction slack of the inner map. Existing deterministic frozen-target contraction and stochastic-gradient-type bounds follow as special cases of our framework, and simulations of TD learning reproduce the predicted contraction rate and scaling of the error floor with the step size.
Chinese Translation
许多迭代算法依赖于自举(bootstrapping):变量使用其冻结副本作为目标进行更新,而该冻结副本会周期性地被更新后的变量所替换。主化-最小化(majorize-minimize)方法和非精确近端点方法均具有这种结构,时序差分(TD)学习亦是如此。然而,针对采样更新与每 $K$ 步才刷新一次的目标相结合的场景,现有的收敛性保证依赖于更新的特定结构,例如线性近似或基于梯度的内层步骤,并要求采样误差一致有界。我们转而将采样更新建模为参数空间上的一个随机算子,从而将分析归结为一个压缩论证,该论证既不需要梯度结构,也允许采样误差随迭代变量增长。在此框架下,我们针对独立同分布样本及任意目标更新周期 $K$ 推导出了有限时间界。我们证明,只要对冻结目标的敏感度小于内层映射的压缩余量,迭代变量就会在均方根意义下几何收敛到不动点周围的一个球域。现有的确定性冻结目标压缩界和随机梯度型界均可作为我们框架的特例,并且TD学习的仿真实验复现了所预测的压缩速率以及误差下限随步长的缩放关系。
cs.LG / 105 / 2609.29974

Diverse Geometries, Frozen Weights: Robust Heterogeneous Treatment-Effect Estimation via Causal Expert Ensembles

多样几何结构,冻结权重:基于因果专家集成(Causal Expert Ensembles)的稳健异质性处理效应估计
Jahromi, Ali Haghpanah, Taheri, Mohammad
Abstract
Estimating heterogeneous treatment effects from observational data is difficult because the most appropriate inductive bias varies with overlap, treatment imbalance, prognostic structure, and sample size. We introduce the Geometry-Diverse Anchor-Correction Expert Ensemble (GeoACE), a five-expert framework that combines a common anchor-correction estimator with complementary overlap-aware and outcome-guided geometries. Its task-level ensemble weights are learned only from internal validation predictions, frozen before test evaluation, and then applied to experts refitted on the complete development sample. The fifth expert, O-Phi-ACE, constructs an outcome-free, overlap-aware statistical projection from covariates and treatment assignment and replaces the anchor input with this lower-dimensional geometry. We evaluate GeoACE against 11 comparators on eight benchmark protocols. Adding O-Phi-ACE reduced mean sqrt(PEHE) relative to the four-expert ensemble on all seven benchmarks with individual-effect truth, winning 998 of 1,225 paired tasks; the change on JOBS policy risk was negligible. The five-expert ensemble ranked first on IHDP100, IHDPA, and IHDPB and second on NEWS, differing from the NEWS leader by 0.13%. Across the seven sqrt(PEHE) benchmarks it obtained the lowest observed average rank (3.714), although the omnibus Friedman and Iman-Davenport tests were not significant (p=0.328 and p=0.330). Using the same five frozen experts, inverse-DR weighting was consistently better than winner-take-all selection, convex DR fitting, R-stacking, and causal Q-aggregation in benchmark-balanced analyses, but was statistically indistinguishable from equal weighting and DR ridge shrinkage. The evidence therefore supports geometry-diverse expert libraries and leakage-free aggregation as a robustness strategy, not universal superiority of either GeoACE or one weighting rule.
Chinese Translation
从观测数据中估计异质性处理效应十分困难,因为最合适的归纳偏置会随重叠程度、处理不平衡性、预后结构以及样本量而变化。我们提出了几何多样锚校正专家集成(Geometry-Diverse Anchor-Correction Expert Ensemble,GeoACE),这是一个由五个专家组成的框架,将一个公共的锚校正估计器与互补的重叠感知几何结构和结果导向几何结构相结合。其任务级集成权重仅从内部验证预测中学习,在测试评估之前被冻结,然后应用于在完整开发样本上重新拟合的专家。第五个专家O-Phi-ACE基于协变量和处理分配构建了一个与结果无关、重叠感知的统计投影,并用这一低维几何结构取代锚输入。我们在八个基准协议上将GeoACE与11个对比方法进行了评估。在所有七个具有个体效应真值的基准上,添加O-Phi-ACE使五专家集成相对于四专家集成的平均sqrt(PEHE)均有所降低,在1,225个配对任务中赢了998个;而在JOBS政策风险上的变化可以忽略不计。五专家集成在IHDP100、IHDPA和IHDPB上排名第一,在NEWS上排名第二,与NEWS上第一名的差距仅为0.13%。在七个sqrt(PEHE)基准上,它获得了最低的观测平均排名(3.714),尽管综合Friedman检验和Iman-Davenport检验均不显著(p=0.328和p=0.330)。使用相同的五个冻结专家,在基准平衡分析中,逆DR加权始终优于胜者全取选择、凸DR拟合、R-stacking和因果Q聚合,但在统计上与等权重和DR岭收缩无显著差异。因此,这些证据支持将几何多样的专家库和无泄漏聚合作为一种稳健性策略,而非证明GeoACE或某一加权规则具有普遍优越性。
cs.LG / 106 / 2609.29988

Let Training Guide Selection: Online Synthetic Data Filtering via Real-Anchored Utility

让训练引导选择:基于真实锚定效用的在线合成数据过滤
Wu, Yanran, Lakdawala, Sana, Miller, Renzo Tassara, Bai, Chongyang, Ciddu, Sharath, Singh, Shivendra Pratap, Li, Kungang, Pandey, Sandeep, Liu, Chunwei
Abstract
Synthetic data can scale training supervision when real-world data are limited, but noise and distribution mismatch can reduce its value. Existing synthetic data selection methods often emphasize fidelity or diversity rather than the learner's evolving needs. We propose FROST, an online framework that estimates synthetic-data utility through gradient feedback anchored in real training data. It calibrates batch utility against recent history to determine when filtering is needed and filters samples only in out-of-band batches to determine what to retain, without an external verifier or held-out validation set. Experiments on two public benchmarks for image classification and LLM fine-tuning for text-to-SQL show that FROST filters out around 20--30% of the synthetic data while improving real-task performance compared with training on the full synthetic data pool. We further apply FROST during training in a large-scale industrial ads re-ranking system, achieving significant performance gains over a highly optimized production baseline, demonstrating its effectiveness and generalizability.
Chinese Translation
当真实世界数据有限时,合成数据可以扩展训练监督的规模,但噪声和分布失配可能降低其价值。现有的合成数据选择方法通常强调保真度或多样性,而忽略了学习器不断变化的需求。我们提出了 FROST,这是一个在线框架,通过以真实训练数据为锚定的梯度反馈来估计合成数据的效用。该框架根据近期历史对批次效用进行校准,以确定何时需要过滤,并仅在带外批次中过滤样本以决定保留哪些内容,无需外部验证器或留出验证集。在图像分类的两个公开基准以及面向文本到 SQL 的 LLM 微调任务上的实验表明,与使用完整合成数据池训练相比,FROST 过滤掉了约 20%–30% 的合成数据,同时提升了真实任务性能。我们进一步将 FROST 应用于大规模工业广告重排序系统的训练过程中,在高度优化的生产基线之上取得了显著的性能提升,证明了其有效性和可泛化性。
cs.LG / 107 / 2609.30017

Canopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits

Canopy:利用分段平滑树先验的多保真度老虎机
Jerge, Michael, Jana, Suman
Abstract
Many LLM inference problems, including model routing, prefix-cache management, prompt trimming, and test-time search, can be viewed as optimization over a tree. This structure arises naturally from autoregressive generation: every prefix defines a node, and its continuations form a subtree below it. Internal nodes of the tree provide cheap but biased estimates of a region's value, while leaf evaluations are expensive but accurate. Hierarchical bandit methods can exploit this structure, but typically require a specific smoothness schedule to be specified in advance, even though real objectives are often only piecewise smooth and their optima may lie near sharp boundaries. We introduce CANOPY, a multi-fidelity tree bandit that learns where the smoothness prior is valid rather than assuming it globally. CANOPY uses cheap random-path probes to construct an online certificate of local aggregation bias, then directs expensive leaf evaluations toward cells where the certificate detects a smoothness violation. We prove fixed-budget and regret guarantees whose additional cost is additive in the number of discontinuities, recovering the smooth-tree rate when no violations are present and approaching structure-blind search as violations become dense. Across routing, top-$k$ identification, test-time search, caching, and prompt trimming, CANOPY consistently improves matched-budget performance, including $2.9\times$ higher top-10 recall on a 1000-model pool, $1.6\times$ more SWE-bench Verified issues resolved than best-of-$N$, and $3.6\times$ lower median time-to-first-token with prefix caching.
Chinese Translation
许多大语言模型(LLM)推理问题,包括模型路由、前缀缓存管理、提示词裁剪和测试时搜索,都可以被视为对树结构的优化。这一结构自然源于自回归生成:每个前缀定义一个节点,其后续延续构成该节点下方的子树。树的内部节点提供廉价但有偏的区域价值估计,而叶子节点评估昂贵但准确。分层老虎机方法可以利用这一结构,但通常需要预先指定特定的平滑性调度,尽管真实目标函数往往只是分段平滑的,且其最优解可能位于剧烈变化的边界附近。我们提出CANOPY,一种多保真度树状老虎机算法,它学习平滑性先验在何处有效,而非全局假设其成立。CANOPY使用廉价的随机路径探测来构建局部聚合偏差的在线证书,然后将昂贵的叶子评估引导至证书检测到平滑性违背的单元。我们证明了固定预算和后悔(regret)保证,其额外开销与不连续点的数量呈可加关系:当无不连续点时恢复平滑树的速率,当违背密集时则趋近于结构盲搜索。在路由、top-$k$识别、测试时搜索、缓存和提示词裁剪等任务上,CANOPY在相同预算下持续提升性能,包括在1000个模型的池上将top-10召回率提高$2.9\times$,在SWE-bench Verified上解决的问题数量比best-of-$N$多$1.6\times$,以及在结合前缀缓存时将首个token的中位等待时间降低$3.6\times$。
cs.LG / 108 / 2609.30036

Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think

瞄准近处以致远行:你的冻结世界模型比想象中更擅长规划
Liu, Xvyuan, Fang, Jianjie, Gao, Chen, Li, Yong
Abstract
Planners built on visual world models commonly score each predicted outcome by its distance to the encoded goal image. We show that this target can limit control even with exact dynamics and globally optimal short-horizon search: reaching a goal may require actions that initially move away from it. With frozen LeWM models, intermediate targets substantially improve action synthesis and recorded-action ranking on Cube, PushT, Reacher, and TwoRoom. Learned targets and targets drawn from observed experience both produce these gains. We introduce Anchored Planning, which retrieves a recorded segment whose start and end resemble the current and goal observations, then aims at an observation shortly after its start. The frozen model scores actions toward this target from the current state. Without additional training, planning toward observed targets outperforms the released LeWM planner on every task in our long-range evaluation. Additional final-goal search falls short of the same gains. Lower successor-prediction error need not translate into better control. Success also depends on how far ahead the target is placed and on shrinking the retrieval span as execution advances. Changing only the target lets the same frozen model and planner reach goals that final-goal scoring misses.
Chinese Translation
基于视觉世界模型构建的规划器通常以预测结果与编码目标图像之间的距离来评估其优劣。我们证明,即使动力学完全精确且短时程搜索全局最优,这一目标设定仍可能限制控制效果:到达目标可能需要最初远离目标的动作。使用冻结的 LeWM 模型,在 Cube、PushT、Reacher 和 TwoRoom 任务上,中间目标显著提升了动作合成和已记录动作排序的性能。学习得到的目标以及从观测经验中提取的目标均能带来这些提升。我们提出了锚定规划(Anchored Planning),它检索一段起始和结束分别与当前观测和目标观测相似的已记录片段,然后将瞄准点设为该片段起始后不久的观测。冻结模型据此从当前状态评估指向该目标的动作。无需额外训练,在长程评估的所有任务上,以观测目标为导向的规划均优于已发布的 LeWM 规划器。额外的最终目标搜索也无法达到同样的提升。更低的后续状态预测误差并不必然带来更好的控制。成功还取决于目标设定的远近,以及随执行推进不断缩小检索范围。仅改变目标,同一冻结模型和规划器就能到达最终目标评分方法无法企及的目标。
cs.LG / 109 / 2609.30079

Reachability-Based Formal Verification of Graph Neural Networks with Node and Edge Features

基于可达性的图神经网络形式化验证(含节点与边特征)
Tumlin, Anne M., Wooding, Ben, Shao, Zhenxuan, Lopez, Diego Manzanas, Derr, Tyler, Johnson, Taylor T.
Abstract
Graph neural networks (GNNs) have become a prominent approach for developing fast, topology-aware surrogates in electric power systems, supporting tasks such as power flow (PF) analysis, optimal power flow (OPF) estimation, and cascading failure analysis (CFA). Despite this growing use, formally verifying GNN-based models remains challenging, with existing methods limited in scope. We extend the neural network verification (NNV) framework to graph-structured inputs through GraphStar sets, a generalization of Star sets that captures uncertainty over both node and edge features. This extension enables the propagation of linear message-passing operations and the sound approximation of ReLU nonlinearities for GNN architectures, including graph convolutional network (GCN) and graph isomorphism network with edge features (GINE) layers. We evaluate GNNV across three power system tasks, PF, OPF, and CFA, on the IEEE-24, IEEE-39, and IEEE-118 test cases, as well as two standard graph classification benchmarks, ENZYMES and PROTEINS. Our results show that GNNV provides tighter robustness guarantees than CORA on graph classification models with ReLU-based activations and, for the first time, delivers edge-aware robustness guarantees for GINE-based PF and OPF models under joint node and edge perturbations.
Chinese Translation
图神经网络(GNN)已成为在电力系统中开发快速、拓扑感知代理模型的重要方法,可支持潮流分析(PF)、最优潮流估计(OPF)以及连锁故障分析(CFA)等任务。尽管应用日益广泛,对基于GNN的模型进行形式化验证仍然极具挑战性,且现有方法的适用范围有限。我们通过GraphStar集合扩展了神经网络验证(NNV)框架。GraphStar集合是Star集合的推广,能够同时刻画节点特征和边特征上的不确定性。该扩展使得线性消息传递操作可以传播,并能对GNN架构(包括图卷积网络(GCN)和带边特征的图同构网络(GINE)层)中的ReLU非线性进行可靠近似。我们在IEEE-24、IEEE-39和IEEE-118测试系统上,针对PF、OPF和CFA三类电力系统任务,以及ENZYMES和PROTEINS两个标准图分类基准,对GNNV进行了评估。结果表明,在采用ReLU激活函数的图分类模型上,GNNV提供了比CORA更紧的鲁棒性保证,并首次为基于GINE的PF和OPF模型在节点与边联合扰动下提供了边感知的鲁棒性保证。
cs.LG / 110 / 2609.30085

Residual Correlation as a Diagnostic for Joint-Uncertainty Gains from GP Coregionalisation

残差相关性作为GP协同区域化联合不确定性增益的诊断工具
Zhou, Fangqin, Vanschoren, Joaquin
Abstract
In multi-target regression, correlated targets are often coupled through multi-output Gaussian processes with an intrinsic model of coregionalisation (GP-ICM), assuming that sharing statistical strength improves overall performance. In practice, the benefits are inconsistent. Across the settings studied, we find that the main benefit of coregionalisation is joint uncertainty quantification rather than point prediction. Raw target correlation does not predict when coupling helps; in the separable GP-ICM settings studied here, residual correlation, the cross-target dependence left unexplained by independent per-target predictors, is the strongest predictor of joint-uncertainty gains. We introduce a lightweight diagnostic, $D_{\rm logdet}=-\frac{1}{2}\log\det R_{\rm res}$, which represents the idealised joint negative log-likelihood (NLL) gain from modelling a full rather than diagonal residual covariance and is computable from independent GPs alone. Across a controlled synthetic study, 16 multi-target benchmarks, and frozen transformer and convolutional neural network representations for keypoint regression, point prediction remains largely unchanged ($\Delta R^2\approx 0$). In contrast, $D_{\rm logdet}$ strongly predicts observed ICM NLL improvements ($\rho_s=-0.83$, $p<0.001$), outperforming heuristics such as the feature-to-sample ratio. We also propose Residual-ICM, which preserves independent marginal variances while adding residual-correlation structure to the joint covariance. Residual-ICM achieves the best average joint NLL among the compared methods, while the diagnostic indicates when covariance coupling is likely to be useful. The diagnostic is specific to global Gaussian residual dependence, the structure captured by separable coregionalisation.
Chinese Translation
在多目标回归中,相关目标通常通过带内在协同区域化模型(GP-ICM)的多输出高斯过程进行耦合,其假设是共享统计强度能提升整体性能。但在实践中,其收益并不一致。在所研究的各种设置中,我们发现协同区域化的主要收益在于联合不确定性量化,而非点预测。原始目标相关性并不能预测耦合何时有帮助;在本文研究的可分离GP-ICM设置中,残差相关性——即独立单目标预测器无法解释的跨目标依赖——是联合不确定性增益的最强预测因子。我们提出了一个轻量级诊断量 $D_{\rm logdet}=-\frac{1}{2}\log\det R_{\rm res}$,它表示通过建模完整的而非对角的残差协方差所获得的理想化联合负对数似然(NLL)增益,且仅可由独立的高斯过程计算得出。在受控合成实验、16个多目标基准数据集,以及基于冻结Transformer和卷积神经网络表示的关键点回归任务中,点预测基本保持不变($\Delta R^2\approx 0$)。相比之下,$D_{\rm logdet}$ 能强烈预测观测到的ICM NLL改进($\rho_s=-0.83$,$p<0.001$),优于诸如特征-样本比等启发式方法。我们还提出了Residual-ICM,它在保持独立边缘方差的同时,向联合协方差中引入残差相关性结构。Residual-ICM在所比较的方法中取得了最佳的联合NLL平均值,而该诊断量则可以指示协方差耦合何时可能有用。该诊断量专门针对全局高斯残差依赖,即可分离协同区域化所捕获的结构。
cs.LG / 111 / 2609.30088

AT-SKM-Net: An Accelerated Trainable Sampling Kaczmarz-Motzkin Framework for Linear Hard-Constraint Feasibility on Dynamic Graphs

AT-SKM-Net:面向动态图上线性硬约束可行性的加速可训练采样Kaczmarz-Motzkin框架
Zhang, Xiaochen, Zhu, Haoyu, Zhang, Yao, Hou, Qingchun
Abstract
Graph-structured optimization with linear constraints is fundamental to critical infrastructure but faces scalability limits due to massive strict hard constraints and high dimensionality. While recent projection-based methods such as Trainable Sampling Kaczmarz-Motzkin Net (T-SKM-Net) guarantee feasibility, they face high computational costs in dynamic environments by processing the entire constraint set and requiring expensive matrix factorizations. To bridge this gap, we propose the Accelerated Trainable-SKM (AT-SKM) Net framework. To concentrate computation on the active constraints and eliminate redundant calculations, we introduce a hybrid sampling strategy guided by a topology-aware heterogeneous GNN model. To efficiently handle topological shifts in graph-based constraints, we employ a Cholesky Update mechanism that theoretically reduces the equality projection complexity from O(N^3) to O(N^2) under low-rank perturbations. Experiments on random geometric graphs, N-1 Security-Constrained DC-OPF, and minimum-cost gas transport problem demonstrate that AT-SKM reduces iteration counts by up to 85% and achieves 2.95x-7.29x SKM layer speedups, while maintaining zero constraint violations.
Chinese Translation
带线性约束的图结构优化对关键基础设施至关重要,但由于海量严格硬约束和高维性,其面临可扩展性瓶颈。尽管近期基于投影的方法(如可训练采样Kaczmarz-Motzkin网络 T-SKM-Net)能够保证可行性,但由于需要处理整个约束集并进行代价高昂的矩阵分解,其在动态环境中的计算成本很高。为弥合这一差距,我们提出了加速可训练SKM(AT-SKM)网络框架。为将计算集中于活跃约束并消除冗余计算,我们引入了一种由拓扑感知异构GNN模型引导的混合采样策略。为高效处理图约束的拓扑变化,我们采用Cholesky更新机制,在理论上将低秩扰动下等式投影的复杂度从O(N^3)降至O(N^2)。在随机几何图、N-1安全约束直流最优潮流(DC-OPF)以及最小成本天然气输运问题上的实验表明,AT-SKM可将迭代次数减少至多85%,实现2.95倍至7.29倍的SKM层加速,同时保持零约束违反。
cs.LG / 112 / 2609.30105

On the SoS Certifiability of Log-Concave Distributions

论对数凹分布的SoS可认证性
Storozhenko, Aleksandr
Abstract
For an arbitrary isotropic log-concave distribution $P$ on $\mathbb{R}^d$, we prove that the polynomial $(Cm)^m\|v\|_2^m - \mathbb{E}_{X\sim P}\langle X,v\rangle^m$ is a sum of squares for every even $m\ge2$, where $C>0$ is a universal constant. This removes the dependence on the Poincar\'e constant in the theorem of Kothari and Steinhardt (arXiv:1711.07465), recovering the optimal moment bounds for log-concave distributions. As an immediate corollary, we obtain computationally efficient algorithms with dimension-free error guarantees for a wide range of high-dimensional statistical estimation problems. Our proof uses stochastic localization to decompose $P$ as an average of random strongly log-concave measures, whose centered moments admit the subgaussian certificates of Diakonikolas, Hopkins, Pensia, and Tiegel (STOC 2025; arXiv:2410.21194). With a covariance-adapted choice of localization, we show that a fourth-moment certificate derived from Letwin's variance inequality for quadratic forms (arXiv:2607.24164) suffices to control this averaging at every even degree.
Chinese Translation
对于$\mathbb{R}^d$上任意各向同性的对数凹分布$P$,我们证明了对每个偶数$m\ge2$,多项式$(Cm)^m\|v\|_2^m - \mathbb{E}_{X\sim P}\langle X,v angle^m$均为平方和(sum of squares),其中$C>0$为一个通用常数。这一结果消除了Kothari和Steinhardt(arXiv:1711.07465)定理中对Poincaré常数的依赖,从而恢复了对数凹分布的最优矩界。作为一个直接推论,我们为一大类高维统计估计问题得到了具有与维度无关误差保证的计算高效算法。我们的证明利用随机局部化(stochastic localization)将$P$分解为随机强对数凹测度的平均值,其中心化矩可由Diakonikolas、Hopkins、Pensia和Tiegel(STOC 2025;arXiv:2410.21194)的次高斯(subgaussian)证书所刻画。通过一种与协方差相适应的局部化选择,我们证明由Letwin关于二次型的方差不等式(arXiv:2607.24164)导出的四阶矩证书足以在每个偶数阶上控制这种平均过程。
cs.LG / 113 / 2609.30150

Graph-Based Inference and Topology-Aware Multi-Agent Reinforcement Learning for Large-Scale Railway Network Management

基于图推理与拓扑感知多智能体强化学习的大规模铁路网络管理
Arcieri, Giacomo, Duthé, Gregory, Muller, Christophe, Papakonstantinou, Konstantinos G., Straub, Daniel, Chatzi, Eleni
Abstract
Modern infrastructure asset management constitutes a complex sequential decision-making problem, characterized by long planning horizons and system-level interactions, such as spatial deterioration correlations and economies of scale. While deep reinforcement learning has shown promise in optimizing maintenance policies, scaling to real-world networks remains challenging. Centralized approaches become computationally intractable in large-scale systems, whereas decentralized approaches often fail to capture essential coordination mechanisms. To address these challenges, we propose a graph-based framework that integrates accurate environment modeling with scalable decision support. First, we employ a hierarchical Bayesian model leveraging a Gaussian Process on Graph kernel to infer a realistic, spatially correlated networked environment of railway maintenance planning from real-world data provided by the Swiss Federal Railways. Second, we introduce a topology-aware Multi-Agent Reinforcement Learning (MARL) framework by integrating graph neural networks and graph Transformers to optimize network-level policies. A central contribution of this work is the demonstration of scalability through zero-shot transfer learning: graph-based agents, trained only on small network portions, are successfully deployed in a zero-shot manner on large-scale unseen networks without any retraining. Numerical results indicate that the proposed method significantly outperforms optimized heuristics and standard MARL baselines, reducing computational training time while maintaining superior performance on large-scale networks.
Chinese Translation
现代基础设施资产管理是一个复杂的序贯决策问题,其特点在于规划周期长以及系统层面的相互作用,例如空间上的劣化相关性和规模经济。尽管深度强化学习在优化维护策略方面已展现出潜力,但将其扩展至现实世界的网络仍具有挑战性。集中式方法在大规模系统中计算上难以处理,而分散式方法往往无法捕捉关键的协调机制。为应对这些挑战,我们提出了一个基于图的框架,将精确的环境建模与可扩展的决策支持相结合。首先,我们采用层级贝叶斯模型,利用图上高斯过程(Gaussian Process on Graph)核,从瑞士联邦铁路(Swiss Federal Railways)提供的真实数据中推断出符合实际且具有空间相关性的铁路维护规划网络化环境。其次,我们通过集成图神经网络和图Transformer,引入了一个拓扑感知的多智能体强化学习(Multi-Agent Reinforcement Learning, MARL)框架,以优化网络级策略。本工作的一个核心贡献是通过零样本迁移学习展示了可扩展性:仅在小型网络片段上训练的基于图的智能体,能够以零样本的方式成功部署于大规模未见过的网络,而无需任何重新训练。数值结果表明,所提出的方法显著优于优化的启发式方法和标准MARL基线,在缩短计算训练时间的同时,在大规模网络上保持了更优的性能。
cs.LG / 114 / 2609.30185

Intrinsic-Extrinsic Coupling in Learning Dynamics

学习动力学中的内在-外在耦合
Wang, Qinyou
Abstract
A learner's current observations need not determine its response to further training. We formulate intrinsic-extrinsic coupling through the continuation-conditioned value of a constrained learning-state intervention, with observation-relative fibers describing present agreement. An executable finite-frame classifier-head write protects current logits while repairing specified historical margins under finite-precision acceptance checks. We distinguish local admissibility, continuation-conditioned intervention value, and complete-policy performance. A matched four-cell contrast identifies readout-specific non-additivity between the same intrinsic intervention and alternative external continuations. In a CLINC-derived class-incremental setting, replay changes the write's 32-update contribution from five correct predictions to zero. Nonzero interactions also occur under output distillation, with a RoBERTa backbone, and under optimizer-native SGDW dynamics. Under SGDW, correct-count interactions are negative in all three activated roots at 128 updates, showing that coupling need not imply positive synergy. The mathematical analysis distinguishes feasible local repairs and favorable terminal outputs from training-reachable repair regions. Separate coordination tests show that content controls match or exceed the development gain, while a five-root fresh-test comparison with Fiber present in every arm shows root-dependent rather than uniformly beneficial correct-count effects. On the secondary cross-entropy readout, guided allocation yields lower mean loss than standard replay in all five pairs. Together, these results make intrinsic-extrinsic coupling operational by connecting executable state geometry to continuation-conditioned value, matched interaction identification, and closed-loop coordination, while separating identified coupling from complete-policy performance.
Chinese Translation
学习者的当前观察结果并不必然决定其对后续训练的响应。我们通过受约束学习状态干预的“延续条件化价值”(continuation-conditioned value)来形式化内在-外在耦合,并以相对于观察结果的纤维(fibers)刻画当前的一致性。一个可执行的有限框架分类头写入(finite-frame classifier-head write)在有限精度接受性检查下,保护当前logits同时修复指定的历史边际。我们区分了局部可容许性、延续条件化干预价值以及完整策略性能。一个匹配的四单元对比实验识别出相同内在干预与不同外部延续之间的读出特异性不可加性。在一个源自CLINC的类增量设置中,重放(replay)使该写入在32次更新内的贡献从五个正确预测变为零。非零交互也出现在输出蒸馏、RoBERTa骨干网络以及优化器原生的SGDW动力学条件下。在SGDW下,128次更新时三个激活根的正确数交互均为负值,表明耦合并不必然意味着正向协同。数学分析区分了可行的局部修复、有利的终端输出与训练可达的修复区域。单独的协调测试表明,内容控制(content controls)能够达到或超过开发集增益;而在每个实验组均包含Fiber的五根新测试对比中,正确数效应依赖于具体根而非普遍有益。在次级交叉熵读出上,引导式分配在全部五对比较中均取得低于标准重放的平均损失。总之,这些结果通过将可执行的状态几何与延续条件化价值、匹配交互识别以及闭环协调相连接,使内在-外在耦合变得可操作化,同时将所识别的耦合与完整策略性能区分开来。
cs.LG / 115 / 2609.30198

Beyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers

超越压缩:面向神经代理求解器稳定长时程演化的潜在表示训练
Robertson, Andreas E., Lenau, Ashley T., Shimanek, John D., Jasperson, Benjamin A., Oommen, Vivek, Damm, David L., Garikipati, Krishna, Dingreville, Remi
Abstract
Latent neural surrogate solvers, or latent dynamics models, accelerate simulations of time-dependent physical systems by evolving a compressed latent space rather than resolving full-resolution fields directly. In principle this reduces computational cost and simplifies learning, but in practice errors often accumulate rapidly during long autoregressive rollouts, limiting predictive utility. We show that this instability does not stem from the latent representation itself, but arises when it is trained solely for reconstruction, producing representations poorly suited to long-horizon forecasting. We systematically evaluate training-level interventions that align latent representations with long-horizon rollout: Koopman operator learning and Hamming noise injection during autoencoder training to improve compression, together with noise injection and multi-step rollout fine-tuning to improve dynamics. Interventions that improve long-horizon rollout stability often degrade conventional training metrics, including reconstruction and one-step prediction accuracy. Collectively, these interventions reduce long-rollout error by approximately 40\% and match or exceed the accuracy of full-resolution models on two physics benchmarks, while requiring 2 orders of magnitude fewer floating point operations and half the GPU memory. Applied to mesoscale crystal-plasticity simulations of high-cycle fatigue, the resulting surrogate achieves stable extrapolation over horizons orders of magnitude beyond those observed during training. More broadly, these results show that neural compression should be designed not merely to reduce dimensionality, but to restructure the solution space for stable dynamical evolution, a key requirement for reliable, efficient neural surrogates in scientific applications.
Chinese Translation
潜在神经代理求解器(或称潜在动力学模型)通过在压缩的潜在空间中演化,而非直接求解全分辨率场,来加速时变物理系统的仿真。原则上,这可以降低计算成本并简化学习,但在实际应用中,误差往往会在长自回归滚动推演过程中快速累积,限制了其预测能力。我们证明,这种不稳定性并非源于潜在表示本身,而是当其仅以重建为目标进行训练时产生——由此得到的表示不适用于长时程预测。我们系统评估了使潜在表示与长时程滚动推演相适配的训练层面的干预措施:包括在自编码器训练中引入Koopman算子学习和Hamming噪声注入以改进压缩,以及通过噪声注入和多步滚动推演微调来改进动力学。这些能提升长时程滚动稳定性的干预措施往往会降低常规训练指标,包括重建精度和单步预测精度。总体而言,这些干预措施将长程滚动误差降低约40%,在两个物理基准上达到或超越全分辨率模型的精度,同时所需浮点运算量降低两个数量级,GPU显存占用减少一半。将该代理模型应用于高周疲劳的介观尺度晶体塑性仿真,其实现了远超训练观测范围数个数量级时程的稳定外推。更广泛地说,这些结果表明,神经压缩的设计不应仅仅为了降维,而应重构解空间以实现稳定的动力学演化——这是在科学应用中构建可靠、高效神经代理模型的关键要求。
cs.LG / 116 / 2609.30218

Minimally Invasive Steering of Language Models

语言模型的最小侵入式引导
Entesari, Taha, Zhang, Jingyu, Khashabi, Daniel, Fazlyab, Mahyar
Abstract
Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states. Unregularized reward optimization can substantially alter the output distribution and degrade generation quality. We propose Minimally Invasive Steering Vector Optimization (MISVO), which penalizes interventions using the local KL geometry of the induced token distribution. The resulting Fisher quadratic measures distributional sensitivity and admits an analytic gradient computed through matrix--vector products with the frozen language-model head. We derive an exact decomposition of the sequence-level KL gradient into an analytic Fisher term and a suffix score-function term. For a fixed generation horizon, we show that the suffix term is second order in the steering magnitude and that three Fisher surrogates agree with the full KL gradient to first order. MISVO uses the frozen-reference surrogate to optimize position-specific interventions without updating model parameters. Across preference and code-generation tasks on models with approximately 1B--14B parameters, MISVO achieves the highest mean reward in six of seven model--task settings, with diversity and coherence scores close to those of Best-of-N.
Chinese Translation
预logit(pre-logit)引导通过在冻结语言模型的最终隐藏状态上添加向量,使其在测试时适应某个奖励信号。未加正则化的奖励优化可能大幅改变输出分布并降低生成质量。我们提出最小侵入式引导向量优化(Minimally Invasive Steering Vector Optimization, MISVO),它利用诱导出的词元分布的局部KL几何结构对干预施加惩罚。所得的Fisher二次型度量了分布敏感性,并可通过与冻结的语言模型输出头进行矩阵-向量乘积来解析地计算其梯度。我们推导了序列级KL梯度的精确分解,将其分解为一个解析的Fisher项和一个后缀得分函数项。对于固定的生成步数,我们证明后缀项关于引导幅值是二阶的,且三种Fisher代理都与完整KL梯度一阶一致。MISVO使用冻结参考代理来优化特定位置的干预,而无需更新模型参数。在约1B至14B参数规模的模型上,跨偏好任务和代码生成任务,MISVO在七个模型-任务设置中的六个取得了最高的平均奖励,其多样性和连贯性得分接近Best-of-N方法。
cs.LG / 117 / 2609.30226

PoEM: Predicting RL Outcomes from Existing Policies

PoEM:基于已有策略预测强化学习结果
Hamidieh, Kimia, Daras, Giannis, Torralba, Antonio
Abstract
Foundation models are post-trained with reinforcement learning (RL) to maximize specific rewards, such as human alignment, correctness, or instruction following. This post-training process is computationally intensive, sometimes unstable, and has to be run from scratch every time the reward model changes or when we want to combine multiple rewards. We hence ask: given a new reward function, is it possible to predict the RL outcomes without actually running RL on it? We answer this in the affirmative by introducing PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards. First, we show that if the new reward function can be written as a linear combination of existing ones, then the new policy in log-space can be written as a linear combination of the existing log-policies. Surprisingly, even in cases where the rewards are not linearly connected, we observe that often log-policies from RL training span an approximately low-rank subspace across rewards. To our benefit, the weighting coefficients for this combination can be estimated using only the reward or basis policy outputs on the samples. We turn these observations into an algorithm that takes post-trained models and a new reward function, and approximates the target RL policy without actually running any additional RL training. We experimentally validate our approach across synthetic and real rewards, spanning both text and image modalities.
Chinese Translation
基础模型通过强化学习(RL)进行后训练,以最大化特定奖励,例如人类对齐、正确性或指令遵循。这一后训练过程计算开销巨大,有时不稳定,并且每次奖励模型发生变化或需要组合多个奖励时,都必须从头开始运行。因此我们提出这样的问题:给定一个新的奖励函数,是否可以在不实际运行强化学习的情况下预测其RL结果?我们通过引入PoEM对此给出了肯定回答,该框架利用一组已在其他奖励上完成后训练的模型,来预测强化学习在新奖励函数上的输出。首先,我们证明如果新的奖励函数可以表示为已有奖励函数的线性组合,那么新策略在对数空间中也可以表示为已有对数策略的线性组合。令人惊讶的是,即使在奖励之间不存在线性关系的情况下,我们也观察到,RL训练得到的对数策略在不同奖励之间往往张成一个近似低秩的子空间。对我们有利的是,该组合的权重系数可以仅通过奖励或基础策略在样本上的输出进行估计。我们将这些观察结果转化为一种算法:输入后训练模型和一个新的奖励函数,即可在不实际运行任何额外RL训练的情况下逼近目标RL策略。我们在合成奖励和真实奖励上,跨越文本和图像两种模态,对所提方法进行了实验验证。
cs.LG / 118 / 2609.30227

To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech

信任还是不信任:语音中的检索增强事实核查
Mazumder, Debajyoti, Mamta, Penamakuri, Abhirama Subramanyam
Abstract
Online misinformation increasingly appears in spoken formats such as news clips, podcasts, interviews, political speeches, and social media videos, creating a need for fact-checking systems that can verify claims directly from speech. We introduce VeriSpeak, a probe benchmark for studying speech-based fact verification in Large Audio Language Models (LALMs). VeriSpeak contains 3,879 spoken claims spanning temporal, geographical, and relational facts, with balanced true and false labels. The benchmark is designed to examine whether factual verification ability transfers from text to speech, and whether retrieval-augmented LALMs can use textual evidence to correctly support or refute spoken claims. Our experiments reveal a consistent text-speech modality gap: LALMs that verify written claims reliably often fail on the same claims when spoken. Moreover, retrieval alone provides limited gains because models frequently conflate retrieved evidence with the spoken claim. In contrast, retrieval combined with explicit reasoning improves claim-evidence comparison, with a thinking-tuned LALM reaching 86.1% accuracy. VeriSpeak highlights that effective speech misinformation detection requires not only speech understanding, but also grounded reasoning over retrieved evidence. The dataset is publicly available via Hugging Face at https://huggingface.co/datasets/abhiram4572/VeriSpeak.
Chinese Translation
网络错误信息日益以新闻片段、播客、访谈、政治演讲和社交媒体视频等语音形式出现,因此需要能够直接从语音中核查声明的系统。我们提出了 VeriSpeak,一个用于研究大型音频语言模型(LALM)中基于语音的事实核查的探测基准。VeriSpeak 包含 3,879 条语音声明,涵盖时间、地理和关系类事实,并具有均衡的真假标签。该基准旨在考察事实核查能力能否从文本迁移到语音,以及检索增强的 LALM 能否利用文本证据来正确支持或驳斥语音声明。我们的实验揭示了一致存在的文本-语音模态鸿沟:能够可靠核查书面声明的 LALM,在面对相同内容的语音声明时往往失败。此外,仅依靠检索带来的提升有限,因为模型经常将检索到的证据与语音声明相混淆。相比之下,检索结合显式推理可以改善声明与证据之间的比对,其中经过思维调优(thinking-tuned)的 LALM 达到了 86.1% 的准确率。VeriSpeak 表明,有效的语音错误信息检测不仅需要语音理解能力,还需要基于检索证据的扎实推理。该数据集已通过 Hugging Face 公开发布:https://huggingface.co/datasets/abhiram4572/VeriSpeak。
cs.LG / 119 / 2609.30258

Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning

面向具身强化学习中隐私轨迹重建的时间梯度反演攻击
Bhujel, Sudip, Shi, Shanghao, Huang, Ruiquan, Zhang, Ning, Xiao, Yang
Abstract
Distributed learning in embodied reinforcement-learning agents offers a degree of privacy by retaining raw sensor data on-device and transmitting only policy gradients to the server. Yet temporal structure can amplify this leakage beyond single-frame attacks. We introduce Temporal Reconstruction Attack on Consecutive Encodings (TRACE), an amortized temporal gradient-inversion attack that autoregressively reconstructs the sequence of private observation-action trajectories from per-step policy-learning gradients. The attack exploits two structural signals ignored by prior single-frame methods: (i) cross-time correlation between successive embodied gradients, which we formalize via a conditional mutual-information bound, and (ii) closed-form action recovery from policy-head gradient structure, which we prove exact when standard entropy regularization is sufficiently small. On held-out embodied scenes, TRACE reaches $18.8$ dB PSNR with near-perfect action recovery at $3$-$4.5$ ms per reconstructed frame, dominating the learning-based baseline across all reconstruction metrics and exceeding optimization attacks while running orders of magnitude faster. Further evaluation demonstrates TRACE's broader applicability across recurrent, residual, and compact transformer victim architectures, multi-modal inputs, and larger discrete action spaces. Defense experiments suggest that protecting temporal gradient streams may require sequence-aware privacy mechanisms.
Chinese Translation
具身强化学习智能体中的分布式学习通过将原始传感器数据保留在设备端、仅向服务器传输策略梯度,从而提供了一定程度的隐私保护。然而,时间结构可能使这种信息泄露超出单帧攻击所能造成的范围。我们提出了针对连续编码的时间重建攻击(Temporal Reconstruction Attack on Consecutive Encodings, TRACE),这是一种摊销式的时间梯度反演攻击,能够从逐步策略学习梯度中自回归地重建私密的观测-动作轨迹序列。该攻击利用了以往单帧方法所忽视的两个结构性信号:(i) 相继具身梯度之间的跨时间相关性,我们通过条件互信息界对其进行了形式化;(ii) 基于策略头梯度结构的闭式动作恢复,我们证明当标准熵正则化足够小时,该恢复是精确的。在留出的具身场景上,TRACE 达到了 18.8 dB 的 PSNR,动作恢复接近完美,每帧重建耗时仅 3-4.5 毫秒,在所有重建指标上全面超越基于学习的基线方法,并在超越优化类攻击的同时运行速度快几个数量级。进一步评估表明,TRACE 在循环、残差及紧凑型 Transformer 受害者架构、多模态输入以及更大的离散动作空间上均具有更广泛的适用性。防御实验表明,保护时间梯度流可能需要具备序列感知能力的隐私保护机制。