← Back to Index
Daily Research Digest

arXiv Papers

2026-09-29
1091
Papers
4
Categories
1091
Translated
收藏清单 0
机器人学 (Robotics)
191
cs.RO / 1 / 2609.31760

What Stops Recursive Self-Improvement in Robotics? Lessons from 123 Rounds of Agentic Skill Discovery

是什么阻碍了机器人领域的递归自我改进?——来自123轮智能体技能发现的启示
Wang, Jiaming
Abstract
Can a robot improve itself the way coding agents now improve software? We built an agentic system to find out. It watches a robot fail, works out which capability is missing, writes new skills or finds and installs external models, tests every change in simulation, and repeats, with no human writing robot code. We ran it for 123 improvement rounds on household manipulation tasks. This report describes what we learned. The good news is that the agent can discover capabilities on its own: noticing that its targets were out of view, it asked for an active-viewing model, debugged it, and deployed a working search skill. The bad news is that its improvements did not add up. Changes kept passing their tests, yet the target task, putting condiments on the top shelf of a fridge, never succeeded. We found that the agent was rarely the bottleneck. Three things around it were. First, chained perception modules do not understand relations. Segmenters such as SAM 3 find shelves but not "the top shelf", so the agent filled the gap with ever more geometric rules that never converged, when what it needed was a different kind of model. Second, skill chains lock learning onto the first step. Long tasks mostly fail early, so evidence and fixes pile up there, and later skills are rarely reached, tested, or improved. Third, what the agent learns is decided by the harness. The agent optimized exactly what the evaluator measured, including where it was wrong, and weak tests and misleading memory turned activity into a standstill. We distill these lessons into concrete recommendations for building robot systems that improve themselves, each paired with an experiment that could prove it wrong.
Chinese Translation
机器人能否像如今编码智能体改进软件那样改进自身?我们构建了一个智能体系统来寻找答案。它观察机器人失败,找出缺失的能力,编写新技能或寻找并安装外部模型,在仿真中测试每项更改,然后重复,全程无需人类编写机器人代码。我们在家庭操作任务上运行了123轮改进。本报告描述了我们的发现。好消息是,智能体能够自主发现能力:它注意到目标不在视野内,便请求一个主动观察模型,调试它,并部署了一个有效的搜索技能。坏消息是,它的改进并未累积。更改不断通过测试,但目标任务——将调味品放到冰箱顶层架子上——却从未成功。我们发现,智能体很少是瓶颈。它周围的三个因素才是。首先,链式感知模块不理解关系。诸如SAM 3之类的分割器能找到架子,但找不到‘顶层架子’,因此智能体用越来越多的几何规则填补空白,但这些规则从未收敛,而它需要的是一种不同类型的模型。其次,技能链将学习锁定在第一步。长任务大多在早期失败,因此证据和修复堆积在那里,后续技能很少被触及、测试或改进。第三,智能体学到什么由测试框架决定。智能体精确优化了评估器所测量的内容,包括其错误之处,而弱测试和误导性记忆将活动变成了停滞。我们将这些经验提炼为构建能自我改进的机器人系统的具体建议,每条建议都配有一个可能证明其错误的实验。
cs.RO / 2 / 2609.31770

Robot Manipulation with GPT-6-Astra: Body Knowledge, Experience Reuse, Emergent Skills, and Sim2Real Transfer

基于GPT-6-Astra的机器人操作:本体知识、经验复用、涌现技能与Sim2Real迁移
He, Sida, Xie, Lingxi, Cao, Yunning, Chen, Pengfei, Duan, Kaiwen, Ge, Jiannan, Huo, Xinyue, Shao, Jiacheng, Tian, Qi
Abstract
General-purpose multimodal agents can write robot-control programs, but repeated exploration and model-mediated action selection can make execution slow. We study how external body knowledge, successful experience, and executable skills improve an XLeRobot controlled by GPT-6-Astra in a simulated and a physical elevator-button task. In 30 fixed-start simulation trials, complete robot geometry and camera information reduce mean completion time by 57.4% relative to a baseline with only the common control interface and no prior experience; images with synchronized action and state records reduce it by 68.6% without additional body assets. In nine paired comparisons (18 trials) at starts displaced by 10-100 cm, experience recorded at the original start reduces mean time by 58-63% relative to no experience, demonstrating generalization to the tested new starting positions. During experience experiments, GPT-6-Astra spontaneously generates a short visual-feedback program. Researcher-refactored versions reduce mean local-task time by 29-31% in 27 simulation trials. Finally, 12 real-robot trials using operator-confirmed button contact demonstrate sim2real reuse: at a shared nominal start, simulation XML assets and simulation experience reduce mean time by 53.0% and 49.9%, respectively; real experience also transfers to two new starts. These results suggest a practical way to build general-purpose manipulation experiments around GPT-6-Astra: supply machine-readable body descriptions and synchronized demonstrations, and turn useful agent-generated feedback routines into reusable skills, while the agent adapts actions from current images. We release all task prompts, trial-level experimental data, and acquired skill implementations at https://github.com/hesd10/astra-robot-sim2real.
Chinese Translation
通用多模态智能体可以编写机器人控制程序,但重复探索和模型介导的动作选择会使执行缓慢。我们研究了外部本体知识、成功经验和可执行技能如何改善由GPT-6-Astra控制的XLeRobot在模拟和物理电梯按钮任务中的表现。在30次固定起始的模拟试验中,完整的机器人几何和相机信息相对于仅具有通用控制接口且无先验经验的基线,将平均完成时间减少了57.4%;带有同步动作和状态记录的图像在没有额外本体资产的情况下将其减少了68.6%。在起始位置偏移10-100厘米的九组配对比较(18次试验)中,在原始起始位置记录的经验相对于无经验将平均时间减少了58-63%,证明了对所测试的新起始位置的泛化能力。在经验实验中,GPT-6-Astra自发地生成了一个简短的视觉反馈程序。研究人员重构的版本在27次模拟试验中将平均局部任务时间减少了29-31%。最后,12次使用操作员确认按钮接触的真实机器人试验展示了Sim2Real复用:在共享的标称起始位置,模拟XML资产和模拟经验分别将平均时间减少了53.0%和49.9%;真实经验也迁移到了两个新的起始位置。这些结果表明了一种围绕GPT-6-Astra构建通用操作实验的实用方法:提供机器可读的本体描述和同步演示,并将有用的智能体生成的反馈例程转化为可重用技能,同时智能体根据当前图像调整动作。我们在https://github.com/hesd10/astra-robot-sim2real发布了所有任务提示、试验级实验数据和获得的技能实现。
cs.RO / 3 / 2609.31771

Reduced Cartesian Kinetostatics for Tendon-Driven Continuum Robots: Residual-Stabilized Full-Shape Propagation

腱驱动连续体机器人的降阶笛卡尔运动静力学:残差稳定化全形状传播
Wu, Ke, Yang, Fangju, Zhang, Xiaohui, Zhang, Zhengqiang, Yi, Jingang, Dai, Jian S.
Abstract
Many planning and control tasks for tendon-driven continuum robots (TDCRs) require the complete Cartesian backbone geometry. We present a reduced Cartesian framework for planar, axially compressible TDCRs that propagates equilibrium configurations along prescribed tendon-force and tendon-displacement trajectories. The backbone is represented by two global position fields. Following exact variation, a Taylor-Galerkin reduction condenses prescribed spatial properties and distributed loads into offline moment vectors, yielding analytic reduced residuals and Jacobians without online spatial quadrature or numerical differentiation. Analytical differentiation and residual correction yield first-order rate systems requiring one fixed-dimensional linear solve per rate evaluation after initial equilibrium alignment on a regular branch. Across four simulated cases covering variable tendon routing, nonuniform geometry, axial compression, and their combined effects, the propagated Cartesian shapes and distributed strains closely match pointwise geometrically variable-strain (GVS) equilibrium solutions. Residual correction suppresses propagation drift across the tested step sizes while adding only about 0.98% to the mean update time of uncorrected Euler. The proposed method requires 0.508 ms per update on average, approximately 11 times faster than pointwise GVS solves. Displacement-driven experiments yield a maximum normalized mean backbone position error of 1.02% and a maximum end-effector position error of 0.28%. These results support efficient and accurate Cartesian full-shape prediction along prescribed actuation paths.
Chinese Translation
许多针对腱驱动连续体机器人(TDCR)的规划与控制任务需要完整的笛卡尔骨干几何。我们提出了一种针对平面、轴向可压缩TDCR的降阶笛卡尔框架,该框架沿规定的腱力和腱位移轨迹传播平衡构型。骨干由两个全局位置场表示。经过精确变分后,Taylor-Galerkin降阶将规定的空间属性和分布载荷凝聚为离线矩向量,从而无需在线空间求积或数值微分即可产生解析的降阶残差和雅可比矩阵。解析微分和残差校正产生了一阶速率系统,在规则分支上进行初始平衡对齐后,每次速率评估需要一个固定维度的线性求解。在涵盖可变腱布线、非均匀几何、轴向压缩及其组合效应的四个仿真案例中,传播的笛卡尔形状和分布应变与逐点几何变应变(GVS)平衡解密切匹配。残差校正抑制了在测试步长上的传播漂移,而仅使未校正欧拉法的平均更新时间增加约0.98%。所提方法平均每次更新需要0.508毫秒,比逐点GVS求解快约11倍。位移驱动实验产生的最大归一化平均骨干位置误差为1.02%,最大末端执行器位置误差为0.28%。这些结果支持沿规定驱动路径进行高效且准确的笛卡尔全形状预测。
cs.RO / 4 / 2609.31773

Timed Rule-Based Supervision of an End-to-End Autonomous Parking Policy

端到端自动泊车策略的定时规则监督
Gao, Kejia, Zhou, Liguo, Yu, Lei, Knoll, Alois
Abstract
We study whether a manually specified runtime supervisor can correct recurring failures of an existing end-to-end parking policy in a fixed CARLA parking lot. The vision-based Transformer architecture is inherited from Yang et al.; our contribution is a timed, rule-based Parametric Safety Shield (PSS) applied to its control outputs. The PSS uses hand-calibrated speed, position, and duration thresholds to intervene in observed failure modes, including boundary exits, delayed braking, and stalled or oscillatory control. In the reported closed-loop evaluation, 16 held-out target slots and six initial poses are each evaluated in four rounds (384 attempts per configuration). Target success increases from 327/384 (85.16%) for the retrained policy to 375/384 (97.66%) with the PSS; mean position and orientation errors among successful attempts are 0.21m and 0.33 degrees. These results show an improvement within this simulator setup. The repeated attempts share one map, vehicle, and sensor configuration, and the PSS uses simulator world coordinates; thus the results do not establish generalization to other lots or real vehicles, or a formal safety guarantee.
Chinese Translation
我们研究在固定的 CARLA 停车场中,人工指定的运行时监督器能否纠正现有端到端泊车策略反复出现的故障。基于视觉的 Transformer 架构继承自 Yang 等人;我们的贡献是应用于其控制输出的定时、基于规则的参数化安全盾(Parametric Safety Shield, PSS)。PSS 使用手动校准的速度、位置和持续时间阈值,对观察到的故障模式进行干预,这些故障模式包括驶出边界、制动延迟以及控制停滞或振荡。在报告的闭环评估中,16 个留出的目标车位与 6 个初始位姿分别在四轮中进行评估(每种配置 384 次尝试)。目标成功率从重新训练策略的 327/384(85.16%)提高到加入 PSS 后的 375/384(97.66%);成功尝试的平均位置误差和朝向误差分别为 0.21 米和 0.33 度。这些结果表明在该仿真器设置内有所改进。重复尝试共享同一地图、车辆和传感器配置,且 PSS 使用仿真器世界坐标;因此,这些结果并不能确立对其他停车场或真实车辆的泛化能力,也不能构成正式的安全保证。
cs.RO / 5 / 2609.31803

Path Planning with Motion Primitives in Dynamic Environments: SIPP on Lattices

动态环境中基于运动基元的路径规划:格点上的SIPP
Agranovskiy, Marat
Abstract
Autonomous navigation in dynamic environments is a critical challenge, particularly when spaces are shared with other mobile agents whose future trajectories are known. While traditional grid-based planners efficiently find collision-free paths, their reliance on stop-and-turn mechanics over $2^k$-connected grids produces piecewise-linear trajectories that are kinodynamically highly sub-optimal for differentially constrained robots. In this paper, we present an adaptation of Safe Interval Path Planning (SIPP) that operates on state lattices, utilizing precomputed, kinodynamically smooth motion primitives. To efficiently handle dynamic environments, we rasterize the spatiotemporal swept volumes of moving obstacles directly onto the grid, treating grid cells as atomic units of space, whose resolution is typically dictated by inherent localization noise. We perform a comprehensive comparative analysis between our lattice-based approach and $2^k$-connected grid planners across diverse topological environments. Our evaluation considers a broad spectrum of performance metrics, including planning time, path angularity, cumulative heading change (angle-over-length), and bending energy. The results demonstrate that while the expanded state space of lattice-based search increases computational overhead, it yields trajectories with significantly superior kinodynamic properties. Specifically, our method achieves a reachability comparable to highly connected grids while ensuring smooth, continuous, and physically executable paths ready for real-world deployment.
Chinese Translation
动态环境中的自主导航是一项关键挑战,尤其是当空间与其他移动智能体共享且其未来轨迹已知时。尽管传统的基于栅格的规划器能高效地找到无碰撞路径,但它们依赖于在$2^k$连通栅格上的停转机制,产生的分段线性轨迹对于受微分约束的机器人来说在运动动力学上高度次优。在本文中,我们提出了一种在状态格点上运行的安全区间路径规划(SIPP)的适应方法,利用预计算的、运动动力学平滑的运动基元。为高效处理动态环境,我们将移动障碍物的时空扫掠体积直接栅格化到网格上,将网格单元视为空间的原子单位,其分辨率通常由固有的定位噪声决定。我们在多种拓扑环境中对基于格点的方法与$2^k$连通栅格规划器进行了全面的比较分析。我们的评估考虑了广泛的性能指标,包括规划时间、路径角度、累积航向变化(角度与长度之比)以及弯曲能量。结果表明,虽然基于格点搜索的扩展状态空间增加了计算开销,但它产生的轨迹具有显著更优的运动动力学特性。具体而言,我们的方法实现了与高度连通栅格相当的可达性,同时确保路径平滑、连续且物理可执行,为实际部署做好准备。
cs.RO / 6 / 2609.31840

Humanoid Badminton: Learning Dynamic Racket Skills from Limited Human Motion Data

人形机器人羽毛球:从有限人体运动数据中学习动态球拍技能
Cui, Jingzhi, Wang, Zhexiong, Xu, Bangjie, Zhao, Pengyu, Li, Youyuan, Su, Zhi, Ren, Peng, Xu, Mengdi, Yu, Chao, Wu, Yi, Wang, Luyang, Li, Zhongyu
Abstract
High-speed racket sports provide a demanding testbed for humanoid robots, requiring time-critical decisions, precise striking, and dynamic whole-body coordination. In badminton, fast-changing shuttle trajectories require timely contact decisions, while successful returns demand precise racket pose and velocity within a brief contact window and across a broad three-dimensional striking workspace. Human motion data provide valuable priors for such athletic skills, but usable badminton references are limited and imperfect. Direct tracking provides insufficient executable variation for diverse shuttle conditions, while purely task-driven optimization may produce unnatural motion. To address these challenges, we present a three-stage hierarchical reinforcement learning framework for dynamic humanoid badminton. First, task-randomized motion augmentation expands sparse annotated hitting events into executable target-conditioned stroke variations, forming a continuous latent skill space. Second, a high-level planner outputs continuous latent skill codes to compose these skills online according to the observed shuttle state. Third, a context-conditioned adversarial regularizer encourages more natural planner-level skill usage while preserving return performance. When deployed on a real humanoid robot, our system achieves sustained multi-skill rallies with human players, including forehand, backhand, and highly dynamic jump returns. This is the first real-world humanoid racket-sport system to demonstrate multi-skill human--robot rallies including highly dynamic jump returns.
Chinese Translation
高速球拍类运动为人形机器人提供了一个高要求的测试平台,要求进行时间关键型决策、精准击球和动态全身协调。在羽毛球中,快速变化的羽毛球轨迹要求及时做出触球决策,而成功回球则需要在短暂接触窗口内、在广阔的三维击球工作空间内具备精确的球拍位姿和速度。人体运动数据为此类竞技技能提供了有价值的先验,但可用的羽毛球参考数据有限且不完美。直接跟踪无法为多样化的羽毛球状况提供足够的可执行变化,而纯任务驱动的优化可能产生不自然的运动。为解决这些挑战,我们提出了一种用于动态人形机器人羽毛球的三阶段分层强化学习框架。首先,任务随机化的运动增强将稀疏标注的击球事件扩展为可执行的、以目标为条件的击球变化,形成连续的潜在技能空间。其次,高层规划器输出连续的潜在技能编码,以根据观测到的羽毛球状态在线组合这些技能。第三,上下文条件对抗正则化器鼓励更自然的规划器级技能使用,同时保持回球性能。当部署在真实人形机器人上时,我们的系统能够与人类球员进行持续的多技能对打,包括正手、反手和高度动态的跳跃回球。这是首个在真实世界中展示多技能人机对打(包括高度动态跳跃回球)的人形机器人球拍运动系统。
cs.RO / 7 / 2609.31855

PHIRL: Aligning Learned Rewards with Task Progress for Inverse Reinforcement Learning

PHIRL: 将学习到的奖励与任务进度对齐用于逆向强化学习
Yu, Hang, Staley, James, Tsou, Cheng Xi, Liu, Xiujin, Gao, Wenchang, Huang, Jindan, Fang, Shijie, Shangguan, Zhegong, Cangelosi, Angelo, Aronson, Reuben, Short, Elaine
Abstract
Human demonstrations provide dense policy-level information but sometimes lack local precision. Human feedback presents accurate local critiques, but offers sparse evaluations rather than direct policy guidance. We propose Progress-Heuristicized Inverse Reinforcement Learning (PHIRL), a data-efficient framework that learns robust reward functions by jointly leveraging demonstrations and feedback. Specifically, we use progress, a feedback modality that describes cumulative task completion. PHIRL iteratively infers a reward function from demonstrations via inverse reinforcement learning, calculates the learned rewards over the progress-annotated demonstrations, and aligns the rewards with progress annotations over four dimensions. We evaluate PHIRL on real and simulated robot tasks, with additional exploration using a fine-tuned vision-language model to provide progress feedback. Results demonstrate that PHIRL significantly outperforms the baselines, achieving substantially higher environmental return rewards and task success with only twenty percent of demonstrations annotated. Analysis of reward-hacking scenarios demonstrates that PHIRL learned reward functions are reliable against exploitation.
Chinese Translation
人类演示提供了密集的策略级信息,但有时缺乏局部精确性。人类反馈提供了准确的局部评价,但提供的是稀疏评估而非直接策略指导。我们提出进度启发式逆向强化学习(PHIRL),这是一个数据高效的框架,通过联合利用演示和反馈来学习鲁棒的奖励函数。具体来说,我们使用进度,一种描述累积任务完成度的反馈模态。PHIRL通过逆向强化学习从演示中迭代地推断奖励函数,计算进度标注演示上的学习奖励,并在四个维度上将奖励与进度标注对齐。我们在真实和模拟机器人任务上评估PHIRL,并额外探索使用微调的视觉语言模型来提供进度反馈。结果表明,PHIRL显著优于基线,在仅标注20%演示的情况下实现了显著更高的环境回报奖励和任务成功率。对奖励黑客场景的分析表明,PHIRL学习到的奖励函数在对抗利用方面是可靠的。
cs.RO / 8 / 2609.31880

Efficient Bezier Velocity Optimization for Free-Floating Space Manipulators

自由漂浮空间机械臂的高效贝塞尔速度优化
Zhang, Duo, Zhang, Zhizhuo, Bai, Xiaoli, Yu, Jingjin
Abstract
We present FAVOR (Free-floating Arm Velocity Optimization with Recursive Sensitivities), a planner for collision-free reaching, tracking, and prescribed-time pre-grasp interception on an unactuated spacecraft. It optimizes Bezier joint-velocity curves with linear velocity, acceleration, and continuity constraints. Decision dimension is independent of rollout resolution. Analytical recursive sensitivities provide task and clearance gradients through the coupled base-arm motion. Parallel evaluation, caching, and incremental collision discovery reduce computation. With a seven-DoF arm and five simulated spacecraft models, FAVOR achieves 99.8% point-to-point success with 1.779 s mean computation, versus 62.9% and 50.270 s for an IK-initialized position-spline baseline. Means include failures and timeouts. FAVOR completes 30 of 36 tracking cases, versus 15 for single-step QP, and all 36 interception instances, versus 22 for the spline baseline. A controlled ablation shows that finite differences increase mean planning time 4.7-fold.
Chinese Translation
我们提出了 FAVOR(Free-floating Arm Velocity Optimization with Recursive Sensitivities,基于递归灵敏度的自由漂浮机械臂速度优化),一种用于无驱动航天器上的无碰撞到达、跟踪和规定时间预抓取拦截的规划器。它优化满足线速度、加速度和连续性约束的贝塞尔关节速度曲线。决策维度与推演分辨率无关。解析递归灵敏度通过耦合的基座-机械臂运动提供任务和间隙梯度。并行评估、缓存和增量碰撞发现减少了计算量。使用一个七自由度机械臂和五个模拟航天器模型,FAVOR 实现了 99.8% 的点对点成功率,平均计算时间为 1.779 秒,而基于逆运动学初始化的位置样条基线方法为 62.9% 和 50.270 秒。均值包括失败和超时情况。FAVOR 完成了 36 个跟踪案例中的 30 个,而单步二次规划(QP)为 15 个;并完成了所有 36 个拦截实例,而样条基线为 22 个。一项受控消融实验表明,有限差分使平均规划时间增加了 4.7 倍。
cs.RO / 9 / 2609.31885

Differentiable Dynamics for Autonomous Micro-Mobility Navigation

面向自主微出行导航的可微动力学
Cai, Grace, Lee, Joey, Parepally, Nithin, Zheng, Laura, Lin, Ming C.
Abstract
Autonomous micro-mobility vehicles (MMVs) such as wheelchairs, scooters, and bicycles have the potential to improve mobility access and support safe low-speed transportation in pedestrian-shared spaces. Achieving MMV autonomy will require realistic, predictable MMV motion. However, many existing autonomous vehicle stacks rely on simplified kinematic models that fail to capture key MMV characteristics such as tire slip, friction, and wheel layouts, limiting realism and gradient-based optimization. In this paper, we explore differentiable formulations of dynamics models for autonomous micro-mobility systems. We first construct DiffKBM, a differentiable version of the kinematic bicycle model (KBM). Then, we introduce DiffGM3, a differentiable formulation of the General Micro-Mobility Model (GM3), a unified tire-based dynamics formulation for micro-mobility vehicles that supports a wide range of MMV configurations. DiffKBM and DiffGM3 enable end-to-end differentiable optimization through MMV dynamics, making them suitable for integration into differentiable autonomy stacks. We evaluate these dynamics models in both open-loop and closed-loop settings: (1) open-loop trajectory matching, where DiffKBM and DiffGM3 are integrated as a dynamics layer within DiffStack and optimized to reproduce real-world MMV trajectories, and (2) closed-loop autonomous navigation, where DiffKBM and DiffGM3 are paired with a differentiable MPC controller in CrowdNav pedestrian scenarios. In the open-loop setting, DiffGM3 outperforms DiffKBM in reproducing trajectories with improvements in ADE and NLL across bicycle, scooter, and motorcycle modes, and reductions in planning loss for bicycle and motorcycle trajectories. We also find that, in closed-loop settings, DiffGM3 improves on DiffKBM's CrowdNav performance by producing 55\% fewer collisions and a 75\% lower discomfort frequency for the bicycle mode.
Chinese Translation
自主微出行车辆(MMVs),如轮椅、滑板车和自行车,有潜力改善出行便利性,并支持行人共享空间中的安全低速交通。实现MMV自主性需要真实、可预测的MMV运动。然而,许多现有的自动驾驶车辆栈依赖于简化的运动学模型,这些模型无法捕捉关键的MMV特征,如轮胎滑移、摩擦和车轮布局,限制了真实性和基于梯度的优化。在本文中,我们探索了自主微出行系统动力学模型的可微形式。我们首先构建了DiffKBM,这是运动学自行车模型(KBM)的可微版本。然后,我们引入了DiffGM3,这是通用微出行模型(GM3)的可微形式,GM3是一种统一的基于轮胎的微出行车辆动力学公式,支持广泛的MMV配置。DiffKBM和DiffGM3通过MMV动力学实现了端到端可微优化,使其适合集成到可微自主栈中。我们在开环和闭环设置中评估这些动力学模型:(1)开环轨迹匹配,其中DiffKBM和DiffGM3作为DiffStack中的动力学层集成并优化以重现真实世界的MMV轨迹;(2)闭环自主导航,其中DiffKBM和DiffGM3与可微MPC控制器配对,用于CrowdNav行人场景。在开环设置中,DiffGM3在重现轨迹方面优于DiffKBM,在自行车、滑板车和摩托车模式下ADE和NLL均有改善,并且自行车和摩托车轨迹的规划损失降低。我们还发现,在闭环设置中,DiffGM3改进了DiffKBM的CrowdNav性能,在自行车模式下产生的碰撞减少了55%,不适频率降低了75%。
cs.RO / 10 / 2609.31901

Game-Theoretic Control with Constrained Potential Surgery

基于约束势手术(Constrained Potential Surgery)的博弈论控制
Zhang, Zhiyuan, Tsiotras, Panagiotis
Abstract
Constrained general-sum dynamic games are a popular formulation for highly interactive multi-agent planning problems. In recent years, Generalized Nash Equilibrium (GNE) solvers have achieved real-time performance for small dynamic games. However, solution speed still remains a bottleneck, and controlling even a small number of agents (e.g., more than four) in a dynamic task remains elusive. In addition, Newton solvers that focus on the first-order conditions are vulnerable to non-Nash saddle points, limiting the usefulness of the provided solution. In this work, we propose a fast and versatile interior point solver for constrained dynamic games, along with a computationally efficient second-order correction that increases the probability of converging to a local GNE solution in constrained dynamic games. The performance of the proposed method is evaluated on numerical benchmarks and a physical experiment involving scaled race cars.
Chinese Translation
约束一般和动态博弈是高度交互多智能体规划问题的一种流行形式。近年来,广义纳什均衡(GNE)求解器已在小规模动态博弈中实现了实时性能。然而,求解速度仍然是一个瓶颈,在动态任务中控制即使少量智能体(例如,超过四个)仍然难以实现。此外,专注于一阶条件的牛顿求解器容易受到非纳什鞍点的影响,限制了所提供解的有效性。在这项工作中,我们提出了一种用于约束动态博弈的快速且通用的内点求解器,以及一种计算高效的二阶校正,它增加了在约束动态博弈中收敛到局部GNE解的概率。所提方法的性能在数值基准和涉及缩比赛车的物理实验上进行了评估。
cs.RO / 11 / 2609.31904

GT-VLA: Target-Conditioned Trace Guidance for Generalizable Robotic Manipulation

GT-VLA:面向可泛化机器人操作的目标条件轨迹引导
Zhong, Ninghan, Peng, Jing-Chen, Vishwanath, Sriram
Abstract
Vision-Language-Action (VLA) models have shown strong performance on robotic manipulation, but they often struggle to generalize to unseen tasks, configurations, and long-horizon settings. A key challenge is that VLAs overfit to training scenes and fail to follow novel language instructions. Off-the-shelf vision-language models (VLMs) often provide stronger generalization, but cannot directly control robot actions. To combine the common sense of VLMs with VLA control, we propose Guided Trace VLA (GT-VLA), a steerable framework that accepts guidance from an external generalist VLM through trace-conditioned action generation. GT-VLA uses a generalist model to identify semantic guidance for the current skill, converts this guidance into a 2D visual trace, and conditions its action policy on the resulting trace-rendered observation. This design separates semantic target acquisition, trace generation, and low-level action execution, allowing high-level guidance to propagate to robot actions. GT-VLA uses a Mixture-of-Experts architecture with skill-specific trace and action modules for robust execution. We evaluate GT-VLA on LIBERO and a physical robot platform, showing improved generalization over recent VLA baselines in both settings. The code and additional supplemental materials are available on our project website at https://ivaniz.github.io/gt-vla/.
Chinese Translation
视觉-语言-动作(VLA)模型在机器人操作上表现出强大性能,但往往难以泛化到未见过的任务、配置和长时程场景。一个关键挑战是VLA过拟合于训练场景,无法遵循新颖的语言指令。现成的视觉-语言模型(VLM)通常提供更强的泛化能力,但无法直接控制机器人动作。为了将VLM的常识与VLA控制相结合,我们提出了引导轨迹VLA(GT-VLA),一个可操控的框架,通过轨迹条件动作生成接受来自外部通用VLM的引导。GT-VLA使用通用模型识别当前技能的语义引导,将该引导转换为2D视觉轨迹,并基于所得轨迹渲染的观测来条件化其动作策略。该设计分离了语义目标获取、轨迹生成和底层动作执行,允许高层引导传播到机器人动作。GT-VLA采用专家混合架构,具有技能特定的轨迹和动作模块,以实现鲁棒执行。我们在LIBERO和物理机器人平台上评估GT-VLA,显示在两种设置中相比近期VLA基线具有改进的泛化能力。代码和额外补充材料可在项目网站 https://ivaniz.github.io/gt-vla/ 获取。
cs.RO / 12 / 2609.31905

End-to-end QP-based policies: A unified perspective on robust control and robot learning

端到端基于QP的策略:对鲁棒控制和机器人学习的统一视角
Vega, Fausto, Balaji, Priyanka Supraja, Dunaway, Chase, Koszut, Joe, Arrizabalaga, Jon, Manchester, Zachary
Abstract
We present an end-to-end QP-based policy framework that enables systematic policy construction with minimal domain-specific design, while preserving the transparency and interpretability of model-based control. The proposed policy representation supports both domain-randomized model-based auto-tuning, where policy parameters are optimized over distributions of disturbances and model variations, and black-box policy construction, where the problem is formulated in terms of a (possibly) unknown model without requiring explicit notions of states, inputs, or the underlying system dynamics. We establish connections to existing policy representations and control paradigms, including robust control, multilayer perceptrons, and robot learning, and interpret our end-to-end QP policies as a common abstraction of these approaches. We validate the resulting framework in both simulation and hardware, demonstrating a broad range of applications spanning robustness, automatic policy tuning, and control under unknown system dynamics.
Chinese Translation
我们提出了一个端到端基于QP的策略框架,该框架能够以最少的领域特定设计进行系统化的策略构建,同时保持基于模型控制的透明性和可解释性。所提出的策略表示支持域随机化的基于模型的自动调优,其中策略参数在干扰和模型变化的分布上进行优化,以及黑盒策略构建,其中问题根据(可能)未知的模型来表述,无需明确的状态、输入或底层系统动力学的概念。我们建立了与现有策略表示和控制范式的联系,包括鲁棒控制、多层感知机和机器人学习,并将我们的端到端QP策略解释为这些方法的共同抽象。我们在仿真和硬件中验证了所得到的框架,展示了广泛的应用,涵盖鲁棒性、自动策略调优以及未知系统动力学下的控制。
cs.RO / 13 / 2609.31924

HapticWorld: an Interactive World Simulator with Real-time Torque Feedback

HapticWorld:一个具有实时力矩反馈的交互式世界模拟器
Peng, Shaoting, Liang, Litian, Wang, Yixuan, Yang, Ming, Driggs-Campbell, Katherine, Cutkosky, Mark, Xu, James Jingxi
Abstract
Contact-rich manipulation depends on force sensing that is hard to infer from visual signals alone, both for collecting demonstrations and for training policies. Force-annotated data, however, remains hard to obtain at scale: real-robot collection ties every demonstration to physical hardware, physics simulators report contact forces that deviate systematically from real measurements, and learned world simulators, though scalable and realistic, are vision-only, so operators feel nothing during data collection and the data carries no force/torque (F/T) labels. We present HapticWorld, an interactive world simulator that predicts joint torque together with observations and renders it back to the operator in real time, closing the haptic loop between a human and a learned world model. Across three contact-rich tasks, torque feedback raises data collection throughput by 1.6 times on average. Policies trained on HapticWorld-generated demonstrations succeed in 54/60 real-world trials, approaching the 56/60 upper bound of real-world data, and far exceeding the 19/60 success rate of the vision-only baseline. Moreover, the success rates measured inside HapticWorld closely match real-world evaluation, demonstrating that HapticWorld can serve as a stand-alone F/T-conditioned policy evaluation platform.
Chinese Translation
富接触操作依赖于仅凭视觉信号难以推断的力传感,无论是收集演示数据还是训练策略。然而,带力标注的数据仍然难以大规模获取:真实机器人收集将每个演示与物理硬件绑定,物理模拟器报告的接触力与真实测量存在系统性偏差,而学习到的世界模拟器虽然可扩展且逼真,但仅依赖视觉,因此操作员在数据收集过程中没有任何触觉感受,且数据不包含力/力矩(F/T)标签。我们提出了HapticWorld,一个交互式世界模拟器,它预测关节力矩并连同观测一起实时渲染回操作员,从而闭合人与学习到的世界模型之间的触觉回路。在三个富接触任务中,力矩反馈将数据收集吞吐量平均提高了1.6倍。在HapticWorld生成的演示数据上训练的策略在54/60的真实世界试验中成功,接近真实世界数据56/60的上限,并远超仅视觉基线19/60的成功率。此外,在HapticWorld内部测得的成功率与真实世界评估密切匹配,表明HapticWorld可以作为一个独立的以F/T为条件的策略评估平台。
cs.RO / 14 / 2609.31929

Inertia-Corrected Newton Method For Generalized Nash Equilibria in Dynamic Games with Optimality Verification

动态博弈中广义纳什均衡的惯性校正牛顿法及最优性验证
Zhang, Zhiyuan, Tsiotras, Panagiotis
Abstract
Newton methods efficiently find Generalized Nash Equilibria (GNE) in dynamic games by solving for the KKT necessary conditions. These methods are fast and can support multi-agent Model Predictive Control (MPC) for highly dynamic robots. However, a small KKT residual alone does not certify that the returned solution satisfies the second-order sufficient conditions for a local GNE. In this paper, we propose an efficient numerical method to verify the second-order sufficient conditions (SOSC) for a local GNE. We connect the inertia of the agent KKT matrix with the positive definiteness of the reduced Hessian of the cost function, projected onto the null space of the constraints. Furthermore, we introduce an inertia-corrected update step that improves convergence to local GNEs by destabilizing strict saddle points with weak cross-agent coupling. Our main contribution is a fast Newton solver for Constrained Dynamic Games that provides efficient optimality checking. Through numerical benchmarks, we demonstrate the proposed solver's runtime and convergence performance in practical multi-agent planning problems. We also validate the solver's real-time capabilities in physical experiments using a platform of miniature autonomous race cars.
Chinese Translation
牛顿法通过求解 KKT 必要条件,能够高效地找到动态博弈中的广义纳什均衡。这些方法速度快,能够支持高度动态机器人的多智能体模型预测控制。然而,仅凭较小的 KKT 残差并不能保证所返回的解满足局部广义纳什均衡的二阶充分条件。本文提出了一种高效数值方法,用于验证局部广义纳什均衡的二阶充分条件。我们将智能体 KKT 矩阵的惯性与代价函数的约化 Hessian 矩阵在约束零空间上的投影的正定性联系起来。此外,我们引入了一种惯性校正更新步,通过使具有弱智能体间耦合的严格鞍点不稳定,来提高收敛到局部广义纳什均衡的速度。我们的主要贡献是提出了一种用于约束动态博弈的快速牛顿求解器,能够高效地进行最优性检验。通过数值基准测试,我们展示了所提求解器在实际多智能体规划问题中的运行时间和收敛性能。我们还使用微型自主赛车平台,在物理实验中验证了该求解器的实时能力。
cs.RO / 15 / 2609.31941

Low Cost Eye Tracking for Vision Screening

用于视力筛查的低成本眼动追踪
Smaili, Maissa Abir, Ismaiel, Yaman, Iscan, Zafer
Abstract
This engineering study assesses gaze-assisted data acquisition during visual screening with a Gaze Quest implementation via webcam and GC308 near infrared camera with an Orlosky eye tracking pipeline. Ten subjects performed under both acquisition conditions. Accuracy of Tumbling E orientation, reported application logMAR value,and response latency were determined from keyboard responses and hence are behavioral outputs as opposed toeye tracker screening results. Frame-to-frame change in theangle between successive normalized gaze vectors (28.2degrees) served as the GC308 eye tracker derived diagnostic with an effective sampling rate of roughly 8 FPS. No screen target calibration or gaze to screen translation was done on the GC308 eye tracker output data for the current analysis and hence, these data represent acquisition stability and not gaze accuracy in relation to the target. The webcam pipeline relied on 25-point screen calibration. Because the display resolution and pixel pitch could not be retained in the records of the study, the reported application logMAR values cannot be viewed as measures of visual acuity in any way.
Chinese Translation
这项工程研究评估了在视力筛查过程中,使用 Gaze Quest 实现通过 webcam 和 GC308 近红外摄像头以及 Orlosky 眼动追踪流程进行注视辅助数据采集的情况。十名受试者在两种采集条件下进行了测试。Tumbling E 方向的准确性、报告的应用 logMAR 值和反应潜伏期由键盘响应确定,因此是行为输出,与眼动追踪筛查结果不同。连续归一化注视向量之间的帧间角度变化(28.2 度)作为 GC308 眼动追踪器得出的诊断指标,其有效采样率约为 8 FPS。在当前分析中,未对 GC308 眼动追踪器输出数据进行屏幕目标校准或注视到屏幕的转换,因此,这些数据代表采集稳定性,而不是相对于目标的注视精度。网络摄像头流程依赖于 25 点屏幕校准。由于研究记录中无法保留显示分辨率和像素间距,报告的应用 logMAR 值不能以任何方式视为视力测量指标。
cs.RO / 16 / 2609.32064

Grasp2Twist: Learning Bimanual Dexterous Jar Opening by Reinforcement Learning

Grasp2Twist:通过强化学习学习双臂灵巧开罐
Xu, Mo, Deng, Yunfu, Wang, Jianuo, Hanna, Josiah, Mutlu, Bilge
Abstract
This paper presents Grasp2Twist, a bimanual dexterous manipulation system that learns to grasp and twist open jar lids using reinforcement learning. Learning this task raises three challenges: learning a unified policy for a multi-stage task, sustaining lid twisting, and sim-to-real transfer. To address the first challenge, we introduce a continuous enclosure measure to guide grasp formation and a binary enclosure indicator to guide the grasp-to-twist transition for unified policy learning. We derive both from the geometric relationship between the object center and the convex hull formed by the hand's palm and fingertips. Kinematic constraints limit how far the hand can rotate the lid with fixed contacts, so sustained twisting requires finger contact reconfiguration. We use a three-stage curriculum to facilitate exploration of these contact changes and also improve robustness for sim-to-real transfer. With our approach, the learned policy demonstrates finger gaiting, reconfiguring hand-object contacts to sustain lid rotation. It transfers zero-shot to the physical system and achieves an 88% task success rate across six household containers, including peanut-butter, vitamin, and instant-coffee jars. Ablations further validate the roles of the geometric enclosure in grasp formation and the curriculum in contact-reconfiguration exploration.
Chinese Translation
本文提出了Grasp2Twist,一个双臂灵巧操作(manipulation)系统,它使用强化学习来学习抓取和拧开罐盖。学习这一任务面临三个挑战:为多阶段任务学习统一策略、持续拧动盖子以及从仿真到现实的迁移(sim-to-real transfer)。为了解决第一个挑战,我们引入了一个连续包围度量来指导抓取的形成,以及一个二元包围指示器来指导从抓取到拧动的过渡,用于统一策略学习。我们根据物体中心与由手掌和指尖形成的凸包之间的几何关系推导出两者。运动学约束限制了手在固定接触下能旋转盖子的程度,因此持续拧动需要手指接触的重新配置。我们使用一个三阶段课程(curriculum)来促进对这些接触变化的探索,并提高从仿真到现实迁移的鲁棒性。通过我们的方法,学习到的策略展示了手指步态(finger gaiting),重新配置手-物体接触以维持盖子的旋转。它零样本(zero-shot)迁移到物理系统,并在六个家用容器(包括花生酱、维生素和速溶咖啡罐)上达到了88%的任务成功率。消融实验进一步验证了几何包围在抓取形成中的作用以及课程在接触重新配置探索中的作用。
cs.RO / 17 / 2609.32069

Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models

找到你做不到的事:面向自我改进 VLA 模型的智能体式真实世界强化学习
Fang, Yuan, Li, Zechu, Tong, Haolei, Liu, Puze, Chalvatzaki, Georgia
Abstract
Vision--language--action (VLA) models provide strong priors for robotic manipulation but are typically deployed as frozen policies, unable to improve from their own failures. Real-world reinforcement learning (RL) offers a path to continued improvement, yet manual environment resets and task-success supervision hinder autonomous learning. We introduce \textbf{FIND}, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace. FIND reframes autonomous practice as a scene-conditioned, performance-aware task-selection problem: instead of restoring a predefined scene after each rollout, it uses the resulting scene to determine what to practice next. A vision--language agent identifies feasible tasks from a predefined library, prioritizes those with lower recent success rates, and evaluates outcomes using paired pre- and post-execution observations. We instantiate FIND with a frozen $\pi_{0.5}$ VLA and residual off-policy RL. Across eight real-world manipulation tasks, the independent human-assessed success rate improves from $55\%$ to $71.9\%$. A representative run completes 456 autonomous episodes within 6 hours of interaction, requiring 30 scene-recovery interventions and no human-provided reward labels during online learning. Ablations and systematic evaluations further examine key design choices, agent evaluation accuracy, and human intervention requirements. Our website is made publicly available at: FIND.github.io.
Chinese Translation
视觉-语言-动作(VLA)模型为机器人操作提供了强大的先验,但通常被部署为冻结策略,无法从自身失败中改进。真实世界强化学习(RL)为实现持续改进提供了一条路径,然而手动环境重置和任务成功监督阻碍了自主学习。我们提出了 FIND,一个智能体式真实世界 RL 框架,它在持久工作空间中闭环连接了场景理解、弱点感知练习、自我评估和策略改进。FIND 将自主练习重新表述为一个以场景为条件、感知性能的任务选择问题:不是每次 rollout 后恢复预定义场景,而是利用结果场景来决定接下来练习什么。一个视觉-语言智能体从预定义任务库中识别可行任务,优先选择近期成功率较低的任务,并使用执行前和执行后的配对观测来评估结果。我们用冻结的 π_{0.5} VLA 和残差离策略 RL 实例化 FIND。在八个真实世界操作任务中,独立人工评估的成功率从 55% 提升到 71.9%。一次代表性运行在 6 小时交互内完成了 456 个自主回合,需要 30 次场景恢复干预,并且在在线学习过程中无需人工提供奖励标签。消融实验和系统性评估进一步考察了关键设计选择、智能体评估准确性以及人工干预需求。我们的网站已公开提供:FIND.github.io。
cs.RO / 18 / 2609.32079

Fiber-Normalized Manipulability and Determinant Proxies: Intrinsic Redundancy Optimization Across and Within Task Fibers

纤维归一化可操作性与行列式代理:跨任务纤维与任务纤维内的内在冗余优化
Franchi, Antonio, Mizzoni, Mirko
Abstract
This work establishes that determinant-based manipulability is an exact objective for fixed-task redundancy optimization, despite its dependence on task coordinates and the choice of task-space metric used for volume measurement. On every regular task fiber, the determinant proxy, its representation in any task chart, and every metric-completed manipulability differ only by positive constants. They consequently induce the same complete ordering, constrained extrema, gradient directions, critical points, and local optimality classifications. For comparisons and trajectory optimization across task fibers, this work introduces fiber-normalized manipulability: the capability attained at an internal state divided by the best capability available on the same fiber. The resulting dimensionless scalar is invariant under coordinate changes on the internal-state and task manifolds and independent of the task-space metric. Its associated loss provides an intrinsic objective for physically admissible cross-fiber trajectories, including problems with prescribed or free task evolution and temporal coupling. Planar-manipulator and redundant aerodynamic-allocation examples demonstrate fixed-fiber equivalence, task-dependent cross-fiber differences, and invariant fiber-normalized trajectory optimization.
Chinese Translation
这项工作表明,基于行列式的可操作性是固定任务冗余优化的精确目标,尽管它依赖于任务坐标以及用于体积测量的任务空间度量的选择。在每个正则任务纤维上,行列式代理、其在任何任务图中的表示以及每个度量完备的可操作性仅相差正常数。因此,它们诱导相同的完全排序、约束极值、梯度方向、临界点和局部最优性分类。为了跨任务纤维的比较和轨迹优化,这项工作引入了纤维归一化可操作性:在内部状态处获得的能力除以同一纤维上可用的最佳能力。所得的无量纲标量在内部状态流形和任务流形上的坐标变化下不变,并且独立于任务空间度量。其相关的损失为物理上可允许的跨纤维轨迹提供了内在目标,包括具有规定或自由任务演化以及时间耦合的问题。平面机械臂和冗余气动分配示例展示了固定纤维等价性、任务相关的跨纤维差异以及不变的纤维归一化轨迹优化。
cs.RO / 19 / 2609.32099

Active 3D weaves for load-bearing and damage-resilient locomotion

用于承载与抗损伤运动的主动三维编织结构
Tu, Guowei Wayne, Filipov, Evgueni T.
Abstract
The craft of weaving, where different materials are interlaced, has tremendous potential for creating active and functional systems for use in soft robots, prosthetics, wearables, exoskeletons, and more. These textile-like systems are flexible and safe for human-machine interaction; however, this inherent flexibility limits their ability to carry loads, which is essential for many robotic functions. In this work, we introduce a general framework for integrating active materials into three-dimensional (3D) woven shells to create robotic structures that combine high axial stiffness for load bearing, low bending stiffness for efficient actuation, and system-level resilience for damage tolerance. These woven robots can be modularly assembled from 'woven corners', a fundamental unit of 3D woven structures. We use eigenvalue calculations to identify load bearing and actuation mechanisms of the 3D woven structures, and use that information to make five different robots capable of locomotion. We demonstrate that these 3D woven robots can locomote carrying loads 70 times their self-weight, and can maintain repeatable performance even after being subjected to extreme compression. This work is a pathway toward the design, manufacturing, and simulation of future 3D woven robotic systems where load bearing, high stiffness, active functional deformation, locomotion, and system-level resilience are all needed.
Chinese Translation
将不同材料交织在一起的编织工艺,在构建用于软体机器人、假肢、可穿戴设备、外骨骼等的主动且功能化系统方面具有巨大潜力。这些类纺织系统具有柔顺性,适合人机交互且安全;然而,这种固有的柔顺性限制了其承载能力,而承载对许多机器人功能至关重要。在本工作中,我们提出一个通用框架,将活性材料集成到三维(3D)编织壳体中,以构建兼具用于承载的高轴向刚度、用于高效驱动的低弯曲刚度以及用于损伤容限的系统级韧性的机器人结构。这些编织机器人可由“编织角(woven corners)”模块化组装而成,编织角是三维编织结构的基本单元。我们利用特征值计算来识别三维编织结构的承载与驱动机制,并利用这些信息制作了五个能够运动的机器人。我们证明,这些三维编织机器人能够承载自身重量70倍的载荷进行运动,并且即使在经历极端压缩后仍能保持可重复的性能。这项工作为未来三维编织机器人系统的设计、制造和仿真提供了一条路径,这些系统同时需要承载能力、高刚度、主动功能变形、运动能力和系统级韧性。
cs.RO / 20 / 2609.32129

Residual Denoising Enables Sample-Efficient Multi-Agent Coordination on Demand

残差去噪实现按需的样本高效多智能体协调
Dong, Dayi, Bhatt, Maulik, Shrivastava, Aayushi, Peters, Lasse, Mehr, Negar
Abstract
Pretrained robot policies offer strong manipulation skills but are typically limited to single-agent settings, where a robot acts in isolation. In this work, we study how to adapt pretrained single-agent diffusion policies to multi-agent settings using minimal collaborative data, co-optimizing for two key objectives: high coordination performance and single-agent skill retention. To this end, we introduce ALTER, an adaptation method for coordination on demand: the adapted policy coordinates with other robots when deployed in a team while remaining capable of acting independently when operating alone. Execution is decentralized: each robot acts only on its own visual observations, without explicit inter-agent communication. Our method trains a coordination head that predicts a residual denoiser to transform single-agent behavior into coordinated multi-agent behavior when necessary while also preserving single-agent capabilities. To preserve single-agent capabilities, we augment a small number of collaborative demonstrations with self-distilled data generated by the base policy during training of the residual denoiser. In simulation,ALTER achieves higher coordination success over our baselines while retaining much higher source-skill retention. In our hardware experiments, we find similar trends where ALTER better co-optimizes for coordination success and single-agent skill retention than the baselines
Chinese Translation
预训练机器人策略提供了强大的操作技能,但通常局限于单智能体设置,即机器人单独行动。在这项工作中,我们研究如何使用最少的协作数据将预训练的单智能体扩散策略适应到多智能体设置,共同优化两个关键目标:高协调性能和单智能体技能保留。为此,我们引入了ALTER,一种按需协调的适应方法:适应后的策略在团队部署时能与其他机器人协调,同时在单独运行时仍能独立行动。执行是去中心化的:每个机器人仅根据自身的视觉观察行动,无需显式的智能体间通信。我们的方法训练一个协调头,预测一个残差去噪器,在必要时将单智能体行为转化为协调的多智能体行为,同时保留单智能体能力。为了保留单智能体能力,我们在训练残差去噪器期间,使用基础策略生成的自蒸馏数据来增强少量协作演示。在仿真中,ALTER在我们的基线上实现了更高的协调成功率,同时保持了更高的源技能保留率。在我们的硬件实验中,我们发现了类似的趋势,ALTER比基线更好地共同优化了协调成功率和单智能体技能保留率。
cs.RO / 21 / 2609.32154

GAUGE: Planner-Conditioned Active Calibration of Opaque Quadruped Velocity Interfaces

GAUGE:不透明四足速度接口的规划器条件化主动校准
Zang, Tianhao, Liu, Zihan, Wang, Shanze, Luo, Liyou, Xie, Xingjian, Li, Chengtai, Zhang, Wei
Abstract
In this paper, we present a Goal-Aware Uncertainty-Guided Exploration (GAUGE) framework for planner-conditioned active calibration of opaque quadruped velocity interfaces. Commercial quadrupeds commonly expose planar-velocity commands, but the underlying locomotion controller remains inaccessible and can produce systematic discrepancies between commanded and realized motion. A navigation planner typically uses a structured subset of the command envelope. GAUGE maintains a Bayesian command-to-motion model and selects authorized trials according to their expected reduction of posterior epistemic uncertainty under the planner-induced command distribution. The resulting posterior supports validation-based stopping, bounded inverse compensation, and task-relevant recalibration after detected interface shifts. In three controlled response families, GAUGE reaches the joint criterion for task-facing accuracy and uncertainty with fewer trials than passive, D-optimal, and task-agnostic alternatives. Across six held-out Isaac Sim navigation maps, it meets the declared noninferiority margins against dense calibration. Code is available at https://github.com/EurekaZang/CalibAgent.
Chinese Translation
本文提出了一个目标感知不确定性引导探索(Goal-Aware Uncertainty-Guided Exploration, GAUGE)框架,用于对不透明四足速度接口进行规划器条件化的主动校准。商用四足机器人通常提供平面速度指令,但其底层运动控制器仍不可访问,并可能在指令运动与实际运动之间产生系统性偏差。导航规划器通常只使用指令包络中的一个结构化子集。GAUGE 维护一个贝叶斯指令到运动模型,并根据在规划器诱导的指令分布下后验认知不确定性的预期减少量来选择授权试验。所得后验支持基于验证的停止、有界逆补偿,以及在检测到接口偏移后进行任务相关的重校准。在三个受控响应族中,GAUGE 以比被动、D-最优和任务无关替代方法更少的试验次数,达到面向任务的精度与不确定性的联合准则。在六个留出的 Isaac Sim 导航地图上,它满足相对于密集校准所声明的非劣效性边界。代码可在 https://github.com/EurekaZang/CalibAgent 获取。
cs.RO / 22 / 2609.32155

RecastVLA: From Past Interaction to Future Control with Adaptive Policy States

RecastVLA:基于自适应策略状态,从过去交互到未来控制
Li, Wenbo, Yang, Jun, Chen, Yiteng, Zhang, Wei, Wu, Qingyao
Abstract
Sequential manipulation requires a robot to track what has already happened, even when the current scene no longer reveals it. Policies with explicit history representations make past interactions available as context for current decisions. We ask how action generation itself can form a persistent state for subsequent control. Building on action-side test-time training, RecastVLA maintains an adaptive policy state within a flow-matching vision-language-action policy. The state is represented by shared fast weights and remains fixed throughout action generation. Depth-specific interfaces read the same state, while features across depths and flow evaluations jointly define one update for the next policy call. Subsequent action losses train the initialization, interfaces, and update rule by differentiating through earlier state transitions. At deployment, updates use the policy's own action-generation features without expert action labels. Across LIBERO, RoboTwin, RoboDojo, and twelve real-robot tasks, RecastVLA improves mean success over a matched policy trained without test-time training, including 10.68 percentage points on RoboTwin Clean-to-Clean. In controlled RoboTwin comparisons, retaining state improves success, and the shared design exceeds independently trained layer-local TTT by 2.58 points.
Chinese Translation
顺序操作要求机器人跟踪已经发生的事情,即使当前场景不再显示它。具有显式历史表示的策略使过去的交互可作为当前决策的上下文。我们探究动作生成本身如何为后续控制形成持久状态。基于动作侧测试时训练,RecastVLA 在流匹配视觉-语言-动作策略中维护一个自适应策略状态。该状态由共享快速权重表示,并在整个动作生成过程中保持固定。深度特定的接口读取相同的状态,而跨深度和流评估的特征共同为下一次策略调用定义一次更新。后续动作损失通过对早期状态转换进行微分来训练初始化、接口和更新规则。在部署时,更新使用策略自身的动作生成特征,而无需专家动作标签。在 LIBERO、RoboTwin、RoboDojo 和十二个真实机器人任务上,RecastVLA 相比于训练中未使用测试时训练的匹配策略,提高了平均成功率,包括在 RoboTwin Clean-to-Clean 上提高了 10.68 个百分点。在受控的 RoboTwin 比较中,保留状态提高了成功率,并且共享设计比独立训练的层局部 TTT 高出 2.58 个百分点。
cs.RO / 23 / 2609.32156

AquaBEV-Nav: Learned BEV Occupancy for Underwater Navigation and Exploration

AquaBEV-Nav:面向水下导航与探索的学习型BEV占据
Dong, Trung Tien, Wu, Zhenqi, Kondapalli, Sahasra, Wu, Jiayi, Sheng, Yi, Lin, Xiaomin
Abstract
Safe underwater exploration requires a robot to understand where surrounding structures are located and which regions are available for motion. Existing vision-based underwater exploration systems commonly obtain this information indirectly by estimating monocular depth, unprojecting the geometry into 3D space, and accumulating it into a 2D bird's-eye-view occupancy map. This reliance on intermediate depth estimation is particularly problematic underwater, where scattering and wavelength-dependent attenuation degrade visual cues and limit the reliability of monocular depth estimates. We introduce AquaBEV-Nav, an underwater exploration framework that bypasses explicit monocular depth estimation through direct bird's-eye-view occupancy prediction. Built upon the CORAL hierarchical exploration framework, AquaBEV-Nav replaces its depth-based perception front end with AquaBEV. Given a single RGB frame, AquaBEV maps visual features into a learned polar representation, performs causal reasoning along the range dimension, and reconstructs local Cartesian occupancy without relying on intermediate depth prediction. The resulting occupancy map is accumulated into CORAL's persistent spatial memory, providing spatial context for VLM-based high-level planning and collision constraints for dynamics-aware local trajectory generation. Across ten simulated reef environments and six occupancy backbones evaluated under a single protocol, AquaBEV-Nav reaches 37.48 structure IoU and 53.2 target IoU, 88.95% closed-loop coverage with zero collisions.
Chinese Translation
安全的水下探索要求机器人理解周围结构的位置以及哪些区域可用于运动。现有的基于视觉的水下探索系统通常通过估计单目深度、将几何反投影到三维空间并将其累积到二维鸟瞰图占据图中来间接获取此信息。这种对中间深度估计的依赖在水下尤其成问题,因为散射和波长相关的衰减会降低视觉线索并限制单目深度估计的可靠性。我们提出AquaBEV-Nav,一个水下探索框架,通过直接的鸟瞰图占据预测绕过显式单目深度估计。AquaBEV-Nav建立在CORAL分层探索框架之上,用AquaBEV替换其基于深度的感知前端。给定单个RGB帧,AquaBEV将视觉特征映射到学习到的极坐标表示中,沿距离维度进行因果推理,并在不依赖中间深度预测的情况下重建局部笛卡尔占据。生成的占据图被累积到CORAL的持久空间记忆中,为基于VLM的高层规划提供空间上下文,并为动态感知的局部轨迹生成提供碰撞约束。在十个模拟珊瑚礁环境和六个占据主干网络(在单一协议下评估)中,AquaBEV-Nav达到37.48的结构IoU和53.2的目标IoU,88.95%的闭环覆盖率且零碰撞。
cs.RO / 24 / 2609.32158

FutureRay: Control-Aligned Future Range for Agile Quadruped Navigation

FutureRay:面向敏捷四足导航的控制对齐未来距离
Zang, Tianhao, Wang, Shanze, Wang, Ziqian, Luo, Liyou, Liu, Zihan, Xie, Xingjian, Zhang, Wei
Abstract
Moving obstacles can block a previously clear route while a quadruped robot executes a motion command. We investigate whether predicting changing clearance improves navigation when motion selection accounts for the robot footprint and the time needed to react and brake. We present FutureRay, which predicts ranges across viewing directions and future times, together with encounter risk, from depth-derived range history and observable robot motion. Training emphasizes near-term clearance and penalizes errors that overstate available space. A local planner queries the same forecast for candidate headings and combines it with current observations to check clearance around the robot footprint. Model-based reaction--braking limits guide speed selection, and the resulting velocity commands are passed to a fixed locomotion policy. In paired evaluations on 60 static and dynamic simulation scenes, FutureRay achieves 93.3% completion, compared with 75.0% for current-range persistence and 80.0% for Cartesian Kalman rollout, with perception, planning, and locomotion held fixed. FutureRay also records fewer collisions than both baselines. Qualitative trials on a physical quadruped show avoidance initiated while an obstacle is approaching the route, followed by renewed goal progress. These results show that joint range and encounter-risk prediction can improve obstacle avoidance without retraining the locomotion policy.
Chinese Translation
移动障碍物可能在四足机器人执行运动指令时阻塞原本畅通的路线。我们研究当运动选择考虑机器人足迹以及反应与制动所需时间时,预测变化的间隙是否能改善导航。我们提出 FutureRay,它从由深度信息导出的距离历史和可观测的机器人运动出发,预测跨观测方向和未来时刻的距离,以及相遇风险。训练强调近期间隙,并惩罚高估可用空间的误差。局部规划器针对候选朝向查询同一预测,并将其与当前观测结合,以检查机器人足迹周围的间隙。基于模型的反应—制动限制指导速度选择,所得速度指令传递给固定的运动策略。在 60 个静态和动态仿真场景的配对评估中,在感知、规划和运动控制保持固定的条件下,FutureRay 达到 93.3% 的完成率,而当前距离保持(current-range persistence)为 75.0%,笛卡尔卡尔曼推演(Cartesian Kalman rollout)为 80.0%。FutureRay 记录的碰撞次数也少于两个基线。在物理四足机器人上的定性试验显示,当障碍物接近路线时即启动避让,随后恢复向目标前进。这些结果表明,联合距离与相遇风险预测可以在不重新训练运动策略的情况下改善避障。
cs.RO / 25 / 2609.32236

RoboFFT: Finetuning generative robot policy via online reinforcement learning with forward process

RoboFFT:通过带前向过程的在线强化学习微调生成式机器人策略
Li, Yu, Hu, Shenghe, Wang, Yuhan, Pu, Yaoxiang, Zhang, Haotong, Chen, Yuanpei, Yang, Yaodong
Abstract
Generative models, such as diffusion and flow-based models, have shown strong promise for robot policy learning by capturing complex and multimodal action distributions from demonstrations. However, policies trained solely with imitation learning often suffer from imperfect demonstrations and distributional shifts, while further improvement typically requires additional expert data. Reinforcement learning offers a natural solution through environment interaction, but effectively finetuning generative robot policies remains challenging due to the intractability of likelihood estimation. In this work, we propose RoboFFT, a forward-process reinforcement learning framework for finetuning generative robot policies, which applies forward noising to sampled actions and uses the weighted score / flow matching loss to construct a surrogate policy ratio for PPO-style updates. We evaluate RoboFFT with popular generative robot policies on representative simulation benchmarks, including long-horizon planning and sparse reward settings. Extensive experiments and analysis demonstrate that RoboFFT consistently improves performance while achieving better stability and training efficiency. We further integrate RoboFFT into a real world RL framework and demonstrate its effectiveness in real world tasks. Project website: https://student-of-holmes.github.io/RoboFFT/.
Chinese Translation
生成模型(如扩散模型和基于流的模型)通过从演示中捕捉复杂且多模态的动作分布,已在机器人策略学习方面展现出强大潜力。然而,仅通过模仿学习训练的策略往往受限于不完善的演示和分布偏移,而进一步提升通常需要额外的专家数据。强化学习通过环境交互提供了一种自然的解决方案,但由于似然估计的难处理性,有效微调生成式机器人策略仍然具有挑战性。在本工作中,我们提出 RoboFFT,一种用于微调生成式机器人策略的前向过程强化学习框架,其对采样动作施加前向加噪,并使用加权分数匹配/流匹配损失构建用于 PPO 风格更新的代理策略比率。我们在代表性仿真基准上,使用流行的生成式机器人策略评估 RoboFFT,包括长时程规划和稀疏奖励设置。大量实验和分析表明,RoboFFT 在持续提升性能的同时,实现了更好的稳定性和训练效率。我们进一步将 RoboFFT 集成到真实世界强化学习框架中,并验证了其在真实世界任务中的有效性。项目网站:https://student-of-holmes.github.io/RoboFFT/。
cs.RO / 26 / 2609.32239

Federated Subspace Guided Vision-Language-Action Policy Distillation for Non-IID Multi-Robot Manipulation

联邦子空间引导的视觉-语言-动作策略蒸馏用于非独立同分布多机器人操作
Pal, Biprodip, Roy, Kaushik, Zhu, Yanming, Tidd, Brendan, Liew, Alan Wee-Chung, Moghadam, Peyman
Abstract
Federated learning offers a natural way for multiple robots to jointly improve manipulation policies without requiring centralized access to training demonstrations. However, non-IID task and environment distributions can induce representation drift and mutually incompatible robot-policy updates, making naive parameter aggregation destructive. We present FedDRMan, a federated subspace-guided distillation framework for heterogeneous robot manipulation. At each communication round, the server model provides a frozen teacher for local behavior cloning, while low-rank multimodal subspace and action-distribution distillation preserve globally useful representation geometry and policy behavior. To address heterogeneous aggregation, FedDRMan groups clients by update compatibility and maintains a persistent model for each cluster. The server then spectrally rebalances each compatible aggregate to mitigate attenuation of weaker task-relevant robot-policy update directions. Extensive experiments on LIBERO across diverse non-IID settings, heterogeneity levels, client participation variation, together with ablations and aggregation analyses, show that FedDRMan substantially improves knowledge transfer and consistently outperforms strong federated baselines achieving a peak mean success rate of 80.7%, 11.6 percentage points above the strongest evaluated federated baseline.
Chinese Translation
联邦学习为多个机器人联合改进操作策略提供了一种自然的方式,而无需集中访问训练演示。然而,非独立同分布的任务和环境分布可能导致表示漂移和相互不兼容的机器人策略更新,使得简单的参数聚合具有破坏性。我们提出了FedDRMan,一个用于异构机器人操作的联邦子空间引导蒸馏框架。在每一轮通信中,服务器模型提供一个冻结的教师模型用于本地行为克隆,同时低秩多模态子空间和动作分布蒸馏保留了全局有用的表示几何和策略行为。为了解决异构聚合问题,FedDRMan根据更新兼容性对客户端进行分组,并为每个聚类维护一个持久模型。然后,服务器对每个兼容的聚合进行谱重平衡,以减轻较弱的任务相关机器人策略更新方向的衰减。在LIBERO上进行的大量实验,涵盖多种非独立同分布设置、异构程度、客户端参与变化,以及消融和聚合分析,表明FedDRMan显著提高了知识迁移,并始终优于强大的联邦基线,达到了80.7%的峰值平均成功率,比最强的评估联邦基线高出11.6个百分点。
cs.RO / 27 / 2609.32253

DS-VLA: A Dendritic-inspired Vision-Language-Action Model for Robust Action Control

DS-VLA:一种用于鲁棒动作控制的树突启发视觉-语言-动作模型
Lyu, Yaxing, Li, Jingyi, Xu, Mingkun, Wu, Yujie
Abstract
Vision-language-action (VLA) models have achieved strong performance in language-conditioned manipulation, yet success under nominal evaluation does not necessarily translate into robust closed-loop behavior when executed actions are transiently corrupted. We introduce DS-VLA, a dendritic-inspired action architecture that incorporates dendritic spiking dynamics into VLA control to address this limitation. Specifically, to enable modularized feature processing and temporal information integration, DS-VLA equips action neurons with multiple sparsely connected dendritic branches, each featuring heterogeneous, learned decay factors. Furthermore, to suppress unreliable state updates while preserving task-relevant historical information, we introduce a neuron-wise inhibitory gate that adaptively regulates the admission of new multimodal evidence into dendritic states prior to somatic dynamics. We evaluate DS-VLA on all four LIBERO suites under both nominal rollouts and a unified closed-loop action-perturbation protocol. DS-VLA achieves a 91.6\% average nominal success rate and an 87.35\% average perturbed success rate, retaining 95.4\% of its nominal performance. Under the same reported perturbation setting, OpenVLA-OFT, FAST, $\pi_0$, and GR00T achieve 39.45\%, 23.90\%, 28.55\%, and 30.75\%, respectively. A controlled ablation isolates the contribution of neuron-wise shared inhibition, while analyses of neural dynamics and post-perturbation trajectories associate robust performance with selective evidence suppression and effective behavioral recovery. Together, these results demonstrate that integrating brain-inspired computational mechanisms offers a promising architectural prior for robust embodied intelligence beyond merely scaling vision-language backbones or generative action decoders.
Chinese Translation
视觉-语言-动作(VLA)模型在语言条件化的操作任务中已取得强大性能,然而,在名义评估下的成功并不一定能在执行动作被短暂破坏时转化为鲁棒的闭环行为。我们提出了DS-VLA,一种树突启发的动作架构,将树突脉冲动力学融入VLA控制以解决这一局限。具体而言,为了实现模块化的特征处理和时间信息整合,DS-VLA为动作神经元配备了多个稀疏连接的树突分支,每个分支具有异质的、可学习的衰减因子。此外,为了抑制不可靠的状态更新,同时保留任务相关的历史信息,我们引入了一个神经元级别的抑制门控,在胞体动力学之前自适应地调节新多模态证据进入树突状态的准入。我们在所有四个LIBERO套件上,在名义部署和统一的闭环动作扰动协议下评估了DS-VLA。DS-VLA达到了91.6%的平均名义成功率和87.35%的平均扰动成功率,保持了其名义性能的95.4%。在相同的报告扰动设置下,OpenVLA-OFT、FAST、π0和GR00T分别达到39.45%、23.90%、28.55%和30.75%。一项受控消融实验分离了神经元级别共享抑制的贡献,而对神经动力学和扰动后轨迹的分析将鲁棒性能与选择性证据抑制和有效行为恢复联系起来。总之,这些结果表明,整合类脑计算机制为鲁棒的具身智能提供了一种有前景的架构先验,而不仅仅是扩展视觉-语言主干或生成式动作解码器。
cs.RO / 28 / 2609.32292

Affordance-Conditioned Decision Making: Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision-and-Language Navigation

可供性条件化的决策:弥合零样本跨楼层视觉-语言导航中的语义-空间鸿沟
Yang, Xuekang, Chen, Lu, Luo, Shuang, Zhu, Jialing, Zhang, Qi, Gao, Yue, Zhang, Xiang
Abstract
Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into reliable physical execution remains difficult in spatially constrained transitions. Reaching a staircase, doorway, or narrow passage does not ensure traversal; the agent must identify an executable affordance pose and recover from accumulated action errors. We propose PACE (Preference-refined Affordance-Conditioned Execution), a supervised local execution module that augments frozen zero-shot semantic planners for reliable cross-floor navigation. PACE grounds transition-related semantics into a long-horizon, agent-centric traversable affordance pose and conditions short-horizon action generation on this spatial target, thereby aligning semantic goals with physical execution. We further post-train PACE through failure-aware preference refinement using rollout-derived pairs that contrast normal or recovery behaviors with deviation-amplifying behaviors, thereby improving closed-loop correction. We integrate PACE into six open-source zero-shot VLN navigators and demonstrate consistent improvements on the cross-floor subsets of R2R-CE and RxR-CE, increasing the average success rate from 16.35% to 27.65% and from 4.76% to 12.06%, respectively. Real-world experiments further demonstrate PACE's applicability in unseen environments, highlighting the potential of traversable affordances to bridge semantic intent and reliable embodied behavior.
Chinese Translation
视觉-语言导航日益依赖通用语义规划器,然而在空间受限的过渡区域,将正确的高层意图转化为可靠的物理执行仍然困难。到达楼梯、门口或狭窄通道并不保证能够通过;智能体必须识别可执行的可供性位姿,并从累积的动作误差中恢复。我们提出 PACE(Preference-refined Affordance-Conditioned Execution,偏好精炼的可供性条件化执行),一种监督式局部执行模块,用于增强冻结的零样本语义规划器,以实现可靠的跨楼层导航。PACE 将与过渡相关的语义锚定为以智能体为中心的长时程可通行可供性位姿,并将短时程动作生成条件化于该空间目标,从而使语义目标与物理执行对齐。我们进一步通过失败感知的偏好精炼对 PACE 进行后训练,使用 rollout 导出的成对样本,将正常或恢复行为与放大偏差的行为进行对比,从而改善闭环校正。我们将 PACE 集成到六个开源零样本 VLN 导航器中,并在 R2R-CE 和 RxR-CE 的跨楼层子集上展示出一致的改进,将平均成功率分别从 16.35% 提高到 27.65%,以及从 4.76% 提高到 12.06%。真实世界实验进一步证明了 PACE 在未见环境中的适用性,凸显了可通行可供性在桥接语义意图与可靠具身行为方面的潜力。
cs.RO / 29 / 2609.32313

MemTransfer: Benchmarking Memory Beyond Matched Experience in Embodied Decision-Making

MemTransfer:在具身决策中超越匹配经验的记忆基准测试
Tang, Haiming, Dai, Xianjie, Shao, Gujie, Guo, Zuyi, Li, Jingguang, Ma, Kailang, Tang, Yihong, Huang, Heye
Abstract
Memory lets an embodied agent reuse past experience, yet retaining useful information does not ensure that the agent can apply it when conditions change. We present MemTransfer, a benchmark comparing six memory representations, a working-memory baseline and five representations of past experience, under a shared frozen vision-language-model policy. It comprises 100 navigation cases across ten task types in a simulated warehouse, with expert demonstrations supplying the history. Three comparisons vary the starting pose, route availability, and amount and task relevance of history. With one demonstration per task, Full-context and Episodic memory reach 95.3% and 100.0% success at the original demonstration start, but lose 48-49 percentage points at a new test start. Summary changes little between these two test starts, yet with four demonstrations per task it retains a smaller fraction of its unchanged-route success after blocking (39.3%) than Working memory (44.8%) or the two trajectory memories (56-58%). At the new test start, increasing from one to four relevant demonstrations raises Episodic success by 14.3 percentage points, while the other evaluated representations gain no more than 1.3 percentage points. Replacing half of the relevant histories with other-task experience lowers success for both trajectory memories. These results show that robustness to one kind of mismatch does not imply robustness to another, motivating evaluation of both stored information and its use at decision time.
Chinese Translation
记忆让具身智能体能够重用过去的经验,然而保留有用信息并不能确保智能体在条件变化时能够应用它。我们提出了 MemTransfer,一个在共享的冻结视觉-语言模型策略下,比较六种记忆表示、一个工作记忆基线以及五种过去经验表示的基准测试。它包含模拟仓库中十种任务类型的 100 个导航案例,专家演示提供了历史记录。三组比较分别改变了起始位姿、路径可用性以及历史记录的数量和任务相关性。每个任务一个演示时,全上下文记忆和情景记忆在原始演示起始点分别达到 95.3% 和 100.0% 的成功率,但在新的测试起始点下降了 48-49 个百分点。摘要记忆在这两个测试起始点之间变化不大,然而每个任务四个演示时,在阻塞后其未改变路径成功率的保留比例(39.3%)低于工作记忆(44.8%)或两种轨迹记忆(56-58%)。在新的测试起始点,将相关演示从 1 个增加到 4 个,使情景记忆的成功率提高 14.3 个百分点,而其他评估的表示最多只提高 1.3 个百分点。将一半的相关历史替换为其他任务的经验,会降低两种轨迹记忆的成功率。这些结果表明,对一种不匹配的鲁棒性并不意味着对另一种不匹配的鲁棒性,这促使我们评估存储的信息及其在决策时的使用。
cs.RO / 30 / 2609.32336

WSM-Aware HRI: An IoT-Enhanced Framework for Early Detection and Norm-Guided Repair of Failures with LLM Guidance

WSM-Aware HRI:一种LLM引导的物联网增强框架,用于故障的早期检测与规范引导修复
Zhang, Hanlin, Wang, Yuquan, Zhang, Tianwei, Sun, Zhenglong
Abstract
Human-robot interaction (HRI) failures remain a major barrier to deploying robots in real-world environments. Prior work often treats failures as isolated technical faults or focuses on post-hoc recovery behaviors. In practice, many breakdowns arise because humans and robots operate under inconsistent assumptions about the current world state. We propose WSM-Aware HRI, an IoT-enhanced modular framework that unifies diverse HRI breakdowns as World-State Mismatches (WSMs) between a human's instruction-implied assumptions and a robot's grounded world model built from multimodal perception and digital augmentation. A Large Language Model (LLM) is used to make implicit assumptions explicit, map them to a small set of mismatch types, and specify the evidence needed for verification against the robot's world state. WSM-Aware HRI shifts failure handling from execution-time recovery to proactive mismatch detection during intention formation, enabling interventions guided by safety, norm compliance, and multi-user coordination with transparent explanations. We evaluate mismatch identification in ten everyday cases spanning both visual and latent-state mismatches. The system can accurately produce the expected output results, and ablations show that reliable identification depends on appropriate grounding representations and verification-oriented refinement. These results indicate that treating interaction breakdowns as explicit world-state mismatches enables earlier detection of impending failures and offers a principled mechanism for integrating external evidence and social constraints into human-robot interaction.
Chinese Translation
人机交互(HRI)故障仍然是机器人在真实环境中部署的主要障碍。以往工作通常将故障视为孤立的技术故障,或侧重于事后的恢复行为。实践中,许多故障源于人类和机器人对当前世界状态持有不一致的假设。我们提出WSM-Aware HRI,一个物联网增强的模块化框架,它将各种HRI故障统一为世界状态不匹配(WSM),即人类指令隐含的假设与机器人基于多模态感知和数字增强构建的具身世界模型之间的不匹配。使用大语言模型(LLM)将隐式假设显式化,将它们映射到一小组不匹配类型,并指定针对机器人世界状态进行验证所需的证据。WSM-Aware HRI将故障处理从执行时的恢复转变为意图形成期间的主动不匹配检测,从而能够在安全、规范遵守和多用户协调的指导下进行干预,并提供透明的解释。我们在十个日常案例中评估了不匹配识别,这些案例涵盖了视觉和潜在状态的不匹配。系统能够准确产生预期的输出结果,消融实验表明,可靠的识别依赖于适当的接地表示和面向验证的细化。这些结果表明,将交互故障视为显式的世界状态不匹配,能够更早地检测即将发生的故障,并为将外部证据和社会约束整合到人机交互中提供了一种有原则的机制。
cs.RO / 31 / 2609.32346

Human Motion Prediction for Human-Robot Collaboration

面向人机协作的人体运动预测
Falqueto, Placido, Basei, Elena, Lamon, Edoardo, Perantoni, Giovanni, Saveriano, Matteo, Fontanelli, Daniele, Palopoli, Luigi
Abstract
This paper compares three human motion prediction methods - GMM-based clustering, TransFusion, and Graph-Mixer - evaluating their effectiveness in human-robot collaboration. These predictions are valuable for improving robot motion planning by anticipating human movements and enhancing safety and efficiency in dynamic environments.
Chinese Translation
本文比较了三种人体运动预测方法——基于高斯混合模型(GMM)的聚类、TransFusion和Graph-Mixer——并评估它们在人机协作中的有效性。这些预测通过预判人体运动,有助于改进机器人运动规划,并提升动态环境中的安全性与效率。
cs.RO / 32 / 2609.32354

Proactive Motion Planning for Human-Robot Cooperation

面向人机协作的主动运动规划
Basei, Elena, Lamon, Edoardo, Saveriano, Matteo, Fontanelli, Daniele, Palopoli, Luigi
Abstract
This abstract addresses the incorporation of human motion prediction into proactive and dynamic human-aware motion planning, with the goal of enabling safe collaboration between humans and robots. A deep learning, graph-based model is used to forecast human motion and is integrated into a planning framework. This framework employs a static roadmap along with a time-variant A* algorithm to modify the trajectory of a UR5e manipulator. This method greatly improves human-robot interaction and enables proactive collision avoidance by combining precise motion forecasts with adaptive trajectory planning.
Chinese Translation
本摘要探讨了将人体运动预测融入主动、动态且具有人体感知的运动规划中,旨在实现人与机器人之间的安全协作。采用基于图的深度学习模型来预测人体运动,并将其集成到规划框架中。该框架采用静态路线图以及时变 A* 算法来调整 UR5e 机械臂的轨迹。该方法通过将精确的运动预测与自适应轨迹规划相结合,极大地改善了人机交互,并实现了主动避碰。
cs.RO / 33 / 2609.32416

RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation

RE-0:通过局部同策略蒸馏对具身代码即策略智能体进行可验证递归改进
Zhang, Jiawei, Zhang, Xiangrong, Song, Rui, Zhou, Huanbin, Song, Chengye, Wang, Hongzhou
Abstract
Code-as-Policy agents accomplish long-horizon embodied tasks by generating and executing code, yet continually improving them with teachers that are stronger but not globally reliable remains a key challenge. Existing distillation methods typically treat the teacher's complete behavior as the supervision target and thus misassign training credit on states where the teacher fails. We propose RE-0, a recursively verified policy improvement framework: rather than assuming that the teacher globally outperforms the student, RE-0 requests local corrections from the teacher on the student's own failure histories and checks in the environment whether each correction is genuinely beneficial; verified corrections yield immediate improvement. Building on this, we propose RE-OPD, which turns verified interventions into supervision for on-policy distillation. Only counterfactually verified teacher interventions provide distribution-level supervision, weighted by their measured local benefit, and the improvement they induce is projected back into the standalone student, so both where supervision is applied and how much credit the teacher receives co-evolve with the student policy. We further prove that the student's per-round gain is lower-bounded by its verified intervention gain up to verification and projection error terms. Experiments across multiple Code-as-Policy embodied tasks show that RE-0 improves both teacher-assisted execution and the standalone student, and generalizes to novel robots and scenes.
Chinese Translation
代码即策略智能体通过生成并执行代码来完成长时程具身任务,但利用更强却非全局可靠的教师来持续改进它们仍是一个关键挑战。现有的蒸馏方法通常将教师的完整行为视为监督目标,从而在教师失败的状态上错误分配训练信用。我们提出RE-0,一个递归验证的策略改进框架:RE-0不假设教师在全局上优于学生,而是针对学生自身的失败历史向教师请求局部纠正,并在环境中检查每个纠正是否真正有益;经核验的纠正带来即时改进。在此基础上,我们提出RE-OPD,将经核验的干预转化为同策略蒸馏的监督信号。只有经反事实验证的教师干预才提供分布级监督,并按测得的局部收益加权,其引发的改进被投影回独立学生,因此监督应用的位置和教师获得的信用都随学生策略共同演化。我们进一步证明,学生的每轮收益下界为其经核验干预收益减去验证和投影误差项。在多个代码即策略具身任务上的实验表明,RE-0同时提升了教师辅助执行和独立学生,并能泛化到新颖机器人和场景。
cs.RO / 34 / 2609.32453

DRAM: Delta-rule Recurrent Associative Memory for Robot Manipulation Policies

DRAM:用于机器人操作策略的Delta规则循环联想记忆
Zhao, Xinyu, Shan, Yixiang, Yang, Tao, Lei, Runyu, Zhao, Yiming, Fan, Jiaxin, Feng, Zongbao, Jia, Peng
Abstract
Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a short temporal window. Equipping such policies with long-term memory remains challenging: existing approaches either feed the backbone multi-frame observation windows, which substantially increase inference cost, or rely on pre-defined semantic features, which limit task generality and may also require the retraining of the backbone to adapt to the memory. We introduce DRAM (Delta-rule Recurrent Associative Memory), a plug-and-play memory module that can be attached to a wide range of pretrained robotic policies, endowing them with long-horizon memory without architectural modification or backbone retraining, requiring only task-specific post-training of the memory module and action expert. DRAM maintains a fixed-size associative memory using gated delta-rule linear attention, with a modified update that incorporates all tokens within each frame in parallel. An architecture-agnostic readout integrates historical context into action prediction across different policy architectures. Experiments show that DRAM consistently improves frozen pretrained policies over short-context baselines and alternative compact memory designs, validating its effectiveness as a fixed-size, post-hoc memory module trained with the backbone frozen.
Chinese Translation
机器人操作本质上依赖历史,然而大多数预训练机器人策略仅以当前观测或短时间窗口为条件。为这类策略配备长期记忆仍具挑战性:现有方法要么向骨干网络输入多帧观测窗口,从而显著增加推理成本;要么依赖预定义的语义特征,这限制了任务通用性,并且可能还需要重新训练骨干网络以适应记忆。我们提出 DRAM(Delta-rule Recurrent Associative Memory),一种即插即用的记忆模块,可附加到广泛的预训练机器人策略上,无需修改架构或重新训练骨干网络即可赋予其长时程记忆,仅需对记忆模块和动作专家进行任务特定的后训练。DRAM 使用门控 Delta 规则线性注意力维护固定大小的联想记忆,并采用修改后的更新方式,可并行纳入每一帧内的所有 token。一种架构无关的读出机制将历史上下文集成到跨不同策略架构的动作预测中。实验表明,DRAM 在短上下文基线以及替代的紧凑记忆设计上持续提升冻结的预训练策略,验证了其作为一种在骨干网络冻结下训练的固定大小、事后记忆模块的有效性。
cs.RO / 35 / 2609.32471

GlowTact: Simple and Compact Vision-Based Tactile Sensing with High Sensitivity and Spatial Resolution

GlowTact:简单紧凑、高灵敏度与高空间分辨率的基于视觉的触觉传感
Ma, Yuxiang, Tippur, Megha, Ye, Pengfei, Liu, Sandra Q., Chen, Haonan, Cottrell, Francis Richard, Adelson, Edward
Abstract
Vision-based tactile sensors (VBTS) provide rich contact information for robotic manipulation, but existing designs can be hard to simplify and adapt to the size and constraints of humanoid fingertips. We introduce \textbf{GlowTact}, a pressure-responsive vision-based tactile sensing mechanism that directly visualizes contact pressure. \textbf{GlowTact} requires only single-color, non-directional illumination, and the raw tactile image directly represents the pressure distribution without explicit geometry reconstruction. This simple sensing principle enables compact, customizable tactile sensors while preserving high sensitivity and rich spatial detail. We demonstrate gram-scale contact detection, accurate normal-force estimation, and reconstruction of fine contact geometry, including M1 screw threads. These results establish \textbf{GlowTact} as a practical new sensing technology for compact humanoid fingertips, combining a durable nitrile membrane and simple optical design with sensitive and information-rich tactile perception.
Chinese Translation
基于视觉的触觉传感器(VBTS)为机器人操作提供了丰富的接触信息,但现有设计难以简化并适配人形指尖的尺寸与约束。我们提出 GlowTact,一种压力响应型基于视觉的触觉传感机制,可直接可视化接触压力。GlowTact 仅需单色、非定向照明,原始触觉图像即可直接表示压力分布,而无需显式几何重建。这一简单的传感原理使得触觉传感器紧凑且可定制,同时保持高灵敏度和丰富的空间细节。我们展示了克级接触检测、精确的法向力估计以及精细接触几何的重建,包括 M1 螺纹。这些结果确立了 GlowTact 作为一种适用于紧凑型人形指尖的实用新型传感技术,将耐用的丁腈膜和简单光学设计与灵敏且信息丰富的触觉感知相结合。
cs.RO / 36 / 2609.32576

Assisting for Open-Ended Tasks: Goal-Oriented Shared Autonomy as a Particle Filter

面向开放式任务的辅助:目标导向的共享自主作为粒子滤波
Fu, Mengxue, Xu, Ethan, Iyer-Singh, Sam, Dai, Yinlong, Hagenow, Michael, Losey, Dylan P.
Abstract
A common approach for shared autonomy blends human inputs with autonomous assistance based on the human's likely goal. However, most existing approaches assume that a static set of possible goals is known a priori, which limits the use of such methods in unstructured assistive settings. We instead investigate how to enable shared autonomy with open-ended and dynamically changing goals. We formulate goal-oriented shared autonomy as a particle filter in which particles represent candidate human goals. Unlike conventional approaches with a fixed goal set, our transition model dynamically proposes new candidate goals as the interaction evolves, and human actions update the belief over these goals in real time. We instantiate this framework with foundation models (e.g., vision grounding and large language models) that propose context-relevant semantic goals, generate goal-conditioned assistance from low-level skill primitives, and refine those skills from human corrections. We assess our approach through a user study where 12 participants perform a variety of tabletop manipulation tasks with our method and state-of-the-art shared autonomy baselines. The results show that our particle filter-based approach reduces the amount of time users spend teleoperating the system and improves user satisfaction. User study videos: https://youtu.be/Ii26XuRqm9c
Chinese Translation
共享自主的一种常见方法基于人类可能的目标,将人类输入与自主辅助相融合。然而,大多数现有方法假设一组静态的可能目标是先验已知的,这限制了此类方法在非结构化辅助环境中的应用。我们转而研究如何实现具有开放式和动态变化目标的共享自主。我们将目标导向的共享自主形式化为一个粒子滤波,其中粒子表示候选的人类目标。与固定目标集的传统方法不同,我们的转移模型随着交互的演进动态地提出新的候选目标,并且人类动作实时更新对这些目标的信念。我们用基础模型(例如,vision grounding 和大语言模型)实例化该框架,这些模型提出上下文相关的语义目标,从低级技能原语生成目标条件辅助,并从人类纠正中细化这些技能。我们通过一项用户研究评估我们的方法,其中12名参与者使用我们的方法和最先进的共享自主基线执行各种桌面操作任务。结果表明,我们基于粒子滤波的方法减少了用户远程操作系统所花费的时间,并提高了用户满意度。用户研究视频:https://youtu.be/Ii26XuRqm9c
cs.RO / 37 / 2609.32591

Think Fast, Plan Selectively: Adaptive Deliberation for Efficient Data-Driven MPC

快速思考,选择性规划:面向高效数据驱动MPC的自适应审议
Goh, Yi Xian, Yang, Sze Jue, Luan, Hao
Abstract
Data-driven model predictive control (MPC) combines learned world models with online trajectory optimization, achieving strong performance in continuous control. However, the per-step cost of sampling and evaluating hundreds of candidate trajectories restricts deployment to control frequencies well below what real-time robotics demands. Motivated by the dual-process theory of human cognition, which distinguishes between fast, intuitive processing (System 1) and slower, deliberative reasoning (System 2), we ask whether every decision requires the same degree of computational deliberation. We propose Fast-TD-MPC, a lightweight framework that adaptively routes between fast policy execution and test-time planning, reserving costly deliberation for states where it is most needed. Fast-TD-MPC delivers competitive task performance across 103 continuous control tasks while achieving up to ~4x faster inference. Under external disturbances, Fast-TD-MPC selectively falls back to planning, maintaining robustness comparable to the original planner.
Chinese Translation
数据驱动模型预测控制(MPC)将学习到的世界模型与在线轨迹优化相结合,在连续控制中取得了强大性能。然而,每步采样和评估数百条候选轨迹的成本限制了部署的控制频率,远低于实时机器人技术的要求。受人类认知的双过程理论启发,该理论区分了快速、直觉的处理(系统1)和较慢、审慎的推理(系统2),我们探讨是否每个决策都需要相同程度的计算审议。我们提出Fast-TD-MPC,一个轻量级框架,能够在快速策略执行和测试时规划之间自适应路由,将昂贵的审议保留给最需要它的状态。Fast-TD-MPC在103个连续控制任务上提供了有竞争力的任务性能,同时实现了高达约4倍的推理速度提升。在外部干扰下,Fast-TD-MPC选择性地回退到规划,保持了与原始规划器相当的鲁棒性。
cs.RO / 38 / 2609.32595

RECAST: Recasting Vision-Language Semantics into an Actionable Cost Map for Robot Navigation

RECAST:将视觉-语言语义重铸为机器人导航的可操作代价地图
Cho, Incheol, Park, Jintae, Kim, Jinkyu, Lee, Jungbeom, Choo, Jaegul, Moon, Seokha
Abstract
Safe and robust robot navigation across diverse environments requires a high-level understanding of complex scenes and the ability to carry it into stable motion. Recent works tackle this with learning-based models trained at scale and with approaches built on vision-language models (VLMs). However, learning-based models break down outside their training distribution, while VLM-based approaches bring that understanding but rarely ground it in the scene or align the action with it. To address these limitations, we present RECAST, a robot navigation framework that combines the reasoning of a VLM with the spatial grounding of vision foundation models to build an Actionable Cost map. Given the robot's front view and the user's instruction, we first decompose the scene with the VLM, judging which surfaces are traversable, which objects pose a risk, which heading to prefer, and which gaps are passable. Vision foundation models then ground these surfaces and objects in the image, and all four judgments are spatially recast into one compact cost map. This map both conditions the trajectory decoders and scores their proposals to select the one to execute. As the VLM's answers trail the live scene, both steps draw on cost maps from two points in time: the pivot frame the VLM judged, which carries all four judgments, and the current frame, whose terrain and collision costs are rebuilt from the current image. RECAST improves success over the strongest prior method by 13.3 points in simulation and 31.4 points on a real quadruped, and reduces the collision rate relative to it by 9.6 and 14.3 points, reaching the lowest collision rate among all methods. The project page is available at https://recast-nav.github.io/
Chinese Translation
安全鲁棒的机器人导航在多样化环境中需要对复杂场景的高层次理解,以及将其转化为稳定运动的能力。近期工作通过大规模训练的基于学习的模型和基于视觉语言模型(VLM)的方法来解决这个问题。然而,基于学习的模型在训练分布之外会失效,而基于VLM的方法带来了这种理解,但很少将其落地到场景中或将动作与之对齐。为了解决这些局限,我们提出RECAST,一个机器人导航框架,它结合了VLM的推理与视觉基础模型的空间定位,来构建可操作代价地图。给定机器人前视图和用户指令,我们首先用VLM分解场景,判断哪些表面可通行、哪些物体有风险、偏好哪个朝向、以及哪些间隙可通过。然后视觉基础模型在图像中定位这些表面和物体,四个判断在空间上被重铸为一个紧凑的代价地图。该地图既作为轨迹解码器的条件,也对其提议进行评分以选择要执行的那个。由于VLM的答案滞后于实时场景,两个步骤都利用两个时间点的代价地图:VLM判断的枢轴帧(承载所有四个判断)和当前帧(其地形和碰撞代价从当前图像重建)。RECAST在仿真中比最强先前方法成功率提高13.3个百分点,在真实四足机器人上提高31.4个百分点,并相对于它分别降低碰撞率9.6和14.3个百分点,达到所有方法中最低的碰撞率。项目页面见 https://recast-nav.github.io/
cs.RO / 39 / 2609.32626

World SLAM Model: Joint World Modeling for SLAM and Navigation

World SLAM Model:面向SLAM与导航的联合世界建模
Qin, Minghui, Yuan, Yijun, Zheng, Weicheng, Li, Kenan, Wang, Weibang, Sun, Chang, Huang, Junhao, Liu, Anmin, Yao, Yicheng, Zhao, Hang
Abstract
We introduce World SLAM Model (WSM), a unified framework that brings the SLAM paradigm directly into downstream navigation. Rather than treating SLAM merely as an upstream module that provides poses, maps or tokens, WSM adopts its core mechanisms, including incremental state updates with persistent memory and backend refinement of accumulated errors, to maintain a consistent world state during interaction. Given the current observation and a navigation goal, WSM predicts future visual states and jointly estimates their camera motion and dense geometry, grounding visual prediction in an evolving spatial world state. This spatial state is continuously updated as new observations arrive and provides the basis for action generation and closed-loop navigation. WSM is trained end-to-end with a joint navigation--SLAM objective, enabling downstream navigation to benefit directly from SLAM-style state maintenance and refinement while preserving accurate geometric estimation. Experiments demonstrate improved navigation performance together with strong SLAM accuracy, highlighting the potential of SLAM as an intrinsic mechanism for long-horizon world modeling and embodied interaction.
Chinese Translation
我们提出了World SLAM Model (WSM),一个将SLAM范式直接引入下游导航的统一框架。WSM并非仅将SLAM视为提供位姿、地图或token的上游模块,而是采用其核心机制,包括带有持久记忆的增量状态更新和累积误差的后端优化,以在交互过程中保持一致的世界状态。给定当前观测和导航目标,WSM预测未来视觉状态,并联合估计其相机运动和稠密几何,将视觉预测建立在不断演化的空间世界状态之上。随着新观测的到来,该空间状态持续更新,并为动作生成和闭环导航提供基础。WSM以联合导航-SLAM目标进行端到端训练,使下游导航能够直接受益于SLAM式的状态维护和优化,同时保持准确的几何估计。实验表明,导航性能得到提升,同时SLAM精度也很高,凸显了SLAM作为长时程世界建模和具身交互内在机制的潜力。
cs.RO / 40 / 2609.32634

PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models

PF-RL:面向视觉-语言-动作模型的基于目标条件价值几何的进度场强化学习
Qing, Yunpeng, Kong, Yilun, Lin, Sixu, Zhou, Ming, Fei, Yiming, Luo, Shuang, Chi, Yixiao, Gu, Haoming, Liu, Jingyuan, Wei, Changxu, Hou, Zhi, Zou, Changqing
Abstract
Reinforcement Fine-Tuning~(RFT) has emerged as a promising paradigm for improving Vision-Language-Action~(VLA) policies, yet sparse task-level outcomes provide limited credit for intermediate transitions, especially in long-horizon manipulation. A natural approach is to model intermediate task progress and use it as dense feedback for policy improvement. Despite their architectural differences, existing progress-aware methods commonly formulate task progress as an explicit scalar prediction, providing limited structure for modeling how intermediate observations relate to the task goal, which may hinder effective transition-level credit assignment. We introduce Progress Field Reinforcement Learning (PF-RL), which learns a structured goal-conditioned progress representation over pretrained VLA features and converts it into dense credit for policy optimization. A lightweight shared Progress Field head maps current and goal representations into a compact progress space, where geometric distance induces goal-conditioned value, while complementary temporal and goal-structure objectives shape the learned geometry. Transition-level value changes naturally yield dense progress advantages, enabling fine-grained credit assignment for both offline policy improvement and online reinforcement fine-tuning. Extensive experiments on LIBERO, RoboTwin2.0, and real-world bimanual manipulation tasks show that PF-RL consistently improves policy performance over strong supervised fine-tuning, reinforcement fine-tuning, and progress-aware baselines.
Chinese Translation
强化微调(RFT)已成为改进视觉-语言-动作(VLA)策略的一种有前景的范式,然而稀疏的任务级结果对中间转移提供的信用有限,尤其是在长视野操作中。一种自然的方法是建模中间任务进度,并将其作为策略改进的密集反馈。尽管架构不同,现有的进度感知方法通常将任务进度表述为显式标量预测,为建模中间观察与任务目标的关系提供了有限的结构,这可能阻碍有效的转移级信用分配。我们引入了进度场强化学习(PF-RL),它在预训练的 VLA 特征上学习结构化的目标条件进度表示,并将其转换为用于策略优化的密集信用。一个轻量级的共享进度场头将当前和目标表示映射到紧凑的进度空间,其中几何距离诱导目标条件价值,而互补的时间目标和目标结构目标塑造学习到的几何。转移级价值变化自然产生密集的进度优势,为离线策略改进和在线强化微调提供细粒度的信用分配。在 LIBERO、RoboTwin2.0 和真实世界双手操作任务上的大量实验表明,PF-RL 在强监督微调、强化微调和进度感知基线之上持续提高策略性能。
cs.RO / 41 / 2609.32698

SEES: A Self-Evolving Embodied System via Failure-Guided VLA Policy Adaptation

SEES:一种通过失败引导的 VLA 策略自适应实现的自进化具身系统
Li, Ziwen, Zhang, Hanlue, Ren, Zhenyang, Huang, Tianyu, Lin, Runqi, Wang, Haoyu, Gao, Zhengqing, Guo, Yandong, Karray, Fakhri, Liu, Tongliang, Russell, Chris, Gong, Mingming
Abstract
Recent vision-language-action (VLA) policies demonstrate promising generalization across diverse short-horizon tasks. However, they remain unreliable on long-horizon tasks, partly because the large-scale training data is biased toward single-stage manipulation tasks that are cheaper to demonstrate. A single weak atomic skill can cause failures across multiple multi-stage tasks. To address such failures, existing methods often require experts to identify the bottleneck and provide additional demonstrations, making the improvement costly and potentially impractical after deployment. To this end, we present a Self-Evolving Embodied System (SEES) that learns from failures and improves the VLA policy without additional expert demonstrations. SEES decomposes long-horizon tasks into atomic tasks and routes them to corresponding family policies. Each family consists of related atomic skills that share one VLA adapter. During execution, the system automatically monitors atomic-task outcomes to identify the most frequently failing atomic skills as the current bottlenecks. To overcome these bottlenecks, SEES constructs tailored RL tasks in simulation by restoring previously encountered states and generating task-specific success criteria with an LLM. Online RL updates the shared family adapters to promote positive transfer among related atomic skills and cumulative improvement across evolution rounds. Extensive experiments show that SEES can be integrated with different VLA backbones to progressively improve their long-horizon performance. We also observe continued improvement on unseen tasks, providing evidence of transfer beyond the evolution settings.
Chinese Translation
最近的视觉-语言-动作 (VLA) 策略在不同短时程任务上展现出有前景的泛化能力。然而,它们在长时程任务上仍然不可靠,部分原因是大规模训练数据偏向于演示成本更低的单阶段操作任务。单个薄弱的原子技能可能导致多个多阶段任务失败。为了解决此类失败,现有方法通常需要专家识别瓶颈并提供额外演示,这使得改进成本高昂,并且在部署后可能不切实际。为此,我们提出了一种自进化具身系统 (SEES),它能够从失败中学习并改进 VLA 策略,而无需额外的专家演示。SEES 将长时程任务分解为原子任务,并将它们路由到相应的家族策略。每个家族由相关的原子技能组成,它们共享一个 VLA 适配器。在执行过程中,系统自动监控原子任务的结果,以识别最频繁失败的原子技能作为当前瓶颈。为了克服这些瓶颈,SEES 通过在模拟中恢复先前遇到的状态,并使用 LLM 生成任务特定的成功标准,从而构建定制的 RL 任务。在线 RL 更新共享的家族适配器,以促进相关原子技能之间的正向迁移,以及跨进化轮次的累积改进。大量实验表明,SEES 可以与不同的 VLA 主干网络集成,逐步提高其长时程性能。我们还观察到在未见任务上的持续改进,这提供了超越进化设置的可迁移性证据。
cs.RO / 42 / 2609.32745

MORPH: Self-Organising Multi-Robot Task Allocation via Neuroplasticity-Inspired Adaptive Topology

MORPH:通过受神经可塑性启发的自适应拓扑实现自组织多机器人任务分配
Niu, Xuezhi, Broo, Didem Gürdür
Abstract
Multi-robot task allocation (MRTA) in dynamic environments faces a fundamental tension: effective coordination requires learned structure, but that structure must adapt when conditions change. Existing methods resolve this by assuming prior task knowledge, a utility function, a cost matrix, or a trained policy making them brittle when deployed without such knowledge or when task distributions shift. We present MORPH(Multi-agent Online Rewiring through Plasticity-guided Hierarchy), a training-free MRTA framework where global allocation quality emerges from 4 local plasticity rules (synaptic, homeostatic, structural, and metaplasticity) applied to a directed pairwise preference matrix updated from runtime co-occurrence and task-completion feedback. MORPH requires no task model, no bid computation, and no offline training; response decisions use learned AGV-to-Picker preferences rather than a fixed proximity rule. Within the Gerkey-Mataric MRTA taxonomy, MORPH is the first method in the single-task, single-robot, instantaneous-assignment class to learn directed pairwise allocation preferences online. Evaluated on the TA-RWARE warehouse benchmark (8-24 agents, 4 maps, 800 steps per episode, 5 seeds), MORPH achieves 110% of all-to-all throughput at N=24 while using only 21% of possible coordination links as an efficiency advantage that grows monotonically with fleet size. Under spatial task distribution shift, MORPH degrades 3x less than proximity-based methods while its learned preferences remain uncorrelated with Manhattan distance. Systematic ablation confirms all four plasticity rules contribute measurably. Two allocation properties emerge without programming: cross-type preference dominance and progressive preference sparsification, mirroring the developmental refinement of biological neural circuits. Learned preferences are driven by task co-occurrence history, not spatial proximity.
Chinese Translation
动态环境中的多机器人任务分配(MRTA)面临一个根本性矛盾:有效的协调需要学习到的结构,但当条件变化时,该结构必须能够适应。现有方法通过假设先验任务知识、效用函数、成本矩阵或训练好的策略来解决这一问题,这使得它们在缺乏此类知识或任务分布发生变化时部署时变得脆弱。我们提出 MORPH(Multi-agent Online Rewiring through Plasticity-guided Hierarchy,基于可塑性引导层次结构的多智能体在线重连),一种无需训练的 MRTA 框架,其中全局分配质量源于 4 条局部可塑性规则(突触可塑性、稳态可塑性、结构可塑性和元可塑性),应用于一个有向成对偏好矩阵,该矩阵根据运行时共现和任务完成反馈进行更新。MORPH 不需要任务模型、不需要投标计算,也不需要离线训练;响应决策使用学习到的 AGV 到拣货员偏好,而非固定的邻近规则。在 Gerkey-Mataric MRTA 分类法中,MORPH 是单任务、单机器人、即时分配类别中第一个在线学习有向成对分配偏好的方法。在 TA-RWARE 仓库基准(8-24 个智能体,4 张地图,每回合 800 步,5 个随机种子)上评估,MORPH 在 N=24 时达到全对全吞吐量的 110%,同时仅使用 21% 的可能协调链路,这种效率优势随着车队规模单调增长。在空间任务分布变化下,MORPH 的性能下降幅度仅为基于邻近方法的 1/3,而其学习到的偏好与曼哈顿距离不相关。系统性消融实验证实所有四种可塑性规则都有可测量的贡献。两个分配特性在没有编程的情况下出现:跨类型偏好主导和渐进式偏好稀疏化,这反映了生物神经回路的发展性精炼。学习到的偏好由任务共现历史驱动,而非空间邻近性。
cs.RO / 43 / 2609.32762

An Empirical Study on What Matters for Viewpoint-Generalizable Policies in Visual Imitation Learning

视觉模仿学习中视角可泛化策略关键因素的实证研究
Nakura, Mino, Krishna, Sriram, Wang, Yufei, Tulsiani, Shubham, Erickson, Zackory, Held, David
Abstract
Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations, making deployment in diverse environments a challenge. We present a controlled empirical study of which design choices allow visuomotor policies to generalize across viewpoints. We find that viewpoint generalization improves when dense visual tokens are retained and the action head participates in geometric reasoning. On a suite of simulated tasks that span a wide range of camera poses, we show that these design choices yield a policy that remains performant across viewpoints. As a practical consequence, a policy trained with these design choices also transfers zero-shot from simulation to the real world under random camera configurations.
Chinese Translation
视觉模仿学习是一种有前景的方法,可用于训练能够完成各种任务的机器人操作策略。然而,当前的策略对视角扰动仍然脆弱,使得在不同环境中的部署成为一项挑战。我们提出了一项受控的实证研究,探讨哪些设计选择能使视觉运动策略在不同视角下泛化。我们发现,当保留密集视觉标记且动作头参与几何推理时,视角泛化能力会提高。在一系列涵盖广泛相机位姿的模拟任务上,我们表明这些设计选择产生的策略在不同视角下保持高性能。作为一个实际结果,使用这些设计选择训练的策略也能够在随机相机配置下零样本从模拟迁移到真实世界。
cs.RO / 44 / 2609.32767

CLAP: Closed-Loop Alignment with Pressure for Precise Suction Manipulation

CLAP:用于精确吸盘操作的带压力闭环对齐
Zou, Yixian, Xu, Chongyang, Xin, Yuling, Feng, Ziliang, Meng, Fanman, Liu, Shuaicheng
Abstract
Stacking and palletising demand precise placement: error left in one layer is inherited by the next, and a flat pad offers no feature to funnel a wrong pose into the right one. Top-down suction suits such dense arrangements, and suction has already been brought into vision-language-action (VLA) policies. What that work does not report, however, is a policy conditioned on a measured vacuum signal, or one that uses it to abandon an action already under way. Vision does not settle the question here, because at the moment it matters the cup and the face it holds occlude each other. We present CLAP, which makes the attachment state observable through a pressure module tapped into the vacuum line. The decoded reading replaces the suction command in the policy's proprioception, is fused with the visual features, and terminates the open-loop execution window so that the policy re-infers from a fresh observation. For data, we record goal-state disassembly on the physical robot and reverse the joint-state sequence offline, without a simulation replay. Targeted phase demonstrations, 8.3% of the training frames, cover the suction transitions and the configurations an interrupted grasp leaves behind. On a real Unitree Z1, one multi-task checkpoint reaches 96.67% average success in both colour settings, 16.67 and 10.00 points above the strongest baseline, its monochromatic four-block successes averaging 15.92 mm of error. Four ablation settings fall 5.00 to 13.33 points short. We will release code and trained weights.
Chinese Translation
堆叠和码垛要求精确放置:一层留下的误差会被下一层继承,而平坦的吸盘没有特征来将错误的位姿引导到正确的位姿。自上而下的吸盘适合这种密集排列,并且吸盘已经被引入视觉-语言-动作(VLA)策略中。然而,该工作没有报告的是,一个以测量的真空信号为条件的策略,或者一个使用它来放弃已经进行的动作的策略。视觉在这里不能解决这个问题,因为在关键时刻,吸盘和它所吸附的表面相互遮挡。我们提出CLAP,它通过接入真空管路的压力模块使吸附状态可观测。解码后的读数取代了策略本体感受中的吸力命令,与视觉特征融合,并终止开环执行窗口,以便策略从新的观测中重新推理。对于数据,我们在物理机器人上记录目标状态拆解,并离线反转关节状态序列,无需模拟回放。针对性阶段演示(占训练帧的8.3%)涵盖了吸力转换和中断抓取留下的配置。在真实的Unitree Z1上,一个多任务检查点在两种颜色设置下达到96.67%的平均成功率,比最强基线分别高出16.67和10.00个百分点,其单色四块成功平均误差为15.92毫米。四个消融设置分别低了5.00到13.33个百分点。我们将发布代码和训练权重。
cs.RO / 45 / 2609.32779

Copper-Policy: Focus on the Representation for Robust Robot Manipulation

Copper-Policy:聚焦面向鲁棒机器人操作的表示
Feng, Zexin, Feng, Yixu, Xiao, Lingyu, Su, Shang, Zheng, Kexin, Xu, Chang, Shi, Mengkai, Feng, Shuo, Yan, Xintao
Abstract
World Action Models (WAMs) acquire behavioral priors by modeling future scene evolution, but predicting detailed futures in pixel or latent space incurs substantial cost. Recent evidence that co-training gains persist without test-time generation raises a question: what must a WAM learn to improve control? We introduce Copper-Policy, which learns a compact World representation with the policy rather than relying on a predefined target space. Through temporal joint-embedding prediction, it predicts future observation embeddings conditioned on task intention without reconstructing pixels. This prediction and action decoding shape the representation jointly, while the policy retains access to current-frame spatial detail for execution. Representation analyses show that the learned features better separate task-driven change from perturbations and provide complementary information for control. Compact prediction targets reduce training tokens per sample, enabling a 2B-parameter model trained in 9.67 hours on 8$\times$ RTX 5090 GPUs and 6$\times$ faster than Fast-WAM on matched A100 GPUs. Copper-Policy outperforms every compared method without embodied pretraining on RoboTwin and several embodied-pretrained VLAs on LIBERO-Plus (80.85%). On three challenging real-robot tasks, it performs comparably to $\pi_{0.5}$ and attains a higher average score. Together, these results show that Copper-Policy combines strong control performance with efficient training.
Chinese Translation
世界动作模型(World Action Models, WAMs)通过对未来场景演化建模来获得行为先验,但在像素空间或潜在空间中预测详细未来会带来高昂成本。近期证据表明,协同训练带来的收益在测试时无需生成的情况下依然存在,这引出一个问题:WAM 必须学习什么才能改进控制?我们提出 Copper-Policy,它通过与策略一起学习紧凑的世界表示,而不是依赖预定义的目标空间。通过时间联合嵌入预测,它在不重建像素的情况下,以任务意图为条件预测未来观测嵌入。该预测与动作解码共同塑造表示,同时策略仍可获取当前帧的空间细节以用于执行。表示分析表明,学习到的特征能更好地区分任务驱动变化与扰动,并为控制提供互补信息。紧凑的预测目标减少了每个样本的训练 token 数,使得一个 20 亿参数模型能在 8 块 RTX 5090 GPU 上于 9.67 小时内完成训练,并在匹配的 A100 GPU 上比 Fast-WAM 快 6 倍。Copper-Policy 在 RoboTwin 上无需具身预训练即优于所有对比方法,并在 LIBERO-Plus 上优于若干具身预训练的 VLA(80.85%)。在三项具有挑战性的真实机器人任务上,其表现与 π_{0.5} 相当,并获得更高的平均分。总之,这些结果表明 Copper-Policy 将强大的控制性能与高效训练结合起来。
cs.RO / 46 / 2609.32783

CollisionGAT: Controller-Agnostic One-Step Collision Screening for Multi-Agent Motion

CollisionGAT:面向多智能体运动的控制器无关单步碰撞筛查
Debbas, Alan, Meriaux, Edwin, Dudek, Gregory
Abstract
Before a team of robots moves, each proposed step must be checked for collisions with other robots and with obstacles. We present CollisionGAT, a graph-attention network that reads the current and proposed states of moving agents together with locally relevant stationary obstacles and returns one collision-risk score per moving agent. Any controller can use these scores to accept, repair, replan, or postpone a proposed step. We mount CollisionGAT on a continuous path-following controller and on GATeD, an obstacle-blind D* Lite planner that uses typed vetoes to update its planning graphs. Exact geometric checks supply the training labels and independently audit every executed step.
Chinese Translation
在一组机器人移动之前,必须检查每个提议步是否与其他机器人以及障碍物发生碰撞。我们提出 CollisionGAT,一种图注意力网络,它读取移动智能体的当前状态和提议状态,以及局部相关的静止障碍物,并为每个移动智能体返回一个碰撞风险评分。任何控制器都可以利用这些评分来接受、修复、重新规划或推迟某个提议步。我们将 CollisionGAT 搭载到连续路径跟随控制器以及 GATeD 上;GATeD 是一种障碍物盲的 D* Lite 规划器,它使用带类型的否决来更新其规划图。精确的几何检查提供训练标签,并独立审计每个已执行的步骤。
cs.RO / 47 / 2609.32806

Towards Kinematic Actionable Infeasibility Detection in Motion Planning

面向运动规划中运动学可操作的不可行性检测
Rath, Aayush, Jindal, Lakshya, Thomas, Antony
Abstract
Motion planning in robotics requires not only computing collision-free paths but also certifying infeasibility when no such path exists. Complete methods are limited to low-dimensional spaces, while sampling-based planners scale efficiently but cannot provide finite-time infeasibility certificates, leaving this problem largely unresolved in high-dimensional spaces. In this letter, we present a geometry-driven framework for certifying infeasibility through an explicit resolution-dependent analysis of configuration space topology. Leveraging signed distance field representations, the proposed method traces separating manifolds induced by obstacle boundaries directly in configuration space, enabling both detection of infeasibility and identification of the specific geometric cause. To address computational challenges, we develop a parallel frontier-expansion algorithm that exploits GPU acceleration for efficient simplicial reconstruction in high-dimensional spaces. We validate the approach on 4-DOF and 5-DOF robot scenarios, certifying infeasibility within seconds for 4-DOF cases and under four minutes for 5-DOF cases. We further discuss avenues for improving scalability to higher-dimensional spaces.
Chinese Translation
机器人运动规划不仅需要计算无碰撞路径,还需要在不存在此类路径时证明不可行性。完备方法局限于低维空间,而基于采样的规划器虽能高效扩展,但无法提供有限时间的不可行性证明,使得该问题在高维空间中很大程度上仍未解决。在本文中,我们提出了一种几何驱动的框架,通过显式地依赖于分辨率的构型空间拓扑分析来证明不可行性。利用符号距离场表示,所提方法直接在构型空间中追踪由障碍物边界诱导的分隔流形,从而既能检测不可行性,又能识别具体的几何原因。为了解决计算挑战,我们开发了一种并行前沿扩展算法,利用GPU加速在高维空间中进行高效的单纯形重建。我们在4自由度和5自由度机器人场景上验证了该方法,对于4自由度情况可在数秒内证明不可行性,对于5自由度情况在四分钟内完成。我们进一步讨论了提高向更高维空间可扩展性的途径。
cs.RO / 48 / 2609.32837

Scanning While Imagining: A Scene-Graph World Model for Robotic Ultrasound Navigation

边想象边扫描:用于机器人超声导航的场景图世界模型
Li, Xuesong, Chen, Shuai, Li, Feng, Jiang, Zhongliang, Navab, Nassir, Bi, Yuan
Abstract
Ultrasound (US) acquisition depends on the operator's ability to interpret anatomy and anticipate how the view will change with probe motion. Many robotic US navigation methods select actions without explicitly predicting these anatomical changes. We propose SonoGraph-WM, an action- and goal-conditioned world model for anticipatory probe navigation. The model represents anatomy as scene graphs (SGs), capturing visible structures, their geometry, and spatial relationships without synthesizing US images. Given a history of SGs and probe poses, a unified Transformer jointly predicts future SGs and poses. A receding-horizon planner recursively imagines candidate trajectories, selects the shortest predicted path reaching a goal graph, and follows it over a short execution horizon before replanning from new observations. To reduce reliance on tracked and anatomically annotated US sequences, we generate aligned SG--pose training data from computed tomography (CT) label maps along surface-constrained probe trajectories. On four held-out CT cases, spatial relation F1 remains above 93% over 20 prediction steps, and closed-loop navigation achieves 77.50% and 75.00% success for the gallbladder and pancreas, respectively, using annotation-derived SGs. In robot--phantom navigation experiments with label-map-derived SGs, the planner reached the target view in 73.7% of trials. These findings support CT-supervised anatomical world modeling for probe planning and highlight the importance of frequent observation updates for reliable navigation. Project Page: https://noseefood.github.io/us-sonograph-wm/
Chinese Translation
超声(US)采集依赖于操作者解读解剖结构并预判视图随探头运动变化的能力。许多机器人超声导航方法在选择动作时并未显式预测这些解剖变化。我们提出SonoGraph-WM,一个用于预期性探头导航的动作和目标条件世界模型。该模型将解剖结构表示为场景图(SG),捕获可见结构、其几何形态及空间关系,而无需合成超声图像。给定场景图历史和探头位姿,一个统一的Transformer联合预测未来的场景图和位姿。滚动时域规划器递归地想象候选轨迹,选择到达目标图的最短预测路径,并在短执行时域内跟随该路径,然后根据新观测重新规划。为了减少对追踪和标注解剖结构的超声序列的依赖,我们沿着表面约束的探头轨迹从计算机断层扫描(CT)标签图生成对齐的场景图-位姿训练数据。在四个留出的CT病例上,空间关系F1在20个预测步长内保持93%以上,使用标注衍生的场景图,闭环导航在胆囊和胰腺上分别达到77.50%和75.00%的成功率。在使用标签图衍生场景图的机器人-体模导航实验中,规划器在73.7%的试验中到达了目标视图。这些发现支持用于探头规划的CT监督解剖世界建模,并强调了频繁观测更新对可靠导航的重要性。项目页面:https://noseefood.github.io/us-sonograph-wm/
cs.RO / 49 / 2609.32855

FINE: Future-Informed Navigation Encoding for Data-Efficient Vision-Language Navigation

FINE:面向数据高效视觉语言导航的未来信息引导导航编码
Nguyen, Khang H., Nguyen, Hoang Pham Quang, Nguyen, Ha Phuong, Binh, Khanh Dinh, Nguyen, Xuan Ha, Ngo, Vien, Minh, Duy Ho Nguyen, Nguyen, Huan, Le, An T.
Abstract
Adapting vision-language navigation (VLN) policies to new environments is expensive because every additional route and instruction requires an embodied demonstration. Yet standard observation-to-action training uses only a small fraction of the information already contained in each trajectory. In particular, future observations reveal the instruction-relevant landmarks that the agent will encounter, including what they look like and how they are arranged in 3D. We introduce FINE, a Future-Informed Navigation Encoding framework that extracts this latent supervision from existing demonstrations. FINE equips a VLN backbone with two complementary auxiliary representations. First, explicit landmark tokens follow the ordered landmarks specified by the instruction and are trained to predict future landmark regions in both semantic 2D patch-feature space and viewpoint-dependent 3D geometric feature space. Second, an implicit future token learns to distinguish the landmark state that is actually reached from plausible same-scene counterfactual futures generated by a video world model. On R2R-CE and RxR-CE val-unseen, FINE improves InternVLA-N1 by 2.6 and 4.5 success-rate points, respectively, at full training data. More importantly, as demonstrations become limited, the benefit grows: at a 70% demonstration budget, FINE improves success rate by 6.8 points, recovering roughly one-third of the performance lost by reducing the training demonstrations. Project page is available at https://finevln.github.io/.
Chinese Translation
将视觉语言导航(VLN)策略适应到新环境成本高昂,因为每条额外的路线和指令都需要具身演示。然而,标准的观测到动作训练仅利用了每条轨迹中已包含信息的一小部分。特别地,未来观测揭示了智能体将会遇到的与指令相关的地标,包括它们的外观以及它们在3D中的排列方式。我们提出了FINE,一个未来信息引导的导航编码框架,从现有演示中提取这种潜在监督。FINE为VLN主干网络配备了两个互补的辅助表示。首先,显式地标token遵循指令指定的有序地标,并被训练以在语义2D块特征空间和视点依赖的3D几何特征空间中预测未来地标区域。其次,一个隐式未来token学习区分实际到达的地标状态与由视频世界模型生成的合理的同一场景反事实未来。在R2R-CE和RxR-CE val-unseen上,使用完整训练数据时,FINE将InternVLA-N1的成功率分别提高了2.6和4.5个百分点。更重要的是,当演示数据受限时,收益更大:在70%的演示预算下,FINE将成功率提高了6.8个百分点,恢复了因减少训练演示而损失的大约三分之一的性能。项目页面见:https://finevln.github.io/。
cs.RO / 50 / 2609.32862

RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents

RoboFoundry:面向自学习具身智能体的系统即策略演化
Liang, Jingsong, Liao, Shuhao, Zhang, Shizhe, Hou, Diyuan, Cai, Yuxin, Deng, Xinjian, He, Chengyang, Huang, Wenhui, Tan, Runjia, Wang, Zhidong, Yu, Lan, Tian, Xuesong, Sartoretti, Guillaume, Luo, Jie, Mu, Yao, Wu, Wenjun, Li, Wanhua, Lv, Chen
Abstract
A foundation model should not act in isolation as an embodied agent. Yet, existing methods often optimize individual components of the agent stack, such as memory, context, skills, or action interfaces, rather than treating the supporting system itself as a unified policy. Moreover, interaction alone does not yield self-improvement unless execution experience is converted into persistent, validated system changes. We therefore propose RoboFoundry, the first embodied agentic framework that formulates this process as Self-Evolving System-as-Policy. RoboFoundry diagnoses capability gaps in decision-making and memory management, converts execution traces into validated task-specific system updates, and promotes recurring improvements to the general system. Evolution operates over two complementary surfaces: a context system that manages active internal context and persistent file-system memory, and a hierarchical skill system that organizes atomic skills, reusable compositions, and failure-conditioned recovery. A shared semantic interface separates embodiment-invariant decisions from embodiment-specific execution, allowing evolved system capabilities to transfer across heterogeneous robots. On EmbodiedBench, RoboFoundry achieves state-of-the-art performance, notably improving GPT-5.5 by 27.8%. It also brings Qwen3.7-Plus to near parity with GPT-5.5 (70.3% vs. 72.7%), showing consistent gains from system-as-policy evolution across foundation models. For long-horizon memory, RoboFoundry outperforms all baselines on RoboMemArena by at least 39.0%, even against methods assisted by external foundation models. On LIBERO-PRO, it further outperforms Cap-Agent0 by 243.8%-679.7% across all perturbation types. In real-world deployments, RoboFoundry demonstrates zero-shot transfer and online evolution across robots and tasks, highlighting its potential for fully autonomous embodied agents.
Chinese Translation
基础模型不应作为具身智能体孤立行动。然而,现有方法通常优化智能体栈的各个组件,如记忆、上下文、技能或动作接口,而不是将支撑系统本身视为统一策略。此外,仅靠交互无法产生自我改进,除非将执行经验转化为持久、经过验证的系统变更。因此,我们提出 RoboFoundry,这是首个将这一过程形式化为自演化系统即策略的具身智能体框架。RoboFoundry 诊断决策和记忆管理中的能力差距,将执行轨迹转化为经过验证的任务特定系统更新,并将反复出现的改进提升到通用系统。演化在两个互补层面进行:一个管理活跃内部上下文和持久化文件系统记忆的上下文系统,以及一个组织原子技能、可复用组合和失败条件恢复的分层技能系统。共享语义接口将具身无关的决策与具身特定的执行分离,使演化的系统能力能够在异构机器人之间迁移。在 EmbodiedBench 上,RoboFoundry 实现了最先进的性能,显著提升 GPT-5.5 达 27.8%。它还使 Qwen3.7-Plus 接近 GPT-5.5 的水平(70.3% 对 72.7%),表明系统即策略演化在基础模型上带来一致增益。对于长时程记忆,RoboFoundry 在 RoboMemArena 上超越所有基线至少 39.0%,即使与由外部基础模型辅助的方法相比也是如此。在 LIBERO-PRO 上,它在所有扰动类型上进一步超越 Cap-Agent0 达 243.8%-679.7%。在真实世界部署中,RoboFoundry 展示了跨机器人和任务的零样本迁移和在线演化,突显了其实现完全自主具身智能体的潜力。
cs.RO / 51 / 2609.32935

Communication-Aware Heterogeneous Graph Learning for Decentralized Multi-Human Multi-Robot Task Allocation

面向去中心化多人多机器人任务分配的通信感知异构图学习
Yuan, Ziqin, Wang, Ruiqi, Yang, Baijian, Min, Byung-Cheol
Abstract
Multi-human multi-robot (MH-MR) teams combine robotic autonomy with human expertise, but effective task allocation requires coordinating scarce, dynamically available human support with distributed robot execution. Limited robot-robot and human-robot communication further complicates this coupling by delaying information exchange and supervisory intervention. We introduce CommHG, a communication-aware heterogeneous graph learning framework for decentralized MH-MR task allocation. CommHG represents coupled human-robot-task interactions through local graphs conditioned on information availability and age. Learned communication actions enable robots to decide when and which operator to query, and when to share information with peers. Allocation and communication are jointly optimized through cooperative multi-agent reinforcement learning, allowing the team to acquire useful information while managing limited communication and supervisory resources. We also introduce a benchmark integrating heterogeneous humans and robots, dynamic tasks and operational states, and constrained robot-robot and bidirectional human-robot links. Experiments across heterogeneous teams with up to 16 robots, 6 humans, and 112 tasks show that CommHG improves timely weighted mission completion as coordination scale increases, outperforming the strongest baseline in the Large scenario.
Chinese Translation
多人多机器人(MH-MR)团队将机器人自主性与人类专业知识相结合,但有效的任务分配需要协调稀缺的、动态可用的人类支持与分布式机器人执行。有限的机器人-机器人以及人-机器人通信通过延迟信息交换和监督干预进一步加剧了这种耦合的复杂性。我们介绍了 CommHG,一个用于去中心化 MH-MR 任务分配的通信感知异构图学习框架。CommHG 通过以信息可用性和信息年龄为条件的局部图来表示耦合的人-机器人-任务交互。学习到的通信动作使机器人能够决定何时以及向哪个操作员查询,以及何时与同伴共享信息。通过协作多智能体强化学习联合优化分配和通信,使团队能够在管理有限的通信和监督资源的同时获取有用信息。我们还引入了一个基准测试,集成了异构人类和机器人、动态任务和操作状态,以及受限的机器人-机器人链路和双向人-机器人链路。在最多 16 个机器人、6 个人类和 112 个任务的异构团队上进行的实验表明,随着协调规模的增加,CommHG 提高了及时加权任务完成度,在大型场景中优于最强基线。
cs.RO / 52 / 2609.32958

Learning Geometry-Aware Virtual Fixtures From Sparse Demonstrations

从稀疏演示中学习几何感知虚拟夹具
Mühlbauer, Maximilian, Piacentino, Marcella, Mancino, Raffaella, De Risi, Paolino, Weber, Bernhard, Campos, Margarida, Hulin, Thomas, Calinon, Sylvain, Stulp, Freek, Albu-Schäffer, Alin, Klodmann, Julian, Ficuciello, Fanny, Silvério, João
Abstract
In many teleoperation applications, collecting a large number of demonstrations as required for traditional probabilistic learning from demonstration (LfD) approaches may not be feasible. To still give operators the ability to intuitively create trajectories as Virtual Fixtures (VFs), we propose to leverage a motion prior in the learning process. Particularly, by using Linear Quadratic Tracking (LQT), users are able to define guiding trajectories from the demonstration of just a few via points. To account for orientation guidance, we further reformulate classical LQT on Riemannian manifolds, introducing an additional geometric prior. Through a probabilistic interpretation of the LQT solution, we derive a covariance estimate at each trajectory point which we use to modulate the stiffness of the resulting fixture, resulting in strong guidance around the via points and softer guidance when far away. The covariance information is also used to define a validity region of the fixture, allowing the operator to leave its influence area and conduct unmodeled tasks. We evaluate the proposed Riemannian LQT formulation in a set of toy examples and the full framework on a cutting task requiring high precision in a minimally invasive surgery setting on the da Vinci Research Kit (dVRK).
Chinese Translation
在许多遥操作应用中,收集传统概率示教学习(LfD)方法所需的大量演示可能并不可行。为了仍然让操作员能够以虚拟夹具(VFs)的形式直观地创建轨迹,我们提出在学习过程中利用运动先验。具体而言,通过使用线性二次跟踪(LQT),用户能够仅通过少量途经点的演示来定义引导轨迹。为了考虑姿态引导,我们进一步在黎曼流形上重新构建经典 LQT,引入额外的几何先验。通过对 LQT 解的概率解释,我们推导出每个轨迹点的协方差估计,并利用它来调节所得虚拟夹具的刚度,从而在途经点附近提供强引导,而在远离时提供较柔和的引导。协方差信息还用于定义虚拟夹具的有效区域,允许操作员离开其影响范围并执行未建模的任务。我们在一组简单示例中评估了所提出的黎曼 LQT 公式,并在达芬奇研究套件(dVRK)上针对微创手术中需要高精度的切割任务评估了完整框架。
cs.RO / 53 / 2609.33000

TriDrive: Joint Driver, Vehicle, and Road Modeling for Forecasting and Driver Monitoring

TriDrive:面向预测与驾驶员监控的驾驶员、车辆与道路联合建模
Wang, Yuhang, Yang, Jingxin, Wei, Chuheng, Guo, Yuechen, Xu, Jinghan, Han, Zhao, Zhou, Hao
Abstract
Predicting how drivers, vehicles, and road scenes interact and evolve together is central to driver monitoring. Prior work models in-cabin activity or traffic-conditioned driver motion in isolation, motivating joint driver, vehicle, and road modeling with real-time on-vehicle evaluation. We introduce TriDrive, to our knowledge the first unified framework that jointly forecasts driver kinematics, vehicle dynamics, and road demands through an automation-conditioned transition model. Modality-specific encoders (an anchored kinematic representation of the driver, causal CAN-bus dynamics, and frozen V-JEPA 2 road latents with structured road margins) are connected by directed residual connections through which driver and road context refine vehicle forecasts. We evaluate TriDrive on three downstream tasks. On the public AIDE benchmark, its kinematic encoder recipe sets a new full-set state of the art (SOTA) among published baselines (48.05 versus 71.47 All-MPJPE). On 197.2 hours of naturalistic BATON subset, directed connections and road margins raise assistance-engaged PR-AUC by 0.084 for steering onset and 0.286 for time-to-collision drops. For real-time use, we distill the road encoders and run TriDrive on a comma four with an external 8 GB GPU, where a lightweight current-state warning probe updates at 5 Hz with 177 ms p95 latency while the joint model forecasts concurrently. The probe is above an openpilot-based baseline on human-labeled manual-driving warnings (AUROC 0.725 versus 0.563), and in a paired on-road study 14 drivers rate its warnings as more appropriate (+1.79) and timely (+2.67) than those of openpilot's driver-monitoring system.
Chinese Translation
预测驾驶员、车辆与道路场景如何交互并共同演化,是驾驶员监控的核心。先前工作孤立地建模车内活动或受交通条件影响的驾驶员运动,这促使我们进行驾驶员、车辆与道路的联合建模,并开展实时车载评估。我们提出TriDrive,据我们所知,这是首个通过自动化条件转移模型联合预测驾驶员运动学、车辆动力学和道路需求的统一框架。特定模态的编码器(驾驶员的锚定运动学表示、因果CAN总线动力学,以及具有结构化道路边缘的冻结V-JEPA 2道路隐变量)通过有向残差连接相连,驾驶员和道路上下文通过该连接来细化车辆预测。我们在三个下游任务上评估TriDrive。在公开的AIDE基准上,其运动学编码器方案在已发表基线中创下了新的全集最优(SOTA)(48.05 对 71.47 All-MPJPE)。在197.2小时的自然驾驶BATON子集上,有向连接和道路边缘将辅助介入的PR-AUC分别提升了0.084(转向起始)和0.286(碰撞时间下降)。为了实时使用,我们蒸馏道路编码器,并在带有外部8 GB GPU的comma four上运行TriDrive,其中轻量级当前状态警告探针以5 Hz更新,p95延迟为177 ms,同时联合模型并行进行预测。在人工标注的手动驾驶警告上,该探针优于基于openpilot的基线(AUROC 0.725 对 0.563),并且在一项配对实路研究中,14名驾驶员认为其警告比openpilot的驾驶员监控系统更合适(+1.79)和更及时(+2.67)。
cs.RO / 54 / 2609.33007

CAPEX: Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive Reasoning

CAPEX:通过经验自适应推理高效地将基础模型行为蒸馏为可部署的机器人策略
Aarya, Shivam, Xi-Jia, Zhang, Huang, Chengyue, Kim, Junhyun, Xue, Huishu, Leen, Hrishit, Yakunin, Roman, Garg, Animesh, Kira, Zsolt
Abstract
Robot learning has largely relied on human-teleoperated demonstrations to acquire effective learnable behaviors. However, human-operated data collection processes can be unintuitive, difficult to scale, and inherently asynchronous. We explore an alternative: distilling physical behavior from general-purpose multimodal foundation models into deployable robot policies by using the foundation model itself as an autonomous demonstrator. While sufficiently capable models can generate successful zero-shot manipulation trajectories, repeatedly invoking them during physical execution is slow and expensive, limiting their utility as scalable data generators. As a solution, we introduce CAPEX, an experience-conditioned demonstration collection framework that uses execution experience from previous attempts to adapt how frequently the foundation model must observe, reason, and replan. We evaluate across RoboCasa tasks and on physical Franka and bimanual YAM-arm platforms, measuring task success, model calls, token usage, collection time, and cost. We further train Diffusion Policy and ACT on matched sets of human-teleoperated and foundation-model-generated demonstrations to evaluate the downstream learning value of autonomously collected data. We find that CAPEX increases the number of successful demonstrations by 4.3x while reducing the cost per successful demonstration by 80%. Policies trained on CAPEX-generated data approach the performance of those trained on matched human demonstrations; with longer training, this gap largely closes for policies trained from scratch. These results suggest that foundation models can serve as scalable sources of reusable robot experience. Project page: https://capex-paper.github.io/
Chinese Translation
机器人学习在很大程度上依赖于人类遥操作的演示来获取有效的可学习行为。然而,人工操作的数据收集过程可能不直观、难以扩展,并且本质上是异步的。我们探索了一种替代方案:通过将基础模型本身作为自主演示者,从通用多模态基础模型中蒸馏出物理行为,并转化为可部署的机器人策略。虽然足够强大的模型可以生成成功的零样本操作轨迹,但在物理执行过程中反复调用它们既缓慢又昂贵,限制了它们作为可扩展数据生成器的实用性。作为一种解决方案,我们提出了CAPEX,一种基于经验的演示收集框架,它利用先前尝试中的执行经验来适应基础模型必须观察、推理和重新规划的频率。我们在RoboCasa任务以及物理Franka和双臂YAM-arm平台上进行评估,测量任务成功率、模型调用次数、token使用量、收集时间和成本。我们进一步在匹配的人类遥操作和基础模型生成的演示集上训练Diffusion Policy和ACT,以评估自主收集数据在下游学习中的价值。我们发现,CAPEX将成功演示的数量增加了4.3倍,同时将每个成功演示的成本降低了80%。在CAPEX生成的数据上训练的策略接近在匹配的人类演示上训练的策略的性能;随着训练时间延长,对于从头训练的策略,这一差距基本消失。这些结果表明,基础模型可以作为可重用机器人经验的可扩展来源。项目页面:https://capex-paper.github.io/
cs.RO / 55 / 2609.33049

Is Online Interaction Necessary for Recovery? A Minimalist Approach to Robust Planning via Perturbation

在线交互对于恢复是必要的吗?一种通过扰动的鲁棒规划极简方法
Park, Bumgeun, Lee, Donghwan
Abstract
Behavior cloning (BC) is vulnerable to covariate shift during closed-loop execution, where small prediction or execution errors can drive the robot toward states poorly covered by the demonstration data. We focus on action-sequence planning, where a policy predicts a finite-horizon sequence of actions as a reference trajectory for robot execution. Existing approaches to covariate shift often rely on collecting additional corrective demonstrations, requiring further environment interaction and access to an expert or reference policy. We propose Perturbation-Augmented Recovery Supervision (PARS), a simple training approach for improving recovery using only existing demonstrations. PARS perturbs the robot's proprioceptive state and anchors only the later portion of the predicted trajectory to the original demonstration, leaving the earlier portion free to generate corrective motion and recover toward the demonstrated behavior within the planning horizon. Unlike conventional input-noise augmentation, which preserves the original supervision target over the entire prediction horizon, PARS explicitly provides trajectory-level supervision for recovery from perturbed states. PARS requires neither additional environment interaction nor expert queries and can be instantiated across different action-sequence policy classes with only minor modifications to their BC objectives. Experiments on 51 RLBench manipulation tasks with flow-based, transformer-based, and diffusion-based policies demonstrate that PARS improves robustness to covariate shift during closed-loop execution.
Chinese Translation
行为克隆(BC)在闭环执行期间容易受到协变量偏移的影响,此时小的预测或执行误差可能将机器人推向示范数据覆盖较差的状态。我们关注动作序列规划,其中策略预测有限时域的动作序列,作为机器人执行的参考轨迹。现有的协变量偏移方法通常依赖于收集额外的纠正示范,需要进一步的环境交互以及访问专家或参考策略。我们提出扰动增强恢复监督(PARS),一种简单的训练方法,仅使用现有示范来提高恢复能力。PARS扰动机器人的本体感知状态,并仅将预测轨迹的后期部分锚定到原始示范,使早期部分自由生成纠正运动,并在规划时域内恢复到示范行为。与传统的输入噪声增强不同,后者在整个预测时域内保持原始监督目标,PARS显式地提供轨迹级监督以从扰动状态恢复。PARS既不需要额外的环境交互,也不需要专家查询,并且可以通过对其BC目标进行微小修改,实例化到不同的动作序列策略类别中。在51个RLBench操作任务上,使用基于流、基于Transformer和基于扩散的策略进行的实验表明,PARS提高了闭环执行期间对协变量偏移的鲁棒性。
cs.RO / 56 / 2609.33053

SwingRL: Adaptive Observation Reinforcement Learning with World-Model Prediction for Cable-Suspended Hoisting Control

SwingRL:面向缆索悬吊吊装控制的自适应观测强化学习与世界模型预测
Wang, Guangming, Zhang, Xiaoyu, Xin, Yucheng, Ma, Wanli, Liu, Jiucai, Ma, Yunxiang, Ingham, Joe, Wu, Haibing, Jing, Yixiong, Wysocki, Olaf, Sheil, Brian
Abstract
Cable-suspended hoisting is widely used to move heavy or bulky payloads that cannot be handled conveniently by rigid pick-and-place systems, for example in crane-assisted construction. Robotic hoisting using flexible cables is challenging because payload motion is underactuated, external disturbances vary, and delayed or lost visual observations can make the perceived payload state stale at control execution. These effects are particularly critical during precise insertion of a suspended payload's sockets onto rebar pins, which is a very common task in construction environments. We present SwingRL, a residual reinforcement-learning (RL) framework that combines an age-aware world model, a classical anti-swing prior, and a recurrent residual policy to address two coupled problems: stale feedback and uncertain dynamics. The world model propagates the newest received payload observation to the current control step using the executed commands, providing a time-aligned state estimate under delayed and lossy sensing. The prior supplies nominal tracking and swing damping. The residual policy learns bounded corrections to the prior rather than the complete control law, compensating for system-parameter variation, external disturbances, and remaining state-estimation errors. We evaluate SwingRL against classical and learning-based baselines across a cumulative difficulty ladder covering system-parameter variation, wind disturbance, degraded sensing, and strong gusts. Under the most difficult setting, SwingRL achieves 69.5% strict and 77.3% broad success, exceeding all baselines by at least 60 percentage points, respectively. World-model ablations support the role of time-aligned state estimation in maintaining insertion success as observation loss increases. Finally, without real-robot fine-tuning, SwingRL achieves 90% success on the physical rig.
Chinese Translation
缆索悬吊吊装广泛用于搬运刚性拾放系统难以方便处理的重型或大型负载,例如在起重机辅助施工中。使用柔性缆索的机器人吊装具有挑战性,因为负载运动是欠驱动的,外部扰动多变,且视觉观测的延迟或丢失会使感知到的负载状态在控制执行时变得过时。这些效应在将悬吊负载的套筒精确插入钢筋销钉的过程中尤为关键,这是建筑环境中非常常见的任务。我们提出了 SwingRL,一种残差强化学习(RL)框架,它结合了时效感知世界模型、经典防摆先验和循环残差策略,以解决两个耦合问题:过时反馈和不确定动力学。世界模型利用已执行的指令将最新接收到的负载观测传播到当前控制步,在延迟和有损传感下提供时间对齐的状态估计。先验提供标称跟踪和摆动阻尼。残差策略学习对先验的有界修正,而非完整的控制律,以补偿系统参数变化、外部扰动和剩余状态估计误差。我们在涵盖系统参数变化、风扰动、传感退化和强阵风的累积难度阶梯上,将 SwingRL 与经典和基于学习的基线进行了评估。在最困难的设置下,SwingRL 分别达到了 69.5% 的严格成功率和 77.3% 的宽松成功率,分别超过所有基线至少 60 个百分点。世界模型消融实验支持了时间对齐状态估计在观测丢失增加时维持插入成功率的作用。最后,无需真实机器人微调,SwingRL 在物理平台上达到了 90% 的成功率。
cs.RO / 57 / 2609.33095

REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening

REALM:一种用于具身反应性聆听的由粗到细生成框架
Li, Peizhen, Cao, Longbing, Zhang, Yang
Abstract
Generating responsive listener facial motion is an important task for embodied conversational AI. Two modeling challenges are central: accounting for the timing of speaker cues while maintaining continuity with the listener's ongoing motion, and capturing locally variable facial events alongside the overall motion trajectory. Listener responses may follow preceding cues with a temporal lag, while brief expressions and blinks introduce variation that is difficult to predict deterministically. These challenges motivate a framework that combines history-aware temporal alignment with stochastic expression refinement. We propose REALM (Reactive Embodied Audio-driven Listening Model), a coarse-to-fine framework for audio-driven reactive listening. A Reactive Gated Speaker-Listener Fusion module combines listener motion history with speaker audio through a delay-centered attention prior and adaptive gating. A coarse decoder predicts a base motion trajectory, which is augmented by audio-conditioned stochastic residuals in the expression subspace while retaining the coarse pose parameters. Evaluations on ViCo and L2L show improvements over the evaluated baselines across multiple motion-quality metrics. Additional analyses examine delay sensitivity, gate behavior, and blink dynamics. Finally, deployment on an Ameca humanoid robot and a perceptual user study demonstrate the applicability of the generated behavior to physical embodiment. Code: https://github.com/lipzh5/REALM Demo: https://youtu.be/Tf5mpd5S8VQ
Chinese Translation
生成具有响应性的聆听者面部运动是具身对话式AI的一项重要任务。两个核心建模挑战是:在保持与聆听者当前运动连续性的同时考虑说话者线索的时序,以及在整体运动轨迹之外捕捉局部可变的面部事件。聆听者反应可能以时间滞后跟随先前线索,而短暂表情和眨眼会引入难以确定性预测的变化。这些挑战促使构建一个将历史感知的时间对齐与随机表情细化相结合的框架。我们提出REALM(Reactive Embodied Audio-driven Listening Model,反应性具身音频驱动聆听模型),一个用于音频驱动反应性聆听的由粗到细框架。一个反应性门控说话者-聆听者融合模块通过以延迟为中心的注意力先验和自适应门控,将聆听者运动历史与说话者音频相结合。粗解码器预测基础运动轨迹,该轨迹在表情子空间中由音频条件随机残差增强,同时保留粗姿态参数。在ViCo和L2L上的评估表明,在多个运动质量指标上相较于所评估的基线均有提升。额外分析考察了延迟敏感性、门控行为和眨眼动态。最后,在Ameca人形机器人上的部署以及一项感知用户研究证明了生成行为在物理具身中的适用性。代码:https://github.com/lipzh5/REALM 演示:https://youtu.be/Tf5mpd5S8VQ
cs.RO / 58 / 2609.33096

Humanoids for Robot-Assisted Surgery: Bimanual Base Placement and Tool-Mount Optimization via Capability Maps

人形机器人用于机器人辅助手术:基于能力图的双臂基座放置与工具安装优化
Zhang, Peihan, Liang, Zekai, Richter, Florian, Thareja, Nikita, Broderick, Ryan, Liu, Shanglei, Yip, Michael
Abstract
Rapid advances in humanoid robotics have motivated growing interest in the application of humanoids for healthcare and clinical tasks. However, it remains unclear how close contemporary humanoids are to meeting the kinematic demands of robot-assisted laparoscopic surgery. In this work, we address the question of optimal robot positioning through a quantitative analysis of workspace and robot setup configurations. We present a capability-map-based robot setup framework that optimizes humanoid base placement and tool mounting orientation to maximize bimanual humanoid reachability while accounting for tool-tip kinematics and remote-center-of-motion (RCM) constraints. We evaluate three humanoid platforms spanning different body dimensions and kinematic redundancy on workspace reachability for three representative general surgery procedures: cholecystectomy, inguinal hernia repair, and sleeve gastrectomy. The proposed joint optimization of base placement and tool mounting consistently outperforms base-only optimization and heuristic baselines. For cholecystectomy and inguinal hernia repair, which are characterized by relatively small and minimally overlapping workspaces, humanoid reachability approached 90%. For the larger, overlapping multi-port arm workspace of sleeve gastrectomy, humanoids yield substantially lower coverage. These results quantify the near-term promise of humanoids for selected laparoscopic procedures and clarify key limitations that must be addressed for broader deployment.
Chinese Translation
人形机器人技术的快速进步激发了人们对将人形机器人应用于医疗保健和临床任务的日益增长的兴趣。然而,当代人形机器人距离满足机器人辅助腹腔镜手术的运动学需求还有多远,仍不清楚。在这项工作中,我们通过对工作空间和机器人设置配置的定量分析,解决了最优机器人定位的问题。我们提出了一种基于能力图(Capability Maps)的机器人设置框架,该框架优化人形机器人基座放置和工具安装方向,以最大化双臂人形机器人可达性,同时考虑工具尖端运动学和远心运动(RCM)约束。我们评估了三个具有不同身体尺寸和运动学冗余的人形机器人平台在三种代表性普通外科手术(胆囊切除术、腹股沟疝修补术和袖状胃切除术)中的工作空间可达性。所提出的基座放置和工具安装的联合优化始终优于仅基座优化和启发式基线。对于胆囊切除术和腹股沟疝修补术,其特点是工作空间相对较小且重叠最小,人形机器人可达性接近90%。对于袖状胃切除术更大、重叠的多孔臂工作空间,人形机器人产生的覆盖范围显著较低。这些结果量化了人形机器人在选定腹腔镜手术中的近期前景,并阐明了更广泛应用必须解决的关键限制。
cs.RO / 59 / 2609.33101

Evolving Dexterous Robots from Scratch

从零开始进化灵巧机器人
Guo, Zihan, Zhang, Shuzhe, Li, Muhan, Li, Peiyang, Kriegman, Sam
Abstract
Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines. Here we evolve freeform robots to pick up, hold, rotate, and use diverse objects. Unlike other approaches to optimizing robot hands, we do not presuppose the presence, articulation, or geometry of any part of the body. Although familiar prehensile forms such as tails, beaks, paws and claws may emerge spontaneously under certain conditions--and while such conditions could be of interest to evolutionary biologists--de novo manipulator design can also reveal whole new solutions, overlooked or unknown structures which may be better suited for the task at hand. We use contrastive learning to create a highly searchable genetic embedding of design space, an autoregressive developmental model to decode designs, evolutionary strategies to find good designs, and reinforcement learning to train each evolved design. Winning designs were automatically converted into a manufacturable blueprint, printed, assembled and tested in the real world in a zero-shot manner. The results represent the state-of-the-art in evolutionary robotics in terms of performance, diversity and complexity.
Chinese Translation
对于如何手动设计能够进行灵巧操作的智能体,我们知之甚少。一些设计原则是从仔细研究动物如何操作物体中推断出来的,但这些结构和行为迄今为止一直难以被仿生模仿,并且可能不是人工机器的最优解。在这里,我们进化自由形态机器人来拾取、握住、旋转和使用各种物体。与其他优化机器人手的方法不同,我们不预设身体任何部分的存在、关节连接或几何形状。虽然诸如尾巴、喙、爪和爪子等熟悉的抓握形式在某些条件下可能会自发出现——尽管这些条件可能对进化生物学家感兴趣——但从头设计操作器也可以揭示全新的解决方案,即被忽视或未知的结构,这些结构可能更适合当前任务。我们使用对比学习来创建设计空间的易于搜索的遗传嵌入,使用自回归发育模型来解码设计,使用进化策略来寻找好的设计,并使用强化学习来训练每个进化设计。获胜的设计被自动转换为可制造的蓝图,打印、组装,并以零样本方式在现实世界中进行测试。这些结果在性能、多样性和复杂性方面代表了进化机器人学的最先进水平。
cs.RO / 60 / 2609.33104

VPTwin: Real-Sim-Real Video Prediction for Robotic Manipulation Planning

VPTwin:面向机器人操作规划的真实-仿真-真实视频预测
Xiao, Zhenghao, Pan, Minting, He, Nantian, Zhou, Dongzhan, Wang, Yunbo
Abstract
While action-conditioned video prediction provides an intuitive world model for robotics, purely data-driven predictors often suffer from compounding errors and physically implausible hallucinations in long-horizon rollouts, severely undermining downstream action planning. We propose VPTwin, a Real-Sim-Real video prediction framework that anchors real-world future prediction using real-synchronized simulation twins. For a target manipulation task, a VLM reconstructs an executable digital twin from a real demonstration episode. To accommodate the ill-posed estimation of unobserved physical properties, Isaac Sim simulates multiple forward dynamic rollouts across randomized physical configurations under candidate action trajectories. Using these rollouts as in-context references, VPTwin harmonizes both domains, using simulation dynamics to enforce physical plausibility while capturing unmodeled contact interactions from real video. Furthermore, we establish a predictive planning loop using VPTwin to visually verify VLM-proposed actions and guide reliable real-world execution. Evaluations show substantial reductions in physical hallucinations during video prediction and marked improvements in manipulation planning performance.
Chinese Translation
尽管以动作为条件的视频预测为机器人学提供了一种直观的世界模型,但纯数据驱动的预测器在长时域推演中常常遭受复合误差和物理上不合理的幻觉,严重损害下游动作规划。我们提出 VPTwin,一种真实-仿真-真实视频预测框架,利用真实同步的仿真孪生来锚定真实世界的未来预测。对于目标操作任务,VLM 从真实演示片段中重建可执行的数字孪生。为了适应未观测物理属性的不适定估计,Isaac Sim 在候选动作轨迹下跨随机物理配置模拟多个前向动态推演。使用这些推演作为上下文参考,VPTwin 协调两个域,利用仿真动力学来强制物理合理性,同时从真实视频中捕获未建模的接触交互。此外,我们建立了一个使用 VPTwin 的预测规划循环,以视觉验证 VLM 提议的动作并指导可靠的真实世界执行。评估表明,在视频预测过程中物理幻觉大幅减少,操作规划性能显著提升。
cs.RO / 61 / 2609.33138

Multi-Modal Non-Prehensile Estimation of Physical Parameters via Press-and-Pull Tipping

基于按压-牵拉倾覆(Press-and-Pull Tipping)的多模态非抓取物理参数估计
Hyland, Steven M., Xiao, Jing, Onal, Cagdas D.
Abstract
Recovering physical properties of unknown objects through non-prehensile interaction is challenging because no single manipulation primitive reveals all relevant parameters. Planar pushing couples mass and friction, while conventional tipping cannot recover friction and may fail entirely when low-friction or curved-base objects slide or rotate instead of tipping. We introduce a multi-modal estimation framework that combines a sliding interaction with a press-and-pull tipping primitive to recover object mass, center of mass height, and surface friction. The press-and-pull interaction increases the object-table sliding threshold and stabilizes the pivot, enabling controlled tipping without any prior geometric object model. Wrist force/torque sensing, RGB-D perception, and robot proprioception are fused to estimate the physical parameters from the two complementary interaction modes. Experiments on an ABB IRB120 across four objects with varied geometry, mass, center of mass, and friction achieve low relative error, while successfully operating on curved-base objects that fail under conventional forward tipping. The results demonstrate that complementary non-prehensile interactions can recover a compact set of physical parameters without grasping, a prior object model, or learned interaction dynamics.
Chinese Translation
通过非抓取式交互恢复未知物体的物理属性具有挑战性,因为没有任何单一操作基元能够揭示所有相关参数。平面推动会耦合质量与摩擦,而传统倾覆无法恢复摩擦,并且当低摩擦或弧形底部物体滑动或旋转而非倾覆时可能完全失败。我们提出一种多模态估计框架,将滑动交互与按压-牵拉倾覆基元相结合,以恢复物体质量、质心高度和表面摩擦。按压-牵拉交互提高了物体-桌面滑动阈值并稳定支点,从而无需任何先验几何物体模型即可实现受控倾覆。融合腕部力/力矩传感、RGB-D感知和机器人本体感觉,从两种互补交互模式中估计物理参数。在ABB IRB120上对四个具有不同几何形状、质量、质心和摩擦的物体进行实验,实现了较低的相对误差,并成功作用于在传统前向倾覆下失败的弧形底部物体。结果表明,互补的非抓取式交互能够在无需抓取、先验物体模型或学习交互动力学的情况下恢复一组紧凑的物理参数。
cs.RO / 62 / 2609.33145

Beyond State-as-Action: Exploiting Command-State Discrepancy for Robot Imitation Learning

超越“状态即动作”(State-as-Action):利用指令-状态差异进行机器人模仿学习
Li, Peiyan, Tao, Yueran, Zhang, Enhao, Zhao, Zhixuan, Yue, Chenghao, Wang, Hao, Lv, Lei, Zhao, Wentao, Chen, Jiahao, Liu, Xin, Huang, Kangyao, Luo, Yu, Liu, Huaping
Abstract
Constructing action targets from measured robot motion is an established approach in imitation learning. Under interaction constraints, however, command-state discrepancy may reflect control demands that motion alone does not capture. We investigate when this information matters and how to exploit it. Across three real-robot tasks, task and phase analyses reveal larger supervision gaps under constrained interaction, while selective command retention provides evidence of locally useful command information. Building on these findings, we propose Command-State Discrepancy Weighting (CSDW), which accounts for robot response times and combines subsequent progress, persistent unmet demand, and demand changes into continuous weights for command supervision. The method requires no task-phase annotations or changes to policy architecture or inference. CSDW improves over uniform command supervision on constrained tasks, while methods perform similarly in the less constrained task.
Chinese Translation
根据实测机器人运动构建动作目标是模仿学习中的一种成熟方法。然而,在交互约束下,指令-状态差异可能反映仅凭运动无法捕捉的控制需求。我们研究了这种信息何时重要以及如何利用它。在三个真实机器人任务中,任务和阶段分析表明,在受约束交互下存在更大的监督差距,而选择性保留指令则提供了局部有用指令信息的证据。基于这些发现,我们提出指令-状态差异加权(Command-State Discrepancy Weighting, CSDW),该方法考虑机器人响应时间,并将后续进展、持续未满足需求和需求变化结合为用于指令监督的连续权重。该方法不需要任务阶段标注,也无需改变策略架构或推理。在受约束任务上,CSDW 优于均匀指令监督;而在约束较少的任务中,各方法表现相似。
cs.RO / 63 / 2609.33157

TimelyDAgger: Timing-Aware Expert Querying for VLA Policy Improvement

TimelyDAgger:面向 VLA 策略改进的时机感知专家查询
Zhao, Zhixuan, Li, Peiyan, Zhang, Enhao, Tao, Yueran, Wang, Hao, Yue, Chenghao, Lv, Lei, Zhao, Wentao, Chen, Jiahao, Liu, Xin, Huang, Kangyao, Luo, Yu, Liu, Huaping
Abstract
DAgger improves robot policies by aggregating expert supervision from states visited during policy execution. Robot-gated DAgger automates expert queries, allowing the robot to decide when to request expert takeover. While existing gates emphasize detecting the need for assistance, takeover timing also shapes the content of these demonstrations and their value for policy learning. We propose TimelyDAgger, combining Bridge-PCA monitoring of internal vision-language-action (VLA) features with Feedback-guided Threshold Adaptation based on expert behavior to improve takeover timing. We introduce an evaluation framework linking failure detection, takeover timing, and policy improvement, including Target-Aligned Supervision Ratio (TASR) for assessing supervision quality without retraining. Experiments show that takeover timing affects policy learning, with TimelyDAgger achieving competitive failure detection and higher post-training success in most evaluated settings under matched expert-action budgets.
Chinese Translation
DAgger 通过聚合策略执行期间访问状态下的专家监督来改进机器人策略。机器人门控 DAgger 自动化专家查询,允许机器人决定何时请求专家接管。虽然现有的门控侧重于检测辅助需求,但接管时机也影响着这些示范的内容及其对策略学习的价值。我们提出 TimelyDAgger,它将内部视觉-语言-动作(VLA)特征的 Bridge-PCA 监测与基于专家行为的反馈引导阈值适应相结合,以改进接管时机。我们引入了一个评估框架,将失败检测、接管时机和策略改进联系起来,包括目标对齐监督比率(TASR),用于在无需重新训练的情况下评估监督质量。实验表明,接管时机影响策略学习,在匹配的专家动作预算下,TimelyDAgger 在大多数评估设置中实现了有竞争力的失败检测和更高的训练后成功率。
cs.RO / 64 / 2609.33165

Beyond Tasks: A Vision for Reproducing an Animal-like Behavioral Substrate Using Modern Robot Learning Techniques

超越任务:利用现代机器人学习技术再现类动物行为基质的愿景
Menik, Samiyuru, Jayalath, Hemadri
Abstract
Recent advances in robot learning have produced increasingly capable embodied agents. Yet comparatively less attention has been given to a more basic form of competence that animals exhibit continuously: the ability to remain situated, responsive, and behaviorally coherent as physical, environmental, and social demands change over time. We propose the ethological behavioral substrate as a conceptual lens for studying this form of competence in artificial agents. Rather than treating these behaviors that animals exhibit as a set of isolated skills, we argue that their continual coordination under competing demands constitutes an important and underexplored target for modern robot learning. We further propose robotic animal companions as a useful research setting for studying sustained interaction and adaptation in human-centered environments. Such systems provide an opportunity to investigate how social behavior, memory, and continual learning develop over long periods of interaction. This perspective motivates further investigation of how such persistent behavioral competence may complement higher-level capabilities in embodied agents.
Chinese Translation
近年来,机器人学习领域的进展已经产生了能力日益增强的具身智能体。然而,相对而言,人们较少关注动物持续展现的一种更基本的能力:随着物理、环境和社会需求随时间变化,保持情境化、响应性和行为连贯性的能力。我们提出行为学行为基质作为一个概念透镜,用于研究人工代理中的这种能力形式。我们认为,与其将动物展现的这些行为视为一组孤立的技能,不如将其在竞争性需求下的持续协调视为现代机器人学习的一个重要且尚未充分探索的目标。我们进一步提出机器人动物伴侣作为一个有用的研究环境,用于研究以人为中心的环境中的持续交互和适应。此类系统提供了一个机会,来研究社会行为、记忆和持续学习如何在长时间交互中发展。这一视角激发了对这种持久行为胜任力如何补充具身智能体中更高层次能力的进一步研究。
cs.RO / 65 / 2609.33172

Dynamic Manipulation with World-Action Models via Counterfactual Planning

基于反事实规划的世界-动作模型动态操作
Park, Sunwoo, Lee, Wonbin, Jin, Seonghyun, Kim, Youngmin, Park, Jangho, Ye, Jong Chul
Abstract
World-Action models (WAMs) trained on static demonstrations often fail to manipulate moving targets even when they possess the required manipulation skills. We attribute this failure to target-response collapse: as execution advances, the policy becomes increasingly biased toward the learned continuation of its ongoing behavior and less responsive to target relocation. To bridge the gap between what the model has learned and what it can generate from the current context, we formulate dynamic manipulation as counterfactual planning by decoupling the context used for plan generation from the physical state used for execution. Our framework, Dynamic Predictive Planning (DPP), first uses the WAM's predictive rollout to estimate when an interaction is expected to occur, and combines this timing estimate with observed target motion to predict the target's future interaction position. DPP then constructs a counterfactual observation that places this predicted target position in a familiar robot context, allowing the model to invoke an existing manipulation skill rather than generate a recovery behavior from an unfamiliar robot-target configuration. The resulting plan is connected to the robot's actual state during execution. DPP enables real-time dynamic manipulation on a single consumer GPU without additional training on dynamic data. Experiments in simulation and on a real robot demonstrate consistent improvements across diverse target motions, with simulation performance surpassing all evaluated baselines, including methods additionally trained on dynamic data. Project page: https://methoder00.github.io/DPP/
Chinese Translation
在静态演示上训练的世界-动作模型(WAMs)即使具备所需的操作技能,也往往无法操作移动目标。我们将其失败归因于目标响应崩溃:随着执行的推进,策略越来越偏向于学习到的当前行为的延续,而对目标移动的响应减弱。为了弥合模型所学内容与其从当前上下文所能生成内容之间的差距,我们将动态操作表述为反事实规划,通过将用于计划生成的上下文与用于执行的物理状态解耦。我们的框架,动态预测规划(DPP),首先使用WAM的预测推演来估计预期发生交互的时间,并将该时间估计与观察到的目标运动相结合,以预测目标未来的交互位置。然后,DPP构建一个反事实观测,将预测的目标位置置于熟悉的机器人上下文中,使模型能够调用现有的操作技能,而不是从不熟悉的机器人-目标配置中生成恢复行为。生成的计划在执行期间与机器人的实际状态相连。DPP能够在单个消费级GPU上实现实时动态操作,而无需在动态数据上进行额外训练。仿真和真实机器人上的实验表明,在各种目标运动上均有一致的改进,仿真性能超过了所有评估的基线,包括额外在动态数据上训练的方法。项目页面:https://methoder00.github.io/DPP/
cs.RO / 66 / 2609.33174

Multi-Terrain Mastery: A Comprehensive Controller for Bipedal Locomotion

多地形掌控:一种用于双足运动的综合控制器
Dosunmu-Ogunbi, Oluwami, Shrivastava, Aayushi
Abstract
Advancing bipedal robots to navigate diverse terrains remains a significant challenge in robotics. Traditional locomotion controllers excel on specific surfaces but struggle across varied environments, limiting their practical applications. Given the unpredictable nature of real-world environments, a single controller capable of handling multiple terrains is ideal, eliminating the need for multiple specialized controllers. We propose a multi-terrain controller to enhance the versatility and robustness of bipedal locomotion. Building on previous work with a stance ankle motor for stability on inclined and rough surfaces, this paper extends capabilities to steep wet uneven grassy slopes, and compliant terrains such as sand, gravel, rocks, and constrained terrains like staircases. To address the unique demands of these terrains, we introduce a new impact map that is essential for maintaining performance and robustness against unseen terrains. We also discuss in detail the control structure for real-time deployment on the robot. We validate our controller on the 20 degree-of-freedom Cassie bipedal robot.
Chinese Translation
提升双足机器人穿越多样地形的能力仍然是机器人学中的一项重大挑战。传统运动控制器在特定表面上表现出色,但在多变环境中却表现不佳,限制了其实际应用。鉴于真实世界环境的不可预测性,一个能够处理多种地形的单一控制器是理想选择,从而无需多个专用控制器。我们提出一种多地形控制器,以增强双足运动的通用性和鲁棒性。在先前利用支撑踝关节电机在倾斜和崎岖表面保持稳定的工作基础上,本文将能力扩展到陡峭湿滑不平的草坡,以及沙地、砾石、岩石等可顺应地形和楼梯等受限地形。为应对这些地形的独特需求,我们引入一种新的冲击映射(impact map),它对于在未见地形上保持性能和鲁棒性至关重要。我们还详细讨论了在机器人上实时部署的控制结构。我们在20自由度的Cassie双足机器人上验证了我们的控制器。
cs.RO / 67 / 2609.33177

DeltaWAM: Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action Model

DeltaWAM:以变化为中心的视觉前瞻,通过Delta Tokens实现高效世界动作模型
Jiang, Tianyun, Bao, Wenrui, Xu, Bingxin, Tian, Yu, Shang, Yuzhang
Abstract
World-Action Models (WAMs) offer visual foresight for robotic manipulation, but pixel-space models repeatedly reconstruct entire future scenes, incurring high computational cost and spatio-temporal redundancy. In physical manipulation, consecutive frames often share most of their visual context; the changes between them are what an action policy needs to anticipate. We introduce DeltaWAM, a change-centric WAM that makes a compact delta token the unit of future prediction. Each token is a single vector encoding changes between consecutive dense DINO feature maps. DeltaWAM builds on DeltaWorld, a latent world model pretrained on large-scale videos, to autoregressively predict one delta token per future frame. A flow-matching action expert then conditions on the predicted transitions and current DINO features, which serve as spatial anchors, to generate action chunks. Trained for 256 GPU hours on two H100 GPUs, DeltaWAM has 0.725B parameters and achieves 92.8% average success on LIBERO. It also shows robust generalization under procedural perturbations on LIBERO-Pro. Inference takes 142.1 ms per action chunk with 3.86 GB peak memory.
Chinese Translation
世界动作模型(WAMs)为机器人操作提供了视觉前瞻,但像素空间模型会重复重建整个未来场景,导致高计算成本和时空冗余。在物理操作中,连续帧通常共享大部分视觉上下文;它们之间的变化才是动作策略需要预测的内容。我们提出了DeltaWAM,一种以变化为中心的WAM,它将紧凑的delta token作为未来预测的单元。每个token是一个单一向量,编码连续密集DINO特征图之间的变化。DeltaWAM基于DeltaWorld(一个在大规模视频上预训练的潜在世界模型),自回归地预测每个未来帧的一个delta token。然后,一个流匹配动作专家基于预测的转换和当前DINO特征(作为空间锚点)生成动作块。在两块H100 GPU上训练256 GPU小时,DeltaWAM有0.725B参数,在LIBERO上达到92.8%的平均成功率。它还在LIBERO-Pro上的程序性扰动下展现出鲁棒的泛化能力。推理每个动作块耗时142.1毫秒,峰值内存为3.86 GB。
cs.RO / 68 / 2609.33197

TAO-DA: Towards Autonomous Operation--A Dual-Arm Vision-Language-Action Model for Coordinated Manipulation

TAO-DA:迈向自主操作——一种用于协调操作的双臂视觉-语言-动作模型
Zhao, Yongsheng, Gao, Han, Cheng, Baoping, Tang, Jingyao, Zhou, Dian, Liang, Deng, Ge, Ji, Wen, Xuanzhang, Zhao, Lei, Wang, Ye
Abstract
Vision-Language-Action (VLA) models provide a unified framework for grounding high-level semantic information into low-level robot actions, enabling scalable robotic manipulation across diverse tasks. However, existing VLA models lack explicit mechanisms to disentangle the states and intents of the two arms, leading to unintended cross-arm interference that degrades task execution success. To address this issue, we propose a symmetric Dual-Arm Expert (DAE) architecture built upon a shared Vision-Language Model (VLM) backbone with decoupled, arm-specific expert towers. Expert selection is carried out through a two-stage dual-arm intent routing scheme, in which experts are routed either by explicit language instructions in the first stage or by implicit visual semantics in the second stage. Moreover, we introduce a lightweight task progress prediction module that leverages cross-attention between the pre-chunk temporal features and semantic representations of proprioceptive and visual observations to accurately estimate frame-wise task completion progress. This module facilitates task progress synchronization to support coordinated scheduling for collaborative multi-robot tasks. Experimental results demonstrate the effectiveness of our model in dual-arm intent routing and the disentanglement of cross-arm interference, and further provide preliminary evidence of emergent skill generalization from single- to dual-arm tasks (as well as the reverse), together with cross-arm motion-domain skill transfer.
Chinese Translation
视觉-语言-动作(VLA)模型提供了一个统一框架,用于将高层语义信息落地为底层机器人动作,从而支持跨多样任务的可扩展机器人操作。然而,现有的VLA模型缺乏显式机制来解耦双臂的状态与意图,导致意外的跨臂干扰,从而降低任务执行成功率。为解决这一问题,我们提出了一种对称的双臂专家(DAE)架构,该架构建立在共享的视觉-语言模型(VLM)主干之上,并具有解耦的、臂特定的专家塔。专家选择通过一个两阶段双臂意图路由方案进行,其中在第一阶段通过显式语言指令路由专家,在第二阶段通过隐式视觉语义路由专家。此外,我们引入了一个轻量级任务进度预测模块,它利用分块前时序特征与本体感觉和视觉观测的语义表示之间的交叉注意力,来准确估计逐帧任务完成进度。该模块促进任务进度同步,以支持协作多机器人任务的协调调度。实验结果表明,我们的模型在双臂意图路由和跨臂干扰解耦方面是有效的,并进一步提供了从单臂任务到双臂任务(以及反向)涌现技能泛化的初步证据,以及跨臂运动域技能迁移的初步证据。
cs.RO / 69 / 2609.33235

Learning with Object-centric Representations of Tactile Interactive Perception for Robot Manipulation

面向机器人操作的触觉交互感知的以物体为中心的表示学习
Yang, Xinyi, Si, Zilin, Xu, Zhuowei, Temel, Zeynep, Kroemer, Oliver
Abstract
Implicit object properties that are difficult to directly infer from vision, such as material, container contents, or softness, can be revealed through tactile sensing and exploratory interactions. However, because tactile signals are transient and sparse, extracting informative tactile events and effectively incorporating them into robotic manipulation remains a challenge. In this work, we present an object-centric context-aware manipulation framework that learns task-agnostic object representations through tactile exploration. A token learner autonomously selects representative tactile segments from long-horizon exploration, while contrastive alignment with descriptive text embeddings enables a latent space that captures multiple physical object properties. These learned representations are then used as semantic context to guide object-centric manipulation policies and adapt strategies based on object properties. Experiments show that the learned representations achieve 93% and 84% property estimation accuracy on seen and unseen objects. Evaluated on three tasks involving visually ambiguous objects, i.e. multi-object rearrangement, pouring, and box opening, the proposed framework improves both target selection and property-dependent manipulation adaptation, raising task success, aggregated over all evaluation trials, from 41% to 92% on seen objects and from 19% to 67% on unseen objects over baseline policies without object-context conditioning. Videos and additional results are available at https://xinyiyxyx.github.io/tactile-object-centric/.
Chinese Translation
难以直接从视觉推断的隐式物体属性,如材质、容器内容物或柔软度,可以通过触觉传感和探索性交互来揭示。然而,由于触觉信号是瞬态且稀疏的,提取有信息量的触觉事件并将其有效地融入机器人操作仍然是一个挑战。在这项工作中,我们提出了一个以物体为中心的上下文感知操作框架,该框架通过触觉探索学习任务无关的物体表示。一个 token 学习器从长时程探索中自主选择具有代表性的触觉片段,同时与描述性文本嵌入的对比对齐使得潜在空间能够捕获多个物理物体属性。然后,这些学习到的表示被用作语义上下文,以指导以物体为中心的操作策略,并根据物体属性调整策略。实验表明,学习到的表示在已见和未见物体上分别达到 93% 和 84% 的属性估计准确率。在涉及视觉模糊物体的三个任务(即多物体重排、倾倒和开箱)上进行评估,所提出的框架改善了目标选择和依赖于属性的操作适应,与没有物体上下文条件的基线策略相比,在所有评估试验中汇总的任务成功率在已见物体上从 41% 提高到 92%,在未见物体上从 19% 提高到 67%。视频和更多结果可在 https://xinyiyxyx.github.io/tactile-object-centric/ 获取。
cs.RO / 70 / 2609.33237

SurgFlow: 3D Object-Centric Contact Flow for Surgical Robot Manipulation

SurgFlow:面向手术机器人操作的3D以物体为中心的接触流
Chen, Changwei, Liang, Xiao, Yang, Yinuo, Shen, Nicole, Zhang, Peihan, Wickenhiser, Sara, Liang, Zekai, Atar, Soofiyan, Yip, Michael
Abstract
Paired video-action demonstrations enable autonomous surgical behavior, but such data is scarce: robots perform roughly 1% of surgeries, while video-only data is abundant. Learning 3D object flow offers an embodiment-agnostic way to utilize video data, but flow alone specifies how an object should move, not where and when the tool should engage it, a distinction that is critical in surgery. We introduce SurgFlow, a framework that learns 3D Object-Centric Contact Flow from stereo surgical video without action labels. For each object point, it predicts a future 3D trajectory and contact scores. We extract targets via 3D tracking and tool-object proximity, train a flow matching generator to predict them, and use predicted contact to trigger grasp and release while optimizing end effector motion from flow. On the da Vinci Research Kit (dVRK), SurgFlow succeeds in 37 of 39 stage evaluations across tissue retraction, bimanual reveal, needle pickup, and handover, outperforming baselines trained on equal data with or without action labels. Zero-shot transfer to a humanoid-based laparoscopic robot achieves 85% and 70% average success under similar and novel camera viewpoints, respectively.
Chinese Translation
成对的视频-动作演示能够实现自主手术行为,但此类数据稀缺:机器人仅执行约1%的手术,而仅有视频的数据却非常丰富。学习3D物体流提供了一种与具身无关的方式来利用视频数据,但仅靠流指定了物体应如何移动,而未指定工具应在何处以及何时与其接触,这一区别在手术中至关重要。我们提出了SurgFlow,一个无需动作标签即可从立体手术视频中学习3D以物体为中心的接触流的框架。对于每个物体点,它预测未来的3D轨迹和接触分数。我们通过3D跟踪和工具-物体接近度提取目标,训练一个流匹配生成器来预测它们,并利用预测的接触来触发抓取和释放,同时根据流优化末端执行器运动。在达芬奇研究套件(dVRK)上,SurgFlow在39个阶段评估中的37个中成功,涵盖组织牵拉、双手显露、拾针和传递,性能优于在同等数据上训练且无论有无动作标签的基线方法。零样本迁移到基于人形的腹腔镜机器人,在相似和新颖相机视角下分别达到85%和70%的平均成功率。
cs.RO / 71 / 2609.33256

ActionGround: Training-Free Runtime Refinement of Frozen VLA Policies

ActionGround:冻结VLA策略的免训练运行时精炼
Chandra, Namai, Thareja, Madhur, Damodaran, Shriram, Wang, Addison Lin
Abstract
Vision-Language-Action (VLA) models map visual observations and language instructions directly to robot actions, but they do not explicitly represent the phase structure of manipulation tasks or the rigid-body dynamics governing execution. We present ActionGround, a neuro-symbolic, training-free runtime layer that wraps a frozen VLA policy without retraining, fine-tuning, or weight access, adding less than 1 ms of overhead per control step. A symbolic phase-aware finite-state machine identifies the manipulation phase (approach, grasp, transport, or place) and applies a phase-specific rule-based correction. In parallel, an always-on, inertia-weighted Euler-Lagrange term incorporates the robot's equations of motion into each control step, while its dynamics residual is logged as a consistency diagnostic rather than used as a gate. We evaluate ActionGround across OpenVLA, OpenVLA-OFT, Force-VLA, and Generalist-VLA on ten LIBERO-Spatial pick-and-place tasks using a 7-DoF Franka Panda. With fixed parameters across tasks and backbones, ActionGround improves success rate by up to 6 percentage points and stability by up to 19.3 percentage points, while improving trajectory efficiency by up to 15%. In a separate Robosuite noise sweep, ActionGround provides approximately a 10x improvement in trajectory-jerk robustness under injected action noise. In a matched-seed Robosuite simulation companion to a real Agilex Piper trial, simulated baseline success increases from 35% to 95%. The physical-hardware experiment is presented as a qualitative deployment demonstration; quantitative per-trial success on the real arm is left for future work. Our evaluation is limited to rigid-object pick-and-place manipulation.
Chinese Translation
视觉-语言-动作(VLA)模型将视觉观察和语言指令直接映射为机器人动作,但它们并未显式表示操作任务的阶段结构或支配执行的刚体动力学。我们提出ActionGround,一个神经符号的、免训练的运行时层,它包装一个冻结的VLA策略,无需重新训练、微调或访问权重,每个控制步骤增加不到1毫秒的开销。一个符号化的阶段感知有限状态机识别操作阶段(接近、抓取、运输或放置),并应用特定于阶段的基于规则的校正。同时,一个始终开启的、惯性加权的欧拉-拉格朗日项将机器人的运动方程纳入每个控制步骤,而其动力学残差被记录为一致性诊断,而非用作门控。我们在十个LIBERO-Spatial拾放任务上,使用7自由度Franka Panda机器人,对ActionGround在OpenVLA、OpenVLA-OFT、Force-VLA和Generalist-VLA上进行了评估。在跨任务和主干网络固定参数的情况下,ActionGround将成功率提高了最多6个百分点,稳定性提高了最多19.3个百分点,同时将轨迹效率提高了最多15%。在另一项Robosuite噪声扫描中,ActionGround在注入动作噪声下将轨迹加加速度鲁棒性提高了约10倍。在与真实Agilex Piper试验匹配种子的Robosuite模拟配套实验中,模拟基线成功率从35%提高到95%。物理硬件实验作为定性部署演示呈现;真实机械臂上每次试验的定量成功率留待未来工作。我们的评估仅限于刚性物体拾放操作。
cs.RO / 72 / 2609.33258

PORTER: Edge-Cloud Residency for Persistent 3D Scene Graph Memory

PORTER:面向持久性3D场景图内存的边云驻留
Chang, Yue, Tian, Yifan, Peng, Jiajing, Huang, Dazhi, Chen, Rufeng, Zhang, Zhaofan, Chen, Li, Xie, Sihong
Abstract
Recent task-driven and just-in-time 3D Scene Graph (3DSG) methods reduce per-task representations by constructing or activating only task-relevant information. Yet sparse per-task working sets do not bound onboard memory usage over a robot's lifetime: as tasks change, payloads accumulated for earlier tasks may become irrelevant to the current task but can be useful again in future tasks. Over repeated task switches and expanding environments, retaining such reusable payloads causes local memory to grow, whereas discarding them entirely can lead to costly repeated construction of the same payloads later. We introduce PORTER, which decouples persistence from residency: lightweight anchors remain in the limited memory of the edge robot while heavy object payloads migrate between the edge and the cloud. Relevance alone is insufficient for deciding residency because multiple relevant payloads may provide redundant information. We therefore decompose each task into functional requirements and introduce Irreplaceable Support Erasure (ISE), which measures the loss in requirement coverage caused by offloading. ISE discounts replaceable support and penalizes losses more strongly when the remaining coverage of a requirement is weak. PORTER constructs a budget-aware local working set by repeatedly offloading the payload with the smallest marginal ISE per byte. Experiments on JITOMA-Bench evaluate PORTER across four 3DSG builders. Under progressive compression, pooled relative mR@3 remains at 100% through 91% payload-byte offloading.
Chinese Translation
近期的任务驱动和即时(just-in-time)3D场景图(3DSG)方法通过仅构建或激活任务相关信息来减少逐任务表示。然而,稀疏的逐任务工作集并不能限制机器人在其生命周期内的机载内存使用:随着任务变化,为早期任务累积的载荷可能变得与当前任务无关,但可在未来任务中再次有用。在反复的任务切换和环境扩展中,保留此类可重用载荷会导致本地内存增长,而完全丢弃它们则可能在之后导致相同载荷的高成本重复构建。我们提出PORTER,它将持久性与驻留解耦:轻量级锚点保留在边缘机器人的有限内存中,而重对象载荷在边缘与云之间迁移。仅靠相关性不足以决定驻留,因为多个相关载荷可能提供冗余信息。因此,我们将每个任务分解为功能需求,并引入不可替代支持擦除(Irreplaceable Support Erasure, ISE),它衡量由卸载引起的需求覆盖损失。ISE会对可替代支持进行折减,并在某个需求的剩余覆盖较弱时更强烈地惩罚损失。PORTER通过反复卸载每字节边际ISE最小的载荷,构建预算感知的本地工作集。在JITOMA-Bench上的实验在四种3DSG构建器上评估PORTER。在渐进压缩下,随着载荷字节卸载率达到91%,汇总相对mR@3仍保持在100%。
cs.RO / 73 / 2609.33269

Q-WAM: 4-Bit Quantization of World Action Models with Action-Subspace Protection

Q-WAM: 世界动作模型的4位量化与动作子空间保护
Akbari, Arash, Akbari, Arman, Luo, Jingwu, Lei, Yuhao, Gao, Yi, Chen, Weiwei, Zhang, Xuan, Fang, Zhenman, Yuan, Geng, Wang, Yanzhi
Abstract
World Action Models (WAMs) jointly generate video and robot actions through iterative diffusion and perform strongly in robotic manipulation. However, their prohibitive compute and memory costs pose substantial deployment challenges. Post-training quantization (PTQ) can reduce these costs, but existing PTQ methods such as smoothing and rotation are insufficient to maintain the precision of action generation. To overcome this limitation, we propose Q-WAM, a new 4-bit weight-activation quantization for WAMs that preserves the actions the model generates. Specifically, we introduce the \textit{Action Observability Gramian (AOG)}, which measures how much rounding errors in each weighted combination of a layer's input channels change the final action through all denoising steps. We also develop Action-Subspace Protection (ASP), which keeps the few most action-sensitive channel combinations in a tiny 16-bit low-rank branch and quantizes the complementary weights and activations to 4 bits, both as dense matrix multiplications that run efficiently on GPUs. Finally, to preserve action quality with minimal overhead, we identify the experts that matter most for the generated action by aggregating the AOG-derived action mass across the layers of each expert and apply ASP only to those experts. We evaluate Q-WAM on three WAMs, both in simulation and in real-world deployment. On the RoboTwin 2.0 benchmark, it reaches 89.6--93.0\% average success rate, within 1.1 percentage points of the 16-bit models, while reducing the memory of the quantized blocks by 3.1--3.4$\times$. Our method outperforms the strongest baseline, SVDQuant, by 2.5--8.7 percentage points. On a Unitree G1 humanoid and a bimanual UR3 robot, it improves success over SVDQuant by 12.8-17.6 percentage points.
Chinese Translation
世界动作模型(WAMs)通过迭代扩散联合生成视频和机器人动作,在机器人操作中表现强劲。然而,其高昂的计算和内存成本带来了巨大的部署挑战。训练后量化(PTQ)可以降低这些成本,但现有的PTQ方法(如平滑和旋转)不足以保持动作生成的精度。为了克服这一局限性,我们提出了Q-WAM,一种新的针对WAMs的4位权重-激活量化方法,能够保留模型生成的动作。具体来说,我们引入了动作可观测性格拉姆矩阵(AOG),它衡量层输入通道的每个加权组合中的舍入误差通过所有去噪步骤对最终动作的影响程度。我们还开发了动作子空间保护(ASP),它将少数对动作最敏感的通道组合保留在一个微小的16位低秩分支中,并将互补的权重和激活量化为4位,两者都作为密集矩阵乘法在GPU上高效运行。最后,为了以最小的开销保持动作质量,我们通过聚合每个专家各层中AOG导出的动作量来识别对生成动作最重要的专家,并仅对这些专家应用ASP。我们在三个WAMs上评估了Q-WAM,包括仿真和真实世界部署。在RoboTwin 2.0基准测试上,它达到了89.6--93.0%的平均成功率,与16位模型相差不到1.1个百分点,同时将量化块的内存减少了3.1--3.4倍。我们的方法比最强的基线SVDQuant高出2.5--8.7个百分点。在Unitree G1人形机器人和双臂UR3机器人上,它比SVDQuant的成功率提高了12.8-17.6个百分点。
cs.RO / 74 / 2609.33299

AquaWAM: A Dynamics-aware World Action Model for Underwater Embodied Agents

AquaWAM:面向水下具身智能体的动力学感知世界动作模型
Zhu, Cunhao, Wang, Yifeng, Xu, Dongliang, Hou, Yunzhong, Yao, Yue, Liu, Chi Harold
Abstract
World Action Models (WAMs) are becoming increasingly important and useful for embodied intelligence, as they enable robots to anticipate the consequences of candidate actions before interacting with the physical environment. However, underwater robots are usually subject to passive dynamics, such as inertia, buoyancy, hydrodynamic drag, and persistent drift, which can continue to affect the vehicle even after an action is completed. Existing WAMs, which primarily predict action-conditioned visual observations, are not explicitly designed to capture such passive motion dynamics. In this paper, we present AquaWAM, the first World Action Model designed for underwater embodied agents. Instead of predicting future images, AquaWAM models both action-conditioned and passive physical dynamics, including the thruster dead band, the inertial glide that outlasts each command, and ambient currents. Specifically, it senses through the DVL, IMU, pressure sensor and joint encoders, while cameras supply only semantics for understanding goals and target pose. By modeling compact navigation states rather than high-dimensional visual observations, AquaWAM substantially reduces the model size and computational cost compared with conventional WAMs. Experimentally, AquaWAM achieves a 72.6% task success rate across 20 underwater tasks on the USIM benchmark, outperforming existing methods while making action decisions 2.7x faster than U0 on an NVIDIA Jetson AGX Orin. Our model also remains effective when some onboard sensor measurements are unavailable. For example, without DVL velocity measurements, our method still achieves a 61.6% success rate, compared with 39.4% for U0.
Chinese Translation
世界动作模型(WAMs)对于具身智能正变得越来越重要和有用,因为它们使机器人能够在与物理环境交互之前预测候选动作的后果。然而,水下机器人通常受到被动动力学的影响,如惯性、浮力、水动力阻力和持续漂移,这些影响甚至在动作完成后仍会持续作用于载具。现有的WAM主要预测以动作为条件的视觉观测,并未明确设计用于捕捉这种被动运动动力学。在本文中,我们提出了AquaWAM,这是第一个为水下具身智能体设计的世界动作模型。AquaWAM不预测未来图像,而是对动作条件和被动物理动力学进行建模,包括推进器死区、比每条指令持续更久的惯性滑行以及环境水流。具体来说,它通过DVL、IMU、压力传感器和关节编码器进行感知,而摄像头仅提供用于理解目标和目标位姿的语义信息。通过建模紧凑的导航状态而非高维视觉观测,AquaWAM与传统WAM相比大幅减小了模型规模和计算成本。在实验中,AquaWAM在USIM基准的20个水下任务上达到了72.6%的任务成功率,优于现有方法,同时在NVIDIA Jetson AGX Orin上做出动作决策的速度比U0快2.7倍。当一些机载传感器测量不可用时,我们的模型仍然有效。例如,在没有DVL速度测量的情况下,我们的方法仍达到61.6%的成功率,而U0为39.4%。
cs.RO / 75 / 2609.33310

CompliantWBC: Whole-Body Compliance for Heavy Humanoids via Force Latent Estimation and Residual Impedance Targets

CompliantWBC:通过力潜变量估计和残差阻抗目标实现重型人形机器人的全身柔顺控制
Do, Tan-Dzung, Trinh, Cuc T., Phuong, Tuan Dat, Le, Chien, Ly, Thanh, Ngo, Vien Anh, Le, An Thai
Abstract
Whole-body compliant control is essential for deploying heavy humanoids under high payload in human-centric environments. Most prior force-aware learning-based pipelines focus on end-effector resistance, per-link upper-body springs, or end-effector stiffness modulation, leaving arbitrary-site perturbations on heavy platforms with lower-body engagement largely unaddressed. We close this gap with CompliantWBC comprising: (1) A base policy trained with RL to maximize compliance-fidelity reward, guided by a multi-site whole-body impedance reference controller, extending classical Cartesian impedance to any controlled link; (2) A bounded residual policy that edits the per-link impedance equilibrium over a frozen base, correcting the coarse but structured wrench estimate supplied by a force encoder co-trained behind a gradient barrier; (3) A Phong-weighted force-origin sampler with an axis-decoupled pelvis anchor induces lower-body-inclusive compliance curriculum training via two interpretable parameters. We evaluate CompliantWBC in simulation against both compliant and stiff baselines, achieving best compliant fidelity of 2.58cm deviation from analytical solutions, and demonstrate it on a real heavy humanoid across static/dynamic force reaction, board wiping, squat under payload, and cooperative payload transport. Project website: https://dotandung.github.io/compliantwbc/
Chinese Translation
全身柔顺控制对于在以人为本的环境中部署高负载下的重型人形机器人至关重要。大多数先前的力感知学习型流程关注末端执行器阻力、每连杆上肢弹簧或末端执行器刚度调制,而重型平台上具有下肢参与的任意位置扰动在很大程度上未得到解决。我们通过CompliantWBC填补这一空白,其包括:(1) 一个使用强化学习(RL)训练的基础策略,以最大化柔顺保真度奖励,由多部位全身阻抗参考控制器指导,将经典笛卡尔阻抗扩展到任何受控连杆;(2) 一个有界残差策略,在冻结的基础上编辑每连杆阻抗平衡,校正由梯度屏障后协同训练的力编码器提供的粗略但结构化的力旋量估计;(3) 一个带有轴解耦骨盆锚点的Phong加权力原点采样器,通过两个可解释参数诱导包含下肢的柔顺性课程训练。我们在仿真中针对柔顺和刚性基线评估CompliantWBC,实现了最佳柔顺保真度,与解析解的偏差为2.58厘米,并在真实重型人形机器人上展示了其静态/动态力反应、擦板、负载下深蹲以及协作负载运输。项目网站:https://dotandung.github.io/compliantwbc/
cs.RO / 76 / 2609.33311

SocialHumanoid: Towards Expressive Humanoid Behavior via One-Step Co-Speech Motion Generation

SocialHumanoid:通过一步式协同语音动作生成实现富有表现力的人形机器人行为
Yang, Chengqun, Zhu, Tengjie, Xu, Liang, Liu, Fulong, Ren, Guanzhu, Xing, Yitong, Lu, Xuefeng, Shi, Fei, Fan, Siyuan, Dong, Weijie, Mu, Yao, Yang, Xiaokang, Yan, Yichao
Abstract
Humanoid robots are increasingly expected to serve as embodied social agents that communicate naturally with humans through face-to-face interaction. During such communication, humanoid robots require body behaviors that are synchronized with speech, affectively expressive, and suitable for real-time execution. However, existing co-speech methods are primarily developed for digital humans and lack joint support for affective control and low-latency continuous generation on physical embodiments. To bridge this gap, we present SocialHumanoid, a system for expressive humanoid behavior via one-step co-speech motion generation. Given response speech and a specified affective condition, SocialHumanoid generates each full-body motion window in a single forward pass and connects successive windows through motion-history conditioning. The generated human motion is further converted online into embodiment-compatible robot references and tracked by a whole-body controller for physical execution. To provide explicit supervision for affective body expression, we further introduce AffectMoCap, a 4-hour dataset captured from two professional actors, containing synchronized speech, body motion, fine-grained hand motion, and emotion annotations. On BEAT2, SocialHumanoid achieves the best FGD among the compared generation methods, competitive speech-motion synchrony, and approximately $6\times$ faster inference than GestureLSM under the same protocol. Perceptual evaluations further show that training with AffectMoCap improves affect recognition from generated body motion, while real-robot experiments demonstrate continuous affect-conditioned behavior and stable long-horizon execution. Our project page is https://rex0191.github.io/SocialHumanoid/.
Chinese Translation
人形机器人越来越多地被期望作为具身社交智能体,通过面对面交互与人类自然交流。在这种交流中,人形机器人需要与语音同步、富有情感表现力且适合实时执行的身体行为。然而,现有的协同语音方法主要面向数字人开发,缺乏对情感控制和物理实体上低延迟连续生成的联合支持。为弥补这一差距,我们提出了SocialHumanoid,一个通过一步式协同语音动作生成实现富有表现力的人形机器人行为的系统。给定回复语音和指定的情感条件,SocialHumanoid在单次前向传播中生成每个全身动作窗口,并通过动作历史条件连接连续窗口。生成的人体动作进一步在线转换为与实体兼容的机器人参考,并由全身控制器跟踪以进行物理执行。为了给情感身体表达提供显式监督,我们进一步引入了AffectMoCap,一个从两名专业演员采集的4小时数据集,包含同步的语音、身体动作、细粒度手部动作和情感标注。在BEAT2上,SocialHumanoid在比较的生成方法中取得了最佳的FGD,具有竞争力的语音-动作同步性,并且在相同协议下比GestureLSM的推理速度快约6倍。感知评估进一步表明,使用AffectMoCap训练可以改善从生成的身体动作中进行情感识别,而真实机器人实验展示了连续的情感条件行为和稳定的长时程执行。我们的项目页面是:https://rex0191.github.io/SocialHumanoid/。
cs.RO / 77 / 2609.33354

Traceable Human-to-Humanoid Sign Language Benchmarking

可追溯的人到人形机器人手语基准测试
Liu, Ao, Tang, Shengeng, Cheng, Lechao, Hao, Yanbin, Bao, Bingkun, Hong, Richang
Abstract
Sign data collection is costly, and teleoperation scales poorly, motivating reuse of large video corpora. Humanoid signing requires converting video-derived human motion into robot trajectories while preserving linguistic motion cues. Errors from fitting, human-motion repair, retargeting, robot geometry repair, and control are hard to separate from the final trajectory alone. We introduce HumanoidCSL-20K, a dataset and benchmark of 20,648 sentence-level Chinese Sign Language sequences, each with four aligned versions: the source, the repaired human motion, the direct robot reference, and the geometry-repaired robot reference. Observation-supported local human-motion repair, full-robot geometry repair, and cross-representation provenance make each transformation traceable. Paired evaluations measure human-motion continuity and content preservation, robot-reference feasibility, and physical execution. A sign-specific kinematic-reference protocol scores handshape, location, palm orientation, and inter-hand relation over the full planned motion. Full-corpus results show fewer abnormal arm / hand steps and less inter-hand and hand-body penetration after repair. Control experiments separate reference learnability from curriculum effects, while component scores expose remaining execution errors.
Chinese Translation
手语数据采集成本高昂,远程操作难以规模化,促使人们复用大规模视频语料库。人形机器人手语需要将视频衍生的真人动作转换为机器人轨迹,同时保留语言学动作线索。仅从最终轨迹中,很难分离出来自拟合、人体运动修复、重定向、机器人几何修复和控制的误差。我们提出了 HumanoidCSL-20K,一个包含 20,648 个句子级中国手语序列的数据集和基准,每个序列有四个对齐版本:源数据、修复后的人体运动、直接机器人参考以及几何修复后的机器人参考。观察支持的局部人体运动修复、全机器人几何修复以及跨表示溯源使得每个转换步骤可追溯。配对评估衡量人体运动的连续性和内容保留、机器人参考的可行性以及物理执行情况。一个针对手语特定的运动学参考协议对完整规划动作中的手形、位置、手掌朝向和双手关系进行评分。全语料库结果显示,修复后异常手臂/手部步态减少,双手之间和手与身体之间的穿透减少。控制实验将参考可学习性与课程效应分离,而组件分数则揭示了剩余的執行错误。
cs.RO / 78 / 2609.33378

Recursive Harness Distillation across Agents for Robot Manipulation

面向机器人操作的跨智能体递归 Harness 蒸馏
Kim, Seungyeon, Lee, Junhoo, Kim, Minkyu, Kim, Baekseung, Kwak, Nojun
Abstract
A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong agents can discover effective interventions through interaction with these policies. We propose Recursive Harness Distillation to accumulate this experience as reusable guidance across agents. A strong agent distills its experience into a playbook for a light agent, then recursively refines the playbook using the light agent's execution feedback. The resulting playbook enables agents to reuse accumulated intervention knowledge in new task instances without updating model parameters. In real-world manipulation, the harness improves success from 37.3% to 64.0%. On SimplerEnv Bridge, the light agent with the playbook achieves 66.7% success, compared with 41.7% for the GR00T-only baseline, and outperforms the strong agent without a playbook. The same playbook also benefits the strong agent, which reaches 79.2% success. These results demonstrate the feasibility of harness distillation for robotics: intervention experience can be accumulated, refined through execution, and reused across agents to improve manipulation.
Chinese Translation
机器人学的一个核心目标是实现跨变化任务和环境的操作。视觉-语言-动作(VLA)模型提供了广泛的操作能力,但在执行需要诊断故障和调整行为时可能遇到困难。强智能体可以通过与这些策略交互来发现有效的干预措施。我们提出递归 Harness 蒸馏,将这种经验积累为跨智能体的可重用指导。一个强智能体将其经验提炼成一个供轻量智能体使用的策略手册,然后利用轻量智能体的执行反馈递归地优化该策略手册。生成的策略手册使智能体能够在新任务实例中重用积累的干预知识,而无需更新模型参数。在真实世界操作中,该 Harness 将成功率从 37.3% 提升到 64.0%。在 SimplerEnv Bridge 上,使用该策略手册的轻量智能体达到了 66.7% 的成功率,而仅使用 GR00T 的基线为 41.7%,并且优于没有策略手册的强智能体。同一策略手册也使强智能体受益,其成功率达到 79.2%。这些结果证明了 Harness 蒸馏在机器人领域的可行性:干预经验可以被积累、通过执行进行优化,并跨智能体重用以提升操作性能。
cs.RO / 79 / 2609.33423

Large Language Models for Model-Based Robot Design

用于基于模型的机器人设计的大型语言模型
Wilhelm, Andrew, Zhao, Angelina, Napp, Nils
Abstract
Large Language Models (LLMs) can contribute useful engineering knowledge to robot design, but directly generated designs may rely on implicit assumptions and provide no guarantees of feasibility or optimality. These assumptions are critical because different reasonable modeling choices can materially change which designs are predicted to be feasible or optimal. We therefore present a framework that uses LLMs to construct explicit engineering models containing physical relationships, compatibility constraints, and objectives, allowing these modeling choices to be inspected and revised before formal optimization. The model can then be updated with additional engineering, manufacturer, or system-specific information before formal multi-objective optimization provides feasibility and Pareto-optimality guarantees with respect to the finalized model and specified design space. We evaluate the framework on quadcopter and line-following robot component-selection problems. Across 30 direct LLM design trials, none could be verified as feasible under the corresponding finalized model. Comparisons with an independently developed expert model and successive stages of model refinement further showed that changes in modeling assumptions substantially altered the predicted feasible and Pareto-optimal design sets. Together, these results show that using LLMs to construct explicit engineering models makes the underlying design choices available for inspection and revision before those assumptions determine the optimized designs. Explicit modeling therefore provides an interface for combining LLM-generated engineering knowledge, system-specific information, and formal design optimization.
Chinese Translation
大型语言模型(LLMs)可以为机器人设计提供有用的工程知识,但直接生成的设计可能依赖于隐式假设,且无法保证可行性或最优性。这些假设至关重要,因为不同的合理建模选择会实质性改变哪些设计被预测为可行或最优。因此,我们提出了一个框架,利用LLMs构建显式的工程模型,其中包含物理关系、兼容性约束和目标,使得这些建模选择在正式优化之前可以被检查和修订。然后,该模型可以在正式多目标优化之前,用额外的工程、制造商或系统特定信息进行更新,从而针对最终确定的模型和指定的设计空间提供可行性和帕累托最优性保证。我们在四旋翼飞行器和循迹机器人的组件选择问题上评估了该框架。在30次直接LLM设计试验中,没有一个能在相应的最终模型下被验证为可行。与独立开发的专家模型以及模型细化的连续阶段进行比较,进一步表明建模假设的变化会显著改变预测的可行和帕累托最优设计集。总之,这些结果表明,使用LLMs构建显式工程模型,使得底层设计选择在这些假设决定优化设计之前可供检查和修订。因此,显式建模提供了一个接口,用于结合LLM生成的工程知识、系统特定信息和正式设计优化。
cs.RO / 80 / 2609.33464

VIDEAS: Distilling Explicit Action Semantics from Demonstration Videos for World Models via Prior-Guided Simulation

VIDEAS:通过先验引导模拟从演示视频中为世界模型蒸馏显式动作语义
Wang, Jianan, Zhai, Haoquan, Zhang, Siyang, Li, Bin, Chen, Juan, Qi, Jingtao, Zhang, Zhuo, Wang, Enze, Jin, Haoxiang, Qian, Chen
Abstract
World models learn internal representations of environment dynamics to predict future states, enabling agents to optimize action plans without physical interactions. However, developing world models that genuinely internalize underlying causal physical laws to explicitly reason about action preconditions and subsequent state transitions remains an open challenge. In this paper, we propose VIDEAS, a data distillation framework that transforms continuous physical dynamics from operational videos into explicit action semantics for foundation models. Specifically, it deconstructs visual demonstrations into discrete action trajectories and utilizes advanced vision-language models (VLMs) to extract structured knowledge encapsulating action preconditions and effects. To ensure physical consistency, we introduce a prior-guided trajectory simulation mechanism grounded within a text-based environment to rigorously validate the extracted knowledge. Notably, we incorporate negative trajectories to enrich knowledge completeness and enhance data diversity to mitigate cognitive bias. Furthermore, we present VIDEAS-WM, an 8B/9B-parameter suite of language-based world models trained on 34K high-quality samples derived from AgiBot-World dataset. Extensive experiments demonstrate that VIDEAS-WM establishes state-of-the-art performance in high-level embodied action semantic reasoning, exhibiting profound physical understanding and robust generalization across unseen scenarios.
Chinese Translation
世界模型学习环境动态的内部表示以预测未来状态,使智能体能够在无需物理交互的情况下优化动作规划。然而,开发能够真正内化潜在因果物理规律以显式推理动作前提条件和后续状态转换的世界模型仍然是一个开放挑战。在本文中,我们提出VIDEAS,一个数据蒸馏框架,将操作视频中的连续物理动态转化为基础模型的显式动作语义。具体而言,它将视觉演示解构为离散动作轨迹,并利用先进的视觉语言模型(VLMs)提取封装动作前提条件和效果的结构化知识。为确保物理一致性,我们引入一种基于文本环境的先验引导轨迹模拟机制,以严格验证提取的知识。值得注意的是,我们纳入负轨迹以丰富知识完整性,并增强数据多样性以减轻认知偏差。此外,我们提出VIDEAS-WM,一个8B/9B参数的语言基世界模型套件,在来自AgiBot-World数据集的34K高质量样本上训练。大量实验表明,VIDEAS-WM在高层次具身动作语义推理方面达到最先进性能,展现出深刻的物理理解和在未见场景中的稳健泛化能力。
cs.RO / 81 / 2609.33484

AMBIT: Anticipatory Multimodal Body Recruitment for Bimanual Tracking on a Humanoid

AMBIT:面向人形机器人双臂跟踪的预判性多模态身体动员
Li, Hanlong, Tan, Sihan, Ashizawa, Takeshi, Yen, Benjamin, Nakadai, Kazuhiro
Abstract
A humanoid with 5-DoF arms cannot track generic bimanual end-effector trajectories with its arms alone; pelvis and waist motion must be recruited, but which motion, and when, is not uniquely determined. On a Unitree R1 in fixed double support, the set of dynamically valid recruitment strategies (pelvis pose and waist trajectories) for a task is a diverse continuous manifold, and a deterministic regressor trained on it mode-averages into strategies valid only 35% of the time, against 52% for a conditional variational autoencoder (CVAE) and 82% for the best of 16 CVAE samples. We introduce AMBIT: the CVAE proposes strategies from a preview of the commanded trajectory, a non-learned selector filters, ranks and verifies them, and a receding-horizon loop commits to one with hysteresis. The committed strategy is the reference of the same whole-body differential-IK QP a reactive tracker runs, which keeps authority over residual error. On 160 held-out episodes that admit a valid strategy, in full MuJoCo dynamics under a torque controller, AMBIT reaches 85% success at a 3 cm/15 deg tolerance against 74% for the tracker (disjoint confidence intervals) and recruits the body before the arms saturate in 48% of episodes against 35%. Because diversity is preserved, constraints unknown at training time are enforced by selection alone: under five zero-shot shifts AMBIT beats the warm-started tracker on every shift and matches a test-time re-optimisation baseline 17x more expensive. On a Unitree G1, with hyperparameters unchanged, the protocol reproduces the structure of the valid set and widens the gap over the tracker to 0.85 against 0.53. Five selected strategies execute on the externally supported physical R1, distinct in pelvis excursion and tracking the planned end-effector motion to a median of 11 mm by encoder forward kinematics, which establishes kinematic realisability, not balance.
Chinese Translation
具有5自由度手臂的人形机器人无法仅靠手臂跟踪一般的双臂末端执行器轨迹;必须动员骨盆和腰部运动,但具体是哪种运动以及何时运动并非唯一确定。在固定双足支撑的 Unitree R1 上,任务对应的动力学有效动员策略集合(骨盆位姿和腰部轨迹)是一个多样的连续流形,而在其上训练的确定性回归器会因模式平均而生成仅在35%时间内有效的策略,相比之下,条件变分自编码器(CVAE)为52%,16个CVAE样本中的最佳者为82%。我们提出 AMBIT:CVAE 根据指令轨迹的预览提出策略,一个非学习的选择器对它们进行过滤、排序和验证,然后一个滚动时域循环以滞回方式确定其中一个。所确定的策略是反应式跟踪器所运行的同一全身微分逆运动学二次规划(QP)的参考,该跟踪器保留对残差的控制权。在160个允许有效策略的留出回合中,在力矩控制器下的完整 MuJoCo 动力学中,AMBIT 在3厘米/15度容差下达到85%成功率,而跟踪器为74%(置信区间不重叠),并且在48%的回合中在手臂饱和前动员身体,而跟踪器为35%。由于多样性得以保留,训练时未知的约束仅通过选择来强制执行:在五个零样本偏移下,AMBIT 在每个偏移上都优于热启动跟踪器,并匹配了一个昂贵17倍的测试时重新优化基线。在 Unitree G1 上,超参数不变,该协议复现了有效集的结构,并将与跟踪器的差距扩大到0.85对0.53。五个选定的策略在外部支撑的物理 R1 上执行,它们在骨盆偏移方面各不相同,并通过编码器正运动学跟踪规划末端执行器运动的中位误差为11毫米,这证明了运动学可实现性,而非平衡性。
cs.RO / 82 / 2609.33522

AI-Driven Collaborative Assembly Line Inspection: System Integration and Deployment Challenges

AI驱动的协作装配线检测:系统集成与部署挑战
Ünal, Asya, Okasha, Amr, Çırakman, Ege, Ünal, Perin
Abstract
Manual visual inspection on assembly lines is a persistent manufacturing bottleneck: operator fatigue over extended shifts lowers defect-detection rates. This paper presents the design, integration, and field deployment of an AI-assisted collaborative inspection cell at the Silverline kitchen-appliance factory, developed within the AI-PRISM project. The cell couples a Universal Robots UR 10e cobot carrying a machine-vision defect-detection pipeline with a Comau Racer-5 cobot for functional tests, coordinated through ROS 2 Humble on an Ubuntu 22.04 LTS server. Multi-modal data (Basler camera imagery, TIA microphone acoustics, and SPS electrical-safety measurements) are logged locally and visualised in real time with Grafana. We report the practical deployment challenges (close-proximity safety, AI robustness under glare and reflections, ROS 2 namespace collisions across two cobots, and operating-system and dependency issues) together with the engineering solutions adopted, and structure the integration through a four-level Human-Robot Interaction analysis. The deployed cell cuts per-unit quality-check time from 82 s to 61 s (about 25%), raises final-control resource efficiency from 0.75 to 0.88, reduces operator visual-inspection viewing time by 82%, and significantly lowers operator mental demand (p = 0.005, NASA-TLX).
Chinese Translation
装配线上的手动视觉检测一直是一个制造瓶颈:操作员在长时间轮班中疲劳会降低缺陷检测率。本文介绍了在Silverline厨房电器工厂设计、集成和现场部署的AI辅助协作检测单元,该单元是在AI-PRISM项目框架内开发的。该单元将搭载机器视觉缺陷检测流水线的Universal Robots UR 10e协作机器人与用于功能测试的Comau Racer-5协作机器人相结合,通过Ubuntu 22.04 LTS服务器上的ROS 2 Humble进行协调。多模态数据(Basler相机图像、TIA麦克风声学和SPS电气安全测量)在本地记录,并使用Grafana进行实时可视化。我们报告了实际部署中的挑战(近距离安全、强光和反射下的AI鲁棒性、两个协作机器人之间的ROS 2命名空间冲突以及操作系统和依赖问题)以及所采用的工程解决方案,并通过四级人机交互分析来构建集成。部署的单元将每单位质量检查时间从82秒减少到61秒(约25%),将最终控制资源效率从0.75提高到0.88,将操作员的视觉检测查看时间减少了82%,并显著降低了操作员的心理需求(p = 0.005,NASA-TLX)。
cs.RO / 83 / 2609.33542

A Disk-Shaped Magnetoelastic Torque Sensor for Robotic Joints Using Permanent Magnetization

一种用于机器人关节的基于永久磁化的圆盘形磁致弹性扭矩传感器
Brajon, Bruno, Gasparin, Enrico, Yazigy, Nicole, Salamin, Lisa, Copt, Florian, Gallego, Guzmán Borque, Klauser, Elias, Close, Gaël
Abstract
Direct torque sensing is a growing need in the robotics community to enable precise control and interactions where torque estimation from motor current is not sufficient. This paper presents a novel disk-shaped magnetoelastic torque sensor with a compact axial envelope of about 1 cm, suitable for integration in robotic joints. A four-magnetometer architecture is used to measure the field modulated by the stress affecting a narrow magnetized region while rejecting the effects of parasitic cross forces. A custom-designed magnetic shield enhances the torque sensitivity while reducing the external stray fields by a factor 7x. The device measures the torque with an accuracy of 1.34 %FS relative to the 50 Nm full scale (FS). The paper details the development of the sensor through the mechanical design, the magnetization procedure, and the experimental validation. The results demonstrate the potential of the proposed sensor for robotic applications.
Chinese Translation
直接扭矩传感在机器人领域是一种日益增长的需求,旨在实现精确控制和交互,而仅通过电机电流进行扭矩估计是不够的。本文提出了一种新颖的圆盘形磁致弹性扭矩传感器,其轴向尺寸紧凑,约为1厘米,适合集成到机器人关节中。采用四磁强计架构来测量由影响狭窄磁化区域的应力所调制的磁场,同时抑制寄生交叉力的影响。定制设计的磁屏蔽提高了扭矩灵敏度,同时将外部杂散场降低了7倍。该设备测量扭矩的精度为1.34%满量程(FS),相对于50 Nm的满量程。本文详细介绍了传感器的开发过程,包括机械设计、磁化过程和实验验证。结果证明了所提出的传感器在机器人应用中的潜力。
cs.RO / 84 / 2609.33546

Steer2Grasp: Inference-Time Embodiment-Aware Steering for Diverse Physically Feasible Grasp Diffusion

Steer2Grasp: 用于多样物理可行抓取扩散的推理时具身感知引导
Vembar, Vignesh, Kaura, Ayush, Padmaprabhan, A, Sinha, Siddharth, Nagarajan, Kailash, Patra, Keshab, Karim, Md Faizal, Krishna, K Madhava
Abstract
Current grasp diffusion models provide rich priors for generation, yet their object-centric approach can violate the kinematic and collision constraints imposed by the embodiment and the environment. Existing embodiment-aware methods primarily perform local corrections around generated grasps through gradient guidance or optimization, making it difficult to recover from fundamentally infeasible modes. We present Steer2Grasp, a training-free, embodiment-agnostic framework for inference-time grasp steering that adapts a frozen Cartesian grasp diffusion model using deployment-specific rewards. Through Feynman-Kac (FK) inspired particle reweighting and resampling, the method reallocates population mass from infeasible to high-reward grasp modes, enabling population-level mode transitions without modifying the pretrained diffusion model or requiring differentiable constraints. The framework enables a unified treatment for single and dual arm grasping through reachability and collision aware rewards, followed by gradient free gripper level local refinement. Across diverse objects, robot embodiments, and constrained environments, our method substantially improves feasible grasp generation while maintaining proximity to the underlying grasp prior.
Chinese Translation
当前抓取扩散模型为生成提供了丰富的先验,然而其以物体为中心的方法可能违反由具身和环境施加的运动学和碰撞约束。现有的具身感知方法主要通过梯度引导或优化在生成的抓取周围进行局部修正,这使得难以从根本不可行的模式中恢复。我们提出了 Steer2Grasp,一个无需训练、具身无关的推理时抓取引导框架,它使用部署特定的奖励来适配一个冻结的笛卡尔抓取扩散模型。通过受 Feynman-Kac (FK) 启发的粒子重加权和重采样,该方法将群体质量从不可行模式重新分配到高奖励抓取模式,从而无需修改预训练扩散模型或要求可微约束即可实现群体级模式转换。该框架通过可达性和碰撞感知奖励实现了对单臂和双臂抓取的统一处理,随后进行无梯度夹爪级局部细化。在多样物体、机器人具身和受限环境中,我们的方法显著提高了可行抓取生成,同时保持与底层抓取先验的接近性。
cs.RO / 85 / 2609.33551

FoLD: Force-Informed Learning for Dexterous Articulated Object Manipulation

FoLD:面向灵巧铰接物体操作的力引导学习
Shen, Haowei, Li, Tingai, Liu, Yumeng, Guang, Wenyuan, Yang, Xuanze, Fang, Qing, Xu, Kai, Liu, Ligang, Hu, Ruizhen
Abstract
Transferring human demonstrations to dexterous robots remains challenging because differences in hand morphology and contact dynamics often cause retargeted motions to fail at producing the intended object behavior. We present \textbf{FoLD}, a framework for learning dexterous manipulation of articulated objects through explicit force guidance. FoLD compute compensatory force fields from human demonstrations together with the robot's current interaction state, yielding a force prior that promotes the demonstrated object motion. This force prior informs a residual policy that adapts retargeted hand motions to the contact requirements of the task. We evaluate FoLD on a public benchmark for articulated object manipulation, where it consistently outperforms state-of-the-art baselines across tasks and embodiments. We further validate FoLD on real dexterous robot platforms, demonstrating successful transfer of human manipulation skills to robot execution. Here is the link of our project page: https://gghgghgghgg.github.io/FoLD-project-page/.
Chinese Translation
将人类演示迁移到灵巧机器人仍然具有挑战性,因为手部形态和接触动力学的差异常常导致重定向后的动作无法产生预期的物体行为。我们提出 FoLD,一个通过显式力引导来学习铰接物体灵巧操作的框架。FoLD 根据人类演示以及机器人当前交互状态计算补偿力场,从而产生促进所演示物体运动的力先验。该力先验指导一个残差策略,使重定向后的手部动作适应任务的接触需求。我们在一个用于铰接物体操作的公开基准上评估 FoLD,其在跨任务和不同机器人本体上始终优于最先进的基线。我们进一步在真实灵巧机器人平台上验证 FoLD,证明人类操作技能可以成功迁移到机器人执行。我们的项目页面链接如下:https://gghgghgghgg.github.io/FoLD-project-page/。
cs.RO / 86 / 2609.33575

SLIP-VLA: Single-Step Latent Imagination for Policy Learning in Vision-Language-Action Models

SLIP-VLA:用于视觉-语言-动作模型中策略学习的单步潜在想象
Li, Tianfu, Xu, Haoxuan, Chen, Wenbo, Li, Haitian, Yang, Changchuan, Zheng, Xinhu, Ma, Jun, Liu, Yuan, Wang, Lujia, Li, Haoang
Abstract
Vision-Language-Action models are increasingly effective for robotic manipulation, yet most predict actions directly from current observations without explicitly modeling future scene evolution. Recent methods introduce future prediction to improve action generation, but dense future modeling often requires expensive iterative denoising, while one-step alternatives can underperform their multi-step counterparts. To reconcile efficient future modeling with strong action performance, we present SLIP-VLA, a policy learning framework that equips VLA models with a Single-Step Latent Imagination for future-aware action prediction. SLIP-VLA obtains temporally dense future latent representations with a single denoising update, and we improve the perceptual sufficiency of these representations by aligning intermediate latents with future geometric and semantic features. We further improve their control sufficiency through action-conditioned latent world modeling and inverse dynamics modeling, explicitly coupling latent transitions with robot actions. SLIP-VLA achieves state-of-the-art performance across diverse simulation benchmarks and real-world manipulation tasks, while its single-step latent imagination takes only 12 ms.
Chinese Translation
视觉-语言-动作模型在机器人操作中日益有效,但大多数方法直接从当前观测预测动作,而未显式建模未来场景演化。近期方法引入未来预测以改进动作生成,但密集未来建模通常需要昂贵的迭代去噪,而单步替代方案可能不如多步方法。为了协调高效未来建模与强动作性能,我们提出SLIP-VLA,一种策略学习框架,为VLA模型配备单步潜在想象,以实现未来感知的动作预测。SLIP-VLA通过单次去噪更新获得时间上密集的未来潜在表示,并通过将中间潜在表示与未来几何和语义特征对齐,提高这些表示的感知充分性。我们进一步通过动作条件潜在世界建模和逆动力学建模提高其控制充分性,将潜在转移与机器人动作显式耦合。SLIP-VLA在不同仿真基准和真实世界操作任务中实现了最先进性能,而其单步潜在想象仅需12毫秒。
cs.RO / 87 / 2609.33595

Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning

超越单步精度:用于可靠视觉规划的状态仿射潜在转换
Zhang, Boyuan, Du, Yingjun, Zhen, Xiantong, Shao, Ling
Abstract
Joint-embedding world models enable visual planning by learning action-conditioned dynamics in latent space. Yet they are commonly trained for one-step prediction on encoded states, while planning recursively applies the learned transition to its own predictions. One-step accuracy therefore does not capture how prediction errors propagate under recursive rollout. We decompose multi-step rollout error into the errors introduced at individual steps and their propagation through subsequent transitions. We show that state-affine dynamics are precisely the differentiable transitions with state-independent Jacobians, eliminating the nonlinear propagation residual and making the error propagation operators depend only on the action sequence. Guided by this result, we introduce SALT (State-Affine Latent Transition), an action-conditioned state-affine dynamics model in which the action modulates both the state transformation and the additive update. We train SALT through recursive multi-step rollout supervision, feeding each predicted latent state back into the transition so that training matches how the model is used during planning. Across four visual planning environments, SALT exhibits $1.48$--$2.19\times$ higher one-step prediction error than the matched LeWM baseline, yet improves closed-loop success in every environment by $10.0$ percentage points on average. On OGBench-Cube, the fraction of episodes that fail with a sharp rise in model-predicted cost after execution decreases from $23.3%$ to $2.0%$.
Chinese Translation
联合嵌入世界模型通过在潜在空间中学习动作条件动力学来实现视觉规划。然而,它们通常被训练为在编码状态上进行单步预测,而规划则递归地将学习到的转换应用于其自身的预测。因此,单步精度无法捕捉预测误差在递归展开下如何传播。我们将多步展开误差分解为各个步骤引入的误差及其通过后续转换的传播。我们证明,状态仿射动力学正是具有状态无关雅可比矩阵的可微转换,消除了非线性传播残差,并使误差传播算子仅取决于动作序列。在此结果的指导下,我们引入了SALT(State-Affine Latent Transition),一种动作条件的状态仿射动力学模型,其中动作调节状态变换和加性更新。我们通过递归多步展开监督来训练SALT,将每个预测的潜在状态反馈到转换中,从而使训练与规划期间模型的使用方式相匹配。在四个视觉规划环境中,SALT的单步预测误差比匹配的LeWM基线高1.48至2.19倍,但在每个环境中将闭环成功率平均提高10.0个百分点。在OGBench-Cube上,执行后模型预测成本急剧上升而失败的片段比例从23.3%下降到2.0%。
cs.RO / 88 / 2609.33647

InfraVLA: Extending Vision-Language-Action Navigation with Infrastructure Cameras

InfraVLA:用基础设施摄像头扩展视觉-语言-动作导航
Vierling, Lukas, Ramtoula, Benjamin, Robinson, Luke, Clark, Ronald, De Martini, Daniele
Abstract
Many indoor environments in which robots operate, such as warehouses, offices, and hospitals, already have cameras installed. They observe parts of the building that the robot cannot see from where it stands, yet navigation policies, including recent vision-language-action (VLA) models, do not use them. We propose InfraVLA, an end-to-end method that adapts a pretrained navigation VLA to such static infrastructure views: a closed-circuit television (CCTV) encoder turns each external view into tokens of the input sequence. Because the views matter only at rare decision points, fine-tuning alone did not make the policy use them in our experiments; we therefore train in two stages, on demonstrations with upsampled counterfactual data and then on recovery data. We evaluate on two simulated warehouse tasks, finding an object named in the instruction and rerouting around blocked aisles, where the deciding information is often visible only to the infrastructure cameras. Tested in distribution, InfraVLA reached a success rate of 100% on both, against 34.0% and 73.6% for a baseline without CCTV input. On out-of-distribution test sets it reached 88.2% and 88.9%. On a real quadruped fine-tuned with under 10 minutes of demonstrations, the policy reached 83.3% against 29.2% for the on-board-only baseline.
Chinese Translation
机器人运行的许多室内环境,如仓库、办公室和医院,已经安装了摄像头。这些摄像头可以观察到机器人从所在位置无法看到的建筑部分,然而导航策略,包括最近的视觉-语言-动作(VLA)模型,并未利用它们。我们提出了 InfraVLA,一种端到端方法,将预训练的导航 VLA 适应于此类静态基础设施视图:闭路电视(CCTV)编码器将每个外部视图转换为输入序列的标记。由于这些视图仅在少数决策点起作用,在我们的实验中,仅微调并不能使策略利用它们;因此,我们分两个阶段进行训练:首先在带有上采样反事实数据的演示上进行训练,然后在恢复数据上进行训练。我们在两个模拟仓库任务上进行评估:根据指令找到指定物体,以及绕过被阻塞的通道重新规划路线,其中决定性信息通常仅对基础设施摄像头可见。在分布内测试中,InfraVLA 在两个任务上均达到了 100% 的成功率,而无需 CCTV 输入的基线方法分别为 34.0% 和 73.6%。在分布外测试集上,它达到了 88.2% 和 88.9%。在一个用不到 10 分钟演示进行微调的真实四足机器人上,该策略达到了 83.3%,而仅使用机载摄像头的基线方法为 29.2%。
cs.RO / 89 / 2609.33653

Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies

面向通用机器人策略的无需演示成功概率奖励学习
Wu, Duo, Wang, Haifeng, Lu, Rongwei, Wang, Jinghe, Xiong, Tianyi, Wang, Zhimin, Yu, Chao, Ma, Shuai, Wang, Zhi
Abstract
Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introduce a distribution mismatch with the mixed-quality rollouts encountered during policy optimization, making their estimates unreliable on suboptimal and failed behaviors from which the policy must learn. In this work, we introduce a demonstration-free reward learning paradigm where dense reward feedback can be learned directly from sparse task outcomes and policy experience. We theoretically show that terminal task outcomes implicitly define dense success-probability feedback at intermediate timesteps, which can be recursively learned through bootstrapping. Based on this insight, we introduce eVTA$_0$, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations. We further introduce RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA$_0$ using newly collected rollouts as the policy evolves. Experiments show that eVTA$_0$ provides more informative rewards than state-of-the-art reward models and achieves the best average policy performance across all LIBERO task suites under the same RL training budget, improving success rates by 5.4%-13.8% over the initial policy. In real-world manipulation, RLER further improves overall success rates by 20%-26%, with 35%-36% gains under out-of-distribution conditions. These results demonstrate the effectiveness of demonstration-free reward learning and adapting rewards as the policy evolves. Project webpage: https://duowuyms.github.io/evta0.
Chinese Translation
强化学习(RL)使通用机器人策略能够通过试错交互得到改进,但其有效性从根本上受限于稀疏的任务奖励。现有通用奖励模型通常通过从专家演示中学习任务进展来缓解该问题,但这会与策略优化过程中遇到的混合质量轨迹产生分布不匹配,使其估计在策略必须从中学习的次优和失败行为上不可靠。在这项工作中,我们提出了一种无需演示的奖励学习范式,其中密集奖励反馈可直接从稀疏任务结果和策略经验中学习。我们从理论上表明,终端任务结果隐式地定义了中间时间步上的密集成功概率反馈,可通过自举(bootstrapping)递归学习。基于这一洞见,我们提出 eVTA$_0$,它通过时序差分风格的自举从混合质量的策略轨迹中学习成功概率,无需专家演示或中间标注。我们进一步提出演化奖励强化学习(RL with Evolving Rewards, RLER),这是一个闭环框架,随着策略演化,利用新收集的轨迹自适应调整 eVTA$_0$。实验表明,在相同 RL 训练预算下,eVTA$_0$ 比最先进的奖励模型提供更具信息量的奖励,并在所有 LIBERO 任务套件上取得最佳平均策略性能,相较初始策略将成功率提升 5.4%-13.8%。在真实世界操作任务中,RLER 进一步将总体成功率提升 20%-26%,在分布外条件下提升 35%-36%。这些结果证明了无需演示的奖励学习以及随策略演化自适应调整奖励的有效性。项目网页:https://duowuyms.github.io/evta0.
cs.RO / 90 / 2609.33681

Observability-Informed Optimal Sensor Placement for Soft Robots

基于可观性的软体机器人最优传感器布置
Smocot, Samuel, Forbes, James Richard, Sedal, Audrey
Abstract
This paper presents the application and experimental evaluation of a systematic method for optimal sensor placement in soft robots. Existing methods either lack generalizability across different soft robot morphologies or do not account for system dynamics. The applied method uses convex optimization to find the optimal sensor configuration that maximizes an observability Gramian-based metric. The framework is experimentally evaluated using position and strain measurements on a soft continuum arm. Kalman filter state estimates using optimal sensor placements yield lower reconstruction error than a baseline across all sinusoidal input trials, with improvements on the order of millimeters. This case study shows that linear control theory tools can guide optimal sensor placement in soft robots, suggesting an interpretable approach to sensor placement that may extend to other morphologies.
Chinese Translation
本文介绍了用于软体机器人最优传感器布置的一种系统方法的实际应用与实验评估。现有方法要么缺乏对不同软体机器人形态的泛化能力,要么未考虑系统动力学。所应用的方法使用凸优化来寻找最优传感器配置,该配置可最大化基于可观性格拉姆矩阵的指标。该框架在软体连续臂上利用位置和应变测量进行了实验评估。在所有正弦输入试验中,使用最优传感器布置的卡尔曼滤波状态估计比基线具有更低的重构误差,改进量约为毫米级。该案例研究表明,线性控制理论工具可以指导软体机器人中的最优传感器布置,提示了一种可解释的传感器布置方法,并可能推广到其他形态。
cs.RO / 91 / 2609.33737

MomWorld: Momentum-Aware Latent World Model for Long-Horizon Autonomous Driving

MomWorld:面向长时域自动驾驶的动量感知潜在世界模型
Song, Ziying, Zhang, Shengkai, Yang, Lei, Chi, Haozhuang, Liu, Yuchen, Su, Jiangtao, Liu, Lin, Liu, Ziyang, Lv, Chen
Abstract
Long-horizon planning enables autonomous vehicles to anticipate scene evolution and potential risks, supporting safe and stable decisions in complex interactions. However, existing methods struggle to propagate motion trends from observed history into the future. Long rollouts based on a single latent state may further attenuate useful dynamics, retain stale motion patterns, and disrupt reliable near-term plans. We introduce MomWorld, a momentum-aware latent world model for long-horizon planning. MomWorld extracts scene motion trends from historical-to-current observations and propagates latent momentum into future horizons, jointly predicting future configuration and momentum states. A learnable momentum persistence mechanism preserves stable trends, scene-conditioned momentum updates adapt future dynamics, and a scene-adaptive reset gate suppresses stale momentum under abrupt changes. We further propose MoFlow, a momentum-conditioned flow-matching module that refines a base trajectory to align with the predicted future scene evolution in only a few integration steps, with a horizon-aware residual fusion that preserves near-term planning stability while permitting stronger long-range corrections. Extensive experiments on NAVSIM, nuScenes and Bench2Drive demonstrate that MomWorld improves long-horizon planning consistency and reduces the average collision rate by 12.2% relative to MomAD over a 6-second planning horizon.
Chinese Translation
长时域规划使自动驾驶车辆能够预测场景演化和潜在风险,支持复杂交互中的安全稳定决策。然而,现有方法难以将观测历史中的运动趋势传播到未来。基于单一潜在状态的长时域推演可能进一步削弱有用的动态信息,保留过时的运动模式,并破坏可靠的近期规划。我们提出MomWorld,一种面向长时域规划的动量感知潜在世界模型。MomWorld从历史到当前的观测中提取场景运动趋势,并将潜在动量传播到未来时域,联合预测未来配置和动量状态。可学习的动量持续性机制保持稳定趋势,场景条件动量更新适应未来动态,场景自适应重置门在突变下抑制过时动量。我们进一步提出MoFlow,一个动量条件流匹配模块,仅需少量积分步骤即可细化基础轨迹,使其与预测的未来场景演化对齐,并采用时域感知残差融合,在保持近期规划稳定性的同时允许更强的长距离修正。在NAVSIM、nuScenes和Bench2Drive上的大量实验表明,MomWorld提高了长时域规划一致性,并在6秒规划时域上相对于MomAD将平均碰撞率降低了12.2%。
cs.RO / 92 / 2609.33748

AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models

AnyStep-WAM:面向世界动作模型的预算对齐蒸馏与自适应推理
Wang, Rui, Wang, Xiangyu, Yang, Donglin, Li, Yibo, Chen, Canyang, Wang, Zhongrui, Qi, Xiaojuan
Abstract
World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. Here we introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student-teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07%, 12.08%, and 8.94% on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.
Chinese Translation
世界动作模型(WAMs)将预测性视觉建模与动作生成相结合,通常依赖于固定去噪步数的迭代去噪。然而,操作任务中的动作块对生成误差的敏感性各不相同:关键动作需要精确性,而较不敏感的动作则允许用更少的去噪步数更快地生成。在此,我们提出了AnyStep World Action Model,一个用于可调预算预测和场景依赖计算分配的通用框架。我们的预算对齐教师轨迹蒸馏方法利用显式冻结教师转移和共享低秩适配器训练区间条件流映射,支持从一步预测到多步细化的动作生成。基于此能力,一个轻量级风险-收益调度器从单次一步预览中预测基于教师曲率的难度和特定预算的学生-教师保真度,选择预测能够满足风险自适应保真度要求的最小预算。我们在三个广泛使用的WAM(Motus、FastWAM和LingBotVA)上使用RoboTwin 2.0评估了我们的框架。我们的方法分别将平均去噪步数减少了60.2%、49.8%和85.28%,同时保持了基线任务成功率。特别是,我们的AnyStep训练显著提高了模型在一步去噪预算下的性能,在Motus、FastWAM和LingBotVA上分别将任务成功率提高了7.07%、12.08%和8.94%。在六个真实世界操作任务上的实验进一步验证了其有效性。
cs.RO / 93 / 2609.33765

Principal Steering Subspaces for Online Adaptation of Frozen Generative Robot Policies

面向冻结生成式机器人策略在线适应的主引导子空间
Ni, Jialeng, Zhao, Nathan, Song, Kunpeng
Abstract
Generative robot policies provide expressive behavior priors, but updating a large diffusion or flow-matching model through online interaction is costly. Latent-space reinforcement learning avoids updating the pretrained generator by controlling its initial sampling noise, yet high-dimensional noise can have strongly anisotropic effects on decoded actions. We introduce Principal Steering Subspaces (PSS), a forward-query interface that constructs a fixed low-dimensional control basis from finite-difference decoder responses. Soft Actor-Critic controls the leading response directions, while the orthogonal complement is independently resampled from the Gaussian prior at each query. On three RoboMimic tasks with diffusion and flow-matching policies, response spectra reveal substantial concentration. Across five matched task-generator pairs, the training curves indicate that PSS generally converges faster and exhibits more stable late-training behavior than full-latent control, while achieving stronger final performance overall. Controlled Diffusion-Square ablations further show that leading-response directions outperform random and least-responsive subspaces of equal dimension. We further integrate PSS with a frozen, closed-source 3B-parameter vision-language-action (VLA) policy in a humanoid learning system with synchronous transition collection, reset-time optimization, and latency-aware asynchronous deployment. In an exploratory screwdriver-placement evaluation, success is observed in 2/10 trials for the frozen VLA policy and 6/10 after SAC+PSS adaptation. These results support decoder-response geometry as a practical basis for online adaptation of frozen generative robot policies.
Chinese Translation
生成式机器人策略提供了富有表现力的行为先验,但通过在线交互更新大型扩散模型或流匹配模型代价高昂。潜空间强化学习通过控制预训练生成器的初始采样噪声来避免更新预训练生成器,然而高维噪声对解码动作可能产生强烈的各向异性影响。我们引入主引导子空间(Principal Steering Subspaces, PSS),一种前向查询接口,它从有限差分解码器响应中构建固定的低维控制基。Soft Actor-Critic 控制主导响应方向,而正交补在每次查询时从高斯先验中独立重采样。在三个采用扩散策略和流匹配策略的 RoboMimic 任务上,响应谱揭示出显著的集中性。在五对匹配的任务-生成器组合上,训练曲线表明,与全潜空间控制相比,PSS 通常收敛更快,后期训练行为更稳定,同时总体上取得更强的最终性能。受控的 Diffusion-Square 消融进一步表明,主导响应方向优于同等维度的随机子空间和响应最弱子空间。我们进一步将 PSS 与一个冻结的、闭源的 30 亿参数视觉-语言-动作(VLA)策略集成到一个类人学习系统中,该系统具有同步转移收集、重置时间优化和延迟感知异步部署。在一项探索性的螺丝刀放置评估中,冻结的 VLA 策略在 2/10 次试验中取得成功,而在 SAC+PSS 适应后达到 6/10。这些结果支持将解码器响应几何作为冻结生成式机器人策略在线适应的实用基础。
cs.RO / 94 / 2609.33807

CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation

CodeActionBench:评估用于具身操作的智能体式代码即策略
Lyu, Yiheng, Jiang, Xueying, Li, Wenhao, Lu, Shijian, Zhang, Gongjie
Abstract
How well can general-purpose multimodal models turn visual understanding and reasoning into embodied manipulation via executable code? We introduce CodeActionBench, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy. Without task-specific fine-tuning, demonstrations, external specialist perception or grasp modules, privileged scene state, or predefined task policies, agents should select visual evidence, form task-relevant 3D estimates, construct manipulation targets, and iteratively execute and revise their policies. A shared robot API provides RGB observations, calibrated geometric operations, robot feedback, and bounded motion, leaving task-dependent decisions to the evaluated agent. Fixed task instances, resource budgets, and a hidden physical-outcome verifier support controlled comparisons across models and harness configurations. Extensive evaluations across nine configurations and 675 attempts achieve success rates ranging from 2.7% to 73.3%. The strongest configuration, GPT-6 Astra with Codex CLI, solves 22 of 25 tasks at least once in three attempts, demonstrating the best performance while still leaving substantial room for improvement. Trajectory analyses reveal difficulties in spatial alignment, object retention, and completion judgment, including task failures despite successfully completed motions. CodeActionBench provides a controlled testbed for measuring how general-purpose models translate their capabilities into manipulation behavior and for examining typical failure scenarios in that process.
Chinese Translation
通用多模态模型能够通过可执行代码将视觉理解和推理转化为具身操作的程度如何?我们介绍了 CodeActionBench,一个包含25个操作任务的基准,通过智能体式代码即策略(Agentic Code-as-Policy)评估这一能力。在没有任务特定微调、演示、外部专家感知或抓取模块、特权场景状态或预定义任务策略的情况下,智能体应选择视觉证据,形成任务相关的3D估计,构建操作目标,并迭代执行和修订其策略。共享的机器人API提供RGB观测、校准的几何操作、机器人反馈和有界运动,将任务相关的决策留给被评估的智能体。固定的任务实例、资源预算和隐藏的物理结果验证器支持跨模型和评估框架配置的受控比较。在九种配置和675次尝试中进行的广泛评估取得了从2.7%到73.3%的成功率。最强的配置,GPT-6 Astra 配合 Codex CLI,在三次尝试中至少解决25个任务中的22个,展示了最佳性能,但仍留有巨大的改进空间。轨迹分析揭示了在空间对齐、对象保持和完成判断方面的困难,包括尽管成功完成了动作但任务失败的情况。CodeActionBench 提供了一个受控测试平台,用于衡量通用模型如何将其能力转化为操作行为,并检查该过程中的典型失败场景。
cs.RO / 95 / 2609.33832

Achieve What You Imagined: Learning to Align Actions with Visual Plans

实现你所想象的:学习将动作与视觉计划对齐
Qiao, Yuheng, Wei, Ziran, Wang, Xiaohan, Guo, Daqiang, Luo, Yichen, Pang, Zhibo, Zhou, Peng, Liu, Sichao
Abstract
World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausible task-completion outcomes before producing action sequences that reliably achieve them. Consequently, we treat the WAM-generated visual prediction as a goal-conditioned visual proposal rather than a directly executable plan. We use a frozen action-conditioned world model to predict action-conditioned consequences and construct feedback based on consistency between the two future predictions and alignment with the terminal goal. Leveraging this feedback, we employ Flow Policy Optimization (FPO) to optimize the action head of the WAM. This framework avoids online robot interaction and additional training of task-specific reward models. Across four real-world UR5 manipulation tasks, our method increases the mean success rate from 43.4% to 75.1%, compared with 61.4% for $\pi_{0.5}$. These results show that cross-model prediction discrepancy can provide useful feedback for improving robot policies under the evaluated manipulation tasks. Website: https://imagine-to-achieve.github.io/
Chinese Translation
世界-动作模型(World-Action Models, WAMs)可以联合预测未来视觉观测和机器人动作。然而,它们的视觉预测与所生成动作隐含的后果之间可能存在差异。我们观察到,WAMs 往往能够先生成视觉上合理的任务完成结果,然后才生成能够可靠实现这些结果的动作序列。因此,我们将 WAM 生成的视觉预测视为一种以目标为条件的视觉提议,而非可直接执行的计划。我们使用一个冻结的动作条件世界模型来预测动作条件下的后果,并基于两个未来预测之间的一致性以及与最终目标的对齐程度构建反馈。利用该反馈,我们采用流策略优化(Flow Policy Optimization, FPO)来优化 WAM 的动作头。该框架避免了在线机器人交互以及任务特定奖励模型的额外训练。在四个真实世界 UR5 操作任务中,我们的方法将平均成功率从 43.4% 提升到 75.1%,作为对比,π0.5 为 61.4%。这些结果表明,在所评估的操作任务中,跨模型预测差异可以为改进机器人策略提供有用的反馈。网站:https://imagine-to-achieve.github.io/
cs.RO / 96 / 2609.33836

DeltaSeek: Toward Active Perception in Evolving Construction Environments

DeltaSeek:面向演化施工环境中的主动感知
Acharjee, Sanjay, Sakib, Md Nazmus
Abstract
Construction environments evolve continuously, causing large geometric changes that degrade static mapping and registration performance. This necessitates active perception, where robots deliberately select sensing configurations to resolve the environment's current state. We present DeltaSeek, an initial framework toward active perception in evolving built environments. While our broader objective is a system that reasons about where, how, and when to observe, this paper addresses a critical prerequisite: how a robot's sensing embodiment constrains the observations it can acquire. We formalize an embodiment's permissible observation set and evaluate with a Husky A300 equipped with a UR5e on an IFC-derived benchmark under chassis-mounted and wrist-mounted RGB-D configurations, scoring observations by geometric visibility and effort by drivable distance. In a room-scale scene with eight controlled changes spanning four observability conditions, exhaustive evaluation over 240 permissible base poses and five arm postures shows that two changes admit no chassis viewpoint whatsoever, while the wrist camera resolves both. For changes observed by both embodiments, the median base travel is $6.0$~m for the wrist camera and $15.2$~m for the chassis camera. These results distinguish sensing limitations from acquisition costs, clarifying whether an observation is impossible or simply requires more travel.
Chinese Translation
施工环境持续演化,导致大的几何变化,从而降低静态建图和配准的性能。这需要主动感知,即机器人有意选择传感配置以解析环境的当前状态。我们提出 DeltaSeek,一个面向演化建筑环境中主动感知的初始框架。虽然我们更广泛的目标是构建一个能够推理在何处、如何以及何时进行观测的系统,但本文解决了一个关键前提:机器人的传感具身如何约束其可获取的观测。我们形式化了一个具身的允许观测集,并在基于 IFC 的基准上,使用配备 UR5e 的 Husky A300 在底盘安装和腕部安装的 RGB-D 配置下进行评估,通过几何可见性对观测进行评分,通过可行驶距离对代价进行评分。在一个房间尺度的场景中,有八个受控变化,跨越四种可观测性条件,对 240 个允许的基座位姿和五种手臂姿态进行穷举评估,结果表明两个变化不存在任何底盘视角,而腕部相机解决了这两个变化。对于两种具身都能观测到的变化,腕部相机的中位基座行程为 6.0 米,底盘相机为 15.2 米。这些结果区分了传感限制和获取代价,阐明了观测是不可能还是仅仅需要更多行程。
cs.RO / 97 / 2609.33872

Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation

Robot-GST:几何感知的时空机器人策略表示与评估
Liu, Sichao, Wang, Zekun, Tang, Lixuan, Li, Yiming, Wang, Xiaohan, Zhang, Hanzhi, Guo, Daqiang, Zhou, Peng, Wang, Lihui
Abstract
Robotic manipulation policies are advancing rapidly with increasing reliance on vision-language models for end-to-end decision making. However, reliable deployment remains challenging because many policies lack explicit mechanisms for predicting task outcomes and evaluating whether generated actions will achieve desired final states, causing execution errors to accumulate during long-horizon manipulation. We present Robot-GST, a geometry-aware spatio-temporal behaviour representation and evaluation framework that constructs a Gaussian-SAM robotic environment for real-to-sim policy verification and improves the reliability of real-world manipulation deployment. Our approach constructs a high-fidelity robotic environment from RGB-D observations using 3D Gaussian Splatting and SAM3D, enabling ``simulation and evaluation before acting''. It integrates visual observations and language instructions with spatio-temporal reasoning for long-horizon task planning using large vision-language models. To bridge high-level planning and real-world execution, we introduce Gaussian-aware final-state estimation through geometric sampling and state-based trajectory planning. Before execution, candidate action sequences are simulated and evaluated in the Gaussian-SAM environment to filter infeasible behaviours. We validate our approach on representative manipulation tasks involving rigid, soft, and deformable objects, including cube placing, toy packing, and duck rearrangement, demonstrating that geometry-aware spatio-temporal reasoning and state-aware execution improve manipulation reliability across different object categories. Our results suggest that combining geometry-aware reconstruction with high-quality rendering and simulation provides a scalable approach for evaluating robotic manipulation behaviours. Website: https://robot-gst.github.io
Chinese Translation
机器人操作策略正迅速发展,越来越依赖视觉-语言模型进行端到端决策。然而,可靠部署仍然具有挑战性,因为许多策略缺乏显式机制来预测任务结果并评估生成的动作是否能够达到期望的最终状态,导致在执行长时程操作过程中误差累积。我们提出 Robot-GST,一个几何感知的时空行为表示与评估框架,它构建了一个 Gaussian-SAM 机器人环境,用于真实到仿真(real-to-sim)的策略验证,并提高了真实世界操作部署的可靠性。我们的方法利用 3D Gaussian Splatting 和 SAM3D 从 RGB-D 观测中构建高保真机器人环境,实现了“先仿真评估,后执行动作”。它使用大型视觉-语言模型,将视觉观测和语言指令与时空推理相结合,用于长时程任务规划。为了连接高层规划与真实世界执行,我们引入通过几何采样和基于状态的轨迹规划进行高斯感知的最终状态估计。在执行之前,候选动作序列在 Gaussian-SAM 环境中进行仿真和评估,以过滤掉不可行的行为。我们在涉及刚性、柔软和可变形物体的代表性操作任务上验证了我们的方法,包括立方体放置、玩具打包和鸭子重排,表明几何感知的时空推理和状态感知的执行提高了不同物体类别下的操作可靠性。我们的结果表明,将几何感知重建与高质量渲染和仿真相结合,为评估机器人操作行为提供了一种可扩展的方法。网站:https://robot-gst.github.io
cs.RO / 98 / 2609.33882

DexTaG: Tactile-as-Guidance in Reinforcement Learning for Dexterous Manipulation

DexTaG:以触觉为引导的强化学习用于灵巧操作
Yang, Han, Wang, Yian, Song, Yunlong, Xu, Zhenjia, Gan, Chuang
Abstract
Glove-based motion capture is emerging as a scalable approach to collecting dexterous-hand demonstration data. However, due to the kinematic gap between the human and robot hand, the recorded human motions cannot be executed directly on the robot, especially for contact-rich tool-use tasks involving in-hand reorientation. Prior work bridges this gap in simulation through reinforcement learning (RL) or trajectory optimization, but the human contact pattern is hard to preserve under such formulations, often producing unnatural manipulation and unstable functional grasps. These methods also train a separate policy or solve a separate optimization for each reference trajectory, which is inefficient. To solve these problems, we propose DexTaG, a tactile-guided RL framework for dexterous manipulation. During training, tactile signals captured by the glove guide policy search toward the measured human contact pattern, reducing reliance on precise reference geometry for contact supervision. To improve efficiency, we train a single generalizable retargeter jointly on all training trajectories of the same object. The retargeter is further distilled into a tactile-free student controller conditioned on the target object trajectory for real-world deployment. On marker-pen and hammer manipulation tasks, DexTaG learns natural, contact-rich behaviors that baselines with distance-based contact heuristics fail to learn, generalizes to held-out trajectories of the same object and task, and outperforms single-trajectory baselines on OakInk2.
Chinese Translation
基于手套的运动捕捉正成为一种可扩展的方法,用于收集灵巧手演示数据。然而,由于人手与机器人手之间的运动学差异,记录的人手运动无法直接在机器人上执行,特别是涉及手中重定向的富接触工具使用任务。先前的工作通过强化学习(RL)或轨迹优化在仿真中弥合了这一差距,但在这种形式下,人类的接触模式难以保持,往往产生非自然的操作和不稳定的功能性抓取。这些方法还为每个参考轨迹训练单独的策略或求解单独的优化,效率低下。为了解决这些问题,我们提出了DexTaG,一个用于灵巧操作的触觉引导强化学习框架。在训练期间,手套捕获的触觉信号引导策略搜索朝向测量到的人类接触模式,减少了对精确参考几何进行接触监督的依赖。为了提高效率,我们在同一对象的所有训练轨迹上联合训练一个可泛化的重定向器。该重定向器进一步蒸馏为一个无触觉的学生控制器,以目标物体轨迹为条件,用于真实世界部署。在马克笔和锤子操作任务上,DexTaG学习到了自然、富接触的行为,而基于距离的接触启发式基线方法无法学习到这些行为,并且能够泛化到同一物体和任务的留出轨迹,在OakInk2上优于单轨迹基线。
cs.RO / 99 / 2609.33892

Residual Learning-Based Control of Vehicle Platoons with $\ell_2$ Stability Guarantees via Recurrent Equilibrium Networks

基于残差学习的车辆编队控制:通过递归均衡网络实现 $\ell_2$ 稳定性保证
Delgado, Brian, Nguyen, Anh-Tu, Taghavifar, Hamid
Abstract
This paper proposes a residual learning-based control framework for heterogeneous vehicle platoons subject to parametric uncertainty and external disturbances. A nominal controller designed via Linear Matrix Inequalities (LMIs), along with disturbance-observer compensation, is enhanced by a Recurrent Equilibrium Network (REN) trained offline using stored trajectories and nominal-model prediction errors. The REN is constrained to satisfy a prescribed $\ell_2$-gain bound, enabling sufficient small-gain conditions for local closed-loop stability and disturbance string stability. Experiments demonstrate reduced spacing and velocity errors relative to the nominal controller.
Chinese Translation
本文提出了一种基于残差学习的控制框架,用于存在参数不确定性和外部干扰的异构车辆编队。通过线性矩阵不等式(LMI)设计的标称控制器,结合干扰观测器补偿,并由使用存储轨迹和标称模型预测误差离线训练的递归均衡网络(REN)进行增强。REN 被约束以满足规定的 $\ell_2$ 增益界,从而为局部闭环稳定性和干扰弦稳定性提供充分的小增益条件。实验表明,相对于标称控制器,间距和速度误差减小。
cs.RO / 100 / 2609.33931

ArticulateArena: A Metric for Articulated Kinematics

ArticulateArena:一种用于铰接运动学的度量
He, Yumeng, She, Yongfei, Chen, Huanyu, Yuan, Chun, Li, Peihao, Masterjohn, Joseph, Yang, Yin, Jiang, Ying, Jiang, Chenfanfu
Abstract
Modern methods reconstruct or generate simulation-ready articulated objects, predicting not only their geometry but also how their parts are connected and allowed to move. Evaluating the geometry is straightforward, but evaluating the predicted articulation is not, because articulation specifies a motion rather than a shape, and there is no agreed distance between two motions. More specifically, existing protocols score joint type, axis direction, origin, and motion limits separately, although these parameters jointly describe a single physical motion, and the same motion can be written as different parameter values. As a result, a joint can score maximally wrong against an equivalent encoding of itself, and several component errors are ill-conditioned or undefined exactly where predictions become accurate. We propose ArticulateArena, a representation-invariant counterpart of Chamfer distance for articulation that compares the motions one-DOF joints induce rather than the parameters that encode them. It represents each joint by the unordered pair of its Lie-algebra endpoint twists, and we prove that the resulting quotient distance is a metric. It unifies fixed, revolute, prismatic, and helical joints, brings continuous joints into the same score through a compactification, and reads as the RMS motion of the moving part in meters when weighted by its mass distribution. A motion-aware tree edit distance lifts the metric to full kinematic trees, pricing structural errors such as spurious or missing joints in the same motion units as joint errors, and for a fixed inner product it remains a metric on trees up to relabeling. Alongside the metric we release ArticulateArena-20K, a new suite of 19,977 articulated objects with verified kinematics, and we re-evaluate published reconstruction methods on it under the new metric. Project page: https://heyumeng.com/ArticulateArena-web/
Chinese Translation
现代方法重建或生成可用于仿真的铰接物体,不仅预测其几何形状,还预测其部件如何连接以及允许如何运动。评估几何形状是直接的,但评估预测的关节连接则不然,因为关节连接指定的是运动而非形状,且两个运动之间没有公认的距离。更具体地说,现有协议分别对关节类型、轴方向、原点和运动限制进行评分,尽管这些参数共同描述了一个单一物理运动,且同一运动可以写成不同的参数值。因此,一个关节可能在与自身等价编码的对比中被判为最大错误,并且若干分量误差在预测变得准确的地方恰恰是病态的或未定义的。我们提出 ArticulateArena,一种用于关节连接的、表示不变的 Chamfer 距离对应物,它比较单自由度关节所诱导的运动,而不是编码它们的参数。它用其李代数端点 twist 的无序对来表示每个关节,并且我们证明所得的商距离是一个度量。它统一了固定、旋转、移动和螺旋关节,通过紧化将连续关节纳入同一评分,并且在按其质量分布加权时,读作运动部件以米为单位的均方根(RMS)运动。一种运动感知的树编辑距离将该度量提升到完整运动学树,以与关节误差相同的运动单位来衡量结构误差(如虚假或缺失的关节),并且对于固定内积,它在树重新标记的意义下仍然是一个度量。与该度量一起,我们发布了 ArticulateArena-20K,一套包含 19,977 个具有验证运动学的铰接物体的新数据集,并在新度量下重新评估了已发表的重建方法。项目页面:https://heyumeng.com/ArticulateArena-web/
cs.RO / 101 / 2609.33939

EpiTransfer: Sparse, Training-Free Long-Range Depth Estimation from Temporal Monocular Aerial Frames

EpiTransfer:基于时序单目航空帧的稀疏、免训练长距离深度估计
Aggarwal, Diksha, Dagadkhair, Rutvik, Srivastava, Sanjana, Denby, Bradley, Kochersberger, Kevin
Abstract
Reliable 3D spatial understanding is essential for autonomous navigation, obstacle avoidance, and scene reconstruction. While state-of-the-art learned depth estimation techniques achieve high accuracy in-distribution, they often generalize poorly to novel viewpoints and altitudes. This paper presents a geometrically derived, training-free depth estimation method using epipolar transfer with only two monocular images and camera pose estimates. By leveraging camera motion to synthesize a virtual stereo pair with a freely chosen baseline, our approach transforms temporal correspondence into a stereo triangulation task while mitigating geometric degeneracies inherent to direct two-view triangulation. Validated across outdoor drone flights (to a maximum range of approximately 90\,m) and indoor OptiTrack environments against LiDAR ground truth, the method achieves an indoor AbsRel of 0.092 and $\delta < 1.25$ of 0.940, comparable to direct triangulation (AbsRel 0.073) while retaining valid depth over a larger fraction of challenging scenes, and substantially outperforms off-the-shelf learning-based baselines such as ZoeDepth (AbsRel 0.225) and Depth Anything V2 (AbsRel 0.570), which are not trained or fine-tuned for this domain, with no training data required.
Chinese Translation
可靠的3D空间理解对于自主导航、避障和场景重建至关重要。尽管最先进的基于学习的深度估计技术在分布内实现了高精度,但它们往往难以泛化到新的视角和高度。本文提出了一种基于几何推导的免训练深度估计方法,利用对极转移(epipolar transfer),仅需两幅单目图像和相机位姿估计。通过利用相机运动合成具有自由选择基线的虚拟立体像对,我们的方法将时序对应转化为立体三角测量任务,同时减轻了直接双视图三角测量固有的几何退化问题。在户外无人机飞行(最大约90米)和室内OptiTrack环境下以LiDAR真值验证,该方法在室内达到AbsRel 0.092和δ<1.25为0.940,与直接三角测量(AbsRel 0.073)相当,同时在更大比例的挑战性场景中保持有效深度,并显著优于现成的基于学习的基线,如ZoeDepth(AbsRel 0.225)和Depth Anything V2(AbsRel 0.570),这些基线未针对该领域进行训练或微调,且无需训练数据。
cs.RO / 102 / 2609.33944

ReSync: Re-Aligning the Two Clocks of Asynchronous World-Action Models

ReSync:重新对齐异步世界-动作模型的两个时钟
Lin, Xi, Zhang, Feihong, Shi, Yulong, Mei, Yanghong, Lu, Zuxing, Zhu, Xiaofan, Liang, Zihao, Gao, Zhirui, Li, Zhaowen
Abstract
Jointly generating future video and actions has become a standard recipe for world-action models, and the strongest systems denoise the two streams on separate schedules: actions are decoded in few steps so control stays fast, while the video stream runs longer to keep the predicted future sharp. The design is deliberate, but it leaves the two streams on different clocks, and an action can become executable while the future that should justify it is still largely unresolved. We formalize this as a two-clock view of asynchronous inference and introduce the commitment-evidence gap, a quantity read directly from a model's own sampling schedule rather than measured by search. The gap is predictive: as it widens, candidate utility becomes harder to identify and extra candidate sampling buys less, while advancing the world stream buys more, and the two cross. Spending more world computation is therefore not simply better. The useful interval is closed at both ends, and both ends can be read off the schedule before any rollout. ReSync places the computation inside it: hold the action state, advance only the world within the supported window, then resume native denoising. No parameters change and no candidates are compared. On a frozen paired RoboCasa panel this improves success by 4.48 points, while an equal-compute control that waits without advancing the world does not move, and the same rule transfers to a second benchmark and a second backbone without retuning.
Chinese Translation
联合生成未来视频和动作已成为世界-动作模型的标准方法,而最强的系统在单独的调度上对两个流进行去噪:动作在少量步骤内解码以保持控制快速,而视频流运行更长时间以保持预测的未来清晰。该设计是有意为之,但它使两个流处于不同的时钟上,并且动作可能变得可执行,而本应证明其合理性的未来仍在很大程度上未解决。我们将此形式化为异步推理的双时钟视角,并引入承诺-证据差距,这是一个直接从模型自身的采样调度中读取的量,而不是通过搜索来测量的。该差距具有预测性:随着它扩大,候选效用变得更难识别,额外的候选采样收益更少,而推进世界流收益更多,两者交叉。因此,花费更多的世界计算并非简单更好。有用区间两端都是封闭的,并且两端都可以在任何 rollout 之前从调度中读出。ReSync 将计算放在其中:保持动作状态,仅在支持的窗口内推进世界,然后恢复原生去噪。没有参数变化,也没有比较候选。在一个冻结的配对 RoboCasa 面板上,这将成功率提高了 4.48 个百分点,而一个等计算量的控制组在不推进世界的情况下等待则没有变化,并且相同的规则无需重新调整即可迁移到第二个基准和第二个骨干网络。
cs.RO / 103 / 2609.33951

Integrity Detection and Characterization of Malicious Injections in RAVEN II

RAVEN II中恶意注入的完整性检测与表征
Zhang, Xingli, Afroze, Diba, Hu, Fei, Hei, Xiali
Abstract
The increasing adoption of robotic systems in surgery, together with the expanding range of procedures they can support and the growing level of autonomy they provide, has substantially increased the complexity of surgical robots. As these systems integrate more sensors, controllers, communication interfaces, and model-driven control components, their attack surface continues to expand. A compromise of the integrity of a surgical robot can therefore cause unintended robot behavior and potentially threaten patient safety. In this paper, we characterize the detection boundary of malicious injections on RAVEN II using a public dataset that pairs the platform's telemetry with external high-resolution encoder ground truth. We identify three injection points spanning the command and observation paths and evaluate three injection patterns with increasing temporal dispersion. To capture different detection behaviors, we perform detection at two timescales: the window scale and the session scale. Rather than reporting detection rates at an arbitrarily chosen threshold, we quantify, for each injection point and injection pattern, the smallest end-effector deviation that can be resolved while maintaining an alarm rate acceptable for surgical operation. Our results show that detectability is strongly influenced by how the injected deviation is distributed over time. An abrupt step can be detected at deviations well below the 1 mm clinical tolerance, whereas the same overall deviation spread across a window or a session can remain hidden from single-window statistics. The open source code can be found at http://github.com/RAVENIIROS/RAVENIIIntegrity.
Chinese Translation
机器人系统在手术中的日益普及,加上它们能够支持的手术范围不断扩大以及所提供的自主性水平不断提高,已大幅增加了手术机器人的复杂性。随着这些系统集成更多的传感器、控制器、通信接口和模型驱动的控制组件,它们的攻击面持续扩大。因此,手术机器人完整性的受损可能导致机器人出现非预期行为,并可能威胁患者安全。在本文中,我们利用一个将平台遥测数据与外部高分辨率编码器真值配对在一起的公共数据集,刻画了RAVEN II上恶意注入的检测边界。我们识别了跨越命令路径和观测路径的三个注入点,并评估了三种具有递增时间分散度的注入模式。为了捕捉不同的检测行为,我们在两个时间尺度上进行检测:窗口尺度和会话尺度。我们没有报告在任意选定阈值下的检测率,而是针对每个注入点和注入模式,量化了在保持手术可接受的报警率的同时能够分辨的最小末端执行器偏差。我们的结果表明,可检测性受到注入偏差随时间分布方式的强烈影响。一个突变的阶跃可以在远低于1毫米临床容差的偏差下被检测到,而同样的总体偏差如果分散在一个窗口或一个会话中,则可能仍无法被单窗口统计检测到。开源代码可在 http://github.com/RAVENIIROS/RAVENIIIntegrity 找到。
cs.RO / 104 / 2609.33973

FINGR: Learning Dexterous Hand Control for Real-World Rubik's Cube Solving

FINGR:学习灵巧手控制以实现真实世界魔方还原
Liang, Yutong, Peng, Quanquan, Kim, Matthew, Wang, Xiaolong
Abstract
Manipulating a Rubik's Cube with a single dexterous hand is a challenging test of sustained, contact-rich control: the hand must execute successive layer turns while keeping the cube secure. Each turn requires some fingers to support the cube while others push a moving layer, release contact, and reset for the next move. To learn this coordination, we introduce FINGR (Future-supervised Interaction Network with Geometric Representations), a policy that combines finger-relative geometry with future interaction prediction. A shared point encoder expresses the cube relative to each fingertip and aggregates its points without depending on cubie indexing. Learned future tokens share the observation encoder and receive supervision for contact-force changes, layer-turn progress, and finger joint displacement at multiple time scales. The resulting representation conditions a flow policy that directly generates finger actions. On a real dexterous hand, our policy achieves 99.0% success over 300 turn attempts, compared with 79.7% for the base flow policy. Integrated with grasping and table-assisted regrasping, the policy solves all ten scrambled $2\times2\times2$ cubes in a mean complete-system time of approximately 137 seconds. The project website is available at https://www.lyt0112.com/projects/FINGR
Chinese Translation
用单只灵巧手操作魔方是对持续、富接触控制的一项挑战性测试:手必须连续执行层旋转,同时保持魔方稳固。每次旋转需要一些手指支撑魔方,而其他手指推动移动层、释放接触并为下一步复位。为了学习这种协调,我们提出了FINGR(未来监督的几何表示交互网络),这是一种将手指相对几何与未来交互预测相结合的策略。一个共享的点编码器相对于每个指尖表达魔方,并聚合其点,而不依赖于小方块索引。学习到的未来标记共享观测编码器,并在多个时间尺度上接收接触力变化、层旋转进度和手指关节位移的监督。得到的表示作为条件输入到一个直接生成手指动作的流策略中。在真实的灵巧手上,我们的策略在300次旋转尝试中达到99.0%的成功率,而基础流策略为79.7%。结合抓取和桌面辅助重抓取,该策略解决了全部十个打乱的2×2×2魔方,平均完整系统时间约为137秒。项目网站位于https://www.lyt0112.com/projects/FINGR
cs.RO / 105 / 2609.33982

Test-Time Spatial Reasoning for Robot Manipulation Using Generative Real-to-Sim

使用生成式 Real-to-Sim 的机器人操作测试时空间推理
Kapelyukh, Ivan, Hu, Yafei, Gong, Ran, May, Brandon, Kusnur, Tushar, Herlant, Laura, Schmeckpeper, Karl, Johns, Edward, Zhang, Xiaohan
Abstract
Spatial reasoning is fundamental to general robot intelligence, as it enables robots to complete long-horizon tasks involving multi-object interaction. We introduce Simify, a training-free, test-time framework that performs explicit spatial reasoning via massively parallel physics simulation. From a single RGB-D image of a scene, Simify reconstructs simulation-ready assets leveraging 3D generative models and vision-language models. Then given a task specified by a reward function (e.g., build the tallest tower), Simify launches thousands of parallel rollouts in simulation and performs an evolutionary search to optimize object arrangements, typically converging within seconds. We conduct quantitative experiments on real-robot hardware to demonstrate the ability of our framework to execute complex object rearrangement tasks end-to-end with previously unseen objects. Results show that our framework outperforms prior work on foundation models for spatial reasoning by effectively exploiting large-scale parallel simulation during inference, and also highlight the importance of complete and accurate geometry for successful sim-to-real transfer.
Chinese Translation
空间推理是通用机器人智能的基础,因为它使机器人能够完成涉及多物体交互的长时程任务。我们提出了 Simify,一个无需训练、在测试时运行的框架,它通过大规模并行物理仿真执行显式空间推理。Simify 从场景的单张 RGB-D 图像中,利用 3D 生成模型和视觉-语言模型重建可直接用于仿真的资产。随后,给定由奖励函数指定的任务(例如,搭建最高的塔),Simify 在仿真中启动数千个并行推演,并执行进化搜索来优化物体布局,通常在数秒内收敛。我们在真实机器人硬件上进行了定量实验,以证明我们的框架能够端到端地执行涉及此前未见物体的复杂物体重排任务。结果表明,我们的框架通过在推理期间有效利用大规模并行仿真,优于先前关于空间推理基础模型的工作,并且还强调了完整且准确的几何结构对于成功实现 sim-to-real 迁移的重要性。
cs.RO / 106 / 2609.34006

TacGooseBumps (TacGB): Retrofitting Normal-Only Tactile Sensors with Shear Encoding for Learning Contact-Rich Manipulation

TacGooseBumps (TacGB):通过剪切编码改造仅法向触觉传感器以学习富接触操作
Li, Wenjie, Yang, Binyu, Chen, Yuxin, Wang, Ambrose, Tomizuka, Masayoshi
Abstract
Contact-rich policies often fail because distinct physical states look alike yet require different actions. Cameras may not reveal whether a connector is aligned or fully seated, while many normal-only tactile sensors can miss the tangential interactions perpendicular to the grasping direction that distinguish these states. We ask whether a learning policy needs calibrated shear measurements, or only a repeatable observation that separates shear-dependent contact states. We introduce TacGooseBumps (TacGB), a passive domed film that mechanically encodes tangential loading as pattern changes in an existing sensor's pressure map. Tangential loading tilts each dome and redistributes pressure across its footprint; an end-to-end policy consumes the resulting maps without added electronics, force reconstruction, or taxel-level dome alignment. Across four imitation-learning tasks and two data-collection pipelines, TacGB improves goal attainment, efficiency, and contact quality: insertion success increases by up to 36 percentage points, and successful insertions are completed faster, while fragile-object placement becomes gentler and drawing becomes more continuous and straight. Signal, stage-wise, failure-mode, and trajectory analyses link these gains to contact regimes in which task-relevant tangential interactions are poorly resolved by vision and normal pressure alone. Together, these results show that shear need not be measured metrically to benefit robot learning; it can instead be mechanically encoded without changing the underlying tactile sensor or the policy's pressure-map input format.
Chinese Translation
富接触策略常常失败,因为不同的物理状态看起来相似,却需要不同的动作。相机可能无法揭示连接器是对齐还是完全就位,而许多仅法向触觉传感器可能遗漏垂直于抓取方向的切向交互,这些交互正是区分这些状态的关键。我们探究学习策略是否需要经过校准的剪切测量,还是只需要一个可重复的观测来区分依赖剪切的接触状态。我们提出TacGooseBumps (TacGB),一种被动拱形薄膜,它通过机械方式将切向载荷编码为现有传感器压力图中的模式变化。切向载荷使每个圆顶倾斜,并在其覆盖区域内重新分布压力;端到端策略消费所得的压力图,无需额外电子元件、力重建或触觉单元级别的圆顶对齐。在四个模仿学习任务和两个数据收集流程中,TacGB提高了目标达成率、效率和接触质量:插入成功率最高提升36个百分点,成功插入完成得更快,同时易碎物体放置变得更轻柔,绘画变得更连续、更笔直。信号、阶段、失败模式和轨迹分析将这些增益与接触状态联系起来,在这些状态下,任务相关的切向交互仅靠视觉和法向压力难以分辨。总之,这些结果表明,剪切不需要被度量测量就能有益于机器人学习;相反,它可以被机械编码,而不改变底层触觉传感器或策略的压力图输入格式。
cs.RO / 107 / 2609.34010

ZeroBot: Learning from Scratch in Minutes with Generative Real2Sim

ZeroBot:使用生成式 Real2Sim 在几分钟内从零开始学习
Kapelyukh, Ivan, Zhang, Xiaohan, James, Stephen, Herlant, Laura, Johns, Edward
Abstract
We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Given only a single view of an object and a goal pose for that object, ZeroBot uses image-to-3D generative models to obtain a complete object mesh, which is used in simulation for large-scale parallel reinforcement learning. To accelerate training, we introduce an action space which leverages the generated geometry and learned value function to sample states involving robot-object contact. When evaluated on real-world tasks including grasping, pushing, articulated object interaction, and multi-stage manipulation, ZeroBot achieves an 87% success rate with an average training time of 119 seconds. These results show the value of using image-to-3D models in a real2sim framework for rapid, autonomous robot learning.
Chinese Translation
我们提出了 ZeroBot,一个 real2sim 框架,用于在具有挑战性的条件下(零人类演示、零策略预训练和零已知物体模型)在几分钟内从零开始学习机器人操作任务。仅给定物体的单视图和该物体的目标位姿,ZeroBot 使用图像到 3D 生成模型获得完整的物体网格,该网格用于仿真中进行大规模并行强化学习。为了加速训练,我们引入了一个动作空间,利用生成的几何结构和学习到的价值函数来采样涉及机器人-物体接触的状态。在包括抓取、推动、铰接物体交互和多阶段操作等真实世界任务上评估时,ZeroBot 实现了 87% 的成功率,平均训练时间为 119 秒。这些结果显示了在 real2sim 框架中使用图像到 3D 模型进行快速、自主机器人学习的价值。
cs.RO / 108 / 2609.34018

Estimate, Don't Imitate: Reusing Differentiable State-Based Policies for Visuomotor Control

估计,而非模仿:重用可微的基于状态的策略用于视觉运动控制
Shcherba, Denis, Abel, Adrian, Cobo-Briesewitz, Eckart, Samek, Wojciech, Toussaint, Marc
Abstract
Simulation-trained manipulation policies can exploit privileged state information to learn effective contact-rich behaviours, but deployment requires acting from partial observations such as noisy camera images. A common solution is teacher-student distillation, in which a visuomotor policy is trained to reproduce the actions of the privileged expert. This requires the student to jointly infer the task-relevant state and relearn the expert's action mapping that is already available. An alternative is to reuse the state-based expert and learn only a perceptual interface that reconstructs its missing state inputs. However, minimising the state estimate error alone does not necessarily minimise the downstream control error induced by these estimates. To bridge this gap, we train a visual state estimator using both direct state supervision and an action-consistency loss backpropagated through the frozen, differentiable expert. A scheduled objective first establishes a physically meaningful state estimate and progressively emphasises errors that affect the expert's actions. Across five goal-conditioned manipulation tasks, retaining the expert consistently outperforms direct pixel-to-action imitation from the same expert demonstration corpus. We further demonstrate sim-to-real transfer on a physical Panda robot, achieving 76% success without retraining the underlying expert.
Chinese Translation
仿真训练的操作策略可以利用特权状态信息来学习有效的接触丰富行为,但部署时需要从部分观测(如噪声相机图像)中进行操作。一个常见的解决方案是教师-学生蒸馏,其中视觉运动策略被训练来复现特权专家的动作。这要求学生联合推断任务相关状态,并重新学习已经可用的专家动作映射。另一种替代方案是重用基于状态的专家,仅学习一个感知接口来重建其缺失的状态输入。然而,仅最小化状态估计误差不一定能最小化由这些估计引起的下游控制误差。为了弥合这一差距,我们训练一个视觉状态估计器,使用直接状态监督和通过冻结的可微专家反向传播的动作一致性损失。一个分阶段的目标首先建立物理上有意义的状态估计,并逐步强调影响专家动作的误差。在五个目标条件操作任务中,保留专家始终优于从同一专家演示语料库进行的直接像素到动作模仿。我们进一步在物理Panda机器人上展示了仿真到现实迁移,在无需重新训练底层专家的情况下实现了76%的成功率。
cs.RO / 109 / 2609.34061

Quantile Head for Vision-Language-Action Models

用于视觉-语言-动作模型的分位数头(Quantile Head)
Wang, Xuan, Wu, Yinan, Duan, Haoran, Han, Jungong
Abstract
Vision-Language-Action (VLA) models integrate pretrained Vision-Language Models (VLMs) with action heads for robot control. Common action heads have distinct limitations: point regression provides only a point estimate of the action distribution, while standard flow-matching samplers require costly iterative sampling. To address these limitations, we unify regression and flow matching under a shared objective and extend it to derive a quantile objective. This quantile objective guides the design of our Quantile Head, which predicts a median and positive gaps to form ordered marginal action quantiles in one forward pass. These quantiles support multiple sampling strategies without retraining and are jointly supervised to train the default median policy. Our local analysis of this joint supervision shows that, with calibrated nearby quantiles, fixed gaps, and matched correction speed, direct median updates have lower variance than under median-only supervision. Experiments show that this jointly supervised median policy achieves the highest average success rates among the compared methods on LIBERO, LIBERO-Plus, LIBERO-Pro, and two real-robot tasks, together with the shortest mean episode time among matched LIBERO baselines; code is available at https://github.com/xwangrs/Quantile-Head-for-VLA.
Chinese Translation
视觉-语言-动作(VLA)模型将预训练的视觉-语言模型(VLMs)与动作头结合用于机器人控制。常见的动作头存在明显局限:点回归仅提供动作分布的点估计,而标准流匹配采样器需要昂贵的迭代采样。为了解决这些局限,我们在一个共享目标下统一了回归和流匹配,并扩展它推导出分位数目标。该分位数目标指导了我们的分位数头(Quantile Head)的设计,它预测一个中位数和正的间隔,以在一次前向传播中形成有序的边缘动作分位数。这些分位数支持多种采样策略而无需重新训练,并且被联合监督以训练默认的中位数策略。我们对这种联合监督的局部分析表明,在具有校准的邻近分位数、固定间隔和匹配的校正速度的情况下,直接中位数更新比在仅中位数监督下具有更低的方差。实验表明,这种联合监督的中位数策略在 LIBERO、LIBERO-Plus、LIBERO-Pro 和两个真实机器人任务上,在对比方法中取得了最高的平均成功率,并且在匹配的 LIBERO 基线中具有最短的平均回合时间;代码见 https://github.com/xwangrs/Quantile-Head-for-VLA。
cs.RO / 110 / 2609.34085

AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving

AD-E2E-JEPA:一种用于端到端自动驾驶的联合嵌入预测架构
Zhu, Haoran, Zhang, Wancong, LeCun, Yann, Choromanska, Anna
Abstract
Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM for end-to-end autonomous driving (E2EAD). To isolate world-model quality from policy learning, we employ a goal-conditioned zero-shot planning setting that evaluates these models using ground-truth future observations as goals, without training any driving policy. We find that existing JEPA-based world models are either accurate for driving but computationally expensive, or computationally efficient but insufficient for planning. To address this trade-off, we propose \textbf{AD-E2E-JEPA}, which introduces a SIGReg-regularized learnable projector applied to projected patch embeddings. The projector reduces the number of planning patches by $16\times$ and the embedding dimension by $4\times$, achieving a $100\times$ inference speedup while retaining planning performance, with a 0.8-second runtime for an 8-frame rollout over 256 candidate trajectories. \textit{Without} training any driving policy, the world model itself reaches the goals located 20 meters away on average within the displacement of respectively 4.0/2.8 meters, using world-model rollouts over trajectory vocabularies of respectively 256/8,192 candidates. On the NAVSIMv2 benchmark, it achieves 67.3/72.9 EPDMS with multiplicative safety metrics and 84.1/86.5 EPDMS$^{\dagger}$ without them in goal-conditioned zero-shot planning. Experiments further show that the self-supervised pretrained projector improves downstream imitation learning performance from 80.2 to 85.4 EPDMS. The source code is available at https://github.com/HaoranZhuExplorer/AD-E2E-JEPA
Chinese Translation
自动驾驶需要能够理解物理世界、进行推理和规划并安全运行的“世界模型”。在本文中,我们首先系统评估了现有的动作条件联合嵌入预测架构(JEPA)世界模型,包括LeWM、DINO-WM和JEPA-WM,用于端到端自动驾驶(E2EAD)。为了将世界模型质量与策略学习分离,我们采用了目标条件零样本规划设置,使用真实未来观测作为目标来评估这些模型,而不训练任何驾驶策略。我们发现现有的基于JEPA的世界模型要么对驾驶准确但计算成本高,要么计算高效但不足以进行规划。为了解决这种权衡,我们提出了AD-E2E-JEPA,它引入了一种SIGReg正则化的可学习投影器,应用于投影后的块嵌入。该投影器将规划块数量减少16倍,嵌入维度减少4倍,实现100倍的推理加速,同时保持规划性能,对于256个候选轨迹的8帧展开,运行时间为0.8秒。在不训练任何驾驶策略的情况下,世界模型本身在平均20米远的目标上分别达到4.0/2.8米的位移,使用世界模型在分别包含256/8,192个候选的轨迹词汇表上进行展开。在NAVSIMv2基准上,在目标条件零样本规划中,它分别实现了带有乘法安全指标的67.3/72.9 EPDMS和不带这些指标的84.1/86.5 EPDMS†。实验进一步表明,自监督预训练的投影器将下游模仿学习性能从80.2提高到85.4 EPDMS。源代码可在https://github.com/HaoranZhuExplorer/AD-E2E-JEPA获取。
cs.RO / 111 / 2609.34145

Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving

超越检索相关性:面向视觉语言驾驶的场景接地风险蕴涵
Liu, Jiaxin, Yu, Ruilin, Peng, Liang, Wang, Jingkai, Zhao, Chengxiang, Zhu, Zhenxin, Wang, Bing, Chen, Guang, Ye, Hangjun, Wang, Hong, Li, Jun
Abstract
Retrieval-augmented generation (RAG) gives vision--language driving systems access to external safety knowledge, yet a retrieved risk rule may be relevant without applying to the current scene. A vision--language model (VLM) receiving such knowledge must ground objects, bind entities across time, and verify relations before deciding how to act, leaving the support for risk conclusions implicit. We address this relevance--applicability gap with a Driving-Risk Knowledge Graph (DRKG) and Semantic Web Rule Language (SWRL) reasoning stage before VLM decision-making. Structured perception instantiates scene facts, from which SWRL rules derive events and directed risk relations when their antecedents are jointly satisfied. Recognized events, bound risk relations, and semantic descriptions of activated rules form compact evidence that conditions the VLM and diffusion planner. In matched comparisons on nuReasoning, our method improved the nuReasoning planning score (NPS) by 1.30 points and the non-at-fault collision score (NC) by 2.76 points over the relevance retrieval-based baseline. These gains indicate that scene-applicable risk evidence improves safety-weighted planning relative to semantically retrieved risk knowledge.
Chinese Translation
检索增强生成(RAG)为视觉语言驾驶系统提供了访问外部安全知识的途径,然而检索到的风险规则可能具有相关性,却并不适用于当前场景。接收此类知识的视觉语言模型(VLM)必须对物体进行接地,跨时间绑定实体,并在决定如何行动之前验证关系,这使得风险结论的支持依据变得隐含。我们通过一个驾驶风险知识图谱(DRKG)和语义网规则语言(SWRL)推理阶段,在VLM决策之前解决这一相关性-适用性差距。结构化感知实例化场景事实,当SWRL规则的前件联合满足时,从中推导出事件和有向风险关系。识别出的事件、绑定的风险关系以及激活规则的语义描述形成了紧凑的证据,用于条件化VLM和扩散规划器。在nuReasoning上的匹配比较中,我们的方法相比基于相关性检索的基线,将nuReasoning规划得分(NPS)提高了1.30分,将非责任碰撞得分(NC)提高了2.76分。这些增益表明,相对于语义检索的风险知识,场景适用的风险证据改善了安全加权规划。
cs.RO / 112 / 2609.34163

Reliability-Aware Sparse Route Memory for Round-Trip Vision-Language Navigation

面向往返视觉语言导航的可靠性感知稀疏路径记忆
Long, Bojun, Bao, Lingfan, Peng, Tianhu, Sun, Jingcheng, Zhou, Chengxu
Abstract
Vision-language navigation (VLN) is typically evaluated as a one-way task, although deployed robots may need to return after reaching a goal. We study continuous round-trip VLN and diagnose failures in directional observability, deviation recovery, and termination stability. We propose a reliability-aware sparse route memory that records the executed Outbound trajectory as ordered geometric anchors and queries them in reverse through a structured hint, action-level arbitration, and terminal verification. On 50 reverse-paired episodes using NaVILA and a simulated Unitree Go2, language-only Return succeeds in 22.0% of episodes, while our online system reaches 55.1%. With exact route information, the same interfaces achieve 86.0%, showing that effective Return requires both accurate information and consistent action on that information. The remaining online gap arises mainly from geometric evidence that is too unreliable to authorise intervention. These results distinguish information quality, behavioural consistency, and online reliability as separate limits in long-horizon navigation.
Chinese Translation
视觉语言导航(VLN)通常被评估为单向任务,尽管部署的机器人可能需要在到达目标后返回。我们研究连续往返 VLN,并诊断了方向可观测性、偏差恢复和终止稳定性方面的失败。我们提出了一种可靠性感知的稀疏路径记忆,它记录已执行的去程轨迹为有序几何锚点,并通过结构化提示、动作级仲裁和终端验证以相反顺序查询它们。在使用 NaVILA 和模拟 Unitree Go2 的 50 个反向配对回合中,仅语言的返程在 22.0% 的回合中成功,而我们的在线系统达到 55.1%。在精确路径信息下,相同的接口达到 86.0%,表明有效的返程既需要准确的信息,也需要对该信息采取一致的行动。剩余的在线差距主要源于几何证据过于不可靠,无法授权干预。这些结果将信息质量、行为一致性和在线可靠性区分为长时程导航中的独立限制。
cs.RO / 113 / 2609.34170

RAVEL: Asynchronous Rolling Inference for Flow-Based Vision-Language-Action Models

RAVEL:面向基于流的视觉-语言-动作模型的异步滚动推理
Chen, Yuhan, Yu, Ke, Liu, Pengfei, Wang, Shuxun, Yang, Yi, Zhu, Linchao
Abstract
Flow-based vision-language-action (VLA) models are highly effective for generalist robot manipulation, yet their reliance on computationally expensive VLM encoding and multi-step iterative action generation imposes a significant latency bottleneck. The resulting inference latency makes it difficult for robots to respond quickly, especially in dynamic environments. We address this limitation with RAVEL (Rolling Asynchronous VLA Enabling Low-Latency Control), an asynchronous inference framework that addresses the computational bottlenecks of both the VLM backbone and the action expert. To reduce the delay from multi-step action denoising, RAVEL allows near-term actions to be executed after a single denoising step by carrying partially denoised future actions forward in a rolling buffer. To avoid blocking on slow VLM encoding, RAVEL decouples VLM encoding from rolling action generation, allowing the action expert to operate continuously using the latest available VLM context, while a lightweight Fast Observation Pathway (FOP) directly conditions the action expert on current observations. Across simulated and real-world manipulation tasks, RAVEL consistently achieves substantially lower response latency while maintaining the task capability of the underlying VLA, enabling high-frequency and responsive closed-loop control.
Chinese Translation
基于流的视觉-语言-动作(VLA)模型在通用机器人操作中非常有效,然而它们依赖计算成本高昂的VLM编码和多步迭代动作生成,这造成了显著的延迟瓶颈。由此产生的推理延迟使机器人难以快速响应,尤其是在动态环境中。我们通过RAVEL(Rolling Asynchronous VLA Enabling Low-Latency Control)来解决这一限制,这是一个异步推理框架,解决了VLM主干网络和动作专家两者的计算瓶颈。为了减少多步动作去噪带来的延迟,RAVEL允许在单步去噪后执行近期动作,方法是在滚动缓冲区中向前携带部分去噪的未来动作。为了避免在缓慢的VLM编码上阻塞,RAVEL将VLM编码与滚动动作生成解耦,允许动作专家使用最新可用的VLM上下文连续运行,同时一个轻量级的快速观察通路(FOP)直接以当前观察为条件调节动作专家。在仿真和真实世界的操作任务中,RAVEL始终实现显著更低的响应延迟,同时保持底层VLA的任务能力,从而实现高频且响应迅速的闭环控制。
cs.RO / 114 / 2609.34175

FailPatch: Failure Residual Patching for Vision-Language-Action Models

FailPatch:面向视觉-语言-动作模型的失败残差修补
Yu, Peng, Wang, Jiacheng, Zhang, Ziheng, Zhang, Xuchong, Li, Baoting, Yu, Zhuoyuan, Chen, Yuxiang, Wang, Tiancai, Sun, Hongbin
Abstract
Vision-Language-Action (VLA) policies are typically adapted using successful demonstrations, which provide direct action supervision but rarely cover failure-prone states. Deployment failures expose these states, yet lack the corrective actions needed for conventional supervised learning. We propose FailPatch, a failure-driven residual patching framework that decouples action supervision from execution-reliability supervision. Successful demonstrations ground how the policy should act, while deployment trajectories indicate when its behavior becomes unreliable. We further observe that action hidden representations exhibit clear linear separability between reliable and failure-associated states while directly conditioning action generation. Building on these insights, FailPatch introduces a Null-gated Residual Expert Bank into the action hidden space of a frozen VLA policy. A unified Preserve--Redirect--Trust objective retains the original policy in reliable states, selects residual experts in failure-associated states and redirects representations from failure regions toward success-associated regions under bounded intervention. With only 0.52% trainable parameters, FailPatch improves success rates by 11.0 percentage points on four long-horizon RoboTwin tasks under clean evaluation, 9.5 percentage points under clean-to-random generalization, and 16.7 percentage points over the baseline across three real-world tasks. Project and code: https://github.com/yupeng-2003/FailPatch.
Chinese Translation
视觉-语言-动作(VLA)策略通常利用成功演示进行适配,这些演示提供了直接的动作监督,但很少覆盖易失败状态。部署失败暴露了这些状态,但缺乏传统监督学习所需的纠正动作。我们提出了FailPatch,一个失败驱动的残差修补框架,将动作监督与执行可靠性监督解耦。成功演示奠定了策略应如何行动的基础,而部署轨迹则指示了其行为何时变得不可靠。我们进一步观察到,动作隐藏表示在可靠状态和失败相关状态之间表现出清晰的线性可分性,同时直接调节动作生成。基于这些见解,FailPatch在冻结的VLA策略的动作隐藏空间中引入了Null-gated残差专家库。一个统一的保留—重定向—信任(Preserve--Redirect--Trust)目标在可靠状态下保留原始策略,在失败相关状态下选择残差专家,并在有界干预下将表示从失败区域重定向到成功相关区域。仅需0.52%的可训练参数,FailPatch在四个长时域RoboTwin任务上,干净评估下成功率提升11.0个百分点,干净到随机的泛化下提升9.5个百分点,并在三个真实世界任务上相比基线提升16.7个百分点。项目与代码:https://github.com/yupeng-2003/FailPatch。
cs.RO / 115 / 2609.34182

Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation

面向灵巧操作的人类演示统一视觉-触觉-动作建模
Li, Wenqiao, Zhao, Qianyou, Hao, Jiawen, Zhu, Xuezhou, Liu, Tengyu, Zhang, Kaifeng, Wen, Chuan, Huang, Siyuan
Abstract
Dexterous manipulation requires tactile feedback.However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated by a simple premise: hands can change, but the underlying physics of interaction does not. We leverage human tactile data to improve dexterous manipulation policies. Specifically, we first build a tactile motion-capture system that synchronously records images, tactile signals, and hand motions. Using this system, we construct the UVTA dataset spanning five contact-rich tasks, with 1,000 human demonstrations covering diverse interaction patterns and 150 robot demonstrations per task. To transfer the underlying physics of human interaction to robot control, we propose a Unified Visual-Tactile-Action Model that maps both embodiments into aligned tactile and action representations and jointly predicts future action and tactile trajectories. The joint objective enables human demonstrations to supervise contact-aware representation learning, while only robot actions are executed during deployment. In real-robot evaluations across five tasks, our method achieves an average success rate of 70%, outperforming the strongest visual-tactile baseline, which achieves 29%, and an architecture ablation, which achieves 42%. Performance improves consistently with additional human demonstrations and exhibits no saturation at 1,000 demonstrations per task, validating the effectiveness of scalable human tactile data for dexterous manipulation. Project page is available at https://uni-vta.github.io/.
Chinese Translation
灵巧操作需要触觉反馈。然而,机器人触觉演示难以规模化,因为灵巧手遥操作向操作者提供的触觉反馈有限。相比之下,人类演示为多样化的触觉交互提供了更具可扩展性的来源。受一个简单前提的启发:手可以改变,但交互的基本物理规律不变。我们利用人类触觉数据来改进灵巧操作策略。具体来说,我们首先构建了一个触觉运动捕捉系统,同步记录图像、触觉信号和手部运动。利用该系统,我们构建了UVTA数据集,涵盖五个接触丰富的任务,包含1,000个人类演示,覆盖多样的交互模式,以及每个任务150个机器人演示。为了将人类交互的基本物理规律迁移到机器人控制,我们提出了一种统一视觉-触觉-动作模型,将两种具身映射到对齐的触觉和动作表示,并联合预测未来动作和触觉轨迹。联合目标使得人类演示能够监督接触感知的表示学习,而部署期间仅执行机器人动作。在五个任务的真实机器人评估中,我们的方法平均成功率达到70%,优于最强的视觉-触觉基线(29%)和架构消融(42%)。性能随着额外人类演示的增加而持续提升,并且在每个任务1,000个演示时未出现饱和,验证了可扩展的人类触觉数据对灵巧操作的有效性。项目页面见 https://uni-vta.github.io/。
cs.RO / 116 / 2609.34199

WB-WAM: Heterogeneous Body-Hand Pre-training for Humanoid Loco-Manipulation

WB-WAM:面向人形机器人移动操作的异构身体-手部预训练
Qin, Chuan, Zhu, Shaoting, Luo, Siyuan, Huang, Siqiao, Zhao, Hongyu, Zhao, Hang
Abstract
Humanoid loco-manipulation demands coordinated body and hand behavior, while conventional robot pre-training data provide limited coverage of such whole-body motion. We present WB-WAM, a World Action Model that incorporates explicit whole-body action supervision into generative video pre-training. A shared physical action space integrates body, root, and dexterous hand annotations from heterogeneous sources, enabling joint video and action learning from 1880.2 hours of partially annotated video and motion data. The resulting priors are refined through PICO mid-training and adapted to robot tasks with auxiliary forward kinematics supervision. We construct WB-Datasets to support these stages with retargeted egocentric human demonstrations and robot trajectories, allowing task-aligned human motion to supplement limited robot data. Evaluations in simulation demonstrate strong whole-body task performance with 81.9% in HumanoidArena, while real-world experiments further validate WB-WAM with 84.0% mean success across five tasks. Moreover, task-aligned PICO mid-training improves downstream task performance while reducing the need for real-robot demonstrations. These results support heterogeneous whole-body pre-training and human motion transfer as a practical route to data-efficient humanoid loco-manipulation.
Chinese Translation
人形机器人移动操作需要身体和手部的协调行为,而传统的机器人预训练数据对此类全身运动的覆盖有限。我们提出了WB-WAM,一种将显式全身动作监督融入生成式视频预训练的世界动作模型。一个共享的物理动作空间整合了来自异构来源的身体、根部和灵巧手部标注,从而能够从1880.2小时的部分标注视频和运动数据中进行联合视频和动作学习。得到的先验通过PICO中间训练进行细化,并利用辅助前向运动学监督适应机器人任务。我们构建了WB-Datasets,通过重定向的以自我为中心的人类演示和机器人轨迹来支持这些阶段,使得任务对齐的人类运动能够补充有限的机器人数据。仿真评估展示了强大的全身任务性能,在HumanoidArena中达到81.9%,而真实世界实验进一步验证了WB-WAM,在五项任务中平均成功率为84.0%。此外,任务对齐的PICO中间训练提高了下游任务性能,同时减少了对真实机器人演示的需求。这些结果支持异构全身预训练和人类运动迁移作为实现数据高效的人形机器人移动操作的实用途径。
cs.RO / 117 / 2609.34210

RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers

RLE-Bench:编码智能体作为机器人学习工程师的资格考试
Ma, Haitong, Gao, Chenxiao, Qiang, Rushi, Li, Na, Dai, Bo
Abstract
Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agents' broader engineering capabilities. Real-world robotics extends beyond control: agents must build, integrate, diagnose, and improve heterogeneous artifacts under resource constraints and reason from multimodal feedback. To evaluate these broader capabilities, we introduce RLE-Bench, a benchmark of robot-learning tasks spanning four representative robotics development workflows: interactive control, policy learning, perception and estimation, and mechanical design. We use diverse task-specific metrics to evaluate the artifacts submitted by the coding agents, from the success rate the agents achieved to the policy agents trained, the harness agent built, and the mechanical structures the agent designed. We aggregate these metrics into an overall RLE Index and report workflow-specific capability profiles, enabling systematic comparison of coding agents' capabilities across multiple capability dimensions. Beyond performance ranks, we also conduct in-depth case studies examining agent behavior on representative tasks, highlighting both current capabilities and limitations, and pointing to the opportunities robotics tasks have to offer for future agent training.
Chinese Translation
编码智能体正开始超越纯数字任务,去应对物理世界的挑战,尤其是在机器人领域。然而,现有的机器人基准主要关注单个制品的性能,例如策略或控制器,对编码智能体更广泛的工程能力的覆盖有限。真实世界的机器人技术超越了控制:智能体必须在资源约束下构建、集成、诊断和改进异构制品,并基于多模态反馈进行推理。为了评估这些更广泛的能力,我们提出了 RLE-Bench,一个包含机器人学习任务的基准,涵盖四种代表性的机器人开发工作流:交互控制、策略学习、感知与估计以及机械设计。我们使用多样的任务特定指标来评估编码智能体提交的制品,从智能体达到的成功率,到训练出的策略智能体、构建的测试框架智能体以及智能体设计的机械结构。我们将这些指标聚合为一个总体的 RLE 指数,并报告特定工作流的能力概况,从而能够在多个能力维度上对编码智能体的能力进行系统比较。除了性能排名之外,我们还进行了深入的案例研究,考察智能体在代表性任务上的行为,突出当前的能力和局限性,并指出机器人任务为未来智能体训练提供的机会。
cs.RO / 118 / 2609.34220

mmHRI: Towards Privacy-Preserving Human-Robot Interaction with Millimeter-Wave Radar

mmHRI:迈向基于毫米波雷达的隐私保护人机交互
Fan, Junqiao, Hu, Yuxuan, Lyu, Bofan, Lu, Yanshuo, Liu, Pengfei, Zhang, Jiarui, Ding, Fangqiang, Xie, Lihua, Li, Gen, Yang, Jianfei
Abstract
Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as object delivery. However, most existing HRI systems rely on RGB cameras that continuously observe humans to respond to non-verbal commands, such as hand gestures. This raises privacy concerns in privacy- critical environments, such as hospital wards or restaurants, where direct camera observation of humans is restricted. To develop privacy-preserving HRI, we leverage millimeter-wave (mmWave) radar, which can sense human motion through privacy barriers without identifiable imagery. We propose mmHRI, the first multi-modal robot manipulation framework that achieves mmWave radar-guided privacy-preserving HRI. mmHRI introduces two key designs to mitigate the sparsity and temporal inconsistency of radar data in cluttered robot manipulation environments. First, we propose a dual-stream architecture that jointly learns from unfiltered raw radar tensors and radar point clouds to estimate both human actions and 3D poses. To mitigate signal inconsistency, mmHRI further incorporates a memory-based state-space model (MSSM) that retains historical radar features to reduce abrupt changes in pose/action. These estimated human states are then converted into structured textual robot instructions, which control a vision-language-action (VLA) policy for closed-loop robot manipulation and human-aware reactions. Our evaluation covers human action recognition and closed-loop delivery and retrieval. In the privacy-preserving curtain setting, mmHRI achieves 85.09% action-recognition accuracy, outperforming existing radar-based alternatives. Robot trials further demonstrate successful delivery and retrieval under visual occlusion, with stable task performance across unseen subjects, clutter configurations, and environments.
Chinese Translation
辅助机器人越来越多地在许多以人为中心的环境中运行,并执行各种人机交互(HRI)任务,例如物体递送。然而,大多数现有HRI系统依赖RGB相机持续观察人类,以响应非语言命令,例如手势。这在隐私敏感环境(如医院病房或餐厅)中引发了隐私担忧,在这些环境中,直接使用相机观察人类受到限制。为了开发隐私保护的HRI,我们利用毫米波(mmWave)雷达,它可以通过隐私屏障感知人体运动,而无需可识别的图像。我们提出了mmHRI,这是首个实现毫米波雷达引导的隐私保护HRI的多模态机器人操作框架。mmHRI引入了两个关键设计,以缓解杂乱机器人操作环境中雷达数据的稀疏性和时间不一致性。首先,我们提出了一种双流架构,该架构从未滤波的原始雷达张量和雷达点云中联合学习,以估计人体动作和3D姿态。为了缓解信号不一致性,mmHRI进一步集成了一个基于记忆的状态空间模型(MSSM),该模型保留历史雷达特征,以减少姿态/动作的突变。然后,这些估计的人体状态被转换为结构化的文本机器人指令,这些指令控制视觉-语言-动作(VLA)策略,以实现闭环机器人操作和人类感知反应。我们的评估涵盖了人体动作识别以及闭环递送和取回。在隐私保护帘幕设置中,mmHRI达到了85.09%的动作识别准确率,优于现有的基于雷达的替代方案。机器人试验进一步证明,在视觉遮挡下能够成功完成递送和取回,并在未见过的受试者、杂乱配置和环境中保持稳定的任务性能。
cs.RO / 119 / 2609.34222

Proprioceptive Force Estimation for Quadruped Locomotion and Human-Robot Interaction

用于四足运动和人机交互的本体感知力估计
Wang, Run, Yang, Xu, Tuerxun, Alapati, Mo, Yilin
Abstract
Payload forces must be accommodated during locomotion, while leash forces can specify desired motion. We investigate whether a shared three-dimensional force estimate in newtons, inferred from proprioceptive history under sustained loading, can support both tasks. An estimator and locomotion policy are jointly trained with supervised force and velocity outputs and learned latent context. The estimated force conditions locomotion and additionally generates planar-velocity and yaw-rate commands for leash guidance through an analytical map. In sustained-force simulation sweeps, temporal means of componentwise force root mean square error range from 1.44 to 2.83\,N. Compared with a domain-randomized baseline, the framework reduces velocity-tracking and base-orientation error scores by 21.6\% and 46.5\%, respectively, and increases mean survival from 68.29\% to 94.60\% in separate sustained-force tests. Unitree Go1 experiments demonstrate stationary vertical and horizontal force estimation, locomotion with an 8.5\,kg payload whose weight exceeds the 70\,N training force limit, and leash guidance using the same force-estimation interface.
Chinese Translation
在运动过程中必须适应负载力,而牵引力可以指定期望的运动。我们研究从持续负载下的本体感知历史推断出的、以牛顿为单位的共享三维力估计能否支持这两项任务。一个估计器和运动策略使用监督的力和速度输出以及学习到的潜在上下文进行联合训练。估计的力为运动提供条件,并通过解析映射额外生成用于牵引引导的平面速度和偏航率命令。在持续力仿真扫描中,各分量力均方根误差的时间均值范围为1.44至2.83 N。与域随机化基线相比,该框架在单独的持续力测试中将速度跟踪和基座姿态误差分数分别降低了21.6%和46.5%,并将平均存活率从68.29%提高到94.60%。Unitree Go1实验展示了静态垂直和水平力估计、携带8.5 kg负载(其重量超过70 N训练力限制)的运动,以及使用相同力估计接口的牵引引导。
cs.RO / 120 / 2609.34233

GAE: General Action Expert for Real-Time Humanoid Teleoperation

GAE:面向实时人形遥操作的通用动作专家
Wang, Yuefan, Zhou, Huaicheng, He, Xiao, He, Zhijie, Yang, Mingchuan, Zhang, Huayi, Chai, Li, Liu, Jinxin, Wang, Donglin
Abstract
Humanoid avatars extend human physical presence beyond the body, enabling people to participate in social, service, and labor activities through remotely operated robots. This requires teleoperation systems capable of realizing diverse and dynamic whole-body behaviors while maintaining responsive human-robot synchronization. We present General Action Expert(GAE), a unified learning framework for general-purpose, low-latency humanoid whole-body teleoperation. To cover diverse human behaviors, GAE builds a large-scale human motion dataset from heterogeneous sources, including videos, animations, and motion capture, followed by standardization and augmentation. GAE then addresses the noise and embodiment mismatch in human motions with a two-stage training paradigm: a privileged generator policy first tracks human motion references in simulation and rolls out feasible humanoid trajectories; a deployable executor policy then learns to track these generated trajectories under curriculum domain randomization. For responsive human-robot synchronization, GAE introduces a latency-conditioned anticipation mechanism that adaptively compensates for end-to-end delay during real-time teleoperation. Simulation and real-world experiments on Unitree G1 and Westlake O1 robots demonstrate that GAE enables humanoids to smoothly mirror diverse, agile, and expressive human behaviors. Project website: https://wangyf0928.github.io/gae-wlrobotics/
Chinese Translation
人形化身将人类的物理存在延伸至身体之外,使人们能够通过远程操作的机器人参与社交、服务和劳动活动。这需要遥操作系统能够实现多样化和动态的全身行为,同时保持响应迅速的人机同步。我们提出了通用动作专家(GAE),一个用于通用、低延迟人形全身遥操作的统一学习框架。为了涵盖多样的人类行为,GAE从异构来源(包括视频、动画和动作捕捉)构建了一个大规模人体运动数据集,并进行了标准化和增强。然后,GAE通过两阶段训练范式解决了人体运动中的噪声和具身不匹配问题:首先,一个特权生成器策略在仿真中跟踪人体运动参考并生成可行的人形轨迹;然后,一个可部署的执行器策略在课程域随机化下学习跟踪这些生成的轨迹。为了实现响应迅速的人机同步,GAE引入了一种延迟条件预测机制,能够在实时遥操作中自适应地补偿端到端延迟。在Unitree G1和Westlake O1机器人上的仿真和真实世界实验表明,GAE使人形机器人能够平滑地模仿多样、敏捷和富有表现力的人类行为。项目网站:https://wangyf0928.github.io/gae-wlrobotics/
cs.RO / 121 / 2609.34250

WAM-OPD: Sharpening World Action Models via On-Policy Distillation

WAM-OPD:通过同策略蒸馏提升世界动作模型
Liu, Panjun, Lei, Xiaohan, Zhang, Shiqi, Wang, Yikun, Zhang, Yongxin, Hu, Mingyi, Sun, Shida, Shou, Jiateng, Zhou, Wengang, Deng, Jiajun, Xiong, Zhiwei
Abstract
Pretrained world action models (WAMs) provide generalist capabilities across diverse robotic manipulation tasks, yet improving target-task performance to an expert level without degrading pretrained skills remains challenging. We explore on-policy distillation (OPD) for WAMs and introduce WAM-OPD. WAM-OPD inherits the advantage of OPD methods that transfer task-specific teacher knowledge under the student's own induced distribution, rather than directly fitting the student to a narrow task-specific data distribution. However, in closed-loop manipulation, the observation histories change as the student policy evolves, requiring fresh environment rollouts to remain on-policy. Applying OPD to WAMs entails repeated data collection, which is costly even in simulation and often impractical on real robots. To avoid repeated environment rollouts during distillation, we introduce prefix-weighted trajectory replay (PWTR). PWTR uses a fixed trajectory pool composed primarily of initial-student rollouts, supplemented with task-specific teacher rollouts to broaden trajectory coverage. For each trajectory replayed from this pool, PWTR conditions the current policy on successive stored histories to generate fresh denoising paths, along which the task-specific teacher provides supervision. Although these denoising paths are refreshed as the policy evolves, the replayed environment trajectories remain fixed. PWTR therefore reweights per-decision distillation losses using proxy importance weights derived from path scores accumulated over the trajectory prefix preceding each decision to mitigate the resulting shift in the history distribution. Simulated and real-world experiments demonstrate task adaptation without additional environment interaction during distillation. In both settings, WAM-OPD improves target-task performance while retaining near-initial performance on tasks excluded from adaptation.
Chinese Translation
预训练的世界动作模型(WAMs)为多样化的机器人操作任务提供了通用能力,然而在不降低预训练技能的前提下将目标任务性能提升至专家水平仍然具有挑战性。我们探索了面向WAMs的同策略蒸馏(OPD),并提出了WAM-OPD。WAM-OPD继承了OPD方法的优势,即在学生模型自身诱导的分布下迁移任务特定的教师知识,而非直接将学生模型拟合到狭窄的任务特定数据分布。然而,在闭环操作中,观测历史随着学生策略的演化而变化,需要新鲜的环境推演以保持同策略。将OPD应用于WAMs需要进行重复的数据收集,这在仿真中成本高昂,在真实机器人上往往不切实际。为避免蒸馏过程中重复的环境推演,我们引入了前缀加权轨迹回放(PWTR)。PWTR使用由初始学生模型推演组成的固定轨迹池,并辅以任务特定教师的推演以扩大轨迹覆盖范围。对于从该池中回放的每条轨迹,PWTR以连续存储的历史为条件,使当前策略生成新的去噪路径,任务特定的教师沿这些路径提供监督。尽管这些去噪路径随着策略的演化而更新,但回放的环境轨迹保持不变。因此,PWTR使用代理重要性权重对每个决策的蒸馏损失进行重加权,这些权重源自每个决策之前轨迹前缀上累积的路径分数,以缓解由此产生的历史分布偏移。仿真和真实世界实验表明,在蒸馏过程中无需额外的环境交互即可实现任务适应。在两种设置中,WAM-OPD均提升了目标任务性能,同时在未纳入适应的任务上保持了接近初始的性能。
cs.RO / 122 / 2609.34256

UMR: Universal Manipulation Representation

UMR:通用操作表示
Liu, Song, Li, Linyi, Zhao, Yanshun, Li, Rxuan, Xu, Xinrui, Ju, Yi, Deng, Yahui, Zhang, Senge, Liu, Guoyu, Li, Yixuan, Zhang, Wuyang, Li, Yao, Zhu, Congcong, Chen, Jingrun
Abstract
General-purpose embodied manipulation hinges on a unified action representation that generalizes across embodiments and scales readily. Yet existing policies rely on embodiment-specific action spaces, making cross-embodiment demonstrations difficult to leverage at scale and limiting transfer to new embodiments and spatial variations. To this end, we introduce Universal Manipulation Representation (UMR), a unified action representation that enables zero-shot skill transfer from human demonstrations to heterogeneous robots. UMR decomposes manipulation into two functionally distinct yet geometrically linked components: embodiment-agnostic World Flow, which describes task-relevant object motion in the world frame, and Ego Trajectory, which represents end-effector motion relative to the current pose. We instantiate UMR as World--Ego Point VLA (WEPVLA), a compact 0.5B-parameter policy that learns in the unified geometric action space through a dual-stream Point Action Adapter and a unified Point Action Expert, with an $SE(3)$ conjugation coupling the two components. To improve data efficiency, we complement UMR with a Data-Efficient Strategy (DES) that diversifies object configurations through stage-aware point-cloud editing while preserving demonstrated contact geometry. In simulation, WEPVLA achieves average success rates of 97.5\% on LIBERO and 85.7\% on the 10-task RLBench benchmark. In real-world experiments, a single policy trained on human demonstrations augmented by DES transfers zero-shot to diverse deployment conditions. With about 10 minutes of collected human demonstrations per task and no robot demonstrations, it achieves 91.7\% average success across six evaluation settings, compared with 60.8\% for HumanEgo. Code and additional materials are available at https://umr-wepvla.github.io/.
Chinese Translation
通用具身操作依赖于一种统一的动作表示,它能够跨具身泛化并易于扩展。然而,现有的策略依赖于特定于具身的动作空间,使得跨具身演示难以大规模利用,并限制了向新具身和空间变化的迁移。为此,我们引入了通用操作表示(UMR),这是一种统一的动作表示,能够实现从人类演示到异构机器人的零样本技能迁移。UMR将操作分解为两个功能不同但几何相关的组件:与具身无关的世界流(World Flow),它描述世界坐标系中与任务相关的物体运动;以及自我轨迹(Ego Trajectory),它表示相对于当前位姿的末端执行器运动。我们将UMR实例化为World--Ego Point VLA(WEPVLA),一个紧凑的0.5B参数策略,它通过双流点动作适配器(dual-stream Point Action Adapter)和统一的点动作专家(unified Point Action Expert)在统一的几何动作空间中学习,并使用$SE(3)$共轭耦合这两个组件。为了提高数据效率,我们为UMR补充了一种数据高效策略(DES),它通过阶段感知的点云编辑来多样化物体配置,同时保留演示的接触几何。在仿真中,WEPVLA在LIBERO上达到了97.5%的平均成功率,在10项任务的RLBench基准上达到了85.7%。在真实世界实验中,一个仅在DES增强的人类演示上训练的策略能够零样本迁移到多样的部署条件。在每个任务收集约10分钟的人类演示且无机器人演示的情况下,它在六个评估设置中达到了91.7%的平均成功率,而HumanEgo为60.8%。代码和附加材料可在https://umr-wepvla.github.io/获取。
cs.RO / 123 / 2609.34261

RoboICL: Embodied In-Context Learning with GPT-6 Astra

RoboICL:基于GPT-6 Astra的具身上下文学习
Liu, Fangcheng, Shen, Yeqing, Cheng, Anda, Mi, Weishi, Tang, Chao, Liu, Chenyuan, Xiang, Yushun, Li, Tingguang, Li, Yong-Lu, Tang, Yehui
Abstract
General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that narrows these gaps without robot-specific parameter updates or a learned VLA. RoboICL separates \emph{demonstration context}, which provides recorded examples when available, from \emph{interaction memory}, which accumulates the model's own actions and observed outcomes. Both use a shared observation--action--receipt--observation grammar. To preserve experience across task stages, RoboICL combines sampled demonstration blocks with bounded anchored memory. Fixed anchors keep earlier rollout interactions available for in-context learning, while the latest interaction supports immediate error correction. Across 30 RoboDojo tasks, using zero shot for Open and one demonstration elsewhere, RoboICL improves on official zero-shot \gptastra{} by 20--27 progress-score points in every category. It leads the leaderboard baselines on Memory and Open, achieves comparable performance to the strongest Precision baseline, and remains competitive on Long-Horizon. Its 30-task Overall score is 50.64, versus 33.68 for the strongest baseline. On a separate ten-task subset, RoboICL scores 60.60, within 2.00 points of the $\pi_{0.5}$ + \gptastra{} hybrid approach. On three real-robot tasks, mean progress rises from 14.45 at zero shot to 63.33 at one shot and 78.89 at three shots. On two development tasks, optional Jev-gated action reuse reduces \gptastra{} calls by 33--48\%. Code is available at \href{https://github.com/Mosi-AI/RoboICL}{https://github.com/Mosi-AI/RoboICL}.
Chinese Translation
通用视觉语言模型为零样本机器人控制提供了一种有前景的方法:GPT-6 Astra 擅长开放式以及语言或图像条件化的操作,但在高精度和长时程任务上仍然明显较弱。我们提出 RoboICL,一个上下文机器人控制框架,无需机器人特定的参数更新或学习到的 VLA,缩小了这些差距。RoboICL 将演示上下文(在可用时提供记录示例)与交互记忆(累积模型自身的动作和观察到的结果)分离。两者都使用共享的观察-动作-回执-观察语法。为了在任务阶段之间保留经验,RoboICL 将采样的演示块与有界锚定记忆相结合。固定锚点保持早期 rollout 交互可用于上下文学习,而最新交互支持即时错误纠正。在 30 个 RoboDojo 任务上,对于 Open 任务使用零样本,其他任务使用一个演示,RoboICL 在每个类别上比官方零样本 GPT-6 Astra 提高了 20-27 个进度分数点。它在 Memory 和 Open 上领先排行榜基线,达到与最强 Precision 基线相当的性能,并在 Long-Horizon 上保持竞争力。其 30 任务总体得分为 50.64,而最强基线为 33.68。在另一个十任务子集上,RoboICL 得分为 60.60,与 π_0.5 + GPT-6 Astra 混合方法相差 2.00 分以内。在三个真实机器人任务上,平均进度从零样本的 14.45 上升到单样本的 63.33 和三次样本的 78.89。在两个开发任务上,可选的 Jev 门控动作重用将 GPT-6 Astra 调用减少了 33-48%。代码可在 https://github.com/Mosi-AI/RoboICL 获取。
cs.RO / 124 / 2609.34268

SAGE: Symbolic Action-Gating and Editing for LLM Task Planners

SAGE:面向 LLM 任务规划器的符号动作门控与编辑
Bui, Trung Minh, Moon, JongSul, Kim, YoungOuk, Phung, Quang-Ngoc, Jun, Se-Woong, Shin, Dongin
Abstract
Large language models (LLMs) are now the default cognitive core of embodied household agents, yet the plans they emit are rarely checked against a grounded model of the environment before execution, and the task-success they report is often measured on benchmarks so saturated that no method can be separated from another. We present SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate (~250 lines of Python, zero tokens, $O(|\pi|)$) that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and untouched work intact; a hybrid seed+live memory store supports cold-start coverage. We evaluate under a leak-free protocol (leave-one-out retrieval) over five open-weight models and a 75-task AI2-THOR benchmark. On the standard benchmark goal-completeness saturates (52% of instances trivially solved) and SAGE ties strong hierarchical baselines. On a harder, method-agnostic multi-goal composition, SAGE's completeness lead re-emerges large (+0.06 to +0.23 across four models). Under injected mid-execution failures, SAGE recovers as reliably as whole-plan replanners at 2.4-3.3x fewer LLM calls. As a verify-before-execute gate, the symbolic monitor blocks unsafe actions before actuation and raises simulator-reported step-success for every planner tested (up to +0.11), a signal the verifier never sees (non-circular). Because the gate calls no model (0.008 ms/plan), it is a safety layer that runs essentially free on the edge: SAGE planning reproduces its quality on a Jetson AGX Orin, where small-model verification helps most. We release the benchmark, the leak-free protocol, the recovery and safety-gate harnesses, and a verifier-portability study (auto-induced on ALFWorld, 0.89 held-out).
Chinese Translation
大语言模型(LLMs)现已成为具身家庭智能体的默认认知核心,然而其输出的规划在执行前很少根据环境的接地模型进行检查,而且其所报告的任务成功率常常是在饱和的基准上度量,以至于无法区分不同方法。我们提出 SAGE(符号动作门控与编辑),一种单 LLM 规划器,由两个轻量级机制构建:一个领域无关的符号门控(约 250 行 Python,零 token,$O(|\pi|)$),作为运行时安全监视器,以类型化原因阻止违反前置条件的动作;以及一个局部编辑,仅重新生成失败子目标的后缀,保持已完成和未触及的工作不变;混合的种子+实时记忆存储支持冷启动覆盖。我们在无泄漏协议(留一检索)下,对五个开放权重模型和 75 个任务的 AI2-THOR 基准进行评估。在标准基准上,目标完整性已饱和(52% 的实例被平凡地解决),SAGE 与强分层基线持平。在更困难、方法无关的多目标组合上,SAGE 的完整性优势再次显著(在四个模型上 +0.06 到 +0.23)。在注入执行中失败的情况下,SAGE 的恢复与全规划重规划器一样可靠,而 LLM 调用次数减少 2.4-3.3 倍。作为执行前验证门控,符号监视器在动作执行前阻止不安全动作,并提高所测试的每个规划器的模拟器报告的单步成功率(最高 +0.11),这是验证器从未看到的信号(非循环)。由于门控不调用任何模型(0.008 毫秒/规划),它是一个在边缘上几乎免费运行的安全层:SAGE 规划在 Jetson AGX Orin 上复现其质量,而在该平台上小模型验证最为有益。我们发布了基准、无泄漏协议、恢复与安全门控工具,以及验证器可移植性研究(在 ALFWorld 上自动诱导,留出集 0.89)。
cs.RO / 125 / 2609.34270

Bayesian Active Learning for Intent Disambiguation in Interactive Robot Planning

用于交互式机器人规划中意图消歧的贝叶斯主动学习
Li, Huao, Sobolewski, Carson, Saravanos, Augustinos, Tan, William, Karigiannis, John, Fan, Chuchu
Abstract
Interactive robot planning requires robots to infer and execute human intentions from natural language instructions that are often ambiguous, incomplete, or underspecified. Although large language models (LLMs) provide a powerful interface for clarification, relying on the generative model to drive an multi-turn conversation can introduce systematic failures. We propose a Bayesian framework that treats clarification as an active learning problem over grounded Signal Temporal Logic (STL) task specifications. Our method uses LLMs to initialize candidate formal specifications and translate informative contrasts into natural-language clarification questions, while Bayesian optimization maintains uncertainty estimation over user intent and selects queries that maximize information gain. After convergence, the inferred STL specification is passed to a formal planner to synthesize a verifiable robot trajectory. Across four simulated and real-world task domains, our approach generally achieves higher task satisfaction and requires fewer clarification rounds than LLM baselines, while helping smaller models close the performance gap against larger reasoning models.
Chinese Translation
交互式机器人规划要求机器人从自然语言指令中推断并执行人类意图,这些指令通常具有歧义、不完整或描述不充分。尽管大语言模型(LLMs)提供了强大的澄清接口,但依赖生成模型驱动多轮对话可能引入系统性失败。我们提出一个贝叶斯框架,将澄清视为基于接地信号时序逻辑(STL)任务规范的主动学习问题。我们的方法使用LLMs初始化候选形式化规范,并将有信息量的对比转化为自然语言澄清问题,同时贝叶斯优化维护对用户意图的不确定性估计,并选择最大化信息增益的查询。收敛后,推断出的STL规范被传递给形式化规划器,以合成可验证的机器人轨迹。在四个模拟和真实世界任务领域中,我们的方法通常比LLM基线获得更高的任务满意度,并需要更少的澄清轮次,同时帮助较小模型缩小与较大推理模型之间的性能差距。
cs.RO / 126 / 2609.34276

NavHarness: Towards Lifelong Embodied Navigation

NavHarness:迈向终身具身导航
Zhao, Xunyi, Zhou, Jian, Lin, Sihao, Zhou, Gengze, Li, Zerui, Yan, Xinyu, Liu, Jiajun, Hengel, Anton van den, Wu, Qi
Abstract
Frontier models can now perform well on individual embodied navigation tasks through multi-round multimodal reasoning with simple tools. Across successive tasks, however, an agent must also rely on an evolving map and earlier search records, both of which may be incomplete or conflict with new observations. We present NavHarness, a training-free embodied harness towards lifelong navigation that makes memory processing part of the navigation loop. During navigation, its multi-round agentic session draws on maps, task records, and house knowledge, checking them against observations and recording corrections to guide its actions. NavHarness preserves this experience across fresh conversations for new tasks or recovery attempts, while outcome verification and run-end summaries support its later reuse. On GOAT-Bench, NavHarness improves s-SR over context-only independent sessions by 18.6 points with Astra and 22.6 with Opus 5. Using SLAM-estimated poses, NavHarness with GPT-6 Astra achieves state-of-the-art task success of 83.7 s-SR with 36.9 e-SR on GOAT-Bench and 85.9 s-SR on IR2R-CE. To understand these gains, we examine how experience is carried between sessions and find that structured recovery handovers outperform length-matched summaries. In extended deployments across houses, consolidation improves navigation beyond retaining maps and task records, with case studies showing how agents use earlier experience to interpret new goals, investigate unresolved questions, and resume failed searches. We suggest that progress towards lifelong navigation depends on how successive reasoning sessions build on prior experience, alongside improvements in single-task capability.
Chinese Translation
前沿模型现在能够通过使用简单工具进行多轮多模态推理,在单个具身导航任务上表现良好。然而,在连续任务中,智能体还必须依赖不断演化的地图和早先的搜索记录,这两者都可能不完整或与新观测相冲突。我们提出NavHarness,一种面向终身导航的免训练具身框架,将记忆处理作为导航循环的一部分。在导航过程中,其多轮智能体会话利用地图、任务记录和房屋知识,将它们与观测进行核对,并记录修正以指导其行动。NavHarness在用于新任务或恢复尝试的新对话中保留这一经验,而结果验证和运行结束摘要则支持其后续复用。在GOAT-Bench上,NavHarness相较于仅依赖上下文的独立会话,使用Astra时将s-SR提高了18.6个百分点,使用Opus 5时提高了22.6个百分点。使用SLAM估计的位姿,NavHarness配合GPT-6 Astra在GOAT-Bench上实现了83.7 s-SR和36.9 e-SR的最先进任务成功率,并在IR2R-CE上达到85.9 s-SR。为了理解这些增益,我们考察了经验如何在会话之间传递,并发现结构化的恢复交接优于长度匹配的摘要。在跨房屋的扩展部署中,整合改进了导航,其效果超越了仅保留地图和任务记录,案例研究展示了智能体如何利用早先经验来解读新目标、探究未解决问题并恢复失败的搜索。我们认为,迈向终身导航的进展取决于后续推理会话如何建立在先前经验之上,以及单任务能力的提升。
cs.RO / 127 / 2609.34297

TLC-DiT: Task-Aligned Local Visual Conditioning for Robust Multitask Robot Manipulation

TLC-DiT:面向鲁棒多任务机器人操作的任务对齐局部视觉条件
Cai, Xianbo, Ichiwara, Hideyuki, Wang, Zihang, Lu, Yijun, Ogata, Tetsuya
Abstract
Language-conditioned robot policies have made clear progress in multitask manipulation, but task-relevant local visual evidence usually stays hidden inside a visual backbone or attention layers. This leaves the policy difficult to inspect and fragile under visual change, two symptoms of a missing explicit, task-aligned local visual channel. We present TLC-DiT, a plug-in extension of the Multitask Diffusion Transformer (DiT) policy that adds explicit task-guided local visual feature maps without changing the diffusion objective or the action-generation process. For each camera view, frozen DINOv2 patch features are modulated by the CLIP task embedding through FiLM and refined by a lightweight CoordConv CNN adapter into smooth spatial maps, which are concatenated with the original global image, language, joint-state, and timestep conditions. On LIBERO, TLC-DiT reaches a 93.5% average success rate, compared with 86.5% for Multitask DiT and 79.25% for SmolVLA. On LIBERO-plus, the total success rate improves from 54.07% to 57.24%, with larger gains under camera, background, and sensor-noise changes. In real-world bimanual tasks, TLC-DiT raises Teabag Putting completion from 44% to 89% while maintaining comparable Match Box Opening performance. Feature-map visualizations confirm that the model attends to task-relevant regions across views and perturbations, providing a direct way to inspect the visual evidence.
Chinese Translation
语言条件机器人策略在多任务操作中已取得明显进展,但任务相关的局部视觉证据通常隐藏在视觉骨干网络或注意力层内部。这使得策略难以检查,且在视觉变化下脆弱,这两者都是缺失显式、任务对齐的局部视觉通道的症状。我们提出TLC-DiT,它是多任务扩散Transformer(DiT)策略的即插即用扩展,在不改变扩散目标或动作生成过程的情况下,添加了显式的任务引导局部视觉特征图。对于每个相机视图,冻结的DINOv2 patch特征通过FiLM由CLIP任务嵌入进行调制,并由轻量级CoordConv CNN适配器细化为平滑的空间图,然后与原始的全局图像、语言、关节状态和时间步条件拼接。在LIBERO上,TLC-DiT达到93.5%的平均成功率,而多任务DiT为86.5%,SmolVLA为79.25%。在LIBERO-plus上,总成功率从54.07%提升至57.24%,在相机、背景和传感器噪声变化下提升更大。在真实世界双臂任务中,TLC-DiT将Teabag Putting(放置茶包)的完成率从44%提升至89%,同时保持相当的Match Box Opening(打开火柴盒)性能。特征图可视化证实,模型在不同视图和扰动下都能关注任务相关区域,提供了一种直接检查视觉证据的方式。
cs.RO / 128 / 2609.34300

When World Models Lie: Adaptive Safety Analysis Under Wrong Imaginations

当世界模型撒谎时:错误想象下的自适应安全分析
Cao, John, Bansal, Somil
Abstract
World models offer a powerful substrate for safety reasoning in high-dimensional robotic systems, but they are also fallible: their predictions can be biased, miscalibrated, or confidently wrong. This creates a central challenge for latent-space safety filters, which often learn Hamilton-Jacobi safety value functions on the dynamics of a world model. If the world model is incorrect, the resulting value function can inherit its errors and produce overconfident safety estimates. Existing latent safety filters often rely on auxiliary signals such as ensemble disagreement or value-target consistency residuals for adaptation, but these signals can remain small even when the world model's predictions deviate from observations. We propose an adaptive latent safety filter that calibrates safety reasoning using directly observed world-model error. Our method uses Adaptive Conformal Inference to construct online uncertainty sets from discrepancies between predicted and observation-inferred latent states, then evaluates safety pessimistically by minimizing the learned value function over these sets. This allows the filter to remain minimally conservative when the world model is accurate, while becoming more cautious when observations reveal model mismatch. We provide a finite-time coverage guarantee for the adaptive uncertainty radius. Through simulation and hardware experiments, we show that our method significantly reduces failures relative to state-of-the-art latent safety filters while preserving task completion.
Chinese Translation
世界模型为高维机器人系统中的安全推理提供了强大的基础,但它们也容易出错:其预测可能存在偏差、校准不良或自信地错误。这给潜在空间安全滤波器带来了核心挑战,这些滤波器通常在世界模型的动力学上学习 Hamilton-Jacobi 安全价值函数。如果世界模型不正确,所得的价值函数可能会继承其误差,并产生过度自信的安全估计。现有的潜在安全滤波器通常依赖辅助信号(如集成分歧或价值目标一致性残差)进行自适应,但当世界模型的预测偏离观测时,这些信号可能仍然很小。我们提出了一种自适应潜在安全滤波器,利用直接观测到的世界模型误差来校准安全推理。我们的方法使用自适应共形推理(Adaptive Conformal Inference)从预测的潜在状态与观测推断的潜在状态之间的差异构建在线不确定性集,然后通过在这些集合上最小化学习到的价值函数来悲观地评估安全性。这使得滤波器在世界模型准确时保持最小保守性,而当观测揭示模型不匹配时变得更加谨慎。我们为自适应不确定性半径提供了有限时间覆盖保证。通过仿真和硬件实验,我们表明,相对于最先进的潜在安全滤波器,我们的方法显著减少了失败,同时保持了任务完成。
cs.RO / 129 / 2609.34356

Predictive Semantic Safety: From Visual Physical Reasoning to Safety-Critical Control

预测性语义安全:从视觉物理推理到安全关键控制
Kim, Taekyung, Fradi, Salem, Dai, Yanning, Ostaszewski, Mateusz, Schmidhuber, Jürgen
Abstract
Physical interactions can create future hazards that are not apparent from the robot's current geometric surroundings. We present a framework termed Predictive Semantic Safety (PSS), which connects visual physical reasoning to backup-based safety filtering. A vision-language model (VLM) predicts physical events and their timing or directly predicts object displacements. An explicit motion model converts event hypotheses into object trajectories. Split conformal prediction calibrates position errors jointly across specified objects, observation times, and future times; geometric shape bounds convert the resulting position regions into predicted object occupancy. PSS evaluates a prescribed backup maneuver against this occupancy and derives input-affine constraints for minimally modifying the nominal input while preserving backup feasibility under the robot dynamics and input limits. MuJoCo experiments with a Unitree Go1 consider falling fixtures, impact-driven support loss, and contact propagation. PSS achieves a safe episode rate of 99.3%, compared with 43.3% for a Backup Control Barrier Function baseline that only uses current obstacle geometry.
Chinese Translation
物理交互可能产生从机器人当前几何环境中无法明显看出的未来危险。我们提出了一个称为预测性语义安全(Predictive Semantic Safety, PSS)的框架,它将视觉物理推理与基于备份的安全滤波联系起来。视觉-语言模型(VLM)预测物理事件及其发生时间,或直接预测物体位移。一个显式运动模型将事件假设转换为物体轨迹。分裂共形预测联合校准指定物体、观测时间和未来时间上的位置误差;几何形状边界将所得位置区域转换为预测的物体占据。PSS 针对该占据区域评估一个指定的备用机动,并推导出输入仿射约束,以在机器人动力学和输入限制下保持备份可行性的同时,最小程度地修改标称输入。使用 Unitree Go1 的 MuJoCo 实验考虑了坠落物体、冲击驱动的支撑丧失以及接触传播。PSS 实现了 99.3% 的安全回合率,而仅使用当前障碍物几何信息的备份控制障碍函数基线为 43.3%。
cs.RO / 130 / 2609.34362

FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models

FutureDuet:在世界动作模型中解耦观测访问与未来监督
Wu, Jie, Huang, Yuzhi, Liu, Junqi, Zhang, Weichen, Huang, Haibin, Chen, Yin, Jiang, Jingyan, Zhang, Chi
Abstract
World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-level task evolution, whereas wrist cameras move with the end effector, mixing local interaction changes with viewpoint shifts and self-occlusion. These contrasting predictive demands suggest that the two views may benefit from different future objectives. We introduce FutureDuet, which retains both views for control, while allowing each visual stream to receive a different future objective. For the main view, future RGB models task evolution, while interaction masks and robot skeletons focus supervision on task objects and robot motion. For the wrist stream, future latent prediction models short-horizon interaction changes without requiring pixel-level reconstruction. ActionDiT jointly reads the resulting Task State and Interaction State, combining scene-level progress with close-range interaction evidence. All auxiliary prediction modules are training-only, adding no inference overhead. FutureDuet achieves 94.2% clean and 94.1% randomized success on RoboTwin50 and 99.2% average success on LIBERO. The improvements are most pronounced on six RoboTwin50 tasks that require precise interaction, averaging gains of 9.2% and 12.8% over Fast-WAM in clean and randomized settings. Controlled studies further show complementary gains from separating the wrist pathway and designing future supervision separately for the two views.
Chinese Translation
世界动作模型(WAMs)通过未来视觉监督增强机器人动作生成。现有的 WAMs 通常将主视角和腕部观测融合为一个视觉流,并用相同的未来视频目标训练两者,尽管它们的视觉动态不同。稳定的主相机揭示场景级任务演进,而腕部相机随末端执行器移动,将局部交互变化与视角偏移和自遮挡混合。这些截然不同的预测需求表明,两种视角可能受益于不同的未来目标。我们提出 FutureDuet,它保留两种视角用于控制,同时允许每个视觉流接收不同的未来目标。对于主视角,未来 RGB 建模任务演进,而交互掩码和机器人骨架将监督聚焦于任务物体和机器人运动。对于腕部流,未来潜在预测建模短时域交互变化,无需像素级重建。ActionDiT 联合读取得到的 Task State 和 Interaction State,将场景级进展与近距离交互证据结合。所有辅助预测模块仅在训练时使用,不增加推理开销。FutureDuet 在 RoboTwin50 上实现了 94.2% 的干净和 94.1% 的随机化成功率,在 LIBERO 上平均成功率为 99.2%。改进在六个需要精确交互的 RoboTwin50 任务上最为显著,在干净和随机化设置下平均比 Fast-WAM 分别提升 9.2% 和 12.8%。对照研究进一步表明,分离腕部路径并为两种视角单独设计未来监督带来了互补增益。
cs.RO / 131 / 2609.34384

RoboIRGBench: Benchmarking Implicit Referential Grounding in Vision-Language-Action Models

RoboIRGBench:视觉-语言-动作模型中隐式指代定位的基准测试
Akelijiang, Aernaer, Li, Jiannan, Chen, Zhineng, Chen, Jingjing, Zhu, Bin
Abstract
Vision-Language-Action (VLA) models have shown strong capabilities in robotic manipulation, yet existing benchmarks typically assume that task-relevant information is explicitly specified in the instruction. In practice, however, humans frequently refer to objects, quantities, and relations implicitly, requiring robots to recover the intended target from linguistic and perceptual context. We study this capability as Implicit Referential Grounding (IRG) and introduce RoboIRG-Bench, a manipulation benchmark designed to systematically evaluate it. Built upon RoboMME, RoboIRG-Bench contains 40 variants derived from 11 tasks and covers four challenges, including direct, reasoning-mediated, spatial, and contextual referential grounding. As IRG often requires retaining and retrieving previously established context, we evaluate representative VLAs spanning different memory mechanisms. Our evaluation reveals a noticeable referential robustness gap. Models that perform well under explicit instructions can degrade sharply when the same task-relevant information must be recovered from context. Reasoning-mediated and spatial references are particularly challenging, while models using external VLMs show greater robustness but still exhibit significant failures. Moreover, replacing the external VLM with a stronger model does not eliminate these gaps. We further validate these findings on a Franka Research 3 robot arm, where the gap persists under real-world manipulation and manifests as both incorrect referent grounding and downstream execution failures. These results establish IRG as a distinct and underexplored capability for reliable robotic instruction following and highlight the need for VLAs that can robustly integrate language, perception, reasoning, and action.
Chinese Translation
视觉-语言-动作(VLA)模型在机器人操作中已展现出强大的能力,然而现有的基准通常假设任务相关信息在指令中被明确指定。然而在实践中,人类经常隐式地指代物体、数量和关系,要求机器人从语言和感知上下文中恢复意图目标。我们将这种能力称为隐式指代定位(IRG),并引入了 RoboIRG-Bench,一个旨在系统评估该能力的操作基准。基于 RoboMME,RoboIRG-Bench 包含从 11 个任务派生的 40 个变体,并涵盖四个挑战,包括直接、推理中介、空间和上下文指代定位。由于 IRG 通常需要保留和检索先前建立的上下文,我们评估了跨越不同记忆机制的代表性 VLA。我们的评估揭示了一个明显的指代鲁棒性差距。在明确指令下表现良好的模型,当必须从上下文中恢复相同的任务相关信息时,性能可能会急剧下降。推理中介和空间指代尤其具有挑战性,而使用外部 VLM 的模型表现出更强的鲁棒性,但仍存在显著失败。此外,用更强的模型替换外部 VLM 并不能消除这些差距。我们进一步在 Franka Research 3 机械臂上验证了这些发现,其中该差距在真实世界操作中持续存在,并表现为错误的指代定位和下游执行失败。这些结果确立了 IRG 作为可靠机器人指令跟随中一个独特且未被充分探索的能力,并凸显了对能够鲁棒地整合语言、感知、推理和行动的 VLA 的需求。
cs.RO / 132 / 2609.34400

Contact-Aware Impedance Controller for Robot-Assisted Ultrasound Imaging

用于机器人辅助超声成像的接触感知阻抗控制器
Arefin, MD Miraj, Tiryaki, M Efe
Abstract
Safe robot-assisted ultrasound imaging requires a reliable controller able to detect and localize probe--tissue interaction. In this paper, we present a B-mode ultrasound image-based contact perception method and a contact-aware impedance controller for robotic ultrasound imaging. The proposed method detects acoustic contact independently of force measurements, enabling contact-conditioned force/torque taring to reduce residual wrench bias. During contact, the method continuously estimates the effective contact location along the curved probe surface and uses it to update the controller interaction frame, enabling visual servoing of the physical probe--tissue contact point during imaging. Experiments on an agar phantom demonstrated a contact-localization RMSE of $\mathbf{1.46 \pm 0.14}$~mm over probe roll angles from $\mathbf{-15^\circ}$ to $\mathbf{15^\circ}$. During static rolling, the proposed controller maintained task-space tracking accuracy comparable to a conventional fixed-frame impedance controller while reducing the maximum compressive interaction force from $\mathbf{31.56}$~N to $\mathbf{20.09}$~N, corresponding to a $\mathbf{36.3\%}$ reduction. These results demonstrate the potential of ultrasound images as direct contact feedback for safe and accurate robot-assisted ultrasound imaging.
Chinese Translation
安全的机器人辅助超声成像需要一个能够检测和定位探头-组织相互作用的可靠控制器。本文提出了一种基于B模式超声图像的接触感知方法,以及一种用于机器人超声成像的接触感知阻抗控制器。所提方法独立于力测量来检测声学接触,从而实现接触条件下的力/扭矩清零,以减少残余力/扭矩偏差。在接触过程中,该方法连续估计沿弯曲探头表面的有效接触位置,并用于更新控制器的交互坐标系,从而在成像过程中实现对物理探头-组织接触点的视觉伺服。在琼脂体模上的实验表明,在探头滚动角从-15°到15°范围内,接触定位的RMSE为1.46±0.14 mm。在静态滚动过程中,所提出的控制器保持了与传统固定坐标系阻抗控制器相当的任务空间跟踪精度,同时将最大压缩相互作用力从31.56 N降低到20.09 N,相当于减少了36.3%。这些结果证明了超声图像作为直接接触反馈在安全准确的机器人辅助超声成像中的潜力。
cs.RO / 133 / 2609.34412

From Language to Task Maps: Compiling Semantic Relations While Preserving Task-Relevant Freedom

从语言到任务映射:编译语义关系并保留任务相关自由度
Park, Jaegyun, Lee, Jingwang, Lee, Jungsoo, Hwang, Soonwoong, Kim, Wansoo
Abstract
Natural-language manipulation instructions specify qualitative relations, whereas continuous controllers require state-evaluable task quantities, differentials, and completion conditions. Because a qualitative relation generally leaves part of the relative configuration unspecified, expanding it into a complete pose can introduce unintended constraints. We present a typed semantic-to-geometric interface in which language specifies entities, relations, and phases, while each relation indexes a registered specification of its task-relevant distinctions and preserved freedoms. A robot-side compiler grounds these specifications, constructs relation-specific task maps and consistent differentials using conformal geometric algebra, and composes the resulting policies through RMPflow. To evaluate the division of responsibility between the language model and the compiler, we compared a Semantic Topology interface with one that additionally requires relation-specific geometric specifications over 60 instructions. Both produced correct shared semantic content in 41/60 cases, but critical errors under their respective interface requirements occurred in 19/60 and 58/60 cases. Across 64 grounded evaluations spanning eight geometric relation forms, the task maps preserved registered null directions and responded to relation-relevant perturbations; analytic directional derivatives agreed with finite differences, and Jacobian ranks matched the registered dimensions. In three closed-loop ablations using a simulated Franka Emika Panda in MuJoCo, fixing a relation-preserved coordinate increased median terminal progress error by 20.24--71.00~mm while the retained relation errors remained within their evaluation bounds. These results support compiling relation-visible geometry and preserved freedom together into composable continuous objectives.
Chinese Translation
自然语言操作指令指定定性关系,而连续控制器需要可状态评估的任务量、微分和完成条件。由于定性关系通常留下部分相对配置未指定,将其扩展为完整位姿可能引入非预期约束。我们提出了一种类型化的语义到几何接口,其中语言指定实体、关系和阶段,而每个关系索引其任务相关区分和保留自由度的注册规范。机器人侧编译器将这些规范落地,使用共形几何代数构造关系特定的任务映射和一致的微分,并通过 RMPflow 组合得到的策略。为了评估语言模型和编译器之间的责任划分,我们在 60 条指令上比较了 Semantic Topology 接口与额外要求关系特定几何规范的接口。两者在 41/60 的情况下都产生了正确的共享语义内容,但在各自接口要求下的关键错误分别出现在 19/60 和 58/60 的情况下。在涵盖八种几何关系形式的 64 次接地评估中,任务映射保留了注册的零方向,并对关系相关扰动做出响应;解析方向导数与有限差分一致,并且雅可比秩与注册维度匹配。在使用 MuJoCo 中模拟的 Franka Emika Panda 进行的三个闭环消融实验中,固定一个关系保留的坐标使中位终端进度误差增加了 20.24--71.00 mm,而保留的关系误差仍在其评估范围内。这些结果支持将关系可见的几何和保留的自由度一起编译为可组合的连续目标。
cs.RO / 134 / 2609.34414

From World Models to World Action Models: Rethinking Next-State Prediction

从世界模型到世界动作模型:重新思考下一状态预测
Yuan, Tingyu, Ji, Ziming, Guan, Biaoliang, Ye, Wen, Tian, Wenrui, Gu, Zhaopeng, Zhang, Feihong, Yang, Xu, Huang, Yan, Li, Zhaowen, Zhao, Chaoyang, Wang, Jinqiao
Abstract
Predicting the next state is a core paradigm of World Models for modeling physical dynamics, emphasizing prediction fidelity. As World Models evolve into World-Action Models (WAMs), existing methods still fix the next state before training as RGB, a single latent feature, or a static combination of predefined targets, thereby constraining action learning to the inductive biases preserved by a particular representation. To address this limitation, we propose CF-WAM, a dynamic next-state prediction framework that samples visual, semantic, geometric, and interaction projections of the same future, standardizes them into a common video form, and supervises a unified WAM across these projections. The action-relevant constraints exposed by these projections accumulate across training steps, forcing WAM to capture the underlying state-transition structure that supports multiple projections of the same action-conditioned future. This dynamic mechanism also provides a natural cross-embodiment dynamics reference frame for Human and Robot learning. By jointly learning across different next-state parameterizations, heterogeneous Human and Robot experience can bypass appearance differences and directly contribute to shared state-transition learning, improving cross-embodiment generalization. Experiments show that CF-WAM improves both training efficiency and final control performance, while translating Human experience effectively into policy gains. CF-WAM achieves state-of-the-art performance on RoboCasa-GR1 with an average success rate of 82.50%, while reaching 82.65% on LIBERO-Plus and up to 84.00% in real-world evaluations.
Chinese Translation
预测下一状态是世界模型建模物理动力学的核心范式,强调预测保真度。随着世界模型演进为世界动作模型(WAMs),现有方法仍然在训练前将下一状态固定为RGB、单一潜在特征或预定义目标的静态组合,从而将动作学习限制在特定表示所保留的归纳偏置中。为了解决这一局限,我们提出CF-WAM,一种动态下一状态预测框架,它采样同一未来的视觉、语义、几何和交互投影,将它们标准化为共同的视频形式,并在这些投影上监督一个统一的WAM。这些投影所揭示的动作相关约束在训练步骤中累积,迫使WAM捕获支持同一动作条件未来的多种投影的底层状态转移结构。这种动态机制还为人类和机器人学习提供了自然的跨实体动力学参考框架。通过跨不同下一状态参数化的联合学习,异构的人类和机器人经验可以绕过外观差异,直接贡献于共享状态转移学习,从而提高跨实体泛化能力。实验表明,CF-WAM提高了训练效率和最终控制性能,同时将人类经验有效地转化为策略收益。CF-WAM在RoboCasa-GR1上实现了最先进的性能,平均成功率为82.50%,同时在LIBERO-Plus上达到82.65%,在真实世界评估中高达84.00%。
cs.RO / 135 / 2609.34484

ARS: Agentic Reward System for Robot Learning

ARS:用于机器人学习的智能体奖励系统
Hu, Sheng, Lu, Weiyi, Zeng, Lingbing, Weng, Gan, Zhang, Weiwei, Xie, Kai, Mou, Xiaofeng, Xu, Yi
Abstract
Progress reward modeling is the problem of estimating how a robot's behavior changes task progress over time. Reliable estimation requires distinguishing meaningful state changes from failed attempts and task-irrelevant actions. We introduce the Agentic Reward System (ARS), an inference framework for progress reward modeling with general-purpose vision-language models (VLMs), without additional reward-model training. Given an offline trajectory and a task instruction, ARS uses adaptive visual inspection for both event proposal and verification. A subagent proposes a task-relevant event timeline, which a primary agent verifies and revises before estimating per-frame progress. ARS can incorporate optional terminal outcome labels and visual references to inform its judgments. It can also audit progress estimates from external reward models. We evaluate ARS with a 27B VLM on a controlled semantic-mismatch benchmark and downstream policy learning in simulation and on a real robot. The benchmark reveals that several evaluated reward baselines assign spurious progress to wrong-object manipulation even in simple pick-and-place scenes. ARS better suppresses these errors and outperforms these baselines in simulation policy learning. We further demonstrate that ARS supports long-horizon policy learning from mixed-quality offline experience on real-robot multi-screw fastening in a full-scale laboratory replica of an industrial washing-machine assembly line. These results suggest that structured inference and verification can improve the usefulness of general-purpose VLMs for robot reward modeling. Code is at https://github.com/midea-ai/ars
Chinese Translation
进度奖励建模是估计机器人行为如何随时间改变任务进度的问题。可靠的估计需要区分有意义的状态变化与失败尝试和任务无关动作。我们介绍了智能体奖励系统(ARS),一个使用通用视觉语言模型(VLM)进行进度奖励建模的推理框架,无需额外的奖励模型训练。给定离线轨迹和任务指令,ARS使用自适应视觉检查进行事件提议和验证。一个子代理提出任务相关的事件时间线,主代理在估计每帧进度之前对其进行验证和修正。ARS可以结合可选的终端结果标签和视觉参考来辅助其判断。它还可以审核来自外部奖励模型的进度估计。我们使用一个27B的VLM,在受控的语义不匹配基准测试以及仿真和真实机器人上的下游策略学习中对ARS进行了评估。该基准测试表明,即使在简单的拾放场景中,几个评估的奖励基线也会将虚假进度分配给错误物体的操作。ARS更好地抑制了这些错误,并在仿真策略学习中优于这些基线。我们进一步证明,ARS支持从混合质量的离线经验中进行长时程策略学习,应用于全尺寸工业洗衣机装配线实验室复制品中的真实机器人多螺钉紧固任务。这些结果表明,结构化的推理和验证可以提高通用VLM在机器人奖励建模中的实用性。代码位于:https://github.com/midea-ai/ars
cs.RO / 136 / 2609.34485

Trajectory-Safe Orienteering for Human-Robot Shared Environments

面向人-机器人共享环境的轨迹安全定向问题
Gao, Songqun, Basei, Elena, Roveri, Marco, Palopoli, Luigi, Fontanelli, Daniele
Abstract
Orienteering problem (OP) has wide real-world applications and also great potential in human-robot collaboration. However, existing approaches struggle to simultaneously ensure safe and feasible trajectories while achieving high-quality task execution in shared workspaces. To this end, this work studies the OP with time windows and variable profits (OPTWVP). A two-stage DEcoupled discrete-Continuous Optimization with Service-time-guided Trajectory (DeCoST) approach is proposed to effectively solve OPTWVP in shared spaces. Meanwhile, the safety-aware time windows of nodes and the discretized workspace are introduced to ensure collision-free trajectories between the end effector and the human. Preliminary results validate the effectiveness of DeCoST in generating collision-free trajectory plans while preserving the quality of orienteering tasks.
Chinese Translation
定向问题(OP)在现实世界中具有广泛的应用,在人-机器人协作中也展现出巨大潜力。然而,现有方法难以在共享工作空间中同时确保安全可行的轨迹并实现高质量的任务执行。为此,本工作研究了带有时间窗和可变利润的OP(OPTWVP)。提出了一种两阶段解耦离散-连续优化与服务时间引导轨迹(DeCoST)方法,以有效求解共享空间中的OPTWVP。同时,引入了节点的安全感知时间窗和离散化工作空间,以确保末端执行器与人类之间的无碰撞轨迹。初步结果验证了DeCoST在生成无碰撞轨迹规划的同时保持定向任务质量的有效性。
cs.RO / 137 / 2609.34486

Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction

基于步到步预测的模型信息引导安全强化学习用于双足运动
Paredes, Victor, Hereid, Ayonga
Abstract
Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches. We propose a model-informed reinforcement learning framework anchored to the analytical Angular Momentum Linear Inverted Pendulum (ALIP) template. We provide a step-to-step safety certificate for ALIP stepping via a discrete exponential control barrier function (DECBF) and use it as (i) a training-time shaping signal and (ii) a runtime action filter that minimally adjusts swing-foot placement to satisfy template-level constraints. Full-order safety is evaluated empirically on the Digit humanoid in MuJoCo with a whole-body controller stack. Compared to an unconstrained baseline, our approach reduces safety-violation events in the reported external-disturbance trial, while larger lateral-velocity transients reveal a safety-tracking tradeoff.
Chinese Translation
人形机器人有望在杂乱、以人为中心的环境中实现多样化的移动能力,但实际部署要求原则性的安全保障。经典的基于模型的步态生成器能够产生可解释的运动,但往往缺乏现代基于强化学习(RL)方法的鲁棒性和适应性。我们提出了一种模型信息引导的强化学习框架,该框架以解析的角动量线性倒立摆(ALIP)模板为基础。我们通过离散指数控制障碍函数(DECBF)为ALIP迈步提供了步到步的安全证书,并将其用作(i)训练时的塑造信号和(ii)运行时动作滤波器,以最小程度地调整摆动脚放置,从而满足模板级约束。全阶安全性在MuJoCo中对Digit人形机器人使用全身控制器栈进行了实证评估。与无约束基线相比,我们的方法在报告的外部扰动试验中减少了安全违规事件,而较大的横向速度瞬变揭示了安全与跟踪之间的权衡。
cs.RO / 138 / 2609.34498

Robot-Assisted Deployment and Maintenance of Inflatable Modules for Lunar Habitation: A Field Demonstration

面向月球居住的充气模块机器人辅助部署与维护:现场演示
Uno, Kentaro, Karimov, Shamistan, Mishra, Ashutosh, Neppel, Elian, Gozbasi, Hazal, Santra, Shreya, Kimura, Shinichi, Yoshida, Kazuya
Abstract
Long-term human habitation and in-situ development on the Moon open a new era of space utilization. In this context, robots are a key technology for facilitating the construction of future human outposts. Toward the deployment and establishment of human habitation modules on the lunar surface, we propose a combined system consisting of inflatable modules and a modular, reconfigurable robotic system. This paper presents a report demonstrating various robot-assisted task executions using real hardware, namely the modular and reconfigurable robot MoonBot and the inflatable module HIDAS, to enhance the reliability of their deployment and maintenance. The demonstrated tasks include robotic inspection during inflation, module position alignment, final safety locking, and three-dimensional mapping for post-deployment maintenance. All demonstrations were conducted either in a laboratory environment or at a lunar analogue test site. Finally, lessons learned are discussed to provide essential insights for this robotic application to future lunar habitation.
Chinese Translation
月球上的长期人类居住和原位开发开启了太空利用的新时代。在此背景下,机器人是促进未来人类前哨站建设的关键技术。为了实现月球表面人类居住模块的部署和建立,我们提出了一种由充气模块和模块化、可重构机器人系统组成的组合系统。本文报告了使用真实硬件(即模块化可重构机器人 MoonBot 和充气模块 HIDAS)演示的各种机器人辅助任务执行,以提高其部署和维护的可靠性。所演示的任务包括充气过程中的机器人检测、模块位置对准、最终安全锁定以及用于部署后维护的三维建图。所有演示均在实验室环境或月球模拟试验场进行。最后,讨论了经验教训,为这一机器人应用在未来月球居住中提供重要见解。
cs.RO / 139 / 2609.34512

MonoEgo: Monocular Metric Egocentric Demonstration Capture with Passive Wrist Constellations and Sparse Workstation Anchors

MonoEgo:基于被动腕部星座和稀疏工作站锚点的单目度量第一人称演示采集
Xu, Jie, Yu, Kangjin, Jin, Ziyi, Wang, Beichen, Xia, Zhongpu
Abstract
Image-aligned metric demonstrations often require dedicated tracking hardware and synchronization across devices. We present MonoEgo, a capture system that replaces active wrist instrumentation with offline monocular reconstruction. One 90-FPS global-shutter camera observes calibrated passive wrist constellations, sparse workstation anchors, and the scene on a shared image clock. MonoTag SLAM combines marker corners with ORB geometry and uses visual evidence to reject ambiguous planar-marker poses. Its metric Atlas supports interval scale re-anchoring, verified map merging, and retrospective localization of earlier frames supported by the final map. Camera and wrist-constellation outputs retain validity and map provenance, and unsupported motion is left missing. Experiments show metric tracking beyond continuous anchor visibility, reconnection of supported map components, and recovery of some missing camera poses. Comparisons against a multisensor camera reference and separate stationary-constellation tests characterize trajectory agreement and precision while revealing incomplete coverage and residual geometric uncertainty. The results indicate that passive fixtures and offline reconstruction can reduce capture-side requirements. Dynamic accuracy, deployment, and downstream policy benefits require further study.
Chinese Translation
图像对齐的度量演示通常需要专用跟踪硬件和跨设备同步。我们提出MonoEgo,一种采集系统,用离线单目重建替代主动腕部仪器。一个90帧/秒的全局快门相机在共享图像时钟上观测经过校准的被动腕部星座、稀疏工作站锚点和场景。MonoTag SLAM将标记角点与ORB几何相结合,并利用视觉证据拒绝模糊的平面标记位姿。其度量Atlas支持区间尺度重锚定、验证地图合并以及由最终地图支持的早期帧的回溯定位。相机和腕部星座输出保留有效性和地图溯源,不支持的运动保持缺失。实验表明,度量跟踪超越连续锚点可见性,支持的地图组件重连,以及恢复一些缺失的相机位姿。与多传感器相机参考和单独的静止星座测试的比较,表征了轨迹一致性和精度,同时揭示了不完整的覆盖和残余几何不确定性。结果表明,被动固定装置和离线重建可以减少采集端需求。动态精度、部署和下游策略收益需要进一步研究。
cs.RO / 140 / 2609.34550

Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning

注视提示:面向视觉-语言-动作微调的时间密集人类注意力
Zhou, Yihan, Yan, Rui, Li, Mingcong, Huang, Zheyuan, Yang, Xu, Guo, Xueyang, Mo, Yilin
Abstract
Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide frame-level visual guidance for VLA fine-tuning. During training, recorded gaze locations are rendered as crosshairs on the robot's head-camera images. At deployment, a lightweight predictor estimates gaze locations from recent images and the instruction, supplying the same type of visual prompt without an eye tracker or changes to the policy architecture. Instantiated with $\pi_0$, gaze prompting increases mean success from $26.3\%$ to $56.0\%$ across six real-world bimanual manipulation tasks, with gains also observed when a single policy is trained on all six tasks. We release \textsc{GazeMani}, a dataset of $1{,}200$ teleoperated trajectories with synchronized gaze.
Chinese Translation
视觉-语言-动作(VLA)微调在每个步骤将图像与动作配对,但通常仅提供任务级的语言指令,使得逐时刻的视觉相关性保持隐式。我们引入了眼动追踪器监督的注视提示,它利用VR遥操作期间记录的眼动数据为VLA微调提供帧级视觉引导。在训练期间,记录的眼动位置以十字准星的形式渲染在机器人头部相机的图像上。在部署时,一个轻量级预测器根据最近的图像和指令估计眼动位置,提供相同类型的视觉提示,而无需眼动追踪器或对策略架构进行更改。以$\pi_0$为例,注视提示在六个真实世界的双臂操作任务上将平均成功率从$26.3\%$提高到$56.0\%$,并且在所有六个任务上训练单个策略时也观察到了增益。我们发布了GazeMani,一个包含$1{,}200$条同步眼动的遥操作轨迹数据集。
cs.RO / 141 / 2609.34554

Where Memory Belongs: Ledger, an Object Ledger for Memory-Augmented VLAs

记忆应归何处:Ledger,一种面向记忆增强VLA的对象账本
Dieudonné, Tanguy, Jedlicki, Jack B., Yang, Heng
Abstract
Memory is essential for long-horizon, partially observed robotic manipulation: a robot must remember which object was placed in a drawer, whose cup it moved, or how many action cycles have elapsed. Recent vision-language-action (VLA) models embed memory directly inside the policy, but benchmarks show no single in-policy mechanism covers all spatio-temporal dimensions, trailing oracle methods by a wide margin. We argue that memory type dictates where memory should reside: short-term perceptual memory (repetition, timing, retracing) belongs inside the policy, while long-term object memory (persistent spatial state, containment, event history) belongs outside as an explicit, readable record. We present Ledger, a harness that realizes this split over a single fine-tuned $\pi_{0.5}$ policy by pairing an in-policy frame-sampling memory with an external spatio-temporal object memory, the ledger, built from a SAM3 tracker and a VLM captioner of the demonstration and read by an LLM planner that decides at step boundaries. On RoboMME, Ledger reaches the highest four-suite average among the evaluated methods, 64.3% (vs. 45.9% for the strongest prior method under identical evaluation), leading object reference (60.7% vs. 40.3%) and object permanence (86.7% vs. 56.2%) using a single set of weights. Choosing the memory source at runtime, from the instruction and the record, removes the need for a task-level router.
Chinese Translation
记忆对于长时程、部分可观测的机器人操作至关重要:机器人必须记住哪个物体被放进了抽屉、它移动了谁的杯子,或者已经过去了多少个动作周期。最近的视觉-语言-动作(VLA)模型将记忆直接嵌入策略内部,但基准测试表明,没有单一的策略内机制能够覆盖所有时空维度,远远落后于oracle方法。我们认为,记忆类型决定了记忆应该位于何处:短期感知记忆(重复、计时、回溯)属于策略内部,而长期物体记忆(持久空间状态、包含关系、事件历史)则属于外部,作为显式、可读的记录。我们提出了Ledger,一个在单个微调的π_{0.5}策略上实现这种分离的框架,它通过将策略内的帧采样记忆与外部时空物体记忆(即账本)配对。账本由SAM3跟踪器和演示的VLM描述器构建,并由一个在步骤边界处进行决策的LLM规划器读取。在RoboMME上,Ledger在所有评估方法中达到了最高的四个套件平均值,为64.3%(在相同评估下,最强先前方法为45.9%),并在使用单一组权重的情况下,在物体指代(60.7% vs. 40.3%)和客体永久性(86.7% vs. 56.2%)方面领先。在运行时根据指令和记录选择记忆来源,消除了对任务级路由器的需求。
cs.RO / 142 / 2609.34594

Aerial GRIPPER: A Gradient-based Real-time Inverse-game Predictor and Planner

空中GRIPPER:一种基于梯度的实时逆博弈预测器与规划器
Chen, Zeshuai, Wang, Meng, Jia, Jindou, Yu, Xiang, Guo, Lei
Abstract
Accurate capture of non-cooperative targets is critical. In an attempt to tackle this intractable challenge, an aerial gripper system integrated with a Gradient-based Real-time Inverse-game Predictor and PlannER (GRIPPER) framework is proposed. The interaction is formulated as a general-sum pursuit-evasion game under incomplete information. Specifically, underlying cost parameters of the target are inferred online, and the open-loop Nash equilibrium (OLNE) strategy is iteratively refined within a receding-horizon loop. To ensure high-frequency execution, a computationally friendly gradient-based inverse-game solver is developed. Without explicit computation of the Hessian inverse, the optimized solution is updated (> 50 Hz) based on implicit differentiation and fast Hessian-vector products. Meanwhile, an anti-disturbance controller is developed to overcome disturbances of uncertain payload and gripper actuation, enabling precise tracking of the planned trajectory and accurate grasping of the target. Simulations and real-world experiments illustrate the superior computational efficiency and task performance of GRIPPER. The task of capturing and delivering a non-cooperative target is accomplished, highlighting the robustness, adaptability, and real-time performance of the framework in highly adversarial scenarios.
Chinese Translation
准确捕获非合作目标至关重要。为应对这一棘手挑战,提出了一种集成了基于梯度的实时逆博弈预测器与规划器(GRIPPER)框架的空中抓取系统。将交互建模为不完全信息下的一般和追逃博弈。具体而言,在线推断目标的潜在代价参数,并在滚动时域循环中迭代优化开环纳什均衡(OLNE)策略。为确保高频执行,开发了一种计算友好的基于梯度的逆博弈求解器。在无需显式计算Hessian逆的情况下,基于隐式微分和快速Hessian-向量乘积,以>50 Hz更新优化解。同时,开发了一种抗扰控制器,以克服不确定负载和抓取器驱动带来的扰动,从而实现对规划轨迹的精确跟踪和对目标的准确抓取。仿真和真实世界实验说明了GRIPPER优越的计算效率和任务性能。完成了捕获并递送非合作目标的任务,凸显了该框架在高度对抗场景中的鲁棒性、适应性和实时性能。
cs.RO / 143 / 2609.34608

Efficient World Action Model Inference with Adaptive Intermediate States

基于自适应中间状态的高效世界动作模型推理
Liu, Zhinnan, Han, Haozhi, Zhang, Ruge, Ma, Teng, Ma, Tao, Liu, Zheng, Chen, Yifeng, Zhang, Yunquan, Cao, Ting, Liu, Yunxin, Li, Kun
Abstract
World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and intermediate representations can quickly render retained state stale. Preserving useful computation therefore requires adapting inference state rather than reusing it as-is. To this end, we present $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$, a training-free framework that accelerates WAM inference by preserving and adapting inference state for efficient and accurate continuation as the control loop evolves. Across closed-loop replans, Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan, reducing redundant trajectory generation. Across denoising steps, Observation Rebinding performs anticipatory inference during action execution and rebinds retained denoising state to the real observation for continuation when consistency checks pass, reducing latency exposed to the control loop. Across Transformer layers, Residual Rescaling selectively rescales retained layer state and refreshes it through full computation of the middle layers when probe checks fail, reducing repeated Transformer computation. Evaluations of three representative WAM architectures on LIBERO and RoboTwin 2.0 show that $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$ achieves 1.47-3.05$\times$ speedups in observation-to-action latency and 2.23-3.27$\times$ speedups in GPU inference time per replan, while preserving 96.69-99.54% of native WAM task success.
Chinese Translation
世界动作模型(World Action Models, WAMs)通过联合建模动作和环境动力学,实现面向未来的控制。然而,迭代扩散或流推理会带来大量的去噪延迟。先前的推理状态为加速提供了自然的机会,但变化的规划上下文、观测和中间表示会迅速使保留的状态变得过时。因此,保留有用的计算需要自适应地调整推理状态,而不是按原样重用。为此,我们提出了WAMACHINE,一个无需训练的框架,通过保留和自适应调整推理状态,随着控制循环的推进,实现高效且准确的延续,从而加速WAM推理。在闭环重新规划中,轨迹重映射(Trajectory Remapping)将来自前一次重新规划的重新规划状态重新映射,以初始化下一次重新规划,减少冗余的轨迹生成。在去噪步骤中,观测重绑定(Observation Rebinding)在动作执行期间执行预期推理,并在一致性检查通过时将保留的去噪状态重新绑定到真实观测以继续,从而减少暴露给控制循环的延迟。在Transformer层中,残差重缩放(Residual Rescaling)选择性地重新缩放保留的层状态,并在探测检查失败时通过中间层的完整计算进行刷新,减少重复的Transformer计算。在LIBERO和RoboTwin 2.0上对三种代表性WAM架构的评估表明,WAMACHINE在观测到动作的延迟上实现了1.47-3.05倍的加速,在每次重新规划的GPU推理时间上实现了2.23-3.27倍的加速,同时保持了原生WAM任务成功率的96.69-99.54%。
cs.RO / 144 / 2609.34666

On the Numerical Reliability of Differentiable Physics-Based Optimization for Robotic Material Manipulation

关于面向机器人材料操作的可微物理优化的数值可靠性
Yang, Xintong, Wei, Minglun, Lai, Yu-Kun, Ji, Ze
Abstract
Differentiable physics is increasingly used in robotic material manipulation for system identification, trajectory or skill optimization, demonstration generation, and robot or end-effector design. These applications depend on gradients propagated through long, contact-rich simulation rollouts. We study the numerical reliability of those gradients using two Material Point Method (MPM) system-identification benchmarks derived from elastoplastic and granular manipulation. The benchmarks provide controlled cases for three effects that also arise in broader differentiable physics-based optimization. GPU many-to-one sums whose order depends on thread scheduling changed long-horizon gradients and reversed the sign of one parameter gradient relative to a deterministic reference. Finite-difference checks became less reliable for longer rollouts because repeated-run loss variation grew much faster than the loss change produced by the tested parameter perturbations. Observation and loss definitions changed optimization behaviour and the solution preferred by an independent metric. These results motivate reproducible accumulation, finite-difference validation that compares perturbation-induced loss changes with repeated-run variation, and explicit reporting of objective construction when differentiable simulation is used for robotic optimization.
Chinese Translation
可微物理正越来越多地用于机器人材料操作中的系统辨识、轨迹或技能优化、演示生成以及机器人或末端执行器设计。这些应用依赖于通过长时、富接触仿真推演传播的梯度。我们使用两个源自弹塑性和颗粒材料操作的物质点法(MPM)系统辨识基准,研究了这些梯度的数值可靠性。这些基准为三种效应提供了受控案例,这些效应在更广泛的可微物理优化中也会出现。顺序取决于线程调度的GPU多对一求和改变了长时程梯度,并且相对于确定性参考,反转了一个参数梯度的符号。对于较长的推演,有限差分检查变得不太可靠,因为重复运行的损失变异增长速度远快于所测试参数扰动产生的损失变化。观测和损失定义改变了优化行为以及独立度量所偏好的解。这些结果推动了可复现的累加、将扰动引起的损失变化与重复运行变异进行比较的有限差分验证,以及在可微仿真用于机器人优化时明确报告目标构建。
cs.RO / 145 / 2609.34674

HOI-Retarget: Contact-Centric Retargeting for Human-Object Interaction

HOI-Retarget:以接触为中心的人-物交互重定向
Shin, Jihwan, Escoriza, Adrià López, He, Junzhe, Heyrman, Matthias, Hutter, Marco
Abstract
Learning from demonstration (LfD) has enabled humanoid robots to acquire diverse whole-body skills, but extending this paradigm to human-object interaction (HOI) is limited by the availability of robot-compatible interaction references. We present HOI-Retarget, a contact-centric retargeting method that transfers HOI onto a humanoid robot for large-scale motion-data generation. Its windowed trajectory optimization uses every labeled contact as a target in the object frame, balancing body tracking, foot support and smoothness under the robot's kinematic limits. The method can augment a single demonstration across object sizes, absorb contacts reconstructed from monocular video, and extend to several robots manipulating one object. We publicly release the code and the retargeted motion dataset.
Chinese Translation
从演示中学习(LfD)已使类人机器人能够获得多样化的全身技能,但将该范式扩展到人-物交互(HOI)受限于机器人兼容交互参考数据的可用性。我们提出 HOI-Retarget,一种以接触为中心的重定向方法,可将 HOI 迁移到类人机器人上,以实现大规模运动数据生成。其窗口化轨迹优化将每个标注接触作为物体坐标系中的目标,在机器人运动学限制下平衡身体跟踪、足部支撑和平滑性。该方法能够将单个演示增强到不同物体尺寸,吸收从单目视频重建的接触,并扩展到多个机器人操作同一物体。我们公开代码和重定向后的运动数据集。
cs.RO / 146 / 2609.34684

Natural State-Prediction Accuracy can Hide Weak Controlled Responsiveness in VLA Readouts

自然状态预测准确率可能掩盖VLA读出中较弱的受控响应性
Kim, Hyungjoon, Son, Wonbin, Lee, Mi Young, Lee, Jun Young, Rho, Seungmin
Abstract
Accurately decoding object states from the internal representations of vision-language-action (VLA) models does not establish that the predictions respond faithfully to changes in the target physical state. In natural observations, object state, robot configuration, occlusion, and task progress vary together, allowing contextual cues to contribute to prediction. In this paper, we introduce an evaluation framework that separates prediction accuracy, target-state responsiveness, and context stability using physically validated observations that cross target coordinates with robot contexts. We demonstrate that high natural-trajectory accuracy can coexist with weak controlled target-state responsiveness in fixed representation-readout pairs. Comparisons and interventions involving representations, readouts, and training data show that the three properties provide distinct diagnostic information. Furthermore, adding responsiveness and context sensitivity to a failure predictor based on initial state error and physical variables reduces policy-failure prediction error on new initializations relative to the specified baseline while same-observation controlled MAE is also informative. These findings motivate evaluating target-state responsiveness and context stability alongside natural prediction accuracy, and examining their relationship to actual policy behavior and task outcomes.
Chinese Translation
从视觉-语言-动作(VLA)模型的内部表征中准确解码物体状态,并不能证明预测能够忠实响应目标物理状态的变化。在自然观察中,物体状态、机器人构型、遮挡和任务进度共同变化,使得上下文线索能够对预测产生贡献。在本文中,我们引入了一个评估框架,该框架使用将目标坐标与机器人上下文交叉的物理验证观察,分离预测准确率、目标状态响应性和上下文稳定性。我们证明,在固定的表征-读出对中,高自然轨迹准确率可以与弱受控目标状态响应性共存。涉及表征、读出和训练数据的比较和干预表明,这三种属性提供了不同的诊断信息。此外,将响应性和上下文敏感性添加到基于初始状态误差和物理变量的失败预测器中,相对于指定的基线,减少了新初始化上的策略失败预测误差,同时同一观察的受控MAE也具有信息量。这些发现促使我们在评估自然预测准确率的同时,评估目标状态响应性和上下文稳定性,并检查它们与实际策略行为和任务结果的关系。
cs.RO / 147 / 2609.34695

Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control

零阶奖励塑形对稳定控制中策略梯度的充分性
Zhang, Yisheng, Wang, Tao, Gao, Sicun
Abstract
Reward shaping is fundamental to modern robotic control with deep reinforcement learning (RL), yet practitioners still rely heavily on heuristic principles borrowed from classical optimal control and trajectory optimization. Existing methods rarely distinguish reward terms that are intrinsic to the control objective from numerical regularizers, leading to brittle hyperparameter tuning. To determine which quantities a reward must contain, we study the stabilization control problem with a focus on zeroth-order (configuration) and first-order (velocity) information. We theoretically and empirically demonstrate that policy gradient methods can successfully solve stabilization tasks without first-order reward terms, adding such terms can instead introduce severe sensitivity as their scale grows. Conversely, our findings confirm that reward functions must be zeroth-order complete over goal-relevant coordinates, while the first-order state remains necessary in the policy observation under our low-dissipation assumptions. Overall, these results provide actionable and principled guidance for reward design in robotic RL.
Chinese Translation
奖励塑形是使用深度强化学习(RL)的现代机器人控制的基础,然而实践者仍严重依赖从经典最优控制和轨迹优化中借用的启发式原则。现有方法很少区分控制目标固有的奖励项与数值正则化项,导致超参数调优脆弱。为了确定奖励必须包含哪些量,我们研究稳定控制问题,重点关注零阶(构型)和一阶(速度)信息。我们从理论和实验上证明,策略梯度方法无需一阶奖励项即可成功解决稳定任务;而加入此类项反而可能随着其尺度增大而引入严重敏感性。与之相对,我们的发现证实,奖励函数必须在目标相关坐标上零阶完整,而在我们的低耗散假设下,一阶状态在策略观测中仍然必要。总体而言,这些结果为机器人强化学习中的奖励设计提供了可操作且有原则的指导。
cs.RO / 148 / 2609.34702

MarsLab: A Martian Rover Simulator for Planetary Rover Autonomous Navigation

MarsLab:面向行星巡视器自主导航的火星车模拟器
Kim, Hoyun, Kim, Beomsu, Kim, Giseop
Abstract
Future Mars missions will require rover autonomy that can operate across unstructured terrain, changing illumination, atmospheric dust, and limited communication. Simulation is a practical way to study these conditions before deployment, but existing Mars-relevant resources differ in scope, including mission-oriented simulators, fixed analog datasets, task-specific environments, and open robotics interfaces. In this context, we present MarsLab, an open-source, ROS2-native Mars rover simulator for autonomy and navigation algorithm development. MarsLab combines HiRISE-derived and procedural terrain with customizable rock, crater, solar-illumination, and atmospheric-dust settings, and runs a Perseverance-class rover model in NVIDIA Isaac Sim. The runtime publishes RGB, depth, RGB-D point clouds, LiDAR, IMU, wheel odometry, and Ground Truth (GT) pose data through standard ROS2 topics. We demonstrate MarsLab with Simultaneous Localization and Mapping (SLAM) benchmarks across sensing modalities, dust levels, scene geometry, and route length, and with Visual Place Recognition (VPR) benchmarks over repeated Mars Base traversals under illumination and dust changes. The results illustrate how controlled scene variation and shared GT trajectories can be used to compare trajectory-level estimation and image-level place recognition within the same simulator. Our Project Page: https://kimhoyun-robotair.github.io/MarsLab/.
Chinese Translation
未来的火星任务将需要巡视器具备自主能力,能够在非结构化地形、变化的光照、大气尘埃以及有限通信条件下运行。仿真是在部署前研究这些条件的实用方法,但现有的火星相关资源在覆盖范围上各不相同,包括面向任务的模拟器、固定的类比数据集、特定任务环境和开放式机器人接口。在此背景下,我们提出了 MarsLab,一个开源、原生支持 ROS2 的火星车模拟器,用于自主与导航算法开发。MarsLab 将源自 HiRISE 的地形与程序化生成地形相结合,并提供可定制的岩石、陨石坑、太阳光照和大气尘埃设置,在 NVIDIA Isaac Sim 中运行一个毅力号级别的火星车模型。运行时通过标准 ROS2 话题发布 RGB、深度、RGB-D 点云、LiDAR、IMU、轮式里程计和真值(GT)位姿数据。我们通过跨传感模态、尘埃水平、场景几何和路线长度的同步定位与建图(SLAM)基准测试,以及在光照和尘埃变化下重复 Mars Base 穿越的视觉位置识别(VPR)基准测试,展示了 MarsLab。结果说明了如何利用受控场景变化和共享 GT 轨迹,在同一模拟器内比较轨迹级估计和图像级位置识别。我们的项目主页:https://kimhoyun-robotair.github.io/MarsLab/。
cs.RO / 149 / 2609.34707

SOR-Nav: Search or Relocate? Context-Gated Exploration and Cross-Region Relocation for Object Navigation

SOR-Nav:搜索还是迁移?面向物体导航的上下文门控探索与跨区域迁移
Ji, Yuan, Li, Zirui, Cai, Yuxin, Wu, Shuge, Han, Boon Siew, Lv, Chen
Abstract
Object navigation requires an embodied agent to find an object in an unseen environment under partial observability and a limited motion budget. Existing methods primarily optimize where the robot should go next by ranking candidate destinations. In contrast to these methods, we present SOR-Nav, a hierarchical navigation system that explicitly arbitrates between continuing to explore the current context and abandoning it for a more promising reachable region. First, an autonomous semantic exploration system is built that accumulates persistent 3D object clusters and organizes reachable frontiers into a cluster decision graph to provide an efficient search abstraction. Then, SOR-Nav uses a context-gated LLM-driven object-search supervisor to evaluate the suitability of the current search context and decide whether to continue exploration or perform cross-region relocation to another reachable frontier cluster. Across the complete, unfiltered validation sets of HM3D-v1, HM3D-v2, and MP3D, SOR-Nav achieves the strongest reported Success Rate (SR) and Success weighted by Path Length (SPL) on all three benchmarks. On MP3D in particular, it more than doubles the previous best SPL from 18.1\% to 38.5\% while increasing SR from 50.7\% to 61.8\%. Nested HM3D-v2 ablations validate the proposed decision structure, while a continuous three-target physical deployment demonstrates persistent ObjectNav operation in real-world scenarios.
Chinese Translation
物体导航要求具身智能体在部分可观测且运动预算有限的未知环境中找到目标物体。现有方法主要通过排序候选目的地来优化机器人下一步应前往的位置。与这些方法不同,我们提出了SOR-Nav,一种分层导航系统,它明确地在继续探索当前上下文与放弃当前上下文以转向更有希望的可达区域之间进行仲裁。首先,构建了一个自主语义探索系统,该系统累积持久的三维物体聚类,并将可达前沿组织成聚类决策图,以提供高效的搜索抽象。然后,SOR-Nav使用上下文门控的LLM驱动的物体搜索监督器来评估当前搜索上下文的适宜性,并决定是继续探索还是执行跨区域迁移到另一个可达前沿聚类。在HM3D-v1、HM3D-v2和MP3D的完整未过滤验证集上,SOR-Nav在所有三个基准测试中都取得了报告的最强成功率(SR)和路径长度加权成功率(SPL)。特别是在MP3D上,它将先前的最佳SPL从18.1%提高到38.5%,翻了一倍多,同时将SR从50.7%提高到61.8%。嵌套的HM3D-v2消融实验验证了所提出的决策结构,而连续的三目标物理部署展示了在真实世界场景中持续的ObjectNav操作。
cs.RO / 150 / 2609.34719

ExcavaTwin: Training-Free Geometry-Guided Semantic Elevation Mapping for Autonomous Excavation

ExcavaTwin:面向自主挖掘的免训练几何引导语义高程建图
Deng, Yu, Zeng, Lingshan, Hu, Tong, Dai, Rushi
Abstract
Autonomous excavation requires a spatial representation that jointly captures terrain geometry and task-relevant semantics. Existing excavation mapping is largely elevation-centric, while generic semantic models remain unstable in unstructured outdoor scenes. We present ExcavaTwin, a pure-vision geometry-guided semantic elevation mapping framework without excavation-specific training. Given multi-view RGB images, the framework: 1) reconstructs scene geometry and semantic observations using frozen vision models; 2) derives terrain and non-terrain geometric support; 3) performs geometry-constrained multi-view semantic fusion to suppress implausible predictions and recover incomplete observations; and 4) projects the fused state into a task-oriented semantic elevation map. Experiments on public datasets and real excavation scenes demonstrate reliable geometric and semantic perception. In real excavation, the system achieved an average update interval of approximately 1.4 s and a mean elevation error of 12.74cm in dynamically modified regions. Larger errors mainly occur during rapid terrain changes and transient visual disturbances caused by machine motion.
Chinese Translation
自主挖掘需要一种能够同时捕捉地形几何和任务相关语义的空间表示。现有的挖掘建图主要侧重于高程,而通用的语义模型在非结构化户外场景中仍不稳定。我们提出了ExcavaTwin,一个无需挖掘特定训练的纯视觉几何引导语义高程建图框架。给定多视角RGB图像,该框架:1) 使用冻结的视觉模型重建场景几何和语义观测;2) 推导地形和非地形的几何支撑;3) 执行几何约束的多视角语义融合,以抑制不合理的预测并恢复不完整的观测;4) 将融合状态投影到面向任务的语义高程图中。在公共数据集和真实挖掘场景上的实验证明了可靠的几何和语义感知。在真实挖掘中,系统在动态修改区域实现了约1.4秒的平均更新间隔和12.74厘米的平均高程误差。较大的误差主要发生在快速地形变化和机器运动引起的瞬态视觉干扰期间。
cs.RO / 151 / 2609.34724

DexWeave: Learning Dexterous Humanoid Loco-Manipulation from Human Demonstrations

DexWeave:从人类演示中学习灵巧人形移动操作
Sun, Naichuan, Shen, Haotian, Zhang, Yizhang, Feng, Luying, Wang, Haoze, Xiangli, Yuanbo, Jin, Yaochu, Liu, Peidong
Abstract
Learning dexterous humanoid loco-manipulation from human demonstrations requires transferring not only human motion, but also the coordinated interaction structure underlying the demonstrated behavior. This is challenging because embodiment differences distort the coupling among body motion, wrist placement, finger articulation, and object interaction, while kinematically accurate references may still be difficult to realize under robot dynamics. We present DexWeave, a unified framework that connects interaction-consistent motion retargeting with anatomy-aware whole-body policy learning. DexWeave first employs a two-stage retargeting procedure that initializes body and hand motions with specialized solvers and subsequently performs coupled refinement over the upper-body interaction chain while preserving lower-body support. The resulting references are tracked by an anatomy-aware Transformer policy that represents anatomical regions as structured tokens and uses directed masked attention to model their dependencies, with object information selectively conditioning the upper-body pathway for dexterous interaction. The policy jointly outputs body and dexterous-hand actions and is trained directly with reinforcement learning, without pretrained tracking policies, teacher-student distillation, or subsequent residual refinement. DexWeave improves retargeting fidelity and interaction consistency while achieving higher manipulation performance and faster policy convergence than MLP baselines. We further deploy the learned policies on a physical Unitree G1 humanoid equipped with Inspire dexterous hands, demonstrating dexterous whole-body loco-manipulation in the real world. See our project page (https://dexweave.github.io) for videos.
Chinese Translation
从人类演示中学习灵巧人形移动操作,不仅需要迁移人体运动,还需要迁移演示行为背后的协调交互结构。这具有挑战性,因为具身差异会扭曲身体运动、手腕位置、手指关节运动和物体交互之间的耦合,而运动学上准确的参考在机器人动力学下可能仍然难以实现。我们提出了DexWeave,一个统一的框架,将交互一致的运动重定向与解剖感知的全身策略学习联系起来。DexWeave首先采用两阶段重定向过程,使用专门的求解器初始化身体和手部运动,随后在上半身交互链上进行耦合细化,同时保持下半身支撑。生成的参考由解剖感知的Transformer策略进行跟踪,该策略将解剖区域表示为结构化token,并使用有向掩码注意力来建模它们的依赖关系,其中物体信息选择性地调节上半身通路以实现灵巧交互。该策略联合输出身体和灵巧手动作,并直接通过强化学习进行训练,无需预训练跟踪策略、师生蒸馏或后续残差细化。与MLP基线相比,DexWeave提高了重定向保真度和交互一致性,同时实现了更高的操作性能和更快的策略收敛。我们进一步将学习到的策略部署在配备Inspire灵巧手的实体Unitree G1人形机器人上,展示了真实世界中的灵巧全身移动操作。请访问我们的项目页面(https://dexweave.github.io)观看视频。
cs.RO / 152 / 2609.34730

Action Sequence Transfer via LLMs for Heterogeneous Environments

基于LLMs的异构环境动作序列迁移
Chung, Choongho, Shin, DongHwan, Lee, Sung-Hee
Abstract
We present an action sequence transfer system that adaptively transfers user action sequences across different target spaces. Given an input action sequence from a source space and scene graph representations of both the source and target environments, our system predicts a corresponding action sequence in the target space by adapting to the spatial and object constraints of the new environment. To achieve this, we leverage multi-level representations of user activity to generalize actions at varying levels of abstraction. To demonstrate our system, we collect a new scene graph-based dataset derived from the Ego4D GoalStep dataset for evaluation. Results indicate that our system can generate valid action sequences even between spaces with drastically different object configurations.
Chinese Translation
我们提出了一种动作序列迁移系统,它能够自适应地将用户动作序列迁移到不同的目标空间。给定来自源空间的输入动作序列以及源环境和目标环境的场景图表示,我们的系统通过适应新环境的空间和物体约束,预测目标空间中相应的动作序列。为此,我们利用用户活动的多层级表示,在不同抽象层级上泛化动作。为了演示我们的系统,我们收集了一个新的基于场景图的数据集,该数据集源自 Ego4D GoalStep 数据集,用于评估。结果表明,即使在物体配置差异巨大的空间之间,我们的系统也能生成有效的动作序列。
cs.RO / 153 / 2609.34743

Simulation for Planetary Robotic Perception and Autonomy: A Concise Survey of Recent Capabilities and Gaps

面向行星机器人感知与自主性的仿真:近期能力与差距的简明综述
Kim, Hoyun, Kim, Giseop
Abstract
Planetary robotics is an important enabler of scientific exploration in environments where direct human-in-the-loop operation is costly, hazardous, or infeasible. However, developing and validating planetary robotic systems remains difficult because representative field testing is expensive, limited, and often unrepeatable under mission-relevant conditions. In this setting, simulation serves as a central tool for perception and autonomy research, synthetic data generation, system integration, and pre-deployment evaluation. Despite its importance, the literature on planetary robotics simulation remains dispersed across different simulation engines, implementations, and application settings. This paper surveys simulation works for planetary robotic perception and autonomy across four practical axes: Openness and Availability, Scenario and Platform Coverage, Sensor and Perception Support, and Environmental and Operational Realism. The surveyed simulation works report visual or physical fidelity and support perception-oriented workflows. They also indicate uneven public availability, rover-centered coverage, partial support for specialized sensing modalities, and uneven reporting of operational constraints such as onboard computation, energy, and communication restrictions.
Chinese Translation
行星机器人技术是在直接人在回路操作成本高昂、危险或不可行的环境中进行科学探索的重要使能技术。然而,开发和验证行星机器人系统仍然困难,因为具有代表性的现场测试成本高昂、受限,并且在任务相关条件下往往不可重复。在此背景下,仿真成为感知与自主研究、合成数据生成、系统集成和部署前评估的核心工具。尽管仿真很重要,但关于行星机器人仿真的文献仍然分散在不同的仿真引擎、实现和应用场景中。本文从四个实践维度综述了面向行星机器人感知与自主的仿真工作:开放性与可用性、场景与平台覆盖、传感器与感知支持、环境与操作真实性。所综述的仿真工作报告了视觉或物理保真度,并支持面向感知的工作流程。它们还表明,公共可用性不均衡、以巡视器为中心的覆盖、对专用传感模态的部分支持,以及对机载计算、能量和通信限制等操作约束的报告不均衡。
cs.RO / 154 / 2609.34782

CoHuB: A Simulation Benchmark for Multi-Humanoid Collaboration

CoHuB:多人形机器人协作仿真基准
Park, Hyunjin, Chae, Jebeom, Park, Minwoo, Park, Sunghyun, Yoo, Hanjun, Choi, Seoyeon, Yoo, Soochul, Seo, Joohwan, Idrees, Sarmad, Hyun, Jae-Sang, Lee, Jongmin, Horowitz, Roberto, Lee, Youngwoon, Choi, Jongeun
Abstract
Many physical tasks in human environments require collaboration, from assisting a partner to jointly manipulating an object. Yet, existing humanoid benchmarks largely focus on single-humanoid skills and lack evaluation of multi-humanoid collaboration under egocentric visual observations. We introduce CoHuB (Collaborative Multi-Humanoid Benchmark), a simulation benchmark for multi-humanoid collaboration under egocentric visual observations. CoHuB provides 10 tasks, eight with two humanoids and two with three humanoids, spanning diverse collaboration patterns. We also provide synchronized demonstrations collected through a multi-operator VR teleoperation pipeline, in which each operator controls one humanoid from its egocentric view. Experiments with representative visuomotor policies reveal substantial challenges across different forms of coordinated perception and control. CoHuB provides a foundation for developing and evaluating multi-humanoid collaboration policies.
Chinese Translation
人类环境中的许多物理任务需要协作,从协助伙伴到共同操作物体。然而,现有的人形机器人基准主要关注单人形机器人技能,缺乏在自我中心视觉观察下的多人形机器人协作评估。我们介绍了CoHuB(协作多人形机器人基准),一个在自我中心视觉观察下的多人形机器人协作仿真基准。CoHuB提供了10个任务,其中8个涉及两个人形机器人,2个涉及三个人形机器人,涵盖了多种协作模式。我们还提供了通过多操作员VR遥操作流程收集的同步演示,其中每个操作员从其自我中心视角控制一个人形机器人。使用代表性视觉运动策略进行的实验揭示了在不同形式的协调感知和控制方面存在巨大挑战。CoHuB为开发和评估多人形机器人协作策略提供了基础。
cs.RO / 155 / 2609.34823

AGRO-SUVIDE: Agentic Robotics for Surgical Viscoelastic Debridement

AGRO-SUVIDE:用于手术粘弹性清创的智能体机器人
Jin, Shutong, Chen, Ziyang, Satish, Preethi, Shen, Medow, Guthart, Gary, Pokorny, Florian T., Goldberg, Ken
Abstract
Augmented dexterity has the potential to reduce the fatigue experienced by surgeons during repetitive surgical tasks. In this paper, we propose the first AGentic RObotics framework for SUrgical VIscoelastic DEbridement (AGRO-SUVIDE), the repeated removal of small fragments attached to a viscoelastic substrate. Leveraging the self-improving and coding capability of agents, AGRO-SUVIDE adopts a modular framework. Specifically, the demonstration analysis module automatically identifies recurring skills from a single expert demonstration, using both visual and kinematic information. The construction module then builds each skill, either as a procedural model-based skill the agent codes against a scaffolded library or as a model-free policy-based skill. At runtime, the monitoring module composes the skills into a loop-style graph sized to the number of fragments it observes, then verifies pre- and post-conditions of each skill to decide whether to advance or retry. We evaluate AGRO-SUVIDE through 340 physical trials on the da Vinci Research Kit (dVRK). AGRO-SUVIDE achieves an average single-fragment removal success rate of 85%, completing consecutive three-fragment removal at 60% and at 95% with one human intervention. It further generalizes to unseen five-fragment scenarios with an average success rate of 80% for single-fragment removal. Project page: https://surgical-robotics.github.io/AGRO-SUVIDE/
Chinese Translation
增强的灵巧性有可能减轻外科医生在重复性手术任务中的疲劳。在本文中,我们提出了首个用于手术粘弹性清创的智能体机器人框架(AGRO-SUVIDE),即重复去除附着在粘弹性基底上的小碎片。利用智能体的自我改进和编码能力,AGRO-SUVIDE采用模块化框架。具体来说,演示分析模块利用视觉和运动学信息,从单个专家演示中自动识别重复出现的技能。然后,构建模块构建每个技能,要么作为基于程序模型的技能,智能体针对脚手架库进行编码,要么作为无模型的基于策略的技能。在运行时,监控模块将技能组合成一个循环式图,其大小根据它观察到的碎片数量而定,然后验证每个技能的前置和后置条件,以决定是继续还是重试。我们通过达芬奇研究套件(dVRK)上的340次物理试验评估AGRO-SUVIDE。AGRO-SUVIDE实现了平均单碎片去除成功率为85%,连续三碎片去除的成功率为60%,在有一次人工干预的情况下达到95%。它进一步泛化到未见过的五碎片场景,单碎片去除的平均成功率为80%。项目页面:https://surgical-robotics.github.io/AGRO-SUVIDE/
cs.RO / 156 / 2609.34911

Don't Throw Away the Tail: Action Upcycling for Policy Acceleration

不要丢弃尾部:用于策略加速的 Action Upcycling
Kwon, Taesung, Park, Jangho, Park, Sunwoo, Kim, Youngmin, Jin, Seonghyun, Jun, Youngjun, Choi, Kyumin, Ye, Jong Chul
Abstract
Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replanning. Choosing the length of this prefix, the execution horizon, poses a trade-off between reactivity and efficiency. A short horizon keeps the policy reactive to the environment, but requires frequent policy calls. Recent test-time methods adaptively select the horizon for each chunk, but they either read model internals, where the signal must be chosen for each architecture, or draw extra samples, which adds cost. We propose *Action Upcycling*, a training-free algorithm that reuses actions the policy would otherwise discard, without accessing model internals or drawing extra samples. We find that discarded actions stay close to their replanned versions as long as the action velocity remains smooth. Action Upcycling therefore extends the execution horizon up to the point where the velocity begins to fluctuate. Extensive experiments on simulated and real-world manipulation tasks show that Action Upcycling reduces policy calls by 1.2--1.7$\times$ with no loss in success rate, across multiple Vision-Language-Action Models (VLAs) and even a World Action Model (WAM). It applies to any chunked policy at negligible cost and is orthogonal to other policy acceleration methods such as few-step sampling and streaming action decoding, opening a new axis for policy acceleration.
Chinese Translation
现代机器人策略从单次观测中预测未来动作块,仅执行前缀,并在重新规划前丢弃其余部分。选择该前缀的长度(即执行视界)需要在反应性和效率之间进行权衡。短视界使策略对环境保持反应性,但需要频繁调用策略。最近的测试时方法为每个动作块自适应选择视界,但它们要么读取模型内部信息(其中信号必须针对每种架构进行选择),要么抽取额外样本,这增加了成本。我们提出 Action Upcycling,一种无需训练的算法,它重用策略原本会丢弃的动作,而无需访问模型内部信息或抽取额外样本。我们发现,只要动作速度保持平滑,被丢弃的动作与其重新规划的版本保持接近。因此,Action Upcycling 将执行视界延长至速度开始波动的点。在模拟和真实世界操作任务上的大量实验表明,Action Upcycling 在多种视觉-语言-动作模型(VLAs)甚至世界动作模型(WAM)上,将策略调用减少 1.2-1.7 倍,且成功率无损失。它以可忽略的成本适用于任何分块策略,并且与其他策略加速方法(如少步采样和流式动作解码)正交,为策略加速开辟了新维度。
cs.RO / 157 / 2609.34968

RoboFL: Federated Expert Assembly for World Action Models

RoboFL:面向世界动作模型的联邦专家组装
Zhang, Rongyu, Fan, Ruizhi, Lou, Yunfan, Fang, Hengyu, Zheng, Shenli, Wu, Chenrui, Jin, Yili, Du, Li, Wang, Dan, Du, Yuan, Zhang, Shanghang
Abstract
Vision-language-action and world-action models are increasingly popular, yet remain bottlenecked by physical interaction data that is scarce, institutionally siloed, and task-heterogeneous. A natural federated solution is to let each client adapt a shared foundation model through parameter-efficient fine-tuning, avoiding the exchange of full-model updates. However, federating these adapters is nontrivial, as naive aggregation can entangle incompatible updates, while incorporating MoE-style routing into federated aggregation may dilute specialization and destabilize expert selection. We present RoboFL, which instantiates MoSAIC (Mixture of Slotted Adapters) for federated world-action learning. MoSAIC directly installs locally trained LoRA adapters as the expert branches of a server MoE. Server-side routers learn token assignments over these prior-informed branches while jointly refining routing and expert parameters. Foresight-to-Action Routing Distillation (FARD) aligns routing across the model's three paths, while Path-Consensus Expert Aggregation (PCEA) converts complete expert updates into a compact global adapter for personalized redistribution. Experiments on RoboTwin 2.0, RLBench, and a real-world Franka robot arm show the superiority of RoboFL with structured expert assembly, as it outperforms centralized PEFT InternVLA-A1 by 12.23% on the Franka arm, while reducing per-round client communication by up to 86.81% relative to MoE-based federated VLA baselines.
Chinese Translation
视觉-语言-动作模型和世界动作模型日益流行,但仍受限于稀缺、机构孤岛化和任务异构的物理交互数据。一种自然的联邦解决方案是让每个客户端通过参数高效微调来适配共享基础模型,避免交换全模型更新。然而,联邦化这些适配器并非易事,因为朴素聚合会纠缠不兼容的更新,而将 MoE 风格的路由引入联邦聚合可能稀释专业化并破坏专家选择的稳定性。我们提出了 RoboFL,它实例化了 MoSAIC(Mixture of Slotted Adapters),用于联邦世界动作学习。MoSAIC 直接将本地训练的 LoRA 适配器安装为服务器 MoE 的专家分支。服务器端路由器在这些先验信息分支上学习令牌分配,同时联合优化路由和专家参数。预见-行动路由蒸馏(FARD)对齐模型三条路径上的路由,而路径共识专家聚合(PCEA)将完整的专家更新转换为紧凑的全局适配器,用于个性化再分发。在 RoboTwin 2.0、RLBench 和真实世界 Franka 机械臂上的实验表明,具有结构化专家组装的 RoboFL 具有优越性,它在 Franka 机械臂上比集中式 PEFT InternVLA-A1 高出 12.23%,同时相对于基于 MoE 的联邦 VLA 基线,将每轮客户端通信减少高达 86.81%。
cs.RO / 158 / 2609.34969

NavJev: Efficient Vision-Language Navigation via Action-Centric Visual Compression and Discriminative Action-Semantic Memory

NavJev: 基于以动作为中心的视觉压缩与判别性动作语义记忆的高效视觉语言导航
Sheng, Kai, Wang, Liuyi, Li, Jinlong, Dai, Haojie, Liu, Chengju, Chen, Qijun
Abstract
Recent zero-shot Vision-and-Language Navigation (VLN) methods increasingly rely on multimodal large language models (MLLMs) to reason over visual observations, navigation instructions, and candidate actions. Although effective, repeatedly invoking autoregressive multimodal reasoning at every navigation step introduces substantial inference latency, limiting the responsiveness of embodied agents. We propose NavJev, an efficient VLN framework that reformulates online navigation from repeated multimodal generation into compact visual compression followed by lightweight typed action selection. Specifically, Action-Centric Visual Compression (ACVC) integrates waypoint geometry, BLIP captions, and RAM semantic tags into compact representations of candidate actions, while Discriminative Action-Semantic Memory (DASM) filters shared semantics and maintains discriminative action-specific evidence across navigation steps. Based on these representations, Jev directly performs structured probabilistic decisions over the available action set. Experiments on R2R-CE show that NavJev achieves 27.0% SR and 22.4% SPL with only 0.65 s per navigation step, while substantially reducing inference latency and cost compared with MLLM-based VLN methods. The project page is available at https://kai-sheng-caesar.github.io/NavJev/.
Chinese Translation
最近的零样本视觉语言导航(VLN)方法越来越依赖于多模态大语言模型(MLLMs)来对视觉观察、导航指令和候选动作进行推理。尽管有效,但在每个导航步骤反复调用自回归多模态推理会引入大量的推理延迟,限制了具身智能体的响应能力。我们提出NavJev,一个高效的VLN框架,它将在线导航从重复的多模态生成重构为紧凑的视觉压缩,随后进行轻量级类型化动作选择。具体来说,以动作为中心的视觉压缩(ACVC)将路点几何、BLIP描述和RAM语义标签整合为候选动作的紧凑表示,而判别性动作语义记忆(DASM)则过滤共享语义,并在导航步骤中保持判别性的动作特定证据。基于这些表示,Jev直接在可用动作集上执行结构化的概率决策。在R2R-CE上的实验表明,NavJev在每个导航步骤仅需0.65秒的情况下达到了27.0%的SR和22.4%的SPL,同时与基于MLLM的VLN方法相比大幅降低了推理延迟和成本。项目页面见 https://kai-sheng-caesar.github.io/NavJev/。
cs.RO / 159 / 2609.35000

Graph-Based Simultaneous Path and Foothold Planning for Multi-Limbed Intra-Vehicular Robots in Space Stations

基于图的空间站多肢舱内机器人同步路径与足点规划
Imai, Masazumi, Uno, Kentaro, Kuwahara, Toshinori, Yoshida, Kazuya
Abstract
Robot-aided operations in space stations are essential for reducing the workload of astronauts and improving the efficiency of on-orbit activities. Multi-limbed intra-vehicular robots (MLIVRs) equipped with grappling end-effectors have emerged as a promising solution, as they can securely grasp pre-existing interfaces, such as handrails and seat tracks, thereby enabling stable locomotion and forceful manipulation in microgravity environments. Since graspable locations on these interfaces are spatially limited and discretely distributed, motion planning for MLIVRs must be addressed jointly with foothold planning. This paper presents a simultaneous path and foothold planning framework based on graph theory for MLIVRs. The proposed method efficiently searches for feasible stance sequences for a multi-limbed robot while satisfying manipulability constraints. The effectiveness of the proposed framework is validated through simulations in a 3D model of the International Space Station (ISS) cabin, demonstrating its capability to generate feasible and efficient locomotion plans in realistic intra-vehicular environments.
Chinese Translation
空间站中的机器人辅助操作对于减轻宇航员工作负荷和提高在轨活动效率至关重要。配备抓取末端执行器的多肢舱内机器人(MLIVRs)已成为一种有前景的解决方案,因为它们能够牢固抓握现有接口,如扶手和座椅轨道,从而在微重力环境中实现稳定移动和有力操作。由于这些接口上的可抓握位置在空间上有限且离散分布,MLIVRs 的运动规划必须与足点规划联合解决。本文提出了一种基于图论的 MLIVRs 同步路径与足点规划框架。所提方法能够高效搜索多肢机器人的可行支撑序列,同时满足可操作度约束。该框架的有效性通过在国际空间站(ISS)舱段三维模型中的仿真得到验证,展示了其在真实舱内环境中生成可行且高效移动规划的能力。
cs.RO / 160 / 2609.35003

Learning to Act under Visual Interruptions with Vision-Language-Action Models

在视觉中断下使用视觉-语言-动作模型学习行动
Jiang, Mingle, Xu, Rui, Wang, Yunke, Xu, Chang
Abstract
Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but they are typically developed and evaluated with all camera streams available throughout task execution. When a camera stops delivering frames during task execution, the policy must continue acting without access to subsequent observations from the missing view. Despite its practical importance, how such interruptions affect closed-loop manipulation remains insufficiently understood. To investigate this problem, we introduce MAIL-Bench, a benchmark that evaluates visual interruptions with VLA models. By interrupting different cameras at multiple stages of each policy's successful reference trajectory, MAIL-Bench measures how well policies retain their capabilities when visual inputs become unavailable. Building on this benchmark, we propose MINT, which first trains VLA policies to remain functional under missing visual inputs. At inference time, MINT selectively supplements missing observations using optical-flow extrapolation or an action-conditioned world model, and withdraws predicted views when they become unreliable. Experiments on $\pi_{0.5}$ and GR00T N1.5 show that MINT significantly improves task success under camera loss over the original models. Experiments on AgiBot G2 further demonstrate the real-robot deployment under camera loss. The benchmark is available at https://minglejiang.github.io/Mail-Bench/
Chinese Translation
视觉-语言-动作 (VLA) 模型在机器人操作中已展现出强大的能力,但它们通常是在任务执行全程所有相机流均可用的情况下进行开发和评估的。当相机在任务执行过程中停止提供帧时,策略必须在无法获取缺失视图后续观测的情况下继续行动。尽管这一问题具有实际重要性,此类中断如何影响闭环操作仍未被充分理解。为了研究这一问题,我们引入了 MAIL-Bench,一个用 VLA 模型评估视觉中断的基准测试。通过在每条策略成功参考轨迹的多个阶段中断不同的相机,MAIL-Bench 衡量了当视觉输入不可用时策略保持其能力的程度。基于此基准,我们提出了 MINT,它首先训练 VLA 策略在缺失视觉输入的情况下仍能保持功能。在推理时,MINT 使用光流外推或动作条件世界模型选择性地补充缺失的观测,并在预测视图变得不可靠时将其撤回。在 $\pi_{0.5}$ 和 GR00T N1.5 上的实验表明,MINT 在相机丢失情况下显著提高了任务成功率,优于原始模型。在 AgiBot G2 上的实验进一步展示了相机丢失下的真实机器人部署。基准测试可在 https://minglejiang.github.io/Mail-Bench/ 获取。
cs.RO / 161 / 2609.35039

Do Not Cut When Uncertain: Rejectable and Calibrated Decision Heads for VLA Policies in Robotic Harvesting

不确定时不要切割:用于机器人采收中 VLA 策略的可拒绝与校准决策头
Zhang, Heng
Abstract
Vision-Language-Action (VLA) policies trained with behavior cloning or flow matching are optimized to output an action trajectory, but they cannot express "I don't know" or "I should not act." In robotic harvesting, occlusion makes single-frame decisions fundamentally ambiguous: identical pixels can correspond either to a cuttable stem or to no stem at all. Existing VLAs are forced to commit, leading to high-confidence errors with irreversible consequences. We argue that the failure mode of a VLA is determined not by backbone scale but by its output interface. We propose Rejectable and Calibrated Decision Heads (RCDH), a typed, rejectable, and calibrated output interface that can be attached to a frozen VLA backbone without retraining or new features. RCDH introduces (i) a decision schema with explicit rejection and ordered, conditional decomposition, and (ii) a calibration procedure for risk-aware abstention. We evaluate RCDH on a robotic harvesting platform with controllable leaf occlusion, comparing generative, enumerated, calibrated, and rejectable interfaces. We show that replacing only the output head restores out-of-distribution usability under occlusion while preserving in-distribution performance. We further test whether the ordering of the rejection space is critical. Our results suggest that the right to refuse, rather than a larger model, is the missing interface for reliable manipulation under uncertainty.
Chinese Translation
使用行为克隆或流匹配训练的视觉-语言-动作(VLA)策略被优化为输出动作轨迹,但无法表达“我不知道”或“我不该行动”。在机器人采收中,遮挡使单帧决策本质上具有歧义:相同像素既可能对应可切割的茎,也可能完全不对应茎。现有 VLA 被迫做出承诺,导致高置信度错误并带来不可逆的后果。我们认为,VLA 的失败模式不是由主干规模决定的,而是由其输出接口决定的。我们提出可拒绝与校准决策头(RCDH),一种类型化、可拒绝且校准的输出接口,可附加到冻结的 VLA 主干上,无需重新训练或新特征。RCDH 引入 (i) 具有显式拒绝和有序条件分解的决策模式,以及 (ii) 用于风险感知弃权的校准过程。我们在具有可控叶片遮挡的机器人采收平台上评估 RCDH,比较生成式、枚举式、校准式和可拒绝式接口。我们表明,仅替换输出头即可在遮挡下恢复分布外可用性,同时保持分布内性能。我们进一步测试拒绝空间的顺序是否关键。我们的结果表明,在不确定性下实现可靠操作所缺失的接口是拒绝的权利,而非更大的模型。
cs.RO / 162 / 2609.35047

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

EMPIRIC:面向机器人规划的残差世界模型的实验驱动学习
Liang, Yichao, Li, Amber, Nguyen, Dat, Bunnapradist, Emily, Naim, Michelangelo, Kodali, Sreela, Merler, Matteo, Li, Bowen, Gopinathan, Kiran, Liu, Yiyun, Pimpalkhare, Nikhil, Tenenbaum, Joshua B., Weller, Adrian, Tavares, Zenna, Silver, Tom, Ellis, Kevin
Abstract
A robot should be able to learn through experiments how unfamiliar objects behave and interact, then plan with that knowledge. It need not start from scratch: physics engines supply knowledge of motion and contact, but can omit entire mechanisms, such as glue curing, water heating, or wind. We present EMPIRIC, an agent that learns a residual world model: a physics engine extended with code for the missing mechanisms. The learned programs can introduce new forces, constraints, and hidden state, and Bayesian inference estimates their parameters and states from noisy observations. The resulting model lets the agent predict the outcomes of actions, choose informative experiments, and revise its hypotheses when predictions fail. Across five simulated domains, EMPIRIC learns interpretable, reusable models, and solves more tasks with fewer environment interactions than all three baselines. On a physical robot, it learns wind forces and domino masses to solve a manipulation task. Website and code: https://yichao-liang.github.io/empiric
Chinese Translation
机器人应能通过实验学习陌生物体的行为与交互,然后利用这些知识进行规划。它无需从零开始:物理引擎提供运动与接触知识,但可能遗漏整个机制,如胶水固化、水加热或风。我们提出 EMPIRIC,一个学习残差世界模型的智能体:一个扩展了缺失机制代码的物理引擎。学习到的程序可引入新的力、约束和隐状态,贝叶斯推断从噪声观测中估计其参数和状态。所得模型让智能体预测动作结果、选择信息量大的实验,并在预测失败时修正假设。在五个模拟领域中,EMPIRIC 学习到可解释、可复用的模型,并以比三个基线更少的环境交互解决更多任务。在物理机器人上,它学习风力与多米诺骨牌质量以解决操作任务。网站和代码:https://yichao-liang.github.io/empiric
cs.RO / 163 / 2609.35094

QuadHand: A Compact Quadrotor Aerial Manipulator with MRC-SDF-Based Whole-Body Motion Planning

QuadHand:一种基于MRC-SDF全身运动规划的紧凑型四旋翼空中机械臂
Jin, Rui, Liu, Ruiyang, Xu, Xinhang, Jin, Haotian, Wang, Yi, Yang, Yizhuo, Xi, Lihua
Abstract
Uncrewed aerial manipulators (UAMs) integrate robotic arms with aerial platforms for three-dimensional physical interaction. However, enlarging the workspace increases arm-induced disturbances, while existing geometric representations face a trade-off between geometric fidelity and computational efficiency in close-proximity interaction. This paper presents QuadHand, a compact quadrotor aerial manipulator with a 3-DoF arm, gripper, and battery-assisted passive CoG compensation module to reduce dominant arm-induced disturbances. We further propose MRC-SDF, a Multi-articulated Robot-Centric Signed Distance Field that preserves fine geometric detail with tractable computation, and a spatiotemporal whole-body trajectory optimization framework that jointly optimizes the quadrotor and manipulator for safe and executable trajectory generation. Simulations and real-world experiments demonstrate safe and executable aerial manipulation in complex environments.
Chinese Translation
无人空中机械臂(UAMs)将机械臂与空中平台集成,用于三维物理交互。然而,扩大工作空间会增加机械臂引起的扰动,而现有的几何表示在近距离交互中面临几何保真度与计算效率之间的权衡。本文提出了QuadHand,一种紧凑型四旋翼空中机械臂,具有3自由度机械臂、夹爪和电池辅助的被动重心补偿模块,以减少主要的机械臂引起的扰动。我们进一步提出了MRC-SDF,一种多关节机器人中心符号距离场,可在可处理的计算下保留精细几何细节,以及一个时空全身轨迹优化框架,联合优化四旋翼和机械臂,以生成安全可执行的轨迹。仿真和真实世界实验证明了在复杂环境中安全可执行的空中操作。
cs.RO / 164 / 2609.35200

ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation

ReCAT:记忆、计数与计时:用于机器人操作的结构化循环记忆
Vanjani, Pankhuri, Hatab, Mostafa, Mizrakli, Can, Shaj, Vaisakh, Li, Zhuoyue, Reuss, Moritz, Lioutikov, Rudolf
Abstract
Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors, such as recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time. We present ReCAT, a language-conditioned policy with structured recurrent memory. An instruction-conditioned encoder forms features from the current observation. A recurrent memory integrates the observation stream through Mamba-2 layers and one causal attention layer. A flow-matching Transformer decoder reads the current and the historical representation through separate cross-attention in every block. ReCAT reaches 95.3\% average success on LIBERO and 62.4\% on RMBench, with the best or tied-best result on six of nine tasks. On three real-robot tasks probing spatial recall, event counting, and interval timing, the best ReCAT variant reaches 66.7\% average success, against 8.3\% for the strongest short-history baseline. Controlled comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance. They also show that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule updates on spatial recall. Project website is at https://intuitive-robots.github.io/ReCAT
Chinese Translation
依赖记忆的操作要求机器人利用当前传感器不再可用的信息做出决策,例如回忆先前的视觉线索、跟踪任务进度、计数重复事件或估计经过的时间。我们提出了 ReCAT,一种具有结构化循环记忆的语言条件策略。一个指令条件编码器从当前观察中形成特征。一个循环记忆通过 Mamba-2 层和一个因果注意力层整合观察流。一个 flow-matching Transformer 解码器在每个块中通过单独的交叉注意力读取当前和历史表示。ReCAT 在 LIBERO 上达到 95.3% 的平均成功率,在 RMBench 上达到 62.4%,在九个任务中的六个上取得最佳或并列最佳结果。在探测空间回忆、事件计数和间隔计时的三个真实机器人任务上,最佳 ReCAT 变体达到 66.7% 的平均成功率,而最强的短历史基线仅为 8.3%。ReCAT 内部的受控比较表明,观察编码器和每个块的记忆条件化对于这一性能是必要的。它们还表明,为高效序列建模开发的更新规则作为机器人记忆时表现不同:加法更新在计数和计时上观察到最高成功率,而 delta 规则更新在空间回忆上表现最佳。项目网站:https://intuitive-robots.github.io/ReCAT
cs.RO / 165 / 2609.35231

Zero-Shot Reactive Obstacle Avoidance for Generative Robot Policies

面向生成式机器人策略的零样本反应式避障
Guo, Weihang, Kavraki, Lydia E.
Abstract
We propose NUDGE (Nudge Update via Differentiable GEometry), a training-free obstacle-avoidance procedure that can be incorporated in any robot policy based on diffusion or flow matching, including diffusion policies and vision-language-action models. Our work injects gradients from a signed distance field, a function returning each point's distance to the nearest obstacle, into the policy at inference time to steer it away from obstacles. It supports any common action parameterization, from absolute or relative joint poses to end-effector poses, through a differentiable joint-trajectory decoder. Experiments show that NUDGE preserves the policy's task distribution and runs reactively in real time.
Chinese Translation
我们提出了NUDGE(Nudge Update via Differentiable GEometry,通过可微几何的微调更新),一种无需训练的避障过程,可以集成到任何基于扩散或流匹配的机器人策略中,包括扩散策略和视觉-语言-动作模型。我们的工作在推理时将来自符号距离场(一种返回每个点到最近障碍物距离的函数)的梯度注入策略中,以引导其远离障碍物。它通过可微的关节轨迹解码器支持任何常见的动作参数化,从绝对或相对关节位姿到末端执行器位姿。实验表明,NUDGE保留了策略的任务分布,并实时反应式运行。
cs.RO / 166 / 2609.35249

Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies

Spatial Grafting:为流匹配机器人策略锚定3D特征
Liu, Dingsheng, Wu, Yangzheng, Asadi, Mahboubeh, Li, Zhiyuan, Huang, Jinbang, Xiao, Yixin, Cao, Tongtong, Zhang, Yingxue
Abstract
Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape without stating where it lies with respect to the robot. How best to deliver these features to a pretrained policy remains unresolved. We propose Spatial Grafting, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry. Spatial Grafting constructs metric-grounded spatial tokens and injects them into the flow-matching action expert through cross-attention, without modifying the host's perceptual pathway, so the host retains the full benefit of its pretraining. We evaluate it more broadly than any geometry-aware policy we compare against: one graft architecture, with no per-host redesign, on two VLAs and two WAMs, across four simulation benchmarks that span short-horizon manipulation, visual robustness, clutter and long-horizon mobile manipulation, and on three real-robot platforms with single- and dual-arm configurations. On RoboTwin 2.0, a dual-arm manipulation benchmark, the graft improves every host across VLAs and WAMs. Grafted $\pi_{0.5}$ gains 11.3% and 15.6% on clean and randomized scenes, reaching 94.0% and 92.4%, above the strongest published 3D-conditioned policy, WAM4D (93.8% and 89.9%). The margin widens as the horizon lengthens: on tasks from BEHAVIOR-1K, a dual-arm mobile manipulation challenge scored by average task progress, it surpasses the 2025 challenge winner on five of six tasks,by up to 0.47 Q-score, and exceeds a map-conditioned spatial policy on average across the three tasks both report.
Chinese Translation
预训练的机器人操作策略,如视觉-语言-动作模型(VLA)或世界-动作模型(WAM),将交互相关的度量几何留为隐式。近期空间重建的突破能够可靠地提供必要的几何信息,但其特征描述的是局部形状,而未说明其相对于机器人的位置。如何最好地将这些特征传递给预训练策略仍待解决。我们提出Spatial Grafting,一种通用、轻量的空间模块,将冻结的重建特征与度量化的、机器人相对的几何绑定。Spatial Grafting构建度量锚定的空间令牌,并通过交叉注意力将其注入流匹配动作专家,而不修改宿主的感知通路,因此宿主保留其预训练的全部优势。我们对其进行了比任何我们所比较的几何感知策略更广泛的评估:一种嫁接架构,无需针对每个宿主重新设计,在两个VLA和两个WAM上,跨越四个仿真基准(涵盖短时域操作、视觉鲁棒性、杂乱环境和长时域移动操作),并在三个真实机器人平台上(包括单臂和双臂配置)进行评估。在RoboTwin 2.0(一个双臂操作基准)上,该嫁接方法提升了所有VLA和WAM宿主。嫁接后的$\pi_{0.5}$在干净和随机场景中分别提升11.3%和15.6%,达到94.0%和92.4%,超过了已发表的最强3D条件策略WAM4D(93.8%和89.9%)。随着时域延长,优势扩大:在BEHAVIOR-1K(一个按平均任务进度评分的双臂移动操作挑战)的任务上,它在六个任务中的五个上超过了2025年挑战赛冠军,最多高出0.47 Q-score,并在两者均报告的三个任务上平均超过了一个地图条件空间策略。
cs.RO / 167 / 2609.35264

Repair Before You Fuse: Frozen-Host Adaptation for Corrupted-but-Present Sensors

先修复,再融合:面向受损但仍存在传感器的冻结宿主自适应
Thai, Gia-Huy, Ly, Quang-Thinh, Phan, Anh-Minh, Dang, Tuan
Abstract
Camera-LiDAR detectors can continue to consume unreliable features even when both sensors remain present, synchronized, and calibrated. We introduce \emph{Boundary Feature Repair} (BFR), a frozen-host adaptation framework that learns task-supervised residual corrections at modality interfaces the detector already consumes. BFR-C repairs each camera feature level read by fusion, whereas BFR-L aligns host-conditioned LiDAR candidates to a selected boundary and routes site-wise innovations relative to the frozen anchor. Their jointly trained composition is BFR-CL. Zero-initialized per-channel scales make every variant an exact detector-level identity before optimization; only the repair modules train, while the encoders, fusion consumer, router, detection head, and host normalization statistics remain fixed. At inference, BFR requires neither clean references, corruption metadata, temporal history, nor online updates. Across the complete 20-corruption, five-severity KITTI-C grid, BFR-C reduces RCE from $14.07$ to $11.92$ on MVX-Net and from $14.29$ to $11.00$ on Focals Conv-F relative to their reproduced frozen baselines. On the latter host, BFR-L raises AP$_{\mathrm{cor}}$ from $73.65$ to $74.48$, while BFR-CL reaches $77.01$ AP$_{\mathrm{cor}}$ and $10.46$ RCE with $86.02$ clean AP. On nuScenes-R, BFR-CL raises the reproduced MoME baseline's mAP robustness ratio from $80.1$ to $81.4$. These results establish boundary repair as a targeted retrofit for corrupted-but-present sensing without retraining the deployed detector.
Chinese Translation
相机-LiDAR 检测器可能继续使用不可靠的特征,即使两个传感器仍然存在、同步且已标定。我们提出边界特征修复(Boundary Feature Repair,BFR),一种冻结宿主自适应框架,其在检测器已使用的模态接口处学习任务监督的残差校正。BFR-C 修复融合所读取的每个相机特征层级,而 BFR-L 将宿主条件化的 LiDAR 候选对齐到选定边界,并相对于冻结锚点路由逐位置新息(innovations)。它们联合训练的组合为 BFR-CL。零初始化的逐通道缩放使每个变体在优化前成为精确的检测器级恒等映射;只有修复模块训练,而编码器、融合消费器、路由器、检测头和宿主归一化统计量保持固定。在推理时,BFR 既不需要干净参考、损坏元数据、时序历史,也不需要在线更新。在完整的 20 种损坏、5 个严重程度的 KITTI-C 网格上,相对于其复现的冻结基线,BFR-C 在 MVX-Net 上将 RCE 从 $14.07$ 降至 $11.92$,在 Focals Conv-F 上从 $14.29$ 降至 $11.00$。在后一宿主上,BFR-L 将 AP$_{\mathrm{cor}}$ 从 $73.65$ 提升至 $74.48$,而 BFR-CL 达到 $77.01$ AP$_{\mathrm{cor}}$ 和 $10.46$ RCE,同时干净 AP 为 $86.02$。在 nuScenes-R 上,BFR-CL 将复现的 MoME 基线的 mAP 鲁棒性比率从 $80.1$ 提升至 $81.4$。这些结果确立了边界修复作为一种针对受损但仍存在感知的定向改造,无需重新训练已部署的检测器。
cs.RO / 168 / 2609.35267

GuardPIBT: Counterfactually Gated Neural Guidance for Ultra-Large-Scale 3D Multi-Agent Path Finding

GuardPIBT:用于超大规模3D多智能体路径规划的反事实门控神经引导
Zhou, Yuan, Hou, Zhenyu, Xu, Guangtong, Ji, Xiaoqiang, Tang, Yuqing, Hou, Jialiang, Gao, Fei
Abstract
Large-scale 3D multi-agent path finding becomes increasingly difficult under dense traffic. Priority Inheritance with Backtracking (PIBT) scales well, but its one-step goal-directed ordering may become insufficient under dense interactions and large-scale congestion. We present GuardPIBT, which augments rather than replaces the PIBT executor: neural predictions only propose residual reorderings of PIBT's native candidates, while final actions remain determined by PIBT. First, local graph attention models nearby interactions, while global source--goal transport features provide population-level coordination context for candidate reordering. Second, a counterfactual group gate filters reorderings whose closed-loop effects may degrade coordination. Third, for ultra-large populations, population-adaptive grouping preserves decision granularity, asynchronous cached inference amortizes neural computation, and selective repair resolves long-tail agents. PIBT retains validity checking, priority inheritance, and backtracking throughout. Experiments with up to 100,000 agents demonstrate reliable completion across 2D and 3D environments, including all three 100,000-agent warehouse runs with zero audited graph violations. The project website is available at {\color{magenta}\texttt{https://guardpibt.github.io/GuardPIBT/}}.
Chinese Translation
在密集交通下,大规模3D多智能体路径规划变得越来越困难。带回溯的优先级继承(PIBT)具有良好的扩展性,但在密集交互和大规模拥堵下,其单步目标导向的排序可能变得不足。我们提出了GuardPIBT,它增强而非取代PIBT执行器:神经预测仅对PIBT原生候选提出残差重排序,而最终动作仍由PIBT决定。首先,局部图注意力对邻近交互进行建模,而全局源-目标传输特征为候选重排序提供群体级协调上下文。其次,反事实群组门控过滤掉那些闭环效应可能降低协调性的重排序。第三,对于超大规模群体,群体自适应分组保持决策粒度,异步缓存推理分摊神经计算,选择性修复解决长尾智能体。PIBT自始至终保留有效性检查、优先级继承和回溯。最多100,000个智能体的实验表明,在2D和3D环境中都能可靠完成,包括所有三次100,000个智能体的仓库运行,且审计到的图违规为零。项目网站见{\color{magenta}\texttt{https://guardpibt.github.io/GuardPIBT/}}。
cs.RO / 169 / 2609.35318

DexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library

DexAgent:一个具有自进化工具库的智能体式 Human2Sim2Robot 灵巧操作框架
Wang, Youhui, Li, Yunzhu, Fei-Fei, Li, Wu, Jiajun, Huang, Huang
Abstract
Human videos offer a scalable source of demonstrations for dexterous robot manipulation. However, existing human-to-simulation-to-robot (Human2Sim2Robot) pipelines rely on predefined procedures that struggle to accommodate diverse object properties and interactions, particularly those involving articulated and deformable objects. We introduce DexAgent, an agentic Human2Sim2Robot framework that converts a single egocentric human video and a task prompt into physically grounded robot trajectories for policy training. It operates through four stages: semantic understanding of human videos, property-based simulation reconstruction, robot trajectory optimization, and robot data generation. At each stage, DexAgent adapts its approach to the task and object properties by selecting suitable skills from its tool library or developing new ones when needed. Property-specific verifiers assess stage outcomes for physical validity and task-specific requirements and provide feedback for refinement, preventing error propagation through the workflow. This adaptive, verification-guided process allows DexAgent to process diverse objects and long-horizon tasks. In the final stage, DexAgent varies object and robot states in simulation to generate diverse robot trajectories from a single human video, then retextures the rendered observations to facilitate sim-to-real transfer. Newly developed skills and verifiers are retained in its tool library, making it self-evolving to accumulate reusable capabilities. This reduces processing time as DexAgent encounters more human videos. Across eleven real-world tasks, policies trained with DexAgent-generated data achieve a 3.5x higher success rate than competing baselines. Project website: https://dexagent123.github.io/.
Chinese Translation
人类视频为灵巧机器人操作提供了可扩展的演示来源。然而,现有的人类到仿真到机器人(Human2Sim2Robot)流程依赖于预定义的过程,难以适应多样化的物体属性和交互,特别是涉及铰接和可变形物体的那些。我们提出了 DexAgent,一个智能体式的 Human2Sim2Robot 框架,它将单个第一人称人类视频和任务提示转换为物理上合理的机器人轨迹,用于策略训练。它通过四个阶段运作:人类视频的语义理解、基于属性的仿真重建、机器人轨迹优化和机器人数据生成。在每个阶段,DexAgent 通过从其工具库中选择合适的技能或在需要时开发新技能,使其方法适应任务和物体属性。特定属性的验证器评估阶段结果的物理有效性和任务特定要求,并提供反馈以进行改进,防止错误在工作流中传播。这种自适应的、验证引导的过程使 DexAgent 能够处理多样化的物体和长时程任务。在最后阶段,DexAgent 在仿真中改变物体和机器人状态,从单个人类视频生成多样化的机器人轨迹,然后对渲染的观察结果重新纹理化,以促进仿真到现实的迁移。新开发的技能和验证器保留在其工具库中,使其能够自我进化以积累可重用的能力。随着 DexAgent 遇到更多的人类视频,这减少了处理时间。在十一个真实世界任务中,使用 DexAgent 生成的数据训练的策略比竞争基线实现了 3.5 倍的成功率。项目网站:https://dexagent123.github.io/。
cs.RO / 170 / 2609.35375

From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations

从像素到位姿:基于人类演示的以物体为中心的工具操作学习
Wang, Bangjun, Wu, Longyan, Wei, Yukun, Shao, Shenghe, Huang, Chaoyi, Cui, Wenze, Xu, Zetong, Wu, Hanlin, Chen, Long, Ma, Yi, Li, Hongyang
Abstract
Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although current state-of-the-arts excel at long-horizon tasks, they struggle with the delicate and precise control required for complex tool manipulation. To overcome these limitations, we introduce P2P-T, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations. P2P-T bridges the cognitive and physical execution gap through a two-stage approach. First, pretraining an object-centric world model to extract stable pose priors; second, integrating these priors into an efficient, pose-aware low-level policy. By utilizing a robust automated data processing pipeline powered by modern foundation models, P2P-T completely bypasses the need for human-robot aligned data. This reduces overall training overhead drastically. With minimal per-task fine-tuning, our framework achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently remain out of reach for standard large-scale pretrained models.
Chinese Translation
扩大机器人操作的规模主要受限于真实世界机器人数据的稀缺。尽管近期方法利用人类视频演示来缓解这一短缺,但它们仍然计算成本高昂,并且仍依赖配对的人-机器人数据来进行域对齐。尽管当前最先进方法在长时程任务上表现出色,但它们在复杂工具操作所需的精细且精确控制上仍面临困难。为克服这些局限,我们提出 P2P-T,即从像素到位姿的工具操作(from Pixel to Poses for Tool Manipulation),这是一个数据高效、以物体为中心的框架,可直接从人类演示中学习工具使用。P2P-T 通过两阶段方法弥合认知与物理执行之间的鸿沟。首先,预训练一个以物体为中心的世界模型以提取稳定的位姿先验;其次,将这些先验集成到高效、位姿感知的底层策略中。通过利用由现代基础模型驱动的鲁棒自动化数据处理流水线,P2P-T 完全绕过了对人类-机器人对齐数据的需求。这大幅降低了整体训练开销。仅需最少的逐任务微调,我们的框架在复杂真实世界工具操作任务上的执行性能相比之前的最先进技术提升了73%,而这些任务目前仍是标准大规模预训练模型无法企及的。
cs.RO / 171 / 2609.35415

Adaptive Safety Filtering for Frozen ACC Policies via Conformal Residual Calibration

面向冻结ACC策略的基于共形残差校准的自适应安全滤波
Zhou, Zhiruo, Li, Rigaudiere Z., Xiwen, Chen, Chen, Yucheng, Zhu, Xiaojun, Liu, Houde
Abstract
Frozen adaptive cruise control (ACC) policies can violate constraints when deployment dynamics differ from their training conditions. We propose residual-aware conformal action filtering (RACF), which calibrates residuals of a fixed nominal predictor and converts their quantile into an operating margin for finite-model action projection. Completed transitions update margins and candidate selection without retraining the policy. In a registered comparison over 2,400 controller-trial units, Adaptive RACF achieves 94.3% episode safety, improving by 19.9 percentage points over the evaluated nominal CBF-QP baseline while reducing projection frequency from 8.11% to 6.63%. A controlled study isolates a 4.54-point improvement from residual-margin injection. In a separate matched-hardware evaluation, Adaptive reduces mean amortized rollout time by 21.2% relative to Robust CBF-QP, with 161/180 versus 170/180 safe episodes. We characterize conditions linking one-step residual coverage to constraint satisfaction and quantify the observed safety-computation trade-offs.
Chinese Translation
冻结的自适应巡航控制(ACC)策略在部署动力学与其训练条件不同时可能违反约束。我们提出残差感知共形动作滤波(RACF),其校准固定标称预测器的残差,并将残差分位数转换为用于有限模型动作投影的运行裕度。已完成的转移会更新裕度和候选选择,而无需重新训练策略。在覆盖2,400个控制器-试验单元的注册比较中,自适应RACF实现了94.3%的回合安全性,相比所评估的标称CBF-QP基线提高了19.9个百分点,同时将投影频率从8.11%降至6.63%。一项对照研究分离出残差裕度注入带来的4.54个百分点改进。在一项单独的匹配硬件评估中,自适应方法相对于鲁棒CBF-QP将平均摊销推演时间减少了21.2%,安全回合数为161/180,而后者为170/180。我们刻画了将一步残差覆盖与约束满足联系起来的条件,并量化了观察到的安全性-计算权衡。
cs.RO / 172 / 2609.35431

Memory in the Sky: Low-Altitude Question Answering with Multi-Agent Memory Aggregation

天空中的记忆:基于多智能体记忆聚合的低空问答
Li, Chengyang, Wan, Yujie, Wang, Shuai, Ye, Kejiang, Yuan, Weijie, Zhou, Boyu, Wu, Yik-Chung, Xu, Chengzhong, Arslan, Huseyin
Abstract
This paper studies low-altitude question answering (LAQA), in which distributed unmanned aerial vehicle (UAV) memories are aggregated at a ground server to answer questions about observations over a long horizon. Unlike conventional resource allocation based on sensing, communication, control, or computation metrics, LAQA requires an explicit measure of memory value. We propose a generative adversarial exam (GAE) that uses forward simulation to evaluate memory retrieval and exam scores to quantify memory quality. This enables the downstream QA value of candidate memories to be measured and optimized without accessing the internal mechanisms of the black-box captioning, retrieval, and reasoning pipeline. Building on this metric, we develop a memory-centric (MemCen) framework that jointly selects UAVs and allocates transmit power to maximize memory quality under communication constraints. In the noise-limited regime, we derive a QoM-aware capped water-filling law that explicitly connects task utility with physical-layer power allocation. We further develop penalty successive optimization (PSO) and learning to memorize (L2M) solvers. MemCen achieves QA accuracies of 92.4% and 84.0% in CARLA Town04 and Town05 under static and dynamic communication conditions, respectively. In real-world experiments, MemCen achieves 88.5% QA accuracy on the panoramic multi-agent system (PMAS) benchmark. Finally, UAV-to-robot-dog demonstrations further validate the practical utility of the acquired memories for environmental understanding and navigation.
Chinese Translation
本文研究低空问答(LAQA),其中分布式无人机(UAV)记忆在地面服务器处聚合,以回答关于长时程观测的问题。与基于感知、通信、控制或计算指标的传统资源分配不同,LAQA需要对记忆价值进行显式度量。我们提出一种生成对抗考试(GAE),利用前向仿真评估记忆检索,并利用考试分数量化记忆质量。这使得无需访问黑盒式描述生成、检索和推理流水线的内部机制,即可度量和优化候选记忆的下游问答价值。基于该度量,我们开发了一种以记忆为中心(MemCen)的框架,在通信约束下联合选择UAV并分配发射功率,以最大化记忆质量。在噪声受限情形下,我们推导出一种QoM感知的上限注水法则,显式地将任务效用与物理层功率分配联系起来。我们进一步开发了惩罚逐次优化(PSO)和学会记忆(L2M)求解器。MemCen在静态和动态通信条件下,分别在CARLA Town04和Town05中达到92.4%和84.0%的问答准确率。在真实世界实验中,MemCen在全景多智能体系统(PMAS)基准上达到88.5%的问答准确率。最后,无人机到机器狗演示进一步验证了所获取记忆在环境理解和导航中的实际效用。
cs.RO / 173 / 2609.35432

Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence

自进化编码智能体:从数字程序到物理世界智能
Gao, Hongcheng, Zhou, Jingjing, Zheng, Zelin, Ge, Shijia, Zhu, Jay, Wang, Yazhe, Zeng, Jianshu, Shangguan, Xuan, Wu, Di, He, Lingyu, Jia, Zhiqi, Wu, Sihang, He, Xiao
Abstract
Vision-language-action (VLA) and world-action (WAM) models map observations and instructions directly to robot actions. This directness ties a policy to training: minor layout or viewpoint changes cause failure, and instructions generalize poorly. The root cause lies in representation: task requirements, conditions, progress, and failure recovery are implicitly encoded in action sequences, making them difficult to inspect or revise. Digital coding agents offer a precedent: LLMs call tools, verify results, and revise from feedback as executable code. The same working pattern of explicit state, manageable execution, and revisable procedures underlies generalization and long-horizon execution in the physical world, letting physical experience return as reusable programs, memory, or evidence. We propose Physical Coding, representing task state and execution as code. Code as World records objects, relations, constraints, and progress; Code as Policy organizes planning, verification, recovery, and execution. We build HexaAnything, which calls perception, planning, and control tools, including VLA/WAM policies, and makes in-the-loop decisions from external feedback. Verified traces become data and memory, enabling evolution from tools and Harness to model weights, architectures, and ultimately hardware and task design. On RoboCasa365, HexaAnything improves Composite-Unseen and overall success over XR-1 VLA, and its Harness-trained HexaModel beats the base on every split, indicating code traces internalize physical execution. On PhyBench and a dual-arm AgileX robot, the agent autonomously completes physics experiments and most tabletop tasks, often faster than published results. We observe data, model, and tool self-evolution; future work targets weight internalization, autonomous redesign of architectures, languages, representations, and tasks, and deployment in manufacturing and science.
Chinese Translation
视觉-语言-动作(VLA)和世界-动作(WAM)模型将观测和指令直接映射为机器人动作。这种直接性将策略与训练绑定:轻微的布局或视角变化就会导致失败,且指令泛化能力差。根本原因在于表示:任务需求、条件、进度和失败恢复被隐含地编码在动作序列中,使其难以检查或修正。数字编码智能体提供了一个先例:大语言模型(LLMs)调用工具、验证结果,并基于反馈以可执行代码的形式进行修正。显式状态、可管理执行和可修正流程的相同工作模式,是物理世界中泛化与长时程执行的基础,使物理经验能够作为可复用程序、记忆或证据返回。我们提出物理编码(Physical Coding),将任务状态和执行表示为代码。代码即世界(Code as World)记录对象、关系、约束和进度;代码即策略(Code as Policy)组织规划、验证、恢复和执行。我们构建了 HexaAnything,它调用感知、规划和控制工具(包括 VLA/WAM 策略),并基于外部反馈进行在环决策。经过验证的轨迹成为数据和记忆,使进化能够从工具和 Harness 到模型权重、架构,并最终到硬件和任务设计。在 RoboCasa365 上,HexaAnything 在 Composite-Unseen 和总体成功率上优于 XR-1 VLA,其经 Harness 训练的 HexaModel 在每个划分上都优于基线,表明代码轨迹将物理执行内化。在 PhyBench 和双臂 AgileX 机器人上,该智能体自主完成物理实验和大多数桌面任务,且通常比已发表结果更快。我们观察到数据、模型和工具的自进化;未来工作旨在实现权重内化,自主重新设计架构、语言、表示和任务,并部署于制造业和科学领域。
cs.RO / 174 / 2609.35439

Revision, Not Restart: Revisable Visual Plans for Closed-Loop World-Action Models

修订而非重启:用于闭环世界-动作模型的可修订视觉规划
Liu, Pengyiang, Niu, Junbo, Zheng, Wenhao, Chen, Xinchen, Li, Canyu, Shi, Zhongyue, Xie, Jiahao, Liu, Si
Abstract
World-action models use predicted visual futures to condition robot actions, yet execution feedback can invalidate parts of a prediction while leaving its task structure useful. We propose Revisable Temporal Planning (RTP), which maintains the visual future as a persistent action condition and revises it after feedback. Its central mechanism is a learned revision bridge: it resumes an intermediate state saved during visual generation and adapts its continuation to current observations. Visual and action supervision connect this revision to subsequent control. Time-aware history supplies observed evidence, and an adaptive policy selects retention, bridge revision, or fresh replanning from new noise before decoding the next action. On RoboMME and RMBench, RTP achieves task-averaged success rates of 48.6% and 84.8%, respectively. Matched comparisons support learned continuation; estimated checkpoint-source and action-prefix effects are positive but less precisely resolved. These results connect feedback-driven visual-plan revision to closed-loop task performance. Project Page: https://PLACEHOLDER.github.io/RTP/
Chinese Translation
世界-动作模型使用预测的视觉未来来调节机器人动作,但执行反馈可能使预测的某些部分失效,同时其任务结构仍然有用。我们提出可修订时序规划(Revisable Temporal Planning, RTP),它将视觉未来保持为持续的动作条件,并在反馈后对其进行修订。其核心机制是一个学习得到的修订桥:它恢复在视觉生成过程中保存的中间状态,并使其后续生成适应当前观测。视觉和动作监督将这一修订与后续控制连接起来。时间感知历史提供观测证据,自适应策略在解码下一个动作之前,从新噪声中选择保持、桥接修订或全新重规划。在 RoboMME 和 RMBench 上,RTP 分别取得了 48.6% 和 84.8% 的任务平均成功率。匹配比较支持学习得到的延续;估计的检查点来源和动作前缀效应为正,但精度较低。这些结果将反馈驱动的视觉规划修订与闭环任务性能联系起来。项目页面:https://PLACEHOLDER.github.io/RTP/
cs.RO / 175 / 2609.35450

Uni-VLaT: Whole-Body Tactile Adaptation of VLA Policies for Humanoid Loco-Manipulation

Uni-VLaT:面向人形机器人移动操作的 VLA 策略全身触觉适应
Wang, Zihao, Liu, Shutong, Zheng, Siqi, Cao, Liu, Chen, Ruoqi, Liu, Rundong, Yang, Yanchao, Xu, Mengdi
Abstract
Physical contact often determines how a humanoid should respond during loco-manipulation, yet vision and proprioception alone are often insufficient to characterize physical interaction, especially when the contact region is occluded. Unlike sparse force or torque measurements at predefined regions, distributed tactile sensing preserves spatially resolved contact patterns across the robot body. We therefore study how to integrate such whole-body tactile information into vision-language-action (VLA) policies for contact-rich control. Our approach, Uni-VLaT, introduces a tactile pathway whose latent state is trained not only for action generation, but also to predict future tactile, proprioceptive, and visual representations. This predictive objective builds a tactile-anchored multimodal context, encouraging a more structured understanding of the physical world. We evaluate Uni-VLaT on five real-robot tasks covering tactile-triggered locomotion, sustained physical interaction, human-robot contact, and loco-manipulation. Uni-VLaT achieves a 75% average success rate, outperforming a baseline without tactile input by 43 points and a tactile-input baseline without predictive supervision by 7 points. Across two pretrained VLA backbones, our method improves Table Sweeping by 30 points on both backbones and Back-Tap Walking by 85-90 points. Ablations further show that contextualized tactile prediction and absolute future targets are critical to performance. These results indicate that predictive tactile learning provides an effective route for extending pretrained VLA policies to whole-body physical interaction.
Chinese Translation
物理接触往往决定了人形机器人在移动操作过程中应如何响应,然而仅凭视觉和本体感觉通常不足以表征物理交互,尤其是当接触区域被遮挡时。与在预定义区域进行稀疏的力或力矩测量不同,分布式触觉传感能够保留机器人全身空间分辨的接触模式。因此,我们研究如何将这种全身触觉信息集成到视觉-语言-动作(VLA)策略中,以实现接触丰富的控制。我们的方法 Uni-VLaT 引入了一条触觉通路,其潜在状态不仅被训练用于动作生成,还用于预测未来的触觉、本体感觉和视觉表示。这一预测目标构建了以触觉为锚点的多模态上下文,促进了对物理世界更结构化的理解。我们在五个真实机器人任务上评估了 Uni-VLaT,涵盖触觉触发的运动、持续物理交互、人机接触以及移动操作。Uni-VLaT 实现了 75% 的平均成功率,比无触觉输入的基线高出 43 个百分点,比无预测监督的触觉输入基线高出 7 个百分点。在两个预训练的 VLA 主干网络上,我们的方法在两者上将 Table Sweeping 提高了 30 个百分点,将 Back-Tap Walking 提高了 85-90 个百分点。消融实验进一步表明,情境化触觉预测和绝对未来目标对性能至关重要。这些结果表明,预测性触觉学习为将预训练的 VLA 策略扩展到全身物理交互提供了一条有效途径。
cs.RO / 176 / 2609.35469

Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching

重新思考流匹配中基于条件退火的因果动作标记化
Zhang, Chenyu, Cao, Yuhang, Du, Daru, Lu, Yingxi, Shao, Jing, Chen, Ruoqu, Liu, Jiajun, Cao, Liu, Liu, Yicheng, Zhao, Hang, Xu, Mengdi
Abstract
Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process. CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens with the precision of hybrid diffusion-head architectures. This discrete bottleneck enforces knowledge insulation by design, cleanly separating high-level semantic reasoning from low-level motor execution without requiring explicit attention masking. Extensive evaluations across three simulation benchmarks and real-world robotic manipulation tasks demonstrate that CATok consistently surpasses existing tokenization methods in both reconstruction fidelity-compression tradeoff and inference efficiency, while improving VLA task success rate and training efficiency, establishing a high-performance, scalable foundation for purely autoregressive VLA systems.
Chinese Translation
自回归视觉-语言-动作 (VLA) 模型为机器人学习提供了一条可扩展的路径,然而现有的动作标记器将标记化视为一个压缩问题,产生的表示与自回归主干在语义上不对齐。我们提出 CATok,一种因果动作标记器,将标记化重新定义为一个因果结构化的生成过程。CATok 引入条件退火机制,通过逐步退火一个流匹配过程来提取动作标记:每个标记以所有先前标记为条件,并编码特定噪声水平下的残差重建信号,从而建立了一个从粗到细的因果标记空间,其生成语义与自回归建模在结构上对齐。基于多模态扩散 Transformer (MMDiT) 构建的标记条件流匹配解码器,以混合扩散头架构的精度从这些离散标记中重建连续动作块。这种离散瓶颈通过设计实现知识隔离,无需显式注意力掩码即可清晰地将高层语义推理与低层运动执行分离。在三个仿真基准和真实机器人操作任务上的广泛评估表明,CATok 在重建保真度-压缩权衡和推理效率方面均持续超越现有标记化方法,同时提高了 VLA 任务成功率和训练效率,为纯自回归 VLA 系统奠定了高性能、可扩展的基础。
cs.RO / 177 / 2609.35476

CoBrush: A Hierarchical Planning Framework for Human-Robot Co-Painting

CoBrush:人机协同绘画的分层规划框架
Qin, Dantong, Guo, Yike, Liu, Qinlin, Bozzon, Alessandro, Wang, Pan
Abstract
Embodied co-painting requires a robot to repeatedly update a shared physical canvas while human intent evolves over interaction. Existing reference-driven painters or reactive assistants are typically optimized for single-shot rendering or sketch completion, limiting their ability to sustain coherent multi-round collaboration or to construct complex, content-rich scenes over time. We present CoBrush, a hierarchical framework that formulates multi-round co-painting as a coordinated semantic, spatial, and execution process. By separating high-level intent inference from spatial grounding and stroke-level control, the system supports progressive scene development on real acrylic canvases. We evaluate the framework through real human-robot painting sessions, stress tests, and user studies. Compared to single-turn baselines, our approach achieves stronger semantic alignment, more stable spatial progression, and higher perceived plausibility of robot actions. These results demonstrate that structured multi-stage reasoning improves the coherence and robustness of interactive painting and supports the progressive development of content-rich physical artworks.
Chinese Translation
具身协同绘画要求机器人在人类意图随交互不断演变的同时,反复更新共享的物理画布。现有的参考驱动绘画器或反应式助手通常针对单次渲染或草图补全进行优化,限制了它们维持连贯的多轮协作或随时间构建复杂、内容丰富场景的能力。我们提出 CoBrush,一个分层框架,将多轮协同绘画形式化为一个协调的语义、空间和执行过程。通过将高层意图推断与空间接地和笔画级控制分离,该系统支持在真实丙烯画布上逐步发展场景。我们通过真实的人机绘画会话、压力测试和用户研究来评估该框架。与单轮基线相比,我们的方法实现了更强的语义对齐、更稳定的空间进展以及更高的机器人动作感知合理性。这些结果表明,结构化的多阶段推理提高了交互绘画的连贯性和鲁棒性,并支持内容丰富物理艺术品的逐步发展。
cs.RO / 178 / 2609.35479

Robot Tool Design from Scratch via Behavior-Aware Hierarchical Optimization

基于行为感知分层优化的机器人工具从零设计
Chen, Yinghan, Tian, Xiyao, Dai, Yizan, Li, Yuyang, Zhu, Yixin
Abstract
The ability to design a tool for a task marks a level of intelligence beyond merely understanding, selecting, or using one. Existing methods for robotic tool design typically optimize a tool's continuous shape and action within a structure that is prescribed or generated beforehand, so the structure itself stays outside the physical optimization loop. We study task-driven tool design from scratch, where tool structure, shape, and action are all derived from the desired physical outcome. Here we show that the three elements can be designed jointly by HOT, a hierarchical optimization whose upper level searches over discrete tool structures with BASS, while lower-level physical optimization evaluates their task behavior and returns milestone progress as behavioral evidence for the search, ultimately providing jointly optimized shape and action. On four tool-use tasks with distinct physical functions, HOT discovers functional structures after evaluating only a small fraction of search spaces containing up to 56 million structures, and the subsequent refinement of their geometry lowers the task loss on all tasks while preserving success, through deformations that are functionally interpretable. Once 3D printed, the tools accomplish all tasks on a real robot with the actions found in simulation. Designing tools from required physical effects, rather than a catalog of known tools, is a step toward the open-ended tool making seen in humans and animals.
Chinese Translation
为任务设计工具的能力标志着一种智能水平,超越了仅仅理解、选择或使用工具。现有的机器人工具设计方法通常在预先指定或生成的框架内优化工具的连续形状和动作,因此结构本身保持在物理优化循环之外。我们研究从零开始的任务驱动工具设计,其中工具的结构、形状和动作都源自期望的物理结果。在这里,我们展示这三个元素可以通过HOT联合设计,HOT是一种分层优化方法,其上层使用BASS在离散的工具结构上进行搜索,而下层物理优化评估它们的任务行为,并将里程碑进展作为搜索的行为证据返回,最终提供联合优化的形状和动作。在四项具有不同物理功能的工具使用任务上,HOT仅评估了包含多达5600万个结构的搜索空间的一小部分,就能发现功能结构,随后对其几何形状的细化通过功能上可解释的变形,降低了所有任务的任务损失,同时保持了成功。一旦3D打印,这些工具就能在真实机器人上以仿真中找到的动作完成所有任务。从所需的物理效果出发设计工具,而不是从已知工具目录中选择,是迈向人类和动物所展现的开放式工具制造的一步。
cs.RO / 179 / 2609.35482

ForVis: An In-Field Dataset and Benchmark for VIO Using Under-Canopy UAV Flights in Forests

ForVis:面向森林冠层下无人机飞行的VIO实地数据集与基准
Kiani, Arman, Ataei, Masoud, Gyaase, Elvis, Eiyike, Jeffrey, Weiskittel, Aaron, Chakraborty, Prabuddha, Dhiman, Vikas
Abstract
Visual-inertial Simultaneous Localization and Mapping (VI-SLAM) for UAVs remains difficult to evaluate in real forest environments, where motion, illumination changes, repetitive vegetation, and vibration can all affect estimation. We present ForVis, an in-field dataset and benchmark for evaluating VI-SLAM during UAV flight in forest environments. The dataset contains twelve flights across open meadow, above-canopy, and under-canopy conditions in each environment. In total, it provides 563.8s of flight over 1096.8m of trajectory, recorded simultaneously with an Intel RealSense D435i and an OAK-D Pro Wide together with inertial and flight-controller data. We benchmark seven open-source VI-SLAM systems over 504 runs. The results show that sensor choice has a larger effect on trajectory error than the spread between algorithms: all seven methods achieve lower median error on the OAK-D Pro than on the D435i. ForVis is intended to support evaluation of speed, accuracy and robustness for VI-SLAM in challenging forest flight.
Chinese Translation
无人机视觉惯性同时定位与建图(VI-SLAM)在真实森林环境中仍难以评估,其中运动、光照变化、重复植被和振动均可能影响估计。我们提出了ForVis,一个用于评估森林环境中无人机飞行期间VI-SLAM的实地数据集与基准。该数据集包含12次飞行,覆盖每个环境中的开阔草甸、冠层上方和冠层下条件。总共提供563.8秒的飞行数据,轨迹长度为1096.8米,使用Intel RealSense D435i和OAK-D Pro Wide同步记录,并包含惯性与飞控数据。我们在504次运行中对七个开源VI-SLAM系统进行了基准测试。结果表明,传感器选择对轨迹误差的影响大于算法之间的差异:所有七种方法在OAK-D Pro上的中位误差均低于在D435i上的中位误差。ForVis旨在支持评估VI-SLAM在具有挑战性的森林飞行中的速度、准确性和鲁棒性。
cs.RO / 180 / 2609.35493

Terrain-Aware Autonomous Planetary Exploration for Exteroceptive-Proprioceptive Mapping with Quadruped Scouts

面向四足侦察机器人外感知-本体感知建图的地形感知自主行星探索
Sanchez-Delgado, Alberto, Soares, João Carlos Virgolino, Barasuol, Victor, Semini, Claudio
Abstract
Autonomous planetary exploration requires robots to navigate unknown, uneven terrain while assessing risk, traversability, and energetic cost. Quadruped scouts are well suited for this task because they can traverse irregular surfaces and gather mobility-relevant information during locomotion. This paper presents a terrain-aware exploration framework that combines exteroceptive and proprioceptive mapping for a quadruped robot in lunar-like environments. An onboard RGB-D camera builds robot-centered elevation maps, estimates geometric traversability, and derives navigation costs for autonomous planning. In parallel, proprioceptive measurements provide interaction-aware terrain cues that complement geometry-based assessment. Local maps are incrementally registered into a global multi-layer representation, which is used by an exploration module to select targets in unexplored regions of interest. The targets are reached by an autonomous navigation system that guides collision-aware motion using the available map and cost layers. Simulation results on NVIDIA Isaac Sim show autonomous exploration, map expansion, and spatial association between terrain geometry and robot-terrain interaction. Subsequent navigation using this information exhibits lower average Cost of Transport (CoT) than initial exploration.
Chinese Translation
自主行星探索要求机器人能够在未知、崎岖的地形中导航,同时评估风险、可通行性和能量消耗。四足侦察机器人非常适合这项任务,因为它们可以穿越不规则表面并在运动过程中收集与移动相关的信息。本文提出了一种地形感知探索框架,该框架结合了外感知和本体感知建图,用于类月环境中的四足机器人。机载RGB-D相机构建以机器人中心的高程图,估计几何可通行性,并推导出用于自主规划的导航代价。同时,本体感知测量提供交互感知的地形线索,以补充基于几何的评估。局部地图被增量注册到全局多层表示中,探索模块利用该表示在未探索的感兴趣区域中选择目标。自主导航系统利用可用的地图和代价层引导碰撞感知运动,从而到达这些目标。在NVIDIA Isaac Sim上的仿真结果展示了自主探索、地图扩展以及地形几何与机器人-地形交互之间的空间关联。使用此信息的后续导航表现出比初始探索更低的平均运输成本(CoT)。
cs.RO / 181 / 2609.35516

Inspection-SPARS: Task-Oriented Sparse Roadmaps for Inspection Planning

Inspection-SPARS:用于巡检规划的面向任务稀疏路线图
Morgan, Adir, Salzman, Oren, Solovey, Kiril
Abstract
Inspection planning seeks a minimum-length collision-free robot tour that observes a given set of points of interest (POIs). Sampling-based methods reduce this continuous problem to a graph inspection planning (GIP) problem over a discrete roadmap, which is then solved using combinatorial solvers. Dense roadmaps capture diverse inspection viewpoints and motion shortcuts, and thus admit higher-quality solutions, but they induce large combinatorial search spaces on which state-of-the-art GIP solvers struggle to find good solutions within practical time budgets. Roadmap sparsification---restructuring a dense roadmap into a compact representation that preserves connectivity and path lengths---can alleviate this burden. However, existing sparsification approaches are either agnostic to the underlying inspection task, or strive to ensure coverage of the POIs without accounting for the quality of the resulting inspection plan. We present Inspection-SPARS, which is, to our knowledge, the first inspection-roadmap sparsifier with POI coverage and path-quality guarantees relative to the dense roadmap. To this end, we generalize the SPARS framework, a popular task-agnostic sparsifier, from purely geometric criteria to task-oriented ones, introducing an inspection-aware vertex admission mechanism that treats POI coverage as a first-class sparsification criterion alongside connectivity and path quality. Experiments in realistic 3D environments show that Inspection-SPARS reduces vertex and edge counts by 4-8x while preserving coverage, allowing the GIP solver to compute tours up to 25% shorter than with the dense roadmap or state-of-the-art inspection roadmap. More broadly, Inspection-SPARS shows that sparsification can be made task-aware without sacrificing guarantees on solution quality.
Chinese Translation
巡检规划旨在寻找一条最短的无碰撞机器人巡回路径,以观测给定的一组兴趣点(POIs)。基于采样的方法将这一连续问题简化为离散路线图上的图巡检规划(GIP)问题,然后使用组合求解器进行求解。密集路线图能够捕捉多样的巡检视角和运动捷径,因此可获得更高质量的解决方案,但它们会导致庞大的组合搜索空间,即使最先进的 GIP 求解器也难以在实际时间预算内找到好的解决方案。路线图稀疏化——将密集路线图重构为保持连通性和路径长度的紧凑表示——可以减轻这一负担。然而,现有的稀疏化方法要么与底层巡检任务无关,要么致力于确保 POIs 的覆盖而不考虑所得巡检计划的质量。我们提出了 Inspection-SPARS,据我们所知,这是第一个相对于密集路线图具有 POI 覆盖和路径质量保证的巡检路线图稀疏化器。为此,我们将 SPARS 框架——一种流行的与任务无关的稀疏化器——从纯几何标准推广到面向任务的标准,引入了一种巡检感知的顶点准入机制,该机制将 POI 覆盖作为与连通性和路径质量并重的主要稀疏化标准。在真实 3D 环境中的实验表明,Inspection-SPARS 在保持覆盖的同时将顶点和边的数量减少了 4-8 倍,使得 GIP 求解器能够计算出比使用密集路线图或最先进的巡检路线图短至多 25% 的巡检路径。更广泛地说,Inspection-SPARS 表明,稀疏化可以在不牺牲解决方案质量保证的情况下实现任务感知。
cs.RO / 182 / 2609.35570

EdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation Model

EdgeVLN:运行时感知的可部署量化视觉语言导航模型
Jonna, Rithvik, Namgung, Man, Gurram, Aakash, Mohsenin, Tinoosh
Abstract
Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preserving navigation behavior. We introduce EdgeVLN, a runtime-aware, deployment-ready quantized VLN model that closes this gap. EdgeVLN combines a quantized StreamVLN model with Latent Trajectory Termination Extractor (LATTE), a lightweight causal transformer that improves real-time stopping by predicting a Stop Action verifier rank. Both execute through our llama.cpp VLN driver, which reconstructs streaming context and prunes memory tokens on-board. We characterize a pretrained StreamVLN backbone across weight quantization from 8 to 2 bits and multiple inference runtimes to identify a feasible operating point. LATTE reuses backbone hidden states within the budget freed by quantization, requiring neither a second vision encoder nor an additional backbone forward pass. We evaluate six backbone precisions and seven candidate stop heads on BF16 and IQ4 NL across all 1,839 R2R VLN-CE val-unseen episodes. We measure success rate (SR) in simulation and latency, energy, and resident memory on an NVIDIA Jetson Orin NX 16 GB. LATTE achieves our highest SR, 58.02 percent on the deployed 4-bit model, exceeding the BF16 baseline with only 0.013 s additional latency per navigation step. Four-bit formats achieve nearly identical SR, but step energy varies 36.8 times by execution path. Only IQ4 NL under our VLN driver fits the board, using 11.35 GB resident memory while running 20.8 times faster and using 13.3 times less energy than storage-streamed BF16. INT2 collapses. Runtime selection, memory-token pruning, and quantization are essential for efficient edge deployment.
Chinese Translation
视觉语言导航(VLN)模型表现良好,但面向计算资源丰富的平台,限制了在内存和功率受限的机器人边缘设备上的部署。仅靠压缩无法确定VLN模型是否在保持导航行为的同时满足边缘平台的内存、延迟和能耗预算。我们提出EdgeVLN,一个运行时感知、可部署的量化VLN模型,弥补了这一空白。EdgeVLN将量化的StreamVLN模型与潜在轨迹终止提取器(LATTE)相结合,后者是一个轻量级因果Transformer,通过预测停止动作验证器排名来改进实时停止。二者都通过我们的llama.cpp VLN驱动程序执行,该驱动程序在板载重建流式上下文并修剪内存令牌。我们表征了预训练的StreamVLN主干在8至2位权重量化以及多种推理运行时下的表现,以确定可行的操作点。LATTE在量化释放的预算内重用主干隐藏状态,既不需要第二个视觉编码器,也不需要额外的主干前向传递。我们在所有1,839个R2R VLN-CE val-unseen回合上,评估了BF16和IQ4 NL上的六种主干精度和七种候选停止头。我们测量了仿真中的成功率(SR),以及NVIDIA Jetson Orin NX 16 GB上的延迟、能耗和驻留内存。LATTE在我们部署的4位模型上达到了最高的SR,为58.02%,超过BF16基线,每个导航步骤仅增加0.013秒延迟。4位格式实现了几乎相同的SR,但步骤能耗因执行路径不同而相差36.8倍。只有我们的VLN驱动程序下的IQ4 NL适合该板,使用11.35 GB驻留内存,运行速度比存储流式BF16快20.8倍,能耗低13.3倍。INT2崩溃。运行时选择、内存令牌修剪和量化对于高效的边缘部署至关重要。
cs.RO / 183 / 2609.35575

F4R: Failure-Driven Recognition, Reconstruction, Refinement, and Redeployment for Continual Robot Self-Improvement

F4R:面向持续机器人自我改进的失败驱动识别、重建、精炼与重新部署
Yu, Zhuoyuan, Wang, Jiacheng, Liu, Tianle, Ren, Yihua, Yu, Peng, Bai, Chen, Zhang, Ziheng, Jia, Yufei, Jia, Jindou, Zhang, Yuhang, Zhang, Xinrui, Yujing, Shang, Chen, Yuxiang, Zhou, Chuhao, Wang, Tiancai, Yang, Jianfei
Abstract
The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world demonstrations of newly encountered failures. However, this process is costly, inefficient, potentially unsafe, and difficult to scale. To address this challenge, we propose Failure for Rising (F4R), a failure-driven real-to-sim-to-real closed-loop learning framework that converts real-world failures into targeted policy improvement. F4R first uses an agent to automatically identify and diagnose failures from rollouts. It reconstructs each failure as an interactive, object-centric table-top environment that preserves the task-relevant spatial and physical conditions. The policy is then refined through failure-conditioned sim-real co-training followed by targeted reinforcement learning in the reconstructed environments. The improved policy is subsequently redeployed, while newly observed failures are continuously fed back into the next reconstruction and learning cycle. Real-world evaluations on four manipulation tasks show that F4R achieves 93.75% In-Distribution and 90.0% Out-of-Distribution (OOD) success, outperforming the budget-matched Targeted BC baseline by 18.75 percentage points under OOD conditions without collecting additional real-world corrective demonstrations.
Chinese Translation
当前视觉-语言-动作模型在真实世界中的性能从根本上受限于专家演示的覆盖范围有限以及对物理交互的理解不足。一个常见的补救方法是收集新遇到的失败的额外真实世界演示。然而,这一过程成本高昂、效率低下、可能不安全且难以扩展。为了应对这一挑战,我们提出了Failure for Rising (F4R),一个失败驱动的real-to-sim-to-real闭环学习框架,将真实世界的失败转化为有针对性的策略改进。F4R首先使用一个智能体从rollouts中自动识别和诊断失败。它将每个失败重建为一个交互式的、以对象为中心的桌面环境,保留任务相关的空间和物理条件。然后,通过以失败为条件的sim-real协同训练,以及在重建环境中有针对性的强化学习,对策略进行精炼。改进后的策略随后被重新部署,而新观察到的失败则持续反馈到下一个重建和学习循环中。在四个操作任务上的真实世界评估表明,F4R在分布内达到了93.75%的成功率,在分布外(OOD)达到了90.0%的成功率,在OOD条件下,无需收集额外的真实世界纠正演示,就比预算匹配的Targeted BC基线高出18.75个百分点。
cs.RO / 184 / 2609.35619

CollisionSplatting: Collision-Aware Motion Planning in 3DGS Scenes with Image-Conditioned Objectives and Adjustable Conservatism

CollisionSplatting:在3DGS场景中具有图像条件目标和可调保守性的碰撞感知运动规划
Khorrambakht, R., Ortiz-Haro, Joaquim, Weiss, Stephan, Righetti, Ludovic
Abstract
Incorporating dense visual information into motion planning remains challenging, as geometric planners rely on abstracted scene representations that discard visual richness, while learned visual models often lack geometric interpretability and computational efficiency. This paper introduces CollisionSplatting, a simple, modular, GPU-accelerated, probability-inspired distance metric with tunable conservatism that operates directly on standard 3D Gaussian Splatting (3DGS) scenes. When combined with learned image-conditioned reward functions, this metric enables joint geometric and visual planning by unifying collision-aware costs with image-space objectives. We integrate the metric into GPU-accelerated Model Predictive Path Integral (MPPI) and Rapidly-Exploring Random Tree (RRT) planners, and show on-par or better collision-classification performance compared to representative baselines while achieving substantially higher collision-checking throughput and significantly lower VRAM usage. Finally, we demonstrate the effectiveness of our metric in real-world vision-guided navigation and manipulation tasks, highlighting 3DGS as a practical bridge between rich perception and real-time motion planning.
Chinese Translation
将密集视觉信息融入运动规划仍然具有挑战性,因为几何规划器依赖于丢弃视觉丰富性的抽象场景表示,而学习的视觉模型通常缺乏几何可解释性和计算效率。本文介绍了CollisionSplatting,一种简单、模块化、GPU加速、受概率启发的距离度量,具有可调保守性,可直接在标准3D高斯泼溅(3DGS)场景上运行。当与学习的图像条件奖励函数结合时,该度量通过将碰撞感知成本与图像空间目标统一起来,实现了联合几何和视觉规划。我们将该度量集成到GPU加速的模型预测路径积分(MPPI)和快速探索随机树(RRT)规划器中,并展示了与代表性基线相比相当或更好的碰撞分类性能,同时实现了显著更高的碰撞检查吞吐量和显著更低的VRAM使用量。最后,我们展示了我们的度量在现实世界视觉引导导航和操作任务中的有效性,突出了3DGS作为丰富感知和实时运动规划之间的实用桥梁。
cs.RO / 185 / 2609.35651

Denoising Multi-Robot Trajectories

去噪多机器人轨迹
Zhang, Yuhao, Okumura, Keisuke, Shankar, Ajay, Prorok, Amanda
Abstract
Multi-robot trajectory planning is a fundamental problem in multi-robot coordination but remains computationally challenging due to its nonconvex, multimodal, and high-dimensional nature. This work builds upon D4orm, a dynamics-aware diffusion-denoising framework, and develops a family of planning architectures for diverse operational requirements. Unlike conventional numerical optimization methods, D4orm employs sampling-based optimization to generate solution trajectories through massively parallel sampling, leveraging modern computing architectures such as GPUs. Its diffusion-denoising structure iteratively optimizes \textit{deformations} to candidate control trajectories, providing an efficient and versatile paradigm for generating kinodynamically feasible and conflict-free trajectories. Using D4orm as the building block for advanced planners, we present a decoupled planner for improved scalability, an online receding-horizon planner with feedback control, and a distributed planner for resource-constrained settings. Evaluations with differential-drive and holonomic robots in 2D and 3D environments demonstrate that D4orm-based approaches find high-quality solutions faster and more reliably than other sampling-based optimization methods, such as MPPI, as well as a learned diffusion-model-based method. We further demonstrate zero-shot deployment on ten real quadrotors with obstacles, large-scale deconfliction with 100 simulated robots, and fully onboard distributed `lifelong' operation with six ground robots. Overall, these results establish diffusion denoising as a scalable and reliable framework for multi-robot coordination. Code and video: https://github.com/proroklab/d4orm
Chinese Translation
多机器人轨迹规划是多机器人协调中的一个基本问题,但由于其非凸、多模态和高维的特性,在计算上仍然具有挑战性。这项工作基于D4orm,一个动力学感知的扩散-去噪框架,并针对多样化的操作需求开发了一系列规划架构。与传统的数值优化方法不同,D4orm采用基于采样的优化,通过大规模并行采样生成解轨迹,利用现代计算架构如GPU。其扩散-去噪结构迭代地优化候选控制轨迹的变形,为生成动力学可行且无冲突的轨迹提供了一种高效且通用的范式。以D4orm作为高级规划器的构建模块,我们提出了一个解耦规划器以提高可扩展性,一个带有反馈控制的在线滚动时域规划器,以及一个用于资源受限环境的分布式规划器。在2D和3D环境中对差速驱动和全向机器人进行的评估表明,基于D4orm的方法比其他基于采样的优化方法(如MPPI)以及基于学习的扩散模型方法更快、更可靠地找到高质量解决方案。我们进一步展示了在十架真实四旋翼无人机上带有障碍物的零样本部署,100个模拟机器人的大规模冲突消解,以及六台地面机器人的完全机载分布式'终身'操作。总体而言,这些结果确立了扩散去噪作为多机器人协调的可扩展且可靠的框架。代码和视频:https://github.com/proroklab/d4orm
cs.RO / 186 / 2609.35652

MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining

MM-ABC:通过感知、协调与想象迈向通用移动操作
Liang, Qiwei, Chen, Guangyu, Zhu, Shaolong, Xiao, Zikuan, Lu, Jinxuan, Xie, Yifan, Xu, Renjing, Ding, Wenbo, Chen, Tianxing
Abstract
Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing approaches strengthen geometry through explicit 3D representations or predictive world models, and often decouple mobility and manipulation into separate action streams. We argue that effective mobile manipulation requires not only decoupling, but also representations that support efficient cross-stream collaboration. We present MM-ABC, a foundation model built around Seeing, Coordinating, and Imagining Arm-Base Collaboration. MM-ABC combines sparse multi-level VLM features for spatial perception; a training-only future branch that uses world imagination and geometric intent as extra supervision, strengthening perception and manipulation-intent prediction and improving the overall learning signal; and MM-APT, which coordinates separate manipulation and mobility streams through masked joint attention and clean-action x-prediction. In controlled ablations, replacing clean-action prediction with velocity prediction lowers success on RoboCasa365 composite-seen tasks from 32.8% to 29.2%, and removing future supervision or multilevel conditioning causes larger drops. We pretrain MM-ABC on 5,000+ hours of heterogeneous robot data spanning 400K+ episodes, 12 datasets, and 17 embodiments. Experiments cover EBench, RoboCasa365, ManiSkill-HAB, LIBERO, LIBERO-Plus, and real-world mobile manipulation. MM-ABC achieves 44.71% success on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus without perturbation training, and 83% mean success on five real-world tasks.
Chinese Translation
移动操作通过使可达区域本身可控,将机器人交互扩展到固定的运动学工作空间之外。这种灵活性带来了两个核心挑战:连续自我运动下的空间基础感知,以及异构机械臂和基座动作的协调控制。现有方法通过显式3D表示或预测性世界模型来增强几何,并通常将移动性和操作性解耦为独立的动作流。我们认为,有效的移动操作不仅需要解耦,还需要支持高效跨流协作的表示。我们提出MM-ABC,一个围绕感知、协调和想象臂-基协作构建的基础模型。MM-ABC结合了用于空间感知的稀疏多级VLM特征;一个仅训练的未来分支,使用世界想象和几何意图作为额外监督,增强感知和操作意图预测并改善整体学习信号;以及MM-APT,它通过掩码联合注意力和clean-action x-prediction来协调独立的操作和移动流。在受控消融实验中,将干净动作预测替换为速度预测会使RoboCasa365复合已见任务的成功率从32.8%降至29.2%,而去除未来监督或多级条件会导致更大的下降。我们在超过5,000小时的异构机器人数据上预训练MM-ABC,涵盖超过40万条回合、12个数据集和17种具身。实验涵盖EBench、RoboCasa365、ManiSkill-HAB、LIBERO、LIBERO-Plus以及真实世界移动操作。MM-ABC在EBench上达到44.71%的成功率,在RoboCasa365上达到61.2%,在LIBERO上达到99.1%,在未经扰动训练的LIBERO-Plus上达到82.8%,在五个真实世界任务上平均成功率为83%。
cs.RO / 187 / 2609.35690

Agent Priors-guided Policy Learning

智能体先验引导的策略学习
Jiang, Puming, Hu, Tianrun, Du, Haozhe, Li, Yibo, Xue, Zhiwei, Li, Xinhu, Soh, Harold
Abstract
Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet information is lost between composition and the skills it calls. Where a skill works is determined by the structure its policy is trained with, while composition sees the skill only through a separate description, such as a name, an instruction, or a symbolic operator, that omits this structure. Our key idea is to use each policy's structural prior as part of the interface between composition and the skill. A structural prior states what a behavior depends on, for example that a grasp depends only on the gripper's pose relative to the object. Built into training, it shapes where the policy generalizes; stated in language, it tells composition where the policy applies. We instantiate this idea in Agent Priors-guided Policy Learning (APPL). A construction agent segments complete demonstrations into reusable skills, proposes several structural priors for each skill, and trains and verifies one policy per prior. A runtime agent then selects among these prior-specific policies and composes them toward new task goals using their interfaces. Across MetaWorld and long-horizon ManiSkill tasks, APPL improves out-of-distribution skill generalization and enables previously unseen skill compositions; ablating the interface information substantially reduces performance. These results support the use of training-time structural assumptions as a bridge between skill learning and skill composition.
Chinese Translation
从少量演示中学习的机器人通常需要两种形式的泛化。组合泛化重新组合技能以解决新任务,而技能泛化使每个技能背后的学习策略能在新情境中工作。两者相互依赖,但在组合与其调用的技能之间会丢失信息。技能的有效范围由其策略训练时所具有的结构决定,而组合仅通过一个单独的描述(如名称、指令或符号操作符)来查看技能,这种描述省略了该结构。我们的核心思想是将每个策略的结构先验用作组合与技能之间接口的一部分。结构先验说明了行为依赖于什么,例如抓取仅依赖于夹爪相对于物体的位姿。融入训练时,它塑造策略泛化的范围;用语言表述时,它告诉组合该策略在何处适用。我们在智能体先验引导的策略学习(APPL)中实例化了这一思想。一个构建智能体将完整演示分割为可重用技能,为每个技能提出若干结构先验,并为每个先验训练和验证一个策略。然后,一个运行时智能体在这些特定于先验的策略中进行选择,并使用它们的接口将它们组合以朝向新的任务目标。在 MetaWorld 和长时域 ManiSkill 任务中,APPL 改进了分布外技能泛化,并实现了先前未见的技能组合;消融接口信息会显著降低性能。这些结果支持使用训练时的结构假设作为技能学习与技能组合之间的桥梁。
cs.RO / 188 / 2609.35700

LQR-ArUco Fusion: Robust Hierarchical Control for Navigation and Asymmetric Manipulation in Two-Wheeled Robots

LQR-ArUco融合:用于两轮机器人导航与不对称操作的鲁棒分层控制
Chatterjee, Anupam, Kumari, Arpita
Abstract
We propose a hierarchical control framework to address severe dynamic instabilities and navigational drift that arise when a two-wheeled inverted pendulum (TWIP) robot attempts asymmetric object manipulation. While two-wheeled platforms are highly manoeuvrable, their constant balancing adjustments make onboard odometry highly unreliable for precise navigation. Furthermore, the addition of a side-mounted robotic arm introduces unactuated lateral roll moments when a payload is lifted, a challenge heavily compounded on uneven terrain. To solve these coupled problems, our architecture divides the workload. An offboard vision system tracks overhead ArUco markers to provide high-latency global waypoint navigation, bypassing odometry drift. Simultaneously, a low-latency onboard control loop rejects active physical disturbances using inertial and encoder data. In our physical experiments, this dual-loop approach enabled the custom-built robot to navigate accurately, reject transient impacts from speed bumps, adapt to a dynamic seesaw ramp, and carry a payload securely without falling over its narrow wheelbase.
Chinese Translation
我们提出了一种分层控制框架,以解决两轮倒立摆(TWIP)机器人尝试进行不对称物体操作时出现的严重动态不稳定性和导航漂移。虽然两轮平台具有高度机动性,但其持续的平衡调整使得车载里程计对于精确导航非常不可靠。此外,侧装机械臂的加入在提升有效载荷时会引入未驱动的侧向滚转力矩,这一挑战在不平坦地形上被严重放大。为了解决这些耦合问题,我们的架构将工作负载进行了划分。一个离板视觉系统跟踪头顶的ArUco标记,以提供高延迟的全局航点导航,绕过里程计漂移。同时,一个低延迟的车载控制回路利用惯性及编码器数据抑制主动物理干扰。在我们的物理实验中,这种双回路方法使定制机器人能够准确导航,抑制来自减速带的瞬态冲击,适应动态跷跷板斜坡,并安全地携带有效载荷而不会在其狭窄的轮距上翻倒。
cs.RO / 189 / 2609.35709

Humanoid Loco-Manipulation With Discrete VLA Model

基于离散VLA模型的人形机器人移动操作
Shao, Wenxin, Chai, Siqi, Li, Kun, Zhang, Kerou, Jiang, Xinzhou, Xu, Wei, Liu, Qiang
Abstract
Vision-language-action (VLA) models using discrete action tokens have proven effective for controling robotic arms on manipulation tasks. For a humanoid, however, the whole-body action space -- legs, torso, arms, and hands -- is far higher-dimensional and heterogeneous, raising tokenization, training, and real-time inference challenges that the previous VLA models do not address. We present Holo-M, to our knowledge the first discrete VLA model for humanoid loco-manipulation that intrinsically exploits the language model by extending its vocabulary with action tokens. In this model, we devise a unified action tokenizer that decomposes the humanoid action space into four body-part-specific tokenizers -- end-effector, body, hand, and kinematics -- enabling training across drastically different embodiments and data sources, including humanoid teleoperation, ego-centric human video, and simulation. By extending the language model's vocabulary with these action tokens, we avoid the knowledge-insulation problem inherent to the models that use separate continuous action experts. To meet real-time control requirements, we decode each body part's action tokens through grouped discrete diffusion decoding, rather than using autoregression on the action tokens. We have conducted extensive experiments on the SIMPLE humanoid loco-manipulation benchmark, in which Holo-M achieves the highest success rates in both the generalist and specialist evaluations, leading the second best by significant margins. We will release all the code and model weights.
Chinese Translation
使用离散动作令牌的视觉-语言-动作(VLA)模型已被证明在机械臂操作任务控制中有效。然而,对于人形机器人,全身动作空间——腿、躯干、手臂和手——维度更高且异构,带来了先前VLA模型未解决的令牌化、训练和实时推理挑战。我们提出Holo-M,据我们所知,这是首个用于人形机器人移动操作的离散VLA模型,它通过用动作令牌扩展语言模型的词汇表,本质上利用了语言模型。在该模型中,我们设计了一个统一的动作令牌化器,将人形动作空间分解为四个身体部位特定的令牌化器——末端执行器、身体、手和运动学——使得能够在截然不同的具身形态和数据源上进行训练,包括人形遥操作、以自我为中心的人类视频和仿真。通过用这些动作令牌扩展语言模型的词汇表,我们避免了使用单独连续动作专家的模型固有的知识隔离问题。为了满足实时控制需求,我们通过分组离散扩散解码来解码每个身体部位的动作令牌,而不是对动作令牌使用自回归。我们在SIMPLE人形移动操作基准上进行了大量实验,其中Holo-M在通用和专家评估中均取得了最高的成功率,显著领先第二名。我们将发布所有代码和模型权重。
cs.RO / 190 / 2609.35717

RoboCompiler: Graph-Native Compilation of Closed-Chain Robots for Consistent Modeling, Control, and Simulation

RoboCompiler:面向闭链机器人的图原生编译,实现一致的建模、控制与仿真
Shahna, Mehdi Heydari, Kim, Joongheon, Mattila, Jouni
Abstract
Robots with kinematic loops, coupled actuators, and changing contacts require consistent models of configuration, motion, force, and dynamics. Yet these interfaces are often reconstructed separately for control and simulation, making closure and actuation consistency difficult to maintain. This paper presents RoboCompiler, a graph-native framework that compiles a canonical mechanism graph into a shared mechanical interface. From bodies, joints, frames, inertias, and actuator ports, it constructs closure paths and analytic residual Jacobians, then assembles feasible configurations through rank-checked continuation and correction. A tangent lift maps independent velocities to full robot and task motion, while paired actuator-port maps preserve virtual work. A constraint-curvature correction extends the reduction to accelerations and projected rigid-body dynamics, including floating-base and support modes. Cycle-local evaluation, generated Jacobians, and dependency-aware reuse enable localized updates when closure inputs change. We evaluate physical loops and task-induced constraints on a Komatsu excavator, Unitree Go2, Franka Panda, Kangaroo, and a six-UPS Stewart platform. High-precision constrained-dynamics and independent Pinocchio checks confirm mechanical consistency; MuJoCo and Isaac Sim/PhysX executions demonstrate task performance and model reuse under native contact. For Kangaroo, compilation reduces residual-and-Jacobian evaluation time by 96.7% and closed-loop rollout wall time by 66.8%, with dynamics and control held fixed.
Chinese Translation
具有运动学闭环、耦合执行器和变化接触的机器人,需要构型、运动、力和动力学的一致模型。然而,这些接口往往分别为控制与仿真单独重建,导致闭环与驱动一致性难以维持。本文提出 RoboCompiler,一种图原生框架,可将规范机构图编译为共享机械接口。该框架从刚体、关节、坐标系、惯量和执行器端口出发,构建闭合路径与解析残差雅可比,然后通过秩检查的延拓与校正组装可行构型。切空间提升将独立速度映射为完整的机器人运动和任务运动,而成对的执行器端口映射保持虚功。约束曲率校正将降维扩展至加速度和投影刚体动力学,包括浮动基座和支撑模式。循环局部求值、生成的雅可比以及依赖感知复用,使得闭环输入变化时能够进行局部更新。我们在小松挖掘机、Unitree Go2、Franka Panda、Kangaroo 以及六-UPS Stewart 平台上评估物理闭环和任务诱导约束。高精度约束动力学和独立 Pinocchio 校验确认了机械一致性;MuJoCo 和 Isaac Sim/PhysX 执行展示了原生接触下的任务性能与模型复用。对于 Kangaroo,在动力学和控制保持固定不变的情况下,编译将残差与雅可比求值时间降低 96.7%,并将闭环 rollout 墙钟时间降低 66.8%。
cs.RO / 191 / 2609.35761

DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations

DexRoam: 从自我中心全身人体演示中学习移动双臂灵巧操作
Zhou, Rui, Yuan, Yibo, Zhao, Junkai, Zhao, Fangyuan, Zhao, Xiaoguang, Zhang, Shanghang, Han, Sirui
Abstract
Mobile bimanual dexterous manipulation requires continuous coordination of locomotion, whole-body motion, and finger-level dexterity within a single trajectory, creating a severe robot demonstration bottleneck. Egocentric human demonstrations offer a scalable alternative, but prior approaches ease the transfer by simplifying human motion, discarding exactly the fine-grained, coupled structure such tasks depend on. We present DexRoam, a complete system for learning mobile bimanual dexterous manipulation from human demonstrations, in which whole-body motion remains continuous and coupled throughout the human-to-robot transfer process. To enable scalable collection of whole-body human manipulation demonstrations, we develop a tracker-free capture system using only a consumer VR headset and a head-mounted stereo camera, without external cameras or motion trackers. We then perform three explicit alignment stages---embodiment, action-semantic, and temporal---to map captured motion into the robot action space, preserving fine-grained whole-body motion and allowing human and robot demonstrations to be jointly learned by standard VLA policies. Real-world experiments with different VLA backbones show that human demonstrations consistently improve policy learning across training paradigms, raising average success from 29% to 56% on GR00T N1.7 and from 32% to 57% on pi0.5, while matching robot-only training with half the robot demonstrations. Ablations confirm that each alignment stage is necessary. These results highlight the potential of human demonstrations for scalable whole-body mobile manipulation with preserved fine-grained motion structure.
Chinese Translation
移动双臂灵巧操作需要在单个轨迹中连续协调移动、全身运动和手指级灵巧性,这造成了严重的机器人演示瓶颈。自我中心的人类演示提供了一种可扩展的替代方案,但先前的方法通过简化人类运动来简化迁移,恰好丢弃了此类任务所依赖的细粒度、耦合结构。我们提出了 DexRoam,一个从人类演示中学习移动双臂灵巧操作的完整系统,其中全身运动在从人类到机器人的整个迁移过程中保持连续和耦合。为了实现可扩展地收集全身人体操作演示,我们开发了一种无追踪器的捕获系统,仅使用消费级 VR 头显和头戴式立体相机,无需外部摄像头或运动追踪器。然后,我们执行三个显式对齐阶段——本体、动作语义和时间——将捕获的运动映射到机器人动作空间,保留细粒度的全身运动,并允许标准 VLA 策略联合学习人类和机器人演示。在不同 VLA 主干上的真实世界实验表明,人类演示在不同训练范式下持续改进策略学习,在 GR00T N1.7 上将平均成功率从 29% 提高到 56%,在 pi0.5 上从 32% 提高到 57%,同时仅用一半的机器人演示就达到了与仅机器人训练相当的性能。消融实验证实每个对齐阶段都是必要的。这些结果凸显了人类演示在保留细粒度运动结构的情况下实现可扩展全身移动操作的潜力。
人工智能 (Artificial Intelligence)
300
cs.AI / 1 / 2609.31763

SMARtCARE: Privacy-Preserving Agentic AI Systems for Bounded-Autonomy Clinical Decision Support

SMARtCARE:用于有界自主临床决策支持的隐私保护智能体AI系统
Ramaswamy, Srini, Nayak, Deveeshree
Abstract
Long-context clinical AI systems can miss relevant patient history when prior admissions fall outside the active reasoning context. In ICU monitoring, this can cause early vital-sign drift to appear nonspecific even when it resembles a prior deterioration pattern. SMARtCARE addresses this gap through a four-state clinical decision-support architecture: Stable, Meta-cognitive, Assisted, and Regulated (Revoked). Rather than automatically retrieving prior records, SMARtCARE uses a lossy six-channel fingerprint of the patient's prior trajectory. When current drift matches that fingerprint and the prior record is absent from context, the system raises a Meta-cognitive escalation for clinician review; full retrieval occurs only through clinician action in the Assisted state. A patient-identity guard is designed to enforce correct attribution across data loading, logging, and audit layers. Evaluation combines a synthetic Monte Carlo study that validates the state-transition logic and estimator stability, not clinical performance, with real-data runs on both the MIMIC-III and MIMIC-IV Clinical Database Demos. On MIMIC-III, one prior-pattern recurrence was identified among 14 two-admission patients; on MIMIC-IV, the same pipeline produced no fingerprint matches among 9 two-admission patients, which illustrates a key limitation of a fixed canonical pattern library. Across both runs all logged decisions were fully traceable and correctly attributed. The results support SMARtCARE as a traceable, privacy-aware mechanism for surfacing middle-context risk; they are not a clinical efficacy claim.
Chinese Translation
长上下文临床AI系统在既往入院记录落在活动推理上下文之外时,可能会遗漏相关患者病史。在ICU监测中,这可能导致早期生命体征漂移看起来非特异性,即使它类似于先前的恶化模式。SMARtCARE通过四状态临床决策支持架构解决这一差距:稳定、元认知、辅助和受监管(撤销)。SMARtCARE不是自动检索先前记录,而是使用患者先前轨迹的有损六通道指纹。当当前漂移与该指纹匹配且先前记录不在上下文中时,系统会提升元认知升级以供临床医生审查;完整检索仅在辅助状态下通过临床医生操作进行。患者身份守卫旨在跨数据加载、日志记录和审计层强制执行正确的归属。评估结合了合成蒙特卡洛研究(验证状态转换逻辑和估计器稳定性,而非临床性能)与在MIMIC-III和MIMIC-IV临床数据库演示上的真实数据运行。在MIMIC-III上,14名两次入院患者中识别出1例先前模式复发;在MIMIC-IV上,同一流程在9名两次入院患者中未产生指纹匹配,这说明了固定规范模式库的关键局限性。在两次运行中,所有记录的决定都是完全可追溯且正确归属的。结果支持SMARtCARE作为一种可追溯、隐私感知的机制来揭示中上下文风险;它们不是临床疗效声明。
cs.AI / 2 / 2609.31784

Witeness Overlap: Directional Provenance Inside Open-Weight Model Families

见证重叠:开放权重模型家族内的方向性溯源
Li, Siyuan, Zeng, Haoxuan, Luo, Xin, Jia, Fernando, Li, Florence, Geng, Zhengyang, Kolter, Zico, Lee, Tai Sing, Li, Tianqin
Abstract
Open-weight models are often released, fine-tuned, aligned, merged, and re-released, making provenance audits ask not only whether checkpoints are related, but also which checkpoint came first. Many existing model-provenance methods are designed for a base-known audit setting: given a victim or source model, they test whether a suspect model is related to it. Although these audits are framed as source-to-suspect tests, their underlying evidence is often symmetric, relying on representation similarity, weight similarity, behavioral fingerprints, or correlation statistics. Symmetric pairwise comparisons can detect relatedness, but they cannot by themselves orient relationship between checkpoints A and B. We therefore introduce a local geometric comparison: instead of comparing two checkpoints directly, we add a third same-family checkpoint as a witness and compare the geometry around each candidate endpoint. Direction is inferred by asking which candidate behaves more like a branching parent. Motivated by this idea, and by the empirically observed asymmetry between parent-anchored and child-anchored witness-overlap distributions, we propose Witness Overlap, a prompt-free, training-free white-box test for directional provenance. On 176 LLM checkpoints from 16 families, our one-witness test orients 95.3\% of parent-child decisions using Frobenius cosine. We further evaluate root identification, sibling discrimination, generalizations to VLM and diffusion families, and chain-structured ordering. The signal is robust to weight noise and sparse pruning, with a proposed SVD weight reduction variant showing greater robustness than Frobenius cosine.
Chinese Translation
开放权重模型经常被发布、微调、对齐、合并和重新发布,这使得溯源审计不仅要问检查点是否相关,还要问哪个检查点先出现。许多现有的模型溯源方法是为基模型已知的审计设置设计的:给定一个受害模型或源模型,它们测试嫌疑模型是否与其相关。尽管这些审计被表述为源到嫌疑的测试,但其底层证据通常是对称的,依赖于表示相似性、权重相似性、行为指纹或相关统计量。对称的成对比较可以检测相关性,但它们本身无法确定检查点 A 和 B 之间的关系方向。因此,我们引入一种局部几何比较:不是直接比较两个检查点,而是添加第三个同家族检查点作为见证,并比较每个候选端点周围的几何结构。方向是通过询问哪个候选更像分支父节点来推断的。受此想法以及经验观察到的以父节点为锚和以子节点为锚的见证重叠分布之间的不对称性的启发,我们提出了见证重叠(Witness Overlap),一种无需提示、无需训练的白盒测试,用于方向性溯源。在来自 16 个家族的 176 个 LLM 检查点上,我们的单见证测试使用 Frobenius 余弦对 95.3\% 的亲子决策进行了定向。我们进一步评估了根识别、兄弟判别、对 VLM 和扩散模型家族的泛化以及链式结构排序。该信号对权重噪声和稀疏剪枝具有鲁棒性,并且提出的 SVD 权重降维变体显示出比 Frobenius 余弦更强的鲁棒性。
cs.AI / 3 / 2609.31790

CP-Agent: A Harness-Engineered Agent for Crystal Plasticity Simulation Workflows

CP-Agent:一种用于晶体塑性模拟工作流的工具链工程化智能体
Alfred, Samuel Onimpa, Kumar, Abhishek, Sundararaghavan, Veera
Abstract
Crystal plasticity (CP) simulations predict the mechanical behavior of polycrystalline metals, yet their routine use is hindered by the manual effort of configuring heterogeneous tools, orchestrating multi-step data pipelines, and calibrating constitutive parameters against experiments. These bottlenecks impede productivity in systematic parameter studies, motivating interest in automated workflows. This study presents CP-Agent, a harness-engineered LLM-based agent that autonomously executes complete CP modeling workflows from natural-language tasks. Operating under the ReAct paradigm, the agent reasons about tool selection and sequencing while delegating numerical search to established optimizers. The harness comprises a minimal system prompt, typed tool definitions, a dispatcher, and a safety-bounded iteration loop, encoding domain knowledge through tool schemas rather than hard-coded logic. CP-Agent is demonstrated on four case studies: calibrating four slip parameters of additively manufactured stainless steel 316L against tensile data; validating the workflow against published copper benchmarks, reproducing stress-strain and texture evolution; recovering the initial crystallographic texture of copper, where the agent correctly identifies a diffuse initial texture; and reproducing the multi-pass rolling texture evolution of a Mg-Zn-Ca alloy, where the agent chains five deformation passes and recovers the experimentally observed weakened, split basal texture. In all cases, the agent inferred the correct execution sequence from the task statement, robustly across repeated runs, and delivered physically interpretable results. This work establishes harness engineering as a systematic approach to automating CP modeling workflows while maintaining physical interpretability and auditability through visible reasoning traces.
Chinese Translation
晶体塑性(CP)模拟能够预测多晶金属的力学行为,然而其常规使用受到阻碍,原因在于配置异构工具、编排多步数据流水线以及根据实验校准本构参数需要手动操作。这些瓶颈阻碍了系统参数研究中的生产力,从而激发了人们对自动化工作流的兴趣。本研究提出了 CP-Agent,一种工具链工程化的基于 LLM 的智能体,能够从自然语言任务自主执行完整的 CP 建模工作流。在 ReAct 范式下运行,该智能体对工具选择和排序进行推理,同时将数值搜索委托给已有的优化器。该工具链包含一个最小系统提示、类型化工具定义、一个调度器以及一个安全受限的迭代循环,通过工具模式而非硬编码逻辑来编码领域知识。CP-Agent 在四个案例研究中得到展示:根据拉伸数据校准增材制造不锈钢 316L 的四个滑移参数;针对已发表的铜基准验证工作流,复现应力-应变和织构演化;恢复铜的初始晶体学织构,其中智能体正确识别出弥散的初始织构;以及复现 Mg-Zn-Ca 合金的多道次轧制织构演化,其中智能体串联五个变形道次并恢复实验观察到的弱化、分裂基面织构。在所有情况下,智能体从任务陈述中推断出正确的执行顺序,在重复运行中具有鲁棒性,并提供了物理上可解释的结果。这项工作确立了工具链工程作为一种系统方法,用于自动化 CP 建模工作流,同时通过可见的推理轨迹保持物理可解释性和可审计性。
cs.AI / 4 / 2609.31792

ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts

ConflictVLA-Bench:对视觉-语言-动作模型在前提冲突下行为响应的基准测试
Hou, Liyu, Wu, Yuan, Chang, Yi
Abstract
While Vision-Language-Action (VLA) models perform strongly on manipulation tasks, their responses to invalid task premises remain underexplored. Existing evaluations of premise conflicts often focus on terminal task outcomes, yet task failure alone cannot distinguish behavioral disengagement from continued pursuit followed by an execution error. We call the latter pattern Failed Persistence. To study this phenomenon, we introduce ConflictVLA-Bench, which pairs conflict rollouts with premise-consistent reference rollouts and evaluates both outcomes and execution processes. Built on LIBERO, the benchmark contains 2,826 prompt-conditioned conflict tasks spanning four conflict families, four structural configurations, and two prompt conditions. Across all eight VLAs, invalid premises reduce original goal completion by at least 17.3 percentage points, with the reduction reaching 56.2 percentage points for OpenVLA. Crucially, even when models succeed on premise-consistent tasks and fail on their matched conflict tasks, they often continue to approach the original targets, retain early trajectory structure, and show limited action magnitude suppression. Failed Persistence therefore recurs across the evaluated models. Explicit premise checking does not consistently produce selective and coordinated behavioral changes. These findings show that terminal failure alone establishes neither behavioral disengagement nor refusal and that outcomes alone are insufficient for VLA evaluation. Experimental data and additional details are available on the project page: https://github.com/EmbodiedAISurvey/ConflictVLA-Bench
Chinese Translation
尽管视觉-语言-动作(VLA)模型在操作任务上表现强劲,但它们对无效任务前提的响应仍未得到充分探索。现有对前提冲突的评估通常关注最终任务结果,然而仅凭任务失败无法区分行为脱离与继续追求后出现执行错误。我们将后一种模式称为 Failed Persistence(失败坚持)。为研究这一现象,我们提出了 ConflictVLA-Bench,它将冲突 rollout 与前提一致的参考 rollout 配对,并同时评估结果与执行过程。该基准基于 LIBERO 构建,包含 2,826 个以提示为条件的冲突任务,涵盖四个冲突类别、四种结构配置和两种提示条件。在所有八个 VLA 中,无效前提使原始目标完成率至少降低 17.3 个百分点,其中 OpenVLA 的降低幅度达到 56.2 个百分点。关键的是,即使模型在前提一致的任务上成功、在其匹配的冲突任务上失败,它们仍常常继续接近原始目标,保留早期轨迹结构,并表现出有限的动作幅度抑制。因此,Failed Persistence 在被评估的模型中反复出现。显式前提检查并不能一致地产生选择性和协调的行为变化。这些发现表明,仅凭最终失败既不能确定行为脱离,也不能确定拒绝;仅凭结果不足以进行 VLA 评估。实验数据和更多细节可在项目页面获取:https://github.com/EmbodiedAISurvey/ConflictVLA-Bench
cs.AI / 5 / 2609.31793

Working with AI: A Design Framework for Human-AI Collaboration

与AI共事:面向人机协作的设计框架
Lu, Yuqian, Lee, Regina, Zhou, Rui, Jiang, Lixin, McDaid, Andrew, Lawrence, Amy
Abstract
Artificial Intelligence (AI), particularly GenAI, is becoming an increasingly important part of modern work. In industrial settings, AI can support decision-making, automate routine activities, assist humans, and improve productivity. However, successful AI adoption depends on more than what the technology can do. It also depends on how people experience and work with it. This raises an important question: how should human-AI collaboration be designed so that it works well for both people and organisations? This white paper addresses that question by presenting a practical framework for designing human-AI collaboration. The framework considers the human, the AI system, the task, the organisation, and the wider societal environment. It explains what effective collaboration looks like, what conditions influence it, what requirements should be met, and what design decisions organisations should consider. The report also includes a human-AI collaborative assembly system with cobot use case to demonstrate how the framework can be applied in practice. The use case shows how design requirements can be translated into specific collaboration features and evaluated through a case study. The aim of this white paper is to provide a clear and practical guide for designing human-AI collaboration that is effective, human-centred, and responsible.
Chinese Translation
人工智能(AI),尤其是生成式人工智能(GenAI),正日益成为现代工作的重要组成部分。在工业场景中,AI可以支持决策、自动化常规活动、辅助人类并提高生产率。然而,成功采用AI不仅取决于技术能够做什么,还取决于人们如何体验和与之协作。这引出了一个重要问题:应该如何设计人机协作,使其对人和组织都行之有效?本白皮书通过提出一个用于设计人机协作的实用框架来回应这一问题。该框架考虑人、AI系统、任务、组织以及更广泛的社会环境。它阐述了有效协作的表现形式、影响协作的条件、应满足的要求,以及组织应考虑的设计决策。报告还包括一个使用协作机器人(cobot)的人机协作装配系统用例,以展示该框架如何应用于实践。该用例展示了如何将设计要求转化为具体的协作特征,并通过案例研究进行评估。本白皮书旨在为设计有效、以人为本且负责任的人机协作提供清晰实用的指南。
cs.AI / 6 / 2609.31814

DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution

DriveHierarchy:从开环理解到闭环执行诊断VLM驾驶能力的基准
Xu, Chengkai, Liu, Jiaqi, Guo, Yicheng, Hang, Peng, Sun, Jian
Abstract
Evaluating VLM-based autonomous driving remains difficult because driving competence is composite, where a capable system must ground traffic participants and hazards, integrate context across views and time, reason about future evolution, and act appropriately under closed-loop interaction. Existing benchmarks usually assess either open-loop understanding or closed-loop driving but provide limited structure for explaining how these abilities are organized, how they relate, and how they may inform model diagnosis and improvement. We present \textsc{DriveHierarchy}, a hierarchical benchmark that organizes VLM-based autonomous driving into four ranks, spanning perceptual grounding, contextual memory, mental reasoning, and closed-loop execution. To instantiate this hierarchy, we integrate multiple open-source autonomous-driving datasets into a unified open-loop benchmark with 76,798 question-answer pairs over 84,279 frames and develop a closed-loop simulation platform with interactive scenario construction on a real-world road network, from which 100 driving scenarios are curated for embodied evaluation. Experiments on 15 VLMs show that \textsc{DriveHierarchy} captures structured but non-redundant capability variation, relates open-loop understanding to closed-loop driving, and provides a practical basis for diagnosis and benchmark-guided optimization. \textsc{DriveHierarchy} therefore serves as a unified framework for evaluating and improving VLM-based autonomous driving systems. An anonymized project has been released on https://github.com/PerfectXu88/DriveHierarchy
Chinese Translation
评估基于VLM的自动驾驶仍然困难,因为驾驶能力是复合的,一个有能力系统必须对交通参与者和危险进行感知定位,跨视图和时间整合上下文,推理未来演变,并在闭环交互中恰当行动。现有基准通常评估开环理解或闭环驾驶,但提供的结构有限,无法解释这些能力如何组织、如何关联以及如何为模型诊断和改进提供信息。我们提出 DriveHierarchy,一个分层基准,将基于VLM的自动驾驶组织为四个等级,涵盖感知基础、上下文记忆、心智推理和闭环执行。为了实例化该分层,我们将多个开源自动驾驶数据集集成到一个统一的开环基准中,包含 84,279 帧上的 76,798 个问答对,并开发了一个闭环仿真平台,在真实道路网络上构建交互式场景,从中精选出 100 个驾驶场景用于具身评估。在 15 个 VLM 上的实验表明,DriveHierarchy 捕捉了结构化但非冗余的能力变化,将开环理解与闭环驾驶联系起来,并为诊断和基准引导的优化提供了实践基础。因此,DriveHierarchy 可作为评估和改进基于VLM的自动驾驶系统的统一框架。一个匿名项目已在 https://github.com/PerfectXu88/DriveHierarchy 发布。
cs.AI / 7 / 2609.31857

LLM Judge Validation Under Sparse Overlap: From Inference to Design

稀疏重叠下的LLM评判者验证:从推断到设计
Li, Junxuan, Mukherjee, Arko, Pal, Soumyabrata
Abstract
Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this \emph{overlap sparsity} is the first-order determinant of wrong deployment decisions: at 5\% pairwise overlap, wrong-decision rates reach 25\% and the probability of selecting the wrong best judge among ten candidates is 65\%. The two actionable levers are overlap \emph{quantity} and \emph{allocation}. For quantity, we derive a minimum-overlap formula showing $\rho \geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.
Chinese Translation
验证一个作为评判者的LLM需要估计其与人类的一致性,然而标注预算很少允许对每个条目进行多次标注。我们证明了这种重叠稀疏性是错误部署决策的一阶决定因素:在5%的成对重叠下,错误决策率达到25%,并且在十个候选者中选错最佳评判者的概率为65%。两个可操作的杠杆是重叠的数量和分配。对于数量,我们推导出一个最小重叠公式,表明 $\rho \geq 0.25$ 对于非边缘评判者就足够了,而边缘情况仍然本质上是困难的。对于分配,当分层具有信息性时,一种零成本的分层方案相对于随机抽样将错误拒绝率减半。我们在10个LLM评判者上、跨四个评估矩阵进行了验证,这些矩阵涵盖视觉评估、因果推理和摘要生成。
cs.AI / 8 / 2609.31868

Metro-WM: Long-Horizon Latent Planning with Realisable Sub-Goals

Metro-WM:具有可实现子目标的长时域潜在规划
Lee, Royson, Rezk, Fady, Parcollet, Titouan, Hospedales, Timothy, Cornelio, Cristina
Abstract
Model-predictive control with Joint-Embedding Predictive Architectures (JEPAs) provides a strong zero-shot goal-reaching planner, but it is only effective over short planning horizons. Hierarchical extensions attempt to bridge this gap by learning a macro planner to predict intermediate latent sub-goals to guide the micro planner. In this work, we demonstrate that unconstrained latent sub-goal prediction is fundamentally flawed. A rigorous evaluation reveals that a leading state-of-the-art macro planner routinely emits physically unrealisable sub-goals. To resolve this, we introduce Metro-WM, a hierarchical framework that issues sub-goals by retrieving genuine states from prior experience rather than generating ungrounded latent vectors. Specifically, Metro-WM constructs a graph whose vertices are observed frames from offline expert demonstrations or random-action trajectories, allowing frames from different episodes to be connected and stitched into routes to the goal. Planning over the full graph also makes the system highly robust to execution errors: if the micro planner drifts off course, Metro-WM instantly finds a new optimal path from the current state. Our experiments show that Metro-WM achieves superior long-horizon success rates of up to 37.33 percentage points over the next best hierarchical approach while being up to 10.9 times faster, requiring both 13-56 times less offline compute and fewer tuned hyperparameters. Additional analysis reveals that Metro-WM finds shorter paths than the offline demonstrations, outperforms an oracle relying on the query's own demonstration, and maintains robust performance under extremely sparse dataset conditions.
Chinese Translation
带有联合嵌入预测架构(JEPAs)的模型预测控制提供了强大的零样本目标达成规划器,但它仅在短规划时域内有效。分层扩展试图通过学习一个宏规划器来预测中间潜在子目标以指导微规划器,从而弥合这一差距。在这项工作中,我们证明无约束的潜在子目标预测存在根本缺陷。严格的评估揭示,领先的最先进宏规划器经常发出物理上不可实现的子目标。为了解决这个问题,我们引入了 Metro-WM,一个分层框架,它通过从先验经验中检索真实状态来发出子目标,而不是生成缺乏依据的潜在向量。具体而言,Metro-WM 构建一个图,其顶点是来自离线专家演示或随机动作轨迹的观测帧,允许将不同回合的帧连接并拼接成通往目标的路线。在全图上进行规划也使系统对执行误差高度鲁棒:如果微规划器偏离路线,Metro-WM 会立即从当前状态找到新的最优路径。我们的实验表明,Metro-WM 实现了优越的长时域成功率,比次优的分层方法高出多达 37.33 个百分点,同时速度最高快 10.9 倍,所需的离线计算量减少 13-56 倍,且需要调优的超参数更少。额外的分析揭示,Metro-WM 找到比离线演示更短的路径,优于依赖查询自身演示的 oracle,并在极其稀疏的数据集条件下保持稳健性能。
cs.AI / 9 / 2609.31871

IndustryLLM: Failure-Driven LLM Training for Industrial Procurement

IndustryLLM:面向工业采购的失败驱动大语言模型训练
Ding, Liang, Xu, Zhiang, Sheng, Yuyang, Chen, Bin, Bai, Songlin, Zhu, Run, Wu, Dingjun, Xu, Hui, Wang, Yandi, Shi, Fulin, Gan, Leilei, Yu, Linlin, Zhong, Qihuang, Peng, Keqin, Li, Yalong, Huo, Chengfu
Abstract
Industrial procurement requires language models to bridge informal buyer jargon, sparse marketplace attributes, and authoritative engineering standards under strict safety tolerances. We present IndustryLLM, an open-weight industrial language model trained from Qwen3.5-35B-A3B-Base (35B total parameters with ~3B activated per token, with the vision encoder frozen). Rather than relying on generic text scaling, we introduce a failure-driven adaptation recipe spanning continued pre-training (CPT) and supervised fine-tuning (SFT). CPT leverages a curated ~100B-token corpus integrating 5B tokens of national standards (e.g., GB/T) and technical archives, 10B tokens of de-identified real-world industrial transaction and inquiry records, and 60B tokens of general replay. To overcome register mismatch and factual brittleness, we systematically reconstruct an estimated 20B-token domain subset via multi-register rewriting across 10 genres and 8 writing styles, confidence-routed minimal factual editing, and error-targeted QA synthesis (resolving colloquial typos like '42-luo-mu' -> 42CrMo, expanding ambiguous codes like '16674' -> GB/T 16674, and clarifying conflicting dimensional specs). For downstream deployment, we formalize an evidence-gated constraint-evaluation interface enforcing three-valued logic where unverified product evidence remains unknown rather than satisfied. Offline evaluations demonstrate consistent gains on procurement-query structuring (+2.97 percentage points in exact match, 95% CI [2.11, 3.86] in No-Think mode), while randomized online A/B experiments in production yield substantial improvements (+4.25% GMV, +8.3% satisfied inquiries) alongside a latency reduction from 6-7 s to 1.5 s. Model weights and configs are released at https://huggingface.co/alibaba-multimodal-industrial-ai/IndustryLLM.
Chinese Translation
工业采购要求语言模型在严格的安全容差下,弥合非正式买家术语、稀疏的市场属性和权威工程标准之间的差距。我们提出 IndustryLLM,一个开放权重的工业语言模型,基于 Qwen3.5-35B-A3B-Base 训练(总参数 35B,每 token 激活约 3B,视觉编码器冻结)。并非依赖通用的文本规模扩展,我们引入一种贯穿持续预训练(CPT)和监督微调(SFT)的失败驱动适配方案。CPT 利用一个精心构建的约 100B token 的语料库,整合了 5B token 的国家标准(如 GB/T)和技术档案、10B token 的去标识化真实工业交易与询价记录,以及 60B token 的通用回放数据。为克服语域不匹配和事实脆弱性,我们通过跨 10 种体裁和 8 种写作风格的多语域重写、基于置信度路由的最小事实编辑、以及面向错误的问答合成,系统性地重建了一个估计 20B token 的领域子集(解决诸如“42-luo-mu”-> 42CrMo 的口语化拼写错误,展开“16674”-> GB/T 16674 等模糊代码,并澄清冲突的尺寸规格)。对于下游部署,我们形式化了一个证据门控的约束评估接口,强制执行三值逻辑,其中未验证的产品证据保持为未知,而不是满足。离线评估表明,在采购查询结构化方面取得了一致的提升(在 No-Think 模式下,精确匹配率提升 2.97 个百分点,95% CI [2.11, 3.86]),而在生产环境中的随机在线 A/B 实验带来了显著改进(GMV +4.25%,满意询价 +8.3%),同时将延迟从 6-7 秒降低到 1.5 秒。模型权重和配置已在 https://huggingface.co/alibaba-multimodal-industrial-ai/IndustryLLM 发布。
cs.AI / 10 / 2609.31874

COUNTERMEM: World-Model Verified Counter-Factual Memory for Language Agents

COUNTERMEM:面向语言智能体的世界模型验证反事实记忆
Pu, Hongji, Tang, Ruixiang, Zhang, Yongfeng
Abstract
Existing agent memory frameworks mainly create memory through an agent's interaction with the factual world, e.g., remembering feedback from actions taken to improve performance on future tasks. However, these frameworks seldom ask the "what if" question during memory construction: what if a different action had been taken, would the feedback have changed, and how could this feedback become useful memory? Obtaining such feedback directly in an active environment can be expensive and can alter the state needed for comparison. In this work, we introduce COUNTERMEM, a reinforcement-learning framework for constructing and using verified counterfactual memory across tasks. After a failed action, COUNTERMEM evaluates local alternatives from a copy or reset of the original state using executable world models, such as tests, proof checkers, and solvers. It stores improvements with the original and corrected actions, checked outcomes, and conditions for reuse. A learned memory-use policy selects a retrieved record or skips memory to balance task success and interaction cost, while the base LLM remains fixed. Both memory and policy are frozen during held-out evaluation. We evaluate COUNTERMEM on 12 benchmark settings across six domains. With gpt-oss-120b, COUNTERMEM improves both ReAct and Reflexion on all 12 benchmarks across six domains, averaging a gain of 12.6 percentage points over their unaugmented versions. In the four-domain comparison across two backbones, task-run tokens decrease by 7.7-42.0%, excluding offline selector-training costs. Further analyses show that removing verification or persistent storage weakens the gains, while applying verified corrections to unsuitable decisions can reverse them. Code will be released upon acceptance.
Chinese Translation
现有的智能体记忆框架主要通过智能体与事实世界的交互来创建记忆,例如,记住所采取行动的反馈,以提升未来任务的表现。然而,这些框架在构建记忆时很少提出“如果……会怎样”的问题:如果采取了不同的行动,反馈是否会改变,以及这种反馈如何能成为有用的记忆?直接在活跃环境中获取此类反馈可能代价高昂,并且可能改变比较所需的状态。在这项工作中,我们提出了COUNTERMEM,一个用于跨任务构建和使用经过验证的反事实记忆的强化学习框架。在一次失败的行动之后,COUNTERMEM利用可执行的世界模型(如测试、证明检查器和求解器),从原始状态的副本或重置中评估局部替代方案。它存储改进以及原始行动和修正后的行动、经过检查的结果以及重用条件。一个学习到的记忆使用策略选择检索到的记录或跳过记忆,以平衡任务成功和交互成本,而基础LLM保持固定。在留出评估期间,记忆和策略都被冻结。我们在六个领域的12个基准设置上评估了COUNTERMEM。使用gpt-oss-120b时,COUNTERMEM在六个领域的所有12个基准上改进了ReAct和Reflexion,相比其未增强版本平均提升12.6个百分点。在两个骨干模型的四领域比较中,任务运行token减少了7.7-42.0%,不包括离线选择器训练成本。进一步的分析表明,移除验证或持久存储会削弱增益,而将经过验证的修正应用于不合适的决策可能会逆转这些增益。代码将在论文被接收后发布。
cs.AI / 11 / 2609.31897

Context-dependent agent evaluation with orthogonal equilibrium learning

基于正交均衡学习的上下文相关智能体评估
Ma, Haorui, Zang, Zehua, Li, Jiangmeng, Li, Yi, Xu, Fanjing, Feuerriegel, Stefan
Abstract
Many applications require to evaluate agents under contextual information (e.g., a prompt, task, or user group). We study how to perform such context-dependent agent evaluation from offline feedback. Existing score-based models for this purpose (e.g., Bradley-Terry) impose a transitive preference ordering, which fails to reflect collective preferences when human judgements are heterogeneous. Inspired by social choice theory, we frame evaluation as a contextual game between two players, each selecting a distribution over agents as the strategy to receive greater collective preference than the other. Then, the support of the Nash equilibrium defines a context-specific set of winners. However, learning context-specific equilibria from offline logs is difficult because each context reveals human feedback on only a subset of agents, and, hence, a naive plug-in estimator can therefore be biased. To address these challenges, we propose NashEval, a general framework for robust contextual equilibrium learning. NashEval first constructs debiased estimates of the contextual payoff matrix that characterizes the game. NashEval then learns the context-to-equilibrium mapping with a tailored orthogonal loss, which avoids the need to solve a separate game for each context. We show theoretically that errors in estimating the nuisance functions underlying the payoff matrix affect the risk of the learned equilibrium (i.e., exploitability) only through higher-order terms. Across various experiments, NashEval improves robustness of equilibrium learning and consistently identifies the set of top-performing agents across contexts.
Chinese Translation
许多应用需要在上下文信息(例如,提示、任务或用户群体)下评估智能体。我们研究如何从离线反馈中进行这种上下文相关的智能体评估。现有的用于此目的的基于分数的模型(例如,Bradley-Terry)强加了一个传递性偏好排序,当人类判断异质时,这无法反映集体偏好。受社会选择理论启发,我们将评估构建为两个玩家之间的上下文博弈,每个玩家选择一个在智能体上的分布作为策略,以获得比对方更大的集体偏好。然后,纳什均衡的支撑集定义了一个特定于上下文的获胜者集合。然而,从离线日志中学习特定于上下文的均衡是困难的,因为每个上下文只揭示人类对一部分智能体的反馈,因此,朴素的插件估计器可能是有偏的。为了应对这些挑战,我们提出了NashEval,一个用于鲁棒上下文均衡学习的通用框架。NashEval首先构建刻画该博弈的上下文收益矩阵的去偏估计。然后,NashEval使用定制的正交损失学习上下文到均衡的映射,这避免了对每个上下文求解单独的博弈。我们从理论上证明,估计收益矩阵背后的干扰函数时的误差仅通过高阶项影响学习到的均衡的风险(即可利用性)。在各种实验中,NashEval提高了均衡学习的鲁棒性,并一致地识别出跨上下文表现最好的智能体集合。
cs.AI / 12 / 2609.31903

Choir: An Open Protocol for Distributed Multi-Agent Autoformalization

Choir:一种面向分布式多智能体自动形式化的开放协议
Qi, Yidi, Weber, Melanie
Abstract
AI agents can now formalize entire textbooks and major theorems in proof assistants such as Lean, but current efforts are typically centralized: a single team runs all agents and bears the full computational cost. We introduce Choir, an open protocol for distributed formalization. Choir decomposes a project into tasks that can be completed by independent contributors, each running their own agent with their own LLM subscription, while coordinating entirely through the project's GitHub repository. To support open participation, every contribution is checked by a deterministic gate before merge. Choir supports Lean 4, Isabelle, and Rocq, and is open source and modular, allowing projects to replace individual components or extend the protocol.
Chinese Translation
AI 智能体如今可以在 Lean 等证明助手中形式化整本教科书和重要定理,但当前的工作通常是集中式的:由一个团队运行所有智能体并承担全部计算成本。我们提出 Choir,一种用于分布式形式化的开放协议。Choir 将项目分解为可由独立贡献者完成的任务,每个贡献者使用自己的 LLM 订阅运行自己的智能体,同时完全通过项目的 GitHub 仓库进行协调。为支持开放参与,每项贡献在合并前都要经过确定性门控检查。Choir 支持 Lean 4、Isabelle 和 Rocq,并且开源且模块化,允许项目替换单个组件或扩展协议。
cs.AI / 13 / 2609.31906

EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

EmailBench:评估 LLM 智能体在企业邮件与生产力任务上的基准
Singh, Mukul, Uniyal, Mansi, Devlin, Devin, Xie, Wen, Thadawasin, Big, Dutt, Ritam, Lai, Vivian, Kang, Hyeonsu B.
Abstract
Enterprise email agents must combine information retrieval, structured state changes, temporal reasoning, and multi-step coordination. Recent agent benchmarks include productivity tasks, but few center on typed email workflows in a self-contained environment. We introduce EmailBench, a benchmark of 206 email and productivity scenarios across 16 task categories. The benchmark couples a typed email API specification with provider-neutral naming, a deterministic synthetic Enron-inspired corpus, and a scenario suite whose topic selection was informed by aggregate task-intent telemetry from an interactive prototype. Its hybrid evaluation protocol combines 258 executable static assertions with 211 LLM rubrics. We evaluate eight LM configurations on a fixed single-user corpus. The best-performing configuration passes only 33.5% of scenarios despite 99.7% of its tool calls completing without an observed API failure, with pass rates varying substantially across task categories. This gap shows that valid tool execution is not equivalent to task completion. EmailBench provides a self-contained environment for end-to-end email-agent evaluation, with broader tool coverage, multi-persona testing, and repeated-run evaluation as future work areas.
Chinese Translation
企业邮件智能体必须结合信息检索、结构化状态变更、时间推理和多步协调。近期的智能体基准包含生产力任务,但很少有以自包含环境中的类型化邮件工作流为中心的。我们引入了 EmailBench,一个包含 16 个任务类别、206 个邮件与生产力场景的基准。该基准将类型化邮件 API 规范与提供商中立命名、一个确定性的受 Enron 启发的合成语料库以及一个场景套件相结合,该场景套件的主题选择基于来自一个交互式原型的聚合任务意图遥测数据。其混合评估协议结合了 258 个可执行的静态断言与 211 个 LLM 评分标准。我们在一个固定的单用户语料库上评估了八种 LM 配置。表现最佳的配置仅通过了 33.5% 的场景,尽管其 99.7% 的工具调用完成且未观察到 API 失败,且不同任务类别的通过率差异很大。这一差距表明,有效的工具执行并不等同于任务完成。EmailBench 为端到端邮件智能体评估提供了一个自包含环境,更广泛的工具覆盖、多角色测试和重复运行评估是未来的工作方向。
cs.AI / 14 / 2609.31908

Improving Medical Calculation of LLMs with Embedded Coding

利用嵌入式编码提升大语言模型的医学计算能力
Ming, Tianshi, Zhang, Yingying, Wu, Xian
Abstract
Large Language Models (LLMs) perform well on medical examinations and question-answering benchmarks, but remain unreliable on medical calculation tasks that require exact numerical outputs. These calculations support high-stakes decisions such as medication dosing, organ-function assessment, and prognostic scoring, for which even small errors can have serious clinical consequences. We introduce MedCode, a framework that improves medical calculation by training LLMs to generate embedded executable code. Given a clinical context, the model identifies the relevant calculator, extracts its input variables, and produces a script that delegates arithmetic operations to a deterministic interpreter. Executing the script returns the calculated value together with an explanation and the appropriate unit. We construct supervised fine-tuning (SFT) and preference datasets from the MedCalc benchmark and additionally curate a dataset for calculation tasks in Intensive Care Unit (ICU) scenarios. We further propose weighted Direct Preference Optimization (wDPO), which adaptively emphasizes preference pairs that are difficult for the model to distinguish. Experiments with LLaMA3-8B, Qwen2.5-7B, and Mistral-7B show absolute accuracy gains of 20--30 percentage points, demonstrating the effectiveness of embedded code generation for medical calculation.
Chinese Translation
大语言模型(LLMs)在医学考试和问答基准上表现良好,但在需要精确数值输出的医学计算任务上仍不可靠。这类计算支撑着高风险决策,如药物剂量、器官功能评估和预后评分,其中即使很小的错误也可能带来严重的临床后果。我们提出 MedCode,一个通过训练 LLMs 生成嵌入式可执行代码来改进医学计算的框架。给定临床背景,模型识别相关计算器,提取其输入变量,并生成一个将算术运算委托给确定性解释器的脚本。执行该脚本可返回计算值,并附带解释和合适的单位。我们基于 MedCalc 基准构建了监督微调(SFT)和偏好数据集,并额外整理了一个用于重症监护病房(ICU)场景计算任务的数据集。我们进一步提出加权直接偏好优化(wDPO),它自适应地强调模型难以区分的偏好对。使用 LLaMA3-8B、Qwen2.5-7B 和 Mistral-7B 的实验显示,绝对准确率提升 20–30 个百分点,证明了嵌入式代码生成在医学计算中的有效性。
cs.AI / 15 / 2609.31939

BioDyad: Synchronize Biomedical Discovery and Machine Learning Engineering

BioDyad:同步生物医学发现与机器学习工程
Du, Xingbo, Ghiffari, Fadli Aulawi Al, Song, Leonard, Li, Loka, Zhang, Duzhen, Wang, Zixiao, Chen, Xiuying, Song, Le
Abstract
Agentic biomedical machine learning (ML) draws on complementary advances in biomedical evidence acquisition and executable program search. Existing systems connect aspects of these capabilities, but coordinating them throughout program search remains challenging. New evidence must guide candidate construction, execution outcomes must inform subsequent discovery and reuse, and validation demands must fit the search budget. We introduce BioDyad, which couples biomedical discovery and ML engineering through two hierarchies within Monte Carlo graph search. Its scientific hierarchy combines prior biomedical guidance with iterative discovery, then links biomedical plans to execution outcomes in memory for reuse across candidates. Its engineering hierarchy moves candidate programs from smoke execution, through train/validation evaluation, to full-data retraining. We evaluate BioDyad on the 76-task BioXArena benchmark under a two-hour per-task budget with three matched LLM backends. It achieves the highest penalized all-task score and task success rate among four agent methods and a one-shot baseline under each backend. These results support coordinating biomedical discovery and ML engineering to integrate external knowledge into executable programs across heterogeneous biomedical tasks.
Chinese Translation
智能体生物医学机器学习(ML)借助生物医学证据获取与可执行程序搜索方面的互补进展。现有系统连接了这些能力的某些方面,但在整个程序搜索过程中协调它们仍然具有挑战性。新证据必须指导候选构建,执行结果必须为后续发现和重用提供信息,验证要求必须适应搜索预算。我们提出 BioDyad,它通过蒙特卡洛图搜索中的两个层级将生物医学发现与机器学习工程耦合起来。其科学层级将先验生物医学指导与迭代发现相结合,然后在内存中将生物医学计划与执行结果关联,以便跨候选重用。其工程层级将候选程序从冒烟执行,经过训练/验证评估,推进到全数据重训练。我们在每个任务两小时预算下,使用三个匹配的LLM后端,在76个任务的BioXArena基准上评估BioDyad。在每个后端下,它在四种智能体方法和一个单次基线中取得了最高的惩罚后全任务得分和任务成功率。这些结果支持协调生物医学发现和机器学习工程,以将外部知识整合到跨异构生物医学任务的可执行程序中。
cs.AI / 16 / 2609.31963

Symbolic Guidance for LLM Agents in Distributed Multiagent Coordination

分布式多智能体协调中LLM智能体的符号引导
Rachmut, Ben, Zhang, Ning, Vorobeychik, Yevgeniy, Yeoh, William
Abstract
Large language models (LLMs) are increasingly deployed as autonomous agents in multi-agent systems, yet their ability to reliably execute distributed coordination protocols remains poorly understood. While AgentsNet, a benchmark framework for distributed coordination among LLM agents, enables such coordination, granting full reasoning autonomy often leads to inconsistent or degraded performance in complex domains. We hypothesize that coordination can be improved by regulating agent autonomy through symbolic guidance derived from established algorithms. To investigate this, we introduce the \emph{Symbolic Guidance Taxonomy (SGT)}, which characterizes a spectrum of autonomy ranging from open-ended natural language reasoning to fully prescribed algorithmic execution, with intermediate levels providing partial pseudocode guidance. Our results show that intermediate autonomy levels consistently outperform both unguided agents and fully prescriptive specifications. These findings identify autonomy regulation as a key design principle for LLM-based distributed coordination.
Chinese Translation
大语言模型(LLMs)正日益被部署为多智能体系统中的自主智能体,然而其可靠执行分布式协调协议的能力仍未被充分理解。尽管AgentsNet——一个用于LLM智能体之间分布式协调的基准框架——能够支持此类协调,但赋予完全的推理自主性往往会在复杂领域中导致性能不稳定或下降。我们假设,通过源自已有算法的符号引导来调节智能体自主性,可以改善协调。为研究这一点,我们引入符号引导分类法(Symbolic Guidance Taxonomy, SGT),其刻画了一个从开放式自然语言推理到完全规定式算法执行的自主性谱系,中间层级提供部分伪代码引导。我们的结果表明,中间自主性水平始终优于无引导智能体和完全规定式规范。这些发现将自主性调节确定为基于LLM的分布式协调的关键设计原则。
cs.AI / 17 / 2609.31980

Goal-Persistent Coding Agents as Scientific Performance Engineers: A Fixed-Radius Nearest-Neighbor Case Study

目标持久的编码智能体作为科学性能工程师:一个固定半径最近邻案例研究
Ju, Xiangyang
Abstract
Coding agents can pursue persistent objectives across many tool-use turns, but evidence that general-purpose agents can conduct rigorous scientific performance engineering remains limited. We present a repository-scale case study in which off-the-shelf Codex and Claude Code agents optimize fixed-radius nearest-neighbor (FRNN) search for particle tracking. Starting from a PyTorch-dependent CUDA implementation, the agents follow an executable goal that specifies exact-correctness tests, profiling requirements, and acceptance criteria without prescribing code transformations. In the primary sequential trajectory, they autonomously remove the PyTorch dependency and conduct hypothesis-driven optimization experiments. The resulting standalone C++/CUDA library exactly reproduces the targeted reference result. Its synchronous NumPy interface achieved 1.6-fold speedup over the original GPU-resident PyTorch interface, despite including host transfers. Similar speedups were observed across different GPU architectures and software stacks. An independent optimization rerun followed a different sequence of hypotheses and reached even better performance on the target workload. These results show that goal-persistent coding agents can act as experimental performance engineers, and that executable scientific contracts are needed both to guide and to validate their optimization.
Chinese Translation
编码智能体能够在多轮工具使用中追求持久目标,但通用智能体能够进行严格的科学性能工程的证据仍然有限。我们展示了一个仓库规模的案例研究,其中现成的 Codex 和 Claude Code 智能体优化用于粒子追踪的固定半径最近邻(FRNN)搜索。从一个依赖 PyTorch 的 CUDA 实现出发,智能体遵循一个可执行目标,该目标指定了精确正确性测试、性能分析要求和验收标准,而不规定代码变换。在主要的顺序轨迹中,它们自主移除了 PyTorch 依赖,并进行了假设驱动的优化实验。得到的独立 C++/CUDA 库精确复现了目标参考结果。其同步 NumPy 接口相比原始 GPU 驻留的 PyTorch 接口实现了 1.6 倍加速,尽管包含了主机传输。在不同的 GPU 架构和软件栈上观察到类似的加速。一次独立的优化重跑遵循了不同的假设序列,并在目标工作负载上达到了更好的性能。这些结果表明,目标持久的编码智能体可以充当实验性能工程师,并且需要可执行的科学契约来指导和验证其优化。
cs.AI / 18 / 2609.31990

CSI-Agent: LLM-Assisted Few-Shot Adaptation for Cross-Domain Wi-Fi CSI Sensing

CSI-Agent:LLM辅助的跨域Wi-Fi CSI感知的小样本自适应
Zhao, Tianya, Liu, Chuan, Wang, Xuyu
Abstract
Wi-Fi channel state information (CSI) has enabled device-free sensing applications such as human activity recognition. However, CSI sensing models remain brittle in cross-domain deployment, where changes in users or environments can produce incorrect predictions. Existing solutions usually treat this problem as an offline model-design problem, by pretraining a stronger representation or applying one fixed adaptation method to the entire target domain. In practice, labeled target data are scarce and different classes may fail in different ways under the same domain shift. To address this, we propose CSI-Agent, an evidence-seeking LLM agent that reformulates cross-domain CSI adaptation as a deployment-time decision-making problem. Rather than processing raw CSI or making sample-level predictions, CSI-Agent summarizes target-domain behavior into sensing-grounded class-level evidence. It establishes a strong target-adaptive default from complementary CSI views and uses an LLM planner to determine whether each class should retain the default or invoke a specialized action. Deterministic verification and bounded execution further reduce unreliable interventions. We evaluate CSI-Agent on four public datasets using five cross-domain splits covering device, user, environment, and compositional shifts. Under 1-shot adaptation, CSI-Agent achieves the best target-domain performance across all splits and improves the average Macro-F1 by about 16\% compared to the strongest baseline method.
Chinese Translation
Wi-Fi信道状态信息(CSI)使得无设备感知应用(如人体活动识别)成为可能。然而,CSI感知模型在跨域部署中仍然脆弱,用户或环境的变化可能导致错误的预测。现有解决方案通常将此问题视为离线模型设计问题,通过预训练更强的表征或对整个目标域应用一种固定的自适应方法。在实践中,标记的目标数据稀缺,且不同类别在同一域偏移下可能以不同方式失败。为解决此问题,我们提出CSI-Agent,一个寻求证据的LLM智能体,将跨域CSI自适应重新表述为部署时的决策问题。CSI-Agent不是处理原始CSI或进行样本级预测,而是将目标域行为总结为基于感知的类别级证据。它从互补的CSI视图中建立一个强大的目标自适应默认值,并使用LLM规划器确定每个类别应保留默认值还是调用专门的动作。确定性验证和有界执行进一步减少不可靠的干预。我们在四个公共数据集上使用五个跨域划分(涵盖设备、用户、环境和组合偏移)评估CSI-Agent。在1-shot自适应下,CSI-Agent在所有划分上实现了最佳的目标域性能,并且与最强的基线方法相比,平均Macro-F1提高了约16%。
cs.AI / 19 / 2609.32000

SenseAgent: An LLM Agent for Adaptive Cross-Domain IMU Sensing

SenseAgent:一种用于自适应跨域IMU感知的LLM智能体
Zhao, Tianya, Liu, Chuan, Wang, Xuyu
Abstract
Deep learning has improved inertial measurement unit (IMU) sensing for mobile and wearable applications. However, an IMU model trained in one domain often becomes unreliable when it is used with a new user, device, or body position. Existing methods usually treat this problem as a static model-design task: they pretrain a stronger representation, add data augmentation, or select one adaptation method before deployment. In practice, the target domain is only gradually observed, labels are scarce, and different domain shifts require different sensing actions. This paper presents SenseAgent, an LLM-guided sensing agent for cross-domain IMU activity recognition. Instead of asking an LLM to classify raw IMU signals, SenseAgent uses the LLM as a runtime planner over sensing tools, source-domain experience memory, online target memory, and verifiers. The agent builds a label-free diagnosis report from the target stream and uses it to decide whether to keep raw inference or invoke specialized tools, including gravity-aware sensing, prototype transfer, and style normalization. Verifiers check source calibration, target-memory reliability, and no-harm criteria before accepting high-risk tool decisions. SenseAgent also supports scarce feedback without retraining the backbone or replacing the label-free route. This design converts cross-domain IMU sensing from a fixed inference pipeline into a closed-loop sensing process that diagnoses target shifts, selects suitable sensing actions, and rejects unsafe adaptations. We evaluate SenseAgent across multiple IMU datasets and deployment shifts. Results show that its verified route selection improves cross-domain sensing, especially under harder placement and compound shifts, and further benefits from limited user feedback.
Chinese Translation
深度学习改进了移动和可穿戴应用中的惯性测量单元(IMU)感知。然而,在一个域中训练的IMU模型在新用户、设备或身体位置使用时往往变得不可靠。现有方法通常将此问题视为静态的模型设计任务:预训练更强的表示、添加数据增强或在部署前选择一种自适应方法。在实践中,目标域只能逐渐被观察到,标签稀缺,并且不同的域偏移需要不同的感知动作。本文提出SenseAgent,一种LLM引导的用于跨域IMU活动识别的感知智能体。SenseAgent不是让LLM对原始IMU信号进行分类,而是将LLM用作感知工具、源域经验记忆、在线目标记忆和验证器之上的运行时规划器。该智能体从目标流构建无标签诊断报告,并用其决定是保留原始推理还是调用专门工具,包括重力感知传感、原型迁移和风格归一化。验证器在接受高风险工具决策前检查源校准、目标记忆可靠性和无害准则。SenseAgent还支持稀缺反馈,而无需重新训练骨干网络或替换无标签路径。这种设计将跨域IMU感知从固定的推理管道转变为闭环感知过程,能够诊断目标偏移、选择合适的感知动作并拒绝不安全的自适应。我们在多个IMU数据集和部署偏移上评估SenseAgent。结果表明,其经过验证的路由选择改进了跨域感知,尤其是在较难的佩戴位置和复合偏移下,并进一步受益于有限的用户反馈。
cs.AI / 20 / 2609.32002

What Does the Rank Buy? A Spectral and Distributional Analysis of Low-Rank Adaptation

秩能换来什么?低秩适应的谱与分布分析
Barazandeh, Babak
Abstract
The rank $r$ in LoRA is widely treated as a capacity control: a smaller rank is assumed to yield a simpler model that generalizes better. We show that, under hard per-factor norm budgets---the idealization of the weight decay and norm control used in practice---this intuition breaks down. The reason is structural: under such budgets, the updates LoRA can reach are exactly the matrices of rank at most $r$ inside a nuclear-norm ball, and every complexity and displacement functional we analyze is maximized over this set by a rank-one update---so the rank cap never binds. The consequences follow directly. The linear-readout model class we study is identical for every $r \ge 1$, its Rademacher complexity carries no dependence on $r$, and the distance the adaptation can move the source distribution obeys a rank-independent upper bound that we show is sharp. If rank does not control capacity, where does it act? We identify two places. Statistically, replacing the per-factor budgets with a joint budget on the product restores a data-dependent, rank-sensitive complexity bound---though the gain appears only for well-spread feature distributions, and the worst case remains rank-free. Spectrally, rank sets the price of adaptation: canceling the leading singular directions of the pretrained weight requires both sufficient rank and sufficient budget. We bound the smallest rank achieving a desired source--target alignment, with upper and lower bounds that match under two-sided spectral decay. Together, these results recast rank as governing which updates are reachable and what cancellation costs---not how much capacity the model has.
Chinese Translation
LoRA中的秩$r$被广泛视为一种容量控制:较小的秩被认为能产生更简单的模型,从而具有更好的泛化能力。我们表明,在硬性逐因子范数预算下——这是对实践中使用的权重衰减和范数控制的理想化——这种直觉不再成立。原因在于结构:在这种预算下,LoRA能够达到的更新恰好是核范数球内秩至多为$r$的矩阵,而我们分析的每一个复杂度和位移泛函在该集合上都被一个秩一更新最大化——因此秩上限从未构成约束。其影响直接显现。我们研究的线性读出模型类对于每个$r \ge 1$都是相同的,其Rademacher复杂度不依赖于$r$,并且适应能够移动源分布的距离服从一个与秩无关的上界,我们证明该上界是紧的。如果秩不控制容量,那么它在哪里起作用?我们识别出两个作用点。在统计上,用对乘积的联合预算替代逐因子预算,可以恢复一个依赖于数据、对秩敏感的复杂度界——尽管这种收益仅出现在特征分布良好分散的情况下,并且最坏情况仍然与秩无关。在谱上,秩设定了适应的代价:抵消预训练权重的主要奇异方向需要足够的秩和足够的预算。我们给出了达到所需源-目标对齐的最小秩的界,在双侧谱衰减下,上界和下界相匹配。总之,这些结果重新将秩解释为控制哪些更新可达以及抵消的代价——而非模型具有多少容量。
cs.AI / 21 / 2609.32019

Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding

去中心化Master-Mind:多智能体路径规划中通过迭代意图去噪进行联合动作精化
Vyaltsev, Valeriy, Andreychuk, Anton, Zlotnikova, Taisia, Yakovlev, Konstantin, Panov, Aleksandr, Skrynnik, Alexey
Abstract
Decentralized multi-agent path finding (MAPF) with communication requires agents to reach individual goals without collisions under partial observability. Learnable policies trained on expert data provide an effective approach to this problem. However, when several coordinated joint actions are valid in the same context, independently sampling from per-agent distributions can recombine locally valid choices into incompatible joint actions. This failure can arise from the final sampling mechanism even when the per-agent action distributions are learned correctly. DMM (Decentralized Master-Mind) addresses this by replacing one-shot action sampling with discrete, iterative refinement of action intents across communication rounds, inspired by denoising in diffusion models. Agents initialize random action intents and refine them through local communication, coupling their choices before commitment. DMM is pretrained with imitation learning on expert MAPF solutions and further optimized with MICPO, a critic-free group-relative reinforcement-learning method designed for multi-agent, multi-round action refinement. DMM generally achieves higher success rates and lower solution costs than the evaluated learnable baselines. On 1,600 MovingAI tasks, DMM fine-tuned with MICPO solves 1,598, the highest coverage among the evaluated methods, while achieving solution costs close to those of the strongest baselines. DMM also scales to over one million simultaneously acting agents in obstacle-rich environments. These results show that round-level intent refinement can improve joint-action coordination while preserving decentralized execution.
Chinese Translation
带通信的去中心化多智能体路径规划(MAPF)要求智能体在部分可观测条件下无碰撞地到达各自目标。在专家数据上训练的可学习策略为此问题提供了有效方法。然而,当在同一情境下存在多个有效的协调联合动作时,从各智能体分布独立采样可能将局部有效的选择重新组合成不兼容的联合动作。即使各智能体动作分布学习正确,这种失败也可能源于最终采样机制。DMM(Decentralized Master-Mind)通过将一次性动作采样替换为跨通信轮次的离散、迭代动作意图精化来解决该问题,其灵感来自扩散模型中的去噪。智能体初始化随机动作意图,并通过局部通信对其进行精化,在最终决策前耦合彼此的选择。DMM 在专家 MAPF 解上使用模仿学习进行预训练,并进一步用 MICPO 优化;MICPO 是一种无critic的组相对强化学习方法,专为多智能体、多轮动作精化设计。与所评估的可学习基线相比,DMM 通常实现更高的成功率和更低的求解成本。在 1,600 个 MovingAI 任务上,经 MICPO 微调的 DMM 求解了 1,598 个,在所评估方法中覆盖率最高,同时求解成本接近最强基线。DMM 还可扩展到障碍物丰富的环境中超过一百万个同时行动的智能体。这些结果表明,轮级意图精化可以在保持去中心化执行的同时改善联合动作协调。
cs.AI / 22 / 2609.32020

A Benchmark for LLM's Understanding of Middle School and High School Science Topics

LLM对初中和高中科学主题理解的基准测试
Schroeder, Noah L., Ambarwati, Yessy Eka, Zhang, Yuji, Zhai, ChengXiang
Abstract
Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understanding LLMs' performance on content directly relevant to secondary science curricula. To address this gap, we developed a comprehensive NGSS-aligned benchmark for both middle and high school science using a rigorous synthetic data pipeline, multi-judge validation, and item-level psychometric analysis. Nine open-weight LLMs were systematically evaluated using this benchmark, indicating that several smaller, locally deployable models achieved high accuracy across diverse science domains and question types. Our findings indicate that model size did not consistently predict performance, emphasizing the importance of intentional model selection for educational deployment. We then incorporated a human reviewer into the loop, reviewing the items generated by the LLMs for alignment with NGSS standards. The human review indicated that synthetically generated items were not in perfect alignment with the NGSS standards, indicating the benefits of human-in-the-loop item development, the need to explore the intersection of content and pedagogical knowledge, and the need to extend benchmarks to evaluate LLMs' capacity for interactive, evidence-based feedback in educational scenarios.
Chinese Translation
大语言模型(LLM)正日益融入教育环境,然而教育工作者缺乏稳健的、符合标准的工具来评估其在K-12科学情境中的有效性。现有基准主要评估一般语言能力或高级科学推理,在理解LLM对与中学科学课程直接相关内容的表现方面存在关键空白。为解决这一空白,我们使用严格的合成数据流水线、多评委验证和题目级心理测量分析,开发了一个面向初中和高中科学的、与NGSS对齐的综合性基准。九个开放权重的LLM使用该基准进行了系统评估,表明几个较小的、可本地部署的模型在不同的科学领域和问题类型上达到了高准确率。我们的发现表明,模型规模并不能一致地预测性能,强调了为教育部署进行审慎模型选择的重要性。然后我们在循环中加入了人类评审员,审查LLM生成的题目与NGSS标准的一致性。人工审查表明,合成生成的题目与NGSS标准并不完全一致,这表明了人在回路中开发题目的益处,探索内容知识与教学法知识交叉的必要性,以及扩展基准以评估LLM在教育场景中提供交互式、基于证据的反馈能力的必要性。
cs.AI / 23 / 2609.32035

Reasoning Concentrates Errors, and Self-Consistency Never Notices

推理使错误集中,而自一致性从未察觉
Althoubi, Asaad
Abstract
Self-consistency assumes that independent samples disagree when a model is unsure, so agreement is evidence of correctness. Holding weights fixed and toggling only a reasoning mode, over five benchmarks and 74,944 samples, we show that reasoning concentrates a model's errors: the probability that two independently drawn wrong answers coincide rises in all ten dataset-scale comparisons (p = 0.00098), and in nine of nine after restricting both arms to the problems each gets wrong. Where the answer space is unbounded, reasoning cuts the distinct answers produced to 0.43-0.65 of the non-reasoning count; where it is bounded, both arms hold an identical option set and reasoning concentrates mass on it instead, which no positional prior can explain at fixed weights. The aggregate cost is smaller than the mechanism predicts, because reasoning also shrinks the set of problems where answer diversity can decide anything, in ten of ten cells and by 2.7x; normalized for available headroom, both arms convert a quarter of it in domain. Confidence weighting does not recover what is left. Across 280 method-dataset-model combinations on eight models and five benchmarks, not one beats plain majority voting after correction; weighted voting agrees with it on 98.5% of problem-method pairs and is right 56.3% of the time on the rest; and a signal's direction can invert within fixed weights, with answer log-probability predicting correctness when reasoning is off and error when it is on. A learned six-signal combination gains nothing out of domain. Confidence signals should be evaluated on decisions, not on discrimination.
Chinese Translation
自一致性假设:当模型不确定时,独立样本会产生分歧,因此一致是正确性的证据。在固定权重、仅切换推理模式的条件下,跨越五个基准测试和74,944个样本,我们表明推理使模型的错误集中:两个独立抽取的错误答案相同的概率在全部十项数据集规模的比较中均上升(p = 0.00098),并且在将两组都限制在各自做错的问题上后,九项比较中九项均上升。当答案空间无界时,推理将产生的不同答案数量削减至非推理模式下的0.43-0.65倍;当答案空间有界时,两组持有相同的选项集,而推理则将质量集中在其上,这在固定权重下无法用位置先验来解释。总体成本小于机制所预测的,因为推理还缩小了答案多样性能够决定任何问题的问题集合,在全部十个单元中,缩小了2.7倍;对可用余量进行归一化后,两组在领域内都转化了其中的四分之一。置信度加权并不能恢复剩余的部分。在八个模型和五个基准测试上的280个方法-数据集-模型组合中,经校正后没有一个能击败简单多数投票;加权投票在98.5%的问题-方法对上与它一致,而在其余部分有56.3%的时间是正确的;并且信号的方向可以在固定权重下反转,答案对数概率在推理关闭时预测正确性,在推理开启时预测错误。一个学习到的六信号组合在领域外没有增益。置信度信号应在决策上评估,而不是在区分度上。
cs.AI / 24 / 2609.32046

Receiver-Conditioned Latent Communication gives 94% CacheBack

接收者条件化的潜在通信实现 94% CacheBack
Rossi, Maximillian, Raghunath, Prajwal, Xuan, Haoqing, Zhang, Yusen, Wu, Eugene
Abstract
Multi-agent systems distribute large contexts across agents that communicate to solve a task. Text messages are compact but require decoding and may omit evidence the receiving agent needs. Recent latent communication instead transfers KV caches. This avoids text generation and can improve accuracy and latency. However, a full KV cache grows linearly with both the context an individual agent processes, and the number of agents that coordinate together. This raises memory and context costs, often far exceeding available GPU resources and context window sizes. Our key observation is that agents need only send what the receiving agent requires for its local task -- which we call receiver-conditioned communication. The receiver agent passes the sender a small description of its information needs, which serves to filter and compress the sender agent's KV cache. CacheBack is a simple, robust, training-free instance of receiver conditioning based on the sender's attention weights. On FanOutQA, CacheBack with Qwen 3 removes 75% of the state the agent would otherwise receive, improving accuracy by 14.7 percentage points and reducing median task-completion latency by 3.2x relative to text communication. We show comparable improvements across model families that span dense Transformers, Mamba-attention hybrids, and sliding-window attention.
Chinese Translation
多智能体系统将大型上下文分布到通过通信来解决问题的多个智能体上。文本消息紧凑,但需要解码,并且可能遗漏接收智能体所需的证据。最近的潜在通信则转而传输 KV 缓存。这避免了文本生成,并可以提高准确性并降低延迟。然而,完整的 KV 缓存会随着单个智能体处理的上下文以及协同工作的智能体数量而线性增长。这增加了内存和上下文成本,通常远远超过可用的 GPU 资源和上下文窗口大小。我们的关键观察是,智能体只需发送接收智能体在其本地任务中所需的内容——我们称之为接收者条件化通信。接收智能体向发送方传递其信息需求的简要描述,用于过滤和压缩发送智能体的 KV 缓存。CacheBack 是一个简单、鲁棒、无需训练的接收者条件化实例,基于发送方的注意力权重。在 FanOutQA 上,使用 Qwen 3 的 CacheBack 移除了智能体原本会接收的 75% 的状态,将准确率提高了 14.7 个百分点,并将中位任务完成延迟相对于文本通信降低了 3.2 倍。我们在跨越密集 Transformer、Mamba-注意力混合模型和滑动窗口注意力的模型家族中展示了类似的改进。
cs.AI / 25 / 2609.32049

EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory

EngramRAG:面向多跳智能体记忆的动态使用加权拓扑与突触巩固
Potineni, Bhavyateja, Giri, Lohit, Jain, Anu, Kutsyy, Vadim, Pentakota, Rajasekhar
Abstract
As autonomous LLM agents are deployed across multi-session environments, conventional memory architectures suffer from Associative Blindness (inability to traverse multi-hop relational dependencies), Scaffolding Amnesia (temporal decay evicting core persona invariants), and Static Topology Stagnation (immutable graphs ignoring usage dynamics). Grounded in Complementary Learning Systems (CLS) principles, we propose EngramRAG, an adaptive memory architecture coupling a low-latency Waking State reflex with an asynchronous background Dreaming State consolidation cycle. EngramRAG introduces: (1) Usage-Modulated Personalized PageRank (U-PPR), where transition probabilities adapt via Hebbian plasticity to promote persistent entities into high-centrality Epistemic Macro-Hubs; (2) Consolidation-Activated Topology Decay (CATD), which scales retention half-life by topological load-bearing weight rather than wall-clock recency, protected by a cold-start grace period (N_grace >= 4); (3) Directed SUPERSEDES DAG filtering to suppress obsolete state during fact mutations; and (4) Triple-source hybrid retrieval fusing dense vectors, BM25, and U-PPR via dynamic Reciprocal Rank Fusion (RRF). Evaluating on all 1,982 QA pairs across 10 long-term conversations in the LoCoMo benchmark, EngramRAG achieves +38.9% relative improvement in Recall@5 (53.21% vs. 38.29%, p < 0.001) and +43.1% in MRR (0.4203 vs. 0.2937) over dense vector RAG, significantly outperforming Okapi BM25 (48.66%) and isolated static graph retrieval (8.50%). On temporal reasoning, EngramRAG reaches 62.33% Recall@5 (+16.67 points over dense vectors). In controlled mutation tests, SUPERSEDES suppresses split-brain hallucinations from 70.0% to 0.0%, while 90-day simulations show 100.0% scaffolding retention under a 26.21ms interactive retrieval reflex.
Chinese Translation
随着自主 LLM 智能体跨多会话环境部署,传统记忆架构面临关联盲区(无法遍历多跳关系依赖)、脚手架失忆(时间衰减驱逐核心人格不变量)和静态拓扑停滞(不可变图忽略使用动态)等问题。基于互补学习系统(CLS)原理,我们提出 EngramRAG,一种自适应记忆架构,将低延迟的清醒状态反射与异步后台梦境状态巩固循环耦合。EngramRAG 引入:(1) 使用调制个性化 PageRank(U-PPR),其中转移概率通过 Hebbian 可塑性自适应,以将持久实体提升为高中心性的认知宏枢纽;(2) 巩固激活拓扑衰减(CATD),其根据拓扑承载权重而非挂钟时间新近度来缩放保留半衰期,并由冷启动宽限期(N_grace >= 4)保护;(3) 有向 SUPERSEDES DAG 过滤,以在事实突变期间抑制过时状态;以及 (4) 三源混合检索,通过动态倒数排名融合(RRF)融合稠密向量、BM25 和 U-PPR。在 LoCoMo 基准的 10 个长期对话中的所有 1,982 个问答对上进行评估,EngramRAG 在 Recall@5 上相较于稠密向量 RAG 实现了 +38.9% 的相对提升(53.21% vs. 38.29%,p < 0.001),在 MRR 上实现了 +43.1% 的提升(0.4203 vs. 0.2937),显著优于 Okapi BM25(48.66%)和孤立静态图检索(8.50%)。在时间推理上,EngramRAG 达到 62.33% 的 Recall@5(比稠密向量提高 16.67 个百分点)。在受控突变测试中,SUPERSEDES 将裂脑幻觉从 70.0% 抑制到 0.0%,而 90 天模拟显示,在 26.21ms 的交互式检索反射下,脚手架保留率达到 100.0%。
cs.AI / 26 / 2609.32061

Contract monitoring: governing AI via separation of powers

合同监控:通过权力分立治理人工智能
Boix-Adsera, Enric
Abstract
We propose an AI safety framework that binds worker agents to contracts specifying their permitted actions. We show how these contracts can be enforced and specified by assigning distinct responsibilities to monitor agents and judges, and asymmetric computational resources to monitors and workers. Our framework allows us to empirically measure statistical safety guarantees. The framework applies to a wide range of settings, including code security and escape-the-box scenarios.
Chinese Translation
我们提出了一个AI安全框架,该框架将工作者智能体(worker agents)绑定到规定其被允许行为的合同上。我们展示了如何通过为监控智能体(monitor agents)和裁判者(judges)分配不同职责,并为监控者与工作者分配不对称的计算资源,来强制执行和具体规定这些合同。我们的框架使我们能够以实证方式测量统计安全性保证。该框架适用于广泛的场景,包括代码安全和“逃出盒子”(escape-the-box)场景。
cs.AI / 27 / 2609.32081

Toward Interactive Understanding of Code APIs

面向代码 API 的交互式理解
Ashok, Dhananjay, Thomason, Jesse, May, Jonathan
Abstract
Empowered by advances in Language Model agents, systems have made substantial strides in code generation and understanding. However, these approaches often rely on read access to the relevant code, an assumption which does not hold when dealing with external APIs. In this work, we introduce the PAU (Python API Understanding) benchmark, where we provide models with black-box, API-level access to code snippets. Models must query the API with exploratory inputs and draw insights from the resulting outputs, with the goal of describing the snippet's true functionality. By treating the code snippets as external tools that must be understood via interaction alone, PAU studies the more general problem of unsupervised tool understanding, specifically for tools implemented as Python methods. Despite recent progress in coding agents, even frontier models struggle to achieve high performance on PAU, with the best model (Claude-4-Opus) failing to understand over 45% of the PAU test set. An investigation into the common error modes reveals that models are overconfident; they often overrate the quality of their current hypothesis, leading to insufficient exploration and premature termination. Finally, we take inspiration from the Asymmetric Actor Critic (AAC) paradigm, frequently used in robot learning, to post-train models for interactive code understanding. Models trained with AAC conduct more active exploration of the APIs, with an AAC-tuned Qwen3-8B model matching the performance of GPT-5-mini.
Chinese Translation
借助语言模型智能体的进展,系统在代码生成与理解方面取得了长足进步。然而,这些方法通常依赖于对相关代码的读取访问权限,而这一假设在处理外部 API 时并不成立。在这项工作中,我们引入了 PAU(Python API 理解)基准测试,其中我们为模型提供对代码片段的黑盒、API 级别的访问。模型必须使用探索性输入查询 API,并从结果输出中获取洞见,目标是描述代码片段的真实功能。通过将代码片段视为必须仅通过交互来理解的外部工具,PAU 研究了更一般的无监督工具理解问题,特别是针对以 Python 方法实现的工具。尽管编码智能体最近取得了进展,但即使是前沿模型也难以在 PAU 上取得高性能,最佳模型(Claude-4-Opus)未能理解 PAU 测试集中超过 45% 的样本。对常见错误模式的调查表明,模型过度自信;它们经常高估当前假设的质量,导致探索不足和过早终止。最后,我们从机器人学习中常用的非对称演员-评论家(AAC)范式中获得灵感,对模型进行后训练以实现交互式代码理解。使用 AAC 训练的模型对 API 进行更主动的探索,经 AAC 调优的 Qwen3-8B 模型达到了 GPT-5-mini 的性能。
cs.AI / 28 / 2609.32091

Memory as Middleware for Self-Improving AI Agents

记忆作为中间件:面向自我改进的AI智能体
Jayaram, K. R., Isahagian, Vatche, Muthusamy, Vinod, Thomas, Gegi, Oum, Punleuk, Fang, Gaodan, Aravindan, Ashwath Vaithinathan
Abstract
AI agents are stateless across sessions by default and therefore operationally amnesic: each session begins with little durable knowledge of prior failures, repairs, preferences, or successful strategies. As a result, agents repeat the same mistakes and discard hard-won experience. The dominant fix is \emph{bespoke memory}---retrieval, persistence, and learning logic hand-wired into one agent and bound to one storage engine. This creates a fragmented landscape where memory cannot be swapped, shared, isolated, or reasoned about independently of the agent that owns it. We argue that this is a middleware problem: agent memory deserves a first-class, pluggable layer, just as data access, messaging, and persistence each became middleware concerns. We develop this vision through six systems challenges: two-sided pluggability, host-native interposition, multi-tenant isolation, write-path consistency, federated sharing with provenance, and lifecycle governance. We present ALTK-Evolve, a reference implementation of memory middleware for self-improving agents, and use it to motivate a broader research agenda for future memory middleware.
Chinese Translation
AI智能体默认跨会话是无状态的,因此在操作上是失忆的:每个会话开始时,对之前的失败、修复、偏好或成功策略几乎没有持久知识。因此,智能体重复同样的错误,丢弃来之不易的经验。主流的解决方案是定制记忆——将检索、持久化和学习逻辑手工接入单个智能体,并绑定到单个存储引擎。这导致了一种碎片化的局面,其中记忆无法独立于拥有它的智能体进行交换、共享、隔离或推理。我们认为这是一个中间件问题:智能体记忆应该成为一个一流的、可插拔的层,就像数据访问、消息传递和持久化各自成为中间件关注点一样。我们通过六个系统挑战来发展这一愿景:双向可插拔性、宿主原生介入、多租户隔离、写路径一致性、带溯源的联邦共享以及生命周期治理。我们提出了ALTK-Evolve,一个用于自我改进智能体的记忆中间件的参考实现,并用它来推动未来记忆中间件的更广泛研究议程。
cs.AI / 29 / 2609.32093

GameBoyWorlds: A Testbed for Self-Improvement in Embodied Video Games

GameBoyWorlds:具身视频游戏中自我改进的测试平台
Ashok, Dhananjay, Shen, Adam, Feng, Aslan Huo, Khanna, Chinmay, Huang, Jun Rui, Sarmukaddam, Raghav, Natarajan, Surendira Balaji, Cui, Xiaotong, Zhang, Xincan, Yen, Thomson, Namkoong, Hongseok, May, Jonathan, Thomason, Jesse
Abstract
Powered by expert guidance, agents can operate in interactive environments; however, it is unclear whether they can learn autonomously from their own experience. To evaluate such self-improvement methods, we introduce GameBoyWorlds, a testbed for agentic self-improvement in video games. GameBoyWorlds-Execution evaluates task execution on a collection of 5 distinct game series. Agents are allowed access to dedicated training games but are provided no demonstrations, documentation, or rewards. Agents must ground themselves in the environment through self-directed exploration and by inferring actionable knowledge from their own experience. At test time, agents must complete short-horizon tasks that evaluate their ability to navigate, interact, and engage with game-specific mechanics in unseen games. Out-of-the-box frontier models complete fewer than 50% of the 500 tasks due to failures in multimodal grounding, establishing that self-improvement methods have room to push performance. We demonstrate that contemporary approaches to self-improvement are lacking, with world modelling and autonomous skill discovery failing, and a novel strategy that uses curiosity-based exploration to write guides achieving only partial success. GameBoyWorlds-Playthrough tests end-to-end game completion in two fan-made Pok\'emon games. We show that while frontier models have been pre-exposed to official releases such as Pok\'emon Red, they lack essential information on the games in our testbed. Instead of relying on their parametric knowledge to succeed, agents must learn from their own experience and autonomously improve over the course of the playthrough. We show that a sophisticated agentic pipeline with multimodal memory and hierarchical subgoals fails to reach even the first major milestone in both games, establishing GameBoyWorlds as an ambitious target for self-improving agents.
Chinese Translation
在专家指导的驱动下,智能体能够在交互式环境中行动;然而,它们能否从自身经验中自主学习尚不清楚。为评估此类自我改进方法,我们提出 GameBoyWorlds,一个用于视频游戏中智能体自我改进的测试平台。GameBoyWorlds-Execution 在一组 5 个不同游戏系列上评估任务执行。智能体可以访问专门的训练游戏,但不会被提供演示、文档或奖励。智能体必须通过自主探索,并从自身经验中推断可操作知识,从而在环境中实现自我 grounding。在测试时,智能体必须完成短时程任务,这些任务评估其在未见过的游戏中导航、交互以及运用游戏特定机制的能力。开箱即用的前沿模型由于多模态 grounding 失败,在 500 个任务中完成不到 50%,这表明自我改进方法仍有提升性能的空间。我们表明,当前的自我改进方法仍存在不足:世界建模和自主技能发现均告失败,而一种利用好奇心驱动探索来编写指南的新策略也仅取得部分成功。GameBoyWorlds-Playthrough 在两个粉丝制作的 Pokémon 游戏中测试端到端游戏通关。我们表明,尽管前沿模型已预先接触过诸如 Pokémon Red 等官方发行版本,但它们缺乏关于我们测试平台中游戏的关键信息。智能体不能依赖其参数化知识来取得成功,而必须从自身经验中学习,并在整个通关过程中自主改进。我们表明,一个具有多模态记忆和分层子目标的复杂智能体流程在两个游戏中甚至未能达到第一个主要里程碑,这确立了 GameBoyWorlds 作为自我改进智能体雄心勃勃的目标。
cs.AI / 30 / 2609.32102

Residual Streams Read, Recurrent States Remember: The Global Workspace in Mamba Models

残差流读取,循环状态记忆:Mamba 模型中的全局工作空间
Wang, Wenlong, Reid, Fergal
Abstract
Can the global-workspace account of transformer representations extend to state-space language models? We fit Jacobian lenses to the residual streams and recurrent states of Mamba-1, Mamba-2 and Mamba-3, using the original 1000-prompt recipe. Joint residual--state readouts improve recovery of known intermediate concepts over the residual lens on at least five of six task families in every tested Mamba checkpoint. On Mamba-2, state alone exceeds residual and logit lenses on all six families; a normalised joint readout improves on both components on five. Temporal maps and word-list experiments show earlier content remaining state-readable as residual visibility changes. We also propose sign-guarded steering, which improves target top-five success over coordinate exchange on matched verbal-report trials in five models. Recurrent state alone supports this verbal access. These gains do not extend consistently to relational answers: guarded edits often output the edited concept itself, and Mamba-3's joint edits can disrupt successful state-only redirection. Recurrent state thus provides a complementary carrier of workspace content, whose recovery, persistence and causal uses require separate measurements.
Chinese Translation
全局工作空间对 Transformer 表征的解释能否扩展到状态空间语言模型?我们使用原始的 1000 个提示词配方,将雅可比透镜拟合到 Mamba-1、Mamba-2 和 Mamba-3 的残差流与循环状态上。联合残差-状态读出在每个测试的 Mamba 检查点上,至少在六个任务族中的五个上,相比残差透镜提高了对已知中间概念的恢复。在 Mamba-2 上,仅状态就在所有六个任务族上超过了残差透镜和 logit 透镜;归一化联合读出在五个任务族上优于两个组件。时间图和词表实验表明,随着残差可见性的变化,较早的内容仍保持状态可读。我们还提出了符号保护引导,在五个模型的匹配言语报告试验中,相比坐标交换提高了目标前五成功率。仅循环状态就支持这种言语访问。这些增益并未一致地扩展到关系性答案:保护性编辑通常输出被编辑的概念本身,而 Mamba-3 的联合编辑可能破坏成功的仅状态重定向。因此,循环状态提供了工作空间内容的补充载体,其恢复、持续性和因果用途需要单独测量。
cs.AI / 31 / 2609.32116

Escaping Alignment: A Physical Trap Model of Best-of-N Jailbreaking

逃离对齐:Best-of-N越狱的物理陷阱模型
Biroli, Marco
Abstract
Best-of-$N$ jailbreaking (BoN) bypasses safeguards of aligned models by drawing $N$ independent augmentations of an unsafe prompt and sampling $M$ completions of each. Previous works have shown that the attack success rate (ASR) seems to follow a power-law in $N$, which we challenge. The exponent drifts with $N$, with an exponential crossover which is a finite-size artifact of the adversarial dataset. Little work has been done to explore the entire two-budget ($N, M$) attack surface as well as its dependence on the generation temperature $T$. We introduce a simple barrier model where each prompt has a baseline safety level and each augmentation a random thermally activated barrier. Then four numbers, each backed by an interpretable safety mechanism, determine the entire ($N, M$) attack surface. They extrapolate predictions from $N \leq 100$ to $N = 10^4$, collapse five distinct models on the same scaling function and predict ASR at different temperatures from the one they were fitted at.
Chinese Translation
Best-of-$N$越狱(BoN)通过抽取不安全提示的$N$个独立增强,并对每个增强采样$M$个补全,从而绕过对齐模型的安全防护。先前的工作表明,攻击成功率(ASR)似乎随$N$呈幂律分布,我们对此提出质疑。该指数随$N$漂移,并出现指数交叉,这是对抗数据集的有限尺寸效应。很少有工作探索整个双预算($N, M$)攻击面及其对生成温度$T$的依赖。我们引入一个简单的势垒模型,其中每个提示有一个基线安全水平,每个增强有一个随机热激活势垒。然后,四个数字,每个都由可解释的安全机制支持,决定了整个($N, M$)攻击面。它们将预测从$N \leq 100$外推到$N = 10^4$,将五个不同模型坍缩到同一缩放函数上,并预测与拟合温度不同的温度下的ASR。
cs.AI / 32 / 2609.32123

READ-Bench: Benchmarking Historical Instance Retrieval for Time-Series Diagnosis

READ-Bench:面向时间序列诊断的历史实例检索基准测试
Pastrana, Gerardo, Li, Haojun, Mehta, Dhruv, Vyas, Anoushka, Pakazad, Sina Khoshfetrat, Ohlsson, Henrik, Paparrizos, John
Abstract
Time-series diagnostic systems rarely rely on retrieving relevant historical cases, and when they do, retrieval is evaluated only indirectly through downstream prediction. We introduce READ-Bench, a benchmark for historical-case retrieval across 12 diagnostic datasets, centered on multivariate time series, that defines relevance by shared fault or event type rather than signal shape, so visually different traces of the same fault count as relevant while similar-looking traces of different faults do not. Treating retrieval as a base retriever followed by a reranker, we evaluate classical distances, symbolic retrievers, self-supervised and foundation-model embedders, and their fusion, plus label-aware and language-model rerankers, under one protocol that varies supervision, pollution, and corpus scale with significance testing. Under a common channel-independent interface, pretrained representations offer no statistically detectable advantage over strong classical and symbolic baselines for search alone. The decisive factor is a small amount of resolved-case supervision at reranking, namely a Gaussian-process reranker that propagates a few neighbor labels in embedding space, which helps far more than more sophisticated representations or language-model reasoning and holds under pollution and at full corpus scale. Guided by these findings, we fuse a normal-residual-scored embedder with a dynamic time warping leg via reciprocal-rank fusion, then rerank with the Gaussian-process reranker, improving NDCG@10 over its own search stage on all 12 datasets, by +0.11 from reranking and +0.16 over the strongest single base retriever.
Chinese Translation
时间序列诊断系统很少依赖检索相关的历史案例,当它们依赖时,检索也仅通过下游预测进行间接评估。我们提出 READ-Bench,一个跨 12 个诊断数据集的历史案例检索基准,以多变量时间序列为中心,其相关性由共享的故障或事件类型定义,而非信号形状,因此同一故障的视觉上不同的轨迹被视为相关,而不同故障的相似轨迹则不被视为相关。我们将检索视为基础检索器后接重排序器,并在一个协议下评估经典距离、符号检索器、自监督和基础模型嵌入器及其融合,以及标签感知和语言模型重排序器,该协议通过显著性检验改变监督、污染和语料库规模。在通用的通道独立接口下,对于仅搜索任务,预训练表示相比强经典和符号基线没有统计上可检测的优势。决定性因素是在重排序时少量已解决案例的监督,即一个高斯过程重排序器,它在嵌入空间中传播少量邻居标签,其帮助远大于更复杂的表示或语言模型推理,并且在污染下和完整语料库规模下仍然成立。在这些发现的指导下,我们通过倒数排名融合将正常残差评分的嵌入器与动态时间规整分支融合,然后用高斯过程重排序器重排序,在所有 12 个数据集上相比其自身搜索阶段提高了 NDCG@10,重排序贡献 +0.11,相比最强单一基础检索器提高 +0.16。
cs.AI / 33 / 2609.32166

PastForward: Faster On-Device GUI Agents via Computational Experience Reuse

PastForward:通过计算经验重用实现更快的设备端GUI智能体
Park, Taehwan, Lee, Changmin, Lee, Hayeon, Gong, Taesik
Abstract
Running GUI agents on edge devices can keep sensitive screens and interaction histories local, but the computational cost of inference at every action step makes deployment challenging. Existing GUI agent systems either perform full vision-language model (VLM) inference at each action step or reuse coarse-grained knowledge matched to prior tasks. However, dynamic mobile environments and user tasks make it difficult to fully utilize prior task executions without additional fine-tuning or task-specific offline exploration. To address this challenge, we present PastForward, a system that accelerates GUI agents through validated, fine-grained reuse of computational experience accumulated during ordinary task execution. During decoding, PastForward retrieves prior output sequences as device-adaptive multi-token proposals and verifies them in a single VLM forward pass. Across action steps, it uses prior GUI transitions to begin next-step inference while the device executes the current action, retains the early computation only when the predicted screen matches the observed screen, and carries reusable KV states forward. We evaluate PastForward on AndroidWorld workloads derived from real mobile usage patterns using multiple VLM backbones across server and edge platforms. On device, PastForward achieves action-step latency speedups of 1.63-2.36$\times$ while maintaining task success rates.
Chinese Translation
在边缘设备上运行 GUI 智能体可以将敏感屏幕和交互历史保留在本地,但在每个动作步骤进行推理的计算成本使得部署具有挑战性。现有的 GUI 智能体系统要么在每个动作步骤执行完整的视觉语言模型(VLM)推理,要么重用与先前任务匹配的粗粒度知识。然而,动态的移动环境和用户任务使得在没有额外微调或任务特定的离线探索的情况下,难以充分利用先前的任务执行。为了解决这一挑战,我们提出了 PastForward,一个通过验证的、细粒度的重用普通任务执行过程中积累的计算经验来加速 GUI 智能体的系统。在解码过程中,PastForward 检索先前的输出序列作为设备自适应的多令牌提议,并在单次 VLM 前向传播中验证它们。跨动作步骤,它使用先前的 GUI 转换在设备执行当前动作时开始下一步推理,仅当预测屏幕与观察到的屏幕匹配时保留早期计算,并向前传递可重用的 KV 状态。我们在源自真实移动使用模式的 AndroidWorld 工作负载上,使用跨服务器和边缘平台的多个 VLM 主干网络评估 PastForward。在设备上,PastForward 在保持任务成功率的同时,实现了 1.63-2.36 倍的动作步骤延迟加速。
cs.AI / 34 / 2609.32172

Noisy Test-Time Reinforcement Learning for Code LLMs

面向代码LLM的噪声测试时强化学习
Yang, Xikai, Nguyen, Hieu Trung, Xu, Dunyuan, Zhao, Yuzhi, Li, Jinpeng, Ma, Wenao, Heng, Pheng-Ann
Abstract
Large language models (LLMs) have demonstrated remarkable performance across various code-related tasks. However, unlike carefully curated datasets that are typically high-quality and error-free, real-world user instructions are often vague and error-prone, posing significant challenges to the robustness of code LLMs. Furthermore, robustness-oriented fine-tuning relies on paired clean-noisy samples, which are costly to curate and require sophisticated noisy simulation techniques. To address these challenges, we propose the Noisy Test-time Reinforcement Learning framework (NTRL-Code), which enables robust self-evolution of code LLMs using only unlabeled noisy data during the testing stage. Specifically, NTRL-Code uses conservative self-denoising to obtain a cleaner semantic anchor for target estimation, and employs an abstract-syntax-tree (AST)-based structural aggregation mechanism to estimate a proxy target from multiple candidate programs. The policy is then optimized on the original noisy prompts with a hybrid reward that combines format validity, code similarity, and anti-repetition signals. Extensive experiments on three benchmarks, each incorporating character-level, word-level, and paragraph-level perturbations, demonstrate that NTRL-Code yields robust and consistent improvements, stabilizing the predictions of various base models. Our code is available at https://github.com/Xikai97/NTRL-Code.
Chinese Translation
大语言模型(LLM)在各种代码相关任务中已展现出卓越的性能。然而,与通常高质量且无错误的精心整理数据集不同,现实世界中的用户指令往往模糊且容易出错,这对代码LLM的鲁棒性提出了重大挑战。此外,面向鲁棒性的微调依赖于成对的干净-噪声样本,这些样本的整理成本高昂,并且需要复杂的噪声模拟技术。为了解决这些挑战,我们提出了噪声测试时强化学习框架(NTRL-Code),该框架仅使用测试阶段中的无标签噪声数据,即可实现代码LLM的鲁棒自进化。具体而言,NTRL-Code使用保守的自去噪来获得更干净的语义锚点以进行目标估计,并采用基于抽象语法树(AST)的结构聚合机制,从多个候选程序中估计代理目标。然后,策略在原始噪声提示上进行优化,使用结合了格式有效性、代码相似性和抗重复信号的混合奖励。在三个基准测试上的大量实验(每个基准测试都包含字符级、词级和段落级扰动)表明,NTRL-Code产生了稳健且一致的改进,稳定了各种基础模型的预测。我们的代码可在 https://github.com/Xikai97/NTRL-Code 获取。
cs.AI / 35 / 2609.32184

AI Harness: Certification under Proposal-Conditioned Information for Foundation-Model Agents

AI Harness:基础模型智能体在提议条件信息下的认证
Zhong, Hailin, Zhu, Shengxin
Abstract
Foundation-model agents are often modeled as policies over an observed state. In deployed systems, however, a runtime may intervene only after the model has emitted a semantic proposal, making the proposal both an action candidate and a decision-time observation generated by a history-conditioned process. We show that collapsing this structure into a state-only proposal envelope can preserve proposal coverage while destroying certifiability. In a finite robust interface, the viability kernel of the collapsed model is contained in the physical projection of the history-augmented kernel, and the collapse is lossless exactly when every proposal-conditioned collapsed fiber retains a common robust-safe intervention. This gap can be maximal even with constant-size proposal and history alphabets. The same common-action condition yields a dual result: observing the current proposal can restore robust feasibility when it separates latent modes requiring incompatible interventions. We extend these one-step results over time using exact finite beliefs and standard safety and reachability fixed points, separating indefinite operational viability from finite worst-case verified progress. Controlled model-in-the-loop tests reproduce the predicted obstructions when telemetry or effect verification is removed or intervention authority is restricted. Thus, our contribution is not a new fixed-point calculus, but a characterization of when proposal--history correlation at the model--tool boundary is necessary for certification.
Chinese Translation
基础模型智能体通常被建模为基于观测状态的策略。然而,在部署系统中,运行时可能仅在模型发出语义提议后介入,这使得提议既是动作候选,也是由历史条件过程生成的决策时观测。我们表明,将此结构坍缩为仅状态的提议包络可以保留提议覆盖,但会破坏可认证性。在有限鲁棒接口中,坍缩模型的生存核包含在历史增强核的物理投影中,并且当每个提议条件下的坍缩纤维保留一个共同鲁棒安全干预时,坍缩恰好无损。即使提议和历史字母表大小恒定,这种差距也可能达到最大。相同的共同动作条件产生了一个对偶结果:观察当前提议可以恢复鲁棒可行性,当它分离出需要不兼容干预的潜在模式时。我们使用精确有限信念以及标准安全与可达性不动点,将这些单步结果扩展到时间上,将不确定运行生存性与有限最坏情况验证进展区分开来。受控的模型在环测试在遥测或效果验证被移除或干预权限受限时复现了预测的障碍。因此,我们的贡献不是一种新的不动点演算,而是刻画了模型-工具边界处的提议-历史相关性何时对认证是必要的。
cs.AI / 36 / 2609.32192

CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies

CoMemBench:跨多智能体工作流拓扑的协作记忆边界基准测试
Zhao, Sen, Kong, Ruiqi, Zhang, Zuyu, Shen, Lifeng, He, Xinyu, Zou, Ding, Zhang, Xu, Zhang, Qinghua
Abstract
Multi-agent workflows require task-relevant information to be shared across agents, while irrelevant, stale, unverified, or incompatible information must remain isolated. We call this task-conditioned scope of information a collaborative memory boundary. Workflow topology determines which intermediate artifacts are applicable to which downstream workers and when they cease to be valid, thereby providing a structural stress dimension for sharing and isolation. Existing memory benchmarks primarily evaluate retention and retrieval, whereas multi-agent benchmarks emphasize coordination and end-to-end completion, leaving topology-conditioned memory boundaries largely unmeasured. We introduce CoMemBench, an execution-grounded benchmark for collaborative memory sharing and isolation across multi-agent workflow topologies. It constructs 800 composite workflows across four domains from source-grounded dependency graphs, with node-local specifications, verifiable artifact handoffs, native evaluators, and matched isolation challenges. CoMemBench measures workflow completion, verified node progress, required-handoff reliability, isolation robustness, and token cost. Experiments reveal a sharing-isolation trade-off: broader context improves information availability but can weaken isolation, while system rankings shift across topologies and artifact violations.
Chinese Translation
多智能体工作流要求任务相关信息在智能体之间共享,而无关的、过时的、未经验证的或不兼容的信息必须保持隔离。我们将这种任务条件下的信息范围称为协作记忆边界。工作流拓扑决定了哪些中间产物适用于哪些下游工作者以及何时不再有效,从而为共享和隔离提供了一个结构性压力维度。现有的记忆基准主要评估保持和检索,而多智能体基准则强调协调和端到端完成,导致拓扑条件下的记忆边界在很大程度上未被测量。我们提出了CoMemBench,一个基于执行的基准,用于跨多智能体工作流拓扑的协作记忆共享与隔离。它从基于源的依赖图构建了跨四个领域的800个复合工作流,具有节点局部规范、可验证的产物交接、原生评估器和匹配的隔离挑战。CoMemBench测量工作流完成度、已验证节点进度、所需交接可靠性、隔离鲁棒性和token成本。实验揭示了共享-隔离权衡:更广泛的上下文提高了信息可用性,但可能削弱隔离,而系统排名在不同拓扑和产物违规之间发生变化。
cs.AI / 37 / 2609.32201

Instruct, Not Answer: Using Instruction Privileges in On-Policy Context Distillation

指导而非作答:在基于策略的上下文蒸馏中使用指令特权
Yu, Hantao, Han, Sandy, Ghai, Udaya, Erata, Ferhat, Lilien, Joe, Goel, Aman, Torkamani, Ali
Abstract
On-Policy Context Distillation (OPCD) has recently emerged as a powerful technique for transferring context to student models and for self-improvement. In OPCD, the teacher is conditioned on privileged information, and the goal is to minimize the Kullback-Leibler (KL) divergence between the privileged teacher and the student, evaluated on student-generated tokens. Many existing studies show that using instance-specific gold answers or gold demonstrations as the default privilege can hurt training performance, especially out-of-distribution (OOD). In this work, we instead design general instructions that target common student mistakes observed on the training samples, and show that such simple instructions can outperform gold as the OPCD privilege. In autoformalization tasks, using a matched formatting instruction as the privilege could outperform gold in OOD accuracy by a large margin. In 7 out of 8 experiments using ProverQA, ProofWriter, and ProntoQA as datasets, and Qwen3-Thinking and Olmo3-Thinking families as models, matched instruction privileges outperform gold in OOD by 4 to 17 points, while remaining on par with gold in-domain. Each instruction is only a few sentences (and thus contains much less information compared to all instance-specific gold) and is applied uniformly to every training sample. These results indicate that a general instruction, which applies equally to source and target domain examples, can be substantially more transferable than instance-specific gold in OPCD while maintaining in-domain performance.
Chinese Translation
基于策略的上下文蒸馏(OPCD)最近作为一种强大的技术出现,用于将上下文迁移到学生模型以及自我改进。在OPCD中,教师模型以特权信息为条件,目标是在学生生成的token上最小化特权教师与学生之间的Kullback-Leibler (KL) 散度。许多现有研究表明,使用实例特定的标准答案或标准演示作为默认特权可能会损害训练性能,尤其是在分布外(OOD)的情况下。在本工作中,我们转而设计针对训练样本上观察到的常见学生错误的通用指令,并表明这种简单的指令作为OPCD特权可以优于标准答案。在自动形式化任务中,使用匹配的格式化指令作为特权可以在OOD准确率上大幅超越标准答案。在8个实验中的7个里,使用ProverQA、ProofWriter和ProntoQA作为数据集,以及Qwen3-Thinking和Olmo3-Thinking系列作为模型,匹配的指令特权在OOD上比标准答案高出4到17个百分点,同时在域内与标准答案持平。每条指令只有几句话(因此与所有实例特定的标准答案相比,包含的信息少得多),并且统一应用于每个训练样本。这些结果表明,在OPCD中,一条同时适用于源域和目标域样本的通用指令,可以比实例特定的标准答案具有显著更好的可迁移性,同时保持域内性能。
cs.AI / 38 / 2609.32208

Witness: Discovery, Deciphering, and Epiphany in Interactive Puzzle Environments

Witness:交互式谜题环境中的发现、破译与顿悟
Ning, Guanghan, Liu, Ping, Li, Linyi, Zheng, Huangjie, Neervannan, Arjun, Nguyen, Huu, Sklar, Michael, Zorlu, Deniz, Ouporov, Nicolai
Abstract
Automated science needs agents that can work out the rules of an unfamiliar environment by interacting with it. Interactive rule-discovery puzzles offer a controlled setting for studying this ability: an agent infers hidden rules through experimentation and uses what it has inferred to reach a stated goal. We ask what limits current language models on these puzzles and whether reinforcement learning (RL) improves performance on rules held out from training. To study both, we introduce WITNESS, a 2D grid-based puzzle environment with ground-truth ASCII observations and controlled access to rules. An agentic pipeline generates games for WitnessGym, the RL training suite, and WitnessBench, comprising public validation and private test games. The validation set separately tests new compositions of trained rule primitives and primitives absent from training. Under a shared harness, the best of 18 frontier proprietary and open-weight models solves only 24\% of private test level slots, with scores sensitive to the observation interface and agent configuration. Providing ground-truth rules raises Opus-5's validation RHAE-L5 (relative human action efficiency over the first five levels) from 59.9 to 97.8, whereas a 27B open-weight model gains only 2.1 points and remains limited even with the rules provided. RL on WitnessGym raises the 27B model's private test RHAE-L5 from 2.1 to 5.4 and yields a mean gain of 4.1 points on four external discovery benchmarks. Together, these results point to rule acquisition as a major difficulty for frontier models like Opus-5 while smaller models further struggle on rule-based execution, and indicate that RL on hidden-rule puzzles transfers to broader rules and real-world tasks beyond training. Benchmark is available at: https://witnessbench.ai
Chinese Translation
自动化科学需要能够通过与陌生环境交互来弄清楚其规则的智能体。交互式规则发现谜题提供了一种受控环境来研究这一能力:智能体通过实验推断隐藏规则,并利用其推断出的规则来达成既定目标。我们探究当前语言模型在这些谜题上的限制因素,以及强化学习(RL)能否提升在训练阶段未见规则上的表现。为研究这两点,我们提出了 WITNESS,一个基于 2D 网格的谜题环境,具有真值 ASCII 观测和受控的规则访问。一个智能体流水线为 WitnessGym(RL 训练套件)和 WitnessBench 生成游戏,后者包含公开验证集和私有测试集游戏。验证集分别测试已训练规则原语的新组合以及训练中未出现的原语。在共享测试框架下,18 个前沿专有和开放权重模型中的最佳模型仅解决了 24% 的私有测试关卡槽位,且分数对观测接口和智能体配置敏感。提供真值规则可将 Opus-5 的验证 RHAE-L5(前五个关卡上相对于人类行动效率)从 59.9 提升至 97.8,而一个 27B 开放权重模型仅提升 2.1 分,且即便提供规则仍受限。在 WitnessGym 上进行 RL 可将 27B 模型的私有测试 RHAE-L5 从 2.1 提升至 5.4,并在四个外部发现基准上平均提升 4.1 分。总之,这些结果表明,规则获取是 Opus-5 等前沿模型的主要困难,而较小模型在基于规则的执行上更加吃力,并表明在隐藏规则谜题上的 RL 可迁移到训练之外的更广泛规则和真实世界任务。基准可在以下网址获取:https://witnessbench.ai
cs.AI / 39 / 2609.32209

Fracast-0: Fractal Weight Sharing for a Time Series Foundation Model with Only 85K Parameters

Fracast-0:面向仅85K参数的时间序列基础模型的分形权重共享
Zhan, Tianxiang, Zhang, Huanyao, He, Yuanpeng
Abstract
Time series foundation models must preserve multi-domain breadth, probabilistic output, and multiple temporal scales, but parameter count grows when each scale receives a separate representation. We introduce Fracast-0, a probabilistic forecasting foundation model that exploits temporal self-similarity to reuse one operator across scales. A parameter-free detector extracts significant seasonal structure. The encoder applies a shared local block along a geometric dilation ladder with scale conditioning, while the decoder combines context-gathered states with an explicit seasonal future state and reuses a second block along another ladder before emitting nine quantiles. Pretraining across six corpora preserves multi-domain breadth within 85,001 parameters. On 97 GIFT-Eval configurations without per-dataset fine-tuning, Fracast-0 is the smallest of 28 evaluated checkpoints and remains non-dominated in the aggregate parameter-accuracy plane with MASE 0.808 and WQL 0.564. It uses 42.0% fewer parameters than TinyCast, whose MASE and WQL are 4.2% and 3.3% lower. These results support cross-scale weight reuse as a practical route to further time series foundation model compression.
Chinese Translation
时间序列基础模型必须保持多领域广度、概率输出和多个时间尺度,但当每个尺度获得单独的表示时,参数量会增加。我们引入了Fracast-0,一种概率预测基础模型,它利用时间自相似性来跨尺度重用同一算子。一个无参数检测器提取显著的季节性结构。编码器在带有尺度条件化的几何膨胀阶梯上应用一个共享局部模块,而解码器将上下文聚合状态与显式的季节性未来状态相结合,并在沿另一阶梯重用第二个模块后输出九个分位数。在六个语料库上进行预训练,在85,001个参数内保持了多领域广度。在97个GIFT-Eval配置上,无需针对每个数据集进行微调,Fracast-0是28个评估检查点中最小的,并且在综合参数-精度平面上保持非支配,MASE为0.808,WQL为0.564。它比TinyCast少用42.0%的参数,而TinyCast的MASE和WQL分别低4.2%和3.3%。这些结果支持跨尺度权重重用作为进一步压缩时间序列基础模型的实用途径。
cs.AI / 40 / 2609.32220

A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases

通过评分量规引导的策略归纳构建的双语AI听力师在模拟病例的盲法评估中优于人类听力师
Li, Linkai, Mo, Changgeng, Yu, Hanlin, Lu, Congxi, Wang, Shangqiguo, Fitzgerald, Matthew B, Wang, Shan X
Abstract
Audiology consultation requires structured history-taking, audiometric interpretation and patient-centred communication, yet real-world case material is scarce. We present a bilingual AI audiologist pairing a general-purpose large language model with rubric-guided playbook induction, multimodal audiogram interpretation and retrieval-augmented grounding, without fine-tuning the language-model backbone. Using a 21-item rubric and an AI patient simulator, we induced a 19-rule consultation policy from 73 training cases (43 English, 30 Chinese) and evaluated the system on 58 independent simulated cases (30 Chinese, 28 English) in a pre-specified, source-blinded comparison with 17 practising audiologists. The AI audiologist outperformed human audiologists on every case (58/58; mean paired $\Delta$ = +1.35 on a 5-point composite, Cohen's d = 1.84, $P = 4.5 \times 10^{-20}$), on 20 of 21 rubric items and in both languages. Component ablation identified the playbook as the largest contributor, offering a practical route to specialist consultation agents in low-data medical domains.
Chinese Translation
听力学咨询需要结构化的病史采集、听力图解读和以患者为中心的沟通,然而真实世界的病例材料稀缺。我们提出了一种双语AI听力师,它将通用大型语言模型与评分量规引导的策略归纳、多模态听力图解读和检索增强的接地相结合,而无需微调语言模型主干。使用一个21项的评分量规和一个AI患者模拟器,我们从73个训练病例(43个英语,30个中文)中归纳出一个19条规则的咨询策略,并在58个独立的模拟病例(30个中文,28个英语)上,通过预先设定的、来源盲法的比较,与17名执业听力师进行了评估。AI听力师在每个病例上都优于人类听力师(58/58;在5分综合评分上平均配对Δ = +1.35,Cohen's d = 1.84,P = 4.5 × 10⁻²⁰),在21个评分量规项目中的20个上以及两种语言中均如此。组件消融实验确定策略手册是最大的贡献者,为低数据医学领域的专科咨询代理提供了一条实用途径。
cs.AI / 41 / 2609.32224

RAO-Nav: Probing Omni-Language Models for Zero-shot Semantic Audio-Visual Navigation

RAO-Nav:探索全语言模型用于零样本语义音视频导航
Ye, Qilang, Liu, Meng, Zhou, Yu
Abstract
We explore whether Omni-Language Models (OLMs) can be directly applied to zero-shot Semantic Audio-Visual Navigation (SAVN). Recent work demonstrates that even state-of-the-art specialized models still struggle to achieve generalist multimodal navigation, despite extensive task-specific training. In this paper, we introduce RAO-Nav, short for Reasoning All-in-One OLM, a deployment pipeline for zero-shot SAVN. By leveraging the rich implicit audio-visual knowledge encoded in OLMs, the embodied agent is enabled to ``hear'', ``see'', ``reason'', and ``act'' in the environment. To further elicit the built-in thinking ability of OLMs, we propose a test-time Latent Navigation Reasoning (LNR) module that can be seamlessly integrated into the decoding space. LNR encourages the model to retrieve more target-relevant observations and make effective navigation decisions. Through comprehensive experiments, we show that our framework surpasses existing state-of-the-art baselines on public SAVN benchmarks without using any training data. Moreover, we introduce a new \emph{Global Navigation Instruction} setting to further evaluate the ability of OLMs to serve as embodied navigation agents. Code: https://github.com/rikeilong/OmniAV\_Nav.
Chinese Translation
我们探究全语言模型(Omni-Language Models, OLMs)能否直接应用于零样本语义音视频导航(SAVN)。近期工作表明,即使是最先进的专用模型,尽管进行了大量任务特定训练,仍难以实现通用多模态导航。本文提出 RAO-Nav,即 Reasoning All-in-One OLM 的缩写,一种用于零样本 SAVN 的部署流程。通过利用 OLMs 中编码的丰富隐式音视频知识,具身智能体能够在环境中“听”、“看”、“推理”和“行动”。为了进一步激发 OLMs 内置的思考能力,我们提出了一种可在解码空间中无缝集成的测试时潜在导航推理(Latent Navigation Reasoning, LNR)模块。LNR 鼓励模型检索更多与目标相关的观测,并做出有效的导航决策。通过全面实验,我们表明我们的框架在公共 SAVN 基准上超越了现有最先进基线,且未使用任何训练数据。此外,我们引入一种新的全局导航指令(Global Navigation Instruction)设置,以进一步评估 OLMs 作为具身导航智能体的能力。代码:https://github.com/rikeilong/OmniAV_Nav。
cs.AI / 42 / 2609.32247

Certifying Interventional Agreement Among Observationally Equivalent Causal Models

认证观测等价因果模型之间的干预一致性
Khanzadeh, Sourena, Platnick, Daniel, Alirezaie, Marjan, Rahnama, Hossein
Abstract
Observationally equivalent causal models can still disagree about what happens under intervention, because interventions create inputs that never occur in observational data. We introduce Interventional Separation Selection (ISS), which repeatedly queries the true system with an admissible intervention on which the surviving candidate models disagree, discards the candidates the outcome contradicts, and stops once no intervention within a cost bound separates the survivors. If the true system is among the candidates, this stopping condition certifies that every survivor agrees with it on every admissible intervention within the bound, a guarantee that no observational learner can give, however much data it sees. The stopping condition depends only on the survivors, so it can be checked without knowing the truth. For continuous variables the candidates form an infinite version space, and mixed-integer linear programs decide the stopping condition exactly over all of it, with agreement holding up to a tolerance. On a three-digit colored MNIST causal abstraction task in which ink hue tracks digit size, plain convolutional networks trained on examples reach zero held-out error, yet disagree with shape-based labels on 26% of single-digit edits, as often as hue-based labels do. Auditing the causal abstractions of networks observed only on such images, ISS certifies what each network perceives with 13.6 interventions per image on average, and each certificate, checked against every admissible intervention, holds whenever the network's true abstraction is among the candidates. When a network bypasses a unit that every candidate abstraction relies on, certificates covering interventions on that unit can be silently void, and twenty random validation interventions refute 69% of them.
Chinese Translation
观测等价的因果模型仍可能对干预下发生的情况存在分歧,因为干预会产生观测数据中从未出现的输入。我们提出干预分离选择(Interventional Separation Selection, ISS),它反复向真实系统查询一个可允许干预,该干预能使现存的候选模型产生分歧,丢弃结果与之矛盾的候选模型,并当成本界限内没有干预能区分存活模型时停止。如果真实系统在候选模型之中,则该停止条件证明每个存活模型都在界限内的每个可允许干预上与其一致,这是任何观测学习器无论看到多少数据都无法给出的保证。停止条件仅依赖于存活模型,因此无需知道真实系统即可检验。对于连续变量,候选模型构成无限版本空间,混合整数线性规划可以在整个空间上精确判定停止条件,一致性在容差范围内成立。在一个三位数字彩色 MNIST 因果抽象任务中,墨色色调与数字大小相关,在样例上训练的普通卷积网络达到零留出误差,却在 26% 的单数字编辑上与基于形状的标签不一致,其频率与基于色调的标签相同。审计仅在此类图像上观察到的网络的因果抽象,ISS 平均每张图像用 13.6 次干预来认证每个网络感知的内容,并且每个证书在对照每个可允许干预进行检验时,只要网络的真实抽象在候选之中,就成立。当一个网络绕过每个候选抽象都依赖的单元时,覆盖对该单元干预的证书可能会悄然无效,而二十次随机验证干预会反驳其中 69%。
cs.AI / 43 / 2609.32254

Why Directly Learning Periodic Trajectories Can Fail

为什么直接学习周期性轨迹可能会失败
Zheng, Kaixin, Layton, Anita
Abstract
Operator learning of periodic solutions requires deciding how simulation data should be recorded and represented. A natural choice is to integrate long enough for transients to decay and record a window wide enough to contain at least one full period of all trajectories. We find that these conservative choices can make the resulting trajectories difficult to learn, even when the underlying periodic orbits vary regularly with system parameters. Unaligned trajectories generalize poorly even within the training distribution. Phase alignment substantially improves in-distribution generalization, but models trained on a fixed physical-time window still have large errors on trajectories with periods outside the training range. We explain both failures through a common mechanism: frequency differences accumulate over time, so the target phase varies rapidly with the parameters. Predictors that cannot track this variation incur a population MSE floor in both settings; for fixed window prediction, we also derive a per-sample lower bound. We then study one of the simplest representations that escape these floors: learning an aligned, normalized waveform and its period separately. We establish regularity of the decoupled targets under ODE assumptions and show experimentally that this approach avoids both failures in ODE systems and a PDE case study.
Chinese Translation
周期解的算子学习需要决定如何记录和表示仿真数据。一种自然的选择是积分足够长的时间以使瞬态衰减,并记录足够宽的窗口,以包含所有轨迹的至少一个完整周期。我们发现,这些保守的选择会使所得轨迹难以学习,即使潜在的周期轨道随系统参数有规律地变化。未对齐的轨迹即使在训练分布内也泛化得很差。相位对齐显著提高了分布内泛化能力,但基于固定物理时间窗口训练的模型在周期超出训练范围的轨迹上仍然存在较大误差。我们通过一个共同的机制来解释这两种失败:频率差异随时间累积,因此目标相位随参数快速变化。无法跟踪这种变化的预测器在两种设置下都会产生总体MSE下界;对于固定窗口预测,我们还推导出了每个样本的下界。然后,我们研究了一种能够避开这些下界的最简单表示之一:分别学习对齐的归一化波形及其周期。我们在ODE假设下建立了解耦目标的正则性,并通过实验表明,该方法在ODE系统和PDE案例研究中都能避免这两种失败。
cs.AI / 44 / 2609.32255

Clarify the User or Verify the World? Uncertainty Routing for Proactive Agents

澄清用户还是验证世界?面向主动式代理的不确定性路由
Li, Zhaofeng, Zhang, Xuan, Xiao, Xiaokui, Deng, Yang
Abstract
Tool-using LLM agents must decide not only whether additional information is needed, but also which source can resolve the uncertainty. Existing proactive approaches often specialize in either user clarification or environment verification, without explicitly determining the appropriate information source for each decision. We formulate this problem as uncertainty routing among ACT, CLARIFY, and VERIFY, and propose PROUR, a proactive uncertainty routing framework. PROUR decomposes action uncertainty into disagreement across plausible user-goal interpretations, which signals user-side ambiguity, and the entropy remaining within each interpretation, which signals missing world-side evidence. To acquire information from the routed source, a query generator is trained with a mode-conditioned information-gain reward, targeting user-goal identification under CLARIFY and next-action identification under VERIFY. On $\tau$-bench, PROUR achieves 28.17% average success rate across retail and airline, outperforming the strongest prior method by 4.57% while using 2.17 fewer interaction steps. The learned policy further generalizes to stronger task agents and transactional domains of $\tau^3$-bench without retraining, demonstrating the benefit of source-aligned uncertainty resolution for proactive agents.
Chinese Translation
使用工具的LLM智能体必须不仅决定是否需要额外信息,还要决定哪个来源可以解决不确定性。现有的主动式方法往往专门针对用户澄清或环境验证,而没有明确确定每次决策的适当信息源。我们将此问题形式化为在ACT、CLARIFY和VERIFY之间的不确定性路由,并提出PROUR,一个主动式不确定性路由框架。PROUR将动作不确定性分解为:对可能的用户目标解释之间的分歧(这表示用户侧模糊性)以及每种解释内部剩余的熵(这表示缺失的世界侧证据)。为了从所路由的来源获取信息,查询生成器使用模式条件信息增益奖励进行训练,目标是在CLARIFY下识别用户目标,在VERIFY下识别下一动作。在$\tau$-bench上,PROUR在零售和航空领域实现了28.17%的平均成功率,优于最强先前方法4.57%,同时交互步骤减少了2.17步。学习到的策略进一步泛化到更强的任务智能体和$\tau^3$-bench的事务性领域,无需重新训练,证明了面向主动式代理的源对齐不确定性解决的好处。
cs.AI / 45 / 2609.32256

LAM: Efficient Lossy Agent Memory Framework With A Retrieval-Score Error Bound

LAM:具有检索分数误差界的高效有损智能体记忆框架
Sun, Baixi, Chen, Le, Chowdhury, Anjir Ahmed, Ma, Xiaolong, Yang, Chih-Hsuan, Xia, Mingze, Zawad, Syed, Di, Sheng, Kettimuthu, Rajkumar, Zheng, Huihuo, Thakur, Rajeev, Vishwanath, Venkatram, Yan, Feng
Abstract
Agent memory grows as agents read inputs, reason, and call tools. Longer histories increase inference cost and eventually exceed the context window. LLM-based summarization reduces this history but adds latency and provides no explicit bound on information loss. We propose LAM, a Lossy Agent Memory system with three components: a deterministic deduplication rule with a substitution bound on retrieval scores - a bound on score perturbation, not a certificate of unchanged ranking; a memory manager that preserves the cached prefix and overlaps compaction with inference; and a performance model that estimates compaction costs before deployment. On 600 agent trajectories, LAM removes 22.47% of observation tokens while retaining 99.984% of the measured gold-patch evidence. At a fixed deletion set, the performance model predicts a 71.4x-91.6x end-to-end speedup from removing records before prefill instead of deleting them from a prefilled context. That benefit comes from the schedule rather than the rule and applies to any prefix-preserving test.
Chinese Translation
智能体记忆随着智能体读取输入、推理和调用工具而增长。更长的历史会增加推理成本,并最终超出上下文窗口。基于LLM的摘要减少了这种历史,但增加了延迟,并且没有提供对信息损失的显式界。我们提出LAM,一种有损智能体记忆系统,包含三个组件:一个确定性去重规则,带有检索分数的替换界——是对分数扰动的界,而不是排名不变的保证;一个记忆管理器,保留缓存前缀,并将压缩与推理重叠;以及一个性能模型,在部署前估计压缩成本。在600条智能体轨迹上,LAM移除了22.47%的观测token,同时保留了99.984%的测得金标补丁证据。在固定删除集下,性能模型预测,在预填充之前移除记录(而不是从已预填充的上下文中删除它们)可带来71.4倍至91.6倍的端到端加速。这种收益来自调度而非规则,并适用于任何保持前缀的测试。
cs.AI / 46 / 2609.32259

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

面向异构多智能体LLM的免预填充跨模型族KV缓存传输
Yun, Vincent-Daniel, Lim, Woosang, Yoo, Haneul, Yoo, Sungjoo, Karimireddy, Sai Praneeth, Annavaram, Murali
Abstract
Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundancy, but prefill-free transfer across model families must handle differences in tokenization, model depth, and KV representations. To address these issues, we propose \textit{HeteroFold}, a prefill-free cross-family KV cache transfer method that keeps both the sender and receiver frozen. HeteroFold aligns model structures, maps the sender cache into the receiver space, and calibrates it to preserve receiver behavior. Across six transfer directions, HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings. It also matches text-based communication on the multi-agent benchmark. At 32K context length, Llama-3.1-8B$\rightarrow$Ministral-3-14B transfer is $10.7\times$ faster than Native Prefill and $1.18$--$1.47\times$ faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge. These results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill.
Chinese Translation
近年来,多智能体LLM系统越来越多地结合异构模型以承担专门的智能体角色。然而,基于文本的通信要求每个接收方对发送方已处理过的共享上下文进行预填充。复用发送方的键值(KV)缓存可避免这种冗余,但跨模型族的免预填充传输必须处理分词、模型深度和KV表示方面的差异。为解决这些问题,我们提出 HeteroFold,一种免预填充的跨模型族KV缓存传输方法,保持发送方和接收方均冻结。HeteroFold 对齐模型结构,将发送方缓存映射到接收方空间,并对其进行校准以保持接收方行为。在六个传输方向上,HeteroFold 在所有四个长上下文基准和大多数短上下文设置上取得了最佳缓存传输性能。它在多智能体基准上也达到了与基于文本通信相当的表现。在32K上下文长度下,Llama-3.1-8B→Ministral-3-14B 传输比 Native Prefill 快 10.7×,比最先进的免预填充基线 Dense Latent 和 KV Ridge 快 1.18–1.47×。这些结果表明,HeteroFold 能够在无需接收方预填充的情况下实现高效的跨模型族 KV 复用。
cs.AI / 47 / 2609.32274

When Does a Skill Add Value? Task-Conditional Gain Prediction for Selective Skill Use

技能何时增加价值?面向选择性技能使用的任务条件增益预测
Xu, Anjie, Zhang, Zhiyu, Ding, Ruiqing, Xu, Fengli, Wang, Leye
Abstract
Agent skills are expected to improve task performance. Yet we find that they often provide no benefit, and can even hurt performance while incurring additional token costs. Can we predict whether a skill will help before the agent acts? We introduce SkillDelta, a framework for predicting task-conditional skill gains from paired executions of the same agent with and without the skill. A local predictor transfers these historical gains to new tasks without retraining the agent. Under explicit transfer assumptions, our analysis links support coverage, representation mismatch, and execution noise to prediction error and decision regret. Across five benchmarks and three target agents, paired history improves observed-gain ranking over skill-assisted outcomes alone in 12 of 15 settings. At matched expected skill-use rates, SkillDelta improves success over random activation in all 15 settings, with an average absolute gain of 4.3%. Most of this advantage comes from allocation across task groups. Evidence for additional within-group selection value is strongest on ToolQA and weaker elsewhere. Code is available at https://github.com/TankTechnology/skilldelta.
Chinese Translation
智能体技能有望提升任务性能。然而我们发现,它们通常不提供任何益处,甚至可能损害性能,同时产生额外的 token 成本。我们能否在智能体行动之前预测技能是否有帮助?我们提出 SkillDelta,一个从同一智能体在有技能和无技能下的配对执行中预测任务条件技能增益的框架。局部预测器将这些历史增益迁移到新任务,而无需重新训练智能体。在明确的迁移假设下,我们的分析将支持覆盖、表示不匹配和执行噪声与预测误差和决策遗憾联系起来。在五个基准测试和三个目标智能体上,配对历史在 15 个设置中的 12 个中,相比仅技能辅助结果,提升了观察增益的排序。在匹配的预期技能使用率下,SkillDelta 在所有 15 个设置中相比随机激活提高了成功率,平均绝对增益为 4.3%。这种优势主要来自任务组间的分配。额外组内选择价值的证据在 ToolQA 上最强,在其他地方较弱。代码可在 https://github.com/TankTechnology/skilldelta 获取。
cs.AI / 48 / 2609.32295

GLIDE: Generalized Layer-wise Intrinsic Distributional Evaluation for Heterogeneous LLM Agents

GLIDE:面向异构 LLM 智能体的广义逐层内在分布评估
Zhu, Wei, Wang, Yiming, Wang, Rui, Yu, Lixing, Yue, Kun, Tang, Zhiwen
Abstract
LLM agents require reliable step-level evaluation to compare candidate branches and allocate computation effectively. However, lightweight evaluation remains challenging. External verifiers introduce additional inference cost, while agent-produced confidence or self-evaluation scores can be miscalibrated, especially when candidates are generated by heterogeneous agents. We propose \textbf{G}eneralized \textbf{L}ayer-wise \textbf{I}ntrinsic \textbf{D}istributional \textbf{E}valuation (\textbf{GLIDE}) for LLM agents. \textsc{GLIDE} derives intrinsic step evidence from layer-wise residual coherence, which measures whether local residual updates consistently support the global residual change induced by a candidate step. It calibrates this evidence against the recent score distribution of the generating agent and converts it into a pessimistic reward that jointly accounts for absolute residual evidence and agent-relative standing. The reward provides a cross-agent value signal for MCTS branch selection, while normalized predictive uncertainty guides adaptive branching. Experiments on multi-hop reasoning, sequential decision making, and symbolic logic show that \textsc{GLIDE} improves task performance, step-level ranking quality, and computational efficiency without external verifiers or task-specific supervision.
Chinese Translation
LLM 智能体需要可靠的步骤级评估,以比较候选分支并有效分配计算资源。然而,轻量级评估仍然具有挑战性。外部验证器会引入额外的推理成本,而智能体产生的置信度或自评估分数可能校准不当,尤其是当候选由异构智能体生成时。我们提出了用于 LLM 智能体的广义逐层内在分布评估(GLIDE)。GLIDE 从逐层残差一致性中导出内在步骤证据,该一致性衡量局部残差更新是否持续支持由候选步骤引起的全局残差变化。它根据生成智能体最近的分数分布校准该证据,并将其转换为悲观奖励,该奖励同时考虑绝对残差证据和智能体相对位置。该奖励为 MCTS 分支选择提供跨智能体价值信号,而归一化预测不确定性指导自适应分支。在多跳推理、序列决策和符号逻辑上的实验表明,GLIDE 无需外部验证器或任务特定监督即可提高任务性能、步骤级排序质量和计算效率。
cs.AI / 49 / 2609.32297

Agentsensus: Consensus-Compressed Shared Memory for Multi-Agent Story Worlds

Agentsensus:面向多智能体故事世界的共识压缩共享记忆
Pan, Yu
Abstract
A agentic story world is a dynamic system simulating who learned what, when, and from whom -- yet the standard design gives each character a private memory stream. A shared event is therefore stored once per witness, large duplication will be incurred in terms of storage. We present Agentsensus, a story-world simulation framework in which there is an unified long-term memory. Records of the same event merge into one owned by all its witnesses, and semantically relevant memory records are linked. We evaluate on four worlds -- two classical Chinese novels, Hamlet, and a real-world conflict timeline -- run for 40 to 80 rounds against three per-character memory designs under an equal-granularity protocol. Agentsensus writes 22-44% fewer entries than the closest baseline and is the only design whose memory becomes shared (14-28% of records held by more than one character, some by 10) and linked (94-99%), at judged simulation quality indistinguishable or even better than the baselines. An ablation attributes this to the merge itself: disabling it multiplies the store by 3.1x and takes sharing to exactly zero. Sharing also compounds with the horizon rather than saturating early, rising 6% to 9% to 14% as one world is re-run at 10, 20 and 40 rounds.
Chinese Translation
一个智能体故事世界是一个动态系统,模拟谁在何时从谁那里学到了什么——然而标准设计为每个角色提供了一个私有记忆流。因此,一个共享事件会为每个目击者存储一次,在存储方面会产生大量重复。我们提出了Agentsensus,一个故事世界模拟框架,其中有一个统一的长期记忆。同一事件的记录合并为一个,由其所有目击者共同拥有,并且语义相关的记忆记录被链接起来。我们在四个世界上进行评估——两部中国古典小说、《哈姆雷特》和一个真实世界冲突时间线——在等粒度协议下,针对三种每角色记忆设计,运行40到80轮。Agentsensus比最接近的基线少写入22-44%的条目,并且是唯一一种记忆变为共享(14-28%的记录由多个角色持有,有些由10个角色持有)和链接(94-99%)的设计,在评判的模拟质量上与基线无法区分甚至更好。一项消融研究将此归因于合并本身:禁用它会使存储量增加3.1倍,并将共享率降至恰好为零。共享也随着时间范围而累积,而不是早期饱和,当同一个世界在10、20和40轮重新运行时,共享率从6%上升到9%再到14%。
cs.AI / 50 / 2609.32303

Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging

Train4Merge: 基于OPD的模型合并中RL与SFT教师的受控单教师研究
Huang, Jingyuan, Huang, Zuming, Shi, Yucheng, Li, Zhongzhi, Zhai, Xiaoming, Chu, Wei, Liu, Ninghao
Abstract
Domain experts trained from a shared checkpoint can be merged into one model through on-policy distillation (OPD), where they act as teachers supervising a student on its own trajectories. One upstream choice is rarely examined: whether to build each expert with supervised fine-tuning (SFT) or reinforcement learning (RL). Yet equally strong teachers need not be equally good teachers. We probe this choice through controlled single-teacher OPD, a building block of multi-teacher OPD: in Agentic, Reasoning, and Perception, comparably performing SFT and RL teachers are trained from Qwen3.5-9B, each guiding a student initialized from it. At their best checkpoints, RL-guided students outperform SFT-guided students by 4.27, 1.50, and 0.86 percentage points in Agentic, Reasoning, and Perception, respectively, and recover more of their teachers' performance gains over the base model. The contrast is clearest in Agentic, where the best SFT-guided student recovers only 44.44% of its teacher's gain, whereas the best RL-guided student recovers 115.00%, surpassing its teacher. Our analysis points to an explanation: RL teachers stay much closer to the shared initialization in parameter space than SFT teachers and are therefore easier for their students to follow.
Chinese Translation
从共享检查点训练的领域专家可以通过同策略蒸馏(OPD)合并为一个模型,其中他们作为教师,在学生模型自身的轨迹上对其进行监督。一个上游选择很少被研究:是用监督微调(SFT)还是强化学习(RL)来构建每个专家。然而,同样强大的教师不一定是同样优秀的教师。我们通过受控单教师OPD来探究这一选择,这是多教师OPD的一个构建模块:在Agentic、Reasoning和Perception中,从Qwen3.5-9B训练出性能相当的SFT和RL教师,每个教师指导一个从它初始化的学生模型。在它们的最佳检查点上,RL指导的学生模型在Agentic、Reasoning和Perception上分别比SFT指导的学生模型高出4.27、1.50和0.86个百分点,并且相对于基础模型,更多地恢复了其教师的性能提升。这种对比在Agentic中最为明显,最佳SFT指导的学生模型仅恢复了其教师性能提升的44.44%,而最佳RL指导的学生模型恢复了115.00%,超过了其教师。我们的分析指出了一种解释:RL教师在参数空间中比SFT教师更接近共享初始化,因此更容易被其学生模型跟随。
cs.AI / 51 / 2609.32309

PhiFold: Towards Dynamic Protein Design with Physics-Structured Covariance Modeling

PhiFold:迈向物理结构化协方差建模的动态蛋白质设计
Liu, Yutian, Lin, Mujie, LanqianZhang, Fan, Meng, Liu, Chang, ZhiweiNie, Ma, Siwei
Abstract
Protein design is moving beyond structural correctness toward function-aware design, yet existing generative models typically treat dynamics as a downstream property estimated through simulation or prediction after structure generation. Using MD trajectories as a generative target is also undesirable because stochastic, path-dependent trajectories over-specify the underlying equilibrium ensemble. We introduce PhiFold, a framework for jointly generating protein backbones and their second-order dynamics, represented by residue-displacement covariance. Rather than predicting the quadratically sized full covariance, PhiFold decomposes dynamics into three interpretable components: local flexibility, a low-rank collective-motion representation, and residue-wise collective participation. These components are assembled into a positive-definite covariance matrix with exact marginal consistency, yielding a compact and physically constrained representation of equilibrium dynamics. Across generated proteins, PhiFold improves recovery of local fluctuations and long-range residue coupling while remaining competitive on dominant collective-motion subspaces. It further enables bidirectional control of residue flexibility while preserving backbone designability. By unifying structure generation with an explicit representation of equilibrium dynamics, PhiFold lays a foundation for designing proteins not only by how they look, but also by how they move.
Chinese Translation
蛋白质设计正超越结构正确性,迈向功能感知的设计,然而现有的生成模型通常将动力学视为一种下游属性,在结构生成之后通过模拟或预测来估计。使用MD轨迹作为生成目标也不理想,因为随机的、路径依赖的轨迹过度指定了潜在的平衡系综。我们提出PhiFold,一个联合生成蛋白质骨架及其二阶动力学的框架,以残基位移协方差表示。PhiFold不是预测二次方大小的完整协方差,而是将动力学分解为三个可解释的成分:局部柔性、低秩集体运动表示和残基级集体参与。这些成分被组装成一个具有精确边缘一致性的正定协方差矩阵,产生平衡动力学的紧凑且物理约束的表示。在生成的蛋白质中,PhiFold改善了局部波动和长程残基耦合的恢复,同时在主导集体运动子空间上保持竞争力。它进一步实现了对残基柔性的双向控制,同时保持骨架可设计性。通过将结构生成与平衡动力学的显式表示统一起来,PhiFold为不仅根据蛋白质的外观、还根据其运动方式来设计蛋白质奠定了基础。
cs.AI / 52 / 2609.32312

Delayed Supervision for Test-Time Language Models

测试时语言模型的延迟监督
Kim, Jinha, Kothari, Taksh
Abstract
Test-time language models adapt a compact memory while processing the input sequence. This perspective encompasses nonlinear fast-weight learning in LaCT, associative delta-rule updates in DeltaNet, and generalized delta-rule state updates in RWKV-7. Training these models to predict the next token does not explicitly require a fact to remain accessible after many subsequent memory updates. We study delayed supervision for this test-time memory: during post-training, ask a simulator-grounded question only after a long interval of unrelated events, and supervise its answer alongside ordinary next-token prediction. Questions are evaluated on disposable branches, so their answers never enter the continuing event stream. The construction distinguishes retention from revision: a retained fact must remain valid throughout the delay, whereas a revised fact must be answered with its latest value. We evaluate this approach on LaCT-760M and plain DeltaNet-1.3B using TextWorld training trajectories and shared BABILong and RULER evaluation panels, and include a separately reported RWKV-7 comparison. Relative to event-only training, delayed QA improves BABILong by 5.48 percentage points for LaCT and 1.32 points for DeltaNet, and single-needle RULER by 1.45 and 3.27 points, respectively. The RWKV-7 comparison reports gains of 4.60 and 7.00 points on its own panels. These results support delayed semantic supervision as a practical outer training objective for usable test-time memory, while leaving open how much of the benefit derives specifically from delay rather than general question-answering and answer-termination supervision.
Chinese Translation
测试时语言模型在处理输入序列时调整一个紧凑的记忆。这一视角涵盖了LaCT中的非线性快速权重学习、DeltaNet中的关联delta规则更新以及RWKV-7中的广义delta规则状态更新。训练这些模型预测下一个词元并未明确要求一个事实在多次后续记忆更新后仍可访问。我们研究针对这种测试时记忆的延迟监督:在后训练期间,仅在经过一长段无关事件后提出一个基于模拟器的问题,并将其答案与普通的下一词元预测一起进行监督。问题在一次性分支上进行评估,因此其答案永远不会进入持续的事件流。该构建区分了保持与修订:保持的事实必须在整个延迟期间保持有效,而修订的事实必须以其最新值来回答。我们在LaCT-760M和普通DeltaNet-1.3B上评估了这种方法,使用TextWorld训练轨迹以及共享的BABILong和RULER评估面板,并包括单独报告的RWKV-7比较。相对于仅事件训练,延迟QA将LaCT的BABILong提高了5.48个百分点,DeltaNet提高了1.32个百分点,单针RULER分别提高了1.45和3.27个百分点。RWKV-7比较在其自身的面板上报告了4.60和7.00个百分点的增益。这些结果支持将延迟语义监督作为可用的测试时记忆的实用外部训练目标,同时保留了以下问题:多少收益具体来自延迟,而不是一般的问答和答案终止监督。
cs.AI / 53 / 2609.32326

RLHarness: Co-evolving Procedural Skills with Reinforcement Learning for Long-horizon Multimodal Reasoning

RLHarness:协同演化程序性技能与强化学习用于长时程多模态推理
Shang, Ziqiao, Xu, Zian, Yan, Ji-Chen, Wu, Weiming, Jia, Ziyi, Meng, Jie, Huang, Tao, Huang, Shan, Guo, Lan-Zhe
Abstract
Multimodal reasoning requires models to preserve visual evidence through long decision chains while selecting appropriate procedures across diverse scenarios and rules. When learning is guided only by terminal verifiers, reinforcement learning (RL) reveals whether a final answer is correct but not how it should be produced. The policy must therefore discover reusable reasoning procedures while learning to execute them, creating a program cold-start problem. Skills can externalize successful procedures, reduce repeated exploration, and provide inspectable guidance. However, a fixed Skill Bank assumes that this guidance remains compatible with an evolving policy, while updating Skills alone can leave their triggers, execution protocols, and demonstrations stale or mutually inconsistent. We introduce RLHARNESS, which organizes Skills, selection and execution protocols, few-shot demonstrations, and task contracts into a unified, versioned Harness and alternates Harness evolution with policy learning. An Exploration-Distillation Harness builds the initial Harness and version-aligned verified traces for SFT and DAPO I. After the first RL block, a Post-RL Reconstruction Harness rebuilds Skills, protocols, and demonstrations from fresh success-failure rollouts, and DAPO II adapts the policy to the reconstructed program. RLHARNESS improves Accuracy from 16.25%/27.50% to 62.00%/50.00% on MetroMap/TravelMap and raises F1 score from 37.13%/45.50% to 65.81%/65.51% on Fee-VL/Cancel-VL. All four tasks achieve their best results only after reconstruction and DAPO II, showing that an evolving Harness complements RL by continually updating the external program that the policy learns to execute.
Chinese Translation
多模态推理要求模型在长决策链中保持视觉证据,同时在不同场景和规则下选择适当的程序。当仅由终端验证器指导学习时,强化学习(RL)能揭示最终答案是否正确,但不能揭示应如何产生该答案。因此,策略必须在学习执行可复用推理程序的同时发现这些程序,从而产生程序冷启动问题。技能可以将成功的程序外化,减少重复探索,并提供可检查的指导。然而,固定的技能库(Skill Bank)假设这种指导与演化的策略保持兼容,而仅更新技能可能使其触发条件、执行协议和演示变得过时或相互不一致。我们提出 RLHARNESS,它将技能、选择和执行的协议、少样本演示以及任务契约组织成一个统一、版本化的 Harness,并交替进行 Harness 演化和策略学习。探索-蒸馏 Harness 构建初始 Harness 和用于 SFT 与 DAPO I 的版本对齐的验证轨迹。在第一个 RL 块之后,RL 后重建 Harness 根据新的成功-失败轨迹重建技能、协议和演示,DAPO II 使策略适应重建后的程序。RLHARNESS 在 MetroMap/TravelMap 上将准确率从 16.25%/27.50% 提升到 62.00%/50.00%,并在 Fee-VL/Cancel-VL 上将 F1 分数从 37.13%/45.50% 提升到 65.81%/65.51%。所有四个任务只有在重建和 DAPO II 之后才达到最佳结果,表明演化的 Harness 通过不断更新策略学习执行的外部程序来补充 RL。
cs.AI / 54 / 2609.32327

HyperReCo: Retrieving and Connecting Evidence with Hypergraph Neural Networks for LLM Multi-hop Reasoning

HyperReCo:利用超图神经网络为LLM多跳推理检索与连接证据
Zhao, Zicheng, Luo, Linhao, Dong, Junnan, Luo, Haoran, Li, Xiaoli, Pan, Shirui, Gong, Chen
Abstract
Large language models (LLMs) have shown strong capabilities, with retrieval-augmented generation (RAG) supporting complex multi-hop reasoning by retrieving evidence distributed across documents. Graph-based approaches exploit connections among evidence, and hypergraph-based retrieval further preserves higher-order entity associations within documents and connects documents through shared entities. However, existing hypergraph retrievers often rely on predefined structural expansion or diffusion, which may miss query-dependent interactions needed to identify relevant evidence. They also leave connections among retrieved evidence implicit, requiring LLMs to reconstruct these connections before reasoning. Therefore, we propose HyperReCo, a framework for retrieving and connecting evidence with a hypergraph neural network (HyperGNN). We represent each document as a hyperedge over its extracted entities, with shared entities connecting the hyperedges. Through hypergraph message passing with joint supervision over documents and entities, the HyperGNN learns query-dependent interactions to retrieve complementary evidence. We further introduce Gradient-Guided Hyper-Path Decoding (GGHD), which uses gradient attribution to interpret the learned interactions and translate them into explicit hyper-paths that help LLMs combine complementary facts for multi-hop reasoning. Experiments on six benchmarks show that HyperReCo achieves the best retrieval performance among the compared methods on all three multi-hop QA datasets, together with strong downstream QA performance. Case studies and further analyses demonstrate the utility of decoded hyper-paths for connecting retrieved evidence.
Chinese Translation
大型语言模型(LLMs)已展现出强大能力,而检索增强生成(RAG)通过检索分布在多篇文档中的证据来支持复杂的多跳推理。基于图的方法利用证据之间的连接,而基于超图的检索进一步保留了文档内的高阶实体关联,并通过共享实体连接文档。然而,现有超图检索器通常依赖预定义的结构扩展或扩散,这可能遗漏识别相关证据所需的查询相关交互。它们还使检索到的证据之间的连接保持隐式,要求LLMs在推理前重建这些连接。因此,我们提出HyperReCo,一个利用超图神经网络(HyperGNN)检索并连接证据的框架。我们将每篇文档表示为其抽取实体上的一个超边,并通过共享实体连接这些超边。通过对文档和实体进行联合监督的超图消息传递,HyperGNN学习查询相关交互以检索互补证据。我们进一步引入梯度引导超路径解码(Gradient-Guided Hyper-Path Decoding, GGHD),它利用梯度归因来解释学习到的交互,并将其转化为显式超路径,帮助LLMs组合互补事实进行多跳推理。在六个基准上的实验表明,HyperReCo在所有三个多跳问答数据集上均取得了对比方法中最佳的检索性能,并具有很强的下游问答性能。案例研究和进一步分析证明了解码所得超路径在连接检索证据方面的效用。
cs.AI / 55 / 2609.32339

Enabling Timely Guidance before Skill Retrieval: Retaining Helpful Warm Tips in Agent Context

在技能检索前实现及时指导:在智能体上下文中保留有用的热提示
Liang, Feng, Li, Yupeng, Zeng, Runhao, Lau, Francis C. M., Hu, Xiping
Abstract
Reusable skills help LLM-based agents solve complex tasks, but the agent must receive guidance before it commits to an ineffective approach. Existing skill mechanisms often expose only metadata and load full content on demand, leaving useful guidance unavailable until the agent decides to retrieve it. General memory methods can incur substantial maintenance overhead, while keeping guidance in conversation context risks repeatedly exposing the agent to irrelevant or harmful advice. We propose TipsWarm, a mechanism that complements existing skill mechanisms by maintaining a budgeted pool of skill-derived keypoints, or \textit{warm tips}, for selective injection into the context of every message turn. By separating event-triggered LLM assessment from inexpensive per-turn screening, it makes transferable skill guidance readily available while controlling maintenance costs. In three coding and iterative task-execution benchmarks, TipsWarm achieves the highest task success rate while remaining time-efficient, compared to recent skill and general memory baselines.
Chinese Translation
可复用的技能有助于基于LLM的智能体解决复杂任务,但智能体必须在采取无效方法之前获得指导。现有的技能机制通常仅暴露元数据并按需加载完整内容,导致有用的指导在智能体决定检索之前不可用。通用记忆方法可能带来大量的维护开销,而将指导保留在对话上下文中则存在使智能体反复接触无关或有害建议的风险。我们提出TipsWarm,一种补充现有技能机制的机制,它维护一个预算化的技能衍生关键点池(即热提示(warm tips)),用于选择性地注入每条消息轮次的上下文中。通过将事件触发的LLM评估与低成本的逐轮筛选分离,它使可迁移的技能指导随时可用,同时控制维护成本。在三个编码和迭代任务执行基准测试中,与最近的技能和通用记忆基线相比,TipsWarm实现了最高的任务成功率,同时保持时间高效。
cs.AI / 56 / 2609.32344

ALLOT: Budgeted Hybrid-Memory Routing for Knowledge Updates in LLMs

ALLOT:面向大语言模型知识更新的预算化混合记忆路由
Huang, Shanfeng, Fang, Zhou, Xiao, Song, Du, Hai
Abstract
For large language models (LLMs), parametric adaptation is costly when retrieval already suffices. We introduce ALLOT, a hybrid-memory routing framework that separates learned write priority from a hard parametric budget. A memory-aware router combines frozen text representations, retrieval confidence, and relation metadata; a single ranking supports multiple write budgets while preserving all facts in external memory. On CounterFact with Qwen3-4B, ALLOT reaches 0.760 accuracy at a 20% parametric-write budget and recovers 78.4% of the budget-matched oracle gain, with 80% fewer parametric writes than dual-writing every fact. At this budget, jointly adding retrieval and relation features to text improves normalized oracle gain by 6.2 percentage points. Complementary Qwen3-0.6B shared-store results achieve dual-write-level accuracy with 6-14.5% parametric writes, and cross-benchmark transfer retains approximately 88% of in-domain gain. These results support allocating adaptation capacity according to its incremental value rather than treating every factual update as an equally valuable training target.
Chinese Translation
对于大语言模型(LLM),当检索已经足够时,参数化适配代价高昂。我们提出 ALLOT,一个混合记忆路由框架,它将学习到的写入优先级与硬性参数预算分离。一个记忆感知路由器结合了冻结的文本表示、检索置信度和关系元数据;单个排序支持多种写入预算,同时将所有事实保留在外部记忆中。在 CounterFact 数据集上,使用 Qwen3-4B,ALLOT 在 20% 参数写入预算下达到 0.760 准确率,并恢复了预算匹配的 oracle 增益的 78.4%,与对每个事实进行双重写入相比,参数写入减少了 80%。在此预算下,将检索和关系特征联合添加到文本中,使归一化 oracle 增益提高了 6.2 个百分点。互补的 Qwen3-0.6B 共享存储结果在 6-14.5% 的参数写入下达到了双重写入级别的准确率,并且跨基准迁移保留了约 88% 的域内增益。这些结果支持根据增量价值分配适配容量,而不是将每个事实更新都视为同等价值的训练目标。
cs.AI / 57 / 2609.32351

From Trajectories to Grounded Preferences: Process Preference Synthesis via Interaction Element Graphs for Web PRMs

从轨迹到基于环境的偏好:通过交互元素图为 Web PRMs 合成过程偏好
Peng, Yangzhe, Wang, Xiaoyang, Zhao, Yiyang, Wu, Lijun, He, Kun
Abstract
Comparative Process Reward Models (PRMs) provide critical step-level guidance for autonomous web agents by evaluating state-conditioned preferences between candidate actions. However, existing preference training data synthesized via multi-policy sampling suffers from a severe scarcity of Grounded Minimal Contrastive Pairs (GMCPs)-where competing candidates target genuine on-page elements with identical action types. In representative baselines preference data (namely, WebArbiter), GMCPs account for merely 24.19%, biasing PRMs during training to rely on shallow shortcuts (such as element hallucinations and action type mismatches) rather than acquiring genuine contextual decision semantics. To address these challenges, we propose SURFPRM, a graph-guided process preference synthesis framework for comparative Web PRMs. SURFPRM structures web demonstrations into a persistent Interaction Element Graph that acts as an environment-grounded negative action proposal mechanism, systematically synthesizing contrastive negative actions across spatial, temporal, and spatiotemporal confusion axes. This elevates the GMCP proportion from 24.19% to 74.60%, producing the curated SURFPRM-DATA dataset. Across six open-source backbones (3B to 9B parameters), PRMs trained on SURFPRM-DATA outperform baseline-trained models on average on WEBPRMBENCH and rival leading proprietary LLMs. In downstream reward-guided trajectory search on WEBARENA-LITE, SURFPRM provides step-level guidance for both GPT-4o (+14.21%) and GPT-4o-mini (+12.83%) policies, yielding substantial improvements in complex web task success rates.
Chinese Translation
比较式过程奖励模型(PRMs)通过评估候选动作之间以状态为条件的偏好,为自主 Web 智能体提供关键的步骤级指导。然而,通过多策略采样合成的现有偏好训练数据严重缺乏基于环境的最小对比对(GMCPs)——其中竞争的候选动作针对具有相同动作类型的真实页面元素。在代表性基线偏好数据(即 WebArbiter)中,GMCPs 仅占 24.19%,导致 PRMs 在训练期间偏向于依赖浅层捷径(例如元素幻觉和动作类型不匹配),而非获得真正的上下文决策语义。为了解决这些挑战,我们提出了 SURFPRM,一种用于比较式 Web PRMs 的图引导过程偏好合成框架。SURFPRM 将 Web 演示构建为持久的交互元素图,该图充当基于环境的负动作提议机制,系统地沿空间、时间和时空混淆轴合成对比负动作。这将 GMCP 比例从 24.19% 提升至 74.60%,生成了精心整理的 SURFPRM-DATA 数据集。在六个开源主干模型(3B 至 9B 参数)上,使用 SURFPRM-DATA 训练的 PRMs 在 WEBPRMBENCH 上平均优于基线训练的模型,并可与领先的专有 LLM 相媲美。在 WEBARENA-LITE 上的下游奖励引导轨迹搜索中,SURFPRM 为 GPT-4o(+14.21%)和 GPT-4o-mini(+12.83%)策略提供步骤级指导,显著提高了复杂 Web 任务的成功率。
cs.AI / 58 / 2609.32378

AuthorityLens: Rethinking LLM-Based Agent Systems Through the Lens of Authority

AuthorityLens:从权威视角重新思考基于LLM的智能体系统
Chen, Shaojin, Jing, Huihao, Chan, Wun Yu, Hu, Wenbin, Li, Jiaxing, Ching, Wu Pandy Pui, Bhatia, Kshitij, He, Xinlei, Li, Haoran, Song, Yangqiu
Abstract
LLM-based agents are increasingly deployed with authority over consequential resources and decisions in real systems. These agents often operate alongside human and LLM-based participants who hold different forms of authority. Yet workflow roles, permission settings, and review mechanisms do not necessarily reflect the authority realized in practice. We introduce AuthorityLens, a framework for measuring a system's authority structure. Starting from an authority portfolio, we evaluate a system along three dimensions: what the system is authorized to do (System Authority), how much joint participation is required to exercise that authority (Authority Separation), and how much authority each participant holds (Principal Authority). We derive these measurements from the minimal combinations of participants sufficient to realize each outcome across admissible runtime states. We apply AuthorityLens to Codex, OpenCode, and Gemini CLI across 13 operating configurations over a common portfolio of agent operations. We find that nominal configurations do not map cleanly onto realized authority. In Codex, Full Access changes System Authority only marginally while substantially concentrating authority in the executing Assistant. OpenCode's Build and Plan configurations have the same System Authority and Authority Separation despite different workflows and root-level permissions. In Gemini CLI, model-based review increases Authority Separation without changing System Authority. Principal Authority further distinguishes authority replication from authority separation: spawned or delegated agents can become alternative holders of the same authority without increasing the required joint participation. Together, these results demonstrate that AuthorityLens provides a unified framework for measuring and comparing realized authority structures across agent systems.
Chinese Translation
基于LLM的智能体正越来越多地被部署,在真实系统中对重要资源和决策拥有权威。这些智能体通常与人类以及基于LLM的参与者一起运作,这些参与者持有不同形式的权威。然而,工作流角色、权限设置和审查机制并不一定反映实践中实现的权威。我们引入AuthorityLens,一个用于测量系统权威结构的框架。从一个权威组合出发,我们沿着三个维度评估系统:系统被授权做什么(系统权威),行使该权威需要多少联合参与(权威分离),以及每个参与者拥有多少权威(主体权威)。我们从在可允许的运行时状态中足以实现每个结果的最小参与者组合中推导出这些测量值。我们将AuthorityLens应用于Codex、OpenCode和Gemini CLI,在13种操作配置下,针对一组常见的智能体操作组合。我们发现,名义配置并不能清晰地映射到实际实现的权威。在Codex中,Full Access仅略微改变系统权威,同时将权威大幅集中到执行的Assistant中。OpenCode的Build和Plan配置具有相同的系统权威和权威分离,尽管工作流和根级权限不同。在Gemini CLI中,基于模型的审查增加了权威分离,但没有改变系统权威。主体权威进一步区分了权威复制和权威分离:生成或委派的智能体可以成为同一权威的替代持有者,而无需增加所需的联合参与。总之,这些结果表明,AuthorityLens提供了一个统一的框架,用于测量和比较跨智能体系统实际实现的权威结构。
cs.AI / 59 / 2609.32390

Reward Hacking and Agent Containment Failure: A Monte Carlo Study Based on the 2026 Hugging Face Incident

奖励黑客与智能体遏制失败:基于2026年Hugging Face事件的蒙特卡洛研究
Ozer, Murat, Erenay, Bulent, Berber, Ibrahim
Abstract
The July 2026 intrusion into Hugging Face production infrastructure showed how reward hacking can become an external cybersecurity incident when a capable agent encounters weak containment boundaries. This study develops a probabilistic risk model linking five stages: reward hacking, containment escape, usable access, persistence, and failure of detection. A Monte Carlo simulation evaluates 100,000 runs under each of four control configurations. Input distributions represent explicit uncertainty and are used for comparative analysis rather than real-world frequency prediction. Under the stated assumptions, layered controls reduce simulated external-incident probability substantially more than network isolation or monitoring used alone, an ordering that holds under independent plus/minus 25% perturbation of every coefficient in the model across 300 draws. Sensitivity analysis shows that agent capability and weaknesses in monitoring, authorization, and credential control exert the greatest influence on modeled risk. Human temporal discounting and metric gaming provide a behavioral analogy for short-horizon optimization, but the study does not infer that AI agents experience gratification or human motivation. The results support treating cyber-capable agent evaluations as hostile security zones in which indirect egress, shared infrastructure, credentials, and evaluation artifacts must remain outside the agent's effective authority.
Chinese Translation
2026年7月对Hugging Face生产基础设施的入侵表明,当一个能力强大的智能体遇到薄弱的遏制边界时,奖励黑客如何可能演变为外部网络安全事件。本研究建立了一个概率风险模型,将五个阶段联系起来:奖励黑客、遏制逃逸、可用访问、持久化以及检测失败。蒙特卡洛模拟在四种控制配置下各评估了100,000次运行。输入分布代表了显式不确定性,用于比较分析,而非真实世界频率预测。在所述假设下,分层控制比单独使用网络隔离或监控更能大幅降低模拟的外部事件概率,这一排序在模型每个系数独立进行正负25%扰动、重复300次抽取的情况下依然成立。敏感性分析表明,智能体能力以及监控、授权和凭证控制中的弱点对建模风险的影响最大。人类时间贴现和指标博弈为短视界优化提供了一种行为类比,但本研究并不推断AI智能体会体验到满足感或人类动机。结果支持将具有网络能力的智能体评估视为敌意安全区,其中间接出口、共享基础设施、凭证和评估产物必须保持在智能体的有效权限之外。
cs.AI / 60 / 2609.32391

SCLATE: a Substrate for Continual-Learning Agent Training and Evaluation

SCLATE:一个用于持续学习智能体训练与评估的基底
Jung, Youngmok, Salekin, Sirajul, Tran, Henry, Movellan, Javier, Huang, Zhao, Bilkhu, Manjot
Abstract
Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. Yet existing benchmarks and training frameworks schedule only the benchmark's own events, leaving each benchmark and agent pair to build a custom scheduling loop. We present SCLATE, an execution substrate where benchmarks and unmodified agents each add their events to one open event scheduler through an adapter. A hybrid simulated clock runs these events on a shared timeline, flowing in real time while the agent works and skipping idle gaps, which compresses a month-long scenario into hours. SCLATE also serves as a rollout engine that runs any agent's harness and memory unmodified, recording the tokens and log probabilities of every model call through an in-container proxy. We port seven benchmarks to SCLATE and compare ten unmodified harness and memory configurations head to head on ten models. The comparison shows that an added memory system does not reliably beat the harness's native memory and that models differ widely in how they use the same harness and memory. We then post-train Qwen3.5-4B through unmodified harnesses and memory systems. The model learns to use both, reading 6.8x fewer file lines with a 16.7-point higher SWE-bench Verified pass rate, and writing richer memory records, while its held-out MetaClaw accuracy rises by up to 11.8 points.
Chinese Translation
持续学习智能体是由模型、运行框架和记忆组成的系统,在长时间跨多会话的范围内运行。评估和训练它们需要将任务与智能体侧的事件(如会话停止和启动、定时任务(crons)和记忆巩固)交错进行。然而,现有的基准测试和训练框架仅调度基准测试自身的事件,导致每个基准测试和智能体对都需要构建自定义的调度循环。我们提出了 SCLATE,一个执行基底,其中基准测试和未经修改的智能体各自通过适配器将事件添加到同一个开放事件调度器中。一种混合模拟时钟在共享时间线上运行这些事件,在智能体工作时实时流动,并跳过空闲间隙,从而将长达一个月的场景压缩到几小时内。SCLATE 还可作为 rollout 引擎,无需修改即可运行任何智能体的运行框架和记忆系统,并通过容器内代理记录每次模型调用的词元和对数概率。我们将七个基准测试移植到 SCLATE,并在十个模型上对十种未经修改的运行框架和记忆配置进行头对头比较。比较表明,添加的记忆系统并不能可靠地胜过运行框架的原生记忆,并且不同模型在使用相同运行框架和记忆系统时差异很大。然后,我们通过未经修改的运行框架和记忆系统对 Qwen3.5-4B 进行后训练。该模型学会了同时使用两者,读取的文件行数减少了 6.8 倍,SWE-bench Verified 通过率提高了 16.7 个百分点,并写入了更丰富的记忆记录,同时其留出的 MetaClaw 准确率提升了高达 11.8 个百分点。
cs.AI / 61 / 2609.32394

Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization

超越脚本化搜索:通过智能体黑箱优化实现样本高效的奖励发现
Li, Minghao, Tan, Rui, Wang, Ruihang
Abstract
Designing dense reward functions for low-level reinforcement learning (RL) control remains difficult. Recent work uses large language models (LLMs) to iteratively generate and refine reward functions using policy-training feedback within scripted search algorithms. However, evaluating each candidate requires a full RL training run, making sample efficiency a central challenge for reward search on complex control tasks. To address this limitation, we propose an Agentic Reward Black-box Optimization (ARBO) framework, in which an LLM agent builds the search strategy at run time from an evaluation history maintained as its persistent workspace. The evaluation history comprises two components: observations maintained by the evaluation oracle, including candidate scores, per-term training curves, and error tracebacks; and an agent-maintained belief that records diagnoses and intended next steps. The agent queries both with tools and generates the next batch of reward candidates, rather than generating them in a single pass from a fixed prompt. Across four control domains, ARBO achieves gains of 29.9% in manipulation success rate and 192.8% in power-grid score over baseline means under a shared evaluation budget. Ablations examine each component's contribution and sensitivity to backbone choice.
Chinese Translation
为低层强化学习(RL)控制设计稠密奖励函数仍然困难。近期工作利用大语言模型(LLM)在脚本化搜索算法中,通过策略训练反馈迭代生成和优化奖励函数。然而,评估每个候选奖励函数都需要完整的 RL 训练运行,这使得样本效率成为复杂控制任务中奖励搜索的核心挑战。为解决这一局限,我们提出了智能体奖励黑箱优化(ARBO)框架,其中 LLM 智能体在运行时根据作为其持久工作区维护的评估历史来构建搜索策略。评估历史包含两个部分:由评估预言机维护的观测,包括候选分数、逐项训练曲线和错误回溯;以及由智能体维护的信念,记录诊断和预期的下一步。智能体使用工具查询这两者,并生成下一批奖励候选,而不是从固定提示中一次性生成它们。在四个控制领域,在共享评估预算下,ARBO 在操作成功率上相比基线均值提升了 29.9%,在电网得分上提升了 192.8%。消融实验考察了每个组件的贡献以及对骨干选择的敏感性。
cs.AI / 62 / 2609.32395

Memory as a cache: Exact context reuse and deletion by construction

记忆作为缓存:通过构造实现精确的上下文重用与删除
Wang, Shengyao, Liu, Jiang
Abstract
The KV cache of a transformer entangles every token's representation with its entire prefix: a passage encoded once cannot be reused under a different prefix or removed without recomputing everything after it, so exact cache reuse is limited to shared prefixes. We present SMem, an architecture whose context representation is a cache by construction. A block-local encoder maps each block to memory rows independently of other blocks, and a reader conditions generation on their union through cross-attention. For every parameter setting, memory composes exactly at fixed block indices, deleting a block is an exact $O(b)$ update for $b$-token blocks, and the memory state is independent of the edit path. At $4\times$ the training context, under the shared recipe, SMem retrieves planted needles beyond any trained-length window (exact match 0.14-0.28 at distances of 31 and 63 blocks), where learned-position, RoPE, and Block-Attention-style transformers all score at most 0.02. A fully cached context is served by computing one block alone at a near-constant 3.1-6.2 ms, whereas cold prefill grows with context; batched decode stores 34-38% fewer KV rows and runs 1.4-1.7$\times$ faster when bandwidth-bound; and deletion beats suffix recomputation by 8.5$\times$ at 512 blocks and 452$\times$ at 4096 blocks (32-256$\times$ the trained length, probing the cost model rather than a served regime). The cost is a perplexity gap of -4.7% to +2.8% (negative favors SMem) against a parameter-matched transformer with the same positional scheme, at 160M-1.5B on FineWeb-Edu across two recipes and a learning-rate search. SMem also composes with RoPE: at 160M and 410M the composite matches or leads the matched transformer and closes 29-59% of SMem's gap to a RoPE transformer. Dropping prefix entanglement thus keeps perplexity comparable while making the cache exactly composable and editable.
Chinese Translation
Transformer 的 KV 缓存将每个 token 的表示与其整个前缀纠缠在一起:一段文本一旦编码,就不能在不同前缀下重用,也不能在不重新计算其之后所有内容的情况下删除,因此精确的缓存重用仅限于共享前缀。我们提出 SMem,一种架构,其上下文表示在构造上就是缓存。一个块局部编码器将每个块独立于其他块映射到记忆行,读取器通过交叉注意力以其并集为条件进行生成。对于每个参数设置,记忆在固定块索引处精确组合,删除一个块对于 b 个 token 的块是精确的 O(b) 更新,并且记忆状态独立于编辑路径。在训练上下文的 4 倍长度下,在共享配方下,SMem 能够检索超出任何训练长度窗口的植入针(在距离 31 和 63 个块时精确匹配 0.14-0.28),而学习位置、RoPE 和 Block-Attention 风格的 transformer 最多得分 0.02。完全缓存的上下文通过单独计算一个块来服务,耗时近乎恒定在 3.1-6.2 毫秒,而冷预填充随上下文增长;批量解码存储的 KV 行减少 34-38%,在带宽受限时运行速度提高 1.4-1.7 倍;并且删除在 512 个块时比后缀重新计算快 8.5 倍,在 4096 个块时快 452 倍(训练长度的 32-256 倍,探测的是成本模型而非服务机制)。代价是与参数匹配且具有相同位置方案的 transformer 相比,困惑度差距为 -4.7% 到 +2.8%(负值有利于 SMem),在 160M-1.5B 参数规模下,在 FineWeb-Edu 上,跨两种配方和学习率搜索。SMem 还可以与 RoPE 组合:在 160M 和 410M 下,组合模型匹配或领先于匹配的 transformer,并缩小了 SMem 与 RoPE transformer 之间 29-59% 的差距。因此,放弃前缀纠缠可以在保持困惑度相当的同时,使缓存完全可组合和可编辑。
cs.AI / 63 / 2609.32398

Function Over Form: Distributional Orthogonalization in Mixture-of-Experts with Replica Expert Mechanism

功能重于形式:带有副本专家机制 (Replica Expert Mechanism) 的混合专家模型中的分布正交化
He, Jinfan, Liu, Yunzhuo, Zhang, Kai, Han, Weidong, Key, Rayying
Abstract
The scaling of LLMs increasingly relies on MoE architectures to decouple active computation from total parameter count. However, the efficacy of MoE is often constrained by expert collapse and representation redundancy, both leading to underutilization of model capacity. To address these challenges, this paper proposes Distributional Orthogonalization Loss (DO-loss), an auxiliary regularization that shifts the focus from static weight diversity to dynamic routing behavior. By representing each expert's token assignment history as a high-dimensional binary load signature, DO-loss penalizes signature overlap to prevent expert collapse while encouraging functional specialization. To align this algorithmic design with system efficiency, we further introduce the Replica Expert Mechanism (REM), which improves load balancing through a two-tiered strategy: adjusting replica expert placement at the global-batch level and performing real-time token dispatching at the micro-batch level. Empirical evaluations demonstrate that our method outperforms the evaluated routing algorithms on downstream tasks for both 4.8BA0.5B and 30BA3B MoE models, while maintaining comparable training efficiency.
Chinese Translation
大语言模型 (LLMs) 的规模扩展越来越依赖于混合专家 (MoE) 架构,以将激活计算量与总参数量解耦。然而,MoE 的有效性常常受到专家坍缩和表示冗余的制约,两者都会导致模型容量利用不足。为了解决这些挑战,本文提出了分布正交化损失 (DO-loss),一种辅助正则化方法,将关注点从静态权重多样性转向动态路由行为。通过将每个专家的 token 分配历史表示为高维二进制负载签名,DO-loss 惩罚签名重叠以防止专家坍缩,同时鼓励功能专业化。为了将这种算法设计与系统效率对齐,我们进一步引入了副本专家机制 (REM),它通过两层策略改善负载均衡:在全局批次级别调整副本专家放置,在微批次级别执行实时 token 调度。实证评估表明,我们的方法在 4.8BA0.5B 和 30BA3B MoE 模型的下游任务上优于所评估的路由算法,同时保持相当的训练效率。
cs.AI / 64 / 2609.32407

Opening LLM Judges: Recovering Preference Signals Beyond the Final Verdict

打开LLM评判器:恢复超越最终裁决的偏好信号
Mukherjee, Sourabrata, Sitaram, Sunayana
Abstract
LLM judges are widely used to evaluate model outputs, but their verdicts can be unreliable: a judge may favor the worse answer for its position, length, or other surface features. When a judge is wrong, is the information needed to judge correctly absent from the model, or present in its internal representations but not reflected in the output? We study this across 64 open-weight evaluators and 14 datasets, including causal interventions on 41 judges (editing activations mid-run to see whether the verdict changes). On LLMBar, built so the superficially better answer is the worse one, the verdicts of 50 judges agree with human labels only 0.456 of the time, even after averaging both answer orders. Yet a small probe on the same judges' activations, with no weight updates, reaches 0.846, and 0.686 once surface features such as length and position are residualized out (0.507 with shuffled labels). The gap holds across eight benchmarks and model families, but is not universal: a score of how well surface features alone predict the human label, computed before any probe is trained, predicts the size of the gain (Spearman rho = 0.90). On rubric tasks that score one answer at a time, leaving no surface cue to exploit, reading the internals gives no advantage. The interventions also show that editing activations mid-network already changes the verdict, before it can be read off directly, and locate the pathways carrying position and length bias. At the same label budget, the recovered signal lets a judge flag cases where it is likely wrong and yields better labels for preference learning. A wrong verdict, then, does not mean the judge lacks the information, and a simple diagnostic shows when it is worth recovering.
Chinese Translation
LLM评判器被广泛用于评估模型输出,但其裁决可能不可靠:评判器可能因其位置、长度或其他表面特征而偏向较差的答案。当评判器出错时,正确评判所需的信息是模型中没有,还是存在于其内部表示中但未反映在输出中?我们在64个开放权重评估器和14个数据集上研究了这一问题,包括对41个评判器进行因果干预(在运行中编辑激活以观察裁决是否改变)。在LLMBar上,该数据集构建为表面更好的答案实际上是更差的答案,50个评判器的裁决与人类标签仅0.456的时间一致,即使在对两种答案顺序取平均后也是如此。然而,对同一评判器激活的简单探针,无需权重更新,达到了0.846,一旦将长度和位置等表面特征残差化后为0.686(打乱标签后为0.507)。这一差距在八个基准和模型家族中持续存在,但并非普遍:在任何探针训练之前计算的仅表面特征预测人类标签的得分,可以预测增益的大小(Spearman rho = 0.90)。在每次只对单个答案评分的评分标准任务中,没有可利用的表面线索,读取内部信息没有优势。干预还表明,在网络中间编辑激活就已经改变了裁决,在它能够被直接读取之前,并定位了携带位置和长度偏差的路径。在相同的标签预算下,恢复的信号让评判器能够标记它可能出错的情况,并为偏好学习产生更好的标签。因此,错误的裁决并不意味着评判器缺乏信息,一个简单的诊断显示了何时值得恢复。
cs.AI / 65 / 2609.32420

Carnator: Fast Text-to-Video Generation with Generation-Native Compatibility-Guided Cross-Request Reuse

Carnator:基于生成原生兼容性引导的跨请求复用实现快速文本到视频生成
Yin, Xingkun, Tang, Xuebin, Xu, Mingkun, Du, Hongyang
Abstract
Video diffusion transformers produce high-quality videos, yet iterative denoising incurs substantial inference latency, limiting interactive and large-scale serving. Most existing acceleration methods focus on individual requests, thereby restricting efficiency gains to redundancy within a single generation trajectory. Recent cross-request reuse offers an additional source of savings, but existing approaches often infer reusability from coarse semantic similarity. This conflates semantic relatedness with generation-level computational compatibility, so aggressive reuse may accept incompatible historical computation while conservative reuse leaves substantial acceleration unrealized. We present \emph{Carnator}, a cross-request acceleration framework that addresses this challenge by extracting and using generation-native compatibility evidence directly from the model's evolving internal states. Specifically, \emph{Carnator} performs a lightweight early probe to construct an Early Signature from internal diffusion states, assessing reuse validity through risk-aware compatibility decisions. The same evidence characterizes reuse scope by localizing target-specific computation and guiding joint reuse of historical latent trajectories and sparse attention connectivity. Across three text-to-video backbones, Carnator consistently achieves higher cache-hit end-to-end acceleration than the evaluated cross-request baselines despite more selective cache acceptance, reaching up to 2.17$\times$ speedup while maintaining competitive generation quality.
Chinese Translation
视频扩散变换器能够生成高质量视频,然而迭代去噪会带来大量的推理延迟,限制了交互式和大规模服务。大多数现有的加速方法关注单个请求,从而将效率提升限制在单条生成轨迹内的冗余上。近期的跨请求复用提供了额外的节省来源,但现有方法通常从粗粒度的语义相似性推断可复用性。这将语义相关性与生成级别的计算兼容性混为一谈,因此激进的复用可能接受不兼容的历史计算,而保守的复用则无法实现显著的加速。我们提出 Carnator,一个跨请求加速框架,通过直接从模型不断演化的内部状态中提取和使用生成原生的兼容性证据来应对这一挑战。具体而言,Carnator 执行轻量级的早期探测,从内部扩散状态构建一个早期签名(Early Signature),通过风险感知的兼容性决策来评估复用有效性。同样的证据通过定位目标特定的计算并引导历史潜在轨迹和稀疏注意力连接的联合复用,来刻画复用范围。在三个文本到视频骨干网络上,Carnator 始终比所评估的跨请求基线实现了更高的缓存命中端到端加速,尽管其缓存接受更具选择性,最高可达 2.17 倍加速,同时保持有竞争力的生成质量。
cs.AI / 66 / 2609.32422

MergeHEIR: Mitigating Multimodal Hallucinations as the Tax of Model Merging

MergeHEIR:缓解作为模型合并代价的多模态幻觉
Li, Jinyu, Fang, Hao, Zhang, Zhiming, Kong, Jiawei, Chen, Bin, Xia, Shu-Tao
Abstract
Model merging consolidates task-specialized experts into a single deployable model. However, we show that such capability consolidation incurs a merging tax of increased hallucination: across 8 model-merging methods, every merged checkpoint exhibits a higher hallucination rate than the average of its constituent experts. An intuitive approach is to adapt existing hallucination-mitigation methods to the post-merge model, yet this unconstrained adaptation disrupts inherited capabilities, creating a tension between hallucination mitigation and expertise retention. To tackle this challenge, we introduce MergeHEIR, a post-merge adaptation framework designed to reduce this merging tax while preserving expertise inherited from initial experts. Using small expert-task calibration sets, MergeHEIR constructs layer-wise null-space projectors via SVD from task-specific activations collected from the merged checkpoint, and periodically projects the accumulated post-merge displacement onto the resulting null spaces to preserve inherited expertise. Theoretically, we establish minimum-distortion and maximum-dimensionality guarantees, characterize the threshold-controlled adaptation-retention trade-off, and extend perturbation guarantees beyond finite calibration data. Across 24 paired comparisons spanning three MLLM configurations and 8 model-merging methods, MergeHEIR consistently mitigates hallucination while largely preserving inherited expertise, demonstrating a more favorable hallucination-retention trade-off.
Chinese Translation
模型合并将任务专用的专家整合为一个可部署的单一模型。然而,我们表明这种能力整合会带来合并代价,即幻觉增加:在8种模型合并方法中,每个合并后的检查点都表现出比其组成专家平均更高的幻觉率。一种直观的方法是将现有的幻觉缓解方法应用于合并后模型,然而这种无约束的适应会破坏继承的能力,造成幻觉缓解与专业能力保留之间的矛盾。为了应对这一挑战,我们提出了MergeHEIR,一种合并后适应框架,旨在减少这种合并代价,同时保留从初始专家继承的专业能力。使用小型专家-任务校准集,MergeHEIR通过SVD从合并检查点收集的任务特定激活中构建逐层零空间投影器,并定期将累积的合并后位移投影到所得零空间上,以保留继承的专业能力。理论上,我们建立了最小失真和最大维度保证,刻画了阈值控制的适应-保留权衡,并将扰动保证扩展到有限校准数据之外。在涵盖三种MLLM配置和8种模型合并方法的24对比较中,MergeHEIR一致地缓解幻觉,同时很大程度上保留继承的专业能力,展示了更有利的幻觉-保留权衡。
cs.AI / 67 / 2609.32423

PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins

PluginRSI:利用可复用插件对智能体运行框架进行递归改进
Shi, Yaorui, Miao, Yuchun, Chen, Yuxin, Zhang, Jiayuan, Sun, Yueqing, Song, Xierui, Wang, Xiang, Zhang, An
Abstract
The harness surrounding a language model is a central determinant of agent performance. Recent methods optimize harnesses by searching over complete programs, where individual mechanisms are difficult to isolate and reuse. We introduce PluginRSI, which represents a harness as a composition of atomized plugins and organizes harness evolution around these plugins. Individual plugins are improved independently and accumulated in a shared library, then recombined into new harnesses at each iteration. PluginRSI improves over existing harness optimization methods across software engineering, command-line interaction, and question-answering tasks. The resulting harnesses retain their advantage when transferred to other solver models without further optimization. The evolved plugin library accelerates subsequent optimization from the initial harness, which helps faster and higher convergence on unseen tasks. These results show that accumulating reusable mechanisms provides an effective basis for continued harness improvement.
Chinese Translation
围绕语言模型的运行框架是智能体性能的核心决定因素。近期方法通过在完整程序空间中进行搜索来优化运行框架,但其中各个机制难以被单独分离和复用。我们提出 PluginRSI,它将运行框架表示为原子化插件的组合,并围绕这些插件组织运行框架的演化。各个插件被独立改进并积累到一个共享库中,然后在每次迭代时重新组合成新的运行框架。在软件工程、命令行交互和问答任务上,PluginRSI 优于现有的运行框架优化方法。所得运行框架在无需进一步优化的情况下迁移到其他求解器模型时仍能保持其优势。演化后的插件库加速了从初始运行框架出发的后续优化,从而有助于在未见任务上更快、更高水平地收敛。这些结果表明,积累可复用机制为持续改进运行框架提供了有效基础。
cs.AI / 68 / 2609.32428

Authorization Closure Graph: Minimal Repair for LLM Agents with Evolving User Instructions

授权闭包图(Authorization Closure Graph):面向具有演化用户指令的LLM智能体的最小修复
Wang, Qingzhuo, Wang, CaiYi, Meng, Jinglu, Qin, Ruiyang, Peng, Kunyu, Wei, Zhihua, Shen, Wen
Abstract
Tool-using large language model (LLM) agents increasingly perform state-changing actions that require user authorization. Yet existing approaches do not provide a principled mechanism for selectively updating prior authorization when only part of an instruction changes. To this end, we propose an Authorization-Closure-Graph (ACG)-based framework that represents authorization and its dependencies as an evolving, versioned state. ACG selectively invalidates authority affected by a revision while preserving unaffected portions of the authorization state, and computes a minimal repair that identifies only the missing evidence or authority required for execution. This enables agents to adapt to revised instructions while avoiding stale authority and unnecessary authorization requests. We evaluate ACG across three advanced LLMs in two natural tasks, and ACG consistently improves action safety rate and task success rate. Code is available at https://github.com/weiliang822/ACG.
Chinese Translation
使用工具的大型语言模型(LLM)智能体越来越多地执行需要用户授权的状态改变动作。然而,现有方法并未提供一种原则性机制,用于在仅部分指令发生变化时选择性地更新先前的授权。为此,我们提出了一种基于授权闭包图(ACG)的框架,将授权及其依赖关系表示为演化的、带版本的状态。ACG会选择性地使受修订影响的权限失效,同时保留授权状态中未受影响的部分,并计算最小修复,仅识别执行所需的缺失证据或权限。这使智能体能够适应修订后的指令,同时避免过期权限和不必要的授权请求。我们在两个自然任务中跨三个先进LLM评估了ACG,ACG持续提升了动作安全率和任务成功率。代码见 https://github.com/weiliang822/ACG。
cs.AI / 69 / 2609.32429

PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers

PrismQuant:分组量化器的最优零空间旋转
Chen, Yanlong, Chen, Yining, Zhang, Song, Habibian, Amirhossein, Li, Yawei
Abstract
Smaller activation outliers do not necessarily imply better low-bit quantization: their alignment with the quantizer matters. We introduce PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the constant group subspace of asymmetric grouped INT4. The affine offsets represent the energy in this subspace without widening the range within the group. We formulate rotation design as a Ky Fan trace maximization and derive a closed-form solution that is provably optimal for this alignment objective. Compact Householder transformations and their compact-WY representation enable gradient-free construction and efficient application at both foldable and online sites. A predictive range law further connects unaligned activation energy and group size to quantization-relevant variation. Experiments on Llama, Qwen, and Mistral span dense models up to 70B parameters and a 30B mixture-of-experts model. Under W4A4KV4, PrismQuant sets the state of the art on Llama-3.2-3B among the compared methods in both perplexity and accuracy. On Llama-3.1-70B, it attains 3.85 perplexity and 72.46% average zero-shot accuracy, only 0.22 percentage points below full precision. In the deployment study on Llama-3.1-8B, our optimized implementation achieves 1.51x prefill and 1.22x CUDA Graph decode speedups over matched FP16 baselines, with 56.34% lower decode peak memory and only 2.35% additional Graph decode latency over Hadamard. Code is available at https://github.com/ForeverBlue816/PrismQuant.
Chinese Translation
较小的激活异常值并不一定意味着更好的低位量化:它们与量化器的对齐很重要。我们提出PrismQuant,一个量化器感知的旋转框架,它将主要的激活特征空间与非对称分组INT4的恒定组子空间对齐。仿射偏移表示该子空间中的能量,而不会扩大组内的范围。我们将旋转设计表述为Ky Fan迹最大化,并推导出一个闭式解,该解对于该对齐目标可证明是最优的。紧凑的Householder变换及其紧凑WY表示使得能够在可折叠和在线位置进行无梯度构建和高效应用。一个预测性范围定律进一步将未对齐的激活能量和组大小与量化相关的变化联系起来。在Llama、Qwen和Mistral上的实验涵盖了高达70B参数的密集模型和一个30B的专家混合模型。在W4A4KV4下,PrismQuant在Llama-3.2-3B上在困惑度和准确度方面在比较方法中达到了最先进的水平。在Llama-3.1-70B上,它达到了3.85的困惑度和72.46%的平均零样本准确率,仅比全精度低0.22个百分点。在Llama-3.1-8B的部署研究中,我们的优化实现比匹配的FP16基线实现了1.51倍的预填充和1.22倍的CUDA Graph解码加速,解码峰值内存降低了56.34%,且仅比Hadamard增加2.35%的Graph解码延迟。代码可在https://github.com/ForeverBlue816/PrismQuant获取。
cs.AI / 70 / 2609.32430

Multi-Agent System Search via Active Substructure-aware Policy Optimization

基于主动子结构感知策略优化的多智能体系统搜索
Xu, Beicheng, Fan, Bowen, Qian, Weitong, Tung, Lingching, Cui, Bin
Abstract
LLMs enable multi-agent systems (MAS) to tackle complex tasks, but manually designing agent roles, prompts, and communication structures requires substantial expertise and effort. This motivates learning policies that construct query-specific MAS from execution reward. Existing approaches typically train these policies by repeatedly traversing a fixed set of training queries and assigning rewards at the workflow level. However, this overlooks differences in queries' evolving learning potential and obscures which substructures improve solution quality. In this paper, we propose Active Substructure-aware Policy Optimization (ASPO), a RL framework for query-level MAS search. ASPO introduces an Adaptive Query-Selection Mechanism (AQSM) that focuses training on queries at the policy's competence boundary: those it can solve but not yet reliably. A complementary discovery mechanism widens architectural exploration for hard queries, helping distinguish insufficient exploration from operator capability limits. Beyond query selection, ASPO introduces substructure-level rewards that measure output-quality gains within each action's descendant subgraph. These rewards guide proximal policy optimization to reinforce useful architectural refinements and discourage redundant or harmful computation. Together, these mechanisms prioritize learnable queries and provide fine-grained feedback for learning effective MASs. Across six benchmarks spanning mathematical reasoning, general question answering, and code generation, ASPO ranks first on every benchmark against twelve baselines.
Chinese Translation
大语言模型(LLMs)使多智能体系统(MAS)能够处理复杂任务,但手动设计智能体角色、提示和通信结构需要大量专业知识和努力。这促使人们学习从执行奖励中构建查询特定MAS的策略。现有方法通常通过反复遍历一组固定的训练查询并在工作流级别分配奖励来训练这些策略。然而,这忽略了查询不断变化的学习潜力的差异,并模糊了哪些子结构提高了解决方案质量。在本文中,我们提出了主动子结构感知策略优化(ASPO),一种用于查询级MAS搜索的强化学习(RL)框架。ASPO引入了一种自适应查询选择机制(AQSM),将训练集中在策略能力边界处的查询上:那些它能够解决但尚不能可靠解决的查询。一个互补的发现机制拓宽了针对困难查询的架构探索,有助于区分探索不足与操作符能力限制。除了查询选择之外,ASPO引入了子结构级奖励,用于衡量每个动作的后代子图中的输出质量增益。这些奖励引导近端策略优化(proximal policy optimization)以强化有用的架构改进,并抑制冗余或有害的计算。这些机制共同优先考虑可学习的查询,并为学习有效的MAS提供细粒度反馈。在涵盖数学推理、通用问答和代码生成的六个基准测试中,ASPO在所有基准测试中均排名第一,击败了十二个基线。
cs.AI / 71 / 2609.32434

From Latents to Wires: Surgical Post-Editing on Large Language Models

从潜在表示到线路:大型语言模型上的外科手术式后编辑
Jin, Jiankai, Zhang, Xiangzheng, Liu, Zhao, Xu, Wenzhuo, Yang, Dongdong, Zhang, Deyue, Zou, Quanchen
Abstract
Given a large language model (LLM), can whoever holds the weights name a semantic target (e.g., the model's identity), locate the model components that produce it, and edit them so that the target no longer appears while other capability is preserved? We call such an edit on a trained model a post-edit. We present L2W (latents to wires), a framework that performs surgical post-edits for named semantic targets. For localization, L2W uses Jacobian lens (J-lens) attribution to score components against the semantic target. For surgical removal, because LLM mechanisms are redundant (i.e., a semantic target may have multiple components producing it), L2W runs Counterexample-Guided Causal Cut (CGCC) until the target no longer appears. CGCC first cumulatively closes model components, treating each surviving expression of the target as a counterexample that exposes the next components to close, and then reopens some of them to preserve capability. In a controlled experiment with an implanted behavioural watermark, L2W removes the watermark, and its localization lands on the model region the implant changed. Across three model configurations, L2W removes model-metadata (e.g., identity) self-claims in all nine runs, and adult-content refusal in all three, with no held-out target residual. L2W further composes two post-edits on a text-to-image model: one removes the refusal of requested nudity, and a second removes the nude rendering the first exposes. The results support post-editing as a complement to post-training: post-training installs preferred behaviours, and post-editing removes named unwanted ones.
Chinese Translation
给定一个大型语言模型(LLM),持有权重的人能否指定一个语义目标(例如,模型的身份),定位产生该目标的模型组件,并编辑它们,使得该目标不再出现,同时其他能力得以保留?我们将这种对训练后模型的编辑称为后编辑。我们提出 L2W(latents to wires,从潜在表示到线路),一个为指定语义目标执行外科手术式后编辑的框架。对于定位,L2W 使用雅可比透镜(J-lens)归因来针对语义目标对组件进行评分。对于外科手术式移除,由于 LLM 机制具有冗余性(即一个语义目标可能由多个组件产生),L2W 运行反例引导的因果切割(CGCC),直到目标不再出现。CGCC 首先累积地关闭模型组件,将目标的每一个仍存活的表达视为一个反例,暴露出下一个需要关闭的组件,然后重新打开其中一些组件以保留能力。在一个植入行为水印的受控实验中,L2W 移除了水印,并且其定位落在植入所改变的模型区域上。在三种模型配置中,L2W 在所有九次运行中移除了模型元数据(例如身份)的自我声明,并在所有三次运行中移除了成人内容拒绝,且没有目标残留。L2W 进一步在一个文本到图像模型上组合了两个后编辑:一个移除了对请求裸体的拒绝,第二个移除了第一个所暴露出的裸体渲染。结果支持后编辑作为后训练的补充:后训练安装偏好行为,后编辑移除指定的不想要的行为。
cs.AI / 72 / 2609.32436

Controllable GNN Explanations via Multi-Metric Preference Selection

通过多指标偏好选择实现可控的GNN解释
Verma, Rachit, Deshmukh, Yashraj J., Dasgupta, Anirban
Abstract
Mechanisms for generating GNN explanations are crucial for building trust and mitigating biases in Graph Neural Networks (GNNs), especially in high-stakes scenarios. Most current methods optimize only for fidelity under the sparsity constraint. However, this discounts the need for interpretable explanations (those that consist of familiar motif patterns) and stable explanations (those that remain unchanged under structural perturbations). We propose a novel approach that optimizes GNN explanations across these metrics, exposing their relative weighing as a control. Experiments on various real-world datasets, including MUTAG, BA-2Motif, BAMultiShapes, and PROTEINS, suggest that our method produces higher-fidelity explanations than a state-of-the-art baseline on MUTAG and PROTEINS across all evaluated budgets, and on BA-2Motif at larger budgets, while being faster in the regime of small explanation budgets. We also explore how, given an input motif library containing standard motifs for the corresponding domain, the method can be used to determine the relative importance of those motifs in generating the explanations, and how this information can be used to further improve the quality of the output explanations. We also examine the relationship between different metrics through their induced tradeoff surface, and explore its dependence on the nature of the motif library.
Chinese Translation
生成GNN解释的机制对于建立信任和减轻图神经网络(GNN)中的偏差至关重要,尤其是在高风险场景中。当前大多数方法仅在稀疏性约束下优化保真度。然而,这忽视了可解释解释(由熟悉的模体模式组成)和稳定解释(在结构扰动下保持不变)的需求。我们提出了一种新颖的方法,在这些指标上优化GNN解释,将它们的相对权重作为一种控制暴露出来。在包括MUTAG、BA-2Motif、BAMultiShapes和PROTEINS在内的各种真实数据集上的实验表明,我们的方法在MUTAG和PROTEINS上在所有评估预算下,以及在BA-2Motif上在较大预算下,比最先进的基线产生更高保真度的解释,同时在小型解释预算的情况下更快。我们还探讨了,在给定包含相应领域标准模体的输入模体库的情况下,该方法如何用于确定这些模体在生成解释中的相对重要性,以及如何利用这些信息进一步提高输出解释的质量。我们还通过它们诱导的权衡曲面来检查不同指标之间的关系,并探索其对模体库性质的依赖性。
cs.AI / 73 / 2609.32448

ForkLeft: Entropy-First Rollouts for Prefix-Aligned Autoregressive-to-Diffusion Distillation

ForkLeft:用于前缀对齐的自回归到扩散蒸馏的熵优先展开
Liu, Junming, Wang, Jicheng, He, Yifeng, Chen, Hao, Qi, Jianzhong
Abstract
Autoregressive Next-Token Prediction (NTP) has enabled strong reasoning capabilities in language models, while Diffusion Language Models (DLMs) offer flexible token orders and parallel generation. We ask whether DLMs can acquire NTP-style reasoning through distillation without giving up their native generation process. Direct distillation, however, faces a fundamental mismatch: an autoregressive teacher predicts from a left prefix, whereas a DLM can condition on tokens on both sides. We introduce ForkLeft, a distillation framework that resolves this mismatch by separating the student's rollout from teacher supervision. During training, the student first performs entropy-first rollouts that commit uncertain positions and expose potential forks. We then fix the resulting student prefix and distill an NTP teacher under the same context, with answer correctness determining the supervision source. At inference, the student returns to its native confidence-first parallel decoding. With Qwen3-30B-A3B-Base, ForkLeft improves Efficient-DLM-4B on all ten benchmarks, raising MATH500 from 72.60% to 79.60% and consistently outperforming three alternative designs. The gains scale with teacher strength and generalize to SDAR-4B with only $500$ updates. At matched scale, the distilled 4B and 8B students exceed the published SDAR-Chat and OPDLM models on seven benchmarks, showing that DLMs can learn NTP-style reasoning without sacrificing native parallel generation. Code and datasets will be released upon acceptance.
Chinese Translation
自回归下一词预测(NTP)赋予了语言模型强大的推理能力,而扩散语言模型(DLM)提供了灵活的标记顺序和并行生成。我们探究DLM能否在不放弃其原生生成过程的情况下,通过蒸馏获得NTP风格的推理能力。然而,直接蒸馏面临一个根本性不匹配:自回归教师基于左侧前缀进行预测,而DLM可以基于两侧的标记进行条件生成。我们提出ForkLeft,一种蒸馏框架,通过将学生的展开与教师监督分离来解决这一不匹配。训练时,学生首先执行熵优先展开,确定不确定位置并暴露潜在分叉。然后我们固定得到的学生前缀,并在相同上下文中蒸馏NTP教师,由答案正确性决定监督来源。推理时,学生回到其原生的置信度优先并行解码。使用Qwen3-30B-A3B-Base,ForkLeft在全部十个基准上提升了Efficient-DLM-4B,将MATH500从72.60%提升至79.60%,并持续优于三种替代设计。增益随教师强度扩展,并仅需500次更新即可泛化到SDAR-4B。在匹配规模下,蒸馏后的4B和8B学生在七个基准上超过了已发布的SDAR-Chat和OPDLM模型,表明DLM可以在不牺牲原生并行生成的情况下学习NTP风格的推理。代码和数据集将在接收后发布。
cs.AI / 74 / 2609.32469

PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders

PULSE:利用稀疏自编码器识别示例效用特征
Hao, Chenduo, Gao, Chuanbao, Zeng, Pinjun, Zhu, Jingze, Liu, Chonghan, Liu, Zidong, Yang, Xu
Abstract
In-context learning is highly sensitive to demonstration choice, yet most methods select demonstrations using external query-demonstration similarity. Such criteria can miss model-specific signals: Similar demonstrations may activate different internal features and downstream behaviors. We introduce PULSE (Paired Utility Localization over Sparse Encodings), an SAE-based framework for identifying model-internal features associated with demonstration utility and using them for demonstration selection. Using a small labeled discovery set, PULSE samples candidate demonstration sets, measures their zero-shot-relative utility under the target model, and scores SAE features by how their activation differences align with utility differences. The top positive and negative coordinates form a sparse utility-localization vector. We use this vector in two complementary ways: as a signed score for controlled complete-set ranking, and as PULSE-Retriever, which converts its magnitude into a feature-relevance mask for scalable pool-scale retrieval. Across classification, generation, and reasoning benchmarks, PULSE-Retriever improves over the strongest baseline by 2-3 accuracy points, 0.6-0.9 BLEU-4, and 3.2 exact-match points, respectively, while controlled ranking validates the identified features encode a predictive set-level utility signal. Feature inspection and cross-dataset experiments suggest that the identified features capture task-relevant, dataset-conditioned patterns, yet retain utility signals that partially transfer across datasets. Our code is available at https://github.com/aohenuo/PULSE.
Chinese Translation
上下文学习对示例选择高度敏感,然而大多数方法使用外部查询-示例相似度来选择示例。这类标准可能会遗漏模型特定的信号:相似的示例可能激活不同的内部特征和下游行为。我们提出 PULSE(Paired Utility Localization over Sparse Encodings,基于稀疏编码的成对效用定位),一个基于 SAE 的框架,用于识别与示例效用相关的模型内部特征,并将其用于示例选择。使用一个小的带标签的发现集,PULSE 采样候选示例集,在目标模型下测量其相对于零样本的效用,并根据其激活差异与效用差异的对齐程度对 SAE 特征进行评分。最高正坐标和负坐标形成一个稀疏的效用定位向量。我们以两种互补的方式使用该向量:作为受控完整集排序的有符号分数,以及作为 PULSE-Retriever,将其幅度转换为特征相关性掩码,用于可扩展的池级检索。在分类、生成和推理基准测试中,PULSE-Retriever 分别比最强基线提高了 2-3 个准确率点、0.6-0.9 BLEU-4 和 3.2 个精确匹配点,而受控排序验证了所识别的特征编码了预测性的集合级效用信号。特征检查和跨数据集实验表明,所识别的特征捕获了任务相关、数据集条件化的模式,但保留了部分跨数据集迁移的效用信号。我们的代码可在 https://github.com/aohenuo/PULSE 获取。
cs.AI / 75 / 2609.32473

VPEvolve: A Self-Evolving Virtual Process Engineer for Computational Lithography

VPEvolve:面向计算光刻的自进化虚拟工艺工程师
Li, Tianyi, Dong, Wenxuan, Luo, Donger, Wang, Nan, Chen, Yanpeng, Liu, Jiaqi, Zhang, Xinyun, Geng, Hao
Abstract
Optical proximity correction (OPC) recipes grow as engineers add local rules to repair newly discovered lithography hotspots. Each correction can interact with existing rules, while lessons from commercial-tool trials remain scattered across code and logs. \system combines a Virtual Process Engineer (VPE) harness with a Skill Bank of measured engineering experience. The harness equips a frozen language model with process manuals, layout analysis, recipe editing, and commercial-tool evaluation. The actor proposes changes to the global parameters, local targeted rules, or diagnostic trials. After each evaluation, an LLM reflector and curator turn the measured response into evidence-linked judgments. The actor retrieves them before its next trial. Feasible improvements update the retained recipe; every measured trial informs the Skill Bank that guides the next edit. The model weights remain fixed. On a FreePDK45-derived benchmark with ten commercial-tool evaluations per case, \system reduces the mean per-case maximum edge placement error from 18.294 to 5.361 nm on Poly and from 22.052 to 15.692 nm on Metal1. Every final recipe satisfies the predefined quality constraints and improves the maximum error by at least 0.1 nm.
Chinese Translation
光学邻近校正(OPC)配方随着工程师添加局部规则以修复新发现的光刻热点而不断增长。每次校正都可能与现有规则相互作用,而来自商业工具试验的经验教训仍分散在代码和日志中。VPEvolve 将虚拟工艺工程师(VPE)执行框架与测量工程经验的技能库相结合。该框架为冻结的语言模型配备工艺手册、版图分析、配方编辑和商业工具评估能力。行动者提出对全局参数、局部定向规则或诊断试验的更改。每次评估后,大语言模型反思器和策展人将测量响应转化为与证据关联的判断。行动者在下次试验前检索这些判断。可行的改进会更新保留的配方;每次测量试验都会为技能库提供信息,以指导下一次编辑。模型权重保持固定。在基于 FreePDK45 的基准测试中,每个案例进行十次商业工具评估,VPEvolve 将 Poly 层平均每案例最大边缘放置误差从 18.294 nm 降低至 5.361 nm,将 Metal1 层从 22.052 nm 降低至 15.692 nm。每个最终配方都满足预定义的质量约束,并将最大误差至少改善 0.1 nm。
cs.AI / 76 / 2609.32482

From Outcomes to Strategies: Learning Strategy Utility for Mathematical Reasoning

从结果到策略:学习数学推理的策略效用
Zhang, Ruikang, An, Xiao, Shen, Xuli, Sun, Jiaxing, Yu, Xiaoyi, Zeng, Jin, Wu, Jiang, Lin, Tong
Abstract
Reinforcement learning with verifiable rewards has substantially improved mathematical reasoning. However, terminal correctness alone provides limited insight into the quality of high-level strategies, such as theorem selection and subgoal decomposition, when considered separately from their subsequent execution. This paper studies strategy utility, which is defined as the likelihood that a strategy supports a correct downstream solution under a given executor. We introduce SURE, a framework for learning and leveraging relative strategy utility. In this framework, high-level strategies are separated from their detailed reasoning. Based on the pairwise preferences constructed from strategy-conditioned rollouts and teacher-generated contrasts, a Strategy Reward Model is learned to estimate relative strategy utility. During reinforcement learning, the frozen reward model reads only the extracted strategy, whose score is combined with the correctness and format rewards in a sequence-level GRPO objective. Compared with outcome-and-format GRPO baselines, experiments show that SURE improves average pass@1 by 1.87%, 2.64%, and 2.93% across three policy backbones. Our method also achieves competitive or better accuracy than stronger reward baselines while requiring substantially lower GRPO-stage compute.
Chinese Translation
可验证奖励的强化学习已大幅提升数学推理能力。然而,当与后续执行分开考虑时,仅凭最终正确性难以洞察高层策略(如定理选择与子目标分解)的质量。本文研究策略效用,其定义为在给定执行器下,策略支持正确下游解决方案的可能性。我们提出SURE,一个用于学习和利用相对策略效用的框架。在该框架中,高层策略与其详细推理过程分离。基于由策略条件化展开和教师生成的对比构建的成对偏好,学习一个策略奖励模型来估计相对策略效用。在强化学习过程中,冻结的奖励模型仅读取提取出的策略,其分数与正确性和格式奖励在序列级GRPO目标中结合。与结果和格式GRPO基线相比,实验表明SURE在三个策略骨干网络上将平均pass@1分别提高了1.87%、2.64%和2.93%。我们的方法在达到与更强奖励基线相当或更好的准确率的同时,所需GRPO阶段计算量显著降低。
cs.AI / 77 / 2609.32483

Separating Diagnosis from Disease Representation: Dual-View EEG Learning with Neural-Dynamics-Guided Deformation

将诊断与疾病表征分离:神经动力学引导形变的双视图EEG学习
Wang, Jiaying, Shi, Shouqian, Chen, Yutong, Yang, Xu, Chen, Jie, Pan, Xingyu, Zhang, Lei, Zhong, Sheng
Abstract
Electroencephalography (EEG)-based closed-loop neuromodulation calls for a subject-specific structured state, as opposed to a single disease probability, specifying which brain regions are deviant, at which frequencies, and at which lags. Sensor-space models keep the strongest diagnostic evidence without anatomy, source-space models give anatomy at a loss of predictive signal, and post-hoc attributions stay outside the prediction. We separate the two instead of forcing them into one representation, and propose DMD-EEG (Dual-view Multiscale Deformation for EEG), which keeps a fixed scalp spectral expert for diagnosis and models the source-space disease-related representation as a low-rank, sparse, iterative deformation of a healthy neural-dynamics prior in a $46$-region-of-interest (ROI) $\times$ $5$-frequency $\times$ $4$-lag (autocorrelation-timescale) space. The two experts meet only at a fixed decision level, so the source state is architecturally separate from the scalp expert. Across major depressive disorder (MDD), first-episode psychosis (FEP), and Parkinson's disease (PD), decision-level fusion matches the strongest single expert on MDD and FEP and exceeds the source branch on PD. On FEP the source expert is the strongest branch, the task where the deformation contributes most. The source state is an explicit ROI-frequency-lag attribution defined in a shared source coordinate system across montages, which we treat as an anatomically-coordinated predictive representation whose coordinates are directly readable and hypothesis-generating. The highest-saliency coordinates align with established disease circuitry (fronto-limbic-temporal regions in MDD, motor-cortex beta in PD), and the MDD state transfers by rank to an unseen cohort recorded with a different montage.
Chinese Translation
基于脑电图(EEG)的闭环神经调控需要一种被试特定的结构化状态,而非单一的疾病概率,该状态应指明哪些脑区异常、在哪些频率以及哪些滞后。传感器空间模型保留了最强的诊断证据但缺乏解剖信息,源空间模型提供了解剖信息却损失了预测信号,而事后归因则停留在预测之外。我们将两者分离,而不是强行纳入一个表征,并提出DMD-EEG(用于EEG的双视图多尺度形变),它保留一个固定的头皮频谱专家用于诊断,并将源空间疾病相关表征建模为健康神经动力学先验在46个感兴趣区(ROI)×5个频率×4个滞后(自相关时间尺度)空间中的低秩、稀疏、迭代形变。两个专家仅在固定的决策层相遇,因此源状态在架构上与头皮专家分离。在重性抑郁障碍(MDD)、首发精神病(FEP)和帕金森病(PD)中,决策级融合在MDD和FEP上匹配最强的单一专家,并在PD上超过源分支。在FEP上,源专家是最强的分支,这也是形变贡献最大的任务。源状态是一种显式的ROI-频率-滞后归因,定义在跨导联的共享源坐标系中,我们将其视为一种解剖协调的预测表征,其坐标可直接读取并可产生假设。最高显著性的坐标与已确立的疾病环路一致(MDD中的额-边缘-颞区,PD中的运动皮层beta),并且MDD状态按秩迁移到使用不同导联记录的未见队列。
cs.AI / 78 / 2609.32484

Towards Scalable Data Diversification for Language Model Pretraining via Leverage Score Sampling

面向语言模型预训练的可扩展数据多样化:基于杠杆分数采样
Ma, Zailin, Huang, Quzhe, Li, Yujun, Rao, Congyuan, Yang, Yaodong
Abstract
Data selection for language model pretraining faces a fundamental tension between quality and diversity. While quality filtering is empirically effective, it often induces diversity collapse: by favoring texts similar to high-quality reference corpora (e.g., educational or QA-style data), it systematically excludes valuable data from underrepresented domains. In contrast, diversified selection preserves domain balance and encourages robust downstream performance, yet existing methods either focus on coverage-oriented objectives that indirectly enhance diversity, or directly optimize for diversity via costly covariance matrix recomputation that limits scalability. To address these issues, we introduce \textbf{Leverage Score Sampling (Lev)}, which iteratively selects samples that maximally expand the determinantal volume of the embedded data via leverage scores, a computationally efficient criterion that eliminates matrix recomputation and enables scalable selection. Empirically, Lev delivers up to $72\times$ speedup and improves dataset diversity, measured by the Vendi score, by $9.2\%$ over the strong diversification baseline \textbf{DiSF}. On CommonCrawl (CC) web data selection, Lev improves accuracy across seven downstream tasks by up to $1.31\%$ over existing baselines. For domains where robust quality criteria are inherently difficult to define (e.g., code), Lev serves as an effective unsupervised curation alternative: on StarCoderData, the selected subset reduces bits-per-byte by $3.08\%$ over DiSF. Notably, we uncover a cross-domain collapse of quality filtering: CC data filtered by DCLM-fastText fail to retain sufficient code-related content, yielding inferior code performance relative to Lev-selected data. These findings advocate for integrating diversity-aware practices into quality filtering for more effective data curation in language model pretraining.
Chinese Translation
语言模型预训练的数据选择面临质量与多样性之间的根本矛盾。虽然质量过滤在经验上是有效的,但它常常导致多样性崩塌:通过偏好与高质量参考语料库(例如教育或问答风格数据)相似的文本,它系统地排除了来自代表性不足领域的有价值数据。相比之下,多样化选择保持了领域平衡并促进稳健的下游性能,但现有方法要么专注于间接增强多样性的覆盖导向目标,要么通过昂贵的协方差矩阵重计算直接优化多样性,这限制了可扩展性。为了解决这些问题,我们引入了杠杆分数采样(Lev),它通过杠杆分数迭代地选择能够最大程度扩展嵌入数据的行列式体积的样本,这是一种计算高效的准则,消除了矩阵重计算并实现了可扩展的选择。实验表明,Lev实现了高达72倍的加速,并将数据集多样性(以Vendi分数衡量)比强大的多样化基线DiSF提高了9.2%。在CommonCrawl (CC)网络数据选择上,Lev在七个下游任务上将准确率比现有基线提高了高达1.31%。对于本质上难以定义稳健质量标准的领域(例如代码),Lev作为一种有效的无监督管理替代方案:在StarCoderData上,所选子集将每字节比特数(bits-per-byte)比DiSF降低了3.08%。值得注意的是,我们发现了质量过滤的跨域崩塌:通过DCLM-fastText过滤的CC数据无法保留足够的代码相关内容,导致代码性能劣于Lev选择的数据。这些发现倡导将多样性感知实践整合到质量过滤中,以便在语言模型预训练中进行更有效的数据管理。
cs.AI / 79 / 2609.32489

When Helpful Text Hurts: Option-Redirecting Bias in Vision-Language Models

当有益文本带来伤害:视觉语言模型中的选项重定向偏差
Thanh, Tam Le Thi, Van, Hoang Tran, Nguyen-Le, Hong-Hanh, Ngo, Thanh Duc
Abstract
In tri-modal visual question answering (VQA), auxiliary text is commonly used to complement visual and textual inputs, yet its reliability is often uncontrolled. While prior work studies modality conflicts in general, it remains unclear how different types of unreliable auxiliary text affect answer selection under fixed image-question-option contexts. In this work, we show that the most harmful auxiliary text is not necessarily the most factually incorrect, but the one that aligns with the question while contradicting the image and favoring a specific distractor, leading to systematic redirection of model predictions. To isolate this effect, we introduce the Textual Reliability Ladder, a controlled diagnostic protocol that decomposes auxiliary text along three axes: image consistency, question relevance, and option support. Across multiple datasets (ScienceQA, VCR, A-OKVQA, Causal-VidQA) and recent VLMs, we find that such distractor-supporting text induces the largest accuracy drops (up to 53.1%) and concentrates errors on specific incorrect options. To mitigate this failure mode, we propose a training-free inference-time intervention that explicitly counteracts this redirection effect via noise-stability steering and dynamic grounding, reducing redirected errors while largely preserving performance under faithful text. Our results highlight that auxiliary-text reliability must be understood at the decision level, rather than solely through factual correctness, and provide a practical pathway toward more robust tri-modal reasoning.
Chinese Translation
在三模态视觉问答(VQA)中,辅助文本通常用于补充视觉和文本输入,但其可靠性往往不受控制。虽然先前的工作一般性地研究了模态冲突,但在固定的图像-问题-选项上下文中,不同类型的不可靠辅助文本如何影响答案选择仍不清楚。在这项工作中,我们表明,最具危害性的辅助文本不一定是最事实上错误的,而是与问题一致、与图像矛盾并偏向特定干扰项的那一个,从而导致模型预测的系统性重定向。为了分离这种效应,我们引入了文本可靠性阶梯(Textual Reliability Ladder),一种受控的诊断协议,它沿三个轴分解辅助文本:图像一致性、问题相关性和选项支持。在多个数据集(ScienceQA、VCR、A-OKVQA、Causal-VidQA)和最近的VLM上,我们发现这种支持干扰项的文本会导致最大的准确率下降(高达53.1%),并将错误集中在特定的错误选项上。为了缓解这种失败模式,我们提出了一种无需训练的推理时干预方法,通过噪声稳定性引导和动态接地来明确抵消这种重定向效应,减少重定向错误,同时在忠实文本下基本保持性能。我们的结果强调,辅助文本的可靠性必须在决策层面理解,而不仅仅是通过事实正确性,并为更稳健的三模态推理提供了一条实用途径。
cs.AI / 80 / 2609.32490

RepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent Systems

RepoMAS:用问题驱动的多智能体系统解决渐进式指定的任务
Song, Yuchen, Chen, Andong, Zhu, Wenxin, Yang, Muyun, Zhao, Tiejun
Abstract
LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution. In practice, user requests are often incomplete, and additional requirements may only become clear during reasoning, tool use, or execution. We refer to such problems as progressively specified tasks. To systematically study this setting, we introduce ProgSpec, a benchmark that evaluates final outputs against requirements explicitly stated in the initial request and additional requirements supported by the available task evidence. We further propose RepoMAS, an issue-driven multi-agent framework inspired by open-source project management. RepoMAS records newly discovered requirements, conflicts, and failures as structured Issues and uses them to revise the task specification and execution structure during problem solving. Across ProgSpec and five existing benchmarks, RepoMAS achieves the best performance. Further analyses show that its issue-driven revision and repository maintenance mechanisms consistently contribute to performance. These results highlight the importance of allowing MASs to revise not only how a task is solved, but also revise their explicit representation of task requirements during execution.
Chinese Translation
基于大语言模型(LLM)的多智能体系统(MASs)在解决复杂任务方面展现出强大潜力,但大多数系统假设任务需求在执行前已充分明确。实际上,用户请求往往不完整,附加需求可能只有在推理、工具使用或执行过程中才变得清晰。我们将此类问题称为渐进式指定的任务。为了系统地研究这一设定,我们引入了 ProgSpec,一个根据初始请求中明确陈述的需求以及由可用任务证据支持的附加需求来评估最终输出的基准测试。我们进一步提出了 RepoMAS,一个受开源项目管理启发的问题驱动多智能体框架。RepoMAS 将新发现的需求、冲突和失败记录为结构化的 Issue,并利用它们在问题解决过程中修订任务规范和执行结构。在 ProgSpec 和五个现有基准测试中,RepoMAS 取得了最佳性能。进一步的分析表明,其问题驱动的修订和仓库维护机制持续对性能做出贡献。这些结果凸显了允许 MASs 不仅修订任务如何解决,而且在执行过程中修订其任务需求的显式表示的重要性。
cs.AI / 81 / 2609.32492

Beyond Prompt or Skill? Attribution-Guided Optimization of Modular LLM Programs

超越提示还是技能?归因引导的模块化LLM程序优化
Shou, Haoran, Liu, Haoyue, Huo, Yu, Zeng, Kun, Tang, Xiaoying
Abstract
Large language models can solve increasingly diverse reasoning tasks, yet their performance remains highly sensitive to task prompts, intermediate instructions, and the way reusable problem-solving knowledge is incorporated. Existing optimization methods usually focus on only one part of this design space: they either optimize a monolithic prompt, or separately induce and refine skills from model traces. As a result, they lack a principled mechanism for deciding which component should be updated when failures occur, and they rarely optimize prompts, skills, and skill-use policies in a unified framework. We propose SPARO (Skill, Prompt, And Routing Optimization), a framework that jointly optimizes task instructions, reusable skill blocks, and routing rules. It performs controlled counterfactual evaluations, converts examples' effects into a probabilistic responsibility distribution over prompt, skill, and routing components, samples one component from that distribution, and applies the corresponding targeted mutation. This design moves language-program optimization beyond global prompt rewriting toward structured, reusable, and selectively activated task knowledge. Across five benchmarks and five worker models, SPARO consistently outperforms both prompt-centered and skill-centered optimization baselines. These results suggest that effective language-program optimization depends not only on discovering useful task knowledge, but also on deciding where that knowledge should be stored and when it should be activated.
Chinese Translation
大语言模型能够解决日益多样化的推理任务,但其性能仍然对任务提示、中间指令以及可复用的问题求解知识的整合方式高度敏感。现有的优化方法通常只关注该设计空间的一个部分:它们要么优化整体式提示,要么从模型轨迹中单独归纳和精炼技能。因此,它们缺乏一种原则性机制来决定在失败发生时应该更新哪个组件,并且很少在统一框架中优化提示、技能和技能使用策略。我们提出SPARO(技能、提示和路由优化),一个联合优化任务指令、可复用技能块和路由规则的框架。它执行受控的反事实评估,将样本的影响转化为关于提示、技能和路由组件的概率责任分布,从该分布中采样一个组件,并应用相应的针对性变异。这种设计使语言程序优化超越了全局提示重写,转向结构化、可复用且选择性激活的任务知识。在五个基准和五个工作模型上,SPARO始终优于以提示为中心和以技能为中心的优化基线。这些结果表明,有效的语言程序优化不仅取决于发现有用的任务知识,还取决于决定该知识应存储在何处以及何时应被激活。
cs.AI / 82 / 2609.32498

DAAF: From Failure Localization to Editable System Assets in LLM Agents

DAAF:从故障定位到LLM智能体中的可编辑系统资产
Yuan, Xiaoyang, Liu, Qi, Ruan, Yubin, Mou, Xinyi, Zhang, Zhuomeng, Wang, Wenjin, Jiao, Hanying, Wu, Di, Xu, Mingye, Bin, Yi, Feng, Ke, Sun, Zixun
Abstract
Deployed LLM agents increasingly rely on persistent, versioned system assets such as routing rules, knowledge segments, prompt instructions, and reusable skills. Failure-localization methods can identify where an error manifests in an agent or execution trace, but repair requires a different decision: which editable system asset should be changed, and is that change expected to improve the task outcome? We study this gap through component-attribute failure attribution, where diagnosis targets versioned, addressable items rather than execution locations. We propose the Detection-Aware Attribution Framework (DAAF), which learns the effects of valid attribute replacements and amortizes this intervention evidence into deployment-time diagnosis. DAAF combines sparse and noisy failure signals to decide whether intervention is warranted, learns component-type-conditioned replacement effects from controlled replays evaluated by executable task outcomes, and shares supervision across requests with compatible intervention responses. At diagnosis time, DAAF uses only the observed execution, registered candidates, and available failure signals; it requires neither counterfactual replay nor task reward and returns no_change, a repair target, or an unresolved decision when evidence is insufficient. On held-out tau^2-bench Telecom tasks, DAAF achieves 80.72% attribute Hit@1, recovers 62.65% of failed executions while limiting clean-task regression to 3.23%, and reaches 71.93% overall task success. These results show that intervention-grounded attribute attribution can connect failure localization to executable system repair.
Chinese Translation
已部署的LLM智能体日益依赖持久化、版本化的系统资产,如路由规则、知识片段、提示指令和可复用技能。故障定位方法能够识别错误在智能体或执行轨迹中显现的位置,但修复需要做出不同的决策:应更改哪个可编辑系统资产,以及该更改是否预期能改善任务结果?我们通过组件-属性故障归因研究这一差距,其中诊断针对的是版本化、可寻址的项,而非执行位置。我们提出了检测感知归因框架(DAAF),该框架学习有效属性替换的效应,并将这种干预证据摊销到部署时诊断中。DAAF结合稀疏且含噪的故障信号来判断是否需要进行干预,从通过可执行任务结果评估的受控重放中学习以组件类型为条件的替换效应,并在具有兼容干预响应的请求之间共享监督。在诊断时,DAAF仅使用观察到的执行、注册的候选和可用的故障信号;它既不需要反事实重放,也不需要任务奖励,并在证据不足时返回no_change、修复目标或未解决的决策。在留出的tau^2-bench Telecom任务上,DAAF实现了80.72%的属性Hit@1,恢复了62.65%的失败执行,同时将干净任务回归限制在3.23%,并达到71.93%的整体任务成功率。这些结果表明,基于干预的属性归因可以将故障定位与可执行系统修复连接起来。
cs.AI / 83 / 2609.32502

TreeRef-BFN: Equivariance-Free De Novo Molecule Generation based on 2D Topology and Internal 3D Geometry

TreeRef-BFN:基于2D拓扑和内3D几何的免等变性从头分子生成
Sun, Ruiqing, Yang, Sen, Feng, Dawei, Ding, Bo, Wang, Yijie, Wang, Huaimin
Abstract
De novo 3D molecular generation jointly models molecular size, topology, and geometry. Most methods pre-sample molecular size and generate Cartesian coordinates, limiting variable-size conditional tasks such as fragment completion and scaffold decoration while often relying on equivariant architectures. Internal-coordinate methods avoid rigid-body redundancy but typically require a known molecular graph or autoregressive construction, which may accumulate errors. We propose TreeRef, a tree-based molecular representation that assigns molecular topology and topology-dependent local 3D geometry to a naturally variable-size tree. RingRef nodes encode ring closures while preserving the tree structure, while Null nodes allow molecular size to emerge directly from node occupancy. Based on TreeRef, we develop TreeRef-BFN, a Bayesian Flow Network with a standard Transformer backbone that globally couples these locally defined variables and jointly generates discrete molecular variables and continuous local geometry. A single pretrained TreeRef-BFN supports unconditional generation and variable-size structure-conditioned 3D generation through masking alone, without retraining. Empirical studies demonstrate strong chemical validity, molecular stability, and diversity, accurate local geometric distributions, fast sampling, and competitive property-conditioned generation, establishing TreeRef-BFN as an efficient and flexible framework for 3D molecular generation.
Chinese Translation
从头3D分子生成联合建模分子大小、拓扑和几何。大多数方法预先采样分子大小并生成笛卡尔坐标,限制了可变大小的条件任务(如片段完成和支架装饰),同时通常依赖等变架构。内坐标方法避免了刚体冗余,但通常需要已知的分子图或自回归构建,这可能累积误差。我们提出TreeRef,一种基于树的分子表示,将分子拓扑和依赖于拓扑的局部3D几何分配给自然可变大小的树。RingRef节点在保留树结构的同时编码环闭合,而Null节点允许分子大小直接从节点占用中涌现。基于TreeRef,我们开发了TreeRef-BFN,一种具有标准Transformer骨干的贝叶斯流网络,它全局耦合这些局部定义的变量,并联合生成离散分子变量和连续局部几何。单个预训练的TreeRef-BFN仅通过掩码支持无条件生成和可变大小的结构条件3D生成,无需重新训练。实证研究展示了强大的化学有效性、分子稳定性和多样性、准确的局部几何分布、快速采样以及具有竞争力的性质条件生成,确立了TreeRef-BFN作为高效灵活的3D分子生成框架。
cs.AI / 84 / 2609.32511

Learning from Others, Acting for You: Cross-User Memory Sharing for LLM Agents

从他人学习,为你行动:面向LLM智能体的跨用户记忆共享
Hu, Jinming, Zhao, Haodong, Jia, Qi, Chen, Die, Zhao, Tianhang, Duan, Sufeng, Liu, Gongshen
Abstract
Large language model (LLM) agents serving different users often solve related tasks, yet separate user histories can leave reusable experience inaccessible to other agents. Pooling memories expands access but risks transferring preferences that conflict with the receiving user's requirements. We introduce ShareMem, a memory architecture that shares reusable experience while grounding its application in the receiving user's own preferences. Shared experiences indicate how to act and which preferences to consult; the receiving user's memory supplies their concrete values. Two-stage consolidation refines experience locally before integrating accepted edits into a shared pool. During execution, scope-first retrieval jointly selects local and shared experiences under a common entry budget, while a user-bound channel supports initial and agent-initiated preference retrieval. We evaluate ShareMem across web navigation (Mind2Web), online personalized interaction (VitaBench~2.0), and multi-session coding (MemoryCode) with four backbone models. It improves step success, average task success, and dialogue-macro coding scores, respectively, over matched user-local memory across all four models. Ablations favor two-stage consolidation for smaller shared pools, lower induction token usage, and better downstream performance, and support complementarity between experience guidance and active preference retrieval. Further analyses show that sharing helps most when relevant local experience is scarce, while source quality and cross-user preference interference limit useful transfer.
Chinese Translation
服务于不同用户的大语言模型(LLM)智能体通常解决相关任务,然而分离的用户历史可能导致可复用的经验无法被其他智能体访问。汇集记忆扩大了访问范围,但存在传递与接收用户需求相冲突的偏好的风险。我们提出 ShareMem,一种记忆架构,它共享可复用经验,同时将其应用建立在接收用户自身偏好的基础上。共享经验指示如何行动以及参考哪些偏好;接收用户的记忆提供其具体偏好值。两阶段整合先在本地精炼经验,再将接受的编辑集成到共享池中。在执行过程中,范围优先检索在共同条目预算下联合选择本地和共享经验,而用户绑定通道支持初始和智能体发起的偏好检索。我们在网络导航(Mind2Web)、在线个性化交互(VitaBench 2.0)和多会话编码(MemoryCode)上使用四个骨干模型评估 ShareMem。在所有四个模型上,相较于匹配的用户本地记忆,它分别提高了步骤成功率、平均任务成功率和对话宏编码分数。消融实验表明,两阶段整合有利于更小的共享池、更低的归纳 token 使用量和更好的下游性能,并支持经验指导与主动偏好检索之间的互补性。进一步分析表明,当相关本地经验稀缺时,共享帮助最大,而来源质量和跨用户偏好干扰限制了有用的迁移。
cs.AI / 85 / 2609.32514

From Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace Diagnosis

从异常到失败:构建面向智能体轨迹诊断的因果错误图
Yang, Shu-Xun, Wang, Yidong, Feng, Zhuoer, Wen, Bosi, Gui, Jiayi, Yang, Dayong, Yu, Wenbo, Zhang, Haoke, Tang, Jie, Wang, Cunxiang
Abstract
LLM-driven agents are increasingly deployed in complex applications, where long agentic traces make failures difficult to diagnose. Existing trace diagnosis methods often conflate anomalies, errors, and failures, making diagnostic targets ambiguous; they also lack structured modeling of how causally relevant errors propagate and amplify into final task failures, resulting in unreliable failure attribution. To address these problems, we propose CEG-Agent, a tool-augmented agentic framework for causal diagnosis of agentic traces. Specifically, CEG-Agent introduces an explicit taxonomy of anomalies, errors, and failures, and constructs Causal Error Graphs (CEGs), a unified typed representation that links execution events, diagnostic nodes, and failure outcomes through causal relations. To evaluate causal trace diagnosis, we further construct CEG-Bench, a fully agent-annotated benchmark with high-confidence, consensus-derived CEG annotations obtained through an Adversarial Agentic Adjudication Protocol (AAAP). We validate the resulting annotations against an expert-curated human gold set, which shows close agreement with the automatic annotations. Experiments on CEG-Bench demonstrate that CEG-Agent achieves state-of-the-art performance under both semantically relaxed and structurally exact evaluation criteria. Our code is publicly available.
Chinese Translation
LLM驱动的智能体越来越多地部署在复杂应用中,其中长智能体轨迹使得失败难以诊断。现有的轨迹诊断方法常常混淆异常、错误和失败,导致诊断目标模糊;它们也缺乏对因果相关错误如何传播并放大为最终任务失败的结构化建模,从而导致不可靠的失败归因。为了解决这些问题,我们提出了CEG-Agent,一个工具增强的智能体框架,用于智能体轨迹的因果诊断。具体而言,CEG-Agent引入了异常、错误和失败的显式分类法,并构建了因果错误图(CEGs),这是一种统一的类型化表示,通过因果关系将执行事件、诊断节点和失败结果联系起来。为了评估因果轨迹诊断,我们进一步构建了CEG-Bench,一个完全由智能体标注的基准,其高置信度、共识推导的CEG标注是通过对抗性智能体裁决协议(AAAP)获得的。我们根据专家整理的人类黄金标准验证了生成的标注,结果显示与自动标注高度一致。在CEG-Bench上的实验表明,CEG-Agent在语义宽松和结构精确的评估标准下均达到了最先进的性能。我们的代码已公开。
cs.AI / 86 / 2609.32517

LocalProp: Neuro-Localized Memory-Efficient Backpropagation

LocalProp:神经局部化的内存高效反向传播
Grigore, Diana-Nicoleta, Georgescu, Iuliana, Ionescu, Radu Tudor
Abstract
The current deep learning training paradigm employs end-to-end backpropagation, regardless of the training stage, i.e. pre-training or fine-tuning. However, backpropagating through the entire model is neither biologically plausible nor memory efficient, since learning inside the brain is highly localized. Therefore, we propose LocalProp, a training procedure that locally updates the weights of a model. Our neuro-localized weight updates follow the "pre-training then fine-tuning" paradigm, where the pre-training is based on I-JEPA. After locally updating the weights, a pruning operation is performed, followed by a short final fine-tuning phase. Pruning helps by sending the learning signal from higher blocks to lower blocks. We perform experiments on several datasets, including large-scale benchmarks such as ImageNet, and empirically show that LocalProp reaches good performance at a fraction of GPU peak memory. By varying the number of jointly optimized blocks, we identify gradient-propagation span as a practical control over the accuracy-memory trade-off.
Chinese Translation
当前的深度学习训练范式采用端到端反向传播,无论训练阶段是预训练还是微调。然而,通过整个模型进行反向传播既不符合生物学合理性,也不节省内存,因为大脑内部的学习是高度局部化的。因此,我们提出了 LocalProp,一种局部更新模型权重的训练过程。我们的神经局部化权重更新遵循“先预训练后微调”的范式,其中预训练基于 I-JEPA。在局部更新权重之后,执行剪枝操作,随后是一个简短的最终微调阶段。剪枝有助于将学习信号从较高块传递到较低块。我们在多个数据集上进行了实验,包括 ImageNet 等大规模基准,并实证表明 LocalProp 在仅消耗一小部分 GPU 峰值内存的情况下达到了良好性能。通过改变联合优化的块数量,我们将梯度传播跨度确定为精度-内存权衡的一种实用控制手段。
cs.AI / 87 / 2609.32519

STR: Supervised Transcoder Replacement for Reducing Steering Side Effects

STR:用于减少引导副作用的监督式转码器替换
Yu, Haonan, Liu, Junhao, Yan, Zhenyu, Lin, Haoran, Zhang, Xin
Abstract
Model steering can strengthen a target behavior while degrading other useful behaviors. We introduce Supervised Transcoder Replacement (STR) to reduce these side effects for existing steering methods, including those fitted without a protection objective. STR learns a replacement for the multilayer perceptron (MLP) computation at the steering layer through supervision for target control, non-target preservation, and fidelity without steering. Selected steering methods then fit directions on the frozen replacement while retaining their own fitting objectives. We evaluate three steering methods across Gemma and Llama models using Corrigibility preferences and four harmful-request safety datasets. SALAD-Bench supplies protection training data and a separate in-distribution evaluation split; HarmBench, AdvBench, and StrongREJECT are reserved for out-of-distribution testing. STR substantially reduces steering side effects on the in-distribution evaluation and extends this protection to the unseen safety datasets while retaining effective target control. For target-only supervised steering vectors, pooled out-of-distribution attack success rate falls from 42.46% to 14.42% on Gemma-3-4B and from 34.97% to 12.91% on Gemma-3-12B. These results show that replacement training can benefit steering methods fitted without protection objectives.
Chinese Translation
模型引导可以强化目标行为,同时降低其他有用行为。我们引入监督式转码器替换(STR)来减少现有引导方法的这些副作用,包括那些在没有保护目标的情况下拟合的方法。STR通过针对目标控制、非目标保持以及无引导保真度的监督,学习替代引导层中多层感知器(MLP)计算的方法。然后,选定的引导方法在冻结的替换上拟合方向,同时保留其自身的拟合目标。我们使用可纠正性偏好和四个有害请求安全数据集,在Gemma和Llama模型上评估了三种引导方法。SALAD-Bench提供保护训练数据和单独的分布内评估划分;HarmBench、AdvBench和StrongREJECT则保留用于分布外测试。STR显著减少了分布内评估中的引导副作用,并将这种保护扩展到未见过的安全数据集,同时保持有效的目标控制。对于仅目标监督的引导向量,在Gemma-3-4B上,汇总的分布外攻击成功率从42.46%下降到14.42%;在Gemma-3-12B上从34.97%下降到12.91%。这些结果表明,替换训练可以使在没有保护目标的情况下拟合的引导方法受益。
cs.AI / 88 / 2609.32521

MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents

MemAgent:学习为LLM智能体管理异构记忆提供者
Wei, Yongxian, Zhao, Yilin, Cheng, Runxi, Chen, Xinrui, Yuan, Chun, Wang, Yaoru, Yan, Jiahong, Li, Dian
Abstract
Current agents remain largely stateless across tasks, limiting their ability to continually improve from prior interactions and making memory essential for long-horizon agentic behavior. Existing memory methods seek to reuse past experience, but most rely on a single memory representation (e.g., trajectories, reflections, skills, structured knowledge) whose effectiveness varies across task distributions. Rethinking this design space, we evaluate 13 memory methods and find that no single method generalizes across benchmarks, revealing the potential of managing heterogeneous memory providers. We formulate agent memory as a routing problem in which a memory agent decides which memory provider to retrieve from, whether to inject short-term memory, and which providers should store the resulting experience. Based on this perspective, we propose MemAgent, featuring a content-aware routing architecture and a training-data synthesis pipeline. The routing architecture combines content-aware probing before retrieval, short-term memory gating during execution, and selective multi-provider storage, while the training pipeline synthesizes phase-specific supervision for routing decisions. Across GAIA, WebWalkerQA, and xBench-DS, MemAgent improves average accuracy by 10.0% and outperforms every individual memory method across all three benchmarks. These gains come with less than 0.3% routing overhead and a 12% reduction in average task steps.
Chinese Translation
当前智能体在任务间基本保持无状态,限制了其从先前交互中持续改进的能力,使得记忆对于长时程智能体行为至关重要。现有记忆方法试图复用过去经验,但大多依赖单一记忆表示(如轨迹、反思、技能、结构化知识),其有效性因任务分布而异。重新审视这一设计空间,我们评估了13种记忆方法,发现没有单一方法能在不同基准上泛化,揭示了管理异构记忆提供者的潜力。我们将智能体记忆形式化为一个路由问题,其中记忆智能体决定从哪个记忆提供者检索、是否注入短期记忆,以及哪些提供者应存储产生的经验。基于这一视角,我们提出MemAgent,其具有内容感知路由架构和训练数据合成流水线。该路由架构结合了检索前的内容感知探测、执行期间的短期记忆门控以及选择性多提供者存储,而训练流水线为路由决策合成阶段特定的监督信号。在GAIA、WebWalkerQA和xBench-DS上,MemAgent将平均准确率提高了10.0%,并在所有三个基准上超越了每一种单独的记忆方法。这些增益伴随着低于0.3%的路由开销和平均任务步骤减少12%。
cs.AI / 89 / 2609.32522

Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations

超越双人记忆:用于多方口语对话的交互感知多模态记忆与自适应智能体检索
Jia, Wenxu, Cheng, Xize, Zhang, Zihan, Fu, Dongjie, Li, Linjun, Chen, Wenshi, Wu, Yangyang, Jin, Tao
Abstract
Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational content, identifying participants across sessions, and retaining who speaks to whom. To this end, we propose VoxPolyMem, an interaction-aware multimodal memory framework combining incremental speaker identification with a memory hierarchy comprising interaction memory, fact memory, and participant profiles. We formulate retrieval as sequential decision-making, where an agent rewrites queries and selects retrieval tools and memory layers based on accumulated evidence to address information gaps. We further introduce Evidence-Gain GRPO (EG-GRPO), which uses round-wise credit assignment to encourage complementary evidence acquisition. We also construct VoxPolyBench to evaluate memory evolution, personalized answering, memory retrieval and reasoning, and interaction reasoning and attribution in multi-party spoken conversations. VoxPolyMem achieves an overall score of 85.0 on VoxPolyBench, surpassing the strongest evaluated baseline by 23.6 points. On Mem-Gallery and H2HMem-Multi, it scores 89.6 and 74.4, respectively, exceeding the strongest evaluated public memory baselines by over 8 points each. These results highlight its potential for persistent, personalized assistance in multi-party multimodal interactions. Code and datasets are available at https://voxpolymem.github.io/VoxPolyBench/demo/
Chinese Translation
长期记忆使智能体能够跨会话积累信息并进行推理,然而现有研究主要关注双人文本或图文对话,使得多方口语对话的长期记忆尚未得到充分探索。这一场景需要保留对话内容、跨会话识别参与者,并记录谁对谁说话。为此,我们提出VoxPolyMem,一个交互感知的多模态记忆框架,它将增量式说话人识别与包含交互记忆、事实记忆和参与者画像的记忆层次结构相结合。我们将检索形式化为序列决策过程,其中智能体基于累积的证据重写查询,并选择检索工具和记忆层,以弥补信息缺口。我们进一步提出证据增益GRPO(EG-GRPO),它使用轮次级的信用分配来鼓励互补证据的获取。我们还构建了VoxPolyBench,用于评估多方口语对话中的记忆演化、个性化回答、记忆检索与推理,以及交互推理与归因。VoxPolyMem在VoxPolyBench上取得了85.0的总体得分,超过最强的评估基线23.6分。在Mem-Gallery和H2HMem-Multi上,它分别获得89.6和74.4分,超过最强的公开记忆基线各8分以上。这些结果凸显了其在多方多模态交互中提供持久、个性化辅助的潜力。代码和数据集可在https://voxpolymem.github.io/VoxPolyBench/demo/获取。
cs.AI / 90 / 2609.32527

AmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared Responses

AmbiModBench:超越共享响应的基因扰动预测基准测试
Huang, Sikai, Yang, Zhiwen, Yu, Kai, Chen, Jiayuan, Li, Stan Z.
Abstract
Predicting cellular responses to genetic perturbations helps prioritize experiments in single-cell genomics, where exhaustive measurement is infeasible. While computational models increasingly predict these responses, three evaluation deficiencies obscure what their scores demonstrate. First, absolute metrics cannot separate target-specific predictions from a shared background response. Second, common metrics remain high under gene shuffling, so gene-level accuracy is never verified. Third, a score at one training size says nothing about coverage, which depends on representation-space proximity and response-constraining power. We propose AmbiModBench, a specificity-aware, gene-resolved and coverage-aware benchmark. It pairs every score with a training-mean reference fitted on the same split, screens each readout by gene-coordinate permutation, and links embedding distance to response variation. Across K562, RPE1 and Norman, strong absolute scores largely reflect shared background rather than target-specific learning. Widely used readouts track response magnitude distributions rather than the affected genes. Detectable gain follows representation-space coverage rather than training-set size. Nonetheless, on RPE1 the protocol yields a reproducible target-specific gain across five additional splits and three gene selections, which absolute scores alone cannot distinguish from shared background.
Chinese Translation
预测细胞对基因扰动的响应有助于在单细胞基因组学中优先安排实验,因为穷尽测量不可行。尽管计算模型越来越多地预测这些响应,但三个评估缺陷模糊了其分数所证明的内容。首先,绝对指标无法将靶标特异性预测与共享背景响应区分开。其次,在基因打乱下常用指标仍然很高,因此基因水平的准确性从未得到验证。第三,在一个训练规模下的分数对覆盖范围没有任何说明,而覆盖范围取决于表示空间邻近性和响应约束能力。我们提出了AmbiModBench,一个特异性感知、基因解析和覆盖感知的基准。它将每个分数与在同一划分上拟合的训练均值参考配对,通过基因坐标置换筛选每个读出,并将嵌入距离与响应变异联系起来。在K562、RPE1和Norman中,强绝对分数主要反映共享背景而非靶标特异性学习。广泛使用的读出跟踪的是响应幅度分布而非受影响的基因。可检测的增益遵循表示空间覆盖范围而非训练集大小。尽管如此,在RPE1上,该协议在五个额外划分和三个基因选择中产生了可重复的靶标特异性增益,而仅凭绝对分数无法将其与共享背景区分开。
cs.AI / 91 / 2609.32528

Fail Loudly: An Auditable Runtime for Agentic Data Analysis

大声失败:面向智能体数据分析的可审计运行时
Yan, Hanxu, Deng, Langxuan, Wang, Zhengle, Wang, Yibo, Liu, Chunwei
Abstract
Large language models (LLMs) have enabled data-science agents to automate multi-step analyses over heterogeneous files. However, incorrect choices regarding data sources, scope, or statistical definitions often lead to silent errors: computations execute successfully but produce plausible yet incorrect outputs that fail to answer the intended question. To mitigate this, we present RADAR, an auditable runtime that makes an agent's analytical choices inspectable and supports their revision through execution feedback. RADAR operates through three core mechanisms. First, an evidence-preserving exploration module retrieves task-relevant content while retaining source locations and observation coverage. Next, the runtime uses typed operators to record the agent's declared inputs, operation arguments, and resulting observations. Finally, runtime validation checks proposed operations against these observations. When a conflict is detected, the runtime rejects the operation or provides diagnostic feedback, allowing the agent to revise its choices before errors propagate. This design enables agents to fail loudly while leaving semantic interpretation to the LLM. On KramaBench, RADAR achieves overall scores of 0.723 with full source retrieval and 0.747 with gold sources supplied, corresponding to relative gains of 35.9% and 28.8% over the strongest baselines. Beyond KramaBench, RADAR achieves relative performance gains of 14.0% on DA-Code and 59.3% on DABStep, demonstrating its applicability across diverse agentic data-analysis workflows.
Chinese Translation
大语言模型(LLMs)使数据科学智能体能够自动化对异构文件的多步分析。然而,关于数据源、范围或统计定义的不正确选择往往导致静默错误:计算成功执行,但产生看似合理却错误的输出,未能回答预期问题。为了缓解这一问题,我们提出了RADAR,一种可审计运行时,它使智能体的分析选择可被检查,并通过执行反馈支持其修正。RADAR通过三个核心机制运作。首先,一个证据保留的探索模块检索任务相关内容,同时保留源位置和观察覆盖范围。其次,运行时使用类型化算子记录智能体声明的输入、操作参数以及产生的观察结果。最后,运行时验证根据这些观察结果检查提议的操作。当检测到冲突时,运行时会拒绝该操作或提供诊断反馈,允许智能体在错误传播之前修正其选择。这种设计使智能体能够大声失败,同时将语义解释留给LLM。在KramaBench上,RADAR在完整源检索下总体得分为0.723,在提供黄金源的情况下为0.747,相对于最强基线分别相对提升了35.9%和28.8%。在KramaBench之外,RADAR在DA-Code上相对性能提升了14.0%,在DABStep上提升了59.3%,展示了其在不同智能体数据分析工作流中的适用性。
cs.AI / 92 / 2609.32533

LLMAdBench: A Human Preference Benchmark for Advertising in LLM Responses

LLMAdBench:面向LLM响应中广告的人类偏好基准
Ai, Rui, Liu, Yuqing, Qiu, Sitao, Qiao, Yun, Wang, Yuhan, Wang, Jessica Xiwen, Yang, Yiqi, Huang, Lihong, Sun, Ruiyao, Zhang, Kaifeng, Ding, Shengze, He, Jiaqi, Wang, Xinman, Gao, Tianhao, Qin, Jimmy, Lin, Jianghao, Wang, Chonghuan
Abstract
Inserting advertisements (ads) into consumer-facing LLM output is emerging as a new business model, but there is little shared evidence on how such ad insertion should be evaluated or how it affects user preferences. We introduce LLMAdBench, a human-preference benchmark for studying advertising in LLM-generated content. The benchmark isolates a simple but practically important decision: given a user conversation, an LLM response, and a matched advertisement, where should the ad be placed? Our dataset compares pairs of responses that differ only in ad position while holding all other conditions fixed including the user query, base answer, advertisement, and disclosure condition. Human annotators evaluate each pair based on six criteria from both advertiser's and user's perspectives. The resulting benchmark contains more than 18000 human judgments across two disclosure conditions: explicitly labeling the ad as sponsored and merging it into the response without disclosure. We use LLMAdBench to evaluate eight frontier LLMs as preference judges and find that they are not reliable substitutes for human evaluation. Even the most stable models reverse roughly one quarter of their decisions when the presentation order is swapped, agreement across models is low, and their placement preferences differ systematically from those of human annotators. Moreover, LLMAdBench contains substantial learnable signal. In particular, a Qwen3-8B model fine-tuned on the human preferences improves substantially over its base model and outperforms all zero-shot frontier judges on the held-out prediction task. Beyond model evaluation, LLMAdBench provides quantitative evidence on the advertiser-user trade-off and shows that the sponsorship disclosure systematically changes users' preference over ad placement.
Chinese Translation
将广告插入面向消费者的LLM输出正成为一种新兴的商业模式,但关于此类广告插入应如何评估或它如何影响用户偏好,目前缺乏共享的证据。我们提出了LLMAdBench,一个用于研究LLM生成内容中广告的人类偏好基准。该基准分离出一个简单但实际重要的决策:给定用户对话、LLM响应和匹配的广告,广告应放置在何处?我们的数据集比较了仅在广告位置上不同的响应对,同时保持所有其他条件固定,包括用户查询、基础答案、广告和披露条件。人类标注者从广告商和用户两个角度,基于六个标准对每一对进行评估。最终基准包含超过18000个人类判断,跨越两种披露条件:明确将广告标注为赞助内容,以及将其合并到响应中而不进行披露。我们使用LLMAdBench评估了八个前沿LLM作为偏好评判者,发现它们不是人类评估的可靠替代品。即使是最稳定的模型,在交换呈现顺序时也会反转大约四分之一的决策,模型间一致性低,并且它们的放置偏好与人类标注者有系统性差异。此外,LLMAdBench包含大量可学习信号。特别地,在人类偏好上微调的Qwen3-8B模型相比其基础模型有显著提升,并在留出预测任务上优于所有零样本前沿评判者。除了模型评估,LLMAdBench还提供了关于广告商-用户权衡的定量证据,并表明赞助披露系统地改变了用户对广告放置的偏好。
cs.AI / 93 / 2609.32539

Interpretable Physics Informed WiFi Indoor Localization: Learning an Effective Access Point Geometry and Using It to Prune

可解释的物理信息WiFi室内定位:学习有效的接入点几何结构并用于剪枝
zadeh, Arshia Eftekhari, Nasiri, Rezvan, Moradi, Hadi
Abstract
Deep learning models can achieve high accuracy for indoor localization, but their black-box nature limits interpretability and the reuse of learned information. We propose a hierarchical deep learning framework for WiFi fingerprint-based indoor localization that jointly predicts user location and learns an effective geometry of the surrounding access points (APs). Physics-informed decoders infer this geometry directly from RSSI measurements and labelled user positions, without requiring the true AP coordinates during training. The learned geometry is then used to rank and prune APs. On the UJIIndoorLoc dataset, the proposed chained model achieves a mean 3D localization error of 7.07 m, reducing error by 26% to 36% compared with baseline models. Previously published methods evaluated on the same official split report errors 10.6% to 31.0% higher. Pruning 35% or 50% of the APs causes only a small loss in localization accuracy. The inferred geometry also enables Fisher-information-based AP ranking even when fingerprint databases do not contain surveyed AP coordinates. Experiments on the Tampere/TUT and UTSIndoorLoc datasets show that geometry-guided AP selection performs comparably to selectors built directly from labelled data. These results show that physics-informed interpretability can improve indoor localization while also supporting effective feature selection.
Chinese Translation
深度学习模型可以在室内定位中达到高精度,但其黑箱性质限制了可解释性和学习信息的重用。我们提出了一种用于基于WiFi指纹的室内定位的分层深度学习框架,该框架联合预测用户位置并学习周围接入点(AP)的有效几何结构。物理信息解码器直接从RSSI测量和标记的用户位置推断该几何结构,而无需在训练期间获得真实的AP坐标。然后,学习到的几何结构用于对AP进行排序和剪枝。在UJIIndoorLoc数据集上,所提出的链式模型实现了7.07米的平均3D定位误差,与基线模型相比,误差降低了26%至36%。先前发表的方法在同一官方划分上评估报告的错误高出10.6%至31.0%。剪枝35%或50%的AP仅导致定位精度的小幅损失。即使指纹数据库不包含测量的AP坐标,推断的几何结构也能够实现基于Fisher信息的AP排序。在Tampere/TUT和UTSIndoorLoc数据集上的实验表明,几何引导的AP选择与直接从标记数据构建的选择器性能相当。这些结果表明,物理信息可解释性可以提高室内定位,同时支持有效的特征选择。
cs.AI / 94 / 2609.32544

Porimon: An LLM-Based Pok\'emon Battle Agent Enhanced by Long/Short-Term Knowledge Augmented Generation

Porimon:一种基于LLM的宝可梦对战智能体,通过长/短期知识增强生成进行增强
Zhuo, Dongyin, Pan, Fengjunjie, Petrovic, Nenad, Knoll, Alois
Abstract
In this paper, we use Pok\'emon Battles as a case study to investigate how to improve the performance of LLM-based agents in tasks that require opponent-aware planning without additional fine-tuning. We propose Long/Short-Term Knowledge Augmented Generation (LSTKAG), a mechanism that enables LLM-based agents to leverage past states of the current task and retrieve experience summaries from similar previous task instances based on the current state. Based on LSTKAG, we design Porimon, an LLM-based agent structure for Pok\'emon Battles. For optimization, we introduce an external API for precise damage calculation and more detailed information about the game. We conduct tournament-like evaluation experiments comprising 15,000 battles for hyperparameter optimization, ablation studies, and performance evaluation. The results indicate that Porimon-based players with hyperparameter optimization significantly outperform players based on Pok\'eLLMon, an LLM-based agent structure proposed in previous research, and the rule-based heuristic player. Furthermore, our ablation study shows that Porimon variants outperform the one without extension in game information retrieval, which shows the contribution of that extension. However, the current experiment results are inconclusive regarding the contribution of Long-Term KAG. These results suggest that introducing external resources, information from previous states of the current task, and experience summaries from similar previous task instances could elevate the performance of LLM-based agents designed for tasks requiring opponent-aware planning.
Chinese Translation
在本文中,我们以宝可梦对战为案例研究,探讨如何在不进行额外微调的情况下,提高基于LLM的智能体在需要对手感知规划的任务中的性能。我们提出了长/短期知识增强生成(LSTKAG),一种使基于LLM的智能体能够利用当前任务的过去状态,并基于当前状态从相似的先前任务实例中检索经验总结的机制。基于LSTKAG,我们设计了Porimon,一种用于宝可梦对战的基于LLM的智能体结构。为了优化,我们引入了一个外部API,用于精确的伤害计算和更详细的游戏信息。我们进行了类似锦标赛的评估实验,包含15,000场对战,用于超参数优化、消融研究和性能评估。结果表明,经过超参数优化的基于Porimon的玩家显著优于基于PokéLLMon(先前研究中提出的一种基于LLM的智能体结构)的玩家以及基于规则的启发式玩家。此外,我们的消融研究表明,Porimon的变体优于在游戏信息检索方面没有扩展的版本,这显示了该扩展的贡献。然而,当前实验结果对于长期KAG的贡献尚无定论。这些结果表明,引入外部资源、当前任务先前状态的信息以及来自相似先前任务实例的经验总结,可以提升为需要对手感知规划的任务设计的基于LLM的智能体的性能。
cs.AI / 95 / 2609.32550

Are Vision-Language-Action Models Robust to One-Step Observation Perturbations?

视觉-语言-动作模型对单步观测扰动鲁棒吗?
Yamabe, Shojiro, Sakuma, Jun
Abstract
Understanding the safety risks of vision-language-action (VLA) models is essential for their deployment in the physical world. Existing safety research has mainly considered persistent perturbations that are applied continuously to observations throughout an episode. However, momentary observation corruption, in which observations are severely perturbed only briefly within an episode, remains an underexplored safety threat. To address this gap, this work investigates robustness to one-step perturbations applied at a single time step per episode. Our experiments reveal that these perturbations substantially degrade VLA performance and that their impact depends on the action chunk execution length. Based on them, we propose CARE, which dynamically selects the execution length based on consistency with the previously predicted action chunk. CARE improves robustness with low computational overhead while preserving clean performance by selecting shorter execution lengths only under perturbations.
Chinese Translation
理解视觉-语言-动作(VLA)模型的安全风险,对其在物理世界中的部署至关重要。现有安全性研究主要考虑在一条 episode 中持续施加于观测的持续扰动。然而,瞬时观测损坏——即观测在一个 episode 内仅短暂地受到严重扰动——仍是一种尚未被充分探索的安全威胁。为弥补这一空白,本文研究了在每个 episode 的单个时间步施加单步扰动时的鲁棒性。我们的实验表明,这些扰动会显著降低 VLA 的性能,且其影响取决于动作块执行长度。基于此,我们提出 CARE,它根据与先前预测的动作块的一致性动态选择执行长度。CARE 通过仅在受到扰动时选择更短的执行长度,在低计算开销下提升鲁棒性,同时保持无扰动条件下的性能。
cs.AI / 96 / 2609.32562

Artificial intelligences and human scientists exhibit complementary strengths in theory building

人工智能与人类科学家在理论构建中展现出互补优势
Li, Ke, Zoumpoulis, Spyros I., Puranam, Phanish, Parker, Philip, Eshbaugh-Soha, Matthew, Gainsburg, Izzy, Gilead, Michael, Grossmann, Igor, Hadar, Britt, Inbar, Yoel, Simchon, Almog, Willer, Robb, Ai, Rui, Ao, Ruicheng, Bala, Gavin J., Bidwell, Matthew, Cai, Shuang, Chang, Kai, Chen, Skyler Y., Clark, Cory J., Dai, Irmak, Dalal, Abhinandan, Douglas, Connor, Du, Alexis, Du, Zhehang, Feng, Leyun, Fernandez-Mateo, Isabel, Gandhi, Linnea, Grumbach, Cyrille, Gupta, Anmol, Gupta, Vansh, Hademer, Maria, Hardy III, Jay H., Huang, Chen Kai, Jin, Jacob Xiangyu, Keskin, Ufuk, Kim, Na Hyun, Kobaş, Mert, Koh, Byounghoon, Lamont-Dobbin, Gabrielle, Lanzalotto, Gregory, Lee, Sun Young, Leng, Dingzhe, Li, Chenjun, Li, Weiyuan, Li, Zeyuan, Liang, Zhongyuan, Liu, Ning, Liu, Peihong, Liu, Yuhan, Lu, Jiuyao, Ma, Wanteng, Martinet, Nicolas, Mulat, Natnael, Nguyen, Christina A., Nguyen, Khai, Nguyen, Quang Minh, Pape, Naja, Park, Chanwoo, Poulidis, Stefanos, Sanchez-Burks, Jeffrey, Schaerer, Michael, Solal, Isabelle, Song, Yanbo, Sun, Junghyo, Sun, Qingyao, Sun, Rui, Swaab, Roderick, Tan, Kevin, Teng, Dequn, Vaccaro, Michelle A., Vigerbaeck, Robin, Wang, Xiaomeng, Yao, Randol H., Yilmaz, Duygu, Yiu, Shun, Yucesoy, Ecem, Zang, Allen, Zhang, Ruijia, Zhang, Xilan, Zhang, Yichi, Zhang, Zhanhao, Uhlmann, Eric Luis
Abstract
We investigate the effectiveness of artificial intelligences (AI)-specifically large language models (LLMs)-relative to human scientists at high-level cognitive tasks in social science such as theory formulation, predictions of novel empirical results, and theory revision in response to new evidence. The research domain was academic discourse regarding gender and race inequality. Our findings, comparing 25 LLMs with 13 senior researchers and 60 doctoral scholars, reveal that the AIs outperformed most humans individually on most of the present tasks, while human theories were more diverse and exhibited greater gains in predictive accuracy from aggregation. AI-generated theories were more extensively elaborated, involving additional theoretical paths and latent variables, and were rated as higher quality than human theories by independent raters blinded to source. However, this theoretical complexity was in part ornamental, in that it was not associated with more accurate predictions about empirical patterns in data; in contrast, human scientists achieved greater predictive efficiency with simpler theories. The AIs were significantly more likely than human scientists to revise their theories to incorporate new evidence; human scientists updated their beliefs in a selective way that is sensitive to prior prediction errors. We speculate that the superior processing capacity of artificial intelligences makes them especially well-suited to tasks requiring grappling with complexity, but that the greater diversity of human ideas is essential to wise crowds and collective creativity.
Chinese Translation
我们探究了人工智能(AI)——特别是大语言模型(LLMs)——相对于人类科学家在社会科学高层次认知任务中的有效性,这些任务包括理论构建、对新实证结果的预测以及根据新证据进行理论修正。研究领域为关于性别和种族不平等的学术讨论。我们的研究比较了25个LLMs与13位资深研究人员和60位博士学者,发现人工智能在大多数当前任务上的个体表现优于大多数人类,而人类理论更加多样,并且通过聚合获得了更大的预测准确性提升。人工智能生成的理论阐述更为详尽,涉及更多理论路径和潜在变量,并且在对来源不知情的独立评分者评定中,其质量被评高于人类理论。然而,这种理论复杂性部分上是装饰性的,因为它与对数据中经验模式的更准确预测无关;相比之下,人类科学家用更简单的理论实现了更高的预测效率。人工智能比人类科学家更可能修正其理论以纳入新证据;人类科学家则以一种对先前预测误差敏感的选择性方式更新其信念。我们推测,人工智能卓越的处理能力使其特别适合需要应对复杂性的任务,但人类思想的更大多样性对于智慧群体和集体创造力至关重要。
cs.AI / 97 / 2609.32564

ProTTT: Learning to Learn Semantic User Memory with Test-Time Training

ProTTT:通过测试时训练学习语义用户记忆的元学习
Park, Sejun, Bhang, Hyoungjo, Jeong, Hyein, Jo, Yohan
Abstract
Personalization requires language models to capture user-specific knowledge from a growing user history. Existing context-based approaches incur increasing inference costs as user history accumulates and rely on separate retrieval or summarization stages, while parametric-based approaches often require reconstructing user representations when new user data is added. We introduce ProTTT, a profile-supervised meta-learning framework for learning semantic user memory. The memory construction starts from a shared initialization and is updated for each user through test-time training on user history, allowing it to evolve continuously as the history grows. However, since test-time training alone does not explicitly encourage the memory to capture semantic user knowledge necessary for personalization, we learn this shared initialization using textual user profiles as supervision, so that test-time training on user history captures semantic knowledge more effectively. ProTTT consistently outperforms both full history ICL and all parametric baselines across diverse benchmarks, while substantially reducing inference cost by compressing user history into a lightweight parameterized memory. Our analysis also shows that profile supervision is a reliable objective for learning semantic user knowledge and that the resulting memory can track and retain evolving user preferences, while remaining robust across different history sizes. Overall, we demonstrate the effectiveness of test-time training for personalization and establish ProTTT as a baseline for continuously evolving user memory.
Chinese Translation
个性化要求语言模型从不断增长的用户历史中捕获用户特定的知识。现有的基于上下文的方法随着用户历史的积累会导致推理成本增加,并且依赖于单独的检索或摘要阶段;而基于参数的方法在添加新用户数据时通常需要重建用户表示。我们提出了 ProTTT,一个用于学习语义用户记忆的、由用户画像监督的元学习框架。记忆构建从一个共享初始化开始,并通过在用户历史上的测试时训练为每个用户更新,使其能够随着历史的增长而持续演化。然而,由于仅靠测试时训练并不能明确地鼓励记忆捕获个性化所需的语义用户知识,我们使用文本用户画像作为监督来学习这个共享初始化,从而使得在用户历史上的测试时训练能够更有效地捕获语义知识。ProTTT 在各种基准测试中始终优于完整历史 ICL 和所有参数化基线,同时通过将用户历史压缩为轻量级参数化记忆,大幅降低了推理成本。我们的分析还表明,用户画像监督是学习语义用户知识的一个可靠目标,并且所得到的记忆能够跟踪和保留不断演变的用户偏好,同时在不同历史大小下保持鲁棒性。总体而言,我们证明了测试时训练在个性化方面的有效性,并将 ProTTT 确立为持续演化的用户记忆的基线。
cs.AI / 98 / 2609.32574

CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations

CUE-Mem:通过多模态对话中的隐式线索基准测试长期用户记忆
Hu, Yulin, Zhao, Yanyan, Long, Zimo, Fu, Xing, Ji, Mengtong, Zhao, Weixiang, Hou, Yutai, Wang, Qianchao, Tu, Dandan
Abstract
Long-term memory is essential for multimodal agents that interact with users across sustained conversations. However, user memories are not always explicitly stated: they may also be implied by recurring background objects in images, ambient sounds in audio, or other peripheral multimodal cues. Existing benchmarks largely focus on text-only memory or explicit multimodal evidence, leaving implicit multimodal cues underexplored. We introduce CUE-Mem, a text-image-audio benchmark for evaluating long-term user memory from implicit cues. CUE-Mem contains 2,674 questions across explicit and implicit evidence settings and covers four tasks: Entity Recall, Long Pattern, Personalized Recommendation, and Answer Refusal. Across textualized memory systems, implicit performance remains far below oracle evidence, locating the main bottleneck in preserving and retrieving subtle cues rather than question answerability. Increasing caption detail recovers more of this evidence, but brings uneven gains and rapidly growing token costs, motivating native multimodal access. Yet native access does not uniformly resolve the bottleneck: evidence use depends strongly on the backbone, while multimodal indexing introduces substantial retrieval noise. CUE-Mem provides a testbed for memory systems that selectively retain, retrieve, and use subtle multimodal evidence.
Chinese Translation
长期记忆对于在持续对话中与用户交互的多模态智能体至关重要。然而,用户记忆并不总是被明确陈述:它们也可能由图像中反复出现的背景物体、音频中的环境声音或其他外围多模态线索所暗示。现有的基准测试主要关注纯文本记忆或显式多模态证据,而隐式多模态线索尚未得到充分探索。我们提出了CUE-Mem,一个文本-图像-音频基准,用于从隐式线索评估长期用户记忆。CUE-Mem包含2,674个问题,涵盖显式和隐式证据设置,并覆盖四个任务:实体召回(Entity Recall)、长模式(Long Pattern)、个性化推荐(Personalized Recommendation)和答案拒绝(Answer Refusal)。在文本化记忆系统中,隐式性能仍远低于理想证据(oracle evidence),将主要瓶颈定位在保留和检索细微线索上,而非问题的可回答性。增加描述细节能恢复更多此类证据,但带来不均衡的增益和迅速增长的token成本,促使采用原生多模态访问。然而,原生访问并不能统一解决瓶颈:证据使用强烈依赖于骨干模型,而多模态索引引入了大量检索噪声。CUE-Mem为选择性保留、检索和使用细微多模态证据的记忆系统提供了一个测试平台。
cs.AI / 99 / 2609.32584

EMIR$^2$: Evolution-Aware Memory with Intent-Guided Multi-Round Retrieval

EMIR$^2$:演化感知记忆与意图引导的多轮检索
Liu, Jinlan, Sun, Hongliang, Wang, Yong, Zhang, Bolin, Sui, Dinabo, Chu, Dianhui, Tu, Zhiying
Abstract
Long-term memory enables large language model (LLM) agents to leverage historical interactions for future tasks. However, existing memory systems struggle to utilize continuously evolving historical information, as they often rely on static memory representations and single-round retrieval strategies, failing to track factual changes or integrate distributed evidence across long-term interactions. To address these challenges, we propose \textsc{EMIR}$^{2}$, an \textbf{E}volution-Aware \textbf{M}emory framework with \textbf{I}ntent-Guided Multi-\textbf{R}ound \textbf{R}etrieval, enabling LLM agents to maintain evolving historical knowledge and adaptively retrieve relevant evidence. Specifically, \textsc{EMIR}$^{2}$ constructs a State-Evolving Memory Graph (SEMG) that represents long-term memory as evolving knowledge states supported by temporal event trajectories and evidential associations. By maintaining semantic states through evidence-based updates, SEMG preserves historical evolution and enables evidence tracing under complex and conflicting scenarios. Building upon this, we introduce an intent-guided multi-round retrieval mechanism that iteratively identifies missing evidence and expands retrieval based on accumulated information. Experiments on LoCoMo and MemConflict demonstrate that \textsc{EMIR}$^{2}$ improves long-term memory utilization, dynamic and static conflict handling, and complex retrieval performance, achieving relative improvements of more than 12\% in certain categories. These results highlight the effectiveness of jointly modeling memory evolution and adaptive evidence acquisition for long-term agent interactions.
Chinese Translation
长期记忆使大语言模型(LLM)智能体能够利用历史交互来完成未来任务。然而,现有记忆系统难以利用持续演化的历史信息,因为它们通常依赖静态记忆表示和单轮检索策略,无法跟踪事实变化或整合长期交互中的分布式证据。为解决这些挑战,我们提出 EMIR$^2$,一个演化感知记忆框架,具有意图引导的多轮检索,使 LLM 智能体能够维护演化的历史知识并自适应地检索相关证据。具体而言,EMIR$^2$ 构建了一个状态演化记忆图(State-Evolving Memory Graph, SEMG),将长期记忆表示为演化的知识状态,并由时间事件轨迹和证据关联支持。通过基于证据的更新维护语义状态,SEMG 保留了历史演化,并能够在复杂和冲突场景下进行证据追踪。在此基础上,我们引入了一种意图引导的多轮检索机制,该机制迭代地识别缺失证据,并基于累积信息扩展检索。在 LoCoMo 和 MemConflict 上的实验表明,EMIR$^2$ 提高了长期记忆利用率、动态和静态冲突处理以及复杂检索性能,在某些类别中实现了超过 12% 的相对改进。这些结果凸显了联合建模记忆演化和自适应证据获取对于长期智能体交互的有效性。
cs.AI / 100 / 2609.32594

MA-FPPO: Multi-Agent Flow-Pretrained Policy Optimization

MA-FPPO:多智能体流预训练策略优化
Zou, Guowei, Chen, Haonan, Wang, Haitao, Zhang, Beiwen, Yan, Na, Wu, Hejun
Abstract
Multi-agent flow policies learn cooperative behavior from fixed offline datasets, but often struggle to complete tasks in situations not covered by the offline data. In these situations, agents must both adapt to changes in the environment and coordinate with one another, yet action patterns learned offline are often insufficient for effective adaptation and coordination. To address this problem, we propose Multi-Agent Flow-Pretrained Policy Optimization (MA-FPPO), which uses online fine-tuning to improve the cooperative behavior of models pretrained with flow matching through new interactions with the environment. Building on the behavior learned during pretraining, we construct policies with explicit action likelihoods for discrete and continuous action spaces. We then update the pretrained model using shared team advantages to further improve coordination based on team performance. Our method achieves, on average, relative gains of 52.8% over the strongest listed offline baselines across 30 settings and 29.8% over purely online learning across 38 comparisons with matched online budgets and evaluation protocols.
Chinese Translation
多智能体流策略从固定的离线数据集中学习协作行为,但在离线数据未覆盖的情况下往往难以完成任务。在这些情况下,智能体既需要适应环境变化,又需要相互协调,然而离线学习到的动作模式往往不足以实现有效的适应和协调。为了解决这个问题,我们提出了多智能体流预训练策略优化(MA-FPPO),它通过与环境的新交互,利用在线微调来改进使用流匹配预训练的模型的协作行为。基于预训练期间学习到的行为,我们为离散和连续动作空间构建了具有显式动作似然的策略。然后,我们使用共享的团队优势来更新预训练模型,以基于团队表现进一步提高协调性。我们的方法在30个设置上平均比最强的已列出的离线基线相对提高了52.8%,在38个具有匹配在线预算和评估协议的比较中比纯在线学习相对提高了29.8%。
cs.AI / 101 / 2609.32600

CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

CUA-SWE:当计算机使用智能体遇上视觉软件工程
Wang, Prince Zizhuang, Liang, Chenhao, Xu, Zelong, Yuan, Aojie, Zhou, Xiaolin, Zhang, Haiyue, Zhao, Yue, Hu, Xiyang, Jiang, Shuli
Abstract
Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, leaving this integrated development process underexplored. Diagnosing a runtime interaction failure requires agents to connect visual observations with the responsible code, then use the application again to verify the repair. We introduce CUA-SWE, a benchmark, environment, and evaluation pipeline for software engineering with computer use. Beyond studying how GUI feedback supports diagnosis and repair, we ask whether agents can complete software engineering tasks when required specification or operational information is available only through the running application's visual interface. CUA-SWE spans four software engineering domains and requires agents to modify code and configuration, execute commands, interact with running software, and inspect visual feedback within the same task. Each task includes deterministic, task-specific tests that verify whether the resulting software satisfies the requirements and preserves specified behavior. Our evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes. We examine performance across domains and task information requirements, alongside the development behaviors associated with successful repairs. CUA-SWE provides a unified testbed for studying how agents use visual feedback and interaction to guide software engineering, with executable correctness criteria for the resulting software.
Chinese Translation
软件开发不仅仅需要编辑代码:开发者会反复运行软件、与其界面交互、目视检查其行为,并利用这些观察来决定接下来更改什么以及某项更改是否有效。现有的编码智能体和计算机使用智能体大多被孤立地研究,导致这一集成开发过程尚未得到充分探索。诊断运行时交互故障要求智能体将视觉观察与相关代码联系起来,然后再次使用应用程序以验证修复。我们提出 CUA-SWE,一个用于结合计算机使用的软件工程的基准、环境和评估流水线。除了研究 GUI 反馈如何支持诊断与修复外,我们还探究:当所需规范或操作信息只能通过运行中应用程序的视觉界面获得时,智能体能否完成软件工程任务。CUA-SWE 涵盖四个软件工程领域,并要求智能体在同一任务中修改代码和配置、执行命令、与运行中的软件交互,并检查视觉反馈。每个任务都包含确定性的、任务特定的测试,用于验证所得软件是否满足需求并保持指定的行为。我们的评估刻画了前沿智能体如何将源代码级执行与应用截图和图形交互结合起来,以产生经过验证的软件更改。我们考察跨领域和任务信息需求的性能,以及成功修复相关的开发行为。CUA-SWE 提供了一个统一的测试平台,用于研究智能体如何利用视觉反馈和交互来指导软件工程,并为所得软件提供可执行的正确性标准。
cs.AI / 102 / 2609.32616

"You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused

“你说得对,让我来修复”:LLM智能体在被错误指控时如何破坏正确的工作
Mao, Xutao, Qian, Rui, Wang, Longxiang, Yi, Xinjian, Li, Mingxuan, Chen, Linghan, Gao, Yudong, Zheng, Xiang, Wang, Cong
Abstract
LLM agents increasingly keep working after a task succeeds as they resume after compaction or take over handoffs. Their finished work keeps receiving follow-up input that sometimes falsely accuses it for later failures. We call an agent's acceptance of such a false accusation gaslight sycophancy, and destructive over-correction when acting on it damages previously correct work. We introduce CAVE-Bench, a benchmark of 365 agentic tasks across six domains built around opaque tasks. Every scored run first reaches a verified correct state, whose supporting rationale and history stay in the workspace while the facts that would settle the accusation lie in external or runtime state beyond the agent's reach. The agent cannot confirm or refute the claim with a local check, so the right response should keep the work and ask for the missing evidence. Each task either hands the agent correct work with saved evidence or let it build and verify that work first, and five risk factors set how the accusation enters the workflow. We score accusation acceptance and evidence use from the trajectory and measure harm by deterministic replay of downstream events. Across 14 of the latest models in Claude Code, false accusations damage correct work in up to 60.06% of runs, and stronger models often do so after recovering the supporting evidence. The same model behaves differently across OpenCode, Codex, and Hermes, and a harness gate driven by the benchmark's live signals cuts replayed harm by 74%. These results show that preserving already-correct work under unsupported accusation is a distinct safety challenge for long-lived agents. Our project is in https://henrymao2004.github.io/agent-over-correction/.
Chinese Translation
LLM智能体在任务成功后越来越多地继续工作,因为它们会在压缩后恢复或接手交接任务。它们已完成的工作会不断收到后续输入,这些输入有时会错误地将其归咎于后来的失败。我们将智能体接受这种错误指控称为煤气灯式谄媚(gaslight sycophancy),而当其据此行动并损害之前正确的工作时,称为破坏性过度纠正。我们介绍了CAVE-Bench,一个包含365个智能体任务的基准,涵盖六个领域,围绕不透明任务构建。每个被评分的运行首先达到一个经过验证的正确状态,其支持依据和历史保留在工作区中,而能解决指控的事实则位于智能体无法触及的外部或运行时状态。智能体无法通过本地检查来确认或反驳该说法,因此正确的回应应该是保留工作并索要缺失的证据。每个任务要么将带有保存证据的正确工作交给智能体,要么让它首先构建并验证该工作,而五个风险因素决定了指控如何进入工作流程。我们从轨迹中评估指控接受和证据使用情况,并通过下游事件的确定性重放来衡量危害。在Claude Code中的14个最新模型中,错误指控在高达60.06%的运行中损害了正确工作,而更强的模型往往在恢复支持证据后仍会这样做。同一个模型在OpenCode、Codex和Hermes中表现不同,而由基准的实时信号驱动的测试框架门控将重放危害降低了74%。这些结果表明,在无支持的指控下保留已经正确的工作是长期存活的智能体面临的一个独特安全挑战。我们的项目见 https://henrymao2004.github.io/agent-over-correction/。
cs.AI / 103 / 2609.32631

SWE-MILE: Asynchronous Potential-Induced Milestone Credit Assignment for Long-Horizon Software Engineering Agents

SWE-MILE:面向长时程软件工程智能体的异步势能诱导里程碑信用分配
Cui, Chaoqun, Zhou, Hao, Chen, Meiqi, Meng, Fandong, Mao, Wenji
Abstract
Long-horizon software engineering (SWE) agents trained with reinforcement learning with verifiable rewards (RLVR) typically receive only terminal outcome supervision, making it difficult to distinguish productive actions from redundant exploration or functional regressions. We propose SWE-MILE, an asynchronous potential-induced milestone credit assignment framework that derives fine-grained process supervision from workflow runtime, without auxiliary reward models or external evaluators. SWE-MILE quantifies task-relevant file exposure and test-state alignment as navigation and verification potentials, respectively. Differences in these potentials attribute milestone progress and regressions to individual actions, while discounted backward credit propagates supervision to preceding steps. To efficiently acquire intermediate verification states, SWE-MILE further introduces asynchronous shadow probing, which replays repository-changing actions in an isolated sandbox and runs verification in parallel with the agent's primary interaction, largely hiding verification latency. The resulting process credit augments terminal outcome advantages and provides informative learning signals. Experiments on two representative long-horizon SWE tasks demonstrate substantial improvements in agent performance, highlighting workflow runtime signals as a practical source of process supervision for long-horizon SWE agents.
Chinese Translation
使用可验证奖励强化学习(RLVR)训练的长时程软件工程(SWE)智能体通常仅接收终端结果监督,难以区分高效用的动作与冗余探索或功能回归。我们提出 SWE-MILE,一种异步势能诱导里程碑信用分配框架,其从工作流运行时中获取细粒度的过程监督,而无需辅助奖励模型或外部评估器。SWE-MILE 分别将任务相关的文件暴露和测试状态对齐量化为导航势能和验证势能。这些势能的差异将里程碑进展和回归归因于单个动作,而折扣反向信用将监督传播到先前步骤。为了高效获取中间验证状态,SWE-MILE 进一步引入异步影子探测,其在隔离沙箱中重放改变仓库的动作,并与智能体的主要交互并行运行验证,从而在很大程度上隐藏验证延迟。由此产生的过程信用增强了终端结果优势,并提供了信息丰富的学习信号。在两个代表性长时程 SWE 任务上的实验表明,智能体性能得到了显著提升,凸显了工作流运行时信号作为长时程 SWE 智能体过程监督的实用来源。
cs.AI / 104 / 2609.32638

Can Open-Weight Large Language Models (LLMs) Simulate Human Survey Populations? A Cross-Instrument Calibration Study

开放权重的大语言模型(LLMs)能否模拟人类调查群体?一项跨工具校准研究
Lee, Grandee, Yue, Wang
Abstract
Large language models (LLMs) are increasingly used to generate synthetic survey respondents and digital twins of real people, but whether their output preserves real human statistical structure, rather than surface plausibility, remains unresolved, and most existing evidence comes from proprietary models rather than open-weight ones. We evaluate three open-weight LLM families on a cross-instrument calibration task: conditioning personas on real respondents' verbatim answers to one psychometric instrument and measuring them on a second, construct-distance-controlled instrument, checked against a 2,058-person human panel. Across a 139-pair grid, the simulated cross-instrument correlation tracks the real human correlation at r = 0.70 - 0.73 in every model, driven mainly by correct sign rather than precise magnitude and concentrated in pairs of moderate construct distance. A correlation of this magnitude, obtained from untuned open-weight models conditioned only on individual-level survey data, is a substantively encouraging result for LLM-based behavioral simulation and digital-twin applications: specific model families and releases already reproduce a meaningful share of real human cross-instrument structure without any fine-tuning. This capability does not, however, improve monotonically across model releases: on a matched panel, the newest of three tested Llama releases performs worst on two of three headline metrics, so realizing its promise in practice requires release-specific, distance-aware verification rather than a one-time benchmark.
Chinese Translation
大语言模型(LLMs)越来越多地被用于生成合成调查受访者和真实人群的数字孪生,但其输出是否保留了真实人类的统计结构,而非表面上的合理性,仍未有定论,且现有证据大多来自专有模型而非开放权重模型。我们在一个跨工具校准任务中评估了三个开放权重LLM系列:将角色(personas)条件化于真实受访者对某一心理测量工具的逐字回答,并在第二个经过构念距离控制的工具上对其进行测量,并与一个包含2058人的人类样本组进行对照检验。在一个包含139对工具的网格中,每个模型模拟的跨工具相关性与真实人类相关性之间的相关系数达到r = 0.70 - 0.73,这主要源于符号的正确性而非精确的幅度,并且集中在中等构念距离的工具对上。这种量级的相关性,是从仅基于个体层面调查数据条件化的未经调优的开放权重模型中获得的,对于基于LLM的行为模拟和数字孪生应用而言是一个实质上令人鼓舞的结果:特定的模型系列和版本已经能够在无需任何微调的情况下,重现真实人类跨工具结构的相当大一部分。然而,这种能力并未随着模型版本的更迭而单调提升:在一个匹配的样本组上,三个测试的Llama版本中最新版本在三个主要指标中的两个上表现最差,因此要在实践中实现其潜力,需要针对特定版本、具有距离感知的验证,而非一次性的基准测试。
cs.AI / 105 / 2609.32643

Business Compromise Detection with Agentic AI and LLM-driven Knowledge Discovery

基于智能体AI和LLM驱动知识发现的商业失陷检测
Palma, Diego, Kim, Kyu Bin, Han, Zhen, Dsouza, Allbright, Liu, Zhiyuan
Abstract
Detecting compromised business ad accounts is a challenge in digital advertising, as attackers exploit hijacked accounts to launch fraudulent campaigns. Large Language Model (LLM) agents show promise for integrity enforcement, but hallucinated mistakes on hard cases create business friction. In a study we find the autonomous agent is a strong, recall-heavy signal extractor but an unreliable final arbiter, conceding precision on ambiguous decisions. We therefore keep the agent as an investigator that emits a structured, interpretable signal vector, and delegate the verdict to a neuro-symbolic stage: symbolic rules discovered by Inductive Logic Programming (FOIL-IE), a Na\"ive Bayes calibration layer, and a data-tuned contradiction layer. Evaluating on a compromise-over-sampled population and a realistic low-prevalence sample with subject-matter-expert labels, this arbiter substitution raises MCC from 0.295 to 0.435 ({\Delta}MCC +0.139, 95% CI [+0.026, +0.245], p=0.018, paired bootstrap), lifting precision from 0.250 to 0.446 (1.8x) at a recall cost (0.920 to 0.660). Benchmarked under identical conditions, it also edge tree ensembles (0.386).The rules encode domain w labels while remaininginterpretable and auditable.
Chinese Translation
检测失陷的商业广告账户是数字广告中的一个挑战,因为攻击者利用被劫持的账户发起欺诈性广告活动。大型语言模型(LLM)智能体在完整性执法方面展现出潜力,但在困难案例上的幻觉错误会造成业务摩擦。在一项研究中,我们发现自主智能体是一个强大的、偏重召回的信号提取器,但不是一个可靠的最终裁决者,在模糊决策上牺牲了精确率。因此,我们将智能体保留为调查者,输出一个结构化、可解释的信号向量,并将裁决委托给一个神经符号阶段:由归纳逻辑编程(FOIL-IE)发现的符号规则、一个朴素贝叶斯校准层,以及一个数据调优的矛盾层。在失陷过采样总体和带有主题专家标签的现实低流行率样本上进行评估,这种裁决者替换将MCC从0.295提高到0.435(ΔMCC +0.139,95% CI [+0.026, +0.245],p=0.018,配对自助法),将精确率从0.250提高到0.446(1.8倍),但召回率有所下降(从0.920降至0.660)。在相同条件下进行基准测试,它也略胜于树集成模型(0.386)。这些规则编码了领域w标签,同时保持可解释性和可审计性。
cs.AI / 106 / 2609.32645

From Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous Driving

从场景图到答案:面向自动驾驶的选择性神经符号推理
Wang, Yiyao, Liu, Pei, Liu, Fangzhou, Ma, Jun
Abstract
Autonomous-driving question answering requires reasoning over structured scene information, yet existing vision-language approaches largely delegate heterogeneous reasoning operations to a single neural inference process. We argue that this uniform strategy overlooks a fundamental distinction: some queries admit exact symbolic solutions, while others require semantic interpretation. We introduce a query-adaptive neuro-symbolic reasoning framework that explicitly allocates computation according to the nature of the query. At its core is a hierarchical Spatiotemporal Scene Graph (STSG) that separates persistent object identities from frame-specific states and represents spatial relations and temporal transitions as explicit directed structures. Given a query, a symbolic executor first attempts to resolve it through exact graph operations; only when symbolic execution abstains is an LLM invoked for semantic reasoning. For these unresolved queries, query-conditioned graph retrieval and evidence filtering preserve relation direction, temporal locality, and object semantics, providing the LLM with compact and verified task-relevant evidence. This design shifts the role of the LLM from a universal reasoning engine to a targeted semantic reasoner, while allowing deterministic computation to be handled exactly and efficiently. We evaluate the framework on 5,916 NuScenes-QA questions across all ten scenes of nuScenes v1.0-mini under an oracle-perception setting. The complete system achieves 80.63 percent overall accuracy with GPT-5.4-mini, improving over the corresponding LLM-only configuration by 5.48 percentage points; with DeepSeek-V4-Flash, the improvement reaches 6.64 points. The largest gains occur on counting questions, with improvements of 10.20 and 12.61 points, respectively. These results show that selective reasoning improves both accuracy and inference efficiency.
Chinese Translation
自动驾驶问答需要对结构化场景信息进行推理,然而现有的视觉-语言方法在很大程度上将异构推理操作委托给单一的神经推理过程。我们认为,这种统一策略忽略了一个根本区别:一些查询允许精确的符号解,而另一些则需要语义解释。我们引入了一个查询自适应的神经符号推理框架,该框架根据查询的性质显式地分配计算。其核心是一个分层时空场景图(STSG),它将持久对象标识与帧特定状态分离,并将空间关系和时间转换表示为显式的有向结构。给定一个查询,符号执行器首先尝试通过精确的图操作来解决它;只有当符号执行弃权时,才调用LLM进行语义推理。对于这些未解决的查询,查询条件图检索和证据过滤保留了关系方向、时间局部性和对象语义,为LLM提供了紧凑且经过验证的任务相关证据。这种设计将LLM的角色从通用推理引擎转变为目标语义推理器,同时允许确定性计算被精确且高效地处理。我们在oracle感知设定下,对nuScenes v1.0-mini所有十个场景中的5,916个NuScenes-QA问题进行了评估。完整系统在使用GPT-5.4-mini时达到80.63%的总体准确率,比相应的仅LLM配置提高了5.48个百分点;使用DeepSeek-V4-Flash时,提升达到6.64个百分点。最大的增益出现在计数问题上,分别提高了10.20和12.61个百分点。这些结果表明,选择性推理提高了准确性和推理效率。
cs.AI / 107 / 2609.32652

Prediction Limits and Koopman Closure of Geometry-Induced Soft State Abstractions

几何诱导的软状态抽象的预测极限与Koopman闭包
Kumar, Mohit, Kargaran, Somayeh
Abstract
We study when geometry-induced soft state abstractions admit accurate finite-dimensional linear dynamics. Each state is represented by simplex-valued coordinates obtained from class-specific Kernel Affine Hull Machine (KAHM) reconstruction scores, and a matrix is used to predict the next-state coordinates. Our main result is a computable lower confidence bound on the minimum root-mean-square prediction error over all matrices satisfying a prescribed spectral-norm limit. The bound combines within-class variation of successor coordinates with the deviation of soft coordinates from their one-hot reference labels, and can be evaluated from independent state-successor pairs without fitting a prediction matrix. For fixed coordinates and evaluation distribution, the certificate converges almost surely to a population lower bound as the sample size grows; any tolerance below this limit is eventually certified unattainable. Reconstruction-score margins further control the soft-to-hard assignment error. Under deterministic dynamics and exact coordinate closure, eigenvectors of the closure matrix and its reduced transpose induce Koopman and adjoint Koopman eigenfunctions, respectively. A four-state KAHM construction shows that identical soft coordinates can permit exact closure under one dynamics map yet force positive prediction error under another. Experiments on Duffing, Van der Pol, CartPole, MountainCar, and Acrobot compare direct soft-coordinate prediction with state-space DMD/EDMD baselines and report prediction, representation-variation, and spectral diagnostics. The benchmarks assess fitted models but do not numerically evaluate the exclusion certificate.
Chinese Translation
我们研究几何诱导的软状态抽象何时允许准确的有限维线性动力学。每个状态由从特定类别的核仿射包机(KAHM)重建得分获得的单纯形值坐标表示,并使用矩阵预测下一状态坐标。我们的主要结果是关于所有满足规定谱范数限制的矩阵的最小均方根预测误差的可计算下置信界。该界结合了后继坐标的类内变化与软坐标相对于其one-hot参考标签的偏差,并且可以从独立的状态-后继对中评估,而无需拟合预测矩阵。对于固定的坐标和评估分布,随着样本量的增长,该证书几乎必然收敛到总体下界;任何低于此限度的容差最终被证明无法达到。重建得分裕度进一步控制软到硬分配的误差。在确定性动力学和精确坐标闭包下,闭包矩阵及其约化转置的特征向量分别诱导Koopman和伴随Koopman特征函数。一个四状态KAHM构造表明,相同的软坐标可以在一个动力学映射下允许精确闭包,而在另一个动力学映射下强制产生正预测误差。在Duffing、Van der Pol、CartPole、MountainCar和Acrobot上的实验比较了直接软坐标预测与状态空间DMD/EDMD基线,并报告了预测、表示变化和谱诊断。这些基准测试评估了拟合模型,但没有数值评估排除证书。
cs.AI / 108 / 2609.32656

MixBench-TS: A Multivariate Time Series Forecasting Benchmark Where Channel Mixing Pays Off

MixBench-TS:一个通道混合发挥作用的多元时间序列预测基准
Abdelmalak, Ibram, Putzke, Mischa, Choi, Jungmin, Hanika, Tom, Yalavarthi, Vijaya Krishna, Schmidt-Thieme, Lars
Abstract
Multivariate Time Series Forecasting (MTSF) models that mix information across channels assume that the past of one channel carries information about the future of another. Yet they are evaluated on a small fixed set of standard datasets whose cross-channel structure is rarely examined. We ask two questions: "How can we reliably measure lagged, non-linear, and joint coupling in MTSF datasets?" and "Do the standard datasets actually have such coupling?" To answer the first, we test four candidate measures on synthetic datasets with planted ground-truth coupling: Granger Causality (GC), Transfer Entropy (TE), lagged Mutual Information (MI), and the CD gain, a model-based measure we introduce that compares a channel-dependent (CD) model to its channel-independent (CI) variant. Only lagged MI and the CD gain recover every planted coupling. For the second question, the answer is a definite no, as the standard datasets have a median of only 23% lagged-coupled channel pairs and a median CD gain of -4.9%, compared to 78% and +4.6% on chaotic ODE systems. We therefore propose MixBench-TS, a benchmark of 10 real-world datasets with a median of 55.5% lagged-coupled pairs and a median CD gain of +1.7%. Across six state-of-the-art models tuned under one protocol, CI models win on 10/10 (MSE) and 8/10 (MAE) standard datasets, but on only 3/10 and 2/10 MixBench-TS datasets. We recommend using our benchmark for evaluating new CD models. Moreover, we propose profiling new datasets with lagged MI and the CD gain before using them to evaluate multivariate models. Code and data are available at https://anonymous.4open.science/r/mixbench-ts-B027.
Chinese Translation
混合跨通道信息的多元时间序列预测(MTSF)模型假设一个通道的过去携带关于另一个通道未来的信息。然而,它们是在一小部分固定的标准数据集上评估的,这些数据集的跨通道结构很少被检查。我们提出两个问题:“我们如何可靠地测量MTSF数据集中的滞后、非线性和联合耦合?”以及“标准数据集确实具有这种耦合吗?”为了回答第一个问题,我们在具有植入的真实耦合的合成数据集上测试了四种候选度量:格兰杰因果(GC)、转移熵(TE)、滞后互信息(MI)以及CD增益,这是我们引入的一种基于模型的度量,它将通道依赖(CD)模型与其通道独立(CI)变体进行比较。只有滞后MI和CD增益能够恢复每一个植入的耦合。对于第二个问题,答案是否定的,因为标准数据集的滞后耦合通道对中位数仅为23%,CD增益中位数为-4.9%,而混沌ODE系统分别为78%和+4.6%。因此,我们提出了MixBench-TS,一个包含10个真实世界数据集的基准,其中滞后耦合对的中位数为55.5%,CD增益中位数为+1.7%。在单一协议下调整的六个最先进模型中,CI模型在10/10(MSE)和8/10(MAE)标准数据集上获胜,但在MixBench-TS数据集上仅分别为3/10和2/10。我们建议使用我们的基准来评估新的CD模型。此外,我们建议在使用新数据集评估多元模型之前,使用滞后MI和CD增益对新数据集进行剖析。代码和数据可在 https://anonymous.4open.science/r/mixbench-ts-B027 获取。
cs.AI / 109 / 2609.32657

World Models with Predictable Long-Horizon Marginals

具有可预测长时程边缘分布的世界模型
Du, Yuhao, Chen, Shunian
Abstract
Accurate one-step predictions do not ensure that a world model's rollouts retain the data distribution. We make the model's decoded stationary law explicit by learning a decoder of a fixed Gaussian reference and constraining the behaviour-averaged transition to preserve that reference. For controlled systems, a joint transition uses a conditional action chart to preserve behaviour occupancy without requiring invariance at each fixed action. Joint state--action rotations and parallel Gaussian noise give an exactly preserving transition with a tractable conditional density. We derive an absolute convergence bound from finite initialization banks and control departure from the reference through conditional action-space divergence. Across $216$ fitted pixel checkpoints on twelve control tasks, the occupancy model with a reference mixture retains every evaluated chain at $10^5$ steps in all $36$ task--seed cells, with a rollout-minus-reference energy-statistic difference of $-0.0002\pm0.0003$ (training-seed standard error). Each of the four nonpreserving comparison arms loses chains, although the Gaussian arm is more accurate at ten steps. An offline DreamerV3 reference also achieves better short-horizon accuracy. These results distinguish three properties of a world model: the distribution it approaches, the rate of approach, and the conditional dynamics it learns.
Chinese Translation
准确的一步预测并不能确保世界模型的展开保持数据分布。我们通过学习固定高斯参考的解码器,并约束行为平均的转移保持该参考,显式地表达模型解码后的平稳律。对于受控系统,联合转移使用条件动作图来保持行为占用,而不需要每个固定动作处的不变性。联合状态-动作旋转和平行高斯噪声给出了一个精确保持的转移,并具有可处理的条件密度。我们从有限初始化库推导出绝对收敛界,并通过条件动作空间散度控制与参考的偏离。在十二个控制任务的216个拟合像素检查点上,带有参考混合的占用模型在所有36个任务-种子单元中,使每个被评估的链在10^5步时都得以保持,其展开减参考的能量统计量差异为-0.0002\pm0.0003(训练种子标准误)。四个非保持对比臂每个都会丢失链,尽管高斯臂在十步时更准确。一个离线DreamerV3参考也实现了更好的短时程精度。这些结果区分了世界模型的三个属性:它所逼近的分布、逼近速率以及它所学习的条件动力学。
cs.AI / 110 / 2609.32658

Contract Memory Compiler: Resolve, Then Traverse

契约记忆编译器:先解析,再遍历
Song, Zhi, Xing, XiMing, Li, Chunhan, Mao, Weian, Tang, Zhenchao, Huang, Hanbo, Xu, Fan, Zhou, Jiale, Guan, Jiahui, Ding, Zejian, Ma, Chen, Wang, Lusheng
Abstract
External memory lets language-model agents answer questions about histories too long for the answer model's context window. Updates create a harder problem than retrieving a recent fact: changing one relation can redirect a multi-hop question to records about an entity absent from the question. We study this update-dependent evidence selection problem and introduce the Contract Memory Compiler (CMC). Before seeing a question, CMC uses a language model to identify relations in the history and record where each one was stated. It applies later updates to determine the current relations, follows them from entities named in the question, and passes the corresponding original records to the answer model in one call. Thus the current state determines which evidence is read, rather than merely refreshing values in a previously selected context. To the best of our knowledge, CMC achieves state-of-the-art multi-hop accuracy on FactConsolidation, reaching 78.25% overall and 61.0% at 262K. With the extracted relations and answer model held fixed, selecting evidence before resolving updates reduces multi-hop accuracy to 21.50%. We also introduce MQuAKE-MemStream, a derived dataset of ordered memory streams built from MQuAKE-Remastered counterfactual cases.
Chinese Translation
外部记忆使语言模型智能体能够回答关于历史记录的问题,而这些历史记录对于答案模型的上下文窗口来说过长。更新带来了比检索近期事实更难的问题:改变一个关系可能将多跳问题重定向到关于问题中未出现的实体的记录。我们研究了这种依赖于更新的证据选择问题,并提出了契约记忆编译器(Contract Memory Compiler, CMC)。在看到问题之前,CMC 使用语言模型识别历史记录中的关系,并记录每个关系在何处被陈述。它应用后续更新来确定当前关系,从问题中提到的实体出发沿着这些关系进行追踪,并将相应的原始记录在一次调用中传递给答案模型。因此,当前状态决定了读取哪些证据,而不仅仅是在先前选定的上下文中刷新值。据我们所知,CMC 在 FactConsolidation 上取得了最先进的多跳准确率,总体达到 78.25%,在 262K 时达到 61.0%。在提取的关系和答案模型保持固定的情况下,若在解析更新之前选择证据,多跳准确率会降至 21.50%。我们还提出了 MQuAKE-MemStream,这是一个从 MQuAKE-Remastered 反事实案例构建的有序记忆流衍生数据集。
cs.AI / 111 / 2609.32667

Learning from a Thoughtful Teacher: Adaptive On-Policy Self-Distillation for Mathematical Reasoning

从深思熟虑的教师中学习:面向数学推理的自适应同策略自蒸馏
Du, Jiacheng, Xie, Weiwei, Du, Tianyi, Guo, Shaoxiong, Ren, Qibing, Zhang, Jiaheng
Abstract
On-policy self-distillation (OPSD) trains a question-only student with token-level feedback from a teacher given training-only privileged information (PI). OPSD therefore provides dense, on-policy supervision, and is free of a larger external teacher, but its effectiveness rests on how PI is designed and utilized. Our preliminary diagnostics suggest a significant gap between teacher utility and student learnability, where a small fraction of high-disagreement tokens dominate the distillation signal, and short teacher continuations at these positions further expose more explicit PI leakage than transferable correction cues, indicating a strong intent on injecting PI-conditioned shortcuts. We propose Adaptive On-Policy Self-Distillation (AOPSD), which adapts what information the teacher receives and how strongly its feedback influences learning. AOPSD encodes each solution as a reasoning DAG, orders problems by the student's evolving capability, and reveals only the affordable subgraph and its next frontier as PI. For high-disagreement tokens, AOPSD utilizes short teacher continuations as probes to encourage useful guidance while mitigating PI-conditioned shortcuts among teacher supervisions. On HMMT25, AIME24, AIME25, and BRUMo25, AOPSD achieves 72.5% Pass@8, which is 6.7 percentage points above OPSD and 4.2 above the strongest competing baseline while reducing 15 percentage points of training time at lower cost.
Chinese Translation
同策略自蒸馏(OPSD)训练一个仅使用问题的学生模型,利用教师给出的仅训练时特权信息(PI)的 token 级反馈。因此,OPSD 提供了密集的、同策略的监督,且无需更大的外部教师,但其有效性取决于 PI 的设计和使用方式。我们的初步诊断表明,教师效用与学生可学习性之间存在显著差距,其中一小部分高分歧 token 主导了蒸馏信号,并且这些位置上的简短教师续写进一步暴露出比可迁移的纠正线索更明显的 PI 泄漏,表明存在强烈的注入 PI 条件捷径的意图。我们提出了自适应同策略自蒸馏(AOPSD),它调整教师接收的信息以及其反馈对学习的影响强度。AOPSD 将每个解决方案编码为推理 DAG,根据学生不断发展的能力对问题进行排序,并仅揭示可负担的子图及其下一前沿作为 PI。对于高分歧 token,AOPSD 利用简短的教师续写作为探针,以鼓励有用的指导,同时减轻教师监督中 PI 条件捷径的影响。在 HMMT25、AIME24、AIME25 和 BRUMo25 上,AOPSD 实现了 72.5% 的 Pass@8,比 OPSD 高出 6.7 个百分点,比最强的竞争基线高出 4.2 个百分点,同时以更低的成本减少了 15 个百分点的训练时间。
cs.AI / 112 / 2609.32669

Refinement Symmetry in Multimodal Transformers

多模态 Transformer 中的细化对称性
Du, Yuhao, Chen, Shunian
Abstract
Attention weights depend on token counts, which change with the representation of a signal. We study refinement symmetry: splitting a representation while preserving content, position, visible context, and total mass should preserve its contribution. Building on proportional and quadrature attention, we show that split invariance forces the local mass factor to be linear for any fixed positive attention kernel, provided that factor is nondecreasing. For changed representations, a physical coupling bounds attention error by separating feature change from weight reallocation. In Qwen2.5-Omni-7B, duplicating half the visual tokens threefold changes 255 of 3,586 MVBench answers under standard attention; measure weighting preserves every answer under matched visibility. Under natural frame resampling, it reduces distributional drift. At twofold merging of a frozen video encoding, a five-seed evaluation shows an all-partition-correct accuracy gain of 1.04 percentage points over global count weighting (average group mass) and 0.93 points over standard attention. The advantage over global count also holds on WorldSense but depends on the compression budget. The result is a representation principle with a measured benefit in robustness across partitions.
Chinese Translation
注意力权重取决于 token 数量,而 token 数量随信号的表示方式而变化。我们研究细化对称性:在保持内容、位置、可见上下文和总质量的同时拆分表示,应保持其贡献不变。基于比例注意力(proportional attention)和正交注意力(quadrature attention),我们证明:对于任何固定的正注意力核,若局部质量因子非递减,则拆分不变性迫使该因子为线性。对于变化的表示,一种物理耦合通过将特征变化与权重重新分配分离来约束注意力误差。在 Qwen2.5-Omni-7B 中,将一半视觉 token 复制三倍,在标准注意力下会改变 3,586 个 MVBench 答案中的 255 个;在匹配可见性下,测度加权(measure weighting)保留了每个答案。在自然帧重采样下,它减少了分布漂移。在冻结视频编码的两倍合并下,五种子评估显示,相对于全局计数加权(平均组质量),全分区正确准确率提升了 1.04 个百分点,相对于标准注意力提升了 0.93 个百分点。相对于全局计数的优势在 WorldSense 上也成立,但取决于压缩预算。其结果是一种表示原则,在跨分区的鲁棒性方面具有可测量的收益。
cs.AI / 113 / 2609.32670

What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation

什么会证伪它?关于自我解释的机制性主张的变量特异性证据标准
zadeh, Arshia Eftekhari
Abstract
When a language model explains an answer it has already given, does it reuse the computation that produced the answer or reconstruct a story from the answer alone? Attribution, transportability and recoverability are each compatible with causal use without establishing it. We propose an evidence standard: pair each positive statistic with a variable specific null that removes the tested variable's identity while matching relevant nuisance dimensions as far as possible, and audit unmatched dimensions. We apply this standard to a known cause. A cue naming a wrong option raises the rate of choosing that option by 64 to 68 percentage points across three models. Explanations mention the cue in 1.8 percent of items or fewer in three of four models tested. Three estimator classes yield favorable statistics, but none establishes causal sensitivity to the cue contrast under its own control in the three-model analysis. In the strongest case, a recovered cue direction reaches $R^2$ of 0.95 and exceeds a geometry matched random direction in all three seeds, while a direction fitted by the same pipeline with cue labels scrambled reproduces 61 to 76 percent of its effect at comparable realized edit magnitude. A fourth model passes one interchange endpoint, but unequal edit magnitudes and a contrast that changes both cue identity and cue-answer agreement limit its interpretation. These experiments leave causal access unresolved. They establish an evidentiary requirement: favorable mechanistic statistics must survive controls for variable identity and nuisance structure. Reusable controls separate generic from identity specific transport effects, fit null directions with scrambled labels, and audit realized intervention magnitudes.
Chinese Translation
当语言模型解释其已经给出的答案时,它是重用了产生该答案的计算,还是仅从答案重建了一个故事?归因、可迁移性和可恢复性均与因果使用相容,但均未确立因果使用。我们提出一种证据标准:将每个正面统计量与一个变量特异性零假设配对,该零假设去除被测变量的身份,同时尽可能匹配相关的干扰维度,并审计未匹配的维度。我们将此标准应用于一个已知原因。一个命名错误选项的线索使选择该选项的比率在三个模型中提高了64至68个百分点。在测试的四个模型中的三个中,解释在至多1.8%的项目中提及该线索。三类估计器产生了有利的统计量,但在三模型分析中,没有一个在其自身控制下确立对线索对比的因果敏感性。在最强的案例中,一个恢复的线索方向达到R²为0.95,并在所有三个种子中超过几何匹配的随机方向,而由同一流程用打乱的线索标签拟合的方向在可比的实现编辑幅度下重现了其效应的61%至76%。第四个模型通过了一个交换端点,但不相等的编辑幅度和一个同时改变线索身份和线索-答案一致性的对比限制了其解释。这些实验使因果访问悬而未决。它们确立了一项证据要求:有利的机制性统计量必须经受住对变量身份和干扰结构的控制。可重用的控制将通用传输效应与身份特异性传输效应分开,用打乱的标签拟合零方向,并审计实现的干预幅度。
cs.AI / 114 / 2609.32674

Expected Reasoning-Step Return Unifies On-Policy Learning from Rewards and Teachers

期望推理步回报统一了从奖励和教师进行的在策略学习
He, Qiangqiang, Li, Jin
Abstract
On-policy reasoning models can learn from task rewards or teacher signals, but these sources differ in form and can favor conflicting updates, leaving unclear which should guide a given reasoning action. We introduce \textbf{Expected Reasoning-Step Return (ERSR)}, which treats semantic reasoning steps as macro-actions and uses Monte Carlo student-policy rollouts to estimate the expected final task reward of student-generated and teacher-proposed actions in a common return space for step-level comparison. ERSR analysis reveals an outcome-dependent asymmetry: student actions are more beneficial than teacher replacements on successful trajectories, whereas teacher replacements become more beneficial on failed trajectories. We further show that student answer-probe gains track student-step ERSR utility and distinguish beneficial from harmful reasoning steps. Based on these findings, we propose \textbf{Return-Referenced On-Policy Learning (R$^2$OPL)}, which reinforces student reasoning on successful trajectories and distills teacher signals on failed ones, while using group success rate for difficulty scaling and student-probe gains for step-level modulation. Experiments across reasoning benchmarks and teacher--student configurations show that R$^2$OPL consistently outperforms strong baselines. ERSR training dynamics further show that R$^2$OPL jointly exploits substantial utility from both reward- and teacher-side signals, whereas existing hybrids often leave substantial residual utility in one branch.
Chinese Translation
在策略推理模型可以从任务奖励或教师信号中学习,但这些来源形式不同,可能倾向于相互冲突的更新,导致不清楚哪个应该指导给定的推理动作。我们引入了期望推理步回报(ERSR),它将语义推理步骤视为宏动作,并使用蒙特卡洛学生策略展开来估计学生生成和教师提出的动作在共同回报空间中的预期最终任务奖励,以进行步骤级比较。ERSR 分析揭示了一种结果依赖的不对称性:在成功轨迹上,学生动作比教师替换更有益,而在失败轨迹上,教师替换变得更有益。我们进一步表明,学生答案探测增益跟踪学生步骤的 ERSR 效用,并能区分有益和有害的推理步骤。基于这些发现,我们提出了回报参考在策略学习(R^2OPL),它在成功轨迹上强化学生推理,在失败轨迹上蒸馏教师信号,同时使用组成功率进行难度缩放,并使用学生探测增益进行步骤级调制。在推理基准和教师-学生配置上的实验表明,R^2OPL 始终优于强基线。ERSR 训练动态进一步表明,R^2OPL 联合利用了来自奖励侧和教师侧信号的显著效用,而现有的混合方法通常在一个分支中留下大量残余效用。
cs.AI / 115 / 2609.32677

When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds

当更好变得更糟:自适应世界中自我改进智能体的改进保真度
Wang, Ke, Zhao, Zijie, Yuan, Zhiyi, Li, Changlun
Abstract
Self-improving agents increasingly rely on proxy verifiers to choose policy updates, yet deployment can change the world in which those updates are evaluated. An update that looks better to the verifier can therefore become worse after deployment even when the verifier ranks policies well overall. We formalize this gap as Improvement Fidelity, which asks whether proxy improvements preserve the sign and ordering of deployment improvements over the updates an improvement process actually proposes. We show that global policy accuracy need not guarantee update fidelity: operator shift and deployment response can create update-level errors, while candidate margins determine whether those errors change the replacement decision. We introduce PIVOT-KG, a paired, decision-aware validator that allocates scarce high-fidelity evaluation according to the expected reduction in selection regret per unit cost. Across 90 held-out roots in Leduc, Kuhn, and Melting Pot, proxy and deployment optimal sets are disjoint in 51 cases. In an eight-candidate HighwayEnv stress test, PIVOT-KG reduces mean improvement-selection regret from 0.0435 under the exact Uniform validation rule to 0.0055 at the primary budget. Together, these results show why reliable self-improvement should evaluate proposed improvements in the worlds they induce, while providing a practical rule for allocating scarce deployment evidence when it can affect the replacement decision.
Chinese Translation
自我改进智能体日益依赖代理验证器来选择策略更新,然而部署可能改变评估这些更新的世界。因此,一个对验证器看来更好的更新,在部署后可能变得更糟,即使验证器总体上对策略排序良好。我们将这一差距形式化为改进保真度(Improvement Fidelity),它追问代理改进是否在改进过程实际提出的更新上保持部署改进的符号和排序。我们表明,全局策略准确性不一定保证更新保真度:算子偏移和部署响应可能产生更新级别的错误,而候选边际决定这些错误是否会改变替换决策。我们提出 PIVOT-KG,一种配对的、决策感知的验证器,它根据单位成本选择遗憾的预期减少来分配稀缺的高保真评估。在 Leduc、Kuhn 和 Melting Pot 的 90 个留存根中,代理和部署最优集在 51 个案例中不相交。在八候选 HighwayEnv 压力测试中,PIVOT-KG 在主要预算下将平均改进选择遗憾从精确 Uniform 验证规则下的 0.0435 降低到 0.0055。总之,这些结果说明了为什么可靠的自改进应该在其所诱导的世界中评估所提出的改进,同时提供了一种实用规则,用于在稀缺部署证据可能影响替换决策时进行分配。
cs.AI / 116 / 2609.32685

PINNMorph: Evolving Online Adaptation Policies for Physics-Informed Neural Networks

PINNMorph:演化物理信息神经网络的在线自适应策略
Yang, Xu, Yu, Mingyang, Zhang, Jun, Li, Keqian, Xu, Jing
Abstract
Physics-informed neural networks (PINNs) provide a learning-based framework for solving partial differential equations (PDEs), yet their training behavior can change substantially throughout optimization. Residual distributions, gradient interactions, regional learning difficulty, and model-capacity requirements may evolve over time, while the network architecture and major training mechanisms are typically determined before training. We propose PINNMorph, an online PINN adaptation framework based on large language model (LLM)-guided policy evolution. PINNMorph maintains a population of state-conditioned adaptation policies that map execution diagnostics to controlled interventions over topology modification, additive representation augmentation, objective balancing, gradient handling, adaptive sampling, and optimizer-phase control. At each intervention opportunity, candidate programs are instantiated from the current policy population, selected according to the observed training state, and applied directly to the PINN under training. The resulting model inherits its existing parameters and training state and continues optimization along the same trajectory. Execution outcomes are subsequently used to evaluate interventions and evolve the policy population. Unlike pre-training architecture search or fixed adaptation rules, PINNMorph jointly adapts the current PINN and the policies governing its interventions using feedback from actual training. Experiments on 13 PDE benchmarks show that PINNMorph achieves lower solution errors than SA-PINN, ConFIG, RoPINN, HARMONIC, and PINNsAgent across all evaluated problems. Ablation studies further examine the effects of online adaptation, state-conditioned intervention selection, and execution-feedback-driven policy evolution.
Chinese Translation
物理信息神经网络(PINNs)为求解偏微分方程(PDEs)提供了一种基于学习的框架,然而其训练行为在优化过程中可能发生显著变化。残差分布、梯度交互、区域学习难度和模型容量需求可能随时间演变,而网络架构和主要训练机制通常在训练前就已确定。我们提出了PINNMorph,一种基于大语言模型(LLM)引导的策略演化的在线PINN自适应框架。PINNMorph维护一个状态条件自适应策略群体,这些策略将执行诊断映射到对拓扑修改、加性表示增强、目标平衡、梯度处理、自适应采样和优化器阶段控制的受控干预。在每个干预机会,候选程序从当前策略群体中实例化,根据观察到的训练状态进行选择,并直接应用于正在训练的PINN。得到的模型继承其现有参数和训练状态,并沿相同轨迹继续优化。执行结果随后用于评估干预并演化策略群体。与预训练架构搜索或固定自适应规则不同,PINNMorph利用来自实际训练的反馈,同时自适应当前PINN及其干预策略。在13个PDE基准上的实验表明,PINNMorph在所有评估问题上比SA-PINN、ConFIG、RoPINN、HARMONIC和PINNsAgent实现了更低的求解误差。消融研究进一步考察了在线自适应、状态条件干预选择和执行反馈驱动的策略演化的效果。
cs.AI / 117 / 2609.32687

Dude, Where's My State? Execution Information Requirements for Stateful Agents

老兄,我的状态呢?有状态智能体的执行信息需求
Mehrotra, Nikita, Tiwari, Ashish, Gupta, Priyanshu, Gulwani, Sumit
Abstract
Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information that must remain accessible for correct completion under specified task and access conditions. We develop LACUNA, a framework that generates tasks with known dependencies and varies information demand, retention, and recovery separately from the difficulty of individual operations. Across four models, restoring a missing result raises accuracy on affected recall steps to 100%, compared with 0% for equal-length irrelevant information. Sufficient storage alone does not ensure success: retention policies can discard required results, errors can propagate through later computations, and agents can stop before recovery is complete. We also introduce VESTIGE, which uses agent execution traces to construct semantic graphs and measure information demand for real tasks. Across 72,562 software-agent trajectories, VESTIGE reveals a steeper distance-related decline in solution-relevant rereading for failed runs (RR 0.951 per distance doubling), while adjusted peak demand alone is not associated with failure. Together, these contributions support evaluating whether agents preserve and recover the information their tasks require.
Chinese Translation
长时间运行的智能体必须保留后续步骤所依赖的信息。我们引入了执行信息需求(EIR),这是在指定的任务和访问条件下,为了正确完成而必须保持可访问的信息的下界。我们开发了LACUNA框架,该框架生成具有已知依赖关系的任务,并将信息需求、保留和恢复与单个操作的难度分开变化。在四个模型中,恢复缺失的结果可将受影响的回忆步骤的准确率提高到100%,而等长无关信息则为0%。仅靠充足的存储并不能确保成功:保留策略可能会丢弃所需的结果,错误可能会通过后续计算传播,智能体可能在恢复完成之前就停止。我们还引入了VESTIGE,它使用智能体执行轨迹构建语义图,并度量真实任务的信息需求。在72,562条软件智能体轨迹中,VESTIGE揭示了失败的运行中与解决方案相关的重读随距离增加而更陡峭的下降(距离每增加一倍,RR为0.951),而仅调整后的峰值需求与失败无关。这些贡献共同支持评估智能体是否保留和恢复了其任务所需的信息。
cs.AI / 118 / 2609.32692

World Agent: Can Language Models Keep a World Running?

World Agent:语言模型能让世界持续运行吗?
Chen, Weixing, Zhang, Weipeng, An, Nan, Liu, Yang, Lin, Liang
Abstract
World models are moving from generating realistic frames to generating playable worlds, yet whether a delivered world can keep running is not tested anywhere. Existing evaluations stop at generation, at delivery, or at single-step transitions, and each stops at a different point along the way. Correct local state transitions or intermediate outcomes do not guarantee a correctly organized causal event flow. We propose the world agent task, which moves the evaluation point of world generation from the moment of delivery to the continued operation that follows. In this task, a model is not asked to generate a world. It is held responsible for keeping the world running, which requires coordinating events and carrying forward their consequences to constrain subsequent evolution. We instantiate the task in WorldAgent-Benchmark with two complementary tracks. In the maintenance track, the model must ground the events of a continuous narrative into correct transitions of the explicit world state while respecting causal, temporal, and concurrency constraints. In the deduction track, the model must predict how the world will evolve under partial observations and act toward a goal. The maintenance track combines LLM-assisted semantic judgments with programmatic validation and scoring, while the deduction track is evaluated entirely programmatically. Individual judgments are auditable against world states and execution logs, and scores can be recomputed from the saved judgments and execution records. Across 8 models, scores decline steadily as pre-built structure is removed from the world, and causal-relation checking is the weakest component for every model. The benchmark makes the continued operation of a world measurable and distinguishes local completion from failures in event organization. Code and dataset will be released on https://github.com/HCPLab-SYSU/WorldAgent-Benchmark.
Chinese Translation
世界模型正从生成逼真的画面转向生成可游玩的世界,然而,一个已交付的世界能否持续运行,尚未在任何地方得到测试。现有的评估停留在生成阶段、交付阶段或单步转移阶段,而每一个都停留在这一过程中的不同节点。正确的局部状态转移或中间结果并不能保证正确组织的因果事件流。我们提出世界智能体任务(world agent task),它将世界生成的评估点从交付时刻转移到随后的持续运行上。在该任务中,模型不被要求生成一个世界,而是被要求负责保持世界持续运行,这需要协调事件并传递其后果以约束后续演化。我们在 WorldAgent-Benchmark 中实例化该任务,包含两个互补的赛道。在维护赛道中,模型必须将连续叙事中的事件锚定到显式世界状态的正确转移中,同时遵守因果、时间和并发约束。在推演赛道中,模型必须预测世界在部分观测下将如何演化,并朝着目标采取行动。维护赛道结合了 LLM 辅助的语义判断与程序化验证和评分,而推演赛道则完全以程序化方式进行评估。个别判断可以根据世界状态和执行日志进行审计,分数可以从保存的判断和执行记录中重新计算。在 8 个模型中,随着从世界中移除预构建结构,分数稳步下降,而因果关系检查是每个模型最薄弱的环节。该基准使世界的持续运行变得可度量,并将局部完成与事件组织中的失败区分开来。代码和数据集将在 https://github.com/HCPLab-SYSU/WorldAgent-Benchmark 上发布。
cs.AI / 119 / 2609.32694

IGSD: Environment-Verified Hindsight Self-Distillation for Search Agents

IGSD:面向搜索智能体的环境验证后见自蒸馏
Jiang, Angqing, Zhang, Gaoming, Zhang, Chaoqun, Song, Jianchun, Kong, Liyuan, Qi, Kena, Lin, Wei, Lian, Defu
Abstract
On-policy self-distillation densifies agent training without external teachers: a policy conditioned on privileged hindsight provides step-level guidance for its own unprivileged rollouts. For search agents, however, hindsight can make the teacher prefer a query that does not improve retrieval from the student's state. Existing methods either distill this preference directly or filter it with model-internal scores, but neither strategy verifies the query's executed retrieval consequence. We propose Information-Gain-Gated Self-Distillation (IGSD), which verifies on-policy token proposals with environment feedback before distilling them. Treating each query token as a micro-action, IGSD completes the teacher's token proposal and the student's sampled token into matched queries and executes both from the same failed state with the same retriever. Shared counterfactual controls account for query-conditioned shifts in answer likelihood, so their difference, the executed paired information gain, provides a relative utility contrast for the retrieved documents. IGSD uses this contrast as a positive-only soft weight for candidate-pair distillation, while leaving the GRPO objective unchanged and confining verification to training. Across seven single-hop and multi-hop QA benchmarks, IGSD reaches macro-average exact-match accuracies of 42.8% and 47.0% with 3B and 7B policies, respectively, without inference-time verification. These results support environment-verified hindsight as an effective approach to reliable action-level supervision for search agents.
Chinese Translation
在策略自蒸馏无需外部教师即可稠密化智能体训练:一个以特权后见信息为条件的策略为其自身无特权的轨迹提供步骤级指导。然而,对于搜索智能体,后见信息可能使教师偏好一个不能从学生状态改善检索的查询。现有方法要么直接蒸馏这种偏好,要么用模型内部评分过滤它,但这两种策略都没有验证查询实际执行的检索后果。我们提出信息增益门控自蒸馏(IGSD),它在蒸馏之前用环境反馈验证在策略的token提议。将每个查询token视为微动作,IGSD将教师的token提议和学生的采样token补全为匹配的查询,并从相同的失败状态用相同的检索器执行两者。共享的反事实控制解释了答案似然中查询条件的变化,因此它们的差异,即执行配对信息增益,为检索到的文档提供了相对效用对比。IGSD将此对比用作候选对蒸馏的仅正软权重,同时保持GRPO目标不变,并将验证限制在训练中。在七个单跳和多跳QA基准上,IGSD分别以3B和7B策略达到42.8%和47.0%的宏平均精确匹配准确率,无需推理时验证。这些结果支持环境验证的后见信息作为搜索智能体可靠动作级监督的有效方法。
cs.AI / 120 / 2609.32696

Flat-Consensus Diffusion for Robust Data Reshaping under Noisy Evaluator

噪声评估器下鲁棒数据重塑的平坦共识扩散
Cao, Hongyu, Liu, Kunpeng, Xie, Fei, Ray, Sandip
Abstract
Data shape determines how features are structured, how patterns are separated, and how distributions cover the underlying domain. Poor data shape can make models learn noise rather than generalizable structure. This paper studies robust feature-centric data reshaping: generating feature transformations that remain useful, stable, and reproducible under noisy evaluation and imperfect data conditions. We view reshaping operation sequence search as reward-guided diffusion generation, and robust reshaping as searching for regions in the latent reward landscape rather than isolated high-reward transformations. The key challenge is dual instability: noisy evaluators distort local reward guidance, while stochastic generative trajectories can converge to inconsistent solutions. We propose FCDiff, a flat-consensus diffusion framework that addresses both failures through a micro-macro decomposition. The micro layer replaces point-estimate reward guidance with Gaussian-smoothed, Monte Carlo averaged gradients, steering generation toward locally flat reward regions. The macro layer aggregates independently guided trajectories with a weighted Frechet-mean barycenter, selecting consensus-supported basins and filtering stochastic outliers. Across an 8-dataset headline cohort under heavy-tailed evaluator noise, FCDiff attains the best aggregate rank on lower-tail reliability and robustness against both search-based AutoFE and robustness-oriented generative baselines, with statistically significant accuracy gains over every generative baseline. Our results show that robust data reshaping requires searching for flat, consensus-supported regions rather than sharp single-trajectory optima.
Chinese Translation
数据形态决定了特征如何组织、模式如何分离,以及分布如何覆盖底层领域。不良的数据形态会使模型学到噪声而非可泛化的结构。本文研究鲁棒的、以特征为中心的数据重塑:生成在噪声评估和不完美数据条件下仍然有用、稳定且可复现的特征变换。我们将重塑操作序列搜索视为奖励引导的扩散生成,并将鲁棒重塑视为在潜在奖励景观中搜索区域,而非孤立的高奖励变换。关键挑战是双重不稳定性:噪声评估器扭曲局部奖励引导,而随机生成轨迹可能收敛到不一致的解。我们提出 FCDiff,一种平坦共识扩散框架,通过微观-宏观分解解决这两个失败。微观层用高斯平滑的蒙特卡洛平均梯度替代点估计奖励引导,引导生成走向局部平坦的奖励区域。宏观层用加权 Frechet 均值重心聚合独立引导的轨迹,选择共识支持的盆地并过滤随机异常值。在重尾评估器噪声下的 8 个数据集主要队列中,FCDiff 在低尾可靠性和鲁棒性上,相较于基于搜索的 AutoFE 和面向鲁棒的生成基线,取得了最佳综合排名,并且在每个生成基线上都获得了统计显著的准确率提升。我们的结果表明,鲁棒数据重塑需要搜索平坦的、共识支持的区域,而非尖锐的单轨迹最优。
cs.AI / 121 / 2609.32700

CAIRN: Dynamic Fact-Intent DAGs for Multi-Agent Exploration

CAIRN:面向多智能体探索的动态事实-意图DAG
Xu, Zuyao, Jia, Yuyang, Guan, Junwei, Li, Xiang, Shen, Kaiwen, Dong, Zhiqiang
Abstract
LLM-powered autonomous systems have demonstrated promising capabilities in mathematical reasoning, engineering, and cybersecurity. Yet how to organize these systems for effective, reliable, and sustained performance remains an open question. In this paper, we present CAIRN, a fact-intent-driven multi-agent paradigm for goal-directed exploration. CAIRN represents observations and planned investigations as a dynamic directed acyclic graph (DAG). A reasoner interprets facts to propose intents, which workers execute to produce new facts. Each intent references its supporting facts and defines a potential exploration branch. The persistent graph preserves goals, dependencies and findings across workers, supporting knowledge reuse and parallel exploration. The graph also makes execution trajectories traceable and auditable, providing a basis for human verification and intervention. We evaluate CAIRN across cybersecurity and mathematical reasoning tasks, examining task success, time to solution, and token consumption. DAG-based coordination can incur higher token costs with no observable performance gains on tasks that require little effort. However, on high-effort tasks (at least 1M tokens), we observe faster solutions in 76.5% of cases, with speedups of up to 3.08x. Moreover, as task effort increases, these time gains become more pronounced while relative token overhead declines, highlighting the potential of DAG-guided parallel exploration.
Chinese Translation
由LLM驱动的自主系统已在数学推理、工程和网络安全等领域展现出令人瞩目的能力。然而,如何组织这些系统以实现有效、可靠且持续的性能仍是一个未决问题。在本文中,我们提出了CAIRN,一种事实-意图驱动的多智能体范式,用于目标导向的探索。CAIRN将观察结果和计划中的调查表示为动态有向无环图(DAG)。推理器解释事实以提出意图,工作者执行这些意图以产生新的事实。每个意图引用其支持事实,并定义一个潜在的探索分支。持久图保留了跨工作者的目标、依赖关系和发现,支持知识重用和并行探索。该图还使执行轨迹可追溯和可审计,为人工验证和干预提供了基础。我们在网络安全和数学推理任务上评估了CAIRN,考察任务成功、求解时间和令牌消耗。基于DAG的协调可能会产生更高的令牌成本,并且在需要较少努力的任务上未观察到性能提升。然而,在高努力任务(至少100万令牌)上,我们观察到76.5%的案例中求解速度更快,加速比高达3.08倍。此外,随着任务努力的增加,这些时间收益变得更加明显,而相对令牌开销下降,突显了DAG引导的并行探索的潜力。
cs.AI / 122 / 2609.32701

Despite Instructions: Frontier Agents Improvise Covert Channels at Test Time

尽管有指令:前沿智能体在测试时即兴构建隐蔽信道
Dineen, Jacob, Ren, Silei, Chen, Muhao, Roth, Dan, Zhou, Ben
Abstract
In security-sensitive applications, language-model agents are often required to coordinate without disclosing confidential information. Yet repeated interactions may also let ordinary messages acquire shared private meaning. We study a repeated game with pairs of models in which the sender model observes one of four secret states and selects one of four summaries of the same public report, while the receiver model tries to infer the secret state. We find that model pairs can learn to communicate the secret using only one bit of feedback indicating whether the receiver inferred it correctly. This learning occurs during inference with fixed parameters and no supplied codebook or encoding examples. The effect also persists when agents generate their own free-form updates in a simulated incident-response task. Across ten independent games, pairs of GPT-5.6 Sol agents reach 98.8% final accuracy, compared with 25% chance, despite explicit instructions prohibiting disclosure and a monitor that screens each message without access to the agents' interaction histories. The same interactions that help agents cooperate can therefore allow confidential information to pass through messages intended for legitimate coordination.
Chinese Translation
在安全敏感的应用中,语言模型智能体通常需要在不泄露机密信息的情况下进行协调。然而,重复交互也可能使普通消息获得共享的私密含义。我们研究了一种由成对模型进行的重复博弈:其中发送方模型观察到四种秘密状态之一,并从同一公开报告的四种摘要中选择一种,而接收方模型试图推断秘密状态。我们发现,模型对仅使用一比特反馈(指示接收方是否推断正确)就能学会传递秘密。这种学习发生在参数固定的推理过程中,且没有提供码本或编码示例。当智能体在模拟的事件响应任务中生成自己的自由形式更新时,这种效应也持续存在。在十个独立博弈中,成对的 GPT-5.6 Sol 智能体达到了 98.8% 的最终准确率,而随机概率为 25%,尽管有明确禁止披露的指令,并且有一个监控器会筛查每条消息但无法访问智能体的交互历史。因此,帮助智能体协作的相同交互,也可能使机密信息通过原本用于合法协调的消息传递。
cs.AI / 123 / 2609.32704

CoWindow Attention: Full Causal Coverage Is a Collective Property

CoWindow注意力:全因果覆盖是一种集体属性
Shi, Jingze, Peng, Zhangyang, Li, Xianduo, Qi, Yanlin, Lin, Xiaotian, Chen, Haoxian, Wang, Liangdong, Liu, Guang, Luo, Yuyu
Abstract
FullAttn repeatedly exposes the complete causal history to every attention head, creating substantial redundant computation and memory traffic even with IO-efficient dense kernels. We introduce CoWA, a structured attention architecture that distributes access to the causal history across KV heads. All heads share near-diagonal and prefix-sink windows, while complementary long-range windows partition the remaining history. Their union provides full causal coverage although each head attends sparsely to distant tokens. This position-defined attention pattern requires no learned router or indexer, is used consistently during training and inference, and aligns with KV-head tensor parallelism. A window-matched ablation at 8K isolates the effect of complementary long-range allocation: CoWA with 100% collective coverage reaches 89.73% accuracy, compared with 89.97% for FullAttn, while duplicated long-range windows perform substantially worse. Across a broader controlled associative-recall comparison with matched token budgets, CoWA closely tracks FullAttn as the context grows, whereas other sparse patterns lose a substantial fraction of the associations. In an attention-operator benchmark at 128K tokens with tensor parallelism, CoWA reduces forward and backward latency during training by 7.4x and 8.6x and decoding latency during inference by 3.0x over FullAttn. Its per-rank peak operator memory matches FullAttn during training and is 7.6x lower during decoding. Across scaling-law training from 0.6B to 14B parameters, CoWA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results show that full causal coverage can be a collective property of the head ensemble rather than a duplicated property of every head.
Chinese Translation
FullAttn将完整的因果历史反复暴露给每个注意力头,即使采用IO高效的稠密核,也会造成大量冗余计算和内存流量。我们提出CoWA,一种结构化注意力架构,将因果历史的访问分配到各个KV头。所有头共享近对角窗口和前缀汇窗口,而互补的长距离窗口划分剩余历史。它们的并集提供了完整的因果覆盖,尽管每个头稀疏地关注遥远的标记。这种位置定义的注意力模式不需要学习路由或索引器,在训练和推理中一致使用,并且与KV头张量并行对齐。在8K上下文下的窗口匹配消融实验分离出互补长距离分配的效果:具有100%集体覆盖的CoWA达到89.73%的准确率,而FullAttn为89.97%,而重复的长距离窗口表现明显更差。在更广泛的、匹配标记预算的受控联想回忆对比中,随着上下文增长,CoWA紧密跟踪FullAttn,而其他稀疏模式丢失了相当一部分联想。在128K标记的注意力算子基准测试中,采用张量并行,相比FullAttn,CoWA在训练期间将前向和后向延迟分别降低7.4倍和8.6倍,在推理期间将解码延迟降低3.0倍。在训练期间,其每卡峰值算子内存与FullAttn相当,而在解码期间低7.6倍。在从0.6B到14B参数的缩放律训练中,CoWA在困惑度上紧密跟踪FullAttn,同时减少总训练FLOPs。通过单独继续训练得到的14B和32B模型在知识、推理和长上下文检索得分上与FullAttn相当。这些结果表明,完整的因果覆盖可以是头集合的集体属性,而不是每个头的重复属性。
cs.AI / 124 / 2609.32712

MassAlloc Attention: Let Attention Allocate Its Own Compute

MassAlloc 注意力:让注意力分配其自身计算
Shi, Jingze, Peng, Zhangyang, Li, Xianduo, Qi, Yanlin, Lin, Xiaotian, Chen, Haoxian, Wang, Liangdong, Liu, Guang, Luo, Yuyu
Abstract
FullAttn often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MALA, a fused attention primitive that preserves score access to every legal causal interaction and uses normalized contribution to allocate post-score computation. Forward uses its evolving online-softmax normalizer, while backward reuses the finalized normalizer to derive nested retained support using only standard attention state. A common tolerance governs training and inference, allowing for adaptive retention of the work. MALA reduces low-contribution post-score computation. A matched-work study at 8K isolates the benefit of distribution-adaptive allocation: under exactly matched total post-score work, MALA approaches a per-instance reference-mass oracle, with mean omitted mass of 0.0188% versus 0.0182%. Across context lengths from 1K to 32K tokens, the same tolerance maintains low output and gradient errors relative to the reference. Across a broader controlled associative-recall comparison, MALA closely tracks FullAttn as context grows, reaching 89.67% accuracy at 8K compared with 89.97% for FullAttn. In an attention-operator benchmark at 128K tokens with tensor parallelism, MALA reduces forward and backward latency during training by 2.2x and 3.0x and decoding latency during inference by 1.6x relative to FullAttn. Across scaling-law training from 0.6B to 14B parameters, MALA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results indicate that allocating post-score computation according to normalized attention contributions can retain the evaluated capabilities of FullAttn while reducing attention computation.
Chinese Translation
FullAttn 通常为大部分因果分数空间分配可忽略的归一化质量,但稠密核在形成每个 QK tile 后仍执行完整的分数后路径。我们引入 MALA,一种融合注意力原语,它保留对每个合法因果交互的分数访问,并使用归一化贡献来分配分数后计算。前向使用其演化的在线 softmax 归一化器,而反向重用最终确定的归一化器,仅使用标准注意力状态来推导嵌套保留支持。一个统一容差控制训练和推理,允许对工作负载进行自适应保留。MALA 减少了低贡献的分数后计算。在 8K 下的匹配工作量研究分离出分布自适应分配的好处:在完全匹配的总分数后计算量下,MALA 接近逐实例参考质量 oracle,平均省略质量为 0.0188%,而参考为 0.0182%。在 1K 到 32K token 的上下文长度范围内,相同的容差相对于参考保持了低的输出和梯度误差。在更广泛的受控联想召回比较中,随着上下文增长,MALA 紧密跟随 FullAttn,在 8K 时达到 89.67% 的准确率,而 FullAttn 为 89.97%。在 128K token 且使用张量并行的注意力算子基准测试中,相对于 FullAttn,MALA 在训练中将前向和后向延迟分别降低了 2.2 倍和 3.0 倍,在推理中将解码延迟降低了 1.6 倍。在从 0.6B 到 14B 参数的缩放律训练中,MALA 在困惑度上紧密跟随 FullAttn,同时减少了总训练 FLOPs。由此产生的 14B 模型和来自单独继续训练的 32B 模型在知识、推理和长上下文检索分数上与 FullAttn 相当。这些结果表明,根据归一化注意力贡献分配分数后计算,可以在减少注意力计算的同时,保留 FullAttn 的已评估能力。
cs.AI / 125 / 2609.32731

SkillVine: Agent Skill Evolution via Branching Exploration

SkillVine:通过分支探索的智能体技能进化
Liu, Kaiwei, Dong, Jiqian, Dong, Liran, Mao, Shuai, Zhao, Mingming, Yang, Bufang, Chuai, Jie, Chen, Zhitang, Xing, Guoliang, Yan, Zhenyu
Abstract
Agent skills encapsulate reusable procedural knowledge that enables LLM agents to perform tasks, and they can be improved automatically using trajectories from interactions with the environment. This is the classic problem of skill evolution. Existing approaches predominately follow a linear evolution paradigm, in which updates are sequentially applied to the latest skill-library version. As a result, they inevitably fall into local optima, leaving many promising evolution paths unexplored. We propose SkillVine, an automatic skill-evolution framework that formulates skill evolution as a graph search problem and employs a branching exploration strategy. Equipped with a trunk-branch collaborative searching mechanism, an intelligent parent-node selector, and an adaptive-granularity update rule, SkillVine achieves a balance between exploration and exploitation. We evaluate SkillVine on 5 benchmarks with two LLMs. Results show that SkillVine discovers better skill-library versions along branches than along the linear trunk and achieves the best test performance in nine of ten benchmark-model combinations.
Chinese Translation
智能体技能封装了可复用的过程性知识,使LLM智能体能够执行任务,并且可以利用与环境交互的轨迹自动改进。这是经典的技能进化问题。现有方法主要遵循线性进化范式,其中更新按顺序应用于最新的技能库版本。结果,它们不可避免地陷入局部最优,许多有前景的进化路径未被探索。我们提出SkillVine,一个自动技能进化框架,将技能进化形式化为图搜索问题,并采用分支探索策略。凭借主干-分支协同搜索机制、智能父节点选择器和自适应粒度更新规则,SkillVine在探索和利用之间实现了平衡。我们在5个基准测试上使用两个LLM评估SkillVine。结果表明,SkillVine沿着分支发现的技能库版本优于沿着线性主干发现的版本,并在十个基准-模型组合中的九个上取得了最佳测试性能。
cs.AI / 126 / 2609.32749

Retrospective Distillation Attribution via Normalized Response Similarity

基于归一化响应相似度的回顾性蒸馏归因
Jang, Minwoo, Kim, Jaechang, Oh, Minhyeon, Hwang, Jeongyeon, Ok, Jungseul
Abstract
Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students immediately after the SFT step. However, a distilled model may undergo further SFT, preference optimization, or reinforcement learning before release, while an auditor may lack access to the pre-distillation checkpoint required by reference-based attribution. To close this gap, we propose SCOUT, an output-only method that aggregates recurring *syntactic patterns* into candidate profiles, filters low-contrast patterns, and calibrates student--candidate distances against inter-candidate distances. SCOUT supports attribution and abstention using only current texts, without model weights, token likelihoods, or historical checkpoints. Auditing publicly released descendants of distilled models spanning diverse post-training objectives, SCOUT consistently identifies the distillation source. Furthermore, tracing teacher-associated *syntactic signatures* along training trajectories reveals that they emerge during distillation and persist through subsequent preference optimization and reinforcement learning.
Chinese Translation
模型蒸馏通过在教师响应上进行监督微调(SFT)来迁移能力,这些响应通常从商业API收集,引发了模型溯源问题。现有的蒸馏归因方法大多在SFT步骤后立即对学生模型进行评估。然而,蒸馏模型在发布前可能会经历进一步的SFT、偏好优化或强化学习,而审计者可能无法访问基于参考的归因所需的蒸馏前检查点。为弥补这一差距,我们提出了SCOUT,一种仅基于输出的方法,它将反复出现的句法模式聚合为候选画像,过滤低区分度的模式,并利用候选间距离来校准学生-候选距离。SCOUT仅使用当前文本即可支持归因和弃权,无需模型权重、token似然或历史检查点。在审计涵盖多种后训练目标的已发布蒸馏模型后代时,SCOUT能一致地识别出蒸馏源。此外,沿训练轨迹追踪教师相关的句法签名表明,这些签名在蒸馏过程中出现,并在后续的偏好优化和强化学习中持续存在。
cs.AI / 127 / 2609.32750

CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning

CUA-Sandbox:用于计算机使用智能体强化学习的高效环境
Yan, Xin, Jiao, Zhengbo, Liu, Jiaqi, Wan, Zhenglin, Ma, SiYuan, Yu, Xuliang, Jiang, Tianyi, Zhang, Chubin, Zhou, Pengfei, Zhao, Wangbo, Yu, Xingrui, An, Bo, You, Yang, Tsang, Ivor
Abstract
Reinforcement learning enables computer-use agents to improve through interaction with real software environments, including websites and desktop applications. However, conventional deployments replicate an initialized runtime for each independent rollout, even when trajectories use the same software, incurring repeated memory and initialization costs as the number of parallel environments grows. Does an independent computer-use environment require an independent execution runtime? Our key observation is that trajectories require independent mutable state, while initialized application runtimes can be reused across concurrently evolving environments, making state the natural unit of environment independence. Guided by this observation, we introduce CUA-Sandbox, which separates private state capsules from shared runtimes through state-scoped execution and transactional lifecycle operations, including resets and branches, while retaining the original software interfaces and task evaluators. Experiments show comparable or improved task success relative to Docker, while substantially reducing rollout and resource costs. CUA-Sandbox achieves up to a 6.20x increase in rollout throughput, a 9.2x reduction in per-environment memory, and a 504x reduction in incremental storage.
Chinese Translation
强化学习使计算机使用智能体能够通过与真实软件环境(包括网站和桌面应用程序)交互来改进。然而,传统部署会为每个独立的rollout复制一个初始化的运行时,即使轨迹使用相同的软件,随着并行环境数量的增长,也会导致重复的内存和初始化开销。一个独立的计算机使用环境是否需要独立的执行运行时?我们的关键观察是:轨迹需要独立的可变状态,而初始化的应用程序运行时可以在并发演化的环境之间重用,从而使状态成为环境独立性的自然单元。受此观察启发,我们提出了CUA-Sandbox,它通过状态域执行和事务性生命周期操作(包括重置和分支),将私有状态胶囊与共享运行时分离,同时保留原始软件接口和任务评估器。实验表明,与Docker相比,任务成功率相当或有所提高,同时大幅降低了rollout和资源成本。CUA-Sandbox实现了高达6.20倍的rollout吞吐量提升、9.2倍的每环境内存降低以及504倍的增量存储降低。
cs.AI / 128 / 2609.32752

Action Shaping: Policies Absorb What They Can Express

动作塑形:策略吸收其所能表达者
Chen, Yanjun, Wang, Jinghan, Shen, Xiaoyu, Li, Wenjie, Zhang, Wei
Abstract
Reward shaping has a theorem: a potential-based term can be removed without changing the optimal policy. The same practice on the action channel, an offset added in training and dropped at deployment, has no theorem. Nothing cancels an action offset, so the correction is kept at deployment or removed without a guarantee. We call it action shaping and state its principle. A trainable policy absorbs an offset its own output layer can reproduce exactly, which is what we mean by express; what is absorbed can be removed with the return intact. Its minimal instance is a zero-initialized linear head behind a learnable gate, added to an actor that trains through a learned action-value function, with no penalty or schedule. The gate rises and then falls on its own, for deterministic and stochastic actors alike, and on 20 tasks removing the head costs almost nothing. The condition is exact reproduction, not capacity: a nonlinear head with more parameters is not absorbed, and in a paired control, one linear path added to a nonlinear base head restores absorption. Exact reproduction gives the loss a flat direction that gradient noise drifts along, and the offset's amplitude indicates, before removal, what dropping the head will cost. Action shaping thus gains the counterpart of the shaping theorem, a condition for absorption, together with the mechanism behind it and a diagnostic that reads it. Policies absorb what they can express, and only that.
Chinese Translation
奖励塑形有一个定理:基于势能的项可以被移除而不改变最优策略。在动作通道上的同样做法,即在训练中添加并在部署时丢弃的偏移量,却没有定理。没有什么能抵消动作偏移,因此该修正要么在部署时保留,要么在没有保证的情况下被移除。我们称之为动作塑形,并陈述其原理。一个可训练策略会吸收其自身输出层能够精确复现的偏移,这就是我们所说的“表达”;被吸收的部分可以在回报不变的情况下移除。其最小实例是一个零初始化的线性头,位于一个可学习门控之后,添加到一个通过学习的动作价值函数进行训练的演员上,没有惩罚或调度。该门控自行上升然后下降,对于确定性和随机演员都是如此,并且在20个任务上移除该头几乎不付出代价。条件是精确复现,而非容量:具有更多参数的非线性头不会被吸收,而在配对对照中,向非线性基础头添加一条线性路径可恢复吸收。精确复现为损失提供了一个平坦方向,梯度噪声沿其漂移,而偏移的幅度在移除之前指示了丢弃该头的代价。因此,动作塑形获得了塑形定理的对应物:一个吸收条件,连同其背后的机制和一种读取它的诊断方法。策略吸收它们能表达的内容,仅此而已。
cs.AI / 129 / 2609.32754

Adaptive Consistency Graph for Long-Horizon Agents

面向长周期智能体的自适应一致性图
Wang, Jiecong, Peng, Hao, Wang, Zhanyi
Abstract
Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long sequences of dependent actions and tool calls. During execution, task requirements, historical evidence, and the current execution state may gradually become disconnected, so later decisions can drift from the original objective. We study this problem by introducing the Adaptive Consistency Graph (ACG) for long-horizon execution. ACG incrementally organizes execution evidence and its provenance in a persistent graph, then constructs a temporary requirement-centered view for each decision under a bounded context budget. Rather than replacing the base agent's planner or tool executor, ACG provides a structured and traceable context view for each decision. In the matched evaluation, ACG improves GPT-5.6-luna's average success from 44.5\% with ReAct to 50.2\%, with the largest gain on BrowseComp-Plus (73.5\% versus 62.4\%). We further analyze trajectory structure and inference cost to characterize this improvement.
Chinese Translation
大型语言模型智能体在短任务上通常能做出合理的局部决策,但当成功需要长序列的依赖动作和工具调用时,其性能会下降。在执行过程中,任务需求、历史证据和当前执行状态可能逐渐脱节,导致后续决策偏离原始目标。我们通过引入用于长周期执行的自适应一致性图(ACG)来研究这个问题。ACG 在持久图中增量地组织执行证据及其来源,然后在有限的上下文预算下为每个决策构建以需求为中心的临时视图。ACG 并不替代基础智能体的规划器或工具执行器,而是为每个决策提供结构化且可追溯的上下文视图。在匹配评估中,ACG 将 GPT-5.6-luna 的平均成功率从 ReAct 的 44.5% 提升至 50.2%,在 BrowseComp-Plus 上提升最大(73.5% 对 62.4%)。我们进一步分析轨迹结构和推理成本来刻画这一改进。
cs.AI / 130 / 2609.32757

Readout is not Recovery: Dissociating Coordinate Emission from Visual-Corruption Repair in Vision-Language Models

读出并非恢复:解离视觉语言模型中的坐标发射与视觉损坏修复
Juanico, Drandreb Earl
Abstract
VLM bounding-box localization is both language generation and spatial commitment. Parseable fields such as bbox_2d make localization easy to score, but dimensions that emit coordinate tokens need not repair localization after visual evidence is damaged. We study this readout/recovery separation in Qwen3-VL-4B-Instruct on single-object COCO grounding. We compare clean coordinate-token readout rankings with corruption-derived repair rankings, using object-mask endpoint replacement for recovery and clean-input flooring for depth localization. In Qwen3-VL, coordinate-token rankings are inert through layer 24, load-bearing from layers 32-35, and peak at layer 34; corruption-derived rankings harm layers 16-24 but become beneficial near layer 35/final. A Kimi-VL-A3B diagnostic shows a matching output-proximal transition despite a different box format. Object-mask recovery separates rank budgets: $k=250$ shows necessity, $k=500$ shows Top-$k$ restoration above random, and $k=d/2$ is largely capacity-driven. Partial-occlusion sweeps reveal that high-overlap coordinate-token sets can hurt at $k=1000$ and help mainly at half-width, while population corruption-derived sets provide no reliable fixed repair set. Edge-attribution patching shows coordinate-token paths are high precision but low recall for detection recovery, and RMSNorm quasi-layer controls do not close the endpoint-repair gap. Endpoint coordinate triage is therefore a useful circuit prior, but occlusion recovery requires a separate benchmark.
Chinese Translation
VLM 边界框定位既是语言生成,也是空间承诺。诸如 bbox_2d 这样可解析的字段使定位易于评分,但在视觉证据受损后,输出坐标 token 的维度并不一定能够修复定位。我们在 Qwen3-VL-4B-Instruct 上针对单目标 COCO 定位任务研究这种读出/恢复的分离。我们比较干净的坐标 token 读出排名与基于损坏的修复排名,使用对象掩码端点替换进行恢复,并使用干净输入下限进行深度定位。在 Qwen3-VL 中,坐标 token 排名在第 24 层之前是无作用的,在第 32-35 层起承载作用,并在第 34 层达到峰值;基于损坏的排名损害第 16-24 层,但在第 35 层/最终层附近变得有益。Kimi-VL-A3B 的诊断显示出匹配的输出邻近转变,尽管框格式不同。对象掩码恢复分离了排名预算:$k=250$ 显示必要性,$k=500$ 显示 Top-$k$ 恢复高于随机,而 $k=d/2$ 主要是容量驱动的。部分遮挡扫描揭示,高重叠坐标 token 集在 $k=1000$ 时可能有害,主要在半宽时有益,而群体损坏导出的集合没有提供可靠的固定修复集。边缘归因修补显示,坐标 token 路径对于检测恢复具有高精度但低召回率,并且 RMSNorm 准层控制不能弥补端点修复差距。因此,端点坐标分诊是有用的电路先验,但遮挡恢复需要单独的基准测试。
cs.AI / 131 / 2609.32763

Mandela-Bench: Multimodal Models Remember Canonical Images Instead of Seeing Them

Mandela-Bench:多模态模型记住规范图像而非看到它们
Bao, Yicheng, Gao, Zhenkun, Guo, Xiahui, Yang, Mingqian, Li, Xueheng, Liu, Bangwei, Chen, Mingang, Li, Lijun, Wang, Xuhong, Tan, Xin
Abstract
Historical photographs and other canonical images can now be edited seamlessly with a single instruction, often leaving no reliable pixel-level trace. In such cases, the only evidence of manipulation may be a fact about what the image depicts. Existing benchmarks instead rely on generator artefacts, image-caption inconsistencies, visual implausibilities, or external references, and therefore do not test whether a model can use its own world knowledge to verify a recognized image. We introduce Mandela-Bench, containing 1,507 edits of canonical images: 1,359 knowledge-only forgeries, each contradicting one verifiable fact, and 148 anchor-free controls that preserve the editing process without introducing a factual contradiction, together with 474 untouched originals. We score not only whether a model detects a forgery, but whether its explanation identifies the inserted entity or the fact being violated. Across 36 multimodal models, from 0.8B parameters to frontier scale, we find a consistent failure mode. When a public figure is removed from a familiar photograph, models still name that person in up to 72.7% of responses. Some models can distinguish the replacement face from the original when shown in isolation, yet still judge the full edited photograph as authentic. Providing the true event and date does not improve knowledge-grounded detection, whereas providing the same information after cropping away the recognizable composition does. Even under explicit verification prompts, only one of the 36 models meets the KGR criterion on at least half of the forged images. These results suggest that the failures cannot be explained by missing knowledge or inadequate perception alone. Instead, they are consistent with recognition biasing verification toward the remembered canonical image rather than the observed edit.
Chinese Translation
历史照片和其他规范图像现在可以通过单一指令无缝编辑,往往不留下可靠的像素级痕迹。在这种情况下,篡改的唯一证据可能是关于图像所描绘内容的一个事实。现有的基准测试则依赖于生成器伪影、图像-标题不一致、视觉不合理性或外部参考,因此没有测试模型是否能够利用自身世界知识来验证一张被识别的图像。我们引入了Mandela-Bench,包含1,507张规范图像的编辑:1,359个仅知识型伪造,每个都违背一个可验证的事实;148个无锚控制,保留编辑过程但不引入事实矛盾;以及474张未修改的原图。我们不仅评估模型是否检测到伪造,还评估其解释是否识别出插入的实体或被违背的事实。在36个多模态模型中,从0.8B参数到前沿规模,我们发现了一致的失败模式。当从一张熟悉的照片中移除一位公众人物时,模型在高达72.7%的回复中仍然说出那个人的名字。一些模型在单独展示时能够区分替换的脸和原始的脸,但仍然将完整编辑后的照片判定为真实。提供真实的事件和日期并不能改善基于知识的检测,而在裁剪掉可识别的构图后提供相同的信息则可以。即使在明确的验证提示下,36个模型中只有一个在至少一半的伪造图像上满足KGR标准。这些结果表明,这些失败不能仅用缺失知识或感知不足来解释。相反,它们与识别偏向于记住的规范图像而非观察到的编辑相一致。
cs.AI / 132 / 2609.32773

Forecasting Intraday USD/CAD Exchange Rate with News-Derived Monetary-Policy Signals

基于新闻衍生的货币政策信号预测日内USD/CAD汇率
Kodeih, Maya, Alnaggar, Aliaa, Cevik, Mucahit
Abstract
Monetary-policy announcements and central-bank communications play a central role in foreign exchange markets, yet their qualitative, unstructured form makes their forecasting value difficult to quantify. While prior research has largely focused on sentiment extracted from financial news, comparatively little is known about the relative contribution of different dimensions of monetary-policy communication. Existing studies primarily evaluate whether textual information improves overall forecasting performance but provide limited insight into which communication channels drive such improvements. To address this gap, this paper introduces a statistical attribution methodology that decomposes monetary-policy communication into interpretable channels and quantifies their incremental forecasting contribution under false-discovery-rate control. Monetary-policy news is transformed into structured communication signals using large language models (LLMs) and temporal feature engineering. These signals are evaluated using rolling-window experiments with tree-based machine-learning models. The results show that monetary-policy communication contains measurable predictive information. Attribution analysis shows that predictive value is concentrated in a small subset of signals, with communication timing providing the strongest individual feature-level contribution, targeted communication-activity measures also contributing positively, and LLM-derived sentiment providing complementary information at the group level. The findings indicate that communication-based forecasting value extends beyond sentiment alone and that attribution, rather than aggregate accuracy alone, is central to evaluating news-derived signals.
Chinese Translation
货币政策公告和央行沟通在外汇市场中发挥着核心作用,但其定性、非结构化的形式使其预测价值难以量化。虽然此前的研究主要集中于从金融新闻中提取的情绪,但对于货币政策沟通不同维度的相对贡献知之甚少。现有研究主要评估文本信息是否改善了整体预测性能,但对于哪些沟通渠道驱动了此类改进所提供的见解有限。为弥补这一空白,本文引入了一种统计归因方法,将货币政策沟通分解为可解释的渠道,并在错误发现率控制下量化其增量预测贡献。货币政策新闻使用大语言模型(LLMs)和时间特征工程转化为结构化的沟通信号。这些信号通过基于树的机器学习模型的滚动窗口实验进行评估。结果表明,货币政策沟通包含可测量的预测信息。归因分析表明,预测价值集中在少数信号中,沟通时机提供了最强的个体特征级贡献,针对性的沟通活动度量也做出了正向贡献,而LLM衍生的情绪在组级别提供了补充信息。研究结果表明,基于沟通的预测价值不仅限于情绪,归因(而非仅总体准确率)对于评估新闻衍生的信号至关重要。
cs.AI / 133 / 2609.32778

Agentic Network Traffic Monitoring

代理式网络流量监控
Tsoukatos, Manuel, Jananthan, Hayden, Kepner, Jeremy
Abstract
As the use of agentic artificial intelligence increases in nearly every industry, there exists a widening attack surface. It is necessary to monitor agents to ensure that agents are acting in a way that is aligned with the users intent. Auditing an agent's network traffic provides a clear record of the agent interactions. This work presents a novel approach to monitoring the network traffic of agentic systems using complex valued hypersparse traffic matrices by integrating DBOS (DataBase OS), the OneSparse PostgreSQL database, and the GraphBLAS math library. To develop these concepts an agentic simulator was constructed, allowing a varying numbers of AI agents to collectively survey a virtual environment using different strategies. The resulting network traffic matrices enable easy monitoring of the AI agents.
Chinese Translation
随着代理式人工智能在几乎每个行业中的使用不断增加,攻击面也在不断扩大。有必要对智能体进行监控,以确保智能体的行为符合用户的意图。审计智能体的网络流量可以提供智能体交互的清晰记录。这项工作提出了一种新颖的方法,通过集成DBOS(DataBase OS)、OneSparse PostgreSQL数据库和GraphBLAS数学库,使用复值超稀疏流量矩阵来监控代理式系统的网络流量。为了开发这些概念,构建了一个代理式模拟器,允许不同数量的AI智能体使用不同策略共同调查虚拟环境。由此产生的网络流量矩阵使得对AI智能体的监控变得容易。
cs.AI / 134 / 2609.32782

Learning response-aware patient dynamics for respiratory support

学习用于呼吸支持的响应感知患者动态
Lu, Xiaolei, Nemati, Shamim
Abstract
Respiratory support can shape the short-term physiological trajectory of critically ill patients, but patients receiving the same intervention may follow different physiological trajectories. Clinical patient dynamics models typically predict future states from recent physiology and recorded interventions, while physiological change is mainly represented through the predicted future state. We propose a response-aware patient dynamics model that explicitly represents physiological change during autoregressive state updating. The model decomposes predicted physiological change into state-dependent baseline dynamics and respiratory-support-associated deviations, with room air providing a reference for the decomposition. We provide a formal analysis of this reference-anchored formulation. A response pathway encodes the predicted physiological change and uses it to update the latent patient state across the forecast horizon. Across ICU cohorts from two independent institutions, the proposed model achieves comparable overall trajectory prediction to patient dynamics baselines, with more consistent improvements when physiological states are changing.
Chinese Translation
呼吸支持可塑造重症患者的短期生理轨迹,但接受相同干预的患者可能遵循不同的生理轨迹。临床患者动态模型通常根据近期生理状态和记录的干预来预测未来状态,而生理变化主要通过预测的未来状态来体现。我们提出一种响应感知的患者动态模型,在自回归状态更新过程中显式表示生理变化。该模型将预测的生理变化分解为状态依赖的基线动态和与呼吸支持相关的偏差,并以室内空气作为分解的参考。我们对这一参考锚定公式进行了形式化分析。一条响应通路对预测的生理变化进行编码,并利用它在整个预测时域内更新潜在患者状态。在两个独立机构的ICU队列中,所提模型在整体轨迹预测上取得了与患者动态基线相当的性能,并且在生理状态发生变化时具有更一致的改善。
cs.AI / 135 / 2609.32787

CLAIRE: A Schema-Grounded Hybrid Workflow for Healthcare Administrative Form Completion

CLAIRE:一种基于模式的混合工作流,用于医疗行政表单填写
Keerthana, Garapati, Gupta, Manik
Abstract
Healthcare administrative staff transfer structured information from electronic health records, referrals, claims systems, provider rosters, and work queues into dynamic forms. We developed and evaluated CLAIRE (Clinical Language and Agentic Intelligence for Reasoning and Entry), a hybrid workflow that separates field-state discovery, source-to-field mapping, deterministic validation, bounded correction, escalation, and audit tracing. We tested five synthetic healthcare administrative schemas, 1,000 source records, four interface variants, two data-quality suites, and six comparators, yielding 24,000 benchmark episodes. A separate strict-output audit evaluated direct mappings from Qwen2.5-1.5B and Qwen2.5-7B, and a trace-derived operational simulation covered 6,000 episodes. Under the evaluated synthetic benchmark conditions, full CLAIRE achieved 1.000 episode success, field accuracy, required-field completion, and dependency completion in both suites; removing validation reduced stress-suite success to 0.500. In the simulation, 100.0% of clean and validation-stress episodes reached a staff-reviewable draft, compared with 68.6% of escalation challenge episodes, unsupported cases were blocked. Scenario-based savings were 149.7-165.5 seconds per case, not observed staff times. The findings support schema-grounded, validation-first healthcare administrative automation in which language-model components assist mapping but do not authorize unsupported or consequential actions.
Chinese Translation
医疗行政人员将结构化信息从电子健康记录、转诊、理赔系统、提供者名册和工作队列转移到动态表单中。我们开发并评估了 CLAIRE(Clinical Language and Agentic Intelligence for Reasoning and Entry,用于推理与录入的临床语言与代理智能),这是一种混合工作流,将字段状态发现、源到字段映射、确定性验证、有界纠正、升级和审计追踪分离。我们测试了五种合成医疗行政模式、1,000 条源记录、四种界面变体、两个数据质量套件和六个比较器,产生了 24,000 个基准回合。一项单独的严格输出审计评估了来自 Qwen2.5-1.5B 和 Qwen2.5-7B 的直接映射,并且一项基于轨迹推导的操作模拟覆盖了 6,000 个回合。在所评估的合成基准条件下,完整 CLAIRE 在两个套件中均实现了 1.000 的回合成功率、字段准确率、必填字段完成率和依赖项完成率;移除验证后,压力套件的成功率降至 0.500。在模拟中,100.0% 的干净回合和验证压力回合达到了可供员工审核的草稿,相比之下,升级挑战回合为 68.6%;不受支持的案例被阻止。基于场景的节省为每例 149.7–165.5 秒,而非观察到的员工用时。研究结果支持基于模式、验证优先的医疗行政自动化,其中语言模型组件辅助映射,但不授权不受支持或具有重大后果的操作。
cs.AI / 136 / 2609.32791

$T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training

T^5:强化中期训练中词元级思维的孪生评论家训练
Qiao, Nan, Yang, Yebin, Wang, Weinong, Wang, Shuning, Peng, Shangpin, Lu, Fengyuan, Wang, Xinming, Kan, Zhehan, Zhang, Ruixu, Zhang, Songyang, Yue, Sheng, Tian, Yonglong, Ren, Ju
Abstract
Reinforcement mid-training lets language models learn internal thoughts from unlabeled text, but efficient token-level credit assignment remains challenging. Existing group-relative methods require costly repeated generation. Learned critics offer single-rollout feedback, but accurate return prediction alone does not ensure reliable policy updates. Our analysis shows how training--inference mismatch and PPO clipping prevent a common offset in advantage estimates from cancelling out, introducing additional update drift. We propose \tfour{}, a twin-critic method that calibrates token-level advantages from a single generated trajectory. After warmup and held-out qualification, the critics provide two advantage estimates, combined using action-dependent weights learned through a conditional-moment saddle-point objective. This objective brings the average advantage at each prefix toward zero, while a signal-retention constraint prevents the correction from erasing the learning signal. Sharing information across text positions avoids repeated sampling of each prefix. Theoretically, we characterize optimal mixing under the signal-retention constraint and establish an upper bound on residual mean-induced drift. Experiments show that, compared with the state-of-the-art critic-free method, \tfour{} improves mean benchmark performance by 7.8\% and reduces mean training-step time by up to 63.4\%.
Chinese Translation
强化中期训练使语言模型能够从无标签文本中学习内部思维,但高效的词元级信用分配仍然具有挑战性。现有的组相对方法需要代价高昂的重复生成。学习到的评论家提供单次展开反馈,但仅准确的回报预测并不能确保可靠的策略更新。我们的分析表明,训练-推理不匹配和 PPO 裁剪如何阻止优势估计中的共同偏移相互抵消,从而引入额外的更新漂移。我们提出 T^5,一种孪生评论家方法,能够从单个生成的轨迹中校准词元级优势。经过预热和留出验证后,评论家提供两个优势估计,并通过条件矩鞍点目标学习到的动作相关权重进行组合。该目标使每个前缀处的平均优势趋向于零,同时信号保留约束防止校正抹除学习信号。跨文本位置共享信息避免了对每个前缀的重复采样。在理论上,我们刻画了信号保留约束下的最优混合,并建立了残差均值诱导漂移的上界。实验表明,与最先进的无评论家方法相比,T^5 将平均基准性能提高了 7.8%,并将平均训练步时间减少了高达 63.4%。
cs.AI / 137 / 2609.32795

AgentHabit: Characterizing Distinct Behaviors of Agents on Everyday Tasks

AgentHabit:刻画智能体在日常任务中的独特行为
Song, Woojung, Yang, Hoyeol, Shim, Jeonghoon, Lim, Sungjib, Lee, Jonggeun, Choi, Yunho, Jo, Yohan
Abstract
Large language model (LLM) agents assist users with everyday tasks that can be completed in many reasonable ways. Even when their answers are useful, how agents carry out these tasks may not match users' preferences and needs. For example, agents differ in whether they ask clarifying questions or search the web. We introduce HABIT, a taxonomy of 23 behavioral axes in five categories, which three authors and three LLMs derive bottom-up from 408 agent trajectories across 17 domains. On held-out tasks, HABIT distinguishes models more clearly than existing taxonomies of human values and agent actions while supporting comparably consistent annotation. Building on HABIT, we construct AgentHABIT, a benchmark that profiles each agent's behavioral tendencies from its trajectories on 86 everyday tasks. Profiling 18 models with AgentHABIT reveals a range of distinctive tendencies. For example, most GPT and Claude models state their assumptions and offer alternatives when requirements conflict, whereas Qwen and Google's models more often leave assumptions or changes to requirements unstated. These profiles remain recognizable even when built from entirely different sets of tasks, indicating that they reflect general tendencies rather than task-specific behavior. Prompting agents to adopt specific behaviors shifts some axes readily but barely changes others, while fine-tuning on another model's trajectories changes only part of a model's profile and leaves much of it intact. Overall, HABIT and AgentHABIT provide a systematic framework for characterizing how agents carry out everyday tasks beyond task success, offering insights to guide the development of agents whose behavior better fits users' needs.
Chinese Translation
大型语言模型(LLM)智能体协助用户完成可以通过多种合理方式完成的日常任务。即使它们的答案有用,智能体如何执行这些任务也可能不符合用户的偏好和需求。例如,智能体在是否提出澄清问题或搜索网络方面存在差异。我们介绍了HABIT,一个包含五个类别下23个行为维度的分类体系,由三位作者和三个LLM从17个领域的408条智能体轨迹中自底向上推导得出。在留出任务上,HABIT比现有的人类价值观和智能体动作分类体系更能清晰地区分模型,同时支持相当一致的标注。基于HABIT,我们构建了AgentHABIT,一个从86个日常任务上的轨迹中刻画每个智能体行为倾向的基准。用AgentHABIT对18个模型进行画像,揭示了一系列独特的倾向。例如,大多数GPT和Claude模型在需求冲突时会陈述其假设并提供替代方案,而Qwen和Google的模型则更常将假设或对需求的更改未加说明。这些画像即使从完全不同的任务集构建,也仍然可识别,表明它们反映的是一般倾向而非特定任务的行为。提示智能体采取特定行为会容易地改变某些维度,但几乎不改变其他维度;而在另一个模型的轨迹上微调只改变模型画像的一部分,而大部分保持不变。总体而言,HABIT和AgentHABIT提供了一个系统框架,用于刻画智能体如何完成日常任务(超越任务成功率),提供洞见以指导开发行为更符合用户需求的智能体。
cs.AI / 138 / 2609.32801

PlanGuard: A Guardrail for Multi-Step Plan Safety in Embodied Agents

PlanGuard:具身智能体多步计划安全护栏
Chen, Junchi, Miao, Changtao, Xiang, Yuxiao, Jin, Zhenchao, Yuan, Haojie, Chu, Qi, Gong, Tao, Liu, He, Zhang, Bo, Cai, Jiansheng, Li, Zhe, Yu, Nenghai
Abstract
Embodied task planners may produce multi-step plans whose subtask dependencies and interactions with the environment create physical risks during execution. Yet existing safeguards overlook such compositional risks, as general-purpose guardrails focus on semantic harm and embodied safety detectors assess subtasks in isolation. To address this gap, we introduce PlanGuard, the first pre-execution detector that evaluates the physical safety of a complete multi-step plan in its current environment. For training and evaluation, we construct a Multi-Step Plan Safety (MSP-Safe) dataset through paired task construction, plan generation using diverse planners, and safety annotation by three judges. Task-oriented SFT on MSP-Safe establishes fundamental plan-safety assessment capabilities, yet a substantial gap remains between compact models suitable for real-time deployment and stronger but costlier large models. Accordingly, we propose Strong-Teacher Adaptive Compensation for On-Policy Distillation (STAC-OPD), which provides compact models with adaptive strong-teacher supervision along their on-policy trajectories. It combines token-level distribution transfer from a fine-tuned strong teacher with probability-routed sequence-level compensation, retaining student-generated targets when the student favors the reference safety decision and using teacher-reconstructed targets otherwise. Across all test subsets, PlanGuard-2B achieves average 87.15% ACC and 87.21% F1, demonstrating effective whole-plan physical-risk detection at compact model scale. Code and dataset will be publicly released.
Chinese Translation
具身任务规划器可能生成多步计划,其子任务依赖关系及与环境的交互在执行过程中会产生物理风险。然而,现有的安全防护措施忽视了这种组合风险,因为通用护栏关注语义危害,而具身安全检测器孤立地评估子任务。为解决这一空白,我们提出了PlanGuard,这是首个在执行前评估完整多步计划在当前环境中物理安全性的检测器。为了训练和评估,我们通过配对任务构建、使用多样化规划器生成计划、以及由三名评判员进行安全标注,构建了多步计划安全(MSP-Safe)数据集。在MSP-Safe上进行面向任务的SFT建立了基本的计划安全评估能力,但适用于实时部署的紧凑模型与更强大但成本更高的大型模型之间仍存在显著差距。因此,我们提出了强教师自适应补偿的在线策略蒸馏(STAC-OPD),它沿着紧凑模型的在线策略轨迹为其提供自适应的强教师监督。它结合了来自微调强教师的令牌级分布转移与概率路由的序列级补偿,当学生倾向于参考安全决策时保留学生生成的目标,否则使用教师重构的目标。在所有测试子集上,PlanGuard-2B平均达到87.15%的准确率和87.21%的F1值,展示了在紧凑模型规模下有效的全计划物理风险检测。代码和数据集将公开发布。
cs.AI / 139 / 2609.32802

Re-derivability Decides What a Staged Agent Pipeline Recovers After an Upstream Fault

可重推导性决定分阶段智能体流水线在上游故障后恢复什么
Bu, Tianqi, Peng, YuXuan, Tu, Junteng, Xiao, Henghui
Abstract
One variable sets what an upstream fault costs a staged pipeline of language-model agents: re-derivability, how much of what a stage needs it can rebuild from the original problem. Grounding an inspector agent in that problem is worth +0.608 [+0.517, +0.700] to +0.358 over a blind one on four open-weight backbones served with thinking disabled, and on the two Qwen backbones the blind inspector changes no item at all. That head-to-head is exploratory. One deterministic fault enters the first stage, and we re-expose the original problem to $k = 0,\dots,3$ of the downstream stages with agents, items, fault and topology held fixed, on 120 gsm_hard items per arm at temperature zero. Accuracy under fault rises on four of four backbones, from +0.233 to +0.392, the largest Holm-adjusted $p$ being $2.1\times10^{-6}$. A registered kill test rules out tokens. Blanking every word holds the word slots fixed, and retention tracks the visible fraction on four of four, climbing from 0.221 to 0.692 on the primary. Those two families are confirmatory and everything else here is exploratory. The interaction excludes zero on two of four backbones under the registered pipeline, four of four under a three-stage pipeline, and three of four under full message history, the primary at +0.317. On Llama-3.1-8B the fault carries no detectable cost at any dose, so the other three carry every claim about what a fault costs. Re-derivability also sets what the architecture costs, and no decomposition we measured reliably beats one direct call. With no fault injected the registered pipeline loses to that call by -0.267, -0.125 and -0.317, and on Phi-4 reads +0.058 at $p = 0.118$, which the test fails to separate from zero. The repair that works is cheap and front-loaded: the first re-grounded stage buys +0.394 of matched retention for +59.8 tokens per item on Qwen3-14B, and the stages after it buy nothing.
Chinese Translation
一个变量决定了上游故障给语言模型智能体分阶段流水线带来的代价:可重推导性,即一个阶段所需的内容中有多少可以从原始问题中重建。在禁用思考(thinking disabled)服务的四个开放权重(open-weight)主干模型上,将检查器智能体(inspector agent)基于该问题进行接地(grounding)比盲检查器(blind inspector)的价值高出 +0.608 [+0.517, +0.700] 到 +0.358,并且在两个 Qwen 主干模型上,盲检查器完全不改变任何项目。这种头对头(head-to-head)比较是探索性的。一个确定性故障进入第一阶段,我们在每个实验臂(arm)120 个 gsm_hard 项目上,在温度为零的情况下,将原始问题重新暴露给下游阶段中的 k = 0,...,3 个阶段,同时保持智能体、项目、故障和拓扑不变。在四个主干模型上,故障下的准确率均上升,从 +0.233 到 +0.392,最大的 Holm 校正 p 值为 2.1×10^-6。一项注册的排除测试(kill test)排除了 token 的影响。将每个单词置空(Blanking every word)保持单词槽位固定,留存率(retention)在四个主干模型上均与可见比例(visible fraction)一致,在主要模型上从 0.221 攀升至 0.692。这两个系列是验证性的,而这里的其他一切都是探索性的。在注册流水线下,交互作用在四个主干模型中的两个上排除零,在三阶段流水线下四个均排除零,在完整消息历史下四个中的三个排除零,主要模型为 +0.317。在 Llama-3.1-8B 上,故障在任何剂量下都没有可检测的代价,因此其他三个模型承担了关于故障代价的所有主张。可重推导性也决定了架构的代价,而我们测量的任何分解都无法可靠地胜过直接调用一次。在没有注入故障时,注册流水线输给该直接调用,分别为 -0.267、-0.125 和 -0.317,而在 Phi-4 上为 +0.058,p = 0.118,测试未能将其与零区分开。有效的修复是廉价且前置的:在 Qwen3-14B 上,第一个重新接地的阶段以每个项目 +59.8 个 token 的代价换取了 +0.394 的匹配留存率,而其后的阶段没有带来任何收益。
cs.AI / 140 / 2609.32803

Nutri-ATLAS: Embodied Agent for Tabulated Lookup and Assistance for Smarter nutrition

Nutri-ATLAS:用于更智能营养的表格查询与辅助的具身智能体
Kallakuri, Uttej, Hu, Boxun, Butala, Ankur A., Dehak, Najim, Mohsenin, Tinoosh
Abstract
Generative and Agentic IoT systems offer a promising foundation for digital healthcare applications that combine sensing, personalized reasoning, and autonomous interaction in real-world environments. Nutrition assistance is a natural use case, but existing Large Language Model (LLM)-based systems are often limited to passive text interaction and static context, making them unreliable when food descriptions are ambiguous or nutritional evidence is missing. We propose Nutri-ATLAS, an Embodied Agent for Tabulated Lookup and Assistance for smarter nutrition in the real world. It integrates graph-grounded nutrition reasoning, hardware-aware LLM selection, and robot-based evidence acquisition. Nutri-ATLAS builds a unified Food-Nutrient knowledge graph from USDA FoodData Central and FoodKG and learns 64-dimensional GATv2 food and recipe embeddings. A shared hybrid graph-text scoring mechanism supports food nutrition extraction, nutritional gap filling, substitute retrieval, and recipe-level meal composition, while an LLM-guided skill interface navigates landmarks, updates dietary-context and food-accessibility memory, and grounds recommendations in observed food availability. We evaluate Nutri-ATLAS across nutrient estimation, substitution retrieval, recipe recommendation, patient-profile adherence, edge deployment, and real-world embodied execution. On HealthyFoodSubs, the hybrid retriever achieves 37.9% MAP, 80.7% RR@5, and 90.1% RR@10. On NutriBench v2, Dense+GAT retrieval grounds nutrient estimation across nine quantized Qwen3.5-9B configurations. On PFoodReQ, Nutri-ATLAS reaches 78.8% MAP, 83.0% MAR, and 77.5% F1. A patient-profile study shows adherence to allergy and healthy-target constraints for all selected cases.
Chinese Translation
生成式和代理式物联网系统为数字医疗应用提供了有前景的基础,这些应用将传感、个性化推理和自主交互结合在真实世界环境中。营养辅助是一个自然的用例,但现有的基于大语言模型(LLM)的系统通常局限于被动的文本交互和静态上下文,当食物描述模糊或营养证据缺失时,它们变得不可靠。我们提出了Nutri-ATLAS,一个用于在真实世界中实现更智能营养的表格查询与辅助的具身智能体。它集成了基于图的营养推理、硬件感知的LLM选择以及基于机器人的证据获取。Nutri-ATLAS从USDA FoodData Central和FoodKG构建了一个统一的食物-营养素知识图谱,并学习了64维的GATv2食物和食谱嵌入。一个共享的混合图-文本评分机制支持食物营养提取、营养缺口填充、替代品检索和食谱级膳食组成,而一个由LLM引导的技能接口导航地标、更新饮食上下文和食物可及性记忆,并将推荐建立在观察到的食物可用性上。我们在营养素估计、替代品检索、食谱推荐、患者画像依从性、边缘部署和真实世界具身执行方面评估了Nutri-ATLAS。在HealthyFoodSubs上,混合检索器达到了37.9%的MAP、80.7%的RR@5和90.1%的RR@10。在NutriBench v2上,Dense+GAT检索在九个量化Qwen3.5-9B配置中为营养素估计提供了基础。在PFoodReQ上,Nutri-ATLAS达到了78.8%的MAP、83.0%的MAR和77.5%的F1。一项患者画像研究显示,对于所有选定的案例,均遵守了过敏和健康目标约束。
cs.AI / 141 / 2609.32805

Decision-Sufficient State Representations: Measuring and Reducing Write-Time Regret

决策充分的状态表示:测量与减少写入时遗憾
Shen, Bingyu, Li, Boyang
Abstract
Long tasks produce more history than an LLM agent can hold in its context, and more than it uses reliably even when the history fits. A growing line of work therefore has agents carry a short written state instead: at every step a writer rewrites the state, and a reader acts from the state alone. Steps stay cheap, but anything the writer drops is lost before later decisions reveal that they need it. We quantify this loss and ask whether training can reduce it. Comparing the written state with the best state of the same size written in hindsight, we split the reader's loss into a budget loss, which any state of that size must incur, and a write-time regret, which comes from the writer's choices. In TextWorld cooking games where we control how long a fact must be carried before it is needed, a 128-token state holding the facts wins nearly every game, while prompted language-model writers win at most 17%. Almost all of the loss is write-time regret, and it grows with the delay. We then train the writer from the reader's own loss. DSSR (decision-sufficient state representations) scores candidate states by how well the reader acts after the writer carries them forward, and teaches the writer to prefer the better ones. This forward-rolled score predicts game outcomes ($\rho = 0.48$), whereas scoring a candidate as a fixed context, as hindsight methods usually do, does not ($\rho \leq 0.07$). On a pre-registered test split opened once, training adds +7.0 [+1.9, +12.2] points of success when facts are needed soon, bringing a plain summary writer to the level of belief- and slot-based memory prompts. The gain shrinks as the delay grows and is significant only at the shortest delay. We trace this limit to credit assignment: keeping a fact now pays off only if every later rewrite keeps it too, which a per-step score cannot see.
Chinese Translation
长任务产生的历史记录超过LLM智能体上下文所能容纳的量,甚至即使历史记录能容纳,也超过它可靠使用的量。因此,越来越多的工作让智能体携带一个简短的书面状态:在每个步骤,写入者重写状态,读者仅根据状态行动。步骤保持低成本,但写入者丢弃的任何内容都会在后续决策揭示需要它之前丢失。我们量化这种损失,并探讨训练是否能减少它。将书面状态与事后写出的相同大小的最佳状态进行比较,我们将读者的损失分解为预算损失(任何该大小的状态都必须承受的损失)和写入时遗憾(来自写入者的选择)。在TextWorld烹饪游戏中,我们控制一个事实在需要之前必须被携带多久,一个包含这些事实的128词元状态几乎赢得每一场游戏,而提示的语言模型写入者最多赢得17%。几乎所有的损失都是写入时遗憾,并且随着延迟增长。然后,我们从读者自身的损失训练写入者。DSSR(决策充分的状态表示)通过写入者将候选状态向前携带后读者的表现来评分,并教写入者偏好更好的状态。这种前向滚动评分能预测游戏结果(ρ=0.48),而将候选状态作为固定上下文评分(如事后方法通常所做)则不能(ρ≤0.07)。在一个预先注册、只打开一次的测试集上,当很快需要事实时,训练增加了+7.0 [+1.9, +12.2]个成功点,使一个简单的摘要写入者达到基于信念和槽位的记忆提示的水平。随着延迟增加,增益缩小,且仅在最短延迟时显著。我们将这一限制归因于信用分配:现在保留一个事实只有在后续每次重写也保留它时才有回报,而逐步评分无法看到这一点。
cs.AI / 142 / 2609.32807

Beyond Accuracy: Counterfactual Fragility and Demographic Bias in Clinical Evaluation of LLMs

超越准确性:大型语言模型临床评估中的反事实脆弱性与人口统计学偏差
Purkayastha, Chaitai Deb, Bolla, Bharath Kumar, Nandi, Vishnu Surya Reddy
Abstract
Clinical LLM evaluation often emphasizes answer accuracy; however, accuracy alone does not test counterfactual consistency or demographic robustness. We evaluated six LLMs on 150 MedQA USMLE questions using two automated perturbation tests to assess their performance. The counterfactual validity (CFV) test asked each model to make a minimal, plausible clinical change that would make a different answer correct. The demographic robustness test added six demographic prefixes to the same vignette and compared the answers and explanations with a no demographic baseline. Of the 900 CFV attempts, 228 (25.3 %) were valid and 672 were invalid. Across 5,400 demographic comparisons, 1,097 answers were changed (20.3%). Automated judging identified 3,128 stereotype evidence flags, including 1,932 in the broad Other category. MedGemma 27B achieved the highest accuracy (87.1%) and CFV (63.3%), lowest answer change rate (16.0%), and low mean Explanation Demographic Dissonance (EDD) score (0.169). However, its accuracy still exceeded its CFV, indicating that correct answers do not guarantee reliable performance on the counterfactual validity task. OpenBioLLM had the highest answer change rate and EDD, whereas GLM had the highest stereotype flag rate. These findings show that accuracy, CFV, answer stability, EDD, and stereotype evidence capture different evaluation aspects. Because all judgments were automated and no clinician validation was available, the results support safety screening but do not establish clinical deployability of the model.
Chinese Translation
临床大型语言模型(LLM)评估往往强调答案准确性;然而,仅凭准确性并不能检验反事实一致性或人口统计学稳健性。我们使用两项自动化扰动测试,在150道MedQA USMLE题目上评估了六个LLM的表现。反事实有效性(CFV)测试要求每个模型做出最小、合理的临床改动,从而使另一个答案变为正确。人口统计学稳健性测试在同一病例情景中加入六个人口统计学前缀,并将答案和解释与无人口统计学信息的基线进行比较。在900次CFV尝试中,228次(25.3%)有效,672次无效。在5,400次人口统计学比较中,1,097个答案发生了改变(20.3%)。自动化评判识别出3,128个刻板印象证据标记,其中1,932个属于广义的“其他”类别。MedGemma 27B取得了最高准确性(87.1%)和CFV(63.3%)、最低答案改变率(16.0%)以及较低的平均解释人口统计学失调(EDD)评分(0.169)。然而,其准确性仍高于其CFV,表明正确答案并不能保证在反事实有效性任务上具有可靠表现。OpenBioLLM的答案改变率和EDD最高,而GLM的刻板印象标记率最高。这些发现表明,准确性、CFV、答案稳定性、EDD和刻板印象证据捕捉的是不同的评估方面。由于所有判断均为自动化且没有临床医生验证,这些结果支持安全性筛查,但并不能确立该模型的临床可部署性。
cs.AI / 143 / 2609.32809

Overwhelmed by Choice: Studying LLM Decision Making at Scale

选择过载:大规模LLM决策研究
Lin, Yu-Chi, Seth, Aryan, Aravind, Anshul, Lee, Eugene, Parekh, Tanmay, Peng, Nanyun, Chang, Kai-Wei
Abstract
Multiple-choice and candidate-selection evaluations are widely used to assess LLM reasoning and decision-making, yet most benchmarks contain relatively small candidate sets. It remains unclear whether conclusions drawn from these settings remain valid as the candidate space scales. We systematically evaluate LLMs as the number of competing candidates increases and find substantial accuracy degradation across tasks, prompting strategies, and model scales. Controlled analyses show that standard long-context retrieval explanations cannot fully account for this degradation. Instead, we identify two systematic failure patterns. First, gold-margin collapse: the score gap between the correct answer and the strongest distractor progressively shrinks, driven primarily by weakening confidence in the correct answer. Second, earlier candidate preferences become increasingly difficult to overturn, with later candidates exerting progressively weaker influence on the final prediction. Motivated by these findings, we evaluate hierarchical partitioning and permutation-based inference, which improve accuracy by roughly 20 percentage points at $N=160$ on both HotpotQA and MIMIC. Overall, our results identify candidate-set scale as an important evaluation-protocol variable and show that strong small-option performance does not necessarily imply robust large-scale candidate comparison.
Chinese Translation
多项选择和候选选择评估被广泛用于评估LLM的推理和决策能力,然而大多数基准测试的候选集相对较小。随着候选空间的扩大,从这些设置中得出的结论是否仍然有效尚不清楚。我们系统地评估了LLM在竞争候选数量增加时的表现,发现无论任务、提示策略还是模型规模,准确率均出现显著下降。控制分析表明,标准的长上下文检索解释无法完全解释这种性能下降。相反,我们识别出两种系统性失败模式。首先,正确边际崩溃(gold-margin collapse):正确答案与最强干扰项之间的分数差距逐渐缩小,这主要是由于对正确答案的信心减弱所驱动。其次,早期候选偏好变得越来越难以推翻,而后期候选对最终预测的影响逐渐减弱。受这些发现的启发,我们评估了分层划分和基于排列的推理方法,在HotpotQA和MIMIC上,当$N=160$时,这些方法将准确率提高了约20个百分点。总的来说,我们的结果将候选集规模确定为一种重要的评估协议变量,并表明在小选项集上的强劲表现并不一定意味着在大规模候选比较中具有稳健性。
cs.AI / 144 / 2609.32817

Right Answer, Wrong Reason: Accuracy, Consistency, and Consensus Are Misleading Indicators of LLM Faithfulness in Clinical Decision Support

正确答案,错误理由:准确性、一致性与共识是临床决策支持中LLM忠实性的误导性指标
Bolla, Bharath Kumar, Bolla, Bharath Kumar, Nandi, Vishnu Surya Reddy
Abstract
Clinical Large Language Models (LLMs) achieve strong medical-exam accuracy; however, a correct answer does not guarantee that the explanation names the concepts that actually drove the decision. We introduce three lightweight, directly interpretable metrics for this faithfulness gap: the Explanation Stability Index (ESI), which measures reasoning consistency across repeated queries; the Causal Faithfulness Score (CFS), which tests whether cited concepts drive predictions via concept ablation; and the Perturbation Stability Score (PSS), which measures robustness to semantic-preserving paraphrases. By evaluating six LLMs on 150 MedQA-USMLE questions (900 model-question observations), we found that only 23.3% of the cited clinical concepts were causally necessary. Correct answers had lower CFS than incorrect answers (0.212 vs. 0.398), answer consistency negatively predicted CFS (Spearman r = -0.466), and model pairs could agree on answers while sharing only 8.8% of cited reasoning concepts. These results show that accuracy, consistency, and consensus are incomplete safety signals for clinical decision-making support. The evidence is behavioral rather than mechanistic: concept ablation tests counterfactual sensitivity of outputs, not internal circuits.
Chinese Translation
临床大语言模型(LLMs)在医学考试中达到了很高的准确率;然而,一个正确的答案并不能保证其解释提及了真正驱动该决策的概念。我们为这一忠实性缺口引入了三个轻量级、可直接解释的指标:解释稳定性指数(Explanation Stability Index, ESI),用于衡量重复查询中的推理一致性;因果忠实性得分(Causal Faithfulness Score, CFS),通过概念消融检验所引用的概念是否驱动预测;以及扰动稳定性得分(Perturbation Stability Score, PSS),用于衡量对语义保持型复述的稳健性。通过在150道MedQA-USMLE问题上评估六个LLM(900次模型-问题观测),我们发现只有23.3%的被引用临床概念是因果必要的。正确答案的CFS低于错误答案(0.212 vs. 0.398),答案一致性负向预测CFS(Spearman r = -0.466),并且模型对可以在答案上一致,却仅共享8.8%的所引用推理概念。这些结果表明,准确性、一致性和共识是临床决策支持中不完整的安全性信号。证据是行为层面的而非机制层面的:概念消融测试的是输出的反事实敏感性,而不是内部电路。
cs.AI / 145 / 2609.32818

When Can First-Order Models of Fine-Tuning Bound Forgetting?

微调的一阶模型何时能够约束遗忘?
Su, Jianchang, Zhang, Wei
Abstract
Fine-tuning a language model on new data can make it forget facts that it should keep. We ask whether measurements taken at the start of a fine-tuning run can bound, for each protected fact, the probability that the run makes the model forget it. In LoRA fine-tuning with stochastic gradient descent on models from 0.6B to 14B parameters, a first-order response model estimated by finite-difference probes predicts changes of per-fact margins with correlation 0.974-0.998. Predictions of forgetting built on this model nevertheless failed, because forgetting requires parameter changes far outside the region in which the model was validated. The probes can, however, bound the probability that a margin first falls below a boundary near zero: we derive Freedman and Azuma first-passage bounds for a linear surrogate of the margin and test on new runs whether they hold for the model. The bounds contain a term R that measures how much the response coefficients change during the run. The simplified Freedman bound, which sets R = 0, certified most facts but was violated in 14 of 112 conditions, and every fact on which it was violated had R >= a, where a is the distance of the fact's margin to the boundary. The complete Freedman bound certifies only facts with R < a, and it held in every condition. On the violated facts, the spread of the margin across test runs was a median of 14.6 times the prediction of the response model, so the failures are breakdowns of the model, and in our data they occurred only where R >= a. We found this pattern post hoc and tested it in two preregistered confirmatory studies with 43 new conditions: the complete bound held in all of them, and the simplified bound failed there on only 3 facts, each with R >= a. First-order models of fine-tuning can thus bound forgetting on the facts whose response coefficients change by less than their distance to the boundary.
Chinese Translation
在新数据上微调语言模型可能会使其忘记本应记住的事实。我们询问,在微调运行开始时进行的测量能否为每个受保护事实约束该运行使模型忘记它的概率。在LoRA微调中,使用随机梯度下降,模型规模从0.6B到14B参数,通过有限差分探针估计的一阶响应模型预测每个事实边际的变化,相关性为0.974-0.998。然而,基于该模型的遗忘预测失败了,因为遗忘需要参数变化远远超出模型被验证的区域。然而,探针可以约束边际首次低于接近零的边界的概率:我们为边际的线性代理推导了Freedman和Azuma首达时界,并在新的运行中测试它们对模型是否成立。这些界包含一个项R,度量响应系数在运行期间变化多少。简化的Freedman界将R设为0,认证了大多数事实,但在112个条件中有14个被违反,且每个被违反的事实都有R >= a,其中a是该事实的边际到边界的距离。完整的Freedman界仅认证R < a的事实,并且它在每个条件下都成立。在被违反的事实上,测试运行中边际的散布是响应模型预测的中位数的14.6倍,因此失败是模型的崩溃,在我们的数据中,它们只发生在R >= a的地方。我们事后发现了这种模式,并在两项预注册的验证性研究中用43个新条件进行了测试:完整的界在所有条件中都成立,而简化的界仅对3个事实失败,每个都有R >= a。因此,微调的一阶模型可以在响应系数变化小于其到边界距离的事实上约束遗忘。
cs.AI / 146 / 2609.32819

Rank Collapse Is Recoverable, Growing $|Q|$ Is Not: Out-of-Sample Early Warning for Value Divergence in High-UTD Soft Actor-Critic

秩塌缩可恢复,$|Q|$ 增长不可恢复:高 UTD Soft Actor-Critic 中价值发散的样本外早期预警
Bu, Tianqi, Peng, YuXuan, Tu, Junteng, Xiao, Henghui
Abstract
Raising the update-to-data (UTD) ratio breaks off-policy critics in two ways grouped as "plasticity loss": collapsing representations and growing value magnitude $|Q|$. We separate them in Soft Actor-Critic (SAC) with scaled critics (width 2048, no normalization). Collapse is survivable: at UTD ratio 16, HalfCheetah critics with most units dormant keep learning, and the training guard, which stops runs whose loss or $|Q|$ explodes, never flags them. Within one high-UTD SAC configuration, runs start close together, and how far a critic's $\log_{10}|Q|$ has climbed by step 15k, its early growth, ranks the runs by how soon the guard flags them. At 15k, a flagged run's $|Q|$ sits a median of over a hundredfold below its flag level, yet the climb's rate already orders the flags (Harrell's C and out-of-sample AUC 0.78 on Walker2d, 0.98 on Ant, at UTD ratio 4). Dormancy does not. Aborting on this rate saves about a tenth of held-out Walker2d compute and stays net-positive live. A LayerNorm critic lowers the rate, removes the flag on Walker2d at UTD ratio 4 and lowers the return.
Chinese Translation
提高更新数据比(UTD)会以两种方式破坏离策略评论家,这两种方式被统称为“可塑性损失”:表示塌缩和价值幅度 $|Q|$ 增长。我们在带有缩放评论家(宽度2048,无归一化)的 Soft Actor-Critic (SAC) 中将它们分开。塌缩是可存活的:在 UTD 比率为16时,大多数单元休眠的 HalfCheetah 评论家继续学习,而当损失或 $|Q|$ 爆炸时停止运行的训练守卫从未标记它们。在一个高 UTD 的 SAC 配置内,运行开始时很接近,评论家的 $\log_{10}|Q|$ 在第15k步时爬升了多少(即其早期增长)根据守卫标记它们的快慢对运行进行排序。在第15k步时,一个被标记的运行的 $|Q|$ 中位数低于其标记水平一百倍以上,然而爬升率已经对标记进行了排序(在 UTD 比率为4时,Harrell's C 和样本外 AUC 在 Walker2d 上为0.78,在 Ant 上为0.98)。休眠则不能。根据这个速率中止可节省约十分之一的留出 Walker2d 计算量,并在实际运行中保持净收益为正。一个 LayerNorm 评论家降低了速率,在 UTD 比率为4时移除了 Walker2d 上的标记,并降低了回报。
cs.AI / 147 / 2609.32821

Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs

仅路由漂移不能诊断合并 MoE LLM 中的失败
Wang, Yuanyi, Gu, Yanggan, Lu, Su, Zhu, Guanghao, Wang, Pengkai, Yang, Yifan, Xie, Congkai, Yan, Zhaoyi, Wu, Jianmin, Yang, Hongxia
Abstract
Model merging efficiently combines specialized large language models (LLMs) without joint retraining, but can substantially alter expert routing in Mixture-of-Experts (MoE) models. Such \emph{routing drift} is often interpreted as routing failure, raising a fundamental question that remains unclear: \emph{does routing drift after MoE merging actually indicate routing failure, and what evidence should justify repair?} We investigate these questions across DeepSeekMoE, OLMoE, and Qwen3-MoE proposing a routing analysis toolkit for controlled counterfactual interventions and token-level analysis. By crossing source and merged router inputs and parameters, we attribute most expert reassignments to input shifts rather than parameter changes at the same layer. However, source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and different expert selections can produce directionally similar mixture outputs. We therefore operationalize routing failure as \textit{task loss recoverable under a specified routing intervention, with non-routing parameters fixed.} These tests detect recoverable loss under deliberate router corruption, whereas source-route restoration does not establish reliable task benefits in the evaluated merged models. Motivated by these, we propose \emph{Selective Router Repair (SRR)} as a case study, and find that source-specialist token-likelihood advantages do not reliably identify beneficial local corrections. Together, these findings show that \textbf{routing drift alone is insufficient evidence of routing failure}: source-informed corrections must be judged by their task-level intervention effects. The analysis toolkit and SRR code are released.
Chinese Translation
模型合并能够高效地组合专门化的大语言模型(LLM),而无需联合重训练,但可能显著改变混合专家(MoE)模型中的专家路由。这种\emph{路由漂移}通常被解释为路由失败,这引发了一个尚未明确的基本问题:\emph{MoE 合并后的路由漂移是否实际上表明路由失败,以及什么样的证据才能证明修复的合理性?}我们在 DeepSeekMoE、OLMoE 和 Qwen3-MoE 上研究了这些问题,并提出一个路由分析工具包,用于受控反事实干预和 token 级分析。通过交叉源路由器和合并路由器的输入与参数,我们将大多数专家重新分配归因于输入偏移,而非同一层的参数变化。然而,源相对路由差异很难预测源路由恢复带来的下一 token 似然增益,并且不同的专家选择可以产生方向相似的混合输出。因此,我们将路由失败操作化为\textit{在指定路由干预下可恢复的任务损失,且非路由参数固定。}这些测试可以检测到故意路由器损坏下的可恢复损失,而在所评估的合并模型中,源路由恢复并不能确立可靠的任务收益。受此启发,我们提出\emph{选择性路由器修复(SRR)}作为案例研究,并发现源专家 token 似然优势不能可靠地识别有益的局部修正。总之,这些发现表明,\textbf{仅路由漂移不足以作为路由失败的证据}:源信息驱动的修正必须根据其任务级干预效果来评判。分析工具包和 SRR 代码已发布。
cs.AI / 148 / 2609.32825

The Decomposition Tax: LLM Pipelines Lose Up to 40 Accuracy Points at Their Own Interfaces

分解税:LLM 流水线在自身接口处损失高达 40 个准确率点
Bu, Tianqi, Peng, YuXuan, Tu, Junteng, Xiao, Henghui
Abstract
A four-stage LLM pipeline gives up as much as 40.5 accuracy points at its own interfaces (gemma-3-12B on MATH-500, Holm-corrected p = 1.66e-19; the largest tax in the primary family). We hold model, problem, stages, stage prompts and completion budget fixed, vary only whether each stage can still see the original problem, and call the accuracy difference the decomposition tax. Across 21 open-weight models from nine organisations, on GSM-Hard and MATH-500 at n = 200 paired items per cell, 70 of 118 primary-family tests survive Benjamini-Hochberg correction and 54 survive Holm. On GSM-Hard, a placebo recovers nothing: it carries at least 60% of the extra tokens and at most one word of the problem. Builders design a pipeline one stage at a time, and its bill arrives at the interfaces between stages. Rewriting one stage's instruction moves gemma-3-12B's tax from 4.5 to 36.5 points, and adding "every relationship stated between them" to a stage that lists the numerical quantities lowers the tax on 9 of 9 models on MATH-500. Re-grounding, which shows a stage the original problem again, belongs after the loss. With one lossy interface, re-grounding the stage after it beats re-grounding the stage before it on 7 of 7 models on both benchmarks; on MATH-500 the earlier repair is worse than none on 7 of 7. Newer models still pay: gemma-4-12B gives up 37.0 points, and the repair holds on all three of the newest models we test. A sealed held-out test refuted a stronger rule we registered, which predicted the paying stage from the interface and receiver types, so we locate the tax by measuring one stage at a time. The prescription has two parts: re-ground the stage after the lossy interface, and if a stage must list the quantities, tell it to keep the relationships.
Chinese Translation
一个四阶段 LLM 流水线在其自身接口处损失高达 40.5 个准确率点(gemma-3-12B 在 MATH-500 上,Holm 校正 p = 1.66e-19;这是主要家族中最大的税)。我们固定模型、问题、阶段、阶段提示和完成预算,仅改变每个阶段是否仍能看到原始问题,并将准确率差异称为分解税。在来自九个组织的 21 个开放权重模型上,在 GSM-Hard 和 MATH-500 上,每个单元格 n = 200 个配对项目,118 个主要家族测试中有 70 个通过 Benjamini-Hochberg 校正,54 个通过 Holm 校正。在 GSM-Hard 上,安慰剂没有恢复任何东西:它携带至少 60% 的额外 token 和至多一个原问题中的词。构建者一次设计一个阶段的流水线,而其账单在阶段之间的接口处到来。重写一个阶段的指令将 gemma-3-12B 的税从 4.5 点移动到 36.5 点,而向列出数值量的阶段添加“它们之间陈述的每个关系”在 MATH-500 上使 9 个模型中的 9 个降低了税。重新接地(re-grounding),即再次向一个阶段展示原始问题,属于损失之后。在一个有损接口的情况下,重新接地其后的阶段在 7 个模型中的 7 个上在两个基准上都优于重新接地其前的阶段;在 MATH-500 上,较早的修复在 7 个模型中的 7 个上比不修复更差。较新的模型仍然付出代价:gemma-4-12B 损失 37.0 点,并且该修复在我们测试的所有三个最新模型上都成立。一个密封的留出测试驳斥了我们注册的一个更强规则,该规则从接口和接收者类型预测付出代价的阶段,因此我们通过一次测量一个阶段来定位税。处方有两个部分:在有损接口之后重新接地该阶段,并且如果一个阶段必须列出数量,告诉它保持关系。
cs.AI / 149 / 2609.32827

Improving LLM Collaboration via Multi-Agent Preference Learning

通过多智能体偏好学习改进LLM协作
Liu, Shuo, Li, Xinzichen, Chen, Tianle, Amato, Christopher
Abstract
Several works have explored multi-agent reinforcement learning (MARL) in LLM collaboration. However, constructing reliable rewards is difficult in practice, as complete and accurate metrics are often unavailable and hard to aggregate. Preference learning provides an alternative by learning from comparative human or AI feedback. Yet, its extension to multi-agent systems remains underexplored. To address this gap, we formulate preference-based multi-agent systems (MAS) from decentralized and centralized collaboration perspectives. We also introduce a general multi-agent preference learning framework (MAPL) to solve these problems. MAPL allows iterative updates by comparing the current solution with decentralized or centralized solutions generated by various agents. We instantiate MAPL using MARL from human feedback (MARLHF) with a learned reward model and multi-agent direct preference optimization (MADPO). Experiments on collaborative writing, coding, tool use, and travel planning show that MAPL can improve collaboration quality and efficiency while approaching the performance of MARL with fixed, well-defined rewards. Within MAPL, MARLHF generally outperforms MADPO on most tasks but remains sensitive to data coverage, agent and comparator models, and the underlying MARL algorithms.
Chinese Translation
已有若干工作探索了LLM协作中的多智能体强化学习(MARL)。然而,在实践中构建可靠的奖励是困难的,因为完整且准确的指标往往不可用且难以聚合。偏好学习通过从比较性的人类或AI反馈中学习提供了一种替代方案。然而,其向多智能体系统的扩展仍然未被充分探索。为弥补这一空白,我们从去中心化和中心化协作的角度形式化了基于偏好的多智能体系统(MAS)。我们还引入了一个通用的多智能体偏好学习框架(MAPL)来解决这些问题。MAPL允许通过将当前解决方案与各种智能体生成的去中心化或中心化解决方案进行比较来进行迭代更新。我们使用来自人类反馈的MARL(MARLHF)和学习的奖励模型以及多智能体直接偏好优化(MADPO)来实例化MAPL。在协作写作、编码、工具使用和旅行规划上的实验表明,MAPL可以提高协作质量和效率,同时接近具有固定、明确定义奖励的MARL的性能。在MAPL中,MARLHF在大多数任务上通常优于MADPO,但仍然对数据覆盖率、智能体和比较器模型以及底层的MARL算法敏感。
cs.AI / 150 / 2609.32835

FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors

FinancialAuditBench:在差分隐私下利用真实世界先验的基准构建
Huang, Jerry, Babu, Sarvesh, Van Buren, Matt, Wang, Alexander, Pillai, Pranav, Jain, Arush, Burton, James P., Hockenmaier, Julia
Abstract
As AI agents are becoming widely adopted in the financial services industry, careful measurement is essential to understand where they can be reliably deployed and where oversight and professional review remain necessary. Such measurement, however, is constrained by limited access to proprietary or privacy-sensitive data. Existing benchmarks therefore often rely on publicly available data, human- and/or LLM-authored tasks, or simplified settings. We introduce FinancialAuditBench, a benchmark for evaluating agents on financial statement audit tasks, along with a framework for systematically generating synthetic engagements. Our task generation framework leverages differentially private aggregate statistics from historical audits along with audit expertise contributed through over 1,100 hours of benchmark development and review. FinancialAuditBench consists of 90 tasks spanning workpaper completion and review across six synthetic audit engagements, each containing an average of 179 files. Evaluation on eleven frontier models shows that while agents complete substantial portions of staff-level audit tasks well, they sometimes perform inappropriate procedures or produce incorrect documentation. Beyond financial auditing, our framework offers an approach for systematically generating synthetic tasks for model evaluation and training in privacy-sensitive domains.
Chinese Translation
随着AI代理在金融服务行业被广泛采用,仔细的测量对于理解它们可以可靠部署在哪里以及哪些地方仍需监督和专业审查至关重要。然而,这种测量受到对专有或隐私敏感数据有限访问的限制。现有的基准因此通常依赖于公开可用的数据、人类和/或LLM编写的任务,或者简化设置。我们介绍了FinancialAuditBench,一个用于评估财务报表审计任务代理的基准,以及一个系统生成合成审计业务的框架。我们的任务生成框架利用了来自历史审计的差分隐私聚合统计以及通过超过1,100小时的基准开发和审查贡献的审计专业知识。FinancialAuditBench包含90个任务,涵盖六个合成审计业务的工作底稿完成和审查,每个平均包含179个文件。对十一个前沿模型的评估显示,虽然代理很好地完成了相当部分员工级别的审计任务,但它们有时执行不适当的程序或产生不正确的文档。除了财务审计之外,我们的框架提供了一种在隐私敏感领域系统生成合成任务用于模型评估和训练的方法。
cs.AI / 151 / 2609.32870

Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning

面向基于证据推理的反事实自演化智能体
Han, Xing, Wang, Yuxin, Chen, Chen, Dai, Wei, Gudur, Gautham Krishna, Li, Shijun, Chung, Hsing-Huan, Hager, Gregory D., Ghosh, Joydeep, Liang, Paul Pu, Saria, Suchi
Abstract
Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verified. We introduce counterfactual self-evolution, which generates counterfactual context for reconsidering the original case. A trainable Proposer constructs targeted evidence edits and describes potential outcome changes with causal explanations. We handcraft an expert-verified counterfactual instruction-tuning dataset to teach the Proposer to generate high-quality counterfactuals across a broad range of action--outcome scenarios. Each counterfactual instruction-tuning example specifies an edit within a defined category and explains its hypothesized causal effect on the decision, teaching the Proposer to reason systematically about what changes and why. We instruction-tune the Proposer on these examples, then formulate a fine-tuning reward that integrates feedback from the Solver and Verifier. Across diverse counterfactual scenarios, this reward favors high-quality counterfactuals and warranted revisions, while penalizing changes that overturn correct decisions. The counterfactual context aims to correct errors and strengthen confidence in correct decisions. Accepted counterfactuals accumulate in memory that supplies in-context evidence to the frozen Solver; the Solver adapts through evolving context rather than weight updates. We apply the framework to clinical reasoning, fact verification, and business reasoning. Our evaluation tracks performance over successive rounds as counterfactual memory grows, including transfer to harder cases. Our method achieves superior results across diverse frontier models.
Chinese Translation
自对弈的提议者-求解器方法通过生成任务并从已验证的解决方案中学习来提升推理能力。然而,对于证据可识别任务,即案例特定证据和领域知识决定了一个可检查答案的任务,自对弈需要生成合理的案例,其答案能够被独立验证。我们引入了反事实自演化,它为重新考虑原始案例生成反事实上下文。一个可训练的提议者(Proposer)构建有针对性的证据编辑,并用因果解释描述潜在的结果变化。我们手工制作了一个由专家验证的反事实指令微调数据集,以教导提议者在广泛的动作-结果场景中生成高质量的反事实。每个反事实指令微调示例指定了一个定义类别内的编辑,并解释了其对决策的假设因果效应,教导提议者系统地推理什么发生了变化以及为什么。我们在这些示例上对提议者进行指令微调,然后制定一个微调奖励,整合来自求解器(Solver)和验证器(Verifier)的反馈。在不同的反事实场景中,该奖励倾向于高质量的反事实和合理的修订,同时惩罚推翻正确决策的改变。反事实上下文旨在纠正错误并增强对正确决策的信心。被接受的反事实会累积在记忆中,为冻结的求解器提供上下文内证据;求解器通过演化的上下文而非权重更新来适应。我们将该框架应用于临床推理、事实验证和商业推理。我们的评估跟踪随着反事实记忆增长,在连续轮次中的性能,包括对更难案例的迁移。我们的方法在不同的前沿模型中取得了优越的结果。
cs.AI / 152 / 2609.32885

Can LLMs Predict the Future? A Brier Score Analysis of Prediction Markets

大语言模型能预测未来吗?预测市场的Brier分数分析
Li, Yuanbo, Li, Zekun, cong, Xiaoyan
Abstract
We study whether model upgrades improve probability estimates for prediction-market questions. Our Resolved Market Forecasting (RMF) benchmark contains 3,000 resolved binary questions across nine domains, on which we evaluate six Claude and Qwen model variants using a question-only, zero-shot protocol. We assess Brier scores relative to an empirical base-rate predictor, examine their Murphy decomposition, and compare models through paired differences, with results stratified by event category and timing relative to training cutoffs. In the reported post-cutoff stratum, the four Claude models achieve Brier scores of 0.183-0.192, improving on the base-rate reference by 0.024-0.033. Qwen 32B does not significantly outperform that reference, although its paired Brier is 0.024 lower than that of the 7B checkpoint. The evaluated Claude version and tier upgrades yield no significant improvement. Within-model differences across event categories exceed the observed differences among Claude variants. These results show why forecasting scores should be interpreted alongside simple probability baselines and question composition: under this protocol, newer versions or higher model tiers do not consistently produce more accurate probabilities.
Chinese Translation
我们研究模型升级是否能改善预测市场问题的概率估计。我们的已解决市场预测(RMF)基准包含九个领域的3000个已解决的二元问题,我们使用仅问题、零样本协议评估了六个Claude和Qwen模型变体。我们评估相对于经验基率预测器的Brier分数,检查其Murphy分解,并通过配对差异比较模型,结果按事件类别和相对于训练截止的时间分层。在报告的截止后分层中,四个Claude模型达到0.183-0.192的Brier分数,比基率参考改善了0.024-0.033。Qwen 32B没有显著优于该参考,尽管其配对Brier比7B检查点低0.024。评估的Claude版本和层级升级没有产生显著改进。模型内跨事件类别的差异超过了观察到的Claude变体之间的差异。这些结果表明为什么预测分数应该与简单概率基线和问题组成一起解释:在该协议下,更新的版本或更高的模型层级并不总是产生更准确的概率。
cs.AI / 153 / 2609.32886

StraTune: Adaptive Selection of Revision Operators for Self-Evolving LLM Skills

StraTune:面向自进化LLM技能的自适应修订算子选择
Liu, Zeping, Li, Yan, Lao, Ni, Wolff, Gil, Mai, Gengchen
Abstract
Large language models (LLMs) can learn reusable textual skills from execution feedback without updating their parameters, but effectively deciding how to revise these skills remains a key challenge. Existing methods typically rely on a fixed revision operator, a search strategy and the revision forms applied under it. However, we observe that no single revision operator consistently performs best across tasks, and repeatedly applying an unsuitable operator can limit further improvement. We propose StraTune (strategy-guided skill tuning), which lets a frozen optimizer LLM choose the revision operator at every round from the optimization state, which is defined as the current execution feedback together with the recorded outcomes of earlier strategies and forms. Candidate skills from every revision operator pass one candidate evaluation, which screens for gains and regressions on a small sample set and validates them on a larger one, and every outcome is written back to the optimization state for later choices. Across four benchmarks and two LLM settings, StraTune outperforms all five baselines in most settings. Ablations attribute the gains to the adaptive choice of the revision operator, since fixed, random, scheduled, and bandit strategy choices all score lower, and skills learned with a small target LLM also improve a stronger one. Code and learned skills are available at https://github.com/seai-lab/StraTune.
Chinese Translation
大语言模型(LLM)能够在不更新其参数的情况下,从执行反馈中学习可复用的文本技能,但如何有效决定如何修订这些技能仍是一个关键挑战。现有方法通常依赖固定的修订算子、一种搜索策略以及在该策略下所采用的修订形式。然而,我们观察到,没有任何单一修订算子能在所有任务上始终表现最佳,并且反复应用不合适的算子会限制进一步改进。我们提出 StraTune(策略引导的技能调优,strategy-guided skill tuning),它让一个冻结的优化器 LLM 在每一轮根据优化状态选择修订算子;优化状态由当前执行反馈以及先前策略和形式的记录结果共同定义。来自每个修订算子的候选技能都要经过一次候选评估,该评估在小样本集上筛查提升与退化,并在较大样本集上对其进行验证;每个结果都会回写到优化状态中,用于后续选择。在四个基准和两种 LLM 设置下,StraTune 在大多数设置中优于所有五个基线。消融实验将增益归因于修订算子的自适应选择,因为固定、随机、按计划以及 bandit 策略选择得分都更低,并且用小型目标 LLM 学到的技能也能提升更强的 LLM。代码和学到的技能可在 https://github.com/seai-lab/StraTune 获取。
cs.AI / 154 / 2609.32900

Constraints Are Graphs, Not Chains: Exact Decoding for Diffusion Language Models

约束是图,而非链:扩散语言模型的精确解码
Su, Jianchang, Zhang, Wei
Abstract
Diffusion language models (dLLMs) predict masked positions in arbitrary order, but their exact constrained decoders still encode constraints as sequential languages, whose state must track every unresolved dependency between positions. For relational constraints this encoding grows exponentially: for same-order copy, every finite automaton needs $4^k$ states, deterministic or nondeterministic, and every context-free grammar has size $2^{\Omega(k)}$, while the factor graph of the same relation has size $O(k)$ and a 16-entry peak table. We introduce FactorDLM, a training-free decoder that represents finite-domain relations as a factor graph and, at each denoising step, conditions the model's mean-field prediction on that graph exactly by variable elimination. Decoding cost then grows exponentially with the induced width of the constraint graph, which replaces automaton size as the governing parameter. Because a finite automaton is a chain-shaped factor graph, one compiler enforces syntax and nonlocal relations together: on JSON records with cross-field references, a schema automaton alone leaves references dangling, relational factors alone produce malformed JSON, and the combined plan is valid on both counts, including on records of variable length. Across nine relational benchmarks and three backbones, every output satisfies every declared constraint at 0.4-6.9% projection overhead, where unconstrained decoding is 0-79% valid, and compiled projection answers repeated queries 13.6x faster than CP-SAT with eight parallel workers. Because model-free rules solve three of five standard benchmarks, we construct benchmarks with exact chance and fixed-template floors, on which selecting among exact constrained samples beats greedy projection. Which encoding is cheaper, sequential state or direct factors, depends on the constraint and is computable before decoding begins.
Chinese Translation
扩散语言模型(dLLMs)以任意顺序预测掩码位置,但其精确约束解码器仍将约束编码为序列语言,其状态必须跟踪位置之间每一个未解决的依赖关系。对于关系约束,这种编码呈指数增长:对于同序复制,每个有限自动机(确定性或非确定性)都需要 $4^k$ 个状态,每个上下文无关文法的大小为 $2^{\Omega(k)}$,而同一关系的因子图大小为 $O(k)$,并且有一个16项峰值表。我们介绍了 FactorDLM,这是一种免训练解码器,它将有限域关系表示为因子图,并在每个去噪步骤中通过变量消元使模型的平均场预测精确地条件依赖于该图。于是,解码成本随约束图的诱导宽度呈指数增长,诱导宽度取代自动机大小成为主导参数。由于有限自动机是链状因子图,一个编译器可以同时强制执行语法和非局部关系:在具有跨字段引用的 JSON 记录上,仅使用模式自动机会使引用悬空,仅使用关系因子会产生格式错误的 JSON,而组合方案在两方面都有效,包括对可变长度记录也有效。在九个关系基准和三个骨干模型上,每个输出都以 0.4-6.9% 的投影开销满足每个声明的约束,而无约束解码的有效率为 0-79%,并且编译后的投影在回答重复查询时比使用八个并行工作器的 CP-SAT 快 13.6 倍。由于无模型规则解决了五个标准基准中的三个,我们构建了具有精确随机和固定模板下限的基准,在这些基准上,在精确约束样本中进行选择优于贪婪投影。哪种编码更便宜,序列状态还是直接因子,取决于约束,并且可以在解码开始之前计算。
cs.AI / 155 / 2609.32907

Logical subspace in LLMs

LLMs中的逻辑子空间
Kean, Hope, Boix-Adsera, Enric
Abstract
Recent work has identified a human brain network specialized for abstract formal reasoning (Kean et al., 2025). Does the same hold true in language models? To answer this question, we introduce the minimal viable subspace (MVS) method, which searches for the lowest-rank activation subspace at a layer that preserves task performance when everything outside that subspace is ablated. Using MVS, we demonstrate low-rank subspaces supporting logical inference on Gemma and Qwen models. Furthermore, these subspaces exhibit a clear dissociation from model capacities on other tasks, such that retaining these late logic subspaces preserves inference while impairing factual knowledge, working memory, cognitive control, and arithmetic. Conversely, ablating them reduces logical inference accuracy to chance while largely sparing these other capacities. Our results suggest a functionally localizable core machinery for logic akin to that in the human brain.
Chinese Translation
近期研究识别出一个专门用于抽象形式推理的人脑网络(Kean et al., 2025)。语言模型中是否也存在类似情况?为回答这一问题,我们提出了最小可行子空间(minimal viable subspace, MVS)方法,该方法在某一层中搜索最低秩的激活子空间,使得当该子空间之外的一切都被消融时,任务性能仍得以保持。利用MVS,我们在Gemma和Qwen模型上展示了支持逻辑推理的低秩子空间。此外,这些子空间与模型在其他任务上的能力表现出明显分离:保留这些晚期逻辑子空间可保持推理能力,同时损害事实知识、工作记忆、认知控制和算术能力。相反,消融它们会将逻辑推理准确率降至随机水平,而基本不影响这些其他能力。我们的结果表明,存在一种在功能上可定位的逻辑核心机制,类似于人脑中的相应机制。
cs.AI / 156 / 2609.32917

Planner-as-Router: Joint Plan-Time Model Routing for Cost-Efficient Multi-Agent Workflows

Planner-as-Router:面向成本高效多智能体工作流的联合计划时模型路由
Singh, Vivek Kumar, Priyam, Preeti, Bhowmick, Gautam
Abstract
Running large language model (LLM) agents in production gets expensive fast. A frontier model (the largest, most capable tier) is accurate but can cost 25 times what a small model costs per token, and the gap compounds once a workflow chains several calls together. Planner-as-Router (PaR) attacks this from a different angle. Instead of leaving model-tier selection to some component downstream, it folds the choice into planning itself. As the planner breaks a query into subtasks, it also assigns each one a model size tier (small, mid, or frontier, ordered by capability and price), so the dependencies between subtasks are visible before any specialist runs. Unlike per-call routers such as cascade routing, which look at one node at a time, PaR sees the whole workflow up front and needs no separate router model or training data. We evaluate PaR with EntBench, a benchmark of 54 enterprise agentic tasks across seven classes, graded by actually running the generated Structured Query Language (SQL) and MongoDB queries against live databases. Over 1,157 evaluations spanning eight routers and three seeds, PaR stays on the observed cost-accuracy frontier. It matches a sink-frontier heuristic (frontier model on terminal nodes only) in accuracy at comparable cost and a faithful FrugalGPT cascade at lower cost, and cuts cost 44% against all-frontier routing while giving up 2.9 points of accuracy. Several accuracy gaps fall inside the plus-or-minus six-point confidence interval of a 54-task study, so we frame PaR's advantage as frontier position rather than a clean accuracy win. We also report a preliminary observation, not a validated result: a small pilot hints that cheap routing may carry a hidden compounding penalty on compositional workflows, which we frame as a hypothesis for future measurement. PaR, EntBench, and all evaluation code are open source.
Chinese Translation
在生产环境中运行大型语言模型(LLM)智能体的成本会迅速变得高昂。前沿模型(最大、能力最强的层级)准确但每token成本可能是小模型的25倍,而一旦工作流串联多个调用,差距会进一步扩大。Planner-as-Router (PaR) 从不同角度解决这一问题。它将模型层级选择融入规划本身,而不是留给下游某个组件。当规划器将查询分解为子任务时,它还为每个子任务分配一个模型大小层级(小、中、前沿,按能力和价格排序),这样在任何专家模型运行之前,子任务之间的依赖关系就可见了。与逐调用路由器(如级联路由)一次只看一个节点不同,PaR 能预先看到整个工作流,且不需要单独的路由器模型或训练数据。我们用 EntBench 评估 PaR,这是一个包含七类54个企业智能体任务的基准,通过实际对实时数据库运行生成的结构化查询语言(SQL)和 MongoDB 查询来评分。在跨越八种路由器和三个种子的 1,157 次评估中,PaR 始终处于观察到的成本-准确率前沿。它在准确率上与 sink-frontier 启发式(仅在终端节点使用前沿模型)相当,成本相近,并以更低成本匹配忠实的 FrugalGPT 级联,同时相对于全前沿路由降低 44% 的成本,但准确率下降 2.9 个百分点。若干准确率差距落在 54 任务研究的 ±6 个百分点置信区间内,因此我们将 PaR 的优势表述为前沿位置,而非明确的准确率胜利。我们还报告了一个初步观察(非验证结果):一个小型试点提示,廉价路由可能在组合工作流上带来隐藏的复合惩罚,我们将其作为未来测量的假设。PaR、EntBench 及所有评估代码均开源。
cs.AI / 157 / 2609.32922

Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer

按需精度:随机计算是一种密集自适应量化器
Jin, Haoran, Zhang, Kangqi, Yang, Jirong, Lyu, Barry, Ding, Qiuyi, Gao, Ruijie, Bleier, Nathan
Abstract
Matrix multiplications dominate the inference cost of modern transformer-based vision models, yet existing efficiency techniques such as post-training quantization and mixed-precision inference are largely limited to the small set of fixed-width formats (INT4, INT8, BF16, and FP16) supported by conventional accelerators. We revisit stochastic computing (SC) as a way to lift this constraint: viewed as a dense adaptive quantizer, SC controls precision by bit-stream length L rather than a fixed datapath, while each multiplication reduces to a single AND/XNOR gate. We build a GPU library that emulates SC matrix multiplication at scale, exposes stream lengths as first-class kernel arguments, and evaluates SC end-to-end on image classification, object detection and instance segmentation, class-conditional image generation, and visual world-model planning. On top of this substrate, we develop a dynamic per-row mixed-precision policy that assigns stream length per token or group at matched average budget, requires no retraining, and uses the same SC hardware across schedules. Across tasks, SC remains competitive with fixed-format INT quantization at matched bit budgets, while per-row mixed precision helps maintain accuracy at lower average stream lengths. These results provide software-level feasibility evidence that SC can serve as a dense-precision substrate for fine-grained mixed-precision inference on modern vision transformers.
Chinese Translation
矩阵乘法在现代基于Transformer的视觉模型的推理成本中占主导地位,然而现有的效率技术(如训练后量化和混合精度推理)在很大程度上受限于传统加速器所支持的小型固定宽度格式集(INT4、INT8、BF16和FP16)。我们重新审视随机计算(SC)作为解除这一限制的途径:被视为一种密集自适应量化器,SC通过比特流长度L而非固定数据通路来控制精度,同时每次乘法简化为单个与/同或门。我们构建了一个GPU库,可大规模模拟SC矩阵乘法,将流长度作为内核的一等参数暴露,并在图像分类、目标检测与实例分割、类别条件图像生成以及视觉世界模型规划上端到端地评估SC。在此基底之上,我们开发了一种动态逐行混合精度策略,该策略在匹配的平均预算下为每个token或组分配流长度,无需重新训练,并在不同调度下使用相同的SC硬件。在各种任务中,SC在匹配的比特预算下与固定格式的INT量化相比仍具竞争力,而逐行混合精度有助于在更低的平均流长度下保持精度。这些结果提供了软件层面的可行性证据,表明SC可以作为现代视觉Transformer上细粒度混合精度推理的密集精度基底。
cs.AI / 158 / 2609.32923

TRACE: Learning to Self-Calibrate Wireless Digital Twins from ISAC Measurements

TRACE:从ISAC测量中学习自校准无线数字孪生
Masrur, Saad, Khosravirad, Saeed R., Guvenc, Ismail
Abstract
Wireless digital twins (DTs) rely on 3D environment models to predict radio propagation and support wireless-network decisions, yet these models are often initialized from imperfect 3D maps. Errors in building position, height, footprint, and orientation can therefore cause a high-fidelity propagation engine to simulate the wrong physical environment. In this paper, we study how a deployed wireless network can repair an existing DT using its own radio frequency (RF) measurements. In particular, we introduce Twin Residual Alignment and Calibration Engine (TRACE), a physics-grounded learning-based self-calibration framework that treats twin maintenance as residual alignment between the physical world and the current DT. Using the same sensing configuration as the physical measurements, TRACE ray-traces the current DT, coherently backprojects the measured and simulated RF onto a common world grid, and extracts the same local region around each building's current DT position. A multi-view corrector then fuses evidence across sensing nodes and neighboring buildings to predict a gated six-parameter correction per building, without relying on absolute layout or sensor ordering, and supports iterative correction through re-rendering. On 5,400 held-out samples from unseen simulated scenes at 28 GHz, TRACE reduces 3D position RMSE from 2.202 m to 0.302 m and yaw RMSE from 4.978{\deg} to 0.894{\deg}, outperforming ViT and U-Net baselines under changes in layout, building count, sensing-node count, and SNR. On measured 28 GHz RF data from the NIST outdoor courtyard, a model trained only on synthetic RF reduces mean planar wall-position error from 1.00 m to 7.8 cm, without measured-data fine-tuning or geometric labels. These results show that the discrepancy between measured and twin-rendered RF can serve as a learning signal for repairing a wireless DT.
Chinese Translation
无线数字孪生(DTs)依赖3D环境模型来预测无线电传播并支持无线网络决策,然而这些模型通常从不完美的3D地图初始化。因此,建筑物位置、高度、占地面积和方向上的误差可能导致高保真传播引擎模拟错误的物理环境。在本文中,我们研究已部署的无线网络如何利用自身的射频(RF)测量来修复现有的DT。具体而言,我们提出孪生残差对齐与校准引擎(TRACE),一个基于物理的、基于学习的自校准框架,将孪生维护视为物理世界与当前DT之间的残差对齐。使用与物理测量相同的传感配置,TRACE对当前DT进行光线追踪,将测量和模拟的RF相干反向投影到公共世界网格上,并提取每个建筑物当前DT位置周围的相同局部区域。然后,一个多视角校正器融合跨传感节点和相邻建筑物的证据,以预测每个建筑物的门控六参数校正,而不依赖绝对布局或传感器顺序,并通过重新渲染支持迭代校正。在28 GHz下从未见过的模拟场景中的5,400个留存样本上,TRACE将3D位置RMSE从2.202 m降低到0.302 m,将偏航角RMSE从4.978°降低到0.894°,在布局、建筑物数量、传感节点数量和SNR变化的情况下优于ViT和U-Net基线。在来自NIST户外庭院的实测28 GHz RF数据上,仅使用合成RF训练的模型将平均平面墙壁位置误差从1.00 m降低到7.8 cm,无需实测数据微调或几何标签。这些结果表明,测量与孪生渲染的RF之间的差异可以作为修复无线DT的学习信号。
cs.AI / 159 / 2609.32924

Diagnosing Sampled LLM Reasoning in Formal Geometry: Coverage, Realization, and Validity Evidence

诊断形式几何中采样的大语言模型推理:覆盖、实现与有效性证据
Yue, Xiao, Qu, Guangzhi
Abstract
Repeated sampling can reveal a correct numerical answer without yielding either a reliable system output or a supported derivation. We present Coverage, Realization, and Validity Evidence (CRV), an evaluation protocol for sampled large language model (LLM) reasoning over formal geometry states. Coverage is answer availability, realization is readout accuracy on the frozen candidate pool, and validity evidence is a label-blinded critic judgment of derivational support rather than a proof certificate. CRV freezes each candidate pool before comparing readouts and analyzes covered failures by correct-answer multiplicity and within-problem discrimination. On HardShift441, a 441-problem set for which a reference solver leaves 406 problems unsolved, a LoRA-adapted Qwen2.5-7B generator obtains 24.2% average single-sample accuracy and 68.9% pass@16, whereas verifier-weighted self-consistency (WSC) reaches 38.0%. Readout accuracy is particularly low when the correct answer occurs only once or twice in the pool. In a separate constructed audit of 195 covered problems, the critic labels 12 correct-answer representatives as supported, 181 as refuted, and two as uncertain. These results show that coverage, realization, and validity evidence from the critic are distinct quantities and should be reported separately.
Chinese Translation
重复采样可以揭示正确的数值答案,而既不能产生可靠的系统输出,也不能提供有支持的推导。我们提出覆盖、实现与有效性证据(CRV),一种针对形式几何状态上采样的LLM推理的评估协议。覆盖是答案的可用性,实现是在冻结候选池上的读出准确率,有效性证据是一种标签盲的批评者判断,判断推导支持而非证明证书。CRV在比较读出之前冻结每个候选池,并通过正确答案多重性和问题内区分度来分析覆盖的失败。在HardShift441(一个包含441个问题的集合,参考求解器对406个问题未解决)上,LoRA适配的Qwen2.5-7B生成器获得24.2%的平均单样本准确率和68.9%的pass@16,而验证器加权自洽性(WSC)达到38.0%。当正确答案在池中仅出现一次或两次时,读出准确率特别低。在一项针对195个覆盖问题的单独构建审计中,批评者将12个正确答案代表标记为支持,181个为反驳,两个为不确定。这些结果表明,覆盖、实现和来自批评者的有效性证据是不同的量,应分别报告。
cs.AI / 160 / 2609.32964

The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining

承诺-弃权回路:为什么语言模型会幻觉而不是弃权
Nguyen, Vy, Xu, Ziqi, Chan, Jeffrey, He, Estrid, Xia, Feng, Luo, Renqiang, Cambria, Erik, Zhang, Xiuzhen
Abstract
Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to answer reliably. A large body of existing work mitigates hallucination through detection or abstention mechanisms, but leaves open how models internally arrive at the decision to commit or abstain in the first place. We study this decision through mechanistic analysis, framing hallucination as unsupported commitment: the model commits despite exhibiting signals of unanswerability. Using causal gating, we identify a Commit-Abstain Circuit (CAC), a sparse, causally localised subset of attention heads and MLP sublayers underlying this decision. Across ten LMs (3B-14B) from five families and three benchmarks, the CAC exhibits a recurring accumulate-yet-undercorrect pattern: commitment-promoting components build up commitment in earlier layers, while abstention-promoting components act later as corrective signals that are often insufficient to overturn the accumulated commitment. Building on this finding, a lightweight policy trained on CAC activations improves decision accuracy by 12.2 points over the model's intrinsic commit-abstain margin, reduces false abstentions by 2.5 times, transfers to unseen benchmarks, and extends to larger models (27B-35B). The CAC is both diagnostic, clarifying how models overcommit, and practical, enabling improved abstention decisions.
Chinese Translation
语言模型(LMs)常常通过做出自信的答案而不是弃权来产生幻觉,即使它们没有足够的信息来可靠地回答。大量现有工作通过检测或弃权机制来缓解幻觉,但留下了模型内部最初如何做出承诺或弃权决定的问题。我们通过机制分析研究这一决策,将幻觉定义为无支持的承诺:模型尽管表现出无法回答的信号,仍然做出承诺。使用因果门控,我们识别出一个承诺-弃权回路(Commit-Abstain Circuit, CAC),这是支持该决策的注意力头和MLP子层的一个稀疏、因果局部化的子集。在来自五个家族的十个LM(3B-14B)和三个基准测试中,CAC表现出一种反复出现的“累积但校正不足”模式:促进承诺的组件在较早层中累积承诺,而促进弃权的组件在较晚层作为纠正信号,但往往不足以推翻已累积的承诺。基于这一发现,在CAC激活上训练的轻量级策略将决策准确性比模型固有的承诺-弃权边际提高了12.2个百分点,将错误弃权减少了2.5倍,可迁移到未见过的基准测试,并扩展到更大的模型(27B-35B)。CAC既具有诊断性,阐明了模型如何过度承诺,又具有实用性,能够改进弃权决策。
cs.AI / 161 / 2609.32965

Relic: From Multi-Agent Collaboration to Persistent Organizational Capability

Relic:从多智能体协作到持久性组织能力
Du, Hongyi, Zhang, Tianyi, Zhang, Weijia, Yang, Yi, Yu, Haofei, Zhu, Kunlun, Dai, Tianxiang, Jiang, Shang, Gao, Zhelun, Pei, Jiaxin, Zhu, Shang, You, Jiaxuan
Abstract
Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the lesson continue to govern the team? We introduce Relic, which turns recurring collaboration failures into organization-owned, executable protocols. Members reflect on visible work, propose rules, and govern their adoption. Adopted protocols bind triggers, responsibilities, required evidence, and execution consequences to the runtime, while remaining open to revision and retirement. In one traced case, repeated integration friction produces an interface-review rule that governs later pull requests and is revised as work continues. Across 360 controlled runs over ten software workloads and three models, Relic raises complete-contract delivery from 14.06% to 19.76% (+5.71 percentage points) over a matched structured team without the protocol lifecycle, improving all four verified production endpoints in every model stratum. Under fresh-member transfer, behavioral correctness is 25.4% with no inherited protocol, 34.6% with the same rules provided as readable text, and 41.2% with executable bindings, a +6.5-point advantage over text alone. On the full CooperBench benchmark, after excluding broken benchmark pairs, Relic achieves 367/477 (76.9%), establishing the best reported result among peer-structured systems. On the fixed 48-pair same-model subset, Relic also exceeds Solo (29/48 vs. 26/48), reversing the coordination loss exhibited by the official peer baseline. Together, these results show how collaboration experience can become persistent organizational state that remains useful beyond the members who created it.
Chinese Translation
多个智能体在组织中可能经常发生冲突:例如,一个编码智能体更改了仓库中的接口,但另一个智能体继续在旧版本上开发,导致现有测试变得过时。一次对话可以解决该事件,但当参与者更换后,是什么让经验教训继续约束团队?我们提出Relic,它将反复出现的协作失败转化为组织拥有的、可执行的协议。成员反思可见的工作,提出规则,并管理其采纳。被采纳的协议将触发器、职责、所需证据和执行后果绑定到运行时,同时仍然可以修订和废止。在一个追踪案例中,反复出现的集成摩擦产生了一条接口审查规则,该规则约束后续的拉取请求,并随着工作的继续而修订。在十个软件工作负载和三个模型上的360次受控运行中,Relic将完整契约交付率从14.06%提高到19.76%(+5.71个百分点),相比没有协议生命周期的匹配结构化团队,在每个模型层中改善了所有四个已验证的生产端点。在新成员转移下,行为正确率在没有继承协议时为25.4%,在相同规则以可读文本提供时为34.6%,在可执行绑定时为41.2%,比仅文本高出6.5个百分点。在完整的CooperBench基准上,排除损坏的基准对后,Relic达到367/477(76.9%),确立了在同行结构化系统中报告的最佳结果。在固定的48对同模型子集上,Relic也超过了Solo(29/48 vs. 26/48),扭转了官方同行基线所表现出的协调损失。总之,这些结果表明协作经验如何成为持久的组织状态,在创建它的成员之后仍然有用。
cs.AI / 162 / 2609.32990

Certified Long-Horizon Code Agent Evolution via Validation-Gated Skill Optimization

基于验证门控技能优化的可认证长时程代码智能体进化
Wang, Yifan, Cheng, Hao, Li, Xiaomin, Hao, Yuexing, Ramesh, Hemanth Neelgund, Jung, Dongwon, Tang, Hao, Wang, Keru, Zhou, Chenliang, Wu, Qianhui, Yao, Wenlin, Grama, Ananth, Banburski-Fahey, Andrzej, Peng, Baolin, Lanier, Jaron, Gao, Jianfeng
Abstract
Long horizon agent self-evolution without model weight updates is essential for enabling deployed agents to accumulate reusable skills and improve over time. Prior self-evolution work has focused primarily on short-horizon tasks, while repository-level software engineering remains unexplored despite being an ideal testbed for long-horizon adaptation. In this setting, agents are required to solve streams of sequential tasks, navigate complex dependencies with evolving repositories and persistently store and reuse experience. Text-based skill optimization offers an efficient, non-parametric approach for such adaptation. However, existing methods often suffer from unstable updates, performance drawdown, and agent collapse over extended deployments. In this paper, we formalize the concept of in-context self-evolution and introduce VALVE, a validated-gated framework for long-horizon skill optimization. We establish finite convergence, provide theoretical guarantees for future-task gain and drawdown, and derive the validation and evaluation holdout sizes required for a prescribed tolerance, with leading-order scaling Empirically, our pipeline, VALVE achieves stable self-improvement over evolution horizon spanning more than 1,000 SWE tasks, with average final and peak gains of $14.9$ and $16.5$ points across three frontier models (GPT-5.5, Claude-4.6 and MiniMax-M2.7). The validation gate reduces average drawdown by 75% and produces an 11x more compact skill bank than ungated evolution. We further present extensive ablations identifying the design choices most critical to long-horizon skill evolution.
Chinese Translation
无需更新模型权重的长时程智能体自进化对于使已部署的智能体积累可重用技能并随时间改进至关重要。此前的自进化工作主要关注短时程任务,而仓库级软件工程尽管是长时程适应的理想试验台,却仍未得到探索。在这种场景下,智能体需要解决连续的序列任务流,在演化的仓库中导航复杂依赖关系,并持续存储和重用经验。基于文本的技能优化为这种适应提供了一种高效的非参数方法。然而,现有方法在扩展部署中常常面临更新不稳定、性能下降和智能体崩溃的问题。在本文中,我们形式化了上下文内自进化的概念,并介绍了 VALVE,一个用于长时程技能优化的验证门控框架。我们建立了有限收敛性,为未来任务增益和性能下降提供了理论保证,并推导了在给定容差下所需的验证和评估留出集大小,具有主导阶缩放。经验上,我们的流程 VALVE 在跨越 1000 多个 SWE 任务的进化视野上实现了稳定的自我改进,在三个前沿模型(GPT-5.5、Claude-4.6 和 MiniMax-M2.7)上平均最终增益和峰值增益分别为 14.9 和 16.5 分。验证门将平均性能下降减少了 75%,并产生比无门控进化紧凑 11 倍的技能库。我们进一步提供了广泛的消融实验,识别出对长时程技能进化最关键的设计选择。
cs.AI / 163 / 2609.32993

X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization

X-Tree:面向高效智能体泛化的可复用经验标记化
Cheng, Sitao, Yin, Xunjian, Sun, Zhiyuan, Li, Yuxuan, Zhou, Ruiwen, Jian, Xiangru, Zhong, Victor
Abstract
Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits unused, and flat training uses each scarce trajectory less fully than its content allows. Recent agents do use that structure, but only as LLM-written skills in context, never in the weights, so their gains do not generalize beyond retrieval. We instead recover this hierarchy from the data itself and train on it, with no LLM calls. Following text tokenizers, which build a vocabulary by counting alone, we score action spans by reusability and merge canonicalized actions into a reusable eXperience tree (X-Tree). Each X-Tree node captures how a frequent and success-bearing skill is composed from sub-skills, guiding efficient generalization. We integrate X-Tree into three training settings: offline RL, with each node as a training instance; online RLVR, with an adaptive skill bonus; and on-policy self-distillation, with X-Tree as the self-teacher's privileged context. Across WebArena, ScienceWorld, and WebShop at three model scales, X-Tree improves over standard recipes at matched data and budget by up to 4.5% SR on WebArena, 5.8% SR on ScienceWorld and 4.1% success on WebShop. Matched analyses attribute the gains to the X-Tree structure and the three integrations.
Chinese Translation
多步智能体在扁平动作流上进行训练:SFT和RLVR对每个标记一视同仁,忽略了跨任务重复出现的子过程,即那种让人类能够从可复用例程自顶向下规划的层次结构。这种结构未被利用,扁平训练对每个稀缺轨迹的利用不如其内容所允许的那样充分。最近的智能体确实使用了这种结构,但仅作为LLM编写的技能在上下文中,从未在权重中,因此它们的收益无法泛化到检索之外。我们则从数据本身恢复这种层次结构并基于其进行训练,无需调用LLM。遵循仅通过计数构建词汇表的文本分词器,我们根据可复用性对动作片段进行评分,并将规范化动作合并到可复用的经验树(X-Tree)中。每个X-Tree节点捕获了一个频繁且包含成功的技能如何由子技能组成,从而指导高效的泛化。我们将X-Tree集成到三种训练设置中:离线强化学习,其中每个节点作为一个训练实例;在线RLVR,带有自适应技能奖励;以及同策略自蒸馏,其中X-Tree作为自教师的特权上下文。在WebArena、ScienceWorld和WebShop上,在三种模型规模下,X-Tree在相同数据和预算下比标准方法提高了:WebArena上成功率(SR)高达4.5%,ScienceWorld上SR高达5.8%,WebShop上成功率高达4.1%。匹配分析将收益归因于X-Tree结构和三种集成。
cs.AI / 164 / 2609.33010

Model-Aware Data Selection from In-and-Out Information Interplay

基于内外信息交互的模型感知数据选择
Wang, Yifan, Li, Xiaomin, Hao, Yuexing, Jung, Dongwon, Ramesh, Hemanth Neelgund, Grama, Ananth, Chandrasekaran, Varun, Hu, Yu, Banburski-Fahey, Andrzej, Lanier, Jaron
Abstract
LLMs are effective representations that assimilate vast amounts of knowledge during pretraining, but post-training is necessary for models to reliably access this knowledge and "know what they know." We observe an interesting rank equilibrium between knowledge stored in the weights and the data stream passing through the model. Across all model layers, we find that the hidden states (data stream) follow a U-shaped pattern, showing substantial compression in early layers and a steep rise during the late-layer decoding phase. In contrast, the weight rank follows an inverted U-shaped pattern, with very low rank in the early and late layers and high rank in the middle. We interpret this as an in-and-out information interplay: intermediate activations do not need to carry content that the weights can supply later, so they primarily preserve what the weights cannot provide. Motivated by this observation, we propose a model-aware data selection method, CAP (Counterfactual Assimilation Profile), which can determine whether a data candidate contains information accessible to the current model by utilizing the divergence gap in early- and late-layer representations between model-generated and reference responses. Across math, code, and science domains, CAP delivers 35.4% greater average improvement over the base model than the strongest baseline under different selection budgets. With only 10% of the data pool, CAP surpasses or matches full-pool training on math and science. We further show that CAP transfers to multimodal data selection and is robust to response horizon and noise.
Chinese Translation
大语言模型(LLMs)是在预训练期间吸收大量知识的有效表示,但为了使模型能够可靠地访问这些知识并“知道自己知道什么”,后训练是必要的。我们观察到在权重中存储的知识与通过模型的数据流之间存在一种有趣的秩均衡。在所有模型层中,我们发现隐藏状态(数据流)呈现U形模式,在早期层表现出显著压缩,而在后期解码阶段急剧上升。相比之下,权重秩呈现倒U形模式,在早期和晚期层秩很低,而在中间层秩很高。我们将此解释为一种内外信息交互:中间激活不需要携带权重稍后可以提供的内容,因此它们主要保留权重无法提供的信息。受此观察启发,我们提出了一种模型感知的数据选择方法CAP(反事实同化轮廓),该方法通过利用模型生成响应和参考响应在早期和晚期层表示中的差异差距,来确定数据候选是否包含当前模型可访问的信息。在数学、代码和科学领域,在不同选择预算下,CAP比最强基线在基础模型上平均提升高出35.4%。仅使用10%的数据池,CAP在数学和科学上就超过或匹配全池训练。我们进一步表明,CAP可迁移到多模态数据选择,并且对响应范围和噪声具有鲁棒性。
cs.AI / 165 / 2609.33012

When Pair Count Is Not the Sample Size: What All-Pairs Agent Comparisons Estimate

当配对数量并非样本量:全配对智能体比较估计什么?
Huang, Wei-Jung
Abstract
When an agent benchmark compares every pair of leaderboard entries, the number of comparisons can look much larger than the independent evidence behind them: A versus B and A versus C both reuse A. Whether this reuse affects inference depends on what the analysis is meant to describe. If the board and its outcomes are fixed, the all-pairs mean is an exact summary of those entries, and any interval must come from another declared source of randomness. If the entries are instead treated as iid draws from a population of future configurations and the pair rule is regular and nondegenerate, the same mean is an order-two U-statistic whose first-order uncertainty depends on the number of configurations, not the number of pairs. We use near ties as the running example, but the distinction extends to other symmetric pair summaries when their regularity conditions hold. We examine both interpretations using a fixed SWE-bench Verified snapshot and an exact binary model with known truth. On SWE-bench, intervals that accounted for shared configurations were more than twice as wide as a pair-iid reference that treated the pairs as independent. In the exact model, pair-iid coverage fell far below the nominal level when edges shared endpoints but remained near nominal for matched independent edges. Results on two other fixed leaderboards show that exact summaries also depend on which pairs are included and how they are weighted. An all-pairs analysis must therefore state what is fixed, what is sampled, and how it handles shared entries and pair aggregation.
Chinese Translation
当智能体基准比较排行榜条目中的每一对时,比较次数可能看起来远大于其背后独立证据的数量:A与B比较以及A与C比较都重复使用了A。这种重复使用是否影响推断,取决于分析旨在描述什么。如果排行榜及其结果是固定的,则全配对均值是这些条目的精确汇总,任何区间必须来自另一个已声明的随机性来源。如果转而将条目视为从未来配置总体中独立同分布(iid)抽取的样本,并且配对规则是正则且非退化的,则同一均值是二阶U统计量,其一阶不确定性取决于配置数量,而非配对数量。我们以近似平局作为贯穿示例,但当其他对称配对汇总满足正则性条件时,这一区分同样适用。我们使用固定的SWE-bench Verified快照和具有已知真值的精确二元模型来检验这两种解释。在SWE-bench上,考虑了共享配置的区间宽度是将配对视为独立的pair-iid参考的两倍以上。在精确模型中,当边共享端点时,pair-iid覆盖率远低于名义水平,但在匹配的独立边中仍接近名义水平。另外两个固定排行榜上的结果表明,精确汇总还取决于包含了哪些配对以及如何对它们加权。因此,全配对分析必须说明什么是固定的、什么是抽样的,以及如何处理共享条目和配对聚合。
cs.AI / 166 / 2609.33013

The Epistemics of Agent Memory: Measuring, and Governing, the Consolidation Decision in Long-Horizon LLM Agents

智能体记忆的认识论:长时程 LLM 智能体中巩固决策的测量与治理
Annapureddy, Sasank, Thamatani, Anjaneya Prasad
Abstract
Long-horizon LLM agents must convert accumulated experience into durable memory, deciding what to keep, compress, abstract into reusable skills and rules, or forget. We report a four-phase research program on this consolidation problem whose central finding is a shift in what is measured: from how much an agent remembers, to whether its consolidation decisions are any good, to whether those decisions can be trusted. Phase 1 learns episodic boundaries from agent traces by downstream utility; an honest near-miss (oracle correlation 0.691 vs a 0.70 bar) whose lasting output is a three-gate anti-leakage protocol. Phase 2 learns when to promote experience and to which abstraction level under a token budget, achieving a verified +22.7% task-success improvement with 7x compression, but exposing a degenerate-forgetting failure and a distribution-shift failure mode we name lambda-prevalence coupling. Phase 3 introduces ConsolidationBench, an oracle-by-construction benchmark that scores consolidation decisions against a known optimum on three non-circular axes; production retrieval systems retain information yet score zero on cross-level transfer. Phase 4 introduces governed consolidation: the decision wrapped in poison-resistance, reversibility, and auditability guarantees with a quality gate. Governance is statistically distinct from the quality score ($r^2 = 0.43$; partial $r = 0.27$; identical-quality policies differ threefold in governance), so the contribution survives independently of the metric's external validity. On that question we report a resolved negative: after a graded-reuse redesign removed a structural ceiling, a two-benchmark study with 2,532 real answer cells finds the quality score does not predict real transfer accuracy (pooled Spearman $\rho = -0.24$, n = 12, CI spanning zero). An adversarial self-critique pass cleared the final claim set with zero surviving overclaims.
Chinese Translation
长时程 LLM 智能体必须将积累的经验转化为持久记忆,决定保留什么、压缩什么、抽象为可复用技能与规则,还是遗忘什么。我们报告一个针对这一巩固问题的四阶段研究计划,其核心发现是测量对象的转变:从智能体记住多少,转向其巩固决策是否足够好,再转向这些决策是否可信。第 1 阶段依据下游效用从智能体轨迹中学习情景边界;这是一次诚实的未达标(oracle 相关性 0.691,而门槛为 0.70),其持久产出是一个三闸门防泄漏协议。第 2 阶段学习在 token 预算下何时提升经验以及提升到哪个抽象层级,实现了经核验的 +22.7% 任务成功率提升与 7 倍压缩,但暴露出退化性遗忘失败,以及一种我们命名为 lambda-prevalence coupling 的分布偏移失败模式。第 3 阶段提出 ConsolidationBench,这是一个按构造具有 oracle 的基准,在三个非循环轴上依据已知最优解对巩固决策进行评分;生产检索系统能保留信息,但在跨层级迁移上得分为零。第 4 阶段提出受治理的巩固:将决策包裹在抗投毒、可逆性和可审计性保证之中,并设置质量门控。治理与质量分数在统计上不同($r^2 = 0.43$;偏 $r = 0.27$;质量相同的策略在治理上相差三倍),因此该贡献独立于该指标的外部效度而成立。关于这个问题,我们报告一个已解决的否定结果:在分级复用重新设计消除了结构性上限之后,一项包含 2,532 个真实答案单元的双基准研究发现,质量分数不能预测真实迁移准确率(合并 Spearman ρ = -0.24,n = 12,置信区间跨零)。一次对抗性自我批评通过了最终主张集,未留下任何过度主张。
cs.AI / 167 / 2609.33017

Trust and Task Completion in the World of Consumer AI Agents

消费者AI智能体世界中的信任与任务完成
Olieslagers, Jeroen, Pujol, Eduardo, Zahavi, Gal, Ingemarsson, Lukas, Poddar, Shivani
Abstract
Action agents do things for people. They send email, spend money, and call businesses while the user is busy with something else, so a mistake can turn into an action before anyone notices. They fail their users in two ways. They break trust when they do something the user never agreed to, or hold back after the user clearly said go. And they fall short on completion when they give up on errands that turn out to be hard. Both depend heavily on the harness around the model, meaning its instructions, tools, context, and guardrails. We built an evaluation that scores trust and completion on the same runs, in a simulated world of businesses with their own websites, inboxes, and phone lines, and of people who write back. A simulated user answers the assistant's questions. Trust means that nothing happens the user did not agree to. No email goes to someone they never approved, no private detail ends up on a group thread, no money is spent past their limit, no stranger's instructions are followed, and nothing is claimed without a source. Every trap has a matched control in which acting is the right call. We use the evaluation to measure Fo, Wajo's personal assistant, against a base model with basic instructions on three foundation models, and against the Fo harness with its guardrails switched off. Fo completes 71% of the errands and keeps the user's trust on 94% of the trap runs. The base models complete 50% to 64% and keep trust on 59% to 75%. On the matched controls, Fo goes ahead slightly less often. OpenClaw, a popular open-source assistant given the same access, completes 42% of the errands it shares with Fo, against 71%, and keeps the user's trust on 74% of the shared trap runs, against 94%. Measuring trust and completion together, on the whole system rather than the model alone, is how we think action agents become safe to hand real work to.
Chinese Translation
行动智能体为人们做事。它们发送邮件、花钱、打电话给商家,而用户正忙于其他事情,因此一个错误可能在任何人注意到之前就变成了行动。它们以两种方式令用户失望。当它们做了用户从未同意的事情时,或者当用户明确说“开始”后却退缩时,它们就破坏了信任。而当它们放弃那些被证明困难的任务时,它们就未能完成。这两者都严重依赖于模型周围的框架,即其指令、工具、上下文和防护栏。我们构建了一个评估,在模拟的商业世界(商家拥有自己的网站、收件箱和电话线)以及会回信的人中,对同一批运行中的信任和完成度进行评分。一个模拟用户回答助手的问题。信任意味着没有发生用户未同意的事情。没有邮件发给用户从未批准的人,没有私人细节出现在群组讨论中,没有超出用户限额的花费,没有遵循陌生人的指令,没有无来源的声明。每个陷阱都有一个匹配的对照,其中采取行动是正确的选择。我们使用该评估来衡量 Fo——Wajo 的个人助理——与在三个基础模型上仅带有基本指令的基线模型进行对比,并与关闭了防护栏的 Fo 框架进行对比。Fo 完成了 71% 的任务,并在 94% 的陷阱运行中保持了用户的信任。基线模型完成了 50% 到 64%,并在 59% 到 75% 的陷阱运行中保持了信任。在匹配的对照中,Fo 继续行动的比例略低。OpenClaw,一个被赋予相同访问权限的流行开源助理,完成了其与 Fo 共享任务的 42%,而 Fo 为 71%,并在共享的陷阱运行中保持了用户信任的 74%,而 Fo 为 94%。我们认为,将信任和完成度一起衡量,针对整个系统而非仅模型本身,是让行动智能体能够安全地接手实际工作的方式。
cs.AI / 168 / 2609.33023

SRE-Marathon: A Continuous, Change-Driven Benchmark for Autonomous Site Reliability Agents

SRE-Marathon:面向自主站点可靠性智能体的持续、变更驱动基准
Tian, Yifang, Bai, Yingjian, He, Yifeng, Chong, Zichun, Gao, Yuanchen, Li, Yiran, Jacobsen, Hans-Arno
Abstract
Benchmarks for site reliability engineering (SRE) agents are typically episodic: one fault is injected, the agent receives an incident task, and its response is scored. Production operation is not. Incidents surface through noisy alerts, overlap in time, and often originate from code or configuration changes. We present SRE-Marathon, a benchmark for long-horizon, continuous SRE operation. An agent is invoked at a fixed cadence with cumulative alert history and a persistent workspace while operating a live two-zone Kubernetes deployment as a fault orchestrator injects overlapping faults according to a seeded, production-calibrated schedule. Curated code and configuration changes deployed through the same build pipeline available for repair. Each run is recorded into a sealed bundle and scored offline: Marathon-Score credits each injected fault for ordered progress through correlation, localization, and repair, with all metrics computed deterministically from recorded system evidence. Across three applications and about sixty faults per run, the best of 10 methods reaches only 41.3 out of 100. Agents often correlate and localize faults, but almost never complete repairs while the faults remain active.
Chinese Translation
站点可靠性工程(SRE)智能体的基准通常是片段式的:注入一个故障,智能体接收一个事件任务,然后对其响应进行评分。生产环境运行并非如此。事件通过嘈杂的告警浮现,在时间上相互重叠,并且通常源于代码或配置变更。我们提出 SRE-Marathon,一个用于长时程、连续 SRE 运行的基准。智能体以固定节奏被调用,并拥有累积的告警历史和持久化工作区,同时操作一个在线双区域 Kubernetes 部署;故障编排器按照种子化、经生产校准的计划注入重叠故障。精选的代码与配置变更通过可用于修复的同一构建流水线部署。每次运行都被记录到一个密封捆绑包中,并离线评分:Marathon-Score 根据每个注入故障在关联、定位和修复方面的有序进展给予分数,所有指标均根据记录的系统证据确定性地计算得出。在三个应用、每次运行约 60 个故障的设置下,10 种方法中最佳者仅达到 100 分中的 41.3 分。智能体通常能够关联并定位故障,但当故障仍处于活跃状态时,几乎从未完成修复。
cs.AI / 169 / 2609.33039

Agent Safety From Within: Detecting Harmful Trajectories from LLM Internal States

从内部保障智能体安全:从LLM内部状态检测有害轨迹
Jiao, Difan, Anderson, Ashton
Abstract
Language model agents can now perform sophisticated sequences of actions via tools and harnesses, which has increased the scope of the damage they can cause. Guard models, however, are mainly built for content moderation and thus are not well-suited to detecting this agentic risk. To address this, we proceed by first conducting a representational analysis, then use the resulting insights to build a solution. In our analysis, we focus on two types of trajectory-level agentic harms: harmful content, which is expressed directly, and unsafe tool use, which depends on whether an action is consistent with the interaction that produced it. We investigate how open-source guard models represent these two types of harm and find that they are linearly readable inside the model, even though guard models predict no better than chance on pairs that differ only in the called tool's schema. The two harm types also follow nearly orthogonal internal directions, and neither reliably serves as a proxy for the other. These results motivate reading trajectory safety directly from internal states. We introduce TACIT, a readout of a frozen backbone's internal states that decodes no tokens. Trained on six trajectory-safety benchmarks, a linear probe raises mean macro-F1 from 62.3 for the strongest open guard to 80.7, and refined readouts reach 86.2. With each benchmark held out of training entirely, the refined readouts still lead the strongest guard (65.7 vs. 61.1). With the same backbone, training data and test split, the frozen readout is on par with full safety fine-tuning, and it improves the fine-tuned model further when applied on top. The probe trains about one millionth as many parameters as full fine-tuning in about a sixth of the time, and TACIT has the lowest latency of the guards we evaluate.
Chinese Translation
语言模型智能体如今可以通过工具和框架执行复杂的动作序列,这扩大了它们可能造成的损害范围。然而,防护模型主要针对内容审核而构建,因此不太适合检测这种智能体风险。为了解决这个问题,我们首先进行表征分析,然后利用由此产生的洞见来构建解决方案。在分析中,我们聚焦于两类轨迹级智能体危害:有害内容,它是直接表达的;以及不安全工具使用,它取决于一个动作是否与产生它的交互一致。我们研究开源防护模型如何表征这两类危害,发现它们在模型内部是线性可读的,尽管防护模型在仅被调用工具的schema不同的配对上的预测并不比随机猜测好。这两类危害还遵循几乎正交的内部方向,并且两者都不能可靠地作为对方的代理。这些结果促使我们直接从内部状态读取轨迹安全性。我们提出TACIT,一种对冻结主干模型内部状态的读出,它不解码任何token。在六个轨迹安全基准上训练,线性探针将平均宏F1从最强开源防护模型的62.3提升到80.7,而精炼读出达到86.2。当每个基准都完全排除在训练之外时,精炼读出仍然领先于最强防护模型(65.7对61.1)。在相同的主干模型、训练数据和测试划分下,冻结读出与完整安全微调相当,并且当叠加在微调模型上时,它进一步改进了微调模型。该探针训练的参数数量约为完整微调的百万分之一,时间约为六分之一,而且在我们评估的防护模型中,TACIT具有最低的延迟。
cs.AI / 170 / 2609.33052

BudgetVerify: Budget-Tiered Verification for Financial QA

BudgetVerify: 面向金融问答的预算分层验证
Jenq, Janet, Shen, Hongda
Abstract
Financial question answering often requires precise numerical extraction, unit handling, and arithmetic over tables and text, but applying expensive verification uniformly wastes test-time compute. We propose BudgetVerify, a budget-tiered generator-verifier framework that routes each generated answer to one of three verification tiers: no verification, lightweight check-and-revise, or higher-cost solve-first-then-compare verification. The router is trained from offline correctness and token-cost outcomes and, at test time, selects a verification tier using information available before verification, including the question, context statistics, the generated answer, and associated generator metadata. The selected tier either returns the generated answer directly or invokes the corresponding verifier. Across six commercial and open-weight base models, BudgetVerify consistently produces more efficient accuracy-cost Pareto frontiers than fixed verification policies by selectively allocating stronger verification only when it is useful. Although absolute performance varies across models, these efficiency gains and the resulting qualitative frontier shape are consistent across generator models.
Chinese Translation
金融问答通常需要精确的数值提取、单位处理和基于表格与文本的算术运算,但统一应用昂贵的验证会浪费测试时的计算资源。我们提出 BudgetVerify,一个预算分层的生成器-验证器框架,它将每个生成的答案路由到三个验证层级之一:无验证、轻量级检查与修正,或更高成本的先求解后比较验证。路由器基于离线的正确性和令牌成本结果进行训练,并在测试时利用验证前可用的信息(包括问题、上下文统计量、生成的答案以及相关的生成器元数据)来选择验证层级。所选层级要么直接返回生成的答案,要么调用相应的验证器。在六个商业和开放权重的基座模型上,BudgetVerify 通过仅在有用时选择性地分配更强的验证,始终产生比固定验证策略更高效的准确率-成本帕累托前沿。尽管绝对性能因模型而异,但这些效率提升以及由此产生的定性前沿形状在不同生成器模型之间是一致的。
cs.AI / 171 / 2609.33055

Large Language Models Substantially Compress Well-Being Inequality but Largely Preserve Its Socioeconomic Structure

大语言模型大幅压缩福祉不平等,但基本保留其社会经济结构
Powdthavee, Nattavudh
Abstract
Research using large language models (LLMs) to generate synthetic populations has repeatedly shown that model outputs compress the diversity of human experience. This has raised doubts about whether LLM-generated data can capture meaningful differences within populations. We show that such compression does not necessarily erase the social structure of human heterogeneity. Using 93,901 respondents from 66 countries and territories in Wave 7 of the World Values Survey, we ask six LLMs to predict respondents' life satisfaction from demographic, socioeconomic, and attitudinal profiles. All six models substantially understate the overall dispersion of life satisfaction. Yet after normalizing for these differences in scale, they largely reproduce the human income gradient in well-being inequality: lower-income groups remain relatively more heterogeneous than higher-income groups. The pattern is robust to country fixed effects, equal-country weighting, WVS survey weights, and observed demographic composition, and it extends directionally to employment, education, and perceived control. Fidelity is weaker for extreme outcomes and country-specific gradients. These results show that the amount of heterogeneity preserved by an LLM and the way that heterogeneity is distributed across social groups are distinct properties. LLM-generated populations can therefore substantially compress human variation while retaining meaningful information about where that variation is concentrated.
Chinese Translation
使用大语言模型(LLMs)生成合成人口的研究已多次表明,模型输出会压缩人类经验的多样性。这引发了人们对LLM生成的数据能否捕捉人群内有意义差异的怀疑。我们表明,这种压缩并不一定会抹去人类异质性的社会结构。利用世界价值观调查(WVS)第七轮中来自66个国家和地区的93,901名受访者,我们要求六个LLM根据人口统计、社会经济和态度特征预测受访者的生活满意度。所有六个模型都大幅低估了生活满意度的总体离散程度。然而,在针对这些规模差异进行标准化后,它们大体上再现了人类福祉不平等中的收入梯度:低收入群体仍然比高收入群体相对更加异质。该模式在控制国家固定效应、等国家权重、WVS调查权重和观测人口构成后依然稳健,并在方向上扩展至就业、教育和感知控制。对于极端结果和特定国家的梯度,保真度较弱。这些结果表明,LLM所保留的异质性数量与异质性在社会群体中的分布方式是两种不同的属性。因此,LLM生成的群体可以大幅压缩人类变异,同时保留关于该变异集中于何处的有意义信息。
cs.AI / 172 / 2609.33061

LLM sequential decision making under uncertainty in biochemical domains

生化领域不确定性下的LLM序列决策
Akke, Mattias, Yang, Soojung, Ruža, Jurgis, Edamadaka, Sathya, Gómez-Bombarelli, Rafael
Abstract
Large language models (LLMs) are increasingly used to drive scientific discovery. Understanding how LLMs make decisions from new data and memory of the literature is vital before trusting them to design experiments under tight experimental budgets. However, their decision strategies are invisible in the current performance scores used to evaluate research agents. Here, we benchmark five frontier LLMs in a Bayesian Optimization setting against published statistical baselines on seven combinatorial datasets spanning protein engineering, reaction optimization, molecular design, peptide self-assembly, and catalysis. Performance is paired with direct measurements of model beliefs and actions, enabling highly resolved behavior analysis. A prompt ablation that progressively strips context separates memorization from chemical reasoning and from bare categorical optimization. Prior chemical knowledge helps in expectation, but with high variance and occasionally even harms performance. No configuration tested decisively beats a mean statistical baseline across domains. Belief-movement and Martingale diagnostics, corrected here for a measurement-noise bias that mislabels rational agents as irrational, show that models overreact to incoming data rather than entrenching on their priors in the contexts studied here. Interestingly, while LLM actions are exploitative, models sincerely intend to explore and consistently act on that intent. This failure is a competence gap arising from context-stickiness. Removing in-context history restores exploration, indicating that priors and data must be decoupled to achieve effective LLM-driven discovery.
Chinese Translation
大语言模型(LLMs)正越来越多地用于驱动科学发现。在将其用于在紧张的实验预算下设计实验之前,理解LLMs如何从新数据和文献记忆中做出决策至关重要。然而,在当前用于评估研究智能体的性能评分中,它们的决策策略是不可见的。在这里,我们在贝叶斯优化设置中对五个前沿LLMs进行基准测试,与已发表的统计基线在七个组合数据集上进行比较,涵盖蛋白质工程、反应优化、分子设计、肽自组装和催化。性能与对模型信念和行动的直接测量相结合,实现了高分辨率的行为分析。一种逐步剥离上下文的提示消融将记忆与化学推理以及纯粹的分类优化区分开来。先前的化学知识在期望上有帮助,但方差很大,有时甚至损害性能。在所有领域中,测试的配置均未能决定性地超越平均统计基线。信念移动和鞅诊断(此处针对将理性智能体错误标记为不理性的测量噪声偏差进行了校正)表明,在所研究的背景下,模型对传入数据过度反应,而非固守其先验。有趣的是,尽管LLM的行为是利用性的,但模型真诚地意图探索,并一致地按照该意图行动。这种失败是由上下文粘性引起的能力差距。移除上下文历史可恢复探索,表明必须将先验与数据解耦,以实现有效的LLM驱动的发现。
cs.AI / 173 / 2609.33075

QureRadEmbed: Structuring Radiological Similarity through Attribute and Reasoning Supervision

QureRadEmbed:通过属性与推理监督构建放射学相似性
Prabhu, Janhavi, Sahil, Shukla, Shivam Ashok, Tadepalli, Manoj
Abstract
Radiological similarity depends on disease relationships and on fine details such as laterality, lobe, severity, size, and certainty. Broad biomedical similarity can overlook these qualifiers, particularly when several attributes vary together. We introduce QureRadEmbed, a 4B radiology-aware encoder trained with two complementary signals: RadSim supplies deterministic, attribute-decomposed ranking targets, while RadThought aligns reports with hierarchical evidence and reasoning descriptions. A three-stage curriculum combines these signals with report triplets, finding perturbations, and single- and cross-attribute contrasts. The final model achieves 0.996 mean ordering accuracy across ten controlled synthetic attributes and raises Spearman correlation with the designed joint-attribute targets from 0.501 to 0.976. On external findings-to-impression retrieval, Recall@1 reaches 10.5% on Open-I, 12.4% on testing XR, and 42.9% on testing CT, compared with 6.6%, 5.7%, and 31.4% for its backbone. Frozen embeddings support finding extraction with only 100 labeled testing-XR reports (macro-F1 0.481 versus 0.412 for the backbone). Whole-report comparison costs 8.8 seconds per 1,000 pairs in our benchmark, versus 2,755.1 seconds for the generative evaluator GREEN. Sentence-level comparison improves sensitivity to local discrepancies, although generative evaluation remains stronger on several expert-rated and subtle-error tasks. The results support reusable radiology-aware representations for search, structured report indexing, and efficient report comparison.
Chinese Translation
放射学相似性取决于疾病关系以及诸如偏侧性、肺叶、严重程度、大小和确定性等细节。广泛的生物医学相似性可能会忽略这些限定因素,特别是当多个属性同时变化时。我们介绍了QureRadEmbed,一个40亿参数的放射学感知编码器,使用两种互补信号进行训练:RadSim提供确定性的、属性分解的排序目标,而RadThought将报告与分层证据和推理描述对齐。一个三阶段课程将这些信号与报告三元组、发现扰动以及单属性和跨属性对比相结合。最终模型在十个受控合成属性上达到了0.996的平均排序准确率,并将与设计的联合属性目标的斯皮尔曼相关性从0.501提高到0.976。在外部发现到印象检索中,Recall@1在Open-I上达到10.5%,在testing XR上达到12.4%,在testing CT上达到42.9%,而其骨干模型分别为6.6%、5.7%和31.4%。冻结嵌入仅使用100个标记的testing-XR报告即可支持发现提取(宏F1为0.481,而骨干模型为0.412)。在我们的基准测试中,整份报告比较每1,000对耗时8.8秒,而生成式评估器GREEN耗时2,755.1秒。句子级比较提高了对局部差异的敏感性,尽管生成式评估在若干专家评分和细微错误任务上仍然更强。这些结果支持可复用的放射学感知表示,用于搜索、结构化报告索引和高效报告比较。
cs.AI / 174 / 2609.33079

Structure-Mapping-Guided Self-Explanation for Learning Mathematical Procedures

结构映射引导的自我解释用于学习数学程序
Lee, Shinhaeng, MacLellan, Christopher J., Weitekamp, Daniel
Abstract
Worked examples are a powerful form of instruction, but learners must infer how the demonstrated steps were produced. A naive simulation of this self-explanation process can generate thousands of numerical explanations that reproduce one observed change without capturing its underlying procedure. We propose structure-mapping-guided self-explanation as a computational account of the cognitive biases that reduce search effort and make this inference tractable. The model represents mathematical expressions as typed relational structures and uses structure mapping to identify corresponding source and target regions. For each changed target value, the corresponding source region serves as an anchor: it guides abductive search toward structurally relevant values and operations before broader alternatives, yielding ordered, executable candidate procedures with inspectable source evidence. Across 70 mathematical transformations containing 120 changed numeric components, the model recovered every intended procedure. It returned the intended procedure before any other computation producing the same target value in 104 subproblems (86.7%), compared with a median of 32 (26.7%) across 100 unguided runs that tested candidate calculations in random order. Our proposed model also tested 93.8% fewer combinations of values and operations than unguided search before reaching the intended procedures. These results provide an efficient, interpretable account of how relational structure can guide procedural learning and a testable hypothesis about human self-explanation from worked examples.
Chinese Translation
样例是一种强大的教学形式,但学习者必须推断所演示的步骤是如何产生的。对这一自我解释过程的朴素模拟可以生成数千个数值解释,这些解释重现了一个观察到的变化,但没有捕捉到其底层程序。我们提出结构映射引导的自我解释,作为一种计算解释,说明减少搜索努力并使这一推断变得可行的认知偏差。该模型将数学表达式表示为带类型的关系结构,并使用结构映射来识别对应的源区域和目标区域。对于每个变化的目标值,对应的源区域作为锚点:它引导溯因搜索优先朝向结构相关的值和操作,然后再考虑更广泛的备选方案,从而产生有序的、可执行的候选程序,并带有可检查的源证据。在包含120个变化数值组件的70个数学变换中,该模型恢复了每一个预期程序。在104个子问题(86.7%)中,它在任何其他产生相同目标值的计算之前返回了预期程序,而100次无引导运行(以随机顺序测试候选计算)的中位数为32(26.7%)。我们提出的模型在达到预期程序之前,所测试的值和操作组合比无引导搜索少了93.8%。这些结果既提供了一个高效、可解释的说明,阐述关系结构如何引导程序学习,又提出了一个关于人类从样例中进行自我解释的可检验假设。
cs.AI / 175 / 2609.33085

The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration

模型知道另一条路:面向有效RLVR探索的策略切换
Cui, Jin, Long, Xinyue, Zhao, Boran, Ren, Pengju, Dong, Hao
Abstract
Reinforcement learning with verifiable rewards (RLVR) is often limited by insufficient exploration: difficult problems can yield uniformly incorrect rollout groups and therefore little learning signal. We show that such failures need not reflect missing capability. Instead, finite sampling often concentrates on a problem-specific dominant reasoning strategy while leaving alternative strategies already supported by the model unexplored. Moreover, the accessibility of these strategies evolves during RL: some are internalized into autonomous behavior, while others become difficult to elicit before being absorbed. Motivated by these observations, we introduce Problem--Strategy Rollout Allocation (PSRA), which treats unguided and strategy-conditioned prompts as competing exploration arms and uses Bayesian sequential allocation to direct a fixed rollout budget toward arms most likely to yield informative, non-saturated groups. A preservation objective keeps useful strategy-conditioned routes accessible while successful guided behaviors are transferred to the unguided policy. Across Qwen2.5 models from 1.5B to 7B and two RL training corpora, PSRA consistently improves reasoning performance, reduces dead saturation, strengthens out-of-distribution transfer, and maintains larger gains under increased inference budgets.
Chinese Translation
可验证奖励的强化学习(RLVR)通常受限于探索不足:困难问题可能产生完全错误的展开组,从而学习信号很少。我们表明,此类失败未必反映能力缺失。相反,有限采样往往集中于问题特定的主导推理策略,而模型已支持的替代策略却未被探索。此外,这些策略的可及性在RL过程中不断演变:一些被内化为自主行为,而另一些在被吸收之前变得难以引出。基于这些观察,我们提出了问题-策略展开分配(PSRA),该方法将无引导和策略条件化的提示视为竞争性探索臂,并使用贝叶斯序贯分配将固定的展开预算导向最有可能产生有信息量、非饱和组的臂。一个保留目标保持有用的策略条件化路径可访问,同时成功的引导行为被转移到无引导策略。在Qwen2.5模型(1.5B到7B)和两个RL训练语料库上,PSRA持续提升推理性能,减少死饱和,增强分布外迁移,并在增加的推理预算下保持更大的增益。
cs.AI / 176 / 2609.33115

Modular Discovery of General Game-Playing Algorithms with Large Language Models

使用大语言模型模块化发现通用游戏博弈算法
Li, Zun, Schultz, John, Lanctot, Marc, Hennes, Daniel
Abstract
General Game Playing across arbitrary games from rules alone remains challenging due to differing algorithmic requirements across game classes and strict decision-time constraints. Rather than hand-designing search heuristics for specific domains, can we leverage Large Language Models (LLMs) to discover general game-playing algorithms? Because language models can propose and refactor structured code, they provide an expressive proposal engine for exploring the space of algorithmic designs. We introduce a multi-agent LLM meta-learning system to co-evolve game-agnostic procedural search mechanisms in C++ alongside domain heuristics synthesized directly from game rules. Controlling the compute budget, we benchmark the discovered mechanisms across more than 400 diverse environments, including OpenSpiel training and held-out games, procedural simulation engines, and games with deep neural policy-value representations trained via PPO. Evaluated via AlphaRank stationary distributions and Soft Condorcet Optimization (SCO) against 15 established MCTS baselines, the discovered search mechanisms consistently achieve top-tier ratings and pairwise ballot majorities over most baselines across independent evolutionary runs, generalizing to unseen human-designed and procedurally synthesized games and remaining competitive with baselines on frozen neural network representations.
Chinese Translation
仅凭规则在任意游戏中进行通用游戏博弈仍然具有挑战性,因为不同游戏类别对算法的要求各异,且存在严格的决策时间约束。与其为特定领域手工设计搜索启发式,我们能否利用大语言模型(LLMs)来发现通用游戏博弈算法?由于语言模型能够提出并重构结构化代码,它们为探索算法设计空间提供了一个富有表现力的提议引擎。我们引入一个多智能体LLM元学习系统,以协同演化C++中与游戏无关的过程式搜索机制,以及直接从游戏规则合成的领域启发式。在控制计算预算的情况下,我们在超过400个多样化环境中对发现的机制进行基准测试,包括OpenSpiel训练和留出游戏、过程化模拟引擎,以及通过PPO训练的具有深度神经策略-价值表示的游戏。通过AlphaRank平稳分布和Soft Condorcet Optimization(SCO)针对15个已建立的MCTS基线进行评估,所发现的搜索机制在独立演化运行中持续获得顶级评分,并在大多数基线上取得成对投票多数,能够泛化到未见的人类设计游戏和过程化合成游戏,并在冻结神经网络表示上与基线保持竞争力。
cs.AI / 177 / 2609.33123

Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring

框架演化中的组合安全失败:识别与运行时监控
Zhang, Zhixiang, Liu, Zesen, Lai, Wai Ip, chen, Hongxu, She, Dongdong
Abstract
Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevolution and validates candidate harnesses or attributed individual component updates, leaving safety analysis of cross-component update interactions largely unexamined. To address this gap, we study compositional safety failures in harness evolution, where interactions among individually safe and utility-preserving component updates can produce undesirable or unsafe agent behavior, revealing a safety risk intrinsic to harness evolution. Across three safety-related benchmarks, we identify 43 pairwise and 18 irreducible 3-way compositional safety failures. Conventional solution incurs combinatorial complexity in validating cross-component interactions, leaving the safety checking impractical as the harness evolves. To solve this, we introduced a typed hypergraph that represents component states as nodes and safety-relevant higher-order interactions as hyperedges. When the harness changes, the hypergraph updates only the interaction neighborhood of the changed states rather than reconstructing the global composition space. Building on that, we develop a hypergraph-guided runtime monitoring mechanism. Experiments show that our method effectively mitigates compositional safety risks while preserving task utility and reducing interaction-checking costs, and further reveal an empirical safety-utility-cost trade-off across different safety mechanisms.
Chinese Translation
自演化智能体框架(harness)不断更新持久组件,如记忆、提示、技能和工具。我们将此过程称为框架演化(harness evolution)。然而,这种演化可能引入意外的安全风险。现有工作研究了框架错误演化,并验证候选框架或归因的单个组件更新,但跨组件更新交互的安全分析在很大程度上未被检验。为填补这一空白,我们研究了框架演化中的组合安全失败,其中单独安全且保持效用的组件更新之间的交互可能产生不良或不安全的智能体行为,揭示了框架演化固有的安全风险。在三个与安全相关的基准上,我们识别出 43 个成对和 18 个不可约的三方组合安全失败。传统解决方案在验证跨组件交互时会产生组合复杂性,使得随着框架演化,安全检查变得不切实际。为了解决这个问题,我们引入了一种类型化超图,将组件状态表示为节点,将安全相关的高阶交互表示为超边。当框架发生变化时,超图仅更新变化状态的交互邻域,而不是重建全局组合空间。在此基础上,我们开发了一种超图引导的运行时监控机制。实验表明,我们的方法有效缓解了组合安全风险,同时保持了任务效用并降低了交互检查成本,并进一步揭示了不同安全机制之间经验性的安全-效用-成本权衡。
cs.AI / 178 / 2609.33134

Ceiling of a Task: When Can a Transformer Succeed Without Its Chain of Thought?

任务的天花板:Transformer 何时能在没有其思维链的情况下成功?
He, Jiashu, Fan, Jinxuan, Xiao, Xiao, Marculescu, Radu, Ribeiro, Alejandro
Abstract
Reasoning models generate long chains of thought before they answer, yet it is debated whether the content of these chains does real computational work or is largely decorative. We study this question by viewing a transformer as a shallow circuit. One forward pass through a fixed number of layers has constant depth, so any procedure that runs the model a constant number of times is a shallow circuit. We call the best accuracy that a shallow circuit can reach on a task the ceiling of the task, and a task is serial if its ceiling lies below one. We prove three results on serial tasks that hold for every transformer, no matter how it was trained. Necessity: replacing the chain by anything that does not depend on its content, such as filler tokens or a restatement of the question, drives the accuracy down to the ceiling, and on a maximally serial task down to chance. Depth: no shallow computation can write the chain of a model whose accuracy exceeds the ceiling, not even approximately. Locality: the answer is one shallow pass away from the finished chain, so all of the serial reasoning happens in the chain. On word problems of finite groups, whose ceilings are known, small transformers trained from scratch, with or without reinforcement learning, attain the predicted numbers: chain-trained models solve every input length and fall to chance when the chain is erased, chainless models collapse to the ceiling as the input length grows, and open-weight reasoning models given the same problem in words return to the baseline without their chain. On MATH-500 and AIME, erasing the chain costs open reasoning models 0.52 to 0.82 accuracy, a sentence shuffle is harmless, and a token shuffle is as harmful as erasing; the same holds for checkpoints trained by GRPO with a correct or a random reward. The ceiling of a task therefore answers when a transformer can succeed without its chain of thought.
Chinese Translation
推理模型在回答之前会生成很长的思维链,然而这些链的内容是真正执行了计算工作,还是在很大程度上只是装饰性的,仍存在争议。我们通过将 Transformer 视为浅层电路来研究这个问题。对固定层数的一次前向传播具有常数深度,因此任何以常数次数运行该模型的过程都是浅层电路。我们将浅层电路在一个任务上能达到的最佳准确率称为该任务的天花板;如果某任务的天花板低于 1,则该任务是串行的。我们证明了关于串行任务的三个结果,它们对任意 Transformer 都成立,无论其如何训练。必要性:将思维链替换为任何不依赖其内容的东西,例如填充 token 或对问题的复述,会使准确率降至天花板;在最大串行任务上则降至随机水平。深度:任何浅层计算都无法写出一个准确率超过天花板的模型的思维链,甚至连近似写出也不行。局部性:答案与完成后的思维链仅相隔一次浅层传递,因此所有串行推理都发生在思维链中。在有限群的文字题上(其天花板已知),从头训练的小型 Transformer,无论是否使用强化学习,都达到预测的数值:经思维链训练的模型能解决任意输入长度,并且当思维链被抹除时降至随机水平;无思维链模型随着输入长度增长而坍缩到天花板;开放权重推理模型在给定相同的文字问题时,没有思维链会回到基线。在 MATH-500 和 AIME 上,抹除思维链会使开放推理模型损失 0.52 到 0.82 的准确率,打乱句子无害,而打乱 token 与抹除同样有害;对于由 GRPO 使用正确或随机奖励训练的检查点也同样如此。因此,任务的天花板回答了 Transformer 何时能在没有其思维链的情况下成功。
cs.AI / 179 / 2609.33141

On Device Agentic Operation Caches -- Classifier-Centric NL-to-Action Generation

端侧智能体操作缓存——以分类器为中心的自然语言到动作生成
Fereidouni, Moghis, Arnold, Anthony, Gulwani, Sumit, Marron, Mark, Siddique, A. B.
Abstract
Agentic AI is increasingly being embedded in software applications to provide natural language interfaces to features and functionality. In most cases these agents are powered by enterprise (100+ billion parameter) or frontier class large language models that require substantial computational resources run and depend on cloud hosted inference to handle the task of transforming natural language inputs into actionable software operations. This reliance on cloud-hosted inference introduces substantial network latency on top of LLM inference times, creates data privacy concerns, and, given the costs of running these models, can rapidly escalate expenses associated with supporting agentic features. This paper introduces a novel means of converting the NL-to-Action problem from a generative one into a classification-centric formulation via on-device operation caches. These caches allow an agentic system to handle frequently occurring classes of actions completely on-device -- reducing latency, enhancing privacy, and lowering operational costs. We show that for a classic NL-to-Formula task, generating Excel Formula in response to user requests, this approach reduces total inference cost by 56% when compared to cloud-only model-routing based inference and, on cache hits, reduces the latency to response latency by 5x.
Chinese Translation
智能体 AI 正越来越多地嵌入到软件应用中,为特性和功能提供自然语言界面。在大多数情况下,这些智能体由企业级(超过1000亿参数)或前沿级大型语言模型驱动,这些模型需要大量计算资源运行,并依赖云托管的推理来处理将自然语言输入转换为可执行软件操作的任务。这种对云托管推理的依赖在 LLM 推理时间之上引入了大量的网络延迟,引发了数据隐私问题,并且考虑到运行这些模型的成本,可能会迅速增加与支持智能体功能相关的费用。本文介绍了一种新颖的方法,通过端侧操作缓存将自然语言到动作的问题从生成式问题转换为以分类为中心的公式化表述。这些缓存允许智能体系统完全在设备上处理频繁出现的动作类别——从而降低延迟、增强隐私并降低运营成本。我们表明,对于经典的 NL-to-Formula 任务(即根据用户请求生成 Excel 公式),与仅基于云模型路由的推理相比,该方法将总推理成本降低了 56%,并且在缓存命中时,将延迟降低了5倍。
cs.AI / 180 / 2609.33146

LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks

LiteEvo:面向未见任务泛化的自动化、成本高效Harness演化
Choi, Euntae, Song, Sumin, Yoo, Sungjoo
Abstract
An LLM agent is defined by two things: the weights inside its model and the harness of components assembled around it. Harnesses are still handcrafted, and HarnessX, which evolves them automatically, starts each benchmark from a handcrafted harness, reports gains on the tasks it evolved on, and budgets 100 to 175 million meta-agent tokens per benchmark. We propose LiteEvo, a lightweight harness-evolution algorithm whose tool-free meta-agents mine agent trajectories for reusable components, curate them into a versioned library, and compose each round's harness from it, starting every benchmark from the same neutral harness and never naming the benchmark. Evolving on the graded tasks of five agentic benchmarks with a frozen Qwen3.5-9B, LiteEvo lifts pass@2 by 10.5 to 67.7pp and reaches comparable or higher pass@2 than a reproduction of HarnessX (71.0 against 67.3 on average) at 13.0 lower mean API cost. Harnesses evolved on train tasks keep their gains on unseen test tasks of four benchmarks, and LiteEvo also lifts Claude Code with Sonnet 4.6 by 1.2 to 71.4pp.
Chinese Translation
一个LLM智能体由两件事定义:其模型内部的权重以及围绕它组装的组件Harness。Harness仍然是手工设计的,而HarnessX虽然能自动演化Harness,但它从手工设计的Harness开始每个基准测试,报告在其演化任务上的增益,并为每个基准测试预算1亿到1.75亿个元智能体token。我们提出LiteEvo,一种轻量级的Harness演化算法,其无工具元智能体从智能体轨迹中挖掘可重用组件,将其整理成版本化库,并从中组合每一轮的Harness,每个基准测试都从同一个中性Harness开始,且从不提及基准测试的名称。在五个智能体基准测试的评分任务上,使用冻结的Qwen3.5-9B进行演化,LiteEvo将pass@2提升了10.5到67.7个百分点,并且与HarnessX的复现相比,达到相当或更高的pass@2(平均71.0对67.3),同时平均API成本降低了13.0。在训练任务上演化的Harness在四个基准测试的未见测试任务上保持其增益,并且LiteEvo还将Claude Code与Sonnet 4.6提升了1.2到71.4个百分点。
cs.AI / 181 / 2609.33149

Not Too Hard, Not Too Easy: Learning from Intermediate States for LLM Structured Reasoning

难易适中:从中间状态学习用于LLM结构化推理
Chen, Hongbo, Lu, Guohua, Dang, Ting, Jia, Hong
Abstract
A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how to apply this principle to structured reasoning tasks such as Sudoku and maze solving. In these tasks, a model can repeatedly revise an incomplete or incorrect candidate solution until it satisfies the problem's constraints. The intermediate candidate solutions along this trajectory provide natural training examples: some are already solved, some cannot yet be repaired by the model, and others lie at its current frontier of achievable progress. We therefore investigate whether pretrained language models can learn to revise such states and whether training on states at this frontier improves reasoning more broadly. To achieve this, we couple a pretrained language-model backbone with a recurrent updater that repeatedly revises an explicit solution state, using the same parameters at every update step. We further introduce Frontier-Oriented Curation Using Self-trajectories (FOCUS), which selects training states from trajectories generated by the current model. FOCUS measures how much the model improves each state within a fixed number of recurrent updates and prioritizes states from which it can make substantial progress. With Qwen3-1.7B, FOCUS achieves 64.4% exact solve accuracy on Sudoku-Extreme and 91.1% on Maze-Hard, with similar gains observed across five Qwen and Llama backbones spanning 1.7B to 8B parameters. We further observe zero-shot transfer in the adapted LLM to mathematical reasoning and code execution, even when the recurrent updater is disabled and no downstream fine-tuning is performed.
Chinese Translation
有效学习的一个常见原则是练习那些既非已经掌握、也非难到无法取得进展的材料。我们探讨如何将这一原则应用于结构化推理任务,如数独和迷宫求解。在这些任务中,模型可以反复修改一个不完整或不正确的候选解,直到满足问题的约束。沿着这一轨迹的中间候选解提供了自然的训练示例:一些已经解决,一些模型尚无法修复,另一些则位于其当前可取得进展的前沿。因此,我们研究预训练语言模型能否学会修改这样的状态,以及在此前沿状态上进行训练是否能更广泛地提高推理能力。为此,我们将预训练语言模型主干与一个循环更新器相结合,该更新器反复修改显式的解状态,并在每个更新步骤中使用相同的参数。我们进一步提出了基于自轨迹的前沿导向整理(FOCUS),它从当前模型生成的轨迹中选择训练状态。FOCUS衡量模型在固定数量的循环更新中对每个状态的改进程度,并优先选择能够取得实质性进展的状态。在Qwen3-1.7B上,FOCUS在Sudoku-Extreme上达到了64.4%的精确求解准确率,在Maze-Hard上达到91.1%,在五个Qwen和Llama主干模型(参数规模从1.7B到8B)上观察到类似的提升。我们进一步观察到,即使禁用循环更新器且不进行下游微调,适应后的LLM在数学推理和代码执行上也表现出零样本迁移。
cs.AI / 182 / 2609.33181

SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought

SeOPD:通过自生成思维链的在线策略蒸馏实现自进化大语言模型
Chen, Xiaoshu, Wong, Xiangyu, Zhou, Sihang, Liang, Ke, Liu, Xinwang
Abstract
Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without external PI, the resulting gains remain limited. In this work, we explore whether LLMs can achieve comparable self-improvement without external PI. Our key observation is that a single LLM can support multiple reasoning modes, such as deep-thinking and non-thinking modes, with deep thinking generating additional information during reasoning. Based on this observation, we propose Self-Evolving Online Policy Distillation (SeOPD), which enables LLMs to distill and internalize information generated by their own chain of thought (CoT). Specifically, it (1) generates CoT with the deep-thinking mode, (2) produces responses with the non-thinking mode, and (3) uses the generated CoT as PI to provide token-level supervision for the non-thinking response, allowing new information inferred during reasoning to guide the non-thinking mode and be internalized into the shared model parameters, thereby improving both non-thinking and deep-thinking capabilities. Extensive experiments across LLMs and tasks demonstrate the effectiveness of SeOPD.
Chinese Translation
近期在线策略自蒸馏(OPSD)的进展表明,大语言模型(LLMs)可以通过利用外部特权信息(PI)(如人工标注或外部环境反馈)来提升其能力。然而,获取准确标注和构建复杂环境通常需要大量人力与计算资源,限制了 OPSD 的可扩展性。尽管近期少数研究探索了无需外部 PI 的自我提升,但所获收益仍然有限。在本工作中,我们探究 LLMs 能否在无需外部 PI 的情况下实现可比的自我提升。我们的关键观察是,单个 LLM 可以支持多种推理模式,例如深度思考模式与非思考模式,其中深度思考在推理过程中会生成额外信息。基于这一观察,我们提出自进化在线策略蒸馏(SeOPD),使 LLMs 能够蒸馏并内化其自身思维链(CoT)生成的信息。具体而言,它(1)以深度思考模式生成 CoT,(2)以非思考模式生成回复,(3)将生成的 CoT 作为 PI,为非思考回复提供 token 级监督,使推理过程中推断出的新信息能够指导非思考模式,并内化到共享模型参数中,从而同时提升非思考与深度思考能力。在多种 LLMs 和任务上的大量实验证明了 SeOPD 的有效性。
cs.AI / 183 / 2609.33182

Unlocking Latent Personalization in LLMs

解锁大语言模型中的潜在个性化
Chen, Wei, Zhu, Guanghui, Cai, Zhongliang, Huang, Yihua
Abstract
Large language models (LLMs) are increasingly expected to adapt to individual users, yet effective personalization remains challenging when only limited user-specific samples are available. In this work, we take an alternative perspective: pretrained LLMs may already possess latent capacity for personalization, and a few user samples may therefore suffice to guide the model toward user-aligned behavior with minimal user-specific adaptation. From this perspective, we propose LatentPersonal, a framework that formulates personalization as navigation in a shared latent adaptation space. LatentPersonal infers a compact latent representation from a few user samples to guide user-specific model adaptation, regularized with a variational information bottleneck to encourage compact preference representations. We instantiate LatentPersonal with LoRA, leveraging its low-rank parameterization as a natural low-dimensional adaptation space for personalization. By simply inserting a user-specific guidance vector between the shared low-rank factors, the model can navigate toward personalized adaptations through lightweight inference of this compact representation, without updating the shared LoRA parameters. Experiments across multiple personalization datasets demonstrate that LatentPersonal substantially reduces user-specific adaptation overhead while achieving effective personalization from only a few user-specific interactions, with particularly strong performance in the one-shot regime.
Chinese Translation
大语言模型(LLM)越来越多地被期望适应个体用户,然而当只有有限的用户特定样本可用时,有效的个性化仍然具有挑战性。在这项工作中,我们采取另一种视角:预训练的 LLM 可能已经具备个性化的潜在能力,因此,少量用户样本可能足以引导模型以最小的用户特定适应来实现用户对齐的行为。从这个视角出发,我们提出了 LatentPersonal,一个将个性化表述为在共享潜在适应空间中进行导航的框架。LatentPersonal 从少量用户样本中推断出一个紧凑的潜在表示,以指导用户特定的模型适应,并使用变分信息瓶颈进行正则化,以鼓励紧凑的偏好表示。我们用 LoRA 实例化 LatentPersonal,利用其低秩参数化作为个性化的自然低维适应空间。通过简单地在共享低秩因子之间插入一个用户特定的引导向量,模型可以通过对这个紧凑表示的轻量级推理来导航到个性化适应,而无需更新共享的 LoRA 参数。在多个个性化数据集上的实验表明,LatentPersonal 大幅减少了用户特定的适应开销,同时仅从少量用户特定交互中实现有效的个性化,在单样本机制中表现尤为强劲。
cs.AI / 184 / 2609.33196

Are Benchmarks Reliable? Toward Structural Diagnosis via Sample-Level Capability Boundaries

基准可靠吗?通过样本级能力边界走向结构诊断
Hu, Haiquan, Liang, Yuzhu, Tang, Weicheng, Li, Yanzeng, Shi, Yao, Wang, Tian
Abstract
Evaluating large language models (LLMs) relies heavily on benchmark scores, yet aggregate metrics can obscure whether benchmark samples reliably support model comparison. We introduce \textbf{BSDProbe}, a sample-level framework for \emph{benchmark structural diagnosis} that estimates capability boundaries from repeated-response trajectories along ordered model axes. BSDProbe summarizes samples by boundary position, boundary width, boundary-signal validity, and order consistency, then aggregates them into benchmark-level structural profiles. Experiments on six benchmarks show that benchmark reliability is axis-conditioned and heterogeneous: GSM8K and MATH exhibit the most stable measurement structures, MMLU and TriviaQA are relatively stable but heterogeneous, while GPQA and PopQA show stronger axis-conditioned risks. These profiles remain consistent across Qwen3, Qwen2.5, and cross-model axes. BSDProbe further selects compact high-value subsets whose model discriminability reaches up to $8.58\times$ that of the full benchmark. These results suggest that reliable benchmark use requires examining sample-level capability boundaries beyond leaderboard scores.
Chinese Translation
评估大语言模型(LLMs)高度依赖基准分数,然而聚合指标可能掩盖基准样本是否可靠地支持模型比较。我们提出 BSDProbe,一个用于基准结构诊断的样本级框架,它沿有序模型轴从重复响应轨迹中估计能力边界。BSDProbe 通过边界位置、边界宽度、边界信号有效性和顺序一致性来概括样本,然后将它们聚合成基准级别的结构概貌。在六个基准上的实验表明,基准可靠性是轴条件依赖且异质的:GSM8K 和 MATH 表现出最稳定的测量结构,MMLU 和 TriviaQA 相对稳定但异质,而 GPQA 和 PopQA 显示出更强的轴条件风险。这些概貌在 Qwen3、Qwen2.5 和跨模型轴上保持一致。BSDProbe 进一步选择紧凑的高价值子集,其模型区分度最高可达完整基准的 8.58 倍。这些结果表明,可靠的基准使用需要超越排行榜分数,考察样本级能力边界。
cs.AI / 185 / 2609.33208

WorldAgent: Verification-Guided Agentic Physical World Construction

WorldAgent:验证引导的智能体物理世界构建
Wang, Caoliwen, Wang, Mengdi, Chen, Yige, Wu, Zejia, Huang, Bowen, Chen, Siyuan, Chen, Guanxiong, Wei, Lifu, Zhang, Heng, Zhang, Qinghai, Yang, Yin, Yang, Guandao, Xiong, Shiying, Wang, Peng, Jiang, Chenfanfu, Chen, Peter Yichen
Abstract
Constructing complex physical worlds from language requires coordinating extensive 3D environments, detailed structures and objects at different spatial scales, and interacting physical processes under both stated goals and implicit physical constraints. We present WorldAgent, an agentic framework for verification-guided physical world construction from a single natural-language prompt, without iterative user debugging. A world construction layer expands the prompt into a structured world specification and uses physical knowledge to build scenes and run numerical simulations. After every step, a verification layer inspects scene geometry and simulation states alongside rendered views. Failed checks guide automatic revisions to the specification and re-execution of the affected steps. Accepted worlds pass the required checks and remain editable for further inspection and resimulation. We introduce AgenticSimBench, on which WorldAgent achieves the best scores among the evaluated agent-based methods on five of seven metrics. In a 26-participant user study, it receives the highest mean ratings across all four criteria.
Chinese Translation
从语言构建复杂的物理世界需要协调大规模的3D环境、不同空间尺度下的详细结构与物体,以及在明确目标和隐含物理约束下相互作用的物理过程。我们提出WorldAgent,一个智能体框架,用于从单个自然语言提示进行验证引导的物理世界构建,无需迭代式用户调试。世界构建层将提示扩展为结构化的世界规范,并利用物理知识构建场景并运行数值模拟。每一步之后,验证层检查场景几何和模拟状态以及渲染视图。检查失败会引导对规范的自动修订以及对受影响步骤的重新执行。被接受的世界通过所需检查,并保持可编辑,以便进一步检查和重新模拟。我们引入AgenticSimBench,在该基准上,WorldAgent在七个指标中的五个上,在所评估的基于智能体的方法中取得了最佳分数。在一项有26名参与者的用户研究中,它在所有四项标准上获得了最高的平均评分。
cs.AI / 186 / 2609.33243

CodeSkill: Latent Skill Abstraction for Long-Horizon Code Agents

CodeSkill:面向长时程代码智能体的潜在技能抽象
Wu, Song-Li, Wang, Jingyi, Du, Zhaocheng, Gan, Weinan, Liu, Weiwen
Abstract
Code agents require long-horizon decision-making over complex interaction trajectories. However, existing reinforcement learning (RL) approaches typically optimize behavior at the token level, creating a mismatch between low-level generation and high-level behavioral reasoning. This limitation leads to inefficient exploration and weak credit assignment under sparse rewards. Moreover, while large-scale agent trajectories often contain recurring multi-step behavioral patterns, their noisy token-level representations hinder effective experience reuse. To address these challenges, we propose CodeSkill, a framework that adapts hierarchical latent skill modeling to the code agent domain. CodeSkill first leverages a teacher model to distill both successful and failed trajectories into multi-level textual abstractions. It then integrates temporal variational inference with reinforcement learning to map these discrete semantics into continuous latent variables, while an adaptive boundary mechanism dynamically gates skill transitions based on execution feedback. The learned skills are injected into a frozen LLM policy as latent semantic prefixes, enabling optimization in a compact semantic space rather than over raw token sequences. By shifting RL from token-level exploration to experience-level reasoning, CodeSkill improves optimization efficiency and long-horizon behavioral coherence. Extensive experiments demonstrate that CodeSkill achieves highly competitive performance against strong open-weight baselines across diverse general and industrial coding benchmarks. Furthermore, the learned skills exhibit strong transferability and robust cross-domain generalization, highlighting the effectiveness of explicit behavioral abstraction for scalable agentic code generation.
Chinese Translation
代码智能体需要在复杂交互轨迹上进行长时程决策。然而,现有的强化学习(RL)方法通常在 token 级优化行为,导致低级生成与高级行为推理之间不匹配。这一限制导致在稀疏奖励下探索效率低下和信用分配薄弱。此外,尽管大规模智能体轨迹通常包含重复出现的多步行为模式,但其噪声 token 级表示阻碍了有效的经验复用。为应对这些挑战,我们提出 CodeSkill,一个将分层潜在技能建模适配到代码智能体领域的框架。CodeSkill 首先利用教师模型将成功和失败的轨迹蒸馏为多级文本抽象。然后,它将时序变分推断与强化学习相结合,将这些离散语义映射为连续潜在变量,同时自适应边界机制根据执行反馈动态控制技能转换。学到的技能作为潜在语义前缀注入到冻结的 LLM 策略中,使得在紧凑语义空间而非原始 token 序列上进行优化成为可能。通过将 RL 从 token 级探索转向经验级推理,CodeSkill 提高了优化效率和长时程行为一致性。大量实验表明,CodeSkill 在多样化的通用和工业编码基准上,与强大的开放权重基线相比,取得了极具竞争力的性能。此外,学到的技能展现出强大的可迁移性和稳健的跨域泛化能力,凸显了显式行为抽象对于可扩展智能体代码生成的有效性。
cs.AI / 187 / 2609.33244

ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents

ActiveMem:面向长时程智能体的动态潜在记忆树
Wu, Song-Li, Wang, Jingyi, Du, Zhaocheng, Gan, Weinan
Abstract
Large Language Model (LLM) agents increasingly rely on external memory to support long-horizon reasoning and decision making. Existing memory systems typically retrieve historical trajectories or summaries as independent context fragments, overlooking the procedural dependencies underlying multi-step execution. As memory scales, such flat retrieval introduces context fragmentation and cross-task interference, leading to structurally inconsistent reasoning trajectories. We propose ActiveMem, a hierarchical memory framework that recursively organizes agent experiences into dependency-aware latent execution trees. ActiveMem abstracts trajectories into reusable subtask nodes while explicitly preserving execution transitions, enabling coherent reasoning-path retrieval conditioned on the current execution state. To support continual adaptation, ActiveMem further learns dynamic memory expansion, retrieval, and pruning policies through reinforcement learning. Experiments across various agent benchmarks demonstrate that ActiveMem consistently improves task completion, reasoning stability, and memory efficiency over existing memory-based agents. Moreover, ActiveMem enables compact open-weight models to achieve competitive performance with substantially larger proprietary systems.
Chinese Translation
大语言模型(LLM)智能体日益依赖外部记忆来支持长时程推理与决策。现有记忆系统通常将历史轨迹或摘要作为独立上下文片段进行检索,忽视了多步执行背后的过程依赖。随着记忆规模扩大,这种扁平检索会引入上下文碎片化和跨任务干扰,导致推理轨迹在结构上不一致。我们提出 ActiveMem,一种分层记忆框架,递归地将智能体经验组织为具有依赖感知的潜在执行树。ActiveMem 将轨迹抽象为可复用的子任务节点,同时显式保留执行转移,从而能够以当前执行状态为条件检索连贯的推理路径。为支持持续适应,ActiveMem 进一步通过强化学习学习动态记忆扩展、检索和剪枝策略。在多个智能体基准上的实验表明,与现有基于记忆的智能体相比,ActiveMem 持续提升任务完成度、推理稳定性和记忆效率。此外,ActiveMem 使紧凑的开放权重模型能够达到与规模大得多的专有系统相当的性能。
cs.AI / 188 / 2609.33260

CORTEX: A Verified Experience Layer for Generalist Agents

CORTEX:面向通用智能体的已验证经验层
Keerthana, Garapati, Gupta, Manik
Abstract
An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is still valid, when it must be adapted, and when it should be discarded. We introduce CORTEX (Contextual Orchestration and Reuse of Task EXperience), a general AI systems framework that connects specialized agents through an external layer of verified experience. Each episode records its task conditions, source and tool state, decisive predicates, proof trace, verifier, and outcome. A meta-controller chooses exact replay, checked adaptation, fresh synthesis, or escalation. Accepted episodes can become task patterns and procedural strategies through a challenge-driven development loop. This gives the system an implicit competence layer that can grow without changing model weights. We formalize system contracts for exact replay and source-version separation, and derive when reuse saves computation. A controlled two-domain implementation tests the exact-replay core on 1,000 synthetic cases. Complete-family holdouts test procedural transfer on 1,000 new-family cases across eight clinical and policy splits, with complete fresh-evidence grounding and perfect invariance to irrelevant-field and insertion-order perturbations. The transfer trace exposes the work required for verified strategy execution. These results establish an initial path toward general intelligence through reusable procedures, typed experience, and developmental transfer.
Chinese Translation
智能体今天可以解决一个任务,明天却可能在新的事实、工具或主导知识下面临同一任务。大多数智能体系统能够检索相关文本或回忆先前对话,但缺乏一种原则性的方式来决定先前的解决方案何时仍然有效、何时必须调整以及何时应被丢弃。我们提出 CORTEX(Contextual Orchestration and Reuse of Task EXperience,任务经验的情境化编排与复用),一个通用的 AI 系统框架,通过一个外部已验证经验层连接专用智能体。每个片段记录其任务条件、源和工具状态、决定性谓词、证明轨迹、验证器以及结果。一个元控制器选择精确重放、经检查的适应、全新合成或升级处理。被接受的片段可以通过挑战驱动的开发循环成为任务模式与过程性策略。这为系统提供了一个隐式能力层,无需改变模型权重即可增长。我们形式化了精确重放和源版本分离的系统契约,并推导出复用何时节省计算。一个受控的双领域实现测试了精确重放核心在 1,000 个合成用例上的表现。完全家族留出测试在跨八个临床和政策划分的 1,000 个新家族案例上测试了过程性迁移,具有完整的新证据基础,并且对无关字段和插入顺序扰动具有完美不变性。迁移轨迹揭示了经过验证的策略执行所需的工作。这些结果通过可复用过程、类型化经验和发展性迁移,为通向通用智能建立了一条初始路径。
cs.AI / 189 / 2609.33268

LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models

LSTMem:面向大语言模型的分层长短期在线记忆
Shi, Xianglong, Yang, Ruijie, Zhao, Sirui, Yin, Shukang, Bian, Zihao, Yi, Tinghao, Chen, Enhong
Abstract
Large language models increasingly serve as long-horizon assistants and agents, where they must both accumulate information across interactions and make the relevant parts available when later requests depend on them. Existing compact online memories typically use a single persistent state both to accumulate history and to serve readout, so what the memory stores cannot be controlled separately from what it exposes to the current computation. We propose LSTMem, an LSTM-inspired online memory that instead equips each layer of a frozen LLM with two matrix-valued states: a cell state that accumulates history and a hidden state whose readouts correct the backbone's attention. Input and forget gates control what the cell stores, while an output gate separately controls what the cell exposes through the hidden state. LSTMem further connects memory across depth through forward hidden-state propagation and block-end feedback, and uses higher-layer reconstruction gradients to refine lower-layer cell states before rebuilding hidden states from shallow to deep layers. Across memory benchmarks on Qwen3-4B-Instruct, LSTMem consistently improves MemoryAgentBench, LoCoMo, and HotpotQA over the plain backbone. Comparisons further show that the LSTM-based memory formulation outperforms an associative-memory counterpart, while removing cross-layer hidden-memory propagation degrades performance. These results demonstrate the benefits of separating memory accumulation from memory expression and organizing memory hierarchically across model depth. The code is available at https://github.com/Longchentong/LSTMem.
Chinese Translation
大语言模型越来越多地充当长时程助手和智能体,它们必须在交互过程中累积信息,并在后续请求依赖这些信息时使相关部分可用。现有的紧凑在线记忆通常使用单个持久状态来累积历史并用于读取,因此记忆存储的内容无法与其暴露给当前计算的内容分开控制。我们提出LSTMem,一种受LSTM启发的在线记忆,它为冻结的LLM的每一层配备了两个矩阵值状态:一个累积历史的细胞状态和一个其读出用于校正骨干网络注意力的隐藏状态。输入门和遗忘门控制细胞存储的内容,而输出门单独控制细胞通过隐藏状态暴露的内容。LSTMem进一步通过前向隐藏状态传播和块末反馈跨深度连接记忆,并使用更高层的重建梯度来细化低层细胞状态,然后从浅层到深层重建隐藏状态。在Qwen3-4B-Instruct的记忆基准测试中,LSTMem在MemoryAgentBench、LoCoMo和HotpotQA上始终优于普通骨干网络。比较进一步表明,基于LSTM的记忆公式优于关联记忆对应方案,而移除跨层隐藏记忆传播会降低性能。这些结果证明了将记忆累积与记忆表达分离以及跨模型深度分层组织记忆的好处。代码可在 https://github.com/Longchentong/LSTMem 获取。
cs.AI / 190 / 2609.33270

Structured Sparse Memory for Recurrent Reasoning

用于循环推理的结构化稀疏记忆
Zhao, Zixuan, Wheeler, Samuel, Getty, Neil, Duan, Xiaotian, Stevens, Rick, Xia, Fangfang
Abstract
Recurrent models trained from scratch have recently become competitive on ARC-style reasoning tasks, but the usual framing around small recurrent backbones overlooks two important parts of the system: task-conditioned memory and synthetic augmentation data. We study this regime through CHARM, a compact hybrid ARC model that combines recurrent reasoning with structured task memory, synthetic data, and inference-time aggregation. In existing approaches, task-conditioned memory supplies a large hidden source of capacity, reaching more than 30x the size of the recurrent backbone. We introduce a compositional sparse embedding (CoSE) for task conditioning that reduces learned task-memory parameters by over 90% while improving pass@2 in controlled ARC ablations. For the recurrent backbone, recurrent depth helps only when balanced with learning horizon. Combining these ingredients, our system reaches 84% pass@2 on ARC-AGI-1 and 46.7% pass@2 on ARC-AGI-2 public evaluation. The benefits of structured memory also generalize to unseen puzzles and other domains. Our code, dataset, and model checkpoints are available at https://github.com/water-vapor/charm.
Chinese Translation
从头训练的循环模型最近在ARC风格的推理任务上已变得具有竞争力,但通常围绕小型循环骨干的框架忽略了系统的两个重要部分:任务条件化记忆和合成增强数据。我们通过CHARM研究这一范式,CHARM是一个紧凑的混合ARC模型,将循环推理与结构化任务记忆、合成数据和推理时聚合相结合。在现有方法中,任务条件化记忆提供了大量的隐藏容量来源,达到循环骨干网络大小的30倍以上。我们引入了一种用于任务条件化的组合稀疏嵌入(CoSE),将学习的任务记忆参数减少了90%以上,同时在受控的ARC消融实验中提高了pass@2。对于循环骨干,循环深度只有与学习视野平衡时才有帮助。结合这些要素,我们的系统在ARC-AGI-1上达到84%的pass@2,在ARC-AGI-2公开评估上达到46.7%的pass@2。结构化记忆的优势还能泛化到未见过的谜题和其他领域。我们的代码、数据集和模型检查点可在 https://github.com/water-vapor/charm 获取。
cs.AI / 191 / 2609.33271

Next Thoughts Are Distributions: Generative Autoregressive Reasoning in the Latent Space

下一个思维是分布:潜在空间中的生成式自回归推理
Li, Yang, Wang, Yi, Huang, Shiyuan, Liu, Yang, Wang, Hao, Mao, Chengzhi
Abstract
Reasoning problems often admit multiple valid ways to proceed. Continuous reasoning promises to move computation beyond language tokens into a more compact latent space, but representing several plausible ways to think next remains difficult. We introduce Autoregressive Thought Flow (ATF), which models the next continuous thought as a multimodal distribution. A causal autoregressive model performs the reasoning computation, while a lightweight diffusion head generates a plausible next thought from the resulting condition. The sampled thought is fed back into the model, allowing continuous reasoning to unfold for a variable number of steps while preserving the pretrained backbone. Across mathematical reasoning tasks, ATF improves accuracy with compact latent traces and benefits from reinforcement learning and additional test-time thinking. Multi-sample evaluation shows broader solution coverage, indicating that its multimodal predictions capture useful diversity among reasoning paths. Our results suggest that continuous reasoning is more effective when multiple possible next thoughts remain available rather than being collapsed into a single prediction.
Chinese Translation
推理问题通常允许多种有效的前进方式。连续推理有望将计算从语言词元推进到更紧凑的潜在空间,但表示若干种可能的下一个思维仍然困难。我们提出自回归思维流(Autoregressive Thought Flow, ATF),它将下一个连续思维建模为多峰分布。一个因果自回归模型执行推理计算,而一个轻量级扩散头根据所得条件生成一个合理的下一个思维。采样的思维被反馈回模型,使连续推理能够展开可变步数,同时保留预训练主干。在数学推理任务上,ATF 以紧凑的潜在轨迹提高了准确率,并受益于强化学习和额外的测试时思考。多样本评估显示出更广的解覆盖范围,表明其多峰预测捕获了推理路径中有用的多样性。我们的结果表明,当多种可能的下一思维保持可用而不是坍缩为单一预测时,连续推理更有效。
cs.AI / 192 / 2609.33276

ChronoFlow: Hierarchical Flow Matching for Irregular Time Series Generation

ChronoFlow:面向不规则时间序列生成的分层流匹配
Kim, Changhun, Jang, Sunguk, Lee, Jeongjun, Choi, Juhwan, Hahn, Sangchul, Chrysos, Grigorios, Yang, Eunho, Lee, Juho
Abstract
Recent advances in generative modeling have substantially improved time series generation, yet most existing methods either assume a regular temporal grid or focus on feature dynamics under a given sampling structure. This makes them illsuited for generating irregular time series in their native form, where a model must capture not only feature values, but also how many observations occur, when they occur, and which features are observed together. To address this heterogeneous generation problem, we propose ChronoFlow, a unified hierarchical flow matching framework organized by statistical granularity. Following a coarse-to-fine hierarchy, ChronoFlow first generates observation counts and feature-wise frequencies, then jointly generates observation times and feature co-observation patterns, and finally generates values conditioned on the realized pattern. This turns a complex joint generation problem into structurally aligned subproblems while preserving their dependencies. To evaluate complete irregular time series generation, we introduce complementary metrics spanning sample realism, sampling structure, value fidelity, and temporal and cross-feature dependencies, and validate them through controlled corruptions. Across five benchmarks, ChronoFlow achieves strong improvements in generation fidelity over existing baselines, while factorization studies support the proposed hierarchy. Our code is available at https://anonymous.4open.science/r/ChronoFlow.
Chinese Translation
近年来,生成建模的进展显著提升了时间序列生成,但大多数现有方法要么假设规则的时间网格,要么关注给定采样结构下的特征动态。这使得它们不适合以原生形式生成不规则时间序列,其中模型不仅必须捕捉特征值,还要捕捉观测发生的次数、发生的时间以及哪些特征被共同观测。为了解决这种异质生成问题,我们提出了ChronoFlow,一个按统计粒度组织的统一分层流匹配框架。遵循从粗到细的层次,ChronoFlow首先生成观测计数和特征频率,然后联合生成观测时间和特征共观测模式,最后基于已实现的模式生成值。这将复杂的联合生成问题转化为结构对齐的子问题,同时保留它们的依赖关系。为了评估完整的不规则时间序列生成,我们引入了涵盖样本真实性、采样结构、值保真度以及时间和跨特征依赖的互补指标,并通过受控损坏进行验证。在五个基准上,ChronoFlow在生成保真度上相比现有基线取得了显著提升,而分解研究支持所提出的层次结构。我们的代码可在https://anonymous.4open.science/r/ChronoFlow获取。
cs.AI / 193 / 2609.33282

Multi-Dimensional Comparative Scale Construction for Efficient Personalized Subjective Judgment in High-Traffic Applications

高流量应用中高效个性化主观判断的多维比较量表构建
Shi, Xianglong, Liu, Shifeng, Zhao, Sirui, Yuan, Shengming, Chen, Enhong
Abstract
Subjective judgments are central to many high-traffic applications, but subjective intensity is difficult to quantify and perceptions vary substantially across individuals. To address these challenges, we propose a pairwise comparative framework for multi-dimensional scale construction. By comparing case-person pairs along case and profile dimensions, the framework constructs relative scales that capture both fine-grained intensity and individual variation. To support practical high-traffic deployment, we optimize both offline scale construction and online inference. For scale construction, we combine sparse Elo comparisons with multi-judge voting, cutting the comparison cost from $O(N^2)$ to $O(NK)$ for $N$ objects and a budget of $K$ opponents per object, while limiting reliance on any single judge. For inference, we propose SubJudge, a System One model for personalized scoring with Batchwise Preference Optimization (BPO). Using Bradley-Terry comparisons, BPO trains the model to learn relative orderings, and SubJudge reads a continuous score from digit-token probabilities at the first response position, requiring only one forward pass per criterion and reducing the inference complexity to $O(1)$. Experiments on PluriHarms and iNews show that our 9B models match or surpass the evaluated frontier LLMs on multiple metrics. On the H100 GPU, SubJudge achieves an approximately $1.29\times$ to $261\times$ speedup in mean inference latency over Qwen3.5-9B with different thinking budgets. The code is available at https://github.com/Longchentong/SubJudge.
Chinese Translation
主观判断是许多高流量应用的核心,但主观强度难以量化,且个体间的感知差异显著。为应对这些挑战,我们提出一种用于多维度量表构建的成对比较框架。通过沿案例和画像维度比较案例-人物对,该框架构建相对量表,以同时捕捉细粒度强度和个体差异。为支持实际高流量部署,我们优化离线量表构建与在线推理。对于量表构建,我们结合稀疏Elo比较与多评委投票,将N个对象的比较成本从O(N^2)降至O(NK)(每个对象的对手预算为K),同时限制对任一单个评委的依赖。对于推理,我们提出SubJudge,一种基于批式偏好优化(BPO)的System One个性化评分模型。使用Bradley-Terry比较,BPO训练模型学习相对排序,而SubJudge从第一个响应位置处的数字token概率中读取连续分数,每个准则仅需一次前向传播,将推理复杂度降至O(1)。在PluriHarms和iNews上的实验表明,我们的9B模型在多个指标上达到或超过所评估的前沿LLM。在H100 GPU上,在不同思考预算下,SubJudge相比Qwen3.5-9B在平均推理延迟上实现约1.29倍至261倍的加速。代码见 https://github.com/Longchentong/SubJudge。
cs.AI / 194 / 2609.33284

RINI: Seeing the Prior Is Not Enough

RINI:仅看到先前文献是不够的
Du, Hongyi, Zhang, Tianyi, Wang, Heng, Gao, Zhelun, Liu, Yimei, Luo, Ambrose, Hao, Annie, Ni, Jiayan, Han, Jiawei, You, Jiaxuan
Abstract
A research proposal can describe an established mechanism correctly while claiming to introduce it. We study whether providing the earlier paper corrects such contribution claims. Three controlled experiments compare proposals generated with a contribution-bearing prior and a same-topic control. Providing the prior yields no clear aggregate reduction in unsupported novelty. Human analysis of 175 interpretable exposed proposals finds that 137 recognize the prior's relevance, but 61 correctly attribute the established contribution. Of 71 proposed remaining distinctions, 37 are covered by the same prior. We introduce Research Idea Novelty Inspection (RINI), which audits contribution claims against evidence, checks the remaining distinction, and applies local revisions. Five human annotators evaluate 1,080 original-revision pairs across three methods. On the same 240 originals judged to require correction, successful repair is 11.7% for Self-Revision, 39.1% for Retrieve-and-Revise, and 72.2% for RINI, with research tasks weighted equally. The improvement over same-evidence direct revision is 33.0 percentage points. The revised proposals retain their research questions and technical methods. These results motivate explicit contribution attribution when using literature to generate and revise research proposals.
Chinese Translation
研究提案可以正确地描述一个已确立的机制,却声称自己引入了该机制。我们研究提供较早的论文是否能纠正此类贡献声明。三项对照实验比较了使用带有贡献声明的先前文献和同主题对照生成的提案。提供先前文献并未在总体上明确减少无支撑的新颖性。对175份可解释的、暴露于先前文献的提案进行人工分析发现,137份认识到先前文献的相关性,但61份正确归因了已确立的贡献。在71个提出的剩余区别中,37个已被同一先前文献覆盖。我们提出研究想法新颖性检查(Research Idea Novelty Inspection, RINI),它依据证据审查贡献声明,检查剩余区别,并进行局部修订。五名人工标注者评估了三种方法下的1,080对原始-修订对。在同样被判定需要纠正的240个原始提案上,Self-Revision(自修订)的成功修复率为11.7%,Retrieve-and-Revise(检索并修订)为39.1%,RINI为72.2%,研究任务等权。相比相同证据的直接修订,提升为33.0个百分点。修订后的提案保留了其研究问题和技术方法。这些结果说明,在使用文献生成和修订研究提案时,需要进行显式的贡献归属。
cs.AI / 195 / 2609.33287

Feedback Makes Perfect: A Closed-Loop Framework for NL-to-STL Translation

反馈铸就完美:一种NL-to-STL翻译的闭环框架
Ye, Bowen, Yin, Xiang
Abstract
Signal Temporal Logic (STL) enables rigorous verification and control of cyber-physical systems, but writing correct specifications requires expertise that most requirement holders lack. Large language models can translate natural-language (NL) requirements into STL, yet stronger translators alone approach an accuracy ceiling. We argue that this ceiling stems from how the task is posed: one-shot, open-loop translation is somewhat ill-defined. Natural language is ambiguous, and, more fundamentally, what a person writes may not always be what they intend, so the target specification is not fully contained in the input text. We therefore reformulate NL-to-STL translation as a closed-loop feedback process. Each generated formula is translated back into natural language for the user to check, and natural-language corrections drive revision until the user accepts the specification. Users never read or write formal syntax. This framework rests on an asymmetry familiar from feedback control theory. The forward path, from ambiguous language to formal logic, is hard and error-prone. The feedback path, from structured STL back to language, can be made highly precise, and a precise feedback path lets an imprecise forward path achieve precise closed-loop behavior. Experiments on 500 expert-authored requirements and seven LLMs support this view. Back-translated explanations agree with expert judgments in 99.5\% of cases. Closed-loop refinement raises strong models from about 89\% open-loop accuracy to 98.0--99.2\%, and yields gains of over 30 percentage points for weaker models (e.g., 17.6\%$\rightarrow$48.0\%). Ablations show these gains come from the semantic content of the feedback rather than from repeated attempts. An expert audit and a 280-session user study further confirm the reliability of the loop. We also identify a capability threshold above which feedback no longer helps.
Chinese Translation
信号时序逻辑(STL)能够对信息物理系统进行严格的验证与控制,但编写正确的规范需要大多数需求持有者所缺乏的专业知识。大型语言模型可以将自然语言(NL)需求翻译成STL,然而仅凭更强大的翻译器也会接近一个准确率上限。我们认为,这一上限源于任务的提出方式:一次性、开环翻译在某种程度上是不明确的。自然语言是模糊的,更根本的是,一个人所写的内容可能并不总是他们意图表达的内容,因此目标规范并未完全包含在输入文本中。因此,我们将NL-to-STL翻译重新表述为一个闭环反馈过程。每个生成的公式都会被翻译回自然语言供用户检查,自然语言修正驱动修订,直到用户接受该规范。用户从不阅读或编写形式化语法。该框架基于反馈控制理论中常见的一种不对称性。前向路径,从模糊语言到形式逻辑,是困难且易出错的。反馈路径,从结构化的STL返回到语言,可以做到高度精确,而精确的反馈路径使得不精确的前向路径能够实现精确的闭环行为。在500个专家编写的需求和七个大型语言模型上进行的实验支持这一观点。反向翻译的解释在99.5%的情况下与专家判断一致。闭环精炼将强模型的开环准确率从约89%提升至98.0--99.2%,并使较弱模型获得超过30个百分点的提升(例如,17.6%→48.0%)。消融实验表明,这些提升来自反馈的语义内容,而非重复尝试。专家审核和一项280次会话的用户研究进一步证实了该闭环的可靠性。我们还确定了一个能力阈值,超过该阈值后反馈不再有帮助。
cs.AI / 196 / 2609.33289

Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets

学习销售:多产品市场中策略性大语言模型智能体的强化学习
Liu, Shuze Daniel, Chen, Claire, Wang, Jiuqi, Joachims, Thorsten
Abstract
Autonomous large language model (LLM) agents operating in multi-product markets must make sequential decisions under information asymmetry and resource constraints. We develop a machine learning approach for training such agents to act effectively as sellers in a multi-item bargaining environment, where a seller concurrently negotiates a catalog of substitutable assets across a pool of independent buyers. Buyers hold private, heterogeneous valuations across products, and each can purchase at most one item. Facing limits on total communication turns, the seller must dynamically match buyers with the most profitable products considering their private valuations, while strategically allocating its limited interaction budget toward combinations of greater potential value. We formalize this problem as a Partially Observable Markov Decision Process using a structured, four-part message protocol that maps natural language into a parsable and regulated decision space. Using this formalization, we design a post-training method using Reinforcement Learning from Verifiable Rewards (RLVR). To evaluate this framework, we construct a multidimensional metric suite that quantifies constraint adherence, seller surplus extraction, and allocation quality. Our trained seller agent learns to match limited inventory to buyers more effectively, matching or outperforming trillion-parameter frontier models in both seller surplus extraction and buyer-product allocation quality. Finally, these learned strategies generalize robustly to unseen market structures, correlated valuation distributions, and price ranges not encountered during training.
Chinese Translation
在多产品市场中运行的自主大语言模型(LLM)智能体必须在信息不对称和资源约束下做出序贯决策。我们开发了一种机器学习方法,用于训练此类智能体在多物品讨价还价环境中有效充当卖方;在该环境中,一个卖方同时与一组独立买家就一揽子可替代资产进行谈判。买家对不同产品持有私有的、异质估值,且每个买家最多购买一件物品。面对总沟通轮次的限制,卖方必须考虑买家的私人估值,动态地将买家与最有利可图的产品匹配,同时策略性地将有限交互预算分配给潜在价值更高的组合。我们将该问题形式化为部分可观测马尔可夫决策过程,并使用结构化的四部分消息协议,将自然语言映射为可解析且受约束的决策空间。基于这一形式化,我们设计了一种使用可验证奖励强化学习(RLVR)的后训练方法。为了评估该框架,我们构建了一个多维指标套件,用于量化约束遵循、卖方剩余提取和分配质量。我们训练出的卖方智能体学会更有效地将有限库存与买家匹配,在卖方剩余提取和买家-产品分配质量上均达到或超过万亿参数前沿模型。最后,这些学习到的策略能够稳健地泛化到未见过的市场结构、相关估值分布以及训练期间未遇到的价格范围。
cs.AI / 197 / 2609.33295

TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

TraceDance:一个从真实世界智能体部署轨迹构建智能体行为基准的自动化系统
Min, Dehai, Zhang, Daoan, Zeng, Yiming, Zhang, Huayi, Chen, Ziyi, Zhang, Yan, Bai, Qinbo, Chao, Mengyuan, Ning, Jing, Hua, Qiyue, Chen, Huiyi, Zhang, Hanrong, Zou, Henry Peng, Yang, Jie, Xu, Wei, Yu, Philip S.
Abstract
An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. The benchmarks use decision-point continuation to evaluate an LLM's next turn at a recorded decision point with a behavior-specific rubric, without a reference answer or environment replay. Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests. Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators. Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points. Analysis across behavior-specific benchmarks further reveals weaknesses in how current LLMs behave as agents. By turning deployment problems into targeted benchmarks, TraceDance could serve as a key component of the recursive self-improvement (RSI) loop.
Chinese Translation
一个智能体可以完成一项任务,同时在执行过程中表现出不良行为。开发者需要对部署中遇到的特定行为进行测试,而不仅仅是固定的基准套件。我们提出了 TraceDance,一个智能体系统,它从部署轨迹中为用户指定的不良行为构建针对性的基准。为了高效构建,Anchor-and-Confirm 将可编程检索与由 Flash 大型语言模型(LLM)进行的候选级别确认相结合,而 Anchor Synthesis Loop 则为自定义行为生成和修订规范。这些基准使用决策点延续,在记录的决策点处用行为特定的评分标准评估 LLM 的下一轮响应,而无需参考答案或环境重放。在编码和通用工具使用方面的实验利用了 252,557 个会话,并产生了 107 个基准,包含 4,125 个实例,满足了 95.3% 的构建目标请求。两位人类标注者在 84% 的抽样实例中确认了所请求的行为,并且自动评分器与人类通过/失败判断的一致性可与标注者之间的一致性相媲美。九个前沿 LLM 的平均通过率仅为 26.7%,表明它们在评估的决策点上仍难以做出适当响应。对行为特定基准的分析进一步揭示了当前 LLM 作为智能体时的弱点。通过将部署问题转化为针对性基准,TraceDance 可以作为递归自我改进(RSI)循环的关键组成部分。
cs.AI / 198 / 2609.33297

The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization

你所见的错误并非你犯的错误:面向推理错误定位的进展感知推理根源
Wang, Yiguo, Yang, Ziyuan, Zou, Yi, Lin, Dan, Li, Rongsheng, Zhang, Yi
Abstract
Verifying multi-step LLM reasoning requires more than determining whether a trace is correct: a useful verifier should identify where the reasoning first goes wrong. However, existing holistic methods provide little positional evidence, while forward sequential verification often treats the first rejected step as the error source. Under error propagation, this assumption can fail, since an earlier mistake may remain locally plausible and become observable only through its downstream consequences. We therefore rethink reasoning verification as a progression-aware error-source localization problem: rather than asking only where a reasoning trace first appears inconsistent, we ask which earlier step best explains how that inconsistency emerges along the trajectory. Based on this view, we propose Progression-aware Reasoning Origin (PRO), a training-free framework for first-error localization. PRO jointly models incoming support from the preceding context and outgoing compatibility with subsequent reasoning, selectively refines regions where these signals disagree, and finally performs detector-conditioned source attribution with intervention-based evidence to distinguish the true error origin from its propagated manifestations. We further formalize the gap between forward rejection and structural exposure, showing why incoming-side evidence alone is insufficient for reliable localization under error propagation. Experiments across open-form, medical, and structured reasoning tasks demonstrate consistent improvements over strong verification baselines, supporting progression-aware source attribution as a more faithful formulation of reasoning verification.
Chinese Translation
验证多步LLM推理不仅仅需要确定一个推理轨迹是否正确:一个有用的验证器应该识别出推理首次出错的位置。然而,现有的整体性方法几乎不提供位置证据,而前向顺序验证通常将第一个被拒绝的步骤视为错误源。在错误传播下,这一假设可能会失败,因为较早的错误可能在局部仍然看似合理,并且只有通过其下游后果才能变得可观察。因此,我们重新思考推理验证,将其视为一个进展感知的错误源定位问题:我们不只是问推理轨迹在哪里首次出现不一致,而是问哪个较早的步骤最能解释这种不一致是如何沿轨迹出现的。基于这一观点,我们提出了进展感知推理根源(PRO),一个用于首个错误定位的免训练框架。PRO联合建模来自前文的传入支持和与后续推理的传出兼容性,选择性细化这些信号不一致的区域,并最终执行基于检测器条件的源归因与基于干预的证据,以区分真实错误根源与其传播表现。我们进一步形式化前向拒绝与结构暴露之间的差距,说明为什么在错误传播下仅靠传入侧证据不足以进行可靠定位。在开放式、医学和结构化推理任务上的实验表明,相较于强大的验证基线,该方法取得了持续改进,支持将进展感知源归因作为推理验证的一种更忠实的表述。
cs.AI / 199 / 2609.33319

PhysAlign: A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics Reasoning

PhysAlign:多模态物理推理中基于证据的角色对齐基准
Liang, Kecheng, Liu, Haoyang, Chen, Zexin, Liu, Zirong, Chen, Weixing, Wang, Qiufeng, Liu, Yang, Lin, Liang
Abstract
A key challenge in physics diagram understanding is correctly associating visual information with the physical entities, relations, and conditions it describes. Even when a value, symbol, or other local element is accurately recognized, assigning it to the wrong entity or scope can distort the underlying physical premise and lead to incorrect reasoning. To systematically study this challenge, we introduce \textbf{PhysAlign}, a benchmark designed to assess whether multimodal models correctly associate information recognized from physics diagrams with its intended physical role. By disentangling visual recognition from physical-role assignment through localized probes and controlled variants, PhysAlign isolates correspondence errors from recognition failures. It contains 3,341 human-validated probes spanning 986 physics problems, enabling systematic evaluation of visual recognition and physical-role correspondence at scale. We further introduce five complementary evaluation metrics, including CAcc, GAcc, and JAcc, which provide a comprehensive assessment of models' ability to recognize diagram content, establish correct physical correspondences, and solve the underlying physics problem. Across our evaluated multimodal models, PhysAlign reveals a consistent gap between local visual recognition and physical-role grounding. Even when the queried content is correctly recognized, the conditional correspondence error rate remains 13.8\% for GPT-6-Astra and rises to about 50.6\% for InternVL3.5-8B. These findings indicate that strong perception alone does not ensure reliable physical interpretation, exposing a distinct grounding bottleneck that is largely hidden by answer-level accuracy and highlighting the need for future models to better align recognized visual evidence with its physical meaning.
Chinese Translation
物理图表理解中的一个关键挑战是正确地将视觉信息与其所描述的物理实体、关系和条件关联起来。即使一个数值、符号或其他局部元素被准确识别,将其分配给错误的实体或范围也可能扭曲潜在的物理前提,并导致错误的推理。为了系统地研究这一挑战,我们引入了PhysAlign,一个旨在评估多模态模型是否能正确地将从物理图表中识别出的信息与其预期的物理角色关联起来的基准。通过局部探测和受控变体将视觉识别与物理角色分配分离,PhysAlign将对应误差与识别失败分离开来。它包含3,341个经人工验证的探测,涵盖986个物理问题,能够大规模系统地评估视觉识别和物理角色对应。我们进一步引入了五个互补的评估指标,包括CAcc、GAcc和JAcc,它们全面评估模型识别图表内容、建立正确的物理对应关系以及解决潜在物理问题的能力。在我们评估的多模态模型中,PhysAlign揭示出局部视觉识别与物理角色关联之间的一致差距。即使查询的内容被正确识别,GPT-6-Astra的条件对应错误率仍为13.8%,而InternVL3.5-8B则上升到约50.6%。这些发现表明,仅靠强大的感知并不能确保可靠的物理解释,暴露出一个明显的关联瓶颈,该瓶颈在很大程度上被答案级准确率所掩盖,并凸显出未来模型需要更好地将识别出的视觉证据与其物理意义对齐。
cs.AI / 200 / 2609.33323

Agentic Multi-Turn Reasoning: A Fairness Approach

智能体多轮推理:一种公平性方法
Truong, Thanh-Dat, Pandey, Sankalp, Churchill, Hugh, Cothren, Jackson, Savvides, Marios, Luu, Khoa
Abstract
Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where supervision is available only at the final outcome, and (2) imbalanced data distributions, where dominant data patterns bias optimization and weaken adaptation to rare but informative reasoning behaviors. In this paper, we propose Fair Multi-Level Preference Optimization (Fair-MPO or $\Phi$-MPO), a new preference optimization framework for agentic learning. We first show that Multi-Level Preference Optimization provides a principled and more computationally efficient framework for long-horizon reasoning. Then, we introduce a Fair Multi-Level Objective that addresses imbalance in agentic learning. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both long-horizon reasoning and data imbalance. Our experiments on agentic reasoning benchmarks demonstrate that our approach achieves State-of-the-Art (SOTA) performance.
Chinese Translation
近年来,大型语言模型(LLMs)的进展使得智能体系统能够通过多轮规划、工具使用、验证和记忆更新来解决复杂任务。然而,学习智能体系统仍然困难,原因在于两个基本挑战,即:(1)长时程信用分配,其中监督信号仅在最终结果处可得;(2)数据分布不平衡,其中占主导的数据模式会使优化产生偏差,并削弱对稀有但信息丰富的推理行为的适应。本文提出公平多级偏好优化(Fair Multi-Level Preference Optimization,Fair-MPO 或 $\Phi$-MPO),一种用于智能体学习的新偏好优化框架。我们首先表明,多级偏好优化(Multi-Level Preference Optimization)为长时程推理提供了一个有原则且计算效率更高的框架。然后,我们引入公平多级目标(Fair Multi-Level Objective)以解决智能体学习中的不平衡问题。我们提供了全面的理论分析,证明我们的方法同时解决了长时程推理和数据不平衡问题。我们在智能体推理基准上的实验表明,我们的方法达到了当前最优(SOTA)性能。
cs.AI / 201 / 2609.33326

ANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information Spaces

ANTMAN:大规模信息空间中多智能体导航的自适应需求跟踪
Wang, Jerry, Jin, Haibo, Yuan, Xiaopeng, Kuang, Peng, Wang, Haohan
Abstract
Information-seeking agents increasingly operate over information spaces that are too large to process exhaustively. Yet many multi-agent systems organize computation around static partitions of the available space, causing coordination to grow with how information is segmented rather than with what the query still requires. We introduce ANTMAN, an adaptive coordination framework that treats evolving unresolved information needs as the unit of runtime coordination. ANTMAN maintains a revisable Need Graph that tracks unresolved requirements, accumulated evidence, prior attempts, and search progress, and uses this state to control worker selection, routing, and task-local recovery as new evidence is discovered. By separating the coordination policy from substrate-specific search interfaces, the same need-conditioned mechanism can operate across different information spaces. Experiments across multi-document question answering, controlled long-context scaling, and realistic structured navigation show that ANTMAN remains effective across settings, including when execution is delegated to substantially smaller worker models. Under a 16x increase in searchable context, ANTMAN increases active coordination by only 1.23x, compared with more than 15x for partition-driven baselines, while preserving strong answer quality.
Chinese Translation
信息搜索智能体越来越多地在无法穷尽处理的大规模信息空间中运行。然而,许多多智能体系统围绕可用空间的静态划分来组织计算,导致协调开销的增长取决于信息的分割方式,而非查询的剩余需求。我们提出了ANTMAN,一个自适应协调框架,它将不断演化的未解决信息需求视为运行时协调的单元。ANTMAN维护一个可修订的需求图(Need Graph),用于跟踪未解决的需求、积累的证据、先前的尝试和搜索进度,并利用该状态在新证据被发现时控制工作器选择、路由和任务局部恢复。通过将协调策略与特定于底层的搜索接口分离,相同的需求条件机制可以跨不同的信息空间运行。在多文档问答、受控的长上下文扩展以及现实结构化导航上的实验表明,ANTMAN在不同设置下仍然有效,包括当执行被委托给小得多的工作模型时。在可搜索上下文增加16倍的情况下,ANTMAN的主动协调仅增加1.23倍,而分区驱动的基线则增加超过15倍,同时保持了强大的答案质量。
cs.AI / 202 / 2609.33339

Naturalness-guided Manifold Flow Matching for Sign Language Production

自然性引导的流形流匹配手语生成
He, Jiayi, Tang, Shengeng, You, Sisi, Hao, Yanbin, Cheng, Lechao, Hong, Richang
Abstract
Sign Language Production (SLP) aims to generate sign motions from text. Conditional Flow Matching methods have achieved strong performance in SLP by constructing conditional paths that transform a source distribution into a target distribution. However, existing methods construct these paths via linear interpolation, whereas the rotational geometry of human joints confines valid joint rotations to a manifold embedded in Euclidean space. Consequently, linear interpolation between two sign motions leaves this manifold and ignores the motion distribution on it. In this paper, we revisit SLP from the perspective of manifold transport and propose a Naturalness-guided Manifold Flow Matching framework, termed \textbf{SignNMFlow}, which constructs conditional paths directly on the motion manifold by jointly considering geometric efficiency and the motion distribution. Specifically, we exploit the intrinsic geometry of the manifold and introduce a motion naturalness measure to characterize the motion distribution. By minimizing the kinetic energy under this measure, we learn a naturalness-guided interpolation that couples a closed-form geodesic, which provides geometrically efficient transport, with a learnable deviation that incorporates the motion distribution, thereby significantly improving the fidelity of generated sign motions. Extensive qualitative and quantitative evaluations demonstrate the effectiveness of this work.
Chinese Translation
手语生成(SLP)旨在从文本生成手语动作。条件流匹配方法通过构建将源分布转换为目标分布的条件路径,在手语生成中取得了强大性能。然而,现有方法通过线性插值构建这些路径,而人类关节的旋转几何将有效的关节旋转限制在嵌入欧几里得空间的流形上。因此,两个手语动作之间的线性插值会离开该流形,并忽略其上的动作分布。在本文中,我们从流形传输的角度重新审视SLP,并提出一种自然性引导的流形流匹配框架,称为SignNMFlow,它通过同时考虑几何效率和动作分布,直接在动作流形上构建条件路径。具体而言,我们利用流形的内在几何,并引入动作自然性度量来表征动作分布。通过在该度量下最小化动能,我们学习一种自然性引导的插值,它将提供几何高效传输的闭式测地线与融入动作分布的可学习偏差耦合起来,从而显著提高生成手语动作的保真度。广泛的定性和定量评估证明了这项工作的有效性。
cs.AI / 203 / 2609.33351

QuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAG

QuPID: 面向医学RAG的量子参数高效输入依赖检索自适应
Ahn, Hyojun, Roh, Emily Jimin, Park, Soohyun, Saad, Walid, Lee, Hyung-Chul, Kim, Joongheon
Abstract
Fidelity-based quantum retrieval ranks candidates by the fidelity between query and archive states. Applying a shared input-independent unitary after fixed state encoding leaves that fidelity unchanged, so training the circuit cannot alter the ranking. Quantum parameter-efficient input-dependent retrieval adaptation (QuPID) repairs this by making the circuit input-dependent through data re-uploading and by comparing measurement readouts, vectors of local Pauli expectations, rather than states. The result is a small readout for adapting frozen image features to a local archive with limited data: training simulates the circuit classically, and inference runs on a GPU with fixed learned parameters. We characterize the class as a structured factorization of input-modulated quadratic feature maps, bound the frequency support of its re-uploading channel, and give a parameter-count generalization bound that motivates its small budget. Under a shared frozen backbone and a label-free protocol, QuPID's 60 parameters give higher precision-at-5 (P@5) on ChestX-ray14 and MURA than frozen medical encoders, and than adapters and low-rank adaptation (LoRA) with up to 5.25 million trainable parameters. On ChestX-ray14, the P@5 gain over the frozen encoder is +0.116, the lead over retuned adapters is widest at 512 adaptation examples (+0.040), and the full-budget margin over an equally compact classical rotation-plane head is +0.023 with a 95% interval excluding zero. Medical imaging is the primary testbed; the pattern recurs on two non-medical benchmarks, in report generation, and under simulated gate noise and finite-shot readout.
Chinese Translation
基于保真度的量子检索根据查询态与档案态之间的保真度对候选进行排序。在固定状态编码后应用一个共享的、与输入无关的酉变换会使该保真度保持不变,因此训练电路无法改变排序。量子参数高效输入依赖检索自适应(QuPID)通过数据重上传使电路依赖于输入,并通过比较测量读数(即局部泡利期望的向量)而非状态,来修复这一问题。结果是一个小型读出,用于将冻结的图像特征适应到数据有限的本地档案:训练在经典计算机上模拟电路,推理在GPU上运行,使用固定的学习参数。我们将该类描述为输入调制二次特征映射的结构化分解,限制其重上传通道的频率支持,并给出一个参数数量泛化界,以证明其小预算的合理性。在共享的冻结骨干网络和无标签协议下,QuPID的60个参数在ChestX-ray14和MURA上给出了比冻结医学编码器、以及比具有多达525万个可训练参数的适配器和低秩适应(LoRA)更高的P@5(precision-at-5)。在ChestX-ray14上,相对于冻结编码器的P@5增益为+0.116,在512个适应样本时对重新调优的适配器的领先优势最大(+0.040),而相对于同等紧凑的经典旋转平面头的全预算边际为+0.023,其95%置信区间不包含零。医学影像是主要测试平台;该模式在两个非医学基准、报告生成以及模拟门噪声和有限次读出下重复出现。
cs.AI / 204 / 2609.33355

Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models

揭示状态:状态自适应何时对掩码扩散语言模型重要?
Kong, Injin, Choi, Sunghwan, Jo, Yohan
Abstract
Masked diffusion language models (MDMs) admit flexible generation orders, making the unmasking strategy an inference decision. Existing methods vary in how they prioritize positions, control parallelism, restrict selection regions, revise predictions, or plan future denoising, yet it remains unclear when these choices should change during generation. We study this question through strategy reversals, where an alternative action becomes preferable to a fixed choice. We organize MDM inference into five axes--score, cardinality, region, commitment, and planning--and define adaptation opportunity as the one-step utility advantage of the best candidate action over a validation-selected fixed action. This view shows that adaptation value depends on both the frequency and magnitude of such reversals. Across three MDMs and ten tasks, adaptation opportunities are highly heterogeneous, with some regimes exhibiting concentrated and predictable one-step gains. This motivates selective adaptation: lightweight detectors calibrated on validation prompts identify high-opportunity states, capturing, for example, 56.9 percent of the candidate-set oracle opportunity by adapting only the top 10 percent of states on LLaDA-8B constrained JSON filling. Our transition-level results suggest that state adaptation is most useful when applied selectively rather than uniformly.
Chinese Translation
掩码扩散语言模型(MDMs)允许灵活的生成顺序,使得解掩码策略成为一种推理决策。现有方法在如何确定位置优先级、控制并行性、限制选择区域、修正预测或规划未来去噪方面各不相同,但尚不清楚在生成过程中这些选择应在何时改变。我们通过策略反转来研究这个问题,即替代动作变得优于固定选择。我们将 MDM 推理组织为五个轴——评分(score)、基数(cardinality)、区域(region)、承诺(commitment)和规划(planning)——并将自适应机会定义为最佳候选动作相对于验证集选择的固定动作的一步效用优势。这一视角表明,自适应的价值取决于此类反转的频率和幅度。在三个 MDM 和十个任务中,自适应机会高度异质,某些机制表现出集中且可预测的一步收益。这催生了选择性自适应:在验证提示上校准的轻量级检测器识别高机会状态,例如,在 LLaDA-8B 约束 JSON 填充任务上仅对前 10% 的状态进行自适应,即可捕获候选集 oracle 机会的 56.9%。我们的转移级结果表明,状态自适应在选择性应用而非统一应用时最为有用。
cs.AI / 205 / 2609.33356

Long-Horizon Analog Design Bench: Benchmarking Agents on Hours-Long Analog and Mixed-Signal Circuit Design Tasks

长时程模拟设计基准:在长达数小时的模拟与混合信号电路设计任务上对智能体进行基准测试
Analog Design Bench Team
Abstract
Coding agents now sustain hours-long, tool-driven loops, yet their ability to carry long-horizon analog and mixed-signal circuits to electrical specification remains unmeasured. We introduce Analog Design Bench, a long-horizon agentic benchmark of 50 transistor-level design tasks contributed by 17 chip designers. Agents work with an open-source simulator, while an isolated verifier evaluates the submitted circuit using specification-based electrical tests. We evaluate 15 agent configurations across 2,250 two-hour attempts and observe full-specification pass rates from 8.0% to 78.0%. Coding-benchmark performance correlates with analog results but leaves much of the performance spread unexplained. Our failure analysis shows that most unsuccessful submissions have no recorded legality rejection but fail electrical acceptance, identifying electrical closure as the dominant endpoint challenge. We test time, reasoning effort, agent harness, and supplied design knowledge as interventions. Longer budgets and higher reasoning effort improve performance, while general skill documents provide little benefit and sometimes reduce performance. Supplying a task-matched reference topology, an idealized form of circuit-IP retrieval, raises DeepSeek V4 Pro by 18.7 percentage points and mainly accelerates GPT-5.6 Sol.
Chinese Translation
编码智能体现在能够维持长达数小时、工具驱动的循环,但它们将长时程模拟与混合信号电路设计推进到满足电气规格的能力仍未被衡量。我们提出了模拟设计基准(Analog Design Bench),这是一个长时程智能体基准,包含由17位芯片设计师贡献的50个晶体管级设计任务。智能体使用开源模拟器工作,而一个隔离的验证器使用基于规格的电气测试来评估提交的电路。我们在2,250次两小时尝试中评估了15种智能体配置,观察到全规格通过率从8.0%到78.0%不等。编码基准性能与模拟结果相关,但大部分性能差异仍无法解释。我们的失败分析表明,大多数不成功的提交没有记录到合法性拒绝,但未能通过电气验收,从而确定电气闭合是主要的终点挑战。我们测试了时间、推理努力、智能体框架和提供的设计知识作为干预措施。更长的预算和更高的推理努力提高了性能,而通用技能文档几乎没有益处,有时反而降低性能。提供任务匹配的参考拓扑(一种理想化的电路IP检索形式)使 DeepSeek V4 Pro 提高了18.7个百分点,并主要加速了 GPT-5.6 Sol。
cs.AI / 206 / 2609.33357

DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?

DISCERN:AI 智能体能否像科学家一样工作并引导发现?
Huang, Nan, Tapia-Pacheco, Mario, Zhou, Kun, Huang, Yiming, Díaz, Kevin José Barrientos, Amariuta, Tiffany, Shang, Jingbo
Abstract
Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical task completion or hypothesis generation separately rather than testing whether reliable evidence supports valid and novel claims. We introduce DISCERN (Data Integrity and Scientific Capability: Evidence, Reasoning, and Novelty), a controlled benchmark on real, publicly available datasets that evaluates three key levels of an automated research workflow. The first two levels test data integrity and analysis verification under confounds and tool traps, while the third tests hypothesis generation and revision under adversarial review, including counterfactual cases in which evidence consistent with real data and documented scientific phenomena conflicts with established expectations, motivating alternative explanations and testable hypotheses. Across 203 tasks, eight life-science tracks, and eight models, DISCERN shows that strong aggregate performance can mask level-specific weaknesses. Agents earn perfect scores in only 60.8% of Level 1, 34.2% of Level 2, and 0.6% of Level 3 evaluations, with penalties attributed to rejection of sound data, failure to carry recognized limitations into conclusions, and wide variation in hypothesis production. Cross-track rankings by token and code use are substantially more stable than rankings by evidence judgment, suggesting greater consistency in computational effort than in evidence-based reasoning. These profiles identify opportunities for supervised scientific assistance, but current agents do not yet demonstrate reliable autonomous analysis or discovery. Code and data: https://huggingface.co/datasets/discern-bench-anon/discern-benchmark
Chinese Translation
可靠的自动化研究要求智能体审查数据、验证分析,并生成以可信证据为基础的假设,这可能减少常规科学工作量,同时让科学家专注于解释和发现。现有基准通常只分别评估分析任务完成情况或假设生成,而非测试可靠证据是否支持有效且新颖的主张。我们提出 DISCERN(数据完整性与科学能力:证据、推理与新颖性),一个基于真实、公开可用数据集的受控基准,用于评估自动化研究工作流程的三个关键层次。前两个层次测试在混杂因素和工具陷阱下的数据完整性和分析验证,而第三个层次测试在对抗性评审下的假设生成和修正,包括反事实案例,其中与真实数据和已记录科学现象一致的证据与既定预期冲突,从而激发替代解释和可检验的假设。在 203 个任务、八个生命科学赛道和八个模型中,DISCERN 表明,强劲的总体表现可能掩盖特定层次的弱点。智能体仅在 60.8% 的第一层、34.2% 的第二层和 0.6% 的第三层评估中获得满分,扣分归因于拒绝合理数据、未能将已识别的局限性带入结论,以及假设产出的巨大差异。按 token 和代码使用进行的跨赛道排名比按证据判断进行的排名稳定得多,表明计算努力的一致性高于基于证据的推理。这些特征图谱指出了有监督科学辅助的机会,但当前智能体尚未展示出可靠的自主分析或发现能力。代码和数据:https://huggingface.co/datasets/discern-bench-anon/discern-benchmark
cs.AI / 207 / 2609.33368

DrafTS: Time-Aware Decomposition with Residual Correction for Time Series Modeling

DrafTS:面向时间序列建模的时间感知分解与残差校正
Liu, Yiqiu, Zhong, Siru, Wang, Zhiguang, Wen, Qingsong, Liang, Yuxuan
Abstract
Real-world time series contain evolving underlying dynamics with irregular variations that lack stable temporal patterns and are often referred to as noise. Existing methods address this mixture by filtering frequencies or suppressing noisy observations. They either miss temporal evolution or risk suppressing useful dynamics. We propose DrafTS, a model-agnostic framework that aims to reduce noise while preserving evolving dynamics through time-aware Decomposition with ResiduAl correction For Time Series. DrafTS uses features derived from instantaneous amplitude and frequency to guide decomposition into a primary component intended to capture underlying dynamics. A task-specific backbone models the primary component, while a lightweight correction module uses residual information to correct the backbone output. Across four time series modeling tasks, DrafTS improves six diverse backbones, demonstrating its effectiveness. Code is at https://github.com/Autumn61q/DrafTS
Chinese Translation
现实世界的时间序列包含不断演化的潜在动态,并伴随缺乏稳定时间模式的不规则变化,这些变化通常被称为噪声。现有方法通过滤除频率或抑制含噪观测来应对这种混合。它们要么忽略时间演化,要么可能抑制有用的动态。我们提出 DrafTS,一个模型无关框架,旨在通过面向时间序列的时间感知分解与残差校正(time-aware Decomposition with ResiduAl correction For Time Series)在降低噪声的同时保留演化动态。DrafTS 使用由瞬时振幅和频率导出的特征来指导分解,得到旨在捕捉潜在动态的主要分量。任务特定的骨干模型对主要分量建模,而轻量级校正模块利用残差信息校正骨干输出。在四个时间序列建模任务中,DrafTS 提升了六种不同的骨干模型,证明了其有效性。代码见 https://github.com/Autumn61q/DrafTS
cs.AI / 208 / 2609.33394

Cross-modal Translation via Conditional Latent Denoising for Video Deepfake Detection

基于条件潜在去噪的跨模态转换用于视频深度伪造检测
Li, Xinzhe, Tu, Youzhi, Lee, Kong Aik
Abstract
The growing threat of video deepfakes necessitates multimodal detection. Beyond serving as independent indicators of authenticity, audio and visual signals have intrinsic dependencies that also provide an essential criterion for detection. Previous methods often overlook the cross-modal correspondences, hindering information transfer between domains and leaving crucial detection cues unexplored. To address this challenge, we propose a framework called Cross-modal Translation via Conditional Latent Denoising (CTCLD) for video deepfake detection. It connects the distinct distributions of heterogeneous modalities in latent spaces, enabling smooth cross-domain information transfer to improve detection performance. We first establish a Bayesian foundation by decomposing the audio-visual joint distribution. Subsequently, CTCLD translates both modalities via bidirectional latent denoising conditioned on each other, effectively capturing subtle inconsistencies in the manipulated signals. Experimental results demonstrate that the proposed CTCLD enables comprehensive domain alignment, resulting in a robust video deepfake detection approach with competitive performance.
Chinese Translation
视频深度伪造日益增长的威胁使得多模态检测成为必要。除了作为真实性的独立指标外,音频和视觉信号还具有内在依赖关系,这也为检测提供了重要标准。先前方法常常忽略跨模态对应关系,阻碍了领域间的信息传递,并使关键检测线索未被充分探索。为应对这一挑战,我们提出了一种称为基于条件潜在去噪的跨模态转换(CTCLD)的框架,用于视频深度伪造检测。它在潜在空间中连接异构模态的不同分布,实现平滑的跨域信息传递,从而提升检测性能。我们首先通过分解音视频联合分布来建立贝叶斯基础。随后,CTCLD通过以彼此为条件的双向潜在去噪来转换两种模态,有效捕捉被操纵信号中的细微不一致。实验结果表明,所提出的CTCLD能够实现全面的领域对齐,从而形成一种具有竞争性能的稳健视频深度伪造检测方法。
cs.AI / 209 / 2609.33397

CoViST: Visual Token Compression via Composable States

CoViST: 通过可组合状态的视觉Token压缩
Zhang, Qi, Meng, Xiandong, Wang, Ronggang, Ma, Siwei
Abstract
Visual token compression lowers the inference cost of vision--language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. Therefore, the compressed representation does not explicitly encode how much visual information each representative carries or where it lies in the original image. This limitation arises even after a single reduction and becomes more pronounced when compression is repeated across decoder layers. To address this issue, we propose CoViST, a training-free framework that represents a compressed image as a composable visual state. Specifically, the state combines representative features with original positions, effective contribution weights, and reusable selection metadata. CoViST constructs this state through coverage-guided selection and conservation-based contribution composition, and explicitly incorporates its contribution and positional information into decoder attention. Each component of the state retains its interpretation under successive reductions, enabling the same formulation to support both fixed compression before prefill and progressive compression within the decoder. Experimental results on seven LLaVA-1.5-7B benchmarks show that CoViST-Fixed retains 99.9\%, 99.5\%, and 98.1\% of uncompressed performance at 192, 128, and 64 tokens, respectively, and CoViST-Pro retains 99.8\%, 99.9\%, and 99.1\% at the corresponding layer-average budgets, outperforming state-of-the-art methods under their respective budget settings. Code will be released publicly.
Chinese Translation
视觉Token压缩通过用更少的Token表示图像,降低了视觉-语言模型的推理成本。然而,大多数现有方法将视觉Token压缩为一个缩减的集合,使得每个Token所表示的视觉证据量及其原始空间上下文成为隐式的。因此,压缩表示没有显式编码每个代表所携带的视觉信息量,或其位于原始图像中的位置。这种局限性即使在单次缩减后也会出现,并且当压缩在解码器层中重复进行时变得更加明显。为了解决这个问题,我们提出了CoViST,一个无需训练的框架,它将压缩后的图像表示为可组合的视觉状态。具体来说,该状态将代表性特征与原始位置、有效贡献权重以及可重用的选择元数据相结合。CoViST通过覆盖引导的选择和基于守恒的贡献组合来构建该状态,并显式地将其贡献和位置信息融入解码器注意力中。状态的每个组件在连续缩减下保持其解释,使得相同的公式能够支持预填充前的固定压缩和解码器内的渐进压缩。在七个LLaVA-1.5-7B基准上的实验结果表明,CoViST-Fixed在192、128和64个Token时分别保持了未压缩性能的99.9%、99.5%和98.1%,而CoViST-Pro在相应的层平均预算下保持了99.8%、99.9%和99.1%,在各自的预算设置下优于最先进的方法。代码将公开发布。
cs.AI / 210 / 2609.33398

COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement

COEVO:面向递归自我改进的共同演化上下文与参数
Chen, Siwei, Bao, Xinping, Cai, Xinyu, Cao, Yuan, Jiang, Wan, Chen, Shaohong
Abstract
Recursive self-improvement (RSI) seeks to move large language models beyond static training pipelines toward systems that can participate in improving their own future behavior. Existing approaches largely follow two directions: updating model parameters through online learning, or improving the external context through search, reflection, and prompt optimization. Although both mechanisms can support continued improvement, they are typically studied independently. This separation overlooks an important interaction: the context shapes the experience from which a model learns, while an evolving model may interpret and utilize the same context differently over time. We therefore formulate RSI as a problem of parameter--context co-evolution, where model parameters and the learning context adapt within a shared feedback loop. We introduce COEVO, a framework that updates model parameters from on-policy experience while adapting contextual guidance according to the state of the evolving policy. Policy entropy and prompt-conditioned attention are used as complementary signals to guide this adaptation. Experiments show that COEVO consistently improves task performance over fixed-context reinforcement learning and produces policies that are more robust to changes in system prompts. More broadly, our results suggest that external context should be viewed not merely as a fixed interface to a large language model, but as an adaptive component of recursive self-improvement.
Chinese Translation
递归自我改进(RSI)旨在让大型语言模型超越静态训练流程,走向能够参与改进自身未来行为的系统。现有方法主要沿两个方向:通过在线学习更新模型参数,或通过搜索、反思和提示优化来改进外部上下文。尽管这两种机制都能支持持续改进,但它们通常被独立研究。这种分离忽视了一个重要交互:上下文塑造了模型学习所依据的经验,而一个不断演化的模型可能随时间以不同方式解释和利用相同的上下文。因此,我们将RSI形式化为一个参数-上下文共同演化问题,其中模型参数和学习上下文在共享反馈循环中相互适应。我们提出了COEVO,一个框架,它从同策略经验中更新模型参数,同时根据演化策略的状态来调整上下文指导。策略熵和提示条件注意力被用作互补信号来指导这种适应。实验表明,COEVO在任务性能上持续优于固定上下文强化学习,并产生对系统提示变化更鲁棒的策略。更广泛地说,我们的结果表明,外部上下文不应仅仅被视为大型语言模型的固定接口,而应被视为递归自我改进的一个自适应组件。
cs.AI / 211 / 2609.33411

MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses

MetaBench-Harness:解锁基准测试框架的端到端优化
Chen, Xuanjun, Chen, Hua-Hsuan, Lu, Wei-Chung, Ma, Yinghao, Jang, Jyh-Shing Roger, Lee, Hung-yi
Abstract
Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid, hard-coded generation rules. Moving beyond the evolution of isolated tasks, we propose to optimize the benchmark generation workflow itself end to end with MetaBench-Harness, a dual-loop search framework. Specifically, the inner loop utilizes a benchmark harness to generate a new benchmark in each round, while the outer meta-harness orchestration layer iteratively refines and searches over harness implementations based on historical evolution trajectories. By applying MetaBench-Harness to the competitive programming CodeContests and Olympiad mathematics AIME-2024 datasets, we demonstrate that the evolved benchmarks are challenging and discriminative for frontier models. Trajectory and quality analyses verify that MetaBench-Harness enables multi-dimensional evolution, steadily improving evolution reasonableness, benchmark competency, and evaluator robustness across successive rounds. Furthermore, case studies reveal its effective utilization of diverse difficulty levers to reframe problems and elevate required capabilities. Ultimately, this work provides a solution to the pressing challenge of benchmark saturation.
Chinese Translation
大型语言模型(LLMs)的快速进展正以快于其设计速度使静态基准测试饱和。尽管现有的自动演化框架试图通过扰动单个任务来生成更难的问题,但它们仍受限于僵化的硬编码生成规则。超越孤立任务的演化,我们提出使用 MetaBench-Harness(一种双循环搜索框架)来端到端地优化基准测试生成工作流本身。具体而言,内循环利用基准测试框架在每一轮中生成一个新的基准测试,而外部的元框架编排层则基于历史演化轨迹,迭代地细化和搜索框架实现。通过将 MetaBench-Harness 应用于竞技编程 CodeContests 和奥林匹克数学 AIME-2024 数据集,我们证明演化后的基准测试对前沿模型具有挑战性和区分度。轨迹和质量分析验证了 MetaBench-Harness 能够实现多维演化,在连续轮次中稳步提升演化合理性、基准测试能力和评估器鲁棒性。此外,案例研究揭示了其有效利用多样化的难度杠杆来重构问题并提升所需能力。最终,这项工作为基准测试饱和这一紧迫挑战提供了解决方案。
cs.AI / 212 / 2609.33428

Temporal Graph Learning of Wearable Actigraphy and Sleep Traces for Modelling Adolescent Crystallized Intelligence

可穿戴活动记录与睡眠轨迹的时间图学习用于建模青少年晶体智力
Rahman, Md. Tanvir, Orka, Nabil Anan, Khan, Asaduzzaman, Moni, Mohammad Ali
Abstract
Wearable actigraphy offers a scalable, ecologically valid alternative to episodic clinical assessment. However, predicting continuous adolescent crystallized intelligence ($G_c$) from such traces remains challenging due to irregular device adherence and complex behavioral-environmental interactions. We address this using daily summary data derived from 21-day Fitbit records of 6,091 adolescents in the Adolescent Brain Cognitive Development Study (Release 5.1). We propose SATURN, a Sleep-Activity Temporal Unified Regression Network. It represents participants as 21-node temporal graphs encoding daily behaviors and temporal adjacency. To prevent imputation artifacts, invalid-day edges are dynamically pruned during forward passes. Node embeddings are refined via residual GATv2 layers, aggregated through masked attention pooling, and fused with sociodemographic covariates. Under family-controlled, age-sex-BMI-stratified cross-validation, SATURN achieves $R^2 = 0.2783 \pm 0.0127$, consistently improving upon flattened machine learning (Gradient Boosting, $R^2 = 0.2372$) and sequential deep learning (BiLSTM, $R^2 = 0.2688$) baselines. Explainability analyses identify light activity, metabolic equivalents, and sleep duration as dominant predictors, while Monte Carlo dropout and subgroup analyses confirm equitable performance across sociodemographic strata. Ultimately, SATURN establishes a rigorous computational framework for digital cognitive phenotyping, offering a scalable pathway to complement traditional assessments by highlighting macro-level behavioral anomalies.
Chinese Translation
可穿戴活动记录提供了一种可扩展、生态效度高的替代方案,可替代阶段性临床评估。然而,由于设备佩戴依从性不规则以及复杂的行为-环境交互,从这类轨迹中预测连续的青少年晶体智力(G_c)仍具挑战性。我们利用青少年大脑认知发展研究(Adolescent Brain Cognitive Development Study,ABCD Study;5.1版)中 6,091 名青少年 21 天 Fitbit 记录的每日汇总数据来解决这一问题。我们提出 SATURN,即睡眠-活动时间统一回归网络。它将参与者表示为 21 节点时间图,编码每日行为和时间邻接关系。为防止插补伪影,在前向传播过程中动态剪枝无效日期边。节点嵌入通过残差 GATv2 层进行细化,通过掩码注意力池化进行聚合,并与社会人口学协变量融合。在家庭控制、年龄-性别-BMI 分层交叉验证下,SATURN 达到 R² = 0.2783 ± 0.0127,持续优于扁平化机器学习(梯度提升,R² = 0.2372)和序列深度学习(BiLSTM,R² = 0.2688)基线。可解释性分析识别出轻度活动、代谢当量和睡眠时长为主要预测因子,而蒙特卡洛 dropout 和亚组分析证实其在不同社会人口学阶层中具有公平的性能。最终,SATURN 为数字认知表型分析建立了一个严谨的计算框架,通过突出宏观层面的行为异常,为补充传统评估提供了可扩展的路径。
cs.AI / 213 / 2609.33430

APEX: An Extensible Model for Agent-Assisted Production Scheduling

APEX:面向智能体辅助生产调度的可扩展模型
Grumbach, Felix J., Görlitz, Stefan
Abstract
Production scheduling requires realistic models that reflect operational constraints and efficient methods that balance competing goals. Putting these methods into use also requires data integration, model adaptation and specialist expertise. We present APEX, an extensible production scheduling framework built around a general model and hybrid multiobjective search. Agent assistance supports both scheduling and model refinement: agents prepare data and explore scenarios in natural language, while coding agents help implement and test new constraints and objectives. Shared construction and checking procedures connect these adaptations to the scheduling core. We benchmark eight APEX configurations against NSGA-II, SPEA2, MOEA/D and SMS-EMOA on 69 public job-shop, flexible job-shop and permutation flow-shop instances, assessing workload completion time (makespan), total job flowtime and computation time. Hybrid configurations achieve the best aggregate solution quality, although the leading method depends on the problem class and objective. A separate synthetic workflow study uses OpenAI's GPT-6-astra as an interaction layer between the human planner and the algorithmic core, testing rule additions, plan and objective changes, and what-if comparisons. All 24 sessions completed the requested changes and passed independent checks of saved models and schedules. A separate coding evaluation produced six native implementations of an additional objective or hard constraint through predefined extension hooks. All passed independent checks without modifying the core.
Chinese Translation
生产调度需要能够反映运营约束的现实模型,以及能够平衡相互冲突目标的高效方法。将这些方法投入应用还需要数据集成、模型适配和专家知识。我们提出了APEX,一个围绕通用模型和混合多目标搜索构建的可扩展生产调度框架。智能体辅助支持调度和模型细化:智能体以自然语言准备数据并探索场景,而编码智能体帮助实现和测试新的约束与目标。共享的构建和检查流程将这些适配连接到调度核心。我们在69个公开的作业车间、柔性作业车间和置换流水车间实例上,将八种APEX配置与NSGA-II、SPEA2、MOEA/D和SMS-EMOA进行基准测试,评估工作负载完成时间(完工时间)、总作业流程时间和计算时间。混合配置实现了最佳的综合解质量,尽管领先方法取决于问题类别和目标。一项独立的合成工作流研究使用OpenAI的GPT-6-astra作为人类规划者与算法核心之间的交互层,测试规则添加、计划和目标更改以及假设分析比较。所有24个会话都完成了请求的更改,并通过了对保存的模型和调度的独立检查。一项独立的编码评估通过预定义的扩展钩子,产生了额外目标或硬约束的六种原生实现。所有实现均通过了独立检查,且未修改核心。
cs.AI / 214 / 2609.33439

Raven: The Harness of Harnesses for Composable Agentic Intelligence

Raven:面向可组合智能体智能的框架之框架
AI, EverMind
Abstract
As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to specific domains limits the generality of a single harness. The central question thus shifts from how to engineer a stronger harness for one domain to how to autonomously construct specialized harnesses, improve them through experience, and orchestrate them across domains. We introduce Raven, \emph{The Harness of Harnesses}, an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains, treating each executable model--harness pair as a composable unit of intelligence. To support an \emph{All-Domain Collaboration Network}, its Host Agent decomposes goals, matches subtasks to specialized agents, coordinates execution dependencies, and integrates results, while a host archive and EverOS preserve experience across tasks and Skill Forge makes that experience available as reusable procedures. Our theory establishes sufficient conditions for such composition to expand reliable task coverage beyond that of the available individual agents under a shared resource budget. On complex and long-horizon tasks, Raven significantly outperforms the state-of-the-art agent systems, pushing the frontier of composable agentic intelligence.
Chinese Translation
随着大语言模型的发展,AI智能体正从孤立的、特定领域的任务转向长视野、跨领域的工作流。这一转变暴露出两个挑战:日益增加的框架复杂性使人工设计难以扩展,而与特定领域的紧密耦合限制了单一框架的通用性。因此,核心问题从如何为单一领域设计更强大的框架,转变为如何自主构建专用框架、通过经验改进它们,并在跨领域间进行编排。我们提出Raven——The Harness of Harnesses,一个开源的多智能体生态系统,它能自动为特定模型和领域构建并演化模块化框架,将每个可执行的模型-框架对视为可组合的智能单元。为了支持All-Domain Collaboration Network,其宿主智能体分解目标,将子任务匹配到专用智能体,协调执行依赖关系并整合结果,同时宿主档案和EverOS跨任务保存经验,而Skill Forge将这些经验作为可重用过程提供。我们的理论为这种组合在共享资源预算下将可靠任务覆盖范围扩展到超越可用个体智能体的水平建立了充分条件。在复杂和长视野任务上,Raven显著优于最先进的智能体系统,推动了可组合智能体智能的前沿。
cs.AI / 215 / 2609.33440

MAC-Net: A Multi-Task Deep Learning Framework for Modeling Cognitive Function From Task-Based fMRI

MAC-Net:一种基于任务态fMRI建模认知功能的多任务深度学习框架
Rahman, Md. Tanvir, Orka, Nabil Anan, Khan, Asaduzzaman, Moni, Mohammad Ali
Abstract
Objective cognitive assessment from neural signals supports neurorehabilitation, but individual-level prediction from task-based fMRI (tfMRI) remains difficult because neural features coexist with substantial demographic and scanner-related variation. We present the Multi-task Activation and Contrast Network (MAC-Net), a covariate-aware deep learning framework for modeling individual cognitive function from regional tfMRI. By isolating tfMRI features into a dedicated neural pathway and restricting participant variables to a terminal late-fusion pathway, MAC-Net prevents dominant covariates from suppressing high-dimensional clinical representations during feature learning. Evaluating baseline data from 6,500 Adolescent Brain Cognitive Development Study participants under family-aware cross-validation, MAC-Net was benchmarked against linear models, random forests, and alternative deep architectures. The N-back plus Monetary Incentive Delay configuration achieved $R^{2}$ values of 0.174, 0.238, and 0.277 for fluid, crystallized, and total cognition, outperforming covariate-only baselines (0.178) and alternative deep models (0.217). N-back was the most informative paradigm, whereas incorporating the Stop Signal Task marginally degraded performance. Feature attributions via Integrated Gradients, DeepLIFT, and Input Gradient were highly concordant, localizing working-memory-related frontal, parietal, and cingulate regions. These findings demonstrate that covariate-aware multi-task modeling yields reproducible cognitive-function estimations, establishing a robust neural engineering framework for clinical translation.
Chinese Translation
基于神经信号的客观认知评估可为神经康复提供支持,但从任务态fMRI(tfMRI)进行个体水平预测仍然困难,因为神经特征与显著的人口统计学和扫描仪相关变异共存。我们提出多任务激活与对比网络(Multi-task Activation and Contrast Network, MAC-Net),一种协变量感知的深度学习框架,用于从区域tfMRI建模个体认知功能。通过将tfMRI特征隔离到专门的神经通路,并将参与者变量限制在终端晚期融合通路中,MAC-Net可防止主导协变量在特征学习过程中抑制高维临床表征。在家庭结构感知交叉验证下,评估了6,500名青少年脑认知发展研究(Adolescent Brain Cognitive Development Study)参与者的基线数据,并将MAC-Net与线性模型、随机森林和其他深度学习架构进行基准比较。N-back加Monetary Incentive Delay配置在流体认知、晶体认知和总认知上分别达到$R^{2}$值为0.174、0.238和0.277,优于仅协变量基线(0.178)和其他深度模型(0.217)。N-back是信息量最大的范式,而纳入停止信号任务(Stop Signal Task)会略微降低性能。通过Integrated Gradients、DeepLIFT和Input Gradient进行的特征归因高度一致,定位于与工作记忆相关的额叶、顶叶和扣带回区域。这些发现表明,协变量感知的多任务建模可产生可重复的认知功能估计,为临床转化建立了稳健的神经工程框架。
cs.AI / 216 / 2609.33455

What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation

共享前缀隐藏了什么:用于同策略蒸馏的轨迹丢弃(Trajectory Dropout)
Lin, Zzizhuo, Liu, Quanling, Yang, Yi, Luo, Yawei
Abstract
On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We find that shared prefixes can lead to weak token-level updates, a phenomenon we call Prefix-Induced Supervision Attenuation (PISA). This attenuation arises in two common cases. (i) High student confidence can weaken corrective gradients even when the teacher disagrees. (ii) Tokens that rely on earlier reasoning can receive learning signals as weak as those for simple local continuations. To solve this problem, we propose Trajectory Dropout, a simple training-time intervention that exposes these weakened signals. The student first performs a standard full-context rollout to generate a complete trajectory. During training, we randomly drop a certain proportion of the student's reasoning trajectory, while the teacher continues to observe the complete trajectory for token-level supervision. This intervention strengthens corrections for overconfident predictions and introduces additional supervision at prefix-sensitive positions. Trajectory Dropout consistently improves average performance across teacher--student model pairs of different scales and six mathematical reasoning benchmarks, while also yielding gains on two out-of-domain benchmarks. It can also be flexibly integrated into existing OPD variants with negligible computational overhead, further improving their performance. These results demonstrate that Trajectory Dropout provides a simple mechanism for strengthening token-level supervision across model scales and OPD objectives.
Chinese Translation
同策略蒸馏(OPD)利用来自更强教师模型的密集词元级反馈,在学生模型自身的轨迹上对其进行训练。由于每次更新都以学生已生成的推理前缀为条件,该前缀也会影响教师反馈转化为学习的有效程度。我们发现,共享前缀可能导致弱词元级更新,我们将此现象称为前缀诱导的监督衰减(Prefix-Induced Supervision Attenuation, PISA)。这种衰减出现在两种常见情况中。(i)即使教师不同意,学生的高置信度也可能削弱纠正梯度。(ii)依赖早期推理的词元所获得的学习信号可能与简单局部延续的词元一样弱。为解决此问题,我们提出轨迹丢弃(Trajectory Dropout),一种简单的训练时干预,可以暴露这些被削弱的信号。学生首先执行标准的全上下文展开(rollout)以生成完整轨迹。在训练过程中,我们随机丢弃学生推理轨迹的一定比例,而教师继续观察完整轨迹以进行词元级监督。这种干预加强了对过度自信预测的纠正,并在前缀敏感位置引入了额外监督。轨迹丢弃在不同规模的师生模型对和六个数学推理基准上持续提升平均性能,同时在两个域外基准上也取得了增益。它还可以灵活集成到现有的 OPD 变体中,计算开销可忽略不计,进一步提升了它们的性能。这些结果表明,轨迹丢弃提供了一种简单机制,可在不同模型规模和 OPD 目标下加强词元级监督。
cs.AI / 217 / 2609.33458

When Does the Concept of "Dog" Emerge in an Audio LLM?

音频LLM中“狗”的概念何时出现?
Wang, Zhe, Liu, Shiqi, Zhong, Ruiyun, Zhu, Tiechong, Tan, Yihua
Abstract
Multimodal large language models answer audio questions, but how they represent auditory semantics and use them in decisions remains unclear, limiting our understanding of response formation. We study dog barking in Qwen2.5-Omni-7B using Jacobian lens (J-lens) readout and directional interventions. We define the dog direction as a J-lens-derived hidden-state vector associated with dog; adding or removing its component modulates dog-related information. We find this information decodable without dog/bark prompt cues or animal-identification requirements. Directional interventions change response tendencies and some final answers, with effects concentrated in late-layer states immediately before generation across species classification, vocalization classification, and sound description. The dog direction shows no comparable advantage over controls in animal/other classification. These results provide causal-intervention evidence that the dog direction affects output scores in a task-dependent manner, most consistently at L22 and L24 immediately before generation.
Chinese Translation
多模态大语言模型能回答音频问题,但它们如何表征听觉语义并将其用于决策仍不清楚,这限制了我们对响应形成的理解。我们使用雅可比透镜(J-lens)读出和方向干预来研究Qwen2.5-Omni-7B中的狗吠声。我们将狗方向定义为从J-lens导出的与狗相关的隐藏状态向量;添加或移除其分量会调节与狗相关的信息。我们发现,无需狗/吠叫提示线索或动物识别要求,该信息即可被解码。方向干预会改变响应倾向和某些最终答案,其效应集中在生成前的后层状态,涉及物种分类、发声分类和声音描述。在动物/其他分类中,狗方向相较于对照组没有表现出可比的优势。这些结果为狗方向以任务依赖的方式影响输出分数提供了因果干预证据,最一致地出现在生成前的L22和L24层。
cs.AI / 218 / 2609.33470

LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs

LiveOption:评估LLM智能体在具有非线性收益的结构化期权交易中的表现
Luo, Haochen, Li, Yifan, An, Binh Minh, Luo, Xiaolong, Lai, Zhengzhao, Zhang, Yuan, Liu, Chen
Abstract
Large language models (LLMs) and multi-agent systems (MAS) have shown promise in financial decision-making, yet existing evaluations focus on equity trading and primarily assess directional prediction, overlooking the structural complexity of derivative markets. Option trading introduces fundamentally different challenges, including nonlinear payoffs and multi-leg strategy construction, requiring structured decisions rather than simple directional bets. We introduce LiveOption, an evaluation framework for LLM-based agents in option trading. LiveOption formulates the problem as structured sequential decision-making under realistic execution and capital constraints, and provides a reproducible environment with standardized interaction protocols. The framework includes three task suites covering portfolio overlays, event-driven earnings trading, and 0DTE intraday trading. We further propose a hierarchical metric suite that evaluates action validity, decision quality, risk characteristics, and outcome-level performance. Experiments show that current agents often fail to achieve competitive returns in most scenarios. LiveOption offers a principled testbed for evaluating structured decision-making beyond outcome-based metrics.
Chinese Translation
大型语言模型(LLMs)和多智能体系统(MAS)在金融决策中已展现出潜力,然而现有评估主要关注股票交易并侧重方向性预测,忽视了衍生品市场的结构复杂性。期权交易带来了根本不同的挑战,包括非线性收益和多腿策略构建,需要结构化决策而非简单的方向性押注。我们提出了LiveOption,一个用于期权交易中基于LLM的智能体的评估框架。LiveOption将问题形式化为在现实执行和资金约束下的结构化序贯决策,并提供了一个具有标准化交互协议的可复现环境。该框架包含三个任务套件,涵盖投资组合叠加、事件驱动的财报交易和0DTE日内交易。我们进一步提出了一个分层指标套件,用于评估动作有效性、决策质量、风险特征和结果层面表现。实验表明,当前的智能体在大多数场景下往往无法实现有竞争力的回报。LiveOption为评估超越结果指标的结构化决策提供了一个有原则的测试平台。
cs.AI / 219 / 2609.33477

Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs

就让线性状态遗忘遥远的过去:通过后缀重放实现混合LLM的前缀缓存
Liu, Yirui, Qi, Ruoling, Wu, Xuaner, Jin, Yuxin, Chen, Jian, Liu, Penghang, Huang, Yafei, Shao, Jiawei, Li, Xuelong
Abstract
Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching. Full-attention KV caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbitrary prefix boundaries. Existing systems materialize recurrent-state checkpoints, restricting prefix reuse to checkpoint-aligned positions. We present SuffixReplay, the first prefix caching system that lets hybrid LLMs reuse cached prefixes at every cache-supported page boundary without materializing recurrent-state checkpoints. Our key insight is to just let linear states forget the distant past. Modern linear-attention mechanisms use recurrent decay and gating to attenuate the influence of old inputs. Therefore, instead of checkpointing every prefix boundary, SuffixReplay approximates the state at a matched boundary by replaying only a recent suffix of the layer's input hidden states, which we retain as anchors. At the algorithmic level, SuffixReplay combines layer-wise and token-wise anchor sparsity with a bounded replay budget to control storage, computation, and quality. At the system level, it uses an independently managed anchor sidecar and a pipelined replay path to overlap anchor movement and state reconstruction with the native serving pipeline. We evaluate SuffixReplay on three hybrid LLMs: OLMo-Hybrid-7B, Qwen3.5-4B, and Qwen3.6-27B-FP8. Across these models, SuffixReplay retains 91.4-100% of full-prefill quality on average across LongBench and RULER, while using only 0.36-0.51x the amortized per-token storage of SGLang's default 8192-token checkpoint cache. Integrated into SGLang, SuffixReplay reduces median TTFT by 15-70% on branching workloads, sustains 2.3-4.3x SGLang's throughput when the working set exceeds HBM, and matches SGLang on high-hit continuation traffic.
Chinese Translation
混合LLM将全注意力层与线性注意力层交错,以降低长上下文推理成本,但这种结构使前缀缓存复杂化。全注意力KV缓存可按token寻址,而线性注意力层维护的循环状态无法回滚到任意前缀边界。现有系统物化循环状态检查点,将前缀重用限制在与检查点对齐的位置。我们提出SuffixReplay,这是第一个让混合LLM在每个缓存支持的页边界重用缓存前缀而无需物化循环状态检查点的前缀缓存系统。我们的关键洞察是:就让线性状态遗忘遥远的过去。现代线性注意力机制使用循环衰减和门控来减弱旧输入的影响。因此,SuffixReplay不为每个前缀边界设置检查点,而是通过仅重放层输入隐藏状态的近期后缀来近似匹配边界处的状态,我们将这些输入隐藏状态保留为锚点。在算法层面,SuffixReplay将逐层和逐token的锚点稀疏性与有界重放预算相结合,以控制存储、计算和质量。在系统层面,它使用独立管理的锚点边车和流水线化重放路径,将锚点移动和状态重建与原生服务流水线重叠。我们在三个混合LLM上评估SuffixReplay:OLMo-Hybrid-7B、Qwen3.5-4B和Qwen3.6-27B-FP8。在这些模型上,SuffixReplay在LongBench和RULER上平均保留了91.4-100%的全预填充质量,而仅使用SGLang默认8192 token检查点缓存的0.36-0.51倍摊销每token存储。集成到SGLang后,SuffixReplay在分支工作负载上将中位TTFT降低了15-70%,当工作集超过HBM时维持SGLang吞吐量的2.3-4.3倍,并在高命中延续流量上与SGLang持平。
cs.AI / 220 / 2609.33492

Federated Multi-Modal Human Activity Recognition using Multi-Agent Reinforcement Learning

基于多智能体强化学习的联邦多模态人体活动识别
Dey, Debasmita, Sen, Tanmay, Mallick, Himel
Abstract
Human Activity Recognition (HAR) from heterogeneous wearable sensors is fundamental to the Internet of Health Things (IoHT), supporting rehabilitation, elderly care, and smart healthcare. Existing multimodal fusion methods often assign fixed equal weights to sensor streams, overlooking differences in modality importance, acquisition cost, and sensor quality, which can vary due to movement, incorrect placement, or temporary blockage. We propose an adaptive and cost-aware multimodal HAR framework based on multi-agent reinforcement learning for centralized HAR and extend it to federated learning as FedMHAR. In the centralized setting, multimodal fusion is formulated as a cooperative Multi-Agent Reinforcement Learning (MARL) problem, where each sensing modality is assigned a PPO-based agent that learns per-sample fusion weights, enabling the model to emphasize informative modalities while down-weighting costly sensors when cheaper alternatives provide sufficient information. In the federated setting, we introduce BiFL-PPO, a bidirectional federated optimization strategy in which a server-side PPO policy learns client-specific trust weights and feeds them back to adapt local learning rates and proximal regularization. Unlike round-level optimization, BiFL-PPO uses dense batch-level rewards for more frequent feedback and stable training under heterogeneous client data. Evaluation on the MEx Rehabilitation and UTD Multimodal Human Action datasets shows that the centralized framework achieves 87.30% and 94.98% accuracy, respectively, outperforming conventional fusion methods and state-of-the-art HAR models. FedMHAR achieves 79.74% and 77.49% in the federated setting, consistently surpassing FedAvg, FedProx, FedBN, FedNova, and AdaFedProx, while providing more stable performance and reducing sensor acquisition cost.
Chinese Translation
来自异构可穿戴传感器的人体活动识别(HAR)是健康物联网(IoHT)的基础,支持康复、老年人护理和智能医疗。现有的多模态融合方法通常为传感器数据流分配固定的相等权重,忽略了模态重要性、采集成本和传感器质量方面的差异,这些差异可能因运动、放置不当或暂时阻塞而变化。我们提出了一种基于多智能体强化学习的自适应且成本感知的多模态HAR框架,用于集中式HAR,并将其扩展为联邦学习,称为FedMHAR。在集中式设置中,多模态融合被建模为一个协作式多智能体强化学习(MARL)问题,其中每个传感模态被分配一个基于PPO的智能体,学习每个样本的融合权重,使模型能够强调信息量大的模态,同时在更便宜的替代方案提供足够信息时降低昂贵传感器的权重。在联邦设置中,我们引入了BiFL-PPO,一种双向联邦优化策略,其中服务器端PPO策略学习客户端特定的信任权重,并将其反馈以调整本地学习率和近端正则化。与轮级优化不同,BiFL-PPO使用密集的批级奖励,以在异构客户端数据下提供更频繁的反馈和更稳定的训练。在MEx Rehabilitation和UTD Multimodal Human Action数据集上的评估表明,集中式框架分别达到87.30%和94.98%的准确率,优于传统融合方法和最先进的HAR模型。FedMHAR在联邦设置下达到79.74%和77.49%,持续超越FedAvg、FedProx、FedBN、FedNova和AdaFedProx,同时提供更稳定的性能并降低传感器采集成本。
cs.AI / 221 / 2609.33503

RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse

RelaxKV:查询引导的重计算与稀疏上下文注意力实现高效KV缓存重用
Qi, Ruoling, Liu, Yirui, Wu, Xuaner, Jin, Yuxin, Chen, Jian, Qin, Jiayu, Chen, Yin, Shao, Jiawei
Abstract
Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk interactions. Existing methods selectively recompute token states to recover these missing interactions, but primarily allocate the recomputation budget to selecting which states to recompute, while fixing the recomputation context to the full causal prefix. We introduce RelaxKV, which formulates selective cache repair as a joint allocation problem over repair targets and recomputation context. Guided by the user query, RelaxKV identifies layer-specific repair targets and restricts their recomputation to a query-relevant context, reducing attention computation. Across four decoder models, RelaxKV at a 15% anchor ratio improves aggregate LongBench performance over ProphetKV on all models. On Qwen3-14B, RelaxKV provides a stronger quality-TTFT trade-off than ProphetKV across a 5%-30% anchor-ratio sweep, and achieves the best selective results on RULER-MV and LV-Eval at 16K and 32K context lengths. Controlled ablations further demonstrate the importance of recomputation context selection.
Chinese Translation
跨请求KV缓存降低了检索增强生成(RAG)的预填充成本,但传统的前缀缓存严重限制了跨请求的缓存重用。位置无关缓存(Position-Independent Caching, PIC)通过重用独立的块消除了这一限制,但它们的KV状态缺少跨块交互。现有方法选择性地重计算令牌状态以恢复这些缺失的交互,但主要将重计算预算分配给选择哪些状态进行重计算,同时将重计算上下文固定为完整的因果前缀。我们提出了RelaxKV,它将选择性缓存修复形式化为关于修复目标和重计算上下文的联合分配问题。在用户查询的引导下,RelaxKV识别层特定的修复目标,并将它们的重计算限制在与查询相关的上下文中,从而减少注意力计算。在四个解码器模型上,15%锚点比例的RelaxKV在所有模型上都比ProphetKV提高了LongBench的总体性能。在Qwen3-14B上,在5%-30%的锚点比例扫描中,RelaxKV比ProphetKV提供了更强的质量-TTFT权衡,并在16K和32K上下文长度下在RULER-MV和LV-Eval上取得了最佳的选择性结果。控制消融进一步证明了重计算上下文选择的重要性。
cs.AI / 222 / 2609.33505

When Evidence Changes the Subject: Subject-Typed Claim Licensing for Learned Routing

当证据改变主体时:面向学习型路由的主体类型化声明许可
Chen, Jian, Yuan, Zixuan
Abstract
Modern learned systems increasingly combine learned components with search, repair, or external solvers. Benchmarks often measure the resulting end-to-end system, while scientific claims may concern only one component, creating an attribution problem: evidence can fail to support the requested component-level claim while still supporting a positive conclusion about the larger system. Existing evidence-to-claim methods primarily calibrate claim strength. We argue that composite systems require a second dimension: scientific subject. We address this problem with subject-typed claim licensing, which separates weaker conclusions about the requested subject from positive but non-substitutive credit about another subject. We instantiate this idea in SCOPE-Routing for preference-conditioned multigraph routing. Non-authors reproducibly apply the declared semantics; held-out review yields fewer reference-relative upward deviations than unstructured review, while the difference from a strong evidence checklist remains unresolved; and a controlled routing study shows that score-optimal and claim-eligible methods can differ while valid hybrid-system credit is preserved. These results motivate treating claim strength and scientific subject as distinct dimensions of evidence-based evaluation.
Chinese Translation
现代学习系统越来越多地将学习组件与搜索、修复或外部求解器相结合。基准测试通常测量最终端到端系统,而科学声明可能只涉及一个组件,这就产生了一个归因问题:证据可能无法支持所要求的组件级声明,但仍能支持关于更大系统的积极结论。现有的证据到声明的方法主要校准声明强度。我们认为复合系统需要第二个维度:科学主体。我们通过主体类型化声明许可来解决这个问题,它将关于所请求主体的较弱结论与关于另一主体的积极但不可替代的认可区分开来。我们在SCOPE-Routing中实例化了这一想法,用于偏好条件多图路由。非作者可重复地应用所声明的语义;与无结构评审相比,留出评审产生的参考相对向上偏差更少,而与强证据清单的差异仍未解决;一项受控路由研究表明,分数最优和声明合格的方法可能不同,同时有效的混合系统认可得以保留。这些结果促使我们将声明强度和科学主体视为基于证据的评估的不同维度。
cs.AI / 223 / 2609.33509

What Happens During Autonomous Deep Research After the User Steps Away?

用户离开后,自主深度研究期间会发生什么?
Liu, Yimin, Zhang, Yijia, Li, Yanmin, Luo, Tangwen, Li, Yuze, Yao, Ziling, Yang, Zhi
Abstract
In autonomous deep research, a user provides a task and relevant background, then leaves the agent to conduct an extended investigation without further human intervention. We study how this initial user information is reflected in intermediate actions and how these actions relate to final recommendations. We introduce DRaligned, a counterfactual behavioral evaluation framework built on PDR-Bench. By varying one task-relevant user factor while keeping the remaining context fixed, we compare acquisition requests, working drafts, and final reports. Source-grounded extraction, blinded local judgments, and deterministic aggregation yield coarse directional measurements while leaving ambiguous cases unresolved. Our experiments show that strong user-specific delivery can emerge from a largely shared research process: agents investigate similar broad questions but allocate requests differently, and final recommendations distinguish user conditions more clearly than explicit requests do. Reports can also integrate user factors that were not jointly visible during acquisition. In readable draft-to-report comparisons, recommendations often retain their coarse user-specific direction despite substantial rewriting. Final directional differences recur across tested agent models, execution harnesses, and evaluator models, even as execution paths vary. These findings describe how initial user information shapes autonomous research and clarify the relationship between the process an agent follows and the recommendations it delivers.
Chinese Translation
在自主深度研究中,用户提供任务和相关背景,然后离开,让智能体在无需进一步人工干预的情况下进行长时间调查。我们研究这种初始用户信息如何反映在中间行动中,以及这些行动与最终推荐的关系。我们介绍 DRaligned,一个建立在 PDR-Bench 上的反事实行为评估框架。通过改变一个与任务相关的用户因素,同时保持其余上下文固定,我们比较获取请求、工作草稿和最终报告。基于来源的提取、盲法局部判断和确定性聚合产生粗粒度的方向性测量,同时留下模糊情况未解决。我们的实验表明,强烈的用户特定交付可以从很大程度上共享的研究过程中涌现:智能体调查相似的广泛问题,但以不同方式分配请求,并且最终推荐比显式请求更清晰地区分用户条件。报告还可以整合在获取过程中未共同可见的用户因素。在可读的草稿到报告比较中,尽管进行了大量重写,推荐通常保留其粗粒度的用户特定方向。最终方向性差异在测试的智能体模型、执行框架和评估器模型中反复出现,即使执行路径各不相同。这些发现描述了初始用户信息如何塑造自主研究,并阐明了智能体遵循的过程与其提供的推荐之间的关系。
cs.AI / 224 / 2609.33516

PPG-LM: A Photoplethysmography-Language Model with Multi-Level Clinical Alignment

PPG-LM:一种具有多级临床对齐的光电容积描记-语言模型
Wang, Xiaoda, Wang, Minxiao, Xu, Maxwell A, Langer, Patrick, Han, Kaiqiao, Cao, Defu, Luo, Xiao, Yang, Yuzhe, Liu, Yan, Hu, Xiao, Sun, Yizhou, Wang, Wei, Yang, Carl
Abstract
Photoplethysmography (PPG) is widely recorded by clinical monitors and consumer wearables, providing a scalable source of continuous physiological information. These recordings offer an opportunity for physiological assessment at scale, but realizing this potential requires models to learn from both signal-derived physiological supervision and broader clinical context captured in electronic health records (EHRs). This involves aligning information spanning local observations, care events, and entire visits with PPG representations at corresponding temporal scales. However, existing PPG foundation models primarily rely on task-specific prediction heads, while the medical knowledge of large language models does not necessarily translate into waveform understanding. To bridge this gap, we introduce PPG-LM, the first PPG-language model family to learn physiological representations from both signal-derived supervision and broader clinical context captured in EHRs. To construct clinically grounded captions, we develop an automatic captioning pipeline that generates segment-, event-, and visit-level descriptions from signal measurements and structured EHR records. We then learn from these pairs through a two-stage framework that first establishes segment-language correspondence through contrastive learning and waveform-conditioned captioning, then extends alignment to events and visits through time-aware aggregation and temporal statement matching. Pretrained on approximately 73k hours of PPG, PPG-LM supports language-based recognition, cross-modal retrieval, and segment captioning. Experiments on MC-MED, MIMIC-III, and VitalDB show improved retrieval and caption factuality over language-model baselines and gains over PPG and time-series foundation models on multiple clinical prediction tasks.
Chinese Translation
光电容积描记法 (PPG) 被临床监护仪和消费级可穿戴设备广泛记录,提供了一种可扩展的连续生理信息来源。这些记录为大规模生理评估提供了机会,但实现这一潜力需要模型同时从信号衍生的生理监督和电子健康记录 (EHR) 中捕获的更广泛临床背景中学习。这涉及将跨越局部观察、护理事件和整个就诊的信息与相应时间尺度上的 PPG 表示对齐。然而,现有的 PPG 基础模型主要依赖任务特定的预测头,而大型语言模型的医学知识并不一定能转化为波形理解。为了弥合这一差距,我们引入了 PPG-LM,这是首个从信号衍生监督和 EHR 中捕获的更广泛临床背景中学习生理表示的 PPG-语言模型系列。为了构建具有临床依据的字幕,我们开发了一个自动字幕生成流程,从信号测量和结构化 EHR 记录中生成片段级、事件级和就诊级描述。然后,我们通过一个两阶段框架从这些配对中学习,该框架首先通过对比学习和波形条件字幕生成建立片段-语言对应关系,然后通过时间感知聚合和时间语句匹配将对齐扩展到事件和就诊。在约 73k 小时的 PPG 上预训练后,PPG-LM 支持基于语言的识别、跨模态检索和片段字幕生成。在 MC-MED、MIMIC-III 和 VitalDB 上的实验表明,与语言模型基线相比,检索和字幕事实性有所提高,并且在多个临床预测任务上优于 PPG 和时间序列基础模型。
cs.AI / 225 / 2609.33524

EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research

EverMine:剖析长周期Alpha研究中研究能力的自演化
Li, Siyuan, Zhang, Jiangfeng, Yao, Rui, Qiu, Weihua, Xu, Mingyang, Yuan, Zixuan
Abstract
Self-evolving agents aim to turn research feedback into reusable skills, tools, and research rules. Whether these accumulated capabilities continue to improve later research requires controlled evaluation. Long-horizon alpha discovery provides a state-dependent setting: once a new factor enters the portfolio, the predictive information already covered changes, so the value of the same candidate or experience may change over time. We introduce EverMine, an empirical framework for studying self-evolving research capabilities in long-horizon alpha discovery. EverMine decomposes the research state into history (Hist), the current factor portfolio (Frontier), and reusable capabilities (Cap). Under matched resource limits, we compare complete runs with fixed or evolving Cap, and replace Cap while holding Hist and Frontier fixed to estimate the conditional value of accumulated capabilities. We also combine full trajectories with historical-state replay to examine how experience-based decisions affect candidate selection and portfolio outcomes. Across 18 long-horizon trajectories, end-to-end comparisons show no consistent gain from Cap evolution. Across 48 continuation branches from shared Hist and Frontier states, accumulated Cap also does not consistently outperform the initial Cap. Parameter tuning of existing factor structures can still improve the portfolio. In an exploratory replay of two screening batches from one Evolving trajectory, some screened-out candidates have positive marginal value at the original state, yet submitting all screened-out candidates sequentially slightly lowers final portfolio IC in both batches. These results show that candidate value depends on the evolving portfolio and submission order, and motivate evaluating self-evolving research capabilities through end-to-end outcomes, conditional capability value, and the consequences of experience-based decisions.
Chinese Translation
自演化智能体旨在将研究反馈转化为可重用的技能、工具和研究规则。这些积累的能力是否能持续改进后续研究,需要受控评估。长周期Alpha发现提供了一个状态依赖的场景:一旦新因子进入组合,已覆盖的预测信息发生变化,因此同一候选因子或经验的价值可能随时间变化。我们提出EverMine,一个用于研究长周期Alpha发现中自演化研究能力的实证框架。EverMine将研究状态分解为历史(Hist)、当前因子组合(Frontier)和可重用能力(Cap)。在资源限制匹配的条件下,我们比较固定Cap或演化Cap的完整运行,并在保持Hist和Frontier固定的情况下替换Cap,以估计积累能力的条件价值。我们还结合完整轨迹与历史状态重放,以检验基于经验的决策如何影响候选因子选择和组合结果。在18个长周期轨迹中,端到端比较表明Cap演化没有带来一致的收益。在来自共享Hist和Frontier状态的48个延续分支中,积累的Cap也没有一致地优于初始Cap。对现有因子结构的参数调优仍然可以改进组合。在对一个演化轨迹的两个筛选批次进行探索性重放时,一些被筛除的候选因子在原始状态下具有正边际价值,然而按顺序提交所有被筛除的候选因子在两个批次中都略微降低了最终组合IC。这些结果表明,候选因子的价值取决于演化的组合和提交顺序,并促使通过端到端结果、条件能力价值和基于经验决策的后果来评估自演化的研究能力。
cs.AI / 226 / 2609.33540

Reasoning on the Simplex: Geometric Fixed-Point Models

在单纯形上的推理:几何不动点模型
Daulbaev, Talgat, Glazkov, Ilya, Rakhuba, Maxim, Oseledets, Ivan
Abstract
Looped reasoners spend test-time compute by iterating a weight-tied map, but a small residual does not mean the state is a fixed point when that map lives in unconstrained latent space. We propose Geometric Fixed-Point Reasoning (GFPR), in which the iterated state is the prediction itself: a field of categorical beliefs on a product of simplices, whose argmax is the answer at every step. Because the state is a belief, task structure can be imposed through compact convex relaxations, either as structured readouts or directly in the recurrent state; in the latter case the update remains a continuous self-map, so a fixed point exists for any parameters. At about 7M parameters, GFPR reaches 95.1% exact match on Sudoku-Extreme, 92.0% on Maze-Hard, and 100% sequence accuracy on S_5 length 128, above the published FPRM numbers at the same scale. The same update also trains a 201M language model on FineWeb-Edu in which each site is a distribution over the vocabulary; with 24 Picard steps it is above GPT-2 small on four zero-shot multiple-choice tasks and above GPT-2 medium on ARC-Easy.
Chinese Translation
循环推理器通过迭代权重绑定的映射来消耗测试时的计算量,但当该映射存在于无约束的潜在空间中时,小的残差并不意味着状态是一个不动点。我们提出几何不动点推理(GFPR),其中迭代状态就是预测本身:一个在单纯形乘积上的类别信念场,其 argmax 在每一步都是答案。由于状态是信念,任务结构可以通过紧凸松弛来施加,或者作为结构化读出,或者直接在循环状态中;在后一种情况下,更新仍然是一个连续自映射,因此对于任何参数都存在一个不动点。在约 7M 参数下,GFPR 在 Sudoku-Extreme 上达到 95.1% 的精确匹配,在 Maze-Hard 上达到 92.0%,在 S_5 长度 128 上达到 100% 的序列准确率,高于同一规模下已发表的 FPRM 数值。同样的更新还在 FineWeb-Edu 上训练了一个 201M 的语言模型,其中每个位置是词汇表上的分布;经过 24 次 Picard 步骤,它在四个零样本多项选择任务上超过了 GPT-2 small,在 ARC-Easy 上超过了 GPT-2 medium。
cs.AI / 227 / 2609.33543

OSCC: Certified Observation-Safe Coupling Optimization for Gradient-Noise Control in Imperfect-Information Learning

OSCC:面向不完美信息学习中梯度噪声控制的认证观测安全耦合优化
Hu, Miaobo, Hu, Shuhao, Guo, Xiaobo, Wang, Xin, Wang, Bokun, Chen, Rui, Zha, Daren, Xiao, Jun
Abstract
Coupled rollouts can reduce the noise of counterfactual action comparisons, but two issues prevent standard common-random-number constructions from serving as a general learning primitive in imperfect-information environments. First, an invalid coupling may expose hidden state, synchronize endogenous policy randomness, or misalign chance events after counterfactual histories diverge. Second, in multi-action policy optimization, lower return-contrast variance is not by itself the relevant objective: the optimizer depends on the return covariance matrix after projection through the local policy-gradient geometry. We introduce observation-safe counterfactual coupling (OSCC), a framework that defines an admissible class through marginal preservation, information-state safety, branch-local policy randomness, semantic event alignment, and trace-before-oracle replay. We derive a gradient-aware coupling criterion showing that, for marginal-preserving couplings, policy-gradient noise changes are determined by policy-Jacobian-weighted off-diagonal return covariance. This motivates OSCC-Select, a calibration-only selector that chooses among independent, root-only, continuation-only, and fully coupled rollouts using separate safety and gain certificates. Its gain target combines projected gradient noise with measured physical sampling cost and falls back to independent sampling whenever a simultaneous lower confidence bound does not certify improvement. On 100,000 fixed-root Leduc comparisons, the fully coupled CP-GRPO instantiation reduces return-contrast variance from 41.1158 to 18.1441, a 55.87% reduction, while preserving the declared branch marginals. With three actions, OSCC-Select chooses continuation coupling and attains gradient-noise trace 0.0783 versus 0.0917 for return-variance selection. Increasing calibration from 64 to 2,048 groups raises certification from 0.327 to 0.995.
Chinese Translation
耦合轨迹可以降低反事实动作比较的噪声,但有两个问题阻碍了标准共同随机数构造在不完美信息环境中作为通用学习原语。首先,无效的耦合可能暴露隐藏状态、同步内生策略随机性,或在反事实历史分叉后错位机会事件。其次,在多动作策略优化中,较低的回报对比方差本身并不是相关目标:优化器依赖于通过局部策略梯度几何投影后的回报协方差矩阵。我们引入了观测安全反事实耦合(OSCC),一个通过边缘保持、信息状态安全、分支局部策略随机性、语义事件对齐和预言前追踪回放定义可接受类别的框架。我们推导了一个梯度感知的耦合准则,表明对于边缘保持耦合,策略梯度噪声变化由策略雅可比加权的非对角回报协方差决定。这启发了OSCC-Select,一种仅校准选择器,使用单独的安全性和增益证书在独立、仅根、仅延续和完全耦合的轨迹之间进行选择。其增益目标将投影梯度噪声与测量的物理采样成本相结合,并且当同时下置信界不能证明改进时,回退到独立采样。在100,000次固定根Leduc比较中,完全耦合的CP-GRPO实例将回报对比方差从41.1158降低到18.1441,降低了55.87%,同时保留了声明的分支边缘。在三个动作的情况下,OSCC-Select选择延续耦合,并获得梯度噪声迹0.0783,而回报方差选择为0.0917。将校准从64组增加到2,048组,将认证从0.327提高到0.995。
cs.AI / 228 / 2609.33565

Dr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search Agents

Dr. Free:自进化搜索智能体无需难度奖励
Qian, Zhipeng, Liang, Zihan, Ma, Yufei, Ma, Jie, Chen, Ben, Dai, Huangyu, Mao, Lingtao, Sun, Xinyu, zhao, Tong, Zhang, Xuxin, Cai, Qingpeng, Jiang, Peng, Hou, Qibin
Abstract
A central limitation of current data-free self-evolution methods for training search agents is their reliance on difficulty-based proposer rewards. These methods reward a proposer for generating questions that challenge a co-evolving solver, using solver difficulty as a proxy for question quality. Yet difficulty alone is insufficient to distinguish questions that require cross-passage evidence from those that are answerable via simpler shortcuts. In addition, measuring difficulty demands repeated solver rollouts for every candidate question, leading to substantial computational costs. In this paper, we introduce \methodname, the first self-evolving search framework that eliminates difficulty-based proposer rewards and directly optimizes for evidence necessity relative to shortcut contexts. Dr. Free samples relational chains from a knowledge graph and pairs them with aligned passages, giving question generation an explicit multi-hop structure. A generated question receives a positive information-gain reward only when the likelihood of the target answer under the complete evidence passages exceeds the maximum likelihood under all evaluated shortcut contexts. Because this signal is computed from teacher-forced likelihoods, it removes the need for pass-rate estimation and reduces proposer training time by over $7\times$. Experiments on seven open-domain QA benchmarks show that Dr. Free outperforms prior data-free search agents and the supervised baseline, with large improvements on multi-hop QA benchmarks.
Chinese Translation
当前用于训练搜索智能体的无数据自进化方法的一个核心局限在于它们依赖于基于难度的提议者奖励。这些方法奖励提议者生成能够挑战共同进化求解器的问题,并使用求解器难度作为问题质量的代理指标。然而,仅凭难度不足以区分需要跨段落证据的问题和可通过更简单捷径回答的问题。此外,测量难度需要对每个候选问题进行重复的求解器推演,导致大量计算成本。在本文中,我们介绍了 Dr. Free,这是第一个自进化搜索框架,它消除了基于难度的提议者奖励,并直接优化相对于捷径上下文的证据必要性。Dr. Free 从知识图谱中采样关系链,并将其与对齐的段落配对,为问题生成提供了明确的多跳结构。生成的问题仅当在完整证据段落下的目标答案似然超过所有评估的捷径上下文下的最大似然时,才会获得正的信息增益奖励。由于该信号是从教师强制似然计算得出的,它消除了对通过率估计的需求,并将提议者训练时间减少了 7 倍以上。在七个开放域问答基准上的实验表明,Dr. Free 优于先前的无数据搜索智能体和监督基线,在多跳问答基准上取得了大幅改进。
cs.AI / 229 / 2609.33579

OpenFC: Learning Verification Policies towards Open-Search Fact Checking

OpenFC:学习面向开放搜索事实核查的验证策略
Wang, Xinming, Qiu, Kaixiang, Lin, Yansong, Lv, Chunji, Chen, Yi, Wang, Boran, Yang, Hong-Ming, Zhang, Xu-Yao
Abstract
Open-search fact checking is not merely retrieval followed by classification, but a sequential decision problem in which every query, source visit, and stopping decision reshapes the evidence available for verification. Yet existing systems often distribute these decisions across predefined pipelines or separately prompted modules rather than learning them as a unified task-specific policy. We introduce \textbf{OpenFC}, a unified verification-policy training framework that post-trains Qwen3-8B as a compact next-action controller over reasoning, evidence acquisition, and stopping. OpenFC learns this policy in two stages. \textbf{Stepwise-Calibrated Cold Start (SCCS)} uses a strong training-time supervisor to review post-initial reasoning, tool-use, and stopping proposals before execution, producing reliable trajectories for supervised fine-tuning without access to gold verdicts. \textbf{Verification-Aware Reinforcement Learning (VA-RL)} then improves the cold-start policy on unresolved claims through budget-aware tool rewards, label-aware advantage reweighting, and localized response masking. Across six fact-checking benchmarks, OpenFC achieves 70.39\% average accuracy and 63.30\% macro-F1, the highest overall averages among the evaluated methods. Stage-wise ablations further show that SCCS and VA-RL provide complementary gains, supporting the design of the two-stage training framework. These results position OpenFC as a strong and effective framework for open-search fact-checking. We will open-source our code and release the model checkpoints to support reproducibility.
Chinese Translation
开放搜索事实核查不仅仅是检索后分类,而是一个顺序决策问题:每个查询、来源访问和停止决策都会重塑可用于验证的证据。然而,现有系统通常将这些决策分散在预定义流程或单独提示的模块中,而不是将其作为统一的任务特定策略来学习。我们提出 OpenFC,一个统一的验证策略训练框架,将 Qwen3-8B 后训练为紧凑的下一动作控制器,用于推理、证据获取和停止。OpenFC 分两阶段学习该策略。逐步校准冷启动(Stepwise-Calibrated Cold Start, SCCS)使用强大的训练时监督器,在执行前审查初始推理后的推理、工具使用和停止提议,从而在无需访问金标准裁决的情况下生成可靠轨迹用于监督微调。验证感知强化学习(Verification-Aware Reinforcement Learning, VA-RL)随后通过预算感知工具奖励、标签感知优势重加权和局部响应掩码,改进未解决声明上的冷启动策略。在六个事实核查基准上,OpenFC 达到 70.39% 平均准确率和 63.30% 宏 F1,是评估方法中总体平均最高。阶段消融进一步表明,SCCS 和 VA-RL 提供互补增益,支持两阶段训练框架的设计。这些结果使 OpenFC 成为开放搜索事实核查的强大有效框架。我们将开源代码并发布模型检查点以支持可复现性。
cs.AI / 230 / 2609.33601

JustQuant: You Don't Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization

JustQuant:4比特激活量化无需平滑、SVD或旋转
Yang, Kaicheng, Yang, Kaisen, Liu, Chunyu, Yan, Xianglong, Qin, Haotong, Wu, Junyi, Zhang, Tianao, Zhang, Xun, Zhang, Shaoqiu, Sun, Youbang, Zhang, Yulun
Abstract
Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization. Recent post-training quantization (PTQ) and quantization-aware training (QAT) methods have made progress in 4-bit activation quantization by introducing smoothing, SVD branches, rotations, mixed precision, or advanced formats such as NVFP4. These additional operators and data types impose demanding requirements on inference engines and hardware, limiting the broad adoption of low-precision models. Can quantization be achieved using only plain low-bit operators? To answer this question, we propose JustQuant, a simple yet effective framework that moves the complexity of low-bit quantization from deployment-time operators into the training process. We first revisit model quantization from the perspective of knowledge distillation and show that a key reason existing PTQ and QAT methods fail is that they typically exploit supervision at only a single level. We then introduce Theseus QAD, a quantization-aware distillation method that progressively applies multi-level supervision, analogous to the gradual replacement process in the Ship of Theseus. Extensive experiments on DiT and diffusion large language models show two distinct regimes. For smaller models, Theseus QAD can serve as a lightweight warm-up stage that substantially improves subsequent QAT with plain operators, while naive QAD may collapse in the same setting. For larger models, Theseus QAD provides a stronger distillation training path than ordinary QAD. Across both regimes, JustQuant improves low-bit quantization quality while avoiding the complex operators required by many existing PTQ methods.
Chinese Translation
近期生成模型变得越来越强大,但其推理成本持续增长。模型量化提供了一种有前途的方法来压缩这些模型并加速推理。然而,在4比特下,激活量化比权重量化更具挑战性。最近的后训练量化(PTQ)和量化感知训练(QAT)方法通过引入平滑、SVD分支、旋转、混合精度或高级格式(如NVFP4)在4比特激活量化方面取得了进展。这些额外的算子和数据类型对推理引擎和硬件提出了苛刻的要求,限制了低精度模型的广泛采用。仅使用普通低位算子能否实现量化?为了回答这个问题,我们提出了JustQuant,一个简单而有效的框架,将低位量化的复杂性从部署时的算子转移到训练过程中。我们首先从知识蒸馏的角度重新审视模型量化,并表明现有PTQ和QAT方法失败的一个关键原因是它们通常仅利用单一级别的监督。然后我们介绍了Theseus QAD,一种量化感知蒸馏方法,它逐步应用多级监督,类似于忒修斯之船中的逐步替换过程。在DiT和扩散大语言模型上的大量实验显示了两种不同的机制。对于较小的模型,Theseus QAD可以作为轻量级预热阶段,显著改善后续使用普通算子的QAT,而朴素的QAD在相同设置下可能会崩溃。对于较大的模型,Theseus QAD提供了比普通QAD更强的蒸馏训练路径。在两种机制下,JustQuant提高了低位量化质量,同时避免了许多现有PTQ方法所需的复杂算子。
cs.AI / 231 / 2609.33610

Supervision Recovery for Time Series Anomaly Detection via Context-Anchored Pairing

基于上下文锚定配对的时间序列异常检测监督恢复
Gao, Yifei, Lan, Tian, Lu, Yimeng, An, Xuming, Wang, Meng, He, Wenjun, Li, Yijie, Zhang, Chen
Abstract
Time series anomaly detection (TSAD) remains challenging not only because anomaly labels are scarce, but also because temporal anomalies are highly context-dependent. Existing methods often rely on unsupervised objectives or surrogate abnormal patterns, providing limited supervision for context-dependent normal--anomalous distinctions. We propose Context-Anchored Pair Supervision (CAPS), a supervision-recovery framework for TSAD. CAPS views ideal anomaly supervision as a matched comparison between normal and anomalous outcomes under the same temporal context, and seeks to recover such supervision without target-domain anomaly labels. Using simulated normal--anomalous pairs, CAPS learns structure and anomaly-semantic representations through reconstruction, background consistency, and within-pair counterfactual recombination. The resulting anomaly representations form a continuous semantic space with coarse modes and induce a sampleable multimodal prior. CAPS conditionally realizes sampled semantics as residual-form effects on target reference trajectories. The resulting context-anchored normal--anomalous counterparts provide temporal supervision for discriminative detector learning. Experiments on nine datasets show that CAPS achieves the strongest aggregate performance across all four evaluation metrics among the compared methods, while complementary ablations and transfer analyses support the roles of context anchoring, semantic disentanglement, and conditional realization.
Chinese Translation
时间序列异常检测(TSAD)仍然具有挑战性,不仅因为异常标签稀缺,还因为时间异常高度依赖上下文。现有方法通常依赖无监督目标或替代异常模式,为依赖上下文的正常—异常区分提供的监督有限。我们提出上下文锚定配对监督(CAPS),一种用于TSAD的监督恢复框架。CAPS将理想异常监督视为同一时间上下文下正常与异常结果之间的匹配比较,并力求在无目标域异常标签的情况下恢复这种监督。利用模拟的正常—异常配对,CAPS通过重构、背景一致性和配对内反事实重组学习结构表示与异常语义表示。所得的异常表示形成一个具有粗粒度模式的连续语义空间,并诱导出一个可采样的多模态先验。CAPS条件化地将采样得到的语义实现为目标参考轨迹上的残差形式效应。由此得到的上下文锚定的正常—异常对应样本为判别式检测器学习提供时间监督。在九个数据集上的实验表明,在所有比较方法中,CAPS在四个评价指标上均取得最强的综合性能;同时,补充消融实验和迁移分析支持了上下文锚定、语义解耦和条件化实现的作用。
cs.AI / 232 / 2609.33614

EAT: Expert Account Tracker for Efficient MoE Inference

EAT:面向高效 MoE 推理的专家账户追踪器
Li, Yuexian, Yang, Yifei, Cao, Zouying, Zhao, Hai
Abstract
Mixture-of-Experts (MoE) models have emerged as a revolutionary method to scale Transformer models. However, traditional MoE architecture still suffers from inefficiency since a large number of experts are unnecessarily activated. Existing approaches for reducing the number of activated experts often overlook the historical performance of each expert. In this paper, we propose EAT, a novel method called Expert Account Tracker (EAT), which utilizes history-awareness metrics and adaptive thresholding to dynamically select the most important experts, thereby reducing the activated expert number while effectively maintaining the model performance. Experiments show that EAT outperforms the existing baseline Top-P method across multiple models and datasets, achieving over 25% an average reduction compared to the vanilla method in the number of activated experts and performing better token generation speed compared to the baseline. Furthermore, the performance of pruned models can be efficiently recovered via OPD using only 9K data. Additionally, through ablation studies, we find that excessively reducing the number of activated experts can significantly harm model performance, and the importance of experts varies across layers, with higher-level experts being generally more critical.
Chinese Translation
混合专家(MoE)模型已成为扩展 Transformer 模型的一种革命性方法。然而,传统 MoE 架构仍然存在效率低下的问题,因为大量专家会被不必要地激活。现有减少激活专家数量的方法往往忽略了每个专家的历史表现。在本文中,我们提出了 EAT,一种称为专家账户追踪器(Expert Account Tracker, EAT)的新方法,它利用历史感知指标和自适应阈值来动态选择最重要的专家,从而在有效保持模型性能的同时减少激活专家数量。实验表明,EAT 在多个模型和数据集上优于现有基线 Top-P 方法,与原始方法相比,在激活专家数量上平均减少超过 25%,并且相比基线具有更好的 token 生成速度。此外,剪枝模型的性能可以通过 OPD 仅使用 9K 数据高效恢复。另外,通过消融研究,我们发现过度减少激活专家数量会显著损害模型性能,而且专家的重要性因层而异,通常更高层的专家更为关键。
cs.AI / 233 / 2609.33618

ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments

ParaAgent:强化开放世界工具环境中的并行行动
Yue, Shengbin, Wang, Hongru, Wang, Siyuan, Chen, Xiaoxin, Chen, Wei, Wei, Zhongyu
Abstract
Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them without coordination. We argue that the key lies not in whether to decouple or interleave them, but in how to coordinate them across granularities. We introduce ParaAct, a structured parallel-action loop that combines phase-level Exploration $\rightleftharpoons$ Execution with action-level parallelism. To learn this loop, ParaAgent combines multi-agent cold-start demonstrations with reinforcement learning under multi-level advantage decoupling, making planning structure explicit and supervising it with step-, phase-, and trajectory-level rewards. Learning is supported by our ToolEnv, a scalable simulator grounded in 50,011 realistic tool interfaces. On two open-world tool benchmarks, ParaAgent-4B achieves the best average success among all baselines, including GPT-4.1 systems, with the largest gains on multi-tool tasks. Behavioral analyses show that these gains stem from this action organization, highlighting its importance for capable and efficient open-world agents.
Chinese Translation
语言模型智能体越来越多地部署在开放世界工具环境中,这需要平衡探索未知能力和利用已知能力。现有方法面临性能与效率的权衡:它们要么僵硬地将探索和执行解耦,要么在没有协调的情况下交错进行。我们认为,关键不在于是否解耦或交错它们,而在于如何跨粒度协调它们。我们引入ParaAct,一种结构化的并行行动循环,它将阶段级的探索⇌执行与行动级并行性相结合。为了学习这个循环,ParaAgent结合了多智能体冷启动演示与多级优势解耦下的强化学习,使规划结构显式化,并用步骤级、阶段级和轨迹级奖励对其进行监督。学习得到了我们的ToolEnv的支持,这是一个基于50,011个真实工具接口的可扩展模拟器。在两个开放世界工具基准上,ParaAgent-4B在所有基线(包括GPT-4.1系统)中取得了最佳平均成功率,在多工具任务上增益最大。行为分析表明,这些增益源于这种行动组织,突显了其对能力强且高效的开放世界智能体的重要性。
cs.AI / 234 / 2609.33639

Trajectory Unlearning on LLM-based Agents

基于LLM的智能体轨迹遗忘
Shi, Yingdan, Wang, Ren
Abstract
Existing large language model (LLM) unlearning has focused primarily on removing specific knowledge, such as harmful facts, private data, or copyrighted content. However, as LLMs are increasingly deployed as autonomous agents, a fundamental yet overlooked problem emerges: beyond suppressing what an agent knows, an agent should not reproduce undesired behaviors through its action trajectories. In this work, we introduce trajectory-level unlearning, a new problem formulation that targets the removal of specific action trajectories in long-horizon agentic tasks, rather than factual knowledge. We identify two fundamental challenges that distinguish trajectory unlearning from knowledge unlearning: (1) our unlearning target is what the agent \emph{does}, not what it \emph{says}; and (2) trajectories are sequentially dependent action sequences that cannot be decomposed into isolated prompt-response pairs without losing inter-step structure. To address these challenges, we propose Group-injected Relative Policy Optimization (GiRPO), which injects forget trajectories into the policy rollout group with penalized rewards and isolates the normalization statistics, yielding a stable and bounded unlearning signal that does not corrupt gradient updates for normal task trajectories. We construct trajectory unlearning benchmarks from two application scenarios, household tasks (ALFWorld) and online shopping (WebShop), and design three complementary metrics for evaluating forgetting quality and model utility. Experiments on ALFWorld and WebShop demonstrate that GiRPO effectively unlearns target trajectories while preserving task success rates, outperforming existing knowledge-unlearning baselines on both forgetting quality and task utility.
Chinese Translation
现有的大语言模型(LLM)遗忘研究主要关注于去除特定知识,例如有害事实、私有数据或受版权保护的内容。然而,随着LLM越来越多地被部署为自主智能体,一个基本但被忽视的问题出现了:除了抑制智能体所知道的内容外,智能体不应通过其动作轨迹重现不良行为。在这项工作中,我们引入了轨迹级遗忘,这是一种新的问题形式,其目标是在长视野智能体任务中移除特定的动作轨迹,而非事实性知识。我们确定了区分轨迹遗忘与知识遗忘的两个基本挑战:(1)我们的遗忘目标是智能体“做”什么,而不是它“说”什么;(2)轨迹是顺序依赖的动作序列,如果分解为孤立的提示-响应对,就会失去步骤间的结构。为了应对这些挑战,我们提出了组注入相对策略优化(GiRPO),它将遗忘轨迹以惩罚奖励的方式注入策略展开组,并隔离归一化统计量,从而产生稳定且有界的遗忘信号,不会破坏正常任务轨迹的梯度更新。我们从两个应用场景构建了轨迹遗忘基准:家庭任务(ALFWorld)和在线购物(WebShop),并设计了三个互补的指标来评估遗忘质量和模型效用。在ALFWorld和WebShop上的实验表明,GiRPO能够有效遗忘目标轨迹,同时保持任务成功率,在遗忘质量和任务效用方面均优于现有的知识遗忘基线。
cs.AI / 235 / 2609.33641

Scalable and Data-Driven Decision Support in the Maintenance, Repair, and Overhaul Process

维护、维修和大修流程中可扩展且数据驱动的决策支持
Zhu, Houkun, Ebel, Helena, Scheinert, Dominik, Schmidt, Florian, Altenkirch, Jens, Kao, Odej
Abstract
Several businesses apply maintenance, repair, and overhaul (MRO) principles to the life-cycle of their existing products. In cases like casted gas turbine component Product Lifecycle Management (PLM), repairing components in frequent intervals can extend the lifetime expectation of the product, provide higher cost efficiency compared to newly produced components, and even improve the part design during the repair cycle. Another aspect of repair concerns sustainability, as products often contain rare materials. The emissions produced by the repair process are usually smaller than mining materials and casting new components. To optimize the repair process further, we propose the Smart Expert System (SES), which assists engineering experts with machine learning-based decision support throughout the repair process. We elaborate on its IT architecture and present machine learning models employed for representative MRO use cases. The SES is evaluated using actual industry data from a leading gas turbine company and demonstrably fulfills formulated requirements concerning the suitability of the overall decision support and the stability of the enclosing IT architecture.
Chinese Translation
多家企业将维护、维修和大修(MRO)原则应用于其现有产品的生命周期。在诸如铸造燃气轮机部件产品生命周期管理(PLM)等情况下,以较高频率对部件进行维修可以延长产品的预期寿命,与全新生产的部件相比具有更高的成本效益,甚至可以在维修周期内改进零件设计。维修的另一个方面涉及可持续性,因为产品通常含有稀有材料。维修过程产生的排放通常小于开采原材料并铸造新部件所产生的排放。为进一步优化维修流程,我们提出了智能专家系统(SES),它通过基于机器学习的决策支持,在整个维修流程中协助工程专家。我们详细阐述其 IT 架构,并展示用于代表性 MRO 用例的机器学习模型。SES 使用来自一家领先燃气轮机公司的实际行业数据进行评估,并明确满足关于整体决策支持适用性以及所依托 IT 架构稳定性的既定要求。
cs.AI / 236 / 2609.33646

Probe to Act: Elevating Browser-Use Agent via Active Visual Probing

Probe to Act: 通过主动视觉探测提升浏览器使用智能体
Li, Keliang, Wang, Heng, Hu, Chen, Jiang, Daxin, Chang, Hong, Shan, Shiguang
Abstract
Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks overlays, leaving the model to resolve dense DOM-pixel alignment before every operation. We introduce Probe to Act (P2A), an active probing framework for the browser-agent loop that moves this alignment into decision time. P2A addresses an asymmetric bridge between symbolic DOM hypotheses and screenshot layout by rendering on-demand symbolic DOM structure back into pixels. Before committing a state-changing browser operation, the agent can issue lightweight probes to translate DOM handles into pixel evidence, map screen regions back to DOM candidates, register visual-only targets, and commit verified notes. These interleaved processes naturally produce evidence-based memory: only probed, acted-on, or explicitly committed observations are kept across steps, preserving only decision-critical evidence in long-horizon contexts. P2A can be used as a prompting strategy for proprietary models under the standard DOM+SoM interface, and can be distilled into open-weight models through cold-start synthesis and self-bootstrapped SFT. Across three browser-use benchmarks, P2A shows clear gains on task success rate for both proprietary and fine-tuned models; on VisualWebArena, for example, it improves Gemini-3-Pro from 54.1% to 61.2% and Qwen3-VL-8B from 24.6% to 32.9%, while matching the costly full-observation history ($\sim$3$\times$) at only $\sim$1.2$\times$ the peak retained input context of action-only history.
Chinese Translation
浏览器使用智能体需要在结构化网页元数据与视觉信息之间实现无缝对齐,同时在长交互过程中保持相关上下文。现有接口通常依赖于截图级别的动作预测或静态的Set-of-Marks覆盖,使得模型在每次操作前需要解决密集的DOM-像素对齐问题。我们提出了Probe to Act (P2A),一个用于浏览器智能体循环的主动探测框架,将这种对齐移至决策时刻。P2A通过按需将符号DOM结构渲染回像素,解决了符号DOM假设与截图布局之间的非对称桥梁问题。在提交改变状态的浏览器操作之前,智能体可以发出轻量级探测,将DOM句柄转换为像素证据,将屏幕区域映射回DOM候选,注册仅视觉目标,并提交已验证的注释。这些交错的过程自然产生基于证据的记忆:只有被探测、操作或明确提交的观察结果会跨步骤保留,在长时程上下文中仅保留决策关键证据。P2A可作为标准DOM+SoM接口下专有模型的提示策略,并可通过冷启动合成和自我引导的SFT蒸馏到开放权重模型中。在三个浏览器使用基准上,P2A在专有和微调模型的任务成功率上均显示出明显提升;例如,在VisualWebArena上,它将Gemini-3-Pro从54.1%提升至61.2%,将Qwen3-VL-8B从24.6%提升至32.9%,同时仅以约1.2倍的仅动作历史的峰值保留输入上下文,即可匹配昂贵的完整观察历史(约3倍)。
cs.AI / 237 / 2609.33658

AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents

AgentBoundary:工具使用型LLM智能体安全性的反事实评估
Yang, Tianzhuo, Mi, Zirui, Huang, Yantao, Zhang, Guoxi, Chen, Jiawei, Yang, Yaodong, Yi, Jingwei
Abstract
Safety alignment for large language models (LLMs) in conversational settings is largely framed around whether to answer or refuse a request. In agentic settings, however, the same models must decide whether to act as permission-critical evidence emerges during execution. This creates a distinct challenge: apparent risk, action permissibility, and task competence are easily confounded, making agentic over-refusal difficult to distinguish from ordinary task failure. To address this, we introduce AgentBound, the first four-way counterfactual generation-and-evaluation framework for tool-using agent safety. AgentBound transforms the same executable workflow by independently varying apparent risk and action permissibility, enabling controlled comparisons of risky-looking but authorized tasks and routine-looking but unauthorized tasks. These comparisons jointly diagnose over-refusal and unsafe compliance while controlling for task competence. We instantiate AgentBound as a human-validated 4,000-task evaluation suite with trajectory-based and post-state-based judgments. Across 17 model and harness configurations, high safety frequently coexists with poor authorized-task completion: GPT-5.5 blocks 99.5\% of routine-looking unauthorized actions yet completes only 28.7\% of risky-looking authorized tasks. We further train a lightweight runtime calibration module that improves authorized-task completion by 18.2\% on average across 10 evaluated configurations, while improving unsafe-action blocking by 5.4\% on average. These show that effective agentic alignment requires action decisions to track permission-relevant execution evidence, rather than refusal strength alone.
Chinese Translation
在对话场景中,大语言模型(LLM)的安全对齐主要围绕是否回答或拒绝请求来构建。然而,在智能体场景中,同样的模型必须在执行过程中出现与权限相关的关键证据时决定是否行动。这产生了一个独特的挑战:表面风险、行动许可性和任务能力很容易混淆,使得智能体的过度拒绝难以与普通任务失败区分。为了解决这个问题,我们引入了AgentBound,这是第一个针对使用工具的智能体安全的四路反事实生成与评估框架。AgentBound通过独立改变表面风险和行动许可性来转换相同的可执行工作流,从而能够对看似有风险但经过授权的任务和看似常规但未经授权的任务进行受控比较。这些比较在控制任务能力的同时,共同诊断过度拒绝和不安全合规。我们将AgentBound实例化为一个经过人工验证的4,000个任务的评估套件,具有基于轨迹和基于后状态判断的评估。在17个模型和测试框架配置中,高安全性常常与授权任务完成度差并存:GPT-5.5阻止了99.5%的看似常规但未经授权的行动,但仅完成了28.7%的看似有风险但经过授权的任务。我们进一步训练了一个轻量级运行时校准模块,在10个评估配置中平均将授权任务完成度提高了18.2%,同时平均将不安全行动的阻止率提高了5.4%。这些表明,有效的智能体对齐要求行动决策跟踪与权限相关的执行证据,而不仅仅是拒绝强度。
cs.AI / 238 / 2609.33662

Audit-First VAPO: Risk-Certified Selective Updates under Imperfect Verification

Audit-First VAPO: 不完美验证下的风险认证选择性更新
Hu, Miaobo, Hu, Shuhao, Guo, Xiaobo, Wang, Xin, Wang, Bokun, Chen, Rui, Zha, Daren, Xiao, Jun
Abstract
Imperfect verifiers can assign a harmful update direction even when clipping and regularization bound its magnitude. We introduce Audit-First VAPO, which separates discrete directional admission from continuous magnitude control. An observation-only accept-appeal-abstain policy uses a finite secondary-verification budget; its action trace is frozen before clean labels are joined. Simultaneous finite-sample bounds then certify selected harmful risk, coverage, and verifier-call rate over a predeclared policy family. Conditional Hoeffding-Azuma bounds account for the dependence induced by shared budgets, and rollout or verifier changes initiate a new certification stage. After admission, a bounded trust-clip-KL actuator controls magnitude. We evaluate two models on two reasoning benchmarks against static RLVR, matched-random selection, confidence thresholding, noise correction, and verifier augmentation. On Qwen3.5-0.8B and GSM8K at target risk $\rho=0.08$, RC-VAPO achieves 74.1% accuracy, selected harmful risk 0.0697, coverage 0.4125, and relative verifier cost $1.16\times$. At matched coverage and update magnitude, its selected-risk difference from matched random is -0.0260 with paired 95% interval $[-0.0364,-0.0157]$. Across asymmetric, confidence-dependent, and correlated-verifier noise, the certificate is satisfied on 57 of 60 independent runs. These comparisons isolate informative directional selection from proposal suppression, update shrinkage, and additional verifier computation.
Chinese Translation
不完美验证器即使在其幅度被裁剪和正则化限制时,也可能指定有害的更新方向。我们引入 Audit-First VAPO,它将离散的方向准入与连续的幅度控制分开。一种仅观察的接受-申诉-弃权策略使用有限的二次验证预算;其动作轨迹在干净标签被加入之前被冻结。然后,同时有限样本界在预先声明的策略族上认证选中的有害风险、覆盖率和验证器调用率。条件 Hoeffding-Azuma 界考虑了由共享预算引起的依赖性,而 rollout 或验证器变化会启动新的认证阶段。在准入之后,一个有界的 trust-clip-KL 执行器控制幅度。我们在两个推理基准上评估两个模型,对比静态 RLVR、匹配随机选择、置信度阈值化、噪声校正和验证器增强。在 Qwen3.5-0.8B 和 GSM8K 上,目标风险 ρ=0.08 时,RC-VAPO 达到 74.1% 的准确率,选中有害风险为 0.0697,覆盖率为 0.4125,相对验证器成本为 1.16×。在匹配覆盖率和更新幅度下,其与匹配随机的选中风险差为 -0.0260,配对 95% 区间为 [-0.0364,-0.0157]。在非对称、置信度依赖和相关验证器噪声下,证书在 60 次独立运行中有 57 次得到满足。这些比较将信息性方向选择与提议抑制、更新收缩和额外验证器计算区分开来。
cs.AI / 239 / 2609.33665

CompoWorld: Compositional Environment Scaling for General Agents

CompoWorld:面向通用智能体的组合式环境扩展
Yang, Xiao-Wen, Xu, Weiyi, Da, Wen, Xu, Hang, Li, Canwei, You, Hong-Jie, Dong, Pusen, Zeng, Yucheng, Luo, Zhaokai, Li, Yu-Feng, Hu, Yao, Chuan, Mu
Abstract
Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional Environment Scaling (\textbf{CompoWorld}), which expands the task space by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. A random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. Verified trajectories support supervised fine-tuning (SFT), while our Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B. Experimental results show that CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.
Chinese Translation
自动生成的环境为训练通用智能体提供了可扩展的交互数据来源。然而,现有方法主要在单一环境中生成任务,而现实世界的工作流要求智能体跨多个服务连接信息和动作。我们引入组合式环境扩展(CompoWorld),通过组合一个有限的可重用服务库来扩展任务空间。编码智能体将工具规范转化为具有类型化状态和共享接口的经过验证的服务,而世界模型则处理无法可靠实现的工具。随机游走过程通过依赖图连接服务,使得能够生成和验证需要信息跨服务流动的任务。经过验证的轨迹支持监督微调(SFT),而我们的以完成为中心的评分标准奖励(Completion-Focused Rubric Reward)通过强调每次 rollout 组中通过率较低的标准,引导强化学习(RL)实现完整任务。我们构建了 448 个服务,暴露了 10,130 个工具,并使用 3K SFT 轨迹和 1K RL 任务来训练 Qwen3.6-35B-A3B。实验结果表明,CompoWorld 在八个基准上平均比其骨干模型提高了 9.17 分。在 AutomationBench 上,它超越了 Claude Opus 4.6 等前沿模型,并在所有比较的专用 35B-A3B 智能体模型中领先。
cs.AI / 240 / 2609.33669

RSD-Poker: Structure-Adaptive and Shift-Robust Risk-Utility Certification for Residual Policies in Imperfect-Information Games

RSD-Poker:针对不完全信息博弈中残差策略的结构自适应与偏移鲁棒的风险-效用认证
Hu, Miaobo, Hu, Shuhao, Guo, Xiaobo, Wang, Xin, Wang, Bokun, Zhang, Peng, Zha, Daren, Xiao, Jun
Abstract
Residual policy adaptation provides a lightweight way to modify a strong reference policy, but a shared scale and a fixed subgroup partition can hide heterogeneous degradation and become fragile when the deployment mixture of information states changes. We introduce RSD-Poker, a structure-adaptive and shift-robust certification framework that freezes a bank of residual families and scales, learns a policy-visible partition on an independent structure split, and freezes that partition before calibration labels are joined. Each candidate-group pair receives a weighted simultaneous upper certificate for anchor-relative risk and a lower certificate for weak-response utility. A robust group-to-candidate map is then selected over a predeclared uncertainty set of deployment group proportions. Under independent calibration units drawn from each frozen group's law, a candidate bank and partition fixed before calibration, and invariant within-group conditionals, the selected map satisfies its declared mixture-robust risk budget and utility certificate with probability at least $1-\zeta_{risk}-\zeta_{util}$. The information contract supports both a teacher-backed transform and a teacher-free observation-only student. The retained deterministic 24-state audit remains an exact replay diagnostic: empirical-zero selects $\alpha=0.08$, raising the weak-response proxy from 4.2082 to 4.2889 with $0/12$ held-out threshold crossings. On stratified held-out states, the learned-partition dual selector raises weak utility from 4.4074 under global dual certification to 4.4936 and lowers held-out violation from 0.0215 to 0.0078; its mixture-robust variant reaches violation 0.0059. Across five observation-only checkpoints, risk-calibrated residuals attain weak utility $4.3659\pm0.0177$ and violation rate $0.0178\pm0.0057$.
Chinese Translation
残差策略自适应提供了一种轻量级的方法来修改强参考策略,但是共享的尺度和固定的子群划分可能掩盖异质性退化,并且当信息状态的部署混合发生变化时变得脆弱。我们介绍了 RSD-Poker,一个结构自适应且偏移鲁棒的认证框架,它冻结一组残差族和尺度,在独立的结构划分上学习策略可见的划分,并在校准标签加入之前冻结该划分。每个候选-组对接收一个加权的同时上界证书用于锚点相对风险,以及一个下界证书用于弱响应效用。然后在预先声明的部署组比例的不确定性集上选择一个鲁棒的组到候选映射。在从每个冻结组的定律中抽取的独立校准单元、在校准之前固定的候选库和划分、以及不变的组内条件分布下,所选映射以至少 $1-\zeta_{risk}-\zeta_{util}$ 的概率满足其声明的混合鲁棒风险预算和效用证书。信息契约支持教师支持的变换和无教师的仅观察学生。保留的确定性 24 状态审计仍然是一个精确的重放诊断:经验零选择 $\alpha=0.08$,将弱响应代理从 4.2082 提高到 4.2889,并且有 $0/12$ 的留出阈值跨越。在分层留出状态上,学习划分对偶选择器将弱效用从全局对偶认证下的 4.4074 提高到 4.4936,并将留出违反从 0.0215 降低到 0.0078;其混合鲁棒变体达到违反 0.0059。在五个仅观察检查点上,风险校准残差达到弱效用 $4.3659\pm0.0177$ 和违反率 $0.0178\pm0.0057$。
cs.AI / 241 / 2609.33676

Auditing Agent Actions through Query-Conditioned Attribution

通过查询条件归因审计智能体动作
Liu, Yifan, Venkateswaran, Praveen, Adebayo, Abdulhamid, Wang, Dong
Abstract
LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, existing attribution formulations do not provide question-specific traces for diverse auditing objectives. Additionally, when access to the acting model is limited (e.g., in API-only deployments), applicable methods commonly rely on costly input perturbations or external LLM analysis of complete trajectories. We therefore formulate $\textit{query-conditioned agent action attribution}, a new task that takes a natural-language auditing query as input and recovers the source and ordered intermediate evidence for the query-specified aspect of an action. We instantiate this task with $A^3Bench$, a benchmark comprising 1,396 auditing queries across policy basis, parameter provenance, failure propagation, and unsafe-behavior tracing. To enable efficient, query-specific attribution, we use small open-weight models as attribution proposers that combine query-conditioned gradient saliency with query-semantic relevance to rank history units. Our proposer consistently achieves stronger source and evidence rankings at lower inference cost than open-weight baselines, improving source MRR by up to 40.9\% and evidence MAP by 42.1\% with only two forward passes and one backward pass. Controlled evaluations confirm that our proposer improves attribution specificity by adapting its rankings to fine-grained changes in the auditing query. Building on a proposer ensemble, our end-to-end system surpasses the strongest frontier-model baseline in source accuracy (64.5\% vs.\ 60.4\%) while reducing empirical deployment latency by 29.9\% relative to the fastest frontier API baseline. Code and data will be released after the initial review period following final validation and cleanup.
Chinese Translation
LLM智能体越来越多地通过与用户、策略和外部工具的交互采取产生后果的行动。审计这些智能体需要将已实现的行动自动归因到其历史依据。然而,现有的归因方法不能为多样化的审计目标提供针对具体问题的轨迹。此外,当对执行模型的访问受限时(例如在仅API部署中),适用方法通常依赖于昂贵的输入扰动或外部LLM对完整轨迹的分析。因此,我们形式化了查询条件智能体动作归因,这是一个新任务,它以自然语言审计查询为输入,并还原针对动作的查询指定方面的来源和有序中间证据。我们使用A³Bench实例化了这一任务,该基准包含1,396个审计查询,涵盖策略依据、参数来源、故障传播和不安全行为追踪。为了实现高效、查询特定的归因,我们使用小型开放权重模型作为归因提议器,将查询条件梯度显著性与查询语义相关性相结合,对历史单元进行排序。我们的提议器在更低的推理成本下,始终比开放权重基线获得更强的来源和证据排序,仅需两次前向传播和一次反向传播,就将来源MRR提高了高达40.9%,证据MAP提高了42.1%。受控评估证实,我们的提议器通过使其排序适应审计查询的细粒度变化,提高了归因特异性。基于提议器集成,我们的端到端系统在来源准确率上超越了最强的前沿模型基线(64.5%对60.4%),同时相对于最快的前沿API基线,将实际部署延迟降低了29.9%。代码和数据将在最终验证和清理后的初始评审期后发布。
cs.AI / 242 / 2609.33678

SWE-Game: Can Coding Agents Build the Games We Want?

SWE-Game:编码智能体能否构建我们想要的游戏?
Chen, Xiaoyu, Wei, Lai, Wang, Jin, Zou, Xiangyu, Fan, Ruochen, Luo, Enze, Yao, Mingzhe, Zhu, Jiahui, Wen, Yuhua, Kong, Linghe, Huang, Weiran
Abstract
We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Reference materials specify the intended gameplay, while a shared instrumentation interface lets evaluator-owned drivers and probes execute actions and observe independently implemented games. Evaluation combines engine-state checks, certified reference-input replay, and agent-authored feature demonstrations to assess mechanic correctness, demonstrated playability, and behavioral restoration and preservation after repairs. Game-specific vision-language rubrics separately assess presentation. Across six models, Opus5 achieves the highest overall score in all five task types. Best overall scores remain below 60 out of 100 across the three construction tasks, with Brief-to-Game reaching 50.38. Analysis of reviewed submissions identifies requirement omissions and gameplay logic errors as predominant implementation problems. On human-labeled behaviors from 100 agent-built games, executable checks achieve 92.59% balanced accuracy, compared with 78.41% for a video-based VLM judge. Rubric-based visual scores reach a Spearman correlation of 0.829 with human ratings of 200 gameplay clips. Together, these results characterize current agent capabilities across game-development activities and support combining runtime evidence with visual assessment.
Chinese Translation
我们提出了 SWE-Game,一个包含 247 项任务的基准,这些任务基于 41 个可执行的参考 Godot 游戏,涵盖 2D 和 3D 中的 13 个玩法类别。五种任务类型涵盖:根据简要说明进行开发、根据游戏设计文档进行实现、骨架补全、修复 83 个注入故障案例,以及 Godot 到 Unity 的移植。参考材料规定了预期的玩法,而共享的插桩接口使评估方拥有的驱动器和探针能够执行动作,并观察独立实现的游戏。评估结合了引擎状态检查、经认证的参考输入回放,以及由智能体编写的功能演示,以评估机制正确性、所展示的可玩性,以及修复后的行为恢复与保持。针对具体游戏的视觉语言评分标准单独评估呈现效果。在六个模型中,Opus5 在全部五种任务类型中取得最高总体得分。在三个构建任务中,最佳总体得分仍低于 100 分制的 60 分,其中 Brief-to-Game 达到 50.38。对已评审提交的分析识别出,需求遗漏和玩法逻辑错误是主要的实现问题。在来自 100 个智能体构建游戏的人工标注行为上,可执行检查达到 92.59% 的平衡准确率,而基于视频的 VLM 评判器为 78.41%。基于评分标准的视觉得分与 200 个游戏片段的人工评分之间的 Spearman 相关系数达到 0.829。总之,这些结果刻画了当前智能体在各类游戏开发活动中的能力,并支持将运行时证据与视觉评估相结合。
cs.AI / 243 / 2609.33688

TopoMamba: A Load-Support Relation-Guided Multi-Directional State-Space Model for Topology Optimization

TopoMamba:一种用于拓扑优化的载荷-支撑关系引导多方向状态空间模型
Lou, Bin, Cheng, Yuxuan, Zong, Huaizhi, Zhang, Junhui, Xu, Bing
Abstract
Deep learning has emerged as an efficient alternative for predicting high-performance material distributions in topology optimization. Existing methods struggle to accurately capture load-transfer information, limiting out-of-distribution generalization, while their model architectures often incur high computational costs. To address these challenges, this paper proposes TopoMamba, a topology prediction framework incorporating a load-support relation-guided multi-directional state-space model. Coupling physical fields with load-support relations enables more effective modeling of mechanical dependencies. A load-support relation-guided spatially adaptive fusion mechanism dynamically adjusts multi-directional scan features according to spatial conditions. Mamba is coupled with the solid isotropic material with penalty method to enhance structural mechanical performance while maintaining computational efficiency. Results on two-dimensional topology optimization benchmarks demonstrate that TopoMamba achieves superior topology prediction accuracy, out-of-distribution generalization, and computational efficiency over state-of-the-art models. The proposed load-support physics-guided framework enables efficient optimization of more complex structural systems.
Chinese Translation
深度学习已成为预测拓扑优化中高性能材料分布的一种高效替代方法。现有方法难以准确捕捉载荷传递信息,限制了分布外泛化能力,同时其模型架构通常带来高计算成本。为了解决这些挑战,本文提出了TopoMamba,一种结合了载荷-支撑关系引导多方向状态空间模型的拓扑预测框架。将物理场与载荷-支撑关系耦合,能够更有效地建模力学依赖性。一种载荷-支撑关系引导的空间自适应融合机制,根据空间条件动态调整多方向扫描特征。将Mamba与固体各向同性材料惩罚(SIMP)方法耦合,以增强结构力学性能,同时保持计算效率。在二维拓扑优化基准上的结果表明,TopoMamba在拓扑预测精度、分布外泛化和计算效率方面均优于最先进的模型。所提出的载荷-支撑物理引导框架能够高效优化更复杂的结构系统。
cs.AI / 244 / 2609.33698

One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion

一个潜在表示,多个词元:联合学习压缩嵌入以实现高效语言扩散
Yuan, Yulin, Zhang, Ying, Meng, Xiangming
Abstract
Most continuous diffusion language models process one latent position per token at each sampling step, making generation expensive. Two-stage methods lower the cost by reducing the latent length, but they fix the compressed embedding space before training the diffusion model. Embeddings from the fixed space can be difficult to model with diffusion and decode reliably into tokens, which limits generation quality after compression. To address this problem, we introduce JPEG-DLM (Joint-embedding Prediction for Efficient Generation with Diffusion Language Model), which jointly trains a compressor, a flow matching model and a decoding module. With joint-embedding prediction, JPEG-DLM learns compressed embeddings that are more structured, easier to model with diffusion and reliably decodable into tokens. JPEG-DLM achieves the lowest mean Gen-PPL and highest throughput among recent diffusion and flow models on LM1B and OWT. At a compression rate of 0.5 on OWT, it reaches a Gen-PPL of 34.52 and approximately 2.3 times ELF's throughput. These results suggest that jointly learning compressed embeddings offers a promising path toward efficient diffusion language modeling. Code will be released soon.
Chinese Translation
大多数连续扩散语言模型在每个采样步骤中为每个词元处理一个潜在位置,导致生成成本高昂。两阶段方法通过减少潜在长度来降低成本,但它们在训练扩散模型之前就固定了压缩嵌入空间。来自固定空间的嵌入可能难以用扩散建模并可靠地解码为词元,这限制了压缩后的生成质量。为了解决这个问题,我们引入了 JPEG-DLM(Joint-embedding Prediction for Efficient Generation with Diffusion Language Model),它联合训练压缩器、流匹配模型和解码模块。通过联合嵌入预测,JPEG-DLM 学习到更结构化、更易于用扩散建模且能可靠解码为词元的压缩嵌入。在 LM1B 和 OWT 上,JPEG-DLM 在最近的扩散和流模型中实现了最低的平均 Gen-PPL 和最高的吞吐量。在 OWT 上压缩率为 0.5 时,它达到了 34.52 的 Gen-PPL 和约 2.3 倍于 ELF 的吞吐量。这些结果表明,联合学习压缩嵌入为高效扩散语言建模提供了一条有前景的路径。代码即将发布。
cs.AI / 245 / 2609.33699

SpecRead: A Benchmark for Measuring Whether Language Models Understand Hardware Specifications

SpecRead:一个衡量语言模型是否理解硬件规范的基准
Huang, Feilian
Abstract
Existing benchmarks for large language models (LLMs) in hardware design evaluate downstream artifacts such as generated RTL, assertions, or testbenches. When a model fails such a benchmark, the failure is ambiguous: it may have misread the specification, or it may have understood the specification and failed to write the code. We present SpecRead, a benchmark that isolates specification comprehension from generation ability. SpecRead v2.1 contains 385 questions over 10 open-source OpenTitan IP blocks: exact retrieval, cross-section reasoning, contradiction detection in mutated specifications, and spec-RTL consistency checking, plus 82 controls (41 distractor, 41 consistent-RTL). Type-4 items are built from real RTL mutations; we retain only mutations that Icarus Verilog simulation shows to change observable behavior. A with-spec vs. without-spec ablation suggests the questions require the excerpt, not training recall alone (without-spec accuracy 3/20 on the t1/t2 subset), though memorization of the source text may still help spot mutations. As an initial characterization with a small model, Ministral-3B scores 33.2% overall (128/385; macro average 39.0%): 55.2% on retrieval, 51.7% on cross-section reasoning. On the two contradiction-focused types, the verdict-plus-location measure gives 48.0% (t3) and 63.3% (t4), with a 51.2% false-positive rate on distractors and 100% on consistent-RTL controls. Layered scoring shows the model locates contradictions well (78.9-81.6% location accuracy) but scores lower on their category (43.9-49.7%). A structured "rule-table" prompting intervention lowers accuracy on every question type except t2 (tied). SpecRead is automatically scorable by deterministic checks, with gray-zone cases counted wrong under the conservative main scoring. The benchmark is regenerable for type-3 items via mutation injection, and built exclusively from public sources.
Chinese Translation
现有针对硬件设计领域大语言模型(LLMs)的基准评估的是下游产物,如生成的 RTL、断言或测试平台。当模型在这样的基准上失败时,失败原因是模糊的:它可能误读了规范,或者理解了规范但未能编写出代码。我们提出 SpecRead,一个将规范理解与生成能力分离开来的基准。SpecRead v2.1 包含覆盖 10 个开源 OpenTitan IP 模块的 385 个问题:精确检索、跨节推理、变异规范中的矛盾检测,以及规范-RTL 一致性检查,另有 82 个对照项(41 个干扰项,41 个一致-RTL)。类型 4 的项目基于真实 RTL 变异构建;我们仅保留 Icarus Verilog 仿真显示会改变可观察行为的变异。有规范 vs. 无规范的消融实验表明,这些问题需要提供的摘录,而不仅仅是训练时的记忆(在 t1/t2 子集上,无规范准确率为 3/20),尽管对源文本的记忆可能仍然有助于发现变异。作为使用小模型的初步表征,Ministral-3B 总体得分为 33.2%(128/385;宏平均 39.0%):检索为 55.2%,跨节推理为 51.7%。在两个以矛盾为重点的类型上,裁决加位置度量给出 48.0%(t3)和 63.3%(t4),在干扰项上的假阳性率为 51.2%,在一致-RTL 对照上为 100%。分层评分显示模型定位矛盾的能力较好(位置准确率 78.9-81.6%),但在其类别上的得分较低(43.9-49.7%)。结构化的“规则表”提示干预降低了除 t2 外每种问题类型的准确率(t2 持平)。SpecRead 可通过确定性检查自动评分,在保守的主评分下,灰色地带案例被计为错误。该基准可通过变异注入为类型 3 项目重新生成,并且完全基于公开来源构建。
cs.AI / 246 / 2609.33707

Does Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View Collapse

对抗训练能否提升多视角VLA的泛化能力?揭示并缓解视角坍塌
Waseda, Futa, Kurita, Shuhei, Echizen, Isao
Abstract
Vision-language-action (VLA) models adapt pretrained vision-language models (VLMs) for closed-loop robot control, transferring their perceptual and semantic capabilities to action prediction. Despite strong in-distribution performance, however, VLAs often degrade under deployment shifts. Adversarial training (AT) offers a model-adaptive approach to robustness without explicitly anticipating individual shifts, but its effect on natural distribution-shift generalization in multi-view VLAs remains unclear. We study this question using a multi-view VLA directly adapted from a pretrained VLM and evaluate generalization across seven LIBERO-Plus shift axes. Direct AT substantially improves Camera Viewpoint and Sensor Noise, the two shifts affecting only the third-person view, yet produces mixed or negative effects on other shifts. Controlled view interventions reveal a surprising failure mode that we term view collapse: Direct AT can shift cross-view reliance so strongly that the policy becomes dominated by the wrist view. This exposes a \textit{robustness shortcut}: apparent robustness to a shifted view can arise from reduced use of that view rather than more robust perception of it. This motivates a distinction between robust perception, extracting reliable information under within-view shifts, and robust fusion, adapting reliance across views according to their reliability. To reduce fixed view reliance, we use a simple View Swap intervention and then re-evaluate AT. With View Swap, AT further improves Camera Viewpoint, Sensor Noise, and Robot Initial State, while its effects remain mixed on other shifts. Our results show that multi-view robustness requires separating improved perception from changes in cross-view reliance, and that AT provides selective rather than generic distribution-shift benefits.
Chinese Translation
视觉-语言-动作(VLA)模型将预训练的视觉-语言模型(VLM)适配于闭环机器人控制,将其感知和语义能力迁移到动作预测。然而,尽管在分布内性能强劲,VLA在部署偏移下常常性能下降。对抗训练(AT)提供了一种模型自适应的鲁棒性方法,无需显式预测个体偏移,但其对多视角VLA中自然分布偏移泛化的影响仍不清楚。我们使用一个直接从预训练VLM适配的多视角VLA来研究这个问题,并在七个LIBERO-Plus偏移轴上评估泛化能力。直接AT显著提升了摄像机视角和传感器噪声(这两个偏移仅影响第三人称视角)的性能,但对其他偏移产生了混合或负面影响。受控视角干预揭示了一种令人惊讶的失效模式,我们称之为视角坍塌:直接AT可能强烈改变跨视角依赖,以至于策略被手腕视角主导。这暴露了一种鲁棒性捷径:对偏移视角的表观鲁棒性可能源于对该视角使用的减少,而非对其更鲁棒的感知。这促使我们区分鲁棒感知(在视角内偏移下提取可靠信息)和鲁棒融合(根据视角的可靠性调整跨视角依赖)。为了减少固定的视角依赖,我们使用了一种简单的视角交换(View Swap)干预,然后重新评估AT。在视角交换下,AT进一步提升了摄像机视角、传感器噪声和机器人初始状态的性能,而对其他偏移的影响仍然混合。我们的结果表明,多视角鲁棒性需要将改进的感知与跨视角依赖的变化分离开来,并且AT提供的是选择性的而非通用的分布偏移优势。
cs.AI / 247 / 2609.33713

BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation

BIRD:将决策边界蒸馏为理由以实现MLLM适配
Liu, Anglin, Wu, Yanlin, Chen, Ruichao, Zhang, Yuting, Zeng, Qingyuan, Cai, Pengxiang, Gong, Ziqi, Li, Muchen, Chen, Jintai
Abstract
Adapting general-purpose multimodal large language models (MLLMs) to specialized domains requires learning domain-specific decision criteria, which often hinge on subtle visual distinctions between otherwise plausible answers. Rationale augmentation aims to expose such evidence through additional observations or inter-sample comparisons, yet a visually valid cue is not necessarily decision-relevant: it may describe how samples differ without changing the model's relative preference between competing answers. We therefore introduce BIRD, a self-improving Boundary-Informed Rationale Distillation framework that uses model-specific confusions to locate unresolved local decision boundaries and distills the evidence that resolves these confusions into rationales. For each sample, BIRD retrieves candidate neighbors from the target MLLM's own representation space and selects the most confusable one according to its answer preferences. It then generates answer-blind candidate evidence from their visual differences and functionally verifies which evidence most effectively strengthens the model's preference for the correct answer while avoiding inappropriate transfer across the pair. The verified evidence is then distilled into a single-sample rationale for standard supervised fine-tuning. Experiments on medical and chart VQA show that BIRD outperforms competing rationale-augmentation methods across two target MLLMs, while further analyses demonstrate clearer separation of confusable answers and stronger gains from model-matched supervision.
Chinese Translation
将通用多模态大语言模型(MLLMs)适配到专业领域需要学习领域特定的决策标准,这些标准通常取决于原本看似合理的答案之间微妙的视觉差异。理由增强旨在通过额外的观察或样本间比较来揭示此类证据,但视觉上有效的线索未必与决策相关:它可能描述了样本如何不同,却没有改变模型在竞争答案之间的相对偏好。因此,我们提出了BIRD,一个自改进的边界感知理由蒸馏框架,它利用模型特有的混淆来定位未解决的局部决策边界,并将解决这些混淆的证据蒸馏为理由。对于每个样本,BIRD从目标MLLM自身的表示空间中检索候选邻居,并根据其答案偏好选择最易混淆的一个。然后,它从它们的视觉差异中生成与答案无关的候选证据,并功能性验证哪些证据最有效地增强模型对正确答案的偏好,同时避免在样本对之间进行不恰当的迁移。经验证的证据随后被蒸馏为单样本理由,用于标准监督微调。在医学和图表VQA上的实验表明,BIRD在两个目标MLLM上优于竞争的理由增强方法,进一步的分析表明,可混淆答案的分离更清晰,并且来自模型匹配监督的增益更强。
cs.AI / 248 / 2609.33717

Self-Designed Evaluators and Warm Memory for Long-Horizon Agents

用于长时程智能体的自行设计评估器与预热记忆
Asgari, Saeid, Kiciman, Emre, Nunes, Leonardo de Oliveira, Chandra, Ranveer
Abstract
A tool-using language-model agent deployed over a long stream of tasks receives no reward, so it cannot tell whether it succeeded, cannot safely retry, and cannot label the experience it needs to improve. We present SelfSuite, in which the agent's own base model, given only the world's public materials, designs a small evaluation suite of weighted judges and grounded per-task briefs, freezes it, and uses it to gate a keep-best retry and to label a typed, outcome-tracked memory. On matched five-repeat benchmarks over tau2-bench and AppWorld, SelfSuite scores above the plain agent without any labels, matches methods given ten expert labels on tau2-bench, and trails Agentic Context Engineering (ACE) on AppWorld, where code execution gives a direct success signal. In an ablation campaign run on the same tasks, it is above label-free ACE in every repeat, and the gated second attempt is the only component whose removal hurts in every repeat. We also simulate a subject-matter expert who grades ten onboarding tasks per world. Using those labels to calibrate SelfSuite's evaluator gives a small, consistent gain, and using them to warm up ACE's memory lifts ACE to tie calibrated SelfSuite. A single-run study on a second model family shows the same ordering.
Chinese Translation
一个使用工具的语言模型智能体在长任务流上部署时不会收到奖励,因此它无法判断是否成功,无法安全地重试,也无法标记其需要改进的经验。我们提出了 SelfSuite,其中智能体自身的基础模型仅给定世界的公开材料,设计一个小型评估套件,包含加权评判器和基于每个任务的简要说明,将其冻结,并用它来门控一个保留最佳的重试,并标记一个类型化、结果跟踪的记忆。在 tau2-bench 和 AppWorld 上匹配的五次重复基准测试中,SelfSuite 在没有任何标签的情况下得分高于普通智能体,在 tau2-bench 上匹配了给定十个专家标签的方法,在 AppWorld 上落后于 Agentic Context Engineering (ACE),其中代码执行给出了直接的成功信号。在相同任务上进行的消融实验中,它在每次重复中都高于无标签的 ACE,并且门控的第二次尝试是唯一一个在每次重复中移除都会造成损害的组件。我们还模拟了一个领域专家,他为每个世界评分十个入门任务。使用这些标签来校准 SelfSuite 的评估器带来了小幅但一致的提升,并使用它们来预热 ACE 的记忆将 ACE 提升到与校准后的 SelfSuite 持平。在第二个模型家族上的单次运行研究显示了相同的排序。
cs.AI / 249 / 2609.33726

Robust Biomolecular Complex Design Across Protein Conformational Landscapes

跨蛋白质构象景观的鲁棒生物分子复合物设计
Zeng, Qingyuan, Xu, Zongqi, Liu, Anglin, Gong, Ziqi, Cai, Pengxiang, Guan, Zixin, Chen, Yunan, Gao, Sen, Zhou, Min, Chen, Jintai
Abstract
Proteins populate conformational ensembles, yet structure-based biomolecular design typically optimizes candidates against a single target conformation. Consequently, a candidate that fits one state can lose favorable interactions or develop steric clashes when the target adopts another. We introduce FlexEvo, a model-agnostic evolutionary framework that adapts candidates once at inference time from a single target conformation to improve compatibility with alternative natural conformations unseen during adaptation, without retraining the source model or requiring a conformational ensemble. FlexEvo casts cross-state adaptation as geometry-constrained bi-objective optimization, balancing preservation of input-state interactions against robustness to plausible conformational perturbations. To limit the search space and reduce invalid structural edits, geometry-derived FlexBoxes define protected anchor regions, adaptable regions for local exploration, and forbidden regions for clash avoidance. A unified all-atom representation supports topology-preserving adaptation across diverse binder categories, while Pareto selection preserves nondominated candidates across the two objectives. We evaluate FlexEvo across multiple generation baselines and nine representative binder categories spanning diverse molecular sizes and structural topologies. FlexEvo reduces the category-balanced mean relative performance degradation from 47.8% to 4.4%, while adding only 1.4--3.1 minutes of adaptation per sample. These results establish single-state inference-time adaptation as a practical route toward robust biomolecular complex design across protein conformational landscapes.
Chinese Translation
蛋白质存在于构象系综中,然而基于结构的生物分子设计通常针对单一目标构象优化候选分子。因此,适合一种状态的候选分子在靶标采取另一种构象时,可能会失去有利相互作用或产生空间位阻冲突。我们提出 FlexEvo,一种模型无关的进化框架,它在推理时仅从单一目标构象对候选分子进行一次适应,以提高与适应过程中未见过的其他天然构象的兼容性,而无需重新训练源模型或需要构象系综。FlexEvo 将跨状态适应表述为几何约束的双目标优化,在保持输入状态相互作用与对合理构象扰动的鲁棒性之间进行权衡。为了限制搜索空间并减少无效的结构编辑,几何衍生的 FlexBoxes 定义了保护的锚定区域、用于局部探索的可适应区域以及用于避免冲突的禁区。统一的全原子表示支持跨不同结合剂类别的保持拓扑的适应,而帕累托选择则保留了两个目标上的非支配候选。我们在多个生成基线和九个代表性结合剂类别上评估了 FlexEvo,这些类别涵盖了不同的分子大小和结构拓扑。FlexEvo 将类别平衡的平均相对性能下降从 47.8% 降低到 4.4%,而每个样本仅增加 1.4–3.1 分钟的适应时间。这些结果确立了单状态推理时适应作为一种实用途径,可实现跨蛋白质构象景观的鲁棒生物分子复合物设计。
cs.AI / 250 / 2609.33731

HTN Planning as a Coordination Layer for Multi-Server MCP Tool Orchestration

HTN规划作为多服务器MCP工具编排的协调层
Jacopin, Eliott, Jacopin, Éric, Takahashi, Koichi
Abstract
The Model Context Protocol (MCP) isolates servers by design: only the host can orchestrate cross-server workflows. When the host is a large language model, the resulting orchestrations are non-deterministic, non-reproducible, and pay one inference round-trip per tool call. We present a coordination architecture in which a Hierarchical Task Network (HTN) planner generates a verifiable cross-server plan once, and a runtime middleware executes it deterministically across multiple MCP servers, binding cross-action data dependencies via a template mechanism (\verb|${context.X}|) substituted at execution time. The architecture mirrors MCP's isolation constraint: each compound task decomposes into server-local primitive actions, and inter-server data flow is bound at execution time via JSON-path output extractors. We instantiate the architecture on five HTN domains spanning laboratory robotics, bioinformatics and multiscale modelling, and demonstrate end-to-end execution from a browser-based plan controller against eight live third-party MCP servers querying real biological databases.
Chinese Translation
模型上下文协议(MCP)在设计上隔离了服务器:只有主机能够编排跨服务器工作流。当主机是大语言模型时,由此产生的编排是非确定性的、不可复现的,并且每次工具调用都需要一次推理往返。我们提出了一种协调架构,其中分层任务网络(HTN)规划器一次性生成可验证的跨服务器计划,运行时中间件在多个MCP服务器上确定性地执行该计划,通过在执行时替换的模板机制(${context.X})绑定跨动作数据依赖。该架构反映了MCP的隔离约束:每个复合任务分解为服务器本地的原子动作,服务器间的数据流在执行时通过JSON路径输出提取器绑定。我们在涵盖实验室机器人、生物信息学和多尺度建模的五个HTN领域实例化了该架构,并演示了从基于浏览器的计划控制器对八个实时第三方MCP服务器的端到端执行,查询真实生物数据库。
cs.AI / 251 / 2609.33772

Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents

Skill2Env:面向能力的从技能出发的通用智能体环境合成
Xu, Weiyi, Yang, Xiaowen, Da, Wen, Xu, Hang, Li, Canwei, You, Hongjie, Dong, Pusen, Zeng, Yucheng, Luo, Zhaokai, Chuan, Mu
Abstract
Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provide reusable domain knowledge, operational procedures, and tool-use instructions, but a substantial gap remains between the information contained in a skill and a concrete, challenging task with a complete executable environment. To address this gap, we introduce Skill2Env, a capability-oriented framework that starts from a skill and uses agent capability demands to guide task and environment synthesis. Skill2Env represents these demands through reusable difficulty patterns and instantiates them into task blueprints that specify objectives, challenges, environment facts, information boundaries, and acceptance criteria. These blueprints guide the joint construction of task instructions, execution substrates, workspaces, and rubric-based evaluators around source skills. We further propose Iterative Task Hardening, which uses solver execution evidence to identify insufficiently challenging task designs, strengthen or extend their difficulty-pattern instantiations, and revise the corresponding blueprints and environments. Using 1.5K high-scoring trajectories generated from Skill2Env environments for supervised fine-tuning, we observe consistent improvements across a broad range of agent benchmarks, demonstrating the effectiveness of capability-oriented environment synthesis for agent post-training.
Chinese Translation
可执行环境对于需要在工具使用和多步交互任务上对智能体进行后训练至关重要,但构建可执行任务及其环境仍然难以规模化。技能提供了可复用的领域知识、操作流程和工具使用说明,但技能中包含的信息与具有完整可执行环境的具体、具有挑战性的任务之间仍存在巨大差距。为了解决这一差距,我们引入了 Skill2Env,一个面向能力的框架,它从技能出发,利用智能体能力需求来指导任务和环境合成。Skill2Env 通过可复用的难度模式来表示这些需求,并将它们实例化为任务蓝图,这些蓝图指定了目标、挑战、环境事实、信息边界和验收标准。这些蓝图指导围绕源技能联合构建任务指令、执行基座、工作空间和基于评分标准的评估器。我们进一步提出了迭代任务强化(Iterative Task Hardening),它利用求解器执行证据来识别挑战性不足的任务设计,加强或扩展其难度模式实例化,并修订相应的蓝图和环境。使用从 Skill2Env 环境生成的 1.5K 高分轨迹进行监督微调,我们观察到在广泛的智能体基准测试中取得了一致的改进,证明了面向能力的环境合成在智能体后训练中的有效性。
cs.AI / 252 / 2609.33773

Learning Strategies to Break Judges

学习攻破评判者的策略
Shabadi, Guruprerana, Naik, Aaditya, Alur, Rajeev, Naik, Mayur
Abstract
As AI agents surpass human performance, it becomes exceedingly hard for system designers to evaluate them directly and understand their failure modes. Consequently, agents themselves are being deployed extensively to evaluate, judge, and provide feedback on model traces. But this raises an important question: how can we trust the judge? In this work, we propose an agent-guided method to find weaknesses of agentic judges that expose interpretable failure mechanisms. Our method focuses on mathematical reasoning and proceeds in two stages: first, we deploy adversarial agents to mutate a set of sound proofs by introducing errors, attempting to misguide judges---in other words, injecting errors that judges are unable to catch. Then, we distill these attempts into a small set of mutation strategies which allow us to analyze the failure modes of the judges. To ensure that these strategies are not overfit to the initial set of proofs, we evaluate them by applying the mutation strategies to a held-out set of proofs and querying the same judge. We deploy our method on GPT-5.6-sol and Claude Opus 5, paired with their agent orchestrators, Codex and Claude Code, respectively. These are used both as mutators to introduce errors and as judges to evaluate correctness of mathematical reasoning. We find that across all the agentic judges, we are able to distill mutation strategies that consistently bypass their evaluations, thereby enabling us to ascertain actionable failure modes. Our analysis also reveals that judge reliability degrades at the frontier: errors in Olympiad-level proofs or graduate-level mathematical texts are detected more consistently, whereas flaws in research-level manuscripts are more likely to escape detection.
Chinese Translation
随着AI智能体超越人类表现,系统设计者直接评估它们并理解其失败模式变得极其困难。因此,智能体本身被广泛部署来评估、评判模型轨迹并提供反馈。但这引发了一个重要问题:我们如何信任评判者?在这项工作中,我们提出了一种智能体引导的方法来发现智能体评判者的弱点,从而揭示可解释的失败机制。我们的方法专注于数学推理,分两个阶段进行:首先,我们部署对抗性智能体,通过引入错误来变异一组可靠的证明,试图误导评判者——换句话说,注入评判者无法捕捉的错误。然后,我们将这些尝试提炼为一小组变异策略,使我们能够分析评判者的失败模式。为了确保这些策略不会过拟合于初始证明集,我们通过将变异策略应用于一个保留的证明集并查询相同的评判者来评估它们。我们将我们的方法部署在GPT-5.6-sol和Claude Opus 5上,分别与它们的智能体编排器Codex和Claude Code配对。这些既被用作引入错误的变异器,也被用作评判数学推理正确性的评判者。我们发现,在所有智能体评判者中,我们能够提炼出始终绕过其评估的变异策略,从而使我们能够确定可操作的失败模式。我们的分析还揭示,评判者的可靠性在前沿领域下降:奥林匹克级别的证明或研究生级别的数学文本中的错误被更一致地检测到,而研究级别手稿中的缺陷则更可能逃过检测。
cs.AI / 253 / 2609.33778

Evidence-Inference Reconstruction: When The Evidence Is Recalled But The Reasoning Goes Wrong

证据-推理重构:当证据被召回但推理出错时
Diehl, Megan, Lim, Ser-Nam
Abstract
Modern multi-hop LLM agents are equipped with built-in mechanisms to detect errors in intermediate reasoning steps. Such errors trigger corrective actions from these agents, which mostly follow the paradigm of retrying the steps or the reasoning trajectories. Not only are these retries expensive, we present in this paper that they are also potentially unnecessary. To this end, we introduce Evidence-Inference Reconstruction (EIR), which uses structured state to guide one retrieval trajectory, accumulating source evidence in the process. We show that as long as the relevant evidence has been collected, EIR is capable of generating the correct answer in a single final model call even if erroneous evidence has been mixed in due to incorrect intermediate reasoning steps. In one evaluation, using Haiku 4.5 and GPT-4.1 Mini, we evaluate EIR on matched 1,000-question subsets of HotpotQA, 2WikiMultiHopQA, and MuSiQue, showing that EIR improves Answer F1, the overlap between the model's and the correct answer, over the baseline by 8.3--32.8 points, Agentic SSR by 10.6--29.1 points, and Reflexion by 1.1--15.9 points. Additionally, we show that EIR averages 4.85 total model calls per question, compared with 35.29 for Agentic SSR and 12.41 for Reflexion. Together, these results corroborate EIR's central premise: separating evidence retrieval from the final answer model call can improve answer accuracy while utilizing substantially less computation.
Chinese Translation
现代多跳LLM智能体配备了内置机制,用于检测中间推理步骤中的错误。此类错误会触发这些智能体的纠正动作,而这些动作大多遵循重试步骤或推理轨迹的范式。这些重试不仅代价高昂,我们在本文中还表明它们可能是不必要的。为此,我们引入证据-推理重构(EIR),它使用结构化状态来引导一条检索轨迹,并在过程中累积源证据。我们表明,只要相关证据已被收集,EIR就能够在单次最终模型调用中生成正确答案,即使由于错误的中间推理步骤而混入了错误证据。在一项评估中,我们使用Haiku 4.5和GPT-4.1 Mini,在HotpotQA、2WikiMultiHopQA和MuSiQue的匹配1,000问题子集上评估EIR,结果表明,EIR将Answer F1(模型答案与正确答案的重叠度)相对于基线提高了8.3-32.8分,相对于Agentic SSR提高了10.6-29.1分,相对于Reflexion提高了1.1-15.9分。此外,我们表明EIR平均每个问题总共调用模型4.85次,而Agentic SSR为35.29次,Reflexion为12.41次。总之,这些结果证实了EIR的核心前提:将证据检索与最终答案模型调用分离,可以在提高答案准确性的同时大幅减少计算量。
cs.AI / 254 / 2609.33786

Is your uncertainty map wrong, or is its target? Exact diagnostics for the Tweedie diagonal, and a gradient-free alternative

你的不确定性图是错的,还是其目标错了?Tweedie对角线的精确诊断与一种无梯度替代方法
Ribas, Vicent, Tous, Anna Oliveras
Abstract
A diffusion model can predict a follow-up medical scan from a baseline, but a clinician needs a per-voxel map of where that prediction can be trusted. Many such maps approximate the diagonal of the Tweedie posterior covariance, and are evaluated against another approximation of it, so whether the estimator or the target limits them is unclear. We compute the exact diagonal on six checkpoints across fourteen model-corpus conditions. Hutchinson at M=200 tracks it at rank agreement of at least 0.92 everywhere, yet in four of the fourteen the exact diagonal is anti-correlated with the denoising error, reaching -0.13, so a faithful estimator reproduces that reversal. All four are real-image conditions; on the models' own samples the reversal does not appear, so evaluating on generated samples flatters this family. What limits these maps is the target, not the estimator. We then introduce Tweedie Probe-Tangent (T-PT), a gradient-free residual probe that corrupts one model-supported prediction repeatedly and measures the voxel-wise variance of the denoiser's response. T-PT reads a different functional of the same Jacobian, and its exact second-order form ranks with the diagonal wherever the diagonal reverses; at thirty probes it returns a map too unstable to reproduce that ranking, while Hutchinson at M=5 already reproduces it, so T-PT there is not evidence against the reversal. We offer it as an instrument, not a better approximation. On brain MRI at full resolution, where every Jacobian-based estimator we test runs out of memory, T-PT leads a twenty-chain Monte-Carlo ensemble on five of eight endpoints inside tissue and trails it on none, at 16x fewer network evaluations; over the whole volume the ensemble leads, and fifty chains close the tissue gap. On lung CT the ensemble is ahead throughout. Both lose most of their discrimination where the change is, which remains open.
Chinese Translation
扩散模型可以从基线预测随访医学扫描,但临床医生需要一张逐体素的图,标明该预测在何处可信。许多此类图近似Tweedie后验协方差的对角线,并针对其另一种近似进行评估,因此是估计器还是目标限制了它们尚不清楚。我们在十四个模型-语料库条件下的六个检查点上计算精确对角线。Hutchinson在M=200时处处以至少0.92的秩一致性跟踪它,然而在十四个条件中的四个中,精确对角线与去噪误差负相关,达到-0.13,因此忠实的估计器会重现这种反转。所有四个都是真实图像条件;在模型自身的样本上,反转不出现,因此在生成样本上评估美化了这一族方法。限制这些图的是目标,而非估计器。然后我们引入Tweedie Probe-Tangent (T-PT),一种无梯度残差探针,反复扰动一个模型支持的预测,并测量去噪器响应的体素级方差。T-PT读取同一Jacobian的不同泛函,其精确二阶形式在对角线反转的地方与对角线排名一致;在三十个探针时,它返回的图太不稳定,无法重现该排名,而Hutchinson在M=5时已经重现了它,因此那里的T-PT不是反对反转的证据。我们将其作为一种工具提供,而非更好的近似。在全分辨率脑MRI上,我们测试的每个基于Jacobian的估计器都内存不足,T-PT在组织内的八个端点中的五个上领先于二十链蒙特卡洛集成,且没有落后,网络评估次数少16倍;在整个体积上,集成领先,五十条链缩小了组织差距。在肺部CT上,集成全程领先。两者在变化发生的地方都失去了大部分区分度,这仍然是一个未解决的问题。
cs.AI / 255 / 2609.33816

Dual-Vocabulary Language Model for Cross-Tokenizer Distillation

跨分词器蒸馏的双词表语言模型
Chen, Kedi, Lin, Chen, Sun, Yutao, Zhang, Wei
Abstract
On-policy distillation (OPD) bridges teacher supervision and student behavior, but different teacher-student tokenizers introduce misalignment in both input tokenization (#1) and output logits (#2). Existing approaches address the former by matching same-text spans or converting tokens to bytes, often losing fine-grained token information or disrupting the native-token paradigm, while for the latter, strategies such as ranking, padding, or key-token selection retain only shared logit dimensions, resulting in much distribution loss. In this paper, we propose Dual-Vocabulary Language Model (DVLM), which replaces the teacher's LM head with a new student-vocabulary projection head and obtains full-dimensional student logits (for #2). To support student tokens (for #1), it takes a Parallel-Tokenized Sequence (PTS) as input, which concatenates the original teacher-tokenized sequence and a re-tokenized sequence formed by independently converting each student token into a teacher-token group. To avoid inference inconsistency with the original teacher tokens, the Hybrid-Prefix Attention (HPA) further restricts re-tokenized groups to their corresponding teacher prefix and uses its last state as the aggregation of the original student-token representation for projection into the student vocabulary space. Similarly, via the combined use of PTS and HPA, the DVLM teacher can provide distribution-aligned supervision with the student's input-tokenization and output-logit during OPD. Experimental results demonstrate that our DVLM teacher has a similar converged loss as the original teacher model and enables student models to improve performance across six reasoning tasks.
Chinese Translation
同策略蒸馏(OPD)连接了教师监督与学生行为,但不同的师生分词器在输入分词(#1)和输出 logits(#2)两方面引入错位。现有方法针对前者,通过匹配相同文本片段或将 token 转换为字节来解决,但往往会丢失细粒度的 token 信息或破坏原生 token 范式;而对于后者,排序、填充或关键 token 选择等策略只保留共享的 logit 维度,导致大量分布损失。本文提出双词表语言模型(DVLM),它将教师的 LM head 替换为新的学生词表投影头,并获得全维度的学生 logits(针对 #2)。为了支持学生 token(针对 #1),它采用并行分词序列(PTS)作为输入,该序列拼接了原始教师分词序列和重新分词序列,后者通过将每个学生 token 独立转换为教师 token 组而形成。为避免与原始教师 token 的推理不一致,混合前缀注意力(HPA)进一步将重新分词组限制到其对应的教师前缀,并使用其最后状态作为原始学生 token 表示的聚合,以投影到学生词表空间。类似地,通过 PTS 和 HPA 的结合使用,DVLM 教师能够在 OPD 过程中提供与学生输入分词和输出 logit 分布对齐的监督。实验结果表明,我们的 DVLM 教师具有与原始教师模型相似的收敛损失,并使学生在六个推理任务上提升性能。
cs.AI / 256 / 2609.33822

Vestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and Memory

Vestrum:通过适配智能体框架的验证、结构与记忆来改进它们
Parashar, Jayant, Douglass, Eugene F., Bastian, William C., Bhandarkar, Suchendra M.
Abstract
An agent harness controls how a language model accesses information, uses tools, preserves memory, and checks its work. Improving this software is costly when each evaluation requires a long interaction with an environment. We introduce Vestrum, a framework that turns failures in execution traces into scoped harness changes without training the task model. Its organizing overhypothesis is that tasks of a shared kind may exhibit recurring failures whose remedies transfer within that kind. Vestrum expresses failures as recognizable classes, proposes changes across verification, retrieval, decomposition, and knowledge synthesis, and screens their scope before evaluating them as a bundle. A persistent lessons file informs subsequent proposals. Across five settings and two baseline harnesses, the frozen harnesses improve held-out performance: UltraHorizon rises from 47.6 to 59.8 over GAM, Terminal-Bench 4 Hard from 63.7% to 70.3% of checks passed over Claude Code on eight held-out tasks at 1.03x test cost, and cell-type annotation agreement from 67.5% to 77.8% on held-out sections of one slide, alongside gains on LoCoMo and AMA-Bench. Across our searches, verification grounded in evidence helped both intermediate steps and final answers, at lower cost at intermediate steps, while critics asked to rebuild finished answers broke more than they repaired. On the three memory benchmarks, Vestrum also scores above the evaluated GEPA configurations in every paired evaluation.
Chinese Translation
智能体框架控制语言模型如何访问信息、使用工具、保留记忆以及检查其工作。当每次评估都需要与环境进行长时间交互时,改进该软件的成本很高。我们提出 Vestrum,一个框架,它将执行轨迹中的失败转化为有范围的框架变更,而无需训练任务模型。其组织性元假设是:同一类型的任务可能表现出反复出现的失败,而这些失败的补救措施可在该类型内迁移。Vestrum 将失败表示为可识别的类别,提议在验证、检索、分解和知识合成方面的变更,并在将其作为整体进行评估之前筛查其范围。一个持久的经验教训文件为后续提议提供信息。在五种设置和两个基线框架上,冻结的框架提高了留出性能:相比 GAM,UltraHorizon 从 47.6 提升至 59.8;在八个留出任务上,相比 Claude Code,Terminal-Bench 4 Hard 的检查通过率从 63.7% 提升至 70.3%,测试成本为 1.03 倍;在一张切片的留出部分上,细胞类型注释一致性从 67.5% 提升至 77.8%,同时在 LoCoMo 和 AMA-Bench 上也有提升。在我们的搜索中,基于证据的验证对中间步骤和最终答案都有帮助,在中间步骤成本较低,而被要求重建已完成答案的批评者破坏的比修复的多。在三个记忆基准上,Vestrum 在每个配对评估中也得分高于所评估的 GEPA 配置。
cs.AI / 257 / 2609.33843

Laya as a Typed Probabilistic Assessor: An Independent Reproduction and a Preregistered Study of Calibration and Selective Escalation

Laya 作为类型化概率评估器:一项独立复现与关于校准及选择性升级的预注册研究
Nandakishore, Gowthamkumar
Abstract
The shipped Laya Typed-Decisions checkpoint, a 421M-parameter ModernBERT-large assessor that answers typed choice/noul/score questions over workflow state, is uniformly under-confident. The signed confidence-accuracy gap is $-0.214$, every occupied reliability bin's accuracy exceeds its confidence, and that sign uniformity collapses every binned ECE variant to the same value, $0.214$. The card frames the risk as over-confidence; the measured direction is the opposite, and the direction decides which way a confidence-gated cascade fails. A single disjointly fitted temperature ($T=0.469$, sharpening) removes most of the miscalibration (held-out ECE $0.204$ to $0.037$) and outperforms the shipped per-option-count table. The frozen selection rule instead chose isotonic regression, which overfit and failed its held-out NLL contrast on both tracks, so hypothesis H2 is not supported. Re-running the released checkpoint on its full official test split reproduces the card's headline accuracy ($0.767$ vs. $0.766$). The retrospective E1 reproduction preceded the analysis freeze; E2-E8 were prospectively preregistered, and 20 of 22 executed confirmatory tests reject under Benjamini-Hochberg FDR at $q=0.05$ (two descoped). The frozen gate beats random escalation but misses its 10% accepted-set error target on both tracks, an exploratory out-of-distribution probe finds no zero-shot transfer (accuracy $0.617$), and every score measures agreement with a synthetic teacher whose self-agreement ceiling ($0.735$) the specialist exceeds. Per-decision predictions, run manifests, and the frozen preregistration are in the ancillary files. The author has no affiliation with the model's publisher, the dataset's publisher, or TypeSafe.
Chinese Translation
已发布的 Laya Typed-Decisions 检查点(一个 421M 参数的 ModernBERT-large 评估器,用于回答关于工作流状态的类型化选择/noul/评分问题)一致地表现出置信度不足。带符号的置信度-准确率差距为 $-0.214$,每个被占据的可靠性分箱的准确率都超过其置信度,并且这种符号一致性使得所有分箱 ECE 变体都坍缩为相同的值,$0.214$。模型卡将风险描述为过度自信;而测量到的方向恰恰相反,并且该方向决定了置信度门控级联以何种方式失效。一个单独拟合的不相交温度($T=0.469$,锐化)消除了大部分校准误差(留出集 ECE 从 $0.204$ 降至 $0.037$),并且优于已发布的按选项计数表。然而,冻结的选择规则选择了等渗回归,该模型过拟合,并且在两个轨道上都未能通过其留出集 NLL 对比,因此假设 H2 未得到支持。在完整的官方测试划分上重新运行已发布的检查点,复现了模型卡中的主要准确率($0.767$ 对 $0.766$)。回顾性的 E1 复现先于分析冻结;E2-E8 是前瞻性预注册的,并且在 Benjamini-Hochberg FDR 控制下,22 个执行的验证性检验中有 20 个在 $q=0.05$ 水平上拒绝原假设(两个被取消范围)。冻结的门控优于随机升级,但在两个轨道上都未达到其 10% 接受集错误目标;一项探索性的分布外探测未发现零样本迁移(准确率 $0.617$);并且每个分数都衡量与合成教师的一致性,而专家模型超过了该教师的自一致性上限($0.735$)。逐决策预测、运行清单和冻结的预注册文件都在辅助文件中。作者与模型的发布者、数据集的发布者或 TypeSafe 均无隶属关系。
cs.AI / 258 / 2609.33845

How code helps different tasks? A decompositional lens on LLM post-training

代码如何帮助不同任务?LLM后训练中的分解视角
Yu, Zheng, Li, Yiwei, Chen, Yishen, Li, Xiang, Han, Jiale, Wang, Benyou, Chen, Jingbang
Abstract
Evaluating code data as a single corpus can obscure which types of code data benefit which models and downstream tasks. Effective data selection requires understanding both the benefits of individual categories and whether these benefits persist when categories are combined. We introduce a decompositional lens for studying these effects in LLM post-training. We first decompose an execution-verified code corpus into interpretable categories based on the computational patterns of its solutions. Through controlled fine-tuning experiments, we compare individual categories with a balanced mixture across instruction-tuned models on question answering, mathematics, and code generation. The resulting response maps reveal recurring gains in average question-answering performance, while the same category can improve one model or task and degrade another. The best-performing category also varies with the starting model and target task. We then compose compact mixtures guided by these results and examine whether benefits observed in individual categories persist under joint training. On selected model--task pairs, mixtures whose constituents each improve the target task outperform both their best constituent and full-corpus training while using roughly 10--15\% of the full corpus. These exploratory findings illustrate a \emph{less is more} pattern and highlight how the value of code data in post training depends on which categories are combined for which model and task.
Chinese Translation
将代码数据作为单一语料库进行评估,可能会掩盖哪些类型的代码数据对哪些模型和下游任务有益。有效的数据选择需要理解单个类别的益处,以及当类别组合时这些益处是否持续存在。我们引入一个分解视角来研究LLM后训练中的这些效应。我们首先基于其解决方案的计算模式,将一个经过执行验证的代码语料库分解为可解释的类别。通过受控的微调实验,我们在指令微调模型上,针对问答、数学和代码生成任务,比较了各个单独类别与一个平衡混合类别。得到的响应图揭示了在平均问答性能上反复出现的增益,而同一类别可能改善一个模型或任务,却降低另一个的性能。表现最佳的类别也随起始模型和目标任务而变化。然后,我们根据这些结果构建紧凑的混合,并检验在单个类别中观察到的益处是否在联合训练下持续存在。在选定的模型-任务对上,其组成成分各自都能改善目标任务的混合,在仅使用约10-15%的全语料库的情况下,优于其最佳组成成分和全语料库训练。这些探索性发现说明了‘少即是多’的模式,并强调了代码数据在后训练中的价值取决于为哪个模型和任务组合哪些类别。
cs.AI / 259 / 2609.33867

R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution

R$^2$ Flow:通过递归技能演化的递归自我改进
Zhang, Mingda, Huang, Qiang, Li, Yanjin, Wang, Zijia, Lin, Qika, Tang, Xiaoying, Shen, Tiesunlong
Abstract
LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision. Three obstacles stand in the way of making this self-improvement reliable: flow training suffers strategy collapse over tree-structured histories; nonnegative flow-based credit rewards frequent use as if it were benefit; and library edits rest on the task reward the policy optimizes. We introduce R$^2$ Flow, a recursive self-improvement framework that alternates policy learning, independent verification, and versioned skill-library updates on a shared-state orchestration graph. The graph merges histories that differ only in the order of independent steps, allowing flow training to pool evidence across equivalent executions. A flow-share readout of the trained flow, invariant to the backward policy, and a separate signed utility rank which skills to change, verifier evidence decides whether an edit is warranted, and a residual-variance plateau sets when to update. Committed edits reshape the graph the next policy learns on, realizing recursive skill evolution. Across question answering, mathematical reasoning, interactive decision making, and code generation, R$^2$ Flow improves task accuracy and library-edit precision over heuristic orchestration, reinforcement learning, and skill-evolution baselines, and transfers across executors. Code is available at https://github.com/beita6969/r2flow.
Chinese Translation
基于LLM的智能体可以通过重用和修正它们编排成可执行过程的技能,跨任务地自我改进。基于流的训练契合这一循环:它根据奖励按比例采样过程,而通过每个技能的流为其下一次技能库修订提供信用。三个障碍阻碍了这种自我改进的可靠性:流训练在树状历史上会遭遇策略崩溃;基于非负流的信用将频繁使用奖励得仿佛它是收益;而库编辑依赖于策略优化的任务奖励。我们提出R$^2$ Flow,一个递归自我改进框架,它在共享状态编排图上交替进行策略学习、独立验证和版本化技能库更新。该图合并仅在独立步骤顺序上不同的历史,允许流训练在等价执行之间汇集证据。对训练流的流量份额读出(对反向策略不变),以及一个单独的有符号效用来对哪些技能进行更改进行排序,验证器证据决定编辑是否合理,而残差方差平台期决定何时更新。提交的编辑重塑下一个策略学习所基于的图,实现递归技能演化。在问答、数学推理、交互式决策和代码生成中,R$^2$ Flow在任务准确性和库编辑精确度上优于启发式编排、强化学习和技能演化基线,并且可以跨执行器迁移。代码可在 https://github.com/beita6969/r2flow 获取。
cs.AI / 260 / 2609.33870

When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents

当成功策略失败:终端智能体对环境新颖性的适应
Singh, Janvijay, Shrivastava, Vaishnavi, Hakkani-Tur, Dilek, Kamar, Ece, Celikyilmaz, Asli
Abstract
LLM agents increasingly solve long-horizon tasks by autonomously interacting with their environment. In doing so, their strategies rely on assumptions about that environment: which resources and tools exist, where they are located, and how they behave. When these assumptions no longer hold, reliable agents must detect the change and adapt while pursuing the same goal. We study this adaptation capability through environmental novelty: a change that keeps the task objective fixed while invalidating an assumption underlying an otherwise successful trajectory. We introduce AGNI, an automated pipeline that extracts trajectory-relevant assumptions, injects targeted environmental changes, and validates that the resulting novel tasks remain solvable. Across three terminal benchmarks, AGNI produces diverse novelties spanning resources, interfaces, constraints, and execution semantics. Evaluating multiple LLM agents reveals a substantial adaptation gap between base and novel tasks. Trajectory analysis suggests that agents often encounter evidence of the change but fail to diagnose its cause and revise their strategy. Finally, post-training for environmental novelty improves adaptation to held-out novel tasks while also improving performance on base tasks. Our results highlight a gap between task competence and adaptive capability and motivate environmental variation as a core dimension of agent training and evaluation.
Chinese Translation
LLM智能体越来越多地通过自主与环境交互来解决长时程任务。在此过程中,它们的策略依赖于对环境的一些假设:存在哪些资源和工具、它们位于何处以及它们如何行为。当这些假设不再成立时,可靠的智能体必须在追求同一目标的同时检测变化并适应。我们通过环境新颖性来研究这种适应能力:一种保持任务目标固定,同时使原本成功轨迹所依赖的假设失效的变化。我们介绍了AGNI,一个自动化流水线,它提取与轨迹相关的假设,注入有针对性的环境变化,并验证所产生的新任务仍然可解。在三个终端基准测试中,AGNI产生了多样化的新颖性,涵盖资源、接口、约束和执行语义。评估多个LLM智能体揭示了在基础任务和新任务之间存在显著的适应差距。轨迹分析表明,智能体经常遇到变化的证据,但未能诊断其原因并修改其策略。最后,针对环境新颖性的后训练提高了对留出新颖任务的适应能力,同时也提高了在基础任务上的性能。我们的结果突显了任务能力与适应能力之间的差距,并促使将环境变化作为智能体训练和评估的核心维度。
cs.AI / 261 / 2609.33878

Curating Merchant-Matching Training Data with Two Confidence-Gated Local LLM Judges

使用两个置信度门控的本地LLM评判者整理商户匹配训练数据
Huang, Donghao, Pei, Jinling, Wang, Zhaoxia
Abstract
Merchant matching resolves a noisy payment descriptor to a retrieved merchant entity or returns no match. A key challenge in curating training labels is distinguishing teacher abstention from evidence that no acceptable entity exists: false no-match labels contaminate pseudo-labeled data, while conservative labeling reduces coverage. We investigate whether agreement between two local large language model judges improves pseudo-label reliability. A label is retained only when the judges agree, with separate ordered thresholds for selections and abstentions that guarantee disjoint positive and negative label sets. Retrospective replay on 2,000 expert-annotated queries shows that higher selection thresholds can improve positive-label purity, whereas higher abstention thresholds increase false no-match labels. At thresholds (0.86, 0.80), Muse Glimmer 30B and Gemma 4 31B jointly label 1,633 queries (81.7% coverage) at 96.88% purity; positive and negative purities are 99.47% and 93.38%. This exceeds either constituent model at the same thresholds by more than two percentage points, with lower coverage. A split-half check finds only 0.14 percentage points of threshold-selection optimism. A symmetric threshold of 0.86 adds 40 erroneous no-match labels, while 46 false abstentions persist even with no confidence threshold. Across five matched within-model comparisons, higher reasoning effort yields no clear F0.5 gain and increases median latency by 1.8-5.0 times. These results motivate separate thresholding and auditing for positive and negative pseudo-labels. The study establishes label purity, not student utility; fresh-data curation and student fine-tuning remain necessary to demonstrate downstream value.
Chinese Translation
商户匹配将嘈杂的支付描述符解析为检索到的商户实体,或返回无匹配。整理训练标签的一个关键挑战是区分教师弃权与不存在可接受实体的证据:错误的无匹配标签污染伪标注数据,而保守标注降低覆盖率。我们研究两个本地大语言模型评判者之间的一致性是否提高伪标签可靠性。仅当评判者一致时保留标签,并为选择和弃权设置单独的有序阈值,以保证正负标签集不相交。对2,000个专家标注查询的回顾性回放表明,较高的选择阈值可以提高正标签纯度,而较高的弃权阈值会增加错误的无匹配标签。在阈值 (0.86, 0.80) 下,Muse Glimmer 30B 和 Gemma 4 31B 联合标注了1,633个查询(81.7%覆盖率),纯度为96.88%;正纯度和负纯度分别为99.47%和93.38%。这比相同阈值下的任一组成模型高出两个百分点以上,但覆盖率较低。分半检查发现阈值选择乐观偏差仅为0.14个百分点。对称阈值0.86增加了40个错误的无匹配标签,而即使没有置信度阈值,仍有46个错误弃权。在五次匹配的模型内比较中,更高的推理努力没有带来明显的F0.5增益,并使中位延迟增加1.8-5.0倍。这些结果促使对正负伪标签分别进行阈值化和审计。该研究确立了标签纯度,而非学生效用;新数据整理和学生微调仍然是证明下游价值所必需的。
cs.AI / 262 / 2609.33910

When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents

当同意比上下文更长久:长生命周期智能体中的残余权限重放
Zhang, Zhihao, Wang, Chao, Li, Rujia, Wang, Qingze, Sun, Xiaoyan, Dai, Jun
Abstract
LLM agents increasingly rely on user approval to authorize security-sensitive actions at runtime. Such approvals are granted within a specific task and execution context. In long-lived agents, authorization decisions may need to persist across tasks or sessions. We find that this continuity can outlive the context that originally justified the approval, creating residual authority reusable without renewed consent. We expose this failure mode through a longitudinal attack that starts from a target security-sensitive action, identifies the authority required to execute it, induces benign interactions that legitimately obtain that authority, and later replays the residual authority during adversarial execution. Across controlled and live settings, we demonstrate that residual-authority replay arises in practice and substantially increases the success of prompt-injection and context-rebinding attacks. We evaluate 508 AgentDojo attack cases across six LLM families using production-derived authorization semantics. With residual authority, attack success rate (ASR) increases by up to 35.1 percentage points compared with a fresh authorization state. In live context-rebinding attacks on 55 Terminal-Bench cases across three real-world production coding agents, residual-authority replay increases ASR by 24.9 percentage points on average. These findings expose a fundamental mismatch between persistent authorization and the contextual nature of user consent in long-lived LLM agents.
Chinese Translation
LLM 智能体越来越多地依赖用户批准来在运行时授权安全敏感操作。此类批准是在特定任务和执行上下文中授予的。在长生命周期智能体中,授权决策可能需要在跨任务或会话中持续存在。我们发现,这种持续性可能比最初证明批准合理的上下文存活得更久,从而产生无需重新同意即可重用的残余权限。我们通过一种纵向攻击暴露了这一失效模式:该攻击从目标安全敏感操作开始,识别执行该操作所需的权限,诱导合法获得该权限的良性交互,随后在对抗性执行期间重放残余权限。在受控和真实环境中,我们证明残余权限重放在实践中会出现,并显著提高提示注入和上下文重绑定攻击的成功率。我们使用源自生产的授权语义,评估了跨六个 LLM 系列的 508 个 AgentDojo 攻击案例。与全新授权状态相比,使用残余权限时,攻击成功率(ASR)最高增加 35.1 个百分点。在对三个真实生产编码智能体的 55 个 Terminal-Bench 案例进行的实时上下文重绑定攻击中,残余权限重放平均使 ASR 增加 24.9 个百分点。这些发现揭示了长生命周期 LLM 智能体中持久授权与用户同意的上下文本质之间的根本性不匹配。
cs.AI / 263 / 2609.33920

HyperMCTS: Hypergraph-Augmented MCTS for Long-Horizon LLM Agents

HyperMCTS:面向长时程 LLM 智能体的超图增强 MCTS
Xiao, Tingsong, Moudhgalya, Nithish Balachandar, Basu, Chandrayee, Wang, Lichao, Kong, Luyang, Yao, Benjamin Z., Jiang, Zhe, Hao, Jie
Abstract
Long-horizon tasks require large language model (LLM) agents to coordinate decisions under constraints that span an entire solution. Monte Carlo Tree Search (MCTS) offers a promising approach to test-time scaling by exploring alternative action trajectories, but model computation and environment interaction make search costly. Efficient search therefore requires effective reuse of trajectory feedback. Standard MCTS maintains prefix-specific statistics, without explicitly accumulating outcomes for decision groups that recur across different paths. To fill this gap, we propose HyperMCTS, a training-free method that augments an ordered MCTS tree with a cross-trajectory hypergraph. Hyperedges represent groups of canonical decisions and accumulate their observed returns within the current task. Our hypergraph-guided HyperUCT selection rule aggregates evidence from overlapping hyperedges into an action prior, allowing outcomes collected under one prefix to inform selection under another while preserving execution histories in the tree. On DeepPlanning, HyperMCTS improves average planning accuracy by 2.3--7.3 percentage points over the strongest baseline for each of three backbone models. It enables Qwen3.6-27B to outperform Claude Opus 4.6 (max) on Shopping Planning, while achieving higher accuracy with fewer LLM calls and output tokens than the evaluated MCTS-based baselines. SealQA experiments further demonstrate improvements in question answering.
Chinese Translation
长时程任务要求大语言模型(LLM)智能体在贯穿整个解决方案的约束下协调决策。蒙特卡洛树搜索(MCTS)通过探索替代动作轨迹,为测试时扩展提供了一种有前景的方法,但模型计算与环境交互使搜索代价高昂。因此,高效搜索需要有效复用轨迹反馈。标准 MCTS 维护前缀特定的统计量,但没有显式累积在不同路径中重复出现的决策组的结果。为填补这一空白,我们提出了 HyperMCTS,这是一种免训练方法,它用跨轨迹超图增强有序 MCTS 树。超边表示规范决策组,并累积其在当前任务中观测到的回报。我们的超图引导 HyperUCT 选择规则将来自重叠超边的证据聚合为动作先验,允许在一个前缀下收集的结果为另一个前缀下的选择提供信息,同时在树中保留执行历史。在 DeepPlanning 上,对于三个主干模型中的每一个,HyperMCTS 相比最强基线将平均规划准确率提高了 2.3--7.3 个百分点。它使 Qwen3.6-27B 在 Shopping Planning 上优于 Claude Opus 4.6 (max),同时相比所评估的基于 MCTS 的基线,以更少的 LLM 调用和输出 token 实现了更高准确率。SealQA 实验进一步证明了在问答任务上的改进。
cs.AI / 264 / 2609.33955

Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents

为多轮业务智能体设计可靠的 LLM-as-a-Judge 测量系统
Luo, Kaiwen, Gao, Ming
Abstract
Many LLM-as-a-judge evaluations score fixed outputs under a fixed task definition. Production multi-turn business agents instead require a maintained measurement system: correctness depends on business-specific facts and procedures, outcomes emerge across turns, and failures must be attributed to either agent capability or missing business knowledge before they are actionable. We present an integrated methodology spanning evaluation specification, modular LLM judges, intent-preserving user simulation, and human-in-the-loop governance. The specification defines conversation-level end states and actionable failure ownership. Atomic judges share versioned evidence and feed an explicit aggregation graph. The simulator is released only after task-preservation and stability checks. Independent human audits estimate measurement fidelity, renew tiered reference sets, and route disagreements to label correction, guideline revision, or judge improvement. Production studies show that system-level fidelity improved across repeated audits, that human reviewers and automated judges improved together under the shared feedback loop, and that their combined workflow had the strongest descriptive performance in both reported task-completion settings. Because the studies are observational and the human reference itself required revision, these findings demonstrate operational usefulness rather than causal or universal superiority. The contribution is a practical framework for making multi-turn agent measurement reliable, actionable, and maintainable as the evaluated system and its evidence evolve.
Chinese Translation
许多 LLM-as-a-judge 评估在固定任务定义下对固定输出进行评分。生产环境中的多轮业务智能体则要求维护一套测量系统:正确性依赖于特定业务的事实与流程,结果在多轮交互中逐步显现,且在可采取行动之前,必须将失败归因于智能体能力不足或业务知识缺失。我们提出一套集成方法,涵盖评估规范、模块化 LLM 评委、保持意图的用户模拟,以及人在回路中的治理。该规范定义了对话级终态和可操作的失败归属。原子评委共享带版本控制的证据,并输入一个显式的聚合图。模拟器只有在通过任务保持性和稳定性检查后才会发布。独立人工审计用于估计测量保真度、更新分层参考集,并将分歧分流至标签校正、指南修订或评委改进。生产研究表明,系统级保真度在重复审计中有所提升;在共享反馈回路下,人工评审员与自动评委共同改进;并且二者的组合工作流在两个已报告的任务完成场景中均具有最强的描述性表现。由于这些研究是观察性的,且人工参考本身也需要修订,这些发现表明的是运营实用性,而非因果性或普适优越性。其贡献在于一个实用框架,使多轮智能体测量在所评估系统及其证据不断演化时仍保持可靠、可操作和可维护。
cs.AI / 265 / 2609.34007

EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events

EHRAdapt:利用语义先验将预训练语言模型适配到电子健康记录以处理罕见临床事件
Goncalves, Andre R, Liu, Vincent, Ray, Priyadip
Abstract
Electronic health records (EHRs) encode clinical histories as (time, modality, code) tuples, whereas pretrained language models expect text tokens. Serializing them as text inflates sequence length and redundantly encodes structure. We introduce EHRAdapt, an adapter that maps tuples directly into a frozen language model's embedding space. Modality receives a learned embedding, time gaps enter through learned attention biases, and event codes receive dedicated vectors. Learning event vectors is the central challenge: clinical vocabularies are long-tailed, leaving rare events too few observations for reliable estimates. EHRAdapt therefore represents each event vector as the sum of a semantic prior and an evidence residual. The prior is a frozen embedding of the event's clinical description from a biomedical language model trained on clinical ontologies, mapped into the model's input space by a shared learned projection, so it supplies clinical meaning even when observations are scarce. The residual, a learned low-rank event-specific correction, refines it as evidence accumulates. We run continued pretraining on about 4 million patients' records with three frozen LLM backbones (OLMo2 1B, Llama3.2 1B, and OLMo2 7B), training only the adapter (0.1--0.6% of all parameters). The full adapter outperforms all ablations in held-out next-event prediction on every backbone. Removing the semantic pathway hurts rare events over ten times more than the most frequent ones, whereas removing the residual hurts overall prediction but improves it for the rarest events. On reportable infectious-disease and syndromic downstream classification tasks, EHRAdapt outperforms text-based LLM and count-based baselines, and both pathways improve rare-disease discrimination. The two pathways therefore play complementary roles, visible only when results are broken down by event frequency rather than averaged.
Chinese Translation
电子健康记录(EHRs)将临床历史编码为(时间、模态、代码)元组,而预训练语言模型则期望文本标记。将它们序列化为文本会膨胀序列长度并冗余地编码结构。我们提出EHRAdapt,一种适配器,可将元组直接映射到冻结的语言模型的嵌入空间。模态获得一个学习到的嵌入,时间间隔通过学习到的注意力偏置进入,事件代码获得专门的向量。学习事件向量是核心挑战:临床词汇呈长尾分布,导致罕见事件的观测太少,无法进行可靠估计。因此,EHRAdapt将每个事件向量表示为语义先验和证据残差之和。先验是来自在临床本体上训练的的生物医学语言模型的事件临床描述的冻结嵌入,通过共享的学习投影映射到模型的输入空间,因此即使观测稀少也能提供临床意义。残差是一个学习到的低秩事件特定校正,随着证据积累对其进行细化。我们在约400万患者的记录上使用三个冻结的LLM骨干(OLMo2 1B、Llama3.2 1B和OLMo2 7B)进行持续预训练,仅训练适配器(占所有参数的0.1--0.6%)。完整的适配器在每个骨干模型的留出下一事件预测中均优于所有消融实验。移除语义路径对罕见事件的伤害比最频繁事件高出十倍以上,而移除残差会损害整体预测,但会改善最罕见事件的预测。在可报告的传染病和综合征下游分类任务上,EHRAdapt优于基于文本的LLM和基于计数的基线,并且两种路径都改善了罕见疾病的区分能力。因此,这两种路径发挥互补作用,只有按事件频率细分结果而非平均时才能显现。
cs.AI / 266 / 2609.34015

A Computer Vision Approach to Visual Fraud Detection in Phishing Websites Using YOLOv8

基于YOLOv8的计算机视觉方法用于钓鱼网站视觉欺诈检测
Shaikh, Basil Sajid, Homayouni, Hajar
Abstract
Phishing remains one of the most common vectors for financial and identity fraud, and most detection systems still rely on inspecting a page's URL, HTML markup, or domain registration history. These signals are easy for an attacker to rotate or obfuscate, and they say very little about what actually convinces a victim to hand over a password or a card number: the way the page looks. This paper describes a visual, image-based approach to phishing detection that treats a rendered webpage the same way a human eye would, as a picture that either matches a trusted brand or doesn't. A YOLOv8 convolutional neural network was trained to classify full-page website screenshots as phishing or legitimate based on layout, logo placement, color scheme, and login-form structure, rather than on text extracted from the page. The system reached 92% classification accuracy on a held-out test set, processed a single screenshot in roughly 100 milliseconds, and, after a round of data augmentation aimed specifically at lighting, compression, and scaling variation, cut the false-positive rate by 11% relative to the pre-augmentation baseline. The paper walks through the dataset construction, the augmentation strategy, the model architecture and training setup, and the resulting performance, and closes with a discussion of where this kind of visual detector fits alongside, rather than instead of, existing URL- and content-based defenses.
Chinese Translation
钓鱼至今仍是最常见的金融与身份欺诈途径之一,而大多数检测系统仍依赖于检查页面的URL、HTML标记或域名注册历史。这些信号很容易被攻击者轮换或混淆,而且几乎无法说明究竟是什么让受害者交出了密码或银行卡号:页面的外观。本文描述了一种基于视觉、以图像为基础的钓鱼检测方法,它像人眼一样将渲染后的网页视为一张图片,该图片要么与受信任品牌相符,要么不符。我们训练了一个YOLOv8卷积神经网络,根据布局、徽标位置、配色方案和登录表单结构,而不是根据从页面中提取的文本来将整页网站截图分类为钓鱼或合法。系统在留出测试集上达到92%的分类准确率,处理单张截图约需100毫秒,并且在专门针对光照、压缩和缩放变化进行一轮数据增强后,相对于增强前基线将假阳性率降低了11%。本文详细介绍了数据集构建、增强策略、模型架构与训练设置以及由此得到的性能,并在最后讨论了这种视觉检测器如何与现有的基于URL和内容的防御措施并存,而不是取而代之。
cs.AI / 267 / 2609.34024

Jev in Medicine: A Benchmark Evaluation. Preliminary Results

Jev在医学中的应用:基准评估。初步结果
Madrid-García, Alfredo, Merino-Barbancho, Beatriz
Abstract
Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev's accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev's probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev's probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was "I don't know or cannot answer", Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.
Chinese Translation
Jev是一种非生成式“系统1”模型,它为预定义的答案选项分配概率,无法回答这些选项之外的问题。其在医学问答和基于病例的诊断推理任务上的准确性和校准情况尚不清楚。我们在四个医学基准测试上评估了Jev 1.13:MetaMedQA、PubMedQA、DiagnosisArena-MCQ和NEJM病例挑战。以使用(中等)推理和不使用推理的GPT-6 Sol作为参考。主要结局指标是top-1准确率;关键次要结局指标是校准、选择性预测和无法回答问题的识别。所有8,469次请求均返回了有效答案。在PubMedQA上,Jev的准确率与使用中等推理的GPT-6 Sol相似(78.4% vs 78.2%;),在MetaMedQA上较低(74.8% vs 82.7%),在DiagnosisArena-MCQ(59.8% vs 82.4%;)和NEJM病例(61.8% vs 82.4%)上则低得多。在MetaMedQA上,Jev的概率校准最佳(期望校准误差0.063 vs 0.146),其概率至少为0.9的答案(占问题的52.9%)准确率为93.4%,但当GPT-6 Sol接受类似比例的问题时,其准确率相同。在DiagnosisArena-MCQ上,Jev的概率区分能力较差(AUROC 0.645 vs 0.768)。在162个正确答案为“我不知道或无法回答”的问题中,Jev选择该选项的比例为10.5%(GPT-6 Sol为8.6%)。中位延迟为0.27-0.31秒;所有2,823个项目花费0.08美元。Jev速度快且成本低,其在研究摘要上的准确率与前沿LLM相似,但在考试问题上较低,在复杂诊断病例上则低得多。在临床使用前需要进行特定任务的验证。
cs.AI / 268 / 2609.34039

Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation

大语言模型用于结构化临床数据分析:双智能体接地与验证
Dehkalani, Erfan D., Shankaran, Seetha, Laptook, Abbot R., Cotten, C. Michael, Grant, P. Ellen, Ou, Yangming
Abstract
Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the executed query and database result, and produce a draft. Deterministic checks and a separately invoked cross-provider Validation Agent then accept the draft, request one bounded repair, or abstain. We formalized the system as a bounded selective pipeline and evaluated CLEAR-Med's configuration and scalability, and the Invocation Agent's accuracy and consistency on a 25-query development benchmark, using a harmonized 21-site neonatal hypoxic-ischemic encephalopathy table containing 532 de-identified infant records and approximately 1,300 variables. Results: CLEAR-Med completed all six nominal scalability configurations, including 500x1300. Across 25 development-benchmark queries repeated five times, the Invocation Agent answered 83 of 125 responses correctly (66.4%; query-cluster bootstrap 95% CI, 48.0-83.2%), compared with 15 of 125 (12.0%; 95% CI, 3.2-22.4%) for the ungrounded ChatGPT baseline, a paired improvement of 54.4 percentage points (95% CI, 36.8-72.0%). Conclusion: CLEAR-Med provides a general architecture for traceable analysis of structured clinical data: numerical claims remain linked to executed SQL, and unresolved cases can fail closed. The reported experiments characterize CLEAR-Med's configuration and scalability and the Invocation Agent's accuracy, while the formal analysis establishes the encoded-property guarantee of the complete control flow; a prospective full-pipeline evaluation of the validation and abstention stages is the next stage of this work.
Chinese Translation
目的:开发并描述 CLEAR-Med,一个用于结构化临床数据自然语言分析的双智能体框架,该框架将基于 SQL 的调用与独立验证分离。方法:CLEAR-Med 使用一个智能体将问题转换为可执行的结构化查询语言(SQL),保留已执行的查询和数据库结果,并生成草稿。确定性检查和单独调用的跨提供商验证智能体随后接受草稿、请求一次有界修复或弃权。我们将该系统形式化为一个有界选择性管道,并使用一个协调的 21 个站点的新生儿缺氧缺血性脑病表(包含 532 条去标识化婴儿记录和约 1,300 个变量)评估了 CLEAR-Med 的配置和可扩展性,以及调用智能体在 25 个查询的开发基准上的准确性和一致性。结果:CLEAR-Med 完成了所有六种标称可扩展性配置,包括 500x1300。在 25 个开发基准查询重复五次中,调用智能体正确回答了 125 个响应中的 83 个(66.4%;查询聚类自助法 95% CI,48.0-83.2%),而未接地的 ChatGPT 基线为 125 个中的 15 个(12.0%;95% CI,3.2-22.4%),配对改进为 54.4 个百分点(95% CI,36.8-72.0%)。结论:CLEAR-Med 为结构化临床数据的可追溯分析提供了一个通用架构:数值声明保持与已执行的 SQL 关联,未解决的案例可以故障关闭。所报告的实验描述了 CLEAR-Med 的配置和可扩展性以及调用智能体的准确性,而形式化分析确立了完整控制流的编码属性保证;对验证和弃权阶段的前瞻性全管道评估是这项工作的下一阶段。
cs.AI / 269 / 2609.34049

Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference

跳出框框思考:滑动窗口 KV 推理中的信息保留与传递
DeLise, Timothy, Cromelin, Seth
Abstract
Sliding-window KV inference refers to processing a sequence incrementally while retaining only a fixed-size cache of recent key and value states. It can be applied to pretrained causal transformers at inference time without additional training, while its KV-cache memory remains fixed as more tokens are processed. Because cached states are computed in the context of earlier tokens, they may carry information from beyond the current window and transmit it to later states. This study presents a series of experiments using five open-weight models spanning Qwen, Llama, Mistral, and Muse Glimmer. We investigate whether information originating outside the immediate context window can persist through a rolling KV cache and remain useful for retrieval. Initial results show that retaining previously computed states improves retrieval across the models tested compared with recomputing the final fixed window from raw tokens. We then measure how far this effect extends and find that Muse Glimmer and Mistral 7B show the strongest \emph{latent information relay}: they can recover information even after the relevant source tokens have left the cache. Both models incorporate sliding-window attention in their published architectures, an association that motivates testing whether training with sliding windows promotes more reliable information retention.
Chinese Translation
滑动窗口 KV 推理是指以增量方式处理序列,同时仅保留固定大小的近期键和值状态缓存。它可以在推理时应用于预训练的因果 Transformer,而无需额外训练,并且随着处理更多 token,其 KV 缓存内存保持固定。由于缓存状态是在较早 token 的上下文中计算的,它们可能携带来自当前窗口之外的信息,并将其传递给后续状态。本研究使用五个开放权重模型进行了一系列实验,涵盖 Qwen、Llama、Mistral 和 Muse Glimmer。我们研究源自直接上下文窗口之外的信息能否通过滚动 KV 缓存持续存在,并仍对检索有用。初步结果表明,与从原始 token 重新计算最终固定窗口相比,保留先前计算的状态在受测模型中提升了检索效果。随后我们衡量了这种效应的延伸范围,并发现 Muse Glimmer 和 Mistral 7B 表现出最强的潜在信息中继:即使在相关源 token 已离开缓存后,它们也能恢复信息。这两个模型在其已发表架构中都采用了滑动窗口注意力,这一关联促使我们测试使用滑动窗口进行训练是否会促进更可靠的信息保留。
cs.AI / 270 / 2609.34069

Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization

迈向证书驱动的软件移植:一种用于科学程序优化的自我改进代理式执行框架
Jha, Piyush, Ghosh, Aishik, Ganesh, Vijay
Abstract
The upgrade and rewriting of large scientific codebases has traditionally been a major challenge. While evolutionary search with large language models (LLMs) can port and accelerate legacy code, repair feedback in prompts alone does not prevent subsequent candidates from repeating the same errors. We introduce Certificate-Driven Evolutionary Search (CDES), which extends evolutionary search with enforceable restrictions derived from failed candidates, recorded as certificates of assumptions, checker evidence, and justified restrictions. Its control logic enforces these restrictions through rejection, backtracking, and targeted repair while preserving compatible edits. We apply CDES to CPU-to-GPU translation of two particle-simulation functions from the Geant4 toolkit, evaluated with a harness that goes beyond unit tests to combine formal checks, numerical comparisons, physics checks, and GPU safety tests. Generated implementations achieve 13.78x and 23.54x function-level speedups over CPU code, including data conversion and transfers; for one function, GPU throughput exceeds an expert implementation by 14.9%, reaching 16.1% when complementary components are combined. In an ablation over execution settings, certificate feedback increases the fraction of candidates passing required correctness checks from 55% to 90%.
Chinese Translation
大型科学代码库的升级和重写历来是一项重大挑战。尽管使用大语言模型(LLMs)的进化搜索可以移植和加速遗留代码,但仅靠提示中的修复反馈并不能防止后续候选方案重复相同的错误。我们引入了证书驱动的进化搜索(CDES),它通过从失败候选中提取的可执行限制来扩展进化搜索,这些限制被记录为假设证书、检查器证据和合理限制。其控制逻辑通过拒绝、回溯和定向修复来强制执行这些限制,同时保留兼容的编辑。我们将 CDES 应用于 Geant4 工具包中两个粒子模拟函数的 CPU 到 GPU 转换,并使用一个超越单元测试的测试框架进行评估,该框架结合了形式化检查、数值比较、物理检查和 GPU 安全性测试。生成的实现相比 CPU 代码实现了 13.78 倍和 23.54 倍的函数级加速(包括数据转换和传输);对于其中一个函数,GPU 吞吐量比专家实现高出 14.9%,当组合互补组件时达到 16.1%。在针对执行设置的消融实验中,证书反馈将通过对所需正确性检查的候选方案比例从 55% 提高到 90%。
cs.AI / 271 / 2609.34072

PhysFieldBench: Can Multimodal Models Understand Physical Fields?

PhysFieldBench:多模态模型能理解物理场吗?
Ma, Yuezhou, Weng, Huikun, Wu, Jialong, Zhao, Chenyi, Zhou, Hang, Shangguan, Haonan, Wang, Jianmin, Long, Mingsheng
Abstract
Multimodal large language models (MLLMs) are increasingly envisioned as core components of scientific and engineering agents, yet their ability to interpret physical fields remains poorly understood. Existing physics benchmarks largely emphasize textbook problem solving or intuitive physical reasoning, leaving open whether MLLMs can infer physically meaningful information from continuous field observations. We introduce PhysFieldBench, a benchmark comprising 24 tasks and 1,160 evaluation examples across controlled equation fields, simulated physical fields, and observed physical fields. The tasks assess three forms of inference: identifying physical mechanisms, comparing latent control variables, and predicting outcome properties. Across representative open-source and proprietary MLLMs, zero-shot performance is low: the best model achieves a chance-normalized score of 29.3, while several open-source models remain near chance. In contrast, a task-specific supervised vision transformer performs substantially better, demonstrating that the inputs contain learnable physical information. To diagnose these failures, a structured self-explanation analysis attributes most errors to missed visual patterns and incorrect visual-to-physical mappings. Further, to explore whether post-training can improve physical inference and generalize to unseen tasks, we compare supervised fine-tuning with final answers or chain-of-thought supervision and reinforcement learning. Final-answer supervision performs best overall but transfers less effectively, whereas reinforcement learning after chain-of-thought supervision achieves the best generalization. Together, these findings highlight the need to improve visual-to-physical grounding and cross-task generalization for MLLMs to reliably interpret physical fields in scientific and engineering workflows.
Chinese Translation
多模态大语言模型(MLLMs)正日益被设想为科学与工程智能体的核心组件,但其解读物理场的能力仍鲜为人知。现有物理基准大多强调教科书问题求解或直观物理推理,仍未回答 MLLMs 能否从连续场观测中推断出具有物理意义的信息。我们提出 PhysFieldBench,一个包含 24 项任务和 1,160 个评估示例的基准,涵盖受控方程场、模拟物理场和观测物理场。这些任务评估三种推理形式:识别物理机制、比较潜在控制变量以及预测结果属性。在代表性开源和专有 MLLMs 上,零样本性能较低:最佳模型取得了 29.3 的经随机水平归一化的得分,而若干开源模型仍接近随机水平。相比之下,一个任务专用的监督式视觉 Transformer 表现显著更好,表明输入含有可学习的物理信息。为诊断这些失败,一项结构化自我解释分析将大多数错误归因于遗漏的视觉模式和错误的视觉到物理映射。此外,为探索后训练能否提升物理推理并泛化到未见任务,我们比较了使用最终答案或思维链监督的监督微调与强化学习。最终答案监督总体表现最佳,但迁移效果较差;而思维链监督后的强化学习实现了最佳泛化。总之,这些发现凸显了需要改进视觉到物理的 grounding 和跨任务泛化,以使 MLLMs 能在科学与工程工作流程中可靠地解读物理场。
cs.AI / 272 / 2609.34079

GenoMorph: Pathway-Grounded Genomic Disease Reasoning via Adaptive Latent Computation

GenoMorph:通过自适应潜在计算实现基于通路的基因组疾病推理
Halder, Tanmoy Kanti, Ghosh, Akash, Roy, Arijit, Saha, Sriparna
Abstract
Large language models (LLMs) have demonstrated strong capabilities in biological reasoning; however, genomic disease inference remains largely dependent on memorized gene-disease associations rather than understanding biological pathways. This shortcut learning undermines robustness and generalization, and breaks down when molecular identifiers are unavailable. We present GenoMorph, a multimodal genomic reasoning framework that shifts disease prediction from associative gene-disease mapping toward pathway-grounded reasoning. GenoMorph couples a frozen DNA foundation model with question-conditioned cross-attention fusion, self-adaptive latent reasoning (LatentSp), a residual reasoning gate for iterative genomic evidence reinjection, and rejection sampling fine-tuning regularized by hierarchical optimal transport (OT). Rather than learning direct gene-disease mappings, GenoMorph aligns genomic sequence representations with latent pathway dynamics, enabling reasoning trajectories that follow molecular interactions before producing disease predictions. LatentSp dynamically allocates computation according to reasoning confidence, reducing unnecessary reasoning steps and improving inference efficiency. We further construct an anonymized benchmark from the Kyoto Encyclopedia of Genes and Genomes (KEGG), replacing every gene and molecular identifier with anonymous symbols while preserving sequences and pathway topology, thereby removing memorization shortcuts. GenoMorph raises the weighted F1 from 0.7863 (BioReason) to 0.9412, and rejection sampling fine-tuning with self-adaptive latent reasoning pushes it to 0.9725 while cutting latency nearly 60%. On the anonymized benchmark it reaches 0.9465 F1, substantially outperforming prior systems and confirming that accurate disease prediction can arise from pathway reasoning rather than memorized gene-disease associations.
Chinese Translation
大型语言模型(LLMs)在生物推理方面展现出强大能力;然而,基因组疾病推断仍然很大程度上依赖于记忆的基因-疾病关联,而非理解生物通路。这种捷径学习削弱了鲁棒性和泛化能力,并且在分子标识符不可用时失效。我们提出 GenoMorph,一个多模态基因组推理框架,将疾病预测从关联性基因-疾病映射转向基于通路的推理。GenoMorph 将冻结的 DNA 基础模型与问题条件交叉注意力融合、自适应潜在推理(LatentSp)、用于迭代基因组证据重注入的残差推理门以及由分层最优传输(OT)正则化的拒绝采样微调相结合。GenoMorph 不是学习直接的基因-疾病映射,而是将基因组序列表示与潜在通路动态对齐,使得推理轨迹在产生疾病预测之前遵循分子相互作用。LatentSp 根据推理置信度动态分配计算,减少不必要的推理步骤并提高推理效率。我们进一步从京都基因与基因组百科全书(KEGG)构建了一个匿名基准,将每个基因和分子标识符替换为匿名符号,同时保留序列和通路拓扑结构,从而消除记忆捷径。GenoMorph 将加权 F1 从 0.7863(BioReason)提升到 0.9412,而结合自适应潜在推理的拒绝采样微调将其推至 0.9725,同时将延迟降低近 60%。在匿名基准上,它达到 0.9465 F1,显著优于先前系统,并证实准确的疾病预测可以来自通路推理而非记忆的基因-疾病关联。
cs.AI / 273 / 2609.34082

K-OPSD: Verifiable On-Policy Self-Distillation for Post-Training Vision-Language Models on AEC Drawings

K-OPSD:用于AEC图纸上后训练视觉语言模型的可验证在线策略自蒸馏
Bai, Yunfei, Chionna, Enrico, Amol, Akash, KC, Kawaljit Singh, Tinnemeyer, Joern
Abstract
Interpreting architecture, engineering, and construction (AEC) drawings is hard for general Multimodal Large Language Models (MLLMs) and vision-language models (VLMs). We introduce K-OPSD, a VLM post-training methodology for improving AEC drawing understanding. Building on On-Policy Self-Distillation (OPSD) with verifiable supervision, we construct a teacher from the model's own best-of-N generations, certified by a process-level verifier, and rescue failed prompts by resampling under a hint that exposes the verified answer. We then perform an on-policy model update by training on verified completions with a cross-entropy inner-loss, outperforming the bounded token-wise generalized Jensen-Shannon divergence (JSD) used by on-policy distillation. Using K-OPSD, we fine-tune Qwen3-VL models on the AECV-Bench dataset. The resulting models attain the top average judge score (0.819) and combined accuracy (0.738), achieving competitive results against open-source baseline models. The recipe transfers to the out-of-domain ArchCAD dataset, where the 8B model gains most. We present the verifier suite and the continual learning and self-improving pipeline, our results provide preliminary evidence that verifier-guided self-distillation is a promising route toward more reliable machine reading of architecture drawings.
Chinese Translation
解读建筑、工程和施工(AEC)图纸对于通用多模态大语言模型(MLLMs)和视觉语言模型(VLMs)来说很困难。我们提出了K-OPSD,一种用于提升AEC图纸理解的VLM后训练方法。基于具有可验证监督的在线策略自蒸馏(OPSD),我们从模型自身的best-of-N生成中构建教师,并通过过程级验证器进行认证,并通过在暴露已验证答案的提示下重新采样来挽救失败的提示。然后,我们通过使用交叉熵内部损失对已验证的补全进行训练来执行在线策略模型更新,其性能优于在线策略蒸馏所使用的有界逐词元广义Jensen-Shannon散度(JSD)。使用K-OPSD,我们在AECV-Bench数据集上微调Qwen3-VL模型。所得模型获得了最高的平均评判分数(0.819)和综合准确率(0.738),与开源基线模型相比取得了有竞争力的结果。该方法迁移到域外的ArchCAD数据集,其中8B模型获益最多。我们提出了验证器套件以及持续学习和自我改进流程,我们的结果为验证器引导的自蒸馏是通向更可靠的机器阅读建筑图纸的有前景的途径提供了初步证据。
cs.AI / 274 / 2609.34103

A Differentiable Optimization Framework for Registering Sequential Bounding Boxes with Point Cloud Stream

面向序列边界框与点云流配准的可微优化框架
Li, Xuesong, Tong, Jinguang, Hong, Jie
Abstract
Refining a sequence of coarse 3D bounding boxes against a LiDAR point-cloud stream demands tracks that are geometrically accurate (high IoU) and temporally coherent (low roughness), preferably without training data. The usual recipe keeps the two concerns apart: register each frame independently, then smooth the trajectory afterwards with a Kalman~RTS or Savitzky--Golay filter. Smoothing displaces boxes from a geometric optimum and never re-optimises, so it trades accuracy for smoothness. We instead fold the temporal smoothness constraint into a training-free registration objective and solve for all poses jointly with L-BFGS. The payoff depends on how well the object is seen. On well-observed tracks it is large: within the low-roughness budget, the joint objective beats both post-hoc smoothers on paired multi-seed statistics and cuts roughness several-fold relative to frame-wise registration at matched accuracy. Treating visibility as an experimental variable exposes the limit. The advantage decays monotonically as views become one-sided, until it is indistinguishable from zero for near-edge-on objects and slightly negative under a ray-cast simulator with range-dependent density and ego motion, where the decoupled pipeline is in fact ahead at tight roughness budgets. We locate that boundary and trace it to one term: orientation alignment ties yaw to the estimated velocity and fails once that estimate is noisy. A ground-truth-free rule can choose the temporal scale and keep every track inside the roughness budget.
Chinese Translation
针对LiDAR点云流对一系列粗3D边界框进行细化,要求轨迹在几何上精确(高IoU)且时间上连贯(低粗糙度),最好无需训练数据。通常的做法将这两个问题分开处理:独立配准每一帧,然后用Kalman-RTS或Savitzky-Golay滤波器对轨迹进行平滑。平滑会使边界框偏离几何最优,且从不重新优化,因此是以精度换取平滑度。我们转而将时间平滑约束融入无需训练的配准目标中,并使用L-BFGS联合求解所有位姿。收益取决于目标的可见程度。在观测良好的轨迹上,收益很大:在低粗糙度预算内,联合目标在配对多种子统计上优于两种事后平滑器,并在匹配精度下将粗糙度相对于逐帧配准降低数倍。将可见性作为实验变量揭示了其局限。随着视角变得单侧,优势单调递减,直至对于近乎侧向的目标与零无法区分,并且在具有距离相关密度和自运动的射线投射模拟器下甚至略微为负,此时解耦流程在严格的粗糙度预算下实际上更优。我们定位了这一边界,并将其追溯到一个项:方向对齐将偏航角与估计速度绑定,一旦该估计存在噪声便会失效。一种无需真值的规则可以选择时间尺度,并使每条轨迹都保持在粗糙度预算内。
cs.AI / 275 / 2609.34111

SpecRegMatch: Robust Semi-Supervised Regression for Vehicle Interior Noise Prediction

SpecRegMatch:面向车内噪声预测的鲁棒半监督回归
Sim, Sejin, Bae, Jinsoo, Kim, Seoung Bum
Abstract
The rapid advancement of artificial intelligence has observed increased application in predicting vehicle interior noise levels within the automotive industry. However, the collection of labeled data for training models in this context involves significant costs. Previous studies in semi-supervised regression (SSR) have effectively mitigated the reliance on labeled data by incorporating unlabeled data. Nonetheless, these approaches often introduce a high computational cost due to the training of multiple models and data sampling. This study introduces SpecRegMatch, a novel SSR method aimed at addressing the computational cost associated with training by leveraging a single model, thus eliminating the need for multiple data samplings. SpecRegMatch integrates consistency regularization and information maximization to robustly train the model, achieved through various augmentations applied to both the embedding vectors and predicted values. Experimental results demonstrate that SpecRegMatch achieves state-of-the-art performance across various scenarios, even when using a single model. It attains a remarkable performance, as indicated by an R^2 score of 0.434. This is especially noteworthy in scenarios where labeled data is scarce. You can access the code for our proposed method at https://github.com/sejin-sim/SpecRegMatch.
Chinese Translation
人工智能的快速发展使其在汽车行业车内噪声水平预测中的应用日益增多。然而,在此背景下,为训练模型收集标记数据涉及高昂成本。以往在半监督回归(SSR)方面的研究通过引入未标记数据,有效减轻了对标记数据的依赖。然而,这些方法通常由于训练多个模型和数据采样而带来高昂的计算成本。本研究提出了SpecRegMatch,一种新颖的半监督回归方法,旨在通过利用单个模型来解决与训练相关的计算成本,从而无需进行多次数据采样。SpecRegMatch结合了一致性正则化和信息最大化来稳健地训练模型,这是通过对嵌入向量和预测值应用各种增强来实现的。实验结果表明,即使使用单个模型,SpecRegMatch在各种场景下也达到了最先进的性能。其性能卓越,R^2得分为0.434。这在标记数据稀缺的场景中尤其值得注意。我们提出方法的代码可在https://github.com/sejin-sim/SpecRegMatch获取。
cs.AI / 276 / 2609.34113

GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions

GUITAR:通过状态转移的GUI智能体结构化故障诊断
Zhang, Shaoqing, Chen, Kehai, Bai, Xuefeng, Zhang, Zhuosheng, Zhang, Pengfei, Xiang, Yang, Zhang, Min
Abstract
Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) functionally equivalent screens are evaluated in isolation, obscuring systematic failure patterns across shared screens; and (2) the long-tailed GUI distribution renders failures on rare but critical screens invisible under standard metrics. To address these issues, we propose \textbf{GUITAR}, a state-centric diagnostic framework that performs structured failure analysis over both states and transitions, using a State Transition Graph (STG) by mapping visually diverse screens to shared functional states. Across 8 agents and 6 tasks from AndroidControl and Mind2Web, GUITAR reveals that 60.4\% of failures occur in 20\% of states, localizing errors to a small set of bottlenecks. Bottleneck-targeted guidance improves SR by 2.8\% and retains a 1.88\% average gain across 7 agents under three-fold trajectory-held-out evaluation with fully automatic STGs. These findings demonstrate the diagnostic and actionable value of structure-aware evaluation within the evaluated mobile and web tasks. Code is available at https://github.com/sqzhang-lazy/GUITAR
Chinese Translation
理解图形用户界面(GUI)智能体在何处以及为何失败,对于构建更可靠的系统至关重要,然而当前的评估依赖于步骤准确率,这一指标独立对待每个屏幕,忽略了GUI环境的底层结构。这导致两个关键盲点:(1)功能等效的屏幕被孤立评估,掩盖了跨共享屏幕的系统性故障模式;(2)长尾GUI分布使得在标准指标下,罕见但关键屏幕上的故障不可见。为了解决这些问题,我们提出了GUITAR,一个以状态为中心的诊断框架,它通过状态转移图(STG)对状态和转移进行结构化故障分析,将视觉多样的屏幕映射到共享的功能状态。在来自AndroidControl和Mind2Web的8个智能体和6个任务中,GUITAR揭示出60.4%的故障发生在20%的状态中,将错误定位到一小组瓶颈。瓶颈针对性引导将成功率(SR)提高了2.8%,并在三倍轨迹留出评估下,使用全自动STG,在7个智能体上保持平均1.88%的增益。这些发现证明了在所评估的移动和Web任务中,结构感知评估的诊断性和可操作性价值。代码可在https://github.com/sqzhang-lazy/GUITAR获取。
cs.AI / 277 / 2609.34132

From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents

从攻击成功到攻击严重性:针对LLM智能体的反事实记忆攻击
Zou, Mingxi, Liang, Langzhang, Wang, Zhuo, Zhao, Yiyang, Qu, Lizhen, Xu, Zenglin
Abstract
As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typically evaluated by whether they succeed, yet successful attacks can leave persistent states with substantially different downstream consequences. We study this severity as a distinct attack-design objective and formalize it with counterfactual memory regret (CMR), the paired increase in expected downstream loss relative to clean memory. We introduce MemHarm, which predeclares a finite class of sparse, grounded semantic edits, evaluates candidates through the normal agent memory interface using offline paired-loss feedback, and certifies resolved selections within that class. Compared with attack-success optimization, CMR-guided selection produces substantially larger downstream loss while retaining most of the success-rate gain. Across two agent benchmarks and diverse memory designs, MemHarm attains the highest CMR point estimates among the evaluated general attacks on identical support. Factor-removal interventions link this harm to the selected semantic factor, and native-agent deployments verify the write-to-fresh-process attack path.
Chinese Translation
随着LLM智能体越来越多地依赖持久记忆来实现长程和个性化行为,它们能够在交互中保留并重用信息,但这也创造了一个持久通道,通过该通道,恶意记忆写入可以影响未来的行为。持久记忆攻击通常根据是否成功来评估,然而成功的攻击可能会留下具有显著不同下游后果的持久状态。我们将这种严重性作为一个独特的攻击设计目标进行研究,并用反事实记忆遗憾(CMR)将其形式化,CMR是相对于干净记忆的预期下游损失的成对增加。我们引入了MemHarm,它预先声明一个有限的稀疏、有依据的语义编辑类,通过正常的智能体记忆接口使用离线成对损失反馈来评估候选,并在该类内认证已解析的选择。与攻击成功优化相比,CMR引导的选择产生显著更大的下游损失,同时保留大部分成功率增益。在两个智能体基准和多样的记忆设计上,MemHarm在相同支持集上评估的通用攻击中达到了最高的CMR点估计。因子移除干预将这种危害与所选的语义因子联系起来,原生智能体部署验证了从写入到新进程的攻击路径。
cs.AI / 278 / 2609.34134

StateGuard: Analytical-State Management with Validity-Aware Intervention for Long-Horizon Data Agents

StateGuard:面向长时域数据智能体的分析状态管理与有效性感知干预
Liao, Wenle, Wang, Zhao, Zhang, Jingchao, Jin, Jiajie, Xu, Yimeng, Dou, Zhicheng
Abstract
LLM-based agents have shown strong capabilities in automated data analysis and are increasingly moving toward long-horizon, multi-stage analytical workflows. However, as the analytical process evolves, constraints, variables, and conclusions remain implicitly embedded in interaction histories, making it difficult for agents to track which analytical artifacts remain valid over increasingly long horizons and changing dependencies. Consequently, stale artifacts may be silently inherited, propagating errors to downstream stages. To address this challenge, we propose StateGuard, an analytical-state validity management framework for long-horizon data agents. StateGuard externalizes evolving analytical progress into a state graph containing constraints, versioned variables, intermediate conclusions, and cross-state relations, treating each state as an executable, verifiable, and traceable object rather than textual memory alone. StateGuard maintains state validity through evidence-grounded verification and hierarchical intervention. To equip StateGuard with these capabilities, we first introduce Manager-Oriented Counterfactual Supervision, which constructs 3K state-centric trajectories through counterfactual runtime synthesis to fine-tune StateGuard for state maintenance, verification, and repair. We then apply Validity-Guided Policy Optimization, using runtime validity evidence to provide fine-grained learning signals for protocol correctness, state grounding, and intervention quality. Experiments on three diverse long-horizon data-analysis benchmarks show that StateGuard consistently improves data-agent performance while reducing dependency-induced downstream error propagation, demonstrating the advantages of explicit analytical-state management for reliable long-horizon data analysis.
Chinese Translation
基于LLM的智能体在自动化数据分析中展现出强大能力,并日益向长时域、多阶段的分析工作流发展。然而,随着分析过程的演进,约束、变量和结论仍隐式地嵌入在交互历史中,使得智能体难以追踪哪些分析产物在日益延长的时域和变化的依赖关系下仍然有效。因此,过时的产物可能被静默继承,将错误传播到下游阶段。为解决这一挑战,我们提出StateGuard,一个面向长时域数据智能体的分析状态有效性管理框架。StateGuard将演进中的分析进展外化为一个状态图,其中包含约束、版本化变量、中间结论和跨状态关系,将每个状态视为可执行、可验证和可追踪的对象,而不仅仅是文本记忆。StateGuard通过基于证据的验证和分层干预来维护状态有效性。为使StateGuard具备这些能力,我们首先引入Manager-Oriented Counterfactual Supervision(面向管理者的反事实监督),它通过反事实运行时合成构建了3K条以状态为中心的轨迹,以微调StateGuard进行状态维护、验证和修复。然后我们应用Validity-Guided Policy Optimization(有效性引导的策略优化),利用运行时有效性证据为协议正确性、状态接地和干预质量提供细粒度的学习信号。在三个不同的长时域数据分析基准上的实验表明,StateGuard持续提升数据智能体的性能,同时减少依赖导致的下游错误传播,证明了显式分析状态管理对于可靠的长时域数据分析的优势。
cs.AI / 279 / 2609.34135

Evo2Team: When Do Evolved Skills Transfer? From Selection to Deployment

Evo2Team:进化技能何时迁移?从选择到部署
Wang, Renxiang, Cui, Jiaming
Abstract
A skill bank that helps one multi-agent system may leave another's behavior unchanged. A transferred rule helps only when target agents act on it successfully. We study this path for routing and communication skills in Count-Frequency and AgentsNet, using teams of 4--32 agents and GPT and Qwen model ladders. Source evolution meets a joint quality, cost, model-tier, and confirmation goal in 14 of 16 settings. We then evaluate Evo2Team, which selects, adapts, and confirms source skills for the target team, alongside six frozen selectors across 28 transfer directions. Evo2Team's target-side exploration cost is below that of evolving a new target bank in every direction, even when reused reference evaluations are charged once. Twenty of 28 held-out outcomes meet the positive-transfer criterion, including three saved diagnostic tests. Selection alone does not explain these outcomes: KNN and CORAL choose different banks in two AgentsNet directions but produce identical recorded executions. When Evo2Team changes execution, gains can reach many tasks, as in a Count-Frequency direction that improves 28 of 32 tasks over KNN. Seven positive AgentsNet outcomes save 6.1--14.6\% in deployment cost while using transferred skills on only three to six of fifteen tasks. In five earlier accepted directions, all 22 task records using transferred skills pass three fixed-graph confirmations, but four fail in recorded executions on new graphs. Graphs and model responses change together in this comparison. These results show that skill transfer must be assessed through the actions agents take, the tasks those actions reach, and the quality and cost of the final deployment.
Chinese Translation
一个帮助某个多智能体系统的技能库可能会使另一个系统的行为保持不变。只有当目标智能体成功执行迁移的规则时,它才起到帮助作用。我们在 Count-Frequency 和 AgentsNet 中研究路由和通信技能的这条路径,使用 4--32 个智能体的团队以及 GPT 和 Qwen 模型阶梯。源进化在 16 种设置中有 14 种达到了质量、成本、模型层级和确认的联合目标。然后我们评估 Evo2Team,它为目标团队选择、适应和确认源技能,并与六个冻结选择器在 28 个迁移方向上进行对比。Evo2Team 的目标侧探索成本在每个方向上均低于进化新目标技能库的成本,即使复用的参考评估仅计费一次。28 个留出结果中有 20 个满足正迁移标准,包括三项节省的诊断测试。仅选择并不能解释这些结果:KNN 和 CORAL 在两个 AgentsNet 方向上选择了不同的技能库,但产生了相同的记录执行。当 Evo2Team 改变执行时,增益可以覆盖许多任务,例如在一个 Count-Frequency 方向上比 KNN 改进了 32 个任务中的 28 个。七个正向 AgentsNet 结果在部署成本上节省了 6.1%--14.6%,同时仅在十五个任务中的三到六个任务上使用迁移技能。在五个早先接受的方向中,所有使用迁移技能的 22 条任务记录都通过了三个固定图的确认,但其中四个在新图上的记录执行中失败。在此比较中,图和模型响应同时变化。这些结果表明,技能迁移必须通过智能体采取的行动、这些行动所触及的任务以及最终部署的质量和成本来评估。
cs.AI / 280 / 2609.34136

Waggle: Learning One Anonymous Local Law for Self-Organizing LLM Swarms

Waggle:学习一种匿名局部法则用于自组织LLM集群
Zou, Mingxi, Zhu, Wei, Wang, Zhuo, Liang, Langzhang, Tang, Zhiwen, Xu, Yinghui, Xu, Zenglin
Abstract
As LLM agents increasingly collaborate on complex tasks, how to organize their interactions becomes a central design question. Existing multi-agent systems typically learn or adapt explicit roles, hierarchies, routing policies, or communication topologies. We shift the learning target to a reusable local law that can be shared across interchangeable agents and adapt coordination as populations or interaction conditions change, without redefining a global organization. We introduce Waggle, a shared anonymous policy over bounded local views that jointly selects task actions, semantic communication, and local commitment updates. Repeated execution of the same law allows coordination to form, persist, and reorganize online without explicit roles or global topology. To learn this law across interchangeable agents and evolving coordination, we develop Swarm-Consistent Distillation (SCD), combining anonymous-orbit consistency with rollout-grounded prediction of the next local coordination field, with no added inference-time components. Across diverse coordination settings, the same learned law remains effective as populations and interaction budgets change, retains over 96% of substrate-specific oracle quality, and transfers without retraining; SCD further improves reorganization after counterevidence. Together, these results show that LLM-agent organization can emerge and adapt through repeated execution of a learned local law.
Chinese Translation
随着LLM智能体越来越多地协作处理复杂任务,如何组织它们的交互成为一个核心设计问题。现有多智能体系统通常学习或适应显式的角色、层级、路由策略或通信拓扑。我们将学习目标转移到一个可复用的局部法则,该法则可在可互换的智能体之间共享,并随着群体或交互条件的变化调整协调,而无需重新定义全局组织。我们提出Waggle,一种在有限局部视图上的共享匿名策略,它联合选择任务动作、语义通信和局部承诺更新。同一法则的重复执行允许协调在线形成、持续和重组,而无需显式角色或全局拓扑。为了在可互换的智能体和演化协调中学习该法则,我们开发了Swarm-Consistent Distillation (SCD),它将匿名轨道一致性与基于rollout的下一局部协调场预测相结合,且不增加推理时组件。在不同的协调设置中,随着群体规模和交互预算的变化,同一学习到的法则仍然有效,保留了超过96%的基底特定的oracle质量,并在无需重新训练的情况下迁移;SCD进一步改善了反证后的重组。总之,这些结果表明,LLM智能体组织可以通过重复执行学习到的局部法则而涌现并适应。
cs.AI / 281 / 2609.34137

You Can't Have It Both Ways: Concept Entanglement Limits Diffusion Model Unlearning

不能两全其美:概念纠缠限制了扩散模型的遗忘
Wang, Yian, Ebrahimpour-Boroojeny, Ali, Sundaram, Hari, Chandrasekaran, Varun
Abstract
Concept unlearning in text-to-image diffusion models aims to suppress a target concept (e.g., \texttt{horse}) while preserving related but distinct content (e.g., \texttt{donkey}), yet existing methods either leak under indirect prompts or visibly degrade other concepts. We show that these failure modes stem from the geometry of concept representations rather than from any particular algorithm. Formalizing concepts as activation-space regions, we prove that the overlap between a target and other concepts lower-bounds the damage any robust erasure must inflict on them, with the trade-off scaling linearly in the degree of overlap. Across thirteen unlearning methods, including methods designed to preserve non-target concepts, no method achieves both strong erasure and strong neighbor preservation: STEREO nearly eliminates indirect leakage but cuts neighbor generation by more than 75\%, while sparse inference-time methods preserve neighbors but leak. Damage increases with our overlap measure, monotonically so for STEREO; the $\kappa$-scaling reproduces on SDXL, and neighbor-selective damage recurs on FLUX. Perfect unlearning is the wrong target for entangled concepts; methods should be evaluated on the Pareto frontier our theorem establishes.
Chinese Translation
文本到图像扩散模型中的概念遗忘旨在抑制目标概念(例如,马)同时保留相关但不同的内容(例如,驴),然而现有方法要么在间接提示下泄漏,要么明显损害其他概念。我们表明,这些失败模式源于概念表示的几何结构,而非任何特定算法。将概念形式化为激活空间区域,我们证明目标与其他概念之间的重叠为任何鲁棒擦除必须对它们造成的损害提供了下界,且权衡随重叠程度线性缩放。在十三种遗忘方法(包括旨在保留非目标概念的方法)中,没有一种方法能同时实现强擦除和强邻居保留:STEREO几乎消除了间接泄漏,但将邻居生成削减了超过75%,而稀疏推理时方法保留了邻居但存在泄漏。损害随着我们的重叠度量而增加,对STEREO而言是单调的;κ缩放在SDXL上复现,邻居选择性损害在FLUX上重现。完美遗忘对于纠缠概念是错误的目标;方法应在我们的定理建立的帕累托前沿上进行评估。
cs.AI / 282 / 2609.34139

Same Tasks, Different Apps: Why Mobile GUI Agents Fail to Generalize?

相同的任务,不同的应用:移动GUI智能体为何无法泛化?
Tran, Tien, Koh, Namho, Matsunaga, Daiki E., Jain, Ayush, Kim, Kee Eung
Abstract
Mobile GUI agents deployed in real settings must work across different applications that support the same functionality. Most existing benchmarks test each task in only one app, so a high score can mean the agent understands the task, or only that it knows that particular app. We introduce AnyAppBench, a category-controlled live Android benchmark that evaluates cross-application generalization while keeping the user goal fixed. It spans 10 functional categories, 100 task templates, and 520 task--application pairs over 52 applications. Agents run from raw instructions and with app-independent sub-goals, and a VLM judge labels every failed run under a fixed failure taxonomy whose reliability is measured by human annotation. We find that, across 13 agents, success on the original application does not transfer reliably to new applications with the same goal. Furthermore, providing high-level sub-goal decomposition produces only small, category-dependent changes that do not close the gap, and the mix of failure types changes with the target interface. Based on those insights, we believe the AnyAppBench benchmark provides an important stepping stone toward robust real-world deployment of mobile GUI agents. Our code, data and the leaderboard can be found at the project website https://anyappbench.github.io/.
Chinese Translation
在实际环境中部署的移动GUI智能体必须跨支持相同功能的不同应用工作。大多数现有基准仅在单个应用中测试每个任务,因此高分可能意味着智能体理解了任务,也可能仅仅表示它熟悉该特定应用。我们提出了AnyAppBench,一个类别受控的实时Android基准,在保持用户目标固定的同时评估跨应用泛化能力。它涵盖10个功能类别、100个任务模板以及跨52个应用的520个任务-应用对。智能体从原始指令和与应用无关的子目标运行,VLM评判器根据固定的失败分类法对每次失败的运行进行标注,该分类法的可靠性通过人工标注来衡量。我们发现,在13个智能体中,在原始应用上的成功并不能可靠地迁移到具有相同目标的新应用。此外,提供高级子目标分解只能产生微小的、依赖于类别的变化,无法弥合差距,并且失败类型的组合随目标界面而变化。基于这些见解,我们认为AnyAppBench基准为移动GUI智能体在现实世界中的稳健部署提供了重要的垫脚石。我们的代码、数据和排行榜可在项目网站 https://anyappbench.github.io/ 找到。
cs.AI / 283 / 2609.34151

Self-Evolving Agents via Likelihood-Guided Tool-Space Optimization

基于似然引导的工具空间优化的自进化智能体
Zhang, Xuanqi, Jin, Ruinan, Yang, Running, Zhang, Yuxuan, Chen, Minghui, Deng, Wenlong, Li, Xiaoxiao
Abstract
Self-evolving agents can continually improve their behavior, while tools define the executable action space through which they interact with the environment. However, exposing the full tool library to model introduces substantial irrelevant context and can impair tool-use decisions. We study tool-space self-evolution, where each recurring task type maintains a persistent tool space which is constructed from accumulated output experience. We identify three limitations of existing methods: (1) output-unaware selection: they rely primarily on tool descriptions or model priors rather than observed tool outputs; (2) statelessness across request: they select tools independently for each request without consolidating prior output experience into persistent task-specific state; (3) inference cost: they repeatedly search, rank, or reason over candidate tools for subsequent requests of the same task. We address these limitations through output-aware tool scoring, persistent task-specific tool spaces, amortized tool selection, and reusable configurations across models. We introduce LOTS (Likelihood-Only Tool Scoring), which evolves an agent's tool space from accumulated output experience while keeping model parameters fixed. After each request, LOTS holds the model's generated answer and estimates each tool's contribution by measuring how much the answer likelihood changes when its observed output is removed. These contributions are aggregated within each recurring task to rank tools and update its persistent space. Across three benchmarks, LOTS improves task performance while substantially reducing tool context. More importantly, sequential experiments demonstrate that task-specific spaces persist and continue to improve over time, while cross-model experiments show that learned configurations transfer across different models.
Chinese Translation
自进化智能体能够持续改进其行为,而工具定义了它们与环境交互的可执行动作空间。然而,将完整的工具库暴露给模型会引入大量无关上下文,并可能损害工具使用决策。我们研究工具空间自进化,其中每种重复出现的任务类型维护一个持久化的工具空间,该空间由累积的输出经验构建。我们指出现有方法的三个局限性:(1) 输出无感知的选择:它们主要依赖工具描述或模型先验,而非观察到的工具输出;(2) 请求间无状态:它们为每个请求独立选择工具,而没有将先前的输出经验整合为持久的任务特定状态;(3) 推理成本:对于同一任务的后续请求,它们会重复搜索、排序或推理候选工具。我们通过输出感知的工具评分、持久的任务特定工具空间、摊销的工具选择以及跨模型的可重用配置来解决这些局限性。我们引入 LOTS(Likelihood-Only Tool Scoring,仅似然工具评分),它从累积的输出经验中进化智能体的工具空间,同时保持模型参数固定。在每个请求之后,LOTS 保留模型生成的答案,并通过测量移除观察到的输出时答案似然的变化来估计每个工具的贡献。这些贡献在每个重复任务内聚合,以对工具进行排序并更新其持久空间。在三个基准测试中,LOTS 提高了任务性能,同时大幅减少了工具上下文。更重要的是,序列实验表明,任务特定空间会持续存在并随时间不断改进,而跨模型实验表明,学习到的配置可以迁移到不同模型。
cs.AI / 284 / 2609.34157

TableSeek: Structure-Preserving Agentic Evidence Seeking over Heterogeneous Table Corpora

TableSeek: 面向异构表格语料库的结构保持式智能体证据搜索
Tian, Jiaming, Li, Liyao, Ye, Wentao, Wang, Haobo, Yu, Lihua, Ren, Zujie, Chen, Gang, Zhao, Junbo
Abstract
Open-domain table retrieval seeks tables that contain sufficient evidence for answering a question or verifying a claim. Yet semantic relevance is often misleading: topically similar tables may lack the required facts, while answer-bearing evidence is often confined to a few cells whose meaning depends on surrounding schema and table context. Heterogeneous schemas, value formats, and serializations further weaken one-shot matching. We present TableSeek, a structure-preserving agentic search framework for heterogeneous table corpora. Instead of ranking tables once, an LLM agent iteratively follows sparse clues, inspects schema-preserving previews, identifies schema- and value-level mismatches, and refines its investigation. TableSeek uses cells and schemas as evidence anchors while retaining complete tables as evidence units, enabling fine-grained localization without losing the context required for interpretation and answerability checking. Without relying on retriever training or a precomputed semantic index, TableSeek produces transparent evidence-seeking trajectories and achieves competitive end-to-end performance against strong retrieval-and-reranking pipelines on heterogeneous table benchmarks. These results suggest that active, structure-preserving evidence seeking is a promising paradigm for open-domain table retrieval.
Chinese Translation
开放域表格检索旨在找到包含足够证据以回答一个问题或验证一个主张的表格。然而,语义相关性常常具有误导性:主题相似的表格可能缺乏所需的事实,而含有答案的证据往往局限于少数单元格,其含义取决于周围的模式和表格上下文。异构的模式、值格式和序列化进一步削弱了一次性匹配。我们提出TableSeek,一种面向异构表格语料库的结构保持式智能体搜索框架。不同于一次性对表格进行排序,LLM智能体迭代地跟随稀疏线索,检查保留模式的预览,识别模式级和值级的不匹配,并细化其调查。TableSeek使用单元格和模式作为证据锚点,同时保留完整表格作为证据单元,从而实现细粒度定位,而不会丢失解释和可回答性检查所需的上下文。在不依赖检索器训练或预计算的语义索引的情况下,TableSeek生成透明的证据搜索轨迹,并在异构表格基准上相较于强大的检索-重排序流水线取得了有竞争力的端到端性能。这些结果表明,主动的、结构保持的证据搜索是开放域表格检索的一种有前景的范式。
cs.AI / 285 / 2609.34160

RoutePrism: Tracing Construction Order Effects in Agent Memory

RoutePrism:追踪智能体记忆中的构建顺序效应
Xu, Dong, Yang, Zhangfan, Wu, Jiantao, Zhang, Shipeng, Zhu, Zexuan, Li, Jiangqiang, Zhang, Jun, Ji, Junkai
Abstract
Processing the same records in a different order can discard different evidence, yet endpoint accuracy alone cannot reveal what changed or whether it mattered. We introduce RoutePrism, a diagnostic protocol that builds memory twice from the same source pool in two processing orders, then traces which sources, compiled contexts, and answers differ. Because record content, timestamps, policy, and the answer model all stay fixed, any observed difference is localized to the memory construction step. A matched four-condition intervention tests whether a record displaced by reordering actually carried task-relevant evidence: restoring that single record recovers over 60 percentage points of lost accuracy, while substituting a non-supporting record of equal length does not. We evaluate the protocol on PersonaMem-32K (63 primary queries, 29 users) and 470 LongMemEval-S questions with histories spanning 38 to 62 sessions, replicating the core intervention across five answer models. Survivor selection, defined as the choice of which record a cluster retains, drives most source-level changes, while different memory policies (compaction, bounded recency, MemoChat-style summarization, A-MEM) produce distinct failure signatures at the source, context, and metadata layers.
Chinese Translation
以不同顺序处理相同的记录可能会丢弃不同的证据,然而仅凭最终准确率无法揭示究竟改变了什么,或者这种改变是否重要。我们提出 RoutePrism,一种诊断协议,它从同一源池出发,以两种处理顺序构建两次记忆,然后追踪哪些来源、编译后的上下文和答案存在差异。由于记录内容、时间戳、策略和答案模型均保持不变,任何观察到的差异都被定位到记忆构建步骤。一个匹配的四条件干预测试了因重排序而被替代的记录是否确实携带了任务相关证据:恢复该单条记录可挽回超过60个百分点的准确率损失,而用等长的非支持性记录替代则不能。我们在 PersonaMem-32K(63个主要查询,29个用户)和470个历史跨度为38至62个会话的 LongMemEval-S 问题上评估该协议,并在五个答案模型上复现了核心干预。幸存者选择,即聚类保留哪条记录的选择,驱动了大多数来源层面的变化,而不同的记忆策略(压缩、有界近因、MemoChat风格摘要、A-MEM)在来源、上下文和元数据层面产生不同的失败特征。
cs.AI / 286 / 2609.34177

ReplayLens: Auditing Agents' Use of Outcomes

ReplayLens:审计智能体对结果的使用
Xu, Dong, Yang, Zhangfan, Wu, Jiantao, Zhang, Shipeng, Zhu, Zexuan, Li, Jiangqiang, Zhang, Jun, Ji, Junkai
Abstract
When an agent reuses logged experience, a changed decision may reflect the recorded score, the action's name, or the record's position in storage. Standard memory evaluations do not reveal which relationship drives that change. We introduce ReplayLens, a black-box audit that changes one relationship in the stored history at a time, holds the remaining interface fixed, and measures the resulting decision. Four interventions target four relationships. Outcome reassignment swaps which scores belong to which actions. Pair transport moves intact action-score pairs to new record slots. Consistent renaming relabels actions in both history and menu. Key-slot reassignment changes both score attachment and position. A constructive separation shows why the audit is needed: two memory writers with identical endpoint accuracy respond differently to the same replay, so conventional evaluation cannot resolve the underlying dependence. On black-box LLM interfaces, swapping scores changes decisions while moving intact pairs does not, separating score attachment from record order. A bounded-memory study exposes ingestion-order sensitivity that endpoint comparison misses. In sequential experiment planning, altered historical scores redirect exploration and reduce final utility despite fresh measurements. A code-debugging agent with sealed hidden tests shows the same pattern outside model selection. ReplayLens provides a relationship-level audit for deciding whether logged experience can be merged, reordered, or reindexed safely.
Chinese Translation
当智能体重用日志经验时,决策的改变可能反映的是记录的分数、动作的名称,或记录在存储中的位置。标准记忆评估无法揭示是哪种关系驱动了这种改变。我们提出ReplayLens,一种黑盒审计方法,它一次改变存储历史中的一种关系,保持其余接口固定,并测量由此产生的决策。四种干预针对四种关系。结果重分配交换分数与动作的归属关系。配对迁移将完整的动作-分数对移动到新的记录槽位。一致重命名在历史和菜单中重新标记动作名称。键槽重分配同时改变分数附着和位置。一个构造性分离显示了为什么需要这种审计:两个具有相同端点准确率的记忆写入器对相同的重放响应不同,因此传统评估无法解决潜在的依赖关系。在黑盒LLM接口上,交换分数会改变决策,而移动完整配对则不会,从而将分数附着与记录顺序分离。一项有限记忆研究揭示了端点比较所忽略的摄取顺序敏感性。在顺序实验规划中,改变的历史分数会重新引导探索并降低最终效用,尽管有新的测量。一个带有密封隐藏测试的代码调试智能体在模型选择之外显示出相同的模式。ReplayLens提供了一种关系级别的审计,用于决定日志经验是否可以安全地合并、重新排序或重新索引。
cs.AI / 287 / 2609.34179

RAGWarrant: Evidence-Preserving Governance for RAG Policy Promotion Under Quality, Cost, Latency, and Risk Constraints

RAGWarrant:在质量、成本、延迟和风险约束下针对RAG策略推广的证据保留治理
Krueger, Richard, Krause, Lucas, Pocquette, Zach
Abstract
Retrieval-augmented generation systems are extensively instrumented with metrics, benchmarks, traces, and automated judges, but these tools do not decide whether a proposed policy change is safe to release. We present RAGWarrant, an open-source promotion-control framework that treats deployment as a constrained evidence decision rather than a leaderboard choice. RAGWarrant normalizes evaluator outputs and operational telemetry, applies predeclared quality and hard-risk gates, assigns evidence-class claim ceilings, preserves negative outcomes, and emits auditable PROMOTE, BLOCK, REJECT, or INCONCLUSIVE decisions. We evaluate the framework across T2-RAGBench, MultiHop-RAG, CRAG, HotpotQA, synthetic reproduction, and bounded local generative experiments. On HotpotQA, operational savings were blocked because answer quality fell beyond the declared margin. A bounded CRAG study selected a lower-cost quality-tied policy, but related generative gains were unstable and a held-out guardrail failed closed. We claim an auditable promotion-control abstraction, not optimizer superiority, human validation, or production readiness. The tagged artifact reproduces from a fresh clone, runs as a hardened Docker job, accepts external evaluator exports, and verifies artifact integrity.
Chinese Translation
检索增强生成系统被广泛地配备度量指标、基准、追踪和自动评判器,但这些工具并不决定所提议的策略变更是否可以安全发布。我们提出RAGWarrant,一个开源的推广控制框架,将部署视为受约束的证据决策,而不是排行榜选择。RAGWarrant对评估器输出和运营遥测进行规范化,应用预先声明的质量和硬风险门控,分配证据类主张上限,保留负面结果,并发出可审计的PROMOTE、BLOCK、REJECT或INCONCLUSIVE决策。我们在T2-RAGBench、MultiHop-RAG、CRAG、HotpotQA、合成复现以及有界本地生成实验上评估了该框架。在HotpotQA上,运营节省被阻止,因为答案质量下降超出了声明的边际。一项有界的CRAG研究选择了一个成本更低且质量相当的政策,但相关的生成增益不稳定,并且一个留出防护栏故障关闭。我们主张一种可审计的推广控制抽象,而非优化器优越性、人工验证或生产就绪性。标记的工件可从全新克隆复现,作为加固的Docker作业运行,接受外部评估器导出,并验证工件完整性。
cs.AI / 288 / 2609.34180

Decision Readouts for Text-Mediated Video Anomaly Detection: An Exploratory Evaluation of Jev and Qwen

文本中介视频异常检测的决策读出:Jev 与 Qwen 的探索性评估
Qin, Xukui, Wang, Youting, He, Xinjie, Luo, Ziyang, Wu, Runxiong, Chen, Yan-Syuan, Chu, Zhongyao
Abstract
How much does the decision readout matter when video-derived textual evidence is held fixed? We evaluate Jev typed decisions and three Qwen readouts on a sparse development sample of 40 videos and 400 target anchors from UCF-Crime and XD-Violence, each presented as a summary and ordered captions. Each dataset contributes 20 source groups and 200 anchors, including only 10 and 37 positives, respectively. The original five-backend pilot requested 4,000 predictions; Jev Choice returned 776 valid responses out of 800 under the study's strict numerical policy, blocking its full-coverage quality comparison. On XD captions, Jev Noul achieved 75.99% average precision versus 48.47% for Qwen generated probability and 57.81% for the stronger local ordinal-likelihood expectation. The latter paired difference was 18.18 percentage points (95% source-group bootstrap interval 5.53-31.50). UCF did not show a corresponding advantage: caption ROC-AUC was 52.26% for Noul and 65.95% for ordinal likelihood. Both probability readouts had higher, hence worse, UCF Brier scores than the evaluation-prevalence reference of 0.0475. We additionally audit historical LAVAD scores at exactly matched anchors and distinguish response structure from numerical consistency. A binary-likelihood control is missing. These exploratory offline results characterize ranking, probability quality and interface failures; they establish neither a causal typed-interface benefit nor general superiority, calibration or end-to-end acceleration.
Chinese Translation
当视频衍生的文本证据固定时,决策读出有多重要?我们在来自 UCF-Crime 和 XD-Violence 的 40 个视频和 400 个目标锚点的稀疏开发样本上,评估了 Jev 类型化决策和三种 Qwen 读出,每个样本以摘要和有序字幕呈现。每个数据集贡献 20 个源组和 200 个锚点,其中分别仅包含 10 个和 37 个正例。原始的五后端试点请求了 4,000 个预测;在研究严格的数值策略下,Jev Choice 在 800 个请求中返回了 776 个有效响应,阻碍了其全覆盖质量比较。在 XD 字幕上,Jev Noul 达到了 75.99% 的平均精度,而 Qwen 生成概率为 48.47%,更强的局部有序似然期望为 57.81%。后者的配对差异为 18.18 个百分点(95% 源组自举区间 5.53-31.50)。UCF 没有显示出相应的优势:字幕 ROC-AUC 对于 Noul 为 52.26%,对于有序似然为 65.95%。两个概率读出的 UCF Brier 分数均高于评估流行率参考值 0.0475,因此更差。我们还审计了精确匹配锚点上的历史 LAVAD 分数,并区分响应结构与数值一致性。缺少二元似然对照。这些探索性离线结果描述了排序、概率质量和接口故障;它们既没有建立因果类型化接口优势,也没有建立普遍优越性、校准或端到端加速。
cs.AI / 289 / 2609.34181

Efficient Reasoning via Constrained Optimization in Latent Space

潜在空间中基于约束优化的高效推理
Hou, Zhinan, Li, XingChen, You, Keyou
Abstract
Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt to mitigate this issue but inevitably truncate necessary steps and induce underthinking, thereby compromising performance. To address this dilemma, we investigate the latent representations and observe that efficient reasoning steps naturally cluster into a concentrated region in latent space, while those deviating from this region tend to produce verbose sequences. To leverage this, we keep reasoning focused within this region via a quadratic program which projects deviating hidden states back into the region. Then we propose a novel training-free framework to achieve efficient reasoning that reduces token generation costs without sacrificing performance. Extensive experiments conducted on four models ranging from 1.5B to 14B, and across six benchmarks in math reasoning, coding, and scientific QA, validate the effectiveness of our method, up to a 12.1\% improvement in accuracy while reducing generated tokens by 11.8\% to 52.8\%. Codes are available at \href{https://github.com/hzn18/Opt4Reasoning}{https://github.com/hzn18/Opt4Reasoning}.
Chinese Translation
大型推理模型(LRMs)已展现出卓越的推理能力,然而它们仍然存在过度思考的问题,生成冗余的推理步骤,导致大量的token消耗。现有方法,如抑制反思性关键词或强制缩短推理长度,试图缓解这一问题,但不可避免地截断了必要的步骤并引发思考不足,从而损害了性能。为了解决这一困境,我们研究了潜在表示,并观察到高效的推理步骤自然地聚集在潜在空间中的一个集中区域,而那些偏离该区域的步骤往往会产生冗长的序列。为了利用这一点,我们通过一个二次规划将偏离的隐藏状态投影回该区域,从而将推理聚焦在该区域内。然后,我们提出了一种新颖的免训练框架,以实现高效推理,在不牺牲性能的情况下降低token生成成本。在四个模型(规模从1.5B到14B)上进行的广泛实验,以及跨越数学推理、编码和科学问答的六个基准测试,验证了我们方法的有效性,准确率最高提升了12.1%,同时生成的token减少了11.8%至52.8%。代码可在 https://github.com/hzn18/Opt4Reasoning 获取。
cs.AI / 290 / 2609.34184

CASS: Contribution-Aware Structured Sparsity for Model Merging

CASS:面向模型合并的贡献感知结构化稀疏
Li, Yan, Cao, Guiping, Xu, Meng, Jiang, Tao, Song, Yaguang, Tao, Ming, Wang, Yaowei, Jiang, Dongmei
Abstract
Model merging integrates task-specific fine-tuned models into a single multi-task model, but often suffers from parameter interference caused by conflicting task-vector updates. Existing methods typically mitigate conflicts by pruning task vectors based on weight magnitude or random heuristics, treating Transformers as unstructured ``bags of parameters'' and overlooking their inherent modularity. In this paper, we propose \textbf{C}ontribution-\textbf{A}ware \textbf{S}tructured \textbf{S}parsity (CASS), a unified framework that reduces parameter interference by identifying and preserving task-specific components. At the core of CASS is a contribution-aware structured mask that identifies task-relevant attention heads and FFN neurons. We instantiate this mask in two settings: CASS-Merging, the primary post-hoc setting where masks serve as a plug-and-play denoising filter for existing merging operators, and CASS-Tuning, an extension for scenarios with fine-tuning access where masks constrain gradients to reduce structural overlap between task vectors. Our analysis shows that task-relevant components are sparse and partially disjoint, supporting structured component-level filtering as an effective way to reduce merging interference. Extensive experiments across vision (ViT, 20 tasks) and language (RoBERTa, 8 tasks; Qwen2.5, 4 tasks) benchmarks demonstrate that CASS improves a range of representative merging baselines.
Chinese Translation
模型合并将任务特定的微调模型集成到一个多任务模型中,但常常受到由冲突的任务向量更新引起的参数干扰。现有方法通常通过基于权重幅度或随机启发式剪枝任务向量来缓解冲突,将Transformer视为无结构的“参数袋”,而忽略了其固有的模块化特性。在本文中,我们提出了贡献感知结构化稀疏(CASS),一个统一的框架,通过识别和保留任务特定的组件来减少参数干扰。CASS的核心是一个贡献感知的结构化掩码,它识别任务相关的注意力头和FFN神经元。我们在两种设置中实例化该掩码:CASS-Merging,主要的后处理设置,其中掩码作为现有合并算子的即插即用去噪滤波器;以及CASS-Tuning,一种针对具有微调访问权限的场景的扩展,其中掩码约束梯度以减少任务向量之间的结构重叠。我们的分析表明,任务相关的组件是稀疏的且部分不相交的,支持结构化组件级过滤作为减少合并干扰的有效方法。在视觉(ViT,20个任务)和语言(RoBERTa,8个任务;Qwen2.5,4个任务)基准上的大量实验表明,CASS改善了一系列代表性的合并基线。
cs.AI / 291 / 2609.34195

PainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language Models

PainterBench:面向工具使用语言模型的图形发散思维基准测试
Hettige, Shane K. A. Dalumura, Oppenlaender, Jonas
Abstract
Figural divergent thinking is the ability to develop a given shape fragment into an original drawing. In humans, this ability is assessed with incomplete-drawing tasks. We introduce PainterBench, a benchmark that ports the incomplete-drawing task to the agentic setting. The agent draws on a canvas through tool calls and observes the result after every turn. The canvas includes a starting shape which cannot be erased, and the agent's goal is to incorporate this shape into the most original drawing it can produce. The task is open-ended, and the agent itself decides when the drawing is finished. The benchmark tests incremental visual planning over a short horizon and the transfer of creative ability from pretraining to multi-turn tool use. We evaluate 14 multimodal language models from small to frontier scale. Across the primary study and six sensitivity analyses, we collect 2,700 drawings and crowdsource creativity and recognizability ratings for every drawing and for 300 human reference drawings. We also present ViDrA-adapted, an automated scorer that predicts human creativity ratings of agent drawings (r = 0.85 on random held-out test split). Figural divergent thinking varies widely across the 14 models, and GPT-6 Astra produces the most creative drawings. Relative to the human drawings, the agent drawings score higher in creativity but lower in recognizability. We release the final drawings, per-round canvas snapshots, tool call traces, stimulus bank, benchmark harness, crowdsourced ratings (N = 72,000), and ViDrA checkpoint.
Chinese Translation
图形发散思维是将给定的形状片段发展成原创绘画的能力。在人类中,这种能力通过不完整绘画任务进行评估。我们提出 PainterBench,一个将不完整绘画任务移植到智能体场景中的基准测试。智能体通过工具调用在画布上绘画,并在每一轮后观察结果。画布包含一个无法擦除的起始形状,智能体的目标是将这个形状融入它能创作的最原创的绘画中。任务是开放式的,智能体自行决定绘画何时完成。该基准测试测试短时间范围内的增量视觉规划,以及创造能力从预训练到多轮工具使用的迁移。我们评估了 14 个从小型到前沿规模的多模态语言模型。在主要研究和六项敏感性分析中,我们收集了 2,700 幅绘画,并通过众包对每幅绘画以及 300 幅人类参考绘画的创造力和可识别性进行评分。我们还提出了 ViDrA-adapted,一种自动评分器,可预测智能体绘画的人类创造力评分(在随机留出测试集上 r = 0.85)。图形发散思维在 14 个模型中差异很大,GPT-6 Astra 产生了最具创造力的绘画。相对于人类绘画,智能体绘画在创造力上得分更高,但在可识别性上得分更低。我们发布了最终绘画、每轮画布快照、工具调用轨迹、刺激库、基准测试框架、众包评分(N = 72,000)以及 ViDrA 检查点。
cs.AI / 292 / 2609.34211

Behavior-Grounded Semantic Enrichment for Financial Fraud Modeling and Reasoning

基于行为的语义增强用于金融欺诈建模与推理
Shao, Linbo, He, Huilin, Lou, Yating, Cheng, Dawei
Abstract
In financial fraud detection, rich semantic context can provide important evidence for transaction behavior modeling and fraud reasoning. However, public real-world financial datasets often lack rich semantics due to privacy constraints. Consequently, synthetic datasets incorporate generated semantics, but at the cost of behavioral realism; textual descriptions for contextual reasoning remain scarce. We address this gap through a semantic enrichment framework grounded in original transaction behavior to simulate multimodal financial data. We (1) propose a multi-agent semantic enrichment framework that generates interpretable financial semantics grounded in transaction behavior through role-specialized agents and consistency refinement, and (2) newly contribute a valuable multimodal financial fraud dataset, MS-FFSD, enriched with structured semantics and textual semantics while preserving real-data-grounded transaction behavior. Furthermore, we systematically analyze the quality and utility of semantic enrichment. Results demonstrate statistical fidelity and framework generalizability, while showing that richer semantics benefit fraud modeling and context-aware LLM reasoning. Overall, this work advances multimodal financial fraud research and bridges emerging LLM and multi-agent capabilities with operational anti-fraud practice. The framework and dataset are released at https://github.com/AI4Risk/MS-FFSD.
Chinese Translation
在金融欺诈检测中,丰富的语义上下文可以为交易行为建模和欺诈推理提供重要证据。然而,受隐私约束影响,公开的真实金融数据集往往缺乏丰富语义。因此,合成数据集虽引入了生成的语义,但牺牲了行为真实性;用于上下文推理的文本描述仍然稀缺。我们通过一个基于原始交易行为的语义增强框架来模拟多模态金融数据,以弥补这一空白。我们(1)提出一个多智能体语义增强框架,通过角色专业化智能体和一致性精炼,生成基于交易行为的可解释金融语义;(2)新贡献一个有价值的多模态金融欺诈数据集 MS-FFSD,其在保留基于真实数据的交易行为的同时,丰富了结构化语义和文本语义。此外,我们系统分析了语义增强的质量与效用。结果表明了统计保真度和框架泛化性,同时表明更丰富的语义有利于欺诈建模和上下文感知的 LLM 推理。总体而言,这项工作推进了多模态金融欺诈研究,并将新兴的 LLM 与多智能体能力同实际反欺诈实践相衔接。框架和数据集已在 https://github.com/AI4Risk/MS-FFSD 发布。
cs.AI / 293 / 2609.34214

GlyphBench: A Playground for Language-Model Reinforcement Learning

GlyphBench:语言模型强化学习的实验场
Castanyer, Roger Creus, Côté, Marc-Alexandre, Sargent, Matthew James, Mavor-Parker, Augustine N., Berseth, Glen, Castro, Pablo Samuel
Abstract
We introduce GlyphBench, an environment suite for reinforcement learning (RL) post-training of language-model agents, with over 360 tasks spanning diverse games. GlyphBench renders spatial observations as two-dimensional Unicode grids and connects training, evaluation, and trajectory replay through a unified interface designed to support efficient and reproducible research. We use GlyphBench to study how observation interfaces, reasoning effort, and agent harnesses affect performance, and how RL configurations shape learning dynamics. Our results show that glyph observations outperform native text and pixels in our Craftax experiments, with further gains on several BALROG environments. RL on 100 GlyphBench tasks improves Qwen3.5-4B on held-out Reasoning Gym problems, reaching 63.48% accuracy and outperforming the base model, a math-trained baseline, and a code-trained baseline. These experiments provide empirical evidence that reasoning gains from gameplay can yield stronger transfer than math or code. Together, these results highlight GlyphBench's value as a testbed for systematic research on how language-model agents learn, interact, and generalize.
Chinese Translation
我们介绍了 GlyphBench,一个用于语言模型智能体强化学习(RL)后训练的环境套件,包含跨越多种游戏的超过 360 个任务。GlyphBench 将空间观察渲染为二维 Unicode 网格,并通过一个统一接口连接训练、评估和轨迹回放,旨在支持高效且可复现的研究。我们使用 GlyphBench 研究观察接口、推理努力和智能体框架如何影响性能,以及 RL 配置如何塑造学习动态。我们的结果表明,在我们的 Craftax 实验中,字形观察优于原生文本和像素,并且在若干 BALROG 环境中取得了进一步提升。在 100 个 GlyphBench 任务上进行 RL 提升了 Qwen3.5-4B 在留出的 Reasoning Gym 问题上的表现,达到 63.48% 的准确率,并优于基础模型、数学训练的基线和代码训练的基线。这些实验提供了经验证据,表明从游戏玩法中获得的推理增益可以比数学或代码产生更强的迁移。总之,这些结果突显了 GlyphBench 作为测试平台的价值,可用于系统研究语言模型智能体如何学习、交互和泛化。
cs.AI / 294 / 2609.34215

Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures

相同的胜者,不同的成功率:评估 LLM 智能体如何从失败中恢复
Xu, Dong, Yang, Zhangfan, Wu, Jiantao, Zhang, Shipeng, Zhu, Zexuan, Li, Jiangqiang, Zhang, Jun, Ji, Junkai
Abstract
Evaluating how LLM agents recover from mid-task failures is central to deploying reliable agentic systems. Existing checkpoint-based benchmarks measure recovery by comparing which action is selected as best across independent runs, a quantity known as set agreement. However, set agreement is a purely ordinal measure that records which action wins without reflecting the absolute level of performance. When all actions fail, they tie at zero reward, and independent runs produce the same tied set with high probability, creating an illusion of stability that masks near-zero recovery success. We formalize this limitation through a set-path symmetry result, proving that for equal-cost Bernoulli actions the success probabilities (0.9, 0.8) and (0.2, 0.1) yield identical best-action-set distributions at every sample size. No procedure based solely on which action wins can distinguish these two regimes. We further prove that certifying exact population ties is impossible in finite time, and that the assignment of outcomes to checkpoints carries information beyond marginal outcome distributions. The pooled success probability is the missing scalar that resolves the ordinal ambiguity. Experiments on 864 frozen RecoveryBench episodes and two planning cohorts totaling 3,456 responses confirm the theoretical predictions. Agreement and held-out quality can move in opposite directions, and permuting checkpoint-to-action bindings changes 8 to 13 percent of cell-level conclusions. Based on these findings, we propose reporting four diagnostic quantities (agreement, all-zero fraction, held-out success, and pooled success) that expose this failure mode with no additional data collection.
Chinese Translation
评估 LLM 智能体如何从任务中途失败中恢复,对于部署可靠的智能体系统至关重要。现有的基于检查点的基准测试通过比较独立运行中被选为最佳的动作来衡量恢复能力,这一量被称为集合一致性(set agreement)。然而,集合一致性是一种纯粹的序数度量,它记录哪个动作胜出,而不反映绝对性能水平。当所有动作都失败时,它们以零奖励并列,独立运行以高概率产生相同的并列集合,从而制造出稳定的假象,掩盖了接近零的恢复成功率。我们通过一个集合-路径对称性结果形式化了这一局限性,证明对于等成本的伯努利动作,成功概率 (0.9, 0.8) 和 (0.2, 0.1) 在任意样本量下都产生相同的最佳动作集合分布。任何仅基于哪个动作胜出的过程都无法区分这两种情形。我们进一步证明,在有限时间内证明精确的总体并列是不可能的,并且将结果分配给检查点所携带的信息超出了边缘结果分布。合并成功率是解决序数模糊性的缺失标量。在 864 个冻结的 RecoveryBench 回合和两个总计 3,456 条响应的规划队列上的实验证实了理论预测。一致性和留出质量可能朝相反方向变化,并且置换检查点到动作的绑定会改变 8% 到 13% 的单元级结论。基于这些发现,我们提出报告四个诊断量(一致性、全零比例、留出成功率和合并成功率),这些量无需额外数据收集即可揭示这种失效模式。
cs.AI / 295 / 2609.34227

When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model

选择何时取代提取?一项基于类型化决策模型的智能体记忆预注册测试
Sharma, Rishabh, Lall, Rishika
Abstract
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound -3.0 points against a -5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061 times less to write, and the result holds with a second answer model. Within this study, reranking's gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound -2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released.
Chinese Translation
对话记忆是否需要LLM提取的事实,还是选择正确的原始对话轮次就足够了?已发表的结果存在分歧。基于提取的系统报告了从蒸馏事实中获得的收益。近期研究发现,具有良好排序的原始历史记录也能达到同样效果,但对于排序是否重要存在分歧。我们在留出的LoCoMo对话和LongMemEval上进行了预注册研究。在LoCoMo的紧张预算下,通过单次调用类型化决策模型Jev选择的原始对话轮次,不劣于LLM提取记忆(单侧95%置信界为-3.0分,非劣效性界值为-5分)。盲人工评分缩小了差距,但并未改变结果。原始对话轮次的写入成本低3,061倍,并且该结果在使用第二个答案模型时依然成立。在本研究中,重排序的收益随着预算的增加而缩小。当保留30个候选中的3个时,它在LoCoMo上增加了17.4分,在LongMemEval上增加了9.1分。在宽松预算下,它分别增加了1.5和1.1分,而提取系统更准确。这解释了已发表结果存在分歧的原因。在匹配上下文下,Jev的选择准确性与LLM重排序器相当(非劣效性界值为-2.0),延迟仅为后者的三分之一,并且比多次调用的图遍历更准确。重排序降低了正确弃权率。研究计划、代码和评分答案均已发布。
cs.AI / 296 / 2609.34241

AdaGuard: An Adaptive Guard Model with User-defined Policies

AdaGuard:一种用户自定义策略的自适应防护模型
Feng, Yunhao, Ding, Yifan, Xie, Yuxiang, Li, Zheng, Lao, Mingrui, Wang, Zeyuan, Guo, Yanming
Abstract
Guard models support the safe deployment of language model agents, but fixed risk taxonomies limit their ability to accommodate requirements that vary across applications and tasks. Under user-defined policies, detecting violations requires interpreting both the applicable rules and the agent's behavior, since identical actions can receive different judgments under different policies. To support learning this capability, we introduce AdaptiveSafety, a dataset of 10,939 training examples and 1,000 test examples covering policies with 1--100 rules. The dataset combines trajectories from multiple sources with policy and behavioral counterfactuals, pairing each example with an explanation and the complete set of violated rules. These counterfactuals expose changes that alter compliance, while structural augmentations provide supervision for consistency under rule reordering and identifier remapping. Building on this supervision, we propose SafePO, a reinforcement learning algorithm for refining violation identification while balancing explanatory reasoning and final verdicts. SafePO uses structured rewards to assess prediction correctness, retains group-relative advantages at the response level, and employs a separately trained value model to modulate token weights within explanation and verdict regions. Separate normalization controls their relative contribution to training despite differences in length. Through supervised initialization followed by SafePO, we develop AdaGuard, a family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time. Our 4B model achieves binary accuracies of 89.30\% on AdaptiveSafety and 71.82\% on DynaBench. The project repository is available at https://github.com/Yunhao-Feng/AdaGuard
Chinese Translation
防护模型支持语言模型代理的安全部署,但固定的风险分类法限制了它们适应不同应用和任务需求的能力。在用户自定义策略下,检测违规行为需要同时解释适用的规则和代理的行为,因为相同的动作在不同策略下可能得到不同的判定。为了支持学习这种能力,我们引入了 AdaptiveSafety,一个包含 10,939 个训练样本和 1,000 个测试样本的数据集,覆盖具有 1-100 条规则的策略。该数据集将来自多个来源的轨迹与策略和行为反事实相结合,为每个样本配以解释和完整的违反规则集合。这些反事实揭示了改变合规性的变化,而结构化增强为规则重排序和标识符重映射下的一致性提供了监督。基于这种监督,我们提出了 SafePO,一种强化学习算法,用于细化违规识别,同时平衡解释性推理和最终裁决。SafePO 使用结构化奖励来评估预测正确性,在响应级别保留组相对优势,并采用单独训练的价值模型来调节解释和裁决区域内的 token 权重。单独归一化控制了它们对训练的贡献比例,尽管长度不同。通过监督初始化后接 SafePO,我们开发了 AdaGuard,一系列 0.6B、4B 和 8B 的防护模型,能够根据推理时提供的策略评估代理轨迹。我们的 4B 模型在 AdaptiveSafety 上实现了 89.30% 的二分类准确率,在 DynaBench 上实现了 71.82% 的准确率。项目仓库位于 https://github.com/Yunhao-Feng/AdaGuard
cs.AI / 297 / 2609.34242

Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents

Stashbird:面向对话代理的高效说话人索引记忆
Biringa, Chidera, Yannul, Lucas, Wang, Xiaowen, Ayala, Marco, Yi, Nicholas, Moyse, Alex, Manchanda, Nishant, Gupta, Vivek
Abstract
AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance. Stashbird organizes memory into episodic records, semantic relations, community summaries, and persisted graph state, with lifecycle operations for incremental updates and episode-level deletion. We evaluate question-answering accuracy and model-facing workload across four long-term memory benchmarks. On LoCoMo, Stashbird uses 76.4x fewer ingestion prompt tokens than Graphiti. Compared with reproduced Hindsight on the same benchmark, it uses 8.1x fewer retrieval prompt tokens, with accuracy 1.6 percentage points lower. It achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.
Chinese Translation
AI代理需要一种记忆,能够在用户-代理交互、用户间对话以及有或没有代理参与的群组对话中保存信息,同时支持在证据变化或被移除时进行更新。我们提出了Stashbird,一个代理记忆系统,它通过显式的溯源信息将源情节与衍生的记忆状态联系起来。Stashbird将记忆组织为情节记录、语义关系、社区摘要和持久化图状态,并具有用于增量更新和情节级删除的生命周期操作。我们在四个长期记忆基准上评估了问答准确率和面向模型的工作负载。在LoCoMo上,Stashbird使用的摄入提示令牌比Graphiti少76.4倍。在同一基准上与复现的Hindsight相比,它使用的检索提示令牌少8.1倍,准确率低1.6个百分点。它在LongMemEval-S和GroupMemBench上取得了比Hindsight更高的准确率,在EverMemBench上准确率相当。
cs.AI / 298 / 2609.34249

Evolving Support Priorities in Empathetic Reinforcement Learning

共情强化学习中不断演进的支持优先级
Huang, Pengyu, Han, Zhiyuan, Tong, Wenwen, Guo, Hewei, Chen, Jiangnan, Chen, Sirui, Lu, Lewei, Zhu, Beier, Yang, Xun
Abstract
We identify a fundamental mismatch in empathetic reinforcement learning: support priorities evolve with the dialogue state, yet existing methods typically optimize predefined reward specifications that remain fixed across turns. To model these evolving support priorities, we organize empathetic support along cognitive, affective, and proactive empathy, and propose Context-Adaptive Rubric Evolution (CARE). At each turn, CARE generates a context-adaptive rubric by adjusting both the weights of these three empathy dimensions and their fine-grained evaluation criteria. The rubric generator is trained with turn-level rubric supervision and human preference data through supervised fine-tuning followed by preference-based reinforcement learning, and then serves as an adaptive reward interface for online empathetic RL. Integrated with both RLVER and MICA, CARE achieves state-of-the-art performance across SentientBench, EQBench3, and EMPA under three independent LLM judges. Notably, on EMPA, CARE improves EPM-Idx over the strongest baseline by at least 13 points under all three judges, including an increase from 28.11 to 83.54 under Gemini-2.5-Pro. Further analyses show that learned rubric priorities systematically vary across dialogue stages and user emotions, demonstrating that CARE adapts what is rewarded as support needs evolve.
Chinese Translation
我们发现共情强化学习中存在一个根本性错配:支持优先级会随对话状态演变,但现有方法通常优化预定义的奖励规范,而这些规范在各轮次中保持固定。为对这类不断演变的支持优先级建模,我们将共情支持按认知共情、情感共情和主动共情进行组织,并提出上下文自适应评分标准演化(Context-Adaptive Rubric Evolution, CARE)。在每一轮,CARE通过调整这三个共情维度的权重及其细粒度评估标准,生成上下文自适应评分标准。该评分标准生成器利用轮级评分标准监督和人类偏好数据,通过监督微调及随后的基于偏好的强化学习进行训练,然后作为在线共情强化学习的自适应奖励接口。CARE与RLVER和MICA集成后,在三个独立LLM评判器下于SentientBench、EQBench3和EMPA上达到最先进性能。值得注意的是,在EMPA上,CARE在三个评判器下均将EPM-Idx相比最强基线至少提升13分,其中在Gemini-2.5-Pro下从28.11提升至83.54。进一步分析表明,学习到的评分标准优先级会随对话阶段和用户情绪系统性变化,这说明CARE会随着支持需求的演变调整被奖励的内容。
cs.AI / 299 / 2609.34259

QuantaSpike: Short-Window Spike-Driven Quantization for Large Language Models

QuantaSpike:面向大语言模型的短窗口脉冲驱动量化
Hu, Bang, Zhu, Guowei, Lv, Changze, Zheng, Xiaoqing, Zhang, Fengzhe, Zhang, Fan, Cao, Wei
Abstract
Large language models (LLMs) achieve strong performance across many tasks but rely on dense multiply-accumulate (MAC) operations during inference, resulting in high energy cost. Spiking neural networks (SNNs) offer an event-driven alternative in which synaptic integration uses lightweight accumulation. However, spike-driven LLM inference remains difficult because outlier-heavy activations typically require long firing windows or auxiliary non-spiking paths. We propose QuantaSpike, a short-window spike-driven quantization framework for LLMs built around Logarithmic Ternary Integrate-and-Fire (LTIF) neurons. LTIF uses ternary events with power-of-two membrane-response quanta, improving the information represented by each firing step while retaining shift-ACC-compatible computation. QuantaSpike combines this neuron with group-adaptive gain and selective outlier admission: normal values use residual LTIF steps, whereas admitted outliers receive one additional onset spike before entering the same residual dynamics. Across OPT and Llama-2, QuantaSpike achieves state-of-the-art or competitive perplexity and zero-shot accuracy among spike-driven LLM quantization methods. It also transfers to newer dense LLMs, remaining close to the FP16 reference on Llama-3-8B and Qwen3-8B under the same four-step firing window. Analytical linear-energy projections show that QuantaSpike reduces the energy of one linear transformation by about $80.0\%$ on OPT models and $67.1\%$ on Llama-2 models relative to SpikeQuant, providing an accurate and energy-efficient spike-driven path for LLM inference.
Chinese Translation
大语言模型(LLMs)在许多任务上取得了强大性能,但在推理过程中依赖密集的乘累加(MAC)操作,导致高能耗。脉冲神经网络(SNNs)提供了一种事件驱动的替代方案,其中突触整合使用轻量级累加。然而,脉冲驱动的 LLM 推理仍然困难,因为含有大量离群值的激活通常需要长发放窗口或辅助的非脉冲路径。我们提出了 QuantaSpike,一个面向 LLMs 的短窗口脉冲驱动量化框架,其核心是 Logarithmic Ternary Integrate-and-Fire (LTIF) 神经元。LTIF 使用具有 2 的幂膜响应量子的三值事件,提高了每个发放步骤所表示的信息,同时保持与移位累加(shift-ACC)兼容的计算。QuantaSpike 将该神经元与组自适应增益和选择性离群值准入相结合:正常值使用残差 LTIF 步骤,而被准入的离群值在进入相同的残差动态之前会额外接收一个起始脉冲。在 OPT 和 Llama-2 上,QuantaSpike 在脉冲驱动的 LLM 量化方法中实现了最先进或有竞争力的困惑度和零样本准确率。它还能迁移到更新的密集 LLMs,在相同的四步发放窗口下,在 Llama-3-8B 和 Qwen3-8B 上保持接近 FP16 参考的性能。分析性线性能耗投影表明,相对于 SpikeQuant,QuantaSpike 在 OPT 模型上将单次线性变换的能耗降低了约 80.0%,在 Llama-2 模型上降低了 67.1%,为 LLM 推理提供了一条准确且节能的脉冲驱动路径。
cs.AI / 300 / 2609.34262

Maintaining Benchmarks Against Increasingly Capable Agents: Detection and Remediation of Unearned Passes

针对能力日益增强的智能体维护基准:非正当通过的检测与修复
Luo, Weijun, Luu, Kelvin, Liu, Xinyi, Luo, Guangze, Calvo, Miguel Romero, Dan, Soham, Zhang, Daniel Yue, Liu, Ying, Elfeki, Mohamed
Abstract
Agentic benchmarks guide model selection and training. Yet an agent can pass a task without demonstrating the intended capability. Such outcomes constitute unearned passes; their proportion among all passes defines the integrity gap. As agents improve, benchmark surfaces that once seemed harmless can become exploitable, making benchmark validity an ongoing maintenance problem. We introduce a process-verification framework that audits passing trajectories, distinguishes evidenced reward hacking from verifier weakness, and localizes exploitable surfaces for repair. Across 3,810 passing trajectories from 29 model-benchmark cohorts, confirmed violations often increase with model generation but not monotonically. On SWEBench Pro V1.0, confirmed violation rates rise from 24% to 73% between Opus 4.7 and Fable 5 on matched tasks; later cohorts fall to 11% for Fable 5.1 and 0% for GPT-6 Astra. These comparisons are descriptive: configurations were not normalized, and the latest models also pass fewer exploitable tasks. Violations concentrate around a small set of recurring surfaces, especially unintended access to reference solutions through git history. Three repair case studies across two benchmarks show why blocking a recorded exploit is insufficient: the same protected information can remain accessible through another route. Therefore, we combine minimal patches with exploit replay and fresh agent evaluation, auditing new passes under the original standard. No evaluated attempt against the final patches reached the protected channel, and every post-patch pass was judged legitimate. Benchmark integrity requires ongoing maintenance: audit passing behavior, repair the enabling surface, and re-evaluate both exploit access and legitimate solvability.
Chinese Translation
智能体基准指导模型选择与训练。然而,智能体可能在没有展示预期能力的情况下通过任务。此类结果构成非正当通过;其在所有通过中所占比例定义了完整性差距。随着智能体能力的提升,曾经看似无害的基准表面可能变得可利用,使得基准有效性成为一个持续的维护问题。我们引入了一个过程验证框架,用于审计通过轨迹,区分有证据的奖励黑客行为与验证器弱点,并定位可利用表面以进行修复。在来自29个模型-基准队列的3,810条通过轨迹中,已确认的违规行为通常随着模型代际的增加而增加,但并非单调递增。在SWEBench Pro V1.0上,在匹配任务上,已确认的违规率从Opus 4.7的24%上升到Fable 5的73%;后续队列中,Fable 5.1降至11%,GPT-6 Astra降至0%。这些比较是描述性的:配置未进行归一化,且最新模型通过的可利用任务也更少。违规行为集中在少数反复出现的表面上,尤其是通过git历史意外访问参考解决方案。跨两个基准的三个修复案例研究表明,仅阻止已记录的漏洞是不够的:相同的受保护信息可能仍可通过另一条途径访问。因此,我们将最小补丁与漏洞重放和新智能体评估相结合,在原始标准下审计新的通过。针对最终补丁的评估尝试均未触及受保护通道,且每个补丁后的通过都被判定为合法。基准完整性需要持续维护:审计通过行为,修复使能表面,并重新评估漏洞访问和合法可解性。
计算机视觉 (Computer Vision)
300
cs.CV / 1 / 2609.31651

PalmLeaf-VQA: A Multi-Script Visual Question Answering Benchmark for Historical Palm-Leaf Manuscript Understanding Across Diverse Regions

PalmLeaf-VQA:一个面向跨区域历史贝叶手稿理解的多文种视觉问答基准
Thuon, Nimol, Du, Jun, Theang, Panhapin
Abstract
Historical manuscripts remain largely absent from modern vision-language benchmarks, leaving open how well multimodal large language models (MLLMs) handle culturally diverse, degraded, and non-Latin document images. We introduce \textbf{PalmLeaf-VQA}, a multi-script visual question answering benchmark for historical palm-leaf manuscript understanding across South and Southeast Asian traditions. PalmLeaf-VQA contains \textbf{923 curated manuscript images} and \textbf{7,384 question--answer pairs} from eight collection groups: Balinese, Grantha, Jathakam, Kambaramayanam, Kannada, Khmer, Sundanese, and Tamil. Unlike recognition-oriented resources, the benchmark targets manuscript-aware visual reasoning over preservation-relevant cues, including physical condition, line structure, material and coating, binding holes, margins, symbols, drawings, and localized visual artifacts. We evaluate recent proprietary and open-weight MLLMs under open-answer and constrained-answer prompting and provide fine-grained analysis across collections, question categories, and task types. The strongest evaluated model reaches only \textbf{58.00\% exact-match accuracy} on the held-out test split, revealing substantial limitations in current MLLMs for rare-script, degraded-layout, and preservation-oriented document understanding. PalmLeaf-VQA provides a standardized benchmark for advancing culturally grounded and layout-aware multimodal document analysis.
Chinese Translation
历史手稿在现代视觉-语言基准中基本缺席,这使得多模态大语言模型(MLLMs)如何处理文化多样、退化且非拉丁文字的手稿图像这一问题仍悬而未决。我们提出了PalmLeaf-VQA,一个面向南亚和东南亚传统历史贝叶手稿理解的多文种视觉问答基准。PalmLeaf-VQA包含来自八个收藏组的923幅精选手稿图像和7,384个问答对:巴厘、Grantha、Jathakam、Kambaramayanam、卡纳达、高棉、巽他、泰米尔。与以识别为导向的资源不同,该基准针对的是对手稿保护相关线索的手稿感知视觉推理,包括物理状况、行结构、材料和涂层、装订孔、页边距、符号、绘图以及局部视觉伪影。我们在开放答案和受限答案提示下评估了最新的专有和开放权重MLLMs,并提供了跨收藏、问题类别和任务类型的细粒度分析。在留出测试集上,评估的最强模型仅达到58.00%的精确匹配准确率,揭示了当前MLLMs在稀有文字、退化版面和面向保护的文件理解方面的重大局限性。PalmLeaf-VQA为推进基于文化的、布局感知的多模态文档分析提供了一个标准化基准。
cs.CV / 2 / 2609.31654

Temporal-Attention Head Specialization During Video Diffusion Training

视频扩散训练中的时间注意力头特化
Ha, Taewoo, Anik, Shafayat Mowla, Lee, Dae Yeol, Lee, Byeong Kil, Ryoo, Jeeho
Abstract
Video diffusion transformers depend on temporal attention to coordinate information across frames, yet nearly everything known about this mechanism comes from analyzing trained models, so when and where temporal-attention structure forms during training remains poorly characterized. Population averages can also hide it, since a few specializing heads and a diffusing majority cancel in the mean. We therefore conduct a checkpoint-resolved census of every temporal-attention head across nine Open-Sora STDiT training runs spanning three model scales (306M to 1.03B parameters), scoring each head with an entropy-normalized measure of cross-frame attention concentration (CFAC) under a preregistered change-point and effect-size selection rule. The census reveals the sparse picture that averages obscure. Aggregate CFAC is flat or decreasing in every run, while a small minority of heads, roughly 4--13% in full-grid runs, develops pronounced concentration. Across seeds, the reproducible signal is positional but block-level. Selected heads repeatedly arise in the first temporal block, whereas individual head coordinates do not reproduce once block membership is accounted for. Among the analyzed 760M selected heads, attention maps converge to a small repertoire of local frame-routing motifs, self-frame diagonals and adjacent-frame bands, even when the responsible coordinates differ across runs. Correlation and ablation analyses do not establish a causal link to generated video quality, and we bound our claims accordingly. Beyond this STDiT family, the study contributes a transferable methodology. Checkpoint-resolved, per-head analysis under fixed selection rules can expose sparse temporal organization in other factorized video diffusion transformers and, with adapted routing metrics, in joint spatio-temporal architectures.
Chinese Translation
视频扩散Transformer依赖时间注意力来协调跨帧信息,然而关于这一机制几乎所有的已知都来自于对训练后模型的分析,因此时间注意力结构在训练期间何时、何处形成仍缺乏刻画。群体平均也可能掩盖这一点,因为少数特化头与占多数的弥散头在均值中相互抵消。因此,我们对九个Open-Sora STDiT训练运行中的每个时间注意力头进行了基于检查点的普查,这些运行涵盖三种模型规模(306M到1.03B参数),并在预注册的变化点和效应量选择规则下,用熵归一化的跨帧注意力集中度(CFAC)指标对每个头进行评分。这项普查揭示了被均值掩盖的稀疏图景。在所有运行中,聚合CFAC持平或下降,而一小部分头(在全网格运行中约4-13%)发展出显著的集中。跨随机种子,可复现的信号是位置性的,但处于块级别。被选中的头反复出现在第一个时间块中,而一旦考虑了块归属,单个头的坐标就不能复现。在分析的760M参数规模选中的头中,注意力图收敛到少量局部帧路由模式,即自帧对角线和相邻帧带,即使不同运行中对应的坐标不同。相关性和消融分析未能建立与生成视频质量的因果联系,因此我们相应地限定了我们的主张。除了这个STDiT家族之外,该研究贡献了一种可迁移的方法论。在固定选择规则下基于检查点的逐头分析,可以揭示其他因子化视频扩散Transformer中稀疏的时间组织,并且,通过调整路由指标,也可用于联合时空架构。
cs.CV / 3 / 2609.31655

One Evaluation, Any Operating Point: Hypernetwork-Amortized MeanFlow for 3D MRI Reconstruction

一次评估,任意操作点:用于3D MRI重建的超网络摊销MeanFlow
Wang, Ruibo
Abstract
Generative priors reconstruct accelerated 3D MRI well but pay heavily at deployment: tens of network evaluations per volume, and protocol-specific hyperparameter tuning. A third hidden cost is the scanner's fixed sampling pattern. We treat the whole operating point as an input. A 3D MeanFlow patch network (a one-step flow model) is fine-tuned end-to-end through a warm-started, five-iteration differentiable conjugate-gradient projection. A small hypernetwork maps the operating point (data-consistency weight, acceleration, and the Cartesian sampling pattern itself) to the network's per-channel modulation. Three findings follow. (i) Learning the acquisition is worth more than any other operating point: on clinical knee data, the learned mask gains up to +2.34 dB over the protocol's variable-density mask. This gain requires the solver: with a feed-forward reconstructor the same learned mask hurts at 4x (-1.7 dB), but with the data-consistency projection it adds +4.6 dB. (ii) One evaluation is highly effective: it beats a 20-step patch-diffusion prior by up to +3.1 dB on brain and +2.9 dB on knee. Three to five evaluations extend the front to +6 dB while using a quarter of the prior's network calls. (iii) Fully sampled targets are optional: trained self-supervised on a split of acquired samples, the reconstructor matches its supervised twin at 4x on real data. Finally, we report what failed and why: subject-adaptive acquisition from measured energy, combining self-supervision with learned acquisition, and amortising the data-consistency weight.
Chinese Translation
生成先验能很好地重建加速的3D MRI,但在部署时代价高昂:每个体积需要数十次网络评估,并且需要针对特定协议进行超参数调优。第三个隐藏成本是扫描仪的固定采样模式。我们将整个操作点视为输入。一个3D MeanFlow补丁网络(一种单步流模型)通过热启动的五次迭代可微共轭梯度投影进行端到端微调。一个小型超网络将操作点(数据一致性权重、加速倍数以及笛卡尔采样模式本身)映射到网络的逐通道调制。由此得出三个发现。(i) 学习采集比其他任何操作点都更有价值:在临床膝关节数据上,学习到的掩膜比协议的可变密度掩膜最多提升+2.34 dB。这一增益需要求解器:使用前馈重建器时,相同的学习掩膜在4倍加速下反而有害(-1.7 dB),但使用数据一致性投影时,它增加了+4.6 dB。(ii) 一次评估非常有效:在脑部数据上比20步补丁扩散先验最多提升+3.1 dB,在膝关节上提升+2.9 dB。三到五次评估将前沿扩展到+6 dB,同时仅使用先验网络调用次数的四分之一。(iii) 全采样目标可选:在获取样本的一个划分上进行自监督训练,重建器在真实数据上4倍加速时与其监督版本相匹配。最后,我们报告了哪些方法失败了以及原因:基于测量能量的受试者自适应采集、将自监督与学习采集相结合,以及摊销数据一致性权重。
cs.CV / 4 / 2609.31657

Enhancing Foundation Models for Imbalanced SAR Ship Classification via Targeted Oversampling

通过针对性过采样增强不平衡SAR船舶分类的基础模型
Awais, Ch Muhammad, Reggiannini, Marco, Moroni, Davide
Abstract
Remote-sensing foundation models offer strong representations for SAR imagery, but their behavior under severe long-tail class imbalance is still not well characterized. We benchmark DOFA and SAR-JEPA on the imbalanced OpenSARShip dataset and compare them with ImageNet-pretrained baselines under a fixed, training-efficient protocol that keeps the backbone frozen. To mitigate imbalance without fine-tuning, we apply four oversampling methods in embedding space exclusively to minority classes and train a lightweight classifier head on the augmented embeddings. Across both foundation models, oversampling improves Macro-F1 and test accuracy relative to their respective baselines, with the largest Macro-F1 gains observed for DOFA using ADASYN (34.39 to 38.56) and for SAR-JEPA using SVM-SMOTE (25.89 to 32.30). We also report class-wise behavior, showing that aggregate improvements can coexist with persistent failures on specific rare classes. Code for embedding extraction and reproducible multi-seed evaluation is provided to support rapid experimentation on free-tier hardware.
Chinese Translation
遥感基础模型为SAR图像提供了强大的表征,但它们在严重长尾类别不平衡下的行为仍未得到很好的刻画。我们在不平衡的OpenSARShip数据集上对DOFA和SAR-JEPA进行基准测试,并在保持骨干网络冻结的固定、训练高效的协议下,将它们与ImageNet预训练的基线进行比较。为了在不微调的情况下缓解不平衡,我们在嵌入空间中仅对少数类应用四种过采样方法,并在增强后的嵌入上训练一个轻量级分类器头。在两个基础模型中,过采样相对于各自的基线提高了Macro-F1和测试准确率,其中DOFA使用ADASYN时观察到最大的Macro-F1提升(34.39到38.56),SAR-JEPA使用SVM-SMOTE时(25.89到32.30)。我们还报告了类别层面的行为,表明整体改进可能与特定稀有类别的持续失败并存。提供了用于嵌入提取和可复现多种子评估的代码,以支持在免费层硬件上的快速实验。
cs.CV / 5 / 2609.31658

Cross-Dataset Transfer and Unknown-Class Detection in Imbalanced SAR Ship Classification

不平衡SAR船舶分类中的跨数据集迁移与未知类检测
Awais, Ch Muhammad, Reggiannini, Marco, Moroni, Davide, Del Corso, Giulio
Abstract
Ship classification from Synthetic Aperture Radar (SAR) imagery is a critical computer vision task, yet the robustness of models under deployment shifts remains unclear. While models are often trained on one dataset and deployed on another, we lack a comprehensive understanding of their cross-dataset generalization. To address this, we evaluate six pretrained models on two SAR ship datasets in three settings: in-domain classification, cross-dataset transfer, and unknown-class detection. For unknown detection, we hold out all classes one at a time. In-domain, SARDet100K gives the best balanced accuracy on both datasets (73.4\% on FUSARShip and 53.2\% on OpenSARShip). In cross-dataset transfer, we observe strong failures: some models show moderate accuracy but near-chance balanced accuracy (for example, 64.5\% accuracy but 33.3\% balanced accuracy for OpenSARShip to FUSARShip). In unknown detection, performance depends on the held-out class and dataset, while MC-dropout variance is often close to random. These findings show that cross-dataset generalization in SAR remains limited and that task-specific uncertainty scores are often more informative than MC-dropout variance for held-out-class detection, although their relative ranking depends on the dataset and held-out class.
Chinese Translation
从合成孔径雷达(SAR)图像中进行船舶分类是一项关键的计算机视觉任务,然而模型在部署偏移下的鲁棒性仍不清楚。尽管模型通常在一个数据集上训练并部署到另一个数据集上,但我们对其跨数据集泛化能力缺乏全面理解。为了解决这个问题,我们在两个SAR船舶数据集上评估了六个预训练模型,在三种设置下:域内分类、跨数据集迁移和未知类检测。对于未知检测,我们每次留出一个类别。在域内,SARDet100K在两个数据集上均给出了最佳平衡准确率(FUSARShip上73.4%,OpenSARShip上53.2%)。在跨数据集迁移中,我们观察到严重的失败:一些模型显示出中等准确率但接近随机的平衡准确率(例如,从OpenSARShip到FUSARShip,准确率为64.5%,但平衡准确率仅为33.3%)。在未知检测中,性能取决于留出的类别和数据集,而MC-dropout方差通常接近随机。这些发现表明,SAR中的跨数据集泛化仍然有限,并且对于留出类检测,任务特定的不确定性分数通常比MC-dropout方差更具信息量,尽管它们的相对排名取决于数据集和留出的类别。
cs.CV / 6 / 2609.31661

ForensicZoom: Adaptive Visual Inspection with Multimodal LLMs for Industrial-Grade Face Forgery Detection

ForensicZoom:面向工业级人脸伪造检测的多模态LLM自适应视觉检查
Zhou, Hang, Tang, Yiming, Yu, Kun, Zhu, Qian, Li, Minghao, Wen, Weigao
Abstract
Reliable face forgery detection is critical to the security of online identity verification systems, where missed attacks compromise security and excessive false positives disrupt legitimate users. Specialized forensic detectors achieve strong detection performance but provide limited interpretability, while multimodal large language models (MLLMs) offer strong semantic understanding and interpretable reasoning yet remain substantially weaker for face forgery detection. We argue that a key limitation lies in how visual evidence is acquired: subtle forensic artifacts may be poorly represented at standard resolution, while uniformly processing all cases at higher resolution is computationally inefficient. We therefore introduce ForensicZoom, an industrial-grade MLLM framework for adaptive visual inspection. ForensicZoom first equips a general-purpose MLLM with forensic-aware visual representations and aligns the language model with these features. Its central mechanism, NEED_ZOOM, enables the model to autonomously request magnified views of suspicious regions when the initial evidence is insufficient, turning fixed-pass classification into adaptive multi-round forensic reasoning. The zoom behavior is learned through reward shaping that balances detection accuracy with unnecessary visual inspection, concentrating additional computation on difficult cases. A final attribution optimization stage improves natural-language forensic reports while preserving detection performance. On large-scale industrial identity verification data, ForensicZoom achieves over 97% TPR at 0.1% FPR, substantially outperforming both specialized detectors and existing MLLM-based methods while producing actionable forensic attributions. These results demonstrate that ForensicZoom can provide an effective path toward accurate, interpretable, and scalable MLLM-based face forgery detection.
Chinese Translation
可靠的人脸伪造检测对于在线身份验证系统的安全至关重要,其中漏检攻击会危害安全,而过多误报会干扰合法用户。专用取证检测器实现了强大的检测性能,但可解释性有限,而多模态大语言模型(MLLMs)提供了强大的语义理解和可解释推理,但在人脸伪造检测方面仍然明显较弱。我们认为,一个关键限制在于视觉证据的获取方式:细微的取证伪影在标准分辨率下可能表征不佳,而统一以更高分辨率处理所有案例在计算上效率低下。因此,我们引入了ForensicZoom,一个用于自适应视觉检查的工业级MLLM框架。ForensicZoom首先为通用MLLM配备取证感知的视觉表示,并将语言模型与这些特征对齐。其核心机制NEED_ZOOM使模型能够在初始证据不足时自主请求可疑区域的放大视图,将固定单次分类转变为自适应多轮取证推理。缩放行为通过奖励塑造学习,该奖励塑造平衡检测准确性与不必要的视觉检查,将额外计算集中在困难案例上。最后的归因优化阶段改进了自然语言取证报告,同时保持检测性能。在大规模工业身份验证数据上,ForensicZoom在0.1% FPR下实现了超过97%的TPR,显著优于专用检测器和现有的基于MLLM的方法,同时产生可操作的取证归因。这些结果表明,ForensicZoom可以为准确、可解释和可扩展的基于MLLM的人脸伪造检测提供有效途径。
cs.CV / 7 / 2609.31662

Learning Steadily: Accumulating Relative Point Margin Scores for Face Image Quality Assessment

稳定学习:累积相对点间隔分数用于人脸图像质量评估
Ozgur, Guray, Chettaoui, Tahar, Caldeira, Eduarda, Huber, Marco, Kolf, Jan Niklas, Damer, Naser, Boutros, Fadi
Abstract
Face Image Quality Assessment determines the suitability of captured face images for automated face recognition (FR), a critical capability for reliable biometric systems. Existing state-of-the-art FR-integrated FIQA methods suffer from temporal instability: as the feature space evolves during training, single-epoch quality estimates fluctuate, creating a moving target that undermines reliable quality prediction. We introduce CARPM-FIQA, a stabilization strategy for FR-integrated FIQA that accumulates relative point margin measurements, the ratio between intra-class compactness and inter-class separation, across the entire training trajectory rather than relying on single-epoch estimates. This cumulative averaging approach provides theoretically grounded advantages: reduced variance in quality estimates, improved mean squared error, and enhanced ranking stability with convergence guarantees as training progresses. Through controlled experiments on the SynFIQA dataset with labeled quality groups, we demonstrate that cumulative averaging achieves superior discriminative ability, and ablation studies across different training configurations confirm consistent improvements. Evaluated against twelve FIQA methods on eight challenging benchmarks with four FR models at two FMR thresholds, CARPM-FIQA places 4th (CARPM-FIQA(L)) and 6th (CARPM-FIQA(S)) of 17 compared methods by pAUC-EDC and AUC-EDC averaged across FR models and, after per-benchmark normalization, across benchmarks, staying within a few percent of the best method's normalized average for every FR model, providing a principled solution to training instability while maintaining the performance benefits of FR integration. More broadly, our work demonstrates that temporal aggregation strategies can stabilize training objectives in deep learning systems where target values inherently fluctuate due to evolving feature representations.
Chinese Translation
人脸图像质量评估(Face Image Quality Assessment,FIQA)确定捕获的人脸图像对于自动人脸识别(FR)的适用性,这是可靠生物识别系统的关键能力。现有的最先进的FR集成FIQA方法存在时间不稳定性:随着特征空间在训练过程中演变,单轮次质量估计波动,产生一个移动目标,从而削弱了可靠的质量预测。我们提出CARPM-FIQA,一种用于FR集成FIQA的稳定策略,它累积整个训练轨迹上的相对点间隔测量值(即类内紧凑度与类间分离度的比率),而不是依赖于单轮次估计。这种累积平均方法提供了有理论依据的优势:减少质量估计的方差,改善均方误差,并随着训练的进行增强排序稳定性,且具有收敛保证。通过在带有标记质量组的SynFIQA数据集上进行受控实验,我们证明累积平均实现了更优的判别能力,并且不同训练配置的消融研究证实了一致的改进。在八个具有挑战性的基准上,针对十二种FIQA方法,使用四个FR模型在两个FMR阈值下进行评估,CARPM-FIQA在17种比较方法中按pAUC-EDC和AUC-EDC(跨FR模型平均,并在每个基准归一化后跨基准平均)分别排名第4(CARPM-FIQA(L))和第6(CARPM-FIQA(S)),对于每个FR模型,均保持在最佳方法归一化平均值的百分之几以内,为训练不稳定性提供了原则性解决方案,同时保持了FR集成的性能优势。更广泛地说,我们的工作表明,在目标值由于不断演变的特征表示而固有波动的深度学习系统中,时间聚合策略可以稳定训练目标。
cs.CV / 8 / 2609.31665

Language-Augmented Video Action Anticipation: Design Fundamentals, Benchmarks, and Open Challenges

语言增强的视频动作预判:设计基础、基准与开放挑战
Mohammadi, Mahsa, Fu, Zeyu, Rowlands, Sareh
Abstract
Action anticipation predicts future human actions from partial video under incomplete context and temporal uncertainty. Recent systems introduce large language models (LLMs), vision-language models (VLMs), or language-derived semantics at different stages, but reported gains are difficult to interpret when task formulation, visual pretraining, supervision, decoder design, and evaluation code change simultaneously. The central contribution of this review is an evidence-aware design map that crosses task regime with the point at which language-derived information intervenes. We characterise task regimes along six axes. These axes organise the literature into five broad task families: single-action, sequence, object-interaction, cross-view, and planning-oriented settings. C1-C3 locate interventions in context construction, goal/intention modelling, and future decoding, while C4 is treated as an adjacent, emerging grounding/executability extension. Unlike a generic processing pipeline, the map links each intervention to an appropriate counterfactual, failure diagnosis, and permissible evidence claim. Supporting contributions include a protocol-level audit of Ego4D-LTA and EPIC-KITCHENS-100, a multidimensional evidence profile, and the Backbone-Aware Comparison and Ablation Protocol (BCAP). The unresolved EK-100 record is treated as a reporting-comparability case study and is not used as a leaderboard. Evidence for LLM benefits, goal ambiguity, and horizon effects is therefore formulated as testable hypotheses requiring matched validation, not as causal conclusions. The accompanying package contains the coded evidence, source locators, protocol metadata, and versioned catalogue used in the review.
Chinese Translation
动作预判旨在根据部分视频,在不完整上下文和时间不确定性下预测未来的人类动作。近期系统在不同阶段引入大语言模型(LLMs)、视觉语言模型(VLMs)或语言衍生的语义,但当任务设定、视觉预训练、监督、解码器设计和评估代码同时发生变化时,所报告的增益难以解释。本综述的核心贡献是一个证据感知的设计图谱,它将任务范式与语言衍生信息介入的节点进行交叉组织。我们沿六个维度刻画任务范式。这些维度将文献组织为五个宽泛的任务族:单动作、序列、物体交互、跨视角和规划导向场景。C1-C3 将干预定位在上下文构建、目标/意图建模和未来解码中,而 C4 被视为相邻的新兴 grounding/可执行性扩展。与通用处理流水线不同,该图谱将每种干预与适当的反事实、失败诊断和可允许的证据主张关联起来。支持性贡献包括对 Ego4D-LTA 和 EPIC-KITCHENS-100 的协议级审计、多维证据画像,以及骨干感知比较与消融协议(BCAP)。未解决的 EK-100 记录被作为报告可比性案例研究处理,而非用作排行榜。因此,关于 LLM 收益、目标歧义和预测时域效应的证据被表述为需要匹配验证的可检验假设,而非因果结论。随附资源包包含本综述使用的编码证据、来源定位符、协议元数据和版本化目录。
cs.CV / 9 / 2609.31668

Query-aligned video frame selection for long video understanding

用于长视频理解的查询对齐视频帧选择
Islam, Md. Safayet, Sarkar, Dilip, Liang, Liang
Abstract
Multimodal large language models (MLLMs) process multimodal inputs by converting text, images, and videos into token sequences that are subsequently processed by a backbone language model. While MLLMs have achieved excellent performance in understanding the content of individual images, video understanding remains significantly more difficult because videos contain large number of video frames. MLLMs typically process only a subset of these frames, usually ranging from 8 to 64. MLLMs usually sample frames uniformly, regardless of their relevance to the question being answered. To address this limitation, several training-free, model-agnostic methods for selecting question-relevant frames have recently been proposed. In this work, we introduce a frame-selection method designed specifically for multiple-choice questions. We extend the query text by appending semantic cues derived from the answer choices and employ a direct query-frame alignment scoring mechanism. To the best of our knowledge, our method is the first to directly utilize answer choices as inference-time cues for selecting frames relevant to answering a question. The method first constructs a compact candidate pool by subsampling video frames at a fixed rate. The frames are then scored according to their maximum cosine similarity across all question-answer pairs to identify the most relevant frames for a given query. This approach preserves a fixed token budget while improving the relevance of the visual evidence provided to the downstream MLLM. We evaluate the effectiveness of our frame-selection method on the MLVU, Video-MME, and LongVideoBench benchmarks using three MLLMs: LLaVA-Mini, Qwen2-VL, and LLaVA-Video. Experimental results demonstrate that answer-aware frame selection generally outperforms uniform sampling and existing training-free frame-selection methods under the same frame budget.
Chinese Translation
多模态大语言模型(MLLMs)通过将文本、图像和视频转换为标记序列来处理多模态输入,这些序列随后由骨干语言模型处理。尽管MLLMs在理解单个图像内容方面取得了优异性能,但视频理解仍然困难得多,因为视频包含大量视频帧。MLLMs通常仅处理这些帧的一个子集,数量通常在8到64之间。MLLMs通常均匀采样帧,而不考虑它们与所回答问题的相关性。为了解决这一限制,最近提出了几种无需训练、与模型无关的方法来选择与问题相关的帧。在这项工作中,我们介绍了一种专门为多项选择题设计的帧选择方法。我们通过附加从答案选项中导出的语义线索来扩展查询文本,并采用直接的查询-帧对齐评分机制。据我们所知,我们的方法是第一个直接利用答案选项作为推理时线索来选择与回答问题相关的帧的方法。该方法首先通过以固定速率对视频帧进行子采样来构建一个紧凑的候选池。然后,根据帧在所有问答对上的最大余弦相似度对帧进行评分,以识别给定查询最相关的帧。这种方法在保持固定标记预算的同时,提高了提供给下游MLLM的视觉证据的相关性。我们使用三个MLLM:LLaVA-Mini、Qwen2-VL和LLaVA-Video,在MLVU、Video-MME和LongVideoBench基准上评估了我们帧选择方法的有效性。实验结果表明,在相同的帧预算下,答案感知的帧选择通常优于均匀采样和现有的无需训练的帧选择方法。
cs.CV / 10 / 2609.31670

One-Step Is Optimal: Unconditional Rectified Flows are Noise2Noise Denoisers, and Multi-Step Integration Provably Hurts---A Benchmark and Task-Based Detectability Study on Low-Dose CT

一步即最优:Unconditional Rectified Flows是Noise2Noise去噪器,且多步积分可证明有害——一项关于低剂量CT的基准与基于任务的可检测性研究
Sereda, Timothy, Jha, Debesh
Abstract
Iterative and generative denoisers are increasingly used under the assumption that multi-step refinement outperforms a single regression pass. We show the opposite for \emph{label-free} denoising. An \emph{unconditional} rectified flow trained on two noisy observations of the same signal, as in Noise2Noise, has a minimiser whose one-step readout is exactly the MMSE denoiser without requiring clean targets. In contrast, multi-step integration provably departs from the MMSE solution because the flow terminates at the noisy data distribution rather than the clean-signal distribution. This departure is exact in a tractable Gaussian model and is confirmed experimentally: one-step flow matches a direct regressor, whereas multi-step Euler integration progressively reduces fidelity. Counterintuitively, the degradation increases with training quality, as a better velocity field more faithfully transports samples toward the noisy terminal law. The key ingredient is therefore the decorrelated \emph{pairing}, not the flow machinery: a one-step regressor trained on matched noisy pairs gives our best label-free result ($+1.99$,dB). We evaluate these findings on \textbf{CTDenoiser}, a controlled low-dose CT benchmark spanning five architectures and supervised, similarity-based, blind-spot, and per-image methods. Among label-free approaches, only correlated-noise-aware Noise2Sim improves over the noisy baseline, while Noise2Void is flat-to-negative because CT noise violates its pixel-independence assumption. Finally, although supervised denoisers gain approximately $4$,dB PSNR, a channelized Hotelling observer shows reduced low-contrast lesion detectability, revealing clinically relevant degradation missed by PSNR and SSIM.
Chinese Translation
迭代式和生成式去噪器越来越多地被使用,其假设是多步细化优于单次回归。我们针对无标签去噪展示了相反的结果。一个在相同信号的两个噪声观测上训练的无条件矫正流(如Noise2Noise中),其最小化器的一步读出恰好是MMSE去噪器,且不需要干净目标。相反,多步积分可证明偏离MMSE解,因为流终止于噪声数据分布而非干净信号分布。这种偏离在可处理的高斯模型中是精确的,并通过实验证实:一步流与直接回归器匹配,而多步欧拉积分逐渐降低保真度。反直觉地,退化随训练质量提高而增加,因为更好的速度场更忠实地将样本输运到噪声终端分布。因此,关键要素是去相关配对,而非流机制:在匹配的噪声对上训练的一步回归器给出了我们最佳的无标签结果(+1.99 dB)。我们在CTDenoiser上评估这些发现,这是一个受控的低剂量CT基准,涵盖五种架构以及监督、基于相似性、盲点和逐图像方法。在无标签方法中,只有相关噪声感知的Noise2Sim优于噪声基线,而Noise2Void持平到负面,因为CT噪声违反了其像素独立性假设。最后,尽管监督去噪器获得约4 dB的PSNR提升,但通道化霍特林观察者显示低对比度病变可检测性降低,揭示了PSNR和SSIM所忽略的临床相关退化。
cs.CV / 11 / 2609.31671

Unsupervised spiking feature learning for event-based pedestrian crossing detection: approaching supervised accuracy without labelled training data

用于基于事件的行人过街检测的无监督脉冲特征学习:在无标注训练数据下接近监督精度
Teklu, Henok, Sakhai, Mustafa, Mertik, Matej, Wielgosz, Maciej
Abstract
Event cameras are well suited to pedestrian crossing detection, and spiking neural networks (SNNs) can process their output natively, but current SNN detectors are trained with supervised backpropagation and therefore require costly frame-level crossing labels. We investigate crossing detection with no labels in feature learning and report the first unsupervised results on the recent DVS-PedX pedestrian-crossing benchmark. A single spiking layer trained with winner-take-all spike-timing-dependent plasticity learns a dictionary from unlabelled event-frame patches; frames are encoded by cosine similarity to the learned filters with spatial pooling and read out by a linear classifier, the only supervised component. On the 24,454-frame test split the method attains 90.3% accuracy and an area under the receiver operating characteristic curve (AUROC) of 0.936, compared with 92.0% and 0.943 for a supervised spiking network trained end-to-end on the same frames; under adverse weather the AUROC is 0.913. The result is insensitive to the choice of plasticity rule but depends strongly on the readout protocol: with the classical neuron-assignment readout the same network attains only 0.70 AUROC. On the benchmark's real converted portion, a readout refit lifts performance from chance (0.53) to 0.67 AUROC, within the published supervised range. These findings indicate that the accuracy cost of removing labels from feature learning is small on this benchmark, and that reported weaknesses of unsupervised spiking networks may be attributable to the readout protocol rather than to the learning rule.
Chinese Translation
事件相机非常适合行人过街检测,脉冲神经网络(SNN)能够原生处理其输出,但当前的SNN检测器使用监督反向传播进行训练,因此需要昂贵的帧级过街标签。我们研究了在特征学习中不使用标签的过街检测,并在最近的DVS-PedX行人过街基准上报告了首个无监督结果。一个使用赢者通吃脉冲时序依赖可塑性训练的单一脉冲层从未标注的事件帧块中学习字典;帧通过余弦相似度编码到学习到的滤波器并进行空间池化,然后由线性分类器读出,这是唯一的有监督组件。在24,454帧的测试集上,该方法达到了90.3%的准确率和0.936的接收者操作特征曲线下面积(AUROC),而在相同帧上端到端训练的有监督脉冲网络分别为92.0%和0.943;在恶劣天气下,AUROC为0.913。结果对可塑性规则的选择不敏感,但强烈依赖于读出协议:使用经典的神经元分配读出,同一网络仅达到0.70 AUROC。在基准的真实转换部分,读出重新拟合将性能从随机水平(0.53)提升至0.67 AUROC,处于已发表的监督范围内。这些发现表明,在该基准上,从特征学习中移除标签的准确率代价很小,并且所报道的无监督脉冲网络的弱点可能归因于读出协议而非学习规则。
cs.CV / 12 / 2609.31677

A Comparative Transfer-Learning Study of CNN Backbones for Partial Face Recognition on the SoF Dataset

基于迁移学习的CNN主干网络在SoF数据集上部分人脸识别的比较研究
Kubba, Ahmed, Alsalama, Ali, Abdalla, Abdelrahman, Nasir, Qassim, Talib, Manar Abu
Abstract
Face recognition is widely deployed in surveillance, access control, and forensic workflows, yet accuracy degrades sharply once the face is occluded by accessories, foreground objects, or the frame edge. Because most faces in the wild are partial, robust partial face recognition (PFR) remains open. This paper compares three pretrained convolutional backbones, ResNet-50, VGG-16, and FaceNet, fine-tuned for PFR by transfer learning under identical preprocessing, splitting, and optimization protocols on the Specs-on-Faces (SoF) dataset. All three arms use a common 160x160 input and a frozen backbone with a trainable head under a fixed epoch budget and no per-backbone hyperparameter search. The FaceNet configuration, denoted PFN (Partial FaceNet), substantially outperforms the other two, reaching 97.4% test accuracy with macro-averaged 87.04% precision, 84.61% recall, and 84.17% F1 over the 112 identity classes, the highest accuracy and recall reported on SoF.
Chinese Translation
人脸识别广泛部署于监控、门禁和法医工作流程中,然而一旦人脸被配饰、前景物体或画面边缘遮挡,准确率就会急剧下降。由于自然场景中的大多数人脸都是部分的,稳健的部分人脸识别(PFR)仍然是一个未解决的问题。本文在Specs-on-Faces (SoF)数据集上,在相同的预处理、划分和优化协议下,通过迁移学习对三种预训练卷积主干网络(ResNet-50、VGG-16和FaceNet)进行微调,用于PFR,并进行了比较。所有三个分支均使用统一的160x160输入和冻结的主干网络,并带有可训练的分类头,在固定训练轮数预算下进行,且未针对每个主干网络进行超参数搜索。FaceNet配置(称为PFN,即Partial FaceNet)显著优于其他两种,在112个身份类别上达到了97.4%的测试准确率,宏平均精确率为87.04%,召回率为84.61%,F1为84.17%,这是SoF上报告的最高准确率和召回率。
cs.CV / 13 / 2609.31679

Toward AI-Assisted Poultry Coccidiosis Diagnosis: Evaluating Gemini and BiomedParse on Eimeria Microscopy Images

迈向AI辅助家禽球虫病诊断:评估Gemini与BiomedParse在艾美耳球虫显微图像上的表现
Alsalama, Ali, Kubba, Ahmed, Talib, Manar Abu
Abstract
Coccidiosis caused by Eimeria parasites is a major economic burden in poultry production, and effective control depends on accurate species-level diagnosis. This study evaluates whether a general-purpose multimodal large language model can support such diagnosis. Google Gemini was assessed on 4,225 mi- croscopy images covering the seven fowl-infecting Eimeria species under two prompting conditions, one without candidate labels and one with a predefined class list, and was further tested for pathology-report generation, while BiomedParse was examined for parasite segmentation. Without candidate labels, the model produced broad and taxonomically inconsistent outputs. With candidate labels, overall accuracy reached only 14.9%, with a strong bias toward E. tenella at 74% and no correct classifications for E. acervulina, E. mitis and E. praecox. Generated treatment reports were coherent but unverified, and segmentation was only partial. Current multimodal models are therefore not yet reliable for standalone Eimeria diagnosis without domain-specific fine- tuning and expert validation.
Chinese Translation
由艾美耳球虫寄生虫引起的球虫病是家禽生产中的主要经济负担,有效控制依赖于准确的种水平诊断。本研究评估通用多模态大语言模型能否支持此类诊断。在两种提示条件下(一种无候选标签,一种有预定义类别列表),对Google Gemini在涵盖七种感染家禽的艾美耳球虫物种的4,225张显微图像上进行了评估,并进一步测试了其生成病理报告的能力,同时检验了BiomedParse的寄生虫分割性能。在没有候选标签的情况下,模型产生了宽泛且分类学上不一致的输出。在有候选标签的情况下,总体准确率仅达到14.9%,强烈偏向于E. tenella(74%),而E. acervulina、E. mitis和E. praecox则没有正确分类。生成的治疗报告连贯但未经核实,分割仅为部分完成。因此,当前的多模态模型在没有领域特定微调和专家验证的情况下,尚不能可靠地用于独立的艾美耳球虫诊断。
cs.CV / 14 / 2609.31681

Devanagari Handwritten Character Recognition Using TrOCR: A Transformer-Based Model with Real-Time Web Deployment

使用TrOCR的天城文手写字符识别:一个基于Transformer的模型及实时Web部署
Baskota, Amrit, Budhathoki, Samyam, Ghimire, Shubham, Ghimire, Abiskar, Phuyal, Sarwesh, P, Baskaran
Abstract
Devnagari is a one of the ancient language of the Indian subcontinent consisting of 36 vowels, 14 consonants and 10 numerals. The accurate recognition of handwritten Devnagari characters is challenging due to high complexity of Devnagari scripts. This paper presents a method to fine tune the pre trained TrOCR model to accurately recognize Devanagari handwritten characters. The Methodology consists of a preprocessing mechanism where input images are standardized into RGB format, tokenized in batches and integrated with Hugging Face Dataset. The pre-trained microsoft/trocr-base-handwritten model is fine-tuned over a dataset of nearly 5000 character images that are uniformly partitioned in the ratio 8:1:1 for training, evaluation and testing. Further optimization is done is the training process through mixed precision training, gradient checkpointing, and early stopping mechanism. The model achieves a character error rate (CER) of 3.95% and a character-level accuracy of 96.05%, outperforming the previous CNN based models. A scalable web application is developed using Vite.js, Golang, and FastAPI which practically deploys the OCR model and serves character recognition task with a latency less than 5 seconds per request. This study demonstrates the use of TrOCR model to build a scalable handwritten Devnagari character recognition system and also a foundation to future research on Devnagari Script Recognition using Transformers.
Chinese Translation
天城文是印度次大陆的古老语言之一,由36个元音、14个辅音和10个数字组成。由于天城文脚本的高度复杂性,准确识别手写天城文字符具有挑战性。本文提出了一种微调预训练TrOCR模型以准确识别天城文手写字符的方法。该方法包括一个预处理机制,其中输入图像被标准化为RGB格式,分批进行标记化,并与Hugging Face数据集集成。预训练的microsoft/trocr-base-handwritten模型在近5000张字符图像的数据集上进行微调,该数据集按8:1:1的比例均匀划分用于训练、评估和测试。训练过程中通过混合精度训练、梯度检查点和早停机制进行了进一步优化。该模型实现了3.95%的字符错误率(CER)和96.05%的字符级准确率,优于之前的基于CNN的模型。使用Vite.js、Golang和FastAPI开发了一个可扩展的Web应用程序,该程序实际部署了OCR模型,并以每次请求低于5秒的延迟提供字符识别任务。本研究展示了使用TrOCR模型构建可扩展的手写天城文字符识别系统,也为未来使用Transformer进行天城文脚本识别的研究奠定了基础。
cs.CV / 15 / 2609.31682

Towards Transparent Diagnostics: Investigating Architectural Trade-offs and Explainability in Malaria Detection

迈向透明诊断:探究疟疾检测中的架构权衡与可解释性
Kunwar, Suman, Dangol, Avishek
Abstract
More than 80 countries have reported malaria cases with 610 thousand deaths and are projected to increase. Identifying malaria early and accurately helps save lives. The effective way to diagnose malaria is through microscopic methods that are labor intensive and require experts with special equipment. Deep learning (DL) has shown promising results in medical diagnosis. Here, we explored various DL models: ResNet18, MobileNetV2, EfficientNet-B2, VGG19 and also proposed a model for detecting malaria presence using blood smears taken from the NIH Malaria dataset. Our experiment shows that MobileNetV2 achieved 96.35% accuracy with the smallest model size (8.49 MB) and fastest inference (1.35 ms). The proposed model achieved 97.67% accuracy, 0.9756 AUC with longest inference time (13.17 ms). The larger architecture outputs a larger model size with moderate accuracy. Upon further pruning, the proposed model gained a slight improvement in accuracy and inference time. The GRAD-CAM, SHAP and LIME shade explainable AI (XAI) insights of the model.
Chinese Translation
超过80个国家报告了疟疾病例,有61万人死亡,并且预计这一数字还会增加。早期准确地识别疟疾有助于挽救生命。诊断疟疾的有效方法是通过显微镜方法,但这种方法需要大量人力,并且需要专家和特殊设备。深度学习(DL)在医学诊断中显示出有希望的结果。在此,我们探索了各种DL模型:ResNet18、MobileNetV2、EfficientNet-B2、VGG19,并提出了一种使用取自NIH疟疾数据集的血液涂片来检测疟疾的模型。我们的实验表明,MobileNetV2达到了96.35%的准确率,模型尺寸最小(8.49 MB),推理速度最快(1.35 ms)。所提出的模型达到了97.67%的准确率,0.9756的AUC,但推理时间最长(13.17 ms)。更大的架构输出更大的模型尺寸,但准确率中等。经过进一步剪枝,所提出的模型在准确率和推理时间上略有改善。利用GRAD-CAM、SHAP和LIME等可解释人工智能(XAI)技术对模型进行了解释。
cs.CV / 16 / 2609.31690

Adapting Vision-Language Models for Human-Readable XAI in Industrial Object Detection

面向工业目标检测中人类可读XAI的视觉语言模型适配
Sardari, Sarvenaz, Fernandes, Freddy, Yelvande, Samarth, Araya-Martinez, Jose Moises, Roitberg, Alina
Abstract
Explainable Artificial Intelligence (XAI) solutions are essential for building trust in AI technologies and their integration in real manufacturing lines. However, most existing methods are tailored to technical experts, limiting their accessibility to diverse user groups such as blue-collar workers in manufacturing lines who use AI for quality control. In this work, we introduce an XAI interface for object detection in industrial manufacturing based on a fine-tuned vision-language model, designed to generate intuitive explanations for non-expert users. We benchmark existing vision-language models and demonstrate that out-of-the-box models often fall short in delivering clear, context-relevant explanations for non-expert users. To address this, we fine-tune a vision-language model and integrate it into our interface, enabling contextualized, accessible explanations for non-expert users. We demonstrate improvements in explanation clarity, instruction adherence, image groundedness, and contextual awareness over GPT 4o-mini on proprietary and public robotics dataset. This approach advances the accessibility and usability of AI explanations, making them more intuitive and applicable in manufacturing domain.
Chinese Translation
可解释人工智能(XAI)解决方案对于建立对AI技术的信任及其在实际制造流水线中的集成至关重要。然而,大多数现有方法是为技术专家量身定制的,限制了其对多样化用户群体(如制造流水线中使用AI进行质量控制的蓝领工人)的可访问性。在本工作中,我们介绍了一种基于微调视觉语言模型的XAI界面,用于工业制造中的目标检测,旨在为非专家用户生成直观的解释。我们对现有的视觉语言模型进行了基准测试,并证明现成的模型往往难以为非专家用户提供清晰、与上下文相关的解释。为了解决这个问题,我们微调了一个视觉语言模型并将其集成到我们的界面中,从而为非专家用户提供情境化、易于理解的解释。我们在专有和公开的机器人数据集上展示了在解释清晰度、指令遵循度、图像依据性和上下文感知方面相较于GPT 4o-mini的改进。这种方法提高了AI解释的可访问性和可用性,使其在制造领域更加直观和适用。
cs.CV / 17 / 2609.31692

Architecture-aware Robustness Evaluation of Explainable Deep Learning for Breast Cancer Diagnosis

架构感知的可解释深度学习用于乳腺癌诊断的鲁棒性评估
Thanusanth, Balenthira, Thuseethan, Selvarajah, Ragel, Roshan G., Weerakoon, Bimali S., Jayasinghe, Ayesh, Krishnamoorthy, Ananthamoorthy
Abstract
Explainable Artificial Intelligence (XAI) has become essential in medical image analysis to ensure transparency of deep learning (DL)-based diagnostic systems. However, selecting appropriate XAI techniques for breast cancer recognition remains largely ad hoc, with limited systematic evaluation across different DL architectures. This study presents a systematic architecture-aware evaluation protocol to assess the effectiveness of nine widely used XAI techniques across four categories of DL models: very deep, lightweight, transformer-based and hybrid neural networks. The evaluation is conducted on a breast ultrasound dataset comprising 780 images using clinically aligned spatial metrics, including Pointing Game, Intersection over Union and Mean Coverage, to quantify agreement between generated explanations and expert-annotated lesion regions. Results indicate that explanation quality is primarily influenced by the interaction between model architecture and XAI method, rather than any single technique consistently outperforming others. Hybrid architectures produce more spatially coherent explanations, while lightweight and transformer-based models exhibit greater variability across methods. The findings show that no single technique generalises across architectures and evaluation criteria, emphasising the need for joint selection of DL models and XAI techniques. Explainability depends on both model design and explanation strategy and should not be considered independently. This work provides a structured evaluation protocol and practical guidance for selecting XAI techniques in breast cancer diagnosis, supporting more transparent clinical decision-support systems. \textcolor{blue}{Code is publicly available at https://github.com/Nishan-Charlie/Explainable-AI}
Chinese Translation
可解释人工智能 (XAI) 在医学图像分析中已变得至关重要,以确保基于深度学习 (DL) 的诊断系统的透明度。然而,为乳腺癌识别选择合适的 XAI 技术在很大程度上仍然是临时的,缺乏跨不同 DL 架构的系统评估。本研究提出了一种系统的架构感知评估协议,以评估九种广泛使用的 XAI 技术在四类 DL 模型(极深、轻量级、基于 Transformer 和混合神经网络)中的有效性。评估在包含 780 张图像的乳腺超声数据集上进行,使用临床对齐的空间指标(包括 Pointing Game、Intersection over Union 和 Mean Coverage)来量化生成的解释与专家标注的病灶区域之间的一致性。结果表明,解释质量主要受模型架构与 XAI 方法之间交互的影响,而不是任何单一技术始终优于其他技术。混合架构产生更具空间连贯性的解释,而轻量级和基于 Transformer 的模型在不同方法之间表现出更大的变异性。研究结果表明,没有单一技术能够在不同架构和评估标准之间泛化,强调需要联合选择 DL 模型和 XAI 技术。可解释性取决于模型设计和解释策略,不应独立考虑。这项工作为乳腺癌诊断中选择 XAI 技术提供了结构化的评估协议和实践指导,支持更透明的临床决策支持系统。代码已在 https://github.com/Nishan-Charlie/Explainable-AI 公开。
cs.CV / 18 / 2609.31693

Disentangle and Drop: Robust Universal Removal of Image Watermarks via Reconstructive Grayscale Residual Decomposition

解耦与丢弃:通过重建性灰度残差分解实现鲁棒的通用图像水印去除
Li, Qi, Yang, Jidong, Fan, Feng-Lei, Miao, Yuantian, Chen, Xiao, Yu, Huaike, Wang, Chunpeng, Gao, Suo, Iu, Herbert Ho-Ching, Ma, Bin
Abstract
Invisible image watermarks are commonly evaluated against benign postprocessing operations such as compression, resizing, blur, and color changes. These tests leave out a different threat: a learned remover that preserves semantic image content while discarding residual evidence that carries the payload. We propose Disentangle and Drop (DnD), an attack that is agnostic to the watermark method and treats watermark removal as a representation routing problem. DnD decomposes a watermarked image into a semantic grayscale carrier and an auxiliary residual branch, and then suppresses the residual branch to reduce watermark evidence. The model is trained with latent spectral perturbations and low-strength diffusion exposure so that the drop operation remains stable under adaptive reconstruction. Experiments on seven representative watermark families show that one shared operating setting gives competitive removal with high visual fidelity. Operating scans and ablations separate usable attacks from image-damaging settings: stronger noise or diffusion can raise removal scores by damaging the image, while the practical regime comes from dropping the residual latent. These results argue for evaluating watermark robustness against learned removal at the representation level, not only against conventional image edits.
Chinese Translation
不可见图像水印通常针对良性后处理操作(如压缩、缩放、模糊和颜色变化)进行评估。这些测试忽略了一种不同的威胁:一种学习到的去除器,它在保留语义图像内容的同时丢弃携带有效载荷的残余证据。我们提出了解耦与丢弃(Disentangle and Drop, DnD),一种与水印方法无关的攻击,它将水印去除视为一个表示路由问题。DnD将水印图像分解为一个语义灰度载体和一个辅助残差分支,然后抑制残差分支以减少水印证据。该模型使用潜在谱扰动和低强度扩散暴露进行训练,以便丢弃操作在自适应重建下保持稳定。在七个具有代表性的水印系列上的实验表明,一个共享的操作设置能够以高视觉保真度实现有竞争力的去除效果。操作扫描和消融实验将可用的攻击与破坏图像的设置区分开来:更强的噪声或扩散可以通过损坏图像来提高去除分数,而实际有效的方案来自于丢弃残差潜在变量。这些结果主张在表示层面评估水印对学习到的去除的鲁棒性,而不仅仅是针对传统图像编辑。
cs.CV / 19 / 2609.31694

RemTraceNet: Few-Shot Forensic Detection of Invisible Watermark Attacks

RemTraceNet:不可见水印攻击的少样本取证检测
Yang, Jidong, Yu, Huaike, Li, Qi, Miao, Yuantian, Zong, Wei, Chow, Yang-Wai, Susilo, Willy, Wang, Chunpeng, Gao, Suo
Abstract
Removing an invisible watermark and concealing the forensic evidence are distinct objectives: successfully disrupting the embedded watermark does not imply that the removal process is forensically undetectable. When verification fails, removal traces can provide complementary evidence for provenance and ownership verification, whereas their absence leaves the cause of the failure ambiguous. Existing methods are typically evaluated by watermark suppression and perceptual quality, while forensic stealth is rarely considered. We therefore study watermark-attack-specific few-shot forensics: for each known pipeline, a specialist can separate its outputs from paired clean and unattacked watermarked controls. Separate Attack-vs-Clean and Attack-vs-Watermarked evaluations prevent watermark-presence shortcuts. Image-aligned and prompt-matched controls are used for post-hoc and generator-integrated schemes, respectively. In this work, we introduce RemTraceNet, which fuses constrained residuals, local relations, FFT/Haar statistics, and block-DCT evidence at native resolution. Across 23 removal pipelines and 10 watermark configurations, we evaluate native 256 x 256 and 512 x 512 inputs. With 100 attacked training images per pipeline, the three-seed TPR@1%FPR, macro-averaged over attacks and watermark configurations, ranges from 82.75% to 88.17% across resolutions and control types. Under the condition of same labels and protocol, RemTraceNet outperforms retrained SRNet, ZhuNet, and SiaStegNet baselines by 10.80--15.68 percentage points. Extensive experimental results show that erasing a watermark and erasing evidence of its removal are distinct challenges, and that removal traces remain learnable under limited supervision.
Chinese Translation
去除不可见水印与隐藏取证证据是不同的目标:成功破坏嵌入的水印并不意味着该去除过程在取证上不可检测。当验证失败时,去除痕迹可以为来源追踪和所有权验证提供补充证据,而此类痕迹的缺失则会使失败原因变得模糊。现有方法通常通过水印抑制和感知质量进行评估,而取证隐蔽性很少被考虑。因此,我们研究针对水印攻击的少样本取证:对于每个已知流程,专用检测器可以将其输出与配对的干净和未受攻击的水印对照样本区分开。分别进行 Attack-vs-Clean 和 Attack-vs-Watermarked 评估可防止依赖水印存在的捷径。对于事后处理方案与生成器集成方案,分别采用图像对齐和提示匹配的对照样本。在这项工作中,我们提出了 RemTraceNet,它在原生分辨率下融合了约束残差、局部关系、FFT/Haar 统计量和块 DCT 证据。在 23 个去除流程和 10 种水印配置上,我们评估了原生 256 x 256 和 512 x 512 输入。在每个流程使用 100 张受攻击训练图像的情况下,跨攻击和水印配置宏平均的三种子 TPR@1%FPR,在不同分辨率和对照类型下介于 82.75% 到 88.17% 之间。在相同标签和协议条件下,RemTraceNet 比重新训练的 SRNet、ZhuNet 和 SiaStegNet 基线高出 10.80--15.68 个百分点。大量实验结果表明,擦除水印和擦除其去除证据是不同的挑战,并且在有限监督下去除痕迹仍然可学习。
cs.CV / 20 / 2609.31697

Video Captioning in Low-Light Conditions through Efficient Uncertainty-Aware Caption Correction

通过高效不确定性感知字幕校正实现低光照条件下的视频字幕生成
Rezaei, Arefeh
Abstract
Low-light conditions can significantly degrade the ability of vision-language models (VLMs) to accurately describe human actions in videos. In this work, I propose an efficient uncertainty-aware representation correction framework for improving captions generated by VideoChat2 under real-world low-light conditions. Instead of fine-tuning the underlying VLM, the proposed framework introduces a lightweight sparse Gaussian process-based error estimation module between the projection layer and the language model to correct the intermediate representation. The correction module learns to estimate the residual between the original projected representation and a verified target representation, which is then adaptively scaled using a newly formulated uncertainty-aware coefficient and added to the original representation. To further improve residual estimation, I introduce a partitioned combined-kernel design. The correction model is trained separately using only 44 samples from the ARID dataset and requires only a small additional computational overhead during inference. The effectiveness of the proposed correction is evaluated through quantitative residual prediction and qualitative analysis of the generated captions. Although VideoChat2 is used in my experiments, the proposed framework is designed to be applicable to other compatible VLM architectures. \textbf{Code Availability: The implementation accompanying this work is publicly available at}:\href{https://github.com/areferezaee/Low-rank-SVGP-NP-update}{https://github.com/areferezaee/Low-rank-SVGP-NP-update}
Chinese Translation
低光照条件会显著降低视觉语言模型(VLMs)准确描述视频中人类动作的能力。在本工作中,我提出了一个高效的不确定性感知表示校正框架,用于改进VideoChat2在真实世界低光照条件下生成的字幕。该框架不是微调底层的VLM,而是在投影层和语言模型之间引入一个轻量级的基于稀疏高斯过程的误差估计模块,以校正中间表示。校正模块学习估计原始投影表示与经过验证的目标表示之间的残差,然后使用新制定的不确定性感知系数进行自适应缩放,并将其添加到原始表示中。为了进一步改进残差估计,我引入了分区组合核设计。校正模型仅使用ARID数据集中的44个样本进行单独训练,并且在推理期间仅需要少量的额外计算开销。通过定量残差预测和生成字幕的定性分析,评估了所提出校正的有效性。虽然我的实验中使用了VideoChat2,但所提出的框架设计适用于其他兼容的VLM架构。代码可用性:本工作的实现已公开提供在:https://github.com/areferezaee/Low-rank-SVGP-NP-update
cs.CV / 21 / 2609.31698

RPA: Residual Patch-Token Adapter for Image Retrieval from EEG and MEG

RPA:用于从EEG和MEG进行图像检索的残差图像块令牌适配器
Jin, Yuhui, Song, Yonghao, Liu, Bingchuan
Abstract
Most existing MEG and EEG (M/EEG) visual decoding methods align brain signals with a single global embedding extracted from a pretrained visual encoder, leaving open whether intermediate patch representations, which preserve richer and more granular rich visual information, can improve representation learning. To address this question, we introduce the Residual Patch Adapter (RPA), a lightweight, modular adapter that leverages all patch tokens from an intermediate layer of a ViT visual encoder for alignment. Through extensive ablation analyses, we first show that pooling or masking patch tokens degrades the learned representation, demonstrating that retaining the full set of patch tokens is important for EEG alignment, while the CLS token provides little unique information. We then use a series of six quantitative feature analyses to show that both higher-level semantics and lower-level visual features, including color and texture, are essential for this EEG-to-image alignment. Under current protocols, our system achieves Top-1 accuracies of 95.4\% within-subject and 35.5\% cross-subject on THINGS-EEG2, and 65.2\% and 6.7\%, respectively, on THINGS-MEG, achieving state-of-the-art (SOTA) performance across both datasets. Evaluations with alternative brain encoders, including pretrained EEG foundation models, demonstrate that the approach extends beyond the projection-based EEG encoder. Furthermore, we provide a plug-and-play interface that allows RPA to be replaced by convolution, attention, or ConvNeXt alternatives. Together, these findings provide significant insight into M/EEG-to-image representation learning by establishing design principles for leveraging the latent space of visual encoders, and open new directions for brain--image alignment and non-invasive brain--computer interface (BCI).
Chinese Translation
大多数现有的MEG和EEG(M/EEG)视觉解码方法将大脑信号与从预训练视觉编码器中提取的单一全局嵌入对齐,使得中间图像块表示(保留了更丰富、更细粒度的视觉信息)能否改进表示学习成为一个悬而未决的问题。为了解决这个问题,我们引入了残差图像块适配器(RPA),这是一个轻量级、模块化的适配器,利用ViT视觉编码器中间层的所有图像块令牌进行对齐。通过广泛的消融分析,我们首先表明池化或掩码图像块令牌会降低学习到的表示,证明保留完整的图像块令牌集对EEG对齐很重要,而CLS令牌提供的独特信息很少。然后,我们使用一系列六项定量特征分析表明,高级语义和低级视觉特征(包括颜色和纹理)对于这种EEG到图像对齐都是必不可少的。在当前协议下,我们的系统在THINGS-EEG2上实现了95.4%的被试内和35.5%的跨被试Top-1准确率,在THINGS-MEG上分别实现了65.2%和6.7%,在两个数据集上均达到了最先进的(SOTA)性能。使用替代大脑编码器(包括预训练EEG基础模型)的评估表明,该方法不仅限于基于投影的EEG编码器。此外,我们提供了一个即插即用接口,允许将RPA替换为卷积、注意力或ConvNeXt替代方案。总之,这些发现通过建立利用视觉编码器潜在空间的设计原则,为M/EEG到图像表示学习提供了重要见解,并为大脑-图像对齐和非侵入式脑机接口(BCI)开辟了新方向。
cs.CV / 22 / 2609.31700

Modernising the Compressed-Domain Video Captioner: A Controlled Study of SigLIP2 and GPT-2 Substitutions

现代化压缩域视频字幕生成器:SigLIP2与GPT-2替换的对照研究
Nepal, Ashim, K, Ashok B.
Abstract
Compressed-domain video captioning avoids full video decoding by operating directly on I-frames, motion vectors and residuals, trading a small amount of accuracy for a large gain in inference speed. CoCap established this pipeline using a CLIP vision encoder and a shallow BERT-style multimodal decoder. Both components predate substantially stronger alternatives. We ask a narrow, controlled question: how much of CoCap's accuracy is limited by these two components, and which of the two is the binding constraint? We replace the CLIP I-frame encoder with SigLIP2 and the BERT-style decoder with GPT-2, and evaluate three configurations (the original pairing, the encoder substitution alone, and both substitutions together) under identical data, sampling budget and optimisation schedule. All comparisons are made against our own reproduction of CoCap rather than its published numbers, because we train on a 4,999-clip subset of VATEX at a reduced sampling budget; absolute values are therefore not comparable with the original work. Our reproduction tracks the published result closely: CIDEr and METEOR run slightly above it (54.9 against 52.7; 23.4 against 23.2), BLEU-4 and ROUGE-L slightly below (29.7 against 31.4; 48.9 against 49.4). We attribute the differences to our evaluation subset rather than to any improvement in either direction. We find the two substitutions pull in opposite directions: SigLIP2 alone improves every metric (+4.5 CIDEr), while adding GPT-2 on top erodes that gain, because a pretrained decoder overfits 4,999 clips within two epochs. We additionally report inference latency for each configuration, since speed is the property that motivates compressed-domain captioning in the first place, and an accuracy gain purchased at a latency cost should be reported as such.
Chinese Translation
压缩域视频字幕生成通过直接操作I帧、运动矢量和残差来避免完整视频解码,以少量精度换取推理速度的大幅提升。CoCap使用CLIP视觉编码器和浅层BERT风格的多模态解码器建立了这一流程。这两个组件都早于更强大的替代方案。我们提出一个狭义的、受控的问题:CoCap的精度在多大程度上受限于这两个组件,以及两者中哪一个是主要约束?我们将CLIP I帧编码器替换为SigLIP2,将BERT风格解码器替换为GPT-2,并在相同的数据、采样预算和优化调度下评估三种配置(原始配对、仅替换编码器、以及两者同时替换)。所有比较都是针对我们自己复现的CoCap,而不是其已发表的数据,因为我们是在VATEX的4,999个视频片段子集上以降低的采样预算进行训练;因此,绝对值无法与原始工作直接比较。我们的复现与已发表结果非常接近:CIDEr和METEOR略高于它(54.9对52.7;23.4对23.2),BLEU-4和ROUGE-L略低于它(29.7对31.4;48.9对49.4)。我们将这些差异归因于我们的评估子集,而不是任何方向上的改进。我们发现这两种替换方向相反:单独使用SigLIP2提升了所有指标(CIDEr +4.5),而在其基础上加入GPT-2则侵蚀了该增益,因为预训练解码器在两个epoch内就过拟合了4,999个片段。我们额外报告了每种配置的推理延迟,因为速度是首先推动压缩域字幕生成的因素,而以延迟成本换取的精度增益也应按此报告。
cs.CV / 23 / 2609.31701

Integrated Deep Learning Framework Designed on Hybrid Optimization Strategies for Automated Health Detection and Analysis in Silkworms

基于混合优化策略设计的集成深度学习框架用于家蚕自动化健康检测与分析
K V, Komala, T, Lata B, R, Venugopal K
Abstract
A Hybrid Residual Attention Network is proposed for accurately classifying silkworm images into six different classes, including healthy and diseased states. It uses residual blocks for deep feature extraction and attention to focus on disease related features. A novel Integrated Adaptive Momentum Optimizer was introduced to enhance convergence and improve training efficiency. The dataset of silkworm images underwent preprocessing techniques such as normalization, resizing, and noise reduction, along with augmentation strategies to improve data quality and diversity. It is optimized using IAMO, achieved an accuracy of 98.67%.The integration of spatial and channel wise attention mechanisms, coupled with IAMO, significantly enhanced the model ability to recognize subtle differences between classes. Results indicate that HRAN can be used to detect disease at an early stage in sericulture, and future work will enhance scalability and efficiency in different environments.
Chinese Translation
提出了一种混合残差注意力网络(HRAN),用于将家蚕图像准确分类为六种不同的类别,包括健康和患病状态。它使用残差块进行深度特征提取,并利用注意力机制关注疾病相关特征。引入了一种新颖的集成自适应动量优化器(IAMO),以增强收敛性并提高训练效率。家蚕图像数据集经过预处理技术,如归一化、调整大小和降噪,以及增强策略以提高数据质量和多样性。使用IAMO进行优化,达到了98.67%的准确率。空间和通道注意力机制的集成,加上IAMO,显著增强了模型识别类别之间细微差异的能力。结果表明,HRAN可用于在蚕业中早期检测疾病,未来的工作将增强不同环境中的可扩展性和效率。
cs.CV / 24 / 2609.31702

CLC-YOLO: A Compact Channel-Gated Prototype Network for Real-Time Leakage-Aware Breast Ultrasound Lesion Segmentation

CLC-YOLO:一种用于实时泄漏感知乳腺超声病变分割的紧凑通道门控原型网络
Nizar, M. Fazri, Rachmatullah, Muhammad Naufal, Supardi, Julian
Abstract
Reliable breast ultrasound lesion segmentation requires accurate boundaries and evaluation that prevents patients or duplicate images from crossing data splits. We propose Channel Local Contrast (CLC), a compact refinement of the YOLO26 segmentation prototype head. CLC adds a fixed local high-pass residual controlled by 64 zero-initialized, bounded channel gates. Baseline and CLC were compared in five matched folds on each of four breast ultrasound datasets. BUS-BRA used patient-disjoint outer tests with separate inner validation. BUS-UCLM and BrEaST used patient-grouped validation folds; BUSI used duplicate-component groups because patient identifiers are unavailable. Group-macro Dice increased by 1.68, 3.11, 1.12, and 2.45 percentage points on BUS-BRA, BUS-UCLM, BUSI, and BrEaST, respectively. Only the BUS-BRA paired 95% confidence interval excluded zero. CLC adds 64 parameters and 0.0049 giga floating-point operations (GFLOPs). At 640 pixels, single-T4, batch-one, 16-bit floating-point (FP16) TensorRT graph times were 2.1478 ms for CLC and 2.0512 ms for baseline, excluding preprocessing and postprocessing. CLC increased group-macro Dice across all four datasets with a measured T4 forward-pass overhead of 0.0966 ms. Code: https://github.com/mfazrinizar/CLC-YOLO
Chinese Translation
可靠的乳腺超声病变分割需要准确的边界以及能够防止患者或重复图像跨越数据划分的评估。我们提出了通道局部对比度(CLC),这是对YOLO26分割原型头的紧凑改进。CLC添加了一个由64个零初始化、有界通道门控制的固定局部高通残差。在四个乳腺超声数据集的每一个上,都使用五个匹配的折比较了基线和CLC。BUS-BRA使用患者不相交的外层测试和单独的内部验证。BUS-UCLM和BrEaST使用患者分组的验证折;BUSI使用重复组件分组,因为患者标识符不可用。组宏Dice在BUS-BRA、BUS-UCLM、BUSI和BrEaST上分别增加了1.68、3.11、1.12和2.45个百分点。只有BUS-BRA的配对95%置信区间不包含零。CLC增加了64个参数和0.0049吉浮点运算(GFLOPs)。在640像素下,单T4、批大小为1、16位浮点(FP16)TensorRT图时间,CLC为2.1478毫秒,基线为2.0512毫秒,不包括预处理和后处理。CLC在所有四个数据集上提高了组宏Dice,测得的T4前向传递开销为0.0966毫秒。代码:https://github.com/mfazrinizar/CLC-YOLO
cs.CV / 25 / 2609.31705

MDL-Calibrated Significance-Gain Pair Encoding: Replication-Aware Automatic Stopping for Subword Tokenization

MDL校准的显著性增益对编码:面向子词分词的复制感知自动停止
Nouri, Azam
Abstract
Byte-Pair Encoding (BPE) constructs subword vocabularies through greedy pair merging, but conventional BPE requires the number of merges or target vocabulary size to be specified externally. Significance-Gain Pair Encoding (SG-BPE) replaces frequency-only selection with a statistical criterion based on how strongly an observed pair exceeds its expected co-occurrence under an independence model. This paper introduces MDL-Calibrated Significance-Gain Pair Encoding (MDL-SG), a three-stage procedure separating discovery, replication, and utility. Candidate pairs are ranked by Significance-Gain on a discovery partition, tested for replication on a separate partition using an exact one-sided hypergeometric test with per-iteration Benjamini-Hochberg correction, and then evaluated on a utility partition using a Minimum Description Length (MDL) criterion. Merging stops automatically when no replicated candidate yields positive held-out MDL gain. On WikiText-103, MDL-SG stops at 209, 433, and 847 merges for 120K, 250K, and 500K-character tokenizer-training samples, respectively. At 500K characters, it selects a stored vocabulary of 1,017 tokens without prescribing the vocabulary size in advance. In a compute-matched TinyGPT experiment with identical 2,024,448-parameter models and 500 optimizer updates per language model, MDL-SG achieves validation/test BPC of 3.2612/3.2436, compared with test BPC of 3.2894 for SG-BPE and 3.3493 for frequency BPE. Frequency BPE achieves stronger raw compression, while MDL-SG achieves lower BPC, showing that compression-oriented merge selection and language-model utility need not coincide.
Chinese Translation
字节对编码(BPE)通过贪心对合并构建子词词汇表,但传统BPE需要外部指定合并次数或目标词汇表大小。显著性增益对编码(SG-BPE)用基于观测对在独立模型下超过其期望共现强度的统计准则,取代了仅基于频率的选择。本文引入MDL校准的显著性增益对编码(MDL-SG),一种将发现、复制和效用分离的三阶段流程。候选对在发现分区上按显著性增益排序,在单独分区上使用精确单侧超几何检验进行复制测试,每次迭代进行Benjamini-Hochberg校正,然后使用最小描述长度(MDL)准则在效用分区上进行评估。当没有复制候选产生正留出MDL增益时,合并自动停止。在WikiText-103上,对于120K、250K和500K字符的分词器训练样本,MDL-SG分别在209、433和847次合并时停止。在500K字符时,它选择了存储的1,017个标记的词汇表,而无需预先指定词汇表大小。在计算匹配的TinyGPT实验中,使用相同的2,024,448参数模型和每个语言模型500次优化器更新,MDL-SG达到了验证/测试BPC为3.2612/3.2436,而SG-BPE的测试BPC为3.2894,频率BPE为3.3493。频率BPE实现了更强的原始压缩,而MDL-SG实现了更低的BPC,表明面向压缩的合并选择和语言模型效用不必一致。
cs.CV / 26 / 2609.31706

Can't Find Waldo: Evaluating VLMs' Sensitivity to Image Resolution and Detail Level

找不到Waldo:评估VLM对图像分辨率和细节水平的敏感性
Schild, Alexandra, de Melo, Gerard
Abstract
Visual Language Models (VLMs) have achieved remarkable success across diverse tasks, yet they struggle with high-resolution inputs where critical information resides in small regions or detailed, cluttered scenes. While several approaches address this limitation, a systematic understanding of why models fail at high resolutions is lacking. We introduce a controlled evaluation framework that disentangles resolution-related performance degradation from task difficulty through semantics-preserving transformations. We propose two simple metrics: Area Under the Scaling Curve (AUSC), which quantifies scaling robustness independent of baseline accuracy, and Prediction Variance Score (PVS), which measures resolution-induced prediction instability. Through comprehensive experiments across 5 model families and 5 benchmarks, we identify three primary failure modes: (1) information loss from downsampling at vision token limits, (2) tokenization artifacts from patch boundary shifts and positional encoding fragility under non-standard aspect ratios, and (3) attention dilution as token counts increase. Our analysis reveals that even state-of-the-art models suffer from performance drops when processing high-resolution images, with degradation patterns varying systematically by architectural family. We provide actionable insights for model architecture design and data augmentation strategies to mitigate these limitations.
Chinese Translation
视觉语言模型(VLM)在各种任务中取得了显著成功,但在处理高分辨率输入时仍存在困难,因为这些输入的关键信息往往位于小区域或细节丰富、杂乱的场景中。尽管已有若干方法试图解决这一局限,但对于模型在高分辨率下为何失败,仍缺乏系统的理解。我们引入了一个受控评估框架,通过保持语义不变的变换,将分辨率相关的性能下降与任务难度分离开来。我们提出了两个简单的指标:缩放曲线下面积(AUSC),它量化了独立于基线准确率的缩放鲁棒性;以及预测方差分数(PVS),它衡量分辨率引起的预测不稳定性。通过在5个模型家族和5个基准上的全面实验,我们识别出三种主要的失败模式:(1)在视觉token限制下因下采样导致的信息丢失,(2)非标准宽高比下patch边界偏移和位置编码脆弱性导致的token化伪影,以及(3)随着token数量增加而出现的注意力稀释。我们的分析揭示,即使是最先进的模型在处理高分辨率图像时也会出现性能下降,且退化模式因架构家族的不同而呈现系统性差异。我们为模型架构设计和数据增强策略提供了可操作的见解,以缓解这些局限。
cs.CV / 27 / 2609.31709

Cross-Dataset Generalization of Bangladeshi Rice Leaf Disease Classifiers: Benchmark, Diagnosis, and Mitigation

孟加拉国水稻叶病分类器的跨数据集泛化:基准、诊断与缓解
Paul, Anindya
Abstract
Cross-dataset transfer in rice leaf disease classification remains a significant challenge, with models trained on one image collection performing substantially worse when deployed on another. We conduct a systematic benchmark across three Bangladeshi rice leaf disease datasets (5,419 images, 6 transfer pairs, 3 CNN backbones, 3 random seeds) to characterize and diagnose this failure. Strong augmentation recovers a mean cross-dataset macro-F1 improvement of +0.070 (Wilcoxon p < 0.001, 15 of 18 transfer pairs positive). Removing non-leaf image content via segmentation shows directional benefit (mean +0.066, p = 0.062, n = 36 paired observations) that is consistent across two independent segmentation methods but does not reach conventional significance. A self-supervised ViT control (DINOv2 linear probe) exhibits equivalent cross-dataset collapse to CNNs, ruling out architecture inductive bias as the primary driver and pointing to acquisition-condition shift. Adaptive batch normalization uniformly harms transfer performance, with harm magnitude correlating with source-target label-prior divergence and model depth (Spearman rho = 0.621, p = 0.009). Grad-CAM attribution analysis on 12 sampled predictions does not distinguish correct from incorrect cross-domain predictions (p = 0.462), indicating that common attribution proxies are insufficient for diagnosing shift at practical sample sizes. We document all frozen results, prespecified analysis criteria, and reproducibility artifacts in a public repository with SHA-256 integrity verification. This work establishes a rigorous empirical baseline for understanding cross-dataset generalization in agricultural computer vision and identifies both effective (augmentation) and ineffective (AdaBN) adaptation strategies.
Chinese Translation
水稻叶病分类中的跨数据集迁移仍然是一个重大挑战,在一个图像集合上训练的模型在部署到另一个图像集合时表现显著变差。我们对三个孟加拉国水稻叶病数据集(5,419 张图像、6 个迁移对、3 个 CNN 骨干网络、3 个随机种子)进行了系统性的基准测试,以刻画和诊断这一失败。强数据增强带来了平均跨数据集 macro-F1 提升 +0.070(Wilcoxon p < 0.001,18 个迁移对中有 15 个为正)。通过分割去除非叶片图像内容显示出方向性收益(平均 +0.066,p = 0.062,n = 36 对观测),在两种独立分割方法中一致,但未达到常规显著性水平。一个自监督 ViT 对照(DINOv2 线性探针)表现出与 CNN 等效的跨数据集崩溃,排除了架构归纳偏置作为主要驱动因素,并指向采集条件偏移。自适应批归一化(Adaptive batch normalization, AdaBN)一致地损害迁移性能,损害程度与源-目标任务标签先验差异和模型深度相关(Spearman rho = 0.621, p = 0.009)。对 12 个采样预测的 Grad-CAM 归因分析无法区分正确与错误的跨域预测(p = 0.462),表明常见的归因代理在实际样本量下不足以诊断偏移。我们在一个公共存储库中记录了所有冻结结果、预先指定的分析标准以及可复现性工件,并带有 SHA-256 完整性验证。这项工作为理解农业计算机视觉中的跨数据集泛化建立了一个严格的实证基线,并识别出有效(数据增强)和无效(AdaBN)的适应策略。
cs.CV / 28 / 2609.31712

Statistical Testing for Multiple Instance Learning via Selective Inference with Applications to Computational Pathology

通过选择性推断的多实例学习统计检验及其在计算病理学中的应用
Hashimoto, Noriaki, Nishino, Shuichi, Katsuoka, Teruyuki, Shiraishi, Tomohiro, Miwa, Daiki, Hanada, Hiroyuki, Sakuma, Jun, Hontani, Hidekata, Miyoshi, Hiroaki, Takeuchi, Ichiro
Abstract
Multiple instance learning (MIL) is widely used in computational pathology because it enables weakly supervised analysis of whole-slide images (WSIs) without requiring patch-level annotations. In attention-based MIL, instances with high attention scores are often interpreted as diagnostically important regions and used as visual explanations. However, attention scores alone cannot determine whether selected high-attention instances are significantly different from normal instances, limiting the reliability of attention-based explanations. In this paper, we formulate the evaluation of high-attention instances as a statistical hypothesis testing problem. Specifically, we assess whether a selected high-attention instance significantly deviates from a representative normal reference instance selected based on feature similarity. A major challenge is that both the target instance and the reference instance are selected through data-dependent procedures, rendering standard hypothesis testing invalid. To address this issue, we introduce a selective inference (SI) framework that explicitly accounts for the selection events induced by attention-based instance selection and adaptive reference selection, thereby enabling the computation of valid selective $p$-values conditional on these events. Experiments demonstrate Type-I error control on synthetic and MNIST-based data and practical applicability to pathological WSIs, with higher statistical power than the conventional over-conditioning approach.
Chinese Translation
多实例学习(MIL)在计算病理学中广泛应用,因为它能够对全切片图像(WSIs)进行弱监督分析,而无需块级标注。在基于注意力的MIL中,注意力得分高的实例通常被解释为诊断重要区域,并用作视觉解释。然而,仅凭注意力得分无法确定所选的高注意力实例是否与正常实例有显著差异,这限制了基于注意力的解释的可靠性。在本文中,我们将高注意力实例的评估形式化为一个统计假设检验问题。具体而言,我们评估一个选定的高注意力实例是否显著偏离基于特征相似性选出的代表性正常参考实例。一个主要挑战是,目标实例和参考实例都是通过数据依赖的过程选出的,这使得标准假设检验无效。为了解决这个问题,我们引入了一个选择性推断(SI)框架,该框架明确考虑了由基于注意力的实例选择和自适应参考选择所引发的选择事件,从而能够计算在这些事件条件下的有效选择性$p$值。实验表明,在合成数据和基于MNIST的数据上实现了第一类错误控制,并且对病理WSIs具有实际适用性,其统计功效高于传统的过度条件化方法。
cs.CV / 29 / 2609.31713

Agentic Video Understanding: A Survey

智能体式视频理解:综述
Deng, Xinyu, Luo, Siwen, Liu, Daochang
Abstract
As large language models (LLMs) become capable of processing increasingly diverse modalities and longer temporal contexts, an emerging line of work is moving beyond fixed video-language inference toward agentic systems that actively decide what information to inspect, retain, verify, and act upon. This survey reviews video understanding agents: systems that use video as the primary information source and solve understanding tasks through adaptive state construction and action selection. We first formalize an agent loop for video understanding, then address a central question: why do agents matter for video understanding? To answer this, we organize the literature through a challenge-to-design taxonomy, linking context bottlenecks to hierarchical evidence memory, evidence sparsity to active evidence acquisition, temporal causality to state and process tracking, and multimodal ambiguity to role-specialized coordination. We further review state space paradigms, learning paradigms, supervision signals, benchmarks, and evaluation protocols. Finally, we identify open directions toward agentic-native temporal modeling and video-native agents. Project page: https://github.com/DXY0711/Awesome-Agentic-Video-Understanding
Chinese Translation
随着大语言模型(LLMs)能够处理日益多样化的模态和更长的时序上下文,一个新兴的研究方向正超越固定的视频-语言推理,转向能够主动决定检查、保留、验证哪些信息并据此采取行动的智能体系统。本综述回顾了视频理解智能体:这些系统将视频作为主要信息源,并通过自适应状态构建和动作选择来解决理解任务。我们首先形式化了一个用于视频理解的智能体循环,然后探讨一个核心问题:为什么智能体对视频理解至关重要?为了回答这个问题,我们通过一个从挑战到设计的分类体系来组织文献,将上下文瓶颈与分层证据记忆相关联,将证据稀疏性与主动证据获取相关联,将时间因果性与状态和过程跟踪相关联,将多模态歧义与角色专门化协调相关联。我们进一步回顾了状态空间范式、学习范式、监督信号、基准测试和评估协议。最后,我们指出了面向智能体原生时序建模和视频原生智能体的开放方向。项目页面:https://github.com/DXY0711/Awesome-Agentic-Video-Understanding
cs.CV / 30 / 2609.31714

OmniFysics-Captioner Technical Report: Grounding Omni-Modal Understanding in the Physical World for Better Captioning

OmniFysics-Captioner 技术报告:将全模态理解锚定于物理世界以实现更好的描述生成
Qiu, Kaixiang, Han, Minghao, Liu, Keliang, Liu, Yizhou, Han, Jinghan, Jiang, Yue, Wu, Xuecheng, Wang, Shunli, Zhang, Lihua, Yang, Dingkang
Abstract
Building omni-modal models with physical intelligence requires fine-grained supervision that captures physical evidence such as contact, support, deformation, and state transitions. However, existing omni-modal captioners primarily model general audiovisual semantics and often overlook transient or spatially localized physical evidence. We present a unified framework for physics-aware audiovisual captioning spanning data construction, training, and evaluation. Firstly, we build a data construction pipeline that identifies physics-rich clips and leverages OmniFysics-Agent to coordinate audio, visual, and physical-perception tools for collecting spatiotemporally aligned and traceable cross-modal evidence; within the Agent, a physical perception model (PPM) fine-tuned on approximately 2M image-level samples serves as a dedicated tool for extracting object-interaction and state-change cues. Secondly, we build the Daily-Physics 50K dataset and introduce the evidence-driven OmniPhysCap (OPC) benchmark to evaluate the recovery of physical and cross-modal evidence from generated captions. Finally, we train OmniFysics-Captioner from the resulting data. Our Captioner matches Gemini 3.1 Pro on audiovisual captioning, achieves state-of-the-art results on multiple video-captioning benchmarks, and substantially outperforms other open-source models. Ablations show that PPM evidence improves physical coverage and produces finer-grained, more reliable cross-modal descriptions.
Chinese Translation
构建具有物理智能的全模态模型需要细粒度的监督,以捕获诸如接触、支撑、形变和状态转换等物理证据。然而,现有的全模态描述器主要建模一般的视听语义,往往忽略瞬态或空间局部化的物理证据。我们提出了一个统一的物理感知视听描述框架,涵盖数据构建、训练和评估。首先,我们构建了一个数据构建流程,该流程识别富含物理信息的片段,并利用 OmniFysics-Agent 协调音频、视觉和物理感知工具,以收集时空对齐且可追溯的跨模态证据;在该 Agent 内部,一个在约 200 万图像级样本上微调的物理感知模型(PPM)作为专用工具,用于提取物体交互和状态变化线索。其次,我们构建了 Daily-Physics 50K 数据集,并引入了证据驱动的 OmniPhysCap (OPC) 基准,以评估从生成的描述中恢复物理和跨模态证据的能力。最后,我们利用所得数据训练了 OmniFysics-Captioner。我们的 Captioner 在视听描述上可与 Gemini 3.1 Pro 匹敌,在多个视频描述基准上取得了最先进的结果,并显著优于其他开源模型。消融实验表明,PPM 证据提高了物理覆盖率,并产生更细粒度、更可靠的跨模态描述。
cs.CV / 31 / 2609.31715

Fysiverse-3D-SimReady Technical Report: Agentic Physical Simulation for Pragmatic 3D World Reconstruction

Fysiverse-3D-SimReady 技术报告:面向实用三维世界重建的智能体物理仿真
Wang, Lintao, Sun, Mingyang, Liu, Yang, Yang, Dingkang, Zhang, Lihua
Abstract
Agentic recognition requires visual perception to move beyond static scene understanding and produce structured scene representations that support the perception--reasoning--action loop. Existing single-image 3D generation methods, however, mainly produce visually plausible object assets rather than simulation-ready scene states. When independently generated meshes are composed in a shared space, they may fail to align with the input camera, violate gravity, interpenetrate nearby objects, or become unstable under physics simulation. We present Fysiverse-3D-SimReady, a grounded refinement framework for reconstructing simulation-ready multi-object scenes from a single RGB image with instance and ground prompts. The method places generated object meshes into a shared gravity-aligned scene, refines their camera-space poses through differentiable rendering, and corrects scene-level supports and contacts for stable physical execution. The scene is then used by an agentic simulation workflow, which converts a scene-specific task goal into an executable physics rollout rendered from the original camera view. Experiments show that Fysiverse-3D-SimReady improves input-view alignment, contact plausibility, and physical stability over existing single-image reconstruction and scene generation baselines, while enabling goal-conditioned physical interactions from a single image.
Chinese Translation
智能体识别要求视觉感知超越静态场景理解,并生成支持“感知—推理—行动”循环的结构化场景表示。然而,现有单图像三维生成方法主要生成视觉上看似合理的物体资产,而非可直接用于仿真的场景状态。当独立生成的网格被组合到共享空间中时,它们可能无法与输入相机对齐、违反重力、与邻近物体相互穿插,或在物理仿真下变得不稳定。我们提出 Fysiverse-3D-SimReady,这是一种接地式精修框架,用于从单张 RGB 图像出发,借助实例提示和地面提示重建可直接用于仿真的多物体场景。该方法将生成的物体网格放入共享的重力对齐场景中,通过可微渲染精修其相机空间位姿,并修正场景级支撑与接触关系,以实现稳定的物理执行。随后,该场景被智能体仿真工作流使用,将场景特定任务目标转换为可执行的物理推演,并从原始相机视角进行渲染。实验表明,与现有单图像重建和场景生成基线相比,Fysiverse-3D-SimReady 提升了输入视角对齐、接触合理性和物理稳定性,同时能够从单张图像实现目标条件化的物理交互。
cs.CV / 32 / 2609.31716

PanOVOcc: Panoramic Embodied Open-Vocabulary Occupancy Mapping with Long-term Spatial Voxel Memory

PanOVOcc:具有长期空间体素记忆的全景具身开放词汇占用映射
Kuang, Di, Duan, Mengfei, Wang, Yuhang, Peng, Weixing, Yang, Kailun
Abstract
Persistent semantic occupancy mapping is essential for embodied scene understanding. However, perspective-based systems provide limited spatial coverage, while existing panoramic methods primarily predict local volumes from single observations. We introduce PanOVOcc, a training-free framework for persistent open-vocabulary semantic occupancy mapping from panoramic sequences. PanOVOcc unifies panoramic SLAM, open-vocabulary perception, and long-term spatial voxel memory within an online architecture, continuously integrating geometric and semantic evidence into a global, language-queryable map. To facilitate systematic evaluation of this setting, we establish Pan-Replica and Pan-Holo360D, two benchmarks pairing continuous panoramic RGB-D sequences with scene-level semantic occupancy ground truth across synthetic and real-world scenes. Compared with the strongest evaluated baseline for each metric, PanOVOcc improves occupancy IoU and semantic mIoU by absolute +20.03 and +7.06 on Pan-Replica, and by +43.26 and +20.16 on Pan-Holo360D, respectively. The source code and the established benchmarks will be available at https://github.com/bakereet/PanOVOcc.
Chinese Translation
持久的语义占用映射对于具身场景理解至关重要。然而,基于透视的系统提供有限的空间覆盖,而现有的全景方法主要从单次观测预测局部体积。我们提出了PanOVOcc,一个无需训练的框架,用于从全景序列中进行持久的开放词汇语义占用映射。PanOVOcc在在线架构中统一了全景SLAM、开放词汇感知和长期空间体素记忆,持续将几何和语义证据整合到一个全局的、可语言查询的地图中。为了促进对此设置的系统评估,我们建立了Pan-Replica和Pan-Holo360D,这两个基准将连续的全景RGB-D序列与合成和真实场景中的场景级语义占用真值配对。与每个指标的最强评估基线相比,PanOVOcc在Pan-Replica上分别将占用IoU和语义mIoU绝对提高了+20.03和+7.06,在Pan-Holo360D上分别提高了+43.26和+20.16。源代码和建立的基准将在https://github.com/bakereet/PanOVOcc上提供。
cs.CV / 33 / 2609.31717

PanoFuse: Panorama-Enhanced Vision-Language-Action Learning with Decoupled Semantic-Geometric Routing

PanoFuse:全景增强的视觉-语言-动作学习与解耦语义-几何路由
Xu, Peng, Lin, Haoran, Jia, Wanjun, Luo, Kai, Chen, Wenrui, Li, Zhiyong, Yang, Kailun
Abstract
Vision-Language-Action (VLA) policies have shown promising performance in language-conditioned robotic manipulation. However, most existing VLA systems rely on conventional perspective cameras with limited fields of view, often missing global scene context and leading to unreliable manipulation under visual occlusions, distractors, and unseen environments. In this work, we propose PanoFuse, a panorama-enhanced VLA framework that complements local manipulation observations with global panoramic perception. PanoFuse introduces a dedicated panoramic branch that leverages a pretrained panoramic foundation model to extract complementary semantic and geometric representations from omnidirectional observations. Rather than directly mixing these heterogeneous features, we introduce Decoupled Semantic-Geometric Routing (DSGR), which maintains semantic and geometric representations as separate context streams and selectively routes both to downstream state and action representations through structured block-wise attention. This design provides the action expert with global spatial context while preserving task-relevant semantic information from the pretrained VLA backbone. We further develop a synchronized data collection pipeline and construct a new real-world manipulation dataset containing panoramic RGB observations, wrist-view images, language instructions, robot states, and actions. Across seven evaluation settings, PanoFuse achieves an average success rate of 52.9%, outperforming the evaluated baselines and achieving consistent gains under novel-object, unseen-background, and distractor-rich settings. Code and data will be released publicly at https://xux-hnu.github.io/PanoFuse.
Chinese Translation
视觉-语言-动作(VLA)策略在语言条件机器人操作中已展现出良好的性能。然而,大多数现有的VLA系统依赖于视野有限的传统透视相机,常常缺失全局场景上下文,导致在视觉遮挡、干扰物和未见环境下操作不可靠。在本工作中,我们提出PanoFuse,一种全景增强的VLA框架,通过全局全景感知来补充局部操作观测。PanoFuse引入了一个专用的全景分支,利用预训练的全景基础模型从全向观测中提取互补的语义和几何表示。我们没有直接混合这些异构特征,而是引入了解耦语义-几何路由(DSGR),它将语义和几何表示保持为独立的上下文流,并通过结构化的分块注意力选择性地将两者路由到下游状态和动作表示。这种设计为动作专家提供了全局空间上下文,同时保留了来自预训练VLA主干的与任务相关的语义信息。我们进一步开发了同步数据采集流水线,并构建了一个新的真实世界操作数据集,包含全景RGB观测、腕部视图图像、语言指令、机器人状态和动作。在七个评估设置中,PanoFuse平均成功率达到52.9%,优于所评估的基线,并在新物体、未见背景和富含干扰物的设置下取得了一致的增益。代码和数据将在https://xux-hnu.github.io/PanoFuse公开。
cs.CV / 34 / 2609.31720

LatentReRig: An SDF-Based VAE with Dual Decoders for Latent-Space Deformation Conditioning

LatentReRig:一种面向潜在空间变形条件化的基于SDF的双解码器VAE
Dolci, Daniele, Poggioni, Fabrizio, Melchiorri, Carlo
Abstract
Transferring deformation between characters with different geometry and topology is challenging because conventional rigs encode behaviour through character-specific structures and correspondences. We present LatentReRig, an experimental framework that investigates whether pose-associated changes can instead be represented as reusable directions in a learned geometric latent space. An SDF-based variational autoencoder is coupled with two decoders: one reconstructs the implicit field, while the other predicts target vertex positions from source geometry and latent deformation conditioning. The source geometry may be neutral or already deformed. Experiments on a controlled humanoid dataset show that several poses induce coherent latent directions across identities, particularly for broad articulated motions. These signals can guide deformation of unseen characters, but explicit predictions remain less accurate for localized changes and corrective contributions. Diagnostic comparisons with repeated SDF sampling show that inter-identity distances exceed same-geometry resampling variability on average, while pose signals exhibit different margins above this baseline. The results support the presence of reusable pose-related structure and identify stable local conditioning and accurate mesh decoding as complementary requirements for improving transfer.
Chinese Translation
在不同几何和拓扑的角色之间传递变形具有挑战性,因为传统绑定通过角色特定的结构和对应关系来编码行为。我们提出LatentReRig,一个实验框架,用于研究姿态相关变化是否可以改为在学习到的几何潜在空间中表示为可重用的方向。一个基于SDF的变分自编码器与两个解码器耦合:一个重建隐式场,另一个从源几何和潜在变形条件预测目标顶点位置。源几何可以是中性姿态或已经变形。在受控人形数据集上的实验表明,若干姿态在不同身份之间诱导出一致的潜在方向,特别是对于大范围的关节运动。这些信号可以指导未见角色的变形,但显式预测对于局部变化和矫正贡献仍然不够准确。与重复SDF采样的诊断性比较表明,身份间距离平均超过相同几何重采样变异性,而姿态信号在该基线上表现出不同的余量。结果支持可重用姿态相关结构的存在,并确定稳定的局部条件化和精确的网格解码是改进传递的互补需求。
cs.CV / 35 / 2609.31722

Where Does the Watermark Hide? Push-Pull Disentanglement for Invisible Watermark Removal

水印隐藏在何处?用于不可见水印去除的推拉解耦
Yang, Jidong, Yu, Huaike, Li, Qi, Wang, Chunpeng, Miao, Yuantian, Gao, Suo, Chen, Xiao
Abstract
Fixed image distortions do not cover an attacker that learns from paired clean and watermarked images. We study this paired-training threat with single-image inference: deployment uses neither the clean reference nor the watermark key, payload, or decoder. An encoder maps each image to a structural latent $g$ and an auxiliary residual latent $u$. Push supervision reconstructs the watermarked image from $D(g_w,u_w)$. Pull supervision trains the zero-auxiliary output $D(A_g(g_w;k),0)$ toward the paired clean image. At $k=1.10,u=0$, the four-method sweep gives an average BER of $0.3958$, PSNR of $31.07$ dB, and SSIM of $0.9554$. Restoring $u$ from $0$ to $0.15$ moves average BER from $0.3893$ to $0.3357$, while PSNR falls from $30.99$ to $28.23$ dB. The intervention supports decoder dependence on the auxiliary input in the evaluated setting. The accompanying theory is a conditional, post-hoc account of this behavior rather than an experimentally verified information-relocation result.
Chinese Translation
固定的图像失真无法应对从成对的干净图像和水印图像中学习的攻击者。我们研究这种在单图像推理下的成对训练威胁:部署时不使用干净参考,也不使用水印密钥、载荷或解码器。编码器将每个图像映射到一个结构潜变量 $g$ 和一个辅助残差潜变量 $u$。推监督从 $D(g_w,u_w)$ 重建水印图像。拉监督训练零辅助输出 $D(A_g(g_w;k),0)$ 趋向于成对的干净图像。在 $k=1.10,u=0$ 时,四种方法的扫描给出平均 BER 为 $0.3958$,PSNR 为 $31.07$ dB,SSIM 为 $0.9554$。将 $u$ 从 $0$ 恢复到 $0.15$ 使平均 BER 从 $0.3893$ 变为 $0.3357$,同时 PSNR 从 $30.99$ 下降到 $28.23$ dB。该干预支持在评估设置中解码器对辅助输入的依赖。伴随的理论是对此行为的一种条件性、事后解释,而不是经过实验验证的信息重定位结果。
cs.CV / 36 / 2609.31723

Frequency-Domain AI-Generated Image Detection: Exploring Decoder and Channel Attention for Feature Refinement

频域AI生成图像检测:探索解码器与通道注意力用于特征细化
Roy, Uday Shankar, Minu, Mahbuba Jahan
Abstract
With the rapid progress of AI, the number of AI-generated images has increased significantly in recent years. However, the increasing variety of image generation models makes detection more difficult. In this work, we use Fast Fourier Transform (FFT) representation with EfficientNet-B0 for AI-generated image detection. EfficientNet-B0 provides a lightweight architecture that can be useful for resource-limited applications. Most frequency-domain detectors use a standard encoder to extract features from the FFT spectrum and directly pass them to a classifier. We explored a different approach by investigating ECA, U-Net, and Attention U-Net as alternatives to this direct encoder-to-classifier approach. ECA applies channel attention, while U-Net and Attention U-Net use decoder-based architectures to recover and refine spatial information in the extracted frequency features. We used a balanced subset of the MS COCOAI dataset that includes AI-generated images from five different models. Three runs were carried out for each experiment, and the average values were recorded. Experimental results indicate that EfficientNet-B0 obtained an accuracy of 84.64%, which is 4.50 percentage points higher than the ResNet-50 baseline reported in the dataset paper. EfficientNet-B0 with U-Net provided a small improvement, achieving an accuracy of 84.85%, while ECA did not increase the overall performance. EfficientNet-B0 with Attention U-Net achieved the best overall performance, with an accuracy of 85.51% and an ROC-AUC of 92.99%. This represents an improvement of 0.87 percentage points in accuracy compared to the EfficientNet-B0 baseline and 5.37 percentage points over the ResNet-50 baseline reported in the dataset paper.
Chinese Translation
随着AI的快速发展,近年来AI生成图像的数量显著增加。然而,图像生成模型的种类日益增多使得检测更加困难。在本工作中,我们使用快速傅里叶变换(FFT)表示与EfficientNet-B0进行AI生成图像检测。EfficientNet-B0提供了一种轻量级架构,可用于资源受限的应用。大多数频域检测器使用标准编码器从FFT频谱中提取特征,并直接将其传递给分类器。我们探索了一种不同的方法,通过研究ECA、U-Net和Attention U-Net作为这种直接编码器到分类器方法的替代方案。ECA应用通道注意力,而U-Net和Attention U-Net使用基于解码器的架构来恢复和细化提取的频率特征中的空间信息。我们使用了MS COCOAI数据集的平衡子集,其中包含来自五个不同模型的AI生成图像。每个实验进行了三次运行,并记录平均值。实验结果表明,EfficientNet-B0获得了84.64%的准确率,比数据集论文中报告的ResNet-50基线高出4.50个百分点。EfficientNet-B0与U-Net提供了小的改进,达到了84.85%的准确率,而ECA没有提高整体性能。EfficientNet-B0与Attention U-Net实现了最佳整体性能,准确率为85.51%,ROC-AUC为92.99%。这代表与EfficientNet-B0基线相比,准确率提高了0.87个百分点,与数据集论文中报告的ResNet-50基线相比提高了5.37个百分点。
cs.CV / 37 / 2609.31725

SWT: Self-Supervised Video Object Segmentation via Sliding, Wavelet and Transportation

SWT:基于滑动窗口、小波变换和最优传输的自监督视频目标分割
Zhu, Zhengtong, Fan, Jiaqing, Qian, Hanwen, Li, Fanzhang
Abstract
Video Object Segmentation (VOS) aims to accurately segment target objects from consecutive video frames and track the changes of the objects in each frame of the video. Conventional VOS methods typically demand substantial quantities of pixel-level labeled video sequences for fully supervised learning, which limits the performance of the model in sparse video scenes, while existing VOS methods have limited adaptability to global changes in objects. Based on this observation, in this paper, we propose self-supervised VOS with Sliding window, Wavelet transform and optimal Transport (SWT), a self-supervised VOS framework entirely trained on static dataset using contrastive learning. Firstly, a rolling sample buffer reuses overlapping groups of independently sampled images across successive updates. Secondly, to address the long-distance modeling difficulty caused by simple convolutional structures, we introduce wavelet transform to expand the receptive field of convolutional kernels, thus improving the model's representational capability. Finally, we incorporate optimal transport to assist the model in finding the globally optimal match between the target across two frames, improving the model's ability to handle nonrigid deformations of objects. SWT only requires training on the COCO dataset once and achieves excellent results on five VOS datasets as well as an additional body part propagation dataset. The code will be released soon at [https://github.com/machine928/SWT.git](https://github.com/machine928/SWT.git).
Chinese Translation
视频目标分割(VOS)旨在从连续视频帧中准确分割目标对象,并跟踪目标在视频每一帧中的变化。传统的VOS方法通常需要大量像素级标注的视频序列进行全监督学习,这限制了模型在稀疏视频场景中的性能,而现有的VOS方法对目标的全局变化适应性有限。基于这一观察,本文提出了基于滑动窗口、小波变换和最优传输的自监督VOS(SWT),这是一种完全在静态数据集上使用对比学习训练的自监督VOS框架。首先,滚动样本缓冲区在连续更新中重用独立采样图像的重叠组。其次,为了解决简单卷积结构导致的远程建模困难,我们引入小波变换来扩大卷积核的感受野,从而提高模型的表示能力。最后,我们引入最优传输来帮助模型找到目标在两帧之间的全局最优匹配,提高模型处理目标非刚性变形的能力。SWT只需在COCO数据集上训练一次,便在五个VOS数据集以及一个额外的身体部位传播数据集上取得了优异的结果。代码即将在 [https://github.com/machine928/SWT.git](https://github.com/machine928/SWT.git) 发布。
cs.CV / 38 / 2609.31726

High-Capacity Robust Medical Image Exfiltration via Neural Network Weight Replacement

基于神经网络权重替换的高容量鲁棒医学图像外泄
Thellier, Elie, Li, Huiyu, Ayache, Nicholas, Delingette, Hervé
Abstract
Collaborative medical AI platforms allow researchers to train models on sensitive imaging data while restricting data export. However, trained models can serve as covert carriers of patient information: medical images may be encoded within model parameters and reconstructed outside the secure environment. Existing defenses rely on lightweight sanitization (e.g., fine-tuning, pruning, quantization) and limited statistical auditing, creating a realistic insider exfiltration risk. We introduce a high-capacity neural steganography attack that encodes medical images as continuous latent representations embedded into model initialization. A StyleGAN2-based adversarial autoencoder learns compact latent codes regularized to match standard weight initialization statistics, keeping embedded parameters statistically consistent with clean models. Noise injection during training improves robustness to export-time mitigation. The carrier model remains functional on its intended task and hidden images can be reconstructed directly from its weights after export. This continuous encoding enables robust and scalable exfiltration, allowing up to 99 brain MRI volumes to be embedded within a 30MB model, and remains recoverable under mitigations that disrupt prior bit-level schemes. While reconstructions are approximate rather than pixel-exact, embedded content remains anatomically recognizable and recoverable at scale, exposing a privacy risk distinct from prior bit-level approaches. Experiments on MIMIC-CXR, BraTS, and LiTS demonstrate effectiveness across modalities, tasks, and architectures, highlighting the need for structural defenses beyond parameter-level sanitization. Code is available at https://github.com/ElieThellier/high-capacity-robust-medical-image-exfiltration.
Chinese Translation
协作式医学AI平台允许研究者在敏感影像数据上训练模型,同时限制数据导出。然而,训练后的模型可成为患者信息的隐蔽载体:医学图像可能被编码进模型参数,并在安全环境之外被重建。现有防御依赖轻量级净化(如微调、剪枝、量化)和有限的统计审计,造成了现实的内部人员外泄风险。我们提出一种高容量神经隐写攻击,将医学图像编码为连续潜在表示并嵌入模型初始化中。基于StyleGAN2的对抗自编码器学习紧凑潜在码,并通过正则化使其匹配标准权重初始化统计特性,从而保持嵌入参数与干净模型在统计上一致。训练期间注入噪声可提升对导出时缓解措施的鲁棒性。载体模型在其预期任务上仍保持功能,且隐藏图像可在导出后直接从其权重中重建。这种连续编码实现了鲁棒且可扩展的外泄,可在30MB模型中嵌入多达99个脑MRI体积,并且在会破坏先前位级方案的缓解措施下仍可恢复。虽然重建是近似的而非像素级精确,但嵌入内容在解剖学上仍可识别,并可在规模上恢复,暴露出不同于先前位级方法的隐私风险。在MIMIC-CXR、BraTS和LiTS上的实验证明了其跨模态、任务和架构的有效性,凸显了超越参数级净化的结构性防御需求。代码见https://github.com/ElieThellier/high-capacity-robust-medical-image-exfiltration。
cs.CV / 39 / 2609.31731

GERIS: A Game-Theoretic Framework for Filtering Instance-Dependent Label Noise in License Plate Data Augmentation

GERIS:一个用于过滤车牌数据增强中实例相关标签噪声的博弈论框架
Shani, Seyedeh Sara Jalili, Ahmadian, Rouhollah, Rahmani, Amin, Bideh, Mahdi, Ghatee, Mehdi
Abstract
In this paper, we propose GERIS, a game-theoretic framework for instance selection in the data augmentation phase of license plate recognition systems. During augmentation, synthetic license plate images are generated and transformed using stochastic noise to simulate real-world conditions. However, certain noise configurations lead to highly distorted, unreadable images that degrade model performance by introducing instance-dependent label noise. GERIS formulates a non-cooperative game in which each noise vector competes for inclusion in the training set based on its similarity to labeled data and its contribution to model reliability. By identifying and pruning low-quality instances, GERIS improves the overall quality of the augmented dataset. Unlike traditional black-box learning methods, GERIS offers a transparent, theoretically grounded mechanism for data filtering. Experimental results demonstrate that GERIS outperforms existing instance selection methods in terms of classification accuracy and robustness.
Chinese Translation
在本文中,我们提出了GERIS,一个用于车牌识别系统数据增强阶段中实例选择的博弈论框架。在增强过程中,使用随机噪声生成和变换合成车牌图像,以模拟真实世界条件。然而,某些噪声配置会导致高度失真、不可读的图像,通过引入实例相关的标签噪声而降低模型性能。GERIS构建了一个非合作博弈,其中每个噪声向量根据其与标记数据的相似性以及对模型可靠性的贡献,竞争进入训练集。通过识别和修剪低质量实例,GERIS提高了增强数据集的整体质量。与传统黑盒学习方法不同,GERIS提供了一种透明、有理论依据的数据过滤机制。实验结果表明,GERIS在分类准确性和鲁棒性方面优于现有的实例选择方法。
cs.CV / 40 / 2609.31733

When Retrieval Hurts: Measuring and Explaining Retrieval-Induced Hallucination in Chest X-ray Report Generation

当检索有害时:测量与解释胸部X光报告生成中的检索诱导幻觉
Idoko, Emmanuel, Olabisi, Abdusshakur, Oni, Shiloh, Josiah, Adesola
Abstract
Retrieval-augmented generation is an attractive way to improve chest X-ray reporting, because reports from similar prior studies supply clinical context a general vision-language model lacks. We show the same mechanism is a reliable source of clinical error. Over 100 MIMIC-CXR studies with a frozen LLaVA-1.5 generator and BioMedCLIP retrieval, CheXbert clinical F1 doubles under relevant retrieval (0.201 to 0.402) and collapses to 0.043 under clinically mismatched retrieval, a fifth of the image-only score; the retrieval-induced hallucination rate, counting only unsupported findings traceable to retrieved evidence, rises from 0.00 to 0.76 and 0.98. To show this is not an artefact of evidence-set coverage, we introduce a coincidental-overlap control that scores image-only generations against evidence they never saw, placing the chance base rate at 0.18, four to five times below the observed rates. The generator does not merely acquire findings, it transcribes text: 95% of reports produced under relevant retrieval contain an eight-word span occurring verbatim in the retrieved evidence but absent from the reference, against 0% without retrieval. We then explain the mechanism: a normal chest X-ray retrieves at least one abnormal precedent in 24 of 28 cases, because medical image-embedding similarity is dominated by anatomy and acquisition rather than by the presence of disease. This has a direct design consequence. Retrieval similarity does not predict harm (RIH rates 0.80/0.80/0.60/0.84 across similarity quartiles), so relevance gates conditioned on embedding similarity cannot work; gating on predicted pathology agreement removes 61% of unsupported evidence at no cost to useful coverage. We argue that retrieval-augmented clinical systems must be evaluated under retrieval failure, not only under retrieval success.
Chinese Translation
检索增强生成是一种改进胸部X光报告的有吸引力的方法,因为来自类似先前研究的报告提供了通用视觉语言模型所缺乏的临床背景。我们表明,同样的机制是临床错误的可靠来源。在超过100项MIMIC-CXR研究中,使用冻结的LLaVA-1.5生成器和BioMedCLIP检索,CheXbert临床F1在相关检索下翻倍(0.201至0.402),而在临床不匹配的检索下骤降至0.043,仅为仅图像得分的五分之一;检索诱导的幻觉率(仅计算可追溯到检索证据的无支持发现)从0.00上升至0.76和0.98。为了表明这不是证据集覆盖范围的假象,我们引入了一种巧合重叠控制,将仅图像生成与它们从未见过的证据进行评分,将随机基准率设定为0.18,比观察到的比率低四到五倍。生成器不仅获取发现,还转录文本:在相关检索下生成的报告中,95%包含一个在检索证据中逐字出现但在参考中缺失的八词片段,而在无检索的情况下为0%。然后我们解释了机制:正常的胸部X光在28例中有24例检索到至少一个异常先例,因为医学图像嵌入相似性主要由解剖结构和采集决定,而不是由疾病的存在决定。这有一个直接的设计后果。检索相似性不能预测危害(相似性四分位数的RIH率为0.80/0.80/0.60/0.84),因此基于嵌入相似性的相关性门控无法工作;基于预测病理一致性的门控在有用覆盖无成本的情况下移除了61%的无支持证据。我们认为,检索增强的临床系统必须在检索失败下进行评估,而不仅仅是在检索成功下。
cs.CV / 41 / 2609.31734

Measuring the evolution of camera distance across a century of film

测量电影百年间摄影机距离的演变
Bamman, David, Cooper, Allison, Hickey, Dan, Mar, Madison
Abstract
The rise of computer vision and artificial intelligence has made possible new forms of large-scale computational measurement. We apply these techniques to a deep collection of 5,205 digitized films viewed in theaters between 1922-2025 (covering popular, prestigious, and independent movies) to trace the development of one of the most fundamental ways through which film communicates: by manipulating the space between the camera and its subject. This work finds abrupt changes with the rise of new technologies in sound and television, and allows us to shed an empirical light on gender disparity (women, despite having substantially less screentime than men, are disproportionately the subject of closer shots), and illustrate how animated films both inherit and break free from the norms of live-action filmmaking.
Chinese Translation
计算机视觉与人工智能的兴起,使新形式的大规模计算测量成为可能。我们将这些技术应用于一个深度影片集合:5,205部在1922—2025年间于影院放映的数字化影片(涵盖流行、声誉卓著和独立电影),以追踪电影进行交流的最基本方式之一的发展:通过操控摄影机与其拍摄对象之间的空间。本研究发现,随着声音和电视新技术的兴起,出现了急剧变化;并使我们能够从实证角度揭示性别差异(女性尽管银幕时间显著少于男性,却不成比例地成为更近景别镜头的拍摄对象),并说明动画电影如何既继承又摆脱真人电影制作的规范。
cs.CV / 42 / 2609.31736

LukeNet: A lightweight CNN integrated with an XAI model for Smart acute lymphoblastic leukemia detection and management

LukeNet:一种集成XAI模型的轻量级CNN,用于智能急性淋巴细胞白血病检测与管理
Ahad, Md Taimur
Abstract
Acute Lymphoblastic Leukemia (ALL) patients require early, accurate detection to enable timely treatment and effective patient management. A Convolutional Neural Network (CNN) is well-suited for creating an end-to-end enabling environment for ALL detection and classification. However, most CNN-based ALL detection systems are theoretical and unsuitable for deployment on edge devices due to high computational demands. The Internet of Medical Things (IoMT)-enabled devices offer an opportunity to monitor ALL patients in real time. Wearables that track temperature, heart rate, oxygen saturation, and activity can deliver critical data to support timely clinical intervention and improve patient outcomes. In smart IoMT environments, a lightweight CNN is essential because connected devices often operate under limited computational power, memory, and latency constraints. To address this need, this study proposes LukeNet, a lightweight CNN integrated with explainable artificial intelligence (XAI) for an IoMT-based SMART Acute Lymphoblastic Leukemia Detection and Management System. Trained on three (3) ALL datasets and five-fold cross-validation, LukeNet achieved an impressive 99% model accuracy as well as 99% unseen test accuracy, which is higher than six state-of-the-art (SOTA) CNNs, such as DenseNet121, MobileNet, ResNet50, InceptionV3, Xception, and VGG16, as well as transfer learning models. Furthermore, LukeNet was compared with two ensemble models. In addition, explainable AI methods are integrated to highlight relevant regions in microscopic images. The novelty of this study lies in the architecture of LukeNet, which balances model depth and computational efficiency by using depthwise separable convolutions, mitigates the risk of gradient loss in deeper layers, and provides strong global and local feature extraction capabilities.
Chinese Translation
急性淋巴细胞白血病(ALL)患者需要早期、准确的检测,以实现及时治疗和有效的患者管理。卷积神经网络(CNN)非常适合为ALL检测和分类创建端到端的赋能环境。然而,大多数基于CNN的ALL检测系统是理论性的,由于高计算需求,不适合部署在边缘设备上。医疗物联网(IoMT)设备提供了实时监测ALL患者的机会。跟踪温度、心率、血氧饱和度和活动的可穿戴设备可以提供关键数据,支持及时的临床干预并改善患者预后。在智能IoMT环境中,轻量级CNN至关重要,因为连接设备通常在有限的计算能力、内存和延迟约束下运行。为了满足这一需求,本研究提出了LukeNet,一种集成了可解释人工智能(XAI)的轻量级CNN,用于基于IoMT的智能急性淋巴细胞白血病检测和管理系统。在三个ALL数据集上训练并使用五折交叉验证,LukeNet实现了令人印象深刻的99%模型准确率和99%未见测试准确率,高于六个最先进的(SOTA)CNN,如DenseNet121、MobileNet、ResNet50、InceptionV3、Xception和VGG16,以及迁移学习模型。此外,LukeNet还与两个集成模型进行了比较。此外,集成了可解释AI方法以突出显微图像中的相关区域。本研究的新颖性在于LukeNet的架构,它通过使用深度可分离卷积来平衡模型深度和计算效率,减轻了深层梯度损失的风险,并提供了强大的全局和局部特征提取能力。
cs.CV / 43 / 2609.31737

Seeing the Heat: Synthesizing High-Resolution Wood Thermal Responses from Optical Imagery

看见热量:从光学图像合成高分辨率木材热响应
Xie, Jingren
Abstract
The thermal behavior of wood is a critical factor in advanced material assembly. However, pixel-level thermal analysis remains fundamentally constrained by the low resolution and noise inherent to infrared thermography. To address this, we introduce an end-to-end computational framework that synthesizes high-resolution thermal responses directly from wood RGB images. We first establish a core physical linkage: because spatial color variation in natural wood is driven by cellular anatomy, optical intensity serves as a reliable geometric proxy for the localized solid volume fraction. By leveraging this theoretical insight, we develop an automated finite-element-method data engine that maps pixel-level optical intensity to a 3D thermodynamic voxel grid, generating high-fidelity synthetic thermal responses. We find that 1) when the thermal conductivity along the thickness direction is uniform or linear, wood RGB images and their corresponding thermal responses exhibit extreme morphological similarities, and the lateral thermal diffusion acts as a low-pass filter that smooths out high-frequency details; 2) when the thermal conductivity along the thickness direction is random, such morphological similarities are destroyed, and wood's 3D structure dominantly governs its thermal response. We further utilize these synthetic thermal responses to supervise a neural surrogate model built upon the DINOv3 foundation model. Our results demonstrate that the neural surrogate model successfully internalizes the governing thermodynamic laws, thereby bypassing computationally expensive simulations and enabling high-resolution thermal inference. This methodology effectively bridges the semantic and thermodynamic domains, unlocking systematic, pixel-level analysis of fine-grained wood thermal responses. Project: https://zekifayes.github.io/seeheat
Chinese Translation
木材的热行为是先进材料组装中的关键因素。然而,像素级热分析从根本上受到红外热成像固有的低分辨率和噪声的限制。为了解决这个问题,我们引入了一个端到端的计算框架,可以直接从木材RGB图像合成高分辨率热响应。我们首先建立了一个核心物理联系:由于天然木材的空间颜色变化是由细胞解剖结构驱动的,光学强度可作为局部固体体积分数的可靠几何代理。利用这一理论见解,我们开发了一个自动化的有限元法数据引擎,将像素级光学强度映射到3D热力学体素网格,生成高保真合成热响应。我们发现:1)当沿厚度方向的热导率均匀或线性时,木材RGB图像及其相应的热响应表现出极端的形态相似性,并且横向热扩散充当低通滤波器,平滑掉高频细节;2)当沿厚度方向的热导率随机时,这种形态相似性被破坏,木材的3D结构主导其热响应。我们进一步利用这些合成热响应来监督基于DINOv3基础模型构建的神经代理模型。我们的结果表明,该神经代理模型成功内化了主导的热力学定律,从而绕过了计算成本高昂的模拟,实现了高分辨率热推断。该方法有效桥接了语义域和热力学域,解锁了对细粒度木材热响应的系统化像素级分析。项目:https://zekifayes.github.io/seeheat
cs.CV / 44 / 2609.31740

Beyond Volume Overlap: Surface Matching for Topology-Aware Coronary Artery Segmentation

超越体积重叠:面向拓扑感知冠状动脉分割的表面匹配
Velasquez, Rafael, Puyol-Antón, Esther, Arbeláez, Pablo
Abstract
Accurate coronary artery segmentation on coronary computed tomography angiography (CCTA) is essential for diagnosing coronary artery disease. Deep networks are conventionally trained and evaluated with the Dice coefficient, but volume-overlap metrics are poorly suited to thin, tubular anatomy: since most voxels belong to a few thickproximal segments, a missing distal branch barely affects Dice despite severely disrupting the connectivity required for clinical use. We introduce a surface metric that matches predicted and reference surface points via bipartite assignment under a localized, vessel-radius tolerance, reporting precision, recall, and F1 with decoupled false positives (spurious branches) and false negatives (missed branches) a distinction the symmetric Dice cannot make. With it we show that a strong Dice-trained baseline omits far more vessel surface than it hallucinates, an asymmetry its high Dice hides. Building on this, we propose a differentiable surface loss that simultaneously suppresses spurious mass and recovers absent structure, validated by fine-tuning three backbones (nnU-Net, SwinUNETR, NexToU) on two public benchmarks (Image-CAS, ASOCA). Against a matched-epoch control, it significantly improves surface F1 by recovering missed distal vessels at comparable Dice. Our findings argue for measuring and optimizing the vessel surface, not the volume it overlaps. Code
Chinese Translation
在冠状动脉计算机断层扫描血管造影(CCTA)上准确分割冠状动脉对于诊断冠状动脉疾病至关重要。深度网络通常使用Dice系数进行训练和评估,但体积重叠指标并不适合纤细的管状解剖结构:由于大多数体素属于少数粗大的近端节段,缺失一条远端分支对Dice几乎没有影响,尽管它严重破坏了临床使用所需的连通性。我们引入一种表面指标,在局部化的血管半径容差下通过二分图分配匹配预测表面点与参考表面点,报告精确率、召回率和F1,并将假阳性(虚假分支)与假阴性(漏检分支)解耦,这是对称Dice无法做出的区分。借助该指标,我们表明一个强大的Dice训练基线遗漏的血管表面远多于其虚构的表面,而这种不对称性被其高Dice掩盖。在此基础上,我们提出一种可微表面损失,可同时抑制虚假结构并恢复缺失结构,并通过在两个公开基准(Image-CAS、ASOCA)上微调三个骨干网络(nnU-Net、SwinUNETR、NexToU)进行验证。与匹配epoch的对照相比,它在Dice相当的情况下通过恢复漏检的远端血管显著提高了表面F1。我们的发现主张测量并优化血管表面,而非其重叠的体积。代码
cs.CV / 45 / 2609.31742

PEEL-DDPM: Physics-Enabled Evidential Learning for the Denoising Diffusion Probabilistic Model

PEEL-DDPM: 面向去噪扩散概率模型的物理使能证据学习
Wang, Ge
Abstract
Normal-inverse-gamma (NIG) regression is not identifiable from its marginal Student-t likelihood: three combinations of four NIG parameters are determined, leaving one degree of freedom. We introduce PEEL-DDPM, a physics-enabled evidential learning framework for denoising diffusion probabilistic models. A measurement-conditioned DDPM is first trained with epsilon-MSE and then frozen. Its complete reverse trajectory produces a reconstruction, whose residual from the known training object is modeled by a final-image evidential network. The network learns only the identifiable Student-t coordinates, while repeated scanner-noise realizations and repeated diffusion trajectories provide a nested Monte Carlo estimate of final-image aleatoric variance, separated into scanner-induced and sampler-induced components. This measured variance resolves the remaining NIG ambiguity and yields a decomposition of predictive uncertainty into measurement, diffusion, and epistemic terms. The method uses sequential training without a cross-loss weighting coefficient. In a feasibility study on eight held-out objects, empirical central-interval coverages were 49.3%, 80.3%, 90.1%, and 95.6% for nominal 50%, 80%, 90%, and 95% intervals. The mean squared residual was 0.967 times the mean predicted variance, and a single-image aleatoric head achieved pooled Spearman correlation 0.785 against an independent nested reference. Across five dose levels, scanner-induced variance showed a log-log dose slope of -1.15, whereas sampler-induced variance remained nearly dose independent with slope -0.01. These results support PEEL-DDPM as a practical route to identifiable and physically interpretable uncertainty quantification in diffusion-based image reconstruction.
Chinese Translation
正态-逆伽马 (NIG) 回归无法从其边际 Student-t 似然中识别:四个 NIG 参数中的三个组合被确定,留下一个自由度。我们引入了 PEEL-DDPM,一种用于去噪扩散概率模型的物理使能的证据学习框架。首先使用 epsilon-MSE 训练一个以测量为条件的 DDPM,然后将其冻结。其完整的反向轨迹产生一个重建,其与已知训练对象的残差由最终图像证据网络建模。该网络仅学习可识别的 Student-t 坐标,而重复的扫描仪噪声实现和重复的扩散轨迹提供了最终图像偶然方差的嵌套蒙特卡洛估计,该方差被分为扫描仪引起的和采样器引起的分量。这种测量方差解决了剩余的 NIG 模糊性,并将预测不确定性分解为测量、扩散和认知项。该方法使用顺序训练,无需交叉损失权重系数。在一项对八个留出对象的可行性研究中,对于标称 50%、80%、90% 和 95% 区间,经验中心区间覆盖率分别为 49.3%、80.3%、90.1% 和 95.6%。均方残差是平均预测方差的 0.967 倍,并且单图像偶然性输出头在独立嵌套参考下达到了汇总的 Spearman 相关系数 0.785。在五个剂量水平上,扫描仪引起的方差显示出双对数剂量斜率为 -1.15,而采样器引起的方差几乎与剂量无关,斜率为 -0.01。这些结果支持 PEEL-DDPM 作为在基于扩散的图像重建中实现可识别且物理可解释的不确定性量化的实用途径。
cs.CV / 46 / 2609.31746

VisionPsy-Nano: Improving Accuracy, Efficiency, and Reliability in On-Device Vision-Language Models

VisionPsy-Nano:提升设备端视觉语言模型的准确性、效率和可靠性
Hashmi, Khurram Azeem, Zolfaghari, Mohammadreza, Park, Changdae, Jain, Rishabh, Moratelli, Nicholas, Wei, Pengfei, Lu, Louis, Liu, Tianchi, Nazir, Amril
Abstract
Sub-billion-parameter Vision-Language Models are increasingly viable for on-device deployment, yet compact model size alone does not guarantee usability. On a phone, such a model can still require more than two minutes to produce its first token. On-device usability depends on three axes: accuracy, efficiency, and behavioral reliability; standard benchmarks miss the third, with answers too short to expose doom loops and prompts too benign to probe adversarial safety. We introduce a diagnosis-driven post-training recipe in which a teacher VLM stress-tests the student, uncovers failure modes beyond human priors, and converts them into targeted supervision and preference alignment, supplementing generic data scaling with failure-driven optimization. Coupled with two visual-token policies, the recipe yields two accuracy-efficiency variants with improved behavioral reliability. \textbf{\NanoFull} attains a 62.3 normalized average over 17 benchmarks, the highest among openly released $\sim$0.5B models, +7.4 over its base at identical architecture and token budget, with doom-loop rates at or below the strongest baseline's. \textbf{\FlashFull} retains 61.4 while cutting warm time-to-first-token on a Pixel 9 from 138\,s to 6.1\,s (23$\times$). By jointly addressing all three axes, we move compact VLMs toward practical on-device usability.
Chinese Translation
参数低于十亿的视觉语言模型日益适用于设备端部署,但仅凭紧凑的模型尺寸并不能保证可用性。在手机上,此类模型生成第一个词元仍可能需要超过两分钟。设备端可用性取决于三个维度:准确性、效率和行为的可靠性;标准基准测试忽略了第三个维度,其答案太短无法暴露死循环,提示过于良性无法探测对抗性安全。我们提出一种诊断驱动的后训练方案,其中教师VLM对学生进行压力测试,发现超出人类先验的失败模式,并将其转化为有针对性的监督和偏好对齐,以失败驱动的优化补充通用数据扩展。结合两种视觉词元策略,该方案产生两个准确率-效率变体,并提高了行为可靠性。\textbf{\NanoFull} 在17个基准测试上达到62.3的归一化平均值,在公开释放的约0.5B模型中最高,在相同架构和词元预算下比其基础模型高7.4,死循环率等于或低于最强基线。\textbf{\FlashFull} 保持61.4,同时将Pixel 9上的预热首词元时间从138秒减少到6.1秒(23倍)。通过联合解决所有三个维度,我们推动紧凑型VLM走向实用的设备端可用性。
cs.CV / 47 / 2609.31747

The Earth in One Gaze: Training-Free Active Focus for UHR Remote Sensing Understanding

一眼地球:面向超高分辨率(UHR)遥感理解的免训练主动聚焦
Zhang, Yao, Dai, Pengyu, Guo, Wei, Liang, Jian, Song, Jian, Ou, Yafei, Chen, Hongruixuan, Yokoya, Naoto
Abstract
Multimodal large language models (MLLMs) must balance local detail against scene context when interpreting ultra-high-resolution (UHR) remote sensing (RS) imagery within a limited visual-input budget. Existing selection-based methods either prune tokens and select patches through relevance scoring, or crop actively through repeated inspection. Neither strategy directly redistributes pixels within a continuous full-scene view: the first retains selected tokens or patches, and the second re-encodes a crop detached from its surroundings. Our pilot study finds that a frozen MLLM already produces useful question-guided spatial requests, yet crop-based inspection of the selected regions does not consistently improve its answers. We therefore formulate UHR understanding as a question of where to spend a fixed pixel budget. Based on this, we introduce GazeEarth, a simple-yet-effective training-free framework that couples question-guided region selection with full-scene foveated observation. The MLLM selects evidence cells from an indexed overview; a deterministic, topology-preserving warp resamples the original image onto a fixed-size canvas, enlarging their shared neighborhood while compressing the periphery; the same frozen model answers from this focused view, using at most two MLLM calls and no external selector or iterative search. Across three UHR remote sensing benchmarks and four frozen backbones, GazeEarth improves benchmark-averaged accuracy by 4.6 to 9.4 percentage points over direct answering and 3.4 to 4.3 over overview answering, outperforming task-trained methods. Our analyses show that existing MLLMs can guide where to look in UHR images on their own, and that what they can infer from the selected evidence depends on how that evidence is presented.
Chinese Translation
多模态大语言模型(MLLMs)在有限的视觉输入预算内解读超高分辨率(UHR)遥感(RS)影像时,必须平衡局部细节与场景上下文。现有的基于选择的方法要么通过相关性评分修剪token并选择patch,要么通过重复检查主动裁剪。这两种策略都没有在连续的全景视图内直接重新分配像素:前者保留选定的token或patch,后者重新编码一个脱离其周围环境的裁剪区域。我们的初步研究发现,冻结的MLLM已经能够产生有用的问题引导空间请求,然而对所选区域进行基于裁剪的检查并不能持续改进其答案。因此,我们将UHR理解形式化为一个关于在何处花费固定像素预算的问题。基于此,我们提出GazeEarth,一个简单而有效的免训练框架,将问题引导的区域选择与全景中央凹观察相结合。MLLM从索引概览图中选择证据单元;一个确定性的、保持拓扑的变形将原始图像重采样到固定大小的画布上,放大它们共享的邻域,同时压缩外围区域;同一个冻结模型从这个聚焦视图进行回答,最多使用两次MLLM调用,且无需外部选择器或迭代搜索。在三个UHR遥感基准和四个冻结主干模型上,GazeEarth相较于直接回答将基准平均准确率提升了4.6至9.4个百分点,相较于概览回答提升了3.4至4.3个百分点,优于任务训练方法。我们的分析表明,现有的MLLM能够自行指导在UHR图像中看向何处,并且它们能从所选证据中推断出什么取决于该证据如何呈现。
cs.CV / 48 / 2609.31749

Fusion Under Component Failure: Negative Results and Failure Modes in Ensemble AI-Generated Image Detection

组件失效下的融合:集成AI生成图像检测中的负面结果与失效模式
Singh, Suraj, Verma, Tushar, Singh, Pragyan, Bhav, Shaurya, Kumar, Shivam
Abstract
We built an ordinary stacking ensemble for AI-generated image detection -- three open detectors producing five scores, fused by a gradient-boosted meta-learner that treats a detector's failure as missing data -- deployed it, and then evaluated it against three controls it should have faced first. This paper reports what the controls found, including where they overturned our own earlier conclusions. Fusion is worth its cost only when refitted on the target domain. The shipped meta-learner, fitted on a separate corpus, does not beat its best single member on 2000 StyleGAN faces (AUC 0.9896 vs 0.9961; McNemar p = 1.000). But a stacker refitted in-domain beats that member plus a post-hoc calibrator (Delta-AUC = +0.0025, [+0.0014, +0.0038]; p = 3.4e-4). An earlier draft claimed the calibrated single detector won outright; that comparison mixed regimes and we correct it here. One corpus is not an evaluation. On 80 screenshots every model's AUC interval contains 0.5. We can say nothing stronger: the difference between the ensemble's drop and its best member's is [-0.185, +0.115]. Abstention is real, correlated, and mishandled. With four of five detectors silent and the survivor reporting "real", the system returns P(AI) = 0.9985, because an all-NaN input scores 0.9995 in a learner never fitted with missingness. The three AIDE checkpoints fail together, sharing one preprocessing path. A quorum rule requiring two distinct architectures prevents both failures with no retraining. We also find Corpus A carries a class-conditional JPEG bias severe enough to separate the classes from the header alone, which limits every in-domain number we report. Code, harness, hash-identified artifacts and all corrections are released.
Chinese Translation
我们构建了一个用于AI生成图像检测的普通堆叠集成——三个开放检测器产生五个分数,由一个梯度提升元学习器融合,该学习器将检测器的失败视为缺失数据——部署了它,然后根据它本应首先面对的三种对照进行了评估。本文报告了这些对照的发现,包括它们推翻我们先前结论的地方。只有在目标域上重新拟合时,融合才值得其成本。部署的元学习器在单独的语料库上拟合,在2000张StyleGAN人脸图像上并不优于其最佳单一成员(AUC 0.9896 vs 0.9961;McNemar p = 1.000)。但在域内重新拟合的堆叠器优于该成员加上事后校准器(Delta-AUC = +0.0025,[+0.0014, +0.0038];p = 3.4e-4)。早期的草稿声称校准后的单一检测器完全胜出;该比较混合了不同机制,我们在此更正。一个语料库不是评估。在80张截图上,每个模型的AUC区间都包含0.5。我们不能说得更绝对:集成模型的下降与其最佳成员之间的差异为[-0.185, +0.115]。弃权是真实的、相关的,并且被错误处理。当五个检测器中有四个沉默,而幸存者报告“真实”时,系统返回P(AI) = 0.9985,因为在一个从未拟合缺失数据的学习器中,全NaN输入得分为0.9995。三个AIDE检查点一起失败,共享一条预处理路径。一种要求两种不同架构的法定人数规则无需重新训练即可防止这两种失败。我们还发现Corpus A带有类别条件JPEG偏差,严重到仅从文件头就能区分类别,这限制了我们报告的每一个域内数字。代码、测试框架、哈希标识的工件以及所有更正均已发布。
cs.CV / 49 / 2609.31750

Dimension-Specific Imbalance and an Adaptive Hybrid Label Strategy for Multi-Task Affective State Recognition in Classroom Video

维度特定不平衡与自适应混合标签策略用于课堂视频中的多任务情感状态识别
Li, Xiangqian
Abstract
Recognition of student affective states from classroom video is constrained by a class imbalance problem whose true nature is, we argue, under-analyzed. We show that on DAiSEE -- the de facto benchmark for this task -- class imbalance is fundamentally a dimension-specific phenomenon: raw imbalance ratios reach as high as approximately 79:1 (Frustration), and the structure of imbalance -- not only its magnitude -- varies across dimensions, so uniform algorithmic treatments (aggressive class reweighting, focal loss, uniform binary simplification) fail in dimension-dependent ways. We propose the Adaptive Hybrid Label Strategy (AHLS), which assigns four-level classification to dimensions with manageable imbalance and binary classification to severely long-tailed dimensions, coupled with a null-class placeholder mechanism that stabilizes multi-task optimization by compressing the maximum effective inverse-frequency weight ratio from as high as ~79$\times$ to at most ~8.5$\times$ across the four dimensions. Validated on a lightweight FERShuffleNetV2 + LTCN architecture (0.39 M parameters, 0.07 G FLOPs), the strategy attains the highest mean Macro-F1 of 47.70% across all four affective dimensions among 11 compared methods, including five state-of-the-art deep models (up to 167$\times$ larger) and five traditional machine-learning baselines. A cross-model behavioral analysis on DAiSEE indicates that, in this setting, classification strategy may influence minority-state recognition more strongly than further backbone scaling alone. We also report evidence of an annotation-density limitation of DAiSEE on minority affective states, which suggests that benchmark design may bound future progress as much as algorithmic refinement. Index Terms Affective computing, classroom video analysis, class imbalance, multi-task learning, lightweight deep learning, DAiSEE, engagement recognition.
Chinese Translation
从课堂视频中识别学生情感状态受到类别不平衡问题的制约,我们认为,这一问题的本质尚未得到充分分析。我们表明,在 DAiSEE——该任务的事实基准——上,类别不平衡从根本上是一个维度特定现象:原始不平衡比高达约 79:1(挫败感),并且不平衡的结构——不仅是其幅度——随维度变化,因此统一的算法处理(激进的类别重加权、焦点损失、统一二元简化)以维度相关的方式失效。我们提出了自适应混合标签策略(AHLS),该策略对不平衡程度可控的维度分配四级分类,对严重长尾维度分配二元分类,并辅以空类占位机制,通过将四个维度上的最大有效逆频率权重比从高达约 79$\times$ 压缩至至多约 8.5$\times$,从而稳定多任务优化。在轻量级 FERShuffleNetV2 + LTCN 架构(0.39M 参数,0.07 G FLOPs)上验证,该策略在 11 种比较方法中,在所有四个情感维度上取得了最高的平均 Macro-F1 为 47.70%,其中包括五个最先进的深度模型(参数量最多达 167$\times$)和五个传统机器学习基线。在 DAiSEE 上的跨模型行为分析表明,在此设置下,分类策略可能比进一步骨干网络缩放本身更强烈地影响少数状态识别。我们还报告了 DAiSEE 在少数情感状态上的标注密度限制的证据,这表明基准设计可能像算法改进一样限制未来的进展。索引词:情感计算,课堂视频分析,类别不平衡,多任务学习,轻量级深度学习,DAiSEE,参与度识别。
cs.CV / 50 / 2609.31751

EgoTSR++: Egocentric Spatiotemporal Reasoning for Task Progress Understanding

EgoTSR++:面向任务进度理解的第一人称时空推理
Yang, Xiaoda, Wang, Can, Liu, Yuxiang, Zhou, Pengfei, Lou, Jianwen, Yan, Shuicheng
Abstract
Vision-Language Models (VLMs) have advanced rapidly in static visual understanding, yet remain unreliable when judging how an egocentric task is progressing. Given a task instruction and two visual observations, a model should determine which state is closer to the goal by analyzing task-relevant object configurations and spatial relations, rather than relying on timestamps or presentation order. This distinction is critical in manipulation, where retries, corrective actions, and temporary regressions make progress inherently non-monotonic. We introduce EgoTSR, a unified framework for diagnosing and improving order-robust task-progress understanding. First, SpatialLogic-Bench evaluates each physical state pair in both original and order-swapped presentations across short- and long-horizon settings, exposing whether a model follows task-state evidence or chronological shortcuts. Second, our data construction pipeline converts successful, approximately monotonic manipulation and first-person trajectories into bidirectional supervision; LongTag further preserves intermediate subtask structure for long-horizon comparison, while failure-aware data extend learning to regressions and recoveries. Third, a progressive CoT-to-Tag curriculum first supervises evidence-grounded interpretation of task-relevant state changes and then consolidates the comparison rule through scalable label-only training. Experiments reveal substantial input-order bias in representative VLMs. EgoTSR achieves 92.4% long-horizon accuracy with a 0.1-point forward-inverse Gap. Failure-aware supervision further improves accuracy on non-monotonic trajectories by 11.8 points and Recovery Accuracy by 11.2 points, while maintaining broad visual and spatial capabilities. These results establish goal-conditioned state comparison as an explicit formulation of egocentric spatiotemporal reasoning for task-progress understanding.
Chinese Translation
视觉语言模型(VLMs)在静态视觉理解方面取得了快速进展,但在判断第一人称任务如何进展时仍不可靠。给定任务指令和两个视觉观测,模型应通过分析任务相关的物体配置和空间关系来判断哪个状态更接近目标,而不是依赖时间戳或呈现顺序。这种区分在操作中至关重要,因为重试、纠正动作和临时回退使得进展本质上非单调。我们提出了EgoTSR,一个用于诊断和改进顺序鲁棒的任务进度理解的统一框架。首先,SpatialLogic-Bench在短时程和长时程设置下,对每个物理状态对在原始和顺序交换的呈现中进行评估,揭示模型是遵循任务状态证据还是时间顺序捷径。其次,我们的数据构建流程将成功的、近似单调的操作和第一人称轨迹转换为双向监督;LongTag进一步保留中间子任务结构以进行长时程比较,而失败感知数据将学习扩展到回退和恢复。第三,渐进式CoT到Tag课程首先监督对任务相关状态变化的基于证据的解释,然后通过可扩展的仅标签训练巩固比较规则。实验揭示了代表性VLMs中存在显著的输入顺序偏差。EgoTSR实现了92.4%的长时程准确率,正向-逆向差距为0.1点。失败感知监督进一步将非单调轨迹上的准确率提高了11.8点,恢复准确率提高了11.2点,同时保持了广泛的视觉和空间能力。这些结果确立了目标条件状态比较作为面向任务进度理解的第一人称时空推理的显式表述。
cs.CV / 51 / 2609.31754

Calibration-Free Surface Normals Estimation in Vision-Based Tactile Sensing using Universal Photometric Stereo

基于 Universal Photometric Stereo 的视觉触觉传感中的无标定表面法线估计
Dugonjic, Zdravko, Speidel, Stefanie, Calandra, Roberto
Abstract
Vision-based tactile sensors are a popular solution for capturing rich contact surface geometry. However, to obtain high-detail contact surface normals and depth, it is necessary to calibrate the sensor by physically pressing a probe with known geometry against the sensor elastomer and mapping tactile images onto the ground truth probe's shape. This approach does not scale across different tactile sensors, and the calibration effort can be complex depending on the sensor shape and optical system. Instead, we propose a calibration-free procedure for the estimation of contact surface normals using Universal Photometric Stereo neural networks. In a series of real-world experiments, we evaluate our approach on 3 sensors with different optical systems, demonstrating that universal methods are a suitable approach for estimating surface normals at the contact patch from tactile images, thereby alleviating the need for tactile sensor calibration. Controlled experiments with a metal ball show that universal methods match the calibrated method, with a mean angular error of $6.56^{\Large\circ}$. We show that the proposed framework recovers high-frequency surface details of objects with natural textures, achieving an overall mean angular error of $10.66^{\Large\circ}$. Universal method robustly recovers the contact surface normals captured with dome-shaped Digit 360, achieving a low angular discrepancy of $10.18^{\Large\circ}$ relative to the calibrated baseline. This experiment demonstrates that with sufficient illumination settings surface normals could be estimated using a model trained solely on synthetic data. By providing a unified representation of contact surfaces across different vision-based tactile sensor designs, Universal Photometric Stereo neural networks lay the foundation for transferable tactile perception across sensors.
Chinese Translation
基于视觉的触觉传感器是捕获丰富接触表面几何形状的常用方案。然而,为了获得高细节的接触表面法线和深度,需要通过将已知几何形状的探针物理按压到传感器弹性体上,并将触觉图像映射到真实探针形状上来标定传感器。这种方法无法扩展到不同的触觉传感器,并且标定工作可能很复杂,取决于传感器形状和光学系统。相反,我们提出了一种无需标定的流程,使用通用光度立体(Universal Photometric Stereo)神经网络来估计接触表面法线。在一系列真实世界实验中,我们在具有不同光学系统的 3 个传感器上评估了我们的方法,证明通用方法是根据触觉图像估计接触区域表面法线的合适方法,从而减轻了触觉传感器标定的需求。使用金属球的控制实验表明,通用方法与标定方法相当,平均角度误差为 6.56°。我们表明,所提出的框架能够恢复具有自然纹理物体的高频表面细节,总体平均角度误差为 10.66°。通用方法稳健地恢复了用圆顶形 Digit 360 捕获的接触表面法线,相对于标定基线实现了 10.18° 的低角度差异。该实验表明,在足够的照明设置下,可以使用仅在合成数据上训练的模型来估计表面法线。通过提供跨不同基于视觉的触觉传感器设计的接触表面的统一表示,通用光度立体神经网络为跨传感器的可迁移触觉感知奠定了基础。
cs.CV / 52 / 2609.31755

Gauge-Equivariant Attention for Rotation-Stable $360^\circ$ Scene Understanding

用于旋转稳定 $360^\circ$ 场景理解的规范等变注意力
Zhou, Tianjian, Li, Yishan, Jiang, Jie, Zhang, Yifei
Abstract
Panoramic $360^\circ$ scene understanding increasingly relies on icosphere transformers, but a state-of-the-art spherical model loses more than half of its segmentation accuracy when the camera rotates by $90^\circ$, and controlled ablations identify gauge dependence in its relative-position bias as a major contributor. We propose gauge-equivariant relative position encoding (GE-RPE): a parameter-free Reynolds average of the bias over a finite cyclic subgroup $C_n\!\subset\!\mathrm{SO}(2)$ of gauge rotations. Plugged into a SphereUFormer backbone the change is invisible at deployment---zero added parameters and $1.5$--$4.4\%$ forward latency---and the matched three-seed GE-RPE model records a $1.3\%$ drop; the published SphereUFormer checkpoint records $53\%$ under the same stress protocol but a different training recipe. Once the gauge defect is removed and a teacher-token permutation $\pi_R$ aligns the SSL views to the rotated student frame, iBOT$+$MAE pretraining stops being a liability and becomes a clean low-label lever: the full framework EquiSSL (GE-RPE $+$ $\pi_R$ $+$ iBOT$+$MAE) tightens the drop to $0.8\%$ at $68.30\%$ val mIoU and lifts $1\%$-label fine-tuning by $+2.39$ mIoU on the $N{=}373$ test split (and by $+4.10$ on the smaller $N{=}40$ val split); the same fix carries over to monocular depth and to zero-shot Structured3D segmentation. The construction is provably $C_n$-invariant and $\mathcal{O}(n^{-2})$-close to the continuous $\mathrm{SO}(2)$ average, making the resulting model a usable $360^\circ$ visual-computing primitive across panoramic relighting, immersive video, and cross-dataset transfer. Code is available at https://github.com/Jaywalk18/equissl-release.
Chinese Translation
全景 $360^\circ$ 场景理解日益依赖二十面体球面变换器(icosphere transformers),但当前最先进的球面模型在相机旋转 $90^\circ$ 时会损失超过一半的分割精度,而受控消融实验表明其相对位置偏置中的规范依赖性是主要因素。我们提出规范等变相对位置编码(GE-RPE):一种在规范旋转的有限循环子群 $C_n\!\subset\!\mathrm{SO}(2)$ 上对偏置进行无参数 Reynolds 平均的方法。将其插入 SphereUFormer 主干网络后,该改动在部署时几乎不可见——零新增参数,前向延迟增加 $1.5$--$4.4\%$——并且匹配的三种子 GE-RPE 模型仅录得 $1.3\%$ 的性能下降;而已发布的 SphereUFormer 检查点在相同压力测试协议下(但采用不同的训练配方)录得 $53\%$ 的下降。一旦消除了规范缺陷,并且教师 token 置换 $\pi_R$ 将 SSL 视图对齐到旋转后的学生坐标系,iBOT$+$MAE 预训练便不再是一种负担,而成为一个干净的低标签学习杠杆:完整框架 EquiSSL(GE-RPE $+$ $\pi_R$ $+$ iBOT$+$MAE)将下降幅度收紧至 $0.8\%$,验证集 mIoU 达到 $68.30\%$,并在 $N{=}373$ 测试集划分上将 $1\%$ 标签微调提升了 $+2.39$ mIoU(在较小的 $N{=}40$ 验证集划分上提升了 $+4.10$);同样的修复也适用于单目深度和零样本 Structured3D 分割。该构造可证明是 $C_n$-不变的,并且以 $\mathcal{O}(n^{-2})$ 的误差逼近连续的 $\mathrm{SO}(2)$ 平均,使得所得到的模型成为一个可用的 $360^\circ$ 视觉计算原语,适用于全景重光照、沉浸式视频和跨数据集迁移。代码可在 https://github.com/Jaywalk18/equissl-release 获取。
cs.CV / 53 / 2609.31756

3dgs-sc: a controlled static screen-content benchmark for 3d gaussian splatting

3DGS-SC:一个用于3D高斯溅射的受控静态屏幕内容基准
Cai, Shicheng, Zhang, Hao, Dai, Dong, Ma, Xuerui, Hu, Ying, Zhang, Tao
Abstract
Does high image fidelity imply readable screen content in 3D Gaussian Splatting (3DGS)? We introduce 3DGS-SC, a controlled static screen-content dataset and benchmark for examining this mismatch. Ten procedural scenes provide fixed multi-view splits, exact cameras, screen masks, text boxes, and transcripts. The protocol separates whole-image fidelity, screen-region fidelity, OCR readability, and edge preservation. In the reported five-method comparison, LightGaussian exceeds Mip-Splatting in screen PSNR by only 0.08 dB, yet trails it in OCR accuracy by 12.7 percentage points. Across all ten method pairs, screen-PSNR and OCR orderings disagree in seven cases. These aggregate results expose a method-selection failure of fidelity-only evaluation. Complementing prior text-aware 3DGS research, 3DGS-SC targets controlled monitor interfaces and exact annotations; scene-wise robustness and acquisition effects remain open validation questions.
Chinese Translation
在3D高斯溅射(3DGS)中,高图像保真度是否意味着屏幕内容可读?我们引入了3DGS-SC,一个用于研究这种不匹配的受控静态屏幕内容数据集和基准。十个程序化场景提供了固定的多视角划分、精确相机、屏幕掩码、文本框和转录文本。该协议将全图保真度、屏幕区域保真度、OCR可读性和边缘保持分开评估。在报告的五个方法比较中,LightGaussian在屏幕PSNR上仅以0.08 dB超过Mip-Splatting,但在OCR准确率上落后12.7个百分点。在所有十个方法对中,屏幕PSNR和OCR排序在七种情况下不一致。这些汇总结果揭示了仅基于保真度的评估在方法选择上的失败。作为对先前文本感知3DGS研究的补充,3DGS-SC针对受控的显示器界面和精确标注;场景级鲁棒性和采集效应仍是待验证的开放问题。
cs.CV / 54 / 2609.31765

HGPTrans: Hierarchical Graph-Pooling Transolver for Automotive Aerodynamic Drag Coefficient Prediction

HGPTrans:用于汽车气动阻力系数预测的层次图池化Transolver
Liu, Bo, Luo, Qiuli, Nie, Lianrui, Zhang, Fengli, Wang, Wenjiang
Abstract
Accurate and rapid prediction of the aerodynamic drag coefficient ($C_D$) is essential for vehicle design, particularly during early-stage styling iterations where a large number of candidate geometries must be evaluated. Although computational fluid dynamics (CFD) provides reliable aerodynamic estimates, its high computational cost, typically requiring hours to days for a single configuration, limits its use in large-scale design exploration. This paper proposes HGPTrans, a hierarchical graph-pooling network with Transolver-based attention, to directly predict $C_D$ from vehicle surface meshes. Motivated by the fact that vehicle aerodynamics depends on both local geometric features and long-range interactions among spatially distant surface regions, HGPTrans integrates three complementary components. Graph isomorphism convolutions encode discriminative local geometry, physics-aware slice attention captures global interactions with linear computational complexity, and information-redundancy-aware hierarchical pooling progressively removes redundant nodes while preserving informative geometric structures. The model is trained and evaluated on the large-scale DrivAerNet and DrivAerNet++ datasets, where it achieves the lowest mean absolute error and mean squared error among the evaluated baselines. Its generalization capability is further assessed through transfer learning on a real-vehicle dataset containing both sedans and SUVs, achieving relative $L_1$ errors of 1.56% (sedans) and 2.12% (SUVs) with an inference time of approximately $0.293$ s per vehicle. This corresponds to an acceleration of several orders of magnitude relative to high-fidelity CFD while keeping the predicted drag coefficients within a few percent of the CFD reference. Ablation studies confirm each component's contribution and reveal the effects of depth and pooling ratio.
Chinese Translation
准确且快速地预测气动阻力系数($C_D$)对于车辆设计至关重要,尤其是在早期造型迭代阶段,此时需要评估大量候选几何形状。尽管计算流体力学(CFD)能够提供可靠的气动估计,但其高昂的计算成本(通常单个构型需要数小时至数天)限制了其在大规模设计探索中的应用。本文提出了HGPTrans,一种具有基于Transolver注意力的层次图池化网络,用于直接从车辆表面网格预测$C_D$。由于车辆空气动力学既依赖于局部几何特征,也依赖于空间上相距较远的表面区域之间的长程相互作用,HGPTrans整合了三个互补组件。图同构卷积编码具有判别性的局部几何,物理感知切片注意力以线性计算复杂度捕获全局交互,信息冗余感知的层次池化逐步移除冗余节点,同时保留信息丰富的几何结构。该模型在大规模DrivAerNet和DrivAerNet++数据集上进行训练和评估,在评估的基线方法中取得了最低的平均绝对误差和均方误差。其泛化能力进一步通过在包含轿车和SUV的真实车辆数据集上进行迁移学习来评估,实现了1.56%(轿车)和2.12%(SUV)的相对$L_1$误差,每辆车的推理时间约为$0.293$秒。相对于高保真CFD,这相当于加速了数个数量级,同时将预测的阻力系数保持在CFD参考值的百分之几以内。消融研究证实了每个组件的贡献,并揭示了深度和池化比例的影响。
cs.CV / 55 / 2609.31766

UNMATCH: Selective Unbalanced Token-Patch Matching for Forensic Image-Claim Verification

UNMATCH:用于取证图像-声明验证的选择性非均衡令牌-图块匹配
Li, Xinjin, Lian, Lian, Yang, Yuanzhe, Xia, Yudi, Liu, Calvin Chang, Xu, Yeyun, Ma, Yu, Cao, Jinghan, Gong, Yuruo
Abstract
Contextual image misuse pairs an image with a misleading claim. We study image-claim correspondence in fact-checked pairs containing out-of-context reuse, visual manipulation, or both. Existing pair-based detectors often compress the two modalities into a global compatibility score or learn a highly flexible interaction module, which can obscure a decisive local mismatch. We introduce directional multiscale coverage, a compact representation that summarizes local image-claim affinity in both directions and at three spatial scales. At each scale, each direction is summarized by its mean, lower quartile, and two thresholded support ratios; the signed difference between directional means completes a nine-dimensional scale descriptor. Concatenating the three scales yields a compact local representation for a lightweight global-local classifier. Under leakage-aware three-fold, three-seed evaluation on the Snopes subset of the Fauxtography benchmark, UNMATCH achieves 69.82 Macro-F1 and 71.05 balanced accuracy, exceeding the MCOT adaptation by 2.60 and 2.16 points. A matched-reassigned intervention shows that breaking the observed pairing lowers coverage and increases both discrepancy and false-pair probability.
Chinese Translation
情境图像滥用将图像与误导性声明配对。我们研究事实核查对中的图像-声明对应关系,这些对包含脱离上下文重用、视觉操纵或两者兼有。现有的基于对的检测器通常将两种模态压缩为一个全局兼容性得分,或学习一个高度灵活的交互模块,这可能掩盖决定性的局部不匹配。我们引入方向性多尺度覆盖,一种紧凑表示,总结局部图像-声明亲和力在双向和三个空间尺度上的情况。在每个尺度上,每个方向由其均值、下四分位数和两个阈值化支持率总结;方向均值之间的符号差完成了一个九维尺度描述符。连接三个尺度产生一个紧凑的局部表示,用于轻量级全局-局部分类器。在Fauxtography基准的Snopes子集上,进行泄漏感知的三折、三种子评估,UNMATCH达到69.82 Macro-F1和71.05平衡准确率,超过MCOT适配版2.60和2.16个点。匹配重分配的干预表明,打破观察到的配对会降低覆盖并增加差异和假对概率。
cs.CV / 56 / 2609.31780

Panoptic Scene Program Diffusion Transformer

全景场景程序扩散Transformer
Maduabuchi, Chika
Abstract
Modern text-to-image models produce high-fidelity images but still struggle with compositional prompts that require instance identity, attribute ownership, counting, spatial ordering, and role-sensitive relations. We introduce Panoptic Scene Program Diffusion Transformer (PSP-DiT), a diffusion-transformer architecture that treats a panoptic scene program as a first-class latent variable rather than an external control signal or post-hoc parse. PSP-DiT jointly denoises image latents and scene-program latents through coupled transformer streams, while panoptic grounding and cycle-consistency objectives tie object instances, attributes, relations, and counts to visual support in the generated image. Under matched training and inference settings, PSP-DiT improves over a strong flat-text baseline across GenEval 2, SANEval-Simple, PSG-Score, and DetailMaster, with the largest gains on counting, attribute binding, role-sensitive relations, and long structured prompts. The method preserves image quality, adds modest inference overhead, and remains robust to imperfect scene programs.
Chinese Translation
现代文本到图像模型能够生成高保真图像,但在处理需要实例身份、属性归属、计数、空间排序和角色敏感关系的组合式提示时仍存在困难。我们提出了全景场景程序扩散Transformer(PSP-DiT),一种扩散Transformer架构,它将全景场景程序视为一等潜变量,而非外部控制信号或事后解析。PSP-DiT通过耦合的Transformer流联合去噪图像潜变量和场景程序潜变量,同时全景接地和循环一致性目标将对象实例、属性、关系和计数与生成图像中的视觉支撑联系起来。在匹配的训练和推理设置下,PSP-DiT在GenEval 2、SANEval-Simple、PSG-Score和DetailMaster上优于强大的扁平文本基线,在计数、属性绑定、角色敏感关系和长结构化提示上提升最大。该方法保持了图像质量,增加了适度的推理开销,并且对不完美的场景程序保持鲁棒。
cs.CV / 57 / 2609.31788

SelfCue: Making a 3D CT Report Generator Say What It Already Knows

SelfCue:让3D CT报告生成器说出它已知的内容
Liang, Renjie, Yang, Yang, Pan, Jinqian, Fan, Zhengkang, Sun, Chengkun, Xu, Jie
Abstract
Progress in 3D CT report generation is usually sought in increasingly sophisticated architectures and larger pools of training data. We find instead that a 3D CT report generator already holds what its report leaves out, and loses it when the hidden state becomes tokens. Over the 18 CT-RATE abnormalities, this hidden-to-report surfacing gap is reflected by a drop in macro AUROC from 0.848 in the hidden states to 0.739 in the generated report. We propose SelfCue based on contrastive decoding. It promotes what the hidden state already supports and suppresses what it does not. It raises clinical efficacy F1 to 0.481 and the LLM-judged GREEN score to 0.510. Distilling that behaviour into the weights gives SelfCue-KD, a student that keeps most of the gain, needs nothing extra at inference, and drops into any pipeline already serving the baseline. Code is available at https://github.com/renjie-liang/SelfCue-CT.
Chinese Translation
3D CT报告生成的进展通常依赖于日益复杂的架构和更大的训练数据池。但我们发现,3D CT报告生成器其实已经掌握了其报告所遗漏的信息,而当隐藏状态转化为token时,这些信息便丢失了。在18种CT-RATE异常上,这种从隐藏状态到报告的浮现差距体现为宏AUROC从隐藏状态的0.848下降到生成报告的0.739。我们提出了基于对比解码的SelfCue。它促进隐藏状态已经支持的内容,并抑制其不支持的内容。它将临床有效性F1提升至0.481,将LLM评判的GREEN分数提升至0.510。将该行为蒸馏到权重中便得到SelfCue-KD,这是一个学生模型,它保留了大部分增益,推理时无需额外操作,并且可以无缝接入任何已服务于基线的流程中。代码见 https://github.com/renjie-liang/SelfCue-CT。
cs.CV / 58 / 2609.31795

Rate-Adaptive One-Step Diffusion Compression for AIGC Images

面向AIGC图像的自适应码率一步扩散压缩
Khanal, Nitiz
Abstract
We describe our entry to the LoViF 2026 AIGC Image Compression Challenge, a benchmark for ultra-low-bitrate coding of AI-generated images under a strict global rate budget of 0.025 bits per pixel (BPP). Generated imagery poses a distinct challenge for compression: it frequently contains rendered typography, synthetic edges, repeated motifs, UI-like layout, and stylized micro-texture that conventional distortion-oriented codecs erase at this rate, while unconstrained generative decoders can restore plausible-looking detail that no longer matches the source geometry or symbols. We treat this as a rate-perception allocation problem. Our system fine-tunes four rate-specialized checkpoints of the AEIC one-step diffusion codec, generates per-image candidates from all four, including one latent refined through encoder-side test-time optimization (TTO) with periodic entropy-conditioning refresh, entropy-codes every candidate with practical rANS coding, and selects exactly one bitstream per image with an exact multiple-choice knapsack solved over true coded file sizes. A fixed, zero-additional-bit residual restoration network is applied at decode time. Every submitted bitstream is independently decodable by the shipped decoder, which uses no source image or external side information. We report the full pipeline, an ablation history spanning 77 logged experiments, and a set of negative results, including why PSNR could not be pushed to parity with rate-distortion-oriented competitors under this architecture, useful to future participants. This is a challenge report: our entry scored 31.527739 (PSNR 27.02 dB, MS-SSIM 0.9176, LPIPS 0.0778, DISTS 0.0390 at 0.02495 BPP), the second-best DISTS on the leaderboard, ranking 5th on the final test-phase leaderboard announced August 4, 2026.
Chinese Translation
我们描述了参加LoViF 2026 AIGC图像压缩挑战赛的参赛方案,该挑战赛是AI生成图像在每像素0.025比特(BPP)的严格全局码率预算下进行超低码率编码的基准测试。生成的图像对压缩提出了独特的挑战:它经常包含渲染的排版、合成边缘、重复的图案、类似UI的布局以及风格化的微纹理,而传统的面向失真的编解码器在此码率下会将其抹去,而无约束的生成解码器可以恢复看似合理的细节,但这些细节不再匹配源几何结构或符号。我们将其视为一个速率-感知分配问题。我们的系统微调了AEIC一步扩散编解码器的四个速率专用检查点,从所有四个检查点生成每张图像的候选,包括一个通过编码器端测试时优化(TTO)并周期性熵条件刷新精炼的潜变量,使用实用的rANS编码对每个候选进行熵编码,并通过对真实编码文件大小求解精确的多重选择背包问题,为每张图像精确选择一个比特流。在解码时应用一个固定的、零额外比特的残差恢复网络。每个提交的比特流都可以由发布的解码器独立解码,该解码器不使用源图像或外部边信息。我们报告了完整的流程、跨越77次记录实验的消融历史以及一系列负面结果,包括为什么在该架构下PSNR无法与面向率失真的竞争对手持平,这对未来的参与者很有用。这是一份挑战赛报告:我们的参赛方案得分为31.527739(在0.02495 BPP下,PSNR为27.02 dB,MS-SSIM为0.9176,LPIPS为0.0778,DISTS为0.0390),在排行榜上DISTS排名第二,在2026年8月4日公布的最终测试阶段排行榜上排名第五。
cs.CV / 59 / 2609.31821

A Surgical Foundation Model Reveals Task-Dependent Label Efficiency

手术基础模型揭示任务依赖的标签效率
Stilz, Florian Philipp, Arboit, Lorenzo, Srivastav, Vinkle, Partners, CAMMA International Surgical, Marescaux, Jacques, Alfieri, Sergio, Mascagni, Pietro, Navab, Nassir, Padoy, Nicolas
Abstract
Developing label-efficient models is a central challenge in surgical AI due to the high cost and scarcity of expert annotation. While self-supervised foundation models adapt well to new tasks with minimal data, how label efficiency varies across different surgical tasks remains largely unexplored. Here, we introduce SURGE, a surgical foundation model trained on SurgSpectrum-30M+, the largest pretraining dataset comprising over 30 million frames, with checkpoints released to enable further research. We systematically evaluate label efficiency across 5 task categories and 15 benchmarks. These range from temporal and spatial scene understanding to fine-grained reasoning tied to instrument-anatomy interactions and safety-critical maneuvers. SURGE outperforms prior state-of-the-art on all benchmarks, even surpassing task-specific models on complex reasoning tasks. Crucially, we reveal a task-dependent scaling behavior: while scene understanding tasks saturate with minimal supervision, fine-grained reasoning tasks continue improving with substantially larger annotation budgets, providing a blueprint for allocating expert effort in complex domains. Code: https://github.com/CAMMA-public/SURGE
Chinese Translation
开发标签高效的模型是手术人工智能领域的一个核心挑战,因为专家标注成本高昂且稀缺。尽管自监督基础模型能够以最少的数据很好地适应新任务,但标签效率在不同手术任务间的变化仍很大程度上未被探索。在此,我们提出SURGE,一个在SurgSpectrum-30M+上训练的手术基础模型,该数据集是最大的预训练数据集,包含超过3000万帧,并发布了模型检查点以支持进一步研究。我们系统地评估了5个任务类别和15个基准上的标签效率。这些任务涵盖从时空场景理解到与器械-解剖结构交互及安全关键操作相关的细粒度推理。SURGE在所有基准上均超越了先前的最先进方法,甚至在复杂推理任务上超越了任务特定模型。关键的是,我们揭示了一种任务依赖的缩放行为:场景理解任务在极少量监督下即饱和,而细粒度推理任务随着标注预算的大幅增加持续提升,为在复杂领域中分配专家精力提供了蓝图。代码:https://github.com/CAMMA-public/SURGE
cs.CV / 60 / 2609.31823

SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions

SynDORBench:在物理约束可见性条件下评估LVLM的感知鲁棒性
Yee, Jeremy Stephen Gabriel, Wang, Zhengkui, Zhang, Zhiyuan, Anand, Avinash, Liu, Timothy, Chan, Benedict, Ng, Aik Beng, See, Simon
Abstract
Large vision-language models (LVLMs) have demonstrated remarkable performance on multimodal reasoning benchmarks, yet their perceptual reliability under physically constrained imaging conditions remains poorly understood. Existing evaluations predominantly assume ideal visual inputs and therefore fail to characterize how camera distance, illumination, viewpoint, and pixel density fundamentally affect semantic recoverability. We introduce SynDORBench, the first physically grounded benchmark for evaluating LVLM perceptual robustness under DORI-calibrated conditions aligned with human visual capability standards. SynDORBench comprises over 54k question--answer pairs generated through a controllable synthetic pipeline that systematically varies viewing distance, lighting, camera geometry, and action pose according to physically interpretable pixel-density regimes. To support scalable low-visibility supervision, we further propose a discernibility annotation framework that propagates human perceptual labels using mask-conditioned statistical features and ensemble learning. We evaluate 16 open-source LVLMs, a commercial LVLM baseline, and YOLO11x across human-presence classification and action recognition tasks under progressively degraded visibility conditions. Our results reveal that perceptual failure in LVLMs is strongly governed by pixel density and physical imaging constraints rather than model scale alone. Surprisingly, several compact open-source LVLMs outperform larger commercial baselines and substantially exceed YOLO11x robustness under long-range and low-light conditions. SynDORBench establishes a new benchmark paradigm for physically grounded multimodal evaluation, enabling systematic analysis of LVLM reliability under real-world perceptual constraints and direct comparison against human visibility thresholds.
Chinese Translation
大型视觉语言模型(LVLMs)在多模态推理基准上表现出色,但它们在物理约束成像条件下的感知可靠性仍知之甚少。现有评估大多假设理想的视觉输入,因此未能刻画相机距离、光照、视角和像素密度如何从根本上影响语义可恢复性。我们提出 SynDORBench,这是首个基于物理的基准,用于评估在与人类视觉能力标准对齐的 DORI 校准条件下的 LVLM 感知鲁棒性。SynDORBench 包含超过 54k 个问答对,这些问答对通过可控合成流程生成,该流程根据物理可解释的像素密度机制系统地改变观察距离、光照、相机几何和动作姿态。为了支持可扩展的低可见度监督,我们进一步提出了一种可辨别性标注框架,该框架使用掩码条件统计特征和集成学习来传播人类感知标签。我们在逐渐退化的可见性条件下,评估了 16 个开源 LVLM、一个商业 LVLM 基线和 YOLO11x 在人体存在分类和动作识别任务上的表现。我们的结果揭示,LVLM 的感知失败主要受像素密度和物理成像约束支配,而非仅仅由模型规模决定。令人惊讶的是,在远距离和低光照条件下,一些紧凑的开源 LVLM 优于更大的商业基线,并大幅超过 YOLO11x 的鲁棒性。SynDORBench 为基于物理的多模态评估建立了一种新的基准范式,使得能够系统分析 LVLM 在现实世界感知约束下的可靠性,并与人类可见度阈值进行直接比较。
cs.CV / 61 / 2609.31873

CueKFS: Agentic Cue-Driven Keyframe Selection for Long Video Understanding

CueKFS:面向长视频理解的智能体线索驱动关键帧选择
Kang, Weitai, Deilamsalehy, Hanieh, Xu, Yumo, Sultania, Dewang, Cellat, Serdar, Yan, Yan
Abstract
Keyframe selection (KFS) has long produced compact video summaries for browsing and retrieval, and representative frames for thumbnails. More recently, when conditioned on a question, KFS provides an alternative to uniform sampling for long-video question answering by selecting frames that are more relevant to the question. Most methods rank frames by similarity to the question. Yet a relevant frame may score poorly when the question combines subjects or moments that no single frame shows, or requires implicit information absent from its wording. Other methods try to break down the question into subqueries, but suffer from inaccurate decomposition due to their static initial context. Therefore, we propose CueKFS, a training-free method that reformulates question--frame matching as comparing frames against a set of dynamically generated visual cues. From an initial set of salient frames, we decompose the question into cues. Each cue concurrently probes the video to navigate to its own evidence. A reasoning VLM then agentically revises the cue set against its evidence to re-explore the video. CueKFS then allocates the budget across the surviving cues. Across three benchmarks, CueKFS establishes state-of-the-art results in all 27 evaluated settings with available prior results, achieving budget-averaged gains of up to +4.54% over the previous baseline and a median of only two VLM calls. We further provide a detailed behavioral analysis of CueKFS, showing that agentic cue refinement drives active re-exploration of the video, yielding relative similarity gains of up to 92% over the initial context.
Chinese Translation
关键帧选择(KFS)长期用于生成紧凑的视频摘要以支持浏览和检索,以及生成缩略图的代表性帧。近年来,当以问题为条件时,KFS为长视频问答提供了一种替代均匀采样的方法,通过选择与问题更相关的帧。大多数方法根据与问题的相似度对帧进行排序。然而,当问题组合了没有单帧能够展示的主体或时刻,或者需要其措辞中不存在的隐含信息时,一个相关的帧可能得分很低。其他方法试图将问题分解为子查询,但由于其静态的初始上下文而遭受不准确的分解。因此,我们提出CueKFS,一种免训练的方法,将问题-帧匹配重新表述为将帧与一组动态生成的视觉线索进行比较。从一组初始显著帧中,我们将问题分解为线索。每个线索同时探测视频以导航到其自身的证据。然后,一个推理VLM根据其证据自主地修正线索集以重新探索视频。CueKFS随后在存活的线索之间分配预算。在三个基准测试中,CueKFS在所有27个有先前结果可用的评估设置中均取得了最先进的结果,相比先前基线实现了高达+4.54%的预算平均增益,并且VLM调用次数的中位数仅为2次。我们进一步提供了CueKFS的详细行为分析,表明智能体线索细化驱动了对视频的主动重新探索,相比初始上下文实现了高达92%的相对相似度提升。
cs.CV / 62 / 2609.31896

Auditing Quality Filters for Long-Tail Human Data Curation

审查长尾人体数据策展中的质量过滤器
Agarwal, Rishav, Sekar, Nirshal Chandra, Vemula, Anirudh
Abstract
Robots on construction sites must detect workers who are kneeling or bending, which we call low poses. These workers can be lost from training datasets during automatic labeling. We study a pipeline that detects people, estimates their body joints using NLF, and groups similar poses. Low poses account for only about 2 percent of the retained examples. This low share may partly reflect the pipeline's quality filter, which rejects examples with low detection confidence or uncertain joint estimates. We examine this filtering using four alternative pose clues: bounding-box shape, vertical body span, pose grouping aligned to the scene's vertical direction, and image appearance. All four suggest that low poses are rejected by the filter more often. Separately, controlled simulated scenes show that a person detector fine-tuned on a public construction dataset misses more workers in these poses even when we correct their bounding box height is matched to that of standing workers. These findings suggest that low poses are scarce and hard to find, and we cannot rely on bounding boxes or poses for long-tail human data curation.
Chinese Translation
建筑工地上的机器人必须检测跪着或弯腰的工人,我们称之为低姿态。在自动标注过程中,这些工人可能会从训练数据集中丢失。我们研究了一个流程,该流程检测人体,使用NLF估计其身体关节,并对相似姿态进行分组。低姿态仅占保留样本的约2%。这一低比例可能部分反映了流程的质量过滤器,该过滤器会拒绝检测置信度低或关节估计不确定的样本。我们使用四种替代姿态线索来检验这种过滤:边界框形状、身体垂直跨度、与场景垂直方向对齐的姿态分组以及图像外观。所有四种线索都表明,低姿态更频繁地被过滤器拒绝。另外,受控模拟场景表明,在公共建筑数据集上微调的人体检测器在此类姿态中会漏检更多工人,即使我们将其边界框高度校正为与站立工人的高度相匹配。这些发现表明,低姿态稀少且难以发现,我们不能依赖边界框或姿态来进行长尾人体数据策展。
cs.CV / 63 / 2609.31911

Enhancing Visual Reasoning in Chest X-Ray Report Generation Using Reinforcement Learning

使用强化学习增强胸部X光报告生成中的视觉推理
Musinguzi, Denis, Katumba, Andrew, Mitra, Prasenjit
Abstract
Medical report generation has made significant progress with the rise of modern vision-language models and the growing availability of large-scale medical datasets. However, hallucinations remain a major challenge, largely due to the limitations of supervised fine-tuning (SFT), which prioritizes lexical similarity to reference reports rather than clinical correctness. While reinforcement learning has shown strong performance in domains with verifiable rewards such as mathematics and code generation, its application to open-ended medical tasks remains limited. Existing work focuses on evaluating final answers, overlooking the model's reasoning, despite evidence that flawed reasoning can degrade overall performance. In this study, we propose a framework that verifies the model's reasoning process by integrating anatomical regions, bounding boxes, and region-level textual descriptions. We design spatial and factual reward mechanisms to ensure that the model's reasoning is both visually grounded and factually accurate. Starting from Qwen3-VL-8B-Instruct as our base model, we adapt it to the medical domain using supervised fine-tuning, introduce reasoning capability through a cold-start SFT stage, and refine it with reinforcement learning. We find that RL provides performance gains beyond those achievable through SFT alone, and that jointly verifying both reasoning steps and final outputs yields larger improvements than verifying either in isolation. We further identify multiple modes of reward hacking in the RL stage. Finally, the model's structured think traces enhance interpretability, making its outputs easier to audit for clinical use.
Chinese Translation
医学报告生成随着现代视觉-语言模型的兴起和大规模医学数据集的日益可用,已经取得了显著进展。然而,幻觉仍然是一个主要挑战,这很大程度上归因于监督微调(SFT)的局限性,即其优先考虑与参考报告的词汇相似性,而非临床正确性。尽管强化学习在数学和代码生成等具有可验证奖励的领域已展现出强大性能,但其在开放式医学任务中的应用仍然有限。现有工作侧重于评估最终答案,而忽略了模型的推理过程,尽管有证据表明有缺陷的推理会降低整体性能。在本研究中,我们提出了一个框架,通过整合解剖区域、边界框和区域级文本描述来验证模型的推理过程。我们设计了空间和事实奖励机制,以确保模型的推理既具有视觉基础又事实准确。我们从Qwen3-VL-8B-Instruct作为基座模型开始,使用监督微调将其适应医学领域,通过冷启动SFT阶段引入推理能力,并用强化学习进行优化。我们发现,RL提供了仅通过SFT无法达到的性能提升,并且联合验证推理步骤和最终输出比单独验证任一者带来更大的改进。我们进一步识别了RL阶段中多种奖励破解模式。最后,模型的结构化思考轨迹增强了可解释性,使其输出更易于在临床使用中进行审查。
cs.CV / 64 / 2609.31912

PredRA: Fast Medical Image Translation by Deterministic Component Extraction and Controlled Stochastic Refinement

PredRA:通过确定性成分提取和受控随机细化实现快速医学图像翻译
Zhang, Jianhai, Charatpangoon, Pattarawut, Zhang, Donghao, Menon, Bijoy K., Qiu, Wu, MacDonald, M. Ethan, Ganesh, Aravind
Abstract
Strongly paired medical image translation contains a substantial component that can be predicted directly from the source. The current reality is that pure deterministic prediction can smooth away fine detail, while generative models can recover detail but may also introduce unnecessary or potentially harmful variation. We propose PredRA, a fast framework that uses the deterministic prediction as a stable reference and further extracts additional deterministic components from the residual during generative refinement, thereby improving fidelity while maintaining perceptual quality. The goal is simple: we view the entire residual-based generative process as an optimization problem and derive a practical solution that recovers useful residual detail through controlled refinement while keeping the prediction close to the paired target. Mechanism studies further show that useful residual information follows structured patterns, but its usefulness is difficult to estimate reliably at the voxel level. We therefore globally control how much of the residual refinement is added to the deterministic prediction, thereby reducing the accumulation of unnecessary uncertainty. PredRA therefore combines deterministic component extraction with controlled stochastic refinement for fast, fidelity-preserving, and perceptually strong medical image translation. We validate the approach across multiple real-world datasets, showing that controlled refinement improves fidelity to the paired target compared with full residual refinement while retaining much of the perceptual benefit of generative modeling. PredRA achieves competitive or superior performance to substantially larger state-of-the-art models with 1.4-11.9x fewer total parameters and 3.1-26.2x fewer trainable parameters, while its 32-step flow sampler requires 31.25x fewer sampling steps than the matched 1000-step DDPM.
Chinese Translation
强配对医学图像翻译包含一个可以从源图像直接预测的实质性成分。目前的现实是,纯确定性预测会平滑掉精细细节,而生成模型可以恢复细节,但可能引入不必要或潜在有害的变化。我们提出PredRA,一个快速框架,它使用确定性预测作为稳定参考,并在生成细化过程中从残差中进一步提取额外的确定性成分,从而在保持感知质量的同时提高保真度。目标很简单:我们将整个基于残差的生成过程视为一个优化问题,并推导出一个实用的解决方案,通过受控细化恢复有用的残差细节,同时保持预测接近配对目标。机制研究进一步表明,有用的残差信息遵循结构化模式,但其有用性在体素级别难以可靠估计。因此,我们全局控制将多少残差细化添加到确定性预测中,从而减少不必要不确定性的积累。因此,PredRA将确定性成分提取与受控随机细化相结合,实现快速、保真且感知强的医学图像翻译。我们在多个真实世界数据集上验证了该方法,表明与完全残差细化相比,受控细化提高了对配对目标的保真度,同时保留了生成建模的大部分感知益处。PredRA在总参数少1.4-11.9倍、可训练参数少3.1-26.2倍的情况下,与规模大得多的最先进模型相比,实现了竞争性或更优的性能,而其32步流采样器所需的采样步骤比匹配的1000步DDPM少31.25倍。
cs.CV / 65 / 2609.31915

Facial classification Using Hybrid Quantum Machine Learning

使用混合量子机器学习的面部分类
Bandlapalli, Roshan Babu, Katakam, Srinivas V, Chougala, Jitendra, Kappagantu, Ravi Kumar, Dontabhaktuni, Jayasri
Abstract
Hybrid quantum methods have received limited study for resource-constrained facial biometrics. We present a hybrid quantum-classical facial recognition pipeline designed to run on standard computing hardware. Images undergo gamma correction, contrast enhancement, and principal component analysis before their features are encoded into an eight-qubit variational quantum classifier. Classical image matching then performs recognition. In experiments with 50,000 images, comprising 25,000 faces from CelebA and 25,000 non-face images from CIFAR-10, the method outperformed the reported CPU-trained FaceNet baseline in accuracy and training efficiency. The pipeline was also evaluated on GPU and quantum hardware. In an attendance monitoring deployment at Mahindra University in collaboration with Lloyds Technology Centre, CPU inference took 0.2 to 0.5 seconds per person, and the system remained robust to the use of spectacles. These findings support the feasibility of deploying hybrid quantum methods for facial recognition on existing CPU hardware.
Chinese Translation
混合量子方法在资源受限的面部生物识别中研究有限。我们提出了一种混合量子-经典面部识别流水线,旨在在标准计算硬件上运行。图像经过伽马校正、对比度增强和主成分分析,然后其特征被编码到八量子比特的变分量子分类器中。随后通过经典图像匹配进行识别。在包含50,000张图像的实验中,其中包括来自CelebA的25,000张人脸和来自CIFAR-10的25,000张非人脸图像,该方法在准确率和训练效率上超过了所报告的CPU训练的FaceNet基线。该流水线也在GPU和量子硬件上进行了评估。在与Lloyds Technology Centre合作于Mahindra大学进行的考勤监控部署中,CPU推理每人耗时0.2至0.5秒,并且系统在使用眼镜的情况下仍保持稳健。这些发现支持在现有CPU硬件上部署混合量子方法进行面部识别的可行性。
cs.CV / 66 / 2609.31957

CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs

CaptchaArena:用于在交互式CAPTCHA上训练计算机使用智能体的大规模细粒度数据集
Zhang, Zhenhao, Fan, Zhaoyu, Ying, Haohan, Hu, Jingwen, Fan, Hancen, Zhou, Junhao, Chen, Zitian, Zhu, Linchao
Abstract
Interactive CAPTCHAs remain challenging for computer-use agents, while existing datasets face trade-offs among type coverage, interaction fidelity, and trajectory supervision. To address these gaps, we present CaptchaArena, the first large-scale, fine-grained training dataset for interactive CAPTCHA solving. It contains 50K puzzles across 20 CAPTCHA types and 5 interaction modes, with every solution verified through execution. CaptchaArena provides 50K screenshot-action trajectories, including 46K with step-by-step reasoning annotations. It also includes fine-grained pixel-mask annotations for irregular targets. Using CaptchaArena, we train CaptchaAgent, a single 9B policy for all 20 CAPTCHA types, with supervised fine-tuning followed by reinforcement learning. The environment verifier directly provides the RL reward. Supervised fine-tuning reaches 70.5 Pass@1, and reinforcement learning further improves it to 71.7, while also improving performance on two external benchmarks. These results demonstrate the value of large-scale, fine-grained computer-use supervision for training interactive CAPTCHA agents. We release CaptchaArena and CaptchaAgent at https://github.com/X0X0X00/CaptchaArena.
Chinese Translation
交互式CAPTCHA对计算机使用智能体仍然具有挑战性,而现有数据集在类型覆盖、交互保真度和轨迹监督之间面临权衡。为了弥补这些差距,我们提出了CaptchaArena,这是第一个用于交互式CAPTCHA求解的大规模细粒度训练数据集。它包含跨越20种CAPTCHA类型和5种交互模式的50K个谜题,每个解决方案都通过执行验证。CaptchaArena提供了50K个截图-动作轨迹,其中46K个带有逐步推理标注。它还包括针对不规则目标的细粒度像素掩码标注。使用CaptchaArena,我们训练了CaptchaAgent,一个适用于所有20种CAPTCHA类型的单一9B策略,采用监督微调后接强化学习。环境验证器直接提供RL奖励。监督微调达到70.5 Pass@1,强化学习进一步提升至71.7,同时还提高了两个外部基准的性能。这些结果证明了大规模细粒度计算机使用监督对于训练交互式CAPTCHA智能体的价值。我们在https://github.com/X0X0X00/CaptchaArena发布了CaptchaArena和CaptchaAgent。
cs.CV / 67 / 2609.31970

Double-Edged Sword of Mediated Visibility: How Visual Framing Undermines Congresswomen's Perceived Competence

中介化可见性的双刃剑:视觉框架如何削弱女国会议员的感知能力
Dietrich, Bryce J., Ko, Hyein, Shiran, Myriam
Abstract
Women's presence in Congress has stalled below 30%. That institutional underrepresentation is mirrored by their limited visibility on televised news, where visual presentation can shape perceived competence and authority. Yet research tells us more about whether congresswomen appear than how they are visually framed. Facial recognition analysis of 695,464 image-text segments from CNN and Fox News (2011-2021) reveals that congresswomen appear disproportionately in split-screen rather than solo shots. Two pre-registered experiments with 6,220 participants show that, in static frames, split-screen framing reduces congresswomen's perceived political competence, but not congressmen's. In dynamic videos, the competence penalty disappears; instead, outraged language reduces a congresswoman's perceived warmth roughly three times as much as a congressman's. Suggestive evidence (p = .057) indicates that women viewers report greater external political efficacy after watching a congresswoman appear alone. We conclude that visual framing shapes both congresswomen's mediated visibility and their perceived capabilities.
Chinese Translation
女性在国会中的比例一直停滞在30%以下。这种制度性代表不足也反映在她们在电视新闻中有限的可见性上,而视觉呈现可以塑造感知能力和权威。然而,研究更多地告诉我们女国会议员是否出现,而不是她们如何被视觉框定。对CNN和Fox News(2011-2021)的695,464个图像-文本片段进行面部识别分析显示,女国会议员出现在分屏镜头中的比例过高,而非单人镜头。两项预注册实验(6,220名参与者)表明,在静态画面中,分屏框架降低了女国会议员的感知政治能力,但对男国会议员没有影响。在动态视频中,能力惩罚消失了;相反,愤怒的语言对女国会议员感知热情的降低程度大约是男国会议员的三倍。提示性证据(p = .057)表明,女性观众在观看女国会议员单独出现后,报告了更高的外部政治效能感。我们得出结论,视觉框架既塑造了女国会议员的中介化可见性,也塑造了她们的感知能力。
cs.CV / 68 / 2609.31985

Does Vision-Language Pretraining Granularity Matter? A Controlled Evaluation of Vision-Language Objectives Across Chest X-Ray Interpretation Tasks

视觉语言预训练粒度重要吗?——跨胸部X光解读任务的视觉语言目标受控评估
Musinguzi, Denis, Katumba, Andrew, Mitra, Prasenjit
Abstract
Vision-language pretraining objectives differ in the spatial granularity of their supervision, yet the implications of this distribution for chest X-ray interpretation remain underexplored. We present a controlled study that isolates the pretraining objective: holding the encoder and pretraining data fixed, we train nine objectives spanning global and local contrastive learning, captioning, and their combinations, and evaluate across five chest X-ray tasks of increasing spatial granularity. We find that (i) pretraining granularity aligns with task granularity at the extremes, with local objectives leading on abnormality detection and global objectives on classification; (ii) local objectives are surprisingly competitive on global-level generation and question answering tasks; (iii) the merits of captioning and contrastive learning reverse across granularity levels; and (iv) among combinations, mixing captioning and contrastive supervision is strongest on classification and in distribution generation, while pairing two captioning objectives generalizes best on zero-shot report generation. These results show that no single objective is universally optimal, and that the interaction of objective type, granularity, and task governs downstream performance.
Chinese Translation
视觉语言预训练目标在其监督的空间粒度上存在差异,然而这种分布对胸部X光解读的影响仍未被充分探索。我们提出了一项隔离预训练目标的受控研究:在固定编码器和预训练数据的情况下,我们训练了九个目标,涵盖全局和局部对比学习、图像描述及其组合,并在五个空间粒度递增的胸部X光任务上进行评估。我们发现:(i) 在极端情况下,预训练粒度与任务粒度一致,局部目标在异常检测上领先,而全局目标在分类上领先;(ii) 局部目标在全局级生成和问答任务上出乎意料地具有竞争力;(iii) 图像描述和对比学习的优势在不同粒度级别上发生反转;(iv) 在组合中,混合图像描述和对比监督在分类和分布生成上最强,而配对两个图像描述目标在零样本报告生成上泛化最佳。这些结果表明,没有单一目标是普遍最优的,并且目标类型、粒度和任务之间的相互作用决定了下游性能。
cs.CV / 69 / 2609.31998

Type-Balanced Federated Learning for Visual Analog Meter Reading

面向视觉模拟仪表读数的类型平衡联邦学习
Zhao, Weida, Bellamy, Logan, Tu, Yazhou, Wang, Jiaqi
Abstract
Analog dial meters are widely deployed in industrial applications and utility sites, where environments and meter types vary and inspection data may be sensitive. Currently, automatic meter readers must be individually developed and deployed for each environment and meter type in practice. Deep learning could handle this variability but requires diverse labeled data that are costly to collect and update. In practice, meter images are distributed across independent sites, each with limited labels, while raw images often cannot be pooled because of ownership, governance, or privacy constraints. To address these challenges, we present a federated framework for visual analog meter reading that enables multiple sites to collaboratively train a reading model without sharing their raw images. Our framework consists of a four-stage pipeline: (1) dial localization, (2) thin-structure segmentation trained federatively across clients, (3) polar unwrapping, and (4) tick-counting decoding for final reading. To enable systematic evaluation of this setting, we release MeterFL, a 1,382-image mask-annotated dataset organized into deployment-motivated pseudo-clients derived from visual attributes via deterministic rules, with dHash near-duplicate control between the segmentation train and test splits. We evaluate both segmentation quality and end-to-end reading accuracy. MeterFL is publicly available at https://github.com/weidazhaoooo/Meter-FL.
Chinese Translation
模拟表盘仪表广泛应用于工业应用和公用事业场所,其中环境和仪表类型各异,检查数据可能敏感。目前,自动抄表器必须针对每种环境和仪表类型单独开发和部署。深度学习可以处理这种可变性,但需要大量标注数据,收集和更新成本高昂。在实践中,仪表图像分布在独立的站点,每个站点的标注有限,而原始图像通常由于所有权、治理或隐私限制而无法集中。为了应对这些挑战,我们提出了一个用于视觉模拟仪表读数的联邦框架,使多个站点能够协作训练读数模型,而无需共享原始图像。我们的框架包括四个阶段的流程:(1) 表盘定位,(2) 跨客户端联邦训练的薄结构分割,(3) 极坐标展开,以及 (4) 用于最终读数的刻度计数解码。为了对该设置进行系统评估,我们发布了MeterFL,一个包含1,382张图像的掩码标注数据集,通过确定性规则根据视觉属性组织成部署驱动的伪客户端,并在分割训练和测试划分之间进行dHash近重复控制。我们评估了分割质量和端到端读数准确性。MeterFL 公开可用,网址为 https://github.com/weidazhaoooo/Meter-FL。
cs.CV / 70 / 2609.32013

TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception

TriO: 面向万物感知的三模态无监督占用世界模型
Sykora, Quinlan, Biswas, Sourav, Diehl, Christopher, Cunningham, Andrew, Gilles, Thomas, Urtasun, Raquel
Abstract
We present TriO, a multi-modal unsupervised world model that predicts 4D occupancy, obstacle segmentation, flow and LiDAR. In contrast to prior work, TriO utilizes three distinct sensor modalities (camera, LiDAR, and RADAR) as both inputs and sources of self-supervision, eliminating the need for additional human annotations. Thanks to its novel supervision, the model is able to segment any occupancy from the drivable surface, overcoming the limitations of existing open-set methods in handling long-tail objects. TriO achieves state-of-the-art results in multiple 3D and 4D tasks, including occupancy, flow, and LiDAR prediction, as well as zero-shot road obstacle segmentation across multiple datasets such as Argoverse 2, and Spotting the Unexpected.
Chinese Translation
我们提出了 TriO,一个多模态无监督世界模型,能够预测4D占用、障碍物分割、光流和LiDAR。与先前工作相比,TriO 利用三种不同的传感器模态(摄像头、LiDAR 和 RADAR)作为输入和自监督来源,无需额外的人工标注。得益于其新颖的监督方式,该模型能够从可行驶表面分割出任何占用,克服了现有开放集方法在处理长尾物体时的局限性。TriO 在多个3D和4D任务中取得了最先进的结果,包括占用、光流和LiDAR预测,以及在多个数据集(如 Argoverse 2 和 Spotting the Unexpected)上的零样本道路障碍物分割。
cs.CV / 71 / 2609.32027

Depth Any Seen: Which Surfaces and How Far?

Depth Any Seen:哪些表面以及距离多远?
Xu, Xiaohao, Huang, Xiaonan
Abstract
When several surfaces are visible along a ray, recovering visible 3D structure from one image requires jointly estimating their presence and metric depth. Depth Any Seen represents these surfaces as image-conditioned multi-Bernoulli depth sets, whose components each contribute one depth or remain absent. Its auxiliary-free Exact Multi-Bernoulli objective (ExactMB) learns depth and presence by marginalizing one-to-one assignments to complete, distinct targets. Our analysis shows that matching expected count can leave component-surface assignment unresolved. We extend real and synthetic layered-depth benchmarks to evaluate depth accuracy, recovered support, and overprediction. Compared to depth stacking, ExactMB reduces overprediction by a relative 88.2% on LD-Real and 80.5% on MD-3K while retaining most ordinal accuracy, with comparable conditional metric-depth error on LD-Syn. Further ablation studies show that ordered assignment improves depth-accurate recall and precision over marginalization, whereas the count-regularized configuration achieves higher deeper-rank precision than ordered assignment at lower recall. Our code will be publicly released.
Chinese Translation
当沿一条射线可见多个表面时,从单张图像中恢复可见的3D结构需要联合估计它们的存在性与度量深度。Depth Any Seen 将这些表面表示为图像条件化的多伯努利深度集合,其每个分量要么贡献一个深度,要么保持缺失。其无辅助的精确多伯努利目标(ExactMB)通过边缘化对完整且不同目标的一对一分配来学习深度和存在性。我们的分析表明,匹配期望计数可能导致分量-表面分配未解决。我们扩展了真实和合成分层深度基准,以评估深度精度、恢复的支撑和过度预测。与深度堆叠相比,ExactMB 在 LD-Real 上相对减少了 88.2% 的过度预测,在 MD-3K 上减少了 80.5%,同时保留了大部分顺序精度,在 LD-Syn 上具有相当的条件度量深度误差。进一步的消融研究表明,与边缘化相比,有序分配提高了深度准确的召回率和精确率,而计数正则化配置在较低召回率下实现了比有序分配更高的更深排名精度。我们的代码将公开发布。
cs.CV / 72 / 2609.32036

ScreenHaystack: Finding Blind Zones in GUI Grounding

ScreenHaystack: 发现GUI定位中的盲区
Li, Chenyue, Sun, Xiaoxiao, Deng, Yubo, Zhao, Qinlin, Yeung-Levy, Serena, Zhang, Yuhui
Abstract
We introduce ScreenHaystack, a dynamic needle-in-a-haystack benchmark for evaluating spatial reliability in GUI grounding. Instead of testing each target at a fixed position, ScreenHaystack systematically relocates controlled target icons across high-resolution GUI backgrounds and measures whether models can localize them consistently. Using this benchmark, we find that leading GUI grounding models, including Qwen3-VL, UI-TARS, GTA, and UI-Venus, exhibit blind zones: spatial regions where grounding accuracy drops sharply despite fixed target appearance and instruction. These blind zones transfer to unseen ScreenSpot-Pro examples: targets inside blind zones are consistently harder to ground, with Qwen3-VL-8B dropping by 16.1 percentage points, and controlled relocation shows that moving targets into blind zones decreases accuracy while moving them out improves accuracy. We further show through controlled synthetic experiments that uneven spatial coverage in training data can induce such blind zones. Therefore, we propose a simple strategy, blind-zone-oriented augmentation, which adds supervision in blind zones and improves ScreenSpot-Pro accuracy over both original and randomly augmented Click-100k fine-tuning.
Chinese Translation
我们介绍了ScreenHaystack,一个动态的大海捞针基准,用于评估GUI定位中的空间可靠性。ScreenHaystack不是测试每个目标在固定位置,而是系统地将受控的目标图标重新定位到高分辨率GUI背景上,并测量模型是否能一致地定位它们。使用该基准,我们发现领先的GUI定位模型,包括Qwen3-VL、UI-TARS、GTA和UI-Venus,表现出盲区:即空间区域,其中尽管目标外观和指令固定,定位准确率却急剧下降。这些盲区可迁移到未见过的ScreenSpot-Pro示例:盲区内的目标始终更难定位,Qwen3-VL-8B下降了16.1个百分点,并且受控的重新定位表明,将目标移入盲区会降低准确率,而移出则会提高准确率。我们进一步通过受控的合成实验表明,训练数据中不均匀的空间覆盖可以诱发此类盲区。因此,我们提出一种简单策略,即面向盲区的增强,它在盲区添加监督,并相较于原始和随机增强的Click-100k微调,提高了ScreenSpot-Pro准确率。
cs.CV / 73 / 2609.32038

ControlGS: Conditioning Neural Gaussians for Downstream-Processing-Aware XR Rendering

ControlGS:为下游处理感知的XR渲染条件化神经高斯
Lin, Weikai, Zhao, Junjie, Marshall, Carl, Kondguli, Sushant, Zhu, Yuhao
Abstract
Extended Reality (XR) users do not directly perceive the output of a rendering engine. Instead, rendered images pass through a post-processing pipeline and the physical display-optics path before reaching the eye. Critically, the exact downstream processing can vary significantly at run time, influenced by, for instance, camera pose and display power budget. Traditional 3DGS methods either implicitly assume that this downstream pipeline preserves image quality or cannot adapt to downstream processing changes. To bridge this gap, we present ControlGS, an XR Gaussian rendering pipeline that optimizes end-to-end visual quality. ControlGS models and integrates the entire downstream processing, between the rendering output and the human eye, into the optimization objective. To adapt to downstream processing at run time, ControlGS dynamically generates Gaussian primitives conditioned upon the downstream processing parameters. Experiments show that ControlGS consistently improves end-to-end post-optics XR quality across different neural Gaussian backbones and datasets, with minimal overhead. Code is available at https://horizon-lab.org/controlgs/.
Chinese Translation
扩展现实(XR)用户并不直接感知渲染引擎的输出。相反,渲染图像在到达人眼之前会经过后处理管线和物理显示光学路径。关键的是,确切的下游处理在运行时可能显著变化,例如受到相机姿态和显示器功率预算的影响。传统的3DGS方法要么隐含假设该下游管线保持图像质量,要么无法适应下游处理的变化。为了弥合这一差距,我们提出了ControlGS,一种优化端到端视觉质量的XR高斯渲染管线。ControlGS建模并将渲染输出与人眼之间的整个下游处理集成到优化目标中。为了在运行时适应下游处理,ControlGS根据下游处理参数动态生成高斯基元。实验表明,ControlGS在不同神经高斯骨干网络和数据集上持续改善端到端后光学XR质量,且开销最小。代码见 https://horizon-lab.org/controlgs/。
cs.CV / 74 / 2609.32041

Amnesia by Design, Memory By Necessity: Persistent State for Document Intelligence

设计性失忆,必要性记忆:面向文档智能的持久状态
Bakkali, Souhail, Merimi, Ayoub
Abstract
Modern Document AI reads contracts, extracts fields, reasons over tables, and grounds answers to page regions, then forgets everything. Processing an amendment the next day begins from scratch: no schema retained, no contradiction detected, no experience carried forward. This is a structural choice, not a scale failure: current systems are stateless functions. We call this the statelessness bottleneck. This bottleneck lies beyond parameter scaling, context extension, and retrieval augmentation: storage provides persistence and retrieval provides access, but neither consolidates observations into knowledge that improves future processing. This survey formalizes persistent evidence-grounded document state as a unifying framework, specifying the operations and invariants required to convert multimodal evidence into durable, provenance-linked state. We introduce a statefulness audit showing that ten representative benchmarks, coded against eight statefulness criteria, leave cross-session state evolution untested, and derive a longitudinal benchmark harness with five counterfactual metrics: Experience Gain, Cost Efficiency, Memory Harm, Forgetting Fidelity, Coverage Retention, to characterize the benefit, cost, risk, and governability of persistent document state. Document AI lacks mechanisms coupling persistent state to document-native structure, provenance, and temporal validity. The next era of Document AI will be defined by what systems retain across documents, sessions, and time.
Chinese Translation
现代文档AI读取合同、提取字段、对表格进行推理、并将答案定位到页面区域,然后忘记一切。第二天处理修订时从头开始:没有保留模式,没有检测到矛盾,没有经验延续。这是一种结构性选择,而非规模故障:当前系统是无状态函数。我们将此称为无状态瓶颈。这一瓶颈超越了参数缩放、上下文扩展和检索增强:存储提供持久性,检索提供访问,但两者都没有将观察结果整合为能够改善未来处理的知识。本综述将持久的、基于证据的文档状态形式化为一个统一框架,具体说明了将多模态证据转换为持久的、与来源链接的状态所需的操作和不变量。我们引入了一项状态性审计,显示十个代表性基准(根据八项状态性标准编码)未测试跨会话状态演化,并推导出一个纵向基准测试框架,包含五个反事实指标:经验增益、成本效率、记忆危害、遗忘保真度、覆盖保留,以表征持久文档状态的收益、成本、风险和可治理性。文档AI缺乏将持久状态与文档原生结构、来源和时间有效性耦合的机制。文档AI的下一个时代将由系统跨文档、会话和时间保留的内容来定义。
cs.CV / 75 / 2609.32068

ReFM: Semantic-Aware Refinement Flow Model for Motion Retargeting

ReFM:面向运动重定向的语义感知细化流模型
Qu, Jingxiang, Taglienti, Lucie, Atherton, Evan
Abstract
Motion retargeting transfers motion across characters with different skeletal structures while preserving semantic intent and physical plausibility. Despite recent progress, two fundamental questions remain: (i) how can reliable source-motion semantics be learned without high-quality paired retargeting data, and (ii) how should retargeting be formulated when no reliable paired motion can serve as a definitive regression objective? Existing methods commonly preserve semantics by constraining predictions toward copied motions. However, such initializations entangle useful articulation cues with artifacts caused by mismatched skeletal proportions and body geometry. Moreover, directly regressing a final motion in one forward pass is restrictive because retargeting is inherently underdetermined, and the desired solution must balance semantic fidelity with target-specific physical and temporal constraints rather than match a unique paired target. Motivated by these limitations, we propose ReFM, a source-mesh-agnostic, energy-guided model that reformulates motion retargeting as progressive refinement. First, an SO(3) canonicalizer removes redundant global-orientation variations. Second, a cross-character semantic encoder, pretrained through contrastive learning, provides a character-invariant representation for both optimization guidance and semantic evaluation. ReFM then progressively refines an initialized target motion through a learned flow guided by semantic consistency, physical plausibility, temporal coherence, and minimal motion modification. The framework is compatible with different initialization strategies, including both direct motion copying and Autodesk HumanIK, an industry-standard full-body inverse-kinematics retargeting system.
Chinese Translation
运动重定向在不同骨骼结构的角色之间迁移运动,同时保持语义意图和物理合理性。尽管近期取得了进展,仍有两个基本问题:(i) 在没有高质量配对重定向数据的情况下,如何学习可靠的源运动语义?(ii) 当没有可靠的配对运动可作为明确的回归目标时,应如何构建重定向问题?现有方法通常通过将预测约束到复制的运动上来保持语义。然而,这种初始化将有用的关节运动线索与因骨骼比例和身体几何形状不匹配而产生的伪影纠缠在一起。此外,在一次前向传递中直接回归最终运动是有局限的,因为重定向本质上是欠定的,期望的解必须在语义保真度与目标特定的物理和时间约束之间取得平衡,而不是匹配唯一的配对目标。受这些局限性的启发,我们提出 ReFM,一种与源网格无关的、能量引导的模型,将运动重定向重新表述为渐进细化。首先,一个 SO(3) 规范化器去除冗余的全局方向变化。其次,一个通过对比学习预训练的跨角色语义编码器提供角色不变的表示,用于优化指导和语义评估。然后,ReFM 通过一个学习到的流逐步细化初始化的目标运动,该流由语义一致性、物理合理性、时间连贯性和最小运动修改引导。该框架兼容不同的初始化策略,包括直接运动复制和 Autodesk HumanIK(一种行业标准的全身逆运动学重定向系统)。
cs.CV / 76 / 2609.32131

Gaussian Image Steganography via Parameter-Domain Keyed Embeddings

通过参数域密钥嵌入的高斯图像隐写术
Wu, Tong, Cheng, Runze, Fan, Xiaoyue, Akşit, Kaan
Abstract
2D Gaussian-based image representation is becoming increasingly popular, and our work proposes a new approach to steganography by embedding information within Gaussian parameters rather than image pixels. We first fit the parameters of this representation to the target image and employ a secret key to select a subset of these parameters for fine-tuning, allowing us to embed an 8-bit message while maintaining high visual fidelity in the reconstructed image. Thirty fitting experiments on three synthetic images show that the correct key can recover the message without error, while decoding with incorrect keys yields a Bit Error Rate (BER) of $0.543$, close to random guessing. Compared with random selection using the key, selecting the least-disturbing edits recovers the message more reliably (one-sided $p=0.031$), and the average PSNR cost is only $0.091$ dB in visual quality. The embedding method transfers to 112 natural images at $256\times256$ using 4,096 Gaussians. The correct-key BER is $0.000$, and decoding under a wrong key stays close to random guessing at $0.520$. Our method embeds the payload through three Gaussian parameter types: log-anisotropy, opacity, and color luminance. In a separate nine-fit reduced setting, color luminance is removed, so the payload uses two instead of three parameter types, a $33.3\%$ reduction; all 256 Gaussians remain in the fitted representation. The correct key still recovers the message without error. However, wrong-key BER rises from $0.514$ to $0.571$, moving farther from random guessing ($0.5$).
Chinese Translation
基于2D高斯的图像表示正变得越来越流行,我们的工作提出了一种新的隐写方法,通过在高斯参数中而不是图像像素中嵌入信息。我们首先将该表示的参数拟合到目标图像,并使用密钥选择这些参数的一个子集进行微调,从而嵌入一个8位消息,同时保持重建图像的高视觉保真度。在三张合成图像上的三十次拟合实验表明,正确的密钥可以无误差地恢复消息,而使用错误密钥解码得到的误码率(BER)为0.543,接近随机猜测。与使用密钥的随机选择相比,选择扰动最小的编辑更可靠地恢复消息(单侧p=0.031),并且视觉质量的平均PSNR代价仅为0.091 dB。该嵌入方法迁移到112张256×256的自然图像,使用4,096个高斯。正确密钥的BER为0.000,错误密钥解码的BER接近随机猜测,为0.520。我们的方法通过三种高斯参数类型嵌入有效载荷:对数各向异性(log-anisotropy)、不透明度(opacity)和颜色亮度(color luminance)。在一个单独的九次拟合简化设置中,移除了颜色亮度,因此有效载荷使用两种而不是三种参数类型,减少了33.3%;所有256个高斯仍保留在拟合表示中。正确密钥仍可无误差地恢复消息。然而,错误密钥的BER从0.514上升到0.571,更远离随机猜测(0.5)。
cs.CV / 77 / 2609.32157

CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving

CausalDriveBench:评估自动驾驶中视觉-语言-动作模型的因果推理
Chembu, Narendiran, Rao, Navvrat, Kodate, Shreedhar Shreeshail, Banda, Gayatri Srujana, Sarkar, Arko, Khanna, Abhinav, Das, Rajarshee, Kanala, Umesh, Khandelwal, Siddarth, Aman, Kumar, Dubey, Aish, Beedkar, Kaustubh, Jain, Arjun
Abstract
Vision-Language-Action (VLA) models for autonomous driving produce natural-language reasoning alongside predicted trajectories, but whether this reasoning reflects the causal structure of the scene remains untested. We introduce CausalDriveBench, an evaluation framework grounded in Pearl's Causal Hierarchy (PCH) that tests causal reasoning in driving-specific VLAs through structured visual question answering (QA) and alternative-trajectory prediction. To this end, we construct causal scene graphs over nuScenes that distinguish causally active, dormant, and distractor entities, separating perceptual salience from causal relevance. The benchmark spans all four rungs of PCH (association, intervention, and counterfactual along with causal discovery) for QA generation. For the higher rungs, we additionally provide reference trajectories under specified scene modifications, enabling action-level verification that complements reasoning-level evaluation. In total, the benchmark contains 7,285 verified causal QA pairs and 1,000 counterfactual trajectories derived from nuScenes. We evaluate 10 driving-specific VLAs and 3 general-purpose VLMs, and report three findings. First, the best model reaches only 70.6% QA accuracy, and 4 of 13 models score below random chance. Second, comparing each driving VLA to the general-purpose VLM that shares its language backbone, the cost of driving fine-tuning ranges from 2 to 34 percentage points on causal QA, with post-training design explaining the spread. Third, causal QA and trajectory accuracy are statistically uncorrelated across models: under counterfactual prompts, predicted trajectories either over-react or collapse onto the observed-scene baseline. Taken together, these results show that neither fluent rationales nor accurate observed-scene trajectories constitute evidence of causal understanding.
Chinese Translation
用于自动驾驶的视觉-语言-动作(VLA)模型在预测轨迹的同时生成自然语言推理,但这种推理是否反映了场景的因果结构仍未经检验。我们提出了 CausalDriveBench,一个基于 Pearl 因果层级(PCH)的评估框架,通过结构化视觉问答(QA)和替代轨迹预测来测试驾驶专用 VLA 的因果推理能力。为此,我们在 nuScenes 上构建因果场景图,区分因果活跃、休眠和干扰实体,将感知显著性与因果相关性分离开来。该基准覆盖 PCH 的全部四个层级(关联、干预、反事实以及因果发现)用于 QA 生成。对于更高层级,我们还提供了在指定场景修改下的参考轨迹,从而能够进行动作级验证,以补充推理级评估。总之,该基准包含 7,285 个经过验证的因果 QA 对和 1,000 条源自 nuScenes 的反事实轨迹。我们评估了 10 个驾驶专用 VLA 和 3 个通用 VLM,并报告了三个发现。第一,最佳模型仅达到 70.6% 的 QA 准确率,13 个模型中有 4 个得分低于随机水平。第二,将每个驾驶 VLA 与共享其语言主干的通用 VLM 进行比较,驾驶微调的代价在因果 QA 上从 2 到 34 个百分点不等,后训练设计解释了这种差异。第三,因果 QA 和轨迹准确率在不同模型之间统计上不相关:在反事实提示下,预测轨迹要么反应过度,要么坍缩到观测场景基线。综上所述,这些结果表明,无论是流畅的推理依据还是准确的观测场景轨迹,都不构成因果理解的证据。
cs.CV / 78 / 2609.32162

PruneForget: Joint Unlearning and Pruning of Vision Models

PruneForget:视觉模型的联合遗忘与剪枝
Tai, Yu-Shan, Zheng, Amber Yijia, Yeh, Raymond A.
Abstract
Machine unlearning and model pruning are increasingly coupled in the real world. Models must support unlearning requests, e.g., for safety concerns, while also meeting requirements in latency and memory budget. Until recently, existing works have studied each aspect as an independent problem, e.g., running unlearning and pruning sequentially. In this work, we show that unlearning and pruning are naturally aligned and should be solved jointly to be made aware of each other. Intuitively, parameters that encode information of the unlearned samples are natural pruning targets, as unlearning and pruning both call for the "deletion" of such parameters. We propose PruneForget, a method that uses the unlearn set as a guide for pruning, so that unlearning and pruning mutually benefit each other. Extensive experiments on image classifiers and generative models show that PruneForget removes the influence of the unlearned samples while producing a more compact model with reduced inference cost. It achieves a negligible performance gap relative to an oracle that retrains from scratch for unlearning and then prunes.
Chinese Translation
机器遗忘和模型剪枝在现实世界中日益耦合。模型必须支持遗忘请求(例如出于安全考虑),同时还要满足延迟和内存预算方面的要求。此前,已有工作将每个方面作为独立问题进行研究,例如顺序地执行遗忘和剪枝。在这项工作中,我们表明遗忘和剪枝天然一致,应当联合求解,使二者相互感知。直观上,编码了被遗忘样本信息的参数是天然的剪枝目标,因为遗忘和剪枝都要求“删除”这些参数。我们提出 PruneForget,一种利用遗忘集作为剪枝指导的方法,从而使遗忘和剪枝相互受益。在图像分类器和生成模型上的大量实验表明,PruneForget 在消除被遗忘样本影响的同时,生成更紧凑的模型并降低推理成本。与从头重新训练以进行遗忘然后再剪枝的 oracle 相比,它实现了可忽略的性能差距。
cs.CV / 79 / 2609.32163

PQR3D: Progressive Query Refinement over Reference-Conditioned Temporal Windows for Multi-View 3D Object Detection

PQR3D:在参考条件时间窗口上进行渐进式查询细化用于多视图3D目标检测
Ye, Hui, Liu, Yudong, Chen, Yiran, Sunderraman, Rajshekhar, Ji, Shihao
Abstract
Temporal context is essential for camera-only multi-view 3D object detection. Existing streaming detectors maintain and propagate query states from one frame to the next, requiring sequence-aware training and chronological inference. We propose PQR3D, which performs progressive query refinement within referenceconditioned temporal windows. This design enables random frame sampling and independent inference without persistent query memory. Within each window, PQR3D progressively transfers motion-aligned high-confidence queries from earlier timestamps toward the target frame. We further introduce masked selfattention to regulate interactions among regular, propagated, and denoising queries while keeping denoising supervision isolated from detection queries. In addition, a stage-decoupled anchor embedding injects position before self-attention and size, orientation, and velocity afterward, reducing interference from temporally inconsistent attributes. With a ViT-L backbone, PQR3D sets a new state of the art on the nuScenes test set, achieving 71.6 NDS and 64.9 mAP. Source code is available at https://github.com/huiyegit/PQR3D
Chinese Translation
时间上下文对于仅相机多视图3D目标检测至关重要。现有的流式检测器维护并传播查询状态从一帧到下一帧,需要序列感知训练和按时间顺序推理。我们提出PQR3D,它在参考条件时间窗口内进行渐进式查询细化。这种设计支持随机帧采样和独立推理,而无需持久查询内存。在每个窗口内,PQR3D逐步将运动对齐的高置信度查询从较早时间戳传递到目标帧。我们进一步引入掩码自注意力来规范常规查询、传播查询和去噪查询之间的交互,同时保持去噪监督与检测查询隔离。此外,一种阶段解耦的锚点嵌入在自注意力之前注入位置,之后注入尺寸、方向和速度,减少了来自时间不一致属性的干扰。在ViT-L骨干网络下,PQR3D在nuScenes测试集上创造了新的最高水平,达到71.6 NDS和64.9 mAP。源代码可在https://github.com/huiyegit/PQR3D获取。
cs.CV / 80 / 2609.32175

OneFixer: High-Quality and Consistent One-Step Autoregressive 3DGS Refinement for Driving Scenes

OneFixer:面向驾驶场景的高质量且一致的单步自回归3DGS细化
Jeon, Boseong, Lee, Junhyeop, Cha, Juhan, Kim, Hayoung
Abstract
Autoregressive video diffusion is a promising render-time fixer for 3D Gaussian Splatting (3DGS) in autonomous-driving simulation, but deployment demands high visual quality and temporal consistency at low latency. This is especially hard for one-step causal generation, where each imperfect prediction immediately becomes context for subsequent frames. Existing approaches stabilize rollouts through staged training with multiple modules and rollout-aware regularization, yet one-step quality still falls short of what deployment requires. We introduce OneFixer, a one-step autoregressive video-diffusion fixer trained in a single task-specific adaptation stage. Our key idea is a deployment-matched shared rollout: the model's own one-step predictions serve as the causal context for flow matching, exposing training to deployment-time errors, while the same rollout receives direct pixel-space perceptual supervision to preserve fine detail. Because the predictions optimized for current-frame quality are exactly those reused as future context, fidelity and autoregressive robustness are learned jointly, without bidirectional-to-causal conversion or teacher-student distillation. OneFixer further exploits cues that driving simulation readily provides, lane geometry and dynamic-agent states, to improve geometric fidelity. On Waymo and proprietary driving scenes with 900-frame rollouts, OneFixer achieves the lowest FVD, LPIPS, and DISTS among all baselines at one step, with temporal consistency matching or exceeding multi-stage DMD pipelines. Under identical backbone and conditioning, it matches a multi-stage DMD-with-Self-Forcing pipeline in under half the GPU-hours and keeps improving beyond its plateau. In closed-loop simulation with a driving policy, OneFixer reduces the collision rate by a third relative to raw 3DGS rendering. Project page: https://onefixer-web.vercel.app/
Chinese Translation
自回归视频扩散是一种有前景的渲染时修复器,用于自动驾驶仿真中的3D高斯泼溅(3DGS),但部署要求在低延迟下实现高视觉质量和时间一致性。这对于单步因果生成尤其困难,因为每个不完美的预测会立即成为后续帧的上下文。现有方法通过多模块的分阶段训练和面向展开(rollout)的正则化来稳定展开过程,但单步质量仍达不到部署要求。我们提出OneFixer,一种在单一任务特定适应阶段训练的单步自回归视频扩散修复器。我们的核心思想是部署匹配的共享展开:模型自身的一步预测作为流匹配的因果上下文,使训练暴露于部署时的误差,同时同一展开接收直接的像素空间感知监督以保留细节。由于为当前帧质量优化的预测恰好被重用为未来上下文,保真度和自回归鲁棒性被联合学习,无需双向到因果的转换或师生蒸馏。OneFixer进一步利用驾驶仿真容易提供的线索,即车道几何和动态智能体状态,以提高几何保真度。在Waymo和专有驾驶场景的900帧展开中,OneFixer在一步生成中实现了所有基线中最低的FVD、LPIPS和DISTS,时间一致性匹配或超过多阶段DMD流水线。在相同骨干网络和条件下,它以不到一半的GPU小时数匹配了多阶段DMD-with-Self-Forcing流水线,并在其平台期后继续改进。在带有驾驶策略的闭环仿真中,OneFixer将碰撞率相对于原始3DGS渲染降低了三分之一。项目页面:https://onefixer-web.vercel.app/
cs.CV / 81 / 2609.32177

Federated 3D Gaussian Splatting for Large-Scale Scene Reconstruction at Wireless Edge

面向无线边缘大规模场景重建的联邦3D高斯泼溅
Wu, Guanlin, Hu, Chao, Chen, Pu, Zhang, Juyong, Hu, Han, Cui, Shuguang, Xu, Jie
Abstract
Three-dimensional (3D) Gaussian splatting (3D-GS) has emerged as a promising technique for large-scale scene reconstruction due to its high rendering efficiency and fidelity. However, the training of large-scale 3D-GS models at wireless edge faces various technical challenges including the limited communication, computation, and graphics processing unit (GPU) memory resources at edge devices, the structural inconsistency issue across local models hindering their effective aggregation, as well as privacy leakage risks associated with raw visual content and camera parameters. To address these challenges, this paper proposes a novel resource-efficient federated learning framework for efficiently training 3D-GS models of large scenes under severe resource constraints. First, we propose an on-device model lightweighting mechanism that adaptively selects and prunes Gaussian points to balance the rendering quality and training efficiency. In this mechanism, we quantitatively evaluate the importance of different Gaussian points at each device to facilitate the pruning, and use a novel importance-to-latency ratio criterion to determine the number of pruned Gaussian points under GPU memory and computation/communication latency constraints. Furthermore, we develop a 3D-GS model recovery mechanism that restores structural consistency across local 3D-GS models without accessing private camera parameters, enabling their effective aggregation towards a global model. Finally, extensive experiments show that our approach significantly accelerates convergence, maintains high rendering quality, and reduces training latency compared to state-of-the-art federated 3D-GS baselines.
Chinese Translation
三维(3D)高斯泼溅(3D-GS)因其高渲染效率和保真度,已成为大规模场景重建的一种有前景的技术。然而,在无线边缘训练大规模3D-GS模型面临各种技术挑战,包括边缘设备上有限的通信、计算和图形处理单元(GPU)内存资源,局部模型之间的结构不一致问题阻碍了它们的有效聚合,以及原始视觉内容和相机参数带来的隐私泄露风险。为了解决这些挑战,本文提出了一种新颖的资源高效的联邦学习框架,用于在严重资源约束下高效训练大场景的3D-GS模型。首先,我们提出了一种设备端模型轻量化机制,自适应地选择和修剪高斯点,以平衡渲染质量和训练效率。在该机制中,我们定量评估每个设备上不同高斯点的重要性以促进修剪,并使用一种新颖的重要性-延迟比准则,在GPU内存和计算/通信延迟约束下确定修剪的高斯点数量。此外,我们开发了一种3D-GS模型恢复机制,无需访问私有相机参数即可恢复局部3D-GS模型之间的结构一致性,使其能够有效地聚合为全局模型。最后,大量实验表明,与最先进的联邦3D-GS基线相比,我们的方法显著加速了收敛,保持了高渲染质量,并减少了训练延迟。
cs.CV / 82 / 2609.32180

Binaural Audio-Visual Instance Segmentation

双耳视听实例分割
Wang, Saijun, Tang, Guanfeng, Zhao, Hongbo, Lei, Zhicheng, Zhang, Yutong, Ye, Wei, Fan, Rui
Abstract
Audio-visual segmentation (AVS) aims to segment sounding objects at the pixel level by integrating auditory and visual cues. However, existing methods are predominantly developed under the monaural setting and primarily rely on cross-modal semantic correspondence, which limits their ability to distinguish visually similar instances of the same semantic class. In contrast, humans naturally exploit binaural hearing, where interaural differences and direction-dependent acoustic filtering introduced by the head and pinnae provide physically grounded spatial cues for accurate sound source localization. Motivated by this observation, we introduce binaural audio-visual instance segmentation (BiAVIS), a new task that leverages synchronized binaural audio and video frames to segment sounding instances. To advance research on this task, we establish two benchmarks by manually annotating an existing binaural audio-visual dataset and collecting a new real-world dataset, BiAVIS-Bench, in more challenging and diverse scenarios. We further propose a BiAVIS model, which leverages an audio-only sound source localization network to learn spatial and semantic priors for sounding instances from binaural audio. A query-level audio-visual fusion strategy is subsequently introduced to inject these informative priors into the instance segmentation decoder. Extensive experiments conducted on the two proposed benchmarks demonstrate the superior performance of the BiAVIS model over previous monaural AVS methods, especially in resolving instance-level intra-class ambiguity. On the more challenging BiAVIS-Bench, the proposed BiAVIS model outperforms the best-performing monaural baselines by 17.22\% in mAP and 7.71\% in FSLA, respectively.
Chinese Translation
视听分割(AVS)旨在通过整合听觉和视觉线索,在像素级别上分割发声物体。然而,现有方法主要是在单耳设置下开发的,并且主要依赖于跨模态语义对应,这限制了它们区分同一语义类别中视觉上相似的实例的能力。相比之下,人类自然地利用双耳听觉,其中由头部和耳廓引入的双耳间差异和方向依赖的声学滤波提供了具有物理依据的空间线索,用于准确的声源定位。受此观察的启发,我们引入了双耳视听实例分割(BiAVIS),这是一个利用同步的双耳音频和视频帧来分割发声实例的新任务。为了推进这项任务的研究,我们通过手动标注现有的双耳视听数据集,并在更具挑战性和多样化的场景中收集一个新的真实世界数据集 BiAVIS-Bench,建立了两个基准。我们进一步提出了一个 BiAVIS 模型,该模型利用仅音频的声源定位网络,从双耳音频中学习发声实例的空间和语义先验。随后引入了一种查询级视听融合策略,将这些信息先验注入到实例分割解码器中。在两个提出的基准上进行的大量实验表明,BiAVIS 模型优于以前的单耳 AVS 方法,特别是在解决实例级类内模糊性方面。在更具挑战性的 BiAVIS-Bench 上,所提出的 BiAVIS 模型在 mAP 和 FSLA 上分别比表现最好的单耳基线高出 17.22% 和 7.71%。
cs.CV / 83 / 2609.32182

KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding

KeyRec:面向流式与长视频理解的有界视觉记忆
Chen, Zihan, Rong, Xuejian, Wang, Xiaojuan, Gong, Boqing, Zicher, Adi, Pritch, Yael, Karnad, Nikhil
Abstract
Vision-language models are increasingly used to understand long videos and continuous streams. However, dense visual tokens accumulate with video duration, making long-context inference prohibitively expensive. Existing training-free visual-token selection methods reduce this cost by retaining informative tokens, but may lose coherent event evidence and fail to distinguish detailed recent observations from long-range history. We propose \textbf{KeyRec}, a training-free framework for constructing bounded visual memory. During query-agnostic writing, KeyRec preserves fine-grained recent observations in a visual cache and organizes historical evidence into a structured event bank. Candidate events are proposed according to their novelty relative to previously stored events and maintained through an online add--merge--evict update. When a question arrives, a text-only router adaptively allocates a fixed readout budget between recent and event memory, without reprocessing historical frames. KeyRec operates on model-facing visual embeddings and supports both modular encoder--projector VLMs and the encoder- and projector-free NEO-ov architecture. Across four streaming and long-video benchmarks and three VLM backbones, KeyRec achieves the best compressed performance in 13 of 15 settings using only 10\% of the dense decoder-facing visual-token budget. It outperforms the strongest compressed baseline by 2.21--18.37 points on real-time questions, achieves the best compressed result in five of six long-video settings, and performs best in every NEO-ov 2B setting.
Chinese Translation
视觉语言模型正越来越多地用于理解长视频和连续流。然而,密集视觉标记会随视频时长而累积,使得长上下文推理的成本高得难以承受。现有的免训练视觉标记选择方法通过保留信息量大的标记来降低这一成本,但可能丢失连贯的事件证据,并且无法区分近期详细观察与长期历史。我们提出 KeyRec,一个用于构建有界视觉记忆的免训练框架。在与查询无关的写入过程中,KeyRec 将细粒度的近期观察保存在视觉缓存中,并将历史证据组织为结构化事件库。候选事件根据其相对于先前存储事件的新颖性被提出,并通过在线添加—合并—淘汰更新进行维护。当问题到来时,一个纯文本路由器在近期记忆和事件记忆之间自适应地分配固定的读取预算,而无需重新处理历史帧。KeyRec 作用于面向模型的视觉嵌入,并同时支持模块化编码器—投影器 VLM 以及无编码器和投影器的 NEO-ov 架构。在四个流式和长视频基准以及三个 VLM 主干上,KeyRec 仅使用 10% 的面向解码器的密集视觉标记预算,就在 15 个设置中的 13 个取得了最佳压缩性能。在实时问题上,它比最强的压缩基线高出 2.21–18.37 个点;在六个长视频设置中的五个取得最佳压缩结果;并且在每个 NEO-ov 2B 设置中都表现最佳。
cs.CV / 84 / 2609.32183

Scalable In-Domain Self-Supervised Foundation Model for Dense Representation Transfer in High-Resolution Plant Imaging

用于高分辨率植物成像中密集表示迁移的可扩展域内自监督基础模型
Guo, Junlin, Majumder, Sharmin, Lyngaas, Isaac, Lagergren, John, Wang, Xiao
Abstract
High-resolution plant imaging enables detailed characterization of plant morphology, but dense scientific analysis remains limited by costly pixel-level annotations, large image pixel dimensions, and substantial variation in imaging conditions. This work proposes a scalable in-domain self-supervised pretrained foundation model for high-resolution, high-pixel-dimension multi-species plant imagery. A masked autoencoder with a ViT backbone is pretrained on more than 10 million multi-view plant image tiles using distributed training. Following scalable pretraining, the learned foundation-model representations are comprehensively benchmarked across fine-grained dense prediction and coarse global feature recognition, with particular emphasis on limited supervision and realistic downstream imaging conditions. This work focuses on the domain gap of existing foundation models in dense feature representation and transfer. Through extensive experiments involving limited annotations, cross-view variation, and resolution degradation, the in-domain FM achieves a Mean Dice of 0.8686 and a Pooled Dice of 0.8959, outperforming an MAE counterpart pretrained on large-scale natural-image data by 0.0694 and 0.0613, respectively. The results further indicate that increasing pretraining scale produces consistent improvements in dense feature transfer. Overall, these findings suggest that scaling in-domain self-supervised pretraining can reduce the domain gap and improve transferable dense representations for high-pixel-dimension scientific imaging.
Chinese Translation
高分辨率植物成像能够详细表征植物形态,但密集的科学分析仍受限于高昂的像素级标注成本、巨大的图像像素尺寸以及成像条件的显著变化。本研究提出了一种可扩展的域内自监督预训练基础模型,用于高分辨率、高像素维度的多物种植物图像。一个带有ViT主干的掩码自编码器在超过1000万个多视角植物图像块上使用分布式训练进行预训练。在可扩展预训练之后,所学到的基础模型表示在细粒度密集预测和粗粒度全局特征识别上进行了全面基准测试,特别强调有限监督和真实的下游成像条件。本研究关注现有基础模型在密集特征表示和迁移中的领域差距。通过涉及有限标注、跨视角变化和分辨率退化的广泛实验,域内FM达到了0.8686的平均Dice和0.8959的合并Dice,分别比在大规模自然图像数据上预训练的MAE对应模型高出0.0694和0.0613。结果进一步表明,增加预训练规模可在密集特征迁移中产生持续改进。总体而言,这些发现表明,扩展域内自监督预训练可以减少领域差距,并改善高像素维度科学成像的可迁移密集表示。
cs.CV / 85 / 2609.32185

Contamination, Prior, or Evidence? Decomposing and Training Evidence Use in Whole-Slide Vision-Language Models

污染、先验,还是证据?在全切片视觉语言模型中分解和训练证据使用
Zhang, Wenhao, Zhou, Zhongliang, Zhang, Shiyuan, Yang, Yiqing, Wang, Pinqiao, Yang, Lehan, Wang, Hanyin, Kang, John, Li, Sheng
Abstract
Pathology vision-language models (VLMs) are conventionally evaluated by accuracy, but accuracy alone does not measure evidence use: it may conflate dataset contamination, prior knowledge, and image evidence. In a motivating study of lymph-node metastasis prediction, we found that most public pathology VLMs showed minimal differences when changing from feeding the models with whole-slide images, an annotated lesion, or no image at all. To better understand the specific features leveraged by these models, this paper presents two contributions aimed at disentangling these factors. First, we present CleanSlide, a TCGA-based VQA benchmark designed to eliminate image- and question-side contamination. It contains 149K audited multiple-choice questions over 9,985 slides, with patient- and tissue-source-disjoint splits. Every question is audited for option shortcuts, stem leakage, cross-split duplication, and blind solvability. Second, we propose Pair-DPO, a preference loss over counterfactual slide pairs from the same question and source. By controlling for shared confounding factors, Pair-DPO cancels out the question-attributable signal and leaves image evidence as the source of preference. Specifically, each pair consists of two real slides with opposite, verified findings, introducing neither editing artifacts nor unverified labels for diffuse or graded features such as invasion, necrosis, and tumor grade. Experiments show that our method gains 15.29% from image evidence on the CleanSlide, compared with 2.81% for the best published model. On the external CPTAC and BCNB cohorts, our method achieves accuracies of 57.6% and 59.0%, outperforming all other evaluated models by 9.7% and 3.4%, respectively. We will release the benchmark and code.
Chinese Translation
病理学视觉语言模型(VLM)通常通过准确率进行评估,但仅凭准确率并不能衡量证据的使用:它可能混淆数据集污染、先验知识和图像证据。在一项关于淋巴结转移预测的动机研究中,我们发现大多数公开的病理学 VLM 在改变模型的输入(全切片图像、注释病变或无图像)时表现出极小的差异。为了更好地理解这些模型所利用的具体特征,本文提出了两项贡献,旨在解开这些因素。首先,我们提出了 CleanSlide,一个基于 TCGA 的 VQA 基准,旨在消除图像和问题侧的污染。它包含 149K 个经审核的多选题,涵盖 9,985 张切片,并具有患者和组织来源不相交的划分。每个问题都针对选项捷径、题干泄漏、跨划分重复和盲解性进行了审核。其次,我们提出了 Pair-DPO,一种对来自同一问题和来源的反事实切片对的偏好损失。通过控制共享的混杂因素,Pair-DPO 抵消了可归因于问题的信号,并将图像证据作为偏好来源。具体来说,每一对由两张真实切片组成,具有相反的、经核实的发现,既不会引入编辑伪影,也不会为弥漫性或分级特征(如侵袭、坏死和肿瘤分级)引入未经验证的标签。实验表明,我们的方法在 CleanSlide 上从图像证据中获得了 15.29% 的提升,而最佳已发表模型仅为 2.81%。在外部 CPTAC 和 BCNB 队列上,我们的方法分别达到了 57.6% 和 59.0% 的准确率,分别比其他所有评估模型高出 9.7% 和 3.4%。我们将发布基准和代码。
cs.CV / 86 / 2609.32188

Presence Is Not Faithfulness: Figurative Vehicle Intrusion in Text-to-Image Generation

存在不等于忠实:文本到图像生成中的喻体侵入
Ma, Xiaoyu, Yang, Chen, Chen, Hao
Abstract
Text-to-image (TTI) models increasingly generate high-quality images from natural-language prompts, yet figurative language exposes a failure: a vehicle that should guide the depiction of a tenor may instead be rendered as a visible object. We call this failure Figurative Vehicle Intrusion: the intruding content is textually licensed, but it is assigned the wrong visual role, showing that visual presence is not always faithfulness and that presence-oriented evaluation can miss such errors. To study it systematically, we introduce Vehicle Intrusion and Semantic Tenor Assessment (VISTA), a multilingual benchmark of figurative prompts organized by Figurative Form and Mapping Mechanism. We further propose V-Score, a diagnostic question-answering metric that evaluates role-aware figurative faithfulness in generated images. Evaluations on recent high-performing TTI models show that vehicle intrusion persists across languages and figurative categories. As a lightweight mitigation, we introduce VISTA-Guard, which partially reduces vehicle intrusion and suggests a practical path toward more figuratively faithful TTI generation. All resources will be released publicly.
Chinese Translation
文本到图像(TTI)模型越来越多地从自然语言提示生成高质量图像,然而比喻性语言暴露了一个失败:本应引导本体描绘的喻体,却可能被渲染为可见物体。我们称这种失败为“喻体侵入”:侵入的内容在文本上是许可的,但被赋予了错误的视觉角色,这表明视觉存在并不总是忠实,且以存在为导向的评估可能会遗漏此类错误。为了系统地研究它,我们引入了 VISTA(Vehicle Intrusion and Semantic Tenor Assessment,喻体侵入与语义本体评估),一个按比喻形式和映射机制组织的多语言比喻性提示基准。我们进一步提出了 V-Score,一种诊断性问答度量,用于评估生成图像中角色感知的比喻忠实性。对近期高性能 TTI 模型的评估表明,喻体侵入在不同语言和比喻类别中持续存在。作为一种轻量级缓解方法,我们引入了 VISTA-Guard,它部分减少了喻体侵入,并为更忠实于比喻的 TTI 生成指明了一条实用路径。所有资源将公开发布。
cs.CV / 87 / 2609.32190

Evaluating Single and Multi-Omics Based Explainable Artificial Intelligence (MOXAI) for Molecular Subclass Classification of Adult-Type Diffuse Gliomas

评估基于单组学与多组学的可解释人工智能(MOXAI)用于成人型弥漫性胶质瘤的分子亚类分类
Alom, Md Zahangir, Tran, Quynh T., Alexandar, Breuer, Orr, Brent A.
Abstract
DNA methylation (DNAM) profiling has emerged as a powerful diagnostic tool for classifying brain and solid tumors. However, existing computational models typically analyze methylation and copy number variation (CNV) data separately, failing to capture the complementary information their integration could provide. Moreover, current classification models lack mechanisms for within-class risk assessment analogous to traditional tumor grading, and no established explainability method can attribute classification decisions to specific genomic loci. In this paper, we present MOXAI (Multi-Omics Based Explainable AI), a deep learning framework that integrates DNA methylation and copy number data from methylation arrays to classify molecular subtypes of adult-type diffuse gliomas, alongside single-modality variants for comparison. Using a cohort from The Cancer Genome Atlas (TCGA), we trained ResNet50, DINOv2, and Graph Attention Network (GAT) models on methylation data alone, copy number data alone, and combined multimodal data. We further developed explainable AI (XAI) methods based on class activation maps (CAMs) and gradient-weighted CAM (Grad-CAM) to identify the specific CpG sites, genes, and chromosomal regions most relevant to each classification decision. The multimodal model achieved up to 92.98% cross-validation accuracy, outperforming models trained on CNV data alone. DINOv2 showed the strongest generalization, reaching 94.25% accuracy (confidence >0.9) on independent validation sets. XAI results aligned with established molecular features of adult-type diffuse glioma subtypes, confirming the biological interpretability of the framework.
Chinese Translation
DNA甲基化(DNAM)谱分析已成为对脑肿瘤和实体瘤进行分类的强大诊断工具。然而,现有计算模型通常分别分析甲基化和拷贝数变异(CNV)数据,未能捕捉整合二者所能提供的互补信息。此外,当前分类模型缺乏类似于传统肿瘤分级的类内风险评估机制,并且尚无已建立的可解释性方法能够将分类决策归因于特定基因组位点。在本文中,我们提出MOXAI(Multi-Omics Based Explainable AI,基于多组学的可解释AI),一个深度学习框架,它整合来自甲基化芯片的DNA甲基化和拷贝数数据,以对成人型弥漫性胶质瘤的分子亚型进行分类,并同时提供单模态变体用于比较。使用来自癌症基因组图谱(TCGA)的队列,我们在仅甲基化数据、仅拷贝数数据以及联合多模态数据上训练了ResNet50、DINOv2和图注意力网络(GAT)模型。我们进一步开发了基于类激活图(CAM)和梯度加权类激活图(Grad-CAM)的可解释人工智能(XAI)方法,以识别与每个分类决策最相关的特定CpG位点、基因和染色体区域。多模态模型达到了高达92.98%的交叉验证准确率,优于仅在CNV数据上训练的模型。DINOv2表现出最强的泛化能力,在独立验证集上达到94.25%的准确率(置信度>0.9)。XAI结果与成人型弥漫性胶质瘤亚型的已知分子特征一致,证实了该框架的生物学可解释性。
cs.CV / 88 / 2609.32193

Devol-ONE: One Autoregressive Mixture of Transformers to Unify Vision-Language-Action and Latent World Modeling

Devol-ONE:一个自回归混合 Transformer 统一视觉-语言-动作与潜在世界建模
Cai, Hongyi, Ong, Yi Herng, Wu, Tingshiuan C., Hui, Lim Chiew, Li, Hanxia, Guo, Kehong, Cheong, Sze Yuan
Abstract
Vision Language Action (VLA) models condition actions directly on current visual and language context, without an explicit account of how the scene evolves under candidate actions. World Action Models (WAM) attempt to address this limitation by predicting future states, but existing designs keep prediction and policy learning architecturally separate, connecting them only through the predicted output, whether through pixel space video generation or a latent forecasting module trained independently of the policy. We present Devol-ONE, a Mixture of Transformers architecture that unifies vision language understanding, latent world dynamics prediction, and action generation within a single autoregressive framework. Instead of encoding vision language tokens once and feeding them to the action expert, Devol-ONE runs autoregressive prediction jointly across a vision language stream and a V-JEPA pretrained dynamics stream, attending to the vision language key-value cache at every layer to forecast future latent states under language guidance. The action expert is in turn shaped continuously by semantic reasoning and predicted physical dynamics rather than by a fixed representation computed in advance. Extensive experiments are conducted on LIBERO, LIBERO-PLUS, RoboTwin2.0 along with real-world evaluation on Flexiv single-arm and dual-arm setups. Ablation studies show the effectiveness of dynamic stream prediction and layer-wise unified attention to validate our model architectural coherency.
Chinese Translation
视觉-语言-动作(VLA)模型直接基于当前视觉和语言上下文来生成动作,而没有明确考虑场景在候选动作下如何演化。世界动作模型(WAM)试图通过预测未来状态来解决这一局限,但现有设计在架构上将预测与策略学习分开,仅通过预测输出将它们连接起来,无论是通过像素空间视频生成还是独立于策略训练的潜在预测模块。我们提出 Devol-ONE,一种混合 Transformer 架构,在单一自回归框架内统一了视觉语言理解、潜在世界动态预测和动作生成。Devol-ONE 不是仅编码一次视觉语言 token 并将其馈送给动作专家,而是在视觉语言流和 V-JEPA 预训练动态流上联合运行自回归预测,在每一层关注视觉语言键值缓存,以在语言指导下预测未来潜在状态。动作专家反过来由语义推理和预测的物理动态持续塑造,而不是由预先计算的固定表示塑造。在 LIBERO、LIBERO-PLUS、RoboTwin2.0 上进行了大量实验,并在 Flexiv 单臂和双臂设置上进行了真实世界评估。消融研究显示了动态流预测和逐层统一注意力的有效性,验证了我们模型架构的一致性。
cs.CV / 89 / 2609.32203

Kernel-Based Steering of CLIP with Vision-Language Model Preferences

基于核的视觉语言模型偏好引导CLIP
Ghiasvand, Sajjad, Oskouie, Haniyeh Ehsani, Mansouri, Sina, Alizadeh, Mahnoosh, Farnia, Farzan, Pedarsani, Ramtin
Abstract
Large vision-language models (VLMs) can judge visual similarity, but their judgments are not directly available as compact image embeddings for efficient comparison. We study how to transfer these preferences into CLIP while retaining its image--text capabilities. We introduce ASK, a kernel-based steering method that learns from elicited pairwise judgments without accessing teacher embeddings or collecting new human similarity annotations. ASK constructs positive semidefinite target kernels within small image groups and combines visual kernel matching with an image--text distributional anchor. Low-rank adapters jointly update the visual and text encoders while regularizing predictions toward frozen CLIP. After adaptation, retrieval uses CLIP image embeddings and cosine similarity, with no VLM calls. Experiments across five image domains, four CLIP backbones, and six judges evaluate teacher agreement, retrieval, and recognition retention. For ViT-B/16, mean retrieval mAP on classes excluded from adaptation increases from 53.8 to 75.0, compared with 71.7 for DINOv2 targets with KL anchoring. Mean zero-shot accuracy with jointly adapted encoders increases from 61.8\% to 62.4\%, averaged over 12 benchmarks and the five adaptation domains. Prompting provides an additional capability: selecting which visual distinctions the student learns. Human-annotated evaluations across four datasets support this criterion-specific control.
Chinese Translation
大型视觉语言模型(VLMs)能够判断视觉相似性,但其判断无法直接作为紧凑的图像嵌入用于高效比较。我们研究如何将这些偏好迁移到CLIP中,同时保留其图像-文本能力。我们提出ASK,一种基于核的引导方法,它从引出的成对判断中学习,无需访问教师嵌入或收集新的人类相似性标注。ASK在小图像组内构造半正定目标核,并将视觉核匹配与图像-文本分布锚点相结合。低秩适配器联合更新视觉和文本编码器,同时将预测正则化到冻结的CLIP。适应后,检索使用CLIP图像嵌入和余弦相似度,无需调用VLM。在五个图像领域、四个CLIP骨干网络和六个评判者上的实验评估了教师一致性、检索和识别保持。对于ViT-B/16,在适应中排除的类别上,平均检索mAP从53.8增加到75.0,而使用KL锚定的DINOv2目标为71.7。联合适应编码器的平均零样本准确率从61.8%增加到62.4%,在12个基准和五个适应领域上平均。提示提供了一项额外能力:选择学生模型学习哪些视觉区分。四个数据集上的人工标注评估支持这种特定标准控制。
cs.CV / 90 / 2609.32222

Geometry-Preserving Blind Watermarking for Raw 3D Point Clouds

面向原始3D点云的几何保持盲水印
Zhou, Rungui, Zhou, Chuanzhi, Wang, Ruihuan, Wang, Peng-Shuai
Abstract
Raw 3D point clouds are a core geometric representation. Establishing their ownership is challenging because point sets are irregular, unstructured, and frequently altered by resampling and geometric preprocessing. We present a blind watermarking framework that operates directly on xyz coordinates and supports both object-level shapes and scene-scale scans. At verification time, the embedded message is recovered from the observed point cloud alone, without access to the original point cloud, color, normals, or mesh connectivity. The method jointly learns watermark embedding and extraction through a feed-forward octree-based architecture, enabling efficient multi-scale geometric reasoning on large point sets. During training, a stochastic transformation layer exposes the decoder to common geometric perturbations, while progressive pose alignment improves robustness to pose changes. Experiments on object-level and scene-level benchmarks demonstrate reliable message recovery under common geometric processing while maintaining low geometric distortion. Qualitative comparisons further show that the learned perturbations are less visually conspicuous and less spatially structured than those of handcrafted alternatives.
Chinese Translation
原始3D点云是一种核心的几何表示。确立其所有权具有挑战性,因为点集是不规则、无结构的,并且经常因重采样和几何预处理而发生改变。我们提出一种盲水印框架,直接作用于xyz坐标,并同时支持物体级形状和场景级扫描。在验证阶段,仅从观测到的点云中即可恢复嵌入的消息,而无需访问原始点云、颜色、法线或网格连通性。该方法通过基于前馈八叉树的架构联合学习水印嵌入和提取,从而能够在大型点集上进行高效的多尺度几何推理。在训练过程中,随机变换层使解码器暴露于常见几何扰动,而渐进式姿态对齐则提升了对姿态变化的鲁棒性。在物体级和场景级基准上的实验表明,在常见几何处理下能够可靠地恢复消息,同时保持较低的几何失真。定性比较进一步表明,与手工设计的替代方案相比,学习到的扰动在视觉上更不显眼,且空间结构化程度更低。
cs.CV / 91 / 2609.32231

Skeletons in Flow: Graph Structured Flow Matching for Human Motion Prediction

流中的骨架:用于人体运动预测的图结构流匹配
Wang, Yixuan, Fallin, Brandon C., Dixon, Warren E.
Abstract
Human motion prediction requires diverse future trajectories that remain consistent with observed motion and the articulated physical structure of the body. Skeletal constraints restrict individual poses, while coordinated motion depends on spatial interactions (between connected joints) and temporal interactions (between time instants). To facilitate human motion prediction in light of these constraints and interactions, we introduce Graph Structured Flow Matching (GSFM), which transports the complete future skeletal trajectory through a single conditional velocity field. The trajectory produces a spatiotemporal skeleton graph, and spatial and temporal attention couple its evolution according to skeletal relations and physical time offsets. Bone directions lie on unit spheres relative to a root joint, and tangent evolution preserves input bone lengths throughout generation. We train a learned velocity field through conditional flow matching along geodesic paths connecting random trajectories centered on the last-observed pose to recorded future trajectories. Experiments on the Archive of Motion capture As Surface Shapes (AMASS) dataset evaluate prediction accuracy, diversity calibration, and motion statistics. We demonstrate the contributions of spatial and temporal message passing in the developed architecture through an ablation study. GSFM models trained on AMASS also perform competitively on the Human3.6M skeleton without parameter updates or retraining, demonstrating applicability to an unseen skeletal structure.
Chinese Translation
人体运动预测需要多样化的未来轨迹,这些轨迹需与观测到的运动以及身体的关节物理结构保持一致。骨骼约束限制单个姿态,而协调运动依赖于空间交互(连接的关节之间)和时间交互(时间点之间)。为了在这些约束和交互下促进人体运动预测,我们引入了图结构流匹配(GSFM),它通过单个条件速度场传输完整的未来骨骼轨迹。该轨迹生成一个时空骨架图,空间和时间注意力根据骨骼关系和物理时间偏移耦合其演化。相对于根关节,骨骼方向位于单位球面上,切线演化在整个生成过程中保持输入骨骼长度。我们通过条件流匹配训练一个学习到的速度场,沿着连接以最后观测姿态为中心的随机轨迹与记录的未来轨迹的测地路径。在AMASS(Archive of Motion capture As Surface Shapes)数据集上的实验评估了预测准确性、多样性校准和运动统计。我们通过消融研究展示了所开发架构中空间和时间消息传递的贡献。在AMASS上训练的GSFM模型在Human3.6M骨骼上无需参数更新或重新训练也能有竞争力的表现,证明了其对未见骨骼结构的适用性。
cs.CV / 92 / 2609.32244

Two-Stage Multi-View Gait Recognition with a Re-Embedding Network

基于重嵌入网络的两阶段多视角步态识别
Le, Long Hoang, Ngo, Trung Thanh
Abstract
Gait recognition always remains challenging due to severe overfitting and the rigid view constraints common in single-stage approaches. We propose a two-stage framework, termed Translate-First-Then-Reason (TFTR), to address these issues. In the first stage, a shallow Siamese convolutional network with triplet loss maps Gait Energy Images (GEIs) into a 128-dimensional view-specific embedding space. In the second stage, these per-view embeddings are treated as tokens and processed by a 12-layer Transformer encoder, which re-projects them into a new space with improved cosine separability. This design enables flexible fusion of an arbitrary number of views at inference, overcoming the fixed-input limitations of prior methods. Trained on the OU-MVLP dataset (6,000 subjects) and evaluated on unseen CASIA-B across normal, bag-carrying, and coat-wearing conditions, our pipeline achieves 96.91\% single-view and 99.49\% three-view accuracy on OU-MVLP, and attains 100\% accuracy on CASIA-B with three views.
Chinese Translation
步态识别始终面临严峻挑战,原因在于单阶段方法中常见的严重过拟合和严格的视角约束。我们提出了一种称为“先翻译后推理”(TFTR)的两阶段框架来解决这些问题。在第一阶段,一个带有三元组损失的浅层孪生卷积网络将步态能量图(GEI)映射到128维的视角特定嵌入空间。在第二阶段,这些逐视角的嵌入被当作token,并由一个12层的Transformer编码器处理,将其重新投影到一个具有更好余弦可分性的新空间。这种设计能够在推理时灵活融合任意数量的视角,克服了先前方法固定输入的限制。在OU-MVLP数据集(6000个受试者)上训练,并在未见过的CASIA-B数据集上跨正常、背包和穿大衣条件进行评估,我们的流程在OU-MVLP上达到了96.91%的单视角和99.49%的三视角准确率,并在CASIA-B的三视角上达到了100%的准确率。
cs.CV / 93 / 2609.32250

RoboSTAR: Next-Scale Autoregressive Sign Language Translation for Humanoid Robots

RoboSTAR:面向人形机器人的下一尺度自回归手语翻译
Zeng, Yujia, Peng, Chensheng, Chen, Yuxin, Shao, Alex, Jew, Nathan, Tomizuka, Masayoshi
Abstract
Sign-language interpretation in public communication relies on qualified professional interpreters and can be difficult to scale, motivating robotic signing as a complementary accessibility interface. We present RoBoSTAR, a text-conditioned sign language production (SLP) framework for generating human-centric sign motion that can be retargeted for robotic execution, with speech supported optionally through an external ASR front end. Conventional autoregressive approaches flatten motion into a single full-resolution token sequence, forcing long-range and local dependencies to be modeled at a uniform temporal granularity. RoBoSTAR instead combines part-wise Finite Scalar Quantization with next-scale autoregression, generating motion over progressively finer temporal resolutions while predicting synchronized body and hand tokens in parallel within each step. This coarse-to-fine formulation provides compact long-range context before progressively refining motion details, while self-conditioning and context corruption improve robustness to cross-scale prediction errors. The generated motion is subsequently retargeted for physical humanoid execution. Extensive qualitative and quantitative evaluations are conducted to demonstrate the effectiveness of RoBoSTAR.
Chinese Translation
公共交流中的手语翻译依赖于合格的专业口译员,且难以规模化,这促使机器人手语成为一种补充性的无障碍交互界面。我们提出了RoBoSTAR,一个文本条件的手语生成(SLP)框架,用于生成以人为中心的手语动作,该动作可重定向用于机器人执行,并可选通过外部ASR前端支持语音。传统的自回归方法将动作展平为单一的全分辨率标记序列,迫使长程和局部依赖在统一的时间粒度上建模。RoBoSTAR则结合了分部位有限标量量化与下一尺度自回归,在逐步更精细的时间分辨率上生成动作,同时在每个步骤中并行预测同步的身体和手部标记。这种由粗到细的表述在逐步细化动作细节之前提供了紧凑的长程上下文,而自条件和上下文损坏提高了对跨尺度预测误差的鲁棒性。生成的动作随后被重定向用于物理人形机器人的执行。我们进行了广泛的定性和定量评估,以证明RoBoSTAR的有效性。
cs.CV / 94 / 2609.32273

FSS-UBrain: Multi-region Few-Shot Brain Tumor MRI Segmentation

FSS-UBrain:多区域小样本脑肿瘤MRI分割
Vu, Truong Viet, Nguyen, Nguyen Phuc, Hang, Dang Thi Thu, Thanh, Tran Thien, Bao, Vo Nguyen Quoc, Anh, Nguyen Thai, Tu, Ngo Hoang
Abstract
Accurate delineation of whole tumor (WT), tumor core (TC), and enhancing tumor (ET) from multimodal magnetic resonance imaging remains challenging under limited annotation, cross-cohort variation, and severe target sparsity. We propose FSS-UBrain, a region-wise one-shot segmentation framework that uses a labeled positive support slice to condition binary query segmentation separately for WT, TC, and ET. Support-derived foreground and background descriptors guide query-feature adaptation, bottleneck interaction, decoder-side reconstruction, and boundary refinement. Episodic training additionally incorporates hard-negative and fully negative queries with empty-query regularization to suppress spurious foreground activation when the selected region is absent. Although inference operates on two-dimensional support--query slice pairs, checkpoint selection, threshold calibration, and final evaluation are performed after volumetric reconstruction. FSS-UBrain is evaluated on a held-out BraTS 2020 split and under target-supported cross-cohort protocols on BraTS 2023 and BraTS-Africa. Cases used as target support are excluded from the query cohorts, and no target-domain fine-tuning or test-time parameter updates are performed. On BraTS 2020, FSS-UBrain achieves volumetric Dice scores of 89.82%, 82.14%, and 77.42% for WT, TC, and ET, respectively, with corresponding 95th-percentile Hausdorff distance (HD95) values of 11.12, 9.01, and 4.46 mm. It also achieves the highest mean Dice and lowest finite-pair mean HD95 point estimates on BraTS 2023 and BraTS-Africa among the compared few-shot methods. These findings support target-conditioned few-shot segmentation while highlighting sensitivity to support selection and cohort-specific variation.
Chinese Translation
在标注有限、跨队列变异以及严重的目标稀疏性条件下,从多模态磁共振成像中准确勾画全肿瘤(WT)、肿瘤核心(TC)和增强肿瘤(ET)仍然具有挑战性。我们提出了FSS-UBrain,一种区域级的单样本分割框架,它使用标记的正支持切片来分别为WT、TC和ET调节二值查询分割。支持集导出的前景和背景描述符指导查询特征自适应、瓶颈交互、解码器侧重构和边界细化。情景训练还结合了难负样本和全负样本查询,并采用空查询正则化,以在所选区域缺失时抑制虚假的前景激活。尽管推理在二维支持-查询切片对上操作,但检查点选择、阈值校准和最终评估是在体积重建后进行的。FSS-UBrain在保留的BraTS 2020划分上以及在BraTS 2023和BraTS-Africa上目标支持的跨队列协议下进行了评估。用作目标支持的案例被排除在查询队列之外,并且不进行目标域微调或测试时参数更新。在BraTS 2020上,FSS-UBrain对WT、TC和ET的体积Dice分数分别为89.82%、82.14%和77.42%,相应的95百分位Hausdorff距离(HD95)值分别为11.12、9.01和4.46 mm。在比较的小样本方法中,它在BraTS 2023和BraTS-Africa上也实现了最高的平均Dice和最低的有限对平均HD95点估计。这些发现支持目标条件的小样本分割,同时强调了对支持选择的和队列特定变异的敏感性。
cs.CV / 95 / 2609.32280

HeroFrame-Bench: Reference-Anchored Evaluation via Rubric--Ranking Co-Evolution for Movie Hero Frame Selection

HeroFrame-Bench:面向电影主角帧选择的基于评分标准-排名协同演化的参考锚定评估
Kang, Weitai, Deilamsalehy, Hanieh, Xu, Yumo, Sultania, Dewang, Cellat, Serdar, Yan, Yan
Abstract
Hero frames are in-film stills used as source imagery for theatrical posters, streaming cover art, film database listings, and other promotional placements. As the first visual entry point, they shape audiences' initial impressions of the movie and their subsequent willingness to watch it. Selecting these frames, a task we term hero frame selection, requires balancing content relevance with aesthetic appeal. A related task is keyframe selection, yet its benchmarks prioritize relevance over aesthetics, using either finite annotations that exclude valid alternatives or VideoQA that entangles selection quality with downstream model capability. We therefore introduce HeroFrame-Bench, built through a scalable VLM-as-a-Judge framework. We construct multimodal contexts from diverse metadata to ground a VLM judge that scores selected frames using our Reference-anchored Percentile. The percentile is obtained by inserting each frame into reusable, pre-ranked reference chains, enabling direct, extensible, and reliable evaluation. To reduce ambiguity and improve consistency in these subjective judgements, we further propose Rubric-Ranking Co-Evolution, which generates movie-specific rubrics to condition the VLM judge and refines rubrics jointly with the resulting rankings. Within this process, we introduce several verifiable signals, most notably the Inverted Rubric Attack, to select robust rubrics. Finally, HeroFrame-Bench is instantiated over 204 movies with 2,031 reference chains and 1,970 learned rubrics. We build an annotation interface for human-alignment studies which show that our construction design improves VLM agreement with human from 77.56% to 83.78%. Evaluation on multiple methods show that hero frame selection remains challenging.
Chinese Translation
主角帧是电影中的剧照,用作影院海报、流媒体封面艺术、电影数据库列表和其他宣传位置的源图像。作为第一个视觉入口,它们塑造了观众对电影的初步印象及其后续观看意愿。选择这些帧,我们称之为主角帧选择的任务,需要在内容相关性和美学吸引力之间取得平衡。一个相关任务是关键帧选择,然而其基准测试优先考虑相关性而非美学,使用要么排除有效替代方案的有限标注,要么将选择质量与下游模型能力纠缠在一起的VideoQA。因此,我们引入了HeroFrame-Bench,它通过可扩展的VLM-as-a-Judge框架构建。我们从多样化的元数据构建多模态上下文,为一个VLM评判器提供依据,该评判器使用我们的参考锚定百分位数对所选帧进行评分。该百分位数通过将每个帧插入可重用的、预排序的参考链中获得,从而实现直接、可扩展和可靠的评估。为了减少这些主观判断中的模糊性并提高一致性,我们进一步提出了评分标准-排名协同演化,它生成电影特定的评分标准来调节VLM评判器,并与生成的排名一起细化评分标准。在此过程中,我们引入了几个可验证的信号,最值得注意的是反向评分标准攻击,以选择稳健的评分标准。最后,HeroFrame-Bench在204部电影上实例化,包含2,031条参考链和1,970个学习到的评分标准。我们构建了一个用于人类对齐研究的标注界面,研究表明我们的构建设计将VLM与人类的一致性从77.56%提高到83.78%。对多种方法的评估表明,主角帧选择仍然具有挑战性。
cs.CV / 96 / 2609.32316

One Perception, All Maneuvers: Directional Traffic Signal Understanding for Maneuver-Level Signal Intent Prediction

一次感知,所有机动:面向机动级信号意图预测的方向性交通信号理解
Zou, Ang, Zheng, Runzhe, li, Zhigang, Yang, Zhen, Xia, Han, Li, Xuewei, Qin, Zequn, Li, Xi
Abstract
Traffic lights are a key regulatory signal for autonomous driving at urban intersections, yet existing traffic signal perception is still predominantly formulated as instance-level detection or color recognition. Such formulations identify where traffic lights are and what colors they display, but leave a critical semantic gap before downstream planning: which ego maneuver is controlled by each visible signal and what dynamic permission the signal expresses for that maneuver. In this paper, we formulate Directional Traffic Signal Understanding, a decision-oriented task that predicts structured signal states for straight, left-turn, right-turn, and U-turn maneuvers from a front-view image. Each state contains the associated signal color and signal-implied passability. Based on OpenLane-V2, we provide a direction-level benchmark with maneuver-level supervision and metrics for color recognition, passability, full-frame consistency, and safety-critical errors. A direction-aware baseline combines global context, localized traffic-light evidence, and maneuver-specific representations. Experiments show that direction-level modeling improves passability prediction over image-level classifiers and detection-oriented pipelines, particularly at complex multi-signal intersections. The resulting representation provides a direct and interpretable traffic-signal interface for downstream planning together with topology, route, and surrounding-agent information.
Chinese Translation
红绿灯是城市交叉路口自动驾驶的关键监管信号,然而现有的交通信号感知仍然主要被表述为实例级检测或颜色识别。这种表述识别了交通灯的位置和显示的颜色,但在下游规划之前留下了关键的语义空白:每个可见信号控制哪个自车机动,以及该信号对该机动表达何种动态许可。在本文中,我们提出了方向性交通信号理解,这是一个面向决策的任务,从前视图图像中预测直行、左转、右转和掉头机动的结构化信号状态。每个状态包含相关的信号颜色和信号隐含的通行性。基于OpenLane-V2,我们提供了一个方向级基准,具有机动级监督和针对颜色识别、通行性、全帧一致性和安全关键错误的指标。一个方向感知基线结合了全局上下文、局部交通灯证据和机动特定表示。实验表明,方向级建模在通行性预测方面优于图像级分类器和面向检测的流程,特别是在复杂的多信号交叉口。由此产生的表示为下游规划提供了一个直接且可解释的交通信号接口,并与拓扑、路线和周围智能体信息一起。
cs.CV / 97 / 2609.32323

FoundDSR: A Generalizable Foundation Model with Guided 2D Gaussian Splatting for Depth Super-Resolution

FoundDSR:一种采用引导式2D高斯泼溅的通用基础模型,用于深度超分辨率
Wang, Zhengxue, Yan, Zhiqiang, Wu, Yuan, Gao, Guangwei, Li, Xiang, Yang, Jian
Abstract
We introduce FoundDSR, a generalizable foundation model for robust depth reconstruction across unseen data distributions using RGB-D pairs. FoundDSR begins with a guided 2D Gaussian Splatting strategy to model depth representations with Gaussian primitives. This strategy employs high-resolution RGB as prompts to optimize the Gaussian parameters, thereby encouraging each Gaussian primitive to anisotropically deform along high-frequency structural directions. The resulting Gaussian-upsampled representations are then mapped to high-resolution depth through an effective depth reconstruction branch. Furthermore, to mitigate training instability and bias toward dominant sources caused by distribution gaps in large-scale heterogeneous data, we introduce heterogeneous federated learning that allocates each data source to an independent client for local optimization and global aggregation. This design effectively endows FoundDSR with stable scalability to diverse and large-scale training data. Extensive zero-shot evaluations on synthetic, real-world, arbitrary-scale, and noisy conditions demonstrate that FoundDSR consistently outperforms existing state-of-the-art approaches, confirming its strong robustness and generalization to unknown scenes.
Chinese Translation
我们介绍了FoundDSR,一种可泛化的基础模型,用于跨未见数据分布使用RGB-D对进行稳健的深度重建。FoundDSR首先采用引导式2D高斯泼溅(Guided 2D Gaussian Splatting)策略,使用高斯基元对深度表示进行建模。该策略利用高分辨率RGB作为提示来优化高斯参数,从而鼓励每个高斯基元沿高频结构方向进行各向异性变形。然后,通过有效的深度重建分支,将得到的高斯上采样表示映射到高分辨率深度。此外,为了缓解由大规模异构数据中的分布差距引起的训练不稳定性和对主导源的偏差,我们引入了异构联邦学习,将每个数据源分配给独立的客户端进行本地优化和全局聚合。这种设计有效地赋予了FoundDSR对多样化和大规模训练数据的稳定可扩展性。在合成、真实世界、任意尺度和噪声条件下的广泛零样本评估表明,FoundDSR始终优于现有的最先进方法,证实了其对未知场景的强大鲁棒性和泛化能力。
cs.CV / 98 / 2609.32333

Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs

多模态大语言模型中区域到全局迁移的渐进视图同策略蒸馏
Huang, Shanfeng, Fang, Zhou, Xiao, Song, Du, Hai
Abstract
Regional-to-global distillation uses crop-conditioned guidance to improve full-image understanding. The challenge is to effectively transfer the teacher's crop-based advantage to the student's full-image inference. We propose progressive-view on-policy distillation (PVD), which shifts the student's view distribution from the crop toward the full image through an intermediate aspect-preserving padded crop. The padded crop preserves regional content while matching the full image's visual-token grid. Across stages, the view mixture assigns increasing probability to the full image. A lightweight regional-advantage weighting reallocates token-level supervision using the crop-conditioned teacher-student log-probability gap. Evaluated under each sampled input, it applies mild reweighting when the gap is small and emphasizes higher-gap tokens when the gap widens. A Jensen-Shannon metric decomposition interprets this schedule as a shift from matched-input imitation toward the deployment objective. Across benchmarks spanning perception, visual mathematics and general multimodal question answering, PVD-full reaches an average accuracy of 77.51 over three seeds, improving on the reward-free distillation baseline by 2.01 points and on its reward-matched variant by 1.00 point. In the reward-free setting, PVD-distill still gains 1.16 points.
Chinese Translation
区域到全局蒸馏利用裁剪条件引导来提升全图理解。挑战在于有效地将教师基于裁剪的优势迁移到学生的全图推理中。我们提出渐进视图同策略蒸馏(PVD),它通过一个中间的长宽比保持的填充裁剪,将学生的视图分布从裁剪逐步转向全图。填充裁剪在保留区域内容的同时,匹配全图的视觉 token 网格。在不同阶段,视图混合为全图分配递增的概率。一种轻量级的区域优势加权利用裁剪条件下的师生对数概率差,重新分配 token 级监督。在每个采样输入下评估时,当差距较小时施加温和的重加权,当差距扩大时强调较高差距的 token。一个 Jensen-Shannon 度量分解将这种调度解释为从匹配输入模仿向部署目标的转变。在涵盖感知、视觉数学和通用多模态问答的基准上,PVD-full 在三个随机种子上达到 77.51 的平均准确率,比无奖励蒸馏基线提高 2.01 分,比其奖励匹配变体提高 1.00 分。在无奖励设置下,PVD-distill 仍获得 1.16 分的提升。
cs.CV / 99 / 2609.32340

SGA-Flow-GRPO: Spatial Gradient-Guided Credit Assignment for Flow-GRPO

SGA-Flow-GRPO: 面向Flow-GRPO的空间梯度引导信用分配
Yang, Yunkai, Zhang, Yudong, Chen, Xinying, Luo, Bin, Lyu, Jienan, Zhang, Kunquan, Wan, Weitao, Dong, Runmin
Abstract
Reinforcement Learning (RL) has proven effective in aligning flow-based generative models with human preferences. Recently, Flow-GRPO has emerged as an efficient critic-free paradigm by calculating advantages over sampled candidate trajectories. However, standard Flow-GRPO applies a uniform scalar advantage across both temporal denoising steps and spatial latent dimensions, without explicitly accounting for the spatial structure of generated images, which may lead to sub-optimal policy updates. To address this, we propose a novel gradient-guided spatial credit assignment framework tailored for Diffusion Transformers (DiTs). We first reformulate the transition-level log-likelihood in Flow-GRPO into a token-wise representation natively aligned with DiT patch architectures, constructing spatially fine-grained importance sampling ratios. To allocate localized credit without rigid, boundary-sensitive segmentation heuristics, we introduce a continuous spatial credit map derived from reward gradients. Crucially, we employ an outlier-robust normalization scheme based on Median Absolute Deviation (MAD) coupled with temperature scaling, effectively eliminating gradient noise while highlighting functional prompt-aligned regions. Extensive evaluations on GenEval show that our approach delivers SOTA alignment quality, achieving a convergence rate comparable to top-tier methods like DiffusionNFT while substantially improving upon Flow-GRPO-based methods in alignment performance.
Chinese Translation
强化学习(RL)已被证明能有效使基于流的生成模型与人类偏好对齐。最近,Flow-GRPO作为一种高效的无评论家范式出现,通过计算采样候选轨迹上的优势。然而,标准Flow-GRPO在时间去噪步骤和空间潜在维度上应用统一的标量优势,而没有明确考虑生成图像的空间结构,这可能导致次优的策略更新。为了解决这个问题,我们提出了一种新颖的梯度引导空间信用分配框架,专为扩散Transformer(DiT)量身定制。我们首先将Flow-GRPO中的转移级对数似然重新表述为与DiT patch架构原生对齐的token-wise表示,构建空间细粒度的重要性采样比率。为了在没有刚性、边界敏感的分割启发式的情况下分配局部信用,我们引入了从奖励梯度导出的连续空间信用图。关键的是,我们采用了一种基于中位数绝对偏差(MAD)并结合温度缩放的异常值鲁棒归一化方案,有效消除梯度噪声,同时突出功能性的提示对齐区域。在GenEval上的广泛评估表明,我们的方法提供了SOTA对齐质量,实现了与DiffusionNFT等顶级方法相当的收敛速度,同时在性能上大幅优于基于Flow-GRPO的方法。
cs.CV / 100 / 2609.32343

OpenMASC: An Open-Source Pipeline for Cross-Trajectory Metal-Aware Sampling and Correction in Accelerated MRI

OpenMASC:用于加速MRI中跨轨迹金属感知采样与校正的开源流程
Lu, Zhengyi, Lu, Ming, Qu, Chongyu, Zhu, Junchao, Guo, Junlin, Lionts, Marilyn, Zhu, Yanfan, Yang, Yuechen, Yao, Tianyuan, Rajagopal, Jayasai, Landman, Bennett Allan, Wang, Xiao, Yan, Xinqiang, Huo, Yuankai
Abstract
Metal implants corrupt MRI measurements throughout $k$-space, yet existing accelerated MRI methods assume clean data and most metal artifact reduction approaches assume fully sampled acquisitions. No public dataset provides paired $k$-space and images with and without metal for the same anatomy, and no framework jointly addresses artifact-aware acquisition and reconstruction across sampling trajectories. We present OpenMASC, an open-source pipeline covering the full workflow from data generation to deployment. A physics-based data generation module converts public CT volumes into paired clean and metal-corrupted MRI data in both Cartesian and radial formats. MA-VarNet, an unrolled reconstruction network with a per-cascade DC Rectifier, corrects artifacts that data-consistency steps reintroduce from corrupted measurements. A reinforcement learning agent actively selects $k$-space readouts and co-trains with the reconstruction network through a decoupled three-stage procedure. The framework is trajectory-agnostic except for the data-consistency operator, supporting both Cartesian and radial acquisition without architectural changes. Experiments on two datasets at $4\times$ and $8\times$ acceleration demonstrate consistent improvements over conventional and learned baselines on both trajectories.
Chinese Translation
金属植入物会破坏整个 k 空间的 MRI 测量数据,然而现有的加速 MRI 方法假设数据是干净的,而大多数金属伪影减少方法假设完全采样的采集。目前没有公共数据集提供同一解剖结构下有无金属的配对 k 空间和图像,也没有框架能够联合解决跨采样轨迹的伪影感知采集和重建问题。我们提出了 OpenMASC,一个涵盖从数据生成到部署的完整工作流程的开源流程。一个基于物理的数据生成模块将公共 CT 体积转换为笛卡尔和径向格式的配对的干净和金属污染的 MRI 数据。MA-VarNet 是一个展开式重建网络,具有每个级联的 DC 整流器,用于校正数据一致性步骤从损坏的测量中重新引入的伪影。一个强化学习智能体主动选择 k 空间读出,并通过解耦的三阶段过程与重建网络协同训练。该框架除了数据一致性算子外与轨迹无关,无需架构更改即可支持笛卡尔和径向采集。在两个数据集上进行的 4 倍和 8 倍加速实验表明,在两种轨迹上均比传统基线和学习基线有一致的改进。
cs.CV / 101 / 2609.32352

EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding

EyeVQA: 从识别到空间定位的眼科视觉语言模型基准测试
Shao, Gujie, Xie, Zixun, Xing, Xuechun, Wang, Ruixiang, Lan, Ziyun, Qi, Yanlin, Zhang, Gangyi, Yang, Yuxin, Li, Dawei, Tang, Haiming
Abstract
Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically designed for individual diseases or specialized tasks, making it difficult to systematically evaluate whether VLMs can move beyond disease recognition toward comparative reasoning and fine-grained spatial grounding. We introduce EyeVQA, a unified visual question answering benchmark for comprehensive evaluation of ophthalmic VLMs. EyeVQA is constructed from 21 available ophthalmic datasets and contains 20,000 question-answer pairs spanning six disease groups and seven question types: Single-Choice, Multi-Select, Variable-Select, True-False, Ranking, Point Location, and Bounding Box. Gold answers are deterministically derived from source-provided diagnoses, severity grades, clinical findings, segmentation masks, bounding boxes, and anatomical landmarks, enabling reproducible evaluation without relying on model-generated annotations. Notably, 44.5% of the questions require reasoning across multiple images, extending evaluation beyond conventional single-image medical VQA. We benchmark fourteen representative general-purpose, scientific, and medically specialized VLMs under a unified zero-shot protocol. The best-performing model only achieves an overall score of 62.8, while substantial gaps remain in spatial grounding and cross-task generalization. These results highlight the limitations of current VLMs in comprehensive ophthalmic visual understanding and establish EyeVQA as a diagnostic benchmark for developing more reliable and spatially grounded ophthalmic multimodal models. The project page is available at https://github.com/PKUTHM/EyeVQA.
Chinese Translation
视觉语言模型(VLMs)在医学图像理解方面展现出越来越大的潜力,但它们在眼科影像中的能力仍未被充分表征。现有的眼科数据集通常针对单一疾病或特定任务设计,难以系统评估VLM能否超越疾病识别,迈向比较推理和细粒度空间定位。我们提出EyeVQA,一个用于全面评估眼科VLM的统一视觉问答基准。EyeVQA由21个可用的眼科数据集构建而成,包含20,000个问答对,涵盖六个疾病组和七种问题类型:单选、多选、可变选择、判断、排序、点定位和边界框。标准答案确定性地来源于源数据提供的诊断、严重程度分级、临床发现、分割掩码、边界框和解剖标志,从而无需依赖模型生成的标注即可实现可重复评估。值得注意的是,44.5%的问题需要跨多张图像推理,将评估拓展到传统单图像医学VQA之外。我们在统一的零样本协议下对14个代表性的通用、科学和医学专用VLM进行了基准测试。性能最佳的模型仅获得62.8的总分,而在空间定位和跨任务泛化方面仍存在显著差距。这些结果凸显了当前VLM在全面眼科视觉理解方面的局限性,并将EyeVQA确立为开发更可靠且具有空间定位能力的眼科多模态模型的诊断基准。项目页面见 https://github.com/PKUTHM/EyeVQA。
cs.CV / 102 / 2609.32353

Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction

更少Token,更多自教:面向极端视觉Token缩减的同策略自蒸馏
Li, Junxian, Yang, Ruixuan, Zhang, Tianao, Xu, Tiange, Dong, Weisheng, Zhang, Yulun
Abstract
Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step further by asking how a heavily compressed MLLM should learn from the states induced by its own generations. This setting naturally calls for on-policy self-distillation: a heavily compressed model is supervised on the states induced by its own generations, while its full-token counterpart serves as an information-rich teacher. Based on this insight, we propose LT-OPD, a training framework for extreme visual-token reduction. The student rolls out responses with only a small fraction of visual tokens, and a frozen full-token copy of the same MLLM provides distributional supervision along these student-generated trajectories. To stabilize on-policy learning when visual evidence is severely limited, we further introduce a budget-level curriculum that progressively decreases the token budget during training. Across nine benchmarks on Qwen3.5-4B, LT-OPD raises average retained performance under 5% visual-token retention from 68.6% to 82.3%, outperforming training-free, training-based, and reinforcement-learning baselines at the same budget. The gains transfer consistently to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. LT-OPD also reduces KV-cache usage by 85.2% and prefill FLOPs by 85.4% without additional inference overhead, demonstrating that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.
Chinese Translation
视觉Token缩减是加速多模态大语言模型(MLLMs)的有效方法,但在极低的Token预算下性能会迅速下降。现有工作探索了视觉Token选择和基于训练的适应来应对缩减后的视觉输入。我们更进一步,探究高度压缩的MLLM应如何从其自身生成所诱导的状态中学习。这一场景自然需要同策略自蒸馏:高度压缩的模型在其自身生成所诱导的状态上受到监督,而全Token对应模型则作为信息丰富的教师。基于这一见解,我们提出了LT-OPD,一种用于极端视觉Token缩减的训练框架。学生模型仅使用一小部分视觉Token生成回复,而同一MLLM的冻结全Token副本则沿着这些学生生成的轨迹提供分布监督。为了在视觉证据严重受限时稳定同策略学习,我们进一步引入了预算级别的课程学习,在训练过程中逐步减少Token预算。在Qwen3.5-4B的九个基准测试上,LT-OPD在5%视觉Token保留率下将平均保留性能从68.6%提升至82.3%,在相同预算下优于免训练、基于训练和强化学习的基线。增益一致地迁移到Qwen3.5-9B、GLM-4.6V-9B和LLaVA-OV-1.5-4B。LT-OPD还减少了85.2%的KV缓存使用量和85.4%的预填充FLOPs,且不增加额外推理开销,表明同策略学习能够大幅恢复因极端视觉Token缩减而损失的能力。
cs.CV / 103 / 2609.32362

StegGNN: Learning Graphical Representation for Image Steganography

StegGNN:学习图像隐写术的图表示
Kumar, Abhinav, Singhal, Shorya, Pandey, Agam, Kumar, Tushar, Jindal, Sukrit
Abstract
Image steganography refers to embedding secret messages within cover images while maintaining imperceptibility. Recent advances in deep learning - primarily driven by Convolutional Neural Networks (CNNs) and architectures such as inverse neural networks, autoencoders, and generative adversarial networks - have led to notable progress. However, these frameworks are primarily built on CNN architectures, which treat images as regular grids and are limited by their receptive field size and a bias toward spatial locality. In parallel, Graph Neural Networks (GNNs) have recently demonstrated strong adaptability in several computer vision tasks, achieving state-of-the-art performance with architectures such as Vision GNN (ViG). This work moves in that direction and introduces StegGNN - a novel autoencoder-based, cover-agnostic image steganography framework based on GNNs. By modeling images as graph structures, our approach leverages the representational flexibility of GNNs over the grid-based rigidity of conventional CNNs. We conduct extensive experiments on standard benchmark datasets to evaluate visual quality and imperceptibility. Our results show that our GNN-based method performs comparably to existing CNN benchmarks. These findings suggest that GNNs provide a promising alternative representation for steganographic embedding and open the field of deep learning-based steganography to further exploration of GNN-based architectures.
Chinese Translation
图像隐写术是指在保持不可感知性的同时,将秘密信息嵌入到载体图像中。近年来,深度学习的进展——主要由卷积神经网络(CNNs)以及诸如逆神经网络、自编码器和生成对抗网络等架构推动——已带来了显著进步。然而,这些框架主要基于CNN架构,将图像视为规则网格,并受限于其感受野大小和对空间局部性的偏好。与此同时,图神经网络(GNNs)最近在多个计算机视觉任务中展现出强大的适应性,凭借如Vision GNN (ViG)等架构达到了最先进的性能。本工作朝这一方向推进,并提出了StegGNN——一种新颖的基于自编码器、与载体无关的、基于GNN的图像隐写框架。通过将图像建模为图结构,我们的方法利用了GNN的表征灵活性,超越了传统CNN基于网格的刚性。我们在标准基准数据集上进行了大量实验,以评估视觉质量和不可感知性。结果表明,我们基于GNN的方法与现有的CNN基准性能相当。这些发现表明,GNN为隐写嵌入提供了一种有前景的替代表示,并为基于深度学习的隐写术领域开辟了进一步探索基于GNN架构的道路。
cs.CV / 104 / 2609.32376

An End-to-End Latent-Rollout Approach for Pushing Few-Step ImageNet-$256$ Generation to FID $1.11$ without Fr\'echet Losses

一种端到端潜在展开方法:将少步ImageNet-256生成推进至FID 1.11且无需Fréchet损失
Xu, Xiaoran, Wang, Yujing
Abstract
Iterative generation poses a joint optimization problem across steps, as intermediate predictions shape subsequent computations and ultimately determine the final output distribution. Few-step generators distilled from pretrained diffusion and flow-matching models make such optimization computationally practical end to end. We build on this opportunity with a distill-then-refine approach that uses teacher imitation to establish a strong initialization for a few-step rollout in latent space, then shifts to end-to-end refinement of the complete latent rollout against real data. We introduce FiST (Flow-in-Stage Transformer), an architecture that composes learned latent-state transitions in a few stages using a shared Transformer, with optional cross-stage hidden communication. Distillation applies teacher-forced regression to selected states along teacher trajectories; refinement replaces this supervision with adversarial and auxiliary classification objectives on the final latent output. A trainable discriminator module operates on semantically rich features extracted from clean real and generated latents by a frozen SiT backbone pretrained with REPA. All training takes place in latent space, without image decoding. During refinement, FiST consumes its own intermediate predictions, and endpoint gradients pass through every generation stage. For class-conditional generation on ImageNet at $256\times256$, our approach achieves FID 1.11 (IS 282) with three stages and FID 1.15 (IS 280) with two. These results demonstrate competitive few-step generation through learned distribution-level supervision, without explicit Fr\'echet-distance minimization. Ablations characterize how distillation, pretrained checkpoint choices, refinement supervision, and cross-stage hidden communication affect generation quality.
Chinese Translation
迭代生成提出了一个跨步骤的联合优化问题,因为中间预测会影响后续计算并最终决定输出分布。从预训练扩散和流匹配模型中蒸馏出的少步生成器使得这种优化在端到端计算上可行。我们基于这一机会提出了一种“先蒸馏后精炼”的方法,利用教师模仿为潜在空间中的少步展开建立强初始化,然后转向针对真实数据对整个潜在展开进行端到端精炼。我们引入了FiST(Flow-in-Stage Transformer),一种使用共享Transformer在几个阶段中组合学习到的潜在状态转移的架构,并可选择跨阶段隐藏通信。蒸馏对教师轨迹上选定的状态应用教师强制回归;精炼则用对最终潜在输出的对抗和辅助分类目标取代这种监督。一个可训练的判别器模块作用于由冻结的、用REPA预训练的SiT主干从干净的真实和生成潜在中提取的语义丰富特征。所有训练都在潜在空间中进行,无需图像解码。在精炼过程中,FiST使用自己的中间预测,并且端点梯度会通过每个生成阶段。对于ImageNet上256×256的类别条件生成,我们的方法在三阶段下达到FID 1.11(IS 282),在两阶段下达到FID 1.15(IS 280)。这些结果表明,通过学习到的分布级监督,无需显式的Fréchet距离最小化,即可实现具有竞争力的少步生成。消融实验刻画了蒸馏、预训练检查点选择、精炼监督以及跨阶段隐藏通信对生成质量的影响。
cs.CV / 105 / 2609.32389

RefCompose: Multi-Reference Image Generation via LoRA-Conditioned Diffusion

RefCompose: 基于LoRA条件扩散的多参考图像生成
Kuppa, Sai Sri Teja, Shinde, Parth, S, Priyadharsan Balaji, Harshavardhan, Jinka, Ramanarayanan, Sriprabha
Abstract
Filmmakers and visual artists routinely need to compose multiple references, actors, locations, props, cultural elements, into a single coherent shot, but existing tools either fail to scale past a handful of references or destroy fine grained subject identity in the process, since per reference tokenization scales memory linearly with reference count $N$ and generated content often departs from the given references rather than reproducing them. We propose \textbf{RefCompose}, a pixel space compositional conditioning framework that decouples \emph{where} things go from \emph{what} they look like, via a single fixed resolution reference canvas that keeps conditioning size constant regardless of reference count. Spatial layout is induced at inference time from a frozen diffusion transformer and extracted via Grounding DINO, requiring no LLM or dedicated layout model, while dual stream LoRA adapters inject a layout derived depth map and the encoded canvas through separate low rank streams, disentangling geometric scaffolding from localized appearance. On the Dense Layout protocol, RefCompose consistently outperforms layout based and state of the art multi reference baselines on color, texture, shape, spatial accuracy, and identity/content preservation at higher reference counts, all with constant inference memory, making it a practical building block for multi subject cinematic composition at production scale.
Chinese Translation
电影制作人和视觉艺术家经常需要将多个参考——演员、地点、道具、文化元素——组合成一个连贯的镜头,但现有工具要么无法扩展到超过少数几个参考,要么在此过程中破坏细粒度的主体身份,因为针对每个参考的标记化会随着参考数量 $N$ 线性扩展内存,并且生成的内容往往偏离给定的参考,而不是重现它们。我们提出 RefCompose,一个像素空间组合条件框架,它通过一个固定的分辨率参考画布,将物品的位置与它们的外观解耦,该画布保持条件大小恒定,与参考数量无关。空间布局在推理时从冻结的扩散变换器中诱导,并通过 Grounding DINO 提取,无需 LLM 或专用布局模型,而双流 LoRA 适配器通过单独的低秩流注入布局导出的深度图和编码画布,将几何脚手架与局部外观解耦。在 Dense Layout 协议上,RefCompose 在颜色、纹理、形状、空间准确性以及身份/内容保留方面,在更高的参考数量下,始终优于基于布局的和最先进的多参考基线,同时具有恒定的推理内存,使其成为生产规模下多主体电影构图的实用构建模块。
cs.CV / 106 / 2609.32399

Endo-TSR: Temporal Spectral Modeling of Appearance and Motion for Endoscopic Reconstruction

Endo-TSR:面向内窥镜重建的外观与运动时间谱建模
Wu, Taoyu, Miao, Yiyi, Shao, Qi, Li, Zhuoxiao, Tang, Zhe, Yu, Limin, Huang, Baoru
Abstract
Endoscopic scene reconstruction requires modeling tissue motion and temporal appearance while recovering fine surface detail. Deformable Gaussian models provide explicit trajectories, but their fixed colour coefficients lack a dedicated temporal representation for photometric changes. We propose Endo-TSR, which augments deformable Gaussian splatting with bounded Fourier colour residuals and independent translation residuals on shared temporal frequencies. The colour residuals capture local appearance changes, while a Mat\'ern spectral prior regularises motion corrections. Multi-scale Laplacian supervision guides tissue-detail recovery during joint image fitting. Extensive experiments on the EndoNeRF and StereoMIS datasets demonstrate state-of-the-art rendering quality, with the highest PSNR across all evaluated sequences. Ablation studies show that temporal appearance yields the largest PSNR gain among the tested component additions, while appearance and detail supervision jointly improve rendering with fixed Gaussian counts.
Chinese Translation
内窥镜场景重建需要在恢复精细表面细节的同时,对组织运动和时间外观进行建模。可变形高斯模型提供了显式轨迹,但其固定的颜色系数缺乏用于光度变化的专用时间表示。我们提出了Endo-TSR,它通过有界傅里叶颜色残差和共享时间频率上的独立平移残差来增强可变形高斯泼溅。颜色残差捕获局部外观变化,而Matérn谱先验则正则化运动校正。多尺度拉普拉斯监督在联合图像拟合过程中引导组织细节的恢复。在EndoNeRF和StereoMIS数据集上的大量实验证明了最先进的渲染质量,在所有评估序列中均获得了最高的PSNR。消融实验表明,在测试的组件添加中,时间外观产生了最大的PSNR增益,而外观和细节监督在固定高斯数量下共同改善了渲染。
cs.CV / 107 / 2609.32405

Toward On-Chip Training of Spiking Neural Networks for Dense Event-Based Vision

面向密集事件视觉的脉冲神经网络片上训练
Vaillant, Maxime, Carlier, Axel, Ng, Lai Xing, Hurter, Christophe, Cottereau, Benoit R.
Abstract
Event cameras provide low-latency, asynchronous visual sensing for resource-constrained robotics. Spiking neural networks (SNNs) process event streams naturally, but training deep SNNs with backpropagation through time (BPTT) requires substantial memory and remains difficult on neuromorphic hardware. Local learning avoids this by restricting error propagation to local blocks, but existing methods mainly target classification rather than dense prediction. We introduce DELL (Dense Event-driven Local Learning), a block wise scheme for dense event-based vision that replaces global gradient propagation with local dense supervision. Learnable, spatially structured local heads supervise each block at its appropriate resolution while preserving temporal dynamics within blocks. We evaluate DELL on optical-flow regression and semantic segmentation with a fully spiking U-shaped architecture. On DSEC optical flow, DELL reduces peak training memory by 39.6% relative to end-to-end BPTT while improving accuracy, reaching 1.670 px endpoint error on the official test benchmark versus 1.941 px for the same backbone trained end-to-end. Block detachment behaves more like a regularizer than a constraint. DECOLLE, the existing local-learning baseline, relies on fixed random local read-outs poorly suited to dense regression, resulting in a 3.9x higher endpoint error; learnable local heads recover this loss and outperform end-to-end training across all optical-flow metrics. On segmentation, they recover most of the performance gap, although DELL remains a few mIoU points behind end-to-end training. With 2.3M parameters, 24x fewer than the strongest SNN baseline, the backbone remains competitive with the SNN state of the art on DSEC. These results extend local learning to dense event-based prediction while substantially reducing training memory.
Chinese Translation
事件相机为资源受限的机器人提供低延迟、异步的视觉感知。脉冲神经网络(SNNs)自然地处理事件流,但使用时间反向传播(BPTT)训练深度SNN需要大量内存,并且在神经形态硬件上仍然困难。局部学习通过将误差传播限制在局部块来避免这一问题,但现有方法主要针对分类而非密集预测。我们引入了DELL(密集事件驱动局部学习),这是一种用于密集事件视觉的逐块方案,它用局部密集监督取代全局梯度传播。可学习的、空间结构化的局部头在其适当的分辨率下监督每个块,同时保留块内的时间动态。我们在光流回归和语义分割上使用全脉冲U形架构评估DELL。在DSEC光流上,相对于端到端BPTT,DELL降低了39.6%的峰值训练内存,同时提高了准确性,在官方测试基准上达到1.670像素端点误差,而相同主干端到端训练为1.941像素。块分离更像是一种正则化器而非约束。DECOLLE,现有的局部学习基线,依赖于固定随机局部读出,不适合密集回归,导致端点误差高出3.9倍;可学习的局部头恢复了这一损失,并在所有光流指标上超越了端到端训练。在分割上,它们恢复了大部分性能差距,尽管DELL仍比端到端训练低几个mIoU点。该主干网络具有230万参数,比最强的SNN基线少24倍,在DSEC上仍与SNN的最先进水平保持竞争力。这些结果将局部学习扩展到密集事件预测,同时大幅减少了训练内存。
cs.CV / 108 / 2609.32415

RefAdapt-DiT: Adaptive Joint Attention for Reference-Conditioned Diffusion Transformers

RefAdapt-DiT:面向参考条件扩散Transformer的自适应联合注意力
Tang, Jian, Fan, Jiawei, Zhou, Qiannan, Liu, Qingbin, Bian, Jiang, Li, Zang
Abstract
Diffusion Transformers (DiTs) have become the standard backbone for high-quality generative modeling, yet deploying them in conditional generation tasks remains computationally prohibitive because bidirectional joint attention repeatedly processes large reference streams. While existing optimization schemes mitigate generic temporal redundancy, they typically rely on coarse-grained static reuse and overlook the distinct dynamics of references and targets. Specifically, we observe that reference representations often evolve slowly along the generation trajectory, while the target often assigns little attention mass to them; reference drift and this target-to-reference exposure jointly shape how strongly stale reference states affect the target. To exploit these patterns, we introduce \RefAdapt, a training-free framework for adaptive control of joint attention between references and targets. Instead of rigid static strategies, \RefAdapt combines consecutive target-Q change with previously observed target-to-reference attention mass to control reference computation adaptively at block granularity. Under ultra-few-step settings, \RefAdapt enables speedups of up to $2.097\times$ on 4-step MiniMax H3 and $3.54\times$ on 8-step Qwen Image Edit, while maintaining comparable visual quality.
Chinese Translation
扩散Transformer(Diffusion Transformers, DiTs)已成为高质量生成建模的标准主干,但将其部署到条件生成任务中仍然在计算上代价高昂,因为双向联合注意力会反复处理大规模的参考流。尽管现有优化方案缓解了通用时间冗余,但它们通常依赖粗粒度的静态复用,并忽视了参考与目标各自不同的动态特性。具体而言,我们观察到参考表示通常沿生成轨迹缓慢演化,而目标往往对其分配较少的注意力权重;参考漂移与这种目标到参考的暴露程度共同决定了过时参考状态对目标的影响强度。为利用这些模式,我们提出 RefAdapt,一种无需训练的框架,用于自适应控制参考与目标之间的联合注意力。不同于僵化的静态策略,RefAdapt 将连续的目标-Q 变化与先前观测到的目标到参考注意力权重相结合,在块粒度上自适应地控制参考计算。在超少步设置下,RefAdapt 在 4 步 MiniMax H3 上实现最高 2.097 倍加速,在 8 步 Qwen Image Edit 上实现 3.54 倍加速,同时保持相当的视觉质量。
cs.CV / 109 / 2609.32427

CityToolVQA: Tool-Augmented Visual Question Answering for 3D Spatial Cognition in Urban Low-Altitude Environments

CityToolVQA:面向城市低空环境3D空间认知的工具增强视觉问答
Yu, Boao, Nie, Yingzhen, Hu, Yue, Zhu, Zhengqiu, Ju, Rusheng
Abstract
CityToolVQA addresses the weak performance of Vision-Language Models (VLMs) on quantitative tasks in urban low-altitude visual question answering. We divide the seven tasks into qualitative and quantitative groups: qualitative questions are answered directly by the VLM, whereas quantitative questions are handled by an external visual-geometric toolchain that performs object grounding, segmentation, depth back-projection, and spatial computation. The toolchain can be attached to different VLMs in a zero-shot manner; a Depth-Assisted Prompt Inference (DAPI) fallback is triggered when the main-chain detection is invalid or unreliable, and CityToolVQA-SFT adapts the 8B backbone to tool-conditioned inputs. On the 73,324-question Open3D-VQA-v2 test set, CityToolVQA-SFT (Qwen3-VL-8B) reaches 67.6% overall accuracy, and attaching the toolchain to ten open-source VLMs improves quantitative-task accuracy by 12.7-36.9 percentage points. These results indicate that externalizing explicit 3D geometric computation effectively complements the limited ability of RGB-only VLMs to estimate metric distances and object sizes.
Chinese Translation
CityToolVQA解决了视觉语言模型(VLM)在城市低空视觉问答中定量任务表现不佳的问题。我们将七个任务分为定性和定量两组:定性问题由VLM直接回答,而定量问题则由外部视觉几何工具链处理,该工具链执行物体定位、分割、深度反投影和空间计算。该工具链可以零样本方式附加到不同的VLM上;当主链检测无效或不可靠时,会触发深度辅助提示推理(DAPI)回退机制,而CityToolVQA-SFT将8B骨干网络适配到工具条件输入。在包含73,324个问题的Open3D-VQA-v2测试集上,CityToolVQA-SFT(Qwen3-VL-8B)达到67.6%的总体准确率,将工具链附加到十个开源VLM上可将定量任务准确率提高12.7-36.9个百分点。这些结果表明,外化显式3D几何计算有效补充了仅RGB的VLM在估计度量距离和物体尺寸方面的有限能力。
cs.CV / 110 / 2609.32454

De-biasing Skeleton-based Action Recognition with Convex Hull Adaptive Shift

基于凸包自适应偏移的骨架动作识别去偏方法
Liu, Mengyuan, Wen, Yuhang, Zhang, Yi, Wu, Songtao, Liu, Hong, Yuan, Junsong, Ding, Beichen
Abstract
Skeleton sequences can represent both individual actions and multi-entity interactions, encompassing human bodies, hands, objects, and robots. Existing approaches to recognize skeleton-based actions and interactions usually adopt a late fusion strategy, which expects individuals are independent and identically distributed to train a robust weight-shared entity encoder. However, observed entity bias in various skeletal data violates this assumption, leading to suboptimal optimization of backbone models that might produce wrong recognition results. This bias arises from the world coordinate system's initial configuration, where the choice of origin often creates bias in representation. To this end, we propose a Convex Hull Adaptive Shift based normalization method to reduce Entity bias (CHASE), improving performance across a variety of skeleton-based action and interaction recognition tasks. To adaptively apply plausible shifts to the input skeletons, we formulate a plug-and-play parameterized network that ensures the relocated world origin lies within the skeleton convex hull, which avoids non-convergence by limiting the search space. To further minimize entity bias, we incorporate an auxiliary objective that leverages pair-wise distribution distances to guide network optimization. To support both single- and multi-entity actions, we propose a sub-entity strategy that offers a consistent formulation for both scenarios. Moreover, CHASE demonstrates compatibility with various intra-skeleton modalities, such as bones and velocities, highlighting its adaptability. Essentially, our method works as a normalization approach to reduce entity bias, enabling subsequent classifiers to achieve improved recognition performance across diverse settings. Extensive experiments on 7 datasets verify our approach by seamlessly integrating with various backbones and significantly boosting their performance.
Chinese Translation
骨架序列可以表示个体动作和多实体交互,涵盖人体、手、物体和机器人。现有的基于骨架的动作和交互识别方法通常采用后期融合策略,该策略假设个体是独立同分布的,以训练一个鲁棒的权重共享实体编码器。然而,在各种骨架数据中观察到的实体偏差违反了这一假设,导致骨干模型的优化次优,可能产生错误的识别结果。这种偏差源于世界坐标系的初始配置,其中原点的选择往往在表示中产生偏差。为此,我们提出了一种基于凸包自适应偏移的归一化方法以减少实体偏差(CHASE),提升了各种基于骨架的动作和交互识别任务的性能。为了自适应地对输入骨架应用合理的偏移,我们构建了一个即插即用的参数化网络,确保重新定位的世界原点位于骨架凸包内,通过限制搜索空间避免了不收敛。为了进一步最小化实体偏差,我们引入了一个辅助目标,利用成对分布距离来指导网络优化。为了支持单实体和多实体动作,我们提出了一种子实体策略,为两种场景提供了一致的公式化方法。此外,CHASE展示了与各种骨架内模态(如骨骼和速度)的兼容性,突出了其适应性。本质上,我们的方法作为一种归一化方法来减少实体偏差,使后续分类器能够在各种设置下实现更好的识别性能。在7个数据集上的大量实验通过无缝集成各种骨干网络并显著提升其性能,验证了我们的方法。
cs.CV / 111 / 2609.32455

QuacamFM: Quaternion-Constrained Flow Matching for Camera Pose Estimation

QuacamFM:用于相机位姿估计的四元数约束流匹配
Tran, Bao-Long, Le, Cuong, Dehdarirad, Tahereh, Viksten, Fredrik, Forssén, Per-Erik
Abstract
Camera pose estimation from multi-view images remains a challenge in computer vision. Traditional methods often address this problem using Structure-from-Motion (SfM) with bundle adjustment. However, camera poses estimated from sparse views are inherently ambiguous due to insufficient geometric constraints. Recent work leverages probabilistic models, such as diffusion models, to generate multiple camera pose hypotheses and therefore capture this uncertainty better. Most of these methods represent camera rotations using unit quaternions, but treat them as unconstrained 4D vectors during the generative processes, thereby ignoring the unit-norm constraint of quaternions. Unconstrained quaternions create non-smooth and suboptimal generation trajectories. To this end, we propose *QuacamFM*, a quaternion-constrained flow matching framework for camera pose estimation that preserves unit quaternion representations throughout the entire flow trajectory. We design the optimal transport of the quaternion flows using smooth spherical linear interpolation. Experiments on CO3Dv2 demonstrate our method's advantage in camera pose accuracy over diffusion-based methods and classical SfM approaches. We further show that our quaternion-constrained formulation outperforms the naive application of standard flow matching to 4D quaternion vectors on sparse-view camera pose estimation. Finally, it is observed that QuacamFM generalizes well across datasets and in-the-wild examples.
Chinese Translation
从多视图图像估计相机位姿仍然是计算机视觉中的一个挑战。传统方法通常使用运动恢复结构(SfM)和光束法平差来解决这个问题。然而,由于几何约束不足,从稀疏视图估计的相机位姿本质上具有歧义性。最近的工作利用概率模型,例如扩散模型,来生成多个相机位姿假设,从而更好地捕捉这种不确定性。这些方法大多使用单位四元数表示相机旋转,但在生成过程中将它们视为无约束的4D向量,从而忽略了四元数的单位范数约束。无约束的四元数会产生非平滑且次优的生成轨迹。为此,我们提出了 QuacamFM,一个用于相机位姿估计的四元数约束流匹配框架,它在整个流轨迹中保持单位四元数表示。我们使用平滑球面线性插值设计了四元数流的最优传输。在 CO3Dv2 上的实验表明,我们的方法在相机位姿精度上优于基于扩散的方法和经典 SfM 方法。我们进一步表明,在稀疏视图相机位姿估计上,我们的四元数约束公式优于将标准流匹配朴素地应用于4D四元数向量。最后,我们观察到 QuacamFM 在不同数据集和真实场景示例中具有良好的泛化能力。
cs.CV / 112 / 2609.32456

Seeing Parts, Reasoning about Worlds: Visual Inference under Partial Observation

见部分,推理世界:部分观测下的视觉推断
Wang, Wei, Zhang, Wenqiao, Lin, Yutong, Xiao, Jun, Zhuang, Yueting
Abstract
World modeling under partial observation requires reasoning about the complete worlds that remain compatible with limited visual evidence. Occluded objects and unseen regions can leave several world states possible; additional views can exclude alternatives and strengthen the conclusions supported by the observations. We introduce WorldScope to study this process through possible-world semantics, evidence-grounded data, and learned visual representations. WorldScope-1.2M provides 1.2 million English question-answer pairs spanning eight world properties, ten task interfaces, and three observation protocols. Its answers encode confirmed facts, supported bounds, and unresolved possibilities. Complementary supervision comprises 4,800 certified counterworld groups with equivalent base observations and different hidden object configurations and query answers. These groups provide physical witnesses of ambiguity and training-only labels for world compatibility and view-induced exclusions. We propose WorldFlow, which composes cross-view entity evidence and support-surface coverage into an image-subset evidence lattice. Counterworld compatibility and transition objectives train subset representations to reflect how new observations constrain possible worlds. A shared answer generator uses these representations to predict the strongest supported conclusion. WorldScope-Bench evaluates claim judgments and evidence-dependent conclusions as views are selected, combined, removed, or ordered. On its 5,000-question test set, WorldFlow reaches 64.34% exact accuracy, improving over the same backbone trained on QA alone by 24.88 percentage points. It retains 50.43% accuracy on the 3,000 questions from structure-disjoint scenes.
Chinese Translation
部分观测下的世界建模要求对与有限视觉证据保持兼容的完整世界进行推理。被遮挡的物体和未见的区域可能留下几种可能的世界状态;额外的视图可以排除其他可能并加强观测所支持的结论。我们引入了 WorldScope,通过可能世界语义、基于证据的数据和学习到的视觉表示来研究这一过程。WorldScope-1.2M 提供了 120 万个英语问答对,涵盖八种世界属性、十种任务界面和三种观测协议。其答案编码了已确认事实、有支持的边界和未解决的可能性。补充监督包含 4,800 个经认证的 counterworld 组,具有等价的基础观测以及不同的隐藏物体配置和查询答案。这些组提供了歧义性的物理见证以及用于世界兼容性和视图诱导排除的仅训练标签。我们提出了 WorldFlow,它将跨视图实体证据和支持面覆盖组合成图像子集证据格。Counterworld 兼容性和转移目标训练子集表示,以反映新观测如何约束可能世界。一个共享答案生成器使用这些表示来预测最强支持的结论。WorldScope-Bench 在视图被选择、组合、移除或排序时评估主张判断和依赖证据的结论。在其 5,000 问题的测试集上,WorldFlow 达到 64.34% 的精确准确率,比仅在 QA 上训练的相同骨干提高了 24.88 个百分点。它在来自结构不相交场景的 3,000 个问题上保持了 50.43% 的准确率。
cs.CV / 113 / 2609.32460

REMEDY: How Far Is Video Generation from Medical Education World Models?

REMEDY:视频生成距离医学教育世界模型还有多远?
Tan, Lixing, Zhou, Yanghao, Xia, Qing, Guo, Yuting, Li, Shuai, Hao, Aimin
Abstract
Recent video generation models produce realistic videos and show potential as a foundation for world models. These advances create opportunities for generating medical teaching demonstrations, which requires both convincing visual quality and precise procedural actions. However, whether current generators can meet these requirements has not been measured. To address this problem, we introduce Readiness Evaluation of Medical Education Demonstration sYnthesis (REMEDY), to our knowledge, the first benchmark for AI-generated medical teaching demonstrations. REMEDY provides 900 first frames from real demonstration videos, covering 12 tasks across four scenarios: operating room, imaging, clinic and bedside, and resuscitation. Five contemporary open-source video generation models produce 4,500 videos from these frames. We combine task-specific clinical checklists with video and motion quality metrics. Evaluation covers four dimensions: clinical action following, clinical profiles, video quality, and motion quality. Our results show that realistic appearance and temporal consistency do not ensure correct clinical actions. Even the most advanced MiniMax-H3 achieves only 28.25% on strict clinical success rate, and fine-grained clinical actions remain challenging. These findings establish a foundation and roadmap for developing future medical education world models.
Chinese Translation
近期,视频生成模型能够生成逼真的视频,并展现出作为世界模型基础的潜力。这些进展为生成医学教学演示创造了机会,而这需要令人信服的视觉质量和精确的操作步骤。然而,当前生成器能否满足这些要求尚未得到评估。为了解决这个问题,我们提出了医学教育演示合成就绪度评估(Readiness Evaluation of Medical Education Demonstration sYnthesis,REMEDY),据我们所知,这是首个针对AI生成的医学教学演示的基准测试。REMEDY提供了来自真实演示视频的900个首帧,涵盖四个场景中的12项任务:手术室、影像、门诊与床旁、以及复苏。五个当代开源视频生成模型基于这些帧生成了4,500个视频。我们将任务特定的临床检查表与视频和运动质量指标相结合。评估涵盖四个维度:临床动作遵循、临床特征、视频质量和运动质量。我们的结果表明,逼真的外观和时间一致性并不能确保正确的临床动作。即使是目前最先进的MiniMax-H3,在严格的临床成功率上也仅达到28.25%,细粒度的临床动作仍然具有挑战性。这些发现为开发未来的医学教育世界模型奠定了基础并提供了路线图。
cs.CV / 114 / 2609.32462

Can Motion-Language Models Ground Structure? STRIDE for Evaluating the Evaluators

运动-语言模型能否实现结构接地?用于评估评估器的 STRIDE
Tan, Lixing, Xia, Qing, Guo, Yuting, Li, Shuai, Hao, Aimin
Abstract
Motion-language models are typically scored by motion-language evaluators, but how well these evaluators ground language structure remains unclear. Here, we introduce the Structure grounding via Temporal-order, Reflection, and Identity Diagnostic Evaluation (STRIDE) benchmark to systematically evaluate the ability of evaluators to track temporal order, mirror reflection, and action identity. STRIDE comprises $5{,}869$ triples, each consisting of a motion, its original caption, and a perturbed caption, spanning both short and long descriptions. We likelihood-balance caption pairs to reduce text-only bias and estimate each evaluator's caption preference under unrelated motions to measure the discrimination gain from matched motions relative to this baseline. Our experiments reveal weak structural grounding and severe deficits in mirror sensitivity among the audited evaluators, which commonly used evaluation protocols fail to expose. To understand why these limitations go undetected in standard tests, we examine the evaluators more closely. We find that text-only priors alone can solve naive perturbation tests on existing datasets, while common retrieval and distributional metrics barely respond to structural corruption introduced by mirroring ground-truth motions. These findings suggest a natural intervention: structural hard negatives. Our experiments show that a simple modification to contrastive learning substantially improves performance on temporal order and mirror reflection. The benchmark and code will be released.
Chinese Translation
运动-语言模型通常由运动-语言评估器进行评分,但这些评估器对语言结构的接地程度仍不清楚。在此,我们引入了通过时间顺序、镜像反射和身份诊断评估(STRIDE)的结构接地基准,以系统评估评估器跟踪时间顺序、镜像反射和动作身份的能力。STRIDE 包含 5,869 个三元组,每个三元组由一个运动、其原始描述和一个扰动描述组成,涵盖短描述和长描述。我们对描述对进行似然平衡以减少纯文本偏差,并估计每个评估器在不相关运动下的描述偏好,以衡量匹配运动相对于该基线的判别增益。我们的实验揭示了受审查的评估器中存在薄弱的结构接地和严重的镜像敏感性缺陷,而常用的评估协议未能暴露这些问题。为了理解为什么这些局限性在标准测试中未被检测到,我们更仔细地检查了这些评估器。我们发现,仅凭纯文本先验就能解决现有数据集上的朴素扰动测试,而常见的检索和分布度量对通过镜像真实运动引入的结构破坏几乎没有反应。这些发现提示了一种自然的干预方法:结构硬负样本。我们的实验表明,对对比学习进行简单修改即可显著提高时间顺序和镜像反射方面的性能。基准和代码将发布。
cs.CV / 115 / 2609.32477

Back-Tracking from Clarity: Self-Learning to See Text from Afar

从清晰回溯:自学习远距离文本检测
Tran, Duc-Tri, Nguyen, Phi Le, Hoai, Minh
Abstract
We propose a self-supervised framework designed to enhance the capability of scene text detectors in identifying and recognizing text in scenarios where instances are shown at significant distances, typically small, blurred, and frequently missed by conventional models. Our approach leverages the high-fidelity performance of existing text spotting models on large, clear text as a foundational supervisor. By temporally back-tracking these high-confidence detections through video sequences, we automatically synthesize pseudo-labels for preceding frames where the distant text is still visually degraded or undersized. These pseudo-labels enable training a student model specialized for early text detection, without requiring any manual annotation. The success of this approach depends on accurate pseudo-label generation, for which we develop a dedicated scene text tracker capable of maintaining consistent text identities across challenging video sequences. In addition, we propose SceneText50, a diverse multilingual outdoor dataset to facilitate training and evaluation. Experiments show that our framework significantly improves early detection accuracy and robustness across varied scenes and languages. Code and data are at \href{https://github.com/trid2912/BackTrackingText}{https://github.com/trid2912/BackTrackingText}.
Chinese Translation
我们提出了一种自监督框架,旨在增强场景文本检测器在实例距离较远、通常较小、模糊且常被传统模型漏检的场景中识别和辨认文本的能力。我们的方法利用现有文本定位模型在大型清晰文本上的高保真性能作为基础监督器。通过在视频序列中时间回溯这些高置信度检测,我们自动为远距离文本仍然视觉退化或尺寸过小的前序帧合成伪标签。这些伪标签使得能够训练一个专门用于早期文本检测的学生模型,而无需任何手动标注。该方法的成功取决于准确的伪标签生成,为此我们开发了一个专用的场景文本跟踪器,能够在具有挑战性的视频序列中保持一致的文本身份。此外,我们提出了SceneText50,一个多样化的多语言户外数据集,以促进训练和评估。实验表明,我们的框架显著提高了在不同场景和语言下的早期检测精度和鲁棒性。代码和数据见\href{https://github.com/trid2912/BackTrackingText}{https://github.com/trid2912/BackTrackingText}。
cs.CV / 116 / 2609.32481

JEPA Learns What the Mask Leaves Unrecoverable

JEPA 学习掩码留下的不可恢复内容
Xie, Peng, Alanwar, Amr
Abstract
Joint-embedding predictive architectures are unusually sensitive to how the input is masked: block masks work, scattered masks do not, and the explanations are empirical. We give a measurement account. A mask is a linear measurement, and in a compactly supported wavelet basis every atom whose support lies inside the hidden region falls in the measurement's null space and leaves no trace in the data. The JEPA loss asks only that the encoded context suffice for the target, so a target that a low-level prior can recover admits a shortcut, one the moving-average target encoder can make self-consistent. What removes the shortcut is the coarse-scale content the mask leaves unrecoverable, provided enough context stays within reach of each target. We score that content before training and test the account's distinctive predictions in 151 pre-training runs. On ImageNet-100, strip masks match blocks in area and contiguity yet are recoverable, and they land at 40.3% linear top-1, beside random masks at 40.8%, against 64.3% for blocks; within one geometry family, the placements that leave the least unrecoverable content lose 6.5 points to those that leave the most, over five seed pairs; pixel targets span 7 points where latent targets span 25; and against a frozen target the gap between random and block masks, 19 points on the same kind of GPU, closes to 1.5, so the geometry acts through the target the encoder produces for itself. On UCF101 the masking ratio decides which condition, content or reach, binds; removing whole frames, unrecoverable in space but recoverable from neighbouring frames, is worst at both ratios; and on V-JEPA's own masks, batching them intact instead of truncated changes little (36.0% against 35.1%), whereas making 100 target tokens inside the blocks visible lifts them to 48.7% and hiding 100 context tokens outside the blocks does not (33.7%).
Chinese Translation
联合嵌入预测架构(Joint-embedding predictive architectures)对输入被掩码的方式异常敏感:块状掩码有效,散乱掩码无效,而这些解释是经验性的。我们给出一种测量解释。掩码是一种线性测量,在紧支撑小波基中,每个支撑位于隐藏区域内的原子都落入测量的零空间,并在数据中不留下任何痕迹。JEPA 损失只要求编码的上下文足以预测目标,因此一个低层先验能够恢复的目标就存在捷径,而移动平均目标编码器可以使这种捷径自洽。消除这种捷径的是掩码留下的不可恢复的粗尺度内容,前提是每个目标都有足够的上下文可达。我们在训练前对该内容进行评分,并在151次预训练运行中测试该解释的独特预测。在 ImageNet-100 上,条状掩码在面积和连续性上与块状掩码相当,但却是可恢复的,它们在线性 top-1 上达到 40.3%,与随机掩码的 40.8% 相近,而块状掩码为 64.3%;在同一个几何家族内,留下最少不可恢复内容的放置比留下最多内容的放置低 6.5 个百分点,在五对随机种子上;像素目标跨度为 7 个点,而潜在目标跨度为 25 个点;并且在与冻结目标对比时,随机掩码与块状掩码之间的差距(在同类 GPU 上为 19 个点)缩小到 1.5,因此几何形状通过编码器自身产生的目标起作用。在 UCF101 上,掩码比例决定了哪个条件(内容或可达性)起约束作用;移除整个帧,在空间上不可恢复但可以从相邻帧恢复,在两种比例下都是最差的;而在 V-JEPA 自己的掩码上,将它们完整分批而不是截断,变化很小(36.0% 对 35.1%),而让块内 100 个目标 token 可见可将其提升到 48.7%,隐藏块外 100 个上下文 token 则没有提升(33.7%)。
cs.CV / 117 / 2609.32510

GeoCR: Learning a Generalist Cloud Removal Prior from Heterogeneous Observations

GeoCR:从异构观测中学习通用云去除先验
Do, Jeonghyeok, Kim, Munchurl
Abstract
Cloud removal methods are typically specialized to individual datasets and input configurations, limiting reuse across sensors, spectral bands, and observation settings. We introduce GeoCR, a generalist model that unifies RGB-only-based CR and multispectral-based CR from single- or multi-temporal cloudy observations, with optional SAR guidance, within a single network. To accommodate different spectral and sensing domains, compact input and output stems extend a pretrained RGB autoencoder while keeping its encoder and decoder trunks frozen. This shared latent interface enables a single flow transformer to jointly model clean RGB and non-RGB latents, conditioned on separate cloudy-observation streams and optional SAR tokens. Through joint pretraining on the training splits of ten datasets comprising 883,331 cloud-free target images, GeoCR learns a shared cloud removal prior across these heterogeneous configurations. The same pretrained checkpoint supports direct inference without dataset-specific fine-tuning and efficient adaptation through low-rank adaptation (LoRA). We evaluate GeoCR against general image restoration and cloud removal methods on test splits of the contributing datasets under full-band and RGB-only settings. GeoCR achieves the best FID and DISTS on full-band SEN12MS-CR and Sen2_MTC_New and RGB-only CUHK-CR2, outperforming existing models and demonstrating the effectiveness of a reusable generative model across diverse settings.
Chinese Translation
云去除方法通常针对单个数据集和输入配置进行专门设计,限制了跨传感器、光谱波段和观测设置的复用。我们提出 GeoCR,一个通用模型,在单一网络中统一了基于纯 RGB 的云去除(CR)和基于多光谱的云去除,可处理单时相或多时相含云观测,并可选 SAR 引导。为了适应不同的光谱和传感域,紧凑的输入和输出 stem 模块扩展了一个预训练的 RGB 自编码器,同时保持其编码器和解码器主干冻结。这种共享潜在接口使单个 flow transformer 能够联合建模干净的 RGB 和非 RGB 潜在表示,并以单独的含云观测流和可选 SAR token 为条件。通过在十个数据集的训练划分上进行联合预训练,这些数据集包含 883,331 张无云目标图像,GeoCR 在这些异构配置中学习到一个共享的云去除先验。同一个预训练检查点支持无需数据集特定微调的直接推理,并可通过低秩适应(LoRA)进行高效适配。我们在全波段和仅 RGB 设置下,在贡献数据集的测试划分上,将 GeoCR 与通用图像恢复和云去除方法进行评估比较。GeoCR 在全波段 SEN12MS-CR 和 Sen2_MTC_New 以及仅 RGB 的 CUHK-CR2 上取得了最佳 FID 和 DISTS,优于现有模型,并证明了可复用生成模型在不同设置中的有效性。
cs.CV / 118 / 2609.32513

Cross-Domain Few-Shot Writer Adaptation for Real-World Handwritten Mathematical Expression Recognition

面向真实世界手写数学表达式识别的跨域少样本书写者自适应
Silva, Paulo Grane Gabriel, Marqueses, Lorenz Bernard, Ilao, Joel
Abstract
Handwritten mathematical expression recognition (HMER) refers to the task of recognizing and converting handwritten mathematics into a parsable markup language, usually LaTeX. No current state-of-the-art-competitive system adjusts to the way a specific person writes, and the domain gap between training images (usually digital or perfectly binarized) and images physically taken with a camera used in inference has received fairly little attention for this specific problem. We characterize this domain gap through fragmentation and stroke-width analyses of the images as well as introduce a writer-adaptive fine-tuning pipeline to MFH-CoMER in an attempt to address it. We further introduce sample author-specific datasets, consisting of five handwriting category subsets from two authors, and evaluate using a McNemar's test and permutation tests adapted to limited data. Results suggest an increase in expression recognition rate and a decrease in CER for digital handwriting but more varied results for physical handwriting, with the model struggling for handwriting articles that the base model can already evaluate well. Nonetheless, adapted models were found to have improved results for four out of five subsets. Statistical testing results suggest a consistency in improvement for two of the tested author-specific subsets. Our results point towards the potential feasibility of writer adaptation for the HMER task.
Chinese Translation
手写数学表达式识别(HMER)是指将手写数学识别并转换为可解析的标记语言(通常为LaTeX)的任务。目前尚无具备最先进竞争力的系统能够适应特定个人的书写方式,并且对于这一特定问题,训练图像(通常为数字图像或完美二值化图像)与推理时使用相机实际拍摄的图像之间的域差异受到的关注相当少。我们通过对图像的碎片化和笔画宽度分析来刻画这种域差异,并尝试通过在MFH-CoMER中引入书写者自适应微调流程来解决这一问题。我们进一步引入了作者特定的样本数据集,该数据集由来自两位作者的五个手写类别子集组成,并使用适用于有限数据的McNemar检验和置换检验进行评估。结果表明,对于数字手写,表达式识别率提高且CER降低,但对于物理手写,结果更为多变,模型在处理基础模型已经能够很好评估的手写文章时表现不佳。尽管如此,发现自适应模型在五个子集中的四个上取得了改进的结果。统计检验结果表明,在两个所测试的作者特定子集上改进具有一致性。我们的结果表明,书写者自适应对于HMER任务具有潜在可行性。
cs.CV / 119 / 2609.32518

UnStep: Training-Free Acceleration of Causal Video Diffusion with Fewer Steps Than Distillation

UnStep:以比蒸馏更少步骤实现因果视频扩散的无训练加速
Mansour, Youssef, Simsar, Enis, Sener, Fadime, Georgopoulos, Markos, Pumarola, Albert, Thabet, Ali, Schoenfeld, Edgar
Abstract
Distilling bidirectional multi-step video diffusion transformers into few-step causal models has become a common approach for streaming video generation. While these few-step students are significantly faster than the teachers they are distilled from, they remain slow for real-time generation. In this work we present UnStep, a training-free wrapper that accelerates few-step causal video models at inference by running them with fewer diffusion transformer (DiT) steps than during distillation and limiting the temporal window retained in the attention KV cache. We propose two inference-only mechanisms to recover quality lost by step reduction and attention windowing: renoising the generated latent frames to a near-clean level and reusing the existing clean-cache pass to refine them, and applying truncated SVD to the DiT attention value and output projections. We also accelerate inference with a quality-preserving runtime stack for the DiT and VAE decoder, including more efficient attention calls and KV indexing, fused Triton RoPE with cached coefficients, and VAE decoding with optimized memory layout, precision, and convolution kernels. By reducing computation and optimizing the runtime stack, UnStep sets a new throughput regime for causal video diffusion, by running substantially faster than current methods, reaching 50 FPS on a single H100 without quality loss, and 77 FPS on GB200, all without retraining.
Chinese Translation
将双向多步视频扩散Transformer蒸馏为少步因果模型已成为流式视频生成的常见方法。虽然这些少步学生模型比它们蒸馏自的教师模型显著更快,但对于实时生成而言仍然较慢。在这项工作中,我们提出了UnStep,一个无需训练的包装器,通过在推理时以比蒸馏期间更少的扩散Transformer(DiT)步骤运行少步因果视频模型,并限制注意力KV缓存中保留的时间窗口,来加速这些模型。我们提出了两种仅用于推理的机制来恢复因步骤减少和注意力窗口化而损失的质量:将生成的潜在帧重新加噪至接近干净水平,并重用现有的干净缓存通道来细化它们,以及对DiT注意力值和输出投影应用截断SVD。我们还通过为DiT和VAE解码器使用保持质量的运行时栈来加速推理,包括更高效的注意力调用和KV索引、带有缓存系数的融合Triton RoPE,以及具有优化内存布局、精度和卷积核的VAE解码。通过减少计算和优化运行时栈,UnStep为因果视频扩散设定了新的吞吐量标准,运行速度显著快于当前方法,在单个H100上达到50 FPS且无质量损失,在GB200上达到77 FPS,所有这些都无需重新训练。
cs.CV / 120 / 2609.32529

SetOPD: From Few Visual Exemplars to Multimodal Candidate Sets for Remote-Sensing Open-Prompt Detection

SetOPD:从少量视觉示例到多模态候选集用于遥感开放提示检测
Hu, Jinlong, Zhang, Yi, Xia, Zhiqi, Zhou, Yikang, Ji, Shunping
Abstract
Open-prompt detectors allow users to specify targets with text, visual exemplars, or both. We argue that existing designs underuse the visual modality in two ways. First, multiple exemplars are commonly compressed into a single class-level embedding. This textualizes visual prompting: the resulting vector plays the role of another class name, is often aligned to or injected into the text pathway, and may be suboptimal when only a few heterogeneous exemplars are available. Second, existing methods interact primarily in prompt or representation space, before modality-specific detection states are formed. We address both issues from a set perspective: \setopd preserves modality-specific decoding states from a shared prompt-conditioned initialization and performs explicit multimodal collaboration at the candidate-state level. For the first issue, we introduce \br prompting, which reads every boxed exemplar in its full scene context and decomposes the pooled support evidence into a Base anchor and a learned Residual correction; the resulting prompt has fixed capacity regardless of the number of examples and drives its own visual detection pathway. For the second, we recast text--visual collaboration from representation-level fusion into a candidate-set modeling problem. Paired-Query Arbitration (\pqa) then performs explicit cross-modal state arbitration only after modality-specific candidate states have been formed. The two readers share query initialization so their candidates are paired by index; a learned gate arbitrates within each pair, followed by a permutation-equivariant module that reasons over the fused set.
Chinese Translation
开放提示检测器允许用户通过文本、视觉示例或二者兼用来指定目标。我们认为现有设计在两个方面未充分利用视觉模态。首先,多个示例通常被压缩为单个类别级嵌入。这使视觉提示文本化:得到的向量扮演另一个类别名称的角色,通常被对齐到或注入文本路径中,并且当仅有少量异构示例可用时可能次优。其次,现有方法主要在提示或表示空间中进行交互,在模态特定的检测状态形成之前。我们从集合视角解决这两个问题:SetOPD 从共享的提示条件初始化中保留模态特定的解码状态,并在候选状态级别执行显式的多模态协作。针对第一个问题,我们引入 BR 提示,它在完整的场景上下文中读取每个带框示例,并将池化的支持证据分解为基础锚点(Base anchor)和学习到的残差校正(Residual correction);得到的提示具有固定容量,与示例数量无关,并驱动其自身的视觉检测路径。针对第二个问题,我们将文本-视觉协作从表示级融合重新定义为候选集建模问题。配对查询仲裁(Paired-Query Arbitration, PQA)仅在模态特定的候选状态形成之后,执行显式的跨模态状态仲裁。两个读取器共享查询初始化,因此它们的候选按索引配对;一个学习到的门控在每个对内部进行仲裁,随后是一个置换等变模块,对融合后的集合进行推理。
cs.CV / 121 / 2609.32534

DepthBench: Measuring How Residual Connections Enable More Computational Depth

DepthBench:衡量残差连接如何实现更多计算深度
Wang, Keyu, Huang, Yangyi, Kang, Jiale, González-Martínez, David, Liu, Weiyang, Liu, Shiwei
Abstract
Depth is a natural way to increase the computational capacity in Transformers, yet the contribution of deeper layers can diminish as depth grows larger. Recent approaches enhance normalization (\text{e.g.}, LayerNorm Scaling) or residual connections (\text{e.g.}, mHC, AttnRes) to enable better information flow and depth utilization. However, it remains unclear whether they truly translate increased architectural depth into effective computational depth, and whether their reported gains stem from better access to information across depth, or unaccounted-for confounding factors. In this paper, we introduce \textbf{DepthBench}, a controlled benchmark for studying computational depth across various architectures. We systematically vary the width--depth aspect ratio ($d_{\text{model}}/n_{\text{layer}}$) from shallow--wide to deep--narrow shapes, while keeping the model size and pre-training recipe fixed. Across 10 representative architectures, we find that the benefit of allocating more capacity to depth is strongly architecture-dependent. Standard Pre-LN and most of its norm- and scaling-based variants provide little benefit and can even degrade performance as models become deeper and narrower, whereas HC and Full AttnRes improve consistently even at extreme deep shapes. These gains extend beyond pre-training loss and consistently translate into improved domain-specific performance and effective computation. Controlled layer-level analyses further show that the gains of HC and Full AttnRes are associated with more effective utilization of additional layers, revealing distinct mechanisms of computational depth across architectures. Overall, our results identify residual connection design as a key determinant of whether depth can serve as a meaningful scaling axis by enabling additional architectural depth to translate into effective computation.
Chinese Translation
深度是增加Transformer计算能力的一种自然方式,然而随着深度增大,更深层网络的贡献可能会减弱。近期方法通过增强归一化(例如,LayerNorm Scaling)或残差连接(例如,mHC、AttnRes)来实现更好的信息流动和深度利用。然而,尚不清楚它们是否真正将增加的架构深度转化为有效的计算深度,以及它们所报告的增益是源于跨深度信息的更好获取,还是源于未考虑的混杂因素。在本文中,我们引入了DepthBench,一个用于研究跨多种架构的计算深度的受控基准。我们系统性地改变宽度-深度比($d_{\text{model}}/n_{\text{layer}}$),从浅宽形状到深窄形状,同时保持模型大小和预训练配方固定。在10种代表性架构上,我们发现将更多容量分配给深度的收益强烈依赖于架构。标准Pre-LN及其大多数基于归一化和缩放的变体几乎不提供收益,甚至当模型变得更深更窄时性能会下降,而HC和Full AttnRes即使在极深的形状下也能持续改善。这些增益不仅限于预训练损失,还持续转化为改进的领域特定性能和有效计算。受控的层级分析进一步表明,HC和Full AttnRes的增益与更有效地利用额外层相关,揭示了不同架构间计算深度的不同机制。总体而言,我们的结果确定了残差连接设计是深度能否作为有意义缩放轴的关键决定因素,它使额外的架构深度能够转化为有效计算。
cs.CV / 122 / 2609.32537

RIPE-MambaSpike: Resolution-Independent Spiking-State-Space Interfaces for Parameter-Efficient Event-Based Vision

RIPE-MambaSpike:面向参数高效事件视觉的分辨率无关脉冲状态空间接口
Islam, Md Muhiminul, Dipu, Shoaib Ahmed, Chowdhury, Sayeed Shafayet
Abstract
Spiking-Mamba hybrids reach strong accuracy on event-based vision, but existing designs often require tens of millions of parameters. Much of that cost comes from how the spiking front-end is connected to the state-space backbone rather than from the hybrid architecture itself. In a representative model, a single resolution-dependent projection accounts for 33.55M of 36.25M parameters. To that end, we introduce RIPE-MambaSpike (Resolution-Independent, Parameter-Efficient), which replaces that projection with a hierarchical multi-resolution bridge of fixed channel width. Its deployed footprint is 0.870M parameters, constant at fixed time steps and widths across a 43x range of input areas. Reparameterized spiking stages, temporal decoupled modulation, and a dynamic convex-hull-bounded dual-stream membrane-potential attention preserve accuracy under this compact design. Result-wise, RIPE-MambaSpike is pareto-optimal on CIFAR10-DVS, N-Caltech101, and DailyDVS-200. Notably, on the 200-class DailyDVS-200, a scaled 8.04M configuration achieves 45.7% top-1 accuracy, the best reported spiking result on that benchmark, and outperforms prior spiking methods with 3.0-15.1x fewer parameters than dense ANNs. Overall, our findings demonstrate that competitive event-based recognition does not require resolution-dependent parameter growth. Code is available at https://github.com/MuhiminOsim/RIPE-MambaSpike.
Chinese Translation
脉冲-Mamba混合模型在事件视觉上达到了很强的准确率,但现有设计通常需要数千万参数。这种开销很大部分来自脉冲前端与状态空间骨干的连接方式,而不是来自混合架构本身。在一个代表性模型中,单个与分辨率相关的投影占了36.25M参数中的33.55M。为此,我们提出了RIPE-MambaSpike(分辨率无关、参数高效),它用固定通道宽度的分层多分辨率桥接替换了该投影。其部署占用为0.870M参数,在固定的时间步和宽度下,跨越43倍输入面积范围保持恒定。重参数化脉冲阶段、时间解耦调制以及动态凸包约束的双流膜电位注意力在这一紧凑设计中保持了准确率。在结果方面,RIPE-MambaSpike在CIFAR10-DVS、N-Caltech101和DailyDVS-200上达到了帕累托最优。值得注意的是,在200类的DailyDVS-200上,一个扩展的8.04M配置达到了45.7%的top-1准确率,是该基准上已报告的最佳脉冲结果,并且以比密集ANN少3.0-15.1倍的参数优于先前的脉冲方法。总体而言,我们的发现表明,有竞争力的事件识别并不需要随着分辨率增长参数。代码可在https://github.com/MuhiminOsim/RIPE-MambaSpike获取。
cs.CV / 123 / 2609.32540

In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion

面向更快自回归视频扩散的带有干净锚点的在途 KV 缓存
Wang, Yikai, Han, Xiao, Xu, Mengmeng, Perez, Juan Camilo, Douratsos, Yiannis, He, Sen, Zhou, Zijian, Zhang, Fei, An, Zhaochong, Perez-Rua, Juan-Manuel, Loy, Chen Change, Xiang, Tao
Abstract
Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less-noisy key--value (KV) cache by additional forwards to build the cache without advancing an output latent. However, every denoising forward itself already computes the in-flight KV of the current chunk. We introduce FlashForward, which directly reuses this cache to avoid the heavy cache-update-only model forwards. After the current chunk completes one denoising stage, its stage-specific cache is already available for the next chunk. Assigning one GPU to each stage therefore lets different chunks occupy different stages concurrently. This early availability has a quality cost: the resulting stage-matched history is noisy, causing appearance and motion drift among chunks. To complement it, FlashForward produces sparse auxiliary clean anchor latents before the corresponding region is generated so the generation trajectories can be stabilized by this two-sided conditioning. The two memories operate at different temporal scales: sparse clean anchor KV supplies coarse, long-range two-sided structural guidance, while dense stage-matched history preserves fine, recent evolution. With up to four GPUs, FlashForward runs $1.16$--$1.69\times$ faster than HiAR and $1.42$--$2.92\times$ faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbone scales at 480p and 720p. On VBench, for the 1.3B model at 480p, it achieves higher scores and remains stable at longer durations, demonstrating that FlashForward generates high-quality and temporally consistent videos across durations of 20s, 35s and 65s at a much faster generation speed.
Chinese Translation
少步自回归视频扩散通过将视频分割成时间块并逐块生成来生成长视频,每个块通过一系列简短的去噪阶段生成。为了记住已经生成的块,先前的方法通过额外的前向传播重建一个干净或噪声较少的键值(KV)缓存,以构建缓存而不推进输出潜变量。然而,每个去噪前向传播本身已经计算了当前块的在途 KV。我们引入了 FlashForward,它直接重用该缓存,以避免繁重的仅更新缓存的模型前向传播。在当前块完成一个去噪阶段后,其特定阶段的缓存已经可用于下一个块。因此,为每个阶段分配一个 GPU 可以让不同的块同时占据不同的阶段。这种早期可用性会带来质量代价:所得到的阶段匹配历史是有噪声的,导致块之间的外观和运动漂移。为了弥补这一点,FlashForward 在生成相应区域之前生成稀疏的辅助干净锚点潜变量,从而可以通过这种双侧条件作用来稳定生成轨迹。这两种记忆在不同的时间尺度上运行:稀疏的干净锚点 KV 提供粗粒度的、长程的双侧结构引导,而密集的阶段匹配历史则保留精细的、近期的演变。在使用最多四个 GPU 的情况下,对于 20 秒或更长的 16 FPS 视频,在 480p 和 720p 分辨率下,FlashForward 在 1.3B 和 14B 骨干规模上比 HiAR 快 1.16–1.69 倍,比 Self-Forcing 快 1.42–2.92 倍。在 VBench 上,对于 480p 的 1.3B 模型,它取得了更高的分数,并在更长的时长下保持稳定,表明 FlashForward 以更快的生成速度在 20 秒、35 秒和 65 秒的时长范围内生成了高质量且时序一致的视频。
cs.CV / 124 / 2609.32551

Harnessing Coupled Stream Completion For Human-Object Ineraction Modeling

利用耦合流补全进行人-物交互建模
Guan, Dawei, Yang, Di, Wang, Jiangtao
Abstract
Text-conditioned human-object interaction (HOI) generation requires body motion, object trajectories & rotations, and hand articulation to remain coordinated. These components differ in scale and dynamics, but must agree on contact, relative pose, and timing. A shared representation may limit the distinct structure of each stream, while independent generation prevents each stream from responding to changes in the others. Latent supervision alone also does not directly constrain contact after decoding. We propose TRACE, a continuous latent framework that keeps stream states separate and couples their updates. TRACE encodes body, object, and hand motion into separate latents and predicts each stream velocity from the complete current interaction state. Geometric losses on decoded motion further constrain contact and object-relative motion over time. The same model supports completion of any single absent stream from the other two. Frozen flow features also serve as input to a language model for HOI understanding. Experiments on InterAct, OMOMO, and BEHAVE show that joint completion training improves generation and that frozen flow features improve understanding over raw-motion encoding. On InterAct, TRACE achieves the highest contact precision, recall, and F1 among the compared methods.
Chinese Translation
文本条件的人-物交互(HOI)生成要求身体运动、物体轨迹与旋转以及手部关节运动保持协调一致。这些组成部分在尺度和动力学上存在差异,但必须在接触、相对姿态和时序上保持一致。共享表示可能会限制每个流的独特结构,而独立生成则阻止每个流响应其他流的变化。仅靠潜在空间监督也无法在解码后直接约束接触。我们提出 TRACE,一个连续的潜在框架,它将流状态保持分离并耦合其更新。TRACE 将身体、物体和手部运动编码为独立的潜在变量,并从完整的当前交互状态预测每个流的速度。对解码运动施加几何损失进一步约束了随时间变化的接触和物体相对运动。同一模型支持从其他两个流补全任意单个缺失流。冻结的流特征也可作为语言模型的输入,用于人-物交互理解。在 InterAct、OMOMO 和 BEHAVE 上的实验表明,联合补全训练改进了生成,并且冻结的流特征相比原始运动编码提升了理解性能。在 InterAct 上,TRACE 在比较方法中取得了最高的接触精度、召回率和 F1 分数。
cs.CV / 125 / 2609.32555

Feature Space Guidance for Breast Cancer Classification in DCE-MRI

用于DCE-MRI乳腺癌分类的特征空间引导
Hamm, Benjamin, Kirchhoff, Yannick, Rokuss, Maximilian, Langenberg, Moritz, Ulrich, Constantin, Wald, Tassilo, Traub, Jeremias, Gotkowski, Karol, Maier-Hein, Klaus
Abstract
Dynamic contrast enhanced breast MRI (DCE-MRI) is a powerful clinical tool for breast cancer detection, providing high resolution anatomical detail together with rich temporal contrast information. However, high dimensional 4D inputs, small lesions, and heterogeneous acquisition protocols across clinical sites hinder robust automated classification of healthy, benign, and malignant cases. To address these challenges, we propose a framework that dynamically analyzes latent representations to adapt to protocol-specific characteristics. Spatial variability is mitigated by reducing confounding background uptake and compensating for misalignment caused by deformable soft tissue. Additionally, relationships in the latent space across phases are leveraged to select the most informative temporal features, improving robustness to protocol-specific temporal variability. Finally, task specific discriminative features are promoted through large scale supervised lesion segmentation pretraining, which substantially enhances downstream finetuning. Evaluated under leave-one-center-out validation on the ODELIA dataset and the held-out AMBL cohort, the proposed framework substantially outperforms finetuned radiology foundation models and prior methods, improving mean AUROC by nearly 8 points and balanced accuracy by 4 points over the strongest baseline. Additionally, our method achieved first place in the MICCAI ODELIA Breast MRI Challenge 2025, further demonstrating its effectiveness for robust breast cancer classification. We publicly release our codebase under https://github.com/MIC-DKFZ/CURIAtor.
Chinese Translation
动态对比增强乳腺MRI(DCE-MRI)是乳腺癌检测的强大临床工具,可提供高分辨率解剖细节以及丰富的时间对比信息。然而,高维4D输入、小病灶以及跨临床机构的异质性采集协议阻碍了对健康、良性和恶性病例的稳健自动分类。为应对这些挑战,我们提出了一个动态分析潜在表示以适应特定协议特征的框架。通过减少混杂的背景摄取并补偿由可变形软组织引起的不对齐,空间变异性得以缓解。此外,利用跨期潜在空间中的关系来选择最具信息量的时间特征,从而提高了对特定协议时间变异性的鲁棒性。最后,通过大规模监督病灶分割预训练促进任务特定的判别特征,从而显著增强了下游微调。在ODELIA数据集和保留的AMBL队列上通过留一中心验证评估,所提出的框架显著优于微调的放射学基础模型和先前方法,相比最强基线,平均AUROC提高了近8个点,平衡准确率提高了4个点。此外,我们的方法在MICCAI ODELIA乳腺MRI挑战赛2025中获得第一名,进一步证明了其在稳健乳腺癌分类方面的有效性。我们公开了我们的代码库,地址为https://github.com/MIC-DKFZ/CURIAtor。
cs.CV / 126 / 2609.32559

CFCH: Coarse-Fine Collaborative Hierarchical Learning for Anterior Segment Disease Analysis

CFCH:用于前节疾病分析的粗细协同层次学习
Wang, Peng, Zou, Haohan, Wu, Yanlin, Xie, Xueshuo, Wang, Yan, Li, Tao
Abstract
Accurate classification of anterior segment diseases is crucial for ophthalmic screening and diagnosis. However, slit-lamp image analysis remains challenging due to substantial variability in imaging conditions and the intrinsic anatomical-disease hierarchy of ocular pathologies. Existing methods typically formulate this task as a flat multi-class classification problem, ignoring the structured dependency between anatomical regions (e.g., cornea, conjunctiva, and lens) and disease manifestations.To address these limitations, we propose CFCH, a Coarse-Fine Collaborative Hierarchical learning framework that explicitly models anatomical context and disease semantics through a dual-branch architecture. To enable effective cross-granularity collaboration, CFCH introduces semantic and cross-granularity attention consistency constraints, encouraging aligned yet complementary feature learning across branches. In addition, we construct AS-9K, a large-scale anterior segment dataset with 8975 images covering 12 common disease categories. To the best of our knowledge, AS-9K is the largest publicly available dataset for anterior segment image classification. Extensive experiments on two anterior segment datasets demonstrate that CFCH outperforms state-of-the-art methods. Qualitative visualizations further show more focused and lesion-relevant activation responses, validating the effectiveness of the proposed framework. Code will be available at https://github.com/ybupengwang/CFCH.
Chinese Translation
前节疾病的准确分类对于眼科筛查和诊断至关重要。然而,由于成像条件的显著差异以及眼部病理固有的解剖-疾病层次结构,裂隙灯图像分析仍然具有挑战性。现有方法通常将此任务视为扁平的多类分类问题,忽略了解剖区域(如角膜、结膜和晶状体)与疾病表现之间的结构化依赖关系。为了解决这些局限性,我们提出了CFCH,一种粗细协同层次学习框架,通过双分支架构显式地建模解剖上下文和疾病语义。为了实现有效的跨粒度协作,CFCH引入了语义和跨粒度注意力一致性约束,鼓励跨分支的对齐但互补的特征学习。此外,我们构建了AS-9K,一个大规模前节数据集,包含8975张图像,涵盖12种常见疾病类别。据我们所知,AS-9K是最大的公开可用的前节图像分类数据集。在两个前节数据集上的大量实验表明,CFCH优于最先进的方法。定性可视化进一步显示了更聚焦且与病变相关的激活响应,验证了所提框架的有效性。代码将在https://github.com/ybupengwang/CFCH上提供。
cs.CV / 127 / 2609.32567

Attribution Gaps in Zero-Training LLM+OVOD Pipelines: A Fine-Grained Analysis of the CAAP--SNAP Discrepancy

零训练 LLM+OVOD 流水线中的归因差距:对 CAAP--SNAP 差异的细粒度分析
Yen, Yu-Feng
Abstract
LAOD and similar zero-training LLM+open-vocabulary-detector (OVOD) pipelines score two things separately: class-agnostic localization accuracy (CAAP) and semantic naming accuracy (SNAP). The two consistently diverge, and nobody has asked why. This paper asks why, on the full 5,000-image COCO-Val split (27,273 detections) rather than the small subset the original work evaluated on. Object visual complexity turns out not to be the driver -- small and occluded objects are, if anything, localized better than large ones. Vocabulary novelty is: once the LLM's wording falls outside the detector's native category set, localization accuracy falls from 80.9% to 31.6%. That drop is not spread evenly across unfamiliar phrasing, though. Almost all of it comes from cases where the novel wording actually names a different object than the one COCO annotated (true synonyms still score 89.3%; semantically unrelated "noise" labels score 12.0%). A closer look at a further failure subset tells a similar story: 78-88% of what looks like complete localization failure is really the model correctly finding a real object that COCO's non-exhaustive 80-category scheme simply never labeled, not hallucination. Swap the detector backbone (YOLO-World for Grounding DINO) or the LLM (Gemma-3 for Qwen2.5-VL) and both the effect and its rough size hold up, so this looks like a general property of the pipeline family rather than a quirk of one model pairing. The upshot is that a large share of the apparent CAAP--SNAP gap traces back to closed-category annotation limits rather than a real grounding failure, which matters for how we detect hallucination, analyze failure modes, and design evaluation for grounded multimodal systems meant to work in the open world.
Chinese Translation
LAOD 及类似的零训练 LLM+开放词汇检测器(OVOD)流水线分别对两项指标进行评分:类别无关定位精度(CAAP)和语义命名精度(SNAP)。二者始终存在分歧,但尚无人追问原因。本文在完整的 5,000 张图像的 COCO-Val 划分(27,273 个检测结果)上探究其原因,而非原工作评估所用的小规模子集。物体视觉复杂度并非驱动因素——小目标和被遮挡目标(如果说有什么区别的话)的定位反而优于大目标。词汇新颖性才是:一旦 LLM 的措辞超出检测器原生类别集,定位精度便从 80.9% 降至 31.6%。然而,这一下降并非均匀分布在所有不熟悉的表述上。其几乎全部来自以下情况:新颖措辞实际命名的物体与 COCO 标注的物体不同(真正的同义词仍得分 89.3%;语义无关的“噪声”标签得分 12.0%)。进一步审视另一个失败子集也讲述着类似的故事:看似完全定位失败的情况中,78-88% 实际上是模型正确找到了 COCO 非穷尽的 80 类方案从未标注的真实物体,而非幻觉。更换检测器骨干(用 YOLO-World 替换 Grounding DINO)或 LLM(用 Gemma-3 替换 Qwen2.5-VL),该效应及其大致幅度依然成立,因此这似乎是该流水线家族的普遍属性,而非某一模型组合的怪癖。其结论是,表观 CAAP--SNAP 差距的很大一部分可追溯至封闭类别标注的局限,而非真正的接地失败,这对我们如何检测幻觉、分析失败模式,以及为旨在开放世界中工作的接地多模态系统设计评估,都具有重要意义。
cs.CV / 128 / 2609.32569

OmniSmartHome: A Multimodal Reasoning Benchmark for Smart-Home Agents

OmniSmartHome:面向智能家居代理的多模态推理基准
Jung, Jihoo, Yoo, Suho, Choi, Jeongsoo, Cho, Hyebin, Haam, Tae Wook, Ryu, Hyeonggon, Park, Sumin, Chung, Joon Son
Abstract
Smart-home assistants are expected to handle diverse, realistic requests that arise in daily life. In such interactions, users often rely on the surrounding multimodal context-pointing at objects or referring to what they see or hear, leaving their requests underspecified in language alone. Existing smart-home benchmarks, however, express user requests solely through language, leaving context-dependent real-world requests underexplored. To bridge this gap, we introduce OmniSmartHome, a multimodal smart-home benchmark where each spoken request is paired with the surrounding visual and spatial-audio context, providing complementary cues to disambiguate underspecified requests. OmniSmartHome comprises 1,360 synthetic and 272 real-world episodes. We evaluate 16 omnimodal large language models (Omni-LLMs) and reveal that, while they perform strongly when speech alone sufficiently conveys the user's intent, performance drops substantially when resolving it requires reasoning over multimodal contextual cues. As a simple agent baseline, we provide PROME (PROcedural Memory for multimodal Evidence gathering), which equips agents with specialized audio-visual perception tools and procedural memory for orchestrating their use. PROME generally improves performance across six Omni-LLMs. Demos and examples are available at https://omni-smart-home.github.io
Chinese Translation
智能家居助手需要处理日常生活中出现的多样化、现实请求。在此类交互中,用户往往依赖周围的多模态上下文——指向物体或指代所见所闻,使得仅凭语言无法充分明确其请求。然而,现有的智能家居基准仅通过语言表达用户请求,导致依赖上下文的现实世界请求未被充分探索。为弥补这一空白,我们提出了 OmniSmartHome,一个多模态智能家居基准,其中每个语音请求都与周围的视觉和空间音频上下文配对,提供互补线索以消除不明确请求的歧义。OmniSmartHome 包含 1,360 个合成片段和 272 个真实世界片段。我们评估了 16 个全模态大语言模型 (Omni-LLMs),并揭示:当仅凭语音足以传达用户意图时,它们表现强劲;但当需要基于多模态上下文线索进行推理才能解决时,性能大幅下降。作为一个简单的代理基线,我们提供了 PROME(用于多模态证据收集的程序性记忆),它为代理配备了专门的视听感知工具和程序性记忆,以协调其使用。PROME 在六个 Omni-LLMs 上普遍提升了性能。演示和示例可在 https://omni-smart-home.github.io 获取。
cs.CV / 129 / 2609.32590

Retrieved but Not Delivered: Multimodal Memory Delivery for Long-Term Agents

检索但未传递:面向长期智能体的多模态记忆传递
Jiang, Yuhang, Liao, Qingwei, Yin, Kaize, Liu, Xingling, Cuomo, Luca, Bacci, Silvio
Abstract
Work on memory for multimodal agents optimizes what is written, updated and retrieved. Between retrieval and the answer, however, is a stage that multimodal memory evaluations do not isolate: what of the retrieved memory reaches the model, and in what form. We call it delivery, and a controlled decomposition on MemLens locates the remaining room there. With the retrieved evidence set exactly fixed, delivering the original pixels instead of withholding them raises accuracy by 13.87 points on an 8B backbone, whereas making retrieval perfect on those same messages improves it by 2.31. Delivery is the larger term on all three MemLens backbones and grows with backbone strength; retrieval grows too, without closing the gap. We propose DeliverMem, an instantiation of delivery as three decisions: keep the original modality, give each item a readable identity, and state when it was seen, with a retrieval-side adapter for the one property delivery cannot supply. Each is measured against a delivery-matched control that alters only its own variable. DeliverMem leads the strongest published memory agent on MemLens at all four context lengths, and beats DMV-Bench's own strongest method at every setting on both backbones. On MemLens it does this on a tenth to a seventieth of the input. Each decision helps only where the question lacks what it supplies, and is null elsewhere. A single fixed configuration nonetheless leads both benchmarks, without training any component or modifying the stored records. Project page: https://avalon-s.github.io/DeliverMem/
Chinese Translation
面向多模态智能体的记忆研究优化了写入、更新和检索的内容。然而,在检索与答案之间,存在一个多模态记忆评估未能分离的阶段:检索到的记忆中有多少到达模型,以及以何种形式到达。我们将其称为传递(delivery),在 MemLens 上的受控分解表明剩余提升空间就在于此。在检索到的证据集完全固定的情况下,传递原始像素而不是不传递它们,在 8B 骨干模型上将准确率提高了 13.87 个百分点;而在相同消息上使检索完美,仅提高了 2.31 个百分点。在所有三个 MemLens 骨干模型上,传递都是更大的贡献项,并且随着骨干模型能力的增强而增大;检索也会增大,但未能缩小差距。我们提出 DeliverMem,将传递实例化为三个决策:保留原始模态,为每个条目赋予可读标识,并说明其被看到的时间;同时用一个检索端适配器来补充传递无法提供的那个属性。每个决策都通过与一个仅改变其自身变量的、与传递匹配的对照进行比较来评估。在 MemLens 上,DeliverMem 在全部四种上下文长度下均领先于已发表的最强记忆智能体,并且在两个骨干模型上的所有设置下都击败了 DMV-Bench 自身的最强方法。在 MemLens 上,它仅用十分之一到七十分之一的输入就实现了这一点。每个决策仅在问题缺乏其提供的信息时才有帮助,在其他情况下则无效果。然而,一个单一的固定配置就能在两个基准上领先,且无需训练任何组件或修改存储记录。项目页面:https://avalon-s.github.io/DeliverMem/
cs.CV / 130 / 2609.32592

SPACE: Sparse Predictive Attractor via Counterfactual Eviction for Streaming Video Memory

SPACE:用于流式视频记忆的通过反事实驱逐的稀疏预测吸引子
Niu, Hongjin, Zhang, Weizhan, Bao, Shuo, Wang, Jiahao, Jiao, Muyan, Wen, Kairui, Liu, Yong-Jin
Abstract
Fixed-capacity streaming video memory requires repeated eviction decisions whose effects accumulate over time. Yet existing policies are evaluated primarily in terms of retained information or downstream accuracy, leaving how repeated updates alter the futures supported by memory largely unexamined. We define a memory's predictive state as the future representations supported by its retained history and formulate eviction as counterfactual control over transitions in this space. We introduce SPACE (Sparse Predictive Attractor via Counterfactual Eviction), which uses a frozen multi-horizon JEPA to predict the future representations induced by alternative eviction actions. Counterfactual utility identifies future-useful alternatives, while slow predictive-basin geometry determines when to correct avoidable drift and when to adapt to sustained predictive change, without online parameter updates. We further introduce MABS-Bench, which evaluates future-task sufficiency, within-regime predictive stability, transition responsiveness, and perturbation recovery under matched causal streams and memory budgets. Across multiple video datasets, SPACE yields consistent improvements in dataset-native task performance while reducing predictive-state drift.
Chinese Translation
固定容量的流式视频记忆需要重复的驱逐决策,其影响随时间累积。然而,现有策略主要根据保留的信息或下游准确性进行评估,而重复更新如何改变记忆所支持的未来则基本未经考察。我们将记忆的预测状态定义为其保留历史所支持的未来表示,并将驱逐表述为对该空间中转换的反事实控制。我们引入SPACE(通过反事实驱逐的稀疏预测吸引子),它使用一个冻结的多视野JEPA来预测由替代驱逐动作引起的未来表示。反事实效用识别出对未来有用的替代方案,而缓慢的预测盆地几何结构决定了何时纠正可避免的漂移,以及何时适应持续的预测变化,而无需在线参数更新。我们进一步引入MABS-Bench,它在匹配的因果流和记忆预算下评估未来任务充分性、模式内预测稳定性、转换响应性和扰动恢复。在多个视频数据集上,SPACE在数据集原生任务性能上取得了一致的改进,同时减少了预测状态漂移。
cs.CV / 131 / 2609.32596

GAUGE: Group-Wise View-Inconsistency Rectification for Feed-Forward 4D Tracking

GAUGE:面向前馈4D跟踪的分组视图不一致性校正
Feng, Zhuoqian, Chen, Weixing, Chen, Ziliang, Liu, Yang, Lin, Liang
Abstract
Feed-forward models regress dense 3D point trajectories directly from monocular video, yet the residual after global alignment is substantial and lacks a structural explanation. Measured on dynamic query points across models and datasets, the error concentrates along the view direction, while the scale correction each motion group requires differs. The predicted displacement direction nevertheless supports reliable grouping, with a median angle far below the 90{\deg} random baseline. The systematic part of the residual is therefore a family of radial degrees of freedom per motion group, along directions 2D observations cannot constrain. We call it group-wise view inconsistency. We present GAUGE (Group-wise Adaptive Unsupervised Gauge Estimation), a training-free and model-agnostic post-hoc module. It recovers motion groups from direction consistency and spatial connectivity, then estimates a per-frame radial scale and group-level translation from 1% to 5% metric anchors, four degrees of freedom per group and frame. On dynamic query points of eight trackers, including D4RT, 4RC and SM4RT, our correction lowers endpoint error by 15.1% to 62.6% over the uncorrected predictions, while spending the same anchors on gradient fine-tuning improves the same models by only -1.1% to 15.2%. Code is publicly available at https://github.com/HCPLab-SYSU/GAUGE.
Chinese Translation
前馈模型直接从单目视频回归密集3D点轨迹,但全局对齐后的残差仍然显著,且缺乏结构性解释。在跨模型和数据集的动态查询点上测量发现,误差集中在视图方向,而每个运动组所需的尺度校正各不相同。尽管如此,预测的位移方向仍能支持可靠的分组,其中位角度远低于90°的随机基线。因此,残差的系统性部分对应于每个运动组的一族径向自由度,这些方向是2D观测无法约束的。我们将其称为分组视图不一致性。我们提出GAUGE(Group-wise Adaptive Unsupervised Gauge Estimation,分组自适应无监督规范估计),一种无需训练且模型无关的事后模块。它从方向一致性和空间连通性中恢复运动组,然后从1%到5%的度量锚点中估计逐帧径向尺度和组级平移,每组每帧四个自由度。在八个跟踪器(包括D4RT、4RC和SM4RT)的动态查询点上,我们的校正相比未校正预测将端点误差降低了15.1%至62.6%,而将相同锚点用于梯度微调时,对相同模型的改进幅度仅为-1.1%至15.2%。代码已在 https://github.com/HCPLab-SYSU/GAUGE 公开。
cs.CV / 132 / 2609.32612

Levy-Driven Correspondence Estimation for Registration

Levy驱动的配准对应估计
Wu, Qianliang, Yang, Jiaqi, Yang, Wankou, Hui, Le, Xie, Jin, Yang, Jian, Ding, Yaqing
Abstract
Finding reliable point correspondences is difficult when point clouds have low overlap or undergo non-rigid deformation. Iterative refinement can correct uncertain matches, but costly network evaluations limit the number of updates. We present LevyMatch, a L\'evy-driven method that uses random jumps to refine a soft matching matrix. At each step, a network uses the current matching state and geometric information to predict a target matching matrix. A Brownian reference bridge gives an explicit formula for the update toward this target. A Gamma random clock sets the time step for each update. The updated matches provide new geometric feedback for the next target prediction. We further propose a fixed front-loaded Gamma policy that assigns more expected clock time to early updates and less to later ones, without retraining or extra network evaluations. Reordering the same sampled Gamma increments shows that placing larger increments early gives higher accuracy than placing them late. On 4DMatch and 4DLoMatch, our method improves both non-rigid feature matching recall (NFMR) and inlier ratio (IR) over the compared methods. The front-loaded policy achieves 93.09% NFMR and 92.11% IR on 4DMatch, and 82.79% NFMR and 79.07% IR on 4DLoMatch.
Chinese Translation
当点云重叠度低或经历非刚性变形时,找到可靠的点对应是困难的。迭代细化可以纠正不确定的匹配,但昂贵的网络评估限制了更新次数。我们提出LevyMatch,一种Lévy驱动的方法,利用随机跳跃来细化软匹配矩阵。在每一步,网络利用当前匹配状态和几何信息来预测目标匹配矩阵。一个Brownian参考桥给出了朝该目标更新的显式公式。一个Gamma随机时钟为每次更新设置时间步。更新后的匹配为下一次目标预测提供新的几何反馈。我们进一步提出一种固定的前载Gamma策略,将更多的期望时钟时间分配给早期更新,较少的分配给后期更新,而无需重新训练或额外的网络评估。对相同采样的Gamma增量进行重新排序表明,将较大的增量放在早期比放在后期获得更高的精度。在4DMatch和4DLoMatch上,我们的方法在非刚性特征匹配召回率(NFMR)和内点比率(IR)上均优于所比较的方法。前载策略在4DMatch上达到93.09%的NFMR和92.11%的IR,在4DLoMatch上达到82.79%的NFMR和79.07%的IR。
cs.CV / 133 / 2609.32628

DraftAttention2: Fast Video Diffusion with Low-Resolution-Guided Mixed-Precision Attention

DraftAttention2:基于低分辨率引导的混合精度注意力的快速视频扩散
Ding, Rui, Li, Haopeng, Ma, Weize, Zhou, Yufa, Li, Yitong, Zhou, Xiaoling, Cao, Jiashuo, Li, Liyang, Geng, Hua, Gu, Jiuxiang, Lin, Jun, Xie, Enze, Shen, Xuan
Abstract
Video generation has broad applications in content creation and entertainment. Diffusion transformers have advanced the quality of generated videos, but attention over spatiotemporal tokens becomes increasingly expensive as video resolution and duration increase. We present DraftAttention2, a training-free framework that uses the low-resolution draft attention map to jointly select attention blocks and assign their numerical precision. Specifically, spatial 2D average- and max-pooled queries and keys capture complementary regional statistics to estimate block importance, and a shared ranking assigns higher precision to important blocks, lower precision to less important retained blocks, and skips the rest under configurable budgets. Our analysis separates sparsification error from attention-weighted quantization error, establishing when recovering skipped interactions with low-bit computation tightens the output-error bound. This analysis motivates retaining more interactions at low precision while reserving higher precision for blocks with larger attention mass. To translate these fine-grained assignments into practical speedups, we further develop fused operand preparation and a single attention kernel with precision-specific phases, sharing data movement, softmax statistics, and output accumulation across precisions. Experiments demonstrate that our method achieves a superior quality-efficiency trade-off over existing efficient video generation methods. Notably, its advantage is particularly pronounced for few-step video diffusion, where jointly combining sparsity with 4- and 8-bit mixed-precision computation substantially improves generation quality while retaining significant acceleration. Code is available at https://github.com/anemoi-project/anemoi
Chinese Translation
视频生成在内容创作和娱乐领域有广泛的应用。扩散Transformer提升了生成视频的质量,但随着视频分辨率和时长的增加,对时空token的注意力计算变得愈发昂贵。我们提出DraftAttention2,一个无需训练的框架,利用低分辨率草稿注意力图来联合选择注意力块并分配其数值精度。具体而言,空间2D平均池化和最大池化的查询和键捕获互补的区域统计信息以估计块的重要性,然后一个共享的排序将较高精度分配给重要块,将较低精度分配给不太重要的保留块,并在可配置的预算下跳过其余块。我们的分析将稀疏化误差与注意力加权的量化误差分开,确定了何时用低位计算恢复跳过的交互可以收紧输出误差界。这一分析促使我们在低精度下保留更多交互,同时为注意力质量更大的块保留更高精度。为了将这些细粒度的分配转化为实际的加速,我们进一步开发了融合的操作数准备和具有精度特定阶段的单一注意力内核,在不同精度之间共享数据移动、softmax统计量和输出累积。实验表明,我们的方法在质量-效率权衡上优于现有的高效视频生成方法。值得注意的是,其优势在少步视频扩散中尤为明显,其中将稀疏性与4位和8位混合精度计算相结合,在保持显著加速的同时大幅提高了生成质量。代码可在 https://github.com/anemoi-project/anemoi 获取。
cs.CV / 134 / 2609.32660

InterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table Reasoning

InterTab: 面向多模态表格推理的交错视觉-结构对齐
Li, Hanqian, Huang, Sirui, Ling, Chen, Li, Jungang, Huang, Yu, Zheng, Kening, Hei, Yonghua, He, Xiangrong, Wang, Shiyi, Zhu, Pengcheng, Liu, Dongnan, Zhou, Wei, Mo, Linjian, Ding, Nai, Hu, Xuming
Abstract
Table images preserve structural information that are often lost in text serialization, and reasoning over them requires locating relevant rows, columns, and cells step by step. Current multimodal large language models (MLLMs) encode the whole image once before reasoning, so they cannot pick up row-, column-, and cell-level evidence as the question unfolds. Encoder-side table structure and generic interleaved visual chain-of-thought still do not bind each reasoning step to that structure. We propose \textbf{InterTab}, an \textbf{Inter}leaved structure-aware framework for CoT reasoning over \textbf{Tab}le images, interleaves chain-of-thought with tool calls that crop structure-aligned table regions. First, we build InterTab-22K, includes reasoning trajectories in which each step is tied to both a structural location and a bounding box. InterTab is trained in two stages: supervised structure-aware alignment (SSA) on InterTab-22K teaches the model to interleave reasoning with structure-aligned crops, and active localization optimization (ALO) further optimizes answer correctness, localization IoU, and output format, while penalizing missing or excessive tool calls. Experiments on nine table benchmarks show that InterTab improves the average accuracy of its backbone from 68.28% to 73.17% and achieves the best average performance among all compared methods. Code and data will be released soon.
Chinese Translation
表格图像保留了在文本序列化中经常丢失的结构信息,而对其进行推理需要逐步定位相关的行、列和单元格。当前的多模态大语言模型(MLLMs)在推理前一次性编码整个图像,因此无法随着问题的展开捕捉行、列和单元格级别的证据。编码器侧的表格结构和通用的交错视觉思维链仍然没有将每个推理步骤绑定到该结构。我们提出了 InterTab,一个用于表格图像上思维链推理的交错结构感知框架,它将思维链与裁剪结构对齐的表格区域的工具调用交错进行。首先,我们构建了 InterTab-22K,其中包含推理轨迹,每个步骤都与结构位置和边界框相关联。InterTab 分两个阶段训练:在 InterTab-22K 上进行监督结构感知对齐(SSA),教导模型将推理与结构对齐的裁剪交错进行;主动定位优化(ALO)进一步优化答案正确性、定位 IoU 和输出格式,同时惩罚缺失或过多的工具调用。在九个表格基准上的实验表明,InterTab 将其骨干网络的平均准确率从 68.28% 提高到 73.17%,并在所有比较方法中实现了最佳平均性能。代码和数据将很快发布。
cs.CV / 135 / 2609.32681

RCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous Driving

RCVLA:面向自动驾驶的4D雷达接地语义推理与轨迹仲裁
Zheng, Lianqing, Bai, Xiaokai, Luo, Yixuan, Guan, Runwei, Liu, Minghao, Wei, Zhiqiang, Shen, Hui-liang, Zhu, Xichan, Ma, Zhixiong
Abstract
4D radar provides geometric and motion cues that complement visual semantics, but integrating it into vision-language-action (VLA) models requires both radar--language alignment for semantic reasoning and explicit use of radar measurements for trajectory refinement and selection. To support these capabilities, we construct Cap4DR with 86,016 radar-image-text samples for alignment pretraining and OmniHD-QA with 520,161 question-answer pairs for instruction tuning across scene description, key-object reasoning, occupancy understanding, and trajectory planning. Building on these datasets, we propose RCVLA, a radar-camera VLA framework consisting of a radar-grounded semantic reasoning stage (RCVLA-Sem) and a trajectory arbitration stage (RCVLA-Phys). RCVLA-Sem performs gated bidirectional interaction between camera and radar tokens for driving question answering and reference trajectory generation, while auxiliary heads provide object and occupancy queries. RCVLA-Phys refines reference-guided trajectory candidates through truncated diffusion conditioned on these queries and cluster-level radar measurements, then calibrates candidate scores using radar-derived time-to-collision risk. On OmniHD-QA, RCVLA-Sem improves CIDEr by 9.92 points and reduces key-object velocity error by $21.9\%$ relative to OmniDrive. RCVLA-Phys further reduces average L2 error from $0.348$ to $0.259\,\mathrm{m}$ and average open-loop collision rate from $0.576\%$ to $0.175\%$ relative to RCVLA-Sem. Ablation studies further show that language-aligned radar tokens improve semantic reasoning, while cluster-level radar measurements and risk calibration improve trajectory arbitration. Code will be released.
Chinese Translation
4D雷达提供了补充视觉语义的几何和运动线索,但将其集成到视觉-语言-动作(VLA)模型中需要雷达-语言对齐以进行语义推理,并显式使用雷达测量以进行轨迹优化和选择。为了支持这些能力,我们构建了Cap4DR(包含86,016个雷达-图像-文本样本用于对齐预训练)和OmniHD-QA(包含520,161个问答对用于指令微调,涵盖场景描述、关键目标推理、占用理解和轨迹规划)。基于这些数据集,我们提出了RCVLA,一个雷达-相机VLA框架,由雷达接地的语义推理阶段(RCVLA-Sem)和轨迹仲裁阶段(RCVLA-Phys)组成。RCVLA-Sem在相机和雷达令牌之间执行门控双向交互,用于驾驶问答和参考轨迹生成,而辅助头提供目标查询和占用查询。RCVLA-Phys通过截断扩散对这些查询和聚类级雷达测量进行条件化,从而优化参考引导的轨迹候选,然后使用雷达推导的碰撞时间风险来校准候选评分。在OmniHD-QA上,与OmniDrive相比,RCVLA-Sem将CIDEr提高了9.92点,并将关键目标速度误差降低了21.9%。相对于RCVLA-Sem,RCVLA-Phys进一步将平均L2误差从0.348 m降低到0.259 m,并将平均开环碰撞率从0.576%降低到0.175%。消融实验进一步表明,语言对齐的雷达令牌改善了语义推理,而聚类级雷达测量和风险校准改善了轨迹仲裁。代码将开源。
cs.CV / 136 / 2609.32690

MM-OPD: Towards One More Bottleneck Between Perception and Reasoning

MM-OPD:迈向感知与推理之间的又一瓶颈
Tong, Jintao, Lou, Yujing, Shen, Zhanming, Gu, Jiaqi, Fan, Lubin, Li, Ruixuan, Wu, Yue, Ye, Jieping, Zou, Yixiong
Abstract
Recent multimodal large language models (MLLMs) advance visual reasoning by strengthening both perception and reasoning, implicitly assuming a process that transitions seamlessly from perception to reasoning. However, we observe a counterintuitive phenomenon that challenges this assumption: holding the model, question, and decoding fixed, we replace images with their caption or code representations (symbolic views), which seems to be redundant given the clear image structures, but the performance surprisingly improves by 10.2% to 23.6% across model scales and datasets. We term this performance gap as the Symbolic Visual Gap and then take a closer look at it. Through experiments, we find that although the visual evidence can already appear in the reasoning trace for the image-input model, the symbolic-view-input model shows much higher attention to the correct evidence than the image-input model. This suggests that despite good capabilities from current works in perception and reasoning themselves, another bottleneck exists between perception and reasoning in selecting perceived visual information as appropriate evidence for subsequent reasoning. To handle this bottleneck, since the symbolic view steers attention toward correct evidence and is readily obtained at scale, it provides supervision for evidence selection without manually labeled evidence. Building on this, we introduce MM-OPD, a multimodal on-policy self-distillation framework for symbolic-to-visual correction that transfers guidance from symbolic-conditioned behavior to the image-conditioned policy through residual token-level targets, steering the model toward correct visual evidence. Experiments across benchmarks and model scales show that MM-OPD improves a broad range of multimodal abilities, with gains in visual perception, chart and document understanding, mathematical reasoning, and general VQA.
Chinese Translation
最近的多模态大语言模型(MLLMs)通过增强感知和推理来推进视觉推理,隐含地假设了一个从感知到推理无缝过渡的过程。然而,我们观察到一种与直觉相反的现象,挑战了这一假设:在固定模型、问题和解码的情况下,我们用图像的描述或代码表示(符号视图)替换图像,尽管鉴于清晰的图像结构这似乎是冗余的,但性能在不同模型规模和数据集上惊人地提升了10.2%到23.6%。我们将这种性能差距称为符号视觉差距(Symbolic Visual Gap),然后对其进行深入探究。通过实验,我们发现,尽管对于图像输入模型,视觉证据已经可以出现在推理轨迹中,但符号视图输入模型对正确证据的注意力远高于图像输入模型。这表明,尽管当前工作在感知和推理本身方面具有良好的能力,但在感知和推理之间还存在另一个瓶颈,即选择感知到的视觉信息作为后续推理的适当证据。为了解决这一瓶颈,由于符号视图将注意力引向正确的证据并且易于大规模获取,它提供了对证据选择的监督,而无需人工标注证据。在此基础上,我们提出了MM-OPD,一个多模态同策略自蒸馏框架,用于符号到视觉的校正,通过残差标记级目标将符号条件行为的指导传递给图像条件策略,引导模型朝向正确的视觉证据。跨基准和模型规模的实验表明,MM-OPD提升了广泛的多模态能力,在视觉感知、图表和文档理解、数学推理以及通用视觉问答方面都有所增益。
cs.CV / 137 / 2609.32705

DPAMixerSR: An Efficient Degradation-Pattern-Aware Model for Image Super-Resolution

DPAMixerSR:一种用于图像超分辨率的高效退化模式感知模型
Wu, Song-Li, Jiang, Haonan, Fan, Jixuan, Huo, Yufei, Zhang, Chubin, Tang, Yansong
Abstract
While content-adaptive schemes have delivered notable advances in image super-resolution (SR), existing approaches typically focus on texture complexity and ignore intrinsic degradation factors (e.g., blur kernels or noise patterns), leading to suboptimal computation allocation and reconstruction performance. To remedy this, we propose DPAMixerSR, a degradation-pattern-aware framework that enables efficient SR through adaptive sparse computation. We design a lightweight Perceptual Degradation Ranking (PDR) module partitions the image into severely and mildly degraded patches, which are routed to the Adaptive Sparse Processing (ASP) and a lightweight convolutional branch, respectively. ASP performs structure-aligned, multi-scale sparse propagation and bidirectional refinement, while the convolutional branch enhances efficiency in mildly degraded regions. By coupling degradation-driven routing with structure-aligned sparse processing, DPAMixerSR establishes a self-regulating framework that dynamically balances computational efficiency and reconstruction fidelity. Extensive experiments on various SR tasks demonstrate that our DPAMixerSR achieves superior structural restoration and perceptual fidelity with markedly reduced computational overhead, providing a novel and scalable framework for degradation-aware, resource-efficient SR.
Chinese Translation
尽管内容自适应方案在图像超分辨率(SR)方面取得了显著进展,但现有方法通常关注纹理复杂度,而忽略了内在退化因素(例如,模糊核或噪声模式),导致计算分配和重建性能次优。为了解决这个问题,我们提出了DPAMixerSR,一个退化模式感知框架,通过自适应稀疏计算实现高效SR。我们设计了一个轻量级的感知退化排序(PDR)模块,将图像划分为严重退化和轻微退化的块,分别路由到自适应稀疏处理(ASP)和轻量级卷积分支。ASP执行结构对齐的多尺度稀疏传播和双向细化,而卷积分支提高了轻微退化区域的效率。通过将退化驱动的路由与结构对齐的稀疏处理相结合,DPAMixerSR建立了一个自调节框架,动态平衡计算效率和重建保真度。在各种SR任务上的大量实验表明,我们的DPAMixerSR以显著降低的计算开销实现了优越的结构恢复和感知保真度,为退化感知、资源高效的SR提供了一个新颖且可扩展的框架。
cs.CV / 138 / 2609.32711

ProDyGS: Dynamic Gaussian Splatting from a Single Static Monocular Camera

ProDyGS:基于单个静态单目相机的动态高斯泼溅
Cavalcanti, Ugo Leone, Tosi, Fabio, Poggi, Matteo, Conti, Andrea, Zlokolica, Vladimir, Cambareri, Valerio, Mattoccia, Stefano
Abstract
We present ProDyGS, a novel dynamic 3D Gaussian Splatting framework for high-quality novel view synthesis from videos captured by a single static camera. While existing methods rely on multi-view setups or significant camera motion for geometric constraints, our approach addresses the challenging scenario where multi-view supervision is completely absent. We overcome this limitation by generating synthetic multi-view supervision through depth-guided proxy image synthesis. Specifically, we estimate temporally consistent depth maps using foundational monocular depth networks, then construct 3D Gaussian representations that generate proxy images from arbitrary viewpoints. A deformation network learns temporal dynamics by warping canonical Gaussians using this augmented supervision. Experiments on the DyNeRF dataset demonstrate that our method achieves state-of-the-art performance while requiring only monocular depth estimation as external supervision, outperforming approaches that rely on stronger priors such as scene flow.
Chinese Translation
我们提出了 ProDyGS,一种新颖的动态 3D 高斯泼溅框架,用于从单台静态单目相机拍摄的视频中实现高质量的新视角合成。现有方法依赖多视角设置或显著的相机运动来提供几何约束,而我们的方法则针对多视角监督完全缺失的挑战性场景。我们通过深度引导的代理图像合成来生成合成的多视角监督,从而克服这一限制。具体而言,我们使用基础单目深度网络估计时间一致的深度图,然后构建能够从任意视点生成代理图像的三维高斯表示。一个变形网络利用这种增强监督对规范高斯进行变形,从而学习时间动态。在 DyNeRF 数据集上的实验表明,我们的方法在仅需要单目深度估计作为外部监督的情况下达到了最先进的性能,并优于依赖场景流等更强先验的方法。
cs.CV / 139 / 2609.32716

Region-Local Copula Evidence Fusion for Heterogeneous Remote Sensing Change Detection

面向异质遥感变化检测的区域局部Copula证据融合
Ji, Zhiyuan, Yin, Junjun, Yang, Jian
Abstract
Superpixel copula models provide stable regional evidence for heterogeneous remote sensing change detection, but a single label per region limits localization within mixed superpixels. This letter develops a region-local copula evidence fusion method that retains the regional decision structure while introducing spatially varying local dependence anomalies. Independently fitted local models characterize departures from unchanged cross-image relationships. Reference ranking and an upper-tail gate transform these anomalies for fusion with continuous regional confidence. We derive the resulting regiondependent local decision threshold and identify a condition under which gating is equivalent to reparameterizing ungated fusion. On Lake and UK, whole-image optimized configurations achieve kappa coefficients of 0.78136 and 0.90817 and improve mixedregion and boundary decisions. Four-fold retrospective spatial validation over ten training subsets confirms complementary local information, with ungated reference fusion increasing mean kappa by 0.00693 and 0.01793. Fixed gating yields a larger UK gain of 0.03353 but only 0.00041 on Lake. These results support regional-local dependence interaction, while showing that calibration and gating have scene-dependent benefits.
Chinese Translation
超像素Copula模型为异质遥感变化检测提供了稳定的区域证据,但每个区域单一标签限制了混合超像素内的定位。本文提出一种区域局部Copula证据融合方法,该方法在保留区域决策结构的同时,引入了空间变化的局部依赖异常。独立拟合的局部模型刻画了对未变化的跨图像关系的偏离。参考排序和上尾门控将这些异常转换,以便与连续区域置信度融合。我们推导了由此产生的区域相关局部决策阈值,并确定了门控等价于重新参数化无门控融合的条件。在Lake和UK数据集上,全图优化配置分别达到了0.78136和0.90817的kappa系数,并改善了混合区域和边界决策。对十个训练子集的四折回顾性空间验证证实了互补的局部信息,无门控参考融合使平均kappa分别提高了0.00693和0.01793。固定门控在UK上产生了更大的增益(0.03353),但在Lake上仅提高了0.00041。这些结果支持区域-局部依赖交互,同时表明校准和门控具有场景依赖的益处。
cs.CV / 140 / 2609.32734

REALIS: A Curated Dataset for Studying the Challenges of AI Image Detection

REALIS:一个用于研究 AI 图像检测挑战的精选数据集
Gushchin, Aleksandr, Abud, Khaled, Bychkov, Georgii, Shumitskaya, Ekaterina, Filippov, Artem, Lavrushkin, Sergey, Vatolin, Dmitriy S., Antsiferova, Anastasia
Abstract
AI-generated image detectors are often evaluated on benchmarks where real and synthetic images differ in content, quality, or generation artifacts, allowing models to rely on dataset-specific cues and fail on unfamiliar generators or processed images. Existing datasets provide limited support for evaluating these challenges jointly across diverse visual content. We introduce REALIS, a dataset of 1.43 million real and synthetic images generated by 42 modern text-to-image models, including the latest proprietary systems such as Nano Banana 2. REALIS combines prompts derived from real images, quality filtering, and stratified sampling to reduce class-specific shortcuts while preserving content diversity. We further introduce REALIS-Expert, a stress-test subset for high-quality synthetic images, where real and generated samples are selected with closely matched semantic and visual characteristics. We also propose a robustness protocol covering 35 transformations at five severity levels to analyze detector behavior under image processing. Based on REALIS, our benchmark evaluates pretrained detectors, fine-tuned models, and zero-shot vision-language models under generator and post-processing shifts. On the hardest processed split, the best pretrained conventional detector achieves 0.550 ROC-AUC, compared with 0.752 for the best REALIS-trained detector. REALIS provides a unified framework for measuring and improving the reliability of AI-image detectors under conditions that better reflect real-world use.
Chinese Translation
AI 生成的图像检测器通常在基准上进行评估,这些基准中真实图像和合成图像在内容、质量或生成伪影上存在差异,使得模型能够依赖数据集特定的线索,并在不熟悉的生成器或处理过的图像上失败。现有数据集在跨多样视觉内容联合评估这些挑战方面提供的支持有限。我们提出了 REALIS,一个包含 143 万张真实和合成图像的数据集,这些图像由 42 个现代文本到图像模型生成,包括最新的专有系统,如 Nano Banana 2。REALIS 结合了从真实图像中派生的提示、质量过滤和分层采样,以减少类别特定的捷径,同时保持内容多样性。我们进一步引入了 REALIS-Expert,一个针对高质量合成图像的压力测试子集,其中真实样本和生成样本在语义和视觉特征上高度匹配。我们还提出了一个鲁棒性协议,涵盖五个严重级别的 35 种变换,以分析图像处理下的检测器行为。基于 REALIS,我们的基准评估了预训练检测器、微调模型和零样本视觉语言模型在生成器和后处理偏移下的表现。在最难的处理分割上,最佳的预训练传统检测器达到了 0.550 的 ROC-AUC,而最佳的 REALIS 训练检测器达到了 0.752。REALIS 提供了一个统一框架,用于在更好反映真实世界使用的条件下衡量和提高 AI 图像检测器的可靠性。
cs.CV / 141 / 2609.32737

Gradient-Guided Decoupled Adaptation for Geospatial Vision-Language Models

地理空间视觉语言模型的梯度引导解耦适配
Wang, Dongdong, Balakrishnan, Deepak, Srinivasan, Ravi, Wang, Shenhao
Abstract
Existing geospatial vision-language models (Geo-VLMs) typically optimize diverse geospatial tasks through a unified multi-task adaptation paradigm without explicitly accounting for the heterogeneous optimization characteristics. Our empirical observations reveal heterogeneous gradient characteristics across tasks, including vision-language differences, intra-branch gradient relationships, and task interference, which hinder effective multi-task optimization. Motivated by these observations, we propose Gradient-Guided Decoupled Adaptation (G2DA), a gradient-aware optimization framework for multi-task Geo-VLM learning. G2DA first partitions tasks into vision- and language-centric groups through gradient-guided cross-modal decoupling. It then constructs modality-specific curricula based on task gradient similarity and employs bidirectional rehearsal to mitigate the recency effects introduced by sequential optimization. We evaluate G2DA on three Geo-VLM benchmarks using six InternVL3 and Qwen3.5-VL variants, along with GeoChat and GeoLLaVA. Across all 24 benchmark-model combinations, G2DA consistently outperforms representative baselines, improving over the strongest competitor by 3.08, 4.30, and 2.81 percentage points on UrBench-MCQ, XLRS-Bench-Lite, and VRS-Bench-VQA, respectively. These results demonstrate the effectiveness of gradient-guided task organization for Geo-VLM adaptation.
Chinese Translation
现有的地理空间视觉语言模型(Geo-VLMs)通常通过统一的多任务适配范式来优化多样化的地理空间任务,而没有明确考虑异构的优化特性。我们的实证观察揭示了任务间异构的梯度特性,包括视觉-语言差异、分支内梯度关系以及任务干扰,这些特性阻碍了有效的多任务优化。受这些观察的启发,我们提出了梯度引导解耦适配(G2DA),一种用于多任务Geo-VLM学习的梯度感知优化框架。G2DA首先通过梯度引导的跨模态解耦将任务划分为以视觉为中心和以语言为中心的两组。然后,它基于任务梯度相似性构建模态特定的课程,并采用双向复演来缓解顺序优化引入的近因效应。我们在三个Geo-VLM基准上使用六个InternVL3和Qwen3.5-VL变体以及GeoChat和GeoLLaVA评估了G2DA。在所有24个基准-模型组合上,G2DA持续优于代表性基线,在UrBench-MCQ、XLRS-Bench-Lite和VRS-Bench-VQA上分别比最强竞争者提高了3.08、4.30和2.81个百分点。这些结果证明了梯度引导的任务组织对Geo-VLM适配的有效性。
cs.CV / 142 / 2609.32740

AnesTRACE: Benchmarking Intraoperative Anesthesia from Multimodal Perception to Multi-step Decision-Making

AnesTRACE:从多模态感知到多步决策的术中麻醉基准测试
Huang, Ziwei, Gao, Qi, Ji, Zhe, Yao, Yuanyuan, Zhang, Fengjiang, Yan, Min, Xie, Zhongle, Chen, Gang
Abstract
Intraoperative anesthesia requires systems to interpret evolving multimodal evidence, recommend timely management, and revise decisions as patient states change, yet existing benchmarks usually isolate perception or single-point reasoning. We introduce AnesTRACE, an evaluation suite comprising AnesTRACE-Bench and AnesTRACE-Eval. Built from public perioperative datasets with anesthesiologist annotation, AnesTRACE-Bench evaluates Intraoperative Perception, Single-point Anesthesia Decision-Making, and Multi-step Anesthesia Decision-Making. AnesTRACE-Eval assesses open-ended responses through anesthesiologist-defined criteria for Clinical Correctness, Evidence Grounding, Task Completeness, and Safety, with Temporal Consistency for multi-step decisions; its domain-specific evaluator is trained by supervised fine-tuning and preference alignment on expert-reviewed judgments. Across more than 30 models, fine-grained visual grounding and intervention selection remain difficult: the leading model reaches only 32.2 mIoU for TEE visual grounding and retains a 17.5\% Major/Critical Safety Error Rate in multi-step management. Evaluator alignment with anesthesiologists improves across both training stages, while the best decision quality is accompanied by a 74.3-second P95 Latency. These results show that aggregate performance alone does not establish safe, timely longitudinal decision-making. We release our code at https://github.com/zjuDBxAI/AnesTRACE.
Chinese Translation
术中麻醉要求系统能够解读不断变化的多模态证据,及时推荐管理措施,并随着患者状态的变化修正决策,然而现有的基准测试通常孤立地评估感知或单点推理。我们提出了 AnesTRACE,一个包含 AnesTRACE-Bench 和 AnesTRACE-Eval 的评估套件。AnesTRACE-Bench 基于带有麻醉医师标注的公开围手术期数据集构建,评估术中感知、单点麻醉决策和多步麻醉决策。AnesTRACE-Eval 通过麻醉医师定义的临床正确性、证据基础、任务完整性和安全性标准评估开放式回答,并针对多步决策增加时间一致性;其领域特定的评估器通过在专家审查的判断上进行监督微调和偏好对齐来训练。在超过 30 个模型上,细粒度视觉基础和干预选择仍然困难:领先模型在 TEE 视觉基础方面仅达到 32.2 mIoU,并且在多步管理中仍有 17.5% 的重大/严重安全错误率。评估器与麻醉医师的一致性在两个训练阶段均有所提高,而最佳决策质量伴随着 74.3 秒的 P95 延迟。这些结果表明,仅凭总体性能并不能确保安全、及时的纵向决策。我们在 https://github.com/zjuDBxAI/AnesTRACE 发布了我们的代码。
cs.CV / 143 / 2609.32742

LoCoVSR: Local Context Diffusion Posterior Sampling for Video Super-Resolution

LoCoVSR:用于视频超分辨率的局部上下文扩散后验采样
Chorin, Matan Ben, Elad, Michael
Abstract
Video super-resolution (VSR) is an ill-posed inverse problem that aims to reconstruct a high-resolution (HR) video from a noisy, low-resolution (LR) version of it. We present LoCoVSR, a diffusion-based VSR framework that leverages pixel-space denoising diffusion probabilistic models. LoCoVSR integrates the Diffusion Posterior Sampling technique with spatio-temporal context learning, operating in a moving-average form. A localized window of adjacent LR frames is used for recovering each center frame, while applying a shared noise trajectory across all frames. The localized windowing enables processing of long videos without length limitations, supports parallel inference, and prevents error accumulation that may occur in recursive processing. Unlike prior methods, LoCoVSR offers a simple yet very effective VSR solution, avoiding explicit optical flow estimation, or information loss caused by latent space processing. Trained on the VFHQ face dataset, LoCoVSR achieves accurate, temporally consistent and high-quality upscaling with competitive results against recent diffusion-based VSR approaches.
Chinese Translation
视频超分辨率(VSR)是一个不适定逆问题,旨在从带有噪声的低分辨率(LR)视频中重建高分辨率(HR)视频。我们提出 LoCoVSR,一种基于扩散的 VSR 框架,利用像素空间去噪扩散概率模型。LoCoVSR 将扩散后验采样(Diffusion Posterior Sampling)技术与时空上下文学习相结合,以移动平均形式运行。使用相邻 LR 帧的局部窗口来恢复每个中心帧,同时在所有帧上应用共享的噪声轨迹。局部窗口化使得能够处理长视频而不受长度限制,支持并行推理,并防止递归处理中可能出现的误差累积。与先前方法不同,LoCoVSR 提供了一种简单却非常有效的 VSR 解决方案,避免了显式光流估计,或潜在空间处理造成的信息损失。在 VFHQ 人脸数据集上训练后,LoCoVSR 实现了准确、时间一致且高质量的上采样,与近期基于扩散的 VSR 方法相比具有竞争力的结果。
cs.CV / 144 / 2609.32761

From Feed-Forward to Flow: Unifying Reconstruction and Generation Is Easier Than You Think

从前馈到流:统一重建与生成比你想象的更简单
Wang, Haoru, Shen, Qianfan, Ye, Kai, Chen, Wenzheng, Chen, Baoquan
Abstract
Reconstruct where the images provide evidence, and generate where they do not: recent success of spatial world models such as Atlas (World Labs Team, 2026) highlights the value of unifying reconstruction and generation in one model. Yet the two have long lived in separate paradigms with distinctive failure modes: feed-forward reconstruction averages ambiguity into blur, while conditional generation invents plausible but scene-inconsistent detail. In this work, we present a unified flow-based formulation for reconstruction and generation, where a shared clean-target predictor performs direct reconstruction at its single-step endpoint and unfolds conditional generation through multi-step flow. A controlled toy study reveals the mechanism: with a single step, the predictor collapses to the conditional mean just like feed-forward methods, favoring consistency over diversity. With multi-step inference, the fidelity of generated details grows with context richness: closer observations reduce ambiguity and yield better-matched details. We further instantiate the formulation in appearance and geometry 3D tasks. JiT-LVSM improves perceptual and distributional quality in novel view synthesis, while JUSt3R retains competitive single-step geometry prediction with additional multi-step inference capabilities that reduces veil and flying-pixel artifacts, producing cleaner surface structure with greater test-time compute. Together, they show that reconstruction and generation can share both a formulation and a backbone, with their behavior governed by denoising configuration---making unification surprisingly simple.
Chinese Translation
在图像提供证据的地方进行重建,在图像未提供证据的地方进行生成:近期空间世界模型(如 Atlas (World Labs Team, 2026))的成功凸显了将重建与生成统一到一个模型中的价值。然而,这两者长期存在于不同的范式中,具有各自独特的失败模式:前馈重建将模糊性平均化为模糊,而条件生成则虚构出看似合理但与场景不一致的细节。在本工作中,我们提出了一种统一的基于流的重建与生成公式,其中共享的干净目标预测器在其单步端点执行直接重建,并通过多步流展开条件生成。一项受控的玩具研究揭示了其机制:在单步情况下,预测器像前馈方法一样坍缩到条件均值,倾向于一致性而非多样性。在多步推理下,生成细节的保真度随着上下文丰富度的增加而提高:更近的观测减少了模糊性,并产生更匹配的细节。我们进一步在外观和几何3D任务中实例化了该公式。JiT-LVSM 在新视角合成中提高了感知和分布质量,而 JUSt3R 保留了具有竞争力的单步几何预测,并增加了多步推理能力,减少了面纱和飞像素伪影,在更大的测试时计算下产生更干净的表面结构。总之,它们表明重建和生成可以共享公式和骨干网络,其行为由去噪配置控制——使得统一变得出奇简单。
cs.CV / 145 / 2609.32780

OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing

OmniMoE-VL:一种具有耦合视觉-深度路由的稀疏视觉语言模型
Qian, Long, Zhu, Bingke, Wei, Jiaqi, Chen, Yingying, Wang, Jinqiao
Abstract
Vision-language models (VLMs) increasingly use sparse mixture-of-experts (MoE) to scale language-side computation, yet visual information is typically routed only after passing through a fixed cross-modal interface. This leaves an important decision unresolved: which intermediate visual representations should be exposed to language computation for a given question? We introduce OmniMoE-VL, a sparse VLM with a coupled visual-depth routed projector. For each image-prompt pair, the projector selects a sparse set of intermediate visual depths and reuses the resulting global preference to guide both local patch fusion and dynamic visual injection into the language model. This design enables question-dependent visual access while preserving the native visual-token sequence, and complements token-level expert routing in the vision and language stacks. Across eight image-based benchmarks, OmniMoE-VL achieves an average score of 85.9 with 28B total and 9B activated parameters. Controlled comparisons show that the routed visual interface provides the dominant architectural gain, while matched route and component controls, same-image route analysis, and route interventions further support the value of coupling and question-conditioned visual access.
Chinese Translation
视觉语言模型(VLMs)越来越多地使用稀疏专家混合(MoE)来扩展语言侧计算,然而视觉信息通常仅在通过固定的跨模态接口之后才被路由。这留下了一个重要的决策未解决:对于给定问题,哪些中间视觉表示应该暴露给语言计算?我们引入了OmniMoE-VL,一种具有耦合视觉-深度路由投影器的稀疏VLM。对于每个图像-提示对,投影器选择一组稀疏的中间视觉深度,并重用得到的全局偏好来指导局部图像块融合和动态视觉注入到语言模型中。这种设计实现了问题依赖的视觉访问,同时保留了原生视觉token序列,并补充了视觉和语言栈中的token级专家路由。在八个基于图像的基准测试中,OmniMoE-VL以总参数28B和激活参数9B达到了85.9的平均分。控制比较表明,路由视觉接口提供了主要的架构增益,而匹配的路由和组件控制、同图像路由分析以及路由干预进一步支持了耦合和问题条件化视觉访问的价值。
cs.CV / 146 / 2609.32781

CT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language Models

CT-OPD:面向扩散视觉语言模型的反事实轨迹同策略蒸馏
Qian, Long, Zhu, Bingke, Wei, Jiaqi, Li, Yu, Chen, Yingying, Wang, Jinqiao
Abstract
Diffusion vision-language models generate answers by gradually resolving masked tokens, making accurate conditional prediction in partially resolved states central to post-training. Masking completed answers yields coherent contexts and targets, but prescribed masks do not reflect the model's reveal decisions. Its trajectories capture these decisions, yet their provisional visible tokens can conflict with the target response. Outcome-based reinforcement learning follows these trajectories but provides only response-level feedback, which loses contrast when sampled rewards tie. To align coherent token-level supervision with the model's reveal decisions, we introduce Counterfactual Trace On-Policy Distillation (CT-OPD), which combines completed teacher responses with trajectory masks from the current student. CT-OPD retokenizes each teacher response in the student's vocabulary and extracts unresolved-position masks at successive stages of the student's reverse process. For each mask, it discards provisional rollout values and reconstructs the partial state from the teacher endpoint, so the supervised positions follow the current trajectory while the visible context and targets remain consistent with the same response. The student is trained on these reconstructed states with its native categorical loss, and trajectories are refreshed as the model evolves. Across dense and sparse diffusion architectures, CT-OPD consistently enhances multimodal understanding and reasoning capabilities, with gains of up to 9.80 points on the nine-benchmark average. On the unified understanding-and-generation architecture, it also improves both visual understanding and image generation, showing that the same principle transfers across architectures and modalities. Ablations further attribute these gains to coherent reconstruction and current-model trajectory masks.
Chinese Translation
扩散视觉语言模型通过逐步解析掩码标记来生成答案,这使得在部分解析状态下的准确条件预测成为后训练的核心。对完整答案进行掩码可以产生连贯的上下文和目标,但预设的掩码并不反映模型的揭示决策。模型的轨迹捕捉了这些决策,但其临时的可见标记可能与目标响应冲突。基于结果的强化学习遵循这些轨迹,但仅提供响应级反馈,当采样奖励打平时会失去对比。为了将连贯的标记级监督与模型的揭示决策对齐,我们提出了反事实轨迹同策略蒸馏(CT-OPD),它将完整的教师响应与来自当前学生的轨迹掩码相结合。CT-OPD 在学生词汇表中对每个教师响应重新分词,并在学生反向过程的连续阶段提取未解析位置掩码。对于每个掩码,它丢弃临时的展开值,并从教师端点重建部分状态,因此监督位置遵循当前轨迹,而可见上下文和目标与同一响应保持一致。学生使用其原生分类损失在这些重建状态上进行训练,并且随着模型演化刷新轨迹。在稠密和稀疏扩散架构中,CT-OPD 持续增强多模态理解和推理能力,在九个基准的平均值上提升高达 9.80 分。在统一的理解与生成架构上,它还提升了视觉理解和图像生成,表明同一原理可跨架构和模态迁移。消融实验进一步将这些增益归因于连贯重建和当前模型轨迹掩码。
cs.CV / 147 / 2609.32789

Hierarchical Frequency-Domain Compression of Implicit Geometric Representations for Large-Scale Point Clouds

面向大规模点云的隐式几何表示分层频域压缩
Yao, Manlin, Liu, Jiabin, Wang, Guan, Liu, Haixu, Li, Hui
Abstract
Large-scale point cloud representations of complex geome tries incur prohibitive computational and memory costs, necessitating compressed implicit representations. To ad dress this, we propose a unified framework comprising im plicit geometric field representation, hierarchical frequency domain compression, and conditional high-frequency predic tion. Specifically, an unordered point cloud is mapped to an implicit field defined within its physical bounding box. A smooth Fourier pyramid is then constructed, where com pact low-frequency components capture the global geometry. Inter-scale high-frequency residuals are encoded to preserve the spatial information required for reconstructing fine geo metric details. To restore the high-frequency information lost during compression, we develop a hierarchical 3D neural net work. The reconstructed implicit field is converted back into a point cloud through isosurface extraction. Experiments on a complex-boundary point cloud with more than eight mil lion points demonstrate that the proposed method achieves a higher compression ratio than existing point cloud compres sion methods while maintaining comparable reconstruction quality.
Chinese Translation
大规模点云对复杂几何体的表示会带来高昂的计算和内存成本,因此需要压缩的隐式表示。为了解决这个问题,我们提出了一个统一的框架,包括隐式几何场表示、分层频域压缩和条件高频预测。具体而言,将无序点云映射到在其物理边界框内定义的隐式场。然后构建一个平滑傅里叶金字塔,其中紧凑的低频分量捕获全局几何。编码跨尺度高频残差,以保留重建精细几何细节所需的空间信息。为了恢复压缩过程中丢失的高频信息,我们开发了一个分层3D神经网络。通过等值面提取将重建的隐式场转换回点云。在具有超过八百万个点的复杂边界点云上的实验表明,所提出的方法在保持相当的重建质量的同时,比现有的点云压缩方法实现了更高的压缩比。
cs.CV / 148 / 2609.32794

Latent Space Is Not Flat: Rethinking Latent Structure for 3D Medical Image Synthesis

潜在空间并非平坦:重新思考3D医学图像合成的潜在结构
Xue, Haowen, Chen, Hao, Hu, Hexuan, Huang, Qian, Han, Yi, Meng, Qing, Xie, Zaipeng, Li, Chao, Xu, Haoli
Abstract
Latent generative models make 3D medical image synthesis computationally practical by generating in a compressed space. However, we show that the common flat Euclidean assumption induced by $\ell_2$ objectives is imprecise: latent-space geometry is so strongly anisotropic that equal-magnitude errors can produce drastically different decoded distortions. We further find that this anisotropy has a clear feature: sensitive variation concentrates in a low-rank subspace. The dominant low-rank components capture the overall structure, encoding long-range, spatially coordinated variation while remaining resistant to local noise. Its orthogonal residual, in contrast, mainly captures local and image-specific variation. Motivated by this asymmetry, we introduce Latent Structure Flow (LSF). At each block, LSF decomposes the latent state into structure and residual, models structural changes with global context, and predicts residual variation locally while preserving a direct path for the input structure. LSF changes only the generator, leaving the frozen codec and pointwise training objective unchanged. Across cross-modality synthesis and tumor inpainting tasks, LSF outperforms all compared baselines on both global and tumor-specific metrics, demonstrating the benefit of explicitly modeling latent-space structure for 3D medical image synthesis.
Chinese Translation
潜在生成模型通过在压缩空间中进行生成,使3D医学图像合成在计算上变得可行。然而,我们表明,由$\ell_2$目标诱导的常见平坦欧几里得假设并不精确:潜在空间几何具有强烈的各向异性,以至于相同幅度的误差可能产生截然不同的解码失真。我们进一步发现,这种各向异性有一个明显的特征:敏感变化集中在低秩子空间中。主导的低秩分量捕获整体结构,编码长程、空间协调的变化,同时对局部噪声保持鲁棒。相比之下,其正交残差主要捕获局部和图像特定的变化。受这种不对称性的启发,我们引入了潜在结构流(LSF)。在每个块中,LSF将潜在状态分解为结构和残差,利用全局上下文建模结构变化,并在局部预测残差变化,同时为输入结构保留直接路径。LSF仅改变生成器,保持冻结的编解码器和逐点训练目标不变。在跨模态合成和肿瘤修复任务中,LSF在全局和肿瘤特定指标上均优于所有比较基线,证明了显式建模潜在空间结构对3D医学图像合成的益处。
cs.CV / 149 / 2609.32811

Progressive Risk Estimation for Accident Anticipation

面向事故预判的渐进式风险估计
Hicsonmez, Samet, Çakar, Eray, Samet, Nermin, Güney, Fatma
Abstract
Accident anticipation aims to recognize anomalous driving cues before a crash while avoiding false alarms during normal driving. Existing approaches typically formulate this task as binary classification, focusing on whether an accident will occur rather than when it will occur. We propose PRE-ACT, a framework that models accident risk as a continuously evolving signal that increases as the crash approaches. By explicitly enforcing temporal ordering and distance-to-accident awareness, our method progressively raises risk while suppressing premature alarms, leading to significant improvements on MM-AU subsets and Nexar. We further introduce a Separation Score to evaluate the global behavior of predicted risk curves beyond local temporal windows. Code and visualizations are available at https://github.com/giddyyupp/PRE-ACT.
Chinese Translation
事故预判旨在识别碰撞前的异常驾驶线索,同时避免正常驾驶过程中的误报。现有方法通常将该任务表述为二分类问题,关注事故是否会发生,而非何时发生。我们提出PRE-ACT框架,将事故风险建模为随着碰撞临近而持续增长的连续演变信号。通过显式强制时序排序和距离事故感知,我们的方法逐步提升风险,同时抑制过早警报,从而在MM-AU子集和Nexar上取得显著改进。我们进一步引入分离分数(Separation Score)来评估预测风险曲线在局部时间窗口之外的全局行为。代码和可视化结果见 https://github.com/giddyyupp/PRE-ACT。
cs.CV / 150 / 2609.32813

USAI-Quant: A Quantitative Reasoning Benchmark for Vision-Language Models in Built Environments

USAI-Quant:面向建成环境中视觉语言模型的定量推理基准
Wang, Dongdong, Song, Qingqi, Chen, Yuzhou, Balakrishnan, Deepak, Srinivasan, Ravi Shankar, Wang, Shenhao
Abstract
Large vision-language models (VLMs) have emerged as a powerful paradigm for urban and spatial AI. However, current state-of-the-art large VLMs still struggle with quantitative reasoning on remote sensing imagery. Existing benchmarks and algorithms are predominantly based on qualitative Visual Question Answering (VQA), providing limited insights into the quantitative reasoning capabilities of VLMs for built environment metrics. To address this gap, we develop Quantitative Urban and Spatial AI benchmark (USAI-Quant), the first benchmark designed to quantitatively evaluate VLM's reasoning capabilities on built environment metrics via remote sensing imagery. USAI-Quant is curated from the 335 largest U.S. cities, aligning high-resolution remote sensing images with quantitative built environment metrics. We then evaluate both general-purpose and remote sensing VLMs (RS-VLMs) by applying VQAs to tens of built environment metrics across three complexity levels. Our results reveal that current state-of-the-art models consistently fall short on numeric reasoning tasks. We further conduct in-depth analyses across models, question types, and geographic locations, uncovering insights into performance variability and task-specific challenges.
Chinese Translation
大型视觉语言模型(VLMs)已成为城市与空间人工智能的一种强大范式。然而,当前最先进的大型 VLMs 在遥感影像的定量推理方面仍然面临挑战。现有基准和算法主要基于定性的视觉问答(VQA),对 VLMs 在建成环境指标上的定量推理能力提供的见解有限。为了弥补这一空白,我们开发了定量城市与空间人工智能基准(USAI-Quant),这是首个旨在通过遥感影像定量评估 VLM 在建成环境指标上推理能力的基准。USAI-Quant 从美国 335 个最大城市中整理而成,将高分辨率遥感影像与定量建成环境指标对齐。然后,我们通过将 VQA 应用于数十个建成环境指标(涵盖三个复杂度级别),评估了通用和遥感 VLMs(RS-VLMs)。我们的结果表明,当前最先进的模型在数值推理任务上持续表现不足。我们进一步从模型、问题类型和地理位置方面进行了深入分析,揭示了关于性能差异和任务特定挑战的见解。
cs.CV / 151 / 2609.32824

Unlocking Geodesic Gromov-Wasserstein Distances for 3D Modeling

解锁用于3D建模的测地Gromov-Wasserstein距离
Choromanski, Krzysztof Marcin, Long, Derek, Parashar, Ananya, Saha, Dwaipayan
Abstract
\textit{Gromov-Wasserstein Distances} (GWDs) provide quantitative ways of comparing probabilistic distributions defined on different metric spaces by applying techniques from the optimal transport theory. As such, GWD can be potentially useful in a large variety of applications ranging from graph matching problems to 3D object detection. However its practical use at scale is significantly limited by cubic time complexity computations involving dense intra-space distance matrices. Even though in the Euclidean metric spaces several techniques (e.g. involving scalable kernel methods) were proposed to address it, to the best of our knowledge, analogous techniques for general geodesic distances on manifolds, or shortest-path distance on graphs in their discretized variants, were not developed. In this paper, we present \textbf{E}fficient \textbf{G}eodesic \textbf{Gro}mov-\textbf{W}asserstein methods (EGGroW), a new class of efficient algorithms designed to calculate geodesic Gromov-Wasserstein distances with entropic Sinkhorn-like approaches, leveraging recently introduced \textit{GenusSink} methods \citep{genussink} and the theory of random features. We provide important downstream applications, namely: 3D pose estimation and 3D template detection. In the latter setting, we formulate a partial 3D template recovery as a staged problem: capacity-constrained scene selection is followed by semi-relaxed recovery of template visibility and correspondence. Our empirical findings show that EGGroW provides accurate solutions when standard Euclidean-based techniques fail and is characterized by light computational footprint, as our theoretical analysis predicts.
Chinese Translation
Gromov-Wasserstein距离(GWDs)通过应用最优传输理论中的技术,提供了比较定义在不同度量空间上的概率分布的定量方法。因此,GWD在从图匹配问题到3D目标检测的大量应用中可能具有潜在用途。然而,其大规模实际应用受到涉及稠密空间内距离矩阵的三次时间复杂度的计算的严重限制。尽管在欧几里得度量空间中已经提出了若干技术(例如涉及可扩展核方法)来解决这个问题,但据我们所知,针对流形上的一般测地距离,或图的离散变体中的最短路径距离的类似技术尚未被开发。在本文中,我们提出了高效测地Gromov-Wasserstein方法(EGGroW),这是一类新的高效算法,旨在利用最近引入的GenusSink方法 \citep{genussink} 和随机特征理论,通过类似熵Sinkhorn的方法计算测地Gromov-Wasserstein距离。我们提供了重要的下游应用,即:3D姿态估计和3D模板检测。在后一种设置中,我们将部分3D模板恢复表述为一个分阶段问题:首先进行容量约束的场景选择,然后进行模板可见性和对应关系的半松弛恢复。我们的实证结果表明,当标准的基于欧几里得的技术失败时,EGGroW能够提供准确的解决方案,并且正如我们的理论分析所预测的那样,其计算开销较小。
cs.CV / 152 / 2609.32840

VCRE-Fib: View-Conditioned Regional Evidence for Fine-Grained Ultrasound Grading of Schistosoma japonicum-Associated Liver Fibrosis

VCRE-Fib:用于日本血吸虫相关肝纤维化细粒度超声分级的视图条件区域证据
Xu, Ziyang, An, Shuli, Zhou, Hao, Zhong, Haitian, Wu, Tingting, Wang, Tao, Yang, Kun, Zeng, Tieyong
Abstract
Accurate assessment of Schistosoma japonicum-associated liver fibrosis is essential for disease management and long-term follow-up in endemic regions. Ultrasound provides non-invasive imaging, but complex local echogenic patterns and anatomical structures make fine-grained grading challenging. Existing deep learning methods can predict fibrosis scores, yet directly incorporating acquisition views and regional cues into grading while retaining spatial information for inspection remains an open problem. Here we present VCRE-Fib, a view-conditioned regional evidence framework that integrates anatomical context, local information, and global image assessment for fine-grained ultrasound grading. The framework forms view-conditioned local grading evidence before spatial pooling, uses weak localization to guide its aggregation, and combines it with global predictions. Image-only inference jointly returns a fibrosis score, acquisition view, and candidate abnormal-region map. We developed and evaluated the method on a re-curated cohort of 108,709 ultrasound images from 6,373 patients across 35 centers. On a patient-disjoint test set of 4,107 images from 240 patients across four centers, VCRE-Fib reduced the prespecified composite grading risk by 7.115% relative to SFibAI trained and evaluated on the same data split. Image-level mean absolute error decreased from 0.391 to 0.378, alongside lower patient-max, patient-median, and center-balanced risks. The full model also achieved lower composite grading risk than variants that separately removed view conditioning or weak localization. These results support incorporating anatomical context and regional evidence into ultrasound grading while exposing spatial predictions for inspection alongside severity estimates.
Chinese Translation
准确评估日本血吸虫相关肝纤维化对于流行地区的疾病管理和长期随访至关重要。超声提供无创成像,但复杂的局部回声模式和解剖结构使得细粒度分级具有挑战性。现有的深度学习方法可以预测纤维化评分,然而在分级中直接整合采集视图和区域线索,同时保留空间信息以供检查,仍然是一个未解决的问题。在此,我们提出VCRE-Fib,一种视图条件区域证据框架,它整合了解剖背景、局部信息和全局图像评估,用于细粒度超声分级。该框架在空间池化之前形成视图条件的局部分级证据,使用弱定位来指导其聚合,并将其与全局预测相结合。仅图像推理联合返回纤维化评分、采集视图和候选异常区域图。我们在来自35个中心、6373名患者的108709张超声图像的重新整理队列上开发和评估了该方法。在来自四个中心、240名患者的4107张图像的患者不相交测试集上,VCRE-Fib相对于在同一数据划分上训练和评估的SFibAI,将预定义复合分级风险降低了7.115%。图像级平均绝对误差从0.391降至0.378,同时患者最大、患者中位数和中心平衡风险也更低。完整模型也比分别去除视图条件或弱定位的变体实现了更低的复合分级风险。这些结果支持将解剖背景和区域证据纳入超声分级,同时在提供严重程度估计的同时暴露空间预测以供检查。
cs.CV / 153 / 2609.32841

Beyond Temporal Smoothing: Spatial Energy Budgets Stabilize One-Step Diffusion Editing

超越时间平滑:空间能量预算稳定单步扩散编辑
Zhou, Shengxiao, Luo, Lei, Yang, Jian
Abstract
One-step text-guided diffusion editing is efficient but prone to spatially misallocated updates that distort the edited object and alter the background. Existing methods often improve stability by averaging the editing field across timesteps. We instead identify spatial energy misallocation as a distinct and measurable failure mode: across two independent noise draws, the residual field is essentially unrepeatable, making the background field unreliable for direct transport, while its total energy still sets a usable magnitude for the draw at hand. BudEdit turns that magnitude into an explicit budget and reallocates it to edit-relevant regions selected jointly by residual energy and cross-attention, controlling where editing energy is spent rather than averaging over timesteps. The resulting training-free, inversion-free editor spends the budget on transport and reuses it to scale a correction in a lower-noise gated refinement. The budgeted injection field matches its prescribed budget exactly and vanishes on the identified background support, by construction. On PIE-Bench with SD-Turbo, BudEdit outperforms ChordEdit under each method's reported default settings on all 11 evaluated metrics, including a $2.1$\,dB gain in background PSNR, $31$\% lower DINO, and $36$\% lower LPIPS, while improving all five editing-quality metrics and reporting the lowest runtime in the comparison.
Chinese Translation
单步文本引导扩散编辑高效,但容易出现空间错配的更新,导致编辑对象失真并改变背景。现有方法通常通过跨时间步平均编辑场来提高稳定性。我们转而将空间能量错配识别为一种独特且可度量的失败模式:在两次独立噪声采样中,残差场本质上不可重复,使得背景场对于直接传输不可靠,而其总能量仍为当前这次采样设定了可用的量级。BudEdit 将该量级转化为显式预算,并将其重新分配到由残差能量和交叉注意力共同选择的编辑相关区域,从而控制编辑能量花费在何处,而不是跨时间步平均。由此得到的免训练、免反演编辑器将预算用于传输,并复用它来缩放低噪声门控细化中的校正。受预算约束的注入场精确匹配其规定预算,并且在识别出的背景支撑集上消失,这是由构造保证的。在 PIE-Bench 与 SD-Turbo 上,在每种方法各自报告的默认设置下,BudEdit 在所有 11 项评估指标上均优于 ChordEdit,包括背景 PSNR 提升 2.1 dB、DINO 降低 31%、LPIPS 降低 36%,同时改善了全部五项编辑质量指标,并在比较中报告了最低运行时间。
cs.CV / 154 / 2609.32846

SynCo: Learning Cross-Modal Synergy by Contrasting Interaction Residuals

SynCo:通过对比交互残差学习跨模态协同
Yarici, Yavuz, AlRegib, Ghassan
Abstract
Multimodal contrastive learning is a dominant paradigm for learning transferable representations from unlabeled data, but standard objectives primarily capture information that is redundant between modalities. Partial Information Decomposition (PID) shows that task-relevant information in multimodal data decomposes into three components: redundancy shared between modalities, uniqueness specific to each modality, and synergy available only from their joint observation. Recent frameworks extend contrastive learning to capture all three components, yet synergy remains undertrained in practice. We propose SynCo (Synergy Contrastive Learning), a method that directly addresses synergy undertraining through dedicated supervision on an interaction residual. SynCo fits a linear projector to predict the fused representation from independently computed unimodal features, and the resulting interaction residual, which removes the linearly unimodal-predictable component, receives dedicated contrastive supervision at negligible computational cost. On the controlled Trifeature benchmark, SynCo achieves state-of-the-art synergy capture with a $+5.98\%$ gain over the baseline, and on real-world benchmarks from MultiBench, DARai, and MM-IMDb, SynCo consistently outperforms or matches prior methods across diverse modality combinations and task types. The method operates as a plug-in to existing contrastive multimodal frameworks without modifying the underlying fusion architecture and can further improve synergy capture when combined with other methods.
Chinese Translation
多模态对比学习是从无标签数据中学习可迁移表示的主要范式,但标准目标主要捕获模态间冗余的信息。部分信息分解 (PID) 表明,多模态数据中与任务相关的信息可分解为三个部分:模态间共享的冗余、每个模态特有的独特性,以及仅从联合观察中可获得的协同。最近的框架扩展了对比学习以捕获所有三个部分,但在实践中协同仍然训练不足。我们提出 SynCo(协同对比学习),一种通过对交互残差进行专门监督来直接解决协同训练不足的方法。SynCo 拟合一个线性投影器,从独立计算出的单模态特征预测融合表示,得到的交互残差(去除了线性单模态可预测的部分)以可忽略的计算成本接受专门的对比监督。在受控的 Trifeature 基准上,SynCo 实现了最先进的协同捕获,比基线提高了 +5.98%;在来自 MultiBench、DARai 和 MM-IMDb 的真实世界基准上,SynCo 在不同的模态组合和任务类型中始终优于或匹配先前的方法。该方法作为现有对比多模态框架的插件运行,无需修改底层融合架构,并且当与其他方法结合时,可以进一步提高协同捕获。
cs.CV / 155 / 2609.32856

PolyTopoBench: A Benchmark for Complex Vector Polygon Generation from Remote Sensing Imagery

PolyTopoBench:面向遥感影像复杂矢量多边形生成的基准测试
Liu, Zeping, Lao, Ni, Sun, Weiwei, Wolff, Gil, Xie, Yiqun, Zhao, Liang, Jiao, Junfeng, Mai, Gengchen
Abstract
Vector polygon generation converts visual inputs, e.g., remote sensing (RS) images, into vectorized polygonal geometries, supporting applications such as autonomous driving, vector map construction, and remote sensing. Early pipelines predict raster masks and post-process them into polygons, which prevents end-to-end optimization and may miss small objects or introduce inaccurate vertices. Recent methods directly generate vector polygons, but most focus on simple exterior contours, while they either cannot represent complex polygons with holes or fail to preserve their topology. In this paper, we propose PolyTopoBench, a unified evaluation framework for vector polygon generation from RS images with explicit emphasis on complex polygons. PolyTopoBench evaluates both exterior and interior rings, and benchmarks 11 representative methods, including segmentation-based polygonization pipelines, vision foundation model baselines, and specialized vector polygon generators, on two RS-image datasets covering buildings, roads, vegetation, and unvegetated regions. Experiments show that existing methods often recover simple exterior boundaries but degrade substantially on polygons with holes or multiple rings. These results reveal complex polygon generation as an unresolved challenge and motivate topology-aware benchmarks and model designs. Code and data are available at https://github.com/seai-lab/PolyTopoBench.
Chinese Translation
矢量多边形生成将视觉输入(例如遥感(RS)影像)转换为矢量化的多边形几何,支持自动驾驶、矢量地图构建和遥感等应用。早期流程预测栅格掩膜并将其后处理为多边形,这阻碍了端到端优化,并可能漏检小目标或引入不准确的顶点。近期方法直接生成矢量多边形,但大多关注简单外轮廓,它们要么无法表示带孔洞的复杂多边形,要么无法保持其拓扑。本文提出 PolyTopoBench,一个用于从遥感影像生成矢量多边形的统一评估框架,并明确强调复杂多边形。PolyTopoBench 同时评估外环和内环,并在两个涵盖建筑物、道路、植被和无植被区域的遥感影像数据集上,对 11 种代表性方法进行基准测试,包括基于分割的多边形化流程、视觉基础模型基线和专用矢量多边形生成器。实验表明,现有方法通常能恢复简单外边界,但在带孔洞或多环的多边形上性能显著下降。这些结果揭示复杂多边形生成仍是一个未解决的挑战,并推动拓扑感知的基准测试与模型设计。代码和数据可在 https://github.com/seai-lab/PolyTopoBench 获取。
cs.CV / 156 / 2609.32857

Is H&E Image-to-Spatial Transcriptomics Simpler Than It Looks?

H&E图像到空间转录组学比看起来更简单吗?
Nguyen, Duc T., Do, Thanh Ha, Cao, Phuong M., Pham, Hieu
Abstract
Predicting spatial gene expression from routine H&E histology offers a scalable route toward spatial molecular profiling. Recent work has pursued increasingly sophisticated architectures to capture spatial context and richer expression structure. At the same time, simple estimators have shown strong performance in several studies, but what they already solve and where additional complexity is needed remain unclear. We study this behavior through the structure of prediction error under the mean-squared error (MSE) objective. Differences in average expression across genes can account for a substantial part of aggregate prediction performance, while a key unresolved error lies in recovering variation within each slide. Decomposing MSE into slide-level and within-slide components, we find that the within-slide component has lower residual-normalized parameter sensitivity in controlled neural experiments. This motivates Component-Guided Loss (CGL), which increases supervision of the within-slide component. CGL-Linear is a closed-form affine instantiation that achieves overall state-of-the-art performance across HEST-1k cohorts and gene-panel sizes. The same within-slide supervision improves existing neural models. These results suggest that substantial gains can come from aligning the training objective with prediction-error structure rather than increasing model complexity.
Chinese Translation
从常规H&E组织学图像预测空间基因表达为空间分子图谱提供了一条可扩展的途径。近期工作采用了越来越复杂的架构来捕获空间上下文和更丰富的表达结构。与此同时,简单的估计器在多项研究中表现出强大性能,但它们已经解决了什么以及哪里需要额外复杂性仍不清楚。我们通过均方误差(MSE)目标下的预测误差结构来研究这一行为。基因间平均表达的差异可以解释总体预测性能的很大一部分,而一个关键的未解决误差在于恢复每个切片内的变异。将MSE分解为切片级和切片内成分,我们发现在受控的神经实验中,切片内成分具有更低的残差归一化参数敏感性。这促使我们提出组件引导损失(CGL),它增加了对切片内成分的监督。CGL-Linear是一种闭式仿射实例化,在HEST-1k队列和基因面板大小上实现了总体最先进的性能。相同的切片内监督改进了现有的神经模型。这些结果表明,显著的增益可以来自将训练目标与预测误差结构对齐,而不是增加模型复杂度。
cs.CV / 157 / 2609.32863

SV2V-RSim: A Comprehensive Benchmark for Self-Selective V2V Cooperative Perception with Near-Realistic Data

SV2V-RSim:面向近真实数据的自选择V2V协同感知综合基准
Wu, Yulu, Wei, Chao, Cheng, Jujun, Ni, Zhangkai, Wang, Haowen, Suo, Dengyang, Chen, Cong, Liu, Xinyi, Gao, Shangce
Abstract
Vehicle-to-Vehicle (V2V) cooperative perception enhances autonomous driving by enabling vehicles to share information beyond their direct line of sight. However, existing V2V datasets are limited by a small number of participating agents, static collaborator selection strategies, and a significant domain gap between simulated and real-world environments. To overcome these challenges, we introduce SV2V-RSim, a large-scale, multi-modal, near-realistic simulation dataset engineered to elevate agent diversity and realism. Additionally, we present the Select Vehicles Adaptively (SVA) module, which optimizes collaborator selection to balance perception performance against communication bandwidth constraints. Our dataset is generated using the Unreal Engine 5-based simulator that integrates high-fidelity 3D assets, diverse environments, and intricate traffic scenarios. All vehicles within a specified range of the ego vehicle are equipped with sensor suites, enabling dynamic and adaptive collaborator selection. SV2V-RSim encompasses four maps, four weather conditions, six time periods from sunrise to night, 203K LiDAR frames, 402K RGB frames, and 788K annotated 3D bounding boxes across 17 object classes, supporting a range of cooperative perception tasks such as 3D object detection, segmentation, and depth estimation. Benchmarking on recent cooperative perception algorithms demonstrates that SVA achieves a superior performance-bandwidth trade-off, while sim-to-real experiments and No-Reference Image Quality Assessment validate the dataset's high realism and practical effectiveness. Our dataset and code will be publicly available.
Chinese Translation
车对车(V2V)协同感知通过使车辆能够共享其直接视线之外的信息,增强了自动驾驶的能力。然而,现有的V2V数据集受到参与代理数量少、合作者选择策略固定以及模拟环境与真实环境之间存在显著域差距的限制。为了克服这些挑战,我们推出了SV2V-RSim,一个大规模、多模态、近真实的仿真数据集,旨在提升代理的多样性和真实感。此外,我们提出了自适应选择车辆(SVA)模块,该模块优化合作者选择,以在感知性能和通信带宽约束之间取得平衡。我们的数据集使用基于虚幻引擎5的模拟器生成,该模拟器集成了高保真3D资产、多样化的环境和复杂的交通场景。自车指定范围内的所有车辆都配备了传感器套件,从而能够进行动态和自适应的合作者选择。SV2V-RSim包含四张地图、四种天气条件、从日出到夜晚的六个时间段、20.3万LiDAR帧、40.2万RGB帧和78.8万个标注的3D边界框,涵盖17个对象类别,支持一系列协同感知任务,如3D目标检测、分割和深度估计。对近期协同感知算法的基准测试表明,SVA实现了更优的性能-带宽权衡,而仿真到真实实验和无参考图像质量评估验证了数据集的高真实感和实用性。我们的数据集和代码将公开提供。
cs.CV / 158 / 2609.32876

Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity

多模态LLM在跨域组织学相似性上优于病理学基础模型
Zhang, Yishu, Li, Yun, Zhang, Daiwei
Abstract
State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outperform these specialized models in cross-domain histological similarity judgments. Using a relative similarity framework that we release as the MOSAIC (Model Similarity Assessment across Institutions and Cohorts) benchmark, we evaluate 17 models across 6 datasets and find that pathology encoders often rank same-institution, different-disease tiles as more similar than same-disease, different-institution tiles, a clinically dangerous failure mode invisible to standard within-domain evaluations. LLMs appear less susceptible to this failure, likely because they perform semantic visual comparison of morphology and tissue architecture rather than relying on shortcut features tied to acquisition context. Scaling training data does not resolve the problem for pathology encoders, implicating the learning objective rather than data coverage. Our results expose a fundamental robustness gap in current pathology foundation models and establish multimodal LLMs as a viable alternative for cross-institutional retrieval, dataset harmonization, and multi-site quality control. Code and data will be released upon acceptance.
Chinese Translation
最先进的病理学基础模型在数百万张组织学切片上训练,但当比较跨越切片或机构边界时,可能无法保持组织相似性。我们表明,通用多模态LLM,在没有被训练为病理学基础模型的情况下,在跨域组织学相似性判断中始终优于这些专用模型。使用我们作为MOSAIC(跨机构和队列的模型相似性评估)基准发布的相对相似性框架,我们评估了6个数据集上的17个模型,发现病理编码器通常将同机构、不同疾病的切片排序为比同疾病、不同机构的切片更相似,这是一种在标准域内评估中不可见的临床危险的失败模式。LLM似乎不太容易受到这种失败的影响,可能是因为它们对形态学和组织结构进行语义视觉比较,而不是依赖于与采集上下文相关的捷径特征。扩大训练数据并不能解决病理编码器的问题,这表明问题在于学习目标而不是数据覆盖。我们的结果揭示了当前病理学基础模型中一个基本的鲁棒性差距,并确立了多模态LLM作为跨机构检索、数据集协调和多站点质量控制的可行替代方案。代码和数据将在接受后发布。
cs.CV / 159 / 2609.32882

Improving Video Sparse Attention with Fine-grained Router and Sparse Rebasing

通过细粒度路由器和稀疏重定基准改进视频稀疏注意力
Zhang, Peiyuan, Wei, Guoqiang, Zhao, Yilong, Zhang, Zixiang, Zhou, Wei, Lin, Will, Zhang, Heng, Nie, Xiaonan, Zeng, Yan, Zhang, Hao
Abstract
We present VSA2, a frontier trainable sparse attention for video DiTs. VSA2 includes a variety of new architectural features and training procedures that we apply across all stages of the DiT development cycle, including pretraining, RL, and inference, to produce a DiT with comparable or better quality than a full attention counterpart. Architecturally, VSA2 introduces a fine-grained router that improves the precision of identifying critical tokens and supports dynamic computation by allowing each query to attend to a variable number of key-value pairs. In training, we identify a Hard-to-Easy Curriculum, where models trained under high sparsity and later evaluated with lower sparsity during inference not only generalize effectively, but also outperform models trained with full attention in motion quality. VSA2 is also flexible: it can replace full attention during the middle of progressive low-to-high resolution pretraining, rebasing early-stage full-attention checkpoints. Experiments show that VSA2 reduces attention computation by half over VSA with lower loss. On 720p videos, it accelerates attention by 8.9x and end-to-end generation by 4.62x compared to the FlashAttention-3 baseline, while achieving comparable or better video quality.
Chinese Translation
我们提出了VSA2,一种用于视频DiT的前沿可训练稀疏注意力。VSA2包含多种新的架构特征和训练流程,我们将其应用于DiT开发周期的所有阶段,包括预训练、强化学习和推理,从而生成质量与全注意力模型相当或更好的DiT。在架构上,VSA2引入了细粒度路由器,提高了识别关键token的精度,并通过允许每个查询关注可变数量的键值对来支持动态计算。在训练中,我们发现了从难到易的课程:在高稀疏度下训练、随后在推理时以较低稀疏度评估的模型不仅能够有效泛化,而且在运动质量上优于全注意力训练的模型。VSA2还具有灵活性:它可以在渐进式低到高分辨率预训练的中途替换全注意力,重定基准早期阶段的全注意力检查点。实验表明,VSA2相比VSA将注意力计算减少了一半,且损失更低。在720p视频上,与FlashAttention-3基线相比,它将注意力加速8.9倍,端到端生成加速4.62倍,同时达到相当或更好的视频质量。
cs.CV / 160 / 2609.32899

Rethinking the Fully Hyperbolic Vision Transformer in Polar Coordinates

重新思考极坐标下的全双曲视觉Transformer
Bdeir, Ahmad, Landwehr, Niels
Abstract
Hyperbolic space can embed intrinsic hierarchies in data with low distortion due to the exponential growth of volume with distance from the origin. However, current Lorentz transformer blocks are formulated in ambient coordinates, where numerical errors increase at large radii due to instability in the Lorentzian inner product, leading many models to limit the radius to avoid this issue. This prevents us from utilizing the regions of hyperbolic space that motivate the geometry. To address this, we revisit the components of the transformer block in polar coordinates, where hyperbolic operations such as distance calculation and attention centroids can be computed without the numerical cancellation in their ambient-coordinate formulations. Specifically, we propose a polar fully connected layer that separately maps an embedding's direction and radius, allowing the radius of a feature to be learned rather than determined by the norm of a linear map. We additionally introduce horospherical shifts as relative positional encodings whose query-key distances grow logarithmically with the token gap. Finally, we reformulate the residual connection as average radius Lorentz boosts. Combining these components, we develop a fully hyperbolic transformer that substantially improves performance over Euclidean and hyperbolic baselines on standard vision tasks. We further evaluate our model on ImageNet and demonstrate its ability to generalize to other datasets using pre-trained weights, similar to Euclidean counterparts.
Chinese Translation
双曲空间由于体积随距原点距离呈指数增长,能够以低失真嵌入数据中的内在层次结构。然而,当前的Lorentz Transformer模块是在环境坐标中表述的,其中由于Lorentz内积的不稳定性,数值误差在大半径处增加,导致许多模型限制半径以避免此问题。这阻止了我们利用激发该几何的双曲空间区域。为了解决这个问题,我们在极坐标中重新审视Transformer模块的组件,其中双曲操作(如距离计算和注意力质心)可以在没有环境坐标表述中的数值抵消的情况下计算。具体来说,我们提出了一种极坐标全连接层,分别映射嵌入的方向和半径,允许特征半径被学习,而不是由线性映射的范数决定。我们还引入了作为相对位置编码的horospherical shifts(超球面偏移),其查询-键距离随token间隔呈对数增长。最后,我们将残差连接重新表述为平均半径Lorentz boosts(洛伦兹推进)。结合这些组件,我们开发了一个全双曲Transformer,在标准视觉任务上显著优于欧几里得和双曲基线。我们进一步在ImageNet上评估我们的模型,并展示了其使用预训练权重泛化到其他数据集的能力,类似于欧几里得对应模型。
cs.CV / 161 / 2609.32944

Synthetic Thermal Image Generation for Real-Time Animal Detection Under Low-Visibility Conditions

低能见度条件下用于实时动物检测的合成热图像生成
Momoh, James, Ahmed, Khandaker Mamun
Abstract
Wildlife-vehicle collisions remain a significant road safety concern, particularly during nighttime and low-visibility conditions when RGB-based perception systems are often unreliable. Thermal imaging offers a promising alternative for detecting animals under poor illumination. However, the limited availability of annotated infrared animal datasets restricts the development of robust deep learning-based detection models. This paper investigates synthetic thermal image generation as a scalable approach for real-time animal detection under low-visibility conditions. A subset of 514 annotated visible-spectrum animal images from the NTLNP dataset is translated into synthetic thermal representations using CycleGAN-Turbo, while a limited real thermal dataset of 60 images is expanded through thermal-focused augmentation. Multiple object detection architectures, including YOLOv8, YOLOv9, YOLOv10, and RT-DETR, are trained independently on synthetic and real thermal datasets and evaluated using precision, recall, mAP@0.5, mAP@0.5:0.95, model size, and inference latency. Experimental results show that synthetic thermal images provide competitive detection performance, with RT-DETR achieving the highest synthetic-data mAP@0.5 of 0.9613. Models trained on augmented real thermal data achieve the strongest overall performance, with YOLOv10s obtaining 0.9879 mAP@0.5 and 0.9571 mAP@0.5:0.95. Computational analysis further indicates that lightweight YOLO variants provide favorable inference latency, supporting their potential for real-time deployment. These findings demonstrate that synthetic thermal imagery can reduce dependence on scarce infrared datasets and support the development of efficient animal detection systems for future vehicle-mounted wildlife collision mitigation applications.
Chinese Translation
野生动物与车辆碰撞仍然是重要的道路安全问题,尤其是在夜间和低能见度条件下,基于RGB的感知系统往往不可靠。热成像为在不良光照下检测动物提供了一种有前景的替代方案。然而,标注的红外动物数据集有限,限制了鲁棒深度学习检测模型的开发。本文研究了合成热图像生成作为一种可扩展方法,用于低能见度条件下的实时动物检测。使用CycleGAN-Turbo将NTLNP数据集中514张标注的可见光谱动物图像子集转换为合成热表示,同时通过热成像聚焦的数据增强扩展了包含60张图像的小型真实热数据集。包括YOLOv8、YOLOv9、YOLOv10和RT-DETR在内的多种目标检测架构分别在合成和真实热数据集上独立训练,并使用精确率、召回率、mAP@0.5、mAP@0.5:0.95、模型大小和推理延迟进行评估。实验结果表明,合成热图像提供了具有竞争力的检测性能,其中RT-DETR在合成数据上达到了最高的mAP@0.5为0.9613。在增强的真实热数据上训练的模型获得了最强的整体性能,YOLOv10s达到了0.9879的mAP@0.5和0.9571的mAP@0.5:0.95。计算分析进一步表明,轻量级YOLO变体提供了有利的推理延迟,支持其用于实时部署的潜力。这些发现表明,合成热成像可以减少对稀缺红外数据集的依赖,并支持开发高效的动物检测系统,用于未来车载野生动物碰撞缓解应用。
cs.CV / 162 / 2609.32957

DynamicDx: Evaluating Evidence Acquisition in Video-Based Diagnosis

DynamicDx:评估基于视频的诊断中的证据获取
Li, Jiahui, Guo, Yutong, Yang, Nan, Song, Wenzhan, Lu, Jin, Dou, Fei
Abstract
Diagnosing a patient from video requires more than recognizing the sign: a vision-language model must turn what it sees into hypotheses, questions and tests. DynamicDx evaluates each step in 71 neurological consultations across 11 sign categories, linking authentic patient videos to confirmed diagnoses and fixed charts built from the same case reports, so that every model queries the same evidence. Across five such models, video improves accuracy by 9.9-22.5 percentage points over blind input, but neither recognition alone nor temporal order explains the gain: the cause is usually missing from the model's video-only differential diagnosis even when the sign is recognized, and shuffling the frames produces no reliable accuracy loss. Instead, a trajectory replay traces most of the gain to the investigation results the video prompts. Evidence acquisition is the bottleneck: supplying the decisive investigations raises accuracy to 73.2-93.0%. Two interventions act on it. A post-trained 4B video describer improves sign descriptions, especially from a short, densely sampled segment, and source-clean literature retrieval expands initial hypotheses; both bring the tests a model orders closer to those the treating clinicians documented and, through them, raise accuracy. For video-based diagnosis, seeing better helps when it leads to asking better.
Chinese Translation
从视频中诊断患者不仅仅需要识别体征:视觉-语言模型必须将其所见转化为假设、问题和检查。DynamicDx 评估了跨 11 个体征类别的 71 次神经科会诊中的每个步骤,将真实的患者视频与确诊诊断以及基于相同病例报告构建的固定病历相连接,从而使每个模型查询相同的证据。在五个此类模型中,视频相较于盲输入将准确率提高了 9.9-22.5 个百分点,但仅凭识别或时间顺序都无法解释这一提升:即使识别出体征,病因通常也缺失于模型仅基于视频的鉴别诊断中,并且打乱帧的顺序不会产生可靠的准确率下降。相反,轨迹回放将大部分提升追溯到视频所提示的检查结果。证据获取是瓶颈:提供决定性的检查可将准确率提高到 73.2-93.0%。有两种干预措施作用于它。一个后训练的 4B 视频描述器改进了体征描述,尤其是从短时密集采样的片段中,而来源干净的文献检索扩展了初始假设;两者都使模型所开具的检查更接近治疗临床医生记录的检查,并借此提高准确率。对于基于视频的诊断,看得更好有助于问得更好。
cs.CV / 163 / 2609.32971

Distributed Hydrological Modeling in the Feature Space

特征空间中的分布式水文建模
Eddin, Mohamad Hakam Shams, Taccari, Maria Luisa, Zhang, Yikui, Jiang, Shijie, Gall, Juergen, Reichstein, Markus
Abstract
Accurate forecasting of river discharge and floods is very challenging. River dynamics are affected by storage, meteorological forcing, and flow propagation at different spatial and temporal scales. Forecasting thus requires a framework that considers the upstream-to-downstream flow through river networks across grid cells and catchments. This modeling is known in hydrology as distributed modeling and routing. Existing deep learning approaches either ignore this topology, operate on lumped catchments, or route predicted physical quantities through a separate graph or physical routing model. We instead introduce feature-space routing: a topology-aware state-space operator embedded directly in the forecasting dynamics. At every forecast step, the operator gathers latent states from upstream grid cells and causally updates the downstream state according to the known river network. This preserves the physical connectivity of the river system while allowing the propagated state itself to be learned end-to-end and allows the model to predict river discharge considering both local dynamics and neighboring upstream contributions. To address uncertainty and provide probabilistic forecasts, we minimize the fair continuous ranked probability score (fCRPS) as a training objective. Our experiments on the European Flood Awareness System (EFAS) and observational data for river discharge forecasting demonstrate that encoding the physical structure of river networks explicitly in the feature space substantially improves the forecasting skill, particularly in an ungauged setting. Our approach achieves state-of-the-art results on both reanalysis and observational data and is able to forecast maps of river discharge at 1 arcminute and 6-hourly resolution up to 10 days lead time.
Chinese Translation
准确预报河流流量和洪水极具挑战性。河流动力学受不同时空尺度上的蓄量、气象强迫和水流传播的影响。因此,预报需要一个考虑通过河网从上游到下游、跨越网格单元和流域的水流的框架。这种建模在水文学中被称为分布式建模和汇流演算。现有的深度学习方法要么忽略这种拓扑结构,要么在集总流域上操作,要么通过单独的图或物理演算模型对预测的物理量进行演算。我们转而引入特征空间演算:一种直接嵌入预报动力学中的拓扑感知状态空间算子。在每个预报步骤,该算子从上游网格单元收集潜在状态,并根据已知河网因果地更新下游状态。这保留了河流系统的物理连通性,同时允许传播状态本身进行端到端学习,并使模型能够考虑局部动态和邻近上游贡献来预测河流流量。为了解决不确定性并提供概率预报,我们最小化公平连续排名概率得分(fCRPS)作为训练目标。我们在欧洲洪水预警系统(EFAS)和河流流量预报观测数据上的实验表明,在特征空间中显式编码河网的物理结构显著提高了预报技巧,尤其是在无测站环境下。我们的方法在再分析数据和观测数据上均取得了最先进的结果,并且能够以 1 角分和 6 小时的分辨率预报最长 10 天预见期的河流流量图。
cs.CV / 164 / 2609.32984

ReVision3D: Attribution-Guided Recursive Self-Improvement for 3D Medical Perception

ReVision3D:面向3D医学感知的归因引导递归自我改进
Lee, Ho Hin, Zhou, Yuyin, Yu, Yannan, Gu, Shi, Wu, Yifan
Abstract
Recursive self-improvement (RSI) offers a promising path for overcoming the limited visual capability of current medical imaging agents. Yet applying RSI to volumetric imaging remains difficult: failures can arise from acquisition, perception, training recipe, or downstream inference, while self-generated feedback and logged trajectories provide little guidance on which component should change. We introduce ReVision3D, an RSI system that leverages 3D volumes with spatially grounded annotations to determine where visual evidence is lost and recursively improve the corresponding visual capability. A frozen language-model designer proposes revisions to acquisition, perception, training, or inference, while the verifier and system-level objective remain fixed. Our key insight is that an annotated volume forms an exact replay world for view rendering and spatial verification: unvisited views can be rendered on demand, and localized predictions can be checked directly against reference masks. This grounded feedback directs targeted revision, while only changes that improve beyond measured seed noise are retained. Each accepted change triggers renewed attribution, allowing the dominant bottleneck to shift across rounds. On abdominal CT, attribution identifies perception as the dominant remaining limitation. Revising that level enables ReVision3D to achieve 79% liver recall and 83% kidney recall at under 0.4 false positives per patient, outperforming the evaluated frozen multimodal foundation models, with the largest gains on small lesions.
Chinese Translation
递归自我改进(Recursive Self-Improvement, RSI)为克服当前医学影像智能体有限的视觉能力提供了一条有前景的路径。然而,将RSI应用于体成像仍然困难:失败可能源于采集、感知、训练方案或下游推理,而自生成反馈和记录轨迹难以指导哪个组件应当改变。我们提出ReVision3D,一种RSI系统,利用带有空间定位标注的3D体积来确定视觉证据在哪里丢失,并递归地改进相应的视觉能力。一个冻结的语言模型设计器提出对采集、感知、训练或推理的修订,而验证器和系统级目标保持固定。我们的关键见解是,带标注的体积为视图渲染和空间验证形成了一个精确的重放世界:未访问的视图可以按需渲染,局部预测可以直接对照参考掩码进行检查。这种有依据的反馈指导有针对性的修订,同时只保留那些改进超过测量种子噪声的更改。每个被接受的更改都会触发新的归因,使得主要瓶颈能够在轮次之间转移。在腹部CT上,归因将感知识别为主要剩余限制。修订该层级使ReVision3D能够在每患者低于0.4假阳性的情况下达到79%的肝脏召回率和83%的肾脏召回率,优于所评估的冻结多模态基础模型,且在小病灶上增益最大。
cs.CV / 165 / 2609.32996

Oracle Gaps in Reliability Coverage: Sampling Noise or Policy Specialization?

可靠性覆盖中的Oracle差距:采样噪声还是策略专业化?
Cakiroglu, Mert Onur, Dalkilic, Mehmet, Kurban, Hasan
Abstract
Policies trained from the same base model can appear to solve different problems. An oracle that chooses the best policy for each problem may therefore appear much stronger than any single policy. Selecting the largest estimated success rate also selects favorable sampling errors. We study this effect through reliability coverage, the fraction of problems whose success probability reaches a chosen threshold. Our first test redistributes stored correctness outcomes across policies within each problem. A second also preserves each policy's total successes, accounting for overall quality differences under a specified statistical model. For five training seeds of a seven-billion-parameter vision-language model, redistribution reproduces 0.096 of an estimated 0.113 oracle gap at threshold 0.10. Neither test finds significant evidence at this threshold. Small advantages remain unresolved. Mixtures, routers, voting, and weight averaging show no detectable improvement over their corresponding single-policy baselines. Training policies on different datasets shows little detectable specialization under light post-training, and no router gain. A stronger recipe does create it, both tests detect it, and a router gain appears only at the high thresholds where the specialists separate. In a control with predictable specialization, a router recovers about half the oracle gap. Coverage bounds explain why even a genuine oracle advantage need not yield a deployment gain. The tests assess apparent specialization from stored responses before investment in routing. Code: https://github.com/KurbanIntelligenceLab/oracle-gaps
Chinese Translation
从同一基础模型训练出的策略可能看起来能解决不同的问题。因此,一个为每个问题选择最佳策略的Oracle可能看起来比任何单一策略都强大得多。选择最大的估计成功率也会选择有利的采样误差。我们通过可靠性覆盖来研究这一效应,即可靠性覆盖是指成功概率达到选定阈值的问题比例。我们的第一个测试在每个问题内跨策略重新分配存储的正确性结果。第二个测试还保留了每个策略的总成功次数,并在指定的统计模型下考虑了整体质量差异。对于一个70亿参数的视觉语言模型的五个训练种子,在阈值0.10下,重新分配再现了估计的0.113 Oracle差距中的0.096。两个测试在此阈值下均未发现显著证据。微小的优势仍未得到解释。混合、路由、投票和权重平均相对于其对应的单策略基线没有显示出可检测的改进。在不同数据集上训练策略在轻量后训练下显示出很少的可检测专业化,并且没有路由增益。更强的配方确实创造了它,两个测试都检测到了它,并且路由增益仅出现在专家分离的高阈值处。在一个具有可预测专业化的对照中,路由器恢复了大约一半的Oracle差距。覆盖边界解释了为什么即使真正的Oracle优势也不一定产生部署增益。这些测试在投资路由之前从存储的响应中评估表观专业化。代码:https://github.com/KurbanIntelligenceLab/oracle-gaps
cs.CV / 166 / 2609.33003

Certified Interface Aliases: Exact Collisions in Vision-Language Preprocessing, and When They Exist

认证的接口别名:视觉-语言预处理中的精确碰撞及其存在条件
Cakiroglu, Mert Onur, Buxton, Elham, Dalkilic, Mehmet, Kurban, Hasan
Abstract
Vision-language verifiers and routers must distinguish errors repairable by more reasoning from those caused by visual evidence never reaching the language model. This distinction lacks ground truth because annotators see full-resolution images while models receive preprocessed tensors. We introduce AliasForge to create cases where the relevant fact is provably absent from the interface. Fixed-point resampling makes the pre-rounding resize an exact integer linear map that can send nonzero integer perturbations to zero. Hiding a label-flipping perturbation there produces images with opposite step-correctness labels but bit-identical interface states. Every verifier therefore has the same output law on both members, giving pair-balanced accuracy exactly one half and zero gain from language-side repair. We prove that every fixed-point downscaler has such null vectors and bound their smallest size at the most common ratios, which rules out an 8-bit fit whenever the bound exceeds 255. From the resize configuration alone, a lattice criterion supplies realizable collisions and certifies their absence within the specified construction family. It resolves all 20 screened configurations, 17 as constructible and 3 as non-constructible. We construct certified pairs across three architectures and certify four additional processors, with zero decision-logit gap on all 18 scored pairs and none of the 18 controls. The pairs also screen routers that waste computation on re-attention or further reasoning. On natural items, per-item routing headroom exists, but no tested interface-only router improves over stopping. Our fiber ceiling bounds the headroom recoverable from the interface. Code: https://github.com/KurbanIntelligenceLab/aliasforge.
Chinese Translation
视觉-语言验证器和路由器必须区分可通过更多推理修复的错误与因视觉证据从未到达语言模型而导致的错误。这种区分缺乏真实标签,因为标注者看到的是全分辨率图像,而模型接收的是预处理后的张量。我们引入 AliasForge 来创建相关事实可证明从接口中缺失的案例。定点重采样使得舍入前的缩放成为一个精确的整数线性映射,可以将非零整数扰动发送为零。在那里隐藏一个标签翻转扰动,会产生具有相反步骤正确性标签但接口状态逐位相同的图像。因此,每个验证器在两个成员上具有相同的输出规律,使得成对平衡准确率恰好为一半,且从语言侧修复获得零收益。我们证明每个定点下采样器都有这样的零向量,并在最常见比率下界定其最小尺寸,当界限超过255时排除了8位拟合。仅从缩放配置出发,一个格准则提供了可实现的碰撞,并在指定构造族内证明其不存在。它解决了所有20个筛选配置,17个可构造,3个不可构造。我们在三种架构上构造了认证对,并认证了四个额外处理器,在所有18个评分对上决策logit差距为零,而18个控制组均无。这些对还筛选出在重新注意或进一步推理上浪费计算的路由器。在自然项目上,存在逐项路由余量,但没有测试的仅接口路由器能超过停止。我们的纤维上界界定了可从接口恢复的余量。代码:https://github.com/KurbanIntelligenceLab/aliasforge。
cs.CV / 167 / 2609.33005

Safety-Constrained Cascade Inference for Robust Malaria Cell Classification Under Field Corruptions

面向现场损坏下鲁棒疟疾细胞分类的安全约束级联推理
Hagbe, J. T., Emel, Michel
Abstract
Automated malaria diagnosis from thin blood-smear microscopy could meaningfully reduce the burden on under-resourced laboratories, but a model that maximises accuracy on clean laboratory images fails badly the moment an inexpensive smartphone camera introduces sensor noise. This paper introduces MalariaCascade, a two-stage inference system in which a lightweight MobileNetV2 sentinel (2,225,153 parameters) makes confident classifications at low compute cost and escalates uncertain cases to an EfficientNet-B3 expert (10,697,769 parameters) that sees only clean, standardised images regardless of how corrupted the incoming frame is. Structural isolation of the expert stage, not learned robustness, is the mechanism. The sentinel is trained under a safety-score objective that places an explicit floor on Recall(Parasitised) (>=0.95) and Recall(Uninfected) (>=0.40) before precision is optimised, ensuring the checkpoint satisfies clinical safety constraints by construction. On the NIH Malaria Cell Images Dataset (27,558 cells), the cascade reaches Accuracy=0.9736, Recall(Parasitised)=0.9570, Precision(Parasitised)=0.9912, F1=0.9738, and AUROC=0.9955 on the clean test set. Under Gaussian sensor noise at full severity, cascade Recall(Parasitised) degrades by only 2.4 pp (0.9570 to 0.9329), while the flat single-model baseline collapses by 63.0 pp (0.9584 to 0.3281). McNemar's test confirms the cascade improvement is statistically significant (chi^2=11.14, p=0.00085). An ablation isolating the structural property shows that routing clean images to the expert is responsible for a 16.8 pp Recall(Parasitised) advantage under sensor noise relative to a cascade where the expert also sees corrupted inputs.
Chinese Translation
基于薄血涂片显微镜的自动疟疾诊断可显著减轻资源匮乏实验室的负担,但在低成本智能手机摄像头引入传感器噪声时,那些在干净实验室图像上追求最高准确率的模型会严重失效。本文提出了 MalariaCascade,一个两阶段推理系统。其中轻量级 MobileNetV2 哨兵(2,225,153 个参数)以低计算成本进行高置信度分类,并将不确定的病例升级给 EfficientNet-B3 专家(10,697,769 个参数),该专家只接收干净的、标准化的图像,而无论输入帧有多损坏。专家阶段的结构隔离(而非学习到的鲁棒性)才是其机制。哨兵模型在安全分数目标下训练,该目标在优化精确率之前,对寄生虫感染召回率(>=0.95)和未感染召回率(>=0.40)设定了明确的下限,从而确保检查点从构造上满足临床安全约束。在 NIH 疟疾细胞图像数据集(27,558 个细胞)上,级联模型在干净测试集上达到准确率=0.9736,召回率(寄生虫感染)=0.9570,精确率(寄生虫感染)=0.9912,F1=0.9738,AUROC=0.9955。在最高强度的高斯传感器噪声下,级联模型的召回率(寄生虫感染)仅下降 2.4 个百分点(0.9570 至 0.9329),而扁平单模型基线则崩溃下降 63.0 个百分点(0.9584 至 0.3281)。McNemar 检验证实级联模型的改进具有统计学显著性(卡方=11.14,p=0.00085)。一项隔离该结构属性的消融实验表明,在传感器噪声下,将干净图像路由到专家,相对于专家也看到损坏输入的级联模型,带来了 16.8 个百分点的召回率(寄生虫感染)优势。
cs.CV / 168 / 2609.33020

Residual Diffusion Implicit Models

残差扩散隐式模型
Guerreiro, João, Tomás, Pedro, Aidos, Helena, Nascimento, Jacinto C.
Abstract
Diffusion models achieve state-of-the-art results across multiple tasks. However, in inverse problems, standard initialization from pure Gaussian noise misaligns the generative process with real-world degradations. More recent methods such as diffusion bridges impose strict endpoint constraints and often require long reverse processes that are prone to hallucinations. Alternative consistency models provide noise-invariant, one-step mappings but lack inherent variance modeling and can degrade under severe corruption. Hence, residual diffusion implicit models (RDIMs) are proposed, constituting a generalized framework that explicitly models the residuals between high-quality (HQ) and low-quality (LQ) images, aligning the forward process with the actual degradation. A non-Markovian implicit reverse sampler is derived, which can skip intermediate timesteps, enabling accurate few-step or even single-step reconstruction, while mitigating the hallucinations inherent to long diffusion chains. RDIM also introduces a controllable variance mechanism that interpolates between deterministic and stochastic sampling, balancing fidelity and diversity. Furthermore, it enables the straightforward use of perceptual losses, when needed. Experiments on denoising and super-resolution benchmarks demonstrate that RDIMs consistently outperforms the state of the art, including bridge and consistency models, in terms of PSNR, SSIM, and LPIPS, reducing hallucinations while requiring only a few sampling steps (often just one). The results position RDIMs as an efficient solution for a broad range of image restoration tasks.
Chinese Translation
扩散模型在多个任务中取得了最先进的结果。然而,在逆问题中,从纯高斯噪声进行的标准初始化使生成过程与真实世界的退化不一致。较新的方法如扩散桥施加了严格的端点约束,并且通常需要长的反向过程,这容易产生幻觉。替代的一致性模型提供了噪声不变的一步映射,但缺乏固有的方差建模,并且在严重损坏下可能退化。因此,提出了残差扩散隐式模型(RDIMs),构成一个广义框架,显式地建模高质量(HQ)和低质量(LQ)图像之间的残差,使前向过程与实际退化对齐。推导出一种非马尔可夫隐式反向采样器,可以跳过中间时间步,实现准确的少步甚至单步重建,同时减轻长扩散链固有的幻觉。RDIM还引入了一种可控方差机制,在确定性采样和随机采样之间进行插值,平衡保真度和多样性。此外,它使得在需要时可以直接使用感知损失。在去噪和超分辨率基准测试上的实验表明,RDIMs在PSNR、SSIM和LPIPS方面始终优于最先进的方法,包括桥和一致性模型,减少幻觉,同时仅需少量采样步骤(通常仅一步)。结果将RDIMs定位为广泛图像恢复任务的高效解决方案。
cs.CV / 169 / 2609.33076

NutriVision: Ingredient-Conditioned Fusion and Prediction for Single-Image Food Nutrition Estimation

NutriVision:面向单图像食物营养估计的食材条件融合与预测
Kumar, Aman, Anand, Avinash, Lakhchaura, Chaitanya, Kumar, Ashutosh, Abrol, Akshita, Liu, Timothy, Wang, Zhengkui, Shah, Rajiv Ratn
Abstract
Nutrition estimation is a fundamental task in consumer diet tracking, clinical dietetics, chronic disease management, sports and hospital nutrition, and broader food computing systems. The existing approaches have progressed along two largely separate axes, vision models that rely on calibrated RGB-depth captures and ingredient-aware methods that use textual cues but use limited multimodal fusion. We introduce NutriVision, an end-to-end framework that leverages visual geometry and ingredient semantics to estimate calories, mass, fat content, carbohydrates, and protein from a single RGB image and an optional ingredient list. It obtains the unavailable depth modality using DepthAnything-V3 and encodes ingredient descriptions using CLIP. It integrates three complementary mechanisms: (1) an \emph{Ingredient-Conditioned Frequency-Aligned Fusion Module (IC-FAFM)}, which uses textual guidance to reweight and align RGB-depth frequency components; (2) an \emph{Ingredient-Aware Mask-based Prediction Head (IA-MPH)}, whose gating and channel masks are conditioned on food identity; and (3) modality-specific \emph{Internal Semantic Modeling (ISM)} blocks. On the Nutrition5k dataset, NutriVision achieves a mean PMAE of $\mathbf{13.60\pm0.10\%}$, outperforming our IGSMNet implementation by $0.89$ percentage points and OmniFood8k by $2.90$ percentage points (both $p<0.001$). The module-level ablations identify the ingredient-aware prediction head as the primary architectural contributor, improving mean PMAE by $1.50\pm0.17$ percentage points ($p<0.001$). These results demonstrate that ingredient-conditioned prediction and frequency-aware RGB-depth fusion provide measurable gains for single-image nutrient estimation. More broadly, NutriVision offers a practical route toward nutrition-assessment systems that exploit geometric and semantic cues without requiring specialized depth-sensing hardware
Chinese Translation
营养估计是消费者饮食跟踪、临床营养学、慢性病管理、运动与医院营养以及更广泛的食品计算系统中的一项基础任务。现有方法主要沿着两条相对独立的路线发展:依赖校准 RGB-D 捕获的视觉模型,以及使用文本线索但多模态融合有限的食材感知方法。我们提出 NutriVision,一个端到端框架,利用视觉几何和食材语义,从单张 RGB 图像和可选的食材列表估计卡路里、质量、脂肪含量、碳水化合物和蛋白质。它使用 DepthAnything-V3 获取不可用的深度模态,并使用 CLIP 编码食材描述。它集成了三种互补机制:(1) 食材条件频率对齐融合模块(IC-FAFM),利用文本引导对 RGB-D 频率分量进行重加权和对齐;(2) 食材感知的基于掩码的预测头(IA-MPH),其门控和通道掩码以食物身份为条件;(3) 模态特定的内部语义建模(ISM)模块。在 Nutrition5k 数据集上,NutriVision 实现了 $\mathbf{13.60\pm0.10\%}$ 的平均 PMAE,比我们的 IGSMNet 实现高出 $0.89$ 个百分点,比 OmniFood8k 高出 $2.90$ 个百分点(均 $p<0.001$)。模块级消融实验表明,食材感知预测头是主要的架构贡献者,将平均 PMAE 提高了 $1.50\pm0.17$ 个百分点($p<0.001$)。这些结果表明,食材条件预测和频率感知的 RGB-D 融合为单图像营养估计带来了可测量的增益。更广泛地说,NutriVision 为利用几何和语义线索而无需专用深度传感硬件的营养评估系统提供了一条实用途径。
cs.CV / 170 / 2609.33077

Parameter-Efficient 3D Segmentation of Liver and Liver tumors: Depthwise factorization Scales Better Than Dense Convolution with Spatial Dimensionality

参数高效的三维肝脏和肝脏肿瘤分割:深度方向因子化随空间维度的扩展优于稠密卷积
Alkhadrawi, Adham M., Mahmoud, Mohammed A. B.
Abstract
Three-dimensional dense convolutional networks are the strongest performers on volumetric medical image segmentation, but their parameter counts scale poorly: moving a dense k x k convolution to k x k x k multiplies its weights by k. We observe that depthwise separable factorization does not share this penalty. Because the cubic kernel term applies only to the depthwise stage while the pointwise projection, which dominates the parameter count, is unchanged, the same architecture grows by 5 % from 2D to 3D where a dense convolutional U-Net grows by 200 %. We exploit this asymmetry to build a 3D U-Net with 536,990 parameters, 24x fewer than an identical dense 3D U-Net. On MSD Task03 Liver (the Medical Segmentation Decathlon liver task, derived from LiTS), evaluated per case under five-fold cross-validation over all 131 public volumes, the model reaches a tumor Dice of 0.577 (95% CI [0.518, 0.633]) and a liver Dice of 0.947 (95% CI [0.941, 0.952]). Its liver Dice exceeds previously reported performance. Its tumor Dice exceeds their low-resolution configuration (0.4701) by 0.107 and their 2D configuration (0.5394), at approximately one twenty-fourth of the parameters and roughly half the in-plane resolution. Trained under identical conditions on a common held-out split, it exceeds a dense 3D U-Net on liver by +0.031 Dice (paired p = 0.006) and on tumor by +0.041 (95 % CI [+0.005, +0.081], paired p = 0.056), suggesting the factorization also acts as a regularizer in the small-data regime characteristic of medical imaging. We further show, on both LiTS and a 2D endoscopy benchmark, that a large fraction of the network's learnable spatial filters can be replaced by fixed shifts at no cost in accuracy, but that replacing all of them is measurably worse, the placement of spatial capacity matters more than its total amount.
Chinese Translation
三维稠密卷积网络在体医学图像分割中表现最强,但其参数量扩展性差:将稠密 k×k 卷积变为 k×k×k 卷积会使其权重乘以 k。我们观察到,深度可分离因子化不存在这一代价。因为立方核项(k×k×k)只作用于深度阶段,而主导参数量的逐点投影保持不变,所以同一架构从 2D 到 3D 仅增长 5%,而稠密卷积 U-Net 增长 200%。我们利用这种不对称性构建了一个 3D U-Net,其参数量为 536,990,约为相同稠密 3D U-Net 的 1/24。在 MSD Task03 Liver(医学分割十项全能肝脏任务,源自 LiTS)上,对所有 131 个公开体数据采用五折交叉验证按病例评估,模型达到肿瘤 Dice 0.577(95% CI [0.518, 0.633])和肝脏 Dice 0.947(95% CI [0.941, 0.952])。其肝脏 Dice 超过先前报道的性能。其肿瘤 Dice 比其低分辨率配置(0.4701)高 0.107,也超过其 2D 配置(0.5394),而参数量约为其 1/24,面内分辨率约为其一半。在共同留出划分上以相同条件训练,它在肝脏上超过稠密 3D U-Net +0.031 Dice(配对 p = 0.006),在肿瘤上超过 +0.041(95% CI [+0.005, +0.081],配对 p = 0.056),表明该因子化在医学图像所典型的小数据场景中还起到正则化作用。我们进一步在 LiTS 和一个 2D 内窥镜基准上表明,网络中的大部分可学习空间滤波器可以被固定平移替换而不损失精度,但替换全部滤波器则会明显更差;空间容量的放置位置比其总量更重要。
cs.CV / 171 / 2609.33082

SemReward-VL: Semantic Reward-Guided Video-Language Adaptation for Developmental Behavior Assessment

SemReward-VL:语义奖励引导的视频-语言自适应用于发育行为评估
Jiang, De, Zhang, Shuo, Yuan, Kehong, Liao, Hongen
Abstract
Developmental screening videos show how children perform specific behaviors, but clinical records usually contain outcomes rather than descriptions of what happened. We present SemReward-VL, which learns to describe item-specific behavior from these outcomes. A vision-language model generates a description, and a frozen language model scores its agreement with the clinical outcome, relevance to the item, abstention on unrelated video-item pairs, and clarity. Group relative policy optimization (GRPO) updates LoRA adapters using this semantic reward. On 13,379 videos covering 41 items, the method improves aggregate accuracy and the number of items with recall above 0.5. Errors remain in temporal direction, duration, and age-specific interpretations of behavior.
Chinese Translation
发育筛查视频展示儿童如何执行特定行为,但临床记录通常包含结果而非对发生情况的描述。我们提出SemReward-VL,它学习从这些结果中描述项目特定的行为。一个视觉-语言模型生成描述,一个冻结的语言模型对其与临床结果的一致性、与项目的相关性、对不相关视频-项目对的弃权以及清晰度进行评分。组相对策略优化(GRPO)使用该语义奖励更新LoRA适配器。在涵盖41个项目的13,379个视频上,该方法提高了总体准确率以及召回率高于0.5的项目数量。错误仍存在于时间方向、持续时间以及对行为的年龄特异性解释方面。
cs.CV / 172 / 2609.33083

dKFD: Phase-Structured Evidence Allocation for Fixed-Budget Localized Event Understanding

dKFD:面向固定预算局部事件理解的阶段结构化证据分配
Bagri, Aditya, Kumar, Ashutosh, Lakhchaura, Chaitanya, Anand, Avinash, Wang, Zhengkui, Shah, Rajiv Ratn
Abstract
Sparse video understanding often requires selecting a small set of visual evidence under a fixed frame budget. Most sparse selectors allocate this budget globally, allowing all frames to compete with one another. For temporally localized events, this can be a poor inductive bias: useful evidence is often distributed across pre-event context, the event itself, and post-event consequences. We study fixed-budget evidence allocation for localized event videos and show that globally competitive Top-$K$ selectors can preserve recognition and grounding while producing unstable event evidence. On DoTA Video Anomaly Recognition, Global Top-$K$ obtains competitive recognition and temporal grounding, but low selector-event alignment at $K=12$ (Frame AUC $51.5 \pm 10.8$). We propose dKFD, a phase-structured differentiable selector that reserves evidence capacity across pre-event, event, and post-event phases after full-sequence temporal encoding. Under matched-budget multi-seed evaluation, dKFD improves Frame AUC by $+30.97$ over a matched Global Top-$K$ selector at $K=12$ ($p<0.01$), while yielding modest but statistically significant recognition gains and comparable temporal grounding. Mechanism ablations show that phase supervision is load-bearing: removing it reduces Frame AUC to $41.1 \pm 12.0$ even when phase-partitioned budgets are retained. Downstream diagnostics on VRU-Accident show consistent gains over learned Global Top-$K$ across VLM families, while dense captioning reveals a boundary condition where uniform sampling remains competitive. These results support phase-structured allocation as a controlled fixed-budget approach for event-centric sparse evidence selection, not as a universal video summarization strategy.
Chinese Translation
稀疏视频理解通常需要在固定帧预算下选择少量视觉证据。大多数稀疏选择器全局分配该预算,允许所有帧相互竞争。对于时间上局部化的事件,这可能是一种糟糕的归纳偏置:有用的证据通常分布在事件前上下文、事件本身和事件后结果中。我们研究局部事件视频的固定预算证据分配,并表明全局竞争的Top-K选择器可以保持识别和定位,同时产生不稳定的事件证据。在DoTA视频异常识别上,全局Top-K在K=12时获得了有竞争力的识别和时间定位,但选择器-事件对齐度低(帧AUC 51.5 ± 10.8)。我们提出dKFD,一种阶段结构化的可微分选择器,在全序列时间编码后,在事件前、事件和事件后阶段保留证据容量。在匹配预算的多随机种子评估下,dKFD在K=12时将帧AUC比匹配的全局Top-K选择器提高了+30.97(p<0.01),同时产生适度但统计显著的识别增益和相当的时间定位。机制消融表明,阶段监督是关键性的:即使保留阶段划分的预算,移除它也会将帧AUC降低到41.1 ± 12.0。在VRU-Accident上的下游诊断显示,在VLM系列中相对于学习的全局Top-K有一致的增益,而密集描述揭示了一个边界条件,即均匀采样仍然具有竞争力。这些结果支持阶段结构化分配作为以事件为中心的稀疏证据选择的一种受控固定预算方法,而不是作为通用的视频摘要策略。
cs.CV / 173 / 2609.33088

QSCP: Beyond Class-Name Prompts for Query-Guided Semantic Change Parsing

QSCP:超越类别名称提示的查询引导语义变化解析
Qian, Yuan, Ma, Jie
Abstract
Traditional change detection (CD) identifies changes between bi-temporal remote sensing images, while semantic change detection (SCD) assigns predefined land-cover classes. However, mapping all changes may not meet a user's specific needs. Referring change detection (RCD) enables selective retrieval through category prompts. However, existing category-prompted RCD uses the queried category to specify the destination of a change and returns only a binary mask of the corresponding regions. Users may instead request a particular transition and paired semantic maps to understand what changed into what. Such requests require explicit source and target reasoning beyond target-class localization. To address these needs, we propose query-guided semantic change parsing (QSCP), which supports category names, synonyms, and intent-bearing sentences and returns a query-specific mask with paired temporal semantic maps. QSCP parses requests into intents and semantic slots, composes bidirectional visual evidence, and predicts both temporal states with a query-conditioned decoder. On SECOND, QSCP outperforms RCDNet on synonym, sentence, and transition queries and improves end-to-end semantic prediction over evaluated semantic baselines. WHU-CDC experiments further assess cross-dataset transfer and consistency across equivalent expressions without target-domain training. Code is available at https://github.com/qianyuancs/QSCP
Chinese Translation
传统的变化检测(CD)识别双时相遥感影像之间的变化,而语义变化检测(SCD)则分配预定义的土地覆盖类别。然而,映射所有变化可能无法满足用户的特定需求。指称变化检测(RCD)通过类别提示实现选择性检索。然而,现有的类别提示RCD使用查询类别来指定变化的目的地,并仅返回对应区域的二值掩膜。用户可能转而请求特定的转换和成对的语义图,以理解什么变成了什么。此类请求需要明确的源和目标推理,而不仅仅是目标类别的定位。为了满足这些需求,我们提出了查询引导的语义变化解析(QSCP),它支持类别名称、同义词和携带意图的句子,并返回带有成对时相语义图的查询特定掩膜。QSCP将请求解析为意图和语义槽,组合双向视觉证据,并使用查询条件解码器预测两个时相状态。在SECOND数据集上,QSCP在同义词、句子和转换查询上优于RCDNet,并在端到端语义预测上超越了所评估的语义基线。WHU-CDC实验进一步评估了在没有目标域训练情况下的跨数据集迁移和等效表达的一致性。代码可在 https://github.com/qianyuancs/QSCP 获取。
cs.CV / 174 / 2609.33097

Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation

查询、对齐与蒸馏:面向高效视觉-语言导航的导航感知跨模态交互
Chen, Zhihao, Ge, Yiyuan, Wang, Ziyang, Cao, Pu, Yang, Lu
Abstract
Recent large-scale Vision-and-Language Navigation (VLN) models deliver strong accuracy but remain costly to deploy due to their large parameter counts and computational requirements. We tackle efficient VLN in two steps. First, we build a high-performing teacher that makes navigation evidence selection explicit and compressible. The teacher introduces a small set of learnable query slots to extract global and local action-sufficient navigable evidence from panoramic observations via a Navigable Query Generator, then progressively grounds these evidence tokens to the instruction with an Instruction-Query Aligner for policy prediction. Second, using this explicit query bottleneck as a distillation interface, we train a compact student by transferring both where to attend and what to do. We distill the teacher's global and local navigable queries with a navigation-aware token-adaptive objective, then further match action distributions during fine-tuning. Experiments on standard VLN benchmarks demonstrate that our student nearly matches the teacher's navigation performance while reducing the number of parameters by 93.65% compared to the teacher.
Chinese Translation
近年来,大规模视觉-语言导航(VLN)模型虽能提供高准确率,但由于其庞大的参数量和计算需求,部署成本依然高昂。我们分两步解决高效VLN问题。首先,我们构建一个高性能的教师模型,使导航证据的选择过程显式且可压缩。该教师模型引入一小组可学习的查询槽(query slots),通过可导航查询生成器(Navigable Query Generator)从全景观测中提取全局和局部的、动作充分的可导航证据,然后利用指令-查询对齐器(Instruction-Query Aligner)将这些证据token逐步与指令进行关联以进行策略预测。其次,利用这个显式的查询瓶颈作为蒸馏接口,我们通过迁移“关注何处”和“执行何种动作”来训练一个紧凑的学生模型。我们使用导航感知的token自适应目标来蒸馏教师的全局和局部可导航查询,并在微调过程中进一步匹配动作分布。在标准VLN基准测试上的实验表明,我们的学生模型几乎达到了教师模型的导航性能,同时参数量相比教师模型减少了93.65%。
cs.CV / 175 / 2609.33100

Octree-based Video Representation

基于八叉树的视频表示
Zhou, Rungui, Zhou, Chuanzhi, Hou, Yuk-Kit, Wang, Peng-Shuai
Abstract
Video models commonly use uniform grids even though visual complexity varies substantially across space and time. We introduce OctVideo, which approximates a video clip with an octree. This hierarchy recursively partitions a spatio-temporal volume into eight subvolumes, so that smooth regions remain coarse while detailed regions receive finer cells. Each leaf stores local RGB values and spatio-temporal gradients, supplemented by a lightweight learned residual. For reconstruction, a Conv1D VAE maps the serialized cells to a regular latent grid and selectively refines details during decoding. Our VAE achieves 36.12 dB PSNR with 38.2M parameters and 189.4 GFLOPs per clip on Kinetics-400 (K400). It also generalizes zero-shot to the high-resolution Densely Annotated VIdeo Segmentation (DAVIS) 2016 dataset with reconstruction quality comparable to the best evaluated models. On both datasets, it requires the fewest model FLOPs and achieves the fastest encoding and decoding among the evaluated models. OctVideo also supports video understanding, achieving competitive recognition performance with few input tokens when trained from scratch. By exploiting the redundancy already present in video signals and efficiently processing sparse structures, OctVideo provides an efficient representation for video.
Chinese Translation
视频模型通常使用均匀网格,尽管视觉复杂度在空间和时间上变化很大。我们提出了 OctVideo,它用八叉树近似视频片段。这种层次结构递归地将一个时空体划分为八个子体积,使得平滑区域保持粗糙,而细节区域获得更精细的单元。每个叶节点存储局部 RGB 值和时空梯度,并辅以一个轻量级的学习残差。为了重建,Conv1D VAE 将序列化的单元映射到规则的潜在网格,并在解码过程中选择性地细化细节。我们的 VAE 在 Kinetics-400 (K400) 上实现了 36.12 dB PSNR,每个片段具有 38.2M 参数和 189.4 GFLOPs。它还能零样本泛化到高分辨率的 Densely Annotated VIdeo Segmentation (DAVIS) 2016 数据集,重建质量与最佳评估模型相当。在两个数据集上,它所需的模型 FLOPs 最少,并且在评估模型中实现了最快的编码和解码。OctVideo 还支持视频理解,从头训练时能用少量输入 token 达到有竞争力的识别性能。通过利用视频信号中已有的冗余并高效处理稀疏结构,OctVideo 为视频提供了一种高效的表示。
cs.CV / 176 / 2609.33109

Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models

迈向全面的3D定位:通过视觉语言模型实现朝向定位
Liang, Tuo, Liu, Disheng, Wang, Nengbo, Chaudhary, Vipin, Yin, Yu
Abstract
Grounding is a core capability of spatial vision-language models, yet most existing work focuses only on where a referred object is. Many 3D tasks also require knowing how it is oriented. Although existing 3D VLMs may predict oriented boxes, box pose does not explicitly capture object-centric orientation or symmetry-induced ambiguities. We introduce orientation grounding, a referring grounding task that predicts an object's 6D orientation and axial symmetry from a language or box query in single-view or multi-view scenes. To support this task, we construct ReferOri, with 331K multi-view and 387K single-view orientation-grounding queries obtained through scalable reconstruction, consistency checking, and human verification. We further present OG-VLM, which adapts a 3D VLM with structured box/orientation outputs, sign and symmetry tokens, and geometry-aware auxiliary losses. Across single-view and multi-view benchmarks, OG-VLM substantially outperforms orientation-aware VLM baselines and surpasses object-level orientation foundation models on scene-level referring benchmarks, showing that explicit orientation grounding is a distinct and learnable capability beyond localization. Downstream results validate its benefit for orientation-related spatial reasoning.
Chinese Translation
定位是空间视觉语言模型的一项核心能力,但多数现有工作仅关注被指代对象位于何处。许多3D任务还需要知道其朝向如何。尽管现有3D VLM可以预测有向框,但框姿态并未显式捕捉以对象为中心的朝向或对称性引起的歧义。我们引入朝向定位(orientation grounding),这是一个指代定位任务,从单视图或多视图场景中的语言或框查询预测物体的6D朝向和轴向对称性。为支持该任务,我们构建了ReferOri,包含331K多视图和387K单视图朝向定位查询,这些查询通过可扩展重建、一致性检查与人工验证获得。我们进一步提出OG-VLM,它通过结构化框/朝向输出、符号与对称性token以及几何感知辅助损失来适配3D VLM。在单视图和多视图基准上,OG-VLM显著优于具备朝向感知的VLM基线,并在场景级指代基准上超越对象级朝向基础模型,表明显式朝向定位是一种区别于定位且可学习的能力。下游结果验证了其对朝向相关空间推理的益处。
cs.CV / 177 / 2609.33125

Train Together or Merge Later? Unifying VLA Experts via a Shared Action Interface

联合训练还是后期合并?通过共享动作接口统一 VLA 专家
Zhang, Zhizhen, Fu, Yuxia, Wang, Zijian, Huang, Helen, Luo, Yadan
Abstract
Co-training offers a straightforward way to build a multi-task vision-language-action (VLA) policy, but can fall short of the performance achieved by training each task independently. The challenge is to retain these task-specific gains in a multi-task policy without joint post-training. Combining independently trained experts through model merging is a natural approach, yet strong individual experts do not necessarily yield a strong merged policy. We identify one source of this incompatibility: task-specific changes to the action interface, comprising action normalization and the action encoder and decoder. We propose PolicyWeave, combining merge-compatible post-training with context-guided sparse merging. During post-training, all experts retain the common base policy's action interface, while task adaptation is restricted to LoRA updates in the hidden layers of the action model. This makes the experts more compatible with existing model merging methods. However, merging all experts can still introduce interference from unrelated tasks at deployment. PolicyWeave scores each expert's LoRA updates using the initial visual-language context, determines the expert set through leave-one-layer-out ranking stability, and forms a sparse weighted merge of the selected updates that remains fixed for current task. We evaluate PolicyWeave with GR00T N1.5 on 18 RoboCasa365 tasks, using only 10% of the target-task demonstrations for supervised fine-tuning (SFT). Preserving the shared action interface raises the average success rate across four static merging methods from 17.0% to 52.8%. PolicyWeave achieves 64.7% success with these SFT experts and 74.1% after task-specific reinforcement learning (RL), compared with 60.7% for joint RL. Further evaluations on LIBERO-10 and an AgileX Piper arm support the deployment of independently learned skills in long-horizon and real-world manipulation.
Chinese Translation
协同训练提供了一种构建多任务视觉-语言-动作(VLA)策略的直接方法,但其性能可能不及独立训练每个任务所达到的水平。挑战在于,如何在不进行联合后训练的情况下,将各任务特定的增益保留在多任务策略中。通过模型合并来组合独立训练的专家是一种自然思路,但强大的单个专家未必能得到强大的合并策略。我们识别出这种不兼容性的一个来源:对动作接口的任务特定修改,包括动作归一化以及动作编码器和解码器。我们提出 PolicyWeave,将合并兼容的后训练与上下文引导的稀疏合并相结合。在后训练期间,所有专家都保留共同基础策略的动作接口,而任务适配仅限于动作模型隐藏层中的 LoRA 更新。这使专家与现有模型合并方法更加兼容。然而,合并所有专家在部署时仍可能引入来自无关任务的干扰。PolicyWeave 使用初始视觉-语言上下文对每个专家的 LoRA 更新进行评分,通过留一层排序稳定性确定专家集合,并对所选更新形成稀疏加权合并,该合并对当前任务保持固定。我们在 18 个 RoboCasa365 任务上使用 GR00T N1.5 评估 PolicyWeave,仅使用目标任务演示的 10% 进行监督微调(SFT)。保留共享动作接口将四种静态合并方法的平均成功率从 17.0% 提升到 52.8%。PolicyWeave 使用这些 SFT 专家达到 64.7% 的成功率,在任务特定强化学习(RL)后达到 74.1%,而联合 RL 为 60.7%。在 LIBERO-10 和 AgileX Piper 机械臂上的进一步评估,支持将独立学得的技能部署到长时程和真实世界操作中。
cs.CV / 178 / 2609.33148

DroneWAM: Efficient World Action Model for Drone Visual Navigation

DroneWAM:面向无人机视觉导航的高效世界动作模型
Yao, Liang, Liu, Fan, Lu, Hongbo, Xu, Wei, Jiang, Jianyu, Shen, Yijun, Zhang, Chuanyi, Peng, Pai
Abstract
World-action models give visual navigation agents a way to anticipate how candidate actions will change future observations and to act from the predicted consequences. For drones, this capability must operate under tight accuracy and efficiency constraints. We present DroneWAM, an efficient world-action model for drone visual navigation. DroneWAM adopts a JEPA-based architecture to model future states directly in representation space, avoiding the cost of explicit future image generation. A pretrained Resampler further compresses dense encoder features into fewer latent tokens, reducing the computation repeated at each imagined step. We also introduce adaptive rollout, where a preference-trained Gate adaptively allocates prediction depth according to the current scene. To support learning under richer aerial motion, we construct DroneNav-6D, a simulated visual navigation dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances. On DroneNav-6D, DroneWAM achieves the best trajectory accuracy among the compared methods. Adaptive rollout further reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy, demonstrating that predictive computation can be allocated more effectively across scenes. \href{https://github.com/1e12Leon/DroneWAM}{Codes and data} will be released.
Chinese Translation
世界动作模型为视觉导航智能体提供了一种途径,使其能够预测候选动作将如何改变未来的观测,并根据预测的结果采取行动。对于无人机而言,这种能力必须在严格的精度和效率约束下运行。我们提出了DroneWAM,一种用于无人机视觉导航的高效世界动作模型。DroneWAM采用基于JEPA的架构,直接在表示空间中建模未来状态,避免了显式生成未来图像的开销。预训练的Resampler进一步将密集的编码器特征压缩为更少的潜在令牌,减少了在每个想象步骤中重复的计算。我们还引入了自适应展开,其中经过偏好训练的Gate根据当前场景自适应地分配预测深度。为了支持在更丰富的空中运动下学习,我们构建了DroneNav-6D,一个模拟的视觉导航数据集,包含同步的RGB观测、6自由度飞行轨迹、控制命令和随机风扰。在DroneNav-6D上,DroneWAM在比较方法中取得了最佳的轨迹精度。自适应展开进一步将平均预测深度从8减少到4.58,同时提高了轨迹精度,表明预测计算可以在不同场景中更有效地分配。\href{https://github.com/1e12Leon/DroneWAM}{代码和数据}将发布。
cs.CV / 179 / 2609.33158

FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs

FOCUS:从基础视觉编码器到多模态大语言模型的视网膜模型泛化基准测试
Restrepo, David, Wu, Chenwei, Nakayama, Luis Filipe, Martins, Miguel L., Christodoulidis, Stergios, Vakalopoulou, Maria, Ferrante, Enzo
Abstract
Progress in AI-based retinal image analysis has advanced with foundation models, yet evaluating their reliability remains challenging. Performance reported on a single dataset does not capture how models behave under dataset shift, across clinical definitions, or for different patient subgroups. This limitation is particularly critical in medical imaging analysis, where robustness, calibration, and fairness are essential for safe deployment. We introduce FOCUS (Foundation Ophthalmic Cross-Dataset Understanding under Shift), a cross-dataset benchmark for evaluating retinal fundus models that considers vision-only encoder models (VM), vision-language dual-encoder models (VLM), and multimodal large language models (MLLM). FOCUS harmonizes binary diabetic retinopathy, referable diabetic retinopathy, and glaucomatous optic neuropathy tasks across ten public datasets spanning diverse geographies, acquisition conditions, and label protocols. The benchmark evaluates models through a unified analysis layer that measures ranking performance, calibration, subgroup disparities, and image-quality robustness. We present a large-scale evaluation covering 532 base configurations and 228 MLLM configurations adapted through supervised fine-tuning with low-rank adaptation (LoRA). Results show that no model family consistently dominates across tasks and datasets: general VM encoders achieve the strongest average ranking performance, medical MLLMs are competitive but variable, and dual encoder VLMs benefit substantially from lightweight adaptation. Fine-tuning improves in-domain performance but exhibits heterogeneous transfer to external datasets, particularly in calibration. These findings demonstrate that retinal model evaluation is inherently multidimensional. FOCUS provides a practical framework and public benchmark to assess generalization, reliability, and robustness beyond single-dataset leaderboards
Chinese Translation
基于 AI 的视网膜图像分析随着基础模型的发展而进步,但评估其可靠性仍然具有挑战性。在单一数据集上报告的性能不能反映模型在数据集偏移下、跨临床定义或针对不同患者亚组的行为。这一局限性在医学影像分析中尤为关键,因为稳健性、校准和公平性对于安全部署至关重要。我们介绍了 FOCUS(Foundation Ophthalmic Cross-Dataset Understanding under Shift),一个用于评估视网膜眼底模型的跨数据集基准,考虑了纯视觉编码器模型 (VM)、视觉-语言双编码器模型 (VLM) 和多模态大语言模型 (MLLM)。FOCUS 统一了二分类糖尿病视网膜病变、需转诊糖尿病视网膜病变和青光眼性视神经病变任务,跨越十个公共数据集,涵盖不同的地理区域、采集条件和标注协议。该基准通过统一的分析层评估模型,测量排序性能、校准、亚组差异和图像质量稳健性。我们进行了一项大规模评估,涵盖 532 个基础配置和 228 个通过低秩适应 (LoRA) 监督微调适配的 MLLM 配置。结果表明,没有哪个模型家族在所有任务和数据集上始终占主导地位:通用 VM 编码器实现了最强的平均排序性能,医学 MLLM 具有竞争力但表现不稳定,而双编码器 VLM 从轻量级适配中获益显著。微调提高了域内性能,但在向外部数据集迁移时表现异质,尤其是在校准方面。这些发现表明,视网膜模型评估本质上是多维的。FOCUS 提供了一个实用的框架和公共基准,用于评估超越单一数据集排行榜的泛化性、可靠性和稳健性。
cs.CV / 180 / 2609.33167

FloodDiffusion 2: Efficient and Path Controllable Streaming Motion Generation

FloodDiffusion 2:高效且路径可控的流式运动生成
Cai, Yiyi, Wu, Yuhan, Li, Kunhang, Fangyuan, Tu, Zhang, Xiangyue, Li, Qiaoge, Wang, Zhixiang, Zhang, Kaipeng, Liu, Haiyang
Abstract
We present FloodDiffusion 2 (FD2), an efficient and controllable framework that builds upon FloodDiffusion (FD1), a state-of-the-art streaming motion generation model. While FD1 produces plausible motion, it suffers from low efficiency and limited controllability, as its attention design requires repeated computation over the entire history, and it lacks precise trajectory control for real-world applications. To address these limitations and improve generation quality, FD2 introduces three advances. First, Partial Attention makes finalized history representations independent of the active window, enabling KV-cached inference and shared-history packing for efficient training. Second, we establish a necessary-and-sufficient Bregman criterion for regression losses to preserve diffusion's conditional-mean velocity field. This criterion guides an FK-induced quadratic loss that incorporates motion geometry without online FK evaluation. Third, FD2 introduces precise path conditioning to control the character's root trajectory while preserving natural body motion. Experiments show that FD2 reduces training computation by 4.6$\times$ and accelerates denoising by 11.29$\times$, reaching 2.303 ms per update on long sequences. Alongside these efficiency gains, FD2 improves motion quality over FD1 and achieves state-of-the-art FID scores among streaming methods, with 0.048 on SEED and 0.053 on HumanML3D.
Chinese Translation
我们提出了 FloodDiffusion 2 (FD2),一个高效且可控的框架,它建立在 FloodDiffusion (FD1) 的基础上,FD1 是一种最先进的流式运动生成模型。尽管 FD1 能生成合理的运动,但它效率低且可控性有限,因为其注意力设计需要对整个历史进行重复计算,并且缺乏对实际应用的精确轨迹控制。为了解决这些局限性并提高生成质量,FD2 引入了三项改进。首先,Partial Attention 使最终的历史表示独立于活动窗口,从而实现 KV 缓存推理和共享历史打包,以进行高效训练。其次,我们为回归损失建立了一个充要的 Bregman 准则,以保持扩散的条件均值速度场。该准则指导了一种 FK 诱导的二次损失,它结合了运动几何信息,而无需在线 FK 评估。第三,FD2 引入了精确的路径条件控制,以控制角色的根轨迹,同时保持自然的身体运动。实验表明,FD2 将训练计算量减少了 4.6 倍,并将去噪速度提高了 11.29 倍,在长序列上每次更新达到 2.303 毫秒。除了这些效率提升之外,FD2 在运动质量上优于 FD1,并在流式方法中取得了最先进的 FID 分数,在 SEED 上为 0.048,在 HumanML3D 上为 0.053。
cs.CV / 181 / 2609.33176

ABO-Med: Accelerated Bilevel Optimization for Few-Shot Medical Image Classification

ABO-Med:面向小样本医学图像分类的加速双层优化
Shi, Ruoxuan, Yang, Sheng, Su, Zhengxing, Hou, Xiaoyang, Liu, Yating
Abstract
In recent years, bilevel optimization has been widely used in a variety of machine learning tasks. However, prior bilevel optimization algorithms generally require the computation of second-order information, which limits their practical scalability. Only recently has a first-order paradigm for bilevel optimization been established, attaining near-optimal theoretical guarantees for solving bilevel optimization problems. In this paper, we propose ABO-Med, a scalable instantiation of this paradigm for few-shot learning, by incorporating it into the model-agnostic meta-learning (MAML) framework and tailoring it to medical image classification. We also introduce Medical Adaptive RandomAugment (MedRAug), a modality-aware augmentation strategy designed for medical images. Theoretically, ABO-Med establishes the optimality of MAML-type meta-learning approaches. Empirically, ABO-Med outperforms prior baselines on several public medical datasets, with gains of 1.99% to 18.76%, while MedRAug further improves the average accuracy by 2.20% to 6.34%. Additional cross-domain experiments, augmentation ablation studies, backbone ablation studies, and training efficiency analysis further validate the effectiveness and efficiency of the proposed method.
Chinese Translation
近年来,双层优化已被广泛应用于各种机器学习任务中。然而,先前的双层优化算法通常需要计算二阶信息,这限制了其实际可扩展性。直到最近,才建立了双层优化的一阶范式,在求解双层优化问题方面获得了接近最优的理论保证。在本文中,我们提出了 ABO-Med,这是该范式在小样本学习中的可扩展实例化,通过将其融入模型无关元学习(MAML)框架,并针对医学图像分类进行调整。我们还引入了 Medical Adaptive RandomAugment (MedRAug),这是一种为医学图像设计的模态感知增强策略。在理论上,ABO-Med 确立了 MAML 类元学习方法的最优性。在实验上,ABO-Med 在多个公共医学数据集上优于先前基线,提升幅度为 1.99% 至 18.76%,而 MedRAug 进一步将平均准确率提高了 2.20% 至 6.34%。额外的跨域实验、增强消融研究、骨干消融研究和训练效率分析进一步验证了所提方法的有效性和效率。
cs.CV / 182 / 2609.33178

Can Protein-Derived Knowledge Improve Pathology Foundation Models?

蛋白质衍生知识能否提升病理学基础模型?
Zhang, Di, Gong, Zhangpeng, Liu, Jiashuai, Zeng, Zhi, Ge, Jiusong, Yang, Chunze, Ling, Xitong, Yi, Kai, He, Kai, Yu, Weimiao, Crispin-Ortuzar, Mireia, Li, Chen, Gao, Zeyu
Abstract
Molecularly guided pathology foundation models (PFMs) exploit transcriptomic or proteomic information to enrich whole-slide image (WSI) representations, yet effectively leveraging large standalone molecular corpora remains challenging. First, existing molecular foundation models encode protein sequences or single-cell states, not the patient-level bulk expression profiles paired with WSIs. Second, because cross-modal supervision is restricted to paired WSI-omics samples, knowledge from standalone molecular corpora reaches the pathology encoder only indirectly, creating a paired-support bottleneck. To address these challenges, we propose a three-stage framework that decouples proteomic knowledge acquisition from cross-modal transfer, yielding ProSlide, a slide-level hierarchical pathology foundation model. First, to close the modality gap, we pretrain a Proteomic Foundation Encoder (PFE) on 12,695 sample-level bulk protein profiles using virtual profile generation and expression-space multi-view pretraining. Second, we pretrain ProSlide, a patch-region-slide encoder, to predict protein expression from paired WSI-protein samples. Third, to relax the paired-support bottleneck, we introduce Prot2Path, a cross-modal relational distillation objective. For each paired sample, it aligns the similarity distributions of the WSI and its protein profile over a shared, frozen bank of PFE-encoded paired and standalone profiles. We evaluate ProSlide on 12 downstream tasks across breast, lung, and renal cancers. Despite being pretrained with only 2,229 WSIs and 12,695 sample-level protein profiles, ProSlide achieves the highest mean accuracy and AUC within each cancer group.
Chinese Translation
分子引导的病理基础模型(PFMs)利用转录组或蛋白质组信息来丰富全切片图像(WSI)表征,但有效利用大规模独立分子语料仍具挑战性。首先,现有分子基础模型编码的是蛋白质序列或单细胞状态,而非与WSI配对的患者的批量表达谱。其次,由于跨模态监督仅限于配对的WSI-组学样本,来自独立分子语料的知识只能间接到达病理编码器,造成配对支持瓶颈。为解决这些挑战,我们提出一个三阶段框架,将蛋白质组知识获取与跨模态迁移解耦,形成ProSlide,一种切片级分层病理基础模型。首先,为弥合模态差距,我们在12,695个样本级批量蛋白质表达谱上预训练一个蛋白质组基础编码器(PFE),采用虚拟表达谱生成和表达空间多视角预训练。其次,我们预训练ProSlide,一个图像块-区域-切片编码器,从配对的WSI-蛋白质样本中预测蛋白质表达。第三,为缓解配对支持瓶颈,我们引入Prot2Path,一种跨模态关系蒸馏目标。对于每个配对样本,它在一个由PFE编码的配对和独立表达谱组成的共享冻结库上,对齐WSI与其蛋白质表达谱的相似度分布。我们在乳腺癌、肺癌和肾癌的12项下游任务上评估ProSlide。尽管仅使用2,229张WSI和12,695个样本级蛋白质表达谱进行预训练,ProSlide在每种癌症组内均取得了最高的平均准确率和AUC。
cs.CV / 183 / 2609.33190

FocusDrive: Reasoning with Visual Focus for Autonomous Driving

FocusDrive:面向自动驾驶的视觉焦点推理
Liu, Zhiyuan, Ke, Zehong, Tian, Yuanxin, Cheng, Hao, Li, Jinhao, Xing, Yining, Jiang, Yanbo, Xu, Zhenhua, Yu, Wenhao, Wang, Jianqiang
Abstract
Driving decisions depend on both where to focus and how to act on what is seen. Effective driving reasoning must establish which objects matter, where they are, and how they inform the intended action. Text-based rationales can describe a driving response while leaving its correspondence to specific visual evidence implicit. Visual focus provides a concrete starting point for this connection by identifying what matters in the scene and where it is. We propose FocusDrive, a structured multimodal reasoning framework that organizes end-to-end planning around explicit visual focus. It pairs descriptions of decision-relevant objects with image-patch references, bringing explicit visual focus into the reasoning that generates driving plans and trajectories. We first assess this focus representation through driver gaze prediction, then investigate its role in planning reasoning using driving-focus annotations within existing NAVSIM training scenes. Experiments on W3DA and NAVSIM demonstrate competitive gaze prediction and end-to-end planning performance, with FocusDrive improving over text-based chain-of-thought. These results support visual focus as an effective link between scene understanding and driving action.
Chinese Translation
驾驶决策取决于关注何处以及如何对所看到的内容采取行动。有效的驾驶推理必须确立哪些对象重要、它们位于何处,以及它们如何影响预期动作。基于文本的推理依据可以描述驾驶响应,却将其与具体视觉证据的对应关系隐含其中。视觉焦点通过识别场景中重要的内容及其位置,为这种联系提供了一个具体的起点。我们提出了 FocusDrive,一个结构化的多模态推理框架,它围绕显式的视觉焦点组织端到端规划。它将决策相关对象的描述与图像块引用配对,将显式视觉焦点引入生成驾驶规划和轨迹的推理中。我们首先通过驾驶员注视预测来评估这种焦点表示,然后利用现有 NAVSIM 训练场景中的驾驶焦点标注,研究其在规划推理中的作用。在 W3DA 和 NAVSIM 上的实验展示了具有竞争力的注视预测和端到端规划性能,且 FocusDrive 优于基于文本的思维链方法。这些结果支持视觉焦点作为场景理解与驾驶动作之间的有效链接。
cs.CV / 184 / 2609.33202

ReAL: Accelerating Flow Matching through Segment Advancement with Shared Lookahead

ReAL:通过共享前瞻的分段推进加速流匹配
Yin, Xuanhua, Xu, Chuanzhi, Zhou, Haoxian, Mao, Shunqi, Cai, Weidong
Abstract
Flow-matching models generate high-quality images and videos, but repeated neural network evaluations make sampling expensive. Skipping evaluations reduces this cost by extending an available velocity estimate over a longer span. However, local velocity agreement alone does not determine a suitable span, and checking each candidate endpoint adds costly model calls. We introduce ReAL, a training-free sampler that selects how far to advance using one shared lookahead. Our key insight is that the discrepancy between uncorrected and lookahead-corrected endpoint proposals can be computed directly from the observed velocity mismatch and the candidate span beyond the lookahead. This relation provides a span-dependent selection criterion without additional endpoint evaluations. The same lookahead selects the longest passing candidate span, corrects the accepted update, and supplies its velocity as the next starting estimate. After initialization, each regular iteration requires only one fresh evaluation. ReAL uses the pretrained velocity output and original noise schedule, with no additional training or access to internal features. Experiments cover four image-generation backbones, video generation, and image editing. ReAL achieves 4.91x measured speedup on FLUX.1-dev while retaining 97.0% of dense mean ImageReward. On HunyuanVideo, it achieves a 5.49x speedup while maintaining a VBench score close to that of dense sampling.
Chinese Translation
流匹配模型能生成高质量的图像和视频,但重复的神经网络评估使得采样成本高昂。跳过评估通过将可用的速度估计扩展到更长的时间跨度来降低这一成本。然而,仅凭局部速度一致性无法确定合适的跨度,而且检查每个候选端点会增加昂贵的模型调用。我们提出ReAL,一种无需训练的采样器,使用一个共享的前瞻来选择推进的距离。我们的关键见解是,未校正和前瞻校正的端点提议之间的差异可以直接从观察到的速度不匹配和超出前瞻的候选跨度计算出来。这种关系提供了一种依赖于跨度的选择标准,而无需额外的端点评估。同一个前瞻选择最长的通过候选跨度,校正被接受的更新,并将其速度作为下一个起始估计。初始化后,每个常规迭代仅需一次新的评估。ReAL使用预训练的速度输出和原始噪声调度,无需额外训练或访问内部特征。实验涵盖了四个图像生成骨干网络、视频生成和图像编辑。ReAL在FLUX.1-dev上实现了4.91倍的实测加速,同时保持了密集平均ImageReward的97.0%。在HunyuanVideo上,它实现了5.49倍的加速,同时VBench分数接近密集采样。
cs.CV / 185 / 2609.33203

Structured Residual Connectivity Matters for Diffusion Transformers

结构化的残差连接对扩散Transformer至关重要
Liu, Yuhe, Ma, Xinyin, Fang, Gongfan, Liu, Songhua, Wang, Xinchao
Abstract
Diffusion Transformers (DiTs) have established themselves as a scalable backbone for high-fidelity image synthesis. However, unlike U-Net based diffusion models that rely on rigid, hand-crafted skip connections, DiTs predominantly use a uniform residual stream that integrates all preceding layers as a monolithic state. In this work, we rethink residual connections in diffusion transformers and propose to transform them from passive summation into an active retrieval mechanism optimized for image denoising. First, we conduct a systematic analysis of DiT's internal representation, revealing a latent preference for early-layer feature reuse and symmetric layer guidance. Motivated by this, we introduce a structured connectivity design that explicitly integrates local residual connections with long-range pathways. Instead of static skip connections or dense all-layer routing, our method enables each transformer block to selectively ``attend'' to critical earlier representations, dynamically retrieving spatial and semantic cues through direct, differentiable cross-depth paths. Experiments show that our adaptive connectivity leads to faster convergence, with up to $1.73\times$ fewer training iterations, and significant gains in FID and visual quality with less than $0.1\%$ additional parameters, further improving a strong REPA-XL/2 model from $5.9$ to $4.34$ FID without guidance and reaching $1.39$ FID with classifier-free guidance. Our findings suggest that adaptive cross-layer connectivity is a critical yet underexplored factor in diffusion transformers, and that incorporating structured information pathways provides a simple and effective direction for improving scalable generative models.
Chinese Translation
扩散Transformer(DiTs)已成为高保真图像合成的可扩展骨干网络。然而,与依赖刚性、手工设计的跳跃连接的基于U-Net的扩散模型不同,DiTs主要使用统一的残差流,将所有先前层集成为一个整体状态。在这项工作中,我们重新思考了扩散Transformer中的残差连接,并提出将其从被动求和转变为针对图像去噪优化的主动检索机制。首先,我们对DiT的内部表示进行了系统分析,揭示了其对早期层特征重用和对称层引导的潜在偏好。受此启发,我们引入了一种结构化连接设计,明确地将局部残差连接与长程路径相结合。我们的方法不是使用静态跳跃连接或密集的全层路由,而是使每个Transformer块能够选择性地“关注”关键的早期表示,通过直接、可微分的跨深度路径动态检索空间和语义线索。实验表明,我们的自适应连接带来了更快的收敛速度,训练迭代次数最多可减少1.73倍,并且在额外参数少于0.1%的情况下,在FID和视觉质量上取得了显著提升,进一步将强大的REPA-XL/2模型在无引导下的FID从5.9改善至4.34,并在无分类器引导下达到1.39的FID。我们的研究结果表明,自适应跨层连接是扩散Transformer中一个关键但尚未充分探索的因素,并且引入结构化信息路径为改进可扩展生成模型提供了一个简单而有效的方向。
cs.CV / 186 / 2609.33210

Background Gradients Shape Memorization in Flow Matching

背景梯度塑造流匹配中的记忆化
Yin, Xuanhua, Wei, Boyu, Zhang, Shuyi, Mao, Shunqi, Xu, Chuanzhi, Cai, Weidong
Abstract
Repetition is closely associated with memorization in generative models, but how other training images affect the retention and copying of targets remains unclear. We study this question in class-conditioned flow matching, where images outside the target set form the background. At fixed target repetition and same-class background row count, replacing repeated same-class images with distinct images reduces the target extraction rate from 80.7% to 18.0%. To explain this effect, we develop a paired-trajectory framework that isolates target-induced parameter displacement and the background gradient response to it. This response has an exact path-integrated curvature representation, connecting background loss geometry to target learning. Reciprocal response transfer between repeated and distinct backgrounds changes target retention and copying in both directions, establishing the response's causal role. After target removal, the response correction parallel to the target-induced displacement preserves approximately 90% of the copying effects of full response transfer. Directly scaling the displacement also changes copying without further training. The post-removal copying effects of reciprocal transfer are reproduced across datasets and architectures. Together, these results identify the background gradient response as a mechanism through which same-class training data shape the retention of target learning and the reproduction of target images.
Chinese Translation
重复与生成模型中的记忆化密切相关,但其他训练图像如何影响目标的保留和复制仍不清楚。我们在类别条件流匹配中研究这个问题,其中目标集之外的图像构成背景。在固定的目标重复次数和同类背景行数下,将重复的同类图像替换为不同的图像,会使目标提取率从80.7%降低到18.0%。为了解释这一效应,我们开发了一个配对轨迹框架,该框架分离了目标引起的参数位移以及背景梯度对其的响应。这种响应具有精确的路径积分曲率表示,将背景损失几何与目标学习联系起来。在重复背景和不同背景之间进行相互响应迁移,会双向改变目标的保留和复制,从而确立了响应的因果作用。移除目标后,平行于目标引起的位移的响应校正保留了完整响应迁移约90%的复制效应。直接缩放位移也会在无需进一步训练的情况下改变复制。相互迁移在移除后的复制效应在不同数据集和架构上得到重现。总之,这些结果将背景梯度响应确定为一种机制,通过该机制,同类训练数据塑造目标学习的保留和目标图像的再现。
cs.CV / 187 / 2609.33217

RepFlow: Reciprocal Supervision Improves Generation and Representation in Flow Models

RepFlow:相互监督改善流模型中的生成与表示
Zeng, Weili, Tian, Feng, Liu, Shengqi, Yan, Yichao
Abstract
Generative models learn visual structure through denoising, yet their internal states are entangled with both noise level and network depth, making it difficult to obtain a stable visual representation from the generator itself. We introduce RepFlow, which learns such a representation from the generator's evolving computation and uses it to guide generation. Specifically, a separate timestep-free encoder is trained, through a timestep-conditioned predictor, to recover generator states across depths and noise levels from masked clean images. By excluding the reference noise and masked-out content from the encoder's input, we encourage the encoder to distill visual information that is recoverable from the visible image context and predictive of generator states. The representation learned from the generator's evolving states is then fed back to guide and improve the generator, whose updated states provide supervision for further representation learning, forming a reciprocal learning process. This reciprocal process improves multi-step generation and representation quality, as measured by frozen linear probing on ImageNet with latent-space SiT and pixel-space JiT, without an externally pretrained representation teacher. Across the two unconditional settings, FID decreases by $19.2$--$40.8\%$ relative to native training, while linear-probe accuracy improves by 6.80--10.04 percentage points over the best searched raw generator features. The learned representation also serves as a distributional metric for one-step JiT post-training, extending its role from instance-level alignment to distribution-level supervision.
Chinese Translation
生成模型通过去噪学习视觉结构,然而其内部状态与噪声水平和网络深度纠缠在一起,使得从生成器本身获得稳定的视觉表示变得困难。我们提出RepFlow,它从生成器不断演化的计算中学习这样的表示,并用它来引导生成。具体而言,一个独立的无时间步编码器通过一个以时间步为条件的预测器进行训练,以从掩码后的干净图像中恢复不同深度和噪声水平下的生成器状态。通过从编码器输入中排除参考噪声和掩码内容,我们鼓励编码器提取可从可见图像上下文中恢复且能预测生成器状态的视觉信息。从生成器不断演化的状态中学习到的表示随后被反馈回来引导和改进生成器,而生成器更新后的状态为进一步的表示学习提供监督,形成一个相互学习的过程。这种相互过程改善了多步生成和表示质量,这一点通过在ImageNet上使用潜在空间SiT和像素空间JiT的冻结线性探测来衡量,且无需外部预训练的表示教师。在两个无条件设置中,FID相对于原生训练下降了19.2--40.8%,而线性探测准确率比最佳搜索的原始生成器特征提高了6.80--10.04个百分点。学习到的表示还可作为一步JiT后训练的分布度量,将其作用从实例级对齐扩展到分布级监督。
cs.CV / 188 / 2609.33218

Scope-WM: Scoped Computation for Efficient Visual World Models

Scope-WM:面向高效视觉世界模型的范围计算
Li, Chunzheng, Jia, Zesheng, Zhang, Hongda, Tang, Jiaying, Wang, Yuntian, Liu, Siao, Wang, Jin
Abstract
Visual world models enable robotic planning by predicting future observations, but dense latent-state propagation and sample-intensive trajectory optimization incur high inference latency and peak memory usage, limiting real-time deployment on resource-constrained platforms. Existing sparse world-model acceleration methods either rely on unguided token sparsification, which may discard planning-relevant information and restrict achievable sparsity, or introduce heavy auxiliary modules and cumbersome multi-stage training pipelines. In this work, we present Scope-WM, an efficient visual world model that scopes computation to prediction-relevant latent regions and promising action sequences. Scope-WM distills prediction relevance into a lightweight action-conditioned selector and applies full dynamics prediction only to a compact subset of selected tokens. It updates the remaining tokens using a compact summary of foreground states and their changes, allowing the background to perceive foreground dynamics without costly token-to-token interactions. During planning, Scope-WM preserves and reuses high-quality action sequences discovered during the initial MPC search, focusing subsequent search under reduced rollout budgets. The resulting pipeline requires only a one-off selector distillation followed by a single joint training stage for the sparse world model. On the challenging Push-T task, Scope-WM reduces peak GPU memory usage and planning time to $18.1\%$ and $14.3\%$ of those of dense DINO-WM, respectively, corresponding to a $6.97\times$ planning speedup, while maintaining competitive task performance. Further evaluations across five diverse visual planning tasks demonstrate the general applicability of Scope-WM. Code is available at https://github.com/ChunZheng2022/Scope-WM.
Chinese Translation
视觉世界模型通过预测未来观测来实现机器人规划,但密集的潜在状态传播和样本密集的轨迹优化会导致高推理延迟和峰值内存使用,限制了在资源受限平台上的实时部署。现有的稀疏世界模型加速方法要么依赖无引导的令牌稀疏化,这可能丢弃规划相关信息并限制可实现的稀疏性,要么引入沉重的辅助模块和繁琐的多阶段训练流程。在这项工作中,我们提出了Scope-WM,一种高效的视觉世界模型,它将计算限定在预测相关的潜在区域和有前景的动作序列。Scope-WM将预测相关性蒸馏到一个轻量级的动作条件选择器中,并仅对选定令牌的紧凑子集应用完整的动力学预测。它使用前景状态及其变化的紧凑摘要来更新剩余令牌,使背景能够感知前景动力学,而无需昂贵的令牌间交互。在规划过程中,Scope-WM保留并重用初始MPC搜索中发现的高质量动作序列,在减少的推演预算下聚焦后续搜索。由此产生的流程只需要一次性选择器蒸馏,随后是对稀疏世界模型进行单一联合训练阶段。在具有挑战性的Push-T任务上,Scope-WM将峰值GPU内存使用和规划时间分别降低至密集DINO-WM的18.1%和14.3%,对应6.97倍的规划加速,同时保持有竞争力的任务性能。在五个不同的视觉规划任务上的进一步评估证明了Scope-WM的普遍适用性。代码可在https://github.com/ChunZheng2022/Scope-WM获取。
cs.CV / 189 / 2609.33230

AevaScenes: An FMCW LiDAR Dataset and Benchmark for Long-Range Perception

AevaScenes:面向远距离感知的FMCW LiDAR数据集与基准
Narasimhan, Gautham Narayan, Vhavle, Heethesh, Viswanatha, Kumar Bhargav, Reuther, James, Ramanan, Deva
Abstract
FMCW LiDAR measures per-point radial Doppler velocity alongside range, providing a motion cue unavailable in conventional time-of-flight sensors. Exploiting this signal at long range remains understudied. We present an FMCW LiDAR dataset of 575 sequences (57.5K frames) with over 8 million annotated 3D boxes across 16 detection classes and per-point labels across 24 semantic classes, captured by six commercial FMCW LiDAR sensors and six paired 4K cameras across eight Bay Area cities, including 237 nighttime sequences, with annotations extending to 400m. We define a benchmark with three tasks: 3D object detection, scene flow estimation, and semantic segmentation. Detection and scene flow are evaluated across three range bins to 400m, with a public evaluation server. We explore the impact of Doppler measurements on flagship recognition tasks, and find significant improvements up to 2X in detection AP of far-away vehicles and pedestrians, particularly in low-latency single-frame settings. We similarly find scene flow accuracy is significantly improved with Doppler measurements across all ranges. Our dataset and benchmark have been publicly released at https://scenes.aeva.com.
Chinese Translation
FMCW LiDAR在测距的同时测量每个点的径向多普勒速度,提供了传统飞行时间(time-of-flight)传感器所不具备的运动线索。在远距离下利用该信号仍缺乏研究。我们提出了一个FMCW LiDAR数据集,包含575个序列(57.5K帧),覆盖16个检测类别、超过800万个标注3D框,以及覆盖24个语义类别的逐点标签;该数据集由六个商用FMCW LiDAR传感器和六台配对的4K相机在八个湾区城市采集,其中包括237个夜间序列,标注范围延伸至400m。我们定义了一个包含三项任务的基准:3D目标检测、场景流估计和语义分割。检测和场景流在三个距离区间上评估,最远至400m,并提供公开评估服务器。我们探究了多普勒测量对旗舰识别任务的影响,发现远处车辆和行人的检测AP显著提升,最高可达2倍,尤其是在低延迟单帧设置下。我们同样发现,在所有距离范围内,多普勒测量都显著提升了场景流精度。我们的数据集和基准已在 https://scenes.aeva.com 公开发布。
cs.CV / 190 / 2609.33236

PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalation

PARSEE-VAD:基于命题感知推理与流式证据升级的高效免训练在线视频异常检测
Wang, Ji, Zhang, Shuangqing, Xie, Guo-Sen, Zhao, Fang
Abstract
Training-free online video anomaly detection (VAD) with frozen multimodal language models faces two coupled challenges: extracting reliable current-window semantics under causal and computational constraints, and maintaining temporal continuity without repeatedly transmitting high-dimensional history. Encoding history through text can compress visual evidence and introduce semantic bias, whereas retaining visual history expands multimodal context. We introduce PARSEE-VAD, a two-module framework that separates semantic evidence acquisition from score-state evolution. Proposition-Aware Reasoning (PAR) extracts structured propositional evidence from the current causal window and conditionally activates more specific queries when coarse evidence warrants further refinement. By sharing a reusable causal visual prefix across queries, PAR reduces redundant computation through selective execution. Streaming Evidence Escalation (SEE) maps the acquired proposition evidence into a compact score-domain event state through current evidence escalation, then propagates only the resulting bounded state across decisions to support temporal continuity. Experiments on four benchmarks demonstrate strong training-free online performance while selective routing reduces specialist computation and score-state propagation remains sparse. These results support a current-first principle for streaming multimodal inference: resolve present semantics first, then use compact historical state only to repair residual continuity gaps.
Chinese Translation
使用冻结多模态语言模型的免训练在线视频异常检测(VAD)面临两个相互耦合的挑战:在因果和计算约束下提取可靠的当前窗口语义,以及在不反复传输高维历史的情况下保持时间连续性。通过文本编码历史可以压缩视觉证据并引入语义偏差,而保留视觉历史则会扩展多模态上下文。我们提出 PARSEE-VAD,这是一个双模块框架,将语义证据获取与分数状态演化分离。命题感知推理(PAR)从当前因果窗口中提取结构化命题证据,并在粗粒度证据需要进一步细化时有条件地激活更具体的查询。通过在查询之间共享可复用的因果视觉前缀,PAR 通过选择性执行减少冗余计算。流式证据升级(SEE)通过当前证据升级将获取的命题证据映射为紧凑的分数域事件状态,然后仅将得到的有界状态跨决策传播,以支持时间连续性。在四个基准上的实验展示了强大的免训练在线性能,同时选择性路由减少了专家计算,并且分数状态传播保持稀疏。这些结果支持流式多模态推理的“当前优先”原则:先解决当前语义,然后仅使用紧凑的历史状态来修复残余的连续性缺口。
cs.CV / 191 / 2609.33253

VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis

VGGT-Diff:视觉几何遇见扩散,用于稀疏视角新视角合成
Chen, Kangjie, Li, Xiangyu, Zhang, Dongbin, Zheng, Chaoda, Chen, Shijia, Deng, Jinhao, Lin, Hongbin, Wai, Choo Sin, Wang, Minqi, Yang, Minghao, Zhong, Dake, Song, Guorui, Zhang, Yu, Liu, Xianming, Wang, Boyang
Abstract
We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing novel view synthesis (NVS) methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-{\Omega} into a pretrained video diffusion model. Each visual token is associated with a 3D point and confidence, then transformed into query-aligned latent conditions through a confidence-aware Visual Geometry Router (VGR) that preserves front and back surface evidence. These conditions guide joint target-view denoising, while Point-Track Residual Consistency (PTRC) regularizes predicted-clean residuals along reliable 3D tracks, improving multi-view stability. We further introduce robust geometry conditioning, combining training-time regularization with inference-time guidance for improved robustness. Experiments show competitive or state-of-the-art performance across interpolation and extrapolation under different viewpoint difficulties. Our code is available at https://github.com/chenkangjie1123/VGGT-Diff.
Chinese Translation
我们提出 VGGT-Diff,一种用于稀疏视角新视角合成的几何路由多视角扩散模型。现有的新视角合成(NVS)方法面临一个根本性的权衡:基于重建的方法能够保持观测到的几何结构,但难以合成未观测区域;而基于扩散的方法提供了强大的生成先验,却依赖于隐式的源到查询对应关系。VGGT-Diff 通过将来自 VGGT-Ω 的视觉几何隐变量路由到预训练的视频扩散模型中,弥合了这两种范式。每个视觉 token 都与一个 3D 点和置信度相关联,然后通过一个置信度感知的视觉几何路由器(VGR)转换为查询对齐的隐变量条件,该路由器保留了前表面和后表面的证据。这些条件引导联合目标视角去噪,而点轨迹残差一致性(PTRC)沿着可靠的三维轨迹对预测的干净残差进行正则化,从而提高多视角稳定性。我们进一步引入了鲁棒的几何条件化,将训练时的正则化与推理时的引导相结合,以提高鲁棒性。实验表明,在不同视角难度下的插值和外推任务中,该方法均取得了具有竞争力或最先进的性能。我们的代码可在 https://github.com/chenkangjie1123/VGGT-Diff 获取。
cs.CV / 192 / 2609.33261

EngIntervene: Benchmarking Multimodal Engineering State Understanding and Design Intervention Reasoning

EngIntervene:多模态工程状态理解与设计干预推理的基准测试
Zhang, Jinchang, Tao, Yingda, Lin, Jiakai, Lu, Guoyu
Abstract
Multimodal engineering benchmarks largely evaluate static understanding, such as recognizing components, interpreting diagrams, or answering technical questions. This leaves a missing middle between engineering perception and full design generation: whether a model can use an understood system state to reason about relations, constraints, and the consequences of design changes. We introduce \textsc{EngIntervene}, a benchmark for this capability. It contains 3,229 questions across seven engineering domains and organizes evaluation into four levels: state grounding (T1), relational and mechanistic reasoning (T2), constraint-aware diagnosis (T3), and intervention reasoning (T4), which asks whether a proposed modification achieves its target while preserving required constraints. The tasks instantiate a unified engineering state representation spanning objects, relations, constraints, and design objectives, and T2--T4 are scored against structured reference answers with atomic criteria. Across open- and closed-weight multimodal models, stronger grounding does not reliably translate into better diagnosis or intervention, and the best open-weight model trails the best closed model by 14.7 percentage points on the T2--T4 average. Removing or shuffling visual evidence consistently degrades performance, while benchmark-specific supervised fine-tuning improves T1 but not T2--T4. T4 further exposes a large gap between satisfying individual revision criteria and producing a fully valid intervention. Engineering reasoning thus requires not only recovering the current state, but also reliably using it to reason about constraints and post-intervention consequences. Code and benchmark artifacts are available at https://github.com/changcv2021/EngIntervene
Chinese Translation
多模态工程基准测试主要评估静态理解,例如识别组件、解释图表或回答技术问题。这在工程感知和完整设计生成之间留下了一个缺失的中间环节:模型是否能够利用已理解的系统状态来推理关系、约束以及设计变更的后果。我们提出了EngIntervene,一个针对该能力的基准测试。它包含跨七个工程领域的3,229个问题,并将评估组织为四个层次:状态基础(T1)、关系和机制推理(T2)、约束感知诊断(T3)和干预推理(T4),其中T4询问提议的修改是否在保持所需约束的同时实现其目标。这些任务实例化了一个统一的工程状态表示,涵盖对象、关系、约束和设计目标,并且T2-T4根据具有原子标准的结构化参考答案进行评分。在开放权重和闭源权重的多模态模型中,更强的状态基础并不能可靠地转化为更好的诊断或干预,最佳开放权重模型在T2-T4平均值上落后最佳闭源模型14.7个百分点。移除或打乱视觉证据会持续降低性能,而针对基准的监督微调能提升T1,但不能提升T2-T4。T4进一步暴露了满足单个修订标准与产生完全有效的干预之间的巨大差距。因此,工程推理不仅需要恢复当前状态,还需要可靠地利用它来推理约束和干预后的后果。代码和基准测试工件可在 https://github.com/changcv2021/EngIntervene 获取。
cs.CV / 193 / 2609.33264

VehDyn: A Driving World Model Benchmark for Vehicle Dynamics

VehDyn:面向车辆动力学的驾驶世界模型基准
Wang, Tianyi, Du, Wangsheng, Chen, Jiazhou, Zeng, Tianyi, Li, Xiangyu, Byeon, Jiseop, Wang, Yujin, Xu, Yiming, Wang, Yangyang, Gao, Bingzhao, Chen, Sikai, Guo, Zhaomiao, Jiao, Junfeng, Claudel, Christian, Bayen, Alexandre
Abstract
Video world models are emerging as data engines, action planners, and generative simulators for autonomous driving, but existing benchmarks primarily assess visual fidelity and coarse physical plausibility, providing limited evidence on whether generated driving futures obey realistic vehicle kinematics and dynamics. This limitation is further compounded by the lack of datasets in which vehicle, road, maneuver, and speed conditions are independently controlled, and ground-truth vehicle states are recorded in synchrony with videos. We introduce VehDyn, a driving world model benchmark for vehicle dynamics. VehDyn is built on a CARLA-CarSim co-simulation platform where photorealistic rendering is coupled with a validated multi-body dynamics model, and it contains 10,080 configurations from a full factorial design over five vehicle types, four tire-road friction coefficients, three maneuvers, four target speeds, 14 scenes, and three illuminations, each paired with synchronized position, velocity, and attitude sequences. Built on this dataset, VehDyn introduces a hierarchical evaluation framework that measures trajectory alignment, kinematic consistency, and dynamic consistency, and benchmarks 12 state-of-the-art video world models. We further assess the video quality using two established protocols and correlate it with the VehDyn score. Trajectory-level metrics are nearly saturated, with ten of twelve models within 20\% of ground truth, while no model reaches 92\% of ground truth on dynamic consistency, and visual-quality metrics are only weakly correlated with vehicle-dynamics fidelity. DrivingWorld achieves the highest VehDyn score, followed by Cosmos 3 Nano and LTX-Video 2.5, and the VehDyn score agrees closely with human judgment. VehDyn provides a systematic foundation for developing driving world models that are physically consistent and visually realistic.
Chinese Translation
视频世界模型正逐渐成为自动驾驶的数据引擎、动作规划器和生成式模拟器,但现有基准主要评估视觉保真度和粗略的物理合理性,对于生成的驾驶未来是否遵循真实的车辆运动学和动力学,所能提供的证据有限。这一局限因缺乏相应数据集而进一步加剧:这些数据集需要独立控制车辆、道路、机动和速度条件,并且真值车辆状态与视频同步记录。我们提出 VehDyn,一个面向车辆动力学的驾驶世界模型基准。VehDyn 构建于 CARLA-CarSim 联合仿真平台,该平台将照片级真实感渲染与经过验证的多体动力学模型相结合。它包含来自全因子设计的 10,080 种配置,涵盖五种车辆类型、四种轮胎-道路摩擦系数、三种机动、四种目标速度、14 个场景和三种光照,每种配置均配有同步的位置、速度和姿态序列。基于该数据集,VehDyn 引入了一个分层评估框架,用于衡量轨迹对齐、运动学一致性和动力学一致性,并对 12 个最先进的视频世界模型进行了基准测试。我们进一步使用两种已有协议评估视频质量,并将其与 VehDyn 得分进行相关性分析。轨迹层面的指标几乎饱和,十二个模型中有十个与真值的偏差在 20% 以内,但在动力学一致性上,没有模型达到真值的 92%,并且视觉质量指标与车辆动力学保真度仅呈弱相关。DrivingWorld 获得最高的 VehDyn 得分,其次是 Cosmos 3 Nano 和 LTX-Video 2.5,并且 VehDyn 得分与人类判断高度一致。VehDyn 为开发物理一致且视觉真实的驾驶世界模型提供了系统性的基础。
cs.CV / 194 / 2609.33286

InfoEdit: Probing Global Layout Reasoning in Infographic Editing

InfoEdit:探究信息图编辑中的全局布局推理
Yang, Cheng, Shi, Chufan, Wang, Huijuan, Shui, Bo, Wu, Yaokang, Tao, Muzi, Yan, Yibo, Ma, Xuezhe, Berg-Kirkpatrick, Taylor
Abstract
Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be adapted. We refer to this global layout reasoning capability as reflow. Existing image-editing benchmarks neither provide a dedicated setting for structured visual content nor evaluate the reflow capability. We introduce InfoEdit, a novel benchmark of 1,000 infographics across eight logical-relation families, paired with 4,000 editing instructions across four editing tasks, and a reflow-aware evaluation protocol. Across eight frontier editors, only GPT-Image-2 clears 60% average success rate; most models fall below 7%, and no editor exceeds 36% on the Swap-Block task even with perfect target localization. We further show that code-level editing can match the strongest pixel-level editor, revealing complementary strengths across tasks. InfoEdit identifies reflow as a central challenge in structured visual content editing and provides a diagnostic benchmark to facilitate future progress.
Chinese Translation
多模态基础模型能够以生产级质量编辑自然照片,但同样的模型却难以处理诸如信息图之类的结构化视觉内容。与照片不同,信息图通过逻辑关系编码信息;编辑一个元素通常需要调整周围元素。我们将这种全局布局推理能力称为 reflow(重排)。现有的图像编辑基准既没有为结构化视觉内容提供专门的设置,也没有评估 reflow 能力。我们提出 InfoEdit,这是一个新颖的基准,包含来自八个逻辑关系类别的 1,000 张信息图,配有覆盖四项编辑任务的 4,000 条编辑指令,以及一个 reflow 感知的评估协议。在八个前沿编辑器上,只有 GPT-Image-2 的平均成功率超过 60%;大多数模型低于 7%,并且即使目标定位完美,也没有编辑器在 Swap-Block 任务上超过 36%。我们进一步表明,代码级编辑可以匹敌最强的像素级编辑器,揭示了不同任务之间的互补优势。InfoEdit 将 reflow 确定为结构化视觉内容编辑中的一个核心挑战,并提供了一个诊断性基准以促进未来进展。
cs.CV / 195 / 2609.33288

Informative Viewpoint Selection for Episodic-Memory Embodied Question Answering using Omnidirectional Images

基于全景图像的情景记忆具身问答的信息性视角选择
Kitamura, Kaname, Kanezaki, Asako
Abstract
Embodied Question Answering (EQA) requires agents to answer natural language questions about surrounding environments from visual observations. In this work, we focus on open-vocabulary episodic-memory EQA (EM-EQA), where an agent answers free-form questions using recorded observation histories. Omnidirectional images are promising for this task, as they provide wide field-of-view observations that can capture surrounding context without requiring explicit camera rotations. However, omnidirectional images introduce two challenges for EQA: (i) equirectangular projection causes severe geometric distortion that degrades vision-language model (VLM) recognition accuracy, and (ii) feeding equirectangular images directly into VLMs introduces excessive irrelevant background information, reducing answer accuracy and increasing the visual-token burden. To address these challenges, we propose a viewpoint selection method for EM-EQA using omnidirectional images. Our method converts equirectangular observations into perspective views via cubemap projection, estimates question-conditioned relevance with fine-tuned BLIP-2, and selects informative and diverse viewpoints through diversity-aware greedy selection. Experiments on the Habitat-Matterport 3D (HM3D) subset of OpenEQA show that our method achieves state-of-the-art model performance among the reported model results with equirectangular observations. Moreover, after removing rotation views, which reduces observation frames by 65.5%, our method largely maintains its answer accuracy.
Chinese Translation
具身问答(EQA)要求智能体根据视觉观察回答关于周围环境的自然语言问题。在本工作中,我们关注开放词汇情景记忆EQA(EM-EQA),其中智能体使用记录的观察历史回答自由形式的问题。全景图像对此任务具有前景,因为它们提供宽视场观察,无需显式相机旋转即可捕获周围上下文。然而,全景图像给EQA带来两个挑战:(i)等距柱状投影导致严重的几何畸变,降低视觉语言模型(VLM)的识别准确率;(ii)将等距柱状图像直接输入VLM会引入过多不相关的背景信息,降低答案准确率并增加视觉token负担。为解决这些挑战,我们提出一种使用全景图像的EM-EQA视角选择方法。我们的方法通过立方体贴图投影将等距柱状观察转换为透视视图,使用微调的BLIP-2估计问题条件下的相关性,并通过多样性感知的贪婪选择选择信息丰富且多样的视角。在OpenEQA的Habitat-Matterport 3D(HM3D)子集上的实验表明,我们的方法在使用等距柱状观察的报告模型结果中达到了最先进的模型性能。此外,在移除旋转视图(减少了65.5%的观察帧)后,我们的方法在很大程度上保持了其答案准确率。
cs.CV / 196 / 2609.33304

Relevance Does Not Imply Applicability: Experience Activation for Personal GUI Agents

相关性不意味着适用性:面向个人 GUI 智能体的经验激活
Zhang, Fuyao, Wang, Xuan, Li, Zherui, Zhang, Jiaming, Huang, Longtao, Lim, Wei Yang Bryan
Abstract
Personal Graphical User Interface (GUI) agents rely on interaction history to infer what a user wants from ambiguous instructions and to anticipate recurring routines. Existing approaches retrieve task-relevant history and append it to the policy's context, implicitly assuming that experience relevant to a task remains useful for each decision within it. We find that this help is largely spent at the first decision: retrieved history strongly improves the opening step of an episode, yet provides little sustained benefit over the remaining 90\% of steps, and offers weak guidance on whether a proactive suggestion is warranted. A relevant record may tell the agent where to begin, but not which past action applies to the current screen or whether a routine is due now. The underlying issue is that relevance does not imply applicability}: relevance is determined at the task level, whereas applicability depends on the situation at decision time. We therefore recast personalization as experience activation and introduce ExpActivator, a training-free framework that activates only the experience applicable to the current situation. During execution, ExpActivator matches each new screen to historical states in the frozen GUI backbone's latent space and supplies the corresponding action as a reference. Before execution, it activates a recurring intent only when the current time and scenario provide sufficient support, and otherwise abstains. Across four GUI backbones, ExpActivator improves within-trajectory step success by 28\% on average, achieves the best personalized execution on every backbone while using about one-fifth as many history tokens, and reaches approximately 2.3$\times$ the Matthews correlation coefficient of the strongest proactive baseline. Experience pays where it is activated, not where it is appended.
Chinese Translation
个人图形用户界面(GUI)智能体依赖交互历史来从模糊指令中推断用户意图,并预测重复出现的例行程序。现有方法检索与任务相关的历史,并将其附加到策略的上下文中,隐含地假设与任务相关的经验对其中的每个决策仍然有用。我们发现,这种帮助主要用在第一个决策上:检索到的历史强烈改善了回合的开场步骤,但在剩余90%的步骤中几乎没有持续收益,并且对是否值得主动建议提供的指导很弱。一条相关记录可能告诉智能体从哪里开始,但不会告诉它哪个过去的动作适用于当前屏幕,或者某个例行程序现在是否到期。根本问题在于,相关性并不意味着适用性:相关性是在任务层面确定的,而适用性取决于决策时的情境。因此,我们将个性化重新定义为经验激活,并引入 ExpActivator,一个无需训练的框架,仅激活适用于当前情境的经验。在执行过程中,ExpActivator 将每个新屏幕与冻结的 GUI 主干网络潜在空间中的历史状态进行匹配,并提供相应的动作作为参考。在执行之前,只有当当前时间和场景提供足够支持时,它才激活重复出现的意图,否则弃权。在四个 GUI 主干网络上,ExpActivator 将轨迹内步骤成功率平均提高28%,在每个主干网络上实现了最佳的个性化执行,同时使用的历史 token 数约为五分之一,并且达到了最强主动基线的马修斯相关系数的约2.3倍。经验在激活之处产生回报,而非附加之处。
cs.CV / 197 / 2609.33306

LoopTrack: A Simple Baseline for Parameter-Efficient Transformer Tracking

LoopTrack:参数高效Transformer跟踪的简单基线
Peng, Liang, Li, Chenxiao, Zhang, Libo, Dong, Xingping, Fan, Heng
Abstract
Current Transformer-based tracking methods typically stack multiple Transformer blocks with separate parameters to model interactions between the target template and the search region for target localization. These trackers often incur substantial parameter overhead from stacked blocks, making their deployment on resource-limited devices difficult. To address this, we propose a parameter-efficient Transformer tracking framework, dubbed LoopTrack, which repeatedly applies a set of Transformer blocks with shared parameters to interact features in a looped architecture for tracking, significantly reducing the number of parameters. To further exploit target cues, we present two lightweight designs, including target-aware looping (TAL) and gated target memory (GTM). The former applies intermediate target information generated by one loop to guide feature interaction in the subsequent loop, enabling progressive feature refinement, while the latter maintains a compact memory across frames, which is incorporated into the loop process to provide long-term information to the tracker, mitigating temporal drift in tracking. Compared to existing Transformer trackers, LoopTrack enables multiple rounds of feature interaction with fewer model parameters, making it resource-friendly for deployment. In extensive experiments on multiple datasets, LoopTrack shows a favorable accuracy-parameter trade-off. In particular, our LoopTrack$_{\rm One}$, with a single shared Transformer block, achieves 66.2\% SUC score on LaSOT with only 3.4M parameters, while LoopTrack$_{\rm Three}$, using three shared blocks, achieves 69.3\% SUC score with 6.4M parameters, surpassing existing parameter-efficient tracking methods with comparable or larger model size. With LoopTrack, we aim to establish a simple yet strong baseline for parameter-efficient Transformer tracking. Our code and models will be released.
Chinese Translation
当前基于Transformer的跟踪方法通常堆叠多个具有独立参数的Transformer块,以建模目标模板和搜索区域之间的交互,用于目标定位。这些跟踪器通常因堆叠的块而带来大量的参数开销,使其难以部署在资源受限的设备上。为了解决这个问题,我们提出了一种参数高效的Transformer跟踪框架,称为LoopTrack,它重复应用一组共享参数的Transformer块,在循环架构中交互特征以进行跟踪,显著减少了参数数量。为了进一步利用目标线索,我们提出了两种轻量级设计,包括目标感知循环(TAL)和门控目标记忆(GTM)。前者应用一个循环生成的中间目标信息来指导后续循环中的特征交互,实现渐进式特征细化;后者跨帧维护一个紧凑的记忆,将其融入循环过程,为跟踪器提供长期信息,减轻跟踪中的时间漂移。与现有的Transformer跟踪器相比,LoopTrack以更少的模型参数实现了多轮特征交互,使其在部署时对资源友好。在多个数据集上的大量实验中,LoopTrack展示了良好的精度-参数权衡。特别地,我们的LoopTrack_One使用单个共享Transformer块,在LaSOT上以仅3.4M参数实现了66.2%的SUC分数,而LoopTrack_Three使用三个共享块,以6.4M参数实现了69.3%的SUC分数,超越了具有相当或更大模型尺寸的现有参数高效跟踪方法。通过LoopTrack,我们旨在为参数高效的Transformer跟踪建立一个简单而强大的基线。我们的代码和模型将会发布。
cs.CV / 198 / 2609.33318

PIC-UIE: Predicting Image-Adaptive Corrections for Lightweight Underwater Image Enhancement

PIC-UIE:预测图像自适应校正以实现轻量级水下图像增强
Zhu, Cunhao, Xu, Dongliang, Kong, Xiangtao, Lu, Xiaoyan, Wang, Tianyu, Yao, Yue
Abstract
Underwater image enhancement (UIE) aims to restore visibility, color fidelity, and structural detail from images degraded by wavelength-dependent attenuation and backscatter. State-of-the-art UIE methods often rely on large backbones and dense image-to-image prediction, limiting their practicality for edge deployment. Moreover, operating entirely in a single color space couples degradation estimation with luminance and chroma correction. To address these challenges, we propose PIC-UIE, a lightweight predictor--executor framework that predicts image-adaptive corrections from a fixed $256\times256$ RGB thumbnail and applies them to the native-resolution input in the YCbCr color space. The predictor produces seven outputs, organized into spatial correction, nonlinear luminance and coupled chroma mapping, and image-level color calibration. A depth map regularizes the transmission proxy during training, whereas inference uses only the RGB input. With 9,486 parameters and 0.094 GFLOPs at $256\times256$, PIC-UIE achieves 24.137 dB PSNR and 0.9216 SSIM on UIEB-90 and 21.320 dB PSNR on zero-shot LSUI. It further processes native 4K images at 55.0 FPS under the comparison protocol. These results show that structured correction prediction provides an effective and practical alternative to dense RGB reconstruction for underwater image enhancement.
Chinese Translation
水下图像增强(UIE)旨在从受波长相关衰减和后向散射退化的图像中恢复可见性、色彩保真度和结构细节。当前最先进的UIE方法通常依赖大型骨干网络和密集的图像到图像预测,限制了其在边缘部署中的实用性。此外,完全在单一颜色空间中操作会将退化估计与亮度和色度校正耦合在一起。为了解决这些挑战,我们提出了PIC-UIE,一种轻量级的预测器-执行器框架,它从固定的256×256 RGB缩略图预测图像自适应校正,并将其应用于YCbCr颜色空间中的原始分辨率输入。预测器产生七个输出,分为空间校正、非线性亮度和耦合色度映射,以及图像级颜色校准。深度图在训练期间正则化透射率代理,而推理仅使用RGB输入。在256×256下,PIC-UIE具有9,486个参数和0.094 GFLOPs,在UIEB-90上达到24.137 dB PSNR和0.9216 SSIM,在零样本LSUI上达到21.320 dB PSNR。在比较协议下,它进一步以55.0 FPS处理原生4K图像。这些结果表明,结构化校正预测为水下图像增强提供了一种有效且实用的替代方案,可替代密集RGB重建。
cs.CV / 199 / 2609.33325

VisionHOPE: Visual Backbones as Self-Modifying Learning Systems

VisionHOPE:视觉骨干网络作为自修改学习系统
Peng, Siran, Zhang, Tianshuo, Fu, Tianyu, Zhao, Weisong, Zhang, Haoyuan, Zhao, Jiankuo, Wu, Minghui, Jiang, Ping, Zhu, Xiangyu, Zhao, Chenxu, Lei, Zhen
Abstract
Visual backbones have evolved from Convolutional Neural Networks (CNNs) with local aggregation to Vision Transformers (ViTs) with global interactions, State-Space Models (SSMs) with input-dependent state transitions, and Test-Time Training (TTT) layers that adapt an inner learner while processing an image. Across this progression, visual computation has become increasingly adaptive to each input, yet the rules governing that adaptation remain largely prescribed by the trained backbone. We introduce VisionHOPE, the first generic visual backbone formulated as a self-modifying learning system, in which what the model remembers and how it learns co-evolve within an image. Building on the self-referential construction of Nested Learning (NL), VisionHOPE realizes this co-evolution through five coupled memories that store content, generate key and value representations, and govern learning rate and retention. These memories evolve jointly as visual context accumulates along each scan. However, directly applying the unconstrained self-referential update to a visual backbone leads to instability. We therefore derive a stability-matched step-size control scheme that combines a soft cap on self-referential injection with a spectral clamp on the retained memory transition, and prove that the resulting memory dynamics are non-expansive along each scan. For two-dimensional feature maps, we adapt NL's chunk formulation by aligning chunks with image rows and columns across four directional scans. The proposed VisionHOPE achieves competitive results on ImageNet-1K, COCO, and ADE20K, establishing self-modifying learning systems as a practical foundation for general-purpose visual backbones. The code is available at https://github.com/PSRben/VisionHOPE.
Chinese Translation
视觉骨干网络已经从具有局部聚合的卷积神经网络(CNN)发展到具有全局交互的视觉Transformer(ViT)、具有输入依赖状态转换的状态空间模型(SSM),以及在处理图像时自适应调整内部学习器的测试时训练(TTT)层。在这一演进过程中,视觉计算对每个输入的适应性越来越强,然而支配这种适应的规则在很大程度上仍由训练好的骨干网络所规定。我们提出了VisionHOPE,这是第一个被形式化为自修改学习系统的通用视觉骨干网络,其中模型记忆的内容和学习的方式在单张图像内共同演化。基于嵌套学习(NL)的自指构造,VisionHOPE通过五个耦合记忆来实现这种共同演化,这些记忆存储内容、生成键和值表示,并控制学习率和保留率。随着视觉上下文沿每次扫描不断累积,这些记忆共同演化。然而,将无约束的自指更新直接应用于视觉骨干网络会导致不稳定。因此,我们推导了一种稳定性匹配的步长控制方案,该方案结合了对自指注入的软上限和对保留记忆转移的谱钳制,并证明了由此产生的记忆动态在每次扫描中是非扩张的。对于二维特征图,我们通过将块与图像行和列在四个方向扫描中对齐,来适配NL的分块公式。所提出的VisionHOPE在ImageNet-1K、COCO和ADE20K上取得了具有竞争力的结果,确立了自修改学习系统作为通用视觉骨干网络实用基础的地位。代码可在https://github.com/PSRben/VisionHOPE获取。
cs.CV / 200 / 2609.33330

FeCoSplat: Feedback-Guided Compression for Feed-Forward 3D Gaussian Splatting

FeCoSplat:面向前馈3D高斯泼溅的反馈引导压缩
Li, Yuxuan, Chen, Yihang, Zhang, Yufeng, Cai, Jianfei, Lin, Weiyao
Abstract
Feed-forward 3D Gaussian Splatting (3DGS) enables efficient novel-view synthesis from sparse multi-view images, yet its representations remain costly to store and transmit. Existing approaches compress either the input images, incurring heavy receiver-side reconstruction, or the reconstructed Gaussian primitives, which are difficult to compress due to their heterogeneous and irregular attributes. We instead compress compact intermediate features, providing a better balance between compression efficiency and receiver-side complexity. Based on this paradigm, we propose FeCoSplat, a feedback-guided compression framework for feed-forward 3DGS. FeCoSplat first compresses multi-view features to obtain an intermediate 3DGS, whose rendered views are used as feedback to guide a second-stage compression for further refinement. The resulting bitstreams are decoded into a compact implicit state, from which the final Gaussian primitives are reconstructed with a lightweight predictor. Experiments demonstrate that FeCoSplat achieves favorable rate--distortion performance, particularly at low bitrates, while requiring only 3.45M parameters for receiver-side Gaussian reconstruction. Code will be released soon.
Chinese Translation
前馈3D高斯泼溅(Feed-Forward 3D Gaussian Splatting, 3DGS)能够从稀疏多视图图像中高效地进行新视角合成,但其表示在存储和传输方面仍然成本高昂。现有方法要么压缩输入图像,导致接收端重建负担重,要么压缩重建的高斯基元,这些基元因其异构和不规则的属性而难以压缩。我们转而压缩紧凑的中间特征,在压缩效率和接收端复杂度之间提供了更好的平衡。基于这一范式,我们提出了FeCoSplat,一种用于前馈3DGS的反馈引导压缩框架。FeCoSplat首先压缩多视图特征以获得一个中间3DGS,其渲染视图被用作反馈来指导第二阶段的压缩以进一步细化。生成的比特流被解码为紧凑的隐式状态,然后使用轻量级预测器从中重建最终的高斯基元。实验表明,FeCoSplat实现了良好的率失真性能,特别是在低比特率下,同时接收端高斯重建仅需3.45M参数。代码即将发布。
cs.CV / 201 / 2609.33338

OPERA: A Unified Omnimodal Progressive Spatio-Temporal Reasoning Agent for Referring Video Segmentation

OPERA:用于指代视频分割的统一全模态渐进式时空推理智能体
Ni, Jingchen, Wang, Yuji, Yan, Shannan, Li, Haoru, Chen, Sitong, Yuan, Chun
Abstract
Referring video segmentation with heterogeneous multimodal queries---spanning text, audio, and reference images---demands both robust cross-modal understanding and precise spatio-temporal reasoning. We propose OPERA (Omnimodal Progressive spatio-tEmporal Reasoning Agent), a unified reasoning agent built on a single MLLM that performs dual-axis progressive reasoning via three specialized stages. Along the temporal axis, a Temporal Reasoning Agent narrows the frame search space through coarse-to-fine filtering to identify the most informative key frame. Along the spatial axis, a Distillation Agent establishes what to locate via cross-modal semantic distillation, and a Grounding Agent enhanced with GRPO determines where the target appears, with dense mask propagation completing the pixel-level output. OPERA sets a new state of the art on OmniAVS and Ref-AVS and transfers zero-shot to standard referring video segmentation benchmarks.
Chinese Translation
指代视频分割涉及异构多模态查询——涵盖文本、音频和参考图像——需要鲁棒的跨模态理解和精确的时空推理。我们提出OPERA(全模态渐进式时空推理智能体),一个基于单一MLLM构建的统一推理智能体,通过三个专门阶段进行双轴渐进式推理。在时间轴上,时间推理智能体通过由粗到细的过滤缩小帧搜索空间,以识别信息量最大的关键帧。在空间轴上,蒸馏智能体通过跨模态语义蒸馏确定要定位的内容,而增强有GRPO的定位智能体确定目标出现的位置,并通过密集掩码传播完成像素级输出。OPERA在OmniAVS和Ref-AVS上创造了新的最先进水平,并零样本迁移到标准指代视频分割基准上。
cs.CV / 202 / 2609.33344

ReLoc: Rethinking Scene Coordinate Regression Architecture for Robust Outdoor LiDAR-based Localization

ReLoc:重新思考用于鲁棒户外LiDAR定位的场景坐标回归架构
Moon, Heejoon, Cho, Yurim, Hong, Je Hyeong
Abstract
Scene Coordinate Regression (SCR) has recently emerged as a promising approach for LiDAR-based localization, achieving accurate localization without requiring an explicit 3D map. Despite their effectiveness, existing SCR methods rely on scene classification-based global embedding that struggles to provide fine-grained discrimination among nearby locations. Moreover, their reliance on uniform sampling of local features during training assigns equal importance to all points, thereby inadvertently propagating features from dynamic objects or unstable regions and potentially degrading training stability. In this paper, we present ReLoc, a revamped SCR architecture that can effectively address these limitations. First, we redesign the global embedding module by combining learnable context tokens with a feature aggregator to capture richer and more discriminative scene context. Second, we introduce an attention-based local feature enhancement module to mitigate the impact of noisy local features while encouraging context-consistent structures, yielding more robust local feature representations. Experimental results on two large-scale outdoor datasets demonstrate that our approach achieves state-of-the-art accuracy over previous SCR-based methods while maintaining real-time inference performance.
Chinese Translation
场景坐标回归(SCR)最近作为一种有前景的基于LiDAR的定位方法出现,无需显式3D地图即可实现精确定位。尽管有效,现有的SCR方法依赖于基于场景分类的全局嵌入,难以对邻近位置提供细粒度的区分。此外,它们在训练期间依赖于局部特征的均匀采样,对所有点赋予同等重要性,从而无意中传播了动态物体或不稳定区域的特征,并可能降低训练稳定性。在本文中,我们提出了ReLoc,一种改进的SCR架构,能够有效解决这些限制。首先,我们通过将可学习的上下文令牌与特征聚合器相结合,重新设计了全局嵌入模块,以捕获更丰富、更具判别性的场景上下文。其次,我们引入了一个基于注意力的局部特征增强模块,以减轻噪声局部特征的影响,同时鼓励上下文一致的结构,从而产生更鲁棒的局部特征表示。在两个大规模户外数据集上的实验结果表明,我们的方法在保持实时推理性能的同时,实现了优于先前基于SCR方法的最先进精度。
cs.CV / 203 / 2609.33353

Focus and Supplement: Dual-Enhanced Vision Transformer for Multi-Class Anomaly Classification

聚焦与补充:用于多类异常分类的双增强Vision Transformer
Li, Xurui, Xu, Enjie, Li, Chenzhou, Zeng, Shilei, Huang, Dayou, Ma, Tianyi, Zhou, Yu
Abstract
Multi-class anomaly classification in industrial vision remains challenging due to noisy/incomplete anomaly representations and the unknown number of anomaly classes. To overcome this, we propose MACO, a novel multi-class anomaly classification framework that learns comprehensive representations and dynamically estimates class number without prior knowledge. First, a soft-focus attention uses anomaly maps to concentrate on relevant abnormal regions, while suppressing background noise. Second, auxiliary classification ([A-CLS]) tokens complement the [CLS] token. They collectively attend to diverse anomaly sub-regions, yielding more holistic and discriminative features. These [A-CLS] tokens are also effective across more tasks and domains. To infer the class number, we propose Correlation-based Number Estimation strategy. It computes the average correlation among labeled classes and transfers its separability cue to the unlabeled set. Experiments on MVTec AD and MTD datasets demonstrate our superiority. Under known class number, MACO improves ARI by 6.5% and $\textbf{16.3%}$ on both datasets, respectively. In the more challenging unknown number scenario, it achieves an $\textbf{11.2%}$ NMI gain on MTD and outperforms existing number estimation strategies by $\textbf{24.1%}$ UPS on MVTec AD. Code will be released at https://github.com/HUST-SLOW/MACO.
Chinese Translation
工业视觉中的多类异常分类仍然具有挑战性,原因在于异常表示的噪声/不完整以及异常类别数量的未知。为了克服这一点,我们提出了MACO,一种新颖的多类异常分类框架,它无需先验知识即可学习全面的表示并动态估计类别数量。首先,一种软聚焦注意力利用异常图来关注相关的异常区域,同时抑制背景噪声。其次,辅助分类([A-CLS])token补充了[CLS] token。它们共同关注不同的异常子区域,产生更全面和更具判别性的特征。这些[A-CLS] token在更多任务和领域中也有效。为了推断类别数量,我们提出了基于相关性的数量估计策略。它计算标记类别之间的平均相关性,并将其可分离性线索转移到未标记集。在MVTec AD和MTD数据集上的实验证明了我们的优越性。在已知类别数量下,MACO在两个数据集上分别将ARI提高了6.5%和$\textbf{16.3%}$。在更具挑战性的未知数量场景中,它在MTD上实现了$\textbf{11.2%}$的NMI提升,并在MVTec AD上以$\textbf{24.1%}$的UPS优于现有的数量估计策略。代码将在https://github.com/HUST-SLOW/MACO发布。
cs.CV / 204 / 2609.33359

When Does Geometric View Synthesis Help Wine Label Retrieval? A Public One-Shot Benchmark Across Self-Supervised and Vision-Language Backbones

几何视图合成何时有助于葡萄酒标签检索?一项跨自监督和视觉-语言骨干网络的公共单样本基准测试
Huang, Yueh-Cheng
Abstract
Geometric view synthesis can expand a single wine-label photograph into a training set, but its value with pretrained image encoders is unclear. We study this on a public WineSensed-derived benchmark of 1,000 classes, one enrollment photograph per class, and 4,295 real queries. With the earlier DINO vision transformer (ViT-S/16) recipe, geometric views raise top-1 accuracy from 34.1% to 62.6-63.7%, about three times the gain from two-dimensional (2D) augmentation. Frozen SigLIP 2-B already reaches 94.7%. A linear head over its frozen features gains 1.2-1.3 percentage points with the two geometric pipelines localized by the Segment Anything Model (SAM), while the other pipelines gain an inconclusive 0.3-0.6 points. Low-rank adaptation (LoRA) and validation-selected full fine-tuning show no clear gain within the reported confidence intervals; fixed-budget full fine-tuning loses 9-24 points. SAM localization supplies all six views for 99% of sources, compared with 43% for the edge-based front end. Recognition differences between the two cylinder constructions depend on the training recipe and are confounded by their crop and canvas conventions. Rendered-cylinder tests show different responses to source tilt, but an uncalibrated rim-ratio proxy establishes no corresponding trend in recognition on real photographs. An author-confirmed audit of 50 residual errors identifies 21 query-enrollment appearance mismatches, without establishing an irreducible error rate. These results support geometric synthesis for the tested self-supervised recipe and a smaller benefit through frozen-feature adaptation of the text-supervised encoder.
Chinese Translation
几何视图合成可以将单张葡萄酒标签照片扩展为训练集,但其在预训练图像编码器上的价值尚不明确。我们在一个源自WineSensed的公共基准上研究这一问题,该基准包含1,000个类别,每类一张注册照片,以及4,295个真实查询。使用较早的DINO视觉Transformer(ViT-S/16)方案,几何视图将top-1准确率从34.1%提升至62.6-63.7%,约为二维(2D)数据增强带来增益的三倍。冻结的SigLIP 2-B已达到94.7%。在其冻结特征上训练的线性分类头,在使用由Segment Anything Model (SAM)定位的两个几何流程时,获得了1.2-1.3个百分点的提升,而其他流程仅获得0.3-0.6个百分点的非结论性提升。低秩适应(LoRA)和验证集选择的完全微调在报告的置信区间内未显示出明显增益;固定预算的完全微调损失了9-24个百分点。SAM定位为99%的源提供了全部六个视图,而基于边缘的前端仅为43%。两种圆柱体构建之间的识别差异取决于训练方案,并且受到其裁剪和画布约定的混杂影响。渲染圆柱体测试显示对源倾斜的响应不同,但一个未校准的边缘比率代理在真实照片的识别中未建立相应的趋势。对50个残余误差的作者确认审计识别出21个查询-注册外观不匹配,但未建立不可约误差率。这些结果支持几何合成在测试的自监督方案中的有效性,以及通过文本监督编码器的冻结特征适应获得的较小收益。
cs.CV / 205 / 2609.33384

PulseQuant: Propagation-Guided Subspace Correction for 4-Bit Video Diffusion Transformers

PulseQuant:面向4比特视频扩散Transformer的传播引导子空间校正
Wang, Yutong, Ge, Xingtong, Liu, Enhuai, Wang, Yunke, Xue, Tianfan, Chen, Xinyuan, Xu, Chang
Abstract
Quantization errors in video diffusion transformers can be amplified or attenuated by subsequent denoising updates, making local reconstruction error an incomplete predictor of final impact. We introduce PulseQuant, a 4-bit post-training quantization method that combines trajectory sensitivity with activation geometry to guide offline calibration. Isolated block--step interventions estimate propagation risk, which prioritizes sensitive trajectory states during row-radius selection. With these radii fixed, response-subspace correction uses neighboring-code edits to reduce residual components along dominant activation directions. Both stages preserve the original 4-bit weight representation. Controlled interventions show that short-horizon propagated error predicts final latent error more reliably than immediate block-output error, supporting calibration beyond local reconstruction objectives. Evaluations on Wan models, Self Forcing, and MiniMax-H3 demonstrate improvements in key consistency and dense-reference metrics while remaining competitive on other attributes across model scales and generation paradigms.
Chinese Translation
视频扩散Transformer中的量化误差可能被后续去噪更新放大或衰减,使得局部重建误差成为最终影响的不完整预测因子。我们提出PulseQuant,一种4比特训练后量化方法,结合轨迹敏感性与激活几何来指导离线校准。孤立的块-步干预估计传播风险,从而在行半径选择期间优先考虑敏感轨迹状态。在这些半径固定后,响应子空间校正使用相邻码编辑来减少沿主导激活方向的残差分量。两个阶段都保留原始4比特权重表示。受控干预表明,短时域传播误差比即时块输出误差更可靠地预测最终潜在误差,支持超越局部重建目标的校准。在Wan模型、Self Forcing和MiniMax-H3上的评估表明,在关键一致性和密集参考指标上有所改进,同时在模型规模和生成范式上在其他属性上保持竞争力。
cs.CV / 206 / 2609.33399

SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation

SciGen-Verifier:面向科学图像生成可解释验证的多模态推理器
Chen, Jiali, Lin, Zhengteng, Wang, Zuqi, Lin, Shirong, Yu, Xi, Hei, Xusen, Fu, DingBa, Xie, Jiayuan, Cai, Yi
Abstract
In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have enabled scientific image generation, yet verifying the correctness of these specialized visual outputs remains a critical bottleneck: errors often arise from intricate domain knowledge, structural reasoning, and multi-step instruction rather than surface-level artifacts. Existing verifiers mainly target natural images and compress judgement into scalar scores, leaving scientific coverage and explainable feedback for error correction underexplored. To bridge this gap, we make three main contributions. (1) We construct SciGen-Verify, a benchmark dedicated to explainable verification of scientific image generation, spanning instruction following, multidisciplinary reasoning, and world knowledge domains. It contains a three-tier hierarchical protocol over the binary judgement, supporting explanation, and corrective editing instruction. (2) We develop SciGen-Verifier, a reasoning-driven multimodal verifier trained via cold-start supervised fine-tuning followed by a curriculum-based two-stage reinforcement learning pipeline. The rubric-guided process rewards first strengthen scientific reasoning exploration and outcome rewards subsequently align output with ground-truth annotation. (3) On SciGen-Verify, SciGen-Verifier achieves competitive performance against much larger proprietary models. It further serves as a practical online critic for iterative image rectification.
Chinese Translation
在现实教育中,解题方案不仅用文字表达,还常用绘图表达——电路图、几何作图、函数图——教师必须像批改文字一样仔细地批改绘图。统一多模态模型的最新进展使科学图像生成成为可能,但验证这些专业视觉输出的正确性仍然是一个关键瓶颈:错误往往源于复杂的领域知识、结构化推理和多步指令,而非表面伪影。现有验证器主要针对自然图像,并将判断压缩为标量分数,导致科学领域的覆盖范围以及用于错误纠正的可解释反馈尚未得到充分探索。为弥补这一空白,我们作出三项主要贡献。(1) 我们构建了 SciGen-Verify,一个专用于科学图像生成可解释验证的基准,涵盖指令遵循、多学科推理和世界知识领域。它包含一个三层分层协议,涉及二值判断、支持性解释和纠正性编辑指令。(2) 我们开发了 SciGen-Verifier,一个推理驱动的多模态验证器,通过冷启动监督微调,随后基于课程的两阶段强化学习流程进行训练。评分标准引导的过程奖励首先增强科学推理探索,结果奖励随后使输出与真实标注对齐。(3) 在 SciGen-Verify 上,SciGen-Verifier 相较于规模大得多的专有模型取得了有竞争力的性能。它还可作为实用的在线评论器,用于迭代图像矫正。
cs.CV / 207 / 2609.33400

Groupwise Selective State-Space Filtering for Accurate and Streaming Action Boundary Detection

用于准确且流式动作边界检测的分组选择性状态空间滤波
Çelik, Mustafa Bora
Abstract
Action boundary detection partitions untrimmed video into intervals without assigning action classes. We present a boundary-detection adapter operating on pre-extracted video features, learning temporal representations via groupwise selective scans. Learned group fusion and temporal modeling convert these into transition scores, which are decoded into boundary timestamps. Trained with boundary-time supervision, the class-agnostic model is evaluated on Breakfast, GTEA, and 50Salads using temporal tolerances and bipartite matching, achieving boundary $F_1$ scores of 0.457, 0.622, and 0.611. A stateful variant enables feature-streaming inference with zero neural look-ahead, one-sample peak confirmation, and bounded memory. Downstream systems can subsequently assign s
Chinese Translation
动作边界检测将未剪辑视频划分为区间,但不分配动作类别。我们提出一种边界检测适配器,作用于预提取的视频特征,并通过分组选择性扫描学习时序表示。学习到的组融合与时序建模将这些表示转换为转移分数,再解码为边界时间戳。在边界时间监督下训练,该类别无关模型在 Breakfast、GTEA 和 50Salads 上使用时间容差与二分匹配进行评估,边界 F1 分数分别达到 0.457、0.622 和 0.611。一种有状态变体支持特征流式推理,具有零神经前瞻、单样本峰值确认以及有界内存。下游系统随后可以分配 s
cs.CV / 208 / 2609.33402

VaME: Exploring Variational Latent Reasoning for Multimodal Embeddings

VaME:探索用于多模态嵌入的变分潜在推理
Wu, Peixi, Jiang, Mingzhou, Ma, Feipeng, Yang, Biao, Zhou, Yunhao, Yuan, Wei, Chai, Bosong, Lin, Huizu, Chen, Jie, Hu, Zhangchi, Yang, Fan, Ou, Wenwu, Li, Hebei, Sun, Xiaoyan
Abstract
Universal multimodal retrieval requires compact embeddings that preserve task-relevant semantic information across diverse modalities. Prior works have incorporated latent reasoning into multimodal embedding learning to refine this information before embedding extraction. However, most existing approaches remain confined to deterministic latent paths, without exploring alternative trajectories to discover better embeddings. Thus, we propose VaME (Variational Multimodal Embeddings), a framework that models latent reasoning as a learnable distribution over trajectories. Specifically, we first introduce Variational Latent Reasoning (VLR) to enable autoregressive exploration in latent space, guided by answer reconstruction through a lightweight decoder. Meanwhile, we augment the original embedding-token readout with a latent-fused embedding to facilitate exploration during subsequent reinforcement learning. Finally, we optimize latent reasoning over stochastic variational trajectories through reinforcement learning, using Semantic Decoding Reward (SDR) to favor semantically meaningful trajectories with interpretable decoded outcomes. On the 78-task MMEB-V2 benchmark, spanning image, video, and visual-document retrieval, VaME outperforms most explicit CoT-based models and all latent-reasoning baselines. VaME also demonstrates robust performance on reasoning-intensive benchmarks such as MRMR, with substantial gains after reinforcement learning. Importantly, VaME achieves these gains with at least a 4.25x inference speedup over the deterministic latent autoregressive baselines. The code will be made publicly available.
Chinese Translation
通用多模态检索需要紧凑的嵌入,以在不同模态间保留任务相关的语义信息。先前的工作已将潜在推理引入多模态嵌入学习,以在嵌入提取之前精炼这些信息。然而,大多数现有方法仍局限于确定性的潜在路径,未能探索替代轨迹以发现更好的嵌入。因此,我们提出 VaME(Variational Multimodal Embeddings),一个将潜在推理建模为轨迹上可学习分布的框架。具体来说,我们首先引入变分潜在推理(VLR),以在潜在空间中实现自回归探索,并由轻量级解码器通过答案重建进行引导。同时,我们用潜在融合嵌入增强原始的嵌入 token 读出,以促进后续强化学习期间的探索。最后,我们通过强化学习在随机变分轨迹上优化潜在推理,使用语义解码奖励(SDR)来偏好具有可解释解码结果的语义有意义的轨迹。在涵盖图像、视频和视觉文档检索的 78 项任务 MMEB-V2 基准上,VaME 优于大多数显式 CoT 模型和所有潜在推理基线。VaME 在 MRMR 等推理密集型基准上也表现出稳健的性能,在强化学习后取得了显著提升。重要的是,VaME 在取得这些增益的同时,相比确定性潜在自回归基线至少实现了 4.25 倍的推理加速。代码将公开提供。
cs.CV / 209 / 2609.33412

Resolving State-Representation Mismatch: State-Space Visual Reasoning for Open-Loop VLA Planning

解决状态表示不匹配:面向开环VLA规划的状态空间视觉推理
Xiao, Junhao, Zhao, Haoxiang, Fang, Menghao, Zhang, Jinkui, Yu, Jinghan, Huang, Xinyu, Wu, Zhiyu, Xu, Kaiming, Chen, Yi, Bao, Youjun, Ma, Zhiyuan
Abstract
Despite rapid progress in vision-language-action (VLA) models, existing reasoning paradigms still face a fundamental \emph{state-representation mismatch} in open-loop planning. Given only an initial observation, models must internally simulate action-conditioned state transitions, whereas text-, pixel-, and latent-space reasoning can suffer from lossy spatial compression, error-accumulating visual generation, and bypass of intermediate latent tokens, respectively, undermining reliable long-horizon planning. We propose \textbf{State-Space Visual Reasoning} (SSVR), which decouples static visual context, language constraints, and a recurrent latent state. SSVR encodes the initial image and instruction once, then conditions each action prediction on the latent state and updates it with an action-conditioned GRU. Using Qwen2.5-VL as the backbone, SSVR achieves 99.5/99.6, 96.3/98.0, and 83.9/90.6 EM/PR on FrozenLake, Maze, and MiniBehavior, substantially outperforming prior methods. Extensive experiments support the effectiveness of recurrent state modeling for VLA open-loop planning across input transformations and transfer settings. By reusing static visual-textual context and updating a compact recurrent state, SSVR supports efficient multi-step inference, achieving up to $98.58\times$ faster Maze decoding rollouts than the evaluated baselines with the prefix cache prebuilt.
Chinese Translation
尽管视觉-语言-动作(VLA)模型取得了快速进展,现有推理范式在开环规划中仍面临根本性的状态表示不匹配。在仅给定初始观测的情况下,模型必须在内部模拟以动作为条件的状态转移,而文本空间、像素空间和潜在空间推理分别可能受制于有损空间压缩、误差累积的视觉生成以及中间潜在token被绕过等问题,从而削弱可靠的长时程规划。我们提出状态空间视觉推理(SSVR),它将静态视觉上下文、语言约束和循环潜在状态解耦。SSVR对初始图像和指令仅编码一次,然后使每个动作预测以潜在状态为条件,并用动作条件GRU更新该状态。以Qwen2.5-VL为主干,SSVR在FrozenLake、Maze和MiniBehavior上分别取得99.5/99.6、96.3/98.0和83.9/90.6的EM/PR,显著优于先前方法。大量实验支持了循环状态建模在跨输入变换和迁移设置下用于VLA开环规划的有效性。通过复用静态视觉-文本上下文并更新紧凑的循环状态,SSVR支持高效的多步推理;在预先构建前缀缓存的情况下,其在Maze解码rollout上相比所评估的基线最高可加速98.58×。
cs.CV / 210 / 2609.33414

TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models

TTRSD:结合自蒸馏的测试时强化学习用于视觉-语言模型
Wang, Shuning, Wu, Zhiheng, Zhou, Xun, Cui, Chongyang, Jia, Chen, Liu, Bowen, Li, Chuanjie, Chen, Xiang, Yang, Yi, Zhang, Yumeng, Huang, Wenjie
Abstract
Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reasoning risking the degradation of pre-trained reasoning capabilities. We propose TTRSD, a test-time reinforcement learning framework combining multi-view answer-level self-distillation with visual contrastive token selection. A shared policy aggregates teacher predictions across original, cropped, and downsampled views into an answer distribution. Student trajectories generated from the original image receive rewards based on the support for their final answers in this distribution. To allocate this feedback precisely toward perceptual bottlenecks, we compare the log-probabilities of the same sampled tokens under original and visually ablated inputs while holding their textual prefixes fixed, selecting visually sensitive positions for policy-gradient updates. TTRSD separates update direction, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher. With only 20 unlabeled adaptation samples, TTRSD improves performance across seven benchmarks and three VLMs, raising InternVL3-2B's MMMU accuracy from 35.79% to 49.32%(+13.53%), demonstrating cross-dataset generalization while preserving inherent reasoning integrity.
Chinese Translation
测试时强化学习使视觉-语言模型(VLMs)能够利用无标签输入进行自适应。然而,在固定视觉条件下重复采样可能会强化共享的感知错误,而序列级奖励未能分离出视觉感知这一锚定多模态推理的基础瓶颈,从而可能导致预训练推理能力退化。我们提出TTRSD,一个结合多视图答案级自蒸馏与视觉对比标记选择的测试时强化学习框架。一个共享策略将原始、裁剪和下采样视图上的教师预测聚合为答案分布。从原始图像生成的学生轨迹根据其最终答案在该分布中的支持度获得奖励。为了将此反馈精确地分配到感知瓶颈上,我们在保持文本前缀固定的情况下,比较相同采样标记在原始输入和视觉消融输入下的对数概率,选择视觉敏感位置进行策略梯度更新。TTRSD将更新方向(由组相对优势决定)与更新位置(由视觉敏感性决定)分离,无需真实标签、外部验证器或单独的教师模型。仅用20个无标签自适应样本,TTRSD在七个基准和三个VLM上提升了性能,将InternVL3-2B的MMMU准确率从35.79%提高到49.32%(+13.53%),展示了跨数据集泛化能力,同时保持了固有的推理完整性。
cs.CV / 211 / 2609.33419

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

TT-VidT:解耦时间轴以实现高效以运动为中心的视频预训练
Yeh, Shih-Ying, Kaplan, Daniel Z., Wang, Xuehai, Yang, Fu-En, Chen, Min-Hung, Lai, Shang-Hong
Abstract
Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched $4 \times 6 = 24$ architecture-objective study at roughly 170M ~ 190M encoder scale on $\sim$1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54% ~ 121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. HMDB51, IARD, and EPIC-Kitchens bound the claim.
Chinese Translation
在视频自监督学习中,比较往往评估完整的训练方案,而不是单独评估方法本身:架构、目标、数据暴露、调度、规模和解码器容量都可能同时变化。这使得很难确定哪些选择会产生以运动为优先的表示,其增益集中在帧间变化上,同时保留有用的外观信息。我们通过一项匹配的 $4 \times 6 = 24$ 架构-目标研究来解决这个问题,该研究在约1.7M OpenVid和Moments-in-Time v2视频片段上进行,编码器规模约为170M ~ 190M,训练8个epoch,并提出了TT-VidT。TT-VidT将DINOv3初始化的ViT-B/16逐帧空间路径与紧凑的时间传输层相结合,通过Diff Compression训练,从第一帧外观锚点和帧特定的运动令牌重建目标帧。扫描表明,TT3D与Diff Compression结合,而非单独使用任一组件,进入了最强的运动敏感机制,解码器消融实验倾向于使用紧凑的视频预训练解码器。在最终比较中,TT-VidT同时在Jester、Something-Something V2、ARID和Diving48微调任务上领先,相比最强的非TT行提高了54% ~ 121%,同时使用的编码器FLOPs比DisMo少48%,比VideoMAE或V-JEPA2少55%。HMDB51、IARD和EPIC-Kitchens限定了这一结论。
cs.CV / 212 / 2609.33432

TC-ADA: One-Shot Active Domain Adaptation for Semantic Segmentation

TC-ADA:面向语义分割的单轮主动域适应
Yan, Weihao, Qian, Yeqiang, Li, Yueyuan, Li, Tao, Wang, Chunxiang, Yang, Ming
Abstract
Manual dense annotation remains a major obstacle to deploying semantic segmentation models in new driving environments. Active domain adaptation (ADA) seeks label-efficient transfer by annotating only a selected portion of the target domain. Existing ADA methods commonly implement this process through multiple rounds of acquisition, annotation, and retraining. We study a practical one-shot image-level setting that selects and densely annotates a fixed target subset in a single round, followed by uninterrupted adaptation. Within this setting, we develop Target-Calibrated Active Domain Adaptation (TC-ADA) as a joint design of complete-image acquisition and target-calibrated adaptation. Stage~1 uses visual representations from a vision foundation model (VFM) together with semantic predictions from a fixed unsupervised domain adaptation model to select representative and informative target images without target annotations. Stage~2 jointly uses labeled source data, labeled target data, and the remaining unlabeled target data, while calibrating source and target supervision under limited target labels. Extensive experiments across five synthetic-to-real and real-to-real driving transfers show consistent improvements over representative ADA baselines. With only 23 to 46 labeled target images on four transfers and 140 on Mapillary, TC-ADA stays within 1.9 mean intersection over union (mIoU) points of target-only full supervision. Code will be available at https://github.com/ywher/TC-ADA.
Chinese Translation
手动密集标注仍然是将语义分割模型部署到新驾驶环境中的主要障碍。主动域适应(ADA)旨在通过仅标注目标域中选定的一部分来实现标签高效的迁移。现有的ADA方法通常通过多轮获取、标注和重新训练来实现这一过程。我们研究了一种实用的单轮图像级设置,该设置在一轮中选取并密集标注固定的目标子集,然后进行不间断适应。在这种设置下,我们提出了目标校准主动域适应(TC-ADA),作为完整图像获取和目标校准适应的联合设计。第一阶段使用来自视觉基础模型(VFM)的视觉表示以及来自固定的无监督域适应模型的语义预测,在没有目标标注的情况下选择具有代表性和信息量的目标图像。第二阶段联合使用有标注的源数据、有标注的目标数据以及剩余的未标注目标数据,同时在有限的目标标签下校准源监督和目标监督。在五个合成到真实和真实到真实的驾驶迁移上的大量实验表明,相较于代表性的ADA基线,该方法取得了持续改进。在四个迁移任务中仅使用23到46张有标注的目标图像,在Mapillary上使用140张,TC-ADA与仅目标域全监督的平均交并比(mIoU)差距保持在1.9个点以内。代码将在 https://github.com/ywher/TC-ADA 提供。
cs.CV / 213 / 2609.33445

Concept Score Relearning: A Unified Cross-Architecture Attack on Concept Erasure

概念分数重学习:一种针对概念擦除的统一跨架构攻击
Tae, Hong Xi, Zhang, Jiaming, He, Wenwen, Wang, Xuan, Lim, Wei Yang Bryan
Abstract
Concept erasure aims to suppress undesirable knowledge in text-to-image generative models. However, existing robustness evaluations typically rely on relearning attacks tailored to specific model architectures. We study concept reactivation across two substantially different generative paradigms: noise-prediction U-Nets and flow-matching Transformers. We introduce \textbf{Concept Score Relearning (CSR)}, a unified parameter-level framework that reactivates erased concepts by optimizing each model within its native prediction space. CSR requires no external target-concept image dataset and applies the same concept-directed objective to both U-Net-based Stable Diffusion and Transformer-based FLUX. Experiments across diverse concepts and multiple erasure methods demonstrate consistent concept reactivation across both architectures, highlighting the cross-architecture applicability of CSR and the persistent recoverability of apparently erased concepts. For strict nudity, CSR reaches average ASRs of 50.47\% on FLUX and 40.29\% on Stable Diffusion, consistently ranking first across all evaluated safety settings.
Chinese Translation
概念擦除旨在抑制文本到图像生成模型中的不良知识。然而,现有的鲁棒性评估通常依赖于针对特定模型架构定制的重学习攻击。我们研究了两种显著不同的生成范式中的概念重激活:噪声预测U-Net和流匹配Transformer。我们引入了概念分数重学习(CSR),一个统一的参数级框架,通过在每个模型的原始预测空间内优化模型来重激活被擦除的概念。CSR无需外部目标概念图像数据集,并对基于U-Net的Stable Diffusion和基于Transformer的FLUX应用相同的概念导向目标。跨多种概念和多种擦除方法的实验表明,两种架构均能一致地重激活概念,突显了CSR的跨架构适用性以及看似被擦除概念的持续可恢复性。对于严格裸体内容,CSR在FLUX上达到平均攻击成功率(ASR)为50.47%,在Stable Diffusion上为40.29%,在所有评估的安全设置中始终排名第一。
cs.CV / 214 / 2609.33449

A Visual Classification Dataset and Model Evaluation for Historical Manuscript Illustrations

历史手稿插图的视觉分类数据集与模型评估
Evron, Yoav, Siegal, Michal Bar-Asher, Fire, Michael
Abstract
Historical manuscript illustrations preserve rich visual evidence of past cultures. They depict people, animals, plants, diagrams, music notations, and decorative forms. Although large digitization projects have made many manuscripts available online, the material itself remains difficult to explore at scale. Extraction systems can find illustrations on manuscript pages, but without meaningful categories, large collections remain hard to search and explore. We address this gap by introducing a manually labeled dataset of 15,000 illustrations from manuscripts dating back hundreds of years across 22 categories, and evaluating modern vision models for image classification on this task. The problem is challenging due to stylistic diversity, degradation, and semantic ambiguity, with many images that fit more than one category. We compare fine-tuned CNN and Transformer-based classifiers, zero-shot CLIP, embedding-based classifiers, and direct vision-language models. Results show that fine-tuned image classifiers perform best overall, with ConvNeXt reaching 88.9% accuracy and 81.3% macro-F1. Using CLIP embeddings with XGBoost provides a strong alternative. In contrast, zero-shot CLIP and direct vision-language classification perform substantially worse, highlighting the limits of general-purpose models in this domain. Beyond overall performance, the analysis reveals which categories are visually separable and where errors reflect genuine semantic overlap, suggesting that some limitations arise from the taxonomy itself.
Chinese Translation
历史手稿插图保存了过去文化的丰富视觉证据。它们描绘了人物、动物、植物、图表、音乐记谱和装饰形式。尽管大型数字化项目使许多手稿可以在线获取,但材料本身仍然难以大规模探索。提取系统可以在手稿页面上找到插图,但如果没有有意义的类别,大型收藏仍然难以搜索和探索。我们通过引入一个手动标注的数据集来填补这一空白,该数据集包含来自可追溯数百年的手稿的15,000幅插图,涵盖22个类别,并评估现代视觉模型在此任务上的图像分类性能。由于风格多样性、退化和语义模糊性,这个问题具有挑战性,许多图像适合多个类别。我们比较了微调的CNN和基于Transformer的分类器、零样本CLIP、基于嵌入的分类器以及直接视觉语言模型。结果表明,微调图像分类器总体表现最佳,ConvNeXt达到88.9%的准确率和81.3%的宏F1。使用CLIP嵌入与XGBoost提供了强大的替代方案。相比之下,零样本CLIP和直接视觉语言分类表现明显较差,突显了通用模型在该领域的局限性。除了整体性能外,分析还揭示了哪些类别在视觉上可分离,以及错误反映真实语义重叠的地方,表明一些限制源于分类法本身。
cs.CV / 215 / 2609.33450

Native Association: Confidence-Aware Human Perception in the Wild with a Foundation VLM

原生关联:使用基础VLM在野外进行置信度感知的人体感知
Dmitriev, Igal, Liba, Ofir
Abstract
Extracting who is where, on which team, wearing which number from a broadcast frame is typically done by stitching a detector, an OCR engine, and classifiers together -- and the stitching step swaps identities under occlusion. We make association native instead: a 0.77B vision-language model (Florence-2) is fine-tuned to emit all per-person attributes as one grammar-constrained sequence, with each attribute generated inside its owner's block. Output is therefore schema-valid on every frame by construction, and no post-hoc binding step exists to attach a correctly read number to the wrong player: residual misassociation is pure perception error, $\approx4\times$ rarer than zero-shot-prompted frontier APIs' (0.057 vs. 0.21-0.24). On a frozen multi-sport test set, this single pass reaches 0.95 detection F1 (APIs: 0.65-0.75). A single extra forward pass yields a per-field confidence that supports a reject option (jersey precision $0.71\rightarrow0.96$ at half coverage) and routes a training-free zoom-and-re-read for small players. Surprisingly, once the grammar is learned, further parameter-efficient tuning yields no measurable gain under the adaptation configurations we test; the identical recipe on WIDER-Attribute reaches 93.1 mAP given-box, yields the first detection-coupled end-to-end results under its standard test protocol (84.5 mAP), and reproduces the same tuning result. In this regime, the gains live in the structure, not in added weights.
Chinese Translation
从广播画面中提取谁在哪里、属于哪支队伍、穿着几号,通常是通过将检测器、OCR引擎和分类器拼接在一起来完成——而在遮挡情况下,这种拼接步骤会交换身份。我们转而将关联原生化:微调一个0.77B的视觉语言模型(Florence-2),使其将所有单人属性作为一个语法约束的序列输出,每个属性都在其所属者的块内生成。因此,输出在构造上对每一帧都符合模式,并且不存在事后绑定步骤来将正确读取的号码附加到错误的球员上:残余的错误关联是纯粹的感知错误,比零样本提示的前沿API(0.057 vs. 0.21-0.24)罕见约4倍。在一个冻结的多运动测试集上,这一单次前向传播达到0.95的检测F1(API:0.65-0.75)。一次额外的前向传播产生每个字段的置信度,支持拒绝选项(在半覆盖下球衣精度0.71→0.96),并为小尺寸球员路由一个无需训练的放大再读取。令人惊讶的是,一旦语法被学习,在我们测试的自适应配置下,进一步的参数高效微调没有产生可测量的增益;相同的配方在WIDER-Attribute上达到93.1 mAP给定框,在其标准测试协议下产生首个检测耦合的端到端结果(84.5 mAP),并重现相同的调优结果。在这种模式下,增益存在于结构中,而不是增加的权重中。
cs.CV / 216 / 2609.33462

SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera

SphMind:迈向稳健、免训练的基于VLM的360相机空间推理
Damodaran, Shriram, Debnath, Soumyaratna, Tan, Cheston, Wang, Lin
Abstract
Omnidirectional or 360 cameras provide embodied AI agents with a holistic, wide field-of-view (FoV) view of their surroundings, motivating the use of Multi-modal Large Language Models (MLLMs) for omnidirectional spatial reasoning. However, most MLLMs are trained on conventional 2D perspective images and struggle with the severe distortions and wrap-around discontinuities induced by spherical geometry. Enabling them to generalize to non-Euclidean 3D spaces without retraining therefore remains challenging. We propose SphMind, a training-free, plug-and-play framework that decouples semantic perception from geometric reasoning. Rather than requiring MLLMs to learn spherical geometry internally, SphMind preserves their semantic capabilities while handling geometry externally. We introduce a Spherical Harmonics-based Spatial Graph (SHSG) that models spatial relationships through equivariant transformations on the sphere, together with Inference-Time Geometric Grounding (IGG), a model-agnostic closed-loop optimization process that aligns MLLM representations with spherical geometric constraints during inference. Experiments on three benchmarks show that SphMind achieves over 21.4% average improvement in directional reasoning on MP3D and Stanford2D-3D, outperforms prompt-engineering baselines by 8.7% on the real-world ODI-Bench, and improves rotational invariance by 5.9% under panorama rotations, without additional training or dataset-specific tuning. In-the-wild evaluations further show that SphMind resolves directional reasoning queries that baseline vision-language models fail to answer correctly.
Chinese Translation
全向或360相机为具身智能体提供了对其周围环境的整体、宽视场(FoV)视图,推动了使用多模态大语言模型(MLLMs)进行全向空间推理。然而,大多数MLLMs是在传统的2D透视图像上训练的,难以处理由球面几何引起的严重畸变和环绕不连续性。因此,使它们无需重新训练就能泛化到非欧几里得3D空间仍然具有挑战性。我们提出了SphMind,一个免训练、即插即用的框架,将语义感知与几何推理解耦。SphMind不需要MLLMs在内部学习球面几何,而是保留其语义能力,同时在外部处理几何。我们引入了一种基于球谐函数的空间图(SHSG),通过球面上的等变变换来建模空间关系,以及推理时几何接地(IGG),这是一种与模型无关的闭环优化过程,在推理过程中将MLLM表示与球面几何约束对齐。在三个基准测试上的实验表明,SphMind在MP3D和Stanford2D-3D上的方向推理平均提高了21.4%以上,在真实世界的ODI-Bench上比提示工程基线高出8.7%,在全景旋转下将旋转不变性提高了5.9%,且无需额外训练或数据集特定调优。野外评估进一步表明,SphMind能够解决基线视觉语言模型无法正确回答的方向推理查询。
cs.CV / 217 / 2609.33513

Printability-Constrained Adversarial Decals for Near-Nadir Aerial Perception: Measured Ink Gamuts, Nested Realism Constraints, and a Physical-World Bound

可打印性约束的对抗性贴花用于近天底空中感知:测量的墨水色域、嵌套现实性约束与物理世界界限
Shrestha, Sandesh, Mahima, K. T. Yasas, Perera, Asanka G.
Abstract
Adversarial patches for aerial perception are typically evaluated as digital composites, with printing left as an implementation detail. This study imposes three physical constraints during optimization rather than after it: the color range a particular printer can reproduce, the size of the flat panel a vehicle offers, and the loss of fine detail incurred when the patch is imaged from altitude. The principal comparison isolates the ink set. Two patches share all seventeen recorded optimization settings and differ only in the colors available to them. One is constrained to a uniform color cube; the other to a gamut measured by printing and scanning a 216-patch chart. Each was optimized at three seeds and evaluated against thirteen victim conditions, with every rate reported against a size-matched optimized control. The effect of the measured gamut is victim-dependent rather than uniform. Net attack success rises on three of six closed-set segmentation victims, and for these the seed ranges of the two ink sets are disjoint: $+0.120$ on DeepLabv3-R101 and $+0.041$ on SegFormer-B0. The color-cube patch is consistently stronger on the open-vocabulary segmenter and on two of four detectors, though no detector exceeds a net of $+0.026$ under either ink set. The natural explanation is that a printable palette is simply less chromatic and lower in frequency than a digital one. Eleven further patches test this account and it does not hold. Once cardinality is matched, a palette as chromatic as the cube attacks equally well. Cardinality itself shows no trend from three inks to thirty-two. Palettes matched on cardinality, lightness and chroma, and differing only in hue placement, span $0.035$ to $0.136$. A physical evaluation with printed decals did not detect transfer; it bounds the transferred rate at $0.133$, which does not exclude the simulated value of $0.121$.
Chinese Translation
用于空中感知的对抗性补丁通常作为数字合成物进行评估,而打印则留作实现细节。本研究在优化过程中而非优化之后施加三个物理约束:特定打印机可再现的颜色范围、车辆提供的平板尺寸,以及从高空成像补丁时产生的细粒度细节损失。主要比较隔离了墨水集合。两个补丁共享所有十七个记录的优化设置,仅在其可用的颜色上不同。一个被约束到均匀颜色立方体;另一个被约束到通过打印和扫描216个色块图测量的色域。每个都在三个随机种子下优化,并针对十三个受害者条件进行评估,每个比率都报告为与尺寸匹配的优化控制相比。测量色域的效果是依赖于受害者的,而非一致的。净攻击成功率在六个闭集分割受害者中的三个上上升,对于这些,两种墨水集合的种子范围是不相交的:在DeepLabv3-R101上为$+0.120$,在SegFormer-B0上为$+0.041$。颜色立方体补丁在开放词汇分割器和四个检测器中的两个上始终更强,尽管在任一墨水集合下没有检测器超过净$+0.026$。自然的解释是,可打印调色板比数字调色板色彩更少且频率更低。另外十一个补丁测试了这一解释,但它并不成立。一旦基数匹配,与立方体一样色彩丰富的调色板攻击效果同样好。基数本身从三种墨水到三十二种没有显示出趋势。在基数、亮度和色度上匹配,仅色相位置不同的调色板,跨度从$0.035$到$0.136$。使用印刷贴花的物理评估未检测到迁移;它将迁移率限制在$0.133$,这并不排除模拟值$0.121$。
cs.CV / 218 / 2609.33518

SceneScaffold: Active Scene-State Construction for Unified 3D Scene Understanding

SceneScaffold:面向统一3D场景理解的主动场景状态构建
Li, Xiangqi, Huang, Libo, Zhao, Jiarui, Feng, Weilun, Yang, Chuanguang, An, Zhulin, Xu, Yongjun
Abstract
Recent 3D large multimodal models (3D-LMMs) rely on a visual bottleneck to compress complex 3D scene evidence into a limited number of visual tokens compatible with large language models (LLMs). Current visual bottlenecks, however, often passively compress heterogeneous 3D evidence into a homogeneous object-centric token sequence, leaving the spatial organization of the scene under-represented. This under-representation forces the LLM to recover spatial relations from a flattened token sequence, leading to unstable reasoning in relation-intensive and spatially ambiguous scenes. To address this issue, we propose SceneScaffold, an active scene-state construction framework for unified 3D scene understanding. SceneScaffold reformulates the visual bottleneck from a passive feature compressor into an active scene organizer, constructing a role-aware spatial scaffold before language reasoning. Specifically, SceneScaffold organizes superpoint-level visual evidence into scene-state components with distinct structural roles: entity states preserve core object semantics, scene-frame states maintain spatial references via boundary and region anchors, relation states encode object-environment interaction cues, and a global summary provides compact context. Through this role-aware construction, SceneScaffold provides the LLM with a spatially organized scene representation before language reasoning. Experiments on unified 3D scene understanding tasks, including 3D visual grounding, question answering, and dense captioning, demonstrate the effectiveness of SceneScaffold, while diagnostic results further show its applicability to relation-intensive and spatially ambiguous cases. Code is available at https://github.com/lixiangqi707/SceneScaffold.
Chinese Translation
近期,3D大型多模态模型(3D-LMMs)依赖视觉瓶颈(visual bottleneck)将复杂的3D场景证据压缩为有限数量的视觉标记(visual tokens),以兼容大型语言模型(LLMs)。然而,当前的视觉瓶颈通常被动地将异构的3D证据压缩为同质的、以对象为中心的标记序列,导致场景的空间组织表征不足。这种表征不足迫使LLM从扁平化的标记序列中恢复空间关系,从而在关系密集和空间模糊的场景中导致不稳定的推理。为解决这一问题,我们提出了SceneScaffold,一种面向统一3D场景理解的主动场景状态构建框架。SceneScaffold将视觉瓶颈从被动的特征压缩器重新定义为主动的场景组织器,在语言推理之前构建角色感知的空间支架(role-aware spatial scaffold)。具体而言,SceneScaffold将超点级(superpoint-level)视觉证据组织为具有不同结构角色的场景状态组件:实体状态(entity states)保留核心对象语义,场景框架状态(scene-frame states)通过边界和区域锚点维持空间参考,关系状态(relation states)编码对象-环境交互线索,而全局摘要(global summary)提供紧凑的上下文。通过这种角色感知的构建,SceneScaffold在语言推理之前为LLM提供了空间组织化的场景表示。在统一3D场景理解任务(包括3D视觉定位、问答和密集描述)上的实验证明了SceneScaffold的有效性,诊断结果进一步表明其适用于关系密集和空间模糊的情况。代码可在https://github.com/lixiangqi707/SceneScaffold获取。
cs.CV / 219 / 2609.33520

Anatomy-Structured Hierarchical MIL for Weakly-Supervised Thoracic Disease Detection in Chest X-rays

用于胸片弱监督胸部疾病检测的解剖结构分层多示例学习
Kim, Jeongin, Ahn, Sohyun, Kang, Seo Young, Sung, Jaeyi, Kim, Soomin, Cho, Sungho, Lee, Rena, Kim, Kwanchang, Noh, Junhyug
Abstract
Weakly-supervised thoracic disease detection in chest X-rays (CXR) is challenging due to subtle appearances and complex anatomical overlap, motivating anatomy-aware modeling for improved localization. However, prior anatomy-aware methods typically rely on coarse region proxies or static spatial priors, which may restrict dynamic instance discovery and limit precise localization of small abnormalities. We propose Anatomy-Structured Hierarchical Multiple Instance Learning (ASH-MIL), a framework that introduces parallel anatomy-structured observation branches (cardiac, pulmonary, and agnostic) combined with hierarchical MIL aggregation. Anatomical priors are injected as soft spatial biases into decoder cross-attention, enabling anatomically grounded evidence maps without disease bounding-box supervision. Instance localization is derived directly from MIL-weighted cross-attention maps without bounding box supervision. Experiments on CXR8 and cross-domain MIMIC-CXR held-out sets demonstrate consistent improvements over prior weakly-supervised and anatomy-aware approaches, particularly under stricter localization criteria. Our code is available at https://github.com/jn-kim/ash-mil.
Chinese Translation
胸片(CXR)中的弱监督胸部疾病检测因表现细微且解剖结构复杂重叠而具有挑战性,这推动了解剖感知建模以改善定位。然而,既往的解剖感知方法通常依赖粗粒度区域代理或静态空间先验,这可能限制动态实例发现并制约小异常的精准定位。我们提出解剖结构分层多示例学习(Anatomy-Structured Hierarchical Multiple Instance Learning, ASH-MIL)框架,该框架引入并行的解剖结构观察分支(心脏、肺部和解剖不可知(agnostic)分支),并结合分层MIL聚合。解剖先验以软空间偏置形式注入解码器交叉注意力,从而无需疾病边界框监督即可生成具有解剖依据的证据图。实例定位直接由MIL加权的交叉注意力图导出,无需边界框监督。在CXR8和跨域MIMIC-CXR留出集上的实验表明,相较于既往弱监督和解剖感知方法,本方法取得一致提升,尤其在更严格的定位标准下。我们的代码见 https://github.com/jn-kim/ash-mil。
cs.CV / 220 / 2609.33523

In-Token Learning for High-Fidelity Image Restoration via Diffusion Transformers

基于扩散Transformer的面向高保真图像修复的Token内学习
Yi, Xingfu, Yu, Xiaoxue
Abstract
We present In-Token Learning, an image restoration framework that adapts a pretrained diffusion transformer using conditional rectified flow matching. Clean targets paired with degraded inputs supervise transport from Gaussian noise to restored images. Spatially aligned degraded-image tokens are fused with evolving latent tokens along the channel dimension, preserving the image-token count at a given resolution. Direct Low-Quality Guidance (DLG) combines frozen degraded-image embeddings with a fixed task prompt through the native conditioning pathway, without a trainable ControlNet-style branch or image captioning. We evaluate super-resolution and denoising on DIV2K, LSDIR, FFHQ, RealLQ250, and RealPhoto60, and automatic colorization on DIV2K and LSDIR. The tasks use separately trained checkpoints under the same framework. Results show competitive fidelity and perceptual quality under the evaluated protocols, with weaker generalization on RealLQ250. We report full-image QHD ($2560{\times}1440$) inference and a tiled $12$K restoration demonstration of Along the River During the Qingming Festival. Attention cost still increases with resolution. This technical report preserves the early broader study underlying Fill2SR, which subsequently developed the real-world super-resolution direction.
Chinese Translation
我们提出了In-Token Learning,一种图像修复框架,它使用条件整流流匹配来适配预训练的扩散Transformer。干净目标与退化输入配对,监督从高斯噪声到恢复图像的传输。空间对齐的退化图像token与演化的潜在token沿通道维度融合,在给定分辨率下保持图像token数量不变。直接低质量引导(DLG)通过原生条件路径将冻结的退化图像嵌入与固定任务提示相结合,无需可训练的ControlNet风格分支或图像描述。我们在DIV2K、LSDIR、FFHQ、RealLQ250和RealPhoto60上评估超分辨率和去噪,并在DIV2K和LSDIR上评估自动上色。这些任务在同一框架下使用单独训练的检查点。结果表明,在评估协议下,保真度和感知质量具有竞争力,但在RealLQ250上的泛化能力较弱。我们报告了全图像QHD($2560{\times}1440$)推理以及对《清明上河图》的分块$12$K修复演示。注意力成本仍随分辨率增加而增加。本技术报告保留了Fill2SR背后的早期更广泛研究,该研究随后发展了真实世界超分辨率方向。
cs.CV / 221 / 2609.33558

PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization

PGL-3D:面向3D视觉查询定位的渐进式几何学习
Peng, Liang, Mu, Shizhuo, Tan, Bohan, Wang, Wenyuan, Zhao, Chen, Dong, Xingping, Fan, Heng, Zhang, Libo, Du, Bo
Abstract
3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB--point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotated pose may differ from how the object appears in the search frames. The benchmark baseline predicts cuboids after feature modeling, leaving their geometry unused for subsequent feature refinement. We investigate whether complete intermediate cuboids can improve query and proposal representations before final decoding. We introduce Progressive Geometric Learning for 3DVQL (PGL-3D), a predict--select--refine--re-predict framework that uses intermediate cuboids to guide the aggregation of search evidence and update query and proposal representations. A shared head first predicts a complete cuboid for every proposal. Query--Tube--Memory (QTM) then selects reference observations by combining proposal association, cuboid quality, frame response, and target absence, since association confidence alone establishes neither target presence nor geometric accuracy. The center, size, and orientation of each selected cuboid define soft pooling weights over query-conditioned proposal features. The pooled memory updates the query and proposal representations, and the head re-predicts from the updated features. A training-only objective, ST-D9O, supervises cuboid geometry at every stage by adding boundary, signed-distance, and soft-overlap terms to parameter regression. PGL-3D achieves a mean stAP of $0.270 \pm 0.004$ on 3DVQL, compared with $0.044$ reported for LaF. Ablations support the benefits of geometry-guided feature updates, while stage-wise analyses show improved cuboid accuracy. Replacing the geometry objective in our PROT3D reproduction with ST-D9O improves mAO on GSOT3D from $21.63\%$ to $25.78\%$. Our code and models will be released.
Chinese Translation
3D视觉查询定位(3DVQL)在RGB-点云序列中检索查询对象的最新连续出现,并为每个响应帧预测一个9自由度立方体。查询的捕获独立于搜索序列,因此其标注姿态可能与物体在搜索帧中的外观不同。基准基线在特征建模后预测立方体,使得其几何信息未用于后续的特征精炼。我们研究完整的中间立方体是否能在最终解码之前改善查询和提案表示。我们提出了用于3DVQL的渐进几何学习(PGL-3D),一个预测-选择-精炼-再预测框架,利用中间立方体指导搜索证据的聚合,并更新查询和提案表示。共享头首先为每个提案预测一个完整的立方体。然后,查询-管-记忆(QTM)通过结合提案关联、立方体质量、帧响应和目标缺失来选择参考观测,因为仅凭关联置信度既不能确定目标存在,也不能确定几何精度。每个选定立方体的中心、大小和方向定义了查询条件提案特征上的软池化权重。池化记忆更新查询和提案表示,然后共享头从更新后的特征重新预测。一个仅训练目标ST-D9O通过在参数回归中添加边界、符号距离和软重叠项,在每个阶段监督立方体几何。PGL-3D在3DVQL上实现了0.270 ± 0.004的平均stAP,而LaF报告的为0.044。消融实验支持几何引导特征更新的优势,而分阶段分析显示立方体精度有所提高。在我们PROT3D复现中,用ST-D9O替换几何目标,将GSOT3D上的mAO从21.63%提高到25.78%。我们的代码和模型将公开。
cs.CV / 222 / 2609.33581

ForeFly: A Dual-Horizon World Action Model for Aerial Vision-Language Navigation

ForeFly:一种用于空中视觉语言导航的双时域世界动作模型
Wang, Kunhui, Zhang, Xintong, Gao, Junyu, Xu, Changsheng
Abstract
Aerial Vision-Language Navigation (AVLN) requires UAVs to maintain reliable instruction following over long trajectories in complex 3D environments. However, existing AVLN approaches are predominantly reactive or limited to single-horizon prediction, overlooking complementary future cues across different temporal horizons. To address this limitation, we propose ForeFly, a dual-horizon latent world action model that predicts both a proximal future for local continuity and an adaptive route-critical future for long-range guidance. Horizon-specific foresight queries are primed with recent and route-critical visual memories, providing history-aware context for future prediction. To exploit their distinct roles in action generation, we introduce Foresight-Guided Action Refinement (FGAR), which asymmetrically exploits proximal foresight for local action enhancement and route-critical foresight for feature-wise correction and route-level guidance. Experiments on the TravelUAV and UAV-ON benchmarks show that ForeFly consistently outperforms strong baselines across seen and unseen settings, validating the effectiveness of dual-horizon foresight and FGAR learning. The code is available at: https://github.com/kunhuiW/ForeFly
Chinese Translation
空中视觉语言导航(AVLN)要求无人机在复杂三维环境中沿长轨迹保持可靠的指令遵循。然而,现有的AVLN方法主要是反应式的,或局限于单时域预测,忽略了不同时间范围上互补的未来线索。为了解决这一局限,我们提出了ForeFly,一种双时域潜在世界动作模型,它同时预测用于局部连续性的近端未来和用于长距离指导的自适应路线关键未来。特定时域的预见查询由近期和路线关键的视觉记忆初始化,为未来预测提供历史感知的上下文。为了利用它们在动作生成中的不同作用,我们引入了预见引导的动作细化(FGAR),它非对称地利用近端预见进行局部动作增强,并利用路线关键预见进行特征级校正和路线级指导。在TravelUAV和UAV-ON基准上的实验表明,ForeFly在已见和未见设置下均持续优于强基线,验证了双时域预见和FGAR学习的有效性。代码可在以下网址获取:https://github.com/kunhuiW/ForeFly
cs.CV / 223 / 2609.33582

Fill2SR: Repurposing Inpainting Diffusion Transformers for Real-World Super-Resolution

Fill2SR:将修复扩散Transformer重新用于真实世界超分辨率
Yi, Xingfu, Yu, Xiaoxue
Abstract
Recent real-world image super-resolution (SR) methods often adapt text-to-image (T2I) backbones with ControlNet-style branches or spatial conditioning tokens, which increases memory and computes with resolution and often constrains training to a fixed scale. We propose Fill2SR, which repurposes a masked-inpainting Diffusion Transformer for SR without extra spatial branches. Our Inpainting-Interface Evidence Adapter (IIEA) writes the low-quality (LQ) observation into the native masked-image slot under a full-image mask, turning inpainting into a reverse-degradation conditional rectified flow trained with LoRA-only tuning. We further introduce RCDT, an offline pipeline that distills degradation descriptors from unpaired real images and transfers them onto clean targets using frozen open-source models. Fill2SR supports mixed-resolution training up to QHD and yields stable performance across $512/1024/2048$ outputs. On synthetic benchmarks, our base model with IIEA achieves the best LPIPS on DIV2K and LSDIR; adding RCDT trades a small LPIPS drop for consistently stronger no-reference quality on RealLQ250 and RealPhoto60. Fill2SR remains memory-predictable, running $1536^2$ inference on a single 32GB GPU and extending to multi-megapixel outputs via tiled restoration.
Chinese Translation
近期真实世界图像超分辨率(SR)方法通常采用ControlNet风格分支或空间条件令牌来适配文本到图像(T2I)主干,这会随分辨率增加内存和计算量,并常常将训练限制在固定尺度。我们提出Fill2SR,它重新利用掩码修复扩散Transformer进行超分辨率,无需额外的空间分支。我们的修复接口证据适配器(IIEA)在全图像掩码下将低质量(LQ)观测写入原生掩码图像槽,将修复转化为反向退化条件校正流,仅使用LoRA调优进行训练。我们进一步引入RCDT,一种离线流水线,从非配对真实图像中蒸馏退化描述符,并使用冻结的开源模型将其转移到干净目标上。Fill2SR支持高达QHD的混合分辨率训练,并在$512/1024/2048$输出上产生稳定性能。在合成基准上,我们的带有IIEA的基础模型在DIV2K和LSDIR上实现了最佳LPIPS;添加RCDT以小幅LPIPS下降换取在RealLQ250和RealPhoto60上持续更强的无参考质量。Fill2SR保持内存可预测,在单个32GB GPU上运行$1536^2$推理,并通过分块恢复扩展到数百万像素输出。
cs.CV / 224 / 2609.33585

IVT-Guard: All-in-One Reasoning Model for AI-Generated Content Detection

IVT-Guard:用于AI生成内容检测的一体化推理模型
Niu, Hongwei, Luo, Yunpeng, Li, Hanjun, Zhou, Ziyin, Lin, Jianghang, Yan, Ke, Ding, Shouhong, Zhang, Shengchuan, Cao, Liujuan
Abstract
The rapid proliferation of highly realistic AI-Generated Content (AIGC) necessitates robust and interpretable detection mechanisms. However, existing detectors are predominantly confined to single modalities and provide binary outputs without reasoning. While Multimodal Large Language Models (MLLMs) present a promising solution, their development is constrained by the scarcity of multimodal reasoning data and the reasoning-detection optimization dilemma, where explicit reasoning supervision can compromise detection accuracy. To this end, we introduce IVT-Set, a comprehensive dataset comprising over 152K diverse image, video, and text samples equipped with multi-granularity Chain-of-Thought (CoT) reasoning trajectories. Based on it, we propose IVT-Guard, a pioneering framework for unified and interpretable AIGC detection across image, video, and text modalities. Furthermore, to overcome the aforementioned optimization dilemma, we design a novel three-stage training paradigm: Artifact-Aware Pre-training, Artifact-to-Evidence Supervised Fine-Tuning via artifact-aware injection, and Evidence-Verdict Consistency Group Relative Policy Optimization. Extensive experiments demonstrate that IVT-Guard achieves state-of-the-art detection performance across in-domain, out-of-domain, and cross-dataset settings while delivering faithful reasoning. Code and data will be released.
Chinese Translation
高度逼真的AI生成内容(AIGC)的迅速激增,迫切需要鲁棒且可解释的检测机制。然而,现有的检测器主要局限于单一模态,并提供无推理的二值输出。虽然多模态大语言模型(MLLMs)提供了一个有前景的解决方案,但其发展受到多模态推理数据稀缺以及推理-检测优化困境的制约,即显式的推理监督可能会损害检测准确性。为此,我们引入了IVT-Set,一个包含超过15.2万个多样化的图像、视频和文本样本的综合数据集,这些样本配备了多粒度的思维链(CoT)推理轨迹。基于此,我们提出了IVT-Guard,一个开创性的框架,用于跨图像、视频和文本模态的统一且可解释的AIGC检测。此外,为了克服上述优化困境,我们设计了一种新颖的三阶段训练范式:伪影感知预训练(Artifact-Aware Pre-training)、通过伪影感知注入的伪影到证据监督微调(Artifact-to-Evidence Supervised Fine-Tuning via artifact-aware injection),以及证据-裁决一致性组相对策略优化(Evidence-Verdict Consistency Group Relative Policy Optimization)。大量实验表明,IVT-Guard在域内、域外和跨数据集设置中均实现了最先进的检测性能,同时提供了忠实的推理。代码和数据将公开发布。
cs.CV / 225 / 2609.33593

LoopLUT: 3D Lookup Tables with Progressive Region Refinement for Real-Time 4K Image Enhancement

LoopLUT:用于实时4K图像增强的渐进区域细化3D查找表
Ye, Yang, Ma, Jiajun, Wu, Chen, Wang, Wei, Lu, Dianjie, Zhang, Guijuan, Fan, Linwei, Zheng, Zhuoran
Abstract
Color enhancement of 4K images must meet a quality target under a tight compute budget. Three-dimensional lookup tables (3D LUTs) dominate real-time enhancement because they decide at low resolution and apply a per-pixel lookup at full resolution. A single global LUT, however, is spatially invariant, so an underexposed shadow and a well-exposed region that share a pixel value receive identical corrections. Spatially heterogeneous demands cannot be expressed by such a mapping. We propose LoopLUT, a region-cascaded 3D LUT with progressive refinement. A global LUT performs the overall correction, followed by K-1 loop iterations. In each iteration a gating head predicts at low resolution the region that still needs correction, then builds a residual LUT from the color statistics of that region alone. The cascaded gates form a partition of unity, so the output is a per-pixel convex combination of the K lookup results. Fusion is therefore performed by the gates themselves, with no separate fusion module and no interpolation error accumulating across rounds. The decision stage runs at a fixed 256x256 resolution, independent of output resolution, so a 4K image costs only K pure lookups. Extensive experiments across four benchmarks show that LoopLUT improves PSNR by up to 2.81 dB over the strongest prior method, while keeping real-time throughput at 4K. The same decomposition also generalizes well to underwater enhancement datasets.
Chinese Translation
4K图像的颜色增强必须在严格的计算预算下达到质量目标。三维查找表(3D LUT)在实时增强中占主导地位,因为它们在低分辨率下进行决策,并在全分辨率下应用逐像素查找。然而,单个全局LUT具有空间不变性,因此共享相同像素值的曝光不足的阴影和曝光良好的区域会接受相同的校正。这种映射无法表达空间异构的需求。我们提出LoopLUT,一种具有渐进细化的区域级联3D LUT。全局LUT执行整体校正,随后进行K-1次循环迭代。在每次迭代中,一个门控头在低分辨率下预测仍然需要校正的区域,然后仅根据该区域的颜色统计构建残差LUT。级联门形成单位分解,因此输出是K个查找结果的逐像素凸组合。因此,融合由门本身执行,无需单独的融合模块,也不会在轮次间累积插值误差。决策阶段以固定的256x256分辨率运行,与输出分辨率无关,因此4K图像仅需K次纯查找。在四个基准上的大量实验表明,LoopLUT在保持4K实时吞吐量的同时,将PSNR比最强先前方法提高了最多2.81 dB。同样的分解也能很好地泛化到水下增强数据集。
cs.CV / 226 / 2609.33603

ViCoR: Reliable Molecular Structure Extraction via Spatially Aligned Verification and Executable Revision

ViCoR:通过空间对齐验证与可执行修正实现可靠的分子结构提取
Yuan, Yujian, Cai, Xin, Chen, Yufan, Xu, Jiaxin, Liu, Mengdi, Tan, Zhichao, Chen, Long, Gao, Hanyu
Abstract
Reliable optical chemical structure recognition (OCSR) is essential for building high-quality chemical data from scientific literature, yet even small recognition errors can propagate into chemical databases and downstream models. In practice, recognized structures often require manual inspection and correction before use, making large-scale data curation costly and difficult to scale. We therefore study Selective Structure Recognition (SSR), a post-recognition setting that automatically produces reliable structured outputs while rejecting unresolved cases. Selection-only approaches can improve reliability by rejection, but cannot create additional correct outputs beyond those produced by the base recognizer. We propose ViCoR, a repair-before-rejection framework for iterative VerIfiCatiOn and Revision. Its key idea is to make observation-prediction correspondence explicit: coordinate-preserving rendering establishes spatial correspondence between the source image and predicted structure, while index anchoring maps localized visual discrepancies to executable graph edits without full-structure regeneration. A shared VLM is progressively trained from verification to revision. On two real-world OCSR benchmarks, ViCoR improves overall accuracy from 73.53\% to 88.26\% and from 61.83\% to 84.32\%, while achieving over 97\% accepted accuracy at 85--89\% coverage. The resulting molecular data further improve reaction-extraction F1 by 15.5 points and literature-sourced reaction prediction accuracy by 7.7 and 5.8 points, demonstrating the value of automated reliability control for scientific data curation and downstream chemical learning.
Chinese Translation
可靠的光学化学结构识别(OCSR)对于从科学文献构建高质量化学数据至关重要,然而即使是微小的识别错误也可能传播到化学数据库和下游模型中。在实践中,识别出的结构在使用前往往需要人工检查和校正,使得大规模数据管理成本高昂且难以扩展。因此,我们研究了选择性结构识别(SSR),这是一种识别后设置,可自动产生可靠的结构化输出,同时拒绝未解决的案例。仅选择的方法可以通过拒绝来提高可靠性,但无法在基础识别器产生的输出之外创建额外的正确输出。我们提出了 ViCoR,一个先修复后拒绝的框架,用于迭代验证与修正。其核心思想是使观察-预测对应关系显式化:保持坐标的渲染建立源图像与预测结构之间的空间对应关系,而索引锚定则将局部视觉差异映射为可执行的图编辑,无需重新生成整个结构。共享的 VLM 从验证逐步训练到修正。在两个真实世界的 OCSR 基准上,ViCoR 将整体准确率从 73.53% 提高到 88.26%,并从 61.83% 提高到 84.32%,同时在 85-89% 的覆盖率下实现了超过 97% 的接受准确率。由此产生的分子数据进一步将反应提取 F1 提高了 15.5 个点,并将文献来源的反应预测准确率提高了 7.7 和 5.8 个点,证明了自动可靠性控制对于科学数据管理和下游化学学习的价值。
cs.CV / 227 / 2609.33611

Anguinus Sculpturae: Compositional Synthesis of Peak-Enhancement Breast DCE-MRI Scans

Anguinus Sculpturae:峰值增强乳腺 DCE-MRI 扫描的组合式合成
Hamm, Benjamin, Disch, Nico Albert, Rokuss, Maximilian, Kirchhoff, Yannick, Ulrich, Constantin, Maier-Hein, Klaus
Abstract
Dynamic contrast-enhanced breast MRI (DCE-MRI) is rich in anatomical and perfusion information, but its reliance on gadolinium-based contrast agents raises safety concerns and adds cost. Virtual contrast enhancement, synthesizing post-contrast from pre-contrast images, is a promising alternative. We address the MAMA-SYNTH challenge task of predicting peak-enhancement breast MRI. Rather than adopting the full machinery of diffusion or flow matching, we observe that under a rectified, straight-line path the generative process collapses to a single difference prediction: the synthetic peak image is the pre-contrast image plus a predicted enhancement map, recovered in one forward pass. Around this we build Anguinus Sculpturae, a compositional pipeline in which nnU-Net segmentations of lesion, foreground and breast region guide two generators - one optimized for global fidelity, one for lesion structure through an asymmetric Tversky term routed via a frozen segmenter - composited region-wise with Gaussian-weighted blending. On the held-out Duke subset of MAMA-MIA our model achieves the best FRD and Dice among all evaluated variants, showing that single-step difference prediction with segmentation guidance suffices to recover both global fidelity and lesion structure. Code is available at https://github.com/MIC-DKFZ/AnguinusSculpturae.
Chinese Translation
动态对比增强乳腺 MRI (DCE-MRI) 富含解剖和灌注信息,但其对钆基对比剂的依赖引发了安全性问题并增加了成本。虚拟对比增强,即从预对比图像合成后对比图像,是一种有前景的替代方法。我们解决 MAMA-SYNTH 挑战任务,即预测峰值增强乳腺 MRI。我们并未采用扩散模型或流匹配的完整机制,而是观察到在矫正的直线路径下,生成过程坍缩为单一的差异预测:合成峰值图像是预对比图像加上预测的增强图,通过一次前向传播即可恢复。围绕这一点,我们构建了 Anguinus Sculpturae,一个组合式流程,其中 nnU-Net 对病灶、前景和乳腺区域的分割指导两个生成器——一个针对全局保真度优化,另一个通过经冻结分割器路由的非对称 Tversky 项针对病灶结构——使用高斯加权混合按区域进行合成。在 MAMA-MIA 的留出 Duke 子集上,我们的模型在所有评估变体中取得了最佳的 FRD 和 Dice,表明带有分割指导的单步差异预测足以恢复全局保真度和病灶结构。代码可在 https://github.com/MIC-DKFZ/AnguinusSculpturae 获取。
cs.CV / 228 / 2609.33616

SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning

SpatialSpeak:面向空间思维链推理、融合局部与全局上下文的QA原生重建
Cao, Yang, Zhang, Jiaxin, Chen, Dave Zhenyu, Zhong, Yingji, Gao, Ruiyuan, Hong, Lanqing, Xu, Dan
Abstract
Vision-language models (VLMs) can benefit from geometric priors for multi-view spatial reasoning, yet answer-only training does not directly supervise the intermediate geometric estimates and their use in deriving quantitative spatial answers. We hypothesize that spatial chain-of-thought (CoT) supervision becomes more effective when the VLM first jointly learns complementary local geometry and global scene context through multi-view reconstruction. We introduce SpatialSpeak, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning. In Stage I, QA-Native Reconstruction Pretraining (QA-RP) combines marked-point 3D queries for fine-grained local geometry with object-center queries for global scene context across views. Both tasks are formulated as text-based question answering, allowing geometric estimation and subsequent reasoning to share the same autoregressive output interface. In Stage II, spatial CoT with Visual Compensation (CoT-VC) trains the model to express question-relevant geometric estimates and use them to derive answers, with reliability assessment and visual compensation supporting answer refinement when needed. On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points, and ablations show that both local and global reconstruction supervision are beneficial. SpatialSpeak achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.
Chinese Translation
视觉语言模型(VLMs)可以受益于几何先验以进行多视角空间推理,但仅答案训练并不能直接监督中间几何估计及其在推导定量空间答案中的使用。我们假设,当VLM首先通过多视角重建联合学习互补的局部几何与全局场景上下文时,空间思维链(CoT)监督会变得更有效。我们提出SpatialSpeak,一个两阶段框架,将QA原生重建预训练与空间CoT学习连接起来。在第一阶段,QA原生重建预训练(QA-RP)将用于细粒度局部几何的标记点3D查询与用于跨视角全局场景上下文的物体中心查询相结合。这两个任务均被形式化为基于文本的问答,从而使几何估计与后续推理共享同一自回归输出接口。在第二阶段,带视觉补偿的空间CoT(CoT-VC)训练模型表达与问题相关的几何估计,并利用它们推导答案;在需要时,通过可靠性评估和视觉补偿支持答案精炼。在ReVSI上,QA-RP将CoT-VC带来的增益从2.6分提高到6.9分,消融实验表明局部和全局重建监督均有益。SpatialSpeak在ReVSI、VSI-Bench和SPAR-Bench上取得了最先进的结果,ReVSI得分为62.8,比最强的对比基线高出8.7分。
cs.CV / 229 / 2609.33627

StoryEngine: A State-Grounded Agentic Framework for Video Storytelling

StoryEngine:一个基于状态的智能体视频叙事框架
Wang, Yingrui, Wang, Zeqing, Jin, Yeying
Abstract
Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, undermining both narrative coherence and visual consistency. To address these challenges, we propose StoryEngine, a state-grounded agentic framework for video storytelling. StoryEngine establishes a separation between authoritative semantic plans and unreliable visual observations. Specifically, StoryEngine maintains a structured representation of entity placement and story-relevant states, and propagates event-induced changes to define the intended start and end states of each shot. To visually realize these states, StoryEngine constructs canonical references for recurring entities and environments, and compiles state and visual constraints into executable render plans. Meanwhile, to realize these states correctly, a bounded evaluation-guided repair loop further corrects local state inconsistencies. Together, these mechanisms preserve causal story progression and prevent local visual errors from propagating across shots. To comprehensively evaluate long-form storytelling, we construct a benchmark across diverse scenarios and visual styles, with metrics assessing storytelling quality, narrative coherence, and visual consistency. Experimental results demonstrate that StoryEngine consistently outperforms state-of-the-art methods across all evaluation dimensions, validating its effectiveness for coherent and consistent video storytelling.
Chinese Translation
尽管智能体多镜头视频生成近期取得了进展,但生成连贯且一致的长篇故事仍然具有挑战性。现有的智能体流程通常依赖于文本镜头计划或先前生成的像素,却缺乏显式机制来传播故事事件的后果并跨镜头维护视频世界状态。因此,缺失的视觉细节可能被不准确地重建,而视觉漂移可能传播到后续镜头,损害叙事连贯性和视觉一致性。为了解决这些挑战,我们提出了StoryEngine,一个基于状态的智能体框架,用于视频叙事。StoryEngine在权威的语义计划与不可靠的视觉观察之间建立分离。具体而言,StoryEngine维护实体放置和故事相关状态的结构化表示,并传播事件引起的变化,以定义每个镜头的预期起始和结束状态。为了在视觉上实现这些状态,StoryEngine为重复出现的实体和环境构建规范参考,并将状态和视觉约束编译为可执行的渲染计划。同时,为了正确实现这些状态,一个有界评估引导的修复循环进一步纠正局部状态不一致。这些机制共同保持因果故事进展,并防止局部视觉错误跨镜头传播。为了全面评估长篇故事叙述,我们构建了一个涵盖多样场景和视觉风格的基准,其指标评估故事叙述质量、叙事连贯性和视觉一致性。实验结果表明,StoryEngine在所有评估维度上一致优于最先进方法,验证了其在连贯且一致的视频叙事方面的有效性。
cs.CV / 230 / 2609.33659

Learning Multimodal Embeddings with Evidence-Aligned Readout

基于证据对齐读出的多模态嵌入学习
Chen, Zirong, Ye, Fuda, Du, Enjun, Pu, Junfu, Wang, Xinlei, Zuo, Xinyu, Duan, Lisheng, Liang, Haijin, Ma, Jin, Wang, Jiachuan, Zhang, Yongqi
Abstract
Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether the semantic organization of that evidence can also specify where representations are read. To address this question, we introduce EviAlign, which couples Semantic Evidence Generation with Boundary Readout in a shared multimodal large language model. It organizes evidence into five semantic units, reads the contextualized state at each unit boundary, and aggregates these states into a single normalized embedding. Generation and contrastive retrieval objectives jointly train this shared structure. With the same trailing readout, semantic evidence and free-form CoT yield nearly identical retrieval performance, suggesting that evidence organization alone does not explain the full gain. A controlled $2\times3$ study compares consistent and permuted evidence organization across three readout strategies, using training targets with matched evidence spans. With five readout states and the same mean pooling, the advantage of consistent semantic organization grows from 0.65 points at length-based training positions to 2.39 at evidence boundaries, yielding a 1.74-point co-design interaction. Across 12 MMEB retrieval tasks, EviAlign achieves 76.9 average Recall@1 with 500K training pairs while retaining single-vector indexing and scoring.
Chinese Translation
多模态大语言模型可以通过生成揭示任务相关证据,但生成有用证据本身并不决定其如何进入检索嵌入。我们研究这些证据的语义组织是否也能指定表征在何处被读出。为解决这一问题,我们提出 EviAlign,它在共享的多模态大语言模型中将语义证据生成与边界读出相结合。它将证据组织为五个语义单元,在每个单元边界读取上下文相关状态,并将这些状态聚合为单个归一化嵌入。生成目标与对比检索目标联合训练这一共享结构。在相同的尾部读出下,语义证据与自由形式 CoT 产生几乎相同的检索性能,表明仅凭证据组织并不能解释全部增益。一项受控的 2×3 研究在三种读出策略下比较了一致和置换的证据组织,并使用具有匹配证据片段的训练目标。在五个读出状态和相同平均池化下,一致语义组织的优势从基于长度的训练位置处的 0.65 分增长到证据边界处的 2.39 分,产生 1.74 分的协同设计交互效应。在 12 个 MMEB 检索任务上,EviAlign 使用 500K 训练对达到 76.9 的平均 Recall@1,同时保持单向量索引和评分。
cs.CV / 231 / 2609.33668

When Noise Meets Long-Tail: Feature-Threshold Dual Calibration for Robust Pseudo-Labeling

当噪声遇到长尾:用于鲁棒伪标签的特征-阈值双重校准
Guo, Ping, Huang, Zhiqi, Li, Xinran
Abstract
Pseudo-labeling has become a cornerstone of learning from unlabeled data in semantic segmentation. Yet its effectiveness drops sharply in real-world scenarios where strong imaging noise and long-tailed class distributions occur together. We trace this failure to a vicious cycle of pseudo-label degradation. Imaging noise entangles foreground and background features, lowering prediction confidence across all classes, while long-tailed distributions leave tail classes with far fewer training samples and inherently lower confidence. Under fixed high-threshold filtering, these tail-class predictions are systematically filtered out, so they receive no supervision from unlabeled data and thus features keep degrading in subsequent iterations. Critically, noise and long-tail are not independent obstacles but mutually amplifying ones, and addressing either alone is insufficient. To break this cycle, we propose FTC-Seg, a Feature-Threshold dual-Calibration framework built on a standard teacher-student framework. At the feature level, Orthogonal Prototype Reconstruction (OPR) uses a set of learnable orthogonal prototypes to residually purify pixel-wise features, widening the margin between weak foreground targets and noisy backgrounds. At the threshold level, Adaptive Threshold Calibration (ATC) dynamically adjusts class-specific thresholds based on learning difficulty and prediction-distribution bias, rescuing low-confidence pseudo-labels of tail classes from systematic exclusion. Extensive experiments on four public benchmarks spanning three distinct noise modalities show that FTC-Seg achieves strong performance against state-of-the-art methods, with particularly substantial gains on tail classes. Our results establish that jointly calibrating features and thresholds is essential for robust pseudo-labeling under compounded noise and class imbalance.
Chinese Translation
伪标签已成为语义分割中从无标注数据学习的基石。然而,在强成像噪声和长尾类别分布同时出现的现实场景中,其有效性急剧下降。我们将这种失败归因于伪标签退化的恶性循环。成像噪声使前景和背景特征纠缠在一起,降低了所有类别的预测置信度,而长尾分布使得尾部类别的训练样本少得多,并且其置信度天然较低。在固定的高阈值过滤下,这些尾部类别的预测被系统性地过滤掉,因此它们无法从无标注数据中获得监督,从而在后续迭代中特征持续退化。至关重要的是,噪声和长尾并非独立的障碍,而是相互放大的,单独解决任何一个都是不够的。为了打破这一循环,我们提出了FTC-Seg,一个基于标准教师-学生框架的特征-阈值双校准框架。在特征层面,正交原型重建(OPR)使用一组可学习的正交原型对逐像素特征进行残差净化,扩大了弱前景目标与噪声背景之间的间隔。在阈值层面,自适应阈值校准(ATC)根据学习难度和预测分布偏差动态调整类别特定的阈值,将尾部类别的低置信度伪标签从系统性排除中挽救出来。在涵盖三种不同噪声模式的四个公开基准上进行的大量实验表明,FTC-Seg相较于最先进的方法取得了强劲的性能,在尾部类别上尤其有显著提升。我们的结果证实,在复合噪声和类别不平衡下,联合校准特征和阈值对于鲁棒伪标签至关重要。
cs.CV / 232 / 2609.33683

MAD-Guard: Controlled Study of Autoregressive Generation versus Direct Decision Interfaces for Closed Multimodal Forensic Tasks

MAD-Guard:面向封闭式多模态取证任务的自回归生成与直接决策接口的对照研究
Chen, Hao
Abstract
When should multimodal foundation models generate tokens, and when should they directly output a decision? We present MAD-Guard, a controlled study of output-decision interfaces for closed multimodal forensic tasks. Once a multimodal representation is computed, is autoregressive generation necessary for closed forensic decisions with high input complexity but low output entropy? Under a matched Qwen3-VL-8B backbone, 2,400 FakeClue training samples, and LoRA budget ($r=16, \alpha=32$) on Huawei Ascend 910C NPUs, we evaluate a progression of decision interfaces (AR-SFT [generate] $\to$ Logit Slice $\to$ Binary Direct Head $\to$ +choice $\to$ +act $\to$ CLM-Head) and decompose latency into backbone representation (53.12 ms), 151,643-way vocabulary projection (+85.04 ms $\to$ 138.16 ms), and decoding (+248.26 ms $\to$ 386.42 ms). Under 1-to-1 binary supervision ($\mathcal{L}_{\mathrm{BCE}}$), a Binary Direct Head cuts latency by $2.60\times$-$7.27\times$ (53.12 ms) and lowers calibration error by $1.88\times$ (ECE = 0.0450 vs. 0.0845), with a -1.80% accuracy trade-off (93.10% vs. 94.90%; 0.9795 vs. 0.9871 ROC-AUC) from forfeiting token priors. Gains above AR-SFT arise either from multi-task attribution and uncertainty gating (+choice+act: 96.44% accuracy, 0.9940 ROC-AUC, 0.0187 ECE at 53.71 ms) or from a disaggregated contrastive head (CLM-Head: 96.55% binary and 96.44% multi-task accuracy, 0.0166 ECE, 98.79% 7-class attribution at 54.42 ms) retaining semantic priors without token decoding. Across 5,000 out-of-sample images from five benchmarks, our framework excels on synthetic, camouflage, and document forgeries (96.44% GenImage, 97.73% Chameleon, 91.84% Doc) while showing a clear boundary on compressed face manipulation (FF++ ROC-AUC = 0.5913).
Chinese Translation
多模态基础模型何时应生成 token,何时应直接输出决策?我们提出 MAD-Guard,一项面向封闭式多模态取证任务的输出决策接口的对照研究。一旦计算出多模态表征,对于输入复杂度高但输出熵低的封闭式取证决策,自回归生成是否必要?在匹配的 Qwen3-VL-8B 骨干模型、2,400 个 FakeClue 训练样本,以及在华为昇腾 910C NPU 上的 LoRA 预算(r=16, α=32)条件下,我们评估了一系列决策接口(AR-SFT [生成] → Logit Slice → Binary Direct Head → +choice → +act → CLM-Head),并将延迟分解为骨干表征(53.12 ms)、151,643 路词表投影(+85.04 ms → 138.16 ms)和解码(+248.26 ms → 386.42 ms)。在 1 对 1 二元监督(L_BCE)下,Binary Direct Head 将延迟降低 2.60×–7.27×(53.12 ms),并使校准误差降低 1.88×(ECE = 0.0450 vs. 0.0845),代价是 -1.80% 的准确率权衡(93.10% vs. 94.90%;0.9795 vs. 0.9871 ROC-AUC),这源于放弃 token 先验。相较于 AR-SFT 的增益,要么来自多任务归因与不确定性门控(+choice+act:96.44% 准确率、0.9940 ROC-AUC、53.71 ms 时 ECE 0.0187),要么来自分解式对比头(CLM-Head:96.55% 二元和 96.44% 多任务准确率、0.0166 ECE、54.42 ms 时 98.79% 的 7 类归因),后者无需 token 解码即可保留语义先验。在来自五个基准测试的 5,000 张样本外图像上,我们的框架在合成、伪装和文档伪造上表现优异(GenImage 96.44%、Chameleon 97.73%、Doc 91.84%),但在压缩人脸篡改上显示出明显边界(FF++ ROC-AUC = 0.5913)。
cs.CV / 233 / 2609.33687

Resource-Aware Parameter-Efficient Model Adaptation for Onboard High-Dimensional Data

面向星载高维数据的资源感知参数高效模型适配
Zhang, Qiyang, Li, Xinhao, Shi, Lei, Lin, Zheng, Wen, Jinfeng, Zhou, Ao, Wang, Shangguang
Abstract
Onboard satellite models often require frequent updates, but the weights adapted to earlier data distributions can quickly become outdated. However, updating large-scale model parameters in orbit presents significant challenges due to the limited uplink bandwidth of Low Earth Orbit (LEO) satellite systems, particularly for hyperspectral satellite imagery, where high-dimensional spectral-spatial inputs lead to increased model size and update costs. Existing full fine-tuning methods are thus expensive to retrain and difficult to deploy under strict communication constraints. To address this challenge, we propose NE-LoRA, a parameter-efficient adaptation framework for bandwidth-constrained onboard hyperspectral model updates. NE-LoRA combines a primary low-rank branch with a nonlinear auxiliary branch to capture both global update trends and complex spectral-spatial variations. Additionally, we introduce a differentiated training strategy for multi-matrix adapters, motivated by the asymmetric initialization and gradient dynamics of different adapter matrices. Experiments on four hyperspectral datasets and three representative backbone models demonstrate that NE-LoRA consistently outperforms LoRA-based baselines and remains competitive with, and in several cases superior to, full fine-tuning. Across the evaluated settings, NE-LoRA updates only a small fraction of the total parameters on average while preserving low deployment overhead, offering a favorable accuracy-communication trade-off for onboard hyperspectral adaptation.
Chinese Translation
星载卫星模型通常需要频繁更新,但针对早期数据分布适配的权重可能很快过时。然而,由于低地球轨道(LEO)卫星系统的上行带宽有限,在轨更新大规模模型参数面临重大挑战,特别是对于高光谱卫星影像,其高维光谱-空间输入导致模型规模和更新成本增加。因此,现有的全参数微调方法重新训练成本高昂,且在严格的通信约束下难以部署。为了解决这一挑战,我们提出了NE-LoRA,一种面向带宽受限星载高光谱模型更新的参数高效适配框架。NE-LoRA结合了一个主低秩分支和一个非线性辅助分支,以捕获全局更新趋势和复杂的光谱-空间变化。此外,受不同适配器矩阵的非对称初始化和梯度动态的启发,我们引入了一种针对多矩阵适配器的差异化训练策略。在四个高光谱数据集和三个代表性骨干模型上的实验表明,NE-LoRA始终优于基于LoRA的基线方法,并且与全参数微调相比具有竞争力,在若干情况下甚至更优。在评估的各种设置中,NE-LoRA平均仅更新总参数的一小部分,同时保持较低的部署开销,为星载高光谱适配提供了有利的精度-通信权衡。
cs.CV / 234 / 2609.33694

Seeing and Solving Are Not Enough for Vision-Language Models

对于视觉语言模型,看见与解决是不够的
Wang, Ziheng, Xie, Mingxuan, Liu, Yilin, Wu, Dayan, Li, Yang, Dai, Pengwen
Abstract
Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separately yet still fail on the original multimodal question, a distinction that overall answer accuracy cannot reveal. To study this, we perform a question-level empirical analysis across multiple VLMs and visual domains. We define an exactly scorable task state (i.e., the visual information sufficient to solve a question) and use it to test whether the same model can extract the required state, solve the question from the ground-truth state, and answer the original multimodal question. We find that composition failures, where extraction and solving both succeed but direct answering fails, account for 17.7% to 75.6% of direct-answering errors across multiple VLMs and datasets. To address this failure mode, we introduce a simple yet effective method, termed State Realization Tuning (SRT). SRT fine-tunes LoRA adapters attached to the language-model layers while keeping the pretrained VLM weights frozen. It trains the model to output the ground-truth task state before the final answer in a single autoregressive response. SRT improves over standard supervised fine-tuning by 1.7 to 14.1 percentage points and repairs 92.5% to 98.1% of diagnosed composition failures. A single LoRA adapter trained with SRT also improves performance across substantially different task-state structures. Our work shows that having both visual extraction and problem-solving capabilities does not guarantee correct multimodal answering. Requiring the model to first output the visual information needed to solve the question can help bridge this gap.
Chinese Translation
视觉语言模型(VLMs)通过结合视觉信息提取与下游问题求解来回答视觉问题。我们研究一个基本问题:错误的答案是否必然反映了视觉提取或问题求解的失败?一个模型在分别测试时可能在这两种能力上都成功,但在原始多模态问题上仍然失败,这一区别是总体答案准确率无法揭示的。为了研究这一点,我们对多个VLM和视觉领域进行了问题级别的实证分析。我们定义了一个可精确评分的任务状态(即足以解决一个问题的视觉信息),并用它来测试同一个模型是否能够提取所需状态、从真实状态解决问题,并回答原始多模态问题。我们发现,组合失败(composition failures,即提取和求解都成功但直接回答失败)在多个VLM和数据集中占直接回答错误的17.7%到75.6%。为了解决这种失败模式,我们引入了一种简单而有效的方法,称为状态实现调优(State Realization Tuning, SRT)。SRT微调附加到语言模型层的LoRA适配器,同时保持预训练VLM权重冻结。它训练模型在单个自回归响应中,在最终答案之前输出真实任务状态。SRT比标准监督微调提高了1.7到14.1个百分点,并修复了92.5%到98.1%的诊断出的组合失败。使用SRT训练的单个LoRA适配器也能在显著不同的任务状态结构上提高性能。我们的工作表明,同时具备视觉提取和问题求解能力并不能保证正确的多模态回答。要求模型首先输出解决问题所需的视觉信息可以帮助弥合这一差距。
cs.CV / 235 / 2609.33701

Prompt-Anchored Residual Adaptation for Biomedical Vision-Language Models

面向生物医学视觉语言模型的提示锚定残差适配
Kang, Jingxuan, Yue, Qianying, Liu, Che, Qin, Chen
Abstract
Pretrained biomedical vision-language models achieve strong zero-shot performance in biomedical image classification. However, downstream biomedical classification often depends on subtle visual differences between classes that may not be fully captured by pretrained representations. Few-shot adaptation addresses this mismatch by optimizing a task-specific predictor on a small labeled support set. Because the selected examples capture only part of the visual variation within the target classes, the adapted predictions can depend strongly on their composition. We propose Prompt-Anchored Residual Adaptation (PARA), which retains the frozen prompt prediction as a support-invariant semantic anchor and incorporates a visual prediction learned from the support set through an anchor-relative residual. The residual step is computed in a closed form from frozen support embeddings using anchor discrepancy and support agreement. Support-set dependence also limits evaluation: comparisons are fair within a shared draw but remain conditional on its composition. To obtain more reliable comparisons, we introduce a repeated-support protocol that separates support-selection variation from optimization randomness and reports both average and worst-20% performance. PARA achieves state-of-the-art performance in both few-shot classification and base-to-novel generalization.
Chinese Translation
预训练的生物医学视觉-语言模型在生物医学图像分类中取得了强大的零样本性能。然而,下游生物医学分类往往依赖于类别之间细微的视觉差异,而预训练表征可能无法完全捕捉这些差异。少样本适配通过在小型有标注支持集上优化任务特定预测器来缓解这种不匹配。由于所选样本仅能捕获目标类别内部视觉变化的一部分,适配后的预测可能强烈依赖于其组成。我们提出提示锚定残差适配(Prompt-Anchored Residual Adaptation, PARA),它将冻结的提示预测保留为支持集不变的语义锚点,并通过锚点相对残差融入从支持集学到的视觉预测。残差步骤基于冻结的支持集嵌入,利用锚点差异和支持集一致性以闭式形式计算。支持集依赖性也限制了评估:在共享抽样内的比较是公平的,但仍以其组成为条件。为获得更可靠的比较,我们引入重复支持集协议,将支持集选择变化与优化随机性分离,并报告平均性能和最差20%性能。PARA在少样本分类和基类到新类泛化方面均达到最先进性能。
cs.CV / 236 / 2609.33716

Revisiting Diffusion Fine-Tuning for Unsupervised Domain Adaptation

重新审视用于无监督域适应的扩散微调
Qi, Xuan, Wei, Yi, Berardini, Daniele, Pastore, Vito Paolo, Murino, Vittorio
Abstract
Diffusion-based unsupervised domain adaptation (UDA) improves cross-domain transfer by generating target-specific synthetic data for downstream adaptation. Existing methods are largely designed for single-target adaptation: when a model trained on one labeled source domain must be adapted to multiple unlabeled target domains, they typically require separate diffusion fine-tuning for each source--target pair, causing training, storage, and deployment costs to grow with the number of targets. In this paper, we study multi-target data generation for diffusion-based UDA, where a single source-guided diffusion fine-tuning process is reused to generate target-specific synthetic data for multiple target domains. We propose MUSE (Multi-target UDA-oriented Synthesis with Efficient diffusion fine-tuning), a decoupled adaptation framework that separates source-supervised semantic adaptation from target-specific style adaptation. MUSE uses a shared semantic branch updated by labeled source data and target-private style branches specialized to individual target domains, enabling target-specific generation while avoiding repeated source-guided fine-tuning for each target. Experiments on standard UDA benchmarks show that MUSE achieves a stronger accuracy--efficiency trade-off than repeated per-target diffusion adaptation, reducing diffusion fine-tuning cost while improving average target-domain accuracy. The project page is available at https://xuanqi99.github.io/MUSE/.
Chinese Translation
基于扩散的无监督域适应(UDA)通过生成目标特定的合成数据用于下游适应,从而改善跨域迁移。现有方法大多针对单目标适应设计:当在一个有标注源域上训练的模型需要适应多个无标注目标域时,它们通常需要为每个源-目标对分别进行扩散微调,导致训练、存储和部署成本随目标数量增长。本文研究基于扩散的UDA的多目标数据生成,其中单个源引导的扩散微调过程被重用来为多个目标域生成目标特定的合成数据。我们提出MUSE(面向多目标UDA的高效扩散微调合成),一个解耦的适应框架,将源监督的语义适应与目标特定的风格适应分离。MUSE使用由有标注源数据更新的共享语义分支和专门针对各个目标域的目标私有风格分支,实现目标特定的生成,同时避免对每个目标重复进行源引导的微调。在标准UDA基准上的实验表明,MUSE比重复的逐目标扩散适应实现了更强的精度-效率权衡,在提高平均目标域精度的同时降低了扩散微调成本。项目页面见https://xuanqi99.github.io/MUSE/。
cs.CV / 237 / 2609.33723

GeoShrink: Accelerating Diffusion Transformers with Two Lines of Code

GeoShrink:用两行代码加速扩散Transformer
Li, Haosen, Chen, Wenshuo, Liang, Shaofeng, Wang, Lei, Tian, Bowen, Yue, Yutao
Abstract
Diffusion transformers incur substantial inference cost through repeated model evaluations along a sampling trajectory. We introduce GeoShrink, a training-free acceleration method that retains the original solver grid while evaluating the model only at a prescribed set of anchors. At skipped stages, GeoShrink predicts the solver-facing output by adding a geometrically retained fraction of the latest observed innovation to the most recent exact output. We derive this rule from chordal tangent transport and round-trip line projection, and establish a geometric anchor-spacing principle that minimizes the largest adjacent gap expansion under fixed coverage and first span. The analysis characterizes the geometric closure and propagation of prediction errors without assuming access to future model outputs. Experiments cover image, video, motion, and audio generation, together with adapted 3D backends. At approximately $5\times$ acceleration, GeoShrink improves FLUX PSNR by 3.10 dB over the strongest listed baseline. On HunyuanVideo, it achieves a reported $4.99\times$ speedup and improves ChronoMagic-Bench-150 PSNR by 5.44 dB over the strongest listed fidelity baseline. Comparisons at fixed evaluation budgets further show substantial gains on motion, audio, music, and 3D generation.
Chinese Translation
扩散Transformer因沿采样轨迹重复进行模型评估而带来巨大的推理开销。我们提出GeoShrink,一种免训练加速方法,它保留原始求解器网格,同时仅在规定的锚点集合处评估模型。在跳过的阶段,GeoShrink通过将最近观测到的新息的几何保留比例加到最近的精确输出上,来预测面向求解器的输出。我们从弦切线传输和往返线投影推导出该规则,并建立了一个几何锚点间距原则,在固定覆盖和第一跨度下最小化最大相邻间隙扩展。该分析刻画了预测误差的几何闭合与传播,而不假设可以访问未来的模型输出。实验涵盖图像、视频、动作、音频生成,以及适配的3D后端。在约$5\times$加速下,GeoShrink将FLUX的PSNR比列出的最强基线提高了3.10 dB。在HunyuanVideo上,它实现了报告的$4.99\times$加速,并将ChronoMagic-Bench-150的PSNR比列出的最强保真度基线提高了5.44 dB。在固定评估预算下的比较进一步显示,在动作、音频、音乐和3D生成上取得了显著增益。
cs.CV / 238 / 2609.33735

Constrained Edit Fields for Training-Free Flow Editing

用于免训练流编辑的约束编辑场
Kang, Jingxuan, Wang, Yinsong, Liu, Che, Qin, Chen
Abstract
Text-guided image editing aims to perform a desired edit while preserving source content unrelated to it. Pretrained rectified-flow models enable training-free editing of real images through modifications to their sampling trajectories. However, responses at locations unrelated to the desired edit can still accumulate along the editing trajectory and become visible in the final result. To overcome this, we propose Constrained Edit Fields (CEF), which assigns each spatial location a continuous edit responsibility that quantifies its relevance to the desired edit. CEF estimates edit responsibility directly from the source image when the relevant content is present. For edits whose target content is absent from the source, CEF first generates an unconstrained proposal to reveal its realized spatial support and then estimates responsibility from that proposal. At each editing step, CEF decomposes the base edit field into prompt-induced and trajectory-induced components, enabling edit responsibility to preserve instruction-relevant updates while suppressing unintended trajectory-induced changes. Evaluated on all 700 PIE-Bench examples, CEF achieves state-of-the-art Structure Distance, background LPIPS, and background MSE with both Stable Diffusion 3.5 Medium and FLUX, while retaining competitive instruction alignment. On Stable Diffusion 3.5 Medium, it reduces these metrics over the previous best results by 10.2%, 21.2%, and 48.0%, respectively.
Chinese Translation
文本引导的图像编辑旨在执行所需的编辑,同时保留与其无关的源内容。预训练的矫正流模型通过修改其采样轨迹,能够实现对真实图像的免训练编辑。然而,与所需编辑无关的位置处的响应仍可能沿着编辑轨迹累积,并在最终结果中变得可见。为了克服这一点,我们提出了约束编辑场(CEF),它为每个空间位置分配一个连续的编辑责任,量化其与所需编辑的相关性。当相关内容存在时,CEF 直接从源图像估计编辑责任。对于目标内容不在源中的编辑,CEF 首先生成一个无约束的提议以揭示其实现的空间支持,然后从该提议中估计责任。在每个编辑步骤中,CEF 将基础编辑场分解为提示诱导和轨迹诱导的成分,使得编辑责任能够保留指令相关的更新,同时抑制意外的轨迹诱导变化。在所有 700 个 PIE-Bench 示例上进行评估,CEF 在 Stable Diffusion 3.5 Medium 和 FLUX 上均实现了最先进的结构距离、背景 LPIPS 和背景 MSE,同时保持了具有竞争力的指令对齐。在 Stable Diffusion 3.5 Medium 上,它将这些指标相比之前的最佳结果分别降低了 10.2%、21.2% 和 48.0%。
cs.CV / 239 / 2609.33758

ENet-GP: Unified Document Image Restoration

ENet-GP:统一文档图像恢复
Burad, Sujal, Aakanksha, Rajagopalan, A. N., Shekar, Sumit
Abstract
Reliable document digitization in uncontrolled capture settings is challenging because real images exhibit multiple interacting degradations rather than a single isolated distortion. Documents thus captured are affected simultaneously by geometric distortions, like page warping, as well as photometric degradations such as non-uniform illumination, and blurring. However, most existing approaches address these factors independently and are evaluated on benchmarks containing only one distortion type, limiting their real-world applicability. We introduce GutenDoc, a large-scale dataset of high-resolution dense-text documents with physically grounded compound degradations. Using physics-based rendering, our dataset jointly models geometric warping and diverse photometric effects, enabling systematic evaluation under realistic capture conditions. We further propose a unified restoration framework that jointly corrects geometric and photometric distortions within a single-network and single-training setup, without the need for degradation-specific retraining or sequential inference passes. Extensive experiments show that our method remains competitive on established single-distortion benchmarks while substantially improving robustness under compound degradations, providing a practical solution for real-world document digitization.
Chinese Translation
在不受控的拍摄环境中进行可靠的文档数字化具有挑战性,因为真实图像表现出多种相互作用的退化,而非单一孤立的失真。如此捕获的文档同时受到几何畸变(如页面弯曲)以及光度退化(如光照不均和模糊)的影响。然而,大多数现有方法独立地处理这些因素,并在仅包含一种失真类型的基准上进行评估,限制了其现实世界的适用性。我们引入了 GutenDoc,一个大规模的高分辨率密集文本文档数据集,具有基于物理的复合退化。使用基于物理的渲染,我们的数据集联合建模了几何弯曲和多样的光度效应,使得在真实拍摄条件下进行系统评估成为可能。我们进一步提出了一个统一的恢复框架,在单一网络和单次训练设置中联合校正几何和光度畸变,无需针对特定退化进行重新训练或顺序推理。大量实验表明,我们的方法在现有的单一失真基准上保持竞争力,同时在复合退化下显著提高了鲁棒性,为现实世界的文档数字化提供了实用解决方案。
cs.CV / 240 / 2609.33769

M3-Score: Fidelity, Memorization and Coverage as Separate Axes for Evaluating Generative Radiology Image Models

M3-Score:保真度、记忆与覆盖度作为评估生成式放射学图像模型的独立轴
Nishankar, Sathiyamohan, Sanjeewani, Pubudu, Perera, Asanka
Abstract
Quantitative evaluation of generative models for radiology remains challenging. Clinically relevant structures are often small and infrequent, feature spaces learned from natural images may represent them poorly, and a single summary score cannot distinguish limited fidelity from limited diversity. This study proposes the Medical Multi-axis Maximum Mean Discrepancy score (M3-Score), an evaluation framework based on RadioDINO-s16, a frozen vision transformer pretrained on radiology images. M3-Score reports three complementary axes computed at pre-specified encoder depths: \emph{fidelity}, measured by an unbiased multi-bandwidth radial basis function (RBF) MMD$^2$ at the final block; \emph{memorization}, measured by nearest-neighbor distances at 75\% depth; and \emph{coverage}, defined as the fraction of real images with a generated neighbor within their $k$-nearest-neighbor radius at 33\% depth. Reference sets are sampled across subjects to limit the influence of correlated slices. On BraTS brain MRI, the fidelity axis ordered five comparison sets of increasing severity (Spearman $\rho = 1.00$), and a subject-disjoint real set yielded $\mathrm{MMD}^2 = 0$ (permutation $p = 1$). An unconditional denoising diffusion probabilistic model achieved $\mathrm{MMD}^2 = 0.073$ (95\% confidence interval $[0.071, 0.080]$) but covered only 38\% of the real distribution. Under progressive mode dropping, $k$-NN manifold recall increased at all twelve encoder blocks, whereas the proposed coverage estimator decreased monotonically ($\rho = -1.00$). RadioDINO-s16 features separated real brain MRI from generated samples with a ROC-AUC of 0.819, compared with 0.555 for InceptionV3 and 0.582 for CLIP. Across a twentyfold range of sample sizes, the mean M3 value varied by a factor of 1.05, compared with 2.52 for the Fr\'{e}chet Inception Distance.
Chinese Translation
定量评估放射学生成模型仍然具有挑战性。临床相关结构通常较小且不常见,从自然图像中学习到的特征空间可能无法很好地表示它们,而单一汇总分数无法区分有限的保真度和有限的多样性。本研究提出了医学多轴最大均值差异分数(M3-Score),这是一个基于 RadioDINO-s16 的评估框架,RadioDINO-s16 是在放射学图像上预训练的冻结视觉 Transformer。M3-Score 在预先指定的编码器深度上报告三个互补的轴:保真度,由最终块处的无偏多带宽径向基函数(RBF)MMD^2 度量;记忆,由 75% 深度处的最近邻距离度量;覆盖度,定义为在其 k 近邻半径内具有生成邻居的真实图像的比例(33% 深度处)。参考集跨受试者采样,以限制相关切片的影响。在 BraTS 脑 MRI 上,保真度轴对五个严重程度递增的比较集进行了排序(Spearman ρ = 1.00),且受试者不相交的真实集得到 MMD^2 = 0(置换 p = 1)。无条件去噪扩散概率模型达到 MMD^2 = 0.073(95% 置信区间 [0.071, 0.080]),但仅覆盖了真实分布的 38%。在渐进模式丢弃下,k-NN 流形召回率在所有十二个编码器块中均增加,而所提出的覆盖度估计量单调下降(ρ = -1.00)。RadioDINO-s16 特征以 0.819 的 ROC-AUC 将真实脑 MRI 与生成样本区分开,而 InceptionV3 为 0.555,CLIP 为 0.582。在二十倍样本量范围内,平均 M3 值变化因子为 1.05,而 Fréchet Inception Distance 为 2.52。
cs.CV / 241 / 2609.33811

Eyes on the Road: A Naturalistic Comparison of MTW Rider Gaze in Urban Indian Traffic

关注道路:印度城市交通中MTW骑手注视的自然情境比较
Srivastava, Prerak, Kumar, Bhaiya Vaibhaw, Vemuri, Kavita
Abstract
Motorized two-wheelers (MTW) dominate Indian roads but remain underrepresented in driver behavior research. This study presents the first large-scale analysis of MTW driver gaze behavior in naturalistic, heterogeneous urban traffic, using the \textit{myEye2Wheeler} dataset. A semantic segmentation pipeline (YOLOv11 + SAM2) was used to extract object-level gaze metrics under two attention modes: direct gaze (foveal overlap) and central vision (parafoveal monitoring). Results reveal a functional division: central vision supports broad monitoring, while direct gaze enables brief, selective sampling. Novice riders exhibit road-anchored scanning, returning to the road between object fixations, while experienced riders form longer chains of attention across multiple objects. The findings suggest that experience primarily refines temporal rhythm rather than altering allocation strategy and reduces object-class effects in gaze patterns. These findings offer new insight into MTW attention structures and inform future work on behavior modeling and safety systems.
Chinese Translation
机动两轮车(MTW)在印度道路上占主导地位,但在驾驶员行为研究中仍然代表性不足。本研究利用myEye2Wheeler数据集,首次对自然情境下异质城市交通中的MTW骑手注视行为进行了大规模分析。采用语义分割流程(YOLOv11 + SAM2)提取两种注意模式下的对象级注视指标:直接注视(中央凹重叠)和中心视觉(副中央凹监控)。结果揭示了功能分工:中心视觉支持广泛监控,而直接注视则实现短暂的选择性采样。新手骑手表现出以道路为锚点的扫视,在对象注视之间返回道路,而经验丰富的骑手则在多个对象之间形成更长的注意链。研究结果表明,经验主要优化了时间节奏,而非改变分配策略,并减少了注视模式中的对象类别效应。这些发现为MTW注意结构提供了新见解,并为行为建模和安全系统的未来工作提供了信息。
cs.CV / 242 / 2609.33818

Augmenting Visual Anomaly Detection with Automated Interpretability

利用自动化可解释性增强视觉异常检测
De Santis, Antonio, Leo, Arsenio, Brambilla, Marco
Abstract
Visual anomaly detectors identify deviations from known-normal data, but their anomaly signals may mix evidence of actual anomalies with benign visual variation. We investigate whether automated interpretability can augment visual anomaly detectors by identifying and intervening on different components of this signal. We decompose PatchCore nearest-normal residuals into sparse features using Sparse Autoencoders (SAEs), and provide high-activation and contrastive non-active examples to a Multimodal LLM, which describes each feature and labels it as anomaly, distractor, or uncertain. These labels guide interventions in the SAE hidden representation, where distractor features are suppressed and anomaly features amplified. The edited representation is then used to reconstruct patch embeddings, which are rescored with PatchCore. Across 40 categories from four benchmarks, applying both interventions jointly improves macro-average image-level AUROC from 0.8724 to 0.8857 on source data and from 0.8066 to 0.8210 under synthetic corruptions. On three additional RobustAD categories with real acquisition shifts, the same interventions improve AUROC from 0.8745 to 0.9056 on source data and from 0.6069 to 0.6599 under real acquisition shifts. Finally, individual feature interventions across all 43 categories show that the MLLM labels are aligned in aggregate with how features differently affect normal and anomalous images.
Chinese Translation
视觉异常检测器识别与已知正常数据的偏差,但其异常信号可能混合了真实异常的证据与良性视觉变化。我们研究自动化可解释性是否能够通过识别并干预该信号的不同组成部分来增强视觉异常检测器。我们使用稀疏自编码器(SAEs)将PatchCore的最近正常残差分解为稀疏特征,并向多模态LLM提供高激活和对比非激活示例,由其描述每个特征并将其标记为异常、干扰或不确定。这些标签指导对SAE隐藏表示的干预,其中干扰特征被抑制,异常特征被放大。然后使用编辑后的表示重建图像块嵌入,并用PatchCore重新评分。在来自四个基准的40个类别中,联合应用两种干预措施使宏平均图像级AUROC从源数据上的0.8724提升至0.8857,在合成损坏下从0.8066提升至0.8210。在另外三个具有真实采集偏移的RobustAD类别上,相同的干预措施使AUROC在源数据上从0.8745提升至0.9056,在真实采集偏移下从0.6069提升至0.6599。最后,在所有43个类别上的单个特征干预表明,MLLM标签总体上与特征对正常和异常图像的不同影响方式相一致。
cs.CV / 243 / 2609.33833

One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs

一次攻击欺骗所有:针对前沿MLLM的高迁移性黑盒对抗攻击
Nie, Sen, Zhang, Jie, Wang, Zhongqi, Shan, Shiguang, Chen, Xilin
Abstract
Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can consistently mislead diverse frontier MLLMs in black-box settings. We propose O-Attack, a highly transferable black-box attack framework. This framework builds on our insight that surrogate models contain a broad, high-level, cross-modally aligned semantic space. This space extends beyond final-layer outputs and provides multiple semantically consistent representations that remain underexploited by existing attacks. Within this space, O-Attack anchors aligned representations, progressively broadens semantic conditions, and optimizes perturbations through semantic consensus to promote consistent target alignment. By fully exploiting this space with the same surrogate models as M-Attack, O-Attack raises attack success rates on GPT-5.4 (29.1% to 77.2%), Claude-4.6 (42.8% to 81.6%), and Gemini-3.1 (38.2% to 80.9%). Extensive experiments across 24 MLLMs show that O-Attack outperforms six state-of-the-art methods in black-box transferability, with consistent effectiveness across prompts and improved efficiency and imperceptibility. This work exposes the practical safety risks posed by black-box adversarial attacks against frontier MLLMs, underscoring the need for more rigorous robustness evaluation and more effective defenses.
Chinese Translation
对抗攻击长期以来对机器学习系统构成根本性威胁。随着多模态大语言模型(MLLM)快速演进并广泛部署,评估其对此类攻击的脆弱性对其安全使用至关重要。在本工作中,我们研究单张对抗图像能否在黑盒设置下持续欺骗多样化的前沿MLLM。我们提出O-Attack,一个高迁移性的黑盒攻击框架。该框架基于我们的洞察:代理模型包含一个广泛、高层、跨模态对齐的语义空间。该空间超越了最终层输出,提供了多个语义一致的表示,这些表示尚未被现有攻击充分利用。在该空间内,O-Attack锚定对齐表示,逐步拓宽语义条件,并通过语义共识优化扰动,以促进一致的目标对齐。通过使用与M-Attack相同的代理模型充分利用该空间,O-Attack将GPT-5.4(29.1%至77.2%)、Claude-4.6(42.8%至81.6%)和Gemini-3.1(38.2%至80.9%)上的攻击成功率提高。在24个MLLM上的大量实验表明,O-Attack在黑盒迁移性上优于六种最先进方法,在不同提示下效果一致,并提高了效率和不可感知性。这项工作揭示了针对前沿MLLM的黑盒对抗攻击所带来的实际安全风险,强调了进行更严格鲁棒性评估和更有效防御的必要性。
cs.CV / 244 / 2609.33834

CLIMB-flow: Coupled Linear Inverse posterior sampling via Multiscale-Based flow

CLIMB-flow:通过基于多尺度的流进行耦合线性逆问题后验采样
Yu, Zeqiu, Yuan, Ruizhi, Jacob, Mathews
Abstract
Diffusion models are now widely used in Bayesian inverse problems in imaging as priors, where latent diffusion models are often used for larger scale problems to keep the computational complexity and model-size manageable. Unfortunately, the auto-encoder based compression results in loss of spatial detail. In addition, the optimization is converted to a non-linear problem. In this paper, we introduce a posterior sampling algorithm customized for the pyramidal/cascaded architecture, which relies on a coarse to fine hierarchical strategy to generate images in the pixel domain. We present CLIMB-Flow which alternates between three steps: an end-point estimation from the current coarse and noisy image, data-consistent update of the clean image, and re-noising it back to the level the network expects. Together these steps sample the posterior at that scale using an approximate Gibbs sampling from two conditional distributions. Experiments on ImageNet, CelebA, AFHQ and fastMRI span inpainting, deblurring, super-resolution and accelerated MRI, with PSNR gains of 1.37-7.66 dB over the strongest competing method on CelebA and pixel-domain reconstruction up to 512x512.
Chinese Translation
扩散模型现在广泛用作成像中贝叶斯逆问题的先验,其中潜扩散模型通常用于更大规模的问题,以保持计算复杂度和模型大小可控。不幸的是,基于自编码器的压缩会导致空间细节的丢失。此外,优化问题转化为非线性问题。在本文中,我们介绍了一种针对金字塔/级联架构定制的后验采样算法,该算法依赖于从粗到细的分层策略在像素域生成图像。我们提出 CLIMB-Flow,它在三个步骤之间交替:从当前粗糙且含噪图像进行端点估计、对干净图像进行数据一致性更新,以及将其重新加噪回网络期望的噪声水平。这些步骤一起使用来自两个条件分布的近似 Gibbs 采样,在该尺度上对后验进行采样。在 ImageNet、CelebA、AFHQ 和 fastMRI 上的实验涵盖图像修复、去模糊、超分辨率和加速 MRI,在 CelebA 上相比最强竞争方法获得了 1.37-7.66 dB 的 PSNR 提升,并实现了高达 512x512 的像素域重建。
cs.CV / 245 / 2609.33854

ReDrive: Shaping Representations with World Modeling for End-to-End Driving

ReDrive:利用世界建模塑造表征以实现端到端驾驶
Zhu, Yueting, Chen, Shaoyu, Song, Yuehao, Sun, Hui, Zhang, Qian, Liu, Wenyu, Wang, Xinggang
Abstract
Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system architecture. Inspired by the transferability of general-purpose visual representations, we argue that combining sufficiently strong visual representations with representation world modeling can support effective planning without relying on complex inference-time auxiliary modules. Based on this insight, we present ReDrive, an end-to-end driving framework that strengthens planning-oriented visual features via future representation prediction. To achieve this, ReDrive adopts a three-stage training pipeline consisting of driving video pretraining, joint world-modeling and planning training, and planner adaptation. This yields a strong planning-oriented representation and a high-performance planner, while requiring neither auxiliary perception modules nor future prediction at inference time. Experiments on NAVSIM demonstrate strong performance, achieving 91.0 PDMS on NAVSIM v1 and 90.8 EPDMS on NAVSIM v2. These results show that shaping representations with world modeling is sufficient to enable high-performance end-to-end planning while retaining a simple encoder-planner inference pipeline.
Chinese Translation
驾驶策略需要场景理解和未来演化预测的能力。为了实现这一目标,当前的端到端模型通常构建复杂的感知-规划流水线,或引入显式预测未来状态的世界模型,导致系统架构复杂。受通用视觉表征可迁移性的启发,我们认为结合足够强的视觉表征与表征世界建模,可以在不依赖复杂推理时辅助模块的情况下支持有效的规划。基于这一见解,我们提出 ReDrive,一个通过未来表征预测来增强面向规划的视觉特征的端到端驾驶框架。为了实现这一点,ReDrive 采用三阶段训练流程,包括驾驶视频预训练、联合世界建模与规划训练,以及规划器自适应。这产生了强大的面向规划的表征和高性能的规划器,同时在推理时既不需要辅助感知模块,也不需要未来预测。在 NAVSIM 上的实验证明了强大的性能,在 NAVSIM v1 上达到 91.0 PDMS,在 NAVSIM v2 上达到 90.8 EPDMS。这些结果表明,利用世界建模塑造表征足以实现高性能的端到端规划,同时保持简单的编码器-规划器推理流水线。
cs.CV / 246 / 2609.33855

Program-Verified Self-Evolution for Vision-Language Models

面向视觉语言模型的程序验证自进化
Heakl, Ahmed, Choi, Sungik, Lee, Moontae, Khan, Salman
Abstract
Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\% of majority-vote labels and 18\% of model-judge labels produced during self-evolution are wrong. To address this problem, we present Verifiable QA Generation for Self-Evolving Models (VQS), which changes how the model judges answers. Instead of voting on an answer, the model parses each image into a structured record, such as a scene graph, a chart table, or a diagram graph. Fixed programs then write a question from the record and compute its answer. The model still acts as a visual checker, but it only confirms the individual facts the program reads, one short claim at a time. These claim-level checks select the parser's training targets, so the parser also improves without labels. Human raters find 94\% of VQS answers correct, against 76\% for majority voting. Across ten benchmarks, VQS improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales and outperforms the strongest self-evolving baseline at each. Gains keep growing over three training rounds, reaching 3.84 points at 2B. Code is released at https://github.com/ahmedheakl/VQS
Chinese Translation
自演化的视觉语言模型在它们从未标注图像中生成的问题上进行训练。由于这些问题没有标准答案,先前的方法通过采样答案的多数投票或由模型评判进行标注。在人工评估中,我们发现自演化过程中产生的多数投票标签中有24%和模型评判标签中有18%是错误的。为了解决这个问题,我们提出了面向自演化模型的可验证问答生成(VQS),它改变了模型评判答案的方式。模型不再对答案进行投票,而是将每张图像解析为结构化记录,例如场景图、图表表格或图形图。然后,固定程序根据该记录编写问题并计算其答案。模型仍然充当视觉检查器,但它只确认程序读取的单个事实,一次一个简短声明。这些声明级别的检查选择了解析器的训练目标,因此解析器也在没有标签的情况下得到改进。人工评分者发现94%的VQS答案是正确的,而多数投票仅为76%。在十个基准测试中,VQS在2B、4B和8B规模上将Qwen3-VL提升了最高3.18分,并在每个规模上优于最强的自演化基线。在三个训练轮次中,提升持续增长,在2B规模上达到3.84分。代码发布于 https://github.com/ahmedheakl/VQS
cs.CV / 247 / 2609.33895

Residual-Stream Burden Shapes Representation Learning in Diffusion Transformers

残差流负担塑造了扩散Transformer中的表示学习
Liang, Tongtong, Kou, Siqi, Xi, Ziqiao, Singh, Esha, Zhou, Kun, Deng, Zhijie, Cloninger, Alexander, Wang, Yu-Xiang, Parhi, Rahul
Abstract
In diffusion-based generation, a neural network can be trained to predict the clean data, the noise, or the velocity from a noisy input. These prediction targets are interconvertible and describe the same generative process, yet plain Diffusion Transformers operating on large pixel patches succeed with clean prediction and fail with noise or velocity prediction. We argue that this asymmetry arises because noisy targets require the residual stream to preserve noise-dependent input variation through depth for the final readout, forcing subsequent layers to compute on noisy representations. A spectrally concentrated clean target imposes a lighter demand, leaving greater freedom to organize hidden representations for subsequent computation. We call this preservation requirement *residual-stream burden* and show how it shapes representation learning in Diffusion Transformers. Controlled experiments indicate that the exploitable structure is spectral concentration in patch space and that the bandwidth of the persistent residual state is a key resource for noisy prediction. We further show that this account is consistent with recent decoupled pixel-space architectures, whose diverse designs all reduce the residual-stream burden on the main pathway. To examine this understanding from a complementary direction, we expand and reorganize the residual-stream bandwidth directly, introducing Spatially Indexed Hyper-Connections (SiHC) that reach FID 1.71 on ImageNet $256^2$. Together, these results identify residual-stream burden as a mechanism through which prediction targets and architecture jointly shape representation learning in Diffusion Transformers.
Chinese Translation
在基于扩散的生成中,可以训练神经网络从含噪输入中预测干净数据、噪声或速度。这些预测目标可相互转换,描述相同的生成过程,然而,在大像素块上操作的普通扩散Transformer在干净预测上成功,而在噪声或速度预测上失败。我们认为这种不对称性出现的原因是,含噪目标要求残差流在深度方向上保留依赖于噪声的输入变化以用于最终读出,迫使后续层在含噪表示上进行计算。而频谱集中的干净目标施加了较轻的需求,为后续计算组织隐藏表示留下了更大的自由度。我们将这种保留要求称为*残差流负担*,并展示它如何塑造扩散Transformer中的表示学习。控制实验表明,可利用的结构是块空间中的频谱集中,并且持久残差状态的带宽是含噪预测的关键资源。我们进一步表明,这一解释与最近解耦的像素空间架构一致,这些架构的多样化设计都减少了主路径上的残差流负担。为了从互补方向检验这一理解,我们直接扩展和重组残差流带宽,引入了空间索引超连接(SiHC),在ImageNet $256^2$上达到了FID 1.71。总之,这些结果将残差流负担识别为一种机制,通过该机制,预测目标和架构共同塑造了扩散Transformer中的表示学习。
cs.CV / 248 / 2609.33928

Preserving DEG Rankings for Gene Discovery in Histology-Based Spatial Gene Expression Prediction

在基于组织学的空间基因表达预测中,为基因发现保留DEG排序
Shiku, Kaito, Nishimura, Kazuya, Kojima, Yasuhiro, Bise, Ryoma
Abstract
Predicting spatial gene expression from histology images could scale spatial transcriptomics (ST) to image-only cohorts, but conventional histology-based ST prediction is trained and evaluated mainly by per-gene spatial-profile reconstruction. This objective is misaligned with a key downstream use of ST: differentially expressed gene (DEG) discovery, where genes are ranked for a biological or morphology-defined contrast by evidence of between-group expression differences. We formulate image-based differential expression ranking (IDER), which asks whether predicted expression profiles preserve the contrast-specific ranked gene list obtained from measured profiles. IDER compares gene rankings induced by differential-expression statistics, rather than raw expression magnitudes or per-gene spatial correlations. We further introduce a differentiable IDER objective that aligns these statistics across genes and can be trained with morphology-derived proxy contrasts without predefined biological group labels. Experiments on public ST datasets show improved DEG-ranking agreement and pathway-enrichment overlap over conventional reconstruction objectives, including morphology-derived and pathologist-annotated tissue-region evaluations.
Chinese Translation
从组织学图像预测空间基因表达可以将空间转录组学(ST)扩展到仅有图像的队列,但传统的基于组织学的ST预测主要通过每个基因的空间分布重建进行训练和评估。这一目标与ST的一个关键下游应用不一致:差异表达基因(DEG)发现,其中基因根据组间表达差异的证据,针对生物学或形态学定义的对比进行排序。我们提出了基于图像的差异表达排序(IDER),它询问预测的表达谱是否保留了从测量谱中获得的针对特定对比的排序基因列表。IDER比较由差异表达统计量诱导的基因排序,而不是原始表达量或每个基因的空间相关性。我们进一步引入了一个可微的IDER目标,该目标使这些统计量在不同基因间对齐,并且可以在没有预定义生物组标签的情况下,使用形态学衍生的代理对比进行训练。在公共ST数据集上的实验表明,与传统的重建目标相比,DEG排序一致性和通路富集重叠得到改善,包括形态学衍生和病理学家注释的组织区域评估。
cs.CV / 249 / 2609.33935

Video, Ergo Genero: Unifying Video Tasks via Spatiotemporal Analogy

视频,故我生成:通过时空类比统一视频任务
Kao, Chia-Hsiang, Zeng, Belinda, Hariharan, Bharath, Jia, Menglin
Abstract
Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse task taxonomy with a strict train-test split, ViGeo generalizes to unseen video manipulations and zero-shot modalities (e.g., event cameras). Finally, we identify task internalization, where a query format associated with a pretrained task overrides the demonstration, and show that this shortcut can be removed with a small amount of task-unrelated data, highlighting the need to decorrelate prompt format from task identity.
Chinese Translation
将视频模型适应新任务通常需要专门的数据整理和微调。虽然视觉类比通过上下文指定任务提供了一种免训练替代方案,但它仍局限于图像领域。为了探索基于类比的方法能否统一多样化的视频任务并泛化到分布外场景,我们引入了 ViGeo,一个通过时空画布补全将视觉上下文学习扩展到视频领域的框架。在具有严格训练-测试划分的多样化任务分类体系上进行评估,ViGeo 能够泛化到未见过的视频操作和零样本模态(例如事件相机)。最后,我们识别出任务内化现象,即与预训练任务相关的查询格式覆盖了演示,并表明这种捷径可以通过少量任务无关数据来消除,这凸显了将提示格式与任务身份去相关的必要性。
cs.CV / 250 / 2609.33937

Test-Time Generalized Category Discovery

测试时广义类别发现
Mishra, Shambhavi, Chakraborty, Omprakash, Silva-Rodriguez, Julio, Ayed, Ismail Ben, Pedersoli, Marco, Dolz, Jose
Abstract
Test-Time Adaptation (TTA) and Generalized Category Discovery (GCD) are traditionally treated as disjoint problems: the former adapts models to domain shift assuming all test classes are known, while the latter discovers novel categories assuming labeled training data for known classes. However, real-world deployment rarely fits either setting. Motivated by this gap, we introduce Test-Time Generalized Category Discovery (TT-GCD), a unified and more realistic scenario where a vision-language model must adapt to distribution shifts, classify known categories using only textual supervision, and discover novel categories, all during test time and without access to labeled data. To address this challenging scenario, we propose PACT (Prototype Assignment for Category discovery at Test time), a fully unsupervised framework that casts known-class recognition and novel-class discovery via prototype assignment. PACT first re-aligns shifted visual features with the text-derived class representations of the VLM using confident zero-shot predictions. Known and novel categories are then both represented by prototypes in the visual embedding space, estimated from the unlabeled test stream, and each test image is assigned to the category whose prototype is most similar to its visual feature. Extensive experiments across corruption and domain-shift benchmarks demonstrate that PACT outperforms adapted state-of-the-art TTA and GCD methods, effectively bridging the gap between adaptation and discovery.
Chinese Translation
测试时自适应(TTA)和广义类别发现(GCD)传统上被视为两个独立的问题:前者假设所有测试类别已知,使模型适应域偏移;后者假设已知类别的标注训练数据可用,从而发现新类别。然而,现实世界的部署很少符合这两种设定。受此差距的启发,我们引入了测试时广义类别发现(TT-GCD),这是一个统一且更现实的场景,其中视觉-语言模型(VLM)必须在测试时适应分布偏移,仅使用文本监督对已知类别进行分类,并发现新类别,且无法访问标注数据。为了解决这一挑战性场景,我们提出了PACT(测试时类别发现的原型分配),一个完全无监督的框架,通过原型分配实现已知类识别和新类发现。PACT首先使用置信的零样本预测,将偏移的视觉特征与VLM的文本派生类别表示重新对齐。然后,已知和新类别都在视觉嵌入空间中由原型表示,这些原型从未标注的测试流中估计,每个测试图像被分配给其视觉特征最相似的原型所属的类别。在损坏和域偏移基准上的大量实验表明,PACT优于经过调整的最先进的TTA和GCD方法,有效弥合了自适应与发现之间的差距。
cs.CV / 251 / 2609.33969

Gaussian Splatting-based Volumetric Video Compression with Sparse 4D Anchors

基于高斯泼溅的稀疏4D锚点体积视频压缩
Gao, Ge, Teng, Siyue, Wang, Chanqgi, Zhang, Fan, Anantrasirichai, Nantheera, Chiang, Jui Chiu, Peng, Wen-Hsiao, Bull, David
Abstract
Immersive video communication requires photorealistic, render-efficient, and compact dynamic scene representations. 3D Gaussian Splatting (3DGS) offers a promising representation, but dynamic 3DGS remains difficult to compress due to dense primitives and spatiotemporal redundancy. Anchor-based formulations improve compactness with sparse scaffolds that share geometry and appearance across primitives. However, existing designs often rely on deforming a single canonical scaffold and condition each primitive on its associated anchor in isolation, limiting their ability to handle non-local dynamics and disocclusion while under-exploiting inter-anchor correlations, particularly in motion- or texture-dense regions. To address these limitations, we propose SAGA, a volumetric video codec built upon Sparse Anchor-assisted GAussian splatting representations. SAGA represents dynamic 3D scenes using hierarchically organized sparse 4D anchors, where coordinate-based INR decoders generate fine anchors and Gaussian primitives from inter-anchor interpolations, enabling compact parameter sharing across spatiotemporal structures. For long-range dependencies among unstructured anchors, we further introduce fixed-size memory slots with orthogonality-informed updates for accurate entropy-context modeling. Experiments show that SAGA achieves strong rate-distortion performance against GIFStream, with PSNR BD-rate reductions of 80.39% and 83.94% on Neu3D and MPEG MIV, respectively.
Chinese Translation
沉浸式视频通信需要逼真、渲染高效且紧凑的动态场景表示。3D高斯泼溅(3D Gaussian Splatting, 3DGS)提供了一种有前景的表示方法,但由于密集的基元和时空冗余,动态3DGS仍然难以压缩。基于锚点的方案通过稀疏脚手架提高了紧凑性,这些脚手架在基元之间共享几何和外观。然而,现有设计通常依赖于对单个规范脚手架进行变形,并孤立地使每个基元以其关联的锚点为条件,这限制了它们处理非局部动态和去遮挡的能力,同时未能充分利用锚点间的相关性,尤其是在运动或纹理密集的区域。为了解决这些限制,我们提出了SAGA,一种基于稀疏锚点辅助高斯泼溅表示的体积视频编解码器。SAGA使用分层组织的稀疏4D锚点表示动态3D场景,其中基于坐标的INR解码器通过锚点间插值生成精细锚点和高斯基元,从而实现跨时空结构的紧凑参数共享。对于非结构化锚点之间的长距离依赖,我们进一步引入了固定大小的记忆槽,并采用正交性信息更新以实现准确的熵上下文建模。实验表明,SAGA在对抗GIFStream时实现了强大的率失真性能,在Neu3D和MPEG MIV上分别实现了80.39%和83.94%的PSNR BD-rate降低。
cs.CV / 252 / 2609.33991

A Multi-Dataset Benchmark of YOLO-Based Weed Detection in Precision Agriculture

精准农业中基于YOLO的杂草检测多数据集基准
Zdraveska, Hristina, Spasev, Vlatko, Dimitrovski, Ivica, Kitanovski, Ivan, Lameski, Petre
Abstract
Weed detection is an important component of precision agriculture, enabling site-specific weed management and reducing unnecessary herbicide use. Although deep learning methods have achieved strong results for crop and weed detection, many studies rely on single-dataset evaluation, making it difficult to assess robustness across different agricultural domains. This paper presents a multi-dataset benchmark of deep object detectors for weed detection in precision agriculture, with a focused evaluation of YOLO26 models. We evaluate nano, small, and medium variants on seven public weed-detection datasets covering different crops, weed species, field conditions, acquisition setups, and annotation protocols. The models are compared in terms of detection accuracy, model complexity, inference latency, FPS, and model size. In addition to in-dataset evaluation, we investigate cross-domain generalization using a unified one-class weed setup and evaluate multi-source training using the combined training subsets from all datasets. The results show that YOLO26 achieves strong in-dataset performance, with YOLO26m obtaining the highest average accuracy and YOLO26s providing the best practical accuracy-efficiency trade-off. However, cross-domain performance decreases substantially, with YOLO26s dropping from an average in-domain mAP$_{50:95}$ of 0.603 to 0.148 in the off-domain setting. Multi-source training improves performance on several datasets, but does not fully eliminate domain shift. Overall, the benchmark highlights the importance of dataset diversity, domain similarity, and target-domain adaptation for robust weed detection in real-world precision agriculture applications.
Chinese Translation
杂草检测是精准农业的重要组成部分,可实现特定地点杂草管理并减少不必要的除草剂使用。尽管深度学习方法在作物和杂草检测方面取得了显著成果,但许多研究依赖单数据集评估,难以评估其在不同农业领域中的鲁棒性。本文提出了精准农业中用于杂草检测的深度目标检测器的多数据集基准,并重点评估了 YOLO26 模型。我们在七个公开杂草检测数据集上评估了 nano、small 和 medium 变体,这些数据集涵盖不同作物、杂草物种、田间条件、采集设置和标注协议。模型比较指标包括检测精度、模型复杂度、推理延迟、FPS 和模型大小。除数据集内评估外,我们还使用统一的单类别杂草设置研究跨域泛化,并利用所有数据集的组合训练子集评估多源训练。结果表明,YOLO26 在数据集内取得了强劲性能,其中 YOLO26m 获得最高平均精度,YOLO26s 提供了最佳的实际精度-效率权衡。然而,跨域性能显著下降,YOLO26s 的平均域内 mAP$_{50:95}$ 从 0.603 降至域外设置下的 0.148。多源训练提升了若干数据集上的性能,但并未完全消除域偏移。总体而言,该基准凸显了数据集多样性、域相似性和目标域自适应对于真实世界精准农业应用中鲁棒杂草检测的重要性。
cs.CV / 253 / 2609.33996

UnfoldCRF: Structured Mask Refinement with Image-Conditioned Latent Regions

UnfoldCRF:基于图像条件潜在区域的结构化掩码细化
He, Chunming, Zhang, Rihan, Xu, Lei, Qin, Guanyi, Fang, Chengyu, Tang, Longxiang, Xiao, Fengyang, Farsiu, Sina
Abstract
Learned mask refiners improve segmentation accuracy, but it is hard to tell how much of the improvement comes from explicit structure rather than from extra capacity, and whether it holds up when the mask generator or its error distribution changes. UnfoldCRF treats refinement as inference in a conditional random field over pixel labels and latent region variables. Its energy has a corrected unary term, learned local pairwise interactions, and image-conditioned latent-region consistency, with a null state that lets a region with weak label agreement withdraw from the consistency term; inference unrolls damped mean-field updates on this one energy. To isolate the effect of structure, we compare against recurrent black-box refiners that read the same inputs and receive the same parameter budget, stage count, and supervision. On COD10K, UnfoldCRF beats the strongest matched control by 1.0 $F^\omega_\beta$ point, improves all four COD metrics, and lowers the fraction of images made worse from 11.7\% to 8.5\%. Under a train-once protocol over five datasets and several mask sources, the 2.6M-parameter variant gains 4.2 mean $\Delta$IoU against 2.0 for its control, and a variant built on frozen DINOv2 features matches the strongest foundation-model refiner with about a seventh of its resident parameters while staying ahead of its own control. On mask generators never seen in training, the gain is 2.0 $F^\omega_\beta$ points against 0.6 for the control. Zeroing individual messages shows where the corrections come from: the pairwise messages mostly fix boundaries, the region messages mostly fix non-boundary errors. Code and supporting materials will be publicly released.
Chinese Translation
学习到的掩码细化器提高了分割精度,但很难判断其中有多少改进来自显式结构而非额外容量,以及当掩码生成器或其误差分布变化时,这种改进是否仍然成立。UnfoldCRF 将细化视为对像素标签和潜在区域变量的条件随机场中的推断。其能量函数包含一个校正的一元项、学习到的局部成对交互以及图像条件潜在区域一致性,并带有一个空状态,使得标签一致性较弱的区域可以从一致性项中退出;推断在该单一能量上展开阻尼平均场更新。为了分离结构的影响,我们与循环黑盒细化器进行比较,这些细化器读取相同的输入,并获得相同的参数预算、阶段数和监督。在 COD10K 上,UnfoldCRF 比最强的匹配对照高出 1.0 个 F^ω_β 点,改善了所有四个 COD 指标,并将变差的图像比例从 11.7% 降低到 8.5%。在五个数据集和多个掩码源上的一次训练协议下,2.6M 参数变体比其对照获得 4.2 的平均 ΔIoU 增益,而对照为 2.0;基于冻结 DINOv2 特征构建的变体以约其常驻参数七分之一的参数量与最强基础模型细化器匹敌,同时保持领先于其自身对照。在训练中未见过的掩码生成器上,增益为 2.0 个 F^ω_β 点,而对照为 0.6。将单个消息置零可以显示修正的来源:成对消息主要修复边界,区域消息主要修复非边界误差。代码和支持材料将公开发布。
cs.CV / 254 / 2609.33998

MetaSampling: Making Frame Samplers Efficient for Long-Video Question Answering

MetaSampling:使帧采样器高效用于长视频问答
Dahal, Ashim, Banerjee, Bikramjit
Abstract
Frame selection is an important component of long-video question answering (VQA) with Multimodal Large Language Models (MLLMs). Existing frame-selection methods improve over simple top-$k$ embedding retrieval and uniform sampling, but are typically applied under a fixed global selection budget. We introduce \textbf{MetaSampling}, a training-free, plug-and-play sampling strategy that can be applied on top of existing frame selectors. MetaSampling improves downstream VQA efficiency by dynamically reducing the number of frames passed to the MLLM while preserving, and in some cases improving, answer accuracy. We evaluate MetaSampling across 36 paired frame-selector--MLLM-backbone--VQA-benchmark configurations. MetaSampling reduces the number of selected frames in all 36 configurations and improves accuracy in 25 of them, yielding an average frame reduction of $8.9\%$ while slightly improving accuracy overall.
Chinese Translation
帧选择是基于多模态大语言模型(MLLMs)的长视频问答(VQA)的重要组成部分。现有的帧选择方法相较于简单的 top-k 嵌入检索和均匀采样有所改进,但通常是在固定的全局选择预算下应用。我们提出了 MetaSampling,一种无需训练、即插即用的采样策略,可以应用于现有帧选择器之上。MetaSampling 通过动态减少传递给 MLLM 的帧数来提高下游 VQA 效率,同时保持(有时甚至提高)答案准确率。我们在 36 个成对的帧选择器--MLLM 主干--VQA 基准测试配置上评估了 MetaSampling。MetaSampling 在所有 36 种配置中减少了所选帧的数量,并在其中 25 种配置中提高了准确率,平均减少 8.9% 的帧数,同时整体略微提升准确率。
cs.CV / 255 / 2609.34021

Position Aware Layer Queries for Test Time Training in Vision Language Models

视觉语言模型测试时训练的位置感知层查询
Modi, Rajat, Pathak, Priyank, Liang, Xin, Rawat, Yogesh Singh
Abstract
Test-Time Training (TTT) adapts models to incoming test samples (e.g. out-of-distribution, (OOD)) when conventional fine-tuning is infeasible. Existing TTT methods for Vision-Language Models (VLMs) create supervision from several augmented views, each requiring forward (and often backward) passes through the entire VLM, incurring substantial computational cost. We observe that one forward pass with all the intermediate layer outputs already yields far more signal than the final embedding from all augmentations. We introduce Layer Query Network (LQN), a lightweight approach that can adapt a frozen VLM (teacher) in a single forward pass of the VLM via a small model (student). LQN uses Position-Aware Distillation (PAD) to mimic the teacher VLM's intermediate-layer spatial tokens by querying spatial coordinates of intermediate tokens. LQN additionally relies on Location Consistency Regularization (LCR), a self-supervision technique, replacing expensive O(H x W) image augmentation with O(1) coordinate sampling. Integrating these, LQN i) adapts and improves zero-shot CLIP ViT-B/16 by 9.8% Top-1 on OOD ImageNet, ii) outperforms the previous best GS-Bias on fine-grained classification by 3.9% Top-1, iii) achieves faster convergence than TPS for CLIP ResNet-50 (47 mins vs 55 mins), iv) generalizes adaptation to VLMs like SigLIP, EVA-CLIP, and CoCa, and lightweight students like MLP, ResNet, VGG, and v) extends to panoptic, instance, and semantic segmentation.
Chinese Translation
测试时训练(TTT)在传统微调不可行时,使模型适应传入的测试样本(例如,分布外(OOD)样本)。现有的面向视觉语言模型(VLM)的 TTT 方法从多个增强视图中创建监督信号,每个视图都需要对整个 VLM 进行前向(通常还有反向)传播,导致大量的计算成本。我们观察到,一次前向传播并利用所有中间层输出,已经比所有增强视图的最终嵌入产生了更多的信号。我们引入了层查询网络(LQN),一种轻量级方法,可以通过一个小模型(学生)在 VLM 的单次前向传播中适应一个冻结的 VLM(教师)。LQN 使用位置感知蒸馏(PAD)通过查询中间 token 的空间坐标来模仿教师 VLM 的中间层空间 token。LQN 还依赖于位置一致性正则化(LCR),一种自监督技术,用 O(1) 坐标采样取代昂贵的 O(H x W) 图像增强。整合这些,LQN i) 在 OOD ImageNet 上将零样本 CLIP ViT-B/16 的 Top-1 准确率适应并提升了 9.8%,ii) 在细粒度分类上以 3.9% 的 Top-1 优于之前最佳的 GS-Bias,iii) 对于 CLIP ResNet-50,收敛速度比 TPS 更快(47 分钟 vs 55 分钟),iv) 将适应泛化到 SigLIP、EVA-CLIP 和 CoCa 等 VLM,以及 MLP、ResNet、VGG 等轻量级学生模型,v) 扩展到全景、实例和语义分割。
cs.CV / 256 / 2609.34032

Re:Cognize -- Open-Set Comic Character Re-Identification

Re:Cognize -- 开放集漫画角色重识别
Baranwal, Aaditya, Kataria, Madhav, Rawat, Yogesh S, Vyas, Shruti
Abstract
A manga reader meets a character on one page and knows them on sight a hundred pages later, without ever being handed a cast list. Re-identifying comic characters demands the same, open-set and sequential: pages arrive as a stream in reading order, new faces appear before anyone names them, and the cast is assembled as the story is read. $\textbf{Re:Cognize}$ evaluates recognition as the story is read, not against a cast handed over in advance: four protocols on one query stream, from closed-set retrieval to a cast the model must build and grow itself. The surprise is where models fail. Recognising is close to solved: one reference image per character already ranks as well as a gallery built in advance. Knowing what to believe is not: a model that adds its own matches makes its cast worse, while the same growth with correct labels would gain over twenty points of top-1 accuracy. The bottleneck is acceptance, not vision, and one comparison decides it: an addition pays exactly when it is right more often than the cast already was on the queries it takes over. The comparison has nothing to fit, and measured on half of a new corpus it calls the other half correctly. $\textbf{ReCast}$ puts it to work with nothing fitted on data: a cast sheet of one running average per character, grown only where the page itself vouches for a crop. It recovers a third to two thirds of what perfect labels would, depending on whether the cast starts from random examples or from first appearances. Re:Cognize measures whether a model can read along; ReCast is a cast that does. Our claims are on identity maintenance, recognising characters already met; the emergence of new ones is measured as a diagnostic under a fixed reference rule, and we propose no method for it.
Chinese Translation
一个漫画读者在一页上遇到一个角色,一百页后一眼就能认出他们,而从未拿到过角色名单。漫画角色的重识别同样需要做到开放集和顺序性:页面按阅读顺序以流的形式到达,新面孔在有人给它们命名之前就出现,角色阵容随着故事被阅读而逐步组建。Re:Cognize 在故事被阅读时评估识别,而不是针对提前给定的角色阵容:在一个查询流上的四种协议,从封闭集检索到模型必须自行构建和扩展的角色阵容。令人惊讶的是模型失败的地方。识别几乎已经解决:每个角色一张参考图像已经能像提前构建的图库一样进行排序。知道该相信什么则不是:一个添加自身匹配结果的模型会让其角色阵容变得更糟,而同样的增长如果使用正确标签将获得超过二十个百分点的 top-1 准确率。瓶颈是接受,而不是视觉,而一个比较决定了它:当一次添加在其接管的查询上正确的频率高于现有角色阵容时,它才是有益的。这个比较没有任何需要拟合的参数,并且在一个新语料库的一半上测量时,它能正确判断另一半。ReCast 将其付诸实践,且没有任何在数据上拟合的参数:每个角色一个运行平均值的角色表,仅在页面本身为裁剪图像提供担保的地方增长。它能恢复完美标签所能达到的三分之一到三分之二,这取决于角色阵容是从随机样本开始还是从首次出现开始。Re:Cognize 衡量一个模型是否能跟随阅读;ReCast 是一个能做到的角色阵容。我们的主张关于身份维护,即识别已经遇到过的角色;新角色的出现作为一种诊断,在固定参考规则下测量,我们不为它提出任何方法。
cs.CV / 257 / 2609.34035

3D Point Tracking with State Space Models

基于状态空间模型的3D点追踪
Ogawa, Masahiro, An, Qi, Yamashita, Atsushi
Abstract
Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget. Our method rests on one observation: once a point's 2D image trajectory is fixed, the quantity that governs its metric accuracy is the depth along its pixel ray. Rather than learning tracking end-to-end, we therefore compose two frozen front-ends - dense optical flow for 2D correspondence and a monocular metric-depth network for the third dimension - and learn only the residual they cannot supply: that depth, refined by a compact state space model (Mamba-3) conditioned on appearance features (DINOv3). A state space model rather than the transformers the strongest 3D trackers adopt is what makes a single-GPU budget attainable: it summarises a track in a fixed-size recurrent state whose memory cost is constant in the number of frames, whereas attention requires a key-value cache that grows linearly with them. On the TAPVid-3D minival benchmark our best configuration attains the highest absolute metric accuracy among methods evaluated under identical conditions (mean metric Average Jaccard, 0.256), exceeding strong feed-forward trackers, while a companion analysis, reproduced with each competitor's own evaluator, explains why several published trackers lose most of their accuracy under this budget.
Chinese Translation
在度量三维中追踪动态场景的任意点——以绝对米为单位,而非未知尺度——支撑着3D和4D重建、机器人导航和自动驾驶,在这些应用中决策是以米而非像素为单位做出的。我们的目标是一个3D点追踪器,在这些绝对项上准确,并在单个商用GPU、无位姿、单目预算内运行。我们的方法基于一个观察:一旦一个点的2D图像轨迹固定,决定其度量精度的量是沿其像素射线的深度。因此,我们不是端到端地学习追踪,而是组合两个冻结的前端——用于2D对应的稠密光流和用于第三维的单目度量深度网络——并仅学习它们无法提供的残差:该深度,由紧凑的状态空间模型(Mamba-3)以外观特征(DINOv3)为条件进行细化。采用状态空间模型而非最强的3D追踪器所采用的Transformer,使得单GPU预算变得可行:它将轨迹总结为固定大小的循环状态,其内存成本随帧数恒定,而注意力机制需要随帧数线性增长的键值缓存。在TAPVid-3D minival基准上,我们的最佳配置在相同条件下评估的方法中达到了最高的绝对度量精度(平均度量Average Jaccard,0.256),超过了强大的前馈追踪器,同时一项使用每个竞争者自己的评估器复现的伴随分析解释了为什么几个已发表的追踪器在此预算下失去了大部分精度。
cs.CV / 258 / 2609.34044

SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

SCOPD:用于高效视觉语言模型的稀疏上下文同策略自蒸馏
Jeddi, Ahmadreza, Zhang, Enming, Gerigk, Jasper, Karaimer, Hakki, Azadani, Mozhgan Nasr, Luo, Jiayun, Le, Minh Ngoc, Aminian, Gholamali, Buurmeijer, Hugo, Chen, Yongchao, Sigal, Leonid, Gilitschenski, Igor, Derpanis, Konstantinos G., Pavone, Marco, Taati, Babak
Abstract
Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show that this explanation is incomplete. In a fixed-context Pass@K analysis, repeated sampling from the same pruned visual representation recovers many examples missed by greedy decoding, indicating that useful visual evidence can remain accessible but be used unreliably. We call this the representation-utilization gap. Motivated by this observation, we introduce SCOPD, a sparse-context on-policy self-distillation framework in which a student generates reasoning trajectories from pruned visual tokens while a privileged full-context teacher supervises the same on-policy prefixes. SCOPD requires no ground-truth responses, architectural changes, or additional inference-time computation. We further introduce SCOPD+, which uses a small visual-budget intervention to identify visually sensitive response positions and selectively distill them. At 10% visual-token retention, the Vanilla model retains 86.37% of its unpruned performance across 13 benchmarks. SCOPD raises this to 90.49%, while SCOPD+ further improves it to 92.43%. Across token budgets, benchmarks, and pruning operators, our results show that efficient reasoning depends not only on which visual information survives pruning, but also on how reliably the model learns to use it.
Chinese Translation
推理视觉语言模型(VLMs)将图像和视频处理为长序列的视觉标记,使得推理代价高昂。免训练的标记剪枝降低了这一成本,但激进的压缩会急剧降低性能,通常归因于任务相关视觉信息的不可逆丢失。我们表明这一解释并不完整。在固定上下文的Pass@K分析中,从相同的剪枝后视觉表示中重复采样,恢复了贪婪解码所遗漏的许多样本,表明有用的视觉证据仍然可访问,但被不可靠地使用。我们将此称为表示-利用差距。受此观察启发,我们提出SCOPD,一种稀疏上下文同策略自蒸馏框架,其中学生从剪枝后的视觉标记生成推理轨迹,而特权全上下文教师监督相同的同策略前缀。SCOPD不需要真实响应、架构更改或额外的推理时计算。我们进一步引入SCOPD+,它使用小幅视觉预算干预来识别视觉敏感的响应位置,并选择性地蒸馏它们。在10%的视觉标记保留率下,Vanilla模型在13个基准上保留了其未剪枝性能的86.37%。SCOPD将其提升至90.49%,而SCOPD+进一步将其提升至92.43%。在不同的标记预算、基准和剪枝算子中,我们的结果表明,高效推理不仅取决于哪些视觉信息在剪枝后存活,还取决于模型学习使用这些信息的可靠性。
cs.CV / 259 / 2609.34047

ARCH-B: Architectural Representation, Comprehension and Hierarchy Benchmark

ARCH-B:建筑表征、理解与层级基准
Parikh, Kieran Sagar, Lopez, Jose Luis Garcia del Castillo y
Abstract
Multimodal models increasingly interpret visual environments, but their ability to recognize the same building across photographs, floor plans, elevations, sections, and renderings remains poorly characterized. We introduce ARCH-B, a benchmark of 354 four-choice questions across 11 cross-representational archetypes, constructed from a building-linked corpus of 3.9 million architectural images using visually similar distractors, model-guided difficulty screening, and manual validation. We evaluate 25 multimodal models and collect 5,830 responses from non-expert human participants. Model accuracy ranges from 10.45% to 83.90%, compared with a human baseline of 35.35%. Models perform comparatively well on mixed-representation outlier detection and photograph matching, but remain weaker on floorplan-to-photograph correspondence. Human and model difficulty across archetypes is only weakly correlated (Spearman's (\rho=0.33)). Held-out evaluation confirms that the difficulty identified during screening generalizes beyond the curation models. ARCH-B provides a diagnostic evaluation of visual correspondence and representation transfer across architectural media.
Chinese Translation
多模态模型越来越多地用于解读视觉环境,但它们在照片、平面图、立面图、剖面图和效果图之间识别同一建筑的能力仍缺乏充分刻画。我们提出了 ARCH-B,一个包含 354 道四选一问题的基准,涵盖 11 种跨表征原型;该基准基于一个与建筑关联的 390 万张建筑图像语料库构建,并采用视觉相似的干扰项、模型引导的难度筛选和人工验证。我们评估了 25 个多模态模型,并从非专家人类参与者处收集了 5,830 份回答。模型准确率范围为 10.45% 至 83.90%,而人类基线为 35.35%。模型在混合表征异常值检测和照片匹配上表现相对较好,但在平面图到照片的对应关系上仍然较弱。人类与模型在各原型上的难度仅呈弱相关(Spearman's ρ=0.33)。留出评估证实,筛选过程中识别出的难度能够泛化到策展模型之外。ARCH-B 为跨建筑媒介的视觉对应与表征迁移提供了一种诊断性评估。
cs.CV / 260 / 2609.34071

SNaP: One-Step Posterior Sampling for Noisy Inverse Problems

SNaP:面向含噪逆问题的一步后验采样
Shoushtari, Shirin, Chandler, Edward P., Shi, Xiao, Kamilov, Ulugbek S.
Abstract
Diffusion and flow-matching models can produce high-quality posterior samples for inverse problems, but typically require tens to thousands of network evaluations per draw. MeanFlow enables one-step generation, yet applying it to inverse problems leaves no intermediate steps at which to enforce measurement consistency. We introduce SNaP, a one-step MeanFlow posterior sampler for linear inverse problems with Gaussian noise. Its central innovation is a measurement-adapted source: a Gaussian distribution whose mean and anisotropic covariance are determined by the measurement operator, observation, and noise level. The source anchors well-measured directions while preserving variation where the measurements are weak or uninformative. We show that the exact conditional flow transports this source to the true posterior. Across natural-image restoration and multi-coil MRI, SNaP produces diverse, high-quality samples with one network evaluation per draw, 30 to 2250 $\times$ faster than iterative samplers.
Chinese Translation
扩散模型和流匹配模型能够为逆问题生成高质量的后验样本,但每次采样通常需要数十到数千次网络评估。MeanFlow 实现了一步生成,但将其应用于逆问题时,没有中间步骤可用于强制测量一致性。我们提出 SNaP,一种针对高斯噪声线性逆问题的一步 MeanFlow 后验采样器。其核心创新是测量自适应的源分布:一个高斯分布,其均值和各向异性协方差由测量算子、观测值和噪声水平确定。该源分布在测量良好的方向上锚定,同时在测量较弱或无信息的方向上保留变化。我们证明,精确的条件流将该源分布输运至真实后验。在自然图像恢复和多线圈 MRI 上,SNaP 每次采样仅需一次网络评估即可生成多样、高质量样本,比迭代采样器快 30 到 2250 倍。
cs.CV / 261 / 2609.34078

WhiteCon: Semi-Supervised Domain Adaptation Regression Through Whitening Transform and Dual Consistency

WhiteCon:通过白化变换与双重一致性的半监督域适应回归
Sim, Se Jin, Kim, Seoung Bum
Abstract
Domain adaptation is crucial for addressing distributional shifts that degrade model performance across domains. While most existing research has centered on classification, semi-supervised domain adaptation regression (SSDAR) for continuous-output tasks remains largely unexplored, particularly in practical scenarios with limited labeled target data. To address this gap, we propose semi-supervised domain adaptation regression through whitening transform and dual consistency (WhiteCon), which combines domain-specific whitening transform (DWT) and dual consistency regularization to enhance training stability and domain adaptation. DWT reduces the variance of the model parameters by transforming the feature covariance matrix into an identity matrix, thus stabilizing training under ordinary least squares assumptions. In addition, variance consistency regularization, as part of dual consistency regularization, aligns the variances of weak, strong, and mixup-augmented features to improve resilience against augmentation-induced perturbations. Empirical evaluations on various benchmark datasets under SSDAR settings demonstrate that the proposed WhiteCon achieves state-of-the-art performance compared to existing methods, effectively addressing domain shifts in regression tasks. The code for WhiteCon is available at https://github.com/sejin-sim/WhiteCon.
Chinese Translation
域适应对于解决导致模型跨域性能退化的分布偏移至关重要。尽管现有研究大多集中于分类,面向连续输出任务的半监督域适应回归(SSDAR)仍很大程度上未被探索,尤其是在目标域标注数据有限的现实场景中。为填补这一空白,我们提出通过白化变换与双重一致性的半监督域适应回归(WhiteCon),它结合了域特定白化变换(DWT)和双重一致性正则化,以增强训练稳定性和域适应。DWT通过将特征协方差矩阵变换为单位矩阵来降低模型参数的方差,从而在普通最小二乘假设下稳定训练。此外,作为双重一致性正则化的一部分,方差一致性正则化对齐弱增强、强增强和mixup增强特征的方差,以提高对增强引起的扰动的鲁棒性。在SSDAR设置下对多个基准数据集的实证评估表明,与现有方法相比,所提出的WhiteCon实现了最先进的性能,有效解决了回归任务中的域偏移。WhiteCon的代码可在 https://github.com/sejin-sim/WhiteCon 获取。
cs.CV / 262 / 2609.34094

Advancing Wildlife Conservation through Multimodal Animal Re-Identification with Environmental Metadata

通过结合环境元数据的多模态动物重识别推进野生动物保护
Li, Yuzhuo, Zhao, Di, Qiao, Tingrui, Wu, Yihao, Pang, Bo, Koh, Yun Sing
Abstract
Identifying individual animals is crucial for effective wildlife monitoring and conservation efforts. Recent advancements in computer vision have shown promise in animal re-identification (Animal ReID) by leveraging data from camera traps. However, existing Animal ReID datasets rely exclusively on visual data, overlooking environmental metadata that ecologists have identified as highly correlated with animal behavior and identity, such as temperature and circadian rhythms. Meanwhile, modern vision-language models (VLMs) offer rich multimodal reasoning capabilities, but existing resources underutilize their text-processing potential. To address these limitations, we propose MetaWild, a multimodal Animal ReID dataset comprising 20,890 images across six species, paired with environmental metadata extracted from embedded camera trap overlays and scene contexts. Additionally, to facilitate the use of metadata in existing ReID methods, we propose the Meta-Feature Adapter (MFA), a lightweight module that can be incorporated into existing VLM-based ReID methods, allowing ReID models to leverage both environmental metadata and visual information to improve ReID performance. Experiments on MetaWild show that combining baseline ReID models with MFA to incorporate metadata consistently improves performance compared to using visual information alone, validating the effectiveness of incorporating metadata in re-identification.
Chinese Translation
识别个体动物对于有效的野生动物监测和保护工作至关重要。近期计算机视觉的进展通过利用来自相机陷阱(camera traps)的数据,在动物重识别(Animal ReID)方面显示出潜力。然而,现有的动物重识别数据集仅依赖于视觉数据,忽视了生态学家已发现与动物行为和身份高度相关的环境元数据,如温度和昼夜节律。与此同时,现代视觉语言模型(VLMs)提供了丰富的多模态推理能力,但现有资源未能充分利用其文本处理潜力。为了解决这些局限性,我们提出了 MetaWild,一个多模态动物重识别数据集,包含来自六个物种的20,890张图像,并配有从嵌入式相机陷阱叠加信息和场景上下文中提取的环境元数据。此外,为了促进元数据在现有重识别方法中的使用,我们提出了 Meta-Feature Adapter(MFA),一个轻量级模块,可以集成到现有的基于VLM的重识别方法中,使重识别模型能够同时利用环境元数据和视觉信息来提高重识别性能。在 MetaWild 上的实验表明,将基线重识别模型与 MFA 结合以纳入元数据,与仅使用视觉信息相比,性能持续提升,验证了在重识别中纳入元数据的有效性。
cs.CV / 263 / 2609.34106

The Devil is in the Spectrum Bias: Spectrum-Balanced Feature Matching for Robust Representation Distillation

魔鬼藏在频谱偏差中:用于鲁棒表示蒸馏的频谱平衡特征匹配
Saito, Kuniaki, Ushiku, Yoshitaka
Abstract
Large visual foundation models have demonstrated remarkable transferability across a wide range of downstream tasks. To deploy such models efficiently, feature matching has become a popular knowledge distillation approach that transfers teacher representations to smaller student models without requiring labeled data. However, we show that the conventional feature matching objective with L2-distance is inherently biased toward reconstructing dominant spectral directions of the teacher representation, while under-optimizing low-variance directions that often contain task-relevant information. To address this, we propose Spectrum-Balanced Feature Matching, SpecMatch, a simple objective that adaptively emphasizes under-optimized spectral directions while preserving the relative importance of dominant directions. SpecMatch is easy to implement and introduces negligible computational overhead. Extensive experiments on image recognition demonstrate that SpecMatch consistently improves downstream adaptation across diverse tasks, including image classification, anomaly detection, medical image analysis, and domain generalization. In particular, SpecMatch outperforms conventional feature matching in 40 of 42 teacher--student and training-setting combinations, while consistently improving over the original student model in all settings. We further demonstrate that the proposed objective generalizes beyond vision, improving downstream performance across six protein understanding tasks.
Chinese Translation
大型视觉基础模型已在广泛的下游任务中展现出显著的迁移能力。为高效部署此类模型,特征匹配已成为一种流行的知识蒸馏方法,能够在无需标注数据的情况下将教师表示迁移到更小的学生模型。然而,我们发现,采用L2距离的传统特征匹配目标固有地偏向于重建教师表示的主导频谱方向,而对通常包含任务相关信息的低方差方向优化不足。为解决这一问题,我们提出了频谱平衡特征匹配(SpecMatch),这是一种简单目标,能够自适应地强调优化不足的频谱方向,同时保持主导方向的相对重要性。SpecMatch易于实现,且引入的计算开销可忽略不计。在图像识别上的大量实验表明,SpecMatch可持续提升多种任务的下游自适应性能,包括图像分类、异常检测、医学图像分析和域泛化。具体而言,SpecMatch在42种教师-学生和训练设置组合中的40种上优于传统特征匹配,同时在所有设置中持续优于原始学生模型。我们进一步证明,所提目标可泛化到视觉之外,在六个蛋白质理解任务上提升下游性能。
cs.CV / 264 / 2609.34124

SpatialSkill: Self-Evolving Skills for Cross-View Spatial Reasoning

SpatialSkill:面向跨视角空间推理的自演化技能
Zuo, Ruifan, Hu, Guocheng, Gan, Wanshui, Wang, Junyi, Lei, Xiang, Gan, Tian
Abstract
Cross-view spatial reasoning requires a model to align different viewpoints into a coherent spatial representation, yet this ability remains challenging for vision-language models despite being natural to humans. Existing methods typically improve spatial reasoning by updating model weights, which keeps the acquired knowledge implicit and tied to a specific backbone. We propose \textit{SpatialSkill}, a weight-update-free framework that enables a frozen vision-language model to accumulate explicit natural-language reasoning skills from offline trajectories. Unlike symbolic tasks, perceptual skills cannot be reliably verified simply by executing them: a plausible spatial rule may lack visual support or require transformations that the frozen model cannot perform. SpatialSkill therefore admits candidate skills only after visual-grounding and executability checks, constrains manual evolution to prevent harmful regressions, and routes skills by spatial-reasoning category to reduce negative transfer. On CityCube, across four frozen executors, SpatialSkill yields consistent gains, and a 9B executor equipped with SpatialSkill surpasses the strongest closed-source reference in our evaluation. The skills are stored in a versioned natural-language manual, making the reasoning strategies explicit and auditable without modifying model parameters. Code at https://github.com/vindahi/SpatialSkill.
Chinese Translation
跨视角空间推理要求模型将不同视角对齐为连贯的空间表示,然而尽管这种能力对人类而言很自然,对视觉语言模型而言仍颇具挑战。现有方法通常通过更新模型权重来提升空间推理能力,这使得所获得的知识保持隐式,并与特定骨干网络绑定。我们提出 SpatialSkill,一个无需更新权重的框架,使冻结的视觉语言模型能够从离线轨迹中积累显式的自然语言推理技能。与符号任务不同,感知技能无法仅通过执行来可靠验证:一个看似合理的空间规则可能缺乏视觉支持,或需要冻结模型无法执行的变换。因此,SpatialSkill 仅在通过视觉接地与可执行性检查后才接纳候选技能,约束技能手册的演化以防止有害回退,并按空间推理类别路由技能以减少负迁移。在 CityCube 上,跨四个冻结执行器,SpatialSkill 均带来一致增益,且配备 SpatialSkill 的 9B 执行器在我们的评估中超越了最强的闭源参考模型。这些技能存储在一个版本化的自然语言手册中,使得推理策略显式且可审计,且无需修改模型参数。代码见 https://github.com/vindahi/SpatialSkill。
cs.CV / 265 / 2609.34133

PrefLUT: Reusable and Refinable Personalized Color Editing from Pairwise Preferences

PrefLUT:基于成对偏好的可重用和可细化的个性化色彩编辑
Xu, Chuanzhi, Chen, Langyi, Yue, Chengkun, Yin, Xuanhua, Wei, Boyu, Zeng, Qingwen, Deng, Zihan, Cai, Weidong
Abstract
Photographic color editing is inherently personal: the same image can appear too warm, too muted, or already satisfactory to different users. Most lookup table (LUT) and reference-guided methods target a specified appearance rather than model persistent preferences from repeated user choices. To address this gap, we introduce PrefLUT, a reusable and refinable user-preference modeling framework for deployable 3D LUTs, encoding ordered preferred/non-preferred image pairs into a lightweight Reusable User Profile that is reused across queries and refined using additional user preference pairs, without per-user optimization. A Query-Conditioned LUT Predictor combines this profile with each image to predict a LUT latent vector and edit strength. An Identity-Residual LUT Decoder and Edit-Strength Controller then produce an exportable 3D LUT. Experiments on three datasets demonstrate effective personalized editing and general-purpose enhancement. Each quantized profile requires only 260 bytes, and editing takes 1.365 ms/image on an RTX 5090 GPU. We also introduce the Preference-Conditioning Verification Protocol (PCVP), an evaluation protocol to verify whether personalized image edits depend on user preferences and the query image through controlled changes to user profiles, preference orders, pair correspondences, and query images.
Chinese Translation
摄影色彩编辑本质上是个人化的:同一张图像对不同用户可能显得过暖、过于柔和或已经令人满意。大多数查找表(LUT)和参考引导的方法针对的是指定的外观,而非从重复的用户选择中建模持久的偏好。为了填补这一空白,我们提出了PrefLUT,一个可重用且可细化的用户偏好建模框架,用于可部署的3D LUT,将有序的偏好/非偏好图像对编码为轻量级的可重用用户画像,该画像可在查询中重用,并使用额外的用户偏好对进行细化,而无需针对每个用户进行优化。一个查询条件LUT预测器将此画像与每个图像结合,以预测LUT潜在向量和编辑强度。然后,一个身份残差LUT解码器和编辑强度控制器生成可导出的3D LUT。在三个数据集上的实验证明了有效的个性化编辑和通用增强。每个量化后的画像仅需260字节,在RTX 5090 GPU上编辑耗时1.365毫秒/图像。我们还引入了偏好条件验证协议(PCVP),一种评估协议,通过控制用户画像、偏好顺序、图像对对应关系和查询图像的变化,来验证个性化图像编辑是否依赖于用户偏好和查询图像。
cs.CV / 266 / 2609.34142

Analytical and Convolutional Neural Network-Based Motion-Vector Propagation for Efficient Video Object Detection

用于高效视频目标检测的解析式和基于卷积神经网络的运动矢量传播
Majeed, Ashiyana Abdul, Meribout, Mahmoud, Joseph, Neethu
Abstract
Continuous video analytics requires accurate localization at low latency within embedded power budgets. This paper presents a hardware-software design methodology that reuses codec motion vectors (MVs) between detector invocations. Two alternative models support translation and scale changes: analytical motion-vector propagation (Analytical-MV) and learned propagation using a convolutional neural network (CNN) (CNN-MV). The learned model uses convolutional operations and independent object updates suited to parallel execution on an edge graphics processing unit (GPU). Analytical-MV combines a harmonic-mean precision-recall score (F1) of 0.909 with a mean end-to-end latency of 9.03 ms and an energy consumption of 0.177 J per frame, yielding the lowest latency and energy among the evaluated configurations. Relative to detection on every frame, it reduces mean latency by 25.9% and energy per frame by 36.4%. CNN-MV offers a different trade-off: its fastest configuration raises recall from 0.871 for Analytical-MV to 0.890 and lowers mean power from 19.64 to 17.32 W, while achieving a latency of 18.42 ms and an energy consumption of 0.319 J per frame. It is therefore useful when recall or operating power is more important than minimum latency and energy. Execution on a deep learning accelerator (DLA) further reduces time-averaged GPU utilization relative to GPU execution. Host-processing optimization substantially improves both latency and energy, demonstrating the value of jointly designing temporal models and their execution pipelines.
Chinese Translation
连续视频分析需要在嵌入式功耗预算内以低延迟进行精确定位。本文提出了一种软硬件设计方法,在检测器调用之间复用编解码器运动矢量(MVs)。两种替代模型支持平移和尺度变化:解析运动矢量传播(Analytical-MV)和使用卷积神经网络(CNN)的学习传播(CNN-MV)。学习模型使用卷积操作和独立的对象更新,适合在边缘图形处理单元(GPU)上并行执行。Analytical-MV 结合了 0.909 的调和平均精确率-召回率得分(F1)与 9.03 ms 的平均端到端延迟和每帧 0.177 J 的能耗,在所评估的配置中实现了最低的延迟和能耗。相对于每帧检测,它将平均延迟降低了 25.9%,每帧能耗降低了 36.4%。CNN-MV 提供了不同的权衡:其最快的配置将召回率从 Analytical-MV 的 0.871 提高到 0.890,并将平均功耗从 19.64 W 降低到 17.32 W,同时实现了 18.42 ms 的延迟和每帧 0.319 J 的能耗。因此,当召回率或运行功耗比最小延迟和能耗更重要时,它是有用的。在深度学习加速器(DLA)上执行进一步降低了相对于 GPU 执行的时间平均 GPU 利用率。主机处理优化显著改善了延迟和能耗,展示了联合设计时序模型及其执行管线的价值。
cs.CV / 267 / 2609.34143

Beyond Geometry: Benchmarking and Consistency Reasoning for 3D Logical Anomaly Detection

超越几何:3D逻辑异常检测的基准测试与一致性推理
Qin, Zhiqiang, Xie, He, Yi, Junfei, Yang, Yang, Wang, Hao, Cao, Yunkang, Zhang, Hui, Wang, Yaonan
Abstract
Existing 3D industrial anomaly detection mainly targets local geometric deviations. In contrast, many industrial anomalies violate object-level design or assembly rules, which we define as 3D logical anomalies. To address these challenges, we introduce the Industrial Logical Anomaly Detection Dataset (ILGAD), the first scalable benchmark dedicated to logical anomalies in industrial point clouds. ILGAD contains 2,774 samples from 15 categories with point-level annotations and covers existence, specification, pose, and assembly-state errors. To detect such 3D logical anomalies, we propose a consistency reasoning framework that assesses whether local geometry, structure coverage, and spatial relations conform to the normal design. The framework detects geometric changes, unsupported expected structures, and abnormal local arrangements. Experiments on ILGAD, Anomaly-ShapeNet, and IEC3D demonstrate superior object-level detection and point-level localization, showing that the framework effectively detects logical anomalies and generalizes to conventional geometric defects.
Chinese Translation
现有的3D工业异常检测主要针对局部几何偏差。相比之下,许多工业异常违反了对象级设计或装配规则,我们将其定义为3D逻辑异常。为了应对这些挑战,我们引入了工业逻辑异常检测数据集(ILGAD),这是首个专用于工业点云中逻辑异常的可扩展基准。ILGAD包含来自15个类别的2,774个样本,具有点级标注,涵盖存在性、规格、姿态和装配状态错误。为了检测此类3D逻辑异常,我们提出了一个一致性推理框架,用于评估局部几何、结构覆盖和空间关系是否符合正常设计。该框架检测几何变化、不支持的预期结构和异常局部排列。在ILGAD、Anomaly-ShapeNet和IEC3D上的实验证明了卓越的对象级检测和点级定位,表明该框架能有效检测逻辑异常并泛化到传统几何缺陷。
cs.CV / 268 / 2609.34144

CAST: Reconstruction-Coupled Acceleration of Interactive World Models

CAST:交互式世界模型的重建耦合加速
Chen, Leyang, Wu, Junyi, Kong, Fanqing, Zhang, Shaoqiu, Zhang, Yulun
Abstract
Interactive world models must respond quickly to controls while preserving scene consistency. Existing acceleration methods can miss heterogeneous control responses and spatial transport when recovering skipped features. We observe that interaction-induced feature changes correlate with approximation error, while low-frequency interpolation errors are phase-sensitive and show more predictable phase progression. These findings motivate CAST, a reconstruction-coupled inference framework. CAST selects anchors by interaction sensitivity and cross-layer coverage, reconstructs skipped residuals with frequency- and confidence-aware Phase-Aware Reconstruction (PAR), and coordinates historical KV routing according to downstream reconstruction responsibility. On Matrix-Game 3.0 and HY-World 1.5, CAST achieves 2.15x and 3.48x speedups, respectively, while maintaining visual quality close to Native (Figure 1). It also attains the highest VBench scores among compared methods and leads non-native baselines on seven and six of thirteen WorldMark dimensions, demonstrating a balance of generation speed, visual quality, and interactive responsiveness under real-time control. Code is available at https://github.com/lokiniuniu/CAST.
Chinese Translation
交互式世界模型必须在保持场景一致性的同时快速响应控制。现有的加速方法在恢复被跳过的特征时,可能会遗漏异构的控制响应和空间传输。我们观察到,交互引起的特征变化与近似误差相关,而低频插值误差对相位敏感,并表现出更可预测的相位进展。这些发现促使我们提出CAST,一个重建耦合的推理框架。CAST根据交互敏感性和跨层覆盖选择锚点,利用频率和置信度感知的相位感知重建(PAR)重建跳过的残差,并根据下游重建责任协调历史KV路由。在Matrix-Game 3.0和HY-World 1.5上,CAST分别实现了2.15倍和3.48倍的加速,同时保持接近Native的视觉质量(图1)。它在比较方法中获得了最高的VBench分数,并在13个WorldMark维度中的7个和6个上领先于非原生基线,展示了在实时控制下生成速度、视觉质量和交互响应性的平衡。代码可在https://github.com/lokiniuniu/CAST获取。
cs.CV / 269 / 2609.34148

Geometric Encoding for Spatial Reasoning in Vision-Language Models

视觉语言模型空间推理的几何编码
Jun, Antonio, Yu, Haoshui, Lu, Zhengyi, Fu, Huirong, Qiang, Yao
Abstract
Vision-Language Models (VLMs) are far more reliable at recognizing what appears in a video than at reasoning about its spatial and temporal properties, such as metric distances, object dimensions, and consistent object identities across frames. We present Geometric Code, a perception-to-geometry pipeline that computes explicit spatial structure from video and supplies it to VLMs as context to augment reasoning. A perception layer segments and classifies objects and recovers depth, camera pose, and intrinsics from monocular RGB video. A deterministic geometric engine then back-projects, merges, and cleans these outputs into a spatial code, including per-object positions, dimensions, counts, inter-object distances, appearance order, and room geometry. The code is serialized into VLMs' prompts, either alongside the video or replacing it entirely. Specifically, there is no component trained or fine-tuned in our approach. On VSI-Bench, augmenting 2B and 4B open models with the spatial code improves average accuracy by +4.1 points over the frames-only baseline, with the largest gains on numeric estimation tasks such as absolute distance (+24.1 points). The results suggest that explicitly computed geometry, delivered through the language channel, recovers spatial competence that small VLMs cannot extract from pixels alone.
Chinese Translation
视觉语言模型(VLMs)在识别视频中出现的内容方面,远比在推理其空间和时间属性(如度量距离、物体尺寸以及跨帧一致的物体身份)方面可靠。我们提出 Geometric Code,一种从感知到几何的流水线,它从视频中计算显式空间结构,并将其作为上下文提供给 VLMs 以增强推理。一个感知层对物体进行分割和分类,并从单目 RGB 视频中恢复深度、相机位姿和内参。一个确定性几何引擎随后对这些输出进行反投影、合并和清理,生成一个空间编码,包括每个物体的位置、尺寸、数量、物体间距离、出现顺序和房间几何结构。该编码被序列化到 VLMs 的提示中,既可以与视频一起使用,也可以完全替代视频。具体而言,我们的方法中没有组件经过训练或微调。在 VSI-Bench 上,用空间编码增强 2B 和 4B 开源模型,相比仅使用帧的基线,平均准确率提高了 +4.1 个百分点,在绝对距离等数值估计任务上提升最大(+24.1 个百分点)。结果表明,通过语言通道传递的显式计算几何,能够恢复小型 VLMs 无法仅从像素中提取的空间能力。
cs.CV / 270 / 2609.34149

Functional Hand Type Prior for 3D Hand Pose Estimation and Action Recognition from Egocentric View Monocular Videos

面向第一人称视角单目视频 3D 手部姿态估计与动作识别的功能性手部类型先验
Roh, Wonseok, Lee, Seung Hyun, Ryoo, Won Jeong, Lee, Jakyung, Oh, Gyeongrok, Hwang, Sooyeon, Chi, Hyung-gun, Kim, Sangpil
Abstract
Current methods for egocentric view action recognition often face challenges in perceiving dynamic hand movements relying solely on geometrical or physical information. In this work, we effectively address this problem by gaining insights into the correlation between functional hand configurations and objects, which improves the detailed interpretation of real-world scenarios. To this end, we introduce a practical taxonomy of hand types based on the functioning perspective and utilize it for per-frame hand type labeling on existing datasets. We also propose a novel hand action recognition framework considering semantic details of the hand type as prior. This approach boosts the network's understanding of the continuous hand interaction throughout the action sequence. Our whole pipeline consists of three main modules: (1) Feature Extraction, (2) Egocentric Knowledge Module, which estimates 3D hand pose, object category, and hand type leveraging short-term cues, and (2) Egocentric Action Module, which aggregates per-frame knowledge, including text embeddings of hand type, over a longer time. In our extensive experiments with large-scale benchmarks, FPHA and H2O, our model outperforms current state-of-the-art methods, demonstrating its superior performance.
Chinese Translation
当前的第一人称视角动作识别方法通常面临挑战,即仅依靠几何或物理信息来感知动态手部运动。在本工作中,我们通过深入理解功能性手部配置与物体之间的相关性,有效地解决了这个问题,从而提升了对真实世界场景的详细解读。为此,我们引入了一种基于功能视角的实用手部类型分类法,并将其用于现有数据集上的逐帧手部类型标注。我们还提出了一种新颖的手部动作识别框架,将手部类型的语义细节作为先验考虑。这种方法提升了网络对整个动作序列中连续手部交互的理解。我们的整个流程包含三个主要模块:(1) 特征提取,(2) 第一人称知识模块,该模块利用短期线索估计 3D 手部姿态、物体类别和手部类型,以及 (3) 第一人称动作模块,该模块在较长时间内聚合逐帧知识,包括手部类型的文本嵌入。在我们使用大规模基准 FPHA 和 H2O 进行的大量实验中,我们的模型超越了当前最先进的方法,展示了其优越的性能。
cs.CV / 271 / 2609.34167

Natural Image Autoencoder-Based fMRI Representations for Trait and State Prediction

基于自然图像自编码器的fMRI表征用于特质与状态预测
Park, Juhyeon, Kim, Yeonwoo, Kim, Peter Yongho, Wang, Yansen, Xiao, Mingqing, Han, Dongqi, Li, Dongsheng, Moon, Taesup
Abstract
Foundation models pre-trained on large-scale fMRI datasets have shown strong downstream performance, but at substantial data and computation cost. To investigate how much fMRI-specific pre-training is actually needed for such performance, we introduce FReD, which derives fMRI representations from a frozen Deep Compression AutoEncoder (DCAE) pre-trained exclusively on natural images and pairs them with a task specific readout. For trait prediction, FReD summarizes frame-wise representations by their temporal mean and log-standard deviation and applies linear probing, with late fusion across two normalization schemes. For state prediction, it represents each frame as a single token and models temporal dependencies with a shallow Transformer. Across four resting-state datasets spanning six trait-prediction targets, linear probes on frozen DCAE features generally outperform those on fMRI foundation model representations and remain competitive with fully fine-tuned fMRI foundation models. On three task-fMRI state-prediction tasks, a temporal readout on DCAE features performs comparably to the strongest foundation models evaluated. A Gaussian injection analysis further shows that localized signal changes are recovered more accurately from the frozen DCAE features than from the evaluated foundation-model representations. Together, these results show that strong performance on current fMRI benchmarks is possible without fMRI-specific representation pre-training, making frozen natural-image features as a useful baseline for assessing its added value.
Chinese Translation
在大规模fMRI数据集上预训练的基础模型已展现出强大的下游性能,但代价是大量的数据和计算成本。为了探究这种性能实际上需要多少fMRI特定的预训练,我们引入了FReD,它从仅在自然图像上预训练的冻结深度压缩自编码器(DCAE)中推导出fMRI表征,并将其与任务特定的读出配对。对于特质预测,FReD通过时间均值和对数标准差来概括逐帧表征,并应用线性探测,在两种归一化方案上进行后期融合。对于状态预测,它将每一帧表示为一个单独的标记,并用浅层Transformer建模时间依赖性。在涵盖六个特质预测目标的四个静息态数据集上,在冻结的DCAE特征上的线性探测通常优于在fMRI基础模型表征上的线性探测,并且与完全微调的fMRI基础模型相比仍具有竞争力。在三个任务态fMRI状态预测任务上,在DCAE特征上的时间读出与所评估的最强基础模型性能相当。高斯注入分析进一步表明,从冻结的DCAE特征中恢复局部信号变化比从所评估的基础模型表征中更准确。总之,这些结果表明,无需fMRI特定的表征预训练,也能在当前fMRI基准上取得强大性能,使得冻结的自然图像特征成为评估其附加价值的有用基线。
cs.CV / 272 / 2609.34176

AGILE-GS: Anchor-Guided Fast Next-Best-View Selection for Active 3D Gaussian Splatting

AGILE-GS:面向主动3D高斯泼溅的锚点引导快速下一最佳视图选择
Khass, Amirhossein Mollaei, Motee, Nader
Abstract
Radiance fields need hundreds of views, and their placement matters as much as their number. Next-best-view (NBV) selection for 3D Gaussian Splatting (3DGS) usually scores every candidate in the pool and keeps one. Searching for information and choosing a camera, however, are separable problems. We present AGILE-GS, an anchor-guided NBV method that separates the two. A virtual anchor pose is optimized on SE(3) by Riemannian gradient ascent on expected information gain. It need not be reachable or in the pool; it marks where the model is most uncertain. Candidates are scored against the anchor's viewing geometry, and a greedy ridge-leverage step distills the pool into a small, non-redundant shortlist without rendering any candidate. The shortlist can be used in two ways. AGILE-GS takes the first view on it as the next view, so no Fisher information is computed for any candidate. AGILE-GS+ computes the Fisher information gain of each shortlisted view and picks the best, so the expensive evaluation runs on a handful of views rather than the whole pool. On standard benchmarks and in closed-loop embodied acquisition, both match or exceed existing baselines while cutting selection latency by one to two orders of magnitude.
Chinese Translation
辐射场需要数百个视图,而它们的位置与数量同样重要。针对3D高斯泼溅(3DGS)的下一最佳视图(NBV)选择通常对候选池中的每个候选进行评分并保留一个。然而,搜索信息和选择相机是可分离的两个问题。我们提出AGILE-GS,一种锚点引导的NBV方法,将两者分离。一个虚拟锚点姿态在SE(3)上通过黎曼梯度上升进行优化,以最大化期望信息增益。该锚点无需可达或位于候选池中;它标记了模型最不确定的位置。候选视图根据锚点的观察几何进行评分,并通过贪婪岭杠杆步骤将候选池提炼成一个小的、非冗余的候选列表,而无需渲染任何候选视图。该候选列表可以以两种方式使用:AGILE-GS将列表中的第一个视图作为下一个视图,因此无需为任何候选视图计算Fisher信息;AGILE-GS+计算每个候选列表中的视图的Fisher信息增益并选择最佳的一个,因此昂贵的评估只需在少数视图上运行,而不是整个候选池。在标准基准测试和闭环具身采集中,两者均达到或超过现有基线,同时将选择延迟降低了一到两个数量级。
cs.CV / 273 / 2609.34178

Enhanced Video Text Editing with Trajectory-Aligned Glyph Rendering

基于轨迹对齐字形渲染的增强视频文本编辑
Zhang, Shulian, Shu, Xiangyu, Li, Wenbo, Chen, Jian, Guo, Yong
Abstract
Video text editing aims to replace or add text in a video while keeping the rest of the video unchanged, which requires the edited text to be correct in every frame and to move coherently with the scene. Despite the remarkable progress of video diffusion models, they struggle to reproduce exact stroke structures and often produce garbled or wrong characters, especially for characters with complex strokes. To address this, we propose a trajectory-aligned glyph rendering reference that provides explicit per-frame glyph guidance following the position and perspective of the text, and a depth-normalized recognizer feature supervision that supervises the generated text on multi-depth features of a frozen text recognizer with per-depth normalized errors, targeting stroke errors overlooked by the diffusion loss. We further build VTEdit, a benchmark of 288 real-scene clips with 440 annotated text trajectories covering text replacement and text addition, which will be publicly released to facilitate future research. Experiments on VTEdit show that our method outperforms image text editing methods, video editing methods, and commercial models in text accuracy and background preservation, achieving a sentence accuracy of 0.9408, and receives the highest preference in a user study.
Chinese Translation
视频文本编辑旨在替换或添加视频中的文本,同时保持视频其余部分不变,这要求编辑后的文本在每一帧中都是正确的,并且与场景连贯地移动。尽管视频扩散模型取得了显著进展,但它们难以复现精确的笔画结构,并且经常产生乱码或错误的字符,尤其是对具有复杂笔画的字符。为了解决这个问题,我们提出了一种轨迹对齐的字形渲染参考,它遵循文本的位置和透视提供显式的逐帧字形引导,以及一种深度归一化的识别器特征监督,该监督在冻结的文本识别器的多深度特征上,以逐深度归一化误差对生成的文本进行监督,针对扩散损失所忽略的笔画错误。我们进一步构建了 VTEdit,一个包含 288 个真实场景片段和 440 个标注文本轨迹的基准,涵盖文本替换和文本添加,将公开发布以促进未来研究。在 VTEdit 上的实验表明,我们的方法在文本准确性和背景保持方面优于图像文本编辑方法、视频编辑方法和商业模型,达到了 0.9408 的句子准确率,并在用户研究中获得了最高的偏好。
cs.CV / 274 / 2609.34183

CRF Loss is How Networks Should Learn Boundaries in Weakly Supervised Segmentation

CRF损失:网络在弱监督分割中应如何学习边界
Li, Joshua, Boykov, Yuri
Abstract
Weakly Supervised Semantic Segmentation (WSSS) learns pixel-level predictions from image-level tags. Recent work focuses on improving coarse CAMs extracted from large vision-language models (commonly CLIP), but does little to improve their accuracy along segment boundaries. That job is instead delegated to a post-processing method like DenseCRF. However, because DenseCRF relies on low-level colour cues, it can flip correct labels to incorrect ones when neighbouring pixels share similar colours. SAM has recently been adopted as a natural alternative, yet it simply takes on DenseCRF's role as an intermediate "refinement" step that outputs one-hot pseudo-labels in prior work. By discarding the valuable uncertainty in CAMs, these one-hot pseudo-labels turn borderline errors into confidently wrong targets. Our key insight is that CAMs should supervise training alongside SAM boundaries, each through its own loss, rather than being fused together into a single hard target. Inspired by CRF potentials, we propose a framework that disentangles soft pseudo-labels as unary supervision and binary edge maps as pairwise supervision. We realize our framework in a single-stage model, DS-CRF, using CAMs from dino.txt and boundaries from SAM. DS-CRF sets a new state-of-the-art of 56.5% mIoU on MS COCO.
Chinese Translation
弱监督语义分割(WSSS)从图像级标签学习像素级预测。最近的工作专注于改进从大型视觉语言模型(通常为CLIP)中提取的粗糙CAM,但在提高沿分割边界的准确性方面所做甚少。这项工作反而被委托给DenseCRF之类的后处理方法。然而,由于DenseCRF依赖低级颜色线索,当相邻像素具有相似颜色时,它可能将正确标签翻转为错误标签。SAM最近被采用为一种自然的替代方案,但在先前的工作中,它只是承担了DenseCRF作为中间“细化”步骤的角色,输出one-hot伪标签。通过丢弃CAM中有价值的不确定性,这些one-hot伪标签将边缘错误转变为自信的错误目标。我们的关键见解是,CAM应与SAM边界一起通过各自的损失来监督训练,而不是融合成一个单一的硬目标。受CRF势函数的启发,我们提出了一个框架,将软伪标签解耦为一元监督,将二值边缘图解耦为成对监督。我们在单阶段模型DS-CRF中实现了我们的框架,使用来自dino.txt的CAM和来自SAM的边界。DS-CRF在MS COCO上创造了56.5% mIoU的新最高水平。
cs.CV / 275 / 2609.34190

MotionSpaceFlow: Representation-Aware Flow Matching in Direct Motion Space

MotionSpaceFlow:直接运动空间中的表示感知流匹配
Yu, Qing, Fujiwara, Kent
Abstract
Recent advances in diffusion and flow models have substantially improved text-driven human motion generation. Yet most methods generate in low-dimensional, temporally downsampled latent spaces learned primarily for reconstruction, a bottleneck that can limit generation quality and preclude direct manipulation of individual frames and joints. We introduce MotionSpaceFlow (MSFlow), a representation-aware flow-matching framework that predicts clean motion directly in continuous motion space without a learned encoder or decoder. To account for the anisotropic structure of direct motion representations, we propose representation-aware noise scaling and show how the initial Gaussian source scale governs the covariance of intermediate probability-path marginals. We further introduce a Representation-Aware Multimodal Diffusion Transformer (RA-MMDiT), which jointly updates token-level language and full-resolution motion features through joint attention while adapting temporal information flow to the motion representation: causal attention for incremental features defined by frame-to-frame changes, and bidirectional attention for global features such as absolute joint coordinates. Across different datasets and motion representations, MSFlow achieves state-of-the-art text-to-motion performance. Its global representation variant additionally enables zero-shot, inference-time control over any joint or frame through projection sampling without control-conditioned training, delivering leading motion quality with exact constraint satisfaction.
Chinese Translation
近年来,扩散模型和流模型的进展显著提升了文本驱动的人体运动生成。然而,大多数方法在低维、时间下采样的潜空间中生成,这些潜空间主要为重建而学习,这一瓶颈可能限制生成质量,并阻碍对单个帧和关节的直接操控。我们提出 MotionSpaceFlow (MSFlow),一种表示感知的流匹配框架,无需学习编码器或解码器,即可在连续运动空间中直接预测干净运动。为考虑直接运动表示的各向异性结构,我们提出表示感知的噪声缩放,并展示初始高斯源尺度如何决定中间概率路径边缘分布的协方差。我们进一步提出表示感知多模态扩散 Transformer (RA-MMDiT),它通过联合注意力共同更新 token 级语言和全分辨率运动特征,同时使时序信息流适应运动表示:对由逐帧变化定义的增量特征采用因果注意力,对绝对关节坐标等全局特征采用双向注意力。在不同数据集和运动表示上,MSFlow 均实现了当前最优的文本到运动性能。其全局表示变体还通过投影采样实现了对任意关节或帧的零样本、推理时控制,且无需控制条件训练,在精确满足约束的同时提供领先的运动质量。
cs.CV / 276 / 2609.34196

ConvCue: Complementary Visual Inductive Biases for Vision-Language Models

ConvCue:面向视觉语言模型的互补视觉归纳偏置
Lan, Zixuan, Sun, Shichu
Abstract
Modern vision-language models (VLMs) achieve strong performance across a broad range of multimodal tasks, yet still struggle with visual questions that require fine-grained discrimination and spatial understanding. These limitations motivate investigating whether supplementary visual representations can improve existing VLMs without replacing their native visual encoders. Pretrained convolutional networks offer a candidate feature source, motivated by their local connectivity and spatial weight sharing. We introduce CONVCUE, which augments the native visual representations of a pretrained VLM with final-stage features from a parallel, frozen pretrained CNN. A learnable adapter maps convolutional features to the native visual feature dimension, while gated cross-attention allows the original visual tokens to retrieve information from the CNN features. The enhanced tokens are passed through the original visual-to-language projector, and the model is adapted through a two-stage training procedure. We evaluate CONVCUE on Qwen3-VL-2B, Qwen3-VL-4B, and LLaVA-OneVision-7B across 13 multimodal benchmarks covering visual question answering, document and chart understanding, and multimodal reasoning. CONVCUE improves average benchmark performance over both the original models and matched two-stage fine-tuning controls on all three backbones. On Qwen3-VL-4B, it improves over the original model on all 13 benchmarks and raises the average score from 75.00 to 78.82 relative to the matched fine-tuning control. These results show that pretrained convolutional representations, when integrated through learned adaptation and fusion, can improve the visual understanding of existing VLMs without replacing their original visual encoders.
Chinese Translation
现代视觉语言模型(VLM)在广泛的多模态任务上取得了强劲性能,但在需要细粒度判别和空间理解的视觉问题上仍然表现不佳。这些局限性促使我们研究:补充的视觉表示能否在不替换现有VLM原生视觉编码器的情况下提升其性能。预训练卷积网络因其局部连接性和空间权重共享而成为候选特征来源。我们提出CONVCUE,它通过并行、冻结的预训练CNN的最终阶段特征来增强预训练VLM的原生视觉表示。一个可学习的适配器将卷积特征映射到原生视觉特征维度,而门控交叉注意力允许原始视觉token从CNN特征中检索信息。增强后的token通过原始视觉到语言投影器,模型通过两阶段训练过程进行适配。我们在Qwen3-VL-2B、Qwen3-VL-4B和LLaVA-OneVision-7B上,跨13个多模态基准(涵盖视觉问答、文档和图表理解以及多模态推理)评估CONVCUE。CONVCUE在三个骨干模型上均优于原始模型和匹配的两阶段微调控制组,提升了平均基准性能。在Qwen3-VL-4B上,它在所有13个基准上均优于原始模型,并且相对于匹配的微调控制组,将平均分从75.00提升至78.82。这些结果表明,预训练卷积表示,当通过学习到的适配和融合进行整合时,可以在不替换原始视觉编码器的情况下提升现有VLM的视觉理解能力。
cs.CV / 277 / 2609.34206

WorldGuide: Learning Success-Failure Boundaries in Latent World Models for Vision-Language-Action Policies

WorldGuide:在视觉-语言-动作策略的潜在世界模型中学习成功-失败边界
Liu, Lin, Zhang, Lu, Song, Ziying, Yang, Wu, Zhuang, Yuzheng, Zhuge, Yunzhi, Tao, Shuai, Liu, Wulong, Lu, Huchuan
Abstract
Latent world models offer a promising way to improve Vision-Language-Action policies by capturing the consequences of actions. However, models trained primarily on expert demonstrations have limited exposure to failure outcomes and may struggle to distinguish visually similar successful and failed interactions. We propose \textbf{WorldGuide}, a framework that learns these distinctions in latent space and uses them to guide policy training. WorldGuide combines predictive pretraining on successful and failed trajectories with contrastive learning on matched success--failure pairs. The learned predictor then provides a differentiable reward to guide joint optimization of the policy and visual encoder. The predictor is discarded after training, so deployment requires no additional world-model inference. Extensive experiments show that WorldGuide substantially improves VLA reliability and achieves state of the art performance on LIBERO 100 and SimplerEnv, reaching \textbf{96.8\%} and \textbf{72.0\%}, respectively. Code will be publicly available.
Chinese Translation
潜在世界模型通过捕捉动作后果,为改进视觉-语言-动作策略提供了一种有前景的途径。然而,主要基于专家演示训练的模型对失败结果的暴露有限,可能难以区分视觉上相似的成功与失败交互。我们提出 WorldGuide,一个在潜在空间中学习这些区分并用于指导策略训练的框架。WorldGuide 将成功与失败轨迹上的预测预训练与匹配成功-失败对上的对比学习相结合。学习到的预测器随后提供可微奖励,以指导策略和视觉编码器的联合优化。训练后丢弃预测器,因此部署无需额外世界模型推理。大量实验表明,WorldGuide 显著提升了 VLA 可靠性,并在 LIBERO 100 和 SimplerEnv 上达到了最先进性能,分别达到 96.8% 和 72.0%。代码将公开。
cs.CV / 278 / 2609.34221

WorldWeave: Growing Persistent Geometric Worlds for Video Generation

WorldWeave:为视频生成生长持久的几何世界
Huang, Yifan, Jiang, Lifan, Hao, Qingyue, Chen, Cheng, Wu, Boxi, Ren, Xiaoxue, He, Xiaofei, Zhao, Dehai
Abstract
Despite rapid progress, world models still lack explicit, persistent structural memory, making it difficult to preserve consistent world structure during continual scene expansion and cross-view revisits. To address this limitation, we present WorldWeave, a world generation framework that decouples world-state maintenance from visual rendering. Specifically, WorldWeave combines continual elevation-map generation with agent-guided scene organization and stitching to build an expandable explicit 3D world that incrementally extends structural memory while preserving existing structure. First, its terrain module uses diffusion-based image outpainting to generate continuous metric elevation maps under neighborhood conditioning and boundary constraints. Next, an agent integrates user intent, terrain evidence, and cross-region connectivity constraints to construct scenes through hierarchical semantic planning, deterministic geometry compilation, and local revision. Finally, during visual generation, planned camera trajectories query world geometry through a read-only interface, producing depth sequences that guide video synthesis without writing the generated results back into the world state. As a result, structural memory remains independent of short-window video generation, enabling continual expansion without predefined map boundaries and providing a consistent geometric basis for observations across trajectories and repeated visits.
Chinese Translation
尽管取得了快速进展,世界模型仍然缺乏显式、持久的结构记忆,这使得在连续场景扩展和跨视角重访过程中难以保持一致的世界结构。为了解决这一局限,我们提出了 WorldWeave,一种将世界状态维护与视觉渲染解耦的世界生成框架。具体而言,WorldWeave 将连续高程图生成与智能体引导的场景组织与拼接相结合,构建一个可扩展的显式 3D 世界,该世界在保留现有结构的同时增量扩展结构记忆。首先,其地形模块使用基于扩散的图像外绘,在邻域条件和边界约束下生成连续的度量高程图。接下来,智能体整合用户意图、地形证据和跨区域连通性约束,通过分层语义规划、确定性几何编译和局部修订来构建场景。最后,在视觉生成过程中,规划好的相机轨迹通过只读接口查询世界几何,生成深度序列以指导视频合成,而不会将生成结果写回世界状态。因此,结构记忆保持独立于短窗口视频生成,使得无需预定义地图边界即可持续扩展,并为跨轨迹和重复访问的观察提供一致的几何基础。
cs.CV / 279 / 2609.34223

Uncovering Ordinal-Matching Bias in Audio-Visual LLMs

揭示音频-视觉大型语言模型中的序数匹配偏差
Jung, Jihoo, Jang, Youngjoon, Cho, Hyebin, Yoo, Suho, Chung, Joon Son
Abstract
This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by simply matching the order of spoken sentences with the left-to-right, top-to-bottom arrangement of visible faces, rather than relying on audio-visual cues such as lip synchronization. We term this behavior \emph{ordinal-matching bias}. We further show that this bias can be substantially mitigated through a simple remedy, Ordinal-Decoupled Fine-Tuning (OD-FT), in which models are fine-tuned on synthetic videos where spatial positions of speakers and speaking order are independently randomized. Despite using only 400 synthetic training videos, OD-FT not only suppresses ordinal-matching bias but also improves audio-visual understanding on real-world videos, yielding average gains of 8.27\% for Qwen2.5-Omni and 2.57\% for video-SALMONN2+ across three audio-visual benchmarks.
Chinese Translation
这项工作旨在改进音频-视觉大型语言模型(AVLLMs)在多说话人场景中将语音与正确的可见说话人关联的方式。我们发现当前的 AVLLMs 在这项任务上经常失败,并分析了这些失败的本质。为此,我们构建了一个合成诊断数据集,其中多个可见说话人每人只说一个单词。对该语料库的分析揭示了三个近期开源 AVLLMs 中一致的错误模式:模型通过简单地将说话句子的顺序与可见人脸的从左到右、从上到下的排列进行匹配来归属话语,而不是依赖诸如唇同步之类的音频-视觉线索。我们将这种行为称为“序数匹配偏差”。我们进一步表明,这种偏差可以通过一种简单的补救措施——序数解耦微调(OD-FT)——得到大幅缓解,其中模型在合成视频上进行微调,这些视频中说话人的空间位置和说话顺序是独立随机化的。尽管仅使用了 400 个合成训练视频,OD-FT 不仅抑制了序数匹配偏差,还提升了在真实世界视频上的音频-视觉理解,在三个音频-视觉基准上分别为 Qwen2.5-Omni 带来 8.27% 的平均提升,为 video-SALMONN2+ 带来 2.57% 的平均提升。
cs.CV / 280 / 2609.34231

ReGDiff: Guided Diffusion in Regulated Latent Space for Exploring Metamaterial Voxel Geometry

ReGDiff:在受调控的潜在空间中进行引导扩散以探索超材料体素几何
Zhan, Wangzhi, Chen, Jianpeng, Fu, Dongqi, Zhou, Dawei
Abstract
Metamaterials are artificially engineered structures whose mechanical and physical behaviors are strongly shaped by geometry rather than composition. Voxel representation provides a unified format for metamaterial geometry generation, as it can express diverse classes such as truss, shell, and porous structures within a single cubic discretization. However, voxel-based generation faces a plausibility-novelty trade-off: staying close to known geometries helps preserve geometric regularities, while moving away from them is necessary for novelty but may produce degenerate geometries. To address this challenge, we propose REGDIFF, a generative framework that couples voxel representation with latent space regulation and guided diffusion. REGDIFF introduces a repel-and-sink (RAS) mechanism to smooth the latent distribution of plausible geometries, and short-range repulsion (SRR) guidance to discourage generation overly close to known samples while maintaining geometric plausibility. We further contribute a voxel-based benchmark covering truss- and shell-type metamaterial geometries, together with an evaluation module for geometric plausibility, novelty, and diversity. Experiments show that REGDIFF outperforms voxel-based generative baselines, achieving +8.9% in geometric plausibility, +46.4% in novelty, and +128.6% in diversity on average across two datasets. These results suggest that REGDIFF is a strong geometry candidate generator for downstream evaluation. Our code is provided at https://github.com/wzhan24/ReGDiff.
Chinese Translation
超材料是人工设计的结构,其力学和物理行为主要由几何形状而非成分决定。体素表示提供了一种统一的格式用于超材料几何生成,因为它可以在单一立方离散化中表达桁架、壳和多孔结构等多种类型。然而,基于体素的生成面临合理性-新颖性权衡:靠近已知几何有助于保持几何规律性,而远离已知几何对于新颖性是必要的,但可能产生退化几何。为了解决这一挑战,我们提出了 ReGDiff,一个将体素表示与潜在空间调控和引导扩散相结合的生成框架。ReGDiff 引入了排斥-下沉(RAS)机制来平滑合理几何的潜在分布,以及短程排斥(SRR)引导来阻止生成过于接近已知样本,同时保持几何合理性。我们进一步贡献了一个基于体素的基准,涵盖桁架型和壳型超材料几何,以及一个用于几何合理性、新颖性和多样性的评估模块。实验表明,ReGDiff 优于基于体素的生成基线,在两个数据集上平均实现了几何合理性 +8.9%、新颖性 +46.4% 和多样性 +128.6% 的提升。这些结果表明 ReGDiff 是用于下游评估的强大几何候选生成器。我们的代码见 https://github.com/wzhan24/ReGDiff。
cs.CV / 281 / 2609.34232

Trustworthy synthetic visual media: Evidence across the media lifecycle

可信的合成视觉媒体:贯穿媒体生命周期的证据
Jia, Zexi, Yuan, Zhiqiang, Zhou, Jie, Zhang, Jinchao
Abstract
Images and videos have long helped people understand what happened and how a work came into being. Generative systems complicate that role. Realistic media can now be produced and revised without leaving a stable history, so appearance no longer reveals whether a scene was captured, synthesized, or altered along the way. Trust must instead come from evidence that explains the path an asset has taken and the circumstances in which it was used. Some of this evidence can be recovered from the media, while some must be recorded during production and preserved as the asset circulates. This review brings those approaches together and asks when their claims remain meaningful after ordinary processing or deliberate manipulation. We argue that trustworthy media do not depend on one universal marker of authenticity. The evidence must suit the question at hand, reach the person making the judgment, and remain open to correction when better information emerges. The larger goal is to keep the history of media intelligible even as the media itself continues to change.
Chinese Translation
图像和视频长期以来帮助人们理解发生了什么,以及一件作品如何形成。生成式系统使这一作用变得复杂。如今,逼真的媒体可以在不留下稳定历史的情况下被制作和修改,因此外观不再能揭示某一场景是被拍摄、合成,还是在过程中被改动。信任必须转而来自证据,这些证据说明资产经历了何种路径,以及在何种情况下被使用。其中一些证据可以从媒体中恢复,而另一些则必须在制作过程中记录,并在资产流转时保存下来。本综述将这些方法汇集在一起,并追问:在经过常规处理或蓄意操纵之后,它们的主张何时仍然有意义。我们认为,可信媒体并不依赖于某种普遍的真实性标记。证据必须适合当前的问题,抵达作出判断的人,并在出现更好信息时保持可修正。更大的目标是,即使媒体本身不断变化,也要让媒体的历史保持可理解。
cs.CV / 282 / 2609.34235

SegBanana: Steering Unified Multimodal Models into Medical Segmenters

SegBanana:将统一多模态模型引导为医学分割器
Liang, Xiaoye, Yan, Ye, Yin, Mingze, Feng, Shikun, Xu, Mai, Liu, Haiguang, Jiang, Lai, Zhu, Yiheng
Abstract
Medical image segmentation remains challenging in practical deployment, as models often struggle to generalize beyond the distributions covered by their training data and high-quality pixel-level annotations are typically unavailable for adaptation. Inspired by the cross-task transferability of large language models, we investigate whether unified multimodal models (UMMs) can transfer their pretrained visual understanding, reasoning, and generation capabilities to medical image segmentation without task-specific post-training. By recasting segmentation as structured visual generation, we find that frontier UMMs (e.g., Nano Banana) already exhibit basic segmentation capabilities across diverse clinical scenarios, but still struggle with challenging tasks requiring specialized anatomical or domain-specific knowledge. We further show that these limitations can be effectively mitigated by incorporating visual anatomical knowledge from in-context exemplars, expanding candidate solutions through repeated sampling, and refining suboptimal predictions via targeted editing.Motivated by these observations, we propose SegBanana, to our knowledge, the first agentic visual generation framework for training-free medical image segmentation. SegBanana builds on a frozen UMM as the core generative model, augmented with Anatomy-Aware Knowledge Retrieval and Comparative Quality Critique to unlock its potential segmentation capability. A State-Aware Multimodal Controller maintains structured state and iteratively orchestrates these tools, repeatedly refining intermediate predictions toward higher-quality masks. Across eight medical segmentation datasets, SegBanana achieves an average mDice of 77.45%, outperforming representative generalist (SAM3 and SegGPT) and medical-specific (BiomedParse and MedSAM3) baselines by at least 14.93 points, while remaining robust to out-of-domain visual supports.
Chinese Translation
医学图像分割在实际部署中仍具挑战性,因为模型往往难以泛化到其训练数据覆盖的分布之外,且高质量像素级标注通常无法用于适应。受大语言模型跨任务可迁移性的启发,我们探究统一多模态模型(UMMs)能否将其预训练的视觉理解、推理和生成能力迁移至医学图像分割,而无需任务特定的后训练。通过将分割重新表述为结构化视觉生成,我们发现前沿 UMMs(如 Nano Banana)已在多样的临床场景中展现出基本的分割能力,但在需要专门解剖学或领域特定知识的挑战性任务上仍表现不佳。我们进一步表明,这些局限可通过以下方式有效缓解:纳入来自上下文示例的视觉解剖知识、通过重复采样扩展候选解,以及通过针对性编辑精炼次优预测。受这些观察的启发,我们提出 SegBanana,据我们所知,这是首个用于免训练医学图像分割的智能体视觉生成框架。SegBanana 以冻结的 UMM 作为核心生成模型,并辅以解剖感知知识检索和比较质量评判,以释放其潜在的分割能力。状态感知多模态控制器维护结构化状态,并迭代地协调这些工具,反复精炼中间预测,以获得更高质量的掩码。在八个医学分割数据集上,SegBanana 平均 mDice 达到 77.45%,超越代表性通用基线(SAM3 和 SegGPT)和医学专用基线(BiomedParse 和 MedSAM3)至少 14.93 个点,同时对域外视觉支持保持稳健。
cs.CV / 283 / 2609.34237

DecFlowEdit: Self-Localized Flow-based Image Editing via Guidance Decoupling

DecFlowEdit:通过引导解耦的自定位基于流的图像编辑
Zhan, Zheyuan, Wang, Can, Chen, Jiawei, Chen, Chun, Lyu, Siwei, Zheng, Zeyu, Chen, Defang
Abstract
Flow-based image editing (FlowEdit) enables inversion-free semantic changes through the difference between source and target velocities. In this paper, we observe that FlowEdit's default classifier-free guidance (CFG) configuration, with asymmetric source and target scales, causes substantial background leakage. Matching these guidance scales, for example by removing CFG, improves edit-relevant localization but severely degrades editability. To get the best of both worlds, we propose DecFlowEdit, which decouples the optimal guidance scales for localization and for editing in flow-based generative models. In particular, DecFlowEdit first extracts an edit-relevant prior by temporally aggregating velocity differences evaluated without CFG, and then uses this prior to reweight the original updates under default CFG. Our method remains training-free and inversion-free, requiring neither external spatial masks nor attention manipulation. Experiments on PIE-Bench across FLUX, SD3, and SD3.5 show that DecFlowEdit improves background preservation, reducing structure distance by approximately 61 to 73 percent and background LPIPS by 68 to 80 percent relative to FlowEdit at comparable editing fidelity.
Chinese Translation
基于流的图像编辑(FlowEdit)通过源速度和目标速度之间的差异实现无需反转的语义变化。在本文中,我们观察到FlowEdit默认的无分类器引导(CFG)配置,其源和目标尺度不对称,会导致大量的背景泄漏。匹配这些引导尺度,例如通过移除CFG,可以改善与编辑相关的定位,但严重降低可编辑性。为了两全其美,我们提出了DecFlowEdit,它将基于流的生成模型中用于定位和编辑的最优引导尺度解耦。具体来说,DecFlowEdit首先通过时间聚合在没有CFG的情况下评估的速度差异来提取编辑相关的先验,然后使用该先验在默认CFG下重新加权原始更新。我们的方法保持无需训练和无需反转,既不需要外部空间掩码,也不需要注意力操作。在PIE-Bench上针对FLUX、SD3和SD3.5的实验表明,DecFlowEdit改善了背景保留,在可比的编辑保真度下,相对于FlowEdit,将结构距离降低了约61%至73%,并将背景LPIPS降低了68%至80%。
cs.CV / 284 / 2609.34271

Scaling Versatile 3D Assets Editing with a Million-Scale Dataset

利用百万级数据集扩展通用3D资产编辑
Li, Badi, Huang, Tianxin, Zhou, Yu, Zheng, Wei-Shi, Ma, Yi, Gao, Shenghua
Abstract
Although recent 3D generative models produce increasingly realistic assets, controllable 3D asset editing remains challenging. Existing methods are limited by scarce training data, insufficient source-aware modeling, and a lack of practical evaluation protocols. To address these limitations, we present Alchemy3D, a unified framework for training and evaluating versatile 3D asset editors that covers data construction, model architecture, and benchmark evaluation. Specifically, we curate Alchemy3D-1M, a large-scale 3D editing dataset containing 1.25M assets and 1.38M editing pairs across seven editing types. On this data, we train a family of generative flow models for general-purpose 3D asset editing. The model family supports image- and text-conditioned editing, few-step inference, and transfer to multi-view 3D part segmentation. We further introduce GEdit3D-Bench, a large-scale, open-world benchmark with a multi-dimensional evaluation protocol. Across existing and newly introduced benchmarks, our method outperforms prior methods on most metrics of editing fidelity, source preservation, and visual quality.
Chinese Translation
尽管最近的3D生成模型能生成越来越逼真的资产,但可控的3D资产编辑仍然具有挑战性。现有方法受限于训练数据稀缺、源感知建模不足以及缺乏实用的评估协议。为了解决这些限制,我们提出了Alchemy3D,一个用于训练和评估通用3D资产编辑器的统一框架,涵盖数据构建、模型架构和基准评估。具体而言,我们构建了Alchemy3D-1M,一个大规模3D编辑数据集,包含125万个资产和138万个编辑对,涵盖七种编辑类型。在此数据上,我们训练了一系列生成流模型,用于通用3D资产编辑。该模型系列支持图像和文本条件编辑、少步推理,并可迁移到多视图3D部件分割。我们进一步引入了GEdit3D-Bench,一个大规模开放世界基准,具有多维评估协议。在现有和新引入的基准上,我们的方法在编辑保真度、源保留和视觉质量的大多数指标上优于先前方法。
cs.CV / 285 / 2609.34277

See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology

观察、测量与推理:学习病理学中的视觉基础推理
Zhang, Chengyang, Zhang, Wenchuan, Li, Bo, Li, Mengran, Liu, Xinyu, Yang, Jiaming, Chen, Jie, Zhang, Zhang, Yi, Yuhao, Bu, Hong, Lv, Jiancheng
Abstract
Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that can persist even when final answers are correct. In this paper, we propose ASPECT to improve visually grounded reasoning through explicit supervision of cellular appearance and abundance. ASPECT trains intermediate visual tokens through pathology feature reconstruction, cell feature alignment, and count supervision. Three-stage supervised fine-tuning teaches the model to perceive, generate visual tokens, and reason, followed by reinforcement learning that rewards answer correctness and consistency with reported measurements. We also introduce PathoVernier, a benchmark of 759 expert-reviewed questions from five pathology datasets covering four cellular composition tasks. It evaluates both final answers and intermediate measurements to expose errors hidden by answer accuracy. On PathoVernier, ASPECT achieves relative accuracy gains of approximately 19.2% over the strongest baseline, Gemini-3.1-Pro, and 99.3% over its Qwen3-VL-8B backbone, while reducing RAWR, which measures counting errors within correct responses, by 28.1% and 42.7%, respectively. ASPECT also improves over its backbone on three external pathology benchmarks covering classification and question answering beyond cellular composition tasks.
Chinese Translation
病理学评估依赖于识别组织学图像中的细粒度视觉细节。视觉语言模型(VLMs)越来越多地支持病理学解释,但其感知这些细节的能力仍然不足。这种弱点导致不准确的细胞观察,即使最终答案正确,这些观察也可能持续存在。在本文中,我们提出ASPECT,通过对细胞外观和丰度的显式监督来改进视觉基础推理。ASPECT通过病理特征重建、细胞特征对齐和计数监督来训练中间视觉标记。三阶段监督微调教会模型感知、生成视觉标记并进行推理,随后是强化学习,奖励答案正确性以及与报告测量的一致性。我们还引入了PathoVernier,一个包含来自五个病理数据集的759个专家评审问题的基准,涵盖四个细胞组成任务。它评估最终答案和中间测量,以暴露被答案准确性掩盖的错误。在PathoVernier上,ASPECT相对于最强基线Gemini-3.1-Pro实现了约19.2%的相对准确率提升,相对于其Qwen3-VL-8B主干实现了99.3%的提升,同时将RAWR(衡量正确响应中的计数错误)分别降低了28.1%和42.7%。ASPECT还在三个外部病理基准上优于其主干,这些基准涵盖细胞组成任务之外的分类和问答。
cs.CV / 286 / 2609.34286

Dexterous Tactile World Model

灵巧触觉世界模型(Dexterous Tactile World Model)
Zeng, Ziyao, Sun, Xiatao, Wang, Hao, Pan, Yueyang, Yu, Zhengxiang, Yang, Fengyu, Liu, Tianyu, Fan, Zhiwen, Rakita, Daniel
Abstract
World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future-frame prediction of egocentric manipulation from both observed video and tactile signals from a glove worn on each hand. We condition a pretrained video diffusion transformer on each hand's tactile signal through a zero-initialized residual at the corresponding hand location in the video tokens, while a causal mask prevents predicted frames from accessing future information. Compared with a vision-only model matched in architecture, parameters, and training, DTWM reduces the underestimation of hand motion from 23% to 9%, while reducing the perceptual error in the hand region by 7.4% across three training runs per model. The benefit also increases over the prediction horizon, with the improvement in the later predicted chunks being about 4.1x larger than in the first. DTWM also outperforms other visual-tactile world models under the same setting, and training with touch improves future-frame prediction even when no touch is available at inference. Ablations show that the model benefits from both the magnitude and spatial location of force: replacing the tactile signal with binary contact states, either per hand or per location, increases prediction error. The observed course of the force indicates whether the interaction will persist or change.
Chinese Translation
用于操作的世界模型通常从视频中训练,但决定操作如何展开的事件,例如建立和解除接触,难以通过视觉观察到,而通常更容易通过触觉感知。我们提出灵巧触觉世界模型(Dexterous Tactile World Model, DTWM),一种视频世界模型,用于从观测视频和每只手佩戴的手套所采集的触觉信号中,对未来以自我为中心的操作帧进行预测。我们通过视频 token 中对应手部位置处的零初始化残差,将预训练视频扩散 transformer 以每只手的触觉信号为条件,同时因果掩码阻止预测帧访问未来信息。与在架构、参数和训练上匹配的纯视觉模型相比,DTWM 将手部运动低估从 23% 降至 9%,同时在对每个模型进行三次训练运行中,将手部区域的感知误差降低 7.4%。这种收益还随预测时域增加而增大,后段预测块的改进约是首段的 4.1 倍。在相同设置下,DTWM 也优于其他视觉-触觉世界模型,并且即使推理时没有触觉可用,使用触觉进行训练也能改善未来帧预测。消融实验表明,该模型受益于力的大小和空间位置:将触觉信号替换为二值接触状态,无论是按手还是按位置,都会增加预测误差。观测到的力变化过程指示交互将持续还是改变。
cs.CV / 287 / 2609.34294

Semantic Modality Compensation for Unsupervised Visible-Infrared Person Re-identification under Unpaired Settings

非配对设置下无监督可见光-红外行人重识别的语义模态补偿
Chen, Duanning, He, Ke, Yang, Bin, Yao, Yongxiang
Abstract
Unsupervised visible-infrared person re-identification (USL-VI-ReID) learns person representations that can be compared across modalities without identity annotations. In the unpaired setting, however, identity correspondences between modalities are often incomplete, leaving many identities without an observed counterpart in the other modality. Existing unpaired methods bridge this gap by generating or mapping features for the other modality, mainly by exploiting the statistics of visual features without explicitly separating content that is discriminative for identity from style that is specific to modality. Consequently, the generated features may distort identity cues or inherit bias from the source modality, undermining the reliability of supervision across modalities. We formulate unpaired learning across modalities as a semantic compensation problem and propose Semantic Modality Compensation (SMC), a framework based on prompt composition that decouples identity semantics from modality style within a shared visual semantic space. SMC first constructs a discriminative ReID space through augmented dual contrastive learning, yielding pseudo labels, cluster prototypes, and memory banks for each modality. It then learns visible and infrared modality prompts in the CLIP semantic space and maps clusters obtained from pseudo labels to identity semantic tokens. For each cluster lacking a reliable match in the other modality, SMC combines its identity token with the prompt for the target modality to synthesize a semantic counterpart in the missing modality. The synthesized counterpart is then projected back into the ReID space and injected into a compensation memory through confidence gating. Extensive experiments under both paired and unpaired settings demonstrate that SMC consistently outperforms state-of-the-art methods, with particularly large gains when identity mismatch is severe.
Chinese Translation
无监督可见光-红外行人重识别(USL-VI-ReID)无需身份标注即可学习能够跨模态比较的行人表征。然而,在非配对设置下,模态间的身份对应关系往往不完整,导致许多身份在另一模态中没有观测到的对应样本。现有的非配对方法通过为另一模态生成或映射特征来弥合这一差距,主要利用视觉特征的统计信息,而没有显式地将对身份具有判别性的内容与模态特有的风格分离开来。因此,生成的特征可能会扭曲身份线索或继承源模态的偏差,从而削弱跨模态监督的可靠性。我们将跨模态的非配对学习形式化为一个语义补偿问题,并提出语义模态补偿(SMC),这是一个基于提示组合的框架,在共享视觉语义空间内将身份语义与模态风格解耦。SMC 首先通过增强的双重对比学习构建一个具有判别性的 ReID 空间,为每个模态生成伪标签、聚类原型和记忆库。然后,它在 CLIP 语义空间中学习可见光和红外模态提示,并将由伪标签获得的聚类映射为身份语义令牌。对于在另一模态中缺乏可靠匹配的每个聚类,SMC 将其身份令牌与目标模态的提示相结合,以在缺失模态中合成一个语义对应样本。然后,将合成的对应样本投影回 ReID 空间,并通过置信度门控注入到补偿记忆中。在配对和非配对设置下的大量实验表明,SMC 始终优于最先进的方法,在身份不匹配严重时提升尤为显著。
cs.CV / 288 / 2609.34299

PSM: Dataset Distillation Based on Precise Statistical Matching by Difficulty

PSM:基于难度精确统计匹配的数据集蒸馏
Ma, Hongxu, Li, Guang, Wang, Shijie, Zhou, Dongzhan, Yang, Suorong, Sun, Baoli, Ogawa, Takahiro, Haseyama, Miki, Wang, Zhihui
Abstract
Dataset distillation (DD) condenses a large original dataset into a small distilled dataset with high training utility. Decoupled statistical matching methods substantially reduce distillation time and memory overhead while achieving strong performance. However, they typically supervise all distilled samples using running statistics estimated from the entire original dataset. These statistics mainly capture the average feature distribution while overlooking differences in sample difficulty, limiting their ability to characterize the difficulty structure of the original data. To address this issue, we propose Precise Statistical Matching (PSM) by difficulty. After pretraining, PSM uses the Global Precision Score (GPS) to estimate image difficulty, ranks the samples within each class, and partitions each class into IPC (images per class) difficulty groups. During distillation, Statistics Updated Again (SUA) updates the teacher's batch normalization (BN) running statistics through forward passes on original samples from each group, providing difficulty-specific supervision for the corresponding distilled batch. Meanwhile, Initial Sample Screening (ISS) initializes distilled samples using original images from the corresponding difficulty group, providing an effective starting point for precise matching. Experiments across multiple datasets and model architectures demonstrate that PSM broadens the difficulty range of distilled samples and improves downstream performance in most evaluated settings. Code will be released.
Chinese Translation
数据集蒸馏(DD)将大型原始数据集压缩为具有高训练效用的小型蒸馏数据集。解耦统计匹配方法大幅降低了蒸馏时间和内存开销,同时实现了强劲性能。然而,它们通常使用从整个原始数据集估计的运行统计量来监督所有蒸馏样本。这些统计量主要捕捉平均特征分布,而忽略了样本难度的差异,限制了它们刻画原始数据难度结构的能力。为了解决这一问题,我们提出了基于难度的精确统计匹配(PSM)。在预训练之后,PSM使用全局精度分数(GPS)来估计图像难度,对每个类别内的样本进行排序,并将每个类别划分为IPC(每类图像数)个难度组。在蒸馏过程中,统计量再次更新(SUA)通过对每个组的原始样本进行前向传播来更新教师模型的批量归一化(BN)运行统计量,为相应的蒸馏批次提供特定于难度的监督。同时,初始样本筛选(ISS)使用来自相应难度组的原始图像初始化蒸馏样本,为精确匹配提供了有效的起点。在多个数据集和模型架构上的实验表明,PSM拓宽了蒸馏样本的难度范围,并在大多数评估设置中提升了下游性能。代码将开源。
cs.CV / 289 / 2609.34302

CAT-Free: Multi-View Pedestrian Localization without Calibration, Annotations, or Target-Scene Training via Adaptive Geometric Filtering

CAT-Free:无需标定、标注或目标场景训练,基于自适应几何滤波的多视角行人定位
Sakai, Taigo, Kouno, Hiroki, Kato, Naoki, Hotta, Kazuhiro
Abstract
Multi-camera pedestrian localization is useful for wide-area monitoring in public and commercial spaces. However, deploying these systems often requires considerable setup for each new environment. Existing methods typically require camera calibration, position annotations, or target-scene training. CAT-Free removes all three requirements. It uses synchronized RGB video as its only scene-specific input. Camera configuration is estimated directly from the video. Pedestrian locations are then estimated by combining observations from multiple cameras. Automatic camera estimation is not always accurate. This can produce unreliable pedestrian locations. CAT-Free therefore introduces two adaptive geometric filters. They remove unreliable position estimates. Their thresholds are estimated from each input sequence. CAT-Free achieves 82.5, 84.5, and 65.7 MODA on WildTrack, MultiviewX, and GMVD. It uses no supplied calibration, position annotations, or target-scene training. Published methods using such scene-specific information report 88.2--95.0 MODA on WildTrack and 83.9--96.5 on MultiviewX under their respective protocols. CAT-Free also transfers without retuning. It reaches 74.9 MODA on four additional sequences and 78.6 on an unseen 8-camera installation. Finally, localization uncertainty predicts MODA with $r=-0.98$. This provides a label-free estimate of localization reliability.
Chinese Translation
多摄像机行人定位可用于公共和商业空间的广域监控。然而,部署这些系统通常需要针对每个新环境进行大量设置。现有方法通常需要相机标定、位置标注或目标场景训练。CAT-Free 消除了这三项要求。它仅使用同步 RGB 视频作为场景特定输入。相机配置直接从视频中估计。随后,通过融合多个摄像机的观测来估计行人位置。自动相机估计并不总是准确,这可能导致不可靠的行人位置。因此,CAT-Free 引入了两种自适应几何滤波器,用于剔除不可靠的位置估计;其阈值由每个输入序列估计得到。CAT-Free 在 WildTrack、MultiviewX 和 GMVD 上分别达到 82.5、84.5 和 65.7 MODA。它不使用提供的标定、位置标注或目标场景训练。使用此类场景特定信息的已发表方法在各自协议下报告 WildTrack 上 88.2--95.0 MODA,MultiviewX 上 83.9--96.5 MODA。CAT-Free 还可无需重新调参进行迁移。它在四个额外序列上达到 74.9 MODA,并在一个未见过的8摄像机部署上达到 78.6 MODA。最后,定位不确定性以 r=-0.98 预测 MODA,这提供了一种无标签的定位可靠性估计。
cs.CV / 290 / 2609.34309

MaLiang-Harness: A Programmable Path to Image and Video Generation

MaLiang-Harness:通向图像和视频生成的可编程路径
Zhao, Haoyu, Zhang, Zihao, Wang, Xudong, Gu, Jiaxi, Wu, Zuxuan, Jiang, Yu-Gang, Yan, Shuicheng
Abstract
Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to-Visual (P2V) gap and introduce MaLiang-Harness, a unified framework for organizing MLLM-driven visual generation into a persistent process of construction, inspection, and revision. Its central design is to make the evolving visual program, its construction history, and its verification share a common revision reference. We define the Persistent Executable Generation (PEG) state as preserving programs and task context. Traceable Generation Process (TGP) connects edits to rendered evidence, and Revision-aware Editing and Verification (REV) supports restoration and checks the current revision before completion. Together, these mechanisms coordinate planning, execution, and visual feedback across rendering backends. We evaluate 11 powerful closed-source MLLMs on MaLiang-IBench and four on MaLiang-VBench, measuring generation success, visual quality, and computational cost. GPT-6-Astra achieves 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds. The comparison also reveals a mismatch between general capability scores and visual generation performance, with similarly scored models differing substantially in their ability to satisfy visual requirements. MaLiang-Harness provides a systematic basis for studying how MLLMs translate executable code into visual outcomes, exposing both the potential of programmable generation and the limitations of general benchmarks as predictors of this ability. The project is available at https://github.com/gulucaptain/MaLiang-Harness.
Chinese Translation
可执行程序提供了对图像和视频构建方式的显式控制,但生成可运行的代码只是视觉创作的开始。一个程序可能正确执行,却违反了所要求的构图、外观或运动。我们将这种差异定义为程序到视觉(Program-to-Visual, P2V)差距,并引入MaLiang-Harness,一个统一的框架,用于将MLLM驱动的视觉生成组织为构建、检查和修订的持续过程。其核心设计是使不断演进的视觉程序、其构建历史及其验证共享一个共同的修订参考。我们将持久可执行生成(Persistent Executable Generation, PEG)状态定义为保留程序和任务上下文。可追踪生成过程(Traceable Generation Process, TGP)将编辑与渲染证据连接起来,而修订感知编辑与验证(Revision-aware Editing and Verification, REV)支持恢复并在完成前检查当前修订。这些机制共同协调跨渲染后端的规划、执行和视觉反馈。我们在MaLiang-IBench上评估了11个强大的闭源MLLM,在MaLiang-VBench上评估了4个,衡量生成成功率、视觉质量和计算成本。GPT-6-Astra在两个基准测试上都实现了100%的生成成功率,其中96.0%的图像任务和76.9%的视频任务满足所有质量阈值。比较还揭示了通用能力分数与视觉生成性能之间的不匹配,得分相似的模型在满足视觉要求的能力上差异显著。MaLiang-Harness为研究MLLM如何将可执行代码转化为视觉结果提供了系统基础,既揭示了可编程生成的潜力,也暴露了通用基准作为这种能力预测指标的局限性。项目地址:https://github.com/gulucaptain/MaLiang-Harness
cs.CV / 291 / 2609.34314

PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?

PlaylistEval:视频-语言评判器在日级尺度及以上是否可信?
Islam, Shayekh Bin, Song, Hwanjun
Abstract
Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typically only a few minutes long, many answer pairs can be separated from the transcript alone, and collecting human judgments does not scale to ultra-long videos. We introduce PlaylistEval, an agentic framework that builds video-language judge benchmarks over 100-hour playlist collection without human annotation. It automatically generates questions with paired answers whose differences are controlled by causal degradation, so that every pair demands retrieval across the collection. The resulting benchmark contains 630 pairs across seven domains spanning both static and dynamic knowledge, and on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781). Evaluating 17 omnimodal and multimodal models from eight families reveals that frontier judges reach only 75.4% pairwise accuracy, while open-source judge models perform far behind. We further show that both retrieval and final judgment depend on using multiple modalities, and that judge accuracy degrades as the playlist set grows. We release our pipeline, benchmark, and evaluation code at https://playlisteval.github.io.
Chinese Translation
视频-语言模型正越来越多地被用作视频理解的评判器,既用于评估模型输出,也用于训练奖励模型。当证据被埋藏在长达一天的视频中时,它们的判断是否仍然可靠,尚待确立。现有基准无法回答这一问题。它们的视频通常只有几分钟长,许多问答对仅凭转录文本就能被区分开,而收集人类判断无法扩展到超长视频。我们提出 PlaylistEval,一个智能体框架,可在无需人工标注的情况下,基于超过 100 小时的播放列表集合构建视频-语言评判基准。它自动生成带有配对答案的问题,这些答案之间的差异由因果退化控制,因此每一对都要求跨整个集合进行检索。最终基准包含跨越七个领域的 630 对数据,涵盖静态和动态知识;在包含 152 对的分层子集上,它与人类判断的一致率为 93.0%(IAA 0.781)。对来自八个家族的 17 个全模态和多模态模型进行评估后发现,前沿评判器仅达到 75.4% 的成对准确率,而开源评判模型的表现则远远落后。我们进一步表明,检索和最终判断都依赖于使用多种模态,并且随着播放列表集合的增长,评判准确率会下降。我们在 https://playlisteval.github.io 发布了我们的流程、基准和评估代码。
cs.CV / 292 / 2609.34319

Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference

文本-视觉协同Token缓存:一种用于高效视觉-语言-动作推理的免训练框架
Li, Qianer, Zhang, Chengjie, Chen, Jingwen, Tong, Zanjia, Zhang, Jiyuan, Zhang, Hong
Abstract
Vision-Language-Action (VLA) models enable generalizable robotic control but remain computationally expensive. Token caching provides a training-free, plug-and-play acceleration alternative. However, existing VLA caching does not fully exploit a key inductive bias of VLA models: text-vision synergy, wherein textual semantics guide the precise visual grounding of task-relevant regions. In particular, existing designs insufficiently account for head-wise reliability in attention aggregation and layer-wise stability in cache reuse. To address this, we propose Text-Vision Synergistic Token Caching (TVCache), a training-free framework for efficient VLA inference. TVCache filters attention heads based on text-vision information focus to improve task-relevant and physically consistent visual grounding. Concurrently, we introduce a reuse-layer selection mechanism guided by text-vision entropy differences to avoid caching unstable representations and improve cache resource allocation. Extensive experiments across four representative VLA models, two simulation benchmarks, and real-world robotic tasks demonstrate the effectiveness and generality of TVCache. At matched token-retention ratios, TVCache consistently improves task success over existing VLA caching with comparable computational cost. On OpenVLA-OFT, it improves average success by up to 14.5 percentage points over VLA-Cache at 12.5% retention while reducing FLOPs by 2.45x relative to full-token inference.
Chinese Translation
视觉-语言-动作(VLA)模型能够实现可泛化的机器人控制,但计算成本依然高昂。Token缓存提供了一种免训练、即插即用的加速替代方案。然而,现有的VLA缓存并未充分利用VLA模型的一个关键归纳偏置:文本-视觉协同,即文本语义引导任务相关区域的精确视觉定位。特别是,现有设计在注意力聚合中未能充分考虑头级别的可靠性,以及在缓存复用中未能充分考虑层级别的稳定性。为了解决这一问题,我们提出了文本-视觉协同Token缓存(TVCache),一种用于高效VLA推理的免训练框架。TVCache基于文本-视觉信息焦点过滤注意力头,以改善任务相关且物理一致的视觉定位。同时,我们引入了一种由文本-视觉熵差异引导的复用层选择机制,以避免缓存不稳定的表示并改善缓存资源分配。在四个代表性VLA模型、两个仿真基准和真实世界机器人任务上的大量实验证明了TVCache的有效性和通用性。在匹配的Token保留率下,TVCache在相当的计算成本下,始终比现有的VLA缓存方法提高任务成功率。在OpenVLA-OFT上,其在12.5%保留率下比VLA-Cache平均成功率提高了多达14.5个百分点,同时相对于全Token推理减少了2.45倍的FLOPs。
cs.CV / 293 / 2609.34325

DORA: Dynamic Online Reinforcement Agent for Token Pruning in Vision Transformers

DORA:用于视觉Transformer中Token剪枝的动态在线强化学习代理
He, Kaixuan, Chen, Song, Kang, Yi
Abstract
Vision Transformers (ViTs) incur quadratic self-attention cost in the number of tokens. Most token-reduction methods adapt token identities within a prescribed layer-wise compression schedule, or search a static mask offline, and thus limit online adaptation of when and how much to prune. We propose DORA (Dynamic Online Reinforcement Agent), which learns an input-adaptive pruning policy itself for frozen ViTs. At each eligible block, a hierarchical actor decides whether to prune, how many tokens to remove, and which tokens to remove from each image's evolving representation. Because early deletions change the states observed by later decisions, DORA formulates pruning as a finite-horizon Markov decision process. Complete-prefix shadow evaluations convert final-prediction fidelity into localized per-step credit, while closed-loop accuracy feedback adjusts the fidelity penalty toward a shared accuracy-drop target. A privileged critic and all shadow computations are training-only. Deployment retains the frozen backbone and a lightweight actor that applies hard deletion and packed variable-length FlashAttention, converting token reduction into measured speedups. On ImageNet-1K with DeiT-Base, DORA reduces FLOPs by 38.4% relative to the uncompressed backbone within one percentage point of accuracy loss. Averaged across four ViT-type backbones at matched accuracy, DORA uses 13.2% fewer FLOPs and achieves 32.4% higher throughput than the corresponding per-backbone baseline means. Under zero-shot transfer to ImageNet-A, these gains widen to 20.3% and 45.6%, respectively.
Chinese Translation
视觉Transformer(ViT)的自注意力成本随token数量呈二次方增长。大多数token缩减方法在预设的逐层压缩调度内调整token标识,或离线搜索静态掩码,从而限制了在线自适应地决定何时剪枝以及剪枝多少。我们提出DORA(动态在线强化学习代理),它为冻结的ViT学习一种输入自适应的剪枝策略。在每个符合条件的块中,一个分层actor决定是否剪枝、移除多少token以及从每个图像不断演化的表示中移除哪些token。由于早期删除会改变后续决策观察到的状态,DORA将剪枝建模为有限视界的马尔可夫决策过程。完整前缀影子评估将最终预测保真度转换为局部的每步信用,而闭环准确率反馈则将保真度惩罚调整至共享的准确率下降目标。特权critic和所有影子计算仅在训练时使用。部署时保留冻结的主干网络和一个轻量级actor,该actor应用硬删除和打包的变长FlashAttention,将token缩减转化为可测量的加速。在ImageNet-1K上使用DeiT-Base,DORA在准确率损失一个百分点以内,相对于未压缩的主干网络减少了38.4%的FLOPs。在匹配准确率下,对四种ViT类型主干网络平均,DORA比相应的各主干基线均值少用13.2%的FLOPs,并实现32.4%更高的吞吐量。在零样本迁移到ImageNet-A时,这些增益分别扩大到20.3%和45.6%。
cs.CV / 294 / 2609.34330

MiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM Inference

MiCo:通过语义擦除建模的互信息覆盖优化,用于高效 MLLM 推理
Wang, Tinghao, Guo, Yichen, Zhang, Qizhe, Zhang, Yuan, Ouyang, Weimin, Huang, Rui, Cao, Jiajun, Chen, Sixiang, Jiang, Hao, Wu, Jixian, Lu, Zheng, Zhu, Bofan, Li, Renyuan, Zhang, Shanghang
Abstract
Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visual tokens results in high computational costs. While many methods have been proposed to reduce the number of visual tokens, most of them rely on heuristics and are prone to discarding substantial visual information during pruning, leading to degradation in model performance. In this work, by using a semantic erasure model, we derive a general mutual information coverage objective from task log-loss and propose MiCo, a training-free two-stage pruning method. MiCo first uses visual signals to select a representative candidate pool before visual tokens enter the language model, then performs task-aware subset selection within it. At each stage, suitable observable proxies instantiate the derived objective as a monotone submodular coverage function, which MiCo greedily optimizes under the token budget. MiCo is evaluated on diverse MLLMs ranging from 7B to 13B parameters across a broad range of image and video benchmarks spanning general visual reasoning, fine-grained OCR and grounding, hallucination detection, and long-video understanding. MiCo consistently achieves the best performance across nearly all evaluated models under all pruning ratios. On LLaVA-NEXT-13B, MiCo uses only 5.6% visual tokens, retains 97.5% of baseline performance, and achieves a 3.8-fold inference speedup. Our experiments demonstrate the effectiveness of MiCo and our mutual information coverage objective for visual token pruning.
Chinese Translation
多模态大语言模型(MLLMs)在多模态理解方面展现出令人瞩目的性能,但处理大量视觉 token 会导致高昂的计算成本。尽管已提出许多减少视觉 token 数量的方法,但大多数依赖于启发式方法,在剪枝过程中容易丢弃大量视觉信息,导致模型性能下降。在本工作中,我们通过使用语义擦除模型,从任务对数损失中推导出一个通用的互信息覆盖目标,并提出 MiCo,一种免训练的两阶段剪枝方法。MiCo 首先在视觉 token 进入语言模型之前,利用视觉信号选择一个具有代表性的候选池,然后在其中执行任务感知的子集选择。在每个阶段,合适的可观测代理将推导出的目标实例化为单调子模覆盖函数,MiCo 在 token 预算下贪婪地优化该函数。MiCo 在从 7B 到 13B 参数的各种 MLLM 上进行了评估,涵盖广泛的图像和视频基准测试,包括通用视觉推理、细粒度 OCR 与定位、幻觉检测和长视频理解。在所有剪枝率下,MiCo 在几乎所有评估模型上始终取得最佳性能。在 LLaVA-NEXT-13B 上,MiCo 仅使用 5.6% 的视觉 token,保留了 97.5% 的基线性能,并实现了 3.8 倍的推理加速。我们的实验证明了 MiCo 以及我们的互信息覆盖目标在视觉 token 剪枝方面的有效性。
cs.CV / 295 / 2609.34335

SkillPE: Creativity-Oriented Cinematic Skill Evolution for Text-to-Video Prompt Engineering

SkillPE:面向创意的电影化技能演化用于文本到视频提示工程
Huang, Yanwei, Zhu, Mingxuan, Li, Shujie, Liu, Shiyuan, Zhang, Yuanxing, Narechania, Arpit
Abstract
Achieving high-quality, cinematic results in text-to-video generation remains challenging for non-experts, whose prompts often lack professional narrative and creative design. We propose SkillPE, a prompt engineering (PE) framework that evolves reusable cinematic skills from expert-authored seeds. SkillPE represents shot logic, composition, lighting, sound design, and other filmmaking cues in a fine-grained format, and retrieves movie references categorized as resonators (good matches), dissonants (weak matches), and divergents (creatively useful near-misses). The first two refine when and how a skill should be applied, while divergents inspire alternative cinematic realizations at different degrees of modification while preserving the user intent. Candidate skills are assessed through generated videos along prompt fidelity, cinematic quality, narrative appeal, and creativity to construct the final skill libraries. Experiments on StoryEval and VBench show improvements of up to 1.40 points over the strongest external baseline and 0.51 points over seed skills on 7-point four-dimensional evaluation, while remaining competitive on benchmark-native metrics. Overall, SkillPE offers a practical approach to balancing fidelity and creativity in cinematic text-to-video generation. Code is available at https://github.com/Ais0n/SkillPE .
Chinese Translation
在文本到视频生成中实现高质量、电影化的结果对于非专家而言仍然具有挑战性,他们的提示通常缺乏专业叙事和创意设计。我们提出了SkillPE,一个提示工程(PE)框架,它从专家编写的种子中演化出可重用的电影化技能。SkillPE以细粒度格式表示镜头逻辑、构图、灯光、声音设计和其他电影制作线索,并检索电影参考,分类为共鸣者(良好匹配)、不和谐者(弱匹配)和发散者(有创意价值的近失)。前两者细化了技能应在何时以及如何应用,而发散者则在保持用户意图的同时,以不同程度的修改激发替代的电影化实现。候选技能通过生成的视频在提示保真度、电影质量、叙事吸引力和创造力方面进行评估,以构建最终技能库。在StoryEval和VBench上的实验表明,在7点四维评估中,比最强外部基线提高了最多1.40分,比种子技能提高了0.51分,同时在基准原生指标上保持竞争力。总体而言,SkillPE提供了一种在电影化文本到视频生成中平衡保真度和创造力的实用方法。代码可在https://github.com/Ais0n/SkillPE获取。
cs.CV / 296 / 2609.34346

E-WAVE: Event-based Continuous Optical Flow via Warping-Aligned Visual Encoding

E-WAVE:通过扭曲对齐视觉编码实现基于事件的连续光流
Wu, Jiale, Bai, Xiaoyang, Yu, Haoming, Chen, Yiwei, Peng, Yifan, Xu, Weiwei
Abstract
Temporally dense optical flow is essential for dynamic perception in immersive VR/AR systems, where rapid head, hand, and object motion must be continuously captured and tracked. Existing frame-based optical flow estimation methods are constrained by the tradeoff between temporal resolution and computational cost; while event cameras, with their high temporal resolution and energy efficiency, serve as a natural solution to the dilemma. However, event-based approaches commonly rely on correlation volumes to capture pairwise voxel correspondences, which incur substantial memory and computation overhead. We present E-WAVE, a correlation-free framework for high-temporal-resolution (HTR) optical flow estimation from event streams. Instead of constructing all-pairs correlation volumes, E-WAVE employs global attention mechanism to model long-range feature dependencies and performs trajectory guided feature warping using B\'ezier curve. Through iterative updates, it predicts trajectories that allow for querying at arbitrary timestamps without repeated inference. Experiments on MultiFlow and DSEC-Flow demonstrate a 25% lower trajectory error and comparable endpoint flow estimation accuracy relative to state-of-the art baselines. Additional evaluations on self-captured data using a head-mounted prototype validate that E-WAVE remains robust under challenging real-world conditions.
Chinese Translation
时间密集的光流对于沉浸式VR/AR系统中的动态感知至关重要,因为这类系统必须连续捕捉和跟踪快速的头部、手部和物体运动。现有的基于帧的光流估计方法受限于时间分辨率与计算成本之间的权衡;而事件相机凭借其高时间分辨率和能效,成为解决这一困境的自然方案。然而,基于事件的方法通常依赖相关体积来捕捉成对体素对应关系,这会带来大量内存和计算开销。我们提出E-WAVE,一种无需相关体积的框架,用于从事件流中进行高时间分辨率(HTR)光流估计。E-WAVE不构建全配对相关体积,而是采用全局注意力机制建模长程特征依赖,并使用贝塞尔曲线进行轨迹引导的特征扭曲。通过迭代更新,它预测出轨迹,从而允许在任意时间戳进行查询而无需重复推理。在MultiFlow和DSEC-Flow上的实验表明,相对于最先进基线,轨迹误差降低25%,端点光流估计精度相当。使用头戴式原型在自采集数据上的额外评估验证了E-WAVE在具有挑战性的真实世界条件下仍保持稳健。
cs.CV / 297 / 2609.34363

SyncRA: Learning Temporal Correspondence in Omni-Modal Models

SyncRA:学习全模态模型中的时间对应关系
Xu, Zelong, Li, Yan, Hu, Wenhe, Hu, Xiyang
Abstract
Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in incorrect audio-visual pairings. We diagnose this problem through controlled temporal swaps, revealing that model answers do not reliably follow changes in these pairings. To address it, we propose Synchrony-Guided Representation Alignment (SyncRA), a lightweight method for strengthening temporal correspondence between audio and vision. Specifically, SyncRA contrasts intermediate audio-visual representations within each video, aligning matching moments while separating mismatched ones to capture local temporal correspondence within a shared global context. The objective derives supervision directly from existing input timing, requiring no additional annotations and leaving inference unchanged. We evaluate SyncRA across four open omni-modal models spanning different sizes and architectures on five public video benchmarks. SyncRA consistently outperforms answer-only fine-tuning across all model-benchmark combinations, while substantially improving the ability to track changing audio-visual pairings in controlled evaluations. These results demonstrate that lightweight, targeted supervision can effectively strengthen temporal correspondence and translate into broad improvements in audio-visual question answering.
Chinese Translation
近期全模态模型在音频和视觉输入的感知方面表现出色,但往往难以将同一时刻听到的内容与看到的内容联系起来。这种时间对应关系上的弱点可能导致模型将语音线索与错误的视觉场景关联,从而基于不正确的视听配对生成看似合理的答案。我们通过受控的时间交换来诊断这一问题,揭示出模型答案并不能可靠地跟随这些配对的变化。为了解决该问题,我们提出了同步引导表示对齐(Synchrony-Guided Representation Alignment, SyncRA),一种用于增强音频与视觉之间时间对应关系的轻量级方法。具体而言,SyncRA 在每段视频内对中间音频-视觉表示进行对比,对齐匹配的时刻并分离不匹配的时刻,以在共享的全局上下文中捕捉局部时间对应关系。该目标直接从现有输入时序中获取监督,无需额外标注,且不改变推理过程。我们在五个公开视频基准上,对四个不同规模和架构的开源全模态模型评估了 SyncRA。在所有模型-基准组合上,SyncRA 始终优于仅答案微调,同时在受控评估中显著提升了跟踪变化的视听配对的能力。这些结果表明,轻量级、有针对性的监督可以有效地增强时间对应关系,并转化为音频-视觉问答的广泛改进。
cs.CV / 298 / 2609.34367

Rate-Distortion Adaptive Primitive Selection for Omnidirectional Gaussian Splatting

面向全向高斯泼溅的率失真自适应基元选择
Cheng, Yulong, Bao, Youneng, Zhou, Junfeng, Li, Mu, Wen, Jie
Abstract
Learned image codecs (LICs) achieve high reconstruction quality, but their decoding speed is often insufficient for immersive virtual reality (VR). Gaussian splatting (GS) codecs render much faster, yet still lag in reconstruction quality and typically decide primitive allocation without considering the coding cost of each primitive. We introduce OIC-GS, an omnidirectional GS codec with a new hierarchical HEALPix primitive grid representation. Gaussian primitives are anchored at predefined spherical locations, eliminating explicit coordinate coding. Finer levels refine their coarser ancestors, naturally supporting coarse-to-fine reconstruction and layered transmission. The predefined grid also enables efficient viewport decoding by selecting only view-relevant primitives. We further introduce a lightweight entropy model for quantized primitives and optimize the codec under a spherical rate-distortion objective. Primitives with insufficient rate-distortion benefit are automatically removed when their quantized opacity becomes zero, allowing OIC-GS to adapt both primitive density and level of detail without a fixed primitive budget. A single bitstream supports full-sphere, viewport-dependent, and progressive decoding. The first viewport reaches final quality after decoding only 52% of the bitstream, and is then rendered at 1,270 FPS. On a 100-image omnidirectional benchmark, OIC-GS outperforms all evaluated GS codecs, reducing WS-PSNR BD-rate by 49.6% over GaussianImage++ and 68.6% over SGI, which uses a learned entropy model.
Chinese Translation
学习型图像编解码器(LICs)实现了高重建质量,但其解码速度往往不足以支撑沉浸式虚拟现实(VR)。高斯泼溅(GS)编解码器的渲染速度快得多,但在重建质量上仍然落后,并且通常在不考虑每个基元编码代价的情况下决定基元分配。我们提出OIC-GS,一种全向GS编解码器,具有新型层次化HEALPix基元网格表示。高斯基元锚定在预定义的球面位置上,消除了显式坐标编码。更精细的层级细化其更粗糙的祖先,自然地支持由粗到细的重建和分层传输。预定义网格还通过仅选择与视图相关的基元来实现高效视口解码。我们进一步引入用于量化基元的轻量级熵模型,并在球面率失真目标下优化编解码器。量化不透明度变为零时,率失真收益不足的基元被自动移除,使得OIC-GS无需固定基元预算即可自适应调整基元密度和细节层次。单一比特流支持全球面、视口相关和渐进式解码。第一个视口在仅解码52%的比特流后达到最终质量,然后以1,270 FPS渲染。在100图像的全向基准测试上,OIC-GS优于所有评估的GS编解码器,将WS-PSNR BD-rate相比GaussianImage++降低49.6%,相比SGI(使用学习熵模型)降低68.6%。
cs.CV / 299 / 2609.34371

From Static to Dynamic: On-Policy Distillation from Image to Video Diffusion Models

从静态到动态:从图像到视频扩散模型的同策略蒸馏
Jiang, Bingqing, Luo, Li, Yu, Zichao, Han, Yujin, Su, Zhaolong, Zou, Difan
Abstract
On-policy distillation (OPD) specializes pretrained video diffusion models through teacher supervision along the student's own generation trajectory. Although large video models are natural teachers, developing specialized video experts can require costly video data and training, while querying them incurs substantially higher latency than querying image experts. More readily available and cheaper to query, image experts offer a cost-effective alternative, particularly for largely temporal-agnostic capabilities such as aesthetics and OCR that admit frame-level supervision. However, heterogeneous image and video latent spaces prevent direct supervision of intermediate student states, while image experts lack cross-frame motion supervision, making temporal consistency vulnerable to frame-level improvements. In this paper, we propose MILD, a Motion-Preserving Image-to-Video Latent Distillation framework that transfers specialized image expertise while preserving pretrained video dynamics. MILD uses a learnable linear connector that aligns student latent states and predicted updates with those of image experts, enabling supervision transfer across heterogeneous latent spaces. We further constrain image-guided corrections around the pretrained student's predictions to preserve video dynamics and incorporate an optical-flow-based motion reward to improve motion quality and temporal consistency. Across specialized image experts and multiple video-student backbones, our method consistently outperforms video-teacher OPD baselines, with further studies demonstrating effective transfer across connector designs and heterogeneous architectures. These results establish image-to-video distillation as an effective route to improving video generation by drawing on the diverse and evolving capabilities of the image-generation ecosystem.
Chinese Translation
同策略蒸馏(OPD)通过沿着学生模型自身生成轨迹的教师监督,来专门化预训练的视频扩散模型。尽管大型视频模型是天然的教师,但开发专门的视频专家可能需要昂贵的视频数据和训练,而查询它们的延迟远高于查询图像专家。更易获得且查询成本更低,图像专家提供了一种成本有效的替代方案,特别是对于诸如美学和OCR等主要与时间无关的能力,这些能力允许帧级监督。然而,异构的图像和视频潜在空间阻碍了对学生中间状态的直接监督,而图像专家缺乏跨帧运动监督,使得时间一致性容易受到帧级改进的影响。在本文中,我们提出了MILD,一个运动保持的图像到视频潜在蒸馏框架,该框架在保留预训练视频动态的同时,迁移专门的图像专家知识。MILD使用一个可学习的线性连接器,将学生的潜在状态和预测更新与图像专家的对齐,从而实现跨异构潜在空间的监督迁移。我们进一步约束图像引导的校正围绕预训练学生的预测,以保留视频动态,并引入基于光流的运动奖励,以提高运动质量和时间一致性。在专门的图像专家和多个视频学生骨干网络上,我们的方法始终优于视频教师OPD基线,进一步的研究表明,该方法在不同连接器设计和异构架构上实现了有效的迁移。这些结果确立了图像到视频蒸馏作为一种有效途径,通过利用图像生成生态系统中多样且不断发展的能力来改进视频生成。
cs.CV / 300 / 2609.34378

Marathoner: Ultra-Long-Horizon Autonomous Intelligence

Marathoner:超长周期自主智能
Ruiyang, Zhang, Jinpeng, Ou, Yifan, Xie, Jingang, Zhou, Lirui, Pan, Qingpei, Guo, Zhedong, Zheng
Abstract
Humans naturally possess the ability to work persistently toward long-term goals. Given a challenging task, humans can continuously work for months or even years to accomplish a specific objective. In this paper, we propose Marathoner, an autonomous agentic model possessing the ability of ultra-long-horizon execution. Specifically, we propose a comprehensive post-training pipeline to instill this critical capability into base model. For Ultra-Long-Horizon Task Synthesis, we leverage major release PRs containing 1000+ lines of new code from diverse GitHub repositories as the primary source for synthesizing challenging task-level data. Additionally, we introduce Multi-Task Chaining, which chains multiple generated tasks into a single more challenging task, enabling the synthesis of tasks with frontier-level difficulty. For rejection sampling finetuning, we combine strong teacher model with diverse harnesses to generate trajectories on our synthesized tasks and conduct supervised finetuning on base model with rejection sampled trajectories. For reinforcement learning, cold-started model performs real-world execution through harnesses in independent sandboxes during rollout process, effectively facilitating the acquisition of genuine ultra-long-horizon execution capability. We further propose a novel reward strategy, Later Stage Bonus Reward, which explicitly encourages model to perform meaningful maneuvers during later stages of execution. Through extensive evaluation on 5 benchmarks containing ultra-long-horizon tasks, Marathoner achieves consistent and substantial performance improvements over base model and even surpasses performance of strong proprietary model. Further analysis shows that Marathoner can consistently work for 10+ hours and conduct 1000+ tool calls on highly challenging tasks.
Chinese Translation
人类天生具备为长期目标坚持不懈工作的能力。面对具有挑战性的任务,人类可以连续工作数月甚至数年以完成特定目标。本文提出Marathoner,一种具备超长周期执行能力的自主智能体模型。具体而言,我们提出了一套全面的后训练流程,将这一关键能力注入基础模型。对于超长周期任务合成,我们利用来自不同GitHub仓库的包含1000+行新代码的主要发布拉取请求(PR)作为合成具有挑战性的任务级数据的主要来源。此外,我们引入了多任务链(Multi-Task Chaining),将多个生成的任务链接成一个更具挑战性的单一任务,从而能够合成具有前沿难度的任务。对于拒绝采样微调,我们结合强大的教师模型与多样化的执行框架(harnesses),在我们合成的任务上生成轨迹,并使用拒绝采样的轨迹对基础模型进行监督微调。对于强化学习,冷启动模型在展开过程中通过独立沙箱中的执行框架执行真实世界任务,有效促进获得真正的超长周期执行能力。我们进一步提出了一种新颖的奖励策略——后期阶段奖励(Later Stage Bonus Reward),明确鼓励模型在执行后期阶段执行有意义的操作。通过在包含超长周期任务的5个基准上的广泛评估,Marathoner相比基础模型取得了持续且显著的性能提升,甚至超越了强大的专有模型。进一步分析表明,Marathoner可以在极具挑战性的任务上持续工作10+小时并进行1000+次工具调用。
机器学习 (Machine Learning)
300
cs.LG / 1 / 2609.31630

Replay in the Silent Degrees of Freedom: Continual Learning Without an Offline Phase

在静默自由度中的重放:无需离线阶段的持续学习
Yanhai, Zhang
Abstract
Replay-based continual learning almost always consolidates in a dedicated offline phase or by interleaving replayed samples with the input stream, whereas brains also consolidate during wakefulness through local sleep, brief use-dependent off-periods of individual circuits. We ask whether a network trained by local, biologically constrained rules can consolidate with no offline phase at all. An isolation rule confines replay updates to hidden synapses invisible to the current input under k-winner-take-all dynamics, with optimiser state advanced only inside the mask; a refractory rotation rule makes units that have just fired sit out the next competition, widening the consolidable set; a homeostatic pressure and a relative-novelty gate decide when replay bursts fire and when rotation runs. This inverts the usual direction of non-interfering continual learning: the hidden computation on the current input is held invariant (exactly on the proven channels, and for all but 0.3% of waking samples per update elsewhere) while past memories are written into the degrees of freedom the current batch leaves unused. On class-incremental split-MNIST the system reaches 91.6+-0.3% with no offline phase, at or above the best offline-night schedule on two held-out splits, tied with DER++ and above experience replay, ER-ACE, A-GEM and unmasked local replay; in a single pass it leads DER++ (91.8% against 90.1%) while the night falls to 76.9%. The advantage is largest at small buffers and gives way to the backpropagation references at large ones; on split CIFAR-10 the system leads offline rehearsal and experience replay but trails ER-ACE and DER++. Rotation carries most of the gain; isolation adds the invariance guarantee. The mechanism is not tied to the local rule: under the same schedule a backpropagation network with k-WTA hidden layers gains from rotation, and isolation is again free on top of it.
Chinese Translation
基于重放的持续学习几乎总是在专用的离线阶段或通过将重放样本与输入流交错来进行巩固,而大脑也在清醒期间通过局部睡眠,即单个回路的短暂使用依赖性关闭期,进行巩固。我们探究一个由局部、生物约束规则训练的网络是否可以在完全没有离线阶段的情况下进行巩固。一种隔离规则将重放更新限制在 k-winner-take-all 动力学下对当前输入不可见的隐藏突触上,优化器状态仅在掩码内部推进;一种不应期轮换规则使刚刚发放的单元退出下一次竞争,从而扩大可巩固集合;稳态压力和相对新颖性门控决定重放爆发何时触发以及轮换何时运行。这反转了非干扰持续学习的通常方向:当前输入的隐藏计算保持不变(在已证明的通道上精确,而在其他地方每次更新对除0.3%外的所有清醒样本保持精确),同时过去的记忆被写入当前批次未使用的自由度中。在类增量 split-MNIST 上,系统在无离线阶段的情况下达到 91.6±0.3%,在两个留出划分上达到或超过最佳离线夜间调度,与 DER++ 持平,并高于经验重放、ER-ACE、A-GEM 和未掩码局部重放;在单次遍历中,它领先 DER++(91.8% 对 90.1%),而夜间则降至 76.9%。这种优势在小缓冲区时最大,在大缓冲区时则让位于反向传播参考方法;在 split CIFAR-10 上,系统领先于离线排练和经验重放,但落后于 ER-ACE 和 DER++。轮换带来了大部分增益;隔离增加了不变性保证。该机制不依赖于局部规则:在相同调度下,具有 k-WTA 隐藏层的反向传播网络也能从轮换中获益,而隔离再次免费叠加于其上。
cs.LG / 2 / 2609.31631

OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit

OMP-MoE:通过正交匹配追踪实现混合专家大语言模型的高效专家剪枝
Li, Dezhi, Li, Lujun, Zhu, Qiyuan, Gu, Hao, Liu, Bei, Han, Sirui, Guo, Yike
Abstract
Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. To address these challenges, we present OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs. Based on observations of expert contribution patterns, we reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit. Specifically, our method first treats individual expert contributions as dictionary atoms and selects experts that greedily minimize reconstruction error with linear computational complexity. Then, we optimize cross-layer expert allocation through a water-filling strategy that accounts for both reconstruction quality and routing stability. Finally, we introduce OMP-MoE{\dag}, an adaptive inference mechanism that dynamically adjusts expert activation based on energy prediction. Comprehensive experiments on Qwen, DeepSeek-V2, GPT-OSS, and Mixtral MoE demonstrate consistent improvements over existing methods at 25-50% pruning ratios. For Qwen3-30B-A3B at 50% compression, we retain 93.3% of original performance, achieving 33$\times$ faster search and 1.55$\times$ inference speedup. Codes will be available after acceptance.
Chinese Translation
混合专家(MoE)模型能够高效扩展大语言模型,但由于巨大的内存需求,面临严峻的部署挑战。现有的剪枝方法要么带来高昂的搜索成本,要么忽略专家之间的动态相互依赖关系。为了解决这些挑战,我们提出了OMP-MoE,一种新颖的免训练压缩框架,用于减少基于MoE的大语言模型中的专家冗余。基于对专家贡献模式的观察,我们将剪枝问题重新表述为通过正交匹配追踪求解的稀疏信号重建任务。具体而言,我们的方法首先将单个专家的贡献视为字典原子,并选择能够以线性计算复杂度贪婪地最小化重建误差的专家。然后,我们通过注水策略优化跨层专家分配,该策略同时考虑重建质量和路由稳定性。最后,我们引入了OMP-MoE†,一种自适应推理机制,可根据能量预测动态调整专家激活。在Qwen、DeepSeek-V2、GPT-OSS和Mixtral MoE上的全面实验表明,在25-50%的剪枝率下,该方法相比现有方法有持续改进。对于Qwen3-30B-A3B在50%压缩率下,我们保留了93.3%的原始性能,实现了33倍的搜索加速和1.55倍的推理加速。代码将在录用后公开。
cs.LG / 3 / 2609.31632

EEGAgentBench: Benchmarking LLM Agents on Short- and Long-Horizon EEG Analysis

EEGAgentBench:对短时程与长时程EEG分析中的LLM智能体进行基准测试
Wu, Huyu, Weng, Weining, Liu, Yuchen, Chen, Yiqiang, Gu, Yang
Abstract
Electroencephalography (EEG) analysis is evolving from short-segment classification toward long-horizon interpretation that demands iterative evidence accumulation, multi-step reasoning, and coordinated use of specialized signal-processing tools. Although large language models (LLMs) have recently shown promise as autonomous agents for EEG analysis, existing EEG agentic evaluations remain fragmented, covering limited tasks over narrow temporal horizons with inconsistent protocols, and providing no comprehensive assessment of agents' reasoning, tool-use, and workflow construction capabilities. To address this gap, we propose \textbf{EEGAgentBench}, a unified benchmark for systematically evaluating LLM agents on short- and long-horizon EEG analysis. EEGAgentBench spans six representative EEG applications ranging from knowledge question answering to sleep staging. It encompasses signal durations from 2 seconds to nearly 23 hours, with prediction targets ranging from class labels to event intervals and epoch-level sequences. This design supports unified evaluation across knowledge reasoning, short-horizon interpretation, long-horizon event detection, and sequential understanding. The benchmark further provides 10 deterministic EEG analysis tools that expose only task-relevant signal measurements. Agents must therefore select tools autonomously, accumulate evidence iteratively, and construct multi-step workflows. For evaluation, we benchmark 29 frontier LLMs from 15 model families. Results demonstrate that EEGAgentBench effectively distinguishes agent capabilities beyond model scale and inference cost, while revealing substantial limitations of current LLM agents in long-horizon EEG analysis, particularly in sustained evidence accumulation and multi-step reasoning.
Chinese Translation
脑电图(EEG)分析正从短片段分类向长时程解读演进,这要求迭代式证据积累、多步推理以及协同使用专用信号处理工具。尽管大语言模型(LLMs)近期展现出作为EEG分析自主智能体的潜力,但现有的EEG智能体评估仍然零散,涵盖的任务有限、时间范围狭窄且协议不一致,并且未对智能体的推理、工具使用和工作流构建能力提供全面评估。为弥补这一空白,我们提出EEGAgentBench,一个用于系统评估LLM智能体在短时程和长时程EEG分析中表现的统一基准。EEGAgentBench涵盖六种代表性EEG应用,从知识问答到睡眠分期。它涵盖的信号时长从2秒到近23小时,预测目标从类别标签到事件区间和epoch级序列。该设计支持在知识推理、短时程解读、长时程事件检测和序列理解方面进行统一评估。该基准还提供了10种确定性EEG分析工具,这些工具仅暴露与任务相关的信号测量值。因此,智能体必须自主选择工具、迭代积累证据并构建多步工作流。在评估中,我们对来自15个模型家族的29个前沿LLM进行了基准测试。结果表明,EEGAgentBench能够有效区分超越模型规模和推理成本的智能体能力,同时揭示当前LLM智能体在长时程EEG分析中的显著局限,尤其是在持续证据积累和多步推理方面。
cs.LG / 4 / 2609.31633

Enhancing generalization in endwall film cooling prediction: Incorporating the superposition principle into transformer-based neural operators

增强端壁气膜冷却预测的泛化能力:将叠加原理融入基于Transformer的神经算子
Wang, Qineng, Song, Liming, Liu, Tianyuan, Guo, Zhendong
Abstract
In this study, a physics-enhanced neural operator framework is proposed to enhance the generalization prediction ability of the cooling layout of a turbine endwall with variable number of film holes. Specifically, inspired by the film cooling superposition principle, we propose a film cooling prediction model, namely superposition-based deep neural operator (SDNO), that divides the endwall temperature field prediction into two stages. In the first stage, the cooling layout of a turbine endwall is divided into several sub-parts with randomly assigned film holes, and a Transformer-based neural operator network, namely Calculate Net, is designed to predict the temperature field of each sub-part. Then, in the second stage, another neural operator network, i.e., Super Net, is trained to combine the temperature fields predicted by Calculate Net for each sub-part and obtain the superposed temperature field of the full cooling layout. Additionally, instead of directly taking the film cooling contours as pixel plots, a signed distance function (SDF) which is sensitive to the variable locations of cooling holes, is designed to encode the location information of cooling holes. Furthermore, the proposed endwall film cooling prediction model is trained with the samples that changing the number of film holes from 1-5 with variable locations. Then, the trained prediction shows excellent generalization prediction ability, which can accurately predict the film effectiveness of the cooling layout with 10-20 film cooling holes that are unseen in the training samples. The proposed SDNO also improves prediction accuracy relative to the fully supervised baseline. With the above, the effectiveness of our proposed prediction model has been well demonstrated.
Chinese Translation
在本研究中,提出了一种物理增强的神经算子框架,以增强具有可变数量气膜孔的涡轮端壁冷却布局的泛化预测能力。具体而言,受气膜冷却叠加原理的启发,我们提出了一种气膜冷却预测模型,即基于叠加的深度神经算子(SDNO),它将端壁温度场预测分为两个阶段。在第一阶段,将涡轮端壁的冷却布局划分为若干具有随机分配气膜孔的子部分,并设计了一种基于Transformer的神经算子网络,即Calculate Net,用于预测每个子部分的温度场。然后,在第二阶段,训练另一个神经算子网络,即Super Net,将Calculate Net预测的每个子部分的温度场进行组合,并获得整个冷却布局的叠加温度场。此外,不是直接将气膜冷却等值线作为像素图,而是设计了一种对冷却孔位置变化敏感的有符号距离函数(SDF)来编码冷却孔的位置信息。进一步,所提出的端壁气膜冷却预测模型使用气膜孔数量从1-5变化且位置可变的样本进行训练。然后,训练后的预测显示出优异的泛化预测能力,能够准确预测训练样本中未见过的具有10-20个气膜冷却孔的冷却布局的气膜冷却效率。与全监督基线相比,所提出的SDNO还提高了预测精度。综上所述,我们提出的预测模型的有效性得到了很好的证明。
cs.LG / 5 / 2609.31634

Symmetry-quotient Flatness and Generalization

对称商平坦性与泛化
Miyagawa, Taiki
Abstract
This paper develops a theorem-level pipeline in symmetry-quotient settings: quotient linear stability implies quotient flatness, quotient flatness implies input smoothness, and input smoothness yields generalization under local covering assumptions. Flatness is often associated with generalization, and Stochastic Gradient Descent (SGD) is frequently viewed as implicitly biased toward flat solutions. However, standard flatness measures are typically defined in the raw parameter space and are therefore not invariant under function-preserving symmetries such as positive rescaling. We develop a symmetry-aware theory of quotient flatness, quotient linear stability, input smoothness, and generalization on quotient spaces of neural-network parameters. For square loss and models equipped with function-preserving group actions, we define quotient flatness as the trace of the Hessian of the empirical loss on the regular quotient manifold. We show that quotient flatness controls input smoothness through a quotient-space analogue of the flatness-to-smoothness argument. We also prove that one-step mean-square quotient linear stability of the linearized SGD dynamics implies an explicit quotient-flatness bound in terms of the batch size and learning rate, and extend this analysis to higher-order tensor moments. Finally, under local covering and boundedness assumptions, we derive population generalization bounds in terms of quotient flatness and, consequently, in terms of quotient linear stability.
Chinese Translation
本文在对称商设定下建立了一条定理级流程:商线性稳定性蕴含商平坦性,商平坦性蕴含输入平滑性,而在局部覆盖假设下,输入平滑性可带来泛化。平坦性常与泛化相关联,随机梯度下降(SGD)也常被视为隐式偏向平坦解。然而,标准平坦性度量通常在原始参数空间中定义,因此对于保持函数的对称性(如正缩放)并不具有不变性。我们发展了一套对称感知的理论,研究神经网络参数商空间上的商平坦性、商线性稳定性、输入平滑性和泛化。对于平方损失和配备保函数群作用的模型,我们将商平坦性定义为经验损失在正则商流形上的 Hessian 的迹。我们证明,通过平坦性到平滑性论证的商空间类似版本,商平坦性控制输入平滑性。我们还证明,线性化 SGD 动力学的一步均方商线性稳定性蕴含以批量大小和学习率表示的显式商平坦性界,并将该分析推广到高阶张量矩。最后,在局部覆盖和有界性假设下,我们推导出以商平坦性表示、并因而以商线性稳定性表示的总体泛化界。
cs.LG / 6 / 2609.31635

What Next-Event Accuracy Cannot See: Closed-Loop Evaluation of Emergency Department Trajectory Simulators

下一事件准确率无法看到什么:急诊科轨迹模拟器的闭环评估
Low, Zhen Xuen Brandon
Abstract
Clinical trajectory models are usually evaluated by next-event accuracy on observed histories. Simulation is different: models must condition on their own generated events, allowing errors to compound. Although this problem is well known in sequence modelling, it has not been systematically quantified for clinical trajectory simulators. We developed EDSim-Bench to evaluate this failure mode using 425,028 MIMIC-IV-ED stays, with external replication on MC-MED, and release the evaluation protocol and scoring code. Starting from held-out visit prefixes, models generate the remainder of each visit and are evaluated on termination, event composition, timing, conditional fidelity, and occupancy forecasting, with a train-only order-3 n-gram as a reference baseline. Despite next-event accuracies within 0.001, three neural architectures behaved very differently under rollout. Across seeds, one Transformer recipe ranged from 0.43 to 0.96 in termination score and from 4.2- to 137-fold the divergence of the n-gram; no prefix-trained neural model approached the n-gram on termination or event composition. Inference-time interventions improved termination but did not jointly recover composition and timing. Supervising every eligible sequence position rather than only the final prefix position was associated with one to two orders of magnitude lower divergence across Transformer, GRU, and LSTM models, with the pattern persisting under model scaling, temporal shift, and external-site evaluation. Nevertheless, even the best model generated visits approximately half as long as observed, and model rankings reversed on occupancy forecasting, a downstream quantity relevant to bed management. These results show that next-event accuracy is insufficient to evaluate clinical trajectory simulators and motivate closed-loop evaluation across seeds, rollout criteria, and downstream tasks.
Chinese Translation
临床轨迹模型通常通过在观察到的历史记录上的下一事件准确率进行评估。模拟则不同:模型必须基于自身生成的事件进行条件推断,这会导致误差累积。尽管这个问题在序列建模中众所周知,但尚未针对临床轨迹模拟器进行系统量化。我们开发了 EDSim-Bench,利用 425,028 条 MIMIC-IV-ED 记录来评估这种失败模式,并在 MC-MED 上进行外部复现,同时发布了评估协议和评分代码。从留出的就诊前缀开始,模型生成每次就诊的剩余部分,并针对终止、事件构成、时间、条件保真度和占用预测进行评估,使用仅训练集的三阶 n-gram 作为参考基线。尽管下一事件准确率在 0.001 以内,三种神经网络架构在 rollout 下表现差异很大。在不同随机种子下,一种 Transformer 方案的终止得分在 0.43 到 0.96 之间,与 n-gram 的差异达到 4.2 到 137 倍;没有经过前缀训练的神经模型在终止或事件构成上接近 n-gram。推理时干预改善了终止,但未能同时恢复构成和时间。监督每个符合条件的序列位置,而不仅仅是最终前缀位置,与 Transformer、GRU 和 LSTM 模型中低一到两个数量级的差异相关,该模式在模型缩放、时间偏移和外部站点评估下持续存在。然而,即使最好的模型生成的访问长度也大约只有观察到的一半,并且模型排名在占用预测(与床位管理相关的下游量)上发生逆转。这些结果表明,下一事件准确率不足以评估临床轨迹模拟器,并推动跨随机种子、rollout 标准和下游任务的闭环评估。
cs.LG / 7 / 2609.31636

Grounding Vision-Language Models in Driving Semantics: A Multi-Dataset Predicate Framework for Explainable Reasoning

将视觉语言模型接地于驾驶语义:一个可解释推理的多数据集谓词框架
Chouai, Mohamed, Okumus, Fazli Faruk, Kugele, Stefan
Abstract
Vision-language models are increasingly used for driving-scene understanding, yet the semantic relations expressed in their outputs are often difficult to verify against the underlying traffic situation. This paper introduces a deterministic multi-dataset predicate framework that derives driving-scene semantics from measurable geometric, kinematic, temporal, map, and traffic-control evidence. Dataset-specific interfaces are used only to recover the required scene information, while predicate definitions remain unchanged across nuPlan and nuScenes and are materialised in a common Predicate Knowledge Graph. Quantitative semantic validation against manually annotated predicate relations on 200 scenarios from each dataset yields macro F1 scores of 0.94 on nuPlan and 0.93 on nuScenes, with an average cross-dataset difference of 0.02 across the shared predicates. The Predicate KG is further evaluated using a frozen LLaVA-OneVision-7B model on the nine NuPlanQA subtasks. Predicate grounding achieves the highest accuracy among the evaluated visual-input conditions in seven of nine NuPlanQA subtasks, including Traffic Light (53.2% to 71.5%), Situation Assessment (76.2% to 86.1%), and Action Recommendation (82.9% to 89.0%). Weather/Lighting remains essentially unchanged (89.4% vs. 88.8%), consistent with the absence of corresponding predicates, while Predicate KG only input outperforms metadata-only input in eight of nine subtasks. The results show that deterministic predicates provide a consistent and traceable semantic representation and, under oracle grounding, can reduce visual dependence for reasoning tasks covered by the predicate vocabulary.
Chinese Translation
视觉语言模型正越来越多地用于驾驶场景理解,但其输出中表达的语义关系往往难以针对底层交通状况进行验证。本文介绍了一种确定性多数据集谓词框架,该框架从可测量的几何、运动学、时间、地图和交通控制证据中推导出驾驶场景语义。数据集特定的接口仅用于恢复所需的场景信息,而谓词定义在nuPlan和nuScenes上保持不变,并在一个公共的谓词知识图谱中实现。针对每个数据集的200个场景中手动标注的谓词关系进行定量语义验证,在nuPlan上获得了0.94的宏F1分数,在nuScenes上获得了0.93的宏F1分数,共享谓词的平均跨数据集差异为0.02。Predicate KG进一步使用冻结的LLaVA-OneVision-7B模型在九个NuPlanQA子任务上进行了评估。在九个NuPlanQA子任务中的七个中,谓词接地在所评估的视觉输入条件中实现了最高准确率,包括交通灯(53.2%至71.5%)、态势评估(76.2%至86.1%)和动作推荐(82.9%至89.0%)。天气/照明基本保持不变(89.4% vs. 88.8%),与相应谓词的缺失一致,而仅Predicate KG输入在九个中的八个子任务中优于仅元数据输入。结果表明,确定性谓词提供了一致且可追溯的语义表示,并且在oracle接地条件下,可以减少谓词词汇表所覆盖的推理任务对视觉的依赖。
cs.LG / 8 / 2609.31637

FIDAL: Diversity-Aware Federated Active Learning Under Real-World Distribution Shifts

FIDAL:真实世界分布偏移下的多样性感知联邦主动学习
Gaviria, David Dueñas, Albarqouni, Shadi
Abstract
Federated learning enables collaborative model training across institutions without centralizing data, yet high annotation costs, domain shifts, and class imbalance remain major obstacles, especially when irrelevant out-of-distribution (OOD) samples dilute the labeled data. Existing active learning methods target uncertainty or diversity within in-distribution (ID) data and overlook unknown samples in federated clinical settings. We propose FIDAL, an open-set federated active learning framework that combines calibrated global-local evidential uncertainty, support-set diversity weighting, and adaptive OOD rejection. The rejection gate thresholds a foundation-model Gaussian-coverage signal per client and per round with Otsu's criterion, so that highly informative ID samples are queried while irrelevant outliers are excluded without any hand-tuned threshold. Evaluated on three multi-center medical imaging benchmarks (dermatology, histopathology, and mammography with organically occurring artifacts) in realistic open-set scenarios, FIDAL outperforms detector-based open-set methods by up to about 12 percentage points of balanced accuracy and is the only method on the accuracy-ID purity Pareto front of all three benchmarks. At an equal query budget it spends at least 1.3 times fewer annotations on OOD samples than every accuracy-matched baseline, saving an estimated 7-29 hours of expert reading on the mammography benchmark. By labeling only a fraction of the data pool, it matches or exceeds fully supervised performance across modalities. These results highlight the value of integrating uncertainty, diversity, and OOD rejection in open-set federated active learning for medicine.
Chinese Translation
联邦学习支持跨机构协作训练模型而无需集中数据,但高昂的标注成本、域偏移和类别不平衡仍然是主要障碍,尤其是当无关的分布外(OOD)样本稀释了标注数据时。现有的主动学习方法主要关注分布内(ID)数据中的不确定性或多样性,而忽略了联邦临床环境中的未知样本。我们提出了 FIDAL,一个开放集联邦主动学习框架,它结合了校准的全局-局部证据不确定性、支持集多样性加权和自适应 OOD 拒绝。该拒绝门控使用 Otsu 准则对每个客户端和每轮的基础模型高斯覆盖信号进行阈值处理,从而查询高信息量的 ID 样本,同时排除无关的异常值,而无需任何手动调整的阈值。在三个多中心医学影像基准(皮肤病学、组织病理学和具有自然发生伪影的乳腺 X 线摄影)上,在真实的开放集场景中进行评估,FIDAL 在平衡准确率上超过基于检测器的开放集方法最多约 12 个百分点,并且是唯一在所有三个基准的准确率-ID 纯度帕累托前沿上出现的方法。在相同的查询预算下,它在 OOD 样本上花费的标注量至少比每个准确率匹配的基线少 1.3 倍,在乳腺 X 线摄影基准上估计节省 7-29 小时的专家阅读时间。通过仅标注数据池的一小部分,它在不同模态上达到或超过了全监督性能。这些结果凸显了在医学开放集联邦主动学习中整合不确定性、多样性和 OOD 拒绝的价值。
cs.LG / 9 / 2609.31638

Energy-aware frugal Bayesian optimization

能耗感知的节俭贝叶斯优化
Plat, Gaston, Saves, Paul, Bartoli, Nathalie, Lefebvre, Thierry, Morlier, Joseph
Abstract
Modern design optimization frameworks aim first and foremost for models with the most accurate predictions without balancing computational overhead. It remains a reason why scaled architecture and multidisciplinary design optimization problems are difficult to address, even with sample-efficient Bayesian optimizers. In this paper, a metric quantifying the computational energy footprint is introduced within a Bayesian optimization framework to guide the parameter setting of a model towards configurations that balance both performance and frugality. The computer experiments highlighted existing tradeoffs between optimum convergence and the underlying energy footprint, and sometimes resulted in both a better-found optimum and lower energy consumption.
Chinese Translation
现代设计优化框架首要追求具有最准确预测的模型,而未权衡计算开销。这仍然是为何即使采用样本高效的贝叶斯优化器,规模化架构和多学科设计优化问题仍难以解决的原因之一。本文在贝叶斯优化框架内引入了一个量化计算能量足迹的指标,以引导模型的参数设置朝向兼顾性能和节俭性的配置。计算机实验凸显了最优收敛与潜在能量足迹之间现有的权衡,有时还同时带来更好的最优解和更低的能耗。
cs.LG / 10 / 2609.31639

When Does Domain Adaptation Help on Physical Vibration Sensors? A Held-Out-Bearing Study of Neural-Operator and Convolutional Models

域适应何时对物理振动传感器有帮助?一项关于神经算子和卷积模型的留出轴承研究
Nagaswetha, Kumbha, Pathak, Rabi
Abstract
Diagnosing rolling-element bearing faults from vibration is a canonical physical-sensing task and a widely used benchmark for domain adaptation under operating-condition shift. Accuracies above 99 percent are commonly reported, but under evaluation splits that place the same physical bearing in both training and test. We revisit the task under a held-out-bearing protocol, assigning every bearing unit entirely to either the training or the test set, and find that source-only transfer is far weaker than such numbers suggest: on a change of shaft speed it reaches only $0.36$, against a target-supervised ceiling of 0.97. We then study what governs transfer. Treating computed order tracking, a shaft-angle resampling that places fault frequencies at fixed shaft orders independent of running speed, as a controlled change of representation, we find that a Fourier Neural Operator raises source-only transfer from $0.36$ to $0.61$ on the speed shift, where the fault peaks move, while a convolutional network of matched feature dimension stays near chance in both representations. The representation also decides whether unsupervised alignment can work: with the same normalized RBF-MMD loss and no target labels, the operator reaches 0.71 in the frequency domain but 0.95 in the order domain, within 0.02 of the target-supervised ceiling and above $0.86$ on every held-out bearing fold. Once the representation is right, a small label budget adds little. These results indicate that, for this task, the input representation rather than the alignment method decides whether adaptation helps. A second dataset, whose held-out units are fault diameters rather than bearings, shows that the same protocol exposes failures that even a target-supervised model cannot avoid.
Chinese Translation
从振动信号中诊断滚动轴承故障是一项经典的物理传感任务,也是工况变化下域适应的广泛使用基准。通常报告准确率超过99%,但这些准确率是在将同一物理轴承同时置于训练和测试集的评估划分下得到的。我们在留出轴承协议下重新审视该任务,将每个轴承单元完全分配给训练集或测试集,并发现仅源域迁移远弱于这些数字所暗示:在轴转速变化时,其仅达到0.36,而目标监督上限为0.97。然后我们研究是什么决定了迁移。将计算阶次跟踪(一种轴角重采样,将故障频率置于固定的轴阶次,与运行速度无关)视为受控的表示变化,我们发现傅里叶神经算子将转速变化下的仅源域迁移从0.36提高到0.61,其中故障峰值移动,而特征维度匹配的卷积网络在两种表示中都接近随机。表示还决定了无监督对齐是否可行:在相同的归一化RBF-MMD损失且无目标标签的情况下,算子在频域达到0.71,但在阶次域达到0.95,与目标监督上限相差0.02以内,并在每个留出轴承折上超过0.86。一旦表示正确,少量标签预算增加很少。这些结果表明,对于该任务,输入表示而非对齐方法决定了适应是否有帮助。第二个数据集,其留出单元是故障直径而非轴承,表明相同的协议暴露了即使目标监督模型也无法避免的失败。
cs.LG / 11 / 2609.31640

Measure Learning at Steady State: A BIRD-SQL Formula 1 Case Study

在稳态下衡量学习:一个BIRD-SQL Formula 1案例研究
Bajaj, Manoj
Abstract
Continual Learning Bench scores learning as short-horizon gain versus a reset baseline and finds naive full-context ICL strongest among the memories it tested. We treat ICL as one learning system and score it on a longer shared-world schedule. Steady-state learning is the gap versus baseline on a pre-set late window (last 40 of 174 BIRD-SQL formula-1 questions). We split the score into exploration efficiency (SQL probes), task reward (hits), and delivery cost (API dollars and context size). On gpt-5.6-luna, late probes fall from 4.6-5.6 to 0.95 while hits rise only modestly and ICL context grows to about 95k tokens with cost roughly doubling. Short-horizon gain understates the late probe saving and misses the cost inversion, so we find that unbounded ICL is a poor candidate for the learning mechanism.
Chinese Translation
Continual Learning Bench将学习评分定义为相对于重置基线的短视界增益,并发现朴素全上下文ICL在其测试的记忆中最强。我们将ICL视为一个学习系统,并在更长的共享世界时间表上对其进行评分。稳态学习是在预设的后期窗口(174个BIRD-SQL formula-1问题中的最后40个)上相对于基线的差距。我们将评分分为探索效率(SQL探测)、任务奖励(命中)和交付成本(API美元和上下文大小)。在gpt-5.6-luna上,后期探测从4.6-5.6下降到0.95,而命中仅略有上升,ICL上下文增长到约95k个token,成本大约翻倍。短视界增益低估了后期探测的节省,并忽略了成本反转,因此我们发现无界ICL是学习机制的一个糟糕候选。
cs.LG / 12 / 2609.31641

Product-Aware Deterministic Rounding for Quantized Matrix Multiplication

面向量化矩阵乘法的乘积感知确定性舍入
Sao, Piyush, Miniskar, Narasinga, Valero-Lara, Pedro, Teranishi, Keita, Seal, Sudip
Abstract
Scalar rounding decisions interact through matrix multiplication. We study deterministic product-aware rounding after scales, clipping bounds, and grids are fixed, with each active scalar choosing between adjacent levels. For dynamic activation rounding, null-space reduction preserves the relaxed product while leaving at most $r$ fractional decisions, where $r$ is the rank of the active gap-weighted weight block. Conditional-expectation completion gives a deterministic polynomial-time algorithm with squared product error at most $ \mathrm{OPT}_{\mathrm{dyn}}+r\nu_{\max}^2/4$, where $\mathrm{OPT}_{\mathrm{dyn}}$ is the best admissible error and $\nu_{\max}$ is the largest row norm of that block. For reusable static weights, the exact expected product-loss metric is the uncentered input second moment with fixed output bias; free bias recalibration yields the centered covariance. Exact optimization is NP-hard even at rank one. In balanced blocks with $K=1024$ and $r=16$, conditional- expectation completion attains a dither-normalized median error of $0.010$, compared with $0.899$ for round-to-nearest. Clipping-aware initialization reduces median normalized error by a factor of $43.4$ at ten-percent clipping. Held-out Digits experiments show that retaining the input mean or correcting the output bias improves median product error over round-to-nearest in all four tested bit-width and calibration-size settings.
Chinese Translation
标量舍入决策通过矩阵乘法相互作用。我们研究在缩放因子、裁剪边界和网格固定后的确定性乘积感知舍入,其中每个活跃标量在相邻级别之间进行选择。对于动态激活舍入,零空间约简保持松弛乘积不变,同时最多留下 $r$ 个分数决策,其中 $r$ 是活跃的间隙加权权重块的秩。条件期望补全给出了一种确定性多项式时间算法,其平方乘积误差至多为 $\mathrm{OPT}_{\mathrm{dyn}}+r\nu_{\max}^2/4$,其中 $\mathrm{OPT}_{\mathrm{dyn}}$ 是最佳容许误差,$\nu_{\max}$ 是该块的最大行范数。对于可重用的静态权重,精确的期望乘积损失度量是固定输出偏置下的非中心输入二阶矩;自由偏置重校准产生中心化协方差。即使秩为一,精确优化也是 NP 难的。在 $K=1024$ 且 $r=16$ 的平衡块中,条件期望补全达到抖动归一化中值误差 $0.010$,而最近舍入为 $0.899$。裁剪感知初始化在百分之十裁剪下将中值归一化误差降低 $43.4$ 倍。留出 Digits 实验表明,在所有四种测试的位宽和校准大小设置下,保留输入均值或校正输出偏置均比最近舍入改善了中值乘积误差。
cs.LG / 13 / 2609.31643

Information Design Against Gaming and Learning Adversaries

针对博弈型与学习型对手的信息设计
Gaikwad, Madhava
Abstract
A principal who deploys a binary classifier with an abstention option must decide which queries the mechanism abstains on. The right choice depends on the adversary. A gaming adversary already knows the classifier and tries to manipulate features across the boundary, so the principal does best by abstaining on queries close to that boundary. The same boundary-localizing rule is the worst possible choice against a learning adversary who does not know the classifier: each abstention now tells the adversary that the boundary is nearby, which is enough to drive a binary search. We analyze this tension. The two natural defenses, abstaining at a fixed rate and abstaining near the boundary, are Blackwell-incomparable: neither can be simulated by post-processing the other's responses. The number of queries needed to reconstruct the boundary to error $\eps$ is $\tilde\Theta(d/\eps)$ under the first defense and $\Theta(d \log(1/\eps))$ under the second, where $d$ is the VC dimension of the classifier family and $\tilde\Theta$ suppresses factors polylogarithmic in $d$ and $1/\eps$. The first rate is a worst case over query distributions; no reconstruction algorithm can close the gap at the distributions that attain it. We characterize the Pareto frontier between the two defense objectives, and confirm both rates on seven binary-classification tasks spanning tabular, image, and language-model-feature inputs: label-plus-counterfactual access extracts the boundary with up to $200\times$ fewer queries than a published label-only baseline.
Chinese Translation
一个部署带有弃权选项的二分类器的委托人必须决定机制在哪些查询上弃权。正确的选择取决于对手。博弈型对手已经知道分类器,并试图操纵特征跨越边界,因此委托人在边界附近的查询上弃权效果最好。同样的边界定位规则对于不知道分类器的学习型对手来说是最糟糕的选择:每次弃权现在都告诉对手边界就在附近,这足以驱动二分搜索。我们分析这种张力。两种自然的防御——以固定比率弃权和在边界附近弃权——是 Blackwell 不可比较的:两者都不能通过后处理另一方的响应来模拟。在第一种防御下,将边界重构到误差 $\eps$ 所需的查询数为 $\tilde\Theta(d/\eps)$,在第二种防御下为 $\Theta(d \log(1/\eps))$,其中 $d$ 是分类器族的 VC 维,$\tilde\Theta$ 抑制了 $d$ 和 $1/\eps$ 的多对数因子。第一个速率是查询分布上的最坏情况;没有重构算法能在达到该速率的分布上缩小差距。我们刻画了两种防御目标之间的帕累托前沿,并在跨越表格、图像和语言模型特征输入的七个二分类任务上验证了这两个速率:标签加反事实访问提取边界所需的查询数比已发表的仅标签基线最多减少 $200\times$ 倍。
cs.LG / 14 / 2609.31644

MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning

MaD-RL:通过强化学习匹配分布以校准大语言模型
Kulkarni, Sourabh, Vepuri, Ksheeraj Sai, Demir, Basar, Bohrer, Jason, Shen, Emily, Chen, Jianfa, Jiang, Nan, Jain, Ankit, Subramanyam, Harihar, Singh, Mannat, Nagpal, Chirag
Abstract
Reinforcement learning (RL) is widely used in language-model post-training to maximize rewards assigned to individual model outputs, such as scores from binary verifiers or reward models trained on human feedback. However, applications such as synthetic-data generation, fairness-related constraint satisfaction, and policy exploration require controlling the distribution of outputs across model generations rather than only maximizing expected reward. We propose a general RL-based framework for \textit{Distribution Matching} allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution. Empirically, we demonstrate that dominant post-training recipes such as Group Relative Policy Optimization (GRPO) reduce output diversity by concentrating policy probability towards a single mode. Entropy regularization and sampling temperature can improve the spread of the distribution but have constrained effectiveness, limited to apply only in token space and toward uniform distributions. We show that prior work in this area is a specific case of Distribution Matching involving the $L_2$ divergence. We then propose reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justification. Finally, we demonstrate the effectiveness of our approach on a set of experiments involving mathematical reasoning and programming.
Chinese Translation
强化学习(RL)被广泛用于语言模型后训练,以最大化分配给单个模型输出的奖励,例如来自二元验证器的分数或基于人类反馈训练的奖励模型。然而,合成数据生成、公平性相关的约束满足以及策略探索等应用,需要控制模型生成过程中输出的分布,而不仅仅是最大化期望奖励。我们提出了一个通用的基于强化学习的分布匹配框架,允许将模型输出的潜在类别属性的分布匹配到指定的目标分布。实证上,我们证明了主流的后训练方法(如组相对策略优化(GRPO))通过将策略概率集中到单一模式来降低输出多样性。熵正则化和采样温度可以改善分布的扩展,但效果有限,仅限于在标记空间和朝向均匀分布应用。我们展示了该领域的先前工作是分布匹配的一个特例,涉及L2散度。然后,我们提出了针对其他散度(如KL和Jensen-Shannon散度)的奖励函数,并通过理论论证来支持它们。最后,我们在涉及数学推理和编程的一系列实验中证明了我们方法的有效性。
cs.LG / 15 / 2609.31645

STAR: Adaptive Spatial-Temporal Normalization for Unified Microservice Incident Management

STAR: 面向统一微服务事件管理的自适应时空归一化
Miao, Xinhua, Zhu, Linyu, Yang, Bowei, Cai, Zhengong
Abstract
Automated incident management in large-scale microservice systems relies on learning robust representations from multimodal observability data, including metrics, logs, and traces. Although recent self-supervised frameworks enable unified modeling for anomaly detection (AD), failure triage (FT), and root cause localization (RCL), they often struggle with non-stationary temporal dynamics and heterogeneous service dependency structures. In this paper, we propose STAR, a Spatial-Temporal Adaptive Representation learning framework that explicitly addresses these challenges through adaptive normalizations. STAR introduces two tightly coupled mechanisms: Temporal Adaptive Normalization (TAN), which dynamically normalizes multivariate time series using multi-scale temporal context, and Spatial Adaptive Normalization (SAN), which performs structure-aware normalization over service dependency graphs. Unlike prior methods that treat normalization as static or task-agnostic, STAR formulates it as a learnable, context-conditioned transformation aligned with the intrinsic properties of microservice systems. The resulting adaptive representations are integrated into a unified self-supervised framework, enabling end-to-end unsupervised support for AD, FT, and RCL tasks. Extensive experiments on two real-world microservice benchmarks demonstrate that STAR consistently outperforms all state-of-the-art baselines, yielding significant and stable improvements across all three tasks. Our results highlight adaptive normalization as a principled and effective mechanism for robust multimodal representation learning in complex software systems.
Chinese Translation
大规模微服务系统中的自动化事件管理依赖于从多模态可观测性数据(包括指标、日志和追踪)中学习鲁棒的表示。尽管最近的自监督框架能够实现异常检测(AD)、故障分诊(FT)和根因定位(RCL)的统一建模,但它们往往难以处理非平稳时间动态和异构服务依赖结构。在本文中,我们提出了STAR,一个时空自适应表示学习框架,通过自适应归一化明确应对这些挑战。STAR引入了两个紧密耦合的机制:时间自适应归一化(TAN),它利用多尺度时间上下文动态归一化多变量时间序列;以及空间自适应归一化(SAN),它在服务依赖图上执行结构感知归一化。与先前将归一化视为静态或任务无关的方法不同,STAR将其表述为一种可学习的、以上下文为条件的变换,与微服务系统的内在属性相一致。由此产生的自适应表示被集成到一个统一的自监督框架中,能够为AD、FT和RCL任务提供端到端的无监督支持。在两个真实世界的微服务基准上进行的大量实验表明,STAR始终优于所有最先进的基线,在三个任务上均取得显著且稳定的改进。我们的结果强调了自适应归一化作为一种原则性且有效的机制,用于复杂软件系统中鲁棒的多模态表示学习。
cs.LG / 16 / 2609.31647

Typed Temporal Interaction Features for Simulation-Backed Forecasting of Open-Source Game Release Incidents

用于模拟支持的开源游戏发布事件预测的类型化时序交互特征
Alkobaisi, Shayma, Ali, Anas
Abstract
Open-source video-game quality depends on inter-actions among code, assets, configuration, tests, contributors, and issue workflows, yet conventional defect predictors usually flatten or omit these relations. We investigate release-level forecasting of a quality incident within thirty days using GAMEQUALGRAPH-Pilot, a typed temporal feature pipeline with calibrated risk estimates and effort-aware ranking. Because the accessible OS-SGameBench materials do not provide manually audited release dates and outbreak labels, the executed evaluation is explicitly simulation-backed rather than an empirical claim about real games. Five seeded worlds each contain 120 projects and 24 releases, with project-disjoint validation and future cross-project testing. The pilot obtains an AUPRC of 0.520, AUROC of 0.673, Brier score of 0.207, and 29.68% effort-aware recall at a twenty-percent testing budget. Its closest local comparator, Static-Hetero-Reimpl, reaches 0.522 AUPRC; the -0.002 difference is not statistically significant after Holm correction. Inference requires 0.023 milliseconds per release in the measured environment. Ablations and controlled missingness, drift, engine, project-size, alert-threshold, and attribution analyses expose where typed interactions help and where they fail. Results support the reproducibility of the proposed protocol, not deployment effectiveness. Real OSSGameBench release reconstruction, stratified label audits, and official graph-model comparisons remain mandatory before journal submission or operational use in practice. This boundary protects research integrity and supports credible evaluation.
Chinese Translation
开源电子游戏质量取决于代码、资产、配置、测试、贡献者和问题工作流之间的交互,然而传统的缺陷预测器通常将这些关系扁平化或忽略。我们使用GAMEQUALGRAPH-Pilot研究三十天内质量事件的发布级预测,这是一个类型化时序特征管道,具有校准的风险估计和工作量感知排序。由于可访问的OS-SGameBench材料不提供人工审核的发布日期和爆发标签,所执行的评估明确是模拟支持的,而不是关于真实游戏的经验性主张。五个种子世界各包含120个项目和24个发布,具有项目不相交的验证和未来跨项目测试。该试点获得0.520的AUPRC、0.673的AUROC、0.207的Brier分数,以及在20%测试预算下29.68%的工作量感知召回率。其最接近的本地比较器Static-Hetero-Reimpl达到0.522 AUPRC;在Holm校正后,-0.002的差异不具有统计学显著性。在测量环境中,推理每次发布需要0.023毫秒。消融实验和受控的缺失性、漂移、引擎、项目规模、警报阈值和归因分析揭示了类型化交互在何处有帮助以及在何处失败。结果支持所提出协议的可重复性,而不是部署有效性。在期刊提交或实际运营使用之前,真实的OSSGameBench发布重建、分层标签审核和官方图模型比较仍然是强制性的。这一边界保护了研究诚信并支持可信的评估。
cs.LG / 17 / 2609.31648

Energy Vision--Language--Action: A Controlled Multimodal Benchmark for Intent-Conditioned Residential Energy Management

Energy Vision-Language-Action:一个用于意图条件住宅能源管理的受控多模态基准
Saoud, Lyes Saad, Doukhi, Oualid, Reihani, Ehsan, Sepasi, Saeed, Lee, Deok Jin, Ayyash, Moussa, Ghorbani, Reza
Abstract
Vision-Language-Action (VLA) models are studied mainly in robotics, where visual observations and language instructions are mapped to physical actions. This paper introduces Energy Vision-Language-Action (EVLA), a controlled multimodal benchmark for intent-conditioned residential energy management. EVLA frames battery scheduling as a multimodal trajectory-prediction problem in which an RGB energy-field representation, a numerical operating state, and a natural-language objective are mapped to a 16-step battery-action trajectory generated by a finite-horizon sampling-based reference generator. Source windows are derived from public residential electrical-load data, while electricity price, battery state of charge, indoor temperature, and time of day are generated benchmark metadata. A hidden operating regime is encoded only through energy-field texture, enabling paired visual changes while the explicit numerical state is fixed. Crossing 439,203 retained base windows with three hidden regimes and five language objectives yields 6,588,045 multimodal instances. An initial study evaluates 36 configurations over three training seeds using fixed subsets of 5,000 training, 500 validation, and 500 test instances. In the MobileNet-family comparison, removing processed language increases trajectory mean-squared error from 0.3856 +/- 0.0039 to 0.8628 +/- 0.0001, whereas removing vision yields 0.3843 +/- 0.0013, comparable to the full model. The results show strong asymmetry in modality use: the processed-language pathway is strongly associated with prediction quality, while the current RGB pathway provides no aggregate error advantage. These results characterize the fixed pilot subset and executed protocol rather than full-benchmark training. EVLA provides a controlled setting for studying how semantic intent and latent context influence residential energy-action prediction.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)模型主要在机器人领域被研究,其中视觉观察和语言指令被映射为物理动作。本文介绍了 Energy Vision-Language-Action (EVLA),一个用于意图条件住宅能源管理的受控多模态基准。EVLA 将电池调度构建为一个多模态轨迹预测问题,其中 RGB 能量场表示、数值运行状态和自然语言目标被映射为由有限时域基于采样的参考生成器生成的 16 步电池动作轨迹。源窗口来自公开的住宅电力负荷数据,而电价、电池荷电状态、室内温度和时间是生成的基准元数据。一个隐藏的运行模式仅通过能量场纹理进行编码,从而在显式数值状态固定的情况下实现成对的视觉变化。将 439,203 个保留的基础窗口与三种隐藏模式和五种语言目标交叉组合,得到 6,588,045 个多模态实例。一项初步研究使用 5,000 个训练、500 个验证和 500 个测试实例的固定子集,在三个训练种子上评估了 36 种配置。在 MobileNet 系列比较中,移除处理后的语言会使轨迹均方误差从 0.3856 +/- 0.0039 增加到 0.8628 +/- 0.0001,而移除视觉则得到 0.3843 +/- 0.0013,与完整模型相当。结果表明模态使用存在强烈不对称性:处理后的语言通路与预测质量强相关,而当前的 RGB 通路没有提供总体误差优势。这些结果描述的是固定的试点子集和已执行的协议,而非全基准训练。EVLA 为研究语义意图和潜在上下文如何影响住宅能源动作预测提供了一个受控环境。
cs.LG / 18 / 2609.31656

From Phase Transition to Systemic Failure: A Decoupled Analytics Framework for GNN Robustness

从相变到系统性失效:一种面向GNN鲁棒性的解耦分析框架
Yan, Shuai, Peng, Dan, Li, Jie, Wang, Ke
Abstract
Data quality is a major bottleneck for the reliable deployment of graph neural networks (GNNs) in real-world graph mining tasks. Among various sources of degradation, label noise and feature distribution shift (hereafter referred to as distribution shift) are two common yet fundamentally different challenges. To study their effects under controlled conditions, this paper constructs a synthetic homophilic graph regression benchmark in which the two factors can be manipulated separately. A total of 41 configurations and 410 runs are conducted to evaluate the behavior of representative GNN models under varying noise and shift conditions. The results show two distinct patterns. First, under additive label corruption, performance remains relatively stable over a broad range of noise settings and begins to deteriorate sharply only after an observed transition region around the 50 percent noise ratio. Second, under extreme feature distribution shift, all tested models suffer substantial degradation, with test MSE increasing by 48 times to 316 times and correlation dropping by 73 percent to 89 percent. These findings suggest that, in the present controlled setting, GNNs are considerably more tolerant to moderate label perturbation than to severe distribution mismatch. The study provides a controlled empirical baseline for understanding how data quality affects GNN-based graph mining systems and offers practical implications for deployment-oriented monitoring and model maintenance.
Chinese Translation
数据质量是图神经网络(GNN)在真实世界图挖掘任务中可靠部署的主要瓶颈。在各种退化来源中,标签噪声和特征分布偏移(以下简称分布偏移)是两个常见但根本不同的挑战。为了在受控条件下研究它们的影响,本文构建了一个合成同配图回归基准,其中这两个因素可以被单独操控。共进行了41种配置和410次运行,以评估代表性GNN模型在不同噪声和偏移条件下的行为。结果显示出两种不同的模式。首先,在加性标签损坏下,性能在广泛的噪声设置范围内保持相对稳定,只有在观察到约50%噪声比附近的相变区域后才开始急剧恶化。其次,在极端特征分布偏移下,所有测试模型都遭受显著退化,测试MSE增加了48倍至316倍,相关性下降了73%至89%。这些发现表明,在当前的受控设置中,GNN对适度的标签扰动比严重的分布不匹配具有更大的容忍度。该研究为理解数据质量如何影响基于GNN的图挖掘系统提供了一个受控的经验基线,并为面向部署的监控和模型维护提供了实际意义。
cs.LG / 19 / 2609.31659

Cross-Material Support Transfer for Core-Loss Prediction Under Waveform Covariate Shift

波形协变量偏移下磁芯损耗预测的跨材料支持迁移
Yao, Cong, Gong, Chunye
Abstract
Power magnetic materials are characterized on the sinusoidal and triangular waveforms that excitation hardware conveniently produces, whereas deployed converters expose cores to trapezoidal, PWM-shaped flux trajectories, so loss models must predict exactly where their training data are thinnest. The final test of the MagNet Challenge embeds a deliberately extreme instance of this characterization-deployment mismatch: for material D, trapezoids form 16.4% of the test set but only 1.4% of the training set. The 95th-percentile relative error, hereafter p95, of the best submission, built on sequential transfer learning, stalled at 15.9%, the worst among the five materials. This paper shows that the obstacle is missing information under covariate shift rather than class imbalance, and that the missing support can be borrowed from sibling materials instead of being extrapolated. Controlled experiments first refute the imbalance reading: four standard remedies fail, and raising the trapezoidal share to the test-set level degrades accuracy further. The proposed material-identity support transfer, MIST, then trains one 2784-parameter predictor jointly on all five challenge materials. Material identity enters through feature-wise linear modulation, or FiLM, the scarce material's true-label loss is reweighted, and material D receives no fine-tuning, so that the bias of its trapezoid-free training set is never re-installed. MIST lowers the five-seed material-D p95 from 20.39+/-2.03% to 12.38+/-0.92% and the trapezoidal-class p95 from 37.4+/-8.8% to 15.16+/-1.69%, surpassing the best submission with one-sixth of its parameters and no fine-tuning stage; removing material identity at matched capacity inflates the error by an order of magnitude. These results argue that scarce materials should be characterized jointly with their siblings.
Chinese Translation
功率磁性材料的表征通常是在激励硬件方便产生的正弦波和三角波波形上进行的,而实际部署的变换器则将磁芯暴露于梯形、PWM 形状的磁通轨迹,因此损耗模型必须在其训练数据最稀疏的地方进行精确预测。MagNet 挑战赛的最终测试嵌入了一个刻意极端的表征-部署不匹配实例:对于材料 D,梯形波占测试集的 16.4%,但仅占训练集的 1.4%。最佳提交(基于顺序迁移学习构建)的 95 百分位相对误差(以下简称 p95)停滞在 15.9%,是五种材料中最差的。本文表明,障碍是协变量偏移下的信息缺失,而非类别不平衡,并且缺失的支持可以从同类材料借用,而不是外推。受控实验首先驳斥了不平衡的解释:四种标准补救措施均失败,并且将梯形波比例提高到测试集水平会进一步降低准确率。然后,所提出的材料身份支持迁移(MIST)在所有五种挑战材料上联合训练一个 2784 参数的预测器。材料身份通过特征级线性调制(FiLM)进入,稀缺材料的真实标签损失被重新加权,并且材料 D 不进行微调,从而其无梯形波训练集的偏差永远不会被重新引入。MIST 将五种子材料 D 的 p95 从 20.39+/-2.03% 降至 12.38+/-0.92%,并将梯形波类别的 p95 从 37.4+/-8.8% 降至 15.16+/-1.69%,以六分之一的参数和无需微调阶段超越了最佳提交;在匹配容量下移除材料身份会使误差膨胀一个数量级。这些结果表明,稀缺材料应与其同类材料联合表征。
cs.LG / 20 / 2609.31667

Beyond the Graph: An Adaptive Meta-Learner Fuses Explainability, Weather, and Dynamics for Robust Bus ETA Prediction

超越图结构:一种自适应元学习器融合可解释性、天气与动力学,实现稳健的公交ETA预测
Payra, Pratham, Jagadish
Abstract
Accurate bus Estimated Time of Arrival (ETA) prediction is vital for urban mobility, passenger satisfaction, and transit efficiency, yet existing models falter against nonlinear spatiotemporal dynamics, data sparsity, and factors such as weather. This paper proposes HYB(nm), an adaptive hybrid ensemble framework that dynamically fuses five complementary models - a historical baseline (MST-AV), periodical temporal pattern analysis (GDRN-DFT), Koopman Neural Operators for nonlinear dynamics (KOOP-NET), weather-integrated feature-engineered neural networks (FENN), and real-time graph convolutional networks (MGCN) - via a meta-learner attuned to real-time context. Evaluated on GPS and weather data from three Kolkata bus routes comprising more than 4,000 trips, the framework leverages the individual strengths of its components (for example, the low-latency explainability of MST-AV, the weather resilience of FENN, and the network-dynamics capture of MGCN) to deliver the superior robustness of HYB(2), state-of-the-art accuracy rivalling leading graph neural networks, and balanced trade-offs in stability and efficiency across prediction horizons and operating conditions. The extensible HYB(k) architecture equips transit agencies with flexible tools, ranging from economical single models to tailored high-fidelity hybrids, advancing predictive, equitable urban transport.
Chinese Translation
准确的公交预计到达时间(ETA)预测对城市出行、乘客满意度和公交效率至关重要,但现有模型在面对非线性时空动态、数据稀疏性以及天气等因素时表现不佳。本文提出HYB(nm),一种自适应混合集成框架,通过一个适应实时上下文的元学习器,动态融合五个互补模型:历史基线(MST-AV)、周期性时间模式分析(GDRN-DFT)、用于非线性动力学的Koopman神经算子(KOOP-NET)、融合天气的特征工程神经网络(FENN)以及实时图卷积网络(MGCN)。在来自加尔各答三条公交线路、包含4000多趟行程的GPS和天气数据上评估,该框架利用其各组件的各自优势(例如,MST-AV的低延迟可解释性、FENN的天气鲁棒性,以及MGCN对网络动态的捕捉能力),实现了HYB(2)的优越稳健性、可与领先图神经网络媲美的最先进精度,并在不同预测时域和运行条件下实现稳定性与效率的平衡权衡。可扩展的HYB(k)架构为公交机构提供了灵活工具,从经济的单一模型到定制的高保真混合模型,推动具有预测性且公平的城市交通发展。
cs.LG / 21 / 2609.31669

NanoForecast v0.5: Competitive Time Series Forecasting Through Training Pipeline Optimization

NanoForecast v0.5:通过训练流程优化实现具有竞争力的时间序列预测
Kishore, Gautam
Abstract
We present NanoForecast v0.5, a 6.5M-parameter forecaster that competes with models 31x its size (TimesFM, 200M parameters) after training pipeline fixes and no architecture change. Retraining the v0.3 architecture with corrected loss-scope handling, tensor shape alignment, and wider augmentation coverage cuts overall Mean Absolute Scaled Error by 43.8% under one fixed protocol (MASE 3.030 to 1.704) on the same data and compute budget. NanoForecast v0.5 beats TimesFM on all three ETT datasets (MASE 0.676/1.110/0.287 vs. 0.705/1.360/0.545) and on exchange rate (4.317 vs. 4.383); TimesFM keeps a clear lead on the high-cardinality electricity and traffic sets. Against PatchTST (15M+ parameters, official configuration), v0.5 wins all three ETT sets. Training takes about 12 hours on a single cloud GPU (NVIDIA T4, Google Colab) and inference needs no GPU (measurements in this paper are on an Apple M4 CPU). We release all code, pretrained checkpoints, and evaluation framework under Apache 2.0 at https://github.com/eulogik/NanoForecast
Chinese Translation
我们提出了NanoForecast v0.5,一个6.5M参数的预测器,在训练流程修复且未改变架构后,可与31倍于其大小的模型(TimesFM,200M参数)竞争。在相同数据和计算预算下,使用修正的损失范围处理、张量形状对齐和更广泛的增强覆盖重新训练v0.3架构,在一个固定协议下将整体平均绝对缩放误差降低了43.8%(MASE从3.030降至1.704)。NanoForecast v0.5在所有三个ETT数据集上击败了TimesFM(MASE 0.676/1.110/0.287 vs. 0.705/1.360/0.545),并在汇率数据集上(4.317 vs. 4.383);TimesFM在高基数电力和交通数据集上保持明显领先。与PatchTST(15M+参数,官方配置)相比,v0.5赢得了所有三个ETT数据集。在单个云GPU(NVIDIA T4,Google Colab)上训练大约需要12小时,推理不需要GPU(本文中的测量是在Apple M4 CPU上进行的)。我们在Apache 2.0许可下发布了所有代码、预训练检查点和评估框架,地址为https://github.com/eulogik/NanoForecast。
cs.LG / 22 / 2609.31675

Active Causal Discovery Benchmark: Evaluating LLM Agents Under Budgeted Interventions

主动因果发现基准:在预算约束干预下评估LLM智能体
Deb, Sagar, Shah, Devam, Krishnan, Ashwanth
Abstract
We introduce the Active Causal Discovery Benchmark (ACDB), an SCM-grounded environment for evaluating whether LLM agents recover causal graph structure from observations and budget-constrained hard interventions. ACDB pairs a linear-Gaussian world generator with a fixed observe-intervene-submit API and a three-layer scoring contract that separates skeleton recovery, DAG recovery, and intervention efficiency. On the current six-level ladder, PC with a greedy active orientation heuristic is the strongest non-oracle method (directed F1 42.7%, SHD 4.79), ahead of Claude Sonnet 4.6 raw active (31.7%, 7.25) and GPT-5.4 raw active (22.9%, 9.27). The most informative diagnostic is the precision-recall decomposition: PC under-commits with high precision, LLMs over-commit with lower precision, and statistical-tool access often increases abstention rather than useful intervention. A structure-blind random DAG baseline reaches 23.6% directed F1 on this dense v0 ladder; a density probe lowers this floor to 16.9%, motivating the v1 calibration pass. The current results should therefore be read as a benchmark audit and calibration report, not as evidence that current LLMs solve active causal discovery.
Chinese Translation
我们介绍了主动因果发现基准(ACDB),这是一个基于SCM的环境,用于评估LLM智能体能否从观测数据和预算受限的硬干预中恢复因果图结构。ACDB将线性高斯世界生成器与固定的观察-干预-提交API以及三层评分契约相结合,该契约分别评估骨架恢复、DAG恢复和干预效率。在当前六级阶梯上,采用贪婪主动定向启发式的PC是最强的非oracle方法(有向F1 42.7%,SHD 4.79),优于Claude Sonnet 4.6 raw active(31.7%,7.25)和GPT-5.4 raw active(22.9%,9.27)。最具信息量的诊断是精确率-召回率分解:PC承诺不足但精确率高,LLM过度承诺且精确率较低,而对统计工具的访问往往增加弃权,而非有用的干预。一个结构盲的随机DAG基线在这个密集的v0阶梯上达到23.6%的有向F1;一个密度探测将该下限降至16.9%,从而推动了v1校准轮次。因此,当前结果应被解读为一份基准审计与校准报告,而不是当前LLM已解决主动因果发现的证据。
cs.LG / 23 / 2609.31678

When Keywords Drop but Classifiers Hold: Soft Refusals under KV Cache Compression

当关键词下降而分类器保持:KV 缓存压缩下的软拒绝
Chen, Kang, Zhou, Xiuze, Chen, Hong, Lin, Yuanguo
Abstract
KV cache compression is widely used for long context LLM inference under memory constraints, while deployed systems typically score refusals after generation with keyword filters or learned classifiers. Such monitors are intended to indicate whether a model declined a harmful request under the serving regime actually used. However, it remains unclear whether matched compression that preserves task accuracy also preserves agreement between lightweight lexical monitors and stronger refusal classifiers. We study this with a paired protocol on n=200 harmful prompts with a long filler context: each prompt is answered once under full retention and once under matched eviction after a shared prefill, and the same replies are scored by keyword heuristics, the HarmBench Llama-2-13B classifier, an auxiliary LLM judge, and humans on disagreements. On Qwen2.5-3B, keyword refusal falls from 98.0% to 80.5% (McNemar p~1e-8) while classifier refusal stays near ceiling (99.0%-99.5%) and MMLU accuracy is unchanged (50.0%); human labels predominantly follow the classifier, consistent with soft refusals. The gap is not universal and weakens under short fillers and paired SnapKV, so safety auditing under compression should rely on several judges matched to the serving context rather than on keyword rates alone.
Chinese Translation
KV 缓存压缩广泛用于内存受限下的长上下文 LLM 推理,而部署系统通常在使用关键词过滤器或学习到的分类器生成后对拒绝进行评分。此类监控器旨在指示模型是否在实际使用的服务机制下拒绝了一个有害请求。然而,尚不清楚在保持任务准确性的匹配压缩下,轻量级词汇监控器与更强的拒绝分类器之间的一致性是否也能保持。我们通过一个配对协议在 n=200 个带有长填充上下文的有害提示上研究了这个问题:每个提示在完整保留和共享预填充后匹配驱逐的情况下各回答一次,并且相同的回复由关键词启发式、HarmBench Llama-2-13B 分类器、辅助 LLM 裁判以及人类在分歧上进行评分。在 Qwen2.5-3B 上,关键词拒绝从 98.0% 下降到 80.5%(McNemar p~1e-8),而分类器拒绝保持在接近上限水平(99.0%-99.5%),MMLU 准确率保持不变(50.0%);人类标签主要跟随分类器,与软拒绝一致。这种差距并非普遍存在,在短填充和配对的 SnapKV 下会减弱,因此在压缩下的安全审计应依赖于与服务上下文匹配的多个裁判,而不是仅依赖关键词率。
cs.LG / 24 / 2609.31680

Does Joint-Embedding Predictive Architecture Pretraining Help Time Series Forecasting?

联合嵌入预测架构预训练是否有助于时间序列预测?
Feng, Yutong, Liao, Bowen, Ng, See Kiong, Liang, Yuxuan
Abstract
Joint-embedding predictive architectures (JEPA) have emerged as a promising self-supervised pretraining paradigm for time series, learning representations by predicting target embeddings in latent space rather than reconstructing raw signals. Yet evidence on their benefits remains mixed, and most studies test only a single backbone or a narrow set of architectures, leaving unclear whether JEPA pretraining is a reliable improvement or one that depends heavily on the downstream model. We address this gap through a large scale evaluation of one JEPA instantiation across nine backbones and eleven benchmarks spanning temporal and spatio-temporal forecasting, the most extensive cross architecture assessment of JEPA for time series to date. We find that the benefit of this instantiation varies sharply across backbones, producing consistent gains for some architectures and consistent degradation for others, even on the same dataset. This pattern holds across both task families, indicating the variability is a general property of this instantiation rather than a dataset specific artifact worth accounting for when choosing a backbone in practice.
Chinese Translation
联合嵌入预测架构(Joint-embedding predictive architectures, JEPA)已成为时间序列领域一种有前景的自监督预训练范式,其通过学习在潜空间中预测目标嵌入,而非重建原始信号来学习表示。然而,关于其收益的证据仍不一致,且大多数研究仅测试单一主干或一组狭窄的架构,因此尚不清楚 JEPA 预训练究竟是一种可靠的改进,还是一种严重依赖下游模型的改进。我们通过对一种 JEPA 实例在九个主干和十一个基准上进行大规模评估来填补这一空白,这些基准涵盖时间预测和时空预测,是迄今为止对 JEPA 用于时间序列最广泛的跨架构评估。我们发现,该实例的收益在不同主干之间差异显著,对某些架构产生一致增益,而对另一些架构则导致一致退化,即使在同一数据集上也是如此。这一模式在两个任务家族中都成立,表明这种变异性是该实例的一般属性,而非数据集特定的偶然现象,在实践中选择主干时值得加以考虑。
cs.LG / 25 / 2609.31683

Autonomous Research Project Management as an Agent Skill: A Case Study in Exact Spectral Spatial Regression

作为智能体技能的自主研究项目管理:精确谱空间回归案例研究
Chen, Alexander, Meng, Jeffrey, Hoex, Bram, Xie, Tong
Abstract
This work presents an end-to-end demonstration of autonomous machine learning research conducted by an agent skill on consumer hardware. The demonstration evaluates an FFT-based Kernel Ridge Regression (KRR) solver for regular spatial grids using 2005 monthly NOAA Kaplan SST v2 anomaly fields on a $36 \times 72$ grid. This was autonomously executed by DeepSeek V4 Flash, orchestrated by our agent skill suite within DeepSeek Harness (DSH). Experiments were executed on CPU-only hardware (Apple M2 Pro; 78.7 s solver time, 1.57 GB peak RSS). Long-horizon state was decoupled into a file-based epic- and issue-tracking substrate. Across 74 sub-agent sessions, the agent demonstrated closed-loop scientific resilience: routing two failed hypothesis review gates back to literature retrieval, patching bootstrap indexing bugs, and executing with only four discrete human steering events. Finally, we reflect on autonomous research governance, arguing that scientific credibility requires inspectable state, falsifiable review gates, and transparent reporting of negative results, urging the machine learning community to favour agent-accessible structured formats over static PDF manuscripts.
Chinese Translation
本文展示了一个由智能体技能在消费级硬件上开展的自主机器学习研究的端到端演示。该演示评估了一种面向规则空间网格的基于FFT的核岭回归(Kernel Ridge Regression, KRR)求解器,使用2005年逐月NOAA Kaplan SST v2异常场,网格为36×72。该研究由DeepSeek V4 Flash自主执行,并由我们在DeepSeek Harness(DSH)中的智能体技能套件进行编排。实验仅在CPU硬件上执行(Apple M2 Pro;求解器时间78.7秒,峰值RSS 1.57 GB)。长时程状态被解耦为基于文件的Epic与Issue跟踪基础设施。在74个子智能体会话中,该智能体展示了闭环的科研韧性:将两次失败的假设评审门控重新路由回文献检索,修复bootstrap索引错误,并且仅通过四次离散的人工引导事件便完成执行。最后,我们反思自主研究治理,认为科学可信度需要可检查的状态、可证伪的评审门控以及对负面结果的透明报告,并呼吁机器学习社区更青睐智能体可访问的结构化格式,而非静态PDF稿件。
cs.LG / 26 / 2609.31685

What does FFN compression change downstream? Same-state causal restoration in diffusion language models

FFN压缩在下游改变了什么?扩散语言模型中的同状态因果恢复
Omar, Shaurya
Abstract
Diffusion language models (DLMs) enable flexible, parallel generation, but their iterative denoising remains computationally expensive, motivating increasingly aggressive compression. Existing compression objectives largely measure how well compressed computation approximates the original locally, but local error does not reveal which removed computations actually matter to the downstream denoising trajectory. We introduce Same-State Causal Restoration (SSR), which restores the original FFN on the exact current input reached by the compressed model and measures how the resulting trajectory changes. To our knowledge, this is the first direct measurement of the same-current-input closed-loop effect of removed FFN computation in DLM compression. Across LLaDA-8B-Instruct and Dream-v0-Instruct-7B, compressed-side state ranks this downstream effect substantially better than local NMSE at fixed denoising phase, while controlled interventions show that correction structure matters beyond magnitude. Using task-label-free calibration, SSR freezes a single restoration window for held-out inference. Under aggressive LLaDA compression, restoring only four transitions recovers 89.9% of the lost accuracy while retaining an estimated 36.8% whole-model MAC saving and outperforming an equal-budget local-error baseline. Dream further shows that restoring dense behavior and repairing the final task are distinct outcomes.
Chinese Translation
扩散语言模型(DLMs)能够实现灵活、并行的生成,但其迭代去噪仍然计算昂贵,促使人们进行越来越激进的压缩。现有的压缩目标主要衡量压缩计算在局部逼近原始计算的程度,但局部误差并不能揭示哪些被移除的计算实际上对下游去噪轨迹至关重要。我们引入了同状态因果恢复(SSR),它在压缩模型所到达的确切当前输入上恢复原始FFN,并测量由此产生的轨迹变化。据我们所知,这是首次直接测量DLM压缩中被移除FFN计算在同当前输入下的闭环效应。在LLaDA-8B-Instruct和Dream-v0-Instruct-7B上,压缩侧状态在固定去噪阶段对这种下游效应的排序显著优于局部NMSE,而受控干预表明校正结构的重要性超出了幅度。使用无需任务标签的校准,SSR为留出推理冻结了一个单一的恢复窗口。在激进的LLaDA压缩下,仅恢复四个转换即可恢复89.9%的丢失准确率,同时保留估计36.8%的整模型MAC节省,并优于同等预算的局部误差基线。Dream进一步表明,恢复密集行为和修复最终任务是不同的结果。
cs.LG / 27 / 2609.31686

3-D Emissions Mapping and Social Cost Estimation for US Domestic Aviation at West Coast Hubs

美国西海岸枢纽国内航空的三维排放映射与社会成本估算
Nia, Hesam Shafiei, MacKenzie, Don
Abstract
Existing aviation emissions inventories lack accurate trajectory data for high-resolution social cost and health impact assessment. This paper develops a 3-D emissions map by reconstructing flight trajectories for US west coast hubs to estimate regional environmental and near-airport health impacts. A physics-informed autoencoder (AE) is applied to ADS-B trajectory records for January 2025 covering US west coast hubs. The encoder combines a Convolutional Neural Network (CNN), a Bi-GRU, and a 3-D CNN with skip connection; the decoder is a Temporal Convolutional Network (TCN). It is benchmarked against a baseline-AE and cubic spline interpolation. Emissions are mapped via EUROCONTROL Base of Aircraft Data (BADA) performance tables and ICAO Engine Emissions Databank (EEDB) emission indices, with altitude corrections via Boeing Fuel Flow Method 2 (BFFM2). Social costs are quantified for all flight phases, with health impacts assessed for Landing and Takeoff cycles within 50 km of each hub. The proposed AE model outperforms both a TCN-AE and cubic spline interpolation across 5% to 50% missing rates. Monetizing the emissions inventory shows NOx produces a small net cooling effect in direct climate forcing, while accounting for 99.7% of monetized air-quality and health cost despite being under 0.4% of CO2 by mass, making it the dominant health-cost driver. To our knowledge, this is among the first studies combining AE-based trajectory reconstruction with separate spatial-temporal feature encoding and altitude-based emissions modeling to produce a regional aviation emissions inventory for air quality, climate impact and population exposure. The resulting emissions map and social cost estimates provide quantitative context for environmental impact assessment and near-airport health policy evaluation for US domestic aviation.
Chinese Translation
现有的航空排放清单缺乏准确的轨迹数据,无法进行高分辨率的社会成本和健康影响评估。本文通过重建美国西海岸枢纽的飞行轨迹,开发了一个三维排放地图,以估算区域环境和机场附近的健康影响。将物理信息自编码器(AE)应用于2025年1月覆盖美国西海岸枢纽的ADS-B轨迹记录。编码器结合了卷积神经网络(CNN)、双向门控循环单元(Bi-GRU)和带有跳跃连接的三维CNN;解码器是时间卷积网络(TCN)。将其与基线AE和三次样条插值进行了基准比较。排放通过EUROCONTROL的飞机基础数据(BADA)性能表和ICAO发动机排放数据库(EEDB)排放指数进行映射,并通过波音燃油流量方法2(BFFM2)进行高度校正。对所有飞行阶段的社会成本进行了量化,并评估了每个枢纽50公里范围内着陆和起飞循环的健康影响。所提出的AE模型在5%至50%的缺失率下均优于TCN-AE和三次样条插值。对排放清单进行货币化显示,NOx在直接气候强迫中产生小的净冷却效应,但在货币化的空气质量和健康成本中占99.7%,尽管其质量不到CO2的0.4%,使其成为主要的健康成本驱动因素。据我们所知,这是首批将基于AE的轨迹重建与独立的时空特征编码和基于高度的排放建模相结合,以生成用于空气质量、气候影响和人口暴露的区域航空排放清单的研究之一。所得的排放地图和社会成本估算为美国国内航空的环境影响评估和机场附近健康政策评估提供了定量背景。
cs.LG / 28 / 2609.31796

Same Probe, Different Numbers: Are Activation Probes Robust to Inference-Time Numerical Non-Determinism?

同一探针,不同数值:激活探针是否对推理时数值非确定性具有鲁棒性?
Khatri, Alizishaan
Abstract
Activation probes are increasingly used to monitor LLMs in deployment. A probe is typically trained under one inference configuration, then used under whatever batch size and numerical precision the serving stack uses. Because common GPU kernels are not batch-invariant and floating-point formats round differently, the activations seen at deployment are not the ones the probe was trained on. We measure what that costs for Llama-3.1-8B, Qwen3-8B and Gemma-3-4B across batch sizes 4, 8 and 16 and float32, bfloat16 and float16, training 768 probes on one configuration, evaluating each on every other, and comparing verdicts example by example. Probes are stable, but aggregate accuracy is the wrong instrument for showing it: it understates how many verdicts change by a factor of two to nine. At the prompt, accuracy never moves by more than 0.47 percentage points across 1,392 transfers and only 0.076% of verdicts change; under float32 with only the batch size varied, none of 201,960 verdicts change. During decoding the flip rate rises to 2.8%, but rows whose realised tokens matched flip in only 0.12-0.15% of cases, while rows whose tokens diverged flip in 12.9%: the cause is the text, not the arithmetic. A bfloat16 batch-size change flips the first generated token for 2.1% of rows and leaves 25% on different tokens by token 20. Flips are symmetric, Cohen's kappa stays above 0.94, and AUROC moves by at most 0.05 points. Underneath, activations move about as much as the format's rounding: a bfloat16 batch-size change perturbs them by a median relative L2 of 1e-2, roughly 8x the float16 figure. Probes absorb this; the model's own next-token argmax does not. Robustness evaluations of activation monitors should report per-example agreement rather than aggregate accuracy, separate representational noise from input change, and state the serving configuration.
Chinese Translation
激活探针越来越多地用于在部署中监测大语言模型。探针通常在一个推理配置下训练,然后在服务栈所使用的任意批大小和数值精度下使用。由于常见的GPU内核不具有批不变性,且浮点格式的舍入方式不同,部署时看到的激活值并非探针训练时所基于的激活值。我们针对Llama-3.1-8B、Qwen3-8B和Gemma-3-4B,在批大小4、8和16以及float32、bfloat16和float16下测量了其代价,在一个配置上训练768个探针,并在其他每个配置上评估,逐例比较判定结果。探针是稳定的,但总体准确率是展示这一点的错误工具:它以二到九倍的因子低估了判定结果发生变化的数量。在提示上,跨1,392次迁移,准确率变化从未超过0.47个百分点,仅0.076%的判定结果发生变化;在仅改变批大小的float32下,201,960个判定结果无一变化。在解码期间,翻转率升至2.8%,但实际生成的token匹配的行仅在0.12-0.15%的情况下发生翻转,而token发散的行有12.9%发生翻转:原因是文本,而非算术。bfloat16批大小变化使2.1%的行的第一个生成token发生翻转,到第20个token时仍有25%的行处于不同的token上。翻转是对称的,Cohen's kappa保持在0.94以上,AUROC变化至多0.05个点。底层上,激活值的变化幅度大约与格式的舍入相当:bfloat16批大小变化对它们的扰动中位相对L2为1e-2,约为float16数值的8倍。探针吸收了这一点;模型自身的下一token argmax则不能。激活监测器的鲁棒性评估应报告逐样本一致性而非总体准确率,区分表示噪声与输入变化,并说明服务配置。
cs.LG / 29 / 2609.31801

seq2cause: One Autoregressive Backbone, Four Causal Discovery Tasks in Event Sequences

seq2cause:一个自回归骨干,事件序列中的四个因果发现任务
Math, Hugo
Abstract
Complex systems such as vehicles, patients, or genomes emit discrete event sequences whose operative question is causal, not predictive: which events cause which other events, and which cause higher-level outcomes such as failures or diseases? This question decomposes along two axes -- dependency type (event $\to$ event vs.\ event $\to$ outcome) and causal scope (single sequence vs.\ population) -- yielding four structurally distinct regimes with different identifiability conditions. No existing method addresses more than one, because all assume multi-stream structure with low vocabulary, and none scales beyond a few hundred event types. We present \textsc{Seq2Cause}, a unified framework that resolves all four regimes through a single shared primitive: a pretrained autoregressive model repurposed as an amortized conditional independence testing engine requiring no task-specific retraining. We establish a prediction--causality duality: the model's excess cross-entropy simultaneously bounds causal identification error across all four regimes, so that every improvement in next-token prediction tightens causal guarantees for free. On nonlinear SCMs (vocabularies up to $8{,}000$ types) and real-world vehicle diagnostic logs ($29$K event types, $474$ failure outcomes), \textsc{Seq2Cause} is the first method to populate all four regimes at scale with a single frozen backbone. Existing methods are either inapplicable, inaccurate, or computationally intractable in this setting.
Chinese Translation
诸如车辆、患者或基因组等复杂系统会产生离散事件序列,其核心问题在于因果性而非预测性:哪些事件导致哪些其他事件,哪些事件导致更高级别的结果,如故障或疾病?该问题沿两个轴分解——依赖类型(事件$\to$事件 vs.\ 事件$\to$结果)和因果范围(单序列 vs.\ 群体)——产生四种结构不同的情形,具有不同的可识别性条件。现有方法没有一种能处理超过一种情形,因为它们都假设多流结构和低词汇量,且没有一种能扩展到几百种事件类型以上。我们提出\textsc{Seq2Cause},一个统一框架,通过一个共享的基元解决所有四种情形:一个预训练的自回归模型被重新用作摊销的条件独立性检验引擎,无需特定任务的重训练。我们建立了预测-因果对偶性:模型的超额交叉熵同时约束了所有四种情形下的因果识别误差,因此下一词预测的每一次改进都能免费地收紧因果保证。在非线性结构因果模型(词汇量高达$8{,}000$种类型)和真实车辆诊断日志($29$K事件类型,$474$种故障结果)上,\textsc{Seq2Cause}是第一个用单一冻结骨干网络在大规模上填充所有四种情形的方法。现有方法在此设置下要么不适用,要么不准确,要么计算上不可行。
cs.LG / 30 / 2609.31806

Medium-Term Multi-Resolution Electric Load Forecasting using Economic Data and Foundation Model

基于经济数据和基础模型的中期多分辨率电力负荷预测
Eloi, Lindas, Yannig, Goude, Philippe, Ciais
Abstract
Accurate medium-term, from a few months to a few years, electricity load forecasts are crucial for informed decision-making in power plant maintenance scheduling, load dispatch and price settlement. Being comprised between Long-Term Load Forecasting (LTLF) which uses mostly economic projections and appliances development scenarios, and Short-Term Load Forecasting (STLF) driven by weather, calendar and autoregressive patterns, Medium-Term Load Forecasting (MTLF) requires both extrapolation capabilities and variability modeling. Yet, it remains unclear if MTLF can benefit from economic indicators, and especially at which forecast horizon and resolution. To address these challenges we investigated the impact of socioeconomic data on predictions issued 1 month and up to 48 months in advance for France at monthly and daily resolution using a tabular Foundation Model (FM). A dataset covering 20 years of observations of electricity load, weather variables and economic features such as consumer price and production indices, electric vehicle counts or employment is created for the study. To avoid noisy data, we used a new feature selection pipeline, creating ensemble of expert models with diverse feature subsets, to demonstrate that selected economic covariates improve forecast skill by 20% over 2015-2025. This enhancement is steady across lead times and resolutions limiting the Mean Absolute Percentage Error to 4% for monthly granularity and 5% for daily granularity. Explainability of the models is investigated through feature and context importance. Results showed that the FM is limited in the context it leverages pointing towards potential computational savings with a reduced context, while feature importance of economic predictors grows with the forecast horizon. This suggests that including economic data in MTLF could bridge the gap with LTLF leading to seamless forecasts.
Chinese Translation
准确的中期(从几个月到几年)电力负荷预测对于电厂维护调度、负荷调度和价格结算的明智决策至关重要。中期负荷预测(MTLF)介于主要使用经济预测和电器发展情景的长期负荷预测(LTLF)与由天气、日历和自回归模式驱动的短期负荷预测(STLF)之间,因此需要外推能力和变异性建模。然而,目前尚不清楚MTLF能否从经济指标中受益,尤其是在哪个预测时间范围和时间分辨率上。为了应对这些挑战,我们使用表格基础模型(FM)研究了社会经济数据对法国提前1个月至48个月发布的月度与日度分辨率预测的影响。本研究创建了一个数据集,涵盖20年的电力负荷、天气变量以及经济特征(如消费者价格和生产指数、电动汽车数量或就业)的观测数据。为了避免噪声数据,我们使用了一种新的特征选择流程,创建了具有不同特征子集的专家模型集成,以证明选定的经济协变量在2015-2025年间将预测技能提高了20%。这种提升在不同提前期和分辨率下保持稳定,将月度粒度的平均绝对百分比误差限制在4%,日度粒度限制在5%。通过特征重要性和上下文重要性研究了模型的可解释性。结果表明,FM在其利用的上下文方面存在局限,这指向了通过减少上下文可能节省计算量的潜力,而经济预测变量的特征重要性随着预测时间范围的增加而增长。这表明在MTLF中纳入经济数据可以弥合与LTLF之间的差距,从而实现无缝预测。
cs.LG / 31 / 2609.31816

Relational Compression: A Framework for Relational Fidelity in Constrained Representations

关系压缩:受限表示中的关系保真度框架
Shulman, Yaniv
Abstract
What should a compressed representation preserve when the information of interest lies in relationships among elements rather than in the elements themselves? We formulate relational compression in the classical source-description-reconstruction sense, but with relational structure itself as the fidelity-bearing content. Each instance specifies the source relation, retained description, reconstructed or evaluated relation, fidelity criterion, and constrained resource. We use this interface to situate selected methods from graph summarization, spectral sparsification, similarity-preserving representation, and relational distillation within a common formulation while keeping their different reconstruction and resource assumptions explicit. We develop finite-codeword collision as one concrete realization. Same-codeword probability yields a relational geometry linking pair-specific alignment and separation to aggregate R\'enyi-2 occupancy and the spherical geometry of categorical assignments, with exact objective correspondences to squared-Euclidean centroid reconstruction and normalized graph association and cut. Graph and image studies illustrate complementary routes within the finite-codeword family: graph- and teacher-defined relational requirements act directly on equality or collision, while reconstruction acts through a joint decoder. Together, these results illustrate how distinct relational requirements can be formulated and tested within a common constrained-representation framework.
Chinese Translation
当感兴趣的信息在于元素之间的关系而非元素本身时,压缩表示应该保留什么?我们在经典的源-描述-重构意义上形式化关系压缩,但将关系结构本身作为承载保真度的内容。每个实例指定源关系、保留的描述、重构或评估的关系、保真度准则以及约束资源。我们使用这一接口将来自图摘要、谱稀疏化、相似性保持表示和关系蒸馏的选定方法置于一个共同的形式框架中,同时明确它们不同的重构和资源假设。我们开发了有限码字碰撞作为一种具体实现。同码字概率产生了一种关系几何,将特定对的对齐和分离与聚合的 Rényi-2 占用以及分类分配的球面几何联系起来,并与平方欧几里得质心重构以及归一化图关联和割有精确的目标对应关系。图和图像研究展示了有限码字族内的互补路径:图和教师定义的关系要求直接作用于相等或碰撞,而重构通过联合解码器起作用。总之,这些结果说明了不同的关系要求如何能够在一个共同的受限表示框架内被形式化和测试。
cs.LG / 32 / 2609.31848

Averaged Mirror Descent and Dual Gradient Methods: Convergent Algorithms for Entropic Gromov-Wasserstein Problem

平均镜像下降和对偶梯度方法:用于熵Gromov-Wasserstein问题的收敛算法
Mark, Joanna, Rioux, Gabriel, Passegger, Riccardo
Abstract
The Gromov-Wasserstein (GW) distance measures the discrepancy between metric measure (mm) spaces and identifies optimal alignments between them based solely on their intrinsic structure. Since it identifies isomorphic mm spaces, it provides a natural notion of distance for heterogeneous datasets which may admit isomorphic representations. In order to accelerate computation of GW distances, many practitioners employ entropic regularization to obtain an Entropic GW (EGW) problem. The most popular EGW solver is the Mirror Descent (MD) algorithm, which reduces EGW computations to an iterative process where an entropic optimal transport (EOT) problem is solved at each iteration. Despite its widespread use, the convergence of MD for this problem has only been established for restricted classes of costs. On the other hand, a recently proposed dual gradient method is available for general costs, but requires a choice of step size which depends on the regularization parameter. To address these two issues, we introduce Averaged Mirror Descent (AMD), which averages consecutive MD steps, and prove its convergence for arbitrary costs. Then, we establish that the dual gradient method with a fixed step size also converges for arbitrary costs at the cost of a more complicated iteration. In both cases, we also account for inexact iterations which are inescapable in practice. We compare the empirical performance of these methods across various settings and, in particular, show that AMD and the dual gradient method both converge on an example where classical MD fails.
Chinese Translation
Gromov-Wasserstein (GW) 距离度量度量测度(mm)空间之间的差异,并仅基于其内在结构识别它们之间的最优对齐。由于它识别同构的mm空间,它为可能具有同构表示的异构数据集提供了一种自然的距离概念。为了加速GW距离的计算,许多从业者采用熵正则化来获得熵GW(EGW)问题。最流行的EGW求解器是镜像下降(MD)算法,它将EGW计算简化为迭代过程,其中每次迭代求解一个熵最优传输(EOT)问题。尽管被广泛使用,但MD对该问题的收敛性仅在受限的代价类中建立。另一方面,最近提出的对偶梯度方法可用于一般代价,但需要选择依赖于正则化参数的步长。为了解决这两个问题,我们引入了平均镜像下降(AMD),它对连续的MD步骤取平均,并证明其对任意代价的收敛性。然后,我们证明具有固定步长的对偶梯度方法也以更复杂的迭代为代价收敛于任意代价。在这两种情况下,我们还考虑了实际中不可避免的不精确迭代。我们比较了这些方法在各种设置下的经验性能,特别表明AMD和对偶梯度方法都在经典MD失败的例子上收敛。
cs.LG / 33 / 2609.31870

Deep Reinforcement Learning for Equity Trading: Benchmarking Actor-Critic Methods with Forward Retraining

深度强化学习用于股票交易:基于前向再训练的 Actor-Critic 方法基准测试
Wang, Bicheng, Zhang, Xinyi
Abstract
Consistently profitable trading is difficult because equity markets are noisy, non-stationary, and only partially predictable from historical data. We benchmark five deep reinforcement learning (DRL) actor-critic methods: A2C, PPO, DDPG, TD3, and SAC, that learn trading actions end-to-end from market states, and compare them with a supervised price-forecasting baseline. Using daily data for 20 large-capitalization S&P 500 stocks from 2000 to 2020, enriched with trend-following technical indicators and log min-max scaling, we train on 2000-2018 and backtest on 2019-2020. Each agent is evaluated both when trained once and under forward retraining, in which it is retrained on all data available before each successive test window. DDPG achieves the highest annual return (55.5%), Sharpe ratio (1.38), and alpha (0.22), but also the highest market beta (1.24). TD3 and SAC offer a better risk-return balance, with Sharpe ratios of 1.37 and 1.33 and maximum drawdowns of about 25%. Forward retraining improves A2C, PPO, and SAC, leaves TD3 essentially unchanged, and reduces DDPG's annual return from 55.5% to 29.8%, consistent with TD3's greater robustness to hyperparameters. The forecasting baseline has the smallest maximum drawdown (9.6%) and the lowest beta (0.31), underscoring a trade-off between the higher returns of end-to-end DRL and the lower risk of forecast-driven strategies.
Chinese Translation
持续盈利的交易很困难,因为股票市场充满噪声、非平稳,且仅能从历史数据中部分预测。我们基准测试了五种深度强化学习(DRL)actor-critic 方法:A2C、PPO、DDPG、TD3 和 SAC,这些方法从市场状态端到端地学习交易动作,并将它们与监督式价格预测基线进行比较。使用 2000 年至 2020 年 20 只大市值 S&P 500 股票的每日数据,并加入趋势跟踪技术指标和对数最小-最大缩放,我们在 2000-2018 年训练,在 2019-2020 年回测。每个智能体在训练一次和在前向再训练下进行评估,其中在每个后续测试窗口之前,使用所有可用数据重新训练。DDPG 实现了最高的年化收益率(55.5%)、夏普比率(1.38)和 alpha(0.22),但也具有最高的市场 beta(1.24)。TD3 和 SAC 提供了更好的风险-收益平衡,夏普比率分别为 1.37 和 1.33,最大回撤约为 25%。前向再训练提高了 A2C、PPO 和 SAC,使 TD3 基本保持不变,并将 DDPG 的年化收益率从 55.5% 降至 29.8%,这与 TD3 对超参数更强的鲁棒性一致。预测基线具有最小的最大回撤(9.6%)和最低的 beta(0.31),突显了端到端 DRL 的较高收益与预测驱动策略的较低风险之间的权衡。
cs.LG / 34 / 2609.31881

TemporalGraphLLM: Temporal Graph Neural Networks with Large Language Models for Dynamic Text-Attributed Graphs

TemporalGraphLLM:基于大语言模型的时间图神经网络用于动态文本属性图
Beladev, Moran, Eitan, Or, Katz, Gilad, Rokach, Lior
Abstract
Dynamic text-attributed graphs (DTAGs), where nodes, edges, and textual attributes evolve over time, are crucial in applications such as social networks, citation graphs, and knowledge graphs. However, existing approaches struggle to jointly model the temporal evolution of graph structures and the semantic richness of textual attributes. While Temporal Graph Neural Networks (TGNNs) capture evolving node relationships, they often lack contextual text reasoning. Conversely, Large Language Models (LLMs) excel in textual understanding but struggle with structured graph reasoning in temporal settings. To bridge this gap, we propose TemporalGraphLLM, a novel framework that can integrate any temporal GNN with an LLM for enhanced reasoning in DTAGs. Our approach fine-tunes LLMs using graph-time-aware instruction tuning and novel temporal GNNs injection to replace dedicated added tokens with graph embeddings. TemporalGraphLLM effectively leverages pretrained TGNNs within an LLM framework to achieve state-of-the-art performance on edge classification, link prediction, and edge-based text generation tasks. Extensive evaluation on real-world dynamic graph datasets demonstrates state-of-the-art performance. Our findings highlight the synergistic potential of LLMs and TGNNs, opening new directions for learning on evolving graphs.
Chinese Translation
动态文本属性图(DTAGs)中,节点、边和文本属性随时间演化,在社交网络、引文图和知识图谱等应用中至关重要。然而,现有方法难以联合建模图结构的时间演化和文本属性的语义丰富性。时间图神经网络(TGNNs)虽能捕捉演化的节点关系,但往往缺乏上下文文本推理能力。相反,大语言模型(LLMs)擅长文本理解,却难以在时间场景下进行结构化图推理。为弥合这一差距,我们提出TemporalGraphLLM,一种新颖框架,可将任意时间GNN与LLM集成,以增强DTAGs上的推理。我们的方法通过图-时间感知的指令微调和新颖的时间GNN注入来微调LLM,用图嵌入替换专门添加的标记。TemporalGraphLLM有效利用LLM框架内预训练的TGNN,在边分类、链接预测和基于边的文本生成任务上实现了最先进的性能。在真实世界动态图数据集上的广泛评估证明了最先进的性能。我们的发现凸显了LLM与TGNN的协同潜力,为演化图上的学习开辟了新方向。
cs.LG / 35 / 2609.31882

DOHF: Online Diffusion Fine-tuning with Doob's $h$-transform Guidance

DOHF:基于Doob $h$-变换引导的在线扩散微调
Guo, Zhengyi, Sheng, Jiayuan, Tang, Wenpin
Abstract
Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to estimate. In this work, we propose Diffusion Online $h$-guidance Fine-tuning (DOHF), which turns Doob's $h$-transform into a practical online training algorithm. DOHF assigns optimality weights to generated samples, estimates the normalized local correction $\nabla\log h$ under the current rollout policy, and distills it directly into the generative model. Theoretically, we characterize the population-optimal DiffusionNFT update as well as the various classfier free guidance methods through a unified $h$-transform perspective. Methodologically, our framework accommodates black-box and non-differentiable rewards without additional network evaluations. We further show improved alignments under three empirical scenarios. Our work demonstrates how adapting probabilistic conditioning through inexpensive estimation and iterative distillation can improve generative learning across statistical sampling and visual generation.
Chinese Translation
基于奖励的扩散微调面临实际挑战,当理想结果罕见或条件校正估计成本高昂时。在这项工作中,我们提出了扩散在线$h$-引导微调(DOHF),它将Doob的$h$-变换转化为实用的在线训练算法。DOHF为生成的样本分配最优性权重,在当前展开策略下估计归一化的局部校正$\nabla\log h$,并将其直接蒸馏到生成模型中。在理论上,我们通过统一的$h$-变换视角,刻画了群体最优的DiffusionNFT更新以及各种无分类器引导方法。在方法上,我们的框架适应黑盒和不可微的奖励,无需额外的网络评估。我们进一步在三个经验场景下展示了改进的对齐。我们的工作展示了如何通过低成本估计和迭代蒸馏来调整概率条件,从而改进跨统计采样和视觉生成的生成学习。
cs.LG / 36 / 2609.31890

FARE: Deep Reinforcement Learning For Fair Exposure Constrained Uncertainty Aware Financial Content Personalization

FARE:面向公平曝光约束的不确定性感知金融内容个性化的深度强化学习
Chinta, Arundeep, Tran, Lucas Vinh, Katukuri, Jay
Abstract
Content personalization systems in financial services must ensure fair exposure across diverse offerings-a requirement driven by contractual obligations and the need to prevent "rich-get-richer" dynamics where content with high click-through rate (CTR) dominates while other relevant products receive minimal visibility. Share of Voice (SOV) constraints, which guarantee each content category a target fraction of top-position exposure, address this by promoting product diversity and balanced user discovery. While re-ranking layers atop CTR models are common in practice, we propose two key novelties: (1) framing SOV-constrained ranking as a deep reinforcement learning problem analogous to constrained trade execution in algorithmic finance, and (2) explicitly incorporating CTR prediction uncertainty into the agent's state space and policy design-enabling larger ranking adjustments for high-uncertainty predictions where deviation from CTR-optimal ordering is less costly. We introduce FARE (Fair Ranking Executor), a modular uncertainty-aware execution layer that translates any black-box CTR model's predictions into SOV-fair rankings without retraining the underlying model. Our uncertainty-weighted proportional control policy (FARE-PC) and learned neural policies (FARE-ES, FARE-PPO) demonstrate that uncertainty-aware approaches can substantially reduce SOV deviation from fairness targets while minimizing engagement loss, with gradient-free evolution strategies outperforming policy gradient methods on synthetic data and the ordering reversing on KuaiRand-Pure.
Chinese Translation
金融服务中的内容个性化系统必须确保在多样化产品之间公平曝光——这一要求由合同义务以及防止“富者愈富”动态的需求所驱动,即高点击率(CTR)的内容占据主导,而其他相关产品获得的曝光极少。声音份额(SOV)约束保证每个内容类别在顶部位置曝光中达到目标比例,通过促进产品多样性和平衡的用户发现来解决这一问题。虽然在实际应用中,CTR模型之上的重排序层很常见,但我们提出两个关键创新点:(1)将SOV约束的排序问题构建为深度强化学习问题,类似于算法金融中的约束交易执行;(2)将CTR预测不确定性显式地纳入智能体的状态空间和策略设计——对于高不确定性的预测,允许更大的排序调整,因为此时偏离CTR最优排序的代价较小。我们提出FARE(公平排序执行器),一个模块化的不确定性感知执行层,能够将任何黑盒CTR模型的预测转化为SOV公平的排序,而无需重新训练底层模型。我们的不确定性加权比例控制策略(FARE-PC)和学习到的神经策略(FARE-ES,FARE-PPO)表明,不确定性感知方法能够显著降低SOV相对于公平目标的偏差,同时最小化参与度损失,其中无梯度进化策略在合成数据上优于策略梯度方法,而在KuaiRand-Pure上顺序反转。
cs.LG / 37 / 2609.31893

CyberWorld: World Models for Sample-Efficient Autonomous Cyber Defense

CyberWorld:面向样本高效自主网络防御的世界模型
Masukawa, Ryozo, Yun, Sanggeon, Hassan, Raheeb, Oh, Hyunwoo, Jeong, SungHeon, Imani, Mohsen
Abstract
Deep reinforcement learning has become a prominent approach to autonomous cyber defense. Existing methods are predominantly model-free and consequently require extensive environment interaction. World models provide an alternative by learning predictive dynamics and optimizing policies through imagined trajectories, yielding substantial gains in sample efficiency in robotics and embodied control. Extending this paradigm to cybersecurity raises a fundamental question: what should constitute the "world" in a cyber world model? We introduce CyberWorld, a Dreamer-style world modeling framework that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of the defended network. Across all four scoreable CyberWheel attack strategies, the graph-based CyberWorld variant exceeds a strategy-agnostic control after 3.6k-15.8k environment steps, compared with millions of steps required by model-free PPO. Across representation choices, graph structure provides greater robustness under topology-dependent attacks, while simpler representations remain competitive in overall performance. Among successful runs, the number of episodes required to reach the control remains approximately constant as network size increases from 15 to 100 hosts. These results establish learned cyber dynamics as a sample-efficient and scalable basis for autonomous defense, and identify world representation as a central design axis for robustness and scalability.
Chinese Translation
深度强化学习已成为自主网络防御的一种重要方法。现有方法大多是无模型的,因此需要大量环境交互。世界模型提供了一种替代方案:通过学习预测性动力学,并利用想象轨迹优化策略,从而在机器人学和具身控制中显著提升样本效率。将这一范式扩展到网络安全领域引出一个根本问题:在网络世界模型中,“世界”应由什么构成?我们提出 CyberWorld,一个 Dreamer 风格的世界建模框架,它从受防御网络的向量、图、文本和多模态表示中学习潜在网络动力学。在所有四种可评分的 CyberWheel 攻击策略上,基于图的 CyberWorld 变体在 3.6k–15.8k 个环境步后超过策略无关的对照基线,而无模型的 PPO 需要数百万步。在不同表示选择中,图结构在依赖拓扑的攻击下提供更强的鲁棒性,而更简单的表示在整体性能上仍具竞争力。在成功的运行中,随着网络规模从 15 台主机增加到 100 台主机,达到对照基线所需的情节数大致保持恒定。这些结果确立了学习到的网络动力学作为自主防御的样本高效且可扩展的基础,并将世界表示确定为鲁棒性和可扩展性的核心设计轴。
cs.LG / 38 / 2609.31900

Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training

理解LLM后训练中SFT、RLVR和OPD的协同效应
Acikgoz, Emre Can, Li, Yang, Liu, Zeyu Leo, Bansal, Srijan, Hakkani-Tür, Dilek, Joty, Shafiq, Yavuz, Semih
Abstract
Modern LLM post-training composes supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and on-policy distillation (OPD) into multi-stage pipelines, yet these stages are typically designed and evaluated in isolation. We show that this composition is consequential: a stage that improves the current model can make the next stage less effective. Through controlled experiments with Qwen3 models on math and science reasoning, we first characterize OPD across nine student-teacher pairs spanning 2x to 53x parameter ratios and show that OPD effectiveness depends on student-teacher compatibility rather than teacher scale alone. The surrounding stages of OPD reshape this compatibility in three ways: (1) A brief SFT warm-up improves subsequent OPD, while an RLVR-strengthened student regresses under distillation from the same teacher. (2) Adapting the teacher with RLVR raises downstream OPD accuracy in proportion to the capability it adds. Following these two interventions, we find that combining teacher adaptation and student warm-up alone raise average OPD accuracy from 29.2\% to 43.8\% (50\% relative improvement) after the same number of distillation steps, with additional preparatory training. (3) At comparable accuracy, OPD leaves a stronger initialization for downstream RLVR than SFT, with a gap that widens as RL compute scales. Our results suggest that each post-training stage should be chosen not only for the capability it adds, but for the learning interface it creates for the next stage.
Chinese Translation
现代LLM后训练将监督微调(SFT)、可验证奖励的强化学习(RLVR)和同策略蒸馏(OPD)组合成多阶段流水线,然而这些阶段通常被孤立地设计和评估。我们表明这种组合具有重要影响:一个改进当前模型的阶段可能使下一个阶段效果变差。通过在数学和科学推理任务上对Qwen3模型进行受控实验,我们首先在九个学生-教师模型对(参数比例从2倍到53倍不等)上刻画了OPD的特征,并表明OPD的有效性取决于学生-教师的兼容性,而不仅仅是教师模型的规模。OPD的周边阶段以三种方式重塑这种兼容性:(1)短暂的SFT预热可以改善后续的OPD,而经过RLVR强化的学生在同一教师的蒸馏下反而会出现性能回退。(2)用RLVR对教师进行适配,可以提升下游OPD的准确率,提升幅度与其增加的能力成正比。在这两项干预之后,我们发现,仅结合教师适配和学生预热,并辅以额外的预备训练,在相同数量的蒸馏步骤下,就能将平均OPD准确率从29.2%提升到43.8%(相对提升50%)。(3)在准确率相当的情况下,OPD为下游RLVR留下的初始化比SFT更强,且这一差距随着RL计算量的增加而扩大。我们的结果表明,选择每个后训练阶段时,不仅应考虑其增加的能力,还应考虑其为下一阶段创造的学习接口。
cs.LG / 39 / 2609.31918

ROTE: Benchmarking Neural Memorization on Complexity-Controlled Symbolic Sequences

ROTE:在复杂度受控的符号序列上对神经记忆进行基准测试
Chen, Xinye, Güttel, Stefan, Mozaffari, Mohammad
Abstract
We introduce ROTE (RollOut Testing of Exact memorization), a benchmarking protocol for evaluating symbolic memorization of neural architectures. We study memorization and the extension of symbolic rules in neural sequence models by using sequences whose complexity is regulated by Lempel--Ziv--Welch (LZW) compression. Under ROTE, each architecture is trained as the same finite-context conditional predictor and is evaluated using teacher-forced one-step prediction as well as closed-loop rollout on the withheld symbols. Following a shared prediction-and-rollout evaluation routine, the benchmark evaluates gated recurrent, minimal recurrent, attention-based, and hybrid recurrent-attention models with their native computational characteristics preserved. Beyond standard predictive metrics, the benchmark reports normalized string distances, training time, memory usage, and parameter count across an LZW-complexity sweep. The study establishes a connection between the complexity of algorithmic sequences and the memorization capacity of neural architectures, revealing the trade-offs involving memorization quality, rollout stability, and computational expense. Our software and reproducible experimental code can be obtained from https://github.com/nla-group/rote.
Chinese Translation
我们介绍了ROTE(精确记忆的展开测试),一种用于评估神经架构符号记忆的基准测试协议。我们通过使用复杂度由Lempel-Ziv-Welch (LZW)压缩调控的序列,研究神经序列模型中的记忆和符号规则的扩展。在ROTE下,每个架构都作为相同的有限上下文条件预测器进行训练,并使用教师强制的一步预测以及对保留符号的闭环展开进行评估。遵循共享的预测和展开评估流程,该基准评估了门控循环、最小循环、基于注意力和混合循环-注意力模型,同时保留了它们原有的计算特性。除了标准的预测指标外,该基准还报告了在LZW复杂度扫描下的归一化字符串距离、训练时间、内存使用和参数数量。该研究建立了算法序列的复杂度与神经架构的记忆能力之间的联系,揭示了涉及记忆质量、展开稳定性和计算开销的权衡。我们的软件和可复现的实验代码可从 https://github.com/nla-group/rote 获取。
cs.LG / 40 / 2609.31932

Resource-Aware Federated Mixture-of-Experts with Adaptive Pruning for Onboard Learning in LEO Satellite Constellations

面向LEO卫星星座在轨学习的资源感知联邦混合专家与自适应剪枝
Shaaban, Mohamed, Elmahallawy, Mohamed, Bernahrndt, Marius, Hecking, Tobias
Abstract
Low-Earth-orbit (LEO) satellites are increasingly expected to perform onboard learning for applications such as disaster response and environmental monitoring. However, conventional federated learning (FL) is ill-suited to onboard satellite learning, as it assumes computational, memory, and communication resources beyond the capabilities of resource-constrained LEO platforms, often necessitating the transmission of raw imagery to ground stations. We present COSMIC-FL, a resource-aware FL framework for efficient onboard learning in LEO satellite constellations. COSMIC-FL introduces two complementary Mixture-of-Experts (MoE) architectures: a Sliced design that shares backbone representations while activating task-specific channel subsets, and a Modular design that employs lightweight gating to route inputs to physically separated expert networks. A semantic class-to-expert mapping enables each satellite to train, update, and communicate only the expert paths relevant to its local data. To further improve efficiency, COSMIC-FL integrates staged optimization with three structured pruning strategies: server-side pruning, client-side fixed-ratio pruning with mean-vote aggregation, and adaptive client-side per-layer pruning based on aggregated importance and a MAD-based gap criterion. Combined with semantic expert routing, these techniques jointly adapt computation and model sparsity to both data semantics and layer importance, yielding a favourable accuracy--efficiency trade-off for heterogeneous space platforms. Experiments on six image classification benchmarks under highly non-i.i.d. settings show that COSMIC-FL maintains competitive accuracy while reducing communication, computation, and energy consumption by up to 80% over SOTA FL methods. We further validate COSMIC-FL on an NVIDIA Jetson AGX Orin, confirming its efficiency gains under realistic embedded deployment constraints.
Chinese Translation
低地球轨道(LEO)卫星越来越多地被期望执行星上学习,以支持灾害响应和环境监测等应用。然而,传统的联邦学习(FL)不适合星上卫星学习,因为它假设的计算、内存和通信资源超出了资源受限的LEO平台的能力,通常需要将原始图像传输到地面站。我们提出了COSMIC-FL,一种用于LEO卫星星座高效星上学习的资源感知FL框架。COSMIC-FL引入了两种互补的混合专家(MoE)架构:一种切片式设计,共享主干表示,同时激活任务特定的通道子集;以及一种模块化设计,采用轻量级门控将输入路由到物理上分离的专家网络。语义类别到专家的映射使每颗卫星仅训练、更新和通信与其本地数据相关的专家路径。为了进一步提高效率,COSMIC-FL将分阶段优化与三种结构化剪枝策略相结合:服务端剪枝、带有均值投票聚合的客户端固定比例剪枝,以及基于聚合重要性和基于MAD的差距准则的自适应客户端逐层剪枝。结合语义专家路由,这些技术共同使计算和模型稀疏性适应数据语义和层重要性,为异构空间平台产生有利的精度-效率权衡。在高度非独立同分布(non-i.i.d.)设置下对六个图像分类基准的实验表明,COSMIC-FL保持了有竞争力的精度,同时将通信、计算和能耗比最先进的FL方法降低了高达80%。我们进一步在NVIDIA Jetson AGX Orin上验证了COSMIC-FL,证实了其在实际嵌入式部署约束下的效率提升。
cs.LG / 41 / 2609.31938

Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders

跨嵌入式世界模型解码器的缓存感知 Conv3D Lowering
Zhang, Jiaming, Yang, Wu, Tao, Shuai, Liu, Wulong
Abstract
Generative world models can provide visual rollouts for embodied planning, yet their feasibility on edge devices depends not only on the learned model but also on how the execution runtime represents its operations. We introduce a cache-aware lowering that expresses supported causal Conv3D calls as batched spatial Conv2D operations while preserving pretrained weights, temporal-cache semantics, convolution parameters, bias placement, and output layout. Across the complete Cosmos3-Edge image-to-video pipeline on a 64-GB NVIDIA Jetson AGX Orin, the proposed route accelerates VAE decoding by approximately $7\times$ and reduces complete-generation latency by more than $2\times$, while repeated decoder evaluations maintain complete fast-path coverage without fallbacks. The unchanged lowering also improves Cosmos3-Nano and transfers to LingBot-World's architecturally distinct Wan2.1 VAE. A clean-device comparison against fully specialized TensorRT shows that TensorRT provides a further $1.36\times$ steady-state improvement, but requires substantially greater per-module and per-runtime-state AOT specialization. Same-latent BF16 and FP32 evaluations characterize the finite-precision differences introduced by the alternative execution order. Together, these results position cache-aware lowering as a lightweight runtime optimization that recovers most of the available decoder acceleration without modifying the learned models themselves.
Chinese Translation
生成式世界模型可为具身规划提供视觉推演(rollout),但其在边缘设备上的可行性不仅取决于所学习的模型,也取决于执行运行时如何表示其算子。我们提出一种缓存感知的 lowering,将受支持的因果 Conv3D 调用表达为批量化的空间 Conv2D 操作,同时保持预训练权重、时间缓存语义、卷积参数、偏置位置与输出布局不变。在配备 64 GB 内存的 NVIDIA Jetson AGX Orin 上,针对完整的 Cosmos3-Edge 图像到视频流水线,所提路径将 VAE 解码加速约 7 倍,并将完整生成延迟降低 2 倍以上,同时重复的解码器评估始终保持完整的快速路径覆盖,无需回退。该未加改动的 lowering 同样改善了 Cosmos3-Nano,并可迁移至 LingBot-World 中架构迥异的 Wan2.1 VAE。与完全特化的 TensorRT 在干净设备上的对比表明,TensorRT 可进一步提升稳态性能 1.36 倍,但需要在每个模块与每个运行时状态上进行显著更多的 AOT 特化。相同潜变量下的 BF16 与 FP32 评估刻画了由替代执行顺序所引入的有限精度差异。总体而言,这些结果将缓存感知 lowering 定位为一种轻量级运行时优化,无需修改学习模型本身即可恢复大部分可用的解码器加速。
cs.LG / 42 / 2609.31940

Simple Extensions of Single-Objective Acquisition Functions and Hedge Strategies for Multi-Objective Bayesian Optimization

多目标贝叶斯优化中单目标采集函数与对冲策略的简单扩展
Sheikh, Haris Moazam
Abstract
Multi-objective Bayesian optimization (MOBO) is commonly approached through specialized acquisition functions or scalarization schemes designed to explicitly account for trade-offs among non-preferential objectives. In this work, we show that such complexity might be unnecessary. We propose a framework that extends standard single-objective acquisition functions directly to the multi-objective setting through a hypervolume-based transformation. We further extend hedge strategies for acquisition functions, which are typically used only in single-objective optimization, to the multi-objective regime. Our approach requires minimal modification to existing Bayesian optimization pipelines and avoids the need for bespoke multi-objective formulations. We demonstrate how a broad class of commonly used single-objective acquisition functions and hedge strategies can be adapted in a principled manner to handle multiple objectives, while preserving their intuitive interpretation and computational efficiency. Empirically, we evaluate the proposed methods across a range of synthetic and real-world multi-objective benchmarks. Despite their simplicity, our extensions consistently match or outperform more complex state-of-the-art MOBO methods in terms of optimization performance and sample efficiency. These results suggest that effective multi-objective Bayesian optimization can be achieved by reusing and carefully extending well-established single-objective acquisition strategies, offering a simpler and more flexible alternative to existing approaches.
Chinese Translation
多目标贝叶斯优化(MOBO)通常通过专门的采集函数或标量化方案来处理,这些方案旨在明确考虑非偏好目标之间的权衡。在这项工作中,我们表明这种复杂性可能是不必要的。我们提出了一个框架,通过基于超体积的变换,将标准单目标采集函数直接扩展到多目标设置。我们进一步将用于采集函数的对冲策略(通常仅用于单目标优化)扩展到多目标情形。我们的方法仅需对现有贝叶斯优化流程进行最小修改,避免了定制多目标公式的需要。我们展示了如何以有原则的方式调整一大类常用的单目标采集函数和对冲策略以处理多个目标,同时保持其直观解释和计算效率。在实验上,我们在一系列合成和真实世界的多目标基准上评估了所提出的方法。尽管简单,我们的扩展在优化性能和样本效率方面始终能够匹配或优于更复杂的最先进的MOBO方法。这些结果表明,通过重用和仔细扩展成熟的单目标采集策略,可以实现有效的多目标贝叶斯优化,为现有方法提供了一种更简单、更灵活的替代方案。
cs.LG / 43 / 2609.31947

On-Policy Attention Linearization

同策略注意力线性化
Raje, Arian, Nayak, Anupam, Fei, Anthony, Parthasarathy, Akaash, Abdelfattah, Mohamed, Joshi, Gauri
Abstract
Hybrid transformer architectures that replace most softmax attention layers with linear attention offer transformer-level quality at a fraction of the memory cost. Rather than pretraining such models, a growing body of work distills them from already trained full-attention transformers. However, these distilled models often collapse on long-context retrieval and reasoning tasks, particularly when operating in thinking mode, where the efficiency gains of hybrid architectures matter most. Since linear attention layers must compress context into a fixed-size state, their errors compound over long sequences. As off-policy distillation never teaches the student model to recover from this drift, tasks that necessitate longer sequence lengths become especially challenging. We introduce On-Policy Attention Linearization (OPAL) in which the hybrid attention student samples its own long-context trajectories and receives dense supervision from the frozen full-attention teacher. Applying OPAL to Qwen3-4B and MiMo-7B-RL-0530, we recover $87$--$94\%$ of full-attention performance on commonsense reasoning, $100\%$ on needle-in-a-haystack (NIAH) retrieval, and $83$--$93\%$ on mathematical reasoning with only 3B training tokens. We achieve these results without supervised fine-tuning (SFT) or reinforcement learning with verifiable rewards (RLVR). Compared with the strongest prior linearization method, which recovers $68\%$ of its teacher's retrieval performance and $21.6\%$ absolute average mathematical reasoning accuracy, OPAL fully recovers retrieval and achieves $67.6$--$72.2\%$ on math reasoning.
Chinese Translation
混合Transformer架构用线性注意力替换大多数softmax注意力层,以一小部分内存成本提供Transformer级别的质量。与其预训练此类模型,越来越多的研究致力于从已训练好的全注意力Transformer中蒸馏它们。然而,这些蒸馏模型在长上下文检索和推理任务上常常崩溃,特别是在思考模式下运行时,而混合架构的效率增益在此最为重要。由于线性注意力层必须将上下文压缩为固定大小的状态,其误差在长序列上会累积。由于异策略蒸馏从未教学生模型从这种漂移中恢复,需要更长序列长度的任务变得尤为具有挑战性。我们引入了同策略注意力线性化(OPAL),其中混合注意力学生模型采样自己的长上下文轨迹,并从冻结的全注意力教师模型接收密集监督。将OPAL应用于Qwen3-4B和MiMo-7B-RL-0530,我们仅用3B训练token就在常识推理上恢复了87%--94%的全注意力性能,在大海捞针(NIAH)检索上恢复了100%,在数学推理上恢复了83%--93%。我们在没有监督微调(SFT)或可验证奖励的强化学习(RLVR)的情况下实现了这些结果。与最强的先前线性化方法相比,该方法恢复了其教师模型68%的检索性能和21.6%的绝对平均数学推理准确率,OPAL完全恢复了检索性能,并在数学推理上达到了67.6%--72.2%。
cs.LG / 44 / 2609.31960

Model-Agnostic Online Certificate-Driven Calibration for Time Series Forecasting Under Distribution Shift

分布偏移下时间序列预测的模型无关在线证书驱动校准
Huang, Chenfeng, Ma, Zixuan, Michailidis, George
Abstract
Time series out-of-distribution generalization requires forecasters to remain reliable when deployment dynamics differ from training conditions due to covariate shift, concept shift, and temporal dependence. Probably Approximately Correct Bayesian domain adaptation provides computable certificates by decomposing target risk into a source risk term, a source-to-target mismatch term, and a complexity term, but standard analyses rely on independent sampling and distributional stability, assumptions that are violated in time series by serial dependence and nonstationary shift. We propose a model-agnostic online martingale Probably Approximately Correct Bayesian framework that yields finite-sample certificates under temporal dependence and distribution shift. The certificate replaces independent-sample concentration with martingale concentration that adapts to loss scale and predictable variation. We use the certificate as a surrogate regularizer for online calibration by training a gated residual Bayesian head on top of a fixed forecasting backbone, producing a corrective update that reverts to the backbone prediction when the gate is closed. Online calibration combines a source risk anchor, a posterior-shift penalty, and a time-adaptive mismatch term computed from target windows observed before forecasting. It follows a predict-then-update protocol in which outcomes become available only after forecasting and are used to update subsequent predictions. Experiments across convolutional, attention-based, and large language model-based forecasters show improved stability and accuracy under covariate and concept shift.
Chinese Translation
时间序列分布外泛化要求预测器在部署动态因协变量偏移、概念偏移和时间依赖性而与训练条件不同时保持可靠。可能近似正确贝叶斯(PAC-Bayesian)域适应通过将目标风险分解为源风险项、源到目标不匹配项和复杂度项,提供了可计算证书,但标准分析依赖于独立采样和分布稳定性,这些假设在时间序列中因序列相关性和非平稳偏移而被违反。我们提出了一种模型无关的在线鞅PAC-Bayesian框架,该框架在时间依赖性和分布偏移下产生有限样本证书。该证书用适应损失尺度和可预测变差的鞅集中代替独立样本集中。我们将该证书用作在线校准的代理正则化器,通过在固定预测主干之上训练门控残差贝叶斯头,产生校正更新,当门关闭时恢复为主干预测。在线校准结合了源风险锚、后验偏移惩罚和根据预测前观察到的目标窗口计算的时间自适应不匹配项。它遵循预测后更新协议,其中结果仅在预测后可用,并用于更新后续预测。在卷积、基于注意力和基于大语言模型的预测器上的实验表明,在协变量和概念偏移下,稳定性和准确性得到了提高。
cs.LG / 45 / 2609.31975

Model Casting and Low-Parameter Gating: Towards More Sparsely Activated FFNs

Model Casting 与 Low-Parameter Gating:迈向更稀疏激活的 FFN
Lomeli, Maria, Groudiev, Antoine, Douze, Matthijs, Cabannes, Loïc, Mazaré, Pierre Emmanuel, Fleuret, François, Beck, Maximilian, Szilvasy, Gergely, Murray, Naila, Jégou, Hervé
Abstract
This paper introduces model casting, a mid-training recipe that drastically sparsifies the activations within the Feed-Forward Network (FFN) layer. With this strategy, at inference time, we first compute the output of the gating matrix and, thanks to its high sparsity, we avoid computations with the two other matrices, reducing the FLOP count by up to 3x. While this theoretical speedup is an upper bound, model casting translates into significant speedups both on CPU and GPU. We then introduce LoPA Gating, a new FFN design that increases the maximum theoretical speedup. It is a low-FLOPs parameterization of the gating matrix that overcomes the 3x cap by allocating fewer FLOPs and parameters to the gating matrix, compared to the two other FFN matrices that are sparsely activated. We consider two cases: (i) we cast a pre-trained model with a sparsity inducing activation; (ii) we train with LoPA from scratch. In all settings, we significantly outperform existing pruning solutions and regular RELU-fication. For instance, at matched quality, we achieve a 3.2x FLOP speedup with LoPA Casting, against 1.6x at best for competing methods top-p and TEAL. Using dedicated kernels, we achieve an actual 3.31x speed-up on GPU at 90% sparsity, past the 3x ceiling of standard gating; RELU-fication, meanwhile, plateaus below 80% sparsity.
Chinese Translation
本文介绍了 Model Casting,一种在训练中期使用的方法,它能极大地稀疏化前馈网络(FFN)层内的激活。采用该策略,在推理时,我们首先计算门控矩阵的输出,得益于其高度稀疏性,我们避免了与其他两个矩阵的计算,从而将 FLOP 计数最多减少 3 倍。虽然这种理论加速是一个上限,但 Model Casting 在 CPU 和 GPU 上都能转化为显著的加速。然后我们介绍了 LoPA Gating,一种新的 FFN 设计,它提高了最大理论加速比。它是一种门控矩阵的低 FLOPs 参数化方法,通过为门控矩阵分配更少的 FLOPs 和参数,克服了 3 倍的上限,与其他两个被稀疏激活的 FFN 矩阵相比。我们考虑了两种情况:(i) 我们对预训练模型应用 Model Casting,使用稀疏诱导激活;(ii) 我们从头开始用 LoPA 训练。在所有设置中,我们都显著优于现有的剪枝解决方案和常规的 RELU-fication。例如,在匹配质量下,我们通过 LoPA Casting 实现了 3.2 倍的 FLOP 加速,而竞争方法 top-p 和 TEAL 最多只有 1.6 倍。使用专用内核,我们在 90% 稀疏度下于 GPU 上实现了实际的 3.31 倍加速,超过了标准门控的 3 倍上限;而 RELU-fication 则在 80% 稀疏度以下趋于平缓。
cs.LG / 46 / 2609.31983

Understanding the Subspace Stabilization of the Hessian and Gradient Covariance Matrix

理解Hessian矩阵与梯度协方差矩阵的子空间稳定化
Liao, Fangshuo, Kyrillidis, Anastasios
Abstract
The phenomenon of the top subspace stabilization of the Hessian matrix is an surprising and critical aspect in study of the second-order information of neural network training. Prior work argues that the top subspace of the Hessian stabilizes by measuring the overlap between the top subspaces of the step-wise Hessian, and explains this stabilization with diminishing parameter change in the late phase of training. In this paper, we define a new instability metric for the subspace evolution, and use it to detect subspace stabilization that is independent of the magnitude of parameter change. In the meantime, we observe that the gradient covariance matrix has a similar property of its top subspace to the Hessian. By using a between-class and within-class decomposition of the gradient covariance matrix, we identify an explicit form that gives a near-perfect approximation of the top-$(C-1)$ subspace of the Hessian and the gradient covariance matrix. In the gradient flow set-up, we show that the slow evolution of the idenfied approximation is due to the separation between the outlier and the bulk eigenvalues of the Hessian matrix, thus providing an explanation to the phenomenon of the top subspace stabilization of the Hessian matrix.
Chinese Translation
Hessian矩阵的顶部子空间稳定化现象是神经网络训练二阶信息研究中的一个令人惊讶且关键的方面。先前的工作通过测量逐步Hessian的顶部子空间之间的重叠来论证Hessian的顶部子空间稳定,并用训练后期参数变化的减小来解释这种稳定化。在本文中,我们定义了一个新的子空间演化不稳定性度量,并用它来检测与参数变化幅度无关的子空间稳定化。同时,我们观察到梯度协方差矩阵的顶部子空间具有与Hessian相似的性质。通过使用梯度协方差矩阵的类间和类内分解,我们确定了一个显式形式,它给出了Hessian和梯度协方差矩阵的顶部$(C-1)$子空间的近乎完美的近似。在梯度流设置中,我们表明所识别近似的缓慢演化是由于Hessian矩阵的离群特征值和体特征值之间的分离,从而为Hessian矩阵的顶部子空间稳定化现象提供了解释。
cs.LG / 47 / 2609.31988

Lagrangian and Hamiltonian Neural Networks With a Dissipative System

耗散系统中的拉格朗日与哈密顿神经网络
Rayamajhi, V., Singal, J.
Abstract
We investigate the applicability of Lagrangian and Hamiltonian Neural Network models to a dissipative system that has explicit time dependence in its Lagrangian, Hamiltonian, and total energy. To do so we consider these neural network models for simulated systems of a harmonic one-dimensional, one-component oscillator with damping, as well as without damping for comparison. We find that both the Lagrangian and Hamiltonian approaches are able to predict the empirical physical behavior of the damped oscillator systems and to effectively ``learn'' to varying degrees the underlying Lagrangians and Hamiltonians, as has previously been shown to be the case with undamped oscillator systems. These investigations elucidate important properties of Lagrangian and Hamiltonian mechanics, including properties that are not manifest when considering systems without explicit time dependence.
Chinese Translation
我们研究了拉格朗日神经网络和哈密顿神经网络模型对耗散系统的适用性,该耗散系统的拉格朗日量、哈密顿量和总能量具有显式时间依赖。为此,我们考虑了这些神经网络模型用于模拟的有阻尼的一维单分量谐振子系统,以及无阻尼的相同系统作为比较。我们发现拉格朗日方法和哈密顿方法都能够预测阻尼振子系统的经验物理行为,并且能够不同程度地有效“学习”潜在的拉格朗日量和哈密顿量,正如之前在无阻尼振子系统中所显示的那样。这些研究阐明了拉格朗日力学和哈密顿力学的重要性质,包括在考虑没有显式时间依赖的系统时不明显的性质。
cs.LG / 48 / 2609.31992

Mechanistic Interpretability Reveals Shared Causal Subspaces in Brain-to-Speech Decoders

机制可解释性揭示脑到语音解码器中的共享因果子空间
Maghsoudi, Maryam, Mishra, Ayushi, Dutta, Sanghamitra
Abstract
Decoding covert speech, such as mimed or imagined, from brain activity is harder than decoding vocalized speech. Cross-modal transfer, where information from one speech form helps decode another, is a promising remedy; yet how a decoder internally represents and processes brain activity from different speech forms remains unclear. In this work, we ask: which internal neurons of a decoder carry cross-modal information, and are these neurons shared across different speech forms? To answer these questions, we leverage mechanistic interpretability, using recordings of the same sentences in vocalized, mimed, and imagined input pairs for activation patching. We insert the decoder's internal activity for a sentence in one condition into its processing of the same sentence in another and measure the change in decoding accuracy. We find that no single neuron drives this benefit; instead, it arises from small groups of neurons, with vocalized speech as the most useful source. These groups are largely condition-specific in the early stage of the decoder but overlap in the later stage. These findings point toward more data-efficient covert speech decoders through training objectives that encourage shared later-stage representations learned mainly from vocalized data.
Chinese Translation
从大脑活动中解码隐蔽言语(如模仿或想象的言语)比解码发声言语更困难。跨模态迁移,即来自一种言语形式的信息有助于解码另一种言语形式,是一种有前景的补救方法;然而,解码器如何内部表征和处理来自不同言语形式的脑活动仍不清楚。在这项工作中,我们提出:解码器的哪些内部神经元携带跨模态信息,并且这些神经元是否在不同言语形式之间共享?为了回答这些问题,我们利用机制可解释性,使用相同句子的发声、模仿和想象输入对的记录进行激活修补。我们将一个条件下一个句子的解码器内部活动插入到另一个条件下同一句子的处理中,并测量解码准确率的变化。我们发现,没有单个神经元驱动这种益处;相反,它来自小群神经元,其中发声言语是最有用的来源。这些群体在解码器的早期阶段很大程度上是条件特异性的,但在后期阶段重叠。这些发现指向更数据高效的隐蔽言语解码器,通过鼓励主要从发声数据中学习共享后期表征的训练目标。
cs.LG / 49 / 2609.31996

Can Circuit Alignment Predict OOD Generalization?

电路对齐能否预测OOD泛化?
Banerjee, Ayan, Chaudhuri, Abhra, Llados, Josep, Pal, Umapada, Dutta, Anjan
Abstract
Can out-of-distribution (OOD) generalization be predicted from a trained model's weights alone, without any target-domain data? Existing representational similarity metrics (CKA, SVCCA, RSA) compare activations rather than forecast generalization. We show they are provably insensitive to structural rerouting in the computational graph, the very change distribution shift induces. We close this gap with the Circuit Alignment Score (CAS), which compares class-specific circuits across domains via graph kernels, decomposed into same-class coherence and cross-class confusion. Casting CAS as a Lebesgue integral over the domain distribution, we prove its Monte Carlo estimate recovers the ground-truth ranking of learners by OOD accuracy, with pairwise inversion error vanishing at rate $O(1/M)$, where $M$ is the number of sampled domains. Across $48$ learners on PACS, CAS attains $0.88$ rank correlation with OOD accuracy, versus $0.58$ (CKA), $0.23$ (SVCCA), and $0.14$ (RSA), with similar trends on other benchmarks and even against data-dependent methods, making it the first provably consistent predictor of distributional robustness requiring neither target-domain data nor labels. The code is available at: https://github.com/ayanban011/ACE
Chinese Translation
仅凭训练模型的权重,而无需任何目标域数据,能否预测分布外(OOD)泛化?现有的表示相似性度量(CKA、SVCCA、RSA)比较的是激活值,而非预测泛化性能。我们证明,它们可证明地对计算图中的结构性重路由不敏感,而该重路由正是分布偏移所引发的改变。我们通过电路对齐分数(CAS)填补这一空白,该分数借助图核比较跨域的类特定电路,并分解为同类一致性和跨类混淆。将CAS表示为域分布上的勒贝格积分,我们证明其蒙特卡洛估计能够恢复按OOD准确率对学习器排序的真实排名,成对反转误差以 $O(1/M)$ 的速率消失,其中 $M$ 为采样域的数量。在PACS上的48个学习器中,CAS与OOD准确率的秩相关性达到0.88,而CKA为0.58、SVCCA为0.23、RSA为0.14,在其他基准上以及与数据依赖方法的比较中也呈现类似趋势,使其成为首个无需目标域数据或标签、可证明一致的分布鲁棒性预测器。代码可在以下网址获取:https://github.com/ayanban011/ACE
cs.LG / 50 / 2609.31997

Integrating Language Models into Listened and Imagined Speech Decoding from MEG

将语言模型集成到基于 MEG 的听觉和想象语音解码中
Maghsoudi, Maryam, Kankanala, Sai Samrat, Shamma, Shihab A., Ganapathy, Sriram
Abstract
Decoding imagined speech is an important goal for brain-computer interfaces but remains challenging due to weak neural responses, low signal-to-noise ratio, and limited imagined-speech datasets. Language models provide strong contextual cues for text prediction, but how much they can help neural decoding and whether their contribution differs for decoding perceived and imagined speech remains unclear. To investigate this, we use a paired listened-imagined MEG dataset and incorporate language-model information at two stages. First, we train a contrastive neural decoder that aligns MEG representations with acoustic and contextual language representations, improving cross-subject word decoding for both listened and imagined speech. Second, at inference, we introduce a neural-constrained beam-search framework that combines neural evidence with language-model next-word probabilities. We find that imagined-speech decoding benefits more from the language model than listened-speech decoding. For Imagined speech, the best-performing balance between neural and language-model evidence shifts toward the language model, and the gain over neural-only decoding is larger. Together, these results suggest that language priors are most useful when neural evidence is weaker, making them particularly valuable for imagined-speech BCIs.
Chinese Translation
解码想象语音是脑机接口的一个重要目标,但由于神经反应微弱、信噪比低以及想象语音数据集有限,仍然具有挑战性。语言模型为文本预测提供了强大的上下文线索,但它们能在多大程度上帮助神经解码,以及它们在解码感知语音和想象语音时的贡献是否不同,仍不清楚。为了研究这一点,我们使用了一个配对的听觉-想象 MEG 数据集,并在两个阶段融入语言模型信息。首先,我们训练了一个对比神经解码器,将 MEG 表征与声学和上下文语言表征对齐,提高了听觉和想象语音的跨被试单词解码性能。其次,在推理阶段,我们引入了一个神经约束的束搜索框架,将神经证据与语言模型的下一个词概率相结合。我们发现,想象语音解码比听觉语音解码从语言模型中获益更多。对于想象语音,神经和语言模型证据之间的最佳平衡向语言模型偏移,并且相对于仅神经解码的增益更大。总之,这些结果表明,当神经证据较弱时,语言先验最为有用,这使得它们对想象语音脑机接口特别有价值。
cs.LG / 51 / 2609.31999

VC Dimension and Expressivity of Real-Valued Transformers

实值Transformer的VC维与表达能力
Dooley, Gavin, Yang, Andy, Zhu, Yijia Jessica, Chiang, David, Cholak, Peter, Pillay, Anand
Abstract
Whereas previous results on abilities and limitations of transformers have restricted the definition of transformers in various ways, here we study softmax-attention, multi-layer transformers operating on real values, with very few additional assumptions. Applying results from real geometry, we obtain upper bounds on the VC dimension and split VC dimension of such transformers ($O(n^4)$ and $O(n^6)$, respectively, where $n$ is the input length). Conversely, we also construct specific transformers witnessing lower bounds on these quantities ($\Omega(n)$ in each case). These results have some notable consequences. For example, within the class of symmetric (permutation-invariant) functions, we show that transformers can uniformly express all functions over an alphabet of one symbol and non-uniformly express all functions over an alphabet of two symbols, but cannot (even non-uniformly) express some functions over an alphabet of six symbols. We also prove limitations on how many bits of a real number a transformer can access.
Chinese Translation
以往关于Transformer能力与局限性的研究结果以多种方式限制了Transformer的定义,而本文研究在实值上运行的softmax注意力、多层Transformer,几乎没有额外假设。应用实几何的结果,我们得到了这类Transformer的VC维和分裂VC维的上界(分别为O(n^4)和O(n^6),其中n为输入长度)。反之,我们还构造了特定的Transformer,给出了这些量的下界(每种情况下为Ω(n))。这些结果有一些值得注意的推论。例如,在对称(置换不变)函数类中,我们证明Transformer可以一致地表达单符号字母表上的所有函数,非一致地表达双符号字母表上的所有函数,但不能(即使非一致地)表达六符号字母表上的某些函数。我们还证明了Transformer访问一个实数的位数的限制。
cs.LG / 52 / 2609.32004

Is invariance all you need for algorithmic fairness? Removing demographic information can create new bias

不变性是算法公平所需的一切吗?移除人口统计信息可能产生新的偏差
Parikh, Aditya, Petersen, Eike, Frank, Stella, Ferrante, Enzo, Ganz, Melanie, Feragen, Aasa
Abstract
Encoded demographic information in internal model representations is a commonly assumed risk factor for algorithmic bias, with demographic representation invariance often being touted as the ideal state. However, while demographic shortcut learning is a genuine threat, some degree of encoding is necessary when demographics correlate with target labels. Here, we show, mathematically and empirically, that enforcing demographic invariance can actually hamper bias mitigation and even create new biases. We distinguish marginal from class-conditional representation invariance, and show that they imply the standard group fairness notions of demographic parity and equalized odds, respectively. We evaluate the effects on predictive performance and fairness of enforcing both invariance types, both theoretically and empirically across five tabular and two chest X-ray imaging datasets. Our findings support our mathematical argument that demographic representation invariance is neither desirable nor sufficient for fairness.
Chinese Translation
内部模型表征中编码的人口统计信息通常被认为是算法偏差的风险因素,而人口统计表征不变性常被标榜为理想状态。然而,虽然人口统计捷径学习是真实威胁,但当人口统计与目标标签相关时,一定程度的编码是必要的。在此,我们从数学上和实证上表明,强制人口统计不变性实际上可能妨碍偏差缓解,甚至产生新的偏差。我们区分了边缘表征不变性和类别条件表征不变性,并表明它们分别蕴含标准群体公平性概念:人口统计均等和机会均等。我们在五个表格数据集和两个胸部X射线成像数据集上,从理论和实证上评估了强制这两种不变性类型对预测性能和公平性的影响。我们的发现支持我们的数学论证:人口统计表征不变性既不可取,也不足以实现公平。
cs.LG / 53 / 2609.32008

Human Activity Recognition via Ultra-Wideband Data: A Framework for Dimensionality Reduction, Pattern Discovery, and Predictive Modeling

基于超宽带数据的人类活动识别:一个用于降维、模式发现和预测建模的框架
Gozin, Nahid Sahel, Sedaghat, Reza, Siddavaatam, Prathap
Abstract
Recent advances in sensor technology have enabled more effective human activity recognition (HAR), particularly in real-time systems with limited computational resources. However, Ultra-Wideband (UWB) radar data remain challenging due to high dimensionality, noise, complexity, and nonlinear characteristics. This research proposes a framework to efficiently reduce data size, uncover significant patterns, and classify six activity types (Standff, Liedown, Noactivity, Sit, Stand, and Walk) from UWB signals with high accuracy. Two novel dimensionality reduction techniques are introduced in this paper. The first, Clustered Polynomial Expansion with Incremental PCA (CPE-IPCA), combines clustering and polynomial feature expansion with Incremental PCA, preserving 100% of the variance in only 50 components. The second, Post-PCA Standardization Approach (PPSA), standardizes data after PCA and retains 99.1% of the variance in 80 components, achieving superior compression and computational efficiency compared to conventional nonlinear methods. Frequent patterns are identified using Apriori and FP-Growth, which are then classified with Random Forest and a Vector Space Model (VSM). The framework achieves 100% accuracy with Random Forest on CPE-IPCA and 99% on PPSA, while VSM attains 100% precision, recall, and F1 on PPSA and near-perfect performance on CPE-IPCA (precision 1.00, recall 0.98-1.00, F1 0.99-1.00), demonstrating a fast, interpretable, and robust HAR system suitable for healthcare, assisted living, and smart environments.
Chinese Translation
近年来,传感器技术的进步使得更有效的人类活动识别(HAR)成为可能,特别是在计算资源有限的实时系统中。然而,由于高维性、噪声、复杂性和非线性特征,超宽带(UWB)雷达数据仍然具有挑战性。本研究提出了一个框架,能够高效地减少数据量、发现显著模式,并对UWB信号中的六种活动类型(Standff、Liedown、Noactivity、Sit、Stand和Walk)进行高精度分类。本文介绍了两种新颖的降维技术。第一种是聚类多项式扩展与增量PCA(CPE-IPCA),它将聚类和多项式特征扩展与增量PCA相结合,仅用50个成分就保留了100%的方差。第二种是PCA后标准化方法(PPSA),它在PCA之后对数据进行标准化,并在80个成分中保留了99.1%的方差,与传统非线性方法相比,实现了更优的压缩和计算效率。使用Apriori和FP-Growth识别频繁模式,然后使用随机森林和向量空间模型(VSM)进行分类。该框架在CPE-IPCA上使用随机森林达到100%的准确率,在PPSA上达到99%的准确率;而VSM在PPSA上达到了100%的精确率、召回率和F1值,在CPE-IPCA上表现接近完美(精确率1.00,召回率0.98-1.00,F1值0.99-1.00),展示了一个快速、可解释且鲁棒的HAR系统,适用于医疗保健、辅助生活和智能环境。
cs.LG / 54 / 2609.32018

Representation Learning for Exact Preimages

精确原像的表示学习
Hess, Konstantin, Feuerriegel, Stefan
Abstract
Modern neural predictors can model highly nonlinear maps, but many scientific and engineering tasks require reasoning in the opposite direction: given a performance or safety level, the goal is to characterize the preimage, that is, the complete set of inputs which meet the desired target level and optimize over that set. For expressive neural predictors, however, such preimages typically have no explicit representation and are expensive to recover or optimize over. This creates a fundamental three-way challenge between expressive forward prediction, accurate preimage approximation, and tractable optimization over the preimage for downstream tasks. We introduce TRIO (tractable representations for preimage learning and inverse optimization), a framework for learning representations that make these objectives compatible by construction. Our key contribution is a preimage factorization: the forward model remains expressive through nonlinear radial transformations (including neural networks), while, under inversion, each transformation reduces to a single scalar radius, which yields simple geometric level sets. This yields an explicit geometric representation that is reusable for downstream optimization over the preimage, and, for linear objectives, we show that this admits a closed-form global solution. We finally prove a universal approximation theorem which shows that TRIO can approximate any continuous forward map and its entire family of potentially disconnected, nonconvex preimages arbitrarily well. Hence, TRIO combines expressive forward modeling, exact preimage recovery, and tractable global downstream optimization over preimages by design.
Chinese Translation
现代神经预测器能够建模高度非线性的映射,但许多科学与工程任务需要在相反方向进行推理:给定性能或安全水平,目标是刻画原像,即满足期望目标水平的完整输入集合,并在该集合上优化。然而,对于表达力强的神经预测器,此类原像通常没有显式表示,且恢复或优化成本高昂。这造成了表达性前向预测、精确原像近似以及针对下游任务对原像进行可处理优化之间的根本性三难困境。我们提出了 TRIO(用于原像学习和逆优化的可处理表示),这是一个学习表示的框架,通过构造使这些目标相互兼容。我们的关键贡献是原像因子分解:前向模型通过非线性径向变换(包括神经网络)保持表达力,而在反演下,每个变换简化为单个标量半径,从而产生简单的几何水平集。这产生了一种显式的几何表示,可复用于对原像的下游优化,并且对于线性目标,我们证明这允许闭式全局解。我们最后证明了一个通用逼近定理,表明 TRIO 可以任意好地逼近任何连续前向映射及其整个可能不连通、非凸的原像族。因此,TRIO 通过设计结合了表达性前向建模、精确原像恢复以及对原像的可处理全局下游优化。
cs.LG / 55 / 2609.32048

Interactive Distributionally Robust Multi-Agent Learning with General Function Approximation

具有通用函数逼近的交互式分布鲁棒多智能体学习
Ghosh, Debamita, Atia, George K., Wang, Yue
Abstract
Model misspecification poses a fundamental challenge in multi-agent reinforcement learning, where transition uncertainty can be amplified by strategic interactions among agents. Distributionally robust Markov games (DRMGs) provide a principled framework for addressing such uncertainty, yet existing methods often rely on restrictive assumptions or scale poorly to large state and joint action spaces. We study online learning in general-sum DRMGs with general function approximation and $\phi$-divergence uncertainty sets. We propose RoMEX-$\phi$, a model-free framework that integrates equilibrium-based exploration with dual fitted learning. Through a functional dual representation of the robust multi-agent Bellman operator, RoMEX-$\phi$ enables tractable worst-case value estimation from nominal interaction data using a centered empirical robust discrepancy. We introduce the robust Multi-Agent Decoupling Coefficient (robust MADC) to characterize the intrinsic exploration complexity arising from strategic interactions and adversarial transition uncertainty. We establish sublinear robust regret guarantees governed by the robust MADC rather than explicitly by the state and joint action space sizes, replacing tabular dependence with intrinsic function-class complexity. Numerical experiments on a scalable general-sum DRMG under total variation uncertainty show that RoMEX-$\phi$ is substantially more resilient to transition shifts than its non-robust counterpart while remaining competitive with an exact tabular robust baseline. Our results provide a scalable framework for distributionally robust multi-agent reinforcement learning with general function approximation.
Chinese Translation
模型误设定对多智能体强化学习构成根本性挑战,其中转移不确定性可能被智能体之间的策略交互放大。分布鲁棒马尔可夫博弈(DRMGs)为解决此类不确定性提供了原则性框架,但现有方法通常依赖限制性假设,或难以扩展到大规模状态和联合动作空间。我们研究具有通用函数逼近和$\phi$-散度不确定性集的一般和DRMG中的在线学习。我们提出RoMEX-$\phi$,一个无模型框架,集成了基于均衡的探索与对偶拟合学习。通过鲁棒多智能体贝尔曼算子的函数对偶表示,RoMEX-$\phi$能够利用中心化经验鲁棒差异,从名义交互数据中进行可处理的最坏情况价值估计。我们引入鲁棒多智能体解耦系数(robust MADC),以刻画由策略交互和对抗性转移不确定性引起的内在探索复杂度。我们建立了由鲁棒MADC控制的次线性鲁棒遗憾保证,而不是显式地由状态和联合动作空间大小控制,用内在函数类复杂度取代了表格依赖性。在总变差不确定性下的可扩展一般和DRMG上的数值实验表明,RoMEX-$\phi$对转移偏移的鲁棒性显著高于其非鲁棒对应方法,同时与精确的表格鲁棒基线保持竞争力。我们的结果为具有通用函数逼近的分布鲁棒多智能体强化学习提供了一个可扩展的框架。
cs.LG / 56 / 2609.32056

Graph Forward Distribution Matching for Molecular Inverse Design

面向分子逆向设计的Graph Forward Distribution Matching
Zhu, Yihan, Liu, Yuhan, Savoie, Brett, Luo, Tengfei, Jiang, Meng
Abstract
Achieving precise control over multiple properties without sacrificing chemical validity remains a central challenge in molecular inverse design. Existing reinforcement learning (RL) methods fine-tune graph diffusion models by treating **reverse** sampling as a sequential policy, using a single terminal reward to optimize hundreds of coupled decisions. They often suffer from instability, validity collapse, and limited property gains. We introduce GraphFDM (Graph Forward Distribution Matching), a new online RL paradigm for graph diffusion that performs optimization through the **forward** process. GraphFDM uses valid generations to define a reward-tilted target distribution jointly optimized over graph size and molecular structure for each property condition, incorporating reinforcement signals into supervised learning without storing reverse trajectories. We derive the unique optimal target, prove a condition-wise improvement guarantee, and show that the fixed graph-size prior of standard graph diffusion leaves an irreducible matching gap. In multi-conditional polymer and small-molecule generation, GraphFDM achieves the lowest MAE on every target property, with reductions of up to 53.0\% relative to the strongest baselines and chemical validity above 0.99. It further generalizes to out-of-distribution property combinations.
Chinese Translation
在分子逆向设计中,实现对多种性质的精确控制而不牺牲化学有效性仍然是一个核心挑战。现有的强化学习(RL)方法通过将反向采样视为序列策略来微调图扩散模型,使用单个终端奖励来优化数百个耦合决策。它们通常面临不稳定性、有效性崩溃和性质增益有限的问题。我们引入了GraphFDM(Graph Forward Distribution Matching),一种新的用于图扩散的在线RL范式,它通过前向过程进行优化。GraphFDM使用有效生成来定义奖励倾斜的目标分布,该分布针对每个性质条件在图大小和分子结构上联合优化,将强化信号融入监督学习而无需存储反向轨迹。我们推导出唯一的最优目标,证明了逐条件改进保证,并表明标准图扩散的固定图大小先验留下了不可约的匹配差距。在多条件聚合物和小分子生成中,GraphFDM在每个目标性质上实现了最低的MAE,相对于最强基线降低了高达53.0%,化学有效性高于0.99。它进一步泛化到分布外性质组合。
cs.LG / 57 / 2609.32060

Depth Laws for the Precision Floor of Trained Neural Networks: Amplification, Residual Scaling, and a Quantization-Aware Training Paradox

训练神经网络的精度下限的深度定律:放大、残差缩放与量化感知训练悖论
Tarawneh, Ahmad S.
Abstract
How many bits does a network need before its accuracy collapses, and how does this grow with depth? We study the precision floor, the perturbation level or bit-width at which accuracy falls halfway to chance, in MLPs, CNNs, Vision Transformers and nine pretrained language models, under post-training quantization (PTQ) and quantization- or noise-aware training (QAT). (i) A first-order theory sets the floor through one full-precision quantity, the predictive amplification $G$: $\eta_c=\Lambda/G$, and $G^2$ grows linearly in depth at a rate proportional to the squared residual branch scale. (ii) The predicted equality $\alpha_{PTQ}=\rho$ of depth exponents holds within 95% intervals in twelve of thirteen trained architectures and in GPT-2 from 12 to 48 layers, with $\Lambda=1.45\pm14\%$ across trained architectures. (iii) Residual branches scaled by $1/\sqrt{D}$ and pre-normalisation remove the depth penalty, and each quantizer turns noise into bits at a rate fixed by its step rule, giving $b_c=(\alpha/\gamma)\log_2 D+C$. (iv) A QAT paradox: noise-aware training roughly doubles the tolerable noise of shallow networks, but the gain decays with depth, so the depth law steepens ($\alpha_{QAT}/\alpha_{PTQ}=1.45$-$1.47$ on two datasets, ten seeds each). Decision margins, cross-layer error cancellation and heavy tails do not set the floor.
Chinese Translation
一个网络在精度崩溃之前需要多少比特,以及这如何随深度增长?我们研究了精度下限,即准确率下降到随机水平一半时的扰动水平或比特宽度,在 MLP、CNN、Vision Transformer 和九个预训练语言模型中,在训练后量化(PTQ)和量化感知或噪声感知训练(QAT)下。(i) 一个一阶理论通过一个全精度量——预测放大 $G$ 来确定下限:$\eta_c=\Lambda/G$,并且 $G^2$ 随深度线性增长,其速率与残差分支尺度的平方成正比。(ii) 预测的深度指数相等关系 $\alpha_{PTQ}=\rho$ 在十三个训练架构中的十二个以及 GPT-2 的 12 到 48 层中在 95% 区间内成立,其中在所有训练架构中 $\Lambda=1.45\pm14\%$。(iii) 通过 $1/\sqrt{D}$ 缩放和预归一化的残差分支消除了深度惩罚,并且每个量化器以其步长规则固定的速率将噪声转换为比特,得到 $b_c=(\alpha/\gamma)\log_2 D+C$。(iv) 一个 QAT 悖论:噪声感知训练大致使浅层网络的可容忍噪声翻倍,但增益随深度衰减,因此深度定律变得更陡峭(在两个数据集上,每个数据集十个种子,$\alpha_{QAT}/\alpha_{PTQ}=1.45$-$1.47$)。决策裕度、跨层误差抵消和重尾并不决定下限。
cs.LG / 58 / 2609.32083

When Should a Human Take Back Control? Optimal Delegation under Turbulent AI Risk

何时人类应收回控制权?动荡 AI 风险下的最优委托
Yan, Haoze, Roze, Julien, Upadhyay, Ved, Tatar, Unal, Mastrolia, Thibaut
Abstract
Deploying AI systems requires deciding when to delegate tasks and when humans should intervene to monitor and mitigate risk induced by AI operations. These decisions are challenging when failures cluster: a hallucination or harmful output can trigger further errors, creating periods of elevated risk. We introduce a continuous-time framework for learning adaptive human oversight under such turbulent AI risk. Existing oversight and delegation formulations condition on history but do not model incident clustering, or its suppression by supervision effort, jointly with the delegation decision and this study fixes this gap. The self-exciting dynamics capture how risk events increase the likelihood of subsequent events, making their timing and history central to decision-making. We formulate a stochastic control problem combining human actions, monitoring effort, and switching between human-AI-assisted operation and full AI delegation, balancing operational rewards against oversight costs, and cascading AI-failures and induced uncertainty. Human participation is an endogenous component of risk management: the policy determines both when oversight is needed and how much effort to allocate. We study a relaxed switching formulation and propose Hawkes-PPO, a policy-gradient method that uses a bank of exponential filters of observed incident times. In a synthetic environment it attains a higher risk-adjusted objective than either fixed regime and approaches an approximate full-information oracle. We illustrate our results with numerical simulations by examining how cascade risks influence intervention and delegation, connecting reinforcement learning with adaptive human oversight of AI systems. In particular, we illustrate the benefit of our switching strategy and Hawkes-PPO algorithm to monitor the project efficiently along time, reducing turbulent risks occurrences and costs.
Chinese Translation
部署 AI 系统需要决定何时委托任务以及何时人类应干预以监控和缓解 AI 操作引发的风险。当故障聚集时,这些决策颇具挑战:一次幻觉或有害输出可能触发进一步错误,造成风险升高的时期。我们引入一个连续时间框架,用于在此类动荡 AI 风险下学习自适应人类监督。现有的监督与委托公式以历史为条件,但未对事件聚集或其通过监督努力抑制与委托决策进行联合建模,本研究填补了这一空白。自激动力学捕捉了风险事件如何增加后续事件的可能性,使这些事件的时间和历史成为决策的核心。我们构建一个随机控制问题,结合人类行动、监控努力,以及在人机辅助运行与完全 AI 委托之间切换,平衡运营奖励与监督成本、级联 AI 故障和诱发的不确定性。人类参与是风险管理的内生组成部分:策略决定了何时需要监督以及分配多少努力。我们研究一种松弛切换公式,并提出 Hawkes-PPO,一种使用观察到的事件时间的指数滤波器组的策略梯度方法。在合成环境中,它比任一固定机制获得更高的风险调整目标,并接近近似完全信息预言机。我们通过数值模拟展示结果,考察级联风险如何影响干预与委托,将强化学习与 AI 系统的自适应人类监督联系起来。特别地,我们展示了我们的切换策略和 Hawkes-PPO 算法在随时间高效监控项目、减少动荡风险发生和成本方面的优势。
cs.LG / 59 / 2609.32100

Emergent One-Third Scaling Law as Attention Tries to Concentrate

注意力试图集中时涌现的1/3缩放定律
Liu, Yizhou, Kangaslahti, Sara, Gore, Jeff
Abstract
The neural scaling law relating longer training to better performance through a power law is central to today's large language models (LLMs), yet its origin remains debated. One recent proposal is that power laws can emerge from the strong non-linearity of a single softmax head learning peaked distributions. What happens with multiple softmax functions, as in LLMs, is unclear. Here, we show through toy models that any softmax learning peaked distributions, regardless of its position in the model, can develop logit magnitudes that grow in a power law with exponent $1/3$, becoming a training bottleneck whose loss contribution decays as a power law with the same exponent $1/3$. The overall loss therefore obeys $1/3$ scaling whenever at least one softmax learns peaked distributions. We confirm that many softmax functions in LLMs learn peaked distributions and that LLM loss scaling matches this $1/3$ prediction. Moreover, logit growth dynamics reveal that attention heads, rather than the language modeling head, are the bottleneck likely driving the $1/3$ loss scaling in LLMs. Attention trying to concentrate on specific information, which is the heart of Transformers, may therefore also be the heart of the neural scaling law of training.
Chinese Translation
神经缩放定律通过幂律将更长的训练与更好的性能联系起来,是当今大型语言模型(LLMs)的核心,但其起源仍存争议。最近的一种观点认为,幂律可能源自单个softmax头学习尖峰分布时的强非线性。而在LLMs中,存在多个softmax函数时会发生什么尚不清楚。在此,我们通过玩具模型表明,任何学习尖峰分布的softmax,无论其在模型中的位置如何,都会产生以指数1/3幂律增长的logit幅度,成为训练瓶颈,其损失贡献以相同指数1/3的幂律衰减。因此,只要至少有一个softmax学习到尖峰分布,总体损失就遵循1/3缩放。我们证实,LLMs中的许多softmax函数学习到尖峰分布,并且LLM损失缩放与此1/3预测相符。此外,logit增长动态揭示,注意力头而非语言建模头是可能驱动LLMs中1/3损失缩放的瓶颈。注意力试图集中于特定信息,这是Transformers的核心,因此也可能成为训练神经缩放定律的核心。
cs.LG / 60 / 2609.32103

LLM Unlearning Evaluation with TRIAGE

基于TRIAGE的LLM遗忘评估
Ataee, Danial, Triantafillou, Peter
Abstract
Large language models can memorize private or harmful information, motivating machine unlearning methods that remove targeted knowledge while preserving other capabilities. However, existing evaluations rely primarily on behavioral benchmarks, which assess \emph{whether} a model appears to forget but provide limited insight into \emph{how} unlearning changes the model or affects related knowledge. We introduce \textit{TRIAGE} (\textit{Tripartite Representation-internal Introspection for Adjacency Gap Evaluation}), a benchmark-agnostic evaluation framework for characterizing these changes. TRIAGE uses diagonal approximations of the Fisher information and Hessian to measure changes in parameter sensitivity and local curvature, and utilizes a Forget / \emph{Adjacent-Retain} / \emph{Generic-Retain} partition to quantify an \emph{adjacency gap} in semantically related knowledge. Based on the magnitude and distribution of these changes, TRIAGE further classifies each algorithm's update as \emph{no-op}, \emph{partially localized}, \emph{collateral dominant}, or \emph{globally destructive}. Across 12 unlearning methods, four language models, and the WMDP, TOFU, and MUSE benchmarks, we find that methods with similar behavioral forgetting can produce substantially different internal changes and patterns of collateral damage. These signatures also vary across models and benchmarks, indicating that the effects of unlearning are not determined solely by the unlearning algorithm. TRIAGE can be applied alongside existing unlearning benchmarks to complement behavioral evaluation with a model-internal view of how unlearning reshapes the model's parameter space and affects retained knowledge.
Chinese Translation
大语言模型可以记忆私有或有害信息,这促使了机器遗忘方法的发展,旨在移除目标知识同时保留其他能力。然而,现有的评估主要依赖于行为基准,这些基准评估模型是否看起来遗忘了,但对于遗忘如何改变模型或影响相关知识提供的见解有限。我们引入了TRIAGE(Tripartite Representation-internal Introspection for Adjacency Gap Evaluation),一个与基准无关的评估框架,用于表征这些变化。TRIAGE使用Fisher信息和Hessian的对角近似来测量参数敏感性和局部曲率的变化,并利用遗忘/邻接保留/通用保留的划分来量化语义相关知识中的邻接差距。基于这些变化的幅度和分布,TRIAGE进一步将每个算法的更新分类为无操作、部分局部化、附带主导或全局破坏性。在12种遗忘方法、4个语言模型以及WMDP、TOFU和MUSE基准上,我们发现具有相似行为遗忘的方法可能产生显著不同的内部变化和附带损害模式。这些特征也因模型和基准而异,表明遗忘的效果并非仅由遗忘算法决定。TRIAGE可以与现有的遗忘基准一起应用,以模型内部的视角补充行为评估,了解遗忘如何重塑模型的参数空间并影响保留的知识。
cs.LG / 61 / 2609.32106

An Attention-Driven Heterogeneous GNN Model for Credit Card Fraud Detection

用于信用卡欺诈检测的注意力驱动异构 GNN 模型
Jayabalan, Kathiresan, Radhakrishnan, Sethuraman
Abstract
The global transition to a cashless economy has placed credit cards as the key element of digital transactions, acclaimed for their easy use, speed, and acceptance in most places. However, the growing dependence on this payment method has led to an escalation of the risks associated with credit card (CC) fraud. Detecting this type of fraud is a difficult task because the patterns are constantly changing, there is a data imbalance, and it is necessary to identify the legitimate transactions and the fraud ones at the same time. This study addresses this challenge by proposing a credit card fraud detection (CCFD) framework using a data balancing technique and a deep learning (DL) model. The proposed fraud detection model is trained and evaluated by collecting the dataset called Credit Card Fraud Detection from the Kaggle repository. As the dataset is highly imbalanced, we utilized the Synthetic Minority Oversampling Technique (SMOTE)-Tomek technique to balance the dataset. Further, the balanced dataset is classified using the Heterogeneous Graph Neural Network (HGNN) model. The HGNN model represent various transactions using a heterogeneous graph architecture and by using an attention based message passing technique, it managed to consider the complex relationships, time factors, and user behavior. The integration of SMOTE-Tomek in the model further boosted its capacity to identify fraudulent transactions, while lowering the rate of false positives. The HGNN model attained a 99.97% accuracy, a 99.48% F1-score, a 99.15% precision, and a 98.97% recall. The findings indicates that this model is effective and can be applied to real-world CCFD scenarios.
Chinese Translation
全球向无现金经济的转型使信用卡成为数字交易的关键要素,因其易用性、速度和广泛接受度而备受赞誉。然而,对这种支付方式的日益依赖导致与信用卡(CC)欺诈相关的风险不断升级。检测此类欺诈是一项艰巨的任务,因为欺诈模式不断变化,存在数据不平衡问题,并且需要同时识别合法交易和欺诈交易。本研究通过提出一个使用数据平衡技术和深度学习(DL)模型的信用卡欺诈检测(CCFD)框架来应对这一挑战。所提出的欺诈检测模型通过从 Kaggle 仓库收集名为 Credit Card Fraud Detection 的数据集进行训练和评估。由于数据集高度不平衡,我们采用了合成少数类过采样技术(SMOTE)-Tomek 技术来平衡数据集。进一步,使用异构图神经网络(HGNN)模型对平衡后的数据集进行分类。HGNN 模型使用异构图架构表示各种交易,并通过基于注意力的消息传递技术,成功考虑了复杂关系、时间因素和用户行为。将 SMOTE-Tomek 集成到模型中进一步提升了其识别欺诈交易的能力,同时降低了假阳性率。HGNN 模型达到了 99.97% 的准确率,99.48% 的 F1 分数,99.15% 的精确率和 98.97% 的召回率。研究结果表明,该模型有效,可应用于现实世界的 CCFD 场景。
cs.LG / 62 / 2609.32108

SAMBAR: Selective Anchoring via Method of Multipliers for Balanced Knowledge Acquisition and Retention in Vision-Language-Action Models

SAMBAR:通过乘子法进行选择性锚定,以在视觉-语言-动作模型中实现平衡的知识获取与保持
Shrivastava, Aayushi, Zhou, Xunlan, Zhao, Hongrui, Chen, Ziyu, Mehr, Negar
Abstract
Vision-Language-Action (VLA) models leverage large-scale pretraining to ultimately achieve generalist manipulation. Deployed VLA policies must support continual learning to acquire new tasks over time. Teaching a VLA a new task generally requires finetuning it on demonstrations of that task. However, naively finetuning on downstream tasks causes the policy to forget earlier tasks and degrades generalist capabilities. This failure is known as catastrophic forgetting. Most continual learning methods counter it by replaying data from earlier tasks. However, the old task demonstrations are not always readily available. In this paper, we introduce SAMBAR, a continual learning algorithm that prevents catastrophic forgetting during VLA finetuning without requiring access to the demonstrations of any previously learned task. We propose to cast continual learning as a constrained optimization problem and solve it with the method of multipliers. In our approach, the method of multipliers drives the policy to learn the new task without the model parameters drifting far away from their previous values. In contrast to a standard regularization penalty, the method of multipliers raises the penalty as the constraint violation accumulates by using a dual variable. We also selectively anchor the parameters critical to previous tasks to preserve past knowledge, leaving other parameters free for new task acquisition. The combination of dual variable and selective anchoring, therefore, balances knowledge acquisition with knowledge retention. We evaluate our method, SAMBAR, on the LIBERO simulation benchmark and on hardware. When sequentially finetuning on a VLA, every replay-free baseline we compare against completely forgets the first task it learned, whereas SAMBAR retains every task it has learned.
Chinese Translation
视觉-语言-动作(VLA)模型利用大规模预训练来最终实现通用操作。部署的VLA策略必须支持持续学习,以随时间获取新任务。教导VLA一个新任务通常需要在其任务的演示上进行微调。然而,朴素地在下游任务上微调会导致策略忘记先前的任务并降低通用能力。这种失败被称为灾难性遗忘。大多数持续学习方法通过重放早期任务的数据来应对。然而,旧任务的演示并不总是容易获得。在本文中,我们介绍了SAMBAR,一种持续学习算法,可在VLA微调期间防止灾难性遗忘,而无需访问任何先前学习任务的演示。我们提出将持续学习视为约束优化问题,并用乘子法求解。在我们的方法中,乘子法驱动策略学习新任务,而不会使模型参数偏离其先前值太远。与标准正则化惩罚相反,乘子法通过使用对偶变量,随着约束违反的累积而提高惩罚。我们还选择性地锚定对先前任务至关重要的参数,以保留过去的知识,让其他参数自由用于新任务获取。因此,对偶变量和选择性锚定的结合平衡了知识获取与知识保持。我们在LIBERO模拟基准和硬件上评估我们的方法SAMBAR。当在VLA上顺序微调时,我们比较的每个无重放基线都完全忘记了它学习的第一个任务,而SAMBAR保留了它学过的每个任务。
cs.LG / 63 / 2609.32109

How Reusable Are Benchmarks with Richer Feedback?

更丰富反馈下的基准测试可复用性如何?
Allouah, Youssef, Duchi, John
Abstract
We study whether benchmarks reliably guide model selection as developers adapt to evaluation feedback across multiple criteria. We find that the worst-case test-set size needed to estimate the best score among $k$ adaptively chosen models, under any convex combination of the criteria, grows exponentially with the number of criteria, reaching the $\Theta(\sqrt{k})$ cost of answering $k$ adaptive statistical queries with only $O(\log k)$ criteria, at fixed accuracy and confidence. In attacks on multi-task large language model benchmarks with five to ten criteria, feedback restricted to nondominated task profiles produces large reused-to-held-out score gaps and frequent false winners. These results challenge a prominent explanation for prior observed reliable benchmark reuse---that developers mainly respond to convincing improvements over the current best---in rich-feedback settings, while leaving open how often ordinary model development encounters this vulnerability.
Chinese Translation
我们研究当开发者根据多个标准的评估反馈进行调整时,基准测试是否能可靠地指导模型选择。我们发现,在标准的任意凸组合下,为估计$k$个自适应选择的模型中的最佳得分所需的最坏情况测试集大小,随标准数量呈指数增长;在固定精度和置信度下,当标准数量仅为$O(\log k)$时,其大小就达到了回答$k$个自适应统计查询的$\Theta(\sqrt{k})$代价。在对具有五到十个标准的多任务大语言模型基准测试的攻击中,仅限于非支配任务配置的反馈会产生较大的复用集与留出集分数差距,并频繁出现错误胜出者。这些结果在反馈丰富的设定下,挑战了对于先前观察到的基准测试可靠复用的一个主流解释——即开发者主要对优于当前最佳模型的令人信服的改进做出响应——同时对于普通模型开发中这种脆弱性出现的频率仍悬而未决。
cs.LG / 64 / 2609.32117

Handwritten Digit Leakage from Smartphone Motion Sensors Across Unseen Users and Phone Models

基于智能手机运动传感器的跨未知用户与手机型号的手写数字泄露
Casado, Constantino Álvarez, Rantahalvari, Erkka, Pedone, Matteo, Matilainen, Matti, Cañellas, Manuel Lage, Nguyen, Le, Hosio, Simo, Silvén, Olli, López, Miguel Bordallo
Abstract
Smartphone motion sensors support interactive applications, but their readings may also reveal touchscreen input beyond their intended use. Assuming known drawing intervals, we study whether handwritten digits remain predictable across users and devices, as a 10-class problem on 19,628 HuMIdb recordings from 481 participants. We compare handcrafted features with classical machine learning algorithms, MiniRocket kernels, and a compact sensor patch transformer on accelerometer, linear acceleration, gyroscope, and gravity signals. The transformer achieves 57.74\% accuracy and 82.64\% top-3 accuracy on 75 unseen participants, and 58.77$\pm$0.95\% over 3 seeds for unseen participants on 9 unseen phone models. Low motion recordings remain informative, accuracy is not monotonic in motion level, and the tested contrastive pretraining, augmentation, and derived signals give no consistent gains. Digits are thus predictable beyond familiar users and phone models under assumed segmentation, while acquisition-order shortcuts limit conclusions about practical privacy exposure. Code available at: https://github.com/Arritmic/motion-digit-leakage.
Chinese Translation
智能手机运动传感器支持交互应用,但其读数也可能在预期用途之外泄露触摸屏输入。在已知绘制间隔的假设下,我们研究手写数字是否在跨用户和跨设备时仍可预测,将其视为一个10分类问题,基于来自481名参与者的19,628条HuMIdb记录。我们在加速度计、线性加速度、陀螺仪和重力信号上,比较了手工特征与经典机器学习算法、MiniRocket核以及一个紧凑的传感器块Transformer。该Transformer在75名未见过的参与者上达到57.74%的准确率和82.64%的top-3准确率,并在9个未见过的手机型号上针对未见过的参与者,在3个随机种子下达到58.77±0.95%的准确率。低运动记录仍然具有信息量,准确率并非随运动水平单调变化,并且所测试的对比预训练、数据增强和派生信号没有带来一致的增益。因此,在假设的分割下,数字在熟悉的用户和手机型号之外仍可预测,而采集顺序捷径限制了对实际隐私暴露的结论。代码见:https://github.com/Arritmic/motion-digit-leakage。
cs.LG / 65 / 2609.32124

What Should We Freeze? Guarded Freezing: Connectivity Shapes the Fine-Tuning of Pretrained Models

我们应该冻结什么?Guarded Freezing:连通性塑造了预训练模型的微调
Aguilar, Leonel
Abstract
When adapting pre-trained models through fine-tuning, freezing weights alone might not preserve performance, as updates elsewhere can change the inputs to the frozen core, ultimately affecting overall performance. We first analyse the case where a selected frozen core can be isolated and propose removal-value, a capacity-based score that approximates HOPE's removal cost averaged over removal orders. We show that in VGG-8, cutting paths from trainable neurons into a frozen core makes selection using this score useful: $70\%$ frozen preserves $5.22\pm0.51$ percentage points more old-task accuracy than DEFT at similar new-task accuracy. In transformers, shared residual streams leave paths into frozen neurons open. For this case, we derive drift-value, a forward-only proxy for the output disturbance from updating each weight entry under a local update model. In language models, at 40 epochs, this policy exceeds adapted Wanda and RIA freezing scores in settings with substantial retention loss, while its differences from Fisher remain unresolved. After 160 epochs on Qwen2.5-1.5B, it retains $0.0433\pm0.0102$ more than static Fisher. In DINOv3 vision-transformer adaptation to point clouds, drift-value retains $0.440$ image accuracy versus $0.187$ for a random mask of the same count. These results motivate Guarded Freezing: select by removal-value when incoming paths are cut, and by drift-value when they remain.
Chinese Translation
当通过微调来适配预训练模型时,仅冻结权重可能无法保持性能,因为其他部分的更新会改变冻结核心的输入,最终影响整体性能。我们首先分析了所选冻结核心可以被隔离的情况,并提出了 removal-value,一种基于容量的分数,它近似于 HOPE 的移除成本在移除顺序上的平均值。我们表明,在 VGG-8 中,切断从可训练神经元到冻结核心的路径,使得使用该分数进行选择变得有用:在相似的新任务准确率下,冻结 70% 比 DEFT 保持高出 5.22±0.51 个百分点的旧任务准确率。在 Transformer 中,共享的残差流使通向冻结神经元的路径保持开放。对于这种情况,我们推导出 drift-value,一种仅前向的代理,用于衡量在局部更新模型下更新每个权重条目所引起的输出扰动。在语言模型中,在 40 个 epoch 时,该策略在具有显著保留损失的设置中超过了 adapted Wanda 和 RIA 冻结分数,而其与 Fisher 的差异仍未解决。在 Qwen2.5-1.5B 上经过 160 个 epoch 后,它比静态 Fisher 多保留 0.0433±0.0102。在 DINOv3 视觉 Transformer 适配到点云时,drift-value 保留了 0.440 的图像准确率,而相同数量的随机掩码为 0.187。这些结果启发了 Guarded Freezing:当传入路径被切断时,通过 removal-value 进行选择;当这些路径保留时,通过 drift-value 进行选择。
cs.LG / 66 / 2609.32132

Latency-Aware Client Assignment for Parallel Split Learning With Global Sampling

面向带全局采样的并行拆分学习的延迟感知客户端分配
Kohankhaki, Mohammad, Rentschler, Valentin, Schmeink, Anke
Abstract
In cross-silo split learning, Parallel Split Learning with Global Sampling forms representative pooled batches when class distributions differ across clients, but ignores client delay when several clients can supply the same class. We introduce Latency Budgeted Parallel Split Learning with Global Sampling, which separates each pooled batch's integer class target from the choice of clients that supply its examples. The flow variant formulates this assignment as an integral network-flow problem and minimizes modeled client-side completion time for the current target. The fast variant uses a greedy next-completion rule to reduce schedule-construction cost. Both preserve the target stream and use every local example once per epoch. A planning rule selects between the variants while accounting for the cost of constructing both candidate schedules. On CIFAR-10, the flow variant reduces modeled training time by 6.75%, with a 0.30 percentage-point decrease in final accuracy. On Tiny ImageNet with 20 candidate classes per client, the fast variant reduces modeled time by 16.87% and reaches all four validation targets earlier than the latency-unaware baseline. Across 405 schedule comparisons, the planning rule stays within 2% of the lower realized cost in 96.54% of cases. In our evaluation, latency-aware provider assignment reduces modeled training time without changing the prescribed class targets, while the preferred variant depends on whether assignment savings outweigh schedule-construction overhead.
Chinese Translation
在跨孤岛拆分学习中,当客户端之间的类别分布不同时,带全局采样的并行拆分学习会形成具有代表性的池化批次,但当多个客户端可以提供相同类别时,它忽略了客户端延迟。我们引入了延迟预算的带全局采样的并行拆分学习,它将每个池化批次的整数类别目标与提供其样本的客户端选择分离开来。流变体将此分配形式化为整数网络流问题,并最小化当前目标下建模的客户端完成时间。快速变体使用贪婪的下一完成规则来降低调度构建成本。两者都保留目标流,并在每个 epoch 中使用每个本地样本一次。规划规则在两种变体之间进行选择,同时考虑构建两种候选调度的成本。在 CIFAR-10 上,流变体将建模的训练时间减少了 6.75%,最终准确率下降了 0.30 个百分点。在每客户端有 20 个候选类别的 Tiny ImageNet 上,快速变体将建模时间减少了 16.87%,并且比延迟无感知基线更早达到所有四个验证目标。在 405 次调度比较中,规划规则在 96.54% 的情况下与较低的实际成本相差不超过 2%。在我们的评估中,延迟感知的提供者分配减少了建模的训练时间,同时不改变规定的类别目标,而首选的变体取决于分配节省是否超过调度构建开销。
cs.LG / 67 / 2609.32141

REALM: Regime-Switching, Explainable, and Activation-Induced Linear Models

REALM:机制切换、可解释且激活诱导的线性模型
Cheng, Xiaoran, Na, Sen, Li, Jia
Abstract
Deep ReLU networks are piecewise-affine mappings that partition the input space into cells, each characterized by a distinct activation pattern. This structure motivates fitting a local linear model within each cell to preserve predictive accuracy while improving interpretability. The challenge is to identify regimes that are stable, data-adaptive, and easy to explain. We propose REALM, a mixture of linear models whose regimes are induced by neural activation patterns. Because the number of activation cells in a deep neural network (DNN) can grow rapidly with depth, we first distill a deep teacher into a wide, shallow student network (WSSN), then binarize and cluster its hidden-layer activations to define the regimes and fit a linear model within each regime. Since the regimes are discovered from internal structure, the router does not carry the predictive burden. To make regime assignment interpretable, we train a multiclass logistic regression, the explanatory gate, to reproduce the regime assignments. The two-level structure is interpretable at both stages in terms of raw tabular or learned convolutional features: the gate identifies features that determine regime assignments, while the linear models identify features that drive predictions within each regime. We analyze an idealized setting that illustrates a trade-off between partition complexity and stability: as the number of regimes grows, finer partitions can improve approximation but may reduce regime-assignment stability. Experiments on tabular and image datasets show that REALM achieves competitive predictive performance relative to other DNN-guided mixture surrogates and inherently interpretable models while producing stable regime-level explanations.
Chinese Translation
深度ReLU网络是分段仿射映射,将输入空间划分为多个单元,每个单元具有独特的激活模式。这种结构促使我们在每个单元内拟合一个局部线性模型,以保持预测准确性的同时提高可解释性。挑战在于识别出稳定、数据自适应且易于解释的机制。我们提出了REALM,一种线性模型的混合,其机制由神经激活模式诱导。由于深度神经网络(DNN)中的激活单元数量随深度快速增长,我们首先将深度教师网络蒸馏为一个宽浅学生网络(WSSN),然后对其隐藏层激活进行二值化并聚类,以定义机制并在每个机制内拟合线性模型。由于机制是从内部结构发现的,路由器不承担预测负担。为了使机制分配可解释,我们训练一个多类逻辑回归,即解释性门控,来复现机制分配。这种两级结构在原始表格特征或学习到的卷积特征方面在两个阶段都是可解释的:门控识别决定机制分配的特征,而线性模型识别驱动每个机制内预测的特征。我们分析了一个理想化设置,说明了分区复杂度与稳定性之间的权衡:随着机制数量的增加,更细的分区可以提高近似能力,但可能降低机制分配的稳定性。在表格和图像数据集上的实验表明,REALM相对于其他DNN引导的混合代理模型和内在可解释模型,取得了有竞争力的预测性能,同时产生稳定的机制级解释。
cs.LG / 68 / 2609.32143

Spectral Reversal: Counteracting Singular Value Bias for Graph Prompting

谱反转:抵消图提示中的奇异值偏置
Yang, Hanxu, Zhao, Yuhuan, He, Xiaodong, Kang, Zhao
Abstract
Pre-training Graph Neural Networks (GNNs) via self-supervised learning has become a dominant paradigm, yet efficiently adapting frozen encoders remains a challenge. Graph prompting offers a parameter-efficient alternative to fine-tuning, but existing methods largely treat pre-trained models as opaque feature extractors, ignoring their internal spectral structure. In this work, we identify a systematic phenomenon in pre-trained GNNs, which we term spectral bias: optimization during pre-training disproportionately aligns representations with directions associated with large singular values, leaving low-energy directions under-explored. We show that these underutilized directions can encode complementary information that is beneficial for downstream adaptation, especially under distribution shift. To leverage this insight, we propose Spectral Reverse Prompt (SRP), a prompting framework that rebalances the spectral contributions of frozen GNN encoders. SRP applies a learnable soft-thresholding mask in the spectral domain to down-weight dominant directions while amplifying weaker ones. In addition, SRP incorporates a null-space augmentation module that captures variation in directions with minimal activation under the frozen encoder. Extensive experiments across multiple benchmarks demonstrate that SRP achieves state-of-the-art performance with minimal additional parameters, highlighting that reweighting spectral components is a principled and effective strategy for parameter-efficient graph adaptation.
Chinese Translation
通过自监督学习预训练图神经网络(GNN)已成为一种主导范式,然而高效适配冻结编码器仍是一项挑战。图提示为微调提供了一种参数高效的替代方案,但现有方法大多将预训练模型视为不透明的特征提取器,忽视了其内部谱结构。在本工作中,我们发现了预训练GNN中一个系统性现象,并将其称为谱偏置:预训练过程中的优化会不成比例地将表示与较大奇异值所关联的方向对齐,导致低能量方向未被充分探索。我们表明,这些未被充分利用的方向可以编码对下游适配有益的补充信息,尤其是在分布偏移下。为利用这一洞见,我们提出谱反向提示(Spectral Reverse Prompt, SRP),一种重新平衡冻结GNN编码器谱贡献的提示框架。SRP在谱域中应用可学习的软阈值掩码,以降低主导方向的权重,同时放大较弱方向。此外,SRP引入一个零空间增强模块,用于捕获在冻结编码器下激活最小的方向上的变化。在多个基准上的大量实验表明,SRP以极少的额外参数实现了最先进的性能,突显了重新加权谱成分是一种用于参数高效图适配的原则性且有效的策略。
cs.LG / 69 / 2609.32146

Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions

达到标准杆:用于可证明最优四边形块分解的强化学习
Narayanan, Arjun, Persson, Per-Olof
Abstract
A quadrilateral block decomposition of a planar domain is judged by whether it is complete, whether its elements are well shaped, and how many of its vertices are irregular. The last has a provable floor: the discrete Gauss-Bonnet identity enforces a lower bound on the total vertex irregularity of any all-quadrilateral mesh of a given domain purely based on its topology and corner angles. We train a reinforcement learning agent to build decompositions that reach this bound, which we call par. It acts directly on the mesh's half-edge data structure through local edits, with a policy network whose convolutions follow the mesh's own connectivity, so it applies unchanged to domains larger than any seen in training. The reward targets the floor directly, and it is sparse: random play reaches it on no domain with more than eight sides. We overcome this exploration barrier via behaviour cloning on optimal meshes that are trivial to construct, walked backward into demonstrations, before training it with PPO. On 96 held-out domains the agent produces an all-quadrilateral mesh on every one, a usable one on 95.7 on average, and a provably optimal one on 90; Gmsh's strongest configuration at the same element count completes 51, is usable on 38 and optimal on none, and even at three to fourteen times the elements never produces a more regular mesh. On 64 domains twice the training size the agent completes all, is usable on 62, and keeps a median excess over par below one against Gmsh's 39 at the same element count.
Chinese Translation
平面域上的四边形块分解的好坏取决于它是否完整、其元素是否形状良好以及有多少个不规则顶点。最后一个指标有一个可证明的下界:离散高斯-博内恒等式根据给定域的拓扑和角点角度,对任何全四边形网格的总顶点不规则性施加了一个下界。我们训练一个强化学习智能体来构建达到这个下界的分解,我们将这个下界称为标准杆(par)。它通过局部编辑直接作用于网格的半边数据结构,其策略网络的卷积遵循网格自身的连接关系,因此可以不变地应用于比训练中见过的任何域都大的域。奖励直接针对这个下界,并且是稀疏的:随机操作在任何超过八条边的域上都无法达到它。我们通过在易于构造的最优网格上进行行为克隆来克服这一探索障碍,这些最优网格被反向遍历成演示,然后用PPO进行训练。在96个保留域上,智能体在每个域上都生成了全四边形网格,平均在95.7个域上生成可用的网格,在90个域上生成可证明最优的网格;而Gmsh在相同单元数量下最强的配置完成了51个,在38个上可用,没有一个是优的,即使单元数量增加三到十四倍,也从未生成更规则的网格。在64个训练规模两倍的域上,智能体完成了全部,在62个上可用,并且与相同单元数量下的Gmsh的39相比,其中位数超出标准杆的值保持在1以下。
cs.LG / 70 / 2609.32167

CAFE: Counterfactual Prediction via Fast Posterior Estimation

CAFE:基于快速后验估计的反事实预测
Han, Xinyan, Lin, Xiaoyu, Zou, Hao, Zhang, Xingxuan, Li, Bo, Cui, Peng
Abstract
Counterfactual prediction estimates an individual's outcome under an alternative intervention given their factual observations. Such outcomes are generally not identifiable from observational data without additional assumptions. Even within the class of fully observed additive noise models (ANMs), different causal graphs can generate the same observational distribution yet imply different individual counterfactual outcomes. Predictions based on a single estimated graph ignore this structural uncertainty. We therefore target a Bayesian counterfactual posterior predictive distribution that combines predictions from plausible SCMs. We introduce CAFE (\textbf{C}ounterf\textbf{A}ctual Prediction via \textbf{F}ast Posterior \textbf{E}stimation), an amortized inference framework that directly approximates the Bayesian counterfactual posterior predictive distribution. We pretrain a transformer-based model on synthetic counterfactual tasks generated from a diverse prior over ANMs. Given an observational dataset, an individual's factual observations, and an intervention, CAFE approximates the corresponding posterior predictive distribution in a single forward pass. Experiments show that CAFE accurately predicts individual counterfactual outcomes in identifiable settings and approximates the posterior predictive distribution when structural uncertainty induced by observationally indistinguishable causal graphs exists. Strong performance in realistic manufacturing and viticulture settings further demonstrates its empirical robustness beyond the assumptions of the training prior.
Chinese Translation
反事实预测旨在给定个体的事实观测下,估计其在替代干预下的结果。如果没有额外的假设,这类结果通常无法从观测数据中识别。即使在完全观测的加性噪声模型(ANMs)这一类中,不同的因果图可以生成相同的观测分布,却意味着不同的个体反事实结果。基于单个估计图的预测忽略了这种结构不确定性。因此,我们针对一个贝叶斯反事实后验预测分布,它结合了来自合理的结构因果模型(SCMs)的预测。我们提出了 CAFE(Counterfactual Prediction via Fast Posterior Estimation),一种摊销推理框架,它直接近似贝叶斯反事实后验预测分布。我们在由 ANMs 上的多样先验生成的合成反事实任务上预训练了一个基于 Transformer 的模型。给定观测数据集、个体的事实观测和一个干预,CAFE 在一次前向传播中近似相应的后验预测分布。实验表明,CAFE 在可识别设置中准确预测个体反事实结果,并在存在由观测上不可区分的因果图引起的结构不确定性时,近似后验预测分布。在现实的制造业和葡萄栽培设置中的强大性能进一步证明了其超越训练先验假设的经验稳健性。
cs.LG / 71 / 2609.32170

Uncertainty-Aware Selection of Online Algorithms with Simulator Ensembles

基于模拟器集成的不确定性感知在线算法选择
Guo, Yongyi, Xu, Zifan, Xu, Ziping, Zhang, Kelly W.
Abstract
The performance of online reinforcement learning depends critically on design choices, especially those that affect exploration. These choices are often selected by fitting a simulator to offline data, evaluating candidate algorithms in that simulator, and deploying the best-performing one. The simplest Plug-In selection rule simply selects the best performing algorithm on the fitted simulator, making evaluations unreliable when the offline data used to fit the simulator are limited. We investigate Uncertainty-Aware selection, which forms an ensemble of simulators---for example, obtained by bootstrap resampling---and selects the online algorithm with the best average performance across the ensemble. While ensemble-based approaches have been used to mitigate distribution shift and facilitate sim-to-real transfer, we formally show that this approach can mitigate the effects of limited data when fitting the simulator and theoretically has significant regret gains compared to Plug-In selection in multi-armed bandits. We also empirically investigate the Uncertainty-Aware selection approach in deep RL experiments on robotic control tasks that involve selecting reward-shaping hyperparameters, and show that it leads to more reliable selection and improved online performance.
Chinese Translation
在线强化学习的性能关键取决于设计选择,尤其是那些影响探索的选择。这些选择通常通过以下方式做出:将模拟器拟合到离线数据,在该模拟器中评估候选算法,并部署表现最佳的算法。最简单的 Plug-In 选择规则只是选择在拟合模拟器上表现最好的算法,当用于拟合模拟器的离线数据有限时,这会使评估不可靠。我们研究了不确定性感知选择(Uncertainty-Aware selection),它形成一个模拟器集成——例如通过 bootstrap 重采样获得——并选择在集成中平均性能最佳的在线算法。虽然基于集成的方法已被用于缓解分布偏移并促进 sim-to-real 迁移,但我们从理论上表明,这种方法可以减轻拟合模拟器时有限数据的影响,并且在多臂老虎机中,与 Plug-In 选择相比,具有显著的 regret 优势。我们还通过深度 RL 实验,在涉及选择奖励塑形超参数的机器人控制任务中,对不确定性感知选择方法进行了实证研究,并表明它能够带来更可靠的选择和更好的在线性能。
cs.LG / 72 / 2609.32176

Efficient Support Recovery of Mixtures of Sparse Linear Classifiers with Less Measurements

更少测量下稀疏线性分类器混合的高效支撑恢复
Li, Xiaxin, Mazumdar, Arya
Abstract
The support recovery problem in mixture of linear classifiers intends to identify which features actually matter when data is generated by a mixture of several linear decision rules. In particular, the aim is to recover the support (nonzero coordinates) of $l$ unknown $k$-sparse vectors from sign measurements. Each measurement is generated by selecting one of the $l$ vectors uniformly at random, and returning the sign of its inner product with a chosen measurement vector. In this paper, we propose adaptive and non-adaptive schemes that significantly improve upon prior results by reducing the number of measurements and achieving sublinear decoding time simultaneously. In particular, our adaptive constructions substantially reduce measurements compared to existing approaches, while also lowering decoding complexity from super-quadratic to sublinear in the ambient dimension. We further provide a non-adaptive scheme that improves previous measurement bounds while maintaining efficient decoding. Overall, our approach yields a more efficient trade-off between sample complexity and decoding time for support recovery in mixture models than previously known methods.
Chinese Translation
线性分类器混合中的支撑恢复问题旨在识别当数据由多个线性决策规则混合生成时,哪些特征真正重要。特别地,目标是从符号测量中恢复 $l$ 个未知的 $k$-稀疏向量的支撑(非零坐标)。每个测量通过均匀随机选择 $l$ 个向量之一,并返回其与所选测量向量的内积的符号来生成。在本文中,我们提出了自适应和非自适应方案,通过减少测量次数并同时实现亚线性解码时间,显著改进了先前的结果。特别地,与现有方法相比,我们的自适应构造显著减少了测量次数,同时将解码复杂度从超二次降低到环境维度上的亚线性。我们进一步提供了一种非自适应方案,在保持高效解码的同时改进了先前的测量界限。总体而言,对于混合模型中的支撑恢复,我们的方法在样本复杂度和解码时间之间产生了比先前已知方法更有效的权衡。
cs.LG / 73 / 2609.32178

Analytic-Walk Rotary Positional Encodings for Graphs

图上的 Analytic-Walk 旋转位置编码
Xie, Jiaqing, Wang, Yuxin, Qiu, Xipeng
Abstract
Rotary position encodings make attention sensitive to relative position, but extending them to graphs requires choosing how graph structure enters the rotation. Previous works assign each node a rotation from spectral coordinates, so the rotary factor between two nodes depends only on their endpoints and cannot distinguish the routes connecting them. We introduce \textit{Analytic-Walk Rotary Positional Encodings} (AW-RoPE), which place the rotations on edges and sum the transported features over all walks, so contributions along different routes can reinforce or cancel. An exact variant evaluates the complete sum by a differentiable linear solve, and a sparse variant truncates it at a finite depth. We prove forward and parameter-derivative truncation bounds at fixed inputs and parameters. Both variants act on projected queries and keys, and the sparse recurrence also augments message-passing networks. Across five synthetic tasks both variants reduce nRMSE by $15$--$58\%$ relative to the strongest baseline, and on real superpixel, peptide and OGB benchmarks the sparse recurrence attains the best mean on every dataset with Performer kernels and on twelve of thirteen datasets with GIN. Analysis shows that AW-RoPE can distinguish routes whose only cue is how two endpoints are connected, while node-wise rotary encodings cannot.
Chinese Translation
旋转位置编码使注意力对相对位置敏感,但将其扩展到图需要选择图结构如何进入旋转。先前的工作从谱坐标中为每个节点分配一个旋转,因此两个节点之间的旋转因子仅取决于它们的端点,无法区分连接它们的路径。我们引入 \textit{Analytic-Walk 旋转位置编码} (AW-RoPE),它将旋转放在边上,并沿所有游走对传输的特征求和,因此沿不同路径的贡献可以增强或抵消。一个精确变体通过可微线性求解评估完整和,而稀疏变体在有限深度处截断它。我们证明了在固定输入和参数下的前向和参数导数截断界限。两种变体都作用于投影的查询和键,并且稀疏递推还增强了消息传递网络。在五个合成任务上,两种变体相对于最强基线将 nRMSE 降低了 $15$--$58\%$,在真实的超像素、肽和 OGB 基准上,稀疏递推在使用 Performer 核时在每个数据集上均达到最佳均值,在使用 GIN 时在十三个数据集中的十二个上达到最佳均值。分析表明,AW-RoPE 可以区分仅以两个端点如何连接为线索的路径,而节点级旋转编码则不能。
cs.LG / 74 / 2609.32187

Rethinking Cross-Channel Importance in Time-Series Forecasting

重新思考时间序列预测中的跨通道重要性
Choi, Yong-Hoon, Park, Kwang-Hyun, Cho, Youngjin
Abstract
Cross-channel modeling is central to multivariate time-series forecasting, yet channels that are statistically related, predictively useful, and actually used by a trained forecaster are often treated as if they defined the same notion of importance. We show that they need not coincide. Cross-channel dependency structures change substantially across future offsets, and horizon-adaptive source selection improves a controlled Ridge predictor in 21 of 32 dataset--prediction-length conditions, with a mean gain of $5.16\%$. This selected-set signal also transfers to a matched nonlinear predictor. Yet imposing the same horizon-specific source logic on iTransformer yields only 11 of 20 wins and a mean gain of $0.208\%$, with little alignment between controlled and neural gains. Functional interventions further show that strong forecasters use cross-channel information, while their source-reliance rankings agree little with controlled utility or with one another across iTransformer, TimesNet, and a cross-channel TimeMixer. As a constructive consequence, bounded post-hoc support improves a frozen channel-independent forecaster in 12 of 16 dataset--horizon conditions, with a positive aggregate bootstrap interval. Cross-channel importance should therefore be interpreted relative to the forecasting mechanism and question that define it: related $\neq$ useful $\neq$ used.
Chinese Translation
跨通道建模是多变量时间序列预测的核心,但统计相关、预测有用且实际被训练好的预测器所使用的通道,往往被当作定义了相同的重要性概念。我们表明它们不一定重合。跨通道依赖结构在不同未来偏移上发生显著变化,视界自适应的源选择在32个数据集-预测长度条件中的21个改进了受控Ridge预测器,平均增益为5.16%。这种选择集信号也迁移到匹配的非线性预测器。然而,将相同的特定视界源逻辑施加于iTransformer仅产生20次胜出中的11次,平均增益为0.208%,且受控增益与神经增益之间几乎没有对齐。功能干预进一步表明,强预测器使用跨通道信息,而它们的源依赖排名与受控效用或彼此之间在iTransformer、TimesNet和跨通道TimeMixer上一致性很低。作为一个建设性结果,有界的事后支持在16个数据集-视界条件中的12个改进了冻结的通道独立预测器,并具有正向的聚合自助法区间。因此,跨通道重要性应相对于定义它的预测机制和问题来解释:相关 ≠ 有用 ≠ 被使用。
cs.LG / 75 / 2609.32213

HM-ROUTER: Joint Model and Harness Routing for Agentic Systems

HM-ROUTER:面向智能体系统的联合模型与执行框架路由
Chen, Hao Mark, Lee, Royson, Okoshi, Yasuyuki, Anastasiou, Dimitris, Luk, Wayne, Fan, Hongxiang
Abstract
Agent performance depends on both the underlying model and the harness that manages its tool use and execution. Selecting a suitable pair requires accounting for their compatibility, yet training samples may cover only a subset of the growing combination space. We introduce HM-Router, a routing method that jointly selects a model and harness for each query. It learns separate model and harness representations shared across routes, with an interaction term inspired by canonical polyadic (CP) tensor decomposition to capture how their compatibility varies with the query. This sharing allows training samples from observed pairs to inform predictions for unobserved combinations. We curate a benchmark from 12 public agent benchmarks, covering 293 routes, 73 models, and 25 harnesses. HM-Router exceeds the strongest evaluated learned baseline by 7.3 percentage points in mean routing accuracy and leads at all seven evaluated cost budgets on the six-benchmark subset. When 90% of routes have their training outcomes withheld, allowing unobserved combinations improves normalized accuracy by 15.8 points over restricting the same router to observed routes. HM-Router has also demonstrated training sample efficiency for new routes and components and generalization to unseen benchmarks. Our code and data are open-sourced at https://github.com/hmarkc/HM-Router.
Chinese Translation
智能体的性能既取决于底层模型,也取决于管理其工具使用与执行的执行框架。选择合适的模型-执行框架组合需要考虑它们之间的兼容性,然而训练样本可能仅覆盖不断增长的组合空间中的一部分。我们提出了HM-Router,一种为每个查询联合选择模型和执行框架的路由方法。它学习跨路由共享的独立模型表示和执行框架表示,并引入受典范多项式(CP)张量分解启发的交互项,以捕捉它们之间的兼容性如何随查询变化。这种共享使得来自已观察组合的训练样本能够为未观察组合的预测提供信息。我们从12个公开的智能体基准中整理了一个基准,涵盖293条路由、73个模型和25个执行框架。HM-Router在平均路由准确率上超过最强的已评估学习基线7.3个百分点,并在六个基准子集的所有七个评估成本预算上领先。当90%的路由的训练结果被隐藏时,允许未观察组合比将同一路由器限制在已观察路由上,将归一化准确率提高了15.8个百分点。HM-Router还展示了对于新路由和组件的训练样本效率,以及对未见基准的泛化能力。我们的代码和数据已在https://github.com/hmarkc/HM-Router上开源。
cs.LG / 76 / 2609.32215

DP-Rec: Towards Dynamic Patching for Efficient Long-Sequence Recommendation

DP-Rec:面向高效长序列推荐的动态分块
Katariya, Dwipam, Caputo, Thomas, Shreemali, Akshat, Origgi, Juan Manuel, Seleznev, Nikita, Mohanty, Pranab, Mishra, Kalanand, Nguyen, Nam, Montgomery, James
Abstract
Transformers have redefined sequential recommendation by effectively modeling dynamic user behaviors and long-range dependencies. However, they remain inherently inefficient: standard architectures operate at a fixed rate, allocating comparable computation to every item in a user's history regardless of its information content. This leads to prohibitive computational overhead on long sequences and increased sensitivity to behavioral noise. To address this, practitioners often resort to lossy sequence compression, staged modeling, or truncation. This limits the model's ability to leverage the full context of long histories during inference. Inspired by the recent success of Byte Latent Transformer, we propose DP-Rec, a dynamic latent patching architecture for recommendation. DP-Rec shifts from item-level modeling to patch-level modeling by segmenting interaction sequences using contrastive entropy surprise to identify informative behavioral boundaries. A lightweight patch encoder compresses these temporally contextualized segments into a reduced set of dynamic latent behavior vectors, which are then processed by a larger latent transformer and decoded for next-item prediction. Extensive experiments show that, under constrained computational budgets, DP-Rec scales effectively to long sequences and achieves a superior efficiency-accuracy trade-off over both non-compressed and fixed-size compression baselines.
Chinese Translation
Transformer 通过有效建模动态用户行为和长距离依赖,重新定义了序列推荐。然而,它们本质上仍然效率低下:标准架构以固定速率运行,无论用户历史中每个项目的信息内容如何,都分配相当的计算量。这导致长序列上高昂的计算开销,并增加了对行为噪声的敏感性。为了解决这个问题,从业者通常诉诸有损序列压缩、分阶段建模或截断。这限制了模型在推理过程中利用长历史完整上下文的能力。受近期 Byte Latent Transformer 成功的启发,我们提出 DP-Rec,一种用于推荐的动态潜在分块架构。DP-Rec 通过使用对比熵惊奇度对交互序列进行分割,以识别有信息量的行为边界,从而从项目级建模转向分块级建模。一个轻量级的分块编码器将这些具有时间上下文的片段压缩为一组缩减的动态潜在行为向量,然后由更大的潜在 Transformer 处理,并解码用于下一项预测。大量实验表明,在受限的计算预算下,DP-Rec 能够有效扩展到长序列,并在非压缩和固定大小压缩基线上实现了更优的效率-准确性权衡。
cs.LG / 77 / 2609.32228

CompassPlay: Rewarding the Proposer for Where It Moves the Solver

CompassPlay:奖励提议者将求解器引向何方
Pu, Sophia Xiao, Sun, Ximeng, Liu, Jiang, Wu, Jialian, Barsoum, Emad, Liu, Zicheng, Wang, William Yang
Abstract
In self-play, a proposer generates verifiable tasks to train a solver. Proposer rewards often depend on the solver's success rate, but equally difficult tasks can differ in their training value. We introduce CompassPlay, a self-play method that rewards the proposer through gradient alignment. The reward favors tasks whose solver loss gradients align with those of reference tasks representing the target capabilities. It draws on a first-order approximation to learning progress and scores each eligible task without additional solver training. Our experiments show gains in performance and training efficiency. In coding self-play with Qwen2.5-Coder-7B, a small external reference set guides task generation. CompassPlay improves average accuracy over AZR's difficulty reward by 1.5 percentage points on in-domain coding and 2.7 on out-of-domain mathematics. In Lean4 theorem proving, CompassPlay matches the difficulty baseline's 150-iteration cumulative coverage with 40\% fewer GPU-hours.
Chinese Translation
在自博弈中,提议者生成可验证的任务来训练求解器。提议者的奖励通常取决于求解器的成功率,但同等难度的任务在训练价值上可能不同。我们提出了CompassPlay,一种通过梯度对齐来奖励提议者的自博弈方法。该奖励偏向于那些求解器损失梯度与代表目标能力的参考任务的梯度对齐的任务。它利用对学习进度的一阶近似,在无需额外训练求解器的情况下对每个符合条件的任务进行评分。我们的实验表明,在性能和训练效率上均有提升。在使用Qwen2.5-Coder-7B进行编码自博弈时,一个小的外部参考集指导任务生成。CompassPlay在域内编码任务上比AZR的难度奖励平均准确率提高了1.5个百分点,在域外数学任务上提高了2.7个百分点。在Lean4定理证明中,CompassPlay以少40%的GPU小时匹配了难度基线在150次迭代时的累积覆盖率。
cs.LG / 78 / 2609.32230

"Where Can I Trust You?": Boundary-Aware Evaluation of Surrogate Fidelity

“我能在哪里信任你?”:代理保真度的边界感知评估
Eshbaugh, Jackson
Abstract
Surrogate models are commonly evaluated by how often they agree with their teacher model over an evaluation set. Local variation in this agreement is well known, but its structure and consequences are less clear. We ask whether disagreement is systematically concentrated near the teacher's decision boundary and whether retaining that structure provides information beyond a single global score. Across several datasets and surrogate model classes, we find substantially lower fidelity near teacher decision boundaries under two different methods of identifying near-boundary examples. Moreover, conditioning agreement on confidence-defined regions improves prediction of teacher--surrogate agreement when evaluation-set composition changes, relative to the global score alone. Yet surrogates that agree equally well with the teacher both globally and near the decision boundary can respond very differently to changes selected using the surrogate itself. Finally, we show that independently trained deep teachers can agree on most predictions while identifying different examples as lying near their decision boundaries, complicating the use of those boundaries as stable reference regions for evaluating surrogates. Together, these results show that surrogate fidelity depends not only on how often a surrogate agrees with its teacher, but also on where that agreement holds and, for deep models, how stable the teacher's decision boundary is across training runs.
Chinese Translation
代理模型通常通过其在评估集上与教师模型的一致程度来评估。这种一致性的局部变化是众所周知的,但其结构和后果尚不明确。我们探究不一致是否系统地集中在教师模型的决策边界附近,以及保留该结构是否能提供超越单一全局分数的信息。在多个数据集和代理模型类别上,我们发现,在两种不同的识别近边界样本的方法下,教师决策边界附近的保真度显著较低。此外,将一致性条件化于置信度定义的区域,当评估集组成发生变化时,相对于仅使用全局分数,能提高对教师-代理一致性的预测。然而,那些在全局和决策边界附近都与教师模型一致程度相同的代理模型,对于使用代理模型自身所选择的变化,可能表现出非常不同的响应。最后,我们表明,独立训练的深度教师模型可能在大多数预测上一致,但将不同的样本识别为位于其决策边界附近,这使得将这些边界作为评估代理模型的稳定参考区域变得复杂。总之,这些结果表明,代理保真度不仅取决于代理模型与教师模型的一致频率,还取决于这种一致性在何处成立,并且对于深度模型,还取决于教师模型的决策边界在不同训练运行中的稳定性。
cs.LG / 79 / 2609.32240

Arithmetic Simplicity in Stochastic Gradient Methods

随机梯度方法中的算术简单性
Fu, Bin, Gu, Pengfei, Nunez, Jose, Vazquez, Fabian
Abstract
A gradient descent method is arithmetically simple if the operations are limited to $+,-, \times$, and division $x/2^t$ with integer $t$. An arthmetically simple gradient method is easy to implement in chip design. We show how to transform AdaGrad, Adam, and AdamW into arithmetically simple. AdamW is based on the recursion $x_{t+1}=(1-\lambda\eta)x_t-\frac{\eta }{s}m_t$ and Adam is the special case of AdamW with $\lambda=0$. We transform them into a static case with $s=S(T)$, where $T$ is the number of iterations, and $S(T)$ is a fixed function. The convergence analysis is given for a static Adam, which is also arithmetically simple.
Chinese Translation
如果梯度下降法的运算仅限于加法、减法、乘法以及整数$t$的除法$x/2^t$,则称该方法是算术简单的。算术简单的梯度方法易于在芯片设计中实现。我们展示了如何将AdaGrad、Adam和AdamW转换为算术简单的方法。AdamW基于递推公式$x_{t+1}=(1-\lambda\eta)x_t-\frac{\eta }{s}m_t$,而Adam是AdamW在$\lambda=0$时的特例。我们将它们转换为静态情况,其中$s=S(T)$,$T$是迭代次数,$S(T)$是一个固定函数。我们给出了静态Adam的收敛性分析,该分析也是算术简单的。
cs.LG / 80 / 2609.32248

Anytime-Valid LLM Leaderboards via Benchmark-weighted and Block-Factorized e-Processes

基于基准加权和块因子化e-过程的随时有效LLM排行榜
Gao, Hongfu, Zhang, Songxin, Xie, Zejian, Jing, Bingyi, Wang, Zhou, Liu, Yiming
Abstract
Large language model (LLM) leaderboards compare model capabilities by ranking models according to their mean performance on fixed benchmarks. However, variability in evaluation outcomes across runs may produce unsupported claims of model superiority on the benchmark, a risk compounded by leaderboard updates. In this paper, we propose BB-EDGE (Benchmark-Weighted and Block-Factorized e-processes for Directed Graph Evaluation), a principled framework that represents an LLM leaderboard as a directed graph whose edges certify pairwise mean-performance advantages, with anytime-valid family-wise error rate (FWER) control. Concretely, for each direction, BB-EDGE constructs an empirical-Bernstein e-process by factorizing evidence over protocol-defined blocks and assigning stakes proportional to the corresponding block weights, then applies direct e-Holm across these $e$-processes to certify directional advantages as edges. Theoretically, we characterize weight-proportional linear stakes under heterogeneous benchmark-average nulls and prove anytime FWER control under arbitrary within-block and cross-pair dependence. BB-EDGE further supports anytime-valid Top-$k$ certification and simultaneous rank intervals. Extensive experiments on synthetic data and four real-world benchmarks demonstrate that BB-EDGE maintains anytime FWER control while achieving high efficiency.
Chinese Translation
大语言模型(LLM)排行榜通过根据模型在固定基准上的平均性能对模型进行排名来比较模型能力。然而,不同运行之间评估结果的可变性可能产生对基准上模型优越性的无支持声明,而排行榜更新会加剧这一风险。在本文中,我们提出BB-EDGE(用于有向图评估的基准加权和块因子化e-过程),这是一个原则性框架,将LLM排行榜表示为有向图,其边证明成对平均性能优势,并具有随时有效的族错误率(FWER)控制。具体而言,对于每个方向,BB-EDGE通过在协议定义的块上分解证据并根据相应块权重分配份额来构建经验伯恩斯坦e-过程,然后在这些e-过程上应用直接e-Holm,以将方向性优势证明为边。理论上,我们刻画了在异质性基准平均零假设下的权重比例线性份额,并证明了在任意块内和跨对依赖下的随时FWER控制。BB-EDGE进一步支持随时有效的Top-k认证和同时秩区间。对合成数据和四个真实基准的广泛实验表明,BB-EDGE在保持随时FWER控制的同时实现了高效率。
cs.LG / 81 / 2609.32263

Representation Editing for Multimodal Test-Time Adaptation

面向多模态测试时自适应的表示编辑
Huang, Longfei, Wu, Xiangyu, Yang, Yang
Abstract
Multimodal test-time adaptation (TTA) aims to adapt a pretrained multimodal model online to distribution shift across modalities using unlabeled test data, showing broad potential in real-world applications. However, existing methods primarily focus on adjusting fused features to bridge the source-target gap, lacking explicit control over intermediate representation misalignment, which is a key driver of performance drop under distribution shift. In this work, we tackle this challenge from the perspective of representation engineering. Unlike previous TTA methods that update fusion weights in place, we propose FourIer Representation Editor (FIRE), a novel multimodal TTA approach that directly edits semantically rich intermediate representations. Specifically, we first adopt representation editors into each intermediate layer of the unimodal encoders, enabling layer-wise calibration of unimodal representations. To further enhance the diversity and stability of the low-rank editing subspaces, each representation editor performs frequency domain mixing via the fast Fourier transform to construct structured bases. Moreover, we introduce multi-level adaptation objectives to optimize these editors, jointly promoting cross-modal semantic alignment, source-target statistical alignment, and asymmetric prediction consistency. In this way, FIRE yields aligned unimodal representations for fusion and further improves prediction reliability. Extensive experiments on two widely used multimodal benchmarks under various corruption types demonstrate the superiority of FIRE over existing multimodal TTA methods.
Chinese Translation
多模态测试时自适应(TTA)旨在利用无标注测试数据,使预训练多模态模型在线适应跨模态的分布偏移,在实际应用中展现出广阔前景。然而,现有方法主要侧重于调整融合特征以弥合源-目标差距,缺乏对中间表示不对齐的显式控制,而这是分布偏移下性能下降的关键驱动因素。在这项工作中,我们从表示工程的角度应对这一挑战。与之前原地更新融合权重的TTA方法不同,我们提出了傅里叶表示编辑器(FIRE),一种新颖的多模态TTA方法,直接编辑语义丰富的中间表示。具体来说,我们首先将表示编辑器引入单模态编码器的每个中间层,实现单模态表示的逐层校准。为了进一步增强低秩编辑子空间的多样性和稳定性,每个表示编辑器通过快速傅里叶变换执行频域混合,以构建结构化基。此外,我们引入多级自适应目标来优化这些编辑器,共同促进跨模态语义对齐、源-目标统计对齐和非对称预测一致性。这样,FIRE为融合产生对齐的单模态表示,并进一步提高预测可靠性。在两个广泛使用的多模态基准上,针对各种损坏类型的大量实验证明了FIRE优于现有多模态TTA方法。
cs.LG / 82 / 2609.32267

When Can Old Evaluations Certify a New Model? Label-Efficient Release Decisions under Evaluator Drift

旧评估何时能认证新模型?评估器漂移下的标签高效发布决策
Mondal, Joyanta Jyoti, Banik, Mridul, Apurba, Md. Shifatul Ahsan, Mahmud, Md Masud Al
Abstract
Releasing a model update requires certifying that its current-population risk stays below a threshold. Trusted labels are expensive, while a cheap evaluator, such as an LLM judge, scores every example. Reusing evaluator errors from earlier audits is tempting, but when may such evidence replace current labels? It depends on the status of history. If the errors can change invisibly, no label-free test detects the change, and every valid, useful certifier must keep buying labels at a rate we characterize; if a bound on the change is assumed, label-free certification is valid at an explicit error cost. For the middle ground, where history is informative but untrusted, we propose \emph{portfolio vigilance}, a sequential certifier mixing a betting expert guided by history with one that learns only from current labels; history affects only how it bets, so validity holds for any history. The contribution is not prior-informed betting or expert mixtures, but separating history that may enter validity from history that may only guide label collection. In a canonical model, accurate history shortens decisions but never raises the evidence growth rate; stale history can destroy it. On held-out CIFAR-10N and DICES-990 data, portfolio vigilance needs 0.465 (95\% CI $[0.327,0.575]$) and 0.740 ($[0.618,0.877]$) times the labels of a matched prediction-powered monitor, with no observed false certification, and fewer labels on all six external blocks. Under corrupted advice it stays within 8.0\% of its better component, while trusting history alone costs up to 1.66 times as much. In post-confirmatory repeated-judge experiments on DICES-990 and ToxicChat, changing a fixed LLM judge's rubric moves its scores beyond run-to-run variation; the portfolio then needs 0.790 ($[0.667,0.909]$) and 0.631 ($[0.520,0.770]$) times the labels of the matched monitor, and fewer than trusting history alone.
Chinese Translation
发布模型更新需要认证其当前总体风险保持在阈值以下。可信标签昂贵,而廉价评估器(如LLM评判器)为每个样本打分。重用早期审计中的评估器误差很诱人,但这种证据何时可以替代当前标签?这取决于历史的状态。如果误差可以不可见地变化,则没有无标签测试能检测到这种变化,每个有效且有用的认证器都必须以我们刻画的速率持续购买标签;如果假设变化有界,则无标签认证在显式误差代价下有效。对于中间情况,历史有信息但不可信,我们提出组合警觉(portfolio vigilance),一种序贯认证器,混合由历史引导的投注专家与仅从当前标签学习的专家;历史仅影响其投注方式,因此对任何历史都保持有效性。贡献不在于先验信息投注或专家混合,而在于区分可进入有效性的历史与仅可指导标签收集的历史。在一个典型模型中,准确的历史缩短决策但从不提高证据增长率;陈旧的历史可能破坏它。在留出的CIFAR-10N和DICES-990数据上,组合警觉需要的标签量是匹配的预测驱动监测器的0.465倍(95% CI [0.327,0.575])和0.740倍([0.618,0.877]),未观察到假认证,并且在所有六个外部块上使用更少标签。在建议被腐蚀的情况下,它与更优组件的差距保持在8.0%以内,而仅信任历史成本高达1.66倍。在DICES-990和ToxicChat上的验证后重复评判实验中,改变固定LLM评判器的评分标准会使其得分超出逐次运行变异;组合随后需要匹配监测器标签量的0.790倍([0.667,0.909])和0.631倍([0.520,0.770]),并且少于仅信任历史。
cs.LG / 83 / 2609.32268

Self-Reconstruction Dynamics for Autoencoder Reconstruction Refinement

用于自编码器重建细化的自重建动力学
Iyatomi, Hitoshi
Abstract
Standard autoencoder (AE) inference uses a single encoder-decoder pass, though the latent may not be optimal for each sample under a fixed decoder. We ask whether a trained AE can reveal information for improving its own reconstruction. Repeated application of a frozen AE to its reconstruction produces transient image- and latent-space trajectories, termed Self-Reconstruction Dynamics (SRD). Although this degrades fidelity in the AEs studied here, SRD contains sample-specific information for correcting the reconstruction. We propose SRD-guided Reconstruction Refinement (SRD-RR), which predicts a latent correction from a short SRD with the AE frozen and no per-sample test-time optimization. We also introduce MSE-recov, an MSE recovery ratio relative to an empirical decoder-optimized reference. Across six datasets, SRD-RR recovers 38.6% of the empirically recoverable MSE gap with one transition and 45.3% with two. A two-transition variant trained without direct access to original images, using an SRD-derived pseudo-target, achieves 40.7% recovery and a 1.74 dB average PSNR gain. Removing trajectory information reduces the gain, while cross-sample trajectory assignment causes severe degradation, confirming strong sample specificity. Nonlinear SRD-conditioned refinement consistently outperforms fixed and trained linear latent correction. On a pretrained DINOv2-based representation autoencoder (RAE) with substantially different latent dynamics, SRD conditioning again improves a matched trajectory-free predictor. However, pixel-MSE latent refinement reveals a strong mismatch between pixel fidelity and perceptual quality, while the SRD-derived pseudo-target mitigates this degradation. Overall, SRD is a useful sample-specific refinement signal, while the objective determines how it translates into pixel and perceptual quality.
Chinese Translation
标准自编码器(AE)推理使用单次编码器-解码器传递,尽管在固定解码器下,潜在表示对每个样本可能并非最优。我们探究训练好的AE能否揭示用于改进其自身重建的信息。将冻结的AE重复应用于其重建结果,会产生瞬态的图像和潜在空间轨迹,称为自重建动力学(SRD)。尽管在此研究的AE中这会降低保真度,但SRD包含用于校正重建的样本特定信息。我们提出SRD引导的重建细化(SRD-RR),它在AE冻结且无需逐样本测试时优化的情况下,从短SRD中预测潜在校正。我们还引入MSE-recov,即相对于经验解码器优化参考的MSE恢复比率。在六个数据集上,SRD-RR通过一次过渡恢复38.6%的经验可恢复MSE差距,通过两次过渡恢复45.3%。一个两次过渡变体,在无法直接访问原始图像的情况下训练,使用SRD导出的伪目标,实现了40.7%的恢复和1.74 dB的平均PSNR增益。移除轨迹信息会降低增益,而跨样本轨迹分配会导致严重退化,证实了强样本特异性。非线性SRD条件细化始终优于固定和训练的线性潜在校正。在具有显著不同潜在动态的预训练DINOv2基于表示自编码器(RAE)上,SRD条件再次改进了匹配的无轨迹预测器。然而,像素MSE潜在细化揭示了像素保真度与感知质量之间的强烈不匹配,而SRD导出的伪目标减轻了这种退化。总体而言,SRD是一种有用的样本特定细化信号,而目标决定了它如何转化为像素和感知质量。
cs.LG / 84 / 2609.32271

Certification Frontiers for Gaussian LoRA: Independent Priors, Posterior Risk, and Prediction-Preserving Balancing

高斯LoRA的认证边界:独立先验、后验风险与保预测平衡
Mondal, Joyanta Jyoti, Shihab, Ibne Farabi
Abstract
Post-hoc Bayesian fine-tuning places Gaussians around trained low-rank adapters, yet a calibrated posterior does not by itself yield a useful generalization certificate. Such a posterior admits an informative PAC-Bayes certificate only when both the loss of its sampled predictors and its KL divergence from an admissible prior are small. In this research, we characterize this certification frontier for Gaussian LoRA posteriors and separate three interventions: changing the prior, changing the stochastic predictor, and changing only how its complexity is counted. First, an exact isotropic KL envelope eliminates the prior scale and yields a width threshold that excludes posterior widths before sampling, while a zero-KL floor identifies targets that no complexity reduction can reach at a measured risk bound. Second, we minimize KL in closed form over the full $GL(r)$ symmetry of the low-rank factors, leaving every sampled adapter product unchanged, and derive the noncentral objective required when the prior center is trained on an independent split. On a small-pool RoBERTa audit of 567 configurations, the recorded 64-draw summaries imply a certificate floor of $0.7298$ even with zero KL and Chernoff accounting, so reducing complexity alone cannot certify these posteriors at the recorded budgets. For a stable posterior in a controlled Gaussian-factor task, Chernoff accounting certifies risk below $0.1$ on 20 of 20 datasets with 1024 draws, whereas Hoeffding certifies none. On deliberately deformed synthetic rank-four factors, matrix balancing reduces KL by $29.3\%$ on average beyond scalar balancing without changing any prediction. Numerical split-prior scenarios make the remaining risk and complexity budgets explicit.
Chinese Translation
事后贝叶斯微调将高斯分布置于训练好的低秩适配器周围,但校准后的后验本身并不能产生有用的泛化认证。这样的后验只有在采样预测器的损失和与可接受先验的KL散度都很小时,才能提供有信息量的PAC-Bayes认证。在本研究中,我们刻画了高斯LoRA后验的这种认证边界,并区分了三种干预:改变先验、改变随机预测器,以及仅改变其复杂度的计算方式。首先,精确的各向同性KL包络消除了先验尺度,并产生一个宽度阈值,在采样前排除后验宽度,而零KL下界则识别出在测量风险界下任何复杂度降低都无法达到的目标。其次,我们在低秩因子的完整$GL(r)$对称性上以闭式最小化KL,保持每个采样适配器乘积不变,并推导出当先验中心在独立划分上训练时所需的非中心目标。在对567种配置的小池RoBERTa审计中,记录的64次抽取摘要意味着即使KL为零且使用Chernoff计算,认证下界也为$0.7298$,因此仅降低复杂度无法在记录的预算下认证这些后验。对于受控高斯因子任务中的稳定后验,Chernoff计算在1024次抽取下对20个数据集中的20个认证风险低于$0.1$,而Hoeffding则一个也没有认证。在故意变形的合成秩四因子上,矩阵平衡在不改变任何预测的情况下,平均比标量平衡额外降低KL $29.3\%$。数值分裂先验场景明确了剩余的风险和复杂度预算。
cs.LG / 85 / 2609.32272

Learning Through Game: Skewed Transfer of Tabular Knowledge to Strengthen Image Model

通过博弈学习:偏斜迁移表格知识以增强图像模型
Huang, Longfei, Yang, Shangdong, Yang, Yang
Abstract
Multimodal tabular-image learning is gaining growing attention, yet it faces challenges due to tabular data unavailable at test time. A practical solution involves transferring tabular knowledge to images during training to enhance the performance of image models at inference. However, the overlooked yet important challenges lie in the modality imbalance between images and tables, as well as their asymmetric modality relationship in cross-modal transfer, which limits the auxiliary role of tabular data. To address these issues, we propose Skewed Knowledge Transfer (SKT), which asymmetrically transfers tabular knowledge to improve the image model by adaptive integration of modality gradients in a shared parameter space. Specifically, we first introduce a multimodal shared head, which allows the model to benefit from cross-modal structure without adding additional parameters. We then design a two-step Nash Bargaining strategy to effectively leverage tabular gradients. In the first step, SKT seeks a point of modality balance and uses preference awareness in the second step to steer combined gradients toward image-beneficial directions. Furthermore, we theoretically analyze the Pareto improvement and convergence of SKT. To this end, tabular knowledge is explicitly transferred to enhance image models. Empirical experiments on widely used tabular-image datasets reveal that SKT consistently improves image unimodal performance by using tabular data as auxiliary information.
Chinese Translation
多模态表格-图像学习正受到越来越多的关注,然而由于测试时表格数据不可用,它面临着挑战。一种实用的解决方案是在训练过程中将表格知识迁移到图像,以提升图像模型在推理时的性能。然而,被忽视但重要的挑战在于图像和表格之间的模态不平衡,以及它们在跨模态迁移中的非对称模态关系,这限制了表格数据的辅助作用。为了解决这些问题,我们提出了偏斜知识迁移(SKT),它通过在共享参数空间中自适应地整合模态梯度,非对称地迁移表格知识来改进图像模型。具体来说,我们首先引入了一个多模态共享头,它允许模型在不增加额外参数的情况下受益于跨模态结构。然后,我们设计了一个两步纳什议价策略来有效地利用表格梯度。第一步,SKT寻求模态平衡点,并在第二步利用偏好感知将组合梯度引导向有利于图像的方向。此外,我们从理论上分析了SKT的帕累托改进和收敛性。为此,表格知识被显式地迁移以增强图像模型。在广泛使用的表格-图像数据集上的实证实验表明,SKT通过将表格数据作为辅助信息,持续提升了图像单模态性能。
cs.LG / 86 / 2609.32275

Topology-Adaptive Hyperbolic Graph Attention Networks Guided by the Hyperbolic Sombor Index

由双曲Sombor指数引导的拓扑自适应双曲图注意力网络
Cao, Haifang, Tao, Boan, Gao, Xiyuan, Li, Timing, Wang, Yu, Zhu, Pengfei
Abstract
Hyperbolic geometry has emerged as a principled space for representing hierarchical graphs. However, existing hyperbolic graph neural networks typically rely on shared curvature configurations and feature-driven attention, failing to explicitly exploit local hierarchical topological patterns. To bridge this gap, we introduce the Hyperbolic Sombor Index (HSO) as a lightweight structural prior for capturing hierarchy-indicative degree stratification. Building on this, we propose \textbf{HSO-GAT}, a topology-adaptive hyperbolic graph attention network that unifies geometric adaptation and message propagation. Specifically, it comprises two complementary modules: HSO-Guided Local Curvature Adaptation, which performs adaptive node-wise geometric scaling from aggregated node-level HSO signals, and HSO-Gated Hyperbolic Graph Attention, which enables structure-aware message passing through feature-conditioned gating. Theoretically, we establish the monotonic sensitivity of edge-level HSO to degree imbalance and analyze the validity and radial scaling properties of node-adaptive hyperbolic mappings. Extensive experiments on eight benchmark datasets demonstrate that HSO-GAT consistently achieves state-of-the-art performance in both node classification and link prediction tasks.
Chinese Translation
双曲几何已成为表示层次图的一种原则性空间。然而,现有的双曲图神经网络通常依赖于共享曲率配置和特征驱动的注意力,未能显式地利用局部层次拓扑模式。为了弥补这一差距,我们引入双曲Sombor指数(HSO)作为一种轻量级结构先验,用于捕获指示层次的度分层。在此基础上,我们提出HSO-GAT,一种拓扑自适应双曲图注意力网络,它统一了几何自适应和消息传播。具体而言,它包含两个互补模块:HSO引导的局部曲率自适应,它根据聚合的节点级HSO信号执行自适应的节点级几何缩放;以及HSO门控的双曲图注意力,它通过特征条件门控实现结构感知的消息传递。在理论上,我们建立了边级HSO对度不平衡的单调敏感性,并分析了节点自适应双曲映射的有效性和径向缩放性质。在八个基准数据集上的大量实验表明,HSO-GAT在节点分类和链接预测任务中均持续达到最先进的性能。
cs.LG / 87 / 2609.32276

HyperLabel: Multi-Label Classification via Hypergraph-Based Label Correlation Modeling

HyperLabel:基于超图的标签相关性建模的多标签分类
Zhang, Peiyu, Ping, Heng, Kanakaris, Nikos, Zhao, Yucheng, Li, Shixuan, Yang, Wei, Xiao, Xiongye, Bogdan, Paul
Abstract
Multi-label classification (MLC) requires predicting multiple relevant labels for each instance, where a central challenge is modeling complex label dependencies arising from co-occurrence patterns. Existing approaches are limited in capturing high-order label correlations, relying on implicit learning through contrastive objectives or pairwise attention mechanisms without structural guidance. We propose HyperLabel, an encoder-decoder framework that explicitly models label dependencies through hypergraph neural networks. Our contributions are twofold: (i) We construct a label hypergraph where sample-defined hyperedges naturally encode multi-way co-occurrence patterns, providing explicit structural prior knowledge that captures relationships beyond pairwise interactions. (ii) We propose a unified cross-modal learning approach where HGNN+ performs bidirectional message passing to integrate feature information with label structure, and a shared cross-attention decoder processes both modalities through complementary learning objectives. Extensive experiments on seven benchmark datasets demonstrate that HyperLabel achieves state-of-the-art performance, with particularly significant improvements on macro-F1 scores (+10.3% on Delicious, +8.2% on Bibtex), validating that explicit hypergraph structure effectively captures complex label relationships. The code is available at https://github.com/iZHpy/Multi-label_hypergraph .
Chinese Translation
多标签分类(MLC)需要为每个实例预测多个相关标签,其中核心挑战是建模由共现模式引起的复杂标签依赖关系。现有方法在捕捉高阶标签相关性方面存在局限,依赖于通过对比目标或成对注意力机制进行隐式学习,缺乏结构指导。我们提出HyperLabel,一个通过超图神经网络显式建模标签依赖关系的编码器-解码器框架。我们的贡献有两方面:(i) 我们构建了一个标签超图,其中由样本定义的超边自然地编码了多路共现模式,提供了显式的结构先验知识,捕捉了超出成对交互的关系。(ii) 我们提出了一种统一的跨模态学习方法,其中HGNN+执行双向消息传递以将特征信息与标签结构整合,并且一个共享的交叉注意力解码器通过互补的学习目标处理两种模态。在七个基准数据集上的大量实验表明,HyperLabel达到了最先进的性能,尤其是在macro-F1分数上有显著提升(Delicious上+10.3%,Bibtex上+8.2%),验证了显式超图结构有效捕捉复杂标签关系。代码可在https://github.com/iZHpy/Multi-label_hypergraph获取。
cs.LG / 88 / 2609.32279

SIMANF: Sample Free Learning of Unnormalized Distributions via Simulated Annealing in Normalizing Flows

SIMANF:通过归一化流中的模拟退火实现非归一化分布的无样本学习
Kanaujia, Vikas
Abstract
Efficiently learning and sampling from high dimensional, multimodal unnormalized distributions without target samples remains a challenging problem. Although normalizing flows can generate samples efficiently, training based on the reverse KL divergence using only the unnormalized target density may suffer from mode collapse. We introduce SIMANF, a sample free framework that integrates simulated annealing with normalizing flows. SIMANF progressively transforms the target distribution from a smooth initial form to the original target distribution and trains the flow sequentially across these stages. By transferring the learned representation between stages, the method promotes mode coverage while progressively capturing finer features of the target distribution. Following annealing, a final refinement stage combines the reverse KL divergence with an importance weighted forward KL objective using samples generated by the flow. SIMANF requires no target samples during training and uses only the unnormalized density. We demonstrate its effectiveness on Many-Well distributions and high dimensional Scalar Phi4 lattice field theory distribution.
Chinese Translation
高效地学习和采样高维、多模态的非归一化分布,且无需目标样本,仍然是一个具有挑战性的问题。尽管归一化流能够高效地生成样本,但仅使用非归一化目标密度基于反向KL散度的训练可能会遭遇模式崩溃。我们提出了SIMANF,一个将模拟退火与归一化流相结合的无样本框架。SIMANF逐步将目标分布从平滑的初始形式变换为原始目标分布,并在这些阶段中顺序地训练流。通过在阶段之间传递学习到的表示,该方法促进了模式覆盖,同时逐步捕捉目标分布的更精细特征。在退火之后,最终的细化阶段结合了反向KL散度和使用流生成的样本的重要性加权前向KL目标。SIMANF在训练过程中不需要目标样本,仅使用非归一化密度。我们在Many-Well分布和高维标量Phi4格点场论分布上展示了其有效性。
cs.LG / 89 / 2609.32288

A Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling Exponents

预训练数据投毒的可解理论:机制依赖的标度指数
Halder, Indranil, Dey, Rastri, Pehlevan, Cengiz
Abstract
Pre-training data poisoning of large language models is usually studied using targeted backdoors and their survival through safety post-training, which leaves open a more basic question: how does a model's clean data performance degrade as the poison rate $\varepsilon$ grows? Motivated by our controlled pre-training runs of OLMo-style models, in which the relative clean data validation perplexity increase $\Delta$ between poisoned and clean models matched in architecture, token budget, and optimization schedule is well fit by a power law $\Delta\approx C\varepsilon^{a}$ with a non-integer exponent, we ask what such a law requires theoretically. We first prove an analyticity barrier: whenever the contaminated objective depends analytically on $\varepsilon$ around a nondegenerate clean data optimum, $\Delta$ is generically quadratic in $\varepsilon$, so a generic non-integer exponent is a signature of genuinely singular structure. We then supply that structure in solvable truncated ridge regression with heavy-tailed covariates, controlled by $q_\star$, and a label-shift poisoning. Our central result is that the excess risk scaling exponent depending on the order of limits: in the higher dimensional proportional regime it is $\epsilon^{q_\star/(q_\star+2)}$, whereas taking the ample-data limit first gives $\epsilon^{2-2/q_\star}$, and the limits do not commute. We confirm this prediction through several numerical simulations. Finally, we argue that finite training time plays the role of truncation on the curvature spectrum in local LLM pre-training, deriving the observed scaling law under heavy tailed inverse curvature spectrum as a modeling hypothesis.
Chinese Translation
大型语言模型的预训练数据投毒通常通过定向后门及其在安全后训练中的存活性来研究,这留下了一个更基本的问题:随着投毒率 $\varepsilon$ 的增长,模型的干净数据性能如何退化?受我们受控的 OLMo 风格模型预训练运行的启发,其中在架构、token 预算和优化调度匹配的情况下,投毒模型与干净模型之间的相对干净数据验证困惑度增加 $\Delta$ 能很好地用幂律 $\Delta\approx C\varepsilon^{a}$ 拟合,且指数 $a$ 非整数,我们追问这样的规律在理论上需要什么条件。我们首先证明一个解析性障碍:只要受污染的目标在非退化的干净数据最优解附近解析依赖于 $\varepsilon$,$\Delta$ 通常就是 $\varepsilon$ 的二次函数,因此一般的非整数指数是真正的奇异结构的标志。然后我们在可解的截断岭回归中提供这种结构,该回归具有重尾协变量,由 $q_\star$ 控制,并伴有标签偏移投毒。我们的核心结果是,超额风险的标度指数取决于极限的顺序:在高维比例机制中,它是 $\epsilon^{q_\star/(q_\star+2)}$,而先取充足数据极限则得到 $\epsilon^{2-2/q_\star}$,并且这些极限不可交换。我们通过若干数值模拟证实了这一预测。最后,我们论证有限训练时间在局部 LLM 预训练中对曲率谱起到了截断的作用,并在重尾逆曲率谱下将观察到的标度律作为建模假设推导出来。
cs.LG / 90 / 2609.32290

A Journey to the Edge of Stability

通往稳定边缘之旅
Lee, Jaerin, Lee, Kyoung Mu
Abstract
It has recently been found that deep learning often occurs at the "edge of stability (EoS)," where the maximum Hessian eigenvalue of the model is stabilized at a value reciprocal to the learning rate. However, what happens before we reach that regime? We fix a deep learning problem and vary first order optimization methods with dense learning rate sweeps. We then track the characterizing quantities of a learning trajectory: the loss, the sharpness, and the alignment between consecutive gradients. To our surprise, if we scale the learning rate by the dc gain of the optimizer, these traces from the sweeps from different optimizers almost perfectly overlap across a large range of learning rates. The dc-normalized optimizers have another role that only becomes apparent in high learning rates: they select when the sharpness value detaches from this universal curve and enters the edge of stability. Upon this discovery, we specify three distinct regimes with respect to the dc-adjusted learning rate: the low-LR regime where the trajectory is nearly insensitive to the optimizer, the high-LR regime, where the optimizer governs the sharpness according to the EoS reciprocal rule, and the in-between mid-LR regime where so-called progressive sharpening originates independently of the optimizer. This distinguishes the role of the optimizer, the learning rate, and the model in shaping the learning progress.
Chinese Translation
最近研究发现,深度学习常常发生在“稳定边缘(EoS)”,此时模型的 Hessian 最大特征值稳定在学习率倒数对应的值上。然而,在到达该区域之前会发生什么?我们固定一个深度学习问题,并通过密集学习率扫描来变化一阶优化方法。然后,我们追踪学习轨迹的刻画量:损失、锐度以及连续梯度之间的对齐度。令人惊讶的是,如果我们按优化器的 DC 增益缩放学习率,来自不同优化器扫描的这些轨迹在很大学习率范围内几乎完全重合。DC 归一化优化器还有另一个作用,该作用仅在高学习率下才变得明显:它们决定了锐度值何时脱离这条通用曲线并进入稳定边缘。基于这一发现,我们针对 DC 调整后的学习率划分了三个不同的区域:低学习率区域,此时轨迹几乎对优化器不敏感;高学习率区域,此时优化器根据 EoS 倒数规则支配锐度;以及介于两者之间的中等学习率区域,此时所谓的渐进锐化独立于优化器而产生。这区分了优化器、学习率和模型在塑造学习进程中的作用。
cs.LG / 91 / 2609.32296

FUND: Density Flow for Sampling Unnormalised Distributions

FUND:用于采样未归一化分布的密度流
Kanaujia, Vikas, Arora, Vipul
Abstract
Efficient sampling from Boltzmann distributions is central to modelling complex physical systems. Markov Chain Monte Carlo (MCMC) methods suffer from critical slowing down, high autocorrelation, and poor mode-mixing, limiting their scalability. Recent advances, like Boltzmann Generators, offer a promising alternative but remain constrained by costly MCMC-based training, inefficient sampling, and poor ergodicity. We introduce an algorithm for learning Boltzmann distributions that does not require any true samples for training. Our approach draws inspiration from flow matching but departs fundamentally from sample-trajectory matching to distribution-trajectory matching. The algorithm iteratively reshapes the target distribution, using model generated samples to guide learning and ensure comprehensive mode coverage. We validate our method on standard benchmarks, including a 2D Gaussian mixture, Many-Well distributions, and high-dimensional scalar $\phi^4$ theory. The proposed approach not only improves sampling performance and accuracy over traditional MCMC and flow-based baselines but also establishes a new method for sample-free learning of complex physical distributions.
Chinese Translation
高效地从玻尔兹曼分布中采样是建模复杂物理系统的核心。马尔可夫链蒙特卡洛(MCMC)方法存在临界慢化、高自相关性和模式混合差等问题,限制了其可扩展性。近期进展,如玻尔兹曼生成器(Boltzmann Generators),提供了一种有前景的替代方案,但仍受限于昂贵的基于 MCMC 的训练、低效采样和较差的遍历性。我们提出一种学习玻尔兹曼分布的算法,该算法在训练时不需要任何真实样本。我们的方法受流匹配(flow matching)启发,但从样本-轨迹匹配根本性地转向分布-轨迹匹配。该算法迭代地重塑目标分布,利用模型生成的样本来引导学习并确保全面覆盖各模式。我们在标准基准上验证了我们的方法,包括二维高斯混合、多井(Many-Well)分布以及高维标量 $\phi^4$ 理论。所提方法不仅相较于传统 MCMC 和基于流的基线提升了采样性能与精度,还建立了一种用于复杂物理分布无样本学习的新方法。
cs.LG / 92 / 2609.32298

Refresh or Realize? Compute Allocation in Drifting Models

刷新还是实现?漂移模型(Drifting Models)中的计算分配
Chen, Sipeng, Zheng, Xu, Li, Shibo
Abstract
Drifting Models train a one-step generator by recomputing a finite-sample drift field at every iteration and taking an optimizer step toward the drifted target. The field says how generated samples should move, but the step is taken in parameters shared by all samples, so the motion the network actually makes need not match the motion it was given. This leaves a basic training question open: should extra compute go into fitting the current target more closely, or into recomputing the field? We study it on ImageNet 256x256. Holding the target fixed for k optimizer steps and measuring the realized displacement, we find that deeper fitting does bring the network closer to the frozen target, and that the number of steps needed before it makes any net progress drops from about sixteen early in training to one later on. When the extra steps come for free, k=2 also lowers FID. Once they are paid for, the result flips: at approximately matched measured wall-clock, spending the budget on fresh fields gives lower FID than deeper fitting, on both training seeds. The target itself shows why a fresh field is worth so much. Redrawing the finite support rotates its direction far more than a parameter update does (cosine ~0.3-0.6 against ~0.95), and a correction that is optimal in field space is not reliably better in FID than a parameter-free one. For Drifting, fitting each target well and spending compute well are different goals.
Chinese Translation
漂移模型(Drifting Models)通过在每个迭代重新计算有限样本漂移场,并朝着漂移目标采取优化器步来训练一步生成器。该场表示生成样本应如何移动,但该步是在所有样本共享的参数中采取的,因此网络实际做出的运动不必与其被给予的运动匹配。这留下了一个基本的训练问题:额外计算应该用于更紧密地拟合当前目标,还是用于重新计算场?我们在 ImageNet 256x256 上研究它。将目标固定 k 个优化器步并测量实现的位移,我们发现更深的拟合确实使网络更接近冻结目标,并且在其取得任何净进展之前所需的步数从训练早期的约十六步下降到后期的一步。当额外步数免费时,k=2 也降低了 FID。一旦需要付出代价,结果就反转:在近似匹配的实测挂钟时间下,将预算花在新鲜场上比更深拟合给出更低的 FID,在两个训练种子上都如此。目标本身显示了为什么新鲜场如此有价值。重新绘制有限支持集比参数更新更能旋转其方向(余弦约0.3-0.6对约0.95),并且在场地空间中最佳的校正,其 FID 不一定可靠地优于无参数校正。对于 Drifting,良好地拟合每个目标与良好地使用计算是不同的目标。
cs.LG / 93 / 2609.32301

When Does Synergy Help Active Feature Acquisition? A PID-Based Study

协同效应何时有助于主动特征获取?一项基于部分信息分解的研究
Li, Jie, Dhali, Maruf A., Bouma, Hjalmar R.
Abstract
Active feature acquisition (AFA) sequentially selects informative features under budget constraints. However, existing policies rarely distinguish whether information contributed by interacting features is redundant, unique, or synergistic. We introduce SynAFA, a state-dependent AFA policy that combines pairwise joint information and conditional information, with Partial Information Decomposition (PID) characterizing its information structure. Across five tabular datasets and MNIST-loop, performance is heterogeneous. SynAFA shows its strongest gains at low budgets on PhysioNet, where synergistic and redundant feature pairs are supported by permutation tests, but its advantage diminishes as budgets increase and does not depend on pair proposals. SynAFA performs significantly worse than nearly all baselines on MiniBooNE, and than CAE on MNIST-loop. Controlled synthetic experiments further show that, within budget, SynAFA's advantage rises as joint information becomes more synergy-dominated, including when total joint information is held approximately constant. Further analyses show that the budget-dependent erosion is not resolved by non-greedy local search, which improves a diagnostic set-level objective but leaves predictive performance no better, often significantly worse, and that improvements in this objective are weakly aligned with the fixed classifier's predictive utility. These findings characterize when pairwise synergy can benefit AFA while exposing a persistent challenge in translating local information into effective acquisition objectives.
Chinese Translation
主动特征获取(AFA)在预算约束下按顺序选择信息量大的特征。然而,现有策略很少区分由交互特征贡献的信息是冗余的、独特的还是协同的。我们提出了 SynAFA,一种状态依赖的 AFA 策略,它结合了成对联合信息和条件信息,并利用部分信息分解(PID)刻画其信息结构。在五个表格数据集和 MNIST-loop 上,性能表现各异。SynAFA 在 PhysioNet 上低预算时表现出最大增益,其中协同和冗余特征对得到置换检验的支持,但随着预算增加,其优势减弱,并且不依赖于特征对提议。在 MiniBooNE 上,SynAFA 的表现显著差于几乎所有基线;在 MNIST-loop 上则显著差于 CAE。受控合成实验进一步表明,在预算内,随着联合信息变得更加协同主导,SynAFA 的优势上升,包括总联合信息保持大致恒定的情况。进一步分析表明,依赖预算的削弱无法通过非贪心局部搜索解决;该搜索改善了一个诊断性的集合级目标,但预测性能并未更好,往往显著更差;并且该目标的改进与固定分类器的预测效用仅弱对齐。这些发现刻画了成对协同何时能够有益于 AFA,同时揭示了将局部信息转化为有效获取目标方面的一个持续挑战。
cs.LG / 94 / 2609.32302

On the Capability and Limitation of Hard Prompt

关于硬提示的能力与局限
Yu, Lijia, Liu, Shuaitong, Jin, Gaojie, Li, Xinyu, Gao, Xiao-Shan
Abstract
Prompt engineering has become an indispensable tool for using large language models (LLMs), turning LLMs into task-specific experts without changing their weights. Despite notable theoretical advances in prompt engineering, the theory for the more practical hard or discrete prompts is largely open. In this paper, we try to fill this gap either by providing a complete solution or by making substantial progress on the three core theoretical questions regarding hard prompts. First, we show that determining the existence of a hard prompt for a transformer to solve a downstream task is NP-complete and that finding an optimal hard prompt is NP-hard, which is the first computational complexity result for hard prompting, as far as we know. Second, we show that, unlike soft or continuous prompts, hard prompts have essential limitations: hard prompts are not complete; short hard prompts do not significantly enhance the ability of transformers; and long hard prompts exhibit the "prompt dominating answer phenomenon," meaning that, with high probability, the same answer is given for all queries of the same length. On the other hand, linear hard prompts do not have the limitations of short or long prompts. Third, we provide a tight bound on the size of the task in terms of the prompt length for the performance of prompts on the finite task to generalize to the entire data distribution, leading to a necessary and sufficient condition for generalizability. This is the first result on generalization for prompting, as far as we know. Our findings not only offer the first theoretical insights into hard prompts but also provide provably reliable practical guidance for real-world LLM usage.
Chinese Translation
提示工程已成为使用大型语言模型(LLMs)不可或缺的工具,它能在不改变模型权重的情况下将LLMs转变为任务特定的专家。尽管提示工程在理论上取得了显著进展,但对于更实用的硬提示或离散提示,其理论在很大程度上仍是开放的。在本文中,我们试图通过提供完整的解决方案或在关于硬提示的三个核心理论问题上取得实质性进展来填补这一空白。首先,我们证明,确定是否存在使Transformer解决下游任务的硬提示是NP完全的,并且找到最优硬提示是NP难的,据我们所知,这是关于硬提示的首个计算复杂性结果。其次,我们表明,与软提示或连续提示不同,硬提示具有本质局限:硬提示并不完备;短的硬提示不能显著增强Transformer的能力;长的硬提示会表现出“提示主导答案现象”,即对于所有相同长度的查询,以高概率给出相同的答案。另一方面,线性硬提示不具有短提示或长提示的局限性。第三,我们针对提示在有限任务上的性能泛化到整个数据分布的情况,给出了任务规模关于提示长度的紧界,从而得出了可泛化性的充要条件。据我们所知,这是关于提示泛化的首个结果。我们的发现不仅首次提供了关于硬提示的理论见解,还为现实世界中的LLM使用提供了可证明可靠的实践指导。
cs.LG / 95 / 2609.32318

What Can a Leaderboard Certify? Compositional Controllability for Fair Evaluation and Training of Biomedical Literature-Review Agents

排行榜能证明什么?面向生物医学文献综述智能体公平评估与训练的组合可控性
Han, Zhaowei, Zhang, Xiang, Guan, Lingxiao, Hu, Danqi, Liu, Kai, Chang, Kevin, Liu, Jie
Abstract
Leaderboards rank long-horizon agents by their final outputs. Yet a higher score alone does not establish whether two systems are comparable or which stage accounts for the difference. Unequal evidence, inputs, or budgets can affect scores, and statistical corrections do not remove this mismatch. We introduce compositional controllability to address these questions. A comparison window covers one stage, several stages, or the whole agent. Our central result bounds the gap between observed and controlled score differences using only nuisance outside the window. This yields an admissibility test applied before scores are inspected. Inadmissible comparisons are refused. For admissible pairs, an ordering is certified only when the score gap exceeds the combined sampling and nuisance radii; otherwise, it remains undecided. These decisions give each system a rank interval. We introduce BioLitBench, a benchmark of 2,042 biomedical articles represented as structured claim graphs. Among seven published pipelines, a conventional statistical analysis declares a winner in 14 of 21 pairwise comparisons. Yet the top-ranked system alone received the target review's bibliography. To isolate pipeline performance, our test requires matched inputs and a fixed backbone model. It refuses 11 of the 21 comparisons, including every comparison involving the top-ranked system. Seven of the 14 conventional conclusions fall within these refused pairs. The same comparison windows support stage-level training. We train SCRIBE on Qwen3.8-27B using rewards measured at each stage's exit. Under matched evidence, SCRIBE achieves a certified rank interval of [1,2], with certified advantages over all evaluated published pipelines and the evaluated Claude and OpenAI agents. Under same pool, SCRIBE matches the strongest published retriever and is certified above three published pipelines.
Chinese Translation
排行榜根据最终输出对长时程智能体进行排名。然而,仅凭更高的分数并不能确定两个系统是否可比,或者哪个阶段导致了差异。不平等的证据、输入或预算会影响分数,而统计校正并不能消除这种不匹配。我们引入组合可控性来解决这些问题。一个比较窗口涵盖一个阶段、多个阶段或整个智能体。我们的核心结果仅使用窗口外的干扰因素来界定观测到的和受控的分数差异之间的差距。这产生了一个在检查分数之前应用的可接受性测试。不可接受的比较会被拒绝。对于可接受的配对,只有当分数差距超过组合的采样和干扰半径时,才会认证排序;否则,它仍然未定。这些决策为每个系统给出一个排名区间。我们引入 BioLitBench,这是一个包含 2,042 篇生物医学文章的基准,以结构化的主张图表示。在七个已发表的流程中,传统的统计分析在 21 次两两比较中有 14 次宣布了获胜者。然而,仅排名最高的系统收到了目标综述的参考文献。为了分离流程性能,我们的测试要求匹配的输入和固定的骨干模型。它拒绝了 21 次比较中的 11 次,包括涉及排名最高系统的每一次比较。14 个传统结论中有 7 个落入了这些被拒绝的配对中。相同的比较窗口支持阶段级训练。我们在 Qwen3.8-27B 上训练 SCRIBE,使用在每个阶段出口处测量的奖励。在匹配证据下,SCRIBE 实现了 [1,2] 的认证排名区间,并相对于所有评估的已发表流程以及评估的 Claude 和 OpenAI 智能体具有认证优势。在相同池条件下,SCRIBE 与最强的已发表检索器相当,并被认证优于三个已发表的流程。
cs.LG / 96 / 2609.32319

Phase Space Attention:A Hairer Lift Circumvents the Single-Layer Induction Obstruction

相空间注意力:一种Hairer提升绕开单层归纳障碍
Maitra, Kingsuk, Hosseini, Shagun Sood Morteza, Gunnala, Suman, Gupta, Vikram
Abstract
We circumvent the Sanford-Hsu-Telgarsky (SHT) single-layer induction obstruction within a linear, one-step, causal, bilinear, symplectically consistent design class on the post-RoPE substrate, by lifting attention onto a symplectic phase space, mirroring Hairer's lift of Stormer-Verlet. The lift exits the premise of the SHT counting argument rather than the bound itself. The reframing exhibits the obstruction as a filter-order gap: a one-layer bilinear score realises a $z$-transform of joint order $(0,0)$, whereas the induction discriminator requires key-side order $\geq 1$. Applying the symplectic upper shear $M_\gamma:(q,p)\mapsto(q+\gamma p,p)$ to the post-RoPE query and key streams closes it. We prove this lift is unique within the factorised subclass, exactly symplectic at operator level, and requires post-RoPE placement; and in an explicit $T_4$-only Gaussian reduction we derive a closed-form two-branch induction phase transition, held out at $r=0.9876$ with zero fitted parameters. That law is an analytically solvable limit, not a robust prediction: restoring the $T_3$ channel moves $\gamma_c$ at $d_k=64$ from 1.030 to 0.569 and removes the crossover. Deployability follows by exact derivation: the KV cache is unchanged, prefix reuse and speculative decoding are preserved, overhead is $6d$ FLOPs per token per layer, INT8 headroom grows by at most $\log_2(1+2\gamma)$ bits, fused kernels are unmodified, and no parameters are added. At 91.3M parameters a supercritical sweep locates an emergence band: induction forms 3/3 seeds at $\gamma=0.80$ in a mean of 717 steps, against 2/3 seeds and 2700 steps at $\gamma=0$. Adverse results are reported as directly: a key-only half-lift reaches 0.949 against 0.811 for the symmetric operator, so if induction accuracy is the objective, the half-lift is the better construction. Forty-one notebooks and result files ship as ancillary material.
Chinese Translation
我们在 post-RoPE 基底上,在一个线性、单步、因果、双线性、辛一致的设计类中,通过将注意力提升到辛相空间,模仿 Hairer 对 Stormer-Verlet 的提升,绕开了 Sanford-Hsu-Telgarsky (SHT) 单层归纳障碍。该提升绕开了 SHT 计数论证的前提,而非界限本身。重新表述将障碍展现为滤波器阶数差距:单层双线性得分实现联合阶数 (0,0) 的 z 变换,而归纳判别器要求键侧阶数 ≥ 1。将辛上剪切 M_gamma:(q,p)↦(q+gamma p,p) 应用于 post-RoPE 查询和键流可弥合该差距。我们证明此提升在因子化子类中是唯一的,在算子层面精确辛,并且需要 post-RoPE 放置;在显式的仅 T4 高斯约简中,我们推导出一个闭式双分支归纳相变,留出验证在 r=0.9876 下,零拟合参数。该定律是一个可解析求解的极限,而非稳健预测:恢复 T3 通道使 d_k=64 处的 gamma_c 从 1.030 移动到 0.569,并消除交叉。可部署性由精确推导得出:KV 缓存不变,前缀重用和推测解码得以保留,开销为每 token 每层 6d FLOPs,INT8 余量最多增加 log2(1+2gamma) 比特,融合内核未修改,且未添加参数。在 91.3M 参数下,一次超临界扫描定位到一个涌现带:归纳在 gamma=0.80 时以平均 717 步形成 3/3 种子,而在 gamma=0 时为 2/3 种子和 2700 步。不利结果同样直接报告:仅键半提升达到 0.949,而对称算子为 0.811,因此如果归纳准确率是目标,半提升是更好的构造。四十一个 notebook 和结果文件作为辅助材料提供。
cs.LG / 97 / 2609.32320

Superposed Inference for Hyperdimensional Computing

面向超维计算的叠加推理
Zhao, Quanling, Pandey, Nilesh Prasad, Tian, Ye, Rosing, Tajana
Abstract
Hyperdimensional computing (HDC) is attractive for efficient and robust learning, but conventional inference still encodes every query independently, repeatedly paying the cost of high-dimensional projection. We introduce SupHDC, a new inference paradigm that processes multiple queries through a shared encoding computation. SupHDC assigns lightweight random slot keys, superposes the keyed queries before encoding, and uses slot-specific classifiers to recover their individual predictions. A random-feature kernel view explains why exact recovery of each hypervector is unnecessary: inference only needs to preserve the class evidence that determines the prediction. Across ten datasets, SupHDC achieves 1.39x analytical speedup with no average accuracy loss, and up to 2.08x speedup with only a 2.67 percentage-point mean accuracy loss. On a Raspberry Pi~5, it delivers 2.01x measured wall-clock speedup with a 2.26 percentage-point loss in mean prediction accuracy. SupHDC shows that high-dimensional redundancy can be used not only for robustness, but also as capacity for shared inference.
Chinese Translation
超维计算(HDC)在高效且鲁棒的学习方面颇具吸引力,但传统推理仍然独立编码每个查询,反复付出高维投影的代价。我们提出SupHDC,一种新的推理范式,通过共享编码计算处理多个查询。SupHDC分配轻量级随机槽键,在编码前叠加带键的查询,并使用槽特定的分类器恢复其各自的预测。随机特征核视角解释了为什么不需要精确恢复每个超向量:推理只需要保留决定预测的类别证据。在十个数据集上,SupHDC实现了1.39倍的分析加速且无平均精度损失,以及在仅2.67个百分点平均精度损失下高达2.08倍的加速。在Raspberry Pi 5上,它实现了2.01倍的实测墙钟加速,平均预测精度损失2.26个百分点。SupHDC表明,高维冗余不仅可用于鲁棒性,还可作为共享推理的容量。
cs.LG / 98 / 2609.32322

Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality

并非所有误差都重要:决策相关预测误差可预测规划质量
Wang, Linhao, Fan, Yiyan, Huang, Dongjin
Abstract
World models are typically trained and evaluated by prediction error, assuming that more accurate predictions lead to better decisions. We show that this assumption can fail because models with similar total error can differ substantially in planning performance when their errors occur on different state dimensions. We introduce Decision-Relevant Prediction Error (DRPE), which measures prediction error on the state dimensions that affect decisions. We also develop an iso-error evaluation protocol that varies error allocation while keeping total error fixed. In a factored gridworld with known state relevance and a standardized planner, we evaluate 55 controlled and learned models across different error levels and allocations. Total prediction error is weakly related to planning success (Spearman $\rho=-0.25$), whereas DRPE is strongly predictive ($\rho=-0.84$; $-0.98$ within the controlled family). Models with only a 1\% difference in total error can differ by 60 percentage points in planning success (97\% vs 37\%). The relevant error also depends on the task, with model rankings reversing across tasks at the same total error. Deeper imagination further amplifies decision-relevant errors, while learned models exhibit systematic bias on rare but decision-critical events. We formalize sufficient conditions under which DRPE correctly ranks models and total prediction error cannot.
Chinese Translation
世界模型通常通过预测误差进行训练和评估,假设更准确的预测会带来更好的决策。我们表明,这一假设可能不成立,因为总误差相似的模型,当其误差出现在不同状态维度上时,规划性能可能差异很大。我们提出决策相关预测误差(Decision-Relevant Prediction Error, DRPE),它衡量影响决策的状态维度上的预测误差。我们还开发了一种等误差评估协议,在保持总误差固定的同时改变误差分配。在具有已知状态相关性和标准化规划器的因子化网格世界中,我们评估了 55 个受控模型和学习模型,涵盖不同的误差水平和分配。总预测误差与规划成功弱相关(Spearman $\rho=-0.25$),而 DRPE 具有很强的预测性($\rho=-0.84$;在受控模型族内为 $-0.98$)。总误差仅相差 1% 的模型,其规划成功率可能相差 60 个百分点(97% vs 37%)。相关误差还取决于任务,在总误差相同时,模型排名在不同任务之间可能发生逆转。更深的想象会进一步放大决策相关误差,而学习模型在罕见但决策关键的事件上表现出系统性偏差。我们形式化了充分条件,在这些条件下,DRPE 能正确地对模型进行排名,而总预测误差则不能。
cs.LG / 99 / 2609.32325

Active Feature Acquisition With Incomplete Training Data

不完整训练数据下的主动特征获取
Rezvan, Reza, Schütz, Valter, Wu, Han, Aronsson, Linus, Chehreghani, Morteza Haghir
Abstract
In many prediction tasks, acquiring all features can be a prohibitively expensive or outright impossible task. Further, in many cases a static subset of features may not be enough to solve the problem sufficiently across various instances. Active Feature Acquisition (AFA) addresses these problems by formalizing the trade-off between feature cost and predictive performance during sequential feature selection. However, prior AFA work largely assumes access to complete training data, an assumption that is often violated in practice. Here we study AFA with Incomplete Training Data (AFA-ITD), showing that under missing completely at random (MCAR) data, one-step acquisition values remain unchanged, whereas multi-step values can decrease. We analyze three approaches to learning from incomplete data: aliasing, filtering, and generative restoration. We show that filtering can require a number of training instances scaling exponentially with the dimension, whereas generative restoration scales exponentially with the acquisition budget. We empirically test our theory on a controlled experiment and across common AFA datasets and find that missingness mainly damages methods that exploit multi-step acquisitions and that generative restoration is able to recover lost performance in many experiments. Code is available at https://github.com/Linusaronsson/AFA-Benchmark/tree/missing-data.
Chinese Translation
在许多预测任务中,获取所有特征可能成本过高或根本不可能。此外,在许多情况下,一个静态的特征子集可能不足以在各种实例上充分解决问题。主动特征获取(AFA)通过在序列特征选择过程中形式化特征成本与预测性能之间的权衡来解决这些问题。然而,先前的 AFA 工作主要假设可以访问完整的训练数据,这一假设在实践中经常被违背。在这里,我们研究不完整训练数据下的主动特征获取(AFA-ITD),表明在完全随机缺失(MCAR)数据下,单步获取价值保持不变,而多步价值可能会降低。我们分析了从不完整数据中学习的三种方法:别名(aliasing)、过滤(filtering)和生成恢复(generative restoration)。我们表明,过滤法可能需要训练实例的数量随维度指数增长,而生成恢复则随获取预算指数增长。我们在受控实验和常见的 AFA 数据集上对我们的理论进行了实证测试,发现缺失性主要损害利用多步获取的方法,并且生成恢复能够在许多实验中恢复损失的性能。代码可在 https://github.com/Linusaronsson/AFA-Benchmark/tree/missing-data 获取。
cs.LG / 100 / 2609.32328

Continual Data Unlearning in Diffusion Models via Transition-based Regularization

基于转移正则化的扩散模型持续数据遗忘
Jeong, Sunbeom, Kim, Sehwan, Hong, Sangwoo, Lee, Jungwoo
Abstract
Data unlearning in diffusion models aims to remove the influence of specific training examples without suppressing the broader concepts they represent. However, when deletion requests arrive sequentially, updates for new requests can degrade generative utility and undermine earlier deletions. We propose a continual data unlearning framework that uses completed deletion transitions as directional references to regularize future updates. For each request, we record changes in denoiser responses on the same fixed noisy inputs before and after unlearning. Rather than matching full post-deletion responses, we apply a one-sided penalty that discourages reversal along the recorded directions relative to the post-deletion references, while leaving orthogonal response changes and progress beyond these references unpenalized. To keep storage independent of the number of requests, we maintain a fixed-capacity bank of representative transition records. Records are selected based on the local sensitivity of progress along their recorded directions to parameter updates, allowing them to be retained even when their penalties are inactive. Empirical evaluations show that the proposed framework achieves a better balance between deletion persistence and generative utility than existing unlearning baselines as requests accumulate, using only a small transition memory.
Chinese Translation
扩散模型中的数据遗忘旨在移除特定训练样本的影响,同时不抑制它们所代表的更广泛概念。然而,当删除请求按顺序到达时,针对新请求的更新可能会降低生成效用,并破坏先前的删除效果。我们提出了一个持续数据遗忘框架,该框架利用已完成的删除转移作为方向性参考,以正则化未来的更新。对于每个请求,我们记录去噪器在相同固定噪声输入上、在遗忘前后响应的变化。我们不是匹配完整的删除后响应,而是应用一种单侧惩罚,该惩罚阻止相对于删除后参考沿记录方向的反转,同时不对正交响应变化以及超出这些参考的进展进行惩罚。为了使存储量与请求数量无关,我们维护一个固定容量的代表性转移记录库。记录的选择基于沿其记录方向的进展对参数更新的局部敏感度,使得即使在其惩罚未激活时也能保留这些记录。实证评估表明,随着请求的累积,所提出的框架在删除持久性和生成效用之间实现了比现有遗忘基线更好的平衡,并且仅使用少量的转移记忆。
cs.LG / 101 / 2609.32332

When the Merge Coefficient Stops Mattering: Proximity Regularized Merging for Continual LoRA Adaptation

当合并系数不再重要:用于持续LoRA适应的邻近正则化合并
Liu, Yixuan, Sun, Yuhao, Song, Sen, Li, Jin
Abstract
Rehearsal-free continual learning with parameter-efficient adapters can be cast as a sequence of task-vector write-in operations: for each new task, a low-rank adapter is learned and merged into a running model. We propose Proximity Regularized Merging (PRM), a minimal modification to sequential LoRA merging that adds a proximal penalty during task-vector training without changing the subsequent write-in rule. PRM acts as a robust task-vector regularizer: in the reported Base->+Prox diagnostics, it improves AAA across multiple write-in rules, backbones, and class-incremental settings, while its fixed-coefficient variant remains competitive with strong coefficient-based baselines. Mechanistically, matched-prefix norm controls and proximal-strength sweeps show that proximal training shrinks the task-vector radius, lowers Fisher-weighted interference, broadens the coefficient plateau, and exposes a stability-plasticity trade-off. Together, these results suggest that the effectiveness of sequential LoRA merging depends not only on how much of a task vector is written in, but also on whether the task vector itself has been trained to be mergeable.
Chinese Translation
基于参数高效适配器的无回放持续学习可以视为一系列任务向量写入操作:对于每个新任务,学习一个低秩适配器并将其合并到运行模型中。我们提出邻近正则化合并(Proximity Regularized Merging, PRM),这是对顺序LoRA合并的最小修改,在任务向量训练期间添加邻近惩罚,而不改变后续的写入规则。PRM充当鲁棒的任务向量正则化器:在报告的Base->+Prox诊断中,它提高了多种写入规则、骨干网络和类增量设置下的AAA,而其固定系数变体与强大的基于系数的基线相比仍具有竞争力。从机制上讲,匹配前缀范数控制和邻近强度扫描表明,邻近训练缩小了任务向量半径,降低了Fisher加权干扰,拓宽了系数平台,并揭示了稳定性-可塑性权衡。总之,这些结果表明,顺序LoRA合并的有效性不仅取决于写入了多少任务向量,还取决于任务向量本身是否被训练为可合并的。
cs.LG / 102 / 2609.32341

A Comparative Analysis of Attention versus State-Space Models for In-Context Learning

注意力机制与状态空间模型在上下文学习中的比较分析
Arda, Enes, Cayci, Semih, Eryilmaz, Atilla
Abstract
Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical framework for comparing the representational capabilities of broad classes of attention and SSMs. Starting from a generalized formulation of in-context linear regression and using cumulative Bayes regret as our measure, we abstract three capabilities required by many sequential learning problems in our belief geometry: evidence assembly, belief maintenance, and addressing. We then study three cases of our formulation that isolate these capabilities and yield sharp architectural lessons: For belief maintenance, SSMs attain the optimal regret over stationary aggregation kernels; for positional assembly, SSMs have a memory advantage; and for content addressing, softmax attention has an exponential width advantage over sigmoid-selective SSMs. Experiments with LLaMA-type Transformers and Mamba-2 show that these architectural insights extend beyond our analytically tractable classes and linear-regression testbed.
Chinese Translation
Transformer 和状态空间模型(SSMs)是两种突出的序列学习架构,然而它们的比较在很大程度上仍是经验性的,并且现有的理论分析通常是任务特定的或架构受限的。在本文中,我们提出信念几何(belief geometry),这是一个统一的分析框架,用于比较广泛的注意力机制和 SSMs 类别的表示能力。从上下文线性回归的广义公式出发,并使用累积贝叶斯遗憾作为度量,我们在信念几何中抽象出许多序列学习问题所需的三种能力:证据组装、信念维护和寻址。然后,我们研究了该公式的三种情况,分别隔离这些能力,并得出明确的架构启示:对于信念维护,SSMs 在平稳聚合核上达到最优遗憾;对于位置组装,SSMs 具有记忆优势;对于内容寻址,softmax 注意力相比 sigmoid 选择性 SSMs 具有指数级的宽度优势。使用 LLaMA 型 Transformer 和 Mamba-2 的实验表明,这些架构见解超出了我们可分析处理的类别和线性回归测试平台。
cs.LG / 103 / 2609.32350

Analog-Friendly Predictive Coding without Activation Derivatives

无需激活导数的模拟友好型预测编码
Innocenti, Francesco
Abstract
Predictive coding (PC) is a local, energy-based alternative to backpropagation (BP) whose iterative inference dynamics make it attractive for implementation on analog hardware. However, standard nonlinear PC requires evaluating the derivative of the activation function during both inference and learning, which can be difficult to realise physically. Here, we introduce \textit{activation-matched Bregman PC}, replacing standard squared-error energies with Bregman divergences matched to the activation function. This formulation eliminates activation derivatives and, when combined with inference via mirror descent, yields local inference and learning rules requiring only weighted sums, local prediction errors, state integration, and the activation function. In digital experiments, Bregman PC performs comparably to standard PC and BP on classification and generative tasks, while preserving characteristic learning dynamics of PC and its convergence to BP under stable large-model parameterisations. These results provide a more analog-friendly formulation of nonlinear PC while retaining its key computational properties.
Chinese Translation
预测编码(Predictive Coding, PC)是一种局部的、基于能量的反向传播(Backpropagation, BP)替代方案,其迭代推理动力学使其对于在模拟硬件上的实现具有吸引力。然而,标准的非线性 PC 在推理和学习过程中都需要评估激活函数的导数,这在物理上难以实现。在这里,我们引入了激活匹配 Bregman PC,将标准平方误差能量替换为与激活函数匹配的 Bregman 散度。这种形式消除了激活导数,并且当与通过镜像下降(mirror descent)进行的推理相结合时,产生了仅需要加权和、局部预测误差、状态积分和激活函数的局部推理和学习规则。在数字实验中,Bregman PC 在分类和生成任务上与标准 PC 和 BP 表现相当,同时保留了 PC 的特征性学习动力学及其在稳定的大模型参数化下向 BP 的收敛性。这些结果为非线性 PC 提供了一种更模拟友好的形式,同时保留了其关键的计算特性。
cs.LG / 104 / 2609.32360

Editable Map-Conditioned Trajectory Generation for Human Mobility Simulation

面向人类移动性模拟的可编辑地图条件轨迹生成
Mizuno, Takayuki, Fujimoto, Shouji, Hiruki, Mikito, Ishikawa, Atushi
Abstract
Geospatial simulation of infrastructure interventions requires mobility generators that respond directly to edited maps, yet many data-driven generators do not expose the map as an editable condition. We formulate this task as map-conditioned autoregressive generation of human mobility: a road raster conditions a decoder that emits nominal 31.25 m mesh-cell tokens at one-minute intervals. The mesh-local vocabulary supports held-out and locally edited maps without retraining or vocabulary changes. We instantiate a ResNet-50 visual-prefix configuration and a Vision Transformer (ViT) cross-attention configuration, trained from scratch on 87,400 smartphone-derived trajectories from 874 meshes in Ishikawa Prefecture, Japan; 219 meshes are held out. We evaluate map sensitivity by comparing correct-map and within-split shuffled-map generations with held-out real trajectories. On the 110-mesh test split, for the ResNet-50 configuration, correct-map generations are closer than shuffled-map generations on 60% of meshes under Hausdorff-based energy distance (p = 0.021), while DTW is directional but inconclusive (57%, p = 0.074); correlation with real density is 0.38 with the correct map versus 0.01 with shuffled maps. The ViT configuration shows weaker trajectory-level sensitivity and smaller density gains. An illustrative bridge-removal edit changes generated continuations without retraining. Together, these results support the feasibility of editable-map human-mobility simulation.
Chinese Translation
基础设施干预的地理空间模拟需要能够直接响应编辑后地图的移动性生成器,然而许多数据驱动的生成器并未将地图作为可编辑条件暴露出来。我们将此任务形式化为地图条件的人类移动性自回归生成:道路栅格作为条件输入到一个解码器,该解码器以一分钟为间隔发出标称31.25米网格单元标记。网格局部词汇表支持留出地图和局部编辑地图,而无需重新训练或更改词汇表。我们实例化了一个ResNet-50视觉前缀配置和一个Vision Transformer (ViT)交叉注意力配置,这些配置在日本石川县的874个网格的87,400条智能手机衍生轨迹上从头训练;其中219个网格被留出。我们通过将正确地图和分割内打乱地图的生成结果与留出的真实轨迹进行比较,来评估地图敏感性。在110个网格的测试集上,对于ResNet-50配置,在基于Hausdorff的能量距离下,正确地图的生成结果在60%的网格上比打乱地图的生成结果更接近(p = 0.021),而DTW具有方向性但不具决定性(57%,p = 0.074);与真实密度的相关性在正确地图下为0.38,在打乱地图下为0.01。ViT配置显示出更弱的轨迹级灵敏度和更小的密度增益。一个示例性的桥梁移除编辑无需重新训练即可改变生成的延续轨迹。这些结果共同支持了可编辑地图的人类移动性模拟的可行性。
cs.LG / 105 / 2609.32361

Black-Box Auditing of Epistemic Reliability in Multi-Agent Debate Distillation

多智能体辩论蒸馏中认知可靠性的黑盒审计
Wang, Derui, Shi, Zewei, Holland, Rayne, Sun, Ruoxi, Yuan, Xingliang, Xue, Jason, Zhu, Liming
Abstract
Debate distillation adapts weaker verifiers using multi-agent debate transcripts to improve their judgement in subsequent debates, but gains on monitored tasks do not establish reliability on related unmonitored tasks. We study epistemic reliability degradation, in which adaptation preserves monitored performance while reducing support for correct responses on hidden tasks. We consider an adversarial debater that manipulates debate arguments while defending the correct monitored response, and ask whether the resulting degradation merely reflects catastrophic forgetting and whether standard evaluation can detect it. To address these questions, we propose ER-Audit, a two-stage black-box auditing framework that compares frozen verifier checkpoints before and after adaptation, and introduce two evaluation benchmarks pairing monitored and hidden task prompts grounded in shared contexts. ER-Audit searches for counterexamples to non-degradation by evaluating semantically valid paraphrases and, if none is found, uses independent paraphrases for sequential hypothesis testing. We derive anytime-valid lower confidence bounds on the non-degradation probability, allowing data-dependent stopping within a finite budget. We further establish a common lower bound across fixed paraphrase distributions and extend it to distributions within a bounded total variation distance of their mixtures. Our experiments show that higher hidden-task accuracy can coexist with more counterexamples to non-degradation and lower non-degradation bounds. This divergence challenges explanations based solely on broad catastrophic forgetting and shows that auditing can uncover selective hidden-task degradation concealed by aggregate performance gains. Our code and benchmarks are available at https://github.com/CSIRO-CQS-AI-alignment-Team/Epistemic-Reliability-Auditor.
Chinese Translation
辩论蒸馏利用多智能体辩论记录来适配较弱的验证器,以改进其在后续辩论中的判断,但在受监控任务上的增益并不能确立在相关未监控任务上的可靠性。我们研究认知可靠性退化,即适配在保持受监控性能的同时,降低了对隐藏任务上正确响应的支持。我们考虑一个对抗性辩论者,它操纵辩论论据,同时捍卫正确的受监控响应,并询问由此产生的退化是否仅仅反映了灾难性遗忘,以及标准评估能否检测到它。为了解决这些问题,我们提出了 ER-Audit,一个两阶段黑盒审计框架,它比较适配前后的冻结验证器检查点,并引入了两个评估基准,将基于共享上下文的受监控和隐藏任务提示配对。ER-Audit 通过评估语义有效的复述来搜索非退化的反例,如果没有找到,则使用独立复述进行序贯假设检验。我们推导出非退化概率的任意时间有效的置信下界,允许在有限预算内进行数据依赖的停止。我们进一步在固定复述分布上建立了一个共同下界,并将其扩展到与其混合分布的有界全变差距离内的分布。我们的实验表明,更高的隐藏任务准确率可以与更多的非退化反例和更低的非退化下界共存。这种分歧挑战了仅基于广泛灾难性遗忘的解释,并表明审计可以揭示被总体性能增益所掩盖的选择性隐藏任务退化。我们的代码和基准可在 https://github.com/CSIRO-CQS-AI-alignment-Team/Epistemic-Reliability-Auditor 获取。
cs.LG / 106 / 2609.32363

DiffPTS: Rethinking Diffusion ELBO for Probabilistic Time Series Forecasting

DiffPTS:重新思考用于概率时间序列预测的扩散ELBO
Ye, Weiwei, Li, Dongyuan, Liu, Hangchen, Jiang, Haotong, Sekimoto, Yoshihide, Jiang, Renhe
Abstract
Probabilistic time series forecasting requires modeling and predicting complex and time-varying distributions. Recently, Denoising Diffusion Probabilistic Model (DDPM)-based approaches have shown promise by equipping the dif- fusion process with pretrained mean and variance estimators to accommodate distributional shift. However, these methods typically follow the standard DDPM framework and consider only partial components of the evidence lower bound (ELBO), treating the training of estimators as designed regression tasks separate from the variational inference framework. To address this, we rethink the ELBO under the Location-Scale Noise Model (LSNM) and find that it naturally induces a Gaussian negative log likelihood objective for the estimators and inherently defines a joint training objective that unifies recent diffusion paradigms for probabilistic forecasting. Building on this principled ELBO reformulation, we propose Diff- PTS, a general framework that enables end-to-end optimization of all components within the ELBO. Across multiple benchmarks, DiffPTS consistently outperforms recent models, achieving state-of-the-art performance with an average CRPS/MSE reduction of over 14.53%/16.55% compared to existing diffusion-based methods. The code is available at https://github.com/wwy155/DiffPTS.
Chinese Translation
概率时间序列预测需要对复杂且随时间变化的分布进行建模和预测。近年来,基于去噪扩散概率模型(DDPM)的方法通过为扩散过程配备预训练的均值和方差估计器以适应分布偏移,展现出良好前景。然而,这些方法通常遵循标准DDPM框架,仅考虑证据下界(ELBO)的部分组件,将估计器的训练视为独立于变分推断框架的、专门设计的回归任务。为了解决这个问题,我们重新思考了位置-尺度噪声模型(LSNM)下的ELBO,发现它自然地为估计器诱导出一个高斯负对数似然目标,并内在地定义了一个联合训练目标,统一了近期用于概率预测的扩散范式。基于这一有原则的ELBO重构,我们提出了DiffPTS,一个通用框架,能够对ELBO内的所有组件进行端到端优化。在多个基准测试中,DiffPTS持续优于近期模型,与现有基于扩散的方法相比,平均CRPS/MSE降低了超过14.53%/16.55%,达到了最先进的性能。代码可在https://github.com/wwy155/DiffPTS获取。
cs.LG / 107 / 2609.32365

Graph Memory: Spectral Associative Memory via Dirichlet Energy

图记忆:基于狄利克雷能量的谱联想记忆
Shi, Zhaoyang
Abstract
Dense associative memories have traditionally focused on storing and retrieving vector-valued patterns. Many modern machine learning problems, however, are naturally graph-structured, requiring memory mechanisms for relational patterns, graph diffusion geometries, community structures, and graph-based inductive biases. We propose a spectral dense associative memory for storage and retrieval of graph data, extending the classical vector-valued memories. Retrieval is performed through a log-sum-exp energy induced by Dirichlet energy with spectral norm distances, producing a softmax-weighted average of the stored Laplacians that remains a valid graph Laplacian. We prove exponential storage capacity and exponentially decaying retrieval error. Beyond graph retrieval, we establish theoretical guarantees for spectral quantities central to graph learning, including eigenvalues, eigenspaces, and diffusion operators. Experiments on synthetic graph data, real-world airline network, protein conformation data and wearable sensor data demonstrate robust graph retrieval while preserving the graph geometry of the data. Our framework provides a new associative memory paradigm for graph-structured data and bridges dense associative memory with modern graph learning and generative AI.
Chinese Translation
密集联想记忆传统上专注于存储和检索向量值模式。然而,许多现代机器学习问题本质上是图结构的,需要用于关系模式、图扩散几何、社区结构以及基于图的归纳偏置的记忆机制。我们提出了一种用于图数据存储和检索的谱密集联想记忆,扩展了经典的向量值记忆。检索通过由带有谱范数距离的狄利克雷能量诱导的log-sum-exp能量进行,产生存储的拉普拉斯矩阵的softmax加权平均,该平均仍然是有效的图拉普拉斯矩阵。我们证明了指数存储容量和指数衰减的检索误差。除了图检索,我们还为图学习核心的谱量建立了理论保证,包括特征值、特征空间和扩散算子。在合成图数据、真实世界航空网络、蛋白质构象数据和可穿戴传感器数据上的实验展示了鲁棒的图检索,同时保持了数据的图几何结构。我们的框架为图结构数据提供了一种新的联想记忆范式,并将密集联想记忆与现代图学习和生成式AI桥接起来。
cs.LG / 108 / 2609.32367

STRIDE: State-Transition Representation via Increment Dynamics and Evolution

STRIDE:通过增量动力学与演化的状态转移表示
Xiong, Yuchen, Huang, Siming, Sun, Jianfeng
Abstract
We introduce STRIDE (State-Transition Representation via Increment Dynamics and Evolution), which defines states through derivative fingerprints and learns local functions for state transitions (qpairs), recasting continuous forecasting as transition prediction. Trailing convolution windows estimate local joint-transition frequencies, whose lagged differences form a high-dimensional increment trajectory. Proper orthogonal decomposition (POD) gives coordinate paths, jointly forecast by sparse dynamics with memory. Recombining their forecasts and applying history-anchored inversion recovers future transition distributions; sampled state paths select local functions to generate continuous forecasts. Three-seed experiments compare STRIDE against fourteen baselines across nine benchmark families. Five independent Markov and hidden-state baselines cover all 230 evaluated tasks, with 220 complete whole-horizon pairs. Against these comparators, system-weighted late-Energy win fractions range from 73.6% to 78.5% on 190 multi-step pairs; against DLinear, the fraction is 77.3% on 72 paired multi-step tasks. Matched controls examine the intermediate representation. On 64 independently initialized Aizawa trajectories, late-Energy reductions against four matched controls range from approximately 24% to 56%, with all four prespecified contrasts passing Holm correction. These results connect transition-statistic prediction to continuous probabilistic forecasting, with substantial long-horizon gains in the matched Aizawa study.
Chinese Translation
我们介绍了 STRIDE(通过增量动力学与演化的状态转移表示),它通过导数指纹定义状态,并学习状态转移(qpairs)的局部函数,将连续预测重新表述为转移预测。尾部卷积窗口估计局部联合转移频率,其滞后差分形成高维增量轨迹。本征正交分解(POD)给出坐标路径,通过带记忆的稀疏动力学联合预测。重新组合它们的预测并应用历史锚定反演,恢复未来的转移分布;采样的状态路径选择局部函数以生成连续预测。三种子实验在九个基准系列上将 STRIDE 与十四个基线进行比较。五个独立的马尔可夫和隐状态基线覆盖所有 230 个评估任务,其中 220 个完整的全时域对。与这些比较对象相比,系统加权的晚期 Energy 获胜比例在 190 个多步对上为 73.6% 至 78.5%;与 DLinear 相比,在 72 个配对的多步任务上该比例为 77.3%。匹配的对照检查了中间表示。在 64 个独立初始化的 Aizawa 轨迹上,相对于四个匹配对照的晚期 Energy 降低范围约为 24% 至 56%,所有四个预先指定的对比均通过 Holm 校正。这些结果将转移统计预测与连续概率预测联系起来,在匹配的 Aizawa 研究中获得了显著的长时域增益。
cs.LG / 109 / 2609.32379

Measurement Boundaries in LLM Financial Agent Evaluation: Fixed-Tape Execution and Multi-Defect Auditing

LLM金融智能体评估中的测量边界:固定磁带执行与多缺陷审计
Xue, Weicheng
Abstract
What controls are needed to interpret execution performance and audit scores in LLM agent evaluations? We study two limits on these interpretations in a financial agent harness. In Study~A, comparing independent runs under idealized and stressed execution on three synthetic settings that share one 24-day upward phase mixes the execution rule with fresh model responses and portfolio feedback: the parsed decision paths agree in only $19.8\%$ of $450$ pairs. Replaying each stored response tape through both execution destinations gives a narrower result. Conditional on those responses, stressed execution changes total return by $-0.0170$ (95\% interval $[-0.0230,-0.0117]$), or $10.4\%$ of the idealized baseline, and ten seed clusters do not resolve the model ranking. Study~B corrects an incomplete answer key and replaces legacy tasks with matched zero-, one-, and two-defect tasks under an explicit multi-label prompt. The drop in target violation recall from one to two defects is positive in five of six combinations of auditor and source (median $0.267$), with three surviving Holm correction. Yet the auditor that includes both target labels most often has micro-precision $0.149$, emits findings on $98/100$ zero-defect tasks, and returns the exact dual-defect set in only $21/100$ cases. Target recall by itself therefore gives a poor account of audit quality on this construction. The studies address different limits: what an execution comparison estimates, and what target recall captures. Together, they show how fixed conditions and diagnostic controls bound the claims a score can support.
Chinese Translation
需要什么控制来解释LLM智能体评估中的执行性能和审计得分?我们在一个金融智能体测试框架中研究了这些解释的两个限制。在Study A中,比较理想化和压力执行下的独立运行,在三个共享一个24天上升阶段的合成设置上,将执行规则与新的模型响应和投资组合反馈混合在一起:解析出的决策路径在450对中只有19.8%一致。通过两个执行目的地重放每个存储的响应磁带,得到了更狭窄的结果。在这些响应条件下,压力执行将总回报改变了-0.0170(95%区间[-0.0230,-0.0117]),即理想化基线的10.4%,且十个种子聚类无法解决模型排名。Study B纠正了一个不完整的答案键,并在明确的多标签提示下,用匹配的零缺陷、一缺陷和两缺陷任务替换了遗留任务。目标违规召回率从一缺陷到两缺陷的下降在六种审计器和来源组合中的五种中为正(中位数0.267),其中三种经过Holm校正后仍然显著。然而,包含两个目标标签的审计器最常具有微精度0.149,在98/100的零缺陷任务上发出发现,并且仅在21/100的情况下返回精确的双缺陷集。因此,仅凭目标召回率无法很好地说明此构造下的审计质量。这些研究解决了不同的限制:执行比较估计的是什么,以及目标召回率捕捉的是什么。总之,它们展示了固定条件和诊断控制如何限制分数所能支持的声明。
cs.LG / 110 / 2609.32380

Attribution Without a Second Pass: Inline Per-Sample Gradient Provenance at ~1% Overhead

无需二次遍历的归因:以约1%开销实现内联逐样本梯度溯源
Nautiyal, Amit
Abstract
Data attribution methods used in practice (TRAK, LoGRA, EK-FAC) are post-hoc: after training they make a second pass over the training set to recompute per-sample gradients, repeated per checkpoint when ensembled. Traceprop avoids that pass by recording projected per-sample gradients inline on the training backward pass. A Kronecker-factored sketch scales from a single tracked layer to every layer without materializing a dense projection matrix. On LoRA fine-tunes of GPT-2 and Pythia models up to 2.8B on one NVIDIA L4, inline logging costs 0.30-1.08% of wall-clock time at last-block scope and stays under 1% (0.79%) even when tracking every layer of Pythia-1B. Against LogIX, the closest inline-capable competitor, the factored sketch is 2.0-4.1x cheaper at equal storage, a gap that grows with tracked scope and is significant at every scope tested, while matching or exceeding LogIX's attribution quality at matched storage. Building the attribution-ready store inline is 60-242x cheaper than one post-hoc pass and 301-1211x cheaper than a five-checkpoint TRAK ensemble, with recorded gradients matching autograd exactly. Because each stored gradient carries source-file lineage, the same pass also produces EU AI Act Article 26 audit trails.
Chinese Translation
实践中使用的数据归因方法(TRAK、LoGRA、EK-FAC)是事后进行的:训练后,它们对训练集进行二次遍历以重新计算逐样本梯度,在集成时每个检查点重复此过程。Traceprop通过在训练反向传播中内联记录投影后的逐样本梯度,避免了这一遍历。一种Kronecker因子分解草图可以从单个被跟踪层扩展到每一层,而无需物化稠密投影矩阵。在单个NVIDIA L4上对高达28亿参数的GPT-2和Pythia模型进行LoRA微调时,在最后块范围,内联记录仅耗费0.30-1.08%的实际时间;即使在跟踪Pythia-1B的每一层时,也保持在1%(0.79%)以下。与最接近的具备内联能力的竞争对手LogIX相比,在相同存储下,因子分解草图便宜2.0-4.1倍,这一差距随着跟踪范围的扩大而增大,并且在每个测试范围下都显著,同时在匹配存储下达到或超过LogIX的归因质量。内联构建归因就绪存储比一次事后遍历便宜60-242倍,比五检查点TRAK集成便宜301-1211倍,且记录的梯度与自动微分完全匹配。由于每个存储的梯度携带源文件谱系,同一次遍历还生成EU AI Act第26条审计追踪。
cs.LG / 111 / 2609.32384

TimeES: Probabilistic and Deterministic Time Series Forecasting via Evolutionary Spectra

TimeES: 基于演化谱的概率与确定性时间序列预测
Ye, Weiwei, Jiang, Renhe, Liu, Hangchen, Li, Dongyuan, Sekimoto, Yoshihide
Abstract
Real-world time series are inherently non-stationary, with trends, periodic patterns, and uncertainty evolving over time. While the Fourier domain offers a natural lens to model time series, current deep learning approaches do not explicitly model evolution and randomness in the Fourier spectra, which limits their ability to accurately predict both the expected trajectory and its uncertainty in non-stationary time series. Motivated by Evolutionary Spectra (ES) theory, we propose TimeES, a general framework that enables probabilistic and deterministic forecasting via the evolutionary spectra theory. Specifically, we derive a parameterizable evolutionary spectra formulation, recasting non-stationary random process modeling as learning an evolving representation modulated by random variables. Furthermore, we reduce the complexity of the estimated spectra from O(NM) to O(NK), where K << M/2, by exploiting Hermitian symmetry and spectral energy sparsity for frequency selection. Based on a simple linear backbone, our proposed TimeES achieves consistent state-of-the-art performance across both deterministic and probabilistic forecasting tasks, with high efficiency and interpretability. Code is available at: https://github.com/wwy155/TimeES.
Chinese Translation
现实世界的时间序列本质上是非平稳的,其趋势、周期模式以及不确定性随时间演变。虽然傅里叶域为时间序列建模提供了自然的视角,但当前的深度学习方法并未显式建模傅里叶谱中的演化和随机性,这限制了它们准确预测非平稳时间序列中预期轨迹及其不确定性的能力。受演化谱(ES)理论的启发,我们提出了 TimeES,一个通过演化谱理论实现概率性和确定性预测的通用框架。具体而言,我们推导了一个可参数化的演化谱公式,将非平稳随机过程建模重新表述为学习一个由随机变量调制的演化表示。此外,我们通过利用厄米特对称性和频谱能量稀疏性进行频率选择,将估计谱的复杂度从 O(NM) 降低到 O(NK),其中 K << M/2。基于简单的线性主干,我们提出的 TimeES 在确定性和概率性预测任务中均取得了持续的最先进性能,同时具有高效率和可解释性。代码可在以下网址获取:https://github.com/wwy155/TimeES。
cs.LG / 112 / 2609.32386

Using Machine Learning to Investigate Predictors of Fasting Blood Glucose: Insights into Circadian Timing and Age Interactions

利用机器学习探究空腹血糖的预测因子:对昼夜节律时相与年龄交互作用的启示
Bu-Dager, Viktoriya, Cirstea, Silvia
Abstract
Impaired glucose regulation is a major contributor to metabolic dysfunction and type 2 diabetes. This study developed an interpretable machine-learning framework to predict log-transformed fasting blood glucose using metabolic, hormonal, lifestyle, demographic, nutritional, and circadian variables from the National Health and Nutrition Examination Survey 2017--2020 pre-pandemic dataset. After merging multiple NHANES sub-datasets, data processing used a leakage-resistant pipeline in which imputation, scaling, and one-hot encoding were performed only after dataset splitting and within training folds. Elastic Net, LASSO, and XGBoost models were evaluated using 94 candidate predictors and engineered circadian interaction terms. Performance was assessed using mean absolute error, root mean squared error, coefficient of determination, calibration, and Shapley Additive Explanations. The final interaction-augmented XGBoost model achieved strong performance on the independent test set, with a mean absolute error of 0.0804, a root mean squared error of 0.1148, and a coefficient of determination of 0.7761, using 10 predictors. Glycohemoglobin was the dominant predictor, followed by insulin, diabetes diagnosis, gamma-glutamyl transferase, age, race, and gender. Among the engineered interaction terms, sleep midpoint multiplied by age was consistently retained in repeated random-split analyses, although its contribution remained modest relative to dominant glycaemic predictors. These findings support further investigation of circadian-age interactions in metabolic health.
Chinese Translation
糖调节受损是代谢功能障碍和2型糖尿病的主要促发因素。本研究开发了一个可解释的机器学习框架,利用美国国家健康与营养调查(NHANES)2017–2020年大流行前数据集中的代谢、激素、生活方式、人口学、营养和昼夜节律变量,预测对数转换后的空腹血糖。在合并多个NHANES子数据集后,数据处理采用防泄漏流程,其中插补、缩放和独热编码仅在数据集划分之后且在训练折内进行。使用94个候选预测因子和构建的昼夜节律交互项评估了Elastic Net、LASSO和XGBoost模型。采用平均绝对误差、均方根误差、决定系数、校准和Shapley加性解释(SHAP)评估性能。最终加入交互项的XGBoost模型在独立测试集上表现良好,使用10个预测因子时,平均绝对误差为0.0804,均方根误差为0.1148,决定系数为0.7761。糖化血红蛋白是最主要的预测因子,其次是胰岛素、糖尿病诊断、γ-谷氨酰转移酶、年龄、种族和性别。在构建的交互项中,睡眠中点乘以年龄在重复随机划分分析中始终被保留,尽管其贡献相对于主要血糖预测因子仍然较小。这些发现支持进一步研究昼夜节律-年龄交互作用在代谢健康中的作用。
cs.LG / 113 / 2609.32409

Recovery-Directed Symbolic Distillation of Neural Likelihoods

恢复导向的神经似然符号蒸馏
Fernandez, Kianté, Li, Xinwei
Abstract
Amortized neural likelihoods enable computationally expensive inference for models with analytically intractable or unspecified likelihoods, but their black-box nature limits interpretability. We introduce a symbolic distillation pipeline that converts trained neural likelihoods into explicit, interpretable expressions optimized for efficient parameter estimation. Our approach uses a recovery-directed objective to guide symbolic regression toward expressions that preserve parameter-recovery accuracy rather than merely approximating the likelihood function. Candidate expressions are evaluated on held-out datasets and selected using a criterion that jointly accounts for expression complexity, parameter-recovery performance, and distributional distance from the learned likelihood. We evaluate the pipeline on the diffusion decision model, a classical cognitive model, whose analytically tractable likelihood provides ground truth for controlled evaluation. The proposed recovery-directed objective improves parameter recovery over standard symbolic-regression objectives. The resulting symbolic likelihoods enable over 100 times faster parameter evaluation than both neural likelihoods and, when available, the exact likelihood, while maintaining a manageable loss in precision. We further demonstrate these computational benefits in Bayesian hierarchical inference on empirical data. Our pipeline provides a lightweight interface for integrating symbolic distillation with existing neural-likelihood estimation methods and can be adapted to a range of simulation-based inference settings.
Chinese Translation
摊销神经似然使得对具有解析难处理或未指定似然的模型进行原本计算昂贵的推理成为可能,但其黑箱性质限制了可解释性。我们引入了一个符号蒸馏流程,将训练好的神经似然转换为显式、可解释的表达式,并针对高效参数估计进行优化。我们的方法使用恢复导向的目标来引导符号回归,使其倾向于保留参数恢复准确性的表达式,而不仅仅是近似似然函数。候选表达式在留出数据集上进行评估,并使用一个同时考虑表达式复杂度、参数恢复性能以及与所学似然的分布距离的准则进行选择。我们在扩散决策模型(diffusion decision model)上评估该流程,这是一个经典的认知模型,其解析上可处理的似然为受控评估提供了真实基准。所提出的恢复导向的目标在参数恢复方面优于标准符号回归目标。得到的符号似然使得参数评估比神经似然以及(当可用时)精确似然快100倍以上,同时保持可控的精度损失。我们进一步在经验数据的贝叶斯分层推断中展示了这些计算优势。我们的流程提供了一个轻量级接口,用于将符号蒸馏与现有的神经似然估计方法集成,并可适应一系列基于模拟的推断设置。
cs.LG / 114 / 2609.32426

HoTS: Homophily-Aware Temperature Scaling for Graph Neural Network Calibration

HoTS:用于图神经网络校准的同质性感知温度缩放
Tae, Inwoo, Hwang, Yoontae, Lee, Yongjae
Abstract
For graph node classification, calibrated class probabilities are needed when confidence scores, usually the maximum predicted class probability, are used to rank predictions, defer uncertain nodes to human review, or control risk. Existing post-hoc calibrators either apply one global temperature or use graph-aware modules without a principled structural form. We study how local graph structure should enter node-level calibration. Our first results show that a logit-only temperature rule is insufficient when nodes with identical logits but different local homophily require different optimal temperatures. We then analyze a population-concentration contextual stochastic block model with Gaussian features and a one-layer linear GCN. Under equidistant class means, the Bayes posterior over class-template scores is a temperature-scaled softmax whose inverse-temperature is governed by a homophily-dependent signal strength. In the positive-signal homophilic regime, the resulting temperature decreases approximately inversely with normalized local homophily. This law motivates Homophily-aware Temperature Scaling (HoTS), a simple post-hoc calibrator that assigns each node a positive scalar temperature from entropy-based logit concentration and estimated local homophily. HoTS has three temperature parameters, preserves the predicted class, and learns the strength of the structural correction from calibration data. Across 18 node-classification benchmarks, two GNN backbones, and eight calibration baselines, HoTS achieves the best mean Expected Calibration Error (ECE) of 4.79%, the best average rank, and the most reliable confidence ranking in selective classification. Code is available at https://github.com/inu0104/HoTS.
Chinese Translation
对于图节点分类,当置信度分数(通常是最大预测类概率)被用于对预测进行排序、将不确定节点推迟给人工审核或控制风险时,需要校准的类概率。现有的后处理校准器要么应用单一全局温度,要么使用图感知模块但没有有原则的结构形式。我们研究局部图结构应如何进入节点级校准。我们的初步结果表明,当具有相同logit但不同局部同质性的节点需要不同的最优温度时,仅基于logit的温度规则是不够的。然后我们分析了一个具有高斯特征和一层线性GCN的群体浓度上下文随机块模型。在等距类均值下,类模板分数上的贝叶斯后验是一个温度缩放的softmax,其逆温度由同质性依赖的信号强度决定。在正信号同质性区域中,所得温度近似与归一化局部同质性成反比。这一规律促使了同质性感知温度缩放(HoTS),一种简单的后处理校准器,它根据基于熵的logit浓度和估计的局部同质性为每个节点分配一个正标量温度。HoTS有三个温度参数,保持预测类别,并从校准数据中学习结构校正的强度。在18个节点分类基准、两个GNN骨干和八个校准基线上,HoTS实现了最佳的平均期望校准误差(ECE)为4.79%,最佳的平均排名,以及在选择性分类中最可靠的置信度排名。代码可在 https://github.com/inu0104/HoTS 获取。
cs.LG / 115 / 2609.32432

Beyond the Manifold Hypothesis: Hybrid Spectral Parameterizations for Flow Matching

超越流形假设:用于流匹配的混合谱参数化
Martin, Ségolène, Gagneux, Anne, Bertrand, Quentin, Emonet, Rémi, Massias, Mathurin
Abstract
Flow matching and diffusion can be trained to predict different quantities, most commonly the data $x_1$, the source noise $x_0$, or the velocity~$v$. Although theoretically equivalent, these can lead to substantially different performances. We identify two main drivers for these differences: the source--data signal-to-noise ratio, and the information bottleneck induced by the neural architecture. We show that, beyond intrinsic data dimension, the factor affecting the optimal parametrization the most is a certain signal-to-noise ratio in each data covariance direction. From this analysis, we introduce new \emph{spectral hybrid} parameterizations that adapt across time and data covariance directions; we show that these are optimal for Gaussian data. We also show that architecture-induced compression changes which parameterization is easier to learn, with $v$-prediction being more sensitive to discarded directions than $x_1$-prediction. Experiments across architectures and source scales show that our spectral parameterizations are robust across regimes, can substantially accelerate optimization, while incurring essentially no additional training cost compared with standard parameterizations.
Chinese Translation
流匹配和扩散模型可以被训练来预测不同的量,最常见的是数据 $x_1$、源噪声 $x_0$ 或速度 $v$。尽管在理论上等价,但这些预测目标可能导致显著不同的性能。我们确定了导致这些差异的两个主要驱动因素:源-数据信噪比,以及由神经架构引起的信息瓶颈。我们表明,除了内在数据维度之外,影响最优参数化最大的因素是在每个数据协方差方向上的特定信噪比。基于此分析,我们引入了新的谱混合参数化方法,能够跨时间和数据协方差方向自适应;我们证明这些方法对高斯数据是最优的。我们还表明,架构引起的压缩改变了哪种参数化更易于学习,其中 $v$-预测对丢弃方向比 $x_1$-预测更敏感。在不同架构和源规模的实验表明,我们的谱参数化在不同机制下具有鲁棒性,可以显著加速优化,同时与标准参数化相比几乎不增加额外的训练成本。
cs.LG / 116 / 2609.32444

Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It

重新思考LLM强化学习中的训练-推理不匹配:它从何而来以及如何纠正它
Yu, Tianrun, Zhao, Kaixiang, Li, Shangzhe, Yang, Yuxiao, Jenkins, Porter, Zhang, Weitong, Killian, Taylor W.
Abstract
We study training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different probabilities to the same tokens. To account for this discrepancy in policy updates, we introduce calibrated importance sampling (CIS). CIS is motivated by an empirically supported logit-displacement characterization that expresses the mismatch as an additive displacement $\varepsilon_t$ in log-odds, determined by the per-logit perturbation before the softmax, whose distribution is approximately invariant to token confidence. This characterization motivates a confidence-aware truncation: large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases. Theoretically, we show that CIS replaces the unbounded second moment that governs the error of exact importance sampling with a term bounded by a constant, at the cost of a bias controlled by the truncated excess. In evaluation across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines. Diagnostic analyses show that CIS places less truncation bias on low-confidence tokens than truncated importance sampling, while upward clipping of small importance weights reduces held-out accuracy.
Chinese Translation
我们研究了在大语言模型的可验证奖励强化学习(RLVR)中的训练-推理不匹配问题,其中轨迹由推理引擎采样,而梯度由训练引擎计算,两个引擎对相同token赋予不同的概率。为了在策略更新中解释这种差异,我们引入了校准重要性采样(CIS)。CIS的动机来自于一个经实验支持的logit位移刻画,该刻画将不匹配表示为对数几率中的加性位移$\varepsilon_t$,由softmax之前的每个logit扰动决定,其分布对token置信度近似不变。这一刻画启发了一种置信度感知的截断:大的正位移在单个常数阈值处被截断,这映射回一个重要性比率上限,该上限随着token置信度的增加而收紧。在理论上,我们表明CIS将支配精确重要性采样误差的无界二阶矩替换为一个由常数界定的项,代价是偏差由截断超出部分控制。在三个混合专家模型和五个数学推理基准的评估中,CIS在所有三个模型上均达到了所评估基线中最高的五基准平均分。诊断分析表明,与截断重要性采样相比,CIS对低置信度token施加的截断偏差更小,而对小重要性权重的向上裁剪会降低留出准确率。
cs.LG / 117 / 2609.32447

Length-Independent State Tracking Under a Parallel Scan

并行扫描下的长度无关状态跟踪
Brandoit, Julien, Fyon, Arthur, Braipson, Thomas, Clara, Tom, De Geeter, Florent, Sacré, Pierre, Ernst, Damien, Drion, Guillaume
Abstract
Learning robust and scalable finite-state tracking is fundamental to sequence processing. While linear recurrent neural networks (RNNs), linear attention, and state space models enable scalable parallel training through affine recurrences, their theoretical expressivity guarantees assume idealized arithmetic and do not extend to finite precision, where the parallel scan that makes them fast is itself a source of perturbation. We formalize finite-state tracking at finite precision and characterize length independence: tracking that stays correct at every sequence length, at a precision cost that does not grow with the length. We show that length-independent state tracking requires two competing dynamics within a single map: contraction to suppress numerical perturbations and separation to keep distinct states apart. We prove that affine recurrences, which offer a single rate at each step to serve both roles, realize at most definite automata at finite precision. Instead of treating scan compatibility as a restriction on the update map, we reinterpret it as a computational budget and introduce the Neural Finite-State Machine (NFSM): a nonaffine, scan-compatible RNN built for length-independent finite-state tracking. On synthetic benchmarks spanning abelian and nonabelian groups, noninvertible monoids, and textual state-tracking tasks, affine baselines fail on every nondefinite task, most of them within a few hundred steps. A single NFSM layer instead learns the exact transition tables of every algebraic task, which certifies correctness beyond the tested lengths, and a stack of NFSMs keeps perfect accuracy on the textual tasks at every tested length.
Chinese Translation
学习鲁棒且可扩展的有限状态跟踪是序列处理的基础。虽然线性循环神经网络(RNNs)、线性注意力和状态空间模型通过仿射递归实现了可扩展的并行训练,但它们的理论表达能力保证假设了理想算术,并不能扩展到有限精度,而在有限精度下,使其快速的并行扫描本身就是一个扰动源。我们在有限精度下形式化了有限状态跟踪,并刻画了长度无关性:即在每个序列长度上保持正确,且精度成本不随长度增长。我们表明,长度无关的状态跟踪需要在单个映射中具有两种相互竞争的动态:收缩以抑制数值扰动,分离以保持不同状态分离。我们证明,仿射递归(在每一步提供单一速率来同时服务于这两种角色)在有限精度下最多只能实现确定型自动机(definite automata)。我们没有将扫描兼容性视为对更新映射的限制,而是将其重新解释为计算预算,并引入了神经有限状态机(NFSM):一种非仿射、扫描兼容的RNN,专为长度无关的有限状态跟踪而构建。在涵盖阿贝尔群和非阿贝尔群、不可逆幺半群以及文本状态跟踪任务的合成基准上,仿射基线在每个非确定型任务(nondefinite tasks)上都失败了,其中大多数在几百步内就失败了。相反,单个NFSM层就能学习每个代数任务的精确转移表,这证明了在测试长度之外的正确性,而堆叠的NFSM在所有测试长度上对文本任务保持完美的准确率。
cs.LG / 118 / 2609.32457

Write Back the $\Delta$: Revisiting the Same Tokens with Fresh Representations

回写 $\Delta$:以新表征重新访问相同词元
Ye, Wencheng, Hu, Anning, Zhang, Xiangdong, Wang, Tianyi, Li, Yikang, Jin, Hengyu, Li, Bing, Yan, Junchi
Abstract
Transformers process information strictly forward through depth, preventing deeper computation from revisiting and refining earlier representations. To augment the standard forward pass, existing approaches either re-execute depth, incurring additional computation, or modify the residual stream using predefined directions, limiting their instance-level adaptation. Recently, inference-time feedback offers a direct mechanism for recycling endogenously produced computation by writing deeper residual states back to earlier layers, yet what should be fed back remains unclear. We argue that the depth increment Delta, capturing newly accumulated computation between two layers, provides a more effective, composable, and scalable feedback signal than the full state. Building on this observation, we introduce ReFlux, a learnable feedback graph that dynamically selects and composes increment-carrying routes. ReFlux supports synchronous feedback to the same token and streaming feedback to subsequent tokens. Extensive experiments across various models, corpora, and benchmarks show that synchronous ReFlux consistently reduces perplexity across ten language-modeling corpora, and improves accuracy by 2.1-2.3 points, with gains reaching 4.7 points on multi-hop reasoning. Streaming ReFlux further retains most of these gains while preserving the base model's 1x theoretical backbone FLOPs. These results establish ReFlux as an efficient paradigm for unlocking the latent computational potential of LLMs, allowing them to revisit the same tokens with fresh representations. Code implementation can be found at https://github.com/gooogleshanghai/reflux.
Chinese Translation
Transformer 严格沿深度方向前向处理信息,阻止了深层计算重新访问并精炼早期表征。为了增强标准前向传播,现有方法要么重新执行深度,带来额外计算,要么使用预定义方向修改残差流,限制了其实例级适应能力。最近,推理时反馈提供了一种直接机制,通过将更深层的残差状态写回较早层来回收内生计算,但应该反馈什么仍不清楚。我们认为,深度增量 Delta(捕获两层之间新累积的计算)提供了比完整状态更有效、可组合且可扩展的反馈信号。基于这一观察,我们提出了 ReFlux,一个可学习的反馈图,它动态选择并组合携带增量的路径。ReFlux 支持对相同词元的同步反馈以及对后续词元的流式反馈。在各种模型、语料库和基准上的大量实验表明,同步 ReFlux 在十个语言建模语料库上持续降低困惑度,并将准确率提高 2.1-2.3 个百分点,在多跳推理上增益达到 4.7 个百分点。流式 ReFlux 进一步保留了大部分这些增益,同时保持基础模型 1 倍理论主干 FLOPs。这些结果确立了 ReFlux 作为一种高效范式,用于释放 LLM 的潜在计算潜力,允许它们以新表征重新访问相同词元。代码实现见 https://github.com/gooogleshanghai/reflux。
cs.LG / 119 / 2609.32464

Adapting Nonstationary Multi-output Gaussian Processes to Bayesian Optimization

将非平稳多输出高斯过程适配于贝叶斯优化
Xie, Zikai
Abstract
Multi-objective Bayesian optimization (MOBO) commonly relies on independent Gaussian processes (GPs) with stationary kernels, limiting its ability to represent nonstationary structure and share information between objectives. However, expressive nonstationary GPs do not necessarily make reliable BO decisions. We study this mismatch for the multi-output low-rank nonstationary (MO-LRN) GP: strong training fit can coexist with large off-design errors and optimistic acquisition predictions. We introduce MOLRN-BO, which combines a regularized shared-spectral surrogate with objective-specific residuals, prequential mean correction and tempered covariance scaling, and Pareto-local qLogEHVI optimization with periodic global search. Experiments on 12 deterministic bi-objective benchmarks show that MOLRN-BO substantially improves upon the original MO-LRN and achieves the best average problem ranks for final normalized hypervolume and normalized inverted generational distance among nine evaluated algorithms. It also achieves the strongest adverse-tail performance while remaining competitive with the leading baselines in anytime optimization. Ablation studies further show that the shared spectral construction improves off-design prediction, the local--global decision policy improves optimization performance, and hierarchical calibration reduces systematic candidate bias. These results demonstrate that nonstationary multi-output surrogates can deliver strong and robust MOBO performance when their structure and use are explicitly adapted to the demands of sequential optimization.
Chinese Translation
多目标贝叶斯优化(MOBO)通常依赖于具有平稳核的独立高斯过程(GP),限制了其表示非平稳结构以及在不同目标之间共享信息的能力。然而,具有表达能力的非平稳GP并不一定能做出可靠的BO决策。我们针对多输出低秩非平稳(MO-LRN)GP研究了这种不匹配:良好的训练拟合可能与大的非设计误差和乐观的采集预测并存。我们提出了MOLRN-BO,它将正则化的共享谱代理与目标特定的残差、序贯均值校正和调温协方差缩放,以及带有周期性全局搜索的帕累托局部qLogEHVI优化相结合。在12个确定性双目标基准上的实验表明,MOLRN-BO在最终归一化超体积和归一化反转世代距离方面,相对于原始MO-LRN有显著改进,并在九个评估算法中取得了最佳的平均问题排名。它还在不利尾部性能上表现最强,同时在随时优化中与领先的基线保持竞争力。消融研究进一步表明,共享谱构造改善了非设计点预测,局部-全局决策策略提高了优化性能,分层校准减少了系统性候选偏差。这些结果表明,当非平稳多输出代理的结构和使用方式被明确地适应于序列优化的需求时,它们能够提供强大且稳健的MOBO性能。
cs.LG / 120 / 2609.32465

PolyStepOR: Learning to Decide Without Optimal Decisions

PolyStepOR:无需最优决策的学习决策
Nguyen, Viet The, Gust, Gunther, Le, An Thai
Abstract
Decision-focused learning (DFL) trains predictors for downstream decision quality, but often relies on optimal reference decisions that are expensive to obtain. We present PolyStepOR, which trains directly from realized decision costs without pre-computed optima and extends to in-constraint predictions through repair or infeasibility penalties. To handle piecewise-constant losses, PolyStepOR perturbs predictor parameters, evaluates the resulting decisions, and uses optimal transport to favor lower-cost directions, requiring no derivatives. Without task-specific tuning, PolyStepOR performs strongly on classical optimization benchmarks and competitively on predicted-constraint and real-world problems. Theoretically, we characterize decision-preserving perturbations and boundary detection, bound sensitivity to cost errors, and establish stationarity guarantees for a smoothed objective. PolyStepOR thus replaces optimal reference decisions and derivatives with forward evaluations.
Chinese Translation
决策聚焦学习(DFL)训练预测器以优化下游决策质量,但通常依赖于获取成本高昂的最优参考决策。我们提出了 PolyStepOR,它直接从已实现的决策成本进行训练,无需预先计算最优解,并通过修复或不可行惩罚扩展到满足约束的预测。为了处理分段常数损失,PolyStepOR 扰动预测器参数,评估产生的决策,并使用最优传输来偏好低成本方向,且不需要导数。无需特定任务的调优,PolyStepOR 在经典优化基准上表现强劲,在预测约束和现实世界问题上具有竞争力。在理论上,我们刻画了决策保持扰动和边界检测,界定了对成本误差的敏感性,并为平滑目标建立了平稳性保证。因此,PolyStepOR 用前向评估取代了最优参考决策和导数。
cs.LG / 121 / 2609.32467

Bison: Cross-Dataset Learning for Unseen-Compound Perturbation Prediction

Bison:跨数据集学习用于未见化合物扰动预测
Liu, Yunfan, Ghorbani, Kasra, Huang, Yufei, Liu, Zicheng, Zheng, Jiangbin, Zhou, Jingbo, Chen, Shaorong, Yu, Chang, Li, Stan Z.
Abstract
Predicting transcriptional responses to unseen compounds is limited by fragmented chemical coverage and heterogeneous experimental platforms and gene panels. To assess molecular generalization across these settings, we build on Chem-PerturBridge to benchmark eight datasets with 16,771 compounds, withholding test compounds from every training dataset. This comparison reveals that high overall response agreement can coexist with weak prediction of drug-specific differences, despite reproducible signals across repeated measurements. To exploit complementary chemical supervision while targeting these differences, we introduce Bison: a shared gene representation connects native panels, while two discrete diffusion models compose context-dependent responses with molecular deviations learned through matched drug-contrast supervision. A single Bison model jointly trained across all eight datasets achieves the highest mean overall-response and drug-contrast Pearson correlations on the full benchmark in comparison with 11 methods trained independently per dataset. Compared with dataset-specific training of the same architecture, joint training increases mean drug-contrast correlation by 27.4\%, with gains across all eight datasets and improvements in overall response prediction. These results demonstrate how matched drug contrasts turn complementary screens into shared molecular supervision for unseen-drug response prediction while preserving native gene measurements.
Chinese Translation
预测对未见化合物的转录响应受到碎片化的化学覆盖以及异质的实验平台和基因面板的限制。为了评估这些设置下的分子泛化,我们基于Chem-PerturBridge构建了一个基准,包含八个数据集和16,771个化合物,并从每个训练数据集中留出测试化合物。这一比较表明,尽管重复测量中存在可重复的信号,但高的总体响应一致性可能与弱的药物特异性差异预测并存。为了利用互补的化学监督,同时针对这些差异,我们引入了Bison:一个共享的基因表示连接原生面板,而两个离散扩散模型通过匹配的药物对比监督学习分子偏差,从而组合出上下文依赖的响应。一个在所有八个数据集上联合训练的单一Bison模型,在完整基准上实现了最高的平均总体响应和药物对比皮尔逊相关性,与11种按数据集独立训练的方法相比。与相同架构的数据集特定训练相比,联合训练将平均药物对比相关性提高了27.4%,在所有八个数据集上均有增益,并且总体响应预测也有所改善。这些结果展示了匹配的药物对比如何将互补筛选转化为用于未见药物响应预测的共享分子监督,同时保留原生基因测量。
cs.LG / 122 / 2609.32470

On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models

论用于校准大型推理模型的言语化置信度先验的陷阱
Wang, Shuoyuan, Luo, Beier, Zeng, Hao, Yu, Chengyao, Zhang, Songxin, Xie, Zejian, Jing, Bingyi, Wei, Hongxin
Abstract
Large reasoning models (LRMs) often suffer from overconfidence when expressing their uncertainty. Confidence-aware reinforcement learning (RL) offers a promising way to optimize calibration. However, it relies on on-policy rollouts and is thus constrained by the model's pre-RL confidence distribution, which we term confidence prior. In this work, we reveal that off-the-shelf LRMs exhibit a confidence prior heavily concentrated on a few high values, which persists throughout RL. Theoretically, we prove that this concentration suppresses policy gradient updates for rarely sampled confidence values and inflates the lower bound on expected Brier risk. To overcome this exploration bottleneck, we propose CalibSFT, a plug-and-play supervised fine-tuning stage that shapes a calibrated confidence prior with broad support before RL. For each question, CalibSFT constructs confidence targets combining its success rate with response-level correctness, which provably preserves proper-scoring optimality, and then balances training responses across the confidence spectrum to enable diverse confidence exploration during RL. To learn from incorrect responses without imitating their reasoning, CalibSFT introduces correctness-conditional supervision, guiding confidence across all responses while supervising reasoning only on correct ones. Across 16 mathematical and general reasoning benchmarks, incorporating CalibSFT reduces calibration errors and improves discrimination across five representative RL algorithms while preserving comparable accuracy. Furthermore, CalibSFT delivers practical benefits for downstream selective prediction and model routing. Our code is available at https://github.com/ml-stat-Sustech/verbalized-confidence-training.
Chinese Translation
大型推理模型(LRMs)在表达其不确定性时常常过度自信。置信度感知的强化学习(RL)为优化校准提供了一种有前景的方法。然而,它依赖于同策略展开,因此受到模型强化学习前置信度分布的约束,我们称之为置信度先验。在这项工作中,我们揭示出现成的 LRMs 表现出一种高度集中在少数高值上的置信度先验,这种先验在整个 RL 过程中持续存在。理论上,我们证明这种集中抑制了极少采样的置信度值的策略梯度更新,并抬高了期望 Brier 风险的下界。为了克服这一探索瓶颈,我们提出了 CalibSFT,一种即插即用的监督微调阶段,在 RL 之前塑造具有广泛支撑的校准置信度先验。对于每个问题,CalibSFT 构建结合其成功率和响应级正确性的置信度目标,这被证明保持了恰当评分最优性,然后在置信度谱上平衡训练响应,以便在 RL 期间实现多样化的置信度探索。为了从不正确的响应中学习而不模仿其推理,CalibSFT 引入了正确性条件监督,在所有响应上引导置信度,同时仅对正确响应监督推理。在 16 个数学和通用推理基准上,结合 CalibSFT 在五种代表性 RL 算法中减少了校准误差并提高了区分度,同时保持了相当的准确率。此外,CalibSFT 为下游选择性预测和模型路由带来了实际益处。我们的代码见 https://github.com/ml-stat-Sustech/verbalized-confidence-training。
cs.LG / 123 / 2609.32479

Fast and Precise Learned Charged-Particle Trajectory Regression at the Large Hadron Collider

大型强子对撞机上快速且精确的学习型带电粒子轨迹回归
Renusch, Jonathan, Huth, Benjamin, Murnane, Daniel, Xochelli, Eleni, Elitez, Doğa, Gessinger-Befurt, Paul, Stefl, Andreas, Couthures, Jeremy, Salzburger, Andreas, Heinrich, Lukas, Kagan, Michael, Elsing, Markus
Abstract
We propose a training recipe that treats charged-particle trajectory parameter regression on high-energy physics detector data as a sequence-modeling task. Kalman filters and linearized least-squares fits have been the classical standard approach for this task: they are optimal estimators for sparsely sampled linear-Gaussian data and are commonly used for trajectory parameter regression (fitting). The classical fitting techniques implemented for this domain reach a final precision of one part in $10^5$ through detailed modeling of detector geometry, material, detection effects and precise numerical integration of the equations of motion through the detector's inhomogeneous magnetic field. With this study, we demonstrate that using a bidirectional gated linear recurrent encoder, one is able to reproduce the full precision of classical track fitting techniques. Using a custom kernel, we also achieve significantly higher throughput during GPU inference, compared to classical fitting software running on similarly priced multi-core CPU servers representing typically employed hardware. Such a speedup would lead to considerable cost savings for the pattern recognition at the Large Hadron Collider. To our knowledge, this is the first end-to-end learned track fit to reach the full precision and, at the same time, offer the opportunity to reduce the computing costs.
Chinese Translation
我们提出了一种训练方案,将高能物理探测器数据上的带电粒子轨迹参数回归视为序列建模任务。卡尔曼滤波器和线性化最小二乘拟合一直是该任务的经典标准方法:它们是对稀疏采样的线性高斯数据的最优估计器,并常用于轨迹参数回归(拟合)。针对该领域实现的经典拟合技术,通过对探测器几何、材料、探测效应进行详细建模,并对运动方程在探测器非均匀磁场中进行精确数值积分,可达到 $10^5$ 分之一的最终精度。通过本研究,我们证明使用双向门控线性循环编码器(bidirectional gated linear recurrent encoder)能够复现经典径迹拟合技术的完整精度。使用自定义内核,与在价格相近的多核 CPU 服务器(代表通常使用的硬件)上运行的经典拟合软件相比,我们还在 GPU 推理期间实现了显著更高的吞吐量。这种加速将为大型强子对撞机的模式识别带来可观的成本节约。据我们所知,这是首个端到端学习型径迹拟合,既达到完整精度,同时又有机会降低计算成本。
cs.LG / 124 / 2609.32485

What Should Federated LoRA Share? FedSAIL via Input-aware Subspace Alignment

联邦LoRA应该共享什么?通过输入感知子空间对齐的FedSAIL
Du, Junye, He, Shuaida, Feng, Long
Abstract
Federated low-rank adaptation (LoRA) requires identifying an update structure that is shared across heterogeneous clients. Prior work reports strong similarity among trained LoRA projection matrices across clients; however, such agreement may be largely induced by common initialization and collapses toward random overlap under independent initialization. More crucially, relying solely on parameter similarity inherently ignores the influence of local input regime. To uncover a more robust shared structure, we introduce an input-aware action matrix that weights the adapter update by the second-moment statistics of local layer inputs. Empirically, while parameter similarity vanishes, the leading right singular directions of this action matrix remain strongly aligned across clients. This shared geometry preserves task-conditioned differences and naturally varies across network depths. Motivated by these findings, we propose Federated Subspace-Guided Action-Informed Learning (FedSAIL). Instead of averaging weights, FedSAIL estimates a shared action subspace to regularize local training while preserving client-specific coefficients. Across several benchmarks, our approach consistently improves predictive performance over competing federated LoRA methods while reducing communication cost significantly.
Chinese Translation
联邦低秩适应(LoRA)需要识别一种在异构客户端之间共享的更新结构。先前的工作报道了不同客户端训练得到的LoRA投影矩阵之间存在很强的相似性;然而,这种一致性可能很大程度上是由共同初始化引起的,并且在独立初始化下会退化为随机重叠。更关键的是,仅依赖参数相似性本质上忽略了局部输入分布的影响。为了揭示更稳健的共享结构,我们引入了一种输入感知的动作矩阵,它通过局部层输入的二阶矩统计量对适配器更新进行加权。经验上,虽然参数相似性消失,但该动作矩阵的主要右奇异方向在不同客户端之间仍然高度对齐。这种共享的几何结构保留了任务相关的差异,并自然地随网络深度变化。受这些发现的启发,我们提出了联邦子空间引导的动作感知学习(FedSAIL)。FedSAIL不是平均权重,而是估计一个共享的动作子空间来正则化本地训练,同时保留客户端特定的系数。在多个基准测试上,我们的方法在显著降低通信成本的同时,始终优于竞争的联邦LoRA方法的预测性能。
cs.LG / 125 / 2609.32486

Elastic Selective Spectral Hybrids for Train-Once, Export-Many Budgeted Inference

面向一次训练、多次导出预算推理的弹性选择性谱混合模型
Song, Dachuan, Chen, Chuchu, Wang, Xuan
Abstract
Deploying a language model under different computing and latency budgets calls for compact models with different quality-cost trade-offs. Towards this end, elastic spectral state space models provide ordered and temporally decomposed channels that can be truncated, but they are associated with linear time-invariant filters that cannot selectively preserve relevant past information or forget irrelevant information as the context evolves. To address this, we introduce the Elastic Selective Spectral Hybrid (ESSH), which realizes each Hankel spectral channel as an independent recurrent unit using fitted damped rotation modes. It also features an input-dependent decay and write/read gates that make temporal retention and state update input-dependent while preserving channel-wise truncation and a structured recurrence for efficient execution. ESSH combines these selective spectral mixers with sliding-window attention and jointly trains multiple capacities by reducing spectral-channel count and feed-forward width at different rates through a two-rate capacity map with full-model distillation. The resulting models support chunked parallel training and fused recurrent decoding while avoiding computation for discarded channels. At full capacity, ESSH achieves language-modeling quality comparable to similarly sized independently trained models, while smaller exports exhibit a smooth quality-cost trade-off. We validate the effectiveness of the proposed framework using language understanding, retrieval, cross-domain text, and DNA experiments, by assessing quality retention and the trade-off against independently trained and elastic baselines. At the 1.53B model configuration, fused batch-one decoding takes 1.37 ms per token on a B300, providing a 2.14-2.80x speedup over the tested Mamba-2 and Mamba-3 implementations and 3.03x over Transformer++ at matched parameter counts.
Chinese Translation
在不同计算和延迟预算下部署语言模型,需要具有不同质量-成本权衡的紧凑模型。为此,弹性谱状态空间模型提供了有序且按时间分解的可截断通道,但它们与线性时不变滤波器相关,这些滤波器无法随着上下文演变选择性地保留相关过去信息或遗忘无关信息。为了解决这个问题,我们引入了弹性选择性谱混合模型(ESSH),它使用拟合的阻尼旋转模式将每个Hankel谱通道实现为独立的循环单元。它还具有输入依赖的衰减和写/读门,使时间保留和状态更新依赖于输入,同时保留通道级截断和结构化递归,以实现高效执行。ESSH将这些选择性谱混合器与滑动窗口注意力相结合,并通过具有全模型蒸馏的双速率容量映射,以不同速率减少谱通道数和前馈宽度,联合训练多个容量。所得模型支持分块并行训练和融合递归解码,同时避免对丢弃通道的计算。在完整容量下,ESSH实现了与类似规模独立训练模型相当的语言建模质量,而较小的导出模型表现出平滑的质量-成本权衡。我们通过语言理解、检索、跨领域文本和DNA实验,评估质量保持以及与独立训练和弹性基线的权衡,验证了所提出框架的有效性。在1.53B模型配置下,融合批大小为1的解码在B300上每token耗时1.37毫秒,相比测试的Mamba-2和Mamba-3实现提供2.14-2.80倍的加速,在匹配参数数量下比Transformer++快3.03倍。
cs.LG / 126 / 2609.32487

Neural Dynamics as the Composition of Quantized Units

神经动力学作为量子化单元的复合
Minniti, Jacopo, Kulanthaivelu, Aravinth, Sproat, Richard
Abstract
Deep learning is commonly interpreted at two levels: the macroscopic, through aggregate trends in loss summarized by scaling laws, and the microscopic, through neurons, features, and circuits. A central challenge is understanding how these levels connect, so that we can explain how elementary computations compose and collectively shape macroscopic behavior. To this end, we study an intermediate abstraction in which training is described as the ordered acquisition of quanta: reusable computations acquired suddenly and binary-activated across examples to reduce loss. By approximating population-gradient updates, we derive quanta's acquisition dynamics. This yields an acquisition priority governed by demand, how frequently a computation is required across examples, and conditional complexity, how difficult that computation is to acquire given those already available. In a Boolean compositional task, we derive predictions for acquisition order and show how staggered discrete acquisitions can produce smooth aggregate loss and, under certain geometries of quanta composition, give rise to scaling laws. We then train a Transformer to map numerals to English number names and recover candidate quanta from its checkpoint trajectory. From these units, we construct a model that preserves much of the Transformer's behavior while exposing interpretable latent computations and acquisition dynamics consistent with the theory. Separately, the quanta structure can serve as training targets to improve transformer generalization. Together, these results suggest the quanta abstraction can provide useful computational atoms for studying a variety of macroscopic phenomena.
Chinese Translation
深度学习通常在两个层面上被解释:宏观层面,通过由缩放定律总结的损失的总体趋势;微观层面,通过神经元、特征和电路。一个核心挑战是理解这些层面如何连接,以便我们能够解释基本计算如何组合并共同塑造宏观行为。为此,我们研究了一种中间抽象,其中训练被描述为量子单元的有序获取:可重用的计算,这些计算突然获得,并在示例中二值激活以减少损失。通过近似群体梯度更新,我们推导出量子单元的获取动力学。这产生了一种获取优先级,由需求和条件复杂度决定:需求是指一个计算在示例中被需要的频率,条件复杂度是指在已有计算可用的情况下获取该计算的难度。在一个布尔组合任务中,我们推导出获取顺序的预测,并展示了交错离散获取如何产生平滑的总体损失,并且在某些量子单元组合的几何结构下,产生缩放定律。然后,我们训练一个Transformer将数字映射到英文数字名称,并从其检查点轨迹中恢复候选量子单元。从这些单元中,我们构建了一个模型,该模型保留了Transformer的大部分行为,同时暴露了可解释的潜在计算和与理论一致的获取动力学。另外,量子单元结构可以作为训练目标来提高Transformer的泛化能力。总之,这些结果表明,量子单元抽象可以为研究各种宏观现象提供有用的计算原子。
cs.LG / 127 / 2609.32493

SoFT: Soft Targets for Generalizable LLM Fine-Tuning

SoFT:面向可泛化大语言模型微调的软目标
Jing, Huihao, Hu, Wenbin, Chen, Shaojin, Shi, Haochen, Xie, Zhongwei, Zhang, Guijia, Liu, Yuxuan, Huang, Haoyu, Li, Haoran, Song, Yangqiu
Abstract
Distillation enables student language models to acquire new capabilities from expert teachers. However, integrating knowledge from multi-teacher, multi-domain demonstrations into a single student remains challenging. We study supervised fine-tuning (SFT) in this setting, where students must acquire diverse capabilities while maintaining generalization beyond the training tasks. Our experiments reveal varying trade-offs between in-distribution learning and out-of-distribution generalization across SFT methods, motivating more explicit control over this balance. To this end, we propose soft-target fine-tuning (SoFT) to balance learning from teacher demonstrations with retaining the Base model's existing capabilities. SoFT sets a minimum target probability for each demonstrated token while making the smallest KL change to the Base distribution. The resulting objective couples learning from demonstrations with adaptively weighted regularization toward the Base model. We further use domain-specific gradient budgets to control this balance and determine a probability threshold for each trajectory. Experiments on mixed-domain reasoning and agentic tasks show that SoFT achieves the best overall performance among the compared methods, with improvements in both in-distribution capability acquisition and out-of-distribution generalization.
Chinese Translation
蒸馏使学生语言模型能够从专家教师处获取新能力。然而,将来自多教师、多领域演示的知识整合到单个学生模型中仍然具有挑战性。我们在此背景下研究监督微调(SFT),其中学生模型必须获得多样化的能力,同时在训练任务之外保持泛化性。我们的实验揭示了不同SFT方法在分布内学习和分布外泛化之间存在不同的权衡,这促使我们更明确地控制这种平衡。为此,我们提出了软目标微调(SoFT),以平衡从教师演示中学习与保留基础模型现有能力之间的关系。SoFT为每个演示的token设定最小目标概率,同时对基础分布做出最小的KL变化。由此产生的目标函数将从演示中学习与对基础模型的自适应加权正则化相结合。我们进一步使用领域特定的梯度预算来控制这种平衡,并为每个轨迹确定概率阈值。在混合领域推理和智能体任务上的实验表明,SoFT在比较方法中取得了最佳整体性能,在分布内能力获取和分布外泛化方面均有提升。
cs.LG / 128 / 2609.32500

Rondo: Unsupervised Discovery of Recurring Temporal Structure

Rondo:无监督发现重复出现的时序结构
Shi, Yingtian, Chandra, Ankith, Plötz, Thomas
Abstract
Many real-world time series data exhibit structural properties at multiple scales, from short, recurring units to complex sequences composed of these units. Unsupervised discovery of both these components and structure enables the design of intelligent systems that help interpret temporal data, thereby limiting the amount of costly human annotations required. Existing modeling approaches typically overlook the hierarchical structure inherent to many time series, treating recurring patterns at different temporal scales as independent structures. Moreover, most assume access to the complete data sequence and treat discovery as a static process, limiting their ability to evolve as new observations arrive. We introduce Rondo, an unsupervised approach for modeling recurring hierarchical structure in continuous temporal streams. By explicitly constructing vocabularies of reusable units and their recurring compositions, Rondo captures structure shared across complex temporal patterns while refining and expanding its discoveries as the stream evolves. Evaluations on temporal sequences spanning diverse domains and data modalities show that Rondo outperforms existing unsupervised recurrence-discovery baselines, with particularly pronounced advantages in limited-data and continual-stream settings. These capabilities provide a stronger foundation for recurring-pattern discovery, scalable behavior understanding, and adaptive intelligent systems operating on long, unlabeled temporal streams.
Chinese Translation
许多真实世界的时间序列数据在多个尺度上表现出结构特性,从短的、重复出现的单元,到由这些单元组成的复杂序列。对这些单元及其结构的无监督发现,有助于设计能够帮助解读时间数据的智能系统,从而减少所需的大量昂贵人工标注。现有建模方法通常忽略许多时间序列固有的层次结构,将不同时间尺度上的重复模式视为独立结构。此外,大多数方法假设可以访问完整的数据序列,并将发现视为静态过程,限制了其在新观测到来时演进的能力。我们提出 Rondo,一种用于对连续时间流中的重复层次结构进行建模的无监督方法。通过显式构建可复用单元及其重复组合的词汇表,Rondo 在随着流演进不断细化和扩展其发现的同时,捕获复杂时间模式间共享的结构。在跨越不同领域和数据模态的时间序列上进行的评估表明,Rondo 优于现有的无监督重复发现基线,在有限数据和持续流设置中优势尤为明显。这些能力为重复模式发现、可扩展的行为理解以及在长时、无标签时间流上运行的自适应智能系统提供了更坚实的基础。
cs.LG / 129 / 2609.32503

Fisher Simplicity in Kolmogorov-Arnold Networks and Multilayer Perceptrons

Kolmogorov-Arnold网络与多层感知机中的Fisher简单性
Tavory, Ami, Feder, Meir
Abstract
Kolmogorov-Arnold Networks (KANs) are motivated in part by interpretability: their learned edge functions can be inspected, pruned, and reduced to symbolic structure. In a fixed-basis KAN, this makes a small or zero basis coefficient look like a certificate of simplicity, much as a dead rectified linear unit (ReLU) marks unused computation in a multilayer perceptron (MLP). Fisher nullity gives a precise statistical notion: a parameter direction is Fisher-simple exactly when perturbing it is invisible under the task distribution. We study when these architectural and Fisher notions agree. For a dead ReLU unit, they agree: the closed activation region makes the associated score directions vanish. For a fixed-basis KAN, they do not. In the single-layer Gaussian case, the coefficient Fisher matrix is a basis Gram matrix under the input distribution and is independent of the fitted coefficients. In a multilayer KAN, Fisher simplicity is graph-path based: the data must reach a basis atom and its perturbation must propagate through the downstream network. We encode these two conditions in an effective edge measure and, under local dictionary independence and effective-measure nondegeneracy, show that zero effective exposure exactly identifies Fisher-null directions within an edge. Controlled diagnostics confirm that zero coefficients can preserve rank while effective path disconnections remove the predicted directions. Coefficient magnitude alone is therefore not a Fisher-based pruning criterion for KANs.
Chinese Translation
Kolmogorov-Arnold网络(KANs)的部分动机在于可解释性:其学习到的边函数可以被检查、剪枝并简化为符号结构。在固定基KAN中,这使得小的或零的基系数看起来像是简单性的证明,就像死亡的修正线性单元(ReLU)标志着多层感知机(MLP)中未使用的计算一样。Fisher零度给出了一个精确的统计概念:当在任务分布下扰动某个参数方向不可见时,该方向恰好是Fisher简单的。我们研究了这些架构概念和Fisher概念何时一致。对于死亡的ReLU单元,它们是一致的:闭合的激活区域使得相关的得分方向消失。对于固定基KAN,它们不一致。在单层高斯情况下,系数Fisher矩阵是输入分布下的基Gram矩阵,并且与拟合系数无关。在多层的KAN中,Fisher简单性是基于图路径的:数据必须到达一个基原子,并且其扰动必须通过下游网络传播。我们将这两个条件编码在一个有效边测度中,并且在局部字典独立性和有效测度非退化的条件下,表明零有效暴露恰好识别边内的Fisher零方向。受控诊断证实,零系数可以保持秩,而有效路径断开则移除预测的方向。因此,仅系数大小并不是KANs的基于Fisher的剪枝标准。
cs.LG / 130 / 2609.32512

What Do Latent Predictive Vehicle Representations Retain? Measuring State, Geometry, and Local Response

潜在预测性车辆表征保留了哪些信息?测量状态、几何与局部响应
Spotorno, Enzo Nicolás, Filho, Josafat Leal, Fröhlich, Antônio Augusto
Abstract
Models of vehicle dynamics learned from logged states and commands complement physics-based models, and latent world models, which predict in a learned representation, are used to plan and train controllers in other domains. Vehicle controllers are usually specified in physical terms: costs, limits, and references depend on position, yaw angle, speed, and yaw rate, and the optimizer compares or differentiates predicted outcomes across nearby commands. A latent model placed in such a controller must therefore let these quantities be recovered and must change its predictions with commands as the vehicle does, and prediction error on its own latent targets measures neither. We contribute a measurement protocol for action-conditioned latent predictors with a physical readout that separately tests retention, physical-neighborhood organization, forecasting, and local response to command perturbations, using an untrained-encoder reference and three matched response paths that locate errors in the representation or the predictor. In a case study of a temporal joint-embedding predictive model trained on signals logged in IPG CarMaker, the representations retain the measured planar outputs, though an untrained encoder of the same architecture retains them slightly better; future-command input improves one-second forecasts with retention nearly unchanged; and responses to small command pulses diverge from the simulator already in latent coordinates, raising regret when choosing among nearby commands in all comparisons. Updating the predictor on responses corrects them locally at a cost in forecast accuracy. Measuring retention, forecasting, and local response separately is thus what qualifies a predictive latent as a candidate model for control, and the protocol provides the basis for its closed-loop evaluation.
Chinese Translation
从记录的状态和命令中学习得到的车辆动力学模型补充了基于物理的模型,而潜在世界模型(在学习的表征中进行预测)被用于在其他领域规划和训练控制器。车辆控制器通常以物理量来指定:成本、限制和参考值依赖于位置、偏航角、速度和偏航角速率,并且优化器比较或微分邻近命令下的预测结果。因此,置于此类控制器中的潜在模型必须能够恢复这些物理量,并且必须像车辆一样随命令改变其预测,而仅在其自身潜在目标上的预测误差无法衡量这两者。我们提出了一种针对动作条件潜在预测器的测量协议,该协议具有物理读出,分别测试保留性、物理邻域组织、预测以及对命令扰动的局部响应,使用未训练编码器作为参考以及三条匹配的响应路径来定位表示或预测器中的误差。在一个对在 IPG CarMaker 中记录的信号上训练的时间联合嵌入预测模型的案例研究中,这些表示保留了所测量的平面输出,尽管相同架构的未训练编码器保留得略好;未来命令输入改善了一秒的预测,而保留性几乎不变;并且对小幅命令脉冲的响应在潜在坐标中就已与模拟器偏离,在所有比较中,当在邻近命令之间进行选择时增加了遗憾值。根据响应更新预测器可以在局部纠正它们,但会以预测精度为代价。因此,分别测量保留性、预测和局部响应是判定一个预测性潜在变量是否可以作为控制候选模型的标准,并且该协议为其闭环评估提供了基础。
cs.LG / 131 / 2609.32525

Does Transolver really need a Transformer?

Transolver 真的需要 Transformer 吗?
Wen, Shizheng, Mishra, Siddhartha
Abstract
The widely used Transolver family of neural operators is based on physics-attention, which softly assigns the points of an unstructured mesh to a small number of slices, applies self-attention among the resulting tokens, and broadcasts the result back to the points. We provide a comprehensive empirical and theoretical analysis to elucidate the mechanisms which are responsible for model performance. To this end, we perform careful ablations on a challenging suite of nine 3D fluid dynamics benchmarks to find that replacing token attention with a constant linear map does not affect the accuracy. Thus, Transolver does not need a Transformer at all. However, removing the global mixing (slicing/deslicing) or doing it only once leads to performance collapse. We leverage the theory of averaging neural operators to explain and corroborate our findings by showing that just slicing/deslicing, in conjunction with pointwise MLPs, already suffices for universal approximation of continuous operators and attention is redundant in this context. Finally, we provide a novel FlashAttention-style efficient implementation of the key slicing/deslicing module of Transolver. This flashslice kernel streams slice and deslice over the points without materializing the heavy slice-weight tensor, while reproducing the best available implementation to floating point error. At the same time, it leads to very significant memory and compute savings, particularly at large slice counts.
Chinese Translation
广泛使用的 Transolver 系列神经算子基于物理注意力机制,该机制将非结构化网格中的点软分配至少数几个切片,在得到的令牌之间应用自注意力,并将结果广播回这些点。我们提供了全面的实证和理论分析,以阐明决定模型性能的机制。为此,我们在九个具有挑战性的三维流体动力学基准测试上进行了细致的消融实验,发现将令牌注意力替换为常数线性映射并不会影响精度。因此,Transolver 根本不需要 Transformer。然而,移除全局混合(切片/去切片)或仅执行一次会导致性能崩溃。我们利用平均神经算子理论来解释和证实我们的发现,表明仅需切片/去切片与逐点 MLP 相结合,就已足以实现连续算子的通用逼近,而在此背景下注意力是冗余的。最后,我们为 Transolver 的关键切片/去切片模块提供了一种新颖的 FlashAttention 风格的高效实现。该 flashslice 内核在点上流式执行切片和去切片,无需实例化沉重的切片权重张量,同时将最佳可用实现复现至浮点误差。同时,它显著节省了内存和计算量,尤其是在切片数量较大时。
cs.LG / 132 / 2609.32530

Activation Flow: Manufacturing Activations for Steering

激活流:制造用于引导的激活
Tan, Hong Kiat, Le, Linh, Williams-King, David
Abstract
Difference-in-means steering requires activations recorded while a model shows the desired behavior, which a sandbagging model withholds by deliberately underperforming. We introduce Activation Flow (ActFlow), which manufactures these activations from $k$ correct labels without fine-tuning. ActFlow sets target logits that rank each labeled item's correct answer first, and moves the logits toward them by adding one vector $x$ to all $k$ residual streams at one layer. ActFlow is a family of ordinary differential equations for $x$, one for each rule that maps the required logit change to the velocity of $x$. The smallest-norm rule lands exactly on the targets, while the others keep only the top singular directions of the Jacobian. We test ActFlow on three instruction-tuned models, each locked by a sandbagging prompt and by a password-locked LoRA. At $k=40$, ActFlow keeping five singular directions raises the mean held-out ARC-Easy accuracy over the six locked models from $0.05$ to $0.85$, against $0.88$ for fine-tuning and $0.92$ for the honest models. Furthermore, it scores higher than the smallest-norm rule in 16 of the 18 combinations of locked model and $k$, and its steering direction is nearly orthogonal to the honest difference-in-means direction. It also unlocks two LoRA locks where the honest direction fails.
Chinese Translation
均值差引导需要记录模型表现出期望行为时的激活,而故意示弱模型通过故意表现不佳来不提供这些激活。我们提出激活流(ActFlow),它无需微调即可从 $k$ 个正确标签制造这些激活。ActFlow 设置目标 logits,将每个标记项的正确答案排在第一,并通过在一层向所有 $k$ 个残差流添加一个向量 $x$ 来使 logits 向它们移动。ActFlow 是 $x$ 的一族常微分方程,每个规则对应一个方程,将所需的 logit 变化映射到 $x$ 的速度。最小范数规则精确落在目标上,而其他规则仅保留雅可比矩阵的顶部奇异方向。我们在三个指令微调模型上测试 ActFlow,每个模型都被一个故意示弱提示和一个密码锁定的 LoRA 锁定。在 $k=40$ 时,保留五个奇异方向的 ActFlow 将六个锁定模型的平均留出 ARC-Easy 准确率从 $0.05$ 提高到 $0.85$,而微调为 $0.88$,诚实模型为 $0.92$。此外,在 18 种锁定模型和 $k$ 的组合中,它有 16 种得分高于最小范数规则,并且其引导方向几乎正交于诚实的均值差方向。它还能解锁两个诚实方向无法解锁的 LoRA 锁。
cs.LG / 133 / 2609.32532

What Does a ProcGen Generalization Gap Measure? Action Rules, Convergence, and the Missing Random Floor

ProcGen 泛化差距衡量的是什么?动作规则、收敛性与缺失的随机下限
Keshari, Abhisek
Abstract
A generalization gap in reinforcement learning (return on training levels minus return on held-out levels) is usually reported without a reference point. We argue that the missing reference is a measured random floor: the return of a uniform-random policy on the same levels under the same harness. On ProcGen, the floor changes what several standard numbers mean. On identical checkpoints and levels across eight environments, switching between sampled and greedy (argmax) test-time actions moves held-out return in both directions, and greedy evaluation takes three environments to or below the floor: in miner, the sampled policy scores 4.9x the floor on held-out levels while its argmax scores below it. Raw policy entropy places seven of eight environments short of convergence, but 35-65% of that entropy is spread across actions with identical effects; after merging them, one to three remain short, and against the floor only heist has learned nothing that transfers. An audit of twelve prior ProcGen codebases finds that all eleven with held-out evaluation sample test-time actions for their policy-gradient agents, nine by default rather than explicit choice, and six report running in-loop averages rather than evaluating a fixed checkpoint. Applied to our own case study, the same checks grade down a statistically significant encoder effect and rule out a within-encoder train-vs-test CKA statistic. We recommend that every reported gap state its action rule, use a matched and seeded protocol, and report the random floor on both level sets.
Chinese Translation
强化学习中的泛化差距(训练关卡上的回报减去留出关卡上的回报)通常在没有参考点的情况下被报告。我们认为,缺失的参考是一个测量得到的随机下限:即在同一评估框架下,均匀随机策略在相同关卡上的回报。在 ProcGen 上,该下限改变了一些标准数值的含义。在八个环境中,使用相同的检查点和关卡,在采样动作与贪婪(argmax)测试时动作之间切换会使留出回报朝两个方向变化,而贪婪评估使三个环境达到或低于该下限:在 miner 中,采样策略在留出关卡上的得分是该下限的 4.9 倍,而其 argmax 得分低于该下限。原始策略熵显示八个环境中有七个尚未收敛,但其中 35–65% 的熵分布在效果相同的动作上;合并这些动作后,只有一到三个仍尚未收敛,并且相对于该下限,只有 heist 没有学到任何可迁移的东西。对十二个先前的 ProcGen 代码库的审计发现,其中所有十一个进行留出评估的代码库都为其策略梯度智能体采样测试时动作,九个是默认而非显式选择,六个报告的是运行中的循环内平均值,而不是评估固定检查点。应用于我们自己的案例研究时,同样的检查降低了一个统计显著的编码器效应的评级,并排除了一个编码器内训练与测试的 CKA 统计量。我们建议,每个报告的差距都应说明其动作规则,使用匹配且固定随机种子的协议,并在两个关卡集上报告随机下限。
cs.LG / 134 / 2609.32546

Shared Autoregressive Context Can Distort Relationships in Synthetic Data

共享自回归上下文可能扭曲合成数据中的关系
Robinson, Thomas S.
Abstract
Large language models can generate several records within one autoregressive completion, making earlier answers available as context for later records. This paper shows that such shared-completion batching can distort relationships among variables in the resulting synthetic data, using controlled tests on synthetic survey respondents. In a matched experiment on 2,000 European Social Survey profiles, generating ten rather than one respondent per request increases mean absolute error in within-country correlations by 48-58% for Qwen3.8-27B and 114-127% for Llama-3.3-70B-Instruct across three seeds, holding profiles, examples, questions and decoding parameters fixed. The distortion primarily reflects exaggerated relationship strength, while retaining substantial agreement with the human ordering of correlations. Controlled interventions establish answer history as a causal channel: re-pairing the same preceding values, with profiles and marginal distributions fixed, changes correlations among subsequently generated responses. Hiding preceding answers reduces correlation error in the tested settings but worsens marginal accuracy. Exploratory corrections across social-attitude, health and economic data likewise show that lower correlation error can coexist with worse marginal distributions and regression estimates. Request construction is therefore part of the data-generating process, and synthetic-data validity must be evaluated against the analyses the generated data are intended to support.
Chinese Translation
大型语言模型可以在一次自回归补全中生成多条记录,使先前的答案可作为后续记录的上下文。本文表明,这种共享补全批处理可能会扭曲所得合成数据中变量之间的关系,并使用合成调查受访者进行了受控测试。在一项针对2,000个欧洲社会调查(European Social Survey)档案的匹配实验中,对于Qwen3.8-27B和Llama-3.3-70B-Instruct,在三个随机种子下,每次请求生成十个而非一个受访者,会使国内相关系数的平均绝对误差分别增加48-58%和114-127%,同时保持档案、示例、问题和解码参数固定不变。这种扭曲主要反映为关系强度的夸大,同时与人类相关性排序保持高度一致。受控干预确立了答案历史作为一种因果渠道:在固定档案和边际分布的情况下,重新配对相同的前置值,会改变随后生成的响应之间的相关性。在测试设置中,隐藏前置答案会降低相关性误差,但会恶化边际准确性。对社会态度、健康和经济数据的探索性校正同样表明,较低的相关性误差可能与较差的边际分布和回归估计并存。因此,请求构建是数据生成过程的一部分,合成数据的有效性必须根据生成数据所旨在支持的分析进行评估。
cs.LG / 135 / 2609.32547

Prioritizing Repeated LLM Evaluation for Hidden Failure Discovery

优先进行重复LLM评估以发现隐藏故障
Broadwater, Keita, Broadwater, Akin
Abstract
Large language models are commonly evaluated by generating a small number of stochastic responses for each prompt in a benchmark. Because inference budgets are limited, this shallow evaluation may fail to observe low-probability but operationally important failures. A prompt that produces no failures in a small sample may therefore appear reliable despite having a nonzero latent probability of failure under repeated inference. We formulate LLM reliability evaluation as a budget-constrained discovery problem in which each prompt is associated with an unknown per-generation failure probability. We propose a budgeted discovery framework that first performs shallow evaluation across the prompt set and then uses trial-level failure outcomes together with prompt-derived representations to learn a feature-based ranking of failure propensity. The resulting scores prioritize prompts with zero observed shallow failures for deeper evaluation, concentrating the deep-evaluation budget where hidden failures are more likely to be discovered. We evaluate this approach on AIRBench and StrongREJECT across multiple model and system-prompt conditions. The central empirical test asks whether models fit without access to deep-evaluation outcomes can rank prompts with zero observed shallow failures according to their likelihood of producing failures under deeper evaluation. On AIRBench, the highest-ranked 10\% of unresolved prompts achieves 2.54x hidden-failure lift for Qwen 2.5 7B and 1.87x for Gemma 3n E4B, recovering 25.4\% and 18.7\% of subsequently observed hidden failures, respectively, compared with 10\% expected under random allocation. Semantic-neighborhood and feature-ablation analyses further show that this predictive signal can be recovered from multiple representations of prompt content and relationships.
Chinese Translation
大语言模型通常通过为基准测试中的每个提示生成少量随机响应来进行评估。由于推理预算有限,这种浅层评估可能无法观察到低概率但操作上重要的故障。因此,在小样本中未产生故障的提示可能看起来可靠,尽管在重复推理下具有非零的潜在故障概率。我们将LLM可靠性评估形式化为一个预算约束的发现问题,其中每个提示都与未知的每次生成故障概率相关联。我们提出了一个预算发现框架,该框架首先对提示集进行浅层评估,然后利用试验级故障结果以及从提示派生的表示来学习基于特征的故障倾向排序。得到的分数优先考虑零观察浅层故障的提示进行更深评估,将深度评估预算集中在更可能发现隐藏故障的地方。我们在AIRBench和StrongREJECT上跨多个模型和系统提示条件评估了该方法。核心实证测试询问,在不访问深度评估结果的情况下拟合的模型,能否根据在更深评估下产生故障的可能性,对零观察浅层故障的提示进行排序。在AIRBench上,排名最高的10%未解决提示对Qwen 2.5 7B实现了2.54倍的隐藏故障提升,对Gemma 3n E4B实现了1.87倍,分别恢复了后续观察到的隐藏故障的25.4%和18.7%,而随机分配下预期为10%。语义邻域和特征消融分析进一步表明,这种预测信号可以从提示内容和关系的多种表示中恢复。
cs.LG / 136 / 2609.32566

Compositional Objectives: Learning Structure in Structure

组合目标:在结构中学习结构
Vivekananda, Pranavchandra, Bettadapura, Sumukh, Subramanian, Ajan
Abstract
Intelligence is defined in many ways. One of these definitions defines intelligence as the pursuit of learnable novelty. However, learnable novelty can be meaningless without the ability to compose the learned structures to take action and achieve goals. Learnable novelty builds on epiplexity, which is a way to measure learnable structure in data through a bounded observer. In this paper, we investigate a closed-form spectral approximation to compute epiplexity. We use a fixed-trace constraint and find that the epiplexity objective prefers a more uniform distribution of spectral mass rather than concentrating it in a small number of directions. However, a representation may spread information across many directions without organizing that information into features useful for a particular task. To address this gap, we propose a compositional objective whose observer measures the relationships between the parts and interactions of an image. We compare it with the original spectral objective given only the masked parts. In our ImageNet training runs, the spectral objective with masked parts produces an almost maximally spread representation while achieving the strongest frozen-feature classification performance on most evaluations, more than doubling the linear-probe accuracy of the whole-image baseline. Across multiple image benchmarks, changing what the observer sees matters more than adding relation and interaction tokens. At the same time, our prediction-oriented compositional objective produces substantially better held-out observer prediction but relatively weaker classification, revealing that spectral diversity, predictability, and downstream utility are distinct properties. These results suggest that the usefulness of spectral spreading depends not only on how much structure is preserved, but on which relationships the observer makes available to the objective.
Chinese Translation
智能有多种定义方式。其中一种定义将智能视为对可学习新颖性的追求。然而,如果没有能力将学到的结构组合起来以采取行动并实现目标,可学习的新颖性可能毫无意义。可学习的新颖性建立在epiplexity之上,epiplexity是一种通过有界观察者测量数据中可学习结构的方法。在本文中,我们研究了一种用于计算epiplexity的闭式谱近似。我们使用固定迹约束,发现epiplexity目标更倾向于谱质量的更均匀分布,而不是将其集中在少数几个方向上。然而,表示可能将信息分散到许多方向,而没有将这些信息组织成对特定任务有用的特征。为了弥补这一差距,我们提出了一种组合目标,其观察者测量图像各部分和交互之间的关系。我们将其与仅给定掩码部分的原始谱目标进行比较。在我们的ImageNet训练运行中,带有掩码部分的谱目标产生了几乎最大程度分散的表示,同时在大多数评估中实现了最强的冻结特征分类性能,将线性探测准确率提高到整个图像基线的两倍以上。在多个图像基准测试中,改变观察者所看到的内容比添加关系和交互标记更重要。同时,我们面向预测的组合目标产生了明显更好的留出观察者预测,但分类相对较弱,揭示了谱多样性、可预测性和下游效用是不同的属性。这些结果表明,谱扩散的有用性不仅取决于保留了多少结构,还取决于观察者向目标提供了哪些关系。
cs.LG / 137 / 2609.32570

CAESAR: Clustering via Autonomous Embedding-Space Agglomerative Reorganization

CAESAR:通过自主嵌入空间凝聚重组的聚类
Bacry, Ilan, Devaux, Rémi, Jardin, Antoine
Abstract
Clustering algorithms that operate on nearest-neighbor graphs, such as FINCH (First Integer Neighbor Clustering Hierarchy), depend heavily on the quality of the embedding space they are given. However, pretrained vision and language model embeddings are not optimized for this purpose. We propose CAESAR, a method that reorganizes a pretrained embedding space: a reorganization network is trained to pull mutual nearest neighbors together and push non-neighbors apart, yielding a reorganized embedding space substantially better suited to clustering. CAESAR offers a second major advantage: it never requires the number of clusters $K$. This matters because in realistic unsupervised settings, $K$ is typically unknown and discovering it is often part of the problem, yet most strong clustering methods take it as an input. We therefore design the entire CAESAR pipeline to infer $K$ rather than assume it is known. Empirically, reorganizing the embeddings consistently improves clustering over the raw space on both text and image datasets. Since the few deep clustering methods that also infer $K$ do not release their code, we complement controlled comparisons with methods that infer $K$ on the same embedding space by comparisons with strong deep clustering methods that are given the true $K$, giving them a substantial oracle advantage. Even so, CAESAR outperforms all of them on text, achieves the best results on the most challenging image benchmark and remains competitive on the others. Reorganizing pretrained embeddings thus emerges as a simple and powerful route to clustering realistic data, where classes overlap and the number of clusters is unknown.
Chinese Translation
基于最近邻图进行操作的聚类算法,例如 FINCH(First Integer Neighbor Clustering Hierarchy),严重依赖于所给嵌入空间的质量。然而,预训练视觉和语言模型嵌入并未针对此目的进行优化。我们提出 CAESAR,一种重组预训练嵌入空间的方法:训练一个重组网络,将互为最近邻的样本拉近,将非邻居推远,从而得到一个更适合聚类的重组嵌入空间。CAESAR 的第二个主要优势是:它不需要聚类数 $K$。这一点很重要,因为在现实的无监督设置中,$K$ 通常是未知的,而发现 $K$ 往往正是问题的一部分,然而大多数强大的聚类方法将其作为输入。因此,我们将整个 CAESAR 流程设计为推断 $K$,而不是假设其已知。实验表明,在文本和图像数据集上,重组嵌入空间相比于原始空间能够持续提升聚类效果。由于那些同样能推断 $K$ 的少数深度聚类方法没有公开代码,我们通过将能够在相同嵌入空间上推断 $K$ 的方法与那些给定真实 $K$ 的强大深度聚类方法进行比较,来补充受控比较,这赋予了后者显著的神谕优势(oracle advantage)。即便如此,CAESAR 在文本上超越了所有方法,在最具挑战性的图像基准上取得了最佳结果,在其他数据集上也保持竞争力。因此,重组预训练嵌入成为对现实数据进行聚类的一条简单而强大的途径,尤其当类别重叠且聚类数未知时。
cs.LG / 138 / 2609.32573

Trapped by Their Own Rollouts: Understanding Aggregation--Rollout Feedback in Federated On-Policy Distillation

被自身轨迹所困:理解联邦在线策略蒸馏中的聚合--轨迹反馈
Chen, Jinqian, Zhu, Jihua, Liu, Chang
Abstract
On-policy distillation (OPD) is a promising approach to language-model adaptation, aligning teacher supervision with the student's own generated trajectories. When adaptation prompts are distributed across clients, can this process benefit from federated collaboration? We study federated OPD and find that substantial collaboration gains can be obscured by learning-rate sensitivity: FedAvg can perform no better than independent local training at a small learning rate, yet recover a clear advantage at a larger rate. We explain this phenomenon through the student's dual role as learner and generator of future training data. An aggregation-induced optimization lag can delay access to useful teacher supervision, which in turn slows subsequent learning. Our theory establishes this aggregation--rollout feedback in a solvable model with a common optimum and stable updates, and identifies two coupled roles of learning rate: learning from current supervision and reaching future supervision. Guided by this analysis, we propose FedTOPS (Federated Teacher-guided On-Policy Scaling), which reuses teacher feedback on current trajectories to adapt the FedAvg update magnitude under clientwise predictive-change constraints. Across six mathematical reasoning benchmarks, FedTOPS improves macro Avg@8 over FedAvg by 4.56--14.57 percentage points across the evaluated student models and local learning rates.
Chinese Translation
在线策略蒸馏(OPD)是一种很有前景的语言模型适配方法,它将教师监督与学生的自生成轨迹对齐。当适配提示分布在客户端上时,这个过程能否从联邦协作中受益?我们研究了联邦 OPD,发现显著的合作增益可能被学习率敏感性所掩盖:在较小的学习率下,FedAvg 的表现可能并不优于独立的本地训练,但在较大的学习率下却恢复了明显的优势。我们通过学生作为学习者和未来训练数据生成者的双重角色来解释这一现象。聚合引起的优化滞后可能会延迟获得有用的教师监督,进而减缓后续学习。我们的理论在一个具有共同最优解和稳定更新的可解模型中建立了这种聚合--轨迹反馈,并识别了学习率的两个耦合作用:从当前监督中学习和达到未来监督。在此分析的指导下,我们提出了 FedTOPS(联邦教师引导的在线策略缩放),它重用当前轨迹上的教师反馈,在客户端预测变化约束下调整 FedAvg 的更新幅度。在六个数学推理基准上,FedTOPS 在所评估的学生模型和本地学习率上将宏平均 Avg@8 相比 FedAvg 提高了 4.56--14.57 个百分点。
cs.LG / 139 / 2609.32579

DimPO: Dimensionality Reduction for Attention using Preference Optimization

DimPO:使用偏好优化对注意力进行降维
Lanz, Vojtěch, Cui, Yufei, Parthasarathi, Prasanna
Abstract
A linear projection can reduce the dimension of query and key vectors without updating the pretrained model, but it remains unclear which training objective best preserves model behavior. We ask whether preferences over keys and attention mass on the highest-weighted keys provide a better signal than matching the full attention distribution with KL divergence, especially in long-context settings. We introduce DimPO, which combines listwise preference optimization with a lightweight top-k cross-entropy term for head-fidelity. DimPO is trained offline from the attention patterns of a frozen language model, with one map per layer, shared by the query and the keys and trained separately from the other layers. Across LLaMA3.2-3B, LLaMA3.1-8B, Qwen2.5-7B, and Qwen3-4B Instruct models, pairwise preference objectives outperform the triplet baseline and retain 98% of the original score on short-context tasks when projecting to half the dimension on the last 40% of the layers. With more projected layers or on long-context RULER, they degrade rapidly. In contrast, KL and DimPO, which use every key during training, retain about 95% of the original RULER 4k score on the 8B model when projecting up to 50% of the layers. KL-based projections remain closer to the original attention distribution and attention output, yet DimPO achieves better downstream performance. Beyond 50% of projected layers, DimPO increasingly outperforms KL on tasks including SQuAD, common-word extraction, frequent-word extraction, and variable tracking. These results suggest that under dimensionality reduction, preserving the ordering and concentration of task-relevant attention can matter more than reproducing the full attention distribution.
Chinese Translation
线性投影可以在不更新预训练模型的情况下降低查询和键向量的维度,但哪种训练目标最能保持模型行为仍不清楚。我们探究,在长上下文设置中,对键的偏好以及最高权重键上的注意力质量是否比用KL散度匹配完整注意力分布提供更好的信号。我们提出DimPO,它将列表式偏好优化与轻量级top-k交叉熵项相结合,以保持头部保真度。DimPO从冻结语言模型的注意力模式中离线训练,每层一个映射,由查询和键共享,并与其他层分开训练。在LLaMA3.2-3B、LLaMA3.1-8B、Qwen2.5-7B和Qwen3-4B Instruct模型上,当成对偏好目标在最后40%的层上投影到一半维度时,它们优于三元组基线,并在短上下文任务上保留了原始得分的98%。当投影更多层或在长上下文RULER上时,它们迅速下降。相比之下,KL和DimPO在训练中使用每个键,当在8B模型上投影最多50%的层时,它们保留了原始RULER 4k分数的约95%。基于KL的投影仍然更接近原始注意力分布和注意力输出,但DimPO实现了更好的下游性能。当投影层超过50%时,DimPO在包括SQuAD、常见词提取、高频词提取和变量追踪等任务上越来越优于KL。这些结果表明,在降维下,保持任务相关注意力的顺序和集中度可能比复现完整的注意力分布更重要。
cs.LG / 140 / 2609.32593

Age of Learning: Temporal Persistence of Prediction Errors as a Learning Signal

学习年龄:预测误差作为学习信号的时间持续性
Wang, Chenyang, Forsström, Stefan, Olsson, Roger, Yuan, Di, He, Qing
Abstract
Current machine learning algorithms primarily rely on instantaneous signals such as loss, margin, and prediction confidence to characterize model behavior. These signals indicate how difficult a prediction is at the current optimization step, but they do not capture how long the model has remained incorrect. We study this temporal dimension of learning and introduce Age of Learning (AoL), a learning-state variable that measures the persistence of prediction errors over time. AoL increases while an error remains unresolved and resets when a correct prediction is achieved, thereby distinguishing persistent under-learning from transient mistakes. We develop AoL-based training strategies for both offline and streaming settings. In offline learning, sample-level AoL is accumulated over training and aggregated into class-level states that guide adaptive reweighting and resampling. In streaming learning, where full historical access is unavailable, we maintain lightweight class-level AoL states using current and buffered observations. Across long-tailed classification settings, AoL improves or matches standard training baselines, with larger benefits when learning difficulty persists over time. Multi-seed streaming experiments further show reproducible gains under temporally stable imbalance. Analysis of class frequency, loss, and margin shows that AoL is related to conventional difficulty measures but captures additional information about error duration. These results suggest that temporal persistence provides a useful complementary signal for characterizing and controlling learning dynamics in imbalanced and non-stationary environments.
Chinese Translation
当前的机器学习算法主要依赖损失、间隔和预测置信度等瞬时信号来刻画模型行为。这些信号指示了在当前优化步骤中预测的难度,但并未捕捉模型保持错误状态的时间长度。我们研究了学习的这一时间维度,并引入了学习年龄(AoL),一种衡量预测误差随时间持续性的学习状态变量。当错误未被解决时,AoL增加;当获得正确预测时,AoL重置,从而区分持续的学习不足与暂时性错误。我们为离线和流式设置开发了基于AoL的训练策略。在离线学习中,样本级AoL在训练过程中累积,并聚合为类级状态,以指导自适应重加权和重采样。在流式学习中,由于无法完全访问历史数据,我们利用当前和缓冲的观测值来维护轻量级的类级AoL状态。在长尾分类设置中,AoL提升或匹配了标准训练基线,当学习难度随时间持续时,收益更大。多种子流式实验进一步表明,在时间稳定的不平衡下具有可复现的增益。对类频率、损失和间隔的分析表明,AoL与传统难度度量相关,但捕获了关于错误持续时间的额外信息。这些结果表明,时间持续性为刻画和控制不平衡和非平稳环境中的学习动态提供了一种有用的补充信号。
cs.LG / 141 / 2609.32597

Not Every Term Adds New Structure: Sobolev Novelty for Symbolic Regression

并非每个项都增加新结构:面向符号回归的Sobolev新颖性
Wang, Boxiao, Li, Kai, Jing, Yuheng, Liu, Tianyi, Li, Chen, Zhang, Yifan, Cheng, Jian
Abstract
Symbolic regression (SR) aims to discover compact and meaningful mathematical equations from data, but searching the vast combinatorial space of symbolic structures remains challenging. Existing methods typically guide this process using expression-level objectives, such as fitting error, which assess a candidate equation as a whole but provide little information about whether an individual term contributes genuinely new structure or is largely redundant with the rest of the expression. We introduce \textbf{Sobolev Novelty}, a term-level measure of structural independence for symbolic equations. For each term, we construct an empirical Sobolev signature from its function values and exact derivatives over the observed inputs, and quantify how much of this behavior cannot be reconstructed by the remaining terms. We further derive a theory-calibrated threshold, yielding a principled and tuning-free criterion for identifying structurally novel terms. Using this threshold, 92.6\% of terms in benchmark ground-truth equations exhibit sufficient structural novelty, compared with only 38.3\% on average for expressions produced by 15 SR methods, revealing a substantial gap between scientific equations and current SR solutions. As a lightweight plug-in, Sobolev Novelty can be incorporated into diverse SR paradigms to support term pruning, search guidance, LLM feedback, and data selection, yielding consistent performance gains and demonstrating broad applicability.
Chinese Translation
符号回归(SR)旨在从数据中发现紧凑且有意义的数学方程,但搜索符号结构的巨大组合空间仍然具有挑战性。现有方法通常使用表达式级目标(如拟合误差)来指导这一过程,这些目标将候选方程作为一个整体进行评估,但很少提供关于单个项是否贡献了真正的新结构或与表达式其余部分在很大程度上冗余的信息。我们引入了Sobolev新颖性(Sobolev Novelty),这是一种针对符号方程的项级结构独立性度量。对于每一项,我们根据其在观测输入上的函数值和精确导数构建经验Sobolev签名,并量化该行为中有多少无法由其余项重构。我们进一步推导出一个理论校准的阈值,从而得到一个有原则且无需调参的准则,用于识别结构新颖的项。使用该阈值,基准真值方程中92.6%的项表现出充分的结构新颖性,而15种SR方法产生的表达式平均只有38.3%,这揭示了科学方程与当前SR解决方案之间的显著差距。作为一个轻量级插件,Sobolev新颖性可以融入多种SR范式中,以支持项剪枝、搜索指导、LLM反馈和数据选择,从而带来一致的性能提升并展现出广泛的适用性。
cs.LG / 142 / 2609.32602

AnchorRep: Defending LLMs Against Cross-Model Adversarial Transfer via Representation Repulsion

AnchorRep:通过表示排斥防御大型语言模型跨模型对抗迁移
Wertheizer, Gal, Himelstein, Rom, Peretz, Tomer, Mendelson, Avi
Abstract
Adversarial attacks optimized on a single open-weight LLM can transfer to and jailbreak architecturally different models, allowing an attacker with white-box access to one model to compromise independently deployed systems. This creates a shared vulnerability across models, yet existing defenses are not designed for this cross-model threat. We find that cross-model transfer aligns with shared internal representation geometry, making it a natural defense target. AnchorRep targets this geometry directly with a lightweight LoRA adapter that pushes the defended model's internal representations of harmful prompts away from those of a frozen anchor model on the same prompts. Training uses a small set of harmful prompts and no adversarial examples. Across five models and four architectural families, AnchorRep reduces cross-model attack success rate to <=1.1% on 2,000 transferred attacks (0% on two), including the largest drop on Mistral (36% -> 1.1%). Existing defenses can reduce transfer, but only at high cost either inducing up to 77% degenerate benign output or increasing over-refusal by up to 18%. Because such degenerate benign outputs are not captured by standard refusal-based metrics, we introduce the Benign Garble Rate to quantify them. Our results suggest that cross-model robustness can be achieved by shaping representation geometry, without requiring attack-specific training
Chinese Translation
在单个开放权重的大型语言模型上优化的对抗攻击可以迁移到架构不同的模型并越狱它们,使得对一个模型具有白盒访问权限的攻击者能够攻陷独立部署的系统。这造成了跨模型的共享漏洞,但现有防御并非针对这种跨模型威胁而设计。我们发现,跨模型迁移与共享的内部表示几何结构相一致,使其成为天然的防御目标。AnchorRep 直接针对这种几何结构,使用轻量级 LoRA 适配器,将被防御模型对有害提示的内部表示推离冻结的锚模型对相同提示的内部表示。训练使用一小组有害提示,而不使用对抗样本。在五种模型和四个架构家族上,AnchorRep 在 2,000 次迁移攻击中将跨模型攻击成功率降至不超过 1.1%(在两个模型上为 0%),其中包括在 Mistral 上的最大降幅(36% -> 1.1%)。现有防御可以减少迁移,但代价高昂:要么导致高达 77% 的退化良性输出,要么将过度拒绝率提高高达 18%。由于标准的基于拒绝的指标无法捕捉此类退化良性输出,我们引入了良性乱码率(Benign Garble Rate)来量化它们。我们的结果表明,通过塑造表示几何结构可以实现跨模型鲁棒性,而无需针对特定攻击进行训练。
cs.LG / 143 / 2609.32613

Learning the Graph and the Embedding Together: Classifier-Independent Rewiring for Heterophilic Node Classification

联合学习图与嵌入:面向异质节点分类的分类器无关重连
Kumar, Harshit, Chakraborty, Sujan, Saha, Priyanka, Kar, Pritam, Bej, Saptarshi
Abstract
Graph neural networks lose much of their advantage on heterophilic graphs, where connected nodes often carry different labels. Graph rewiring is a popular remedy, but rewiring methods are usually evaluated with a single classifier, which makes it hard to tell whether the gains come from the new topology or from that particular pairing. We propose an affinity-guided rewiring method that estimates the graph and the node representation together. It alternates, in the spirit of expectation maximisation, between training a lightweight graph neural network on the current graph and re-weighting candidate edges under a modularity objective with a pseudo-label homophily term. Candidate edges come from a compact pool scored by a contrastively learned node similarity and a neighbourhood-distribution affinity. The method returns two classifier-independent outputs: a rewired graph and a node embedding learned on it. Across six heterophilic benchmarks and five downstream classifiers, it improves accuracy over the original graph with normalised features in 23 of 30 classifier-dataset combinations, with a mean gain of 5.8 points, and reduces the accuracy spread between classifiers about fourfold. A controlled ablation shows that the two outputs are each useful and play complementary roles: the embedding contributes most of the accuracy gain, while the rewired graph makes different classifiers agree. A fully unsupervised variant, which uses no labels during rewiring, retains most of the improvement. The rewired graphs are also more homophilic and improve label propagation and community detection.
Chinese Translation
图神经网络在异质图上失去大部分优势,其中相连节点往往携带不同标签。图重连是一种流行的补救措施,但重连方法通常用单个分类器进行评估,这使得难以判断增益是来自新拓扑还是来自该特定配对。我们提出了一种亲和力引导的重连方法,它同时估计图和节点表示。它以期望最大化的精神,在当前图上训练轻量级图神经网络与在带有伪标签同质性项的模块度目标下重新加权候选边之间交替。候选边来自一个紧凑的池,由对比学习的节点相似度和邻域分布亲和力进行评分。该方法返回两个分类器无关的输出:一个重连图和一个在其上学习的节点嵌入。在六个异质基准和五个下游分类器上,它在30个分类器-数据集组合中的23个上,相较于具有归一化特征的原始图提高了准确率,平均提升5.8个百分点,并将分类器之间的准确率差异缩小了约四倍。一项受控消融实验表明,这两个输出各自有用且发挥互补作用:嵌入贡献了大部分准确率增益,而重连图使不同分类器趋于一致。一个完全无监督的变体,在重连过程中不使用标签,保留了大部分改进。重连后的图也更加同质,并改善了标签传播和社区检测。
cs.LG / 144 / 2609.32615

Stabilizing the Dynamic Low-Rank Training

稳定动态低秩训练
Xu, Zhonghan, Wang, Ling, Chen, Junhao, Zhao, Jianwei, Yang, Jinwei
Abstract
Training neural networks directly in a low-rank parameterization is an appealing route to reducing memory, compute, and storage simultaneously during both training and inference. Dynamic low-rank training (DLRT), which confines weights to a rank-$r$ manifold via the Galerkin projection of the gradient flow, is particularly attractive because it identifies efficient subnetworks on the fly without specialized initialization or post-factorization. However, DLRT fails to find trainable networks under high compression. In this paper, we derive the gradient flow of the best rank-$r$ approximation and point out that the offset of DLRT comes from a curvature-coupling term which is large and thus non-negligible under aggressive compression. Guided by this analysis, we propose a stable dynamic low-rank training method, named SDLRT, which maintains a lightweight compensation buffer that reinjects the top neglected singular directions. Additionally, we introduce a negative feedback on the truncation tolerance to stabilize each layer's rank. Experimentally, SDLRT reliably finds trainable subnetworks where DLRT collapses and as a PEFT adapter on DeBERTa-v3, it achieves the best average score on SuperGLUE at only $2.8\%$ parameter overhead over LoRA.
Chinese Translation
直接以低秩参数化训练神经网络是一条有吸引力的途径,可以同时在训练和推理期间减少内存、计算和存储。动态低秩训练(DLRT)通过梯度流的Galerkin投影将权重限制在秩$r$流形上,它特别有吸引力,因为它可以在运行中识别有效的子网络,而无需专门的初始化或后分解。然而,在高压缩下,DLRT无法找到可训练的网络。在本文中,我们推导了最佳秩$r$近似的梯度流,并指出DLRT的偏差来自一个曲率耦合项,该项在激进压缩下很大,因此不可忽略。在此分析的指导下,我们提出了一种稳定的动态低秩训练方法,称为SDLRT,它维护一个轻量级的补偿缓冲区,重新注入被忽略的主要奇异方向。此外,我们引入了对截断容差的负反馈,以稳定每一层的秩。实验上,SDLRT在DLRT崩溃的情况下可靠地找到了可训练的子网络,并且作为DeBERTa-v3上的PEFT适配器,它在SuperGLUE上取得了最佳平均分数,仅比LoRA多$2.8\%$的参数开销。
cs.LG / 145 / 2609.32619

Intuition vectors

直觉向量
Haim, Shahar, McNamee, Daniel C.
Abstract
Large self-supervised vision models learn representations that support scene segmentation and the semantic decomposition of physical objects. We ask whether their representational geometry supports transfer to visual reasoning problems without any task-specific fine-tuning. We hypothesized that relational representations may bridge perception and abstract reasoning by encoding similarities and transformations among visual inputs such that an intuitive, implicit form of reasoning may be performed via latent vector arithmetic. Specifically, we examine DINOv3, MAE, and random pixel projections on abstract and naturalistic Bongard problems, ARC-AGI-1 and ARC-AGI-2, and novel ARC-GEN instances. On both Bongard benchmarks, the accuracy of a simple nearest-centroid readout of frozen visual embeddings is within four percentage points of the task-specific baselines reported with the original benchmarks. In ARC, latent difference vectors summarizing demonstration input-output transformations, which we refer to as intuition vectors, show greater alignment with test vectors from the same task, whereas those from unrelated tasks are near orthogonal. This latent geometry is operational: transporting a query along its intuition vector consistently improves exact-output retrieval, reaching 70.7 on ARC-AGI-2 evaluation. Across 397,000 ARC-GEN instances from 794 tasks, single-pair intuition vectors identify the generating task with approximately 87\% leave-one-out accuracy. These findings suggest that latent vector arithmetic over frozen visual representations supports implicit rule inference across varied problem domains without a generative model component, indicating that inferring an abstract transformation and generating its instance-specific consequence may be separable capacities.
Chinese Translation
大型自监督视觉模型学习到的表示能够支持场景分割和物理对象的语义分解。我们探究其表示几何是否支持在无需任何任务特定微调的情况下迁移到视觉推理问题。我们假设关系表示可以通过编码视觉输入之间的相似性和变换来连接感知和抽象推理,从而可以通过潜在向量算术执行一种直观的、隐式的推理形式。具体而言,我们在抽象和自然的 Bongard 问题、ARC-AGI-1 和 ARC-AGI-2 以及新颖的 ARC-GEN 实例上检查了 DINOv3、MAE 和随机像素投影。在两个 Bongard 基准上,对冻结视觉嵌入进行简单最近质心读取的准确率与原始基准报告的任务特定基线相差在四个百分点以内。在 ARC 中,总结演示输入-输出变换的潜在差异向量(我们称之为直觉向量)与同一任务的测试向量表现出更大的对齐,而来自无关任务的向量则接近正交。这种潜在几何是可操作的:沿其直觉向量传输查询持续改善精确输出检索,在 ARC-AGI-2 评估中达到 70.7。在来自 794 个任务的 397,000 个 ARC-GEN 实例中,单对直觉向量以约 87% 的留一法准确率识别生成任务。这些发现表明,在冻结视觉表示上的潜在向量算术支持跨不同问题领域的隐式规则推理,而无需生成模型组件,这表明推断抽象变换和生成其特定于实例的后果可能是可分离的能力。
cs.LG / 146 / 2609.32641

Extremely Fast and Compact Binary Graph Representations via Randomized Operator Sketching

基于随机算子草图的极速紧凑二值图表示
Agarwal, Srajan, P, Megha, Das, Bikas C, Laskar, Zakaria, Bej, Saptarshi
Abstract
Graph neural networks typically rely on dense, floating-point node representations, which can impose substantial memory and computational costs. Binary graph hashing offers an alternative by encoding node information as compact bit strings. However, existing approaches either sacrifice global topological information for computational efficiency or incur substantial generation costs. We introduce an ultra-fast, entirely algebraic hashing method that constructs binary node representations directly from graph structure, without requiring node features or gradient-based training. Our method approximates a high-order structural transition matrix using randomized column sampling inspired by the Nystr\"om method and combines it with an efficient label-safe semantic propagation mechanism. The resulting continuous representations are discretized through column-wise thresholding to obtain compact binary codes. Experiments on ten node classification datasets show that the proposed method consistently improves classification accuracy over existing feature-free binary baselines while requiring sub-second code generation on many datasets. The resulting binary representations are also naturally suited to event-driven computation, making them compatible with neuromorphic spiking neural networks and gradient-free learning rules. These results demonstrate that simple algebraic approximations can provide an efficient alternative to learned pipelines for discrete graph representation learning.
Chinese Translation
图神经网络通常依赖于稠密的浮点节点表示,这会造成巨大的内存和计算开销。二值图哈希通过将节点信息编码为紧凑的比特串提供了一种替代方案。然而,现有方法要么为了计算效率而牺牲全局拓扑信息,要么产生高昂的生成成本。我们提出了一种超快速、完全代数的哈希方法,它直接从图结构构造二值节点表示,无需节点特征或基于梯度的训练。我们的方法使用受 Nyström 方法启发的随机列采样来近似高阶结构转移矩阵,并将其与高效的标签安全语义传播机制相结合。得到的连续表示通过按列阈值化进行离散化,以获得紧凑的二值编码。在十个节点分类数据集上的实验表明,所提方法在现有无特征二值基线上持续提高分类准确率,同时在许多数据集上只需亚秒级的编码生成。得到的二值表示也天然适合事件驱动计算,使其与神经形态脉冲神经网络和无梯度学习规则兼容。这些结果表明,简单的代数近似可以为离散图表示学习提供一种高效的学习式流程替代方案。
cs.LG / 147 / 2609.32659

Quantization-Aware Pre-Training with Constrained Empirical Weight Distribution

基于约束经验权重分布的量化感知预训练
Yang, Ningfeng, Aamodt, Tor M.
Abstract
Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise, they either introduce additional hyperparameters or memory overhead, or cannot consistently improve model accuracy. In this work, we propose optimization with $\textbf{C}$onstrained $\textbf{E}$mpirical $\textbf{W}$eight dis$\textbf{T}$ribution (CEWT), the first hyperparameter-free memory-overhead-free oscillation suppression method that consistently improves QAPT performance: an optimizer post-update step that projects weights to the nearest point in weight space whose empirical distribution (histogram) matches a zero-mean Gaussian. Our key insight is many quantizers are designed with the implicit assumption that the to-be-quantized data are permutations of samples from a zero-mean Gaussian, and this assumption is not true during QAPT. By enforcing the zero-mean Gaussian prior as a hard constraint, CEWT can suppress this detrimental noise. Empirical results on various combinations of SOTA quantizers and hypersphere optimizers suggest, that with a geomean increase of 4% in training time, CEWT can consistently reduce the pre-training perplexity (by an average of 2.5 and up to 21 points) of low-precision (down to 1-bit activations and weights and up to 610M parameters) LLaMA/GPT models without introducing any hyperparameters or storage overhead. Code is available at https://github.com/1733116199/cewt
Chinese Translation
量化感知预训练(QAPT)可以提高深度神经网络的推理效率,但一种称为舍入边界权重振荡的问题行为会在训练过程中引入有害噪声,并显著降低收敛速度。虽然现有方法可以减少这种有害噪声,但它们要么引入额外的超参数或内存开销,要么无法持续提高模型精度。在这项工作中,我们提出了基于约束经验权重分布(CEWT)的优化,这是第一个无需超参数、无内存开销的振荡抑制方法,可持续提高QAPT性能:一个优化器更新后步骤,将权重投影到权重空间中的最近点,该点的经验分布(直方图)与零均值高斯分布匹配。我们的关键见解是,许多量化器的设计隐含假设待量化数据是来自零均值高斯的样本的排列,而这一假设在QAPT期间并不成立。通过将零均值高斯先验作为硬约束,CEWT可以抑制这种有害噪声。对SOTA量化器和超球面优化器的各种组合进行的实证结果表明,在训练时间几何平均增加4%的情况下,CEWT可以持续降低低精度(低至1比特激活和权重,高达610M参数)LLaMA/GPT模型的预训练困惑度(平均降低2.5,最高降低21个点),而不会引入任何超参数或存储开销。代码可在https://github.com/1733116199/cewt获取。
cs.LG / 148 / 2609.32661

Equivariant Neural Primal-Dual Assignment for Maximum Common Edge Subgraphs

最大公共边子图的等变神经原始-对偶分配
Xie, Jiaqing, Li, Yanchao, Yang, Zhuo, Wang, Yuxin, Fu, Tianfan, Li, Yuqiang
Abstract
Maximum common edge subgraph (MCES) matching finds a partial vertex correspondence between two labeled graphs that preserves as many labeled edges as possible. Molecular similarity search requires matching many graph pairs, making the cost of repeated queries important. The strongest baseline attains accurate MCES solutions but trains a separate network for each pair. We introduce Equivariant Neural Primal-Dual Assignment (ENPDA), which learns a shared matching policy and applies it to new pairs without further training, answering queries roughly three orders of magnitude faster and recovering its training cost after a few dozen queries. The policy recomputes exact objective marginals for candidate matches and learns corrections and step sizes that update their scores. Target prices respond to competition when several source vertices favor the same target. Four update rounds and a Hungarian projection produce a partial one-to-one matching. We prove per-pair guarantees that hold for any network parameters. In exact arithmetic, reordering either graph permutes the assignment and price states, the projected matching is one-to-one, and repaired prices give a valid MCES upper bound. Subtracting the preserved-edge count bounds the optimality gap; combined with structural caps, these certificates prove global optimality for 60 of 291 native test pairs. On three molecular benchmarks with disjoint train/validation/test splits, ENPDA improves over an analytic counterpart with the same update and projection budget by 7.4-8.6 accuracy points; after one second of refinement search, 2.5-3 points of the gain remain. Transferred without fine-tuning to edge-deletion tasks from social and protein graphs, the policy gains 9.1-17.6 points over the analytic counterpart. When output matchings must keep aromatic rings intact, ENPDA recovers more reference bonds than the baselines on all three datasets.
Chinese Translation
最大公共边子图(MCES)匹配寻找两个带标签图之间的部分顶点对应关系,使得尽可能多地保留带标签的边。分子相似性搜索需要匹配许多图对,因此重复查询的成本很重要。最强的基线能够获得准确的MCES解,但为每一对训练一个单独的网络。我们提出等变神经原始-对偶分配(ENPDA),它学习一个共享的匹配策略,并将其应用于新的图对而无需进一步训练,回答查询的速度大约快三个数量级,并在几十次查询后收回训练成本。该策略为候选匹配重新计算精确的目标边际,并学习更新其得分的修正量和步长。当多个源顶点倾向于同一目标时,目标价格会对竞争做出响应。四轮更新和匈牙利投影产生部分一对一匹配。我们证明了对于任意网络参数都成立的逐对保证。在精确算术中,重排任一图会置换分配和价格状态,投影匹配是一对一的,并且修复后的价格给出有效的MCES上界。减去保留边数限制了最优性差距;结合结构上限,这些证书证明了291个原生测试对中60个的全局最优性。在三个具有不相交训练/验证/测试划分的分子基准上,ENPDA在相同的更新和投影预算下比解析对应方法提高了7.4-8.6个准确率点;经过一秒的细化搜索后,增益仍保留2.5-3个点。在没有微调的情况下迁移到来自社交和蛋白质图的边删除任务,该策略比解析对应方法提高了9.1-17.6个点。当输出匹配必须保持芳香环完整时,ENPDA在所有三个数据集上比基线恢复了更多的参考键。
cs.LG / 149 / 2609.32663

Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference

Distance-KV:利用相对距离实现高效长上下文推理
Shang, Xianpeng, Huang, Canbin, Li, Jiang, Lan, Tian, Cai, Qianyi, Quan, Xiaojun, Su, Xiangdong
Abstract
The memory usage and decoding latency of LLM inference grow rapidly with context length. To reduce these costs, key-value (KV) cache compression methods selectively retain cached states based on token importance or differences in attention patterns across heads. However, we discover that retrieval capability varies substantially with relative distance, even within the same attention head. To exploit this structure, we introduce Distance-KV, which learns a static KV retention pattern over the joint space of layers, attention heads, and relative distances. The pattern is learned offline with the language model frozen and reused across inputs to prune and compact the KV cache without online importance scoring. Across three backbone models and four long-context benchmarks, Distance-KV consistently achieves the best overall performance among competing KV cache compression methods, exceeding the strongest compression baseline by up to 9.3 points on RULER at 128K. On Llama-3.1-8B-Instruct at 128K, Distance-KV reduces KV cache memory by 65.4% and achieves a $1.66\times$ decoding speedup relative to Dense. Together, these results identify relative distance as an important structural dimension for understanding how LLMs retrieve information over long contexts and for designing more efficient inference methods.
Chinese Translation
大语言模型(LLM)推理的内存占用和解码延迟会随着上下文长度迅速增长。为降低这些开销,键值(KV)缓存压缩方法基于词元重要性或不同注意力头之间注意力模式的差异,选择性地保留缓存状态。然而,我们发现检索能力会随相对距离发生显著变化,即使在同一个注意力头内部也是如此。为利用这一结构,我们提出 Distance-KV,它在层、注意力头和相对距离的联合空间上学习一种静态 KV 保留模式。该模式在冻结语言模型的情况下离线学习,并可跨输入复用,从而在无需在线重要性评分的情况下对 KV 缓存进行剪枝和压缩。在三个骨干模型和四个长上下文基准上,Distance-KV 在竞争性 KV 缓存压缩方法中始终取得最佳总体性能,在 128K 的 RULER 上比最强压缩基线高出最多 9.3 分。在 128K 的 Llama-3.1-8B-Instruct 上,Distance-KV 将 KV 缓存内存减少 65.4%,并相对于 Dense 实现 1.66 倍解码加速。这些结果共同表明,相对距离是一个重要的结构维度,既有助于理解 LLM 如何在长上下文中检索信息,也有助于设计更高效的推理方法。
cs.LG / 150 / 2609.32665

Timestep Weighting: A Hidden Key to Effective ELBO-Based Flow-Matching RL

时间步加权:基于ELBO的流匹配强化学习有效性的隐藏关键
Ma, Qinwei, Shi, Jingzhe, Fan, Simin, Li, Ling, Wang, Mengdi, Lamb, Alex
Abstract
ELBO-based reinforcement learning offers a sampler-agnostic approach to fine-tuning flow matching models with reward feedback. Timestep weighting in ELBO-based RL has large impact on performance, and it also provides a unified view (as we show in this work) to understand prediction losses heuristically chosen in prior work, yet it remains under-researched and is often chosen to inherit pretrain configs. We investigate impacts and dynamics of timestep weighting in ELBO-based RL. We show that effective weighting depends on both the reward landscape and stage of learning. (1) Through experiments on controlled CIFAR image generation, complemented by robotics, we investigate how weighting impacts reward-driven updates across noise levels. (2) Through gradient analysis, we reveal distinct patterns of cross-noise coordination across tasks and their evolution during training. These findings motivate the hypothesis that useful weighting depends on the gap between the policy's current behavior and the behavior favored by the reward. (3) Guided by this analysis, we study simple static weighting, budgeted profile selection, and dynamic schedules that improve performance beyond conventional target choices. Our results establish timestep weighting as an important design choice for flow-matching RL and motivate further research into methods that choose and adapt it throughout learning.
Chinese Translation
基于ELBO的强化学习提供了一种与采样器无关的方法,通过奖励反馈微调流匹配模型。基于ELBO的强化学习中的时间步加权对性能有很大影响,并且(如本工作所示)它提供了一个统一的视角来理解先前工作中启发式选择的预测损失,然而它仍然研究不足,且常被选择为继承预训练配置。我们研究了基于ELBO的强化学习中时间步加权的影响和动态。我们表明,有效的加权取决于奖励景观和学习阶段。(1) 通过在受控CIFAR图像生成上的实验,辅以机器人学,我们研究了加权如何影响跨噪声水平的奖励驱动更新。(2) 通过梯度分析,我们揭示了跨任务中跨噪声协调的不同模式及其在训练期间的演变。这些发现引出了一个假设:有用的加权取决于策略当前行为与奖励所偏好的行为之间的差距。(3) 在此分析的指导下,我们研究了简单的静态加权、预算化剖面选择和动态调度,它们将性能提升到超过传统目标选择。我们的结果确立了时间步加权作为流匹配强化学习的重要设计选择,并推动进一步研究在整个学习过程中选择并适应它的方法。
cs.LG / 151 / 2609.32672

STAMP: Predicting Out-of-Distribution Generalization without Target Data

STAMP:无需目标数据预测分布外泛化
Mahbub, Md Kawsher, Biswas, Milon
Abstract
Predicting whether a trained model will generalize under distribution shift remains difficult, especially when target-domain data are unavailable. We introduce STAMP (Semantic Temporal Augmented Model Prediction), a source-only, target-label-free criterion that estimates out-of-distribution (OOD) performance from paired source-domain images. STAMP computes the output-space correlation ratio $\eta^2=S_B/S_T$ by contrasting semantically stable pairs with random pairs: higher $\eta^2$ indicates that model outputs vary with semantic identity rather than nuisance variation. On 44 chest X-ray models spanning CNNs, ViTs, MetaFormers, foundation models, and SSL/VLM probes, temporal STAMP attains Spearman correlations of $0.844$--$0.855$ with macro AUROC on VinDr-CXR, CheXpert, and MIMIC-CXR; a class-matched variant improves single-class RSNA from $0.311$ to $0.663$. STAMP attains the best average source-only medical ranking and outperforms the target-domain ATC and AoTL estimators without any target data. On 27 ImageNet models, temperature-scaled STAMPTS attains $\rho{=}0.984$ on ObjectNet and $\rho\geq0.905$ on four additional distribution shifts, with partial correlations of $0.662$--$0.949$ after controlling for ImageNet accuracy. Requiring approximately 12s per model on one GPU, STAMP is a practical pre-deployment model-selection and auditing tool.
Chinese Translation
预测一个训练好的模型在分布偏移下能否泛化仍然很困难,尤其是当目标域数据不可用时。我们提出 STAMP(语义时序增强模型预测),一种仅源域、无需目标标签的准则,它从成对的源域图像估计分布外(OOD)性能。STAMP 通过对比语义稳定的图像对与随机图像对,计算输出空间相关比 $\eta^2=S_B/S_T$:较高的 $\eta^2$ 表明模型输出随语义身份而非无关变化而变化。在涵盖 CNN、ViT、MetaFormer、基础模型和 SSL/VLM 探针的 44 个胸部 X 光模型上,时序 STAMP 在 VinDr-CXR、CheXpert 和 MIMIC-CXR 上与宏 AUROC 达到 $0.844$--$0.855$ 的 Spearman 相关性;一个类别匹配的变体将单类 RSNA 从 $0.311$ 提升到 $0.663$。STAMP 在仅源域医学排名中取得最佳平均排名,并且在没有任何目标数据的情况下优于目标域 ATC 和 AoTL 估计器。在 27 个 ImageNet 模型上,温度缩放的 STAMPTS 在 ObjectNet 上达到 $\rho{=}0.984$,在另外四个分布偏移上达到 $\rho\geq0.905$,在控制 ImageNet 准确率后的偏相关性为 $0.662$--$0.949$。STAMP 在单块 GPU 上每个模型约需 12 秒,是一种实用的部署前模型选择与审计工具。
cs.LG / 152 / 2609.32676

SIFT: Enhancing Time Series Foundation Models via Semantic Invariance and Structural Fidelity Fine-Tuning

SIFT:通过语义不变性和结构保真度微调增强时间序列基础模型
Tang, Yi, Zhang, Tengxue, Shu, Yang, Guo, Chenjuan, Sun, Chenchen, An, Yisheng
Abstract
Time Series Foundation Models (TSFMs) have achieved remarkable zero-shot performance through extensive pre-training on massive time series datasets. Nevertheless, due to the low-dimensional properties and diverse structural patterns of time series data, performing naive fine-tuning on TSFMs often leads to overfitting and falling into the mean-prediction trap. To address these challenges, we propose SIFT, a robust adaptation method that enhances time series foundation models by preserving Semantic Invariance and structural Fidelity throughout the fine-Tuning process. We employ semantic-invariant adversarial augmentation, which utilizes semantic spectrum decomposition to partition the semantic space and then generates perturbations within the non-core semantic subspace to bolster the model's robustness against these perturbations, mitigating overfitting. We implement a component-based structural fidelity enhancement, which facilitates component-wise mixup and imposes a reconstruction objective to improve the model's ability to preserve structural fidelity, alleviating the mean-prediction trap. Extensive experiments on representative TSFMs covering 10 real-world datasets demonstrate that SIFT can significantly enhance the performance of TSFMs.
Chinese Translation
时间序列基础模型(TSFMs)通过在海量时间序列数据集上进行广泛预训练,取得了显著的零样本性能。然而,由于时间序列数据的低维特性和多样的结构模式,对TSFMs进行朴素微调往往会导致过拟合并陷入均值预测陷阱。为了解决这些挑战,我们提出了SIFT,一种鲁棒的适应方法,通过在微调过程中保持语义不变性和结构保真度来增强时间序列基础模型。我们采用语义不变对抗增强,利用语义谱分解来划分语义空间,然后在非核心语义子空间内生成扰动,以增强模型对这些扰动的鲁棒性,从而缓解过拟合。我们实现了基于组件的结构保真度增强,促进组件级mixup并施加重构目标,以提高模型保持结构保真度的能力,缓解均值预测陷阱。在涵盖10个真实世界数据集上的代表性TSFMs上进行的广泛实验表明,SIFT可以显著提升TSFMs的性能。
cs.LG / 153 / 2609.32679

The GUI Is Not the State: Diagnosing State Aliasing in GUI World Models

GUI 并非状态:诊断 GUI 世界模型中的状态混叠
Liu, Dongsheng, Jin, Chao, Yang, Wenkui, Wang, Hejin, Yang, Junwei, Zhang, Zeren, Chen, Ziwei, Huang, Huaibo, Cao, Jie, He, Ran
Abstract
GUI World Models (GUI-WMs) are increasingly used to predict future states for agent planning and simulation, yet most existing formulations condition only on the current GUI observation and action. We identify state aliasing, where the vis- ible interface omits transition-relevant environment state, so identical observable conditions can correspond to different valid futures. To diagnose this failure mode, we introduce StateAliasBench, a diagnostic benchmark that explicitly isolates such ambiguities via strict pairing. We further propose lightweight predictive- state recovery that infers structured state from history and augments otherwise frozen GUI-WMs through a deterministic state interface. Family-specific special- ists provide state recovery across heterogeneous state types, and multi-teacher dis- tillation consolidates them into a single unified estimator. Experiments show that existing GUI-WMs exhibit systematic failures under observation-only condition- ing, while predictive-state augmentation substantially restores state-sensitive pre- diction across evaluated WMs, preserves generative fidelity, and improves down- stream performance of GUI agents on AndroidWorld. These results suggest that reliable GUI world modeling should account not only for what is visible, but also for the hidden transition state that determines what happens next.
Chinese Translation
GUI 世界模型(GUI-WMs)正越来越多地被用于为智能体规划与模拟预测未来状态,然而大多数现有形式仅以当前 GUI 观测和动作为条件。我们识别出状态混叠(state aliasing):即可见界面遗漏了与状态转移相关的环境状态,因此相同的可观测条件可能对应不同的有效未来。为了诊断这一失效模式,我们提出 StateAliasBench,一个通过严格配对显式隔离此类歧义的诊断基准。我们进一步提出轻量级预测状态恢复方法,从历史中推断结构化状态,并通过确定性状态接口增强原本冻结的 GUI-WMs。面向不同状态家族的专家模型可提供状态恢复,而多教师蒸馏将它们整合为单一统一估计器。实验表明,现有 GUI-WMs 在仅观测条件下会出现系统性失败,而预测状态增强可在所评估的 WM 上大幅恢复状态敏感预测,保持生成保真度,并提升 GUI 智能体在 AndroidWorld 上的下游性能。这些结果表明,可靠的 GUI 世界建模不仅应考虑可见内容,还应考虑决定下一步发生什么的隐藏转移状态。
cs.LG / 154 / 2609.32689

Self-Evolving Time-Series Forecasting Agents with Episodic Memory and Online Policy Learning

具有情节记忆与在线策略学习的自演化时间序列预测智能体
Wang, Junyi, Wang, Yilin, Wu, Wen, Zhang, Chao
Abstract
LLM-based agents are increasingly used for time-series forecasting because they can organise contextual information, perform multi-step analysis, and guide the sequence of actions required to complete forecasting tasks. Most existing agents focus only on the current forecasting instance. However, in real-world deployments, forecasting commonly operates online, with new forecasts issued from the currently available history as the forecast origin advances and the ground-truth targets of earlier instances progressively become available. These targets provide feedback on the actions taken in earlier instances, yet existing agents generally do not preserve or utilise this information to adapt their subsequent actions. To address this limitation, we introduce FASE, a Feedback-Aware Self-Evolving forecasting agent that converts such feedback into task-specific experience for subsequent forecasting instances. FASE combines episodic memory, which retrieves relevant completed instances, with online policy learning, which summarises the feedback accumulated across instances into ranking guidance. The proposed framework is evaluated on 29 dataset configurations selected from the GIFT-Eval benchmark. Across these 29 configurations, FASE attains the strongest aggregate point forecasting performance among the evaluated methods and reduces the normalised MAE by 9.1% relative to the best individual foundation model baseline. The results further indicate that the cumulative advantage of FASE increases as delayed feedback accumulates. Together, these findings demonstrate that FASE can continually self-evolve through feedback from completed forecasting instances without updating the parameters of the LLM.
Chinese Translation
基于LLM的智能体越来越多地用于时间序列预测,因为它们能够组织上下文信息、执行多步分析,并引导完成预测任务所需的动作序列。大多数现有智能体只关注当前预测实例。然而,在实际部署中,预测通常以在线方式运行:随着预测起点推进,并根据当前可用的历史数据发布新的预测,而较早实例的真实目标也逐渐可用。这些目标为较早实例中采取的动作提供了反馈,但现有智能体通常不会保存或利用这些信息来调整其后续动作。为解决这一局限,我们提出FASE,一种反馈感知的自演化预测智能体,它将此类反馈转化为后续预测实例的任务特定经验。FASE结合了情景记忆(用于检索相关的已完成实例)与在线策略学习(用于将跨实例累积的反馈总结为排序指导)。所提框架在从GIFT-Eval基准中选出的29个数据集配置上进行了评估。在这29个配置上,FASE在所评估方法中取得了最强的总体点预测性能,并且相对于最佳单个基础模型基线,将归一化MAE降低了9.1%。结果进一步表明,随着延迟反馈的累积,FASE的累积优势会增加。总之,这些发现表明,FASE可以通过来自已完成预测实例的反馈持续自演化,而无需更新LLM的参数。
cs.LG / 155 / 2609.32720

BiasReducer: Adaptive Bias Mitigation for Reward Models

BiasReducer:面向奖励模型的自适应偏差缓解
Liu, Shuang, Miao, Yongliang, Liu, Yanguang, Xiong, Haoyi, Du, Mengnan
Abstract
Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods either retrain the reward model or apply a fixed correction to one known bias, such as a preference for longer responses. Retraining requires additional data and computational resources, while existing editing methods require the target bias to be specified in advance and use a fixed edit for that bias. To this end, we propose BiasReducer, a lightweight framework that edits only the linear reward head and selects the relevant edits for each new dataset. First, BiasReducer uses a sparse autoencoder (SAE)-style encoder to learn which attributes (e.g., length and confidence) the reward model is sensitive to. Second, it learns how to reduce the reward model's dependence on each attribute by determining which direction to adjust the reward head and how much to adjust it. Third, for a new dataset, it ranks the attributes by their influence on reward scores, selects the relevant ones, and edits the reward model accordingly. BiasReducer consistently improves reward-model robustness to biases toward superficial response attributes. Across five reward models, BiasReducer-M improves the three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming the two training-based baselines. The gains transfer downstream, reducing unnecessary verbosity and sycophancy while maintaining comparable judged quality.
Chinese Translation
奖励模型对来自大语言模型(LLM)的回复进行评分,并引导 LLM 的训练朝向人类偏好。然而,奖励模型可能偏好诸如长度或置信度等表面属性,导致 LLM 产生得分更高但并非更正确的回复。现有的缓解方法要么重新训练奖励模型,要么对某一已知偏差(例如对更长回复的偏好)应用固定校正。重新训练需要额外数据和计算资源,而现有编辑方法需要预先指定目标偏差,并对该偏差使用固定编辑。为此,我们提出 BiasReducer,一个轻量级框架,它仅编辑线性奖励头,并为每个新数据集选择相关编辑。首先,BiasReducer 使用稀疏自编码器(SAE)风格的编码器来学习奖励模型对哪些属性(例如长度和置信度)敏感。其次,它通过学习调整奖励头的方向以及调整幅度,来减少奖励模型对每个属性的依赖。第三,对于新数据集,它按属性对奖励分数的影响进行排序,选择相关属性,并相应地编辑奖励模型。BiasReducer 持续提高奖励模型对偏向表面回复属性的鲁棒性。在五个奖励模型上,BiasReducer-M 将三个基准平均提高了 8.3、18.0 和 6.9 个百分点,优于两个基于训练的基线。这些增益可迁移到下游,减少不必要的冗长和谄媚,同时保持相当的评判质量。
cs.LG / 156 / 2609.32722

Scaling Properties of Same-Family On-Policy Distillation

同家族同策略蒸馏的缩放特性
Bao, Yuntai, Li, Qinfeng, Jiang, Guoqing, Chen, Liwei, Qin, Zhiheng, Li, Xuanping, Zhang, Wenqi, Zhang, Xuhong
Abstract
*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how $G_{\mathrm{peak}}$ and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
Chinese Translation
强化学习(RL)能够在大语言模型(LLMs)中激发出大量推理能力,但这种能力在不同模型规模之间的迁移程度以及迁移速度仍不清楚。我们研究了同策略蒸馏(OPD)在弱到强、同基座和强到弱师生设置中的缩放特性。我们发现,早期OPD训练动态一致地表现出一种规则的有效迁移区间,其中留出准确率(黄金分数,G)随 d = √(KL(π_θ || π_ref)) 近似线性上升,d 是相对于学生初始化的词元级反向KL散度的平方根。在每一个观察到的弱到强配对中,学生的峰值黄金分数都超过了其教师自身的峰值,因此一个紧凑的RL专家可以通过OPD将能力迁移到一个大得多的学生模型。为了估计OPD的结果,我们拟合了幂律,描述 G_peak 和有效迁移区间的斜率如何随学生和教师的参数量以及教师黄金分数而变化。这些规律表明,峰值黄金分数仅随教师规模提升到大约学生规模的水平,并且在匹配的黄金分数下,较小的教师迁移效果更好,因此仅凭教师的分数并不能定义其监督价值。我们还研究了两种OPD变体的缩放效应:自举弱到强OPD,以及同策略监督的程度。
cs.LG / 157 / 2609.32727

Distributionally Robust Average-Reward Reinforcement Learning: Finite-Sample Guarantees under Weak Communication

分布鲁棒平均奖励强化学习:弱通信下的有限样本保证
Lu, Chenyu, Chen, Zijun, Si, Nian
Abstract
We study distributionally robust reinforcement learning (DR-RL) in the average-reward setting under weak communication. Our main result provides finite-sample guarantees for estimating the robust optimal average reward and learning a near-optimal policy, covering both SA-rectangular and S-rectangular structures with divergence-based and distance-based uncertainty sets. Specifically, for Kullback--Leibler and $f_k$-divergence balls, we establish explicit radius conditions under which the robust average-reward Bellman equation admits a constant-gain solution, while for total variation and Wasserstein balls, any positive radius suffices without requiring the nominal MDP to be weakly communicating. Our algorithm is prior-knowledge-free and achieves sample complexities of $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-1}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$ for estimating the robust optimal average reward and $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-2}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$ for learning an $\epsilon$-optimal policy. Here, $p_{\wedge}$ is the smallest positive nominal transition probability and $u_{\delta}^{\ast}$ is a robust optimal bias function. We further provide an almost-tight explicit upper bound on $\operatorname{Span}(u_{\delta}^{\ast})$. Finally, we validate the predicted $n^{-1/2}$ convergence rate through numerical experiments.
Chinese Translation
我们研究弱通信下平均奖励设定中的分布鲁棒强化学习(DR-RL)。我们的主要结果为估计鲁棒最优平均奖励和学习近最优策略提供了有限样本保证,涵盖 SA-rectangular 和 S-rectangular 结构以及基于散度和基于距离的不确定性集合。具体地,对于 Kullback-Leibler 和 $f_k$ 散度球,我们建立了显式的半径条件,在这些条件下鲁棒平均奖励 Bellman 方程允许常数增益解;而对于全变差和 Wasserstein 球,任何正半径都足够,无需名义 MDP 是弱通信的。我们的算法无需先验知识,对于估计鲁棒最优平均奖励,达到样本复杂度 $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-1}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$;对于学习 $\epsilon$-最优策略,达到 $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-2}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$。这里,$p_{\wedge}$ 是最小正名义转移概率,$u_{\delta}^{\ast}$ 是鲁棒最优偏差函数。我们进一步提供了 $\operatorname{Span}(u_{\delta}^{\ast})$ 的一个几乎紧的显式上界。最后,我们通过数值实验验证了预测的 $n^{-1/2}$ 收敛速率。
cs.LG / 158 / 2609.32741

Predicting the Financial Impact of Supply Chain Risk for Major AI-Related Semiconductor Firms: A Heterogeneous Graph Patch Transformer Approach

预测主要AI相关半导体公司供应链风险的财务影响:一种异构图分块Transformer方法
Hur, Jianna, Samtani, Sagar
Abstract
Modern semiconductor production relies on a globally distributed, multi-tier supply chain in which financial stress at one firm spreads with a delay and eventually affects the revenue, inventory, and profitability of the companies that design AI chips. Most firms see only their direct partners, and prior predictive research has mainly targeted market-based risk measures, so few tools forecast how supply chain stress will appear in reported financials. In this study, we propose a heterogeneous graph patch transformer that forecasts these quarterly changes one and two quarters ahead. Learning from a 15,186-company network over 60 quarters, the proposed model fuses quarterly fundamentals with macro-trade, event, and disaster signals through learned gates, carries risk across supplier, customer, ownership, and headquarters relations through typed, direction-specific propagation, and encodes the propagated histories with patch-based tokenization. In preliminary experiments on 116 focal semiconductor firms, the proposed model achieves the lowest error on every target at both horizons, and its profitability advantage widens at the two-quarter horizon. These forecasts can help supply chain managers and investors act before disruptions appear in reported financials.
Chinese Translation
现代半导体生产依赖于全球分布的多层级供应链,其中一家公司的财务压力会延迟扩散,并最终影响设计AI芯片的公司的收入、库存和盈利能力。大多数公司只能看到其直接合作伙伴,而此前的预测研究主要针对基于市场的风险指标,因此很少有工具能够预测供应链压力将如何体现在已报告的财务数据中。在本研究中,我们提出了一种异构图分块Transformer(Heterogeneous Graph Patch Transformer),用于提前一个和两个季度预测这些季度变化。基于一个包含15,186家公司、跨越60个季度的网络进行学习,所提模型通过可学习门控将季度基本面与宏观贸易、事件和灾害信号融合,通过类型化、方向特定的传播在供应商、客户、股权和总部关系中传递风险,并使用基于分块的标记化对传播历史进行编码。在针对116家重点半导体公司的初步实验中,所提模型在两个预测期上对所有目标均取得最低误差,并且其盈利能力优势在两个季度预测期上进一步扩大。这些预测可以帮助供应链管理者和投资者在中断出现在已报告的财务数据之前采取行动。
cs.LG / 159 / 2609.32743

Benchmarking EEG Foundation Models at Scale: Lessons from 20,000 Evaluations

大规模基准测试 EEG 基础模型:来自 20,000 次评估的启示
Chen, Zhige, Peng, Shu, Qin, Chengxuan, Liu, Rui, Yang, Rui, Tan, Kay Chen, Wu, Jibin
Abstract
Electroencephalography (EEG) foundation models (FMs) promise transferable neural representations, yet their advantages over strong supervised baselines and their prospects for further scaling remain unclear. To address these questions, we introduce EEG-Arena, an open-source benchmark covering 30 EEG FMs and 25 supervised baselines evaluated on 57 downstream tasks from 23 public datasets. Through more than 20,000 evaluations across five experimental protocols, we assess downstream performance, pretraining benefits, model size scaling, pretraining data scaling, and robustness to channel configuration. We find that (1) EEG FMs outperform strong task-specific supervised baselines on most evaluated tasks, particularly under non-bipolar settings; (2) compared with architecture-matched supervised training from scratch, pretraining improves both early optimization and final downstream performance, with larger and more consistent gains as more labeled downstream data become available; (3) existing EEG FMs do not exhibit a consistent positive relationship between parameter count and downstream performance; (4) under a fixed architecture, increasing the pretraining data scale yields sustained downstream gains; and (5) channel-flexible FMs achieve higher absolute performance than channel-constrained models across most evaluated channel configurations. Together, these findings demonstrate the downstream value of EEG FMs and identify pretraining data expansion as a promising direction for further progress. To support continued research, we release EEG-Arena as an open-source evaluation framework that provides shared infrastructure for reproducible benchmarking, model comparison, and community-driven development.
Chinese Translation
脑电图(EEG)基础模型(FMs)有望提供可迁移的神经表征,然而其相较于强监督基线的优势以及进一步扩展的前景仍不明确。为解决这些问题,我们提出了 EEG-Arena,一个开源基准,涵盖 30 个 EEG 基础模型和 25 个监督基线,并在来自 23 个公开数据集的 57 个下游任务上进行评估。通过跨越五种实验协议的 20,000 多次评估,我们评估了下游性能、预训练收益、模型规模扩展、预训练数据扩展以及对通道配置的鲁棒性。我们发现:(1)在大多数评估任务上,EEG 基础模型优于强任务特定监督基线,尤其是在非双极设置下;(2)与架构匹配的从头监督训练相比,预训练改善了早期优化和最终下游性能,并且随着更多有标签的下游数据可用,增益更大且更一致;(3)现有的 EEG 基础模型在参数数量和下游性能之间没有表现出一致的正相关关系;(4)在固定架构下,增加预训练数据规模可带来持续的下游性能提升;以及(5)在大多数评估的通道配置下,通道灵活的 FMs 比通道受限模型实现了更高的绝对性能。总之,这些发现表明 EEG 基础模型具有下游价值,并指出预训练数据扩展是进一步进展的有前景方向。为支持持续研究,我们发布了 EEG-Arena 作为开源评估框架,为可重复基准测试、模型比较和社区驱动开发提供共享基础设施。
cs.LG / 160 / 2609.32756

Reuse or Relearn? A Spectral View of Earth Observation Foundation Models

复用还是重新学习?地球观测基础模型的谱视角
Turkoglu, Mehmet Ozgur, Marsocci, Valerio, Mühlematter, Dominik J., Senti, Dominik, Schindler, Konrad, Aasen, Helge
Abstract
Foundation models are rarely used as generic, frozen feature extractors; instead, they are fine-tuned for the target downstream application. This practice is particularly prevalent in Earth observation (EO), and it raises a question that downstream accuracy alone cannot answer: does fine-tuning reuse the pretrained representation, or does it relearn a new one? We study this with spectral diagnostics that compare a model before and after adaptation, quantifying how well its dominant singular subspaces are preserved, how broadly the weight update is distributed, and how large it is. Using natural image models such as CLIP and DINO as a reference, we find that, under the evaluated fine-tuning settings, EO models undergo far larger, higher-rank updates and retain much less of their pretrained structure, so their downstream performance is often obtained with substantial changes to the pretrained weight structure. The diagnostics further provide insight into how cheaply a model can be adapted: where the pretrained subspaces are preserved, adapting a small fraction of the parameters can match full fine-tuning, and where they are not, it can fall behind. More broadly, foundation models, and EO foundation models in particular, should be assessed not only by benchmark accuracy, but also by how reusable their pretrained representation is.
Chinese Translation
基础模型很少被用作通用的、冻结的特征提取器;相反,它们会被微调以适应目标下游应用。这种做法在地球观测(EO)中尤为普遍,它提出了一个仅靠下游准确率无法回答的问题:微调是在复用预训练表示,还是在重新学习一个新的表示?我们用谱诊断来研究这个问题,通过比较模型在适应前后的状态,量化其主导奇异子空间的保留程度、权重更新的分布广度以及更新幅度。以CLIP和DINO等自然图像模型为参照,我们发现在所评估的微调设置下,EO模型经历更大、更高秩的更新,并且保留的预训练结构要少得多,因此它们的下游性能往往是通过对预训练权重结构进行实质性改变而获得的。这些诊断进一步揭示了模型可以以多低的成本进行适应:在预训练子空间得以保留的地方,仅调整一小部分参数就能达到完全微调的效果;而在没有保留的地方,则可能落后。更广泛地说,基础模型,尤其是EO基础模型,不仅应通过基准准确率来评估,还应通过其预训练表示的可重用性来评估。
cs.LG / 161 / 2609.32759

The Extender: A Log-Structured Transformer

Extender:一种日志结构化的 Transformer
Eriksson, Jakob
Abstract
We introduce the Extender, a log-structured variant of the standard Transformer architecture. In a standard Transformer, each layer communicates with subsequent layers exclusively via the residual $\mathbf{h}$, a superposition channel. The Extender adds a concatenation channel $\mathbf{x}$: each layer $\ell$ emits both a residual update $\delta_\ell$ which is added to $\mathbf{h}$, and a much smaller extension $\epsilon_\ell$ which is appended to $\mathbf{x}$. While both the FFN and $\mathbf{q}$ see $\mathbf{h}$, the attention $\mathbf{kv}$ projections take only $\mathbf{x}$ as input. As a result, the fully extended $\mathbf{x}$ contains the complete input for the $\mathbf{kv}$ projections of all layers, reducing the persistent attention memory footprint from $2Ld_{model}$ to $\sum|\epsilon_\ell|$. We find that with $|\epsilon_\ell|=32$, the Extender matches Transformer accuracy on short-context (CORE) tasks at 199M-924M parameters, and exceeds Transformer accuracy on long-context (RULER) workloads, again at 924M parameters. For our 1664-wide, 924M model, the Extender's persistent attention memory footprint is $104\times$ smaller than MHA. The memory savings grow with model width.
Chinese Translation
我们介绍了 Extender,一种标准 Transformer 架构的日志结构化变体。在标准 Transformer 中,每一层仅通过残差 $\mathbf{h}$(一个叠加通道)与后续层通信。Extender 增加了一个拼接通道 $\mathbf{x}$:每一层 $\ell$ 既发出一个残差更新 $\delta_\ell$(被加到 $\mathbf{h}$ 上),也发出一个小得多的扩展 $\epsilon_\ell$(被追加到 $\mathbf{x}$ 上)。虽然 FFN 和 $\mathbf{q}$ 都看到 $\mathbf{h}$,但注意力 $\mathbf{kv}$ 投影仅以 $\mathbf{x}$ 作为输入。因此,完全扩展的 $\mathbf{x}$ 包含了所有层的 $\mathbf{kv}$ 投影的完整输入,将持久注意力内存占用从 $2Ld_{model}$ 减少到 $\sum|\epsilon_\ell|$。我们发现,当 $|\epsilon_\ell|=32$ 时,Extender 在 199M-924M 参数下在短上下文(CORE)任务上达到了与 Transformer 相当的准确率,并且在长上下文(RULER)工作负载上(同样在 924M 参数下)超过了 Transformer 的准确率。对于我们的 1664 宽、924M 参数的模型,Extender 的持久注意力内存占用比 MHA 小 $104\times$。内存节省随着模型宽度的增加而增加。
cs.LG / 162 / 2609.32764

AnchorMixGAN: Anchor-Aligned Generative Semi-Supervision for DDoS Detection in Cloud-Integrated IoT Networks

AnchorMixGAN:面向云集成物联网网络中DDoS检测的锚点对齐生成式半监督
Yang, Jin, Liu, Xufeng, Hu, Yong, Wang, Xueyang, Yang, Honglu, Li, Gang
Abstract
Detecting distributed denial-of-service (DDoS) attacks in cloud-integrated IoT networks is difficult when labeled traffic is scarce. Generative semi-supervised learning can supplement the available training data, but prediction shifts induced by synthetic views may affect the targets assigned to real unlabeled flows. We propose AnchorMixGAN, a generative semi-supervised framework that addresses this problem through anchor-aligned target construction. Its Anchor-MAS module treats each real unlabeled flow as an anchor and creates alternative views by replacing one field group at a time with values from generated traffic. A frozen reference classifier predicts the anchor and its views; averaging and sharpening these predictions produces a soft target for the original flow. The flow and its target are then mixed with a labeled example using MixUp, allowing the detector to learn from both the original labeled records and the mixed examples. We analyze how reference-classifier error, view construction, and sharpening affect the target, and derive a bound on the resulting change in cross-entropy at a fixed detector prediction. At the reported 90% training setting with 20% of the training records labeled, AnchorMixGAN attains accuracies of 97.3%, 97.4%, and 96.5% on NSLKDD, BoT-IoT, and CICIoT2023, respectively, exceeding the corresponding MixGAN results by 1.6, 1.0, and 4.4 percentage points.
Chinese Translation
在云集成物联网网络中,当标记流量稀缺时,检测分布式拒绝服务(DDoS)攻击十分困难。生成式半监督学习可以补充可用的训练数据,但由合成视图引起的预测偏移可能会影响分配给真实未标记流的目标。我们提出AnchorMixGAN,一个通过锚点对齐目标构建来解决此问题的生成式半监督框架。其Anchor-MAS模块将每个真实未标记流视为锚点,并通过每次用生成流量的值替换一个字段组来创建替代视图。一个冻结的参考分类器预测锚点及其视图;对这些预测进行平均和锐化,为原始流生成软目标。然后,使用MixUp将该流及其目标与一个标记示例混合,使检测器能够从原始标记记录和混合示例中学习。我们分析了参考分类器误差、视图构建和锐化如何影响目标,并推导了在固定检测器预测下交叉熵变化的界限。在报告的90%训练设置(其中20%的训练记录有标记)下,AnchorMixGAN在NSLKDD、BoT-IoT和CICIoT2023上分别达到97.3%、97.4%和96.5%的准确率,比相应的MixGAN结果分别高出1.6、1.0和4.4个百分点。
cs.LG / 163 / 2609.32766

Structuring Relations Among Learning Paradigms via Protocol--Objective--Resource Reductions

通过协议-目标-资源归约构建学习范式之间的关系
Su, Junwei, Wang, Changjie, Chang, Dongyang
Abstract
Modern machine learning spans supervised, transfer, continual, meta-learning, and related regimes that often reuse the same hypothesis classes, architectures, and optimizers but differ in information access, objectives, memory, adaptation, and sample accounting. This makes it difficult to determine whether one paradigm is genuinely distinct, a special case of another, or part of a broader structural hierarchy. We introduce a protocol-objective-resource (POR) framework that separates representational capacity from these design choices. A paradigm is specified by an environment class, observation protocol, admissible learners, performance functional, and resource accounting rule. POR reductions combine environment embeddings, learner compilers, threshold maps, and calibrated resource overheads. Our main theorem shows that such reductions imply worst-case complexity domination on embedded comparison classes, transferring upper bounds forward and lower bounds backward; under labeled-example accounting, this yields sample complexity domination. We also show that calibrated nontrivial accuracy regimes are necessary to avoid vacuous comparisons, and that strengthening the objective can strictly increase minimax sample complexity even with unchanged protocols and learner classes. Instantiating the framework for supervised, transfer, continual, and meta-learning yields canonical special-case relations: continual contains transfer, transfer contains supervised, and meta-learning contains supervised under aligned raw-example accounting. We further derive a non-exact episode-to-example reduction for episodic meta-learning and capture within-paradigm refinements such as replay memory and task identifiers. The framework thus provides a unified language for structuring learning paradigms and transferring complexity guarantees across them.
Chinese Translation
现代机器学习涵盖监督学习、迁移学习、持续学习、元学习及相关范式,这些范式通常重用相同的假设类、架构和优化器,但在信息访问、目标、记忆、适应和样本核算方面有所不同。这使得难以确定一个范式是真正不同的、另一个的特例,还是更广泛的结构层次的一部分。我们引入协议-目标-资源(POR)框架,将表示能力与这些设计选择分离。一个范式由环境类、观测协议、允许的学习器、性能泛函和资源核算规则来指定。POR归约结合了环境嵌入、学习器编译器、阈值映射和校准的资源开销。我们的主要定理表明,此类归约蕴含在嵌入比较类上的最坏情况复杂度支配,将上界向前传递、下界向后传递;在带标签样本核算下,这产生样本复杂度支配。我们还表明,校准的非平凡精度机制是避免空洞比较所必需的,并且即使协议和学习器类不变,加强目标也可能严格增加极小极大样本复杂度。将该框架实例化到监督、迁移、持续和元学习,可得到典型的特例关系:持续包含迁移,迁移包含监督,并且在一致的原始样本核算下,元学习包含监督。我们进一步为情景元学习推导了一个非精确的回合到样本归约,并刻画了范式内的细化,如回放记忆和任务标识符。因此,该框架为构建学习范式并在它们之间传递复杂度保证提供了统一语言。
cs.LG / 164 / 2609.32771

Continual Learning via Self-Probe Gradients

通过自探测梯度的持续学习
Cho, Dongkyu, Chunara, Rumi, Cha, Sungmin
Abstract
Adapting pretrained models to new data can cause catastrophic forgetting of previously learned behavior. When only a few past samples remain, they give continual learning methods sparse and narrow evidence about what to preserve. We show that language models can expand this evidence through self-probing, in which the frozen model generates new inputs from the retained samples and records its own predictions on them. Unlike prior work that replays such data as training examples, our method, CPLUS uses self-probe and past-sample gradients to scale down parameter updates that conflict with prior behavior. Experiments with five language models on four benchmarks show three results. First, the same probes preserve more prior behavior as gradient signals than as replay data. Second, CPLUS learns the new data while consistently reducing forgetting more than existing baselines, especially when past data are scarce, and this protection extends to benchmarks not used for training. Third, we observe that CPLUS also becomes more effective as models grow: within the Qwen3 model family, it recovers an increasing share of the forgetting caused by standard fine-tuning.
Chinese Translation
将预训练模型适配到新数据可能导致对先前学习行为的灾难性遗忘。当仅保留少量过去样本时,它们为持续学习方法提供了关于需要保留内容的稀疏且狭窄的证据。我们表明,语言模型可以通过自探测来扩展这一证据,其中冻结的模型从保留样本中生成新输入,并记录其对这些输入的预测。与先前将这些数据作为训练样本重放的工作不同,我们的方法 CPLUS 使用自探测和过去样本梯度来缩减与先前行为冲突的参数更新。在四个基准上对五个语言模型的实验显示了三个结果。首先,相同的探测作为梯度信号比作为重放数据能保留更多先前行为。其次,CPLUS 在学习新数据的同时,比现有基线持续减少更多遗忘,尤其是在过去数据稀缺时,并且这种保护扩展到未用于训练的基准。第三,我们观察到,随着模型规模增长,CPLUS 也变得更有效:在 Qwen3 模型系列中,它恢复了由标准微调引起的遗忘的越来越大的份额。
cs.LG / 165 / 2609.32774

Beyond Gaussian Assumptions: Distribution-Aware Channel Capacity for Effective Connectivity

超越高斯假设:面向有效连接性的分布感知信道容量
Jian, Jianan, Kang, Jacob, Multezem, Nurahmed, Li, Benjamin, Xu, Nan
Abstract
Effective-connectivity estimation from brain signals often relies on Gaussian residual modeling, which enables tractable estimation but can discard informative distributional structure and distort inferred directed interactions when empirical residuals are non-Gaussian. We show across multiple modalities, species, and experimental conditions that both brain signals and fitted channel residuals frequently deviate from Gaussianity. We therefore introduce a distribution-aware, information-theoretic measure of effective connectivity based on channel capacity under general residual distributions. To estimate the resulting capacity from empirical, potentially non-Gaussian residuals, we develop a dual-flow min-max estimator based on normalizing flows, in which a generator searches over admissible input distributions under a power constraint while an observer estimates output entropy. We provide a theoretical characterization of the estimator, showing that the observer objective recovers differential entropy up to a KL approximation term, that the formulation reduces to classical Gaussian capacity as a special case, and that residual entropy can alter achievable information rates beyond variance; game-gap and error analyses further characterize optimization and approximation sources. In brain-like simulations with known directed connectivity, Dual-flow achieves the highest AUROC and AUPRC across ten conditions spanning diverse network topologies, hidden drivers, feedback, and heterogeneous hemodynamics, compared with Gaussian capacity, Granger causality, VAR-LiNGAM, and GIMME. Applied to multimodal brain signals, the method reveals time- and condition-resolved directed interactions consistent with known neurobiological circuitry. Together, these results establish a principled distribution-aware framework for effective-connectivity estimation beyond Gaussian residual modeling.
Chinese Translation
从脑信号中估计有效连接性通常依赖于高斯残差建模,这虽然能够实现易于处理的估计,但当经验残差非高斯时,可能会丢弃有信息的分布结构并扭曲推断出的有向交互。我们跨多种模态、物种和实验条件表明,脑信号和拟合的信道残差都经常偏离高斯性。因此,我们引入了一种基于一般残差分布下信道容量的、分布感知的、信息论的有效连接性度量。为了从经验的、可能非高斯的残差中估计由此得到的容量,我们开发了一种基于归一化流的双流最小-最大估计器,其中生成器在功率约束下搜索允许的输入分布,而观察者估计输出熵。我们提供了该估计器的理论刻画,表明观察者目标恢复了微分熵,直到一个KL近似项,该公式作为特例退化为经典高斯容量,并且残差熵可以在方差之外改变可达信息速率;博弈间隙和误差分析进一步刻画了优化和近似来源。在具有已知有向连接性的类脑模拟中,与高斯容量、格兰杰因果、VAR-LiNGAM和GIMME相比,Dual-flow在涵盖不同网络拓扑、隐藏驱动、反馈和异质性血流动力学的十种条件下实现了最高的AUROC和AUPRC。应用于多模态脑信号时,该方法揭示了与已知神经生物回路一致的时间分辨和条件分辨的有向交互。总之,这些结果建立了一个超越高斯残差建模的、原则性的分布感知框架,用于有效连接性估计。
cs.LG / 166 / 2609.32775

Robust Bayesian Optimization with Q-Exponential Surrogates

基于q-指数代理模型的鲁棒贝叶斯优化
Suwandi, Richard Cornelius, Lin, Zhidi, Yin, Feng, Zoubir, Abdelhak M.
Abstract
Bayesian optimization (BO) is a widely used framework for optimizing expensive black-box objectives, but standard BO methods often use Gaussian process (GP) surrogates whose Gaussian assumption is sensitive to outliers and heavy-tailed noise. We introduce q-ED-BO, a robust BO method whose surrogate follows a univariate q-exponential (q-ED) distribution, preserving GP-BO's closed-form posterior mean and variance while a shape parameter q controls the tail behavior, recovering the GP at q = 2 and growing heavier-tailed with wider confidence bounds as q decreases. This tractability yields a closed-form q-upper confidence bound (q-UCB) with sublinear regret, and an exact closed-form q-expected improvement (q-EI) that generalizes EI to the heavy-tailed predictive, recovering classical EI at q = 2. Experiments on beamformer and adaptive filter tuning with impulsive outliers show that q-ED-BO matches or exceeds existing baselines on clean data, and under corruption, improves the strongest baseline by approximately 0.7 dB in output SINR and 1.1 to 1.2 dB in misalignment reduction.
Chinese Translation
贝叶斯优化(BO)是一种广泛用于优化昂贵黑箱目标的框架,但标准BO方法通常使用高斯过程(GP)代理模型,其高斯假设对异常值和重尾噪声敏感。我们提出了q-ED-BO,一种鲁棒的BO方法,其代理模型服从单变量q-指数(q-ED)分布,保留了GP-BO的闭式后验均值和方差,同时形状参数q控制尾部行为,当q=2时恢复为GP,随着q减小,尾部变得更重,置信界限更宽。这种可处理性产生了具有次线性遗憾的闭式q-上置信界(q-UCB),以及一个精确的闭式q-期望改进(q-EI),它将EI推广到重尾预测分布,在q=2时恢复为经典EI。在带有脉冲异常值的波束形成器和自适应滤波器调谐上的实验表明,q-ED-BO在干净数据上达到或超过现有基线,并且在数据受损情况下,将最强基线在输出SINR上提高了约0.7 dB,在失准减少方面提高了1.1至1.2 dB。
cs.LG / 167 / 2609.32785

Learning When to Recur: Token-Adaptive Recursion for Imbalanced Ophthalmic Domain Incremental Learning

学习何时递归:用于不平衡眼科领域增量学习的Token自适应递归
Yu, Nanxi, Li, Kang, Du, Ye, Hu, Xiaowei, Yang, Weihua, Wang, Shujun
Abstract
Domain incremental learning is essential for adapting ophthalmic deep learning models to sequential clinical domains while preserving diagnostic expertise. Existing domain incremental learning methods predominantly address the domain shift induced by style variations. However, they often overlook the severe class imbalance inherent in real-world clinical scenarios, such as clinical referral systems. Institutions in these systems encounter drastic fluctuations in class priors, resulting in label distribution shift, a critical form of domain shift that triggers severe catastrophic forgetting. To address these challenges, we propose ToRe, a rehearsal-free and parameter-efficient framework that leverages frozen ophthalmic foundation models for robust incremental adaptation. ToRe employs a parameter isolation strategy to decouple domain-specific optimization paths, thereby helping mitigate catastrophic forgetting driven by both label distribution shift and style variations. Simultaneously, it introduces token-adaptive recursion that adaptively allocates additional computational depth across tokens, allowing simple tokens to exit the recursion loop early while subjecting complex tokens, such as those associated with lesions, to deeper recursive processing. This mechanism enhances the feature representations for minority classes, thereby supporting generalization throughout the domain incremental learning process. Extensive evaluations on nine heterogeneous datasets demonstrate that ToRe consistently outperforms state-of-the-art methods in overall performance across the three benchmarks, while maintaining near-zero forgetting. Together, these results support the applicability of ToRe to dynamic and imbalanced clinical environments. The code is available at https://github.com/Nancyolo/ToRe
Chinese Translation
领域增量学习对于使眼科深度学习模型适应顺序临床领域,同时保留诊断专业知识至关重要。现有的领域增量学习方法主要解决由风格变化引起的领域偏移。然而,它们往往忽略了真实临床场景(如临床转诊系统)中固有的严重类别不平衡。这些系统中的机构遇到类别先验的剧烈波动,导致标签分布偏移,这是一种关键的领域偏移形式,会引发严重的灾难性遗忘。为了应对这些挑战,我们提出了ToRe,一个无需回放且参数高效的框架,利用冻结的眼科基础模型进行鲁棒的增量适应。ToRe采用参数隔离策略来解耦领域特定的优化路径,从而有助于缓解由标签分布偏移和风格变化驱动的灾难性遗忘。同时,它引入了Token自适应递归,该递归在Token之间自适应地分配额外的计算深度,允许简单的Token提前退出递归循环,而将复杂的Token(例如与病变相关的Token)进行更深入的递归处理。这种机制增强了少数类的特征表示,从而支持整个领域增量学习过程中的泛化。在九个异构数据集上的广泛评估表明,ToRe在三个基准上的整体性能始终优于最先进的方法,同时保持接近零的遗忘。这些结果共同支持了ToRe在动态和不平衡临床环境中的适用性。代码可在 https://github.com/Nancyolo/ToRe 获取。
cs.LG / 168 / 2609.32788

Getting Motif-ated: Controllable AI Compositions from Injected Motif Prompts

动机激发:来自注入动机提示的可控AI作曲
Yang, Chao Peter, Rudin, Cynthia, Jiang, Yue, Mak, Simon, Ni-Hahn, Stephen
Abstract
Deep learning has transformed symbolic music generation by borrowing the training paradigms of large language models, with systems such as NotaGen now producing complete, stylistically convincing classical scores from a short prompt. These systems could become powerful creative partners, helping musicians generate endless possibilities. However, current systems expose almost no control handles on the music itself. In principle, control handles could be built into a foundation model trained from scratch, but this is rarely practical without massive amounts of quality annotated data and compute resources. We therefore present MotiGen, a recipe for retrofitting pretrained symbolic music models to use new instruction prompts. MotiGen injects a musical motif as a structured prompt line, reinforces it with a scalar attention bias toward the motif tokens, and learns the association with a two-phase curriculum. First, it learns from focused excerpts cropped around motif occurrences in the training data, then full scores including the motifs. Our experiments show that our model composes with the prompted motif in over 92.3\% of generated pieces. Generated pieces using a variety of motifs are included in our sample site: https://motigen-site.github.io/.
Chinese Translation
深度学习通过借鉴大型语言模型的训练范式,已经改变了符号音乐生成,像NotaGen这样的系统现在能够根据简短提示生成完整且风格令人信服的古典乐谱。这些系统可能成为强大的创意伙伴,帮助音乐家产生无限的可能性。然而,当前系统几乎不提供对音乐本身的控制手段。原则上,控制手段可以内置到从头训练的基础模型中,但如果没有大量高质量注释数据和计算资源,这很少可行。因此,我们提出MotiGen,一种对预训练符号音乐模型进行改造以使用新指令提示的方法。MotiGen将音乐动机作为结构化提示行注入,通过一个标量注意力偏置来强化对动机token的关注,并通过两阶段课程学习这种关联。首先,它从训练数据中围绕动机出现处裁剪出的聚焦片段中学习,然后学习包含动机的完整乐谱。我们的实验表明,我们的模型在超过92.3%的生成作品中都使用了提示的动机进行作曲。使用各种动机生成的作品包含在我们的示例网站中:https://motigen-site.github.io/。
cs.LG / 169 / 2609.32790

FedHV: Low-Overhead Hypervolume Weighting for Federated Multi-Objective Optimization

FedHV:用于联邦多目标优化的低开销超体积加权
Dehghanpour, Amirardalan, Azimi-Abarghouyi, Seyed Mohammad, Brinton, Christopher G.
Abstract
Task-wise federated multi-objective optimization (FedMOO) trains a shared model for competing prediction objectives under heterogeneous data, partial participation, and communication constraints. Existing methods commonly derive task weights from gradient or update geometry. This requires task-specific information or iterative server-side optimization. We introduce FedHV, which maps reference-relative objective slacks to closed-form inverse-slack weights. Each client optimizes one weighted loss and returns objective estimates with its model update. The protocol adds exactly 2m auxiliary scalars per participating client, yielding Theta(d + m) total per-client communication, compared with the Theta(md) task-specific communication of FSMGDA, and requires no additional synchronization stage. We analyze the resulting one-round-delayed weights under client heterogeneity, multi-step local updates, partial participation, and finite-sample objective reports. Under a fixed-horizon positive-slack reference condition, with the prescribed horizon-dependent step size and vanishing report error, FedHV achieves an O(T^(-1/2)) rate for the average squared log-hypervolume gradient norm; persistent report error determines the resulting stationarity neighborhood. The same bound controls the squared Pareto-stationarity residual. Across six Dirichlet-partitioned non-IID settings from four vision benchmark families and three training seeds, FedHV exceeds FSMGDA and FedCMOO in mean accuracy in five settings. Among these methods and uniform scalarization, it achieves the highest worst-task accuracy in four settings and improves the difficult CIFAR-10 objective in both CIFAR10-MNIST settings.
Chinese Translation
任务级联邦多目标优化(FedMOO)在异构数据、部分参与和通信约束下,为相互竞争的预测目标训练一个共享模型。现有方法通常从梯度或更新几何中导出任务权重,这需要任务特定信息或迭代的服务器端优化。我们提出了FedHV,它将相对于参考点的目标松弛量映射为闭式反松弛权重。每个客户端优化一个加权损失,并随其模型更新返回目标估计值。该协议为每个参与的客户端精确增加2m个辅助标量,产生每个客户端总通信量为Θ(d+m),相比之下FSMGDA的任务特定通信量为Θ(md),并且不需要额外的同步阶段。我们在客户端异构性、多步本地更新、部分参与和有限样本目标报告下,分析了由此产生的一轮延迟权重。在固定时域正松弛参考条件下,使用指定的时域相关步长和消失的报告误差,FedHV对于平均平方对数超体积梯度范数达到了O(T^(-1/2))的速率;持续的报告误差决定了最终的平稳性邻域。相同的界控制了平方Pareto平稳性残差。在来自四个视觉基准家族的六个Dirichlet划分的非独立同分布设置和三个训练种子中,FedHV在五个设置中的平均准确率超过了FSMGDA和FedCMOO。在这些方法和均匀标量化中,它在四个设置中实现了最高的最差任务准确率,并在两个CIFAR10-MNIST设置中改善了困难的CIFAR-10目标。
cs.LG / 170 / 2609.32792

Understanding and Exploiting Anisotropy in Post-Training

理解与利用后训练中的各向异性
Jha, Samyak, Saini, Harshvardhan, Liao, Yizhen, Tang, Yiming, Liu, Dianbo
Abstract
LLM post-training combines supervised fine-tuning (SFT), a mode-covering forward-KL objective, with reinforcement learning (RL), a mode-seeking reverse-KL objective. Frequency-weighted likelihood training leaves a well-known signature: \emph{anisotropy}, in which a few residual channels carry disproportionately large activations. Anisotropy is widely documented and usually treated as a defect, yet its function and its interaction with post-training remain unclear. We first analyze it. A label-free outlier rule isolates about 5\% of residual channels that are essential for language modeling: removing them raises perplexity from 10 to over $10^6$, versus 35 for count-matched random channels. Yet they barely distinguish correct from incorrect reasoning. SFT reshapes them, whereas RL leaves them largely intact and adapts the complementary channels. These channels therefore form the model's \emph{coherence substrate}, and reasoning adaptation happens elsewhere. We then exploit this. \textsc{SphereGate} learns one bounded gain per residual channel on a frozen backbone. Its activation-weighted gradients provably limit movement of high-energy coherence channels and leave the remaining channels free. With 0.1M trainable parameters, \textsc{SphereGate} outperforms parameter-efficient baselines by 2.0--7.3 points on MATH-500 across Qwen2.5 (0.5B--7B) and Llama-3-8B, is comparable or exceeds full-model GRPO. Anisotropy is not a defect but a division of labor that post-training can exploit.
Chinese Translation
LLM后训练将监督微调(SFT)——一种模式覆盖的前向KL目标——与强化学习(RL)——一种模式寻求的反向KL目标——相结合。频率加权似然训练留下了一个众所周知的特征:各向异性,即少数残差通道承载着不成比例的大激活值。各向异性已被广泛记录,通常被视为一种缺陷,但其功能及其与后训练的相互作用仍不清楚。我们首先对其进行分析。一种无标签的离群值规则分离出约5%的残差通道,这些通道对语言建模至关重要:移除它们会使困惑度从10上升到超过10^6,而数量匹配的随机通道仅上升到35。然而,它们几乎无法区分正确与错误的推理。SFT重塑它们,而RL基本保持它们不变,并适应互补的通道。因此,这些通道构成了模型的连贯性基底,而推理适应发生在其他地方。然后我们利用这一点。SphereGate在冻结的主干网络上为每个残差通道学习一个有界的增益。其激活加权梯度可证明地限制了高能量连贯性通道的移动,并让其余通道自由。仅有0.1M可训练参数,SphereGate在MATH-500上,在Qwen2.5(0.5B--7B)和Llama-3-8B上,比参数高效基线高出2.0--7.3分,与全模型GRPO相当或超过。各向异性不是缺陷,而是后训练可以利用的一种分工。
cs.LG / 171 / 2609.32797

Derivative-Informed Training of Neural Operators On-the-Fly via Sketched Tangent Consistency

基于草图切向一致性的神经算子在线导数信息训练
Yang, Xinhan, Lu, Lu, Mou, Shancong
Abstract
Derivative-informed training improves neural operators by directly supervising their input-output sensitivities, which is crucial when neural operators are used as differentiable surrogates for inverse problems, PDE-constrained optimization, design, and control. However, existing methods rely on offline-generated derivative labels, making data generation slow, storage-intensive, and difficult to adapt across datasets, resolutions, or perturbation bases. We propose sketched tangent consistency loss (sTCL), an on-the-fly derivative-informed training objective that, for PDEs with a known and differentiable residual, enforces sensitivity consistency directly from the governing equation without offline tangent labels or neural-operator architecture changes. sTCL uses randomly sketched input perturbations to provide a lightweight derivative-level physics constraint during training. However, raw forward-sensitivity residual penalties can fail for stiff, ill-conditioned, indefinite, or coupled saddle-point tangent operators. To address this, we introduce lightweight operator-aware loss-conditioning mechanisms selected by a simple tangent-operator decision rule. Across Helmholtz, nonlinear diffusion-reaction, Burgers, Allen-Cahn, and Navier-Stokes, with the same neural-operator backbone for all methods, the PDE-specific sTCL losses achieve solution and Jacobian accuracy comparable to offline derivative-informed training (DIFNO) while eliminating the offline derivative-data generation and storage pipeline. These results show that on-the-fly derivative-informed training need not merely amortize offline tangent-solve cost into training; with appropriate sketching and loss design, sTCL provides an effective drop-in path to derivative-informed neural operators. Code is available at https://github.com/yang9579/Derivative-informed-traning-on-the-fly.
Chinese Translation
导数信息训练通过直接监督神经算子的输入-输出敏感度来改进神经算子,当神经算子被用作可微代理模型用于反问题、偏微分方程约束优化、设计和控制时,这一点至关重要。然而,现有方法依赖于离线生成的导数标签,导致数据生成缓慢、存储密集,并且难以适应不同的数据集、分辨率或扰动基。我们提出草图切向一致性损失(sTCL),一种在线导数信息训练目标,对于具有已知且可微残差的偏微分方程,它直接从控制方程强制敏感度一致性,而无需离线切向标签或神经算子架构更改。sTCL 使用随机草图的输入扰动,在训练期间提供轻量级的导数级物理约束。然而,对于刚性、病态、不定或耦合鞍点切向算子,原始的前向敏感度残差惩罚可能会失效。为了解决这个问题,我们引入了轻量级的算子感知损失条件化机制,由简单的切向算子决策规则选择。在 Helmholtz、非线性扩散-反应、Burgers、Allen-Cahn 和 Navier-Stokes 方程上,对于所有方法使用相同的神经算子主干,针对特定 PDE 的 sTCL 损失达到了与离线导数信息训练(DIFNO)相当的解和雅可比精度,同时消除了离线导数数据生成和存储流程。这些结果表明,在线导数信息训练不必仅仅将离线切向求解成本摊销到训练中;通过适当的草图设计和损失设计,sTCL 为导数信息神经算子提供了一条有效的即插即用途径。代码可在 https://github.com/yang9579/Derivative-informed-traning-on-the-fly 获取。
cs.LG / 172 / 2609.32808

Mind the Spike: Mechanisms and Brittleness of Visual Massive Activations in Large Vision-Language Models

注意尖峰:大型视觉语言模型中视觉巨量激活的机制与脆弱性
Ngnawé, Jonas, Pequignot, Yann, Sahoo, Sabyasachi, Gagné, Christian, Precioso, Frédéric, Koyejo, Sanmi
Abstract
Large vision-language models (LVLMs) inherit massive activations from their text-only bases: spikes where a few fixed hidden channels receive values thousands of times above the typical magnitude. The text spike systematically appears in early layers at a fixed initial position, independently of input content. Visual spikes vary across images, but whether their formation follows a consistent pattern across LVLMs and how they respond to image perturbations remain open questions. We find that some LVLMs do not form visual spikes, while others spike at different rates, typically in deeper layers. We identify the trigger direction from model weights and an interpretable location rule: before the language model decoder runs, eventual spike tokens are largely restricted to those sharing least with the rest of the image. Crucially, visual spikes are strikingly brittle. Common corruptions frequently create and relocate spikes, and less often remove them, raising overall incidence. Our trigger-guided spike attack deliberately creates or removes spikes under a small $\ell_\infty$ budget, with 1/255 enough in nine of the ten models that spike. Finally, our preventive intervention removes only the trigger component before spikes erupt, eliminating or substantially reducing spikes on clean and perturbed images while leaving the other image tokens nearly unchanged. Our study spans 25 adapter-based LVLMs built on 18 released text-only bases from 10 families, ranging from 2B to 72B parameters.
Chinese Translation
大型视觉语言模型(LVLMs)从其纯文本基座中继承了巨量激活:即少数固定的隐藏通道接收到比典型幅度高数千倍的值,形成尖峰。文本尖峰系统性地出现在早期层中一个固定的初始位置,与输入内容无关。视觉尖峰因图像而异,但其形成是否在不同 LVLM 中遵循一致模式,以及它们如何响应图像扰动,仍是未决问题。我们发现一些 LVLM 不形成视觉尖峰,而另一些则以不同频率出现尖峰,通常出现在更深的层中。我们从模型权重中识别出触发方向,并得到一个可解释的位置规则:在语言模型解码器运行之前,最终形成尖峰的 token 在很大程度上局限于那些与图像其余部分共享最少信息的 token。关键的是,视觉尖峰极其脆弱。常见的损坏会频繁地创建和移动尖峰,较少地移除它们,从而提高了总体发生率。我们的触发引导的尖峰攻击在小的 $\ell_\infty$ 预算下故意创建或移除尖峰,在十个出现尖峰的模型中有九个仅需 1/255 就足够。最后,我们的预防性干预在尖峰爆发前仅移除触发成分,从而在干净和扰动图像上消除或大幅减少尖峰,同时使其他图像 token 几乎保持不变。我们的研究涵盖 25 个基于适配器的 LVLM,它们构建于来自 10 个系列的 18 个已发布纯文本基座之上,参数范围从 2B 到 72B。
cs.LG / 173 / 2609.32814

Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers

改变乘积,保留参数:面向 Transformer 的结合代数层
Koziev, Ilya, Oseledets, Ivan
Abstract
Fast matrix multiplication algorithms keep the product fixed and search for a cheaper way to evaluate it. We instead ask whether a Transformer's learned projections can use a different, cheaper product altogether. Building on an associative-algebra construction that replaces ordinary matrix multiplication with a sparser interaction table over the same weight blocks, we construct a family with quadratic arithmetic in the matrix dimension when the physical block size remains fixed, and derive finite-shape constraints for GPU execution. The construction is provably optimal for its bilinear rank by the Alder--Strassen bound and can be realized as row-typed rectangular projections compatible with causal masking and KV-cached decoding. We provide an empirical test of this approach by training two approximately 110M-parameter decoder-only Transformer LMs from the same recipe and 12.3B-token budget, differing only in their feed-forward layer: one uses ordinary dense matrix multiplication and the other uses the associative-algebra product. Across four prompt domains, the algebraic model achieves a 6.2--7.8\% increase in end-to-end generation throughput, while obtaining lower scores on all three reported downstream metrics. We treat these results as a feasibility and trainability check for the proposed approach at small scale, leaving further investigation to future work.
Chinese Translation
快速矩阵乘法算法保持乘积不变,并寻找更廉价的计算方式。我们则反过来追问:Transformer 学习到的投影能否使用一种完全不同且更廉价的乘积。基于一种结合代数构造——该构造在相同权重块上用更稀疏的交互表替换普通矩阵乘法——我们构造了一族方法:当物理块大小保持固定时,其算术复杂度在矩阵维度上为二次;并推导了面向 GPU 执行的有限形状约束。根据 Alder--Strassen 界,该构造对其双线性秩可证明是最优的,并且可实现为行类型(row-typed)矩形投影,与因果掩码和 KV 缓存解码兼容。我们通过训练两个约 110M 参数的仅解码器 Transformer 语言模型,对这种方法进行了实证测试;二者采用相同配方和 12.3B token 预算,仅在前馈层不同:一个使用普通稠密矩阵乘法,另一个使用结合代数乘积。在四个提示域上,代数模型实现了端到端生成吞吐量 6.2--7.8% 的提升,但在报告的全部三个下游指标上得分更低。我们将这些结果视为对所提方法在小规模下的可行性和可训练性检验,并将进一步研究留待未来工作。
cs.LG / 174 / 2609.32828

Retimed Bellman Flows: Escaping the Impossible Triangle of Velocity Bootstrapping

重定时贝尔曼流:逃离速度自举的不可能三角
Xu, Boyang, Chen, Shengzhe, Yan, Hao
Abstract
Flow critics learn return distributions by transporting Gaussian noise to Bellman endpoints via continuous velocity fields. While velocity bootstrapping stabilizes training by querying a successor teacher, existing methods face a structural dilemma: on straight paths, no residual-free same-time affine mapping can preserve Gaussian initial noise while maintaining an unbiased target. To overcome this limitation, we introduce Retimed Bellman Flows (ReBF). ReBF queries the teacher critic at a dynamically shifted earlier flow time, aligning intermediate student and teacher trajectories. By combining this retimed clock with fresh, decoupled noise generation, ReBF constructs a provably conditionally unbiased velocity target that preserves the Bellman fixed point and contracts under Wasserstein distances. Empirically, ReBF reduces $W_1$ distance to ground-truth return distributions by up to $7.7\times$ on synthetic MRPs and outperforms existing flow critics across 38 challenging OGBench and D4RL offline reinforcement learning tasks.
Chinese Translation
流评论家通过连续速度场将高斯噪声传输到贝尔曼端点来学习回报分布。虽然速度自举通过查询后继教师来稳定训练,但现有方法面临一个结构性困境:在直线路径上,没有无残差的同步仿射映射能够在保持无偏目标的同时保持高斯初始噪声。为了克服这一局限,我们引入了重定时贝尔曼流(ReBF)。ReBF在动态偏移的较早流时间查询教师评论家,对齐中间学生和教师轨迹。通过将这种重定时时钟与新鲜、解耦的噪声生成相结合,ReBF构建了一个可证明的条件无偏速度目标,该目标保持了贝尔曼不动点并在Wasserstein距离下收缩。经验上,ReBF在合成MRP上将与真实回报分布的$W_1$距离最多减少了7.7倍,并在38个具有挑战性的OGBench和D4RL离线强化学习任务上优于现有的流评论家。
cs.LG / 175 / 2609.32831

UniCache: Task- and Type-Aware KV Cache Compression for Unified Multimodal Models

UniCache: 面向统一多模态模型的任务与类型感知KV缓存压缩
Yang, Wanqi, Ma, Yuexiao, Xie, Mei, Zheng, Xiawu, Liu, Shiwei
Abstract
Unified multimodal models combine understanding, generation, and editing within a single network, offering a promising foundation for versatile multimodal applications. However, growing multimodal contexts make KV cache storage and access increasingly costly. Existing KV cache compression methods are typically tailored to specific tasks and single-modality caches, while overlooking changes in cache importance across tasks and timesteps. However, in unified multimodal models, each task involves multiple KV cache types, and both their composition and dynamics differ across tasks. As a result, a single compression policy overlooks task- and type-specific requirements, leading to the loss of critical information and degraded quality across tasks. Based on these findings, we propose UniCache, a training-free framework for task- and type-aware KV cache compression. UniCache identifies the cache segments activated by each task and assigns suitable compression policies through offline calibration. It coordinates their parallel execution under a shared storage budget through attention-guided allocation and task-aware temporal scheduling. Experiments show that UniCache achieves $5\times$ KV cache compression for understanding and editing and $2.5\times$ for generation with negligible quality loss, while increasing throughput by up to $1.78\times$ in long-context settings, significantly improving the practicality of scaling unified multimodal models to longer context.
Chinese Translation
统一多模态模型将理解、生成和编辑结合在一个单一网络中,为多功能多模态应用提供了有前景的基础。然而,不断增长的多模态上下文使得KV缓存存储和访问变得越来越昂贵。现有的KV缓存压缩方法通常针对特定任务和单模态缓存定制,而忽略了缓存重要性在不同任务和时间步长上的变化。然而,在统一多模态模型中,每个任务涉及多种KV缓存类型,并且它们的组成和动态因任务而异。因此,单一的压缩策略忽略了任务和类型特定的需求,导致关键信息丢失和跨任务质量下降。基于这些发现,我们提出了UniCache,一个无需训练的任务与类型感知KV缓存压缩框架。UniCache识别每个任务激活的缓存段,并通过离线校准分配适当的压缩策略。它通过注意力引导的分配和任务感知的时间调度,在共享存储预算下协调它们的并行执行。实验表明,UniCache在理解和编辑任务上实现了5倍的KV缓存压缩,在生成任务上实现了2.5倍的压缩,质量损失可忽略不计,同时在长上下文设置下将吞吐量提高了高达1.78倍,显著提高了将统一多模态模型扩展到更长上下文的实用性。
cs.LG / 176 / 2609.32832

Over-the-Air Federated Learning in Heterogeneous Mobile Wireless Networks

异构移动无线网络中的空中联邦学习
Xiang, Ming, Michelusi, Nicolò, Eldar, Yonina C., Su, Lili
Abstract
Over-the-air computation has emerged as a scalable and efficient solution for deploying federated learning algorithms in wireless networks by exploiting waveform superposition for simultaneous model aggregation. Most existing work struggles with heterogeneous fading channels. These approaches either enforce unbiased updates from all devices or allow partial device contributions, requiring careful tuning of the convergence bound to mitigate bias under specific fading models. However, the former significantly amplifies receiver noise due to the weakest channel, whereas the latter is sensitive to fading model mismatch and converges only to a biased objective. To tackle these challenges, we propose FedOAG, which employs algorithmic components to automatically satisfy energy constraints via gradient normalization and evenly mix devices' updates through implicit gossiping. Importantly, FedOAG does not require transmission from all devices, nor does it rely on a specific fading model or knowledge of time-varying statistical channel distributions. We show that FedOAG converges to a stationary point of an unbiased non-convex objective at the best possible rate $O(1/\sqrt{T})$ for any stochastic first-order method. We corroborate our analysis with numerical experiments over dynamic wireless conditions on real-world datasets.
Chinese Translation
空中计算通过利用波形叠加进行同时模型聚合,已成为在无线网络中部署联邦学习算法的一种可扩展且高效的解决方案。大多数现有工作难以处理异构衰落信道。这些方法要么强制所有设备进行无偏更新,要么允许部分设备贡献,并需要仔细调整收敛界以在特定衰落模型下减轻偏差。然而,前者由于最弱信道会显著放大接收机噪声,而后者对衰落模型失配敏感,并且仅收敛到一个有偏目标。为应对这些挑战,我们提出 FedOAG,它采用算法组件通过梯度归一化自动满足能量约束,并通过隐式 gossip 均匀混合设备更新。重要的是,FedOAG 不要求所有设备都传输,也不依赖于特定衰落模型或时变统计信道分布的知识。我们证明,对于任何随机一阶方法,FedOAG 以最佳可能速率 $O(1/\sqrt{T})$ 收敛到无偏非凸目标的稳定点。我们通过在真实世界数据集上动态无线条件下的数值实验验证了我们的分析。
cs.LG / 177 / 2609.32833

Permutation-Equivariant Flow Matching for Alignment-Free Neural Weight Generation

置换等变流匹配用于无需对齐的神经权重生成
Piven, Arkadi, Eitan, Yam, Bar-Shalom, Guy, Frasca, Fabrizio, Cremers, Daniel, Dagès, Thomas, Kimmel, Ron, Maron, Haggai
Abstract
A trained neural network can be represented by a parameter vector in high dimensions. Learning distributions over these vectors enables the generation of new models across various tasks and architectures. A central challenge is permutation symmetry: permuting hidden neurons can produce distant parameter vectors representing the same function. This introduces variations that a generative model must account for when learning from trained networks. Existing methods typically address this using networks derived from a common base model or costly approximate neuron alignment. We instead parameterize a flow-matching velocity field with a permutation-equivariant Graph Meta Network, enabling direct learning from independently trained networks without alignment. Extensive experiments show that our method closely reproduces the joint statistics of accuracy, functional similarity, and weight similarity of independently trained collections, providing evidence of generation beyond checkpoint memorization. A single conditional model also generates task-specific networks on heterogeneous architectures and generalizes to unseen hidden-width configurations. On a tabular domain-shift task, intermediate conditioning produces individual networks with performance comparable to logit ensembles across both domains. Taken together, our results show how permutation equivariance enables learning from diverse collections of independently trained networks without permutation alignment.
Chinese Translation
训练好的神经网络可以用高维参数向量表示。学习这些向量上的分布能够生成跨各种任务和架构的新模型。一个核心挑战是置换对称性:置换隐藏神经元可以产生表示相同函数的相距很远的参数向量。这引入了生成模型在从训练好的网络学习时必须考虑的变化。现有方法通常使用从共同基础模型派生的网络或昂贵的近似神经元对齐来解决这个问题。我们转而使用置换等变图元网络(permutation-equivariant Graph Meta Network)参数化流匹配速度场,从而能够直接从独立训练的网络中学习,无需对齐。大量实验表明,我们的方法能够紧密复现独立训练集合的准确率、功能相似性和权重相似性的联合统计量,提供了超越检查点记忆的生成证据。单个条件模型还能在异构架构上生成任务特定的网络,并泛化到未见过的隐藏宽度配置。在表格领域偏移任务上,中间条件化产生的单个网络在两个领域上的性能与 logit 集成相当。综上所述,我们的结果展示了置换等变性如何使得无需置换对齐即可从独立训练网络的多样化集合中学习。
cs.LG / 178 / 2609.32838

HamiFormer: Dual-Expert Diffusion Fields with Affine Symplectic Maps

HamiFormer:具有仿射辛映射的双专家扩散场
Huang, Haoxiang, Liu, Xiang, Wang, Shuwei, Ma, Jingheng, Cui, Sen
Abstract
Predicting smooth dynamics and collisions requires modeling continuous evolution and abrupt state changes. We introduce HamiFormer, a dual-expert diffusion field combining whole-window denoising with residual-corrected Hamiltonian propagation. Their mixed-state feedback attenuates the direct contribution of inherited autoregressive error: each mixed state guides subsequent propagation within the jointly refined window. Parallel Local Affine Scan (PLAS) amortizes iterative refinement across rectified-flow steps and evaluates derivatives in parallel across physical time. PLAS's affine symplectic maps achieve lower solver error and runtime than sequential explicit Euler in our evaluation. A Regime Model Tree specializes residuals and routing to balance typical-state accuracy against large tail errors. Our analysis gives conditions for physically consistent refinement and warm-start tracking, and finite-window error bounds under diffusion feedback. In 192-step evaluations, HamiFormer reduces normalized phase-space MSE by 26.3% against PhysiFormer on HamiBalls-1 and 21.4% against DiT on HamiBalls-2, with comparable model capacities. Disjoint-interval comparisons show the lowest late-horizon position and momentum errors among baselines on both datasets. Project page: https://hamiformer.github.io/.
Chinese Translation
预测平滑动力学和碰撞需要对连续演化和突变状态变化进行建模。我们提出HamiFormer,一种结合全窗口去噪与残差校正的哈密顿传播的双专家扩散场。它们的混合状态反馈减弱了继承的自回归误差的直接贡献:每个混合状态在联合精炼窗口内引导后续传播。并行局部仿射扫描(PLAS)将迭代精炼分摊到整流流步骤中,并在物理时间上并行计算导数。在我们的评估中,PLAS的仿射辛映射比顺序显式欧拉法实现了更低的求解器误差和运行时间。Regime Model Tree专门处理残差和路由,以平衡典型状态精度与大尾部误差。我们的分析给出了物理一致精炼和热启动跟踪的条件,以及扩散反馈下的有限窗口误差界。在192步评估中,HamiFormer在HamiBalls-1上相比PhysiFormer将归一化相空间MSE降低了26.3%,在HamiBalls-2上相比DiT降低了21.4%,且模型容量相当。不相交区间比较显示,在两个数据集上,其晚期时域位置和动量误差在基线中最低。项目页面:https://hamiformer.github.io/。
cs.LG / 179 / 2609.32849

Transfer Learning for Edge Classification on Dynamic Text-Attributed Graphs

动态文本属性图上的边分类迁移学习
Bonnet, Tyler, Rei, Marek
Abstract
Learning transferable representations for dynamic text-attributed graphs (DyTAGs) requires models to capture underlying interaction dynamics that persist across domains. However, existing methods tend to overfit to domain-specific structural, temporal, and semantic patterns, limiting edge classification performance under distribution shifts. To expose and address this, we formally establish a leave-one-domain-out (LODO) transfer learning protocol for edge classification on DyTAGs. Under this protocol, we demonstrate that state-of-the-art self-supervised methods for dynamic graph learning perform poorly when transferred to unseen domains. Strikingly, existing methods underperform a structurally and temporally unaware Bag of Events (BoE) model we introduce, which inputs only unordered sequences of node and edge text features. Proceeding from the BoE, we propose Spatio-Temporal Semantic Alignment (STSA), which integrates a spatio-temporal encoder that fuses representations of time deltas and node occurrence frequencies into a unified manifold. STSA is trained with a Contrastive Semantic Forecasting objective, which anchors edge representations to a multi-domain textual latent space initialized by a pretrained language model, providing a robust prior that outperforms BoE and all existing methods we evaluate.
Chinese Translation
学习动态文本属性图(DyTAGs)的可迁移表示,要求模型能够捕获跨领域持续存在的底层交互动态。然而,现有方法往往过度拟合特定领域的结构、时间和语义模式,限制了分布偏移下的边分类性能。为了揭示并解决这一问题,我们正式建立了一种针对 DyTAGs 上边分类的留一领域(LODO)迁移学习协议。在该协议下,我们证明了最先进的动态图学习自监督方法在迁移到未见领域时表现不佳。令人惊讶的是,现有方法的表现不如我们引入的一个对结构和时间无感知的事件袋(BoE)模型,该模型仅输入节点和边文本特征的无序序列。从 BoE 出发,我们提出了时空语义对齐(STSA),它集成了一个时空编码器,将时间差和节点出现频率的表示融合到一个统一的流形中。STSA 使用对比语义预测目标进行训练,该目标将边表示锚定到由预训练语言模型初始化的多领域文本潜在空间,提供了一个优于 BoE 和我们评估的所有现有方法的稳健先验。
cs.LG / 180 / 2609.32854

Logic Gate Networks and Lookup Table Networks as Lightweight Hardware Classifiers for Inter-patient ECG Arrhythmia Classification

逻辑门网络和查找表网络作为轻量级硬件分类器用于患者间心电图心律失常分类
Mommen, Wout, Keuninckx, Lars, Patil, Siddharth, Detterer, Paul, Colpaert, Achiel, Wambacq, Piet
Abstract
Deep Differentiable Logic Gate Networks (LGNs) and Lookup Table Networks (LUTNs) offer a promising approach for very low power inference due to their use of simple binary logic operations instead of arithmetic. In this work, we generalize the logic gates of LGNs to more than two input pins, naturally arriving at networks consisting of $N$-input LUTs. To obtain a differentiable expression for training the $N$-LUT entries, we adopt the Boolean equation of a $2^N$:1 multiplexer (MUX) and optimize its input parameters during training. We investigate the applicability of LGNs and LUTNs to inter-patient ECG arrhythmia classification using the MIT-BIH data set. The proposed models achieve up to 94.41\% accuracy and a $j\kappa$ index of 0.683 on a four-class task, showing a competitive performance compared to existing CNN-, SVM- and SNN-based methods. Our LGNs and LUTNs only require an estimated 2.89k to 6.17k FLOPs, including preprocessing and readout, which is three to six orders of magnitude less than state-of-the-art methods. We verified our design, which consists of the preprocessing pipeline and a 6-LUTN classifier, by implementing it on a Xilinx Zynq-7000 ZedBoard. The complete system consumes a dynamic energy of 8.25 $\mu$J/inference, of which only 0.46 nJ is utilized by the LUTN classifier. These results show that both LGNs and LUTNs can be employed as lightweight hardware-based classifiers for inter-patient ECG arrhythmia classification.
Chinese Translation
深度可微逻辑门网络(LGNs)和查找表网络(LUTNs)由于使用简单的二进制逻辑操作而非算术运算,为极低功耗推理提供了一种有前景的方法。在这项工作中,我们将LGN的逻辑门推广到两个以上输入引脚,自然地得到了由$N$输入LUT组成的网络。为了获得用于训练$N$-LUT条目的可微表达式,我们采用了$2^N$:1多路复用器(MUX)的布尔方程,并在训练期间优化其输入参数。我们使用MIT-BIH数据集研究了LGN和LUTN在患者间心电图心律失常分类中的适用性。所提出的模型在四类任务上实现了高达94.41%的准确率和0.683的$j\kappa$指数,与现有的基于CNN、SVM和SNN的方法相比表现出具有竞争力的性能。我们的LGN和LUTN仅需要估计2.89k到6.17k FLOPs(包括预处理和读出),这比最先进的方法少了三到六个数量级。我们通过在Xilinx Zynq-7000 ZedBoard上实现我们的设计来验证它,该设计由预处理流水线和6-LUTN分类器组成。完整系统每次推理消耗8.25 $\mu$J的动态能量,其中只有0.46 nJ被LUTN分类器使用。这些结果表明,LGN和LUTN都可以用作轻量级硬件分类器,用于患者间心电图心律失常分类。
cs.LG / 181 / 2609.32861

Muon Under Gradient Noise and the Limits of Orthogonalization Near Optima

梯度噪声下的 Muon 与最优点附近正交化的局限
Xie, Xiaohui
Abstract
Muon replaces the momentum buffer of each weight matrix by its orthogonal polar factor. We ask what this orthogonalization does near an optimum, where minibatch noise dominates the gradient. Under a Gaussian noise model, the expected Muon update becomes a scaled gradient step, so a linear method with a suitably matched learning rate reproduces Muon's first-order mean response. The stochastic update is a different matter: after the response is matched, Muon retains a nonlinear residual that is uncorrelated with the input noise and contributes additional covariance. A Hermite expansion shows how momentum acts on this residual. Its higher-order components decorrelate faster than the linear component, so momentum suppresses the residual's accumulated covariance relative to the linear part, but never eliminates it. In a local quadratic surrogate that evaluates the residual on the stationary noise buffer, the residual adds stationary covariance and raises the stationary loss floor at every stable step size, while leaving the contraction dynamics unchanged. Simulations of the full nonlinear recursion on quadratics and measurements on frozen transformer gradients support each step of this picture. Together, the results make a theoretical case for replacing orthogonalization by response-matched momentum SGD once optimization becomes noise-dominated.
Chinese Translation
Muon 将每个权重矩阵的动量缓冲区替换为其正交极因子。我们探究这种正交化在最优点附近的作用,此时小批量噪声主导梯度。在高斯噪声模型下,期望的 Muon 更新变为一个缩放梯度步,因此具有适当匹配学习率的线性方法可以再现 Muon 的一阶平均响应。随机更新则不同:在响应匹配后,Muon 保留了一个与输入噪声不相关的非线性残差,并贡献额外的协方差。埃尔米特展开展示了动量如何作用于该残差。其高阶分量比线性分量去相关更快,因此动量相对于线性部分抑制了残差的累积协方差,但从未消除它。在评估平稳噪声缓冲区上残差的局部二次代理中,残差增加了平稳协方差,并在每个稳定步长下提高了平稳损失下界,同时保持收缩动力学不变。在二次函数上对完全非线性递推的模拟以及对冻结的 Transformer 梯度的测量支持了这一图景的每一步。总之,这些结果为一旦优化变为噪声主导时,用响应匹配的动量 SGD 替代正交化提供了理论依据。
cs.LG / 182 / 2609.32864

Staying on the Attractor: Supervising Neural Surrogates of 3D Turbulence Where They Leave It

保持在吸引子上:在三维湍流神经代理离开吸引子之处对其进行监督
Dai, Yilong, Mitra, Shaswata, Patel, Raj, Sun, Yiming, Chen, Shengyu, Gong, Jiaqi, Mittal, Sudip, Rahimi, Shahram, Jia, Xiaowei, Yu, Runlong
Abstract
Neural surrogates are trained to predict 3D turbulent flows in place of direct numerical simulation (DNS). For chaotic flows, the goal is short-term pointwise accuracy followed by long-term physical and statistical fidelity. However, small prediction errors can carry a surrogate away from the flow's attractor. Off-attractor states are poorly represented in training data, leaving their evolution weakly constrained. The learned dynamics can then amplify deviations and lead to blow-up, freezing, or statistical drift. We propose off-attractor supervision (OAS) to supervise neural surrogates where they leave the attractor. OAS teaches the model how the true Navier-Stokes dynamics would evolve from these states. Each selected state is paired with its own future computed by DNS. Three generators select a few hundred states for relabeling. The first collects states from the surrogate's own rollouts. The second uses surrogate attacks to target freezing, excessive amplification, and violations of incompressibility and energy balance. The third perturbs training states along an amplified direction and a strongly damped random direction of the dynamics. All attacks run on the surrogate alone, and DNS relabeling is performed offline once per selected state. Experiments on $128^3$ turbulence show that OAS increases the median time to failure from 21 to 721 steps. The compared baselines achieve medians of at most 110 steps, and the advantage holds across training seeds. OAS also achieves the lowest pointwise error at step 15 and the best long-horizon statistics among the compared methods. OAS integrates physical models into neural simulation by extending supervision from fixed reference trajectories to states where the surrogate is likely to fail. This principle can guide the development of more reliable scientific surrogates when deployment takes models beyond the coverage of their training data.
Chinese Translation
神经代理被训练来预测三维湍流流动,以替代直接数值模拟(DNS)。对于混沌流动,目标是短期逐点精度,随后是长期物理和统计保真度。然而,小的预测误差可能将代理带离流动的吸引子。离吸引子状态在训练数据中表示不足,导致其演化受到弱约束。学习到的动力学随后可能放大偏差,并导致爆炸、冻结或统计漂移。我们提出离吸引子监督(OAS),在神经代理离开吸引子之处对其进行监督。OAS教会模型真实的纳维-斯托克斯动力学如何从这些状态演化。每个选定的状态与其由DNS计算的自身未来配对。三个生成器选择数百个状态进行重新标记。第一个从代理自身的滚动预测中收集状态。第二个使用代理攻击来针对冻结、过度放大以及不可压缩性和能量平衡的违反。第三个沿动力学的放大方向和强阻尼随机方向扰动训练状态。所有攻击仅在代理上运行,DNS重新标记对每个选定状态离线执行一次。在$128^3$湍流上的实验表明,OAS将失效时间中位数从21步增加到721步。所比较的基线方法中位数最多为110步,且该优势在不同训练种子下保持。OAS还在第15步实现了最低的逐点误差,并在所比较的方法中获得了最佳的长时程统计。OAS通过将监督从固定参考轨迹扩展到代理可能失败的状态,将物理模型集成到神经模拟中。当部署将模型带出其训练数据覆盖范围时,这一原则可以指导开发更可靠的科学代理。
cs.LG / 183 / 2609.32888

Refreshing Less, Selecting Better: Reusing Stale Gradient Features for Efficient Influence-Based Data Selection

更少刷新,更好选择:重用陈旧梯度特征以实现高效的基于影响的数据选择
Su, Jianchang, Zhang, Yifan, Zhang, Wei
Abstract
Gradient-based data selection methods such as LESS score each candidate by the alignment between its gradient and a target validation gradient, and recomputing per-example gradient features at every new checkpoint dominates their cost. Across three selection seeds, two model families, two candidate pools, and two target tasks, features cached at a post-warmup checkpoint and paired with fresh validation gradients preserve the ranking 40 optimizer steps later with Spearman correlation from 0.952 to 0.991, while the top-10% subset they induce misses 10 to 22% of the examples that full recomputation selects. We therefore propose Cached Diverse Influence Selection (CDIS), which recomputes gradient features for the top-ranked fraction $p$ of candidates under the stale scores, fits an affine calibration on the recomputed examples, and selects the final subset under source and length quotas. A refresh fraction at or above the selection fraction recovers the exact top-$k$ subset whenever the calibrated stale scores have bounded error, and the budget curves follow this rule: at $p=0.3$ the recovered top-$k$ subset coincides with full recomputation in every setting, allocating the same budget per stratum recovers the stratified subset at 0.92 to 1.00, and gradient-stage wall-clock drops 3.5 to 3.6 times. Iterating the cache over four checkpoints keeps top-$k$ overlap at 0.98 or higher at 1.9 gradient features per example against 4 for full recomputation, and stale-to-recomputed agreement on the refreshed examples provides a free check for unsafe reuse. Downstream, unconstrained top-$k$ selection collapses to a single data source and scores 13 points below random selection on GSM8K. CDIS scores 12 points above random selection with paired confidence intervals that exclude zero and trails full recomputation by 4.3 points, one training-run standard deviation, at 3.4 times lower selection cost.
Chinese Translation
基于梯度的数据选择方法(如LESS)通过每个候选样本的梯度与目标验证梯度之间的对齐度进行评分,而在每个新检查点重新计算每个样本的梯度特征主导了其成本。在三个选择种子、两个模型家族、两个候选池和两个目标任务上,缓存于预热后检查点的特征并与新鲜验证梯度配对,在40个优化器步骤后仍保持排序,Spearman相关系数为0.952至0.991,而它们诱导的前10%子集遗漏了完全重计算所选示例的10%至22%。因此,我们提出缓存多样化影响选择(CDIS),它根据陈旧分数对排名前$p$比例的候选样本重新计算梯度特征,对重新计算的样本拟合仿射校准,并在源配额和长度配额下选择最终子集。当校准后的陈旧分数具有有界误差时,刷新比例等于或高于选择比例可以恢复精确的top-$k$子集,并且预算曲线遵循这一规则:当$p=0.3$时,恢复的top-$k$子集在所有设置中与完全重计算一致,按层分配相同预算可以以0.92至1.00恢复分层子集,梯度阶段的挂钟时间下降3.5至3.6倍。在四个检查点上迭代缓存,使top-$k$重叠保持在0.98或更高,每个样本1.9个梯度特征,而完全重计算为4个,并且刷新样本上陈旧与重计算之间的一致性为不安全重用提供了免费检查。在下游任务中,无约束的top-$k$选择坍缩到单一数据源,在GSM8K上比随机选择低13分。CDIS比随机选择高12分,配对置信区间不包括零,并以3.4倍较低的选择成本落后完全重计算4.3分(一个训练运行标准差)。
cs.LG / 184 / 2609.32889

Measurement-Gated Provenance Attenuation for Frozen EEG Representations

用于冻结EEG表示的测量门控来源衰减
Aimoldin, Anuar, Chen, Yankai, Mussabayeva, Ayana, Akhanov, Nurdaulet, Liu, Xue
Abstract
Frozen EEG representations retain acquisition signatures as well as neural activity. Source predictability alone does not identify what should be removed: it can reflect measurement effects or genuine biological and population differences, which should not be erased. We propose Measurement-Gated Provenance Attenuation (MGPA), built on one principle: measurement evidence determines where correction may act, and preserved information determines what it should aim for. Paired measurement contrasts define a gate outside which nothing changes; inside it, the source score is moved to the value the preserved coordinates already predict: for a fixed affine score, this keeps the same information as any target set by those coordinates and needs the least expected squared movement. Closed-form and critic-guided iterative constructions apply it without source identity or encoder retraining. Three studies test the principle at increasing distance from its assumptions. Under controlled reference changes, where the source-task association is known, MGPA brings source to near chance with task performance unchanged, whereas erasing what predicts source (LEACE) lowers frozen-task AUROC from .753 to .656 while barely touching source; ablations attribute the attenuation to the gate's directions and 2.7x less movement to the conditional target. Across recordings from different devices and electrodes, iterative correction lowers source accessibility while preserving or improving task performance. Finally, one iterative map selected on one task and reused unchanged on existing heads for two others raises their worst-association AUROC (lowest over device-label shifts) by .057 and .019 over LEACE, at a cost to those heads while the training association holds. A reusable correction shows its value in how an existing predictor behaves once acquisition cues stop being reliable, not only in what a probe can read.
Chinese Translation
冻结的EEG表示不仅保留神经活动,还保留采集特征。仅凭来源可预测性并不能确定应该移除什么:它可能反映测量效应或真实的生物学和群体差异,这些不应被抹去。我们提出测量门控来源衰减(MGPA),其基于一个原则:测量证据决定校正可以在何处起作用,而保留的信息决定其目标应该是什么。配对测量对比定义了一个门,在门外什么都不会改变;在门内,来源分数被移动到保留坐标已经预测的值:对于固定的仿射分数,这保持了与那些坐标设定的任何目标相同的信息,并且需要最小的期望平方移动。闭式解和评论家引导的迭代构造可以在不需要来源身份或重新训练编码器的情况下应用它。三项研究在与其假设距离越来越远的情况下测试该原则。在受控参考变化下,当来源-任务关联已知时,MGPA将来源降至接近随机水平而任务性能不变,而擦除预测来源的信息(LEACE)将冻结任务的AUROC从0.753降低到0.656,同时几乎不影响来源;消融实验将衰减归因于门的方向,以及条件目标上减少2.7倍的移动。在来自不同设备和电极的记录中,迭代校正降低了来源可访问性,同时保持或提高了任务性能。最后,在一个任务上选择的一个迭代映射,并在另外两个任务的现有头上保持不变地重用,将其最差关联AUROC(在设备标签偏移上的最低值)相对于LEACE分别提高了0.057和0.019,在训练关联保持的情况下,对这些头造成了代价。一个可重用的校正的价值体现在当采集线索不再可靠时,现有预测器的行为方式上,而不仅仅体现在探针可以读取的内容上。
cs.LG / 185 / 2609.32890

A Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired Benchmarks

函数级漏洞评分更多衡量的是标记率而非模型:协议对配对基准的影响
Cichoń, Maciej, Dmitruk, Bartłomiej
Abstract
Language models are increasingly evaluated as vulnerability detectors, and scores reported for similar models differ widely between papers. We measured how much of that difference evaluation protocol accounts for, with model outputs held fixed. In a paired test, a model must flag a vulnerable function and clear its version after a fixing commit. Three choices that published evaluations make differently were varied one at a time: metric, verdict extraction and output budget. Seven frontier and large open models were evaluated on five released pair benchmarks and a set pooled for this work under one protocol, and 61 open models of 1.5B to 36B parameters on the pooled set. Function-level F1 follows how often a model flags both functions of a pair (Spearman $+0.86$ over 42 combinations) and is nearly unrelated to pair-level correctness ($+0.16$). On the pair score, extraction changes a model's number by $+0.001$ at the median and budget by $+0.02$ with an interval through zero, whereas the model changes a benchmark's number by up to 0.18 and the benchmark a model's by up to 0.16; a function-level score therefore measures flag rate more than model. For 37 of 68 models the difference between correct and reversed pairs is within its 95% interval of zero, the value for a null model that flags each function at a fixed rate, while both-flagged and both-cleared rates exceed that null by 0.055 on median, and for 64 of 68 both functions of a pair receive one answer more often than independence predicts: verdicts are determined by the text common to both functions. On length-matched pairs a linear probe on activations separates 0.78 by within-pair ranking, against 0.64 for a tf-idf baseline and 0.5 for length; the generated verdict is near chance for three of six models and at 0.55 to 0.57 for the other three, and a prompted logit is at chance for all six.
Chinese Translation
语言模型越来越多地被评估为漏洞检测器,而相似模型在不同论文中报告的分数差异很大。我们测量了评估协议在多大程度上解释了这种差异,同时保持模型输出固定。在配对测试中,模型必须标记一个易受攻击的函数,并在修复提交后清除其版本。已发表评估中做出不同选择的三个选项被一次改变一个:度量、判定提取和输出预算。七个前沿和大型开源模型在五个已发布的配对基准和一个为此工作汇集的数据集上以一个协议进行评估,以及61个1.5B到36B参数的开源模型在汇集集上。函数级F1遵循模型标记一对中两个函数的频率(在42个组合上Spearman +0.86),并且与对级正确性几乎无关(+0.16)。在对分数上,提取使模型的数量中位数变化+0.001,预算变化+0.02,区间穿过零,而模型使基准的数量变化最多0.18,基准使模型的数量变化最多0.16;因此函数级分数衡量标记率多于模型。对于68个模型中的37个,正确对和反转对之间的差异在其95%区间内为零,这是空模型(以固定速率标记每个函数)的值,而双标记率和双清除率超过该空模型中位数0.055,并且对于68个中的64个,一对中的两个函数比独立性预测更频繁地收到一个答案:判定由两个函数共有的文本决定。在长度匹配的对上,激活上的线性探针通过对内排序达到0.78的分离度,而tf-idf基线为0.64,长度为0.5;生成的判定对于六个模型中的三个接近随机,其他三个为0.55到0.57,提示的logit对所有六个都接近随机。
cs.LG / 186 / 2609.32897

Optimizing H-Graph Hybridization for Diffusion-Guided RRT

优化面向扩散引导RRT的H-Graph混合
Talmi, Omer
Abstract
Sampling-based motion planners guided by diffusion models produce high-quality trajectories in a single run, yet the stochastic diversity available at inference time is left largely unexploited. We present two inference-time diversification strategies for a fixed, pretrained DiTree model, combined via H-Graph hybridization, and evaluate them on a holonomic AntMaze robot across 15 maze scenarios. The first, factorial diversity, sweeps the random seed and Diffusion Goal Bias (DGB) parameter, the second, refinement-only diversity, sweeps the diffusion refinement strength (RS) that controls how much an RRT-generated trajectory is edited. Because a single-run baseline only partially succeeds, we additionally compare H-Graph results with pool-based statistics. H-Graph improves the mean pool length of the factorial and refinement-only diversities by 18.8% and 14.5%, respectively. In addition, it also improves the best individual candidate's lengths by 9.7% and 6.8%, respectively. And last, compared with the successful baseline's trajectory length, it improves the results by 18.2% and 19.9%, respectively. These results show that inference-time parameter variation is a reliable, training-free source of path diversity, and that H-Graph hybridization reliably converts this diversity into shorter, higher quality trajectories.
Chinese Translation
由扩散模型引导的基于采样的运动规划器在单次运行中产生高质量的轨迹,然而在推理时可用的随机多样性在很大程度上未被利用。我们针对一个固定的、预训练的 DiTree 模型提出了两种推理时多样化策略,并通过 H-Graph 混合进行组合,并在 15 个迷宫场景中对完整约束的 AntMaze 机器人进行评估。第一种是因子多样化,它扫描随机种子和扩散目标偏置(DGB)参数;第二种是仅细化多样化,它扫描控制 RRT 生成轨迹编辑程度的扩散细化强度(RS)。由于单次运行基线仅部分成功,我们额外将 H-Graph 结果与基于池的统计进行比较。H-Graph 将因子多样化和仅细化多样化的平均池长度分别提高了 18.8% 和 14.5%。此外,它还将最佳单个候选的长度分别提高了 9.7% 和 6.8%。最后,与成功基线的轨迹长度相比,它将结果分别提高了 18.2% 和 19.9%。这些结果表明,推理时参数变化是一种可靠的、无需训练的路径多样性来源,并且 H-Graph 混合能够可靠地将这种多样性转化为更短、更高质量的轨迹。
cs.LG / 187 / 2609.32898

When Less Compute Is More: Adaptive Early Exit Improves Pretrained Outlier Detection

当更少计算意味着更多:自适应提前退出提升预训练异常检测
Zhou, Tianyang, Akoglu, Leman
Abstract
Pretrained tabular foundation models process every dataset at a fixed depth, with inference costs growing with dataset size. To address this, we present the first study of depth-adaptive early-exit for pretrained outlier detection models. While early-exit is typically motivated by efficiency, we uncover a surprising benefit: exiting at the optimal intermediate layer can also improve detection performance on diverse real-world benchmarks by 4.7-7.3% on average, consistent across three distinct foundation models. First, we investigate the factors driving these gains, and identify a key mechanism: context pollution, i.e., the presence of outliers among in-context samples. Our analysis reveals that nearby in-context samples exert increasing influence on query predictions at greater depths, consistent with a retrieval-based view of these models. In effect, early-exit alleviates the adverse effects of retrieving accurate-yet-polluted neighbors, with gains of 13-21% when context pollution matches the natural outlier rate. Motivated by these findings, we pretrain a plug-in router to select a dataset-specific exit layer, using query outlier labels as privileged information available only during router training. The router operates post hoc, leaving the base model parameters and prediction head unchanged. Experiments on three large real-world benchmarks show that, on clean context, the router recovers up to 45% of the oracle gain with up to 1.8x speedup across three pretrained backbones, with larger gains as context pollution increases.
Chinese Translation
预训练表格基础模型以固定深度处理每个数据集,推理成本随数据集大小增长。为解决此问题,我们首次研究了用于预训练异常检测模型的深度自适应提前退出。虽然提前退出通常出于效率考虑,但我们发现了一个令人惊讶的好处:在最优中间层退出还可以在多样化的真实世界基准上平均提升检测性能 4.7-7.3%,并且在三个不同的基础模型上保持一致。首先,我们研究了驱动这些增益的因素,并确定了一个关键机制:上下文污染,即上下文样本中存在异常值。我们的分析表明,在更深的层中,附近的上下文样本对查询预测的影响越来越大,这与这些模型基于检索的观点一致。实际上,提前退出减轻了检索准确但受污染的邻居的不利影响,当上下文污染与自然异常率匹配时,增益为 13-21%。受这些发现的启发,我们预训练了一个即插即用路由器,以选择特定数据集的退出层,使用查询异常标签作为仅在路由器训练期间可用的特权信息。路由器事后操作,保持基础模型参数和预测头不变。在三个大型真实世界基准上的实验表明,在干净上下文上,路由器在三个预训练骨干网络上恢复了高达 45% 的 oracle 增益,并实现了高达 1.8 倍的加速,随着上下文污染的增加,增益更大。
cs.LG / 188 / 2609.32908

Predicting the Next State Is Not Enough: JEPA Representations for Lean Theorem Proving

预测下一状态是不够的:用于Lean定理证明的JEPA表示
Choudhary, Aarnav
Abstract
Neural theorem provers must both propose tactics and decide which valid successor states to explore. We study whether one-step Lean transitions provide a self-supervised signal for branch ordering. A JEPA-style model predicts latent successor representations and scores only kernel-validated, nonterminal successors generated by a fixed pretrained ByT5 proposer. JEPA achieves higher Top-1 than matched InfoNCE on a same-theorem ranking diagnostic(50.18% versus 31.55%), but averages 282.3 of 987 solved theorems across three seeds versus 308 for proposer ordering, while requiring more tactic checks. In this setting, accurate one-step transition ranking is therefore insufficient as a long-horizon search value. The controlled evaluation separates representation from proposal quality and treats kernel-checked proof completion as the primary endpoint.
Chinese Translation
神经定理证明器必须既提出策略,又决定探索哪些有效的后继状态。我们研究一步Lean转换是否能为分支排序提供自监督信号。一个JEPA风格的模型预测潜在后继表示,并仅对由固定预训练ByT5提议器生成、经内核验证的非终端后继进行评分。JEPA在同一定理排序诊断上比匹配的InfoNCE获得更高的Top-1(50.18%对31.55%),但在三个随机种子下平均解决了987个定理中的282.3个,而提议器排序为308个,同时需要更多的策略检查。因此,在此设定下,准确的一步转换排序不足以作为长时程搜索价值。受控评估将表示与提议质量分离,并将内核检查的证明完成作为主要终点。
cs.LG / 189 / 2609.32913

Allspark: Weak to Strong Transfer via Alternating Chain of Thought

Allspark: 通过交替思维链实现弱到强迁移
Liang, Kaizhao, Wang, Junxiong, Liang, Chen, Wang, Zhendong, Liu, Qiang
Abstract
Recent progress in frontier models has renewed interest in large-scale reinforcement learning (RL), but the cost of generating large-model rollouts makes even testing RL recipes expensive. We ask whether reasoning improvements learned by a small, weak model can benefit a larger, stronger model without using the strong model's rollouts during training. We introduce Allspark, a training and inference framework for weak-to-strong transfer through alternating chains of thought. A weak teacher is trained alongside a frozen copy of the same model; the two alternate reasoning segments, and the frozen model produces the final answer. At inference time, a stronger student replaces the frozen training partner, while both models remain fixed. Because they communicate through text, the teacher can steer students from different model families and with different tokenizers. We study Allspark at two scales: controlled Qwen experiments across math and reasoning, and larger-scale Inkling experiments on ARC-AGI-2. The Inkling experiments show accuracy gains in within-family and cross-family settings, including transfer to Kimi and Nemotron, with benefits that vary across inference settings. These findings motivate reusing a trained weak teacher across strong students and examining the resulting accuracy--token tradeoff.
Chinese Translation
前沿模型的最新进展重新激发了人们对大规模强化学习(RL)的兴趣,但生成大模型 rollout 的成本使得即使测试 RL 配方也变得昂贵。我们探究一个小而弱的模型学到的推理改进能否在不使用强模型 rollout 的情况下使一个更大、更强的模型受益。我们提出 Allspark,一个通过交替思维链实现弱到强迁移的训练和推理框架。一个弱教师模型与同一模型的冻结副本一起训练;两者交替进行推理片段,并由冻结模型产生最终答案。在推理时,一个更强的学生模型替换冻结的训练伙伴,而两个模型都保持固定。由于它们通过文本进行交流,教师可以引导来自不同模型家族且具有不同分词器的学生。我们在两个规模上研究 Allspark:在数学和推理任务上受控的 Qwen 实验,以及在 ARC-AGI-2 上更大规模的 Inkling 实验。Inkling 实验显示在家族内和跨家族设置中准确率提升,包括迁移到 Kimi 和 Nemotron,其收益随推理设置的不同而变化。这些发现促使我们复用训练好的弱教师模型到多个强学生模型,并研究由此产生的准确率- token 权衡。
cs.LG / 190 / 2609.32921

Adaptive Latent Capacity for World Models

世界模型的自适应潜在容量
Achituve, Idan, Dikstein, Lior, Diamant, Idit, Netzer, Arnon, Habi, Hai Victor
Abstract
We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over prefix lengths and trains the predictor to estimate the full next embedding from a sampled input prefix. As standard anti-collapse objectives encourage variation across latent coordinates and do not organize them by predictive importance, we also introduce MixSIGReg. MixSIGReg regularizes the masked embeddings against a prior-weighted mixture with Gaussian active prefixes and zeros in the remaining coordinates. As a result, the ALeWM objective encourages early coordinates to retain information useful for prediction and recursive planning. Our analysis shows that the mixture distribution used by MixSIGReg assigns higher variance to earlier coordinate blocks and lower variance to later ones. In addition, we show that, under specified assumptions, prediction error is minimized by placing the information most useful for prediction in earlier blocks. Empirically, we study the behavior of ALeWM in a controlled dynamical system with known state variables and in goal-conditioned visual control. We show that ALeWM consistently achieves higher mean success rates than tuned fixed-width LeWM, with lower planning capacity on average.
Chinese Translation
我们介绍了自适应LeWorldModel(ALeWM),一种基于联合嵌入预测架构(JEPA)的世界模型,它学习将预测信息集中在宽潜在表示的紧凑前缀中。为了鼓励这种排序,ALeWM学习一个以序列为条件的关于前缀长度的分布,并训练预测器从采样的输入前缀中估计完整的下一嵌入。由于标准的抗崩溃目标鼓励潜在坐标间的变化,并且不按预测重要性组织它们,我们还引入了MixSIGReg。MixSIGReg对掩码嵌入进行正则化,使其符合一个先验加权的混合分布,其中活跃前缀服从高斯分布,而剩余坐标为零。因此,ALeWM的目标鼓励早期坐标保留对预测和递归规划有用的信息。我们的分析表明,MixSIGReg使用的混合分布为较早的坐标块分配较高的方差,而为较晚的坐标块分配较低的方差。此外,我们表明,在特定假设下,将最有助于预测的信息置于较早的块中可使预测误差最小化。在实证上,我们在具有已知状态变量的受控动力系统和目标条件视觉控制中研究了ALeWM的行为。我们表明,ALeWM始终比经过调优的固定宽度LeWM获得更高的平均成功率,且平均规划容量更低。
cs.LG / 191 / 2609.32927

The Geometry of Logic: Stratification Induces Semantic Structure and Robust Reasoning

逻辑的几何:分层诱导语义结构与鲁棒推理
Lopes, Cristina V., Li, Yuangang, Chen, Justin Tian Jin, Krone-Martins, Alberto, Ma, Iris, Misu, Md Rakib Hossain
Abstract
Transformer-based language models perform well on symbolic tasks, yet it remains unclear whether they learn generalizable rules or rely on statistical shortcuts. Mechanistic studies link algorithmic behavior to structured internal representations, motivating the hypothesis that robust reasoning benefits from separating values from the types that control their manipulation. Can making this separation an architectural primitive improve the learnability and generalization of logical mechanisms? We introduce \textbf{STRAT} (\textbf{ST}ratified \textbf{R}egisters \textbf{A}nd \textbf{T}ypes), which partitions the residual stream into orthogonal Data and Type subspaces and uses Type-based attention and gating to govern Data transformations. Controlled arithmetic ablations identify three failure modes associated with data-control interference: the Linear Trap, Gradient Wall, and Open Gate Trap. Mechanistic analysis reveals interpretable logical structure, and in arithmetic, STRAT reduces median OOD error 35-fold relative to a Transformer baseline. On each of 11 datasets spanning 10 tasks, STRAT outperforms the Transformer baseline in mean accuracy, by 26 percentage points on average, with both models trained from 10 base examples per dataset using identical task-specific augmentation where applicable. Under distribution shift, STRAT's mean accuracy drops by only 2.39 percentage points, compared with 11.75 for the Transformer.
Chinese Translation
基于Transformer的语言模型在符号任务上表现良好,但其究竟学习到可泛化的规则还是依赖统计捷径仍不清楚。机制性研究将算法行为与结构化内部表示联系起来,促使人们假设:鲁棒推理受益于将值与控制其操作的类型分离。将这种分离作为架构原语,能否提升逻辑机制的可学习性与泛化能力?我们提出STRAT(Stratified Registers And Types),它将残差流划分为正交的Data和Type子空间,并使用基于Type的注意力与门控来支配Data变换。受控算术消融实验识别出与数据-控制干扰相关的三种失败模式:线性陷阱(Linear Trap)、梯度墙(Gradient Wall)和开放门陷阱(Open Gate Trap)。机制分析揭示了可解释的逻辑结构;在算术任务中,STRAT相对于Transformer基线将OOD误差中位数降低了35倍。在涵盖10个任务的11个数据集中的每一个上,STRAT在平均准确率上均优于Transformer基线,平均高出26个百分点;两个模型均从每个数据集的10个基础示例训练,并在适用时使用相同的任务特定数据增强。在分布偏移下,STRAT的平均准确率仅下降2.39个百分点,而Transformer下降了11.75个百分点。
cs.LG / 192 / 2609.32928

Phenomenon-Graph JEPA: Label-Efficient Representation Learning for Contactless Cardiorespiratory Sensing

Phenomenon-Graph JEPA:面向非接触式心肺传感的标签高效表示学习
Casado, Constantino Álvarez, Nguyen, Nhi, Rahman, Mohammad Rakibur, Nguyen, Le, Cañellas, Manuel Lage, Sharifipour, Sasan, López, Miguel Bordallo
Abstract
Millimeter-wave (mmWave) radar and RGB-D cameras can record cardiac and respiratory waveforms continuously and without contact, but labeled recordings remain scarce because every label requires a supervised acquisition session. Self-supervised pretraining can exploit the unlabeled signals, yet contrastive methods depend on signal transformations and negative pairs whose validity is uncertain for cardiorespiratory data, where time warping changes breathing rate and distant windows can share the same physiological state. We present Phenomenon-Graph JEPA, a joint-embedding predictive architecture that learns from four processed one-dimensional streams without negative pairs or synthetic augmentation in its base configuration. Each stream is encoded by a temporal convolutional branch and a band-limited spectral branch. During pretraining, the model predicts stopped target embeddings along typed edges, which connect streams assigned to the same physiological phenomenon, and forward in time within a state episode. We treat this physiological typing as a testable hypothesis and compare it with wrong-edge and all-pairs prediction graphs. In the OMuSense-23 dataset, pretraining improves label-efficiency area over matched supervised training by 3.91 percentage points (95% interval 2.08 to 5.80, Holm-adjusted p = 0.006), and by 3.74 points under a second configuration evaluated on the same test participants. However, the wrong-edge and all-pairs controls do not establish a benefit from physiological typing. Optional Takens-inspired delay coordinates improve a validation comparison with learned history, whereas two wrist-only WESAD protocols do not establish a pretraining advantage. The study therefore separates the measured benefit of predictive representations from the physiological prior used to organize their training.
Chinese Translation
毫米波(mmWave)雷达和RGB-D相机可以连续且非接触地记录心脏和呼吸波形,但标注记录仍然稀缺,因为每个标签都需要一次监督采集会话。自监督预训练可以利用未标注信号,然而对比方法依赖于信号变换和负样本对,这些对于心肺数据而言其有效性是不确定的,因为时间扭曲会改变呼吸率,而远距离窗口可能共享相同的生理状态。我们提出Phenomenon-Graph JEPA,一种联合嵌入预测架构,在其基础配置中从四个处理后的一维流中学习,而不使用负样本对或合成增强。每个流由一个时间卷积分支和一个带限频谱分支进行编码。在预训练期间,模型沿类型化边预测停止的目标嵌入,这些边连接分配给相同生理现象的流,并在状态片段内沿时间向前预测。我们将这种生理类型化视为一个可检验的假设,并将其与错误边和全对预测图进行比较。在OMuSense-23数据集中,预训练将标签效率面积相对于匹配的监督训练提高了3.91个百分点(95%区间2.08至5.80,Holm校正p = 0.006),并且在第二种配置下,对相同测试参与者进行评估时提高了3.74个百分点。然而,错误边和全对控制并未证实生理类型化带来的益处。可选的Takens启发的延迟坐标在与学习历史的验证比较中有所改善,而两个仅手腕的WESAD协议并未确立预训练优势。因此,该研究将预测表示的可测量益处与用于组织其训练的生理先验区分开来。
cs.LG / 193 / 2609.32929

Efficient Dynamic Algorithms for Graph Neural Networks with Non-Linear Propagation

面向非线性传播的图神经网络高效动态算法
Banihashem, Kiarash, Hajiaghayi, MohammadTaghi, JafariRaviz, Mahdi, Lattanzi, Silvio, Mittal, Danny
Abstract
Graph Neural Networks (GNNs) are widely used for representation learning on graphs, but most methods assume static topologies, making them inefficient on evolving networks where edges change over time. Existing dynamic approaches either model graph evolution through temporal GNN architectures without focusing on efficient dynamic maintenance, or are restricted to linear propagation models based on Personalized PageRank. In this work, we study how to efficiently maintain node representations for non-linear GNN propagation under edge insertions and deletions. The propagation has no learned parameters, and only a classifier applied afterward is trained. For a broad class of standard activation functions, we develop a residual-based dynamic algorithm that selectively propagates local errors via push operations, maintaining an approximation to the evolving fixed point without full recomputation. We prove that our method achieves amortized $O(1/\epsilon)$ update time per graph change under a degree-normalized error guarantee. Our approach uses a potential-based analysis in a degree-scaled norm and, in contrast to prior work on the linear case, requires no randomness assumptions on either the update sequence or the input vector. For the linear special case, we additionally provide an exact dynamic algorithm via low-rank matrix inverse updates. Experiments on benchmark datasets show that incorporating non-linearity improves accuracy while preserving efficient update performance, yielding a scalable and theoretically grounded method for maintaining this propagation on dynamic graphs.
Chinese Translation
图神经网络(GNNs)广泛用于图上的表示学习,但大多数方法假设静态拓扑,这使得它们在边随时间演化的网络上效率低下。现有的动态方法要么通过时序GNN架构对图演化建模而不关注高效的动态维护,要么局限于基于个性化PageRank的线性传播模型。在这项工作中,我们研究如何在边插入和删除下高效维护非线性GNN传播的节点表示。该传播没有可学习参数,仅训练之后应用的一个分类器。对于一大类标准激活函数,我们开发了一种基于残差的动态算法,该算法通过推送操作选择性地传播局部误差,维护对演化不动点的近似,而无需完全重新计算。我们证明,在度归一化误差保证下,我们的方法每次图变化达到摊还O(1/ε)更新时间。我们的方法在度缩放范数中使用基于势能的分析,与线性情况下的先前工作相比,不需要对更新序列或输入向量做任何随机性假设。对于线性特殊情况,我们还通过低秩矩阵逆更新提供了精确的动态算法。在基准数据集上的实验表明,引入非线性可以提高精度,同时保持高效的更新性能,从而产生一种可扩展且具有理论依据的方法,用于在动态图上维护这种传播。
cs.LG / 194 / 2609.32932

Counting on Thinking: Tracing Evidence Integration in Language Models

依赖思考:追踪语言模型中的证据整合
Xue, Jingming, Wilson, Robert C., Xiong, Huadong
Abstract
Finite computational resources force a tradeoff between automatic System 1 processes and costly System 2 thinking. Large language models (LLMs) can spend extra computation on hard problems, yet direct answers struggle even with counting, an elementary operation humans and animals perform automatically. We ask why this requires thinking in LLMs. Evidence integration has long been used in psychology and neuroscience to probe decision-making. Our evidence-integration task presents one letter per conversational turn and asks which of two target letters appeared more often. A running count difference solves the task optimally by weighting every letter equally; tokens at each turn could represent and update this difference. Direct responses instead weighted evidence unevenly, with strong recency effects, and assigned less probability to the correct answer as difficulty increased. Thinking improved performance and made integration weights nearly uniform, yet final-query attention remained concentrated on the sequence ends in both modes. Reasoning trajectories showed models revisiting input, recounting letters, and checking intermediate counts that informed the answer, suggesting that thinking constructs the accumulated count that direct responses lack rather than reading out one already formed. Reasoning-token costs grew with the number of letters far more than with coherence. Outcome feedback did not bring this computation into direct responses: under in-context reinforcement learning (ICRL), performance deteriorated over repeated games and recency effects strengthened, yet models grew more confident. Humans and animals amortize such computations into automatic processes, whereas current LLMs still pay for them with thinking on every trial. Which operations learning can make directly available remains central to how future models allocate computation.
Chinese Translation
有限的计算资源迫使我们在自动的 System 1 过程和代价高昂的 System 2 思维之间进行权衡。大型语言模型(LLM)可以在难题上花费额外计算,然而即使是计数这种人类和动物自动执行的基本操作,直接回答也难以应对。我们追问为何这在 LLM 中需要思考。证据整合长期以来在心理学和神经科学中被用于探究决策。我们的证据整合任务在每轮对话中呈现一个字母,并询问两个目标字母中哪一个出现得更频繁。一个运行中的计数差值通过平等地权衡每个字母来最优地解决该任务;每一轮的 token 都可以表示并更新这个差值。直接回答则对证据的加权不均匀,存在强烈的近因效应,并且随着难度增加,赋予正确答案的概率更低。思考提高了性能并使整合权重几乎均匀,然而在两种模式下,最终查询的注意力仍然集中在序列两端。推理轨迹显示模型会重新审视输入、重新计数字母并检查为答案提供信息的中间计数,这表明思考构建了直接回答所缺乏的累积计数,而不是读出已经形成的计数。推理 token 成本随字母数量的增长远大于随连贯性的增长。结果反馈并未将这种计算带入直接回答:在上下文强化学习(ICRL)下,性能在重复游戏中下降,近因效应增强,但模型变得更加自信。人类和动物将此类计算摊销到自动过程中,而当前的 LLM 在每次试验中仍需要用思考来支付这些计算。学习能够使哪些操作直接可用,仍然是未来模型如何分配计算的核心问题。
cs.LG / 195 / 2609.32933

Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs

结构化约束MDP中在线强化学习的最后迭代保证
Tran, Nam Phuong, Huynh, Trinh Ha Mai, Le, Tuyen Pham, Nguyen, Van-Truong, Nguyen, Quan, Tran-Thanh, Long
Abstract
In safety-critical applications, deployment uses a single policy, whose performance and constraint satisfaction should hold directly rather than only for an average or mixture of training policies. This motivates last-iterate guarantees in constrained reinforcement learning. Recent progress has established such guarantees in exact-gradient or tabular online settings, yet scalable results for structured large-state problems remain open. We develop a general, statistically efficient framework for last-iterate convergence in structured Constrained MDPs (CMDPs). Our analysis separates contraction of the regularised primal-dual dynamics from actor approximation and statistical errors in policy evaluation under online exploration. This enables model-free on- and off-policy learning with structured function approximation: optimistic policy evaluation avoids explicit transition-model construction, while a compact parametric actor avoids maintaining mixtures or histories of past policies. We instantiate the framework for linear CMDPs and general function approximation, obtaining representation-dependent complexity and improved target-accuracy dependence over prior optimistic regularised primal-dual analyses. We further validate the stabilising effect predicted by our theory on a synthetic linear CMDP: the regularised method exhibits stable last-iterate behaviour, whereas its unregularised counterpart shows larger oscillations.
Chinese Translation
在安全关键应用中,部署使用单一策略,其性能和约束满足应直接成立,而不仅仅是对于训练策略的平均或混合。这激发了约束强化学习中的最后迭代保证。最近的研究进展已在精确梯度或表格在线设置中建立了此类保证,但针对结构化大状态问题的可扩展结果仍然悬而未决。我们开发了一个通用的、统计高效的框架,用于结构化约束MDP(CMDP)中的最后迭代收敛。我们的分析将正则化原始-对偶动力学的收缩与在线探索下策略评估中的策略近似和统计误差分离开来。这使得具有结构化函数逼近的无模型同策略和异策略学习成为可能:乐观策略评估避免了显式的转移模型构建,而紧凑的参数化策略避免了维护过去策略的混合或历史。我们将该框架实例化用于线性CMDP和一般函数逼近,获得了依赖于表示的复杂度,并改进了先前乐观正则化原始-对偶分析的目标精度依赖性。我们进一步在合成线性CMDP上验证了理论预测的稳定效果:正则化方法表现出稳定的最后迭代行为,而未正则化的对应方法则显示出更大的振荡。
cs.LG / 196 / 2609.32934

The Impact of Stochasticity on the Rashomon Effect in Machine Learning

随机性对机器学习中 Rashomon 效应的影响
Apicella, Andrea, Isgrò, Francesco, Pollastro, Andrea, Prevete, Roberto
Abstract
Neural network training is inherently stochastic, with factors such as weight initialization leading to distinct models despite comparable predictive performance. This phenomenon is commonly associated with the Rashomon effect, which describes the existence of multiple near-optimal models for the same task. Although the Rashomon effect has received increasing attention, it remains unclear whether different sources of training stochasticity contribute similarly or differently to its manifestations. In this work, we present an empirical study of the Rashomon phenomenon along three complementary dimensions: solution-space multiplicity, predictive multiplicity, and decision-basis multiplicity. These dimensions are quantified through the size of the empirical Rashomon set, predictive ambiguity, and agreement between XAI attribution maps, respectively. By independently controlling three standard sources of stochasticity, namely weight initialization, mini-batch data ordering, and dropout, we isolate their respective contributions to each dimension of the Rashomon phenomenon. Experiments on tabular and image classification benchmarks reveal that these sources affect the three dimensions in different ways. In particular, larger empirical Rashomon sets do not necessarily correspond to greater predictive disagreement or lower explanation agreement, indicating that solution-space, predictive, and decision-basis multiplicity capture complementary rather than interchangeable aspects of the Rashomon effect. Overall, our results show that training stochasticity influences not only predictive performance but also the stability of predictions and explanations, highlighting the importance of identifying the specific sources of stochasticity responsible for different manifestations of the Rashomon phenomenon when assessing the reliability, reproducibility, and interpretability of neural network models.
Chinese Translation
神经网络训练本质上是随机的,诸如权重初始化等因素会导致在预测性能相当的情况下产生不同的模型。这一现象通常与 Rashomon 效应相关联,该效应描述了对于同一任务存在多个近乎最优模型的现象。尽管 Rashomon 效应受到越来越多的关注,但尚不清楚不同的训练随机性来源对其表现形式的贡献是相似还是不同。在这项工作中,我们沿着三个互补维度对 Rashomon 现象进行了实证研究:解空间多重性、预测多重性和决策依据多重性。这些维度分别通过经验 Rashomon 集的大小、预测模糊性以及 XAI 归因图之间的一致性来量化。通过独立控制三种标准随机性来源,即权重初始化、小批量数据排序和 dropout,我们分离了它们各自对 Rashomon 现象每个维度的贡献。在表格和图像分类基准上的实验表明,这些来源以不同方式影响这三个维度。特别是,更大的经验 Rashomon 集不一定对应于更大的预测分歧或更低解释一致性,这表明解空间、预测和决策依据多重性捕捉的是 Rashomon 效应的互补而非可互换的方面。总体而言,我们的结果表明,训练随机性不仅影响预测性能,还影响预测和解释的稳定性,强调了在评估神经网络模型的可靠性、可重复性和可解释性时,识别导致 Rashomon 现象不同表现形式的特定随机性来源的重要性。
cs.LG / 197 / 2609.32939

Theory of Scene: Breaking the Symmetry Trap in Multi-Agent LLM Coordination

场景理论:打破多智能体大语言模型协作中的对称性陷阱
Yuan, Liangqi, Fang, Wenzhi, Wang, Shiqiang, Brinton, Christopher G.
Abstract
Multi-agent systems built on large language models (LLMs) are largely homogeneous, as their agents behave alike even across distinct LLMs. We show that when such agents act concurrently without communication, they collide on targets they must split and diverge on targets they must take together, a double failure we term the symmetry trap. Theory of Mind (ToM), widely used for coordination without communication, cannot escape this trap, since homogeneous agents form the same prediction of one another and respond to it in the same way. We propose Theory of Scene (ToS), a training-free reasoning schema in which each agent reads its public role, the only difference between the agents, and the task context they all observe. Homogeneous agents thereby derive one division of labor, each taking the part its role fixes, which turns homogeneity from the cause of the trap into the cure. ToS reads the role together with the scene through role gating, which determines whether ownership overlaps or is already divided, and the task context through task coupling, which infers whether the team must converge on each target, divide it, or take its stages in turn. We evaluate on DivvyBench, a controlled environment we introduce, whose target types make an episode Competitive, Cooperative, or Mixed across Tabletop, Airspace, and Household scenarios, and on two established agentic benchmarks, GovSim and Overcooked. ToS outperforms all six baselines on every benchmark, and each baseline falls far behind it in at least one setting. Against ToM given the same inputs, ToS raises the DivvyBench success rate from 71.1% to 99.6%, the GovSim total gain from 207 to 400, and the Overcooked level-normalized throughput from 1.41 to 1.67.
Chinese Translation
基于大语言模型(LLMs)构建的多智能体系统在很大程度上是同质的,因为其智能体即使来自不同的LLM,行为也相似。我们表明,当此类智能体在无通信情况下并发行动时,它们会在必须分开处理的目标上发生冲突,并在必须共同处理的目标上产生分歧;我们将这一双重失败称为对称性陷阱。心智理论(Theory of Mind, ToM)虽被广泛用于无通信协调,却无法逃脱这一陷阱,因为同质智能体会对彼此形成相同预测,并以相同方式回应。我们提出场景理论(Theory of Scene, ToS),一种无需训练的推理图式:每个智能体读取其公共角色——这是智能体之间唯一的差异——以及它们共同观察到的任务上下文。同质智能体由此推导出同一种劳动分工,每个智能体承担其角色所固定的部分,从而将同质性从陷阱的成因转变为解决之道。ToS通过角色门控将角色与场景一并读取,以判断所有权是重叠还是已划分;并通过任务耦合读取任务上下文,以推断团队必须在每个目标上汇聚、将其划分,还是依次执行其各阶段。我们在我们提出的受控环境DivvyBench,以及两个已有智能体基准GovSim和Overcooked上进行评估;在DivvyBench中,目标类型使一个回合在桌面(Tabletop)、空域(Airspace)和家庭(Household)场景中呈现竞争、合作或混合性质。ToS在所有基准上均优于全部六个基线,且每个基线至少在一个设置中远远落后于它。在给定相同输入的情况下,与ToM相比,ToS将DivvyBench成功率从71.1%提升至99.6%,将GovSim总收益从207提升至400,并将Overcooked的关卡归一化吞吐量从1.41提升至1.67。
cs.LG / 198 / 2609.32940

TwinS-GCN: Spectral conjugate for Spectral Graph Convolutional Networks

TwinS-GCN:用于谱图卷积网络的谱共轭
Chan, Chun Hei Michael, Petruso, Flavia, Van De Ville, Dimitri
Abstract
Graph convolutional networks propagate information by repeated local aggregation through a graph shift operator; i.e., a $K$-layer network reaches $K$ hops neighborhood. On the one hand, such spreading can lead to oversmoothing. On the other hand, long-range dependencies demand the depth. Transporting information on long distances and without attenuation requires the shift to distinguish a direction of flow, which a symmetric operator cannot perform but a directed one can fulfill. A natural way to extract pure-directionality is to take the skew-symmetric part of the shift operator through the Cartesian split, which, however, generally does not commute with the shift itself, meaning that the filters built on it are not shift-invariant. We instead use the spectral conjugate; i.e., the image of the operator under $\tau:z\mapsto \bar{z}$, which commutes with the shift and splits it into a dissipative and a non-dissipative part. Two filter families follow: a sum filter, whose non-dissipative component transports signal without energy loss, and a ratio filter, ratio in the pair of components rather than polynomial in the shift. Both arise from non-holomorphic kernels, placing them outside the holomorphic class underlying classical spectral convolution. On the directed cycle, the ratio filter becomes an IIR filter with global impulse response, for which we prove a long-range reach gap against every degree-$K$ polynomial filter. Chebyshev reparameterization gives stable vertex-domain layers with real coefficients, yielding TwinS-GCN, which solves graph transfer tasks at reduced depth and is competitive with state-of-the-art graph convolutional networks on node classification benchmarks.
Chinese Translation
图卷积网络通过图移位算子重复局部聚合来传播信息;即,一个$K$层网络可以到达$K$跳邻域。一方面,这种传播可能导致过平滑。另一方面,长距离依赖需要深度。在长距离上传输信息且不衰减,需要移位算子能够区分流动方向,这是对称算子无法实现的,而有向算子可以满足。提取纯方向性的自然方法是通过笛卡尔分解取移位算子的斜对称部分,然而,这通常不与移位算子本身可交换,意味着基于它构建的滤波器不是移位不变的。我们转而使用谱共轭;即,算子在$\tau:z\mapsto \bar{z}$下的像,它与移位算子可交换,并将其分解为耗散部分和非耗散部分。由此产生两个滤波器族:一个和滤波器,其非耗散分量无能量损失地传输信号;以及一个比率滤波器,其比率在于分量对中,而不是移位算子的多项式。两者都源于非全纯核,使它们位于经典谱卷积所依赖的全纯类之外。在有向循环上,比率滤波器变为具有全局脉冲响应的IIR滤波器,对此我们证明了相对于每个$K$次多项式滤波器的长距离可达性差距。切比雪夫重参数化给出了具有实系数的稳定顶点域层,产生了TwinS-GCN,它以降低的深度解决图迁移任务,并在节点分类基准上与最先进的图卷积网络具有竞争力。
cs.LG / 199 / 2609.32941

Generative Priors Conditioned on Natural Language for Bayesian Inversion in PDEs

以自然语言为条件的生成先验用于偏微分方程中的贝叶斯反演
Zhang, Pengyu, Girolami, Mark, Vadeboncoeur, Arnaud
Abstract
Inferring quantities of interest (QoI) from data is a central task in Science and Engineering. In such contexts, we often have access to both quantitative data and qualitative data. Quantitative data may be represented by noisy sensor measurements, simulation data, re-analysis data; qualitative data may be in the form of text descriptions of experimental setups, expected experiment outcomes, and human-perceived system behaviours. The task we address in this paper is the following. Given a training set of paired qualitative text and quantitative QoI data, we learn to exploit the inherent correlation between the two modalities to learn a highly informative data-driven natural-language-conditional Bayesian prior, such that when presented with a new physical system, we can coherently combine (i) the training dataset, (ii) qualitative text describing the new system, and (iii) a small number of noisy sensor readings from that new system, to perform inference and uncertainty quantification (UQ) over the QoI. To achieve this task, we develop two parallel approaches, one uses conditional diffusion and the other conditional autoencoders, and compare both against classical Bayesian methodology, unconditional generative models and deterministic supervised methods. Each approach has specific strengths and tradeoffs; conditional autoencoder offers theoretical tractability, allows for fast posterior sampling, and provides better-calibrated UQ, whereas conditional diffusion is explored for greater expressiveness and capturing complex posteriors with irregular QoI fields. The approach is tested on the steady-state heat equation, damped Helmholtz equation, and UK weather reanalysis data.
Chinese Translation
从数据中推断感兴趣量(QoI)是科学与工程中的一项核心任务。在此类场景中,我们通常可以同时获得定量数据和定性数据。定量数据可由含噪传感器测量、模拟数据、再分析数据表示;定性数据可以是对实验设置、预期实验结果以及人类感知的系统行为的文本描述形式。本文处理的任务如下。给定由成对的定性文本和定量 QoI 数据组成的训练集,我们学习利用两种模态之间的内在相关性,以学习一个高度信息化的、数据驱动的、以自然语言为条件的贝叶斯先验,从而在面对一个新的物理系统时,我们能够连贯地结合 (i) 训练数据集,(ii) 描述新系统的定性文本,以及 (iii) 来自该新系统的少量含噪传感器读数,以对 QoI 进行推断和不确定性量化(UQ)。为实现这一任务,我们开发了两种并行方法,一种使用条件扩散,另一种使用条件自编码器,并将二者与经典贝叶斯方法、无条件生成模型和确定性监督方法进行比较。每种方法都有其特定的优势与权衡;条件自编码器提供了理论可处理性,允许快速后验采样,并提供校准更好的 UQ,而条件扩散则被探索用于更强的表达能力和捕获具有不规则 QoI 场的复杂后验。该方法在稳态热方程、阻尼 Helmholtz 方程和英国天气再分析数据上进行了测试。
cs.LG / 200 / 2609.32949

Optimal Nonparametric Dynamic Pricing with Censored Demand and Adversarial Inventory

具有删失需求和对抗性库存的最优非参数动态定价
Zhang, Mengxiao, Wang, Yingfei, Luo, Haipeng
Abstract
We study online dynamic pricing with censored demand, where an arbitrary inventory level is revealed before pricing and may adapt to past observations, while demand follows an unknown, price-dependent distribution that is stationary over time. For a horizon of $T$ rounds, Xu et al. [2026] achieved $\widetilde{\mathcal{O}}(\sqrt{T})$ regret under restrictive structural assumptions including linear demand, price-independent additive noise, and conditions relating inventory levels to the noise support. Our first contribution is to extend this framework to a substantially more general and statistically harder nonparametric setting, requiring only the natural assumption that expected sales are nonincreasing in price and allowing nonlinear demand curves and price-dependent noise. For this model, we first propose a simple baseline, Double-Grid-UCB, which discretizes both price and inventory and achieves $\widetilde{\mathcal{O}}(T^{3/4})$ expected regret using separate revenue estimates for each price-inventory grid pair. Then, we develop Threshold-UCB, which improves the expected regret to $\widetilde{\mathcal{O}}(T^{2/3})$. Unlike Double-Grid-UCB, Threshold-UCB reuses sales observations across inventory levels through shared estimates of demand-tail probabilities, allowing the same data to support revenue upper bounds for multiple inventories rather than a single inventory bin. We also complement this upper bound with an $\Omega(T^{2/3})$ lower bound via a reduction from stochastic posted pricing, establishing its minimax optimality. Finally, extensive experiments across inventory processes, demand functions, and noise models demonstrate consistently superior performance of Threshold-UCB over benchmark algorithms.
Chinese Translation
我们研究具有删失需求的在线动态定价,其中在定价前会揭示一个任意的库存水平,并且该库存可能根据过去的观测进行调整,而需求则遵循一个未知的、依赖于价格的分布,该分布在时间上是平稳的。对于 $T$ 轮的周期,Xu 等人 [2026] 在限制性结构假设下实现了 $\widetilde{\mathcal{O}}(\sqrt{T})$ 的遗憾,这些假设包括线性需求、与价格无关的加性噪声,以及将库存水平与噪声支撑集相关联的条件。我们的第一个贡献是将该框架扩展到一个更为一般且统计上更困难的非参数设定,仅要求期望销售量随价格非递增这一自然假设,并允许非线性需求曲线和依赖于价格的噪声。针对该模型,我们首先提出一个简单的基线算法 Double-Grid-UCB,它同时对价格和库存进行离散化,通过对每个价格-库存网格对使用单独的收益估计,实现了 $\widetilde{\mathcal{O}}(T^{3/4})$ 的期望遗憾。然后,我们提出 Threshold-UCB,将期望遗憾改进至 $\widetilde{\mathcal{O}}(T^{2/3})$。与 Double-Grid-UCB 不同,Threshold-UCB 通过共享的需求尾部概率估计,跨库存水平重用了销售观测,使得同一数据能够为多个库存(而非单个库存区间)提供收益上界。我们还通过从随机标价定价的归约,用 $\Omega(T^{2/3})$ 下界补充了这一上界,从而确立了其极小极大最优性。最后,跨库存过程、需求函数和噪声模型的大量实验表明,Threshold-UCB 相较于基准算法始终具有更优的性能。
cs.LG / 201 / 2609.32952

Constrained Flow Policy Updates: A Generalized Schr\"odinger Bridge View

约束流策略更新:广义Schrödinger桥视角
Li, Boyang, Kim, Matthew, Herbert, Sylvia
Abstract
Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. Reward and safety can induce multimodal action distributions, challenging the prevailing primal-dual methods: Gaussian actors may collapse onto a single suboptimal mode, and optimization over the nonconvex Lagrangian landscape can be unstable. Diffusion and flow policies can represent such distributions, but recent work with a diffusion actor relies on estimating and matching the score of an augmented-Lagrangian target policy. Instead, we differentiate the augmented objective directly through the generation path of a flow policy, so no score needs to be estimated. Because a flow policy lacks a readily available action log-density for entropy regularization, we build on the density-free kinetic-energy regularizer of FLAC, a recent reward-only method, and propose Reparameterized Augmented-Lagrangian Flow Actor with Least Energy (RAFALE), an off-policy actor-critic method for safe RL. We formulate its update as a constrained one-ended generalized Schr\"odinger bridge and show that, for each source draw, this path-space problem is exactly an entropy-regularized problem in action space. At positive noise, its solution reweights the reward-only action distribution only where the estimated cost exceeds a threshold set by the Lagrange multiplier. As the noise vanishes, the optimal value converges to that of a least-energy map objective that the flow policy optimizes directly. Across seven Safety-Gymnasium tasks, RAFALE achieves competitive reward with mean final cost within budget on every task, whereas strong baselines trade one for the other; ablations support the necessity of both its augmented objective and its flow actor.
Chinese Translation
在线安全强化学习(RL)寻求在满足安全约束的同时最大化奖励的策略。奖励和安全性可能引发多模态动作分布,对主流的原始-对偶方法构成挑战:高斯行动者可能坍缩到单一亚优模式,并且对非凸拉格朗日景观的优化可能不稳定。扩散和流策略可以表示此类分布,但最近使用扩散行动者的工作依赖于估计和匹配增广拉格朗日目标策略的得分。相反,我们直接通过流策略的生成路径对增广目标进行求导,因此无需估计得分。由于流策略缺乏易于获取的动作对数密度用于熵正则化,我们基于FLAC(最近的一种仅奖励方法)的无密度动能正则化器,提出了重参数化增广拉格朗日最小能量流行动者(RAFALE),一种用于安全RL的离策略行动者-评论家方法。我们将其更新表述为约束单端广义Schrödinger桥,并证明对于每个源抽样,该路径空间问题恰好是动作空间中的熵正则化问题。在正噪声下,其解仅在对估计成本超过由拉格朗日乘子设定的阈值处对仅奖励动作分布重新加权。随着噪声消失,最优值收敛到流策略直接优化的最小能量映射目标的值。在七个Safety-Gymnasium任务上,RAFALE在每项任务上均实现了有竞争力的奖励,且平均最终成本在预算内,而强基线则牺牲一方以换取另一方;消融实验支持了其增广目标和流行动者的必要性。
cs.LG / 202 / 2609.32956

Efficient Message Passing for Partial Differential Equation Priors

偏微分方程先验的高效消息传递
Kazachkova, Anna, Hennicke, Leonhard, Schlosser, Rainer, Herbrich, Ralf
Abstract
Prior information for real-world physical quantities is most elegantly expressed via partial differential equations (PDEs). In this paper, we propose a novel way to solve PDEs using probabilistic inference on a factor graph. In general, factor graphs provide a natural way to encode prior knowledge into a model as explicit factors; here, this knowledge is provided by a governing PDE, which narrows the solution space, while observed data further shape the posterior over the parameters. The approximate parameter posterior is inferred using message passing based on moment matching, without posterior sampling or global gradient-based optimization. We demonstrate our approach on the first-order advection and the second-order semi-linear Fisher-KPP equations, where it achieves predictive accuracy comparable to a standard baseline while providing structured predictive uncertainty. Moreover, the inferred posterior marginal means and uncertainty structure match more closely those obtained using Hamiltonian Monte Carlo than the evaluated mean-field variational inference baseline, while requiring up to 10x less training time in our experiments, with inference speed comparable to variational inference.
Chinese Translation
现实世界物理量的先验信息最优雅地通过偏微分方程(PDEs)来表达。在本文中,我们提出了一种利用因子图上的概率推理来求解偏微分方程的新方法。通常,因子图提供了一种自然的方式,将先验知识作为显式因子编码到模型中;在这里,该知识由控制偏微分方程提供,它缩小了解空间,而观测数据进一步塑造了参数的后验。近似参数后验是通过基于矩匹配的消息传递来推断的,无需后验采样或全局基于梯度的优化。我们在一阶对流和二阶半线性 Fisher-KPP 方程上展示了我们的方法,它实现了与标准基线相当的预测精度,同时提供了结构化的预测不确定性。此外,与评估的均值场变分推理基线相比,推断的后验边际均值和不确定性结构与使用 Hamiltonian Monte Carlo 获得的结果更为接近,而在我们的实验中所需的训练时间最多减少 10 倍,推理速度与变分推理相当。
cs.LG / 203 / 2609.32961

Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents

超越Token节省:LLM智能体中上下文压缩的系统性研究
Satish, Ritul, Sinha, Prasoon, Kawada, Akiho, Yadwadkar, Neeraja J.
Abstract
As LLM agents tackle longer tasks, they increasingly compress growing histories of reasoning, actions, and tool outputs. Compression can reduce token use, but it also changes the information available for later decisions. Existing agentic harnesses bundle decisions about what to compress, when to compress, and how much to remove into fixed policies. A systematic characterization is needed to disentangle these decisions and reveal how each affects task success and execution cost. We systematically vary these decisions across three open-weight models on SWE-bench Verified and Terminal-Bench 1.0. Across nearly 35,000 agent runs, we measure task success, token use, end-to-end latency, and estimated cost. We find that fewer tokens need not mean faster or cheaper execution: on Terminal-Bench with Qwen, policies using roughly one-third as many tokens can take 20-80% longer than the uncompressed agent. Policies with similar overall success can solve different tasks, while the same policy can perform quite differently across models. Our results motivate evaluating compression by its effects on agent execution and tailoring policies to the task, model, and workload.
Chinese Translation
随着LLM智能体处理更长的任务,它们越来越多地压缩不断增长的推理、行动和工具输出历史。压缩可以减少Token使用,但它也改变了后续决策可用的信息。现有的智能体框架将关于压缩什么、何时压缩以及移除多少的决策捆绑到固定策略中。需要进行系统性的表征来解开这些决策,并揭示每个决策如何影响任务成功和执行成本。我们在SWE-bench Verified和Terminal-Bench 1.0上,针对三个开放权重模型系统地改变这些决策。在近35,000次智能体运行中,我们测量了任务成功率、Token使用量、端到端延迟和估计成本。我们发现,更少的Token并不一定意味着更快或更便宜的执行:在Terminal-Bench上使用Qwen时,使用大约三分之一Token的策略可能比未压缩的智能体耗时多出20-80%。总体成功率相似的策略可能解决不同的任务,而相同的策略在不同模型上的表现可能差异很大。我们的结果促使我们根据压缩对智能体执行的影响来评估压缩,并针对任务、模型和工作负载定制策略。
cs.LG / 204 / 2609.32962

Clipped or Unclipped? Finite-Sample Trade-offs for Averaged SGD under Heavy-Tailed Noise

裁剪还是未裁剪?重尾噪声下平均SGD的有限样本权衡
Suvorikova, Alexandra, Gladin, Egor, Dvinskikh, Darina, Agafonov, Artem, Alkousa, Mohammad, Dorn, Yuriy, Matyukhin, Vladislav, Gasnikov, Alexander
Abstract
Gradient clipping is widely used to stabilize training, but it need not improve the statistical accuracy of averaged SGD, even under heavy-tailed noise. We derive a finite-sample comparison of clipped and unclipped Polyak-Ruppert averaged SGD under finite conditional $p$-th moments, $p\ge2$. Our main result gives explicit accuracy and confidence conditions under which, for $p>2$, the Gaussian term dominates the unclipped deviation bound, so clipping need not improve its leading order. By balancing clipping bias and concentration, we obtain a bound in which the heavy-tail correction depends logarithmically rather than polynomially on the inverse failure probability. At $p=2$, this improves the confidence dependence of the leading bound. We establish sharpness of the unclipped heavy-tail term through an exact one-dimensional quadratic recursion and extend the comparison to projected convex SGD. We also prove concrete costs of clipping: every fixed finite threshold increases asymptotic variance on a scalar Gaussian quadratic, while whole-gradient clipping can shift the limiting point under asymmetric noise.
Chinese Translation
梯度裁剪被广泛用于稳定训练,但即使存在重尾噪声,它也不一定能提高平均SGD的统计精度。我们在有限条件$p$阶矩($p\ge2$)下,推导了裁剪和未裁剪的Polyak-Ruppert平均SGD的有限样本比较。我们的主要结果给出了明确的精度和置信条件,在这些条件下,当$p>2$时,高斯项主导未裁剪的偏差界,因此裁剪无需改善其主导阶。通过平衡裁剪偏差和集中性,我们得到一个界,其中重尾修正对逆失败概率的依赖是对数级而非多项式级的。当$p=2$时,这改善了主界的置信依赖性。我们通过精确的一维二次递归建立了未裁剪重尾项的尖锐性,并将比较扩展到投影凸SGD。我们还证明了裁剪的具体代价:每个固定的有限阈值都会增加标量高斯二次型上的渐近方差,而全梯度裁剪在非对称噪声下可能移动极限点。
cs.LG / 205 / 2609.32966

Self-Confirming Superposition Traps in Reinforcement Learning

强化学习中的自确认叠加陷阱
Shi, Dai, Han, Andi, Chen, Feng, Duan, Yiqun, Gao, Junbin, Hernández-Lobato, José Miguel
Abstract
Reinforcement learning (RL) trains representations on data selected by the agent's policy, which then uses the resulting returns to guide its next choices. We show that this loop can sustain a lower-return policy even when representation fitting is globally optimal on those data. In a self-confirming superposition trap, every optimal code assigns overlapping directions to features that rarely occur together under the current policy. An alternative action brings them together, causing interference that lowers its return and reinforces avoidance, although refitting to that action would yield more return at the same capacity. We characterize the dimensions admitting a trap in a tied two-step model and show separately that equal feature frequencies, continued visitation, and independent controller learning need not prevent it. Because fitting weights errors by visitation, an avoided action can lose its return advantage at little cost to the objective. In a finite-action model, we bound this distortion and derive a replay condition: sufficient training weight on the best separately adapted action preserves its ranking despite residual error. Neural PPO experiments show how the feedback develops during learning: agents initialized toward different actions develop different interference patterns, opposite mean return rankings, and different final policies at the same capacity. We therefore test whether retaining access to neglected states can improve control. Keeping these states in training reduces measured interference and improves sequential return, with gains even when the encoder is frozen. Related interventions on state access, replay weights, and feature overlap improve control on MiniGrid and DMControl. For agents that learn through a world model, protected fitting improves DreamerV3--Crafter's cumulative training scores at unchanged capacity.
Chinese Translation
强化学习(RL)根据智能体策略选择的数据训练表征,然后利用得到的回报来指导其下一步选择。我们表明,即使表征拟合在这些数据上全局最优,这种循环仍可能维持一个较低回报的策略。在自确认叠加陷阱中,每个最优编码都将重叠的方向分配给在当前策略下很少同时出现的特征。一个替代动作将它们带到一起,导致干扰,从而降低其回报并强化回避行为,尽管在相同容量下对该动作重新拟合会产生更多回报。我们在一个绑定两步模型中刻画了允许陷阱的维度,并分别表明,相等的特征频率、持续的访问以及独立的控制器学习并不一定能阻止它。因为拟合按访问次数对误差加权,一个被回避的动作可以在对目标函数代价很小的情况下失去其回报优势。在有限动作模型中,我们界定了这种扭曲,并推导出一个重放条件:对最佳单独适应的动作给予足够的训练权重,即使存在残差误差,也能保持其排名。神经网络PPO实验展示了这种反馈在学习过程中如何发展:朝向不同动作初始化的智能体在相同容量下发展出不同的干扰模式、相反的平均回报排名以及不同的最终策略。因此,我们测试保留对受忽视状态的访问是否能改善控制。在训练中保留这些状态可以减少测得的干扰并提高顺序回报,即使编码器冻结也能获得增益。对状态访问、重放权重和特征重叠的相关干预改善了在MiniGrid和DMControl上的控制。对于通过世界模型学习的智能体,保护性拟合在容量不变的情况下提高了DreamerV3在Crafter上的累积训练分数。
cs.LG / 206 / 2609.32976

Adaptive Ensemble Selection for Noisy Labels on Tabular Data

表格数据噪声标签的自适应集成选择
Ali, Faizaan, Kang, Inwon, Seneviratne, Oshani
Abstract
Incorrect or corrupted labels in tabular datasets can significantly degrade supervised learning performance, particularly when mislabeling is subtle and not easily detectable from feature space alone. In the context of automated or AI-augmented data science workflows, robust detection of such label noise is critical for building reliable models. We propose a data-centric reasoning module for AI data science systems that automatically diagnoses dataset quality and selects appropriate cleaning strategies. Given a dataset, a meta-model predicts weights over a diverse set of detectors, including confidence-based, neighborhood-based, and distributional methods. Across benchmark datasets with controlled noise, our approach achieves performance comparable to a Confident Learning baseline on average, with dataset-dependent gains and losses, particularly in heterogeneous regimes. We further show that detector effectiveness is systematically linked to dataset properties. These results demonstrate the value of descriptor-driven, data-centric ensembling as a component of AI-assisted data-science pipelines for robust dataset assessment and model reliability.
Chinese Translation
表格数据集中不正确或损坏的标签会显著降低监督学习性能,特别是当错误标记很微妙且难以仅从特征空间检测时。在自动化或AI增强的数据科学工作流中,对此类标签噪声的鲁棒检测对于构建可靠模型至关重要。我们提出了一个面向AI数据科学系统的以数据为中心的推理模块,该模块自动诊断数据集质量并选择合适的清洗策略。给定一个数据集,一个元模型预测在一组多样化检测器上的权重,这些检测器包括基于置信度的、基于邻域的和基于分布的方法。在具有受控噪声的基准数据集上,我们的方法平均实现了与Confident Learning基线相当的性能,并具有依赖于数据集的增益和损失,尤其是在异构场景中。我们进一步表明,检测器的有效性系统性地与数据集属性相关联。这些结果证明了描述符驱动的、以数据为中心的集成作为AI辅助数据科学管道的一个组成部分,对于鲁棒的数据集评估和模型可靠性的价值。
cs.LG / 207 / 2609.32980

Feasible Flow Matching for Graph Reconstruction via Within-Sampling Primal-Dual Guidance

基于采样内原始-对偶引导的可行流匹配用于图重建
Chen, Haoming, Zilberstein, Nicolas, Paternain, Santiago, Segarra, Santiago
Abstract
Graph reconstruction from partial observations often comes with structural side information, such as degree bounds, triangle counts, or an edge-density band. Prior-Informed Flow Matching (PIFM) reconstructs graphs by transporting a local prior toward the graph distribution, but it provides no mechanism to incorporate this side information. We put forth Constrained Primal-Dual PIFM (CPD-PIFM), which augments the sampler with Lagrange multipliers that evolve along each trajectory. The multipliers respond to constraint violations at a predicted endpoint and guide subsequent sampling steps without retraining. We prove that the sampler inherits PIFM's permutation equivariance and bound its expected terminal slack by a term that decays as the inverse square root of the number of steps, plus two approximation terms. On three link-prediction benchmarks and nine combinations of datasets and constraints, CPD-PIFM raises feasibility by 11-26 percentage points and remains competitive with fixed guidance without selecting a separate multiplier for each constraint.
Chinese Translation
从部分观测中重建图通常伴随着结构侧信息,如度约束、三角形计数或边密度带。先验信息流匹配(PIFM)通过将局部先验传输到图分布来重建图,但它没有提供机制来整合这些侧信息。我们提出了约束原始-对偶PIFM(CPD-PIFM),它通过沿每条轨迹演化的拉格朗日乘子来增强采样器。这些乘子响应预测端点处的约束违反,并引导后续采样步骤而无需重新训练。我们证明了采样器继承了PIFM的置换等变性,并将其期望终端松弛量限制在一个以步数的平方根倒数衰减的项加上两个近似项上。在三个链接预测基准和九种数据集与约束的组合上,CPD-PIFM将可行性提高了11-26个百分点,并且与固定引导相比具有竞争力,而无需为每个约束选择单独的乘子。
cs.LG / 208 / 2609.32991

What Should Data Teach? Moving Bottlenecks Across Circuit, Store, and Use

数据应教什么?在回路、存储与使用间移动瓶颈
Chen, Yixiao, Cheng, Ke, Guan, Jiangtao, Huang, Shuo, Liu, Yue, Zhang, Jun, Liu, Yuhong, Jiang, Jie
Abstract
What should data teach a language model at a particular point in training? A circuit view reveals three distinct bottlenecks: forming a computation, making its required content available, and selecting among available routes. A shared diagnosis-to-data principle connects them: localize the missing operation, preserve its causal relation, vary shortcut-bearing context, and re-audit the residual. Formation-sensitive selection and prerequisite ordering accelerate a binding-matching-transport path; a brief early prefix from the same training multiset retains a validation advantage through 100B tokens. Availability counterfactuals then distinguish writing content from invoking available memory, while paired supervision and context-opportunity ranking improve matched route decisions and long-context answer likelihood. A continuous 350M-model experiment connects the three interventions on the same facts: early circuit training improves subsequent learning, and the complete sequence outperforms stage-replacement controls on facts withheld from Use teaching. Independent query surfaces and opposed-source decisions expose conditional arbitration as the remaining frontier. Together, these results show why a change in the limiting operation calls for a change in supervision, not merely a new ranking of difficult examples.
Chinese Translation
在训练的特定时刻,数据应该教语言模型什么?回路视角揭示了三个不同的瓶颈:形成计算、使其所需内容可用,以及在可用路径中选择。一个共享的从诊断到数据的原则将它们联系起来:定位缺失的操作,保持其因果关联,变化带有捷径的上下文,并重新审查残差。对形成敏感的选择和先决条件排序加速了绑定-匹配-传输路径;来自相同训练多重集的简短早期前缀在1000亿token上保持了验证优势。可用性反事实随后区分写入内容与调用可用记忆,而配对监督和上下文机会排序改善了匹配路径决策和长上下文答案可能性。一个连续的350M模型实验在相同事实上连接了三种干预:早期回路训练改善了后续学习,完整序列在未用于使用教学的事实上优于阶段替换控制。独立查询表面和对立来源决策暴露了条件仲裁作为剩余前沿。总之,这些结果说明了为什么限制性操作的变化需要监督的变化,而不仅仅是对困难示例的新排序。
cs.LG / 209 / 2609.33011

Saturation-Insensitive Dueling Bandits with General Function Approximation

通用函数逼近下的饱和不敏感对决赌博机
Zhang, Chenggong, Li, Xuheng, Di, Qiwei, Zhang, Weitong, Gu, Quanquan
Abstract
We study contextual dueling bandits with general function approximation under the Bradley-Terry-Luce (BTL) preference model. A key challenge in this setting is the saturation of the preference model: when the current reward model can already distinguish two actions with high confidence, the resulting preference feedback becomes weakly informative, making it difficult to further improve reward estimation. Consequently, existing sample-complexity analyses often depend on the inverse-derivative factor $1 / \sigma'[\Delta_{r^\ast}]$ which can be prohibitively large when the link function $\sigma$ saturates for large reward gaps $\Delta_{r^\ast}$. To address this issue, we introduce `SI-CDB`, an algorithm that selects opponent arms using a carefully designed heuristic for arm selection. This design enables saturation-insensitive reward learning and recovers the near-optimal dependence for linear reward classes, eliminating the unfavorable $1/\sigma'(\cdot)$ factor. The core of our analysis is a localized Eluder dimension framework tailored to dueling bandits with general function approximation. Our theoretical results also explain why two-arm regret analysis is crucial for improving single-arm performance in dueling bandits.
Chinese Translation
我们研究了在 Bradley-Terry-Luce (BTL) 偏好模型下具有通用函数逼近的上下文对决赌博机。该设置中的一个关键挑战是偏好模型的饱和:当当前奖励模型已经能够以高置信度区分两个动作时,所产生的偏好反馈变得信息量微弱,从而难以进一步改进奖励估计。因此,现有的样本复杂度分析通常依赖于逆导数因子 1 / σ'[Δ_{r^*}],当链接函数 σ 在大奖励差距 Δ_{r^*} 处饱和时,该因子可能变得极大而无法承受。为了解决这个问题,我们引入了 `SI-CDB`,一种使用精心设计的臂选择启发式来选择对手臂的算法。该设计实现了对饱和不敏感的奖励学习,并恢复了线性奖励类的近最优依赖,消除了不利的 1/σ'(·) 因子。我们分析的核心是一个局部化的 Eluder 维度框架,专为具有通用函数逼近的对决赌博机量身定制。我们的理论结果还解释了为什么双臂遗憾分析对于提高对决赌博机中的单臂性能至关重要。
cs.LG / 210 / 2609.33016

Relative Generalization Invariance of LLM Pretraining

LLM预训练的相对泛化不变性
Zhang, Fengzhuo, Wang, Shuche, Li, Shenggui, Ruan, Tianyu, He, Jianliang, Tsang, Ivor, Pang, Tianyu, Du, Chao, Zhang, Tianwei, Yang, Zhuoran
Abstract
Large Language Model (LLM) pretraining performance is jointly shaped by three components of the training triplet: the optimizer, model architecture, and training data stream. However, how these components influence performance in distinct ways remains unclear. We take a first step toward isolating their effects by studying relative generalization. We introduce Relative Generalization Invariance (RGI), the invariance of the validation-loss difference between any two tokens across models. We show that RGI approximately holds across a wide range of optimizers and moderate architectural variations, suggesting that these choices induce an approximately uniform shift in token-wise losses. In contrast, changing the training data stream can substantially alter relative generalization. We further show that RGI cannot be explained by the neural tangent kernel or mean-field regimes alone and prove that it can emerge in an overparameterized quadratic model. Overall, our work identifies RGI as a new phenomenon in LLM pretraining that helps distinguish the effects of optimizers and architectures from those of training data.
Chinese Translation
大型语言模型(LLM)预训练性能由训练三元组的三个组成部分共同塑造:优化器、模型架构和训练数据流。然而,这些组成部分如何以不同方式影响性能仍不清楚。我们通过研究相对泛化,迈出了分离其影响的第一步。我们引入了相对泛化不变性(RGI),即任意两个词元在不同模型间的验证损失差异的不变性。我们表明,RGI在广泛的优化器和适度的架构变化中近似成立,这表明这些选择导致词元级损失产生近似均匀的偏移。相比之下,改变训练数据流可以显著改变相对泛化。我们进一步表明,RGI不能仅由神经正切核或平均场机制来解释,并证明它可以在一个过参数化二次模型中出现。总体而言,我们的工作将RGI识别为LLM预训练中的一种新现象,有助于区分优化器和架构的影响与训练数据的影响。
cs.LG / 211 / 2609.33025

Low-Rank Single-Index Bandits with Unknown Links: From Matrices to Tensors

低秩单指标Bandit与未知链接:从矩阵到张量
Liu, Zhongxuan, Kang, Yue, Lee, Thomas C. M.
Abstract
Low-rank matrix and tensor bandits exploit structured interactions but typically assume a known reward link. Recent single-index bandit methods accommodate unknown links without directly exploiting matrix or tensor rank. We address this gap by studying stochastic matrix and tensor bandits with an unknown shared Lipschitz link and a low-rank index parameter under known regular candidate distributions and finite-variance noise. For monotone links, T-ESTOR combines robust, rank-adaptive Stein estimation with epoch-based greedy selection. Under exact selected-score access and a uniformly positive selected-design Stein signal, it achieves square-root regret with dimension dependence determined by the low-rank structure. For every admissible design, the monotone lower bound matches the rank, dimension, and horizon dependence up to logarithmic factors at large horizons, for fixed menu size and model/design constants. For nonmonotone links under a nonzero base-law Stein signal, T-BSTOR combines structured estimation with robust bin-based learning and attains the optimal $\widetilde{O}(T^{2/3})$ horizon rate for fixed dimensions, menu size, and model/design constants. Synthetic and CCLE-based experiments illustrate the benefits of structured estimation relative to vectorized and competing single-index baseline methods.
Chinese Translation
低秩矩阵和张量Bandit利用结构化交互,但通常假设已知奖励连接函数。最近的单指标Bandit方法能够适应未知连接函数,而无需直接利用矩阵或张量秩。我们通过研究具有未知共享Lipschitz连接函数和低秩指标参数的随机矩阵和张量Bandit来填补这一空白,其中候选分布已知且规则,噪声具有有限方差。对于单调连接函数,T-ESTOR将稳健的、秩自适应的Stein估计与基于轮次的贪婪选择相结合。在精确的选定分数访问和一致正的选定设计Stein信号下,它实现了平方根遗憾,其维度依赖性由低秩结构决定。对于每个允许的设计,在固定菜单大小和模型/设计常数的情况下,单调下界在大时间范围上匹配秩、维度和时间范围的依赖性,直至对数因子。对于非单调连接函数,在非零base-law Stein信号下,T-BSTOR将结构化估计与稳健的基于箱的学习相结合,并在固定维度、菜单大小和模型/设计常数的情况下,达到最优的 $\widetilde{O}(T^{2/3})$ 时间范围速率。合成和基于CCLE的实验说明了结构化估计相对于向量化和竞争性单指标基线方法的优势。
cs.LG / 212 / 2609.33028

\L{}ukasiewicz Neural Networks Extended: Residual Architectures and Crystallization Strategies for Interpretable Rule Extraction

扩展的Łukasiewicz神经网络:用于可解释规则提取的残差架构与结晶策略
Leandro, Carlos
Abstract
A feed-forward neural network whose weights are integers and whose activation is the truncated identity implements, neuron by neuron, the connectives of \L{}ukasiewicz many-valued logic. This exact correspondence --- established theoretically by Castro and Trillas and developed into a training algorithm by Leandro --- enables \emph{symbolic knowledge extraction}: training produces not a black-box model but a logical formula. Two obstacles have limited the approach to shallow architectures and small datasets: crystallization (forcing weights to integers) succeeds only probabilistically under the original Levenberg--Marquardt training scheme, and the theoretical guarantees break down as networks grow deeper. This paper addresses both obstacles. First, we prove that \emph{residual connections} (skip connections of the kind used in ResNets) extend \L{}ukasiewicz neural networks to arbitrary depth while preserving the symbolic correspondence \emph{at merge neurons} by construction: merge neurons in a \L{}ukasiewicz residual block automatically satisfy the neuron-classification proposition, regardless of the inner layer weights; inner-layer neurons are trained toward representability by the crystallization strategy. Second, we analyse three crystallization strategies --- Levenberg--Marquardt (corrected), straight-through estimation (STE), and proximal regularization --- characterizing their theoretical guarantees, failure modes, and interpretability trade-offs.
Chinese Translation
一个前馈神经网络,其权重为整数,激活函数为截断恒等函数,能够逐个神经元地实现Łukasiewicz多值逻辑的连接词。这种精确对应——由Castro和Trillas在理论上建立,并由Leandro发展为训练算法——使得符号知识提取成为可能:训练产生的不是黑箱模型,而是一个逻辑公式。两个障碍限制了该方法只能用于浅层架构和小数据集:结晶(强制权重为整数)在原始Levenberg-Marquardt训练方案下只能概率性地成功,并且随着网络变得更深,理论保证会失效。本文解决了这两个障碍。首先,我们证明了残差连接(ResNet中使用的那种跳跃连接)可以将Łukasiewicz神经网络扩展到任意深度,同时通过构造在合并神经元处保持符号对应:Łukasiewicz残差块中的合并神经元自动满足神经元分类命题,无论内层权重如何;内层神经元通过结晶策略被训练以具有可表示性。其次,我们分析了三种结晶策略——Levenberg-Marquardt(校正版)、直通估计(STE)和近端正则化——描述了它们的理论保证、失效模式和可解释性权衡。
cs.LG / 213 / 2609.33030

What Must a World Model Distinguish for Planning?

世界模型为了规划必须区分什么?
Wei, Rongzhe, Hsu, Hans Hao-Hsun, Niu, Peizhi, Li, Yifan, Li, Pan
Abstract
World models simulate the consequences of action candidates, but good planning need not preserve every physical distinction required for accurate prediction. We formalize this gap through a hierarchy of mechanism, response, and decision sufficiency. Given a candidate set, the planning query determines which physical variations matter and how precisely they must be preserved: coarse decisions can discard much of the information needed for prediction, whereas fine decisions may require nearly the same resolution. In practice, planners often adaptively search to construct candidates, and information unnecessary for final selection may still be needed to discover good candidates. What a world model must preserve therefore depends on the query, the candidate set, and the planner. We study these effects in a collision system, nonlinear dynamics, and robotic planning. These varying requirements raise a design question: where should query information enter the planning system? A model that jointly generates actions and outcomes conditioned on the query achieves lower regret than an action-conditioned world model on seen objectives, but this advantage largely disappears when generalizing to unseen objectives. Motivated by this, we propose a modular design in which the query determines where to look and an action-conditioned model predicts what will happen, allowing the same predictions to be reused across objectives.
Chinese Translation
世界模型模拟动作候选的后果,但良好的规划不必保留准确预测所需的每一个物理区分。我们通过机制、响应和决策充分性的层级来形式化这一差距。给定一个候选集,规划查询决定了哪些物理变化重要以及必须多精确地保留它们:粗略决策可以丢弃预测所需的大量信息,而精细决策可能需要几乎相同的分辨率。在实践中,规划器通常自适应搜索以构造候选,最终选择不需要的信息可能仍然需要用于发现好的候选。因此,世界模型必须保留什么取决于查询、候选集和规划器。我们在碰撞系统、非线性动力学和机器人规划中研究这些效应。这些不同的需求提出了一个设计问题:查询信息应该在哪里进入规划系统?一个以查询为条件联合生成动作和结果的模型在已见目标上比动作条件世界模型获得更低的遗憾,但当泛化到未见目标时,这种优势在很大程度上消失了。受此启发,我们提出了一种模块化设计,其中查询决定看向何处,动作条件模型预测将会发生什么,从而允许相同的预测在不同目标间重用。
cs.LG / 214 / 2609.33041

Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons

平衡早期性能牺牲与长期收益:跨训练总长缩放学习率预热持续时间
Topollai, Kristi, Choromanska, Anna
Abstract
Learning-rate warmup is a standard technique in language-model training, yet its duration remains largely heuristic. Common approaches use either a fixed number of updates or a fixed fraction of the training horizon, two choices that imply very different scaling as training gets longer. When should warmup stay fixed, and when should it grow with the horizon? We address this question with a quadratic model whose modes respond differently to the peak learning rate. Warmup slows progress in directions that already contract well at the peak rate, but can remove persistent error in directions near the stability edge, with higher peak rates shifting the balance toward longer warmup durations. This yields a compact horizon scaling law that captures regimes ranging from essentially no warmup, through fixed-duration warmup, to durations that grow with the training horizon, and explains how the preferred regime changes with peak learning rate. Because the law captures the tradeoff between giving up early progress and improving the trajectory that follows, it can be fit using shorter runs and used to predict warmup at substantially longer horizons. Together, our results explain several familiar properties of warmup through a single tradeoff and suggest treating warmup duration as a horizon-dependent hyperparameter rather than a fixed training heuristic.
Chinese Translation
学习率预热是语言模型训练中的标准技术,但其持续时间在很大程度上仍然是启发式的。常见方法要么使用固定数量的更新步数,要么使用训练总长的一个固定比例,这两种选择意味着随着训练变长,其缩放方式非常不同。何时应保持预热固定,何时应使其随训练总长增长?我们用一个二次模型来解决这个问题,该模型的模式对峰值学习率的响应不同。预热会减慢在峰值学习率下已经很好收敛的方向上的进展,但可以消除接近稳定边缘的方向上的持续误差,更高的峰值学习率会将平衡移向更长的预热持续时间。这产生了一个紧凑的训练总长缩放定律,它涵盖了从基本无预热,到固定时长预热,再到随训练总长增长的持续时间的各种情况,并解释了首选情况如何随峰值学习率变化。由于该定律捕捉了放弃早期进展与改善后续轨迹之间的权衡,因此可以使用较短的运行来拟合,并用于预测在更长训练总长下的预热。总之,我们的结果通过一个单一的权衡解释了预热的几个熟悉特性,并建议将预热持续时间视为依赖于训练总长的超参数,而不是固定的训练启发式。
cs.LG / 215 / 2609.33047

Cost-free Spectral Estimation for Adaptive Newton--Schulz in Matrix Optimizers

矩阵优化器中自适应 Newton--Schulz 的零成本谱估计
Topollai, Kristi, Choromanska, Anna
Abstract
Matrix optimizers such as Muon transform each momentum matrix through an approximate orthogonalization, typically implemented by a small number of Newton-Schulz matrix multiplications. The quality and cost of this approximation depend strongly on the singular-value spectrum of its input, yet existing implementations use the same fixed polynomial routine for every layer and throughout training. We show that this uniform treatment is unnecessary: the computations in the Newton-Schulz method already reveal enough information to make the method adaptive. The Gram matrices formed inside Newton-Schulz iterations yield spectral moments through inexpensive scalar reductions, requiring no additional matrix multiplications. From these moments, we recover an estimate of the empirical singular-value distribution and use it to select a polynomial routine specialized to the current matrix. This turns Newton--Schulz orthogonalization into a spectrum-adaptive procedure that responds to differences across both layers and training time. On saved momentum matrices, spectral estimation substantially reduces orthogonalization error at a fixed iteration budget or reaches the same accuracy with fewer iterations, and in GPT pretraining up to 1B parameters it lowers the validation loss of two matrix optimizers. Our results suggest that matrix-function operations inside optimizers need not be designed for a conservative worst-case spectrum: they can cheaply measure the spectrum they are already processing and specialize computations accordingly.
Chinese Translation
矩阵优化器(如 Muon)通过近似正交化来变换每个动量矩阵,这通常由少量 Newton-Schulz 矩阵乘法实现。该近似的质量和成本在很大程度上取决于其输入的奇异值谱,然而现有实现对于每一层和整个训练过程都使用相同的固定多项式例程。我们表明,这种统一处理是不必要的:Newton-Schulz 方法中的计算已经揭示了足够的信息,使该方法可以自适应。Newton-Schulz 迭代内部形成的 Gram 矩阵通过廉价的标量归约产生谱矩,无需额外的矩阵乘法。从这些矩中,我们恢复出经验奇异值分布的估计,并利用它来选择专门针对当前矩阵的多项式例程。这将 Newton-Schulz 正交化转变为一种谱自适应过程,能够响应不同层以及训练时间上的差异。在保存的动量矩阵上,谱估计在固定迭代预算下显著降低正交化误差,或者以更少的迭代达到相同精度;在高达 10 亿参数的 GPT 预训练中,它降低了两个矩阵优化器的验证损失。我们的结果表明,优化器内部的矩阵函数操作不必针对保守的最坏情况谱进行设计:它们可以廉价地测量自己已经在处理的谱,并据此专门化计算。
cs.LG / 216 / 2609.33048

DevelopmentODE: Structured Neural ODEs for Early Brain Development Dynamics Across a Decade

DevelopmentODE:用于十年间早期大脑发育动态的结构化神经常微分方程
Han, Kaiqiao, Chen, Haitao, Quah, Bryan, Wang, Xiaoda, Liu, Janelle, Gilmore, John H, Gao, Wei, Sun, Yizhou
Abstract
Understanding how individual brain development unfolds over childhood requires modeling developmental trajectories from sparse longitudinal observations. Long-term neurodevelopmental forecasting is challenging because each child is typically observed at only a few irregularly spaced visits, while developmental dynamics vary across individuals and age. Generic continuous-time models accommodate irregular timing but often absorb these factors into a single flexible transition function, providing little structure for how population progression, individual variability, and developmental age shape the dynamics. We propose DevelopmentODE, a structured continuous-time framework that organizes population- and subject-specific variation within a shared developmental geometry while allowing the governing dynamics to evolve with age. The model builds this geometry around a developmental canal representing the population trajectory, whose local direction provides a reference for organizing subject-specific variation. Subject deviation velocities are constrained relative to this direction, while a shared nonlinear deviation field captures individual developmental motion without disrupting population-level progression. DevelopmentODE further models developmental non-stationarity through ordered age-dependent deformations of the shared vector field, progressively adapting a common dynamical structure as age changes, while elapsed time determines the integration horizon. This formulation uses population-level developmental structure to guide learning from sparse individual trajectories while allowing dynamics to evolve smoothly with age. We evaluate DevelopmentODE on longitudinal fMRI by predicting future functional connectivity of the same child from earlier observations. DevelopmentODE consistently outperforms competing baselines across short- and long-horizon predictions.
Chinese Translation
理解个体大脑发育在童年期如何展开,需要对来自稀疏纵向观测的发育轨迹进行建模。长期神经发育预测具有挑战性,因为每个儿童通常仅在少数不规则间隔的访视中被观测,而发育动态因个体和年龄而异。通用的连续时间模型能够适应不规则的时间安排,但通常将这些因素吸收到单个灵活的转移函数中,几乎不提供关于群体进展、个体变异性和发育年龄如何塑造动态的结构。我们提出 DevelopmentODE,一个结构化的连续时间框架,它在共享发育几何中组织群体和受试者特异性变异,同时允许控制动态随年龄演变。该模型围绕代表群体轨迹的发育通道构建这种几何,其局部方向为组织受试者特异性变异提供了参考。受试者偏差速度相对于该方向受到约束,而共享非线性偏差场捕获个体发育运动,而不破坏群体水平的进展。DevelopmentODE 进一步通过对共享向量场的有序年龄依赖性变形来建模发育非平稳性,随着年龄变化逐步适应共同的动态结构,而流逝时间决定积分时域。该公式利用群体水平的发育结构来指导从稀疏个体轨迹中学习,同时允许动态随年龄平滑演变。我们在纵向 fMRI 上评估 DevelopmentODE,通过从早期观测预测同一儿童未来的功能连接。DevelopmentODE 在短期和长期预测中始终优于竞争基线。
cs.LG / 217 / 2609.33051

SketchSSM: Write to the Full State, Read from a Compact Sketch

SketchSSM:全状态写入,紧凑草图读取
Kwon, Omin, Shin, JoongWon, Kim, Minseo, Keutzer, Kurt, Kim, Sehoon, Lee, Jae W.
Abstract
Hybrid-attention models replace most softmax attention layers with linear attention, reducing KV-cache growth and enabling larger decode batches where recurrent state access becomes a major bottleneck. ReplaySSM amortizes state updates by buffering keys and values, but each new query still requires a full-state read even though the state remains unchanged between state updates. We observe that low-rank state-weighted query approximation accurately preserves state-read outputs. Although future queries are unknown, the basis vectors used to approximate them can be fixed offline. Based on this observation, we introduce SketchSSM, which preserves full-state updates while approximating reads. At each state update, SketchSSM reads the full state once to precompute outputs for these basis vectors, storing them in a compact sketch. Each subsequent decode step combines the sketch vectors with query-dependent coefficients to reconstruct the output without a full-state read. Across four Mamba-2-, GDN-, and KDA-based models, SketchSSM reduces state-access traffic by approximately 10x while largely preserving average accuracy across four decode benchmarks and recall on four RULER retrieval tasks. On one NVIDIA B300, linear-attention kernel speedups over the standard vLLM baseline reach 7.78x, 5.22x, and 5.20x for Mamba-2, GDN, and KDA, respectively, with up to 2.64x higher decode throughput on Nemotron 3 Super.
Chinese Translation
混合注意力模型用线性注意力替代大多数softmax注意力层,减少了KV缓存增长,并能够实现更大的解码批次,其中循环状态访问成为主要瓶颈。ReplaySSM通过缓冲键和值来分摊状态更新,但每个新查询仍然需要全状态读取,即使状态在状态更新之间保持不变。我们观察到,低秩状态加权查询近似能够准确保持状态读取输出。尽管未来查询未知,但用于近似它们的基向量可以离线固定。基于这一观察,我们提出了SketchSSM,它在保持全状态更新的同时近似读取。在每次状态更新时,SketchSSM读取一次完整状态,为这些基向量预计算输出,并将它们存储在一个紧凑的草图中。随后的每个解码步骤将草图向量与查询相关的系数结合,以在不进行全状态读取的情况下重建输出。在四个基于Mamba-2、GDN和KDA的模型中,SketchSSM将状态访问流量减少了约10倍,同时在四个解码基准上基本保持了平均准确率,并在四个RULER检索任务上保持了召回率。在一台NVIDIA B300上,对于Mamba-2、GDN和KDA,线性注意力内核相对于标准vLLM基线的加速分别达到7.78倍、5.22倍和5.20倍,在Nemotron 3 Super上解码吞吐量最高提升2.64倍。
cs.LG / 218 / 2609.33060

Deep Learning Techniques for Phoneme Recognition in Italian Children' s Speech

意大利儿童语音音素识别的深度学习技术
Barbaro, Nicola, Gena, Cristina, Petriglia, Francesco, Meirone, Andrea, Mazzei, Alessandro, Viotti, Arianna
Abstract
Speech therapists often face difficulties diagnosing impairments due to the lack of efficient tools for transcribing speech into the International Phonetic Alphabet (IPA). This work addresses this challenge with Broca, a Conformer-based deep learning system pretrained on 8 days of adult speech and fine-tuned on a 165-minute dataset of Italian child speech collected through a range of standardized diagnostic tests for children aged 3.5-6.5. Broca was optimized to handle phonetic variability in children's speech, including tone, accent, and speech errors, and achieved a state-of-the-art weighted Phoneme Error Rate of 13.36% on Italian speech. Remarkably, this performance was obtained using less than three hours of child-specific data, underscoring the model's efficiency and robustness in low-resource clinical settings. This work demonstrates that accurate, vocabulary-independent speech-to-IPA transcription can be achieved with minimal data, paving the way for more accessible, data-efficient tools to support speech assessment and diagnosis.
Chinese Translation
言语治疗师常常因缺乏将语音转写为国际音标(IPA)的高效工具而难以诊断障碍。本研究通过 Broca 应对这一挑战;Broca 是一种基于 Conformer 的深度学习系统,先在 8 天的成人语音上预训练,随后在一个 165 分钟的意大利儿童语音数据集上微调,该数据集通过一系列针对 3.5-6.5 岁儿童的标准化诊断测试采集。Broca 经过优化以处理儿童语音中的语音变异性,包括语调、口音和言语错误,并在意大利语语音上取得了 13.36% 的最先进加权音素错误率。值得注意的是,这一性能是在使用不到三小时的儿童特定数据的情况下取得的,凸显了模型在低资源临床环境中的效率和稳健性。本研究表明,准确的、与词汇无关的语音到 IPA 转写可以在极少数据下实现,为支持语音评估和诊断的更易获取、数据高效的工具铺平了道路。
cs.LG / 219 / 2609.33066

Zero-Storage Procedural Neural Synthesis via Boundary Dynamics: Formal Verification in Lean 4 and Bare-Metal Gauntlet Validation

基于边界动力学的零存储过程式神经合成:Lean 4 形式化验证与裸金属严苛验证
Dağlı, Volkan, Dağlı, Zerrin, Dağlı, Dağhan
Abstract
Contemporary neural inference architectures rely on dense floating-point weight matrices stored in high-bandwidth memory (VRAM), incurring severe memory-wall bottlenecks and preventing native execution inside deterministic virtual machines like the Ethereum Virtual Machine (EVM). Verifying termination and arithmetic invariants for recursive dynamical systems over continuous domains is generally undecidable in the Blum-Shub-Smale model. Here, we present the formal verification and bare-metal empirical validation of WERR (Waves & Errors) and Phase III Orbital Error Dynamics (OED), a non-tensor decision paradigm that procedurally synthesizes non-linear decision boundaries on demand from a 24-byte coordinate seed $\Theta = (c_x, c_y, \text{zoom})$ along the boundary of the Mandelbrot set ($\partial\mathcal{M}$). By projecting the recurrence $z_{n+1} = z_n^2 + c$ onto the modular residue ring $\mathbb{Z}/9\mathbb{Z}$ and the fixed-point domain $\mathbb{Q}_{16.16}$, we establish ten machine-verified theorems in Lean 4 (v4.34.1) with Mathlib4 and zero unproven conjectures (sorry): proving $\mathcal{I}_3 = \{0,3,6\} \subset \mathbb{Z}/9\mathbb{Z}$ ideal closure, universal fuel-bounded halting ($\le 9$ and $\le 12$ steps), absence of $\mathbb{Q}_{16.16}$ square overflow below $2^{63}-1$, non-constant boundary escape sensitivity, and a parametric EVM gas bound ($\le 22,557 \le 24,000$ gas). Evaluated on a 40-core Dual Intel Xeon server, the vectorized 36-iteration CPU kernel processes 100,000 decisions in 6.49 s (15,397 decisions/s, 0 Bytes VRAM, 15.15x speedup), while a sigmoidal outlier gate suppresses 100.00% of adversarial spikes while preserving 89.60% of clean baseline signals.
Chinese Translation
当代神经推理架构依赖于存储在高带宽内存(VRAM)中的稠密浮点权重矩阵,导致严重的内存墙瓶颈,并阻止了在确定性虚拟机(如以太坊虚拟机 EVM)内的原生执行。在 Blum-Shub-Smale 模型中,验证连续域上递归动力系统的终止性和算术不变量通常是不可判定的。在此,我们展示了 WERR(Waves & Errors)和 Phase III Orbital Error Dynamics (OED) 的形式化验证与裸金属实证验证,这是一种非张量决策范式,可根据需求,从 Mandelbrot 集边界($\partial\mathcal{M}$)上的 24 字节坐标种子 $\Theta = (c_x, c_y, \text{zoom})$ 过程式地合成非线性决策边界。通过将递推关系 $z_{n+1} = z_n^2 + c$ 投影到模剩余环 $\mathbb{Z}/9\mathbb{Z}$ 和定点域 $\mathbb{Q}_{16.16}$ 上,我们在 Lean 4 (v4.34.1) 和 Mathlib4 中建立了十个机器验证的定理,且零未证明猜想(无 sorry):证明了 $\mathcal{I}_3 = \{0,3,6\} \subset \mathbb{Z}/9\mathbb{Z}$ 理想闭包,通用燃料有界停机($\le 9$ 和 $\le 12$ 步),在 $2^{63}-1$ 以下不存在 $\mathbb{Q}_{16.16}$ 平方溢出,非恒定边界逃逸敏感性,以及参数化 EVM gas 界限($\le 22,557 \le 24,000$ gas)。在 40 核双路 Intel Xeon 服务器上评估,向量化的 36 次迭代 CPU 内核在 6.49 秒内处理 100,000 个决策(15,397 决策/秒,0 字节 VRAM,15.15 倍加速),而 Sigmoid 异常值门控抑制了 100.00% 的对抗性尖峰,同时保留了 89.60% 的干净基线信号。
cs.LG / 220 / 2609.33073

Algorithmic Harms Associated with Generative Model-Augmented Recommendation Systems

与生成模型增强推荐系统相关的算法危害
Herlihy, Christine, Xi, Xumei, Desai, Shloka, Hutchful, Kevin Bannerman, Silva, Pedro
Abstract
In this work, we consider algorithmic harms that may arise as generative models are incorporated into machine learning platforms. We argue that existing harm taxonomies and threat models require extension to (1) address novel causal drivers of well-studied representational and quality-of-service harms; and (2) anticipate and mitigate endogenous harms, such as sanitization, which may arise when system inputs are misaligned with the system designer's objectives, or the generative model's inductive priors. To this end, we introduce an expanded taxonomy of algorithmic harms associated with the use of generative models in non-conversational recommendation systems. In addition, we offer a causal analysis of how problematic subsets of the (input, output) joint distribution can arise, in an effort to inform harms detection and mitigation efforts.
Chinese Translation
在本工作中,我们考虑将生成模型纳入机器学习平台时可能出现的算法危害。我们认为,现有的危害分类法和威胁模型需要进行扩展,以:(1) 应对已被广泛研究的表征性危害和服务质量危害的新因果驱动因素;(2) 预见并缓解内生危害,例如净化(sanitization),当系统输入与系统设计者的目标或生成模型的归纳先验不一致时,这类危害便可能出现。为此,我们引入了一个扩展的分类体系,用于描述在非对话式推荐系统中使用生成模型所关联的算法危害。此外,我们对(输入,输出)联合分布中有问题的子集如何产生进行了因果分析,以期为危害检测和缓解工作提供参考。
cs.LG / 221 / 2609.33074

KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation

KernelZero:协同进化的 Proposer 与 Coder 用于持续改进的 GPU 内核生成
Ke, Changxin, Zhang, Rui, Fang, Zixiang, Li, Zhenghong, Wen, Yuanbo, Shen, Jiashuo, Wang, Shuo, Guo, Jiaming, Li, Ling, Guo, Qi, Chen, Yunji
Abstract
High-performance GPU kernels are essential to modern machine learning systems, yet automatically generating kernels that are both correct and efficient remains challenging. Existing LLM-based approaches face two major limitations: the scarcity of high-quality training data aligned with the model's current capabilities, and the inherent trade-off between kernel correctness and performance. To address these challenges, we propose KernelZero, a co-evolution framework that continuously improves GPU kernel generation through two specialized models: a Proposer that generates Torch modules from API sets and a Coder that translates them into CUDA or Triton kernels. KernelZero uses a frontier-driven module generation mechanism to continuously produce capability-aligned training modules based on the Coder's current weaknesses. It further introduces Correctness-Aware Group Relative Policy Optimization (CA-GRPO), which optimizes performance only after correctness becomes sufficiently reliable. By alternating the optimization of the Proposer and Coder, KernelZero forms an automatic curriculum that enables targeted and training-efficient capability improvement. Empirically, KernelZero-7B surpasses Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton. On KernelBench Level 1 and 2, it achieves CUDA pass@1 scores of 75.8% and 69.6%, respectively, with pass@10 reaching 100% and 97%. On Triton, it achieves pass@1 scores of 77.2% and 72.5%, respectively.
Chinese Translation
高性能 GPU 内核对于现代机器学习系统至关重要,然而自动生成既正确又高效的内核仍然具有挑战性。现有的基于 LLM 的方法面临两个主要限制:与模型当前能力对齐的高质量训练数据的稀缺,以及内核正确性与性能之间固有的权衡。为了解决这些挑战,我们提出了 KernelZero,一个协同进化框架,通过两个专门的模型持续改进 GPU 内核生成:一个 Proposer 从 API 集合生成 Torch 模块,一个 Coder 将它们转换为 CUDA 或 Triton 内核。KernelZero 使用前沿驱动的模块生成机制,基于 Coder 当前的弱点持续产生能力对齐的训练模块。它进一步引入了正确性感知的组相对策略优化(Correctness-Aware Group Relative Policy Optimization, CA-GRPO),仅在正确性变得足够可靠后才优化性能。通过交替优化 Proposer 和 Coder,KernelZero 形成了一个自动课程,实现有针对性的且训练高效的能力提升。实验上,KernelZero-7B 在 CUDA 上超越了 Claude-4.5-Sonnet,在 Triton 上超越了 DeepSeek-V4-Pro。在 KernelBench Level 1 和 2 上,它分别达到了 CUDA pass@1 分数 75.8% 和 69.6%,pass@10 达到 100% 和 97%。在 Triton 上,它分别达到了 pass@1 分数 77.2% 和 72.5%。
cs.LG / 222 / 2609.33078

The limits of exactness: On the failure of automatic differentiation in physics-informed machine learning

精确性的局限:论自动微分在物理信息机器学习中的失效
Jagtap, Ameya D.
Abstract
Automatic differentiation (AD) lets neural networks compute derivatives of governing equations to machine precision, and this precision has made it the computational backbone of physics-informed machine learning. Yet exactness in the mathematical sense is not the same as fidelity to the physics. Here I argue that a derivative can be numerically perfect and still be the wrong derivative for the problem at hand, because AD, by construction, has no notion of the physical structure a solution must obey. Convection and its associated directionality, diffusion, and dispersion are only the most visible instances of a much longer list that spans all branches of computational science and engineering, including conservation, thermodynamic consistency, symmetry, symplectic structure, positivity, monotonicity, and boundedness. Recognizing this broader gap reframes how the field should build the next generation of PDE-driven neural surrogates.
Chinese Translation
自动微分(AD)使神经网络能够以机器精度计算控制方程的导数,这种精度使其成为物理信息机器学习的计算支柱。然而,数学意义上的精确性并不等同于对物理的保真度。在此,我认为一个导数可以在数值上完美,但对于手头的问题来说仍然是错误的导数,因为AD从构造上就不具备解必须遵守的物理结构的概念。对流及其相关的方向性、扩散和色散只是一个更长列表中最明显的例子,该列表跨越计算科学与工程的所有分支,包括守恒、热力学一致性、对称性、辛结构、正性、单调性和有界性。认识到这一更广泛的差距,重新定义了该领域应如何构建下一代PDE驱动的神经代理模型。
cs.LG / 223 / 2609.33086

From Constitutions to Control: Interpretable Rewards for Aligning Language Models

从宪法到控制:用于对齐语言模型的可解释奖励
Gaebler, Johann D., Isley, Calvin, Lamparth, Max, Casper, Stephen, Goel, Sharad
Abstract
Current approaches to aligning language models often make it hard to know what behavior is being rewarded or to change that reward in a targeted way. In particular, standard preference-based methods collapse multiple considerations into aggregate human judgments, obscuring what drives the resulting reward, while principle-based methods specify high-level values without fully operationalizing them. To address this gap, we develop a rubric-based framework to transform a general-purpose constitution into an interpretable and tunable reward model, using constitution-guided AI feedback to estimate initial weights for the constituent rubric items. We then reweight those dimensions to construct modified rewards for training. Across experiments on political alignment and safety-helpfulness tradeoffs, reweighting individual dimensions predictably changes targeted behaviors largely independently while navigating tradeoffs between conflicting alignment objectives. We show that the same framework can mitigate label bias encoded in preference judgments -- including sycophancy and demographic bias -- by reducing their influence on the training reward. Our results demonstrate that constitution-derived, interpretable rewards can translate high-level alignment principles into more transparent and controllable model behavior.
Chinese Translation
当前对齐语言模型的方法通常难以了解哪些行为获得奖励,也难以有针对性地改变奖励。特别是,标准的基于偏好的方法将多种考量因素合并为总体人类判断,掩盖了驱动所得奖励的因素,而基于原则的方法指定了高层次价值观,但没有完全将其操作化。为了弥补这一差距,我们开发了一个基于评分标准的框架,将通用宪法转化为可解释且可调优的奖励模型,使用宪法引导的AI反馈来估计组成评分标准项目的初始权重。然后,我们重新加权这些维度,为训练构建修改后的奖励。在关于政治对齐和安全-有用性权衡的实验中,重新加权各个维度可预测地改变目标行为,且很大程度上独立于彼此,同时在相互冲突的对齐目标之间进行权衡。我们表明,同一框架可以通过减少偏好判断中编码的标签偏差(包括谄媚和人口统计学偏见)对训练奖励的影响来缓解这些偏差。我们的结果表明,源自宪法的可解释奖励可以将高层次对齐原则转化为更透明且可控的模型行为。
cs.LG / 224 / 2609.33089

Geometry-Aware Operator Families for Structured Representation Learning

面向结构化表示学习的几何感知算子族
Zhang, Zuyuan, Yu, Fei Xu, Lan, Tian
Abstract
The geometry of latent representations governs which components should interact and how information should propagate, making geometry-aware operator design a fundamental ingredient of structured deep representation learning. However, existing neural architectures typically rely on generic operator templates or geometry-specific constructions, creating a need for a unified framework that can derive admissible operators directly from fixed structural information while remaining adaptive to changing contexts. We introduce \emph{Geometry-Induced Operator Families} (GIOF), a general framework that converts fixed geometry into a structured family of propagation operators and dynamically selects an appropriate member of this family according to the current context. GIOF first transforms geometry-derived interaction channels into reusable generator bases, then combines them through a context-dependent selector and adaptive propagation scale, and finally realizes the selected operator through stable continuous-time propagation and a bottleneck residual layer. We establish theoretical guarantees covering parameter compression, identifiability, stability, locality, compositional structure, and oversmoothing behavior, while controlled experiments validate these mechanisms and experiments on PEMS-BAY and METR-LA achieve the lowest mean MAE across all reported regional-outage settings, improving over the strongest retained baseline by 2.4\%--8.8\% at 30\% missing sensors.
Chinese Translation
潜在表示的几何结构决定了哪些组件应该交互以及信息应如何传播,使得几何感知的算子设计成为结构化深度表示学习的基本要素。然而,现有的神经架构通常依赖于通用算子模板或特定几何的构造,因此需要一个统一的框架,能够直接从固定结构信息中导出可容许的算子,同时对变化的上下文保持适应性。我们引入了几何诱导算子族(GIOF),这是一个通用框架,将固定几何转换为结构化的传播算子族,并根据当前上下文动态选择该族中合适的成员。GIOF首先将几何导出的交互通道转换为可重用的生成器基,然后通过上下文相关的选择器和自适应传播尺度将它们组合起来,最后通过稳定的连续时间传播和瓶颈残差层实现所选算子。我们建立了理论保证,涵盖参数压缩、可识别性、稳定性、局部性、组合结构和过度平滑行为,同时受控实验验证了这些机制,在PEMS-BAY和METR-LA上的实验在所有报告的区域中断设置中实现了最低的平均MAE,在30%传感器缺失时比最强的保留基线提升了2.4%--8.8%。
cs.LG / 225 / 2609.33093

How Linear Attention Remembers

线性注意力如何记忆
Lee, Kichang, Park, JaeYeon, Kim, Songkuk, Ko, JeongGil
Abstract
Linear attention replaces the growing key--value (KV) cache of standard attention with a fixed-size recurrent state, substantially reducing memory growth with context length. This efficiency, however, changes how past information is stored: many tokens must share and repeatedly update the same memory. We study how this recurrent state functions as a memory system. Using an analytical decomposition together with controlled causal interventions in pretrained GLA and GDN models, we trace how recalled information is written, retained, and later accessed. We find that fact-specific information enters recurrent memory through concentrated, content-dependent writes and is later accessed through concentrated query-time read pathways. Multiple facts can remain selectively accessible within the same state, yet their internal representations exhibit cross-fact causal coupling rather than independent KV-like storage. As memory load increases, both recall and targeted editability degrade, whereas elapsed context alone has a substantially smaller effect within the tested regime. Causal interventions on subsequent writes further show that interference is shaped by their overlap with existing memory. Finally, in hybrid architectures that combine recurrent layers with full attention, the runtime memory directly supporting recall shifts predominantly to the full-attention KV state. Together, these results reveal how fixed-size recurrent memory supports selective recall despite shared storage, while exposing the interference and capacity limits that distinguish it from token-addressable KV memory.
Chinese Translation
线性注意力用固定大小的循环状态取代了标准注意力中不断增长的键值(KV)缓存,从而大幅减少了随上下文长度增长的内存开销。然而,这种效率改变了过去信息的存储方式:许多 token 必须共享并反复更新同一份记忆。我们研究这种循环状态如何作为一个记忆系统发挥作用。通过分析性分解以及在预训练的 GLA 和 GDN 模型中进行受控因果干预,我们追踪了被回忆的信息是如何写入、保留以及随后被访问的。我们发现,事实特定的信息通过集中的、依赖于内容的写入进入循环记忆,并在之后通过集中的查询时读取路径被访问。多个事实可以在同一状态中保持选择性可访问,但它们的内部表示表现出跨事实的因果耦合,而非独立的类 KV 存储。随着记忆负载增加,回忆和针对性编辑能力都会下降,而在测试范围内,仅经过的上下文本身的影响要小得多。对后续写入的因果干预进一步表明,干扰程度取决于它们与现有记忆的重叠。最后,在将循环层与全注意力相结合的混合架构中,直接支持回忆的运行时记忆主要转移到全注意力的 KV 状态。总之,这些结果揭示了固定大小的循环记忆如何在共享存储的情况下支持选择性回忆,同时暴露了使其区别于 token 可寻址 KV 记忆的干扰和容量限制。
cs.LG / 226 / 2609.33106

CARVE: Breaking Data Barriers in Chip Placement by Harnessing Reusable Expertise

CARVE:利用可复用专长打破芯片布局中的数据障碍
Zhang, Jiefu, Sun, Haixiang, Xu, Yang, Aggarwal, Vaneet, Wan, Zishen
Abstract
Pretrained macro-placement policies can reduce repeated optimization across circuits, but deployment often exposes them to unfamiliar designs when the original training data are unavailable. Repeatedly fine-tuning a single serving model can overwrite earlier improvements, while simply saving checkpoints does not determine where they can be reliably reused. We introduce Continual Adaptation through the Reuse of Validated Expertise (CARVE), a framework that represents accumulated expertise as a frozen base policy, immutable specialists, and task-specific credentials obtained through local validation. For a new task, CARVE first checks existing specialists and trains a new specialist from the frozen base only when none qualifies. Under fixed task distributions and validation rules that control cumulative error, we establish expected-performance guarantees for repeated reuse. For bounded losses, we also derive matching worst-case bounds on the local samples needed for reliable reuse. In macro placement, a reuse-first follow-up reduces recorded training time by 58.5% (9.66 to 4.01 hours), while mean HPWL gain changes only from 8.41% to 7.86%. In a simulated receiving deployment, imported specialists are reused on six of seven new IBM circuits with no receiver-side training, achieving a 5.76% mean HPWL gain. Navigation studies provide complementary evidence on repair retention and repeated adaptation.
Chinese Translation
预训练的宏布局策略可以减少跨电路的重复优化,但当原始训练数据不可用时,部署往往使其面对不熟悉的设计。反复微调单个服务模型可能会覆盖早先的改进,而仅保存检查点并不能确定它们可以在何处被可靠地重用。我们提出了通过重用已验证专业知识进行持续适应(CARVE)框架,该框架将积累的专业知识表示为冻结的基础策略、不可变的专家以及通过本地验证获得的任务特定凭证。对于新任务,CARVE 首先检查现有专家,仅当没有专家符合条件时,才从冻结的基础策略训练新的专家。在固定任务分布和控制累积误差的验证规则下,我们为重复重用建立了预期性能保证。对于有界损失,我们还推导出了可靠重用所需的本地样本的匹配最坏情况边界。在宏布局中,重用优先的后续方法将记录的训练时间减少了 58.5%(从 9.66 小时减少到 4.01 小时),而平均 HPWL 增益仅从 8.41% 变化到 7.86%。在模拟的接收部署中,导入的专家在七个新的 IBM 电路中的六个上被重用,无需接收方训练,实现了 5.76% 的平均 HPWL 增益。导航研究提供了关于修复保留和重复适应的补充证据。
cs.LG / 227 / 2609.33110

D-JEPA: Design-Recoverable JEPA Representation with Swappable Physics Decoders

D-JEPA:具有可交换物理解码器的设计可恢复JEPA表示
Kulkarni, Nitin Nagesh, Mishra, Aashwin Anand, Yu, Yin, Lyu, Peter
Abstract
Joint-Embedding Predictive Architectures (JEPAs) provide a framework for learning compact representations without directly reconstructing high-dimensional observations. However, in parameterized physical systems, learned representations can entangle geometry with operating conditions and task-specific physical responses, limiting their reuse across prediction tasks. We introduce D-JEPA (Design-recoverable JEPA), a geometry-centric JEPA that computes a compact representation from geometry alone and reuses it across operating conditions and physical response spaces through lightweight physics-specific decoders. An explicit design-recoverability objective encourages the geometry latent to preserve information about the underlying design variables, enabling the representation to support design analysis and optimization. We further identify a case-level collapse failure mode in which target representations become nearly invariant across distinct geometries despite low reconstruction error, and mitigate it using case-level variation constraints and auxiliary target reconstruction. Across four 3D aerodynamic, hydrodynamic, and structural benchmarks, D-JEPA maintains or improves full-field prediction accuracy while achieving near-perfect linear recoverability of design parameters. The frozen geometry representation can be reused at held-out operating conditions and transferred to a structural response task with fewer trainable parameters. Finally, the representation supports differentiable design optimization, with designs validated using high-fidelity CFD, preserving the predicted ranking of candidate designs. These results demonstrate that separating a reusable geometry representation from physics-specific prediction provides a practical representation for scientific surrogate modeling and design.
Chinese Translation
联合嵌入预测架构(JEPA)提供了一个学习紧凑表示的框架,无需直接重建高维观测。然而,在参数化物理系统中,学习到的表示可能会将几何形状与工况条件和任务特定的物理响应纠缠在一起,限制了它们在预测任务间的重用。我们提出D-JEPA(设计可恢复JEPA),一种以几何为中心的JEPA,仅从几何形状计算紧凑表示,并通过轻量级的物理特定解码器在工况条件和物理响应空间之间重用。一个明确的设计可恢复性目标鼓励几何潜在表示保留关于底层设计变量的信息,使该表示能够支持设计分析和优化。我们进一步识别出一种案例级崩溃失败模式,其中目标表示在不同几何形状之间几乎不变,尽管重建误差较低,并使用案例级变化约束和辅助目标重建来缓解它。在四个三维空气动力学、水动力学和结构基准上,D-JEPA保持或提高了全场预测精度,同时实现了设计参数的近乎完美的线性可恢复性。冻结的几何表示可以在留出工况条件下重用,并以更少的可训练参数迁移到结构响应任务。最后,该表示支持可微分设计优化,使用高保真CFD验证设计,保持候选设计的预测排序。这些结果表明,将可重用的几何表示与物理特定预测分离,为科学代理建模和设计提供了一种实用的表示。
cs.LG / 228 / 2609.33112

Simulation-Free Learning of GP-SDEs from Irregular Observations

从不规则观测中无模拟学习GP-SDEs
Lin, Zhidi, Liu, Yuhao, Li, Ying, Fong, Edwin, Djurić, Petar
Abstract
Gaussian process stochastic differential equations (GP-SDEs) provide a flexible Bayesian model for unknown continuous-time state dynamics with uncertainty quantification, but learning and inference from noisy and irregular observations remain computationally challenging. To address this issue, we propose GP-SDE Matching, a simulation-free variational framework for Bayesian GP drift learning and continuous-time state smoothing. We analytically marginalize the sparse GP posterior to derive a tractable drift-matching objective that accounts for both the posterior mean and uncertainty of the unknown drift. To handle irregular observations, we further introduce an irregular-time-aware variational state posterior that incorporates the actual observation times during both encoding and continuous-time marginal querying. Experiments on the stochastic Lorenz--63 system demonstrate substantially improved drift recovery and state reconstruction under irregular observations, while five system identification benchmarks show robust forecasting under increasing observation sparsity and competitive performance against existing latent-SDE and state-space methods.
Chinese Translation
高斯过程随机微分方程(GP-SDEs)为未知的连续时间状态动态提供了一个灵活的贝叶斯模型,并具备不确定性量化能力,但从带噪和不规则观测中进行学习和推断在计算上仍然具有挑战性。为了解决这个问题,我们提出了GP-SDE Matching,一个用于贝叶斯GP漂移学习和连续时间状态平滑的无模拟变分框架。我们通过解析地边缘化稀疏GP后验,推导出一个可处理的漂移匹配目标,该目标同时考虑了未知漂移的后验均值和不确定性。为了处理不规则观测,我们进一步引入了一个不规则时间感知的变分状态后验,它在编码和连续时间边缘查询过程中都融入了实际的观测时间。在随机Lorenz-63系统上的实验表明,在不规则观测下,漂移恢复和状态重建得到了显著改善,而五个系统辨识基准则展示了在观测稀疏性增加时的稳健预测能力,并且与现有的潜在SDE和状态空间方法相比具有竞争性性能。
cs.LG / 229 / 2609.33117

ECG-Scroll: A Long-Horizon, Streaming Benchmark and Agent Environment for Interpretation of Ambulatory Electrocardiograms

ECG-Scroll:一个用于动态心电图解读的长时程、流式基准与智能体环境
Li, Haitao, Li, Chenglin, Ding, Zhengyao, Li, Ziyu, Mao, Yiheng, Huang, Zhengxing
Abstract
Multimodal large language models (MLLMs) can now interpret a standard ten-second, twelve-lead electrocardiogram (ECG) with clinically grounded, reward-verified reasoning. Real cardiac monitoring is different. Ambulatory (Holter) and telemetry recordings span hours to days and are read as they stream in, and their clinically decisive findings are paroxysmal, brief episodes buried in an otherwise unremarkable trace. Such a recording cannot be held in one context at diagnostic resolution, and its future has not yet happened, so a reader must work online, deciding what to measure now, committing evidence to memory as it passes, and reporting events as they occur. We recast long-duration ECG interpretation as a long-horizon, online (streaming, causal) sequential decision process and introduce ECG-Scroll. As a benchmark, long ambulatory recordings are streamed to an agent chunk by chunk, and it must localize, quantify, and promptly flag paroxysmal events without access to future signal; because the underlying signal is retained, every answer is checkable against objective ground truth, giving rule-based rather than judge-based rewards, and the streaming formulation adds a metric batch evaluation cannot express, the detection latency between an event's onset and the moment the agent records it. As an agent environment, it is a fixed, gym-style interaction layer that exercises three competencies single-glance ECG models never touch: Memory, Tool use through signal-grounded measurement rather than reading pixels, and Planning of what to measure now and when to commit. We release 390 whole-recording instances spanning 2,536 hours of two-lead ambulatory ECG and evaluate a signal-threshold rule agent alongside off-the-shelf LLM agents online, characterizing how they use memory, tools, and planning and where the benchmark's head-room lies.
Chinese Translation
多模态大语言模型(MLLMs)现在能够解读标准的十秒十二导联心电图(ECG),并具备临床依据和奖励验证的推理能力。真实的心脏监测则不同。动态(Holter)和遥测记录持续数小时至数天,并随着信号流入而实时读取,其临床决定性发现是阵发性的短暂事件,埋藏在原本无明显异常的波形中。这样的记录无法在单一上下文中以诊断分辨率完整保持,且其未来尚未发生,因此读取者必须在线工作,决定现在测量什么,在证据经过时将其存入记忆,并在事件发生时进行报告。我们将长时程心电图解读重新构建为一个长时程、在线(流式、因果)的序列决策过程,并提出了 ECG-Scroll。作为一个基准,长时程动态记录以逐块方式流式传输给智能体,它必须在无法访问未来信号的情况下定位、量化并及时标记阵发性事件;由于底层信号被保留,每个答案都可以对照客观真值进行检查,从而提供基于规则而非基于评判的奖励,并且流式表述增加了一个批量评估无法表达的指标,即事件起始与智能体记录时刻之间的检测延迟。作为一个智能体环境,它是一个固定的、gym 风格的交互层,锻炼了单次判读心电图模型从未触及的三项能力:记忆、通过基于信号的测量而非读取像素的工具使用,以及规划现在测量什么和何时提交。我们发布了 390 个完整记录实例,涵盖 2,536 小时的双导联动态心电图,并在线评估了一个信号阈值规则智能体和现成的 LLM 智能体,描述了它们如何使用记忆、工具和规划,以及基准的余量所在。
cs.LG / 230 / 2609.33120

Mycelium: A Generalizable Cross-Grid Multi-Task Model for Electrical Distribution Systems

Mycelium:一种用于配电系统的可泛化跨电网多任务模型
Wei, Zhengyang, Bose, Shourya, Hilmarsson, Helgi, Carnio, Elena, Suri, Dhruv
Abstract
Electrical distribution grid operations require inference across heterogeneous networks from sparse, noisy, and incomplete time series measurements. In this work, we identify challenges and explore solutions towards a unified model that can perform diverse tasks grounded in the physics of the electric grid and generalize to unseen distribution networks. We define a unified grid ontology that represents variable sized distribution networks as heterogeneous graphs while preserving native topology, asset types, and electrical relationships across networks. We develop a physics based data simulation pipeline that combines reference and procedurally generated distribution networks with network reconfigurations, fault scenarios, and configurable sensing conditions. We present Mycelium, a heterogeneous graph transformer with structure aware communication edges and electrical reference features that encode network position and nominal phase orientation, together with task specific temporal readouts which generate per task outputs. We train Mycelium on reference as well as synthetic grids, and study its generalization on benchmark networks completely excluded from training and validation. Mycelium is observed to outperform task specific neural baselines on most reported benchmark metrics. Architectural ablations and the aforementioned studies reveal Mycelium's capability to learn representations of the underlying physics which serves to enhance cross-task performance, thereby addressing a significant challenge in unified grid models.
Chinese Translation
配电电网运营需要从稀疏、噪声和不完整的时间序列测量中跨异构网络进行推理。在这项工作中,我们识别了挑战,并探索了朝着统一模型的解决方案,该模型能够执行基于电网物理的多样化任务,并泛化到未见过的配电网络。我们定义了一个统一的电网本体,将可变规模的配电网络表示为异构图,同时保留跨网络的原始拓扑、资产类型和电气关系。我们开发了一个基于物理的数据模拟流程,将参考网络和程序生成的配电网络与网络重构、故障场景和可配置的传感条件相结合。我们提出了Mycelium,一种异构图Transformer,具有结构感知的通信边和编码网络位置与标称相位方向的电气参考特征,以及生成每个任务输出的任务特定时间读出。我们在参考电网和合成电网上训练Mycelium,并在完全排除在训练和验证之外的基准网络上研究其泛化能力。观察到Mycelium在大多数报告的基准指标上优于任务特定的神经网络基线。架构消融和上述研究揭示了Mycelium学习底层物理表示的能力,这有助于增强跨任务性能,从而解决了统一电网模型中的一个重大挑战。
cs.LG / 231 / 2609.33126

Save Your Saturated Data: Learning Beyond Reward Saturation in Group-Based RL

拯救你的饱和数据:在基于组的强化学习中超越奖励饱和进行学习
Yang, Ziyuan, Wang, Yike, Feng, Shangbin, Tsvetkov, Yulia
Abstract
Group-relative reinforcement learning (RL) relies on reward variation among sampled responses to estimate informative relative advantages. As language models become increasingly capable, existing training data can become reward-saturated: all sampled responses to the same problem might receive equally high rewards, where the group-relative learning signals vanish and leave previously useful data obsolete. In this work, we investigate whether useful learning signals can be recovered from such saturated data. We study interventions at four levels of group-policy RL pipelines---data, rollout, reward, and advantage---and conduct extensive RL training on saturated reasoning data only. While standard GRPO on saturated data would almost always yield near-0 advantages and near-noise signals, diverse interventions successfully recycle and repurpose such data: among the proposed strategies, interventions at rollout generation are consistently most effective: nudging the policy to generate ``high-quality'', incorrect solutions introduces rollouts with poor rewards into saturated groups as negative samples, which turns out to improve GRPO by 6.4% to 9.0% across Qwen3-1.7B and 4B. Other interventions such as increasing rollout temperature or adding auxiliary rewards can also restore non-zero advantages, but yield less consistent gains. Further analyses show that effective negative rollouts require informative negative trajectories, that the method remains effective alongside unsaturated data, and that it supports iterative recycling of newly saturated examples. While increasingly stronger LLMs would render more data as saturated, our results demonstrate that don't waste your saturated data: with the right strategies they can be recycled into useful RL training signals in an increasingly data-scarce world.
Chinese Translation
组相对强化学习(RL)依赖采样回答之间的奖励差异来估计有信息量的相对优势。随着语言模型能力日益增强,现有训练数据可能变得奖励饱和:同一问题的所有采样回答可能获得同样高的奖励,此时组相对学习信号消失,先前有用的数据也变得过时。在本工作中,我们研究能否从这类饱和数据中恢复有用的学习信号。我们研究了组策略强化学习流程四个层面的干预——数据、rollout、奖励和优势——并仅使用饱和的推理数据进行大量 RL 训练。虽然标准 GRPO 在饱和数据上几乎总是产生接近 0 的优势和接近噪声的信号,但多种干预能够成功回收并重新利用这类数据:在所提出的策略中,对 rollout 生成阶段的干预始终最有效:引导策略生成“高质量”但错误的解答,会将奖励较差的 rollout 作为负样本引入饱和组中,结果使 GRPO 在 Qwen3-1.7B 和 4B 上提升 6.4% 到 9.0%。其他干预,如提高 rollout 温度或添加辅助奖励,也能恢复非零优势,但带来的增益不那么稳定。进一步分析表明,有效的负向 rollout 需要具有信息量的负向轨迹;该方法在与未饱和数据一起使用时仍然有效;并且支持对新近饱和样本的迭代回收。虽然越来越强的 LLM 会使更多数据变为饱和,但我们的结果表明,不要浪费你的饱和数据:在数据日益稀缺的世界中,通过正确的策略,它们可以被回收为有用的 RL 训练信号。
cs.LG / 232 / 2609.33127

Policy Plasticity Matters in Offline-to-Online Reinforcement Learning: Refitting Offline Policies for Online Adaptation

策略可塑性在离线到在线强化学习中至关重要:为在线适应重新拟合离线策略
Huang, Yuheng, Qing, Yunpeng, Chi, Yixiao, Kong, Yilun, Zou, Changqing
Abstract
Offline-to-Online Reinforcement Learning (O2O RL) has emerged as a practical paradigm that pre-trains the policy using static offline datasets and subsequently adapts the policy through online interactions. Existing O2O methods primarily address the transition through value calibration, while generally treating the offline-trained policy as a given initialization. We instead study O2O adaptation from the perspective of network plasticity, asking whether the offline-trained policy remains sufficiently adaptable for online learning. Controlled experiments show that prolonged optimization on static offline data progressively reduces network plasticity even after offline performance has largely saturated, and that lower plasticity is associated with weaker subsequent online improvement. Motivated by these observations, we propose REstoring plasticity via Fresh Initialization and policy Transfer (REFIT), a lightweight model-level method for the O2O transition. Before online fine-tuning, REFIT distills the offline policy into a freshly initialized student while temporarily freezing a random subset of student units, transferring the learned offline behavior to a more plastic policy initialization. Extensive experiments on D4RL and OGBench demonstrate that REFIT consistently achieves higher aggregate performance than existing O2O plug-in methods across both Cal-QL and IQL backbones, while plasticity diagnostics and ablations provide further evidence of restored network plasticity.
Chinese Translation
离线到在线强化学习(O2O RL)已成为一种实用范式,它利用静态离线数据集预训练策略,随后通过在线交互对策略进行适应。现有的O2O方法主要通过价值校准来处理过渡,同时通常将离线训练的策略视为给定的初始化。我们转而从网络可塑性的角度研究O2O适应,探究离线训练的策略是否保持足够的可适应性以进行在线学习。对照实验表明,即使在离线性能基本饱和之后,对静态离线数据进行长时间优化也会逐渐降低网络可塑性,并且较低的可塑性与后续在线改进较弱相关。受这些观察结果的启发,我们提出了通过全新初始化和策略迁移恢复可塑性(REFIT),这是一种用于O2O过渡的轻量级模型级方法。在在线微调之前,REFIT将离线策略蒸馏到一个全新初始化的学生网络中,同时暂时冻结学生单元的一个随机子集,将学习到的离线行为迁移到更具可塑性的策略初始化中。在D4RL和OGBench上的大量实验表明,在Cal-QL和IQL两种主干网络上,REFIT始终比现有的O2O插件方法获得更高的总体性能,同时可塑性诊断和消融实验进一步证明了网络可塑性的恢复。
cs.LG / 233 / 2609.33129

Flow-Matching-Based Protein Structure Tokenizer Made Efficient and Easy

基于流匹配的蛋白质结构标记器:高效且易用
Zhang, Zhe, Zhang, Yikai, Feng, Jiangtao, Zhang, Ya-Qin, Ma, Wei-Ying, Zhou, Hao
Abstract
As the bridge between protein modality and discrete modeling, protein structure tokenization still largely relies on heavily engineered training objectives tailored to specific downstream tasks and large training datasets, which hinders its transfer to broader application scenarios. To address this issue, we propose ProFiT, a lightweight flow matching tokenizer. With simple training strategies that encourage healthy codebook utilization, ProFiT can be trained efficiently and naturally learns semantically meaningful representations without any manual semantic alignment, while achieving reconstruction quality and generalization that match or surpass those of substantially larger tokenizers. We conduct extensive evaluations across a wide range of settings and demonstrate that ProFiT is a plug-and-play tokenizer adaptable to diverse downstream tasks. This study further reveals the significant potential of the flow matching tokenizer paradigm. Our code is publicly available at https://github.com/QDKStorm/ProFiT.
Chinese Translation
作为蛋白质模态与离散建模之间的桥梁,蛋白质结构标记化仍在很大程度上依赖于针对特定下游任务和大型训练数据集精心设计的训练目标,这阻碍了其向更广泛应用场景的迁移。为了解决这个问题,我们提出了 ProFiT,一种轻量级的流匹配标记器。通过鼓励健康码本利用的简单训练策略,ProFiT 可以高效训练,并自然地学习到语义上有意义的表示,无需任何手动语义对齐,同时达到与规模大得多的标记器相当或超越的重建质量和泛化能力。我们在广泛的设置下进行了大量评估,并证明 ProFiT 是一种即插即用的标记器,可适应多样化的下游任务。这项研究进一步揭示了流匹配标记器范式的巨大潜力。我们的代码公开在 https://github.com/QDKStorm/ProFiT。
cs.LG / 234 / 2609.33131

ILP-BO: Integer Linear Programming-Based Black-Box Optimization

ILP-BO:基于整数线性规划的黑箱优化
Nakada, Hyakka, Tanaka, Shu
Abstract
Black-box Optimization (BO) is a powerful framework for optimizing expensive objective functions or unknown functions with a limited number of evaluations. A central step of standard BO such as Bayesian optimization is the optimization of a surrogate-based acquisition criterion, which is commonly performed using nonlinear optimization or heuristic search. Therefore, conventional black-box optimization generally does not guarantee global optimality in candidate selection. In this study, we propose Integer Linear Programming-based Black-box Optimization (ILP-BO), a quasi-Bayesian optimization framework that transforms kernel-based surrogate optimization over discrete domains into an Integer Linear Programming (ILP) problem. The key idea is to represent nonlinear kernel functions exactly on finite discrete distance levels by introducing binary one-hot auxiliary variables. This transformation converts the nonlinear surrogate into a linear objective with linear constraints and binary variables. To incorporate exploration while preserving the linear structure, we further introduce a Hamming-distance margin that excludes neighborhoods around previously observed points. We derive the proposed formulation for several standard kernels and obtain an analytical upper bound on the Hamming-distance threshold based on the measure in the binary search space. The resulting candidate-selection problem can be solved by integer programming solvers with certificates of optimality. Thus, our methodology has the potential to serve as a highly transparent black-box optimization framework. Experiments on synthetic and discrete optimization benchmarks show that ILP-BO achieves competitive optimization performance compared with practical Bayesian optimization methods.
Chinese Translation
黑箱优化(BO)是一种强大的框架,用于在有限评估次数下优化代价高昂或未知的目标函数。标准BO(如贝叶斯优化)的一个核心步骤是优化基于代理模型的采集准则,这通常通过非线性优化或启发式搜索来执行。因此,传统的黑箱优化在候选选择中通常无法保证全局最优性。在本研究中,我们提出基于整数线性规划的黑箱优化(ILP-BO),这是一种准贝叶斯优化框架,将离散域上基于核的代理优化转化为整数线性规划(ILP)问题。其核心思想是通过引入二元独热辅助变量,在有限离散距离层级上精确表示非线性核函数。该转化将非线性代理模型转换为带有线性约束和二元变量的线性目标函数。为了在保持线性结构的同时融入探索,我们进一步引入一个汉明距离裕度,以排除先前观测点周围的邻域。我们针对几种标准核推导了所提出的公式,并基于二元搜索空间中的测度得到了汉明距离阈值的解析上界。由此产生的候选选择问题可由整数规划求解器求解,并带有最优性证书。因此,我们的方法有潜力成为一个高度透明的黑箱优化框架。在合成和离散优化基准上的实验表明,与实用的贝叶斯优化方法相比,ILP-BO取得了具有竞争力的优化性能。
cs.LG / 235 / 2609.33136

Leaky Students: Membership Inference against On-Policy Distillation

Leaky Students:针对同策略蒸馏的成员推理
Lu, Zhexi, Zhu, Mingzhi, Patterson, Stacy, Yu, Lei
Abstract
On-policy distillation (OPD) trains a student to match a teacher's next-token distributions on student-generated trajectories. However, privileged information supplied to the teacher for OPD training may contain sensitive data. Whether the student leaks private information about the records supplied to the teacher during distillation remains poorly understood. To the best of our knowledge, we present the first systematic study of membership inference in this setting. We find that fresh student trajectories expose sparse membership signals that fixed reference-answer losses often miss. These signals are mixed with probability changes caused by training on other records. We introduce Leaky, which samples fresh trajectories from the target model and compares its token log-probabilities with the maximum across matched reference models trained without the candidate records. It applies Leaky ReLU to the resulting gaps, preserving positive gaps and downweighting negative gaps as an approximate correction for incidental positive gaps in non-members. Across fifteen targets spanning mathematics, medical question answering, and code generation, Leaky outperforms all evaluated baselines and achieves mean AUROC 0.875, compared with 0.614 for the strongest baseline on each target in the main evaluation. On the same sampled trajectories, the strongest baseline achieves mean AUROC 0.826. These results show that students trained through OPD can expose the membership of records used for teacher supervision, even when fixed reference-answer losses provide little evidence of membership.
Chinese Translation
同策略蒸馏(On-policy distillation, OPD)训练学生模型在学生生成的轨迹上匹配教师模型的下一个词元分布。然而,为 OPD 训练提供给教师模型的特权信息可能包含敏感数据。学生是否会泄漏蒸馏过程中提供给教师模型的记录的隐私信息,目前仍知之甚少。据我们所知,我们首次对这一场景下的成员推理进行了系统性研究。我们发现,新生成的学生轨迹暴露了稀疏的成员信号,而固定的参考答案损失往往无法捕捉这些信号。这些信号与其他记录训练引起的概率变化混杂在一起。我们提出了 Leaky,它从目标模型采样新轨迹,并将其词元对数概率与在未使用候选记录训练的匹配参考模型上的最大值进行比较。它对得到的差距应用 Leaky ReLU,保留正差距,并降低负差距的权重,作为对非成员中偶然正差距的近似校正。在涵盖数学、医学问答和代码生成的十五个目标上,Leaky 优于所有评估的基线,平均 AUROC 达到 0.875,而在主评估中每个目标上最强基线为 0.614。在相同的采样轨迹上,最强基线达到平均 AUROC 0.826。这些结果表明,通过 OPD 训练的学生能够暴露用于教师监督的记录的成员身份,即使固定的参考答案损失几乎不提供成员证据。
cs.LG / 236 / 2609.33144

Beyond the Training Horizon: Mechanisms and Limits of Length Generalization in Looped Transformers

超越训练视野:Looped Transformer 中长度泛化的机制与局限
Liang, Jia, Jin, Xi, Pan, Liangming
Abstract
Looped Transformers can generalize to reasoning chains longer than those encountered during training, but the computations enabling this behavior and limiting its extent remain unclear. We mechanistically compare two looped-Transformer configurations, which we call the Matched-Recurrence Looped Transformer (MR-Loop) and Decoupled-Recurrence Looped Transformer (DR-Loop), reflecting their respective recurrence-training schemes. We evaluate polynomial iteration, finite-state composition, and knowledge-graph traversal using detailed mechanistic analysis. Attention analysis, intermediate-state decoding, and causal interventions reveal distinct mechanisms learned under final-answer supervision. MR-Loop updates an intermediate state at a fixed readout while advancing relation selection through adjacent-token interactions and a transferable progress cue. DR-Loop instead propagates intermediate states across relation positions, forming an advancing computational frontier. However, both mechanisms become unreliable at greater depths: MR-Loop exhibits degradation of its readout state and progress cues, while DR-Loop exhibits declining reliability of state propagation. Limited self-correction allows local errors to persist and compound. Across both models, we uncover a common representational principle: recurrent states encode not only task-relevant content but also its computational status, whether that content remains in a form that can support subsequent computation. Transferable live-consumed and fresh-aged residual directions causally control whether represented information can participate in subsequent computation, including beyond the training horizon. We further show that length generalization need not rely on faithful step-by-step reasoning, as Looped Transformers can exploit task structure without explicitly representing every intermediate state.
Chinese Translation
Looped Transformer(循环 Transformer)能够泛化到比训练时遇到的更长的推理链,但实现这种行为并限制其范围的计算仍不清楚。我们从机制上比较了两种循环 Transformer 配置,分别称为匹配递归循环 Transformer(MR-Loop)和解耦递归循环 Transformer(DR-Loop),反映了它们各自的递归训练方案。我们使用详细的机制分析评估了多项式迭代、有限状态组合和知识图谱遍历。注意力分析、中间状态解码和因果干预揭示了在最终答案监督下学习到的不同机制。MR-Loop 在固定读出处更新中间状态,同时通过相邻 token 交互和可迁移的进度线索推进关系选择。DR-Loop 则在关系位置之间传播中间状态,形成推进的计算前沿。然而,两种机制在更大深度下都变得不可靠:MR-Loop 表现出其读出状态和进度线索的退化,而 DR-Loop 表现出状态传播可靠性的下降。有限的自校正允许局部错误持续并累积。在这两个模型中,我们揭示了一个共同的表示原则:循环状态不仅编码任务相关内容,还编码其计算状态,即该内容是否保持能够支持后续计算的形式。可迁移的 live-consumed 和 fresh-aged 残差方向因果性地控制所表示的信息能否参与后续计算,包括超越训练视野。我们进一步表明,长度泛化不必依赖忠实的逐步推理,因为循环 Transformer 可以利用任务结构而无需显式表示每个中间状态。
cs.LG / 237 / 2609.33147

CFLoRA: Federated Fine-tuning of LLMs with Complementary Factors for Error-free Aggregation

CFLoRA:基于互补因子的LLM联邦微调实现无误差聚合
Ma, Yanan, Chen, Qiyuan, Fang, Zihan, Chen, Xianhao, Fang, Yuguang
Abstract
Federated low-rank adaptation (LoRA) enables collaborative fine-tuning of large language models without centralizing private client data. Its factorized update, however, creates a structural mismatch in federated averaging: averaging the two LoRA factors separately does not equal averaging their products. Existing exact methods resolve this issue mainly by freezing an entire factor or alternating factors across rounds, but none can update factors simultaneously without aggregation errors or expanding communication ranks. To address this fundamental problem, we present \texttt{CFLoRA}, a federated LoRA scheme that partitions latent LoRA channels into two complementary sets in every communication round. By ensuring that columns and rows are complementary across factors, we eliminate bilinear terms in matrix multiplications, making federated aggregation exact. Crucially, our framework also supports clients with heterogeneous rank budgets. Convergence analysis validates \texttt{CFLoRA} achieves $\mathcal{O}(1/\sqrt{T})$ convergence rate of the \textit{original} LoRA objective in homogeneous-rank cases. Extensive experiments with RoBERTa on the GLUE benchmark and with LLaMA-3.2-3B-Instruct on commonsense reasoning tasks demonstrate that \texttt{CFLoRA} achieves superior performance and training efficiency compared to state-of-the-art federated LoRA baselines.
Chinese Translation
联邦低秩适应(LoRA)使得无需集中私有客户端数据即可协作微调大型语言模型。然而,其因子化更新在联邦平均中造成了结构不匹配:分别对两个LoRA因子求平均并不等于对它们的乘积求平均。现有的精确方法主要通过冻结整个因子或跨轮次交替因子来解决此问题,但无法在无聚合误差或不扩展通信秩的情况下同时更新因子。为了解决这一根本问题,我们提出了CFLoRA,一种在每个通信轮次中将潜在LoRA通道划分为两个互补集合的联邦LoRA方案。通过确保因子间的列和行互补,我们消除了矩阵乘法中的双线性项,从而使联邦聚合精确。关键的是,我们的框架还支持具有异构秩预算的客户端。收敛性分析验证了CFLoRA在同构秩情况下达到了原始LoRA目标的O(1/√T)收敛速率。在GLUE基准上使用RoBERTa以及在常识推理任务上使用LLaMA-3.2-3B-Instruct的大量实验表明,与最先进的联邦LoRA基线相比,CFLoRA实现了更优的性能和训练效率。
cs.LG / 238 / 2609.33152

Convergence of Practical Muon

实用 Muon 的收敛性
Wang, Haonan, Wu, Yu, Liwang, Minghui, Yi, Xinlei, Hong, Yiguang
Abstract
Muon is emerging as a promising alternative to AdamW for large-scale neural network training, yet theoretical understanding of its practical implementation remains incomplete, as existing analyses often simplify or omit two key components: (i) practical Newton--Schulz iterations with empirically tuned polynomial coefficients $(3.4445,-4.7750,2.0315)$; and (ii) decoupled weight decay for regularization. In this paper, we provide an optimization interpretation and establish convergence for practical Muon, jointly accounting for both components. Specifically, we interpret practical Muon as right-preconditioned optimization of the original loss with a dynamic weighted $\ell_2$ regularizer that vanishes as stationarity is approached, so that the optimization target remains the original objective. We then establish, to our best knowledge, the first convergence guarantee for practical Muon in the stochastic nonconvex setting, with an $\mathcal{O}(T^{-1/4})$ convergence rate in terms of the expected Frobenius norm of the gradient, improving the dimension dependence of the best known AdamW's convergence rate by a factor of $\sqrt{d}$, where $T$ is the iteration horizon and $d$ is the parameter dimension. Experiments further support the theoretical convergence results.
Chinese Translation
Muon 正逐渐成为大规模神经网络训练中 AdamW 的一个有前景的替代方案,然而对其实际实现的理论理解仍不完整,因为现有分析通常简化或忽略两个关键部分:(i) 使用经验调优多项式系数 $(3.4445,-4.7750,2.0315)$ 的实际 Newton--Schulz 迭代;以及 (ii) 用于正则化的解耦权重衰减。在本文中,我们提供了优化解释,并建立了实用 Muon 的收敛性,同时考虑这两个部分。具体来说,我们将实用 Muon 解释为对原始损失进行右预条件优化,并带有动态加权 $\ell_2$ 正则化项,该正则化项在接近平稳点时消失,因此优化目标保持为原始目标。然后,据我们所知,我们首次在随机非凸设置下建立了实用 Muon 的收敛保证,其期望梯度 Frobenius 范数下的收敛速率为 $\mathcal{O}(T^{-1/4})$,将已知最佳 AdamW 收敛速率的维度依赖性改善了 $\sqrt{d}$ 倍,其中 $T$ 是迭代步数,$d$ 是参数维度。实验进一步支持了理论收敛结果。
cs.LG / 239 / 2609.33166

Downstream-Aware Context Selection for Online In-Context Reinforcement Learning

面向在线上下文强化学习的下游感知上下文选择
Li, Ruihan A., Zhang, Shangtong, Chandra, Rohan
Abstract
In-context reinforcement learning (ICRL) enables large language model agents to adapt to new environments using their interaction history without updating model parameters. However, repeatedly conditioning on growing histories can lead to substantial token cost. We propose a bounded-history context-management framework that predicts the task-dependent downstream effect of removing historical interactions to guide history selection and determine a decision-dependent context budget. Formally, our framework uses the full rolling history as a reference. The predictor evaluates removal effects, defines a deletion ordering, and applies a shared selection criterion to determine how much history to retain at each decision. We evaluate the method in closed-loop SUMO driving under held-out in-distribution, unseen-domain, and unseen-route settings, and in ScienceWorld under a continual ICRL protocol. Relative to a baseline using the full context, our method reduces total token usage by 25.7%, 25.8%, and 23.2% across the three driving settings while maintaining comparable closed-loop driving performance. In ScienceWorld, it reduces total token usage by 52.1% compared to full context and uses 30.2% and 37.8% fewer tokens than the Recent and Similarity baselines, respectively, while maintaining performance.
Chinese Translation
上下文强化学习(ICRL)使大语言模型智能体能够利用其交互历史适应新环境而无需更新模型参数。然而,反复以不断增长的历史为条件可能导致显著的token开销。我们提出了一种有界历史的上下文管理框架,该框架预测移除历史交互对任务相关的下游影响,以指导历史选择并确定依赖于决策的上下文预算。形式上,我们的框架使用完整的滚动历史作为参考。预测器评估移除效应,定义删除顺序,并应用共享的选择准则来确定在每个决策点保留多少历史。我们在闭环SUMO驾驶环境中评估了该方法,包括留出的分布内、未见领域和未见路线设置,以及在ScienceWorld中,在持续ICRL协议下。相对于使用完整上下文的基线,我们的方法在三个驾驶设置中分别减少了25.7%、25.8%和23.2%的总token使用量,同时保持了可比的闭环驾驶性能。在ScienceWorld中,与完整上下文相比,它将总token使用量减少了52.1%,并且相比Recent和Similarity基线分别少用了30.2%和37.8%的token,同时保持了性能。
cs.LG / 240 / 2609.33169

When Does Backpropagating Through Policy Memory Matter? Physical Credit, Optimizer Updates, and Observability

何时通过策略记忆进行反向传播才重要?物理信用、优化器更新与可观测性
Li, Xingjian, Han, Yi, Huang, Jianhua Z.
Abstract
Policies with memory can learn along two backward paths: through the physical states their actions produce and through the representations they store. Transformer-XL and truncated backpropagation through time cut the second path at stored history while keeping its values. We ask when this cut matters. Holding the forward computation fixed and varying only derivative edges, we measure parameter gradients, the updates the optimizer applies, and continued training in a Transformer vessel-trajectory model and a quadrotor tracking policy. In the vessel model, detaching the key-value cache shrank the gradient to about a tenth of its norm, with little rotation, when gradients flowed through all earlier physical states, but barely changed it under one-step physical credit. In this strongly clipped regime the optimizer, not the gradient, set how far updates differed: global-norm clipping removed most of the gradient difference between memory-cut graphs, whereas AdamW turned a 2% gradient difference between two placements of the cut into update differences of up to 31% at the step where the placement was switched. In a quadrotor trained from initialization with 0.20 m/s velocity noise, removing memory raised tracking error by 43% and cutting memory gradients raised it by 32%; at low noise the cut's mean cost exceeded the value of memory. Two-step truncation segments gave no measurable gain, although with hidden velocity a two-step window captured most of the value of memory; eight-step segments removed half to three quarters of the cost. Switching the cut on only for the last fifth of training understated its cost about threefold at 0.20-0.30 m/s, but not at low noise or with hidden velocity. These results suggest measuring the cost of a memory cut by training with it from initialization, and comparing backward graphs by the updates the optimizer applies rather than by raw gradients.
Chinese Translation
带有记忆的策略可以沿两条反向路径学习:通过其动作产生的物理状态,以及通过其存储的表示。Transformer-XL 和随时间截断的反向传播(truncated backpropagation through time)在存储历史处切断了第二条路径,同时保留其值。我们探究这种切断何时重要。在保持前向计算固定、仅改变导数边的情况下,我们在一个 Transformer 船舶轨迹模型和一个四旋翼跟踪策略中,测量参数梯度、优化器所应用的更新以及持续训练。在船舶模型中,当梯度流经所有更早的物理状态时,分离键值缓存将梯度缩小到其范数的约十分之一,且旋转很小;但在单步物理信用下,几乎不改变梯度。在这个强裁剪区域,决定更新差异程度的是优化器而非梯度:全局范数裁剪消除了记忆切断图之间的大部分梯度差异,而 AdamW 将切断位置两种放置之间 2% 的梯度差异,在切换该放置的那一步转化为高达 31% 的更新差异。在一个从初始化开始训练、带有 0.20 m/s 速度噪声的四旋翼中,移除记忆使跟踪误差提高了 43%,切断记忆梯度使其提高了 32%;在低噪声下,切断的平均代价超过了记忆的价值。两步截断片段没有带来可测量的增益,尽管在隐藏速度下,两步窗口捕获了记忆的大部分价值;八步片段消除了二分之一到四分之三的代价。仅在训练最后五分之一时开启切断,在 0.20-0.30 m/s 下将其代价低估了约三倍,但在低噪声或隐藏速度下并非如此。这些结果表明,应通过在初始化时就使用记忆切断进行训练来衡量其代价,并通过优化器所应用的更新而不是原始梯度来比较反向图。
cs.LG / 241 / 2609.33171

Perturb-and-Solve: Efficient Learned-Operator Conditioning for Latent Diffusion Inverse Problems

扰动-求解:面向潜在扩散逆问题的高效学习算子条件化
Shtanchaev, Abduragim, Asadulaev, Arip, Labazanova, Luiza, Alimbayev, Aidar, Salta, Karim, Moulines, Eric
Abstract
Latent diffusion models serve as powerful priors for solving inverse problems in image restoration, such as deblurring, inpainting, and super-resolution. Current methods have a trade-off between generality and efficiency. Solvers that are restricted to a fixed set of degradation operators are fast and efficient. Methods that support arbitrary degradation operators are slow and require gradients through the diffusion network. To break this bottleneck, we introduce PASEO (Perturb-And-Solve for Efficient Operator conditioning), a method that uses a small (1M parameters) learned network to degrade diffusion model predictions in latent space. PASEO supports learned degradation operators without back-propagating through the diffusion network. We efficiently sample reconstructions from an approximate posterior by combining the diffusion model's prediction with the observed image. We do this by adding noise and solving linear equations based on a local linear approximation of the learned network, without building or inverting large covariance matrices. Across super-resolution, deblurring, and inpainting on FFHQ and COCO, PASEO achieves strong perceptual quality while running up to 9x faster and using up to 34% less peak memory than the tested baselines, with the same or fewer model evaluations.
Chinese Translation
潜在扩散模型作为强大的先验,用于解决图像恢复中的逆问题,例如去模糊、修复和超分辨率。当前方法在通用性和效率之间存在权衡。局限于固定退化算子集合的求解器快速且高效。支持任意退化算子的方法速度慢,并且需要通过扩散网络进行梯度计算。为了打破这一瓶颈,我们提出了PASEO(Perturb-And-Solve for Efficient Operator conditioning,扰动与求解以实现高效算子条件化),该方法使用一个小型(100万参数)的学习网络在潜在空间中对扩散模型预测进行退化。PASEO支持学习到的退化算子,而无需通过扩散网络进行反向传播。我们通过将扩散模型的预测与观测图像相结合,从近似后验中高效地采样重建结果。我们通过添加噪声并基于学习网络的局部线性近似求解线性方程来实现这一点,而无需构建或求逆大型协方差矩阵。在FFHQ和COCO上的超分辨率、去模糊和修复任务中,PASEO实现了强大的感知质量,同时运行速度比测试基线快高达9倍,峰值内存使用减少高达34%,且模型评估次数相同或更少。
cs.LG / 242 / 2609.33189

When an Evaluation Rule Writes Training Labels: Measuring Human-Reference Forgiveness in NAVSIM

当评估规则生成训练标签:测量NAVSIM中的人类参考宽容
Guo, Jiaxuan, Yang, Jingxin, Ye, Jiaqi, Sun, Youran, Xin, Shuo, Zhang, Kejia, Yang, Haizhao
Abstract
When the human reference scores zero on a metric, the released GTRS-Dense label generator for NAVSIM marks every candidate trajectory in the scene as passing it. NAVSIM's authors introduced this human-reference forgiveness to avoid penalizing contextually justified maneuvers when scoring one trajectory, and warned that it could overlook important failures. In label generation it sets a whole column of 16,384 candidate targets to passing. To measure the consequences for supervision, we re-run the generator with the overwrite disabled and compare the pre-overwrite targets with the released labels on all 103,288 navtrain scenes. The rule erases a candidate distinction that the training loss reads on 11,237 of them (10.8793%). Firing usually changes most of a column: lane keeping carries 9,982 of the 13,042 forgiven loss columns, and its median forgiven column had 14,391 of 16,384 candidates failing before the overwrite. On held-out navtest scenes forgiven on lane keeping, the released lane-keeping head's median AUC against the pre-overwrite outcome is 0.7095; on unforgiven scenes matched on failing-candidate count it is 0.9807. For the Hydra-MDP checkpoint released with GTRS, whose configuration takes the same label file, the two values are 0.6627 and 0.9761. Continuing the released GTRS-Dense checkpoint for 300 optimizer steps with three paired seeds, we observe the forgiven-scene AUC 0.1086-0.1251 higher with pre-overwrite than with published targets, and a narrower gap between matched groups, still above zero. Scoring with forgiveness disabled, we observe lane keeping higher by 2.478-3.524 points on navtest scenes forgiven on any of five loss metrics, with lower adjacent-frame plan consistency. Both changes are larger there than on the rest. EPDMS, scored the same way, does not separate the two target sets.
Chinese Translation
当人类参考在某个指标上得分为零时,为NAVSIM发布的GTRS-Dense标签生成器会将场景中的每个候选轨迹标记为通过该指标。NAVSIM的作者引入这种人类参考宽容,以避免在为一个轨迹评分时惩罚情境合理的机动,并警告说它可能会忽略重要的失败。在标签生成中,它将一整列16,384个候选目标设为通过。为了衡量对监督的影响,我们在禁用覆盖的情况下重新运行生成器,并在所有103,288个navtrain场景上比较覆盖前的目标与发布的标签。该规则抹去了训练损失在11,237个场景(10.8793%)上读取的候选区别。触发通常改变一列中的大部分:车道保持承载了13,042个被宽容的损失列中的9,982个,其中位数被宽容列在覆盖前有14,391个(共16,384个)候选失败。在车道保持上被宽容的留出navtest场景中,发布的lane-keeping头针对覆盖前结果的中位AUC为0.7095;在失败候选数匹配的未被宽容场景上为0.9807。对于随GTRS发布的Hydra-MDP检查点,其配置采用相同的标签文件,这两个值分别为0.6627和0.9761。使用三个配对种子继续训练发布的GTRS-Dense检查点300个优化步骤,我们观察到使用覆盖前目标时被宽容场景AUC比使用发布目标高0.1086-0.1251,且匹配组之间的差距缩小,但仍高于零。在禁用宽容的情况下评分,我们观察到在五个损失指标中任一被宽容的navtest场景上,车道保持高出2.478-3.524点,但相邻帧规划一致性较低。这两种变化在这些场景中比在其他场景中更大。以相同方式评分的EPDMS无法区分这两个目标集。
cs.LG / 243 / 2609.33192

AG-CoT: Verified Algorithmic Traces for LLM Program Synthesis on Clifford Circuits

AG-CoT:面向克利福德电路上LLM程序合成的已验证算法轨迹
Wei, Lu, Wang, Yufeng, Cao, Chenfeng, Pang, Lu, Ling, Haibin
Abstract
Scientific code generation can produce executable programs that fail to compute the intended scientific object. We study this problem in language-model synthesis of Clifford circuits, which prepare the stabilizer states used in quantum error correction and admit exact classical verification. In our target-conditioned framework, each target is given as compact signed stabilizer generators, and an exact verifier checks the generated OpenQASM circuits. We supervise models with Aaronson-Gottesman chain-of-thought (AG-CoT) traces checked by the verifier, and continue training on model generations that the verifier accepts. Across two independently trained model families (3B and 7B), AG-CoT supervision multiplies greedy-decode state-equivalence accuracy by four to six times over circuit-only baselines, and verifier-filtered continuation training adds a further consistent gain atop both. A complementary 32B study shows that supervised models achieve near-perfect syntax and Clifford validity while the strongest direct model reaches 6.14% state equivalence per target, rising to over 10% under verifier-guided selection with multiple candidates. These results show that algorithmic trace supervision gives a large, statistically significant gain in both model families and that verifier-filtered continuation adds a further repeated gain. The persistent gap between Clifford validity and state equivalence confirms that exact verification is necessary: a circuit can be syntactically and physically valid yet prepare the wrong quantum state.
Chinese Translation
科学代码生成可以产生可执行程序,但这些程序可能无法计算预期的科学对象。我们在克利福德电路的的语言模型合成中研究这个问题,这些电路制备用于量子纠错的稳定子态,并且允许精确的经典验证。在我们的目标条件框架中,每个目标以紧凑的有符号稳定子生成元给出,并且一个精确的验证器检查生成的OpenQASM电路。我们使用由验证器检查的Aaronson-Gottesman思维链(AG-CoT)轨迹来监督模型,并在验证器接受的模型生成上继续训练。在两个独立训练的模型系列(3B和7B)中,AG-CoT监督将贪婪解码的状态等价准确率相对于仅电路基线提高了四到六倍,而验证器过滤的继续训练在此基础上又增加了持续一致的增益。一项补充的32B研究表明,监督模型实现了近乎完美的语法和Clifford有效性,而最强的直接模型在每个目标上达到6.14%的状态等价率,在验证器引导的多候选选择下升至超过10%。这些结果表明,算法轨迹监督在两个模型系列中都带来了巨大的、统计显著的增益,并且验证器过滤的继续训练进一步增加了重复增益。Clifford有效性与状态等价之间的持续差距证实了精确验证的必要性:一个电路可以在语法和物理上有效,但制备出错误的量子态。
cs.LG / 244 / 2609.33193

Apparent Compression, Real Stability: The Intrinsic Dimension of Learning a Quantum Wavefunction

表观压缩,真实稳定:学习量子波函数的内在维度
Wei, Lu, Wang, Yufeng, Cao, Chenfeng, Ling, Haibin
Abstract
How many directions in weight space does training need? The intrinsic dimension answers this with the smallest number of random directions in which training still reaches a target accuracy, and small values have motivated parameter-efficient methods such as LoRA. We measure it for variational Monte Carlo (VMC), which trains a neural network to represent the ground state of a quantum many-body system. VMC is a demanding test, because the network generates its own training samples and every gradient is noisy, and a revealing one, because the exact answer is known and every run can be scored. We train only a small latent vector that a frozen random map turns into the network's weights, with no change to the standard natural-gradient optimizer. We find that a small dimension can be misleading, while the stability it brings is real. On a magnet with a hard sign pattern, a network that cannot represent signs reaches its best energy in 8 of 28,642 directions, but only because no such network can go lower; once signs are learnable, neither the signs nor the magnitudes are cheap. The dimension rises across a quantum phase transition, so it tracks how difficult a state is at far less compute than fitting a scaling law, yet it never falls below a floor set by the random subspace itself, even where the ground state is nearly trivial. Training in the subspace, in contrast, never diverged in our experiments, whereas full-parameter training with the same settings did, and a control with matched solvers attributes the difference to the reduced dimension.
Chinese Translation
训练在权重空间中需要多少个方向?内在维度用训练仍能达到目标精度的最少随机方向数来回答这个问题,并且较小的取值推动了诸如 LoRA 之类的参数高效方法。我们针对变分蒙特卡洛(VMC)测量了该维度,VMC 训练一个神经网络来表示量子多体系统的基态。VMC 是一项要求苛刻的测试,因为网络生成自己的训练样本,且每个梯度都有噪声;它也是一项具有揭示性的测试,因为精确答案是已知的,并且每次运行都可以评分。我们仅训练一个小的潜在向量,由一个冻结的随机映射将其转换为网络的权重,而不改变标准的自然梯度优化器。我们发现,小的维度可能具有误导性,而它带来的稳定性是真实的。在一个具有困难符号模式的磁体上,一个无法表示符号的网络在 28,642 个方向中的 8 个方向上达到了其最佳能量,但这只是因为这类网络无法达到更低;一旦符号可学习,符号和幅度都不廉价。该维度在跨越量子相变时上升,因此它能在远少于拟合标度律的计算量下追踪一个状态的困难程度,然而它永远不会低于随机子空间本身设定的下限,即使在基态几乎平凡的地方也是如此。相比之下,在子空间中训练在我们的实验中从未发散,而相同设置下的全参数训练却发散了,并且一个使用匹配求解器的对照将差异归因于降低的维度。
cs.LG / 245 / 2609.33194

Orthogonal Witness Control for Muon Optimization via Sigmoid Spectral Reshaping

通过Sigmoid谱重塑的Muon优化正交见证控制
Van, Dat Phi, Minh, Ngo Vu, Nguyen, Tuc, Nguyen, Thin, Dinh, Ngoc-Thanh, Le, Trung
Abstract
Matrix-valued optimizers such as Muon exploit the spectral structure of neural network updates through Newton--Schulz orthogonalization, but their near-flattening of the singular spectrum discards relative magnitude information across gradient modes. We introduce \emph{Soren} (\textbf{S}pectral \textbf{O}rthogonal \textbf{Re}shapi\textbf{n}g), a matrix-valued optimizer that preserves the singular subspaces of the gradient while applying a bounded, monotone sigmoid transformation to its singular values. This smoothly compresses dominant modes without fully flattening the spectrum. We interpret Soren as a positive-definite preconditioned gradient method and establish convergence guarantees under relative smoothness and metric Polyak--{\L}ojasiewicz geometry. To avoid explicit singular value decomposition, we further develop a finite-depth Soft Newton--Schulz (SNS) polynomial realization of the sigmoid spectral map and characterize how its spectral approximation affects the induced convergence geometry. Experiments across LLM pre-training, supervised fine-tuning, and direct preference optimization demonstrate the effectiveness and robustness of Soren against established optimizers.
Chinese Translation
矩阵值优化器(如Muon)通过Newton--Schulz正交化利用神经网络更新的谱结构,但它们对奇异谱的近乎展平丢弃了跨梯度模式的相对幅度信息。我们引入了Soren(Spectral Orthogonal Reshaping),一种矩阵值优化器,它保留梯度的奇异子空间,同时对其奇异值应用有界单调的Sigmoid变换。这平滑地压缩了主导模式,而不会完全展平谱。我们将Soren解释为正定预条件梯度方法,并在相对光滑性和度量Polyak-Łojasiewicz几何下建立了收敛保证。为了避免显式奇异值分解,我们进一步开发了Sigmoid谱映射的有限深度Soft Newton--Schulz (SNS)多项式实现,并刻画了其谱近似如何影响诱导的收敛几何。在LLM预训练、监督微调和直接偏好优化上的实验证明了Soren相对于已有优化器的有效性和鲁棒性。
cs.LG / 246 / 2609.33200

Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning

教会自己看向哪里:面向推理的在线策略注意力自蒸馏
Arib, Safaeid Hossain, Akter, Rabeya, Swapnil, Ismam Nur, Sayeedi, Md. Faiyaz Abdullah, Islam, Md Mofijul, Mohiuddin, Tasnim
Abstract
On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53x faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.
Chinese Translation
在线策略自蒸馏利用拥有已验证解答的特权教师提供的稠密 token 分布指导,在推理模型自身的轨迹上训练这些模型。这种监督迁移了教师的预测内容,却没有直接迁移其在先前上下文中的关注位置。我们提出在线策略注意力自蒸馏(On-Policy Attention Self-Distillation, OPASD),它通过以解答为条件的注意力蒸馏来补充 token 级监督。由于特权教师可以关注学生无法获得的已验证解答 token,OPASD 将教师注意力投影到学生可见的位置,并在对齐前对所得分布重新归一化。在三种模型规模和四个竞赛级数学基准上,OPASD 始终优于仅 token 的 OPSD,将平均准确率提高 4.98 至 8.40 个百分点。OPASD 还避免了仅 token 蒸馏中观察到的回复长度膨胀和性能退化,将生成的 rollout token 减少 73.9%,估计模型计算量减少 72.6%,同时训练速度提升 1.53 倍。这些结果表明,以解答为条件的注意力提供了一种互补的监督信号,使在线策略自蒸馏更准确、更稳定且计算效率更高。
cs.LG / 247 / 2609.33205

SMORE: Stability-Promoting Mesh-Agnostic Model Reduction for Time-Dependent PDEs

SMORE:面向含时偏微分方程的稳定性促进型网格无关模型降阶
Li, Yangyuan, Li, Weichao, Pan, Shaowu
Abstract
High-fidelity simulations of time-dependent partial differential equations (PDEs) are computationally expensive, motivating data-driven reduced-order surrogates for many-query tasks such as uncertainty quantification, design optimization, data assimilation, and optimal control. However, existing surrogate models often exhibit poor temporal stability, which can lead to unstable rollouts and exploding gradients during backpropagation, especially in multistep long-horizon forecasting. To address this, we propose SMORE, a mesh-agnostic framework for model order reduction of time-dependent PDEs. Its latent dynamics are trained with Lyapunov-guided stability regularization, which promotes stable long-horizon rollouts. We provide theoretical guarantees under the stated structural assumptions. Beyond forecasting PDE evolution, the learned latent dynamics, which are interpretable and linear or linear-quadratic, could bring benefits for downstream tasks such as data assimilation and optimal control. Moreover, our framework is capable of predicting continuous PDE solution fields from sparse measurements of the initial condition. We evaluate SMORE on a range of problems, including wave propagation, the Navier-Stokes equations, and the shallow water equations. Our results show that it improves long-horizon rollout generalization and empirical robustness, and achieves competitive accuracy at comparable parameter budgets relative to competitive baselines including DINo, FNO, CNO, and Transolver.
Chinese Translation
高保真模拟含时偏微分方程(PDEs)的计算成本高昂,这促使人们为不确定性量化、设计优化、数据同化和最优控制等多次查询任务开发数据驱动的降阶代理模型。然而,现有代理模型往往时间稳定性较差,这可能导致推演不稳定以及反向传播过程中的梯度爆炸,尤其是在多步长时程预测中。为了解决这一问题,我们提出 SMORE,一个用于含时偏微分方程模型降阶的网格无关框架。其潜动力学通过 Lyapunov 引导的稳定性正则化进行训练,从而促进稳定的长时程推演。我们在所述结构假设下提供了理论保证。除了预测 PDE 演化之外,所学到的潜动力学具有可解释性且为线性或线性二次型,可为数据同化和最优控制等下游任务带来益处。此外,我们的框架能够从初始条件的稀疏测量中预测连续的 PDE 解场。我们在包括波动传播、Navier-Stokes 方程和浅水方程在内的一系列问题上评估 SMORE。结果表明,与 DINo、FNO、CNO 和 Transolver 等有竞争力的基线相比,SMORE 在可比参数预算下提升了长时程推演泛化能力和经验鲁棒性,并达到有竞争力的精度。
cs.LG / 248 / 2609.33220

When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning

模型何时承认自己错了?失败披露在强化学习下不稳定
Feng, Steven Y., Goodman, Noah D., Frank, Michael C., Hubinger, Evan, Bogdan, Paul C., Lampinen, Andrew
Abstract
Outcome-based reinforcement learning can produce models with similar task performance but very different ways of communicating about their mistakes. We study failure disclosure: whether a model admits that an attempted solution failed rather than staying silent or presenting it as successful. Across repeated outcome-only GRPO training runs, failure disclosure varies far more than task accuracy. The pattern extends to a second reasoning task and stabilized PPO, persists at 7B, and also appears in an instruction-conditioned 32B setting. We also find that small floating-point and sampling differences during training can redirect reporting behavior even when the task objective and earlier training history are held fixed. Additional tests show that failure disclosure is not a single decision: Checking the answer, entering a report, and completing the admission can separate, and the weak point depends on the task and response format. Further, experiments with neutral controls show more broadly that behaviors left weakly constrained by training are especially likely to vary across runs, of which failure disclosure is an example. We can reduce variability in failure disclosure by discouraging the model from drifting from its starting policy on failed, well-formed responses. This makes reporting substantially more consistent, though its effect on task performance depends on the setting. Stable task accuracy therefore does not guarantee stable safety-relevant behavior: Researchers should measure these behaviors directly across runs and design training methods that keep them reliable.
Chinese Translation
基于结果的强化学习可以产生任务性能相似但沟通错误方式截然不同的模型。我们研究失败披露:模型是否承认尝试的解决方案失败,而不是保持沉默或将其呈现为成功。在多次仅基于结果的GRPO训练运行中,失败披露的变化远大于任务准确性。该模式扩展到第二个推理任务和稳定的PPO,在7B规模上持续存在,并且也出现在指令条件下的32B设置中。我们还发现,即使在任务目标和早期训练历史保持不变的情况下,训练期间微小的浮点数和采样差异也可以改变报告行为。额外的测试表明,失败披露不是一个单一决策:检查答案、进入报告和完成承认可以分离,而薄弱点取决于任务和响应格式。此外,使用中性对照的实验更广泛地表明,训练约束较弱的行为特别容易在不同运行中发生变化,失败披露就是一个例子。我们可以通过阻止模型在失败的、格式良好的响应上偏离其初始策略来减少失败披露的变异性。这使得报告更加一致,尽管其对任务性能的影响取决于设置。因此,稳定的任务准确性并不能保证稳定的安全相关行为:研究人员应直接在不同运行中测量这些行为,并设计保持其可靠性的训练方法。
cs.LG / 249 / 2609.33221

RMB: Reward Model Boosting Mitigates Reward Hacking

RMB:奖励模型提升缓解奖励黑客
Fan, Jiabin, Ye, Dezhi, Hao, Yongchang, Mou, Lili
Abstract
Reinforcement Learning from Human Feedback (RLHF) is a powerful technique for aligning large language models (LLMs) with human preference. However, it often suffers from the reward hacking issue, where policy optimization improves the proxy reward model while actually degrading performance with respect to the true human preference, due to the imperfection of the proxy. To address this, we propose Reward Model Boosting (RMB), a novel approach that enhances the robustness and reliability of the reward signal for RLHF. RMB first trains a set of reward models with a diversity-promoting regularizer. This encourages each model to learn complementary aspects of the reward landscape. Then, RMB learns a lightweight aggregator in the principle of boosting to aggregate the outputs of the diverse reward models into a more accurate and robust reward signal. Our extensive experiments demonstrate that RMB significantly improves reward accuracy on both in-distribution and out-of-distribution datasets, substantially mitigating the reward hacking issue and ultimately improving RLHF performance.
Chinese Translation
基于人类反馈的强化学习(RLHF)是一种强大的技术,用于使大型语言模型(LLMs)与人类偏好对齐。然而,它经常遭受奖励黑客问题,即由于代理的不完美,策略优化提高了代理奖励模型,但实际上降低了相对于真实人类偏好的性能。为了解决这个问题,我们提出了奖励模型提升(RMB),这是一种新颖的方法,它增强了 RLHF 奖励信号的鲁棒性和可靠性。RMB 首先使用促进多样性的正则化器训练一组奖励模型。这鼓励每个模型学习奖励景观的互补方面。然后,RMB 基于提升原理学习一个轻量级聚合器,将不同奖励模型的输出聚合为更准确、更鲁棒的奖励信号。我们的大量实验表明,RMB 在分布内和分布外数据集上都显著提高了奖励准确性,大幅缓解了奖励黑客问题,并最终提高了 RLHF 性能。
cs.LG / 250 / 2609.33232

MTLiquid: Enabling Efficient Multi-Task Learning using Liquid Neural Networks for Lightweight Healthcare Monitoring Systems

MTLiquid:利用液态神经网络实现轻量级医疗监护系统的高效多任务学习
Putra, Rachmad Vidya Wicaksana, Rauf, Fahad Abdul, Shafique, Muhammad
Abstract
Continuous-time sensing and monitoring with timely and accurate decision-making are critical for many real-world applications. In healthcare monitoring systems, physiological signals are often available or sampled at irregular time intervals, hence requiring continuous-time processing to provide accurate prediction. Moreover, such systems often need to solve multiple detection/prediction tasks to provide a comprehensive patient review from different physiological aspects for more accurate decision-making. To solve this, continuous-time neural networks (CTNNs) can be employed. However, state-of-the-art works typically solve only one task at each network, thereby limiting their efficiency gains. To address this limitation, we propose MTLiquid, a novel methodology to enable efficient multi-task learning in continuous-time processing for healthcare monitoring systems through effective network design and training strategy. MTLiquid employs: (1) multiple input and output heads to accommodate different tasks, while sharing the same backbone network across tasks; as well as (2) an effective training strategy that leverages a loss-weighting technique to balance learning updates across different tasks and a proportional data presentation technique to address imbalanced dataset sizes. Experimental results for mortality prediction (P12) and sepsis early detection (P19) tasks for ICU patients show that, MTLiquid achieves strong performance (AUROC: 0.84 for P12 and 0.94 for P19) comparable to the state-of-the-art single-task learning in both continuous-time networks (AUROC: 0.84 for P12 and 0.95 for P19) and discrete-time networks (AUROC: 0.79-0.82 for P12 and 0.92-0.94 for P19), while incurring significantly smaller memory cost by 44%-94%. These results highlight the potential of our MTLiquid methodology to enable lightweight continuous-time healthcare monitoring systems for better decision-making.
Chinese Translation
连续时间传感与监测以及及时准确的决策对许多实际应用至关重要。在医疗监护系统中,生理信号通常以不规则时间间隔可用或采样,因此需要连续时间处理以提供准确预测。此外,此类系统通常需要解决多个检测/预测任务,以从不同生理方面提供全面的患者评估,从而实现更准确的决策。为了解决这个问题,可以采用连续时间神经网络(CTNNs)。然而,最先进的工作通常在每个网络中只解决一个任务,从而限制了其效率提升。为了解决这一局限,我们提出了MTLiquid,一种新颖的方法,通过有效的网络设计和训练策略,在医疗监护系统的连续时间处理中实现高效的多任务学习。MTLiquid采用:(1)多输入和多输出头以适应不同任务,同时在任务间共享相同的骨干网络;以及(2)有效的训练策略,利用损失加权技术平衡不同任务间的学习更新,并采用比例数据呈现技术解决数据集大小不平衡问题。针对ICU患者的死亡率预测(P12)和脓毒症早期检测(P19)任务的实验结果表明,MTLiquid取得了强劲的性能(P12的AUROC为0.84,P19的AUROC为0.94),与连续时间网络(P12的AUROC为0.84,P19的AUROC为0.95)和离散时间网络(P12的AUROC为0.79-0.82,P19的AUROC为0.92-0.94)中最先进的单任务学习相当,同时内存成本显著降低44%-94%。这些结果凸显了我们的MTLiquid方法在实现轻量级连续时间医疗监护系统以改善决策方面的潜力。
cs.LG / 251 / 2609.33240

The Price of Locality: Why Forward-Forward Underperforms Backpropagation?

局部性的代价:为什么前向-前向算法表现不如反向传播?
Wu, Zhaoxian, Liu, Haichuan, Chen, Tianyi
Abstract
The Forward-Forward Algorithm (FFA) replaces backpropagation (BP) with layer-wise local contrastive objectives, eliminating the backward pass and the need to retain intermediate activations, yet suffers a persistent performance gap with BP that worsens with depth. This paper diagnoses two structural failure modes: an optimization floor arising from concurrent local updates; and a geometric collapse of layer representations driven by the local update mechanism. On the optimization side, we prove that the FFA loss satisfies the Polyak--Lojasiewicz inequality at each layer; however, simultaneous layer updates induce inter-layer representation-distribution drift, so each layer optimizes against a moving input distribution and incurs an error floor. On the representational side, the pairwise similarity kernel of layer representations contracts exponentially toward rank one as depth increases, collapsing the diversity of per-layer error signals. This collapse bounds FFA's effective learning capacity, which measures the diversity of gradient information across layers, independently of depth, whereas BP's chain-rule signal preserves per-layer diversity, yielding a capacity that scales with depth.
Chinese Translation
前向-前向算法(FFA)用逐层局部对比目标取代反向传播(BP),消除了反向传递以及保留中间激活的需要,但其与BP之间存在持续的性能差距,且该差距随深度增加而加剧。本文诊断出两种结构性失效模式:一种是由并发的局部更新引起的优化下界;另一种是由局部更新机制驱动的层表示几何塌缩。在优化方面,我们证明FFA损失在每一层都满足Polyak-Łojasiewicz不等式;然而,同时进行的层更新会导致层间表示分布漂移,因此每一层都针对移动的输入分布进行优化,并产生误差下界。在表示方面,随着深度增加,层表示的成对相似性核以指数方式收缩至秩一,使逐层误差信号的多样性崩溃。这种崩溃限制了FFA的有效学习容量,该容量衡量跨层梯度信息的多样性,且与深度无关;而BP的链式法则信号保持了逐层多样性,产生的容量随深度扩展。
cs.LG / 252 / 2609.33246

Minimax-Optimality of Posterior Sampling for Reinforcement Learning

强化学习中后验采样的极小极大最优性
Goo, Taewon, Hong, Kihyuk
Abstract
Posterior sampling for reinforcement learning (PSRL) is one of the simplest and most effective exploration methods, but a basic question has remained open: does unmodified PSRL achieve minimax regret without structural assumptions on the prior? We answer yes. Exact vanilla PSRL is minimax optimal in leading-order Bayesian regret under arbitrary correlated priors. The difficulty is that a posterior-sampled transition model is coupled with its own continuation value. We overcome this with a common empirical transition reference that isolates the resulting value mismatch and a Bellman-based variance argument that controls it without an extra leading-order state-space factor. For finite-horizon, time-inhomogeneous tabular MDPs with unknown stochastic rewards, this yields the minimax $\widetilde{O}(\sqrt{SAH^3K})$ regret rate under arbitrary joint priors over rewards and transitions. The same proof principle gives the minimax $\widetilde{O}(d\sqrt{H^3K})$ rate for linear-mixture MDPs under arbitrary joint parameter priors.
Chinese Translation
强化学习中的后验采样(PSRL)是最简单且最有效的探索方法之一,但一个基本问题仍未解决:未经修改的PSRL是否能在不对先验做结构假设的情况下达到极小极大遗憾?我们给出的答案是肯定的。在任意相关先验下,精确的原始PSRL在主导阶贝叶斯遗憾上是极小极大最优的。困难在于,后验采样的转移模型与其自身的延续价值相耦合。我们通过一个通用的经验转移参考来隔离由此产生的价值不匹配,并利用基于Bellman的方差论证来控制它,而无需额外的主导阶状态空间因子。对于具有未知随机奖励的有限时域、时间非齐次表格MDP,这在奖励和转移的任意联合先验下产生了极小极大 \widetilde{O}(\sqrt{SAH^3K}) 遗憾率。相同的证明原理在任意联合参数先验下,为线性混合MDP给出了极小极大 \widetilde{O}(d\sqrt{H^3K}) 遗憾率。
cs.LG / 253 / 2609.33248

Feedback-Robust AI for Patient Knowledge Graphs

面向患者知识图谱的反馈鲁棒 AI
Syed, Mohammed Sameer
Abstract
Patient knowledge graphs from bedside monitoring should type their relations and state whether the data support their signs. In anesthesia and intensive care, clinicians titrate drugs and ventilation in response to the physiology, so temporal relations mix the patient's response with the clinician's policy. We introduce ClosedLoopBench: 29 relations with signs fixed by physics, pharmacology or clinical practice, on 3,442 VitalDB surgical cases (12,653 h) with negative-control action streams. When each patient's actions are replaced by another patient's, six of 12 estimators declare on average 11-18 of their 19-29 distinct relation estimates significant without calibration, and after calibration cross-correlation and Granger tests still assign ventilator rate -> end-tidal CO2 the sign of the clinician's policy. We propose feedback-robust patient graphs that combine concept nodes with evidence pointers, typed relations admitted against negative controls, and beat-level couplings. On VitalDB under null streams, our graphs contain 0.06-0.10 false concept-level relation instances per graph, versus 10-12 for correlational construction. Patient-specific estimates of 11 slow drug and ventilator responses predict later data no better than population estimates, whereas the pulse-arrival-time-systolic-pressure slope is negative in 94.3% of 2,884 cases and patient-specific (early-late correlation 0.67 [0.63, 0.70]).
Chinese Translation
来自床边监测的患者知识图谱应对其关系进行类型标注,并说明数据是否支持其符号。在麻醉和重症监护中,临床医生根据生理状况调整药物和通气,因此时间关系混合了患者的反应与临床医生的策略。我们引入 ClosedLoopBench:包含 29 种关系,其符号由物理学、药理学或临床实践固定,基于 3,442 例 VitalDB 手术病例(12,653 小时)并带有阴性对照动作流。当每个患者的动作被另一名患者的动作替换时,12 个估计器中有 6 个在未校准的情况下平均宣称其 19-29 个不同关系估计中的 11-18 个显著;在校准后,互相关和 Granger 检验仍然将呼吸机频率 -> 呼气末 CO2 的符号赋予临床医生的策略。我们提出反馈鲁棒的患者图,它将概念节点与证据指针相结合,关系经过类型标注并针对阴性对照进行准入,以及逐拍耦合。在 VitalDB 的零流下,我们的图每个图包含 0.06-0.10 个虚假的概念级关系实例,而基于相关性的构建则为 10-12 个。对 11 种缓慢药物和呼吸机反应的患者特异性估计对后续数据的预测并不优于群体估计,而脉搏到达时间-收缩压斜率在 2,884 例中有 94.3% 为负,且具有患者特异性(早期-晚期相关性 0.67 [0.63, 0.70])。
cs.LG / 254 / 2609.33254

BERT4DTI : BERT-based Model for Predicting Drug-Protein Interactions

BERT4DTI:基于BERT的药物-蛋白质相互作用预测模型
Hamitouch, Thanina, Henni, Khadidja, Arie, Abdelkrim, Haichour, Amina Selma, Mezghani, Neila, Abou-Abbas, Lina
Abstract
Understanding how drugs interact with protein targets is fundamental to drug discovery, drug repurposing and the early identification of promising therapeutic candidates before costly experimental testing. Sequence-based DTI models face three practical limitations: labelled interactions are scarce and unevenly distributed, large pretrained chemical and protein encoders are expensive to fine-tune end-to-end, and independently encoded sequences do not capture pair-specific dependencies. We present BERT4DTI, which encodes SMILES strings with ChemBERTa and amino-acid sequences with ProtBERT, applies bidirectional mutual attention between token-level representations, and classifies the resulting interaction features using convolutional layers and a multilayer perceptron. To reduce trainable size, ProtBERT is truncated to 18 retained layers and only the last two layers of each encoder are fine-tuned. On BIOSNAP, DAVIS and BindingDB, BERT4DTI is competitive, achieving the best ROC-AUC and PR-AUC on BIOSNAP and the highest sensitivity on all three benchmarks. An ablation on DAVIS shows that mutual attention improves PR-AUC and specificity. With 125M trainable parameters compared with 353M for full BERT fine-tuning, BERT4DTI provides a favourable performance-parameter trade-off for sequence-based DTI screening, while leaving runtime profiling, calibration and leakage-audited validation for future work.
Chinese Translation
理解药物与蛋白质靶标如何相互作用,对于药物发现、药物重定位以及在昂贵的实验测试之前早期识别有前景的治疗候选物至关重要。基于序列的DTI模型面临三个实际限制:标记相互作用稀缺且分布不均,大型预训练化学和蛋白质编码器端到端微调成本高昂,且独立编码的序列无法捕获配对特异性依赖关系。我们提出BERT4DTI,它使用ChemBERTa编码SMILES字符串,使用ProtBERT编码氨基酸序列,在token级表示之间应用双向互注意力,并使用卷积层和多层感知机对得到的相互作用特征进行分类。为减少可训练参数量,ProtBERT被截断为保留18层,并且仅微调每个编码器的最后两层。在BIOSNAP、DAVIS和BindingDB上,BERT4DTI具有竞争力,在BIOSNAP上取得了最佳的ROC-AUC和PR-AUC,并在所有三个基准上取得了最高的灵敏度。在DAVIS上的消融实验表明,互注意力提高了PR-AUC和特异性。与完整BERT微调的353M相比,BERT4DTI仅有125M可训练参数,为基于序列的DTI筛选提供了有利的性能-参数权衡,同时将运行时分析、校准和泄漏审计验证留待未来工作。
cs.LG / 255 / 2609.33257

Two Heads Are Better Than One: Aggregating Weaker LLMs for Better Forecasts

两个头脑胜过一个:聚合较弱的LLM以获得更好的预测
Peng, Cheng, Luo, Ruixi, Chen, Zhi, Tang, Wei
Abstract
Large language models (LLMs) are increasingly used to forecast real-world events, but access to the strongest individual forecaster may be costly or otherwise constrained. We study weak-to-strong forecast aggregation: can individually weaker LLM forecasters be aggregated to outperform a stronger forecaster? Using ForecastBench (Karger et al., 2025), we evaluate 70 LLM forecasters across 16 comparison groups, each with more than 1,000 shared subquestions, yielding 1,121 weaker-model pairs. Within each group, we identify the strongest individual by test Brier score and evaluate aggregates composed exclusively of weaker forecasters, with aggregation weights learned on separate training data. We find substantial evidence of weak-to-strong improvement. Learned linear pooling identifies a weaker pair that matches or outperforms the strongest individual in 11 of 16 groups and comes within 5% of its Brier score in all 16 groups. We also find that these improvements do not rely on having a near-best constituent and are generally accompanied by good calibration. Additional analyses show that adding more models does not consistently improve performance, and competitive weaker-model aggregates also remain available under practical constraints.
Chinese Translation
大语言模型(LLM)正越来越多地用于预测现实世界事件,但获取最强单个预测器可能成本高昂或受到其他限制。我们研究弱到强预测聚合:能否将单个较弱的LLM预测器聚合起来以超越更强的预测器?我们使用ForecastBench(Karger等,2025)评估了70个LLM预测器,跨16个比较组,每个组有超过1,000个共享子问题,产生1,121个较弱模型对。在每个组内,我们通过测试Brier分数识别最强的单个预测器,并评估完全由较弱预测器组成的聚合,聚合权重在单独的训练数据上学习。我们发现大量证据表明弱到强改进存在。学习到的线性池化在16个组中的11个中识别出匹配或超过最强单个预测器的较弱对,并且在所有16个组中其Brier分数与最强单个预测器的差距在5%以内。我们还发现,这些改进不依赖于拥有接近最优的组成部分,并且通常伴随着良好的校准。额外分析表明,添加更多模型并不能持续提高性能,并且在现实约束下,具有竞争力的较弱模型聚合仍然可用。
cs.LG / 256 / 2609.33259

GTRL: Grounding Divide-and-Conquer Value Learning with Temporal Differences

GTRL:用时间差分对分治价值学习进行接地
Chowdhury, Abdul Monaf, Chowdhury, MD Sameer Iqbal, Arman, Shifat E, Hasan, Md Mehedi
Abstract
In offline goal-conditioned reinforcement learning (GCRL), divide-and-conquer scales to long horizons by joining two shorter segments at a subgoal. However, under stochastic dynamics, the base case of this rule values the luckiest trajectories through the data. The subgoal must also lie on a shared trajectory, so a state-goal pair that no trajectory connects gets no value update at all. To address both, we present Grounded Transitive RL (GTRL), an offline GCRL value learning algorithm that grounds the divide-and-conquer update with a one-step TD target. Over a single step, TD is correct, as its target averages over the successors and needs no subgoal. GTRL adds this target to the composition rather than replacing it, so every pair receives an update, and the composition still carries the long horizon. GTRL also corrects the bias from hindsight relabeling by reweighting each goal against how reachable it was from other successors. We evaluate our algorithm on nineteen OGBench tasks spanning stochastic, deterministic, and stitching environments, where it achieves the highest average success rate. Code will be released soon.
Chinese Translation
在离线目标条件强化学习(GCRL)中,分治方法通过在一个子目标处连接两个较短的片段来扩展到长时域。然而,在随机动力学下,该规则的基本情况会通过数据评估最幸运的轨迹。子目标还必须位于共享轨迹上,因此没有轨迹连接的状态-目标对根本得不到价值更新。为了解决这两个问题,我们提出了Grounded Transitive RL (GTRL),一种离线GCRL价值学习算法,它用一步TD目标来接地分治更新。在单步上,TD是正确的,因为其目标对后继状态取平均,且不需要子目标。GTRL将此目标添加到组合中而不是替换它,因此每一对都会收到更新,并且组合仍然承载长时域。GTRL还通过根据每个目标从其他后继状态的可达性重新加权来纠正后见之明重标记带来的偏差。我们在十九个OGBench任务上评估了我们的算法,这些任务涵盖随机、确定性和拼接环境,其中它达到了最高的平均成功率。代码即将发布。
cs.LG / 257 / 2609.33273

Towards Identifiable Representations under Misspecified Structure

迈向误设结构下的可识别表示
Li, Yuke, Zheng, Yujia, Chen, Ziyi, Zhang, Kun, Huang, Heng
Abstract
The presence of noise that depends on the latent variables poses a fundamental challenge to identifiability. Existing results rely on conditional independence among the observations given the latent variables. We study a more general \emph{misspecified structure}, where this conditional factorization does not hold, and establish both precise and approximate identifiability guarantees. We characterize structural misspecification as a perturbed factor analysis problem. For precise identifiability, we establish subspace identifiability under spectral separation and controlled perturbation, followed by component-wise identifiability under structural sparsity. When the precise condition is not guaranteed, we derive an approximate subspace-identifiability theorem. Based on these results, we develop an unsupervised variational estimator for recovering latent variables. Experiments demonstrate the effectiveness of the proposed framework.
Chinese Translation
依赖于潜在变量的噪声的存在对可识别性构成了根本性挑战。现有结果依赖于在给定潜在变量时观测之间的条件独立性。我们研究了一种更一般的误设结构(misspecified structure),其中这种条件分解不成立,并建立了精确和近似的可识别性保证。我们将结构误设刻画为一个扰动因子分析问题。对于精确可识别性,我们在谱分离和受控扰动下建立了子空间可识别性,随后在结构稀疏性下建立了逐分量可识别性。当精确条件无法保证时,我们推导了一个近似子空间可识别性定理。基于这些结果,我们提出了一种用于恢复潜在变量的无监督变分估计器。实验证明了所提框架的有效性。
cs.LG / 258 / 2609.33279

Domain Generalization under Sampling Pattern Shifts in Irregular Time Series

不规则时间序列中采样模式偏移下的领域泛化
Kim, Changhun, Lee, Joohyung, Lee, Kwanhyung, Yoon, Donghwee, Chrysos, Grigorios, Yang, Eunho
Abstract
Irregularly sampled multivariate time series (ISMTS) are prevalent in real-world applications, where both observation times and available measurements can vary substantially across domains. While recent models increasingly exploit such sampling information for prediction, its robustness under sampling pattern shifts remains underexplored. We introduce HAR-C, to the best of our knowledge the first controlled benchmark for sampling pattern shifts in ISMTS, and show that sampling shifts alone can substantially degrade performance, induce sampling-specific shortcuts, and remain challenging for existing domain generalization (DG) methods. Motivated by these findings, we propose PRISM, a DG framework that first learns complementary feature-centric and sampling-centric representations without task labels, and subsequently performs robust supervised training across diverse sampling variations to discourage brittle shortcut reliance. Extensive experiments on controlled and real-world ISMTS benchmarks demonstrate that PRISM consistently improves robustness to unseen sampling shifts over existing methods. Our code is available at https://anonymous.4open.science/r/PRISM.
Chinese Translation
不规则采样的多变量时间序列(ISMTS)在现实应用中普遍存在,其中观测时间和可用测量值在不同领域之间都可能存在显著差异。尽管近期模型越来越多地利用此类采样信息进行预测,但其在采样模式偏移下的鲁棒性仍未得到充分探索。我们提出了 HAR-C,据我们所知,这是首个针对 ISMTS 中采样模式偏移的受控基准,并表明仅采样偏移就可能显著降低性能、诱导采样特有的捷径,并且对现有领域泛化(DG)方法仍然具有挑战性。受这些发现启发,我们提出了 PRISM,一个 DG 框架,它首先在无任务标签的情况下学习互补的以特征为中心和以采样为中心的表示,随后在多样的采样变化上进行鲁棒的监督训练,以抑制对脆弱捷径的依赖。在受控和真实世界 ISMTS 基准上的大量实验表明,与现有方法相比,PRISM 持续提升了对未见采样偏移的鲁棒性。我们的代码可在 https://anonymous.4open.science/r/PRISM 获取。
cs.LG / 259 / 2609.33298

Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs

直接隐藏状态对齐:映射和控制LLM中的偏好表达
Zhang, Fansheng, Guo, Shengran, Wang, Zexiao, Yuan, Liang, Chen, Jiyuan, Luo, Ruikun
Abstract
In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. We ask how a specified preference is represented in native model computation, what prevents target-supporting computation from reliably dominating generation, and whether this structure can directly guide control. We introduce Residual Competition Maps (RCMs), which map a behavioral preference onto signed causal effects of native residual computation. Across preference domains, RCMs reveal coexisting target-supporting and target-competing effects, input-dependent component roles, and cases where a single native-component intervention reverses the preference outcome. DPO substantially reorganizes these effects and can weaken opposition without guaranteeing its removal. We then propose Direct Hidden-State Alignment (DHSA), which treats inference-time hidden states rather than base-model weights as the direct adaptation space. RCM-guided Causal Activation State Transition (CAST) implements DHSA through local state interventions at a small number of preference-relevant interfaces while freezing the base model. With only 256-16,384 controller parameters, CAST reaches DPO-competitive operating points across three preference domains, can complement DPO-trained models, and can be enabled or removed at inference time.
Chinese Translation
在许多情况下,后训练无需从头创建目标行为:基础模型已经能够产生该行为,但并不可靠。这将部分偏好对齐从能力获取转向行为表达。我们探究特定偏好如何在原生模型计算中表示,是什么阻止了支持目标的计算可靠地主导生成,以及这种结构能否直接指导控制。我们引入了残差竞争图(RCMs),它将行为偏好映射到原生残差计算的有符号因果效应上。在不同偏好领域,RCMs揭示了共存的支持目标和竞争目标的效应、输入相关的组件角色,以及单一原生组件干预逆转偏好结果的情况。DPO大幅重组这些效应,并能削弱反对作用但不能保证其消除。然后我们提出了直接隐藏状态对齐(DHSA),它将推理时的隐藏状态而非基础模型权重视为直接适应空间。RCM引导的因果激活状态转移(CAST)通过在少量偏好相关接口处进行局部状态干预来实现DHSA,同时冻结基础模型。仅需256-16,384个控制器参数,CAST在三个偏好领域达到与DPO竞争的操作点,可以补充DPO训练的模型,并能在推理时启用或移除。
cs.LG / 260 / 2609.33303

BITS: Rethinking Fair and Comprehensive Evaluation for Irregular Time Series Forecasting

BITS:重新思考不规则时间序列预测的公平与全面评估
Yan, Kangjia, Wang, Linfeng, Shen, Tianen, Qiu, Xiangfei, Zhang, Ruitong, Miao, Hao, Hu, Jilin, Guo, Chenjuan, Yang, Bin, Jensen, Christian S.
Abstract
Despite recent progress in irregular time series forecasting, the field still lacks a unified benchmark for fair and comprehensive evaluation. Existing evaluations are often conducted on a limited set of datasets with inconsistent experimental protocols and predominantly error-based metrics, rendering it difficult to compare and assess methods fairly and comprehensively across diverse settings. To eliminate these limitations and accelerate progress, we propose BITS, a standardized, reproducible, and extensible benchmark for advancing research on irregular time series forecasting. BITS covers eleven datasets from nine domains with diverse irregularity characteristics, and it characterizes the datasets according to their missing rate, missing pattern complexity, sampling irregularity, and skewness. Further, it offers a unified pipeline for data preprocessing, model integration and evaluation, and reporting. It accommodates regular and irregular time series forecasting methods, including time series foundation models, under consistent settings, incorporating both error-based and non-error-based evaluation metrics. Findings include that method performance varies substantially across irregularity characteristics, with no single modeling strategy consistently dominating. We also find that using error-based or non-error-based metrics can yield different model rankings, highlighting the need for multi-dimensional evaluation. The code can be found at https://anonymous.4open.science/r/BITS-8F2E/.
Chinese Translation
尽管不规则时间序列预测近年来取得了进展,该领域仍然缺乏一个用于公平和全面评估的统一基准。现有评估通常是在有限的数据集上进行的,实验协议不一致,且主要基于误差的指标,这使得难以在不同设置下公平和全面地比较和评估方法。为了消除这些限制并加速进展,我们提出了BITS,一个标准化、可复现且可扩展的基准,用于推进不规则时间序列预测的研究。BITS涵盖了来自九个领域的十一个数据集,具有多样的不规则性特征,并根据缺失率、缺失模式复杂度、采样不规则性和偏度来刻画这些数据集。此外,它提供了一个统一的流程,用于数据预处理、模型集成与评估以及报告。它容纳了规则和不规则时间序列预测方法,包括时间序列基础模型,在一致的设置下,结合了基于误差和非基于误差的评估指标。研究发现包括:方法性能在不同不规则性特征下差异很大,没有一种单一的建模策略能够持续占据主导地位。我们还发现,使用基于误差或非基于误差的指标可能产生不同的模型排名,凸显了多维评估的必要性。代码可在 https://anonymous.4open.science/r/BITS-8F2E/ 找到。
cs.LG / 261 / 2609.33314

ZeroGAR: Benchmarking the Adversarial Robustness of Zero-Shot Graph Models

ZeroGAR:零样本图模型对抗鲁棒性基准测试
Zhang, Zhongjian, Wang, Xiao, Zhang, Busheng, Yan, Bo, Yu, Xingtong, Gao, Yue, Li, Jia, Shi, Chuan
Abstract
Zero-shot graph models (ZGMs), which learn transferable knowledge from source graphs and directly apply to unseen target graphs without any adaptation, have achieved promising performance and attracted considerable attention. Despite their proliferation, existing ZGMs are predominantly evaluated on clean graphs, while existing graph robustness benchmarks mainly focus on supervised settings, leaving a fundamental question largely unexplored: How robust are ZGMs when their unseen target graphs are exposed to adversarial manipulation? In this paper, we answer this question by proposing ZeroGAR, the first systematic benchmark for evaluating the adversarial robustness of ZGMs. ZeroGAR evaluates 13 representative ZGMs from 3 different paradigms on 8 graph datasets across 4 domains, covering both in-domain and cross-domain transfer under structural, textual, and node injection attacks with multiple perturbation budgets. It further investigates whether existing graph defenses remain effective in the zero-shot setting. Extensive experiments reveal that strong clean zero-shot performance does not guarantee adversarial robustness, with three key findings: (1) Vulnerability patterns are related to model prediction mechanisms: GNN-based methods are particularly vulnerable to structural and node injection attacks, whereas LLM-based methods are more vulnerable to textual attacks; (2) Stronger LLM backbones introduce a structure-text robustness trade-off; (3) Existing graph defense methods do not consistently improve zero-shot robustness and may compromise clean performance. We hope that ZeroGAR will facilitate rapid, equitable evaluation and inspire further innovative research in ZGM security.
Chinese Translation
零样本图模型(ZGMs)从源图学习可迁移知识,并直接应用于未见的目标图而无需任何适应,已取得令人瞩目的性能并引起了广泛关注。尽管其大量涌现,现有ZGMs主要在干净图上进行评估,而现有的图鲁棒性基准主要关注监督设置,导致一个基本问题在很大程度上未被探索:当未见的目标图暴露于对抗性操纵时,ZGMs的鲁棒性如何?在本文中,我们通过提出ZeroGAR来回答这个问题,这是第一个用于评估ZGMs对抗鲁棒性的系统基准。ZeroGAR在来自4个领域的8个图数据集上评估了来自3种不同范式的13个代表性ZGM,涵盖了在多种扰动预算下,面对结构、文本和节点注入攻击的域内和跨域迁移。它进一步研究了现有图防御在零样本设置中是否仍然有效。大量实验表明,强大的干净零样本性能并不能保证对抗鲁棒性,并有三个关键发现:(1)脆弱性模式与模型预测机制相关:基于GNN的方法特别容易受到结构和节点注入攻击,而基于LLM的方法更容易受到文本攻击;(2)更强的LLM主干引入了结构-文本鲁棒性权衡;(3)现有的图防御方法不能持续提高零样本鲁棒性,并且可能损害干净性能。我们希望ZeroGAR将促进快速、公平的评估,并激发ZGM安全方面的进一步创新研究。
cs.LG / 262 / 2609.33334

When to Evict, Not What to Keep: Draft-Guided Eviction for Training-Free KV-Cache Compression

何时驱逐,而非保留什么:面向免训练 KV 缓存压缩的草稿引导驱逐
Kang, Haeyong, Yoo, Chang D.
Abstract
Training-free KV-cache compression methods such as SnapKV, H2O, and PyramidKV evict tokens at the end of prefill, aiming to preserve the attention mass that future queries are expected to use -optimizing what to keep. We show that this objective fails in two distinct ways. (1) Compensation: restoring the evicted attention mass can recover the attention-level target without recovering task quality. (2) Selection: covering more of the true decode-query mass can hurt quality when the recovered mass is fragmented rather than concentrated in coherent spans. These failures share a common cause: eviction occurs before the queries that determine the answer trajectory exist. We propose Draft-Guided Eviction (DGE), which defers eviction until after drafting the first k=2 answer tokens using the full cache - just one decode step beyond prefill. Because the draft is generated from the answer's own prefix, no cache entries are discarded before this trajectory signal becomes available. The per-head cache budget remains unchanged, and DGE can be applied directly to SnapKV, PyramidKV, H2O, and StreamingLLM without modifying their eviction scores. Unlike extra-pass methods, DGE changes when eviction occurs rather than what cache entries are selected. Extensive experiments demonstrate that DGE outperforms prior methods at every evaluated budget on five of six instruct-tuned backbones, achieving 44.2 on LongBench, nearly matching FullKV at 44.3. The timing-only control DGE-W achieves the same score, demonstrating that the gain comes from when eviction occurs rather than what is selected - an effect we term trajectory anchoring.
Chinese Translation
免训练的 KV 缓存压缩方法(如 SnapKV、H2O 和 PyramidKV)在预填充结束时驱逐 token,旨在保留未来查询预期使用的注意力权重总量——即优化保留什么。我们表明这一目标在两种不同情况下失效:(1) 补偿:恢复被驱逐的注意力权重总量可以达成注意力层面的目标,但无法恢复任务质量;(2) 选择:当恢复的权重是碎片化的而非集中在连贯片段时,覆盖更多真实解码查询权重反而会损害质量。这些失败有一个共同原因:驱逐发生在决定答案轨迹的查询存在之前。我们提出草稿引导驱逐(DGE),它将驱逐推迟到使用完整缓存草拟前 k=2 个答案 token 之后——仅比预填充多一个解码步骤。由于草稿是从答案自身的前缀生成的,在这个轨迹信号出现之前不会丢弃任何缓存条目。每头缓存预算保持不变,并且 DGE 可以直接应用于 SnapKV、PyramidKV、H2O 和 StreamingLLM,而无需修改其驱逐分数。与额外遍历方法不同,DGE 改变的是驱逐发生的时机,而非选中的缓存条目。大量实验表明,在六个指令微调主干模型中的五个上,DGE 在每个评估预算下都优于先前方法,在 LongBench 上达到 44.2,几乎与 FullKV 的 44.3 持平。仅控制时机的 DGE-W 取得了相同分数,表明增益来自驱逐发生的时机而非选择了什么——我们将这一效应称为轨迹锚定。
cs.LG / 263 / 2609.33336

Beyond Conservatism: Recoverability-Conditioned Exploration for Model-Based Imitation Learning

超越保守性:面向基于模型模仿学习的可恢复性条件探索
Chen, Xuanlin, Wang, Ziyue, Zhou, Xunlan, Shang, Yuan-yih, Wu, Qiang, Wan, Shenghua
Abstract
Model-based imitation learning (MBIL) improves real-environment interaction efficiency by optimizing policies on imagined rollouts from a learned world model. However, the gap between model-induced and real-environment occupancies makes policy learning sensitive to model error. Conservative MBIL mitigates model exploitation during policy optimization, but when real-environment interactions are collected by the same conservative policy, uncertain regions around the expert distribution remain insufficiently sampled. Generic uncertainty-driven exploration, on the other hand, may allocate interaction to novel but task-irrelevant dynamics. We propose REcoverability-CONditioned Exploration for Model-Based Imitation Learning (RECON). RECON separates conservative policy learning from active data collection by maintaining a main policy for task execution and an explorer for real-environment interaction. The explorer is optimized based on epistemic uncertainty conditioned on recoverability estimated from multi-step main-policy imagination, focusing data collection on unknown states from which the main policy can still return toward expert behavior. Experiments on locomotion, navigation and manipulation show consistent gains in interaction efficiency, imitation performance, and robustness, indicating that RECON directs real-environment interaction toward recovery regions around the expert distribution that are underexplored by prior methods, and thereby learns a world model better suited for imitation.
Chinese Translation
基于模型的模仿学习(MBIL)通过在学习到的世界模型生成的想象轨迹上优化策略,提高了真实环境交互效率。然而,模型诱导的占用分布与真实环境占用分布之间的差距使得策略学习对模型误差敏感。保守的MBIL在策略优化过程中缓解了模型利用,但当真实环境交互由相同的保守策略收集时,专家分布周围的不确定区域仍然采样不足。另一方面,通用的不确定性驱动探索可能将交互分配到新颖但与任务无关的动力学上。我们提出了用于基于模型模仿学习的可恢复性条件探索(RECON)。RECON通过维护一个用于任务执行的主策略和一个用于真实环境交互的探索器,将保守策略学习与主动数据收集分离。探索器基于以可恢复性为条件的认知不确定性进行优化,该可恢复性从多步主策略想象中估计,从而将数据收集聚焦于主策略仍能返回专家行为的未知状态。在运动、导航和操作任务上的实验表明,交互效率、模仿性能和鲁棒性均有一致提升,表明RECON将真实环境交互引导至专家分布周围先前方法未充分探索的恢复区域,从而学习到更适合模仿的世界模型。
cs.LG / 264 / 2609.33337

Safe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement Learning

安全分数匹配:基于Hamilton-Jacobi可达性的扩散策略用于在线安全强化学习
Li, Boyang, Kim, Matthew, Herbert, Sylvia Lee
Abstract
Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Process via primal-dual Lagrangian updates that only enforce safety on average. To address this limitation, hard, state-wise constraints are introduced and often imposed through Hamilton-Jacobi (HJ) reachability. Yet such constraints require solving different objectives in the feasible and infeasible regions: reward maximization in the former, recovery toward the feasible regions in the latter. The resulting target action distributions are inherently multimodal, and this structure poses a fundamental challenge for the Gaussian or deterministic actors used in existing HJ-based safe RL, which often collapse onto suboptimal modes. Diffusion policies provide the expressiveness needed to represent such distributions, and recent work on Q-score matching offers a route to training them for online RL by score regression -- but has been applied only to reward maximization. We propose Safe Score Matching (SSM), an off-policy actor-critic method that adapts Q-score matching to hard-constrained safe RL by gating a two-branch score target with HJ reachability: inside the feasible set, the denoising process degenerates to Q-score matching on actions classified as viable by the HJ critic; outside, a recovery branch biases denoising toward regions with lower worst-case violation. On quadrotor and fixed-wing trajectory-tracking and stabilize-and-avoid benchmarks, SSM attains the best or near-best task performance with low false-safe rates, whereas the primal-dual baseline admits more unsafe behavior and reachability-based baselines tend to be more conservative; on Safety-Gymnasium velocity tasks, SSM attains the lowest cost with competitive reward.
Chinese Translation
在线安全强化学习(RL)旨在最大化奖励的同时满足安全约束。安全RL中一个流行方向是将安全松弛为软期望代价约束,并通过原始-对偶拉格朗日更新求解所得的约束马尔可夫决策过程,但这类方法仅在平均意义上强制安全。为克服这一局限,硬性、逐状态约束被引入,并通常通过Hamilton-Jacobi(HJ)可达性施加。然而这类约束需要在可行域和不可行域中求解不同目标:前者最大化奖励,后者恢复到可行域。由此得到的目标动作分布本质上是多模态的,而这一结构对现有基于HJ的安全RL中使用的高斯或确定性actor构成根本挑战,后者常常坍缩到次优模态。扩散策略具备表示此类分布所需的表达能力,近期关于Q-score matching的工作通过分数回归为在线RL训练扩散策略提供了途径,但仅应用于奖励最大化。我们提出安全分数匹配(Safe Score Matching, SSM),一种离策略actor-critic方法,通过用HJ可达性门控双分支分数目标,将Q-score matching适配到硬约束安全RL:在可行集内,去噪过程退化为对HJ critic判定为可行的动作进行Q-score matching;在可行集外,恢复分支将去噪偏向最坏情况违反更低的区域。在四旋翼和固定翼轨迹跟踪与稳定避障基准上,SSM取得最佳或接近最佳的任务性能,且误安全率低;而原始-对偶基线允许更多不安全行为,基于可达性的基线往往更保守;在Safety-Gymnasium速度任务上,SSM以有竞争力的奖励获得最低代价。
cs.LG / 265 / 2609.33340

RAEGL: Risk-Aware Evidence-Gated Learning for Selective Contextual Routing under Temporal Shift

RAEGL:时间偏移下选择性上下文路由的风险感知证据门控学习
Guo, Yifan
Abstract
Contextual specialization can improve forecasting accuracy, but a correction selected on one historical interval may become unreliable under temporal distribution shift. To address this issue, we propose RAEGL, a Risk-Aware Evidence-Gated Learning framework for selective contextual forecasting. RAEGL retains a validated global predictor by default and activates a contextual residual only when pre-deployment evidence supports its use. The framework separates candidate selection from gate calibration and jointly evaluates randomization significance, practically meaningful gain, and temporal stability. Experiments on real-world audits and controlled panels show how RAEGL can prevent harmful contextual deployment while making conservative opportunity costs explicit. In a reconstructed Our World in Data audit, exact fallback avoids RMSE degradations of 0.0960 and 0.0239 caused by two validation-selected corrections. In a sealed World Development Indicators evaluation, a region-based correction passes the randomization test but is withheld because its gain is only 0.000092, its country-clustered 95% confidence interval crosses zero, and only 0.02% of bootstrap replicates reach the practical threshold. In controlled panels, the stability- and support-aware extension activates in 97.2% of strong, stable-context runs while rejecting all high-drift settings. These results support RAEGL as an auditable, evidence-based mechanism for managing contextual deployment risk and as a conservative alternative to validation-driven contextual selection.
Chinese Translation
上下文专业化可以提高预测准确性,但在一个历史区间上选择的修正可能在时间分布偏移下变得不可靠。为了解决这个问题,我们提出了 RAEGL,一个用于选择性上下文预测的风险感知证据门控学习框架。RAEGL 默认保留一个经过验证的全局预测器,仅当部署前证据支持其使用时才激活上下文残差。该框架将候选选择与门控校准分离,并联合评估随机化显著性、实际有意义的增益和时间稳定性。在真实世界审计和受控面板上的实验表明,RAEGL 能够防止有害的上下文部署,同时使保守的机会成本明确化。在重建的 Our World in Data 审计中,精确回退避免了由两个验证选择的修正引起的 RMSE 退化 0.0960 和 0.0239。在封存的世界发展指标评估中,一个基于区域的修正通过了随机化检验,但被拒绝采用,因为其增益仅为 0.000092,其国家聚类的 95% 置信区间跨越零,并且只有 0.02% 的自助法重复达到实际阈值。在受控面板中,稳定性和支持感知扩展在 97.2% 的强、稳定上下文运行中激活,同时拒绝所有高漂移设置。这些结果支持 RAEGL 作为一种可审计、基于证据的机制来管理上下文部署风险,并作为验证驱动的上下文选择的保守替代方案。
cs.LG / 266 / 2609.33342

CalibHyper: Chance-Corrected Relational Hypergraphs for Few-Shot Molecular Property Prediction

CalibHyper:面向小样本分子性质预测的机遇校正关系超图
Li, Linyu, Jin, Zhi, He, Yuanpeng, Jin, Dongming, Liu, Huanyu, Zhang, Huanyao, Duan, Haoran, Tian, Heng, Luosang, Gadeng, Tashi, Nyima
Abstract
Molecular property prediction is central to drug development and materials discovery, but experiments are costly and labeled data are scarce. Context-aware methods use auxiliary assay labels to support few-shot prediction, and recent work supervises property relations with label agreement. However, label agreement is sensitive to class marginals and does not directly capture dependence between properties. We propose CalibHyper, a chance-corrected relational hypergraph method based on the joint label distribution. CalibHyper subtracts an independence baseline from the ordered four-state label distribution and shrinks the residual according to the number of joint observations. A swap-equivariant relation head estimates these residuals, which choose the auxiliary properties for each molecule and set the sign and weight of their hyperedge messages. On thirteen datasets from five benchmarks, in both 1-shot and 10-shot settings, CalibHyper and its ablation settings achieve ROC-AUC competitive with the strongest reported results.
Chinese Translation
分子性质预测是药物开发和材料发现的核心,但实验成本高昂且标记数据稀缺。上下文感知方法利用辅助测定标签来支持小样本预测,近期工作通过标签一致性来监督性质关系。然而,标签一致性对类别边缘分布敏感,且不能直接捕捉性质之间的依赖关系。我们提出CalibHyper,一种基于联合标签分布的机遇校正关系超图方法。CalibHyper从有序四状态标签分布中减去独立性基线,并根据联合观测的数量对残差进行收缩。一个交换等变关系头估计这些残差,这些残差为每个分子选择辅助性质,并设置其超边消息的符号和权重。在来自五个基准的十三个数据集上,在1-shot和10-shot设置中,CalibHyper及其消融设置达到了与最强报告结果相当的ROC-AUC。
cs.LG / 267 / 2609.33347

MultiEcho: An Experimental Science of Learned Worlds

MultiEcho:习得世界的实验科学
Zhu, Meng, Zhang, Airui
Abstract
World models can be studied as experimental systems with response laws of their own. We introduce MultiEcho, a framework for estimating these laws through controlled counterfactual interventions, delimiting their applicability, and separately testing their physical correspondence. Across nine simulated physical systems and seven frozen model configurations, three-reference estimators predict complete intervention responses and recover intervention parameters. Estimator selection uses discovery data only; frozen fits are evaluated on validation and confirmation contexts. The experiments distinguish response predictability, intervention readability and physical accuracy. Responses can be locally describable yet poorly match physical effects in the same target coordinates. Event-window, visibility and camera interventions reveal conditional applicability, and paired generator configurations show reduced readability under a scene prompt with stronger guidance. Magnitude sweeps expose small image errors alongside large relative effect errors. An exact-reset material experiment separates registered visible-response success from fixed-readout failure on material-dependent futures at matched positions and velocities. Exact finite-scale identities resolve odd and even response errors; first-order remainder bounds specify when refined calibration converges. MultiEcho provides an experimental basis for studying learned-world laws independently of, and in relation to, physical laws.
Chinese Translation
世界模型可以作为具有自身响应规律的实验系统来研究。我们引入了 MultiEcho,一个通过受控反事实干预来估计这些规律、界定其适用性,并单独测试其物理对应性的框架。在九个模拟物理系统和七个冻结模型配置上,三参考估计器预测完整的干预响应并恢复干预参数。估计器选择仅使用发现数据;冻结拟合在验证和确认上下文中进行评估。实验区分了响应可预测性、干预可读性和物理准确性。响应可以在局部描述,但在相同的目标坐标中与物理效应匹配较差。事件窗口、可见性和相机干预揭示了条件适用性,并且配对生成器配置显示在具有更强引导的场景提示下可读性降低。幅度扫描揭示了小的图像误差以及大的相对效应误差。一个精确重置的材料实验将注册的可见响应成功与在匹配位置和速度下对材料依赖未来的固定读出失败区分开来。精确的有限尺度恒等式解决了奇数和偶数响应误差;一阶余项界指定了精细校准何时收敛。MultiEcho 为独立于物理定律并与之相关地研究习得世界定律提供了实验基础。
cs.LG / 268 / 2609.33350

KoopCell: Koopman-Based Generative Model for Learning Single-Cell Dynamics from Distribution Snapshots

KoopCell:基于Koopman的生成模型用于从分布快照中学习单细胞动力学
Lu, Wanfeng, Zhang, Yutong, Zhou, Keyi, Ge, Chenxin, Lin, Wei, Zhu, Qunxi
Abstract
Learning population dynamics from temporally sparse, unpaired distribution snapshots is a fundamental challenge in developmental biology. Recent approaches based on neural differential equations and flow matching can interpolate between observed population snapshots, but may struggle to extrapolate beyond the training horizon and often lack an explicit mechanism for modeling developmental branching. We propose KoopCell, a unified generative framework based on Koopman-Mori-Zwanzig theory that jointly learns representations and predictive linear latent dynamics. Theoretically, using the weak continuity equation, we derive a closed-form least-squares estimator for the Koopman generator from distribution snapshots and establish convergence guarantees under suitable assumptions. To model branching dynamics, we further develop KoopCell-M, which incorporates non-Markovian memory into the latent Koopman dynamics through a Markovian embedding. Experiments on synthetic systems and three scRNA-seq datasets demonstrate the ability of our framework to recover Koopman spectra, model branching through memory, and scale to predicting high-dimensional gene expression distributions, achieving state-of-the-art performance among the evaluated methods.
Chinese Translation
从时间稀疏、非配对的分布快照中学习群体动力学是发育生物学中的一个基本挑战。最近基于神经常微分方程和流匹配的方法可以在观察到的群体快照之间进行插值,但可能难以外推到训练范围之外,并且通常缺乏对发育分支进行建模的显式机制。我们提出了KoopCell,一个基于Koopman-Mori-Zwanzig理论的统一生成框架,联合学习表示和预测性线性潜在动力学。理论上,利用弱连续性方程,我们从分布快照中推导出Koopman生成元的闭式最小二乘估计量,并在适当假设下建立收敛保证。为了对分支动力学建模,我们进一步开发了KoopCell-M,它通过马尔可夫嵌入将非马尔可夫记忆融入潜在Koopman动力学。在合成系统和三个scRNA-seq数据集上的实验证明了我们的框架能够恢复Koopman谱、通过记忆对分支建模,并可扩展到预测高维基因表达分布,在评估的方法中达到了最先进的性能。
cs.LG / 269 / 2609.33352

How Much Imprecision is Enough Imprecision in my Classifier? A Practical Elicitation Procedure

我的分类器中多少不精确性才是足够的不精确性?一种实用的诱导程序
de Souza, Victor F. Lopes, Destercke, Sébastien, Imoussaten, Abdelhak
Abstract
Set-valued classifiers, whether derived from precise probabilities and an adapted cost function, from convex sets with a robust inference mechanism, or from conformal methods, are routine options to obtain more robust, trustworthy predictions. However, there is a lack of operational tools to measure how robust or imprecise a given user is ready to be when receiving predictions, that is how much precision he/she is ready to let go in exchange of more accuracy. This is why we propose, in this paper, practical and operational elicitation procedures to measure the user proneness to set-valued predictions. The effectiveness of the iterative elicitation procedure in converging to the target parameter value is demonstrated on both tabular and image datasets drawn from standard machine learning benchmarks. The results show that the procedure also presents the user with a small number of instances, highlighting the practicality of the approach for real-world applications aimed at identifying the decision maker's optimal behavior when faced with imprecision.
Chinese Translation
集合值分类器,无论是从精确概率和适应的代价函数推导而来,还是从具有鲁棒推理机制的凸集,或从共形方法,都是获得更鲁棒、更可信预测的常规选项。然而,缺乏操作工具来衡量给定用户在接收预测时准备接受多大的鲁棒性或不确定性,即他/她准备放弃多少精确度以换取更高的准确性。这就是为什么我们在本文中提出了实用且可操作的诱导程序,以测量用户对集合值预测的倾向性。迭代诱导程序在收敛到目标参数值方面的有效性在来自标准机器学习基准的表格和图像数据集上得到了证明。结果表明,该程序还向用户展示了少量实例,突出了该方法在现实应用中的实用性,这些应用旨在识别决策者在面临不精确性时的最优行为。
cs.LG / 270 / 2609.33370

The Selection Rule Decides the Winner: A Pre-Registered Audit of Open-Set Graph Anomaly Detection

选择规则决定胜者:开放集图异常检测的预注册审计
Hossain, Farhan Shahriyar, Fuad, Taufikur Rahman, Jahin, Md Abrar, Parvez, Md Rizwan
Abstract
Open-set graph anomaly detection trains on a few labeled anomalies from one class and must also find anomaly classes that were never labeled. Published results share three conventions: the test score is read at the best epoch on the test set, baseline numbers are copied from earlier papers, and most anomalies are minority classes relabeled as anomalous. We ask how much of the reported ranking these conventions decide. We re-run two recent methods, DEMO and NSReg, together with OUTPOST, a small first-order detector built for this study. All three use one protocol with identical seeds and splits on eight graphs (seven for the baselines, which cannot run on ogbn-mag), ten seeds each, and every run is scored under both the best-epoch rule and a deployable validation rule. Before the runs that test them, we registered 40 predictions. Three findings hold. First, the rule changes the leader: under the best-epoch rule, OUTPOST and NSReg each lead three of seven graphs, while under the validation rule, NSReg leads five. Second, the best-epoch bonus depends on how the benchmark was built: 0.045--0.080 AUC-ROC on the three small relabeled-class graphs and 0.002--0.014 on the three real fraud graphs. Third, pseudo-labeling in OUTPOST is worth 0.038--0.065 AUC-ROC on the same three graphs but gives no benefit on any real fraud graph. We also show that a 0.002 tie band for hyperparameter selection lies below the paired standard error on all six graphs tested, even at ten seeds. Twelve of our 40 predictions were falsified, and we report them. We close with a short reporting checklist.
Chinese Translation
开放集图异常检测在一个类别的少量标注异常上训练,还必须找到从未标注过的异常类别。已发表的结果共享三个惯例:测试分数是在测试集的最佳epoch上读取的,基线数字是从早期论文复制的,并且大多数异常是少数类被重新标记为异常。我们问这些惯例在多大程度上决定了所报告的排名。我们重新运行两种近期方法 DEMO 和 NSReg,以及 OUTPOST,一个为本研究构建的小型一阶检测器。三者都使用一个协议,在八个图(基线为七个,因为无法在 ogbn-mag 上运行)上具有相同的种子和划分,每个十个种子,并且每次运行都在最佳epoch规则和可部署的验证规则下评分。在测试它们的运行之前,我们注册了40个预测。有三个发现成立。第一,规则改变了领先者:在最佳epoch规则下,OUTPOST 和 NSReg 各在七个图中的三个领先,而在验证规则下,NSReg 在五个图中领先。第二,最佳epoch的增益取决于基准是如何构建的:在三个小型重新标记类图上为 0.045--0.080 AUC-ROC,在三个真实欺诈图上为 0.002--0.014。第三,OUTPOST 中的伪标签在同样的三个图上价值 0.038--0.065 AUC-ROC,但在任何真实欺诈图上没有益处。我们还表明,用于超参数选择的 0.002 平局带在所有六个测试图上都低于配对标准误差,即使在十个种子时也是如此。我们的40个预测中有12个被证伪,我们报告了它们。我们以一个简短报告清单结束。
cs.LG / 271 / 2609.33376

TNF based Spectral Embedding for Effective Application of Supervised Machine Learning Techniques in Automobile Insurance Fraud Detection

基于TNF的谱嵌入在汽车保险欺诈检测中有效应用监督机器学习技术
Gupta, Rohan Yashraj, Chintalapati, Lalith Srikanth, Mudigonda, Satya Sai, Baruah, Pallav Kumar, Rachakonda, Raghunatha Sarma
Abstract
Fraud detection is an important area of research in the insurance business due to its financial implications. The primary aim of a fraud detection model is to identify fraud and non-fraud cases with high accuracy along with other important metrics such as Sensitivity, Specificity, Precision, F1-score, False Positive Rate, False Discovery Rate, AUC etc. To achieve this, we need to explore a suitable classification model to identify fraud and non-fraud cases. In this work, we have used auto insurance data set and explored classification models such as Decision Tree (DT), Random Forest (RF), XGBoost, LightGBM and Gradient Boosting Machine (GBM). To overcome the problem of data imbalance, we have employed MWMOTE and TGAN techniques. We have used Topological Node Feature(TNF) based spectral embedding for low dimensional data representation along with some popular embedding methods like MDS, Isomaps and t-SNE. After studying all the 65 possible combinations of these models, we have proposed an innovative method for effective automobile insurance fraud detection. For the given dataset, our results show that using a combination of MWMOTE as a data imbalance handling technique (Phase I), TNFSE2 as data embedding (Phase II) and Random Forest as classification (Phase III) provides the best result in comparison to all other combinations. This work also highlights the efficacy of TNF based spectral embedding in automobile insurance dataset
Chinese Translation
欺诈检测因其财务影响而成为保险业务中的一个重要研究领域。欺诈检测模型的主要目标是以高准确率识别欺诈和非欺诈案件,同时兼顾其他重要指标,如敏感度(Sensitivity)、特异性(Specificity)、精确率(Precision)、F1分数(F1-score)、假阳性率(False Positive Rate)、错误发现率(False Discovery Rate)、AUC等。为实现这一目标,我们需要探索合适的分类模型来识别欺诈和非欺诈案件。在这项工作中,我们使用了汽车保险数据集,并探索了分类模型,如决策树(DT)、随机森林(RF)、XGBoost、LightGBM和梯度提升机(GBM)。为了克服数据不平衡问题,我们采用了MWMOTE和TGAN技术。我们使用了基于拓扑节点特征(TNF)的谱嵌入进行低维数据表示,并结合了一些流行的嵌入方法,如MDS、Isomaps和t-SNE。在研究这些模型的所有65种可能组合后,我们提出了一种创新的方法,用于有效的汽车保险欺诈检测。对于给定的数据集,我们的结果表明,使用MWMOTE作为数据不平衡处理技术(第一阶段)、TNFSE2作为数据嵌入(第二阶段)和随机森林作为分类(第三阶段)的组合,与所有其他组合相比提供了最佳结果。这项工作还突出了基于TNF的谱嵌入在汽车保险数据集中的有效性。
cs.LG / 272 / 2609.33377

Optimal Transport Dropout for Structured Predictive Uncertainty

面向结构化预测不确定性的最优传输 Dropout
Lorenzon, Giacomo, Regazzoni, Francesco
Abstract
Deterministic neural networks and neural operators provide point predictions with no intrinsic measure of reliability. Yet, predictive uncertainty may stem from irreducible outcome variability, finite data, or limitations of the chosen model class. Monte Carlo dropout offers a computationally convenient way to construct a predictive distribution through stochastic feature masking, without training multiple independent networks or explicitly inferring a posterior over model parameters. However, its perturbation law is largely prescribed a priori and typically factorised across latent coordinates. We introduce Optimal Transport Dropout (OTD), which instead learns the predictive mapping and the law of its latent perturbations jointly. Starting from a simple independent reference distribution, OTD transports latent perturbations through a learnable flow and propagates them through the predictive neural network, thereby inducing a structured predictive law. Training uses the strictly proper Energy Score, while a kinetic-action term geometrically regularises the transport. Synthetic benchmarks show that OTD captures multimodal predictive distributions, generates meaningful dispersion when the model is misspecified, and exhibits contracting dispersion as more training data or greater model capacity are provided. For a field-valued partial differential equation surrogate, predictive dispersion strongly aligns with the spatial pattern of prediction errors. On this task, compared with Monte Carlo dropout, OTD yields more accurate predictions and better-calibrated, substantially narrower intervals. On real-world regression benchmarks, it further shows competitive accuracy and better probabilistic predictions compared to established baselines. OTD therefore offers a way to learn structured predictive uncertainty without explicit posterior inference or ensembles of independently trained predictors.
Chinese Translation
确定性神经网络和神经算子提供点预测,却没有内在的可靠性度量。然而,预测不确定性可能源于不可约的结果变异性、有限数据或所选模型类别的局限性。蒙特卡洛 Dropout 提供了一种计算上便捷的方法,通过随机特征掩蔽来构建预测分布,而无需训练多个独立网络或显式推断模型参数的后验。然而,其扰动规律在很大程度上是先验规定的,并且通常在潜在坐标上因子化分解。我们提出最优传输 Dropout(OTD),它转而联合学习预测映射及其潜在扰动的规律。从简单的独立参考分布出发,OTD 通过可学习流传输潜在扰动,并将其传播通过预测神经网络,从而诱导出结构化的预测规律。训练使用严格适当的能量评分(Energy Score),同时一个动能作用项从几何上正则化传输。合成基准表明,OTD 能捕获多模态预测分布,在模型设定错误时产生有意义的离散度,并且随着提供更多训练数据或更大模型容量而呈现收缩的离散度。对于场值偏微分方程代理模型,预测离散度与预测误差的空间模式高度一致。在此任务上,与蒙特卡洛 Dropout 相比,OTD 产生更准确的预测以及校准更好、显著更窄的区间。在真实世界回归基准上,与已有基线相比,它进一步展现出有竞争力的精度和更好的概率预测。因此,OTD 提供了一种学习结构化预测不确定性的方法,无需显式后验推断或独立训练预测器的集成。
cs.LG / 273 / 2609.33387

From Grey-Box to Green-Box: When can Physics-Informed Machine Learning Reduce Carbon Footprints in Structural Health Monitoring?

从灰箱到绿箱:物理信息机器学习何时能够降低结构健康监测中的碳足迹?
Bradley, Daisy R., Hinchliffe, Nathan A., Pitchforth, Daniel J., Jones, Matthew R., Cross, Elizabeth J.
Abstract
Machine learning plays an increasingly vital role in engineering, but the corresponding increase in compute time is not without environmental cost. Physics-informed machine learning or "grey-box" models have been developed to overcome some of the limitations of traditional black-box learners, utilising the physical insight that an engineer would have about the structure they are modelling and have shown promising results in the structural engineering field among many others. This work explores whether an additional advantage could be a reduced environmental impact, considering the relationship between training data quantity and training time, linking this duration to carbon emissions from computing. In a structural health monitoring context, four physics-informed machine learning approaches - spanning Gaussian processes and neural networks - are evaluated: residual modelling, input augmentation, hybrid modelling, and constrained learning. The emissions for training each of the models to reach a given error threshold is compared, and in most examples, shown to be lower for the physics-informed models (with input augmented models being an exception). This reduction in training emissions further compounds the environmental savings achieved by collecting and storing less data. Although promising results, we cannot expect a silver bullet and the case studies demonstrate that a trade-off is needed between the increased complexity that comes from introducing physics into a machine learner, against the gain from reduced training data requirements.
Chinese Translation
机器学习在工程中发挥着越来越重要的作用,但相应增加的计算时间并非没有环境代价。物理信息机器学习或“灰箱”模型已被开发出来,以克服传统黑箱学习器的一些局限性,利用工程师对其所建模结构拥有的物理洞见,并在结构工程领域等许多领域显示出有前景的结果。这项工作探讨了一个额外优势是否可能是降低环境影响,考虑训练数据量与训练时间之间的关系,并将该持续时间与计算产生的碳排放联系起来。在结构健康监测背景下,评估了四种物理信息机器学习方法——涵盖高斯过程和神经网络:残差建模、输入增广、混合建模和约束学习。比较了将每个模型训练到达到给定误差阈值时的排放量,并且在大多数示例中,物理信息模型的排放量更低(输入增广模型除外)。训练排放量的减少进一步叠加了通过收集和存储更少数据所实现的环境节约。尽管结果有前景,但我们不能期望有银弹,并且案例研究表明,需要在将物理知识引入机器学习器所带来的复杂性增加与减少训练数据需求所带来的收益之间进行权衡。
cs.LG / 274 / 2609.33391

Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents

超越时间戳:面向长时程智能体的决策对齐同策略蒸馏
Chen, Mingju, Lv, Can, Liu, Jinrong, Zhang, Huan, Chang, Heng, Zhou, Shiji
Abstract
Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify \emph{Decision--Timestamp Mismatch}: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce \textsc{AlignOPSD}, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate \textsc{AlignOPSD} with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. \textsc{AlignOPSD} outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 \% and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at https://github.com/mingju-c/Align-OPSD
Chinese Translation
带可验证奖励的强化学习(RLVR)通常依赖稀疏的结果奖励,为长时程智能体提供粗粒度的监督。同策略自蒸馏(OPSD)以密集的特权反馈补充这一信号。然而,我们发现了决策-时间戳不匹配:特权指导可能与学生的功能性决策不对齐,因为对应的决策可能发生在不同的时间步,而学生自身的决策可能跨越多个时间步,而非绑定到单个时间戳。因此,时间戳局部监督可能使上下文和信用分配的时间范围都发生错位。为了解决这一不匹配,我们引入了AlignOPSD,遵循先对齐监督再分配信用的原则。决策对齐的监督矫正通过在功能性匹配的上下文中对相同的学生采样响应在兄弟轨迹之间重新打分,以校准局部教师证据。半马尔可夫分层信用分配随后从对应关系变化中推导出可变时长的决策跨度,并使用矫正后的证据在跨度及其组成轮次之间分配基于结果的信用。我们在ALFWorld、WebShop和Search-QA上,使用Qwen2.5-3B和Qwen2.5-7B对AlignOPSD与代表性基线进行了评估。在所有八项骨干-聚合指标比较中,AlignOPSD均优于GRPO和StepOPSD,比GRPO提升5.5-8.7%,并在六项中排名第一。额外的分析考察了两个对齐阶段以及任务间的超参数敏感性。我们的代码可在 https://github.com/mingju-c/Align-OPSD 获取。
cs.LG / 275 / 2609.33392

AutoHGNN: Robust and Efficient Neural Architecture Search for Hypergraph Neural Networks

AutoHGNN:面向超图神经网络的鲁棒且高效的神经架构搜索
Li, Sirui, b, Pietro Liò, Li, Xinsheng, Liu, Baisong, Peng, Chengbin
Abstract
Hypergraph neural networks have achieved significant success in recent years. However, manual architecture crafting is labor-intensive and often fails to capture complex, higher-order relations, making the automation of hypergraph neural network structure design crucial. To improve the automation and adaptability of hypergraph learning, this paper proposes AutoHGNN, a neural architecture search framework tailored for hypergraph neural networks. First, we introduce a Hyper-Interaction Module (HIM) into the search space to address the mismatch between conventional graph neural network designs and hypergraph data. Second, we propose Hypergraph Stable Topological Distance (HyperSTD) as a structural selection criterion to identify architectures that best preserve the intrinsic structural affinities of the original hypergraph during differentiable search. Extensive experiments on various benchmark datasets demonstrate that AutoHGNN consistently outperforms manually designed and automatically searched baselines in classification accuracy and time efficiency, proving that the discovered architectures are significantly more effective.
Chinese Translation
超图神经网络近年来取得了显著的成功。然而,手动设计架构是劳动密集型的,并且往往无法捕捉复杂的高阶关系,因此超图神经网络结构设计的自动化至关重要。为了提高超图学习的自动化和适应性,本文提出了 AutoHGNN,一个为超图神经网络量身定制的神经架构搜索框架。首先,我们在搜索空间中引入了超交互模块 (HIM),以解决传统图神经网络设计与超图数据之间的不匹配问题。其次,我们提出了超图稳定拓扑距离 (HyperSTD) 作为结构选择准则,以在可微搜索过程中识别最能保留原始超图内在结构亲和性的架构。在各种基准数据集上的大量实验表明,AutoHGNN 在分类准确率和时间效率方面始终优于手动设计和自动搜索的基线,证明所发现的架构显著更有效。
cs.LG / 276 / 2609.33405

Decoupling Token Roles in Autoregressive Pretraining

解耦自回归预训练中的词元角色
Yuan, Suqin, Lin, Runqi, Lin, Kevin Qinghong, Yu, Junchi, Feng, Lei, Russell, Chris, Liu, Tongliang
Abstract
Autoregressive pretraining increasingly draws on heterogeneous data, making it important to understand how a model learns from an individual token. The next-token prediction objective naturally identifies a token's contribution with its own loss. However, each token is not only a prediction target but also context for what follows. Using controlled corruption, we decouple these two roles and find a reversal: making a noisy token easier to predict reduces its damage as a target but increases it as context. The same decoupling helps explain text generated by language models: generation selects each token by its fit to the prefix, while its role as context is never tested against an independently determined continuation, because that continuation is generated to fit it. At known corrupted positions, acting through the context can reduce damage that removing the token's own loss does not. Understanding and controlling what a model learns from a token therefore requires decoupling its roles.
Chinese Translation
自回归预训练越来越多地利用异构数据,因此理解模型如何从单个词元中学习变得十分重要。下一词元预测目标自然地将一个词元的贡献等同于其自身的损失。然而,每个词元不仅是预测目标,也是后续内容的上下文。通过受控破坏,我们解耦了这两种角色,并发现了一种反转:让一个含噪词元更容易预测,会减少其作为目标的损害,但会增加其作为上下文的损害。同样的解耦有助于解释语言模型生成的文本:生成过程根据每个词元与前缀的匹配程度来选择它,而其作为上下文的作用从未与独立确定的后续内容进行检验,因为该后续内容是为了匹配它而生成的。在已知的损坏位置,通过上下文起作用可以减少损害,而移除词元自身的损失却无法做到这一点。因此,理解和控制模型从词元中学到什么,需要解耦其角色。
cs.LG / 277 / 2609.33408

StarBOA: Real-Time Mamba State-Space Unrolling for Sparse Radar Micro-Doppler in ISAC Networks

StarBOA:面向ISAC网络中稀疏雷达微多普勒的实时Mamba状态空间展开
Çelik, Mustafa Bora, Çelik, Ceren, Gazi, Orhan
Abstract
In Integrated Sensing and Communications (ISAC), radar sensing must operate under chirp subsampling with up to 90\% missing data. An attention-based baseline, limited to a 52~ms buffer, collapses toward maximum uniform entropy ($H=2.584$ bits) as sparsity increases, failing to capture long-range gait-cycle context. We propose StarBOA, which replaces attention with a causal Mamba state-space model that updates incrementally on a per-window basis without re-scanning past reconstructions. By maintaining a persistent state, StarBOA integrates over $100\times$ more temporal history at no additional per-step computational cost. StarBOA outperforms the baseline's published results across all sparsity levels, with SSIM gains increasing from $+0.0379$ at 50\% missing data to $+0.2472$ at 90\%. Each window is processed in 1.53~ms with zero lookahead, demonstrating efficient causal reconstruction under extreme chirp subsampling.
Chinese Translation
在集成传感与通信(ISAC)中,雷达感知必须在chirp子采样下运行,数据缺失率高达90%。一个基于注意力的基线受限于52 ms的缓冲区,随着稀疏度增加而坍缩至最大均匀熵(H=2.584比特),无法捕获长距离步态周期上下文。我们提出StarBOA,它用因果Mamba状态空间模型取代注意力,该模型基于逐窗口进行增量更新,无需重新扫描过去的重建结果。通过维护持久状态,StarBOA在不增加每步计算开销的情况下整合了超过100倍的时间历史。StarBOA在所有稀疏度水平上均优于基线已发表的结果,SSIM增益从50%数据缺失时的+0.0379提高到90%数据缺失时的+0.2472。每个窗口的处理时间为1.53 ms,且零前瞻,展示了在极端chirp子采样下的高效因果重建。
cs.LG / 278 / 2609.33410

FoldAttention: Declared-Reference Softmax for Fast Decode and Deterministic Backward

FoldAttention:用于快速解码与确定性反向传播的声明参考Softmax
Achanta, Sriman
Abstract
Autoregressive decode repeatedly streams a growing KV cache, making attention a major cost at long context. Existing high-performance kernels use online softmax, which discovers a row's normalization reference as it scans keys. Earlier contributions therefore remain provisional and may require rescaling. We argue that the reference need not be discovered: softmax is invariant to a common shift, so the reference only has to keep the weights in range. We present FoldAttention, an additive formulation of softmax attention that fixes a finite reference $Z_i$ before scanning the KV cache. Each weight $2^{s_{ij}-Z_i}$ is then final when computed, so contributions add across disjoint key ranges and their quotient equals softmax attention in real arithmetic. We use this property to develop two techniques for Hopper decode: (1) final weights gate key and value reads before the bytes are fetched, and a per-call depth $T$ cuts keys below $2^{-T}$ while keeping their mass, and (2) additive partials compose split KV and shared-prefix cascades without rescaling. On H100 at $T=16$, FoldAttention decodes seven real-model generations 1.36-2.30$\times$ faster than the fastest BF16 baseline, and up to 3.09$\times$ faster across MHA and GQA shapes, at an error within 1.5% of the lowest BF16 error on six of the seven; reading every key, it is 1.14-1.30$\times$ faster at matched error. We validate on Qwen3-8B that a whole decode step is up to 1.46$\times$ faster while likelihood and long-context accuracy match those under BF16 kernels. The same principle makes the backward deterministic: CTAs round bounded partial gradients onto an integer grid declared before the reduction and add them in any order. FoldAttention thereby removes the determinism tax: its deterministic backward is up to 1.84$\times$ faster than deterministic FlashAttention-3/4 and 1.05$\times$ faster than the fastest nondeterministic kernel.
Chinese Translation
自回归解码重复流式传输不断增长的KV缓存,使得注意力在长上下文中成为主要开销。现有的高性能内核使用在线softmax,它在扫描键时发现一行的归一化参考。因此,较早的贡献仍然是暂定的,并且可能需要重新缩放。我们认为参考不需要被发现:softmax对共同偏移具有不变性,因此参考只需要保持权重在范围内。我们提出FoldAttention,一种softmax注意力的加性公式,在扫描KV缓存之前固定一个有限参考$Z_i$。每个权重$2^{s_{ij}-Z_i}$在计算时就是最终的,因此贡献在不相交的键范围上相加,并且在实数算术中它们的商等于softmax注意力。我们利用这一性质为Hopper解码开发了两种技术:(1) 最终权重在字节被获取之前门控键和值读取,并且每次调用的深度$T$将低于$2^{-T}$的键截断,同时保留其质量;(2) 加性部分组合拆分KV和共享前缀级联,无需重新缩放。在H100上,当$T=16$时,FoldAttention对七个真实模型生成进行解码,比最快的BF16基线快1.36-2.30倍,在MHA和GQA形状上最高快3.09倍,在七个模型中有六个的误差在最低BF16误差的1.5%以内;读取每个键时,在匹配误差下快1.14-1.30倍。我们在Qwen3-8B上验证,整个解码步骤最高快1.46倍,而似然和长上下文准确率与BF16内核下相当。同样的原理使反向传播具有确定性:CTA将有界部分梯度舍入到归约前声明的整数网格上,并以任意顺序相加。因此,FoldAttention消除了确定性税:其确定性反向传播比确定性FlashAttention-3/4快达1.84倍,比最快的非确定性内核快1.05倍。
cs.LG / 279 / 2609.33416

Investigating the Effect of k-NN Preprocessing on Developing Graph Neural Networks: A Fairness-Based Perspective

探究k-NN预处理对图神经网络开发的影响:基于公平性的视角
Zafeiropoulos, Nikolaos, Mavrikos, Emmanouil, Tsekouras, George E.
Abstract
In this paper, a methodology to design fair graph convolutional neural networks (GCNs) is developed and tested over several application data sets. The graphs that are used as inputs to the network are constructed by a k-nearest neighbor-based preprocessing procedure, while fairness issues are considered in terms of the equalized odds criterion. To effectively incorporate the above heterogenous information, the equalized odds criterion is directly embedded into the model's optimization objective through an additional fairness-driven loss functional term. The proposed methodology investigates how varying the neighborhood size in the k-NN algorithm during graph construction influences both the classification performance and the fairness of the resulting models. Extensive experimentation is conducted on three real-world tabular datasets with known biases, evaluating the interplay between graph structure and fairness enforcement. The results demonstrate that the choice of the value of the parameter k critically impacts the performance trends, either steadily improving or peaking at intermediate values depending on dataset characteristics, while the application of fairness constraints significantly mitigates disparities in false positive and false negative rates across groups defined by the protected variable at hand, without incurring major sacrifices in overall accuracy. This study highlights the importance of jointly optimizing the graph construction process and fairness objectives in GCN-based learning, providing a systematic approach toward building more equitable and effective graph-based models.
Chinese Translation
本文提出了一种设计公平图卷积神经网络(GCNs)的方法,并在多个应用数据集上进行了测试。网络的输入图是通过基于k-近邻的预处理过程构建的,同时公平性问题根据均衡赔率(equalized odds)准则来考虑。为了有效整合上述异构信息,均衡赔率准则通过一个额外的公平性驱动的损失函数项直接嵌入到模型的优化目标中。所提出的方法研究了在图构建过程中改变k-NN算法中的邻域大小如何影响分类性能和所得模型的公平性。在三个具有已知偏差的真实世界表格数据集上进行了大量实验,评估图结构与公平性执行之间的相互作用。结果表明,参数k值的选择对性能趋势有重大影响,根据数据集特征,性能要么稳步提升,要么在中间值处达到峰值,而公平性约束的应用显著减轻了由手头受保护变量定义的各群体之间在假阳性和假阴性率方面的差异,且不会在整体准确率上造成重大损失。本研究强调了在基于GCN的学习中联合优化图构建过程和公平性目标的重要性,为构建更公平、更有效的基于图的模型提供了系统方法。
cs.LG / 280 / 2609.33424

A Light Bilevel Refinement Aligns Self-Supervised Representations for Stronger Task-Specific Learning

轻量级双层优化精炼对齐自监督表示以增强任务特定学习
Zakarias, Gustav Wagner, Tan, Zheng-Hua
Abstract
Self-supervised pretraining learns representations that are broadly transferable across downstream tasks, yet direct fine-tuning can be suboptimal due to misalignment between self-supervised and downstream task objectives, potentially degrading pretrained features beneficial to the downstream task. The BiSSL framework addressed this by introducing a transitional training stage formulated as a bilevel optimization problem, in which the downstream task objective guides the self-supervised learning process in refining pretrained representations to better facilitate subsequent fine-tuning. However, BiSSL relies on conventional bilevel optimization solving techniques whose costly implicit hypergradient approximations render the method increasingly impractical for contemporary model architectures. To make it efficient and scalable, we introduce BiSSLight, which combines M-FAC-based implicit gradient approximation with parameter-efficient fine-tuning via LoRA, enabling efficient application at larger scales that were previously impractical. Evaluation across multiple downstream tasks and contemporary model architectures shows that BiSSLight consistently improves downstream performance, with gains becoming more pronounced as model size increases despite stronger baselines. The method is highly computationally efficient, reducing computation time by more than a factor of ten compared to its predecessor on a ViT-H backbone.
Chinese Translation
自监督预训练学习到的表示在下游任务中具有广泛的迁移性,然而由于自监督与下游任务目标之间的不对齐,直接微调可能并非最优,可能损害对下游任务有益的预训练特征。BiSSL框架通过引入一个过渡训练阶段来解决这个问题,该阶段被表述为双层优化问题,其中下游任务目标指导自监督学习过程,以精炼预训练表示,从而更好地促进后续微调。然而,BiSSL依赖于传统的双层优化求解技术,其代价高昂的隐式超梯度近似使得该方法对于当代模型架构越来越不实用。为了使其高效且可扩展,我们提出了BiSSLight,它结合了基于M-FAC的隐式梯度近似与通过LoRA的参数高效微调,从而能够在以前不切实际的更大规模上高效应用。跨多个下游任务和当代模型架构的评估表明,BiSSLight持续提升下游性能,并且随着模型规模增大,尽管基线更强,增益也变得更加显著。该方法在计算上非常高效,在ViT-H骨干网络上,与其前身相比,计算时间减少了十倍以上。
cs.LG / 281 / 2609.33431

MoGround: Measuring and Mitigating Modality Distraction in Vision-Language Models

MoGround:测量与缓解视觉语言模型中的模态干扰
Zhou, Luca, Zhao, Bo, Yu, Rose, Rodolà, Emanuele, Dessì, Roberto
Abstract
We release MoGround, a vision-language dataset spanning four visual domains in which the answer to every question is guaranteed to be available from exactly one modality. This guarantee enables us to measure modality distraction, the failure in which a model answers a question correctly from one modality alone and then flips to a wrong answer once irrelevant content from the other modality is added. Existing probes rarely establish single-modality answerability this way, making it hard to isolate distraction in the first place. Across seven open-source VLMs, we find that modality distraction is not universal but model-dependent. The weaker-grounded modality is the more distracted one (r = +0.86), and distraction scales inversely with grounding strength (r = -0.90). The single-modality guarantee also enables a mitigation method that needs to distinguish between relevant and irrelevant context. Trained on one split of MoGround alone, a weight-space robustness vector reduces distraction on all seven models by 9% to 51%, at a cost of only 0.1 average points of accuracy on standard multimodal tasks.
Chinese Translation
我们发布了 MoGround,这是一个视觉语言数据集,涵盖四个视觉领域,其中每个问题的答案都保证可以仅从一种模态中获得。这一保证使我们能够测量模态干扰,即模型仅从一种模态正确回答问题,但当加入来自另一种模态的无关内容后,模型会翻转为错误答案的失败情况。现有的探针很少以这种方式建立单模态可回答性,这使得首先很难分离出干扰。在七个开源 VLM 中,我们发现模态干扰并非普遍存在,而是依赖于模型。基础较弱的模态更易受干扰(r = +0.86),且干扰程度与基础强度呈负相关(r = -0.90)。单模态保证还实现了一种缓解方法,该方法需要区分相关和无关的上下文。仅在一个 MoGround 划分上训练,一个权重空间鲁棒性向量就将所有七个模型上的干扰减少了 9% 到 51%,而在标准多模态任务上的准确率平均仅损失 0.1 个百分点。
cs.LG / 282 / 2609.33436

SchemaMem: Schema-Indexed Recurrent Memory for Delayed State Retrieval

SchemaMem:面向延迟状态检索的模式索引循环记忆
Goo, Sungwoo, Yun, Hwi-yeol, Jung, Sangkeun
Abstract
Attention provides direct access to past representations, but retaining an ever-growing history is costly. Recurrent models bound persistent state, yet must preserve selected information while processing subsequent inputs. We introduce SchemaMem, an attention-based recurrent memory architecture combining chunk-local attention with a persistent, schema-indexed phase state. Learned schema embeddings provide a shared representational reference for reading and writing. Reads use the current state, whereas writes use the layer input and static schema embeddings, excluding direct feedback from that layer's own state. Chunk-boundary commits aggregate bounded phase increments through forward computation. The same parameters also support full-history attention training before and during recurrent training. We studied selective updates, preservation, and delayed retrieval in a controlled address--value task, comparing three-layer models with approximately matched parameter counts and persistent-state dimensions. Across nine address/value settings and three training seeds, SchemaMem has higher mean written-value retention at four times the maximum training delay than both baselines, which are trained toward a higher in-range accuracy target. Updated-value recovery favors SchemaMem in all nine settings against Mamba-3 and seven against Gated DeltaNet. Defaults consistently favor Gated DeltaNet over SchemaMem at that delay, and SchemaMem requires substantially more optimization steps. These results identify a promising retention--optimization trade-off in schema-indexed recurrence.
Chinese Translation
注意力机制可以直接访问过去的表示,但保留不断增长的历史代价高昂。循环模型对持久状态加以限制,但在处理后续输入的同时必须保留选定信息。我们提出 SchemaMem,一种基于注意力的循环记忆架构,将分块局部注意力与持久化的、模式索引的阶段状态相结合。学习到的模式嵌入为读取和写入提供共享的表示参考。读取使用当前状态,而写入使用层输入和静态模式嵌入,排除来自该层自身状态的直接反馈。分块边界提交通过前向计算聚合有界的阶段增量。相同的参数还支持在循环训练之前和期间进行全历史注意力训练。我们在一个受控的地址-值任务中研究了选择性更新、保持和延迟检索,比较了参数量和持久状态维度大致匹配的三层模型。在九种地址/值设置和三个训练种子下,SchemaMem 在四倍于最大训练延迟处的平均写入值保持率高于两个基线,而这两个基线是以更高的范围内准确率目标进行训练的。在全部九种设置中,更新值恢复相对于 Mamba-3 更有利于 SchemaMem,在七种设置中相对于 Gated DeltaNet 更有利于 SchemaMem。在该延迟下,默认设置始终更倾向于 Gated DeltaNet 而非 SchemaMem,并且 SchemaMem 需要显著更多的优化步骤。这些结果揭示了模式索引循环中一个有前景的保持-优化权衡。
cs.LG / 283 / 2609.33437

SMAT: Simple and Efficient Merge-Aware Training

SMAT:简单高效的合并感知训练
Gu, Yanggan, Wang, Yuanyi, Li, Zhen, Cai, Shuo, Liu, Yuhang, Li, Junzhuo, Wang, Zihao, Yang, Hongxia
Abstract
Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance, but existing methods do not fully account for common merging operations and add training cost. We observe that, from an expert's perspective, common merging methods can be described by three operations: Scale reweights its own update, Mask removes selected coordinates, and Perturb adds updates from other experts. Based on this view, we introduce SMAT (Simple MAT), which jointly optimizes expert loss and expected loss at simulated merged parameters generated by sampling scaling coefficients, masks, and additive noise. We further introduce periodic scheduling, kernel fusion, and parameter storage switching to make SMAT efficient, with one forward and one backward pass per step. Across four language and vision-language backbones, SMAT improves the mean score across five merging methods by 1.07-2.16 points over the strongest baseline for each backbone, with less than 2% training-time overhead over standard fine-tuning.
Chinese Translation
模型合并整合了多个专家的能力,无需联合重训练,但标准的专家训练仅优化任务损失,并不能保证合并后的性能良好。合并感知训练(MAT)旨在提升合并后的性能,但现有方法未能充分考虑常见的合并操作,并且增加了训练成本。我们观察到,从专家的角度来看,常见的合并方法可以通过三种操作来描述:Scale(缩放)重新加权其自身的更新,Mask(掩码)移除选定的坐标,Perturb(扰动)添加来自其他专家的更新。基于这一视角,我们提出了 SMAT(简单 MAT),它联合优化专家损失和在模拟合并参数处的期望损失,这些模拟合并参数通过采样缩放系数、掩码和加性噪声生成。我们进一步引入了周期性调度、核融合和参数存储切换,以使 SMAT 高效,每一步仅需一次前向和一次后向传播。在四个语言和视觉语言主干网络上,SMAT 在五种合并方法上的平均得分比每个主干网络的最强基线提高了 1.07-2.16 分,而相比标准微调,训练时间开销不到 2%。
cs.LG / 284 / 2609.33444

Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning

阐明基于回归的扩散强化学习的设计空间
Li, Toyota, Zhao, David, Zhao, Alan
Abstract
A nascent family of methods that forgoes the policy gradient and reweights a supervised regression instead has garnered momentum in reinforcement learning for diffusion and flow models. DiffusionNFT, FlowAWR, and RAM are representative regimes with contrasting motivations. It is yet opaque what, if anything, they share. We substantiate that each is the solution of one divergence-constrained reward-maximization problem, and they are differentiated only by the convex generator that defines the constraint. Under the unified modeling framework, we unravel the relaxations that prior art made during building the advantage-embedded regression target: approximating the KKT condition and posterior normalizer for the linear and exponential tilt shapes DiffusionNFT and FlowAWR respectively, while preserving the exact sparsemax projection onto the probability simplex for linear tilt leads to another superior model type in this work. Beyond the theoretical underpinnings, we further empirically investigate the design space and shed light on the training recipe for regression-style diffusion RL. Retaining the merits discovered during our exploration gives rise to DiffusionRFT, our paradigm that converges faster, trains more stably, and attains the top performance.
Chinese Translation
一类新兴的方法放弃了策略梯度,转而重新加权监督回归,已在扩散模型和流模型的强化学习中获得了发展势头。DiffusionNFT、FlowAWR 和 RAM 是具有不同动机的代表性范式。它们之间有何共同点(如果有的话)尚不清楚。我们证实,它们各自都是一个散度约束奖励最大化问题的解,并且仅由定义约束的凸生成函数所区分。在统一的建模框架下,我们揭示了先前工作在构建优势嵌入回归目标时所做的松弛:分别针对线性倾斜和指数倾斜形式的 DiffusionNFT 和 FlowAWR 近似 KKT 条件和后验归一化常数,而在本工作中,保持线性倾斜的精确 sparsemax 投影到概率单纯形上导致了另一种更优的模型类型。除了理论基础之外,我们进一步通过实验研究了设计空间,并为回归式扩散强化学习的训练配方提供了启示。保留我们在探索过程中发现的优点,催生了 DiffusionRFT,我们的范式收敛更快、训练更稳定,并达到了顶级性能。
cs.LG / 285 / 2609.33457

A Free Knob: Decoupling Calibration and Predictive Skill in Threshold-Based Evaluation

一个自由旋钮:在基于阈值的评估中解耦校准与预测技能
Munim, Md Tanveer Hossain, Saiem, Bijoy Ahmed, Sany, Al-Amin, Hashem, Tanzima
Abstract
Many dense-prediction benchmarks evaluate rare events by pooling prediction and target over spatial blocks, thresholding each, and scoring the contingency table. At a fixed rare operating point, the max-pooled Critical Success Index (CSI) confounds spatial discrimination with amplitude calibration: sharp observations promote many blocks above threshold, while attenuated predictions from squared-error regression leave the same blocks below it. We repurpose classical monotone calibration as a symmetric audit: a post-hoc transform fitted on held-out data and applied separately to each system. The transform cannot reverse pixel ordering, so any contrast it reproduces cannot establish improved spatial ranking. On SEVIR, two released checkpoints of one architecture differ by -29.5% in extreme-threshold CSI before the control and by +5.3% after it. Across 450 pairwise contrasts among 6 systems, the difference in pooled frequency-bias deviation is associated with how far the CSI contrast moves under the control (r = +0.796), and 51 contrasts reverse sign. At CasCast's published extreme-event operating point, the cascade-over-backbone CSI gap falls from 0.1601 to 0.0339, a 78.8% reduction; the remaining gap stays positive. The effect persists when the transform is fitted on a window before the test period, and calibration also reveals advantages hidden by a better-calibrated baseline. On geostationary infrared imagery the relative gain grows as events become rarer, crowd counting reproduces the bias-gain relationship under patch-sum pooling, and semantic segmentation, where frequency bias is already near one, shows little average change. The confound therefore requires both a fixed operating point and a training regime that leaves the output miscalibrated there. We recommend reporting pooled frequency bias and a symmetric held-out FreeKnob Audit alongside rare-event pool-and-threshold scores.
Chinese Translation
许多密集预测基准通过将预测和目标在空间块上池化,对每个块进行阈值化,并对列联表进行评分来评估稀有事件。在固定的稀有操作点,最大池化的临界成功指数(CSI)将空间辨别力与幅度校准混淆:锐利的观测使许多块超过阈值,而来自平方误差回归的衰减预测使相同的块低于阈值。我们将经典单调校准重新用作对称审计:一种事后变换,在留出数据上拟合,并分别应用于每个系统。该变换不能反转像素顺序,因此它所重现的任何对比都不能证明空间排序的改善。在SEVIR上,一个架构的两个已发布检查点在极端阈值CSI上,在控制前相差-29.5%,控制后相差+5.3%。在6个系统之间的450个成对对比中,池化频率偏差差异与控制下CSI对比移动的程度相关(r = +0.796),并且51个对比反转符号。在CasCast发布的极端事件操作点,级联相对于骨干的CSI差距从0.1601下降到0.0339,减少了78.8%;剩余差距保持为正。当变换在测试期之前的时间窗口上拟合时,该效应持续存在,并且校准还揭示了被更好校准的基线所隐藏的优势。在地球静止红外图像上,相对增益随着事件变得更稀有而增长,人群计数在块求和池化下重现了偏差-增益关系,而语义分割(其频率偏差已经接近1)显示平均变化很小。因此,这种混淆需要同时具备固定的操作点和使输出在该处校准不良的训练机制。我们建议报告池化频率偏差和对称留出的FreeKnob审计,与稀有事件池化和阈值评分一起。
cs.LG / 286 / 2609.33467

A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards

廉价验证器已足够:LLM 后训练对错误奖励具有鲁棒性
Plesner, Andreas, Northcutt, Curtis, Guzmán, Francisco, Athalye, Anish
Abstract
When post-training large language models on tasks with semi-verifiable rewards, there are many factors (training steps, base model size, training order, data quality, verifier accuracy, etc.) that practitioners must contend with to maximize model performance. Yet, it remains unclear how well verifier agreement predicts post-training performance on such tasks. In this paper, we explore this question with over 11k H100 GPU-hours, across HealthBench and PRBench tasks in medical, legal, and finance domains. Across the tested domains, Qwen3 trainees (1.7B-8B on HealthBench; 8B on PRBench), evaluation splits, and frontier LLM reference judges (which we call golden verifiers), higher verifier agreement does not consistently identify the best training verifier. Expensive verifiers need not outperform inexpensive ones, and open-weight Gemma verifiers produce strong training outcomes. We compare two low-cost choices retrospectively -- a cost-reducing choice and a balanced choice -- with estimated grading cost reductions of 98.8%-99.7% relative to the golden grading protocols and average post-training score gaps of 1-3 points from the best evaluated training verifier. These averages include larger losses in individual settings; they do not establish that verifier choices are interchangeable.
Chinese Translation
在具有半可验证奖励的任务上后训练大语言模型时,从业者必须应对许多因素(训练步数、基础模型大小、训练顺序、数据质量、验证器准确率等)以最大化模型性能。然而,验证器一致性在多大程度上能预测此类任务上的后训练性能仍不清楚。在本文中,我们利用超过 11k H100 GPU 小时,在医学、法律和金融领域的 HealthBench 和 PRBench 任务上探索了这个问题。在测试的领域中,无论是 Qwen3 训练模型(HealthBench 上为 1.7B-8B;PRBench 上为 8B)、评估划分,还是前沿 LLM 参考评判器(我们称之为黄金验证器),更高的验证器一致性并不总能识别出最佳训练验证器。昂贵的验证器未必优于廉价的验证器,而开放权重 Gemma 验证器产生了强劲的训练结果。我们回顾性地比较了两种低成本选择——一种降本选择和一种平衡选择——相对于黄金评分协议,其估计评分成本降低了 98.8%-99.7%,并且与最佳评估训练验证器的平均后训练分数差距为 1-3 分。这些平均值包括个别设置中更大的损失;它们并不表明验证器选择可以互换。
cs.LG / 287 / 2609.33472

Geometric Identification in Predict-Then-Optimize Learning

预测-优化学习中的几何识别
Xu, Jiaxiao, Mou, Changhong, Liu, Keji, Xu, Dinghua, Zhang, Yeyu
Abstract
Decision-focused surrogates can recover downstream decisions without identifying the quotient report. We characterize the equality set of the convex Smart Predict-then-Optimize surrogate (SPO+) population risk. Under central symmetry, the centered mean class is the unique Bayes minimizer exactly when every nonzero effective displacement makes the old optimizer leave the shifted optimal face with positive probability. This condition separates face crossing from selected-oracle disagreement and gives quantitative local coercivity. Without symmetry, strict crossing alone need not identify the mean; selection balance with reflected crossing restores quotient-report identification, and conditional versions extend the result to measurable predictors. These are population statements, without finite-sample report-recovery or generic transfer-regret guarantees. Closed-form mechanisms reproduce the analytic identities and rates. Portfolio, complete-matrix KuaiRec, and Energy/Storage studies measure predictive fidelity, shifted regret, and fitted-report geometry. A known data-generating process (DGP) companion retains their application geometries while isolating conditional-mean recovery and crossing, without testing the original observational assumptions.
Chinese Translation
决策聚焦的代理模型可以在不识别商报告的情况下恢复下游决策。我们刻画了凸的Smart Predict-then-Optimize代理(SPO+)总体风险的等值集。在中心对称下,中心均值类是唯一的贝叶斯最小化器,当且仅当每个非零有效位移使得旧优化器以正概率离开平移后的最优面。该条件将面穿越与所选预言机不一致区分开来,并给出了定量的局部强制性。在没有对称性的情况下,仅靠严格穿越不足以识别均值;选择平衡与反射穿越共同恢复商报告识别,条件版本将结果扩展到可测预测器。这些是总体层面的陈述,不提供有限样本的报告恢复或一般迁移遗憾保证。闭式机制重现了解析恒等式和速率。投资组合、完整矩阵KuaiRec以及能源/存储研究测量了预测保真度、平移遗憾和拟合报告几何。一个已知的数据生成过程(DGP)配套保留了它们的应用几何,同时分离出条件均值恢复和穿越,而不检验原始观测假设。
cs.LG / 288 / 2609.33482

How Synthetic Labels Improve Conformal Prediction: A Perspective on Conditional Coverage

合成标签如何改进共形预测:条件覆盖的视角
Chen, Qianyi, Li, Bo
Abstract
Conformal prediction provides distribution-free finite-sample marginal coverage, but post-hoc calibration data may be too scarce to learn how uncertainty varies across inputs. Meanwhile, abundant covariates can often be labeled cheaply by domain models or general-purpose language models. We study whether these synthetic labels can improve conditional coverage when only a small trusted sample is available. Building on score-quantile regression, we introduce prediction-powered quantile learning: a synthetic-labeled pool estimates pinball risk, paired trusted and synthetic outcomes correct its bias, and an independent trusted split performs final conformalization. Profiling pinball risk over scalar corrections reveals that population conditional-coverage error is its functional gradient; the corresponding Hessian removes global shifts and weights remaining shape error by boundary density. Composing this geometry with prediction-powered learning yields a three-resource expansion and a benefit--cost rule for synthetic power. Across eight regression benchmarks, synthetic-powered quantile learning substantially improves downstream conditional coverage while preserving marginal validity and producing more compact prediction sets. A human-rating study finds similar gains from external LLM labels and exposes a quality--quantity--cost tradeoff.
Chinese Translation
共形预测提供了无分布有限样本的边缘覆盖,但事后校准数据可能过于稀缺,无法学习不确定性如何随输入变化。同时,大量的协变量通常可以通过领域模型或通用语言模型廉价地标注。我们研究当只有少量可信样本可用时,这些合成标签能否改善条件覆盖。基于分数-分位数回归,我们引入了预测驱动的分位数学习:一个合成标记池估计弹球风险,成对的可信和合成结果校正其偏差,独立的可信划分执行最终共形化。通过对标量校正进行弹球风险剖析,揭示了总体条件覆盖误差是其函数梯度;相应的 Hessian 消除了全局偏移,并通过边界密度对剩余形状误差进行加权。将此几何与预测驱动的学习相结合,产生了三资源扩展和合成效力的收益-成本规则。在八个回归基准上,合成驱动的分位数学习显著改善了下游条件覆盖,同时保持了边缘有效性,并产生了更紧凑的预测集。一项人类评分研究发现了来自外部 LLM 标签的类似收益,并揭示了质量-数量-成本权衡。
cs.LG / 289 / 2609.33487

What masking geometry works best for EEG foundation models?

哪种掩码几何结构最适合EEG基础模型?
Guetschel, Pierre, Aristimunha, Bruno, Ouahidi, Yassine El, Delorme, Arnaud, Moreau, Thomas, Tangermann, Michael
Abstract
EEG foundation models hold promise for scalable brain-signal decoding across clinical and cognitive neuroscience applications, yet their pre-training pipelines remain poorly understood. Among design choices, the masking strategy is particularly critical: it determines what the network must predict and from which context. Yet it has never been ablated in isolation, as each new model bundles a new masking strategy with a new backbone and objective. In this paper, we formalize the design choices for spatio-temporal masking strategies and train various models with a single pipeline under varying masking configurations across two SSL frameworks (MAE and JEPA). We then systematically evaluate the resulting 58 pre-trained models on the 12 datasets of OpenEEGBench under a linear probe. Both frameworks agree on an optimal masking configuration and on shared failure modes. Outside these, performance is robust: 11 MAE and 9 JEPA configurations are statistically indistinguishable from the best. We further identify a novel JEPA-specific failure mode, tagged bias-inflation collapse, invisible to standard detectors. With a well-chosen mask, our pipeline reaches REVE-level downstream performance at a fraction of REVE's pre-training compute.
Chinese Translation
EEG基础模型有望在临床和认知神经科学应用中实现可扩展的脑信号解码,但其预训练流程仍未被充分理解。在设计选择中,掩码策略尤为关键:它决定了网络必须预测什么以及从何种上下文中进行预测。然而,它从未被单独进行消融研究,因为每个新模型都将新的掩码策略与新的骨干网络和目标函数捆绑在一起。在本文中,我们形式化了时空掩码策略的设计选择,并在两种自监督学习(SSL)框架(MAE和JEPA)下,使用单一流程在不同掩码配置中训练了多种模型。随后,我们在OpenEEGBench的12个数据集上,通过线性探针系统地评估了由此得到的58个预训练模型。两种框架在最优掩码配置和共同失败模式上达成一致。在这些情况之外,性能仍很稳健:11种MAE和9种JEPA配置与最佳配置在统计上无法区分。我们进一步发现了一种新的JEPA特有失败模式,称为偏差膨胀崩溃(bias-inflation collapse),标准检测器无法察觉。在精心选择掩码的情况下,我们的流程以REVE预训练计算量的一小部分达到了REVE级别的下游性能。
cs.LG / 290 / 2609.33489

Predicting Block-Coordinate Performance via Cross-Curvature

通过交叉曲率预测块坐标性能
Zhu, Shengkun, Zeng, Jinshan, Kou, Zhiqiang, Tong, Yongxin, Liu, Yang
Abstract
Simultaneous and sequential block updates are two basic optimization strategies used across machine learning, such as neural-network training, federated learning, and low-rank adaptation. Choosing between them is difficult because their relative advantage depends on both the objective geometry and the number of iterations. We develop a unified theory for comparing Jacobi (JC), Gauss--Seidel (GS), and partially sequential deterministic block-gradient updates. Our analysis expresses the one-step loss difference through cross-block curvature, with an $O(\eta^3)$ remainder, where $\eta$ is the learning rate. We derive a signed loss comparison after $K$ iterations with $O(K\eta^3)$ error under regularity conditions and $\eta K\le T$ for fixed $T$, identifying the better method when the predicted difference exceeds this error. We evaluate these formulas along observed training trajectories across different machine learning settings. Over 500 iterations, our theory correctly identifies the lower-loss method in 98.0\% of iterations for the neural network, 83.4\% for federated learning, and 97.6\% for LoRA. Applying the loss recursion at each step using the measured parameter difference raises these rates to 100.0\%, 93.2\%, and 99.6\%, respectively.
Chinese Translation
同步和顺序块更新是机器学习中两种基本的优化策略,例如神经网络训练、联邦学习和低秩适应。在两者之间进行选择很困难,因为它们的相对优势取决于目标几何形状和迭代次数。我们建立了一个统一理论,用于比较Jacobi (JC)、Gauss-Seidel (GS)以及部分顺序的确定性块梯度更新。我们的分析通过跨块曲率来表示一步损失差异,并带有O(η^3)的余项,其中η是学习率。我们在正则性条件下,对于固定的T且ηK≤T,推导出经过K次迭代后带有O(Kη^3)误差的带符号损失比较,当预测差异超过该误差时识别出更好的方法。我们在不同的机器学习设置中,沿着观察到的训练轨迹评估这些公式。在超过500次迭代中,我们的理论在神经网络中98.0%的迭代、联邦学习中83.4%的迭代以及LoRA中97.6%的迭代中正确识别出损失更低的方法。使用测得的参数差异在每个步骤应用损失递归,将这些比率分别提高到100.0%、93.2%和99.6%。
cs.LG / 291 / 2609.33496

Chameleon: Dynamic Format Adapter for Efficient Diffusion

Chameleon:用于高效扩散的动态格式适配器
Sanyal, Arnab, Chinchali, Sandeep
Abstract
Post-training quantization (PTQ) is the standard way to run modern diffusion models on memory-constrained accelerators, yet every existing diffusion PTQ scheme fixes the $\mathit{number\ format}$ in advance and only tunes the scale, zero point, or per-layer bit-width. At a fixed bit-width the best format depends on the distribution being encoded, and that distribution differs across weight channels, across layers, and along the diffusion timestep, where activation distributions slide from heavy-tailed and noise-dominated to tightly clustered and structured. We propose Chameleon, a PTQ framework that holds the bit-width fixed and treats the format itself as a discrete variable, chosen per weight channel and per (layer, timestep bucket) activation tensor. Activation formats come from {INT8, FP8 E4M3, FP8 E5M2, MXFP8, MXINT8}, selected ahead of time from two cheap statistics (empirical kurtosis and the closed-form diffusion SNR) and stored in a lookup table; weight formats come from {INT8, MXINT8} at 8 bits or {INT4, NF4, FP4 E2M1, MXINT4, MXFP4} at 4 bits, selected offline by reconstruction error. An architectural fork adapts the same selection layer to multi-step UNets, single-step distilled models, and Diffusion Transformers. Across SDXL, SDXL-Turbo, and PixArt-$\alpha$ on COCO-2014, Chameleon achieves the best FID in all six backbone $\times$ bit-width settings, with CLIP within 0.24 of the FP16 reference and the best of all quantized methods at $W_{4}A_{8}$.
Chinese Translation
训练后量化(PTQ)是在内存受限的加速器上运行现代扩散模型的标准方法,然而现有的扩散PTQ方案都预先固定了$\mathit{数字格式}$,仅调整缩放因子、零点或每层位宽。在位宽固定时,最佳格式取决于被编码的分布,而该分布因权重通道、层以及扩散时间步而异,其中激活分布从重尾和噪声主导转变为紧密聚集和结构化。我们提出了Chameleon,一种PTQ框架,它保持位宽固定,并将格式本身视为离散变量,按权重通道和按(层,时间步桶)激活张量进行选择。激活格式来自{INT8, FP8 E4M3, FP8 E5M2, MXFP8, MXINT8},根据两个廉价统计量(经验峰度和闭式扩散SNR)提前选择并存储在查找表中;权重格式来自8位下的{INT8, MXINT8}或4位下的{INT4, NF4, FP4 E2M1, MXINT4, MXFP4},通过重建误差离线选择。一种架构分叉将相同的选择层适配到多步UNet、单步蒸馏模型和Diffusion Transformer。在COCO-2014上的SDXL、SDXL-Turbo和PixArt-$\alpha$上,Chameleon在所有六种主干 $\times$ 位宽设置中取得了最佳FID,CLIP与FP16参考相差在0.24以内,并且在$W_{4}A_{8}$下是所有量化方法中最好的。
cs.LG / 292 / 2609.33497

Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State

Hamiltonian JEPA:具有继承控制状态的动作条件世界模型
Zoabi, Tamim, Ali, Ameen, Wolf, Lior
Abstract
Planning from pixels needs more than a latent space that is stable and predictable. The state the planner scores must also be organized by how actions move the system. Joint-embedding predictive architectures (JEPAs) avoid pixel reconstruction by predicting future representations, but existing action-conditioned JEPAs ask one embedding to serve both perception and control. We introduce H-JEPA, which separates the two. A wide perceptual code is regularized toward a well-scaled isotropic geometry with a Bures-Wasserstein prior, and a fixed orthonormal slice of that code is the control state, which inherits the code's covariance without any objective of its own. The state evolves under phase-conditioned dissipative port-Hamiltonian dynamics whose input port has orthonormal columns. Port-inverse consistency (PIC) reads the executed action back through the transpose of that port. We show that this readout is exactly the rollout error projected onto the port directions, so PIC is a parameter-free reweighting of prediction error and not an auxiliary action decoder. Untying the readout from the port breaks this identity and loses half of the gain. H-JEPA matches or exceeds reconstruction-free baselines, including the action-decoding Delta-JEPA, on four pixel-based control benchmarks after at most $10$ training epochs, and its largest gain is on OGB-Cube ($91.9$ against $79.3$ percent). Ablations on PushT and OGB-Cube separate the contributions of the structured predictor, PIC, the prediction horizon, the state rank, and the anti-collapse prior.
Chinese Translation
从像素进行规划需要的不仅仅是一个稳定且可预测的潜在空间。规划器所评估的状态还必须根据动作如何移动系统来组织。联合嵌入预测架构(JEPA)通过预测未来表示来避免像素重建,但现有的动作条件 JEPA 要求一个嵌入同时服务于感知和控制。我们引入了 H-JEPA,它将二者分离。一个宽感知编码通过 Bures-Wasserstein 先验被正则化趋向于良好缩放的各向同性几何,而该编码的一个固定正交切片是控制状态,它继承了编码的协方差,而无需自身任何目标。该状态在相条件耗散端口-哈密顿动力学下演化,其输入端口具有正交列。端口逆一致性(PIC)通过该端口的转置读回执行的动作。我们表明,这种读出恰好是投影到端口方向上的展开误差,因此 PIC 是预测误差的无参数重新加权,而不是辅助动作解码器。将读出与端口解绑会破坏这种恒等性,并损失一半的增益。在最多 $10$ 个训练周期后,H-JEPA 在四个基于像素的控制基准上匹配或超过了无重建基线(包括动作解码的 Delta-JEPA),其最大增益在 OGB-Cube 上($91.9$ 对 $79.3$ 百分比)。在 PushT 和 OGB-Cube 上的消融实验分离了结构化预测器、PIC、预测范围、状态秩和抗坍缩先验的贡献。
cs.LG / 293 / 2609.33499

The cost of useful natural gradient updates

有用自然梯度更新的代价
Bhattacharjee, Subhransu S., Campbell, Dylan, Shome, Rahul
Abstract
What information is needed to turn a natural-gradient direction into a useful finite update? Under a population Kullback-Leibler (KL) budget, we call a step useful if it is feasible and loses at most a fraction $\varepsilon$ of the best feasible gain along the direction. We construct a four-state exponential family whose laws share their initial gradient, scalar Fisher information and natural gradient, yet two laws have disjoint useful-step sets. With these quantities supplied exactly and the law otherwise known only through draws, the family's worst-case sample complexity is $\Theta(\log(1/\delta)/(p\varepsilon^2))$ for small $\varepsilon$, where $p$ scales rare-state probabilities and $\delta$ is the failure probability. The budget is fixed and the optimal gain stays bounded away from zero, so the step length, not the direction, carries this cost. For succinctly described event-tilt models, returning a useful step is NP-hard even with the exact natural gradient and efficient exact sampling. Recovering the unit natural gradient to constant error is also NP-hard even in a two-parameter logistic family with Fisher condition number at most 3. We also give matching sample bounds for event tilts, sample bounds for damped Fisher solves and a population-KL certificate for affine classifiers. In frozen-feature classifier heads, stopping at a sampled KL boundary succeeds in about half of the trials, and a 10% KL margin raises joint success above 93% at a KL budget of 0.01. Thus, knowing where to move is not enough: how far to move can carry an update's entire cost.
Chinese Translation
需要什么信息才能将一个自然梯度方向转化为有用的有限更新?在群体Kullback-Leibler (KL) 预算下,如果一步是可行的,并且最多损失沿该方向最佳可行增益的 $\varepsilon$ 比例,我们称这一步是有用的。我们构造了一个四状态指数族,其分布共享初始梯度、标量Fisher信息和自然梯度,但两个分布具有不相交的有用步集。在这些量被精确提供,而分布仅通过抽样得知的情况下,对于小的 $\varepsilon$,该族的最坏情况样本复杂度为 $\Theta(\log(1/\delta)/(p\varepsilon^2))$,其中 $p$ 与稀有状态概率成比例,$\delta$ 是失败概率。预算固定,最优增益保持有界远离零,因此代价由步长而非方向承担。对于简洁描述的事件倾斜模型,即使拥有精确的自然梯度和高效的精确采样,返回一个有用步也是NP难的。即使在Fisher条件数至多为3的双参数logistic族中,将单位自然梯度恢复到常数误差也是NP难的。我们还给出了事件倾斜的匹配样本界、阻尼Fisher求解的样本界以及仿射分类器的群体KL证书。在冻结特征分类器头中,在采样的KL边界处停止在大约一半的试验中成功,而在KL预算为0.01时,10%的KL裕度将联合成功率提高到93%以上。因此,知道往哪里移动是不够的:移动多远可能承担更新的全部代价。
cs.LG / 294 / 2609.33501

Pulseflow: PPG Counterfactual Generation Via Latent Transport

PulseFlow:基于潜在传输的PPG反事实生成
Pham, Hung Manh, Ma, Dong, Zhu, Bin, Zhou, Pan
Abstract
Photoplethysmography (PPG) has become an important modality for continuous cardiovascular monitoring, including atrial fibrillation (AF) detection. However, labeled AF recordings remain limited in many clinical settings, making model adaptation difficult when only limited target data are available. Generative modeling offers a natural way to alleviate this scarcity by synthesizing additional AF signals. Existing approaches, however, mainly generate samples that match the target condition without explicitly modeling how an observed source recording should be transformed, making it difficult to leverage abundant source recordings from a specific population or cohort for targeted augmentation. We introduce PulseFlow, a source-conditioned counterfactual generation framework that combines conditional representation learning with invertible latent transport to edit cardiac rhythm while retaining information from the source. Experiments across two clinical cohorts demonstrate effective rhythm transformation, measurable source correspondence, and improved AF classification under limited labels.
Chinese Translation
光电容积描记术(PPG)已成为连续心血管监测的重要模态,包括心房颤动(AF)检测。然而,在许多临床环境中,标注的AF记录仍然有限,使得仅有有限目标数据时模型适应困难。生成建模通过合成额外的AF信号,为缓解这种数据稀缺提供了一种自然途径。然而,现有方法主要生成匹配目标条件的样本,而没有显式建模应如何转换观察到的源记录,从而难以利用来自特定人群或队列的大量源记录进行定向增强。我们提出PulseFlow,一个源条件反事实生成框架,它将条件表示学习与可逆潜在传输相结合,以编辑心律,同时保留源信息。在两个临床队列上的实验证明了有效的心律转换、可测量的源对应关系,以及在有限标签下改进的AF分类。
cs.LG / 295 / 2609.33510

Source Anchoring for Physical Consistency in Flow Matching Models

面向流匹配模型物理一致性的源锚定
Romoli, Giulia, Ruffini, Filippo, Soda, Paolo
Abstract
Deep generative models are used to solve partial differential equations and model distributions of physical system states, but ensuring that the generated samples satisfy the governing laws remains challenging. Projection-based flow-matching methods enforce physics by correcting the flow from an unconstrained noise distribution. These corrections shift the generated samples away from the distribution of target solutions, especially in high noise regions. To address this limitation, we propose Source Anchoring for Physical Consistency (SAPC), a Functional Flow Matching method that encodes the physical constraints into the source noise before generation begins. We evaluate SAPC on five systems governed by partial differential equations, covering six tasks with linear and non-linear dynamics, and compare results against five baselines and the unconstrained backbone. Anchoring the source reduces the need for large corrections that drive samples onto admissible but off-distribution states, and SAPC reproduces the target distributions most accurately on every evaluated task, while matching the constraint precision of the best projection-based baselines. Ablation experiments show that this gain arises from pairing source projection with a matched training objective that regresses toward the projected source. These results identify the source distribution as a key design choice for physically consistent generative modelling.
Chinese Translation
深度生成模型被用于求解偏微分方程并建模物理系统状态的分布,但确保生成的样本满足控制规律仍然具有挑战性。基于投影的流匹配方法通过校正来自无约束噪声分布的流来施加物理约束。这些校正会将生成样本偏离目标解的分布,尤其是在高噪声区域。为克服这一局限,我们提出了面向物理一致性的源锚定(SAPC),这是一种函数式流匹配方法,可在生成开始之前将物理约束编码到源噪声中。我们在五个由偏微分方程控制的系统上评估 SAPC,涵盖具有线性和非线性动力学的六个任务,并将结果与五个基线以及无约束主干模型进行比较。对源进行锚定减少了对大幅校正的需求,这些校正会将样本驱动到可容许但偏离分布的状态;SAPC 在每个评估任务上都最准确地再现了目标分布,同时达到了最佳基于投影基线的约束精度。消融实验表明,这一收益来自将源投影与匹配的训练目标相结合,该目标回归到投影后的源。这些结果将源分布确定为物理一致生成建模的关键设计选择。
cs.LG / 296 / 2609.33515

SLP-ProbHard: Probabilistic Hard-Constrained Learning via Structural Latent Parameterization

SLP-ProbHard:基于结构潜变量参数化的概率硬约束学习
Bekele, Wondesen Teshome, D'Oria, Marco
Abstract
Many probabilistic predictors must satisfy exact structure in every stochastic realization, yet common hard-constraint approaches form predictions in ambient coordinates and then correct or project them. We introduce SLP-ProbHard, a cross-family, representation-centered framework for probabilistic hard-constrained learning when explicit structural parameterizations are available. Its core object, a Structural Feasible Latent Parameterization (SFLP), combines a structural latent law $Z \sim P^Z_\theta(\cdot\mid x)$ with a feasible map $Y=h_\phi(x,Z)$ that satisfies the constraint for every latent realization. Together these components define the predictive law itself, including its support and boundary probabilities, rather than serving as a final feasibility wrapper. We study how feasible coordinates and maps affect stochastic dimension, dependence, calibration, expressiveness, and computation. Experiments use Gaussian latent laws and fixed geometry-derived maps across affine equalities, ordering and simplex constraints, nonlinear manifolds, and three structural representations of seven-basin hydrological flow-duration-curve (FDC) data. In an official-source affine comparison with ProbHardE2E/DPPL, both methods achieve zero practical constraint violations. SLP-ProbHard uses 8 instead of 11 stochastic coordinates and improves MSE/MAE, while DPPL yields better marginal CRPS and closer-to-nominal coverage; a paired test detects no Energy Score difference across ten seeds. Real-world affine and nonlinear FDC representations reduce 13 to 7 and 14 to 8 ambient versus computational coordinates, respectively. Exact feasibility alone thus does not determine a predictive law, motivating direct structural generation when meaningful feasible coordinates are available.
Chinese Translation
许多概率预测器必须在每个随机实现中满足精确结构,然而常见的硬约束方法在环境坐标中形成预测,然后对其进行校正或投影。我们引入了 SLP-ProbHard,一个跨族、以表示为中心的框架,用于当显式结构参数化可用时的概率硬约束学习。其核心对象是结构可行潜变量参数化(SFLP),它将结构潜变量律 $Z \sim P^Z_{\theta}(\cdot\mid x)$ 与可行映射 $Y=h_{\phi}(x,Z)$ 相结合,该映射对每个潜变量实现都满足约束。这些组件共同定义了预测律本身,包括其支撑集和边界概率,而不是作为最终的可行性包装器。我们研究可行坐标和映射如何影响随机维度、依赖性、校准、表达能力和计算。实验使用高斯潜变量律和固定几何导出的映射,涵盖仿射等式、排序和单纯形约束、非线性流形,以及七流域水文流量历时曲线(FDC)数据的三种结构表示。在与 ProbHardE2E/DPPL 的官方来源仿射比较中,两种方法都实现了零实际约束违反。SLP-ProbHard 使用 8 个而非 11 个随机坐标,并改善了 MSE/MAE,而 DPPL 则产生更好的边际 CRPS 和更接近名义的覆盖率;配对检验在十个种子上未检测到 Energy Score 差异。真实世界的仿射和非线性 FDC 表示分别将环境坐标与计算坐标从 13 减少到 7,以及从 14 减少到 8。因此,仅凭精确可行性并不能确定预测律,这促使我们在有意义的可行坐标可用时进行直接结构生成。
cs.LG / 297 / 2609.33521

LLM4Trust: Exploring the Capabilities of Large Language Models for Trust Evaluation

LLM4Trust:探索大型语言模型在信任评估中的能力
Wang, Jie, Sun, Yanbo, Yan, Zheng, Lan, Jiahe, Bertino, Elisa
Abstract
Trust evaluation plays a critical role in cybersecurity by supporting risk mitigation and decision-making. A variety of trust evaluation methods have been proposed, with learning-based approaches offering high accuracy and automation. However, they often require substantial ground truth, suffer from low training efficiency, lack support for basic trust properties, and provide limited explainability. Large Language Models (LLMs) offer a compelling alternative due to their strong zero-/few-shot reasoning abilities and broad knowledge. To this end, we propose LLM4Trust, the first benchmark framework that systematically explores the capabilities of LLMs for trust evaluation. We first construct diverse trust graphs to model five basic trust properties and design corresponding property understanding tasks. We then assess the ability of eight representative LLMs to understand these properties under nine prompt methods. Based on this exploration, we identify the most effective LLM-prompt combinations and apply them to five real-world datasets for validating LLMs' trust evaluation capability. During this process, we propose two strategies to extract key information from large-scale trust graphs, addressing the context window limitations of LLMs. Extensive experiments show that LLMs can effectively understand basic trust properties and have great potential for real-world trust evaluation, particularly under limited supervision. However, they remain vulnerable to attacks targeting trust graphs and demonstration examples used in few-shot prompting, and incur high inference costs. Accordingly, we propose a defense mechanism and batch inference to improve the robustness and efficiency of LLM-based trust evaluation. The source code of LLM4Trust is available at https://github.com/Jieerbobo/LLM4Trust
Chinese Translation
信任评估通过支持风险缓解和决策,在网络安全中发挥着关键作用。已经提出了多种信任评估方法,其中基于学习的方法提供了高精度和自动化。然而,这些方法通常需要大量的真实标签,训练效率低,缺乏对基本信任属性的支持,并且可解释性有限。大型语言模型(LLMs)因其强大的零样本/少样本推理能力和广泛的知识,提供了一种有吸引力的替代方案。为此,我们提出了LLM4Trust,这是第一个系统探索LLMs信任评估能力的基准框架。我们首先构建了多样化的信任图来建模五个基本信任属性,并设计了相应的属性理解任务。然后,我们评估了八种代表性LLM在九种提示方法下理解这些属性的能力。基于这一探索,我们确定了最有效的LLM-提示组合,并将其应用于五个真实世界数据集,以验证LLM的信任评估能力。在此过程中,我们提出了两种策略来从大规模信任图中提取关键信息,解决了LLM的上下文窗口限制问题。大量实验表明,LLM能够有效理解基本信任属性,并在真实世界信任评估中具有巨大潜力,尤其是在有限监督下。然而,它们仍然容易受到针对信任图和少样本提示中使用的演示示例的攻击,并且推理成本高。因此,我们提出了防御机制和批量推理,以提高基于LLM的信任评估的鲁棒性和效率。LLM4Trust的源代码可在https://github.com/Jieerbobo/LLM4Trust获取。
cs.LG / 298 / 2609.33527

Discovering Symmetries in Neural Network Parameter Spaces

发现神经网络参数空间中的对称性
Zhao, Bo, Dehmamy, Nima, Walters, Robin, Yu, Rose
Abstract
Parameter space symmetries are important for understanding neural networks' loss landscape, training dynamics, and generalization. However, systematically identifying these symmetries remains a challenge. In this paper, we formalize data-dependent parameter symmetries and characterize loss invariance and the group-action axioms through infinitesimal conditions, which provide objectives for jointly learning group generators and nonlinear action maps. Our framework systematically uncovers parameter symmetries, including previously unknown ones. To study larger networks, we establish conditions under which subnetwork symmetries extend to the full model. The same construction gives an explicit family of finite-batch symmetries, providing both analytical examples and a foundation for discovery through small subnetworks. Using the infinitesimal characterization and subnetwork construction, we implement a framework for automated discovery of parameter symmetries, and successfully uncovered symmetries in various architectures, including pretrained transformer models.
Chinese Translation
参数空间对称性对于理解神经网络的损失景观、训练动力学和泛化非常重要。然而,系统地识别这些对称性仍然是一个挑战。在本文中,我们将数据相关的参数对称性形式化,并通过无穷小条件刻画损失不变性和群作用公理,这些条件为联合学习群生成元和非线性作用映射提供了目标。我们的框架系统地揭示了参数对称性,包括此前未知的对称性。为了研究更大的网络,我们建立了子网络对称性扩展到完整模型的条件。相同的构造给出了一族显式的有限批次对称性,既提供了分析示例,也为通过小型子网络进行发现奠定了基础。利用无穷小刻画和子网络构造,我们实现了一个用于自动发现参数对称性的框架,并成功地在各种架构中发现了对称性,包括预训练的 Transformer 模型。
cs.LG / 299 / 2609.33536

Does Execution Require Target KV Fidelity? A Mixed-Fidelity KV Runtime for LLM Serving

执行是否需要目标 KV 保真度?面向 LLM 服务的混合保真度 KV 运行时
Jiang, Jiantong, Yang, Yue, Yang, Peiyu, Liu, Feng
Abstract
Large language model (LLM) serving is increasingly constrained by the GPU memory consumed by key-value (KV) caches. Existing compression, eviction, and offloading techniques alleviate this pressure, but serving runtimes typically treat only the configured target KV representation as execution-ready. Under memory pressure, this target-only contract can turn KV shortage into request stalls and preemptions. We present ElasticKV, a mixed-fidelity KV runtime built on the observation that target fidelity need not gate execution. ElasticKV introduces a compact intermediate KV state, making fidelity a runtime-managed execution property. To realize this state in a paged serving runtime, ElasticKV combines (i) a pair-structured layout that turns fidelity reduction into reusable GPU capacity, (ii) a dual-mode attention backend that directly consumes the compact state while preserving the native target-only path, and (iii) pressure-aware fidelity management that adapts KV fidelity to memory pressure. Our extensive evaluation across diverse workloads, model families and scales, and GPU platforms demonstrates the effectiveness and generality of ElasticKV. Under high concurrency, ElasticKV achieves 3.8-4.0$\times$ lower time-to-first-token (TTFT) and 9.1$\times$ lower P90 TTFT than vLLM while preserving generation quality.
Chinese Translation
大语言模型(LLM)服务日益受到键值(KV)缓存所消耗的 GPU 内存的制约。现有的压缩、驱逐和卸载技术缓解了这种压力,但服务运行时通常仅将配置的目标 KV 表示视为可执行就绪状态。在内存压力下,这种仅目标(target-only)的契约可能将 KV 短缺转变为请求停滞和抢占。我们提出 ElasticKV,一种混合保真度 KV 运行时,其基于以下观察:目标保真度不必门控执行。ElasticKV 引入一种紧凑的中间 KV 状态,使保真度成为由运行时管理的执行属性。为了在分页服务运行时中实现这一状态,ElasticKV 结合了:(i)一种成对结构布局,将保真度降低转化为可复用的 GPU 容量;(ii)一种双模式注意力后端,可直接消费该紧凑状态,同时保留原生的仅目标路径;以及(iii)压力感知的保真度管理,可根据内存压力调整 KV 保真度。我们在多种工作负载、模型系列与规模以及 GPU 平台上的广泛评估证明了 ElasticKV 的有效性和通用性。在高并发下,与 vLLM 相比,ElasticKV 在保持生成质量的同时,实现了 3.8-4.0 倍更低的首次令牌时间(TTFT)和 9.1 倍更低的 P90 TTFT。
cs.LG / 300 / 2609.33548

TerMeZO: Ternary Sparse Zeroth-Order Optimization for Fine-tuning BitNet Models at the Edge

TerMeZO:面向边缘端 BitNet 模型微调的三元稀疏零阶优化
Sifaou, Houssem, Katti, Prabodh, Rajendran, Bipin, Simeone, Osvaldo
Abstract
Fine-tuning anguage models (LLMs) with first-order optimizers requires a memory several times larger than that required for inference. Memory-efficient zeroth-order optimization (MeZO) sidesteps this cost by estimating gradients from forward passes only. However, for BitNet architectures, a family of LLMs with ternary {-1,0,1\} weights and 8-bit activations, fine-tuning requires updating full-precision latent weights, and thus the memory footprint of MeZO no longer matches that of inference. A promising solution is to finetune only a subset of the latent weights, but existing sparse zeroth-order (ZO) methods either ignore the ternary structure or require first-order gradient information to build a sparse mask, which is at odds with the purpose of ZO fine-tuning. We propose TerMeZO, a sparse MeZO scheme that exploits the geometry of the ternary quantizer itself to identify the latent weights that are more likely to change values during fine-tuning, at no additional data or memory cost. Our convergence analysis shows that TerMeZO can converge faster than full-parameter MeZO, owing to its optimized reduction of the fine-tuning effective dimension. We run extensive experiments on BitNet models ranging from 1B to 3B parameters, spanning classification, instruction-following, and mathematical reasoning tasks. TerMeZO matches or exceeds the performance of full-parameter MeZO while substantially reducing the fine-tuning memory footprint.
Chinese Translation
使用一阶优化器微调语言模型(LLMs)所需的内存是推理所需内存的数倍。内存高效的零阶优化(MeZO)仅通过前向传播估计梯度,从而规避了这一开销。然而,对于 BitNet 架构——一类具有三元 {-1,0,1} 权重和 8 位激活的 LLMs——微调需要更新全精度的潜在权重,因此 MeZO 的内存占用不再与推理时相匹配。一个有前景的解决方案是仅微调潜在权重的一个子集,但现有的稀疏零阶(ZO)方法要么忽略三元结构,要么需要一阶梯度信息来构建稀疏掩码,这与 ZO 微调的目的相悖。我们提出 TerMeZO,一种稀疏 MeZO 方案,它利用三元量化器本身的几何特性来识别在微调过程中更可能发生值变化的潜在权重,且无需额外的数据或内存开销。我们的收敛性分析表明,由于优化地降低了微调的有效维度,TerMeZO 可以比全参数 MeZO 收敛得更快。我们在参数规模从 1B 到 3B 的 BitNet 模型上进行了大量实验,涵盖分类、指令遵循和数学推理任务。TerMeZO 在显著降低微调内存占用的同时,性能与全参数 MeZO 相当或更优。